WO2020163388A1 - Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems - Google Patents

Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems Download PDF

Info

Publication number
WO2020163388A1
WO2020163388A1 PCT/US2020/016654 US2020016654W WO2020163388A1 WO 2020163388 A1 WO2020163388 A1 WO 2020163388A1 US 2020016654 W US2020016654 W US 2020016654W WO 2020163388 A1 WO2020163388 A1 WO 2020163388A1
Authority
WO
WIPO (PCT)
Prior art keywords
sensing
genetic
seq
molecular component
sensitive
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2020/016654
Other languages
French (fr)
Inventor
Dan Mcfarland Park
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Lawrence Livermore National Security LLC
Original Assignee
Lawrence Livermore National Security LLC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Lawrence Livermore National Security LLC filed Critical Lawrence Livermore National Security LLC
Publication of WO2020163388A1 publication Critical patent/WO2020163388A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12YENZYMES
    • C12Y207/00Transferases transferring phosphorus-containing groups (2.7)
    • C12Y207/13Protein-histidine kinases (2.7.13)
    • C12Y207/13003Histidine kinase (2.7.13.3)
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6897Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids involving reporter genes operably linked to promoters
    • CCHEMISTRY; METALLURGY
    • C07ORGANIC CHEMISTRY
    • C07KPEPTIDES
    • C07K14/00Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof
    • C07K14/195Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof from bacteria
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N1/00Microorganisms; Compositions thereof; Processes of propagating, maintaining or preserving microorganisms or compositions thereof; Processes of preparing or isolating a composition containing a microorganism; Culture media therefor
    • C12N1/20Bacteria; Culture media therefor
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/11DNA or RNA fragments; Modified forms thereof; Non-coding nucleic acids having a biological activity
    • C12N15/52Genes encoding for enzymes or proenzymes
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/635Externally inducible repressor mediated regulation of gene expression, e.g. tetR inducible by tetracyline
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/70Vectors or expression systems specially adapted for E. coli
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/74Vectors or expression systems specially adapted for prokaryotic hosts other than E. coli, e.g. Lactobacillus, Micromonospora
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/10Transferases (2.)
    • C12N9/12Transferases (2.) transferring phosphorus containing groups, e.g. kinases (2.7)
    • GPHYSICS
    • G21NUCLEAR PHYSICS; NUCLEAR ENGINEERING
    • G21FPROTECTION AGAINST X-RADIATION, GAMMA RADIATION, CORPUSCULAR RADIATION OR PARTICLE BOMBARDMENT; TREATING RADIOACTIVELY CONTAMINATED MATERIAL; DECONTAMINATION ARRANGEMENTS THEREFOR
    • G21F9/00Treating radioactively contaminated material; Decontamination arrangements therefor
    • G21F9/04Treating liquids
    • G21F9/06Processing
    • G21F9/18Processing by biological processes
    • GPHYSICS
    • G21NUCLEAR PHYSICS; NUCLEAR ENGINEERING
    • G21FPROTECTION AGAINST X-RADIATION, GAMMA RADIATION, CORPUSCULAR RADIATION OR PARTICLE BOMBARDMENT; TREATING RADIOACTIVELY CONTAMINATED MATERIAL; DECONTAMINATION ARRANGEMENTS THEREFOR
    • G21F9/00Treating radioactively contaminated material; Decontamination arrangements therefor
    • G21F9/28Treating solids
    • G21F9/30Processing
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2830/00Vector systems having a special element relevant for transcription
    • C12N2830/001Vector systems having a special element relevant for transcription controllable enhancer/promoter combination
    • C12N2830/002Vector systems having a special element relevant for transcription controllable enhancer/promoter combination inducible enhancer/promoter combination, e.g. hypoxia, iron, transcription factor
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2830/00Vector systems having a special element relevant for transcription
    • C12N2830/55Vector systems having a special element relevant for transcription from bacteria
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2840/00Vectors comprising a special translation-regulating system
    • C12N2840/002Vectors comprising a special translation-regulating system controllable or inducible
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2840/00Vectors comprising a special translation-regulating system
    • C12N2840/10Vectors comprising a special translation-regulating system regulates levels of translation
    • C12N2840/105Vectors comprising a special translation-regulating system regulates levels of translation enhancing translation
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N2520/00Use of whole organisms as detectors of pollution

Definitions

  • the present disclosure relates to uranium (U) biosensors and related U-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems.
  • U biosensors and related methods and systems to detect and/or neutralize bioavailable uranium and more particularly bioavailable uranyl oxycation.
  • U biosensors and related U-sensing genetic molecular components, gene cassettes, genetic circuits, compositions, methods and systems which in several embodiments can be used to detect and/or neutralize uranium and in particular bioavailable UF 6 , or its stable hydrolysis product UO2F2 , which is typically produced in U enrichment operations.
  • a U0 2 F 2 -biosensor comprising a U-sensing genetic molecular component and an F-sensing riboswitch configured to report and/or neutralize uranium in presence of bioavailable Uranium and Fluoride, in a single output or dual output configuration.
  • the U0 2 F 2 -biosensor comprises a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 herein also UrpS, and U sensitive transcriptional regulator 1362 herein also UrpR.
  • the genetically modified bacterial cell is an engineered bacterial cell comprising a 1362 U-sensing reportable genetic molecular component and/or a 1362 U- sensing/U-neutralizing genetic molecular component, each comprising a U-sensitive promoter in a configuration wherein the U-sensitive promoter directly initiates expression of the 1362 U- sensing reportable molecular component and/or of the 1362U-neutralizing molecular component in presence of bioavailable U.
  • the 1362 U-sensitive promoter comprises a 1362 (UrpR) binding site having a DNA sequence
  • Ni is C or T, preferably C
  • N 2 is G or A, preferably G;
  • N 3 is T or C, preferably T;
  • N 4 is C
  • Ns is A or G, preferably A;
  • N ⁇ is G or C, preferably G;
  • Ns is any nucleotide
  • N 9 is any nucleotide
  • Nio is any nucleotide
  • N 11 is any nucleotide
  • N12 is T or C
  • Ni 4 is T or C, preferably T;
  • Nis is C
  • N 16 IS A or C, preferably A;
  • Nn G
  • Nis is C or G, and wherein Ni to Nn are selected independently.
  • the genetically modified bacterial cell is an engineered bacterial cell further comprising an F-sensing riboswitch within at least one of
  • an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F.
  • the F-sensing riboswitch and the 1362 U-sensing reportable genetic molecular component are in a single output configuration.
  • the F-sensing riboswitch and the 1362 U-sensing reportable genetic molecular component are in a dual output configuration.
  • the genetically modified bacteria are bacteria incapable of natively expressing the histidine kinase 1363 (UrpS) and the U-sensitive transcriptional regulator 1362 (UrpR) (e.g. E. Coli bacteria).
  • the bacterial cell is further engineered to include a 1362 U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363(UrpS), and a gene encoding response regulator 1362 are expressed upon activation of a controllable promoter.
  • the genetically modified bacteria are bacteria capable of natively expressing histidine kinase 1363 (UrpS), and U-sensitive transcriptional regulator 1362 (UrpR) (e.g. proteobacteria such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure).
  • UrpS histidine kinase 1363
  • UrpR U-sensitive transcriptional regulator
  • the endogenous genes encoding the histidine kinase 1363, and the U- sensitive transcriptional regulator 1362 (UrpR), are knocked out and the genetically engineered bacterial cell is further engineered to include a 1362 U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363 (UrpS), and a gene encoding response regulator 1362 (UrpR) are expressed upon activation of a controllable promoter.
  • a 1362 U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362
  • the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure).
  • the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure).
  • the endogenous genes encoding the F-sensing riboswitch are preferably knocked out.
  • a UC Fi-biosensor comprising a U-sensitive F-sensitive genetic circuit wherein a U-sensing genetic molecular component and an F-sensing riboswitch are configured to report and/or neutralize uranium in presence of bioavailable Uranium and Fluoride, in a single output or dual output configurations.
  • the UO2F2 biosensor comprises a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR).
  • the genetically modified bacterial cell is an engineered bacterial cell comprising a U-sensitive F-sensitive genetic circuit in which molecular components are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components.
  • At least one molecular component is a 1362 U- sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U-sensitive 1362 (UrpR) binding site having a DNA sequence
  • Ni is C or T, preferably C
  • N2 is G or A, preferably G;
  • N 3 is T or C, preferably T;
  • N 4 is C
  • Ns is A or G, preferably A;
  • N is G or C, preferably G; N 7 C or G;
  • Ns is any nucleotide
  • N9 is any nucleotide
  • Nio is any nucleotide
  • N11 is any nucleotide
  • N12 is T or C
  • Ni4 is T or C, preferably T;
  • Nis is C
  • N16IS A or C preferably A
  • Nn G
  • Nis is C or G
  • Ni to Nn are selected independently.
  • At least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), and/or a U-neutralizing molecular component, (in particular one or more U-neutralizing genetic molecular component and/or one or more U-neutralizing cellular molecular component).
  • the 1362 U-sensing F-sensing genetic circuit further comprises an F sensing riboswitch within at least one genetic molecular component of the molecular components of the U-sensitive F-sensitive genetic circuit, in a configuration wherein the at least one genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
  • the at least one reportable molecular component and/or the U- neutralizing molecular component are expressed when the genetic circuit operates according to the circuit design in presence of an effective amount of bioavailable U and an effective amount of bioavailable F in a single output or dual output configurations.
  • the genetically modified bacteria are bacteria not capable of natively expressing the histidine kinase 1363 (UrpS) and the U-sensitive transcriptional regulator 1362 (UrpR) (e.g. E. Coli bacteria).
  • the bacterial cell is further engineered to include a 1362 U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363 (UrpR), and a gene encoding response regulator 1362(UrpR) are expressed upon activation of a controllable promoter.
  • the genetically modified bacteria are bacteria capable of natively expressing histidine kinase 1363(UrpS), and U-sensitive transcriptional regulator 1362(UrpR) (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure).
  • UrpS histidine kinase 1363
  • UrpR U-sensitive transcriptional regulator
  • the endogenous genes encoding the histidine kinase 1363(UrpS), and the U- sensitive transcriptional regulator 1362(UrpR), can be preferably knocked out and the genetically engineered bacterial cell is further engineered to include a 1362 U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 in a configuration wherein the a gene encoding histidine kinase 1363(UrpS), and a gene encoding response regulator 1362 (UrpR) are expressed upon activation of a controllable promoter.
  • a 1362 U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS), and an endogenous or exogenous gene encoding U-
  • the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure).
  • the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure).
  • the endogenous genes encoding the F-sensing riboswitch are preferably knocked out.
  • a UC Fi-biosensor comprising a U-sensing genetic molecular component and an F-sensing riboswitch configured to report and/or neutralize uranium in presence of bioavailable Fluoride, in a single output or double output configuration.
  • the UC Fibiosensor comprises a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase UzcS and U-sensitive transcriptional response regulator UzcR.
  • the genetically modified bacterial cell is an engineered bacterial cell comprising a UzcR U-sensing reportable genetic molecular component and/or a UzcR U-sensing U-neutralizing genetic molecular component, each comprising a U-sensitive promoter in a configuration wherein the U-sensitive promoter directly initiates expression of the UzcR U- sensing reportable molecular component and/or of the UzcR U-sensing U-neutralizing molecular component in presence of bioavailable U.
  • the U-sensitive promoter comprises an UzcR binding site having a DNA sequence:
  • N 7 -N 12 is independently any nucleotide, and in some embodiments any one of N 7 -N 11 can independently be A.
  • the genetically modified bacterial cell is an engineered bacterial cell further comprising an F-sensing riboswitch within at least one of
  • an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F.
  • the F-sensing riboswitch and the UzcR U-sensing reportable genetic molecular component are in a single output configuration.
  • the F-sensing riboswitch and the UzcR U-sensing reportable genetic molecular component are in a dual output configuration.
  • the genetically modified bacteria are bacteria not capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. E. Coli bacteria), and the bacterial cell is further engineered to include a UzcR U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding the histidine kinase UzcS, and an endogenous or exogenous gene encoding the U- sensitive transcriptional response regulator UzcR, in a configuration wherein the gene encoding the histidine kinase UzcS, and the gene encoding the U-sensitive transcriptional response regulator UzcR, are expressed upon activation of a controllable promoter.
  • a UzcR U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding the histidine kinase UzcS, and an endogenous or exogenous gene encoding the U- sensitive
  • the genetically modified bacteria are bacteria capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure).
  • the U-sensitive transcriptional response regulator UzcR e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure.
  • the endogenous genes encoding the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR can be preferably knocked out and the genetically engineered bacterial cell can be further engineered to include a UzcR U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
  • the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure).
  • the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure).
  • the endogenous genes encoding the F-sensing riboswitch are preferably knocked out.
  • a UC Fi-biosensor comprising a U-sensitive F- sensitive genetic circuit wherein a UzcR U-sensing genetic molecular component and an F- sensing riboswitch are configured to report and/or neutralize uranium in presence of bioavailable Uranium and Fluoride, in a single output or dual output configurations.
  • the U biosensor comprises a genetically modified bacterial cell natively and/or heterologously expressing histidine kinase UzcS, and response regulator UzcR.
  • the genetically modified bacterial cell is an engineered bacterial cell comprising a U-sensitive F-sensitive genetic circuit in which molecular components are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components.
  • the UzcR U-sensitive F-sensitive genetic circuit at least one molecular component is a UzcR U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a UzcR binding site having DNA sequence:
  • N 7 -N 12 is independently any nucleotide, and in some embodiments any one of N 7 -N 11 can independently be A.
  • At least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), and/or a U-neutralizing molecular component, (in particular one or more U-neutralizing genetic molecular component and/or one or more U-neutralizing cellular molecular component),
  • the genetically modified bacterial cell further comprises an F-sensing riboswitch within at least one genetic molecular component of the molecular components of the U-sensitive F-sensitive genetic circuit in a configuration wherein the at least one genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
  • the at least one reportable molecular component and/or the a U- neutralizing molecular component are expressed when the genetic circuit operates according to the circuit design in presence of an effective amount of bioavailable U and an effective amount of bioavailable F in a single output or dual output configurations.
  • the genetically modified bacteria are bacteria not capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. E. Coli bacteria).
  • the bacterial cell is further engineered to include a UzcR U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding the histidine kinase UzcS, and an endogenous or exogenous gene encoding the U-sensitive transcriptional response regulator UzcR, in a configuration wherein the gene encoding the histidine kinase UzcS, and the gene encoding the U-sensitive transcriptional response regulator UzcR, are expressed upon activation of a controllable promoter.
  • the genetically modified bacteria are capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure).
  • the endogenous genes encoding the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR are knocked out.
  • the genetically engineered bacterial cell is further engineered to include a UzcR U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
  • a UzcR U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS
  • U-sensitive transcriptional response regulator UzcR U-sensitive transcriptional response regulator
  • the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure).
  • the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure).
  • the endogenous genes encoding the F-sensing riboswitch are preferably knocked out.
  • a UC Fi-biosensor comprising a U-sensing genetic circuit and an F-sensing riboswitch configured to report and/or neutralize uranium in presence of bioavailable Uranium and Fluoride in a dual output configuration.
  • the UO2F2- biosensor comprises a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 herein also UrpS, and U sensitive transcriptional regulator 1362 herein also UrpR and/or natively and/or heterologously expressing histidine kinase UzcS, and response regulator UzcS.
  • the genetically modified bacterial cell is an engineered bacterial cell comprising a U-sensitive genetic circuit in which molecular components are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components.
  • At least one molecular component is a U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U.
  • a U-sensitive promoter when the cell is capable of natively and/or heterologously expressing histidine kinase 1363 herein also UrpS, and U sensitive transcriptional regulator 1362 herein also UrpR, at least one molecular component is a 1362 U-sensing genetic molecular component in which the U sensitive promoter comprises a U-sensitive 1362 (UrpR) binding site having a DNA sequence
  • Ni is C or T, preferably C
  • N2 is G or A, preferably G;
  • N 3 is T or C, preferably T;
  • N 4 is C
  • Ns is A or G, preferably A;
  • N is G or C, preferably G; N 7 C or G;
  • Ns is any nucleotide
  • N9 is any nucleotide
  • Nio is any nucleotide
  • N11 is any nucleotide
  • N12 is T or C
  • Ni4 is T or C, preferably T;
  • Nis is C
  • N16IS A or C preferably A
  • Nn G
  • Nis is C or G
  • At least one molecular component is a UzcR U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a UzcR binding site having DNA sequence:
  • N7-N12 is independently any nucleotide, and in some embodiments any one of N7-N11 can independently be A.
  • At least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), and/or a U-neutralizing molecular component, (in particular one or more U-neutralizing genetic molecular component and/or one or more U- neutralizing cellular molecular component), the reportable molecular component and/or the a U- neutralizing molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable U.
  • the genetically modified bacterial cell is an engineered bacterial cell further comprising an F-sensing riboswitch within at least one of
  • a genetic molecular component of an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F; and
  • an F-sensing reportable genetic molecular component in a configuration wherein the F- sensing reportable genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
  • the F-sensing reportable genetic molecular component and the F-sensing genetic circuit are in a dual output configuration with the U-sensing genetic circuit.
  • the genetically modified bacteria are bacteria incapable of natively expressing the histidine kinase 1363 (UrpS) and the U-sensitive transcriptional regulator 1362 (UrpR) (e.g. E. Coli bacteria).
  • the bacterial cell is further engineered to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363(UrpS), and a gene encoding response regulator 1362 are expressed upon activation of a controllable promoter.
  • the genetically modified bacteria are bacteria capable of natively expressing histidine kinase 1363 (UrpS), and U-sensitive transcriptional regulator 1362 (UrpR) (e.g. proteobacteria such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure).
  • UrpS histidine kinase 1363
  • UrpR U-sensitive transcriptional regulator
  • the endogenous genes encoding the histidine kinase 1363, and the U- sensitive transcriptional regulator 1362 (UrpR), are knocked out and the genetically engineered bacterial cell is further engineered to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363 (UrpS), and a gene encoding response regulator 1362 (UrpR) are expressed upon activation of a controllable promoter.
  • a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362
  • the genetically modified bacteria are bacteria not capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. E. Coli bacteria), and the bacterial cell is further engineered to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding the histidine kinase UzcS, and an endogenous or exogenous gene encoding the U-sensitive transcriptional response regulator UzcR, in a configuration wherein the gene encoding the histidine kinase UzcS, and the gene encoding the U-sensitive transcriptional response regulator UzcR, are expressed upon activation of a controllable promoter.
  • a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding the histidine kinase UzcS, and an endogenous or exogenous gene encoding the U-sensitive transcriptional response regulator
  • the genetically modified bacteria are bacteria capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure).
  • the U-sensitive transcriptional response regulator UzcR e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure.
  • the endogenous genes encoding the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR can be preferably knocked out and the genetically engineered bacterial cell can be further engineered to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
  • the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure).
  • the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure).
  • the endogenous genes encoding , encoding a fluoride efflux pump are preferably knocked out.
  • the UC Fi-biosensor comprises a U- sensing genetic molecular component in which a U-sensitive promoter comprising a UzcR binding site
  • the genetically modified bacterial cell is a bacterial cell capable of natively expressing MarR family repressors such as marRi (CCNA_03498) and marR2 (CCNA_02298) genes (e.g.
  • the bacterial cell is preferably further engineered to knock out at least one endogenous MarR family repressors such as marRi and marR genes to provide an amplified UC Fi-biosensor configured to provide an amplified signal following activation of the UO2F2- sensitive genetic circuit.
  • the U0 2 F 2 -biosensor comprises a U-sensing genetic molecular component in which a U-sensitive promoter comprises a UzcR binding site
  • the UO2F2- biosensor or the U-sensitive F sensitive genetic circuit further comprises an amplifier genetic molecular component comprising a U-sensitive promoter and UzcY and/or UzcZ in a configuration wherein the U-sensitive promoter directly initiates expression of the amplifier molecular component.
  • a method to provide a UC Fi-biosensor comprising
  • a bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and U sensitive response regulator 1362 (UrpR) and/or U sensitive response regulator UzcR in combination with a heterologous F-sensing riboswitch, the genetically engineering performed by introducing into the cell
  • U-sensing genetic molecular components configured to report and/or neutralize U herein described
  • UC Fi-biosensor optionally operatively connecting the UC Fi-biosensor so provided to an electronic signal transducer adapted to convert a UOiFibioscnsor reportable molecular component output into an electronic output.
  • the method further comprises genetically engineering the cell to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362(UrpR) and/or response regulator UzcR in a configuration wherein the a gene encoding histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and a gene encoding response regulator 1362 (UrpR) response regulator UzcR are expressed upon activation of a controll
  • the bacterial cell is a cell capable of natively expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR (e.g.
  • the genetically engineering can preferably further comprises knocking out the natively expressed histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR of the bacterial cell, and introducing in the bacterial cell a U-sensing regulator component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) and/or response regulator UzcR in a configuration wherein the a gene encoding histidine kinase 1363 and/or histidine kinase Uz
  • the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure).
  • the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure).
  • the endogenous genes encoding the F-sensing riboswitch are preferably knocked out.
  • the UO2F2- sensing gene cassette comprises one or more U-sensing genetic molecular components herein described, one or more U-sensing regulator genetic molecular components herein described and/or one or more reportable genetic component herein described.
  • the U0 2 F 2 -sensing gene cassette further comprises an F-sensitive riboswitch within the one or more U-sensing genetic molecular components configured to report U, within the one or more U-sensing regulator genetic molecular components and/or within an additional reportable genetic molecular component, in a configuration wherein the U-sensing genetic molecular components, the one or more U-sensing regulator genetic molecular components, and the additional reportable genetic molecular component are transcribed in presence of an effective amount of bioavailable fluoride.
  • the U-sensing gene cassette is an expression cassette.
  • the gene cassette is comprised within a vector.
  • a vector comprising a polynucleotide encoding for one or more U-sensing genetic molecular components herein described, one or more F- sensing genetic molecular components herein described, one or more U-sensing and/or F sensing regulator genetic molecular components herein described and/or one or more genetic molecular components of a UC Fi-biosensor herein described.
  • the one or more vectors are configured to introduce one or more U-sensitive genetic molecular components, one or more F-sensing genetic molecular components and/or one or more genetic molecular components of a U-sensitive and/or F-sensitive genetic circuit into a bacterial cell of a plurality of bacterial cells.
  • a UC Fi-sensing system comprises one or more vectors herein described and/or a plurality of bacterial cells natively and/or heterologously expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR in combination with one or more F-sensitive riboswitches.
  • the bacterial cell is further genetically engineered to include a U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS) and/or histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362(UrpR) and/or response regulator UzcR in a configuration wherein the a gene encoding histidine kinase 1363(UrpS) and/or histidine kinase UzcS, and a gene encoding response regulator 1362 (UrpR) response regulator UzcR are expressed upon activation of a controllable
  • the bacterial cell is a cell capable of natively expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR (e.g.
  • the bacterial cells is preferably further genetically engineered to comprises knocking out the natively expressed histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR of the bacterial cell, and introducing in the bacterial cell a U-sensing regulator component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362(UrpR) and/or response regulator UzcR in a configuration wherein the a gene encoding histidine kinase 1363 and/or histidine kinase UzcS, and a gene
  • the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure).
  • the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure).
  • the endogenous genes encoding the F-sensing riboswitch are preferably knocked out.
  • a UC Fi-sensing system comprises one or more of the UO2F2 biosensors herein described operatively connected to an electronic signal transducer adapted to convert a U biosensor reportable molecular component output into an electronic output.
  • composition comprises one or more UO2F2 biosensors, U0 2 F 2 -sensing gene cassettes and/or vectors herein described together with a suitable vehicle.
  • a system comprising an electronic signal transducer adapted to convert a UO2F2 biosensor reportable molecular component output into an electronic output.
  • the system comprises an electronic signal transducer and one or more UO2F2 biosensors herein described operatively connected to the electronic signal transducer.
  • a method of detecting, reporting and/or neutralizing bioavailable UO 2 F 2 is described. The method comprises:
  • UO 2 F 2 biosensors herein described contacting one or more UO 2 F 2 biosensors herein described, or a system comprising an electronic transducer operatively connected to one or more UO 2 F 2 biosensors herein described, with a target environment comprising one or more target ranges of U concentration in combination with one or more target F concentration for a time and under conditions to detect, report and/or neutralize bioavailable UO 2 F 2 in the target environment.
  • one or more U0 2 F 2 -sensing genetic reportable components are also described, the UO 2 F 2 sensing genetic reportable components comprising a U sensitive promoter comprising a U-sensitive 1362 (UrpR) binding site and/or a U sensitive promoter comprising a U sensitive UzcR binding site together with
  • the U-sensitive promoter directly initiates expression of the U- sensing reportable molecular component and/or of the U- sensing U-neutralizing molecular component in presence of bioavailable U and the U- sensing reportable molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
  • one or more U-sensitive and/or F sensitive genetic circuits are described wherein at least one molecular component is a U sensing genetic molecular component herein described and wherein at least one molecular component is a reportable molecular component, and/or a U-neutralizing molecular component (and in particular possibly one or more U-neutralizing genetic molecular components),
  • the U-sensitive and/or F sensitive genetic circuits further comprises an F-sensing riboswitch within at least one of the genetic molecular components of the genetic circuit in a configuration wherein the at least one of genetic molecular components of the genetic circuit is transcribed in presence of an effective amount of bioavailable fluoride.
  • the reportable molecular component and/or the U-neutralizing molecular component are expressed when the genetic circuit operates according to the circuit design in presence of bioavailable U and/or bioavailable Fluoride.
  • the biosensors and related genetic molecular components, genetic circuits, compositions, methods and systems herein described are configured for selective and sensitive detection, reporting and/or neutralizing UO 2 F 2 a bioavailable environmental decomposition product of UF 6 and toxic form of U typically produced during enrichment operations.
  • UO 2 F 2 biosensors and related genetic molecular components, genetic circuits, compositions, methods and systems herein described provide in several embodiments a selective, sensitive, portable, easy to use, high-throughput measurement and or neutralizing bioavailable UO 2 F 2 , with little or no sample preparation required.
  • UO 2 F 2 biosensors and related genetic molecular components, genetic circuits, compositions, methods and systems herein described allow in several embodiments construction of consolidated bioremediators comprising bacterial systems that possess all the necessary components for deployment in environmental cleanup efforts, for example by coupling UO 2 F 2 sensing with activation of one or more U-neutralizing components.
  • UO 2 F 2 biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described allow in several embodiments detection, reporting and/or neutralization of U in its bioavailable form UO 2 F 2 with low cost approaches as various proteobacterial cells, such as Caulobacter, can be inexpensively grown to high densities as will be understood by a skilled person.
  • the UO 2 F 2 biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described can be used in connection with various applications wherein detection and/or neutralizing of uranium is desired.
  • the U0 2 F 2 biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described can be used in biodefense and in particular to be used for non-proliferation purposes, in environmental monitoring and/or cleanup by regulatory agencies or communities, and in mining in particular for toxicology and safety concerns, as well as diagnostic applications.
  • Additional exemplary applications include uses of the UO 2 F 2 biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described in several fields including basic biology research, applied biology, bio- engineering, medical diagnostics, and in additional fields identifiable by a skilled person upon reading of the present disclosure.
  • Figure 1 shows a graph of pairwise comparison of fold-change in expression in Caulobacter crescentus of exemplary genes activated by Zn or U, analyzed using RNA-Seq.
  • CCNA_01362 and CCNA_01353 (phytase) (dark grey data points shown as respectively labeled circles) are induced more than 100-fold by U but not induced by Zn.
  • the promoters regulating these genes, Pi 36i and P 1353 (Pphyt) respectively represent promising exemplary promoters for use in a whole-cell U biosensor.
  • urcA dark grey data point indicated as labeled circle
  • a gene regulated by UzcRS Figure 24A-C
  • Figure 23A is strongly induced by both U ( Figure 24A) and Zn
  • Figure 2 shows graphs reporting determination of metal specificity of native U- responsive promoters in Caulobacter crescentus.
  • the metal specificity of chromosomal P m cA- lacZ (Panel A), P phyt -lacZ (Panel B), and P nei-lacZ (Panel C) was determined by treating mid exponential phase cells with a range of concentrations of various metal salts for two hours before determining b-galactosidase activity using the method of Miller [1].
  • Cell growth was performed in peptone yeast extract (PYE) media supplemented with 50 mM MES pH 6.1.
  • Figure 3A shows schematics of two exemplary independent two component systems (TCS) that are sensitive to U in Caulobacter crescentus. Shown on the left side of Figure 3A is a schematic of the putative mechanism of action of the UzcRS TCS, which was characterized as a transcriptional activator of at least 40 operons in response to U/Zn/Cu [3].
  • FIG. 3A Shown on the right side of Figure 3A is a schematic of the putative mechanism of action of the 1363/1362 TCS, wherein a histidine kinase 1363 (encoded by CCNA_01363) and a response regulator 1362 (encoded by CCNA_01362) activate promoters comprising the 1362 regulator direct repeat DNA binding site, e.g. P phyt and Pi36i, in response to U. 1363 is membrane-bound and senses U either through a direct (U binding to 1363) or indirect mechanism. Upon U sensing, it is expected that 1363 autophosphorylates and then transphosphorylates 1362, activating 1362 for DNA binding of the 1362 direct repeat, e.g. in P phyt or Pi36i. Possible additional stimuli for 1363/1362 remain unknown (indicated by the circled question mark).
  • a histidine kinase 1363 encoded by CCNA_01363
  • a response regulator 1362 encoded by
  • Figure 3B shows a schematic of an exemplary U-sensitive AND gate that incorporates two independent points of uranyl sensing inputs (the two-component systems 1363/1362 AND UzcRS), which are both required to affect an output (such as a reportable molecular component and/or a U-neutralizing molecular component).
  • the 1363/1362 two-component system is specifically activated by uranyl.
  • the UzcRS two component system is activated for transcriptional regulation by uranyl, as well as zinc, copper and cadmium [3].
  • Panels A-C shows schematics illustrating the stepwise genetic engineering of an exemplary U-sensitive genetic circuit with incremental improvements from Panel A to Panel C to enhance specificity for U, resulting in a genetic circuit comprising an‘in series’ AND gate comprising two points of U-sensing by (1) P phyt or Pi36i and (2) UzcRS two component system.
  • Panel A shows a schematic of a U-sensitive genetic circuit comprising uzcRS under the control of the native P UZ cR promoters PI and P2 and GFP expression under the control of UzcR-regulated promoter P1968.
  • Panel B shows a schematic of a U-sensitive genetic circuit where Pi and Pn are replaced with P phyt or Pi36i such that uzcRS expression is now dependent on activation by these U- specific promoters.
  • This construct requires two points of U sensing for reporter activation, (1) activation of uzcRS transcription by P phyt or P i ; 1 ⁇ 2 i and (2) stimulation of UzcRS transcriptional regulatory activity.
  • This sensor shows greater signal in response to U compared to the genetic circuit shown in Panel A, as shown in Figure 5 Panel A.
  • Figure 4 Panel B shows a schematic of a U- sensitive genetic circuit where a negative feedback loop was incorporated into the circuit, whereby UzcR represses its own expression from P phyt or Pi36i. Specifically, an m_5 UzcR binding site was placed downstream of the P pi , yL or Pi36i transcription start site. This genetic circuit shows minimized basal expression of uzcRS, while maintaining strong responsiveness to U, and further shifted ratio of U response to that of Zn to 5.5 as shown in Figure 5 Panel C.
  • Figure 5 shows graphs of exemplary GFP reporter fluorescence produced by the U- sensitive genetic circuits shown in Figure 4, comprised in the host organism C. crescentus NA1000, upon exposure to U, Zn or Cu.
  • Figure 5 Panel A shows graphs reporting quantification of GFP fluorescence produced by C. crescentus NA1000 comprising the U-sensitive genetic circuit shown in Figure 4 Panel A in response to exposure to 10 and 20 mM U (Figure 5 Panel A left graph), 10, 20 and 40 pM Zn ( Figure 5 Panel A middle graph), or 80, 120 and 200 pM Cu ( Figure 5 Panel A right graph), from 0 to 4 hours after exposure.
  • Figure 5 Panel B shows graphs reporting quantification of GFP fluorescence produced by C.
  • crescentus NA1000 comprising the U-sensitive genetic circuit shown in Figure 4 Panel B in response to exposure to 10 and 20 pM U (Figure 5 Panel B left graph), 10, 20 and 40 pM Zn (Figure 5 Panel B middle graph), or 80, 120 and 200 pM Cu (Figure 5 Panel B right graph), from 0 to 4 hours after exposure.
  • Figure 5 Panel C shows graphs reporting quantification of GFP fluorescence produced by C. crescentus NA1000 comprising the U-sensitive genetic circuit shown in Figure 4 Panel C in response to exposure to 10 and 20 mM U (Figure 5 Panel C left graph), 10, 20 and 40 mM Zn (Figure 5 Panel C middle graph), or 80, 120 and 200 pM Cu (Figure 5 Panel C right graph), from 0 to 4 hours after exposure.
  • P phyt was used for the GFP fluorescence measurements shown in Figure 5 Panels B-C.
  • U were performed in modified M5G medium (10 mM PIPES, pH 7, I mM NaCl, I mM KC1, 0.05 % NH 4 C1, 0.01 mM Fe/EDTA, 0.2% glucose, 0.5 mM MgS0 4 , 0.5 mM CaCh) supplemented with 5 mM glycerol-2- phosphate as the phosphate source (M5G-G2P).
  • Cells were grown to mid-exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated U concentration.
  • Zn and Cu were performed in PYE.
  • Figure 6A and Figure 6B show schematics of two exemplary U-sensitive genetic circuits with‘in parallel’ AND gates comprising two points of U-sensing by (1) 1363/1362 two- component system (exemplified by U-sensitive transcriptional regulator direct repeat-containing promoters P phyt or Pi36i) and (2) UzcRS two component system (exemplified by UzcRS- responsive promoter P UTC B).
  • Figure 6A shows a schematic of an exemplary‘in parallel’ AND gate, comprised of an HRP AND gate system, in which hrpS is placed under the control of P phyt or Pi 36i and hrpR is placed under the control of PurcB, a promoter activated by UzcRS.
  • FIG. 6B shows a schematic of an exemplary‘in parallel’ AND gate, comprised of a tripartite GFP system, in which gfplO subunit is placed under the control of P phyt or Pi36i, gfpll is placed under the control of PurcB, and gfpl-9 is placed under the control of the Caulobacter S layer promoter, P rS aA, which is a strong, constitutive promoter [4].
  • Gfp-10 is shown fused to K1 and gfpl l is shown fused to El, wherein K1 and El are exemplary interacting protein partners comprised of oppositely charged coiled-coils [5].
  • K1 and El interact, GFP10 and GFP11 self-associate with GFP1- 9 to constitute a functional GFP reporter.
  • Figure 6B shows a more detailed version of Figure 6B, depicting how the two independent U sensing systems are integrated into the AND gate.
  • Figure 7 shows schematics of regulatory sequences within P p h y t ( Figure 7 Panel A), showing the sequence
  • Figure 8 shows DNA sequences of full-length Pp hyt (Panel A), Pp hyt with a mutation of four nucleotides in the second direct repeat sequence (DR2, shown in bold, Panel B), P p h y t with a mutation of two nucleotides in the first direct repeat sequence (DR1, shown in bold, Panel C), Pphyt with a mutation of two nucleotides in the second direct repeat sequence (DR2, shown in bold, Panel D), a shortened Pphyt (Pphyt short, Panel E), full-length Pi36i (Panel F), Pi36i with a mutation of four nucleotides in the second direct repeat sequence (DR2, shown in bold, Panel G), P 1361 with a mutation of two nucleotides in the second direct repeat sequence (DR2, shown in bold, Panel H), and a shortened Pi36i (Pi36i short, Panel I).
  • TR large tandem repeat
  • the direct repeat sequence that is likely bound by the U-sensitive transcriptional response regulator 1362 is shown in uppercase, italic (with direct repeat sequences underlined).
  • Putative Transcription start sites are shown in lowercase, underlined.
  • Figure 9 shows graphs reporting quantification of exemplary fluorescence levels of P p h y t- gfp (Figure 9 Panel A) and Pi36i-gfp (Figure 9 Panel B) variants in response to U at 10 mM or 25 pM in the host organism C. crescentus NA1000.
  • Panel A fluorescence levels are shown for variants comprising full length P phyt promoter (Full length), a shortened P phyt (Short), Pphyt with a mutation of four nucleotides in the second direct repeat sequence (DR2 GTCA -> CAGT), P phyt with a mutation of two nucleotides in the first direct repeat sequence (DR1, GT -> CA), and P phyt with a mutation of two nucleotides in the second direct repeat sequence (DR2, GT -> CA).
  • Panel B fluorescence levels are shown for variants comprising full length Pi36i promoter (Full length), a shortened Pi36i (Short), Pi36i with a mutation of four nucleotides in the second direct repeat sequence (DR2 GTCA -> CAGT), and Pi36i with a mutation of two nucleotides in the second direct repeat sequence (DR2, GT -> CA).
  • Cells were grown to mid exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated U concentration. Fluorescence was quantified following a two- hour exposure of cells to each U concentration and normalized to the ODeoo .
  • FIG 10 illustrates the results of an electrophoretic mobility assay (EMSA) showing UrpR binding to wild type and mutant P p h yt fragments.
  • the mutant P p h yt DNA (P1361) contains a GTCA -> CAGT mutation of DR2 (see Figure 7B).
  • the assays were performed with 50 nM 6-FAM-labeled DNA and UrpR, phosphorylated with carbamoyl phosphate. The concentrations indicate the total UrpR used in the assay. A representative example of three biological replicates is depicted.
  • Figure 11 shows graphs reporting exemplary data corresponding to the exemplary U- sensitive genetic circuit in Figure 6 Panel B.
  • gfpl-9 is controlled by P AV/ and gfpl-9 expression is induced with 10 mM xylose.
  • Fluorescence output for both P phyt and Pi36i sensor variants is plotted as a function of time following metal exposure and was normalized to the fluorescence of a strain lacking the UzcR regulator. Cells were grown to mid-exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated metal concentration.
  • Figure 12 shows a graph reporting data showing activity of an exemplary native signal amplifier module for UzcRS.
  • the graph shows increase in fluorescence at various U concentrations in a CCNA_03499 mutant (+ UzcY) or in wild type cells (- UzcY), relative to conditions in absence of metal. Fluorescence was normalized to cell density (ODeoo).
  • Figure 13 shows a schematic of an exemplary signal amplifier module incorporated within a U sensing circuit. The signal amplifier uzcY is placed under the control of the U- specific promoters R,, / , ,/R i % i such that signal amplification is restricted to conditions of U exposure.
  • FIG 14 shows an exemplary embodiment of a U-responsive AND gate design.
  • Panel A shows a schematic of an example of an‘in series’ AND gate combined with an‘in parallel’ AND gate in a U sensitive genetic circuit.
  • the P mcB promoter in the exemplary‘in parallel’ tripartite GFP AND gate is activated by UzcR, and expression of UzcR and UzcS is under transcriptional regulation of P phyt in an‘in series’ AND gate.
  • Grey arrows depict regulator modifications made to the base‘in parallel’ AND gate to generate the combined‘in series’,‘in parallel’ AND gate circuit.
  • gfplO subunit is placed under the control of P phyt -short
  • gfpll is placed under the control of P U rcB
  • gfpl-9 is placed under the control of the Caulobacter S layer promoter, P rsa A, which is a strong, constitutive promoter.
  • P rsa A which is a strong, constitutive promoter.
  • UC, UDC and UTC represent the aqueous complexes UO2CO3 0 , U02(C03)2 2- and U02(C03)3 4- ⁇
  • the shaded area represents the range of conditions of common natural waters [9] as presented in Newsome et al., (2014) [7]
  • Figure 16A is from Figure 2 of Newsome et al. (2014) [7] showing a schematic illustrating exemplary mechanisms of microbe-uranium interactions, such as bioreduction [10-
  • Figure 16B is from Figure 2 of Newsome et al. (2014) [7] showing a schematic illustrating exemplary mechanisms of microbe-uranium interactions, such as biomineralisation [14-16]
  • Figure 16C is from Figure 2 of Newsome et al. (2014) [7] showing a schematic illustrating exemplary mechanisms of microbe-uranium interactions, such as biosorption [17, 18].
  • Figure 16D is from Figure 2 of Newsome et al. (2014) [7] showing a schematic illustrating exemplary mechanisms of microbe-uranium interactions, such as bioaccumulation [19] as presented in Newsome et al., (2014) [7].
  • Figure 17 shows graphs reporting quantification of exemplary fluorescence levels of a shortened P Phyt -gfp (Figure 17 Panel A) and shortened Pi 36i -gfp (Figure 17 Panel B) variants in response to U at 5 mM, 10 mM or 20 pM.
  • Figure 17 Panels A and B fluorescence levels are shown for wild type C. crescentus and strains deleted for CCNA_01362 (1362 response regulator) and CCNA_01363 (1363 histidine kinase).
  • Figure 18 shows schematics illustrating exemplary tripartite GFP U-sensitive genetic circuits together with graphs reporting quantification of exemplary GFP reporter fluorescence produced by the respective genetic circuits under the conditions indicated.
  • Figure 18 Panel A shows a schematic of the U-sensitive genetic circuit shown in Figure 6 Panel B
  • Figure 18 Panel B shows a schematic of a control circuit that incorporates input from only the UzcRS TCS, comprised in the host organism C. crescentus NA1000, upon exposure to U, Zn, Cu or Cd.
  • Figure 17 Panel A shows a graph reporting exemplary quantification of GFP fluorescence produced by C. crescentus NA1000 comprising the U-sensitive genetic circuit shown in Figure 6 Panel B upon exposure to U, Zn, Cu and Cd.
  • FIG. 17 Panel B shows a graph reporting exemplary quantification of GFP fluorescence produced by C. crescentus NA1000 comprising a control U-sensitive genetic circuit that incorporates input only from UzcRS.
  • Figure 17 demonstrates the enhanced selectivity of an exemplary‘in parallel’ AND gate comprising two points of U-sensing by (1) 1363/1362 and (2) UzcRS two component systems.
  • Figure 19 shows a graph reporting exemplary data indicating the limit of U detection for the U biosensor described in Figure 4 Panel B.
  • Mid-exponential phase cells were washed twice in 10 mM Pipes pH 7 and then resuspended in 10 mM Pipes pH 7 containing uranyl nitrate. As shown in the graph, a linear response was observed for U concentrations in the low micromolar range.
  • Figure 20 shows a graph reporting exemplary data showing the ratio of the fluorescence output in response to 10 and 20 mM of U and Zn for the“In-parallel” AND-gate shown in Figure 6 Panel C and the combined“In-series” plus“in-parallel” genetic circuit shown in Figure 14.
  • Cells were grown to mid-exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated metal concentration. Fluorescence was quantified following a three-hour exposure of cells to each metal concentration and normalized to the OD600 . The U/Zn ratio was calculated by dividing the normalized fluorescence with U by that with Zn.
  • Figure 21 shows a schematic showing exemplary U-sensitive genetic circuits having exemplary U-neutralizing outputs to allow U bioprecipitation or bioadsorption.
  • Figure 22 shows a schematic showing an exemplary AND gate comprised in a U- sensitive genetic circuit wherein an alpha-gall IP fusion and a lambda repressor-gal4 fusion are expected to be driven by a combination of P phyt fP i36i and any UzcRS regulated promoter.
  • Figure 23 shows in some embodiments the determination of metal specificity of P urcA ⁇ Panel (A): Time course of chromosomal P mcA-lacZ induction after treatment of early exponential phase cells with or without Zn. Cells were grown in PYE and the b-galactosidase activity at each time point is depicted. Error bars represent the standard deviation of biological triplicates.
  • Figure 24 shows an exemplary direct activation of P mcA by UzcRS according to some embodiments herein described.
  • Panel (A) Time course of P mcA-lacZ expression following treatment with uranyl nitrate in M5G media supplemented with 5 mM glycerol-2-phosphate as the phosphate source media for wild type (WT), D uzcR and D uzcS.
  • UzcR-P was phosphorylated with carbamoyl phosphate and the concentrations indicate the total UzcR used in the assay.
  • FIG. 25 shows an exemplary identification of the UzcR sequence recognition motif in some embodiments herein described.
  • Panel (A) The 18-bp UzcR (m_5) sequence logo was constructed from the alignment of 49 UzcR boxes identified within the sequence regions bound by UzcR in vivo using MEME [21]. The sequence conservation (bits) is depicted by the height of the letters with the relative frequency of each base depicted by its relative height.
  • TSS Transcription Start Site
  • Promoter -gfpmut3 fusions with the wild type and mutant Pi968 fragments described in Panel (B) were constructed and fluorescence was quantified following a two-hour treatment with or without 20 mM Zn using a Biotek plate reader (ex: 480/ em: 516).
  • the fluorescence signal was normalized to the OD600 and fold activation was calculated by dividing the normalized fluorescence in the presence of Zn by the fluorescence in the uninduced condition.
  • Error bars represent that standard deviation calculated using a formula for propagation of standard error [20] This figure is taken from the Appendix B of U.S. provisional application No. 62/587,753 the disclosure of which is incorporated herein by reference in its entirety.
  • Figure 26 shows exemplary MarR regulators repressing expression of the membrane proteins UzcY and UzcZ according to some embodiments herein described.
  • the Fluorescence of promoter-g p fusions was quantified at mid-exponential phase and normalized to the OD600. Error bars represent the standard deviation of biological triplicates.
  • Figure 27 shows exemplary effect of UzcRS regulator mutants in minimal media and U induction and TFInfer activity of UzcR in negative regulator mutants according to some embodiments herein described.
  • Figure 28 shows diagrams illustrating the results of experiments showing that transposon insertions within CCNA_01362 and CCNA_01363 abrogate U-dependent induction of Pphyt.
  • Pphyt -lacZ activity was assayed in strains containing transposons inserted within CCNA_01362 and CCNA_01363 using b-galactosidase assays. Error bars represent the standard deviation of triplicate measurements.
  • Figure 29 shows diagrams illustrating the results of experiments showing metal selectivity of UrpRS and UzcRS in an exemplary embodiment.
  • Panel A Functional independence of UrpRS and UzcRS.
  • U response curves of Pphyt-short-g ⁇ ? and P urcB-gfp reporters were determined by quantifying fluorescence following a two-hour exposure to a range of U concentrations in wild type C. crescentus and strains deleted for uzcS, urpS and the cognate response regulator. Error bars represent the average of biological triplicates.
  • Figure 30 shows diagrams illustrating the results of experiments showing in an exemplary embodiment functional independence of UrpRS and UzcRS. Fluorescence of a Pphyt- short -gfp and PurcB-gfp were quantified following a two-hour exposure to a range of U concentrations in wild type C. crescentus and strains deleted for uzcR and urpR. Error bars represent the average of biological triplicates.
  • Figure 31 illustrates a schematic of U-sensing AND gate configuration that was integrated at the chromosomal urcA locus. Details of AND gate construction and chromosomal integration are outlined in the methods sections.
  • Figure 32 shows diagrams illustrating the results of experiments showing in an exemplary embodiment U sensing AND gate variants with different mechanisms of gfpl-9 expression.
  • Panel A shows the fluorescence output profile of the core sensor and a sensor variant with gfpl-9 expression driven by P p h y t following a three-hour exposure to the indicated U concentration.
  • Panel B shows the fluorescence output of a U-sensor variant with gfpl-9 expression driven by the xylose-inducible promoter P xyi as a function of time .
  • Figure 33 shows a graph reporting that core sensor is not responsive to nitrate.
  • HNO3 100 mM UO2NO3 in 100 mM HNO3
  • Relevant concentrations of HNO3 failed to induce fluorescence, confirming that nitrate alone is not responsible for the core sensor output signal.
  • Figure 34 shows in a graph the U response curves for the core sensor and control variants in which the expression of gfplO-Kl and El-gfpl l components is driven by UzcRS or UrpRS alone. Fluorescence was quantified following a three-hour exposure to a range of U concentrations. The data were fit with a Hill equation to determine the U concentration that yields half-maximal fluorescence induction and the hill slope as an indicator of U sensitivity. Data points with diminished fluoresescence output were not included in the curve fitting.
  • Figure 35 shows diagrams illustrating the results of experiments showing metal selectivity profile of a core sensor. Fluorescence was quantified following a two-hour exposure to each metal. Error bars represent the average of biological triplicates.
  • Figure 36 shows graphs reporting incorporation of a signal amplifier module within the core sensor circuitry.
  • Panel A schematic for the integration of the UzcY signal amplifier within the U-sensing AND gate. In this variant, uzcY is placed under the control of the U-specific promoter P P h y t-short such that the signal amplification is restricted to conditions of U exposure.
  • Panel B The fluorescence output of the core sensor and signal amplifier variants following a three-hour exposure to a range of U concentrations. The data were fit with a Hill equation to determine the U concentration that yields half-maximal fluorescence induction and the hill slope as an indicator of U sensitivity. Signal amplifier data points with diminished fluoresescence output were not included in the curve fitting.
  • Panel C The fluorescence output of the core sensor and signal amplifier variants following a three-hour exposure to a range of Zn, Pb, and Cu concentrations. Error bars represent the average of biological triplicates.
  • Figure 37 shows a zoomed in version of the U/Zn/Cu/Pb response curves for the core sensor, control variants, and uzcY amplifier variants. Fluorescence was quantified following a three-hour exposure to a range of metal concentrations. The data represent a zoomed in version of data depicted in Figure 40 and Figure 36C. Error bars represent the average of biological triplicates.
  • Figure 38 shows graphs reporting fluorescence output of sensor variants in ground water samples without nutrient supplementation.
  • Panel A Fluorescence output of the core sensor and control variants lacking critical regulatory components as a function of time in three distinct site 300 samples without nutrient supplementation.
  • Panel B Fluorescence output of the core sensor and a control strain lacking a UrpR binding site as a function of time in three distinct site 300 samples supplemented with glucose (0.2%), orthophosphate (1 mM), and ammonium chloride (0.05%).
  • the plot legend is depicted below the plots.
  • FIG 39 shows detection of U in ground water samples in an exemplary embodiment.
  • Panel A Fluorescence output of the core sensor and control variants lacking critical regulatory components as a function of time in three distinct site 300 samples supplemented with glucose (0.2%), glycerol-2-phosphate (5 mM), and ammonium chloride (0.05%). Ground water samples from W-815-2621, W-812-01, and W-6C are referred to as 1, 2, and 3, respectively, in the text. The plot legend is depicted below the plots.
  • Panel B The fluorescence of the core sensor and control variants was quantified six hours after exposure to site 300 samples supplemented with uranyl nitrate and the nutrients described in panel A. The x-axis concentrations represent the total U concentration in the ground water sample after addition of U in 1 mM increments.
  • Figure 40 shows a metal selectivity profile of core sensor and control variants.
  • the top panel depicts simplified schematics of the core sensor and control variants in which the expression of gfplO-Kl and El-gfpl l components is driven by UzcRS or UrpRS alone.
  • the bottom panel depicts the U/Zn/Cu/Pb response curves for the core sensor and control variants. Fluorescence was quantified following a three-hour exposure to a range of metal concentrations. Error bars represent the average of biological triplicates.
  • Figure 41 shows a schematic illustration of an exemplary transducer device configured to convert an optical signal from a biosensor herein described into an electrical current using inexpensive, commercially available components (e.g., a blue LED for excitation, optical filters configured for excitation/emission of GFP, and a photodiode for detection).
  • inexpensive, commercially available components e.g., a blue LED for excitation, optical filters configured for excitation/emission of GFP, and a photodiode for detection.
  • Figure 42 shows the output voltage as a function of U concentration at three different currents using the exemplary transducer device depicted in Figure 41. Measurements were taken two hours after uranium exposure.
  • Figure 43 shows the conserved nucleotides in naturally occurring Fluoride sensing riboswitchs from a gapped alignment of the 2138 fluoride riboswitch sequences from the Rfam database
  • Figure 44 shows a schematic illustration of a design of a bacteria-based sensor of uranyl fluoride products.
  • Figure 45 shows a schematic illustration of a design to integrate and optimize a fluoride- detection capability in C. crescentus to provide a UCLFibioscnsor herein described.
  • Figure 46 shows a schematic representation of exemplary single output configurations of a UO2F2 biosensor herein described.
  • Panel A describes a schematic of uranium- and fluoride- sensing components integrated in series.
  • the promoter can be any UzcR- or UrpR-regulated promoter (defined in prior patent app) such that U-dependent transcription is initiated by UzcR or UrpR.
  • transcription will be prematurely terminated by the fluoride sensing riboswitch in the absence of fluoride (i.e., not detectable output). Binding of fluoride to the riboswitch will mediate transcriptional read-through and ultimately, production of the reporter.
  • Figure 46 Panel B describes a schematic of uranium- and fluoride-sensing components integrated in series where U-dependent transcriptional activation requires the function of both UrpR and UzcR. This sensor will provide greater selectivity for uranium compared to the sensors described in panel A. However, in this configuration, the sensitivity for U would be limited by the UzcRS component.
  • Figure 46 Panel C describes a schematic of the integration of the fluoride riboswitch within the AND gate circuit such that expression of component three requires fluoride exposure. In this configuration, reconstitution of GFP fluorescence requires activation of the two uranium-responsive pathways and fluoride binding to the fluoride riboswitch. Transcription of component three can theoretically be controlled by any constitutive promoter. The assignment of each component with the given regulatory promoter is arbitrary and easily swapped. For example, the fluoride riboswitch could be used to control expression of component one or two.
  • Figure 47 shows a schematic representation of an exemplary dual output configurations of a UO 2 F 2 biosensor herein described.
  • Figure 48A is Figure 1A of US 9, 580,713 and, as indicated in US 9, 580,713 shows a schematic representation of “the consensus sequence and structural model based on the comparison of 2188 representatives from bacterial and archaeal species”. As indicated in US 9, 580,713 incorporated herein by reference in its entirety“PI, P2, P3 and pseudoknot labels identify base-paired substructures. Note that the bottom of PI carries a possible G-C pair. However, because noncomplementary nucleotides occur in these positions in some representatives, these nucleotides are depicted as unpaired.”
  • Figure 48B is Figure IB of US 9, 580,713 and, as indicated in US 9, 580,713 shows the ”[s]equence and secondary structure model for the WT 78 Psy RNA (SEQ ID NO:349). Numbers 1, 2, 3, 4, 5, and 6depict the sites of the in-line probing analysis; results presented in C. The two G residues preceding nucleotide 1 5 were added to facilitate RNA production by in vitro transcription.”
  • Figure 49 shows charts reporting the effect of crcB deletion on the growth of Caulobacter crescentus in the presence of NaF and NaCl.
  • Figure 49 panel A reports results showing growth in Caulobacter crescentus WT at NaF concentrations above 4 mM and
  • Figure 49 panel B reports results showing growth in Caulobacter crescentus AcrcB at NaF concentrations above 126.5 mM,.
  • Figure 49 panel C reports results showing growth in Caulobacter crescentus WT at NaCl concentrations above 4 mM.
  • Figure 49 panel D reports results showing growth in Caulobacter crescentus AcrcB at NaCl concentrations above 4 mM.
  • Figure 50 shows a schematic and charts illustrating the result of testing the function of fluoride riboswitches from three different bacteria in Caulobacter crescentus.
  • Figure 50 Panel A shows a schematic configuration of the exemplary construct used to perform the testing using b- galactosidase as a reporter.
  • Figure 50 Panel B shows a chart reporting the b-galactosidase activity for the Sphingomonas MM-1 in WT and the crcB deletion strain over a range of NaF concentrations
  • Figure 50 Panel C shows a chart reporting the b-galactosidase activity for the Pseudomonas syringae riboswitch in WT and the crcB deletion strain over a range of NaF concentrations
  • Figure 50 Panel D shows a chart reporting the b-galactosidase activity for the Sphingomonas 67-36 riboswitch in WT and the crcB deletion strain over a range of NaF concentrations.
  • Figure 51 shows a schematic and charts illustrating fluoride responsiveness of fluoride riboswitches from three different bacteria tested in C. crescentus using the xylose inducible promoter.
  • Figure 51 Panel A shows a schematic representation of a construct wherein each of three tested riboswitches was included in construction between to the xylose inducible promoter (Pxyl) and the reporter mCherry.
  • Figure 51 Panel B shows a chart reporting the effect of xylose administration ad different concentration on promoter activity.
  • Figure 51 Panel C shows a chart reporting the effect of 200 uM xylose administration on mCherry expression in the constructs including each of the three riboswitches as indicated.
  • Figure 51 Panel D shows a chart reporting the effect of 2 mM xylose administration on mCherry expression in the constructs including each of the three riboswitches as indicated.
  • Figure 51 Panel E shows a chart reporting the effect of 200 uM xylose administration on mCherry expression in the constructs including the Sphingomonas MM-1 fluoride riboswitch in C. Crescentus WT and C. Crescentus with a crcB deletion.
  • Figure 51 Panel F shows a chart reporting the effect of 2 mM xylose administration on mCherry expression in the constructs including the Sphingomonas MM-1 fluoride riboswitch in C Crescentus WT and C. Crescentus with a crcB deletion.
  • Figure 52 shows a schematic representation of the primary genetic parts involved in the construction of a fluoride responsive reporter.
  • Figure 53 shows an alignment of RNA encoded by 287 sequences (SEQ ID NO: 1989- 1998 and SEQ ID NO: 2232 to SEQ ID NO: 2508) from the 2017 exemplary F-sensing riboswitch sequences shown in Appendix I(SEQ ID NO: 205-1988 and SEQ ID NO: 1999- 2231) wherein the nucleotide sequences shown are sequences forming a same structure in corresponding aligned sequences and the symbols“—” indicates gaps between the aligned sequences shown.
  • Figure 54 show the effect of translational fusion length on Sphingomoms MM- 1 fluoride riboswitch function
  • Figure 55 shows charts reporting the testing of an in -series uranyi fluoride sensing circuit.
  • Figure 55 Panel A shows fluoride detection performed with a control, MM-1 fluoride sensing circuit with native crcB promoter not responsive to U.
  • Figure 55 Panel B shows fluoride detection performed with a uranyi fluoride sensing circuit constructed by combining the UrpRS- responsive Pphyt promoter with the MM- 1 riboswitch.
  • Figure 55 Panel € shows the fluorescence of the uranyi fluoride sensing circuit in the presence of U alone, F alone, and both U and F. Tests-were performed with F-, added as NaF, and uranyi, added as uranyi nitrate.
  • sequences from Appendix I and Sequence Listing from SEQ ID NO: 205 to SEQ ID NO: 1988 and SEQ ID No: 1999 to SEQ ID NO: 2231 are sequences from Rfam database reported with their related Genome file ID, bacteria and additional information concerning the position of the sequence in the genome of the bacteria as will be understood by a skilled person Appendix I together with the detailed description section, the Example section and the Drawings, serve to explain the principles and implementations of the disclosure. Other features, objects, and advantages will be apparent from the entire description and drawings, and from the claims.
  • UO2F2 biosensors and related U-sensing and/or F sensing genetic molecular components, genetic circuits, compositions, methods and systems which in several embodiments can be used to detect, report and/or neutralize U and in particular bioavailable Uranium which is produced in connection with enrichment programs UO2F2.
  • bioavailable refers to a molecule in particular a soluble molecule that is able to cross an organism's cellular membrane from the environment, or is otherwise able to exert a biological effect on an organism, if the organism has access to the molecule.
  • a bioavailable toxic molecule is a toxic molecule that is able to exert toxicity on an organism contacted with the organism and/or with a toxic molecule sensing system of the organism.
  • the bioavailability can be inferred based on toxicity or activation of a tress response in an organism as will be understood by a skilled person.
  • bioavailable U refers to a soluble molecular form of U that can cross an organism's cellular membrane from the surrounding environment or is otherwise able to exert a biological effect on an organism, e.g. following contact with the organism and/or with an organism U-sensing system.
  • the term“bioavailable uranium” comprises uranyl ion, which has a linear structure with short U-0 bonds, indicative of the presence of multiple bonds between uranium and oxygen and can bind four or more ligands in an equatorial plane.
  • the uranyl ion forms many complexes, particularly with ligands that have oxygen donor atoms. Complexes of the uranyl ion are important in the extraction of uranium from its ores and in nuclear fuel reprocessing.
  • ‘naked’ or‘uncomplexed’ uranyl oxycation is a bioavailable form of U.
  • uranyl oxycation complexed with inorganic phosphate is not considered to be bioavailable.
  • uranyl oxycation refers to the predominant form of U in oxygenated environments, comprising the +6 oxidation state (UO + 2 2 .), which has high chemical toxicity [24].
  • the US Environmental Protection Agency’s maximum contaminant limit for U in drinking water is 30 pg/L (-0.13 mM), however, groundwater concentrations in the US frequently exceed this limit [25, 26].
  • UO2F2 biosensors and related U-sensing and/or F-sensing genetic molecular component, gene cassettes, genetic circuits, compositions, methods and systems described herein can be used in several embodiments to detect and report and/or neutralize bioavailable U, and in particular UO2F2 which is a derivative of UF 6 .
  • Uranium hexafluoride is a compound used in the process of enriching uranium, which produces fuel for nuclear reactors and nuclear weapons.
  • Uranium hexafluoride (UF 6 ) is almost always produced as a precursor in any U-enrichment operation, is routinely released during the conversion process, and is not expected to occur naturally in the environment [27, 28]
  • environmental detection of UF 6— or the more stable hydrolysis product UO2F2 which is rapidly formed when atmospheric UF 6 reacts with water vapor [29] . does strongly suggest an enrichment program.
  • UO2F2 aerosol and HF gas
  • UO2F2 is expected to have reasonable stability in the environment such that its detection may be feasible for months after release [27, 31].
  • UO2F2 aerosols which are rapidly formed when atmospheric UF 6 reacts with water vapor [29], are expected to be deposited on vegetation, soil, or into aquatic systems [29, 30]. While UO2F2 aerosols are expected to have reasonable stability in low moisture environments [27, 31], UO2F2 exhibits relatively high solubility in aqueous environments (up to 2 M [32]);
  • a solution-based detection approach is therefore expected to be most relevant for UO2F2 monitoring in aquatic systems or when coupled with a sampling device that collects and solubilizes UO2F2 aerosols. Additionally, the rationale for separate uranium and fluoride detection components is supported by the solution chemistry of UO2F2 and prior studies on bacterial uranium interactions.
  • the physicochemical form, or speciation, of UO2F2 is dependent on the geochemical conditions [33, 34].
  • the naked uranyl oxycation (U0 2 2+ ) and uranyl fluoride species (U0 2 F + , UO2F2, UO2F3 ) are expected to predominate [33].
  • uranyl hydroxide and/or carbonate species are expected to predominate with concomitant formation of the anion, F [33, 34].
  • uranium and fluoride are expected to largely exist as separate species under geochemical conditions most relevant to aqueous environmental sampling (pH 6-8 range).
  • Insoluble forms of uranyl for example uranyl phosphate minerals formed when uranyl nitrate is added to solutions containing high orthophosphate levels— are not detected by the biosensor. This is an advantageous feature for environmental detection since natural U commonly occurs in the form of insoluble U minerals [37] and aqueous phosphate concentrations are typically very low ( ⁇ 10 ppb) [34].
  • U0 2 F 2 biosensor herein described are bacteria-based U0 2 F 2 -sensor, which are engineered to comprise a fluoride- sensing riboswitch in combination a U-sensing genetic molecular components, U neutralizing genetic molecular component, reportable genetic molecular components, and/or additional components possibly configured in genetic circuits directed to detect and/or neutralize U in presence of bioavailable Uranium and Fluoride.
  • the UO2F2 biosensors herein described are whole-cell biosensors comprising a genetically engineered bacterial cell.
  • bacteria used herein interchangeably with the terms“cell” or “host” indicates a large domain of prokaryotic microorganisms.
  • prokaryotic is used herein interchangeably with the terms“cell” or“host” and refers to a microbial species which contains no nucleus or other organelles in the cell.
  • Exemplary prokaryotic cells include bacteria. Typically, a few micrometers in length, bacteria have a number of shapes, ranging from spheres to rods and spirals, and are present in several habitats, such as soil, water, acidic hot springs, radioactive waste, the deep portions of Earth's crust, as well as in symbiotic and parasitic relationships with plants and animals.
  • Bacteria in the sense of the disclosure refers to several prokaryotic microbial species which comprise Gram-positive bacteria, Proteobacteria, Cyanobacteria, Spirochetes and related species, Planctomyces, Bacteroides, Flavobacteria, Chlamydia, Green sulfur bacteria, Green non-sulfur bacteria including anaerobic phototrophs, Radioresistant micrococci and related species, Thermotoga and Thermosipho thermophiles as would be understood by a skilled person.
  • Gram positive bacteria refers to cocci, nonspomlating rods and spomlating rods, such as, for example, Actinomyces, Bacillus, Clostridium, Corynebacterium, Erysipelothrix, Lactobacillus, Listeria, Mycobacterium, Myxococcus, Nocardia, Staphylococcus, Streptococcus and Streptomyces.
  • proteobacteria refers to a major phylum of Gram-negative bacteria. Many move about using flagella, but some are nonmotile or rely on bacterial gliding. As understood by skilled persons, taxonomic classification as proteobacteria is determined primarily in terms of ribosomal RNA (rRNA) sequences. The Proteobacteria are divided into six classes, referred to by the Greek letters alpha through epsilon and the Acidithiobacillia and Oligoflexia, including alphaproteobacteria, betaproteobacteria and gammaproteobacteria as will be understood by a skilled person.
  • rRNA ribosomal RNA
  • alphaproteobacteria refers to bacteria identifiable by those skilled in the art in the phylogenetic Class Alphaproteobacteri, in the Phylum Proteobacteria.
  • Alphaproteobacteria is a diverse taxon and comprises several phototrophic genera, several genera metabolising Cl-compounds (e.g., Methylobacterium spp.), symbionts of plants (e.g., Rhizobium spp.), endosymbionts of arthropods ( Wolbachia ) and intracellular pathogens (e.g. Rickettsia).
  • taxonomic classification of alphaproteobacteria can be identified by reference to publicly available online databases such as the List of Prokaryotic names with Standing in Nomenclature (LPSN) and National Center for Biotechnology Information (NCBI) and the phylogeny is based on 16S rRNA-based LTP release 106 by 'The All-Species Living Tree' Project.
  • the Class Alphaproteobacteria is divided into three subclasses Magnetococcidae, Rickettsidae and Caulobacteridae [38].
  • the Caulobacteridae is a subclass composed of the orders Holosporales, Rhodospirillales, Sphingomonadales, Rhodobacterales, Caulobacterales, Rhizobhiales, Kiloniellales, Kordiimonadales, Parvularculales and Sneathiellales.
  • Betaproteobacteria refers to a class of gram-negative bacteria, and one of the classes of the phylum Proteobacteria.
  • the Betaproteobacteria comprise more than 75 genera and 220 species of bacteria identifiable by persons skilled in the art. [40] Seven orders of betaproteobacteria have been described: Burkholderiales, Hydrogenophilales, Methylophilales, Neisseriales, Nitrosomonadales, Rhodocyclales, and Sulfuricellales. Examples of Betaproteobacteria genera comprise Bordetella, Ralstonia, Neisseria and Nitrosomonas, among others identifiable by skilled persons.
  • Betaproteobacteria While many Betaproteobacteria identifiable by skilled persons are found in environmental soil and water, others are obligate pathogens and can cause disease in a variety of hosts. Some members of betaproteobacteria can cause disease in various eukaryotic organisms. Several cause diseases in humans, such as members of the genus Neisseria: N gonorrhoeae and N. meninngitides which cause gonorrhea and meningitis respectively, as well as Bordetella pertussis which causes whooping cough. Other members infect plants, such as Burkholderia cepacia which causes bulb rot in onions as well as Xylophilus ampelinus which causes necrosis of grapevines. [40]
  • gammaproteobacteria refers to a class of gram-negative bacteria, and one of the classes of the phylum Proteobacteria.
  • exemplary taxonomic orders, families and genera belonging to the class gammaproteobacteria comprise Acidithiobacillus, Xanthomonadales, Chromatiales, Methylococcus, Beggiatoa, Legionellales, Ruthia, Vesicomyosocius, Thiomicrospira, Dichelobacter, Francisella, Moraxellaceae, Alcalinovorax, Saccharophagus, Reinekea, Oceanospirillaceae, Marinobacter, Pseudomonadaceae, Aeromonas, Vibrionales, Pasteurellales, and Enterobacteriales among others.
  • a number of bacteria have been described as members of gammaproteobacteria, but have not yet been assigned an order or family. These comprise bacteria of the genera Alkalimarinus, Alkalimonas, Arenicella, Gallaecimonas, Ignatzschineria, Litorivivens, Marinicella, Methylohalomonas, Methylonatrum, Plasticicumulans, Pseudohongiella, Sedimenticola, Thiohalobacter, Thiohalomonas, Thiohalorhabdus, Thiolapillus, and Wohlfahrtiimonas among others identifiable by skilled persons.
  • gammaproteobacteria genera comprise Escherichia, Shigella, Salmonella, Yersinia, Buchnera, Haemophilus, Vibrio, and Pseudomonas, among others identifiable by skilled persons.
  • Some members of gammaproteobacterial are pathogenic in humans, for example some strains of the species Salmonella spp., Yersinia pestis, Vibrio cholerae, Pseudomonas aeruginosa, and Escherichia coli, among others identifiable by skilled persons.
  • Some members of gammaproteobacteria are pathogenic in plants, such as Xanthomonas axonopodis pv. citri, Pseudomonas syringae pv. actinidiae, and Xylella fastidiosa, among others identifiable by skilled persons.
  • the UOiFibioscnsors herein described are whole-cell biosensors comprising a genetically engineered alphaproteobacterial cell of the subclass Caulobacteridae.
  • the U biosensors described herein can comprise a cell of any genus, species and/or strain of Caulobacteridae identifiable by those skilled in the art.
  • Exemplary Caulobacteridae that can be used in U-biosensors herein described comprise species of the Families Bradyrhizobiaceae, Sphingomonadaceae , Caulobacteraceae, Hyphomicrobiaceae and Rhodobacteraceae which include species naturally comprising, as well as others identifiable by persons skilled in the art.
  • UC Fi-biosensor s herein described can comprise species from the order Caulobacterales, the family Caulobacteraceae , the genus Caulobacter and the species Caulobacter crescentus which is described herein as one of the representative species of the subclass Caulobacteridae.
  • the bacterial cell of the U-biosensor is capable of natively and/or heterologously expressing a U-sensitive histidine kinase 1363, and cognate response regulator 1362.
  • the term“histidine kinase P1363” or“UrpS” as used herein refers to a histidine kinase having the amino acid sequence MS GGS LRWRLIVGGMLAILA AL A V A WL AMT WLFERHI VRRET ADLTRAGQ VL V AGLR LEPN G AP VID ATLS DPRLS KA AGGFY W Q V STTSGSERSVS LWDQ ALKPPQT AP AEGW S S RIA AGPFDDR VLL VERS VRPDRDGP A VLIQ V AS DEKVLRA ARREF GRELAIFLGGLW AIL S G A A ALQ V VLGLS PLTRVR ADLARLRKS PS ARMS LDHPREIAPLAE AIN AL AE ARE ADL ARARRR AGDLAHS LKTPL A ALS AQS RRAREDG A V A ADGLD A AIAS V A A ALE AEL AR ARAAAAREAVFAAETAPLAVAERLVAVLERTAD
  • crescentus NA1000 or a sequence that when aligned with sequence SEQ ID NO: 3 has a BLAST score between 240 and 300, between 300 and 500, or preferably between 500 and 800, or more preferably over 800 but less than 100% homology or even more preferably having a BLAST Score of 851 and 100% homology with the sequence SEQ ID NO: 3.
  • U sensitive response regulator 1362 or “response regulator 1362” or “UrpR” refers to a response regulator having amino acid sequence
  • crescentus NA1000 or a sequence that when aligned with sequence SEQ ID NO: 4 has a BLAST score between 200 and 250, or preferably between 250 and 300, or more preferably over 300 but less than 100% homology or even more preferably having a BLAST Score of 429 and 100% homology with the sequence SEQ ID NO: 4.
  • BLAST Basic Local Alignment Search Tool
  • a BLAST search enables a researcher to compare a query sequence with a library or database of sequences, and identify library sequences that resemble the query sequence above a certain threshold.
  • BLAST or Basis Local Alignment Search Tool uses statistical methods to compare a DNA or protein input sequence, also referred to as a query sequence to a database of nucleotide and protein (subject sequences) and returns sequences hits that have a level of similarity to the query sequence ranked based on the score.
  • score in the context of sequence alignments, indicates a numerical value that describes the overall quality of an alignment. Higher scores correspond to higher similarity and lower scores correspond to lower similarity. The score scale depends on the scoring system used for conducting the sequence alignment.
  • a BLAST score also referred to as bit score or max score in the BLAST output is a normalized score with respect to the scoring system provided by the BLAST algorithm.
  • the BLAST score defines the highest alignment score of a set of aligned segments from the same subject (database) sequences. The score is calculated from the sum of the match rewards and the mismatch, gap open an extend penalties independently for each segment.
  • the BLAST score normally gives the same sorting order as the expect value (E value) in the BLAST alignment output.
  • the“histidine kinase P1363” or“UrpS” and the “response regulator 1362” or“UrpR” typically form a two-component system herein also indicated as“1363/1362 TCS”, UrpRS TCS” or“UrpRS”.
  • two component system refers to a stimulus-response coupling mechanism that allows organisms to sense and respond to changes in many different environmental conditions [41].
  • Two-component systems typically consist of a membrane-bound histidine kinase that senses a specific environmental stimulus and a corresponding response regulator that mediates the cellular response, mostly through differential expression of target genes [42].
  • two-component signaling systems are found in all domains of life, they are most common in bacteria, particularly in Gram-negative and cyanobacteria [43].
  • Two- component systems accomplish signal transduction through the phosphorylation of a response regulator (RR) by a histidine kinase (HK).
  • Histidine kinases are typically homodimeric transmembrane proteins containing a histidine phosphotransfer domain and an ATP binding domain.
  • Response regulators can consist only of a receiver domain, but usually are multi-domain proteins with a receiver domain and at least one effector or output domain, often involved in DNA binding [43].
  • the HK Upon detecting a particular change in the cellular environment, the HK performs an autophosphorylation reaction, transferring a phosphoryl group from adenosine triphosphate (ATP) to a specific histidine residue.
  • the cognate response regulator (RR) then catalyzes the transfer of the phosphoryl group to an aspartate residue on the response regulator's receiver domain [44, 45]. This typically triggers a conformational change that activates the RR's effector domain, which in turn produces the cellular response to the signal, usually by activating or repressing expression of target genes [43].
  • a representative example of histidine kinase P1363 (UrpS) and response regulator 1362 (UrpR) in a two component system herein described are provided by the histidine kinase encoded by 1363 gene CCNA_01363 in Caulobacter crescentus (SEQ ID NOG), and the response regulator 1362 encoded by 1362 gene CCNA_01362 (SEQ ID NO:4), forming a two component systems respectively as will be understood by a skilled person.
  • the histidine kinase PI 363 (UrpS) and response regulator pi 362 (UrpR) can be heterologously expressed in the bacterial cell through genetic engineering of the cell performed to include in the cell a U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS) and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the gene encoding histidine kinase 1363 and the gene encoding response regulator 1362 (UrpR) response regulator UzcR are expressed upon activation of a controllable promoter.
  • a U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS) and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a
  • the histidine kinase P1363 (UrpS) and response regulator pl362 (UrpR) can be natively expressed in the bacterial cell.
  • the host cell of the U-biosensors herein described is capable of natively expressing the proteins of the U-sensing two-component system 1363 (UrpS) and 1362 (UrpR) described herein.
  • Those embodiments typically comprise certain proteobacterial cell such as alphaproteobacteria, betaproteobacteria or gammaproteobacteria comprising an endogenous 1363/1362 TCS (UrpRS) which can be identified and selected by methods to detect 1363 (UrpS) and/or 1362 (UrpR) genes in a candidate bacterial cell identifiable by a skilled person.
  • proteobacterial cell such as alphaproteobacteria, betaproteobacteria or gammaproteobacteria comprising an endogenous 1363/1362 TCS (UrpRS) which can be identified and selected by methods to detect 1363 (UrpS) and/or 1362 (UrpR) genes in a candidate bacterial cell identifiable by a skilled person.
  • presence of a 1363/1362 TCS or UrpRS in a proteobacterial cell can be identified by wet bench experiments, such as PCR, Southern blotting and additional techniques identifiable by a skilled person performed with histidine kinase PI 363 (UrpS) and response regulator pl362 (UrpR) and/or fragments thereof used as primers or probes for the related detection, followed by isolation and sequencing of the identified 1363 gene and/or 1362 gene as will be understood by a skilled person.
  • a 1363/1362 TCS (UrpRS) in a proteobacterial cell
  • presence of a 1363/1362 TCS (UrpRS) in a proteobacterial cell can be identified by performing a sequence alignment using BLASTP or PST BLAST or other alignment algorithms known to persons skilled in the art with the 1363 (UrpS) amino acid sequence of C. crescentus NA1000 (SEQ ID NO:3) and/or the 1362 (UrpR) protein sequence of Caulobacter crescentus NA1000 (SEQ ID NO:4) as a query sequence against protein sequences of a given proteobacterial cell, as would be understood by a skilled person.
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing a 1363 (UrpS) protein having 100% homology to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3), and/or having a BLAST Score of 851 when aligned with SEQ ID NO: 3, (herein also 1363/1362 Tier 1 proteobacteria or UrpRS Tier 1 proteobacteria) such as proteobacteria C. crescentus NA1000 and C. crescentus CB 15.
  • proteobacteria capable of natively expressing a 1363 (UrpS) protein having 100% homology to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3), and/or having a BLAST Score of 851 when aligned with SEQ ID NO: 3, (herein also 1363/1362 Tier 1 proteobacteria or UrpRS Tier 1 proteobacteria) such as proteobacteria C. crescentus
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score over 800 when aligned to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3) and with less than 100% homology to C. crescentus NA1000 1363) SEQ ID NO: 3, (herein also 1363/1362 Tier 2 proteobacteria or UrpRS Tier 2 proteobacteria) such as exemplary proteobacterium C. crescentus CB2 among others identifiable by persons skilled in the art.
  • proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score over 800 when aligned to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3) and with less than 100% homology to C. crescentus NA1000 1363) SEQ ID NO: 3, (herein also 1363/1362 Tier 2 proteobacteria or UrpRS Tier 2 prote
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score of 500 - 800 when aligned to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3) (herein also 1363/1362 Tier 3 or UrpRS Tier 3) such as proteobacteria Caulobacter henricii, Caulobacter sp. CCH5-E12, Caulobacter sp. OV484, Caulobacter sp. Root487D2Y, Caulobacter sp. Rootl455, Caulobacter sp. 12-67-6, Caulobacter sp.
  • proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score of 500 - 800 when aligned to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3) (herein also 1363/1362 Tier
  • Root487D2Y Caulobacter sp. Rootl455, Caulobacter sp. UNC358MFTsu5.1, Caulobacter sp. AP07 and Caulobacter , among others identifiable by persons skilled in the art.
  • the U sensing bacterial cell can comprise proteobacteria capable of natively comprising a 1362 (UrpR) binding site (SEQ ID NO: 1) in a phytase or 1361 promoter, also referred to herein as“Pphyt” and“P1361”, respectively (herein also indicated as 1363/1362 Tier 4 proteobacteria or UrpRS Tier 4 proteobacteria).
  • Exemplary proteobacteria within these embodiments comprise Caulobacter sp. Root342, Phenylobacterium sp. Root700, Caulobacter crescentus NA1000, Caulobacter sp. Rootl455, Caulobacter sp.
  • Root487D2Y Paracoccus sp. 228, Caulobacteraceae bacterium OTSz_A_272, Novosphingobium sp. AP12 PMI02, Hyphomicrobium sp. MCI, Hyphomicrobium denitrificans, Brevundimonas sp. Root 1279
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score of 300 - 500 when aligned to 1363 protein of C. crescentus NA1000 (SEQ ID NO: 3) (herein also indicated as 1363/1362 Tier 5 proteobacteria or UrpRS Tier 5 proteobacteria) such as exemplary proteobacteria Phenylobacterium sp. Root700, Phenylobacterium sp. Root700, Caulobacter sp. 39-67-4, Sphingopyxis sp. SCN 67-31, Phenylobacterium sp. SCN 70-31, Sphingopyxis flava, Caulobacteraceae bacterium OTSz_A_272, Sphingobium baderi, Caulobacterales bacterium 68-
  • proteobacteria capable of natively expressing 1363 (UrpS) proteins having a
  • alpha proteobacterium U9-li Caulobacter sp. 35-67-4, Sphingopyxis granuli, Sphingopyxis macro goltabida, Brevundimonas sp. Rootl279, Sphingopyxis macrogoltabida, Brevundimonas sp.
  • Rootl279 Sphingopyxis macrogoltabida, Hyphomonas polymorpha, Porphyrobacter mercurialis, Caulobacteraceae bacterium TH1-2, Hyphomonadaceae bacterium UKL13-1, Sphingopyxis macrogoltabida, Porphyrobacter mercurialis, and Novo sphingobium sp. PASSN1, among others identifiable by persons skilled in the art.
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score of 240-300 when aligned to 1363 protein of C. crescentus NA1000 (SEQ ID N03), (herein indicated also as 1363/1362 Tier
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 proteins having a BLAST Score of 200-240 when aligned to 1363 protein of C. crescentus NA1000 (SEQ ID NO:3) (herein also indicated as 1363/1362 Tier 7 proteobacteria or UrpRS Tier 7 proteobacteria), identifiable by persons skilled in the art.
  • proteobacteria capable of natively expressing 1363 proteins having a BLAST Score of 200-240 when aligned to 1363 protein of C. crescentus NA1000 (SEQ ID NO:3) (herein also indicated as 1363/1362 Tier 7 proteobacteria or UrpRS Tier 7 proteobacteria), identifiable by persons skilled in the art.
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score of less than 200 when aligned to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3) (herein also indicated as 1363/1362 Tier 8 proteobacteria or UrpRS Tier 8 proteobacteria), identifiable by persons skilled in the art (herein also indicated as Tier 8 proteobacteria).
  • the host cell is a proteobacterial cell of any one of 1363/1362 Tiers 1 to 6 (UrpRS Tiers 1 to 6)
  • the host proteobacterium can comprise a bacterial cell with a natively and/or heterologously expressed 1363/1362 TCS (UrpRS)endogenous to the host proteobacterium.
  • the host cell is a bacterial cell other than proteobacteria or a proteobacteria of 1363/1362 Tiers 7 and
  • the host is engineered to include a heterologous 1363/1362 TCS system in a configuration capable of heterologous expression in the host bacteria as will be understood by a skilled person.
  • the host cell is a proteobacteria of 1363/1362 Tier 6
  • the host can be firstly tested for the presence of a natively expressed 1363/1362 TCS endogenous to the host proteobacteria according to procedure identifiable by a skilled person. The test can be performed by transforming the cell with a plasmid or other vector containing Pphyt or PI 361 -regulated gfp fusion and assaying the system for U-dependent induction of GFP as will be understood by a skilled person.
  • the host cell does not possess a natively expressed 1363/1362 TCS, the host can be engineered to include a heterologous 1363/1362 TCS system in a configuration capable of heterologous expression in the host bacteria.
  • the heterologous 1363/1362 TCS can be a 1363/1362 TCS system from a 1363/1362 Tier 1, a 1363/1362 Tier 2, a 1363/1362 Tier 3, a 1363/1362 Tier 4 or a 1363/1362 Tier 5 proteobacteria, and it is preferably a 1363/1362 TCS from a 1363/1362 Tier 2 proteobacteria and more preferably from a 1363/1362 Tier 1 proteobacteria.
  • the native 1363/1362 TCS of the host cell is preferably knocked out, in particular in embodiments wherein the host organism is a proteobacteria of any one of 1363/1362 Tiers 1 to 6.
  • the native 1363/1362 TCS of the host cell can be knocked out by deleting or inactivating the 1362 gene cluster only or by deleting or otherwise inactivating both the 1362 and 1363 gene clusters.
  • UC Fi-biosensors herein described the proteobacterial cell is capable of natively and/or heterologously expressing a U-sensitive histidine kinase UzcS, and a transcriptional response regulator UzcR.
  • E SEQ ID NO: 5
  • a sequence that when aligned with sequence SEQ ID NO: 5 has a BLAST score has a BLAST score greater than 300 and less than 500, or preferably greater than 500 and less than 767, or more preferably a BLAST score greater than 800 and a homology with SEQ ID NO: 5 less than 100%, or even more preferably BLAST score of 925 and an homology of 100% with SEQ ID NO: 5.
  • transcriptional response regulator UzcR or“UzcR’ as used herein indicates a transcriptional regulator having the amino acid sequence: MRILIIEDDLEAAGAMAHGLKEAGYDVAHAPDGEAGLAEAQKGGWDVLVVDRMMPK MDGVTVVETLRREGDQTPVLFLSALGEVNDRVVGLKAGADDYLVKPYAFPELMARVE ALS RRRETG A V ATTLKV GELEMNLINRT VHRQGKEIDLQPREF QLLEFMMRHAGQS VT RTMLLEKVWEYHFDPQTNVIDVHISRLRSKIDKGFDRAM LQTVRGAGYRLDP (SEQ ID NO: 6).
  • sequence SEQ ID NO: 6 has a BLAST score greater than 250 and less than 300, or preferably greater than 300 and less than 400, or more preferably a BLAST score greater than 400 and a homology with SEQ ID NO: 6 less than 100%, or even more preferably a BLAST score of 452 and an homology of 100% with SEQ ID NO: 6.
  • UzcRS two-component system in the sense of the disclosure, also referred to herein as“UzcRS two-component system” or“UzcRS TCS” which is similar to the 1363/1362 TCS system herein described and exemplified by the schematics of
  • UzcRS two component system refers to a regulatory system responsible for U, Zn, and Cu-dependent regulation of numerous genes in Caulobacter crescentus [3].
  • the UzcRS two component system comprises an OmpR/PhoB family response regulator (RR) and a histidine kinase (HK) containing a 123 amino acid periplasmic domain, placing it in the periplasmic- sensing class of histidine kinases [42].
  • a representative example of histidine kinase UzcS and transcriptional response regulator UzcR in a UzcRS two components system herein described are provided by histidine kinase encoded by UzcS gene CCNA_02842in C. crescentus NA1000 (SEQ ID NO: 5) and by a transcriptional response regulator, for example encoded by UzcR gene CCNZ_02485in C. crescentus NA1000 (SEQ ID NO: 6) as will be understood by a skilled person.
  • the histidine kinase UzcS and transcriptional response regulator UzcR can be heterologously expressed in the bacterial cell through genetic engineering of the cell performed to include in the cell a U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the gene encoding histidine kinase UzcS and the gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
  • the histidine kinase UzcS and transcriptional response regulator UzcR are natively expressed in the bacterial cell.
  • the host cell of the U-biosensors herein described is capable of natively expressing the proteins of the U- sensing UzcS/UzcR TCS described herein.
  • Those embodiments typically comprise certain proteobacterial cell such as alphaproteobacteria, betaproteobacteria or gammaproteobacteria comprising an endogenous UzcS/UzcR TCS which can be identified and selected by methods to detect UzcS and/or UzcR genes in a candidate bacterial cell identifiable by a skilled person.
  • UzcS/UzcR TCS presence of a UzcS/UzcR TCS in a proteobacterial cell can be identified by wet bench experiments, such as PCR, Southern blotting and additional techniques identifiable by a skilled person performed with histidine kinase UzcS and response regulator UzcR and/or fragments thereof used as primers or probes for the related detection, followed by isolation and sequencing of the identified UzcS gene and/or UzcR gene as will be understood by a skilled person.
  • the presence of an UzcS/UzcR TCS can also be identified by introducing in the cell a UzcR-regulated GFP fusion promoter and detecting GFP fluorescence thus testing for U- dependent fluorescence as will be understood by a skilled person.
  • a UzcS/UzcR TCS in a proteobacterial cell can be identified by performing a sequence alignment using BLASTP or PSI-BLAST or other alignment algorithms known to persons skilled in the art with the UzcS amino acid sequence of C. crescentus NA1000 (SEQ ID NO: 5) and/or the UzcR protein sequence of Caulobacter crescentus NA1000 (SEQ ID NO: 6) as a query sequence against protein sequences of a given proteobacterial cell, as would be understood by a skilled person.
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having 100% homology to UzcS protein of C. crescentus NA1000 (SEQ ID NO: 5) and a BLAST score of 925 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcS/UzcR Tier 1 proteobacteria).
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a homology of less than 100% to UzcS protein of C. crescentus NA1000 (SEQ ID NO: 5) and a BLAST score greater than 800 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 2 proteobacteria).
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins a BLAST score greater than 767 and lower than 800 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 3 proteobacteria).
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score greater than 500 and less than 767 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 4 proteobacteria).
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score greater than 300 and less than 500 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 5 proteobacteria).
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score greater than 250 and less than 300 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 6 proteobacteria).
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score greater than 200 and less than 250 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 7 proteobacteria).
  • the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score less than 200 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 8 proteobacteria).
  • the host cell is a proteobacterial cell of any one of UzcRS Tiers 1 to 6, the host proteobacterium can comprise a bacterial cell with a natively and/or heterologously expressed UzcRS TCS endogenous to the host proteobacterium.
  • the host cell is a bacterial cell other than proteobacteria or a proteobacteria of UzcRS Tiers 7 and 8 the host is engineered to include a heterologous UzcRS TCS system in a configuration capable of heterologous expression in the host bacteria as will be understood by a skilled person.
  • the heterologous UzcRS TCS can be a UzcRS TCS system from a 1363/1362 Tier 1, a UzcRS Tier 2, a UzcRS Tier 3, a UzcRS Tier 4 or a UzcRS Tier 5 proteobacteria, and it is preferably a UzcRS TCS from a UzcRS Tier 2 proteobacteria and more preferably UzcRS TCS from a UzcRS Tier 1 proteobacteria.
  • the native UzcRS TCS of the host cell is preferably knocked out, in particular in embodiments wherein the host organism is a proteobacteria of any one of UzcRS Tiers 1 to 6.
  • the native UzcRS TCS of the host cell can be knocked out by deleting or otherwise inactivating the UzcR gene cluster only or by deleting or otherwise inactivating both the UzcR and UzcS gene clusters according to techniques identifiable by a skilled person (e.g.
  • a bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363, and transcriptional response regulator 1362, and/or capable of natively and/or heterologously expressing a U-sensitive histidine kinase UzcS, and a transcriptional response regulator UzcR is genetically engineered to include a U-sensitive genetic molecular component configured to report and/or neutralize U.
  • molecular component indicates a chemical compound comprised in a cellular environment.
  • exemplary molecular components thus comprise polynucleotides, such as ribonucleic acids or deoxyribonucleic acids, polypeptides, polysaccharides, lipids, amino acids, peptides, sugars and/or other small or large molecules and/or polymers that can be found in a cellular environment.
  • the term“genetic molecular component” as used herein indicates a molecular unit formed by a gene, an RNA transcribed from the gene or a portion thereof and optionally a polypeptide or a protein translated from the transcribed RNA.
  • a genetic molecular component comprises a promoter operatively connected to the gene of the genetic molecular component, so that the promoter is configured to initiate transcription of said gene.
  • promoters are typically located adjacent to the transcription start sites of genes, on the same strand and upstream on a DNA sequence (towards the 5' region of the sense strand), and for transcription to occur, the enzyme that synthesizes RNA, known as RNA polymerase, attaches to the promoter.
  • Promoters contain DNA sequences identifiable by those skilled in the art and described herein, such as those that provide binding sites for RNA polymerase and also for proteins that function as transcription regulatory factors that can either activate or repress gene transcription.
  • transcription regulatory factor refers to any type of factors that can function by acting on a regulatory DNA element such as a promoter or enhancer sequence.
  • the transcription regulatory factors can be broadly classified into a transcription repression factor (also referred to as “repressor”) and a transcription activation factor (also referred to as“activator”).
  • the transcription repression factor acts on a regulatory DNA element to repress the transcription of a gene, thereby reducing the expression level of the gene.
  • the transcription activation factor acts on a regulatory DNA element to promote the transcription of a gene, thereby increasing the expression level of the gene. Both the transcription repression factors and the transcription activation factors can be used as one or more components in the gene circuits herein described.
  • a transcription regulatory factor has typically at least one DNA-binding domain that can bind to a specific sequence of enhancer or promoter sequences. Some transcription factors bind to a DNA promoter sequence near the transcription start site and help form the transcription initiation complex. Other transcription factors bind to other regulatory sequences, such as enhancer sequences, and can either stimulate or repress transcription of the related gene. Examples of specific transcription repression factors include TetR, Lacl, LambdaCI, PhlF, SrpR, Qacl, BetR, LmrA, AmeR, LitR, met, and others identifiable by a skilled person, as well as homologues of known repression factors, that function in both prokaryotic and eukaryotic systems.
  • transcription activation factors include AraC, LasR, LuxR, IpgC, MxiE, Gal4, GCN4, GR, SP1, CREB, and additional activation factors identifiably by a skilled person as well as homologues of known activation factors that function in both prokarayotic and eukaryotic systems also identifiable by a skilled person.
  • Exemplary inducible regulators that can be used in Caulobacter comprise VanR (regulated by vanillate) and XylR (regulated by xylose), as well as others identifiable by those skilled in the art.
  • a gene comprised in a genetic molecular component is a polynucleotide that can be transcribed to provide an RNA and typically comprises coding regions as well as one or more regulatory sequence regions which is a segment of a nucleic acid molecule which is capable of increasing or decreasing transcription or translation of the gene within an organism either in vitro or in vivo.
  • coding regions of a gene herein described can comprise one or more protein coding regions which when transcribed and translated produce a polypeptide, or if RNA is the final product only a functional RNA sequence that is not meant to be translated.
  • RNA of a genetic molecular component comprises any RNA that can be transcribed from a gene, such as a messenger ribonucleic acid (mRNA), short interfering ribonucleic acid, and ribonucleic acid capable of acting as regulating factors in the cell.
  • mRNA messenger ribonucleic acid
  • ribonucleic acid capable of acting as regulating factors in the cell.
  • mRNA comprised in a genetic molecular component comprise regions coding for the protein as well as regulatory regions e.g. ribosome binding site domains (“RBS”), which is a segment of the upstream (5’) part of an mRNA molecule to which the ribosomal machinery of a cell binds to position the message correctly for the initiation of translation.
  • RBSs control the accuracy and efficiency with which the translation of mRNA begins.
  • mRNA can have additional control elements encoded, such as riboregulator sequences or other sequences that form hairpins, thereby blocking the access of the ribosome to the Shine-Delgarno sequence and requiring an external source, such as an activating RNA, to obtain access to the Shine-Delgarno sequence.
  • a protein comprised in a molecular component can be proteins with activating, inhibiting, binding, converting, or reporting functions. Proteins that have activating or inhibiting functions typically act on operator sites encoded on DNA, but can also act on other molecular components. Proteins that have binding functions typically act on other proteins, but can also act on other molecular components. Proteins that have converting functions typically act on small molecules, and convert small molecules from one small molecule to another by conducting a chemical or enzymatic reaction. Proteins with converting functions can also act on other molecular components.
  • Proteins with reporting functions have the ability to be easily detectable by commonly used detection methods (absorbance, fluorescence, for example), or otherwise cause a reaction on another molecular component that causes easy detection by a secondary assay (e.g. adjusts the level of a metabolite that can then be assayed for).
  • the activating, inhibiting binding, converting, or reporting functions of a protein typically form the interactions between genetic components of a genetic circuit.
  • Exemplary proteins that can be comprised in a genetic molecular component comprise monomeric proteins and multimeric proteins, proteins with tertiary or quaternary structure, proteins with linkers, proteins with non-natural amino acids, proteins with different binding domains, and other proteins known to those skilled in the art.
  • Specific exemplary proteins include TetR, Lacl, LambdaCI, PhlF, SrpR, Qacl, BetR, LmrA, AmeR, LitR, met, AraC, LasR, LuxR, IpgC, MxiE, Gal4, GCN4, GR, SP1, CREB, and others known to a skilled person in the art.
  • a “U-sensing genetic molecular component” or “U-sensitive genetic molecular component” as used herein indicates a genetic molecular component wherein the gene of the genetic molecular component is under control of a U-sensing or U-sensitive promoter.
  • At least one U-sensitive promoter comprises a U-sensitive 1362 binding site having a DNA sequence
  • Ni is C or T, preferably C
  • N 2 is G or A, preferably G;
  • N 3 is T or C, preferably T;
  • N 4 is C; Ns is A or G, preferably A;
  • N ⁇ is G or C, preferably G;
  • Ns is any nucleotide
  • N 9 is any nucleotide
  • Nio is any nucleotide
  • N 11 is any nucleotide
  • N12 is T or C
  • Ni 4 is T or C, preferably T;
  • Nis is C
  • Ni 6 is A or C, preferably A; Nn is G; and
  • Ni 8 is C or G
  • Ni to Nn are selected independently.
  • nucleotide Ni of the regulator direct repeat is in a position from about 16 nucleotides downstream of the transcription start site of the genetic molecular component as described herein to about 40 nucleotides upstream of the transcription start site of the genetic molecular component or genetic molecular component.
  • the U- sensitive promoter further comprises nucleotides N19N20N21 (SEQ ID NO: 83), downstream of SEQ ID NO: 1 wherein each of N19 to N21 can independently be any nucleotide, and therefore N19 is any nucleotide N20 is any nucleotide; and N21 is G.
  • the 1362 binding site has a DNA sequence CGTCAGCNNNNTGTCAGC (SEQ ID NO:7), C GTC AGGNNNNTGTC AGG (SEQ ID NO: 8), CGTCAGCNNNNTGTCAGG (SEQ ID NO:9),
  • CGTCAGCNNNNCGTCAGG (SEQ ID NO: 10), TGTCAGCNNNNTGTCAGC (SEQ ID NO: 11), CGCCTGCNNNNCGTCAGC (SEQ ID NO: 12), C GTC AGGNNNN C GTC AGC (SEQ ID NO: 13), CGTCAGCNNNNTGTCAGC (SEQ ID NO: 14), TGTCAGGNNNNTGTCAGC (SEQ ID NO: 15), CGTCAGCNNNNCGTCAGT (SEQ ID NO: 16),
  • CCGCGGGNNNNTGTCAGG (SEQ ID NO: 17), CGTCGGGNNNNAGACCGG (SEQ ID NO: 18), CGTCCGGNNNNCGTCAGA (SEQ ID NO: 19), CAACGCCNNNNCGTCAGC (SEQ ID NO: 20), CATCAGGNNNNCGTCAGC (SEQ ID NO: 21), CGCAGGGNNNNTGCAAGC (SEQ ID NO: 22), CATCAGCNNNNCGTCAGC (SEQ ID NO: 23),
  • N can be any nucleotide.
  • the U-biosensor comprises a genetically engineered proteobacterial cell capable of natively and/or heterologously expressing histidine kinase UzcS, and U-sensitive transcriptional response regulator UzcR
  • at least one U-sensitive promoter comprises a UzcR binding site with an m_5 site configured for binding UzcR, having a DNA sequence:
  • N 7 -N 12 is independently any nucleotide, and in some embodiments N 7 -N 11 can independently be A. In an embodiment, each of N 7 -N 12 can be A.
  • the endogenous genes encoding the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR are knocked out and the genetically engineered proteobacterial cell is further engineered to include a U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
  • the U-sensitive promoter comprises a UzcR binding site can be either a constitutively active promoter or an inducible promoter. In preferred embodiments, the promoter is constitutively active.
  • m_5 site or““UczR binding site” as used herein refers to a semi- palindromic consensus DNA binding site of sequence CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) [3], wherein N7-N12 is independently any nucleotide, and in some embodiments any one of N7-N11 can independently be A.
  • One UzcR dimer likely to binds one m_5 site [3].
  • a variant of an m_5 site wherein CATTAC (SEQ ID NO:29) is mutated to CAATAG (SEQ ID NO:30) is not bound by UzcR and a variant of an m_5 site wherein TTAA (SEQ ID NO:31) is mutated to TAAT (SEQ ID NO:32) is no longer activated by UzcR [3].
  • UzcRS -regulated promoters comprise those having naturally-occurring m_5 sites or m_5 sites that are introduced into a promoter through genetic engineering. Accordingly, UzcRS -regulated promoters comprise DNA sequence elements required for RNA Polymerase binding, as well as one or more m_5 sites, such that the promoter is configured to be regulated by the UzcRS two-component system. Similar to promoters comprising 1362 binding sites, in UzcRS -regulated promoters, the s-RNAP biding sites typically have low sequence homology to the canonical G 73 -RNAP -10 and -35 hexamer sequences. Accordingly, typically transcriptional activation of native UzcRS -regulated promoters occurs through binding of UzcR to the promoter, consistent with little observed transcriptional activation in absence of UzcR.
  • an UzcRS -regulated promoter can comprise 1 - 3 copies of the m_5 site.
  • one or more m_5 sites are located at a position from about -50 to about -100 upstream of the TSS, preferably at a position -52/53 or - 62/63 upstream of the TSS, considering the first nucleotide of Seq ID NO: 2 as the first nucleotide of the m_5 site upstream from the TSS (see the position of the first nucleotide (C on the 5’ end) of the binding site in Figure 25 A) wherein the m_5 site is configured for activation of the UzcRS -regulated promoter (see e.g.
  • one or more m_5 sites are located at a position 52 to53 bp upstream of the TSS or at a position 62 or-63 bp upstream of the TSS considering the first nucleotide of Seq ID NO: 2 as the first nucleotide of the m_5 site upstream from the TSS (see the position of the first nucleotide (C on the 5’ end) of the binding site in Figure 25 A) wherein the m_5 site is configured for activation of the UzcRS -regulated promoter (see e.g. Figure 25 B).
  • one or more m_5 sites are located at a position within 100 nucleotides upstream of the TSS considering the first nucleotide of Seq ID NO: 2 as the first nucleotide of the m_5 site upstream from the TSS (see the position of the first nucleotide (C on the 5’ end) of the binding site in Figure 25 A).
  • the m_5 site when one or more m_5 sites are located at a position from about -49 upstream of the TSS to +25 downstream of the TSS, in particular, -33 upstream of the TSS to +15 nt downstream of the TSS, the m_5 site is configured for repression of the UzcRS -regulated promoter.
  • the U-sensitive promoter is configured such that upon binding of the response regulator UzcR to the m_5 site, the U-sensitive promoter is activated and transcription of a gene operatively connected to the U-sensitive promoter within the related genetic molecular component is initiated.
  • the U-sensitive promoter is configured such that upon binding of the U-sensitive transcriptional regulator to the U-sensitive transcriptional UzcR binding site, the U-sensitive promoter is repressed and transcription of a gene operatively connected to the U- sensitive promoter within the related U-sensing genetic molecular component is not initiated.
  • one or more m_5 sites are located at a position -49 to +25 bp from the TSS considering the first nucleotide of Seq ID NO: 2 as the first nucleotide of the m_5 site upstream from the TSSThe approximate range is -49 to +25 bp from the TSS (see the position of the first nucleotide (C on the 5’ end) of the binding site ( Figure 25A).
  • a m_5 site is engineered downstream of a transcription start site of a U-sensitive promoter such as a P 1361 promoter or P phyt promoter (see Example 2).
  • a m_5 site downstream of the TSS of e.g. P phyt minimizes activation of P phyt in both the presence and absence of U.
  • the latter is preferable as it minimizes UzcRS expression levels when no U is present, minimizing cross -reactivity with Zn and Cu.
  • Examples of promoters regulated by UzcRS comprise P U r cA , P u r cB , Pi968, and others identifiable by those skilled in the art, such as those described in Park et ah, 2017 [3] herein incorporated by reference in its entirety (see also Example 11).
  • one or more U-sensing promoters herein described are comprised within a U sensing genetic molecular component which is a genetically engineered polynucleotide construct configured to regulate expression of one or more RNA and/or protein encoding genes through the one or more U-sensing promoter.
  • the U-sensitive promoter includes a 1362 binding site functioning as a binding site for natively expressed U-sensitive transcriptional regulator 1362 and/or a UzcR binding site for natively expressed UzcR.
  • the histidine kinase 1363, and U-sensitive transcriptional regulator 1362 are therefore encoded respectively by 1363 and 1362 genes natively encoded in the genome of the proteobacterial cell and the encoded 1363 and 1362 proteins can be natively expressed in the proteobacterial cell.
  • the histidine kinase UzcR, and U-sensitive transcriptional regulator UzcS are therefore encoded respectively by uczR and uczS genes natively encoded in the genome of the proteobacterial cell and the encoded UczR and UczS proteins are natively expressed in the proteobacterial cell.
  • 1363 and 1362 genes and/or the uczR and uczS genes are introduced into the proteobacterial cell of (e.g. a species of the subclass Caulobacteridae ) within one or more U-sensing regulator genetic molecular components configured to express the proteins 1363 and 1362 and/or UczR and UczS proteins upon activation of a controllable promoter.
  • proteobacterial cell of e.g. a species of the subclass Caulobacteridae
  • U-sensing regulator genetic molecular components configured to express the proteins 1363 and 1362 and/or UczR and UczS proteins upon activation of a controllable promoter.
  • the genetic molecular components comprise the 1363 1362, uczR and/or uczS genes together with one or more regulatory regions configured to directly initiate expression of operatively connected 1363 and 1362 genes and/or operatively connected uczR and uczS genes.
  • a genetic molecular component introduced into the proteobacterial cell can comprise 1363 and 1362 genes and/or uczR and uczS genes in a same genetic molecular component, while in other embodiments the 1363 gene, the 1362 gene, the uczR gene and/or the uczS gene are comprised in different genetic molecular components.
  • the one or more regulatory regions operatively connected to the 1363 and/or 1362 genes and/or to the uczR and uczS genes can comprise any promoter identifiable by skilled persons that is capable of initiating gene expression in a Caulobacteridae cell.
  • Exemplary promoters that can be used to express 1363 and/or 1362 in Caulobacteridae comprise inducible promoter systems such as VanR (regulated by vanillate) and XylR (regulated by xylose), as well as others identifiable by those skilled in the art.
  • any constitutive promoter identifiable by those skilled in the art that has been characterized as functional to express an operatively linked gene of interest in a proteobacterial species of interest can be used to express 1363 and 1362 and/or uczR and uczS genes in the proteobacterial species of interest.
  • the proteobacterial cell is further genetically engineered so that expression of its native a 1363 gene, 1362 gene, a uzcS and/or uzcR gene is inactivated by gene knockout.
  • the endogenous genes encoding histidine kinas 1363, transcriptional regulator 1362, histidine kinase UzcS, and/or the U-sensitive transcriptional regulator UzcR are knocked out and the genetically engineered proteobacterial cell is further engineered to include a U-sensing regulator component comprising a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR in a configuration wherein the a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
  • a U-sensing regulator component comprising a gene encoding histidine kinase UzcS
  • a gene encoding response regulator UzcR in a configuration wherein the a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR
  • the promoter controlling the expression of the heterologous 1363 gene, the 1362 gene, the UczR gene and/or the UczS gene can be either a constitutively active promoter or an inducible promoter. In preferred embodiments, the promoter is constitutively active.
  • the proteobacterial cell is further genetically engineered so that expression of the host endogenous 1363 gene, 1362 gene, uczR gene and/or uczS gene is inactivated by gene knockout.
  • Methods for performing genetic knockout are identifiable by persons skilled in the art, such as gene targeting using techniques such as homologous recombination, or transposon-mediated mutagenesis, or gene editing techniques such as those using CRISPR/Cas9 among others known to those skilled in the art.
  • the U sensing genetic molecular component herein described is a genetically engineered polynucleotide construct configured to regulate expression of one or more RNA and/or protein-encoding genes through one or more U-sensing promoter.
  • the U-sensitive promoter is configured such that upon binding of the U-sensitive transcriptional regulator 1362 or UzcR to the U-sensitive corresponding binding site, the U-sensitive promoter is activated and transcription of a gene operatively connected to the U-sensitive promoter within the related genetic molecular component is initiated.
  • the U-sensing promoter comprises a 1362 binding site is located in a position wherein nucleotide Nis of the regulator direct repeat (SEQ ID NO: 1) is from about 17 nucleotides upstream of the transcription start site of the genetic molecular component to about 40 nucleotides upstream of the transcription start site of the genetic molecular component
  • the regulator direct repeat is configured to function as a transcriptional activator binding site.
  • the U-sensitive promoter is configured such that upon binding of the U-sensitive transcriptional regulator to the U-sensitive transcriptional 1362 binding site, the U-sensitive promoter is repressed and transcription of a gene operatively connected to the U- sensitive promoter within the related U-sensing genetic molecular component is not initiated.
  • the 1362 binding site can be located in a position wherein nucleotide Ni of the regulator direct repeat (SEQ ID NO:l) is from about 16 nucleotides downstream of the transcription start site of the genetic molecular component to about 16 nucleotides upstream of the transcription start site of the genetic molecular component or genetic molecular component, the 1362 binding site configured to function as a transcriptional repressor binding site.
  • SEQ ID NO:l nucleotide Ni of the regulator direct repeat
  • substitution of a regulator repeat binding site polynucleotide sequence described herein for a promoter polynucleotide sequence comprising a -35 and/or a -10 hexamer sequence of the promoter, or one or more nucleotides at the transcriptional start site or downstream of the transcriptional start site, is expected to provide a promoter comprising a 1362 binding site configured to repress transcription of the promoter.
  • the regulator direct repeat is located in a position wherein nucleotide Nis of the regulator direct repeat (SEQ ID NO: l) is from about 17 nucleotides upstream of the transcription start site of the U-sensing genetic molecular component to about 40 nucleotides upstream of the transcription start site of the genetic molecular component or genetic molecular component, such that the regulator direct repeat is configured to function as a transcriptional activator binding site
  • positions in the promoter are designated relative to the transcriptional start site, where transcription of DNA begins for the gene of interest. Positions upstream (towards the 5’ end of the promoter) are negative numbers counting back from -1, for example -10 is a position 10 base pairs upstream of the transcription start site.
  • transcription start site refers to the location where transcription starts at the 5’ end of an encoded gene sequence.
  • the location of the transcription start site is typically referred to as +1 relative to the 3’ end of a promoter operatively connected to the gene.
  • a putative transcription start site can be detected using techniques such as differential RNA-seq (dRNA-seq) [47], which can differentially detect primary transcripts having triphosphorylated 5' ends, and processed RNAs which do not.
  • Additional techniques to detect transcription start sites known in the art comprise bioinformatic analysis to identify enrichment of promoter elements upstream of a putative transcriptional start site, and experimental validation of selected putative transcription start sites, for example using primer extension methods [48], or by using Northern blots to detect the associated RNAs, among other techniques identifiable by those skilled in the art [49] .
  • the U-sensitive promoter comprising the 1362 binding site is a Pi36i promoter.
  • the term“Pi36i promoter” as used herein refers to the promoter that natively regulates expression of an operon comprising CCNA_01361, CCNA_1362 and CCNA_1363 genes in Caulobacter crescentus.
  • the DNA sequence of Pi36i is shown in Table 4.
  • the U-sensitive promoter comprising the 1362 binding site is a P p h y t promoter.
  • the term“Pphyt promoter” as used herein refers to the promoter that natively regulates expression of an operon comprising CCNA_01353, CCNA_01352_, and CCNA_01351 genes in Caulobacter crescentus.
  • the DNA sequence of P pi , yL is shown in Table 4.
  • Caulobacter crescentus (Poindexter 1964) refers to a Gram- negative, oligotrophic bacterium widely distributed in fresh water lakes and streams.
  • Caulobacter is an obligate aerobe with a ubiquitous presence in aqueous environments where it is well-adapted to life under low-nutrient conditions [50]. Caulobacter species tolerate high concentrations of U [51, 52], are found in U- contaminated sites [53], and can mineralize U through the formation of uranyl phosphate precipitates [54]. Multi-omics studies to elucidate the U stress response pathways in C. crescentus have revealed many highly-induced genes that are not induced by Cd, Cr, Pb or Se
  • the U-sensitive promoter is a UzcRS -regulated promoters comprising DNA sequence elements required for RNA Polymerase binding, as well as one or more sequence elements for binding UzcR known as an m_5 site [3], as understood by those skilled in the art and herein also identified as UczR binding site.
  • Examples of promoters regulated by UzcRS comprise PurcA, PurcB, Pi968, and others identifiable by those skilled in the art, such as those described in Park et ah, 2017 [3].
  • UzcRS -regulated promoters comprise those having naturally-occurring m_5 sites or m_5 sites that are introduced into a promoter through genetic engineering. Accordingly, UzcRS -regulated promoters comprise DNA sequence elements required for RNA Polymerase binding, as well as one or more m_5 sites, such that the promoter is configured to be regulated by the UzcRS two-component system. Similar to promoters comprising 1362 binding sites, in UzcRS -regulated promoters, the s-RNAP biding sites typically have low sequence homology to the canonical G 73 -RNAP -10 and -35 hexamer sequences. Accordingly, typically transcriptional activation of native UzcRS -regulated promoters occurs through binding of UzcR to the promoter, consistent with little observed transcriptional activation in absence of UzcR.
  • an UzcRS -regulated promoter can comprise 1 - 3 copies of the m_5 site.
  • the m_5 site when one or more m_5 sites are located at a position from about -34 to about -100 upstream of the TSS, preferably at a position -43 to -53 upstream of the TSS, the m_5 site is configured for activation of the UzcRS-regulated promoter [3] .
  • the m_5 site when one or more m_5 sites are located at a position from about -33 upstream of the TSS to +15 nt downstream of the TSS, the m_5 site is configured for repression of the UzcRS-regulated promoter.
  • a promoter can be either activated or repressed by UzcR.
  • a m_5 site is engineered downstream of a transcription start site of a U- sensitive promoter such as a P1361 promoter or P phyt promoter (see Example 2).
  • insertion of a m_5 site downstream of the TSS of e.g. P phyt minimizes activation of P phyt in both the presence and absence of U.
  • the latter is preferable as it minimizes UzcRS expression levels when no U is present, minimizing cross -reactivity with Zn and Cu.
  • the U-sensitive promoter can be a promoter genetically engineered to comprise the U-sensitive 1362 binding site.
  • a promoter can be engineered comprising a U-sensitive 1362 binding site having Ni 8 of SEQ ID NO:l at a position of about -40 to -42 upstream of the transcription start site, as in exemplary promoters Pi36i and P phyt (see e.g. Figure 7).
  • the U-sensitive promoters described herein can further comprise a holo-RNA Polymerase (RNAP) binding site.
  • RNAP holo-RNA Polymerase
  • promoters in bacteria such as members of Caulobacteridae require DNA sequence elements for s-RNAP binding for initiation of transcription.
  • the s-RNAP biding site typically have low sequence homology to the canonical G 73 -RNAP -10 and -35 hexamer sequences.
  • a promoter comprising a 1362 binding site at a position configured for transcriptional activation (e.g. wherein nucleotide Nis of the regulator direct repeat is located about -17 to about -40 upstream of the TSS), such as exemplary promoters Pi36i and P phyt , s- RNAP binding likely requires binding of the U-responsive transcriptional factor to the 1362 binding site, consistent with the observed low level of transcriptional activation in absence of U (see Examples section).
  • RNAP RNA polymerase
  • holoenzyme refers to enzymes that contain multiple protein subunits, such as RNA polymerases, wherein the holoenzyme is a complete complex containing all the subunits needed for activity.
  • the term“holoenzyme” can also refer to an enzyme together with one or more cofactors required for activity.
  • a promoter is recognized by RNA polymerase (RNAP) and an associated sigma factor, and the complex is referred to as an“RNAP holoenzyme” or“holo-RNAP”.
  • RNAP holoenzyme An example of a RNAP holoenzyme in Caulobacteridae is RNAP holoenzyme containing s - 73 .
  • the UO2F2 biosensors comprising the U-sensitive molecular component and/or U- sensitive genetic circuits comprising 1363/1362 TCS and/or UczRS TCS described herein in several embodiments can show improved selectivity for U compared to those previously described, such as in Hillson et ah, 2007 [55], as illustrated in the Examples.
  • the previously unknown exemplary U-responsive Caulobacter crescentus promoters P1361 and P p h yt show highly selective gene expression upregulation in response to U, in contrast to Caulobacter crescentus promoter P U rcA [55], which also shows upregulation of gene expression in response to other metals such as Zn, Cu and Cd.
  • a U-sensitive promoter such as a P1361 or a P p h yt promoter, or a U-sensitive promoter genetically engineered to comprise the U-sensitive 1362 binding site can be used as a stand-alone selective U-sensing promoter.
  • the U-Sensing promoter of the present disclosure directly or indirectly controls the expression of a reportable molecular component and/or a U-neutralizing molecular component.
  • the term“reportable molecular component” as used herein indicates a molecular component capable of detection in one or more systems and/or environments.
  • the terms“detect” or“detection” as used herein indicates the determination of the existence, presence or fact of a target in a limited portion of space, including but not limited to a sample, a reaction mixture, a molecular complex and a substrate.
  • The“detect” or“detection” as used herein can comprise determination of chemical and/or biological properties of the target, including but not limited to ability to interact, and in particular bind, other compounds, ability to activate another compound and additional properties identifiable by a skilled person upon reading of the present disclosure.
  • the detection can be quantitative or qualitative.
  • a detection when it refers, relates to, or involves the measurement of quantity or amount of the target or signal (also referred as quantitation), which includes but is not limited to any analysis designed to determine the amounts or proportions of the target or signal.
  • a detection is“qualitative” when it refers, relates to, or involves identification of a quality or kind of the target or signal in terms of relative abundance to another target or signal, which is not quantified.
  • the reportable molecular component can be a molecular component linked to or comprising a label
  • label refers to a compound capable of emitting a labeling signal, including but not limited to radioactive isotopes, fluorophores, chemiluminescent dyes, chromophores, enzymes, enzymes substrates, enzyme cofactors, enzyme inhibitors, dyes, metal ions, nanoparticles, metal sols, ligands (such as biotin, avidin, streptavidin or haptens) and the like.
  • fluorophore refers to a substance or a portion thereof which is capable of exhibiting fluorescence.
  • the genetic molecular component comprises a“reporter gene”, which can be any genetically- encoded reportable molecular component.
  • the terms“genetically-encoded reportable molecular component”, “reportable genetic molecular component”, “genetically encoded reporter” or“reporter gene” comprises polynucleotide-encoded RNA and/or proteins having reportable characteristics identifiable by those skilled in the art and as described herein.
  • a reporter gene can be placed under the regulatory control of a promoter, and expression of the genetically-encoded reportable molecular component thereby serves as an indication of activation of the promoter in a host organism comprising the reporter gene.
  • a reporter gene can be fused to another gene under the regulatory control of the promoter, such as a gene encoding a protein natively regulated by the promoter, so that the promoter regulates the expression of a fusion gene encoding a fusion protein comprised of the natively regulated protein covalently linked to the reportable molecular component.
  • a reporter gene that is not natively expressed in the host organism, since the expression of the genetically- encoded reportable molecular component is used as a marker of activation of the promoter in the host organism.
  • Exemplary genetically encoded reportable molecular components comprise fluorescent proteins such as green fluorescent protein (GFP) from Aequorea victoria or Renilla reniformis, red fluorescent protein from Discosoma species (dsRED), and variants thereof, beta galactosidase encoded by lacZ gene, luciferase and others identifiable to those skilled in the art.
  • exemplary reporters comprise a fluorescent protein which is a mutant variant of GFP referred to as‘gfpmut3’ [56].
  • the UC Fi-biosensors comprising U-sensitive genetic molecular components and/or U-sensitive genetic circuits described herein are configured to produce a U- neutralizing molecular component in presence of U.
  • U-neutralizing molecular component refers to any component capable of decreasing the bioavailable U concentration.
  • the U-neutralizing molecular component is a U-neutralizing genetic molecular component comprising a“U-neutralizing gene”.
  • the term“U-neutralizing genetic molecular component” as used herein refers to a genetic molecular component in which the gene of the genetic molecular component is a U-neutralizing gene and wherein polynucleotide-encoded RNA and/or proteins have U-neutralizing characteristics identifiable by those skilled in the art upon reading of the present disclosure, such as proteins having enzymatic functions capable of allowing bioreduction, biomineralization, biosorption, or bioaccumulation of bioavailable U, as described herein.
  • bioreduction refers to altering the redox state of uranium from aqueous U (VI) to insoluble U (IV).
  • VI aqueous U
  • IV insoluble U
  • some bacteria are able to respire different electron acceptors to gain energy for metabolism.
  • anoxia progresses, the most energetically favorable electron acceptors are used in sequence, starting with the reduction of nitrate, then proceeding through Mn(IV), Fe(III) and sulfate, and finally the reduction of carbon dioxide to produce methane.
  • U(VI) has a similar redox couple to Fe(III), and natively Fe(III)-reducing bacteria are able to respire U(VI) as an alternative electron acceptor, reducing it to insoluble U(IV) [12].
  • Other groups natively capable of U(VI) reduction comprise bacteria such as sulfate- reducing bacteria [57], fermentative bacteria [58], acid-tolerant bacteria [59] and myxobacteria [60].
  • the host cell is preferably selected among cells natively expressing the components required to perform uranium bioreduction, which can be engineered to include one or more U-sensing genetic molecular component and/or other components of the U-sensing genetic circuit herein described.
  • a host cell can be E.Coli or other facultative anaerobe genetically engineered to include one or more U-sensing genetic molecular component and/or other components of the U-sensing genetic circuit herein described as well as genetic molecular components required to perform U-bioreduction as will be understood by a skilled person
  • uranium bioreduction can be used as a bioremediation technique, stimulated by adding an electron donor to promote enzymatic reduction of aqueous U(VI) to insoluble U(IV) [61-66].
  • Enzymatic reduction of U(VI) can be catalyzed using U-neutralizing genes such as those expressing cytochrome c [57, 67, 68].
  • chelators can be used to solubilize U(VI) and/or electron shuttles to mediate extracellular electron transfer, such as U-neutralizing genes expressing flavin mononucleotide or riboflavin [11, 69-72].
  • exemplary U-neutralizing genes comprise cytochrome c genes, flavin mononucleotide genes, or riboflavin genes which can be natively expressed in suitable host possibly further engineered to include one or more U-sensing genetic molecular components and/or U-sensing genetic circuit herein described.
  • biomineralization and“bioprecipitation” as used herein refers to a process by which metals precipitate with microbially generated ligands such as sulfide or phosphate, or as carbonates or hydroxides in response to localized alkaline conditions at the cell surface. Thereafter, the uranium precipitate can be removed.
  • the U can be comprised into a stable mineral that sediment out and not be re leached over time.
  • U-removal can be performed by some form of on-site filtration identifiable by a skilled person.
  • sequestration of uranium as insoluble uranyl U(VI) phosphate biominerals can be used for in situ biomineralization for sites where bioreduction may not be feasible due to high nitrate concentrations or where there is risk of reoxidation reoccurring, e.g., in sites comprising carbonate [73, 74].
  • the UC Fi-biosensors described herein can be engineered to catalyze precipitation of uranium such as uranyl phosphates.
  • bacteria can be engineered to precipitate uranyl phosphates [75] by expressing U-neutralizing genes such as acid-phosphatase genes [76] or alkaline-phosphatase genes [77] or phytase genes.
  • U-neutralizing genes such as acid-phosphatase genes [76] or alkaline-phosphatase genes [77] or phytase genes.
  • an exogenous source of phosphate can be added such as glycerol phosphate [16] or tributylphosphate [78].
  • exemplary U-neutralizing genes comprise acid-phosphatase genes, or alkaline-phosphatase genes.
  • the U-neutralizing gene is phoY, encoding an alkaline phosphatase that has been shown to allow the coupling of release of inorganic phosphorus (Pi) from organophosphates with U-Pi precipitation in Caulobacter crescentus [54], or phytase, that can be used to liberate phosphate from phytate, an environmental source of phosphate (see Example 10).
  • Pphyt or P1361 can be used to drive expression of an alkaline phosphatase.
  • the U-biosensors described herein can be engineered to comprise additional genetic molecular components configured to express one or more genes encoding proteins having enzymatic functions to catalyze release of inorganic phosphate from organophosphates (via hydrolytic cleavage catalyzed by phosphatases), inorganic phosphite (via enzymatic oxidation) and phosphonates (via cleavage catalyzed by C-P lyases), or from nucleic acids [79], phytate [80] or phospholipids [81].
  • exemplary U-neutralizing genes comprise ppk genes.
  • exemplary U-neutralizing genes comprise genes encoding proteins configured to bind U to the cell surface of the U biosensor.
  • the U-neutralizing gene is an ompA-SUP fusion gene encoding a rationally engineered super uranyl binding protein (SUP) having femtomolar affinity [86] (see Example 10).
  • the encoded ompA fusion is configured to anchor SUP to the outer membrane, allowing adsorption of U to the cell surface.
  • the U-neutralizing gene is a fusion gene comprising the ompA protein fused to a Calmodulin EF-Hand Peptide (CaM) [87] that has been engineered for high U selectivity.
  • the U-neutralizing gene is a fusion gene comprising the Calmodulin EF-Hand Peptides (CaM) or SUP fused with the rsaA (S-layer) gene (see Example 10), for example using the method outline in Nomellini [88].
  • proteobacteria can be engineered to provide U biosensors comprising U-neutralizing components configured to produce a U biosorption output, following a methodology such as has been used in Caulobacter for rare earth element adsorption [89].
  • UCEFi-biosensors comprising U-neutralizing molecular components described herein can be used in several embodiments for U bioremediation.
  • bioremediation refers to a waste management technique that involves the use of organisms to neutralize pollutants from a contaminated site. Bioremediation can be performed in situ or ex situ. In situ bioremediation involves treating the contaminated material at the site, while ex situ involves the removal of the contaminated material to be treated elsewhere.
  • uranium in the environment depends on its speciation and redox state (e.g., see Figure 15). It is present as mobile U(VI) in oxidizing conditions, predominantly as the uranyl ion (U0 2 2+ ) or hydroxyl complexes below -pH 6.5, or as uranyl carbonate complexes at higher pH [90]. In the absence of carbonate, the uranyl ion and its complexes sorb strongly onto the surface of iron oxides and organics [73, 91, 92] and onto the edge sites of clay minerals [93, 94].
  • Biogeochemical interactions play a key role in controlling the speciation and mobility of uranium, through direct metabolic processes such as microbial respiration, or indirectly by changing ambient redox/pH conditions, producing ligands or new biominerals, or altering mineral surfaces.
  • these biogeochemical processes can be stimulated to accelerate clean-up of contaminated environments through bioremediation.
  • the preferred U-neutralizing molecular component output is a U-neutralizing molecular component having a U biomineralization function.
  • the reporter gene or U-neutralization gene is contiguous with the U-sensitive promoter, wherein the 5’ end of the reporter gene is immediately adjacent to the 3’ end of the U-sensitive promoter.
  • the reporter gene is not contiguous with the promoter, such that one or more nucleotides are located between the 3’ end of the U-sensitive promoter and the 5’ end of the reporter gene.
  • a ribosome binding site can be inserted between the 3’ end of the U-sensitive promoter (downstream of the TSS) and the 5’ end of the reporter gene.
  • the U biosensors described herein comprise any non-pathogenic member of Caulobacteridae.
  • the U biosensor comprises Caulobacteridae such as C. crescentus strains NA1000, CB15, and OR37, an environmental isolate from a U-contaminated site that exhibits high heavy metal tolerance [97].
  • an exemplary host organism is C. crescentus strain NA1000.
  • the cell can be any Caulobacteridae having a genome that natively comprises promoters having 1362 binding sites, such as exemplary promoters P 1361 or Pphyt, or a homolog thereof.
  • the U biosensor herein described is further engineered to include an F-sensing riboswitch.
  • riboswitch in the sense of the disclosure indicates a regulatory segment of a messenger RNA molecule that is configured to bind a target compound and to provide upon binding with the target compound a change in production of the proteins encoded by the mRNA. Accordingly, a riboswitch sense concentrations of the target compound. [98].
  • Riboswitches in the sense of the disclosure comprise an aptamer domain which is configured to specifically and selectively bind the target compound and an expression platform domain configured for genetic control of the mRNA expression.
  • a riboswitch is configured so that in absence of the target compound the aptamer domain and the expression platform domain are configured to inhibit the expression of the mRNA.
  • the riboswitch changes configuration and in the expression platform domain’s inhibition of the mRNA expression is removed, thus providing in metabolite-dependent allosteric control of gene expression.
  • aptamer domains are typically configured to form a stem structure in absence of the target compound.
  • the hybridizing strands forming the stem are referred to as the aptamer strand and the control strand which are configured to complementarily bind to each other.
  • the stem structure which is either formed or be disrupted in absence of target compound is an aptamer domain-control strand stem structure.
  • expression platform domains In a riboswitch, expression platform domains generally have at least a portion identified as regulated strand configured to complementarily bind with the control strand of a linked aptamer domain to form a control strand-regulated strand structure, which can be a control strand-regulated strand stem. This control strand regulated strand structure will either form or be disrupted upon binding of the target compound.
  • control strand of the aptamer domain can complementarily bind to with the aptamer strand and the regulated strand to form alternative stem structures with the aptamer strand and the regulated strand of the expression platform domain depending on the presence of the target domain.
  • nucleic acids indicates the two nucleotides on opposite polynucleotide strands or sequences that are connected via hydrogen bonds.
  • adenine (A) forms a base pair with thymine (T) and guanine (G) forms a base pair with cytosine (C).
  • adenine (A) forms a base pair with uracil (U) and guanine (G) forms a base pair with cytosine (C).
  • base pairing indicates formation of hydrogen bonds between base pairs on opposite complementary polynucleotide strands or sequences following the Watson-Crick base pairing rule as will be applied by a skilled person to provide duplex polynucleotides. Accordingly, when two polynucleotide strands, sequences or segments are noted to be binding to each other through complementarily binding or complementarily bind to each other, this indicate that a sufficient number of bases pairs forms between the two strands, sequences or segments to form a thermodynamically stable double- stranded duplex, although the duplex can contain mismatches, bulges and/or wobble base pairs as will be understood by a skilled person.
  • complementary binding between the aptamer strand and the control strand is more or less thermodynamically stable than complementary base paring between the aptamer strand and other sequences of the riboswitch, and complementary binding between the control strand with the regulated strand of the riboswitch, depending on the presence of the target compound.
  • thermodynamic stability indicates a lowest energy state of a chemical system.
  • Thermodynamic stability can be used in connection with description of two chemical entities (e.g. two molecules or portions thereof) to compare the relative energies of the chemical entities.
  • a chemical entity is a polynucleotide
  • thermodynamic stability can be used in absolute terms to indicate a conformation that is at a lowest energy state, or in relative terms to describe conformations of the polynucleotide or portions thereof to identify the prevailing conformation as a result of the prevailing conformation being in a lower energy state.
  • Thermodynamic stability can be detected using methods and techniques identifiable by a skilled person.
  • thermodynamic stability can be determined based on measurement of melting temperature T m , among other methods, wherein a higher T m can be associated with a more thermodynamically stable chemical entity as will be understood by a skilled person.
  • Contributors to thermodynamic stability can include, but are not limited to, chemical compositions, base compositions, neighboring chemical compositions, and geometry of the chemical entity.
  • control strand-regulated strand structure affects expression of the RNA molecule containing the riboswitch.
  • the stem structure generally either is, or prevents formation of, an expression regulatory structure.
  • An expression regulatory structure is a structure that allows, prevents, enhances or inhibits expression of an RNA molecule containing the structure. Examples include Shine-Dalgarno sequences, initiation codons, transcription terminators, and stability and processing signals, such as splice sites and sequences.
  • Riboswitches in the sense of the disclosure can be naturally occurring, isolated and recombinant riboswitches.
  • Microbes, in particular, bacteria and archea have evolved riboswitches to selectively detect over a dozen small molecules/metabolites (e.g., purine nucleobases) [99], several of which have been exploited for the construction of whole-cell biosensors [100].
  • Riboswitches in the sense of the disclosure can be comprised in sequences encoding proteins or peptides of interest, including reporter proteins or peptides, which can naturally occurring or synthetic as well as endogenous or heterologous to the host cell.
  • F-sensing riboswitch indicates a riboswitch for which Fluoride is the target compound. Accordingly, F-sensing riboswitches in the sense of the disclosure function as riboswitches that sense fluoride ions. These F-sensing riboswitches increase expression of downstream genes when fluoride levels are elevated, and the downstream genes can modulate the toxic effects of high levels of fluoride.
  • An F-sensing riboswitch comprises a fluoride aptamer domain which is configured to specifically and selectively bind fluoride and an expression platform domain configured for genetic control of the mRNA expression.
  • An F-sensing riboswitch in accordance with the disclosure is configured so that in the absence of fluoride the fluoride aptamer domain and the expression platform domain are configured to inhibit the expression of the mRNA.
  • the F-sensing riboswitch changes configuration and the inhibition of the mRNA expression imposed by the expression platform domain is thus removed.
  • the fluoride aptamer is expected to bind fluoride anions with a dissociation constant of between 50 mM and 60 pM, inclusive and to be able to detect an amount of fluoride in the environment, which can be calculated accordingly.
  • the dissociation constant is particularly useful for determining the amount of fluoride detectable in the environment if the fluoride export capability of the host is abolished (when fluoride is retained inside the cell rather than pumped back out) as will be understood by a skilled person.
  • Fluoride detection capability of specific riboswitches is expected to be identified using approaches such as the ones exemplified in Example 20 and Example 22 as will be understood by a skilled person.
  • the host genome can be searched to identify presence or absence of a native fluoride riboswitch such as a CrcB or EricF homolog. If a native fluoride riboswitch is present, the detection limit for a biosensor of the disclosure is expected to be close to 1 mM. In those instances, deletion of the fluoride transporter will likely enable a 100-fold decrease in the detection limit.
  • host cell not expressing a native fluoride riboswitch or with a deleted native riboswitch is expected to be able to detect about lOpM or higher. It is expected that levels such as 50 pM and 60 pM fluoride will not substantially affect the viability of cells lacking fluoride efflux activity and that toxicity for a host will require mM levels of Fluoride depending on the host as will be understood by a skilled person.
  • fluoride aptamer domains or“fluoride aptamers” indicate nucleic acid molecules that can specifically bind to fluoride ions.
  • the fluoride aptamer domain is typically configured in a stem structure formed by the fluoride aptamer strand complementarily binding the control strand of the fluoride sensing riboswitch or in an alternative stem structure with the fluoride regulated strand of the expression platform domain or other sequences designed to complementarity bind the fluoride control strand depending on the presence of absence or fluoride according to the riboswitch design.
  • Fluoride aptamers typically are configured to form a stem structure in absence of fluoride.
  • Fluoride aptamers include nucleic acid molecules that bind fluoride as the anion alone or fluoride with a counterion.
  • Fluoride aptamers generally can be naturally occurring fluoride aptamers, such as fluoride aptamers in naturally-occurring F-sensing riboswitches, and fluoride aptamers derived from naturally-occurring fluoride aptamers.
  • R indicates a purine (A or G).
  • Y indicates a pyrimidine (C or U).
  • N indicates any nucleotide
  • W indicates a weak nucleotide (A or U).
  • M indicates an amino nucleotide (A or C).
  • This sequence like any other riboswitch, is encoded by a corresponding DNA sequence wherein U is replaced by T as will be understood by a skilled person. This sequence further provides an indication of the conserved nucleotides in naturally occurring Fluoride sensing riboswitches as will be understood by a skilled person.
  • Exemplary F-sensing riboswitches in the sense of the disclosure comprise a fluoride- sensing riboswitch called a‘ crcB motif [102, 103].
  • crcB motif RNAs are typically located upstream of genes encoding proteins of diverse functions and presumably regulate these genes. Some of the gene products are annotated as ion transporters (for example, chloride, sodium, proton) and some others are involved in various physiological (e.g., universal stress adaptation, DNA repair) or metabolic (e.g., enolase, formate-hydrogen lyase) processes.
  • the crcB riboswitch in particular, enables a mechanism for sensing and detoxifying environmental fluoride by coupling fluoride binding [104] with the activation of genes encoding enzymes that mitigate fluoride toxicity, most commonly a fluoride exporter (CrcB) that expels internal fluoride anions [102, 105] (see Example 26).
  • a fluoride exporter CrcB
  • a key feature of the crcB motif is the high fluoride selectivity; binding has not been observed for other halides (including concentrations of chloride up to 2.5 M), small anions, gases and 36 other relevant cellular metabolites [102].
  • crcB motif can be applied toward the environmental detection of fluoride, and ultimately, function as an integral detection component within a whole-cell UO2F2 compound sensor. .
  • Figure 48A shows a consensus sequence and structural model based on the comparison of 2188 representatives from bacterial and archaeal species.
  • PI, P2, P3 and pseudoknot labels of FIG. 1A of US 9580713 [102] identify base-paired substructures.
  • a pseudoknot is a nucleic acid secondary structure containing at least two stem-loop structures in which half of one stem is intercalated between the two halves of another stem as will be understood by a person skilled in the art.
  • Exemplary F-sensing riboswitches with a consensus sequence and structure schematically illustrated in Figure 48A has sequence
  • Ni N4 N10 N15 Nis Nis and N36 are independently any amino acid
  • Ni N 4 Nio Nis Nis Nis and N36 are independently present or absent; anyone of N 1 G 2 G 3 N 4 Rs is linked with any one of Y i4 N 15 C i6 Cn Nis y by a pseudoknot
  • a 6 is linked with Lkg by a pseudoknot
  • a first insertion segment of variable length is located between Nis and A19, the first insertion segment comprising a stem loop structure of variable length and having sequence starting with NR and ending with YN in a 5’ to 3’ direction
  • a second insertion segment of variable length is located between Y25 and R26; second insertion segment configured to form a stem loop structure R is A or G; and
  • Y is C or U as also indicated in Figure 48A
  • any one of nucleotides G 3 , R5, U7, Gs , R9, Y 14 , Ci6, A20, A31, Y37 , C38 , U39 , and R40 is a 97% conserved nucleotide (present in 97% of the F-sensing riboswitches encompassed by SEQ ID NO: 2510)
  • any one of nucleotides G 2 , G 3 , R5, A ⁇ , Lb, Gs, R9, Y 14, Ci6 , Ci7, Ai9 , A20, G23, C24, G27, U29, A3i,Y37, C38, U39, R40 and Y 41 is a 90% conserved nucleotide (present in 90% of the F-sensing riboswitches encompassed by SEQ ID NO: 2510)
  • any one of nucleotides Gn, U12 , Yi3 , C21, C22, Y25, R26, C28, G3O,G33, A34, and C35 is a 75% conserved nucleotide (present in 75% of the F-sensing riboswitches encompassed by SEQ ID NO: 2510W)
  • FIG. 48B reproduces FIG. IB of US 9580713 [102] and shows an exemplary sequence and secondary structure model for the WT 78 Psy RNA comprising a crcB motif sequence.
  • the 78-nucleotide RNA encompasses the crcB motif from Pseudomonas syringae.
  • exemplary crcB motifs are described in (WO 2011/088076 and in [108]) also describing related structure and locations in organisms. Comparative genomics reveals 104 candidate structured RNAs from bacteria, archaeal, and their metagenomes as described in WO 2011/088076 and in [108] which are both hereby incorporated by reference in their entirety. In particular, incorporated by reference in its entirety is section 24 of the Additional File 3 of [108], which shows the sequences of numerous crcB motifs, as well as related genes, organisms, alignments, and consensus structures that can be used in connection with the biosensors herein described.
  • the crcB RNA motif sequences herein described is a crcB RNA motif of the RF01734 family annotated by Rfam database which includes 2138 crcB RNA motif sequences of this family, their secondary and 3-D structures, and predicted phylogenetic tree for the sequence alignment (see http://rfam.org/family/RF01734 at the date of filing of the present disclosure which is also incorporated herein by reference in its entirety).
  • Ni N 2 Ns N 11 Ni 5 N 18 Ni 9 and N 37 are independently any amino acid
  • Ni N 2 Ns N 11 Ni 5 N 18 Ni 9 and N 37 are independently present or absent; anyone of N 1 N 2 G 3 G 4 N 5 G 6 A 7 can be linked with any one of C 15 N 16 C 16 C 17 N 18 N 19 by a pseudoknot of sequence NNCCNC a first insertion segment of 0 to 48 nucleotides is located between N 19 and A 20 , wherein when the first insertion segment is > 12 nucleotides long, the first insertion comprises a stem loop structure of variable length having sequence starting with NRR and ending with YYN in a 5’ to 3’ direction
  • a second insertion segment of variable length structure is located between Y26 and R26 , the second insertion segment configured to form a stem loop ;
  • R is A or G
  • Y is C or U
  • any one of Ni N2 N5 Nn N15 N18 and N37 is a 97% conserved nucleotide (present in 97% of the F-sensing riboswitches encompassed by sequence SEQ ID NO: 2512), and N19 is a 50% conserved nucleotide (present in 50% of the riboswitches encompassed by sequence SEQ ID NO: 2512) ( Figure 48C)
  • exemplary F-sensing riboswitch sequence in the sense of the disclosure can have sequence
  • Ni, N2, N3, N4, Ns, N ⁇ , N7, NS, N9, N 12, Nis, N23, N26, N27, N28, N36, N37, N38, N39, N40, N41, N42, N53, N59, N ⁇ o, N ⁇ , N ⁇ 2, N 63 , N ⁇ 4 are interdependently any amino acid ;
  • R is A or G
  • Y is C or U.
  • each of nucleotides Ni, N2, N3, N4, Ns, N ⁇ , N7, Ns, N9, G11N12, A14 U15 Gi6, Nis, N23 , C24, N26 N28 , A29 A30 ,N37 N38 U46 A48 U49 N53 Y54 C55 U56 Css N59 N ⁇ o N ⁇ N ⁇ 2 N 63 N 6 is a 97% conserved nucleotide (conserved in 97% of the F-sensing riboswitches) each of nucleotides G10 , Y21, C25, G33 N39 N40 R57 is a 90% conserved nucleotide (conserved in 90% of the F-sensing riboswitches) each of nucleotides Go Gn G19 C22 U20 C31 C32 C34 Y35 R43 G44 C
  • F-sensing riboswitches of sequence SEQ ID NO: 2513 and Figure 48D the positions of the riboswitch Ni, N2, N3, N4, Ns, N ⁇ , N7, Ns, N9, G11N12, A14 U15 G1 ⁇ 2, NIS, N23 , C24, N26 N28 , A29 A30 ,N37 N38 U46 A48 U49 N53 Y54 C55 U56 Css N59 N ⁇ o N ⁇ N ⁇ 2 N ⁇ 3 N 6 [are at least 97% conserved nucleotides, positions G10 , Y21, C25, G33 N39 N40 R57 are at least 90% conserved nucleotides, positions G13 Gn G19 C22 U20 C31 C32 C34 Y35 R43 G44 C45 G47 G50 A51 C 52 are at least 75% conserved nucleotides, and positions N 27 N 36 N 41 N 42 are at least 50% conserved nucleotides.
  • the crcB RNA motif of the RF01734 family can have a dissociation constant (K ⁇ ) of ⁇ 60 mM with respect to binding of fluoride.
  • Exemplary F-sensing riboswitches in the sense of the disclosure also comprise a fluoride sensing riboswitch that can be found in regulatory regions of a fluoride efflux pump called Eric F ’ a F _ /H + antiporter), performing an function ( F- efflux) of CrcB efflux pump that is commonly associated with fluoride riboswitches. CrcB and Eric F are expected to carryout the same function - fluoride export. It seems that organisms have one or the other.
  • Additional exemplary Fluoride sensing riboswitches are sequences from SEQ ID NO: 205 to SEQ ID NO; 1988 and SEQ ID No: 1999 to SEQ ID NO: 2231 reported in Appendix I and SEQ ID NO: 1989-1998 and SEQ ID NO: 2232-2508 in Figure 53 of the disclosure which are incorporated herein by reference in its entirety, which contains 2017 exemplary F-sensing riboswitch sequences.
  • FIG. 53 An alignment of 287 sequences from the 2017 exemplary F-sensing riboswitch sequences is shown in Figure 53 (SEQ ID NO: 1989-1998 and SEQ ID NO: 2232- 2508) wherein the nucleotide sequences shown are sequences forming a same structure in corresponding aligned sequences and the symbols“—” indicates gaps between the aligned sequences shown (see Example 20) is also for further guidance concerning sequences and structures of F-sensing riboswitches suitable in constructs in accordance with the present disclosure.
  • the F-sensing riboswitch is the F-sensing riboswitch from Sphingomonas sp, 67-36 encoded by
  • the F-sensing riboswitch is the F-sensing riboswitch from Pseudomonas Syringae encoded by
  • the F-sensing riboswitch is MM-1 from Sphingomonas encoded by sequence
  • an F-sensing riboswitch can be comprised within any one of the U-sensing genetic molecular components herein described.
  • an F-sensing riboswitch can be comprised in an F-sensing molecular component which is a genetically engineered polynucleotide construct configured to regulate expression of one or more RNA and/or protein-encoding genes through one or more promoter in combination with one or more F-sensing riboswitches.
  • an F-sensing riboswitch can be comprised in an F-sensing genetic reportable molecular component such as the one exemplified in Example 21.
  • any fluoride sensing riboswitch is expected to be functional in the host bacterium of interest.
  • the well-characterized fluoride riboswitches from P. syringae DC3000 and Bacillus subtilis [102] represent exemplary riboswitches for fluoride sensor construction.
  • the native fluoride riboswitch from the host bacterium (if present) or a closely related bacterium can be used in the biosensors herein described, particularly if the promoter associated with the fluoride riboswitch is to be employed.
  • Group 1 From the species of interest (e.g., Caulobacter crescentus)
  • species of interest e.g., Caulobacter crescentus
  • Group 2 From the genus of interest (e.g., Caulobacter)
  • Group 3 From the family of interest (e.g., Caulobacteraceae)
  • Group 4 From the order of interest (e.g. Caulobacterales)
  • Group 5 From the subclass of interest (e.g., Caulobacteridae)
  • Group 6 From the class of interest (e.g. Alphaproteobacteria)
  • “Closely related” bacteria encompass bacteria within a same subclass of interest (group 5) preferably within a same order of interest, more preferably within a same family of interest (Group 2) and more preferably within the same genus/species of interest (Group I).
  • C. crescentus a preferred host strain for the U sensor, lacks a native fluoride riboswitch, the uncharacterized fluoride riboswitch from Sphingomonas sp. MM-1 are also expected to be usable.
  • This bacterium is closely related to C. crescentus and possesses similarly high genomic G+C content, minimizing compatibility risks with C. crescentus.
  • the function of the Sphingomonas sp. crcB motif in fluoride sensing is supported by its homology with characterized crcB riboswitches [Pfam database [109]] and its genomic location upstream of a fluoride exporter.
  • the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch. In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch. In some of those embodiments, the endogenous F-sensing riboswitch, are preferably knocked out or disabled through mutagenesis of conserved nucleotides critical for riboswitch function. [102] (see Example 26)
  • the detection limit of the colorimetric crcB reporter ( ⁇ 1 mM) can be adversely affected.
  • abolishing fluoride export by deletion of the crcB gene yielded a 100-fold improvement in the fluoride detection limit (sub 10 mM) [102]. Achieving a low fluoride detection limit will likely require deletion of the crcB gene or the analogous eriC F gene in the host organism.
  • RBS promoter and ribosome binding sequence
  • Any promoter and Ribosome Binding Sequence from a bacteria from the above Groups 1 to Group 6 which is encompassed by a known consensus and has a location upstream of crcB/ericF can be used in connection with each riboswitch sequence from all groups, with preference for promoters from bacteria closely related to the host bacteria of the UO2F2- biosensors herein described.
  • An F-sensing riboswitch can be included in any F-sensing molecular component herein described in combination with promoter and RBS in a configuration identifiable by a skilled person.
  • Reference is made in this connection to the schematic of Figure 52 which shows a schematic representation of the primary genetic parts involved in the construction of a fluoride sensing reporter include a (2) promoter to initiate transcription, a (1) fluoride riboswitch to terminate transcription in the absence of fluoride, a (3) ribosome binding site (RBS) to initiate translation, and a (5) reporter for detection.
  • a portion of the (4) native crcB or eric F gene may be required for riboswitch function (i.e., coupling fluoride binding with transcriptional termination) of some fluoride riboswitches.
  • the crcB gene is not required for the function of the Sphingomonas MM-1 fluoride riboswitch (Example 28).
  • a skilled person will be able to identify a combination and configuration of an F sensing riboswitch based on the host bacteria by selecting the promoter and ribosome binding sequence (RBS) that are functional, and preferably optimized, for the host, such as a promoter and RBS native to the selected host bacterial cell or the closely related bacteria.
  • RBS promoter and ribosome binding sequence
  • a biosensor can be engineered in the preferred host strain for the U sensor C. crescentus , which lacks a native fluoride riboswitch with the uncharacterized fluoride riboswitch from Sphingomonas sp. MM-1.
  • Sphingomonas sp is closely related to C. crescentus and possesses similarly high genomic G+C content, minimizing compatibility risks with C. crescentus.
  • a suitable F-sensing riboswitch and the related configuration to obtain a desired F-sensitivity in a selected host bacteria can be identified with a method wherein the F-sensing riboswitch, promoter and RBS are tested in the host bacteria to identify a functional F-sensing construct to be used in F-sensing genetic molecular components herein described.
  • the method comprises selecting an F-sensing riboswitch from the host bacteria or from a closely related bacteria; selecting a promoter native to the selected host bacterial cell or the closely related bacteria , selecting a ribosome binding sequence (RBS) native to the selected host bacteria or the closely related bacteria, and selecting a portion of the crcBleric F coding region.
  • F-sensing riboswitch from the host bacteria or from a closely related bacteria
  • selecting a promoter native to the selected host bacterial cell or the closely related bacteria selecting a ribosome binding sequence (RBS) native to the selected host bacteria or the closely related bacteria, and selecting a portion of the crcBleric F coding region.
  • RBS ribosome binding sequence
  • the method further comprises providing a candidate F-sensing construct wherein the selected F- sensing riboswitch, the selected native promoter, the selected RBS, and the selected portion of the crcBleric F coding region are included in an gene expression cassette together with a reporter in a candidate configuration allowing expression of the reporter in the host bacteria in presence of Fluoride (see configuration schematically illustrated in Figure 52)
  • the method further comprises introducing the candidate F-sensing construct in the host bacteria for a time and under condition to allow expression of the reporter; and detecting expression of the reporter in presence of a selected F amount to identify an F-sensitivity of the candidate F-sensing construct.
  • the method can further comprise performing the method with additional candidate F-sensing constructs and selecting the candidate F-sensing construct having the desired F-sensitivity.
  • testing of different F-sensing riboswitches in F- sensing construct can be performed to identify the combination of F-sensing riboswitches promoter and RSB sequence, number of riboswitches, length of the construct, presence of spacers and additional structural features of the configuration of the construct, resulting in an F-sensing cassette with a desired F sensitivity to the U-biosensor over additional combinations with less desired or no F-sensitivity (see e.g. testing of Examples 27-33).
  • an F-sensing construct in C. crescentus preferably comprises the Sphingomonas MM-1 fluoride riboswitch module which outperformed an analogous module from the more distantly related strain Pseudomonas syringae in terms of both signal amplitude and dynamic range (see Example 27 and Example 28).
  • an F-sensing in C. crescentus comprises the fluoride riboswitch from the closely related Sphingomonas 67-36 in combination with a xylose-inducible promoter (see Example 27 and Example 28).
  • an F-sensing construct can comprise two or more F-sensing riboswitches (see Example 29).
  • the F-sensing riboswitch can be comprised in an F-sensing construct in combination with the Pxyl promoter.
  • the F-sensing riboswitch can be anyone of the F-riboswitches from Sphingomonas SP., in particular MM-1 (see e.g. SEQ ID NO: 2514) and Sphingomonas 67-36 (see e.g. SEQ ID NO: 2525).
  • the F-sensing riboswitch can also be anyone of the F-riboswitches from Pseudomonas Syringae.
  • the F-sensing construct can have sequences SEQ ID NO: 2528, SEQ ID NO: 2519 or SEQ ID NO: 2520, preferably SEQ ID N02518 (see e.g. Example 28).
  • the F-sensing riboswitch can be comprised in an F-sensing construct in combination with a promoter native to the host bacteria, In those embodiments, the F-sensing riboswitch can preferably be an F-sensing riboswitch MM-lfrom Sphingomonas Sp. (see e.g. SEQ ID NO: 2514).
  • the F-sensing riboswitch can be comprised in an F-sensing construct in combination with Pphyt promoter.
  • the F-sensing riboswitch can preferably be an F-sensing riboswitch MM-lfrom Sphingomonas Sp. (see e.g. SEQ ID NO: 2514).
  • an F-sensing riboswitch is comprised together with other elements such promoters, RBS, Shine Dalgamo sequences and others identifiable by a skilled person in F-sensing constructs in a configuration that has a high signal amplitude (e.g. 25,000 - 200,000 for a normalized fluorescence signal)and large dynamic range with respect to fluorescence (higher than 100, preferably higher then 1000 and more preferably higher than 10,000)).
  • a high signal amplitude e.g. 25,000 - 200,000 for a normalized fluorescence signal
  • large dynamic range with respect to fluorescence higher than 100, preferably higher then 1000 and more preferably higher than 10,000
  • the wording“signal amplitude” as used herein with respect to a construct, cassette or component herein described, indicates the difference between the maximum (concentration of analyte such as U or F- that yields the highest signal) and minimum (no analyte, in particular no U and/or F) signal produced by the construct, cassette or component herein described
  • a high signal amplitude is typically a normalized fluorescence signal of 25,000 - 200,000, wherein the wording normalized fluorescence indicates the fluorescence signal (in arbitrary units) /cell density (OD600).
  • the wording“dynamic range” as used herein with respect to a construct, cassette or component herein described, indicates is defined as the ratio between the maximum (concentration of analyte such as U and/or F- that yields highest signal) and minimum (no analyte, in particular no U and/or F) signal produced by the construct, cassette or component herein described.
  • a high value for the dynamic range are 100, 1000, or 10,000, wherein a higher dynamic range indicates a better analyte detection as will be understood by a skilled person.
  • At least one Fluoride sensing riboswitch herein described can be comprised in combination of at least one a U-sensing genetic molecular component herein described, in a U-sensitive F-sensitive genetic circuit together with a reporter molecular component and/or a U-neutralizing molecular component.
  • a Fluoride sensing riboswitch can be comprised in an F-sensitive genetic circuit together with a reporter molecular component, the F-sensing genetic circuit to be comprise in a biosensor herein described in combination with a U-sensitive genetic circuit and/or component as will be understood by a skilled person upon reading of the present disclosure.
  • the U-sensing genetic molecular component, the reporter molecular component, and/or the U-neutralizing molecular components as well as possibly other components are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components.
  • an F-sensing genetic molecular component comprising a fluoride sensing riboswitch and a reporter molecular component (as well as possibly other components) are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components when the genetic circuit operates according to the circuit design in presence of bioavailable Fluoride.
  • At least one molecular component is a U-sensing genetic molecular component in which a U-sensitive promoter having a regulator direct repeat sequence of SEQ ID NO: 1 or any of SEQ ID NO:7-28) is activated or repressed in presence of bioavailable U.
  • at least one molecular component is a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design in presence of bioavailable U.
  • At least one of the genetic molecular components of the U-sensitive genetic circuit herein described comprises a Fluoride sensing riboswitch
  • the genetic circuit operates according to the circuit design in presence of bioavailable U and in presence of bioavailable Fluoride.
  • the term “genetic circuit” as used herein indicates a collection of molecular components connected one to another by biochemical reactions according to a circuit design.
  • the molecular components are connected one to another by the biochemical reactions so that the collection of molecular components is capable to provide a specific output in response to one or more inputs.
  • the molecular components forming parts of the genetic circuit can be genetic molecular components or cellular molecular components.
  • cellular molecular component indicates a molecular component not encoded by a gene, or indicates a molecular component transcribed and/or translated by a gene but comprised in the circuit without the corresponding gene.
  • exemplary cellular components comprise polynucleotides, polypeptides, polysaccharides, small molecules and additional chemical compounds that are present in a cellular environment and are identifiable by a skilled person.
  • Polysaccharides, small molecules, and additional chemical compounds can include, for example, NAD, FAD, ATP, GTP, CTP, TTP, AMP, GMP, ADP, GDP, Vitamin Bl, B 12, citric acid, glucose, pyruvate, 3-phosphoglyceric acid, phosphoenolpymvate, amino acids, PEG-8000, FiColl 400, spermidine, DTT, b-mercaptoethanol maltose, maltodextrin, fructose, HEPES, Tris- Cl, acetic acid, aTc, IPTG, 30C12HSL, 30C6HSL, vanillin, malachite green, Spinach, succinate, tryptophan, and others known to those skilled in the art.
  • Polynucleotides can include RNA regulatory factors (small activating RNA, small interfering RNA), or“junk” decoy DNA that either saturates DNA-binding enzymes (such as exonuclease) or contains operator sites to sequester activator or repressor enzymes present in the system (for example, as in [110]).
  • Polypeptides can include those present in the genetic circuit but not produced by genetic components in the circuit, or those added to affect the molecular components of the circuit.
  • one or more molecular components is a recombinant molecular component that can be provided by genetic recombination (such as molecular cloning) and/or chemical synthesis to bring together molecules or related portions from multiple sources, thus creating molecular components that would not otherwise be found in a single source.
  • a genetic circuit comprises at least one genetic molecular component or at least two genetic molecular components, and possibly one or more cellular molecular components, connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components.
  • activating refers to a reaction involving the molecular component which results in an increased presence of the molecular component in the cellular environment.
  • activation of a genetic molecular component indicates one or more reactions involving the gene, RNA and/or protein of the genetic molecular component resulting in an increased presence of the gene, RNA and/or protein of the genetic molecular component (e.g. by increased expression of the gene of the molecular component, and/or an increased translation of the RNA).
  • An example of“activating” described herein comprises the initiation of expression of a gene regulated by a UzcRS -regulated promoter by a UzcR protein (e.g., see Example 2).
  • Activation of a molecular component of a genetic circuit by another molecular component of the circuit can be performed by direct or indirect reaction of the molecular components.
  • Examples of a direct activation of a genetic molecular component comprised in a circuit the production of an alternate sigma factor (molecular component of the circuit) that drives the expression of a gene controlled by the alternate sigma factor promoter (other molecular component of the circuit), or the production of a small ribonucleic acid (molecular component of the circuit) that increases expression of a riboregulator-controlled RNA (molecular component of the circuit).
  • an alternate sigma factor molecular component of the circuit
  • a small ribonucleic acid molecular component of the circuit
  • a riboregulator-controlled RNA molecular component of the circuit
  • Examples of indirect activation of a genetic molecular component comprise the production of a first protein that inhibits an intermediate transcriptional repressor protein, wherein the intermediate transcriptional repressor protein represses the production of a target gene, such that the first protein indirectly activates expression of the target gene.
  • inhibitorting refers to a reaction involving the molecular component of the genetic circuit and resulting in a decreased presence of the molecular component in the cellular environment.
  • inhibition of a genetic molecular component indicates one or more reactions involving the gene, RNA and/or protein of the genetic molecular component resulting in a decreased presence of the gene, RNA and/or protein (e.g. by decreased expression of the gene of the molecular component, and/or a decreased translation of the RNA).
  • Inhibition of a cellular molecular component indicates one or more reactions resulting in a decreased production or increased conversion, sequestration or degradation of the cellular molecular components (e.g. a polysaccharide or a metabolite) in the cellular environment.
  • a cellular molecular component e.g. a polysaccharide or a metabolite
  • Inhibition can be performed in the genetic circuit by direct reaction of a molecular component of the genetic circuit with another molecular component of the circuit or indirectly by reaction of products of a reaction of the molecular components of the genetic circuit with another molecular component of the circuit.
  • binding refers to the connecting or uniting two or more molecular components of the circuit by a bond, link, force or tie in order to keep two or more molecular components together, which encompasses either direct or indirect binding where, for example, a first molecular component is directly bound to a second molecular component, or one or more intermediate molecules are disposed between the first molecular component and the second molecular component another molecular component of the circuit.
  • bonds comprise covalent bond, ionic bond, van der Waals interactions and other bonds identifiable by a skilled person.
  • the binding can be direct, such as the production of a polypeptide scaffold that directly binds to a scaffold-binding element of a protein.
  • the binding may be indirect, such as the co-localization of multiple protein elements on one scaffold.
  • binding of a molecular component with another molecular component can result in sequestering the molecular component, thus providing a type of inhibition of said molecular component.
  • binding of a molecular component with another molecular component can change the activity or function of the molecular component, as in the case of allosteric interactions between proteins, thus providing a type of activation or inhibition of the bound component.
  • An example of“binding” as described herein comprises the binding of UzcR to an m_5 site in a UzcRS -regulated promoter (e.g., see Example 2).
  • converting refers to the direct or indirect conversion of the molecular component into another molecular component.
  • An example of this is the conversion of chemical X by protein A to chemical Y that is then further converted by protein B to chemical Z.
  • An example of “converting” as described herein comprises the cleavage of o-nitrophenyl-P-D-galactoside (ONPG) by beta-galactosidase encoded by the lacZ gene (e.g. see Example 1).
  • the molecular components are connected one with another according to a circuit design in which a molecular component is an input and another molecular component is an output.
  • a genetic circuit typically has one or more input or start molecular component which activates, inhibits, binds and/or convert another molecular component, one or more output or end molecular component which are activated, inhibited, bound and/or converted by another molecular ⁇ component, and intermediary molecular components each inhibiting, binding and/or converting another molecular component and being activated, inhibited, bound and/or converted by another molecular component.
  • the input is bioavailable U and/or bioavailable F and the output is a reportable molecular component and/or a U -neutralizing molecular component.
  • the U-sensitive and/or F-sensitive genetic circuits can be comprised together within a same biosensor to detect or to detect and neutralize bioavailable UO 2 F 2 .
  • the U-sensitive and/or F-sensitive genetic circuits can be comprised in combination with a U-sensing genetic molecular component and/or a F-sensing genetic molecular component of the disclosure to the detect or to detect and neutralize bioavailable UO 2 F 2 .
  • the U-sensitive F- sensitive genetic circuit can comprise at least one F-sensing genetic molecular component in which RNA is expressed in presence of bioavailable F, and at least one U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U-sensitive transcriptional 1362 binding site or a U-sensitive transcriptional UzcR binding site.
  • a U-sensing genetic molecular component comprises at least one Fluoride sensing riboswitch.
  • At least one molecular component is a reportable molecular component and/or a U- neutralizing molecular component when the genetic circuit operates according to the circuit design in presence of bioavailable U and bioavailable Fluoride.
  • the UO 2 F 2 comprise a U-sensing genetic molecular component and/or a U-sensing genetic circuit in combination with a separate F-sensing genetic molecular component and/or a F-sensing genetic circuit to provide a dual output configuration wherein one or more fluoride riboswitches control the expression of a reportable molecular component configured to be independent of U.
  • Exemplary dual output configurations of the UO 2 F 2 biosensor herein described are showing in Example 25.
  • the UO 2 F 2 biosensor comprise a U-sensing genetic molecular component and/or a U-sensing genetic circuit in combination with a F-sensing riboswitch integrated in the U-sensing genetic molecular component and/or a U-sensing genetic circuit to provide a single output configuration wherein one or more fluoride riboswitches control the output of the U-sensing genetic molecular component and/or the U-sensing genetic circuit.
  • the fluoride riboswitch functions to prematurely terminate transcription from a U-activated promoter in the absence of fluoride. Fluoride binding mitigates this termination, enabling gene expression thus providing an output in presence of bioavailable U and Fluoride.
  • Exemplary single output configurations of the UO 2 F 2 biosensor herein described are shown in Example 24.
  • an exemplary genetic circuit described herein comprises, a U sensing genetic molecular component in which a U-sensitive promoter such as P phyt or a P 1361 is configured to initiate expression of a lacZ gene encoding the beta-galactosidase enzyme (U- sensing genetic molecular component), wherein the beta-galactosidase enzyme converts the substrate ONPG (cellular molecular component) to yield galactose and o-nitrophenol which has a yellow color (reportable molecular component).
  • the genetic molecular components can further comprise a Fluoride sensing riboswitch (e.g.
  • the UO 2 F 2 biosensor can further comprise the F-sensing riboswitch in a separate F-sensing reportable genetic molecular component and/or a separate F-sensing genetic circuit to provide a dual output configuration.
  • the U-sensitive genetic circuit comprises at least one U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U- sensitive transcriptional 1362 binding site and the U sensitive genetic circuit is comprised in the biosensor together with an F-sensitive reportable genetic molecular component comprising an F- sensitive riboswitch herein described configured to be expressed in presence of an effective amount of bioavailable Fluoride in a dual output configuration.
  • the U-sensitive genetic circuit in the U- sensitive genetic circuit at least one molecular component is a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design in presence of bioavailable U as will be understood by a skilled person.
  • the U-sensitive genetic circuit further comprises at least one genetic molecular component in which a UzcRS two-component system regulated promoter is activated or repressed in presence of bioavailable U.
  • At least one U-sensitive genetic molecular component comprises a U-sensitive promoter such as P phyt or a Pi36i configured to initiate expression of a uzcS gene (CCNA_02842) and a uzcR gene (CCNZ_02485), encoding proteins UzcS and UzcR, respectively (first U-sensing genetic molecular component) and further comprises a second U-sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of a reporter gene (e.g.
  • a U-sensitive promoter such as P phyt or a Pi36i configured to initiate expression of a uzcS gene (CCNA_02842) and a uzcR gene (CCNZ_02485), encoding proteins UzcS and UzcR, respectively (first U-sensing genetic molecular component) and further comprises a second U-sensing genetic molecular component in which
  • GFP an exemplary reportable molecular component
  • binding of UzcR protein to the UzcRS -regulated promoter activates the UzcRS -regulated promoter (see Example 2).
  • uzcS and uzcR are genes that are natively comprised in a uzcRS operon in Caulobacter and other alphaproteobacteria, as described in Park et ah, 2017 [3].
  • uzcRS operon refers to the genetically encoded UzcRS two- component system, comprising the uzcS gene (CCNA_02842) and the uzcR gene (CCNZ_02485), and operatively linked promoters and regulatory elements [3].
  • uzcR and uzcS are physically separated by genes parD3 and parE3 encoding the ParDE3 toxin antitoxin (TA) system, together forming a putative four-gene operon [3, 112].
  • uzcR and uzcS are conserved throughout alphaproteobacteria, the insertion of parDE3 between uzcR and uzcS is unique to a subset of the Caulobacter genus; uzcR and uzcS are adjacently located in the majority of closely related alphaproteobacteria [3] including C. crescentus OR37, an environmental isolate from a U-contaminated site [97].
  • the U biosensor can be any genetically engineered proteobacteria and in particular a genetically engineered Caulobacteridae which comprises a UzcRS two component system.
  • a U- sensitive genetic circuit further comprises one or more genetic molecular components comprising one or more negative regulators of UzcRS that function to maintain UzcRS in an OFF state in absence of metal (see Example 9).
  • the UO2F2 biosensor described herein comprises a chromosomal copy of UzcRS negative regulators 1 and 2 (Example 9) under transcriptional regulation of their native promoters.
  • one or more genetic molecular components comprising exemplary UzcRS negative regulators 1 and 2 are placed under control of inducible transcriptional regulatory elements in order to desensitize UzcS to a particular signal.
  • overexpression of CCNA_03680-CCNA_03681 reduces the sensitivity of UzcRS for Zn and Cu.
  • the MarR-type regulators CCNA_03498 and/or CCNA_02289 can be deleted in the host Caulobacteridae genome to increase sensitivity of uzcRS for U.
  • a fluoride riboswitch can be integrated between the UzcRS -regulated promoter and the GFP gene and/or between the P phyt or P1361 promoter and the uzcS ATG to provide a U- sensing F-sensing genetic circuit in a single output configuration.
  • a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
  • riboswitch performance is expected to be adversely affected by changes in genomic context, such as fusing the riboswitch to a reporter gene. Accordingly, in any embodiments herein described, if an initial transcriptional fusion to a reporter (such as GFP) fails, a Ribo-attenuator, a recently developed genetic element that insulates the riboswitch from genomic context and enhances the tunability (e.g., sensitivity, cell-to-cell variability, etc.), will be used to build a riboswitch reporter (e.g. a CrcB reporter).
  • a reporter such as GFP
  • the U-sensitive genetic circuit comprises a U-sensitive promoter such as P phyt or a P i % i configured to initiate expression of a hrpS gene encoding an HrpS protein (first U-sensing genetic molecular component) and further comprises a second U-sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of an hrpR gene encoding an HrpR protein.
  • a U-sensitive promoter such as P phyt or a P i % i configured to initiate expression of a hrpS gene encoding an HrpS protein (first U-sensing genetic molecular component) and further comprises a second U-sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of an hrpR gene encoding an HrpR protein.
  • the U-sensitive F-sensitive genetic circuit further comprises a fourth genetic molecular component comprising a hrpL promoter (P hrpL ) configured to initiate expression of a reporter gene (e.g. GFP, an exemplary reportable molecular component), in which binding of both HrpS and HrpR are required for s 54 dependent activation of the P hrpL and expression of HrpS or HrpR alone is not sufficient for transcriptional activation (see Example 3).
  • P hrpL hrpL promoter
  • GFP an exemplary reportable molecular component
  • a fluoride riboswitch could be integrated in three positions 1) Between the P phyt or P i % i promoter and hrpS, 2); 2) between the UzcRS- regulated promoter and the hrpR; and/or between the PhrpL promoter and gfp to provide a U- sensitive F-sensitive genetic circuit with a single output configuration.
  • a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
  • hrpR and hrpS refer to genes that are natively comprised in a s 54 dependent hrpR/hrpS hetero-regulation module from the hrp (hypersensitive response and pathogenicity) system for Type III secretion in Psuedomonas syringae [114-116], wherein both HrpS and HrpR are required for s 54 dependent activation of the hrpL promoter (P hi L ) and expression of HrpS or HrpR alone is not sufficient for transcriptional activation.
  • At least two genetic molecular components comprise complementary protein fragments of a transcription factor, which are configured to associate together to form a functional transcription factor, such as those based on a bacterial two-hybrid system.
  • the term“bacterial two-hybrid” as used herein refers to a technique used to detect protein-protein interactions and protein-DNA interactions by testing for physical interactions (such as binding) between two proteins or a single protein and a DNA molecule, respectively.
  • the bacterial two-hybrid system relies on the activation of downstream reporter gene(s) upon binding of a transcription factor onto an upstream activating sequence (UAS), wherein the transcription factor is split into two separate fragments, called the binding domain (BD) and activating domain (AD).
  • UAS upstream activating sequence
  • BD binding domain
  • AD activating domain
  • the bacterial two-hybrid system is a protein-fragment complementation assay that requires both the BD and the AD for reporter expression.
  • An exemplary bacterial two-hybrid system utilizes an E. coli omega protein, which copurifies with RNA polymerase, and can function as a transcriptional activator when linked covalently to a DNA-binding protein.
  • the E. coli omega protein can function as an activation target when this covalent linkage is replaced by a pair of interacting polypeptides fused to the DNA-binding protein and to omega, respectively [117].
  • the U-sensitive genetic circuit comprises a U-sensitive promoter such as P phyt or a P i % i configured to initiate expression of a BD gene encoding an BD protein (first U- sensing genetic molecular component), and further comprises a second U- sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of an AD gene encoding an AD protein.
  • the U-sensitive genetic circuit further comprises a fourth genetic molecular component comprising a promoter comprising binding sites for the BD and the AD, configured to activate expression of the reporter gene, e.g.
  • GFP or the U- neutralizing gene in which binding of both BD and AD are required activation of the third genetic molecular component and expression of the reportable molecular component and/or the U-neutralizing molecular component (see Example 5).
  • a 1363/1362 regulated promoter e.g., P phyt
  • a UzcRS regulated promoter e.g., P UTCB
  • k-gal4 an exemplary BD
  • P iacOR2-62 that contains the UAS to initiate expression of a reporter gene, such as GFPmut3 (see Example 5).
  • a fluoride riboswitch could be integrated in three positions: 1) Between the P phyt or P i % i promoter and BD; 2) Between the UzcRS -regulated promoter and AD and/or 3) ) Between the BD/AD-regulated promoter and gfp to provide a U-sensitive F-sensitive genetic circuit with a single output configuration.
  • a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
  • the U-sensitive genetic circuit comprises at least one U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U-sensitive transcriptional 1362 binding site and/or a UzcR binding site together.
  • the U-sensitive genetic circuit at least one molecular component is a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design in presence of bioavailable U, wherein the reportable genetic component and/or a U-neutralizing molecular component is formed by an assembly of two or more subunits of the reportable molecular component and/or a U-neutralizing molecular component.
  • the U-sensitive genetic circuit further preferably comprises at least one genetic molecular component in which an UzcRS two-component system regulated promoter is activated or repressed in the presence of bioavailable U and bioavailable F.
  • a Fluoride riboswitch could be integrated in two positions: 1) between the P phyt or P1361 promoter and U neutralizing gen; and/or 2) Between the UzcRS -regulated promoter and U neutralizing gene to provide a U-sensitive F- sensitive genetic circuit with a single output configuration.
  • a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
  • the U-sensitive F-sensitive genetic circuit comprises a U-sensitive promoter such as Pp hyt or a Pi36i configured to initiate expression of a gfplO-Kl fusion gene encoding a GFP10-K1 fusion protein (first U-sensing genetic molecular component) and further comprises a second U- sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of an El-gfpll gene encoding E1-GFP11 fusion protein.
  • a U-sensitive promoter such as Pp hyt or a Pi36i configured to initiate expression of a gfplO-Kl fusion gene encoding a GFP10-K1 fusion protein (first U-sensing genetic molecular component) and further comprises a second U- sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of an El-gfpll gene
  • the U-sensitive genetic circuit further comprises a third genetic molecular component comprising a gfpl-9 gene regulated by a non-U-responsive, xylose inducible promoter (P xyi ), a constitutively active promoter (e.g., P rSaA ) or by P phyt /Pi36i .
  • P xyi non-U-responsive, xylose inducible promoter
  • P rSaA constitutively active promoter
  • P phyt /Pi36i e.g., P rSaA
  • a Fluoride riboswitch can be integrated in three positions: 1) between the P phyt or P i % i promoter and gfplO-Kl 2) between the UzcRS -regulated promoter and El-gfpll and/or 3) Between constitutive promoter and gfp-1-9 to provide a U-sensitive F-sensitive genetic circuit with a single output configuration.
  • a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
  • tripartite GFP refers to a GFP reporter that requires expression and assembly of GFP10, GFP11 and GFP1-9 together for GFP reporter function [5].
  • tripartite GFP assembly and reporter function is based on tripartite association between two twenty amino-acids long GFP tags, GFP10 and GFP11, which are fused to interacting protein partners, in addition to a complementary GFP 1-9 detector.
  • GFP 10 and GFP11 self-associate with GFP1-9 to form a functional GFP [5].
  • any protein interaction pair can be used for the tripartite system.
  • Exemplary interacting protein partners comprise oppositely charged Kl/El coiled coils, FKBP12-FRB rapamycin inducible protein interaction [5], or the leucine zipper of GCN4 [118] among others known to those skilled in the art.
  • a U-sensitive genetic circuit can comprise at least one U- sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in the presence of bioavailable U, the U sensitive promoter comprising a U-sensitive transcriptional 1362 binding site and/or a UzcR binding site.
  • the U-sensitive genetic circuit at least one molecular component is a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design in the presence of bioavailable U, wherein the reportable molecular component and/or the U-neutralizing molecular component is post-transcriptionally and/or post-translationally converted by the U-sensitive F-sensitive genetic circuit in presence of bioavailable U.
  • the U-sensitive genetic circuit further preferably comprises at least one genetic molecular component in which a UzcRS two-component system regulated promoter is activated or repressed in presence of bioavailable U and bioavailable F.
  • the U-sensitive genetic molecular component comprises a U- sensitive promoter such as P phyt or a P i % i configured to initiate expression of a protease configured to cleave at a cleavage sequence comprised in a linker peptide in a Forster resonance energy transfer (FRET) sensor protein (first U-sensing genetic molecular component), and further comprises a second U-sensing genetic molecular component in which a UzcRS two- component system regulated promoter is configured to initiate expression of the FRET sensor protein.
  • a U- sensitive promoter such as P phyt or a P i % i configured to initiate expression of a protease configured to cleave at a cleavage sequence comprised in a linker peptide in a Forster resonance energy transfer (FRET) sensor protein (first U-sensing genetic molecular component)
  • FRET forster resonance energy transfer
  • the U-sensitive genetic circuit is configured to express and cleave the FRET sensor protein.
  • a fluoride sensing riboswitch can be integrated in two positions: 1) between the P phyt or Pi36i promoter and FRET protein 1 and/or 2) between the UzcRS -regulated promoter and FRET protein 2 to provide a U- sensitive F-sensitive genetic circuit with a single output configuration.
  • a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
  • the terms“Forster resonance energy transfer”, “FRET”, “fluorescence resonance energy transfer”, “resonance energy transfer”, “RET” or “electronic energy transfer”, and“EET” as used herein refers to a mechanism describing energy transfer between two light-sensitive molecules (chromophores) [119].
  • a donor chromophore initially in its electronic excited state, can transfer energy to an acceptor chromophore through nonradiative dipole-dipole coupling [120].
  • the efficiency of this energy transfer is inversely proportional to the sixth power of the distance between donor and acceptor, making FRET extremely sensitive to small changes in distance.
  • FRET efficiency can be used to determine if two fluorophores are within a certain distance of each other [121]. Such measurements can be used as a research tool in fields such as biology and chemistry.
  • fluorophores for biological use is a cyan fluorescent protein (CFP) - yellow fluorescent protein (YFP) pair [122]. Both are color variants of green fluorescent protein (GFP).
  • CFP cyan fluorescent protein
  • YFP yellow fluorescent protein
  • GFP green fluorescent protein
  • a genetically-encoded fusion of CFP and YFP covalently linked by a protease cleavage sequence can be used as a cleavage assay, wherein if the linker is intact, excitation at the absorbance wavelength of CFP (414nm) causes emission by YFP (525nm) due to FRET. If the linker is cleaved by a protease, FRET is abolished and emission is at the CFP wavelength (475nm) [123].
  • a UOiFibioscnsor herein described comprises two or more U- sensitive genetic molecular components and/or U-sensitive genetic circuits, wherein each of the U-sensitive genetic molecular components expresses a different reporter gene and/or a U- neutralizing gene, and/or each of the U-sensitive genetic circuits comprise a different reportable molecular component and/or U-neutralizing molecular component, in presence of bioavailable U.
  • exemplary different reportable molecular components can comprise a first genetically-encoded reporter (e.g., GFP) and a second, different genetically-encoded reporter (e.g., dsRED).
  • a Fluoride riboswitch can be integrated in two positions: 1) between the P phyt or P i % i promoter and reportable molecular component 1 and/or U-neutralizing molecular component 1 ; and/or 2) between the UzcRS -regulated promoter and reportable molecular component 2 and/or U-neutralizing molecular component 2 to provide a U-sensitive F- sensitive genetic circuit with a single output configuration.
  • a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
  • the U0 2 F 2 biosensors described herein can comprise a UzcRS U- sensitive genetic molecular component comprising a reporter gene or a U-neutralizing gene operatively connected to a UzcRS-regulated promoter, such as PurcA, PurcB, and others identifiable by those skilled in the art, such as those described in Park et ah, (2017) [3].
  • the UzcRS U-sensitive genetic molecular component can be comprised in the biosensor alone or in combination with a 1363/1362 U sensitive genetic molecular component comprising a reporter gene or a U neutralizing gene operatively connected to a 1362 regulated promoter.
  • a Fluoride riboswitch can be integrated in two positions: 1) Between the P p h y t or P1361 promoter and reportable molecular component 1 and/or U-neutralizing molecular component; and/or 2) Between the UzcRS-regulated promoter and reportable molecular component 2 and/or U-neutralizing molecular component 2.
  • one or more genetic molecular components of the U-sensitive F-sensitive genetic circuits described herein can comprise genomic DNA of the proteobacterial cell and in particular of the Caulobacteridae cell.
  • the one or more genetic molecular components can comprise native genomic DNA in the Caulobacteridae cell or can be introduced into the genome of the Caulobacteridae cell through genetic engineering, or comprised in the Caulobacteridae cell in one or more extra-genomic polynucleotides or vectors, using standard genetic engineering methods known to those skilled in the art and described herein.
  • the U0 2 F 2 biosensors-can detect uranium in a range dependent on the composition of the growth media since media components influence bioavailability.
  • the bioavailability of both compounds will be dictated by the composition of the growth media / the chemical composition of the environmental sample
  • a U-sensing genetic molecular component and/or U sensing genetic circuits herein described comprising UzcR binding site and a UzcRS TCS can detect Uranium in ⁇ 1 micromolar concentrations as will be understood by a skilled person upon reading of the present disclosure.
  • a U-sensing genetic molecular component and/or U sensing genetic circuits herein described comprising 1362 binding site and a 1363/1362 TCS can detect Uranium at -500 nM in aqueous conditions lacking Pi or glycerol-phosphate as will be understood by a skilled person upon reading of the present disclosure.
  • the biosensor U sensitivity is expected to be about the same and therefore can detect Uranium at -500 nM in aqueous conditions lacking Pi or glycerol-phosphate as will be understood by a skilled person upon reading of the present disclosure
  • UOiFibioscnsors herein described comprising a U-sensitive and/or F-sensitive genetic circuit comprising one or more F sensing riboswitches activated in presence of bioavailable F, one or more U-sensing genetic molecular components in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U-sensitive transcriptional 1362 binding site and/or a UczR binding site, in a first target range of bioavailable U the endogenous proteobacteria U-sensitive transcriptional regulator is not bound to the 1362 binding site and/or the UczR binding site of the U-sensitive promoter and the U-sensitive F-sensitive genetic circuit does not comprise a reportable molecular component and/or a U-neutralizing molecular component, when the genetic circuit operates according to the circuit design.
  • the endogenous proteobacteria U-sensitive transcriptional regulator is bound to the 1362 binding site and/or to the UzcR binding site of the U-sensitive promoter and the U-sensitive genetic circuit comprises a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design.
  • the fluoride riboswitch will not bind F and transcription will be prematurely terminated.
  • the fluoride riboswitch will bind F and transcription initiation will occur, resulting in activation of the reportable element and the single output of the circuit.
  • a U-sensitive genetic circuit which comprises a 1362 binding site further comprises at least one genetic molecular component comprising a UzcRS- regulated promoter comprising a UzcR binding site.
  • activation or repression of the UzcRS -regulated promoter is also required as an input for the U-sensitive genetic circuit to comprise a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design, wherein activation or repression of the U-sensitive promoter comprising the regulator direct repeat together with activation or repression of the UzcRS- regulated promoter is herein referred to as an“AND gate”, wherein the term“AND” is an operation of Boolean logic.
  • Boolean logic is a branch of algebra in which the values of the variables are the truth values ‘true’ and ‘false’, usually denoted by the digital logic terms‘G and ⁇ ’ respectively.
  • the main operations of Boolean logic are the conjunction‘AND’, the disjunction OR’, and the negation‘NOT’.
  • AND gate refers to a digital logic gate that implements logical conjunction - it behaves according to the truth table shown in Table 1.
  • A‘true’ output (1) results only if both the inputs to the AND gate are‘true’ (1). If neither or only one input to the AND gate is‘true’ (1), a‘false’ (0) output results. Therefore, the output is always 0 except when all the inputs are 1.
  • the term“AND gate” as used herein refers to the logical relation between two genetic molecular components in a U-sensitive genetic circuit, wherein inputs‘A’ and ⁇ ’ in Table 1 are two independently activated or repressed genetic molecular components, wherein a first independently activated or repressed genetic molecular component comprises a first promoter having a U-sensitive transcriptional 1362 binding site, and a second independently activated or repressed genetic molecular component comprises a UzcRS-two-component system regulated promoter, and the output ‘A AND B’ in Table 1 is the reportable molecular component and/or a U-neutralizing molecular component of the U-sensitive genetic circuit.
  • any‘AND gate’ genetic system can be employed in the U-sensitive genetic circuits described herein, such as those described in [124, 125]
  • the U-sensitive genetic circuits described herein comprise U- sensing genetic molecular components whose expression is regulated independently by (1) a U- sensitive promoter comprising a 1362 binding site, such as P phyt or P i % i and (2) UzcRS two- component system, and comprise an AND gate wherein‘inputs’ of activation or repression of both (1) a U-sensitive promoter comprising a 1362 binding site, such as P phyt or P i % i and (2) UzcRS two-component system-regulated promoter is required for the reportable molecular component‘output’ and/or the U-neutralizing molecular component‘output’ according to the U- sensitive genetic circuit design (Figure 3).
  • an F-sensing component and/or a U-sensing genetic molecular component, independently activated or repressed are arranged‘in series’ in the U-sensitive genetic circuit.
  • output of the reportable molecular component and/or the U-neutralizing molecular component according to the genetic circuit design requires two or more of the independently activated or repressed F-sensing genetic molecular component and U- sensing genetic molecular components to be activated or repressed in temporal succession, wherein the activation or repression of the F-sensing genetic molecular component precedes the activation or representation of a first independently activated or repressed U-sensing genetic molecular component which on its turn can precede the activation or repression of a second independently activated or repressed U-sensing genetic molecular component.
  • the term“in series” as used herein refers to a genetic circuit in which genetic molecular components are connected through biochemical reactions along a single linear circuit path.
  • Table 1 in an ‘in series’ AND gate, the temporal sequence of the activation or repression of the independently activated U-sensing genetic molecular components denoted by inputs ‘A’ and‘B’ is such that the second input‘B’ is dependent on a prior activation or repression of the first input‘A’ in linear succession according to the genetic circuit design.
  • an in- series AND gate comprised in a U- sensitive genetic circuit comprises, a first U-sensing genetic molecular component comprising a U-sensitive promoter having a 1362 binding site, such as a Pi36i promoter or a P p h y t promoter, and optionally further comprises a UzcRS two component system-dependent promoter, wherein expression of a uzcR and uzcS genes are under the transcriptional control of a U-sensitive promoter comprising a 1362 binding site, such as a Pi36i promoter or a P phyt promoter (see Example 2).
  • At least two independently activated or repressed U-sensing genetic molecular components are arranged‘in parallel’ in the U-sensitive genetic circuit, herein referred to as an ‘in parallel’ AND gate.
  • an ‘in parallel’ AND gate With regard to Table 1, in contrast to the‘in series’ AND gate, in an‘in parallel’ AND gate, the temporal sequence of inputs‘A’ and‘B’ is such that a second input‘B’ is not dependent on a prior first input‘A’ in linear succession, but rather inputs‘A’ and‘B’ can occur simultaneously.
  • the term“in parallel” as used herein refers to a genetic circuit in which genetic molecular components are connected through biochemical reactions along more than one circuit path.
  • output of the reportable molecular component and/or the U-neutralizing molecular component according to the genetic circuit design requires two or more of the independently activated or repressed F-sensing genetic molecular component and/or one or more U-sensing genetic molecular components to be activated or repressed in parallel.
  • the in-parallel AND gate comprised in a U- sensitive genetic circuit comprises at least two independently activated or repressed U-sensing genetic molecular components, each comprising a different promoter, (1) a U-sensitive promoter comprising a 1362 direct repeat binding site, such as a Pi36i or Pphyt, and (2) a UzcRS -regulated promoter, wherein promoters (1) and (2) act as independently activated or repressed parallel inputs into the U-sensitive genetic circuit, functioning as two independent points of U-sensing in response to their respective U-sensitive transcriptional regulators.
  • the reportable molecular component output and/or a U-neutralizing molecular component output of the U-sensitive F-sensitive genetic circuit is present only when both of the two independently activated or repressed F-sensing genetic molecular component and U-sensing genetic molecular component are activated or repressed according to the U-sensing genetic circuit design.
  • a U-sensitive F-sensitive genetic circuit comprising an ‘in parallel’ AND gate for more than one U sensing genetic molecular component, can provide improved selectivity for U, as output of the reportable molecular component and/or the U- neutralizing molecular component is dependent on two independent points of U-sensing in response to two different U-sensitive transcriptional regulators.
  • This also reduces the probability of a false positive output, in the first target range of bioavailable U concentration, such as in response to non-U stimuli (such as Zn or Cu).
  • an‘in parallel’ AND gate comprises an HRP AND gate.
  • the term“HRP AND gate” as used herein refers to an AND gate system from Pseudomonas syringae that was developed in E. coli [126].
  • the HRP AND gate system comprises an orthogonal s 54 dependent hrpR/hrpS hetero -regulation module from the hrp (hypersensitive response and pathogenicity) system for Type III secretion in Psuedomonas syringae [114-116], as described above.
  • both HrpS and HrpR are required for s 54 dependent activation of the hrpL promoter (P hrpL ) and expression of HrpS or HrpR alone is not sufficient for transcriptional activation.
  • two different promoters (1) a U-sensitive promoter comprising a 1362 binding site, such as a P 1361 or Pphyt, and (2) a UzcRS -regulated promoter, act as independently activated inputs to initiate the transcription of hrpR and hrpS, respectively, functioning as two independent points of U-sensing in response to their respective U-sensitive transcriptional regulators (Figure 6 Panel A).
  • transcription of the output hrpL promoter is activated only when both proteins HrpR and HrpS bind the upstream activator sequence to remodel a closed o54-RNAP-hrpL transcription complex to an open one through ATP hydrolysis [126].
  • the output shown is GFP reporter expression ( Figure 6 Panel A).
  • the performance of the HRP AND gate can be described using a Hill function for the promoter steady-state input-output response (transfer function) in the form:
  • [/] is the concentration of the inducer, such as bioavailable U
  • Ki and m are the Hill constant and coefficient, respectively, relating to the promoter- regulator/inducer interaction
  • k is the maximum expression level due to induction
  • a is a constant relating to the basal level of the promoter due to leakage [126] and further in the form:
  • Eq. (2) which describes the normalized output of the AND gate as a function of the levels of the two activator proteins ([L’] for HrpR, [5] for HrpS) at steady state.
  • [G] m ax is the maximum activity observed for the output.
  • KR, KS and HR, Tls are the Hill constants and coefficients for HrpR and HrpS, respectively [126].
  • the expression of hrpS is placed under the control of a U- sensitive promoter comprising a 1362 binding site, such as a Pp hyt or Pi36i, while hrpR is placed under the control of P U rcB, a UzcRS -dependent promoter that has lower basal activity compared to P U rcA [3].
  • the PhrpL promoter regulating gfp expression requires Pp hyt /Pi36i and P urc B to be active to generate a fluorescent signal.
  • an‘in parallel’ AND gate comprises a tripartite GFP AND gate.
  • reporter function is based on tripartite association between two twenty amino-acids long GFP tags, GFP10 and GFP11, which are fused to interacting protein partners, in addition to a complementary GFP1-9 detector.
  • GFP10 and GFP11 self-associate with GFP1-9 to form a functional GFP [5].
  • a gfplO-Kl fusion gene is placed under control of a U-sensitive promoter comprising a 1362 direct repeat binding site, such as a P phyt or Pi36i
  • expression of a El-gfpll fusion gene is placed under control of a UzcRS-responsive promoter, such as PurcB
  • expression of gfpl-9 is placed under control of a non-U-responsive, strong, constitutively active promoter, P rS aA ⁇ .
  • an‘in parallel’ AND gate system comprises a bacterial two- hybrid AND gate.
  • an‘in parallel’ AND gate system comprises a FRET sensor AND gate.
  • the U-sensitive genetic circuits described herein comprise a combination of two or more‘in series’ and/or‘in parallel’ AND gates as described herein, wherein the two or more AND gates are connected by activating, inhibiting, binding or converting reactions.
  • Figure 14 shows an exemplary combination of an‘in series’ AND gate and an‘in parallel’ AND gate (see Example 8).
  • the exemplary ‘in series’ AND gate shown in Figure 4 Panel C and the exemplary‘in parallel’ tripartite GFP AND gate shown in Figure 6 Panel B are connected, such that the P phyt -regulated UzcR no longer activates expression of GFP regulated by P1968 as in Figure 4 Panel C, but rather activates expression of El-gfpll regulated by P urcB within the tripartite GFP‘in parallel’ AND gate.
  • circuit contains the native uzcRS genes placed under the control of the Pphyt promoter - the chromosomal PI and P2 promoters are swapped with Pphyt as described herein. This could be done by deleting uzcRS and using a plasmid- based system to reintroduce these genes into the circuit.
  • the circuit leverages the improved selectivity of U over Zn observed in the sensor depicted in Figure 4 Panel C - A U/Zn ratio of 5.5.
  • a Fluoride sensing riboswitch can be included within any one of the genetic molecular components of any one of the U- sensitive genetic circuit.
  • the fluoride riboswitch will not bind F and transcription will be prematurely terminated.
  • the fluoride riboswitch will bind F and transcription initiation will occur, resulting in activation of the reportable element of the circuit and in providing the single output of the circuit.
  • a F sensing riboswitch can be included within a separate F-sensing reportable genetic molecular component and/or F-sensing genetic circuit.
  • the fluoride riboswitch will bind F and transcription initiation will occur, resulting in activation of the separate F-sensing reportable genetic molecular component and/or the separate F-sensing genetic circuit to provide the F-sensing output of the dual output configuration of the UOiFibioscnsors herein described.
  • the U-sensitive and/or F-sensitive genetic circuits described herein comprise one or more genetic molecular components that are orthogonal to the proteobacteria cell of the UO2F2 biosensor.
  • the circuit components of U-sensing and/or F-sensing circuit herein described are stably integrated in the host.
  • the‘in series’ AND gate described herein was constructed using the native uzcR and uzcS genes.
  • the chromosomal PI and P2 promoters have been replaced with Pphyt.
  • the U- sensing genetic molecular component and/or U-sensing genetic circuit can further comprise an amplifier genetic molecular component comprising a U-sensitive promoter and/or a controllable promoter operatively connected to the UzcY and/or UzcZ in a configuration wherein the U- sensitive promoter and/or a controllable promoter directly initiates expression of the amplifier molecular component (see Figure 13).
  • MTRDQDTLRML AE VE A AN ADL ARRAKAPLW YHP ALGLLV G ALIA V QGQPT S ILLVFY A AYIAGLALLVRAYKRHTGLWVSGYRAGRTRWVALGLATLTMIGGVIAVWLLRERGLT A APLIFG AIV A VIVT V GGF VWE
  • a AF RADLRDGRPL SEQ ID NO: 33 or a sequence that when aligned with sequence SEQ ID NO: 33 has a BLAST score has a BLAST score greater than 50 and less than 100, or preferably greater than 100 and less than 200, or more preferably a BLAST score greater than 200 and a homology with SEQ ID NO: 33 less than 100%, or even more preferably BLAST score of 289 and an homology of 100% with SEQ ID NO: 33.
  • the amplifier molecular component acts as a ‘genetic signal amplifier’ configured to increase an output, e.g. expression or levels of a reportable molecular component and/or a U- neutralizing molecular component at a given bioavailable U concentration, thus enabling more sensitive detection and reporting and/or neutralizing of bioavailable U at lower concentrations.
  • the U-sensitive genetic circuit comprises one or more amplifier genetic molecular components comprising uzcY gene CCNA_03497 encoding Caulobacter Crescentus N100 UzcY (SEQ ID NO: 33) and/or uzcZ gene CCNA_02291 encoding for Caulobacter Crescentus N100 UzcZ (SEQ ID NO: 34) under a promoter that can be either a constitutively active promoter or an inducible promoter. In preferred embodiments, the promoter is constitutively active.
  • the uzcY and/or uzcZ gene are under the transcriptional regulation of a promoter comprising a regulator direct repeat, such as exemplary promoters P phyt or Pi36i.
  • the U-sensitive genetic circuit comprises CCNA_03497 placed under the control of P phyt so that U exposure enhances the sensitivity of UzcRS for U (see Example 7).
  • the bacteria are capable of natively expressing endogenous MarR family repressors such as marRl and/or marR2 genes.
  • VPE A (SEQ ID NO: 35) or a sequence that when aligned with sequence SEQ ID NO: 35 has a BLAST score greater than 65 and less than 100, or preferably greater than 100 and less than 150, or more preferably a BLAST score greater thanl50 and a homology with SEQ ID NO: 35 less than 100%, or even more preferably BLAST score of 199 and a homology of 100% with SEQ ID NO: 35.
  • MarR2 indicates a protein having amino acid sequence: M APRFDIS GLDD VIHGRVRLGI V A YLAS AE V ADFTELKD VLE VT QGNLS IHLRKLEE AG YVSIDKSFVGR KPLTRVRLTDTGRAAFSSYLRAMGQLVEQAGGG (SEQ ID NO: 36) or a sequence that when aligned with sequence SEQ ID NO: 36 has a BLAST score has a BLAST score greater than 75 and less than 100, or preferably greater than 100 and less than 150, or more preferably a BLAST score greater than 150 and a homology with SEQ ID NO: 36 less than 100%, or even more preferably BLAST score of 206 and a homology of 100% with SEQ ID NO: 36.
  • U0 2 F 2 -biosensors herein described comprising UzcR binding and a UzcRS TCS
  • the host is capable of natively expressing endogenous MarR family repressors such as marRl and/or marR2 genes
  • at least one gene of the endogenous MarR family is knocked out to provide an amplified U-biosensor configured to provide an amplified signal following activation of the U- sensitive genetic circuit (see Example 7 and Figure 12).
  • uzcY expression is induced (see Figure 26), thus leading to the stimulation of UzcS activity and causing a hypersensitive output in response to low metal concentration.
  • the protein MarRl and MarR2 are the protein encoded by marRl gene CCNA_03498 found in Caulobacter crescentus NA1000 (SEQ ID NO: 35) and protein encoded by marR2 gene CCNA_02289 found in Caulobacter crescentus NA1000 (SEQ ID NO: 36) respectively.
  • the bacteria are capable of natively expressing an UzcRS Regulating Transporter Atpase aminoPeptidase (herein also urtAP).
  • UrtAP is an ATP-binding cassette transporter (ABC transporters) comprising UrtA an ATPase and UrtP a peptidase which contains 13 transmembrane domains and a C-terminal peptidase domain.
  • ABSC transporters ATP-binding cassette transporter
  • the term“UrtA” as used herein indicates a protein encoded by a gene having a protein coding region adjacent in the genome and cotranscribed with urtP, the protein having amino acid sequence:
  • IDRNGDDNT VKVGG (SEQ ID NO: 144) or a sequence that when aligned with sequence SEQ ID NO: 144 has a BLAST score greater than 600 and less than 800, or preferably greater than 800 and less than 1200 or greater than 1000 and less than 1200, or more preferably a BLAST score greater thanl200 and less than 2000, or more preferably a BLAST score equal to or higher than 2000 or even more preferably BLAST score of 2436 and a homology of 100% with UrtP sequence from Caulobacter crescentus NA1000 or Caulobacter crescentus CB 15 and in particular with SEQ ID NO: 144.
  • U0 2 F 2 -biosensors herein described comprising UzcR binding and a UzcRS TCS
  • the host is capable of natively expressing endogenous urtAP
  • at least one gene of the endogenous urtAP preferably all genes of the endogenous urtAP
  • the UzcRS TCS will be stimulated and exhibit greater uranium sensitivity.
  • UrtA and UrtP Genes encoding for UrtA and UrtP herein described are herein also indicated as urtA gene or UrtA and urtP gene or UrtP, respectively as will be understood by a skilled person.
  • Detection of a host capable of expressing endogenous urtAP can be performed by detecting UrtP or urtP as urtA is more conserved as will be understood by a skilled person.
  • the protein UrtA and UrtP are the protein encoded by UrtA gene CCNA_03681 found in Caulobacter crescentus NA1000 (SEQ ID NO: 143) and protein encoded by UrtP gene CCNA_03680 found in Caulobacter crescentus NA1000 (SEQ ID NO: 144) respectively.
  • UO2F2 UO2F2
  • pathogen detection and eradication 131
  • organic pollutant detection 132
  • standoff detection of landmines 133, 134
  • microfluidic chemostat platform that maintains cell viability and sensor function for at least a week [137]
  • optical transducer components that convert cell fluorescence into an electronic signal that is transmitted on a mobile phone network [138]
  • automated water sampling feature [138].
  • Additional exemplary features comprise standoff detection functions [133, 134]and couples sensor cells, which are encapsulated within a polymeric matrix and placed in proximity of a suspicious region, with a novel optical scanning system to enable remote detection at a distance of up to 50 meters in optically noisy aqueous and soil environments [133, 134].
  • a method to provide a U biosensor comprising
  • U-sensing genetic molecular components configured to report and/or neutralize U herein described
  • U biosensors optionally operatively connecting one or more of the U biosensors to an electronic signal transducer adapted to convert a U biosensor reportable molecular component output into an electronic output.
  • Fluoride sensing riboswitch, a promoter comprising the U-sensitive 1362 binding site and/or an UzcR binding site can be genetically engineered by introducing into a polynucleotide comprising a promoter DNA sequence a polynucleotide comprising the Fluoride sensing riboswitch, the U-sensitive 1362 binding site and/or a UzcR binding site, as described herein.
  • a Fluoride sensing riboswitch, a U-sensitive promoter comprising the 1362 binding site and/or a UzcR binding site can be genetically engineered by de novo designing a synthetic promoter DNA sequence comprising the Fluoride sensing riboswitch, the 1362 binding site and/or the UzcR binding site.
  • nucleic acid molecules such as “polynucleotide” and“nucleotide sequence” comprise any polynucleotides such as DNA and RNA molecules and include both single-stranded and double-stranded molecules whether it is natural or synthetic in origin.
  • polynucleotide indicates an organic polymer composed of two or more monomers including nucleotides, or analogs thereof.
  • the isoelectric point of a polynucleotide in the sense of the disclosure is less than 7 as will be understood by a skilled person.
  • the term“nucleotide” refers to any of several compounds that consist of a ribose or deoxyribose sugar joined to a purine or pyrimidine base and to a phosphate group and that is the basic structural unit of nucleic acids.
  • nucleotide analog refers respectively to a nucleotide in which one or more individual atoms have been replaced with a different atom or with a different functional group.
  • polynucleotide includes nucleic acids of any length, and in particular DNA, RNA, analogs and fragments thereof.
  • a polynucleotide of three or more nucleotides is also called “nucleotidic oligomer” or “oligonucleotide”.
  • polynucleotides in the sense of the disclosure comprise biological molecules comprising a plurality of nucleotides.
  • nucleic acids include deoxyribonucleic acids, ribonucleic acids, and synthetic analogues thereof, including peptide nucleic acids.
  • Polynucleotides can typically be provided in single-stranded form or double- stranded form and in liner or circular form as will be understood by a person of ordinary skill in the art.
  • the polynucleotide is a DNA molecule that can be in a linear or circular form, and encodes one or more proteins under the control of a promoter recognizable by an enzyme such as an RNA polymerase, that is capable of transcribing the encoded DNA.
  • the term“protein” as used herein indicates a polypeptide with a particular secondary and tertiary structure that can interact with another molecule and in particular, with other biomolecules including other proteins, DNA, RNA, lipids, metabolites, hormones, chemokines, and/or small molecules.
  • the term“polypeptide” as used herein indicates an organic linear, circular, or branched polymer composed of two or more amino acid monomers and/or analogs thereof.
  • the term“polypeptide” includes amino acid polymers of any length including full- length proteins and peptides, as well as analogs and fragments thereof. A polypeptide of three or more amino acids is also called a protein oligomer, peptide, or oligopeptide.
  • the terms“peptide” and“oligopeptide” usually indicate a polypeptide with less than 100 amino acid monomers.
  • the polypeptide provides the primary structure of the protein, wherein the term“primary structure” of a protein refers to the sequence of amino acids in the polypeptide chain covalently linked to form the polypeptide polymer.
  • a protein “sequence” indicates the order of the amino acids that form the primary structure. Covalent bonds between amino acids within the primary structure can include peptide bonds or disulfide bonds, and additional bonds identifiable by a skilled person.
  • Polypeptides in the sense of the present disclosure are usually composed of a linear chain of alpha-amino acid residues covalently linked by peptide bond or a synthetic covalent linkage.
  • the two ends of the linear polypeptide chain encompassing the terminal residues and the adjacent segment are referred to as the carboxyl terminus (C-terminus) and the amino terminus (N-terminus) based on the nature of the free group on each extremity.
  • counting of residues in a polypeptide is performed from the N-terminal end (Nth-group), which is the end where the amino group is not involved in a peptide bond to the C-terminal end (-COOH group) which is the end where a COOH group is not involved in a peptide bond.
  • Proteins and polypeptides can be identified by x-ray crystallography, direct sequencing, immunoprecipitation, and a variety of other methods as understood by a person skilled in the art. Proteins can be provided in vitro or in vivo by several methods identifiable by a skilled person. In some instances where the proteins are synthetic proteins in at least a portion of the polymer two or more amino acid monomers and/or analogs thereof are joined through chemically-mediated condensation of an organic acid (- COOH) and an amine (-NH2) to form an amide bond or a“peptide” bond.
  • - COOH organic acid
  • -NH2 an amine
  • amino acid refers to organic compounds composed of amine and carboxylic acid functional groups, along with a side-chain specific to each amino acid.
  • alpha- or a- amino acid refers to organic compounds composed of amine (-NH2) and carboxylic acid (-COOH), and a side-chain specific to each amino acid connected to an alpha carbon.
  • Different amino acids have different side chains and have distinctive characteristics, such as charge, polarity, aromaticity, reduction potential, hydrophobicity, and pKa.
  • Amino acids can be covalently linked to form a polymer through peptide bonds by reactions between the amine group of a first amino acid and the carboxylic acid group of a second amino acid.
  • Amino acid in the sense of the disclosure refers to any of the twenty naturally occurring amino acids, non-natural amino acids, and includes both D an L optical isomers.
  • the sequence of a polynucleotide encoding a genetic molecular component described herein can be homologous to the polynucleotide sequence of the genetic molecular component described herein.
  • two polynucleotide (RNA or DNA) sequences are substantially homologous when at least 80% (preferably at least 85% and most preferably at least 90%) of the nucleotides match over the defined length of the sequence using algorithms such as CLUSTAL or PHILIP. Sequences that are substantially homologous can be identified in a polynucleotide hybridization experiment under stringent conditions as is known in the art. See, for example, Sambrook et al. [139].
  • stringent conditions can be adjusted to screen for moderately similar fragments, such as homologous sequences from distantly related organisms, to highly similar fragments, such as genes that duplicate functional enzymes from closely related organisms.
  • stringent conditions are selected to be about 5°C lower than the thermal melting point (Tm) for the specific sequence at a defined ionic strength and pH.
  • Tm thermal melting point
  • stringent conditions encompass temperatures in the range of about 1°C to about 20°C, depending upon the desired degree of stringency as otherwise qualified herein.
  • substantially similar refers to polynucleotides wherein changes in one or more nucleotide bases can result in substitution of one or more amino acids, but do not affect the functional properties of the polypeptide or protein encoded by the nucleotide sequence.
  • substantially similar also refers to modifications of the nucleic acid fragments of the instant disclosure such as deletion or insertion of nucleotides that do not substantially affect the functional properties of the resulting polynucleotide or transcript. It is therefore understood that the disclosure encompasses more than the specific exemplary nucleotide or amino acid sequences and includes functional equivalents thereof. Alterations in a nucleic acid fragment that result in the production of a chemically equivalent amino acid at a given site, but do not affect the functional properties of the encoded polypeptide, are well known in the art.
  • Computer implementations of these mathematical algorithms can be utilized for comparison of sequences to determine sequence identity. Such implementations include, but are not limited to: CLUSTAL in the PC/Gene program (available from Intelligenetics, Mountain View, Calif.); the ALIGN program (Version 2.0) and GAP, BESTFIT, BLAST, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Version 8 (available from Genetics Computer Group (GCG), 575 Science Drive, Madison, Wis., USA). Alignments using these programs can be performed using the default parameters.
  • sequence homology “sequence homology”,“homology”, “sequence identity” or “identity” in the context of two nucleic acid or polypeptide sequences makes reference to the nucleotides or residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window.
  • sequence identity When percentage of sequence identity is used in reference to proteins, it is recognized that residue positions which are not identical often differ by conservative amino acid substitutions, where amino acid residues are substituted for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity) and therefore do not change the functional properties of the molecule.
  • “percentage homology” means the value determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide sequence in the comparison window may comprise additions or deletions (gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison, and multiplying the result by 100 to yield the percentage of sequence identity.
  • reference sequence is a defined sequence used as a basis for sequence comparison.
  • a reference sequence may be a subset or the entirety of a specified sequence; for example, as a part of a full-length cDNA or partial genomic DNA sequence, or the complete cDNA or gene sequence.
  • a reference sequence can comprise, for example, a sequence identifiable a database such as GenBank and others known to those skilled in the art.
  • substantially homology or “substantial identity” of polynucleotide sequences means that a polynucleotide comprises a sequence that has at least 80% sequence identity, preferably at least 85%, more preferably at least 90%, most preferably at least 95% sequence identity compared to a reference sequence using one of the alignment programs described using standard parameters.
  • sequence identity preferably at least 85%, more preferably at least 90%, most preferably at least 95% sequence identity compared to a reference sequence using one of the alignment programs described using standard parameters.
  • Substantial homology or identity of amino acid sequences for these purposes normally means sequence identity of at least 80%, preferably at least 85%, more preferably at least 90%, and most preferably at least 95%.
  • polypeptides and proteins of the disclosure can be altered in various ways including amino acid substitutions, deletions, truncations, and insertions. Novel proteins having properties of interest can be created by combining elements and fragments of proteins of the present disclosure, as well as with other proteins. Methods for such manipulations are generally known in the art.
  • the polynucleotides described herein comprise both the naturally occurring sequences as well as genetically engineered forms.
  • proteins of the disclosure encompass naturally occurring proteins as well as variations and modified forms thereof.
  • polynucleotides of the present disclosure comprise both synthetic molecules and molecules obtained through recombinant DNA techniques known in the art.
  • the bacterial cells described herein can be genetically engineered using methods known to those skilled in the art.
  • the polynucleotides, genetic molecular components and molecular components comprised in vectors described herein can be introduced into the cells using transformation techniques such as electroporation, heat shock, and others known to those skilled in the art and described herein.
  • the U-sensitive genetic molecular components and/or genetic molecular components of the U-sensitive genetic circuits are introduced into the organism to persist as a plasmid or integrate into the genome.
  • the cells can be engineered to chromosomally integrate a polynucleotide comprising one or more U-sensitive genetic molecular components and/or genetic molecular components comprised in the U-sensitive genetic circuits described herein, using methods such as sacB counterselection procedure [147].
  • a polynucleotide comprising one or more U-sensitive genetic molecular components and/or genetic molecular components comprised in the U-sensitive genetic circuits described herein, using methods such as sacB counterselection procedure [147].
  • the U- sensitive genetic molecular components or genetic circuit components are inserted into the Caulobacter chromosome (see Examples).
  • a custom designed low copy vector such as with a pBBRl oriV origin and a cat (chloramphenicol acetyltransferase) gene can be used for reporter expression or to introduce the signal amplifier gene (CCNA_03497) (see Examples).
  • a system comprising a plurality of proteobacterial cells and one or more vectors comprising one or more U-sensitive genetic molecular components and/or one or more genetic molecular components of a U- sensitive genetic circuit.
  • the one or more vectors are configured to introduce one or more U- sensitive genetic molecular components and/or one or more genetic molecular components of a U-sensitive genetic circuit into an alphaproteobacterial cell.
  • vectors comprising a U-sensitive genetic molecular components or genetic molecular components such as promoters or RNA- or protein coding genes described herein or fragments thereof can be engineered using techniques such as In-Fusion cloning and other methods identifiable by those skilled in the art, to generate vectors suitable for genetically engineering the proteobacterial cells described herein.
  • Polynucleotides encoding genetic molecular components such as promoters and genes encoding RNA and proteins described herein can be isolated from genomic DNA or cDNA comprising the polynucleotides of interest, such as polynucleotides isolated from organisms such as Caulobacter or other Caulobacteridae, using standard Polymerase Chain Reaction (PCR)-based methods known in the art. Plasmids comprising reporter genes and/or U-neutralizing genes described herein are commercially available from vendors such as Thermo-Fisher and Clontech, and other sources such as Addgene, among others known to those skilled in the art. Polynucleotides can also be designed and synthesized de novo , such as using gBlock synthesis (IDT Technologies) as described herein.
  • IDT Technologies gBlock synthesis
  • a composition comprising one or more UC Fi-biosensor or vectors herein described together with a suitable vehicle.
  • vehicle indicates any of various media acting usually as solvents, carriers, binders or diluents for the one or more U-sensitive genetic molecular components, vectors, or cells herein described that are comprised in the composition as an active ingredient.
  • the composition including the one or more U-sensitive genetic molecular components, vectors, or cells herein described can be used in one of the methods or systems herein described.
  • a system comprising an electronic signal transducer adapted to convert a UCLFi-biosensor reportable molecular component output into an electronic output.
  • the system comprises an electronic signal transducer and one or more U biosensors herein described operatively connected to the electronic signal transducer.
  • the system comprises an electronic signal transducer and one or more U biosensors herein described comprised in a composition together with a suitable vehicle.
  • the term“electronic signal transducer” as used herein refers to an electronic device typically comprising a bio-recognition component, a biotransducer component, and an electronic system which can comprise a signal amplifier, processor, data display, and data communicator. Transducers and electronics can be combined, such as in CMOS-based microsensor systems [148]
  • the transducer can be exemplarily fashioned with a blue LED (light emitting diode) (ex. 5 mM, emitting excitation light centered at 470 nm wavelength) 4110 that shines excitation light into a cuvette holder 4105.
  • a curvette holder is a containing device for the filters and curvette, opaque so that the only light in is from the LED and the only light striking the photodetector is the filtered emission light).
  • the blue light passes through an excitation filter (ex filter) 4115 with a center wavelength (CWL) at 469 nm (blue) and a full width at half maximum of 35 nm (narrow bandwidth).
  • the cuvette is a transparent container for the sample (cells).
  • the emitted light (emission light) from the sample passes through an emission filter (em filter) 4125 that has a CWL of 525 nm and a FWHM of 39 nm (a narrow bandwidth of green light).
  • This emission light is read by a silicon amplified photodetector 4130 (with switchable gain for sensitivity control), turning the emission light into an electrical signal for analysis.
  • the recognition component can use biomolecules from organisms or receptors modeled after biological systems to interact with the reportable molecular component output comprising a target analyte of interest. This interaction is measured by the biotransducer which outputs a measurable signal proportional to the presence of the target analyte in the sample.
  • a biotransducer is the recognition-transduction component of the device. In some embodiments, it can comprise a bio-recognition layer and a physicochemical transducer, which acting together converts a biochemical signal to an electronic or optical signal.
  • the bio-recognition layer typically can contain an enzyme or another binding protein such as antibody.
  • polynucleotides can also comprise the bio-recognition layer.
  • the physicochemical transducer is typically in contact with the recognition layer.
  • a physico-chemical change is produced within the biorecognition layer that is measured by the physicochemical transducer producing a signal that is proportionate to the concentration of the analyte.
  • the physicochemical transducer can be electrochemical, optical, electronic, gravimetric, pyroelectric or piezoelectric, as understood by those skilled in the art.
  • a quantitative, field-portable UC Fi-biosensor system comprises one or more UC Fi-biosensors described herein coupled with an electronic signal transducer to convert the cellular output reportable molecular component signal (e.g., fluorescence) into an electronic output signal.
  • one or more UC Fi-biosensors described herein are coupled in conjunction with established, inexpensive, commercially available transducer devices known to those skilled in the art, using one of several immobilization methods such as those utilizing carbon nanotubes or nanoparticles to adhere a UC Fi-biosensor host organism cells to the transducer.
  • the holdfast organelle that facilitates irreversible adhesion to surfaces can be used to couple the cells to the electronic signal transducer, eliminating the need for exogenous immobilization substrates.
  • a method of detecting and reporting and/or neutralizing bioavailable UO2F2 comprises:
  • UO2F2 biosensors contacting one or more UO2F2 biosensors, or a system comprising an electronic transducer operatively connected to one or more U0 2 F 2 biosensors, with a target environment comprising one or more target ranges of bioavailable U concentration and bioavailable F for a time and under conditions to detect and report and/or neutralize bioavailable U02F2in the target environment.
  • a method to detect bioavailable U0 2 F 2 with an F-sensing and a U-sensing genetic reportable molecular component and/or with a U sensing and/or F-sensing genetic circuit genetic including a fluorescent label such as GFP is with a fluorometer to quantify GFP or other fluorophore’s fluorescence. This can be accomplished with high sensitivity in the laboratory using a microplate reader or in the field using a mini-fluorometer.
  • the method can comprise adding an environmental sample to a 96-well plate or cuvette containing a U-biosensor herein described.
  • target environment indicates the aggregate of components and related conditions wherein a U biosensor can be operated.
  • the target environment comprises a sample obtained from a field site.
  • the sample is provided by means of an operator, such as a human or a machine, to the host organism optionally operatively connected to the electronic transducer.
  • the sample is provided, in absence of an operator, to the host organism optionally operatively connected to the electronic transducer.
  • the host organism, optionally operatively connected to the electronic transducer can be in situ in a field site comprising the target environment.
  • the host is Caulobacter crescentus NA1000
  • the biosensor is expected to work within a pH range of 6-8, temperature range of ⁇ RT— 37 C. Additional growth nutrients other than inorganic phosphate can be added.
  • the host can be Caulobacter crescentus OR37 strain isolated from the Oak Ridge Field site as an environmentally robust host: including a greater pH, heavy metal, and U tolerance with respect to Caulobacter crescentus NA1000.
  • the reporting of bioavailable UO2F2 can be observed directly, such as by visualizing the output of a reportable molecular component, such as fluorescence of a reportable molecular component, e.g., GFP, wherein the expression and/or function of the reportable molecular component is activated by the U0 2 F 2 biosensor.
  • the reporting of bioavailable UO2F2 can be observed indirectly and/or remotely, such as through transduction of reportable molecular component output into an electronic output, which can be quantified by a computer, and which can optionally be communicated to a location at a distance from the target environment by a data communicator, either through wired or wireless communication.
  • UO2F2 biosensors herein described provide selective and sensitive detection and reporting of UO2F2 which is a bioavailable product of environmental decomposition of UF6, a toxic form of U.
  • the U0 2 F 2 biosensors, and related U-sensitive F-sensitive genetic molecular components, genetic circuits, compositions, methods and systems described herein provide a cost-effective, selective, sensitive, portable, easy to use, high-throughput measurement of bioavailable U, with little or no sample preparation required.
  • U0 2 F 2 biosensors described herein can be used in the construction of consolidated bioremediators comprising bacterial systems that possess all the necessary components for deployment in environmental cleanup efforts.
  • Applications of the U biosensors described herein comprise uses in biodefense (e.g., to be used for non-proliferation purposes), environmental monitoring, and mining (for toxicology and safety concerns), among other uses identifiable by those skilled in the art.
  • Site-directed mutagenesis was performed by amplifying the entire plasmid with the primer sets listed in Table 3. Chromosomal integration and counter selection were performed as described above, and successful substitutions were confirmed by sequencing.
  • the P p h y t -lacZ fusion was integrated at the chromosomal urcA locus in lacA mutant strain JOE2321, yielding strain DMP470.
  • DMP470 was electroporated with YMCS2::Tn5Pvan[l52 ] and plated onto PYE agar plates containing 25 pg ml 1 kanamycin. -12,000 colonies were scraped into a PYE master solution that was frozen and stored at -80 °C.
  • the Transposon library was diluted and spread on M5G-G2P agar containing 40 pg ml 1 Xgal and 25 pM uranyl nitrate, yielding a total of -36,000 colonies. Colonies exhibiting a white colony phenotype were selected and nested semi-arbitrary PCR was used to map the location of each transposon as described previously. [152]
  • Plasmid-borne P Phyt -gfp and Pi36i -gfp fusions were generated by amplifying P phyt and P i % i fragments from the Caulobacter genome with the primer pairs BamHI_Pphyt_F /EcoRI_Pphyt_R and BamHI_P1362/ EcoRI_P1362, respectively, digested with BainH] and CcoRI and cloned into the similarly digested pDMP450, generating pDMP460 and pDMP463.
  • the promoter-g p fusions were shortened and/or mutated by amplifying pDMP460 and pDMP463 with the primer pairs described in Table 3 and re-ligating using infusion cloning.
  • Tripartite GFP AND gate sensor construction A gblock (Tripartite GFP gblock; Table 6) was synthesized (Integrated DNA Technologies, gBlock) with the following components listed in 5’ to 3’ orientation: promoterless gfpl0-m2_kl, rrnBTl and T7Te transcription terminators (BBa_B0015), P UrcB -El-gfpll-M4, lamda To terminator, and gfpl-9 under the control of the rsaA promoter (P rSaA [4]).
  • gfp 10-m2_k 1 was placed under the control of P phyt by digesting the Tripartite GFP gblock with Bglll and Xhol and ligating into the similarly digested pDMP460 to form pDMP791.
  • DNA sequence encompassing the entire tripartite DNA and an insulating upstream rmbTl transcription terminator was amplified with primers HRP_chrome_int F and HRP_chrome_int R and cloned into pDMP82 that was amplified with the primers urcA_loc_DR_F and urcA-loc_UR_R using infusion cloning to form pDMP792.
  • a variant containing GFP10-m2_kl under the control of a shortened version of P p h y t (he., the core sensor) was constructed by amplifying pDMP792 with Pphyt_TR_elim_F and Pphyt_promoter_shorten_R and re-ligating using infusion cloning, forming pDMP883.
  • a variant of this AND gate containing gfpl-9 under the control of the xylose inducible promoter was constructed by amplifying the 360 bp P xyi fragment using Pxyl_F and Pxyl_R and cloning into pDMP792 that were amplified with gfpl-9_amp_for_pxyl_F gfpl- 9_amp_for_pxyl_R using infusion cloning to form pDMP664. All tripartite GFP AND gate variants were then integrated into the chromosomal urcA locus using a two-step sacB counterselection procedure, [155] forming DMP804, DMP895, and DMP683
  • a control tripartite variant containing both GFP10-m2_kl and E1-GFP11-M4 under the control of P mCB was constructed by directionally cloning a P mCB fragment, generated with primers XbaI_PurcB and BglII_PurcB_R and digested with Xbal and Bglll, into the similarly digested pDMP712.
  • a tripartite variant containing both GFP10-m2_kl and E1-GFP11-M4 under the control of UrtAP was constructed by directionally cloning a P nei fragment, generated with primers KpnI_P1362 and AvrII_P1362 and digested with Kpnl and AvrII, into the similarly digested pDMP712.
  • Both control AND gates were cloned into the pNPTS138 double recombination plasmid by amplifying with primers HRP_chrome_int F and R and cloning into pDMP82 that was amplified with the primers urcA_loc_DR_F and urcA-loc_UR_R using Infusion cloning.
  • P xyi promoter in both constructs was swapped for P rsa A by cloning P rsaA , amplified with PrsaA_frag_F and PrsaA_frag_R, into the template vectors that were linearized with the primers gfp_for_PrsaA_F and gfp_for_PrsaA_R, forming pDMP932 and pDMP952.
  • tripartite GFP AND gate variants were then integrated into the chromosomal urcA locus using a two-step sacB counterselection procedure, [155] to form DMP911 and DMP994.
  • a U sensing AND gate strain (DMP993) with constitutive UzcY expression was generated by integrating the core sensor (pDMP 883) into the urcA locus of a strain deleted for CCNA_03498 and CCNA_03499 (DMP863).
  • UzcY expression was restricted to conditions of U exposure by placing uzcY expression under the control of Pphyt-short as follows.
  • uzcY was amplified from the C. crescentus chromosome with primers 3497_for_Pphyt_plas_F and 3497_for_Pphyt_plas_R, digested with Bglll and Xhol, and cloned into the similarly digested pDMP746.
  • the DNA containing the P phyt-short -CCNA_03497 fusion was then amplified with 3497_for_Pphyt_short_chrome_R and Pphyt_short_for_chrom_3497 and cloned into pDMP1113 that was linearized using the primers chrome_3497_vect_amp_F and chrome_vector_amp_R to
  • the resulting suicide vector (form pDMP1114) was used to integrate P phyt-short -nzcF into DMP993 to form DMP1009.
  • IPTG Isopropyl- 1-thio-b-D-galactopyranoside
  • Strep-UrpR was isolated from cell lysates using a Strep-tactin column as described in the manufacturer’s protocol (IB A Lifesciences). The protein concentration of UrpR (reported here as monomers) was determined using a Bradford protein assay (Biorad) with lysozyme as a standard.
  • Electrophoretic mobility shift assays A P phyt promoter fragment containing the region from 13 to 245 with respect to the translation initiation site was amplified from pDMP460 and pDMP475 with primers BamHI_Pphyt_F and Pphyt_EMSA_R, the latter of which was labeled with fluorescein on the 5’ end. Prior to the EMSA, the Strep Tag was removed from UzrpR using HRV 3C protease (Thermo Fisher) according to the manufacturer’s protocol.
  • UrpR was phosphorylated by incubation in phosphorylation buffer (50 mM Tris, pH 7.9, 150 mM NaCl, 10 mM MgCE) with 50 mM disodium carbamyl phosphate (Sigma-Aldrich) for 1 h at 30°C (Lynch and Lin, 1996) and immediately used in the binding assays.
  • phosphorylation buffer 50 mM Tris, pH 7.9, 150 mM NaCl, 10 mM MgCE
  • disodium carbamyl phosphate Sigma-Aldrich
  • EMSAs were performed by incubating phosphorylated UrpR with P phyt DNA (50 nM) for 10 min at 37°C in buffer containing 50 mM Tris-HCl (pH 7.9), 200 mM NaCl, 10 mM MgCE, 0.1 mg ml 1 BSA, 5% glycerol, 1 mM DTT, and 50 pg ml 'poly- dl-dC.
  • a 5% TBE mini-protean polyacrylamide gel was pre- run with 0.5x TBE at 120 V for 30 min in a Mini-PROTEAN tetra cell (Bio-Rad) prior to loading samples. Samples were run at 100V for 45 min, and the reaction products were visualized using a Biorad Gel Doc XR1 System.
  • Site 300 sample collection Standard operating procedures for sampling and sample handling at LLNL Site 300 have been described in detail [156] and are consistent with the guidance and requirements of the U.S. EPA.
  • the groundwater samples used in this study were collected from each well with either an electrical submersible pump or a bailer. Well samples were placed on ice, filtered using a 0.2 Dm filter, and stored at 4°C.
  • Site 300 samples were diluted in 2% (v/v) nitric acid (trace metal grade) and spiked with an internal holmium standard. U was quantified using a Thermo XSeriesII ICP-MS run in standard mode. The sample introduction system was an ESI PFA-ST nebulizer pumped at 120 m ⁇ /min. Zn, Pb, Cu, Cd, and Cr were quantified using a Thermo iCAP 7400 radial ICP-OES in standard operating mode. Standard curves were generated using a 100 mM uranyl nitrate stock solution and 10 ppm Zn, Pb, Cu, Cd, and Cr ICP-MS standards (Inorganic Ventures).
  • Pphyt -lacZ and Pi36i -lacZ reporter constructs Chromosomally integrated P Phyt -lacZ and Pi36i-lacZ translational fusions were constructed using a two-step sacB counterselection procedure [147] to swap the P U r cA promoter in pDMP82[3] with either P phyt or Pi36i.
  • pDMP82 contains the necessary sequence to generate a translational V WcK -lacZ fusion at the P urcA locus in which the sequence 24 nt downstream of the urcA start codon is fused to E. coli lacZ.
  • the P phyt and Pi36i fragments were amplified from the Caulobacter genome with the primer pairs Phyt-138_F/ Phyt-138_R and 1362_138_F/1362_138_R (Table 4), respectively and cloned into pDMP82 that was linearized using the primers 138_PurcA_F and 138_PurcA_F (Table 4) using In-Fusion cloning.
  • Pphyt -gfp and Pi36i -gfp constructs Plasmid-borne Pphyt -gfp and Pnei-gfp fusions were generated by amplifying P phyt and Pi36i fragments from the Caulobacter genome with the primer pairs Phyt-138_F/ Phyt-138_R and 1362_138_F/1362_138_R (Table 4), respectively, digested with BamHI and Z oRI and cloned into the similarly digested pDMP450 (Table 5), generating pDMP460 and pDMP463 (Table 5).
  • the pl5A origin was then swapped with a pBRR-repl origin that was amplified from pPROBE-GFP’ using the primers pBBRl-rep_F and pBBRl-rep_R and the restriction enzymes Nhel and Hindlll, generating pDMP460.
  • the DNA sequence region from 170 to -9 with respect to the translation initiation site of CCNA_01968 was amplified with P1968_BamHI_F and P1968_EcoRI_R, digested with EcoRI and BamHI, and cloned into the similarly digested pDMP460 to construct a CCNA_01968-promoter gfp fusion (pDMP558).
  • Site directed mutagenesis was performed with the primers P1968_HSl_mut_F and P1968_HSl_mut_R to mutate UzcR half site one from 5’-CATTAC-3' to 5'-CAATAG -3' and primers P1968_HS2_mut_F and P1968_HS2_mut_R and to mutate the half site two from 5’-TTAA-3' to 5'-TAAT-3', generating pDMP559 and pDMP560, respectively.
  • a synthetic DNA containing the rmb T1 and T7Te transcription terminators (BBA_0B0015) followed by a Pi36i fusion with the first 168 nucleotides of uzcR was prepared (Integrated DNA Technologies, Inc.)
  • pDMP499 Table 5
  • a pNPTS138-based vector containing the DNA sequence for substituting the aspartate residue at position 51 for alanine was amplified with primers pNTPS138_urcR_F and pNTPS138_urcR_R (Table 4) and the 531 bp region upstream of the uzcR promoters was amplified with urcR_UR_F and urcR_UR_R (Table 4).
  • Plasmids pDMP609 and pDMP614 were amplified with the primer pairs P1362_rsaFb_F/ P1362_rsaFb_F and Pphyt_rsaFb_F/ Pphyt_rsaFb_F, respectively (Table 5) and re-ligated using InFusion cloning.
  • the resulting suicide vectors pDMP673 and pDMP621 (Table 5) were electroporated into Caulobacter strain FC922 and the P1361m_5-uzcRS (DMP679) and Pphytm_5 -uzcRS (DMP643) strains (Table 5) were obtained by the two-step sacB counterselection procedure. All strains were transformed with pDMP558, encoding a CCNA_01968 promoter gfpmut3 fusion (Table 5). [00468] Engineering of constructs to place expression of hrpS under control of Pphyt or
  • P h rp L DNA (SEQ ID 80) was synthesized (IDT), then digested with BamHI and BglU and ligated into the similarly digested pDMP450, generating pDMP610.
  • the synthetic Pphyt-hrpS_PurcB-hrpR DNA fragment was digested with Xbal and BamHI and cloned into the similarly digested pDMP610 to generate pDMP612.
  • pDMP612 was cloned into NA1000 to produce DMP681.
  • the UzcRS two-component system was identified as the regulatory system responsible for U, Zn, and Cu-dependent activation of P urcA and 41 other promoters in the Caulobacter genome [3]. Together, these data suggest that the P urcA does not have satisfactory selectivity to function as a standalone sensor of environmental U. Nevertheless, since UzcRS exhibits a U- concentration dependence in a wide range of media conditions, a sensor that incorporates UzcRS as one component within a more advanced U- sensitive genetic circuit comprising an additional point of U sensing that is independent of the UzcRS system could produce an effective U sensor.
  • RNA-seq Figure 1
  • Pi36i promoter of operon containing CCNA_01362
  • P phyt promoter of CCNA_01353
  • Example 2 Engineering of a U-sensitive genetic circuit comprising an‘in series’ AND gate wherein the uzcRS operon is placed under the transcriptional control of P D hvt/Pi36i [00472]
  • uzcR and uzcS are physically separated by genes encoding the ParDE3 toxin anti-toxin (TA) system, together forming a putative four-gene operon [112].
  • TA ParDE3 toxin anti-toxin
  • uzcR and uzcS are conserved throughput much of alphaproteobacteria
  • the insertion of parDE3 between uzcR and uzcS is unique to a subset of the Caulobacter genus; uzcR and uzcS are adjacently located in the majority of closely related alphaproteobacteria [3] including C. crescentus strain OR37, an environmental isolate from a U-contaminated site [97].
  • the parDE3 system does not contribute to the metal-dependent regulation by UzcRS [3]. Given this result and the potential toxicity associated with parDE3 overexpression, the parDE3 TA system was deleted, so that uzcR and uzcS are adjacently located.
  • uzcR The expression of uzcR is controlled by two promoters (Pi and P2) in C. crescentus [3], which enables sufficient basal expression of uzcRS to activate transcription in response to metal (U, Zn, Cu).
  • uzcR binding site located upstream of Pi that likely yields a positive feedback loop.
  • UzcR protein levels increase in a wzcS-dcpcndcnt manner in response to metal sensing. Deletion of the parDE3 TA system and the parD promoter places uzcS expression under the exclusive control of Pi and P2 ( Figure 4 Panel A).
  • A“control strain” of Caulobacter was generated comprising a U-sensitive genetic circuit in which uzcR is under the control of Pi and P2 and uzcS under the control of Pi and P2, ( Figure 4 Panel A). As expected, this strain produces a high fluorescence signal in response to U, Zn, Cu ( Figure 5 Panel A left, middle, right graphs, respectively). Incremental improvements were made to this circuit to enhance specificity, as described below.
  • a negative feedback loop was incorporated into the circuit, whereby UzcR represses its own expression from P phyt or Pi36i, in order to minimize the basal expression of the uzcRS operon.
  • a UzcR binding site was placed downstream of the P phyt or Pi36i transcription start site ( Figure 4 Panel C).
  • any m_5 site is suitable.
  • Caulobacter comprising this sensor showed strong responsiveness to U ( Figure 5 Panel C, left graph) and further shifted ratio of U response to that of Zn to 5.5 ( Figure 5 Panel C, left and right graphs).
  • Example 3 Engineering of a U-sensitive genetic circuit comprising an‘in parallel’ HRP AND gate
  • This Example describes the engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ AND gate that utilizes the HRP AND gate system from Pseudomonas syringae that was recently developed in E. coli [126].
  • HRP AND gate system from Pseudomonas syringae that was recently developed in E. coli [126].
  • both HrpS and HrpR are required for s 54 dependent activation of the hrpL promoter (Ri, Gr[ J. Expression of HrpS or HrpR alone is not sufficient for transcriptional activation.
  • the genetic circuit described in this Example contains genetic components whose expression is controlled independently by (1) P phyt or Pi36i and (2) UzcRS systems, and reporter expression requires both HrpS and HrpR to be expressed.
  • Example 4 Engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ tripartite GFP AND gate
  • This Example describes the engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ AND gate that utilizes the tripartite GFP system [5], which requires expression of gfplO, gfpl 1 and gfpl-9 for reporter expression (Figure 6 Panel B).
  • gblock Tripartite GFP gblock
  • Integrated DNA Technologies, gBlock Integrated DNA Technologies, gBlock
  • E1-GFP11-M4 under the control of the UzcRS promoter (P Mrc s)
  • lamdaTo terminator gfpl-9 under the control of the rsaA promoter (P rsa A [4]).
  • GFP10-m2_kl was placed under the control of P phyt or P i % i by digesting Tripartite GFP gblock with Bglll and Xhol and ligating into the similarly digested pDMP460 and pDMP463, respectively, to form pDMP791 and pDMP881.
  • DNA sequence encompassing the entire R,, / ,g,/R I % i tripartite DNA and an insulating upstream rrnbTl transcription terminator was amplified with primers HRP_chrome_int F and R and cloned into pDMP82 that was amplified with the primers urcA_loc_DR_F and urcA-loc_UR_R using infusion cloning to form pDMP792 and pDMP692.
  • a variant of this AND gate containing gfpl-9 under the control of the xylose inducible promoter was constructed by amplifying the 360 bp P xyi fragment using Pxyl_F and Pxyl_R and cloning into pDMP792 (R,, / ,g, version) and pDMP692 ( P i % i version) that were amplified with gfpl-9_amp_for_pxyl_F gfpl-9_amp_for_pxyl_R using infusion cloning to form pDMP664 and pDMP 665, respectively.
  • a variant of this AND gate containing gfpl-9 under the control of the shortened P p h yt promoter was constructed by amplifying pDMP736 with Pphyt_TR_elim_F and Pphyt_SD_amp_R and cloning this 137 bp fragment into pDMP712 ( R,,I, M version) or pDMP713 (Pi36i version) that was amplified with gfpl-9_for_pphyt-short_F and gfpl-9_for_phyt_R using infusion cloning to form pDMP808 and pDMP809.
  • a variant containing GFP10-m2_kl under the control of a shortened version of P phyt and gfpl-9 under the control of P rS aA was constructed by amplifying pDMP792 with Pphyt_TR_elim_F and Pphyt_promoter_shorten_R and re-ligating using infusion cloning, forming pDMP883.
  • a control tripartite variant containing both GFP10-m2_kl and E1-GFP11-M4 under the control of P mCB was constructed in three parts.
  • a P mCB fragment generated with primers XbaI_PurcB and BglII_PurcB_R was digested with Xbal and Bglll and directionally cloned into the similarly digested pDMP712, forming pDMP741.
  • DNA sequence encompassing the entire P mCB tripartite DNA and an insulating upstream rrnbTl transcription terminator was amplified with primers HRP_chrome_int F and R and cloned into pDMP82 that was amplified with the primers urcA_loc_DR_F and urcA- loc_UR_R using infusion cloning to form pDMP748.
  • the P AV/ promoter was swapped for P rsaA by cloning PrsaA, amplified with PrsaA_frag_F and PrsaA_frag_R, into pDMP748 that was amplified with gfp_for_PrsaA_F and gfp_for_PrsaA_R, forming pDMP932.
  • Figure 11 shows graphs reporting exemplary data corresponding to the exemplary U- sensitive tripartite GFP genetic circuit.
  • Tripartite GFP system with PrsaA-gfpl-9 Can detect U in the 2-20 uM range. When a growth media containing Glycerol-2-phosphate as the P source is used, the signal amplitude is higher but the responsive range is shifted to 8 uM-30 uM. Higher concentrations have a diminished signal output. The shifted range likely reflects U coordination by glycerol-2- phosphate that reduces the bioavailability.
  • Example 5 Engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ Bacterial two-hybrid system AND gate.
  • This Example describes the engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ AND gate that utilizes the bacterial two-hybrid system [117].
  • the alpha-gall IP fusions and lamda repressor-gal4 fusion are expected to be driven by a combination of P phy / nei and any UzcRS regulated promoter (see Figure 22).
  • a putative regulator binding site was identified in proximity to the transcription start site in P phyt and Pi36i that is comprised of two direct repeat elements (e.g., CGTCAGC (SEQ ID NO: 202)); Figure 7 Panels A and B). This binding site is conserved amongst Caulobacteridae Bradyrhizobiaceae, Sphingomonadaceae,
  • Hyphomicrobiaceae, and Rhodobacteracea facilitating a phylogenetic footprinting approach to construct a putative regulator DNA-binding motif (Figure 7 Panel C), using 26 DNA sequences of Pphyt and Pi36i from Caulobacter sp. Root342, Phenylobacterium sp. Root700, Caulobacter crescentus NA1000, Caulobacter sp. Rootl455, Caulobacter sp. Root487D2Y, Paracoccus sp. 228, Caulobacteraceae bacterium OTSz_A_272, Novosphingobium sp. AP12 PMI02, Hyphomicrobium sp.
  • MCI Hyphomicrobium denitrificans , Brevundimonas sp. Rootl279, Sphingopyxis sp. Rootl497, Afipia sp. P52-10, Caulobacter sp. Root342, Hyphomicrobium denitrificans, Sphingobium sp. YBL2, Sphingobium baderi LL03, Sphingobium indicum B90A, and Roseovarius indicus strain DSM 26383.
  • Pp hyt (but not Pi36i) also contains a large tandem repeat (TR;32 bp) located further upstream from the putative regulator site ( Figure 8 Panel A).
  • TR;32 bp large tandem repeat
  • Figure 8 Panel A To test the function of this TR in U-dependent induction of P phyt the size of the P phyt DNA was reduced from 238 bp to 81 bp, eliminating the TR and all upstream DNA (Figure 8 Panel C).
  • the data shown in Figure 9 Panel A indicate that the TR is not required for U induction.
  • a shortened version of Pi36i that included only nine bp upstream of the regulator binding site (Figure 8 Panel I) retained U-dependent regulation ( Figure 9 Panel B).
  • the consensus direct repeat UrpR binding site can be integrated within a promoter region, such that UrpR DNA binding will interfere with RNA polymerase binding and/or transcription.
  • the UrpR binding site should be integrated at a location that overlaps, but does not alter the sequence of the -35 and/or -10 promoter elements or the TSS; disturbing -35 and/or -10 promoter elements will yield a promoter with low basal activity. Ideally, multiple locations will be tested to optimize results. While this promoter can be used to control transcription of any biological reporter, destabilized gfp (e.g., GFP-LVA[159]) is expected to yield the best results given the enhanced degradation rate. Highly stable reporters (e.g., GFP) will require several rounds of cell division to observe a uranium (e.g. UrpR) dependent decrease in reporter activity.
  • destabilized gfp e.g., GFP-LVA[159]
  • a positive regulator protein UzcY encoded by CCNA_03497, was identified, which functions as a “natural” signal amplifier for the UzcRS system.
  • uzcY is repressed by the MarR family transcription factor (CCNA_03498), and thus has no effect on UzcRS activity.
  • UzcY expression is induced through relief of CCNA_03498 repression or by ectopic expression, it stimulates UzcS activity through a direct interaction, causing a hypersensitive output in response to the metal inducers U, Zn, and Cu. This has the effect of dramatically increasing the output signal amplitude in response to low U (or Zn/Cu) concentrations, thus increasing sensitivity, and lowering the U detection limit of UzcRS by over 4-fold ( Figure 12).
  • this“natural” signal amplifier module can be integrated into U sensor circuitry.
  • the U-specific promoter P phyt or P/.3 ⁇ 4/ is used to drive expression of uzcY (e.g., see Figure 13) such that UzcY levels are modulated in a U-concentration dependent manner.
  • the low U detection limit of P phyt and Pi36i (-500 nM) is expected to allow signal amplification at environmentally relevant U concentrations, which is expected to improve U sensitivity and lower the detection limit in view of exemplary data demonstrating that both of these properties can be achieved with the native UzcRS system (e.g., see Figure 12).
  • UzcY-mediated signal amplification e.g. by placing UzcY under regulatory control of a U- selective promoter such as Pi36i or Pphyt, the selectivity for U is expected to be further enhanced.
  • a U- selective promoter such as Pi36i or Pphyt
  • FIG 14 An example of combining‘in series’ and‘in parallel’ AND gate circuits within the same cell to enhance selectivity is shown in Figure 14.
  • An advantage of this exemplary circuit is that the UzcRS input is U selective whereas in the original‘in parallel’ circuit as shown in Figure 6 Panel B, the UzcRS -regulated promoter P urcB is cross -reactive with Zn and Cu.
  • Negative regulators of UzcRS The following four different negative regulators of UzcRS that are encoded in the Caulobacter chromosome were identified (see Figure 27). None of these regulators are required for U sensing by UzcRS, however, their levels modulate the sensitivity to U.
  • Negative regulator 1 CCNA_03681 and CCNA_03680 encode an ABC transporter ATPase and an ABC-2 family transporter fused to a C-terminal aminopeptidase N domain, respectively (urtAP). Together these proteins form an ABC transporter with a C-terminal aminopeptidase domain.
  • Negative regulator 2 CCNA_02866 (also referred to herein as uzcX), encodes a membrane protein of unknown function that is located within a prophage region of the genome and part of the UzcR direct regulon ( ⁇ 8-fold activated by UzcR (Park el al, 2017) [3].
  • Negative regulator 3 A MarR family regulator CCNA_03498 that represses expression of an operon containing CCNA_03497, CCNA_03498 and CCNA_03499 (see Figure 26). Expression of CCNA_03497 (occurs when repression mediated by CCNA_03498 is lifted) hypersensitizes UzcS to metal inducers.
  • Negative regulator 4 A second, paralogous MarR family regulator CCNA_02289 that represses expression of an operon containing CCNA_02291, CCNA_02290 and CCNA_02289 (see Figure 26). Expression of CCNA_02291 (occurs when repression mediated by CCNA_02289 is lifted) hypersensitizes UzcS to metal inducers.
  • This Example describes a U biosensor having exemplary U-neutralization outputs in response to bioavailable U.
  • FIG. 21 shows a schematic showing exemplary U-sensitive genetic circuits having exemplary U-neutralizing outputs to allow U bioprecipitation or bioadsorption.
  • an exemplary AND gate comprises a P phyt promoter configured to initiate expression of UzcR and UzcS in presence of bioavailable U, and a P urcB promoter (activated by UzcR) configured to initiate expression of exemplary U-neutralizing genes phoY or phytase (to provide a U bioprecipitation output), or fusion genes of rsaA-SUP, rsaA-CaM, ompA- SUP, or ompA-CaM (to provide a bioadsorption output).
  • Table 7 shows a list of UzcR regulated promoters. Bolded regions are putative UzcR m_5 sites and single boded nucleotide is TSS.
  • Example 12 Application of a combinatorial input logic towards U detection
  • a combinatorial sensor approach using multiple regulators with broad specificity profiles can be adopted for the selective detection of compounds lacking specific regulators (e.g., butanol).
  • the operons are furthermore likely regulated by the same transcription factor based on the presence of nearly identical DNA sites comprised of two tandem direct repeats (5’-GTCAG- 3’; Figure 7A-B) with 11-bp center-to-center (etc) spacing within the CCNA_01353 (Pp hyi ) and CCNA_01361 promoters (Pi36i).
  • This direct repeat site is conserved within the promoters of closely related alpha proteobacteria (Figure 7C) and required for U-dependent induction of P pi , y , and Pi36i; mutations away from consensus in both P phyr and Pi36i -gfp fusions abrogate U- induction ( Figure 9A-B).
  • the CCNA_01353-CCNA_01352-CCNA_01351operon encodes a phytase enzyme that confers U tolerance [162] and an uncharacterized response regulator and histidine kinase pair
  • the CCNA_01361-CCNA_01362-CCNA_01363 operon encodes a PepSY superfamily protein and another uncharacterized response regulator and histidine kinase pair.
  • UrpR and UrpS Uranium Responsive Phytase Regulator and Sensor, respectively
  • UrpR was purified, and its binding to P p h y t was tested using an electrophoretic mobility shift assay. As expected, UrpR bound to P phyt in a concentration-dependent manner (Figure 10). Binding was not observed to a P phyt fragment containing a 5’-GTCA-3’ to 5’-CAGT-3’ mutation at DR2 ( Figure 10), supporting the functional role of the direct repeat site in UrpR DNA binding.
  • TCS comprised of UrpR and UrpS is the U-dependent activator of P p h y t and Pi36i, revealing a positive feedback loop within this regulatory system.
  • UrpRS exhibits improved metal selectivity compared to UzcRS
  • UrpRS functions independently of UzcRS with respect to U perception (a property that is important for its dual integration with UzcR in a synthetic U sensing pathway)
  • the expression of P p h y t and the UzcR-regulated promoter P mCB were tested in strains lacking the non-cognate TCS.
  • Example 15 Construction of a U-responsive AND gate pathway in C. crescentus
  • GFP is split into three parts (gfplO, gfpl l and gfpl-9) that interact to reconstitute active GFP when co-expressed; the synthetic K1 and El coiled-coils[172] were fused to the C-terminus of gfplO and the N-terminus of gfpl l, respectively, to mediate dimerization.
  • An advantage of the tripartite system is the low basal level of GFP fluorescence, ensuring a robust OFF state.
  • Example 16 U-sensing AND gate exhibits improved selectivity relative to UzcRS alone
  • the core sensor was unresponsive to Ni, Cd, Th, Al, Fe(III), Fe(II), Mn, Co, arsenate, Se, and chromate (Figure 35).
  • the core sensor and the sensor variant driven by UrpRS alone were also unresponsive to Cu, in contrast to the sensor driven exclusively by UzcRS ( Figure 40).
  • the sensitivity of the core sensor to Zn and Pb was significantly diminished relative to the UzcRS control; high Pb concentrations (greater than 20 uM) were required for weak fluorescence induction while the Zn-induced fluorescence was reduced relative to U for every tested concentration.
  • Example 17 Integration of UzcY signal amplifier improves U sensitivity and selectivity.
  • a notable limitation of this AND gate approach is that the improved selectivity comes at a cost to U sensitivity in the low micromolar range; incorporation of the less sensitive UzcRS TCS yielded a sensor with lower sensitivity compared to the control sensor driven by UrpRS alone ( Figure 34).
  • swapping the UzcRS -regulated P mC B promoter with P urC A a promoter that is highly induced by UzcR and sensitive to low UzcR-P concentrations, failed to significantly improve sensor sensitivity (Data not shown). This suggests that simply swapping P urcB with an alternative UzcR-regulated promoter is unlikely to remedy the sensitivity limitation.
  • uzcY expression is silenced by the MarR family regulator MarRi (CCNA_03498) and has no effect on UzcRS activity.
  • CCNA_03498 MarR family regulator MarRi
  • the signal amplification mechanism was incorporated within the core sensor by either deleting marRi , which yields a constitutive amplifier function, or by swapping the native uzcY promoter with the UrpRS-regulated Pphyt-short ( Figure 36A), such that UzcY levels are modulated in a U-concentration dependent manner.
  • Example 18 Whole-cell U sensor detects as low as 1.0 uM in groundwater
  • the organophosphate G2P was employed that serves the dual function of providing a phosphate source for cell growth while maintaining initial U solubility through complexation.[161, 173] Since the physicochemical form— or speciation of U— is dependent on the geochemical conditions and strongly influences bioavailability, and consequently the detectability by this biosensor, addition of G2P may be an effective means of conditioning the environmental samples for U detection. Evaluating this hypothesis, and ultimately, the utility of the sensor for environmental monitoring, will require systematic characterization of the solution matrix composition.
  • Example 19 Naturally occurring Fluoride sensing riboswitches and related consensus sequence
  • the Rfam database was queried to identify naturally occurring F-sensing riboswitches comprising a crcB motif.
  • Example 20 Establishing a fluoride-detection capability in C. crescentus
  • Figure 44 shows a schematic illustration of an approach which will be used to develop a bacteria-based sensor of uranyl fluoride products.
  • uranium- and fluoride- detection components are expected to be integrated and optimized within C. crescentus (right panel) and the UO2F2 detection performance under aqueous conditions systematically characterized.
  • the left schematic depicts a foreseeable application of the whole cell sensor: autonomous environmental monitoring for aqueous UO2F2 species by a C. crescentus monolayer biofilm.
  • the fluoride sensing crcB riboswitch [102] will be engineered to control the expression of a fluorescent reporter (e.g ., mCherry) such that fluoride perception leads to cell fluorescence.
  • a fluorescent reporter e.g ., mCherry
  • Initial sensor strains will be built with the crcB motifs from three distinct bacteria, including the well-characterized crcB motifs in P. syringae DC3000 and Bacillus subtilis [102] and an uncharacterized crcB motif from Sphingomonas sp. MM-1. Sphingomonas sp. MM-1 represents a particularly promising option since this bacterium is closely related to C.
  • crcB-mCherry reporters will be integrated within the C. crescentus chromosome.
  • mCherry fluorescence will be quantified in a high-throughput (96-well format) manner over a range of NaF concentrations. Then, the fluoride selectivity will be characterized using commonly encountered anions (e.g., Cl , NO3 , SO4 2 , PO4 3 ).
  • Figure 45 shows a schematic illustration of the approach.
  • Phase I fluoride sensor crcB-mCherry fusion built in wild type strain (Right).
  • Phase II fluoride sensor crcB-mCherry fusion built in strain deleted for fluoride export (D crcB).
  • Deletion of the crcB gene in E. coli improved the detection limit of a crcB-lacZ reporter by over 100-fold [102].
  • this approach is expected to yield a fluoride-sensing capability that, when coupled with a U- sensing component herein described, will enable UO2F2 detection by C. crescentus.
  • Example 22 Reduction of the fluoride detection limit by manipulating the expression levels of the native fluoride detoxification svstem/s in C. crescentus.
  • C. crescentus possesses a putative fluoride ion transporter (CrcB) and withstands high mM concentrations of fluoride, suggesting native mechanism/s of fluoride detoxification.
  • CrcB putative fluoride ion transporter
  • Prior data in E. coli indicate that the CrcB fluoride exporter functions to maintain low intracellular fluoride concentrations, adversely affecting the detection limit of the colorimetric crcB reporter ( ⁇ 1 mM).
  • the detection limit of phase 1 fluoride sensors development is expected to be significantly higher than the detection limit of our U sensor ( ⁇ 2 mM), and thus unlikely to be acceptable for nuclear effluent detection.
  • the fluoride tolerance will be partly rescued by expressing the crcB fluoride exporter over a range of concentrations to identify levels that yield a desirable balance between the organism’s fluoride tolerance and the sensitivity of the crcB-mCherry.
  • Example 23 Integration of uranyl- and fluoride-sensing components and evaluate detection performance under aqueous conditions.
  • a U-sensing AND gate circuit will be integrated with the optimal performing crcB- mCherry circuit within C. crescentus to enable UO2F2 detection. Two distinct configurations will be constructed and further characterized.
  • Figures 46 and 47 show a schematic illustration of the related approach.
  • the configurations of the exemplary U0 2 F 2 -biosensors of Figure 46 and Figure 47 will provide an individual readout of uranium and fluoride levels and provides a safeguard for environments where natural uranium or fluoride occur at elevated levels.
  • Uranium occurs naturally at concentrations of -10-100 ppb in soils and -5- 100s of ppb in ground water [37]while fluoride levels in groundwater, sea, and soil are typically in the 10-100 mM range [176].

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Genetics & Genomics (AREA)
  • Chemical & Material Sciences (AREA)
  • Organic Chemistry (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • General Engineering & Computer Science (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Biomedical Technology (AREA)
  • Biotechnology (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • Biochemistry (AREA)
  • Physics & Mathematics (AREA)
  • Microbiology (AREA)
  • Biophysics (AREA)
  • Plant Pathology (AREA)
  • High Energy & Nuclear Physics (AREA)
  • Medicinal Chemistry (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Virology (AREA)
  • Tropical Medicine & Parasitology (AREA)
  • Immunology (AREA)
  • Analytical Chemistry (AREA)
  • Gastroenterology & Hepatology (AREA)
  • Micro-Organisms Or Cultivation Processes Thereof (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

UO2F2 biosensors, and related U-sensing and/or F-sensing genetic molecular components, genetic circuits, compositions, methods and systems are described, which in several embodiments can be used to detect and/or neutralize uranium and in particular bioavailable UO2F2.

Description

BIOSENSORS FOR DETECTING AND/OR NEUTRALIZING BIOAVAILABLE URANIUM AND RELATED U-SENSITIVE GENETIC MOLECULAR COMPONENTS, GENE CASSETTES, VECTORS, GENETIC CIRCUITS, COMPOSITIONS, METHODS
AND SYSTEMS
CROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims priority to US provisional application number 62/801,077 enfitled“ Biosensors for detecting and/or Neutralizing Bioavailable Uranium and Related U- sensitive Genetic Molecular components, Gene Cassettes, Vectors, Genetic Circuits, Compositions, Methods and Systems’ filed on February 4, 2019 with docket number IL13081-2, which is incorporated by reference in its entirety. The present application is further related to US
Application S/N _ entitled “Biosensors for detecting and/or Neutralizing
Bioavailable Uranium and Related U-sensitive Genetic Molecular components, Gene Cassettes, Vectors, Genetic Circuits, Compositions, Methods and Systems” filed on February 4, 2020 with Docket Number IL13081-2 herein incorporated by reference in its entirety. The present application is also related to International patent application PCT/US2018/061667, entitled “Biosensors for Detecting and/or Neutralizing Bioavailable Uranium And Related U-Sensitive Genetic Molecular Components, Gene Cassettes, Vectors, Genetic Circuits, Compositions, Methods And Systems” filed on November 16, 2018 with Docket No. IL13081 which claims priority to U.S. provisional application No. 62/587,753, entitled“Biosensors for Detecting and/or Neutralizing Bioavailable Uranium And Related U-Sensitive Genetic Molecular Components, Gene Cassettes, Vectors, Genetic Circuits, Compositions, Methods And Systems” filed on November 17, 2017, with docket number IL-13081, the disclosure of each of which is incorporated by reference in its entirety.
STATEMENT OF GOVERNMENT GRANT
[0002] This United States Government has rights in this invention pursuant to Contract No. LDRD# 16-LW-055 between the United States Department of Energy and Lawrence Livermore National Security, LLC, for the operation of Lawrence Livermore National Laboratory. FIELD
[0003] The present disclosure relates to uranium (U) biosensors and related U-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems. In particular, the present disclosure relates to U biosensors and related methods and systems to detect and/or neutralize bioavailable uranium and more particularly bioavailable uranyl oxycation.
BACKGROUND
[0004] Various methods and systems such as spectroscopy as well as antibodies and DNA enzymes are available to monitor environmental U concentrations which is important to minimize human exposure and inform remediation strategies.
[0005] In particular, detection of U in an environment in situ , and more particularly detection of bioavailable U that can transverse the cell membrane and exert toxicity, is an important part of an evaluation of the potential risk of environmental U exposure.
[0006] However, despite availability of various approaches, development of sensing technologies that provide sensitive, selective and/or cost-effective detection of bioavailable U is still challenging.
SUMMARY
[0007] Provided herein are U biosensors, and related U-sensing genetic molecular components, gene cassettes, genetic circuits, compositions, methods and systems which in several embodiments can be used to detect and/or neutralize uranium and in particular bioavailable UF6, or its stable hydrolysis product UO2F2, which is typically produced in U enrichment operations.
[0008] According to a first aspect, a U02F2-biosensor is described comprising a U-sensing genetic molecular component and an F-sensing riboswitch configured to report and/or neutralize uranium in presence of bioavailable Uranium and Fluoride, in a single output or dual output configuration. The U02F2-biosensor comprises a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 herein also UrpS, and U sensitive transcriptional regulator 1362 herein also UrpR. In the UC Fi-biosensor, the genetically modified bacterial cell is an engineered bacterial cell comprising a 1362 U-sensing reportable genetic molecular component and/or a 1362 U- sensing/U-neutralizing genetic molecular component, each comprising a U-sensitive promoter in a configuration wherein the U-sensitive promoter directly initiates expression of the 1362 U- sensing reportable molecular component and/or of the 1362U-neutralizing molecular component in presence of bioavailable U.
In the U-biosensor, the 1362 U-sensitive promoter comprises a 1362 (UrpR) binding site having a DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO:l), wherein
Ni is C or T, preferably C;
N2 is G or A, preferably G;
N3 is T or C, preferably T;
N4 is C;
Ns is A or G, preferably A;
Nό is G or C, preferably G;
N7 C or G;
Ns is any nucleotide;
N9 is any nucleotide;
Niois any nucleotide;
N11 is any nucleotide;
N12 is T or C;
Ni3 is G;
Ni4 is T or C, preferably T;
Nis is C;
N16 IS A or C, preferably A;
Nn is G; and
Nis is C or G, and wherein Ni to Nn are selected independently.
In the U02F2-biosensor, the genetically modified bacterial cell is an engineered bacterial cell further comprising an F-sensing riboswitch within at least one of
- the 1362 U- sensing reportable genetic molecular component wherein the 1362 U- sensing reportable genetic molecular component, is transcribed in presence of an effective amount of bioavailable fluoride,
- an F-sensing reportable genetic molecular component in a configuration wherein the F- sensing reportable genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride, and
- an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F.
In the U02F2-biosensor comprising the F-sensing riboswitch within the 1362 U-sensing reportable genetic molecular component, the F-sensing riboswitch and the 1362 U-sensing reportable genetic molecular component are are in a single output configuration.
In the U02F2-biosensor comprising the F-sensing reportable genetic molecular component and/or the F-sensing genetic circuit, the F-sensing riboswitch and the 1362 U-sensing reportable genetic molecular component are in a dual output configuration.
In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing the histidine kinase 1363 (UrpS) and the U-sensitive transcriptional regulator 1362 (UrpR) (e.g. E. Coli bacteria). In those embodiments the bacterial cell is further engineered to include a 1362 U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363(UrpS), and a gene encoding response regulator 1362 are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing histidine kinase 1363 (UrpS), and U-sensitive transcriptional regulator 1362 (UrpR) (e.g. proteobacteria such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding the histidine kinase 1363, and the U- sensitive transcriptional regulator 1362 (UrpR), are knocked out and the genetically engineered bacterial cell is further engineered to include a 1362 U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363 (UrpS), and a gene encoding response regulator 1362 (UrpR) are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure). In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding the F-sensing riboswitch, are preferably knocked out.
[0009] According to a second aspect, a UC Fi-biosensor is described, comprising a U-sensitive F-sensitive genetic circuit wherein a U-sensing genetic molecular component and an F-sensing riboswitch are configured to report and/or neutralize uranium in presence of bioavailable Uranium and Fluoride, in a single output or dual output configurations.
The UO2F2 biosensor according to the second aspect comprises a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR). In the U02F2-biosensor, the genetically modified bacterial cell is an engineered bacterial cell comprising a U-sensitive F-sensitive genetic circuit in which molecular components are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components. In the 1362 U-sensitive F-sensitive genetic circuit at least one molecular component is a 1362 U- sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U-sensitive 1362 (UrpR) binding site having a DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO:l), wherein
Ni is C or T, preferably C;
N2 is G or A, preferably G;
N3 is T or C, preferably T;
N4 is C;
Ns is A or G, preferably A;
N , is G or C, preferably G; N7 C or G;
Ns is any nucleotide;
N9 is any nucleotide;
Niois any nucleotide;
N11 is any nucleotide;
N12 is T or C;
Ni3 is G;
Ni4 is T or C, preferably T;
Nis is C;
N16IS A or C, preferably A;
Nn is G; and
Nis is C or G,
and wherein Ni to Nn are selected independently.
In the U-sensitive F-sensitive genetic circuit, at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), and/or a U-neutralizing molecular component, (in particular one or more U-neutralizing genetic molecular component and/or one or more U-neutralizing cellular molecular component).
In the UC Fi-biosensor, the 1362 U-sensing F-sensing genetic circuit further comprises an F sensing riboswitch within at least one genetic molecular component of the molecular components of the U-sensitive F-sensitive genetic circuit, in a configuration wherein the at least one genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
In the U02F2-biosensor, the at least one reportable molecular component and/or the U- neutralizing molecular component are expressed when the genetic circuit operates according to the circuit design in presence of an effective amount of bioavailable U and an effective amount of bioavailable F in a single output or dual output configurations.
In some embodiments, the genetically modified bacteria are bacteria not capable of natively expressing the histidine kinase 1363 (UrpS) and the U-sensitive transcriptional regulator 1362 (UrpR) (e.g. E. Coli bacteria). In those embodiments, the bacterial cell is further engineered to include a 1362 U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363 (UrpR), and a gene encoding response regulator 1362(UrpR) are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing histidine kinase 1363(UrpS), and U-sensitive transcriptional regulator 1362(UrpR) (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding the histidine kinase 1363(UrpS), and the U- sensitive transcriptional regulator 1362(UrpR), can be preferably knocked out and the genetically engineered bacterial cell is further engineered to include a 1362 U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 in a configuration wherein the a gene encoding histidine kinase 1363(UrpS), and a gene encoding response regulator 1362 (UrpR) are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure). In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding the F-sensing riboswitch, are preferably knocked out.
[0010] According to a third aspect, a UC Fi-biosensor is described, comprising a U-sensing genetic molecular component and an F-sensing riboswitch configured to report and/or neutralize uranium in presence of bioavailable Fluoride, in a single output or double output configuration.
The UC Fibiosensor comprises a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase UzcS and U-sensitive transcriptional response regulator UzcR.
In the U-biosensors, the genetically modified bacterial cell is an engineered bacterial cell comprising a UzcR U-sensing reportable genetic molecular component and/or a UzcR U-sensing U-neutralizing genetic molecular component, each comprising a U-sensitive promoter in a configuration wherein the U-sensitive promoter directly initiates expression of the UzcR U- sensing reportable molecular component and/or of the UzcR U-sensing U-neutralizing molecular component in presence of bioavailable U.
In the U-biosensor, the U-sensitive promoter comprises an UzcR binding site having a DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide, and in some embodiments any one of N7-N11 can independently be A.
In the U02F2-biosensor, the genetically modified bacterial cell is an engineered bacterial cell further comprising an F-sensing riboswitch within at least one of
- the UzcR U-sensing reportable genetic molecular component wherein the UzcR U- sensing reportable genetic molecular component, is transcribed in presence of an effective amount of bioavailable fluoride,
- an F-sensing reportable genetic molecular component in a configuration wherein the F- sensing reportable genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride, and
- an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F.
In the UC Fi-biosensor comprising the F-sensing riboswitch within the UzcR U-sensing reportable genetic molecular component, the F-sensing riboswitch and the UzcR U-sensing reportable genetic molecular component are in a single output configuration.
In the UC Fi-biosensor comprising the F-sensing reportable genetic molecular component and/or the F-sensing genetic circuit, the F-sensing riboswitch and the UzcR U-sensing reportable genetic molecular component are in a dual output configuration.
In some embodiments, the genetically modified bacteria are bacteria not capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. E. Coli bacteria), and the bacterial cell is further engineered to include a UzcR U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding the histidine kinase UzcS, and an endogenous or exogenous gene encoding the U- sensitive transcriptional response regulator UzcR, in a configuration wherein the gene encoding the histidine kinase UzcS, and the gene encoding the U-sensitive transcriptional response regulator UzcR, are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, in the bacterial cell, the endogenous genes encoding the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR, can be preferably knocked out and the genetically engineered bacterial cell can be further engineered to include a UzcR U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure). In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding the F-sensing riboswitch, are preferably knocked out.
[0011] According to a fourth aspect, a UC Fi-biosensor is described, comprising a U-sensitive F- sensitive genetic circuit wherein a UzcR U-sensing genetic molecular component and an F- sensing riboswitch are configured to report and/or neutralize uranium in presence of bioavailable Uranium and Fluoride, in a single output or dual output configurations.
The U biosensor comprises a genetically modified bacterial cell natively and/or heterologously expressing histidine kinase UzcS, and response regulator UzcR.
In the UC Fi-biosensor, the genetically modified bacterial cell is an engineered bacterial cell comprising a U-sensitive F-sensitive genetic circuit in which molecular components are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components. In the UzcR U-sensitive F-sensitive genetic circuit at least one molecular component is a UzcR U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a UzcR binding site having DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide, and in some embodiments any one of N7-N11 can independently be A.
In the UzcR U-sensitive F-sensitive genetic circuit at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), and/or a U-neutralizing molecular component, (in particular one or more U-neutralizing genetic molecular component and/or one or more U-neutralizing cellular molecular component),
In the U02F2-biosensor the genetically modified bacterial cell further comprises an F-sensing riboswitch within at least one genetic molecular component of the molecular components of the U-sensitive F-sensitive genetic circuit in a configuration wherein the at least one genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
In the U02F2-biosensor, the at least one reportable molecular component and/or the a U- neutralizing molecular component are expressed when the genetic circuit operates according to the circuit design in presence of an effective amount of bioavailable U and an effective amount of bioavailable F in a single output or dual output configurations.
In some embodiments, the genetically modified bacteria are bacteria not capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. E. Coli bacteria). In those embodiments, the bacterial cell is further engineered to include a UzcR U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding the histidine kinase UzcS, and an endogenous or exogenous gene encoding the U-sensitive transcriptional response regulator UzcR, in a configuration wherein the gene encoding the histidine kinase UzcS, and the gene encoding the U-sensitive transcriptional response regulator UzcR, are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, in the bacterial cell, the endogenous genes encoding the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR, are knocked out. and the genetically engineered bacterial cell is further engineered to include a UzcR U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure). In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding the F-sensing riboswitch, are preferably knocked out.
[0012] According to a fifth aspect, a UC Fi-biosensor is described comprising a U-sensing genetic circuit and an F-sensing riboswitch configured to report and/or neutralize uranium in presence of bioavailable Uranium and Fluoride in a dual output configuration. The UO2F2- biosensor comprises a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 herein also UrpS, and U sensitive transcriptional regulator 1362 herein also UrpR and/or natively and/or heterologously expressing histidine kinase UzcS, and response regulator UzcS.
In the U02F2-biosensor, the genetically modified bacterial cell is an engineered bacterial cell comprising a U-sensitive genetic circuit in which molecular components are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components.
In the U-sensitive genetic circuit at least one molecular component is a U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U. In the U-sensitive genetic circuit, when the cell is capable of natively and/or heterologously expressing histidine kinase 1363 herein also UrpS, and U sensitive transcriptional regulator 1362 herein also UrpR, at least one molecular component is a 1362 U-sensing genetic molecular component in which the U sensitive promoter comprises a U-sensitive 1362 (UrpR) binding site having a DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO:l), wherein
Ni is C or T, preferably C;
N2 is G or A, preferably G;
N3 is T or C, preferably T;
N4 is C;
Ns is A or G, preferably A;
N , is G or C, preferably G; N7 C or G;
Ns is any nucleotide;
N9 is any nucleotide;
Niois any nucleotide;
N11 is any nucleotide;
N12 is T or C;
Ni3 is G;
Ni4 is T or C, preferably T;
Nis is C;
N16IS A or C, preferably A;
Nn is G; and
Nis is C or G,
and wherein Ni to Nn are selected independently. In addition or in the alternative, when the cell is natively and/or heterologously expressing histidine kinase UzcS, and response regulator UzcS, in the U-sensitive genetic circuit at least one molecular component is a UzcR U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a UzcR binding site having DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide, and in some embodiments any one of N7-N11 can independently be A.
In the U-sensitive genetic circuit, at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), and/or a U-neutralizing molecular component, (in particular one or more U-neutralizing genetic molecular component and/or one or more U- neutralizing cellular molecular component), the reportable molecular component and/or the a U- neutralizing molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable U.
In the U02F2-biosensor the genetically modified bacterial cell is an engineered bacterial cell further comprising an F-sensing riboswitch within at least one of
a genetic molecular component of an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F; and
an F-sensing reportable genetic molecular component in a configuration wherein the F- sensing reportable genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
In the U02F2-biosensor, the F-sensing reportable genetic molecular component and the F-sensing genetic circuit are in a dual output configuration with the U-sensing genetic circuit.
In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing the histidine kinase 1363 (UrpS) and the U-sensitive transcriptional regulator 1362 (UrpR) (e.g. E. Coli bacteria). In those embodiments the bacterial cell is further engineered to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363(UrpS), and a gene encoding response regulator 1362 are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing histidine kinase 1363 (UrpS), and U-sensitive transcriptional regulator 1362 (UrpR) (e.g. proteobacteria such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding the histidine kinase 1363, and the U- sensitive transcriptional regulator 1362 (UrpR), are knocked out and the genetically engineered bacterial cell is further engineered to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS), and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the a gene encoding histidine kinase 1363 (UrpS), and a gene encoding response regulator 1362 (UrpR) are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria not capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. E. Coli bacteria), and the bacterial cell is further engineered to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding the histidine kinase UzcS, and an endogenous or exogenous gene encoding the U-sensitive transcriptional response regulator UzcR, in a configuration wherein the gene encoding the histidine kinase UzcS, and the gene encoding the U-sensitive transcriptional response regulator UzcR, are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, in the bacterial cell, the endogenous genes encoding the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR, can be preferably knocked out and the genetically engineered bacterial cell can be further engineered to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
[0013] In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure). In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding , encoding a fluoride efflux pump are preferably knocked out.In some embodiments of the UC Fi-biosensors according to the third a fourth and a fifth aspect wherein the UC Fi-biosensor comprises a U- sensing genetic molecular component in which a U-sensitive promoter comprising a UzcR binding site, the genetically modified bacterial cell is a bacterial cell capable of natively expressing MarR family repressors such as marRi (CCNA_03498) and marR2 (CCNA_02298) genes (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure). In those embodiments, the bacterial cell is preferably further engineered to knock out at least one endogenous MarR family repressors such as marRi and marR genes to provide an amplified UC Fi-biosensor configured to provide an amplified signal following activation of the UO2F2- sensitive genetic circuit.
[0014] In some preferred embodiments of the U02F2-biosensor s according to the third a fourth aspect and a fifth aspect wherein the U02F2-biosensor comprises a U-sensing genetic molecular component in which a U-sensitive promoter comprises a UzcR binding site,, the UO2F2- biosensor or the U-sensitive F sensitive genetic circuit, further comprises an amplifier genetic molecular component comprising a U-sensitive promoter and UzcY and/or UzcZ in a configuration wherein the U-sensitive promoter directly initiates expression of the amplifier molecular component.
[0015] According to a sixth aspect, a method to provide a UC Fi-biosensor is described, the method comprising
genetically engineering a bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and U sensitive response regulator 1362 (UrpR) and/or U sensitive response regulator UzcR in combination with a heterologous F-sensing riboswitch, the genetically engineering performed by introducing into the cell
one or more U-sensing genetic molecular components configured to report and/or neutralize U herein described,
an F-sensitive riboswitch within the one or more U-sensing genetic molecular components configured to report U,
an F-sensitive riboswitch within an F-sensing reportable genetic molecular component;
one or more genetic molecular components of an F sensitive genetic circuit described herein, and/or
one or more genetic molecular components of the U-sensitive F-sensitive genetic circuits described herein,
to provide a UC Fi-biosensor according to the first aspect, the second aspect, the third aspect the fourth aspect, and/or the fifth aspect herein described and
optionally operatively connecting the UC Fi-biosensor so provided to an electronic signal transducer adapted to convert a UOiFibioscnsor reportable molecular component output into an electronic output.
In some embodiments wherein the bacterial cell is a cell incapable of natively expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR, the method further comprises genetically engineering the cell to include a U-sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362(UrpR) and/or response regulator UzcR in a configuration wherein the a gene encoding histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and a gene encoding response regulator 1362 (UrpR) response regulator UzcR are expressed upon activation of a controllable promoter.
In some embodiments wherein the bacterial cell is a cell capable of natively expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR (e.g. proteobacteria, such as certain alpha proteobacteria, beta proteobacteria and/or gamma proteobacteria identifiable by a skilled person upon reading of the present disclosure), the genetically engineering can preferably further comprises knocking out the natively expressed histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR of the bacterial cell, and introducing in the bacterial cell a U-sensing regulator component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) and/or response regulator UzcR in a configuration wherein the a gene encoding histidine kinase 1363 and/or histidine kinase UzcS, and a gene encoding response regulator 1362 (UrpR) response regulator UzcR are expressed upon activation of a controllable promoter.
[0016] In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure). In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding the F-sensing riboswitch, are preferably knocked out.
[0017] According to a seventh aspect a UC Fi-sensing gene cassette described. The UO2F2- sensing gene cassette comprises one or more U-sensing genetic molecular components herein described, one or more U-sensing regulator genetic molecular components herein described and/or one or more reportable genetic component herein described. The U02F2-sensing gene cassette further comprises an F-sensitive riboswitch within the one or more U-sensing genetic molecular components configured to report U, within the one or more U-sensing regulator genetic molecular components and/or within an additional reportable genetic molecular component, in a configuration wherein the U-sensing genetic molecular components, the one or more U-sensing regulator genetic molecular components, and the additional reportable genetic molecular component are transcribed in presence of an effective amount of bioavailable fluoride. In some embodiments the U-sensing gene cassette is an expression cassette. In some embodiments, the gene cassette is comprised within a vector.
[0018] According to an eighth aspect a vector is described comprising a polynucleotide encoding for one or more U-sensing genetic molecular components herein described, one or more F- sensing genetic molecular components herein described, one or more U-sensing and/or F sensing regulator genetic molecular components herein described and/or one or more genetic molecular components of a UC Fi-biosensor herein described. The one or more vectors are configured to introduce one or more U-sensitive genetic molecular components, one or more F-sensing genetic molecular components and/or one or more genetic molecular components of a U-sensitive and/or F-sensitive genetic circuit into a bacterial cell of a plurality of bacterial cells.
[0019] According to a ninth aspect, a UC Fi-sensing system is described. The UOiFiscnsing system comprises one or more vectors herein described and/or a plurality of bacterial cells natively and/or heterologously expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR in combination with one or more F-sensitive riboswitches.
In some embodiments wherein the bacterial cell is a cell incapable of natively expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR, the bacterial cell is further genetically engineered to include a U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363(UrpS) and/or histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362(UrpR) and/or response regulator UzcR in a configuration wherein the a gene encoding histidine kinase 1363(UrpS) and/or histidine kinase UzcS, and a gene encoding response regulator 1362 (UrpR) response regulator UzcR are expressed upon activation of a controllable promoter.
In some embodiments wherein the bacterial cell is a cell capable of natively expressing histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR (e.g. proteobacteria, such as alphaproteobacteria, beta proteobacteria and/or gamma proteobacteria), the bacterial cells is preferably further genetically engineered to comprises knocking out the natively expressed histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) and/or histidine kinase UzcS and response regulator UzcR of the bacterial cell, and introducing in the bacterial cell a U-sensing regulator component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362(UrpR) and/or response regulator UzcR in a configuration wherein the a gene encoding histidine kinase 1363 and/or histidine kinase UzcS, and a gene encoding response regulator 1362(UrpR) response regulator UzcR are expressed upon activation of a controllable promoter.
In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch (e.g. Caulobacter crescentus) identifiable by a skilled person upon reading of the present disclosure). In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch (e.g. Sphingomonas sp. MM- 1, and Caulobacterales bacterium) identifiable by a skilled person upon reading of the present disclosure). In some of those embodiments, the endogenous genes encoding the F-sensing riboswitch, are preferably knocked out.
[0020] According to a tenth aspect, a UC Fi-sensing system is described. The system comprises one or more of the UO2F2 biosensors herein described operatively connected to an electronic signal transducer adapted to convert a U biosensor reportable molecular component output into an electronic output.
[0021] According to an eleventh aspect, a composition is described. The composition comprises one or more UO2F2 biosensors, U02F2-sensing gene cassettes and/or vectors herein described together with a suitable vehicle.
[0022] According to a twelfth aspect, a system comprising an electronic signal transducer adapted to convert a UO2F2 biosensor reportable molecular component output into an electronic output is described. The system comprises an electronic signal transducer and one or more UO2F2 biosensors herein described operatively connected to the electronic signal transducer. [0023] According to a thirteenth aspect, a method of detecting, reporting and/or neutralizing bioavailable UO2F2 is described. The method comprises:
contacting one or more UO2F2 biosensors herein described, or a system comprising an electronic transducer operatively connected to one or more UO2F2 biosensors herein described, with a target environment comprising one or more target ranges of U concentration in combination with one or more target F concentration for a time and under conditions to detect, report and/or neutralize bioavailable UO2F2 in the target environment.
[0024] According to a fourteenth aspect, one or more U02F2-sensing genetic reportable components are also described, the UO2F2 sensing genetic reportable components comprising a U sensitive promoter comprising a U-sensitive 1362 (UrpR) binding site and/or a U sensitive promoter comprising a U sensitive UzcR binding site together with
an F-sensing riboswitch,
in a configuration wherein the U-sensitive promoter directly initiates expression of the U- sensing reportable molecular component and/or of the U- sensing U-neutralizing molecular component in presence of bioavailable U and the U- sensing reportable molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
[0025] According to a fifteenth aspect, one or more U-sensitive and/or F sensitive genetic circuits are described wherein at least one molecular component is a U sensing genetic molecular component herein described and wherein at least one molecular component is a reportable molecular component, and/or a U-neutralizing molecular component (and in particular possibly one or more U-neutralizing genetic molecular components),
The U-sensitive and/or F sensitive genetic circuits further comprises an F-sensing riboswitch within at least one of the genetic molecular components of the genetic circuit in a configuration wherein the at least one of genetic molecular components of the genetic circuit is transcribed in presence of an effective amount of bioavailable fluoride.
In the U-sensitive and/or F- sensitive genetic circuit the reportable molecular component and/or the U-neutralizing molecular component are expressed when the genetic circuit operates according to the circuit design in presence of bioavailable U and/or bioavailable Fluoride. [0026] In particular, in several embodiments the biosensors and related genetic molecular components, genetic circuits, compositions, methods and systems herein described are configured for selective and sensitive detection, reporting and/or neutralizing UO2F2 a bioavailable environmental decomposition product of UF6 and toxic form of U typically produced during enrichment operations.
[0027] The UO2F2 biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described provide in several embodiments a selective, sensitive, portable, easy to use, high-throughput measurement and or neutralizing bioavailable UO2F2 , with little or no sample preparation required.
[0028] The UO2F2 biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described allow in several embodiments construction of consolidated bioremediators comprising bacterial systems that possess all the necessary components for deployment in environmental cleanup efforts, for example by coupling UO2F2 sensing with activation of one or more U-neutralizing components.
[0029] The UO2F2 biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described allow in several embodiments detection, reporting and/or neutralization of U in its bioavailable form UO2F2 with low cost approaches as various proteobacterial cells, such as Caulobacter, can be inexpensively grown to high densities as will be understood by a skilled person.
[0030] The UO2F2 biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described can be used in connection with various applications wherein detection and/or neutralizing of uranium is desired. For example, the U02F2biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described can be used in biodefense and in particular to be used for non-proliferation purposes, in environmental monitoring and/or cleanup by regulatory agencies or communities, and in mining in particular for toxicology and safety concerns, as well as diagnostic applications. Additional exemplary applications include uses of the UO2F2 biosensors, and related genetic molecular components, genetic circuits, compositions, methods and systems herein described in several fields including basic biology research, applied biology, bio- engineering, medical diagnostics, and in additional fields identifiable by a skilled person upon reading of the present disclosure.
[0031] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.
BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate one or more embodiments of the present disclosure and, together with the detailed description and the examples, serve to explain the principles and implementations of the disclosure.
[0033] Figure 1 shows a graph of pairwise comparison of fold-change in expression in Caulobacter crescentus of exemplary genes activated by Zn or U, analyzed using RNA-Seq. CCNA_01362 and CCNA_01353 (phytase) (dark grey data points shown as respectively labeled circles) are induced more than 100-fold by U but not induced by Zn. As such, the promoters regulating these genes, Pi36i and P1353 (Pphyt) respectively, represent promising exemplary promoters for use in a whole-cell U biosensor. urcA (dark grey data point indicated as labeled circle), a gene regulated by UzcRS (Figure 24A-C), is strongly induced by both U (Figure 24A) and Zn (Figure 23A).
[0034] Figure 2 shows graphs reporting determination of metal specificity of native U- responsive promoters in Caulobacter crescentus. The metal specificity of chromosomal P mcA- lacZ (Panel A), P phyt-lacZ (Panel B), and P nei-lacZ (Panel C) was determined by treating mid exponential phase cells with a range of concentrations of various metal salts for two hours before determining b-galactosidase activity using the method of Miller [1]. Cell growth was performed in peptone yeast extract (PYE) media supplemented with 50 mM MES pH 6.1. For Panel A, the relative expression was determined by normalizing the b-galactosidase activity observed by treatment with each metal to that of Zn. For Panels A and B, fold activation was calculated by dividing the activity with no added metal from the activity following metal exposure. Error bars represent the standard deviation calculated using a formula for propagation of standard error [2]. See Figure 23 for more comprehensive metal specificity plot for P urcA-lacZ.
[0035] Figure 3A shows schematics of two exemplary independent two component systems (TCS) that are sensitive to U in Caulobacter crescentus. Shown on the left side of Figure 3A is a schematic of the putative mechanism of action of the UzcRS TCS, which was characterized as a transcriptional activator of at least 40 operons in response to U/Zn/Cu [3]. Shown on the right side of Figure 3A is a schematic of the putative mechanism of action of the 1363/1362 TCS, wherein a histidine kinase 1363 (encoded by CCNA_01363) and a response regulator 1362 (encoded by CCNA_01362) activate promoters comprising the 1362 regulator direct repeat DNA binding site, e.g. Pphyt and Pi36i, in response to U. 1363 is membrane-bound and senses U either through a direct (U binding to 1363) or indirect mechanism. Upon U sensing, it is expected that 1363 autophosphorylates and then transphosphorylates 1362, activating 1362 for DNA binding of the 1362 direct repeat, e.g. in Pphyt or Pi36i. Possible additional stimuli for 1363/1362 remain unknown (indicated by the circled question mark).
[0036] Figure 3B shows a schematic of an exemplary U-sensitive AND gate that incorporates two independent points of uranyl sensing inputs (the two-component systems 1363/1362 AND UzcRS), which are both required to affect an output (such as a reportable molecular component and/or a U-neutralizing molecular component). The 1363/1362 two-component system is specifically activated by uranyl. The UzcRS two component system is activated for transcriptional regulation by uranyl, as well as zinc, copper and cadmium [3].
[0037] Figure 4 Panels A-C shows schematics illustrating the stepwise genetic engineering of an exemplary U-sensitive genetic circuit with incremental improvements from Panel A to Panel C to enhance specificity for U, resulting in a genetic circuit comprising an‘in series’ AND gate comprising two points of U-sensing by (1) Pphyt or Pi36i and (2) UzcRS two component system. Panel A shows a schematic of a U-sensitive genetic circuit comprising uzcRS under the control of the native P UZcR promoters PI and P2 and GFP expression under the control of UzcR-regulated promoter P1968. In this genetic circuit, binding of U directly or through indirect stimulation of UzcS causes the UzcS -mediated phosphorylation and activation of UzcR, leading to the homodimerization of UzcR and DNA binding of the UzcR dimer at the m_5 binding sites in the Pi968 promoter. The genetic circuit shown in Panel A only requires one point of U sensing, by UzcRS. As expected, this genetic circuit produces a high fluorescence signal in response to U, Zn, and Cu, as shown in Figure 5. Incremental improvements of this circuit are shown in Panel B and Panel C to enhance selectivity for U. Panel B shows a schematic of a U-sensitive genetic circuit where Pi and Pn are replaced with Pphyt or Pi36i such that uzcRS expression is now dependent on activation by these U- specific promoters. This construct requires two points of U sensing for reporter activation, (1) activation of uzcRS transcription by Pphyt or P i ;½ i and (2) stimulation of UzcRS transcriptional regulatory activity. This sensor shows greater signal in response to U compared to the genetic circuit shown in Panel A, as shown in Figure 5 Panel A. Importantly, in the genetic circuit shown in Figure 4 Panel B, Cu-sensing has been completely abolished while Zn induction with the range of inducing Zn concentrations narrowed compared to the UzcRS sensor alone; also, the ratio of the U signal output to that of Zn has been increased from 1.6 to 3.5 as shown in Figure 5 Panel B. Figure 4 Panel C shows a schematic of a U- sensitive genetic circuit where a negative feedback loop was incorporated into the circuit, whereby UzcR represses its own expression from Pphyt or Pi36i. Specifically, an m_5 UzcR binding site was placed downstream of the Ppi,yL or Pi36i transcription start site. This genetic circuit shows minimized basal expression of uzcRS, while maintaining strong responsiveness to U, and further shifted ratio of U response to that of Zn to 5.5 as shown in Figure 5 Panel C.
[0038] Figure 5 shows graphs of exemplary GFP reporter fluorescence produced by the U- sensitive genetic circuits shown in Figure 4, comprised in the host organism C. crescentus NA1000, upon exposure to U, Zn or Cu. Figure 5 Panel A shows graphs reporting quantification of GFP fluorescence produced by C. crescentus NA1000 comprising the U-sensitive genetic circuit shown in Figure 4 Panel A in response to exposure to 10 and 20 mM U (Figure 5 Panel A left graph), 10, 20 and 40 pM Zn (Figure 5 Panel A middle graph), or 80, 120 and 200 pM Cu (Figure 5 Panel A right graph), from 0 to 4 hours after exposure. Figure 5 Panel B shows graphs reporting quantification of GFP fluorescence produced by C. crescentus NA1000 comprising the U-sensitive genetic circuit shown in Figure 4 Panel B in response to exposure to 10 and 20 pM U (Figure 5 Panel B left graph), 10, 20 and 40 pM Zn (Figure 5 Panel B middle graph), or 80, 120 and 200 pM Cu (Figure 5 Panel B right graph), from 0 to 4 hours after exposure. Figure 5 Panel C shows graphs reporting quantification of GFP fluorescence produced by C. crescentus NA1000 comprising the U-sensitive genetic circuit shown in Figure 4 Panel C in response to exposure to 10 and 20 mM U (Figure 5 Panel C left graph), 10, 20 and 40 mM Zn (Figure 5 Panel C middle graph), or 80, 120 and 200 pM Cu (Figure 5 Panel C right graph), from 0 to 4 hours after exposure. Pphyt was used for the GFP fluorescence measurements shown in Figure 5 Panels B-C. Experiments with U were performed in modified M5G medium (10 mM PIPES, pH 7, I mM NaCl, I mM KC1, 0.05 % NH4C1, 0.01 mM Fe/EDTA, 0.2% glucose, 0.5 mM MgS04, 0.5 mM CaCh) supplemented with 5 mM glycerol-2- phosphate as the phosphate source (M5G-G2P). Cells were grown to mid-exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated U concentration. Experiments with Zn and Cu were performed in PYE.
[0039] Figure 6A and Figure 6B show schematics of two exemplary U-sensitive genetic circuits with‘in parallel’ AND gates comprising two points of U-sensing by (1) 1363/1362 two- component system (exemplified by U-sensitive transcriptional regulator direct repeat-containing promoters Pphyt or Pi36i) and (2) UzcRS two component system (exemplified by UzcRS- responsive promoter PUTCB). Figure 6A shows a schematic of an exemplary‘in parallel’ AND gate, comprised of an HRP AND gate system, in which hrpS is placed under the control of Pphyt or Pi36i and hrpR is placed under the control of PurcB, a promoter activated by UzcRS. In this U- sensitive genetic circuit, U exposure stimulates production of both HrpS and HrpR, leading to activation of Pi,rp[. and expression of GFP. Figure 6B shows a schematic of an exemplary‘in parallel’ AND gate, comprised of a tripartite GFP system, in which gfplO subunit is placed under the control of Pphyt or Pi36i, gfpll is placed under the control of PurcB, and gfpl-9 is placed under the control of the Caulobacter S layer promoter, P rSaA, which is a strong, constitutive promoter [4]. This system requires expression of subunits gfplO, gfpl l and gfpl-9 for tripartite GFP reporter assembly and fluorescence. Gfp-10 is shown fused to K1 and gfpl l is shown fused to El, wherein K1 and El are exemplary interacting protein partners comprised of oppositely charged coiled-coils [5]. When K1 and El interact, GFP10 and GFP11 self-associate with GFP1- 9 to constitute a functional GFP reporter. Other exemplary embodiments of the genetic circuit shown in Figure 6B comprise placement of gfpl-9 under the control of R^/Rqόΐ, or under the control of a different non-U-sensitive promoter, such as Pxyi that is responsive to xylose [6]. In addition, Figure 6C shows a more detailed version of Figure 6B, depicting how the two independent U sensing systems are integrated into the AND gate. [0040] Figure 7 shows schematics of regulatory sequences within Pphyt (Figure 7 Panel A), showing the sequence
CCCAAAGAGGGTGTGGCCCAAAGAGGGTGTGGATTTCTCTTCGCGCCACCCGTTTCG TCAGCCGGACGTCAGGTCCAGACGGCTAAGCTAGCTGCGA (SEQ ID NO: 203) and Pi36i (Figure 7 Panel B), showing the sequence
ATGTTCAGCGCCTGGTTACCGGCGATGGCGCGGTGTCAGCGTTCGGGCGTTGCGATG CGTCAGGAGCGTGTCAGGATGCCTGTGGAATCCTAAGCGC (SEQ ID NO: 204) with arrows above nucleotides indicating putative transcription start sites. In Figure 7 Panel C, a phylogenetic footprinting approach was used to construct a 1362 DNA-binding motif. An alignment of 26 Pphyt or Pi36i DNA sequences from exemplary members of the subclass Caulobacteridae, Bradyrhizobiaceae, Sphingomonadaceae, Hyphomicrobiaceae, and Rhodobacteracea are indicated, with larger letters representing a higher level of consensus between aligned sequences.
[0041] Figure 8 shows DNA sequences of full-length Pphyt (Panel A), Pphyt with a mutation of four nucleotides in the second direct repeat sequence (DR2, shown in bold, Panel B), Pphyt with a mutation of two nucleotides in the first direct repeat sequence (DR1, shown in bold, Panel C), Pphyt with a mutation of two nucleotides in the second direct repeat sequence (DR2, shown in bold, Panel D), a shortened Pphyt (Pphyt short, Panel E), full-length Pi36i (Panel F), Pi36i with a mutation of four nucleotides in the second direct repeat sequence (DR2, shown in bold, Panel G), P 1361 with a mutation of two nucleotides in the second direct repeat sequence (DR2, shown in bold, Panel H), and a shortened Pi36i (Pi36i short, Panel I). The sequence of the large tandem repeat (TR) in Pphyt is shown in uppercase, underlined. The direct repeat sequence that is likely bound by the U-sensitive transcriptional response regulator 1362 is shown in uppercase, italic (with direct repeat sequences underlined). Putative Transcription start sites (based on RNA-seq data) are shown in lowercase, underlined.
[0042] Figure 9 shows graphs reporting quantification of exemplary fluorescence levels of Pphyt- gfp (Figure 9 Panel A) and Pi36i-gfp (Figure 9 Panel B) variants in response to U at 10 mM or 25 pM in the host organism C. crescentus NA1000. In Figure 9 Panel A, fluorescence levels are shown for variants comprising full length Pphyt promoter (Full length), a shortened Pphyt (Short), Pphyt with a mutation of four nucleotides in the second direct repeat sequence (DR2 GTCA -> CAGT), Pphyt with a mutation of two nucleotides in the first direct repeat sequence (DR1, GT -> CA), and Pphyt with a mutation of two nucleotides in the second direct repeat sequence (DR2, GT -> CA). In Figure 9 Panel B, fluorescence levels are shown for variants comprising full length Pi36i promoter (Full length), a shortened Pi36i (Short), Pi36i with a mutation of four nucleotides in the second direct repeat sequence (DR2 GTCA -> CAGT), and Pi36i with a mutation of two nucleotides in the second direct repeat sequence (DR2, GT -> CA). Cells were grown to mid exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated U concentration. Fluorescence was quantified following a two- hour exposure of cells to each U concentration and normalized to the ODeoo.
[0043] Figure 10 illustrates the results of an electrophoretic mobility assay (EMSA) showing UrpR binding to wild type and mutant Pphyt fragments. The mutant Pphyt DNA (P1361) contains a GTCA -> CAGT mutation of DR2 (see Figure 7B). The assays were performed with 50 nM 6-FAM-labeled DNA and UrpR, phosphorylated with carbamoyl phosphate. The concentrations indicate the total UrpR used in the assay. A representative example of three biological replicates is depicted.
[0044] Figure 11 shows graphs reporting exemplary data corresponding to the exemplary U- sensitive genetic circuit in Figure 6 Panel B. Graphs reporting quantification of GFP fluorescence produced by C. crescentus NA1000 comprising the U-sensitive genetic circuit shown in Figure 6 Panel B with gfplO under the control of Pphyt (Panel A) or Pi36i (Panel B) in response to exposure to 20, 50, and 100 mM U (Figure 11 Panels A and B, upper graphs), or 10, 20 pM Zn and 6 pM Cu (Figure 11 Panels A and B, lower graphs). In this version of the circuit, gfpl-9 is controlled by PAV/ and gfpl-9 expression is induced with 10 mM xylose. Fluorescence output for both Pphyt and Pi36i sensor variants is plotted as a function of time following metal exposure and was normalized to the fluorescence of a strain lacking the UzcR regulator. Cells were grown to mid-exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated metal concentration.
[0045] Figure 12 shows a graph reporting data showing activity of an exemplary native signal amplifier module for UzcRS. The graph shows increase in fluorescence at various U concentrations in a CCNA_03499 mutant (+ UzcY) or in wild type cells (- UzcY), relative to conditions in absence of metal. Fluorescence was normalized to cell density (ODeoo). [0046] Figure 13 shows a schematic of an exemplary signal amplifier module incorporated within a U sensing circuit. The signal amplifier uzcY is placed under the control of the U- specific promoters R,,/, ,/R i % i such that signal amplification is restricted to conditions of U exposure.
[0047] Figure 14 shows an exemplary embodiment of a U-responsive AND gate design. Panel A shows a schematic of an example of an‘in series’ AND gate combined with an‘in parallel’ AND gate in a U sensitive genetic circuit. In the exemplary combined circuit, the P mcB promoter in the exemplary‘in parallel’ tripartite GFP AND gate is activated by UzcR, and expression of UzcR and UzcS is under transcriptional regulation of Pphyt in an‘in series’ AND gate. Grey arrows depict regulator modifications made to the base‘in parallel’ AND gate to generate the combined‘in series’,‘in parallel’ AND gate circuit. In the tripartite GFP system, gfplO subunit is placed under the control of Pphyt-short, gfpll is placed under the control of P UrcB, and gfpl-9 is placed under the control of the Caulobacter S layer promoter, P rsaA, which is a strong, constitutive promoter. [4] This genetic circuit was integrated into the chromosomal urcA locus and integrates regulatory input from the native, autoregulatory UzcRS and UrpRS TCS, depicted in the gray box. Panel B shows the fluorescence output profile of the core sensor and control variants lacking critical regulatory components over a range of U concentrations. Error bars represent the average of biological triplicates.
[0048] Figure 15 is Figure 1 from Newsome et al. (2014) [7] showing an exemplary Eh-pH diagram (which maps out possible stable (equilibrium) phases of an aqueous electrochemical system) for aqueous species in the U-O2-CO2-H2O system in pure water at 25 °C and 1 bar total pressure for åU = 10-8 M and a typical groundwater CO2 pressure of PCO2 = 10-20 bar [8]. UC, UDC and UTC represent the aqueous complexes UO2CO30, U02(C03)22- and U02(C03)34-· The position of the U02(c) solid solution boundary for åU = 10-8 M is stippled. The shaded area represents the range of conditions of common natural waters [9] as presented in Newsome et al., (2014) [7]
[0049] Figure 16A is from Figure 2 of Newsome et al. (2014) [7] showing a schematic illustrating exemplary mechanisms of microbe-uranium interactions, such as bioreduction [10-
13]. [0050] Figure 16B is from Figure 2 of Newsome et al. (2014) [7] showing a schematic illustrating exemplary mechanisms of microbe-uranium interactions, such as biomineralisation [14-16]
[0051] Figure 16C is from Figure 2 of Newsome et al. (2014) [7] showing a schematic illustrating exemplary mechanisms of microbe-uranium interactions, such as biosorption [17, 18].
[0052] Figure 16D is from Figure 2 of Newsome et al. (2014) [7] showing a schematic illustrating exemplary mechanisms of microbe-uranium interactions, such as bioaccumulation [19] as presented in Newsome et al., (2014) [7].
[0053] Figure 17 shows graphs reporting quantification of exemplary fluorescence levels of a shortened PPhyt-gfp (Figure 17 Panel A) and shortened Pi36i-gfp (Figure 17 Panel B) variants in response to U at 5 mM, 10 mM or 20 pM. In Figure 17 Panels A and B, fluorescence levels are shown for wild type C. crescentus and strains deleted for CCNA_01362 (1362 response regulator) and CCNA_01363 (1363 histidine kinase). U induction of both P phyt and P i % i is abolished by deletion of either CCNA_01362 (indicated by D CCNA_01362) or CCNA_01363 (indicated by D CCNA_01363), confirming that these genes encode the U sensitive transcriptional regulatory system required for U-responsive activation of the exemplary promoters Pphyt and Pi36i. Cells were grown to mid-exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated U concentration. Fluorescence was quantified following a two-hour exposure of cells to each U concentration and normalized to the ODeoo.
[0054] Figure 18 shows schematics illustrating exemplary tripartite GFP U-sensitive genetic circuits together with graphs reporting quantification of exemplary GFP reporter fluorescence produced by the respective genetic circuits under the conditions indicated. In particular, Figure 18 Panel A shows a schematic of the U-sensitive genetic circuit shown in Figure 6 Panel B and Figure 18 Panel B shows a schematic of a control circuit that incorporates input from only the UzcRS TCS, comprised in the host organism C. crescentus NA1000, upon exposure to U, Zn, Cu or Cd. Figure 17 Panel A shows a graph reporting exemplary quantification of GFP fluorescence produced by C. crescentus NA1000 comprising the U-sensitive genetic circuit shown in Figure 6 Panel B upon exposure to U, Zn, Cu and Cd. In this circuit, the expression of gfplO-Kl is initiated by the shortened 1363/1362-regulated Pphyt (Pphyt-short) and the expression of El-gfpl l is initiated by the UzcRS -regulated promoter PurcB . Figure 17 Panel B shows a graph reporting exemplary quantification of GFP fluorescence produced by C. crescentus NA1000 comprising a control U-sensitive genetic circuit that incorporates input only from UzcRS. Figure 17 demonstrates the enhanced selectivity of an exemplary‘in parallel’ AND gate comprising two points of U-sensing by (1) 1363/1362 and (2) UzcRS two component systems. Cells were grown to mid-exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated metal concentration. Fluorescence was quantified following a three-hour exposure of cells to each metal concentration and normalized to the ODeoo.
[0055] Figure 19 shows a graph reporting exemplary data indicating the limit of U detection for the U biosensor described in Figure 4 Panel B. Mid-exponential phase cells were washed twice in 10 mM Pipes pH 7 and then resuspended in 10 mM Pipes pH 7 containing uranyl nitrate. As shown in the graph, a linear response was observed for U concentrations in the low micromolar range.
[0056] Figure 20 shows a graph reporting exemplary data showing the ratio of the fluorescence output in response to 10 and 20 mM of U and Zn for the“In-parallel” AND-gate shown in Figure 6 Panel C and the combined“In-series” plus“in-parallel” genetic circuit shown in Figure 14. Cells were grown to mid-exponential phase in M5G-G2P media, washed once with fresh media, then resuspended in fresh media containing the indicated metal concentration. Fluorescence was quantified following a three-hour exposure of cells to each metal concentration and normalized to the OD600. The U/Zn ratio was calculated by dividing the normalized fluorescence with U by that with Zn.
[0057] Figure 21 shows a schematic showing exemplary U-sensitive genetic circuits having exemplary U-neutralizing outputs to allow U bioprecipitation or bioadsorption.
[0058] Figure 22 shows a schematic showing an exemplary AND gate comprised in a U- sensitive genetic circuit wherein an alpha-gall IP fusion and a lambda repressor-gal4 fusion are expected to be driven by a combination of PphytfP i36i and any UzcRS regulated promoter. [0059] Figure 23 shows in some embodiments the determination of metal specificity of P urcA· Panel (A): Time course of chromosomal P mcA-lacZ induction after treatment of early exponential phase cells with or without Zn. Cells were grown in PYE and the b-galactosidase activity at each time point is depicted. Error bars represent the standard deviation of biological triplicates. Panel (B): Metal specificity of P mcA was determined by treating mid-exponential phase cells with various metal cations in modified M5G media supplemented with 1.3 mM inorganic phosphate for two hours before determining b-galactosidase activity. Expression values were normalized to the level of expression with 50 mM Zn and error bars represent that standard deviation calculated using a formula for propagation of standard error [20]. The depicted metal concentrations are in units of mM except for CaCh and MgSCC that were added at mM concentrations.
[0060] Figure 24 shows an exemplary direct activation of P mcA by UzcRS according to some embodiments herein described. Panel (A): Time course of P mcA-lacZ expression following treatment with uranyl nitrate in M5G media supplemented with 5 mM glycerol-2-phosphate as the phosphate source media for wild type (WT), D uzcR and D uzcS. Panel (B): EMSA assay of UzcR binding to a 5' 6-FAM (Fluorescein)-labeled urcA promoter fragment. UzcR-P was phosphorylated with carbamoyl phosphate and the concentrations indicate the total UzcR used in the assay. Arrows depict shifted complexes. Panel (C): In vivo binding of UzcR to P urcA using ChIP-qRT PCR. Data are plotted as the fold enrichment at P urcA relative to the control region ( sodA ) for wild type (WT) with or without 40 pM Zn and D uzcR with 40 pM Zn.
[0061] Figure 25 shows an exemplary identification of the UzcR sequence recognition motif in some embodiments herein described. Panel (A): The 18-bp UzcR (m_5) sequence logo was constructed from the alignment of 49 UzcR boxes identified within the sequence regions bound by UzcR in vivo using MEME [21]. The sequence conservation (bits) is depicted by the height of the letters with the relative frequency of each base depicted by its relative height. Panel (B): Location of the predicted UzcR binding sites with respect to the previously determined Transcription Start Site (TSS) [22] for the directly activated operons. For urcA and urcB, the approximate TSSs as determined by tiled microarray analysis [23] were used. Note that multiple copies of the m_5 motif were found within some binding regions. For urcA and urcB, the approximate TSSs as determined by tiled microarray analysis [23] were used. The length of the line is representative of the length of the binding site with the line color denoting a directional orientation on the coding strand (black) or noncoding strand (gray). Panel (C): EMSA assays of UzcR-P binding to wild type and mutant P1968 fragments. Each UzcR half site was individually eliminated by mutation away from consensus (bolded nucleotides). The Assays were performed with 5' 6-FAM-labeled DNA and UzcR-P, generated by phosphorylation of UzcR with carbamoyl phosphate. The concentrations indicate the total UzcR-P used in the assay and arrow depicts the shifted complex. A representative example of three biological replicates is depicted Panel (D): Effects of mutations in each UzcR half site on CCNA_01968 promoter activity in wild type and A uzcR backgrounds. Promoter -gfpmut3 fusions with the wild type and mutant Pi968 fragments described in Panel (B) were constructed and fluorescence was quantified following a two-hour treatment with or without 20 mM Zn using a Biotek plate reader (ex: 480/ em: 516). The fluorescence signal was normalized to the OD600 and fold activation was calculated by dividing the normalized fluorescence in the presence of Zn by the fluorescence in the uninduced condition. Error bars represent that standard deviation calculated using a formula for propagation of standard error [20] This figure is taken from the Appendix B of U.S. provisional application No. 62/587,753 the disclosure of which is incorporated herein by reference in its entirety.
[0062] Figure 26 shows exemplary MarR regulators repressing expression of the membrane proteins UzcY and UzcZ according to some embodiments herein described. Panel (A): Operon diagram and RNA-seq expression profile of the autoregulatory marRi (top) and marRz (bottom) operons. The RNA-seq data for a A marRi AuzcR (top), A marRz AuzcR (middle) and A uzcR (bottom) is depicted for each operon. The black arrows depict the location of the P UZcY and P UZcz transcription start sites. Panel (B): Effect of marRi and marRz deletions on the expression of the uzcY, uzcZ and CCNA_02290 promoters. The Fluorescence of promoter-g p fusions was quantified at mid-exponential phase and normalized to the OD600. Error bars represent the standard deviation of biological triplicates.
[0063] Figure 27 shows exemplary effect of UzcRS regulator mutants in minimal media and U induction and TFInfer activity of UzcR in negative regulator mutants according to some embodiments herein described. Panel (A): Genes containing transposon insertions that led to high basal P cA-lacZ activity were deleted and the resulting strains were transformed with plasmid-bome P mcA-lacZ (pNJH123). Cells were grown to mid-exponential phase in M5G supplemented with glycerol-2-phosphate, washed once with fresh media, then resuspended in fresh media containing 50 mM U. b-galactosidase assays were performed following 1 h exposure to U. Error bars represent the standard deviation of triplicate measurements. Panel (B): The activity of UzcRS was inferred from the log2 fold change values for 37 UzcR regulon members determined in each mutant relative to WT. PurcA-lacZ expression data from Figure 1A is plotted for comparison.
[0064] Figure 28 shows diagrams illustrating the results of experiments showing that transposon insertions within CCNA_01362 and CCNA_01363 abrogate U-dependent induction of Pphyt. Pphyt -lacZ activity was assayed in strains containing transposons inserted within CCNA_01362 and CCNA_01363 using b-galactosidase assays. Error bars represent the standard deviation of triplicate measurements.
[0065] Figure 29 shows diagrams illustrating the results of experiments showing metal selectivity of UrpRS and UzcRS in an exemplary embodiment. Panel A) Functional independence of UrpRS and UzcRS. U response curves of Pphyt-short-g^? and P urcB-gfp reporters were determined by quantifying fluorescence following a two-hour exposure to a range of U concentrations in wild type C. crescentus and strains deleted for uzcS, urpS and the cognate response regulator. Error bars represent the average of biological triplicates. Metal selectivity profiles for Pphyt-short-g^? (panel B) and P urcB-gfp (panel C). Reporter fluorescence was quantified following exposure to 16 metals. The furthest right plot depicts representative data for metals that failed to induce either promoter. Raw fluorescence values can be found in Table 10. Error bars represent the average of biological triplicates.
[0066] Figure 30 shows diagrams illustrating the results of experiments showing in an exemplary embodiment functional independence of UrpRS and UzcRS. Fluorescence of a Pphyt- short -gfp and PurcB-gfp were quantified following a two-hour exposure to a range of U concentrations in wild type C. crescentus and strains deleted for uzcR and urpR. Error bars represent the average of biological triplicates.
[0067] Figure 31 illustrates a schematic of U-sensing AND gate configuration that was integrated at the chromosomal urcA locus. Details of AND gate construction and chromosomal integration are outlined in the methods sections. [0068] Figure 32 shows diagrams illustrating the results of experiments showing in an exemplary embodiment U sensing AND gate variants with different mechanisms of gfpl-9 expression. Panel A shows the fluorescence output profile of the core sensor and a sensor variant with gfpl-9 expression driven by Pphyt following a three-hour exposure to the indicated U concentration. Panel B shows the fluorescence output of a U-sensor variant with gfpl-9 expression driven by the xylose-inducible promoter P xyi as a function of time.
[0069] Figure 33 shows a graph reporting that core sensor is not responsive to nitrate. Given the use of a uranyl nitrate stock in HNO3 (100 mM UO2NO3 in 100 mM HNO3) to characterize the U sensing performance of the core sensor, control experiments were conducted with HNO3 nitrate alone. Relevant concentrations of HNO3 failed to induce fluorescence, confirming that nitrate alone is not responsible for the core sensor output signal.
[0070] Figure 34 shows in a graph the U response curves for the core sensor and control variants in which the expression of gfplO-Kl and El-gfpl l components is driven by UzcRS or UrpRS alone. Fluorescence was quantified following a three-hour exposure to a range of U concentrations. The data were fit with a Hill equation to determine the U concentration that yields half-maximal fluorescence induction and the hill slope as an indicator of U sensitivity. Data points with diminished fluoresescence output were not included in the curve fitting.
[0071] Figure 35 shows diagrams illustrating the results of experiments showing metal selectivity profile of a core sensor. Fluorescence was quantified following a two-hour exposure to each metal. Error bars represent the average of biological triplicates.
[0072] Figure 36 shows graphs reporting incorporation of a signal amplifier module within the core sensor circuitry. Panel A: schematic for the integration of the UzcY signal amplifier within the U-sensing AND gate. In this variant, uzcY is placed under the control of the U-specific promoter PPhyt-short such that the signal amplification is restricted to conditions of U exposure. Panel B: The fluorescence output of the core sensor and signal amplifier variants following a three-hour exposure to a range of U concentrations. The data were fit with a Hill equation to determine the U concentration that yields half-maximal fluorescence induction and the hill slope as an indicator of U sensitivity. Signal amplifier data points with diminished fluoresescence output were not included in the curve fitting. Panel C: The fluorescence output of the core sensor and signal amplifier variants following a three-hour exposure to a range of Zn, Pb, and Cu concentrations. Error bars represent the average of biological triplicates.
[0073] Figure 37 shows a zoomed in version of the U/Zn/Cu/Pb response curves for the core sensor, control variants, and uzcY amplifier variants. Fluorescence was quantified following a three-hour exposure to a range of metal concentrations. The data represent a zoomed in version of data depicted in Figure 40 and Figure 36C. Error bars represent the average of biological triplicates.
[0074] Figure 38 shows graphs reporting fluorescence output of sensor variants in ground water samples without nutrient supplementation. (Panel A) Fluorescence output of the core sensor and control variants lacking critical regulatory components as a function of time in three distinct site 300 samples without nutrient supplementation. (Panel B) Fluorescence output of the core sensor and a control strain lacking a UrpR binding site as a function of time in three distinct site 300 samples supplemented with glucose (0.2%), orthophosphate (1 mM), and ammonium chloride (0.05%). The plot legend is depicted below the plots.
[0075] Figure 39 shows detection of U in ground water samples in an exemplary embodiment. Panel A: Fluorescence output of the core sensor and control variants lacking critical regulatory components as a function of time in three distinct site 300 samples supplemented with glucose (0.2%), glycerol-2-phosphate (5 mM), and ammonium chloride (0.05%). Ground water samples from W-815-2621, W-812-01, and W-6C are referred to as 1, 2, and 3, respectively, in the text. The plot legend is depicted below the plots. Panel B: The fluorescence of the core sensor and control variants was quantified six hours after exposure to site 300 samples supplemented with uranyl nitrate and the nutrients described in panel A. The x-axis concentrations represent the total U concentration in the ground water sample after addition of U in 1 mM increments.
[0076] Figure 40 shows a metal selectivity profile of core sensor and control variants. The top panel depicts simplified schematics of the core sensor and control variants in which the expression of gfplO-Kl and El-gfpl l components is driven by UzcRS or UrpRS alone. The bottom panel depicts the U/Zn/Cu/Pb response curves for the core sensor and control variants. Fluorescence was quantified following a three-hour exposure to a range of metal concentrations. Error bars represent the average of biological triplicates. [0077] Figure 41 shows a schematic illustration of an exemplary transducer device configured to convert an optical signal from a biosensor herein described into an electrical current using inexpensive, commercially available components (e.g., a blue LED for excitation, optical filters configured for excitation/emission of GFP, and a photodiode for detection).
[0078] Figure 42 shows the output voltage as a function of U concentration at three different currents using the exemplary transducer device depicted in Figure 41. Measurements were taken two hours after uranium exposure.
[0079] Figure 43 shows the conserved nucleotides in naturally occurring Fluoride sensing riboswitchs from a gapped alignment of the 2138 fluoride riboswitch sequences from the Rfam database
[0080] Figure 44 shows a schematic illustration of a design of a bacteria-based sensor of uranyl fluoride products.
[0081] Figure 45. shows a schematic illustration of a design to integrate and optimize a fluoride- detection capability in C. crescentus to provide a UCLFibioscnsor herein described.
[0082] Figure 46 shows a schematic representation of exemplary single output configurations of a UO2F2 biosensor herein described. In particular Figure 46 Panel A describes a schematic of uranium- and fluoride- sensing components integrated in series. The promoter can be any UzcR- or UrpR-regulated promoter (defined in prior patent app) such that U-dependent transcription is initiated by UzcR or UrpR. In this circuit, transcription will be prematurely terminated by the fluoride sensing riboswitch in the absence of fluoride (i.e., not detectable output). Binding of fluoride to the riboswitch will mediate transcriptional read-through and ultimately, production of the reporter. Figure 46 Panel B describes a schematic of uranium- and fluoride-sensing components integrated in series where U-dependent transcriptional activation requires the function of both UrpR and UzcR. This sensor will provide greater selectivity for uranium compared to the sensors described in panel A. However, in this configuration, the sensitivity for U would be limited by the UzcRS component. Figure 46 Panel C describes a schematic of the integration of the fluoride riboswitch within the AND gate circuit such that expression of component three requires fluoride exposure. In this configuration, reconstitution of GFP fluorescence requires activation of the two uranium-responsive pathways and fluoride binding to the fluoride riboswitch. Transcription of component three can theoretically be controlled by any constitutive promoter. The assignment of each component with the given regulatory promoter is arbitrary and easily swapped. For example, the fluoride riboswitch could be used to control expression of component one or two.
[0083] Figure 47 shows a schematic representation of an exemplary dual output configurations of a UO2F2 biosensor herein described.
[0084] Figure 48A is Figure 1A of US 9, 580,713 and, as indicated in US 9, 580,713 shows a schematic representation of “the consensus sequence and structural model based on the comparison of 2188 representatives from bacterial and archaeal species”. As indicated in US 9, 580,713 incorporated herein by reference in its entirety“PI, P2, P3 and pseudoknot labels identify base-paired substructures. Note that the bottom of PI carries a possible G-C pair. However, because noncomplementary nucleotides occur in these positions in some representatives, these nucleotides are depicted as unpaired.”
[0085] Figure 48B is Figure IB of US 9, 580,713 and, as indicated in US 9, 580,713 shows the ”[s]equence and secondary structure model for the WT 78 Psy RNA (SEQ ID NO:349). Numbers 1, 2, 3, 4, 5, and 6depict the sites of the in-line probing analysis; results presented in C. The two G residues preceding nucleotide 1 5 were added to facilitate RNA production by in vitro transcription.”
[0086] Figure 48C shows a schematic representation of a consensus secondary structure for F- sensing riboswitches in the sense of the disclosure from https://rfam.org/family/RF01734#tabview=tab0
[0087] Figure 48D shows a schematic representation of an exemplary fluoride sensing riboswitch herein described which correspond to a grayscale version of of schematic reported in https://rfam.xfam.org/family/RF01734#tabview=tab3 at the filing date of the present application.
[0088] Figure 49 shows charts reporting the effect of crcB deletion on the growth of Caulobacter crescentus in the presence of NaF and NaCl. Figure 49 panel A reports results showing growth in Caulobacter crescentus WT at NaF concentrations above 4 mM and Figure 49 panel B reports results showing growth in Caulobacter crescentus AcrcB at NaF concentrations above 126.5 mM,. Figure 49 panel C reports results showing growth in Caulobacter crescentus WT at NaCl concentrations above 4 mM. Figure 49 panel D reports results showing growth in Caulobacter crescentus AcrcB at NaCl concentrations above 4 mM.
[0089] Figure 50 shows a schematic and charts illustrating the result of testing the function of fluoride riboswitches from three different bacteria in Caulobacter crescentus. Figure 50 Panel A shows a schematic configuration of the exemplary construct used to perform the testing using b- galactosidase as a reporter. Figure 50 Panel B shows a chart reporting the b-galactosidase activity for the Sphingomonas MM-1 in WT and the crcB deletion strain over a range of NaF concentrations Figure 50 Panel C shows a chart reporting the b-galactosidase activity for the Pseudomonas syringae riboswitch in WT and the crcB deletion strain over a range of NaF concentrations and Figure 50 Panel D shows a chart reporting the b-galactosidase activity for the Sphingomonas 67-36 riboswitch in WT and the crcB deletion strain over a range of NaF concentrations.
[0090] Figure 51 shows a schematic and charts illustrating fluoride responsiveness of fluoride riboswitches from three different bacteria tested in C. crescentus using the xylose inducible promoter. Figure 51 Panel A shows a schematic representation of a construct wherein each of three tested riboswitches was included in construction between to the xylose inducible promoter (Pxyl) and the reporter mCherry. Figure 51 Panel B shows a chart reporting the effect of xylose administration ad different concentration on promoter activity. Figure 51 Panel C shows a chart reporting the effect of 200 uM xylose administration on mCherry expression in the constructs including each of the three riboswitches as indicated. Figure 51 Panel D shows a chart reporting the effect of 2 mM xylose administration on mCherry expression in the constructs including each of the three riboswitches as indicated. Figure 51 Panel E shows a chart reporting the effect of 200 uM xylose administration on mCherry expression in the constructs including the Sphingomonas MM-1 fluoride riboswitch in C. Crescentus WT and C. Crescentus with a crcB deletion. Figure 51 Panel F shows a chart reporting the effect of 2 mM xylose administration on mCherry expression in the constructs including the Sphingomonas MM-1 fluoride riboswitch in C Crescentus WT and C. Crescentus with a crcB deletion. [0091] Figure 52 shows a schematic representation of the primary genetic parts involved in the construction of a fluoride responsive reporter.
[0092] Figure 53 shows an alignment of RNA encoded by 287 sequences (SEQ ID NO: 1989- 1998 and SEQ ID NO: 2232 to SEQ ID NO: 2508) from the 2017 exemplary F-sensing riboswitch sequences shown in Appendix I(SEQ ID NO: 205-1988 and SEQ ID NO: 1999- 2231) wherein the nucleotide sequences shown are sequences forming a same structure in corresponding aligned sequences and the symbols“—” indicates gaps between the aligned sequences shown.
[0093] Figure 54 show the effect of translational fusion length on Sphingomoms MM- 1 fluoride riboswitch function
[0094] Figure 55 shows charts reporting the testing of an in -series uranyi fluoride sensing circuit. Figure 55 Panel A shows fluoride detection performed with a control, MM-1 fluoride sensing circuit with native crcB promoter not responsive to U. Figure 55 Panel B shows fluoride detection performed with a uranyi fluoride sensing circuit constructed by combining the UrpRS- responsive Pphyt promoter with the MM- 1 riboswitch. Figure 55 Panel€ shows the fluorescence of the uranyi fluoride sensing circuit in the presence of U alone, F alone, and both U and F. Tests-were performed with F-, added as NaF, and uranyi, added as uranyi nitrate.
APPENDIX I
[0095] The accompanying Appendix I which is incorporated into and constitute a part of this specification, illustrates one or more fluoride sensing riboswitch that can be used in any one of the UO2F2 biosensors of the present disclosure which are also reported in the enclosed Sequence Listing from SEQ ID NO: 205 to SEQ ID NO: 1988 and SEQ ID No: 1999 to SEQ ID NO: 2231. In particular the sequences from Appendix I and Sequence Listing from SEQ ID NO: 205 to SEQ ID NO: 1988 and SEQ ID No: 1999 to SEQ ID NO: 2231 are sequences from Rfam database reported with their related Genome file ID, bacteria and additional information concerning the position of the sequence in the genome of the bacteria as will be understood by a skilled person Appendix I together with the detailed description section, the Example section and the Drawings, serve to explain the principles and implementations of the disclosure. Other features, objects, and advantages will be apparent from the entire description and drawings, and from the claims.
DETAILED DESCRIPTION
[0096] Provided herein are UO2F2 biosensors and related U-sensing and/or F sensing genetic molecular components, genetic circuits, compositions, methods and systems which in several embodiments can be used to detect, report and/or neutralize U and in particular bioavailable Uranium which is produced in connection with enrichment programs UO2F2.
[0097] The term“bioavailable” as used herein refers to a molecule in particular a soluble molecule that is able to cross an organism's cellular membrane from the environment, or is otherwise able to exert a biological effect on an organism, if the organism has access to the molecule. In particular, with regard to a toxic molecule, a bioavailable toxic molecule is a toxic molecule that is able to exert toxicity on an organism contacted with the organism and/or with a toxic molecule sensing system of the organism. In some scenarios, the bioavailability can be inferred based on toxicity or activation of a tress response in an organism as will be understood by a skilled person. Thus, the term“bioavailable U” as used herein refers to a soluble molecular form of U that can cross an organism's cellular membrane from the surrounding environment or is otherwise able to exert a biological effect on an organism, e.g. following contact with the organism and/or with an organism U-sensing system.
[0098] In particular, the term“bioavailable uranium” comprises uranyl ion, which has a linear structure with short U-0 bonds, indicative of the presence of multiple bonds between uranium and oxygen and can bind four or more ligands in an equatorial plane. The uranyl ion forms many complexes, particularly with ligands that have oxygen donor atoms. Complexes of the uranyl ion are important in the extraction of uranium from its ores and in nuclear fuel reprocessing. As would be understood by persons skilled in the art,‘naked’ or‘uncomplexed’ uranyl oxycation is a bioavailable form of U. In contrast, for example, uranyl oxycation complexed with inorganic phosphate is not considered to be bioavailable.
[0099] The term“uranyl oxycation” as used herein refers to the predominant form of U in oxygenated environments, comprising the +6 oxidation state (UO+ 2 2.), which has high chemical toxicity [24]. The US Environmental Protection Agency’s maximum contaminant limit for U in drinking water is 30 pg/L (-0.13 mM), however, groundwater concentrations in the US frequently exceed this limit [25, 26].
[00100] The UO2F2 biosensors and related U-sensing and/or F-sensing genetic molecular component, gene cassettes, genetic circuits, compositions, methods and systems described herein can be used in several embodiments to detect and report and/or neutralize bioavailable U, and in particular UO2F2 which is a derivative of UF6.
[00101] The term“UF6” or“uranium hexafluoride” as used herein indicates is a compound used in the process of enriching uranium, which produces fuel for nuclear reactors and nuclear weapons. In particular Uranium hexafluoride (UF6) is almost always produced as a precursor in any U-enrichment operation, is routinely released during the conversion process, and is not expected to occur naturally in the environment [27, 28] As such, environmental detection of UF6— or the more stable hydrolysis product UO2F2, which is rapidly formed when atmospheric UF6 reacts with water vapor [29] . does strongly suggest an enrichment program.
[00102] In particular, conversion facilities, which produce the UF6 precursor that is almost always required for the production of highly enriched uranium, represent a likely source of leaked UF6 as a consequence of the higher-than atmospheric pressures employed. [27, 28]
[00103] When gaseous UF6 is released into the atmosphere, it is rapidly hydrolyzed by ambient moisture to form an UO2F2 aerosol (and HF gas) that is dispersed and ultimately deposited on vegetation, soil, or an aquatic system [29, 30]. UO2F2 is expected to have reasonable stability in the environment such that its detection may be feasible for months after release [27, 31].
[00104] UO2F2 aerosols, which are rapidly formed when atmospheric UF6 reacts with water vapor [29], are expected to be deposited on vegetation, soil, or into aquatic systems [29, 30]. While UO2F2 aerosols are expected to have reasonable stability in low moisture environments [27, 31], UO2F2 exhibits relatively high solubility in aqueous environments (up to 2 M [32]);
[00105] A solution-based detection approach is therefore expected to be most relevant for UO2F2 monitoring in aquatic systems or when coupled with a sampling device that collects and solubilizes UO2F2 aerosols. Additionally, the rationale for separate uranium and fluoride detection components is supported by the solution chemistry of UO2F2 and prior studies on bacterial uranium interactions.
[00106] The physicochemical form, or speciation, of UO2F2, is dependent on the geochemical conditions [33, 34]. In particular, at pH below 5, the naked uranyl oxycation (U02 2+) and uranyl fluoride species (U02F+, UO2F2, UO2F3 ) are expected to predominate [33]. At circumneutral (6.5-7.5) and alkaline pH, uranyl hydroxide and/or carbonate species are expected to predominate with concomitant formation of the anion, F [33, 34]. Under those conditions, uranium and fluoride are expected to largely exist as separate species under geochemical conditions most relevant to aqueous environmental sampling (pH 6-8 range).
[00107] In addition, studies of bacterial uranium interactions suggest that cell surface functional groups could have a dissociative effect on uranyl fluoride species, yielding detectable forms of both components; cell surface functional groups have been shown to compete with environmental ligands (e.g., carbonate and citrate) for uranyl coordination, effectively freeing the coordinating ligand [35, 36].
[00108] Insoluble forms of uranyl— for example uranyl phosphate minerals formed when uranyl nitrate is added to solutions containing high orthophosphate levels— are not detected by the biosensor. This is an advantageous feature for environmental detection since natural U commonly occurs in the form of insoluble U minerals [37] and aqueous phosphate concentrations are typically very low (<10 ppb) [34].
[00109] U02F2biosensor herein described are bacteria-based U02F2-sensor, which are engineered to comprise a fluoride- sensing riboswitch in combination a U-sensing genetic molecular components, U neutralizing genetic molecular component, reportable genetic molecular components, and/or additional components possibly configured in genetic circuits directed to detect and/or neutralize U in presence of bioavailable Uranium and Fluoride.
[00110] In particular, the UO2F2 biosensors herein described are whole-cell biosensors comprising a genetically engineered bacterial cell.
[00111] The term“bacterial cell”, bacteria” used herein interchangeably with the terms“cell” or “host” indicates a large domain of prokaryotic microorganisms. The term“prokaryotic” is used herein interchangeably with the terms“cell” or“host” and refers to a microbial species which contains no nucleus or other organelles in the cell. Exemplary prokaryotic cells include bacteria. Typically, a few micrometers in length, bacteria have a number of shapes, ranging from spheres to rods and spirals, and are present in several habitats, such as soil, water, acidic hot springs, radioactive waste, the deep portions of Earth's crust, as well as in symbiotic and parasitic relationships with plants and animals. Bacteria in the sense of the disclosure refers to several prokaryotic microbial species which comprise Gram-positive bacteria, Proteobacteria, Cyanobacteria, Spirochetes and related species, Planctomyces, Bacteroides, Flavobacteria, Chlamydia, Green sulfur bacteria, Green non-sulfur bacteria including anaerobic phototrophs, Radioresistant micrococci and related species, Thermotoga and Thermosipho thermophiles as would be understood by a skilled person. More specifically, the wording “Gram positive bacteria” refers to cocci, nonspomlating rods and spomlating rods, such as, for example, Actinomyces, Bacillus, Clostridium, Corynebacterium, Erysipelothrix, Lactobacillus, Listeria, Mycobacterium, Myxococcus, Nocardia, Staphylococcus, Streptococcus and Streptomyces.
[00112] The term“proteobacteria” as used herein refers to a major phylum of Gram-negative bacteria. Many move about using flagella, but some are nonmotile or rely on bacterial gliding. As understood by skilled persons, taxonomic classification as proteobacteria is determined primarily in terms of ribosomal RNA (rRNA) sequences. The Proteobacteria are divided into six classes, referred to by the Greek letters alpha through epsilon and the Acidithiobacillia and Oligoflexia, including alphaproteobacteria, betaproteobacteria and gammaproteobacteria as will be understood by a skilled person.
[00113] The term“alphaproteobacteria” as used herein refers to bacteria identifiable by those skilled in the art in the phylogenetic Class Alphaproteobacteri, in the Phylum Proteobacteria. As understood by those skilled in the art, Alphaproteobacteria is a diverse taxon and comprises several phototrophic genera, several genera metabolising Cl-compounds (e.g., Methylobacterium spp.), symbionts of plants (e.g., Rhizobium spp.), endosymbionts of arthropods ( Wolbachia ) and intracellular pathogens (e.g. Rickettsia). As understood by those skilled in the art, taxonomic classification of alphaproteobacteria can be identified by reference to publicly available online databases such as the List of Prokaryotic names with Standing in Nomenclature (LPSN) and National Center for Biotechnology Information (NCBI) and the phylogeny is based on 16S rRNA-based LTP release 106 by 'The All-Species Living Tree' Project. The Class Alphaproteobacteria is divided into three subclasses Magnetococcidae, Rickettsidae and Caulobacteridae [38]. In particular, the Caulobacteridae is a subclass composed of the orders Holosporales, Rhodospirillales, Sphingomonadales, Rhodobacterales, Caulobacterales, Rhizobhiales, Kiloniellales, Kordiimonadales, Parvularculales and Sneathiellales.
[00114] The term“betaproteobacteria” as used herein refers to a class of gram-negative bacteria, and one of the classes of the phylum Proteobacteria. [39] The Betaproteobacteria comprise more than 75 genera and 220 species of bacteria identifiable by persons skilled in the art. [40] Seven orders of betaproteobacteria have been described: Burkholderiales, Hydrogenophilales, Methylophilales, Neisseriales, Nitrosomonadales, Rhodocyclales, and Sulfuricellales. Examples of Betaproteobacteria genera comprise Bordetella, Ralstonia, Neisseria and Nitrosomonas, among others identifiable by skilled persons. While many Betaproteobacteria identifiable by skilled persons are found in environmental soil and water, others are obligate pathogens and can cause disease in a variety of hosts. Some members of betaproteobacteria can cause disease in various eukaryotic organisms. Several cause diseases in humans, such as members of the genus Neisseria: N gonorrhoeae and N. meninngitides which cause gonorrhea and meningitis respectively, as well as Bordetella pertussis which causes whooping cough. Other members infect plants, such as Burkholderia cepacia which causes bulb rot in onions as well as Xylophilus ampelinus which causes necrosis of grapevines. [40]
[00115] The term“gammaproteobacteria” as described herein refers to a class of gram-negative bacteria, and one of the classes of the phylum Proteobacteria. As would be identifiable by skilled persons, exemplary taxonomic orders, families and genera belonging to the class gammaproteobacteria comprise Acidithiobacillus, Xanthomonadales, Chromatiales, Methylococcus, Beggiatoa, Legionellales, Ruthia, Vesicomyosocius, Thiomicrospira, Dichelobacter, Francisella, Moraxellaceae, Alcalinovorax, Saccharophagus, Reinekea, Oceanospirillaceae, Marinobacter, Pseudomonadaceae, Aeromonas, Vibrionales, Pasteurellales, and Enterobacteriales among others. A number of bacteria have been described as members of gammaproteobacteria, but have not yet been assigned an order or family. These comprise bacteria of the genera Alkalimarinus, Alkalimonas, Arenicella, Gallaecimonas, Ignatzschineria, Litorivivens, Marinicella, Methylohalomonas, Methylonatrum, Plasticicumulans, Pseudohongiella, Sedimenticola, Thiohalobacter, Thiohalomonas, Thiohalorhabdus, Thiolapillus, and Wohlfahrtiimonas among others identifiable by skilled persons. Other examples of gammaproteobacteria genera comprise Escherichia, Shigella, Salmonella, Yersinia, Buchnera, Haemophilus, Vibrio, and Pseudomonas, among others identifiable by skilled persons. Some members of gammaproteobacterial are pathogenic in humans, for example some strains of the species Salmonella spp., Yersinia pestis, Vibrio cholerae, Pseudomonas aeruginosa, and Escherichia coli, among others identifiable by skilled persons. Some members of gammaproteobacteria are pathogenic in plants, such as Xanthomonas axonopodis pv. citri, Pseudomonas syringae pv. actinidiae, and Xylella fastidiosa, among others identifiable by skilled persons.
[00116] In some embodiments, the UOiFibioscnsors herein described are whole-cell biosensors comprising a genetically engineered alphaproteobacterial cell of the subclass Caulobacteridae. In particular in some embodiments, the U biosensors described herein can comprise a cell of any genus, species and/or strain of Caulobacteridae identifiable by those skilled in the art. Exemplary Caulobacteridae that can be used in U-biosensors herein described comprise species of the Families Bradyrhizobiaceae, Sphingomonadaceae , Caulobacteraceae, Hyphomicrobiaceae and Rhodobacteraceae which include species naturally comprising, as well as others identifiable by persons skilled in the art.
[00117] In particular, in some embodiments, UC Fi-biosensor s herein described can comprise species from the order Caulobacterales, the family Caulobacteraceae , the genus Caulobacter and the species Caulobacter crescentus which is described herein as one of the representative species of the subclass Caulobacteridae.
[00118] In some embodiments of the UC Fi-biosensors herein described, the bacterial cell of the U-biosensor is capable of natively and/or heterologously expressing a U-sensitive histidine kinase 1363, and cognate response regulator 1362.
[00119] The term“histidine kinase P1363” or“UrpS” as used herein refers to a histidine kinase having the amino acid sequence MS GGS LRWRLIVGGMLAILA AL A V A WL AMT WLFERHI VRRET ADLTRAGQ VL V AGLR LEPN G AP VID ATLS DPRLS KA AGGFY W Q V STTSGSERSVS LWDQ ALKPPQT AP AEGW S S RIA AGPFDDR VLL VERS VRPDRDGP A VLIQ V AS DEKVLRA ARREF GRELAIFLGGLW AIL S G A A ALQ V VLGLS PLTRVR ADLARLRKS PS ARMS LDHPREIAPLAE AIN AL AE ARE ADL ARARRR AGDLAHS LKTPL A ALS AQS RRAREDG A V A A ADGLD A AIAS V A A ALE AEL AR ARAAAAREAVFAAETAPLAVAERLVAVLERTADGERLIFDIDVPADLKAPASEDVVTE MLGALIENAARHARRQVRISGAVVGQGAVLIVEDDGPGLDKGRAEAALARGARLDEA GPGHGLGLAIVRD LAE AS G A VLS MDRGDLGGLR AM VS WT APG AGP (SEQ ID NOG), found in C. crescentus NA1000, or a sequence that when aligned with sequence SEQ ID NO: 3 has a BLAST score between 240 and 300, between 300 and 500, or preferably between 500 and 800, or more preferably over 800 but less than 100% homology or even more preferably having a BLAST Score of 851 and 100% homology with the sequence SEQ ID NO: 3.
[00120] The term “U sensitive response regulator 1362” or “response regulator 1362” or “UrpR” refers to a response regulator having amino acid sequence
MMR AL V VEDDP V V GPDLAKALS AS GFV VDI ARDGED AS FKGE VED Y AL V VLDLGLPR LDGLSVLRRWRANDRAFPVLILSARGDWTEKVEGIEAGADDYLAKPFEMGELLARARG L VRRA AGRT S P VIG AGRL ALDTRRMS ATLDG APIRLS PLEFRLLDCL AHNPGRA V S AGE L AEQLY G V ADT ADTN AIE AL V ARLRRKIG AD VIETR RGFGYLLAGGTA (SEQ ID NO: 4) of C. crescentus NA1000 or a sequence that when aligned with sequence SEQ ID NO: 4 has a BLAST score between 200 and 250, or preferably between 250 and 300, or more preferably over 300 but less than 100% homology or even more preferably having a BLAST Score of 429 and 100% homology with the sequence SEQ ID NO: 4.
[00121] The term “BLAST” or “Basic Local Alignment Search Tool” is an algorithm for comparing primary biological sequence information, such as the amino-acid sequences of proteins or the nucleotides of DNA sequences. A BLAST search enables a researcher to compare a query sequence with a library or database of sequences, and identify library sequences that resemble the query sequence above a certain threshold. Accordingly, BLAST or Basis Local Alignment Search Tool uses statistical methods to compare a DNA or protein input sequence, also referred to as a query sequence to a database of nucleotide and protein (subject sequences) and returns sequences hits that have a level of similarity to the query sequence ranked based on the score.
[00122] The term“score” in the context of sequence alignments, indicates a numerical value that describes the overall quality of an alignment. Higher scores correspond to higher similarity and lower scores correspond to lower similarity. The score scale depends on the scoring system used for conducting the sequence alignment.
[00123] A BLAST score, also referred to as bit score or max score in the BLAST output is a normalized score with respect to the scoring system provided by the BLAST algorithm. The BLAST score defines the highest alignment score of a set of aligned segments from the same subject (database) sequences. The score is calculated from the sum of the match rewards and the mismatch, gap open an extend penalties independently for each segment. The BLAST score normally gives the same sorting order as the expect value (E value) in the BLAST alignment output.
[00124] A BLAST score can be obtained using BLAST software suite at the NCBI website and related references and in particular at the website https://blast.ncbi. nlm.nih.gov/Blast.cgi?PROGRAM=blastp&PAGE_TYPE=BlastSearch&LINK _LOC=blasthome at the date of filing of the present disclosure, as will be understood to a person skilled in the art.
[00125] In embodiments herein described the“histidine kinase P1363” or“UrpS” and the “response regulator 1362” or“UrpR” typically form a two-component system herein also indicated as“1363/1362 TCS”, UrpRS TCS” or“UrpRS”.
[00126] The term“two component system” as used herein refers to a stimulus-response coupling mechanism that allows organisms to sense and respond to changes in many different environmental conditions [41]. Two-component systems typically consist of a membrane-bound histidine kinase that senses a specific environmental stimulus and a corresponding response regulator that mediates the cellular response, mostly through differential expression of target genes [42]. Although two-component signaling systems are found in all domains of life, they are most common in bacteria, particularly in Gram-negative and cyanobacteria [43]. Two- component systems accomplish signal transduction through the phosphorylation of a response regulator (RR) by a histidine kinase (HK). Histidine kinases are typically homodimeric transmembrane proteins containing a histidine phosphotransfer domain and an ATP binding domain. Response regulators can consist only of a receiver domain, but usually are multi-domain proteins with a receiver domain and at least one effector or output domain, often involved in DNA binding [43]. Upon detecting a particular change in the cellular environment, the HK performs an autophosphorylation reaction, transferring a phosphoryl group from adenosine triphosphate (ATP) to a specific histidine residue. The cognate response regulator (RR) then catalyzes the transfer of the phosphoryl group to an aspartate residue on the response regulator's receiver domain [44, 45]. This typically triggers a conformational change that activates the RR's effector domain, which in turn produces the cellular response to the signal, usually by activating or repressing expression of target genes [43].
[00127] An exemplary illustration of two-component system formed by the U-sensitive histidine kinase 1363 (UrpS) and the cognate response regulator 1362 (UrpR) is shown in the schematics of Figure 3A.
[00128] Genes encoding histidine kinase 1363 (UrpS) and response regulator 1362 (UrpR) herein described are herein also indicated as 1363 gene or 1363 and 1362 gene or 1362, respectively as will be understood by a skilled person.
[00129] A representative example of histidine kinase P1363 (UrpS) and response regulator 1362 (UrpR) in a two component system herein described are provided by the histidine kinase encoded by 1363 gene CCNA_01363 in Caulobacter crescentus (SEQ ID NOG), and the response regulator 1362 encoded by 1362 gene CCNA_01362 (SEQ ID NO:4), forming a two component systems respectively as will be understood by a skilled person.
[00130] In some embodiments the histidine kinase PI 363 (UrpS) and response regulator pi 362 (UrpR) can be heterologously expressed in the bacterial cell through genetic engineering of the cell performed to include in the cell a U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase 1363 (UrpS) and an endogenous or exogenous gene encoding U-sensitive transcriptional regulator 1362 (UrpR) in a configuration wherein the gene encoding histidine kinase 1363 and the gene encoding response regulator 1362 (UrpR) response regulator UzcR are expressed upon activation of a controllable promoter.
[00131] In some embodiments the histidine kinase P1363 (UrpS) and response regulator pl362 (UrpR) can be natively expressed in the bacterial cell. In particular in some embodiments, the host cell of the U-biosensors herein described is capable of natively expressing the proteins of the U-sensing two-component system 1363 (UrpS) and 1362 (UrpR) described herein. Those embodiments typically comprise certain proteobacterial cell such as alphaproteobacteria, betaproteobacteria or gammaproteobacteria comprising an endogenous 1363/1362 TCS (UrpRS) which can be identified and selected by methods to detect 1363 (UrpS) and/or 1362 (UrpR) genes in a candidate bacterial cell identifiable by a skilled person.
[00132] For example, presence of a 1363/1362 TCS or UrpRS in a proteobacterial cell can be identified by wet bench experiments, such as PCR, Southern blotting and additional techniques identifiable by a skilled person performed with histidine kinase PI 363 (UrpS) and response regulator pl362 (UrpR) and/or fragments thereof used as primers or probes for the related detection, followed by isolation and sequencing of the identified 1363 gene and/or 1362 gene as will be understood by a skilled person.
[00133] In addition or in the alternative, presence of a 1363/1362 TCS (UrpRS) in a proteobacterial cell can be identified by performing a sequence alignment using BLASTP or PST BLAST or other alignment algorithms known to persons skilled in the art with the 1363 (UrpS) amino acid sequence of C. crescentus NA1000 (SEQ ID NO:3) and/or the 1362 (UrpR) protein sequence of Caulobacter crescentus NA1000 (SEQ ID NO:4) as a query sequence against protein sequences of a given proteobacterial cell, as would be understood by a skilled person.
[00134] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing a 1363 (UrpS) protein having 100% homology to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3), and/or having a BLAST Score of 851 when aligned with SEQ ID NO: 3, (herein also 1363/1362 Tier 1 proteobacteria or UrpRS Tier 1 proteobacteria) such as proteobacteria C. crescentus NA1000 and C. crescentus CB 15.
[00135] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score over 800 when aligned to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3) and with less than 100% homology to C. crescentus NA1000 1363) SEQ ID NO: 3, (herein also 1363/1362 Tier 2 proteobacteria or UrpRS Tier 2 proteobacteria) such as exemplary proteobacterium C. crescentus CB2 among others identifiable by persons skilled in the art.
[00136] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score of 500 - 800 when aligned to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3) (herein also 1363/1362 Tier 3 or UrpRS Tier 3) such as proteobacteria Caulobacter henricii, Caulobacter sp. CCH5-E12, Caulobacter sp. OV484, Caulobacter sp. Root487D2Y, Caulobacter sp. Rootl455, Caulobacter sp. 12-67-6, Caulobacter sp. Root487D2Y, Caulobacter sp. Rootl455, Caulobacter sp. UNC358MFTsu5.1, Caulobacter sp. AP07 and Caulobacter , among others identifiable by persons skilled in the art.
[00137] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively comprising a 1362 (UrpR) binding site (SEQ ID NO: 1) in a phytase or 1361 promoter, also referred to herein as“Pphyt” and“P1361”, respectively (herein also indicated as 1363/1362 Tier 4 proteobacteria or UrpRS Tier 4 proteobacteria). Exemplary proteobacteria within these embodiments comprise Caulobacter sp. Root342, Phenylobacterium sp. Root700, Caulobacter crescentus NA1000, Caulobacter sp. Rootl455, Caulobacter sp. Root487D2Y, Paracoccus sp. 228, Caulobacteraceae bacterium OTSz_A_272, Novosphingobium sp. AP12 PMI02, Hyphomicrobium sp. MCI, Hyphomicrobium denitrificans, Brevundimonas sp. Root 1279
Sphingopyxis sp. Rootl497, Afipia sp. P52-10, Caulobacter sp. Root342 Hyphomicrobium denitrificans, Sphingobium sp. YBL2, Sphingobium baderi LL03, Sphingobium indicum B90A, and Roseovarius indicus strain DSM 26383, among others identifiable by persons skilled in the art.
[00138] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score of 300 - 500 when aligned to 1363 protein of C. crescentus NA1000 (SEQ ID NO: 3) (herein also indicated as 1363/1362 Tier 5 proteobacteria or UrpRS Tier 5 proteobacteria) such as exemplary proteobacteria Phenylobacterium sp. Root700, Phenylobacterium sp. Root700, Caulobacter sp. 39-67-4, Sphingopyxis sp. SCN 67-31, Phenylobacterium sp. SCN 70-31, Sphingopyxis flava, Caulobacteraceae bacterium OTSz_A_272, Sphingobium baderi, Caulobacterales bacterium 68-
7, alpha proteobacterium U9-li, Caulobacter sp. 35-67-4, Sphingopyxis granuli, Sphingopyxis macro goltabida, Brevundimonas sp. Rootl279, Sphingopyxis macrogoltabida, Brevundimonas sp. Rootl279, Sphingopyxis macrogoltabida, Hyphomonas polymorpha, Porphyrobacter mercurialis, Caulobacteraceae bacterium TH1-2, Hyphomonadaceae bacterium UKL13-1, Sphingopyxis macrogoltabida, Porphyrobacter mercurialis, and Novo sphingobium sp. PASSN1, among others identifiable by persons skilled in the art.
[00139] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score of 240-300 when aligned to 1363 protein of C. crescentus NA1000 (SEQ ID N03), (herein indicated also as 1363/1362 Tier
6 proteobacteria or UrpRS Tier 6 proteobacteria) identifiable by persons skilled in the art.
[00140] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 proteins having a BLAST Score of 200-240 when aligned to 1363 protein of C. crescentus NA1000 (SEQ ID NO:3) (herein also indicated as 1363/1362 Tier 7 proteobacteria or UrpRS Tier 7 proteobacteria), identifiable by persons skilled in the art.
[00141] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing 1363 (UrpS) proteins having a BLAST Score of less than 200 when aligned to 1363 (UrpS) protein of C. crescentus NA1000 (SEQ ID NO: 3) (herein also indicated as 1363/1362 Tier 8 proteobacteria or UrpRS Tier 8 proteobacteria), identifiable by persons skilled in the art (herein also indicated as Tier 8 proteobacteria).
[00142] In embodiments of the U biosensor herein described wherein the host cell is a proteobacterial cell of any one of 1363/1362 Tiers 1 to 6 (UrpRS Tiers 1 to 6), the host proteobacterium can comprise a bacterial cell with a natively and/or heterologously expressed 1363/1362 TCS (UrpRS)endogenous to the host proteobacterium. In embodiments wherein the host cell is a bacterial cell other than proteobacteria or a proteobacteria of 1363/1362 Tiers 7 and
8, the host is engineered to include a heterologous 1363/1362 TCS system in a configuration capable of heterologous expression in the host bacteria as will be understood by a skilled person. In some embodiments wherein the host cell is a proteobacteria of 1363/1362 Tier 6, the host can be firstly tested for the presence of a natively expressed 1363/1362 TCS endogenous to the host proteobacteria according to procedure identifiable by a skilled person. The test can be performed by transforming the cell with a plasmid or other vector containing Pphyt or PI 361 -regulated gfp fusion and assaying the system for U-dependent induction of GFP as will be understood by a skilled person. If the host cell does not possess a natively expressed 1363/1362 TCS, the host can be engineered to include a heterologous 1363/1362 TCS system in a configuration capable of heterologous expression in the host bacteria.
[00143] In embodiments wherein a heterologous 1363/1362 TCS (UrpRS) is introduced into the host cell, the heterologous 1363/1362 TCS can be a 1363/1362 TCS system from a 1363/1362 Tier 1, a 1363/1362 Tier 2, a 1363/1362 Tier 3, a 1363/1362 Tier 4 or a 1363/1362 Tier 5 proteobacteria, and it is preferably a 1363/1362 TCS from a 1363/1362 Tier 2 proteobacteria and more preferably from a 1363/1362 Tier 1 proteobacteria. In those embodiments wherein a heterologous 1363/1362 TCS is introduced into a host cell, the native 1363/1362 TCS of the host cell is preferably knocked out, in particular in embodiments wherein the host organism is a proteobacteria of any one of 1363/1362 Tiers 1 to 6. The native 1363/1362 TCS of the host cell can be knocked out by deleting or inactivating the 1362 gene cluster only or by deleting or otherwise inactivating both the 1362 and 1363 gene clusters.
[00144] In some embodiments, UC Fi-biosensors herein described the proteobacterial cell is capable of natively and/or heterologously expressing a U-sensitive histidine kinase UzcS, and a transcriptional response regulator UzcR.
[00145] The term“histidine kinase UzcS”,“UzcS” in the sense of the disclosure refers to a histidine kinase having the amino acid sequence:
MRLPRLLRTTPFRLTLLFL ALFA A A AS AFLG YIY V AT AGE VNRRAQ AEIS REFES LE A A Y RQGG VD ALN QTIVER AT S ERPFLYFLAD KDGKRIS GS IEES P V S GFT GDGPEW AS FKVTE TDLDGAEVKAAARG V QQRLDN GEILFV GAD VD ASE A YVRKIVRALW GAGALVILLGM AGG VLIS RN V S RS MQGL VD V VN A VRGGDLH ARAR VRGTRDE YDELAEGLNDMLDRIE RLMGGLRH AGD AIAHDLRS PLTRLRARME V ALID AEN GKGDP V A ALET ALQD ADG VL KTFNAVLAIARLQAAGSAPDQRQFDASELAGDMAELYELSCEDKGLDFKAEIVPALTIK GNREFL AQ AL ANILDN AIKYTPEGG AIMLR ARRT S S GELEF S VTDT GPG VPE ADR ARV V QRF VRLEN S RS EPG AGLGLS L V S A V AT S HGGRLEL AEGPGE YN GMGPGLR V ALVLPRV
E (SEQ ID NO: 5). or a sequence that when aligned with sequence SEQ ID NO: 5 has a BLAST score has a BLAST score greater than 300 and less than 500, or preferably greater than 500 and less than 767, or more preferably a BLAST score greater than 800 and a homology with SEQ ID NO: 5 less than 100%, or even more preferably BLAST score of 925 and an homology of 100% with SEQ ID NO: 5.
[00146] The term“transcriptional response regulator UzcR” or“UzcR’ as used herein indicates a transcriptional regulator having the amino acid sequence: MRILIIEDDLEAAGAMAHGLKEAGYDVAHAPDGEAGLAEAQKGGWDVLVVDRMMPK MDGVTVVETLRREGDQTPVLFLSALGEVNDRVVGLKAGADDYLVKPYAFPELMARVE ALS RRRETG A V ATTLKV GELEMNLINRT VHRQGKEIDLQPREF QLLEFMMRHAGQS VT RTMLLEKVWEYHFDPQTNVIDVHISRLRSKIDKGFDRAM LQTVRGAGYRLDP (SEQ ID NO: 6). or a sequence that when aligned with sequence SEQ ID NO: 6 has a BLAST score greater than 250 and less than 300, or preferably greater than 300 and less than 400, or more preferably a BLAST score greater than 400 and a homology with SEQ ID NO: 6 less than 100%, or even more preferably a BLAST score of 452 and an homology of 100% with SEQ ID NO: 6.
[00147] In U02F2-biosensors herein described, the U-sensitive histidine kinase UzcS, and transcriptional response regulator UzcR form a two-component system in the sense of the disclosure, also referred to herein as“UzcRS two-component system” or“UzcRS TCS” which is similar to the 1363/1362 TCS system herein described and exemplified by the schematics of
Figure 3.
[00148] In particular, the term“UzcRS two component system” or UzcRS TCS” as used herein refers to a regulatory system responsible for U, Zn, and Cu-dependent regulation of numerous genes in Caulobacter crescentus [3]. The UzcRS two component system comprises an OmpR/PhoB family response regulator (RR) and a histidine kinase (HK) containing a 123 amino acid periplasmic domain, placing it in the periplasmic- sensing class of histidine kinases [42].
[00149] Genes encoding histidine kinase UzcS and response regulator UczR herein described are herein also indicated as UzcS gene or UzcS and UczR gene or UczR, respectively as will be understood by a skilled person.
[00150] A representative example of histidine kinase UzcS and transcriptional response regulator UzcR in a UzcRS two components system herein described are provided by histidine kinase encoded by UzcS gene CCNA_02842in C. crescentus NA1000 (SEQ ID NO: 5) and by a transcriptional response regulator, for example encoded by UzcR gene CCNZ_02485in C. crescentus NA1000 (SEQ ID NO: 6) as will be understood by a skilled person.
[00151] In some embodiments the histidine kinase UzcS and transcriptional response regulator UzcR can be heterologously expressed in the bacterial cell through genetic engineering of the cell performed to include in the cell a U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the gene encoding histidine kinase UzcS and the gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
[00152] In some embodiments the histidine kinase UzcS and transcriptional response regulator UzcR are natively expressed in the bacterial cell. In particular in some embodiments, the host cell of the U-biosensors herein described is capable of natively expressing the proteins of the U- sensing UzcS/UzcR TCS described herein. Those embodiments typically comprise certain proteobacterial cell such as alphaproteobacteria, betaproteobacteria or gammaproteobacteria comprising an endogenous UzcS/UzcR TCS which can be identified and selected by methods to detect UzcS and/or UzcR genes in a candidate bacterial cell identifiable by a skilled person.
[00153] For example, presence of a UzcS/UzcR TCS in a proteobacterial cell can be identified by wet bench experiments, such as PCR, Southern blotting and additional techniques identifiable by a skilled person performed with histidine kinase UzcS and response regulator UzcR and/or fragments thereof used as primers or probes for the related detection, followed by isolation and sequencing of the identified UzcS gene and/or UzcR gene as will be understood by a skilled person. The presence of an UzcS/UzcR TCS can also be identified by introducing in the cell a UzcR-regulated GFP fusion promoter and detecting GFP fluorescence thus testing for U- dependent fluorescence as will be understood by a skilled person.
[00154] In addition or in the alternative, a UzcS/UzcR TCS in a proteobacterial cell can be identified by performing a sequence alignment using BLASTP or PSI-BLAST or other alignment algorithms known to persons skilled in the art with the UzcS amino acid sequence of C. crescentus NA1000 (SEQ ID NO: 5) and/or the UzcR protein sequence of Caulobacter crescentus NA1000 (SEQ ID NO: 6) as a query sequence against protein sequences of a given proteobacterial cell, as would be understood by a skilled person.
[00155] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having 100% homology to UzcS protein of C. crescentus NA1000 (SEQ ID NO: 5) and a BLAST score of 925 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcS/UzcR Tier 1 proteobacteria).
[00156] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a homology of less than 100% to UzcS protein of C. crescentus NA1000 (SEQ ID NO: 5) and a BLAST score greater than 800 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 2 proteobacteria).
[00157] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins a BLAST score greater than 767 and lower than 800 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 3 proteobacteria).
[00158] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score greater than 500 and less than 767 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 4 proteobacteria).
[00159] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score greater than 300 and less than 500 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 5 proteobacteria).
[00160] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score greater than 250 and less than 300 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 6 proteobacteria).
[00161] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score greater than 200 and less than 250 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 7 proteobacteria).
[00162] In some embodiments, the U sensing bacterial cell can comprise proteobacteria capable of natively expressing UzcS proteins having a BLAST score less than 200 when aligned to sequence SEQ ID NO: 5 (herein also indicated as UzcRS Tier 8 proteobacteria).
[00163] In embodiments of the U biosensor herein described wherein the host cell is a proteobacterial cell of any one of UzcRS Tiers 1 to 6, the host proteobacterium can comprise a bacterial cell with a natively and/or heterologously expressed UzcRS TCS endogenous to the host proteobacterium. In embodiments wherein the host cell is a bacterial cell other than proteobacteria or a proteobacteria of UzcRS Tiers 7 and 8, the host is engineered to include a heterologous UzcRS TCS system in a configuration capable of heterologous expression in the host bacteria as will be understood by a skilled person.
[00164] In embodiments wherein a heterologous UzcRS TCS is introduced into the host cell, the heterologous UzcRS TCS can be a UzcRS TCS system from a 1363/1362 Tier 1, a UzcRS Tier 2, a UzcRS Tier 3, a UzcRS Tier 4 or a UzcRS Tier 5 proteobacteria, and it is preferably a UzcRS TCS from a UzcRS Tier 2 proteobacteria and more preferably UzcRS TCS from a UzcRS Tier 1 proteobacteria. In those embodiments wherein a heterologous UzcRS TCS is introduced into a host cell, the native UzcRS TCS of the host cell is preferably knocked out, in particular in embodiments wherein the host organism is a proteobacteria of any one of UzcRS Tiers 1 to 6. The native UzcRS TCS of the host cell can be knocked out by deleting or otherwise inactivating the UzcR gene cluster only or by deleting or otherwise inactivating both the UzcR and UzcS gene clusters according to techniques identifiable by a skilled person (e.g. by microdeletion, clean deletion via double recombination, recombineering (e.g., Wanner method [46]) insertional inactivation, CRISPRi, CRIS PR-mediate recombination, transposon insertion, mutational inactivation, methylation and/or epigenetic inactivation as well as other techniques identifiable by a skilled person). [00165] In some embodiments of the UC Fi-biosensors herein described, a bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363, and transcriptional response regulator 1362, and/or capable of natively and/or heterologously expressing a U-sensitive histidine kinase UzcS, and a transcriptional response regulator UzcR, is genetically engineered to include a U-sensitive genetic molecular component configured to report and/or neutralize U.
[00166] The term “molecular component” as used herein indicates a chemical compound comprised in a cellular environment. Exemplary molecular components thus comprise polynucleotides, such as ribonucleic acids or deoxyribonucleic acids, polypeptides, polysaccharides, lipids, amino acids, peptides, sugars and/or other small or large molecules and/or polymers that can be found in a cellular environment.
[00167] The term“genetic molecular component” as used herein indicates a molecular unit formed by a gene, an RNA transcribed from the gene or a portion thereof and optionally a polypeptide or a protein translated from the transcribed RNA.
[00168] In embodiments herein described, a genetic molecular component comprises a promoter operatively connected to the gene of the genetic molecular component, so that the promoter is configured to initiate transcription of said gene. As would be understood by those skilled in the art, promoters are typically located adjacent to the transcription start sites of genes, on the same strand and upstream on a DNA sequence (towards the 5' region of the sense strand), and for transcription to occur, the enzyme that synthesizes RNA, known as RNA polymerase, attaches to the promoter. Promoters contain DNA sequences identifiable by those skilled in the art and described herein, such as those that provide binding sites for RNA polymerase and also for proteins that function as transcription regulatory factors that can either activate or repress gene transcription.
[00169] The term“transcription regulatory factor” or“transcription factor” as used herein refers to any type of factors that can function by acting on a regulatory DNA element such as a promoter or enhancer sequence. The transcription regulatory factors can be broadly classified into a transcription repression factor (also referred to as “repressor”) and a transcription activation factor (also referred to as“activator”). The transcription repression factor acts on a regulatory DNA element to repress the transcription of a gene, thereby reducing the expression level of the gene. The transcription activation factor acts on a regulatory DNA element to promote the transcription of a gene, thereby increasing the expression level of the gene. Both the transcription repression factors and the transcription activation factors can be used as one or more components in the gene circuits herein described. In particular, a transcription regulatory factor has typically at least one DNA-binding domain that can bind to a specific sequence of enhancer or promoter sequences. Some transcription factors bind to a DNA promoter sequence near the transcription start site and help form the transcription initiation complex. Other transcription factors bind to other regulatory sequences, such as enhancer sequences, and can either stimulate or repress transcription of the related gene. Examples of specific transcription repression factors include TetR, Lacl, LambdaCI, PhlF, SrpR, Qacl, BetR, LmrA, AmeR, LitR, met, and others identifiable by a skilled person, as well as homologues of known repression factors, that function in both prokaryotic and eukaryotic systems. Examples of transcription activation factors include AraC, LasR, LuxR, IpgC, MxiE, Gal4, GCN4, GR, SP1, CREB, and additional activation factors identifiably by a skilled person as well as homologues of known activation factors that function in both prokarayotic and eukaryotic systems also identifiable by a skilled person. Exemplary inducible regulators that can be used in Caulobacter comprise VanR (regulated by vanillate) and XylR (regulated by xylose), as well as others identifiable by those skilled in the art.
[00170] A gene comprised in a genetic molecular component is a polynucleotide that can be transcribed to provide an RNA and typically comprises coding regions as well as one or more regulatory sequence regions which is a segment of a nucleic acid molecule which is capable of increasing or decreasing transcription or translation of the gene within an organism either in vitro or in vivo. In particular coding regions of a gene herein described can comprise one or more protein coding regions which when transcribed and translated produce a polypeptide, or if RNA is the final product only a functional RNA sequence that is not meant to be translated. Regulatory regions of a gene herein described comprise promoters, transcription factor binding sites, operators, activator binding sites, repressor binding sites, enhancers, protein-protein binding domains, RNA binding domains, DNA binding domains, silencers, insulators and additional regulatory regions that can alter gene expression in response to stimuli as will be recognized by a person skilled in the art. [00171] An RNA of a genetic molecular component comprises any RNA that can be transcribed from a gene, such as a messenger ribonucleic acid (mRNA), short interfering ribonucleic acid, and ribonucleic acid capable of acting as regulating factors in the cell. mRNA comprised in a genetic molecular component comprise regions coding for the protein as well as regulatory regions e.g. ribosome binding site domains (“RBS”), which is a segment of the upstream (5’) part of an mRNA molecule to which the ribosomal machinery of a cell binds to position the message correctly for the initiation of translation. RBSs control the accuracy and efficiency with which the translation of mRNA begins. mRNA can have additional control elements encoded, such as riboregulator sequences or other sequences that form hairpins, thereby blocking the access of the ribosome to the Shine-Delgarno sequence and requiring an external source, such as an activating RNA, to obtain access to the Shine-Delgarno sequence. Other RNAs that serve regulatory roles that can comprise the genetic molecular component include riboswitches, aptamers (e.g. malachite green, Spinach), aptazymes, guide CRISPR RNAs, and other RNAs known to those skilled in the art.
[00172] A protein comprised in a molecular component can be proteins with activating, inhibiting, binding, converting, or reporting functions. Proteins that have activating or inhibiting functions typically act on operator sites encoded on DNA, but can also act on other molecular components. Proteins that have binding functions typically act on other proteins, but can also act on other molecular components. Proteins that have converting functions typically act on small molecules, and convert small molecules from one small molecule to another by conducting a chemical or enzymatic reaction. Proteins with converting functions can also act on other molecular components. Proteins with reporting functions have the ability to be easily detectable by commonly used detection methods (absorbance, fluorescence, for example), or otherwise cause a reaction on another molecular component that causes easy detection by a secondary assay (e.g. adjusts the level of a metabolite that can then be assayed for). The activating, inhibiting binding, converting, or reporting functions of a protein typically form the interactions between genetic components of a genetic circuit. Exemplary proteins that can be comprised in a genetic molecular component comprise monomeric proteins and multimeric proteins, proteins with tertiary or quaternary structure, proteins with linkers, proteins with non-natural amino acids, proteins with different binding domains, and other proteins known to those skilled in the art. Specific exemplary proteins include TetR, Lacl, LambdaCI, PhlF, SrpR, Qacl, BetR, LmrA, AmeR, LitR, met, AraC, LasR, LuxR, IpgC, MxiE, Gal4, GCN4, GR, SP1, CREB, and others known to a skilled person in the art.
[00173] A “U-sensing genetic molecular component” or “U-sensitive genetic molecular component” as used herein indicates a genetic molecular component wherein the gene of the genetic molecular component is under control of a U-sensing or U-sensitive promoter.
[00174] In particular, in some embodiments herein described, wherein the host cell is capable of natively and/or heterologously expressing the histidine kinase P1363 and response regulator pl362, at least one U-sensitive promoter comprises a U-sensitive 1362 binding site having a DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO: 1), wherein
Ni is C or T, preferably C;
N2 is G or A, preferably G;
N3 is T or C, preferably T;
N4 is C; Ns is A or G, preferably A;
Nό is G or C, preferably G;
N7 C or G;
Ns is any nucleotide;
N9 is any nucleotide;
Niois any nucleotide;
N11 is any nucleotide;
N12 is T or C;
Ni3 is G;
Ni4 is T or C, preferably T;
Nis is C;
Ni6is A or C, preferably A; Nn is G; and
Ni8 is C or G
and wherein Ni to Nn are selected independently.
[00175] In some embodiments of the U-sensing promoter comprising SEQ ID NO: 1, nucleotide Ni of the regulator direct repeat is in a position from about 16 nucleotides downstream of the transcription start site of the genetic molecular component as described herein to about 40 nucleotides upstream of the transcription start site of the genetic molecular component or genetic molecular component.
[00176] In some embodiments, the U- sensitive promoter further comprises nucleotides N19N20N21 (SEQ ID NO: 83), downstream of SEQ ID NO: 1 wherein each of N19 to N21 can independently be any nucleotide, and therefore N19 is any nucleotide N20 is any nucleotide; and N21 is G.
[00177] In some embodiments of the U biosensors described herein, the 1362 binding site has a DNA sequence CGTCAGCNNNNTGTCAGC (SEQ ID NO:7), C GTC AGGNNNNTGTC AGG (SEQ ID NO: 8), CGTCAGCNNNNTGTCAGG (SEQ ID NO:9),
CGTCAGCNNNNCGTCAGG (SEQ ID NO: 10), TGTCAGCNNNNTGTCAGC (SEQ ID NO: 11), CGCCTGCNNNNCGTCAGC (SEQ ID NO: 12), C GTC AGGNNNN C GTC AGC (SEQ ID NO: 13), CGTCAGCNNNNTGTCAGC (SEQ ID NO: 14), TGTCAGGNNNNTGTCAGC (SEQ ID NO: 15), CGTCAGCNNNNCGTCAGT (SEQ ID NO: 16),
CCGCGGGNNNNTGTCAGG (SEQ ID NO: 17), CGTCGGGNNNNAGACCGG (SEQ ID NO: 18), CGTCCGGNNNNCGTCAGA (SEQ ID NO: 19), CAACGCCNNNNCGTCAGC (SEQ ID NO: 20), CATCAGGNNNNCGTCAGC (SEQ ID NO: 21), CGCAGGGNNNNTGCAAGC (SEQ ID NO: 22), CATCAGCNNNNCGTCAGC (SEQ ID NO: 23),
CGTCATCNNNNTGTCACG (SEQ ID NO: 24), CGTCAGCNNNNCATCAGC (SEQ ID NO: 25), CTTCGCGNNNNCGTCCGG (SEQ ID NO: 26), CGTC AGGNNNN GGTC AGG (SEQ ID NO: 27), or TGTCAGCNNNNATCCTGC (SEQ ID NO: 28), wherein N can be any nucleotide.
[00178] In some embodiments, wherein the U-biosensor comprises a genetically engineered proteobacterial cell capable of natively and/or heterologously expressing histidine kinase UzcS, and U-sensitive transcriptional response regulator UzcR, at least one U-sensitive promoter comprises a UzcR binding site with an m_5 site configured for binding UzcR, having a DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide, and in some embodiments N7-N11 can independently be A. In an embodiment, each of N7-N12 can be A.
[00179] In some embodiments, in the proteobacterial cell, the endogenous genes encoding the histidine kinase UzcS, and the U-sensitive transcriptional response regulator UzcR, are knocked out and the genetically engineered proteobacterial cell is further engineered to include a U- sensing regulator genetic molecular component comprising an endogenous or exogenous gene encoding histidine kinase UzcS, and an endogenous or exogenous gene encoding U-sensitive transcriptional response regulator UzcR in a configuration wherein the a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
[00180] In some embodiments, the U-sensitive promoter comprises a UzcR binding site can be either a constitutively active promoter or an inducible promoter. In preferred embodiments, the promoter is constitutively active.
[00181] The term “m_5 site” or““UczR binding site” as used herein refers to a semi- palindromic consensus DNA binding site of sequence CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) [3], wherein N7-N12 is independently any nucleotide, and in some embodiments any one of N7-N11 can independently be A. One UzcR dimer likely to binds one m_5 site [3]. For example, a variant of an m_5 site wherein CATTAC (SEQ ID NO:29) is mutated to CAATAG (SEQ ID NO:30) is not bound by UzcR and a variant of an m_5 site wherein TTAA (SEQ ID NO:31) is mutated to TAAT (SEQ ID NO:32) is no longer activated by UzcR [3].
[00182] In embodiments herein described, UzcRS -regulated promoters comprise those having naturally-occurring m_5 sites or m_5 sites that are introduced into a promoter through genetic engineering. Accordingly, UzcRS -regulated promoters comprise DNA sequence elements required for RNA Polymerase binding, as well as one or more m_5 sites, such that the promoter is configured to be regulated by the UzcRS two-component system. Similar to promoters comprising 1362 binding sites, in UzcRS -regulated promoters, the s-RNAP biding sites typically have low sequence homology to the canonical G73-RNAP -10 and -35 hexamer sequences. Accordingly, typically transcriptional activation of native UzcRS -regulated promoters occurs through binding of UzcR to the promoter, consistent with little observed transcriptional activation in absence of UzcR.
[00183] In some embodiments, an UzcRS -regulated promoter can comprise 1 - 3 copies of the m_5 site. In particular, in some embodiments, when one or more m_5 sites are located at a position from about -50 to about -100 upstream of the TSS, preferably at a position -52/53 or - 62/63 upstream of the TSS, considering the first nucleotide of Seq ID NO: 2 as the first nucleotide of the m_5 site upstream from the TSS (see the position of the first nucleotide (C on the 5’ end) of the binding site in Figure 25 A) wherein the m_5 site is configured for activation of the UzcRS -regulated promoter (see e.g. configurations of Figure 25B); [3]). In some embodiments, one or more m_5 sites are located at a position 52 to53 bp upstream of the TSS or at a position 62 or-63 bp upstream of the TSS considering the first nucleotide of Seq ID NO: 2 as the first nucleotide of the m_5 site upstream from the TSS (see the position of the first nucleotide (C on the 5’ end) of the binding site in Figure 25 A) wherein the m_5 site is configured for activation of the UzcRS -regulated promoter (see e.g. Figure 25 B).
[00184] In some embodiments, one or more m_5 sites are located at a position within 100 nucleotides upstream of the TSS considering the first nucleotide of Seq ID NO: 2 as the first nucleotide of the m_5 site upstream from the TSS (see the position of the first nucleotide (C on the 5’ end) of the binding site in Figure 25 A). In some embodiments, when one or more m_5 sites are located at a position from about -49 upstream of the TSS to +25 downstream of the TSS, in particular, -33 upstream of the TSS to +15 nt downstream of the TSS, the m_5 site is configured for repression of the UzcRS -regulated promoter.
[00185] In some embodiments herein described, the U-sensitive promoter is configured such that upon binding of the response regulator UzcR to the m_5 site, the U-sensitive promoter is activated and transcription of a gene operatively connected to the U-sensitive promoter within the related genetic molecular component is initiated. [00186] In other embodiments, the U-sensitive promoter is configured such that upon binding of the U-sensitive transcriptional regulator to the U-sensitive transcriptional UzcR binding site, the U-sensitive promoter is repressed and transcription of a gene operatively connected to the U- sensitive promoter within the related U-sensing genetic molecular component is not initiated. In particular, in some embodiments one or more m_5 sites are located at a position -49 to +25 bp from the TSS considering the first nucleotide of Seq ID NO: 2 as the first nucleotide of the m_5 site upstream from the TSSThe approximate range is -49 to +25 bp from the TSS (see the position of the first nucleotide (C on the 5’ end) of the binding site (Figure 25A).
[00187] For example, in an exemplary embodiment wherein a promoter is repressed by UzcR, a m_5 site is engineered downstream of a transcription start site of a U-sensitive promoter such as a P 1361 promoter or Pphyt promoter (see Example 2). In these exemplary embodiments, insertion of a m_5 site downstream of the TSS of e.g. Pphyt minimizes activation of Pphyt in both the presence and absence of U. As would be understood by skilled persons, the latter is preferable as it minimizes UzcRS expression levels when no U is present, minimizing cross -reactivity with Zn and Cu.
[00188] Examples of promoters regulated by UzcRS comprise PUrcA, PurcB, Pi968, and others identifiable by those skilled in the art, such as those described in Park et ah, 2017 [3] herein incorporated by reference in its entirety (see also Example 11).
[00189] In several embodiments, one or more U-sensing promoters herein described are comprised within a U sensing genetic molecular component which is a genetically engineered polynucleotide construct configured to regulate expression of one or more RNA and/or protein encoding genes through the one or more U-sensing promoter.
[00190] In some embodiments, the U-sensitive promoter includes a 1362 binding site functioning as a binding site for natively expressed U-sensitive transcriptional regulator 1362 and/or a UzcR binding site for natively expressed UzcR. In some embodiments of the U- biosensors herein described, the histidine kinase 1363, and U-sensitive transcriptional regulator 1362 are therefore encoded respectively by 1363 and 1362 genes natively encoded in the genome of the proteobacterial cell and the encoded 1363 and 1362 proteins can be natively expressed in the proteobacterial cell. Similarly, in some embodiments of the U-biosensors herein described, the histidine kinase UzcR, and U-sensitive transcriptional regulator UzcS are therefore encoded respectively by uczR and uczS genes natively encoded in the genome of the proteobacterial cell and the encoded UczR and UczS proteins are natively expressed in the proteobacterial cell.
[00191] In other embodiments of UC Fi-biosensors herein described, 1363 and 1362 genes and/or the uczR and uczS genes are introduced into the proteobacterial cell of (e.g. a species of the subclass Caulobacteridae ) within one or more U-sensing regulator genetic molecular components configured to express the proteins 1363 and 1362 and/or UczR and UczS proteins upon activation of a controllable promoter.
[00192] In those embodiments, the genetic molecular components comprise the 1363 1362, uczR and/or uczS genes together with one or more regulatory regions configured to directly initiate expression of operatively connected 1363 and 1362 genes and/or operatively connected uczR and uczS genes. In some embodiments, a genetic molecular component introduced into the proteobacterial cell can comprise 1363 and 1362 genes and/or uczR and uczS genes in a same genetic molecular component, while in other embodiments the 1363 gene, the 1362 gene, the uczR gene and/or the uczS gene are comprised in different genetic molecular components. In some embodiments, the one or more regulatory regions operatively connected to the 1363 and/or 1362 genes and/or to the uczR and uczS genes, can comprise any promoter identifiable by skilled persons that is capable of initiating gene expression in a Caulobacteridae cell. Exemplary promoters that can be used to express 1363 and/or 1362 in Caulobacteridae comprise inducible promoter systems such as VanR (regulated by vanillate) and XylR (regulated by xylose), as well as others identifiable by those skilled in the art. In some embodiments, any constitutive promoter identifiable by those skilled in the art that has been characterized as functional to express an operatively linked gene of interest in a proteobacterial species of interest can be used to express 1363 and 1362 and/or uczR and uczS genes in the proteobacterial species of interest.
[00193] In some embodiments of the UO2F2 biosensors described herein wherein one or more genetic molecular components comprising a 1363 gene, 1362 gene, a uzcS and/or uzcR genes are introduced into a proteobacterial cell, the proteobacterial cell is further genetically engineered so that expression of its native a 1363 gene, 1362 gene, a uzcS and/or uzcR gene is inactivated by gene knockout. [00194] In some embodiments, in the proteobacterial cell, the endogenous genes encoding histidine kinas 1363, transcriptional regulator 1362, histidine kinase UzcS, and/or the U-sensitive transcriptional regulator UzcR are knocked out and the genetically engineered proteobacterial cell is further engineered to include a U-sensing regulator component comprising a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR in a configuration wherein the a gene encoding histidine kinase UzcS, and a gene encoding response regulator UzcR are expressed upon activation of a controllable promoter.
[00195] In some embodiments, the promoter controlling the expression of the heterologous 1363 gene, the 1362 gene, the UczR gene and/or the UczS gene can be either a constitutively active promoter or an inducible promoter. In preferred embodiments, the promoter is constitutively active.
[00196] In some embodiments of the UO2F2 biosensors described herein wherein one or more genetic molecular components comprising the 1363 and/or 1362 genes and/or the uczR and/or uczS genes are introduced into a proteobacterial cell, the proteobacterial cell is further genetically engineered so that expression of the host endogenous 1363 gene, 1362 gene, uczR gene and/or uczS gene is inactivated by gene knockout. Methods for performing genetic knockout are identifiable by persons skilled in the art, such as gene targeting using techniques such as homologous recombination, or transposon-mediated mutagenesis, or gene editing techniques such as those using CRISPR/Cas9 among others known to those skilled in the art.
[00197] In several embodiments, the U sensing genetic molecular component herein described is a genetically engineered polynucleotide construct configured to regulate expression of one or more RNA and/or protein-encoding genes through one or more U-sensing promoter.
[00198] In some embodiments herein described, the U-sensitive promoter is configured such that upon binding of the U-sensitive transcriptional regulator 1362 or UzcR to the U-sensitive corresponding binding site, the U-sensitive promoter is activated and transcription of a gene operatively connected to the U-sensitive promoter within the related genetic molecular component is initiated.
[00199] In particular, in some embodiments, when the U-sensing promoter comprises a 1362 binding site is located in a position wherein nucleotide Nis of the regulator direct repeat (SEQ ID NO: 1) is from about 17 nucleotides upstream of the transcription start site of the genetic molecular component to about 40 nucleotides upstream of the transcription start site of the genetic molecular component, the regulator direct repeat is configured to function as a transcriptional activator binding site.
[00200] In other embodiments, the U-sensitive promoter is configured such that upon binding of the U-sensitive transcriptional regulator to the U-sensitive transcriptional 1362 binding site, the U-sensitive promoter is repressed and transcription of a gene operatively connected to the U- sensitive promoter within the related U-sensing genetic molecular component is not initiated.
[00201] In particular, in some embodiments, when the U-sensing promoter comprises a 1362 binding site, the 1362 binding site can be located in a position wherein nucleotide Ni of the regulator direct repeat (SEQ ID NO:l) is from about 16 nucleotides downstream of the transcription start site of the genetic molecular component to about 16 nucleotides upstream of the transcription start site of the genetic molecular component or genetic molecular component, the 1362 binding site configured to function as a transcriptional repressor binding site. As would be understood by those skilled in the art, typically, within a given promoter polynucleotide sequence, substitution of a regulator repeat binding site polynucleotide sequence described herein for a promoter polynucleotide sequence comprising a -35 and/or a -10 hexamer sequence of the promoter, or one or more nucleotides at the transcriptional start site or downstream of the transcriptional start site, is expected to provide a promoter comprising a 1362 binding site configured to repress transcription of the promoter.
[00202] In preferred embodiments of the U-sensitive genetic molecular component described herein comprising the 1362 binding site, the regulator direct repeat is located in a position wherein nucleotide Nis of the regulator direct repeat (SEQ ID NO: l) is from about 17 nucleotides upstream of the transcription start site of the U-sensing genetic molecular component to about 40 nucleotides upstream of the transcription start site of the genetic molecular component or genetic molecular component, such that the regulator direct repeat is configured to function as a transcriptional activator binding site
[00203] As understood by those skilled in the art, positions in the promoter are designated relative to the transcriptional start site, where transcription of DNA begins for the gene of interest. Positions upstream (towards the 5’ end of the promoter) are negative numbers counting back from -1, for example -10 is a position 10 base pairs upstream of the transcription start site.
[00204] The term“transcription start site” or“TSS” as used herein refers to the location where transcription starts at the 5’ end of an encoded gene sequence. The location of the transcription start site is typically referred to as +1 relative to the 3’ end of a promoter operatively connected to the gene. As would be understood by persons skilled in the art, a putative transcription start site can be detected using techniques such as differential RNA-seq (dRNA-seq) [47], which can differentially detect primary transcripts having triphosphorylated 5' ends, and processed RNAs which do not. Additional techniques to detect transcription start sites known in the art comprise bioinformatic analysis to identify enrichment of promoter elements upstream of a putative transcriptional start site, and experimental validation of selected putative transcription start sites, for example using primer extension methods [48], or by using Northern blots to detect the associated RNAs, among other techniques identifiable by those skilled in the art [49] .
[00205] In some exemplary embodiments of the U biosensors described herein, the U-sensitive promoter comprising the 1362 binding site is a Pi36i promoter. The term“Pi36i promoter” as used herein refers to the promoter that natively regulates expression of an operon comprising CCNA_01361, CCNA_1362 and CCNA_1363 genes in Caulobacter crescentus. The DNA sequence of Pi36i is shown in Table 4.
[00206] In some exemplary embodiments of the U biosensors described herein, the U-sensitive promoter comprising the 1362 binding site is a Pphyt promoter. The term“Pphyt promoter” as used herein refers to the promoter that natively regulates expression of an operon comprising CCNA_01353, CCNA_01352_, and CCNA_01351 genes in Caulobacter crescentus. The DNA sequence of Ppi,yL is shown in Table 4. As understood by those skilled in the art, Caulobacter crescentus (Poindexter 1964) refers to a Gram- negative, oligotrophic bacterium widely distributed in fresh water lakes and streams. Caulobacter is an obligate aerobe with a ubiquitous presence in aqueous environments where it is well-adapted to life under low-nutrient conditions [50]. Caulobacter species tolerate high concentrations of U [51, 52], are found in U- contaminated sites [53], and can mineralize U through the formation of uranyl phosphate precipitates [54]. Multi-omics studies to elucidate the U stress response pathways in C. crescentus have revealed many highly-induced genes that are not induced by Cd, Cr, Pb or Se
[51].
[00207] In some embodiments, the U-sensitive promoter is a UzcRS -regulated promoters comprising DNA sequence elements required for RNA Polymerase binding, as well as one or more sequence elements for binding UzcR known as an m_5 site [3], as understood by those skilled in the art and herein also identified as UczR binding site.. Examples of promoters regulated by UzcRS comprise PurcA, PurcB, Pi968, and others identifiable by those skilled in the art, such as those described in Park et ah, 2017 [3].
[00208] In embodiments herein described, UzcRS -regulated promoters comprise those having naturally-occurring m_5 sites or m_5 sites that are introduced into a promoter through genetic engineering. Accordingly, UzcRS -regulated promoters comprise DNA sequence elements required for RNA Polymerase binding, as well as one or more m_5 sites, such that the promoter is configured to be regulated by the UzcRS two-component system. Similar to promoters comprising 1362 binding sites, in UzcRS -regulated promoters, the s-RNAP biding sites typically have low sequence homology to the canonical G73-RNAP -10 and -35 hexamer sequences. Accordingly, typically transcriptional activation of native UzcRS -regulated promoters occurs through binding of UzcR to the promoter, consistent with little observed transcriptional activation in absence of UzcR.
[00209] In some embodiments, an UzcRS -regulated promoter can comprise 1 - 3 copies of the m_5 site. In particular, in some embodiments, when one or more m_5 sites are located at a position from about -34 to about -100 upstream of the TSS, preferably at a position -43 to -53 upstream of the TSS, the m_5 site is configured for activation of the UzcRS-regulated promoter [3] . In other embodiments, when one or more m_5 sites are located at a position from about -33 upstream of the TSS to +15 nt downstream of the TSS, the m_5 site is configured for repression of the UzcRS-regulated promoter.
[00210] Accordingly, in some embodiments described herein, a promoter can be either activated or repressed by UzcR. For example, in an exemplary embodiment wherein a promoter is repressed by UzcR, a m_5 site is engineered downstream of a transcription start site of a U- sensitive promoter such as a P1361 promoter or Pphyt promoter (see Example 2). In these exemplary embodiments, insertion of a m_5 site downstream of the TSS of e.g. Pphyt minimizes activation of Pphyt in both the presence and absence of U. As would be understood by skilled persons, the latter is preferable as it minimizes UzcRS expression levels when no U is present, minimizing cross -reactivity with Zn and Cu.
[00211] In some embodiments, the U-sensitive promoter can be a promoter genetically engineered to comprise the U-sensitive 1362 binding site. In some exemplary embodiments, it is expected that a promoter can be engineered comprising a U-sensitive 1362 binding site having Ni8 of SEQ ID NO:l at a position of about -40 to -42 upstream of the transcription start site, as in exemplary promoters Pi36i and Pphyt (see e.g. Figure 7).
[00212] The U-sensitive promoters described herein can further comprise a holo-RNA Polymerase (RNAP) binding site. As understood by those skilled in the art, promoters in bacteria, such as members of Caulobacteridae require DNA sequence elements for s-RNAP binding for initiation of transcription. In particular, in Caulobacteridae promoters comprising the 1362 binding site such as exemplary promoters Pi36i and Ppi,yL, the s-RNAP biding site typically have low sequence homology to the canonical G73-RNAP -10 and -35 hexamer sequences. As such, activation of a promoter comprising a 1362 binding site at a position configured for transcriptional activation (e.g. wherein nucleotide Nis of the regulator direct repeat is located about -17 to about -40 upstream of the TSS), such as exemplary promoters Pi36i and Pphyt, s- RNAP binding likely requires binding of the U-responsive transcriptional factor to the 1362 binding site, consistent with the observed low level of transcriptional activation in absence of U (see Examples section).
[00213] The term“holoenzyme” as used herein refers to enzymes that contain multiple protein subunits, such as RNA polymerases, wherein the holoenzyme is a complete complex containing all the subunits needed for activity. The term“holoenzyme” can also refer to an enzyme together with one or more cofactors required for activity. For example, in bacteria, a promoter is recognized by RNA polymerase (RNAP) and an associated sigma factor, and the complex is referred to as an“RNAP holoenzyme” or“holo-RNAP”. An example of a RNAP holoenzyme in Caulobacteridae is RNAP holoenzyme containing s -73. [00214] The UO2F2 biosensors comprising the U-sensitive molecular component and/or U- sensitive genetic circuits comprising 1363/1362 TCS and/or UczRS TCS described herein in several embodiments can show improved selectivity for U compared to those previously described, such as in Hillson et ah, 2007 [55], as illustrated in the Examples. As shown in Figures 1 and 2, the previously unknown exemplary U-responsive Caulobacter crescentus promoters P1361 and Pphyt show highly selective gene expression upregulation in response to U, in contrast to Caulobacter crescentus promoter PUrcA [55], which also shows upregulation of gene expression in response to other metals such as Zn, Cu and Cd. Accordingly, a U-sensitive promoter, such as a P1361 or a Pphyt promoter, or a U-sensitive promoter genetically engineered to comprise the U-sensitive 1362 binding site can be used as a stand-alone selective U-sensing promoter.
[00215] In a U-sensitive genetic molecular component herein described, the U-Sensing promoter of the present disclosure directly or indirectly controls the expression of a reportable molecular component and/or a U-neutralizing molecular component.
[00216] The term“reportable molecular component” as used herein indicates a molecular component capable of detection in one or more systems and/or environments. The terms“detect” or“detection” as used herein indicates the determination of the existence, presence or fact of a target in a limited portion of space, including but not limited to a sample, a reaction mixture, a molecular complex and a substrate. The“detect” or“detection” as used herein can comprise determination of chemical and/or biological properties of the target, including but not limited to ability to interact, and in particular bind, other compounds, ability to activate another compound and additional properties identifiable by a skilled person upon reading of the present disclosure. The detection can be quantitative or qualitative. A detection is“quantitative” when it refers, relates to, or involves the measurement of quantity or amount of the target or signal (also referred as quantitation), which includes but is not limited to any analysis designed to determine the amounts or proportions of the target or signal. A detection is“qualitative” when it refers, relates to, or involves identification of a quality or kind of the target or signal in terms of relative abundance to another target or signal, which is not quantified.
[00217] In some embodiments, the reportable molecular component can be a molecular component linked to or comprising a label wherein the term label refers to a compound capable of emitting a labeling signal, including but not limited to radioactive isotopes, fluorophores, chemiluminescent dyes, chromophores, enzymes, enzymes substrates, enzyme cofactors, enzyme inhibitors, dyes, metal ions, nanoparticles, metal sols, ligands (such as biotin, avidin, streptavidin or haptens) and the like. The term“fluorophore” refers to a substance or a portion thereof which is capable of exhibiting fluorescence.
[00218] In embodiments of the U- sensitive genetic molecular component described herein, the genetic molecular component comprises a“reporter gene”, which can be any genetically- encoded reportable molecular component.
[00219] As would be understood by persons skilled in the art, the terms“genetically-encoded reportable molecular component”, “reportable genetic molecular component”, “genetically encoded reporter” or“reporter gene” comprises polynucleotide-encoded RNA and/or proteins having reportable characteristics identifiable by those skilled in the art and as described herein. Using genetic engineering techniques known to those skilled in the art, a reporter gene can be placed under the regulatory control of a promoter, and expression of the genetically-encoded reportable molecular component thereby serves as an indication of activation of the promoter in a host organism comprising the reporter gene. A reporter gene can be fused to another gene under the regulatory control of the promoter, such as a gene encoding a protein natively regulated by the promoter, so that the promoter regulates the expression of a fusion gene encoding a fusion protein comprised of the natively regulated protein covalently linked to the reportable molecular component. As would be understood by those skilled in the art, it is typical to use a reporter gene that is not natively expressed in the host organism, since the expression of the genetically- encoded reportable molecular component is used as a marker of activation of the promoter in the host organism. Exemplary genetically encoded reportable molecular components comprise fluorescent proteins such as green fluorescent protein (GFP) from Aequorea victoria or Renilla reniformis, red fluorescent protein from Discosoma species (dsRED), and variants thereof, beta galactosidase encoded by lacZ gene, luciferase and others identifiable to those skilled in the art. In exemplary embodiments described herein, exemplary reporters comprise a fluorescent protein which is a mutant variant of GFP referred to as‘gfpmut3’ [56]. [00220] In some embodiments, the UC Fi-biosensors comprising U-sensitive genetic molecular components and/or U-sensitive genetic circuits described herein are configured to produce a U- neutralizing molecular component in presence of U.
[00221] The term “U-neutralizing molecular component” as used herein refers to any component capable of decreasing the bioavailable U concentration.
[00222] In some embodiments, the U-neutralizing molecular component is a U-neutralizing genetic molecular component comprising a“U-neutralizing gene”. The term“U-neutralizing genetic molecular component” as used herein refers to a genetic molecular component in which the gene of the genetic molecular component is a U-neutralizing gene and wherein polynucleotide-encoded RNA and/or proteins have U-neutralizing characteristics identifiable by those skilled in the art upon reading of the present disclosure, such as proteins having enzymatic functions capable of allowing bioreduction, biomineralization, biosorption, or bioaccumulation of bioavailable U, as described herein.
[00223] The term“bioreduction” as used herein refers to altering the redox state of uranium from aqueous U (VI) to insoluble U (IV). As would be understood by persons skilled in the art, in the absence of oxygen, some bacteria are able to respire different electron acceptors to gain energy for metabolism. As anoxia progresses, the most energetically favorable electron acceptors are used in sequence, starting with the reduction of nitrate, then proceeding through Mn(IV), Fe(III) and sulfate, and finally the reduction of carbon dioxide to produce methane. At circumneutral pH, U(VI) has a similar redox couple to Fe(III), and natively Fe(III)-reducing bacteria are able to respire U(VI) as an alternative electron acceptor, reducing it to insoluble U(IV) [12]. Other groups natively capable of U(VI) reduction comprise bacteria such as sulfate- reducing bacteria [57], fermentative bacteria [58], acid-tolerant bacteria [59] and myxobacteria [60]. In some embodiments where a bioreduction component is comprised in the U-biosensor herein described, the host cell is preferably selected among cells natively expressing the components required to perform uranium bioreduction, which can be engineered to include one or more U-sensing genetic molecular component and/or other components of the U-sensing genetic circuit herein described. In other embodiments, a host cell can be E.Coli or other facultative anaerobe genetically engineered to include one or more U-sensing genetic molecular component and/or other components of the U-sensing genetic circuit herein described as well as genetic molecular components required to perform U-bioreduction as will be understood by a skilled person
[00224] Accordingly, in some embodiments uranium bioreduction can be used as a bioremediation technique, stimulated by adding an electron donor to promote enzymatic reduction of aqueous U(VI) to insoluble U(IV) [61-66]. Enzymatic reduction of U(VI) can be catalyzed using U-neutralizing genes such as those expressing cytochrome c [57, 67, 68]. In addition, chelators can be used to solubilize U(VI) and/or electron shuttles to mediate extracellular electron transfer, such as U-neutralizing genes expressing flavin mononucleotide or riboflavin [11, 69-72]. Therefore, in some embodiments of the U biosensors described herein, exemplary U-neutralizing genes comprise cytochrome c genes, flavin mononucleotide genes, or riboflavin genes which can be natively expressed in suitable host possibly further engineered to include one or more U-sensing genetic molecular components and/or U-sensing genetic circuit herein described.
[00225] The terms“biomineralization” and“bioprecipitation” as used herein refers to a process by which metals precipitate with microbially generated ligands such as sulfide or phosphate, or as carbonates or hydroxides in response to localized alkaline conditions at the cell surface. Thereafter, the uranium precipitate can be removed. In some embodiments. In some embodiments, the U can be comprised into a stable mineral that sediment out and not be re leached over time. In addition or in the alternative U-removal can be performed by some form of on-site filtration identifiable by a skilled person. In particular, sequestration of uranium as insoluble uranyl U(VI) phosphate biominerals can be used for in situ biomineralization for sites where bioreduction may not be feasible due to high nitrate concentrations or where there is risk of reoxidation reoccurring, e.g., in sites comprising carbonate [73, 74].
[00226] Accordingly, in some embodiments, the UC Fi-biosensors described herein can be engineered to catalyze precipitation of uranium such as uranyl phosphates. For example, bacteria can be engineered to precipitate uranyl phosphates [75] by expressing U-neutralizing genes such as acid-phosphatase genes [76] or alkaline-phosphatase genes [77] or phytase genes. In some embodiments, an exogenous source of phosphate can be added such as glycerol phosphate [16] or tributylphosphate [78]. Therefore, in some embodiments of the U biosensors described herein, exemplary U-neutralizing genes comprise acid-phosphatase genes, or alkaline-phosphatase genes. In an exemplary embodiment, the U-neutralizing gene is phoY, encoding an alkaline phosphatase that has been shown to allow the coupling of release of inorganic phosphorus (Pi) from organophosphates with U-Pi precipitation in Caulobacter crescentus [54], or phytase, that can be used to liberate phosphate from phytate, an environmental source of phosphate (see Example 10). In some embodiments, Pphyt or P1361 can be used to drive expression of an alkaline phosphatase. In some embodiments, the U-biosensors described herein can be engineered to comprise additional genetic molecular components configured to express one or more genes encoding proteins having enzymatic functions to catalyze release of inorganic phosphate from organophosphates (via hydrolytic cleavage catalyzed by phosphatases), inorganic phosphite (via enzymatic oxidation) and phosphonates (via cleavage catalyzed by C-P lyases), or from nucleic acids [79], phytate [80] or phospholipids [81].
[00227] The term“bioaccumulation” as used herein refers to accumulation of metals such as uranium in bacteria. For example, intracellular uranium accumulation occurs as uranyl phosphates in bacteria such as Pseudomonas species (Kazy et al., 2009; [82], VanEngelen et al., 2010; [83], Choudhary and Sar, 2011, [84]). In particular, overexpression of the polyphosphate kinase gene (ppk) encoding the PPK enzyme can be used to produce high levels of polyphosphate, a phosphate polymer with chain lengths of two to a few hundred, to allow precipitation of uranyl phosphate at the cell membrane [85]. Accordingly, in some embodiments of the U biosensors described herein, exemplary U-neutralizing genes comprise ppk genes.
[00228] The terms“biosorption” or“bioadsorption” as used herein refers to the passive uptake of uranium to the surface of microbial cells, wherein bacterial cell envelopes possess an electronegative charge, so are able to attract metal cations which sorb to the surface. In some embodiments of the U biosensors described herein, exemplary U-neutralizing genes comprise genes encoding proteins configured to bind U to the cell surface of the U biosensor. In an exemplary embodiment, the U-neutralizing gene is an ompA-SUP fusion gene encoding a rationally engineered super uranyl binding protein (SUP) having femtomolar affinity [86] (see Example 10). The encoded ompA fusion is configured to anchor SUP to the outer membrane, allowing adsorption of U to the cell surface. In other exemplary embodiments, the U-neutralizing gene is a fusion gene comprising the ompA protein fused to a Calmodulin EF-Hand Peptide (CaM) [87] that has been engineered for high U selectivity. In other exemplary embodiments, the U-neutralizing gene is a fusion gene comprising the Calmodulin EF-Hand Peptides (CaM) or SUP fused with the rsaA (S-layer) gene (see Example 10), for example using the method outline in Nomellini [88]. Accordingly, in some embodiments, proteobacteria can be engineered to provide U biosensors comprising U-neutralizing components configured to produce a U biosorption output, following a methodology such as has been used in Caulobacter for rare earth element adsorption [89].
[00229] Accordingly, UCEFi-biosensors comprising U-neutralizing molecular components described herein can be used in several embodiments for U bioremediation.
[00230] The term“bioremediation” as used herein refers to a waste management technique that involves the use of organisms to neutralize pollutants from a contaminated site. Bioremediation can be performed in situ or ex situ. In situ bioremediation involves treating the contaminated material at the site, while ex situ involves the removal of the contaminated material to be treated elsewhere.
[00231] As would be understood by persons skilled in the art, mobility of uranium in the environment depends on its speciation and redox state (e.g., see Figure 15). It is present as mobile U(VI) in oxidizing conditions, predominantly as the uranyl ion (U02 2+) or hydroxyl complexes below -pH 6.5, or as uranyl carbonate complexes at higher pH [90]. In the absence of carbonate, the uranyl ion and its complexes sorb strongly onto the surface of iron oxides and organics [73, 91, 92] and onto the edge sites of clay minerals [93, 94]. Sorption decreases in the presence of complexing ligands such as humic and fulvic acids, and in the presence of competing cations such as Ca2+ and Mg2+ [95]. Under reducing conditions, relatively insoluble and immobile U(IV) predominates, typically as the mineral uraninite, or as other U(IV) minerals [10, 96]
[00232] Biogeochemical interactions play a key role in controlling the speciation and mobility of uranium, through direct metabolic processes such as microbial respiration, or indirectly by changing ambient redox/pH conditions, producing ligands or new biominerals, or altering mineral surfaces. In addition to controlling uranium mobility via“natural attenuation”, these biogeochemical processes can be stimulated to accelerate clean-up of contaminated environments through bioremediation.
[00233] Preventing uncontrolled dispersion and transport of uranium in groundwater is a primary remediation goal at contaminated sites. Accordingly, stimulating bacterial interactions to fix aqueous uranium into insoluble minerals in situ can provide a relatively inexpensive and non- intmsive solution to remediating uranium contamination. Exemplary mechanisms of different microbe-uranium interactions are illustrated in Figures 16A-D, comprising bioreduction, biomineralization, biosorption, and bioaccumulation [7], among other identifiable by those skilled in the art.
[00234] In embodiments of the UChFi-biosensors described herein configured to have a U- neutralizing molecular component output to perform a U bioremediation function in response to bioavailable U, the preferred U-neutralizing molecular component output is a U-neutralizing molecular component having a U biomineralization function.
[00235] In some embodiments of a U- sensing genetic molecular component, the reporter gene or U-neutralization gene is contiguous with the U-sensitive promoter, wherein the 5’ end of the reporter gene is immediately adjacent to the 3’ end of the U-sensitive promoter. In other embodiments, the reporter gene is not contiguous with the promoter, such that one or more nucleotides are located between the 3’ end of the U-sensitive promoter and the 5’ end of the reporter gene. For example, in some embodiments, a ribosome binding site can be inserted between the 3’ end of the U-sensitive promoter (downstream of the TSS) and the 5’ end of the reporter gene.
[00236] In some embodiments, the U biosensors described herein comprise any non-pathogenic member of Caulobacteridae. In some embodiments described herein, the U biosensor comprises Caulobacteridae such as C. crescentus strains NA1000, CB15, and OR37, an environmental isolate from a U-contaminated site that exhibits high heavy metal tolerance [97]. In Examples provided herein, an exemplary host organism is C. crescentus strain NA1000.
[00237] In some embodiments of the U biosensor described herein, the cell can be any Caulobacteridae having a genome that natively comprises promoters having 1362 binding sites, such as exemplary promoters P1361 or Pphyt, or a homolog thereof.
[00238] In the UC Fi-biosensor of the instant disclosure, the U biosensor herein described is further engineered to include an F-sensing riboswitch.
[00239] The term“riboswitch” in the sense of the disclosure indicates a regulatory segment of a messenger RNA molecule that is configured to bind a target compound and to provide upon binding with the target compound a change in production of the proteins encoded by the mRNA. Accordingly, a riboswitch sense concentrations of the target compound. [98].
[00240] Riboswitches in the sense of the disclosure comprise an aptamer domain which is configured to specifically and selectively bind the target compound and an expression platform domain configured for genetic control of the mRNA expression. In particular, a riboswitch is configured so that in absence of the target compound the aptamer domain and the expression platform domain are configured to inhibit the expression of the mRNA. In a riboswitch in the sense of the disclosure upon binding of the target compound with the aptamer domain the riboswitch changes configuration and in the expression platform domain’s inhibition of the mRNA expression is removed, thus providing in metabolite-dependent allosteric control of gene expression.
[00241] In particular, aptamer domains are typically configured to form a stem structure in absence of the target compound. The hybridizing strands forming the stem are referred to as the aptamer strand and the control strand which are configured to complementarily bind to each other. In a riboswitch the stem structure which is either formed or be disrupted in absence of target compound is an aptamer domain-control strand stem structure.
[00242] In a riboswitch, expression platform domains generally have at least a portion identified as regulated strand configured to complementarily bind with the control strand of a linked aptamer domain to form a control strand-regulated strand structure, which can be a control strand-regulated strand stem. This control strand regulated strand structure will either form or be disrupted upon binding of the target compound.
[00243] Thus, the control strand of the aptamer domain can complementarily bind to with the aptamer strand and the regulated strand to form alternative stem structures with the aptamer strand and the regulated strand of the expression platform domain depending on the presence of the target domain.
[00244] The term‘complementary bind”,“base pair”,“complementary base pair” as used herein with respect to nucleic acids indicates the two nucleotides on opposite polynucleotide strands or sequences that are connected via hydrogen bonds. For example, in the canonical Watson-Crick DNA base pairing, adenine (A) forms a base pair with thymine (T) and guanine (G) forms a base pair with cytosine (C). In RNA base paring, adenine (A) forms a base pair with uracil (U) and guanine (G) forms a base pair with cytosine (C). Accordingly, the term“base pairing” as used herein indicates formation of hydrogen bonds between base pairs on opposite complementary polynucleotide strands or sequences following the Watson-Crick base pairing rule as will be applied by a skilled person to provide duplex polynucleotides. Accordingly, when two polynucleotide strands, sequences or segments are noted to be binding to each other through complementarily binding or complementarily bind to each other, this indicate that a sufficient number of bases pairs forms between the two strands, sequences or segments to form a thermodynamically stable double- stranded duplex, although the duplex can contain mismatches, bulges and/or wobble base pairs as will be understood by a skilled person.
[00245] In particular, in the riboswitch, complementary binding between the aptamer strand and the control strand is more or less thermodynamically stable than complementary base paring between the aptamer strand and other sequences of the riboswitch, and complementary binding between the control strand with the regulated strand of the riboswitch, depending on the presence of the target compound.
[00246] The term "thermodynamic stability" as used herein indicates a lowest energy state of a chemical system. Thermodynamic stability can be used in connection with description of two chemical entities (e.g. two molecules or portions thereof) to compare the relative energies of the chemical entities. For example, when a chemical entity is a polynucleotide, thermodynamic stability can be used in absolute terms to indicate a conformation that is at a lowest energy state, or in relative terms to describe conformations of the polynucleotide or portions thereof to identify the prevailing conformation as a result of the prevailing conformation being in a lower energy state. Thermodynamic stability can be detected using methods and techniques identifiable by a skilled person. For example, for polynucleotides thermodynamic stability can be determined based on measurement of melting temperature Tm, among other methods, wherein a higher Tm can be associated with a more thermodynamically stable chemical entity as will be understood by a skilled person. Contributors to thermodynamic stability can include, but are not limited to, chemical compositions, base compositions, neighboring chemical compositions, and geometry of the chemical entity.
[00247] Typically, in a riboswitch, the formation of the control strand-regulated strand structure affects expression of the RNA molecule containing the riboswitch.
[00248] The stem structure generally either is, or prevents formation of, an expression regulatory structure. An expression regulatory structure is a structure that allows, prevents, enhances or inhibits expression of an RNA molecule containing the structure. Examples include Shine-Dalgarno sequences, initiation codons, transcription terminators, and stability and processing signals, such as splice sites and sequences.
[00249] Riboswitches in the sense of the disclosure can be naturally occurring, isolated and recombinant riboswitches. Microbes, in particular, bacteria and archea, have evolved riboswitches to selectively detect over a dozen small molecules/metabolites (e.g., purine nucleobases) [99], several of which have been exploited for the construction of whole-cell biosensors [100].
[00250] Riboswitches in the sense of the disclosure can be comprised in sequences encoding proteins or peptides of interest, including reporter proteins or peptides, which can naturally occurring or synthetic as well as endogenous or heterologous to the host cell.
[00251] The term“F-sensing riboswitch” as used herein indicates a riboswitch for which Fluoride is the target compound. Accordingly, F-sensing riboswitches in the sense of the disclosure function as riboswitches that sense fluoride ions. These F-sensing riboswitches increase expression of downstream genes when fluoride levels are elevated, and the downstream genes can modulate the toxic effects of high levels of fluoride.
[00252] An F-sensing riboswitch comprises a fluoride aptamer domain which is configured to specifically and selectively bind fluoride and an expression platform domain configured for genetic control of the mRNA expression. An F-sensing riboswitch in accordance with the disclosure is configured so that in the absence of fluoride the fluoride aptamer domain and the expression platform domain are configured to inhibit the expression of the mRNA. Upon binding of fluoride with the aptamer domain the F-sensing riboswitch changes configuration and the inhibition of the mRNA expression imposed by the expression platform domain is thus removed.
[00253] In embodiments herein described the fluoride aptamer is expected to bind fluoride anions with a dissociation constant of between 50 mM and 60 pM, inclusive and to be able to detect an amount of fluoride in the environment, which can be calculated accordingly. In particular, the dissociation constant is particularly useful for determining the amount of fluoride detectable in the environment if the fluoride export capability of the host is abolished (when fluoride is retained inside the cell rather than pumped back out) as will be understood by a skilled person.
[00254] Fluoride detection capability of specific riboswitches is expected to be identified using approaches such as the ones exemplified in Example 20 and Example 22 as will be understood by a skilled person. In an example, following identification of a suitable host, the host genome can be searched to identify presence or absence of a native fluoride riboswitch such as a CrcB or EricF homolog. If a native fluoride riboswitch is present, the detection limit for a biosensor of the disclosure is expected to be close to 1 mM. In those instances, deletion of the fluoride transporter will likely enable a 100-fold decrease in the detection limit. Accordingly, host cell not expressing a native fluoride riboswitch or with a deleted native riboswitch is expected to be able to detect about lOpM or higher. It is expected that levels such as 50 pM and 60 pM fluoride will not substantially affect the viability of cells lacking fluoride efflux activity and that toxicity for a host will require mM levels of Fluoride depending on the host as will be understood by a skilled person.
[00255] As used herein,“fluoride aptamer domains” or“fluoride aptamers” indicate nucleic acid molecules that can specifically bind to fluoride ions. In an F-sensing riboswitch the fluoride aptamer domain is typically configured in a stem structure formed by the fluoride aptamer strand complementarily binding the control strand of the fluoride sensing riboswitch or in an alternative stem structure with the fluoride regulated strand of the expression platform domain or other sequences designed to complementarity bind the fluoride control strand depending on the presence of absence or fluoride according to the riboswitch design.
[00256] In particular, Fluoride aptamers typically are configured to form a stem structure in absence of fluoride. Fluoride aptamers include nucleic acid molecules that bind fluoride as the anion alone or fluoride with a counterion. Fluoride aptamers generally can be naturally occurring fluoride aptamers, such as fluoride aptamers in naturally-occurring F-sensing riboswitches, and fluoride aptamers derived from naturally-occurring fluoride aptamers.
[00257] A general sequence for F-sensing riboswitches is
N1N2N3N4N5N6N7N8N9G10G11N12R13A14U15G16R17N18U19R20U21Y22C23Y24N25C26C27N28R29N30
N31N32N33N34N35C36C37A38U39C40N41R42R43N44S45N46N47N48N49N50N51N52N53N54N55N56N57N58
N59N60W6lN62N63N64N65N66N67N68M69N70R7lR72N73N74M75U76G77M78Y79R80G8'C82M83R84Ys5M
86W87R88A89A90S91N92N93N94N95N96N97N98N99N100N101N102N103N104N105M106N107N108N109N110N1 I1N112N113Y114N115N1I6A117A1I8U119U120C121C122G123C124Y125N126N127N128N129N130N131N132G133 N 134A135G136N 137N138N139N140N141N142N143N144N145N146N147N148N149N150N151N152N153N154W 155N 156N 157N ISSN 159N I60N I61M162N163N I64A165N 166NI67NI68NI69NI70NI7INI72NI73NI74NI75NI76NI77NI7 8N179N180N181N182N183N184N185N186N187N188N189N190N191N192C193U194W195M196N197N198N199N200 N201N202N203N204N205N206N207N208N209N210N2I1N212N213R214G215C2I6U217R2I8A219U220G221R222Y2 23K224Y225C226U227R228Y229N230N231N232N233N234N235 (SEQ ID NO: 2509), wherein any one of N2, N3, N4, Ns, Nό, N7, Ns, N9, N31, N32, N33, N34, N35, N47, N48, N49, N50, N51, N52, N53, N54, N55, N56, N57, N58, N59, Nόo, N63, N64, N65, Nee, Nό7, N68, N93, N94, N95, N96, N97, N98, N99, N100, Nioi, N102, Nl03, NI04, NI05, N108, N109, N110, Nll l, N112, Nll3, NI27, NI28, N129, Nl30, Nl31, Nl32, Nl38,
Nl39, Nl40, NI41, NI42, Nl43, Nl44, Nl45, Nl46, Nl47, Nl48, Nl49, NlSO, NlSl, Nl52, Nl53, Nl54, , Nl67, Nl68, Nl69, Nl70, Nl71, Nl72, Nl73, Nl74, Nl75, Nl76, Nl77, Nl78, Nl79, Nl80, Nl81, Nl82, Nl83, Nl84, Nl85, Nl86, Ni87, N188, Ni89, Ni90, Ni9iNi92, N231, N232, N233, N234, and N235 indicated in the sequence in bold fonts, can be independently present or absent;
R indicates a purine (A or G). Y indicates a pyrimidine (C or U).
N indicates any nucleotide
W indicates a weak nucleotide (A or U).
S, indicates a strong nucleotide (G or C).
M, indicates an amino nucleotide (A or C).
K, Keto (G or U).
as also illustrated in Figure 43, This sequence, like any other riboswitch, is encoded by a corresponding DNA sequence wherein U is replaced by T as will be understood by a skilled person. This sequence further provides an indication of the conserved nucleotides in naturally occurring Fluoride sensing riboswitches as will be understood by a skilled person.
[00258] Exemplary F-sensing riboswitches in the sense of the disclosure comprise a fluoride- sensing riboswitch called a‘ crcB motif [102, 103]. crcB motif RNAs are typically located upstream of genes encoding proteins of diverse functions and presumably regulate these genes. Some of the gene products are annotated as ion transporters (for example, chloride, sodium, proton) and some others are involved in various physiological (e.g., universal stress adaptation, DNA repair) or metabolic (e.g., enolase, formate-hydrogen lyase) processes. The crcB riboswitch, in particular, enables a mechanism for sensing and detoxifying environmental fluoride by coupling fluoride binding [104] with the activation of genes encoding enzymes that mitigate fluoride toxicity, most commonly a fluoride exporter (CrcB) that expels internal fluoride anions [102, 105] (see Example 26). A key feature of the crcB motif is the high fluoride selectivity; binding has not been observed for other halides (including concentrations of chloride up to 2.5 M), small anions, gases and 36 other relevant cellular metabolites [102]. Furthermore, examination of crcB riboswitch function within a modified Escherichia coli strain revealed a fluoride detection limit below 10 mM (sub-500 ppb) and a dynamic range that spanned two orders of magnitude [102] (see Example 26) . Collectively, these data support the conclusion that crcB motif can be applied toward the environmental detection of fluoride, and ultimately, function as an integral detection component within a whole-cell UO2F2 compound sensor. .
[00259] Exemplary fluoride riboswitches, as well as related structures functions and locations in genomes are described in US patent US 9,580,713 incorporated herein by reference in its entirety. .[102] as well as in the papers Zao et al, 2017 [106], Ren et al 2012 [104] and Park and Taffet 2019 [107] each of which is herein also incorporated by reference in its entirety.
[00260] In particular, an exemplary consensus sequence and structure for crcB motif RNAs is shown in Figure 48A which reproduces FIG 1A of US patent US 9,580,713 [102] incorporated herein by reference in its entirety.
[00261] In particular, Figure 48A shows a consensus sequence and structural model based on the comparison of 2188 representatives from bacterial and archaeal species. PI, P2, P3 and pseudoknot labels of FIG. 1A of US 9580713 [102] identify base-paired substructures. As used herein, a pseudoknot is a nucleic acid secondary structure containing at least two stem-loop structures in which half of one stem is intercalated between the two halves of another stem as will be understood by a person skilled in the art.
[00262] Exemplary F-sensing riboswitches with a consensus sequence and structure schematically illustrated in Figure 48A has sequence
N1G2 G3 N4 R5 Aό U7 Gs R9 N10 G11 U12 Y13 Y14 N15 C½ C17 Nis A19 A20 C21 C22 G23 C24 Y25 R26 G27 C28 U29 G30 A31 U32 G33 A34 C35 N 6 Y37 C 8 U39 R40 Y41 (SEQ ID NO: 2510) wherein
Ni N4 N10 N15 Nis Nis and N36 are independently any amino acid;
Ni N4 Nio Nis Nis Nis and N36 are independently present or absent; anyone of N1G2 G3 N4 Rs is linked with any one of Yi4 N15 Ci6 Cn Nisy by a pseudoknot
A6is linked with Lkg by a pseudoknot
a first insertion segment of variable length is located between Nis and A19, the first insertion segment comprising a stem loop structure of variable length and having sequence starting with NR and ending with YN in a 5’ to 3’ direction
a second insertion segment of variable length is located between Y25 and R26; second insertion segment configured to form a stem loop structure R is A or G; and
Y is C or U as also indicated in Figure 48A
[00263] In F-sensing riboswitches of sequence SEQ ID NO: 2510, any one of nucleotides G3 , R5, U7, Gs , R9, Y 14 , Ci6, A20, A31, Y37 , C38 , U39 , and R40 is a 97% conserved nucleotide (present in 97% of the F-sensing riboswitches encompassed by SEQ ID NO: 2510)
[00264] In F-sensing riboswitches of sequence SEQ ID NO: 2510, any one of nucleotides G2, G3 , R5, Aό, Lb, Gs, R9, Y 14, Ci6 , Ci7, Ai9 , A20, G23, C24, G27, U29, A3i,Y37, C38, U39, R40 and Y41 is a 90% conserved nucleotide (present in 90% of the F-sensing riboswitches encompassed by SEQ ID NO: 2510)
[00265] In F-sensing riboswitches of sequence SEQ ID NO: 2510, any one of nucleotides Gn, U12 , Yi3 , C21, C22, Y25, R26, C28, G3O,G33, A34, and C35 is a 75% conserved nucleotide (present in 75% of the F-sensing riboswitches encompassed by SEQ ID NO: 2510W)
[00266] An additional, exemplary consensus sequence and structure for crcB riboswitches is reported in Figure 48. Figure 48B reproduces FIG. IB of US 9580713 [102] and shows an exemplary sequence and secondary structure model for the WT 78 Psy RNA comprising a crcB motif sequence. The 78-nucleotide RNA encompasses the crcB motif from Pseudomonas syringae.
[00267] An additional exemplary F-sensing riboswitch that is schematically illustrated in Figure 48B has sequence
G1G2A3U4C5G6G7C8G9C10A11U12U13G14G15A16G17A18U19G20G21C22A23U24U25C26C27U28C29C30A 31U32U33A34A35C36A37A38A39C40G41C42U43G44C45G46C47C48C49G50U51A52G53C54A55G56C57U58G5 9A60U61G62A63U64G65C66C67U67A68C69A70G71A72A73A74C75C76U77G78 (SEQ ID NO: 2511) wherein anyone of 2Ui3Gi4GisAi6Gi7is linked with any one of C27U28C29C30A by a pseudoknot and the F-sensing riboswitch Additional exemplary sequences are indicated in US 9580713. In particular as indicated in US 9580713 exemplary crcB motifs are described in (WO 2011/088076 and in [108]) also describing related structure and locations in organisms. Comparative genomics reveals 104 candidate structured RNAs from bacteria, archaeal, and their metagenomes as described in WO 2011/088076 and in [108] which are both hereby incorporated by reference in their entirety. In particular, incorporated by reference in its entirety is section 24 of the Additional File 3 of [108], which shows the sequences of numerous crcB motifs, as well as related genes, organisms, alignments, and consensus structures that can be used in connection with the biosensors herein described.
[00268] In some embodiments, the crcB RNA motif sequences herein described is a crcB RNA motif of the RF01734 family annotated by Rfam database which includes 2138 crcB RNA motif sequences of this family, their secondary and 3-D structures, and predicted phylogenetic tree for the sequence alignment (see http://rfam.org/family/RF01734 at the date of filing of the present disclosure which is also incorporated herein by reference in its entirety).
[00269] Additional consensus sequence and structures for F-sensing riboswitches are identifiable by a skilled person upon reading of the disclosure. In particular, similar to the consensus sequence and structure of FIG. 1A of US 9580713, a consensus secondary structure of the crcB RNA motif sequences of the RF01734 family is illustrated in Figure 48C. 8 out of 10 base pairs shown in Figure 48C are significant at E-value=0.05.
[00270] An F-sensing riboswitch with a consensus sequence and structure schematically illustrated in Figure 48C has sequence
N1N2G3G4N5G6A7U8G9R10N11G12U13Y14C15N16C16C17N18N19A20A21C22C23G24 C25Y26R27G28
C25U26G27A28U29G30A231C32N33Y34C35U36R37C38 (SEQ ID NO: 2512) wherein
Ni N2 Ns N11 Ni5 N18 Ni9 and N37 are independently any amino acid;
Ni N2 Ns N11 Ni5 N18 Ni9 and N37 are independently present or absent; anyone of N1N2G3G4N5G6A7 can be linked with any one of C15N16C16C17N18N19 by a pseudoknot of sequence NNCCNC a first insertion segment of 0 to 48 nucleotides is located between N19 and A20, wherein when the first insertion segment is > 12 nucleotides long, the first insertion comprises a stem loop structure of variable length having sequence starting with NRR and ending with YYN in a 5’ to 3’ direction
a second insertion segment of variable length structure is located between Y26 and R26, the second insertion segment configured to form a stem loop;
R is A or G; and
Y is C or U,
as also indicated in Figure 48C.
[00271] In F-sensing riboswitches of sequence SEQ ID NO: 2512, any one of Ni N2 N5 Nn N15 N18 and N37 is a 97% conserved nucleotide (present in 97% of the F-sensing riboswitches encompassed by sequence SEQ ID NO: 2512), and N19 is a 50% conserved nucleotide (present in 50% of the riboswitches encompassed by sequence SEQ ID NO: 2512) (Figure 48C)
[00272] In particular, exemplary F-sensing riboswitch sequence in the sense of the disclosure can have sequence
Ni N2 N3 N4 Ns Nό Ng Ns N9 G10 Gn N12 G13 A 14 U15 G½ G17 Nis G19 U20 Y21 C22 N23 C24 C25 N26 N27 N28 A29 A30 C31 C32 G33 C34 Y35 N36 N37 N38 N39 N40 N41 N42 R43 G44 C45 U46 G47 A48 U49 Gso Asi C52 N53 Y54 C55 Use R57 Css N59 N6O Net N62 N63 N64 (SEQ ID NO: 2513)
wherein
Ni, N2, N3, N4, Ns, Nό, N7, NS, N9, N 12, Nis, N23, N26, N27, N28, N36, N37, N38, N39, N40, N41, N42, N53, N59, Nόo, Nόΐ, Nό2, N 63 , Nό4 are interdependently any amino acid;
R is A or G; and
Y is C or U.
[00273] In particular, in F-sensing riboswitches of sequence SEQ ID NO: 2513, each of nucleotides Ni, N2, N3, N4, Ns, Nό, N7, Ns, N9, G11N12, A14 U15 Gi6, Nis, N23 , C24, N26 N28 , A29 A30 ,N37 N38 U46 A48 U49 N53 Y54 C55 U56 Css N59 Nόo Nόΐ Nό2 N63 N6 is a 97% conserved nucleotide (conserved in 97% of the F-sensing riboswitches) each of nucleotides G10 , Y21, C25, G33 N39 N40 R57 is a 90% conserved nucleotide (conserved in 90% of the F-sensing riboswitches) each of nucleotides Go Gn G19 C22 U20 C31 C32 C34 Y35 R43 G44 C45 G47 G50 A51 C52 is a 75% conserved nucleotide (conserved in 75% of the F-sensing riboswitches); and each of nucleotides N27 N36 N41 N42 is a 50% conserved nucleotide (conserved in 50% of the F-sensing riboswitches) as illustrated in Figure 48D
[00274] In particular, , F-sensing riboswitches of sequence SEQ ID NO: 2513 and Figure 48D the positions of the riboswitch Ni, N2, N3, N4, Ns, Nό, N7, Ns, N9, G11N12, A14 U15 G½, NIS, N23 , C24, N26 N28 , A29 A30 ,N37 N38 U46 A48 U49 N53 Y54 C55 U56 Css N59 Nόo Nόΐ Nό2 Nό3 N6 [are at least 97% conserved nucleotides, positions G10 , Y21, C25, G33 N39 N40 R57 are at least 90% conserved nucleotides, positions G13 Gn G19 C22 U20 C31 C32 C34 Y35 R43 G44 C45 G47 G50 A51 C52are at least 75% conserved nucleotides, and positions N27 N36 N41 N42 are at least 50% conserved nucleotides. It has been shown that the most highly conserved nucleotides of the RF01734 family undergo structural change on addition of fluoride, which suggests that these nucleotides help form a ligand-binding aptamer for fluoride [102]. Mutations that alter the aptamer’s conserved substructures or conserved nucleotides adversely affect fluoride binding, suggesting that the features common to all crcB motif RNAs are necessary for the selective recognition of fluoride.
[00275] In some embodiments, the crcB RNA motif of the RF01734 family can have a dissociation constant (KΏ) of ~60 mM with respect to binding of fluoride. [102]
[00276] Exemplary F-sensing riboswitches in the sense of the disclosure also comprise a fluoride sensing riboswitch that can be found in regulatory regions of a fluoride efflux pump called EricF’ a F_/H+ antiporter), performing an function ( F- efflux) of CrcB efflux pump that is commonly associated with fluoride riboswitches. CrcB and EricF are expected to carryout the same function - fluoride export. It seems that organisms have one or the other.
[00277] Additional exemplary Fluoride sensing riboswitches are sequences from SEQ ID NO: 205 to SEQ ID NO; 1988 and SEQ ID No: 1999 to SEQ ID NO: 2231 reported in Appendix I and SEQ ID NO: 1989-1998 and SEQ ID NO: 2232-2508 in Figure 53 of the disclosure which are incorporated herein by reference in its entirety, which contains 2017 exemplary F-sensing riboswitch sequences. An alignment of 287 sequences from the 2017 exemplary F-sensing riboswitch sequences is shown in Figure 53 (SEQ ID NO: 1989-1998 and SEQ ID NO: 2232- 2508) wherein the nucleotide sequences shown are sequences forming a same structure in corresponding aligned sequences and the symbols“—” indicates gaps between the aligned sequences shown (see Example 20) is also for further guidance concerning sequences and structures of F-sensing riboswitches suitable in constructs in accordance with the present disclosure.
[00278] In another, preferred embodiment the F-sensing riboswitch is the F-sensing riboswitch from Sphingomonas sp, 67-36 encoded by
CATGGTGACGGGGATGGAGTTCCCCGATAACCGCCGTTCCGGGCTGATGACTCCTAC CAACAC (SEQ ID NO: 2525).
[00279] In another, preferred embodiment the F-sensing riboswitch is the F-sensing riboswitch from Pseudomonas Syringae encoded by
CGGCGCATTGGAGATGGCATTCCTCCATTAACAAACCGCTGCGCCCGTAGCAGCTGA TGATGCCTACAGAAAC (SEQ ID NO: 2526).
[00280]
[00281] In a more preferred embodiment, the F-sensing riboswitch is MM-1 from Sphingomonas encoded by sequence
GTCGGCAACGGCAATGGATTCCTGCCGGGCCTCGCGCCGAACCGCCATTGAGGGCT GATGATTCCTACCTGCGG (SEQ ID NO: 2514) (See Example 27 and Example 28).
[00282]
[00283] In some embodiments of the U02F2-biosensors herein described an F-sensing riboswitch can be comprised within any one of the U-sensing genetic molecular components herein described.
[00284] In some embodiments of the U02F2-biosensors herein described an F-sensing riboswitch can be comprised in an F-sensing molecular component which is a genetically engineered polynucleotide construct configured to regulate expression of one or more RNA and/or protein-encoding genes through one or more promoter in combination with one or more F-sensing riboswitches. For example, an F-sensing riboswitch can be comprised in an F-sensing genetic reportable molecular component such as the one exemplified in Example 21.
[00285] In embodiments, herein described, any fluoride sensing riboswitch is expected to be functional in the host bacterium of interest. The well-characterized fluoride riboswitches from P. syringae DC3000 and Bacillus subtilis [102] represent exemplary riboswitches for fluoride sensor construction.
[00286] In some embodiments the native fluoride riboswitch from the host bacterium (if present) or a closely related bacterium can be used in the biosensors herein described, particularly if the promoter associated with the fluoride riboswitch is to be employed.
[00287] The relatedness among bacteria is defined based on taxonomic rank as will be understood by a person skilled in the art. In this connection relatedness among bacteria suitable to be used in biosensor of the present disclosure can be from the following groups
Group 1: From the species of interest (e.g., Caulobacter crescentus)
Group 2: From the genus of interest (e.g., Caulobacter)
Group 3: From the family of interest (e.g., Caulobacteraceae)
Group 4: From the order of interest (e.g. Caulobacterales)
Group 5 From the subclass of interest (e.g., Caulobacteridae)
Group 6: From the class of interest (e.g. Alphaproteobacteria)
[00288]“Closely related” bacteria encompass bacteria within a same subclass of interest (group 5) preferably within a same order of interest, more preferably within a same family of interest (Group 2) and more preferably within the same genus/species of interest (Group I).
[00289] For example, while C. crescentus , a preferred host strain for the U sensor, lacks a native fluoride riboswitch, the uncharacterized fluoride riboswitch from Sphingomonas sp. MM-1 are also expected to be usable. This bacterium is closely related to C. crescentus and possesses similarly high genomic G+C content, minimizing compatibility risks with C. crescentus. Additionally, despite the lack of biochemical characterization, the function of the Sphingomonas sp. crcB motif in fluoride sensing is supported by its homology with characterized crcB riboswitches [Pfam database [109]] and its genomic location upstream of a fluoride exporter.
[00290] In some embodiments, the genetically modified bacteria are bacteria incapable of natively expressing an F-sensing riboswitch. In some embodiments, the genetically modified bacteria are bacteria capable of natively expressing an F-sensing riboswitch. In some of those embodiments, the endogenous F-sensing riboswitch, are preferably knocked out or disabled through mutagenesis of conserved nucleotides critical for riboswitch function. [102] (see Example 26)
[00291] In particular in some embodiment where the native CrcB fluoride exporter functions are maintained the detection limit of the colorimetric crcB reporter (~1 mM) can be adversely affected. In those embodiments, abolishing fluoride export by deletion of the crcB gene yielded a 100-fold improvement in the fluoride detection limit (sub 10 mM) [102]. Achieving a low fluoride detection limit will likely require deletion of the crcB gene or the analogous eriCF gene in the host organism. [102]
[00292] The utility of a fluoride riboswitch for applications in a whole-cell fluoride sensor is supported by the finding that the crcB motif can be appended to the lacZ reporter gene in both E. coli and B. subtilis, enabling a colorimetric output response over a 100-fold range of fluoride concentrations. [102]. Additional elements and configuration of F-sensing components herein described (including U-sensing F-sensing reportable genetic molecular component, an F-sensing reportable genetic molecular component) can be identified by a skilled person based on the component and the host bacterial cell.
[00293] Any fluoride sensing riboswitch and in particular any fluoride sensing riboswitch reported in Appendix I or encompassed by any one of SEQ ID NO: 2509 to SEQ ID NO: 2520, is expected to be employed for fluoride sensor construction, assuming that the promoter and ribosome binding sequence (RBS) are optimized for the organism of interest [00294] Any promoter and Ribosome Binding Sequence from a bacteria from the above Groups 1 to Group 6 which is encompassed by a known consensus and has a location upstream of crcB/ericF can be used in connection with each riboswitch sequence from all groups, with preference for promoters from bacteria closely related to the host bacteria of the UO2F2- biosensors herein described.
[00295] An F-sensing riboswitch can be included in any F-sensing molecular component herein described in combination with promoter and RBS in a configuration identifiable by a skilled person. Reference is made in this connection to the schematic of Figure 52 which shows a schematic representation of the primary genetic parts involved in the construction of a fluoride sensing reporter include a (2) promoter to initiate transcription, a (1) fluoride riboswitch to terminate transcription in the absence of fluoride, a (3) ribosome binding site (RBS) to initiate translation, and a (5) reporter for detection. A portion of the (4) native crcB or ericF gene may be required for riboswitch function (i.e., coupling fluoride binding with transcriptional termination) of some fluoride riboswitches. The crcB gene is not required for the function of the Sphingomonas MM-1 fluoride riboswitch (Example 28).
[00296] In embodiments, herein described a skilled person will be able to identify a combination and configuration of an F sensing riboswitch based on the host bacteria by selecting the promoter and ribosome binding sequence (RBS) that are functional, and preferably optimized, for the host, such as a promoter and RBS native to the selected host bacterial cell or the closely related bacteria.
[00297] An example of closely related bacteria is provided by C. crescentus, and Sphingomonas sp. Accordingly, in an exemplary embodiments, a biosensor can be engineered in the preferred host strain for the U sensor C. crescentus , which lacks a native fluoride riboswitch with the uncharacterized fluoride riboswitch from Sphingomonas sp. MM-1. Sphingomonas sp is closely related to C. crescentus and possesses similarly high genomic G+C content, minimizing compatibility risks with C. crescentus.
[00298] Following identification of a promoter and RBS functional for the host bacteria, a suitable F-sensing riboswitch and the related configuration to obtain a desired F-sensitivity in a selected host bacteria can be identified with a method wherein the F-sensing riboswitch, promoter and RBS are tested in the host bacteria to identify a functional F-sensing construct to be used in F-sensing genetic molecular components herein described.
[00299] The method comprises selecting an F-sensing riboswitch from the host bacteria or from a closely related bacteria; selecting a promoter native to the selected host bacterial cell or the closely related bacteria , selecting a ribosome binding sequence (RBS) native to the selected host bacteria or the closely related bacteria, and selecting a portion of the crcBlericF coding region. The method further comprises providing a candidate F-sensing construct wherein the selected F- sensing riboswitch, the selected native promoter, the selected RBS, and the selected portion of the crcBlericF coding region are included in an gene expression cassette together with a reporter in a candidate configuration allowing expression of the reporter in the host bacteria in presence of Fluoride (see configuration schematically illustrated in Figure 52) The method further comprises introducing the candidate F-sensing construct in the host bacteria for a time and under condition to allow expression of the reporter; and detecting expression of the reporter in presence of a selected F amount to identify an F-sensitivity of the candidate F-sensing construct. The method can further comprise performing the method with additional candidate F-sensing constructs and selecting the candidate F-sensing construct having the desired F-sensitivity.
[00300] In embodiments herein described, testing of different F-sensing riboswitches in F- sensing construct according to methods herein described can be performed to identify the combination of F-sensing riboswitches promoter and RSB sequence, number of riboswitches, length of the construct, presence of spacers and additional structural features of the configuration of the construct, resulting in an F-sensing cassette with a desired F sensitivity to the U-biosensor over additional combinations with less desired or no F-sensitivity (see e.g. testing of Examples 27-33).
[00301] For example, in some embodiments, an F-sensing construct in C. crescentus preferably comprises the Sphingomonas MM-1 fluoride riboswitch module which outperformed an analogous module from the more distantly related strain Pseudomonas syringae in terms of both signal amplitude and dynamic range (see Example 27 and Example 28).
[00302] In another exemplary application, an F-sensing in C. crescentus comprises the fluoride riboswitch from the closely related Sphingomonas 67-36 in combination with a xylose-inducible promoter (see Example 27 and Example 28).
[00303] In some embodiments, an F-sensing construct can comprise two or more F-sensing riboswitches (see Example 29).
[00304] In some embodiment, the F-sensing riboswitch can be comprised in an F-sensing construct in combination with the Pxyl promoter. In some of those embodiments the F-sensing riboswitch can be anyone of the F-riboswitches from Sphingomonas SP., in particular MM-1 (see e.g. SEQ ID NO: 2514) and Sphingomonas 67-36 (see e.g. SEQ ID NO: 2525). In some of these embodiments the F-sensing riboswitch can also be anyone of the F-riboswitches from Pseudomonas Syringae. In some of embodiments the F-sensing construct can have sequences SEQ ID NO: 2528, SEQ ID NO: 2519 or SEQ ID NO: 2520, preferably SEQ ID N02518 (see e.g. Example 28).
[00305] In some embodiments, the F-sensing riboswitch can be comprised in an F-sensing construct in combination with a promoter native to the host bacteria, In those embodiments, the F-sensing riboswitch can preferably be an F-sensing riboswitch MM-lfrom Sphingomonas Sp. (see e.g. SEQ ID NO: 2514).
[00306] In some embodiments, the F-sensing riboswitch can be comprised in an F-sensing construct in combination with Pphyt promoter. In those embodiments, the F-sensing riboswitch can preferably be an F-sensing riboswitch MM-lfrom Sphingomonas Sp. (see e.g. SEQ ID NO: 2514).
[00307] In embodiments, herein described, preferably an F-sensing riboswitch is comprised together with other elements such promoters, RBS, Shine Dalgamo sequences and others identifiable by a skilled person in F-sensing constructs in a configuration that has a high signal amplitude (e.g. 25,000 - 200,000 for a normalized fluorescence signal)and large dynamic range with respect to fluorescence (higher than 100, preferably higher then 1000 and more preferably higher than 10,000)).
[00308] The wording“signal amplitude” as used herein with respect to a construct, cassette or component herein described, indicates the difference between the maximum (concentration of analyte such as U or F- that yields the highest signal) and minimum (no analyte, in particular no U and/or F) signal produced by the construct, cassette or component herein described For signal amplitude detected by fluorescence a high signal amplitude is typically a normalized fluorescence signal of 25,000 - 200,000, wherein the wording normalized fluorescence indicates the fluorescence signal (in arbitrary units) /cell density (OD600).
[00309] The wording“dynamic range” as used herein with respect to a construct, cassette or component herein described, indicates is defined as the ratio between the maximum (concentration of analyte such as U and/or F- that yields highest signal) and minimum (no analyte, in particular no U and/or F) signal produced by the construct, cassette or component herein described. A high value for the dynamic range are 100, 1000, or 10,000, wherein a higher dynamic range indicates a better analyte detection as will be understood by a skilled person.
[00310] In some embodiments, at least one Fluoride sensing riboswitch herein described can be comprised in combination of at least one a U-sensing genetic molecular component herein described, in a U-sensitive F-sensitive genetic circuit together with a reporter molecular component and/or a U-neutralizing molecular component. In addition or in the alternative, a Fluoride sensing riboswitch can be comprised in an F-sensitive genetic circuit together with a reporter molecular component, the F-sensing genetic circuit to be comprise in a biosensor herein described in combination with a U-sensitive genetic circuit and/or component as will be understood by a skilled person upon reading of the present disclosure.
[00311] In particular, in a U-sensitive genetic circuit, the U-sensing genetic molecular component, the reporter molecular component, and/or the U-neutralizing molecular components as well as possibly other components are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components.
[00312] Similarly for an F-sensing genetic circuit an F-sensing genetic molecular component comprising a fluoride sensing riboswitch and a reporter molecular component (as well as possibly other components) are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components when the genetic circuit operates according to the circuit design in presence of bioavailable Fluoride. [00313] In a U-sensitive genetic circuit herein described, at least one molecular component is a U-sensing genetic molecular component in which a U-sensitive promoter having a regulator direct repeat sequence of SEQ ID NO: 1 or any of SEQ ID NO:7-28) is activated or repressed in presence of bioavailable U. In a U-sensitive genetic circuit herein described at least one molecular component is a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design in presence of bioavailable U. In embodiments wherein at least one of the genetic molecular components of the U-sensitive genetic circuit herein described comprises a Fluoride sensing riboswitch the genetic circuit operates according to the circuit design in presence of bioavailable U and in presence of bioavailable Fluoride.
[00314] The term “genetic circuit” as used herein indicates a collection of molecular components connected one to another by biochemical reactions according to a circuit design. In particular, in a genetic circuit the molecular components are connected one to another by the biochemical reactions so that the collection of molecular components is capable to provide a specific output in response to one or more inputs.
[00315] In genetic circuits in the sense of the present disclosure, the molecular components forming parts of the genetic circuit can be genetic molecular components or cellular molecular components.
[00316] The term“cellular molecular component” indicates a molecular component not encoded by a gene, or indicates a molecular component transcribed and/or translated by a gene but comprised in the circuit without the corresponding gene. Exemplary cellular components comprise polynucleotides, polypeptides, polysaccharides, small molecules and additional chemical compounds that are present in a cellular environment and are identifiable by a skilled person. Polysaccharides, small molecules, and additional chemical compounds can include, for example, NAD, FAD, ATP, GTP, CTP, TTP, AMP, GMP, ADP, GDP, Vitamin Bl, B 12, citric acid, glucose, pyruvate, 3-phosphoglyceric acid, phosphoenolpymvate, amino acids, PEG-8000, FiColl 400, spermidine, DTT, b-mercaptoethanol maltose, maltodextrin, fructose, HEPES, Tris- Cl, acetic acid, aTc, IPTG, 30C12HSL, 30C6HSL, vanillin, malachite green, Spinach, succinate, tryptophan, and others known to those skilled in the art. Polynucleotides can include RNA regulatory factors (small activating RNA, small interfering RNA), or“junk” decoy DNA that either saturates DNA-binding enzymes (such as exonuclease) or contains operator sites to sequester activator or repressor enzymes present in the system (for example, as in [110]). Polypeptides can include those present in the genetic circuit but not produced by genetic components in the circuit, or those added to affect the molecular components of the circuit.
[00317] In some embodiments of genetic circuits herein described, one or more molecular components is a recombinant molecular component that can be provided by genetic recombination (such as molecular cloning) and/or chemical synthesis to bring together molecules or related portions from multiple sources, thus creating molecular components that would not otherwise be found in a single source.
[00318] In embodiments herein described, a genetic circuit comprises at least one genetic molecular component or at least two genetic molecular components, and possibly one or more cellular molecular components, connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components.
[00319] The term“activating” as used herein in connection with a molecular component of a genetic circuit refers to a reaction involving the molecular component which results in an increased presence of the molecular component in the cellular environment. For example, activation of a genetic molecular component indicates one or more reactions involving the gene, RNA and/or protein of the genetic molecular component resulting in an increased presence of the gene, RNA and/or protein of the genetic molecular component (e.g. by increased expression of the gene of the molecular component, and/or an increased translation of the RNA). An example of“activating” described herein comprises the initiation of expression of a gene regulated by a UzcRS -regulated promoter by a UzcR protein (e.g., see Example 2).
[00320] Activation of a molecular component of a genetic circuit by another molecular component of the circuit can be performed by direct or indirect reaction of the molecular components. Examples of a direct activation of a genetic molecular component comprised in a circuit the production of an alternate sigma factor (molecular component of the circuit) that drives the expression of a gene controlled by the alternate sigma factor promoter (other molecular component of the circuit), or the production of a small ribonucleic acid (molecular component of the circuit) that increases expression of a riboregulator-controlled RNA (molecular component of the circuit). Specific examples of this include the activity of sigma28 or sigma54 as demonstrated in [111]. Examples of indirect activation of a genetic molecular component comprise the production of a first protein that inhibits an intermediate transcriptional repressor protein, wherein the intermediate transcriptional repressor protein represses the production of a target gene, such that the first protein indirectly activates expression of the target gene.
[00321] The term“inhibiting” as used herein in connection with a molecular component of a genetic circuit refers to a reaction involving the molecular component of the genetic circuit and resulting in a decreased presence of the molecular component in the cellular environment. For example, inhibition of a genetic molecular component indicates one or more reactions involving the gene, RNA and/or protein of the genetic molecular component resulting in a decreased presence of the gene, RNA and/or protein (e.g. by decreased expression of the gene of the molecular component, and/or a decreased translation of the RNA). Inhibition of a cellular molecular component indicates one or more reactions resulting in a decreased production or increased conversion, sequestration or degradation of the cellular molecular components (e.g. a polysaccharide or a metabolite) in the cellular environment.
[00322] Inhibition can be performed in the genetic circuit by direct reaction of a molecular component of the genetic circuit with another molecular component of the circuit or indirectly by reaction of products of a reaction of the molecular components of the genetic circuit with another molecular component of the circuit.
[00323] The term“binding” as used herein in connection with molecular components of a genetic circuit refers to the connecting or uniting two or more molecular components of the circuit by a bond, link, force or tie in order to keep two or more molecular components together, which encompasses either direct or indirect binding where, for example, a first molecular component is directly bound to a second molecular component, or one or more intermediate molecules are disposed between the first molecular component and the second molecular component another molecular component of the circuit. Exemplary bonds comprise covalent bond, ionic bond, van der Waals interactions and other bonds identifiable by a skilled person. [00324] In some embodiments, the binding can be direct, such as the production of a polypeptide scaffold that directly binds to a scaffold-binding element of a protein. In other embodiments, the binding may be indirect, such as the co-localization of multiple protein elements on one scaffold. In some instances, binding of a molecular component with another molecular component can result in sequestering the molecular component, thus providing a type of inhibition of said molecular component. In some instances, binding of a molecular component with another molecular component can change the activity or function of the molecular component, as in the case of allosteric interactions between proteins, thus providing a type of activation or inhibition of the bound component. An example of“binding” as described herein comprises the binding of UzcR to an m_5 site in a UzcRS -regulated promoter (e.g., see Example 2).
[00325] The term“converting” as used herein in connection with a molecular component of the circuit refers to the direct or indirect conversion of the molecular component into another molecular component. An example of this is the conversion of chemical X by protein A to chemical Y that is then further converted by protein B to chemical Z. An example of “converting” as described herein comprises the cleavage of o-nitrophenyl-P-D-galactoside (ONPG) by beta-galactosidase encoded by the lacZ gene (e.g. see Example 1).
[00326] In embodiments of the U-sensitive and/or F-sensitive genetic circuits described herein, the molecular components are connected one with another according to a circuit design in which a molecular component is an input and another molecular component is an output. In particular, a genetic circuit typically has one or more input or start molecular component which activates, inhibits, binds and/or convert another molecular component, one or more output or end molecular component which are activated, inhibited, bound and/or converted by another molecular· component, and intermediary molecular components each inhibiting, binding and/or converting another molecular component and being activated, inhibited, bound and/or converted by another molecular component. In embodiments of the U-sensitive and/or F-sensitive genetic circuits herein described, the input is bioavailable U and/or bioavailable F and the output is a reportable molecular component and/or a U -neutralizing molecular component.
[00327] In some embodiments of the UO2F2 biosensor herein described the U-sensitive and/or F-sensitive genetic circuits can be comprised together within a same biosensor to detect or to detect and neutralize bioavailable UO2F2.
[00328] In some embodiments of the UO2F2 biosensor herein described the U-sensitive and/or F-sensitive genetic circuits can be comprised in combination with a U-sensing genetic molecular component and/or a F-sensing genetic molecular component of the disclosure to the detect or to detect and neutralize bioavailable UO2F2.
[00329] In some embodiments of the UO2F2 biosensor described herein, the U-sensitive F- sensitive genetic circuit can comprise at least one F-sensing genetic molecular component in which RNA is expressed in presence of bioavailable F, and at least one U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U-sensitive transcriptional 1362 binding site or a U-sensitive transcriptional UzcR binding site. In addition or in the alternative, in the U- sensing F-sensing genetic circuit, a U-sensing genetic molecular component comprises at least one Fluoride sensing riboswitch. In these embodiments, in the U-sensitive F-sensitive genetic circuit at least one molecular component is a reportable molecular component and/or a U- neutralizing molecular component when the genetic circuit operates according to the circuit design in presence of bioavailable U and bioavailable Fluoride.
[00330] Exemplary configurations of an F-sensing bioswitch and a U-sensing genetic molecular component in a UO2F2 biosensors are shown in Examples 23 to 25.
[00331] In some embodiments, the UO2F2 comprise a U-sensing genetic molecular component and/or a U-sensing genetic circuit in combination with a separate F-sensing genetic molecular component and/or a F-sensing genetic circuit to provide a dual output configuration wherein one or more fluoride riboswitches control the expression of a reportable molecular component configured to be independent of U. Exemplary dual output configurations of the UO2F2 biosensor herein described are showing in Example 25.
[00332] In some embodiments, the UO2F2 biosensor comprise a U-sensing genetic molecular component and/or a U-sensing genetic circuit in combination with a F-sensing riboswitch integrated in the U-sensing genetic molecular component and/or a U-sensing genetic circuit to provide a single output configuration wherein one or more fluoride riboswitches control the output of the U-sensing genetic molecular component and/or the U-sensing genetic circuit. In particular in those embodiments, the fluoride riboswitch functions to prematurely terminate transcription from a U-activated promoter in the absence of fluoride. Fluoride binding mitigates this termination, enabling gene expression thus providing an output in presence of bioavailable U and Fluoride. Exemplary single output configurations of the UO2F2 biosensor herein described are shown in Example 24.
[00333] Exemplary configurations of UO2F2 biosensors herein described are reported below and in the Examples sections.
[00334] In particular, an exemplary genetic circuit described herein comprises, a U sensing genetic molecular component in which a U-sensitive promoter such as Pphyt or a P1361 is configured to initiate expression of a lacZ gene encoding the beta-galactosidase enzyme (U- sensing genetic molecular component), wherein the beta-galactosidase enzyme converts the substrate ONPG (cellular molecular component) to yield galactose and o-nitrophenol which has a yellow color (reportable molecular component). In some of these embodiments, the genetic molecular components can further comprise a Fluoride sensing riboswitch (e.g. any one of the crcB riboswitches herein described) in a configuration which allows the expression of the lacZ gene in presence of an effective amount of bioavailable fluoride in a single output configuration as will be understood by a skilled person. In addition or in the alternative, the UO2F2 biosensor can further comprise the F-sensing riboswitch in a separate F-sensing reportable genetic molecular component and/or a separate F-sensing genetic circuit to provide a dual output configuration.
[00335] In some embodiments of the UO2F2 biosensor, the U-sensitive genetic circuit comprises at least one U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U- sensitive transcriptional 1362 binding site and the U sensitive genetic circuit is comprised in the biosensor together with an F-sensitive reportable genetic molecular component comprising an F- sensitive riboswitch herein described configured to be expressed in presence of an effective amount of bioavailable Fluoride in a dual output configuration. In these embodiments, in the U- sensitive genetic circuit at least one molecular component is a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design in presence of bioavailable U as will be understood by a skilled person. In some embodiments, the U-sensitive genetic circuit further comprises at least one genetic molecular component in which a UzcRS two-component system regulated promoter is activated or repressed in presence of bioavailable U.
[00336] In an exemplary embodiment of the U-sensitive genetic circuit described herein, at least one U-sensitive genetic molecular component comprises a U-sensitive promoter such as Pphyt or a Pi36i configured to initiate expression of a uzcS gene (CCNA_02842) and a uzcR gene (CCNZ_02485), encoding proteins UzcS and UzcR, respectively (first U-sensing genetic molecular component) and further comprises a second U-sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of a reporter gene (e.g. GFP, an exemplary reportable molecular component), in which binding of UzcR protein to the UzcRS -regulated promoter activates the UzcRS -regulated promoter (see Example 2). As understood by those skilled in the art, uzcS and uzcR are genes that are natively comprised in a uzcRS operon in Caulobacter and other alphaproteobacteria, as described in Park et ah, 2017 [3].
[00337] The term“uzcRS operon” as used herein refers to the genetically encoded UzcRS two- component system, comprising the uzcS gene (CCNA_02842) and the uzcR gene (CCNZ_02485), and operatively linked promoters and regulatory elements [3]. In an exemplary C. crescentus NA1000 genome, uzcR and uzcS are physically separated by genes parD3 and parE3 encoding the ParDE3 toxin antitoxin (TA) system, together forming a putative four-gene operon [3, 112]. Although uzcR and uzcS are conserved throughout alphaproteobacteria, the insertion of parDE3 between uzcR and uzcS is unique to a subset of the Caulobacter genus; uzcR and uzcS are adjacently located in the majority of closely related alphaproteobacteria [3] including C. crescentus OR37, an environmental isolate from a U-contaminated site [97].
[00338] Accordingly, in some embodiments of the UO2F2 biosensors described herein, the U biosensor can be any genetically engineered proteobacteria and in particular a genetically engineered Caulobacteridae which comprises a UzcRS two component system. [00339] In some preferred embodiments, a U- sensitive genetic circuit further comprises one or more genetic molecular components comprising one or more negative regulators of UzcRS that function to maintain UzcRS in an OFF state in absence of metal (see Example 9). In particular, in some exemplary embodiments, the UO2F2 biosensor described herein comprises a chromosomal copy of UzcRS negative regulators 1 and 2 (Example 9) under transcriptional regulation of their native promoters. In other embodiments, one or more genetic molecular components comprising exemplary UzcRS negative regulators 1 and 2 are placed under control of inducible transcriptional regulatory elements in order to desensitize UzcS to a particular signal. For example, overexpression of CCNA_03680-CCNA_03681 reduces the sensitivity of UzcRS for Zn and Cu. In some embodiments, the MarR-type regulators CCNA_03498 and/or CCNA_02289 can be deleted in the host Caulobacteridae genome to increase sensitivity of uzcRS for U.
[00340] In some of the embodiments of the U- sensitive genetic circuit in which at least one U- sensitive genetic molecular component comprises a U-sensitive promoter such as Pphyt or a P1361 configured to initiate expression of a uzcS gene (CCNA_02842) and a uzcR gene (CCNZ_02485), a fluoride riboswitch can be integrated between the UzcRS -regulated promoter and the GFP gene and/or between the Pphyt or P1361 promoter and the uzcS ATG to provide a U- sensing F-sensing genetic circuit in a single output configuration. In addition or in the alternative, a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
[00341] In some cases of these embodiments or other embodiments herein described, riboswitch performance is expected to be adversely affected by changes in genomic context, such as fusing the riboswitch to a reporter gene. Accordingly, in any embodiments herein described, if an initial transcriptional fusion to a reporter (such as GFP) fails, a Ribo-attenuator, a recently developed genetic element that insulates the riboswitch from genomic context and enhances the tunability (e.g., sensitivity, cell-to-cell variability, etc.), will be used to build a riboswitch reporter (e.g. a CrcB reporter). This approach was successfully applied with various elements such as in construction of GFP fusions with the addA (2-aminopurine) and btuB (adenosylcobalamin) riboswitches [113]. [00342] In an exemplary embodiment of the U-sensitive F-sensitive genetic circuits described herein, the U-sensitive genetic circuit comprises a U-sensitive promoter such as Pphyt or a P i % i configured to initiate expression of a hrpS gene encoding an HrpS protein (first U-sensing genetic molecular component) and further comprises a second U-sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of an hrpR gene encoding an HrpR protein. The U-sensitive F-sensitive genetic circuit further comprises a fourth genetic molecular component comprising a hrpL promoter (PhrpL) configured to initiate expression of a reporter gene (e.g. GFP, an exemplary reportable molecular component), in which binding of both HrpS and HrpR are required for s54 dependent activation of the PhrpL and expression of HrpS or HrpR alone is not sufficient for transcriptional activation (see Example 3). In these embodiments a fluoride riboswitch could be integrated in three positions 1) Between the Pphyt or P i % i promoter and hrpS, 2); 2) between the UzcRS- regulated promoter and the hrpR; and/or between the PhrpL promoter and gfp to provide a U- sensitive F-sensitive genetic circuit with a single output configuration. In the addition or in the alternative, a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
[00343] As understood by those skilled in the art, hrpR and hrpS refer to genes that are natively comprised in a s54 dependent hrpR/hrpS hetero-regulation module from the hrp (hypersensitive response and pathogenicity) system for Type III secretion in Psuedomonas syringae [114-116], wherein both HrpS and HrpR are required for s54 dependent activation of the hrpL promoter (Phi L) and expression of HrpS or HrpR alone is not sufficient for transcriptional activation.
[00344] In some embodiments of the U-sensitive and/or F-sensitive genetic circuits described herein, at least two genetic molecular components comprise complementary protein fragments of a transcription factor, which are configured to associate together to form a functional transcription factor, such as those based on a bacterial two-hybrid system.
[00345] As would be understood by those skilled in the art, the term“bacterial two-hybrid” as used herein refers to a technique used to detect protein-protein interactions and protein-DNA interactions by testing for physical interactions (such as binding) between two proteins or a single protein and a DNA molecule, respectively. As understood by those skilled in the art, the bacterial two-hybrid system relies on the activation of downstream reporter gene(s) upon binding of a transcription factor onto an upstream activating sequence (UAS), wherein the transcription factor is split into two separate fragments, called the binding domain (BD) and activating domain (AD). The BD is a domain configured to bind to the UAS and the AD is a domain configured to activate transcription of the operatively linked gene. Thus, the bacterial two-hybrid system is a protein-fragment complementation assay that requires both the BD and the AD for reporter expression. An exemplary bacterial two-hybrid system utilizes an E. coli omega protein, which copurifies with RNA polymerase, and can function as a transcriptional activator when linked covalently to a DNA-binding protein. The E. coli omega protein can function as an activation target when this covalent linkage is replaced by a pair of interacting polypeptides fused to the DNA-binding protein and to omega, respectively [117].
[00346] Accordingly, in an exemplary embodiment of a U-sensitive genetic circuits described herein, the U-sensitive genetic circuit comprises a U-sensitive promoter such as Pphyt or a P i % i configured to initiate expression of a BD gene encoding an BD protein (first U- sensing genetic molecular component), and further comprises a second U- sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of an AD gene encoding an AD protein. The U-sensitive genetic circuit further comprises a fourth genetic molecular component comprising a promoter comprising binding sites for the BD and the AD, configured to activate expression of the reporter gene, e.g. GFP or the U- neutralizing gene (third genetic molecular component), in which binding of both BD and AD are required activation of the third genetic molecular component and expression of the reportable molecular component and/or the U-neutralizing molecular component (see Example 5). For example, it is expected that a 1363/1362 regulated promoter (e.g., Pphyt) can be used to initiate expression of alpha-gall lp (an exemplary AD), and a UzcRS regulated promoter (e.g., P UTCB) can be used to initiate expression of k-gal4 (an exemplary BD) and P iacOR2-62 that contains the UAS to initiate expression of a reporter gene, such as GFPmut3 (see Example 5). In those embodiments a fluoride riboswitch could be integrated in three positions: 1) Between the Pphyt or P i % i promoter and BD; 2) Between the UzcRS -regulated promoter and AD and/or 3) ) Between the BD/AD-regulated promoter and gfp to provide a U-sensitive F-sensitive genetic circuit with a single output configuration. In the addition or in the alternative, a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
[00347] In some embodiments of the U biosensor, the U-sensitive genetic circuit comprises at least one U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U-sensitive transcriptional 1362 binding site and/or a UzcR binding site together. In these embodiments, in the U-sensitive genetic circuit at least one molecular component is a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design in presence of bioavailable U, wherein the reportable genetic component and/or a U-neutralizing molecular component is formed by an assembly of two or more subunits of the reportable molecular component and/or a U-neutralizing molecular component. In some embodiments where a 1362 binding site is present, the U-sensitive genetic circuit further preferably comprises at least one genetic molecular component in which an UzcRS two-component system regulated promoter is activated or repressed in the presence of bioavailable U and bioavailable F. In those embodiments, a Fluoride riboswitch could be integrated in two positions: 1) between the Pphyt or P1361 promoter and U neutralizing gen; and/or 2) Between the UzcRS -regulated promoter and U neutralizing gene to provide a U-sensitive F- sensitive genetic circuit with a single output configuration. In the addition or in the alternative, a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
[00348] In an exemplary embodiment of the U-sensitive genetic circuits described herein (see Example 4), the U-sensitive F-sensitive genetic circuit comprises a U-sensitive promoter such as Pphyt or a Pi36i configured to initiate expression of a gfplO-Kl fusion gene encoding a GFP10-K1 fusion protein (first U-sensing genetic molecular component) and further comprises a second U- sensing genetic molecular component in which a UzcRS two-component system regulated promoter is configured to initiate expression of an El-gfpll gene encoding E1-GFP11 fusion protein. The U-sensitive genetic circuit further comprises a third genetic molecular component comprising a gfpl-9 gene regulated by a non-U-responsive, xylose inducible promoter (Pxyi), a constitutively active promoter (e.g., PrSaA) or by Pphyt/Pi36i. In some of those embodiments, a Fluoride riboswitch can be integrated in three positions: 1) between the Pphyt or P i % i promoter and gfplO-Kl 2) between the UzcRS -regulated promoter and El-gfpll and/or 3) Between constitutive promoter and gfp-1-9 to provide a U-sensitive F-sensitive genetic circuit with a single output configuration. In the addition or in the alternative, a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
[00349] Accordingly, the term“tripartite GFP” as used herein refers to a GFP reporter that requires expression and assembly of GFP10, GFP11 and GFP1-9 together for GFP reporter function [5]. As understood by those skilled in the art, tripartite GFP assembly and reporter function is based on tripartite association between two twenty amino-acids long GFP tags, GFP10 and GFP11, which are fused to interacting protein partners, in addition to a complementary GFP 1-9 detector. When the interacting protein partners interact, GFP 10 and GFP11 self-associate with GFP1-9 to form a functional GFP [5]. In embodiments described herein, any protein interaction pair can be used for the tripartite system. Exemplary interacting protein partners comprise oppositely charged Kl/El coiled coils, FKBP12-FRB rapamycin inducible protein interaction [5], or the leucine zipper of GCN4 [118] among others known to those skilled in the art.
[00350] In some embodiments of the UO2F2 biosensor, a U-sensitive genetic circuit can comprise at least one U- sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in the presence of bioavailable U, the U sensitive promoter comprising a U-sensitive transcriptional 1362 binding site and/or a UzcR binding site. In these embodiments, in the U-sensitive genetic circuit at least one molecular component is a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design in the presence of bioavailable U, wherein the reportable molecular component and/or the U-neutralizing molecular component is post-transcriptionally and/or post-translationally converted by the U-sensitive F-sensitive genetic circuit in presence of bioavailable U. In some embodiments wherein at least one of the U sensing promoter is 1362 binding sites, the U-sensitive genetic circuit further preferably comprises at least one genetic molecular component in which a UzcRS two-component system regulated promoter is activated or repressed in presence of bioavailable U and bioavailable F.
[00351] Accordingly, in an exemplary embodiment of the U sensing genetic molecular components described herein, the U-sensitive genetic molecular component comprises a U- sensitive promoter such as Pphyt or a P i % i configured to initiate expression of a protease configured to cleave at a cleavage sequence comprised in a linker peptide in a Forster resonance energy transfer (FRET) sensor protein (first U-sensing genetic molecular component), and further comprises a second U-sensing genetic molecular component in which a UzcRS two- component system regulated promoter is configured to initiate expression of the FRET sensor protein. Thus, in presence of bioavailable U, the U-sensitive genetic circuit is configured to express and cleave the FRET sensor protein. In those embodiments, a fluoride sensing riboswitch can be integrated in two positions: 1) between the Pphyt or Pi36i promoter and FRET protein 1 and/or 2) between the UzcRS -regulated promoter and FRET protein 2 to provide a U- sensitive F-sensitive genetic circuit with a single output configuration. In the addition or in the alternative, a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
[00352] As understood by those skilled in the art, the terms“Forster resonance energy transfer”, “FRET”, “fluorescence resonance energy transfer”, “resonance energy transfer”, “RET” or “electronic energy transfer”, and“EET” as used herein refers to a mechanism describing energy transfer between two light-sensitive molecules (chromophores) [119]. A donor chromophore, initially in its electronic excited state, can transfer energy to an acceptor chromophore through nonradiative dipole-dipole coupling [120]. The efficiency of this energy transfer is inversely proportional to the sixth power of the distance between donor and acceptor, making FRET extremely sensitive to small changes in distance. Measurements of FRET efficiency can be used to determine if two fluorophores are within a certain distance of each other [121]. Such measurements can be used as a research tool in fields such as biology and chemistry. For example, one common pair fluorophores for biological use is a cyan fluorescent protein (CFP) - yellow fluorescent protein (YFP) pair [122]. Both are color variants of green fluorescent protein (GFP). Thus, for example, a genetically-encoded fusion of CFP and YFP covalently linked by a protease cleavage sequence can be used as a cleavage assay, wherein if the linker is intact, excitation at the absorbance wavelength of CFP (414nm) causes emission by YFP (525nm) due to FRET. If the linker is cleaved by a protease, FRET is abolished and emission is at the CFP wavelength (475nm) [123].
[00353] In some embodiments, a UOiFibioscnsor herein described comprises two or more U- sensitive genetic molecular components and/or U-sensitive genetic circuits, wherein each of the U-sensitive genetic molecular components expresses a different reporter gene and/or a U- neutralizing gene, and/or each of the U-sensitive genetic circuits comprise a different reportable molecular component and/or U-neutralizing molecular component, in presence of bioavailable U. For example, exemplary different reportable molecular components can comprise a first genetically-encoded reporter (e.g., GFP) and a second, different genetically-encoded reporter (e.g., dsRED). In those embodiments, a Fluoride riboswitch can be integrated in two positions: 1) between the Pphyt or P i % i promoter and reportable molecular component 1 and/or U-neutralizing molecular component 1 ; and/or 2) between the UzcRS -regulated promoter and reportable molecular component 2 and/or U-neutralizing molecular component 2 to provide a U-sensitive F- sensitive genetic circuit with a single output configuration. In the addition or in the alternative, a riboswitch can be comprised in the UO2F2 biosensor in a separate F-sensing reportable genetic molecular component and/or in a separate F-sensing genetic circuit herein described in a dual output configuration.
[00354] In some embodiments, the U02F2biosensors described herein can comprise a UzcRS U- sensitive genetic molecular component comprising a reporter gene or a U-neutralizing gene operatively connected to a UzcRS-regulated promoter, such as PurcA, PurcB, and others identifiable by those skilled in the art, such as those described in Park et ah, (2017) [3]. In those embodiments, the UzcRS U-sensitive genetic molecular component can be comprised in the biosensor alone or in combination with a 1363/1362 U sensitive genetic molecular component comprising a reporter gene or a U neutralizing gene operatively connected to a 1362 regulated promoter. In those embodiments, a Fluoride riboswitch can be integrated in two positions: 1) Between the Pphyt or P1361 promoter and reportable molecular component 1 and/or U-neutralizing molecular component; and/or 2) Between the UzcRS-regulated promoter and reportable molecular component 2 and/or U-neutralizing molecular component 2. [00355] In the UO2F2 biosensors described herein, one or more genetic molecular components of the U-sensitive F-sensitive genetic circuits described herein can comprise genomic DNA of the proteobacterial cell and in particular of the Caulobacteridae cell. The one or more genetic molecular components can comprise native genomic DNA in the Caulobacteridae cell or can be introduced into the genome of the Caulobacteridae cell through genetic engineering, or comprised in the Caulobacteridae cell in one or more extra-genomic polynucleotides or vectors, using standard genetic engineering methods known to those skilled in the art and described herein.
[00356] In embodiments described herein, the U02F2biosensors-can detect uranium in a range dependent on the composition of the growth media since media components influence bioavailability. The bioavailability of both compounds will be dictated by the composition of the growth media / the chemical composition of the environmental sample
[00357] In the U02F2biosensors herein described comprising one or more U-sensitive genetic molecular components, in the absence of bioavailable U, 1362 is not bound to the 1362 binding site of a 1362 U-sensitive genetic molecular component and/or UzcR is not bound to an m_5 site of the UczR U-sensitive genetic molecular component, and reporter gene and/or U-neutralizing gene is not expressed. In a second target range, in presence of bioavailable U, 1362 is bound to the 1362 binding site of the U-sensitive genetic molecular component and/or UzcR is not bound to an m_5 site of the U-sensitive genetic molecular component, and the reporter gene and/or U- neutralizing gene is expressed.
[00358] In the U02F2biosensors herein described comprising one or more F-sensitive genetic molecular components, in the absence of bioavailable F, the fluoride riboswitch will not bind F and transcription will be prematurely terminated. In the presence of bioavailable F, the fluoride riboswitch will bind F and transcription initiation will occur.
[00359] In some embodiments, a U-sensing genetic molecular component and/or U sensing genetic circuits herein described comprising UzcR binding site and a UzcRS TCS can detect Uranium in ~ 1 micromolar concentrations as will be understood by a skilled person upon reading of the present disclosure. In some embodiments a U-sensing genetic molecular component and/or U sensing genetic circuits herein described comprising 1362 binding site and a 1363/1362 TCS can detect Uranium at -500 nM in aqueous conditions lacking Pi or glycerol-phosphate as will be understood by a skilled person upon reading of the present disclosure. In embodiments herein described wherein the U-sensing genetic circuit is integrated with a F-sensing riboswitch in a single output configuration and/or an F sensing molecular component or circuit in a dual output configuration, the biosensor U sensitivity is expected to be about the same and therefore can detect Uranium at -500 nM in aqueous conditions lacking Pi or glycerol-phosphate as will be understood by a skilled person upon reading of the present disclosure
[00360] In the UOiFibioscnsors herein described comprising a U-sensitive and/or F-sensitive genetic circuit comprising one or more F sensing riboswitches activated in presence of bioavailable F, one or more U-sensing genetic molecular components in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U-sensitive transcriptional 1362 binding site and/or a UczR binding site, in a first target range of bioavailable U the endogenous proteobacteria U-sensitive transcriptional regulator is not bound to the 1362 binding site and/or the UczR binding site of the U-sensitive promoter and the U-sensitive F-sensitive genetic circuit does not comprise a reportable molecular component and/or a U-neutralizing molecular component, when the genetic circuit operates according to the circuit design. In a second target range of bioavailable U and bioavailable F, the endogenous proteobacteria U-sensitive transcriptional regulator is bound to the 1362 binding site and/or to the UzcR binding site of the U-sensitive promoter and the U-sensitive genetic circuit comprises a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design. In those embodiments, the fluoride riboswitch will not bind F and transcription will be prematurely terminated. In the presence of bioavailable F, the fluoride riboswitch will bind F and transcription initiation will occur, resulting in activation of the reportable element and the single output of the circuit.
[00361] In some of those embodiments, a U-sensitive genetic circuit which comprises a 1362 binding site further comprises at least one genetic molecular component comprising a UzcRS- regulated promoter comprising a UzcR binding site. In these embodiments, in the second target range of bioavailable U concentration, activation or repression of the UzcRS -regulated promoter is also required as an input for the U-sensitive genetic circuit to comprise a reportable molecular component and/or a U-neutralizing molecular component when the genetic circuit operates according to the circuit design, wherein activation or repression of the U-sensitive promoter comprising the regulator direct repeat together with activation or repression of the UzcRS- regulated promoter is herein referred to as an“AND gate”, wherein the term“AND” is an operation of Boolean logic.
[00362] As would be understood by persons skilled in the art, Boolean logic is a branch of algebra in which the values of the variables are the truth values ‘true’ and ‘false’, usually denoted by the digital logic terms‘G and Ό’ respectively. Instead of elementary algebra where the values of the variables are numbers, and the main operations are addition and multiplication, the main operations of Boolean logic are the conjunction‘AND’, the disjunction OR’, and the negation‘NOT’. As understood by those skilled in the art, it is thus a formalism for describing logical relations in the same way that ordinary algebra describes numeric relations. The term “AND gate” refers to a digital logic gate that implements logical conjunction - it behaves according to the truth table shown in Table 1. A‘true’ output (1) results only if both the inputs to the AND gate are‘true’ (1). If neither or only one input to the AND gate is‘true’ (1), a‘false’ (0) output results. Therefore, the output is always 0 except when all the inputs are 1.
Table 1.‘AND gate’ truth table:
Figure imgf000115_0001
[00363] In particular, the term“AND gate” as used herein refers to the logical relation between two genetic molecular components in a U-sensitive genetic circuit, wherein inputs‘A’ and Έ’ in Table 1 are two independently activated or repressed genetic molecular components, wherein a first independently activated or repressed genetic molecular component comprises a first promoter having a U-sensitive transcriptional 1362 binding site, and a second independently activated or repressed genetic molecular component comprises a UzcRS-two-component system regulated promoter, and the output ‘A AND B’ in Table 1 is the reportable molecular component and/or a U-neutralizing molecular component of the U-sensitive genetic circuit.
[00364] As would be understood by those skilled in the art, any‘AND gate’ genetic system can be employed in the U-sensitive genetic circuits described herein, such as those described in [124, 125]
[00365] In some embodiments, the U-sensitive genetic circuits described herein comprise U- sensing genetic molecular components whose expression is regulated independently by (1) a U- sensitive promoter comprising a 1362 binding site, such as Pphyt or P i % i and (2) UzcRS two- component system, and comprise an AND gate wherein‘inputs’ of activation or repression of both (1) a U-sensitive promoter comprising a 1362 binding site, such as Pphyt or P i % i and (2) UzcRS two-component system-regulated promoter is required for the reportable molecular component‘output’ and/or the U-neutralizing molecular component‘output’ according to the U- sensitive genetic circuit design (Figure 3).
[00366] In some embodiments of the U-sensitive and/or F-sensing genetic circuit, an F-sensing component and/or a U-sensing genetic molecular component, independently activated or repressed, are arranged‘in series’ in the U-sensitive genetic circuit. In an‘in-series’ AND gate of a U-sensitive F-sensitive genetic circuit, output of the reportable molecular component and/or the U-neutralizing molecular component according to the genetic circuit design requires two or more of the independently activated or repressed F-sensing genetic molecular component and U- sensing genetic molecular components to be activated or repressed in temporal succession, wherein the activation or repression of the F-sensing genetic molecular component precedes the activation or representation of a first independently activated or repressed U-sensing genetic molecular component which on its turn can precede the activation or repression of a second independently activated or repressed U-sensing genetic molecular component.
[00367] The term“in series” as used herein refers to a genetic circuit in which genetic molecular components are connected through biochemical reactions along a single linear circuit path. With regard to Table 1, in an ‘in series’ AND gate, the temporal sequence of the activation or repression of the independently activated U-sensing genetic molecular components denoted by inputs ‘A’ and‘B’ is such that the second input‘B’ is dependent on a prior activation or repression of the first input‘A’ in linear succession according to the genetic circuit design.
[00368] In exemplary embodiments described herein, an in- series AND gate comprised in a U- sensitive genetic circuit comprises, a first U-sensing genetic molecular component comprising a U-sensitive promoter having a 1362 binding site, such as a Pi36i promoter or a Pphyt promoter, and optionally further comprises a UzcRS two component system-dependent promoter, wherein expression of a uzcR and uzcS genes are under the transcriptional control of a U-sensitive promoter comprising a 1362 binding site, such as a Pi36i promoter or a Pphyt promoter (see Example 2).
[00369] In some embodiments of the U-sensitive and/or F-sensing genetic circuits herein described, at least two independently activated or repressed U-sensing genetic molecular components are arranged‘in parallel’ in the U-sensitive genetic circuit, herein referred to as an ‘in parallel’ AND gate. With regard to Table 1, in contrast to the‘in series’ AND gate, in an‘in parallel’ AND gate, the temporal sequence of inputs‘A’ and‘B’ is such that a second input‘B’ is not dependent on a prior first input‘A’ in linear succession, but rather inputs‘A’ and‘B’ can occur simultaneously. In other words, the term“in parallel” as used herein refers to a genetic circuit in which genetic molecular components are connected through biochemical reactions along more than one circuit path.
[00370] In an‘in parallel' AND gate system, output of the reportable molecular component and/or the U-neutralizing molecular component according to the genetic circuit design requires two or more of the independently activated or repressed F-sensing genetic molecular component and/or one or more U-sensing genetic molecular components to be activated or repressed in parallel.
[00371] In several embodiments herein described, the in-parallel AND gate comprised in a U- sensitive genetic circuit comprises at least two independently activated or repressed U-sensing genetic molecular components, each comprising a different promoter, (1) a U-sensitive promoter comprising a 1362 direct repeat binding site, such as a Pi36i or Pphyt, and (2) a UzcRS -regulated promoter, wherein promoters (1) and (2) act as independently activated or repressed parallel inputs into the U-sensitive genetic circuit, functioning as two independent points of U-sensing in response to their respective U-sensitive transcriptional regulators. In these embodiments, the reportable molecular component output and/or a U-neutralizing molecular component output of the U-sensitive F-sensitive genetic circuit is present only when both of the two independently activated or repressed F-sensing genetic molecular component and U-sensing genetic molecular component are activated or repressed according to the U-sensing genetic circuit design.
[00372] Thus, in several embodiments, a U-sensitive F-sensitive genetic circuit comprising an ‘in parallel’ AND gate for more than one U sensing genetic molecular component, can provide improved selectivity for U, as output of the reportable molecular component and/or the U- neutralizing molecular component is dependent on two independent points of U-sensing in response to two different U-sensitive transcriptional regulators. Persons skilled in the art will recognize that this also reduces the probability of a false positive output, in the first target range of bioavailable U concentration, such as in response to non-U stimuli (such as Zn or Cu).
[00373] In some embodiments described herein, an‘in parallel’ AND gate comprises an HRP AND gate. The term“HRP AND gate” as used herein refers to an AND gate system from Pseudomonas syringae that was developed in E. coli [126]. The HRP AND gate system comprises an orthogonal s54 dependent hrpR/hrpS hetero -regulation module from the hrp (hypersensitive response and pathogenicity) system for Type III secretion in Psuedomonas syringae [114-116], as described above. In the HRP AND gate system, both HrpS and HrpR are required for s54 dependent activation of the hrpL promoter (PhrpL) and expression of HrpS or HrpR alone is not sufficient for transcriptional activation. In exemplary embodiments described herein, two different promoters, (1) a U-sensitive promoter comprising a 1362 binding site, such as a P 1361 or Pphyt, and (2) a UzcRS -regulated promoter, act as independently activated inputs to initiate the transcription of hrpR and hrpS, respectively, functioning as two independent points of U-sensing in response to their respective U-sensitive transcriptional regulators (Figure 6 Panel A). In exemplary embodiments described herein, transcription of the output hrpL promoter is activated only when both proteins HrpR and HrpS bind the upstream activator sequence to remodel a closed o54-RNAP-hrpL transcription complex to an open one through ATP hydrolysis [126]. In an exemplary HRP AND gate described herein, the output shown is GFP reporter expression (Figure 6 Panel A). [00374] The performance of the HRP AND gate can be described using a Hill function for the promoter steady-state input-output response (transfer function) in the form:
Figure imgf000119_0001
Eq. (1)
where [/] is the concentration of the inducer, such as bioavailable U; Ki and m are the Hill constant and coefficient, respectively, relating to the promoter- regulator/inducer interaction; k is the maximum expression level due to induction; and a is a constant relating to the basal level of the promoter due to leakage [126] and further in the form:
/(M,I5]) = [0I/[01ΪM1£ = ([K]/¾)Bk((5:|/ίί5)¾
({l+(|gl/ICR)n*Xl + (fS]/%)¾))
Eq. (2) which describes the normalized output of the AND gate as a function of the levels of the two activator proteins ([L’] for HrpR, [5] for HrpS) at steady state. [G]max is the maximum activity observed for the output. KR, KS and HR, Tls are the Hill constants and coefficients for HrpR and HrpS, respectively [126].
[00375] In an exemplary HRP AND gate (see Example 3), the expression of hrpS is placed under the control of a U- sensitive promoter comprising a 1362 binding site, such as a Pphyt or Pi36i, while hrpR is placed under the control of PUrcB, a UzcRS -dependent promoter that has lower basal activity compared to PUrcA [3]. In Example 3, the PhrpL promoter regulating gfp expression requires Pphyt/Pi36i and PurcB to be active to generate a fluorescent signal.
[00376] In some embodiments, an‘in parallel’ AND gate comprises a tripartite GFP AND gate. As described herein, in the tripartite GFP AND gate system, reporter function is based on tripartite association between two twenty amino-acids long GFP tags, GFP10 and GFP11, which are fused to interacting protein partners, in addition to a complementary GFP1-9 detector. When the interacting protein partners interact, GFP10 and GFP11 self-associate with GFP1-9 to form a functional GFP [5].
[00377] In an exemplary tripartite GFP AND gate (see Example 4), expression of a gfplO-Kl fusion gene is placed under control of a U-sensitive promoter comprising a 1362 direct repeat binding site, such as a Pphyt or Pi36i , expression of a El-gfpll fusion gene is placed under control of a UzcRS-responsive promoter, such as PurcB, and expression of gfpl-9 is placed under control of a non-U-responsive, strong, constitutively active promoter, PrSaA·. Thus, in Example 4, assembly and function of tripartite GFP requires both U binding and activation independently of a U-sensitive promoter comprising a 1362 binding site, such as a Pphyt/Pi36i and a UzcRS- responsive promoter.
[00378] In some embodiments, an‘in parallel’ AND gate system comprises a bacterial two- hybrid AND gate.
[00379] In an exemplary bacterial two-hybrid AND gate (see Example 5). In some embodiments, an‘in parallel’ AND gate system comprises a FRET sensor AND gate.
[00380] In some embodiments, the U-sensitive genetic circuits described herein comprise a combination of two or more‘in series’ and/or‘in parallel’ AND gates as described herein, wherein the two or more AND gates are connected by activating, inhibiting, binding or converting reactions.
[00381] For example, Figure 14 shows an exemplary combination of an‘in series’ AND gate and an‘in parallel’ AND gate (see Example 8). In this exemplary embodiment, the exemplary ‘in series’ AND gate shown in Figure 4 Panel C and the exemplary‘in parallel’ tripartite GFP AND gate shown in Figure 6 Panel B are connected, such that the Pphyt-regulated UzcR no longer activates expression of GFP regulated by P1968 as in Figure 4 Panel C, but rather activates expression of El-gfpll regulated by PurcB within the tripartite GFP‘in parallel’ AND gate. As a result, in the exemplary combined configuration shown in Figure 14 is that activation of PurcB in the‘in parallel’ tripartite GFP AND gate is restricted to U-selective activation of UzcR and UzcS expression by Pphyt in the‘in series’ AND gate, whereas in the‘in parallel’ tripartite GFP AND gate shown in Figure 6 Panel B PurcB can be activated by U, Zn or Cu. Thus, an advantage of the combination of these exemplary AND gates is increased selectivity for U.
[00382] In particular in some embodiments, circuit contains the native uzcRS genes placed under the control of the Pphyt promoter - the chromosomal PI and P2 promoters are swapped with Pphyt as described herein. This could be done by deleting uzcRS and using a plasmid- based system to reintroduce these genes into the circuit. In those embodiments, the circuit leverages the improved selectivity of U over Zn observed in the sensor depicted in Figure 4 Panel C - A U/Zn ratio of 5.5.
[00383] In any one of the above embodiments a Fluoride sensing riboswitch can be included within any one of the genetic molecular components of any one of the U- sensitive genetic circuit. The fluoride riboswitch will not bind F and transcription will be prematurely terminated. In the presence of bioavailable F, the fluoride riboswitch will bind F and transcription initiation will occur, resulting in activation of the reportable element of the circuit and in providing the single output of the circuit. In the addition or in the alternative, a F sensing riboswitch can be included within a separate F-sensing reportable genetic molecular component and/or F-sensing genetic circuit. In those embodiments in the presence of bioavailable F, the fluoride riboswitch will bind F and transcription initiation will occur, resulting in activation of the separate F-sensing reportable genetic molecular component and/or the separate F-sensing genetic circuit to provide the F-sensing output of the dual output configuration of the UOiFibioscnsors herein described.
[00384] As would be understood by persons skilled in the art, in order for the UC Fi-biosensor s described herein to operate in response to bioavailable U, the U-sensitive and/or F-sensitive genetic circuits described herein comprise one or more genetic molecular components that are orthogonal to the proteobacteria cell of the UO2F2 biosensor.
[00385] In some embodiments, the circuit components of U-sensing and/or F-sensing circuit herein described are stably integrated in the host. In some embodiments, the‘in series’ AND gate described herein was constructed using the native uzcR and uzcS genes. In some of these embodiments the chromosomal PI and P2 promoters have been replaced with Pphyt.
[00386] In some embodiments of the U-sensitive genetic molecular components and/or U- sensitive genetic circuit herein described comprising UzcR binding and a UzcRS TCS, the U- sensing genetic molecular component and/or U-sensing genetic circuit can further comprise an amplifier genetic molecular component comprising a U-sensitive promoter and/or a controllable promoter operatively connected to the UzcY and/or UzcZ in a configuration wherein the U- sensitive promoter and/or a controllable promoter directly initiates expression of the amplifier molecular component (see Figure 13).
[00387] The term“UzcY” as used herein indicates a protein having the amino acid sequence:
MTRDQDTLRML AE VE A AN ADL ARRAKAPLW YHP ALGLLV G ALIA V QGQPT S ILLVFY A AYIAGLALLVRAYKRHTGLWVSGYRAGRTRWVALGLATLTMIGGVIAVWLLRERGLT A APLIFG AIV A VIVT V GGF VWE A AF RADLRDGRPL (SEQ ID NO: 33) or a sequence that when aligned with sequence SEQ ID NO: 33 has a BLAST score has a BLAST score greater than 50 and less than 100, or preferably greater than 100 and less than 200, or more preferably a BLAST score greater than 200 and a homology with SEQ ID NO: 33 less than 100%, or even more preferably BLAST score of 289 and an homology of 100% with SEQ ID NO: 33.
[00388] The term“UzcZ” as used herein indicates a protein having the amino acid sequence:
MRRGS QPMH APIS ATERS TIAQIGGRN APIG ALRAF VTLL VIAHHT VLA YTPNPPPIGDF S Q AP YLW Q AFP VRDPQKFELFGLLTLINDLFFMS LMFFIS GLFV AD GLRAKGN GGLLS GR AARLGVPFVLAAGLLAPLAYFPAWLQAGGDVSIAGFASAWLDLPSWPSGPAWFLWVL LAF G AI VTLLNLI APG VID ALGRLVRG ADRKPGLFFLGLVIAS A V A YIPMS ATFTFMH WT QLGPFT V QTS RV VH YF V YFLAG V A V G A AG V GQGLTDS EGKLAKRWW AW Q A APILP V V GVIAVIIMAFSPKPPPRVALDIGGGVMFALACATLSFAALATFLRFVKKTGPVAAS LQ AN AY GM YLTH Y VFTTWL A WLLLPQ A W GGLAKG A A VF V G ATLLS WILTM ALRRLP LLGRIL (SEQ ID NO: 34) or a sequence that when aligned with sequence SEQ ID NO: 34 has a BLAST score has a BLAST score greater than 200 and less than 475, or preferably greater than 470 and less than 700, or more preferably a BLAST score greater than 700 and a homology with SEQ ID NO: 34 less than 100%, or even more preferably BLAST score of 801 and an homology of 100% with SEQ ID NO: 34.
[00389] In particular, in U-biosensors herein described comprising UzcR binding and a UzcRS TCS, the amplifier molecular component acts as a ‘genetic signal amplifier’ configured to increase an output, e.g. expression or levels of a reportable molecular component and/or a U- neutralizing molecular component at a given bioavailable U concentration, thus enabling more sensitive detection and reporting and/or neutralizing of bioavailable U at lower concentrations.
[00390] Genes encoding activators UzcY and UczZ herein described are herein also indicated as uzcY gene or uzcY and uczZ gene or uczZ, respectively as will be understood by a skilled person.
[00391] In particular, in some embodiments the U-sensitive genetic circuit comprises one or more amplifier genetic molecular components comprising uzcY gene CCNA_03497 encoding Caulobacter Crescentus N100 UzcY (SEQ ID NO: 33) and/or uzcZ gene CCNA_02291 encoding for Caulobacter Crescentus N100 UzcZ (SEQ ID NO: 34) under a promoter that can be either a constitutively active promoter or an inducible promoter. In preferred embodiments, the promoter is constitutively active.
[00392] In some embodiments the uzcY and/or uzcZ gene are under the transcriptional regulation of a promoter comprising a regulator direct repeat, such as exemplary promoters Pphyt or Pi36i. In an exemplary embodiment described herein, the U-sensitive genetic circuit comprises CCNA_03497 placed under the control of Pphyt so that U exposure enhances the sensitivity of UzcRS for U (see Example 7).
[00393] In some embodiments of the U-sensitive genetic molecular components, and/or U- sensitive genetic circuit herein described comprising UzcR binding and a UzcRS TCS, the bacteria are capable of natively expressing endogenous MarR family repressors such as marRl and/or marR2 genes.
[00394] The term“MarRl” as used herein indicates a protein of amino acid sequence:
MSAALDPVIHAPNRLQMCCMLAAVDTIDFATVREALDVSESVLSKHVKTLEEAGYVKV KKAASDGRQRTW LS LS KPGRE ALKGHLA ALKAMM AG VPE A (SEQ ID NO: 35) or a sequence that when aligned with sequence SEQ ID NO: 35 has a BLAST score greater than 65 and less than 100, or preferably greater than 100 and less than 150, or more preferably a BLAST score greater thanl50 and a homology with SEQ ID NO: 35 less than 100%, or even more preferably BLAST score of 199 and a homology of 100% with SEQ ID NO: 35.
[00395] The term“MarR2” as used herein indicates a protein having amino acid sequence: M APRFDIS GLDD VIHGRVRLGI V A YLAS AE V ADFTELKD VLE VT QGNLS IHLRKLEE AG YVSIDKSFVGR KPLTRVRLTDTGRAAFSSYLRAMGQLVEQAGGG (SEQ ID NO: 36) or a sequence that when aligned with sequence SEQ ID NO: 36 has a BLAST score has a BLAST score greater than 75 and less than 100, or preferably greater than 100 and less than 150, or more preferably a BLAST score greater than 150 and a homology with SEQ ID NO: 36 less than 100%, or even more preferably BLAST score of 206 and a homology of 100% with SEQ ID NO: 36.
[00396] In preferred embodiments, of U02F2-biosensors herein described comprising UzcR binding and a UzcRS TCS, wherein the host is capable of natively expressing endogenous MarR family repressors such as marRl and/or marR2 genes, at least one gene of the endogenous MarR family, preferably all genes of the endogenous MarR family, is knocked out to provide an amplified U-biosensor configured to provide an amplified signal following activation of the U- sensitive genetic circuit (see Example 7 and Figure 12). In particular, by deleting MarR family repressor, uzcY expression is induced (see Figure 26), thus leading to the stimulation of UzcS activity and causing a hypersensitive output in response to low metal concentration.
[00397] Genes encoding for repressors MarRl and MarR2 herein described are herein also indicated as marRl gene or MarRl and marR2 gene or MarRl , respectively as will be understood by a skilled person.
[00398] In some embodiments, wherein the host is Caulobacter crescentus NA1000 the protein MarRl and MarR2 are the protein encoded by marRl gene CCNA_03498 found in Caulobacter crescentus NA1000 (SEQ ID NO: 35) and protein encoded by marR2 gene CCNA_02289 found in Caulobacter crescentus NA1000 (SEQ ID NO: 36) respectively.
[00399] In some embodiments of the U- sensitive genetic molecular components and/or U- sensitive genetic circuit herein described comprising UzcR binding and a UzcRS TCS, the bacteria are capable of natively expressing an UzcRS Regulating Transporter Atpase aminoPeptidase (herein also urtAP).
[00400] In particular, UrtAP is an ATP-binding cassette transporter (ABC transporters) comprising UrtA an ATPase and UrtP a peptidase which contains 13 transmembrane domains and a C-terminal peptidase domain.
[00401] The term“UrtA” as used herein indicates a protein encoded by a gene having a protein coding region adjacent in the genome and cotranscribed with urtP, the protein having amino acid sequence:
MLIIENLTH V Y GN GTR ALDE V S LTIPRGM Y GLLGPN G AGKS TLMRTI ATLQ APT S GHIRF GDID VLKHPEELRKTLG YLPQDF G V YPR V S A YDMLDHM A VLKGIS GGKERK AT VEHLL N QVNLWD VRKKAIAGFS GGMRQRFGIAQALIGDPRLIIVDEPT AGLDPEERNRFLNLLAE IGENVVVILSTHIVEDVSDLCPAMAIICNGAIVREGAPADLVAQLKGRIWKKIIDKAELEA AKARYKVIS TRLL AGRT VIHIES ETDPGDGFT A VEGGLED V YF S TLS S TRS RQ A A (SEQ ID NO: 143) or a sequence that when aligned with sequence SEQ ID NO: 143 has a BLAST score greater than 250 and less than 500, or preferably greater than 500 and less than 599, and a homology with SEQ ID NO: 143 less than 100%, or even more preferably BLAST score of 599 and a homology of 100% with SEQ ID NO: 143. The term“urtP” as used herein indicates a protein of amino acid sequence:
MFGKI AGFELRY QLKS P VFW V V A VIFFLMTF G A ATIDQIRIGGGGNIHKN AP Y AI AQTHLI L AIFYMF VTT AF V AN V V VRDDETGFGPILRS TRIRKFD YL Y GRFT G AFLA A AIS FLV VPL AIF V GS FMPW VDPERLGPNDLN A YLFS YF ALALP AILLT S AIFF AL AT VTRS MM WT Y V G VIAFLVLYIIAGIALDRPEYEKGAALWEPLGTAAFGLATKYWTASERNSLTPPLAGALLF NRVF VLVL A AGFL AL A Y S LFRF QS AELS GQRKS AKKTKA APTE A AP A AS GPLPTP VFDR RT AW AQLV VRTRLDMGQ VFKS P AFFVLLFLGL AN AMG ALWFATE AGRY GG V V YP VT RILLFPLLGS F GLIPIII AIY Y S GEL VWREREKKTHEIID ATP VPD W AF V APKTL AIS L VLIS T LLISVVAAMLSQVFHGYFNFELEKYLLWYVLPQALDFILLAVLAVFLQTISPHKFIGWAL MVIYIVSTITFTNLGFEHKLYNYGATTETPFSDMNGLGKFWMGAWWLRAYWTAFALV LLVLA Y GLWRRGTES RLLPRLRRLPLRLN GG AG ALMG V S L V AF AGLGGFIY VNTN VWN EYRTNIDGEKWQAEYEKTLLPFENTPQPKIIAQTLDIDIQPHAPSLETKGSYVLENKTGAA LKEIHVRFDRDLEVKGLSIEGARPKKTFEKFNYRIFAFDTPMAPGEQRKMSFITLRAQRG FPN S G AETR V VDN GTF VNNLEI APILGMS RDGLLTDRAKRRKY GLPPEQRM AKLGD V S S MQFNGLRKDADFIQSDITVTTVADQTPIAPGYKVSDSVRNGRRTARFVTEAPIMPFVSIQ S ARYKV AEET YKG V QLA V Y YDPQH A WNIDRMKTS MKRS LD YMGTNF S P Y QFRQLR Y
QEFPDYAQFAQSFANTIPWSEGMFFISDYRDPTKIDMVTYVGAHEIGHQWWAHQVIGA
NQQGGAMFSETFAQYSAFMVMKHTYGEDQIRKFFKFEFDSYFRARGGDVIDEQPFYKV
ENQPYIYYRKGSLVMYRLQDQIGEEAVNRALRKLIADHAFKGAPYPTTLDFMAALRAE
APADKQALITDLFEKITLYDLKTKSAAVKKRADGKFDVTVVVEAQKKYADGKGKETV
AALNETMEIGLFT AKPGDKGFV AKNVVLY QRRPIRS GENTFTFIVDKAPTFAGIDPYNTV
IDRNGDDNT VKVGG (SEQ ID NO: 144) or a sequence that when aligned with sequence SEQ ID NO: 144 has a BLAST score greater than 600 and less than 800, or preferably greater than 800 and less than 1200 or greater than 1000 and less than 1200, or more preferably a BLAST score greater thanl200 and less than 2000, or more preferably a BLAST score equal to or higher than 2000 or even more preferably BLAST score of 2436 and a homology of 100% with UrtP sequence from Caulobacter crescentus NA1000 or Caulobacter crescentus CB 15 and in particular with SEQ ID NO: 144.
[00402] In preferred embodiments, of U02F2-biosensors herein described comprising UzcR binding and a UzcRS TCS, wherein the host is capable of natively expressing endogenous urtAP, at least one gene of the endogenous urtAP, preferably all genes of the endogenous urtAP, is knocked out to provide an amplified U-biosensor configured to provide an amplified signal following activation of the U-sensitive genetic circuit. In particular, by deleting the endogenous urtAP, it is expected that the UzcRS TCS will be stimulated and exhibit greater uranium sensitivity.
[00403] Genes encoding for UrtA and UrtP herein described are herein also indicated as urtA gene or UrtA and urtP gene or UrtP, respectively as will be understood by a skilled person.
[00404] Detection of a host capable of expressing endogenous urtAP can be performed by detecting UrtP or urtP as urtA is more conserved as will be understood by a skilled person.
[00405] In some embodiments, wherein the host is Caulobacter crescentus NA1000 the protein UrtA and UrtP are the protein encoded by UrtA gene CCNA_03681 found in Caulobacter crescentus NA1000 (SEQ ID NO: 143) and protein encoded by UrtP gene CCNA_03680 found in Caulobacter crescentus NA1000 (SEQ ID NO: 144) respectively. [00406] Additional features of the UO2F2 which are derived from whole cell biosensors in the sectors of heavy metal detection in aqueous systems [127-130], pathogen detection and eradication [131], organic pollutant detection [132], and standoff detection of landmines [133, 134] comprise genetic modification to boost sensor performance (sensitivity, detection limit, etc.
[135, 136]), a microfluidic chemostat platform that maintains cell viability and sensor function for at least a week [137], optical transducer components that convert cell fluorescence into an electronic signal that is transmitted on a mobile phone network [138], and an automated water sampling feature [138].
[00407] Additional exemplary features comprise standoff detection functions [133, 134]and couples sensor cells, which are encapsulated within a polymeric matrix and placed in proximity of a suspicious region, with a novel optical scanning system to enable remote detection at a distance of up to 50 meters in optically noisy aqueous and soil environments [133, 134].
[00408] In some embodiments herein described, a method to provide a U biosensor is provided, the method comprising
genetically engineering a bacterial cell capable of natively and/or heterologously expressing a histidine kinase 1363 and/or histidine kinase UzcS, and U sensitive response regulator 1362 and/or U sensitive response regulator UzcR in combination with a heterologous F-sensing riboswitch, the genetically engineering performed by introducing into the
one or more U-sensing genetic molecular components configured to report and/or neutralize U herein described,
an F-sensitive riboswitch within the one or more U-sensing genetic molecular components configured to report U,
an F-sensitive riboswitch within an F-sensing reportable genetic molecular component;
one or more genetic molecular components of an F sensitive genetic circuit described herein, and/or
one or more genetic molecular components of the U-sensitive F-sensitive genetic circuits described herein,
to provide a U02F2-biosensor according to any one of the embodiments herein described as will be understood by a skilled person upon reading of the present disclosure, and
optionally operatively connecting one or more of the U biosensors to an electronic signal transducer adapted to convert a U biosensor reportable molecular component output into an electronic output.
[00409] Polynucleotide and protein molecules as described herein can be genetically engineered using recombinant techniques known to those of ordinary skill in the art. In particular, in some embodiments, Fluoride sensing riboswitch, a promoter comprising the U-sensitive 1362 binding site and/or an UzcR binding site can be genetically engineered by introducing into a polynucleotide comprising a promoter DNA sequence a polynucleotide comprising the Fluoride sensing riboswitch, the U-sensitive 1362 binding site and/or a UzcR binding site, as described herein. In other embodiments, a Fluoride sensing riboswitch, a U-sensitive promoter comprising the 1362 binding site and/or a UzcR binding site can be genetically engineered by de novo designing a synthetic promoter DNA sequence comprising the Fluoride sensing riboswitch, the 1362 binding site and/or the UzcR binding site.
[00410] Production and manipulation of the polynucleotides described herein are within the skill in the art and can be carried out according to recombinant techniques described, for example, in Sambrook et al. 1989. Molecular Cloning: A Laboratory Manual, 2d ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. [139] and Innis et al. (eds). 1995. PCR Strategies, Academic Press, Inc., San Diego, [140] which are incorporated herein by reference.
[00411] It is understood that terms herein referring to nucleic acid molecules such as “polynucleotide” and“nucleotide sequence” comprise any polynucleotides such as DNA and RNA molecules and include both single-stranded and double-stranded molecules whether it is natural or synthetic in origin.
[00412] The term“polynucleotide” as used herein indicates an organic polymer composed of two or more monomers including nucleotides, or analogs thereof. The isoelectric point of a polynucleotide in the sense of the disclosure is less than 7 as will be understood by a skilled person. The term“nucleotide” refers to any of several compounds that consist of a ribose or deoxyribose sugar joined to a purine or pyrimidine base and to a phosphate group and that is the basic structural unit of nucleic acids. The term“nucleotide analog” refers respectively to a nucleotide in which one or more individual atoms have been replaced with a different atom or with a different functional group. Accordingly, the term“polynucleotide” includes nucleic acids of any length, and in particular DNA, RNA, analogs and fragments thereof. A polynucleotide of three or more nucleotides is also called “nucleotidic oligomer” or “oligonucleotide”. In particular, polynucleotides in the sense of the disclosure comprise biological molecules comprising a plurality of nucleotides. Exemplary nucleic acids include deoxyribonucleic acids, ribonucleic acids, and synthetic analogues thereof, including peptide nucleic acids. Polynucleotides can typically be provided in single-stranded form or double- stranded form and in liner or circular form as will be understood by a person of ordinary skill in the art.
[00413] In some embodiments herein described, the polynucleotide is a DNA molecule that can be in a linear or circular form, and encodes one or more proteins under the control of a promoter recognizable by an enzyme such as an RNA polymerase, that is capable of transcribing the encoded DNA.
[00414] The term“protein” as used herein indicates a polypeptide with a particular secondary and tertiary structure that can interact with another molecule and in particular, with other biomolecules including other proteins, DNA, RNA, lipids, metabolites, hormones, chemokines, and/or small molecules. The term“polypeptide” as used herein indicates an organic linear, circular, or branched polymer composed of two or more amino acid monomers and/or analogs thereof. The term“polypeptide” includes amino acid polymers of any length including full- length proteins and peptides, as well as analogs and fragments thereof. A polypeptide of three or more amino acids is also called a protein oligomer, peptide, or oligopeptide. In particular, the terms“peptide” and“oligopeptide” usually indicate a polypeptide with less than 100 amino acid monomers. In particular, in a protein, the polypeptide provides the primary structure of the protein, wherein the term“primary structure” of a protein refers to the sequence of amino acids in the polypeptide chain covalently linked to form the polypeptide polymer. A protein “sequence” indicates the order of the amino acids that form the primary structure. Covalent bonds between amino acids within the primary structure can include peptide bonds or disulfide bonds, and additional bonds identifiable by a skilled person. Polypeptides in the sense of the present disclosure are usually composed of a linear chain of alpha-amino acid residues covalently linked by peptide bond or a synthetic covalent linkage. The two ends of the linear polypeptide chain encompassing the terminal residues and the adjacent segment are referred to as the carboxyl terminus (C-terminus) and the amino terminus (N-terminus) based on the nature of the free group on each extremity. Unless otherwise indicated, counting of residues in a polypeptide is performed from the N-terminal end (Nth-group), which is the end where the amino group is not involved in a peptide bond to the C-terminal end (-COOH group) which is the end where a COOH group is not involved in a peptide bond. Proteins and polypeptides can be identified by x-ray crystallography, direct sequencing, immunoprecipitation, and a variety of other methods as understood by a person skilled in the art. Proteins can be provided in vitro or in vivo by several methods identifiable by a skilled person. In some instances where the proteins are synthetic proteins in at least a portion of the polymer two or more amino acid monomers and/or analogs thereof are joined through chemically-mediated condensation of an organic acid (- COOH) and an amine (-NH2) to form an amide bond or a“peptide” bond.
[00415] As used herein the term“amino acid”,“amino acid monomer”, or“amino acid residue” refers to organic compounds composed of amine and carboxylic acid functional groups, along with a side-chain specific to each amino acid. In particular, alpha- or a- amino acid refers to organic compounds composed of amine (-NH2) and carboxylic acid (-COOH), and a side-chain specific to each amino acid connected to an alpha carbon. Different amino acids have different side chains and have distinctive characteristics, such as charge, polarity, aromaticity, reduction potential, hydrophobicity, and pKa. Amino acids can be covalently linked to form a polymer through peptide bonds by reactions between the amine group of a first amino acid and the carboxylic acid group of a second amino acid. Amino acid in the sense of the disclosure refers to any of the twenty naturally occurring amino acids, non-natural amino acids, and includes both D an L optical isomers.
[00416] In some embodiments, the sequence of a polynucleotide encoding a genetic molecular component described herein can be homologous to the polynucleotide sequence of the genetic molecular component described herein. For purposes of the present disclosure, two polynucleotide (RNA or DNA) sequences are substantially homologous when at least 80% (preferably at least 85% and most preferably at least 90%) of the nucleotides match over the defined length of the sequence using algorithms such as CLUSTAL or PHILIP. Sequences that are substantially homologous can be identified in a polynucleotide hybridization experiment under stringent conditions as is known in the art. See, for example, Sambrook et al. [139]. Sambrook et al. describe highly stringent conditions as a hybridization temperature 5-10° C below the Tm of a perfectly matched target and probe; thus, sequences that are“substantially homologous” would hybridize under such conditions. Stringency conditions can be adjusted to screen for moderately similar fragments, such as homologous sequences from distantly related organisms, to highly similar fragments, such as genes that duplicate functional enzymes from closely related organisms. Generally, stringent conditions are selected to be about 5°C lower than the thermal melting point (Tm) for the specific sequence at a defined ionic strength and pH. However, stringent conditions encompass temperatures in the range of about 1°C to about 20°C, depending upon the desired degree of stringency as otherwise qualified herein.
[00417] As used herein, "substantially similar" refers to polynucleotides wherein changes in one or more nucleotide bases can result in substitution of one or more amino acids, but do not affect the functional properties of the polypeptide or protein encoded by the nucleotide sequence. "Substantially similar" also refers to modifications of the nucleic acid fragments of the instant disclosure such as deletion or insertion of nucleotides that do not substantially affect the functional properties of the resulting polynucleotide or transcript. It is therefore understood that the disclosure encompasses more than the specific exemplary nucleotide or amino acid sequences and includes functional equivalents thereof. Alterations in a nucleic acid fragment that result in the production of a chemically equivalent amino acid at a given site, but do not affect the functional properties of the encoded polypeptide, are well known in the art.
[00418] Methods of alignment of sequences for comparison are well known in the art. Thus, the determination of percent identity between any two sequences can be accomplished using a mathematical algorithm. Non-limiting examples of such mathematical algorithms are the algorithm of Myers and Miller [141], the local homology algorithm of Smith et al. [142]; the homology alignment algorithm of Needleman and Wunsch [143]; the search-for- similarity- method of Pearson and Lipman [144]; the algorithm of Karlin and Altschul [145], modified as in Karlin and Altschul [146].
[00419] Computer implementations of these mathematical algorithms can be utilized for comparison of sequences to determine sequence identity. Such implementations include, but are not limited to: CLUSTAL in the PC/Gene program (available from Intelligenetics, Mountain View, Calif.); the ALIGN program (Version 2.0) and GAP, BESTFIT, BLAST, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Version 8 (available from Genetics Computer Group (GCG), 575 Science Drive, Madison, Wis., USA). Alignments using these programs can be performed using the default parameters.
[00420] As used herein,“sequence homology”,“homology”, "sequence identity" or "identity" in the context of two nucleic acid or polypeptide sequences makes reference to the nucleotides or residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. When percentage of sequence identity is used in reference to proteins, it is recognized that residue positions which are not identical often differ by conservative amino acid substitutions, where amino acid residues are substituted for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity) and therefore do not change the functional properties of the molecule.
[00421] As used herein,“percentage homology” "percentage of sequence identity" means the value determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide sequence in the comparison window may comprise additions or deletions (gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison, and multiplying the result by 100 to yield the percentage of sequence identity.
[00422] As used herein, "reference sequence" is a defined sequence used as a basis for sequence comparison. A reference sequence may be a subset or the entirety of a specified sequence; for example, as a part of a full-length cDNA or partial genomic DNA sequence, or the complete cDNA or gene sequence. A reference sequence can comprise, for example, a sequence identifiable a database such as GenBank and others known to those skilled in the art.
[00423] The term“substantial homology” or "substantial identity" of polynucleotide sequences means that a polynucleotide comprises a sequence that has at least 80% sequence identity, preferably at least 85%, more preferably at least 90%, most preferably at least 95% sequence identity compared to a reference sequence using one of the alignment programs described using standard parameters. One of skill in the art will recognize that these values can be appropriately adjusted to determine corresponding identity of proteins encoded by two nucleotide sequences by taking into account codon degeneracy, amino acid similarity, reading frame positioning, and the like. Substantial homology or identity of amino acid sequences for these purposes normally means sequence identity of at least 80%, preferably at least 85%, more preferably at least 90%, and most preferably at least 95%.
[00424] The polypeptides and proteins of the disclosure can be altered in various ways including amino acid substitutions, deletions, truncations, and insertions. Novel proteins having properties of interest can be created by combining elements and fragments of proteins of the present disclosure, as well as with other proteins. Methods for such manipulations are generally known in the art. Thus, the polynucleotides described herein comprise both the naturally occurring sequences as well as genetically engineered forms. Likewise, the proteins of the disclosure encompass naturally occurring proteins as well as variations and modified forms thereof.
[00425] It is furthermore to be understood that the polynucleotides of the present disclosure comprise both synthetic molecules and molecules obtained through recombinant DNA techniques known in the art.
[00426] The bacterial cells described herein can be genetically engineered using methods known to those skilled in the art. The polynucleotides, genetic molecular components and molecular components comprised in vectors described herein can be introduced into the cells using transformation techniques such as electroporation, heat shock, and others known to those skilled in the art and described herein. In some embodiments, the U-sensitive genetic molecular components and/or genetic molecular components of the U-sensitive genetic circuits are introduced into the organism to persist as a plasmid or integrate into the genome. In some embodiments, the cells can be engineered to chromosomally integrate a polynucleotide comprising one or more U-sensitive genetic molecular components and/or genetic molecular components comprised in the U-sensitive genetic circuits described herein, using methods such as sacB counterselection procedure [147]. For example, in some embodiments described herein, the U- sensitive genetic molecular components or genetic circuit components are inserted into the Caulobacter chromosome (see Examples). In some exemplary embodiments, a custom designed low copy vector, such as with a pBBRl oriV origin and a cat (chloramphenicol acetyltransferase) gene can be used for reporter expression or to introduce the signal amplifier gene (CCNA_03497) (see Examples).
[00427] In embodiments herein described, a system is provided. The system comprises a plurality of proteobacterial cells and one or more vectors comprising one or more U-sensitive genetic molecular components and/or one or more genetic molecular components of a U- sensitive genetic circuit. The one or more vectors are configured to introduce one or more U- sensitive genetic molecular components and/or one or more genetic molecular components of a U-sensitive genetic circuit into an alphaproteobacterial cell.
[00428] As understood by those skilled in the art, vectors comprising a U-sensitive genetic molecular components or genetic molecular components such as promoters or RNA- or protein coding genes described herein or fragments thereof can be engineered using techniques such as In-Fusion cloning and other methods identifiable by those skilled in the art, to generate vectors suitable for genetically engineering the proteobacterial cells described herein. Polynucleotides encoding genetic molecular components such as promoters and genes encoding RNA and proteins described herein can be isolated from genomic DNA or cDNA comprising the polynucleotides of interest, such as polynucleotides isolated from organisms such as Caulobacter or other Caulobacteridae, using standard Polymerase Chain Reaction (PCR)-based methods known in the art. Plasmids comprising reporter genes and/or U-neutralizing genes described herein are commercially available from vendors such as Thermo-Fisher and Clontech, and other sources such as Addgene, among others known to those skilled in the art. Polynucleotides can also be designed and synthesized de novo , such as using gBlock synthesis (IDT Technologies) as described herein.
[00429] In embodiments herein described, a composition is provided. The composition comprises one or more UC Fi-biosensor or vectors herein described together with a suitable vehicle. [00430] The term "vehicle" as used herein indicates any of various media acting usually as solvents, carriers, binders or diluents for the one or more U-sensitive genetic molecular components, vectors, or cells herein described that are comprised in the composition as an active ingredient. In particular, the composition including the one or more U-sensitive genetic molecular components, vectors, or cells herein described can be used in one of the methods or systems herein described.
[00431] In embodiments herein described, a system comprising an electronic signal transducer adapted to convert a UCLFi-biosensor reportable molecular component output into an electronic output is provided. The system comprises an electronic signal transducer and one or more U biosensors herein described operatively connected to the electronic signal transducer. In some embodiments, the system comprises an electronic signal transducer and one or more U biosensors herein described comprised in a composition together with a suitable vehicle.
[00432] The term“electronic signal transducer” as used herein refers to an electronic device typically comprising a bio-recognition component, a biotransducer component, and an electronic system which can comprise a signal amplifier, processor, data display, and data communicator. Transducers and electronics can be combined, such as in CMOS-based microsensor systems [148]
[00433] As shown in Figure 41, the transducer can be exemplarily fashioned with a blue LED (light emitting diode) (ex. 5 mM, emitting excitation light centered at 470 nm wavelength) 4110 that shines excitation light into a cuvette holder 4105. A curvette holder is a containing device for the filters and curvette, opaque so that the only light in is from the LED and the only light striking the photodetector is the filtered emission light). The blue light passes through an excitation filter (ex filter) 4115 with a center wavelength (CWL) at 469 nm (blue) and a full width at half maximum of 35 nm (narrow bandwidth). This causes a narrow spectrum of excitation light to hit upon the sample in the cuvette 4120. The cuvette is a transparent container for the sample (cells). The emitted light (emission light) from the sample, caused by excitation from the excitation light, passes through an emission filter (em filter) 4125 that has a CWL of 525 nm and a FWHM of 39 nm (a narrow bandwidth of green light). This emission light is read by a silicon amplified photodetector 4130 (with switchable gain for sensitivity control), turning the emission light into an electrical signal for analysis. In some embodiments, the recognition component, often called a bioreceptor, can use biomolecules from organisms or receptors modeled after biological systems to interact with the reportable molecular component output comprising a target analyte of interest. This interaction is measured by the biotransducer which outputs a measurable signal proportional to the presence of the target analyte in the sample. A biotransducer is the recognition-transduction component of the device. In some embodiments, it can comprise a bio-recognition layer and a physicochemical transducer, which acting together converts a biochemical signal to an electronic or optical signal. The bio-recognition layer typically can contain an enzyme or another binding protein such as antibody. For example, polynucleotides, sub-cellular fragments such as organelles (e.g. mitochondria) and receptor carrying fragments (e.g. cell wall), single whole cells, or a plurality of cells optionally on synthetic scaffolds, can also comprise the bio-recognition layer. The physicochemical transducer is typically in contact with the recognition layer. In some embodiments, as a result of the presence and biochemical action of the target analyte of interest, a physico-chemical change is produced within the biorecognition layer that is measured by the physicochemical transducer producing a signal that is proportionate to the concentration of the analyte. The physicochemical transducer can be electrochemical, optical, electronic, gravimetric, pyroelectric or piezoelectric, as understood by those skilled in the art.
[00434] In some embodiments, a quantitative, field-portable UC Fi-biosensor system comprises one or more UC Fi-biosensors described herein coupled with an electronic signal transducer to convert the cellular output reportable molecular component signal (e.g., fluorescence) into an electronic output signal. In some embodiments, one or more UC Fi-biosensors described herein are coupled in conjunction with established, inexpensive, commercially available transducer devices known to those skilled in the art, using one of several immobilization methods such as those utilizing carbon nanotubes or nanoparticles to adhere a UC Fi-biosensor host organism cells to the transducer. In particular embodiments, wherein the UC Fi-biosensor host organism is C. crescentus, the holdfast organelle that facilitates irreversible adhesion to surfaces can be used to couple the cells to the electronic signal transducer, eliminating the need for exogenous immobilization substrates.
[00435] In embodiments herein described, a method of detecting and reporting and/or neutralizing bioavailable UO2F2 is provided. The method comprises:
contacting one or more UO2F2 biosensors, or a system comprising an electronic transducer operatively connected to one or more U02F2biosensors, with a target environment comprising one or more target ranges of bioavailable U concentration and bioavailable F for a time and under conditions to detect and report and/or neutralize bioavailable U02F2in the target environment.
[00436] In some embodiments, a method to detect bioavailable U02F2with an F-sensing and a U-sensing genetic reportable molecular component and/or with a U sensing and/or F-sensing genetic circuit genetic including a fluorescent label such as GFP is with a fluorometer to quantify GFP or other fluorophore’s fluorescence. This can be accomplished with high sensitivity in the laboratory using a microplate reader or in the field using a mini-fluorometer. In some of those embodiments, the method can comprise adding an environmental sample to a 96-well plate or cuvette containing a U-biosensor herein described.
[00437] The term“target environment” as used herein indicates the aggregate of components and related conditions wherein a U biosensor can be operated.
[00438] In some embodiments, the target environment comprises a sample obtained from a field site. In some embodiments, the sample is provided by means of an operator, such as a human or a machine, to the host organism optionally operatively connected to the electronic transducer. In other embodiments, the sample is provided, in absence of an operator, to the host organism optionally operatively connected to the electronic transducer. In some embodiments, the host organism, optionally operatively connected to the electronic transducer, can be in situ in a field site comprising the target environment.
[00439] In an exemplary embodiment, wherein the host is Caulobacter crescentus NA1000 the biosensor is expected to work within a pH range of 6-8, temperature range of ~RT— 37 C. Additional growth nutrients other than inorganic phosphate can be added. In some embodiments, the host can be Caulobacter crescentus OR37 strain isolated from the Oak Ridge Field site as an environmentally robust host: including a greater pH, heavy metal, and U tolerance with respect to Caulobacter crescentus NA1000. [00440] In some embodiments, the reporting of bioavailable UO2F2 can be observed directly, such as by visualizing the output of a reportable molecular component, such as fluorescence of a reportable molecular component, e.g., GFP, wherein the expression and/or function of the reportable molecular component is activated by the U02F2biosensor. In some embodiments, the reporting of bioavailable UO2F2 can be observed indirectly and/or remotely, such as through transduction of reportable molecular component output into an electronic output, which can be quantified by a computer, and which can optionally be communicated to a location at a distance from the target environment by a data communicator, either through wired or wireless communication.
[00441] Therefore, in several embodiments UO2F2 biosensors herein described provide selective and sensitive detection and reporting of UO2F2 which is a bioavailable product of environmental decomposition of UF6, a toxic form of U.
[00442] In several embodiments, the U02F2biosensors, and related U-sensitive F-sensitive genetic molecular components, genetic circuits, compositions, methods and systems described herein provide a cost-effective, selective, sensitive, portable, easy to use, high-throughput measurement of bioavailable U, with little or no sample preparation required.
[00443] Additionally, in some embodiments U02F2biosensors described herein can be used in the construction of consolidated bioremediators comprising bacterial systems that possess all the necessary components for deployment in environmental cleanup efforts. Applications of the U biosensors described herein comprise uses in biodefense (e.g., to be used for non-proliferation purposes), environmental monitoring, and mining (for toxicology and safety concerns), among other uses identifiable by those skilled in the art.
EXAMPLES
[00444] The U biosensors, and related U-sensitive genetic molecular components, genetic circuits, compositions, methods and systems herein disclosed are further illustrated in the following examples, which are provided by way of illustration and are not intended to be limiting.
[00445] In particular, the following examples illustrate exemplary methods and protocols for providing and using U biosensors, and related U-sensitive genetic molecular components, genetic circuits, compositions, methods and systems. A person skilled in the art will appreciate the applicability and the necessary modifications to adapt the features described in detail in the present section, to additional U biosensors, and related U-sensitive genetic molecular components, genetic circuits, compositions, methods and systems according to embodiments of the present disclosure.
[00446] The following methods were used:
[00447] Bacterial strains, media, and materials. All strains were derived from wild type C. crescentus strain NA1000 (ATCC 19089) and listed in Table 2.
Table 2. Strain and Plasmid Table
Figure imgf000139_0001
Table 2. Strain and Plasmid Table
Figure imgf000140_0001
Table 2. Strain and Plasmid Table
Figure imgf000141_0001
[00448] Cell growth and fluorescence characterization experiments were performed in modified M5G medium (10 mM PIPES, pH 7, I mM NaCl, I mM KC1, 0.05 % NH4C1, 0.01 mM Fe/EDTA, 0.2% glucose, 0.5 mM MgSCU, 0.5 mM CaCl2) supplemented with 5 mM glycerol-2- phosphate as the phosphate source (M5G-G2P) to facilitate uranium solubility at the start of growth assays. U stocks were prepared in nitric acid as previously described [153]. A 10,000 ppm ThCU stock was prepared in 5% nitric acid. 50-100 mM stock solutions of Pb(N03)2, NiS04, ZnCl2, CuS04, CaCl2, MnCk, MgS04, K2Cr04, Na2Se03, FeCb, CoCl2, FeS04, NaAs02 were prepared in Milli-Q H20 and a 1 g l 1 AlCb ICP-MS standard in HC1 was used for A1 addition. All strains were grown at 30°C with shaking at 220 RPM in Erlenmeyer flasks or at 1000 RPM in 96-well plates in a PHMP-4 Thermoshaker (Grant Instruments). Strain manipulation was performed in PYE medium, containing 0.2% (wt/vol) Bacto peptone (Difco), 0.1% yeast extract (Difco), 1 mM MgS04, and 0.5 mM CaCl2 and appropriate antibiotics. Escherichia coli HST08 (Clontech) was used for cloning following standard procedures.
[00449] Clean deletions and site-directed mutagenesis. In frame deletions of CCNA_01362 and CCNA_01363 were obtained by a two-step sacB counterselection procedure [154] as described previously[3] using the primers depicted in Table 3.
Figure imgf000142_0001
Figure imgf000143_0001
Figure imgf000144_0001
Figure imgf000145_0001
Figure imgf000146_0001
Figure imgf000147_0001
**Synthesized (Integrated DNA Technologies, gBIock) with the following components listed in 5’ to 3’ orientation: promoterless gfpl0-m2_kl (bolded), rrnBTl and T7Te transcription terminators (BBa_B0015; (lowercase)), P urCB (uppercase)-untraslated region and RBS (lowercase) El-gfpll-M4( bolded), lamda To terminator (lowercase), the rsaA promoter (PrsaA[4]; lowercase) controlling expression of gfpl-9 (bolded). This DNA region is visually depicted in Figure 31.
[00450] Site-directed mutagenesis was performed by amplifying the entire plasmid with the primer sets listed in Table 3. Chromosomal integration and counter selection were performed as described above, and successful substitutions were confirmed by sequencing.
[00451] Transposon screen for regulators of Pphyt. A chromosomally integrated Pphyt -lacZ fusion was constructed using the two-step sacB counterselection procedure [147] to swap the P urcA promoter in pDMP82[3] with Pphyt. To accomplish this, a Pphyt fragment was amplified from the C. crescentus genome with the primer pair Phyt-138_F/ Phyt-138_R (Table 3) and cloned into pDMP82 that was linearized using the primers 138_PurcA_F and 138_PurcA_F using In-Fusion cloning. The Pphyt -lacZ fusion was integrated at the chromosomal urcA locus in lacA mutant strain JOE2321, yielding strain DMP470. For the transposon screen, DMP470 was electroporated with YMCS2::Tn5Pvan[l52 ] and plated onto PYE agar plates containing 25 pg ml 1 kanamycin. -12,000 colonies were scraped into a PYE master solution that was frozen and stored at -80 °C. The Transposon library was diluted and spread on M5G-G2P agar containing 40 pg ml 1 Xgal and 25 pM uranyl nitrate, yielding a total of -36,000 colonies. Colonies exhibiting a white colony phenotype were selected and nested semi-arbitrary PCR was used to map the location of each transposon as described previously. [152]
[00452] Construction of promoter-g/p transcriptional fusions. Plasmid-borne PPhyt -gfp and Pi36i -gfp fusions were generated by amplifying Pphyt and P i % i fragments from the Caulobacter genome with the primer pairs BamHI_Pphyt_F /EcoRI_Pphyt_R and BamHI_P1362/ EcoRI_P1362, respectively, digested with BainH] and CcoRI and cloned into the similarly digested pDMP450, generating pDMP460 and pDMP463. The promoter-g p fusions were shortened and/or mutated by amplifying pDMP460 and pDMP463 with the primer pairs described in Table 3 and re-ligating using infusion cloning.
[00453] Tripartite GFP AND gate sensor construction. A gblock (Tripartite GFP gblock; Table 6) was synthesized (Integrated DNA Technologies, gBlock) with the following components listed in 5’ to 3’ orientation: promoterless gfpl0-m2_kl, rrnBTl and T7Te transcription terminators (BBa_B0015), P UrcB-El-gfpll-M4, lamda To terminator, and gfpl-9 under the control of the rsaA promoter (PrSaA[4]). gfp 10-m2_k 1 was placed under the control of Pphyt by digesting the Tripartite GFP gblock with Bglll and Xhol and ligating into the similarly digested pDMP460 to form pDMP791. DNA sequence encompassing the entire tripartite DNA and an insulating upstream rmbTl transcription terminator was amplified with primers HRP_chrome_int F and HRP_chrome_int R and cloned into pDMP82 that was amplified with the primers urcA_loc_DR_F and urcA-loc_UR_R using infusion cloning to form pDMP792. A variant containing GFP10-m2_kl under the control of a shortened version of Pphyt (he., the core sensor) was constructed by amplifying pDMP792 with Pphyt_TR_elim_F and Pphyt_promoter_shorten_R and re-ligating using infusion cloning, forming pDMP883. A variant of this AND gate containing gfpl-9 under the control of the xylose inducible promoter (Pxyl [6]) was constructed by amplifying the 360 bp P xyi fragment using Pxyl_F and Pxyl_R and cloning into pDMP792 that were amplified with gfpl-9_amp_for_pxyl_F gfpl- 9_amp_for_pxyl_R using infusion cloning to form pDMP664. All tripartite GFP AND gate variants were then integrated into the chromosomal urcA locus using a two-step sacB counterselection procedure, [155] forming DMP804, DMP895, and DMP683
[00454] A control tripartite variant containing both GFP10-m2_kl and E1-GFP11-M4 under the control of P mCB was constructed by directionally cloning a P mCB fragment, generated with primers XbaI_PurcB and BglII_PurcB_R and digested with Xbal and Bglll, into the similarly digested pDMP712. Similarly, a tripartite variant containing both GFP10-m2_kl and E1-GFP11-M4 under the control of UrtAP was constructed by directionally cloning a P nei fragment, generated with primers KpnI_P1362 and AvrII_P1362 and digested with Kpnl and AvrII, into the similarly digested pDMP712. Both control AND gates were cloned into the pNPTS138 double recombination plasmid by amplifying with primers HRP_chrome_int F and R and cloning into pDMP82 that was amplified with the primers urcA_loc_DR_F and urcA-loc_UR_R using Infusion cloning. Finally, the P xyi promoter in both constructs was swapped for P rsaA by cloning P rsaA, amplified with PrsaA_frag_F and PrsaA_frag_R, into the template vectors that were linearized with the primers gfp_for_PrsaA_F and gfp_for_PrsaA_R, forming pDMP932 and pDMP952. These tripartite GFP AND gate variants were then integrated into the chromosomal urcA locus using a two-step sacB counterselection procedure, [155] to form DMP911 and DMP994.
[00455] A U sensing AND gate strain (DMP993) with constitutive UzcY expression was generated by integrating the core sensor (pDMP 883) into the urcA locus of a strain deleted for CCNA_03498 and CCNA_03499 (DMP863). UzcY expression was restricted to conditions of U exposure by placing uzcY expression under the control of Pphyt-short as follows. uzcY was amplified from the C. crescentus chromosome with primers 3497_for_Pphyt_plas_F and 3497_for_Pphyt_plas_R, digested with Bglll and Xhol, and cloned into the similarly digested pDMP746. The DNA containing the Pphyt-short-CCNA_03497 fusion was then amplified with 3497_for_Pphyt_short_chrome_R and Pphyt_short_for_chrom_3497 and cloned into pDMP1113 that was linearized using the primers chrome_3497_vect_amp_F and chrome_vector_amp_R to
The resulting suicide vector (form pDMP1114) was used to integrate Pphyt-short-nzcF into DMP993 to form DMP1009.
[00456] Metal induction experiments. Unless otherwise specified, all metal induction experiments were performed with M5G-G2P. Cells were grown to early-exponential phase, washed once with M5G-G2P, and then resuspended in the same volume of fresh M5G-G2P. Washed cells were added in 195 pi aliquots to black 96- well clear bottom plates containing 5 mΐ of the appropriate metal solution. Cell fluorescence and ODeoo were determined using a Biotek plate reader (Ex: 480 / Em: 516) 2-3 hours post-metal exposure. The mean fluorescence and Oϋόoo value for each data point was subtracted by the respective values for a M5G-G2P blank and then the fluorescence was normalized to the ODeoo. The relationship between the U concentration and the GFP output rate was modeled using a Hill function:
Ymax Ymin
y = ymin +
1 + "
[00457] where y min represents basal normalized fluorescence, ymax is the maximum normalized fluorescence, x represents the U concentration, h is the Hill coefficient, and K is the concentration of U required for half maximal fluorescence expression. The best fit values were found by using the lsqcurvefit function in Matlab.
[00458] Cloning, overexpression, and purification of Strep-UrpR. urpR was amplified with primers 1362_gene_F and 1362_gene_R_GC and cloned into pET 52-b that was amplified with pET52b_F and pET52b_R using infusion cloning to generate plasmid pDMP1118. E. coli BL21(DE3) plys, containing pDMP1118, was grown at 37°C until an OD600 of 0.5 was reached. Isopropyl- 1-thio-b-D-galactopyranoside (IPTG) was added to 0.5 mM, and the cells were shifted to 30°C for four hours of induction. Cells were harvested and stored at -20°C. The cells pellet was thawed and resuspended in 1.4 ml B-PER™ Complete Bacterial Protein Extraction Reagent (ThermoFisher), followed by addition of EDTA to ImM, and rocking for 15 min at room temperature. Insoluble cell debris was pelleted via centrifugation (20,000xg, 10 min at 4°C). Strep-UrpR was isolated from cell lysates using a Strep-tactin column as described in the manufacturer’s protocol (IB A Lifesciences). The protein concentration of UrpR (reported here as monomers) was determined using a Bradford protein assay (Biorad) with lysozyme as a standard.
[00459] Electrophoretic mobility shift assays. A Pphyt promoter fragment containing the region from 13 to 245 with respect to the translation initiation site was amplified from pDMP460 and pDMP475 with primers BamHI_Pphyt_F and Pphyt_EMSA_R, the latter of which was labeled with fluorescein on the 5’ end. Prior to the EMSA, the Strep Tag was removed from UzrpR using HRV 3C protease (Thermo Fisher) according to the manufacturer’s protocol. UrpR was phosphorylated by incubation in phosphorylation buffer (50 mM Tris, pH 7.9, 150 mM NaCl, 10 mM MgCE) with 50 mM disodium carbamyl phosphate (Sigma-Aldrich) for 1 h at 30°C (Lynch and Lin, 1996) and immediately used in the binding assays. EMSAs were performed by incubating phosphorylated UrpR with Pphyt DNA (50 nM) for 10 min at 37°C in buffer containing 50 mM Tris-HCl (pH 7.9), 200 mM NaCl, 10 mM MgCE, 0.1 mg ml 1 BSA, 5% glycerol, 1 mM DTT, and 50 pg ml 'poly- dl-dC. A 5% TBE mini-protean polyacrylamide gel was pre- run with 0.5x TBE at 120 V for 30 min in a Mini-PROTEAN tetra cell (Bio-Rad) prior to loading samples. Samples were run at 100V for 45 min, and the reaction products were visualized using a Biorad Gel Doc XR1 System.
[00460] Site 300 sample collection. Standard operating procedures for sampling and sample handling at LLNL Site 300 have been described in detail [156] and are consistent with the guidance and requirements of the U.S. EPA. The groundwater samples used in this study were collected from each well with either an electrical submersible pump or a bailer. Well samples were placed on ice, filtered using a 0.2 Dm filter, and stored at 4°C.
[00461] Inductively coupled - plasma mass spectrometry / optical emission spectrometry.
Site 300 samples were diluted in 2% (v/v) nitric acid (trace metal grade) and spiked with an internal holmium standard. U was quantified using a Thermo XSeriesII ICP-MS run in standard mode. The sample introduction system was an ESI PFA-ST nebulizer pumped at 120 mΐ/min. Zn, Pb, Cu, Cd, and Cr were quantified using a Thermo iCAP 7400 radial ICP-OES in standard operating mode. Standard curves were generated using a 100 mM uranyl nitrate stock solution and 10 ppm Zn, Pb, Cu, Cd, and Cr ICP-MS standards (Inorganic Ventures).
[00462] Pphyt -lacZ and Pi36i -lacZ reporter constructs. Chromosomally integrated PPhyt-lacZ and Pi36i-lacZ translational fusions were constructed using a two-step sacB counterselection procedure [147] to swap the PUrcA promoter in pDMP82[3] with either Pphyt or Pi36i. pDMP82 contains the necessary sequence to generate a translational VWcK-lacZ fusion at the PurcA locus in which the sequence 24 nt downstream of the urcA start codon is fused to E. coli lacZ. To accomplish this, the Pphyt and Pi36i fragments were amplified from the Caulobacter genome with the primer pairs Phyt-138_F/ Phyt-138_R and 1362_138_F/1362_138_R (Table 4), respectively and cloned into pDMP82 that was linearized using the primers 138_PurcA_F and 138_PurcA_F (Table 4) using In-Fusion cloning.
Table 4. DNA primers:
Figure imgf000152_0001
Table 4. DNA primers:
Figure imgf000153_0001
Figure imgf000154_0001
[00463] Both fragments were cloned into the Hindlll and BamHl- digested pNPTS 138 (Table 5) using In-Fusion cloning and integrated into the chromosome of a Caulobacter crescentus lacA mutant strain JOE2321 (Table 5) or NA1000 strain (Table 5) using the double-crossover allele replacement as described above, yielding strains DMP89 and DMP90, respectively (Table 5), integrated at the PUrcA locus.
Table 5. Caulobacter crescentus strains and plasmids:
Figure imgf000155_0001
Table 5. Caulobacter crescentus strains and plasmids:
Figure imgf000156_0001
Table 5. Caulobacter crescentus strains and plasmids:
Figure imgf000157_0001
Table 5. Caulobacter crescentus strains and plasmids:
Figure imgf000158_0001
Table 5. Caulobacter crescentus strains and plasmids:
Figure imgf000159_0001
[00464] Pphyt -gfp and Pi36i -gfp constructs. Plasmid-borne Pphyt -gfp and Pnei-gfp fusions were generated by amplifying Pphyt and Pi36i fragments from the Caulobacter genome with the primer pairs Phyt-138_F/ Phyt-138_R and 1362_138_F/1362_138_R (Table 4), respectively, digested with BamHI and Z oRI and cloned into the similarly digested pDMP450 (Table 5), generating pDMP460 and pDMP463 (Table 5).
[00465] Construction of a CCNA_01968 promoter-g/p transcriptional fusion. A synthetic vector with a P15A origin, kanamycin resistance cassette and gfpmut3 gene insulated with an upstream rrnB terminator 1 (DNA2.0) was used as a template to construct promoter-gfp fusions. First, the cat gene (chloramphenicol acetyltransferase) from pNJH123 was amplified with the primers cat_F and cat_R and cloned into pDMP462 that was linearized with the primers kan_elim_F and kan_elim_R using infusion cloning. The pl5A origin was then swapped with a pBRR-repl origin that was amplified from pPROBE-GFP’ using the primers pBBRl-rep_F and pBBRl-rep_R and the restriction enzymes Nhel and Hindlll, generating pDMP460. The DNA sequence region from 170 to -9 with respect to the translation initiation site of CCNA_01968 was amplified with P1968_BamHI_F and P1968_EcoRI_R, digested with EcoRI and BamHI, and cloned into the similarly digested pDMP460 to construct a CCNA_01968-promoter gfp fusion (pDMP558). Site directed mutagenesis was performed with the primers P1968_HSl_mut_F and P1968_HSl_mut_R to mutate UzcR half site one from 5’-CATTAC-3' to 5'-CAATAG -3' and primers P1968_HS2_mut_F and P1968_HS2_mut_R and to mutate the half site two from 5’-TTAA-3' to 5'-TAAT-3', generating pDMP559 and pDMP560, respectively.
[00466] Engineering of constructs to place the uzcRS operon under the transcriptional control of Pphyt/Pi36i. The two mapped promoters of uzcR were replaced with P i % i and Pphyt as follows. First, a synthetic DNA containing the rmb T1 and T7Te transcription terminators (BBA_0B0015) followed by a Pi36i fusion with the first 168 nucleotides of uzcR was prepared (Integrated DNA Technologies, Inc.) Then, pDMP499 (Table 5), a pNPTS138-based vector containing the DNA sequence for substituting the aspartate residue at position 51 for alanine was amplified with primers pNTPS138_urcR_F and pNTPS138_urcR_R (Table 4) and the 531 bp region upstream of the uzcR promoters was amplified with urcR_UR_F and urcR_UR_R (Table 4). All three DNAs were ligated together to make pDMP610 (Table 5). Pphyt was swapped for Pi36i using infusion cloning with the fragments generating by amplifying pDMP610 (Table 5) with Pphyt-uzcR_vect_F/Pphyt-uzcR_vect_R and pDMP614 with Pphyt_urcR-loc_F/ Pphyt_urcR-loc_R (Tables 4 and 5).
[00467] The resulting suicide vectors pDMP609 and pDMP614 (Table 5) were electroporated into FC922 (Table 5) and the PI 361 -uzcRS (DMP690) and Pphyt-uzcRS strains (DMP601) (Table 5) were obtained by a two-step sacB counterselection procedure [147]. The UzcR binding site from the rsaFb promoter TGCGTGAAAAAAGCTTAACT (SEQ ID NO: 201) was inserted downstream of the transcriptional start site of Pphyt and Pi36i as follows. Plasmids pDMP609 and pDMP614 and were amplified with the primer pairs P1362_rsaFb_F/ P1362_rsaFb_F and Pphyt_rsaFb_F/ Pphyt_rsaFb_F, respectively (Table 5) and re-ligated using InFusion cloning. The resulting suicide vectors pDMP673 and pDMP621 (Table 5) were electroporated into Caulobacter strain FC922 and the P1361m_5-uzcRS (DMP679) and Pphytm_5 -uzcRS (DMP643) strains (Table 5) were obtained by the two-step sacB counterselection procedure. All strains were transformed with pDMP558, encoding a CCNA_01968 promoter gfpmut3 fusion (Table 5). [00468] Engineering of constructs to place expression of hrpS under control of Pphyt or
Pi36i, and hrpR under control of PurcB. PhrpL DNA (SEQ ID 80) was synthesized (IDT), then digested with BamHI and BglU and ligated into the similarly digested pDMP450, generating pDMP610. Next, the synthetic Pphyt-hrpS_PurcB-hrpR DNA fragment was digested with Xbal and BamHI and cloned into the similarly digested pDMP610 to generate pDMP612. pDMP612 was cloned into NA1000 to produce DMP681.
Table 6. DNA sequences of genetic molecular components:
Figure imgf000161_0001
Table 6. DNA sequences of genetic molecular components:
Figure imgf000162_0001
Table 6. DNA sequences of genetic molecular components:
Figure imgf000163_0001
Table 6. DNA sequences of genetic molecular components:
Figure imgf000164_0001
Example 1: Identification of U-selective promoters Pi36i and Pphvtin Caulobacter crescentus
[00469] The highest U-induced gene urcA (uranium response in Caulobacter) has been exploited as a U sensor that can detect micromolar U concentrations in contaminated ground water [55]. However, previously the molecular mechanisms governing PurcA regulation were not examined, nor its cross -reactivity with environmentally relevant metal cations. To address this, the specificity was characterized and the transcriptional regulatory mechanism governing expression of the PurcA was identified [3] . Although most metals failed to induce PurcA, significant induction was observed with the metal ions Zn and Cu and to a lesser degree, Cd (Figure 2 Panel A).
[00470] The UzcRS two-component system was identified as the regulatory system responsible for U, Zn, and Cu-dependent activation of PurcA and 41 other promoters in the Caulobacter genome [3]. Together, these data suggest that the PurcA does not have satisfactory selectivity to function as a standalone sensor of environmental U. Nevertheless, since UzcRS exhibits a U- concentration dependence in a wide range of media conditions, a sensor that incorporates UzcRS as one component within a more advanced U- sensitive genetic circuit comprising an additional point of U sensing that is independent of the UzcRS system could produce an effective U sensor.
[00471] To identify additional U responsive genes that are not cross-reactive with other metal cations, gene expression was monitored in the presence of U and Zn using RNA-seq (Figure 1). Two additional promoters (Pi36i (promoter of operon containing CCNA_01362) and Pphyt (promoter of CCNA_01353)) that are strongly responsive to U but not to Zn. To further characterize the specificity of these promoters, the DNA sequences corresponding to each promoter were cloned upstream of lacZ and gfpmut3 reporter genes and reporter expression was monitored following exposure to 10 different metals. Surprisingly, the data revealed that both promoters lack cross -reactivity with metal ions commonly encountered in the environment (Figure 2 Panels B-C). Furthermore, U-dependent induction of these promoters was not dependent on UzcRS; Pi36i-gfp expression was induced 10.5 (±0.8) for wild type and 9.6 (±0.5) for a strain deleted for uzc during exposure to 20 mM U, suggesting that that regulation of these promoters is governed by a regulatory mechanism distinct from PurcA. In other words, U-sensing by P phyt/P 1361 and PurcA is mediated through independent mechanisms. As such, these sensors are suitable for U sensor construction.
Example 2: Engineering of a U-sensitive genetic circuit comprising an‘in series’ AND gate wherein the uzcRS operon is placed under the transcriptional control of PDhvt/Pi36i [00472] In the C. crescentus NA1000 genome, uzcR and uzcS are physically separated by genes encoding the ParDE3 toxin anti-toxin (TA) system, together forming a putative four-gene operon [112]. Although uzcR and uzcS are conserved throughput much of alphaproteobacteria, the insertion of parDE3 between uzcR and uzcS is unique to a subset of the Caulobacter genus; uzcR and uzcS are adjacently located in the majority of closely related alphaproteobacteria [3] including C. crescentus strain OR37, an environmental isolate from a U-contaminated site [97]. The parDE3 system does not contribute to the metal-dependent regulation by UzcRS [3]. Given this result and the potential toxicity associated with parDE3 overexpression, the parDE3 TA system was deleted, so that uzcR and uzcS are adjacently located. The expression of uzcR is controlled by two promoters (Pi and P2) in C. crescentus [3], which enables sufficient basal expression of uzcRS to activate transcription in response to metal (U, Zn, Cu). There is also a putative UzcR binding site located upstream of Pi that likely yields a positive feedback loop. Indeed, UzcR protein levels increase in a wzcS-dcpcndcnt manner in response to metal sensing. Deletion of the parDE3 TA system and the parD promoter places uzcS expression under the exclusive control of Pi and P2 (Figure 4 Panel A). A“control strain” of Caulobacter was generated comprising a U-sensitive genetic circuit in which uzcR is under the control of Pi and P2 and uzcS under the control of Pi and P2, (Figure 4 Panel A). As expected, this strain produces a high fluorescence signal in response to U, Zn, Cu (Figure 5 Panel A left, middle, right graphs, respectively). Incremental improvements were made to this circuit to enhance specificity, as described below.
[00473] To enhance the selectivity of UzcRS for U, PI and P2 were replaced with Pphyt or P1361 such that uzcRS expression is dependent on activation by these U- specific promoters (Figure 4 Panel B). This genetic circuit requires two points of U sensing for reporter activation, (1) activation of uzcRS transcription by Pphyt or P1361 and (2) stimulation of UzcRS transcriptional activity. Caulobacter comprising this sensor showed greater reporter expression signal in response to U (Figure 5 Panel B, left graph) compared to the control (Figure 5 Panel A, left graph). Importantly, Cu-sensing has been completely abolished (Figure 5 Panel B, right graph) while Zn induction with the range of inducing Zn concentrations narrowed (Figure 5 Panel B, middle graph) compared to the PurcA sensor alone (Figure 5 Panel A, middle graph). The ratio of the U signal output to that of Zn has been increased from 1.6 to 3.5 (Figure 5 Panel B, left and right graphs).
[00474] To further improve specificity of U sensing, a negative feedback loop was incorporated into the circuit, whereby UzcR represses its own expression from Pphyt or Pi36i, in order to minimize the basal expression of the uzcRS operon. Specifically, a UzcR binding site was placed downstream of the Pphyt or Pi36i transcription start site (Figure 4 Panel C). Although the UzcR binding site from the rsaFb promoter was used, any m_5 site is suitable. Caulobacter comprising this sensor showed strong responsiveness to U (Figure 5 Panel C, left graph) and further shifted ratio of U response to that of Zn to 5.5 (Figure 5 Panel C, left and right graphs).
Example 3: Engineering of a U-sensitive genetic circuit comprising an‘in parallel’ HRP AND gate
[00475] This Example describes the engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ AND gate that utilizes the HRP AND gate system from Pseudomonas syringae that was recently developed in E. coli [126]. In this system, both HrpS and HrpR are required for s54 dependent activation of the hrpL promoter (Ri,Gr[ J. Expression of HrpS or HrpR alone is not sufficient for transcriptional activation.
[00476] The genetic circuit described in this Example contains genetic components whose expression is controlled independently by (1) Pphyt or Pi36i and (2) UzcRS systems, and reporter expression requires both HrpS and HrpR to be expressed.
[00477] To generate this U-sensing AND gate, the expression of hrpS was placed under the control of Pphyt or Pi36i , while hrpR was placed under the control of PUrcB, a UzcRS -dependent promoter that was recently identified that has lower basal activity compared to PurcA [3] (Figure 6 Panel A). A PhrpL -gfp fusion was generated as a reporter and requires Pphyt /Pi36i and PurcB to be active to generate a fluorescent signal.
Example 4: Engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ tripartite GFP AND gate
[00478] This Example describes the engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ AND gate that utilizes the tripartite GFP system [5], which requires expression of gfplO, gfpl 1 and gfpl-9 for reporter expression (Figure 6 Panel B).
[00479] To construct an in parallel AND gate comprising the tripartite GFP system a gblock (Tripartite GFP gblock) was synthesized (Integrated DNA Technologies, gBlock) with the following components listed in 5’ to 3’ orientation: GFP10-m2_kl, rmBTl and T7Te transcription terminators, E1-GFP11-M4 under the control of the UzcRS promoter (PMrcs), lamdaTo terminator, gfpl-9 under the control of the rsaA promoter (PrsaA [4]). GFP10-m2_kl was placed under the control of P phyt or P i % i by digesting Tripartite GFP gblock with Bglll and Xhol and ligating into the similarly digested pDMP460 and pDMP463, respectively, to form pDMP791 and pDMP881. DNA sequence encompassing the entire R,,/,g,/R I % i tripartite DNA and an insulating upstream rrnbTl transcription terminator was amplified with primers HRP_chrome_int F and R and cloned into pDMP82 that was amplified with the primers urcA_loc_DR_F and urcA-loc_UR_R using infusion cloning to form pDMP792 and pDMP692. A variant of this AND gate containing gfpl-9 under the control of the xylose inducible promoter (Pxyl [6]) was constructed by amplifying the 360 bp P xyi fragment using Pxyl_F and Pxyl_R and cloning into pDMP792 (R,,/,g, version) and pDMP692 ( P i % i version) that were amplified with gfpl-9_amp_for_pxyl_F gfpl-9_amp_for_pxyl_R using infusion cloning to form pDMP664 and pDMP 665, respectively. A variant of this AND gate containing gfpl-9 under the control of the shortened Pphyt promoter was constructed by amplifying pDMP736 with Pphyt_TR_elim_F and Pphyt_SD_amp_R and cloning this 137 bp fragment into pDMP712 ( R,,I,M version) or pDMP713 (Pi36i version) that was amplified with gfpl-9_for_pphyt-short_F and gfpl-9_for_phyt_R using infusion cloning to form pDMP808 and pDMP809. A variant containing GFP10-m2_kl under the control of a shortened version of Pphyt and gfpl-9 under the control of PrSaA was constructed by amplifying pDMP792 with Pphyt_TR_elim_F and Pphyt_promoter_shorten_R and re-ligating using infusion cloning, forming pDMP883. Lastly, a control tripartite variant containing both GFP10-m2_kl and E1-GFP11-M4 under the control of P mCB was constructed in three parts. First, a P mCB fragment generated with primers XbaI_PurcB and BglII_PurcB_R was digested with Xbal and Bglll and directionally cloned into the similarly digested pDMP712, forming pDMP741. Next, DNA sequence encompassing the entire P mCB tripartite DNA and an insulating upstream rrnbTl transcription terminator was amplified with primers HRP_chrome_int F and R and cloned into pDMP82 that was amplified with the primers urcA_loc_DR_F and urcA- loc_UR_R using infusion cloning to form pDMP748. Finally, the PAV/ promoter was swapped for P rsaA by cloning PrsaA, amplified with PrsaA_frag_F and PrsaA_frag_R, into pDMP748 that was amplified with gfp_for_PrsaA_F and gfp_for_PrsaA_R, forming pDMP932. These tripartite GFP AND gate variants were then integrated into the chromosomal urcA locus using a two-step sacB counterselection procedure [155].
[00480] Figure 11 shows graphs reporting exemplary data corresponding to the exemplary U- sensitive tripartite GFP genetic circuit.
[00481] Tripartite GFP system with PrsaA-gfpl-9: Can detect U in the 2-20 uM range. When a growth media containing Glycerol-2-phosphate as the P source is used, the signal amplitude is higher but the responsive range is shifted to 8 uM-30 uM. Higher concentrations have a diminished signal output. The shifted range likely reflects U coordination by glycerol-2- phosphate that reduces the bioavailability.
Example 5: Engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ Bacterial two-hybrid system AND gate.
[00482] This Example describes the engineering of a U-sensitive genetic circuit comprising an ‘in parallel’ AND gate that utilizes the bacterial two-hybrid system [117]. In such an embodiment, the alpha-gall IP fusions and lamda repressor-gal4 fusion are expected to be driven by a combination of Pphy/ nei and any UzcRS regulated promoter (see Figure 22).
Example 6. Identification of a putative U-sensitive transcriptional regulatory DNA binding site in Pphvt and Pi36i promoter sequences.
[00483] Using a bioinformatic approach, a putative regulator binding site was identified in proximity to the transcription start site in Pphyt and Pi36i that is comprised of two direct repeat elements (e.g., CGTCAGC (SEQ ID NO: 202)); Figure 7 Panels A and B). This binding site is conserved amongst Caulobacteridae Bradyrhizobiaceae, Sphingomonadaceae,
Hyphomicrobiaceae, and Rhodobacteracea, facilitating a phylogenetic footprinting approach to construct a putative regulator DNA-binding motif (Figure 7 Panel C), using 26 DNA sequences of Pphyt and Pi36i from Caulobacter sp. Root342, Phenylobacterium sp. Root700, Caulobacter crescentus NA1000, Caulobacter sp. Rootl455, Caulobacter sp. Root487D2Y, Paracoccus sp. 228, Caulobacteraceae bacterium OTSz_A_272, Novosphingobium sp. AP12 PMI02, Hyphomicrobium sp. MCI, Hyphomicrobium denitrificans , Brevundimonas sp. Rootl279, Sphingopyxis sp. Rootl497, Afipia sp. P52-10, Caulobacter sp. Root342, Hyphomicrobium denitrificans, Sphingobium sp. YBL2, Sphingobium baderi LL03, Sphingobium indicum B90A, and Roseovarius indicus strain DSM 26383.
[00484] The functional role of the putative regulator site within each promoter was tested by mutating conserved nucleotides within each direct repeat away from consensus, as shown in Figure 8 Panels B, C, D, G and H. Mutations within either direct repeat abrogated U- dependent activation of both the Pphyt and Pi36i promoters in the host organism C. crescentus NA1000 suggesting that the U-responsive regulator is natively a transcriptional activator (Figure 9 Panels A and B).
[00485] Pphyt (but not Pi36i) also contains a large tandem repeat (TR;32 bp) located further upstream from the putative regulator site (Figure 8 Panel A). To test the function of this TR in U-dependent induction of Pphyt the size of the Pphyt DNA was reduced from 238 bp to 81 bp, eliminating the TR and all upstream DNA (Figure 8 Panel C). The data shown in Figure 9 Panel A indicate that the TR is not required for U induction. Similarly, a shortened version of Pi36i that included only nine bp upstream of the regulator binding site (Figure 8 Panel I) retained U-dependent regulation (Figure 9 Panel B). Collectively, these data indicate that the DNA sequence extending beyond the regulator binding site within Pi36i and Pphyt are not required for U-dependent activation.
[00486] To create a synthetic uranium repressed promoter, the consensus direct repeat UrpR binding site can be integrated within a promoter region, such that UrpR DNA binding will interfere with RNA polymerase binding and/or transcription. The UrpR binding site should be integrated at a location that overlaps, but does not alter the sequence of the -35 and/or -10 promoter elements or the TSS; disturbing -35 and/or -10 promoter elements will yield a promoter with low basal activity. Ideally, multiple locations will be tested to optimize results. While this promoter can be used to control transcription of any biological reporter, destabilized gfp (e.g., GFP-LVA[159]) is expected to yield the best results given the enhanced degradation rate. Highly stable reporters (e.g., GFP) will require several rounds of cell division to observe a uranium (e.g. UrpR) dependent decrease in reporter activity.
Example 7. Signal amplifier module
[00487] A positive regulator protein UzcY, encoded by CCNA_03497, was identified, which functions as a “natural” signal amplifier for the UzcRS system. Under normal growth conditions, uzcY is repressed by the MarR family transcription factor (CCNA_03498), and thus has no effect on UzcRS activity. However, when UzcY expression is induced through relief of CCNA_03498 repression or by ectopic expression, it stimulates UzcS activity through a direct interaction, causing a hypersensitive output in response to the metal inducers U, Zn, and Cu. This has the effect of dramatically increasing the output signal amplitude in response to low U (or Zn/Cu) concentrations, thus increasing sensitivity, and lowering the U detection limit of UzcRS by over 4-fold (Figure 12).
[00488] To enhance the sensitivity of U sensors, this“natural” signal amplifier module can be integrated into U sensor circuitry. To accomplish this, in an exemplary circuit the U-specific promoter P phyt or P/.¾/ is used to drive expression of uzcY (e.g., see Figure 13) such that UzcY levels are modulated in a U-concentration dependent manner. The low U detection limit of P phyt and Pi36i (-500 nM) is expected to allow signal amplification at environmentally relevant U concentrations, which is expected to improve U sensitivity and lower the detection limit in view of exemplary data demonstrating that both of these properties can be achieved with the native UzcRS system (e.g., see Figure 12). Additionally, by restricting UzcY-mediated signal amplification to conditions of U exposure (e.g. by placing UzcY under regulatory control of a U- selective promoter such as Pi36i or Pphyt, the selectivity for U is expected to be further enhanced.
Example 8. Combination of‘in parallel’ and‘in series’ AND gate circuits
[00489] An example of combining‘in series’ and‘in parallel’ AND gate circuits within the same cell to enhance selectivity is shown in Figure 14. An advantage of this exemplary circuit is that the UzcRS input is U selective whereas in the original‘in parallel’ circuit as shown in Figure 6 Panel B, the UzcRS -regulated promoter PurcB is cross -reactive with Zn and Cu.
Example 9. Negative regulators of UzcRS [00490] The following four different negative regulators of UzcRS that are encoded in the Caulobacter chromosome were identified (see Figure 27). None of these regulators are required for U sensing by UzcRS, however, their levels modulate the sensitivity to U.
[00491] Negative regulator 1: CCNA_03681 and CCNA_03680 encode an ABC transporter ATPase and an ABC-2 family transporter fused to a C-terminal aminopeptidase N domain, respectively (urtAP). Together these proteins form an ABC transporter with a C-terminal aminopeptidase domain.
[00492] Negative regulator 2: CCNA_02866 (also referred to herein as uzcX), encodes a membrane protein of unknown function that is located within a prophage region of the genome and part of the UzcR direct regulon (~8-fold activated by UzcR (Park el al, 2017) [3].
[00493] Negative regulator 3: A MarR family regulator CCNA_03498 that represses expression of an operon containing CCNA_03497, CCNA_03498 and CCNA_03499 (see Figure 26). Expression of CCNA_03497 (occurs when repression mediated by CCNA_03498 is lifted) hypersensitizes UzcS to metal inducers.
[00494] Negative regulator 4: A second, paralogous MarR family regulator CCNA_02289 that represses expression of an operon containing CCNA_02291, CCNA_02290 and CCNA_02289 (see Figure 26). Expression of CCNA_02291 (occurs when repression mediated by CCNA_02289 is lifted) hypersensitizes UzcS to metal inducers.
Example 10. U biosensor having U-neutralization output
[00495] This Example describes a U biosensor having exemplary U-neutralization outputs in response to bioavailable U.
[00496] Figure 21 shows a schematic showing exemplary U-sensitive genetic circuits having exemplary U-neutralizing outputs to allow U bioprecipitation or bioadsorption. In the exemplary U-sensitive genetic circuits, an exemplary AND gate comprises a Pphyt promoter configured to initiate expression of UzcR and UzcS in presence of bioavailable U, and a PurcB promoter (activated by UzcR) configured to initiate expression of exemplary U-neutralizing genes phoY or phytase (to provide a U bioprecipitation output), or fusion genes of rsaA-SUP, rsaA-CaM, ompA- SUP, or ompA-CaM (to provide a bioadsorption output).
Example 11. UzcR regulated promoter
[00497] Table 7 shows a list of UzcR regulated promoters. Bolded regions are putative UzcR m_5 sites and single boded nucleotide is TSS.
Figure imgf000173_0001
Figure imgf000174_0001
Figure imgf000175_0001
Figure imgf000176_0001
Figure imgf000177_0001
Figure imgf000178_0001
Example 12: Application of a combinatorial input logic towards U detection
[00498] A combinatorial sensor approach using multiple regulators with broad specificity profiles can be adopted for the selective detection of compounds lacking specific regulators (e.g., butanol).
[00499] To apply the combinatorial input logic towards U detection, an AND gate circuit was developed by integrating two functionally independent, native U-responsive regulatory pathways into a single synthetic pathway in Caulobacter crescentus. Prior studies revealed that C. crescentus tolerates high U concentrations [160, 161] and exhibits a robust and specific gene expression response following U exposure. [160, 162, 163] Using the promoter of the highest U- induced gene, urcA,{Pim,\) to drive expression of UV-excitable gfp, Hillson el al generated a whole-cell U sensor that was responsive to sub-micromolar U concentrations and successfully detected U in a groundwater sample. [163] Subsequent genetic and biochemical analyses revealed that the majority of U-dependent gene induction in this bacterium, including regulation of PurcA, is mediated by the TCS UzcRS, [153] supporting a role for UzcRS as a U-responsive master regulator. However, characterization of the specificity of UzcRS revealed strong cross reactivity with the common environmental metals Zn and Cu.[3]
[00500] To improve upon the selectivity of the UzcRS system, the selectivity of a second U- responsive regulatory system (UrpRS) in C. crescentus was identified and characterized. Leveraging the distinct selectivity profiles of the UzcRS and UrpRS TCS and the tripartite GFP genetic framework, [164] an AND gate circuit that integrates signaling input from both pathways was constructed and characterized.
Example 13: Identification of the U-responsive UrpRS TCS
[00501] Examination of U and Zn transcriptomic data [3, 160] in C. crescentus revealed a small subset of highly U-induced genes that are not regulated by UzcRS and only weakly induced by Zn (Figure 1). Two operons (CCNA_01361-CCNA_01362-CCNA_01363 and CCNAJ31353- CCNA_01352-CCNA_01351) were of particular interest based on several observations. First, both operons are highly induced by U in minimal and complex media, but lack induction by other known stress responses (e.g., DNA damage, heat shock, heavy metals; Table 8A to Table 8C).
Figure imgf000179_0001
Figure imgf000180_0001
Figure imgf000180_0002
Figure imgf000180_0003
[00502] The operons are furthermore likely regulated by the same transcription factor based on the presence of nearly identical DNA sites comprised of two tandem direct repeats (5’-GTCAG- 3’; Figure 7A-B) with 11-bp center-to-center (etc) spacing within the CCNA_01353 (Pphyi) and CCNA_01361 promoters (Pi36i). This direct repeat site is conserved within the promoters of closely related alpha proteobacteria (Figure 7C) and required for U-dependent induction of Ppi,y, and Pi36i; mutations away from consensus in both Pphyr and Pi36i -gfp fusions abrogate U- induction (Figure 9A-B). The CCNA_01353-CCNA_01352-CCNA_01351operon encodes a phytase enzyme that confers U tolerance [162] and an uncharacterized response regulator and histidine kinase pair, while the CCNA_01361-CCNA_01362-CCNA_01363 operon encodes a PepSY superfamily protein and another uncharacterized response regulator and histidine kinase pair.
[00503] This direct repeat binding site bears striking resemblance to the binding sites of OmpR/PhoB family response regulators. [169, 170] As such, the OmpR/PhoB -family response regulators CCNA_01351 and CCNA_01362, whose expression is governed by Ppi,yL and Pi36i, respectively, are potential candidates for the U-responsive regulator. Deletion of CCNA_01351 or the histidine kinase CCNA_01352 had no effect on U-dependent induction of either promoter (data not shown), consistent with the results of transcriptomics performed with both mutants during U exposure. [160] In contrast, deletion of CCNA_01362 or the histidine kinase CCNA_01363 abolished U-dependent induction of both promoters (Figure 17A-B).
[00504] Furthermore, a forward genetic screen for mutants that failed to induce Pphyt -lacZ in response to U resulted in five transposons that mapped to unique locations within CCNA_01362 and CCNA_01363 (Figure 28; Table 9), providing independent validation for a functional role of both proteins in U-dependent stimulation of Pphyt.
Figure imgf000181_0001
[00505] These proteins have been putatively named UrpR and UrpS (Uranium Responsive Phytase Regulator and Sensor, respectively), since this TCS strongly activates a gene encoding a phytase enzyme.
[00506] To confirm that regulation by UrpR is direct, UrpR was purified, and its binding to Pphyt was tested using an electrophoretic mobility shift assay. As expected, UrpR bound to Pphyt in a concentration-dependent manner (Figure 10). Binding was not observed to a Pphyt fragment containing a 5’-GTCA-3’ to 5’-CAGT-3’ mutation at DR2 (Figure 10), supporting the functional role of the direct repeat site in UrpR DNA binding. Collectively, these data suggest that the TCS comprised of UrpR and UrpS is the U-dependent activator of Pphyt and Pi36i, revealing a positive feedback loop within this regulatory system.
Example 14: UrpRS exhibits improved metal selectivity compared to UzcRS
[00507] To test whether UrpRS functions independently of UzcRS with respect to U perception (a property that is important for its dual integration with UzcR in a synthetic U sensing pathway), the expression of Pphyt and the UzcR-regulated promoter P mCB were tested in strains lacking the non-cognate TCS. The U-induction profile of a shortened Pphyt reporter (PPhyt-short; diagrammed in Figure 7A-B) was largely unaffected by deletion of uzcR or uzcS (Figure 29A; Figure 30), while the U-induction profile for P urcB-gfp was unaffected by deletion of urpR or urpS (Figure 29A; Figure 30). Collectively, these data indicate that C. crescentus possesses at least two independent U-responsive TCS.
[00508] To determine the metal selectivity of UrpRS, PPhyt-short-g/p fluorescence was quantified following exposure to 16 metals of environmental relevance (Figure 29B). To encourage metal solubility and thus bioavailability, minimal media containing glycerol-2-phosphate (G2P) was employed as the sole source of phosphate. As a control, the metal selectivity of the P urcB-gjp reporter was also tested (Figure 29C). For both promoters, the strongest induction was observed in response to U; however, PPhyt-short exhibited greater U sensitivity, resembling a switch-like activation response (Figure 29A) that can likely be attributed to the strong positive auto regulation within the UrpRS system. Most metals, including Th, a radionuclide that is part of the U decay chain and likely co-occurs with U, failed to induce either reporter (Table 10).
Table 10. Metal selectivity of UrpRS and UzcRS
Figure imgf000182_0001
Figure imgf000183_0001
Figure imgf000184_0001
[00509] Importantly, Pphyt-short was unresponsive to Cu and only minimally responsive to Zn or Pb (Figure 29B). In contrast, strong P mCB induction was observed with Zn, Pb, and within a narrow Cu range. The Pb-dependent induction of P mCB was surprising and contrasts with prior reports. [153, 163] It is suspected that the Pb induction likely reflects the higher initial Pb bioavailability in the presence of G2P compared to orthophosphate, given the low solubility of lead phosphate. [171] These data suggest that while UrpRS is not exclusively selective for U, it exhibits an improved metal selectivity profile compared to UzcRS.
Example 15: Construction of a U-responsive AND gate pathway in C. crescentus
[00510] Given the distinct metal selectivity profiles and functional independence of UzcRS and UrpRS, it is expected that a combinatorial approach that incorporates the U-responsive functionality of both UrpRS and UzcRS will yield a whole-cell U sensor with enhanced specificity. Accordingly, the recently developed Tripartite GFP system[164] was used as a template to construct a U sensing AND gate. In this system, GFP is split into three parts (gfplO, gfpl l and gfpl-9) that interact to reconstitute active GFP when co-expressed; the synthetic K1 and El coiled-coils[172] were fused to the C-terminus of gfplO and the N-terminus of gfpl l, respectively, to mediate dimerization. [164] An advantage of the tripartite system is the low basal level of GFP fluorescence, ensuring a robust OFF state.
[00511] Pphyt-short and P UrcB, which both exhibit low basal activity and a large U-dependent fold change, were used to drive expression of gfplO-Kl and El-gfpll, respectively. The expression of gfpl-9 was driven by the strong, constitutive rsaA promoter (P,«,,t)[4J in order to produce high GFP1-9 levels at all stages of growth (Figure 14A; diagrammed in Figure 31). Hereafter, this sensor configuration in an otherwise wild type (WT) strain background is referred to as the core sensor. Sensor variants where gfpl-9 expression was driven by the xylose-inducible promoter (P*p)[6] or the UrpR-regulated Pphyt were also tested, but exhibited a lower signal amplitude and were not pursued further (Figure 32). Additionally, attempts to employ the hrp (hypersensitive response and pathogenicity) amplifier from Pseudomonas syringae as the AND gate template for U sensor construction by placing hrpR and hrpS under the control of Pphyt-short and P UrcB, respectively, and using P/„7,/. to drive gfp expression, did not result in a functional sensor.
[00512] Initial characterization of the tripartite U sensor was performed in M5G G2P, given the robust U-dependent induction of both Pphyt and P mCB under these conditions and the lack of induction in growth media containing orthophosphate, where U bioavailability is low as a result of uranyl phosphate mineralization. [161, 173] Notably, the basal fluorescence output of the sensor was indistinguishable from a control strain lacking GFP 1-9 expression, supporting a very low OFF state (Figure 14B). Addition of uranyl nitrate, but not nitrate alone (Figure 33), yielded a nonlinear fluorescence response curve that deviated from background at concentrations as low as 5 mM and plateaued at ~15 mM (Figure 14B). The hypersensitivity of the U response curve was quantified by fitting the data with a Hill function. It is noted that this approach is semiempirical, and it is not used as a basis to derive insight on the mechanisms of the reactions occurring in the system. The model revealed a hill coefficient hypersensitive nature of the U response curve suggests that the core sensor is well poised as a qualitative YES/NO digital sensor of environmental U.
[00513] Importantly, U-dependent fluorescence was not observed with control sensor variants that lacked either a functional UrpR binding site in PPhyt-short or gfpl-9 expression (Figure 14B). Additionally, deletion of uzcR severely impaired, but did not completely abolish, U-dependent fluorescence (Figure 14B). Since the data in Figure 29A indicate that P urCB activity is abolished by uzcR deletion, the minor induction of sensor fluorescence in the A uzcR strain may reflect low- level transcriptional read-through of the transcription terminator separating Ppim-^mn-gfplO-KI and P urcB-El-gfpll modules (diagramed in Figure 31). Collectively, these data highlight the requirement for expression from all three promoters for a U-dependent fluorescent output.
Example 16: U-sensing AND gate exhibits improved selectivity relative to UzcRS alone
[00514] To characterize the selectivity of the core sensor, fluorescence was quantified in response to known inducers of either TCS and compared to results obtained with control sensor variants where either UzcRS or UrpRS governs the expression of both gfplO and gfpll (Figure 40). Consistent with the U response curves for Pphyt-short- and PMrcs-gfp fusions (Figure 29A), the control sensor driven by UrpRS alone yielded a U response curve with greater sensitivity compared to the sensor driven by UzcRS alone (nH of 12.5, K of 8.8 mM compared to nH of 7.9 and K of 17.4 mM; Table 11; Figure 34).
Figure imgf000186_0001
[00515] As expected, the core sensor was unresponsive to Ni, Cd, Th, Al, Fe(III), Fe(II), Mn, Co, arsenate, Se, and chromate (Figure 35). The core sensor and the sensor variant driven by UrpRS alone were also unresponsive to Cu, in contrast to the sensor driven exclusively by UzcRS (Figure 40). Critically, the sensitivity of the core sensor to Zn and Pb was significantly diminished relative to the UzcRS control; high Pb concentrations (greater than 20 uM) were required for weak fluorescence induction while the Zn-induced fluorescence was reduced relative to U for every tested concentration. Despite the improved selectivity for U, low Zn concentrations (5 mM Zn) yielded a fluorescence response in the core sensor that exceeded the U response (Figure 37). Nevertheless, the lack of Cu responsiveness and the weakened Zn/Pb responsiveness of the core sensor relative to a sensor constructed with UzcRS alone highlights the selectivity improvement of the AND gate approach.
Example 17: Integration of UzcY signal amplifier improves U sensitivity and selectivity. [00516] A notable limitation of this AND gate approach is that the improved selectivity comes at a cost to U sensitivity in the low micromolar range; incorporation of the less sensitive UzcRS TCS yielded a sensor with lower sensitivity compared to the control sensor driven by UrpRS alone (Figure 34). Notably, swapping the UzcRS -regulated P mCB promoter with P urCA, a promoter that is highly induced by UzcR and sensitive to low UzcR-P concentrations, failed to significantly improve sensor sensitivity (Data not shown). This suggests that simply swapping P urcB with an alternative UzcR-regulated promoter is unlikely to remedy the sensitivity limitation.
[00517] As an alternative approach to improve the coupling and matching of the UzcRS and UrpRS inputs, the use of a signal amplifier module was considered to boost the sensitivity of UzcRS. While the hrp (hypersensitive response and pathogenicity) amplifier from Pseudomonas syringae appeared to be a logical choice given its impressive ability to increase the sensitivity and output dynamic range of the ArsR-based arsenic sensor, it was unable to generate a functional U sensor using hrpR, hrpS, and P hrPL components in C. crescentus. Instead, the recently identified membrane protein UzcY that functions as a native signal amplifier for the UzcRS system was leveraged. Under normal growth conditions, uzcY expression is silenced by the MarR family regulator MarRi (CCNA_03498) and has no effect on UzcRS activity. However, when UzcY expression is induced by deleting marRi, the sensitivity of UzcRS to its metal inducers is enhanced. The signal amplification mechanism was incorporated within the core sensor by either deleting marRi , which yields a constitutive amplifier function, or by swapping the native uzcY promoter with the UrpRS-regulated Pphyt-short (Figure 36A), such that UzcY levels are modulated in a U-concentration dependent manner.
[00518] Constitutive UzcY expression significantly enhanced the U sensitivity of the core sensor in the low U concentration range (5-10 mM), yielding a U-dependent fluorescence profile with comparable sensitivity (nH of 11, K of 6.3 mM; Figure 36B; Table 11) to the sensor driven exclusively by UrpRS alone (nH of 12.5; K of 8.8 pM). A diminished fluorescence output was observed at U concentrations above 10 pM, and may reflect reduced tolerance of the A marRi strain to U toxicity compared to WT as was previously observed for Zn. Placing uzcY under the control of Pphyt-short, and, thus, conditionally restricting UzcY function to conditions of UrpRS stimulation, improved sensitivity (nH of 9.5, K of 8.4) in the 7.5-12.5 pM range, but not to the same degree as constitutive uzcY expression (Figure 36B). Importantly, neither amplifier configuration altered the Zn-, Cu-, or Pb-dependent induction profile of the core sensor (Figure 36C). By enhancing U sensitivity without affecting Zn sensitivity, the expression of UzcY significantly improved the U to Zn output ratio in the low concentration range (5-10 mM; Figure 37).
[00519] Collectively, these data suggest that in M5G G2P, integration of the UzcRS- specific signal amplifier UzcY overcame the sensitivity limitations of the UzcRS TCS, and enhanced the selectivity for U compared to the core sensor. While the sensitivity and selectivity of this amplified sensor are comparable to the sensor driven by UzcRS alone, the use of a combinatorial sensing approach is expected to be beneficial for minimizing core sensor cross-reactivity with yet unidentified inducers of UrpRS. This is based on the expectation that UrpRS is directed to detect a stress that is likely to be encountered in the oligotrophic freshwater environment of this bacterium.
Example 18: Whole-cell U sensor detects as low as 1.0 uM in groundwater
[00520] To test the efficacy of the sensor to detect U in ground water, samples from three distinct locations were obtained from LLNL site-300, a high-explosives test facility in the Altamont Hills of California where ground water concentrations of U exceed the EPA MCL. Quantification of the heavy metal content of each sample using ICP-MS or ICP-OES revealed U concentrations that range from 1.0 to 1.24 mM (238-295 ppb; Table 12), representing a challenging test for the whole-cell sensor given the ~5 pM detection limit observed in M5G G2P medium. The trace metals Zn, Pb, Cu, Cd, and Cr were either undetectable or in the low nanomolar range in all samples (Table 12), and thus not expected to affect sensor performance.
Figure imgf000188_0001
[00521] The fluorescence output of the core sensor, UzcY amplifier variants, and negative controls lacking UzcRS, UrpRS, or gfpl-9 input were monitored as a function of time in the site- 300 samples with and without growth nutrient supplementation. Exposure of the sensor strains to unmodified site-300 samples yielded no detectable increase in fluorescence and no cell growth (Figure 38A; Data not shown). In contrast, supplementation with glucose, glycerol-2-phosphate (P source), and ammonium chloride (N source) together, but not glucose alone (data not shown), enabled cell growth and yielded a detectable increase in fluorescence for the core sensor within six hours of exposure to each sample (Figure 39A). The fluorescence induction of the core sensor was significantly reduced in all samples when G2P was replaced with orthophosphate, which reduces U bioavailability (Figure 38B). For samples 1 and 3 (~ 1 mM U), the PPhyt-short- uzcY amplifier variant yielded a fluorescence induction profile with slightly lower amplitude compared to the core sensor, while the constitutive amplifier failed to produce an output that differed from the negative controls (Figure 39A). Future efforts will seek to optimize the constitutive level of UzcY to improve this sensor variants performance in environmental samples. In contrast, in sample 2 (-1.24 mM U), both signal amplifier variants yielded a fluorescence output signal that exceeded that of the core sensor (Figure 39A), suggesting that UzcRS is limiting U sensor sensitivity in sample 2. Fastly, compared to the negative controls lacking a UrpR binding site or gfpl-9 expression, the D uzcR mutant yielded a fluorescence induction profile with similar kinetics, but lower signal amplitude compared to the core sensor. This result suggests that while P urcB-El-gfpll is incompletely insulated from upstream UrpRS- mediated transcriptional activity, the function of both TCS is required to produce the fluorescence response of the core sensor in the ground water samples.
[00522] To further confirm U detectability in the ground water samples, uranyl nitrate was added in small, incremental amounts (1 uM increments), and fluorescence was quantified following a six-hour exposure (Figure 39B). The rational is that if U is responsible for the fluorescence induction of the whole-cell sensors then small, incremental U additions should further boost sensor fluorescence. A nearly identical linear increase in fluorescence as a function of U concentration was observed for the core sensor in all three ground water samples. The Pphyt- short-nzcF amplifier variant yielded a comparable fluorescence output as the core sensor in sample 1, but enhanced the signal amplitude for all U concentrations in samples 2 and 3, supporting its ability to increase sensitivity in an environmental context. Collectively, these data confirm the functionality of the AND gate sensor to detect U in environmental samples and support a lower limit of detection for U of ~1 mM (-238 ppb). This result is in agreement with the -0.5 mM detection limit reported for the UzcRS -regulated P cA-gfpUV sensor[163] and comparable with other field-portable U detection methodologies such as gamma spec and X-ray fluorescence. While additional work will be required to achieve a detection limit on par with the EPA MCL (30ppb), the current sensitivity of the AND gate sensor is well suited as a screening mechanism (e.g., yes/no) for elevated U concentrations in regions with known or suspected anthropogenic activities.
[00523] The finding that cell growth and U solubility are required for robust U detection has important implications for further sensor development. The inability of the biosensor to detect insoluble forms of uranyl (e.g., uranyl phosphate minerals) may be exploited as a mechanism to distinguish natural from anthropogenic U since natural U commonly occurs in the form of insoluble minerals, [37] and aqueous phosphate concentrations are typically very low (<10 ppb). [34] To circumvent the U-Pi incompatibility, the organophosphate G2P was employed that serves the dual function of providing a phosphate source for cell growth while maintaining initial U solubility through complexation.[161, 173] Since the physicochemical form— or speciation of U— is dependent on the geochemical conditions and strongly influences bioavailability, and consequently the detectability by this biosensor, addition of G2P may be an effective means of conditioning the environmental samples for U detection. Evaluating this hypothesis, and ultimately, the utility of the sensor for environmental monitoring, will require systematic characterization of the solution matrix composition. Nevertheless, given the requirement for nutrient supplementation, it is expected that the path toward a fieldable sensor will entail encapsulation of the whole-cell sensor within an integrated detection device that maintains cells in an active state of growth and automates sampling and solution conditioning (e.g., addition of G2P) for detection. Recent efforts have yielded promising results for integrating cell sensors into field-applicable autonomous devices. [137, 138, 174]
[00524] By identifying and integrating two independent, U-responsive TCS within a synthetic AND gate circuit, a selective U- sensing functionality was developed in C. crescentus. The results highlight the value of a combinatorial approach for selective detection of compounds for which there are no known evolved regulators. This approach is expected to be generalizable and to drive the development of additional, bio-based modules for environmental toxin detection.
Example 19: Naturally occurring Fluoride sensing riboswitches and related consensus sequence
[00525] The Rfam database was queried to identify naturally occurring F-sensing riboswitches comprising a crcB motif.
[00526] 2138 fluoride riboswitch sequences were identified and aligned and conserved nucleotides identified from a gapped alignment as illustrated in Figure 43, which shows the conserved nucleotides in naturally occurring Fluoride sensing riboswitches.
[00527] The sequences of the exemplary fluoride sensing riboswitches from the Rfam database are reported in Appendices I and II incorporated herein by references in their entirety.
Example 20: Establishing a fluoride-detection capability in C. crescentus
[00528] Environmental detection of UF6— or the more stable hydrolysis product UO2F2, which is rapidly formed when atmospheric UF6 reacts with water vapor [29] strongly suggests an enrichment program. A bacteria-based U02F2-sensor will be developed by integrating and optimizing a fluoride-sensing functionality within the engineered whole-cell U biosensor.
[00529] Figure 44 shows a schematic illustration of an approach which will be used to develop a bacteria-based sensor of uranyl fluoride products. Using a synthetic biology approach, uranium- and fluoride- detection components are expected to be integrated and optimized within C. crescentus (right panel) and the UO2F2 detection performance under aqueous conditions systematically characterized. The left schematic depicts a foreseeable application of the whole cell sensor: autonomous environmental monitoring for aqueous UO2F2 species by a C. crescentus monolayer biofilm.
Example 21: Construction of a crcB-mCherry fusion and test performance in C. crescentus
[00530] To establish a fluoride detection capability in C. crescentus , the fluoride sensing crcB riboswitch [102] will be engineered to control the expression of a fluorescent reporter ( e.g ., mCherry) such that fluoride perception leads to cell fluorescence. Initial sensor strains will be built with the crcB motifs from three distinct bacteria, including the well-characterized crcB motifs in P. syringae DC3000 and Bacillus subtilis [102] and an uncharacterized crcB motif from Sphingomonas sp. MM-1. Sphingomonas sp. MM-1 represents a particularly promising option since this bacterium is closely related to C. crescentus and possesses similarly high genomic G+C content (6x %), minimizing compatibility risks with the C. crescentus transcriptional machinery. Additionally, despite the lack of biochemical characterization, the function of the Sphingomonas sp. crcB motif in fluoride sensing is supported by its homology with characterized crcB riboswitches [Pfam database [109]] and its genomic location upstream of a fluoride exporter.
[00531] To accomplish a targeted integration of crcB riboswitches within C. crescentus , all crcB-mCherry reporters will be integrated within the C. crescentus chromosome. mCherry fluorescence will be quantified in a high-throughput (96-well format) manner over a range of NaF concentrations. Then, the fluoride selectivity will be characterized using commonly encountered anions (e.g., Cl , NO3 , SO42 , PO43 ).
[00532] Based on prior biochemical characterization of the P. syringae crcB riboswitch [102], we expect the crcB-mCherry reporters to exhibit high fluoride selectivity. Each sensor will be evaluated based on the detection limit, sensitivity, selectivity, dynamic range, and signal amplitude with the best performing crcB-mCherry reporter used for subsequent studies.
[00533] Figure 45 shows a schematic illustration of the approach. (Left) Phase I fluoride sensor: crcB-mCherry fusion built in wild type strain (Right). Phase II fluoride sensor: crcB-mCherry fusion built in strain deleted for fluoride export (D crcB). Deletion of the crcB gene in E. coli improved the detection limit of a crcB-lacZ reporter by over 100-fold [102]. As such, this approach is expected to yield a fluoride-sensing capability that, when coupled with a U- sensing component herein described, will enable UO2F2 detection by C. crescentus.
Example 22 Reduction of the fluoride detection limit by manipulating the expression levels of the native fluoride detoxification svstem/s in C. crescentus.
[00534] C. crescentus possesses a putative fluoride ion transporter (CrcB) and withstands high mM concentrations of fluoride, suggesting native mechanism/s of fluoride detoxification. Prior data in E. coli indicate that the CrcB fluoride exporter functions to maintain low intracellular fluoride concentrations, adversely affecting the detection limit of the colorimetric crcB reporter (~ 1 mM). As such, the detection limit of phase 1 fluoride sensors development is expected to be significantly higher than the detection limit of our U sensor (~2 mM), and thus unlikely to be acceptable for nuclear effluent detection.
[00535] Accordingly, an analogous approach will be implemented as what allowed E. coli to yield a 100-fold improvement in the fluoride detection limit (sub 10 mM) [102]. Specifically, the function of the fluoride ion exporter will be abolished by deleting the crcB gene (encodes the CrcB fluoride exporter) in a strain containing the crcB-mCherry reporter (Figure 45). Subsequently, crcB-mCherry fluorescence will be quantified over a range of NaF concentrations.
[00536] An analogous reduction in the F- detected limit in C. crescentus or other potential host organisms is expected be achieved through deletion of the crcB homolog or eriCF, an analogous fluoride exporter. Both the CrcB and EricF proteins have been previously defined. [102]
[00537] If deletion of crcB fails to improve crcB-mCherry performance or reduces the ability of C. crescentus to withstand high fluoride concentrations, the genes responsible for fluoride tolerance will be experimentally identified using one of two unbiased whole genome approaches: 1) RNA-seq to identify genes that increase in expression following fluoride exposure, and 2) transposon mutagenesis to screen for genes that, when inactivated, enhance the fluoride sensitivity of the crcB-mCherry reporter. The team successfully applied both approaches to identify the U-responsive regulators that form the basis of the whole-cell sensor [175].
[00538] One caveat of eliminating the native mechanism of fluoride efflux in C. crescentus is the reduced tolerance of cells to environmental fluoride. For example, the minimal inhibitory concentration of fluoride was reduced from 200 to slightly above 1 mM in an analogous E. coli crcB mutant. Since fluoride levels in groundwater, sea, and soil are typically in the 10-100 mM range [176], the reduced tolerance is not expected to be problematic except in the most highly concentrated environments.
[00539] If the C. crescentus crcB mutant will be too sensitive to fluoride (MIC < 1 mM), the fluoride tolerance will be partly rescued by expressing the crcB fluoride exporter over a range of concentrations to identify levels that yield a desirable balance between the organism’s fluoride tolerance and the sensitivity of the crcB-mCherry.
Example 23, Integration of uranyl- and fluoride-sensing components and evaluate detection performance under aqueous conditions.
[00540] A U-sensing AND gate circuit will be integrated with the optimal performing crcB- mCherry circuit within C. crescentus to enable UO2F2 detection. Two distinct configurations will be constructed and further characterized.
[00541] Figures 46 and 47 show a schematic illustration of the related approach.
[00542] The configurations of the exemplary U02F2-biosensors of Figure 46 and Figure 47 will provide an individual readout of uranium and fluoride levels and provides a safeguard for environments where natural uranium or fluoride occur at elevated levels. Uranium occurs naturally at concentrations of -10-100 ppb in soils and -5- 100s of ppb in ground water [37]while fluoride levels in groundwater, sea, and soil are typically in the 10-100 mM range [176].
[00543] Continuous environmental monitoring with the biosensor is expected to identify significant deviations from background levels for either or both analytes that it is expected to signify an anthropogenic release. This circuit configuration also enables the flexibility to in dividually adjust the sensitivity of the U and F detection components to suit the specific applica tion or region of interest and in particular a lower fluoride detection limit (compared to U) may be acceptable if the whole cell sensor is integrated with an air-sampling device that collects UO2F2 aerosols and HF formed from atmospheric hydrolysis of UF6.
[00544] Specific configurations are exemplified in Example 24 and 25.
Example 24: UO2F2 single output biosensor configuration: Integrated uranium- and fluoride-sensing components within an AND gate circuit
[00545] A UO2F2 sensor 2 will be constructed with an integrated uranium- and fluoride-sensing components within an AND gate circuit such that GFP fluorescence is produced only when the cell perceives both U and fluoride. [00546] In particular, uranium- and fluoride- sensing components can be integrated within a configuration such that a single reporter molecule (e.g., mCherry, GFP) is produced only when the cell perceives both U and fluoride.
[00547] Examples of single output UOiFibioscnsor configuration are shown in Figure 46
[00548] In particular in Figure 46 panel A a configuration is schematically illustrated wherein uranium- and fluoride-sensing components are integrated in series. The promoter can be any UzcR- or UrpR-regulated promoter (defined in prior patent app) such that U-dependent transcription is initiated by UzcR or UrpR. In this circuit, transcription will be prematurely terminated by the fluoride sensing riboswitch in the absence of fluoride (i.e., not detectable output). Binding of fluoride to the riboswitch will mediate transcriptional read-through and ultimately, production of the reporter.
[00549] In Figure 46 panel B a configuration is schematically illustrated wherein uranium- and fluoride-sensing components integrated in series where U-dependent transcriptional activation requires the function of both UrpR and UzcR. Will provide greater selectivity for uranium compared to the sensors described in panel A. However, in this configuration, the sensitivity for U would be limited by the UzcRS component.
[00550] In Figure 46 panel C a configuration is schematically illustrated wherein integration of the fluoride riboswitch within the AND gate circuit such that expression of component three requires fluoride exposure. In this configuration, reconstitution of GFP fluorescence requires activation of the two uranium-responsive pathways and fluoride binding to the fluoride riboswitch. Transcription of component three can theoretically be controlled by any constitutive promoter. The assignment of each component with the given regulatory promoter is arbitrary and easily swapped. For example, the fluoride riboswitch could be used to control expression of component one or two.
[00551] The single output of this configuration simplifies detection and will be most useful for an all-or-none screening function that fluoresces when threshold U and fluoride levels are encountered. The sensitivity of detection would be theoretically limited by the least sensitive component. [00552] The single output of this configuration also simplifies standoff detection and will be most useful for an all-or-none screening function that fluoresces when threshold U and fluoride levels are encountered.
[00553] While the sensitivity of detection would be theoretically limited by the least sensitive component, this configuration will focus on improving the sensitivity of the U and fluoride components. To construct the UO2F2 sensor 2, the crcB riboswitch will be integrated within the state-of-the art AND gate circuit exemplified in the present disclosure such that expression of component three requires fluoride exposure.
[00554] In this configuration, reconstitution of GFP fluorescence requires activation of the two uranium-responsive pathways and fluoride binding to the crcB riboswitch.
Example 25 UO2F2 dual output biosensor configuration with uranium and fluoride-sensing riboswitch in separate circuits with different output reporter.
[00555] Uranium- and fluoride-sensing components can be integrated within a configuration such that integrated uranium- and fluoride- sensing components are provided as separate components or separate circuits with different output reporters (Dual Output UO2F2 Sensor)
[00556] Exemplary configurations of a dual output U02F2biosensor are shown in Figure 47
[00557] In particular, in Figure 47 an exemplary sensor 1 is schematically described in comprising a U-sensing genetic circuit and a separate a Fluoride sensing reportable genetic molecular components.
[00558] In the illustration of Figure 47, the U-sensing genetic circuit comprises a reportable genetic molecular component which is expressed when the genetic circuit operates according to the circuit design in presence of bioavailable U. in the exemplary illustration of Figure 47, the reportable component is the GFP protein which is expressed upon expression of a first U-sensing genetic molecular component comprising a PUzcrB promoter encoding for a first component of the GFP protein, a second U sensing component comprising a Pphyt promoter and encoding for a second component of the GFP protein and a genetic molecular component a third component of the GFP protein under an exemplary active constitutive promoter (rsA). [00559] In the illustration of Figure 47, the Fluoride sensing reportable genetic molecular component is a genetic reportable molecular component expresses the fluorescent reporter mCherry within a gene comprising a Fluoride sensing riboswitch under the control of an crcB promoter to provide a Fluoride sensing component.
[00560] To construct the biosensor in the output configuration schematically illustrated in exemplary Figure 47, the optimal performing crcB-mCherry fusion will be integrated into the chromosome of a C. crescentus strain that already contains the tripartite GFP U-sensing AND gate. As such, mCherry and GFP fluorescence can be quantified to provide a separate readout of environmental fluoride and U levels.
[00561] In particular the variation of the biosensor described therein is shown in Figure 47 wherein in the fluoride sensing reportable genetic molecular component the fluorescent reporter mCherry within a gene comprising a Fluoride sensing riboswitch is under the control of an active constitutive promoter to provide a Fluoride sensing component.
[00562] This configuration will provide an individual readout of uranium and fluoride levels and provides a safeguard for environments where natural uranium or fluoride occur at elevated levels. Uranium occurs naturally at concentrations of -10-100 ppb in soils and -5- 100s of ppb in ground water [37] while fluoride levels in groundwater, sea, and soil are typically in the 10-100 mM range. [176]
[00563] Continuous environmental monitoring with the biosensor could identify significant deviations from background levels for either or both analytes that could signify an anthropogenic release. This circuit configuration also enables the flexibility to individually adjust the sensitivity of the U and F detection components to suit the specific application or region of interest— a lower fluoride detection limit (compared to U) may be acceptable if the whole cell sensor is integrated with an air-sampling device that collects UO2F2 aerosols and HF formed from atmospheric hydrolysis of UF6. To construct this sensor, the fluoride riboswitch will be placed downstream of a constitutive promoter and upstream of a reporter (e.g., mCherry), such that transcription of the reporter occurs only in the presence of fluoride. This crcB-mCherry fusion will be employed in tandem with the tripartite GFP U-sensing AND gate. As such, mCherry and GFP fluorescence can be quantified to provide a separate readout of environmental fluoride and U levels.
Example 26: F-sensing riboswitch of host bacteria knocked out
[00564] To mitigate fluoride toxicity, many bacteria employ a fluoride export protein (e.g., CrcB or EricFthat functions to maintain low cytoplasmic fluoride concentrations.
[00565] Prior data in E. coli indicate that this CrcB -mediated fluoride exporter adversely affects the detection limit of the colorimetric crcB reporter (~1 mM) as a consequence of the low intracellular fluoride concentrations.
[00566] A similar fluoride export activity is expected in C. crescentus since C. crescentus possesses an annotated CrcB transporter with high amino acid sequence similarity to previously characterized CrcB fluoride- specific ion channels [177] and tolerates high mM concentrations of fluoride.
[00567] Accordingly, to achieve a fluoride detection limit on par with the U detection limit of our sensor (~1 mM), we will implement an analogous approach as that in E. coli, which yielded a 100-fold improvement in the fluoride detection limit (sub 10 mM) and maintained a dynamic range that spanned two orders of magnitude. [102]
[00568] Specifically, the function of the fluoride ion exporter in C. crescentus was tested by monitoring the growth of a strain deleted for crcB in the presence of NaF and NaCl (Figure 49).
[00569] The results illustrated in Figure 49 panels A and B show growth defect is observed for WT at NaF concentrations above 4 mM and for the crcB deletion strain at concentrations above 126.5 pM, marking a greater than 30-fold reduction in Fluoride tolerance in the absence of the fluoride efflux pump.
[00570] The results illustrated in Figure 49 panels C and D show that growth of C. crescentus in the presence of NaCl is unaffected by deletion of crcB. These data confirm the role of C. crescentus CrcB in fluoride efflux.
[00571] Importantly, the growth profile of the crcB deletion strain was indistinguishable from WT over a range of NaCl concentrations (Figure 49 panel B ), supporting a specific role for C. crescentus CrcB in fluoride efflux.
Example 27: Comparison of riboswitches performance in a LacZ fusion construct
[00572] To enable detection of fluoride in C. crescentus, a fluoride sensing riboswitch [102] will be employed to control the expression of a fluorescent reporter ( e.g ., mCherry) such that fluoride detection results in cell fluorescence (Figure 55).
[00573] Since C. crescentus lacks a native fluoride riboswitch, a library of riboswitch reporters (transcriptional fusions to mCherry) was built and screened that included the well-characterized fluoride riboswitches from P. syringae DC3000 [102] and previously uncharacterized fluoride riboswitches from Sphingomonas sp. MM-1 and Sphingomonas sp. 67-36, which are closely related to C. crescentus and possesses similarly high genomic G+C content (66%) (Figure 50 Panels A to D).
[00574] The sequences of the riboswitches and related constructs are reported in Table 13 below.
Table 13: Sequences of tested riboswitches
Figure imgf000199_0001
sequence ATTTCGACGACCTGAGGTACCTGC
Figure imgf000200_0001
Figure imgf000200_0002
[00575] The function of the fluoride riboswitches was tested in C. crescentus initially using translational lacZ fusions with the associated native crcB or ericF promoters schematically illustrated in Figure 50 panel A.
[00576] In particular the function of fluoride riboswitches from three different bacteria was tested in Caulobacter crescentus with a construct in which each riboswitch was included between a Pcrcb and a LacZ reporter (Figure 50 Panel A)
[00577] In particular, DNA sequence spanning from 197 and 180 nucleotides upstream of the start codon for Sphingomonas MM-1 and 67-36, respectively, through the first 7 amino acids of the crcB gene was fused to lacZ to produce a translational fusion. A DNA sequence range from 236 nucleotides upstream of the start codon through the first 11 amino acids of the EricF efflux pump was used for Pseudomonas syringae.
[00578] The results reported in Figure 50 Panels B to D indicate a difference in performance of the different riboswitches in the construct schematically described in Figure 50, Panel A.
[00579] In particular, the fluoride riboswitch from Sphingomonas MM-1 was responsive to fluoride and had the largest dynamic range and signal amplitude (Figure 50 Panel B), while the riboswitch from Sphingomonas 67-36 was non-responsive to fluoride (Figure 50 Panel D).
[00580] In particular, the construct for Sphingomonas 67-36 lacked fluoride induction, which is expected to reflect a lack of a promoter sequence in the construct or weak promoter activity in C. crescentus.
[00581] The riboswitch from Pseudomonas syringae was also fluoride responsive, but had a lower dynamic range compared to the riboswitch from MM-1 (Figure 50 Panel C).
[00582] Importantly, this set of experiments further evidenced that the crcB deletion strain significantly improved the F sensitivity compared to WT with induction occurring at almost two orders of magnitude lower F concentration.
An additional FacZ fusion including Sphingomonas MM-1 riboswitch sequence with a reduced G+C content is
CAGAGATCTAACGCCGGTGACGGCGCGCGTGGCGAAGGGCGCGATCGAGGTGGCCG
CGCCGGCTGGGTGACGGGCTGCCATTGACGCGCTGCCCCACCCCCGCCATATCGTCG
GCAACGGCAATGGATTCCTGCCGGGCCTCGCGCCGAACCGCCATTGAGGGCTGATG
ATTCCTACCTGCGGCCGCCCTGCGGCAAGGAGGAATCATGCGTTCGGTATTTCTCGT
CTTGGT ACCTGC (SEQ ID NO: 2527).
Example 28: Modularity of exemplary riboswitches
[00583] The function of each ribo switch outside of the context of the native promoter was tested to determine the modularity with respect to the promoter sequence.
[00584] Modularity with respect to the promoter is critical feature for a fluoride riboswitch- based sensor since this would allow the use of a strong constitutive promoter (to improve the signal amplitude) and a U-sensitive promoter to enable detection of UO2F2.
[00585] To this end, DNA sequence 97, 146, and 80 bp upstream of the start codon for Sphingomonas MM-1, Pseudomonas syringae, and Sphingomonas 67-36, respectively was fused to the xylose inducible promoter (Pxyl) and to mCherry in a construct schematically illustrated in Figure 51 Panel A.
Figure imgf000201_0001
was employed since it allows a wide range of promoter activities to be assayed (Figure 51Panel A).
[00586] The sequences of each riboswitch is reported in Table 14 below
Table 14: Sequences of tested riboswitches
Figure imgf000201_0002
Figure imgf000202_0001
[00587] The fluoride responsiveness of the fluoride riboswitches in such construct was tested in C. crescentus to test the compatibility of each riboswitch with a non-native promoter to test the modularity of the riboswitch.
[00588] Fluorescence was monitored as a function of xylose and NaF concentrations.
[00589] The results illustrated in Figure 51 Panel B show that expression of mCherry remains low in the absence of fluoride at xylose concentrations up to 200 uM. However, high concentrations of xylose (2-20 mM) lead to partial fluoride-independent activation. This result suggests that a single fluoride riboswitch is insufficient to halt all transcription initiation events when promoter activity is high.
[00590] The results illustrated in Figure 51 Panels C and D show that while all three riboswitches were responsive to fluoride, the riboswitch from Sphingomonas MM-1 exhibited the largest dynamic range, which is in agreement with the lacZ fusion data reported in Example 27 and Figure 50.
[00591] These results suggest that the fluoride riboswitch from Sphingomonas MM-1 is modular with respect to the promoter sequence and the most suitable for application in C. crescentus. It is expected that a further optimization of the native Shine Dalgamo sequence will enhance the dynamic range and signal amplitude of the Pseudomonas syringae riboswitch construct.
[00592] The results of the experiments related to placement of the MM-1 riboswitch downstream of the xylose-inducible promoter (Pxyl), also indicated that at high concentrations of xylose (2-20 mM) lead to partial fluoride-independent activation.
[00593] This result suggests that a single fluoride riboswitch can be insufficient to halt all transcription initiation events when promoter activity is high. In those cases, a second MM-1 riboswitch, added in tandem, is expected to increase the probability of terminating transcription in the absence of fluoride, which will reduce the level of fluoride-independent activation and is also expected to increase the dynamic range. Preferably, when two riboswitches are placed in tandem, the first riboswitch does not contain the native RBS or downstream protein coding sequence.
[00594] In addition, it is also expected that addition of a spacer DNA region between the DNA regions encoding each riboswitch will increase performance of the F-sensing element with MM- 1 riboswitches. A non- structured DNA of 10-50 bp is expected to be sufficient (preferably with a shorter sequence to minimize the length of the this 5’ untranslated region). If the RBS and downstream protein coding sequence cannot be removed without losing fluoride binding activity, then the DNA encoding the first 50 amino acids can be fused to the first riboswitch and terminated with a stop codon.
[00595] The results illustrated in Figure 51 Panels E and F further show that the crcB deletion strain has significantly improved F- sensitivity compared to WT. These results further confirm that the crcB deletion background can be employed for all subsequence experiments to maximize fluoride sensitivity.
[00596] The above results support the use of the Sphingomonas MM-1 fluoride riboswitch in a crcB deletion background for subsequent experiments.
Example 29: Effects of different fusion lengths and attenuators on fluoride detection of F sensing riboswitches
[00597] In order to test the length of the fusion including a fluorescence protein [113] further experiment have been performed to identify the lengths and associated fluoride detection to identify an optimized length.
[00598] In particular, the effect of two different translational fusion lengths (7 and 50 amino acids) and two so-called ribo-attenuators (ATT2 and ATT3) [113], which are genetic elements designed for increased ribosome modularity through predictable tuning, insulation from contextual changes, and a reduction in expression variation.
[00599] Accordingly, the DNA encoding 0, 7, and 50 amino acids of Sphingomonas MM-I CrcB was fused to mCherry and the native crcB promoter was used to drive expression of the riboswltch reporter construct
[00600] A transcriptional mCherry fusion was also tested, where the crcB gene and the shine Dalgamo sequence were deleted from the construct.
[00601] The results reported in Figure 54 show that the constructs containing either a 50 AA translational fusion or a transcriptional fusion exhibited a fluoride-dependent increase in fluorescence, whereas the 7 AA translational fusion was only very weakly responsive to fluoride f Figure 54).
[00602] These data suggest that the MM-l ribos witch can function independently of its native crcB gene and that a longer translational fusion with mCherry is preferred vs a shorter fusion in the context of MM-I riboswltch function.
Example 30: Selection of promoters and F sensing riboswitches for an F-sensing genetic reportable molecular component
[00603] To develop a uranyl fluoride sensor, the UrpRS regulated promoter p phyt-sbort was integrated upstream of the MM-I riboswltch, replacing the native crcB promoter. The native shine Dalgamo sequence for the crcB gene was not modified and a translational mCherry promoter was constructed with the first 150 nucleotides (50 amino acids) of the crcB gene. While the presence of U or fluoride alone yielded a minor increase in fluorescence, the presence of both U and fluoride yielded the highest fluorescence increase, confirming the intended functionality of the circuit (Figure 55 Panels B and C). in contrast, the control fluoride sensor containing the native crcB promoter was unresponsive to uranium
MM-1 fluoride sensing circuit with native crcB promoter is not responsive to U. (B) A uranyi fluoride sensing circuit constructed by combining the IJrpRS- responsive Pphyt promoter with ihe MM-1 ribos itcii. (C) Depicts the fluorescence of the uranyi fluoride sensing circuit in the presence of U alone, F alone, and both U and F Tests-were performed with F-, added as NaF, and uranyi, added as uranyi nitrate. The presence of both U and fluoride yields the highest signal amplitude.
Example 31: Selection of promoters and F sensing riboswitches for an F-sensing genetic reportable molecular component
[00604] The following approach can be followed to optimize promoter for a bacterial host and application of interest for cassettes where U and F sensing components are placed in different circuits (e.g., Figure 46C; Figure 47).
[00605] The consensus promoter sequence for the housekeeping sigma factor (sigma73) in Caulobacter crescentus has been previously determined [178] .
[00606] For optimal fluoride riboswitch reporter function, mutating the native crcB/ericF promoter toward the C. crescentus consensus is recommended. For example, it is expected that mutating the -10 region of the Sphingomonas MM-1 crcB promoter from 5’-GCCATATC-3’ (SEQ ID NO: 2521)to 5OCTATATC-3’ (SEQ ID NO: 2522) will enhance the signal amplitude of the fluoride sensing reporter. The optimal functional fluoride reporter can be integrated within a C. crescentus strain containing a EG-sensing promoter (any from prior patent app) for a two-color output (e.g., Figure 47) or integrated within the EG sensing AND gate (two EG- responsive components and one F-responsive component (e.g., Figure 46C)) for a single-color output.
Example 32: Selection of promoters and F sensing riboswitches for a U sensing F-sensing genetic reportable molecular component
[00607] For application wherein the optimization is desired when combining a ET-responsive promoter with a fluoride riboswitch (e.g., Figure 46B-C) to place the EG-responsive and F- responsive components within a single circuit, the native crcB promoter can be replaced with a uranium responsive promoter (any UzcRS or UrpRS promoter). For example, The Pphyt-short promoter was added upstream of the Sphingomonas sp MM-1 fluoride riboswitch and exhibited the highest fluorescent output signal in the presence of uranyl and fluoride. It is anticipated that the signal amplitude can be further improved by optimizing the RBS as described below. Furthermore, the F-independent output signal is expected be reduced by adding a second riboswitch sequence, resulting in an improve dynamic range.
Example 33: Selection of RBS sequence for inclusion in F-sensing component and U- sensing F-sensing components
[00608] 3Since the fluoride riboswitch has been shown to function through a mechanism of transcriptional termination, the ribosome binding site is likely amenable to optimization. The RBS can be optimized using the standard registry of parts (http://parts.igem.org/Main Page at the date of filing of the instant application), which contains several well-characterized RBS sequences that are likely to work in a broad range of host organisms. For example, it is expected that swapping the native Sphingomonas crcB RBS (5’-AAGGAGGAATCATG-3’) (SEQ ID NO: 2523) with the potent 5’ -AAGGAGGAAAAACATATG-3’ RBS (SEQ ID NO; 2524) will improve the dynamic range.
[00609] In summary described herein are U biosensors, and related U-sensing genetic molecular components, genetic circuits, compositions, methods and systems are described, which in several embodiments can be used to detect and/or neutralize uranium and in particular bioavailable U.
[00610] Described herein are UO2F2 biosensors, and related U-sensing and/or F-sensing genetic molecular components, genetic circuits, compositions, methods and systems are described, which in several embodiments can be used to detect and/or neutralize uranium and in particular bioavailable UO2F2..
[00611] In particular , in a first set of embodiments a U02F2-biosensor comprising a genetically engineered bacterial cell capable of heterologously and/or natively expressing histidine kinase 1363, and U-sensitive transcriptional regulator 1362. In the U02F2-biosensor of the first set of embodiments, the bacterial cell is an engineered bacterial cell including a U-sensing genetic reportable molecular component and/or a U-sensing/U-neutralizing genetic molecular component, each comprising a U-sensitive promoter in a configuration wherein the U-sensitive promoter directly initiates expression of the U-sensing reportable molecular component and/or of the U-sensing U-neutralizing molecular component in presence of bioavailable U;
In the UC Fi-biosensor of the first set of embodiments, the U-sensitive promoter comprises a U- sensitive transcriptional 1362 binding site having a DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO:l), in which
Ni is C or T, preferably C;
N2 is G or A, preferably G;
N3 is T or C, preferably T;
N4 is C;
Ns is A or G, preferably A;
Nό is G or C, preferably G;
N7 C or G;
Ns is any nucleotide;
N9 is any nucleotide;
Niois any nucleotide;
N11 is any nucleotide;
N12 is T or C;
Ni3 is G;
Ni4 is T or C, preferably T;
Nis is C;
N16 IS A or C, preferably A;
Nn is G; and Ni8 is C or G,
and in which Ni to Nn selected independently
In the UC Fi-biosensor of the first set of embodiments, the genetically modified bacterial cell is an engineered bacterial cell further comprising an F-sensing riboswitch within at least one of
- the U- sensing reportable genetic molecular component in a configuration wherein the U-sensing reportable genetic molecular component, is transcribed in presence of an effective amount of bioavailable fluoride,
- an F-sensing reportable genetic molecular component in a configuration wherein the F- sensing reportable genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride, and
- an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F.
[00612] In some embodiments of the first set of embodiments, the UO2F2 biosensor further comprises an UzcR binding site having a DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide,
the UzcR binding site inserted at a location downstream of a transcription start site of the U- sensitive promoter.
[00613] In some embodiments of the first set of embodiments, the UO2F2 biosensor further comprises a UzcY gene and/or UzcZ gene inserted at a location downstream of a transcription start site of the U-sensitive promoter.
[00614] In some embodiments of the first set of embodiments, the bacteria of the UO2F2 biosensor are capable of natively expressing endogenous MarR family repressors such as MarRl and/or MarR2 genes and at least one gene of the endogenous MarR family, preferably all genes of the endogenous MarR family, is knocked out.
[00615] In some embodiments of the first set of embodiments, the bacteria of the UO2F2 biosensor are capable of natively expressing endogenous urtAP genes and at least one gene of the endogenous urtAP, preferably all genes of the endogenous urtAP, is knocked out.
[00616] In a second set of embodiments, UO2F2 biosensor is described comprising a genetically engineered bacterial cell capable of heterologously and/or natively expressing histidine kinase 1363, and response regulator 1362. In the U02F2-biosensor of the second set of embodiments, wherein the bacterial cell is an engineered bacterial cell comprising a U-sensitive F-sensitive genetic circuit in which molecular components are connected one to another in accordance to a circuit design by activating, inhibiting, binding, or converting reactions to form a fully connected network of interacting components.
[00617] In the U02F2-biosensor of the second set of embodiments, at least one molecular component is a U- sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U- sensitive transcriptional 1362 binding site having a DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO:l), wherein
Ni is C or T, preferably C;
N2 is G or A, preferably G;
N3 is T or C, preferably T;
N4 is C;
Ns is A or G, preferably A;
Nό is G or C, preferably G;
N7 is C or G;
Ns is any nucleotide;
N9 is any nucleotide; Niois any nucleotide;
Nil is any nucleotide;
Ni2 is T or C;
Ni3 is G;
Ni4 is T or C, preferably T;
Nis is C;
N16 IS A or C, preferably A;
Nn is G; and
N18 is C or G;
wherein Ni to Nn selected independently,
[00618] In the U02F2-biosensor of the second set of embodiments, at least one molecular component of the UO2F2 biosensor, is a reportable molecular component, and/or a U- neutralizing molecular component, the reportable molecular component and/or the a U- neutralizing molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable U.
[00619] In the U02F2-biosensor of the second set of embodiments, the U-sensing F-sensing genetic circuit of the UO2F2 biosensor further comprises an F sensing riboswitch within at least one genetic molecular component of the molecular components of the U-sensitive F-sensitive genetic circuit in a configuration wherein the at least one genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
[00620] In the U02F2-biosensor of the second set of embodiments, the reportable molecular component and/or the U-neutralizing molecular component of the UO2F2 biosensor, are expressed when the genetic circuit operates according to the circuit design in presence of an effective amount of bioavailable U and an effective amount of bioavailable F.
[00621] In some embodiments of the second set of embodiments, the UO2F2 biosensor further comprises a UczRS U sensing genetic molecular component comprising an UzcR binding site having a DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide, the UzcRS U sensing genetic molecular component operatively connected to the reportable molecular component, and/or the U- neutralizing molecular component.
[00622] In some embodiments of the second set of embodiments, the UO2F2 biosensor further comprises at least one activator molecular component comprising a UzcY gene and/or UzcZ gene, the at least one activator genetic molecular component operatively connected to the additional U sensing genetic molecular component.
[00623] In some embodiments of the second set of embodiments, the bacteria of the UO2F2 biosensor are capable of natively expressing endogenous MarR family repressors such as MarRl and/or MarR2 genes and at least one gene of the endogenous MarR family, preferably all genes of the endogenous MarR family, is knocked out.
[00624] In some embodiments of the second set of embodiments, the bacteria of the UO2F2 biosensor are capable of natively expressing endogenous urtAP genes and at least one gene of the endogenous urtAP, preferably all genes of the endogenous urtAP, is knocked out.
In some embodiments of the second set of embodiments, the U sensing genetic molecular component and UzcRS U sensing genetic molecular component of the UO2F2 biosensor are operatively connected to a same or a different reportable molecular component in the circuit.
[00625] In some embodiments of the second set of embodiments, the U-sensitive genetic circuit comprises one or more AND gates. In some of those embodiments, at least one of the AND gates is an in-series AND gate. In some of those embodiments, at least one of the AND gates is an in parallel AND gate. In some of those embodiments two or more in series AND gates and/or in parallel AND gates are connected by activating, inhibiting, binding, or converting reactions. In some of those embodiments, at least one of the AND gates is selected from the group consisting of an HRP AND gate, a bacterial two-hybrid AND gate, a tripartite GFP AND gate, and a FRET sensor AND gate. [00626] In some embodiments of the second set of embodiments, the UO2F2 biosensor further comprises a genetic signal amplifier configured to increase an output of a reportable molecular component and/or a U-neutralizing molecular component at a given concentration of bioavailable U.
[00627] In some embodiments of the second set of embodiments, the reportable molecular component and/or the U-neutralizing molecular component of the UO2F2 biosensor is formed by an assembly of two or more subunits of the reportable molecular component and/or the U- neutralizing molecular component when the U- sensitive genetic circuit operates according to the circuit design in presence of bioavailable U.
[00628] In some embodiments of the second set of embodiments, the reportable molecular component and/or the U-neutralizing molecular component of the UO2F2 biosensor is post- transcriptionally and/or post-translationally converted when the U-sensitive genetic circuit operates according to the circuit design in presence of bioavailable U.
[00629] In some embodiments of the first and second sets of embodiments, the engineered bacterial cell of the UO2F2 biosensor comprises endogenous genes encoding the histidine kinase 1363, and the U-sensitive transcriptional regulator 1362.
[00630] In some embodiments of the first and second sets of embodiments, the engineered bacterial cell of the UO2F2 biosensor comprises a U-sensing regulator component comprising endogenous and/or exogenous gene encoding histidine kinase 1363, and an endogenous or exogenous gene encoding response regulator 1362 in a configuration wherein the a gene encoding histidine kinase 1363, and a gene encoding response regulator 1362 are expressed upon activation of a controllable promoter.
[00631] In some embodiments of the first and second sets of embodiments, in the sequence SEQ ID NO:l:
Ni is C; and/or
N2 is G; and/or
N3 is T; and/or Ns is A; and/or
Nό is G; and/or
Ni4 is T; and/or
Ni6is A.
[00632] In some embodiments of the first and second sets of embodiments, the U-sensitive transcriptional 1362 binding site of the UO2F2 biosensor has a sequence selected from the group consisting of SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26 and SEQ ID NO:27.
[00633] In some embodiments of the first and second sets of embodiments, the U-sensitive promoter further comprises nucleotides N19N20N21, downstream of SEQ ID NO: 1 wherein N19 is any nucleotide; N2o is any nucleotide; and N21 is G (SEQ ID NO: 83).
[00634] In some embodiments of the first and second sets of embodiments, Nis of the regulator direct repeat is located about -17 to about -40 upstream of a transcription start site.
[00635] In some embodiments of the first and second sets of embodiments, the U-sensitive promoter is P1361 or Pphyt.
[00636] In some embodiments of the first and second sets of embodiments, the U02F2biosensor is configured to detect and/or neutralize bioavailable U present in a target environment at a concentration of 100 nM or great, between 100 nM and 1 mM, or greater than 1 pM.
[00637] In some embodiments of the first and second sets of embodiments, the U02F2biosensor is configured to detect bioavailable F present in a target environment at a concentration 10 pM, or greater, between 50 pM and 60 pM, greater than 1 nanomolar and between 100 nanomolar and 1 mM.
[00638] In a third set of embodiments, a U02F2-biosensor is described comprising a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase UzcS and U-sensitive transcriptional response regulator UzcR.
In the UC Fi-biosensor of the third set of embodiments, the genetically modified bacterial cell is an engineered bacterial cell comprising a U-sensing reportable molecular component and/or a U- sensing/U-neutralizing molecular component, each comprising a U-sensitive promoter in a configuration wherein the U-sensitive promoter directly initiates expression of the U-sensing reportable molecular component and/or of the U-sensing U-neutralizing molecular component in presence of bioavailable U,
[00639] In the UC Fi-biosensor of the third set of embodiments, the U-sensitive promoter comprises an UzcR binding site having a DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide.
[00640] In the U02F2-biosensor of the third set of embodiments, the genetically modified bacterial cell is an engineered bacterial cell further comprising an F-sensing riboswitch within at least one of
- the U-sensing reportable genetic molecular component in a configuration wherein the U-sensing reportable genetic molecular component, is transcribed in presence of an effective amount of bioavailable fluoride,
- an F-sensing reportable genetic molecular component in a configuration wherein the F- sensing reportable genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride, and
- an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F.
[00641] In some embodiments of the third set of embodiments, the U02F2-biosensor further comprises a UzcY gene and/or UzcZ gene inserted at a location downstream of a transcription start site of the U-sensitive promoter. [00642] In a fourth set of embodiments, a UO2F2- biosensor comprising a genetically modified bacterial cell natively and/or heterologously expressing histidine kinase UzcS, and response regulator UzcS,
[00643] In the U02F2-biosensor of the fourth set of embodiments, the genetically modified bacterial cell is an engineered bacterial cell comprising a U-sensitive genetic circuit in which molecular components are connected one to another in accordance to a circuit design by activating, inhibiting, binding, or converting reactions to form a fully connected network of interacting components.
[00644] In the U02F2-biosensor of the fourth set of embodiments, in the U-sensitive genetic circuit at least one molecular component is a U-sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a UzcR binding site having DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide.
[00645] In the U02F2-biosensor of the fourth set of embodiments, at least one molecular component is a reportable molecular component, and/or a U-neutralizing molecular component, the reportable molecular component and/or the U-neutralizing molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable U.
[00646] In the U02F2-biosensor of the fourth set of embodiments, the genetically modified bacterial cell further comprises an F-sensing riboswitch within at least one genetic molecular component of the molecular components of the U-sensitive F-sensitive genetic circuit in a configuration wherein the at least one genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride.
[00647] In some embodiments of the fourth set of embodiments, the UO2F2 biosensor further comprises at least one activator molecular component comprising a UzcY gene and/or UzcZ gene, the at least one activator genetic molecular component operatively connected to the U sensing genetic molecular component. [00648] In some embodiments of the fourth set of embodiments, the molecular component of the UC Fi-biosensor is a genetic molecular component wherein UzcY and/or UzcZ are comprised under control of a U- sensitive promoter and/or a controllable promoter operatively connected to the UzcY and/or UzcZ in a configuration wherein the U-sensitive promoter and/or a controllable promoter directly initiates expression of the amplifier molecular component.
[00649] In some embodiments of the fourth set of embodiments, the bacteria of the UO2F2- biosensor are capable of natively expressing endogenous MarR family repressors such as MarRl and/or MarR2 genes and at least one gene of the endogenous MarR family, preferably all genes of the endogenous MarR family, is knocked out.
[00650] In some embodiments of the fourth set of embodiments, the bacteria of the UO2F2- biosensor are capable of natively expressing endogenous urtAP genes and at least one gene of the endogenous urtAP, preferably all genes of the endogenous urtAP, is knocked out.
[00651] In some embodiments of the fourth set of embodiments, the UO2F2 biosensor further comprises a genetic signal amplifier configured to increase an output of a reportable molecular component and/or a U-neutralizing molecular component at a given concentration of bioavailable U , and/or to increase an output of a reportable component at a given concentration of bioavailable F.
[00652] In some embodiments of the fourth set of embodiments, the reportable molecular component and/or the U-neutralizing molecular component of the UO2F2 biosensor, is formed by an assembly of two or more subunits of the reportable molecular component and/or the U- neutralizing molecular component when the U-sensitive genetic circuit operates according to the circuit design in presence of bioavailable U and bioavailable Fluoride.
[00653] In some embodiments of the fourth set of embodiments, the reportable molecular component and/or the U-neutralizing molecular component of the UO2F2 biosensor is post- transcriptionally and/or post-translationally converted when the U-sensitive genetic circuit operates according to the circuit design in presence of bioavailable U and bioavailable Fluoride.
[00654] In some embodiments of the fourth set of embodiments, the UO2F2 biosensor further comprises a U-sensitive genetic molecular component comprising a reporter gene or a U- neutralizing gene operatively connected to a UzcRS -regulated promoter.
[00655] In some embodiments of the fourth set of embodiments, the reportable molecular component of the UO2F2 biosensor comprises at least a first reportable molecular component and a second reportable molecular component, wherein the first reportable molecular component is different from the second reportable molecular component.
[00656] In some embodiments of the fourth set of embodiments, the reportable molecular component of the UO2F2 biosensor is capable of being detected using fluorescence, luminescence, chemiluminescence, colorimetric analysis, radioactivity, or electrical.
[00657] In some embodiments of the fourth set of embodiments, the U-neutralizing molecular component of the UO2F2 biosensor is configured to decrease or eliminate toxicity of U by bioreduction, biomineralization, bioaccumulation, and/or biosorption.
[00658] In some embodiments of the fourth set of embodiments, the bacterial cell of the UO2F2 biosensor is a proteobacterial cell.
[00659] In some embodiments of the fourth set of embodiments, the proteobacterial cell of the UO2F2 biosensor is an alphaproteobacteria, a betaproteobacteria, or a gammaproteobacteria.
[00660] In some embodiments of the fourth set of embodiments, the proteobacterial cell of the UO2F2 biosensor is a Caulobacteridae cell.
[00661] In some embodiments of the fourth set of embodiments, the proteobacterial cell of the UO2F2 biosensor is a Caulobacter crescentus cell, possibly a member of a strain selected from the group consisting of NA1000, CB15, and OR37.
[00662] In some embodiments of the fourth set of embodiments, the U02F2biosensor is configured to detect and/or neutralize bioavailable U present in a target environment at a concentration of 100 nM or greater, between 100 nM and 1 mM, or greater than 1 pM.
[00663] In some embodiments of the fourth set of embodiments, the U02F2biosensor is configured to detect bioavailable F present in a target environment at a concentration of 10 pM, or greater, between 50 pM and 60 pM, greater than 1 nanomolar and between 100 nanomolar and 1 mM.
[00664] In a fifth set of embodiment, a U02F2-biosensor is described comprising
a genetically modified bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 herein also UrpS, and U sensitive transcriptional regulator 1362 herein also UrpR, and/or natively and/or heterologously expressing histidine kinase UzcS, and response regulator UzcS, the genetically modified bacterial cell comprising a U-sensitive genetic circuit in which molecular components are connected one to another in accordance to a circuit design by activating, inhibiting, binding or converting reactions to form a fully connected network of interacting components. at least one molecular component is a U- sensing genetic molecular component in which a U-sensitive promoter is activated or repressed in presence of bioavailable U, the U sensitive promoter comprising a U-sensitive 1362 (UrpR) binding site having a DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO:l), in which
Ni is C or T, preferably C;
N2 is G or A, preferably G;
N3 is T or C, preferably T;
N4 is C;
Ns is A or G, preferably A;
N , is G or C, preferably G; N7 C or G;
Ns is any nucleotide; N9 is any nucleotide;
Niois any nucleotide;
N11 is any nucleotide;
N12 is T or C;
Ni3 is G;
Ni4 is T or C, preferably T;
Nis is C;
N16 IS A or C, preferably A;
Nn is G; and
N18 is C or G,
and in which Ni to N17 are selected independently,
when the cell is capable of natively and/or heterologously expressing histidine kinase 1363 herein also UrpS, and U sensitive transcriptional regulator 1362 herein also UrpR and/or
a UzcR binding site having DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2)
wherein N7-N12 is independently any nucleotide, and in some embodiments any one of N7-N11 can independently be A, when the cell is natively and/or heterologously expressing histidine kinase UzcS, and response regulator UzcS.
[00665] In the U02F2-biosensor of the fifth set of embodiments, in the U-sensitive genetic circuit at least one molecular component is a reportable molecular component and/or a U- neutralizing molecular component, the reportable molecular component and/or the a U- neutralizing molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable U.
[00666] In the U02F2-biosensor of the fifth set of embodiments, the U02F2-biosensor further comprises an F-sensing riboswitch within at least one of a genetic molecular component of an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F; and
an F-sensing reportable genetic molecular component in a configuration wherein the F- sensing reportable genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride; and
[00667] In the UC Fi-biosensor of the fifth set of embodiments,, the F-sensing reportable genetic molecular component and the F-sensing genetic circuit are in a dual output configuration with the U-sensing genetic circuit.
[00668] In some embodiments of the fifth set of embodiments, the U- sensitive genetic circuit and/or the F-sensitive genetic circuit of the UO2F2 biosensor comprise one or more AND gates.
[00669] In some embodiments of the fifth set of embodiments, at least one of the AND gates of the UO2F2 biosensor is an in-series AND gate.
[00670] In some embodiments of the fifth set of embodiments, at least one of the AND gates of the UO2F2 biosensor is an in-parallel AND gate.
[00671] In some embodiments of the fifth set of embodiments, two or more in series AND gates and/or in parallel AND gates of the UO2F2 biosensor are connected by activating, inhibiting, binding, or converting reactions.
[00672] In some embodiments of the fifth set of embodiments, at least one of the AND gates of the UO2F2 biosensor is selected from the group consisting of an HRP AND gate, a bacterial two- hybrid AND gate, a tripartite GFP AND gate, and a FRET sensor AND gate .
[00673] In some embodiments of the fifth set of embodiments, the UO2F2 biosensor further comprises a genetic signal amplifier configured to increase an output of a reportable molecular component and/or a U-neutralizing molecular component at a given concentration of bioavailable U, and/or to increase an output of a reportable component at a given concentration of bioavailable F.
[00674] In some embodiments of the fifth set of embodiments, the reportable molecular component and/or the U-neutralizing molecular component of the U- sensing circuit of the UO2F2 biosensor, is formed by an assembly of two or more subunits when the U-sensitive genetic circuit operates according to the circuit design in presence of bioavailable U.
[00675] In some embodiments of the fifth set of embodiments, the reportable molecular component and/or the U-neutralizing molecular component of the UO2F2 biosensor is post- transcriptionally and/or post-translationally converted when the U-sensitive genetic circuit operates according to the circuit design in presence of bioavailable U.
[00676] In some embodiments of the fifth set of embodiments, the reportable molecular component of the F-sensing circuit of the UO2F2 biosensor, is formed by an assembly of two or more subunits when the F-sensitive genetic circuit operates according to the circuit design in presence of bioavailable F.
[00677] In some embodiments of the fifth set of embodiments, the reportable molecular component of the F-sensitive circuit of the UO2F2 biosensor is post-transcriptionally and/or post- translationally converted when the F-sensitive genetic circuit operates according to the circuit design in presence of bioavailable F.
[00678] In some embodiments of the fifth set of embodiments, in the sequence SEQ ID NO:l: of the UO2F2 biosensor
Ni is C; and/or
N2 is G; and/or
N3 is T; and/or
Ns is A; and/or
Nό is G; and/or
Ni4 is T; and/or
N16IS A. [00679] In some embodiments of the fifth set of embodiments, the U-sensitive transcriptional 1362 binding site of the UO2F2 biosensor has a sequence selected from the group consisting of SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26 and SEQ ID NO:27.
[00680] In some embodiments of the fifth set of embodiments, the U-sensitive promoter of the UO2F2 biosensor further comprises nucleotides N19N20N21, downstream of SEQ ID NO: 1 wherein N19 is any nucleotide; N20 is any nucleotide; and N21 is G (SEQ ID NO: 83).
[00681] In some embodiments of the fifth set of embodiments, Nis of the regulator direct repeat of the UO2F2 biosensor is located about -17 to about -40 upstream of a transcription start site.
[00682] In some embodiments of the fifth set of embodiments, the U-sensitive promoter of the UO2F2 biosensor is P1361 or Pphyt.
[00683] In some embodiments of the fifth set of embodiments, the U02F2biosensor is configured to detect and/or neutralize bioavailable U present in a target environment at a concentration of 100 nM or grater, between 100 nM and 1 mM, and greater than 1 pM.
[00684] In some embodiments of the fifth set of embodiments, the U02F2biosensor is configured to detect bioavailable F present in a target environment at a concentration of 10 pM, or greater, between 50 pM and 60 pM, greater than 1 nanomolar and between 100 nanomolar and 1 mM.
[00685] In some embodiments of the fifth set of embodiments, the U02F2biosensor further comprises a UzcY gene and/or UzcZ gene inserted at a location downstream of a transcription start site of the U-sensitive promoter of the UO2F2 biosensor comprising a UzcR binding site.
[00686] In some embodiments of the fifth set of embodiments, the bacteria of the UO2F2 biosensor are capable of natively expressing endogenous MarR family repressors such as MarRl and/or MarR2 genes and at least one gene of the endogenous MarR family, preferably all genes of the endogenous MarR family, is knocked out. [00687] In some embodiments of the fifth set of embodiments, the bacteria of the UO2F2 biosensor are capable of natively expressing endogenous urtAP genes and at least one gene of the endogenous urtAP, preferably all genes of the endogenous urtAP, is knocked out.
[00688] In some embodiments of the fifth set of embodiments, the U-sensing circuit of the UO2F2 biosensor comprises a UrpR U-sensing genetic molecular component comprising a U- sensitive 1362 (UrpR) binding site and UzcRS U sensing genetic molecular component comprising a U-sensitive UzcR binding site, the UrpR U-sensing genetic molecular component and the UzcR U-sensing genetic molecular component are operatively connected to a same or a different reportable molecular component in the circuit.
[00689] In some embodiments of any one of the first to the fifth set of embodiments, the Fluoride sensing riboswitch of the UO2F2 biosensor comprises a crcB motif.
[00690] In some embodiments of any one of the first to the fifth set of embodiments, the Fluoride sensing riboswitch of the UO2F2 biosensor comprises an ericFmotif.
[00691] In some embodiments of any one of the first to the fifth set of embodiments, the Fluoride sensing riboswitch comprises any one of the riboswitches and/or sequences in
Appendix I and Figure 53.
[00692] In a sixth set of embodiments, a method is described to provide a U biosensor comprising:
genetically engineering a bacterial cell capable of natively and/or heterologously expressing histidine kinase 1363 (UrpS) and/or histidine kinase UzcS, and U sensitive response regulator 1362 (UrpR) and/or U sensitive response regulator UzcR in combination with a heterologous F-sensing riboswitch, the genetically engineering performed by introducing into the cell
one or more U-sensing genetic molecular components configured to report and/or neutralize U herein described,
an F-sensitive riboswitch within the one or more U-sensing genetic molecular components configured to report U, an F-sensitive riboswitch within an F-sensing reportable genetic molecular component;
one or more genetic molecular components of an F sensitive genetic circuit described herein, and/or
one or more genetic molecular components of the U-sensitive F-sensitive genetic circuits described herein
the genetically engineering performed to obtain a UO2F2 biosensor according to any one of the embodiments, of any one of the first to the fifth set of embodiments.
[00693] In some embodiments of the sixth set of embodiments, the method further comprising: operatively connecting one or more of the U biosensors to an electronic signal transducer adapted to convert a U biosensor reportable molecular component output into an electronic output.
[00694] In a seventh set of embodiments, a UO2F2- sensing gene cassette is described comprising one or more genetic molecular components described in any one of the first to the fifth set of embodiments, and an F-sensitive riboswitch within one or more genetic molecular components in a configuration wherein one or more genetic molecular components are transcribed in presence of an effective amount of bioavailable fluoride.
[00695] In some embodiments of the seventh set of embodiments, the UO2F2- sensing gene cassette is an expression cassette.
[00696] In some embodiments of the seventh set of embodiments, the UO2F2- sensing gene cassette is comprised within a vector.
[00697] In an eighth set of embodiments, a vector is described comprising:
a polynucleotide encoding one or more genetic molecular components described in any one of the first to the fifth set of embodiments, or a U02F2-sensing gene cassette of any one of the embodiments of seventh set of embodiments,
wherein the vector is configured to introduce the one or more genetic molecular components into a proteobacterial cell. [00698] In some embodiments of the eighth set of embodiments, the one or more genetic molecular components form a U-sensitive genetic circuit into the proteobacterial cell.
[00699] In a ninth set of embodiments, a system is described comprising:
a plurality of bacterial cells; and
one or more of the vectors of any one of the embodiments of the eighth set of embodiments.
[00700] In a tenth set of embodiments, a composition is described comprising:
one or more of the UO2F2- biosensors of any one of the embodiments of the first to the fifth set of embodiments and/or one or more of the vector of the eight set of embodiments together with a suitable vehicle.
[00701] In an eleventh set of embodiments, a U02F2-sensing system is described comprising: one or more of the UO2F2- biosensors of any one of the embodiments of any one of the first to the fifth set of embodiments, operatively connected to an electronic signal transducer adapted to convert a UO2F2- biosensor reportable molecular component output into an electronic output.
[00702] In a twelfth set of embodiments, a method is described of detecting and reporting and/or neutralizing bioavailable U comprising:
contacting one or more of the U biosensors of any one of the embodiments of any one of the first to the fifth set of embodiments, or the UO2F2— sensing system of the eleventh set of embodiments, with a target environment comprising one or more target ranges of U concentration, the contacting performed for a time and under conditions to detect and report and/or neutralize bioavailable U in the target environment.
[00703] In some embodiments of the twelfth set of embodiments, the U02F2biosensor is configured to detect and/or neutralize bioavailable U present in a target environment at a concentration of 100 nM or greater, between 100 nM and 1 mM, or greater than 1 pM.
[00704] In some embodiments of the twelfth set of embodiments, the U02F2biosensor is configured to detect bioavailable F present in a target environment at a concentration of 10 mM, or greater, between 50 mM and 60 pM, greater than 1 nanomolar and between 100 nanomolar and 1 mM.
[00705] In a thirteenth set of embodiments a UC Fi-sensing genetic reportable component is described comprising:
one or more U sensitive promoters comprising a U- sensitive 1362 binding site and/or a U sensitive UzcR binding site,
together with
a U- sensing reportable molecular component,
[00706] In the UC Fi-sensing genetic reportable component of the thirteenth set of embodiments, at least one of the one or more U-sensitive promoters and the U-sensing reportable molecular component, comprises an F-sensing riboswitch in a single output configuration wherein the U-sensing reportable genetic molecular component, is transcribed in presence of an effective amount of bioavailable fluoride,
[00707] In the UC Fi-sensing genetic reportable component of the thirteenth set of embodiments, the one or more U sensitive promoters and the U-sensing reportable molecular component are in a configuration wherein the one or more U sensitive promoters directly initiate expression of the U-sensing reportable molecular component in presence of bioavailable U and bioavailable Fluoride.
[00708] In the UC Fi-sensing genetic reportable component of the thirteenth set of embodiments, the U sensitive promoter comprises a U-sensitive 1362 binding site having a
DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO:l), wherein
Ni is C or T, preferably C;
N2 is G or A, preferably G; N3 is T or C, preferably T;
N4 is C;
Ns is A or G, preferably A;
Nό is G or C, preferably G;
N7 is C or G;
Ns is any nucleotide;
N9 is any nucleotide;
Niois any nucleotide;
N11 is any nucleotide;
N12 is T or C;
Ni3 is G;
Ni4 is T or C, preferably T;
Nis is C;
N16 IS A or C, preferably A;
Nn is G; and
Nis is C or G,
wherein Ni to Nn selected independently,
and optionally further comprises an UzcR binding site having a DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide
[00709] In the U02F2-sensing genetic reportable component of the thirteenth set of embodiments, the UzcR binding site inserted at a location downstream of a transcription start site of the U- sensitive promoter.
[00710] In some embodiments of the thirteenth set of embodiments, the U02F2sensing genetic reportable component further comprises a UzcY gene and/or UzcZ gene inserted at a location downstream of a transcription start site of the U-sensitive promoter.
[00711] In some embodiments of the thirteenth set of embodiments, in the sequence SEQ ID NO: l of the UChFisensing genetic reportable component:
Ni is C; and/or
N2 is G; and/or
N3 is T; and/or
Ns is A; and/or
Nό is G; and/or
Ni4 is T; and/or
N16 IS A.
[00712] In some embodiments of the thirteenth set of embodiments, the U-sensitive transcriptional 1362 binding site of the U02F2sensing genetic reportable component, has a sequence selected from the group consisting of SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26 and SEQ ID NO:27.
[00713] In some embodiments of the thirteenth set of embodiments, the U-sensitive promoter of the U02F2sensing genetic reportable component, further comprises nucleotides N19N20N21, downstream of SEQ ID NO: 1 wherein N19 is any nucleotide; N20 is any nucleotide; and N21 is G (SEQ ID NO: 83)
[00714] In some embodiments of the thirteenth set of embodiments, Nis of the regulator direct repeat of the U02F2sensing genetic reportable component, is located about -17 to about -40 upstream of a transcription start site.
[00715] In some embodiments of the thirteenth set of embodiments, the U-sensitive promoter of the UC Fisensing genetic reportable component, is P i % i or Pphyt.
[00716] In some embodiments of the thirteenth set of embodiments, the U- sensitive promoter of the UC Fisensing genetic reportable component comprises an UzcR binding site having a DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide.
[00717] In some embodiments of the thirteenth set of embodiments, the UO2F2- sensing genetic reportable component further comprises a UzcY gene and/or UzcZ gene inserted at a location downstream of a transcription start site of the U- sensitive promoter of the U02F2sensing genetic reportable component.
[00718] In some embodiments of the thirteenth set of embodiments, the Fluoride sensing riboswitch of the U02F2sensing genetic reportable component comprises a crcb motif
[00719] In some embodiments of the thirteenth set of embodiments, the Fluoride sensing riboswitch of the U02F2sensing genetic reportable component comprises an ericFmotif.
[00720] In some embodiments of the thirteenth set of embodiments, the Fluoride sensing riboswitch of the U02F2sensing genetic reportable component comprises any one of the riboswitches and/or sequences in Appendices I to IV.
[00721] In some embodiments of the thirteenth set of embodiments, the UO2F2- sensing genetic reportable component is configured to detect and/or neutralize bioavailable U present in a target environment at a concentration of 100 nM or grater, between 100 nM and 1 mM, and greater than 1 pM.
[00722] In some embodiments of the thirteenth set of embodiments, the U02F2-sensing genetic reportable component is configured to detect bioavailable F present in a target environment at a concentration of 10 pM, or greater, between 50 pM and 60 pM, greater than 1 nanomolar and between 100 nanomolar and 1 mM.
[00723] In a fourteenth set of embodiments, a U02F2-sensitive genetic circuit is described comprising at least one UO2F2 sensing genetic molecular component of any one of claims 88 to 102 and wherein at least one molecular component is a reportable molecular component, and/or a U-neutralizing molecular component, the reportable molecular component, and/or the U- neutralizing molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable U and bioavailable Fluoride.
[00724] In some embodiments of the fourteenth set of embodiments, the UO2F2- sensitive genetic circuit is configured to detect and/or neutralize bioavailable U present in a target environment at a concentration of 100 nM or grater, between 100 nM and 1 mM, and greater than 1 pM.
[00725] In some embodiments of the fourteenth set of embodiments, the UO2F2- sensitive genetic circuit is configured to detect bioavailable F present in a target environment at a concentration of 10 pM, or greater, between 50 pM and 60 pM, greater than 1 nanomolar and between 100 nanomolar and 1 mM.
[00726] In a fifteenth set of embodiments, a U02F2-sensing system is described comprising one or more vectors of any one of the embodiments of the eighth set of embodiments and/or a plurality of bacterial cells natively and/or heterologously expressing histidine kinase 1363 and response regulator 1362 and/or histidine kinase UzcS and response regulator UzcR.
[00727] In some embodiments of the fifteenth set of embodiments, the UO2F2- sensing system is configured to detect and/or neutralize bioavailable U present in a target environment at a concentration of 100 nM or grater, between 100 nM and 1 pM, and greater than 1 pM.
[00728] In some embodiments of the fifteenth set of embodiments, the UO2F2- sensing system is configured to detect bioavailable F present in a target environment at a concentration of 10 pM, or greater, between 50 pM and 60 pM, greater than 1 nanomolar and between 100 nanomolar and 1 mM.
[00729] The examples set forth above are provided to give those of ordinary skill in the art a complete disclosure and description of how to make and use the embodiments of the U biosensors, and related U-sensitive genetic molecular components, genetic circuits, compositions, methods and systems of the disclosure, and are not intended to limit the scope of what the inventors regard as their disclosure. Those skilled in the art will recognize how to adapt the features of the exemplified U biosensors, and related U-sensitive genetic molecular components, genetic circuits, compositions, methods and systems herein disclosed to additional U biosensors, and related U-sensitive genetic molecular components, genetic circuits, compositions, methods and systems, and related genetic molecular components, sets of polynucleotides, polypeptides, proteins, and/or metabolites, in according to various embodiments and scope of the claims.
[00730] All patents and publications mentioned in the specification are indicative of the levels of skill of those skilled in the art to which the disclosure pertains.
[00731] The entire disclosure of each document cited (including patents, patent applications, journal articles, abstracts, laboratory manuals, books, or other disclosures) in the Background, Summary, Detailed Description, and Examples is hereby incorporated herein by reference. All references cited in this disclosure are incorporated by reference to the same extent as if each reference had been incorporated by reference in its entirety individually. However, if any inconsistency arises between a cited reference and the present disclosure, the present disclosure takes precedence. Further, the computer readable form of the sequence listing of the ASCII text file IL13081-2-P2326-PCT-Seq-Listing_ST25 filed on February 4,, 2020 concurrently with the present application, is incorporated herein by reference in its entirety and forms integral part of the present description.
[00732] The terms and expressions which have been employed herein are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the disclosure claimed. Thus, it should be understood that although the disclosure has been specifically disclosed by embodiments, exemplary embodiments and optional features, modification and variation of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this disclosure as defined by the appended claims.
[00733] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. The term“plurality” includes two or more referents unless the content clearly dictates otherwise. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosure pertains.
[00734] When a Markush group or other grouping is used herein, all individual members of the group and all combinations and possible sub-combinations of the group are intended to be individually included in the disclosure. Every combination of components or materials described or exemplified herein can be used to practice the disclosure, unless otherwise stated. One of ordinary skill in the art will appreciate that methods, system elements, and materials other than those specifically exemplified may be employed in the practice of the disclosure without resort to undue experimentation. All art-known functional equivalents, of any such methods, device elements, and materials are intended to be included in this disclosure. Whenever a range is given in the specification, for example, a temperature range, a frequency range, a time range, or a composition range, all intermediate ranges and all subranges, as well as, all individual values included in the ranges given are intended to be included in the disclosure. Any one or more individual members of a range or group disclosed herein may be excluded from a claim of this disclosure. The disclosure illustratively described herein suitably may be practiced in the absence of any element or elements, limitation or limitations which is not specifically disclosed herein.
[00735] A number of embodiments of the disclosure have been described. The specific embodiments provided herein are examples of useful embodiments of the disclosure and it will be apparent to one skilled in the art that the disclosure can be carried out using a large number of variations of the genetic circuits, genetic molecular components, and methods steps set forth in the present description. As will be obvious to one of skill in the art, methods and systems useful for the present methods and systems may include a large number of optional composition and processing elements and steps.
[00736] In particular, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.
REFERENCES
1. Miller, J., Experiments in molecular genetics. Cold Spring Laboratory Press. 1972, Cold Spring Harbor, NY.
2. Ku, H., Notes on the use of propagation of error formulas. Journal of Research of the National Bureau of Standards, 1966. 70(4).
3. Park, D.M., et al., Identification of a U/Zn/Cu responsive global regulatory two- component system in Caulobacter crescentus. Mol Microbiol, 2017.
4. Fisher, J.A., J. Smit, and N. Agabian, Transcriptional analysis of the major surface array gene of Caulobacter crescentus. J Bacteriol, 1988. 170(10): p. 4706-13.
5. Cabantous, S., et al., A new protein-protein interaction sensor based on tripartite split- GFP association. Scientific reports, 2013. 3: p. 2854.
6. Meisenzahl, A.C., L. Shapiro, and U. Jenal, Isolation and characterization of a xylose- dependent promoter from Caulobacter crescentus. J Bacteriol, 1997. 179(3): p. 592-600.
7. Newsome, L., K. Morris, and J.R. Lloyd, The biogeochemistry and bioremediation of
uranium and other priority radionuclides. Chemical Geology, 2014. 363: p. 164-184.
8. Langmuir, D., Agueous environmental geochemistry. 1997.
9. Ewing, R.C., Environmental impact of the nuclear fuel cycle. Geological Society, London, Special Publications, 2004. 236(1): p. 7-23.
10. Bernier-Latmani, R., et al., Non-uraninite products of microbial U (VI) reduction.
Environmental science & technology, 2010. 44(24): p. 9456-9462.
11. Brutinel, E.D. and J.A. Gralnick, Shuttling happens: soluble flavin mediators of
extracellular electron transfer in Shewanella. Applied microbiology and biotechnology, 2012. 93(1): p. 41-48.
12. Lovley, D.R. and E.J. Phillips, Microbial reduction of uranium. Nature, 1991. 350(6317): p.
413.
13. Williams, K.H., et al., Bioremediation of uranium-contaminated groundwater: a systems approach to subsurface biogeochemistry. Current opinion in biotechnology, 2013. 24(3): p. 489-497.
14. Beazley, M.J., et al., The effect of pH and natural microbial phosphatase activity on the speciation of uranium in subsurface soils. Geochimica et Cosmochimica Acta, 2011. 75(19): p. 5648-5663.
15. Macaskie, L.E., et al., Enzymically mediated bioprecipitation of uranium by a Citrobacter sp.: a concerted role for exocellular lipopolysaccharide and associated phosphatase in biomineral formation. Microbiology, 2000. 146(8): p. 1855-1867.
16. Macaskie, L.E., et al., Uranium Bioaccumulation by a Citrobacter sp. as a Result of
Enzymically Mediated Growth of Polycrystalline HUO_2 PO_4. Science, 1992: p. 782- 784.
17. Beveridge, T. and R. Murray, Sites of metal deposition in the cell wall of Bacillus subtilis.
Journal of bacteriology, 1980. 141(2): p. 876-887. Gadd, G.M., Biosorption: critical review of scientific rationale, environmental importance and significance for pollution treatment. Journal of Chemical Technology and
Biotechnology, 2009. 84(1): p. 13-28.
Choudhary, S. and P. Sar, Uranium biomineralization by a metal resistant Pseudomonas aeruginosa strain isolated from contaminated mine waste. Journal of hazardous materials, 2011. 186(1): p. 336-343.
Ku, H.H., Notes on the use of propagation of error formulas. Journal of Research of the National Bureau of Standards— C. Engineering and Instrumentation, 1966. 70C(4): p. 263-273.
Bailey, T.L. and C. Elkan, Fitting a mixture model by expectation maximization to discover motifs in biopolymers. Proc Int Conf Intell Syst Mol Biol, 1994. 2: p. 28-36.
Zhou, B., et al., The global regulatory architecture of transcription during the
Caulobacter cell cycle. PLoS Genet, 2015. 11(1): p. el004831.
McGrath, P.T., et al., High-throughput identification of transcription start sites, conserved promoter motifs and predicted regulons. Nat Biotechnol, 2007. 25(5): p. 584- 92.
Markich, S.J., Uranium speciation and bioavailability in aguatic systems: an overview.
The Scientific World Journal, 2002. 2: p. 707-729.
Focazio, M.J., et al., The Chemical Quality of Self-Supplied Domestic Well Water in the United States. Groundwater Monitoring & Remediation, 2006. 26(3): p. 92-104.
Hoover, J., et al., Elevated Arsenic and Uranium Concentrations in Unregulated Water Sources on the Navajo Nation, USA. Exposure and Health, 2016: p. 1-12.
Kemp, R.S., Environmental Detection of Clandestine Nuclear Weapon Programs. Annual Review of Earth and Planetary Sciences, 2016. 44(1): p. 17-35.
Kemp, R.S., Initial Analysis of the Detectability of U02F2 Aerosols Produced by UF6 Released from Uranium Conversion Plants. Science & Global Security, 2008. 16(3): p. 115-125.
Bostick, W., et al., Sampling and characterization of aerosols formed in the atmospheric hydrolysis of UF/sub 6. 1983.
Wogman, N.A., Prospects for the introduction of wide area monitoring using
environmental sampling for proliferation detection. Journal of Radioanalytical and Nuclear Chemistry, 2013. 296(2): p. 1071-1077.
Kips, R.S. and M.J. Kristo, Investigation of chemical changes in uranium oxy fluoride particles using secondary ion mass spectrometry. Journal of Radioanalytical and Nuclear Chemistry, 2009. 282(3): p. 1031.
Myers, W.L., A literature review on the chemical and physical properties of uranyl fluoride (UO sub 2 F sub 2 ). 1990: United States.
Langmuir, D., Uranium solution-mineral eguilibria at low temperatures with applications to sedimentary ore deposits. Geochimica et Cosmochimica Acta, 1978. 42(6, Part A): p. 547-569.
Markich, S.J., Uranium speciation and bioavailability in aguatic systems: an overview. ScientificWorldJournal, 2002. 2: p. 707-29.
Sheng, L. and J.B. Fein, Uranium adsorption by Shewanella oneidensis MR-1 as a function of dissolved inorganic carbon concentration. Chemical Geology, 2013. 358: p. 15-22. Bencheikh-Latmani, R. and J.O. Leckie, Association of uranyl with the cell wall of
Pseudomonas fluorescens inhibits metabolism. Geochimica et Cosmochimica Acta, 2003. 67(21): p. 4057-4066.
Nolan, J. and K.A. Weber, Natural Uranium Contamination in Major U.S. Aquifers Linked to Nitrate. Environmental Science & Technology Letters, 2015. 2(8): p. 215-220.
Ferla, M.P., et al., New rRNA gene-based phytogenies of the Alphaproteobacteria provide perspective on major groups, mitochondrial ancestry and phylogenetic instability. PLoS One, 2013. 8(12): p. e83383.
Slonczewski JL, F.J., Microbiology: An Evolving Science W. W. Norton & Company, 2014: p. 742-3.
Dworkin M, F.S., Rosenberg E, Schleifer KH, Stackebrandt E, The Prokaryotes:
Proteobacteria: Alpha and Beta Subclasses. 2006. 5: p. 15-18.
Stock, A.M., V.L. Robinson, and P.N. Goudreau, Two-component signal transduction. Annual review of biochemistry, 2000. 69(1): p. 183-215.
Mascher, T., J.D. Helmann, and G. Unden, Stimulus perception in bacterial signal- transducing histidine kinases. Microbiology and Molecular Biology Reviews, 2006. 70(4): p. 910-938.
Capra, E.J. and M.T. Laub, Evolution of two-component signal transduction systems. Annual review of microbiology, 2012. 66: p. 325-347.
Sanders, D., et al., Phosphorylation site of NtrC, a protein phosphatase whose covalent intermediate activates transcription. Journal of bacteriology, 1992. 174(15): p. 5117- 5122.
Sanders, D.A., et al., Identification of the site of phosphorylation of the chemotaxis response regulator protein, CheY. Journal of Biological Chemistry, 1989. 264(36): p. 21770-21778.
Datsenko, K.A. and B.L. Wanner, One-step inactivation of chromosomal genes in
Escherichia coli K-12 using PCR products. Proc Natl Acad Sci U S A, 2000. 97(12): p. 6640- 5.
Sharma, C.M., et al., The primary transcriptome of the major human pathogen
Helicobacter pylori. Nature, 2010. 464(7286): p. 250.
Carey, M.F., C.L. Peterson, and S.T. Smale, The primer extension assay. Cold Spring Harbor Protocols, 2013. 2013(2): p. pdb. prot071902.
Wade, J.T., Where to begin? Mapping transcription start sites genome-wide in
Escherichia coli. Journal of bacteriology, 2015. 197(1): p. 4-6.
Poindexter, J.S., The caulobacters: ubiquitous unusual bacteria. Microbiological reviews, 1981. 45(1): p. 123.
Hu, P., et al., Whole-genome transcriptional analysis of heavy metal stresses in
Caulobacter crescentus. Journal of bacteriology, 2005. 187(24): p. 8437-8449.
Park, D.M. and Y. Jiao, Modulation of medium pH by Caulobacter crescentus facilitates recovery from uranium-induced growth arrest. Applied and environmental microbiology, 2014. 80(18): p. 5680-5688.
Bollmann, A., et al., Isolation and physiology of bacteria from contaminated subsurface sediments. Applied and environmental microbiology, 2010. 76(22): p. 7413-7419. Yung, M.C. and Y. Jiao, Biomineralization of uranium by PhoY phosphatase activity aids cell survival in Caulobacter crescentus. Applied and environmental microbiology, 2014. 80(16): p. 4795-4804.
Hillson, N.J., et al., Caulobacter crescentus as a whole-cell uranium biosensor. Applied and environmental microbiology, 2007. 73(23): p. 7615-7621.
Cormack, B.P., R.H. Valdivia, and S. Falkow, FACS-optimized mutants of the green fluorescent protein (GFP). Gene, 1996. 173(1): p. 33-38.
Lovley, D.R. and E. Phillips, Reduction of uranium by Desulfovibrio desulfuricans. Applied and environmental microbiology, 1992. 58(3): p. 850-856.
Francis, A.J., et al., XPS and XANES studies of uranium reduction by Clostridium sp.
Environmental science & technology, 1994. 28(4): p. 636-639.
Shelobolina, E.S., et al., Isolation, characterization, and U (Vl)-reducing potential of a facultatively anaerobic, acid-resistant Bacterium from Low-pH, nitrate-and U (VI)- contaminated subsurface sediment and description of Salmonella subterranea sp. nov. Applied and Environmental Microbiology, 2004. 70(5): p. 2959-2965.
Wu, Q., R.A. Sanford, and F.E. Loffler, Uranium (VI) reduction by Anaeromyxobacter dehalogenans strain 2CP-C. Applied and environmental microbiology, 2006. 72(5): p. 3608-3614.
Begg, J.D., et al., Bioreduction behavior of U (VI) sorbed to sediments. Geomicrobiology Journal, 2011. 28(2): p. 160-171.
Istok, J., et al., In situ bioreduction of technetium and uranium in a nitrate-contaminated aguifer. Environmental Science & Technology, 2004. 38(2): p. 468-475.
Law, G.T., et al., Uranium redox cycling in sediment and biomineral systems.
Geomicrobiology Journal, 2011. 28(5-6): p. 497-506.
Wilkins, M., et al., The influence of microbial redox cycling on radionuclide mobility in the subsurface at a low-level radioactive waste storage site. Geobiology, 2007. 5(3): p. 293- 301.
Williams, K.H., et al., Acetate availability and its influence on sustainable bioremediation of uranium-contaminated groundwater. Geomicrobiology Journal, 2011. 28(5-6): p. 519- 539.
Wu, W.-M., et al., In situ bioreduction of uranium (VI) to submicromolar levels and reoxidation by dissolved oxygen. Environmental Science & Technology, 2007. 41(16): p. 5716-5723.
Lovley, D.R., et al., Geobacter metallireducens gen. nov. sp. nov., a microorganism capable of coupling the complete oxidation of organic compounds to the reduction of iron and other metals. Archives of microbiology, 1993. 159(4): p. 336-344.
Richter, K., M. Schicklberger, and J. Gescher, Dissimilatory reduction of extracellular electron acceptors in anaerobic respiration. Applied and environmental microbiology, 2012. 78(4): p. 913-921.
Lovley, D.R., D.E. Holmes, and K.P. Nevin, Dissimilatory fe (Hi) and mn (iv) reduction. Advances in microbial physiology, 2004. 49: p. 219-286.
Marsili, E., et al., Shewanella secretes flavins that mediate extracellular electron transfer. Proceedings of the National Academy of Sciences, 2008. 105(10): p. 3968-3973. Suzuki, Y., et al.. Flavin mononucleotide mediated electron pathway for microbial U (VI) reduction. Physical Chemistry Chemical Physics, 2010. 12(34): p. 10081-10087.
Von Canstein, H., et al., Secretion of flavins by Shewanella species and their role in extracellular electron transfer. Applied and environmental microbiology, 2008. 74(3): p. 615-623.
Anderson, R.T., et al., Stimulating the in situ activity of Geobacter species to remove uranium from the groundwater of a uranium-contaminated aguifer. Applied and environmental microbiology, 2003. 69(10): p. 5884-5891.
Senko, J.M., et al., The effect of U (VI) bioreduction kinetics on subseguent reoxidation of biogenic U (IV). Geochimica et Cosmochimica Acta, 2007. 71(19): p. 4644-4654.
Martinez, R.J., et al., Aerobic uranium (VI) bioprecipitation by metal-resistant bacteria isolated from radionuclide-and metal-contaminated subsurface soils. Environmental microbiology, 2007. 9(12): p. 3122-3133.
Basnakova, G., et al., The use of Escherichia coli bearing a phoN gene for the removal of uranium and nickel from agueous flows. Applied microbiology and biotechnology, 1998. 50(2): p. 266-272.
Powers, L.G., et al., Introduction of a plasmid-encoded phoA gene for constitutive overproduction of alkaline phosphatase in three subsurface Pseudomonas isolates. FEMS microbiology ecology, 2002. 41(2): p. 115-123.
Thomas, R.A. and L. Macaskie, Biodegradation of tributyl phosphate by naturally occurring microbial isolates and coupling to the removal of uranium from agueous solution. Environmental science & technology, 1996. 30(7): p. 2371-2375.
Siuda, W. and R. Chrost, Utilization of selected dissolved organic phosphorus compounds by bacteria in lake water under non-limiting orthophosphate conditions. Polish Journal of Environmental Studies, 2001. 10(6): p. 475-484.
Lim, B.L., et al., Distribution and diversity of phytate-mineralizing bacteria. The ISME journal, 2007. 1(4): p. 321.
Ko, W.-h. and F.K. Hora, Production of phospholipases by soil microorganisms. Soil Science, 1970. 110(5): p. 355-358.
Kazy, S.K., S.F. D'Souza, and P. Sar, Uranium and thorium seguestration by a
Pseudomonas sp.: mechanism and chemical characterization. J Hazard Mater, 2009. 163(1): p. 65-72.
Vanengelen, M.R., et al., UO(2) 2+ speciation determines uranium toxicity and bioaccumulation in an environmental Pseudomonas sp. isolate. Environ Toxicol Chem, 2010. 29(4): p. 763-9.
Choudhary, S. and P. Sar, Uranium biomineralization by a metal resistant Pseudomonas aeruginosa strain isolated from contaminated mine waste. J Hazard Mater, 2011. 186(1): p. 336-43.
Renninger, N., et al., Uranyl precipitation by Pseudomonas aeruginosa via controlled polyphosphate metabolism. Appl Environ Microbiol, 2004. 70(12): p. 7404-12.
Zhou, L., et al., A protein engineered to bind uranyl selectively and with femtomolar affinity. Nat Chem, 2014. 6(3): p. 236-41.
Pardoux, R., et al., Modulating uranium binding affinity in engineered calmodulin EF- hand peptides: effect of phosphorylation. PLoS One, 2012. 7(8): p. e41922. Nomellini, J.F., et al., S-layer-mediated display of the immunoglobulin G-binding domain of streptococcal protein G on the surface of Caulobacter crescentus: development of an immunoactive reagent. Appl Environ Microbiol, 2007. 73(10): p. 3245-53.
Park, D.M., et al., Bioadsorption of Rare Earth Elements through Cell Surface Display of Lanthanide Binding Tags. Environ Sci Technol, 2016. 50(5): p. 2735-42.
Choppin, G., J. Liljenzin, and J. Rydberg, Behavior of Radionuclides in the Environment. Radiochemistry and Nuclear Chemistry, 1995.
Hsi, C.-k.D. and D. Langmuir, Adsorption of uranyl onto ferric oxyhydroxides: application of the surface complexation site-binding model. Geochimica et Cosmochimica Acta,
1985. 49(9): p. 1931-1941.
Koch-Steindl, H. and G. Prohl, Considerations on the behaviour of long-lived
radionuclides in the soil. Radiation and environmental biophysics, 2001. 40(2): p. 93-104. Davis, J.A., et al., Approaches to surface complexation modeling of uranium (VI) adsorption on aguifer sediments. Geochimica et Cosmochimica Acta, 2004. 68(18): p. 3621-3641.
Pabalan, R.T., et al., Uranium (VI) sorption onto selected mineral surfaces: Key geochemical parameters. 1996, American Chemical Society, Washington, DC (United States).
Siegel, M. and C. Bryan, Radioactive Contamination. Environmental Geochemistry, 2005. 9: p. 205.
Bargar, J.R., et al., Uranium redox transition pathways in acetate-amended sediments. Proceedings of the National Academy of Sciences, 2013. 110(12): p. 4506-4511.
Utturkar, S.M., et al., Draft genome seguence for Caulobacter sp. strain OR37, a bacterium tolerant to heavy metals. Genome announcements, 2013. 1(3): p. e00322-13. Garst, A.D., A.L. Edwards, and R.T. Batey, Riboswitches: structures and mechanisms. Cold Spring Harb Perspect Biol, 2011. 3(6).
Breaker, R.R., Riboswitches and the RNA world. Cold Spring Harb Perspect Biol, 2012. 4(2).
Hallberg, Z.F., et al., Engineering and In Vivo Applications of Riboswitches. Annu Rev Biochem, 2017. 86: p. 515-539.
! ! ! INVALID CITATION ! ! ! [101]
Baker, J.L., et al., Widespread genetic switches and toxicity resistance proteins for fluoride. Science, 2012. 335(6065): p. 233-235.
Breaker, R.R., New insight on the response of bacteria to fluoride. Caries Res, 2012.
46(1): p. 78-81.
Ren, A., K.R. Rajashankar, and D.J. Patel, Fluoride ion encapsulation by Mg2+ ions and phosphates in a fluoride riboswitch. Nature, 2012. 486(7401): p. 85-9.
Stockbridge, R.B., et al., Fluoride resistance and transport by riboswitch-controlled CLC antiporters. Proc Natl Acad Sci U S A, 2012. 109(38): p. 15289-94.
Zhao, B., et al., An excited state underlies gene regulation of a transcriptional riboswitch. Nat Chem Biol, 2017. 13(9): p. 968-974.
Park, D.M. and M.J. Taffet, Combinatorial Sensor Design in Caulobacter crescentus for Selective Environmental Uranium Detection. ACS Synth Biol, 2019. 8(4): p. 807-817. Weinberg, Z., et al.. Comparative genomics reveals 104 candidate structured RNAsfrom bacteria, archaea, and their metagenomes. Genome Biol, 2010. 11(3): p. R31.
Finn, R.D., et al., The Pfam protein families database: towards a more sustainable future. Nucleic Acids Res, 2016. 44(D1): p. D279-85.
Brewster, R.C., et al., The transcription factor titration effect dictates level of gene expression. Cell, 2014. 156(6): p. 1312-23.
Shin, J. and V. Noireaux, An E. coli cell-free expression toolbox: application to synthetic gene circuits and artificial cells. ACS Synth Biol, 2012. 1(1): p. 29-41.
Procaccini, A., et al., Dissecting the specificity of protein-protein interaction in bacterial two-component signaling: orphans and crosstalks. PloS one, 2011. 6(5): p. el9729.
Folliard, T., et al., Ribo-attenuators: novel elements for reliable and modular riboswitch engineering. Sci Rep, 2017. 7(1): p. 4599.
Buttner, D. and U. Bonas, Who comes first? How plant pathogenic bacteria orchestrate type III secretion. Current opinion in microbiology, 2006. 9(2): p. 193-200.
Hutcheson, S.W., et al., Enhancer-Binding Proteins HrpR and HrpS Interact To Regulate hrp-Encoded Type III Protein Secretion inPseudomonas syringae Strains. Journal of Bacteriology, 2001. 183(19): p. 5589-5598.
Jin, Q., et al., Type III protein secretion in Pseudomonas syringae. Microbes and
Infection, 2003. 5(4): p. 301-310.
Dove, S.L. and A. Hochschild, Conversion of the w subunit of Escherichia coli RNA polymerase into a transcriptional activator or an activation target. Genes &
development, 1998. 12(5): p. 745-754.
Blondel, A. and H. Bedouelle, Engineering the guaternary structure of an exported protein with a leucine zipper. Protein Eng, 1991. 4(4): p. 457-61.
Cheng, P.-C, The contrast formation in optical microscopy, in Handbook of Biological Confocal Microscopy. 2006, Springer p. 162-206.
Helms, V., Principles of computational cell biology. 2008: John Wiley & Sons.
Zheng, J., Spectroscopy-based guantitative fluorescence resonance energy transfer analysis. Ion channels: methods and protocols, 2006: p. 65-77.
Periasamy, A., Fluorescence resonance energy transfer microscopy: a mini review.
Journal of biomedical optics, 2001. 6(3): p. 287-291.
Nguyen, A.W. and P.S. Daugherty, Evolutionary optimization of fluorescent proteins for intracellular FRET. Nature biotechnology, 2005. 23(3): p. 355.
Buchler, N.E., U. Gerland, and T. Hwa, On schemes of combinatorial transcription logic. Proceedings of the National Academy of Sciences, 2003. 100(9): p. 5136-5141.
Silva-Rocha, R. and V. de Lorenzo, Mining logic gates in prokaryotic transcriptional regulation networks. FEBS letters, 2008. 582(8): p. 1237-1244.
Wang, B., et al., Engineering modular and orthogonal genetic logic gates for robust digital-like synthetic biology. Nature communications, 2011. 2: p. 508.
Park, M., S.L. Tsai, and W. Chen, Microbial biosensors: engineered microorganisms as the sensing machinery. Sensors (Basel), 2013. 13(5): p. 5777-95.
Yagi, K., Applications of whole-cell bacterial sensors in biotechnology and environmental science. Appl Microbiol Biotechnol, 2007. 73(6): p. 1251-8. Dai, C. and S. Choi, Technology and Applications of Microbial Biosensor. Open Journal of Applied Biosensor, 2013. 2(3).
Bereza-Malcolm, L.T., G. Mann, and A.E. Franks, Environmental Sensing of Heavy Metals Through Whole Cell Microbial Biosensors: A Synthetic Biology Approach. ACS Synth Biol, 2014.
Hwang, I.Y., et al., Engineered probiotic Escherichia coli can eliminate and prevent Pseudomonas aeruginosa gut infection in animal models. Nature Communications,
2017. 8: p. 15028.
King, J.M., et al., Rapid, sensitive bioluminescent reporter technology for naphthalene exposure and biodegradation. Science, 1990. 249(4970): p. 778-81.
Belkin, S., et al., Remote detection of buried landmines using a bacterial sensor. Nat Biotechnol, 2017. 35(4): p. 308-310.
Kabessa, Y., et al., Standoff detection of explosives and buried landmines using fluorescent bacterial sensor cells. Biosensors and Bioelectronics, 2016. 79: p. 784-788. Wang, B., M. Barahona, and M. Buck, Engineering modular and tunable genetic amplifiers for scaling transcriptional signals in cascaded gene networks. Nucleic Acids Res, 2014. 42(14): p. 9484-92.
Berset, Y., et al., Mechanistic Modeling of Genetic Circuits for ArsR Arsenic Regulation. ACS Synth Biol, 2017. 6(5): p. 862-874.
Buffi, N., et al., An automated microreactor for semi-continuous biosensor
measurements. Lab Chip, 2016. 16(8): p. 1383-92.
Truffer, F., et al., Compact portable biosensor for arsenic detection in agueous samples with Escherichia coli bioreporter cells. Rev Sci Instrum, 2014. 85(1): p. 015120.
Sambrook, J., E. Fritsch, and T. Maniatis, Molecular cloning: a laboratory manual, 2nd edn. Cold Spring Laboratory Press. New York, 1989.
Innis, M.A., D.H. Gelfand, and J.J. Sninsky, PCR strategies. 1995: Academic Press.
Myers, E.W. and W. Miller, Optimal alignments in linear space. Computer applications in the biosciences: CABIOS, 1988. 4(1): p. 11-17.
Smith, T.F. and M.S. Waterman, Comparison of bioseguences. Advances in applied mathematics, 1981. 2(4): p. 482-489.
Needleman, S.B. and C.D. Wunsch, A general method applicable to the search for similarities in the amino acid seguence of two proteins. Journal of molecular biology, 1970. 48(3): p. 443-453.
Pearson, W.R. and D.J. Lipman, Improved tools for biological seguence comparison. Proceedings of the National Academy of Sciences, 1988. 85(8): p. 2444-2448.
Karlin, S. and S.F. Altschul, Methods for assessing the statistical significance of molecular seguence features by using general scoring schemes. Proceedings of the National Academy of Sciences, 1990. 87(6): p. 2264-2268.
Karlin, S. and S.F. Altschul, Applications and statistics for multiple high-scoring segments in molecular seguences. Proceedings of the National Academy of Sciences, 1993. 90(12): p. 5873-5877.
Stephens, C., et al., A cell cycle-regulated bacterial DNA methyltransferase is essential for viability. Proceedings of the National Academy of Sciences, 1996. 93(3): p. 1210- 1214. Hierlemann, A. and H. Baltes, CMOS-based chemical microsensors. Analyst, 2003. 128(1): p. 15-28.
Evinger, M. and N. Agabian, Envelope-associated nucleoid from Caulobacter crescentus stalked and swarmer cells. J Bacteriol, 1977. 132(1): p. 294-301.
Arellano, B.H., et al., Identification of a dehydrogenase required for lactose metabolism in Caulobacter crescentus. Appl Environ Microbiol, 2010. 76(9): p. 3004-14.
Skerker, J.M., et al., Two-component signal transduction pathways regulating growth and cell cycle progression in a bacterium: a system-level analysis. PLoS Biol, 2005. 3(10): p. e334.
Christen, B., et al., High-throughput identification of protein localization dependency networks. Proc Natl Acad Sci U S A, 2010. 107(10): p. 4681-6.
Park, D.M., et al., Identification of a U/Zn/Cu responsive global regulatory two- component system in Caulobacter crescentus. Mol Microbiol, 2016.
Stephens, C, et al., A cell cycle-regulated bacterial DNA methyltransferase is essential for viability. Proc Natl Acad Sci U S A, 1996. 93(3): p. 1210-4.
Fiebig, A., et al., Interaction specificity, toxicity and regulation of a paralogous set of ParE/RelE-family toxin-antitoxin systems. Mol Microbiol, 2010. 77(1): p. 236-51.
Goodrich, R., Lorega, G., LLNL Livermore Site and Site 300 Environmental Restoration Project Standard Operating Procedures (SOPs). Lawrence Livermore National Laboratory Livermore, Calif, 2016. (UCRL-MA-109115 Rev. 15).
Arellano, B.H., et al., Identification of a dehydrogenase required for lactose metabolism in Caulobacter crescentus. Applied and environmental microbiology, 2010. 76(9): p. 3004-3014.
Fiebig, A., et al., Interaction specificity, toxicity and regulation of a paralogous set of ParE/RelE-family toxin-antitoxin systems. Molecular microbiology, 2010. 77(1): p. 236- 251.
Andersen, J.B., et al., New unstable variants of green fluorescent protein for studies of transient gene expression in bacteria. Appl Environ Microbiol, 1998. 64(6): p. 2240-6. Hu, P., et al., Whole-genome transcriptional analysis of heavy metal stresses in
Caulobacter crescentus. J Bacteriol, 2005. 187(24): p. 8437-49.
Park, D.M. and Y. Jiao, Modulation of medium pH by Caulobacter crescentus facilitates recovery from uranium-induced growth arrest. Appl Environ Microbiol, 2014. 80(18): p. 5680-8.
Yung, M.C., et al., Shotgun proteomic analysis unveils survival and detoxification strategies by Caulobacter crescentus during exposure to uranium, chromium, and cadmium. J Proteome Res, 2014. 13(4): p. 1833-47.
Hillson, N.J., et al., Caulobacter crescentus as a whole-cell uranium biosensor. Appl Environ Microbiol, 2007. 73(23): p. 7615-21.
Cabantous, S., et al., A new protein-protein interaction sensor based on tripartite split- GFP association. Sci Rep, 2013. 3: p. 2854.
Britos, L., et al., Regulatory response to carbon starvation in Caulobacter crescentus.
PLoS One, 2011. 6(4): p. el8179.
Jonas, K., et al., Proteotoxic stress induces a cell-cycle arrest by stimulating Lon to degrade the replication initiator DnaA. Cell, 2013. 154(3): p. 623-36. Modell, J.W., A.C. Hopkins, and M.T. Laub, A DNA damage checkpoint in Caulobacter crescentus inhibits cell division through a direct interaction with FtsW. Genes Dev, 2011. 25(12): p. 1328-43.
da Silva Neto, J.F., R.F. Lourenco, and M.V. Marques, Global transcriptional response of Caulobacter crescentus to iron availability. BMC Genomics, 2013. 14: p. 549.
Blanco, A.G., et al., Tandem DNA Recognition by PhoB, a Two-Component Signal Transduction Transcriptional Activator. Structure, 2002. 10(5): p. 701-713.
Park, D.M. and P.J. Kiley, The influence of repressor DNA binding site architecture on transcriptional control. MBio, 2014. 5(5): p. e01684-14.
Nriagu, J.O., Lead orthophosphates. I. Solubility and hydrolysis of secondary lead orthophosphate. Inorganic Chemistry, 1972. 11(10): p. 2499-2503.
Tripet, B., et al., Engineering a de novo designed coiled-coil heterodimerization domain for the rapid detection, purification and characterization of recombinantly expressed peptides and proteins. Protein Eng, 1997. 10(3): p. 299.
Yung, M.C. and Y. Jiao, Biomineralization of uranium by PhoY phosphatase activity aids cell survival in Caulobacter crescentus. Appl Environ Microbiol, 2014. 80(16): p. 4795- 804.
Roggo, C. and J.R. van der Meer, Miniaturized and integrated whole cell living bacterial sensors infield applicable autonomous devices. Curr Opin Biotechnol, 2017. 45: p. 24-33. Park, D.M., et al., Identification of a U/Zn/Cu responsive global regulatory two- component system in Caulobacter crescentus. Mol Microbiol, 2017. 104(1): p. 46-64. Weinstein, L.H. and A. Davison, Fluorides in the Environment: Effect on Plants and Animals. 2004.
Stockbridge, R.B., et al., A family of fluoride-specific ion channels with dual-topology architecture. Elite, 2013. 2: p. e01084.
Malakooti, J., S.P. Wang, and B. Ely, A consensus promoter seguence for Caulobacter crescentus genes involved in biosynthetic and housekeeping functions. J Bacteriol, 1995. 177(15): p. 4372-6.
APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>ALXI01000102.1/18602-18543 Clostridium sp . Maddingley MBC34-26
contig_203_1 , whole genome shotgun sequence.
ATATTTTCAGGTGATGGAGTTCGCCTTTAACTGCGTAATGCTAATGACTCCTACAAAATA
>KB944507.1/1472202-1472268 Plesiomonas shigelloides 302-73 genomic scaffold scaffold00002, whole genome shotgun sequence.
TGCCATCGTGGAGATGGCATTCCTCCCTTAACCGTCGGGTTTCCCGTCTGATGATGCCTG CTAGATT
>MGOG01000202.1/8032-8093 Chloroflexi bacterium RBG_19FT_COMBO_47_9 rbg_19ft_combo_scaffold_45167 , whole genome shotgun sequence.
ACAAATTAAGGCGATGAGGCTCGCCGCAATTGTCTATTAGGACTGATAGCCTCTGTAGAA GA
>FUZT01000010.1/96793-96866 Clostridium halophilum strain Ml genome assembly, contig: G325DRAFT_scaffoldOOOlO .10
TTTAAGATAGGGAATGAAGTTCTCCCTGGATTAATATTCCAAACCGCTGATAAGCTAATG
ACTTCTAAGCACCG
>MKWG01000037.1/377216-377147 Sphingobacteriales bacterium 44-15
scnpilot_expt_750_bf_scaffold_50, whole genome shotgun sequence.
TTGCTAAAAGGAAATGGTGTCTTCCTAATTGAACCGCTCATCTTTTTGGGCTGATGGCGC CTGCAGGTTA
>CP005085.1/264607-264537 Sphingobium sp . TKS plasmid pTKl, complete sequence .
GGCATCAACGGCGATGGATTTCCGCCTGGCTTCGGCCGAACCGCCTCGGGGTTGATGATT
CCTACCTGCTG
>MERH01000282.1/3229-3080 Betaproteobacteria bacterium
RIFCSPL0W02_12_FULL_62_13b rifcsplowo2_12_scaffold_93587 , whole genome shotgun sequence.
CAGTCGCAAGGAGATGGCATTCCTCCTTTAACNNNNNNNNNNNNNNNNNNNNNNNNNNNN NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNCTCCTTTAAC CGCCGCAAGGCTGATGATGCCTACCTGACC
>MGDD01000266.1/9395-9456 Candidatus Schekmanbacteria bacterium
RBG_13_48_7 RBG_13_scaffold_4973 , whole genome shotgun sequence.
ATGATTATAGGTGATGGAGTTCGCCATTAACTGCCGTAAAGGCTGATAACTCCTACTAAA AA
>AZDZ01000001.1/10642-10710 Lactobacillus nodensis DSM 19682 = JCM 14932 = NBRC 107160 strain DSM 19682 NODE_3, whole genome shotgun sequence.
AAGTTAATCGGCGATGACGTTCGCCACTAAAATTAATTGGTAAATCAATTTGATGACGTC TACTGCACA
>CP003646.1/2696294-2696358 Gloeocapsa sp. PCC 7428, complete genome.
TCTGAACTGGGTGATGGAGCTCACCCTAACCGCCTATTATTAAGGCTGATGGCTCCTACT
ATTCT
>CP001958.1/840150-840069 Segniliparus rotundus DSM 44985, complete genome .
CCCAGACACGGCGATGGATCTCCGCCGGGACAGACTTCCTGTCCGAACCGCCCCCAACGA
GGCTGATGGCTCCTTCGCATGA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>GG658170.1/1322864-1322793 Oxalobacter formigenes OXCC13 genomic scaffold supercontl .1, whole genome shotgun sequence.
CTCTTTCAAGGTGATGGTGATCCACCTTTTCCAACCGCCACGTTTTTTGTGGATGATGAC
GCCTGATATGCT
>LDSG01000006.1/39539-39613 Pantoea dispersa strain SA5 contig_6, whole genome shotgun sequence.
TGCCTCTGAGGTGATGGCGTGCCACCTTACCCAACCGCCCTCTGGTTACAGACGGCTGAT
GACGCCTGACAACTT
>AP012046.1/687349-687410 Tetragenococcus halophilus NBRC 12172 DNA, complete genome.
ATTATTTATAGAGATGGTGCTCTCTATAACTGCCAGGGCTGGCTGATGGCGCCTGTTATA
TA
>MNXB01000100.1/20990-20920 Elusimicrobia bacterium CG1_02_37_114 cg_0.2_subl0_scaffold_372_c, whole genome shotgun sequence.
TATATGTTAGGAGATGGAGTTCTCCATTAATCGCTTCAACGTTTTTAAAGGCTGATAACT CCTACGATTAA
>MKSV01000007.1/1496035-1495976 Acidobacteriales bacterium 59-55
SCNpilot_cont_750_bf_scaffold_4, whole genome shotgun sequence.
CGGAGAGAAGGGAATGGGGTCTCCCGAAACCGCCGAGAGGCTGATGACTCCTGCCAATTC
>LGCL01000045.1/10653-10588 Ornatilinea apprima strain P3M-1 contig_4, whole genome shotgun sequence .
TGATGATTGAGCGATGAGGCTCGCTCTTGAACCAAACGCTGATAAGCTGATGGCCTCTAC
CAGATT
>AJFI01000022.1/160240-160300 Mycobacterium xenopi RIVM700367 contig22, whole genome shotgun sequence .
GTTGACGTTGGCGATGAAGCTCGCCTTGATTGCCGCACCGGCTGATGGCTTCTACCGCGT
G
>CP016557.1/373614-373550 Vibrio coralliilyticus strain 58 chromosome II, complete sequence.
ATTACGCAAGGTGATGGGGTTCCACCTACTTAACCGCCACTCTGGCTGATGACTCCTACA
GAATA
>MGDE01000163.1/4593-4530 Candidatus Schekmanbacteria bacterium
RBG_16_38_10 RBG_16_scaffold_45780, whole genome shotgun sequence.
TTTAAAAAAGGCGATGGGGTTCGCCATTAAGCATCCATAAGGGATTAATGACCCCTACTG AAAA
>JMFG01000038.1/10101-10037 Thermoanaerobaculum aquaticum strain MP-01 contig_38, whole genome shotgun sequence.
ACTCTCGCTGGCGATGGAGTCCGCCGTTAACCGCCCACCCTTGGGCTGATTACCCCTACC
CCACA
>MGNU01000076.1/2309-2394 Chloroflexi bacterium RBG_16_57_9
RBG_16_scaffold_27514 , whole genome shotgun sequence.
ATTGCCGTTGGCGATGAGGCTCGCCAGGGGCCGGAGTGGCGGCCCCGAACAACTGTCGGT GAACGACTGATAGCCTCTACCTTGTG
>LK028559.1/250889-250960 Acholeplasma oculi genome assembly, chromosome:
1
ATAGATAAAGGGAATGAAGTGCTCCCTTCGTACATACGTAAACCGCTTATTGCTGATGAC
TTCTACAGATTT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>LGCK01000014.1/710379-710315 Leptolinea tardivitalis strain YMTK-2 contig_l, whole genome shotgun sequence.
AATAGTTCTGGCGATGAAGCTCGCCCTTTAAAAGTGCCAATCTGGCTGATAGCTTCTGAC
CTGGT
>CP000481.1/804142-804061 Acidothermus cellulolyticus 11B, complete genome .
CCTGATGACGGCGATGGAGTCCGCCTGACCGGTGTGATCGACACCGGCGAAGTGTCACCT
GGCTGATGGCTCCTACCTCCCA
>AZQP01000012.1/15223-15164 Fervidicella metallireducens AeB contig00012, whole genome shotgun sequence .
TATTATTAAGGTGATGGGGTTCACCTTTAAGTGCTTTGTGCTGATGACTCCTGGATTGAA
>CP001656.1/4006535-4006595 Paenibacillus sp . JDR-2, complete genome.
AGTGAATAAGGCGATGGAGTTCGCCATTAACCGCCTGCGGGCTAATGACTCCTACCAATG
T
>HF986505.1/7297-7365 Bacteroides sp . CAG:598 genomic scaffold, scf205
TTAATCGGAGGGAATGGGACTTCCCTGTTATAAACCGCTAATCAAGTAGCTGATAGTTCC
TGTTTGTAG
>JRLY01000013.1/111920-111856 Flavobacterium subsaxonicum WB 4.1-42 = DSM 21790 strain WB 4.1-42 contigl4, whole genome shotgun sequence.
TTGCATACCGGCAATGATGTCTGCCCTGAACCGCTGCCCGCGCAGCTGATGACGTCTGTT TAATG
>MHYC01000027.1/2817-2754 Planctomycetes bacterium RBG_13_44_8b
RBG_13_scaffold_16883, whole genome shotgun sequence.
TTAAAACAAGGGGATGGAGTTCCCCCAAAACCGCCTGAAAAAGGCTGATAACTCCTACCA ATAT
>FR887854.1/2121-2192 Eubacterium sp . CAG:202 genomic scaffold, scfl4
TATATGTATGGGAATGAGGTTCTCCCACGGATTTTTTCCGAAACCGCTTTTGCTGATGAC
TTCTGTTTTATA
>LZPM01000007.1/52149-52086 Methanobacterium sp . A39 contig_16, whole genome shotgun sequence.
TTTTGTATCTGCGATGGGGTTCGCATTAACCGCTTAAAAATAAGCTGATGACTCCTAACC
TATT
>JFHN01000066.1/21329-21398 Erwinia mallotivora strain BT-MARDI contig_66, whole genome shotgun sequence .
CAGAACCAGGGTGATGGCGTTCCACCTTTCCCAACCGCTCCATTCCGGAGCTGATGACGC
CTGATGTAAC
>CP000759.1/989075-989138 Ochrobactrum anthropi ATCC 49188 chromosome 2, complete sequence.
GTACATGATGGGAATGGGGTTCTCCCGAAACTGCCAGCAATTGGCTGATGACTCCTGCTT
GAAT
>MZGT01000023.1/13112-13048 Clostridium chromiireducens strain DSM 23318 CLCHR_contig000023, whole genome shotgun sequence.
AATTTTATAGGTGATGAAGTTCGCCTTTAAACATCTCTTTAGAGATTGATGACTTCTACT ACAAT
>MKVZ01000019.1/56053-56113 Rhizobiales bacterium 63-22
SCNpilot_cont_500_bf_scaffold_96, whole genome shotgun sequence.
GATGTATGAGGGGATGGAGTTCCCCGCGACCGCCCTCGAGGCTGATGACTCCTGCCAGAA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
A
>CP001941.1/738406-738345 Aciduliprofundum boonei T469, complete genome.
TCTGCTTCGGGTGATGGCGTCCACCCCGAAACGCCCGCGTGGCTGATGACGCCTTTTCGC
AT
>ADBS01000001.1/1408516-1408593 Photobacterium damselae subsp. damselae CIP 102761 Contig58, whole genome shotgun sequence.
ATCTTAACGGGAGATGATGTTCCTCCTTTAACCGCCTTTCTGATCTTTTTCTCTAAGGAT GATGACGTCTAACAACAG
>AVBI01000012.1/64911-64844 Flavobacterium cauense R2A-7 FCR2A7T_12, whole genome shotgun sequence.
CAGACAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTCACAAAAGCTGATGACGCCT
GATTAATA
>FN806773.1/1914034-1914105 Propionibacterium freudenreichii subsp.
shermanii CIRM-BIAl, complete genome
GATGAACGCGGCGATGGATTCCGCCTGGGCCTCGGCCCGAACCGCCGTCCGGCTGATGAT
TCCTGTCCCATG
>JH815223.1/216852-216924 Clostridium sp . 7_2_43FAA genomic scaffold supercont2.4 , whole genome shotgun sequence.
ATAATTTTTGGGAATGAAGTTCTCCCATGAGAAATCTAAACCGTCAAGTAAGACTAATGA
CTTCTATGATTTA
>JNGB01000017.1/21296-21371 Robinsoniella sp . RHS robin_cl7, whole genome shotgun sequence.
AAAATAAACGGGAATGAGGTTCTCCCCAGGTACAACAATACCTAAACTGCTATTAGCTGA
TGACTTCTGTTTTAGG
>MRAE01000024.1/40018-40077 Clostridium sp . IEH 97212 scaffold_23, whole genome shotgun sequence.
ATAATTTTAGGCGATGGGGTTCGCCATTAAATGCGTAAAGCTAATGACTCCTACAAATAG
>MGPS01000070.1/30145-30206 Deltaproteobacteria bacterium GWC2_42_11 gwc2_scaffold_898 , whole genome shotgun sequence.
CATCAAATAGGTGATGGAGTTCACCTTTAACCGCTTTAATGGCTGATGACTCCTACAAAT TA
>AJFI01000022.1/153221-153282 Mycobacterium xenopi RIVM700367 contig22, whole genome shotgun sequence .
TCGCGGACAGGCGATGAGGCTCGCCTTGAACTGCCACACCGGCTAATGGCTTCTACCCGT
GA
>LKEU01000030.1/118640-118581 Acetobacterium wieringae strain DSM 1911 ACWI_contig000030, whole genome shotgun sequence.
TTCATAGAAGGTGATGGAGTTCACCAAAATTGCTTATCAGCTGATGACTCCTGCAGGATT
>JXNU01000003.1/4160542-4160614 Erwinia tracheiphila strain BuffGH
Etrach_2, whole genome shotgun sequence.
GTTTCGCAAGGTGATGGTGTTCCACCTTTCCCAACCGCCCTGTTCGCAAGGGGCTGATGA
CGCCTGATAACAC
>FPBV01000014.1/4475-4412 Alicyclobacillus macrosporangiidus strain DSM 17980 genome assembly, contig: Ga0104483_114
ATCGAATCCGGTGATGGAGCTCACCGTATAAATGCCCATTGAGGATGATGGCTCCTGTGC
ATCT
>ABGR01000008.1/95568-95493 Vibrio sp . AND4 1103602000417, whole genome APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
shotgun sequence.
ATTGGGGTTGGTGATGGGGTGCCACCGGAATCGAATGATTCAAACCGCTTGTATAGCTGA
TGACTCCTACAGAGAC
>MKSQ01000014.1/173748-173824 Bosea sp. 67-29
SCNpilot_expt_1000_bf_scaffold_308, whole genome shotgun sequence.
TTCGGCGATGGGGATGGAGTTCCCCCGAAACCGCCCGCCGGTCGCCATTCCGGTCGGCTG
ATGACTCCTACCGCACG
>BA000016.3/731613-731672 Clostridium perfringens str. 13 DNA, complete genome .
ATTTTTATAGGCGATGGAGTTCGCCATTAAATGCTTTATGCTAATGACTCCTACAAAATA
>CP001785.1/1464504-1464567 Ammonifex degensii KC4, complete genome.
ACCACATAAGGCGATGGAGCCCGCCTTAACCGCCCGGAGAAGGGCTGATGGCTCCTGGTC
TTTC
>JRMY01000022.1/50889-50947 Clostridiales bacterium S7-1-4 contig22, whole genome shotgun sequence.
AGTTTAATAGGGAATGAAGTTTCCCGTTAATTGCAACTGCTGATGACTTCTACAATTTT
>CP002770.1/2630559-2630622 Desulfotomaculum kuznetsovii DSM 6115, complete genome.
ATCACACCAGGCGATGGAGCTCGCCTTTAAGCGCCTCTTTCGGGCTAATAGCTCCTACCA
GAAT
>CP006850.1/5427340-5427403 Nocardia nova SH22a, complete genome.
TGATCTGTAGGCGATGAAGCTCGCCTTCGACCGCGCCCCCGGCGCTGATGGCTTCTACCA
CGGT
>KQ969397.1/92070-92008 Streptococcus sp . DD11 genomic scaffold
scaffold00006, whole genome shotgun sequence.
AAGAGTCAAAGGGATGGTGCACCCCTATAACCGCTAGTGATAGCTGATGGCGCCTGTTAG
AAA
>MENL01000062.1/7077-7147 Bacteroidetes bacterium GWB2_41_8
gwb2_scaffold_l 940 , whole genome shotgun sequence.
TTTGTGGGCGGCAATGAGGTCTGCCTTTAACCGCCTGTCCTCCTGTTCAGGCTGATGACA TCTACCAACGA
>LCVM01000117.1/8488-8556 Providencia rettgeri strain MR4
P_rettgeri_contig_117 , whole genome shotgun sequence.
AAACCTTTGGGAGATGGCATTCCTCCTTATATAAAACCGCCCGTAGAGGCTGATGATGCC TACGTTAAC
>FQUN01000016.1/10512-10572 Leeuwenhoekiella marinoflava DSM 3653 genome assembly, contig: EJ63DRAFT_scaffold00014.14
ACATACAAAGGCGATGGAGTTCGCCAAAACCGCCCAAAAAGCTAATGACTCCTGCTCAAT
T
>LMSL01000040.1/140723-140662 Frateuria sp . Soil773 contig_8, whole genome shotgun sequence.
TCGATGTCGGGAGATGGCATGCCTCCCTGAACCGCCGCAAGGCTGATGATGCCTACCATC
CC
>CTEN01000007.1/29015-29096 Streptococcus sp . FF10 genome assembly, contig: CONTIG000007
TTTTATCAAGGGAATGAGGCACTCCCTAAATTTCTGTAAAGAATTTAAACCGCGATTTTT
CGCTGATGGCTTCTATGACGCT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>CP002545.1/4356416-4356351 Pseudopedobacter saltans DSM 12145 chromosome, complete genome.
TAATCAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTTTAAGCTGATGACGCCTGA
TTAATG
>MKZW02000001.1/3834718-3834787 Kluyvera intestini strain GT-16
tigOOOOOOOO, whole genome shotgun sequence.
TGCATGCAAGGTGATGGCGTTCCACCTTACCCAACCGCTTCCCAGCGAAGCTGATGACGC
CCCCGGTATG
>AYLZ02000183.1/23975-23911 Shinella sp . DD12 SHLA_40c, whole genome shotgun sequence.
GCGTGGGATGGGGATGGAGCTCCCCGACAACCGCCGTTTGACCGGCTGATAGCTCCTGCC
AAGAC
>CP001739.1/3350399-3350323 Sebaldella termitidis ATCC 33386, complete genome .
TAAACTTAAGGAGATGAAGCTCTCCGCGGTGTATGCCGGAACCACATGTTTTTCATGTTA
ATAGCTTCTGCAGCGTT
>AVOG02000067.1/146393-146465 Alcaligenes sp . EGD-AK7 Contig7, whole genome shotgun sequence.
TTTCTTCCTGGAGATGACGTTCCTCCTTTAACCGCCGTTTCGTTCTGAAGCAGCTGATGA
CGTCTACAGTGAC
>MKVK01000019.1/64767-64705 Magnetospirillum sp . 64-120
scnpilot_p_inoc_scaffold_239, whole genome shotgun sequence.
GCGTGGGATGGGGATGGAGCTCCCCGACAACCGCCCGCGAGGGCTGATAGCTCCTGCTGC GAG
>FOOE01000012.1/96932-96851 Clostridium cadaveris strain NLAE-zl-G419 genome assembly, contig: Ga0070261_112
AAAAGTATTGGGAATGAAGTTCTCCCATAGTATTCTTTTATGCTAAAACCGCTGAATTTA
GGCTGATGACTTCTGTGATTAT
>CP002109.1/3176874-3176949 Clostridium saccharolyticum WM1, complete genome .
AATAGCCAAGGGAATGAAGTTCTCCCTGCGCTGTTAAAGCCGAACCGCTTATTAAGCTGA
TGACTTCTGCGATTAT
>HF570958.1/4210628-4210703 Tetrasphaera japonica T1-X7 genomic scaffold, 1540_scaffoldl
GTGGGCCCCGGCGATGGATCCCGCCAGGACGCGCAGCGTCCGAACTGCCGCCCCGGCTGA
TGGCTCCTGCTCCGTC
>CP002206.1/3251024-3250949 Pantoea vagans C9-1, complete genome.
CCGGCTCAAGGTGATGGCGTGCCACCTTTCCTAACCGCCGGTGGGCTTAACACCGGCTGA
TGACGCCTGACACTCA
>LDZF01000021.1/90063-89994 Pluralibacter gergoviae strain JS81F13 contig_21, whole genome shotgun sequence.
TTAATGCAAGGTGATGGTGTTCCACCTTTCCCAACCGCTTCCCTGAGAAGATGATGACGC
CTGGTATGTC
>LBDA02000090.1/15520-15584 Streptomyces malaysiense strain MUSC 136 90, whole genome shotgun sequence .
GAACGCGCCGGTGATGGGGCTCACCGCAACCGCGGCGACATGCCGCTGACGGTCCCTGGT
CGAAC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>KB850961.1/2036-1973 Clostridium colicanis 209318 genomic scaffold acBRg- supercontl .5, whole genome shotgun sequence.
ATAATAATAGGTTATGAAGTTAACCTTTAACTGCTTATTAAAAGCTAATGACTTCTACAA
ATTT
>LZLJ01000064.1/71689-71630 Mycobacterium sp . 1274756.6 contig_67, whole genome shotgun sequence.
CGCGGGTAAGGCGATGAGGCTCGCCTGAACCGCCGTTTGGCTGATGGCCTCTACCACGGA
>BAND01000009.1/12481-12543 Acidomonas methanolica NBRC 104435 DNA, contig: Amme_009.
ACGGGTAAGGGAGATGGAGTTCCTCCCGTAACCGCCCTCAGGGCTGATGACTCCTGCCGA
AGT
>AZEC01000003.1/146176-146240 Lactobacillus perolens DSM 12744 Scaffold3, whole genome shotgun sequence .
CATCTAATCGGTGATGACGTTCGCCGTAACCATTACAGAACGTAATTAATGACGTCTACA
GTTGC
>AQQS01000011.1/89624-89562 Aurantimonas sp . 22II-16-19i contigll, whole genome shotgun sequence.
GGGAGTGATGGGGATGGGGTTCCCCCGATAACCGCCATGTTGGCTGATGACTCCTACCGA
ACG
>JGVW01000053.1/43962-43897 Burkholderia sp . Iig30 ctg20, whole genome shotgun sequence.
TCGCGCTTCGGAGATGGCATGCCTCCCTTAACCGCCGGTAGCCCGGCTGATGATGCCTGC
GCGTTC
>MGSS01000136.1/35936-35874 Deltaproteobacteria bacterium
RIFCSPLOWO2_12_FULL_60_19 rifcsplowo2_12_scaffold_703 , whole genome shotgun sequence.
TCCATTCGCGGAGATGGCGTTCTCCCCTAACCGCCAGCCAAGGCTGATGACGCCTGCTTT
CAT
>FWXW01000002.1/266583-266656 Papillibacter cinnamivorans DSM 12816 genome assembly, contig: EJ70DRAFT_scaffold00002.2
CTCAAATGCGGGAATGAAGTTCTCCCCAGGCAATGGCCTGAAGTGCCTTAACGGCTGATG
ACTTCTGCAGCGTA
>KB849803.1/86858-86792 Acinetobacter haemolyticus CIP 64.3 genomic scaffold acLrj-supercontl .14, whole genome shotgun sequence.
TAGGGCAAAGGAGATGGCATTCCTCCTGTAACAAACCGCCATTGTGGCTAATGATGCCTA CGTTACC
>AOTI010612873.1/11168-11096 Triticum urartu cultivar G1812 contig612874, whole genome shotgun sequence .
TCGCTATAAGGTGATGGTGTTCCCCCTTTCCCAACCGCCTCGTCCGTAAGGGGCTGATGA
CGCCTGATAACCC
>MIHD01000015.1/16307-16245 Mycobacterium triviale strain HMC_M3
PROKKA_contigO 00015, whole genome shotgun sequence.
AATCAAATCGGCGATGAAGCTCGCCCTTAACCGCCGAACCCGGCTGATGGCTTCTACCCG AAG
>CP002865.1/773260-773339 Zymomonas mobilis subsp. pomaceae ATCC 29192, complete genome.
TTTAGCACCGGTAATGGATTCTGCCTGATCCTTTGAATAAGGATCGAACCGCCAATCAGG
CTAATGACTCCTGCTCAAAC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>JWHM01000006.1/60788-60858 Burkholderia sp . MR1 ctgl5, whole genome shotgun sequence.
CGCCCTTCTGGAGATGGCATTCCTCCAAACAACCGCCGGCTGTCCTGCCGGCTGATGATG
CCTACGAGTTC
>LDZF01000020.1/46839-46778 Pluralibacter gergoviae strain JS81F13 contig_20, whole genome shotgun sequence.
CCCGCCGCAGAGGATGGCATACCCCCTTGAACCGCCATGTGGCTGATGATGCCTGCTTTT
TC
>CP001069.1/247334-247411 Ralstonia pickettii 12J chromosome 2, complete sequence .
CCCCGTTGTGGAGATGGCATGCCTCCCTTAACCGCCGGTTAGCGCGTAACGATACCGGCT
GATGATGCCTACAAGTTC
>JYGZ01000003.1/195845-195912 Flavobacterium sp . 316 scaffold_2, whole genome shotgun sequence.
GAAATAGAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTCTAAAAGCTGATGACGCCT
GATTAATA
>MEQT01000143.1/7059-6990 Betaproteobacteria bacterium
RIFCSPLOWO2_02_64_14 rifcsplowo2_02_subl 0_scaffold_2315, whole genome shotgun sequence.
CGGCAATAGGGAGATGGCATTCTCCTCCAACAGGAAACCGCCGGAAACGGCTGATGATGC
CTGCTGGGTC
>HF995201.1/10844-10915 Firmicutes bacterium CAG:194 genomic scaffold, scf25
TACAGCGTAGGGAATGAAGTTCTCCCTTGGGAAACCTAAACTGCTTACAGAGCTGATGAC
TTCTACGATATT
>LTGD01000049.1/199-133 Anaerosphaera sp . HMSC064C01
Anaerosphaera_spHMPREF2800-l .0_Cont75.5, whole genome shotgun sequence.
TAAATACATGGGAATGAAGTTCTCCCACGAAATGAAACCGCTATAAGCTGATGACTTCTG
TTTTTAC
>KI391948.1/803048-802975 Ruminococcaceae bacterium D16 genomic scaffold acsPy-supercont2.2 , whole genome shotgun sequence.
AATTCTCTTGGGAATGAAGTTCTCCCTTGGTGAATACCTAAACCGCTGATAAGGCTGATG ACTTCTGTGAGCAC
>ADVG01000001.1/1763333-1763468 Ktedonobacter racemifer DSM 44963 strain SOSP1-21 Krac_Contig206, whole genome shotgun sequence.
AGAATAACTGGTGATGGGGTTCACCACGTCATATAATGCCTCTTCAAGAGAAACGTTCAC TGACCGATTTTTTATCTTAAGAAGAATTGGCTATACCTGACGAAACCGCAAAGAAGCTAA TAACTCCTGATAACGC
>CP004050.1/1233497-1233564 Methanobrevibacter sp. AbM4, complete genome.
TAACATTAAGGTGATGAGGTTCCACCTTTAACCGTCAGTGATTACTGAATGATGACTTCT
ATTTTTCA
>MGRA01000126.1/3058-2996 Deltaproteobacteria bacterium
RBG_19FT_COMBO_46_12 rbg_19ft_combo_scaffold_1869, whole genome shotgun sequence .
ATTAGTAAGGGTGATGGAGTTCACCCAATAACCGCCATTGAGGCTGATAACTCCTGCAGG
GAG
>MNVU01000044.1/2332-2265 Candidatus Omnitrophica bacterium CG1_02_49_16 cgl_0.2_scaffold_14807_c, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GTAGTTTTAGGCGATGGGGTTCGCCTGATAGCTAGACTGCCAATGCGGCTGATAACTCCT
ATTGTTCC
>CP017749.1/2317300-2317385 Cupriavidus sp . USMAA2-4 chromosome 2, complete sequence.
TGTATGCGTGGAGATGGCATACCTCCACCCCAGAAGTAGAGCTGGGAAAACCGCACTTTC
AAAGTGCTGATGATGCCTACAGGTCC
>CP003169.1/6228224-6228150 Mycobacterium rhodesiae NBB3, complete genome.
TGGCGCAGTGGTGATGGATCTCGCCCGAGCTCGGGTCGCTCGAACCGCCAACCGGCTGAT
AGTTCCTACCGTCAT
>MKTL01000016.1/298909-298834 Clostridiales bacterium 38-18
SCNpilot_bf_inoc_scaffold_50, whole genome shotgun sequence.
ACCAATGTTGGGAATGAGGTTCTCCCTTAGTCTTTAGGACTTAAACCGCGTGTACGCTGA TGACTTCTGTGATATC
>LGHM01000032.1/20467-20391 Bradyrhizobium sp. AS23.2
Bradyrhizobium_AS_23.2_c32 , whole genome shotgun sequence.
CACCCAGATGGGGATGGAGTCCCCCGACAACCGCCCGGTGGTCGAGACTGTATTGGGCTG ATGACTCCTGCTTGAGG
>LSKU01000001.1/1852967-1852903 Tepidibacillus sp. Z9 scaffoldOOOOl, whole genome shotgun sequence.
TAAATAATGGGCGATGGAGTTCGCCTTAAACTGCCTAATCATAGGCTAATGACTCCTACC
AGAAC
>ATBB01000155.1/387-313 Lactococcus lactis subsp. cremoris TIFN6
scatfoldl2.1_39, whole genome shotgun sequence.
AATAATGATGGGTATGGTGCACACCCGAAACCGCTTTAAGAATAAAATCTTAAAACTAAT GGCGCCTACAAACAA
>JQCR01000003.1/1098358-1098297 Paenibacillus wynnii strain DSM 18334 unitig_3_lr, whole genome shotgun sequence.
AGTTGATAAGGCGATGGAGTTCGCCAAAACCGCCGGTAACGGCTAATGACTCCTACCAGC
GA
>GG663535.1/988748-988826 Bifidobacterium angulatum DSM 20098 genomic scaffold ScfldO, whole genome shotgun sequence.
CTTACCGGTGGCGATGGGACTCGCCTGGGGCCTGCAAGGATTCCCGAACCGCAATTCCGC TGATGGTTCCTATTGACGC
>ACNP01000013.1/37301-37241 Leptospirillum ferrodiazotrophum UBAL3_4481_9, whole genome shotgun sequence .
CCGGCGGATGGCGATGGGGTTCGCCTAATCCCGCATTTCGGGTGATGACTCCTACCATGG
G
>AFBX01000321.1/11509-11430 Nitrospirillum amazonense Y2 contig00386, whole genome shotgun sequence .
GCTGCGTCCGGCAATGGATTTCTGCCTGGCCGTCTCACAGCCGAACCGCCCCATCCGGGG
CTGATGATTCCTACCTTCGG
>FONA01000033.1/2272-2211 Thermophagus xiamenensis strain DSM 19012 genome assembly, contig: Ga0131197_133
TTTGTACAGGGCAATGAGGTCTGCCTTTAACCGCCAAAATGGCTGATGACTTCTACTTTA
GT
>KB976103.1/684369-684288 Butyricicoccus pullicaecorum 1.2 genomic scaffold acBRa-supercontl .1, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AACAGCCGAGGGAATGAATGTCTCCCGCGGCCTTTGCGCCGAAACCGCTTGAATGACGAG
AGCTGATGACTTCTGCAACGAT
>MIBH01000036.1/29669-29604 Spirochaetes bacterium RBG_13_51_14
RBG_13_scaffold_263 , whole genome shotgun sequence.
CTTTTTATTGGGGATGGAGTCCCCCAAAAAACGCCTTTTAAAAAGGCTGATGACTCCTAC CGCTTC
>MSXN01000005.1/102347-102277 Nitrospira sp . ST-bin5 254, whole genome shotgun sequence.
ATTCCTTACGGTGATGGGGTTCACCGGAACCGCCTGCGGTCTCTACCCAGGCTGATAACT
CCTACGCACCC
>AL935263.2/1741232-1741167 Lactobacillus plantarum WCFS1 complete genome
AAGTTAATCGGCGATGACGTTCGCCACATAATAATTGATAATCAATTTGATGACGTCTAC
TGTTTG
>AJHJ01000011.1/95451-95513 Serratia sp . M24T3 contigOll, whole genome shotgun sequence.
CCGTTTTCTGGAGATGACATTCCTCCCAAACCGCCCTGACCGGCTAATGATGTCTACGCA
ATC
>BAVL01000003.1/2472-2405 Chryseobacterium indologenes NBRC 14944 DNA, contig: CIN01S03.
GTACCAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTACAGAAGCTGATGACGCCT GATTGAGA
>ABYL01000251.1/1837-1776 Burkholderia sp. H160 ctg00372, whole genome shotgun sequence.
TTCGCGGCTGGAGATCGCATTCTCCATTAACCGCCCCTGGGGCTGATGATGCCTGCTACG
CC
>MKSQ01000024.1/27769-27707 Bosea sp . 67-29
SCNpilot_expt_1000_bf_scaffold_559, whole genome shotgun sequence.
TCACCGTATGGGGATGGAGTCCCCCGATAACCGCCCTAACCGGCTGATGACTCCTACAAG
CGC
>LADT01000026.1/67639-67568 Clostridiaceae bacterium BRH_c20a
BRHa_1003380, whole genome shotgun sequence.
ATATGGATAGGCGATGGAGCTCGCCTTTAACCGCTAAATGTTTGTTAATTAGCTAATGGC
TCCTACTGGTCA
>CP019948.1/2037217-2037156 Methylocystis bryophila strain S285, complete genome .
TGCCATCCAGGGGATGGGGTGCCCCCGTAACCGCCGCAAAGGCTGATGACTCCTACTTGA
GA
>MERP01000018.1/5185-5246 Betaproteobacteria bacterium
RIFCSPL0W02_12_FULL_68_19 rifcsplowo2_12_scaffold_13509, whole genome shotgun sequence.
CCGCGGCAGGGAGATGGCATTCCTCCTTTAACCGCCGCGAGGCTGATGATGCCTATCTGA
CA
>AYZJ01000062.1/5414-5473 Lactobacillus camelliae DSM 22697 = JCM 13995 strain DSM 22697 NODE_98, whole genome shotgun sequence.
AACATGATAGGCGAAGGTGTTCGCCATAACCGTTATTCAACTAATGACACCTAGTTTCGG
>KE557320.1/1289108-1289048 Rubellimicrobium thermophilum DSM 16684 genomic scaffold Rl_scaffold2 , whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GTCCGCCCGCGGGATGGGGCTTCCCGACAACCGCCTTGTGGCTGATGGCTCCTGCTGCGG
G
>KB946293.1/189807-189872 Enterococcus asini ATCC 700915 genomic scaffold acvKe-supercontl .3 , whole genome shotgun sequence.
TTGTGTAAGAGGGATGGTGCTCCCTCCTAACTGCTGAGTGATTCAGCTGATGGCGCCTGT TATAAG
>AEXE01000084.1/6900-6834 Burkholderia sp. TJI49 contig072, whole genome shotgun sequence.
TTTCGCGCTGGAGATGGCATTCCTCCTCATAACCGCTGGTCAGCCAGCTGATGATGCCTA
CGCGTTC
>MGQI01000010.1/5244-5173 Deltaproteobacteria bacterium RBG_13_58_19 RBG_13_scaffold_11141, whole genome shotgun sequence.
AGTTTATTAGGCGATGAAGTCCGCCAAAACCGCCCCGCTTTCCTGGGCTGGGCTGATGAC TTCTGTCCAACG
>LOES01000122.1/6105-6044 Clostridia bacterium BRH_c25 BRHa_1003505 , whole genome shotgun sequence.
ATAACAACTGGCGAAGGAGTTCGCCATTAATCGTCCTAGTGACTAATGACTCCTGTATTA
AG
>KI392031.1/1315201-1315129 Oxalobacter formigenes HOxBLS genomic scaffold supercont2.2 , whole genome shotgun sequence.
CAGACTCAAGGTGATGGTGTTCCACCTTTCCCAACCGCCACGGTTATTTGTGGCTGATGA
CGCCTGATTTTTT
>AODF01000029.1/2014-2075 Listeria floridensis FSL S10-1187 c26, whole genome shotgun sequence.
ATCTGTAACGGCGATGGAGTTCGCCAAAAACGCGAGTATTCGCTGATGACTCCTATTTAA
AT
>MKVA01000011.1/434077-434016 Devosia sp . 67-54
SCNpilot_cont_1000_bf_scaffold_9, whole genome shotgun sequence.
GCTGATGATGGGGATGGGGCTCCCCGATAACCGCCGTCGTGGCTGATAACTCCTGCTGCA
AC
>CP017267.1/450598-450675 Vagococcus teuberi strain DSM 21459, complete genome .
AACTTAAACGGGAATGATGTTCTCCCATAGTTGAAATTTAACTAAAACTGCTAAATAGCT
AATGACGTCTACATGCTT
>CP001043.1/2012557-2012617 Burkholderia phymatum STM815 chromosome 1, complete sequence.
GCGCGCGCCGGAGATGGCATTCTCCTTAACCGCCCTCGTGGCTGATGATGCCTGCTTCGC
C
>CP002745.1/4925097-4924999 Collimonas fungivorans Ter331, complete genome .
ACACGCTTCGGAGATGGCATTCCTCCCTGAACCGCCGCTTCTGGCCTTGAATCCTTGCGG
GTTCGCCGGCCAGACGCAGCTGATGATGCCTACAAGTTC
>JPF001000020.1/242113-242051 Yersinia ruckeri strain 37551 YR01b0000020, whole genome shotgun sequence .
TAAAGCACTGGAGATGACATTCCTCCATAACCGCCCCTCAAGGCTGATGATGTCTACGTA
ACC
>CP002541.1/1343920-1343997 Sphaerochaeta globosa str. Buddy, complete APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
genome .
ATCTGAATGGGCGATGGGGTTCGCCTTCATAATGAAAACCGCAGAGTTCTGAAAGCTGCT GATAACTCCAGAAATAAA
>CP002038.1/1376618-1376557 Dickeya dadantii 3937, complete genome.
GCCAGTAACGGAGATGACATTCCTCCATAACCGCCCTTCCGGCTGATGATGTCTGCGAGT
CC
>CP001825.1/1684987-1685047 Thermobaculum terrenum ATCC BAA-798 chromosome 1, complete sequence.
ATAACATAAGGCGATGAGGCTCGCCTTGAACCGCCGTTGTGCTGATGGCCTCTACCCTTC
T
>MRCD01000046.1/30478-30406 Calothrix sp . HK-06 NIES-2101_Scaffold_46, whole genome shotgun sequence .
TGATATCTGGGCAATGGAGCTTGCCATTAACCGCCATTACATATAATAAAAGGCTGATAG
CTCCTACTTTCCT
>MLFQ01000454.1/710-784 Pantoea gaviniae strain LMG 25382 Contig0454_, whole genome shotgun sequence .
CGGTGATAAGGTGATGGCGTACCACCTTGCCCAACCGCGACTCCGCTTGCGAGCGCTGAT
GACGCCTGACACGAC
>LWMN01000010.1/106954-107019 Enterococcus thailandicus strain F0711D 46 Ent-16_S35_L001_001_2 , whole genome shotgun sequence.
TTATGTGTGAGGGATGGTGCTCCCTCCTAAACGCTGAGTAAATCAGCTGATGGCGCCTAT TGTAAA
>JYOE01000003.1/272616-272692 Terrabacter sp . 28 contig3, whole genome shotgun sequence.
GTGGTGCTCGGCGATGGATCCCGCCAGGTCACCTCGAGTGCCCGAACCGCCGGCCGGCTG
ATGGTTCCTGTCCGTGC
>KV441297.1/76989-76906 Clostridiales bacterium KLE1615 genomic scaffold Scaffoldl33, whole genome shotgun sequence.
ATATGTTTTGGGAATGAAGTTCTCCCATTGATACATTCCTTGTATCACAAACCGCTGGTT
AAGGCTGATGACTTCTGCGAATTT
>LIDG01000339.1/243-312 OM182 bacterium BACL3 MAG-120619-bin32
contig04497_l_120619-bin32, whole genome shotgun sequence.
TAATTAAAAGGAAATGGTGTTCTTCCTTGCCCAACCGCTTTTAAAAAAAGCTGATGACGC CTGATTAAAT
>BBJM01000037.1/4461-4516 Lactobacillus oryzae DNA, contig: sequence37, strain: SG293.
TTATCTTATGGCGATGGTGCTCGCCTACGCGAAACGCTGATGGCGCCTACTCTGTG
>JNVT01000097.1/26197-26284 Gammaproteobacteria bacterium MFB021 Contig- 100, whole genome shotgun sequence.
TGGCAGCCGGGAGATGGCATTCCTCCCGCGGCATGGGCCCCGCCGCCCCGGCCGCAAACC
GCAGCCCGCTGATGATGCCTGCCAACGC
>AEXE01000084.1/24437-24507 Burkholderia sp . TJI49 contig072, whole genome shotgun sequence.
TCGATCTACGGAGATGGCATTCTCCGCCAACCGCCGCGCCCTCAGGCCCGGCTGATGATG
CCTGCAGCCGT
>LMRS01000059.1/46491-46567 Sphingomonas sp . Leaf339 scaffold7.1, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GTCGGGCACGGCGATGGATTTCCGCCGGGCCATTCGGCCGAACCGCCCTCGCAAGGGCCG
ATGATTCCTACCAGGCG
>MHEU01000028.1/21625-21561 Nitrospirae bacterium
RIFCSPLOWO2_02_FULL_62_14 rifcsplowo2_02_scaffold_21280, whole genome shotgun sequence.
AGATACACAGGGGATGGAGTCCCCCGGGTAACCGCCCTGGCCGGGCTGATGACTCCTACG
CTACC
>HF993815.1/50413-50342 Clostridium sp. CAG:127 genomic scaffold, scflll
ACAGAAAAAGGGAATGAAGTTCTCCCTCGGGTAACCGGAACTGCTTATAAAGCTGATGAC
TTCTACGATATT
>NGMM01000001.1/316799-316877 Enterococcus sp. 9E7_DIV0242 scaffoldOOOOl, whole genome shotgun sequence .
AGAGCAATAGGGAATGAAGTTCTCCCTAGACAGTGTATGTCAAAACCGCAATTATTTTGC
TGATGACTTCTACCACCAT
>FR883718.1/1287-1360 Blautia sp . CAG:52 genomic scaffold, scfl44
ATCCAGAACGGGAATGAGGTTCTCCCACGATTTTTGATCGAAACCGCCAAAAGGCTGATG
ACTTCTGTGCATGG
>AOID01000048.1/66745-66635 Natrinema versiforme JCM 10478 contig_48, whole genome shotgun sequence .
CAGTATCCGGGCGATGGGGCCCGCCTGATCCAACCGCCGGAGAAACGCCGACCCGCTGGC
ATCGCCGCCGGCAGGGTCGAGCGTCGGTCGGCTGACGGTCCCTGCGACCAG
>MKRJ01000017.1/33766-33828 Rhizobiales bacterium 65-79
SCNpilot_cont_1000_bf_scaffold_222 , whole genome shotgun sequence.
CGTTCAGGCGGGGATGGAGTCCCCCGCTAACCGCCTTTCCAGGCTGATGACTCCTACTGT CGG
>MKIM01000027.1/828078-828002 Rhizobium oryziradicis strain N19 scaffold3, whole genome shotgun sequence .
TCTCCCATCGGTAATGGATTCTGCCTGGCATCAAGCCGAACCGCTTCCTTGAAGAAGCTG
ATGACTCCTACTCAACG
>CYZX01000019.1/43096-43038 Clostridium disporicum strain 2789STDY5834856 genome assembly, contig: SCcontig000019
AATAAATAAGGCGATGGAGTTCGCCATAATCGCGTAAAGCTAATGACTCCTACTCTTTT
>KN196470.1/1239508-1239433 Xanthomonadaceae bacterium 3.5X genomic scaffold scaffoldl, whole genome shotgun sequence.
GGAGGCCGGGGAGATGGCATGCCTCCCGTCCGCCATGCGGATGAACCGCCCCAGGGCTGA TGATGCCTGTTCCGTC
>AZEC01000003.1/140411-140499 Lactobacillus perolens DSM 12744 Scaffold3, whole genome shotgun sequence .
AAGTAAATAGGCGATGACGTTCGCCAGTGGTGCTGTATTTAACAGCACTTCAAATATTAC
ACTCCGTAATTGATGACGTCTGCAATATT
>FTNT01000006.1/163404-163489 Williamsia sterculiae strain CPCC 203464 genome assembly, contig: Ga0104459_106
TCAGTCTCTGGCGATGGATCTCGCCGGACCTTGCTGATCTGAGCTGAGGTCGAACCGCCC
TCTCGGCTGATAGTTCCTACTCGTCG
>MFR001000022.1/66621-66535 Candidatus Melainabacteria bacterium
RIFCSPL0W02_12_FULL_35_11 rifcsplowo2_12_scaffold_2329, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GAAATTCCTGGCGATGAAGCTCGCCTTGTCTTAAGTAAAAGACAAAACCACTACTAAAGA AGAGTAGTTGATAGCTTCTACTGAAAA
>CP009278.1/1755095-1755033 Sphingobacterium sp . ML3W, complete genome.
CCTAAATAAGGCAATGATGTCTGCCCTAAACCGCTTTAAGAAGCTGATGACGTCTACTTT
AGA
>FUKX01000060.1/309916-309854 Sphingobacterium faecium PCAi_F2.5 genome assembly, contig: Scaffold3
GCTTAATAAGGCCATGATGTCTGCCCTGAACCGCTTGATAAAGCTGATGACGTCTACTCC
GAT
>MGUR01000059.1/3351-3280 Elusimicrobia bacterium
RIFCSPLOWO2_01_FULL_59_12 rifcsplowo2_01_scaffold_75370, whole genome shotgun sequence.
CTGCTGGACGGCGATGGGGTTCGCCGCTAACCGCCTGAAAGTCCTCTCGAGGCTGATGAC
TCCTGCCAATCC
>CM001195.1/1820945-1820885 Bradyrhizobiaceae bacterium SG-6C chromosome, whole genome shotgun sequence .
AATTTTGATGGGGATGGAGTCCCCCGATACCCGCCGGATGGCTGATGACTCCTACCGGGC
G
>AP011532.1/2428555-2428484 Methanocella paludicola SANAE DNA, complete genome .
ACATACAGCGGCGATGAAGCTCGCCGTGGCTCATGGCAAAACCGCCGCAAGGCTGATAGC
TTCTATTTTAAA
>MLQM01000022.1/24289-24363 Mycobacterium sp . NE-TNMC-100812
PROKKA_contig000022, whole genome shotgun sequence.
ACGAGCACCGGCGATGGGGCCCGCCCGGAAGTCTGACTTCCGAACCGCCGCAAGGATGAT GGCCTCTGCGAACGA
>CP000678.1/1167738-1167808 Methanobrevibacter smithii ATCC 35061, complete genome.
TAATAACTTGGTGATGGGGTTCACCAGAAACTTATTTTCAAACCGCAATTGCTGATAACT
CCTATGTTATA
>CP002919.1/1075339-1075412 Leptospirillum ferriphilum ML-04, complete genome .
ATGACGAGGGGTGATGGAGTCCACCTGAACGGCCCTTTTTCAGACGAGAAAGGGATGATG
ACTCCTGTGAGAGA
>JJNX02000059.1/399-458 Peptococcaceae bacterium SCADC1_2_3 contig_59, whole genome shotgun sequence .
TTTCGGGATGGCGATGGAGTTCGCCACTAAGTCTGTAAGGCTGATAACTCCTACCACTTT
>CP014517.1/1247441-1247513 Variovorax sp. PAMC 28711, complete genome.
TTTTTCGACGGAGATGGCATTCCTTCCGTGAACCGCCCTGGCATCGCCCGGGGCTAATGA
TGCCTACAGACAC
>MHXU01000070.1/3359-3426 Planctomycetes bacterium GWC2_45_44
gwc2_scaffold_20600 , whole genome shotgun sequence.
AAATATATAGGCGATGAGGTTCGCCAAGTAACTGCCTTTATAAAAAGGACGATGACCTCT ACCAACAA
>MKRW01000012.1/13718-13781 Alphaproteobacteria bacterium 62-8
SCNpilot_cont_300_bf_scaffold_1308, whole genome shotgun sequence.
GATGGCGATGGGGATGGGGTTCCCCCGAAACCGCCCTTTCCGGGCTGATGACTCCTGTTA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TTTC
>AMZ001000044.1/254-317 Photobacterium marinum strain AK15 contig44, whole genome shotgun sequence.
TTCATGACGGGTGATGGAGTTCCACCTTTAACCGCTCTACTGAGATGATGACTCCTGCTG
TAGT
>CP014687.1/2546985-2547050 Acetobacter persici strain TMW2.1084, complete genome .
AGACCGCAAGGGGATGGGGTTCCCCCTTTAACCGCCCTTGCAAGGGCTGATGACTCCTGC
CGTAAC
>JXMW01000015.1/11268-11192 Methanobrevibacter arboriphilus JCM 13429 =
DSM 1125 strain DH1 MBBAR_15c, whole genome shotgun sequence.
TTATAATAAGGTGATGAGGTTCACCTATTTAAACTGCCAGAATACTATTTTTTCTGGATG
ATGACCTCTACTAATAA
>GL379781.1/3872744-3872814 Chryseobacterium gleum ATCC 35910 genomic scaffold SCAFFOLD1, whole genome shotgun sequence.
TTAGTAAAAGGAAATGGTGTTCTTCCTTTCCCAACCGCTCCAAAAAAAGAGCTGATGACG CCTGATTAAAT
>CP000780.1/914945-915010 Candidatus Methanoregula boonei 6A8, complete genome .
TCTGTCAGGGGCGATGAAGTCCGCCGTAACCGCCTTACCCTAAAGGCTGATGACTTCTGT
TTTATG
>JRVI01000046.1/170562-170638 Sphingomonas sp. Ant20
Sphingomonas_Ant20_46 , whole genome shotgun sequence.
CTCGGGCACGGCGATGGATTTCCGCCGGGCCATTCGGCCGAACCGCCCTCGCAAGGGCCG ATGATTCCTACCAGGCG
>AAMY01000002.1/145285-145345 Nitrobacter sp . Nb-311A 1099473104484, whole genome shotgun sequence.
GAAATCGATGGGGATGGAGTTCCCCGATACCCGCCGCAAGGCTGATGACTCCTACCGGGC
G
>JRLX01000004.1/77708-77780 Flavobacterium rivuli WB 3.3-2 = DSM 21788 strain WB 3.3-2 contig4, whole genome shotgun sequence.
GTAGTATATGGCAATGATGTCTGCCCTAAACCGCCCGCAAATTTTATTTGTAGCTGATGA CGTCTACTTTAAT
>CP001818.1/1537683-1537744 Thermanaerovibrio acidaminovorans DSM 6589, complete genome.
CAAAGAGAGGGTGATGGGGCTCACCCGAAACTGCTTTTCAAGCTGATGGCCTCTGCGGAC
CT
>MNDG01000032.1/6573-6634 Acidobacteriales bacterium 13_2_20CM_55_8
13_2_20cm_scaffold_3902, whole genome shotgun sequence.
CCTATGGGTGGCGATGGAGTTCGCCATCAACCGCCCACAGGGCTGATGACTCCTACCAGT TC
>MRBZ01000073.1/36257-36173 Nostoc calcicola FACHB-389 FACHB- 389_Scaffold_73, whole genome shotgun sequence.
CAAATCAATGGCGATGGAGCTCGCCGAAACCGCCTTTTCAACCCGATAAATCATCACTTG AAAGGCTGATGGCTCCTACTTTCTT
>CP000822.1/1484941-1485003 Citrobacter koseri ATCC BAA-895, complete genome . APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TTTTTTACAGGAGATGGCATTCCTCCTCATAACCGCCCTCTGGCTGATGATGCCTGCACA
TTC
>MVOL01000018.1/15197-15136 Pseudomonas sp . MF4836 contig018, whole genome shotgun sequence.
ATCAGATCAGGAGATGGCATTCCTCCTTTAACGGCCTTCGTGCTGATGATGCCTACGTGA
AC
>MUI001000009.1/15607-15679 Pseudomonas floridensis strain GEV388 GEV- 388_contig_9, whole genome shotgun sequence.
CGGCGCATTGGAGATGGCATTCCTCCATTAACCAACCGCTGCGCCCGTAGCAGCTGATGA
TGCCTACAGAAAC
>CP002216.1/2041037-2040963 Caldicellulosiruptor owensensis OL, complete genome .
AAATATCAGGGCGATGGAGTTCGCCTTAAAGTGCCACGGGGAATACCCTTGTTGGCTGAT
AACTCCTGTCCGTAT
>MGTY01000072.1/123-190 Elusimicrobia bacterium GWA2_69_24
gwa2_scaffold_3432 , whole genome shotgun sequence.
TAGGCTCAAGGCGATGGAGTTCGCCTGAACCGCCTCGTCCACCCGGGGCTGATGACTCCT ACCAAACA
>CP000319.1/237736-237676 Nitrobacter hamburgensis X14, complete genome.
GACATCGATGGGGATGGAGTTCCCCGATACCCGCCGCAAGGCTGATGACTCCTACCGGGC
G
>BBMI01000017.1/94305-94230 Vibrio ponticus DNA, contig: contig00012_17 , strain: JCM 19238.
TAATCACACGGTGATGGGGTACCACCGGAATCAAACGATTCAAACCGCTTTTAATGCTGA TGACTCCTACAGAAAC
>MELW01000158.1/3682-3618 Alicyclobacillus sp. RIF0XYA1_FULL_53_8 rifoxyal_full_scaffold_2892, whole genome shotgun sequence.
AAATGAATCGGTGATGGAGCTCACCTTAAACCGCCAACGTGTTGGCTAATGGCTCCTGCG AATGG
>CP000909.1/429462-429523 Chloroflexus aurantiacus J-10-fl, complete genome .
AATCGATTGGGTGATGAGGCTCACCCTCAACTGCCATTACGGCTGATAGCCTCTACAGGG
AA
>LMLT01000001.1/571022-571092 Stenotrophomonas sp. Leaf70 contig_l, whole genome shotgun sequence.
TACGTGGTAGGAGATGGCGTTCCTCCTTTAACCGCAGCGATTTCTTAGCTGCTGATGACG
CCTACAGACAC
>JRLV01000009.1/149236-149300 Flavobacterium beibuense F44-8 contiglO, whole genome shotgun sequence .
TTTGGTACTGGCAATGATGTCTGCCTTGAACCGCCCTTAACTGGGCTGATGGCGTCTACA
GTTAC
>MKRX01000024.1/15032-14962 Armatimonadetes bacterium 55-13
SCNpilot_expt_500_bf_scaffold_558 , whole genome shotgun sequence.
CTATTAATAGGGCATGGAGTTCGCCCTACACAAGAACCGCCCAAAGACGGGCTGATGACT CCTACGGTGGC
>LNJC01000025.1/8487-8416 Arc I group archaeon BMIXfssc0709_Meth_Bin006 APG12_contig000025, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ATTCTCTTTGGTGATGAAGTCCGCCAAGGCCCTAGCCTAAACTGCCTTCAAGCTGATGAC
TTCTTTTATCAA
>BATM01000007.1/13923-13990 Vibrio ezurae NBRC 102218 DNA, contig:
VEZ01S07.
AATTGCACAGGTGATGGGGTTCCACCTACTTAACCGCCAAATCATTGGCTGATGACTCCT
ACTAAAAT
>MERA01000302.1/2780-2840 Betaproteobacteria bacterium
RIFCSPLOWO2_02_FULL_65_24 rifcsplowo2_02_scaffold_64491, whole genome shotgun sequence.
CCGCAGCAGGGAGATGGCATTCTCCCCTAACCGCCGCAAGGCTGATGATGCCTACTCGAG
A
>ALVC01000192.1/11830-11760 Sphingomonas sp . LH128 Contigl92, whole genome shotgun sequence.
CGCCGGCAAGGCGATGGATTTCCGCCGGGCTTCGGCCGAACCGCCTTCGGGCTGATGATT
CCTACCTGATG
>KI273145.1/1739300-1739359 Clostridium intestinale URNW genomic scaffold scaffoldl, whole genome shotgun sequence.
TAATTTATAGGTGATGGAGTTCACCTTAATCGCAATTTAGCTAATGACTCCTACAGGTCC
>LMK001000022.1/74417-74345 Xylophilus sp. Leaf220 contig_3, whole genome shotgun sequence.
TTTTTCGACGGAGATGGCATTCCTTCCGTGAACCGCCCTGGCATCGCCCGGGGCTAATGA
TGCCTACAGACAC
>LSBP01000002.1/235766-235677 Halomonadaceae bacterium T82-2 scaffold_l, whole genome shotgun sequence .
TATCGCCCGGGAGATGGCATGCCTCCCGGCGGCACGGTGTTTCCGGCCCGCCAAACCACC
GTACTCACGGTTAATGATGCCTGCGAGGCG
>LDP001000008.1/6838-6915 Mycobacterium heraklionense strain Davo contig_8, whole genome shotgun sequence.
CTGGGGACGCGCGATGGATCTCCGCCGGGGCACTTGAGTGCCCAAACCGCCGTTCTGGCT
GATGGTTCCTGCCCGGAC
>JH976207.1/895295-895360 Acidocella sp . MX-AZ02 genomic scaffold Contigl, whole genome shotgun sequence .
CAGAGGCACGGGGATGGAGTCCCCCCAAAACCGCTCCTTCGGGGAGCTGATGACTCCTGC
ACGGAT
>CP009962.1/5381437-5381372 Collimonas arenae strain Cal35, complete genome .
TGAGAAGCGGGAGATGGCATGCCTCCCCGAACCGCCGCGCAAGCAGCTGATGATGCCTAC
AAGTTC
>MQZY01000028.1/60641-60578 Chromobacterium aquaticum strain CC-SEYA-1 NODE_28_length_65912_cov_30.846569, whole genome shotgun sequence.
GCTTCAGATGGAGATGGCATTCCTCCGCCAACCGCCCGCAAGGGCTGATGATGCCTACAG CGTT
>CP000230.1/293075-293134 Rhodospirillum rubrum ATCC 11170, complete genome .
GACGATACTGGGAATGGGGTCTCCCGAAACCGCCGAAAGGCTGATGACTCCTACCGCGTA
>AZFW01000008.1/50439-50347 Lactobacillus harbinensis DSM 16991 NODE_ll, whole genome shotgun sequence . APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AAGTGAATAGGCGATGACGTTCGCCAGGGGTAAGCAGGTTTTCTGCATACCGCACAAACA
TTACAGTATGTAATTAATGACGCCTACAATATT
>MHZZ01000169.1/3615-3678 Rhodospirillales bacterium
RIFCSPLOWO2_01_FULL_65_14 rifcsplowo2_01_scaffold_42705, whole genome shotgun sequence.
CCCCGCGCCGGTGATGAAGTCCACCGCAACCGCTCCTTCCGGGGCTGATGACTTCTGGCA
GAAC
>LZIT01000178.1/31099-31160 Mycobacterium sp . E2978 contig_259, whole genome shotgun sequence.
TCGCGGATAGGCGATGAAGCTCGCCGTAACCGCCGCACCTGGCTGATGGCTTCTACCCGT
GC
>CP003926.1/1806942-1807005 Gluconobacter oxydans H24, complete genome.
ATTCGGATAGGAGATGGGGTTCCTCCTTTTAACCGCCGCATCGGCTGATGACTCCTGGCA
TGAA
>CP005974.1/365233-365296 Photobacterium gaetbulicola Gung47 chromosome 2, complete sequence.
GGGATGACGGGTGATGGAGTTCCACCTTTAACCGCTCAACCGAGATAATGACTCCTGCTG
TAGT
>MLFS01000010.1/32071-32147 Pantoea wallisii strain LMG 26277 Contig0010_, whole genome shotgun sequence .
TGCGCCTGAGGTGATGGCGTGCCACCTTACCCAACCGCCGTTCTGGTAACCAGACGGCTG
ATGACGCCTGACAACAC
>CP003184.1/651765-651825 Thermoanaerobacterium saccharolyticum JW/SL- YS485, complete genome.
ATTAAGATTGGTGATGGAGCTCACCATTAAATGCGATGATGCTGATGGCTCCTATTGGGT
A
>AOMD01000005.1/1612-1545 Halococcus saccharolyticus DSM 5350 contig_5, whole genome shotgun sequence .
ACGGTATCGGGCGATGGAGTCCGCCTGCCCCAACCGCTGGTTACCCAGCTGATGACTCCT
TTTCGGAG
>CP011862.1/1288559-1288498 Pseudonocardia sp. AL041005-10, complete genome .
GCCACCGATGGCGATGGGGCTCGCCGAGAACCGCACCGCGTGCTGATAGCTCCTACCGAC
GA
>MKVL01000010.1/455530-455592 Mesorhizobium sp . 65-26
SCNpilot_cont_1000_bf_scaffold_20, whole genome shotgun sequence.
GGCCCCGATGGGAATGGGGTTCTCCCGAAACCGCCAGTGATGGCTGATGACTCCTGCCGC GAA
>ATDV01000005.1/32087-32148 Thermoplasmatales archaeon Gpl
AMDU5_GPLC00005, whole genome shotgun sequence.
TTCAAGAAAGGCTATGACGTTAGCCTTTAAGCGCCCTCCGGGCTGATAACGTCTGATCTT TA
>LGJH01000114.1/2787-2717 Novosphingobium sp . ST904 contig_53, whole genome shotgun sequence.
CGCCGGCAAGGCGATGGATTTCCGCCGGGCTTCGGCCGAACCGCCTTCGGGCTGATGATT
CCTACCTGATG
>CP001219.1/2606884-2606947 Acidithiobacillus ferrooxidans ATCC 23270, APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
complete genome.
AAAACGTGTGGCGATGGAGTCCGCCCAATAACCGCCTTTATAGGCTGATGACTGCTACTG
ATCC
>CP017269.1/1008507-1008444 Geosporobacter ferrireducens strain IRF9, complete genome.
AATGGTAAAGGTGATGGAGCTCACCATAAACATTGCTGCGGCAATTAATGGCTCCTACTA
AGGA
>CYHF01000012.1/45411-45480 Thiomonas bhubaneswarensis strain DSM 18181 genome assembly, contig: Ga0061069_112
AACAAATACGGCGATGGAGTTCGCCAAAACCGCGCGGGGCTGTCCCGTAGCTAATGACTC
CTACCTGCCG
>AP014808.1/1618185-1618122 Lactobacillus acetotolerans DNA, complete genome, strain: NBRC 13120.
AATTAAATAGGTGATGACGTTCACCAATAATAAATGGGCAAATTCTAATGACGTCTACTA
TTAC
>MLQM01000022.1/24196-24119 Mycobacterium sp . NE-TNMC-100812
PROKKA_contig000022, whole genome shotgun sequence.
TGGCGGCCAGGCGATGGATCTCCGCCGGGGTGTGTGCACACCCAAACCGCCGACCGGGCT GATGGTTCCTGCTGGAAC
>LWBP01000234.1/1623-1551 Niastella populi strain 208 contig71, whole genome shotgun sequence.
ACTCACAAAGGAAATGGTGTCTTCCTATTTGAACCGTCCGATACACTTCGCGACTGATGG
CGCCTGCAAATAG
>ALAN01000099.1/12847-12908 Bacillus vireti LMG 21834 contig99, whole genome shotgun sequence.
ATGATAAAAGGCGATGGAGTTCGCCATCAACCGTCTTTGTGACTAATGACTCCTACCAGT
GG
>MHFJ01000058.1/3061-3137 Omnitrophica bacterium
RIFCSPHIGHO2_02_FULL_63_14 rifcsphigho2_02_scaffold_4915 , whole genome shotgun sequence.
CCACGTAAGGGCGATGGAGTTCGCCGGATAGTCCTCGTGCTTTTTAGCGATGTGGGAATG
ATGACTCCTACCGAGAA
>FR900143.1/8193-8114 Roseburia sp. CAG:303 genomic scaffold, scf291
ATATTATAAGGGAATGAAGTTCTCCCCTAATGTAATAAACATTAAAACCGCTTATTTAAG
CTGATGACTTCTGCGTAATA
>LZEK01000001.1/154640-154701 Macellibacteroides sp . HH-ZS contig_00001, whole genome shotgun sequence .
AGAAAATTTGGCAATGGAATCTGCCCGAAACCGCCCAAAAGGCTGATGATACCTACCTAA
CT
>MKIP01000044.1/4442-4521 Rhizobium oryzae strain 1.7048 scaffoldl8, whole genome shotgun sequence.
CAAGTTTCCGGTAATGGATTCTGCCTGGCCATGGTGGCCGAACCGCTTCGCAGGATGAAG
CTGATGACTCCTACTCAACT
>AL954747.1/1756266-1756341 Nitrosomonas europaea ATCC 19718, complete genome
TGAAGATAAGGAGATGGTGTTCCTCCTTTTGAAGAAACCGCAGCCGTTTAGCGCTGCTGA
TGACGCCTACAGGACC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>MGTT01000015.1/119197-119258 Elusimicrobia bacterium GWA2_56_46
gwa2_scaffold_209, whole genome shotgun sequence.
GGATATACCGGCGATGGGGTTCGCCATAAACCGCCTCAACGGCTGATGACTCCTATAATC GC
>CP000884.1/1926512-1926446 Delftia acidovorans SPH-1, complete genome.
GCAGCGCAAAGAGATGGCGTGCCTCTAAAACCGCCGCACTAGTGCGGCTGATGACGCCTA
CAGAACC
>AURB01000182.1/9519-9581 Alicyclobacillus acidoterrestris ATCC 49025 contig_76, whole genome shotgun sequence.
AATCGCATCGGTGATGGAGCTCACCATTAACCGCCCGTAAGGGATAATGGCTCCTGCAAG
GAA
>MFGX01000109.1/22374-22303 Candidatus Fraserbacteria bacterium
RBG_16_55_9 RBG_16_scaffold_6090, whole genome shotgun sequence.
GGCAGAAATGGCGATGGAGTTCGCTATAACCACTGCTCGGCAGTTGGAGCAGTTGATAAC TCCTACCTCGTC
>LZRT01000091.1/1028-1087 Bacillus thermozeamaize isolate ZCTH02-B2
BIN02_NODE49710855993, whole genome shotgun sequence.
CGCGAAATAGGCGATGGAGTTCGCCATAACCGCCTGACGGCTGATGACTCCTACCGATAC
>KQ954263.1/126366-126289 Novosphingobium sp . FSW06-99 genomic scaffold PRJNA299315_s013, whole genome shotgun sequence.
GGCGCGCCCGGCGATGGATTTCCGCCGGGCGCTATGCCGAACCACCCGTTTTTTCGGGTT GATGATTCCTACCTTTGG
>LGEW01000001.1/140405-140473 Euryarchaeota archaeon 55_53 PW_scaffold_0 , whole genome shotgun sequence .
GTGTGTATGGGCGATGAGGCTCGCCCCAACCGCCCCACATCCCATGGGGATGATGGCCTC
TTTTATGAG
>CP007806.1/3317303-3317244 Brevibacillus laterosporus LMG 15441, complete genome .
GTAAATAGAGGCGATGGAGTTCGCCCTAAGCGCCAATAGGCTAATGACTCCTACCAGTAA
>MNBX01000082.1/9029-8966 Sodalis sp . TME1 V4_final_sodalis . fa . Contig82 , whole genome shotgun sequence .
GACCCTTCAGGAGATGACATTCCTCCTTATAACCGCCGTTCTGGCTGATGATGTCTACGT
TCGC
>LVXG01000032.1/37818-37748 Niastella yeongjuensis strain 17621 contig38, whole genome shotgun sequence .
TACGCGCAAGGAAATGGTGTCTTCCTACTTAACCGTCCGCCTCTGGCGAGACTGATGGCG
CCTACAAATAC
>FR894597.1/5344-5272 Firmicutes bacterium CAG:308 genomic scaffold, scf43
CTTCTTTGTGGAAATGAAGTTCTTCCATGGGATAACCTAAACCGCTTATTAAGCTGATGA
CTTCTACGAGTTT
>AKZI01000012.1/7720-7798 Rhodovulum sp . PH10 contigl2, whole genome shotgun sequence.
TCGCCCGATGGGGATGGAGTCCCCCGACAACCGCCTCGCCCGGTCGTCCGGTTCCGAGGC
TGACGACTCCTGCCGCGCC
>MGDE01000163.1/2906-2838 Candidatus Schekmanbacteria bacterium
RBG_16_38_10 RBG_16_scaffold_45780, whole genome shotgun sequence.
ATTAGAAACGGCGATGGAGTTTGCCTATAACTGCTTACAAAAATTGAGGCTGATAACTCC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TGCTAAAGG
>AP012337.1/2643146-2643086 Caldilinea aerophila DSM 14535 = NBRC 104270 DNA, complete genome.
ATCGATTCAGGCGATGAGGCTCGCCCTGAACTGCCGTTTGGCTGATAGCCTCTACGGAGA
A
>AEXE01000084.1/9332-9398 Burkholderia sp. TJI49 contig072, whole genome shotgun sequence.
TCTCGCCATGGAGATGGCATTCTCCAAAACCGCTCAAGATTGATGAGCCGATGATGCCTA
CGACCCG
>ASZT02000005.1/28896-28958 Sphingobacterium sp . IITKGP-BTPF85
Sphingobacterium_sp_IITKGP-BTPF85_Contig_00005, whole genome shotgun sequence .
CTATGTTAAGGCAATGATGTCTGCCACTAACCGCTTTAATGAGCTGATGACGTCTACTTT
TGC
>MGSY01000056.1/1894-1962 Deltaproteobacteria bacterium RIFOXYB2_FULL_66_7 rifoxyb2_full_scaffold_12479_curated, whole genome shotgun sequence.
TGGAATCCCGGCGATGGAGTTCGCCGTGAAATGCCTGTGTCGATAGAGGCTGATGACTCC TGCCCGACA
>MUNP01000040.1/35704-35794 Rhodanobacter sp . C06
NODE_l 8_length_93590_cov_8.59036_ID_35, whole genome shotgun sequence.
ATGCATGAGGGAGATGGCATGCCTCCCGCTCCAGCCATGGAGCGAACCGCCGAAGTGCAG
CCGCACCGACGCTAATGATGCCTGCACCACC
>MKWJ01000020.1/46647-46588 Sphingomonas sp . 67-41
SCNpilot_cont_300_bf_scaffold_301 , whole genome shotgun sequence.
TCGCTGGATGGGGATGGAGTCCCCCATAACCGCCGACAGGCTGATGACCCCTACTGCGTC
>ASXJ01000196.1/5605-5664 Ochrobactrum intermedium 229E Contig269, whole genome shotgun sequence.
ATGGGAATGGGGGGGTTCTCCCGAAACTGCCAGCAATTGGCTGATGACTCCTGCTTGAAT
>CP000423.1/2430592-2430656 Lactobacillus casei ATCC 334, complete genome.
AGAAGAACAGGCGATGATGTTCGCCGCAAATGATTGTGTAGCAATCTGATGACGTCTACT
GAAAC
>LHUR01000005.1/89231-89172 Clostridium homopropionicum DSM 5847
CLHOM_contig000009, whole genome shotgun sequence.
ATTAATATAGGTGATGGAGTTTACCTTAACTGCTATAAAGCTAATGACTCCTACAAGTAA
>MNXJ01000057.1/171770-171830 Anaerolineae bacterium CG2_30_57_67
cg2_3.0_scaffold_15_c, whole genome shotgun sequence.
ACTCTTCCTGGCGATGAGGCTCGCCCAAACCGTCTTAACGACTGATAGCCTCTACCGAAA
C
>MAST01000001.1/449973-449899 Humibacillus sp. DSM 29435 contigl, whole genome shotgun sequence.
CGCAAAGCAGGTGATGGATCCCACCTGAGCGACAGCGCCCAAACCGCCCTACCGGCTGAT
GGTTCCTACCGAAAC
>FWXR01000008.1/90116-90050 Fulvimarina manganoxydans strain CGMCC 1.10972 genome assembly, contig: Ga0171588_108
ATGATGGATGGGGATGGAGTCCCCCGATAACCGCCTCGTCTTCGAGGCTGATGGCTCCTA
CCGCACG
>MNDH01000070.1/19596-19671 Verrucomicrobia bacterium APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
13_2_20CM_2_54_15_9cls 13_2_20cm_2_scaffold_933 , whole genome shotgun sequence .
TCCGGCAACGGCAATGGAGTTTGCCTTCCTAGGTAAAACTGGGAAAACCGCGTGAGCTAA
TGACTCCTTCCAAGCG
>AWXH01000002.1/1713938-1713878 Serratia sp . ATCC 39006 ctg2, whole genome shotgun sequence.
ACCAGTAACGGAGATGACATTCCTCCATAACCGCCTTCCGGCTAATGATGTCTACGAGCT
C
>LMZQ01000003.1/592572-592634 Pedobacter ginsenosidimutans strain KACC 14530 contig03, whole genome shotgun sequence.
TGGAACAATGGCGATGGAGTACCGCCAAAACCGCCCAAACAGGCTGATGACTCCTACGAT TTT
>CP003155.1/2504807-2504743 Sphaerochaeta pleomorpha str. Grapes, complete genome .
ATAGTCACTGGTGGTGGAGTTCGCCGTTAACCGCTGTTGTCACAGCCAATGACTCCTGTA
TAAGA
>MHYY01000014.1/1131-1071 Planctomycetes bacterium RIFCSPL0W02_12_38_17 rifcsplowo2_12_subl0_scaffold_1193, whole genome shotgun sequence.
ATTTTAATTGGTGATGGAGTTCGCCTTTAATTGCTAAATAGCTGATAACTCCTATTGAAG
C
>CP003346.1/1855847-1855916 Echinicola vietnamensis DSM 17526, complete genome .
ATGGTACAAGGTGATGGGGTGCCACCTATAACCGCTCTGATTTTTCAGCGCTGATGACTC
CTGCTTCGAC
>AP012342.1/1504277-1504355 Leptospirillum ferrooxidans C2-3 DNA, complete genome .
CGAATCATGGGCGATGGGGTTCGCCTGAACGGCAGCGGTGTTTGATATGACACCGGATGA
TGATGACCCCTACTTTAAA
>AJXT01000008.1/99215-99319 Rhodanobacter spathiphylli B39 contig008, whole genome shotgun sequence .
TTCGCCATGGGAGATGGCATGCCTCCCGTCCCGGAATGCTCGCTCCCCCCAGGACAAACC
GCTTCTCGCGCTCCTCGTGTTGAAGTTGATGATGCCTGCTCCCCC
>MDZE01000018.1/1522-1606 Suitobacillus thermosulfidooxidans strain ZBY contig3, whole genome shotgun sequence.
TAAAATCATTGCAATGGAGTTCTGCTCTCTTTTGGTTGTGACCAAAAGAAAAACTGTCCC
CAGGACTAATGACTCCTACCCATAT
>CP002021.1/1838401-1838324 Thiomonas intermedia K12, complete genome.
TCAACAAACGGCGATGGAGTTCGCCGTGCAATTGCCTCAATAGGGGTTTCGCCCCAGGCT
GATGACTCCTACCCATGT
>AE000666.1/688929-688867 Methanothermobacter thermautotrophicus str.
Delta H, complete genome.
CATCCTGCAGGTGATGGAGTCCACCTTTAACCGCTTTTTCTGGCTGATGACTCCTGCATA
ATA
>BAMU01000018.1/172763-172686 Acetobacter aceti NBRC 14818 DNA, contig: Abac_018.
GACCCATCAGGCAATGGATTCTGCCGGACTTCCCTGGAAGTCGAACCGCCCGTCAGGGCT
GATGATTCCTACCTCGCT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>CP000826.1/1310091-1310029 Serratia proteamaculans 568, complete genome.
TCGAACGACGGAGATGACATTCCTCCTTAACCGCCTTCACGGGCTGATGATGTCTACGTA
ACC
>KB849514.1/275-208 Acinetobacter gerneri DSM 14967 = CIP 107464 genomic scaffold acLZs-supercontl .1, whole genome shotgun sequence.
CAGCTTAAAGGAGATGGCTATCCTCCTTATACAAACCGCCAATTCTGGCTGATGATACCT ACGTTTCC
>LPLU01000048.1/36118-36222 Burkholderia ubonensis strain MSMB782WGS MSMB782WGS_26, whole genome shotgun sequence.
GCGCGCTTCGGAGATGGCATGCCTCCCCGGGGTCCGGGCTTGCTCGCCACGAGCACGCAC AGGCCCTCAACCGCCGGTCGTCCGGCTGATGATGCCTGCGCATTC
>ARYN01000004.1/81966-81902 Zunongwangia atlantica 22II14-10F7 contig4, whole genome shotgun sequence .
AAGAAATACGGCAATGAGGTCTGCCTTAAACCGTTCCACCTGGAGCTAATGACATCTACC
TTTAA
>ALJF01000010.1/188103-188179 Agrobacterium albertimagni AOL15 Contig_10, whole genome shotgun sequence .
GTTTTGTACGGTAATGGATTCTGCCGGGCATCAAGCCAAACCGCTTCCCCGATGAAGCTG
ATGACTCCTACTCAGCG
>MGNT01000094.1/1589-1671 Chloroflexi bacterium RBG_16_57_8
RBG_16_scaffold_30472 , whole genome shotgun sequence.
CTCATCATTGGCGATGAAGCTCGCCCGAGGATATTACCAGCTTCCTGAACCGCCCCAACC AGGCTGATAGCTTCTACTGGCCC
>CP000854.1/4256142-4256216 Mycobacterium marinum M, complete genome.
ACTAGCAGCGGCGATGGGGCTCGCCAGGAAGCTCGACTTCTGAACCGCCAGCCGGCTGAT
GGCACCTGCGAACAA
>FR900948.1/41495-41422 Clostridium sp. CAG:277 genomic scaffold, scflOO
TGCAGAAAAGGGAATGAAGTTCTCCCCCGATTTGATTCGAAACCGCTTTATAAGCTGATG
ACTTCTGTGCTTCA
>JH414691.1/284832-284771 Synergistes sp . 3_l_synl genomic scaffold supercontl .2 , whole genome shotgun sequence.
GCATACGAGGGCGATGGAGTTCGCCCGAAATTCGTTTTTACGATGATGACTCCTGTCCAA
AG
>FRAB01000019.1/15466-15527 Paraburkholderia terricola strain LMG 20594 genome assembly, contig: Ga0075244_1019
TTGCGCGCTGGAAATGGCATTCTCCCTTAACCGCCGTGGCGGCTGATGATGCCTGCTTCG
CC
>JMIY01000007.1/448709-448634 Candidatus Methanoperedens nitroreducens strain ANME-2d ANME2D_Contig_7.7 , whole genome shotgun sequence.
AGGAAATCTGGCGATGAGGTTCGCCTTGGTATTTACCTAAACTGTCTCATTCGAGACTGA TAACTTCTATTTTAAG
>GL379781.1/3119783-3119850 Chryseobacterium gleum ATCC 35910 genomic scaffold SCAFFOLD1, whole genome shotgun sequence.
TATCAAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTACAAAAGCTGATGACGCCT GATTAAAA
>MLQM01000022.1/22919-22858 Mycobacterium sp . NE-TNMC-100812
PROKKA_contig000022, whole genome shotgun sequence.
TTCCGACTAGGCGATGAGGCTCGCCGTAACCGCCGCACCTGGCTAATGGCCTCTACCGAA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TG
>JH114303.1/128052-128113 Desulfovibrio sp . 6_1_46AFAA genomic scaffold supercontl .4 , whole genome shotgun sequence.
TATCTTTACGGGGATGGAGTCCCCCTTGAACCGCATGTATTGCTGATGACTCCTGCCTGA
CG
>LMSL01000040.1/129165-129095 Frateuria sp . Soil773 contig_8, whole genome shotgun sequence.
TCGAACCAGGGAGATGGCATTCCTCCCGGAACAGAAGTAACCGCCGCAAGGCTAATGATG
CCTACCAAACG
>FUWM01000003.1/189103-188987 Selenihalanaerobacter shriftii strain ATCC BAA-73 genome assembly, contig: EI55DRAFT_scaffoldOOOOl .1
AAAAAAATAGGTGATGAAGCTCACCTCTATAACTGAAAAACAAATATAAATTATTAAGTT TATATTTTAAAATCAGTTAAGAAACTGTCTATATAGGCTAATGGCTTCTACTTTCTA
>MGCA01000107.1/222-157 Candidatus Rokubacteria bacterium
RIFCSPHIGHO2_02_FULL_73_26 rifcsphigho2_02_scaffold_183162 , whole genome shotgun sequence.
CGCCGTCAAGGCGATGGGGTTCGCCTGAACCGCCCGACGCGTCGGGCTGATGACCCCTAC
CGGGAG
>MHZC01000276.1/449-510 Planctomycetes bacterium RIFOXYB12_FULL_42_10 rifoxyb3_full_scaffold_20047_curated, whole genome shotgun sequence.
TTACGTCCATGCGATGGAGTTCGCTTAACCGCTATAACATAGCTGATAACTCCTATTGAA GC
>CP000697.1/677221-677283 Acidiphilium cryptum JF-5, complete genome.
TCCGATGCCGGTGATGGAGTTCACCGTGTAACCGCCCGTGGGGCCGATGACTCCTGTTCA
TAT
>ALW002000052.1/297773-297707 Indibacter alkaliphilus LW1
Indibacter_contig_22 , whole genome shotgun sequence.
CAAACTTAAGGTGATGGGGTTCCGCCTAAAACCGCTTTTCTTGAAAGCTGATGACTCCTA CTTCAAC
>CP011801.1/4168483-4168553 Nitrospira moscoviensis strain NSP M-l, complete genome.
TAAATCCACGGTGATGGGGTTCACCGGAACCGCCTTGGAGAGGGTCTGAGGCTGATAACT
CCTACCTTGCC
>MKSP01000341.1/21678-21612 Acinetobacter sp . 38-8
SCNpilot_cont_750_p_scaffold_900, whole genome shotgun sequence.
TAGAGCAAAGGAGATGGCATTCCTCCTATTGCAAACCGCCGTTTTGGCTAATGATGCCTA CGTTACC
>MNVK01000036.1/188-257 Nitrospirae bacterium CG1_02_44_142
cgl_0.2_scaffold_15996_c, whole genome shotgun sequence.
AATTTAAAAGGCGATGGAGTTCGCCGCTAACCGCTCCGATAATATCGGGGCCGATAACTC CTGCCACCCC
>AWXB01000058.1/76022-75953 Peptoniphilus sp . BV3C26 contig00003, whole genome shotgun sequence.
ATATTTTTTGGGAATGAAGTTCTCCCAAGATTTTTCTGAACTGCTATAAGCTGATGACTT
CTGTATTTTA
>CP000360.1/3066978-3066906 Candidatus Koribacter versatilis Ellin345, complete genome.
CAACTTCATGGCGATGGAGTTCGCCCTAACCGCCGCAGCCACGTTGCTGCGTGCTGATAA
CTCCTGCCGACAC
>AZFW01000008.1/44645-44521 Lactobacillus harbinensis DSM 16991 NODE_ll, whole genome shotgun sequence .
TATCGAATCGGTGATGACGTTCGCCGAGCAGCCTTCGGTAAAAGCAGTTGATGTCGCCCT
TTTAAAGGGGCATGAAATGAGGGCTGCAACTATTACGGAATGTAATTAATGACGTCTACA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CGCGC
>CP017749.1/2328568-2328646 Cupriavidus sp . USMAA2-4 chromosome 2, complete sequence.
GTTGACACAGGAGATGGTGTTCCTCCTCTTTTGAAGAAACCGCAGCCGAAGATGCGCTGC
TGATGACGCCTACAGAACC
>GL988000.1/79407-79481 Fusobacterium varium ATCC 27725 genomic scaffold supercont2.6, whole genome shotgun sequence.
TAATACTTTGGGAATGAGGTTCTCCCTTTGTGACAACACAAAAACCGCTTATAAGCTGAT
GACTTCTGCATATTT
>LT592170.1/1353496-1353425 Thiomonas delicata strain DSM 16361 genome assembly, scaffold: BBA-D_THIARS_scaffold4
AAACGAGACGGCGATGGAGTTCGCCGCAACCGCCATCGGCGCCAGCCGGGGGCTGATGAC
TCCTACCCAGTG
>NJIH01000001.1/194242-194308 Candidimonas nitroreducens strain SC-089 Scaffoldl_l, whole genome shotgun sequence.
TGCGCACAAGGAGATGGCATTCCTCCTTCAACCGCTGATCCCATCGGCTGATGATGCCTA
CGAAGCT
>AORV01000025.1/80742-80801 [Clostridium] termitidis CT1112
Ct_contig00028, whole genome shotgun sequence.
ATAAATTCAGGTGATGGAGTTCACCTTTAAATGCATTGAGCTAATGACTCCTACTTTAAT
>LTDM01000011.1/42823-42764 Tissierella creatinophila DSM 6911
TICR_contig000011, whole genome shotgun sequence.
AGAATAGATGGTGATGGAGTTCACCTTTAACCGCATAAAGCTAATGACTCCTACAAAAAT
>CP010024.1/101616-101546 Burkholderia fungorum strain ATCC BAA-463 plasmid pBIL, complete sequence.
GCTCAATCTGGAGATGGCGTTCCTCCTTTAACCATCGCGCTTACCGCGCGGTTAATGACG
CCTACAGTTAC
>GL397067.1/1167953-1167887 Pediococcus acidilactici DSM 20284 genomic scaffold SCAFFOLD1, whole genome shotgun sequence.
AACTAAATAGGCGATGACGTTCGCCGTAAAAATAATTAAACATTAATTTGATGACGTCTA TTATTTC
>FR888464.1/3-67 Eubacterium sp . CAG:274 genomic scaffold, scf34
GGGAATGATGTTCTCCCTAGGACTTTGTCCTGAACTGCCGAAGGGCTGATGGCGTCTGCG
>MGWF01000024.1/12941-13011 Firmicutes bacterium GWF2_51_9
gwf2_scaffold_2190 , whole genome shotgun sequence.
AATACCGACGGGAATGAAGTACTCCCCTAGCTCTGCTAAAACCGCCGAAGGCTGATGACT TCTACGATCTT
>AEXE01000084.1/25781-25851 Burkholderia sp . TJI49 contig072, whole genome shotgun sequence.
GTCGTTACGGGAGATGGCATACCTCCACGAACCGCCGCCCGACCGGTGCGGCTGATGATG
CCTACGGTTCC
>NHOT01000001.1/540633-540695 Pedobacter sp . AJM contigOl, whole genome shotgun sequence.
TGAGCTTAAGGCGATGGGGTACCGCCAAAACCGCCCAGACAGGCTGATGACTCCTACGAT
TTA
>JWTA01000004.1/441586-441653 Chryseobacterium taiwanense strain TPW19 Contig_4, whole genome shotgun sequence.
AATACAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTACAAAAGCTGATGACGCCT
GATTAATA
>AZEM01000127.1/200252-200310 Lactobacillus kefiranofaciens subsp.
kefirgranum DSM 10550 = JCM 8572 strain DSM 10550 NODE_383, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ATGTAAATAGGTGATGACGTTCACCATTAACCGAGAAATCTAATGATGTCTACTTTATC
>MNDI01000050.1/530-606 Verrucomicrobia bacterium 13_2_20CM_2_54_15
13_2_20cm_2_scaffold_1464 , whole genome shotgun sequence.
TCCGCTCAAGGCAATGGAGTTTGCCGTCCCACCTGAAGGCTGGGAAAACCGCGCGAGCTA ATGACTCCTTCCAAACG
>LSTC01000089.1/8506-8442 Nitrospira sp . SCGC AG-212-E16 AG-212-E16.32, whole genome shotgun sequence .
CCGCCTACAGGGGATGGAGTCCCCCGTATAACCGCCTCGGCCAGGCTGATGACTCCTACG
CTACA
>CP000679.1/782565-782643 Caldicellulosiruptor saccharolyticus DSM 8903, complete genome.
GAATATAGAGGCGATGGAGTTCGCCTTAAATTGCCACAAGAATGAAAAATACCTTGTGGC
TGATAACTCCTGTCCGTAT
>BAKI01000030.1/17748-17682 Lactobacillus farraginis DSM 18382 = JCM 14108 DNA, contig: JCM14108. contigO 0030.
AATTGATATGGCGATGACGTTCGCCTAAACTTTATGTCCACAAAGAGTTGATGACGTCTA
CCTTTAA
>KI543089.1/1492-1425 Leptospirillum sp . Group IV &apos;UBA BS&apos;
genomic scaffold D084_Lepto4S00232 , whole genome shotgun sequence.
ATTCCTGGCGGCGATGGGGTTCGCCTTTCCCTGTTCCTGACGGACAGGGTGATGACTCCT ACCATGGG
>MEGM01000093.1/4345-4410 Rubrivivax sp . SCN 70-15 ABT20_C0093, whole genome shotgun sequence.
CGCGGCCCTGGAGATGGCATGCCTCCATGAACCGCCGCGTCAGCGGCTGATGATGCCTGC
GCTGGG
>GG695970.1/494453-494392 Paenibacillus sp . oral taxon 786 str. D14 genomic scaffold supercontl .1 , whole genome shotgun sequence.
ACGCAATAAGGCGATGGAGTTCGCCTAAACCGCTCGTTTGAGCTGATGACTCCTACCAGA CT
>AENN01000001.1/23664-23593 Eremococcus coleocola ACS-139-V-Col8
contig00015, whole genome shotgun sequence.
ATAGAAAAAGGGAATGAAGTTCTCCCTGGCATTAGCCAAACCGCTTGTTAAGCTGATGAC
TTCTGCAACTTA
>KQ959596.1/130248-130327 Clostridiales bacterium KA00274 genomic scaffold ScaffoldlOO, whole genome shotgun sequence.
ACAAATAAAGGGAATGAAGTTCTCCCTAAGTAATATATTTACTTAAAACGCCTATTAAAG
CTGATGACTTCTGCGATTAA
>CYGY02000009.1/35870-35810 Burkholderia sp . STM 7183 genome assembly, contig: CYGY01000009
GCGCGCGCCGGAGATGGCATTCTCCTGAACCGCCCTCGTGGCTGATGATGCCTGCTTCGC
C
>MKI001000015.1/6805-6884 Rhizobium sp. MH17 C775, whole genome shotgun sequence .
CAAGTTTCCGGTAATGGATTCTGCCTGGCCATGGTGGCCGAACCGCTTCGCAGGATGAAG
CTGATGACTCCTACTCAACT
>CP014504.1/3076196-3076132 Pedobacter cryoconitis strain PAMC 27485, complete genome.
TGCTGACACGGAAATGGTGTCTTCCGACTTAACCGCCCTGAAAAGCTGATGGCGCCTGCA
AAATT
>JEMV01000013.1/77002-76926 Sphingobium sp . Ant17 Contig_2057, whole genome shotgun sequence.
ATCGACCTCGGCAATGGATTTCTGCCGGGCAATTTGGCCGAACCGCTCTGACATGAGCTG
ATGATTCCTACTTGGCG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>AKZN01000023.1/24873-24934 Alicyclobacillus hesperidum URH17-3-68
URH17368ctg23, whole genome shotgun sequence.
TAATCAAGTGGTGATGGGGCTCACCAAAACCGCCCGCGAGGGATGATGGCTCCTGCAACG
TA
>LIUQ01000028.1/83647-83743 Planktothricoides sp . SR001 contig_114, whole genome shotgun sequence.
AATAATGATGGCGATGGGGCTCGCCAAAACCGCCGTAAGTTTGCCAAGAATTTGCCAAGA
ATTTGACCAGACCAAGGCTGATGGCTCCTACTATTCC
>GL637644.1/113943-114016 Solobacterium moorei F0204 genomic scaffold Scfldl5, whole genome shotgun sequence.
ATTTCTGACGGAAATGGGGTTTTCCAAGATTATAAATCTAAACTGCTATAAGTGCTGATA
ACTCCTGTAACTTC
>JAOC01000010.1/320154-320228 Mycobacterium xenopi 3993 GMX3993. contig .9, whole genome shotgun sequence .
ACTTGCACCGGCGATGGAGCCCGTCCGGAAGTGCGACTTCCGAACCGCCACAAGGCTGAT
GGCTTCTACGAACGA
>JTDJ01000052.1/49642-49539 Mumia flava strain MUSC 201 Contig52, whole genome shotgun sequence.
GCGGCCAGCGGAGATGGCATGCCTCCCTGCGGTCGTAGTCGGCGCGTTGCGCACGGTGCG
GCCCTCAACCGCCGGTCAGCCCGGCTGATGATGCCTGCGTGTTC
>AM889285.1/3849897-3849819 Gluconacetobacter diazotrophicus PA1 5 complete genome
GTCGGCGCAGGCAATGGATTTCTGCCTGACCGTCCGGGTCGAACCGCCTCTTTCGGGGGC
TGATGATTCCTACCCACCG
>LMHV01000041.1/1557-1633 Rhizobium sp. Root708 contig_6, whole genome shotgun sequence.
ACCGGGCACGGTAATGGATTCTGCCGGGCCGTGGTGGCCGAACCGCTTGCTTGCAAGCTG
ATGACTCCTACTCGTGA
>KQ959855.1/1372-1303 Gemella asaccharolytica strain KA00071 genomic scaffold Scaffoldl, whole genome shotgun sequence.
AGTACTATTGGGAATGAAGTTCTCCCAAGAAAAATAAATGCTTATTTTAGCTGATGACTT CTACAAATTT
>CP017269.1/4333368-4333435 Geosporobacter ferrireducens strain IRF9, complete genome.
AGAAAAATCGGTGATGGAGCTCACCTTTAAGCATTGCAGGATATGCAATTAATGGCTCCT
ACTTGGCC
>BAFN01000001.1/1082359-1082420 Candidatus Brocadia sinica JPN1 DNA, contig: brosiA.
AATATTTAGAGCGATGGAGTTCGCTTAACTGCTACATTGTAGCTGATAACTCCTGTTGAA
GT
>FR893892.1/35353-35285 Bacteroides sp. CAG:714 genomic scaffold, scfl02
CTGCTAAAAGGGAATGGGACTTCCCTGTCATAAACCGCTAAGAAAATAGCTGATAGTTCC
TGTCGTGAG
>ALXI01000102.1/10533-10597 Clostridium sp . Maddingley MBC34-26
contig_203_1 , whole genome shotgun sequence.
AATTTCATAGGTGATGAAGTTCGCCTTTAAACATCTCTTTGGAGATTAATGACTTCTACT
ATAAT
>CP003130.1/3651628-3651693 Granulicella mallensis MP5ACTX8, complete genome .
AGCATTCATGGCGATGGGGTTCGCCTGCCGCTCCGCACGCGCGTAGCTGATGACTCCTAC
ATCTCG
>CP000930.2/1203106-1203045 Heliobacterium modesticaldum Icel, complete genome . APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AAGGGAATACGCGATGGAGTTCGCGGTAAAGCCGTTTTGCGGATGATGACTCCTACTCGA
TT
>MHYC01000102.1/4812-4874 Planctomycetes bacterium RBG_13_44_8b
RBG_13_scaffold_73617 , whole genome shotgun sequence.
AATTTCTGCGGAGATGGAGTTCTCCGTATAACCGCCCAAAAGGCTGATAACTCCTGCAGT ATT
>CP002606.1/1302749-1302816 Hippea maritima DSM 10411, complete genome.
CAGCTTAAAGGCGGTGGAGTTCGCCTTAAACTGCCCTCTCTCTTTGGGCCAATAACTCCT
GCAAAAAC
>DS996361.1/121355-121416 Desulfovibrio piger ATCC 29098 ScfldlO genomic scaffold, whole genome shotgun sequence.
GAAGGTCAAGGGGATGGAGTCCCCCATGAACCGCATGTTTTGCTGATGACTCCTGCCGGA
CC
>CH478194.1/77118-77180 Aedes aegypti strain Liverpool supercontl .1010 genomic scaffold, whole genome shotgun sequence.
CCGACAGACGGAGATGACATTCCTCCATAACCGCCCTCACCGGCTGATGATGTCTACGTA ACC
>JPLB01000025.1/17289-17224 Luteibacter rhizovicinus DSM 16549 Contig_25, whole genome shotgun sequence .
CCGTCCACAGGAGATGACATAACTCCTGACCCAACCGCCGCCGTGGCTGATGATGCCTAC
CATATG
>CBML010000006.1/1375951-1375890 Clostridium chauvoei JF4335, WGS project CBML01000000 data, contig: 00001
TATTATATAGGTGATGGAGTTCACCATTAAATGCTTAAGAAGCTAATGACTCCTACTCTT
TT
>CBLY010000006.1/104416-104504 Saccharibacter sp . AM169, AM168, WGS project CBLY01000000 data, contig: NODE_2
AGAGGGAACGGTGATGGATTGCCACCGGGTTCCTTGTGGGCCGAACCGCCACCCGGTTGC
CGGGGAAGGCTGATGATTCCTGCCAACAG
>CP015436.1/780010-779951 Anoxybacillus sp . B7M1, complete genome.
AATTGAAGAGGCGATGGAGTTCGCCGTAATTGCCTTTGAGCTGATGACTCCTACCGGTAT
>MEFR01000172.1/7376-7315 Acetobacteraceae bacterium SCN 69-10
ABS99_C0172, whole genome shotgun sequence.
GATTGACGAGGGGATGGAGTCCCCCTTGAACCGCCGCAGAGGCTGATGACTCCTGTCGCG
CA
>MKSH01000238.1/6187-6124 Solirubrobacterales bacterium 70-9
SCNpilot_bf_inoc_scaffold_5980, whole genome shotgun sequence.
CGAAACTGAGGCGATGAGGCTCGCCCTCGACCGCGGCGCCGCCGCTGATGGCCTCTACGG AACG
>BAMY01000027.1/11025-11090 Acetobacter orleanensis JCM 7639 DNA, contig: Abol_030.
AAGCCGCAAGGGGATGGGGTTCACCCTTTAACCGCCCCAGTAAGGGCTGATGACTCCTGC
CGTATC
>BAEH01000015.1/34807-34726 Gordonia effusa NBRC 100432 DNA, contig:
GOEFSO 15.
GCTGTTTTCGGCGATGGATCTCGCCTGGTCCGATTCTTCGGCCCGAACCGCCGGCTCTCC
GGCTGATAGTTCCTCCTCGTAT
>AYKS02000112.1/26161-26223 Serratia sp . DD3 SRDD_contig000112, whole genome shotgun sequence.
CTGACAAACGGAGATGACATGCCTCCATAACCGCCCCATTAGGCTGATGGTGTCTACGCT
ACC
>LOHZ01000015.1/55129-55190 Thermovenabulum gondwanense strain R270 ATZ99_contig000016, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AAAATCCAAGGTGATGGAGCTCACCGAAAACCGCCTTTTAGGCTAATAGCTCCTACGTAT
CG
>LDJP01000006.1/10388-10323 Stenotrophomonas daejeonensis strain JCM 16244 contig_6, whole genome shotgun sequence.
CCGGGCAACGGAGATGGCATTCCTCCGCGAACCGCCGCGACAGCGGCTGATGATGCCTAC
GCTTGG
>CP002216.1/521273-521336 Caldicellulosiruptor owensensis OL, complete genome .
ATTGTAAAAGGCGATGGAGTCCGCCGTACAAATGCCAATGATGGCTGATGACTCCTACAG
ATGC
>MKRK01000154.1/4187-4115 Clostridiales bacterium 43-6
SCNpilot_BF_INOC_scaffold_3713, whole genome shotgun sequence.
AATACCTGCGGGAATGAAGTACTCCCTTGGGTAACCTAAACCGCTTATTGAAGCTGATGA CTTCTGTGATTTT
>JSYN01000016.1/81103-81041 Pedobacter kyungheensis strain KACC 16221 contig016, whole genome shotgun sequence.
TGAAATCAGGGCGATGGGGTGCCGCCAAAACCGCCAAAACAGGCTGATGACTCCTACGAT
TTT
>FQXE01000023.1/6418-6352 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_123
TGCACCGACGGAGATGGCATTCCTCCGCTAACCGTCGCTATCAGCGACTGATGATGCCTA
CACCTTC
>HG422565.1/3993555-3993621 Candidatus Microthrix parvicella RN1 genomic scaffold, 2605
GCCGTTGGTGGCGATGGGGCTCGCCCTCAACTGCCGGACACCTCCGGCTGACGGTCCCTG
CGAACGA
>LMFS01000001.1/3352987-3353048 Pseudolabrys sp . Rootl462 contig_l, whole genome shotgun sequence.
CCGCACGATGGGGATGGGGTCCCCCGATAACCGCCGGAATGGCTGATGACTCCTGCCGGA
TG
>JMPR01000012.1/20844-20921 Tatumella ptyseos ATCC 33301
GTPT . assembly .012 , whole genome shotgun sequence.
TTGTTACGGGGTGATGGTGTTCCACCTCAACCAACCGCCAGTCCGTTTCATGGAATGGCT GATGACGCCTGACATAAT
>FQVU01000002.1/320536-320617 Jatrophihabitans endophyticus strain DSM 45627 genome assembly, contig: Ga0131187_102
GTCGGGCGAGGCGATGGATCCCGCCGGGGTCGTCCGATCGGACGACCCGAACCGCCGGCC
GGCTGATGGTTCCTTCTCACGA
>LUUM01000312.1/8605-8667 Methylosinus sp. R-45379 contig_92, whole genome shotgun sequence.
GGCGCGCCAGGGGATGGGGTGCCCCCTACAACCGCCGAAGAGGCTGATGACTCCTGCTCA
ATG
>MGQS01000129.1/9858-9922 Deltaproteobacteria bacterium RBG_16_54_11 RBG_16_scaffold_77520, whole genome shotgun sequence.
ATACTAAGAGGGGATGGAGTTCCCCTGACAATTGCCCAGGCCGGGCTGATGACTCCTACC GCGCA
>CP002273.2/3698397-3698455 Eubacterium limosum KIST612, complete genome. TGGTCTCAGGGCGATGGAGTTCGCCTAAACCGCATAATGCTGATGACTCCTGTTTCTTA
>MEEQ01000471.1/3206-3146 Bacterium SCN 62-11 ABS71_C0471, whole genome shotgun sequence.
AGCTCCAATGGGGATGGAGTCCCCCCAAACCGCCTCCCCGGCTGATGACTCCTACCACCT
T
>KQ965575.1/5893-5824 Clostridiales bacterium KA00134 genomic scaffold APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
Scaffold22, whole genome shotgun sequence.
ATGTTTAGAGGAAATGAGGTACTTCCTTGGGAAACCTAAATGCAAATATGCTGATGACTT
CTACTTTTGT
>MICJ01000091.1/164-127 Syntrophobacterales bacterium GWC2_56_13
gwc2_scaffold_5779, whole genome shotgun sequence.
GGGGTTCACCGCAACCGCCGCTCGGCTGATAACTCCTA
>CP014168.1/2698639-2698565 Sphingomonas panacis strain DCY99, complete genome .
GTGCGCCACGGTGATGGATTTCCGCCGGGCATTCGCCGAACCGCTCTCACAAGAGCTGAT
GATTCCTACTGGGCG
>AJYA01000016.1/88384-88450 Nitritalea halalkaliphila LW7 Contigl6, whole genome shotgun sequence.
TTTGTTTCAGGTGATGGGGTTCCACCTAAAAAAAACCGCCGTATCGGCTGATGACTCCTA
CACAGTA
>AM180355.1/782179-782255 Clostridium difficile 630 complete genome
AATAATTATGGGAATGAGGTTCTCCCTTGGATGATAATATCCTAAACTGCCAATGAGCTG
ATGACTTCTATGATATC
>CP002160.1/406583-406525 Clostridium cellulovorans 743B, complete genome. AAATTTATAGGTGATGGAGTTCACCTTAATCGCTTGATGCTAATGACTCCTACAGGTAA
>MIHC01000025.1/13855-13928 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
TGATCGTGAGGCGATGGATCTCCGCCGGGGTATTGCACCCGAACCGCCAACCGGCTGATG GCTCCTTCGACGTC
>MHZZ01000169.1/3491-3555 Rhodospirillales bacterium
RIFCSPLOWO2_01_FULL_65_14 rifcsplowo2_01_scaffold_42705, whole genome shotgun sequence.
TCGGCACACGGGGATGGAGTCCCCCGCTTGTAACCGCCGAAGAGGCTGATGACTCCTGCT
GCGCG
>DS990260.1/2249356-2249281 Clostridiales bacterium 1_7_47_FAA
supercont2.1 genomic scaffold, whole genome shotgun sequence.
TGTTTACAGGGGAATGAGGTTCTCCCTTAGTGTCATACTAAAACCGCTTATTTAAGCTGA TGACTTCTGTGTGATT
>KK073875.1/1423111-1423050 Saccharibacillus sacchari DSM 19268 genomic scaffold SacsacDRAFT_Scaffoldl .1, whole genome shotgun sequence.
GTTCAATACGGCGATGGAGTTCGCCATAACCGCCTTTCGGGGCTAATGACTCCTACCCAT GA
>LOED01000018.1/32047-31987 Fervidicola ferrireducens strain Y170
AN618_contig000018, whole genome shotgun sequence.
AAGTGAAAAGGTGATGGAGCTCACCGAAACTGCCTTTCAGGCTAATAGCTCCTACAATAT
C
>CP000724.1/4518368-4518305 Alkaliphilus metalliredigens QYMF, complete genome .
ATGCTCATAGGCGATGGAGTTCGCCATTAACCACTGGATAACAGTTAATGACTCCTGCGA
AACA
>LMGN01000024.1/146807-146886 Rhizobium sp . Root564 contig_9, whole genome shotgun sequence.
CAAGTTTTTGGTAATGGATTCTGCCGGGCCATGGTGGCCGAACCGCTTCACGGGATGAAG
CTGATGACTCCTACTCAATT
>LQYN01000018.1/1414-1355 Bacillus sporothermodurans strain B4102 NODE_20, whole genome shotgun sequence .
AATGTAAAGGGCGATGGAGTTCGCCAGATGTGCTACTAGGCTAATGACTCCTACCGGAAA
>AAWL01000036.1/2586-2647 Thermosinus carboxydivorans Norl ctg23, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TAAGTTTGCGGCGATGGAGTTCGCCATAAGTCCGCATAAGCGGTAATGACTCCTACCAGG
AC
>LWSA01000079.1/6883-6945 Acidithiobacillus thiooxidans strain A02 contig079, whole genome shotgun sequence.
AATTGACATGGTGATGAAGTCCACCAAATAACCGCCATTGTGGCTGATGACTTCTACATG
ACT
>FPIR01000010.1/55645-55730 Ruminococcus sp . YE71 genome assembly, contig: IE40DRAFT_scaffoldOOOlO .10
TCTCATATCGGGAATGATGTTCTCCCGCATACGGCTCCGAGCGAGCCTGTGAAACCGCTT
TCAAAGCTGATGACTTCTGCGTTCCT
>DF968182.1/1455626-1455563 Bacteroidales bacterium TBC1 DNA, scaffold: TBCl_scaffold_l .
TTTGCCGCAGGCTATGGGGTTAGCCTTTAACCGCCTTCAACAAGCTGATAACTCCTGGAT
GATT
>CP006763.1/431726-431785 Clostridium autoethanogenum DSM 10061, complete genome .
TATTTTATAGGTGATGGAGCTCACCTTTAATTGCTTATGGCTGATAGCCCCTACTTTAAT
>AE009950.1/1333263-1333185 Pyrococcus furiosus DSM 3638, complete genome.
GTTCTTTGGGGTGATGGCGCCGCCCGGGGCTTCGAGCCGAACCGCCTCCCGTTGGGAGGC
TGATGGCGCCTATTCTGTA
>ACJN02000001.1/209444-209506 Desulfonatronospira thiodismutans AS03-1 ctg21, whole genome shotgun sequence.
CGCCTGAAAGGGGATGGAGTTCCCCCTGGTAACCGCCGTGATGCTGATGACTCCTGTTGT
TTT
>NJIH01000001.1/186315-186381 Candidimonas nitroreducens strain SC-089 Scaffoldl_l, whole genome shotgun sequence.
GGCCGCAAAGGAGATGGCATTCCTCCTGTAAACCGCCACCATACCGGCTAATGATGCCTA
CCTGTTC
>MIDC01000012.1/36114-36189 Tenericutes bacterium RIFOXYB2_FULL_36_25 rifoxyb2_full_scaffold_21, whole genome shotgun sequence.
CATAGAAGAGGGAATGAACGTCTCCCTTAGTCAACCATGACTAAAACCGCAACATGCTAA TGACTTCTACAACACG
>CP007516.1/143851-143917 Rubrobacter radiotolerans strain RSPS-4 plasmid 2, complete sequence.
AGAGCAACGGGTGATGAAGCTCACCTCTATAACCGCATCGAAAGATGCTGATGGCTTCTA
CCAGGGC
>CP003326.1/2952530-2952589 [Clostridium] acidurici 9a chromosome, complete genome.
AAAACTCACGGCGATGGAGTTCGCCATTAAATGCGTAAAGCTGATGACTCCTACAAAGTA
>AP012551.1/1598812-1598881 Plautia stali symbiont DNA, complete genome.
GGACGGCAGGGTGATGGCGTTCCACCTTTCCCAACCGTTCCGTTCCGGAACTGATGACGC
CTGATGTGAC
>CP003926.1/1813274-1813336 Gluconobacter oxydans H24, complete genome.
TCGTGTAAGGGAGATGGAGTTCCTCCCGTAACCGCCCTAAGGGCTGATGACTCCTGCCGA
AGC
>CP001358.1/1325490-1325548 Desulfovibrio desulfuricans subsp.
desulfuricans str. ATCC 27774, complete genome.
TTGAACGCAGGGGATGGAGTTCCCCTTGAACCGCATTTGCTGATGACTCCTGCCAGACC
>HF995081.1/61766-61694 Coprobacillus sp . CAG:235 genomic scaffold, scf229
TTAATCCATGGGAATGAGGTTCTCCCTCGATTTATATCGAAACCGCTATTTGGCTAATGA
CTTCTGTGTAATA
>LSRN01000054.1/3034-3102 Alkalibacterium sp . 20 contig48, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TAACAAATAGGCGATGGTGTTCGCCTTTAACTATCATAAATAGATATGATTAATGACGCC
TACTTAGAC
>MEFT01000106.1/933-1004 Clostridium sp . SCN 57-10 ABT01_C0106, whole genome shotgun sequence.
ACAAAGCAAGGGAATGAAGTTCTCCCGCGGGCGCTGCCCGAAACCGCTTTTGCTGATGAC
TTCTGCAATTTG
>AYZJ01000062.1/6126-6187 Lactobacillus camelliae DSM 22697 = JCM 13995 strain DSM 22697 NODE_98, whole genome shotgun sequence.
AAGTGAAATGGTGATGACGTTCACCCTAATACAAGCTAATTGTTGATGACGTCTGTCTTA AT
>KK106988.1/2432078-2432141 Streptomyces sp . Tu 6176 genomic scaffold scaffoldOOOOl, whole genome shotgun sequence.
GACGGAGACGGCGATGAGGCTCGCCCTTGACCGCACACCCCGTGCTGATGGCCTCTGTGA
ACGA
>CP002400.1/470084-470160 Ethanoligenens harbinense YUAN-3, complete genome .
CGCATATCGGGGAATGAAGTGCTCCCCGGGCGTTTGCCATAACTGCCAATACACAGGCTG
ATGACTTCTGTGATCGA
>LAD001000039.1/63932-63875 Peptococcaceae bacterium BRH_c4b BRHa_1004828, whole genome shotgun sequence .
TATAATTTTGGCGATGGGGCTCGCCTAATGCTGCGAAGCTGATGGCTCCTACCACAAG
>MNYH01000071.1/4-55 Deltaproteobacteria bacterium CG2_30_66_27
cg2_3.0_scaffold_l 0978_c, whole genome shotgun sequence.
CGATGGAGTTCGCCGTGAACCGCCCGAACGGGCTGATGACTCCTGGTTGTAT
>HG322950.1/4730010-4729948 Pseudomonas knackmussii B13 complete genome
CGCGATTCAGGAGATGGCATTCCTCCTTTAACCGCCCCTGGGGCTGATGATGCCTACGCA
TTC
>LCVM01000141.1/4442-4374 Providencia rettgeri strain MR4
P_rettgeri_contig_141 , whole genome shotgun sequence.
AAACCTTTGGGAGATGGCATTCCTCCTTATATAAAACCGCCCGTAGAGGCTGATGATGCC TACGTTAAC
>AE017198.1/798633-798695 Lactobacillus johnsonii NCC 533, complete genome .
ATTTAATATGGTGATGGTGTTCACCAATTTAACCGATTAATATCTGACGACGCCTACTTT
CTT
>AE005176.1/1105212-1105286 Lactococcus lactis subsp. lactis 111403, complete genome.
ATAAATGATGGGTATGGTGCACACCCGAAACCGCTTTAAGAATAAAATCTTAAAACTAAT
GGCGCCTACAAACAA
>CP003943.1/2856403-2856485 Calothrix sp . PCC 7507, complete genome.
AACGACTATGGCGATGGAGCTCGCCAAAACCGCCTCTTCAGTAAAAGATTTATAACTGCA
AGGCTGATGGCTCCTACTTTTCC
>DS981450.1/8984-9052 Bacteroides coprocola DSM 17136 Scfld_02_23 genomic scaffold, whole genome shotgun sequence.
CAAAGTCAGGGGAATGGGACTTCCCTGTTATAAACCGCTAATAAAGTAGCTGATAGTTCC
TGCCGGAGT
>AOCG01000006.1/242070-242010 Listeria aquatica FSL S10-1188 c5, whole genome shotgun sequence.
TACTGAAAAGGTGATGGAGTTCGCCTAAAATGTGATATTCGCTGATGACTCCTATTTAAA
T
>AAXG02000029.1/31858-31932 Bacteroides capillosus ATCC 29799
B_capillosus-2.0. l_Cont263 , whole genome shotgun sequence.
GTGAATCAGGGGAATGAAGTGCTCCCCGGGGCGTGGCCCCGAACCGCGTTTGACGCTGAT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GACTTCTGCATTTTG
>MWUE01000006.1/31346-31270 Pantoea sp. AS1 5_len_86521, whole genome shotgun sequence.
ACGCCTTAAGGTGATGGCGTTCCACCTTTGCCAACCGCCTGCTGGCGATCCCAGCGGCTG
ATGACGCCTGACATTAA
>LGGP01000002.1/11694-11617 Mesotoga prima MPI_scaffold_981 , whole genome shotgun sequence.
AATTCTCTAGGCAATGAAGCTTGCCTTGGAGAACAGTCTTCCGGAACCGCTTTTTTGGCT
GATAGTTTCTGCGATTAT
>AE000782.1/1391410-1391471 Archaeoglobus fulgidus DSM 4304, complete genome .
ATCGGTATGGGTGATGGCGTCCACCTTCAACCGCCGCTAAGGCTGATGACGCCTGTTCTC
AG
>CP003119.1/3751759-3751684 Gordonia polyisoprenivorans VH2 , complete genome .
TCGGCCGCGGGTGATGGATCTCGCCCGGGCCGCGCCGGCCCGAACCGCCGACAAGGCTGA
TGGCTCCTGTTCCGAT
>MHYT01000009.1/1931-1991 Planctomycetes bacterium RIFCSPHIGH02_12_39_6 rifcsphigho2_12_subl0_scaffold_10301, whole genome shotgun sequence.
ATTTTAATTGGTGATGGAGTTCGCCTTTAATTGCTAAATAGCTGATAACTCCTATTGAAG
C
>CP002606.1/544616-544551 Hippea maritima DSM 10411, complete genome.
TCAAAAAAAGGCTATGAGGTTAGCCTAAAAAGCCCCGCCAAAGGGGATGATAACCTCTAC
TCGATA
>CP002435.1/144988-144917 Pantoea sp . At-9b plasmid pPAT9B02, complete sequence .
TGAGTCGAAGGTGATGGCGTTCCACCTTATCCAAACCGCTCCCTCGCTGGAGCTGATGAC
GCCTGGTATGTC
>CBIH010000205.1/2901-2971 Ruminococcus sp . CAG:330, WGS project
CBIHOIOOOOOO data, contig: 9
TTCAAGATCGGGAATGATGTTCTCCCCTGGGATTCCAAAACCGCTGTTTCGCTGATGACG
TCTGCTTTTTT
>AMGM01000055.1/23054-22988 Cecembia lonarensis LW9 contig000055, whole genome shotgun sequence.
CAAGATCTAGGTGATGGGGTACCACCTTTAACCGCCCTTTTTTGGTGCTGATGACTCCTG
CTTCAAC
>MHXS01000153.1/2999-2938 Planctomycetes bacterium GWB2_41_19
gwb2_scaffold_6965 , whole genome shotgun sequence.
TTACGTCCATGCGATGGAGTTCGCTTAACCGCTATAACATAGCTGATAACTCCTATTGAA GC
>MPHG01000248.1/37571-37630 Bacillus sp . VT-16-70
Bacillus_obstructivus_contig232 , whole genome shotgun sequence.
TAAATATGCGGCGATGGAGTTCGCCCTTAAGTGCTTAAAGCTAATGACTCCTACCCTATT
>MFPR01000009.1/26239-26178 Candidatus Lindowbacteria bacterium
RIFCSPL0W02_12_FULL_62_27 rifcsplowo2_12_scaffold_1449, whole genome shotgun sequence.
AGATGTGCAGGCGATGGAGTTCGCCTTAACCGCCCCTACCGGCTGATGACTCCTACCACA
TG
>MKSU01000016.1/336643-336581 Sphingomonas sp. 67-36
SCNpilot_cont_300_bf_scaffold_251 , whole genome shotgun sequence.
CATGGTGACGGGGATGGAGTTCCCCGATAACCGCCGTTCCGGGCTGATGACTCCTACCAA CAC
>DF820472.1/159652-159588 Bacterium UASB270 DNA, scaffold: APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
UASB270_scaffold_l 0.
GCCTTTCGAGGGGATGGAGTTCCCCTGACAAATGCCCGGCAAGGGCTGATGACTCCTGCG
ATGCG
>MVDC01000122.1/5278-5350 Desulfobacteraceae bacterium 4484_190.3 ex4484_190.3_scaffold_20784 , whole genome shotgun sequence.
TAGCAATAAGGGGATGGAGTCCCCCGCATAATTGTCCAACCGCTTCAAAAGGGCTAATGA CTCCTGTCGCGTG
>LKUC01000117.1/3567-3500 Smithella sp. SDB scaffold-15358, whole genome shotgun sequence.
ACATGAATAGGTGATGGAGTTCACCTAAGATTAACCGCTTTTATATAGCTGATAACTCCT
GCAGAAAA
>MAST01000002.1/541-622 Humibacillus sp . DSM 29435 contiglO, whole genome shotgun sequence.
GTCGGAGACGGCGATGGATCCCGCCTGGGCGTGCCGGCTTCACCCCGAACTGCCAGCTCC
GGCTGATGGTTCCTGACTCGTC
>CP000628.1/2724502-2724415 Agrobacterium radiobacter K84 chromosome 1, complete sequence.
AAATCCAACGGTAATGGATTTCTGCCGGGCCATGGTGGCCGAACCGCTGCTCTGAAGATT
CAGGGAAGCTGATGACTCCTACTCAACG
>MEFA01000007.1/60862-60923 Pseudonocardia sp. SCN 72-86 ABS81_C0007, whole genome shotgun sequence .
TCAGCACACGGCGATGGGGCTCGCCTTTGACCGCCTAACCGGCTGATAGCTCCTACCAAC
GG
>FQZG01000024.1/3478-3557 Tessaracoccus bendigoensis DSM 12906 genome assembly, contig: EK12DRAFT_scaffold00022.22
GTTCGACCCGGTGATGGACCCCGCCGGGGCCGGTTGATCCGGCCCGAACCGCCTCTTTGG
CTGATGGCTCCTGTCCACCG
>HF998978.1/5102-5031 Blautia sp . CAG:257 genomic scaffold, scf4
AATGAATTTGGGAATGAAGTTCTCCCAAGGGAAACCTGAACTGCTTATTAAGCTGATGAC
TTCTGAGATTAT
>MKSC01000083.1/2627-2690 Shinella sp . 65-6
SCNpilot_expt_1000_bf_scaffold_14020, whole genome shotgun sequence.
CCGTGGGATGGGGATGGAGCTCCCCGACAACTGCCGTTCATCGGCTGATAGCTCCTGCCA
GACC
>LOPU01000003.1/23382-23314 Halobacteriaceae archaeon SB9 contig_ll, whole genome shotgun sequence.
TATCACTCGGGCAATGGAGTCTGCCCCACCCAACCGCCGAACACGTCGGCTGATGACTCC
TTCTCGTCA
>BAJS01000008.1/110947-110883 Bacteroides graminisolvens DSM 19988 = JCM 15093 DNA, contig: JCM15093. contigOOO 08.
TCCGATTTCGGCAATGGATTCTGCCTTTAACCGCCCTTCGTGGGGCTGATGACTCCTACT
TGGAA
>FR890736.1/26782-26711 Ruminococcus sp . CAG:57 genomic scaffold, scf82
CTCATATCGGGGAATGATGTTCTCCCCTGGGAAACCTAAACCGCTTATTAAGCTGATGAC
TTCTGTTTTTTA
>CBXV010000007.1/153141-153206 Pyrinomonas methylaliphatogenes , K22, WGS project CBXV01000000 data, contig: PYK22_K22_C11
AAAGTTCAGGGTGATGGAGTCCACCATCAACCACTGCGATCGGCAGTTGATGACTCCTAT TTCGAA
>MIDS01000081.1/1568-1491 Verrucomicrobia bacterium
RIFCSPL0W02_12_FULL_64_8 rifcsplowo2_12_scaffold_126916, whole genome shotgun sequence.
GTTTGGCCCGGCGATGGATTCCGCCTCCCCGCGAGGGAAAACCGCCGCGCTTGCGTGGCT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AATGACTCCTGTTGATCC
>FJNB01000023.1/1869-1802 Trichococcus sp. R210 isolate Trichococcus_R210 genome assembly, contig: TR210DRAFT_scaffold-23
TAGTATGTCGGCGATGGTGTTCGCCTTTAACTATCATAATAACTATGATTAATGACACCT ACTTAAAT
>AFVJ01000019.1/16616-16691 Mycoplasma anatis 1340 contig26, whole genome shotgun sequence.
AATAATAACGGGAATGATGTTCTCCCTGTTGGATTAAACAAACCGCATTTTTAATGCTAA
TGACGTCTACGACTTT
>CP000557.1/2739339-2739277 Geobacillus thermodenitrificans NG80-2, complete genome.
AACATACTAGGCGATGGGGTTCGCCATAACCACCGGCTTCCTGCTGATGACTCCTGCTGC
GTT
>CP000682.1/1397941-1398019 Metallosphaera sedula DSM 5348, complete genome .
TAGGTTACTGGGGATGGCGTCCCCCAGGGCTTGTTCCCGAACCGCCTGATACATGAAGGC
TGATGACGCCTATCTCTTG
>KE992666.1/27697-27770 Clostridium sp. KLE 1755 genomic scaffold
Scaffold204, whole genome shotgun sequence.
ATAAGGTCCGGGAATGAAGTTCTCCCATGGCTATTGCCAAAACCGCTGATAAAGCTGATG
ACTTCTGTGATTTT
>CP000957.1/84698-84619 Synechococcus sp . PCC 7002 plasmid pAQ7, complete sequence .
TTTAGCGATGGTGATGGAGCTCACCGAAACCGCCCAGGATTTTTCTTAATTGCACCTCGG
CTAATGGCTCCTACGCTTCC
>FP236826.1/96294-96225 Erwinia billingiae strain Eb661 complete plasmid pEB102.
ATATGACAAGGTGATGGCGTTCCACCTTTCCCAACCGCTCCGTTCCGGAGCTGATGACGC
CTGATGCGAC
>CP002105.1/948502-948436 Acetohalobium arabaticum DSM 5501, complete genome .
ACATATTATGGTGATGGAGCTCGCCGTTAACTATTGCATTCTTGCAATTGATGGCTCCTA
TTTGGGT
>MHYP01000280.1/11687-11749 Planctomycetes bacterium
RIFCSPHIGHO2_02_FULL_38_41 rifcsphigho2_02_scaffold_83486, whole genome shotgun sequence.
TGCGTCAAGGGCGATGGAGTTCGCCATAATTGCTACAAAATAGCTGATAACTCCTGTAGA
AGT
>CALN01000517.1/3428-3365 Stenotrophomonas maltophilia SKK35, WGS project CALN01000000 data, contig: 517
GGCCGCGATGGGGATGGAGCTCCCCCGATAACCGCCTGCAAGGGCTGATGGCTCCTGCCA
AGAC
>JXUW01000032.1/3544-3605 Ferrimicrobium acidiphilum DSM 19497 strain T23 FEAC_contig000032, whole genome shotgun sequence.
AGTCCAAAGGGTGATGGAGTCCACCTTGCAACGCCGAATTGGCTGATGACTCCTCTGATA AG
>CP003281.1/1599165-1599231 Belliella baltica DSM 15883, complete genome.
TGATGCTTAGGTGATGGGGTTCCACCTTCAACCGCCCATTTTTGGTGCTGATGACTCCTA
CTTCAAC
>LUUB01000092.1/109091-109014 Bradyrhizobium sp . BR 10245 contig61, whole genome shotgun sequence.
CACACAGATGGGGATGGAGTCCCCCGATAACCGCCCGGCGGCCAACGACCGTAGTGGGCT
GATGACTCCTGCTCGAGG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>LADR01000066.1/7173-7238 Desulfatitalea sp . BRH_cl2 BRHa_1003278, whole genome shotgun sequence.
AAATAGATTGGGGATGGAGTCCCCCATATAAACCGCCCGGGCTGGGCTGATGACTCCTAC
CGGCAG
>MVHR01000005.1/158203-158143 Mycobacterium heidelbergense strain DSM 44471 NODE_5_length_243907_cov_56.1658, whole genome shotgun sequence. ATCGAGGCAGGCGATGAAGCTCGCCCCAATTGCCACACCGGCTGATGGCTTCTACCACGT
C
>CP002194.1/206403-206336 Deinococcus gobiensis 1-0 plasmid P3, complete sequence .
AGCTCTATAGGCGATGGCATTCGCCCTCCACAACCGCCCCGTGTCCGGCTGATGATGCCT
ACTCCTCA
>FP929053.1/3518373-3518445 Ruminococcus sp . SRI/5 draft genome.
AGGCAGAAAGGGAATGAAGTTCTCCCCCGATCTTTATCGAAACCGCCACTGAGCTGATGA
CTTCTGTGCAATG
>MEDQ01000035.1/10801-10866 Bordetella sp. SCN 67-23 ABS43_C0035, whole genome shotgun sequence.
TATGTTCCTGGAGATGGCATTCCTCCATGAACCGCCGCGCAAGCAGCTGATGATGCCTAC
AGATCC
>MHXS01000153.1/6879-6819 Planctomycetes bacterium GWB2_41_19
gwb2_scaffold_6965 , whole genome shotgun sequence.
ACAAATCAAGGCGATGGAGTTCGCCTAACGCTACGTTGTAGCTGATAACTCCTATTGAAG
C
>FR896389.1/3937-3870 Dialister sp. CAG:357 genomic scaffold, scf55
AAGAAAGACGGCGATGGAGTTCGCCGGGCATATGCCGAACGCTTCGCGCTGATGACTCCT
GTCCCTTA
>KB946293.1/246879-246805 Enterococcus asini ATCC 700915 genomic scaffold acvKe-supercontl .3 , whole genome shotgun sequence.
TGTTAATAGGGGAATGATGTCTCCCTCAGTTTTTACTGAAACCGCAATGAACTTGCTAAT GACTTCTGCCTCAGT
>FWXI01000019.1/99800-99863 Sporomusa malonica strain DSM 5090 genome assembly, contig: GaO 070592_119
GATATTTGGGGCGATGGAGTTCGCCATTAACCCCGTGTCGAACGGTGATGACTCCTACCA
GGTG
>MGYU01000065.1/3913-3975 Gammaproteobacteria bacterium
RIFCSPL0W02_12_47_11 rifcsplowo2_12_subl 0_scaffold_394 , whole genome shotgun sequence.
TTTCTGAAAGGAAATGGCATTCTTCCTTGAACCGTCCCATGGACTGATGATGCCTGCGAG
TTT
>LZYV01000054.1/19198-19137 Clostridium roseum strain DSM 7320
CROST_contig000054 , whole genome shotgun sequence.
AATTTCATAGGTGATGAGGTTCGCCGTAAATATCCTAATGGATTGATGACTTCTGCATAT TA
>MUIN01000007.1/366458-366530 Pseudomonas sp . Bc-h sc_007, whole genome shotgun sequence.
ATCGGGCCTGGAGATGGCATTCCTCCACTTACCAACCGCTGCGCCCGTAGCAGCTGATGA
TGCCTACAGACAC
>KI260332.1/2642-2572 Ruminococcus callidus ATCC 27760 genomic scaffold Scaffoldl390 , whole genome shotgun sequence.
GTTTCCTTAGGGAATGATGTTCTCCCTTGGGATTCCAAAACCGCTTGAGTGCTGATGACG
TCTGCTTTTTA
>FM954972.2/699217-699134 Vibrio splendidus LGP32 chromosome 1
TAGCCATCAGGTGATGGGGTTCCACCTAAGCTTTTGCTTCAACCGCCCGTTCTTTTGAAC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CGTGCTAATGACTCCTACAGAATC
>AEXE01001175.1/192-296 Burkholderia sp . TJI49 contig634, whole genome shotgun sequence.
GTCGGCTGCGGAGATGGCATGCCTCCCTGCGGTCGCCGGCGGCGCGTGGCGCGCGGCCCG
GACCCTCAACCGCCGGTGCGTCCGGCTGATGATGCCTGCGTGTTC
>CP015772.1/1950443-1950517 Niabella sp . BS26, complete genome.
TTATAAACAGGAAATGGTGTCTTCCTGAACCAACCGTTCACGCCTTTTGCGGGATCTGAT
GGCGCCTACAAATTT
>MGQS01000055.1/380-316 Deltaproteobacteria bacterium RBG_16_54_11
RBG_16_scaffold_261725, whole genome shotgun sequence.
ATACTAAGAGGGGATGGAGTTCCCCTGACAATCGCCCAGGCCGGGCTGATGACTCCTACG ACGCA
>MBSV01000068.1/16362-16420 Clostridium sp . W14A NODE_46, whole genome shotgun sequence.
TAATAGATTGGTGATGAGGCTCACCACAATGCATTTATGCTGATGGCTTCTACTTTAAG
>ADAD01000052.1/93049-92975 Leptotrichia goodfellowii F0264 contig00067, whole genome shotgun sequence .
TTATTTCATGGAAATGATGTCTCCCGAGGCAAGTGCCTAAACCGCTTTATTTGAGCTGAT
GACTTCTGCATATTT
>MGMG01000019.1/23548-23607 Chloroflexi bacterium GWB2_54_36
gwb2_scaffold_1426 , whole genome shotgun sequence.
CCCCGCTGTGGTGATGGAGTCCACCCTTAACCGCTTCGTGCTGATGACTCCTGCATTTCA
>MLY001000053.1/67998-68063 Streptomyces sp . MUSC 1 53, whole genome shotgun sequence.
CTGGTATCCGGCGATGGGGCTCGCCGCCAACCGCACCCCTGGGGTGCTGATGGCCTCTAC
CCGCTG
>JPKR02000005.1/497787-497711 Tatumella morbirosei strain LMG 23360
Contig5, whole genome shotgun sequence.
TGACTTCGGGGTGATGGTGTTCCACCTCAACCAACCGCCGCTCTGAGACCTGAGCGGCTG ATGACGCCTGACATAAT
>HF987456.1/100469-100539 Ruminococcus sp. CAG:108 genomic scaffold, scf76
AATATGAATGGGAATGAGGTTCTCCCACGGCTTTTGCCGAAACCGCTTTTGCTGATGACT
TCTGTTTTTTG
>CP014675.1/5813-5873 Kozakia baliensis strain DSM 14400 plasmid
pKB14400_l, complete sequence.
CCTGGACAAGGGGATGGGGTTCCCCTGAAACCGCCGCAAGGCTGATGACTCCTCTGTGCC
G
>GL876933.1/301939-301837 Novosphingobium nitrogenifigens DSM 19370 genomic scaffold scaffoldOOOlO, whole genome shotgun sequence.
TGCGATGACGGCGATGGATTTCCGCCGGGCCTGCTGTGCCGCACAGGCCGAACCGCCTTT CGACGGGCTCTCTTGTCGAAAGGCTGATGATTCCTGCCTTCGC
>MAST01000004.1/33923-33850 Humibacillus sp . DSM 29435 contigl2, whole genome shotgun sequence.
CTGAGAGAAGGTGATGGACCCCGCCTGGGTGGTGACAGCCCTAACTGCATTTTGCTGATG
GTTCCTACGACGTT
>MGDF01000100.1/3502-3565 Candidatus Schekmanbacteria bacterium
RBG_16_38_11 RBG_16_scaffold_39291, whole genome shotgun sequence.
TTTAAAAATGGCGATGGGGTTCGCCTCTAACCATCCATAAAGGATTGATGACCCCTACTG AAAC
>CP008796.1/645368-645297 Thermodesulfobacterium commune DSM 2178, complete genome.
CAAAGAATTGGGGATGGCGTCCCCCATAAATCGCCTCTAAAGTCTTTAGGGGCTGATGAC
GCCTACCTCTCA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>FRAM01000002.1/331939-332004 Chryseobacterium molle strain DSM 18016 genome assembly, contig: Ga0131196_102
TAAAAAAAAGGGAATGGTGTCTCCCTTACCCAACCGCCCTAAAAAGCTGATGGCGCCTGA
TTAAAA
>FXAT01000002.1/736027-736088 Paraburkholderia susongensis strain LMG 29540 genome assembly, contig: Ga0139082_102
TTCGCCGCTGGAGATGGCATTCTCCATTAACCGCCCTCGGGGCTGATGATGCCTGCTACG
CC
>JXOF01000129.1/2647-2586 Bradyrhizobium elkanii strain UASWS1015
Contigl29, whole genome shotgun sequence.
GTATCCCTCGGGGATGGAGTTCCCCATTAAGCGCCAATCCGGCTGATGACTCCTGCAACG
GG
>CP000568.1/234427-234485 Clostridium thermocellum ATCC 27405, complete genome .
TGAATAAAAGGTGATGGAGTTCACCTGAAATGCGTAATGCTGATGACTCCTACAGGAAT
>MNIS01000023.1/10346-10283 Gemmatimonadetes bacterium 13_1_20CM_4_66_11 13_1_20cm_4_scaffold_534 , whole genome shotgun sequence.
CTTACAGCAGGCGATGGAGTTCGCCTCAAACCGCCCTTTTCGGGCTGATGACTCCTACCC GAAC
>CP003942.1/1168935-1168860 Serratia marcescens FGI94, complete genome.
ATCGACCACGGAGATGACATTCCTCCCGCCCCGATAGGGATAAACCGCCTGACCGGCTGA
TGATGTCTGCGTAACC
>CP009313.1/723035-722972 Streptomyces nodosus strain ATCC 14899 genome.
ATGGAGGTGGGCGATGAGGCTCGCCCTCGACCGCACATCCCGTGCTGATGGCCTCTGCTG
CCGG
>CP012996.1/2979607-2979542 Pedobacter sp. PACM 27299, complete genome.
CACAGAAAAGGAAATGGTGTCTTCCTGATTGAACCGCCCTAAAAAGCTGATGGCGCCTGG
TTAATT
>CBTY010000009.1/201754-201680 Thaumarchaeota archaeon N4, WGS project CBTY01000000 data, contig: APF_1883_Contig_4
ACAGAACAAGGCTATGACATTAGCCGTGGCGCAAGCCAGAACCGCATCGCGCATGCTGAT
AATGTCTACACAAAG
>ATHV01000008.1/6866-6807 Desulfovibrio sp . X2 ctgl7, whole genome shotgun sequence .
GGCCGCCACGGGGATGGAGTCCCCCGAGAACCGCACCATGCTGATGACTCCTGCCGGACG
>JXUW01000001.1/2579-2636 Ferrimicrobium acidiphilum DSM 19497 strain T23 FEAC_contig000001, whole genome shotgun sequence.
ATGGCTTGCGGCGATGGAGTTCGCTGATGCGTCTCAGACTGATGACTCCTACTAAAGC
>LHUR01000005.1/88557-88496 Clostridium homopropionicum DSM 5847
CLHOM_contig000009, whole genome shotgun sequence.
AACTAAATAGGTGATGAAGTTCGCCATAATCATCCTCATGGATTAATGACTTCTACAATA AG
>AP009386.1/408531-408637 Burkholderia multivorans ATCC 17616 genomic DNA, complete genome, chromosome 2.
CCGCCTTGCGGAGATGGCATGCCTCCGTGCGGCGTTTTCCCGCGCATCGTGCGCGCCGGT
CCGCTGCTCAACCGCCGGTTCGTCCGGCTGATGATGCCTGCGTGTTC
>MBSV01000068.1/16190-16118 Clostridium sp . W14A NODE_46, whole genome shotgun sequence.
CAGATCATTAGGGATGAAGCTCCCCTTGAAAATGTTCGAAATGCTGGAAACGGCTGATGG
CTTCTACAGGAAT
>AKFU01000014.1/12710-12638 Clostridium sp . MSTE9 ctgl20006482998, whole genome shotgun sequence.
AATTTTTCTGGGAATGAAGTTCTCCCGCGGCATTGCCGAAACTGCTTATTTAGCTGATGA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CTTCTGCGCTGGC
>MNGV01000003.1/10075-10135 Acidobacteria bacterium 13_1_40CM_3_55_6 13_1_40cm_3_scaffold_1207 , whole genome shotgun sequence.
AGTTCTACCGGCGATGGAGTCCGCCATCAAATGCCTGCGTGCTGATGACTCCTACCAAAT
C
>CP018309.1/1749379-1749315 Vibro shilonii strain QT6D1 chromosome 2, complete sequence.
TGCAGCCAAGGTGATGGGGTTCCACCTACTTAACCGCCAAATCGGCTGATGACTCCTACA
GTTTC
>KB946874.1/2597044-2596978 Enterococcus gilvus ATCC BAA-350 genomic scaffold acEQl-supercontl .3, whole genome shotgun sequence.
TTTGGTAAGAGGGATGGTGCTCCCCTCCTAACCGCTGACTCAATCAGCTGATGGCGCCTG TTATAAG
>LMQK01000023.1/308864-308792 Methylobacterium sp. Leaf399 contig_3, whole genome shotgun sequence.
ATGCACCACGGCAATGGATTCTGCCGGGCCACATGGCCGAACCGCCCTTCGGGCTGATGA
TTCCTACCTTTTG
>KI543234.1/518618-518554 uncultured archaeon A07HR60 genomic scaffold scf7180000001783, whole genome shotgun sequence.
GTCTGTTCAGGCGATGGAGTCCGCCTGTCCCAACCGCCATTCTGGCTGATGACTCCTACT GAACA
>JOKG01000001.1/596090-596005 Endozoicomonas montiporae strain LMG 24815 scaffoldOOOOl, whole genome shotgun sequence.
ATGGCTCAGGGTGATGGAGTTCCACCTGCTAACCGCTCCCATAGCTGATTCACATAGCTT
TGGCAGCTGATGACTCCTGTGTATTG
>BAMZ01000022.1/60532-60611 Acetobacter syzygii 9H-2 DNA, contig:
Absy_022.
ACTGAGGGTGGCAATGGACTCTGCCGGATTTTGGTAACCAAAATCGAACCGCCCTGTTGG
CTGATGATTCCTACCCCGCA
>GG705267.1/72722-72791 Providencia rettgeri DSM 1131 genomic scaffold Scfld5, whole genome shotgun sequence.
AAACCTTTGGGAGATGGCATTCCTCCTTATATAAAACCGCCCGTAGAGGGCTGATGATGC
CTACGTTAAC
>MNJQ01000085.1/11364-11292 Acidobacteria bacterium 13_1_20CM_3_58_11 13_l_20cm_3_scaffold_5366, whole genome shotgun sequence.
TTTCGATTTGGCGATGGAGTTCGCCATAACCGCCCCGGTGATGTTCCCCGGTGCTGATGA CTCCTGGCAGCCC
>MEKW01000051.1/13961-14026 Acidobacteria bacterium
RIFCSPLOWO2_02_FULL_68_18 rifcsplowo2_02_scaffold_3607 , whole genome shotgun sequence.
CCGAGTGTCGGGAATGGTGTCAATCCTAAACCACATCTTGAAGATGCTGATGACGCCTTC
GGACAC
>MIA001000077.1/1710-1646 Spirochaetes bacterium GWB1_36_13
gwbl_scaffold_2872 , whole genome shotgun sequence.
TTTTTATTTGGAGATGGAGTTCTCCCTTAAACCGTCCTGTAAAGACTGATAACTCCTACT AAAAT
>BBSB01000017.1/87810-87906 Vibrio sp . JCM 19236 DNA, contig00017.
GCGTCGAAAGGTGATGGGGTTCCACCTAATTGAACCGCCGAATCGTTGAAGCGCACAGCG
CTGAACAACCCTCTCGGCTGATGACTCCTACTAAAAT
>FTNC01000013.1/57314-57247 Halanaerobium kushneri strain ATCC 700103 genome assembly, contig: Ga0104713_113
ATGAATTAAGGGGATGGAGCTCCCCCTTTAACCACTGTTAAATAGCAGTTGATGGCTCCT
ATCTTTTT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>JSYZ01000007.1/136286-136356 Pseudomonas fuscovaginae strain IRRI 6609 Pf_IRRI_6609_contig_7.7, whole genome shotgun sequence.
CGGCGCTCCGGAGATGGCATTCCTCCGTTCCCAACCGCTGTGCTGCTGCAGCTGATGATG CCTACAGACAC
>KI518170.1/6377-6319 Cetobacterium somerae ATCC BAA-474 genomic scaffold Scaffold37, whole genome shotgun sequence.
TAATAAATAGGCGATGGAGTTCGCCATTAAACGCGAAAGCTAATGACTCCTACTCTTTT
>CP002293.1/2935878-2935817 Geobacillus sp . Y4.1MC1, complete genome.
TTAAAATTAGGCGATGGGGTTCGCCGTAATTGCCTTTTATGGCTGATGACTCCTACCAGT
AA
>MKVZ01000019.1/51236-51296 Rhizobiales bacterium 63-22
SCNpilot_cont_500_bf_scaffold_96, whole genome shotgun sequence.
ATGATCGCAGGGAATGGAGTCTCCCGCAGACCGCGCCACCGCTGATGACTCCTGCCAATG
G
>MIDI01000122.1/16449-16382 Thermodesulfovibrio sp . RBG_19FT_COMBO_42_12 rbg_19ft_combo_scaffold_289, whole genome shotgun sequence.
TGAAGTAAAGGCGATGGAGTTCGCTTAATCGTCCCGATTAAATCGGGACCGATAACTCCT ACCACCCT
>CP017253.1/1274876-1274817 Clostridium taeniosporum strain 1/k, complete genome .
TATATATTAGGTGATGGAGTTCGCCTTTAACTGCTTAATGCTAATGACTCCTACAAAAAC
>KB456680.1/1125-1190 Drosophila ficusphila unplaced genomic scaffold scf7180000450230, whole genome shotgun sequence
CGCGGCACAGGGGATGGGGTTCCCCCTTATAACCGCCCGCTCGGGGCTGATGACTCCTGC CGTTAT
>ACJM01000018.1/29060-28989 Dethiobacter alkaliphilus AHT 1 ctgl6, whole genome shotgun sequence.
TTTAATTAAGGCGATGGAGCTCGCCGGGGCATTGCCCAAACCGCACAATGTGCTAATAGC
TCCTGCCATGTA
>AOMB01000031.1/71848-71781 Halococcus hamelinensis 100A6 contig_31, whole genome shotgun sequence.
GCTTCATCGGGCGATGGAGTCCGCCTGCCCCAACCGCTGGGTTTCCAGCTGATGACTCCT
TCTCCAGC
>AP006878.1/440492-440567 Thermococcus kodakarensis KOD1 DNA, complete genome .
TTCAGCGCGGGCGATGGCGTCCGCCCGGGGCTTCGAGCCCGAACCGCTCGAAAGGGCTGA
TGACGCCTGTTCTACT
>LOEE01000028.1/49751-49818 Thermotalea metallivorans strain B2-1
AN619_contig000035, whole genome shotgun sequence.
TTTATCCTAGGTGATGGAGCTCACCTTTAACTATTGCAAGGAAGGCAATTGATGGCTCCT GCTATAAC
>LADR01000006.1/15176-15112 Desulfatitalea sp. BRH_cl2 BRHa_1002920, whole genome shotgun sequence.
AAGTTGATCGGGGATGGAGCCCCCCGTATAACCGCCCGGGCTGGGCTGATGACTCCTACC
GGCAG
>MKRV01000110.1/11303-11243 Alphaproteobacteria bacterium 65-37
SCNpilot_cont_300_bf_scaffold_2027, whole genome shotgun sequence.
GTCTCCGATGGGGATGGAGTCCCCCGACAACCGCCGCAAGGCTGATGACTCCTACCGAAC
G
>FR903093.1/21636-21561 Eubacterium sp. CAG:581 genomic scaffold, scf56
TAACCCTTAGGGAATGAATGTCTCCCTTAGTATTATTACTTAAACTGCTTTTTAAGCTGA
TGACTTCTGCATTTTG
>BAFN01000001.1/1079928-1079988 Candidatus Brocadia sinica JPN1 DNA, APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
contig: brosiA.
TAAAATTCAGGCGATGGAGTTCGCCAATTGCTATATGATAGCTGATAACTCCTATTGAAG
C
>AP011532.1/2421352-2421430 Methanocella paludicola SANAE DNA, complete genome .
ACGTATATCGGCGATGAAGCTCGCCGTGCTCTTATGCAGAACCGCCGCACGGACTGCGAC
TAATGGCTTCTGTTCGTTA
>CP000615.1/985285-985390 Burkholderia vietnamiensis G4 chromosome 2, complete sequence.
GCCGGCTGCGGAGATGGCATGCCTCCCTGCGGTCGAGCACGGCGCGTTGCGCGCGCGAGT
CGGCCCTCAACCGCCGGTGAGCCCGGCTGATGATGCCTGCGTGTTC
>FQZL01000032.1/30867-30950 Dethiosulfatibacter aminovorans DSM 17477 genome assembly, contig: EJ55DRAFT_scaffoldO 0029.29
ATAATATAAGGCGATGGAGTTCGCCAGGCAAGAAAACATAACAAGATTGCCTAACCGCCA AGTGCTAATGACTCCTGTTGACTT
>JH414706.1/1587485-1587559 Subdoligranulum sp . 4_3_54A2FAA genomic scaffold supercont 1.9 , whole genome shotgun sequence.
TCCATTACAGGGAATGAAGTTCTCCCTTGGATTTTATCCAAAACCGCTTTTCAAGCTGAT GACTTCTGCGATTTT
>FQXE01000008.1/20864-20798 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_108
GCACCCAACGGAGATGGCATTCCTCCGCTAACCGCCGCGCTCTGCGGCTGATGATGCCTA
CAAGTTC
>BA000019.2/2723228-2723309 Nostoc sp . PCC 7120 DNA, complete genome.
ATTACTAATGGCGATGGAGTTCGCCAAAACCGCTATTCTAGTCAAAAAATTAATCTAGAA
GGCTGATGACTCCTACTTTTCT
>FQXP01000005.1/99007-98936 Clostridium collagenovorans DSM 3089 genome assembly, contig: EJ35DRAFT_scaffold00003.3
GAGATATGAGGAGATGAAGCTCTCCTGATAAACTTTATCAAACTGCTAAATGCTAATAGC
TTCTACGACCTG
>CP003926.1/1806829-1806753 Gluconobacter oxydans H24, complete genome.
TTCGAGATCGGCAATGGACTCTGCCTGACCCGAACAAGGTCGAACCGCCCGCAAGGGCTG
ATGATTCCTACCTCGCC
>JYMT01000038.1/519073-519133 Bradyrhizobium sp . LTSP885 NODE_38, whole genome shotgun sequence.
CCTGTTGATGGGGATGGAGTCCCCCGATAACCGCCGCAAGGCTGATGACTCCTACCGGGC
G
>MNJV01000118.1/1655-1590 Candidatus Rokubacteria bacterium
13_1_20CM_2_70_7 13_l_20cm_2_scaffold_3979, whole genome shotgun sequence.
AACTGAAAAGGTGATGGGGTTCGCCTAAACCGCCCAGACCATCGGGCTGATGACCCCTGC
CGGACC
>JPMD01000022.1/41832-41889 Clostridium sulfidigenes strain 113A c22, whole genome shotgun sequence .
ATAATTTATGGTGATGGAGTTCACCTTAATCGCAATTGCTAATGACTCCTACAGACAA
>CP001793.1/4378898-4378958 Paenibacillus sp . Y412MC10, complete genome.
ATTGAATGAGGCGATGGAGTTCGCCATAACCGCTTCCATAGCTAATGACTCCTACCAGAG
G
>KB405046.1/1712591-1712668 Photobacterium damselae subsp. piscicida DI21 genomic scaffold scaffoldOOOOl, whole genome shotgun sequence.
ATCTTAACGGGAGATGATGTTCCTCCTTTAACCGCCTTTCTGATCTTTTTCTCTAAGGAT GATGACGTCTAACAACAG
>MICJ01000091.1/301-369 Syntrophobacterales bacterium GWC2_56_13
gwc2_scaffold_5779, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CAAACTGGTGGGGATGGAGTCCCCCGCTGAAAAATGCTTCTTAACAGAGCTGATGACTCC
TGCCCCGCC
>FXAT01000003.1/189831-189759 Paraburkholderia susongensis strain LMG 29540 genome assembly, contig: Ga0139082_103
GTCGAAATGGTAGATGGCATTCCTCCCTGAATCGCTGCTTGCATAGGCCGCAGCTGATGA
TGCCTACAGTTCC
>KI273145.1/1773152-1773210 Clostridium intestinale URNW genomic scaffold scaffoldl, whole genome shotgun sequence.
AAAAATAAAGGTGATGGAGTTCACCATAATCGCTAAATGCTAATGACTCCTATAAGTAA
>GL379781.1/3877121-3877188 Chryseobacterium gleum ATCC 35910 genomic scaffold SCAFFOLD1, whole genome shotgun sequence.
AACAAAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTCTAAAAGAGCTGATGACGCCT GATTAAAT
>FR898703.1/17474-17546 Clostridium sp. CAG:307 genomic scaffold, scf82
AAATATTTAGGAAATGATGTTCTTCCTTGTTATTTAACAAAACCGCTATTATGCTGATGA
CGTCTATAGTTTT
>LQOM01000035.1/130087-130147 Mycobacterium celatum strain DSM 44243 contig_40, whole genome shotgun sequence.
TCGGCCTCAGGCGATGAAGCTCGCCTTAACCGCCGAACCGGCTGATGGCTTCTACCCGTG
A
>MNQW01000009.1/7224-7285 Butyricimonas synergistica isolate 43_13
Ley3_66761_scaffold_4522 , whole genome shotgun sequence.
CCAACTTATGGTAATGGTATCTACCTCGAACCGCTTGAAAAGCTGATGATGCCTACATGT AA
>AE015928.1/5995735-5995673 Bacteroides thetaiotaomicron VPI-5482, complete genome.
CTTTGCACCGGCTATGGGATTAGCCTTTAACCGCCCTTTGGAGCTGATGATCCCTACAAT
TTG
>AP013035.1/752509-752574 Thermosulfidibacter takaii ABI70S6 DNA, complete genome .
CTGGCTAGAGGCGATGGAGCCCGCCTTTAACCGCCGCACTCAGCGGCTGATGGCTCCTAC
CTCAAA
>CM001774.1/2070305-2070375 Providencia rettgeri Dmell chromosome, whole genome shotgun sequence.
ATATATATGGGAGATGGCATTCCTCCTTTTTACAAAACCGCCCGTAGAGGGCTAATGATG
CCTGCGTTAAC
>CP013854.1/1763673-1763612 Pseudonocardia sp. HH130630-07, complete genome .
CCCGGTGAAGGCGATGGAGCTCGCCGACAACCGCACCCGCTGCTGATGGCTCCTACCGAT
CC
>CP006850.1/5444076-5444147 Nocardia nova SH22a, complete genome.
CTGAGAGCAGGTGATGGATCCCGCCCGGGAAATTCCCGAACCGCCACCTCGGCTGATGGG
TCCTGCTGGCGT
>MHBM01000161.1/11806-11885 Lentisphaerae bacterium RIFOXYA12_FULL_48_11 rifixya3_full_scaffold_335, whole genome shotgun sequence.
TCCTCCCCGAGCGATGGAGTTCGCTCAGAACTGCCCTGATGAATCTTTGTGAACTAAGGG ATGATGACTCCTGCTGTATC
>MNRF01000071.1/1702-1624 Clostridiales bacterium 42_27
Ley3_66761_scaffold_27397 , whole genome shotgun sequence.
AGAATAAACGGGAATGAAGTTCTCCCGAAGTAACCAGATTACTTGAACTGCTTTGAAAGC TGATGACTTCTGCGACGAA
>MPCW01000005.1/143348-143417 Flavobacterium sp . MedPE-SWcel
Flavobacterium-sp-MedPE-SWcel-C5 , whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ACTATCAATGGCAATGATGTCTGCCTCGAACCGCTTTTTATAACTAGAAGCTGATGACGT
CTATTTAATG
>JXKH01000002.1/476995-476934 Enterococcus canis strain DSM 17029
Scaffold2, whole genome shotgun sequence.
TTTCTTGTAAGGGATGGTGCTCCCGATAACCGCTAAAATTAGCTGATGGCGCCTGTTGTA
AG
>LGK001000003.1/69724-69656 Thermanaerothrix daxensis strain GNS-1 contig_5, whole genome shotgun sequence.
GCAGCTTGCGGTGATGAGGCTCACCGGAGACCCAAACCGCCCATCAGGGCTGATAGCCTC
TGTCGTTTC
>DF820455.1/1608747-1608687 Bacterium UASB14 DNA, scaffold:
UASB14_scaffold_l .
TTTTTATTCGGTGATGGGGTTCACCGCTAACCGCCGACGTGCTGATAACTCCTACCAGCG
T
>HG916765.1/3012383-3012452 Castellaniella defragrans 65Phen complete genome
CCCTCATCTGGAGATGGCATTCCTCCACGAAACCGCCGCGCGCGCCGCGGCTGATGATGC
CTACGGCCCC
>FUWX01000006.1/107811-107886 Cetobacterium ceti strain ATCC 700028 genome assembly, contig: EI52DRAFT_scaffold00003.3
TAATAAAGAGGGAATGAAGTTCTCCCTTAGTTTTGCAACTAAAACCGCTTATTTAGCTGA
TGACTTCTACGACTTT
>AZFF01000008.1/81061-80998 Lactobacillus rossiae DSM 15814 Scaffold8, whole genome shotgun sequence .
AACTTAATAGGCGATGACGTTCGCCGCAAACAATTATTGATAATTTGATGACGTCTACTG
CATG
>LJC001000041.1/20596-20655 Alicyclobacillus ferrooxydans strain TC-34 contig_36, whole genome shotgun sequence.
GTTGTGAAAGGTGATGGAGCTCACCATAACCGCCTCGAGGCTAATGGCTCCTGTGGATGA
>CP009286.1/604006-604065 Paenibacillus stellifer strain DSM 14472, complete genome.
CTGATGAACGGCGATGGAGTTCGCCATAACCGCCGCAAGGCTGATGACTCCTACCAGTGG
>LOES01000122.1/5172-5114 Clostridia bacterium BRH_c25 BRHa_1003505 , whole genome shotgun sequence.
AATAGTTATGGTGATGGGGCTCACCATAACCGCTGTAAGCTGATGGCTCCTATTTTGAC
>MEKP01000380.1/13627-13564 Acidobacteria bacterium
RIFCSPLOWO2_02_FULL_59_13 rifcsplowo2_02_scaffold_8252 , whole genome shotgun sequence.
TAAATCTTCGGCGATGGAGTTCGCCCAAACCGTCCGGTCACGGACTGATGACTCCTGGCG
GCCC
>MJIL01000045.1/762-825 Photobacterium sp. 13-12 C2015, whole genome shotgun sequence.
TTAGTAACGGGTGATGGAGTTCCACCTTTAACCGCTCGACTGAGATGATGACTCCTGCTG
TAGT
>FR903040.1/40256-40157 Eubacterium sp. CAG:76 genomic scaffold, scf98
AATATTACTGGGAATGAGGTTCTCCCAGGAACGAGATAGAAATAGCCAAATTTTTCGTTC
ATAACCGCTTGATTATAAAGCTGATGACTTCTGCAAAGTG
>AP011532.1/2434049-2434128 Methanocella paludicola SANAE DNA, complete genome .
TCCCATACTGGTGATGAAGCTCACCGTGCTCTTATGCAGAACCGCCGCAAGGCAATGCGA
CTAATAGCTTCTCCAGACAT
>LGGZ01000182.1/1820-1745 Thermotogales bacterium 46_20 MPJ_scaffold_5313 , whole genome shotgun sequence . APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AATTCTCTAGGCAATGAAGCTTGCCTTGGAGAACAGTCCTCCAGAACCGCCACCAGCTGA
TAGCTTCTACCAAAAT
>MNQS01000021.1/67471-67537 Bacteroides sp . 43_108 Ley3_66761_scaffold_49, whole genome shotgun sequence .
ATTAAAAAAGGGAATGGGACTTCCCTGTTATAAACCGCTGAAAAAAGCTGATAGTTCCTG
TTGAGAG
>NGKU01000001.1/3091201-3091129 Enterococcus sp . 8G7_MSG3316
scatfoldOOOOl, whole genome shotgun sequence.
GGAGAATAGGGGAATGAATGTCTCCCTCAGTGATTACTGAAACCGCAATTTTGCTAATGA
CTTCTACAAGCTT
>CP002657.1/360259-360193 Alicycliphilus denitrificans K601, complete genome .
CGGCCGCCAGGAGATGGCGTTCCTCCTGAAACCACCGCGTCGCGCGGTTGATGACGCCTG
CGGAGCC
>JXAK01000060.1/37628-37688 Paenibacillus sp . VKM B-2647 B-2647_060, whole genome shotgun sequence.
AGCTTATACGGCGATGGAGTTCGCCTTTAACCGCCTTTTTGCTAATGACTCCTACCAGTT
T
>MDSU01000018.1/319158-319094 Desulfurella amilsii strain TR1
opera_scaffold_l, whole genome shotgun sequence.
ATAATGGCAGGCTATGAGGTTAGCCTTTATACGCTCTTTTTTGAGCTGATGACCTCTACT TTTTA
>LXGI01000116.1/46999-47091 Paraburkholderia tropica strain P-31
P3 l_contig_7 , whole genome shotgun sequence.
TCGTGCGCAGGAGATGGCATTCCTCCTGCATTTTCCCCAACCGCCGGACCGGCAGCGCTG
CTTCCTGCTCCGGCTGATGATGCCTGCGAGTCC
>CP002021.1/1823327-1823396 Thiomonas intermedia K12, complete genome.
AACCATTAAGGCGATGGAGTTCGCCAAAACCGCGTGGGGCTGCCCCCAAGCTAATGACTC
CTACCTGCCG
>CP002826.1/2716030-2715969 Oligotropha carboxidovorans OM5, complete genome .
ATTGTTGATGGGGATGGAGTTCCCCCGATAACCGCCGCAAGGCTGATGACTCCTACCGGA
CG
>CP000102.1/1672674-1672743 Methanosphaera stadtmanae DSM 3091, complete genome .
AACAAATAAGGTGATGAGGTTCCACCTATTTAACTGCCAGTGATTACTGGATGATGACTT
CTATTTTTAA
>AKJA01000121.1/27981-28047 Herbaspirillum sp. YR522 PMI40_contig_322.322, whole genome shotgun sequence .
ATTCGGCCCGGAGATGGCATTGCTCCCCGAACCGCCGCCTGTTGCGGCTGATCATGCCTA
CGGTGAC
>MEFB01000036.1/3026-2924 Rhodanobacter sp . SCN 67-45 ABS82_C0036, whole genome shotgun sequence.
GGCCCGATGGGAGATGGCATGCCTCCCGTTGCGGGCGCGCGACGCGTCCGCAACCAACCA
CCCCGGGTGCTTCCGCCACGAGGTTGATGATGCCTGCAACGCC
>LMLC01000004.1/376439-376515 Sphingomonas sp. Leaf34 contig_4, whole genome shotgun sequence.
ATCGACTTTGGCAATGGATTTCTGCCGGGCAATTTGGCCGAACCGCTCTGATAAGAGCTG
ATGATTCCTACTTGGCG
>FQZV01000005.1/59151-59217 Geosporobacter subterraneus DSM 17957 genome assembly, contig: EJ58DRAFT_scaffold00003.3
TAGAAAGACGGTGATGGAGCTCACCTTTAAGTATTGCATCCGTGCAATTAATGGCTCCTA
CTTGGCT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>MGQI01000065.1/2753-2674 Deltaproteobacteria bacterium RBG_13_58_19 RBG_13_scaffold_17853 , whole genome shotgun sequence.
TAAATAGTTGGCGATGGAGTCCGCCGGGGTGGAGAGGGCATCACCCAAACCGCATTGATG CTGATGGCTCCTGATCTCAC
>CP000033.3/965933-965991 Lactobacillus acidophilus NCFM, complete genome.
AAGTAAAAAGGTGATGACGTTCACCGTTAATCGAGAAATCTAATGACGTCTACTTTATC
>FR887768.1/8595-8527 Bacteroides sp . CAG:443 genomic scaffold, scf45
CGAAATCAAGGGAATGGGACTTCCCTGTTATAAACCGCTAATTGAATAGCTGATAGTTCC
TGCCGGAGT
>CP000975.1/1995740-1995815 Methylacidiphilum infernorum V4, complete genome .
CTCTTATTGGGCGATGGGGCTCGCCGACCGTTAACCGCCGTTCTTTTCAAAGAAGGCTGA
TAGCTCCTACCTCTAA
>LSLJ01000003.1/344447-344509 Sporomusa sphaeroides DSM 2875
SPSPH_contig000003, whole genome shotgun sequence.
AAAGTTTGGGGCGATGGAGTTCGCCCGAATCCTGGCTGGCCAGGTGATGACTCCTACCGT AAA
>BBMT01000001.1/428748-428811 Vibrio maritimus DNA, contig: contigOOOOl, strain: JCM 19240.
TGTCATCAAGGTGATGGGGTTCCACCTACTTAACCGCCAACTGGCTGATGACTCCTACAG
TTTC
>AM412317.1/1975371-1975313 Clostridium botulinum A str. ATCC 3502 complete genome
TTAAATATAGGCGATGGAGTTCGCCATAATTGCGTAATGCTAATGACTCCTACAAATAA
>CR543861.1/406046-406112 Acinetobacter sp . ADP1 complete genome.
CGACAGATAGGAGATGACATTCCTCCTCGCAAAAACCGCCGTTCTGGCTAATGATGTCTA CGTTACC
>NFI001000003.1/155466-155536 Alistipes sp . An31A An31A_contig_3, whole genome shotgun sequence.
AGCCGCAAAGGCAATGGTATCTGCCTCAAACCGCCGCCCGTTCGGGATGTGCTGATGATA
CCTGCATAAAC
>AEXE01000085.1/4565-4500 Burkholderia sp. TJI49 contig072.1, whole genome shotgun sequence.
AGCTGAAACGGAGATGGCATTCCTCCAAAACCGCCGGTATCTCCGGCTGATGATGCCTAC
GGATTC
>MKUQ01000030.1/7789-7852 Burkholderiales bacterium 70-64
SCNpilot_expt_1000_bf_scaffold_318, whole genome shotgun sequence.
GCTGTCAGTGGAGATGGCATACCTCCCTTAACCGCCGGTCACGGCTGATGATGCCTACGC AAGG
>CP003548.1/698262-698337 Nostoc sp . PCC 7107, complete genome.
AAGAACAGTGGTGATGAAACTCGCCTTCTGGAATAACCAGAAAACCGCCTGTGGGGCTGA
TAGTTTCTAGTTCAAT
>CP011254.1/3454057-3453994 Serratia fonticola strain DSM 4576, complete genome .
CCAACAAACGGAGATGACATTCCTCCATAACCGCCTTTCACAGGCTGATGATGTCTACGT
AACC
>FSRD01000001.1/3621429-3621489 Bradyrhizobium erythrophlei strain GAS478 genome assembly, contig: Ga0132009_ll
CATCCCGATGGGGATGGGGTCCCCCGATAACCGCCGCAAGGCTGATGACTCCTGCCAGGC
G
>JXMW01000003.1/205756-205832 Methanobrevibacter arboriphilus JCM 13429 = DSM 1125 strain DH1 MBBAR_3c, whole genome shotgun sequence.
ATACAGTAAGGTGATGAGGTTCCACCTATAACTGCCATTTTTTTTAAAAAAATCTGGATG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ATGACTTCTACTATCAA
>KE136516.1/288643-288710 Enterococcus durans ATCC 6056 genomic scaffold acyDB-supercont2.4 , whole genome shotgun sequence.
AAGTGAATAGGTGATGGTGTTCGCCTTTAAACGTTGATCGTAAGATAACTAATGACGCCT ACTAAAGA
>JMIY01000007.1/448820-448880 Candidatus Methanoperedens nitroreducens strain ANME-2d ANME2D_Contig_7.7 , whole genome shotgun sequence.
ATAGAAACAGGTGATGGAGTTCACCTTCAACCGCTTGATAGCTGATAACTCCTACCGATC
A
>LN831776.1/1329372-1329310 Paenibacillus riograndensis SBR5 isolate SBR5 (T) genome assembly, chromosome: I
ATTCAATACGGCGATGGAGTTCGCCATTAACCGTCCCTAGAGACTAATGACTCCTACCAG
CGG
>CP021848.1/831245-831319 Thermococcus sp. 5-4, complete genome.
GGTCTTTCGGGCGATGGCGTCCGCCCGGGCTTCGAGCCGAACCGCCCGTCGAGGGCTGAT
GACGCCTGTTCTCCG
>MERQ01000020.1/82209-82148 Betaproteobacteria bacterium
RIFCSPLOWO2_12_FULL_68_20 rifcsplowo2_12_scaffold_1499, whole genome shotgun sequence.
GCGAAGCAGGGAGATGGCATTCCTCCCATAACCGCCGCAAGGCTGATGATGCCTACGCGA
GA
>MGRA01000126.1/7329-7258 Deltaproteobacteria bacterium
RBG_19FT_COMBO_46_12 rbg_19ft_combo_scaffold_1869, whole genome shotgun sequence .
ATCGACAGAGGCGATGAGGTTCGCCTTTAACTGTCCTTTTGTTGAGAAAGGACTGATAAC
TTCTACCGGGTT
>MEGJ01000001.1/237916-238006 Rhodanobacter sp . SCN 69-32 ABT17_C0001, whole genome shotgun sequence .
CTCCATGAGGGAGATGGCATGCCTCCCGTTCCCCGGCGGAACGAACCGCCCCGCACGACC
CATCGCAGGGGCTGATGATGCCTGCACCACC
>BAU001000002.1/383903-383842 Paenibacillus sp . JCM 10914 DNA, contig: contig_2.
ATCAAATGAGGCGATGGAGTTCGCCATAACCGCTGTTTGCAGCTAATGACTCCTACCAGA
GG
>MLQM01000022.1/19259-19199 Mycobacterium sp . NE-TNMC-100812
PROKKA_contig000022, whole genome shotgun sequence.
TCGAGGTCAGGCGATGGAGTTCGCCTCAACTGCCACACCGGCTGATGGCTTCTACGACGT
G
>FQXE01000008.1/6418-6352 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_108
TGCACCGACGGAGATGGCATTCCTCCGCTAACCGTCGCTATCAGCGACTGATGATGCCTA
CACCTTC
>MGUD01000025.1/66406-66329 Elusimicrobia bacterium GWC2_61_19
gwc2_scaffold_560 , whole genome shotgun sequence.
GCGTCCCCCGGCGATGATGTCCGCCGCCCTGTAAAGACAGGGAAACCGTCCGCAGGGACT AATGACCTCTACTCAAGG
>MFKF01000231.1/5169-5237 Candidatus Handelsmanbacteria bacterium
RIFCSPLOWO2_12_FULL_64_10 rifcsplowo2_12_scaffold_31199, whole genome shotgun sequence.
TCGCCTGCGGGAGATGGCATTCCTCCCGTAAAAAGAACCGCCGGCCCGGCTGATGATGCC
TACGTGATC
>CP002919.1/1745084-1745154 Leptospirillum ferriphilum ML-04, complete genome . APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AGGAAGACAGGCGATGGGGTTCGCCAAAACCGCCCCGGAGAACAGCGGGAGCTGATGACC
CCTACTCCATT
>CCNA01000034.1/638081-638020 Mesorhizobium sp . SOD10 genome assembly, contig: MPLSOD_Contig_4
CTCAAAAATGGGGATGGGGTTCCCCCGAAACCGCCCTTGTGGCTGATGACTCCTGCCAGG
CG
>CP015736.1/1032604-1032668 Shinella sp . HZN7, complete genome.
GTGTGGGATGGGGATGGAGCTCCCCGACAACTGCCGTGTGACCGGCTGATAGCTCCTGCC
AAGAC
>FUWZ01000006.1/324561-324634 Chitinophaga eiseniae strain DSM 22224 genome assembly, contig: Ga0070509_106
TGCGACACAGGAAATGGTGTCTTCCTGCTCAGAACCGTTCCGGGTCCTCCGGAACTGATG
GCGCCTACAAAACA
>LMLI01000016.1/324384-324322 Serratia sp. Leaf50 contig_4, whole genome shotgun sequence.
AAAGCTGTTGGAGATGACATTCCTCCCAAACCGCCTTCACCGGCTGATGATGTCTACGAA
ATC
>CP014854.1/931547-931620 Thermococcus celer strain Vu 13, complete genome .
GTTCCACCGGGCGATGGCGTCCGCCCGGGCTTCGAGCCGAACCGCCCTCCAGGGCTGATG
ACGCCTGTTCTCCG
>CP000141.1/1888854-1888795 Carboxydothermus hydrogenoformans Z-2901, complete genome.
AAATTTAATGGCGATGGGGCTCGCCTAACGCTCATCTGAGCTAATAGCTCCTACCGAAAG
>MTKC01000003.1/5836372-5836309 Streptomyces griseofuscus strain NG1-7 NG1-17_1_1, whole genome shotgun sequence.
ACGGAGATGGGCGATGAGGCTCGCCCTTGAGCGCACGTTCCGTGCTGATGGCCTCTGCGA
GCAC
>FR898362.1/2736-2661 Dorea longicatena CAG:42 genomic scaffold, scfl81
CTTTTAGAAGGGAATGAAGTTCTCCCTTAGTGATCATACTAGAACCGCTTATAAAGCTGA
TGACTTCTGCGAATAA
>KI530763.1/81247-81181 Acinetobacter brisouii CIP 110357 genomic scaffold adgTE-supercont2.2 , whole genome shotgun sequence.
ACCTGATCAGGAGATGGCATTCCTCCTGTAACAAACCGCCATATCGGCTAATGATGCCTA CGTCACC
>FTMG01000002.1/526443-526507 Mucilaginibacter lappiensis strain ATCC BAA- 1855 genome assembly, contig: Ga0104715_102
CGTATCAAAGGAAATGGTGTCTTCCTACTTAACCGCCCTGAAAAGCTGATGGCGCCTACA
ATGAA
>LMPL01000001.1/915709-915780 Williamsia sp . Leaf354 contig_l, whole genome shotgun sequence.
GACGAACACGGTGATGGATCCCGCCGGAGCATCTGCTCGAACCGCCACACGGCTGATGGT
TCCTACCCATGT
>KE340353.1/613687-613620 Acinetobacter rudis CIP 110305 genomic scaffold acVBL-supercontl .4 , whole genome shotgun sequence.
ATTAACAAAGGAGATGGTAATCCTCCTTAAACAAACCGCCAATTCTGGCTGATGATGCCT ACGTTTCC
>MJHW01000002.1/695258-695328 Roseburia sp . 831b contig000002, whole genome shotgun sequence.
AAATTGACAGGGAATGAAGTTCTCCCTTGGGAAACCTAAACTGCTTATTGGCTGATGACT
TCTACGATTAT
>AJFI01000022.1/152647-152573 Mycobacterium xenopi RIVM700367 contig22, whole genome shotgun sequence . APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ACTTGCACCGGCGATGGAGCCCGTCCGGAAGTGCGACTTCCGAACCGCCACAAGGCTGAT
GGCTTCTACGAACGA
>MTCM01000001.1/2219788-2219850 Sphingobacterium sp . CZ-UAM QB.A_0, whole genome shotgun sequence.
TGAATATATGGCAATGATGTCTGCCTTGAACTGCCCTAATAAGCTGATGACGTCTACTTT
AAG
>AM181176.4 /2081856-2081784 Pseudomonas fluorescens SBW25 complete genome
CGGCTCATTGGAGATGGCATTCCTCCATTAACCAACCGCTGTGCCCGTAGCAGCTGATGA
TGCCTACAGAAAC
>LGCL01000023.1/39163-39094 Ornatilinea apprima strain P3M-1 contig_9, whole genome shotgun sequence .
ATATCAATAAGCGATGAGGCTCGCTCAAAAGATATAAATCGCTCAAACAGCTAATGGCCT
CTACCAGGTT
>AASG02041893.1/549-637 Ricinus communis cultivar Hale ctg_1100012285214, whole genome shotgun sequence .
AAATCCAACGGTAATGGATTTCTGCCGGGCCTTGGTGGCCGAACCGCTGCTCTGATTTTC
CCAGGGAAGCTGATGACTCCTACTCAACA
>MVHR01000005.1/163180-163106 Mycobacterium heidelbergense strain DSM 44471 NODE_5_length_243907_cov_56.1658, whole genome shotgun sequence. TGGTCGCGAGGCGATGGATCTCCGCCTGGGTAATCCGCCCAAACCGCCGGCCCGGCTGAT GGCTCCTGTGGTGCT
>FQXE01000023.1/20864-20798 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_123
GCACCCAACGGAGATGGCATTCCTCCGCTAACCGCCGCGCTCTGCGGCTGATGATGCCTA
CAAGTTC
>LMI001000004.1/100573-100491 Leifsonia sp . Root227 contig_4, whole genome shotgun sequence.
TGGGCGGACGGCGATGGATCTCGCCTGGGCAGCACCCGTGCCGTGCCCGAACCGCGTCGT
CCGCTGATAGTTCCTACGCTTCC
>CP001638.1/886398-886457 Geobacillus sp . WCH70, complete genome.
AGAGTAATAGGCGATGGAGTTCGCTAAACTATCGTTCGGGCTGATGACTCCTACCAGTAA
>LMSK01000001.1/52013-51910 Rhodanobacter sp . Soil772 contig_l, whole genome shotgun sequence.
GGCCGCGCGGGAGATGGCATGCCTCCCGTTCCGGAGCCTGCGGGCACACCGGGACAAACC
GCTTCTCGTGATCCCGCGTTGAAGTTGATGATGCCTGCTCCCCC
>MENT01000026.1/88105-88167 Bacteroidetes bacterium GWD2_45_23
gwd2_scaffold_79, whole genome shotgun sequence.
AACATTTTTGGCGATGGTATTCGCCCTAAACCGCCAGTTTTGGCTAATAATACCTACTAA TAA
>JGZR01000003.1/181277-181162 Bifidobacterium subtile strain LMG 11597 Contig03, whole genome shotgun sequence.
GGGCAGAACGGTGATGGGCCCCGCCTGGACGCATGTCCGAACCGCACAGTGTTGACAAGC
ATGGCGTTGATGAATGCAATGCCGACAAACACAGTGCTGATGGTTCCTACTCAATC
>FXAT01000003.1/191226-191297 Paraburkholderia susongensis strain LMG
29540 genome assembly, contig: Ga0139082_103
GATTTTTGCGGAGATGGCATTCCTCCCTAACTGCCGCTAGCAAGCTTTGCGGCTAATGAT
GCCTACAGGGTC
>LMSL01000010.1/6484-6391 Frateuria sp. Soil773 contig_18, whole genome shotgun sequence.
GTCGCACAGGGAGATGGCATGCCTCCCGTTCCCTCGTGGAACGAACCGCCCCACGCCTTG
ACCGATCGACGGGGTTGATGATGCCTGCATCACC
>MGDB01000134.1/599-662 Candidatus Schekmanbacteria bacterium GWA2_38_11 gwa2_scaffold_6985 , whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CCTGAAAATGGCGATGGGGTTCGCCAATAACCATCCATAATGGATTGATGACCCCTACTG
AAAC
>JXYS01000110.1/6185-6245 Acidithrix ferrooxidans strain Py-F3
AXFE_contig000110, whole genome shotgun sequence.
TTATGGACTGGCGATGGAGTTCGCCATAACTGCCACAATGGCTAATAACTCCTACATAGC
G
>HF991186.1/156700-156777 Coprococcus sp . CAG:782 genomic scaffold, scf290
TATAGACCAGGGAATGAGGCACTCCCTTGACAAGATTTGTCTAAACCGCCCGAAATGGCT
GATGGCTTCTGCATTTTA
>MEZG01000067.1/124-191 Candidatus Bathyarchaeota archaeon RBG_13_60_20 RBG_13_scaffold_5248 , whole genome shotgun sequence.
CTTGATCGAGGCGATGAGGCTCGCCTTTAACCGCCCTCTCGGTGAGGACTGACAGCTTCT ACACAAAC
>MGQ001000116.1/9048-9114 Deltaproteobacteria bacterium RBG_16_47_11 RBG_16_scaffold_21686 , whole genome shotgun sequence.
TTGAAACCTGGCGATGAAGTTCGCCTTAAATGTCCCCACGGAAGGGAATGATAACTTCTA CTGGTTA
>LEKT01000054.1/20510-20571 Megasphaera cerevisiae DSM 20462 scaffold_53, whole genome shotgun sequence .
AATTTCATAGGCGATGGAGTTCGCCTGTACATGACTATCGTGTTGATGACTCCTACCTTT
TA
>AYZL01000008.1/51973-52034 Lactobacillus floricola DSM 23037 = JCM 16512 strain DSM 23037 NODE_14, whole genome shotgun sequence.
AAAATAAACGGTGATGGTGTTCGCCATTAAATGCAAGTATTGCTGATGACACCTGTTGAA GC
>AFCH01000049.1/79202-79124 Gluconacetobacter sp . SXCC-1 Contig2957, whole genome shotgun sequence.
AATACGGGCGGCAATGGACTCTGCCTGATTTCAGTGATGAAATCGAACCGCCCGCCGGGC
TGATGATTCCTACCCCGCC
>GL834309.1/286554-286475 Clostridium symbiosum WAL-14163 genomic scaffold supercontl .5, whole genome shotgun sequence.
AATACAGTCGGGAATGAAGTTCTCCCCCAGTATTGTTTATACTGAAACCGCTTATTTGAG
CTGATGACTTCTGCGGATAC
>JXKH01000005.1/130635-130562 Enterococcus canis strain DSM 17029
Scaffold5, whole genome shotgun sequence.
TCTAGGACAGGGAATGAGGTTCTCCCTTGGTTCATACCTAAACTGCTTACAAAGCTGATG
ACTTCTGTTTTTTA
>FQVM01000008.1/12844-12783 Clostridium fallax strain DSM 2631 genome assembly, contig: Ga0131101_108
ATTTTAATAGGTTATGGAGTTAACCGTAAACCGCGAAAAAAGCTGATGACTCCTACATAA
AA
>JTHF01000066.1/5252-5325 Methylobacterium platani strain SE2.11
contig_66, whole genome shotgun sequence.
CACCGTCACGGCAATGGATTTCTGCCGGGCCATCGTGGCCGAACCGCCCTCGGGCTGATG
ATTCCTACCTTTTG
>CP007051.1/592759-592698 Desulfurella acetivorans A63, complete genome.
TACAAAAGTGGCGATGGAGTTCGCCAAAAATGCCTACCTAGGCTGATAACTCCTACAAAT
TA
>MEGL01000229.1/173-112 Rhodanobacter sp . SCN 68-63 ABT19_C0229, whole genome shotgun sequence.
TGACAGATGGGAGATGGCATGCCTCCCCGAACCGCCGCAAGGCTGATGATGCCTGGTTGA
CC
>CP002282.1/640124-640200 Ilyobacter polytropus DSM 2926 plasmid pILYOPOl, APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
complete sequence.
AAATAAAAGGGGGATGAAGCTCTCCCGGAAAAATTCCAAATGCTTCTAATTACGGAGATG
ATGGCTTCTACAGAAAA
>LGCK01000014.1/708435-708365 Leptolinea tardivitalis strain YMTK-2 contig_l, whole genome shotgun sequence.
AATCAATTAAGCGATGAAGCTCGCTCGGAAATAATCAAACTGTCTTTATGGCTGATAGCT
TCTGCCTTAAA
>CP014674.1/1095373-1095310 Kozakia baliensis strain DSM 14400, complete genome .
CGAGACAAAGGAGATGGGGTTCCTCCTTGTAACCGCCCGTAGGGCTGATGACTCCTATTC
GATT
>LGTC01000001.1/6340601-6340661 Pseudobacteroides cellulosolvens ATCC 35603 = DSM 2933 ctgl, whole genome shotgun sequence.
AATTAAATAGGTGATGGAGTTCGCCATTAACTGCTACTATGCTAATGACTCCTACATAAT
T
>FPJ001000018.1/52009-51945 Streptomyces atratus strain OK807 genome assembly, contig: Ga0007934_1018
GACACATCGGGCGATGAGGCTCGCCGTAACCGCGCCGATGCGGCGCTGATGACCTCTGCG
TGACC
>CP002039.1/2800602-2800675 Herbaspirillum seropedicae SmRl, complete genome .
TGCGCCGCAGGAGATGGCATTCCTCCTCAAACCGCCGCCAAGTCCAGCTTGGGGCTGATG
ATGCCTACAGTGAC
>LN681234.1/1342269-1342208 [Clostridium] sordellii strain JGS6382 genome assembly, chromosome: 1
AAACAAATCGGCGATGGAGTTCGCCATTAAATGCACAAGATGCTAATGACTCCTACTCTT
AT
>DF820472.1/170173-170112 Bacterium UASB270 DNA, scaffold:
UASB270_scaffold_l 0.
GTAGTATCAGGTGATGGAGTTCACCTCAAACCGCATGTTCTGCTGATAACTCCTACCAAT
TT
>MDCF01000005.1/56436-56498 Acidiferrobacter thiooxydans strain ZJ
Contig_102, whole genome shotgun sequence.
CGTCTTGCGGGAGATGGCAGCCTACCTATAACCGCCCTCGCGGCTGATGATGCCTACGAT
TTC
>CP009313.1/1026203-1026276 Streptomyces nodosus strain ATCC 14899 genome.
GGAGGGGGTGGCGATGAGGCTCGCCGCAACCGCACGACCGGAGCACCGACCGTGCTGATG
GCCTCTCCCTGACG
>CP000780.1/ 914829-914764 Candidatus Methanoregula boonei 6A8, complete genome .
TTTTTTAACGGCGATGAAGTCCGCCTCCAACCGCCACGTATCGTGACTGATGACTTCTAC
CGGGGA
>AZGA01000066.1/29758-29824 Lactobacillus composti DSM 18527 = JCM 14202 strain DSM 18527 NODE_114, whole genome shotgun sequence.
AAGTTAATAGGCGATGATGTTCGCCATAAAATAATTGTTCAAACAATTTGATGACGTCTA CTGTAAT
>LJGW01000451.1/13279-13342 Streptomyces nanshensis strain SCSIO 10429 scaffold449, whole genome shotgun sequence.
CAGTGTGCAGGTGATGGGGCTCACCGCAACCGCGGCCTGTGCCGCTGACGGTCCCTGGTC
CACG
>MNRC01000242.1/4607-4680 Clostridiales bacterium 36_14
Ley3_66761_scaffold_l 1528 , whole genome shotgun sequence.
AAAGAGAAAGGGAATGAAGTTCTCCCTCGAAGAGATTCGAAACCGCTTATTAAGCTGATG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ACTTCTGTGTAATG
>MNJR01000078.1/16660-16590 Acidobacteria bacterium 13_1_20CM_3_53_8 13_l_20cm_3_scaffold_30, whole genome shotgun sequence.
CACTACATAGGCGATGGAGTTCGCCACATAACCGTCCTGTTTTTGACAGGGCTGATAACT CCTGGTCATAT
>LAQU01000008.1/136473-136575 Burkholderia andropogonis strain ICMP2807 contig00008, whole genome shotgun sequence.
AGCCGCGCAGGAGATGGTGCTCCTCCTGGTTTAGACAAACCGCCCCGTCGATCGCTGAAT
ATATTATTCCGCAGGACGCGAGGCCGATGGCGCCTACAGTTTC
>MNHH01000267.1/1038- 976 Gemmatimonadetes bacterium 13_1_40CM_2_60_3 13_1_40cm_2_scaffold_105334 , whole genome shotgun sequence.
CTTTCAGCAGGCGATGGAGTTCGCCTTCAACCGCCCATTCGGGCTGATGACTCCTACCGA ATA
>CP001398.1/1580180-1580105 Thermococcus gammatolerans EJ3, complete genome .
AGGTGGTTGGGCGATGGCGTCCGCCCGGGGCTTCGGGCCCGAACCGCCCGTAGGGGCTGA
TGACGCCTGTTCTCCG
>CP002433.1/4140667-4140592 Pantoea sp. At-9b, complete genome.
CCCCCTTAAGGTGATGGCGTGCCACCTTCCCCAACCGCCGTTCTGGCAACAGACGGCTGA
TGACGCCTGACAACAC
>CP000099.1/4790369-4790299 Methanosarcina barkeri str. Fusaro, complete genome .
TATCCTATAGGTGATGAAGTCCGCCAAGATTTACAGTCTAAACCGCTTATGCTGATGACT
TCTACATTAAT
>CP013909.1/4555498-4555573 Hymenobacter sedentarius strain DG5B
chromosome, complete genome.
TCCAAACCCGGTGATGGGGTACCACCTTGAACCGCCAGTGCCTTTGGGGCGCTGCGCTGA
TGACTCCTACGCCTAC
>MSDV01000018.1/183941-183880 Burkholderia sp. SRS-W-2-2016 contig42914, whole genome shotgun sequence .
TGCGCGGCTGGAGATGGCATTCTCCATTAACCGCCCTTGGGGCTGATGATGCCTGCTACG
CC
>MGDB01000157.1/8681-8618 Candidatus Schekmanbacteria bacterium GWA2_38_11 gwa2_scaffold_9836 , whole genome shotgun sequence.
TTTAAAAATGGCGATGGGGTTCGCCTCTAACCATCCATAAAGGATTGATGACTCCTACTG AAAC
>JH651379.1/2048345-2048285 Joostella marina DSM 19592 genomic scaffold Joomascaffold_l , whole genome shotgun sequence.
ACATACAAAGGCGATGGAGTTCGCCAAAACCGCCCAAAAAGCTAATGACTCCTACTCAAT
T
>AGUD01000045.1/10920-10984 Patulibacter medicamentivorans strain Ill contig45, whole genome shotgun sequence.
CGGCCGCAGAGCGATGAGGCCCGCTCCACAACCGCGGCGTGACCGCTGATGGCTTCTACG ACCAC
>CP016282.1/3754118-3754190 Cryobacterium arcticum strain PAMC 27867 chromosome 1, complete sequence.
CTTTTGCGTGGTGATGGATCCCACCAGAGCACCCGCTCGAACCGCCGATTGCGCTGATGG
TTCCTACCGACTT
>KE136354.1/938710-938782 Enterococcus dispar ATCC 51266 genomic scaffold acpMG-supercontl .1 , whole genome shotgun sequence.
GAATATTCAGGGAATGATGATCTCCCATGGACTTATCCAAAACCGCTATTAAGCTAATGA CTTCTACAAGTTT
>NJGU01000005.1/267616-267544 Herbaspirillum sp . HZ10 Scaffold3_l, whole APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
genome shotgun sequence.
ATGCGCGCCGGAGATGGCATTCCTCCCCGAACCGCCGCCGGCATCCGCCGGGGCTGATGA
TGCCTACGGATAC
>MKRJ01000013.1/39700-39638 Rhizobiales bacterium 65-79
SCNpilot_cont_1000_bf_scaffold_187 , whole genome shotgun sequence.
TCGCCCGATAGGGATGGGGTTCCCCCGACAACCGCCGAAAAGGCTGATGACTCCTGCCAG GGC
>LGPB01000103.1/27152-27094 Bacillus galactosidilyticus strain PL133 scaffold000103, whole genome shotgun sequence.
TTTCGTTTAGGCAATGGAGTTCACCAAAACTGCTGAACGCTAATGACTCCTGCCGAAAA
>CYSP01000007.1/25074-25012 Propionispora sp . 2/2-37 isolate 2/2-37 genome assembly, contig: 2/2_contig7
TGAATAGCGGGCGATGGAGTTCGCCTTGAATTCCGTATAAACGGTAATGACTCCTACCAC
CTA
>LGCL01000045.1/11910-11982 Ornatilinea apprima strain P3M-1 contig_4, whole genome shotgun sequence .
AATTAAAAAAGCGATGAGGCTCGCTTGAGATGTAATCAAACCGCCGTTTTTGGCTGATAG
CCTCTGTCTTATT
>ABVL01000003.1/298778-298860 Chthoniobacter flavus Ellin428 ctg77, whole genome shotgun sequence.
GCCCGTAACGGCAATGGATTTCTGCCGCCCGAGTGATCGGGGAACCGCTTGGTAACGGCC
AGGCTGATGATTCCTGTTTTCAC
>MGTS01000012.1/84755-84676 Elusimicrobia bacterium GWA2_51_34
gwa2_scaffold_159, whole genome shotgun sequence.
TAAATTAACGGCGATGATGTCCGCCTCCCTGGCCTCACCGGGGAAACCGCTCACAATGGG CTGATGACCTCTATCTAAAA
>AYS001000017.1/378466-378525 Clostridium argentinense CDC 2741
U732. Contig249 , whole genome shotgun sequence.
ATATCACAAGATGATGGAGTTCACCATAATCGCAGAAATGCTAATGACTCCTACAAGATG
>MEFQ01000335.1/3112-3173 Xanthomonadaceae bacterium SCN 69-48
ABS98_C0335, whole genome shotgun sequence.
GCCCATACGGGAGATGGCATCCTCCCCGAACCGCCGTCCCGGCTGATGATGCCTGCCAAC
CC
>CELZ01000050.1/60745-60807 Moorella glycerini strain NMP genome assembly, contig : M_glycerini_NMP_DRAFT_scaffold-50
ATAGTTTTAGGCGATGGAGTTCGCCTTAAATCTGCAGCAAGCAGTGATGACTCCTACCAA
AAA
>LT838272.1/2866416-2866355 Thermanaeromonas toyohensis ToBE genome assembly, chromosome: I
TATAAAAGTGGCGATGGAGCTCGCCTAATGCTCAGGTTTGAGCTAATGGCTCCTACCGTA
AT
>DS995355.1/552971-552900 Clostridium hiranonis DSM 13275 Scfldl genomic scaffold, whole genome shotgun sequence.
ATACAACTAGGGAATGAATGTCTCCCTTGGGAAACCTAAACCGCTTATCGAGCTGATGAC
TTCTGCAACCAC
>CP001699.1/4998418-4998502 Chitinophaga pinensis DSM 2588, complete genome .
CCTTTTACAGGAAATGGTGTCTTCCTGCACAGAACCGTCTTCGCCCTGATCAATCTCCGG
CAAGTCTGATGGCGCCTACAAATAA
>BATJ01000006.1/80403-80467 Vibrio proteolyticus NBRC 13287 DNA, contig: VPR01S06.
CCTAAGCAAGGTGATGGGGTTCCACCTATTTAACCGCCGCTTGGGCTGATGACTCCTACA
GAAAC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>AP008232.1/1344619-1344556 Sodalis glossinidius str. 'morsitans ' DNA, complete genome.
CTTTCCCCAGGAGATGAAATTACTCTTTATAACCGCCGCTCTGGCTGATGATGTCTACGT
TCAT
>MAST01000004.1/30569-30488 Humibacillus sp . DSM 29435 contigl2, whole genome shotgun sequence.
GTCGGAGACGGCGATGGATCCCGCCAGGGTGTGAGAGCCTCACCCCCAAACTGCCACCTC
GGCTGATGGTTCCTGACTCGTC
>HE774682.1/1106424-1106358 Flavobacterium indicum GPTSA100-9 complete genome
GACAAAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTAAAAAGCTGATGACGCCTG
ATTAATT
>LMJN01000001.1/1254218-1254143 Lysobacter sp. Root96 contig_l, whole genome shotgun sequence.
ATCGGTGCAGGAGATGGCATTCCTCCTGTCCCCACCGGGGCAAACCGCCGCCTTGGCTGA
TGATGCCTGCTGACCC
>CP003372.1/3516099-3516203 Natrinema pellirubrum DSM 15624, complete genome .
TCGTATCCGGGCGATGGGGCCCGCCTGACCCAACTGCCGACCGCCGTTCCGAAGCCGTGA
CGACCCACATCGCGGCGACGGTCGGCTGACGGTCCCTGCCACCAC
>AGFR01000009.1/332139-332206 Commensalibacter intestini A911 74_9, whole genome shotgun sequence.
TAAAGTTTAGGAGATGATATCCTCCTTGTAAACCGCCGCATTTTGCGGATGATGATATCT
ACGCGGAG
>JZQY01000069.1/46772-46848 Nitrospira sp. OLB3 UZ03_NOB00100CONTIG000070, whole genome shotgun sequence .
GCTTCAACAGGTGATGGAGTTCACCGAGAACCGCCGGTCCACGATGCACATGGACGGCTG
ATAACTCCTACTGGCAC
>KB233222.1/1770946-1770876 Providencia burhodogranariea DSM 19968 genomic scaffold scaffoldOOOOl, whole genome shotgun sequence.
TTAGTTTGAGGCGATGGCATTCCTCCCGTTTAAGAACTGTCCTAATAAGGACTGATGATG CCTACGTTAAC
>FQUN01000003.1/227081-227017 Leeuwenhoekiella marinoflava DSM 3653 genome assembly, contig: EJ63DRAFT_scaffoldOOOOl .1
TCAATTACAGGCAATGAGGTCTGCCTCAAACCGCTTCTTATGAAGCTGATGACTTCTACT
AAAAA
>LMLF01000003.1/12954-12878 Sphingomonas sp . Leaf42 contig_ll, whole genome shotgun sequence.
ATCGACTTTGGCAATGGATTTCTGCCGGGCAATTTGGCCGAACCGCTCTGATAAGAGCTG
ATGATTCCTACTTGGCG
>GG657592.1/978902-978818 Clostridium asparagiforme DSM 15981 genomic scaffold Scfld6, whole genome shotgun sequence.
CAGCGCAAAGGGAATGAGGTTCTCCCGTGGCGAAAGCCAGAACCGCTTGTAAGGCAGATT ATGAGCTGATGACTTCTGTGATGAT
>KB946300.1/215687-215748 Enterococcus pallens ATCC BAA-351 genomic scaffold acvKb-supercontl .1, whole genome shotgun sequence.
TCGTGAATAAGGGATGGTGTTCCCGATAACTGCTAAAATTAGCTGATGACACCTGTTTAA CT
>CYGY02000045.1/572-637 Burkholderia sp . STM 7183 genome assembly, contig: CYGY01000045
TCGATTGACGGAGATGGCATTCCTCCCGTAACCGCCGGTTGGCCGGCTAATGATGCCTAC
GGTTCC
>CP001875.2/4203225-4203149 Pantoea ananatis LMG 20103, complete genome. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CGCGCATAAGGTGATGGCGTTCCACCTTTCCCAACCGCCTGCAGGCGTAACCTGCGGCTG ATGACGCCTGACATTAT
>LAKY01000069.1/17972-18036 Clostridiales bacterium PH28_bin88 ph28_412, whole genome shotgun sequence .
AACTATCATGGCGATGGGGTTCGCCATAAACACTGCCCCTGGCGGTTGATAACCCCTACC
AAATT
>AP012204.1/4277444-4277518 Microlunatus phosphovorus NM-1 DNA, complete genome .
GTGAATCGCGGCGATGGCGTCCGCCAGGGGTAGCTCCCCGAACCGCCCCGCAGGGCTGAT
GACTCCTGAGACGAT
>MGZG01000181.1/10510-10581 Gemmatimonadetes bacterium
RIFCSPL0W02_12_FULL_68_9 rifcsplowo2_12_scaffold_21457, whole genome shotgun sequence.
CTCGGCGCCGGCGATGGAGTTCGCCGTCAACCGTCCCCTGCCATTCAGTGGACTGATGAC
TCCTGGAACCCT
>CP002588.1/1355815-1355752 Archaeoglobus veneficus SNP6, complete genome.
CTGTATGCGGGTGATGGCGTCCACCCTTAACCGCCCGGCAAGGGCTGATGACGCCTGTTT
TCGG
>MEOG01000126.1/84306-84366 Bacteroidetes bacterium GWF2_35_48
gwf2_scaffold_871 , whole genome shotgun sequence.
TTCCATTCTGGTGATGGAGTTCGCCATAAACTGCAATTGTGCTGATAACTCCTGCAATTT
T
>MFGX01000109.1/13660-13738 Candidatus Fraserbacteria bacterium
RBG_16_55_9 RBG_16_scaffold_6090, whole genome shotgun sequence.
TTATCGCTGGGCGATGGAGTCCGCCGGGGAGCGCCTGCTCCCAAACCGTCCCAACTGGAC TGATGGCTCCTACCTGAAC
>FTPU01000060.1/8776-8845 Chryseobacterium bovis DSM 19482 genome
assembly, contig: LX71DRAFT_scaffold00060.60
ATTCGCAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTTTCAAAAAGCTGATGACGC
CTGATTAATG
>LJJC01000004.1/3650072-3650013 Bacillus shackletonii strain LMG 18435 superll, whole genome shotgun sequence.
AACTCAATAGGCGATGGGGTTCGCCAAAAACGCTGAAAAGCTAATGACTCCTACCAGTAT
>FWXS01000009.1/123184-123250 Moheibacter sediminis strain CGMCC 1.12708 genome assembly, contig: Ga0171597_109
AAAATAAAAGGAAATGGTGTTCTTCCTTACCCAAACCGCGTAAATCGCTGATGACGCCTG
ATTTTTA
>CP014675.1/5635-5570 Kozakia baliensis strain DSM 14400 plasmid
pKB14400_l, complete sequence.
TCCCGACCAGGGGATGGAGCTCCCCCTTCAACCGCCCTCGCAAGGGCTGATGGCTCCTAC
CGCGAC
>MDER01000075.1/31306-31368 Paenibacillus sp . TI45-13ar contig00075, whole genome shotgun sequence.
AATAAAATCGGCGATGGAGTTCGCCATTGACTGCTGATTTCAGCTAATGACTCCTACTTG
TTT
>LXGI01000116.1/48114-48185 Paraburkholderia tropica strain P-31
P3 l_contig_7 , whole genome shotgun sequence.
GGCACTTTCGGAGATGGCATTCCTCCCGTAACCGCACCGCTTCCCGGCGGCGCTGATGAT
GCCTACGAGTTC
>MGQI01000010.1/2334-2261 Deltaproteobacteria bacterium RBG_13_58_19 RBG_13_scaffold_11141, whole genome shotgun sequence.
AATCAAACAGGCGATGAAGTCCGCCGTAACCGCCCTGCCAGCTATGACGCGGGGCTGATG ACTTCTGTCCATTT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>MDET01000028.1/31339-31401 Pseudaminobacter manganicus strain JH-7 JH- 7_scaffold34 , whole genome shotgun sequence.
GGCTTCGATGGGAATGGAGTTCTCCCGAAAGCGCCTGTGAGAGCTGATGACTCCTATCGC
GGA
>KK073872.1/385211-385150 Cloacibacillus evryensis DSM 19522 genomic scaffold CloevDRAFT_Scaffoldl .1 , whole genome shotgun sequence.
GCATACGAGGGCGATGGAGTTCGCCCGAAATTCGTTTTTACGATGATGACTCCTGTCCAA AG
>DF820460.1/ 629916-629977 Bacterium UASB14 DNA, scaffold:
UASB14_scaffold_6.
ACAAAATAAGGTGATGGGGTTCACCTCTAACCGCACGCATCGCTGATAACTCCTACAAAC
TA
>CP002573.1/2024788-2024849 Acidithiobacillus caldus SM-1, complete genome .
CTTTCCAACGGCGATGAGGTCCGCCTGGAACCGCCTTCGTGGCTAATGACTTCTACCGAA
TC
>LHOX01000018.1/204801-204729 Enterococcus sp. RIT-PI-f contig_l, whole genome shotgun sequence.
AGCGAAAAAGGAAATGAATGTCTTCCTCAGTCTTAACTGAAACCGCAAACTTGCTAATGA
CTTCTACCACTTA
>LMSL01000040.1/117726-117792 Frateuria sp . Soil773 contig_8, whole genome shotgun sequence.
TTGGAAACGGGAGATGGCATTCCTCCCCCGACCCAACCGCCGCAAGGCTGATGATGCCTA
CCCGTCG
>MJID01000007.1/419613-419688 Roseburia sp . 499 contig000007, whole genome shotgun sequence.
ATCATTGTTGGGAATGAGGTTCTCCCTTAGTGATAACACTAAAACCGCTTATTAAGCTGA
TGACTTCTGCGAATTA
>FUKX01000034.1/78618-78549 Sphingobacterium faecium PCAi_F2.5 genome assembly, contig: Scaffold8
TAGCAACATGGTGATGGTGTTCACCACGAACTATTTATTGGACGAATAAATTAATGACGC
CTACCAAACG
>CP000789.1/3337036-3337110 Vibrio harveyi ATCC BAA-1116 chromosome I, complete sequence.
AATGGGCACGGTGATGGGGTGCCACCGGAATCGAAAGATTCGAACCGCTTTAAAGCTAAT
GACTCCTACAGAAAC
>FXAT01000003.1/190021-190087 Paraburkholderia susongensis strain LMG 29540 genome assembly, contig: Ga0139082_103
GAATTCAGTGGAGATGGCATTCCTCCCGAACCGCCGGTCAAGTCCGGCTGATGATGCCTA
CGAGTTC
>CP001634.1/455111-455170 Kosmotoga olearia TBF 19.5.1, complete genome.
ATTTCTTTGGGCGATGGGGTCCGCCCTAATTGCCAAAAGGCTGATGACTCCTATCTCTAT
>CP002629.1/3040640-3040713 Desulfobacca acetoxidans DSM 11109, complete genome .
GTGCAAATAGGTGATGAGGTCCACCATAACTGCCTTGCCGAGGAATTGGCCGGGATGATG
ACTTCTGCCCCACA
>BAMX01000001.1/9377-9445 Acetobacter orientalis 21F-2 DNA, contig:
Abor_001.
AACCTTACAGGGGATGGAGTACCCCCTTATAACCGCCCGCTTGTAAGGGCTGATGACTCC
TGCCGTAGT
>AYZV02266241.1/1351-1289 Spinacia oleracea cultivar SynViroflay
scaffold87624. con0007.1, whole genome shotgun sequence.
TGGGACAATGGCGATGGAGTACCGCCAAAACCGCCCAAACAGGCTGATGACTCCTATGAT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ttt
>MNDV01000274.1/341-401 Alphaproteobacteria bacterium 13_2_20CM_2_64_7 13_2_20cm_2_scaffold_5513 , whole genome shotgun sequence.
GAAACGAACGGGGATGGGGTCCCCCACAACCGCCGCGCAGGCTGATGACTCCTACCGGTT
A
>CP014167.1/3902176-3902116 Paenibacillus yonginensis strain DCY84 chromosome, complete genome.
AACACGATAGGCGATGGAGTTCGCCATAACCGCCCTTTGGGCTGATGACTCCTACCAGTT
T
>FWXH01000002.1/821201-821143 Clostridium acidisoli DSM 12555 genome assembly, contig: EJ29DRAFT_scaffoldOOOOl .1
TTTTTTATAGGTGATGGAGTTCACCTTAAATGCTTAGAGCTGATGACTCCTATAGGTTT
>MGUI01000007.1/127981-127858 Elusimicrobia bacterium GWD2_63_28
gwd2_scaffold_122 , whole genome shotgun sequence.
ATTAATAACGGCGATGATGTCCGCCGCCCCGGCCGCCCGGGGAAACCGCCCGCGAGGCCG TCTGAGGCGCATAGCGCTTGTTACGGCACGGCCATAGGCCAGGGCTGATGACCTCTACCC CCTA
>LMQJ01000004.1/88217-88293 Bradyrhizobium sp. Leaf396 contig_12, whole genome shotgun sequence.
CACACAGATGGGGATGGAGTCCCCCGACAACCGCCCGGCGGTCGATGACCTCGTGGGCTG
ATGACTCCTGCTTGAGG
>MGSY01000755.1/450-382 Deltaproteobacteria bacterium RIFOXYB2_FULL_66_7 rifoxyb2_full_scaffold_71842_curated, whole genome shotgun sequence.
CGTGAACCCGGCGATGGAGTTCGCCGTGAACCGCCTATGTCGATAGAGGCTGATGACTCC TACCCTACA
>MEDJ01000005.1/88809-88741 Acidobacteria bacterium SCN 69-37 ABS36_C0005, whole genome shotgun sequence .
GTGTCGTCTGGCGGTGGAGTCCGCCTGAACCGCCGTTCCGACCGGACGGCCAATGACTCC
TGCGGGCCC
>MNXM01000330.1/10724-10651 Armatimonadetes bacterium CG2_30_59_28 cg2_3.0_scaffold_2174_c, whole genome shotgun sequence.
GGCGTCAATGGCGATGGAGTTCGCCCGTCCGAACAGACAAACGGCCTCTGATGACTGATG ACTCCTGATCCCGA
>LGUG01000004.1/3070818-3070756 Aneurinibacillus migulanus strain DSM 2895 super7, whole genome shotgun sequence.
AGTTGTATAGGCGATGGAGTTCGCCGTAACCGCTGACTGACAGCTAATGACTCCTACCGG
ACG
>CP006965.1/512354-512279 Thermococcus paralvinellae strain ESI
chromosome, complete genome.
GTTAACTCGGGCGATGGCGTCCGCCCGGGCTTCGAGCCCGAACCGCCCGCAAAGGGCTGA
TGACGCCTGTTCTAAC
>CP000020.2/35443-35510 Vibrio fischeri ES114 chromosome I, complete sequence .
CCAAATACGGGTGATGGATTTCCACCTTTAACCGCTCTATTTTTAGAGATAATGATTCCT
ACTATAGC
>KZ248224.1/8923-8847 Rhizobium yanglingense strain CCBAU 01603 genomic scaffold C573, whole genome shotgun sequence.
ACGCGCAACGGTAATGGATTCTGCCGGGCATTACGCCAAACCGCTTCCCCGACGAAGCTG
ATGACTCCTACTCAACG
>MHXP01000105.1/6885-6824 Planctomycetes bacterium GWA2_39_15
gwa2_scaffold_29798, whole genome shotgun sequence.
TGATAATAAGGCGATGGAGTTCGCCCAACTGCTACGTTGTAGCTGATAACTCCTATTGAA GC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>MNXE01000063.1/20102-20035 Hydrogenophilaceae bacterium CG1_02_62_390 cg_0.2_subl0_scaffold_23_c, whole genome shotgun sequence.
TTCAAAGCAGGAGATGGCATACCTCCTAATAACCGCCTGGGTTCCCGGCTGATGATGCCT ACGCGAAC
>AKVJ01000022.1/194416-194354 Pelosinus fermentans B4 ctg2, whole genome shotgun sequence.
GTTACTGGGGGCGATGGGGTTCGCCATTAAATCTGCTGATGCAGTAATGACTCCTACCAA
CAT
>MFIX01000182.1/183-247 Candidatus Glassbacteria bacterium
RIFCSPL0W02_12_FULL_58_11 rifcsplowo2_12_scaffold_44438, whole genome shotgun sequence.
CAATAAACAGGGGATGGAGTCCCCCTTCAAGCGCTCTTTCCAGGGCTGATGGCTCCTACC
GCGGG
>CP000673.1/304406-304465 Clostridium kluyveri DSM 555, complete genome. AAATAATAAGGTGATGGAGTTCACCATAACCGCAGAAATGCTTATGACTCCTACAAATAA
>AEXM01000013.1/22000-21927 Anaerococcus prevotii ACS-065-V-Coll3 contig00044, whole genome shotgun sequence.
ATATATTTAGGAAATGAAGTTCTTCCCAAAATTTATATTTTAAACCGCTATAAGCTAATG
ACTTCTGCACTTGC
>LT853885.1/2227198-2227273 Xanthomonas fragariae strain PD5205 genome assembly, chromosome: 1
TAGATGCTCGGAGATGGCGTTCCTCCGCGCTGTTCCATCAACCGTAGCGATCGCTACTGA
TGACGCCTACAAAAAC
>MKWJ01000020.1/48107-48046 Sphingomonas sp . 67-41
SCNpilot_cont_300_bf_scaffold_301 , whole genome shotgun sequence.
ATAGGGTTCGGGGATGGAGTTCCCCATCAACCGCCGCAAAGGCTGATAACTCCTACCAGA AC
>LGTC01000001.1/3350936-3350876 Pseudobacteroides cellulosolvens ATCC 35603 = DSM 2933 ctgl, whole genome shotgun sequence.
AATTAAATAGGTGATGGAGTTCGCCATTAACTGCTACTATGCTAATGACTCCTACATAAT
T
>LGHL01000085.1/63834-63757 Bradyrhizobium sp. NAS80.1
Bradyrhizobium_NAS_80. I_c85, whole genome shotgun sequence.
CGCACAGATGGGGATGGAGTCCCCCGACAACCGCCCGGTGGTCGATGACCGTAGTGGGCT
GATGACTCCTGCTTGAGG
>MGVC01000049.1/51974-51896 Elusimicrobia bacterium RIFOXYA2_FULL_39_19 rifoxya2_full_scaffold_350, whole genome shotgun sequence.
TGTTTCACTGGTGATGGATCTCCACCTAACCATTCTCTCAGGTTAAACCGCCCTAAAAGC
TGATGGTTCCTACTAAAAA
>CP007440.1/1559223-1559139 Rhodoplanes sp . Z2-YC6860, complete genome.
TGCGCCGATGGGAATGGGGTTCTCCCGATAACCGCCCGGCGAACAGACAATCCAGGTTCG
TTGCGCTGATGACTCCTACCGCTGA
>BATA01000004.1/44861-44927 Halarchaeum acidiphilum MH1-52-1 DNA, contig: HALO04.
CACGGTAGCGGCGATGGAGTCCGCCCTTTAACTGCCGGGGACGCCGGCTGATGACTCCTG
TGGAGAA
>MGVP01000586.1/3611-3672 Elusimicrobia bacterium RIFOXYB2_FULL_62_6 rifoxyb2_full_scaffold_5602, whole genome shotgun sequence.
ATTTAGCAAGGCGATGGAGTTCGCCTTAACCGCCCGAAAGGGCTGATGACTCCTGCGATA
CG
>GG729830.1/772508-772439 Bifidobacterium breve DSM 20213 genomic scaffold ScfldO, whole genome shotgun sequence.
CTGGGATTCGGCGATGGGACCCGCCTGAACTTGTTTCGAACCGCACACTGCTGATAGTTC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CTACACACGA
>CP000855.1/737518-737592 Thermococcus onnurineus NA1, complete genome.
GGCTTTTCGGGCGATGGCGTCCGCCCGGGCTTCGAGCCGAACCGCCCGTTTTGGGCTGAT
GACGCCTGTTCTCTG
>CP003591.1/2313962-2314073 Geitlerinema sp . PCC 7407, complete genome.
TAGAATCTAGGCGATGGAGCTCGCCATAACCGCCGTCCGGCGATCGCGCTAAAATTTGCG
CCGACCTTTTTTCTAAACTGGCTTTGGACTGGCTGATGGCTCCTACTTGTTT
>APMY01000064.1/29715-29815 Rhodococcus rhodnii LMG 5362
LMG_5362_contig070 , whole genome shotgun sequence.
CTTGGCTCCGGCGATGGATCCCGCCAGGGTTCGATTGTCGGCCGCGTCGGCGGCCGGCGC
GGCCCGAACCGCCGACACCGGCTGATGGTTCCTTTCCGATC
>LBHV01000005.1/188693-188766 Rhizobium sp . LC145 LC145_contig5 , whole genome shotgun sequence.
TCGCTCCACGGCGATGGATTTCTGCCGGGCCACACGGCCAAACCGCCCAGAGGGCGGATG
ATTCCTACTCGGTT
>NHOT01000002.1/395628-395696 Pedobacter sp . AJM contig02, whole genome shotgun sequence.
TTGCAATACGGAAATGGTGTCCTTCCGATTCAACCGCTTTTTTCAAAAGCTGATGGCGCC
TACAATAGA
>MPPJ01000001.1/323127-323065 Ochrobactrum sp. P6BS-III
NODE_l_length_445701_cov_32.571, whole genome shotgun sequence.
ATGCATGATGGGAATGGGGTTCTCCCGAAACCGCCGCTTCTGGCTGATGACTCCTGATCG AAT
>JYMR01000028.1/78932-79009 Bradyrhizobium sp. LTSP849 NODE_28, whole genome shotgun sequence.
CGCACAGATGGGGATGGAGTCCCCCGACAACCGCCCGACGGTCGATGATCGTGCTGGGCT
GATGACTCCTGCTTGAGG
>MEGC01000066.1/1080-1174 Novosphingobium sp . SCN 63-17 ABT10_C0066, whole genome shotgun sequence.
CGCGTCCAAGGCGATGGATTTCCGCCGGGCCTTTGGGCCGAACCGCCCGGCCCTAGGGCA
ACTTGGGGTCACAGGCTGATGATTCCTACCTTTCG
>FCNS01000019.1/7032-7129 Clostridiales bacterium CHKCI001 isolate CHKC1 genome assembly, contig: {contigl9}
ATGATGCAAGGGAATGAAGTTCTCCCTTAGTGAATGATATTGGAACATAAAATAACCAAA
CACTAGAACCGCTTAATGCTGATGACTTCTGCTATAAA
>MKTM01000033.1/104869-104931 Flavobacterium sp . 40-81
scnpilot_p_inoc_scaffold_71 , whole genome shotgun sequence.
TACAAATACGGCAATGATGTCTGCCTTGAACCGCCTTAAAAAGCTGATGACGTCTATTGA
ATG
>FUXA01000010.1/84794-84717 Eubacterium ruminantium strain ATCC 17233 genome assembly, contig: EI46DRAFT_scaffold00008.8
ATTGGTATTGGGAATGAGGTCTCCCAAAGCAGCATGCCTGCTTAAACCGCTTATTAGGCT GATGACTTCTGCATCTTT
>CP002351.1/495496-495436 Pseudothermotoga thermarum DSM 5069 chromosome, complete genome.
ATTAATTTGGGCGATGGAGTCCGCCTTTAATAGCCATGAGGCTGATGACTCCTACAGGTT
G
>JRNV01000011.1/18234-18295 Paenibacillus sp . P1XP2 CM49_contig000011, whole genome shotgun sequence .
CGAATAAAAGGCGATGGAGTTCGCCATAACCGCCTGAAAAGGCTGATGACTCCTACCGGT
GA
>MRCE01000059.1/24343-24431 Phormidium ambiguum IAM M-71 NIES- 2119_Scaffold_59, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AATAAAGATGGCGATGGAGCTCGCCAAAACCGCCTTTGATTATTGAATTAATTCATAAAA
TCTAAAAGGCTGATGGCTCCTACTATTCC
>MHDL01000053.1/7947-8008 Nitrospinae bacterium RIFCSPLOWO2_02_39_17 rifcsplowo2_02_subl0_scaffold_2327, whole genome shotgun sequence.
AAATAAACAGGTGATGGAGTTCACCTTTAACCGACTGAATTGCTGATAACTCCTTACATT AC
>FR893375.1/102776-102705 Firmicutes bacterium CAG:345 genomic scaffold, scf22
AATATCTATGGAAATGATGTCTTCCATCTAGTGATAGAGAACTGCAATATTGCTGATGAC
GTCTATAGTTTT
>CP007044.2 /4758383-4758446 Chania multitudinisentens RB-25, complete genome .
ACGACCAACGGAGATGACATTCCTCCATTAACCGCCCTTTCAGGCTGATGATGTCTACGT
AACC
>CYGX02000040.1/25150-25210 Burkholderia sp . STM 7296 strain STM7296 genome assembly, contig: BN2475_Contig_040
TGCGGCGCTGGAGATGGCATTCTCCATTAACCGCCTCGAAGCTGATGATGCCTGCTTCGC
C
>KI530723.1/120678-120613 Acinetobacter nectaris CIP 110549 genomic scaffold adgTG-supercont2.2 , whole genome shotgun sequence.
TGAGATATAGGAAATGGCATCTTCCTTTTAAAAACCGCCAATATGGCTAATGATGCCTAC GTAACC
>BANC01000005.1/58240-58304 Acidocella aminolytica 101 = DSM 11237 DNA, contig: Aam_005.
CAGGCAGACGGGGATGGAGTCCCCCAAAACCGCTCCTTCGGGGAGCTGATGACTCCTGCA
GGCGT
>CBXI010000023.1/134848-134912 Clostridium tyrobutyricum DIVETGP , WGS project CBXIOIOOOOOO data, contig: NODE_74
TGTATCTTAGGTGATGAAGTTCACCTTTAACCATCTCTTCCGGGATTGATGACTTCTACT
ATACA
>CP007453.1/365503-365563 Peptoclostridium acidaminophilum DSM 3953 plasmid EAL2_808p, complete sequence.
ATGAGAATAGGTGATGGGGTTCGCCTTTAAACGCATTTATGCTGATGACTCCTGCAATGA
A
>JANZ01000004.1/80066-80146 Mycobacterium kansasii 732 gmk732. contig .3 , whole genome shotgun sequence .
GGGTGTCGCGGCGATGGATGTCCGCCGGGATCATGTTGACTATCCGAACCGCCCACACGG
GATGATGACTCCTACCGAATG
>JAOC01000010.1/312580-312520 Mycobacterium xenopi 3993 GMX3993. contig .9, whole genome shotgun sequence .
GTTGACGTTGGCGATGAAGCTCGCCTTGATTGCCGCACCGGCTGATGGCTTCTACCGCGT
G
>MBSV01000068.1/14114-14041 Clostridium sp . W14A NODE_46, whole genome shotgun sequence.
TGTCATTAAAGGGATGAAGCTCCCCTTTGGAAATATCCGAATCGCTGGCAACGGCTGATG
GCTTCTACAGGAGT
>MGUV01000020.1/11423-11346 Elusimicrobia bacterium
RIFCSPLOWO2_02_FULL_61_ll rifcsplowo2_02_scaffold_l 6467 , whole genome shotgun sequence.
AATTAGAGCGGCGATGATGTCCGCCGCCCCGGAAAGCCGGGGGAACCGTCCGCGTGGACT
GATGACCTCTACCCTGAT
>MKSU01000008.1/3888-3950 Sphingomonas sp. 67-36
SCNpilot_cont_300_bf_scaffold_208, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GCCGGCGATGGGGATGGAGTACCCCCGATAACCGCCGTACCGGCTGATGACTCCTGTCAT
GCG
>AP011532.1/2430327-2430249 Methanocella paludicola SANAE DNA, complete genome .
TGCTATTCTGGTGATGAAGCTCACCGTGCTCTTATGCAGAACCGCCGCGCAGCGGGCGAC
TAATGGCTTCTGATAGTTC
>MKRJ01000017.1/31079-31140 Rhizobiales bacterium 65-79
SCNpilot_cont_1000_bf_scaffold_222 , whole genome shotgun sequence.
GTCGTCCCCGGGGATGGAGTCCCCCTCCAACCGCCTCGCCGGCTGATGACTCCTGTCGCA AC
>MKST01000003.1/260765-260834 Solirubrobacterales bacterium 67-14
SCNpilot_bf_inoc_scaffold_165, whole genome shotgun sequence.
ATAGCCCCGAGCGATGGGGCTCGCTCCTTCATAACCGCGGTTTCAATCCGCTGATGGCTC CTGTCAACCG
>FP929050.1/1488806-1488733 Roseburia intestinalis XB6B4 draft genome.
GTATGAAACGGGAATGAGGTTCTCCCACGATTAATTCGAAACCGCTAAATTTAGCTGATG
ACTTCTGTGTAGAC
>CP019948.1/4130853-4130792 Methylocystis bryophila strain S285, complete genome .
CCGCGTCCAGGGGATGGGGTGCCCCCTAAACCGCCGCAAAGGCTGATGACTCCTGCTTGT
TC
>CP001510.1/1276466-1276388 Methylobacterium extorquens AMI, complete genome .
GCCTTCGACGGTAATGGTGTCTGCCGGGCTTTGAGCCGAACCGCCGATCCTTGATGCGGC
TGATGACTCCTACGATTCG
>MKUN01000004.1/3411-3347 Bacteroidetes bacterium 46-16
SCNpilot_cont_1000_p_scaffold_1313, whole genome shotgun sequence.
GCGGTAAAAGGGAATGGAGTTTCCCTTGCCCAACCGCCTTAATGGCTGATGACTCCTGTT CAATC
>BAFH01000003.1/99123-99062 Candidatus Jettenia caeni DNA, contig: KSU1_C.
GAATTCTAAGGCGATGGAGTTCGCCAAACTGCTACACGGTAGCTGATAACTCCTATTGAG
GT
>AE008691.1/1847557-1847497 Thermoanaerobacter tengcongensis MB4, complete genome .
AATATTTAAGGTGATGAAGCCCACCTTAATTGCCGTAAAGGCTGATGGCTTCTACGAAGA
T
>MDSU01000018.1/304632-304694 Desulfurella amilsii strain TR1
opera_scaffold_l, whole genome shotgun sequence.
GTAAAGAAAGGCGATGGAGTTCGCCTTAAAATGCCCACCTTGGCTGATGACTCCTGCATA AAA
>FZOC01000002.1/459747-459808 Desulfovibrio mexicanus strain DSM 13116 genome assembly, contig: Ga0070557_102
CAGGGCGCAGGGGATGGAGTCCCCCTTGAACCGCGAATGCCGCTGATGACTCCTGCTGTA
CG
>MGOY01000004.1/4788-4860 Clostridiales bacterium GWF2_38_85
gwf2_scaffold_1205 , whole genome shotgun sequence.
TAAGTATAAGGGAATGAAGTTCTCTCTTACACATTGTAAAACCGCTTATTAAGCTGATGA CTTCTGCAACACC
>MNFC01000140.1/3171-3106 Candidatus Rokubacteria bacterium
13_1_40CM_4_69_5 13_l_40cm_4_scaffold_1285, whole genome shotgun sequence.
AACTGAAAAGGTGATGGGGTTCGCCTAAACCGCCCAGACCATCGGGCTGATGACCCCTGC
CGGACC
>GL987988.1/ 64193- 64251 Fusobacterium mortiferum ATCC 9817 genomic APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
scaffold supercont2.2 , whole genome shotgun sequence.
TAAATAATAGGCGATGGAGTTCGCCTTTAAGTGCGAAAGCTAATGACTCCTACTCTTTA
>LDXJ01000062.1/3052-3110 Peptococcaceae bacterium CEB3 CEB3_contig000062 , whole genome shotgun sequence .
AAAAACCCAGGCGATGGGGCTCGCTTATTTGCCTGAAGGCTGATGGCTCCTACCTTCAA
>JFHN01000018.1/140995-140923 Erwinia mallotivora strain BT-MARDI contig_18, whole genome shotgun sequence.
GTTGCGCAAGGTGATGGTGTTCCACCTTTCCCAACCGCCCCGCCTGGCCGGGGCTGATGA
CGCCTGATATCGC
>MGSQ01000223.1/5820-5882 Deltaproteobacteria bacterium
RIFCSPL0W02_12_FULL_57_22 rifcsplowo2_12_scaffold_44076, whole genome shotgun sequence.
CCAATTCGCGGAGATGGCGTTCTCCCTTAACCGCCTACCAGGGCTGATGACGCCTGCTTC
ATG
>FUZA01000014.1/23719-23785 Dyadobacter psychrophilus strain DSM 22270 genome assembly, contig: LX75DRAFT_scaffold00014.14
ACAATTACAGGAAATGGTGTCTTCCTGCGGAAACCGCCTTAACACGGCTGATGGCGCCTA TTAGGAT
>GL890571.1/895054-895144 Lachnospiraceae bacterium 9_1_43BFAA genomic scaffold supercont 1.1 , whole genome shotgun sequence.
CTTTAGCAAGGGAATGAAGTTCTCCCTTAGTAATACTAAAACCGCCAGGTGGAGCGTGCC TCCCCTGATTGCTGATGACTTCTGCTTTTTG
>CP015243.1/731191-731095 Halotalea alkalilenta strain IHB B 13600, complete genome.
CTGGGAACTGGAGATGGCATGCCTCCAGCCTTCGGCCATTCACCCAGCGAATCACGAAGC
GAACCGCCCTCACCGGGCCGATGATGCCTACGCACTG
>JH932292.1/1116847-1116767 Facklamia hominis CCUG 36813 genomic scaffold supercontl .1, whole genome shotgun sequence.
CTTTATTGAGGAAATGAAGTGCTTCCTGGCTTACAAAGTGTGAGTCGAACCGTTCTAAGA
ACTAATGACTTCTACAAGCTT
>CP002131.1/442379-442319 Thermosediminibacter oceani DSM 16646, complete genome .
AATCAAATAGGCGATGGAGTTCGCCCTAACCGTCCGTGCGACTAATGACTCCTACCAGCG
A
>CP022282.1/4852083-4852016 Chryseobacterium sp . T16E-39 chromosome, complete genome.
AACTGAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTACATAAGCTGATGACGCCT
GATTAATA
>MHYP01000280.1/7052-7112 Planctomycetes bacterium
RIFCSPHIGHO2_02_FULL_38_41 rifcsphigho2_02_scaffold_83486, whole genome shotgun sequence.
ATTTTAATTGGTGATGGAGTTCGCCTTTAATTGCTAAATAGCTGATAACTCCTATTGAAG
C
>CP003178.1/1777183-1777116 Niastella koreensis GR20-10, complete genome.
TACCCGCAAGGAAATGGTGTCTTCCTACTTAACCGTTCCAACCAGGAACTGATGGCGCCT
ACAAATAC
>LT629740.1/2938316-2938254 Mucilaginibacter mallensis strain MP1X4 genome assembly, chromosome: I
TTAGCAGCAGGAAATGGTGTCTTCCTTTTAAAAACCGCGATTGCTGATGGCGCCTGCAAA
TTC
>CP012184.1/4360830-4360893 Pseudonocardia sp. EC080619-01, complete genome .
CCCCTCGACGGCGATGGGGCTCCGCCGACAACCGCCCGCCCGGGCTGATGGCTCCTACCC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GACC
>MNFC01000140.1/4478-4410 Candidatus Rokubacteria bacterium
13_1_40CM_4_69_5 13_l_40cm_4_scaffold_1285, whole genome shotgun sequence.
GTAGACGACGGCGATGGGGTTCGCCCACAACCGCCTTGGAGTCGCGAGGCTGATGACCCC
TACCAGGCA
>MKRJ01000017.1/44346-44406 Rhizobiales bacterium 65-79
SCNpilot_cont_1000_bf_scaffold_222 , whole genome shotgun sequence.
ATTTCACATGGCGATGGAGTTCGCCTTCGAGCGCCGATCGGCTGATGACTCCTTGACTGT
C
>JGZE01000017.1/23163-23092 Bifidobacterium mongoliense DSM 21395
Contigl7, whole genome shotgun sequence.
GCAGCCAGTGGTGATGGATCCCGCCCGGACTTCGTGTCCAAACCGCATAACGCTGATGGT
TCCTACTCAATC
>MKTE01000006.1/11315-11253 Bacteroidia bacterium 44-10
SCNpilot_bf_inoc_scaffold_1083, whole genome shotgun sequence.
CACTGTTTCGGCAATGGGATCTGCCTTTAACCGTCCGCCAGGACTGATGATACCTACTTT TTC
>MERJ01000097.1/5928-6003 Betaproteobacteria bacterium
RIFCSPL0W02_12_FULL_63_13 rifcsplowo2_12_scaffold_l 9614 , whole genome shotgun sequence.
TGCCCCGAAGGAGATGGCATTCCTCCTTCAATAAGCAACAAGAACCGCCGCTACGGCTGA
TGATGCCTACAGATTT
>MEQK01000145.1/1118-1180 Bdellovibrionales bacterium RIF0XYD1_FULL_55_31 rifoxydl_full_scaffold_5117, whole genome shotgun sequence.
GGAACCTACGGCAATGGAGTCTGCCTTTTAACCGCCCCTTCGGCTGATGACTCCTACTCA AAC
>LMMS01000008.1/43799-43728 Methylobacterium sp . Leaf93 contig_16, whole genome shotgun sequence.
GAAGGCTGCGGCAATGGATTCTGCCGGGCGCAAAGCCGAACCGCCATAAGGGCTGATGAT
TCCTAACCTTTT
>GL379772.1/308520-308587 Sphingobacterium spiritivorum ATCC 33861 genomic scaffold SCAFFOLD3, whole genome shotgun sequence.
AATTAAAAAGGAAATGGTGTTCTTCCTTAACCAAACCGCTTTAGCAAGCTGATGACGCCT GATTGATA
>CP003598.1/6135-6014 Chroococcidiopsis thermalis PCC 7203 plasmid pCHRO.Ol, complete sequence.
ATTCTATATGGCAATGGAGCTTGCCAAAATCGCCGCTTCAGTTCGTGCTAAATGTTTTAC
CAGTTATGGCTGTAAAACAAAAGCTTGTAACTATCTGAAAGGCTGATGGCTCCTACTTTC
CA
>CP000478.1/3591388-3591322 Syntrophobacter fumaroxidans MPOB, complete genome .
TTGAATTGCGGCGATGAAGTTCGCCCTGAACCGCCCCGGGCCAGGGGATGATGACTTCTA
CTTCGAA
>AKZN01000023.1/27052-26986 Alicyclobacillus hesperidum URH17-3-68
URH17368ctg23, whole genome shotgun sequence.
CAATTTGTTGGTGATGGAGCTCACCAGAACCGTCCACCTGACGTGGACTGATAGCCCCTA
CCCGTGT
>FQUF01000034.1/23465-23404 Atopostipes suicloacalis DSM 15692 genome assembly, contig: EJ15DRAFT_scaffold00033.33
TTAATCAAAAGGGATGGTACTCCCTATAACCGCCAGAAATGGCTGATGGTGCCTGTTATG
TA
>AJWE01000036.1/40479-40540 Rhizobium sp . CF142 PMIll_contig_37.37, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CAGCCCGATGGGGATGGAGTGCCCCCGAAACCGCCCTCAAGGCTGATGACTCCTACCGTG
TT
>CP016895.1/3079667-3079734 Acinetobacter larvae strain BRTC-1 chromosome, complete genome.
AATTCTTCAGGAGATGGTCATCCTCCTCTAACAAACCGCCACGATCGGCTAATGATGCCT
ACGTTTCC
>MNXB01000100.1/23160-23091 Elusimicrobia bacterium CG1_02_37_114 cg_0.2_subl0_scaffold_372_c, whole genome shotgun sequence.
AAATTAAAAGGTGATGAGGTTCGCCTACGAACCACCCCGATAAATCGGGGTTGATGACCT CTACTCTGAA
>FWXF01000014.1/27835-27903 Desulfacinum hydrothermale DSM 13146 genome assembly, contig: EJ40DRAFT_scaffold00014.14
ATTGGGAAAGGCAAAGAAGTCTGCCTGAACCACCTGCCCAACGTGCAGGTTGATGACTTC
TACCCCTGC
>FR883363.1/841-769 Clostridium sp. CAG:221 genomic scaffold, scf26
ATTAACGTAGGGAATGAAGTTCTCCCTTGATTAATATCAAAACTGCGTAATAGCTGATGA
CTTCTACAACTTT
>FQ859181.1/1438924-1438986 Hyphomicrobium sp. MCI chromosome, complete genome .
GTAAGCGATGGGGATGGGGTTCCCCCGATAACCGCCATTGTGGCTGATGACTCCTACGAT
AGC
>FUZZ01000004.1/97869-97940 Chitinophaga ginsengisegetis strain DSM 18108 genome assembly, contig: LX61DRAFT_scaffold00004.4
CAGATAACAGGAAATGGTGTCTTCCTGACCCAAACCGTCAGCCCAGGGCTGTCTGATGGC GCCTACAAATAT
>CP004371.1/4957776-4957845 Flammeovirgaceae bacterium 311, complete genome .
GGGCGTAAAGGCAATGGTGTCTGCCTCAAACCGCCTGTAGTATTGTCAGGATGATGGCGC
CTGCTCATCG
>LQQU01000013.1/156749-156688 Crenobacter luteus strain CN10 contig_20, whole genome shotgun sequence .
CGCAGCGTGGGAGATGGCATGCCTCCCTAACCGCCTTTACGGCTGATGATGCCTACACGC
GA
>CP016211.1/3076023-3076093 Minicystis rosea strain DSM 24000, complete genome .
AGTCGGGATGGCGGTGGAGTCCGCCGACAACCGTCGTCGCTTCCGCGACGGCCAATGACT
CCTACGAGCAG
>FTOP01000001.1/138850-138783 Belliella pelovolcani strain DSM 46698 genome assembly, contig: Ga0111625_101
TTTTTCATAGGCGATGGGGTTCCGCCTTACAACCGCCCATTTTTGGTGCTGATGACTCCT
ACTTCAAC
>CP006568.1/1167260-1167196 Candidatus Sodalis pierantonius str. SOPE, complete genome.
GACCCTTCAGGAGATGACATTCCTCCTTAATAACCGCCGTTCTGGCTGATGATGTCTACG
TTCGC
>LGCL01000023.1/36324-36265 Ornatilinea apprima strain P3M-1 contig_9, whole genome shotgun sequence .
CAATAATATGGCGATGAGGCTCGCTAAAACTGTCACCCGACTGATGGCCTCTACTGAGCG
>ABCH01000002.1/2875-2939 Vibrio shilonii AK1 1103207002041, whole genome shotgun sequence.
TGCAACCAAGGTGATGGGGTTCCACCTACTTAACCGCCAAATCGGCTGATGACTCCTACA
GTTTC
>MKTW01000012.1/129407-129339 Sphingobacteriales bacterium 44-61 APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
SCNpilot_expt_1000_bf_scaffold_47 , whole genome shotgun sequence.
TGCCTGCAAGGAAATGGTGTCTTCCTACTGAACCGTTCTTTCTTAAGAACTAATGGCGCC
TACAAATTA
>JUEI01000047.1/36496-36435 Paenibacillus sp . IHB B 3415 contig_047, whole genome shotgun sequence.
ATTCAATACGGCGATGGAGTTCGCCATAACCGTCCTTCGGGACTAATGACTCCTACCAGT
GA
>CP002772.1/1977338-1977278 Methanobacterium sp . SWAN-1, complete genome.
ATTTTTACAGGCGATGGAGTTCACCTTTAACCGCGTCAACGCTGATGACTCCTGCTATCT
A
>CP001674.1/8629-8565 Methylovorus glucosetrophus SIP3-4, complete genome.
ACAGGCGGTGGAGATGGCATTCCTCCCTTAACCGCTGCCTTGCAGCTAATGATGCCTACG
ATAAC
>JPQW01000201.1/1074-1148 Cellulosimicrobium sp . MM contig201, whole genome shotgun sequence.
GCTGGGAGCGGTGATGGATCCCGCCGGGGCGCAGCCGCCCGAACCGCCGCCAGGGCTGAT
GGTTCCTGACTCGTC
>CP002480.1/2557621-2557691 Granulicella tundricola MP5ACTX9, complete genome .
CACAATCCTGGCGATGGGGTTCGCCTACGATGCCCGGCTGTAAAGCCGGCGCTGATGACT
CCTACATCTCG
>MNIL01000065.1/10116-10040 Verrucomicrobia bacterium 13_1_20CM_4_54_11 13_l_20cm_4_scaffold_189, whole genome shotgun sequence.
TCCGCTCAAGGCAATGGAGTTTGCCGTCCCACCTGAAGGCTGGGAAAACCGCGCGAGCTA ATGACTCCTTCCAAACG
>ADKM02000081.1/126-53 Ruminococcus albus 8 contig00134, whole genome shotgun sequence.
AGTTCATACGGGAATGATGTTCTCCCACGGGAAACCGAAACCGCTTATCATAAGCTGATG
ACTTCTGCGTTTTA
>MGQI01000065.1/2868-2939 Deltaproteobacteria bacterium RBG_13_58_19 RBG_13_scaffold_17853 , whole genome shotgun sequence.
GCAGACGAAGGCGATGAAGTCCGCCGCAAACGCCCCGGGATAGGTTCCAGGGCTGATGAC TTCTGTCCAACC
>HF570958.1/3235724-3235802 Tetrasphaera japonica T1-X7 genomic scaffold, 1540_scaffoldl
GTCGGACTCGGTGATGGATCCCACCGGGGCGTGAGGCTCACGCCCGAACCGCCTCCCGGA
TGATGGTTCCTGACTCGTC
>MICN01000022.1/2896-2964 Syntrophus sp . GWC2_56_31 gwc2_scaffold_15257 , whole genome shotgun sequence .
TTCATTGATGGGGATGGAGTCCCCCGCTGTAAAACGCTTCTTAACAGAGCTGATGACTCC
TGCTGCACG
>MNYV01000167.1/9181-9114 Nitrospirae bacterium CG2_30_53_67
cg2_3.0_scaffold_6689_c, whole genome shotgun sequence.
AAAAAAGTCGGCGATGGAGTTCGCCGTTGACTGCTTCCCCAAACGAAGCTGATGACTCCT ACTCTTTA
>CP007444.1/824586-824505 Dyella j iangningensis strain SBZ 3-12, complete genome .
GCCCGCCAGGGAGATGGCATGCCTCCCGCTCCGACTTGCGTACGGGGCGAACCACCTTCG
GGTTGATGATGCCTGCACCACC
>FQXE01000023.1/8132-8065 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_123
ACTCTTTACGGAGATGGCATTCCTCCTTTAACCATCGGCGCAAGCCGATTGATGATGCCT
ACAGGTCC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>FWXD01000003.1/194430-194367 Andreprevotia lacus DSM 23236 genome assembly, contig: BR64DRAFT_scaffold00003.3
GCAGGGTACGGAGATGGCATTCCTCCCCAAACCGCCCTGACCGGCTGATGATGCCTGCAC
ACGT
>LAKY01000073.1/82222-82160 Clostridiales bacterium PH28_bin88 ph28_926, whole genome shotgun sequence .
AACGAAATAGGCGATGGAGCTCGCCTTAACCACTGCCAAGCAGTTGATGGCTCCTACCGG
GGA
>LJYG01000104.1/286389-286467 Bradyrhizobium manausense strain BR3351 contig2104, whole genome shotgun sequence.
CCGACAGATGGGGATGGAGTCCCCCGACAACCGCCCGGCGGTCGCGCGACTGCAATGGGC
TGATGACTCCTGCCGAAGG
>CYSP01000014.1/96818-96882 Propionispora sp . 2/2-37 isolate 2/2-37 genome assembly, contig: 2/2_contigl4
CGACAATCGGGCGATGGAGTTCGCCGTAACCCCGCAGTGAAAGCGGTAATAACTCCTACC
TTCCC
>BAVZ01000004.1/149938-149878 Paenibacillus pini JCM 16418 DNA, contig: contig00004.
ATAAAGAAAGGCGATGGAGTTCGCCATAACTGCCTTCGTGGCTAATGACTCCTACCGGGG
A
>CP019875.1/3028593-3028672 Komagataeibacter nataicola strain RZS01, complete genome.
AACGAACACGGCAATGGATTCTGCCGGACCATGCTGGTCGAACCGCTGCCCGCAAGGGCG
CTGATGATTCCTGCCCACTG
>JGZI01000010.1/395195-395266 Bifidobacterium psychraerophilum strain LMG 21775 ContiglO, whole genome shotgun sequence.
ATCGGGTCTGGCGATGGAACCCGCCGGGAGCTCTGCTCCGAACCGCATACCGCTGATAGT TCCTACCCGTCC
>LN890280.1/1321626-1321549 Candidatus Nitrosotalea devanaterra genome assembly, chromosome: 1
CTCAAAGAAGGCTATGACATTAGCCTTGGAAGATTCCTAAACCGCTCCTAGAAAGGTGCT
GATAATGTCTACCTAATG
>HF987625.1/77121-77193 Eubacterium sp. CAG:86 genomic scaffold, scf291
TTACAAGTTGGGAATGAAGTTCTCCCATGGGAAACCTAAACTGCTTATTATAGCTGATGA
CTTCTACGATTTT
>LQQU01000013.1/158638-158573 Crenobacter luteus strain CN10 contig_20, whole genome shotgun sequence .
GCTGCTGCAGGAGATGGCATTCCTCCTCATAACCGCCCCGATCGGGCTGATGATGCCTAT
GGTATT
>CP002581.1/2902006-2901938 Burkholderia glumae PG1 chromosome 2, complete sequence .
CCGCGCTTCGGAGATGGCATGCCTCCGCCCCCAACCGCCGGTCAGCCGGCTGATGATGCC
TACGGATCC
>CP000812.1/982299-982238 Thermotoga lettingae TMO, complete genome.
AATAAGCCTGGTGATGGAGTCCACCGTTATTTGCCTTAAAGGCTGATGACTCCTACTTGA
AT
>CP005587.1/218662-218724 Hyphomicrobium denitrificans 1NES1, complete genome .
GGCGGCGATGGGGATGGGGTTCCCCCGATAACCGCCCATATGGCTGATGACTCCTAGCAA
AGC
>JH815222.1/1329237-1329163 Clostridium sp . 7_2_43FAA genomic scaffold supercont2.3 , whole genome shotgun sequence.
AGACTTTAAGGGAATGAAGTTCTCCCTTGACTTATGTCTAAACCGCTTTGGTAAGCTAAT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GACTTCTGCAAGTTT
>MHAY01000034.1/50235-50296 Ignavibacteria bacterium RIFOXYC2_FULL_35_21 rifoxyc2_full_scaffold_216, whole genome shotgun sequence.
TGTAACATTGGCGATGGAGTTCGCCATTGAACGCTGATAAAGCTGATGACTCCTGTTAAA AG
>AP009256.1/922963-923040 Bifidobacterium adolescentis ATCC 15703 DNA, complete genome.
CTGGACTGCGGTGATGGGACTCGCCTGAAGCCGTTACAGGCTTCGAACCGCAAACCCGCT
GATGGTTCCTACGACATA
>LAD001000039.1/59871-59815 Peptococcaceae bacterium BRH_c4b BRHa_1004828, whole genome shotgun sequence .
AATACAAGCGGCGATGGAGCTCGCCTAATGCGTTGCGCTGATGGCTTCTGCCGGGCT
>CP000481.1/805841-805902 Acidothermus cellulolyticus 11B, complete genome .
ACTGAACGCGGCGATGAAGCCCGCCGACAATGCGCCGCCCGGCTGATGGCTTCTACCGAG
TT
>CP003364.1/4855682-4855762 Singulisphaera acidiphila DSM 18658, complete genome .
CTGAATAATGGCGATGGGGTTCGCCAATCAATAACCGCCGCAGCCTCGTGGAAGGTTGCG
GCTGATGACCCCTACCTTTGG
>LQYX01000074.1/163851-163913 Geobacillus sp . B4113_201601 NODE_219, whole genome shotgun sequence.
AACTTCATAGGCGATGGAGTTCGCCATAACCGCCGGTTTCCGGCTGATGACTCCTGCTGC
GAA
>AEYM01000898.1/830-776 Lactobacillus rhamnosus MTCC 5462 contigl806, whole genome shotgun sequence .
GAATTAAATGGCGATGGTGTTCGCCTATACGTAAGTTGATGACACCTACCTTGTA
>JGZE01000020.1/9653-9563 Bifidobacterium mongoliense DSM 21395 Contig20, whole genome shotgun sequence .
GGTGAGCATGGGTATGAAGTTACCCAAGATAGTTTGTCGCGGTCGTGATGAGCGTCTGAA
CCGCTGTTAGGCTGATGACTTCTGCAGAGAT
>MDET01000013.1/67506-67568 Pseudaminobacter manganicus strain JH-7 JH- 7_scaffold20, whole genome shotgun sequence.
GACAATAATGGGGATGGGGTTCCCCCGATAACCGCCGCGAAGGCTGATGACTCCTACCGG
GCG
>HF990711.1/3779-3703 Clostridium sp . CAG:7 genomic scaffold, scfl72
ATAAGCGCCGGGAATGAAGTTCTCCCAAGGGAAACCTGAACCGCTTAGAAATGTAAGCTG
ATGACTTCTGTGATGAA
>MIHC01000025.1/15383-15322 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
ATGGGAGAAGGCGGTGGGGCTCGCCGAAACCGCAGACTAGTGCCAATAGCCCCTGCGGTC TG
>KE150405.1/570283-570197 Lachnospiraceae bacterium 3_1_57FAA_CT1 genomic scaffold acsNU-supercontl .1, whole genome shotgun sequence.
TCGAAGATTGGAAATGAAGTTCTTCCATGGTACCGGATACCAACGGATTACCTAAACCGC TTGACAGCTGATGACTTCTGCGACAAA
>AE008922.1/2322664-2322589 Xanthomonas campestris pv. campestris str.
ATCC 33913, complete genome.
GGCCATCTCGGAGATGGCGTTCCTCCGCGCTGTACCATCAACCGTAGCGATCGCTACTGA
TGACGCCTACAAGAAC
>MVHR01000005.1/164269-164343 Mycobacterium heidelbergense strain DSM 44471 NODE_5_length_243907_cov_56.1658, whole genome shotgun sequence. ACTGGCATCGGCGATGGCGCTCGCCGGAAAGCTCTGCTTTCGAACCGCCACACGGCTGAT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GGCACCTGCGAACGC
>CP000679.1/751276-751339 Caldicellulosiruptor saccharolyticus DSM 8903, complete genome.
TGCTGCAACGGCGAGGGAGTCCGCCGAACAAATGCCAATGATGGCTGATGACTCCTACAA
ATAT
>CP000319.1/1049491-1049431 Nitrobacter hamburgensis X14, complete genome.
CTGGAAACAGGGAATGGTGTCTCCCTTTCACCGCCGACTGGCTGATGACTCCTGCTGACT
G
>HG326223.1/1779972-1779897 Serratia marcescens subsp. marcescens Dbll, complete genome
CGTTCACAAGGTGATGGTGTTCCACCTTTCCCAACCGCCGCGTTCGCTCGAACCGGATGA
TGACGCCTGATATACA
>CP004856.1/2179184-2179256 Enterococcus casseliflavus EC20, complete genome .
ATGCTAATAGGGAATGAATGTCTCCCTCAGTTCTTACTGAAACCGCAATTTTGCTAATGA
CTTCTACCACTTT
>CU468135.1/2891609-2891537 Erwinia tasmaniensis strain ET1/99 complete chromosome .
GTATCACAAGGTGATGGTGCTCCACCTTTCCCAACCGCCCTGATTTTTTCGGGCTGATGA
CGCCTGATAAAAC
>JQAZ01000005.1/41654-41591 Lactobacillus selangorensis strain DSM 13344 Scaffold5, whole genome shotgun sequence.
TAGAAAATAGGGCATGGTGTTCGCCCGATGAAACTGCATGAATGCTGATGACACCTATCC
GGAG
>MNEM01000075.1/10149-10212 Gemmatimonadetes bacterium 13_1_40CM_66_11 13_l_40cm_scaffold_3648, whole genome shotgun sequence.
TTGCTCGCAGGCGATGGAGTTCGCCTTCAACCGCCCTTTACGGACTGATGACTCCTACCT GAAC
>CP001147.1/1737215-1737284 Thermodesulfovibrio yellowstonii DSM 11347, complete genome.
TTTAAAAAAGGCGATGGAGTTCGCCTGTAAGTGCCTCTGATTTTCAGGGGCTGATAACTC
CTACCTTAAA
>MNYH01000071.1/560-627 Deltaproteobacteria bacterium CG2_30_66_27 cg2_3.0_scaffold_l 0978_c, whole genome shotgun sequence.
CGAAAAAACGGCGATGGAGTTCGCCGATAACCGCTGTCGTTCGGACGGCTGATGACTCCT GCCCGGTC
>CP009284.1/5388823-5388763 Paenibacillus sp . FSL R7-0331, complete genome .
GAAGATAAGGGCGATGGAGTTCGCCATTAACTGCCGCGAGGCTAATGACTCCTACCAGCG
A
>CP001034.1/285558-285483 Natranaerobius thermophilus JW/NM-WN-LF, complete genome.
GGATTTGACGGCGATGGAGCTCGCCAGGATTAAATTCCAAACTGCTTTCTTAGAAGCTAA
TGGCTCCTACCCTATC
>LGCK01000014.1/708466-708524 Leptolinea tardivitalis strain YMTK-2 contig_l, whole genome shotgun sequence.
ACTTAGTTTGGCGATGAGGCTCGCTCAACCGTCATGTGACTGATAGCTTCTACTATTTT
>MGPZ01000034.1/11145-11207 Deltaproteobacteria bacterium GWD2_55_8 gwd2_scaffold_1402 , whole genome shotgun sequence.
TCTATTCGCGGAGATGGCGTTCTCCTTTAACCGCCTGCCAGGGCTGATGACGCCTGAATT CAT
>GG669604.1/320287-320352 Lactobacillus brevis subsp. gravesensis ATCC 27305 genomic scaffold SCAFFOLD1, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AATTAATATGGCGATGACGTTCGCCTATACTCGTTTAATAAAAGAGTTGATGACGTCTAC
CTTTAT
>FRAF01000044.1/3382-3318 Alicyclobacillus sp. USBA-503 genome assembly, contig: Ga0105849_144
AAAATAAAAGGTGATGGAGCTCACCATAAACTGCCAACTGGTTGGCTAATGGCTCCTACG
GATAA
>FSGV01000012.1/99309-99237 Mycobacterium abscessus subsp. abscessus strain 57 genome assembly, contig: ERS244815SCcontig000012
TTCGTCAGGCGCGATGGATCTCGCCAGGGCTTGTCCCGAACCGCCACTGACGGCTGATAG TTCCTGTGTTGAT
>CP003360.1/869588-869519 Desulfomonile tiedjei DSM 6799, complete genome.
CTCGCACGAGGCAATGGAGTCTGCCTAATTAACCGCCGAGCCGGATTCGGCTGATGACTC
CTACTGTTCA
>MRUL01000002.1/549750-549816 Izhakiella sp . D4N98
NODE_2_length_554444_cov_106.828_ID_1989, whole genome shotgun sequence.
CCCACCACGGGTGATGGCGTTCCACCTGCCCCAACCGCCCTTCTGGGCTGATGACGCCTG
GCAAACT
>LVYD01000065.1/51211-51278 Niastella vici strain DJ57 contig36, whole genome shotgun sequence.
TTCAAGTAGGGTGATGGCGTTCCACCTAATTAACCGCCTTTCGGAAGGACGATGACGCCT
ACAGGTTG
>AE008692.2/542148-542071 Zymomonas mobilis subsp. mobilis ZM4, complete genome .
TATAGACAAGGTAATGGAATCTACCTGATCCTTTCTATGGATCGAACCGCCCTTTAGGCT
GATGACTCCTGCTCAAAC
>GL379781.1/3860594-3860660 Chryseobacterium gleum ATCC 35910 genomic scaffold SCAFFOLD1, whole genome shotgun sequence.
AAGAAAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTCTAAAGAGCTGATGACGCCTG ATTAAAT
>AZEM01000025.1/51232-51172 Lactobacillus kefiranofaciens subsp.
kefirgranum DSM 10550 = JCM 8572 strain DSM 10550 NODE_34, whole genome shotgun sequence.
AAATGAATAGGCGATGACGTTCGCCATTAAATGAGCAAAATCTAATGACGTCTACTATTT
A
>FR887307.1/81021-80949 Clostridium sp. CAG:253 genomic scaffold, scfl63
AACGAGTGAGGGAATGAAGTTCTCCCGAGGAAAACCTAAATTGCTTATAATAGCTGATGA
CTTCTACGATTAT
>BX950851.1/1468958-1468853 Erwinia carotovora subsp. atroseptica
SCRI1043, complete genome
CACTAAAGCGGAGATGACATTCCTCCCTACCCTAACGCCGCCAAAAATGCTGGCAGCCTG
AAGTACAAGGTACAACCGCCACCCGGCTGATGATGTCTACGTCCAC
>MUGK01000091.1/13066-12999 Chromatiales bacterium USCg_Taylor
scaffold_90, whole genome shotgun sequence.
AACGCTACAGGAGATGGCATGCCTCCTGACCAACAACCGCCGTTTCGGCTGATGATGCCT
ACGTGTTC
>MUNU01000019.1/461039-461139 Rhodanobacter sp . B05
NODE_l_length_645680_cov_ll .9856_ID_1, whole genome shotgun sequence.
GGCCGCGTGGGAGATGGCATGCCTCCCGTTCCGGCGCGACTGTCGCGGGACAAACCACCT
CCCGTGCGACCCACGTCGAGGTTGATGATGCCTGCATCGCC
>MNFC01000140.1/1400-1331 Candidatus Rokubacteria bacterium
13_1_40CM_4_69_5 13_l_40cm_4_scaffold_1285, whole genome shotgun sequence.
GCACGAAGCGGCGATGGGGTTCGCCTGAAACCGCCCCGCCGGATGCGAGGCTGATGACCC
CTACCGGAAG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>AQHR01000110.1/102878-102980 Lunatimonas lonarensis strain AK24
S14_contig_28, whole genome shotgun sequence.
TGCCCATCGGGTGATGGGGTGCCACCCAAATATCAGGGACTTATCCGAAGTTTTCTGATT
TTTTGAACCGCTTCGCAAGAATGCTGATGACTCCTACTTCAAC
>HF951689.1/2127795-2127711 Chthonomonas calidirosea T49 complete genome
TTGAACCTTGGCGATGGAGGCCGCCAGGGCGCTTGTTTTTGCTGCCCAAACCGCTCTCTT
TTGAGCTGATGGCTCCTAGCTCACA
>BAND01000009.1/6045-5970 Acidomonas methanolica NBRC 104435 DNA, contig: Amme_009.
GTCGGAATTGGCAATGGACTCTGCCTGACCTGCAAAGGTCGAACCGCCCGCGAGGGCTGA
TGATTCCTACCTCGCC
>LZPM01000042.1/12528-12634 Methanobacterium sp . A39 contig_5, whole genome shotgun sequence.
ATAATAATAGGTGATGGAGTTCACCTTTAACTGCTTATTTAAGCTGGATCAATAAATTGA
TCAAAATATAAACCGCTTATTTTTAAGCTGATGACTCCTACCAAGAT
>MKUQ01000060.1/73481-73542 Burkholderiales bacterium 70-64
SCNpilot_expt_1000_bf_scaffold_93 , whole genome shotgun sequence.
CTTCACATGGGAGATGGCATGCCTCCCCGAACCGCCTGACGGCTGATGATGCCTACCGAT
CG
>AYYX01000035.1/16618-16560 Lactobacillus vini DSM 20605 Scaffold35, whole genome shotgun sequence.
TTGATTTTAGGTGATGACGTTCACCGTTAAACCAATTTATTAATGACGTCTGTTTAAGG
>CP003379.1/1270075-1270145 Terriglobus roseus DSM 18391, complete genome.
CACCTGCGTGGCGTTGGGGTTCCGCCAGAACCGCACGCGGCCTGGCCGCTGCTGATGACT
CCTACGAATCG
>FXAZ01000009.1/141555-141494 Paenibacillus sp . 11 genome assembly, contig: Ga0139009_109
CATTATATAGGCGATGGAGTTCGCCGTATAACCGCCGCGAGGCTAATGACTCCTACCAGA
GA
>CELZ01000050.1/69316-69258 Moorella glycerini strain NMP genome assembly, contig : M_glycerini_NMP_DRAFT_scaffold-50
AACCCCAACGGCGATGGAGCTCGCCTAACGCTCAGTGAGCTAATAGCTCCTACCGAAAG
>CP011280.1/1051708-1051637 Sneathia amnii strain SN35, complete genome.
AATGAAATTGGAGATGAAGTTCTCCATGATTAAATATCAAAACTGCTTTATGCTAATGAC
TTCTACACTAAG
>HF985409.1/884-792 Clostridium sp. CAG:1024 genomic scaffold, scfl48
AAAAAAGATGGGAATGAGGTTCTCCCGCGGCAGGACGCAGCTTGCGCCTTGCCGGAACCG
CCGGCACTTCCGACTGATGACTTCTCATGCGCC
>CP009282.1/1122616-1122555 Paenibacillus sp . FSL R5-0912, complete genome .
ATTCAATACGGCGATGGAGTTCGCCATAACCGTCCTTCGGGACTAATGACTCCTACCAGT
GA
>MIAV01000034.1/16563-16483 Spirochaetes bacterium GWC2_52_13
gwc2_scaffold_337 , whole genome shotgun sequence.
TGCGCTACAGGCGATGGGGTTCGCCTCTACCTGACAAGGTCGCAACCGTGGACACTGTCC ACTGATAACTCCTACCCATCA
>JAOC01000010.1/317039-316978 Mycobacterium xenopi 3993 GMX3993. contig .9, whole genome shotgun sequence .
ATCAGGCCAGGCGATGAAGCTCGCCTTGACCGCCACCAAAGGCTGATGGCTTCTACCACA
GA
>LVHF01000033.1/31222-31285 Photobacterium jeanii strain R-40508
scaffold9, whole genome shotgun sequence.
TTTGGCAAAGGTGATGGGGTTCCACCTACTTAACCGCCATGTGGCTGATGACTCCTACAG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AAAA
>CP002403.1/1537238-1537164 Ruminococcus albus 7, complete genome.
ATATCGTATGGGAATGATAGTCTCCCGAGGTTTTACCTGAACTGCTTATTCTGAGCTGAT
GACATCTGCAATACG
>CP000853.1/1242223-1242295 Alkaliphilus oremlandii OhILAs, complete genome .
ATTTCTATAGGGAATGAAGTTCTCCCTTGGCAATAGCCTAAACCGCAAATATGCTGATGA
CTTCTGCGAAAAT
>AL954747.1/1838968-1838902 Nitrosomonas europaea ATCC 19718, complete genome
ATAATAATTGGAGATGGCATTCCTCCATTAACCGCCGCAATCTGTGGCTGATGATGCCTA
CGACTGG
>CP002629.1/3040536-3040463 Desulfobacca acetoxidans DSM 11109, complete genome .
ATGTAAACTGGCGATGGGGTCCGCCTTCAACCGCCTGAAGCAAATGTTTTCGGGATGATG
ACTTCTACTGGGAT
>DS264286.1/287218-287292 Eubacterium ventriosum ATCC 27560 Scfld0224 genomic scaffold, whole genome shotgun sequence.
TAAATAAAAGGGAATGAGGTTCTCCCTCGATTTAATCGAAACCGCTTATAACAAGCTGAT GACTTCTGTGCAATA
>MGTA01000258.1/277-339 Deltaproteobacteria bacterium RIFOXYD12_FULL_50_9 rifoxyd3_full_scaffold_874 , whole genome shotgun sequence.
ATGCAATGCGGTGATGGAGTCCACCTTTAACCGCCGGAAATGGCTGATGACTCCTACCAG GAT
>JH601133.1/698518-698445 Facklamia languida CCUG 37842 genomic scaffold supercontl .1, whole genome shotgun sequence.
TGAAACCAAGGAAATGAAGTGCTTCCTTTCGTTGACGATAAACCGCCGTGTTGGCTGATG
ACTTCTGTAGGCTT
>FR896817.1/33673-33594 Ruminococcus obeum CAG:39 genomic scaffold, scfl84
AAATATAAAGAGAATGAGGTTCTCCCTAAGTAATAAAATTACTTAAACCGCTTATGAAAG
CTGATGACTTCTGCGAGTAA
>AWUE01002742.1/2094-2171 Corchorus olitorius cultivar 0-4 contig02744, whole genome shotgun sequence .
GAGCCTTGCGGAGATGGCATGCCTCCCTTAACCGCCGGTCTGCGGATGACGCGTCCGGCT
GATGATGCCTACAAGTTC
>HG764817.1/4224730-4224788 Clostridium ultunense Esp genomic scaffold, CULT_3033
ATAAATAATGGTGATGGGGTTCACCAAAATCGCTTAATGCTAATGACTCCTACAGGAAT
>FPBV01000005.1/42076-42140 Alicyclobacillus macrosporangiidus strain DSM
17980 genome assembly, contig: Ga0104483_105
ACGCGAATCGGCGATGAGGCTCGCCACAACCGTCTGCCAGACAGACTGATGACTCCTGCC
CGGCG
>JJNX02000028.1/24007-24068 Peptococcaceae bacterium SCADC1_2_3 contig_28, whole genome shotgun sequence .
TATGACATAGGGAATGAAGTCTCCCTCATAAATGCCAATTGGCTGATGACTTCTGCAGAA
AA
>JQBK01000090.1/3619-3552 Lactobacillus acidipiscis strain DSM 15353 Scaffold90, whole genome shotgun sequence.
ATGTTTTCTGGTGATGACGTTCACCAGTCCAAAAATCCGTGCTCCCACGTAATGACGTCT
AAGTGAAT
>FR886704.1/8190-8263 Blautia sp . CAG:237 genomic scaffold, scf90
ACCCGAATTGGGAATGAAGTTCTCCCATGGGAAACCTAAACTGCTTATTCAGAGCTGATG
ACTTCTACGATTAT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>FQXE01000008.1/8211-8275 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_108
TGTGTTACAGGAGATGGCATTCCTCCTTATAACCGCCCAAACCGGCTGATGATGCCTACC
AAACT
>MHYG01000119.1/1324-1262 Planctomycetes bacterium RBG_13_62_9
RBG_13_scaffold_24815 , whole genome shotgun sequence.
ACTGATCGTGGGGATGGAGTCCCCCGGACAACCGCCTGCAGGGCTGATGACTCCTACCGG TTT
>MEKZ01000322.1/5576-5514 Acidobacteria bacterium
RIFCSPLOWO2_12_FULL_60_22 rifcsplowo2_12_scaffold_52111, whole genome shotgun sequence.
NTACCCGGGGGTGATGGGGTTCACCCAAACCGTCCGCTAGGGACTGATGACCCCTGCCGA
CTG
>APHK02000086.1/1858-1751 Gordonia sp . NB4-1Y contig_86, whole genome shotgun sequence.
ATTGGTCGCGGCGATGGATCTCGCCGGAGGCCGGCCCCGATGACGGGACCCCGGCCCGAA
CCGCCTCCACCGGGGTGACCCGGACCGGCTGATAGTTCCTGTTCCGAT
>LOEW01000239.1/52785-52725 Desulfosporosinus sp . BRH_c37 BRHa_1005612 , whole genome shotgun sequence .
GTTATTTCAGGTGATGGGACTCACCTTATAAAGGCCAACAGCTAATGGCTCCTACTTTCA
A
>MGVH01000034.1/1081-1160 Elusimicrobia bacterium RIFOXYA2_FULL_58_8 rifoxya2_full_scaffold_1728, whole genome shotgun sequence.
CCGTACATCGGCGATGATGTCCGCCGCACCGCCGGAACGCGGTGAAACCGTCCGCGGGGA CTGATGACCTCTACCGAAGA
>LXKN01000075.1/60798-60710 Rhizobium sp . AC27/96 AS1I6_9, whole genome shotgun sequence.
AAATCCAACGGTAATGGATTTCTGCCGGGCCATGGTGGCCGAACCGCTGCTCTGAATTTT
CCAGGGAAGCTGATGACTCCTACTCAACA
>CP001802.1/3683321-3683245 Gordonia bronchialis DSM 43247, complete genome .
TGATCGTCGGGTGATGGATCTCCGCCCGGAGGGTTTCCTCCGAACCGCCGATGTCGGCTG
ATGGCTCCTGTTCCGAT
>AE013598.1/2682944-2683020 Xanthomonas oryzae pv. oryzae KACC10331, complete genome.
TCGGGATTCGGAGATGGCGTTCCTCCGCGCAGTTCCATCAAACCGTAGCGATTGCTACTG
ATGACGCCTACAAGAAC
>MNFG01000237.1/15519-15455 Gemmatimonadetes bacterium 13_1_40CM_4_69_8 13_l_40cm_4_scaffold_11475, whole genome shotgun sequence.
CTTTGGCCAGGCGATGGAGTTCGCCTCGAACCGCCGGTGCGCCGGCTGATAACTCCTACC CGAAC
>JMIH01000028.1/181850-181915 Anditalea andensis strain LY1 Contig_28, whole genome shotgun sequence .
ATCATTATAGGTGATGGGGTTCCACCTTTAACCGCGCAAACAGCTGCTGATGACTCCTAC
TTCAAC
>FSRA01000001.1/1529604-1529670 Chitinophaga niabensis strain DSM 24787 genome assembly, contig: Ga0070009_ll
CCTCAAACAGGAAATGGTGTCTTCCTGACCCAACCGCTTCTCCGAAGCTAATGGCGCCTA
CAAATTA
>FR899086.1/37911-37836 Ruminococcus sp . CAG:403 genomic scaffold, scf93
TTGATTATTGGGAATGAATGTCTCCCACGGAATCACTCCGAAACCGCTTTTCTAAGCTGA
TGACTTCTGTTTGCAT
>CP000822.1/1490183-1490109 Citrobacter koseri ATCC BAA-895, complete APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
genome .
AGTTCACAAGGTGATGGTGTTCCACCTTTCCCAACCGCCGGTTCGCTCGAATCGGATGAT
GACGCCTGATATACA
>KB946329.1/254744-254677 Enterococcus phoeniculicola ATCC BAA-412 genomic scaffold acvKl-supercontl .7 , whole genome shotgun sequence.
CGAATAATAGGTGATGGTGTTCACCTTTAACCATTTATCACAGATAAATTAATGACGCCT ACTAAAAG
>CSXB01000035.1/2984-3053 Mycobacterium abscessus strain PAP053 genome assembly, contig: ERS075544SCcontig000035
TAGCAACATGGTGATGGTGTTCACCACGAACTATTTATTGGACGAATAAATTAATGACGC
CTACCAAACG
>AXUN02000005.1/18307-18235 Youngiibacter fragilis 232.1 contig_102, whole genome shotgun sequence.
TTTCTACATGGGAATGAACGTCTCCCTTGGCTTTGCCAGAACCGCATTTTATGCTGATGA
CTTCTGCGATCTC
>CP014160.1/937362-937445 Aerococcus sanguinicola strain CCUG43001, complete genome.
CAAGACACAGGAAATGAATGTCTTCCTTGGGCCTAGCCCTAAACCGCTATTGATGATTCA
AAAGCTGATGACTTCTGCTGTAAA
>CP013213.1/467579-467651 Erysipelothrix larvae strain LV19 chromosome, complete genome.
GTACACAAAGGGAATGAAGTTCTCCCTGTTCTGATTAAGACAAACCGCATATGCTGATGA
CTTCTATACTTTT
>MGZC01000119.1/11088-11150 Gemmatimonadetes bacterium GWC2_71_10 gwc2_scaffold_28429, whole genome shotgun sequence.
TCAGTGATGGGTGATGGGGCTCACCCCGAACCGCCGGCGACGGCTGATGGCTCCTGCGCA ACA
>FR880482.1/2074-1999 Firmicutes bacterium CAG:56 genomic scaffold, scf405
ATGACTGCCGGGAATGAGGTTCTCCCTGGTTTTTAACCAAACCGCTTGTATATAAGCTGA
TGACTTCTGCGATTTC
>LIYD01000005.1/2797649-2797582 Flavobacterium akiainvivens strain IK-1 scaffoldOOOOl, whole genome shotgun sequence.
AAGACAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTTTAAGAGCTGATGACGCCT
GATTAAAC
>MIHC01000025.1/40346-40427 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
GTAACTATCGGCGATGGATGTCCGCCTGGGTGAACACTCGTCCGAACCGCCTCGTATCGA GGCTGATGGCTCCTGCTGCATT
>AP012204.1/145014-145086 Microlunatus phosphovorus NM-1 DNA, complete genome .
GGTAGCCGTGGCGATGGTGTCCGCCTGAGGATCGTCCTCGAACCGCCGGCCGGCTGATGA
CCCCTGCGACGAA
>MIAU01000034.1/7562-7469 Spirochaetes bacterium GWC1_61_12
gwcl_scaffold_l 929 , whole genome shotgun sequence.
TCGGCCCACGGCGATGGAGTTCGCCGAGCCTTCCCGATTCCGGGTGAACTAAACTGCCTA CCCTCCGGTACTGGCTGATGACTCCTGCCAACAT
>MKWE01000021.1/65340-65401 Rhodospirillales bacterium 70-18
SCNpilot_cont_750_bf_scaffold_241, whole genome shotgun sequence.
GATTGACGAGGGGATGGAGTCCCCCTTGAACCGCCGCAGAGGCTGATGACTCCTGTCGCG CA
>LWAE01000004.1/73755-73693 Clostridium magnum DSM 2767
CLMAG_contig000004 , whole genome shotgun sequence.
TAATAATTAGGTGATGAAGTTCGCCTTTAAATATCCTTTGGGATTGATGGCTTCTACTAT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ATG
>CP002987.1/3185660-3185718 Acetobacterium woodii DSM 1030, complete genome .
AATTAAGAAGGTGATGGAGTTCACCTGAATCACTTTTAGTTAATGACTCCTGTAGAATT
>AE014295.3/925740-925809 Bifidobacterium longum NCC2705, complete genome.
TGGAGAACCGGCGATGGAACCCGCCTGAACTCGTTTCGAACCGCACACTGCTGATAGTTC
CTACACACGA
>CP003926.1/1810121-1810037 Gluconobacter oxydans H24, complete genome.
ATTTAGGAAGGCAATGGACTCTGCCTGATGGCCTACTGGCGATCGAACCGCTTCCGCGAC
GGAGGCTGATGATTCCTACCCGACG
>FODD01000012.1/10355-10437 Streptomyces rubidus strain CGMCC 4.2026 genome assembly, contig: Ga0079901_1012
TCACGAACTGGCGATGGACCCCGCCAGGGACGGCGCGATGCCGCCCCGAACCGCCCCTCC
GGGCTGATGGCTCCTGACCGAGT
>FQXE01000008.1/8132-8065 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_108
ACTCTTTACGGAGATGGCATTCCTCCTTTAACCATCGGCGCAAGCCGATTGATGATGCCT
ACAGGTCC
>KB946287.1/569530-569463 Enterococcus villorum ATCC 700913 genomic scaffold acvIF-supercontl .8, whole genome shotgun sequence.
AAGTGAATAGGTGATGGTGTTCGCCTTTAAACGTTATTCTTTAGATAACTAATGACGCCT ACTAAAAA
>CP010026.1/62545-62607 Burkholderia fungorum strain ATCC BAA-463
chromosome 1, complete sequence.
TTGCGCGCTGGAGATGGCATTCTCCATTAACCGCCCTTGATGGCTGATGATGCCTGCTTC
GCC
>MNDE01000149.1/15745-15673 Acidobacteria bacterium 13_2_20CM_57_17
13_2_20cm_scaffold_1184, whole genome shotgun sequence.
ATTTCGGAAGGCGATGGAGTTCGCCATAACCGTCCCGGCACTGTTCACCGGGTCTGATGA CTCCTGGCAGACG
>MGNH01000132.1/8045-7963 Chloroflexi bacterium RBG_16_48_7
RBG_16_scaffold_64961 , whole genome shotgun sequence.
TTTATCTAGGGCGATGAAGCTCGCCGAGGAAGTTCGGAACTTTCCTGAACTGCTTTGAGA AAGCTAATAGCTTCTACCGGTAT
>AP012279.1/5492285-5492362 Bradyrhizobium sp. S23321 DNA, complete genome .
CGCACAGATGGGGATGGAGTCCCCCGATAACCGCCCGACGGCCCGTGACCGTACTGGGCT
GATGACTCCTGCTCGAGG
>FTNE01000002.1/116552-116490 Acidiphilium rubrum strain ATCC 35905 genome assembly, contig: Ga0104712_102
TCTGCCTGCGGCGATGGAGTTCGCCATGTAACCGCCCGCGGGGCTGATGACTCCTGTTCC
TTG
>HF996938.1/32238-32157 Clostridium hathewayi CAG:224 genomic scaffold, scf126
CTATACCCCGGGAATGAGGTCTCCCGTGGTAGTGAACATGTTACCAGAACCGCTTATTAC
AGCTGATGGCTTCTGCATTGTG
>FZOC01000008.1/25292-25353 Desulfovibrio mexicanus strain DSM 13116 genome assembly, contig: Ga0070557_108
ATGACCAGAGGGGATGGAGTTCCCCTTGAACCGCGACAGACGCTGATGACTCCTGCTGCC
CG
>BAND01000009.1/9337-9253 Acidomonas methanolica NBRC 104435 DNA, contig: Amme_009.
ATTCAGCAAGGCAATGGATTCTGCCTGATGGCCTGATGGCGATCGAACCGCTTCCGCGAC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GGAGGCTGATGATTCCTACCCGACG
>MERI01000159.1/4172-4246 Betaproteobacteria bacterium
RIFCSPL0W02_12_FULL_62_58 rifcsplowo2_12_scaffold_32846, whole genome shotgun sequence.
CCAGCACAAGGAGATGGCATTCCTCCTTTAAGCTCCAACAAAACCGCCTGTCAGGCTGAT
GATGCCTACGGTTTT
>CP013729.1/208392-208469 Roseateles depolymerans strain KCTC 42856, complete genome.
TCGCCGCACGGAGATGGCATTCCTCCGCGAACCGCCGAACGACGCCCTCGTCGTGACGCT
GATGATGCCTACATCGTC
>CP002582.1/2204933-2205007 Clostridium lentocellum DSM 5427, complete genome .
ATAAATAAGGGGAATGAAGTACTCCCTAAGTTTGATACTTAAACCGCTTATAAAGCTGAT
GACTTCTGTGAAACA
>CP000270.1/1757402-1757341 Burkholderia xenovorans LB400 chromosome 1, complete sequence.
TGCTGCGCTGGAGATGGCATTCTCCTTTAACCGCCCTTATGGCTGATGATGCCTGCTCGC
CC
>CP002206.1/1030223-1030154 Pantoea vagans C9-1, complete genome.
ATATGACAAGGTGATGGCGTTCCACCTTTCCCAACCGCTCCGTTCCGGAGCTGATGACGC
CTGATGCGAC
>MKZT01000024.1/330276-330356 Kushneria sp . YCWA18 contig24, whole genome shotgun sequence.
ATTGATAAGGGAGATGGCATGCCTCCCGAATGCGCCCGTATTCAAACCATCACGTCCGTG
ATTGATGATGCCTGCAGTACG
>CP015971.1/1426412-1426479 Arachidicoccus sp. BS20, complete genome.
TTAAAAAAAGGTAATGGCATTCTACCTTATCAACCGCCTTTAAGAAGGACGATGATGCCT
GTAAAAAT
>FPJE01000012.1/99004-98938 Sinomicrobium oceani strain CGMCC 1.12145 genome assembly, contig: IQ32DRAFT_scaffold00012.12
AAAATCAGTGGCAATGGTGTCTGCCCTGAACCGCCCTTTTACAAGGGATAATGGCGCCTG TTATCGT
>MGUX01000142.1/8815-8748 Elusimicrobia bacterium RIFCSPL0W02_12_FULL_59_9 rifcsplowo2_12_scaffold_6648 , whole genome shotgun sequence.
TCGAAAAAAGGCGATGGAGTTCGCCTAAACCGCCCCGGAAGCCCGGGGATGATGACTCCT GCCGACGC
>CP000822.1/1409469-1409541 Citrobacter koseri ATCC BAA-895, complete genome .
CCTGTGCAGGGTGATGGCGTTCCACCTTATTCCAACCGCCTCGTTTCAGTAGGCTGATGA
CGCCTGATATGAC
>MNJV01000118.1/2962-2894 Candidatus Rokubacteria bacterium
13_1_20CM_2_70_7 13_l_20cm_2_scaffold_3979, whole genome shotgun sequence.
GTAGACGACGGCGATGGGGTTCGCCCACAACCGCCTTGGAGTCGCGAGGCTGATGACCCC
TACCAGGCA
>JH114303.1/129496-129555 Desulfovibrio sp . 6_1_46AFAA genomic scaffold supercontl .4 , whole genome shotgun sequence.
CAGGGCAACGGGGATGGGGTCCCCCATAAAACCGCATTCGTTGATGACTCCTGCCAGCGC
>AE017334.2/4823859-4823918 Bacillus anthracis str. 'Ames Ancestor', complete genome.
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>JYHN01000069.1/66814-66741 Clostridium sp . FS41 CLFS41_contig000076, whole genome shotgun sequence .
CGTGAGTATGGGAATGAGGTTCTCCCTTAGTGAATACTAAAACCGCTTATCAAGCTGATG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ACTTCTGCGTAAAT
>GL870809.1/1469640-1469711 Clostridium sp . D5 genomic scaffold
supercontl .1, whole genome shotgun sequence.
AGATTTTAAGGGAATGAGGTTCTCCCGTGGAGTATTCCAAAAGCGCATTTTGCTGATGAC
TTCTGTGTAAAA
>CP000021.2/1258800-1258735 Vibrio fischeri ES114 chromosome II, complete sequence .
ATACGTAAAGGTGATGGGGTTCCACCTACTTAACCGCCAATGTTGGCTGATGACTCCTAC
AGAATA
>CAHT01000010.1/3917-3838 Methylacidiphilum fumariolicum SolV, WGS project CAHT01000000 data, contig: 90-2114
AAAGCAATTGGTAATGGAGTTTACCATCCAATCTTTTGGAAAACCCTTCTTCTTTAAGAA GTGATAACTCCTACGACAAC
>CP003583.1/1868350-1868284 Enterococcus faecium DO, complete genome.
AAGTGAATAGGTGATGGTGTTCGCCTTTAAATGCTAACGAAAGATAGCTAATGACGCCTA
CTAAAGT
>FRAD01000003.1/165246-165161 Clostridium proteolyticum DSM 3090 genome assembly, contig: EJ39DRAFT_scaffold00001.1
AAAAGATAAGGGAATGAAGTTCTCCCTAGTATAATACTGTTGTATTATATTTAAACCGCA
TTTATGCTGATGACTTCTGCTATTAT
>MKWR01000013.1/156999-156925 Xanthomonadales bacterium 63-13
scnpilot_p_inoc_scaffold_240, whole genome shotgun sequence.
CCCGTCCGCGGAGATGGCATTCCTCCGGCCCCATTCGGGCAAACCGCCGCCTCGGCTGAT GATGCCTGCCCCACC
>BAND01000009.1/12015-11942 Acidomonas methanolica NBRC 104435 DNA, contig: Amme_009.
ATCAACGAAGGCAATGGACTCTGCCTGGCCGCCGTGGCCGAACCGCCCCCAGGGCTGATG
ACTCCTACCTCGCC
>CP000285.1/2902118-2902253 Chromohalobacter salexigens DSM 3043, complete genome .
GTGGATGCCGGAGATGGCATTCCTCCGGTCGCCGAGCCCCGTCATGTGCGGTGCCACCGC
GACAAACCGCCTGGCGTGATCCGAGACCATGACGTCTCGACCACCGGCCGCCAGCGCTAA
TGATGCCTGCCAACGC
>HF998200.1/85537-85605 Bacteroides plebeius CAG:211 genomic scaffold, scf611
TCGAATCAAGGGAATGGGACTTCCCTGTGATAAACCGCTAATGAAGTAGCTGATAGTTCC
TACCGGAGC
>LFMX01000014.1/29847-29783 Candidatus Burkholderia humilis strain UZHbot5 BHUMctgl4, whole genome shotgun sequence.
TTTGACATAGGAGATGGCATTCTCCTTCAACCGCTCGCCCTCGAGCCGATGATGCCTACC
CCGCG
>FP929047.1/1868297-1868221 Gordonibacter pamelaeae 7-10-1-b draft genome.
AATGCCTGCGGGAATGAGGTTCTCCCTAAGCCGTTTCAAGGCTTGAACCGCGTAAAGCTG
ATGACTTCTGCAAACGG
>MNZR01000173.1/200-282 Syntrophobacteraceae bacterium CG2_30_61_12 cg2_3.0_scaffold_29238_c, whole genome shotgun sequence.
CATGCGACAGGCAATGAAGTCTGCCTGAAACCGCCCGGCTCTGCCCGCGAGCCGCCTGCA GGGCTGATGACTTCTACCTCGTT
>MIHC01000025.1/17004-17065 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
ATTGACCCAGGCGATGAAGCTCGCCTCAACCGCCGCACCCGGCTAATGGCTTCTACCGCG CG
>CP004036.1/3514107-3514034 Sphingomonas sp . MM-1, complete genome. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GTCGGCAACGGCAATGGATTCCTGCCGGGCCTCGCGCCGAACCGCCATTGAGGGCTGATG
ATTCCTACCTGCGG
>LGK001000002.1/1173186-1173246 Thermanaerothrix daxensis strain GNS-1 contig_l, whole genome shotgun sequence.
ATACTTGGTGGCGATGAGGCTCGCCGTAAGTGTCGGTATGACTAATAGCCTCTACTTGAA
C
>MGQE01000007.1/105-44 Deltaproteobacteria bacterium RBG_13_51_10
RBG_13_scaffold_11549_curated, whole genome shotgun sequence.
ATTTATTACGGTGATGGAGTCCACCGAAGAGTGCCTCAAAGGCTGATGACTCCTACCAAG GT
>MEOC01000004.1/19599-19659 Bacteroidetes bacterium GWE2_42_42
gwe2_scaffold_115, whole genome shotgun sequence.
TATTTTTCGGGCGATGAGGTTCGCCCTGAACCGCCGAAAGGCTGATGACCTCTGCAGATA
A
>MKSM01000141.1/6408-6349 Nitrobacter sp . 62-23
SCNpilot_cont_300_bf_scaffold_967 , whole genome shotgun sequence.
GACCAAGATGGGGATGGAGTCCCCCGAAACCGCCGCAAGGCTGATGACTCCTGCCGCGCG
>LQYT01000056.1/35438-35503 Caldibacillus debilis strain B4135 NODE_72, whole genome shotgun sequence .
TTATCATCCGGTGATGGAGTTCCGCCGTTAACCTCCTTGCAAAAGGCTGATGACTCCTGC
TGAAAA
>MKSH01000031.1/2971-2907 Solirubrobacterales bacterium 70-9
SCNpilot_bf_inoc_scaffold_1372, whole genome shotgun sequence.
CCTGTTCCGAGCGATGGGGCTCGCTCCATAACCGCGGCGCGACCGTTGATGGCTCCTACC GACGT
>CP006019.1/502086-502161 Palaeococcus pacificus DY20341, complete genome.
CCTAAGCTGGGCGATGGCGTCCGCCTGGGCTTCGAGCCCAAACCGCCCGCAATGGGCTAA
TGACGCCTGTTCTCAC
>CP017269.1/1004610-1004672 Geosporobacter ferrireducens strain IRF9, complete genome.
ATTTCAATAGGTGATGGGGCTCACCACTAATATTGCTTAGCAATTGATGGCTCCTACTAA
AGT
>CP003360.1/2604039-2603977 Desulfomonile tiedjei DSM 6799, complete genome .
ACAGATGACGGGGATGGAGTCCCCCCAAAATGCTCGTAATGAGCTGATGACTCCTACCGC
TAC
>CP017749.1/2273623-2273694 Cupriavidus sp . USMAA2-4 chromosome 2, complete sequence.
GCATCACGCGGAGATGGCATTCCTCCTCTAACCGCCGTCGTTTTATATGCGGCTGATGAT
GCCTACAGGTCC
>CP002987.1/1840156-1840226 Acetobacterium woodii DSM 1030, complete genome .
CTATATAAAGGGGATGAAGTTCGCCTTGGCCAATGCCAAAATGCTGATAAGCTGATGACT
TCTGCAAATAG
>FO000002.1/10207-10266 Methylocystis sp . strain SC2 plasmid 2.
CCTTGTCCCGGCGATAGAGTTCGCGCTCAACCGCCTGCGGCTGATGACCCCTGCTCTACA
>JQBX01000016.1/12805-12868 Pediococcus stilesii strain DSM 18001
Scaffoldl6, whole genome shotgun sequence.
AATTGAATAGGCGATGACGTTCGCCGTAAATAATTGATAATAATTTGATGACGTCTATTA
CTTT
>MEZF01000053.1/3627-3699 Candidatus Bathyarchaeota archaeon RBG_13_52_12 RBG_13_scaffold_25772 , whole genome shotgun sequence.
ATGTTTTTCGGCGATGAAGTCCGCCTAAGCTGTCGCTTAAACCGCCCGTGGGGCTGATGA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CCTCTACAACCCA
>CP000764.1/3708919-3708978 Bacillus cereus subsp. cytotoxis NVH 391-98, complete genome.
AATGAATACGGCGATGGAGTTCGCCACAACCGCTGCTTAGCTAATGACTCCTACCAGGAT
>NGKU01000001.1/2071454-2071526 Enterococcus sp . 8G7_MSG3316
scatfoldOOOOl, whole genome shotgun sequence.
CTCGACAAAGGAAATGAATGTCTTCCTCAGTTTTGACTGAAACCGCAAATTTGCTAATGA
CTTCTACCACTTA
>MPPL01000001.1/5613334-5613400 Mucilaginibacter polytrichastri strain RG4-7 Scaffoldl, whole genome shotgun sequence.
GTACTGAAAGGCAATGGTGAGCTGCCTATTTGAACCGCCTTAAAAAGCTGATGACGCCTA CTCGAAT
>FR883297.1/5583-5511 Eubacterium sp . CAG:192 genomic scaffold, scf283
TTAAATAAAGGGAATGAGGTTCTCCCTTGATTTTATCAAAACCGCTAGTAAAGCTGATGA
CTTCTGTGTAACA
>CP000254.1/1236994-1237053 Methanospirillum hungatei JF-1, complete genome .
AATCCTGCCGGTGATGAGATCCACCGTAAACCGCCCGCTGCTAATGACTTCTACCCGGAA
>FQXE01000023.1/21421-21489 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_123
TTCAGATACGGAGATGGCATACCTCCTATCCCAACCGCTGGATTCCCAGCTGATGATGCC
TACGAATTT
>LBBT01000022.1/3006-3067 Clostridium sp . JC272 Contig23, whole genome shotgun sequence.
AATTAAATCGGCGATGGAGTTCGCCATTAAATGCATAAGATGCTAATGACTCCTACTCTT
AT
>MHIY01000004.1/11965-12026 Candidatus Colwellbacteria bacterium
RIFCSPLOWO2_01_FULL_48_10 rifcsplowo2_01_scaffold_11379, whole genome shotgun sequence.
AGTTACAATGGCGATGGAGTTCGCCAAAACCGCCTTCACAGGCTGATGACTCCTGCCAAT
TA
>AP012027.1/1689511-1689454 Erysipelothrix rhusiopathiae str. Fujisawa DNA, complete genome.
TTACAAAACGGTGATGGGGTTTGTCATAACTGCGACAGCTGATAACTCCTACATAAGA
>CP009928.1/3333143-3333076 Chryseobacterium gallinarum strain DSM 27622, complete genome.
GTATCAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTACAGAAGCTGATGACGCCT GATTGAGA
>AOIV01000036.1/89301-89233 Halogeometricum pallidum JCM 14848 contig_36, whole genome shotgun sequence .
GGTAGAACAGGCGATGGAGTCCGCCTAATCCAACCGCCGAATCCGTCGGCTGATGACTCC
TGTCTATGA
>GL636579.1/244140-244213 Coprobacillus sp . 29_1 genomic scaffold supercontl .3, whole genome shotgun sequence.
CTATAGCAAGGGAATGAATGTCTCCCTTAGTTTAAGAACTAAAACCGCTACATGCTGATG
ACTTCTGCATTTTT
>MEQZ01000078.1/2332-2268 Betaproteobacteria bacterium
RIFCSPLOWO2_02_FULL_65_20 rifcsplowo2_02_scaffold_l 8574 , whole genome shotgun sequence.
ATTCTGCCGGGAGATGGCATTCCTCCTTTAACCGCCGCGCGCCGGCTAATGATGCCTACA
CGGAA
>CP001560.1/792938-792867 Shimwellia blattae DSM 4481 = NBRC 105725, complete genome. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ATTCCGGGCGGAGATGGCATTCCTCCCGTGTAACAAACCGCCTTTGCGAAGGCTGATGAT
GCCTGCATGATT
>CP000115.1/166253-166313 Nitrobacter winogradskyi Nb-255, complete genome .
GAAATCGATGGGGATGGAGTCCCCCGATAACCGCCGCAAGGCTGATGACTCCTACCGGGC
G
>AGIZ01000001.1/281997-282071 Fischerella sp . JSC-11 ctgll8, whole genome shotgun sequence.
GATGTTAAAGGTGATGAAACTCACCTTCTGGCATTTCCAGAGAACCGCCACTCGGCTAAT
AGTTTCTAGCTCAAC
>GL982453.1/6334526-6334589 Achromobacter insuavis AXX-A genomic scaffold scaffold00003, whole genome shotgun sequence.
AACGCAAATGGGAATGGAGCTTCCCAAAACCGCCGATCCGTCGGCTGATGGCTCCTGCTG
ATCC
>MHEE01000007.1/18381-18474 Nitrospirae bacterium GWF2_44_13
gwf2_scaffold_206, whole genome shotgun sequence.
TTTAGTATCGGCGATGGAGTTCGCCATTAACTCCGTTAAAAAATACGGAGTAAACCGTCC CGATCTTATCGGAACTGATGACTCCTACTCTCAC
>ADMN01000037.1/24546-24488 Turicibacter sanguinis PC909 contig00012, whole genome shotgun sequence .
TTATCAATAGGCGATGGAGTTCGCCATTAACCGCGTAAGCTAATGACTCCTACTCTTTA
>CP000240.1/171311-171206 Synechococcus sp . JA-2-3B ' a (2-13) , complete genome .
CCGGCAACCGGCGATGGAGTTCGCCGCAACCGCCTCACTTGGAATTTTTCTATAGAGGAA
GAGGGTTCTAACCTCAAAAGCTGAGGCTGATGACTCCTACACTTCT
>LT629799.1/1480323-1480250 Friedmanniella sagamiharensis strain DSM 21743 genome assembly, chromosome: I
TTGTCCCCTGGCGATGGATTCCGCCGGAGGTCGCGACCTCGAACCGCCGCACGGCTGATG
ACTCCTGGACGCTC
>MKRJ01000017.1/40298-40357 Rhizobiales bacterium 65-79
SCNpilot_cont_1000_bf_scaffold_222 , whole genome shotgun sequence.
TATCCAACAGGGAATGGAGTCTCCCTTCAACCGCCGTCGGCTGATGACTCCTGCTGCGTA
>MGRA01000126.1/7736-7664 Deltaproteobacteria bacterium
RBG_19FT_COMBO_46_12 rbg_19ft_combo_scaffold_1869, whole genome shotgun sequence .
ATCAATCTCGGTGATGAAGTTCGCCATTAACCGCTCCCCCCTTTTATCGGGGGCTGATAA
CTTCTACTAGTTT
>LBGU01000041.1/56834-56772 Sphingobacterium sp . Agl JX05_contig_20, whole genome shotgun sequence.
AATAGATATGGCAATGATGTCTGCCCTTAACCGCCTTAATAAGCTGATGACGTCTACTAT
TTT
>LMLS01000005.1/17281-17205 Sphingomonas sp . Leaf231 contig_5, whole genome shotgun sequence.
GCTCGGCACGGCGATGGATTTCCGCCGGGCCATTCGGCCGAACCGCCCTCGCAAGGGTCG
ATGATTCCTACCAGGCG
>JNBQ01000001.1/173215-173288 Cellulosimicrobium funkei strain Ull contigOOOOl, whole genome shotgun sequence.
GCTGAGGGCGGTGATGGATCCCGCCAGGGCTAGCCGCCCGAACCGCCGCCAGGGCTGATG
GTTCCTGTCTCGTC
>MESK01000107.1/33086-33014 Burkholderiales bacterium
RIFCSPL0W02_12_FULL_64_99 rifcsplowo2_12_scaffold_76, whole genome shotgun sequence .
CCACACGTGGGAGATGGCATGCCTCCCTGAACCGCCCCGATCCCCTTCGGGCGCTGATGA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TGCCTACAGAAGT
>CP009687.1/513822-513885 Clostridium aceticum strain DSM 1496, complete genome .
TAATTAAATGGCGATGGAGTTCGCCCTTAACCACTGAAATTTAGTTGATGACTCCTGTAA
AACA
>AP012224.1/865518-865455 Pseudogulbenkiania sp . NH8B DNA, complete genome .
TCCGCGACAGGAGATGGCATGCCTCCTTGTAACCGCCCTCGGGGCTGATGATGCCTACAC
GGAT
>AODQ01000006.1/6334-6265 Cesiribacter andamanensis AMV16 contig000006, whole genome shotgun sequence .
GCAGCCATAGGCAATGGTGTCTGCCTTGAACCGCCTGGATACAGCTCAGGATGATGGCGC
CTGCTCCCAG
>AKFU01000006.1/20391-20332 Clostridium sp . MSTE9 ctgl20006482988, whole genome shotgun sequence.
AATATTCAAAGTGATGAGGCTCACTTGAATTGCTGAAAAGCTGATGGCTTCTACCTTAAG
>LJGW01000451.1/16452-16528 Streptomyces nanshensis strain SCSIO 10429 scaffold449, whole genome shotgun sequence.
CTGAGCCGAGGCGATGGAACCCGCCCGGGCCGCTTCCGGCCCGAACCGCCCTCCGGGCTG
ATGGTCCCTGACCGAGG
>FPBV01000009.1/77626-77565 Alicyclobacillus macrosporangiidus strain DSM 17980 genome assembly, contig: Ga0104483_109
TGGAACATGGGTGATGGGGCTCACCTAAACCGCCTGCCAAGGCTGATGGCTCCTGCGAAC
CG
>FNYK01000003.1/65728-65798 Sharpea azabuensis strain DSM 20406 genome assembly, contig: Ga0059086_1003
TAAGCACAAGGGAATGAATGTCTCCCGCGAGCAATCGAAACCGCTATTAAGCTGATGACT
TCTATCTTATT
>MIBJ01000074.1/1274-1209 Spirochaetes bacterium RBG_16_49_21
RBG_16_scaffold_826, whole genome shotgun sequence.
TGTTATCAAGGGGATGGAGTCCCCCGTATAAATGCACAATGTTCGGCTGATGACTCCTAC CGCATC
>AJXU01000007.1/18977-19073 Rhodanobacter fulvus Jip2 contig007, whole genome shotgun sequence.
GCCGCTCCGGGAGATGGCATGCCTCCCGTTCCCGCGACAGGGTCGAGGGAACCAACCGCC
CCGCATCACTGCAGGGGCTGATGATGCCTGCACCACC
>ABOX02000009.1/3464-3546 Pedosphaera parvula Ellin514 ctgl90, whole genome shotgun sequence.
ATATGCATTGGCAATGGGGTTTGCCTTCTCGGTCAACGAGAGAACCGCCTGATTATTTCG
GTGCTGATAACTCCTATTGGTAC
>MJMJ01000034.1/323335-323260 Vibrio panuliri strain CAIM 703 CAIM703_4, whole genome shotgun sequence .
TAGCGACAAGGTGATGGGGTGCCACCGGAATCAAAAGATTCGAACCGCTTATAAAGCTGA TGACTCCTACAGAAAC
>LT607756.1/1813946-1813868 Methanobacterium congolense isolate Buetzberg genome assembly, chromosome: I
CCATTAACAGGTGATGGAGTTCACCTTTAATTGCCAAAGATATATCTTTTATCCTTTGGA
TGATGACTCCTACCGTAAG
>MNEA01000136.1/8806-8734 Acidobacteria bacterium 13_2_20CM_2_57_6
13_2_20cm_2_scaffold_458, whole genome shotgun sequence.
CATTCTGGAGGCGATGGAGTTCGCCATAACCGTCCCGGCGCTGTTCACCGGGTCTGATGA CTCCTGGCAGCAG
>CYGY02000045.1/1800-1869 Burkholderia sp. STM 7183 genome assembly, APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
contig: CYGY01000045
GTGTCTGCCGGAGATGGCATTTCTGCCTAACCGCCGCTGCTCGTGTGCGGCTGATGATGC
CTACTGGGCC
>MVHR01000005.1/136280-136345 Mycobacterium heidelbergense strain DSM 44471 NODE_5_length_243907_cov_56.1658, whole genome shotgun sequence. TTCGCGGCCGGCGATGAAGCTCCGCCGCCTTTCGCGCCGTCCGGCGCTGATGGCTTCTAC CCCGCA
>MEMX01000112.1/17983-18063 Anaeromyxobacter sp . RBG_16_69_14
RBG_16_scaffold_4847, whole genome shotgun sequence.
CGACATCGAGGCGATGAAGCTCGCCCGTCGTTCGGGACGCCGCCAACCAGCCCGAGAGGG GTTGATAGCTTCTGCGAGGCC
>CP010817.1/973882-973815 Myroides profundi strain D25, complete genome.
TAAATCATGGGCAATGATGTCTGCCTCGAACCGCTTCTCTATAGAATGCTGATGACGTCT
AACATAAT
>AKKE01000136.1/14252-14328 Novosphingobium sp . AP12 PMI 02_contig_294.294 , whole genome shotgun sequence .
CAAGGCGACGGCGATGGATTTCCGCCGGGCCACTTGGCCGAACCGCCCTCGCAAGGGTCG
ATGATTCCTACCAGGCG
>AP008971.1/1641037-1640963 Finegoldia magna ATCC 29328 DNA, complete genome .
ATAGATTTGGGGAATGAAGTGCTCCCTTGTTATTAACTAAAACGCAAAAATTTTGCTGAT
GACTTCTGTTTTTTA
>MNFK01000059.1/3686-3748 Deltaproteobacteria bacterium 13_1_40CM_4_54_4 13_1_40cm_4_scaffold_4977 , whole genome shotgun sequence.
TCCATTCGCGGAGATGGCGTTCTCCTCTAACCGCCCACCAGGGCTGATGACGCCTACTTC AAT
>MZGT01000023.1/10005-10064 Clostridium chromiireducens strain DSM 23318 CLCHR_contig000023, whole genome shotgun sequence.
TATATTTTAGGTGATGGAGTTCGCCTTTAACTGCGTAATGCTAATGACTCCTACAAAATA
>FRAR01000004.1/46252-46311 Desulfotomaculum aeronauticum DSM 10349 genome assembly, contig: EJ44DRAFT_scaffoldOOOOl .1
AATTAATACGGCGATGGAGTTCGCCATAACTGCCCTTGGGCTAATGACTCCTACCTAATT
>FOVK01000003.1/204991-205066 Proteiniclasticum ruminis strain ML2 genome assembly, contig: GaO 073291_103
TAGACTTCTGGAAATGAAGTGCTTCCATGGCAAAAACTGCCAAAACCGCAAGTATGCTGA TGACTTCTACGACGAA
>LWMT01000242.1/9864-9792 Methanobrevibacter filiformis strain DSM 11501 MBFIL_contig000242 , whole genome shotgun sequence.
TTAATTAAAGGTGATGAGGTTCACCTATTTAACTGCCAACTTTTTAAAAGTGGCTGATGA CCTATATAAATAC
>LWDV01000010.1/757336-757414 Orenia sp . Z6 2524618313, whole genome shotgun sequence.
CATAAAATTGGTGATGAGGCTTACCTTGATATTCTTGATATATCAAAACCGTCTTTAGAC
TAATAGCTTCTACTCCCTA
>MKSM01000141.1/15485-15423 Nitrobacter sp . 62-23
SCNpilot_cont_300_bf_scaffold_967 , whole genome shotgun sequence.
TTTGTACATGGCGATGGAGTTCGCCTTTTAACCGCCCTCGGGGCTGATGACTCCTGTCGA
GCC
>MKSZ01000133.1/24711-24647 Bacteroidales bacterium 36-12
SCNpilot_bf_inoc_scaffold_680 , whole genome shotgun sequence.
ATAGAAAAAGGGAATGGGACTTCCCTGAAACCGCTTGAAATATAGCTGATAGTTCCTACT TAATG
>CP014578.1/1835529-1835458 Burkholderia sp . OLGA172 chromosome 1, APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
complete sequence.
TTGCGCGCTGGGGATGGCATTCTCCACATGCTCCATTCGCCGCCCTTGATGGCTGGTGAT
GCCTGCTTCGCC
>MGSF01000017.1/5216-5151 Deltaproteobacteria bacterium
RIFCSPLOWO2_02_FULL_50_16 rifcsplowo2_02_scaffold_13928, whole genome shotgun sequence.
GAAAACTTTGGCGATGGAGTTCGCCGTATAACCGCCTCAACACGGGCTGATAACTCCTAC
TCGTTA
>MNVK01000062.1/4795-4860 Nitrospirae bacterium CG1_02_44_142
cgl_0.2_scaffold_17662_c, whole genome shotgun sequence.
TTTAGTGTCGGCGATGGAGTTCGCCGTTAACTGTCCTTTTAAGTGGCTGATGACTCCTAC CACCCC
>BABS01000034.1/15894-15829 Acetobacter tropicalis NBRC 101654 DNA, contig: 0034.
CGCGGCACAGGGGATGGGGTTCCCCCTTATAACCGCCCGCTCGGGGCTGATGACTCCTGC
CGTTAT
>LVZN01000005.1/54206-54135 Erysipelotrichaceae bacterium MTC7 Contig_13, whole genome shotgun sequence .
CTATGCGCAGGGAATGAAATCTCCCTCGATTACATTATCGAAACCGCGCAAGCTAATGAT
TTCTGCATTTTT
>AYZK01000011.1/8809-8871 Lactobacillus thailandensis DSM 22698 = JCM 13996 strain DSM 22698 Scaffoldll, whole genome shotgun sequence.
TAAGTTCACGGCGATGACGTTCGCCTTTAAATATTCGTAAGGATTGATGACGTCTGTTTA TTC
>CM001633.1/2572890-2572776 Oscillatoriales cyanobacterium JSC-12 chromosome, whole genome shotgun sequence.
TAGGGGGATGGCGATGGAGATCGCCAAAACCGCCTTTTGTTTTGCCTAAAGTCTAAAATC
TGGTGATTTGAAACCTAAAAGCTGAAAACATAAGGCTGATGGCTCCTACGATTCC
>JWKP01000014.1/24045-24114 Archaeon GW2011_AR4 QS99_C0014, whole genome shotgun sequence.
TTATCTCTTGGCGATGAAGTCCGCCATGGATATGCCCAAAACCGCCTCAGCTGATGACTT
CTAACAAAGG
>AE008691.1/1840761-1840702 Thermoanaerobacter tengcongensis MB4, complete genome .
TGCATAAAAAGTGATGGAACCCACTTTAACCGCCGAAAGGCTGATGGTTCCTACTTGTGA
>LRVM01000010.1/55507-55434 [Clostridium] neopropionicum strain DSM-3847 CLNEO_contig000010, whole genome shotgun sequence.
TAAATTCTTGGGAATGAAGTTCTCCCAAGGGTTAAACCTAAACCGCTTATTTAGCCGATG ACTTCTGCAATTTA
>CELZ01000010.1/89700-89758 Moorella glycerini strain NMP genome assembly, contig : M_glycerini_NMP_DRAFT_scaffold-10
AACCCCAACGGCGATGGAGCTCGCCTAACGCTCAGTGAGCTAATAGCTCCTACCGAAAG
>LM995447.1/668610-668671 [Clostridium] cellulosi genome assembly, chromosome: I
CTTTGTGTTGGTGATGGGGCTCACCATAATGCATTACAGATGCTAATGGCTCCTGCTTTA
AT
>CP017749.1/2322684-2322754 Cupriavidus sp . USMAA2-4 chromosome 2, complete sequence.
GTCCTCATAGGAGATGGCGTTCCTCCTCATAACCGCCACGCGTTCGCGTGGCTGATGACG
CCTACAGACTT
>CP002408.1/2823807-2823887 Candidatus Nitrososphaera gargensis Ga9.2, complete genome.
AAAATGACTGGCTATGGCATTAGCCTTGGCGTGAGCCAGAACCGCCCTTTTTCTAAAGGA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GCTGATAATGCCTACCAACAA
>CP011270.1/4418827-4418743 Planctomyces sp . SH-PL14, complete genome.
GAATACCTTGGCGATGGGGTTCCGCCATAACCGCCGGGAGAATGCCGCAAGGTCTGCCTC
CCGCGCTGATGACTCCTACCGATTG
>AE001437.1/1730397-1730460 Clostridium acetobutylicum ATCC 824, complete genome .
AATTTCATAGGTGATGAGGTTCGCCCGTAAATATCCTGTTTGGATTGATGACTTCTGCAT
AATT
>CP009415.1/653563-653636 Candidatus Izimaplasma sp . HR1 chromosome, complete genome.
TTATTTTTAGGGAATGAAGTACTCCCTTAGTATTTTTACTAAAACCGCTATAAGCTGATG
ACTTCTACGTTATT
>FP565809.1/130293-130214 Clostridium sticklandii str. DSM 519 chromosome, complete genome.
TATATTTATGGGAATGAAGTTCTCCCCTAGCAGTATTTGCTAAAACCGCTTTTTCTAGAG
CTGATGACTTCTGTGAGTAT
>FTMS01000015.1/78533-78471 Spirochaeta americana strain ASpGl genome assembly, contig: Ga0071122_115
TTGGCGCCTGGCGATGGGGCTCGCCACATAACCGCCTCTCCGGCTGACAGTCCCTGTACA
ACA
>MGP001000019.1/6796-6728 Deltaproteobacteria bacterium GWA2_65_63 gwa2_scaffold_16628, whole genome shotgun sequence.
AGCAACCAAGGCGATGGAGTTCGCCGATAACCGCCTGTCGAACAGACGGCTAATGACTCC TGCGATCCC
>BBJS01000043.1/250767-250843 Sphingomonas paucimobilis NBRC 13935 DNA, contig: SP643.
CGCAGGCACGGCGATGGATTTCCGCCGGGCCATTCGGTCGAACCGCCCTCGCAAGGGCCG
ATGATTCCTACCAGGCG
>FQUT01000001.1/582597-582530 Chryseobacterium arachidis strain DSM 27619 genome assembly, contig: Ga0131169_101
GCTATAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCCTTAAAAAAGCTGATGACGCCT
GATTAAAA
>AL646052.1/1941892-1941826 Ralstonia solanacearum GMI1000 chromosome complete sequence
CCGGCCTGCGGAGATGGCATGCCTCCCTTAACCGCCGATCCCATCGGCTGATGATGCCTA
CAGGTTC
>MAQW01000012.1/326540-326477 Bacillus sp. FJAT-27264 scaffold2, whole genome shotgun sequence.
GATGAATCGGGCGATGGAGTTCGCCATTAACCGTCTTTTGCAGACTAATGACTCCTACCA
GTGA
>AOKF01002523.1/2016-2088 Pseudomonas syringae pv. actinidiae ICMP 19096 scaffold37, whole genome shotgun sequence.
CGGCGCATTGGAGATGGCATTCCTCCATTAACCAACCGCTGCGCCCGTAGCAGCTGATGA
TGCCTACAGAAAC
>LMSL01000040.1/129296-129357 Frateuria sp . Soil773 contig_8, whole genome shotgun sequence.
TTCCGATGGGGCGATGGTGTCCGCCTTCGACTGCCCTCAGGGCTGATGACACCTACAGGC
TA
>MVHR01000005.1/161584-161523 Mycobacterium heidelbergense strain DSM 44471 NODE_5_length_243907_cov_56.1658, whole genome shotgun sequence. TGCGAGTAAGGCGATGAAGCTCGCCTCAAATGCCGCCCCTGGCTAATGGCTTCTACCCAG TG
>MGQR01000021.1/6131-6068 Deltaproteobacteria bacterium RBG_16_50_11 APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
RBG_16_scaffold_125085, whole genome shotgun sequence.
AATTGATTTGGTGATGGAGTTCACCGTATAACCGCCTAAGAAGGCTGATAACTTTTGTTT
TGGG
>CP001771.1/142912-142849 Spirosoma linguale DSM 74 plasmid pSLIN02, complete sequence.
CCTATCAAAGGTGATGGAGTTCGCCTGGAACCGCCCTCAAAAAGCTGATAACTCCTGTAT
TCGT
>HE956757.1/1008434-1008493 Methylocystis sp . SC2 complete genome
TTCCCTACGGGATGGGGCTCCCGGTAACCGCCTGATGAGGCTGATGGCTCCTGCCGCAAG
>MKIN01000018.1/220870-220794 Rhizobium taibaishanense strain 14971 C181, whole genome shotgun sequence .
ATCGGCAACGGTAATGGATTCTGCCGGGCATTAAGCCGAACCGCTTCCCCGATGAAGCTG
ATGACTCCTACTCAACG
>MNXM01000330.1/9225-9166 Armatimonadetes bacterium CG2_30_59_28
cg2_3.0_scaffold_2174_c, whole genome shotgun sequence.
GAGAAGGCTGGCGATGGAGTTCGCCAATTGTCCGTGAGGACTGATGACTCCTATCCGATT
>CP006569.1/3268057-3268120 Sodalis praecaptivus strain HS1, complete genome .
GACCCATCAGGAGATGACATTCCTCCTTATAACCGCCGTTCTGGCTGATGATGTCTACGT
TCGC
>HF986880.1/58759-58689 Firmicutes bacterium CAG:555 genomic scaffold, scf21
CATCATATAGGGAATGAGGTCTCCCGTGGAGGTATCCAAAATCGCTTAGTGCTGATGACT
TCTGAAAAATC
>CP000975.1/1993796-1993891 Methylacidiphilum infernorum V4, complete genome .
GAAATCCCTGGCAATGGAGTTTGCCGTCCCGCCGGTTCGAAGTACTTGGGAAAACCTCCT
TTCTTGAAAAGAAAGGCTGATAACTCCTACTTCAAC
>CP002868.1/2576532-2576609 Treponema caldarium DSM 7334, complete genome.
AAGGTATGAGGTGATGGAGTTCACCAACGAGCGCATTCGCTCTGAACCGCCTTCCTGGCC
GATAACTCCTATGAAGGA
>HF993782.1/19382-19452 Eubacterium sp. CAG:603 genomic scaffold, scfl02
ATACCACTTGGGAATGAAGTTCTCCCTTGGATAACCAGAACCGCTTATTAGCTGATGACT
TCTGTAATTAA
>FPBV01000009.1/76554-76616 Alicyclobacillus macrosporangiidus strain DSM 17980 genome assembly, contig: Ga0104483_109
AATTCGATAGGTGATGGAGCTCACCGTTTACTTCCTGCAAAGGATGATGGCTCCTACGGA
TAG
>AE016830.1/2384514-2384445 Enterococcus faecalis V583, complete genome.
TAGCAACATGGTGATGGTGTTCACCACGAACTATTTATTGGACGAATAAATTAATGACGC
CTACCAAACG
>AP006620.1/61342-61253 Nocardia farcinica IFM 10152 plasmid pNF2 DNA, complete sequence.
GAGTCCGCAGGCGATGGATCTCCGCCGAGACGACCGACCTATGGGCCGTCTGAACCGCCC
CGGTCCGGGGCTGATGGTTCCTTCCCCACG
>MGBU01000051.1/1810-1744 Candidatus Rokubacteria bacterium GWA2_73_35 gwa2_scaffold_17261, whole genome shotgun sequence.
CGCCGTCAAGGCGATGGGGTTCGCCTGAAACCGCCCGACGCGTCGGGCTGATGACCCCTA CCGGAGT
>CCES01000002.1/45276-45339 Serratia symbiotica genome assembly, type strain CWBI-2.3T, contig: SYMBAF_Contig_l0
ACAGCCCCTAGAGATATCATTCCTCCCAAACCGCCTTCGCAAGGCTAATGATGTTTACGT
ATAG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>CP002541.1/2695201-2695118 Sphaerochaeta globosa str. Buddy, complete genome .
CTGCATTCAGGCGATGGGGTTCGCCGAAACCTATTTGGTATAAACCGCACGTTATAAACG
TTTGCTGATAACTCCTACAAATTG
>CP000633.1/960988-960912 Agrobacterium vitis S4 chromosome 1, complete sequence .
CACGGTAACGGCAATGGATTCTGCCGGGCACGACGCCGAACCGCTTCCTTGATGAAGCCG ATGACTCCTACTCAACA
>JGZP01000012.1/331381-331455 Bifidobacterium stellenboschense strain DSM 23968 Contigl2, whole genome shotgun sequence.
TTGCACGCAGGCGATGGAACCCGCCCGGAAACCCGTTTTCCGAACCGCATTACCGCTGAT AGTTCCTACACGATC
>MKCR01000003.1/233446-233511 Chromobacterium amazonense strain DSM 26508 Camzn_DSM26508.11, whole genome shotgun sequence.
ATCTGATGCGGAGATGGCATTCCTCCCGTAACCGCCTTTGCCAAGGCTGATGATGCCTAC AGTATT
>MHYI01000054.1/4844-4783 Planctomycetes bacterium RBG_16_41_13
RBG_16_scaffold_16981, whole genome shotgun sequence.
GAATAAAGAGGCGATGGAGTTCGCCTGATTGCTATGTGATAGCTGATAACTCCTATTGAA GT
>LGET01000097.1/121-197 Thermococcales archaeon 44_46 PW_scaffold_576, whole genome shotgun sequence .
AAGGTCTTGGGCGATGGCGTCCGCCTGGGGCTTCGAGCCCGAACCGCCCGCTGAGGGCTG
ATGACGCCTGTTCTCCG
>LMPU01000008.1/234928-234996 Pedobacter sp . Leafl94 contig_2, whole genome shotgun sequence.
TTGCAATACGGAAATGGTGTCCTTCCGATTCAACCGCTTTTTTCAAAAGCTGATGGCGCC
TACAATAGA
>FUKX01000032.1/617494-617561 Sphingobacterium faecium PCAi_F2.5 genome assembly, contig: Scaffold2
AAACAAAAAGGAAATGGTGTTCTTCCTTAACCAACCGCCCTCAAAAAGCTGATGACGCCT
GATTAAAA
>MDET01000045.1/79833-79771 Pseudaminobacter manganicus strain JH-7 JH- 7_scaffold5, whole genome shotgun sequence.
TGGAAACAAGGGGATGGGGTTCCCCAAAACCGCCCGACAAGGGCTGATGACTCCTGCTGT
GTG
>CP001275.1/1997016-1996938 Thermomicrobium roseum DSM 5159, complete genome .
GAGTCGAACGGCGATGGAGTCCGCCGGGGCTCGTCGAGAGCCCGAAGCGCCAGAGAAGGC
TGATGGCTCCTGCTCGCAG
>BA000043.1/2726409-2726347 Geobacillus kaustophilus HTA426 DNA, complete genome .
AATCGAATAGGCGATGGAGTTCGCCATAACCGCCGGCTTCCGGCTGATGACTCCTGCTGC
AAA
>AJFI01000022.1/155770-155831 Mycobacterium xenopi RIVM700367 contig22, whole genome shotgun sequence .
ATCAGCCCAGGCGATGAAGCTCGCCTTGACCGCCACCAAAGGCTGATGGCTTCTACCACA
GA
>CP010796.1/1214934-1215003 Carnobacterium sp. CPI, complete genome.
AAGTAAATAGGCGATGGTGTTCGCCCTTAACTATCATAGTAAGTTTATGATTAATGACGC
CTACTTAAGA
>MEOY01001160.1/628-696 Bacteroidetes bacterium RIFCSPHIGHO2_02_FULL_44_7 rifcsphigho2_02_scaffold_415318, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TAATCAGAAGGCGATGAGGTTCGCCGAAAACCGCCCAAGATAAGTTGGGATAATAACCTC
TACTGAAAA
>MIHC01000025.1/46527-46463 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
TTCGCGGTCGGCGATGAGGCTCCGCCCGCTACGCGTCGCCCGACGCTGATGGCTTCTACC CCGCA
>MHEU01000060.1/5182-5261 Nitrospirae bacterium RIFCSPLOWO2_02_FULL_62_14 rifcsplowo2_02_scaffold_5531, whole genome shotgun sequence.
CATCGTTTAGGTGATGGAGTTCACCTAAAACTGTCCGTAACTTTGGTACAAGGTTGCGGA CTAATAACTCCTACCAGCAC
>CT573326.1/3719897-3719836 Pseudomonas entomophila str. L48
chromosome , complete sequence.
CGTGGCCCAGGAGATGGCATTCCTCCTTGAACCGCCTCAGCGCTGATGATGCCTACGCAA
CC
>LMJY01000017.1/142940-143016 Novosphingobium sp . Leaf2 contig_4, whole genome shotgun sequence.
CCGCACCACGGTGATGGATTTCCGCCGGGCTATTTGGCCGAACCGCCCTCGCAAGGGCCG
ATGATTCCTACTTGGCG
>AJYK02000040.1/277-347 Vibrio rumoiensis 1S-45 lS-45_contig_19, whole genome shotgun sequence.
CAAATCACGGGTGATGGGGTTCCACCTTTAACCGCTCAAAGTCTTTTTGAGATGATGGCT
CCTGCTAGGTG
>JPRJ01000053.1/18246-18179 Chryseobacterium piperi strain CTM contig53, whole genome shotgun sequence .
ATGATAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTACATAAGCTGATGACGCCT
GATTAATA
>MDZE01000022.1/430989-430909 Suitobacillus thermosulfidooxidans strain ZBY contig7, whole genome shotgun sequence.
TAAACACCAGGCGATGGATTCCGCCTTGGACCTAGGTCCAGAACCTCCGTGCATAGTGCG
GATAATGACTCCTGCCAGATT
>LGFR01000150.1/2134-2071 Clostridia bacterium 62_21 S_scaffold_2037, whole genome shotgun sequence .
GGTTCCACAGGCGATGGGGTTCGCCTGAACTACCGCGCGAGCGGTTGATAACCCCTACCT
GAAC
>LT838272.1/2862748-2862812 Thermanaeromonas toyohensis ToBE genome assembly, chromosome: I
TGTAATTGCGGTGATGGAGTCCACCGATAAAATACCCGTCAAGGGTTGATGACTCCTACC
CGAAA
>LNTD01000270.1/13282-13354 Cellulomonas sp . B6 contig_90, whole genome shotgun sequence.
GCTGGTCGCGGCGATGGATCCCGCCGGGGCGTGACGCCCGAACCGCCGCACGGCTGATGG
TTCCTGCTCCGGT
>CP001962.1/1051801-1051739 Thermus scotoductus SA-01, complete genome.
TAAAGAGAGGGCGATGGAGTCCGCCTGAACCGCCCGAAACGGGCTGATGACTCCTACCGG
GAA
>LGVB01000015.1/3022246-3022306 Clostridium sp . DMHC 10 contig_15, whole genome shotgun sequence.
AACTGAATAGGTGATGAAGTTCGCCTTAACTATCTTTTAGATTAATGACTTCTACATTTT
A
>MKVZ01000019.1/51156-51096 Rhizobiales bacterium 63-22
SCNpilot_cont_500_bf_scaffold_96, whole genome shotgun sequence.
CATCGATGCGGGGATGGAGTTCCCCTTCACAAGCCGGCCGGCTGATGACTCCTACCGTTG
G APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>FR888936.1/21289-21361 Firmicutes bacterium CAG:466 genomic scaffold, scf203
AGATTGTAAGGGAATGAGGTTCTCCCGCGGCAATGCCGGAAACGCTGAAACAGCTGATGA
CTTCTGTAAAATC
>MKSQ01000024.1/26007-25946 Bosea sp . 67-29
SCNpilot_expt_1000_bf_scaffold_559, whole genome shotgun sequence.
TTGCCGAGGAGGGATGGAGTCCCCCCAAACCGCCTTAACCGGCTGATGACTCCTACAAGC
GC
>LGTC01000001.1/6367484-6367424 Pseudobacteroides cellulosolvens ATCC 35603 = DSM 2933 ctgl, whole genome shotgun sequence.
AATTAAATAGGTGATGGAGTTCGCCATTAACTGCTATTATGCTAATGACTCCTACATAAT
T
>FP929041.1/1733794-1733866 Eubacterium cylindroides T2-87 draft genome.
CTTCTTTGTGGAAATGAAGTTCTTCCATGGGATAACCTAAACCGCTTATTAAGCTGATGA
CTTCTACGAGTTT
>JXAK01000039.1/9322-9384 Paenibacillus sp . VKM B-2647 B-2647_039, whole genome shotgun sequence.
TGTTCATAAGGCGATGGAGTTCGCCATAAACCGCCGGTAACGGCTAATGACTCCTACCAG
ATT
>BX571966.1/2539005-2538939 Burkholderia pseudomallei strain K96243, chromosome 2, complete sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>AEXE01000084.1/3368-3306 Burkholderia sp. TJI49 contig072, whole genome shotgun sequence.
GAACCGCAGGGAGATGGCATCCTCCCCAAACCGCCCGATCGGGCTGATGACGCCTGCCCA
CGC
>MDTU01000001.1/1022125-1022198 Piscirickettsia sp . Y2 scaffoldOOOOl, whole genome shotgun sequence .
TCTCTGGCAGGTGATGGAGTTCACCTTAAGCAATCGCTTTAAACCGCACTCATGCTAATG
ACTCCTGCTACAAT
>AZEY01000007.1/5773-5699 Lactobacillus diolivorans DSM 14421 NODE_13, whole genome shotgun sequence .
TATTAAAACGGCGATGACGTTCGCCCTTACTCAGCAGTGTTAACACACCAAAGAGTTGAT
GACGTCTACTTTAAC
>CP011129.1/5704756-5704821 Lysobacter antibioticus strain 76, complete genome .
CTTAGGGAAGGAGATGGCATTCCTCCTTGAACTGCCGGTTTTCTGGCTGATGATGCCTAC
CGCACC
>CP006721.1/434167-434231 Clostridium saccharobutylicum DSM 13864, complete genome.
ATTTTCATAGGTGATGAAGTTTGCCTTTAAACATCTCTTAGAAGATTAATGATTTCTACT ATAAA
>CP000924.1/2271016-2270957 Thermoanaerobacter pseudethanolicus ATCC 33223, complete genome.
ATACCATAAGGTGATGGAGCTCGCCATAAATGCGGAAACGCTGATGGCTCCTATTAGGAG
>HF988337.1/77505-77577 Clostridium bolteae CAG:59 genomic scaffold, scf126
ATAAGTACCGGGAATGAAGTTCTCCCTTAGTCTGACTAGAACCGCTTATAAAGCTGATGA
CTTCTGCATTATG
>MKUQ01000030.1/14795-14725 Burkholderiales bacterium 70-64
SCNpilot_expt_1000_bf_scaffold_318, whole genome shotgun sequence.
TTGGACGAGGGAGATGGCATTCCTCCCGCATTCGCGAAACCGCCGCAACGGCTGATGATG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CCTGCACCGCC
>CP016434.1/935257-935317 Burkholderia sp. AD24 chromosome I sequence.
TTGCGCGCTGGAGATGGCATTCTCCCTTAACCGCCCTCGGGCTGATGATGCCTGCTACGC
C
>JH597773.1/966425-966499 Leptonema illini DSM 21528 genomic scaffold Lepilscaffold_l , whole genome shotgun sequence.
TTTAGCTCAGGCGATGGGGTTCCGCCTTATAACCGCCCTTCTCTTTTGAAGCGTGCTGAT GACTCCTACCACTCA
>KQ969157.1/201392-201453 Streptococcus sp . DD10 genomic scaffold
scaffoldOOOOl, whole genome shotgun sequence.
TATCCATAAAGGGATGGTGCTCCCTATAAACGCTAGAAATAGCTGATGGCGCCTGTTAGA
AA
>MKRQ01000189.1/476-537 Rhizobiales bacterium 68-8
SCNpilot_expt_1000_bf_scaffold_8828, whole genome shotgun sequence.
ATGCGTGGCGGGGATGGGGTCCCCCGTATAACCGCCCTGTGGCTGATGACTCCTGCCAGC
GC
>CP000386.1/554775-554692 Rubrobacter xylanophilus DSM 9941, complete genome .
GACGATACGGGCGATGAGGCCCGCCTTCCGGGGGCTCCGCCTCCGGAGAACCGCGCCGGC
GGCGCTGATGGCTTCTGCCCCGCA
>MGVM01000041.1/13941-14016 Elusimicrobia bacterium RIFOXYB2_FULL_48_7 rifoxyb2_full_scaffold_1989, whole genome shotgun sequence.
ATTAGCAACGGAGATGGAGTTCTCCTTTAACCCAACCGCTTTAACTTCATTAAAAGCTGA TAACTCCTACAATTAA
>AYZV02264506.1/1975-2037 Spinacia oleracea cultivar SynViroflay
scaffold87291. con0006.1, whole genome shotgun sequence.
ATAAATAGTGGCAATGATGTCTGCCTTGAACCGCCTTAATAAGCTGATGACGTCTACTTT TAG
>CP010056.1/30164-30238 Hymenobacter sp . DG25B plasmid, complete sequence.
CCACAATCTGGTGATGGGGTACCACCACGAACCGCCGAAGCCGTTCCGCTTTGCGCTGAT
GACTCCTACGTTTAC
>MKSH01000219.1/12690-12752 Solirubrobacterales bacterium 70-9
SCNpilot_bf_inoc_scaffold_4893, whole genome shotgun sequence.
CTGTATCTGAGCGATGAGGCTCGCTCAAACCGTGGCTCAGCCGCTGATGGCCTCTGCCCC CAG
>AE016877.1/4975681-4975742 Bacillus cereus ATCC 14579, complete genome.
ATAATCCTAGGCGATGGAGTTCGCCATAAACGCGCTGCTTATCTTATGACTCCTACCAGT
AT
>MGQO01000116.1/7675-7746 Deltaproteobacteria bacterium RBG_16_47_11 RBG_16_scaffold_21686 , whole genome shotgun sequence.
TATAATCTTGGCGATGAAGTTCGCCATTAAATGCTCCCTCTCTTATCGGGGGCTGATAAC TTCTACTGGTTT
>FXAT01000003.1/185330-185392 Paraburkholderia susongensis strain LMG 29540 genome assembly, contig: Ga0139082_103
AATTGACGTGGAGATGGCATTCCTCCCTAACCGCCCGCAAGGGCTAATGATGCCTACAGG
CCT
>LMPU01000001.1/488329-488269 Pedobacter sp . Leafl94 contig_l, whole genome shotgun sequence.
TTGAACCATGGCGATGGGGTACCGCCAAAACCGCCAACAGGCTGATGACTCCTACGATTT
T
>LGPD01000045.1/73083-73025 Virgibacillus soli strain PL205
scaffold000045, whole genome shotgun sequence.
TTTCGTTTAGGCAATGGAGTTCACCAAAACTGCTGAACGCTAATGACTCCTGCCGAAAA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>LVYD01000065.1/59800-59869 Niastella vici strain DJ57 contig36, whole genome shotgun sequence.
TAATCAATAGGCGATGGAGTTCGCCTTTGAGCGCCACTTTCCACGCGTGGCTAATGACTC
CTGTTTCATG
>BAHC01000126.1/5797-5709 Gordonia rhizosphera NBRC 16068 DNA, contig: GORHZ126.
AAGGCCCGCGGTGATGGATCTCGCCGGGGTGAGCGTCACCGACGTGGTCGCCCGAACCGC
CGACATCGGCTGATAGTTCCTGTTCCTGG
>APJS01000002.1/101420-101479 Clostridium tetanomorphum DSM 665 contig2, whole genome shotgun sequence .
GAAATCTGTGGTGATGGAGTTCGCCATTAAATGCGTAAAGCTAATGACTCCTACAAAATA
>LMOZ01000003.1/513183-513109 Leifsonia sp . Leaf336 contig_3, whole genome shotgun sequence.
TGATCCAGCGGCGATGGATCTCGCCGGGGCGTGAACGCCCGAACCGCGACGGTCGCTGAT
AGTTCCTGCGCTTCA
>FR887608.1/8393-8318 Firmicutes bacterium CAG:41 genomic scaffold, scf25
TTAAATAAAGGGAATGAGGTTCTCCCTTAGTATAATACTAAAACCGCTTATTTAAGCTGA
TGACTTCTGCACTATT
>CP006570.1/280332-280398 Sodalis praecaptivus strain HS1 plasmid pHSl, complete sequence.
ATTTCATAAGGTGATGGCGTTCCACCTTACCTTAACCGCCCTCAGGGCTGATGACGCCTG
TCATAAC
>CP013293.1/1082699-1082632 Chryseobacterium sp . IHB B 17019, complete genome .
GCAATACAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTTCAGAAGCTGATGACGCCT
GATTAAGA
>MUNT01000010.1/156311-156209 Rhodanobacter sp . C03
NODE_6_length_218089_cov_16.1895_ID_11, whole genome shotgun sequence. GACCACTCGGGAGATGGCATGCCTCCCGTTTCGCGGAAGCTTTCGGCATCCCGGAACAAA CCACCTCTCGCCCGGCGTTGAGGTTGATGATGCCTGCACCACC
>LWHJ01000030.1/14380-14301 Pedobacter sp. CCM 8644
NODE_6_length_330457_cov_122.052_ID_11, whole genome shotgun sequence.
ATCAAAAAAGGTGATGGTGTACCACCTATTTAAACCGCCCTCAAGTATATTGTTTGAGGT
ATGATGACACCTACTTATAA
>JH417874.1/1731346-1731275 Yokenella regensburgei ATCC 43003 genomic scaffold Scfldl57, whole genome shotgun sequence.
ACTCTACAAGGTGATGGTGTTCCACCTTTCCCAACCGCCGGCATTCAGCCGGATGATGAC GCCTGATATACA
>CAKP01000082.1/165666-165724 Caloramator australicus RC3, WGS project CAKP01000000 data, contig: 2
GAAATTATAGGTGATGGGGTTCACCTTAATTGCCTTGTGCTGATAACTCCTACTTTTGT
>MIC001000116.1/2059-2127 Syntrophus sp . RIFOXYC2_FULL_54_9
rifoxyc2_full_scaffold_15789, whole genome shotgun sequence.
TTCATTGATGGGGATGGAGTCCCCCGCTGTAAAACGCTTCTTAACAGAGCTGATGACTCC TGCTGCACG
>CU468230.2/2879674-2879608 Acinetobacter baumannii str. SDF, complete genome .
TTAAAACAGGGAGATGGCATTCCTCCCTTGAAAAACCGCCGTATTGGCTAATGATGCCTA
CGTTACC
>MRCG01000023.1/67227-67144 Phormidium tenue NIES-30 NIES-30_Scaffold_23 , whole genome shotgun sequence .
CATAGAAAAGGCGATGGAGCTCGCCAAAACCGCCCTACGTGAGTCAGATTTGACCGCCTA
AGGGCTAATGGCTCCTACTCATTC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>CP009706.1/2399395-2399325 Hafnia alvei FBI, complete genome.
TTGCGGCGAGGTGATGGCGTTCCACCTTACCCAACCGCTTCCCCTGAGAAGCTGATGACG
CCTGATATTAC
>CP011144.1/3480073-3480013 Pseudoxanthomonas suwonensis strain Jl, complete genome.
TGGGCCAACGGGAATGGGGCTTCCCGCGAACCGCCTTCGGGCTGATAGCTCCTACACACC
G
>AP013058.1/774434-774318 Burkholderia sp. RPE64 DNA, chromosome 1, complete genome.
ATCGCTCACGGAGATGGCATGCCTCCCGCAGATCATCGCAAACCGCTGGATTGATTACGT AGCTGACTTACCCGATTGAGTCGCCTAATTAAGCCAGCTAATGATGCCTACGCGTTC
>LJOD01000002.1/294502-294435 Chryseobacterium indologenes strain CI_885 contig_2, whole genome shotgun sequence.
ATCAAAAAAGGAAATGGTGCTCTTCCTTACCCAACCGCTTTACAAAAGCTGATGACGCCT
GATTAAGA
>AGIZ01000005.1/353585-353669 Fischerella sp . JSC-11 ctgll2, whole genome shotgun sequence.
AATAAAAAAGGCGATGGAACTCGCCATAATCGCCTTCTGGTGATATTTAATACTTAATCA
AAAGGCTGATGGCTCCTACTTCTCC
>FR890317.1/35643-35571 Dorea formicigenerans CAG:28 genomic scaffold, scf145
TCTCAGACAGGGAATGAGGTTCTCCCTCAATTACGATTGAAACCGCCAAATGGCTGATGA
CTTCTGTGAATCG
>FRAC01000031.1/82127-82051 Anaerocolumna jejuensis DSM 15929 genome assembly, contig: EJ3 lDRAFT_scaffoldO 0026.26
TTGCAGAGTGGGAATGAAGTTCTCCCCTGGTAGACAATACCTAAACTGCTTATAAAGCTA
ATGACTTCTGTTTAAAT
>FQXR01000016.1/22863-22805 Sporanaerobacter acetigenes DSM 13106 genome assembly, contig: EJ96DRAFT_scaffold00015.15
ATAAATAATGGTGATGGAGTTCACCAAAACTGCATTATGCTAATGACTCCTACAGGTAA
>CP009621.1/3032854-3032790 Pontibacter korlensis strain X14-1T, complete genome .
GCTACAAACGGCAATGATGTCTGCCTTGAACCGCTCTACCCAGAGCTGATGACTTCTACT
TTAAG
>CP003325.1/1273195-1273267 Bifidobacterium asteroides PRL2011, complete genome .
TGAGCCTCGGGCGATGGAACCCGCCCGGGGCTTGAGCCCCGAACCGCAAACCGCTGACAG
TTCCTGCACCATG
>HF570958.1/2643127-2643057 Tetrasphaera japonica T1-X7 genomic scaffold, 1540_scaffoldl
GATAGGGCGGGTGATGGAACCCACCCGGGGCGACGCTCCGAACTGCGACCGCTGATGGCT
CCTACGACGAC
>MHEH01000185.1/2976-3044 Nitrospirae bacterium RBG_13_43_8
RBG_13_scaffold_6510, whole genome shotgun sequence.
TAAAAATAAAGCGATGGAGTTCGCTTAACCGTTCCGATAGATATCGGGACTGATAACTCC TACCAGGAT
>KI535313.1/618130-618057 Enterococcus cecorum DSM 20682 = ATCC 43198 genomic scaffold adfme-supercont2.1 , whole genome shotgun sequence.
CCAAATATAGGGAATGAGGTTCTCCCTTGGTTTTTACCTAAACTGCTATTACAGCTGATG ACTTCTATGATTTT
>FP929061.1/1879515-1879588 Anaerostipes hadrus draft genome.
AAAGAAAAAGGGAATGAAGTTCTCCCTCGAAGAGATTCGAAACCGCTTATTAAGCTGATG
ACTTCTGTGCGATG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>MGTH01000159.1/17659-17500 Desulfobacca sp . RBG_16_60_12
RBG_16_scaffold_4360, whole genome shotgun sequence.
AGGAAGAATGGCGATGGAGCTCGCCTGGGAAACTTTTCATACTGAGTTGCATTCCAAGAG TTTTTGATGGCCGCCGCAGGCTTTAGCCGGCGCCGGATGTCCGATTGCGAGTCGGTATCA TGGCCCGAACCGCGCTGACGCTGATGGCTCCTGATCTCAC
>AJJU01000010.1/56840-56902 Imtechella halotolerans K1 ContiglO, whole genome shotgun sequence.
CCTTGTTATGGCAATGATGTCTGCCTTGAACCGCCTTATAAAGCTGATGACGTCTATAAA
GAG
>AE016853.1/5215709-5215637 Pseudomonas syringae pv . tomato str. DC3000, complete genome.
CGGCGCATTGGAGATGGCATTCCTCCATTAACAAACCGCTGCGCCCGTAGCAGCTGATGA
TGCCTACAGAAAC
>MGQW01000036.1/1664-1597 Deltaproteobacteria bacterium RBG_16_64_85 RBG_16_scaffold_20158, whole genome shotgun sequence.
AACATAGCGGGCGATGGAGTTCGCCGAAACCGCCTGTCCGCAAGACGGCTGATAACTCCT GCGAAGAA
>GG753627.1/966677-966757 Streptomyces sp. el4 genomic scaffold
supercontl .2 , whole genome shotgun sequence.
CTCGACCGTGGCGATGGATCCCGCCTGGACCGCCTCCGGTCCGAACCGCCCCGATCCCGG
GCTGATGGTTCCTGACCATTG
>CP018811.1/3024504-3024565 Paraburkholderia sp . SOS3 chromosome 1, complete sequence.
TTGCGCACCGGAGATGGCATTCTCCTTTAACCGCCCTTGTGGCTGATGATGCCTGCTTCG
TC
>LBIC01000020.1/54156-54226 Sphingobium chungbukense strain DJ77
contig00020, whole genome shotgun sequence.
GGCATGGACGGCGATGGATTTCCGCCTGGCTTCGGCCGAACCGCCTCCGGGCTGATGATT
CCTACCTGCTG
>LSYZ01000275.1/8929-8845 Leptolyngbya valderiana BDU 20041 contig00275, whole genome shotgun sequence .
TTCAAAACGGGCGATGGAGCTCGCCCCAACCGCCTGCAAACCGCCGCAAACGGCAGAGTT
TAAGGCTGATGGCTCCTACACTTCC
>FQUM01000002.1/657912-657850 Mariniphaga anaerophila strain DSM 26910 genome assembly, contig: Ga0131162_102
ATTTGCATAGGCAATGAGGTCTGCCATTAACTGCCGTAAAAAGCTGATGACTTCTACCAG
CAA
>LGI001000031.1/91888-91812 Burkholderia sp . ST111 contig_31, whole genome shotgun sequence.
TGGCGCGCTGGAGATGGCATTCTCCCGATGGCATTCTCACTTAACCGCCCTCGATGGCTG
ATGATGCCTGCTTCGCC
>FMJL01000016.1/7643-7570 Clostridium sp . N3C genome assembly, contig: N3C_contigl6
ATCATCAAGGGGAATGAAGTTCTCCCTTGGTTTATACCTAAACCGCATTTAACGCTAATG
ACTTCTGTTTTCTT
>MEEB01000351.1/1777-1717 Hyphomicrobium sp . SCN 65-11 ABS54_C0351, whole genome shotgun sequence.
TCGTGGGATGGGGATGGGGTCCCCCGACAACTGCTGAAAGGCTGATGGCTCCTGCCAGAC
C
>FWXF01000009.1/42569-42507 Desulfacinum hydrothermale DSM 13146 genome assembly, contig: EJ40DRAFT_scaffold00009.9
TGCGATCTAGGCGATGAGGCTCGCCTTGGAACCGCCCTTCGGGCTGATGGCCTCTAGGGA
AGC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>CP007547.1/929790-929723 Elizabethkingia anophelis NUHP1, complete genome .
ACAAAAAGAGGAAATGGTGTTCTTCCTCACCCAACCGCTTCTGAGGAGCTGATGACGCCT
GATTAAAT
>FUWH01000001.1/578761-578824 Sediminibacterium ginsengisoli strain DSM 22335 genome assembly, contig: Ga0070520_101
AGTAATACAGGAAATGGTGTCTTCCTGACCCAACCGCTACAGAGCTGATGACGCCTGCAA
ATTG
>MUNV01000026.1/22600-22497 Rhodanobacter sp . B04
NODE_2_length_497009_cov_12.5251_ID_3, whole genome shotgun sequence.
TGGGCAACGGGAGATGGCATGCCTCCCGTTTCGCGAGCGCCTGCGGCATTCCCGGAACAA
ACCACCTCTCGTCCGTCGTTGAGGTTGATGATGCCTGCACCACC
>LNTU01000001.1/355243-355164 Paramesorhizobium deserti strain A-3-E Scaffoldl, whole genome shotgun sequence.
AATCGATGCGGCAATGGATTCTGCCGGGCCGGAAATGGCCGAACCGCTTCCCCGATGAAG
CTGATGACTCCTGCTCATTT
>MVHR01000007.1/97693-97632 Mycobacterium heidelbergense strain DSM 44471 NODE_7_length_175488_cov_52.3162, whole genome shotgun sequence.
ACTGACAGCGGCGATGGAGCCCGCCCTGAACCGCCGATCCGGCTGATGGCCCCTGCGAAC AT
>AOTI010596493.1/2825-2758 Triticum urartu cultivar G1812 contig596494, whole genome shotgun sequence .
AGATAAAAAGGAAATGGTGCTCTTCCTTAACCAACCGCATTATAAATGCTGATGACGCCT
GATTAATT
>AJWX01000005.1/215217-215149 Pseudomonas sp . M47T1 contig05, whole genome shotgun sequence.
TTCGTGCCCGGAGATGGCATTCCTCCCCGAACCGCCGCGCTGGCCGCAGCTGATGATGCC
TACAGACAC
>CP002028.1/582315-582385 Thermincola sp . JR, complete genome.
TGTAAAAAAGGCGATGGGGCTCGCCATAATTGCTGATACATAATTGCTCAGCTGATAGCT
CCTACCGGTGA
>FWXI01000013.1/79855-79919 Sporomusa malonica strain DSM 5090 genome assembly, contig: GaO 070592_113
CAATCTATAGGCGATGGAGTTCGCCATTAAATCCGTTATTATGCGGTGATGACTCCTACT
AGTAA
>JH370371.1/1073259-1073190 Alistipes indistinctus YIT 12060 genomic scaffold supercont 1.1 , whole genome shotgun sequence.
GGAAGTACAGGCAATGGTATCTGCCTTGAACCGCCGCCCGGCCGGGTGCGCTGATGATAC CTACAGGGCT
>LMKA01000101.1/104821-104749 Chryseobacterium sp. Leaf201 contig_6, whole genome shotgun sequence.
GCAACACAGGGAAATGGTGTTCTTCCTTACCCAACCGCCCTGTTTCAACGGGACTGATGA
CGCCTGATTAAGA
>CYSP01000014.1/95346-95408 Propionispora sp . 2/2-37 isolate 2/2-37 genome assembly, contig: 2/2_contigl4
TTCAATAACGGCGATGGAGTTCGCCCTGAACCCCGTGAAAACGGTAATGACTCCTACCAC
GTA
>FN806773.1/1882150-1882077 Propionibacterium freudenreichii subsp.
shermanii CIRM-BIAl, complete genome
CCTGAAAACGGCGATGGGACCCGCCGGGGCATAAGAGCCCGAACCGCAGACCTGCTGATA
GTTCCTGTTCGTTG
>JFZT01000020.1/7527-7607 Candidatus Acidianus copahuensis strain ALE1 ctg7180000000532, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AGTGAGACGGGGGATGGCGTCCCCCTGAGGGTTTTCCCTCAAACTGCCTAAACAATGAGG
GCTGATGACGCCTATCTCTAC
>MAST01000004.1/48439-48366 Humibacillus sp . DSM 29435 contigl2, whole genome shotgun sequence.
TGAGGCTCAGGTGATGGCGCCCGCCCGGAGCGTGAGCTTCGAACCGCAGCCCTGCTGATG
GCCCCTGTTGATGG
>CP002770.1/2632433-2632494 Desulfotomaculum kuznetsovii DSM 6115, complete genome.
CTGCTCCACGGCGATGGAGCTCGCCTGGAATGCCCGGAAAGGCTGATGGCTCCTACCAAA
AT
>LYOS01000002.1/254544-254613 Candidatus Syntrophoarchaeum caldarius isolate BOX2 BOX2_8_reassembly_contig02 , whole genome shotgun sequence. GATGATAGTGGTGATGAGGTCCACCTTGGTTTATGCCAAAACCGCTACAGCTAATGACTT CTAAATAAAA
>CP000254.1/1238338-1238398 Methanospirillum hungatei JF-1, complete genome .
TGTCATACTGGTGATGGAGTTCACCAAAACCGCCGTTCGGGCTGATAACTCCTACAGGAT
G
>AUPA01000193.1/6290-6348 Clostridium sp . BL8 Contig44, whole genome shotgun sequence.
AAGTTAAAAGGTGATGGAGTTCACCATAACCGCAATATGCTAATGACTCCTACAGGTAA
>MVHR01000005.1/139483-139543 Mycobacterium heidelbergense strain DSM 44471 NODE_5_length_243907_cov_56.1658, whole genome shotgun sequence. CCCTGACAAGGCGATGAAGCTCGCCTTAACCGCCGAACCGGCTGATGGCTTCTACCCGTG
G
>LBIC01000001.1/127094-127024 Sphingobium chungbukense strain DJ77 contigOOOOl, whole genome shotgun sequence.
CGCCGGCAAGGCGATGGATTTCCGCCGGGCTTCGGCCGAACCGCCTTCGGGCTGATGATT
CCTACCTGATG
>CP001990.1/86509-86568 Bacillus megaterium QM B1551 plasmid pBM700, complete sequence.
AATATAAATGGCGATGGAGTTCGCCAAAATTGTCGTTGTGCTAATGACTCCTACCCTTTT
>GG666045.1/116305-116378 Anaerococcus lactolyticus ATCC 51172 genomic scaffold SCAFFOLD2, whole genome shotgun sequence.
ATATTTTTAGGGAATGAAGTACTCCCTTGATAATTATAATCTAAAACGCGTTGGCTGATG ACTTCTACATTTTT
>FWXY01000034.1/22182-22243 Desulfobacterium vacuolatum DSM 3385 genome assembly, contig: Ga0002809_134
AACTGCTCAGGGAATGGGGTTTCCCTTTAACCGCTTGTAAAGCTGATAACTCCTGGTTGA
GA
>CP020705.1/1280264-1280187 Ruminococcaceae bacterium CPB6, complete genome .
ATAGAATATGGGAATGAAGTTCTCCCATGGTGAAAAACCGGAACTGCTTTCAAAAGAGCT
GATGACTTCTGCGGTTTT
>KB291615.1/28458-28530 Clostridium celatum DSM 1785 genomic scaffold Scfld22, whole genome shotgun sequence.
TTAAACATAGGGAATGAAGTTCTCCCTTGATATCTATCAAAATCGCTAATAAGCTAATGA
CTTCTACAACCTT
>KI391947.1/598628-598554 Ruminococcaceae bacterium D16 genomic scaffold acsPy-supercont2.1 , whole genome shotgun sequence.
CGCAATAAAGGGAATGAGGTCTCCCTTGTCCTTCGGGACTAAACCGCCTACAAGGCTGAT GGCTTCTGCGATTTT
>HF951689.1/2549689-2549768 Chthonomonas calidirosea T49 complete genome APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ATGGATAACGGTGATGGAGGCCGCCGAGACGTTTTGACGTCTGAACCGCCGCTTTGTAGG
CTAATGGCTCCTGGTCGGAT
>CP000943.1/4458865-4458926 Methylobacterium sp . 4-46, complete genome.
TGGCGGTCTGGGGGTGGAGTTCCCCATAACCGCCCACCCGGGCCGATGACTCCTACCGCG
TG
>MHXP01000105.1/6401-6339 Planctomycetes bacterium GWA2_39_15
gwa2_scaffold_29798, whole genome shotgun sequence.
TTTATTTCAGGCGATGGAGTTCGCCCAACTGCTACGCTTGTAGCTGATAACTCCTATTGA AGC
>FRAG01000049.1/6608-6535 Clostridium caminithermale DSM 15212 genome assembly, contig: EJ27DRAFT_scaffoldO 0049.49
TTTATAGTAGGGAATGAAGTTCTCCCTGGAATGATGTTCCAAACCGCTGATGAGCTAATG
ACTTCTATGCATAT
>LMDP01000005.1/166756-166693 Afipia sp . Rootl23D2 contig_5, whole genome shotgun sequence.
GTTACAGATGGGGATGGAGTCCCCCGATACCCGCCCGTCGTGGGCTGATGACTCCTGCTG
GGAT
>LELH02000025.1/36824-36912 Massilia sp . WF1 contig25, whole genome shotgun sequence.
ATAATCAGCGGAGATGGCATTCCTCCCCGTATCGCTCGACACCGTTCGACACCAAACCGC
CCTTCGGGGCTGATGATGCCTGCACGTTC
>LTBB01000007.1/122122-122181 Clostridium colicanis DSM 13634
CLCO_contig000007, whole genome shotgun sequence.
AATATTAGTGGCGATGGAGTTCGCCTTTAAATGCGCAAAGCTGATGACTCCTACAAAATG
>CP011602.1/127305-127374 Kluyvera intermedia strain CAV1151, complete genome .
GATCAGCAAGGTGATGGCGTTCCACCTTACCCAACCGCTTCCCTGTGAAGCTGATGACGC
CTGGTATGAC
>LAQT01000037.1/32472-32546 Amantichitinum ursilacus strain IGB-41 contig37, whole genome shotgun sequence.
CGCCCTGCAGGAGATGGCATTCCTCCCAGAACCGCCGCGCCGCCCTCCGGTGCAGCTGAT
GATGCCTACGTAACC
>MJGC01000053.1/54710-54642 Desertifilum sp . IPPAS B-1220
NODE_2_length_240057_cov_33.7793, whole genome shotgun sequence.
ATTTGTAAAGGTAATGGAGCTTACCTTAAACCGCCAAGCTTTAGCTCGGCTGATGGCTCC TACTATTCT
>KE356576.1/ 907962-907898 Halorubrum sp . J07HR59 genomic scaffold scf7180000001382, whole genome shotgun sequence.
GTCTGTTCAGGCGATGGAGTCCGCCTGTCCCAACCGCCATTCTGGCTGATGACTCCTACT GAACA
>MRTP01000006.1/212158-212218 Paenibacillus rhizosphaerae strain FSL R5- 0378 NODE_6_length_352458_cov_l .76218_ID_3980, whole genome shotgun sequence .
AACAACATAGGCGATGGAGTTCGCCTTTAACCGCCCGTGAGCTAATGACTCCTACCAGTG
C
>MNEG01000108.1/20597-20528 Candidatus Rokubacteria bacterium
13_1_40CM_68_15 13_l_40cm_scaffold_771, whole genome shotgun sequence. GACGAAAAAGGCGATGGGGTTCGCCTGAAACCGCCTCACCGACCGTGAGGCTGATGACCC CTACCGGCAG
>LT629737.1/2606979-2606919 Gillisia sp . Hell_33_143 genome assembly, chromosome: I
ACATACAAAGGCGATGGAGTTCGCCAAAACCGCCCAAAAAGCTAATGACTCCTGCTCAAT
T APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>FZOC01000008.1/42168-42104 Desulfovibrio mexicanus strain DSM 13116 genome assembly, contig: Ga0070557_108
CTCCAACAAGGCGAAGGAGTTCGCCTCTAACCGCTTCCGCTGAAGCTGATGACTCCTGTC
CAGGC
>LILC01000023.1/4351-4292 Bacillus koreensis strain DSM 16467 scaffold3, whole genome shotgun sequence .
AATACATACAGCGATGGAGTTCGCTATAAGTGCCTTCGGGCTGATGACTCCTACTCATTC
>LMSH01000002.1/1244479-1244566 Arthrobacter sp . Soil763 contig_2, whole genome shotgun sequence.
TCTGTTCACGGTGATGGATCCCGCCGGAACGCCTTGCACCCGCTGCGGACGTTCGAACCG
CCTCCCGGCTGATGGTTCCTACCTTTGC
>AZDA01000072.1/1632-1696 Lactobacillus bifermentans DSM 20003 NODE_106, whole genome shotgun sequence .
AACAGAACAGGCGATGACGTTCGCCTTCGAAAACCACATTTGTGATTGATGACGTCTACT
TTTTG
>MHYK01000142.1/160-225 Planctomycetes bacterium RBG_16_55_9
RBG_16_scaffold_24606, whole genome shotgun sequence.
ATAAGACTTGGGGATGGAGTCCCCCTTTATAATCGCCTTAGCCGGGCTGATGACTCCTAC CGCGCA
>LMOH01000001.1/41121-41184 Aureimonas sp. Leaf324 contig_l, whole genome shotgun sequence.
CATCGCTACGGGGATGGGGTTCCCCCGCAACCGCCCGCGTCGGGCTGATGACTCCTATCG
AAAC
>MDFL01000126.1/40994-40932 Mesorhizobium sp . ORS3428 contig_4, whole genome shotgun sequence.
CCAAAAAATGGGGATGGGGTTCCCCCGAAACCGCCCTTAAGGGCTGATGACTCCTGCCAG
GCG
>MDJS01000005.1/13684-13749 Mucilaginibacter sp . PPCGB 2223
NODE_5_length_405491_cov_l 65.226_ID_646, whole genome shotgun sequence. TACTAAGGGGGTGATGGTGTACCACCTATTTGAACCGCTGCAAAAGCTGATGACGCCTGC CAAATA
>CP000004.1/152206-152146 Gluconobacter oxydans 621H plasmid pGOXl, complete sequence.
CCTGGACAAGGGGATGGGGTTCCCCTGAAACCGCCGCAAGGCTGATGACTCCTCTGTGCC
G
>BBRZ01000042.1/20346-20250 Vibrio sp . JCM 19231 DNA, contig00042.
TATGGCGAAGGTGATGGGGTTCCACCTAATTGAACCGCCGAATTGTTGAGACGCTCAGCG
TCAAACAACCCTCTCGGCTGATGACTCCTACTAAAAT
>FP929003.1/4028302-4028377 Candidatus Nitrospira defluvii chromosome, complete genome.
AAATCCTCAGGTGATGGAGTTCACCGAGAACCGCCGGTCGCGATCCTCGCCTTCGGCTGA
TAACTCCTACCAGCGC
>CP000884.1/3908908-3908971 Delftia acidovorans SPH-1, complete genome.
CTGCACGATGGGAATGGAGCTTCCCGAAACCTCCCTGCGAAGGGCTGATGGCTCCTGCCT
TGGT
>LLXU01000106.1/25546-25613 Stenotrophomonas panacihumi strain JCM 16536 contig_67, whole genome shotgun sequence.
CCGCGCACAGGAGATGGCATTCCTCCTCGAACCGCCCGCACGCGTGGGCTGATGATGCCT
GCCAAGCC
>MGTB01000013.1/8703-8637 Deltaproteobacteria bacterium
RIFOXYD12_FULL_53_23 rifoxyd3_full_scaffold_1091, whole genome shotgun sequence .
ATACATGACGGCGATGGAGTTCGCCTTTAAACGCCTTGCCCGCAAGGCTGATGACTCCTA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CTCTTGA
>JMSOO 1000176.1/25457-25385 Bacteria symbiont BFol of Frankliniella occidentalis strain BFol contigl66, whole genome shotgun sequence.
TCGCTATAAGGTGATGGTGTTCCACCTTTCCCAACCGCCTCGTCCGTAAGGGGCTGATGA CGCCTGATAACCC
>LT629690.1/2302770-2302839 Terriglobus roseus strain GAS232 genome assembly, chromosome: I
CAACCAAAAGGCGATGGGGTTCGCCTTGATGCCCAGCCCACGGGCTGGAGCTGATGACTC
CTACGACAGT
>CP007035.1/698836-698908 Niabella soli DSM 19437, complete genome.
TATTAAACAGGAAATGGTGTCTTCCTGAACCAAACCGTTCCCACTGGCGGGATCTGATGG
CGCCTGCAAATAG
>AEXE01000084.1/9167-9104 Burkholderia sp. TJI49 contig072, whole genome shotgun sequence.
TTGGCGTCAGGAGATGGCATTCTCCTTAACCGCCGAGACACCGGATGATGATGCCTACTG
CGTT
>CP014304.1/4021829-4021901 Hymenobacter sp . PAMC26628, complete genome.
CCACCCAACGGTGATGGATTTCCACCGCGAACCGCCGGTGCTTGTGCGCCGCGCTGATGA
TTCCTACTGAACT
>FQUW01000005.1/6749-6812 Desulfotomaculum australicum DSM 11792 genome assembly, contig: EJ60DRAFT_scaffold00002.2
AACACACCGGGCGATGGAGCTCGCCTTTAAGCGCCTCTTTCGGGCTAATAGCTCCTACCA
GAGT
>AM412317.1/3416950-3416891 Clostridium botulinum A str. ATCC 3502 complete genome
AGAATTTTGGGCGATGGAGTTCGGCATTAAATGCGTGAAGCTAATGACTCCTACAAATAG
>MHFY01000089.1/39817-39898 Omnitrophica WOR_2 bacterium GWA2_47_8 gwa2_scaffold_2699, whole genome shotgun sequence.
TTAAATTTAGGCGATGGAGTTCGCCGGATTTTTTATAAAACCGCTAACATATCTTTTTAA AGCTGATAACTCCTACGGTTCT
>ALVC01000281.1/2711-2632 Sphingomonas sp. LH128 Contig281, whole genome shotgun sequence.
TCGGTGCCAGGCAATGGATTTCTGCCTGTCCCTCTCTTGAGGGCCGAACCGCCATTTTGG
CTGATGATTCCTACCTTTCG
>LZKJ01000166.1/72737-72798 Mycobacterium kyorinense strain E861 contig_66, whole genome shotgun sequence.
ATCCCGCTAGGCGATGAAGCTCGCCTTAACCGCCGCACCTGGCTGATGGCTTCTACCCGT
GA
>CP001825.1/675546-675485 Thermobaculum terrenum ATCC BAA-798 chromosome 1, complete sequence.
CTGAATACAGGCGATGAAGCTCGCCTTAAGCGCCGCACAAGGCTGATAGCTTCTACCAAG
AT
>MFYY01000268.1/813-874 Candidatus Riflebacteria bacterium GWC2_50_8 gwc2_scaffold_9095 , whole genome shotgun sequence.
CTCATCACAGGGGATGAAGTCCCCCGAAAGCGCCAAACCAGGCTAATGACTTCTACCAGA TT
>FR901374.1/15762-15832 Ruminococcus sp . CAG:353 genomic scaffold, scf223
CTTAACAAAGGGAATGAAGTGCGCCCTTGGGAAACCAAAACCGCTTATTAGCTGATGACT
TCTGCGATATT
>CBXV010000007.1/151298-151365 Pyrinomonas methylaliphatogenes , K22, WGS project CBXV01000000 data, contig: PYK22_K22_C11
TCGCGAAAAGGCGATGGAGTTCGCCTAAGAAGCGTCCAAGAATTTGGACTGATGACTCCT GCATCGTG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>MNRI01000145.1/1526-1448 Clostridiales bacterium 52_15
Ley3_66761_scaffold_18972 , whole genome shotgun sequence.
AGAATAAACGGGAATGAAGTTCTCCCGAAGTAACCAGATTACTTGAACCGCTTTGAAAGC TGATGACTTCTGCGACGAA
>MBTF01000023.1/290123-290059 Mucilaginibacter pedocola strain TBZ30 scaffold3, whole genome shotgun sequence.
TTCACGAAGGGCGATGGTGCTCCGCCTATTTGAACCGCTGCAAGGCTGATGGCTCCCACT
CAAAT
>LWMN01000013.1/126318-126382 Enterococcus thailandicus strain F0711D 46 Ent-16_S35_L001_001_5, whole genome shotgun sequence.
AAGTGAATAGGTGATGGTGTTCGCCTTTAAACGTTTATCAATAAACTAATGACGCCTACT AAAAA
>FR898162.1/4839-4916 Clostridium sp . CAG:288 genomic scaffold, scf4
ATATTTAATGGAAATGATGTCTTCCAAGTTTTACTATTAAACAAAACCGCTTATTTAGCT
AATGACGTCTATGATTTT
>MHZQ01000303.1/3179-3076 Rhodanobacter sp . RIF0XYA1_FULL_67_6
rifoxyal_full_scaffold_25782, whole genome shotgun sequence.
TGGCCCCTCGGAGATGGCATGCCTCCCGTTCCGGAACCCTCGGGCACACCGGAACAAACC GCTTCTCGTGATCCCGCGTTGAAGTTGATGATGCCTGCACCCCG
>AWXR01000011.1/52946-52878 Rhodonellum psychrophilum GCM71 = DSM 17998 strain GCM71 EFN26G1Q, whole genome shotgun sequence.
ATTTCATTCGGCGATGGGGTTCCGCCTTACAACCGCTCCTTTTTTGGTGCTGATGACTCC TACTTCAAC
>CP017603.1/1723046-1723109 Clostridium formicaceticum strain ATCC 27076, complete genome.
TAATTAAATGGCGATGGAGTTCGCCCTTAACCACTGAAATTCAGTTGATGACTCCTGTAA
AACA
>FQZ001000013.1/31190-31132 Clostridium amylolyticum strain DSM 21864 genome assembly, contig: Ga0131114_113
TTTTAAATAGGTGATGGAGTTCACCATAATTGCGGAATGCTGATGACTCCTACAGGTAT
>MKVX01000018.1/234308-234368 Rhizobiales bacterium 62-47
SCNpilot_cont_300_bf_scaffold_6, whole genome shotgun sequence.
AATGTTGATGGGGATGGGGTCCCCCGACAACCGCCGCGAGGCTGATGACTCCTACCGGGC
G
>MGQR01000021.1/7235-7165 Deltaproteobacteria bacterium RBG_16_50_11 RBG_16_scaffold_125085, whole genome shotgun sequence.
TGATCTCTTGGCGATGAGGTTCGCCCTCAACCGTCCTATATTTAAGAAGGACTAATAACT TCTACTGGATT
>BDL001000001.1/1393809-1393867 Coriobacteriaceae bacterium EMTCatBl DNA, scaffoldl .
TGGAGCACAGGCGATGGAGCTCGCCTTGAACCGCTTCGGCTGATGGCTCCTGCACCGGA
>MNRC01000153.1/2783-2856 Clostridiales bacterium 36_14
Ley3_66761_scaffold_17690 , whole genome shotgun sequence.
AAAGAAAAAGGGAATGAAGTTCTCCCTCGAAGGGATTCGAAACTGCTTATTAAGCTGATG ACTTCTGTGTAATG
>HG326223.1/463488-463426 Serratia marcescens subsp . marcescens Dbll, complete genome
CGGCACGCTGGAGATGACATTCCTCCCTAACCGCCTTCACAGGCTGATGATGTCTACGTA
ACC
>CP013140.1/2956134-2956056 Lysobacter enzymogenes strain C3 genome.
ATGACCGGCGGAGATGGCATTCCTCCGGTCCTGTCGCGACAGGCGAACCGCCGCCAAGGC
TGATGATGCCTGCCCACCC
>MNXU01000046.1/15457-15519 Betaproteobacteria bacterium CG2_30_68_42 APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
cg2_3.0_scaffold_2887_c, whole genome shotgun sequence.
CTGCCCGCAGGAGATGGCAGTCCTCCTCCAACCGCCCGCGAGGCTGATGATGCCTGCAAT
GAC
>FQTV01000020.1/20780-20718 Bacteroides luti strain DSM 26991 genome assembly, contig: Ga0131163_120
GCCTGTTTTGGCAATGGCATCTGCCTTAAACCGCTCAATTGAGCTGATGATACCTACTTA
AAG
>MELC01000038.1/206115-206185 Acidobacteria bacterium
RIFCSPL0W02_12_FULL_66_21 rifcsplowo2_12_scaffold_259, whole genome shotgun sequence.
TGGGCCAACGGCGGTGGAGTTCGCCGAAACCGTCCCCGCTCATTGGGGGGACTGATGACT
CCTACTCGAGG
>LGCL01000013.1/9108-9168 Ornatilinea apprima strain P3M-1 contig_26, whole genome shotgun sequence .
TTTTTTATTGGTGATGGAGTCCACCGTTAATTGCTTTATTGCTGATGACTCCTGGATTTT
T
>MKTH01000014.1/142601-142527 Cellulomonas sp. 73-92
SCNpilot_cont_500_bf_scaffold_247 , whole genome shotgun sequence.
GTCGCCGTCGGCGATGGATCCCGCCCGCGTCGATGACGACGGAACCGCCGGCCGGCTGAT GGTTCCTGCTCCCAG
>CP002736.1/343658-343715 Desulfotomaculum nigrificans CO-l-SRB
chromosome, complete genome.
AAATTGAACGGCGATGGAGTTCGCCTAACGCTGCCAAGCTGATGACTCCTACCTTTGT
>MNUQ01000167.1/18676-18615 Desulfovibrionaceae bacterium CG1_02_65_16 cgl_0.2_scaffold_l 191_c, whole genome shotgun sequence.
TCCAGCAAAGGGGATGGAGTCCCCCTTGAACCGCGACAGCCGCTGATGACTCCTGCTGCC
CG
>LLEU01000028.1/8512-8440 Acidobacteria bacterium OLB17
UZ17_ACD001CONTIG000028, whole genome shotgun sequence.
CCAATTCCGGGAAATGGTGTCTTCCCGAACCGAACCGCGAGCCTGACGCGCCGCTGATGA CGCCTGTTTTTAG
>FRAF01000044.1/2027-1966 Alicyclobacillus sp. USBA-503 genome assembly, contig: Ga0105849_144
TGATGGTAGGGTGATGGGGCTCACCCAAACTTCCCAGAAGGGATAATGGCTCCTGCGAAT
GG
>CP000859.1/3573766-3573832 Desulfococcus oleovorans Hxd3, complete genome .
AGAAACGATGGGGATGGGGTCCCCCGACAACCGCCGCACTGGACAGGCTGATGACTCCTG
CCGATAC
>AL590842.1/2918233-2918170 Yersinia pestis C092 complete genome
TCAAGATCTGGAGATGACACGCCTCCATAACCGCCCTAACAAGGCTAATGATGTCTACGT
GACC
>CP001463.1/457388-457463 Thermococcus sibiricus MM 739, complete genome.
TTGCCGTTGGGCGATGGCGTCCGCCCGGGGCTTCGGGCCCGAACCGCCCACAAGGGCTGA
TGACGCCTGTTCTCCG
>LAZH01000020.1/53148-53209 Bacilli bacterium VT-13-104 contig020, whole genome shotgun sequence.
TTCTAAATAGACGATGGAGTTCGCCATTAACCGTCTTTTCGACTAATCACTCCTGCCAGT
TT
>LPWG01000013.1/305197-305139 Methyloceanibacter methanicus strain R-67174 contig_8, whole genome shotgun sequence.
CAATTCGAAGGGGATGGAGTCCCCCTGAACCGCCCTTGGCTGATGACTCCTGCCGTGTC
>CM001165.1/4044875-4044798 Streptomyces sp . Tu6071 chromosome, whole APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
genome shotgun sequence.
CTCGCCACAGGCGATGAATCCCGCCGGGGCCCCTTACGGCCCGAACCGCCCTCCCGGGCT
GATGGCTTCTGACACCCA
>LWHG01000001.1/82066-82130 Haladaptatus sp . R4
NODE_l 0_length_139766_cov_26.9624_ID_1043 , whole genome shotgun sequence.
GCTATCGCGGGAGATGGGGTCCTCCTTCAACCGCCACGAAACTGGCTGACGATCCCTGCG
AAGGT
>AJXS01000164.1/4499-4588 Rhodanobacter sp . 115 contigl64, whole genome shotgun sequence.
GGCGGGAAAGGAGATGGCATGCCTCCGGTTCCACCTCGGAACAAACCGCCCCGATGTCAT
CAAACAGGGGCTGATGATGCCTGCACCACC
>AOIV01000036.1/89716-89647 Halogeometricum pallidum JCM 14848 contig_36, whole genome shotgun sequence .
GATACTCGTGGCGATGGAGTCCGCCTTACCCCAACCGCTGGACCCGCCAGCTGATGACTC
CTGTTTCCAC
>NGLL01000001.1/2165996-2166068 Enterococcus sp . 2G9_DIV0600
scatfoldOOOOl, whole genome shotgun sequence.
ATGCTAATCGGGAATGAATGTCTCCCTCAGTTCTTACTGAAACCGCAATTTTGCTAATGA
CTTCTACCACTTT
>FP 929048.1/1688235-1688299 Megamonas hypermegale ART12/1 draft genome.
TAAATAGATGGTGATGGAGTTCACCGATAAATCTGTATAGATACAGTAATGACTCCTGTT ATATT
>AMCX01000055.1/33045-33106 Sinorhizobium fredii GR64 contig00055, whole genome shotgun sequence.
TAGGAGAGTGGGAATGGTGTCTCCCCTTTAACCGCGCTGATGCTGATGACGCCTGCTGAC
GA
>MGVR01000325.1/2434-2373 Elusimicrobia bacterium RIFOXYD12_FULL_66_9 rifoxyd3_full_scaffold_33646, whole genome shotgun sequence.
GGGTCGTACGGCGATGGAGTTCGCCACAACCGCCCGCGGGGGCTGATGACTCCTACCGAA AA
>LJC001000051.1/58159-58221 Alicyclobacillus ferrooxydans strain TC-34 contig_19, whole genome shotgun sequence.
TCCATAGACGGTGATGGTGCTCACCACAACTGCCTCAAGAAGGATAATGGCGCCTGCAAA
TCA
>HF952022.1/57931-57991 Thermobrachium celere DSM 8682 genomic scaffold, scaffold37
ATAAACTAAGGTGATGGAGTTCACCTTTAACCGCCTTTTGGCTAATGACTCCTACAGATA
A
>CP000697.1/351714-351653 Acidiphilium cryptum JF-5, complete genome.
GGCCGCGCCGGCGATGGGGTTCGCCGTGAACCGCCCGCGGGGCTGATGACTCCTGTTTCG
AT
>AE017198.1/1172633-1172571 Lactobacillus johnsonii NCC 533, complete genome .
AATTGAATAGGTGATGACGTTCACCAATTAAACTGAGCAAAATCTAACGACGTCTACTTC
CTA
>LTFS01000131.1/2559-2483 Globicatella sp. HMSC072A10
Globicatella_spHMPREF2811-l .0_Cont338.1, whole genome shotgun sequence.
TTAACGAAAGGGAATGAGGTTCTCCCTAGGTAAGTTTACCTAAACCGCAGTTATTTGCTG
ATGACTTCTGTGACGCT
>LFNF01000003.1/849128-849061 Chryseobacterium sp. BLS98 contig03, whole genome shotgun sequence.
GCAACATAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTACCAAAGCTGATGACGCCT
GATTAACA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>LJC001000051.1/58048-58111 Alicyclobacillus ferrooxydans strain TC-34 contig_19, whole genome shotgun sequence.
AACAAATCAGGTGATGGAGCTCACCGTTAACTTCCTGCAAAAGGATAATGGCTCCTGCGG
ACGG
>MIHE01000032.1/32834-32895 Mycobacterium shimoidei strain HMC_M2
PROKKA_contig000032, whole genome shotgun sequence.
TCGCCGAAAGGCGATGAAGCTCGCCTTAACCGCCGCACCTGGCTGATGGCTTCTACCCGT TG
>JQCA01000090.1/1056-1124 Lactobacillus paucivorans strain DSM 22467 NODE_144, whole genome shotgun sequence.
AAGTTAATCGGCGATGACGTTCGCCACTAAAACTAATTGGTAAATCAATTTGATGACGTC
TACTGCACA
>AE010299.1/4831358-4831427 Methanosarcina acetivorans str. C2A, complete genome .
TGTAACAATGGTGATGAGGTCCACCCTGATAATTGTCAAAACCGCTTATGCTGATGACTT
CTACACAACT
>FRCY01000013.1/68704-68811 Cyclobacterium lianum strain CGMCC 1.6102 genome assembly, contig: Ga0070154_113
CCCTATGATGGTGATGGGGTGCCACCAATGCTTGCAGATTTCCGGGATTAATACTGGTTT
TCACAACACAACCGCCGGGACTTTCTGGCTGATGACTCCTGTTTCAAG
>CP003614.1/5761519-5761607 Oscillatoria nigro-viridis PCC 7112, complete genome .
AGTAATGATGGCGATGGAGCTCGCCGAAACCGCCCTTCGTTATTAATAATCTTCATAACG
TTTAAAAGGCTGATGGCTCCTACTGTTCC
>LMPG01000003.1/12954-12878 Sphingomonas sp . Leaf343 contig_ll, whole genome shotgun sequence.
ATCGACTTTGGCAATGGATTTCTGCCGGGCAATTTGGCCGAACCGCTCTGATAAGAGCTG
ATGATTCCTACTTGGCG
>MHEJ01000073.1/1243-1343 Nitrospirae bacterium RBG_16_43_8
RBG_16_scaffold_45301, whole genome shotgun sequence.
CTTAATAACGGCGATGGGGTTCGCCGTTCATCCCATTAGAGAAAAATTTCTAACGGGATA AACCGTTCCGATGCTTCGGGACTGATGACTCCTACTCTCAA
>CP009279.1/3101747-3101808 Paenibacillus sp . FSL H7-0737, complete genome .
ATATAACAAGGCGATGGAGTTCGCCACAACCGTCTCTTCAGACTAATGACTCCTACCAGA
GT
>LYVF01000013.1/69725-69664 Desulfotomaculum sp . LMal
LMA27_trimmed_contig_12 , whole genome shotgun sequence.
AACGGTTGGGGCGATGGAGCTCGCCATTAACTGCTCTTGCAGCTGATAGCTCCTGTTGGA CG
>HF972668.1/15540-15611 Clostridium sp. CAG:1013 genomic scaffold, scfl98
GAATTTCACGGGAATGAAGTGCTCCCTTGGGAAACCTAAACCGCTTTTTAAGCTGATGAC
TTCTGCGAGACA
>CP019948.1/4131812-4131752 Methylocystis bryophila strain S285, complete genome .
CACCGAAATGGGGATGGAGTCCCCCGAAAGCGCCGCGATGGCTGATGACTCCTACAACCC
C
>MEFU01000001.1/26650-26718 Comamonadaceae bacterium SCN 68-20
ABT02_C0001, whole genome shotgun sequence.
GTGCCCGCAGGAGATGGCGTTCCTCCCCTAACCACCGCGCTGCGCGCGGTTGATGACGCC
TACAGACAC
>FUWX01000010.1/6757-6831 Cetobacterium ceti strain ATCC 700028 genome assembly, contig: EI52DRAFT_scaffold00007.7 APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AATAAAAAAGGGAATGATGATCTCCCTTAGTTATAAACTAAAACTGCTTATAAAGCTGAT
GACGTCTACATTTTT
>MKVZ01000019.1/52844-52904 Rhizobiales bacterium 63-22
SCNpilot_cont_500_bf_scaffold_96, whole genome shotgun sequence.
CCAGTCGTAGGGGATGGAGTTCCCCCTGAACCGCCGTAAGGCTGATGACTCCTGAACGGT
C
>NGMM01000001.1/1091604-1091534 Enterococcus sp . 9E7_DIV0242
scatfoldOOOOl, whole genome shotgun sequence.
GTCTATGAGGGCGATGGTGTTCGCCGATAAATATTTATTTTGAGGAATAAATTGATGACG
CCTACCAAACA
>FRCH01000003.1/261927-261826 Rhodanobacter sp . OK091 genome assembly, contig: Ga0115501_103
GTTGCCAGGGGAGATGGCATGCCTCCCGTTTCGGGCATTCTTCGGCATCCCGTGACAAAC
CACCTCTCGTCCGTCGTTGAGGTTGATGATGCCTGCACCACC
>FRBL01000003.1/389150-389077 Chitinophaga j iangningensis strain DSM 27406 genome assembly, contig: Ga0131168_103
AGTCTACCAGGAAATGGTGTCTTCCTGCTTTGAAACCGTTTCGCATCGCCGAAACTGATG
GCGCCTACAAATAT
>JMPL01000110.1/7822-7883 Kluyvera ascorbata ATCC 33433 GKAS . assembly .110, whole genome shotgun sequence .
TTATGCGTAGGGGATGGCATTCCCCCTTGAACCGCCGTGTGGCTGATGATGCCTGTTTGT
AT
>HE956757.1/2510539-2510478 Methylocystis sp . SC2 complete genome
TTAACGGATGGGGATGGGGTTCCCCCGATAACCGCCGAAAGGCTGATGACTCCTGCGAAC
TC
>JPME01000010.1/371729-371805 [Clostridium] celerecrescens strain 152B clO, whole genome shotgun sequence.
ATAGCTCAAGGGAATGAAGTTCTCCCTGTGCTGATTATAGCCGAACCGCTTATTAAGCTG
ATGACTTCTGCAATAAT
>CP014841.1/1961461-1961522 Dyella thiooxydans strain ATSB10, complete genome .
GGACGCAAGGGAGATGGCATGCCTCCCCGAACCGCCGCAAGGCTGATGATGCCTGGTCGA
AC
>CP001276.1/661070-661147 Thermomicrobium roseum DSM 5159 plasmid, complete sequence.
GTGGTGCACGGCGATGAGGCTCGCCAGGTGGCTTCGGCCACCGAAGTGCCCGCTGCGGCT
GATAGCCTCTGCTCGACC
>LLWP01000002.1/417118-417183 Pedobacter sp . Hvl Contig02, whole genome shotgun sequence.
AGACAAAAAGGAAATGGTGTCTTCCTACTTGAACCGCCTTAAAAAGCTGATGACGCCTGG
TTATTA
>CP004143.1/1260953-1261015 Pseudomonas denitrificans ATCC 13867, complete genome .
TGCGCGGCAGGAGATGGCATTCCTCCTTCAACCGCCCTCGTGGCTGATGATGCCTACGCA
TAA
>HF989520.1/57872-57949 Eubacterium sp. CAG:786 genomic scaffold, scf497
ATATCACCGGGGAATGACGTCCTCCCCGTGGGCATAGCCCCGAACCGCAAATCTTTTGCT
GATGGCATCTGCGATATC
>MICZ01000050.1/72332-72405 Tenericutes bacterium GWF2_57_13
gwf2_scaffold_822 , whole genome shotgun sequence.
ATCGGGTCGGGGAATGAAGTGCTCCCTTCGACATGTTCGAAAACCGCCTGACGGCTGATG ACTTCTACGCGTTT
>CAOS01000009.1/20614-20672 Desulfotomaculum hydrothermale Lam5 = DSM APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
18033, WGS project CAOSOIOOOOOO data, contig: ABX_2590_2
AAAGTGAATGGCGATGGAGTTCGCCTAACGCTGTTCAAGCTGATGACTCCTACCCCTTG
>MGTG01000089.1/3250-3327 Desulfobacca sp. RBG_16_58_9
RBG_16_scaffold_326703_curated, whole genome shotgun sequence.
AGCAATTACGGCGATGGAGTCCGCCGGGGCGGCCAGAATCCGCCCGAACCGCGCTACGCT
GATGGCTCCTGCTCACAC
>FQYX01000026.1/15509-15445 Arenibacter nanhaiticus strain CGMCC 1.8863 genome assembly, contig: GaOO 66783_126
GCTACAATTGGCAATGATGTCTGCCTTAAACCGTTCCTAATGGAGCTAATGACTTCTACT
TTTAA
>CP002690.1/1077454-1077516 Thermodesulfobium narugense DSM 14796, complete genome.
AAAAATATCGGTGATGGAGTTCACCGTAATGCCCCATAGATGGCTGATGACTCCTAGTAT
ATT
>BATA01000039.1/8119-8051 Halarchaeum acidiphilum MH1-52-1 DNA, contig: HALO39.
GACGAGCGCGGCGATGGCGTCCGCCCACCACAACCGCCGGACACGCCGGCTGATGACGCC
TTCTCACCG
>CP015405.2/2570795-2570723 Blautia sp. YL58, complete genome.
TAGTTTTAAGGGAATGAGGTTCTCCCGTGGATGATTCCAAAAGTGCGTGTATGCTGATGA
CTTCTGTGTAAAA
>LDP001000008.1/9003-9064 Mycobacterium heraklionense strain Davo contig_8, whole genome shotgun sequence.
ATCCAGGCAGGCGATGGAGCTCGCCTTTAACCGCCGCACCGGCTGATGGCTTCTACCAGG
GC
>JGZR01000006.1/227799-227872 Bifidobacterium subtile strain LMG 11597 Contig06, whole genome shotgun sequence.
ATGAACAGCGGGTATGAGGTTCACCCTGGGCACACGGCCCGAACCGCCTTCGGGCTGATG
ACTTCTGCAGGGAT
>NGMO01000001.1/ 906763-906696 Enterococcus sp. 10A9_DIV0425 scaffoldOOOOl, whole genome shotgun sequence .
ATGTAAATAGGCGATGGTGTTCGCCATGAAATGTTAATCAGAAGATAACTAATGACACCT
ACTAAAAT
>KQ758490.1/5050-5110 Bacillus enclensis strain SGD-1123 genomic scaffold Scaffold9, whole genome shotgun sequence.
AACCCTACAGGTGATGGAGTTCACCTTTAACCGCCGGAAGGCTAATGACTCCTGTCAGTG
T
>MGDP01000031.1/488-552 Candidatus Tectomicrobia bacterium
RIFCSPL0W02_12_FULL_69_37 rifcsplowo2_12_scaffold_14005, whole genome shotgun sequence.
CTGTTGAACGGGGATGGAGTCCCCCTCACAATCGCCTGTTTTGGGCTGATGACTCCTGTC
GCGAG
>MDCF01000056.1/969-1030 Acidiferrobacter thiooxydans strain ZJ
Contig_149, whole genome shotgun sequence.
GAGCCAAAGGGCGATGGGGTTCGCCGTATAACCGCCTCTGGGCTGATGACTCCTGAATGA
CC
>CP006938.2/344694-344771 Pandoraea pnomenusa strain RB-44, complete genome .
CTGCTGCACGGAGATGGCATACCTCCGGCGTTCCTGCGCGAACCGCTCCTCCCCGGGGCT
GATGATGCCTGCAAGATC
>MKRJ01000017.1/20749-20812 Rhizobiales bacterium 65-79
SCNpilot_cont_1000_bf_scaffold_222 , whole genome shotgun sequence.
TGACATATCGGGGATGGAGTCCCCCCTAAACCGCCGTTGGACGGCTGATGACTCCTGCCA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TGCT
>AFXZ01000044.1/4163-4227 Bizionia argentinensis JUB59 contig00057, whole genome shotgun sequence.
CAAATATTTGGCAATGAGGTCTGCCTTAAACCGTTCTATCTAGAGCTAATGACATCTACT
TTAAT
>JJMM01000002.1/387613-387549 [Clostridium] litorale DSM 5388 strain W6 CLIT_2c, whole genome shotgun sequence.
TCAATATCAGGCGATGGAGCTCGCCTTCAAACATTGCTGTTGCAATTGATGGCTCCTGCT
GGCCT
>CR354531.1/1717779-1717696 Photobacterium profundum SS9
ATTATGACGGGAGATGATGTTCCTCCTTTAACTGCCTTACTGAAGTTTACCCTTCTTAAT
AAGGATGATGACGTCTAACAACAT
>MUYX01000049.1/5769-5706 Bacterium AM6 strain AM6 2017-02-03- Sterotrophomonas_trimmed_contig_49, whole genome shotgun sequence.
GCCGTCGATGGGGATGGGGCTCCCCCGATAACCGCCTGAGATGGCTGATGGCTCCTGCTG AGAC
>CP000254.1/1244505-1244443 Methanospirillum hungatei JF-1, complete genome .
GTAAACCAGGGTGATGGGGTTCACCTGTAACCGCTTTTTCCAGCTGATGACTCCTCTTCC
TGT
>CP001338.1/1058916-1058849 Candidatus Methanosphaerula palustris El-9c, complete genome.
GATGGATGAGGTGATGGGGTCCACCTCAATTAACCGCTCGTGTACGAGCTGATGACTCCT
ACTAATCT
>MGPO01000019.1/8285-8221 Deltaproteobacteria bacterium GWA2_65_63 gwa2_scaffold_16628, whole genome shotgun sequence.
ACGCGTGAAGGGGATGGAGTCCCCCTCATAATCGCCTACGCCCGGCTGATGACTCCTACC GCGCA
>KB850956.1/10839-10900 Clostridium colicanis 209318 genomic scaffold acBRg-supercontl .1 , whole genome shotgun sequence.
AATTTTATAGGTCATGGAGTTGACCTTTAACCGCTTTCAAAGCTGATGACTCCTGTAGAA CT
>FQUW01000005.1/8614-8675 Desulfotomaculum australicum DSM 11792 genome assembly, contig: EJ60DRAFT_scaffold00002.2
CTGCTCCATGGCGATGGAGCTCGCCTGGAATGCCCGGAAAGGCTGATGGCTCCTACCAAA
AT
>MHYT01001093.1/1884-1945 Planctomycetes bacterium RIFCSPHIGH02_12_39_6 rifcsphigho2_12_subl0_scaffold_9112, whole genome shotgun sequence.
TTACGTCCATGCGATGGAGTTCGCCTAACCGCTATAACATAGCTGATAACTCCTATTGAA GC
>LMQP01000002.1/551522-551598 Sphingomonas sp. Leaf412 contig_2, whole genome shotgun sequence.
CTCGGGCACGGCGATGGATTTCCGCCGGGCCATTCGGCCGAACCGCCCTCGCAAGGGCCG
ATGATTCCTACCAGGCG
>MZGW01000001.1/276335-276269 [Clostridium] thermoalcaliphilum strain DSM 7309 CLOTH_contig000001, whole genome shotgun sequence.
AAATATATAGGTGATGGAGTTCACCTTTAACTATTGCAATTTTGCAATTAATGACTCCTA CTTGGAC
>MGPZ01000034.1/12910-12985 Deltaproteobacteria bacterium GWD2_55_8 gwd2_scaffold_1402 , whole genome shotgun sequence.
TTTTGAACAGGTGATGGGGTTCACCTTTTAACCATCCTGGTCTCATTCAGACGGGATTGA TAACCCCTACCAGTGG
>MELC01000038.1/206050-205980 Acidobacteria bacterium APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
RIFCSPL0W02_12_FULL_66_21 rifcsplowo2_12_scaffold_259, whole genome shotgun sequence.
TGTGCGACGGGTGTTGGAGTTCACCCTCAACCGCCCTACCTCTTGACGGGGCAGATGACT
CCTACGGGCTT
>MKSQ01000024.1/22465-22403 Bosea sp . 67-29
SCNpilot_expt_1000_bf_scaffold_559, whole genome shotgun sequence.
TTCATAAATGGTGATGGAGTTCGCCATTTAACCGCCCCCAGGGCTGATGACTCCTGTCGA
GCC
>FNHF01000002.1/161501-161561 Sediminibacillus halophilus strain CGMCC 1.6199 genome assembly, contig: Ga0079860_102
GGATCGATAGGTGATGGAGTTCACCTTTAACCGCCGACAGGCTGATGACTCCTGCCAGTA
A
>CP021112.1/2665108-2665048 Pseudorhodoplanes sinuspersici strain RIPI110, complete genome.
ATCGATGATGGGGATGGAGTCCCCCGAAACTGCCGCAGAGGCTGATGACTCCTATCGCGT
T
>FXAH01000001.1/515764-515826 Paraburkholderia caryophylli strain Ballard 720 genome assembly, contig: Ga0139045_101
CGTGCGACAGGAGATGGCATGCCTCCTCGAACCGCCCTCGTGGCTGATGATGCCTACCTT
GTC
>MKUQ01000060.1/73427-73366 Burkholderiales bacterium 70-64
SCNpilot_expt_1000_bf_scaffold_93 , whole genome shotgun sequence.
TCCGTCATGGGAGATGGCATGCCTCCCCGAACCGCCTGACGGCTGATGATGCCTACCGGC CA
>MGYB01000092.1/204-280 Gammaproteobacteria bacterium
RIFCSPHIGH02_12_FULL_41_15 rifcsphigho2_12_scaffold_6720, whole genome shotgun sequence.
CTTAAATCGGGAGATGGCATTCCTCCTTTAAACCGCTTCTTATCTCACCGATCAGAGCTG
ATGATGCCTACCATGAG
>LCVM01000112.1/16442-16510 Providencia rettgeri strain MR4
P_rettgeri_contig_112 , whole genome shotgun sequence.
AAACCTTTGGGAGATGGCATTCCTCCTTATATAAAACCGCCCGTAGAGGCTGATGATGCC TACGTTAAC
>CAHT01000061.1/20583-20678 Methylacidiphilum fumariolicum SolV, WGS project CAHT01000000 data, contig: 77-2114
TAATAATCTGGCAATGGAGTTTGCCAGCCCGCGTTTTCGAAGTGTTTTGGGCAAACCCCT
TTTCTTTTAGGAAAAGGCGATAACTCCTACTTCAAA
>CP012024.1/1739208-1739270 Bacillus smithii strain DSM 4216, complete genome .
AGCTGAATAGGCGATGGAGTTCGCCGTATAACCGTCCATGCGACTAATGACTCCTACCAG
TAA
>AM889285.1/3856660-3856597 Gluconacetobacter diazotrophicus PA1 5 complete genome
CGTCCGGTCGGAGATGGAGTACCTCCGTATAACCGCCCCCAGGGCTGATGACTCCTGCTC
GCAT
>LGFR01000059.1/2947-3011 Clostridia bacterium 62_21 S_scaffold_859, whole genome shotgun sequence.
CGTACTGACGGCGATGGAACCCGCCTAACGCCGCCGAACGAACGGCTGATGGTTCCTACC
ATTTC
>CP009313.1/721578-721515 Streptomyces nodosus strain ATCC 14899 genome.
AGGGGAGTGGGCGATGAGGCTCGCCCTCGACCGCACATCCCGTGCTGATGGCCTCTGGGA
GCGA
>AE016822.1/68491-68420 Leifsonia xyli subsp. xyli str. CTCB07, complete APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
genome .
TGGCTCTGCGGCGATGGATCTCGCCGGGGCGCTGACCGAACCGTGATGGCCGCTGATAGT
TCCTGCGCTTCT
>AOMC01000135.1/53685-53618 Halococcus morrhuae DSM 1307 contig_135, whole genome shotgun sequence.
CGCCCATCGGGCGATGGGGTCCGCCTTTCACAACCGCTGGGCAGCCAGCTGATGACTCCT
CCCGGAGA
>HF998441.1/92082-92166 Coprococcus eutactus CAG:665 genomic scaffold, scf195
ATTATAAAAGGGAATGAGGTTCTCCCTTGGAAAACCTGAATGGGCAACCTAAACTGCTAA
TATAGCTGATGACTTCTACGATTAT
>CP000650.1/81552-81622 Klebsiella pneumoniae subsp . pneumoniae MGH 78578 plasmid pKPN5, complete sequence.
TTGCGGCGAGGTGATGGCGTTCCACCTTACCCAACCGCTTCCCCTGAGAAGCTGATGACG
CCTGATATTAC
>HF987873.1/11923-11999 Clostridium sp. CAG:590 genomic scaffold, scf386
TTATAGGAAGGGAATGAAGTTCTCCCTAAGTAGTGAATACTTGAACCGCTTATTAAGCTG
ATGACTTCTGTGTTAAG
>LJYF01000012.1/170404-170480 Bradyrhizobium yuanmingense strain BR3267 contigl019, whole genome shotgun sequence.
GTCGCAGATGGGGATGGAGTCCCCCGACAACCGCCCGGCGATCGATGACCTCATGGGCTG
ATGACTCCTGCTCGAGG
>CP012398.1/3843722-3843662 Chelatococcus sp . CO-6, complete genome.
GAGACCGCAGGGGATGGGGTCCCCCGTAACCGCCGCAGAGGCTGATGACTCCTACCGGGC
G
>MNZR01000273.1/3832-3748 Syntrophobacteraceae bacterium CG2_30_61_12 cg2_3.0_scaffold_3560_c, whole genome shotgun sequence.
ATTGCGCTAGGCAATGAAGTCTGCCTGAACCGCCCTCCTTTTTGTCTGCCTGCAAGCAGC AAGGGCTGATGACTTCTACCGTTTC
>JH417874.1/2154575-2154645 Yokenella regensburgei ATCC 43003 genomic scaffold Scfldl57, whole genome shotgun sequence.
AAGGGTTAAGGTGATGGCGTTCCACCTTCCCCAACCGCCCACAATGATGGGCTGATGACG CCTGGTATATC
>MHXP01000105.1/4565-4504 Planctomycetes bacterium GWA2_39_15
gwa2_scaffold_29798, whole genome shotgun sequence.
GCTAAAAAGGGCGATGGAGTTCGCCATAATTGCTATGTATAGCTGATAACTCCTATTGAA GC
>CP003065.1/3427437-3427379 [Clostridium] clariflavum DSM 19732
chromosome, complete genome.
TGAATATATGGTGATGGAGTTCACCTGAATTGCTTAATGCTGATGACTCCTACAGGAAT
>MKTE01000006.1/13414-13350 Bacteroidia bacterium 44-10
SCNpilot_bf_inoc_scaffold_1083, whole genome shotgun sequence.
ATTTGTAAAGGCTATGGGATTAGCCGTTAACCGCCCGATAAAGAGCTGATGATCCCTACT ATACC
>CELZ01000050.1/62441-62505 Moorella glycerini strain NMP genome assembly, contig : M_glycerini_NMP_DRAFT_scaffold-50
TCTCCAAACGGTGATGGAGTTCACCGGAAAAACCACCCGTAAGGGTTGATGACTCCTACC
AAGGA
>MGQO01000116.1/8078-8149 Deltaproteobacteria bacterium RBG_16_47_11 RBG_16_scaffold_21686 , whole genome shotgun sequence.
ATCAGTCAAGGCGATGAGGTTCGCCTTTAACCGTCCTTTTATTGAGAAAGGATTGATAAC TTCTACCGGGTC
>CP001130.1/516815-516883 Hydrogenobaculum sp. Y04AAS1, complete genome. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TAACCACAAGGAGATGGCGTTCTCCTTTAAGCGCCTCCTTCAAAAGAGGCTGATAACGCC
TACAAGCTC
>BANB01000284.1/4316-4377 Acidisphaera rubrifaciens HS-AP3 DNA, contig: Asru_0284.
TGCCGCGATGGGGATGGGGTCCCCCGAAACCGCCCGTCCGGGCTGATGACTCCTGCGCTT
TT
>FUYH01000004.1/17531-17469 Caloramator quimbayensis strain USBA 833 genome assembly, contig: Ga0105845_104
AAGTAAATAGGTGATGAAGTTCGCCATAATTATCTTTTTTAGATTGATAACTTCTGCAGT
ATT
>MGPV01000033.1/5026-4958 Deltaproteobacteria bacterium GWC2_65_14 gwc2_scaffold_2584 , whole genome shotgun sequence.
GAAGATCGGGGTGATGGGGTTCACCCGGAACCGTCCGCCGGGAGGCGGGCTGATGACTCC TGCGATGGC
>CP009880.2/1007548-1007475 Pantoea sp. PSNIH1, complete genome.
CTCAGCACGGGTGATGGCGTACCACCCGATCCAACCGCCGCTGTTTTCAGACGGCTGATG
ACGCCTGACTATAC
>CP006850.1/5425559-5425647 Nocardia nova SH22a, complete genome.
CGGCGGTCAGGCGATGAAGCTCCGCCTGACCGGTGCCTCCGGACACCGGTGCCCCCGTGC
CGAACGGCACTGATGGCTTCTACCCACAG
>CBXV010000007.1/141761-141688 Pyrinomonas methylaliphatogenes , K22, WGS project CBXV01000000 data, contig: PYK22_K22_C11
AAAATTCGGGGCGATGGAGTTCGCCGTATAACCATCTCGGTCGCCAGATCGAGATTGATG ACTCCTTTTCAAGC
>CP000517.1/1061364-1061422 Lactobacillus helveticus DPC 4571, complete genome .
AACTAAATAGGTGATGACGTTCACCATTAACCGAGAGATCTAATGACGGCTACTTTATC
>KB849435.1/1926175-1926109 Acinetobacter sp . CIP 102637 genomic scaffold acLrO-supercontl .3 , whole genome shotgun sequence.
TAGGATAAAGGAGATGGCATTCTTCCTGTTACAAACCGCCATAGTGGCTAATGATGCCTA
CGTGACC
>MTKC01000003.1/5832327-5832263 Streptomyces griseofuscus strain NG1-7 NG1-17_1_1, whole genome shotgun sequence.
CGACGCGCCGGTGATGGGGCTCACCGCAACCGCGGCACCATGCCGCTGACGGTCCCTGGT
CGAAC
>JH591188.1/96457-96526 Dialister succinatiphilus YIT 11850 genomic scaffold supercont 1.2 , whole genome shotgun sequence.
CCCGCAAACGGCGATGGAGTTCGCCGGGCGTAGGCCGAATGCTCTTTGAGCTGATGACTC CTGTCCCCCT
>KB849539.1/187621-187554 Acinetobacter gerneri DSM 14967 = CIP 107464 genomic scaffold acLZs-supercontl .32 , whole genome shotgun sequence.
TAAATTAAAGGAAATGGCTATCTTCCTTCTACAAACCGCCACTAACGGCTGATGATACCT ACGTTACC
>AP017295.1/143881-143970 Nostoc sp . NIES-3756 DNA, complete genome.
TAAATACAAGGCGATGGAGCTCGCCACAACCGCCTTCTGATTTTTACATACTTAATACTC
GATCAAAAGGCTAATGGCTCCTACTTTTCC
>MNCN01000036.1/12631-12706 Verrucomicrobia bacterium 13_2_20CM_55_10 13_2_20cm_scaffold_1270, whole genome shotgun sequence.
TCCGGCAACGGCAATGGAGTTTGCCTTCCTAGGTAAAACTGGGAAAACCGCGTGAGCTAA TGACTCCTTCCAAGCG
>AE016825.1/2613318-2613254 Chromobacterium violaceum ATCC 12472, complete genome .
TTCTATTGCGGAGATGGCATTCCTCCCGTAACCACCCCTAGAGGGTTGATGATGCCTACG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
GTATT
>JZRC01000050.1/8689-8590 Aquitalea magnusonii strain SM6 contig_50, whole genome shotgun sequence.
TTTTCTCCGGGAGATGGCATTCCTCCCGCCTGTTGGCTGCTTCCTGCCCGGAAGTCGCTG
CAGCGAACCGCCCGTCCGGGCTGATGATGCCTGCAGGTAT
>AP014924.1/2829973-2830047 Limnochorda pilosa DNA, complete genome, strain: HC45.
CGCTGCCCCGGCGATGGAGCTCGCCGTTGCCCTCAGGCAAGGTCCCGGGATCCGGGTGAT
GGCTCCTACCAGCAC
>MEDI01000308.1/12732-12794 Kaistia sp. SCN 65-12 ABS35_C0308, whole genome shotgun sequence.
GGCGGCGATGGGGATGGAGTTCCCCCGAAACCGCCCGCGAGGGCTGATGACTCCTACCGC
AAC
>JZSG01001632.1/861-791 Raoultella planticola strain GCSL-DIFS-295
CFSAN029441_contigl631, whole genome shotgun sequence.
TTGCGGCGAGGTGATGGCGTTCCACCTTACCCAACCGCTTCCCCTGAGAAGCTGATGACG CCTGATATTAC
>NATN01000018.1/11848-11909 Candidate division KSB1 bacterium 4484_87 ex4484_87_scaffold_1483, whole genome shotgun sequence.
ACTGATATGGGCGATGGAGTCCGCCTGAGACTGCTGTCAGAGCTAATGACTCCTTTTTGT TA
>CP002584.1/186490-186555 Sphingobacterium sp. 21, complete genome.
TTGCGAAACGGAAATGATGCCTTCCGACTGAACCGCCTTAAAAAAGCTGATGGCGTCTAC
TAAAAT
>MVBK01000001.1/109763-109702 Thioalkalivibrio denitrificans strain ALJD Ga0073317_1001, whole genome shotgun sequence.
ACAGAGCCGGGAGATGGCATCCCTCCCAAACCGCCCTCGCGGCTGATGATGCCTACACTG CT
>CP002467.1/957934-957861 Terriglobus saanensis SP1PR4, complete genome.
AACAACCTGGGCGATGGGGTTCCGCCACTAACCGCACGCGCTGTATGCGCGGAGCTGATG
ACTCCTACGAGACA
>MKVZ01000012.1/36693-36755 Rhizobiales bacterium 63-22
SCNpilot_cont_500_bf_scaffold_404 , whole genome shotgun sequence.
CATAATCAAGGAGATGGAGCTCTCCTGGTAACCGCCTTGTCGGCTGATGGCTCCTGCTGT TGA
>LDOT01000023.1/122968-122877 Photobacterium aquae strain CGMCC 1.12159 contig0023, whole genome shotgun sequence.
CGCTTGGCGGGAGATGATGTTCCTCCTTTAACCGCCTTTCTTGATCGGCCTTGTGCCATT
TGCTCAAAAAGGATAATGACGTCTAGCAACAC
>CP000100.1/2263916-2263989 Synechococcus elongatus PCC 7942, complete genome .
TACAAGCAGGGTGATGGAGCTCACCTCAACCGCGCACTCGGCATGGCCCGATCGCTGATG
GCTCCTACAGTTCT
>AWXH01000003.1/755831-755900 Serratia sp. ATCC 39006 ctg3, whole genome shotgun sequence.
CCTATAAAAGGTAATGGCGTACTACCTTCCCCAACCGCTATAGAATATAGCTGATGACGC
CTGGCTTACG
>CP001104.1/1954632-1954704 [Eubacterium] eligens ATCC 27750 chromosome, complete genome.
AATATCAAAGGGAATGAAGTACTCCCTTGGGAAACCTAAACTGCTTATTCAAGCTGATGA
CTTCTACGATTTT
>LT828648.1/1869705-1869769 Nitrospira japonica isolate Genome sequencing of Nitrospira japonica strain NJ11 genome assembly, chromosome: I APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ATGTCTACAGGGGATGGGGTCCCCCGTATAACCGCCACGGCCAGGCTGATGACTCCTACG
TCGTA
>CP020477.1/2370032-2370113 Acidianus manzaensis strain YN-25, complete genome .
TTATCTTATGGGAATGGCGTCTCCCAGAGGGTAACTCCCTCAAACCGCCAGATTAATGAT
GGCTGATGACGCCTATCTCTAT
>LVYD01000057.1/27458-27385 Niastella vici strain DJ57 contig29, whole genome shotgun sequence.
CATCGCAAAGGAAATGGTGTCTTCCTATTTGAACCGTCCCGATTCAACTCGCGACTGATG
GCGCCTACAAATAG
>AHIE01000030.1/174482-174406 Pantoea stewartii subsp. stewartii DC283 WGS_Sequence30 , whole genome shotgun sequence.
CCTGCGCAAGGTGATGGCGTTCCACCTTCCCCAACCGCCAGCAGGTTTCCCCTGCGGCTG ATGACGCCTGACACTAT
>MIEK01000018.1/16817-16746 Enterococcus rivorum strain LMG 25899 25, whole genome shotgun sequence .
TAATCTGAGGGCGATGGTGTTCGCCGATAACTATTTATTGAATAACAATAGATTAATGAC
GCCTATTGAAGG
>FPIZ01000003.1/253946-253872 Chitinophaga sancti strain DSM 784 genome assembly, contig: LX65DRAFT_scaffold00003.3
TCAGAACCAGGAAATGGTGTCTTCCTGCACAAGAACTGTCCCGTATTCTTCGGGTCTGAT
GGCGCCTACAAATAG
>AJYW02000195.1/134-200 Vibrio genomosp. F6 str. FF-238 FF-238_contig_43, whole genome shotgun sequence .
TCTAGCTCAGGTGATGGATTTCCACCTTTAACCGCTCAATTAATGAGATAATGATTCCTA
CTATAAC
>CP003552.1/6414284-6414366 Nostoc sp . PCC 7524, complete genome.
ATAATTAATGGCGATGGAGCTCGCCAAAACCGCTATTCCAGTGAAAAGATTGCATCTGGA
AGGCTGATGGCTCCTACTTTCCT
>METF01000325.1/19416-19314 Candidate division KSB1 bacterium RBG_16_48_16 RBG_16_scaffold_9849, whole genome shotgun sequence.
CATTCTTTTGGCGATGGAGTTCGCCGCAACTATCTATACCGACTTTTCTCAGCAGAGGTA CTCTCTGCACTAGTCGGAGAAGATTGATGACTCCTATTTTTAA
>MHYC01000027.1/2940-3013 Planctomycetes bacterium RBG_13_44_8b
RBG_13_scaffold_16883, whole genome shotgun sequence.
AGATAAGGCGGTGATGAGGTTCACCAAATAACCGCCCTTAAAATAATAATAGGGATGATA ACCTCTACTGACAG
>FR881878.1/30693-30617 Clostridium sp. CAG:230 genomic scaffold, scf242
GATAAAAAAGGGAATGAAGTTCTCCCATAGTAATGATACTAAAACCGCTTATATAAGCTG
ATGACTTCTGCATTTTT
>CP010777.1/1638732-1638801 Rufibacter sp. DG31D, complete genome.
TTTGGCCCGGGTGATGGGGTACCACCAATTCAACCGCCTCTTTGCAAAGGCTGATGACTC
CTACATGGGT
>AJLS01000043.1/12129-12069 Bacillus bataviensis LMG 21833 contig43, whole genome shotgun sequence.
TAAAAAATTGACGATGGAGTTCGCCATAACCGTCCTAGTGATTAATGACTCCTACCAGTG
G
>CP013111.1/1761328-1761398 Bordetella sp. N genome.
TCACGCCCTGGAGATGGTGTTCCTCCATTAACCACCGCGCGCTAAACGCGGTTGATGACG
CCTACAGAAAC
>MGPV01000057.1/12406-12338 Deltaproteobacteria bacterium GWC2_65_14 gwc2_scaffold_6009, whole genome shotgun sequence.
CCGAACCAGGGCGATGGAGTTCGCCCGGAACCGCCTGTCGAACGGACGGCTGATGACTCC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TGCTTGCGC
>CP003590.1/2938298-2938383 Pleurocapsa sp . PCC 7327, complete genome.
AGAGACGACGGCGATGGAGCTCGCCGAAACCGCCCAGCGTTTTGAGCATTATCACAACGT
ATTTGGCTGATGGCTCCTACTGTTCC
>CP000924.1/673881-673940 Thermoanaerobacter pseudethanolicus ATCC 33223, complete genome.
TTAAATAACGGTGATGAAGCTCACCTTATTGCCGTGAAGGCTAATGGCTTCTACGCGCTT
>LXWX01000002.1/30496-30558 Methylosinus sp . 3S-1 3SlcontiglO, whole genome shotgun sequence.
TCCGCGCCAGGGGATGGGGTGCCCCCTCTAACCGCCGAAGAGGCTGATGACTCCTGTCGG
ATT
>CP021358.1/2600117-2600204 Kushneria marisflavi strain SW32, complete genome .
TTTTCCAAAGGAGATGGCATGCCTCCCGAAGCATGTGCTTCGAACCATCGCGTTCTGACA
CGTCGTGATTGATGATGCCTGCAACGCC
>CP002042.1/2809184-2809267 Meiothermus silvanus DSM 9946, complete genome .
TGAACTCGAGGCGATGGAGCTCGCCATAACCGCCTTCCAGAAGGGCCGCTGACACCCGCT
GAGGCTGATAGCTCCTACCCAAGG
>JWHR01000142.1/307-380 Terrisporobacter sp . 08-306576 Contigl42, whole genome shotgun sequence.
TAAAAACACGGGAATGAAGTTCTCCCAAAGTATTTACTTAAACCGCATTTTATGCTGATG
ACTTCTGTGATTAT
>CP001720.1/2876469-2876412 Desulfotomaculum acetoxidans DSM 771, complete genome .
AAAATGTATGGCGATGGAGCTCGCCTAATGCTGTTAAGCTGATGGCTCCTACCAGGAC
>FUXB01000012.1/96728-96655 Vibrio cincinnatiensis DSM 19608 genome assembly, contig: BS67DRAFT_scaffold00012.12
CTCAATTATGGTGATGATGTTCCACCAAGCACGTTGTGCTAAACCGCTTTAAAGCTGATG
ACGTCTGTATATAT
>JRUN01000024.1/7367-7308 Bacillus ginsengihumi strain M2.11 NODE_24, whole genome shotgun sequence .
CATAATTGCGGCGATGGAGTTCGCCTTTAACTGCTTTAAGCTGATGACTCCTACCTTTAA
>GL397071.1/ 909736-909677 Peptoniphilus duerdenii ATCC BAA-1640 genomic scaffold SCAFFOLD1, whole genome shotgun sequence.
AATATTATAGGGAATGAAGTACTCCCTTTAAACGCGAAAGCTGATGACTTCTACTTAAAT
>CP003167.1/709714-709778 Methanoregula formicica SMSP chromosome, complete genome.
TTATCATCTGGTGATGAGGTCCACCATCAATCGCCCGGATCCGGGCTGATGACTTCTTGT
TTTGA
>ATMF01000017.1/70283-70212 Thermoplasmatales archaeon A-plasma
AMDU1_APLC00017 , whole genome shotgun sequence.
TGTAGTCCGTGGAATGGCGTTTCCACTAATTCAAACCGCCGGCTGAAGCTGGCTGATGAC GCCTGCTTGAAA
>MERD01000069.1/5163-5245 Betaproteobacteria bacterium
RIFCSPLOWO2_02_FULL_67_26 rifcsplowo2_02_scaffold_299, whole genome shotgun sequence.
GATCCCAAGGGAGATGGCATTCCTCCACGGAGTCTCTGACGGAGACTCTCAACCGCCGCA
AGGCTGATGATGCCTACGTAGCT
>MEFA01000045.1/7729-7662 Pseudonocardia sp . SCN 72-86 ABS81_C0045, whole genome shotgun sequence.
GCCCACTACGGCGATGGGGCTCGCCGTCGACCGCCCGGCTCGCCCGGGCTGATAGCTCCT
ACCAACCG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>CM000920.1/1662506-1662424 Gluconacetobacter hansenii ATCC 23769
chromosome, whole genome shotgun sequence.
CTGTCACCCGGCAATGGACTCTGCCAGGTGCAGCCTTCACGGCGAACCGAACCGCCCCTT
GGGCTGATGATTCCTACCTCGCC
>JGZ001000023.1/39670-39741 Bifidobacterium scardovii strain LMG 21589 Contig23, whole genome shotgun sequence.
TTGGGATGCGGCGATGGGACCCGCCTGAAACTTGTTTTCGAACCGCATCACGCTGATAGT
TCCTGCACACGA
>HF995145.1/6392-6317 Ruminococcus torques CAG:61 genomic scaffold, scf53
TTTGACTATGGGAATGAAGTTCTCCCTTAGTGATGAAACTAAAACCGCTGATTTGGCTGA
TGACTTCTGCATTTTG
>CP002049.1/1696664-1696726 Truepera radiovictrix DSM 17093, complete genome .
AACGGTCGTGGCGATGGGTCCCGCCCTCAACCGCCCCTCGAGGCTGATGGTCCCTACGTC
GTT
>JRMZ01000009.1/53591-53662 Clostridiales bacterium S5-A11 contig009, whole genome shotgun sequence .
AGTATGAAAGGGAATGAAGTTCTCCCTTGGGTAACCTAAACCGTTTATTAAGCTGATGAC
TTCTACAATTAC
>FR898699.1/5701-5629 Clostridium sp . CAG:307 genomic scaffold, scf76
TAATAACCAGGAAATGATGTTCTTCCTTGTTTTATAACAAAACCGCAATTATGCTGATGA
CGTCTATAGTTTT
>ALXI01000102.1/18824-18896 Clostridium sp . Maddingley MBC34-26
contig_203_1 , whole genome shotgun sequence.
ATAAATATTGGGAATGAAGTACTCCCGTGGTTTATACCTAAAACGCTAACAAGCTAATGA
CTTCTGTGATGTT
>CM001633.1/2375952-2375855 Oscillatoriales cyanobacterium JSC-12
chromosome, whole genome shotgun sequence.
CTGGACTATGGCAATGGAGCTTGCCACAACCGCCATTGGATCTTGACATGCCCTCAACTG
GGTTGCCAACGTCGATGGATGATGGCTCCTACTTTCCC
>CP003255.1/3702362-3702421 Thermobacillus composti KWC4, complete genome. GATGAAATAGGCGATGGAGTTCGCCATAACTGCCGAAAGGCTGATGACTCCTACCCGTGT
>ACZ001000002.1/5807-5880 Vibrio metschnikovii CIP 69.14 VIB . Contigl39, whole genome shotgun sequence .
ATGCGGCTAGGTGATGATGTTCCACCTAGCAGTTATGCTAAACCGCTTTTCAAGCTGATG
ACGTCTGTACATGC
>MIHC01000025.1/12748-12672 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
ACTGACATCGGCGATGGAGTCCGCCGGAAAGCTTTGCTTTCGAACCGCCGCATAAGGCTG ATGGCGCCTGCGAACCG
>MHEJ01000036.1/13287-13356 Nitrospirae bacterium RBG_16_43_8
RBG_16_scaffold_21261 , whole genome shotgun sequence.
TTTAGTATCGGCGATGGAGTTCGCCATTAACCGCCCCGATCTTATCGGGACTGATGACTC CTACTCTCAG
>LGSS01000006.1/104041-104099 [Clostridium] purinilyticum strain DSM 1384 CLPU_6c, whole genome shotgun sequence.
CTTAATAATGGTGATGGAGTTCACCAGAATCGCTGAATGCTAATGACTCCTACAAGAAA
>NJIH01000001.1/175757-175692 Candidimonas nitroreducens strain SC-089 Scaffoldl_l, whole genome shotgun sequence.
GCGCGCAACGGAGATGGCATTCCTCCTTGAACCGCCGGTACGCCGGCTGATGATGCCTAC
AGGTTT
>KK036467.1/145049-144985 Lactobacillus fabifermentans T30PCM01 genomic scaffold scaffoldl3, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AAATTAAATGGCGATGACGTTCGCCGCAAATAATTGTTGATCAATTTGATGACGTCTACT
ACAAA
>LFJJ01000312.1/124-200 Candidatus Burkholderia verschuerenii strain UZHbot4 BVERctg319, whole genome shotgun sequence.
TTGTCGATAGGAGATGGCATTCTCCTGAGGGTCTCCCGCTAACCGCTCCTCGATAAGCCG ATGATGCCTACCCACAC
>MEDI01000593.1/8694-8631 Kaistia sp . SCN 65-12 ABS35_C0593, whole genome shotgun sequence.
CAGCGGGATGGTGATGGGCTCCACCTTCAACCGCCTGTTCTGGGCTGATGATCCCTACGC
AACA
>JHUR01000048.1/115196-115134 Geobacillus sp . CAMR12739 Contig48, whole genome shotgun sequence.
AATCGAATAGGCGATGGAGTTCGCCATAACCGCCGGCTTCCGGCTGATGACTCCTGCTGC
GTG
>CP004078.1/511381-511441 Paenibacillus sabinae T27, complete genome.
TTCATACAGGGCGATGGAGTTCGCCTTTAACCGCCGCAAGGCTGATGACTCCTACCAGAG
G
>CH672395.1/2226-2162 Leeuwenhoekiella blandensis MED217 scf_1099517004314 genomic scaffold, whole genome shotgun sequence.
TCAATCACAGGCAATGAGGTCTGCCTCAAACCGCTTCTTAGCAAGCTGATGACTTCTACT TAACC
>CP002433.1/1819207-1819139 Pantoea sp. At-9b, complete genome.
CGCTAACAAGGTGATGGCGTTCCACCTTACCCAACCGCCCCATTAAGGGCTGATGACGCC
TGATATGAC
>JRFF01000026.1/33643-33716 Mollicutes bacterium HR2 KQ78_contig000026, whole genome shotgun sequence .
ATAATATTTGGGAATGATGTTCTCCCTTAGTTAATTCTAAAACCGCATTTAGTGCTGATG
ACTTCTACCTATAT
>LKEU01000030.1/118781-118850 Acetobacterium wieringae strain DSM 1911 ACWI_contig000030, whole genome shotgun sequence.
TATAAAACAGGTGATGAAGTTCGCCCTGGCAACCGCCAAAATGCGTTAAGCTGATGACTT CTGTTAAAAT
>MKSH01000238.1/6276-6342 Solirubrobacterales bacterium 70-9
SCNpilot_bf_inoc_scaffold_5980, whole genome shotgun sequence.
CGCCGTGGTGGCGATGAGGCTCGCCGTAACCGCCGTGAAGCTCGCGGCTGATGGCCTCTA CGGAAGT
>AE004091.2 /4914637-4914575 Pseudomonas aeruginosa PAOl, complete genome.
TTGCCGACAGGAGATGGCATTCCTCCTTCAACCGCCCCTGGGGCTGATGATGCCTACGCA
TGA
>LSFY01000001.1/775110-775044 [Clostridium] paradoxum JW-YL-7 = DSM 7308 strain JW-YL-7 ctgl, whole genome shotgun sequence.
ATTAGTTTAGGTGATGGAGTTCGCCTTTAACCATTGCGTCTTTGCAATTAATGACTCCTA CTTAGAA
>LFJI01000454.1/683-606 Candidatus Burkholderia brachyanthoides strain UZHbot7 BRCHctg457, whole genome shotgun sequence.
TTGTGGATAGAAAATGGCATTCTCCTGTGGTTCTCCCGCTTAACCGCTCCTCGATGAGCC GATGATGCCTACACACAC
>MKRP01000098.1/535-469 Acinetobacter sp . 39-4
SCNpilot_BF_INOC_scaffold_12059, whole genome shotgun sequence.
TAGAGAAAAGGAGATGGCATTCCTCCTATTGCAAACCGCCGTTTTGGCTAATGATGCCTA
CGTTACC
>LMGJ01000004.1/88871-88809 Mesorhizobium sp . Rootl57 contig_12, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AGCCGCGATGGGGATGGAGTTCCCCCGAAACTGCCCGCAAGGGCTGATGACTCCTACCGC
AAC
>FRBH01000007.1/124440-124508 Chishuiella changwenlii strain DSM 27989 genome assembly, contig: Ga0131172_107
AACAAAATAGGAAATGGTGTTCTTCCTAAACCAAACCGCTTTATAATAGCTGATGACGCC TGATTATTA
>LMNW01000031.1/376710-376781 Duganella sp . Leafl26 contig_4, whole genome shotgun sequence.
CTGAGCATGGGAGATGGCATTCCTCCTGCAACTATCCCAACCGCCGCTCCGGCTGATGAT
GCCTGCTGATTC
>MKTE01000074.1/12152-12083 Bacteroidia bacterium 44-10
SCNpilot_bf_inoc_scaffold_454 , whole genome shotgun sequence.
GATCATAAAGGGAATGGGACTTCCCTGTTATAAACCGCTACCATACATAGCTGATAGTTC CTACTTAACT
>MHEH01000185.1/1598-1666 Nitrospirae bacterium RBG_13_43_8
RBG_13_scaffold_6510, whole genome shotgun sequence.
TAAACAACTGGCGATGGGGTTCGCCATTAACCGTCCTTAAAATTAAGGAATGATAACCCC TACTGGGAC
>FZOU01000002.1/249716-249630 Granulicella rosea strain DSM 18704 genome assembly, contig: Ga0111638_102
ATAATTCAAGGTGATGGGGTTCCACCGCACAACCGCCCAGTTTTTCGGCAACGAGCCAAA
GACGAAGCTGATGACTCCTACGATTTG
>FRAS01000007.1/ 92672-92745 Hymenobacter psychrotolerans DSM 18569 genome assembly, contig: EJ61DRAFT_scaffold00007.7
TGCGCATACGGTGATGGATTTCCACCGCGAACCGCCGAAGCTTCGTGCTTTGCGCTGATG
ATTCCTACTGAGCG
>DS996844.1/92835-92759 Eubacterium biforme DSM 3989 Scfld8 genomic scaffold, whole genome shotgun sequence.
AATCTTTTTGGGAATGAAGTTCTCCCTTAGTTTAAAACTAAAACCGCTTTTTATAGGCTG
ATGACTTCTGTGTTTTT
>CP009355.1/584421-584486 Vibrio tubiashii ATCC 19109 chromosome 2, complete sequence.
ATTTCCCAAGGTGATGGGGTTCCACCTACTTAACCGCCAATTCTGGCTGATGACTCCTAC
AGATAC
>LMS001000011.1/147408-147332 Arthrobacter sp. Soil782 contig_6, whole genome shotgun sequence.
GTCGAATCCGGCGATGGATCCCGCCGGGACCCACTTGGGTCCGAACCGCCGCAGTGGCTG
ATGGTTCCTGCCTTCGA
>LECU01000009.1/160315-160250 Pedobacter sp . BMA contig09, whole genome shotgun sequence.
TTGACTTACGGAAATGGTGTCCTTCCGATTCAACCGCTTTTCAAAGCTGATGGCGCCTGC
AAAGTA
>MFRQ01000032.1/11321-11222 Candidatus Melainabacteria bacterium
RIFOXYA2_FULL_32_9 rifoxya2_full_scaffold_1691, whole genome shotgun sequence .
TTAAATTTAGGTGATGAGGCTCGCCAGTGTTCGCAAGAACCGAACTGCCAGAATAAGTGG
GTAATGCTCGCTTTAATGGGCTGATAGCTTCTACTTCATT
>AWWV01005986.1/4472-4400 Corchorus capsularis cultivar CVL-1 contig06004, whole genome shotgun sequence .
TGCGCGCTTGGAGATGGCGTTCCTCCTGGGAAACCGAACCGCAGCATTCGCTGCCGATGA
CGCCTACAGTTAC
>CP015438.1/1600163-1600222 Anoxybacillus amylolyticus strain DSM 15939, complete genome. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ACTTGAATAGGCGATGGAGTTCGCCATAATTGCCGGAAGGCTGATGACTCCTACCAGTTG
>MGFH01000188.1/16403-16464 Candidatus Wallbacteria bacterium GWC2_49_35 gwc2_scaffold_4870 , whole genome shotgun sequence.
TGAAATAATGGCGATGGAGTTCGCCAATAACCGCATTTTGTGCTGATGACTCCTACTTTT
TG
>AEXE01000085.1/1430-1366 Burkholderia sp. TJI49 contig072.1, whole genome shotgun sequence.
TTGAACCCAGGAGATGGCATTCCTCCCGACAACCGCCCTTCGGGGCTAATGATGCCTACA
GGCCT
>MKVW01000001.1/773297-773358 Rhizobiales bacterium 62-17
SCNpilot_cont_300_bf_scaffold_105 , whole genome shotgun sequence.
CGTAGCGATGGGGATGGAGTTCCCCGATACCCGCCGCGAAGGCTGATGACTCCTACTGCG AC
>MNRF01000414.1/23270-23343 Clostridiales bacterium 42_27
Ley3_66761_scaffold_1716, whole genome shotgun sequence.
TGCGCAATCGGGAATGAGGTTCTCCCACGGCACTGCCGAAACCGCTTTTTTGAGCTGATG ACTTCTGTGATTTT
>JANZ01000004.1/78261-78188 Mycobacterium kansasii 732 gmk732. contig .3 , whole genome shotgun sequence .
CGACTGACTGGCGATGGTGCCCGCCCGGAAGACGACTTCCGAACCGCCACAAGGCTGATG
GCCCCTGCGAACGA
>AGSN01000124.1/10893-10954 Mesorhizobium amorphae CCNWGS0123 contig00132, whole genome shotgun sequence .
ATAGCGCTCGGGAATGGTGTCTCCCTGAAAACCGCGCTGATGCTGATGACGCCTACTGAC
GA
>ATIB01000027.1/93607-93537 Sphingobium baderi LL03 contig027, whole genome shotgun sequence.
GGCATCAACGGCGATGGATTTCCGCCTGGCTTCGGCCGAACCGCCTCGGGGTTGATGATT
CCTACCTGCTG
>AWUE01004991.1/274-200 Corchorus olitorius cultivar 0-4 contig05010, whole genome shotgun sequence .
TGCCTCTGAGGTGATGGCGTGCCACCTTACCCAACCGCCCTCTGGTTACAGACGGCTGAT
GACGCCTGACAACTT
>MKRI01000013.1/30949-31027 Caulobacterales bacterium 68-7
SCNpilot_cont_300_bf_scaffold_1229, whole genome shotgun sequence.
CCCCGCCTCGGCGATGGATTTCCGCCTGGCGTCATGCCGAACCGCTCCTCGGGAAGGAGC CGATGATTCCTACTTTGAC
>CP013476.1/1492496-1492438 Turicibacter sp . H121, complete genome.
TATTGAATAGGCGATGGAGTTCGCCATTAACCGCGAAAGCTAATGACTCCTACTCTTTA
>CP015971.1/1423680-1423742 Arachidicoccus sp. BS20, complete genome.
ATCGCAAAAGGAAATGGTGTCTTCCTTCTCCAACCGCTAAAAGCTGATGGCGCCTACAAA
TTT
>GG753567.1/57762-57824 Serratia odorifera DSM 4582 genomic scaffold SCAFFOLD1, whole genome shotgun sequence.
TCTTCCCTTGGAGATGGCATTCCTCCATAACCGCCCTTACCGGCTGATGATGCCTACGTT
ACC
>AOTI010848925.1/18020-17949 Triticum urartu cultivar G1812 contig848926, whole genome shotgun sequence .
CTCTATTAAGGTGATGGTGCTCCACCTTTCTCAACCGCCAAAATTCTCTTGGATGATGAC
GCCTGATATATA
>AP012551.1/3160364-3160290 Plautia stali symbiont DNA, complete genome.
TGCCTCTGAGGTGATGGCGTGCCACCTTACCCAACCGCCCTCAGGTCACAGACGGCTGAT
GACGCCTGACAACTT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>AP013036.1/1991234-1991167 Enterococcus mundtii QU 25 DNA, complete genome .
AAATGAATAGGTGATGGTGTTCGCCTTTAAATGTCGATCGATAGATGACTAATGACGCCT
ACTAAAAA
>CP001001.1/4759637-4759698 Methylobacterium radiotolerans JCM 2831, complete genome.
CGTCCCGATGGGGATGGAGTTCCCCGAAACCGCCCGTCAGGGCTGATGACTCCTACCGCG
AA
>CBWP010000050.1/12734-12660 Escherichia coli ISCll, WGS project
CBWP01000000 data, contig: c64
GATTTACAAGGTGATGGTGCTCCACCTTTCCCAACCGCCGGTTTTTCAACACCGGATGAT
GGCGCCTGATATACA
>MNYM01000024.1/77452-77520 Flavobacteriaceae bacterium CG2_30_31_66 cg2_3.0_scaffold_84_c, whole genome shotgun sequence.
ATTTAAGAAGGAAATGGTGTCCTTCCTTTTCTAACCGCTTTTTTAAAAGCTGATGACGCC TGATTAAGC
>AWTC01000004.1/68382-68315 Sporolactobacillus laevolacticus DSM 442 Contig004, whole genome shotgun sequence.
TAGTAAATAGGCGATGACGTTCGCCTTCAACCGCTGTACTAGAATCAGCTAATAACGTCT
ACCTTGAT
>FQXV01000011.1/100368-100426 Sporobacter termitidis DSM 10068 genome assembly, contig: EK05DRAFT_scaffoldOOOll .11
AAAACGACAAGCGATGGAGTTCGCTTTAAATGCTTCTGGCTGATGACTCCTACAAACGG
>HG425166.1/734874-734937 Methanobacterium sp. MB1 complete sequence
TCTATTGCTGGTGATGGGGTTCACCATTAACCGCCTATCCAAGGCTGATGACTCCTAGAA
ACTA
>KE136389.1/620807-620745 Enterococcus saccharolyticus ATCC 43076 genomic scaffold acpML-supercontl .1, whole genome shotgun sequence.
TTTGGGCATAGGGATGGTGCTCCCCTATAACCGCCAGAAATGGCTGATGGCGCCTGTTAT ATA
>FR894469.1/6216-6145 Clostridium sp . CAG:299 genomic scaffold, scf326
ATAAAACTAGGGAATGAAGTTCTCCCATGGCAAGCCAAAACCGCTGATAGAGCTGATGAC
TTCTGTGACAAC
>CP009573.1/2691-2767 Sphingomonas taxi strain ATCC 55669 plasmid STP2, complete sequence.
CTCGGGCACGGCGATGGATTTCCGCCGGGCCATTCGGCCGAACCGCCCTCGCAAGGGCCG
ATGATTCCTACCAAGCG
>FPIQ01000006.1/72610-72543 Nitrosovibrio sp . Nvl7 genome assembly, contig: Ga0111729_106
TTCCCCCACGGAGATGGCATGCCTCCGACAACCGCCGCGCCCGCGCGGCTGATGATGCCT
ACGCAACA
>NJIH01000001.1/186182-186114 Candidimonas nitroreducens strain SC-089 Scaffoldl_l, whole genome shotgun sequence.
ATTGCTTACGGAGATGGCATTCCTCCGATAAACGCCGGCGCTCAGCCGGATGATGATGCC
TGCATGATC
>AOTI010641095.1/136-202 Triticum urartu cultivar G1812 contig641096, whole genome shotgun sequence .
TCAAAAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTAAAAAGCTGATGACGCCTG
GTTAATC
>AVPD02000038.1/13376-13288 Arthrobacter sp . AK-YN10 contig038, whole genome shotgun sequence.
TCTGTTCACGGTGATGGATCCCGCCGGAACAGCCCGCACGCGCGTGGGCTTGTTCGAACC
GCCTCCTGGCTAATGGTTCCTACCTTTGC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>FP565176.1/1682310-1682380 Xanthomonas albilineans GPE PC73 complete genome
TGCGCCGCAGGAGATGGCGTTCCTCCTTTAACCACCGTGCGATCCGCATGGTTGATGACG
CCTACAGCCGC
>AAEW02000001.1/156146-156076 Desulfuromonas acetoxidans DSM 684 ctg66, whole genome shotgun sequence .
CATGTTAACGGGAATGGAGTTCTCCCGGTAACCGCCTGTCGATTTTTCAGGCTGATGACT
CCTGTATGAGG
>MEW001000012.1/3987-4048 Candidate division Zixibacteria bacterium
RBG_16_50_21 RBG_16_scaffold_16836, whole genome shotgun sequence.
AGTATTAATGGCGATGGGGTTCGCCTTTAACCGCCTAAACGGCTGATAACTCCTACCGGT TA
>MHEI01000122.1/8060-7997 Nitrospirae bacterium RBG_16_43_11
RBG_16_scaffold_91172 , whole genome shotgun sequence.
CATGGAGATGGTGATGGGGTTCACCTTATAACCGCCTGAGAAGGCTGATGACTCCTACCT GTAG
>LGIP01000019.1/54203-54270 Chryseobacterium sp . Hurlburt 01
HurlbutContig21 , whole genome shotgun sequence.
TGTAAACAGGGAAATGGTGTTCTTCCTTACCCAACCGCTTTAGCAAAGCTGATGACGCCT GATTAATT
>LGEZ01000003.1/9477-9413 Clostridia bacterium 41_269 MPF_scaffold_6, whole genome shotgun sequence .
TATAAAAAAGGCGATGGGGCTCGCCAGAATTACCCAAACGAAGGGTTGATGGCTCCTAGT
TTTTC
>MKRJ01000017.1/22706-22768 Rhizobiales bacterium 65-79
SCNpilot_cont_1000_bf_scaffold_222 , whole genome shotgun sequence.
TTCCCGTGCGGTGATGGAGTTCACCATGTAACCGCCCCCAGGGCTGATGACTCCTCTGTT GCT
>MKVF01000006.1/236977-237051 Flavobacteriia bacterium 40-80
scnpilot_p_inoc_scaffold_14 , whole genome shotgun sequence.
ATTAAAAAAGGAAATGGTGTTCTTCCTTACCCAACCGTCCCGATAGCTATCGGAGCTGAT GACGCCTGATTAATT
>MELW01000360.1/7682-7620 Alicyclobacillus sp. RIF0XYA1_FULL_53_8 rifoxyal_full_scaffold_975, whole genome shotgun sequence.
AATTGCATCGGTGATGGAGCTCACCATTAAACTCCTGCAAAGGATAATGGCTCCTGCGAA TAG
>FQXE01000023.1/8211-8275 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_123
TGTGTTACAGGAGATGGCATTCCTCCTTATAACCGCCCAAACCGGCTGATGATGCCTACC
AAACT
>FR886277.1/3436-3520 Eubacterium sp . CAG:252 genomic scaffold, scf53
ACTTAAACAGGGAATGAAGTTCTCCCTTGGAATCTGGAAGCAGAAACCTAGACTGCTTAT
TATAGCTGATGACTTCTACGATTAT
>CP003221.1/3960656-3960716 Desulfovibrio africanus str. Walvis Bay, complete genome.
CGCCGCACAGGGGATGGAGTCCCCCTTGAACCGCCAAGTTGCTGATGACTCCTACTGTCG
A
>CP016895.1/890108-890042 Acinetobacter larvae strain BRTC-1 chromosome, complete genome.
ACATTCTAAGGAGATGGTATTCCTCCGATTACAAACCGCCTTTTTGGCTAATGATGCCTA
CGTTACT
>KQ758640.1/49975-50034 Bacillus sp . SGD-V-76 genomic scaffold Scaffoldl4, whole genome shotgun sequence . APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AAAATATACGGCGATGGAGTTCGCCAAAACTGCTGCAGAGCTAATGACTCCTACCCATAT
>FTNT01000006.1/163316-163237 Williamsia sterculiae strain CPCC 203464 genome assembly, contig: Ga0104459_106
AGCAGGTGCGGCGATGGATCTCGCCGGATTTCCTGTTCGGGAAATCGAACCGCCTCTCGG
CTGATAGTTCCTGCTGATGA
>MIHC01000025.1/44643-44579 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
TTAGCGGTCGGCGATGAGGCTCCGCCCCCTTCGCGTCATCTGACGCTGATGGCTTCTACC CCGCA
>LVYV01000014.1/51254-51331 Tardiphaga sp. Vaf07 contig_21, whole genome shotgun sequence.
AAACGACACGGTAATGGATTCTGCCGGGCCATCGTGGCCGAACCGCTTCCGAGAGAAGCT
GATGACTCCTACTCAACT
>LN885086.1/39594-39523 Candidatus Nitrospira inopinata isolate ENR4 genome assembly, chromosome: 1
CAGAGCATTGGTGATGGAGTTCACCAAGAACCGTCCCACCGTGAAAGCGGGACTGATGAC
TCCTACCGAAGA
>CYSP01000003.1/154139-154078 Propionispora sp . 2/2-37 isolate 2/2-37 genome assembly, contig: 2/2_contig3
CTGTATATCGGCGATGGGGTTCGCCATACATACCGTACATCGGTTATGACTCCTACCACA
TC
>MKSQ01000024.1/24728-24650 Bosea sp . 67-29
SCNpilot_expt_1000_bf_scaffold_559, whole genome shotgun sequence.
AGCGAGCCCGGCAATGGATTCTGCCTGAGCTGTTCTTCGGCTCGAACCGCCTCGGAAGGC
TGATGACTCCTACTCGACT
>FSRU01000002.1/319558-319495 Paraburkholderia phenazinium strain GAS95 genome assembly, contig: Ga0132017_12
TCGCCTGCTGGAGATGGCATTCTCCTTTAACCGCCCAGACAGGGCTGATGATGCCTGCAT
GCCC
>MKKK01000021.1/12660-12592 Acinetobacter qingfengensis strain ANC 4671 Contig28, whole genome shotgun sequence.
GATCCATCGGGAAATGGCATACTTCCCTATACAAACCGCCAATTGTTGGCTAATGATGCC
TACGTAACC
>HF952018.1/1488085-1488027 Thermobrachium celere DSM 8682 genomic scaffold, scaffold33
AAGTTAATAGGCGATGGAGTTCGCCTAAATGCTGATAAGCTAATGACTCCTACAAATGT
>LGFH01000009.1/2976-3038 Thermotoga sp . 50_1627 MPF_scaffold_34 , whole genome shotgun sequence.
AACAGTTGGGGCGATGGAGTCCGCCCACAAGTGCCCGTCCGGGCTGATGACTCCTACCTT
CGG
>BBMQ01000019.1/1678-1758 Vibrio sp . C7 DNA, contig: contig00024, strain: JCM 19233.
CCAACTCACGGAGATGATGTTCCTCCATGCTAATTTTTTGATTAGCAAAACCGCTTAAAC
GCTGATGACGTCTGTTAATAC
>LGJH01000116.1/18218-18294 Novosphingobium sp . ST904 contig_51, whole genome shotgun sequence.
CGCAGGCACGGCGATGGATTTCCGCCGGGCCATTCGGTCGAACCGCCCTCGCAAGGGCCG
ATGATTCCTACCAGGCG
>MTKC01000003.1/5829667-5829733 Streptomyces griseofuscus strain NG1-7 NG1-17_1_1, whole genome shotgun sequence.
GAGGCGTACGGCGATGAGGCTCGCCGCAACCGCACCCATGGCCGGTGCTGATGGCCTCTA
CGTGATG
>JXKQ01000003.1/71030-71104 Enterococcus hermanniensis strain DSM 17122 APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
Scaffold3, whole genome shotgun sequence.
ATTAAATACGGGAATGAGTGTCTCCCTAAGATTTTTTCTTAAACCGCTAATCAAGCTGAT
GACTTCTACAAATTT
>JH651379.1/3012680-3012743 Joostella marina DSM 19592 genomic scaffold Joomascaffold_l , whole genome shotgun sequence.
ATCACATTTGGCAATGAGGTCTGCCTTTAACCGCTCAAAATGAGCTGATGACTTCTACTG ATAA
>MGVQ01000022.1/32399-32329 Elusimicrobia bacterium RIFOXYC2_FULL_34_12 rifoxyc2_full_scaffold_1509, whole genome shotgun sequence.
TAATTAACTGGAGATGGGGTTCTCCGTAAAGCGTCCCGCTTAAAATTGGGGCTGATGACT CCTACTTAAAC
>BAMW01000022.1/78547-78612 Acetobacter indonesiensis 5H-1 DNA, contig: Abin_022.
CGCCGCATAGGGGATGGGGTTCCCCCTTATAACCGCCCGTACAGGGCTGATGACTCCTGC
CGTTGC
>LGY001000007.1/208624-208683 Acetobacterium bakii strain DSM 8239 contig_7, whole genome shotgun sequence.
ATTAACGATGGTGATGGAGTTCACCCAAATTGCTACTTAGCTAATGACTCCTGCAGGCTT
>CP013023.1/4660124-4660056 Paenibacillus bovis strain BD3526, complete genome .
AATTGAATCGGTGATGGAGTTCGCCGTTAACCGCTGGCATAATAATCAGCTGATGACTCC
TACCTGTTG
>FQXE01000008.1/21421-21489 Candidimonas bauzanensis strain CGMCC 1.10190 genome assembly, contig: Ga0070552_108
TTCAGATACGGAGATGGCATACCTCCTATCCCAACCGCTGGATTCCCAGCTGATGATGCC
TACGAATTT
>JJMI01000057.1/53371-53301 Mangrovibacter sp. MFB070 Contig-058, whole genome shotgun sequence.
AAAGCACAAGGTGATGGCGTTCCACCTTACCCAACCGCTCCATTCAGGGGGCTGATGACG
CCTGGCGTAAC
>CP013347.1/2218741-2218801 Paraburkholderia caribensis strain BcrslW chromosome 1, complete sequence.
GCGCGCGCCGGAGATGGCATTCTCCTAAACCGCCCTCGTGGCTGATGATGCCTGCTTCGC
C
>HF570958.1/3227824-3227894 Tetrasphaera japonica T1-X7 genomic scaffold, 1540_scaffoldl
GTGGCCACAGGTGATGGACCCCACCTGGGGCGAGGCCCCGAACCGCATCCGCTAATGGCT
CCTGCGACGTT
>CP003597.1/2805389-2805307 Chroococcidiopsis thermalis PCC 7203, complete genome .
ATAATCTGAGGTGATGGGGCTCGCCTTAACTGCCATCTGGAAATTGACATGACAACCAGA
AGGCTGATGGCTCCTACTGTTCC
>CP003653.1/502946-502859 Stanieria cyanosphaera PCC 7437, complete genome .
CATAGAGATGGCAATGGAGCTTGCCAAAACCGCCCTTTGTTTTTAAGTTTTTCATAATCT
TTAAGAGGCTGATAGCTCCTACTCTTCC
>KN125580.1/3295453-3295393 Paenibacillus macerans strain 8244 genomic scaffold scaffoldl, whole genome shotgun sequence.
AGTATTAAAGGCGATGGAGTTCGCCTGAACCGCTCCCCGAGCTAATGACTCCTACCAGTT
G
>CP000423.1/2429871-2429817 Lactobacillus casei ATCC 334, complete genome.
AAATTAAATGGCGATGGTGTTCGCCTATACGCAAGTTGATGACACCTACCTTGAG
>MEMZ01000092.1/8658-8723 Armatimonadetes bacterium RBG_16_58_9 APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
RBG_16_scaffold_26212 , whole genome shotgun sequence.
TCGACTAACGGGGATGGAGTCCCCCTTGAAAACCGCCTTGGCCATGCTGATGACTCCTAC
CGCGCG
>KK106988.1/2431965-2431899 Streptomyces sp . Tu 6176 genomic scaffold scaffoldOOOOl, whole genome shotgun sequence.
ACGTGCCGGGGCGATGGGGCTCGCCCGTGACCGCACGGACCGCCGTGCTGATGGCCTCTG
CGTGGAC
>MEQV01000269.1/15415-15487 Betaproteobacteria bacterium
RIFCSPLOWO2_02_FULL_62_17 rifcsplowo2_02_scaffold_6838, whole genome shotgun sequence.
AATATCAGTGGAGATGGCATTCCTCCCTTAACCGCCTCACGTTTCTTTCGGGGCTGATGA
TGCCTGCGTTAGT
>LMPF01000022.1/26482-26598 Burkholderia sp . Leafl77 contig_5, whole genome shotgun sequence.
TTGCGCTACGGAGATGGCATTCCTCCTGCAGCTCATCGCAAACCGCTGGATTAGTTACGT
AGCTGACTTATCTGATTGAGTCGCCTAATTAAGCCAGCTAATGATGCCTACGCGTTC
>MNIJ01000235.1/7576-7504 Acidobacteria bacterium 13_1_20CM_58_21
13_1_20cm_full_scaffold_l 091 , whole genome shotgun sequence.
TTTCGATTTGGCGATGGAGTTCGCCATAACCGCCCCGGTGATGTTCCCCGGTGCTGATGA
CTCCTGGCAGCCC
>LISW01000001.1/1383246-1383309 Streptomyces sp . CB01249 scaffoldl, whole genome shotgun sequence.
ATGGCGAACGGCGATGAGGCCCGCCATAACCGCGGAATTCCCCGCTGACGGTCTCTGTTT
CTTG
>FN692037.1/1065514-1065574 Lactobacillus crispatus ST1 complete genome, strain ST1
AACTGAATAGGTGATGACGTTCGCCATTAACCGAGTAAAATCTAATGACGTCTACTTTAT
C
>MNXN01000067.1/6742-6660 Armatimonadetes bacterium CG2_30_66_41
cg2_3.0_scaffold_l 6220_c, whole genome shotgun sequence.
CTGACTTTCGGCGATGGAGTTCGCCGGCTCCGGTTTTCAGGCTGGAGCGAACTGTTCGGA ACGCTGATGACTCCTGACCAGTA
>LMSL01000040.1/145353-145274 Frateuria sp . Soil773 contig_8, whole genome shotgun sequence.
TGTACGCTTGGAGATGGCATTCCTCCGCCCTCGGCGTTCCGGGGGGAACCGCCATTGTGG
CCGATGATGCCTGCTCGTTT
>MGQS01000118.1/1171-1108 Deltaproteobacteria bacterium RBG_16_54_11 RBG_16_scaffold_68799, whole genome shotgun sequence.
ATTGTTATTGGTGATGGAGTTCGCCATAACCGCCCTCAAGAGGGCTGATGACTCCTATCT TGAA
>LSTD01000273.1/5338-5265 Planctomycetaceae bacterium SCGC AG-212-F19 AG- 212-F19.6, whole genome shotgun sequence.
TACGGAAAAGGCGATGGAGTTCGCCATAACCGCGCGGTTCCGAAGGAGCCGTTGCTGATA
ACTCCTACCGCGAG
>CP001727.1/264912-264974 Alicyclobacillus acidocaldarius subsp.
acidocaldarius DSM 446, complete genome.
CAGTATGTCGGTGATGGAGCTCACCGGAACCGCCCACGACGGGCTGATGGCTCCTGCGAT
GCA
>MWPH01000003.1/579874-579786 Natronolimnobius baerhuensis strain CGMCC 1.3597 ZB100002, whole genome shotgun sequence.
ACGTATTCGGGCGATGGGGCCCGCCTGACCCAACTGCCGCTCTCGGCGACCGACACCCCT CGAGCGTGGCTGACGGTCCCTGCCTTCAA
>MBTF01000001.1/592373-592438 Mucilaginibacter pedocola strain TBZ30 APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
scaffoldl, whole genome shotgun sequence.
GATTTAAACGGAAATGGTGTCTTCCGTTCAAAATAACCGCCTCGTGCTGATGACGCCTAC
AATTAT
>MKVZ01000012.1/35950-36011 Rhizobiales bacterium 63-22
SCNpilot_cont_500_bf_scaffold_404 , whole genome shotgun sequence.
TTGTCGCAAGGTGATGGAGTTCACCTTTAACTGCCTTCGAGGCTGATGACTCCTGTTTCG CC
>LSLJ01000003.1/352869-352931 Sporomusa sphaeroides DSM 2875
SPSPH_contig000003, whole genome shotgun sequence.
TTGAATAGCGGCGATGGAGTTCGCCCTGAATCCCGTGCAAACGGTAATGACTCCTACCAC ATA
>BBMS01000037.1/42981-43044 Vibrio variabilis DNA, contig: contig00037, strain: JCM 19239.
TACCACCAAGGTGATGGGGTTCCACCTACTTAACCGCCATTTGGCTGATGACTCCTACAG
TTTC
>CP001504.2 /2120579-2120511 Burkholderia glumae BGR1 chromosome 2, complete sequence.
CCGCGCTGCGGAGATGGCATGCCTCCGCCCCCAACCGCCGCTCAGCCGGCTGATGATGCC
TACGGATTC
>LATZ01000021.1/34082-34023 Bacillus sp . SA1-12 scf7180000003373, whole genome shotgun sequence.
GTTAACATAGGTGATGGAGTTCACCTTTAACCGCCGTTGGCTAATGACTCCTGCCAGTAA
>NFLE01000014.1/2802-2879 Lachnoclostridium sp . Anl4 Anl4_contig_14 , whole genome shotgun sequence.
CACACATTCGGGAATGAAGTGCTCCCGTAACCGTAATTGGTAAAACCGCTTGAGAGAGCT
GATGACTTCTGTGAATAC
>KE150017.1/129019-128943 Lachnospiraceae bacterium 2_1_46FAA genomic scaffold aclXT-supercont-complete , whole genome shotgun sequence.
GGTGATAGAGGGAATGAAGTTCTCCCTTAGTTATAAAACTAAAACCTCTAATCTTAGCTG ATGACTTCTGCATTTAG
>JJMU01000002.1/30791-30852 Sphingobacterium sp . ACCC 05744 h261, whole genome shotgun sequence.
TACATCTGTGGCAATGATATCTGCCCTTAACCGCCCAAAAAGCTGATGATGTCTACTTAA
TT
>LDRW01000016.1/291186-291116 Novosphingobium barchaimii strain NS277 contig_16, whole genome shotgun sequence.
CGCTGGCAAGGCGATGGATTTCCGCCGGGCTTCGGCCGAACCGCCTCCGGGCTGATGATT
CCTACCTGATG
>LPWA01000099.1/48188-48112 Mesorhizobium loti strain UFLA 01-765 contig_20, whole genome shotgun sequence.
CCGAAAAATGGGGATGGGGTTCCCCCGAAACCGCCCACGCTCAGTACGAGTATCGGGCTG
ATGACTCCTGCCAGGCG
>CP000004.1/152384-152449 Gluconobacter oxydans 621H plasmid pGOXl, complete sequence.
TCCCGACAAGGGGATGGAGCTCCCCCTTCAACCGCCCTCGCAAGGGCTGATGGCTCCTAC
CGCGAC
>MEKU01000048.1/29444-29375 Acidobacteria bacterium
RIFCSPLOWO2_02_FULL_67_21 rifcsplowo2_02_scaffold_2758, whole genome shotgun sequence.
CGTCTCGTTGGTGGTGGAGTCCACCTGAACCGCCGGCCCCGACTGGACGGCCAATGACTC
CTGCGGCGAT
>CP019433.1/173107-173182 Jeotgalibaca sp. PTS2502, complete genome.
AATAAAAGAGGGAATGATCGTCTCCCAAAGTTGAAATAAACTTGAACCGCTTAAAGCTGA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TGACTTCTACCTGTTT
>CP001941.1/602419-602477 Aciduliprofundum boonei T469, complete genome. CATTTATAAGGTGATGGGGCCCACCAAACCGCCGCAAGGCTGATGGCTCCTTTTGAAAA
>CAHT01000061.1/22514-22587 Methylacidiphilum fumariolicum SolV, WGS project CAHT01000000 data, contig: 77-2114
CAACTATTGGGCGATGGGGCTCGCCGTCCGTTAACCGCCACTCCTTTGAGAAGGCTGATA
GCTCCTACGATTAC
>AYZJ01000062.1/4412-4344 Lactobacillus camelliae DSM 22697 = JCM 13995 strain DSM 22697 NODE_98, whole genome shotgun sequence.
AAGTTAATCGGCGATGACGTTCGCCACTAAAACTAATTGGTAAATCAATTTGATGACGTC TACTGCACA
>AP012159.1/1360953-1361031 Komagataeibacter medellinensis NBRC 3288 DNA, complete genome.
AATAAAGATGGCAATGGACTCTGCCTGATTTCAGCCATGAAATCGAACCGCCCGCCGGGC
TGATGATTCCTACCCCGCC
>CP004360.1/907036-907115 Komagataeibacter xylinus E25 chromosome, complete genome.
AAGACCGACGGCAATGGACTCTGCCGGACCATGATGGTCGAACCGTCTCCTGCATGGGGG
CTGATGATTCCTACCCAACG
>FRAF01000024.1/21577-21643 Alicyclobacillus sp . USBA-503 genome assembly, contig: Ga0105849_124
ATTGAATCCGGTGATGGAGCTCACCGTGAACTGCCGGAATGAACTGGCTAATGGCCCCTA
CTGTTCT
>MNRE01000171.1/67202-67285 Clostridiales bacterium 41_2 l_two_genomes Ley3_66761_scaffold_497 , whole genome shotgun sequence.
ATATGTTTTGGGAATGAAGTTCTCCCATTGATACATTCCTTGTATCACAAACCGCTGGTT AAGGCTGATGACTTCTGCGAATTT
>CM001437.1/1181018-1181078 Myroides odoratus DSM 2801 chromosome, whole genome shotgun sequence.
ATCATAATAGGCAATGATGTCTGCCTCGAACCGCCTTGTAGCTGATGACGTCTAGTATAC
T
>CP002868.1/2636695-2636774 Treponema caldarium DSM 7334, complete genome.
ATTTAAAGCGGTGATGGAGTCCACCTTGGCACCAGCCTGAACCGCCCGGGATATCCCCTG
CTGATGACTCCTTCCGATAC
>LXKA01000349.1/9186-9125 Burkholderia ginsengiterrae strain DCY85 NODE_5, whole genome shotgun sequence .
TGTTGCGCTGGAGATGGCATTCTCCTTTAACCGCCCTTGTGGCTGATGATGCCTGCTCGC
CC
>ADNY01000026.1/2434-2494 Lactobacillus amylolyticus DSM 11664
contig00030, whole genome shotgun sequence.
AATCAAATAGGCGATGACGTTCGCCATTAACTGAGTAAAATCTGATGACGTCTACTACCT
C
>JH932301.1/637310-637385 Facklamia ignava CCUG 37419 genomic scaffold supercontl .2 , whole genome shotgun sequence.
TTGAAGATAGGAAATGAGGTTCTTCCTTAGCCGTATGGCTTAAACTGCTTTGAATGCTGA
TGACTTCTACCTGTTC
>AURB01000141.1/44632-44571 Alicyclobacillus acidoterrestris ATCC 49025 contig_39, whole genome shotgun sequence.
AATTGCATCGGTGATGGAGCTCACCAACATCACCCGAAAGGGTTAATGGCTCCTGCGAAG
GT
>AVPD02000038.1/14384-14302 Arthrobacter sp . AK-YN10 contig038, whole genome shotgun sequence.
TGTTGTAATGGTGATGGATCCCACCGGGGCCGATTGATTGGCTGGCCTGGACCGCCGACG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TGGCTGATGGTTCCTGCTCCGAC
>FQVH01000001.1/130089-130029 Caldanaerobius fijiensis DSM 17918 genome assembly, contig: EJ21DRAFT_scaffold00001.1
AAATTAATAGGTGATGGGACTCACCATTATGTGCCTTAAGGCTGATGGTTCCTACTAAGT
T
>CSXB01000013.1/23266-23194 Mycobacterium abscessus strain PAP053 genome assembly, contig: ERS075544SCcontig000013
TTCGTCAGGCGCGATGGATCTCGCCAGGGCTTGTCCCGAACCGCCACTGACGGCTGATAG
TTCCTGTGTTGAT
>MASI01000011.1/38499-38437 Methyloligella halotolerans strain VKM B-2706 A7A08_contig000011, whole genome shotgun sequence.
ACCAATGAAGGGGATGGGGTTCCCCGACAACCGTCCATGAGGACTGATGACTCCTGTTTG CTA
>KQ758903.1/235653-235729 Dehalogenimonas alkenigignens strain IP3-3 genomic scaffold scaffoldOOOOl, whole genome shotgun sequence.
TCGTAAAGCGGCGATGAAGCTCGCCAAGGAATTTTCACCTGAACCGCTTGCTTAAGGCTG ATAGCTTCTACCGGTAT
>KK106994.1/337598-337535 Streptomyces sp. Tu 6176 genomic scaffold scaffold00007, whole genome shotgun sequence.
CGTCGCCATGGCGATGGGGCTCGCCACAACCGCAGGACCTCCTGCTGATGGCCTCTACCG
GTAA
>FWXI01000013.1/81729-81793 Sporomusa malonica strain DSM 5090 genome assembly, contig: GaO 070592_113
GATGTTTGGGGCGATGGAGTTCGCCATTAAATCCGTAAACAGGCGGTGATGACTCCTACC
AGGTA
>MNFK01000012.1/837-775 Deltaproteobacteria bacterium 13_1_40CM_4_54_4 13_1_40cm_4_scaffold_l 9741 , whole genome shotgun sequence.
ATTGCCTGCGGAGATGGCGTTCTCCATTAACTGCCCGCCAGGGCTGATGACGCCTACTTC AAT
>MERG01000081.1/42967-43036 Betaproteobacteria bacterium
RIFCSPL0W02_12_FULL_62_13 rifcsplowo2_12_scaffold_1558, whole genome shotgun sequence.
GCTCGGTAGGGAGATGGCATTCCTCCCGGAATAAGAACCGCCTGAAATGGCTGATGATGC
CTACTGTGTC
>JAOC01000010.1/319581-319520 Mycobacterium xenopi 3993 GMX3993. contig .9, whole genome shotgun sequence .
TCGCGGACAGGCGATGAGGCTCGCCTTGAACTGCCACAACGGCTAATGGCTTCTACCCGT
GA
>AEXE01000085.1/6764-6696 Burkholderia sp. TJI49 contig072.1, whole genome shotgun sequence.
TGCTCCGAAGGAGATGGCATTCCTCCCACCAAACCGCCGCGCAAGTCGGCTAATGATGCC
TACGTGATC
>MDVT01000007.1/146532-146604 Archaeon Odin LCB_4 OdinLCB4_contig000007 , whole genome shotgun sequence .
ATAAACGTCAGCGATGGAGTCCGCTGTGGGTGTAAGCCTAAACCGCTTTTAAGCTGATGA
CTCCTATCTTTTT
>LVCV01000312.1/20328-20240 Rhodococcus sp . EPR-157 Scaffold28_4, whole genome shotgun sequence.
ATTGCGGCTGGCGATGGATCTCCGCCTAGATCGCCCGTCAAGGGTTGTCTGAACCGCCCC
AGCCCGGGGCTGATGGCTCCTTTCCCTGA
>KQ960178.1/16010-16078 Peptoniphilus coxii strain DNF00729 genomic scaffold Scaffold207, whole genome shotgun sequence.
AAGAAAATAGGGAATGAAGTTCTCCCTTGGGCAACCTAAACCGCAACAGCTGATGACTTC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TACGATTTT
>MDCF01000112.1/526-463 Acidiferrobacter thiooxydans strain ZJ Contig_2, whole genome shotgun sequence .
ACGAGAACAGGAGATGGCATTCCTCCTCATAACCGCCCTACCGGCTGATGATGCCTACAT
AGTC
>MTKC01000003.1/5837480-5837417 Streptomyces griseofuscus strain NG1-7 NG1-17_1_1, whole genome shotgun sequence.
TATGGAGAGGGCGATGAGGCTCGCCCTTGACCGCACATCCTGTGCTGATGGCCTCTACCT
CTCC
>CP002600.1/2792828-2792760 Burkholderia gladioli BSR3 chromosome 2, complete sequence.
GTGCTTGCCGGAGATGGCATTCCTCCGTACACAACCGCCGGTCAGCCGGCTGATGATGCC
TACGCATCC
>AZMC01000279.1/8695-8624 Negativicoccus succinicivorans DORA_17_25
Q612_NSC00279, whole genome shotgun sequence.
TGTATGAAAGGGAATGAAGTTCTCCCTTGGGTAACCTAAACCGTTTATTAAGCTGATGAC TTCTACAATTAC
>AEXE01000084.1/14637-14698 Burkholderia sp . TJI49 contig072, whole genome shotgun sequence.
TTCCAGCGAGGAGATGGCATTCTCCGTTGAACCGCCTCACGGCTGATGATGCCTACGAAC
CC
>MIDI01000122.1/18129-18059 Thermodesulfovibrio sp . RBG_19FT_COMBO_42_12 rbg_19ft_combo_scaffold_289, whole genome shotgun sequence.
CTTAATATGAGCGATGGGGTTCGCTGTTAACCGTCCTTGATTGTTTAAGGAATGATAACC CCTACTGAACC
>MGQE01000007.1/10900-10839 Deltaproteobacteria bacterium RBG_13_51_10 RBG_13_scaffold_11549_curated, whole genome shotgun sequence.
ATTATCTCCGGCGATGGAGTCCGCCAGAACCGCCTTTCCAGGCTGATGACTCCTGCCAAG GT
>FMIF01000028.1/ 62135-62076 Sporanaerobacter sp . PP17-6a isolate PP176A genome assembly, contig: PP17-6a_contig28
TAAAATTCAGGTGATGGAGTTCACCTTTAAATGCTTTTTGCTAATGACTCCTCCTTAATA
>LUUM01000113.1/172427-172488 Methylosinus sp. R-45379 contig_20, whole genome shotgun sequence.
GTGCGGCCGGGGGATGGGGTCCCCCTTCAACCGCCGCAAAGGCTGATGACTCCTGCTGAT
GT
>CP002551.1/1432273-1432210 Methanobacterium sp . AL-21, complete genome.
ATATGGAATGGTGATGGGGTTCACCTTTAACCGCTTATTTTAAGCTGATGACTCCTGCAC
AATA
>MVIM01000009.1/203384-203456 Mycobacterium tusciae strain DSM 44153 NODE_9_length_212353_cov_68.0015, whole genome shotgun sequence.
CTTCGAAGTGGTGATGGATCTCGCCCGAGCTCGTCGCTCGAACCGCCAATTGGCTGATAG TTCCTGCCGTCAT
>FWXF01000003.1/222602-222668 Desulfacinum hydrothermale DSM 13146 genome assembly, contig: EJ40DRAFT_scaffold00003.3
AAAACGAAAGGCAATGAAGTCTGCCTGAAACGCCCTCACGTTCTGGGATGATGACTTCTA
CTCACGC
>CP003350.1/2397153-2397217 Frateuria aurantia DSM 6220, complete genome.
TTTTCGATGGGAGATGGCATGCCTCCTGTCCCAACCGCCGCAAGGCTGATGATGCCTACG
CCACC
>AP009389.1/615295-615357 Pelotomaculum thermopropionicum SI DNA, complete genome .
ATATTTTCTGGCGATGGGGCTCGCCTTAATTTTGCTGAAAAAGCTGATAGCTCCTACCCG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TAA
>MIHC01000025.1/15457-15517 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
TCGCTGACAGGTGATGAGGCTCACCTTCAACCGCCATACGGCTAATGGCTTCTACCCACT
G
>CP000358.2/504027-504091 Deinococcus geothermalis DSM 11300 plasmid pDGEOOl, complete sequence.
AGCAACACAGGCGATGGCATGCGCCTTGAACCGCCCCGCGTCCGGCTGATTATGCCTACC
TTCCC
>CP006850.1/5443985-5443887 Nocardia nova SH22a, complete genome.
GAAGATGTAGGCGATGGAGCTCGCCTCAACCGCGCCAGCCGAGGTCCTGTGACGATGCGG
GCCTTTACAGCGGTTCGCGCTGATAGCTCCTACCGATCC
>FTOV01000006.1/110401-110334 Chryseobacterium gambrini strain DSM 18014 genome assembly, contig: Ga0111674_106
AGTACAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTTCAAAAGCTGATGACGCCT
GATTAAAA
>MKU001000007.1/339798-339731 Bacteroidetes bacterium 43-93
SCNpilot_expt_1000_bf_scaffold_15 , whole genome shotgun sequence.
TACGCAAAAGGAAATGGTGTCTTCCTGTAATAAACCGCTTTAAAGAAGCTGATGGCGCCT GCAGACCG
>CACA01000037.1/3863-3958 [Oscillatoria] sp . PCC 6506, WGS project
CACA01000000 data, contig: Contig41
AATCCCAACGGCAATGGAGCTCGCCAAAACCGCCTGCTAGTGATTCTTTCCTAACGTTAA
CTAAAACACTAAAAGGCTGATGGCTCCTACTTTTCT
>JZSR01000015.1/16083-16157 Photobacterium iliopiscarium strain ATCC 51761 CFSAN029430_contig0014, whole genome shotgun sequence.
TTTATAGCGGGAGATGATGTTCCTCCTTTAACCGCCTTAATGACTTTCTTTAAGGATAAT GACGTCTGACAACAT
>FQVH01000013.1/28639-28579 Caldanaerobius fijiensis DSM 17918 genome assembly, contig: EJ21DRAFT_scaffold00013.13
CTTTAATAAGGCGATGGAGCTCGCCTTAAATGCCTGAGATGCTGATGGCTCCTGCTTTTA
A
>CP001997.1/35922-35856 Aminobacterium colombiense DSM 12261, complete genome .
TAAATAGCGGGTGATGGGGTTCACCATAAATACTATAAATGCGAAAGCTGATGACTCCTA
CAAGTAT
>MAST01000004.1/35803-35889 Humibacillus sp . DSM 29435 contigl2, whole genome shotgun sequence.
TGGTATACGGGTGATGGAGCTTGCCCGAGAAGCCAGGTGACCCTGGCTTTTCAAACCGCC
AAGATGGCTGATAGCCCCTGAGATCTT
>JH417664.1/1346-1274 Flavonifractor plautii ATCC 29863 genomic scaffold Scfld64, whole genome shotgun sequence.
CCGGAACACGGGAATGAGGTTCTCCCGGGGCGCTGGCCCGAACCGCCGCTTGGCTGATGA
CTTCTGCAAGAGA
>CP000009.1/2355903-2355981 Gluconobacter oxydans 621H, complete genome.
GAGGCAGAAGGCAATGGATTCTGCCTGACCGTTCCCGGTCGAACCGCTTCATCCGGAAGC
TGATGATTCCTGTCCCTCG
>MKSM01000129.1/35964-36024 Nitrobacter sp . 62-23
SCNpilot_cont_300_bf_scaffold_793 , whole genome shotgun sequence.
GAAATCGATGGGGATGGGGTCCCCCGACAACCGCCGCGAGGCTAATGACTCCTACCGGGC
G
>FQZV01000012.1/15866-15803 Geosporobacter subterraneus DSM 17957 genome assembly, contig: EJ58DRAFT_scaffoldOOOlO .10 APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AACCATTAAGGTGATGGAGCTCACCGTAAACATTGCTGCAGCAATTAATGGCTCCTACTA
AGGA
>LASZ01000009.1/144060-143988 Luteimonas sp . FCS-9 scf7180000000232 , whole genome shotgun sequence.
GCCGCCGACGGGAATGGGGCTTCCCGATAACCGCCGGCCGGTGCACGGTCCGGCTGATGG
CGCCTGCTGCCAA
>CP003926.1/1812802-1812729 Gluconobacter oxydans H24, complete genome.
ATAAACGAAGGCAATGGACTCTGCCTGGCCATCGTGGCCGAACCGCCCCCAGGGCTGATG
ACTCCTACCTCGCC
>AFCE01000056.1/840-781 Caldalkalibacillus thermarum TA2 ,A1 ctgl49, whole genome shotgun sequence.
CGCGAAATAGGCGATGGAGTTCGCCATAACCGCCTGACGGCTGATGACTCCTACCGATAC
>KQ954289.1/11573-11496 Novosphingobium sp . Fuku2-ISO-50 genomic scaffold
PRJNA299316_s005 , whole genome shotgun sequence.
CTGTCATACGGCGATGGATTTCCGCCGGGTTTTCGGGCCGAACCGCCTTTACCAAAGGCT
GATGATTCCTACCTGACG
>JZEJO 1000001.1/4961423-4961483 Paenibacillus sp . D9 Scaffoldl, whole genome shotgun sequence.
AACCGCATAGGCGATGGAGTTCGCCTTTAAACGCCCAAGGGCTGATGACTCCTACCAGCG
C
>MERH01000157.1/5469-5529 Betaproteobacteria bacterium
RIFCSPL0W02_12_FULL_62_13b rifcsplowo2_12_scaffold_3504 , whole genome shotgun sequence.
CCGCAGCAGGGAGATGGCATTCTCCCCTAACCGCCGCAAGGCTGATGATGCCTACTCGAG
A
>CYZR01000001.1/463859-463924 Clostridium ventriculi strain
2789STDY5834858 genome assembly, contig: SCcontigOOOOOl
AATTTAATGGGTTATGGAGTTAACCCTAAACCGCTTTATTTTAAAGCTAATGACTCCTAC ACAAAA
>JMIY01000007.1/446381-446306 Candidatus Methanoperedens nitroreducens strain ANME-2d ANME2D_Contig_7.7 , whole genome shotgun sequence.
TTATATGTGGGCGATGAGGTTCGCCTTGGTATTAGCCTAAATTGTCTCGATCGAGACTGA TAACTTCTATTTTCAA
>AE006641.1/885812-885728 Sulfolobus solfataricus P2, complete genome.
AATGAAGTAGGGGATGGCGTCCCCCTGGGGATAACCCCGAACCGCCTCTTGATTAATGAT
AGGGGCTGATGACGCCTACTTCACC
>CP002083.1/228539-228600 Hyphomicrobium denitrificans ATCC 51888, complete genome.
GCCGTCGATGGGGATGGGGTTCCCCGATAACCGCCGCTATGGCTGATGACTCCTGGCGAG
GC
>LGK001000006.1/133282-133345 Thermanaerothrix daxensis strain GNS-1 contig_4, whole genome shotgun sequence.
GATCTCGCTGGCGATGAGGCTCGCCTGAAACCGTCGGTTCACGACTGATAGCCTCTGCCC
CTGA
>MBSV01000072.1/130026-130087 Clostridium sp . W14A NODE_4, whole genome shotgun sequence.
GTAAATACCGGTGATGAGGCTCACCGTAATCGCATCTTAATGCTGACGGCTTCTGCTTTA
AG
>CP005587.1/724703-724643 Hyphomicrobium denitrificans 1NES1, complete genome .
TCGCGCGATGGGGATGGAGTCCCCCGATAACCGCCATAAGGCTGATGGCTCCTACCGAGC
G
>ADKM02000108.1/1643-1570 Ruminococcus albus 8 contig00006, whole genome APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
shotgun sequence.
ATTTCATACGGGAATGATGTTCTCCCACGGGAAACCGAAACCGCTTATCATAAGCTGATG
ACTTCTGCGTTTTA
>LDX001000001.1/181398-181468 Deltaproteobacteria bacterium CSPl-8
XU12_C0001, whole genome shotgun sequence.
AGGAAAATAGGCGATGGAGTTCGCCGAAAAAACCGCCTGTCCTTCCGACGGCTGATGACT
CCTGCGACGGA
>AFGF01000078.1/30229-30167 Acetonema longum DSM 6540 Contig00078, whole genome shotgun sequence.
TACAATATCGGCGATGGAGTTCGCCTTAAACCCCGTGCAAACGGTAATGACTCCTACCAC
TGA
>LDOV01000006.1/43313-43245 Photobacterium aphoticum strain DSM 25995 contig0006, whole genome shotgun sequence.
GCCCGTGCGGGAGATGATGTTCCTCCTTTAACCGCCCGTTTCATTAGGGATGATGACGTC
TAACAACAA
>CP001962.1/1044224-1044286 Thermus scotoductus SA-01, complete genome.
TGTCCTCGAGGCGATGGGGCCCGCCTGAAACCGCCCAGCCGGGCTGATGGCTCCTACCCA
GAA
>BA000040.2/2887392-2887315 Bradyrhizobium japonicum USDA 110 DNA, complete genome.
CGCACAGATGGGGATGGAGTCCCCCGACAACCGCCCGGCGGTAGGTGACCGTGATGGGCT
GATGACTCCTGCTTGAGG
>LBMM01008604.1/1742-1680 Lasius niger Lnig_2.1_8699, whole genome shotgun sequence .
CGGCACGCTGGAGATGACATTCCTCCCTAACCGCCTTCACAGGCTGATGATGTCTACGTA
ACC
>JMSP01000269.1/1616-1545 Bacteria symbiont BFo2 of Frankliniella occidentalis strain BFo2 contig416, whole genome shotgun sequence.
CTCTATTAAGGTGATGGTGCTCCACCTTTCTCAACCGCCAAAATTCTCTTGGATGATGAC GCCTGATATATA
>CP003493.1/112663-112734 Acidipropionibacterium acidipropionici ATCC 4875 chromosome, complete genome.
TGTGGGCAGGGCGATGGAACCCGCCCGGGGATTCTCCCGAACCGCCACTCGGCTGATGGT
CCCTGCACGTCG
>LMSL01000040.1/127758-127688 Frateuria sp . Soil773 contig_8, whole genome shotgun sequence.
CAAACCAAGGGAGATGGCATTCCTCCCGGGCCGTGCCCAATCGCCGCAAGGCTGATGATG
CCTATCGAACG
>KE993512.1/158596-158668 Clostridium sp . ATCC 29733 genomic scaffold Scaffoldl, whole genome shotgun sequence.
AGAAGATTGGGGAATGAAGTGCTCCCCCGGGCAACCGAAACCGCTTTTGTAAGCTGATGA
CTTCTGTGATTTT
>FNDN01000013.1/66675-66751 Rhodococcus triatomae strain DSM 44892 genome assembly, contig: Ga0116918_113
AATCGTGACGGCGATGGATCTCCGCCGGGGCACCTGCCCGAACCGCCTCACCCGAGGCTG
ATGGCTCCTTTCCCACC
>CGIG01000001.1/172208-172268 Brenneria goodwinii strain OBR1 genome assembly, contig: Contig
GTGCAAAACGGAGATGACATTCCTCCATAACCGCCATTCGGCTGATGATGTCTATGCACT
T
>ADNT01000040.1/13066-12970 Aerococcus viridans ATCC 11563 = CCUG 4311 strain ATCC 11563 contig00050, whole genome shotgun sequence.
ATAAAAAGAGGGAATGAGGTTCTCCCTGGTCAAGTAGATGGTGATTTTAGGCCATCCCAT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ACCGAACCGCTGAAAAGCTGATGACTTCTGCAAGTAA
>MAW01000020.1/15342-15274 Thermodesulfovibrio sp . N1 contig00020, whole genome shotgun sequence.
ATTAAAAAAGGCGATGGAGTTCGCCTTTAAGTGCCTCTGATTTTCAGGGCTGATAACTCC
TACTTGAAA
>LHOX01000015.1/364503-364576 Enterococcus sp. RIT-PI-f contig_2, whole genome shotgun sequence.
CGTAAATAAGGGAATGAATGTCTCCCTCAGTGATGCTGGAACCGCAAATGATTGCTAATG
ACTTCTACAGGCTT
>JGZI01000010.1/76150-76242 Bifidobacterium psychraerophilum strain LMG 21775 ContiglO, whole genome shotgun sequence.
GGCGTGGTTGGGTATGAAGTTACCCTAGGCACCAAAATGAAACACAGTACATACGCCTGA ACCGCTCAACGAGCTGATGACTTCTGCAGAGAT
>GL892006.1/282867-282935 Dysgonomonas mossii DSM 22836 genomic scaffold supercontl .3, whole genome shotgun sequence.
CGAAATTAAGGGAATGGGACTTCCCTGTTATAAACCGCTAATAAAGTAGCTGATAGTTCC
TGCTTAACT
>MKSE01000017.1/211467-211528 Altererythrobacter sp . 66-12
SCNpilot_expt_1000_bf_scaffold_653, whole genome shotgun sequence.
GCCCCCGCAGGGAATGGGGTCTCCCTTGTAACCGCGATCGTGCTGATGACTCCTGCTTTA TC
>AGIZ01000001.1/283970-283896 Fischerella sp . JSC-11 ctgll8, whole genome shotgun sequence.
AAGAGTATCGGTGATGAAACTCACCATTTGGTTTTTCCAGAGAACCACCTCCGGGCTGAT
AGTTTCTAGCTCTCA
>JH379030.1/280351-280280 Clostridium hathewayi WAL-18680 genomic scaffold supercontl .4 , whole genome shotgun sequence.
TGGGCGGGAGGGAATGAAGTTCTCCCGGGGGAAACCTGAACCGCTTTTTGAGCTGATGAC
TTCTGTGATATG
>FYEW01000001.1/407799-407873 Hymenobacter gelipurpurascens strain DSM 11116 genome assembly, contig: Ga0170436_ll
CCATAACCCGGTGATGGGGTACCACCATAAACCGCCGAAGCTCCAGGGCTTTGCGCTGAT
GACTCCTGCGTTGCC
>LFLF01000036.1/26287-26363 Candidatus Burkholderia calva strain UZHbot6 BUMBctg_36, whole genome shotgun sequence.
TGTAACATAGGAGATGGCATTTTCCTGTGGTTCTCCCGCTAATCGCTCCTCGATGAGCCG
ATGATACCTACCCACAC
>KK106988.1/2431313-2431233 Streptomyces sp . Tu 6176 genomic scaffold scaffoldOOOOl, whole genome shotgun sequence.
CTCGGACTCGGCGATGGATCCCGCCTGGGCCGTTCTCCGGCCCGAACCGCCCTCCTCCGG
GCTGATGGTTCCTGACCGACG
>CM001403.1/3147625-3147561 Mucilaginibacter paludis DSM 18603 chromosome, whole genome shotgun sequence .
TTAAACAAAGGTGATGGTGTTCCACCTAATTGAACCGCTGCAAAGCTGATGACGCCTGCC
AAATA
>MHBH01000023.1/31060-31127 Lentisphaerae bacterium GWF2_49_21
gwf2_scaffold_l 613 , whole genome shotgun sequence.
GATGATCTAGGCGATGAGGTTCGCCTTATAACCGTGGCTCGAATGCCTCTGATGACCTCT ACCTTTGC
>JTJC02000078.1/5340-5266 Scytonema millei VB511283 scaffold_192, whole genome shotgun sequence.
TGAAACATAGGTGATGAGACTCACCTTCTGGAATCTCCAGAGAACCGCCTCCTGGCTGAT
AGTCTCTACCTCAAT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>BAND01000009.1/6158-6221 Acidomonas methanolica NBRC 104435 DNA, contig: Amme_009.
ATTCGGATAGGAGATGGGGTTCCTCCTTTTAACCGCCGCATCGGCTGATGACTCCTGGCA
TGAC
>CP014773.1/1898664-1898728 Mucilaginibacter sp . PAMC 26640, complete genome .
TAAAAACAAGGCAATGGTGTTCTGCCTATTTGAACCGCTGCAAAGCTGATGACGCCTACC
CAAAT
>AE005176.1/274025-274097 Lactococcus lactis subsp. lactis 111403, complete genome.
ATAGTAGATGGGTATGGTGCACACCCGAAACCGTTTCAAGAAATAGCTTGAATCTAATGG
CGCCTACAAAGCA
>MKWA01000009.1/237030-237116 Rhizobiales bacterium 64-17
SCNpilot_cont_500_bf_scaffold_256, whole genome shotgun sequence.
CGCCGCGATGGGGATGGGGTTCCCCCGAAACCGCCCAACGCCGCGTCCAGAACGATTGGC GTAAGGGCTGATGACTCCTGCCGAACA
>CP003261.1/2076238-2076298 Clostridium pasteurianum BC1, complete genome.
TAAGTCATAGGTGATGGAGTTCGCCTTTAAATGCTTGAATGCTGATGACTCCTACAGGTC
T
>MKWJ01000020.1/50891-50816 Sphingomonas sp . 67-41
SCNpilot_cont_300_bf_scaffold_301 , whole genome shotgun sequence.
GCCGATCACGGCAATGGATTCCTGCCGGGCTCGTCCGAGCCGAACCGCCCGCAAGGCTGA TGATTCCTACCTGTCG
>MKRJ01000017.1/40177-40115 Rhizobiales bacterium 65-79
SCNpilot_cont_1000_bf_scaffold_222 , whole genome shotgun sequence.
AGGACAACAGGGGATGGAGTCCCCCTCTGAACTGCCGAAGAGGCTGATGACTCCTGCCGC GAC
>AEXE01000084.1/23803-23733 Burkholderia sp . TJI49 contig072, whole genome shotgun sequence.
TGAACAAGCGGAGATGGCATTCTCCCTGAACCGCCGTGCCTTCAGGCGCGGCTGATGATG
CCTACCAGTCC
>MLBF01000005.1/61041-61101 Desulfosporosinus sp . OL contig00005, whole genome shotgun sequence.
GTTATTTCAGGTGATGGGGCTCACCTTATAAATGCCAACAGCTAATGGCTCCTACTTTCA
A
>CH399130.1/3910-3972 Oryza sativa Indica Group scatfoldOOl 625 genomic scaffold, whole genome shotgun sequence.
CGGCACGCTGGAGATGACATTCCTCCCTAACCGCCTTCACAGGCTGATGATGTCTACGTA
ACC
>FRCM01000010.1/62473-62411 Mucilaginibacter sp . OK098 genome assembly, contig: Ga0066759_110
CTTGCATCAGGAAATGGTGTCTTCCTTTTAAAAACCGCATTCGCTGATGACGCCTGCAAC
ATC
>JXOF01000156.1/24392-24329 Bradyrhizobium elkanii strain UASWS1015
Contigl56, whole genome shotgun sequence.
AATATAGATGGGGATGGAGTCCCCCGATACCCGCCCGTCGTGGGCTGATGACTCCTGCTG
GGAT
>FQYN01000002.1/205852-205779 Hymenobacter daecheongensis DSM 21074 genome assembly, contig: EJ57DRAFT_scaffold00002.2
CCTGCTTACGGTGATGGATTTCCACCGCGAACCGCCGGTGCCCCGCGCCCCGCGCTGATG
ATTCCTACGCAACG
>MTKC01000003.1/5840543-5840614 Streptomyces griseofuscus strain NG1-7 NG1-17_1_1, whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CTCGACCTCGGCGATGGATCCCGCCTGGACCTCGGTCCGAACCGCCCCCGGGCTGATGGT
TCCTGACCGACG
>MDDV01000059.1/288092-288017 Ensifer sp . LCM 4579 contig_7, whole genome shotgun sequence.
TGGCACGTCGGCAATGGATTCTGCCAGGCATCAAGCCGAACCGCTTCCCCACGAAGCTGA
TGACTCCTGCTCATGA
>AQHN01000001.1/36899-36812 Rhizobium freirei PRF 81 Cont658, whole genome shotgun sequence.
AAATTCAACGGTAATGGATTTCTGCCGGGCCATGGTGGCCGAACCGCTGCTCTGATTTTC
CAGGGAAGCTGATGACTCCTACTCAAAA
>MKSF01000006.1/50415-50348 Bacteroidetes bacterium 47-18
scnpilot_p_inoc_scaffold_176, whole genome shotgun sequence.
GATAAAAAAGGAAATGGTGTGCTTCCTTACCCAACCGCTTTCCACAAGCTGATGACGCCT GATCATTA
>HF997767.1/2167-2244 Roseburia sp. CAG:18 genomic scaffold, scf259
TCTGTCAATGGGAATGAAGTTCTCCCAAGGGAATTTCCCTGAACTGCTGAACGATAAGCT
GATGACTTCTACGATTGT
>LLWF02000032.1/24311-24373 Roseomonas mucosa strain SAVE376 contig00032, whole genome shotgun sequence .
GTTCCCGATGGGGATGGGGTTCCCCCGAAACCGCCCAGATGGGCTGATGACTCCTACCGC
AAC
>AWQS01000020.1/21083-21159 Intrasporangium chromatireducens Q5-1 contig28, whole genome shotgun sequence.
TGTCGCCACGGCGATGGATCCCGCCGGGACGCCACGCGCGCCCGAACCGCCACCCGGCTG
ATGGTTCCTGTTCATCA
>AM747721.1/913754-913857 Burkholderia cenocepacia J2315 chromosome 2, complete genome
GCGGCCTGCGGAGATGGCATGCCTCCCTGCGGTCGTTGCCGGCGCGTAGCGCGCGGTGCG
GCCCTCAACCGCCGGTTTGCCCGGCTGATGATGCCTGCGTGTTC
>LJAJ01000001.1/63803-63865 Geobacillus sp . BC02 LR69_contig000001, whole genome shotgun sequence.
AATCGAATAGGCGATGGAGTTCGCCATAACCGCCGGCTTCCGGCTGATGACTCCTGCTGC
AGA
>FTM001000001.1/123874-123811 Bosea sp. TND4EK4 genome assembly, contig: Ga0136738_101
TGAGCCGATGGGGATGGGGTTCCCCCGAAACCGCCCTTGAAGGGCTGATGACTCCTGCCG
AGCC
>FR899336.1/4361-4285 Eubacterium sp . CAG:38 genomic scaffold, scf36
ATAGAACAAGGGAATGAAGTACTCCCGAAGTGAAAAACACTTAAACCGCTTATTAAGCTG
ATGACTTCTGCGATGAA
>MEQV01000269.1/20307-20375 Betaproteobacteria bacterium
RIFCSPLOWO2_02_FULL_62_17 rifcsplowo2_02_scaffold_6838, whole genome shotgun sequence.
TGATTGTAGGGAGATGGCATACCTCCCGCAAGAATAACCGCCGTCTCGGCTGATGATGCC
TACAGGATC
>GG657557.1/343061-343134 Holdemania filiformis DSM 12042 genomic scaffold Scfld6, whole genome shotgun sequence.
TAAAGGACAGGGAATGATTTCTCCCCTGGCAATCCGCCTAAACCGCTGTTTCAGCTGATG
ACTTCTATGTGTTT
>CP001338.1/1058733-1058667 Candidatus Methanosphaerula palustris El-9c, complete genome.
CTTCGTCAGGGTGATGGGGTTCACCCACAATAAACCGCGTTTCTTCGCTGATGACCCCTA
TCCATCA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>FMAF01000003.1/33283-33195 Rhizobium lusitanum strain Pl-7 genome assembly, contig: Ga0061101_103
AAATCCAACGGTAATGGATTTCTGCCGGGCCATGGTGGCCGAACCGCTGCCTTGAATTTT
CCAGGGAAGCTGATGACTCCTACTCAAAG
>LDP001000008.1/5770-5695 Mycobacterium heraklionense strain Davo contig_8, whole genome shotgun sequence.
CTTGAGCTCGGCGATGGGTCCCGCCCGGAAGCTTTGACTTCCGAACCGCCACACGGCTGA
TGGCCCCTGCGAACGA
>MHZC01000921.1/589-649 Planctomycetes bacterium RIFOXYB12_FULL_42_10 rifoxyb3_full_scaffold_5445_curated, whole genome shotgun sequence.
ACAAATCAAGGCGATGGAGTTCGCCTAACGCTACGTTGTAGCTGATAACTCCTATTGAAG
C
>MNUR01000060.1/8249-8328 Elusimicrobia bacterium CG1_02_56_21
cgl_0.2_scaffold_1755_c, whole genome shotgun sequence.
TACAGTCAAGGCGATGATGTCCGCCGCCCCGTGACCGCACGGGGGAACCGTCCGAGGGGA CTAATGACCTCTACCTTAAG
>FXAU01000001.1/1275664-1275725 Sphingobacterium psychroaquaticum strain DSM 22418 genome assembly, contig: LY64DRAFT_scaffold00001.1
TTAACTATAGGCAATGATGTCTGCCTTGAACCGCTACTGGAGCTGATGACGTCTACTTTG AA
>LVHF01000010.1/260-325 Photobacterium jeanii strain R-40508 C1841, whole genome shotgun sequence.
CCGATCACGGGTGATGGAGTTCCACCTTTAACCGCTCATCATTGAGATAATGACTCCTGC
TGTAGT
>CP002028.1/579265-579329 Thermincola sp . JR, complete genome.
TACAAAAAAGGCGATGGGGCTCGCCTTGAACCGCTGTTTTTACAGATGATGGCTCCTACT
GGCTT
>LT629759.1/1456296-1456370 Olsenella umbonata strain DSM 22620 genome assembly, chromosome: I
TGTTCCGTGCGGAATGAGGTTCTCCGCGGGGCAGAAGCCCCAAACCGCCTTCGGGATGAT
GACCTCTGCCACTGT
>CP001130.1/523442-523357 Hydrogenobaculum sp. Y04AAS1, complete genome.
CAGGTTTAAGGAGATGGCGTTCTCCTTGATGCTCATAAAAGCATCAAAACCGCCTCACAT
ATAAGGCTAATAACGCCTACCTGCTC
>CP002026.1/3910535-3910474 Starkeya novella DSM 506, complete genome.
CCTGTCTGCGGGAATGGTGTCTCCCGTAAAACCGCGCTGATGCTGATGACGCCTGCTTTT
CG
>JZRB01000011.1/34318-34378 Luteibacter yeojuensis strain SU11 contig_ll, whole genome shotgun sequence .
GCGGCCGATGGGAATGGGGCTTCCCGATAACCGCCAGTTGGCTGATAGCTCCTGCCGTAG
T
>FTOR01000005.1/380028-380097 Filimonas lacunae strain DSM 21054 genome assembly, contig: Ga0111651_105
TCGCAGTAAGGCGATGGTGCTCCGCCTTATCAACCGCCCTAACCAGAGGGACGATGGCGC
CTGCAATAAC
>CH902600.1/668074-668141 Vibrio angustum S14 1099604003226 genomic scaffold, whole genome shotgun sequence.
TTTATGGCGGGAGATGATGTTCCTCCTTCAACCGCCTTTCAATAAAGGATAATGACGTCT
AACAACAT
>CP000885.1/3806199-3806127 Clostridium phytofermentans ISDg, complete genome .
CTTACTGACGGGAATGAAGTTCTCCCAAGGTTTTACCTAAACTGCTTATAAAGCTGATGA
CTTCTGTGAATTA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>MIHC01000025.1/20762-20822 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
CGCGGTTACGGCGATGAAGCTCGCCGCAACCGCCGCGCCGGCTGATGGCTTCTACCAACG
A
>MKTT01000036.1/28217-28147 Planctomycetales bacterium 71-10
scnpilot_p_inoc_scaffold_l 61 , whole genome shotgun sequence.
TTCGCGGAGGGTGATGGAGTTCACCCGACGATTGGACCGCCCCTCGCGGGGCTGATGACT CCTACCTGGAG
>FR894075.1/2447-2525 Roseburia sp. CAG:309 genomic scaffold, scf388
ATAGAACAAGGGAATGAAGTTCTCCCTAAGTAGTTAGATACTTAAACCGCTTATAAAAGC
TGATGACTTCTGCGATATT
>CP002780.1/1695616-1695675 Desulfotomaculum ruminis DSM 2154, complete genome .
ACAAATAATGGCGATGGAGCTCGCCATAACCGCCCTTGGGCTAATGGCTCCTACCTAGAA
>CM001167.1/2015568-2015502 Bacteroides coprosuis DSM 18011 chromosome, whole genome shotgun sequence .
AAGCAAATAGGGAATGGAACTTCCCTGTCATAAACCGCTAATAATAGCTGATAGTTCCTA
CGATATC
>DF820472.1/161849-161788 Bacterium UASB270 DNA, scaffold:
UASB270_scaffold_l 0.
ATATTTTCAGGTGATGGAGTTCACCTCGAACCGCATGTTCTGCTAATAACTCCTACCAGT
TT
>MGDI01000031.1/92400-92337 Candidatus Schekmanbacteria bacterium
RIFCSPL0W02_12_FULL_38_15 rifcsplowo2_12_scaffold_364 , whole genome shotgun sequence.
TTTAAAAAAGGCGATGGGGTTCGCCATTAAGCATCCATAAGGGATTAATGACCCCTACTG
AAAA
>AP008230.1/2441450-2441510 Desulfitobacterium hafniense Y51 DNA, complete genome .
AATCAGATAGGCGATGGAGTTCGCCTTTAACCGCCGCTTGGCTAATGACTCCTGCCGGTT
A
>MAPZ01000011.1/252388-252329 Clostridium paraputrificum strain 373-A1 CP373A1_19, whole genome shotgun sequence.
AATAATAAGGGCGATGGAGTTCGCCATTAAATGCTTTATGCTAATGACTCCTACAAAGTA
>CP001056.1/2720217-2720276 Clostridium botulinum B str. Eklund 17B, complete genome.
ATATATTTAGGTGATGGAGTTCACCTTTAACTGCGTAATGCTAATGACTCCTACAAAAAA
>MCAP01000024.1/14886-14802 Salinicola sp. MIT1003
NODE_31_length_23181_cov_92.2337_ID_61, whole genome shotgun sequence.
CTTGCCGTGGGAGATGGCATTCCTCCCGAGCGGCCTTCTCGTTCGCTCGAACCACCGTCT
ATCGGTTAATGATGCCTGCGGGCGC
>MHD001000021.1/4226-4165 Nitrospinae bacterium RIFCSPL0W02_12_39_16 rifcsplowo2_12_subl0_scaffold_141, whole genome shotgun sequence.
TTAAAAGGAGGTGATGGAGTTCACCTTTAACCGCCTTAATGGCTGATAACTCCTACAAAT TT
>HF990157.1/9033-9105 Ruminococcus sp . CAG:60 genomic scaffold, scfl74
AATAAAAAGGGGAATGAATGTCTCCCCGGGATCATCCCGAACCGCTTATTAAGCTGATGA
CTTCTGTGATAAC
>FR897630.1/11154-11065 Anaerotruncus sp . CAG:390 genomic scaffold, scfl73
TATCGCCGAGGGAATGAAGTTCTCCCGAAGCACCGGGGCATAAGCCCCGTGCCTGAACCG
CTTATAAAAGCTGATGACTTCTGCATTATT
>CP003137.1/1745706-1745645 Pediococcus claussenii ATCC BAA-344, complete genome . APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AATCGAATCGGCGATGACGTTCGCCACAAAATTATAATTTAATTGATGACGTCTACCACA
AG
>LWSA01000081.1/1744-1681 Acidithiobacillus thiooxidans strain A02 contig081, whole genome shotgun sequence.
ACAAATTATGGCGATGGATTCCGCCCAATAACCGCCTTTATAGGCTAATGACTCCTACTG
ATCC
>CP002329.1/931000-931061 Mycobacterium sp . JDM601, complete genome.
GTTGATGGCGGCGATGAAGCTCGCCATGAACCGCCGCCCCGGCTGATGGCTTCTACCACG
TG
>DF820472.1/161032-160966 Bacterium UASB270 DNA, scaffold:
UASB270_scaffold_l 0.
GATGAGTAAGGCAATGAAGTCTGCCTGAACCTCCTGGAATTTCCAGGATGATGACTTCTA
TTGTATA
>AFWD01000095.1/32709-32782 Vibrio sp . N418 VIBRN418_89, whole genome shotgun sequence.
TGCCTACACGGTGATGGGGTTCCACCGTGGCCTTTCGGTCAAAACCGCACGATGCTGATG
ACTCCTACAGAAAC
>FOOE01000004.1/180052-180111 Clostridium cadaveris strain NLAE-zl-G419 genome assembly, contig: Ga0070261_104
AAATATAAAGGTGAAGGAGTTCACCTATAAATGCTATAAGCTGATGACTCCTATAAGATA
>MGTV01000029.1/111483-111406 Elusimicrobia bacterium GWA2_62_23
gwa2_scaffold_252 , whole genome shotgun sequence.
TTGAACCACGGCGATGATGTCCGCCGCCCCGGCCCCCCGGGGGAACCGCGCGAAGGCGCT GATGACCTCTACCCCTCA
>MUMY01000002.1/148269-148183 Nocardia donostiensis strain X1655 2, whole genome shotgun sequence.
CTTGTCGTCGGCGATGGATCTCCGCCAAGGCACCCCAGGTGGTGTCTGAACCGCCTCGCT
GCCGGGGCTGATGGTTCCTTTCTCGTC
>LJHO01000001.1/467664-467576 Novosphingobium sp . AAP1 AAPIContigsl , whole genome shotgun sequence.
CGCGGCCAAGGCGATGGATTTCCGCCGGGCATTTTGCCGAACCGCCTGCCGCGCCAAACG
GCGCCAGGGCTGATGATTCCTACCTGAAA
>CP013438.1/500555-500662 Burkholderia latens strain AU17928 chromosome 2, complete sequence.
GCCGCCTGCGGAGATGGCATGCCTCCCTGCGGTCGAGTCCGGCGCGTCGCGCGCGCGGTC
TTCGGCCCTCAACCGCCGGTTCGCCCGGCTGATGATGCCTGCGTGTTC
>BBJM01000037.1/4345-4282 Lactobacillus oryzae DNA, contig: sequence37, strain: SG293.
AAATGAATCGGCGATGACGTTCGCCATTAATAATTGACTATCAATTGATGACGTCTACTG
AATC
>AYZV02278730.1/438-505 Spinacia oleracea cultivar SynViroflay
scaffold91480. con0027.1, whole genome shotgun sequence.
TGAACAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTACAAAAGCTGATGACGCCT GATTAAAA
>MHGQ01000025.1/13844-13776 Omnitrophica WOR_2 bacterium
RIFCSPHIGH02_02_FULL_50_17 rifcsphigho2_02_scaffold_1877 , whole genome shotgun sequence.
ATTTTAAAAGGCGATGAGGTTCGCCAAGACTGCCCGGGGATAAGTCGGGATGATAACCTC
TGCTGGAAA
>ATMD01000004.1/321499-321572 Thermoplasmatales archaeon I-plasma
AMDU3_IPLC00004 , whole genome shotgun sequence.
CATTTGCAAGGAGATGGAATCCTCCAGCGCTCTCTGGGCCGAACTTTCTAATGAATGATG ATTCCTGCTTGGAG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>LISW01000001.1/1383979-1384043 Streptomyces sp . CB01249 scaffoldl, whole genome shotgun sequence.
ATGCGTACCGGTGATGGGGCTCACCGAAACCGCGGCCGATGGCCGCTGACGGTCCCTGGT
TGTTT
>GG700527.1/977325-977253 Anaerococcus vaginalis ATCC 51170 genomic scaffold SCAFFOLD1, whole genome shotgun sequence.
ATACTAAAAGGGAATGAAGTTCTCCCTTAGATTATTTCTAAAACCGCATTTTGCTGATGA CTTCTGTTACTAT
>FRAF01000007.1/ 62851-62790 Alicyclobacillus sp . USBA-503 genome assembly, contig: Ga0105849_107
AACTGAACAGGTGATGGAGCCCACCTTAACCGCCTCTATAGGATAATAGCTCCTACAAAT
TT
>CP018082.1/3527307-3527394 Nocardia soli strain Y48, complete genome.
CTCTTCCGTGGCGATGGATCTCCGCCGAGGCAGCCCCGTGGGAGTGTCTGAACCGCCCCG
GACCGGGGCTGATGGTTCCTTCCCTCGC
>LVYD01000065.1/48506-48581 Niastella vici strain DJ57 contig36, whole genome shotgun sequence.
CACCACAAAGGAAATGGTGTCTTCCTATTTGAACCGTCCCGATGATTAACTAAGGACTGA
TGGCGCCTACAAATAG
>MIHC01000025.1/53604-53666 Mycobacterium sherrisii strain BC1_M4
PROKKA_contig000025, whole genome shotgun sequence.
CCCGCGGCTGGCGATGAGGCTCGCCACTTTCGCGCTTATGGCGCTGATGGCTTCTACCCC GCA
>ABCM01000022.1/54798-54866 Pedobacter sp. BAL39 1103467000501, whole genome shotgun sequence.
TTGCAATACGGAAATGGTGTCCTTCCGATTCAACCGCTTTTTTCAAAAGCTGATGGCGCC
TACAATAGA
>KB849533.1/123115-123180 Acinetobacter gerneri DSM 14967 = CIP 107464 genomic scaffold acLZs-supercontl .27 , whole genome shotgun sequence.
TGAAGCTTAGGAGATGGTATTCCTCCCATTACAAACCGCCCTTAGGCTGATGATACCTAC GTCTTC
>FR903987.1/154-80 Roseburia sp . CAG:182 genomic scaffold, scf226
TTTAGAAAAGGGAATGAAGTTCTCCCTCGATTGATTTCGAAACCGCTGATTCAGGCTGAT
GACTTCTGTGCATTT
>MGBT01000235.1/2204-2273 Candidatus Rokubacteria bacterium GWA2_70_23 gwa2_scaffold_94712, whole genome shotgun sequence.
CGGCCACAGGGCGATGGAGTTCGCCCTCCCCAACCGCCTGGCGTCACGGGCTGATAACTC CTACCGAAGC
>LMET01000002.1/86989-87060 Aeromicrobium sp . Root472D3 contig_2, whole genome shotgun sequence.
ATCGACCTTGGCGATGGACCCCGCCTGAGTTTCAGACTCGAACCGCATTGTGCTGATGGC
TCCTGCTATTCG
>FR880840.1/41178-41249 Clostridium sp. CAG:510 genomic scaffold, scfl06
ATGATTTCAGGGAATGAAGTTCTCCCTTGGGAAACCTAAACCGCTTAGTAAGCTGATGAC
TTCTGTGATGCA
>JAOL01000124.1/37066-37140 Mycobacterium ulcerans str. Harvey
Harvey . contig .123 , whole genome shotgun sequence.
ACTAGCAGCGGCGATGGGGCTCGCCAGGAAGCTCGACTTCTGAACCGCCAGCCGGCTGAT GGCACCTGCGAACAA
>ABCM01000001.1/151944-151879 Pedobacter sp . BAL39 1103467000516, whole genome shotgun sequence.
TTGCCTAAGGGAAATGGTGTCCTTCCAATTTAACCGCTTTTAGAAGCTGATGGCGCCTGC
AATGAT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>MRCT01000006.1/6591-6503 Methanobrevibacter sp . 87.7 contig6, whole genome shotgun sequence.
TTATAATTTGGTGATGAGGTTCGCCATGTATTAATTTACAGAACCGTTAAACTTTTTTTA
AGATTTAACTGATGACTTCCTATTTAACT
>AGQV01000005.1/14506-14586 Gluconobacter morbifer G707 75_5, whole genome shotgun sequence.
TTTTGTAGTGGCAATGGACTCTGCCTGACCTCATTACGTCGAACCGCGATCCTTCGGGGC
GCTGATGATTCCTACCTCGCA
>CP003504.1/1769319-1769386 Enterococcus hirae ATCC 9790, complete genome.
AAGTGAATAGGTGATGGTGTTCGCCTTTAAACGTTAATCTTAAGATAACTAATGACGCCT
ACTAAGAA
>FQVH01000038.1/5568-5626 Caldanaerobius fijiensis DSM 17918 genome assembly, contig: EJ21DRAFT_scaffold00038.38
TATAAATATGGTGATGGGGCTCACCTTAAATGCGAAATGCTAATGGCTCCTACTTTGAT
>KE351688.1/3523-3456 Enterococcus faecalis 13-SD-W-01 genomic scaffold
Scaffold959, whole genome shotgun sequence.
AAGGAATAAGGTGATGGTGTTCGCCTTTAAACGTTGATCGAAAGACAGCTAATGACGCCT
ACTAGAGA
>FQV001000006.1/133983-133915 Chryseobacterium takakiae strain DSM 26898 genome assembly, contig: Ga0131160_106
CTAACAAAAGGAAATGGTGTTCTTCCTTACCCAACCGCTTTTCAGAAAGCTGATGACGCC
TGATTAAAA
>MHBJ01000043.1/8102-8172 Lentisphaerae bacterium GWF2_52_8
gwf2_scaffold_14617 , whole genome shotgun sequence.
AAATAAAAAGGCTATGAGGTCAGCCGAAACCGTCTGAGCGAAGAACTCGGAATGATGACT TCTATTTAGGG
>JPME01000010.1/357217-357144 [Clostridium] celerecrescens strain 152B clO, whole genome shotgun sequence.
TTGAGAAAAGGGAATGAAGCTCTCCCATGATAGAATTATCAAAACCGCAATTTGCTGATG
GCTTCTGTATATTT
>JSZN01000007.1/281463-281358 Burkholderia sp. A9 Contig7, whole genome shotgun sequence.
GCGCCCTGCGGAGATGGCATGCCTCCCTGCGGTCGATCCCGGCGCGTGGCGCGTGCGGTG CGGCCCTCAACCGCCGGGTTGCCCGGCTGATGATGCCTGCGTGTTC
>MGRA01000126.1/1306-1246 Deltaproteobacteria bacterium
RBG_19FT_COMBO_46_12 rbg_19ft_combo_scaffold_1869, whole genome shotgun sequence .
GATTGATGTGGCGATGGAGTCCGCCGGAACCGCCTTCTTGGCTGATGGCTCCTATCAAAN
N
>CYHF01000010.1/53222-53296 Thiomonas bhubaneswarensis strain DSM 18181 genome assembly, contig: Ga0061069_110
ACATCAAACGGCGATGGAGTTCGCCGTTCAAATGCCTCTGAGGGCTTGCCCCGGGCTGAT
GACTCCTACCCGTCT
>CP011568.2 /I 647342-1647269 Pandoraea thiooxydans strain DSM 25325, complete genome.
CTCGCGGCAGGAGATGGCATGCCTCCTTTAACCGCCACCGGCCTCGTGCCGGGGCTGATG
ATGCCTACGCAACC
>LZLV01000067.1/32325-32264 Mycobacterium sp . 1245111.1 contig_159, whole genome shotgun sequence.
TCCCGTTCAGGCGATGAAGCCCGCCTTCAACTGCCGCAACGGCTGATGGCTTCTACCCGT
CC
>CP006721.1/3419704-3419763 Clostridium saccharobutylicum DSM 13864, complete genome. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TATATTTCAGGTGATGGAGTTCGCCTTTAACTGCGTAATGCTAATGACTCCTACAAAATA
>CP006936.1/2090582-2090495 Mycobacterium neoaurum VKM Ac-1815D, complete genome .
CTCGACGACGGCGATGGATCTCCGCCGAGACGCCCGGTGACAGGTGTCTGAACCGCCTCG
TCCGGAGGCTGATGGCTCCTTTCCCCGA
>FR882595.1/181344-181419 Firmicutes bacterium CAG:145 genomic scaffold, scf50
ATTGCTTAAGGGAAAGAGGTTCTCCCATGGCGTGACTTGCCTGAACTGCCGAAAGGCTGA
TGACTTCTGCAATCGC
>MEKR01000017.1/22217-22145 Acidobacteria bacterium
RIFCSPLOWO2_02_FULL_61_28 rifcsplowo2_02_scaffold_11887, whole genome shotgun sequence.
ATGAATCTTGGCGATGGAGTTCGCCCTCAACCGCCCAGAGGACTGCTGACGGGCTGATGA
CTCCTACCTAACT
>LVJU01000016.1/2690-2587 Rhodanobacter sp . FW510-R10 contigll2, whole genome shotgun sequence.
TGGCCGCTCGGAGATGGCATGCCTCCCGTTCCGGCACCTTCGGATGCGCCGGGACAAACC
GCTACTCGTGATCCCGCGTTGAAGCTGATGATGCCTGCACCTCC
>LGGF01000044.1/13727-13666 Desulfotomaculum sp . 46_80 MPI_scaffold_814 , whole genome shotgun sequence .
ATTATAAAAGGTGATGGGGTTCCGCCGGAAATACCTGATTGGTTGATGACTCCTGCTATG
GA
>CP009268.1/2643261-2643201 Clostridium pasteurianum DSM 525 = ATCC 6013, complete genome.
GTAAAAATAGGTGATGGAGTTCACCTTTAAGTGCTGAAATGCTGATGACTCCTACAGATT
T
>FZOQ01000002.1/203758-203692 Pontibacter ummariensis strain NKM1 genome assembly, contig: Ga0139000_102
GTATTACATGGTGATGGGGTACCACCAGAAACCGCCTAAAGCAAAGGCTGATGACTCCTA
CATGAAG
>LMKC01000001.1/301789-301865 Sphingomonas sp. Leaf9 contig_l, whole genome shotgun sequence.
GCTCGGCACGGCGATGGATTTCCGCCGGGCCATTCGGCCGAACCGCCCTCGCAAGGGTCG
ATGATTCCTACCAGGCG
>KK073879.1/66063-66124 Aquamicrobium defluvii strain W13Z1 genomic scaffold V9_scaffold_03, whole genome shotgun sequence.
TAAATGGATGGGGATGGAGCTCCCCGAAACCGCCTGAAATGGCTGATGGCTCCTGTCGCA
AC
>CP001720.1/2367796-2367739 Desulfotomaculum acetoxidans DSM 771, complete genome .
TAACTTTATGGCGATGGAGTTCGCCTAATGCTGTGAAGCTAATGGCTCCTACCTGACA
>FRAF01000024.1/35081-35018 Alicyclobacillus sp . USBA-503 genome assembly, contig: Ga0105849_124
AGTGATATCGGTGATGGAGCTCACCAGAAAGCCTCTACAAGAGGATGATGGCTCCTGCAA
AAGC
>MBSV01000068.1/14235-14298 Clostridium sp . W14A NODE_46, whole genome shotgun sequence.
AATTAAAACGGCGATGAAGTTCGCCTTTTATCCCAGCTCCGGGGTTAATGACTTCTATTA
TTTT
>CP002696.1/200047-200162 Treponema brennaborense DSM 12168, complete genome .
AGTTGAATCGGGAATGAAGTTCTCCCGAGGAAAATCCTGAACCGGGACCGGCCGCCGCGC
TGCAGAACGTAACTGCGCACGCACGGCGCGCGGTCTCTGATGACTTCTGTGATTAC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>CP008714.1/2345838-2345768 Xanthomonas translucens pv. undulosa strain Xtu 4699, complete genome.
GCGCCCACAGGAGATGGCGTTCCTCCTTCAACCACTGTGCCGACGGCGCAGTTGATGACG
CCTACAGACAC
>CP003607.1/3969870-3969954 Oscillatoria acuminata PCC 6304, complete genome .
AAGATTGAAGGCGATGGAGCTCGCCATAACCGCCGATTCAGCGATCCCGATGGGAGGCTG
AAGGGCTGATGGCTCCTACCTGTCC
>LCVM01000143.1/4442-4374 Providencia rettgeri strain MR4
P_rettgeri_contig_143 , whole genome shotgun sequence.
AAACCTTTGGGAGATGGCATTCCTCCTTATATAAAACCGCCCGTAGAGGCTGATGATGCC TACGTTAAC
>CP001801.1/1928728-1928661 Halothiobacillus neapolitanus c2, complete genome .
CAACGAAATGGCGATGGAGTTCGCCGTAACCGCCCCCACTTCTGGTGGCTGATGACTCCT
GAAGGCGC
>APWR01000094.1/18921-18860 Paeniclostridium sordellii ATCC 9714 gcs 9714. contig .94 , whole genome shotgun sequence.
AAACAAATCGGCGATGGAGTTCGCCATTAAATGCACAAGATGCTAATGACTCCTACTCTT AT
>FUYH01000022.1/42644-42702 Caloramator quimbayensis strain USBA 833 genome assembly, contig: Ga0105845_122
CGTAATTTTGGTGATGGAGTTCGCCATAAATGCTTTAAGCTGATGACTCCTGCAAATAA
>MENF01000040.1/1522-1464 Bacteroidetes bacterium GWA2_32_17
gwa2_scaffold_14764 , whole genome shotgun sequence.
AAATATTTTGGCGATGGAGTTCGCCATAAACCGCATAAGCTAATGACTCCTACTAATTT
>LJWX01000002.1/320404-320475 Ferrovum sp. JA12 FERRO_contig000002, whole genome shotgun sequence.
TGAGTCACAGGAAATGGCATTTTCCTACTCAACCGCCTGCTTAAAGTCAAGGCTGATAAT
GCCTACAATTTT
>CP015102.1/1424415-1424488 Thermococcus pacificus strain P-4, complete genome .
GGTTCCATGGGCGATGGCGTCCGCCCGGGCTTCGAGCCGAACCGCCCTCATGGGCTGATG
ACGCCTGTTCTCCG
>LMKL01000002.1/3129-3053 Sphingomonas sp. Leafl7 contig_10, whole genome shotgun sequence.
CTCGGGCACGGCGATGGATTTCCGCCGGGCCATTCGGCCGAACCGCCCTCGCAAGGGCCG
ATGATTCCTACCAAGCG
>FZOC01000008.1/20850-20911 Desulfovibrio mexicanus strain DSM 13116 genome assembly, contig: Ga0070557_108
GATGGCTCAGGGGATGGAGTCCCCCGTGAACCGCGACAACCGCTGACGACTCCTGCCGCC
GA
>AOI001000039.1/19717-19789 Natrialba asiatica DSM 12278 contig_39, whole genome shotgun sequence.
CCCCCTTCGGGCGATGGGGCCCGCCCGACCCAACCGCCGGCACCGACCGCCGGCTGACGG
TCCCTGCCACTAC
>GL635751.1/207851-207791 Bacillus sp . 2_A_57_CT2 genomic scaffold supercontl .2 , whole genome shotgun sequence.
AAAATAATAGGTGATGGAGTTCACCTTTAACCGCCTTTAGGCTAATGACTCCTGTCAGTT
G
>LQMP01000030.1/18008-18084 Hadesarchaea archaeon YNP_N21 contig_4, whole genome shotgun sequence.
AAGAAAAAAGGCGATGGAGCTCGCCGTTAACCGCTCCTGCTTTGCATTATGCGGGGGCTG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
ATAGCTCCTACTTTTCG
>CP001739.1/594403-594325 Sebaldella termitidis ATCC 33386, complete genome .
ATTAAAAAAGGAGATGAAGCTCTCCGTAGTATAAAAACTAAAACTACAATTATAGATTGT
TAATAGCTTCTGCAGTATT
>AAJM01000092.1/4654-4595 Bacillus thuringiensis serovar israelensis ATCC 35646 sql863, whole genome shotgun sequence.
ATAATCAGAGGCGATGGAGTTCGCCATAACTGCTGCTTAGCTAATGACTCCTACCAGTAT
>LWHG01000021.1/731273-731341 Haladaptatus sp. R4
NODE_2_length_1049977_cov_26.4246_ID_1027, whole genome shotgun sequence.
GACTAACCGGGCGATGGAGTCCGCCTCCCCCAACCGCCGAACACGTCGGCTGATGACTCC
TTTTCGAGA
>MKVY01000039.1/120888-120949 Rhizobiales bacterium 65-9
SCNpilot_cont_300_bf_scaffold_522 , whole genome shotgun sequence.
CGGGCCGATGGGGATGGAGTCCCCCGAAACCGCCCGAGAGGGCTGATGACTCCTGCCGCA GT
>KE136492.1/346941-347012 Enterococcus columbae DSM 7374 = ATCC 51263 genomic scaffold acyDG-supercont2.2 , whole genome shotgun sequence.
ATACGAACAGGGAATGAAGTTCTCCCTTGGGTAACCAAAACCGCGTAAAACGCTGATGAC TTCTATGGCTTT
>JH597761.1/584494-584415 Metallosphaera yellowstonensis MK1 genomic scaffold MetMKlscaffold_l 0 , whole genome shotgun sequence.
AAGAGGACGGGGGATGGCGTCCCCCTGGGCTCATAGCCCAAACCGCCGGAATCATGAGGG CTGATGACGCCTATCTCATG
>CP002962.1/126944-127008 Emticicia oligotrophica DSM 17448 plasmid pEMTOLOl, complete sequence.
GCCACATTTGGCAATGATGTCTGCCTTGAACCGCCTCATTTGAGGCTGATGGCGTCTGTT
TCTGA
>LOHZ01000033.1/40695-40753 Thermovenabulum gondwanense strain R270
ATZ99_contig000034 , whole genome shotgun sequence.
TAAAAGTATGGTGATGGAGCTCACCATAAAAGCAGTATGCTGATGGCTCCTGCTATATC
>MGVG01000006.1/24302-24229 Elusimicrobia bacterium RIFOXYA2_FULL_53_38 rifoxya2_full_scaffold_138, whole genome shotgun sequence.
TATTGTGGAGGCTATGAGGTTCGCCTTGGATATTTCTGAAACTGCCTTAATGGGCTGAGA ACCTCTACCGTGAA
>AYYX01000035.1/18045-17976 Lactobacillus vini DSM 20605 Scaffold35, whole genome shotgun sequence.
TGAAATAAAGGTGATGACGTTCACCTATTAAATGATCGTAAGCAAATGACCTGATGACGT
CTACTAAACT
>FR885232.1/633-718 Phascolarctobacterium sp . CAG:207 genomic scaffold, scf110
GTAATAAAAGGGAATGGGGTACTCCCTAAATGACTTTAGTCGTTTGAACCGCTTAGAGTT
GATAAGCTGATGACTTCTGCCGTGTA
>MGTD01000007.1/168961-168895 Deltaproteobacteria bacterium
RIFOXYD12_FULL_56_24 rifoxyd3_full_scaffold_128 , whole genome shotgun sequence .
CACAATGATGGCGATGGAGTTCGCCTTTGCGCGCTTCCCCGACGAAGCTAATGACTCCTA
CTCTTGA
>MQUF01000018.1/20436-20372 Desulfobulbaceae bacterium DB1 DBl_contig_18, whole genome shotgun sequence .
AAATATTAAGGCGATGGAGTTCGCCATTAACCGCTTGCCTGCAAGCTGATGACTCCTGCC
CTCCG
>LGHF01000110.1/169-230 Methanomicrobiales archaeon 53_19 APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
MPJ_scaffold_3394, whole genome shotgun sequence.
TCAGATGCAGGCGATGAAGTCCACCTCCATCCGCTGACGAGGCCGATGACTTCTACACCT
CG
>FQVB01000016.1/92143-92211 Desulfacinum internum DSM 9756 genome assembly, contig: EJ47DRAFT_scaffold00014.14
CAGAGAAAAGGCAATGAAGTCTGCCTGAACCACCTGCGCCCGGCGCCGGTTGATGACTTC
TACCTCAGG
>LGHJ01000002.1/27361-27302 Bellilinea caldifistulae strain GOMI-1 contig_21, whole genome shotgun sequence.
ATTCACAAAAGCGATGAGGCTCGCTTTAACCGTCGAATGACTGATAGCCTCTACTGAATC
>NATG01000104.1/9296-9233 Bacteroidetes bacterium 4484_276
ex4484_276_scaffold_6541, whole genome shotgun sequence.
TTCGCCATTGGCGATGGAGTTCGCCATAAAATGTCATAAACGGACTGATGACCCCTTTAA TAAT
>MKWF01000046.1/267110-267182 Sphingobacteriales bacterium 40-81
SCNpilot_expt_1000_bf_scaffold_30 , whole genome shotgun sequence.
TGCGGCAAAGGAAATGGTGTCTTCCTATTTGAACCGTCCCGCTAAAGGCAGGACTGATGG CGCCTACAGATTA
>CP002351.1/1604059-1603999 Pseudothermotoga thermarum DSM 5069
chromosome, complete genome.
AGTTTTATCGGCGATGGAGTCCGCCGTAAAATGCCGTAAGGCTGATGACTCCTACCCAAT
T
>MKJX01000184.1/170251-170189 Pseudonocardia sp . CNS-139
NODE_3_length_171736_cov_120.769_ID_221 , whole genome shotgun sequence. CCGGCCACAGGCGATGGGGCTCGCCTTCGACCGCCCGCCCGGGCTGACGGCTCCTACCGA CAG
>AZGF01000012.1/7088-7152 Lactobacillus suebicus DSM 5007 = KCTC 3549 strain DSM 5007 Scaffoldl2, whole genome shotgun sequence.
AATTAAATAGGCGATGACGTTCGCCGTAATTAAACTGATTGTAATCTGATGACGTCTACT ATTCC
>CP001850.2/690883-690954 Clostridiales genomosp. BVAB3 str. UPII9-5, complete genome.
ATGAAAAAAGGGAATGAAGTTCTCCCTTGGGCAACCTGAACCGCTTTTTAAGCTGATGAC
TTCTGCGATAAA
>LPVJ01000070.1/129704-129768 Acidibacillus ferrooxidans strain ITV001 contig0088, whole genome shotgun sequence.
ACAAAAATGGGTGATGGAGCTCACCCAAACTGCCCTCACCAAGGGATAATGGCTCCTGCG
AATCG
>MEGK01000014.1/76455-76520 Rhodanobacter sp . SCN 66-43 ABT18_C0014, whole genome shotgun sequence.
GGGGTTTCAAGAGATGGCATGCCTCTTTGAACCGCCGGCGATCCGGCTGATGATGCCTAC
CCACCG
>MGQI01000010.1/5321-5400 Deltaproteobacteria bacterium RBG_13_58_19 RBG_13_scaffold_11141, whole genome shotgun sequence.
CTTTAATATGGCGATGGAGTCCGCCGGGGTGGAGGTGGCATCACCCAAACCGCATTGACG CTGATGGCTCCTGATCTCAC
>CP003412.1/181993-182103 Natrinema sp. J7-2, complete genome.
AAACGACCGGGCGATGGGGCCCGCCCGACCCAACCGCCGACAGCGACGTCGACCGCGTCT
GGAGCGCACAGACTTACCGAGCGTCGTTCGGCTGACGGTCCCTGCGTGAGT
>CM000735.1/4744050-4743991 Bacillus cereus Rock4-18 chromosome, whole genome shotgun sequence.
ATAATTATAGGCGATGGAGTTCGCCATAACCGCTGCTTAGCTAATGACTCCTATCAGTAT
>LYBW01000040.1/63579-63504 Ensifer sp. YIC4027 C593, whole genome shotgun APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
sequence .
TGGCACGTTGGCAATGGATTCTGCCAGGCATCGAGCCGAACCGCTTCCCCACGAAGCTGA
TGACTCCTGCTCATGA
>KI517368.1/144871-144810 Pseudomonas sp . HPB0071 genomic scaffold aczli- supercontl .2 , whole genome shotgun sequence.
TGGTGTACAGGAGATGGTATTTCTCCTTTAACTGCTCTTGGGGTAATGATGCCTAAGAAC
CC
>HF999089.1/3110-3034 Blautia sp . CAG:257 genomic scaffold, scfl63
ATTTTATTGGGGAATGAATGTCTCCCTAAGTATTAATACTTAAACCGCTTAGAAAAGCTG
ATGACTTCTGCAATCAA
>MKV001000004.1/560861-560787 Micrococcales bacterium 73-15
SCNpilot_expt_1000_bf_scaffold_204, whole genome shotgun sequence.
TGACGACACGGCGATGGCGCTCGCCGGGGCAGACAGCCCGAACCGCCGTTCTCGGCTGAT AGCACCTGTGACGAG
>MGQE01000007.1/13655-13593 Deltaproteobacteria bacterium RBG_13_51_10 RBG_13_scaffold_11549_curated, whole genome shotgun sequence.
GATACAAATGGTGATGGAGTTCACCATTGAACCGCCCGCGAGGCTGATAACTCCTGTTTT TAG
>FQUY01000004.1/108760-108703 Desulfotomaculum putei DSM 12395 genome assembly, contig: EJ51DRAFT_scaffold00004.4
AAATTGAACGGCGATGGAGTTCGCCTAACGCTGCCAAGCTGATGACTCCTACCCTTGT
>CP003532.1/60231-60156 Mesotoga prima MesGl .Ag .4.2 , complete genome.
GATTCTCTAGGCAATGAAGCTTGCCTTGGAGAACAATCCTCCAGAACCGCCACCAGCTAA TAGCTTCTACCAAAAT
>LGCK01000010.1/389365-389305 Leptolinea tardivitalis strain YMTK-2 contig_3, whole genome shotgun sequence.
ATTTCCGTTGGTGATGAGGCTCACCCGAAACCGTCCTTTGACTGATAGCCTCTACTGAAT
T
>BAW001000010.1/58961-58901 Geobacillus caldoxylosilyticus NBRC 107762 DNA, contig: GCA01S010.
AGTGTAATAGGCGATGGAGTTCTCCTAAACTTCCTTTCGGGCTGATGGCTCCTACCAATG
A
>JPRP01000001.1/2233190-2233125 Chryseobacterium formosense strain LMG 24722 contigOl, whole genome shotgun sequence.
TTCGAAAAAGGAAATGGTGTCTTCCTTACCCAACCGCCTTAAAAAGCTGATGGCGCCTGA TTAATT
>MIBK01000196.1/1903-1839 Spirochaetes bacterium RBG_16_67_19
RBG_16_scaffold_64826 , whole genome shotgun sequence.
TGCTGGCCTGGGGATGGAGTCCCCCGCATAAATACCCACCAAGGGCTGATGACTCCTACG ACGCA
>MGNM01000040.1/11337-11423 Chloroflexi bacterium RBG_16_51_9
RBG_16_scaffold_21187 , whole genome shotgun sequence.
TTGTTTAGTGGTGATGAAGCTCGCCAAGGATTATCTTTAATTTCCTGAACCGTCCTTCTA AGAAGGTCTGATAGCTTCTACTGGCTT
>JPRH01000001.1/1006855-1006788 Chryseobacterium soli strain DSM 19298 ContigOl, whole genome shotgun sequence.
CAGGATCAAGGAAATGGTGTTCTTCCTTCCCCAACCGCTTTACAAAAGCTGATGACGCCT
GATTAACA
>CP011390.1/2573175-2573108 Flavisolibacter tropicus strain LCS9, complete genome .
CTATATACAGGAAATGGTGTCTTCCTGCCCCAACCGCTTCTTCTAAAGCTAATGGCGCCT
ACAAGTAG
>FUYA01000003.1/163995-163933 Desulfovibrio bizertensis DSM 18034 genome APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
assembly, contig: BR26DRAFT_scaffold00003.3
TCAAACTTCGGGGATGAAGTCCCCCGTGTAACTGCTCGATGAGCTGATGACTTCTGCTGT
GAA
>LBUE01000009.1/8013-7945 Microgenomates (Woesebacteria) bacterium
GW2011_GWC1_38_13 US75_C0009, whole genome shotgun sequence.
GAAGCTATAAGCGATGAGGCTCGCTATACAACCGCTAAAACTCCTTTAGCTGATAGCTTC TACAAATTT
>MHZE01000558.1/428-489 Planctomycetes bacterium RIFOXYD12_FULL_42_12 rifoxyd3_full_scaffold_33258_curated, whole genome shotgun sequence.
TTACGTCCATGCGATGGAGTTCGCTTAACCGCTATAACATAGCTGATAACTCCTATTGAA GC
>CP005290.1/1481651-1481590 Archaeoglobus suitaticallidus PM70-1, complete genome .
AAAGAGAAAGGTGATGGCGTTCACCTCTAACTGCCGTAAAGGCTGATGACGCCTGTTTTC
AG
>CP000863.1/439864-439930 Acinetobacter baumannii ACICU, complete genome.
CTAAAACAGGGAGATGGCATTCCTCCCTTGAAAAACCGCCGTATTGGCTAATGATGCCTA
CGTTACC
>AM743169.1/1735698-1735635 Stenotrophomonas maltophilia K279a complete genome, strain K279a
GACCGCGATGGGGATGGGGCTCCCCCGATAACCGCCTGAGAAGGCTGATGGCTCCTGCCA
GGAC
>BAAZ01017612.1/714-637 Human gut metagenome DNA, contig sequence: F2- X_017612.
CTGGACTGCGGTGATGGGACTCGCCTGAAGCCGTTACAGGCTTCGAACCGCAAACCCGCT
GATGGTTCCTACGACATA
>CP000903.1/4869306-4869365 Bacillus weihenstephanensis KBAB4, complete genome .
ATAATCATAGGTGATGGAGTTCGCCATAACCGCTGCTTAGCTAATGACTCCTACCAGTAT
>ABLC01000046.1/24103-24037 Burkholderia ambifaria IOP40-10 ctg00616, whole genome shotgun sequence .
GTCACGCGTGGAGATGGCATTCCTCCTTTAACCGCCGATTCGCTCGGCTGATGATGCCTA
CGTGCCC
>CP000545.1/1653068-1653134 Burkholderia mallei NCTC 10229 chromosome II, complete sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>AE009952.1/1314365-1314302 Yersinia pestis KIM, complete genome.
TCAAGATCTGGAGATGACACGCCTCCATAACCGCCCTAACAAGGCTAATGATGTCTACGT
GACC
>AM406671.1/272338-272410 Lactococcus lactis subsp. cremoris MG1363, complete genome
ATAATAGATGGGTATGGTGCACACCCGAAACCGTTTCAAGAAATAACTTGAATCTAATGG
CGCCTACAAAGCA
>AAOU01000011.1/78540-78473 Photobacterium sp. SKA34 1099521381155, whole genome shotgun sequence.
ATTATGGCGGGAGATGATGTTCCTCCTTAAACCGCCTTTCAATCAAGGATAATGACGTCT
AACAACAT
>AAFX01080011.1/664-577 Environmental sequence
2662324_fasta . screen . Contig7507 , whole genome shotgun sequence.
TTTGGGCGCGGTGATGGAGTTCACCATCTACTACAAATATGGAAAACCGTCCAACTTCTT
TAGCCGGACTAATAACTCCTACCTATAA
>CP001132.1/2508708-2508771 Acidithiobacillus ferrooxidans ATCC 53993, APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
complete genome.
AAAACGTGTGGCGATGGAGTCCGCCCAATAACCGCCTTTATAGGCTGATGACTGCTACTG
ATCC
>CU914166.1/1142118-1142184 Ralstonia solanacearum strain IPO1609 Genome Draft .
CCGGCTTGCGGAGATGGCATACCTCCCTTAACCGCCGATCACCTCGGCTGATGATGCCTA
CAAGTTC
>ACNB01000166.1/6496-6555 Bacillus thuringiensis serovar sotto str. T04001 contig00032, whole genome shotgun sequence.
ATAATCAGAGGCGATGGAGTTCGCCATAACTGCTGCTTAGCTAATGACTCCTACCAGTAT
>CP001176.1/5025071-5025130 Bacillus cereus B4264, complete genome.
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>CU458896.1/2498027-2498099 Mycobacterium abscessus chromosome, complete sequence .
TTCGTCAGGCGCGATGGATCTCGCCAGGGCTTGTCCCGAACCGCCACTGACGGCTGATAG
TTCCTGTGTTGAT
>BABC01003766.1/468-402 Human gut metagenome DNA, contig sequence: In- B_003766.
AAGTGAATAGGTGATGGTGTTCGCCTTTAAATGCTAACGAAAGATAGCTAATGACGCCTA
CTAAAGT
>CP001283.1/4890016-4890075 Bacillus cereus AH820, complete genome.
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>AAQL01009062.1/467-395 Environmental sequence s8_174661, whole genome shotgun sequence.
AGGCAGAAAGGGAATGAAGTTCTCCCCCGATTTTTATCGAAACCGCCAATGAGCTGATGA
CTTCTGTGCAATG
>ABQQ01000001.1/79958-80027 Bifidobacterium longum subsp . infantis CCUG 52486 cont2.1, whole genome shotgun sequence.
TAAGGAACCGGCGATGGAACCCGCCTGAACTCGTTTCGAACCGCACACTGCTGATAGTTC
CTATACACGA
>CP001052.1/4147651-4147580 Burkholderia phytofirmans PsJN chromosome 1, complete sequence.
GAAGCTTTCGGAGATGGCATTCCTCCCGTAACCGCCCTGCATTCCCGCAGAGCTGATGAT
GCCTACGAGTTC
>ABJD02000101.1/724041-724112 Providencia stuartii ATCC 25827 P_stuartii- 2.0. l_Cont 1786 , whole genome shotgun sequence.
CTAACAAAAGGAGATGGCATTCCTCCTGTTTAAAAACCGTCCAAATAAAGGGCTGATGAT GCCTGCGTTCAC
>ABXV02000011.1/470942-470875 Providencia rustigianii DSM 4541
P_rustigianii-1.0. l_Cont0.7 , whole genome shotgun sequence.
ATTTATTTGGGAGATGGCATTCCTCCTACCCAAACCGTCCATATTGGACTGATGATGCCT ACGTAAAC
>AAXB02000001.1/235634-235709 Dorea longicatena DSM 13814 D_longicatena- MSIQ_Cont368, whole genome shotgun sequence.
CTTTTAGAAGGGAATGAAGTTCTCCCTTAGTGATCATACTAGAACCGCTTATAAAGCTGA
TGACTTCTGCGAATAA
>AADL01001839.1/2172-2233 Thermoplasmatales archaeon Gpl AMC_Contl839, whole genome shotgun sequence .
TTCAAGAAAGGCTATGACGTTAGCCTTTAAGCGCCCTCCGGGCTGATAACGTCTGATCTT
TA
>ABHH01000009.1/50257-50182 Carnobacterium sp. AT7 1101238000988, whole genome shotgun sequence.
TACCATTTTGGGAATGAATGTCTCCCTTAGCATACTGCTTAAACCGCTTATTAAAGCTGA APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TGACTTCTGCGCATAG
>ABSD01000051.1/73291-73230 Aciduliprofundum boonei T469
ctg_ll 04670154630 , whole genome shotgun sequence.
TCTGCTTCGGGTGATGGCGTCCACCCCGAAACGCCCGCGTGGCTGATGACGCCTTTTCGC AT
>AADL01000991.1/122-193 Leptospirillum sp. Group III AMC_Cont991, whole genome shotgun sequence.
TGTAGTCCGTGGAATGGCGTTTCCACTAATTCAAACCGCCGGCTGAAGCTGGCTGATGAC
GCCTGCTTGAAA
>BABG01014606.1/762-690 Human gut metagenome DNA, contig sequence: In- R_014606.
CGAATGAGAGGGAATGAAGTTCTCCCTAGGGAAACCAGAACTGCTGAATATAGCTGATGA
CTTCTACGATTAT
>ABBF01000636.1/6946-6880 Burkholderia oklahomensis E0147
PMP6xxBPSxxE0147-636, whole genome shotgun sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA CGCGTTC
>ABQC02000019.1/404312-404380 Bacteroides plebeius DSM 17135 B_plebeius- 2.0. l_Cont2578, whole genome shotgun sequence.
TCGAATCAAGGGAATGGGACTTCCCTGTGATAAACCGCTAATGAAGTAGCTGATAGTTCC TACCGGAGC
>ABSD01000048.1/33401-33462 Aciduliprofundum boonei T469
ctg_ll 04670154627 , whole genome shotgun sequence.
TCTGCTTCGGGTGATGGCGTCCACCCCGAAACGCCCGCGTGGCTGATGACGCCTTTTCGC AT
>AE016879.1/4823733-4823792 Bacillus anthracis str. Ames, complete genome. ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>BABB01015950.1/548-449 Human gut metagenome DNA, contig sequence: In- A_015950.
AATATTACTGGGAATGAGGTTCTCCCAGGAACGAGATAGAAATAGCTAAATCTTTCGTTC
ATAACCGCTTGATTATAAAGCTGATGACTTCTGCAAAGCG
>ABCC02000039.1/237860-237932 Clostridium bolteae ATCC BAA-613 C_bolteae- 3.0. l_Cont299, whole genome shotgun sequence.
ATAAGTACCGGGAATGAAGTTCTCCCTTAGTCTGACTAGAACCGCTTATAAAGCTGATGA
CTTCTGCATTATG
>AAHW02000001.1/301150-301084 Burkholderia pseudomallei S13
ctg_1100107453550, whole genome shotgun sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA CGCGTTC
>ACIZ01000090.1/22576-22630 Lactobacillus rhamnosus LMS2-1 contig00098, whole genome shotgun sequence .
GAATTAAATGGCGATGGTGTTCGCCTATACGTAAGTTGATGACACCTACCTTGTA
>CP000547.1/225172-225238 Burkholderia mallei NCTC 10247 chromosome II, complete sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>BABE01001448.1/1961-1891 Human gut metagenome DNA, contig sequence: In- E_001448.
GTCCGTGACGGCGATGGATCCCGCCTGGCCGATGCCGAACTGCCGATCTGGCTGATGGTT
CCTGCCCCGTC
>ABYW01000007.1/311285-311355 Methanobrevibacter smithii DSM 2375
M_smithiDSM2375-l .0_Cont0.8, whole genome shotgun sequence.
TAATAACTTGGTGATGGGGTTCACCAGAAACTTATTTTCAAACCGCAATTGCTGATAACT APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
CCTATGTTATA
>CP001177.1/4829881-4829940 Bacillus cereus AH187, complete genome.
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>CP000413.1/283109-283169 Lactobacillus gasseri ATCC 33323, complete genome .
AATTGAATAGGCGATGACGGTCGCCATTAAACGAGCAAAATCTAATGACGTCTACTATTT
A
>AE017355.1/4836662-4836721 Bacillus thuringiensis serovar konkukian str. 97-27, complete genome.
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>ABSD01000050.1/82708-82766 Aciduliprofundum boonei T469
ctg_ll 04670154629, whole genome shotgun sequence.
CATTTATAAGGTGATGGGGCCCACCAAACCGCCGCAAGGCTGATGGCTCCTTTTGAAAA
>CP000413.1/1279279-1279217 Lactobacillus gasseri ATCC 33323, complete genome .
ATTATAATAGGTGATGGCGTTCACCAATTTAACCGATTAAGATCTGACGACGCCTACTTT
CTT
>CP000250.1/3360939-3360999 Rhodopseudomonas palustris HaA2, complete genome .
AACGGCGATGGGGATGGAGTCCCCCGACAACCGCCGCGAGGCTGATGACTCCTACCGGGC
G
>ABVY01000154.1/6708-6604 Pectobacterium carotovorum subsp. carotovorum WPP14 064_059_Sequence00223, whole genome shotgun sequence.
TACTAAAGCGGAGATGACATTCCTCCCTACCCTCGCGCAGCCAAAACGCTGACAGCCTTA AAAACAGGGTACAACCGCCACTCGGCTAATGATGTCTACGTCCAC
>ABGC03000023.1/52156-52229 Clostridium sp . SS2/1 Clostridium_sp_SS21- 1.0. l_Cont 675 , whole genome shotgun sequence.
AAAGAAAAAGGGAATGAAGTTCTCCCTCGAAGAGATTCGAAACCGCTTATTAAGCTGATG
ACTTCTGTGCGATG
>AAAK03000106.1/3475-3542 Enterococcus faecium DO ctg519, whole genome shotgun sequence.
AAGTGAATAGGTGTATGGTGTTCGCCTTTAAATGCTAACGAAAGATAGCTAATGACGCCT
ACTAAAGT
>BAAV01013884.1/268-196 Human gut metagenome DNA, contig sequence: Fl- T_013884.
AATATCAAAGGGAATGAAGTACTCCCTTGGGAAACCTAAACTGCTTATTCAAGCTGATGA
CTTCTACGATTTT
>BAAW01000138.1/2195-2134 Human gut metagenome DNA, contig sequence: Fl- U_000138.
AGCCTCACAGGGGATGGCATTCCCCCTTGAACCGCCATGTGGCTAATGATGCCTGCTTTT
TA
>AF359557.1/22906-22967 Pseudomonas syringae pv. maculicola plasmid pFKN, complete sequence.
ATCTGCACAGGAGATGGCATTCCTCCTTGAAACCGCCTCGTGCTGATGATGCCTACGTAA
AC
>AAFX01046182.1/76-12 Environmental sequence XZS73310.bl, whole genome shotgun sequence.
AACCCCAAGGGAAATGGTGTCTTCCCTGCACAAACCGCTTGAAAGCTGATGACGCCTTCA
AATCG
>CP000125.1/1222860-1222794 Burkholderia pseudomallei 1710b chromosome II, complete sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>ABYS02000004.1/646163-646241 Bifidobacterium angulatum DSM 20098
B_angulatum-2.0. l_ContO .2 , whole genome shotgun sequence.
CTTACCGGTGGCGATGGGACTCGCCTGGGGCCTGCAAGGATTCCCGAACCGCAATTCCGC TGATGGTTCCTATTGACGC
>BABE01000053.1/15068-15137 Human gut metagenome DNA, contig sequence: In- E_000053.
TGGAGAACCGGCGATGGGACCCGCCTGAACTTGTTTCGAACCGCACACTGCTGATAGTTC
CTACACACGA
>CP000325.1/3045327-3045401 Mycobacterium ulcerans Agy99, complete genome.
ACTAGCAGCGGCGATGGGGCTCGCCAGGAAGCTCGACTTCTGAACCGCCAGCCGGCTGAT
GGCACCTGCGAACAA
>CU695240.1/565243-565309 Ralstonia solanacearum strain MolK2 Genome Draft .
CCGGCTTGCGGAGATGGCATACCTCCCTTAACCGCCGATCACCTCGGCTGATGATGCCTA
CAAGTTC
>AAQL01004413.1/1469-1396 Environmental sequence s8_165414, whole genome shotgun sequence.
AAAGAGAAAGGGAATGAAGTTCTCCCTCGAAGAGATTCGAAACCGCTTATTAAGCTGATG
ACTTCTGTGTAATG
>AAQL01000312.1/1879-1949 Environmental sequence s8_178954.2, whole genome shotgun sequence.
TAATAACTTGGTGATGGGGTTCACCAGAAACTTATTTTCAAACCGCAATTGCTGATAACT
CCTAGGTTATA
>AM039952.1/2350991-2351059 Xanthomonas campestris pv. vesicatoria complete genome
GGCCGCACAGGAGATGGCATTCCTCCTCGAACCGCACGCACCCTGCGCGCTGATGATGCC
TGCCCACCC
>CP001364.1/429603-429664 Chloroflexus sp. Y-400-fl, complete genome.
AATCGATTGGGTGATGAGGCTCACCCTCAACTGCCATTACGGCTGATAGCCTCTACAGGG
AA
>BAAY01017851.1/254-182 Human gut metagenome DNA, contig sequence: F2- W_017851.
TAGCTGGCAGGGAATGAAGTTCTCCCCAGGCAGTGCCTAAACCGCTATTACAGCTGATGA
CTTCTGTTGTTTA
>CP001598.1/4823759-4823818 Bacillus anthracis str. A0248, complete genome .
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>BABF01001201.1/3795-3726 Human gut metagenome DNA, contig sequence: In-
M_001201.
TGGAGAACCGGCGATGGGACCCGCCTGAACTTGTTTCGAACCGCACACTGCTGATAGTTC
CTACACACGA
>CP001146.1/51886-51949 Dictyoglomus thermophilum H-6-12, complete genome.
ATCACAAGTGGCGATGGAGTCCGCCATATAATTGCCGATAAAGGCTGATGACTCCTACTT
ATGT
>CP000485.1/4850198-4850257 Bacillus thuringiensis str. A1 Hakam, complete genome .
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>FM177140.1/2571522-2571586 Lactobacillus casei BL23 complete genome, strain BL23
AGAAGAACAGGCGATGATGTTCGCCGCAAATGATTGTGTAGCAATCTGATGACGTCTACT
GAAAC
>AAW001000020.1/7863-7793 Leptospirillum rubarum
Leptol I_Scaffold_8049_Cont20 , whole genome shotgun sequence. APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AGGAAGACAGGCGATGGGGTTCGCCAAAACCGCCCCGGAGAACAGCGGGAGCTGATGACC
CCTACTCCATT
>CR378667.1/326415-326332 Photobacterium profundum SS9; segment 5/12
ATTATGACGGGAGATGATGTTCCTCCTTTAACTGCCTTACTGAAGTTTACCCTTCTTAAT
AAGGATGATGACGTCTAACAACAT
>FM209186.1/5248344-5248282 Pseudomonas aeruginosa LESB58 complete genome sequence
TTGCCGACAGGAGATGGCATTCCTCCTTCAACCGCCCCTGGGGCTGATGATGCCTACGCA
TGA
>CP001407.1/4853207-4853266 Bacillus cereus 03BB102, complete genome.
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>CP001139.1/34842-34908 Vibrio fischeri MJ11 chromosome I, complete sequence .
CCAAATACGGGTGATGGATTTCCACCTTTAACCGCTCTATTTTAGAGATAATGATTCCTA
CTATAGC
>AE008923.1/2352963-2353030 Xanthomonas axonopodis pv. citri str. 306, complete genome.
GGCCGCACAGGAGATGGCATTCCTCCTCGAACCGCACGCACCCGCGCGCTGATGATGCCT
GCCCACCC
>BABC01000066.1/7792-7723 Human gut metagenome DNA, contig sequence: In- B_000066.
CTGGGATTCGGCGATGGGACCCGCCTGAACTTGTTTCGAACCGCACACTGCTGATAGTTC
CTACACACGA
>CP000117.1/118149-118230 Anabaena variabilis ATCC 29413, complete genome.
GTTACTAATGGCGATGGAGTTCGCCCAAACCGCTATTCTAGTCAAAAAATAAATCTAGAA
GGCTGATGACTCCTACTACTCT
>CP002372.1/388527-388452 Thermococcus barophilus MP, complete genome.
GTTTAGTCGGGCGATGGCGTCCGCCCGGGCTTCGAGCCCGAACCGCCCGTTAAGGGCTGA
TGACGCCTGTTCTACC
>CP000239.1/520035-519935 Synechococcus sp . JA-3-3Ab, complete genome.
CCGGCAACGGGCGATGGGGTTCGCCGCAACCGCCTCAGGCAGATCCTAGGGATAAAGTTT
CTCTAAGACTTAGAGCTGGGGCTGATGACTCCTACTCTTCT
>CP002222.1/1678154-1678089 Lactobacillus plantarum subsp. plantarum ST- III, complete genome.
AAGTTAATCGGCGATGACGTTCGCCACATAATAATTGATAATCAATTTGATGACGTCTAC
TGTTTG
>AAPH01000002.1/303167-303249 Photobacterium profundum 3TCK 1099451005285, whole genome shotgun sequence .
ATTATGACGGGAGATGATGATCCTCCTTTAACTGCCTTACTGAAGTTTACCCTTCTTATA AGGATGATGACGTCTAACAACAT
>AAZW01000073.1/10424-10341 Vibrionales bacterium SWAT-3 1101732140416, whole genome shotgun sequence .
TAGCCATCAGGTGATGGGGTTCCACCTAAGCTTTTGCTTCAACCGCCCGTTCTTTTGAAC
CGTGCTAATGGCTCCTACAGAATC
>CP000227.1/4763720-4763779 Bacillus cereus Ql, complete genome.
ATAATTATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>AADL01002490.1/18401-18475 Leptospirillum sp. Group II '5-way CG' ctgl0933, whole genome shotgun sequence.
ATGACGAGGGGTGATGGAGTCCACCTGAACGGCCCTTTCTTTCATAAGGAAAGGGATGAT
GACTCCTGTGAGAGA
>ABQV01000049.1/12339-12285 Lactobacillus paracasei subsp. paracasei 8700:2 cont2.49, whole genome shotgun sequence.
AAATTAAATGGCGATGGTGTTCGCCTATACGCAAGTTGATGACACCTACCTTGAG APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>AAPS01000002.1/29498-29423 Vibrio alginolyticus 12G01 1100007009696, whole genome shotgun sequence .
AATGGGTTCGGTGATGGGGTACCACCGGAATCGAAAGATTCGAACCGCTGTTAAAGCTAA TGACTCCTACAGAAAC
>AAZP01000053.1/29998-29932 Burkholderia mallei PRL-20
gcontig_1105338602060, whole genome shotgun sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA CGCGTTC
>BABG01014343.1/547-474 Human gut metagenome DNA, contig sequence: In- R_014343.
ATCCAGAACGGGAATGAGGTTCTCCCACGATTTTTGATCGAAACCGCCAAAAGGCTGATG
ACTTCTGTGCATGG
>BABB01014491.1/724-793 Human gut metagenome DNA, contig sequence: In- A_014491.
TGGAGAACCGGCGATGGGACCCGCCTGAACTTGTTTCGAACCGCACACTGCTGATAGTTC
CTACACACGA
>ABLC01000083.1/20871-20937 Burkholderia ambifaria IOP40-10 ctg00076, whole genome shotgun sequence .
GTTGGCCGTGGAGATGGCATTCCTCCCTTAACTGCCGATTCGTTCGGCTGATGATGCCTA
CACACCC
>AM920689.1/2646740-2646665 Xanthomonas campestris pv. campestris complete genome, strain B100
GGCCATCTCGGAGATGGCGTTCCTCCGCGCTGTACCATCAACCGTAGCGATCGCTACTGA
TGACGCCTACAAGAAC
>BABF01015109.1/43-109 Human gut metagenome DNA, contig sequence: In- M_015109.
AAGTGAATAGGTGATGGTGTTCGCCTTTAAATGCTAACGAAAGATAGCTAATGACGCCTA
CTAAAGT
>BA000001.2/1278365-1278285 Pyrococcus horikoshii OT3 DNA, complete genome .
GGTTCCTTGGGCGATGGCGCCCGCCCGGGGCTTCGAGCCCGAACCGCCTCCCGTTGGGAG
GCTGATGGCGCCTATTCTGTA
>CP001078.1/1342444-1342385 Clostridium botulinum E3 str. Alaska E43, complete genome.
TATATATTAGGTGATGGAGTTCGCCTTTAAACGCGTAATGCTAATGACTCCTACAAAAAA
>CP000438.1/5078189-5078127 Pseudomonas aeruginosa UCBPP-PA14, complete genome .
TTGCCGACAGGAGATGGCATTCCTCCTTCAACCGCCCCTGGGGCTGATGATGCCTACGCA
TGA
>ABXU01000089.1/23581-23642 Desulfovibrio piger ATCC 29098
D_piger_ATCC29098-l .0_Contl0.5, whole genome shotgun sequence.
GAAGGTCAAGGGGATGGAGTCCCCCATGAACCGCATGTTTTGCTGATGACTCCTGCCGGA CC
>AP009049.1/304406-304465 Clostridium kluyveri NBRC 12016 DNA, complete genome .
AAATAATAAGGTGATGGAGTTCACCATAACCGCAGAAATGCTTATGACTCCTACAAATAA
>ABVK02000003.1/30981-31043 Thermus aquaticus Y51MC23 ctg73, whole genome shotgun sequence.
TGCCCTTAGGGCGATGGAGTCCGCCTTAACCGCCCGCTTCGGGCTGATGACTCCTACCGG
GTT
>AAFX01005291.1/667-735 Environmental sequence
2662324_fasta . screen . Contig38989, whole genome shotgun sequence.
ACGTCGCGTGGTGGTGGAGTCCACCTGAACCGCCGCTCTGACTGGACGGCCAATGACTCC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TACGCATAG
>ABYT01000096.1/44476-44400 Eubacterium biforme DSM 3989 E_biforme- 1.0. l_Cont 8.4 , whole genome shotgun sequence.
AATCTTTTTGGGAATGAAGTTCTCCCTTAGTTTAAAACTAAAACCGCTTTTTATAGGCTG
ATGACTTCTGTGTTTTT
>CP000702.1/1794817-1794880 Thermotoga petrophila RKU-1, complete genome.
AACCCAACGGGCGATGAGGCCCGCCCAAACTGCCCTGAAGAGGGCTGATGGCCTCTACTG
GCTT
>AM406671.1/1450312-1450238 Lactococcus lactis subsp. cremoris MG1363, complete genome
AATAATGATGGGTATGGTGCACACCCGAAACCGCTTTAAGAATAAAATCTTAAAACTAAT
GGCGCCTACAAACAA
>CP000668.1/3010272-3010335 Yersinia pestis Pestoides F, complete genome.
TCAAGATCTGGAGATGACACGCCTCCATAACCGCCCTAACAAGGCTAATGATGTCTACGT
GACC
>AAMR01000022.1/13784-13701 Vibrio splendidus 12B01 1099451319040, whole genome shotgun sequence.
TAGCCATCAGGTGATGGGGTTCCACCTAAGCTTTTGCTTCAACCGCCCGTTCTTTCGAAC
AGTGCTAATGACTCCTACAGAATC
>CU459141.1/3434387-3434321 Acinetobacter baumannii str. AYE, complete genome .
TTAAAACAGGGAGATGGCATTCCTCCCTTGAAAAACCGCCGTATTGGCTAATGATGCCTA
CGTTACC
>CP000923.1/40525-40583 Thermoanaerobacter sp. X514, complete genome. ATACCATAAGGCGATGGAGCTCGCCATAAATGCGGAACGCTGATGGCTCCTATTAGGAG
>BAAV01021889.1/203-131 Human gut metagenome DNA, contig sequence: Fl- T_021889.
GAAACAAATGGGAATGAATGTCTCCCCGGGATCATCCCGAACCGCTTATTAAGCTGATGA
CTTCTGTGATAAC
>AAV002000004.1/13914-13993 Ruminococcus obeum ATCC 29174 R_obeum- MSIQ_Contl59, whole genome shotgun sequence.
AAATATAAAGGGAATGAGGTTCTCCCTAAGTAATAAAATTACTTAAACCGCTTATGAAAG
CTGATGACTTCTGCGAGTAA
>CP002952.1/ 944770-944695 Thermococcus sp. AM4, complete genome.
AGGGGGACGGGCGATGGCGTCCGCCCGGGGCTTCGAGCCCGAACCGCCCGCTCGGGCTGA
TGACGCCTGTTCTACC
>CP001186.1/4973458-4973517 Bacillus cereus G9842, complete genome.
ATAATCAGAGGCGATGGAGTTCGCCATAACTGCTGCTTAGCTAATGACTCCTACCAGTAT
>CP000950.1/3331346-3331409 Yersinia pseudotuberculosis YPIII, complete genome .
TCAAGATCTGGAGATGACACGCCTCCATAACCGCCCTAACAAGGCTAATGATGTCTACGT
GACC
>BAAY01001852.1/271-348 Human gut metagenome DNA, contig sequence: F2- W_001852.
CTGGACTGCGGTGATGGGACTCGCCTGAAGCCGTTACAGGCTTCGAACCGCAAACCCGCT
GATGGTTCCTACGACATA
>CP000744.1/5109630-5109568 Pseudomonas aeruginosa PA7, complete genome.
TAGCCGGCAGGAGATGGCATTCCTCCTTCAACCGCCCCTGGGGCTGATGATGCCTACGCA
TGA
>CP001052.1/1800329-1800268 Burkholderia phytofirmans PsJN chromosome 1, complete sequence.
TTGCGCGCTGGAGATGGCATTCTCCCTTAACCGCCCTCGTGGCTGATGATGCCTGCTTCG
TC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>ABSD01000048.1/169379-169321 Aciduliprofundum boonei T469
ctg_ll 04670154627 , whole genome shotgun sequence.
CATTTATAAGGTGATGGGGCCCACCAAACCGCCGCAAGGCTGATGGCTCCTTTTGAAAA
>AALF02000002.1/247760-247698 Yersinia intermedia ATCC 29909 contig01174, whole genome shotgun sequence .
GCAAGATCTGGAGATGACATTCCTCCATAACCGCCCTTCAAGGCTGATGATGTCTACGTA
ACC
>CP000525.1/1388196-1388262 Burkholderia mallei SAVP1 chromosome II, complete sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>ABLD01000041.1/16968-17073 Burkholderia graminis C4D1M ctg26, whole genome shotgun sequence.
TTGCGTGCTGGAGATGGCATTCTCCAATGGGTGTCCGGCGTGCGCTTCAACACACGCGCT
CGCGCGAACCCATTAACCGCCTCGTGCTGATGATGCCTGCTTCCCC
>BAAZ01021286.1/777-699 Human gut metagenome DNA, contig sequence: F2- X_021286.
AGAATAAACGGGAATGAAGTTCTCCCGAAGTAACCAGATTACTTGAACCGCTTTAAAAGC
TGATGACTTCTGCGACGAA
>AE017225.1/4825068-4825127 Bacillus anthracis str. Sterne, complete genome .
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>CP001133.1/1365444-1365379 Vibrio fischeri MJ11 chromosome II, complete sequence .
ATACGTAAAGGTGATGGGGTTCCACCTACTTAACCGCCAATGTTGGCTGATGACTCCTAC
AGAATA
>ABEY02000020.1/92291-92375 Coprococcus eutactus ATCC 27759 C_eutactus- 2.0. l_Cont273, whole genome shotgun sequence.
ATTATAAAAGGGAATGAGGTTCTCCTCTGGAAAACCTGAATGGGCAACCTAAACTGCTAA
TATAGCTGATGACTTCTACGATTAT
>ABXA01000043.1/134014-134086 Anaerococcus hydrogenalis DSM 7454
A_hydrogenalis-l .0_Contl65.1, whole genome shotgun sequence.
ATATTAGAAGGGAATGAAGTTCTCCCTTAGAAAAATCTAAAACCGCATTTATGCTGATGA CTTCTGTAACTAT
>CP001337.1/1659433-1659494 Chloroflexus aggregans DSM 9485, complete genome .
TAGCTCATAGGTGATAAGGCTCACCGTTGAACTGCCCGCGGGCTGATGGCCTCTACAGAA
GT
>AAKV01000114.1/24659-24721 Pseudomonas aeruginosa C3719 contl.114, whole genome shotgun sequence.
TTGCCGACAGGAGATGGCATTCCTCCTTCAACCGCCCCTGGGGCTGATGATGCCTACGCA
TGA
>BABF01003476.1/1904-1985 Human gut metagenome DNA, contig sequence: In- M_003476.
CTGTTACCCGGGAATGAGGTCTCCCATGGTAGTGTAATATTTACCAGAACCGCTTATGAC
AGCTGATGGCTTCTGCATTGTG
>AAHS03000001.1/321518-321584 Burkholderia pseudomallei 1710a
gcontig_1105229757784 , whole genome shotgun sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA CGCGTTC
>ABXW01000051.1/16406-16474 Providencia alcalifaciens DSM 30120
P_alcalifaciens-1.0_Contl37.1 , whole genome shotgun sequence.
GTCTTGTTAGGAGATGGCATTCCTCCTAGTTTAACCGTCCTTTTGTGGACTGATGATGCC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
TACGTAAAC
>ABXP01000064.1/2857-2798 Carboxydibrachium pacificum DSM 12653
ctg_1106511212226, whole genome shotgun sequence.
TGCATAAAAAGTGATGGAACCCACTTTAACCGCCGAAAGGCTGATGGTTCCTACTTGTGA
>CP000967.1/2767544-2767620 Xanthomonas oryzae pv. oryzae PX099A, complete genome .
TCGGGATTCGGAGATGGCGTTCCTCCGCGCAGTTCCATCAAACCGTAGCGATTGCTACTG
ATGACGCCTACAAGAAC
>FM177140.1/2570801-2570747 Lactobacillus casei BL23 complete genome, strain BL23
AAATTAAATGGCGATGGTGTTCGCCTATACGCAAGTTGATGACACCTACCTTGAG
>ABVX01000036.1/35509-35405 Pectobacterium carotovorum subsp. brasiliensis PBR1692 0044_0031_Sequence00019, whole genome shotgun sequence.
TACTAAAGCGGAGATGACATTCCTCCCTACCCTCGCGCAGCCAAAACGCTGACAGCCTTA AAAACAGGGTACAACCGCCACTCGGCTAATGATGTCTACGTCCAC
>BABG01003409.1/431-358 Human gut metagenome DNA, contig sequence: In- R_003409.
AAAGAAAAAGGGAATGAAGTTCTCCCTCGAAGAGATTCGAAACCGCTTATTAAGCTGATG
ACTTCTGTGTAATG
>CP000605.1/969434-969503 Bifidobacterium longum DJO10A, complete genome.
TGGAGAACCGGCGATGGGACCCGCCTGAACTTGTTTCGAACCGCACACTGCTGATAGTTC
CTACACACGA
>ABXP01000064.1/9651-9591 Carboxydibrachium pacificum DSM 12653
ctg_1106511212226, whole genome shotgun sequence.
AATATTTAAGGTGATGAAGCCCACCTTAATTGCCGTAAAGGCTGATGGCTTCTACGAAGA
T
>CP001189.1/2909225-2909147 Gluconacetobacter diazotrophicus PA1 5, complete genome.
GTCGGCGCAGGCAATGGATTTCTGCCTGACCGTCCGGGTCGAACCGCCTCTTTCGGGGGC
TGATGATTCCTACCCACCG
>L39876.1/4762-4825 Caldicellulosiruptor saccharolyticus alpha-dextrin 6- glucanohydrolase (pulA) and pepX genes, complete cds and pepY gene, partial cds.
TGCTGCAACGGCGAGGGAGTCCGCCGAACAAATGCCAATGATGGCTGATGACTCCTACAA
ATAT
>ABYN01000048.1/132548-132614 Acinetobacter sp . ATCC 27244 contig00048, whole genome shotgun sequence .
TAGGGCAAAGGAGATGGCATTCCTCCTGTAACAAACCGCCATTGTGGCTAATGATGCCTA
CGTTACT
>ABBJ01001663.1/1415-1349 Burkholderia pseudomallei 14 PMP6xxBPSxxl4-1663, whole genome shotgun sequence .
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>AANC01000009.1/2226-2162 Leeuwenhoekiella blandensis MED217
1099517004302, whole genome shotgun sequence.
TCAATCACAGGCAATGAGGTCTGCCTCAAACCGCTTCTTAGCAAGCTGATGACTTCTACT
TAACC
>BABE01006493.1/378-447 Human gut metagenome DNA, contig sequence: In- E_006493.
TAAGGAACCGGCGATGGAACCCGCCTGAACTCGTTTCGAACCGCACACTGCTGATAGTTC
CTACACACGA
>AAWP01000027.1/44793-44868 Vibrio harveyi HY01 gcontig_1104549816479, whole genome shotgun sequence . APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
AATGCACACGGTGATGGGGTGCCACCAGAATCGAAAGATTCAAACCGCTTTTACAGCTGA TGACTCCTACAGAAAC
>AACY020346117.1/2623-2561 Marine metagenome 1096626735748, whole genome shotgun sequence.
TTGCGCGCTGGAGATGGCATTCTCCATTAACCGCCCTTGATGGCTGATGATGCCTGCTTC
GCC
>ABBM01000878.1/872-806 Burkholderia thailandensis MSMB43 PMP6xxBPSxx381- 878, whole genome shotgun sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>AAQK01000307.1/4537-4610 Environmental sequence s7_178954.3, whole genome shotgun sequence.
TAATAACTTGGTGATGGGGTTCACCAGAAACTTATTTTCAAACCGCAAAAATTGCTGATA
ACTCCTATGTTATA
>ABBQ01000905.1/4347-4281 Burkholderia pseudomallei NCTC 13177 strain NCTC13177 PMP6xxBPSxxNCTC13177-905, whole genome shotgun sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA CGCGTTC
>AAFX01053991.1/352-284 Environmental sequence XZS55564.bl, whole genome shotgun sequence.
AACGATAACGGCGATGGGGTTCGCCAGTCAAACCACTCGCGGCGGCGAGTTGATGACCCC
TACTTCGAG
>FM178380.1/1142244-1142180 Aliivibrio salmonicida LFI1238 chromosome 2 complete genome
GGCGTACAAGGTGATGGGGTGCCACCTACTTAACCGCCAATTTGGCTGATGACTCCTACA
GTAAA
>ACCE01000005.1/360072-360138 Burkholderia pseudomallei 576 BUC . Contigl 85 , whole genome shotgun sequence .
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>AAVL02000037.1/287218-287292 Eubacterium ventriosum ATCC 27560
E_ventriosum-MSIQ_Cont233 , whole genome shotgun sequence.
TAAATAAAAGGGAATGAGGTTCTCCCTCGATTTAATCGAAACCGCTTATAACAAGCTGAT GACTTCTGTGCAATA
>BABD01000872.1/1124-1056 Human gut metagenome DNA, contig sequence: In- D_000872.
CAAAGTCAGGGGAATGGGACTTCCCTGTTATAAACCGCTAATAAAGTAGCTGATAGTTCC
TGCCGGAGT
>ACKA01000027.1/241427-241361 Burkholderia pseudomallei Pakistan 9
BUH . Contig272 , whole genome shotgun sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>ACGY01000069.1/12451-12387 Lactobacillus paracasei subsp. paracasei ATCC 25302 contigOOHO, whole genome shotgun sequence.
AGAAGAACAGGCGATGATGTTCGCCGCAAATGATTGTGTAGCAATCTGATGACGTCTACT GAAAC
>AJ248285.1/89834-89914 Pyrococcus abyssi GE5 complete genome; segment 3/6
GTTCTCTTGGGCGATGGCGCCCGCCCGGGGCTTCGAGCCCGAACCGCCTCCCATGGGGAG
GCTGATGGCGCCTATTCTTGA
>ABBP01001115.1/6777-6711 Burkholderia pseudomallei 112 PMP 6xxBPSxxll2- 1115, whole genome shotgun sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC APPENDIX I
Fluoride Sensing Riboswitches Sequences unaligned
>CP000 Oil .2 /224010-224076 Burkholderia mallei ATCC 23344 chromosome 2, complete sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>AAIR02000026.1/4415-4481 Burkholderia mallei JHU ctg_1099471777536, whole genome shotgun sequence.
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>ABQR01000024.1/52165-52090 Clostridiales bacterium 1_7_47_FAA cont2.24, whole genome shotgun sequence .
TGTTTACAGGGGAATGAGGTTCTCCCTTAGTGTCATACTAAAACCGCTTATTTAAGCTGA
TGACTTCTGTGTGATT
>CP001215.1/4826181-4826240 Bacillus anthracis str. CDC 684, complete genome .
ATAATCATAGGCGATGGAGTTCGCCATAAACGCTGCTTAGCTAATGACTCCTACCAGTAT
>AAFX01041918.1/509-577 Environmental sequence XZS67508.xl, whole genome shotgun sequence.
AAGTCCGAAGGCGGTGGAGTCCGCCACAACCGTCATGCGCCTGCCGTGGCCAATGACTCC
TGCGGGGTT
>AAH001000134.1/6134-6200 Burkholderia mallei GB8 horse 4 contig_456, whole genome shotgun sequence .
GCGGCGACCGGAGATGGCATGCCTCCGTACAACCGCCGGCGAGCCGGCTGATGATGCCTA
CGCGTTC
>CP001196.1/1408427-1408488 Oligotropha carboxidovorans OM5 strain OM5, complete genome.
ATTGTTGATGGGGATGGAGTTCCCCCGATAACCGCCGCAAGGCTGATGACTCCTACCGGA
CG
>ACAN01000070.1/46719-46650 Enterococcus faecalis HIP11704 contl.70, whole genome shotgun sequence.
TAGCAACATGGTGATGGTGTTCACCACGAACTATTTATTGGACGAATAAATTAATGACGC
CTACCAAACG
>AACY023333559.1/192-117 Marine metagenome ctg_1101668140910, whole genome shotgun sequence.
AATAGTGGCGGGAATGAAGTGCTCCCTTCATATGATATGAAAACCGCAAACACGTGCTGA
TGACTTCTACGATTTT
>BAAV01007000.1/1090-1152 Human gut metagenome DNA, contig sequence: Fl- T_007000.
CTTTGCACCGGCTATGGGATTAGCCTTTAACCGCCCTTCGGAGCTGATGATCCCTACAAT
TTG
>BAAV01024370.1/684-614 Human gut metagenome DNA, contig sequence: Fl- T_024370.
TCAGNNATCGGGAATGATGTTCTCCCCTGGGATTCCAAAACCGCTGTTTCGCTGATGACG
TCTGCTTTTTT
>CP001111.1/1611287-1611224 Stenotrophomonas maltophilia R551-3, complete genome .
GGCGACGATGGGGATGGGGCTCCCCCGATAACCGCCTGAGAAGGCTGATGGCTCCTGCCA
AGAC
>AADL01000572.1/2844-2904 Leptospirillum sp . Group III AMC_Cont572, whole genome shotgun sequence.
CCGGCGGATGGCGATGGGGTTCGCCTAATCCCGCATTTCGGGTGATGACTCCTACCATGG
G

Claims

1. A UC Fi-biosensor comprising a genetically engineered bacterial cell capable of heterologously and/or natively expressing histidine kinase 1363, and U-sensitive transcriptional regulator 1362;
wherein the bacterial cell is an engineered bacterial cell including a 1362 U-sensing genetic reportable molecular component and/or a 1362 U-sensing/U-neutralizing genetic molecular component, each comprising a U-sensitive promoter in a configuration wherein the U-sensitive promoter directly initiates expression of the U-sensing reportable molecular component and/or of the U-sensing U-neutralizing molecular component in presence of bioavailable U;
wherein the U-sensitive promoter is
a 1362 U-sensitive prompter comprising a U-sensitive transcriptional 1362 binding site having a DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO:l), in which
Ni is C or T, preferably C;
N2 is G or A, preferably G;
N3 is T or C, preferably T;
N4 is C;
Ns is A or G, preferably A;
Nό is G or C, preferably G;
N7 C or G;
Ns is any nucleotide;
N9 is any nucleotide;
Niois any nucleotide;
N11 is any nucleotide;
N12 is T or C;
Ni3 is G; Ni4 is T or C, preferably T;
Nis is C;
N16 IS A or C, preferably A;
Nn is G; and
N18 is C or G,
and in which Ni to N17 selected independently
and wherein the genetically modified bacterial cell is an engineered bacterial cell further comprising an F-sensing riboswitch within at least one of
- the U- sensing reportable genetic molecular component in a configuration wherein the U-sensing reportable genetic molecular component, is transcribed in presence of an effective amount of bioavailable fluoride,
- an F-sensing reportable genetic molecular component in a configuration wherein the F- sensing reportable genetic molecular component is transcribed in presence of an effective amount of bioavailable fluoride, and
- an F sensitive genetic circuit in which at least one molecular component is a reportable molecular component (and in particular one or more reportable genetic molecular component and/or one or more reportable cellular molecular component), the reportable molecular component expressed when the genetic circuit operates according to the circuit design in presence of bioavailable F.
2. The UO2F2 biosensor of claim 1, wherein the U-sensitive promoter further comprises a UzcR binding site having a DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide,
the UzcR binding site inserted at a location downstream of a transcription start site of the U- sensitive promoter.
3. The UO2F2 biosensor of claim 2, further comprising a UzcY gene and/or UzcZ gene inserted at a location downstream of a transcription start site of the UzcR U-sensitive promoter.
4. The UO2F2 biosensor of claims 2 or 3, wherein the bacteria are capable of natively expressing endogenous MarR family repressors such as MarRl and/or MarR2 genes and at least one gene of the endogenous MarR family, preferably all genes of the endogenous MarR family, is knocked out.
5. The UO2F2 biosensor of any one of claims 2 to 4, wherein the bacteria are capable of natively expressing endogenous urtAP genes and at least one gene of the endogenous urtAP, preferably all genes of the endogenous urtAP, is knocked out.
6. A UO2F2 biosensor of any one of claims 1 to 5, wherein the F-sensing riboswitch comprises a crcB motif or an ericFmotif.
7. A UO2F2 biosensor of any one of claims 1 to 6, wherein the F-sensing riboswitch comprises a F-sensing riboswitch having sequence SEQ ID NO: 2509.
8. A UO2F2 biosensor of any one of claims 1 to 7, wherein the F-sensing riboswitch comprises a F-sensing riboswitch having sequence SEQ ID NO: 2510 or SEQ ID NO: 2512.
9. A UO2F2 biosensor of any one of claims 1 to 8, wherein the F-sensing riboswitch comprises a F-sensing riboswitch having sequence SEQ ID NO: 2511 or SEQ ID NO: 2513.
10. A UO2F2 biosensor of any one of claims 1 to 9, wherein the F-sensing riboswitch comprises a F-sensing riboswitch having any one of SEQ ID NO: 205 TO SEQ ID NO 1988 and and SEQ ID No: 1999 to SEQ ID NO: 2231..
11. The UO2F2 biosensor of any one of claims 1 to 10, wherein the F-sensing riboswitch comprises a F-sensing riboswitch Sphingomonas sp. MM-1, having sequence SEQ ID NO: 2514 ,an F-sensing riboswitch from Sphingomonas sp, 67-36 havng sequence SEQ ID NO: 2525 or the F-sensing riboswitch is the F-sensing riboswitch from Pseudomonas Syringae having sequence SEQ ID NO: 2526.
12. A UO2F2 biosensor of any one of claims 1 to 6 wherein the bacterial cell comprises an endogenous F-sensing riboswitch and the endogenous F-sensing riboswitch is knocked out.
13. The UO2F2 biosensor of any one of claims 1 to 12, wherein in the sequence SEQ ID NO:l:
Ni is C; and/or
N2 is G; and/or N3 is T; and/or
Ns is A; and/or
Nό is G; and/or
Ni4 is T; and/or
Ni6 is A.
14. The UO2F2 biosensor of any one of claims 1 to 13, wherein the U-sensitive transcriptional 1362 binding site has a sequence selected from the group consisting of SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26 and SEQ ID NO:27.
15. The UO2F2 biosensor of anyone of claims 1 to 14, wherein the U-sensitive promoter further comprises nucleotides N19N20N21, downstream of SEQ ID NO: 1 wherein N19 is any nucleotide; N2o is any nucleotide; and N21 is G (SEQ ID NO: 83).
16. The UO2F2 biosensor of claim 15, wherein Nis of the regulator direct repeat is located about - 17 to about -40 upstream of a transcription start site.
17. The UO2F2 biosensor of any one of claims 1 to 16, wherein the U-sensitive promoter is P1361 or Pphyt-
18. The UO2F2 biosensor of any one of claims 1 to 17, wherein biosensor comprises, the F- sensing reportable genetic molecular component and the 1362 -sensing genetic reportable molecular component and/or a 1362 U-sensing/U-neutralizing genetic molecular component, and the F- sensing reportable genetic molecular component are operatively connected to a same or a different reportable molecular component in U- sensing genetic genetic circuit.
19. The UO2F2 biosensor of claim 15, wherein the cell is further natively and/or heterologously expressing histidine kinase UzcS, and response regulator UzcS, and the U-sensing genetic circuit further comprises a UzcR U-sensing genetic reportable molecular component and/or a UzcR U-sensing/U-neutralizing genetic molecular component, each comprising a UzcR U-sensitive promoter in a configuration wherein the U-sensitive promoter directly initiates expression of the U-sensing reportable molecular component and/or of the U- sensing U-neutralizing molecular component in presence of bioavailable U,
the UzcR U sensitive promoter comprising a UzcR binding site having DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide, and in some embodiments any one of N7-N11 can independently be A.
20. The UO2F2 biosensor of claim 18 or 19, wherein the U-sensitive genetic circuit comprises one or more AND gates.
21. The UO2F2 biosensor of claim 20, wherein at least one of the AND gates is an in-series AND gate.
22. The UO2F2 biosensor of claim 20 or 21, wherein at least one of the AND gates is an in parallel AND gate.
23. The UO2F2 biosensor of any one of claims 20 to 22, wherein two or more in series AND gates and/or in parallel AND gates are connected by activating, inhibiting, binding, or converting reactions.
24. The UO2F2 biosensor of any one of claims 20 to 23, wherein at least one of the AND gates is selected from the group consisting of an HRP AND gate, a bacterial two-hybrid AND gate, a tripartite GFP AND gate, and a FRET sensor AND gate.
25. The UO2F2 biosensor of any one of claims 1 to 24, wherein the UO2F2 biosensor is configured to detect and/or neutralize bioavailable U present in a target environment at a concentration of 100 nM or greater, between 100 nM and 1 mM, or greater than 1 pM.
26. The method of any one of claims 1 to 24, wherein the U02F2biosensor is configured to detect bioavailable F present in a target environment at a concentration 10 pM, or greater, between 50 pM and 60 pM, greater than 1 nanomolar and between 100 nanomolar and 1 mM.
27. The UO2F2 biosensor of any one of claims 1 to 26, wherein the reportable molecular component is capable of being detected using fluorescence, luminescence, chemiluminescence, colorimetric analysis, radioactivity, or electrical.
28. The UO2F2 biosensor of any one of claims 1 to 27, wherein the U-neutralizing molecular component is configured to decrease or eliminate toxicity of U by bioreduction, biomineralization, bioaccumulation, and/or biosorption.
29. The UO2F2 biosensor of any one of claims 1 to 28, wherein the bacterial cell is a proteobacterial cell.
30. The UO2F2 biosensor of claim 29, wherein the proteobacterial cell is an alphaproteobacteria, a betaproteobacteria, or a gammaproteobacteria.
31. The UO2F2 biosensor of claim 29, wherein the proteobacterial cell is a Caulobacteridae cell.
32. The UO2F2 biosensor of claim 29, wherein the proteobacterial cell is a Caulobacter crescentus cell.
33. The UO2F2 biosensor of claim 32 wherein the Caulobacter crescentus cell is a member of a strain selected from the group consisting of NA1000, CB15, and OR37.
34. A UO2F2- sensing system comprising:
one or more of the UO2F2- biosensors of any one of claims 1 to 33 operatively connected to an electronic signal transducer adapted to convert a UO2F2- biosensor reportable molecular component output into an electronic output.
35. A method of detecting and reporting and/or neutralizing bioavailable U comprising:
contacting one or more of the UO2F2 biosensors of any one of claims 1 to 33, or the UO2F2— sensing system of claim 34, with a target environment comprising one or more target ranges of U concentration, the contacting performed for a time and under conditions to detect and report and/or neutralize bioavailable U in the target environment.
36. A UO2F2- sensing genetic reportable component comprising:
one or more U sensitive promoters comprising a U-sensitive 1362 binding site having a DNA sequence
N1N2N3N4N5N6N7N8N9N10N11N12N13N14N15N16N17N18 (SEQ ID NO:l), in which
Ni is C or T, preferably C; N2 is G or A, preferably G;
N3 is T or C, preferably T;
N4 is C;
Ns is A or G, preferably A;
Nό is G or C, preferably G;
N7 C or G;
Ns is any nucleotide;
N9 is any nucleotide;
Niois any nucleotide;
N11 is any nucleotide;
N12 is T or C;
Ni3 is G;
Ni4 is T or C, preferably T;
Nis is C;
N16 IS A or C, preferably A;
Nn is G; and
Nis is C or G,
and in which Ni to N17 selected independently
together with
a U- sensing reportable molecular component,
wherein at least one of the one or more U- sensitive promoters and the U- sensing reportable molecular component, comprises an F-sensing riboswitch in a single output configuration wherein the U-sensing reportable genetic molecular component, is transcribed in presence of an effective amount of bioavailable fluoride,
and wherein the one or more U sensitive promoters and the U-sensing reportable molecular component are in a configuration wherein the one or more U sensitive promoters directly initiate expression of the U-sensing reportable molecular component in presence of bioavailable U and bioavailable Fluoride.
37. UC Fi-sensing genetic reportable component of claim 36, wherein the one or more U sensitive promoters comprise
a U sensitive UzcR binding site, having a DNA sequence:
CATTACN7N8N9N10N11N12TTAA (SEQ ID NO:2) wherein N7-N12 is independently any nucleotide,
38. A UO2F2- sensing genetic reportable component of claim 36 wherein the one or more U- sensitive promoters and the U- sensing reportable molecular component are configured to provide a UO2F2- sensing gene cassette is an expression cassette.
39. A UO2F2- sensing genetic reportable component of claim 38, wherein the U02F2-sensing gene cassette is comprised within a vector.
PCT/US2020/016654 2019-02-04 2020-02-04 Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems Ceased WO2020163388A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201962801077P 2019-02-04 2019-02-04
US62/801,077 2019-02-04

Publications (1)

Publication Number Publication Date
WO2020163388A1 true WO2020163388A1 (en) 2020-08-13

Family

ID=71837401

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2020/016654 Ceased WO2020163388A1 (en) 2019-02-04 2020-02-04 Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems

Country Status (2)

Country Link
US (2) US11898211B2 (en)
WO (1) WO2020163388A1 (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11608536B2 (en) 2017-11-17 2023-03-21 Lawrence Livermore National Security, Llc Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems
US12571058B2 (en) 2019-02-04 2026-03-10 Lawrence Livermore National Security, Llc Biosensors for detecting and/or neutralizing bioavailable uranium and related U-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11541105B2 (en) 2018-06-01 2023-01-03 The Research Foundation For The State University Of New York Compositions and methods for disrupting biofilm formation and maintenance
CN111808874B (en) * 2019-04-11 2022-10-14 中国科学院微生物研究所 Encoding gene of phosphotriesterase 8047-PTE and application thereof
CN120424841A (en) * 2025-04-21 2025-08-05 兰州大学 A recombinant bacterium overexpressing uranyl binding protein and its application

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20050114923A1 (en) * 2003-07-01 2005-05-26 Edenspace Systems Corporation Plant biosensor systems
US8697388B2 (en) * 2007-02-08 2014-04-15 U.S. Department Of Energy Heavy metal biosensor
WO2019099934A2 (en) * 2017-11-17 2019-05-23 Lawrence Livermore National Security, Llc Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9580713B2 (en) 2011-09-17 2017-02-28 Yale University Fluoride-responsive riboswitches, fluoride transporters, and methods of use
WO2020163388A1 (en) 2019-02-04 2020-08-13 Lawrence Livermore National Security, Llc Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20050114923A1 (en) * 2003-07-01 2005-05-26 Edenspace Systems Corporation Plant biosensor systems
US8697388B2 (en) * 2007-02-08 2014-04-15 U.S. Department Of Energy Heavy metal biosensor
WO2019099934A2 (en) * 2017-11-17 2019-05-23 Lawrence Livermore National Security, Llc Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
HILLSON, N. J. ET AL.: "Caulobacter crescentus as a whole-cell uranium biosensor", APPL. ENVIRON. MICROBIOL., vol. 73, no. 23, 2007, pages 7615 - 7621, XP002711438, DOI: 10.1128/AEM.01566-07 *
PARK, D. M. ET AL.: "Identification of a U/Zn/Cu responsive global regulatory two-component system in Caulobacter crescentus", MOLECULAR MICROBIOLOGY, vol. 104, no. 1, 2017, pages 46 - 64, XP055631295, DOI: 10.1111/mmi.13615 *

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11608536B2 (en) 2017-11-17 2023-03-21 Lawrence Livermore National Security, Llc Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems
US12571058B2 (en) 2019-02-04 2026-03-10 Lawrence Livermore National Security, Llc Biosensors for detecting and/or neutralizing bioavailable uranium and related U-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems

Also Published As

Publication number Publication date
US20200248277A1 (en) 2020-08-06
US12571058B2 (en) 2026-03-10
US20240376556A1 (en) 2024-11-14
US11898211B2 (en) 2024-02-13

Similar Documents

Publication Publication Date Title
WO2020163388A1 (en) Biosensors for detecting and/or neutralizing bioavailable uranium and related u-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems
Baker et al. Diversity, ecology and evolution of Archaea
Abby et al. Candidatus Nitrosocaldus cavascurensis, an ammonia oxidizing, extremely thermophilic archaeon with a highly mobile genome
Herbold et al. Ammonia‐oxidising archaea living at low pH: insights from comparative genomics
Zhou et al. How sulphate-reducing microorganisms cope with stress: lessons from systems biology
Alves et al. Ammonia oxidation by the arctic terrestrial thaumarchaeote Candidatus Nitrosocosmicus arcticus is stimulated by increasing temperatures
Scanlan et al. Ecological genomics of marine picocyanobacteria
Riesenfeld et al. Metagenomics: genomic analysis of microbial communities
Ghosh et al. Microbial metagenomics: current advances in investigating microbial ecology and population dynamics
Schübbe et al. Complete genome sequence of the chemolithoautotrophic marine magnetotactic coccus strain MC-1
Jarrell et al. Major players on the microbial stage: why archaea are important
Dumont et al. Identification of a complete methane monooxygenase operon from soil by combining stable isotope probing and metagenomic analysis
Palenik et al. Coastal Synechococcus metagenome reveals major roles for horizontal gene transfer and plasmids in population diversity
Youssef et al. Partial genome assembly for a candidate division OP11 single cell from an anoxic spring (Zodletone Spring, Oklahoma)
Melnyk et al. The perchlorate reduction genomic island: mechanisms and pathways of evolution by horizontal gene transfer
Rowe et al. In situ electrochemical enrichment and isolation of a magnetite‐reducing bacterium from a high pH serpentinizing spring
US12297513B2 (en) Biosensors for detecting and/or neutralizing bioavailable uranium and related U-sensitive genetic molecular components, gene cassettes, vectors, genetic circuits, compositions, methods and systems
Schwarzenlander et al. Characterization of DNA transport in the thermophilic bacterium Thermus thermophilus HB27
Yoon et al. Parallel evolution of transcriptome architecture during genome reorganization
Zhou et al. StressChip as a high-throughput tool for assessing microbial community responses to environmental stresses
US20210332350A1 (en) Recombinase Genome Editing
Zhong et al. Pan‐genome study of Thermococcales reveals extensive genetic diversity and genetic evidence of thermophilic adaption
Lindner et al. A synthetic glycerol assimilation pathway demonstrates biochemical constraints of cellular metabolism
Jin et al. The role of three‐tandem Pho Boxes in the control of the C‐P lyase operon in a thermophilic cyanobacterium
Zik et al. Dual transposon sequencing profiles the genetic interaction landscape in bacteria

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20752667

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20752667

Country of ref document: EP

Kind code of ref document: A1