WO2025019889A1 - Novel hydrogenases - Google Patents
Novel hydrogenases Download PDFInfo
- Publication number
- WO2025019889A1 WO2025019889A1 PCT/AU2024/050777 AU2024050777W WO2025019889A1 WO 2025019889 A1 WO2025019889 A1 WO 2025019889A1 AU 2024050777 W AU2024050777 W AU 2024050777W WO 2025019889 A1 WO2025019889 A1 WO 2025019889A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- nucleic acid
- sequence
- seq
- acid molecule
- amino acid
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/11—DNA or RNA fragments; Modified forms thereof; Non-coding nucleic acids having a biological activity
- C12N15/52—Genes encoding for enzymes or proenzymes
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
- C12N9/0004—Oxidoreductases (1.)
- C12N9/0067—Oxidoreductases (1.) acting on hydrogen as donor (1.12)
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12P—FERMENTATION OR ENZYME-USING PROCESSES TO SYNTHESISE A DESIRED CHEMICAL COMPOUND OR COMPOSITION OR TO SEPARATE OPTICAL ISOMERS FROM A RACEMIC MIXTURE
- C12P3/00—Preparation of elements or inorganic compounds except carbon dioxide
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Y—ENZYMES
- C12Y112/00—Oxidoreductases acting on hydrogen as donor (1.12)
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Y—ENZYMES
- C12Y112/00—Oxidoreductases acting on hydrogen as donor (1.12)
- C12Y112/99—Oxidoreductases acting on hydrogen as donor (1.12) with other acceptors (1.12.99)
- C12Y112/99006—Hydrogenase (acceptor) (1.12.99.6)
-
- H—ELECTRICITY
- H01—ELECTRIC ELEMENTS
- H01M—PROCESSES OR MEANS, e.g. BATTERIES, FOR THE DIRECT CONVERSION OF CHEMICAL ENERGY INTO ELECTRICAL ENERGY
- H01M8/00—Fuel cells; Manufacture thereof
- H01M8/16—Biochemical fuel cells, i.e. cells in which microorganisms function as catalysts
-
- H—ELECTRICITY
- H01—ELECTRIC ELEMENTS
- H01M—PROCESSES OR MEANS, e.g. BATTERIES, FOR THE DIRECT CONVERSION OF CHEMICAL ENERGY INTO ELECTRICAL ENERGY
- H01M8/00—Fuel cells; Manufacture thereof
- H01M8/22—Fuel cells in which the fuel is based on materials comprising carbon or oxygen or hydrogen and other elements; Fuel cells in which the fuel is based on materials comprising only elements other than carbon, oxygen or hydrogen
Definitions
- Novel hydrogenases Field of the invention [0001] The invention relates to enzymes and polypeptide complexes for generating energy from hydrogen or for generating hydrogen, nucleic acid molecules encoding the same, devices and systems comprising the same, and uses thereof.
- Related application [0002] This application claims priority from Australian provisional application no. 2023902336, the entire contents of which are incorporated herein by reference.
- Sequence listing [0003] A sequence listing in ST.26 format is filed herewith, the entire contents of which are incorporated herein by reference.
- Background of the invention [0004] Molecular hydrogen (H2) is heralded as a future green energy carrier. In a biological context, this energy-rich gas already plays a central role in bioenergetics and evolution has driven an elaborate hydrogen economy.
- H2 ⁇ 2 H + + 2 e- hydrogen gas
- metalloenzymes called hydrogenases.
- Three hydrogenases have independently evolved in microorganisms, namely the [FeFe]-, [NiFe]-, and [Fe]-hydrogenases, which differ in their metal cofactors and catalytic mechanism.
- H 2 serves multiple roles in microbial physiology.
- Microorganisms produce H2 to dispose of electrons during fermentation.
- Numerous bacteria and archaea also use electrons derived from H2 oxidation for respiration and carbon fixation.
- H2 was likely 1005373994 the primordial electron donor, but continues to have a central role in microbiology both as a desirable energy source and diffusible electron sink.
- H 2 exchange between bacteria and archaea underlies eukaryogenesis, as described in various syntrophy hypotheses.
- [FeFe]-hydrogenases are typically fast-acting, but oxygen-sensitive, and are best known for their roles in obligate anaerobes. These enzymes currently comprise four phylogenetically distinct groups (groups A to D), which can be further subdivided through two different schemes based on domain architecture and genetic organisation.
- [NiFe]-hydrogenases are extraordinarily structurally and functionally diverse enzymes encoded by bacteria and archaea across all ecosystems. They are presently subdivided into four major groups (groups 1 to 4) and 29 subgroups that each differ in their phylogeny, genetic organisation, and physiological roles.
- the catalytic (large) subunit and electron-relaying iron-sulfur (small) subunit of the [NiFe]-hydrogenase associate with other subunits depending on the subgroup; the different complexes formed can mediate respiration, fermentation, energy-conversion, electron-bifurcation, carbon fixation, and H2 sensing processes.
- [Fe]- hydrogenases are a much narrower lineage that contribute to archaeal methanogenesis.
- the three hydrogenase classes are phylogenetically unrelated, despite having some similar structural features, and are not thought to genetically or structurally associate.
- [NiFe]-hydrogenases were present in the last universal common ancestor (LUCA), whereas [FeFe]-hydrogenases are proposed to have evolved later in fermentative bacteria. [0008] Several genomic studies have also suggested that [FeFe]-hydrogenases may be encoded by uncultivated DPANN archaea.
- the present invention provides an isolated, synthetic or purified nucleic acid molecule encoding a hydrogenase from archaea of a lineage Thermoplasmatota, Asgardarchaeota, Thermoproteota, EX4484-52, Aenigmarchaeota / QMZS01, Nanoarchaeota, Altarchaeota, lainarchaeota, or Micrarchaeota.
- the archaea is one shown in Figure 1.
- the nucleic acid molecule encodes a hydrogenase from group A1, A3, B, C1, E, F or G.
- the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 1 to 150, or a sequence at least about 80% identical thereto.
- the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 300, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 151 to 300.
- the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 1 to 1005373994 150 (e.g. in Table 1), or a sequence at least about 80% identical thereto, and one or more HydB, HydC, HydD, HyhL or HyhS sequences as set forth in any one of SEQ ID Nos: 301 to 463, 627 to 629 (e.g. in Table 3), or a sequence at least about 80% identical thereto.
- the HydA, HydB, HydC, HydD, HyhL and/or HyhS combinations are any one shown in the row 1 to 73 of Table 5.
- the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 300 (e.g. in Table 2), or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 151 to 300, and and one or more HydB, HydC, HydD, HyhL or HyhS sequences as set forth in any one of SEQ ID NOs: 464 to 626, 630 to 632 (e.g.
- the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 1 to 26, or a sequence at least about 80% identical thereto.
- the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 176, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 151 to 176.
- the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 27 to 70, or a sequence at least about 80% identical thereto.
- the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 27 to 70, or a sequence at least about 80% identical thereto, and one or both of a HydB and HydC sequence as set forth in any one of SEQ ID NOs: 301 to 334, 336 to 341, 343 to 386, 389 to 390, 393 to 396, 399 to 400, 404 to 407, 410 to 411, 415 to 416, 422 to 425, 1005373994 428 to 429, 432 to 433, 438 to 439, 443 to 444, 447 to 450, 452 to 453, and 457 to 458, or a sequence at least about 80% identical thereto.
- the HydA, HydB and HydC combinations are any one shown in rows 2 to 39 in Table 5.
- the combination shown in row 20 of Table 5 includes the HyhL.
- the combination shown in row 22 of Table 5 includes the HyhL.
- the combination shown in row 23 includes either HydB (SEQ ID NO: 343 or 344).
- the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 177 to 220, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 177 to 220.
- the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 177 to 220, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 177 to 220, and and one or both of a HydB and HyC sequence as set forth in any one of SEQ ID NOs: 464 to 497, 499 to 504, 506 to 549, 552 to 553, 556 to 559, 562 to 563, 567 to 570, 573 to 574, 578 to 579, 585 to 588, 591 to 592, 595 to 596, 601 to 602, 606 to 607, 610 to 613, 615 to 616, and 620 to 621, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence
- the HydA, HydB and HydC combinations are any one shown in rows 2 to 39 in Table 5.
- the combination shown in row 20 of Table 5 includes HyhL.
- the combination shown in row 22 of Table 5 includes HyhL.
- the combination shown in row 23 includes HydB as set forth in SEQ ID NO: 506 or 507.
- the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 71 to 82, or a sequence at least about 80% identical thereto.
- the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 71 to 82 (e.g. in Table 1), or a sequence at least about 80% identical thereto, and one HydC sequence as set forth in in Table 3, or a sequence at least about 80% identical thereto.
- the HydA and HydC combinations are any one shown in rows 40 to 50 in Table 5.
- the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 221 to 232, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 221 to 232.
- the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 221 to 232 (e.g.
- the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in SEQ ID NOs: 83, or a sequence at least about 80% identical thereto.
- the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in SEQ ID NOs: 233, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of SEQ ID NO: 233.
- the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 84 to 112, or a sequence at least about 80% identical thereto.
- the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 234 to 262, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 234 to 262.
- the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 113 to 133, or a sequence at least about 80% identical thereto.
- the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 113 to 133, or a sequence at least about 80% identical thereto, and one or more HydB, HydC, HydD or HyhL sequences as set forth in SEQ ID NOs: 301 to 463, or a sequence at least about 80% identical thereto.
- the HydA, HydB, HydC, HydD and HyhL combinations are any one shown in rows 51 to 70 in Table 5.
- the HydA, HydB, HydC, HydD and HyhL are shown in row 69.
- the HyhL shown in row 61 is SEQ ID NO: 419 or 420.
- the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 263 to 283, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 263 to 283.
- the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 263 to 283, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 263 to 283, and at least one or more HydB, HydC, HydD or HyhL sequences as set forth in SEQ ID NOs: 464 to 626 or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the HydB, HydC, HydD or HyhL sequence as set forth in SEQ ID NOs: 464 to 626.
- the HydA, HydB, HydC, HydD and HyhL combinations are any one shown in rows 51 to 70 in Table 5.
- the amino acid 1005373994 sequence of the HydA, HydB, HydC, HydD and HyhL polypeptide is shown in row 69.
- the HyhL shown in row 61 is SEQ ID NO: 582 or 583.
- the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 134 to 136, or a sequence at least about 80% identical thereto.
- the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 134 to 136, or a sequence at least about 80% identical thereto, and one or both of a HyhL and HyhS sequence as set forth in SEQ ID NOs: 335, 388, 392, 398, 402, 409, 413 to 414, 418 to 420, 427, 431, 435 to 436, 440 to 441, 445, 455 to 456, 460 to 463, and 627 to 629, or a sequence at least about 80% identical thereto.
- the HydA, HyhL and HyhS combinations are any one shown in rows 71 to 73 in Table 5.
- each of rows 71 to 73 in Table 5 further include a HyhS.
- Row 71 further includes nucleotide SEQ ID NO: 628.
- Row 72 further includes nucleotide SEQ ID NO: 629.
- Row 73 further includes nucleotide SEQ ID NO: 627.
- the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 284 to 286, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 284 to 286.
- the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 284 to 286, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 284 to 286, and at least one or both of a HyhL and HyhS sequence as set forth in SEQ ID NOs: 498, 551, 555, 561, 565, 572, 576 to 577, 581 to 583, 590, 594, 598 to 599, 603 to 604, 608, 618 to 619, 623 to 626, and 630 to 632, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to a HyhL sequence as set forth in SEQ ID NOs: 498, 551, 555, 561, 565, 572, 576 to
- the HydA, HyhL and HyhS combinations are any one shown in rows 71 to 73 in Table 5.
- rows 71 to 73 in Table 5 further include a HyhS.
- Row 71 further includes nucleotide SEQ ID NO: 628 and amino acid SEQ ID NO: 631.
- Row 72 further includes nucleotide SEQ ID NO: 629 and amino acid SEQ ID NO: 632.
- Row 73 further includes nucleotide SEQ ID NO: 627 and amino acid SEQ ID NO: 630.
- the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 137 to 143, or a sequence at least about 80% identical thereto.
- the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 287 to 293, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 287 to 293.
- the nucleic acid molecule does not include any other nucleotide sequence or encode for any other protein that occurs in the native archea organism from which the HydA was derived.
- a nucleic acid construct comprising a nucleic acid molecule as described herein.
- the construct is synthetic, recombinant or isolated. More preferably, the construct may comprise one or more heterologous promoters for enabling expression of the nucleic acid(s) comprised in the construct.
- the construct comprises a heterologous sequence encoding a tag for enabling the purification of one or more proteins encoded by the nucleic acid sequences as described herein.
- the construct is in the form of a vector or plasmid.
- a nucleic acid molecule or construct may comprise a codon optimised sequence for enabling expression of the nucleic acid in a heterologous host.
- the nucleic acid molecule may comprise a codon optimised sequence for enabling expression of the nucleic acid(s) or operon in a cell that is not an archaea, or not an archaea of a lineage as shown in Figure 1.
- the heterologous host is E. coli. 1005373994 [0047]
- a nucleic acid molecule comprising, consisting essentially of or consisting of a nucleotide sequence set forth in any one of SEQ ID Nos: 144 to 150.
- the nucleic acid molecule comprises, consists essentially of or consists of a nucleotide sequence set forth in any one of SEQ ID Nos: 144 to 150 without a sequence encoding the C-terminal SSGWSHPQFEK.
- the nucleic acid molecule is isolated, recombinant or synthetic.
- the nucleic acid molecule is codon optimised for expression in a heterologous host, such as E.
- nucleic acid molecule does not encode the C-terminal SSGWSHPQFEK as depicted in SEQ ID Nos: 294 to 300.
- a cell comprising a nucleic acid molecule as described herein, or comprising a nucleic acid construct as described herein.
- the cell is a microorganism.
- the cell is E. coli.
- a particularly preferred strain of E. coli is DE3.
- the cell is a recombinant cell.
- an isolated microorganism comprising a nucleic acid as described herein, or a construct as described herein, wherein the isolated microorganism is capable of oxidising hydrogen.
- the isolated microorganism is any strain shown in Figure 1 or described herein.
- the invention also provides a method of producing a hydrogenase as described herein, the method including culturing a cell or microorganism comprising a nucleic acid molecule as described herein, or comprising a nucleic acid construct as described herein under conditions to allow expression of the hydrogenase.
- the expression is inducible expression.
- the expressed protein is harvested from the cell under anaerobic conditions to minimise or prevent inactivation of the produced hydrogenase by atmospheric oxygen.
- the cell or microorganism is cultured at a temperature of less than 37 ⁇ C after induction of expression, preferably the temperature is less than about 30°C, or less than about 20 ⁇ C or about 20 ⁇ C.
- the polypeptide comprises, consists essentially of or consists of an amino acid sequence selected from any one of SEQ ID NOs: 151 to 293, or a functional equivalent, homology or derivative thereof having a sequence at least 80% identical thereto.
- the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in any one of rows 1 to 73 of Table 5.
- the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 20 without the HyhL shown.
- the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 22 without the HyhL shown.
- the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 23 with either HydB shown.
- the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 22 without the HyhL shown. In one embodiment, the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 61 with either HyhL shown. In one embodiment, the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 69 with either HydB, HydC, HydD or HyhL shown. In relation to HyhS each of rows 71 to 73 in Table 5 further include a HyhS. Row 71 further includes amino acid SEQ ID NO: 631. Row 72 further includes amino acid SEQ ID NO: 632.
- Row 73 further includes amino acid SEQ ID NO: 630.
- an isolated, recombinant or purified polypeptide capable oxidising hydrogen wherein the polypeptide comprises, consists essentially of or consists of an amino acid sequence selected from any one of SEQ ID NOs: 287 to 293, or a functional equivalent, homology or derivative thereof having a sequence at least 80% identical thereto.
- an isolated, recombinant or purified polypeptide capable oxidising hydrogen wherein the polypeptide comprises, consists essentially of or consists of an amino acid sequence selected from any one of SEQ ID NOs: 294 to 300, or a functionally equivalent, homology or derivative thereof having a sequence at least 80% identical thereto.
- the polypeptide is capable of H + reduction catalysis and/or H 2 gas oxidation.
- an anaerobically matured enzyme, or enzyme complex for oxidising hydrogen
- the enzyme or enzyme complex comprises proteins comprising the amino acid sequences as set forth in any one or more of SEQ ID NOs: 151 to 300, or functional equivalent, homologs or derivatives having at least 80% sequence identity thereto.
- the enzyme complex may also be referred to herein as a multiprotein enzyme complex.
- the enzyme or enzyme complex has been matured using a synthetic mimic.
- An example of a synthetic mimic is [2Fe] adt ([Fe2(azadithiolate)(CO)4(CN)2] 2 ⁇ ).
- the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of the subunits or amino acid sequences as referred to in any one of the rows 1 to 73 of Table 5.
- the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of amino acid sequences referred to in row 20 without the HyhL shown.
- the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 22 without the HyhL shown.
- the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 23 with either HydB shown.
- the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 22 without the HyhL shown. In one embodiment, the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 61 with either HyhL shown. In one embodiment, the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 69 with either HydB, HydC, HydD or HyhL shown. In relation to HyhS each of rows 71 to 73 in Table 5 further include a HyhS.
- Row 1005373994 71 further includes amino acid SEQ ID NO: 631.
- Row 72 further includes amino acid SEQ ID NO: 632.
- Row 73 further includes amino acid SEQ ID NO: 630.
- the enzyme or enzyme complex is capable of H + reduction catalysis and H2 gas oxidation.
- the complex comprises a HydA subunit and a HydC subunit; or a HydA subunit, a HydB subunit and a HydC subunit; or a HydA subunit and HyhL subunit; a HydA subunit, a HydB subunit, a HydC subunit, a HydD subunit and a HyhL subunit; or a HydA subunit, a HyhL and a HyhS subunit.
- the complex is arranged substantially as depicted in the Examples and Figures herein.
- the present invention provides a method of activating a polypeptide, enzyme or enzyme complex as described herein.
- activating a polypeptide, enzyme or enzyme complex includes contacting the polypeptide, enzyme or enzyme complex with iron and sulphur sources under anaerobic conditions.
- iron and sulfur sources are ferrous ammonium sulfate and L-cysteine.
- the iron and sulfur sources are both added in 1.5-fold, or about 1.5-fold, molar excess to the desired number of Fe-atoms to be added.
- a method of activating a hydrogenase is outlined in Example 1 as described below.
- a functional equivalent, homolog or derivative of a HydA protein is a protein which retains the same or substantially the same function of a HydA protein as herein described; and/or a functional equivalent, homolog or derivative of an enzyme complex is a protein which retains the same or substantially the same function of an enzyme complex as herein described.
- the phrase “at least 80% sequence identity” should be understood to provide basis for at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity.
- a sequence having “at least 80% sequence identity” may consist of a sequence which is about 80%, about 81%, about 82%, about 1005373994 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98% or about 99% sequence identity.
- a functional equivalent of any amino acid sequence disclosed herein may include an equivalent amino acid sequence wherein the N-terminal methionine residue is cleaved in the final protein product.
- any of the amino acid sequences set forth in Tables 2 and 4 should be understood to provide basis for the same amino acid sequences but which do not comprise an N terminal methionine.
- an amino acid sequence disclosed herein will be understood to provide basis for an identical amino acid sequence wherein the N- terminal methionine residue is cleaved in the final protein product.
- any of the amino acid sequences set forth in Tables 2 and 4 should be understood to provide basis for the same amino acid sequences but which do not comprise an N terminal methionine.
- a method for converting hydrogen to electrons comprising contacting a source of hydrogen with an isolated or recombinant microorganism as described herein, or with an isolated, recombinant or purified polypeptide, enzyme, or enzyme complex as described herein.
- a method for producing hydrogen (H 2 ) comprising contacting a source of protons with an isolated or recombinant microorganism as described herein, or with an isolated, recombinant or purified polypeptide, enzyme, or enzyme complex as described herein.
- a method of generating energy from a source of hydrogen comprising contacting a source of hydrogen with an isolated or recombinant microorganism as described herein, or with an isolated, recombinant or purified polypeptide, enzyme or enzyme complex as described herein.
- an isolated or recombinant microorganism as described herein, or an isolated, recombinant or purified polypeptide, enzyme or enzyme complex as described herein for generating energy from a source of hydrogen.
- the isolated or recombinant microorganism, or the isolated, recombinant or purified polypeptide, enzyme or enzyme complex is immobilised or encapsulated.
- the enzyme may be immobilised on an electrode. Methods for generating protein films for immobilising proteins on electrodes are well known in the art and are further described herein.
- the enzyme is covalently bound to the surface of an electrode.
- the electrode is comprised of pyrolytic graphite edge (PGE), encapsulated in epoxy.
- PGE pyrolytic graphite edge
- the microorganism may be encapsulated or immobilised.
- a device for converting hydrogen to electrons comprising an immobilised polypeptide, enzyme or enzyme complex as described herein.
- the device may otherwise be referred to herein as an electrode.
- the polypeptide, enzyme or enzyme complex is provided in the device/electrode covalently bound to the surface of an electricity-conducting material.
- the material is comprised of pyrolytic graphite edge (PGE), encapsulated in epoxy.
- a system for oxidising hydrogen comprising an isolated or recombinant microorganism of the invention or an isolated, purified or recombinant polypeptide, enzyme or enzyme complex of the invention, a source of hydrogen, and means for detecting the oxidation of hydrogen.
- the system further comprises one or more cofactors for accepting the electrons produced by the oxidation of hydrogen.
- the co-factor may be a small molecule, or an enzyme that accepts the electrons and transfers to molecular oxygen.
- the means for enabling detection of the oxidation of hydrogen comprises a means for direct measurement of electrical current (such as an amperometer).
- the present invention further contemplates the provision of a hydrogen sensor comprising a polypeptide, enzyme or enzyme complex described herein, wherein the sensor is useful for measuring/detecting hydrogen.
- the sensor device comprises a polypeptide, enzyme or enzyme complex immobilised on an electrode, such that upon oxidation of hydrogen by the polypeptide, enzyme or enzyme complex upon contact with hydrogen, and electrons are produced which enter an electrical circuit associated with the electrode. The magnitude of the electrical current generated could be correlated with the amount of hydrogen present.
- the system may also be calibrated using known amounts of hydrogen to enable the determination of unknown quantities of hydrogen in test samples.
- a fuel cell comprising a device or electrode of the invention.
- an air-powered device comprising a fuel cell of the invention.
- the fuel cell generates energy from hydrogen present in the ambient air thereby providing energy to the device.
- the use comprises: - obtaining or having obtained a reference data set in the form of data comprising a measure of electron production by the polypeptide or enzyme present in the system, complex or device in response to known quantities of hydrogen, - contacting the polypeptide or enzyme in the system, complex or device with a test sample comprising an unknown quantity of hydrogen, 1005373994 - comparing the electron production resulting from the contacting, to the electron production in the reference data set to thereby determine the concentration of hydrogen present in the test sample.
- the left portion of the figure shows a maximum-likelihood phylogenomic tree (model LG+F+G4) based on the concatenated 15 ribosomal marker proteins of archaeal genomes that encode [FeFe]-hydrogenases. Results are shown for the 118 (out of 130) genomes that are at least 60% complete, less than 5% contaminated, and contain at least 75% of the 15 syntenic proteins. Branches are colour-coded encoding according to the respective phylum. Black circles indicate bootstrap support values over 80%. The middle portion shows the presence of key metabolic genes involved in different metabolic processes.
- Carbon fixation ATP- citrate lyase beta-subunit (AclB), acetyl- CoA synthase beta subunit (AcsB), propionyl- CoA synthetase (PrpE), 4-hydroxybutyryl- CoA dehydratase / vinylacetyl-CoA-delta- isomerase (AbfD), CODH/ACS complex subunit delta (CdhD), CODH/ACS complex subunit gamma (CdhE), anaerobic carbon monoxide dehydrogenase catalytic subunit (CooS), type II/III ribulose-bisphosphate carboxylase (RbcL II/II), type III ribulose- bisphosphate carboxylase (RbcL III); respiration: reductive dehalogenase (RdhA), formaldehyde activating enzyme (Fae), glutathione-independent formaldehyde dehydrogenase (FdhA),
- FIG. 1 The right portion shows the diverse environments the archaeal genomes were retrieved from.
- Phylum QMZS01 was classified as Aenigmatarchaeota in GTDB R06- RS207 while Thermoproteota class EX4484 ⁇ 205 was proposed as Brockarchaeia.
- Figure 2 Archaea encode genetically and structurally diverse [FeFe]- hydrogenases. Catalytic domain structure, genetic organisation, and AlphaFold2- based structural modelling of representative [FeFe]-hydrogenases encoded in archaeal genomes.
- Predicted cofactors are positioned based on the structures of homologous proteins.
- a zoomed view of the H- cluster and conserved coordinating cysteine residues (C 1 to C 5 ) is shown for each group.
- C 1 to C 5 conserved coordinating cysteine residues
- FeS clusters within plausible electron transfer distance are connected by dashed lines.
- hydA [FeFe]-hydrogenase
- hydB diaphorase
- hydC thioredoxin
- hydD nuoG-like conduit protein
- hyd6TM uncharacterised 4 to 6-helix transmembrane protein associated with group A [FeFe]-hydrogenases; (His)[4Fe4S], (Cys)3His-ligated [4Fe4S] cluster binding domain; [2Fe2S], [2Fe2S] cluster binding domain; [4Fe4S], [4Fe4S] cluster binding domain; 2[4Fe4S], bacterial ferredoxin-like 2[4Fe4S] cluster binding domain; 6Cys, putative iron-sulfur cluster binding domain.
- HydC protein in group A1 gene cluster was not predicted to form a complex with HydA.
- Surface structures are used for the multisubunit group B and A3 [FeFe]-hydrogenases, with ribbon diagram versions provided in Figure 8.
- Figure 3 Three classes of [FeFe]-hydrogenases are catalytically active in archaea. (a) H2 gas production monitored from cell lysates in E. coli BL21(DE3) cells expressing group A1, B, and E [FeFe]-hydrogenases from archaea. The cell lysates were activated by addition of [2Fe] adt .
- H 2 was measured by GC after addition of methyl viologen 1005373994 and dithionite to activated cell lysates, set to pH 6.8 with 100 mM KPi buffer. Activities are normalized for number of cells used (nmol H 2 min -1 OD600 -1 ) and error bars reflect standard deviation from biological triplicates.
- the strain expressing prototypical CrHydA1 was used as a positive control while “Blank” represents the same strain but containing an empty vector.
- (b) FTIR spectra of the group E [FeFe]-hydrogenase from Ca. Forterrea multitransposorum (Fm) after heterologous expression, semisynthetic maturation with [2Fe] adt , and purification.
- the absorbance spectrum (top) indicates a CO inhibited di- ferrous H-cluster state (Hsox-CO).
- the difference spectrum (bottom) illustrates the transitions of Fm into catalytically active states through photoreduction (illumination after the addition of eosin Y as a photosensitizer and triethanolamine as a sacrificial electron donor).
- illumination bands associated with the highly oxidized CO-inhibited state decreased (grey bands), while new bands reflecting reduced and catalytically active H- cluster states appear, assigned to H ox H (cyan), H ox (blue) and H red (red) (spectra arranged chronologically from top to bottom).
- FIG. 4 [FeFe]- and [NiFe]-hydrogenases form unique complexes in archaea. (a) Catalytic domain structure and predicted operon encoding a putative complex of a group F [FeFe]-hydrogenase and group 3 [NiFe]-hydrogenase in Thermoplasmatota UBA147 sp002496385.
- hydA [FeFe]-hydrogenase
- hydB diaphorase
- hydC thioredoxin
- hydD nuoG-like conduit protein
- hyhL group 3 [NiFe]- hydrogenase catalytic subunit
- hyhS group 3 [NiFe]-hydrogenase small subunit.
- [FeS] clusters are numbered and labelled according to their subunit of origin (e.g. A1, A2, A3 originate from the HydA subunit).
- H2 levels were 1005373994 measured every 15 mins for 2 hours by gas chromatography after addition of methyl viologen and dithionite to activated cell lysates, set to pH 6.8 with 100 mM KPi buffer. Activities are normalized for number of cells used (nmol H 2 OD 600 -1 ) and error bars reflect standard deviations from two biological triplicates. “Blank” represents the same strain but containing an empty vector.
- Figure 5 [FeFe]-hydrogenases are diverse, ancient, and potentially ancestral in archaea. Maximum-likelihood phylogenetic tree of the catalytic subunit (HydA) of [FeFe]-hydrogenases and novel hybrid hydrogenases.
- the tree was constructed based on 3,677 amino acid sequences using the LG+F+G4 model. The numbers at the branches indicate the ultrafast bootstrap support values. The tree was rooted using the NADH-quinone oxidoreductase subunit D (NuoD) and formate dehydrogenase alpha chain (FdhA) from Methylorubrum extorquens.
- NuoD NADH-quinone oxidoreductase subunit D
- FdhA formate dehydrogenase alpha chain
- Subclass (protein domain structure-based scheme) and subgroup (protein phylogeny-based scheme) classification of [FeFe]- hydrogenases are based on Land et al.2020 and Greening et al.2016, respectively, with proposed modifications.
- FIG. 7 Genetic organization of 136 archaeal [FeFe]-hydrogenases. Up to 10 genes upstream and downstream of the [FeFe]-hydrogenase (hydA) are shown. Gene length is shown to scale.
- hydA [FeFe]-hydrogenase
- hydS [FeFe]-hydrogenase small 1005373994 subunit
- hydB [FeFe]-hydrogenase diaphorase subunit
- hydC [FeFe]-hydrogenase thioredoxin subunit
- hydD [FeFe]-hydrogenase nuoG-like conduit protein
- hydF [FeFe]- hydrogenase H-cluster maturation GTPase
- hyd6TM uncharacterised 4 to 6- helix transmembrane protein associated with group A [FeFe]-hydrogenases; hyhL / hoxH, group 3 [NiFe]-hydrogenase catalytic subunit; hyhS / hoxY, group 3 [NiFe]- hydrogenase small subunit; hyhD
- Figure 8 AlphaFold2 models of archaeal [FeFe]-hydrogenases.
- FIG. 10 SDS-PAGE visualising the molecular weights of the heterologously expressed archaeal [FeFe]-hydrogenases.
- Expression constructs with verified sequences were transformed in chemically competent E. coli BL21(DE3).
- 1005373994 Protein bands are shown from before induction with IPTG (B), after induction (Name- Subclass and with the expected kDa size in parenthesis), and lysate or supernatant after cell lysis and centrifugation (L).
- the bands in each after-induction lane corresponded well with the expected molecular weights in kDa.
- Both group A1 [FeFe]-hydrogenases (Mu and Ia) had the highest expression and solubility levels.
- FIG. 11 Isolation and reconstitution of [4Fe-4S]+ cluster of Fm.
- Figure 12 FTIR difference spectra and redox state kinetics of Fm [FeFe]- hydrogenase.
- the super-oxidised CO inhibited H sox -CO species grey bands
- HoxH cyan bands
- H ox blue bands
- the oxidised species get further reduced to the [4Fe4S] cluster reduced state H red ’ (red bands). Illumination that facilitates photoreduction was applied for 88 seconds (compare b).
- Figure 13 Phylogenetic tree of HydA sequences constructed with LG+FO+R model.
- the maximum-likelihood phylogenetic tree was constructed based on 3,677 amino acid sequences of the catalytic subunit (HydA) of [FeFe]-hydrogenases and novel hybrid hydrogenases.
- the model finder was used with default parameters and no mixed model testing; the best-fit model identified was LG+FO+R. Ultrafast bootstrap support values are denoted on the branches.
- FIG. 14 Phylogenetic tree of HydA sequences constructed with LG+C60 model.
- the maximum-likelihood phylogenetic tree was constructed based on 3,677 amino acid sequences of the catalytic subunit (HydA) of [FeFe]-hydrogenases and novel hybrid hydrogenases.
- the best model including the LG+C60 was utilised to test the amino acid replacement rate to vary across different sites. Ultrafast bootstrap support values are denoted on the branches.
- Figure 15 Phylogenetic tree of the [FeFe]-hydrogenase maturase HydE. Different colours show archaeal (red), bacterial (black), and eukaryotic (green) sequences. Evolutionary history was inferred using the Maximum Likelihood method and JTT matrix-based model with 50 bootstrap replicates and midpoint rooting.
- Figure 16 Phylogenetic tree of the [FeFe]-hydrogenase maturase HydF. Different colours represent archaeal (red), bacterial (black), and eukaryotic (green) sequences. Evolutionary history was inferred using the Maximum Likelihood method and JTT matrix-based model with 50 bootstrap replicates and midpoint rooting.
- Figure 17 Phylogenetic tree of the [FeFe]-hydrogenase maturase HydG. Different colours represent archaeal (red), bacterial (black), and eukaryotic (green) sequences. Evolutionary history was inferred using the Maximum Likelihood method and JTT matrix-based model with 50 bootstrap replicates and midpoint rooting.
- Figure 18 Phylogenetic tree of the [NiFe]-hydrogenase large / catalytic subunit (HyhL) with focus on group 3 [NiFe]-hydrogenases. The subunits predicted to associate with [FeFe]-hydrogenases are shown in red (for group F [FeFe]- 1005373994 hydrogenases) and purple (for group G [FeFe]-hydrogenases). Evolutionary history was inferred using the Maximum Likelihood method and JTT matrix-based model with 50 bootstrap replicates and midpoint rooting. [0108] Figure 19.
- Tables comprising sequence information Table 1: Nucleotide sequences of HydA SEQ Description Subgroup, Sequence ID subclass NO: 1 AQRS01000037.1 A1, M1 ATGGGTTCGATTGAGGACGTGAATGCCGCGCTCGCGGATGAAGGGA _10, nucleotide AAATGGTCATGGCGCAGGTAGCGCCCGCGGTTAGGGTTACTATCGG CGAGGAGTTCGGCCTTCCGGCGGGAACAATTGTGACGAAAAAGCTC GTGGGCGCGTTGAGGCAGGCCGGCTTTGAAAAGGTGTTTGACACCT CCGTTGCCGCGGATATTGTAACAATTGAGGAAGGAACGGAATTCCTG AACAGGCTCGAGGACCAGGAGGACCTTCCATTGCTGACTTCCTGCTG CCCTGCATCGGTTTTTTTTGTTGAGAACACTTTCCCGAAATTTTTGCAC CACTTCTGCACTGTTAAAAGCCCGCAGCAGGGCATGGGCTCGCTCAT AAAAACCTATTACGCGCAGGA
- Row 71 further includes nucleotide SEQ ID NO: 628 and amino acid SEQ ID NO: 631.
- Row 72 further includes nucleotide SEQ ID NO: 629 and amino acid SEQ ID NO: 632.
- Row 73 further includes nucleotide SEQ ID NO: 627 and amino acid SEQ ID NO: 630 .
- DPANN have notably evolved ultraminimal fermentative [FeFe]-hydrogenases, providing them with a genomically streamlined way to efficiently dispose excess electrons produced during carbohydrate fermentation. It is also remarkable that even the Asgard archaeon expresses a hitherto-overlooked [FeFe]-hydrogenase, which likely enables it to mediate syntrophic hydrogen exchange with its methanogenic partner.
- the inventors heterologously purified and biochemically, electrochemically, and spectroscopically characterised unique groups of [FeFe]-hydrogenases from archaea only known for their genomes.
- nucleic acids As used herein an "isolated" nucleic acid molecule is a nucleic acid molecule that is identified and separated from at least one contaminant nucleic acid molecule with which 1005373994 it is ordinarily associated in the natural source of the polypeptide encoding nucleic acid. An isolated nucleic acid molecule is other than in the form or setting in which it is found in nature. Isolated nucleic acid molecules therefore are distinguished from the nucleic acid molecule as it exists in natural cells.
- nucleic acid molecule includes nucleic acid molecules contained in cells that ordinarily express the nucleic acid where, for example, the nucleic acid molecule is in a chromosomal location different from that of natural cells.
- nucleic acid molecule and “polynucleotide” may be used interchangeably herein and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogues thereof.
- Non-limiting examples of polynucleotides include a gene, a gene fragment, messenger RNA (mRNA), cDNA, recombinant polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers.
- a nucleic acid sequence which “encodes” a selected polypeptide is a nucleic acid molecule which is transcribed (in the case of DNA) and translated into a polypeptide in vivo when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a start codon at the 5′ (amino) terminus and a translation stop codon at the 3′ (carboxy) terminus.
- a transcription termination sequence may be located 3′ to the coding sequence.
- Polynucleotides of the invention can be synthesised according to methods well known in the art, as described by way of example in Sambrook et al (1989, Molecular Cloning—a laboratory manual; Cold Spring Harbor Press).
- “codon optimised” refers to optimisation of the DNA sequence to resemble the codon usage of genes in host microorganism. In preferred embodiments, the codon usage in the sequence is optimised to resemble that of highly expressed E. coli genes.
- the polynucleotide molecules of the present invention may be provided in the form of an expression cassette which includes control sequences operably linked to the inserted sequence, thus allowing for expression of the polypeptide.
- These expression cassettes are typically provided within vectors (e.g., plasmids or recombinant vectors).
- a suitable vector may be any vector which is capable of carrying a sufficient amount of genetic information, and allowing expression of a polypeptide of the invention.
- the present invention thus includes expression vectors that comprise such polynucleotide sequences.
- Expression vectors are routinely constructed in the art of molecular biology and may for example involve the use of plasmid DNA and appropriate initiators, promoters, enhancers and other elements which may be necessary, and which are positioned in the correct orientation, in order to allow for expression of a desired polypeptide.
- Other suitable vectors would be apparent to persons skilled in the art.
- a polypeptide of the invention may be provided by delivering such a vector to a cell and allowing transcription from the vector to occur.
- the skilled person will be familiar with standard techniques for delivery such expression vectors to a cell, including transformation techniques and the like.
- the vector may be a plasmid.
- the plasmid is a high copy number plasmid or a low copy number plasmid.
- Vectors are well known in the art and may include cloning vectors, expression vectors, etc.
- a cloning vector is a recombinant nucleic acid construct which is able to replicate autonomously or integrated in the genome in a host cell, and which is further characterized by one or more endonuclease restriction sites at which the vector may be cut in a determinable fashion and into which a desired DNA sequence may be ligated such that the new recombinant vector retains its ability to replicate in the host cell.
- An expression vector is a recombinant nucleic acid construct into which a desired DNA sequence may be inserted by restriction and ligation such that it is operably joined to regulatory sequences and may be expressed as an RNA transcript.
- Vectors may further contain one or more marker sequences suitable for use in the identification of cells which have or have not been transformed or transfected with the vector.
- Markers include, for example, genes encoding proteins which increase or decrease either resistance or sensitivity to antibiotics or other compounds, genes which encode polypeptides or enzymes whose activities are detectable by standard assays known in the art (e.g., ⁇ - galactosidase, luciferase or alkaline phosphatase), and genes which visibly affect the phenotype of transformed or transfected cells, hosts, colonies or plaques (e.g., fluorescent proteins such as green fluorescent protein).
- Preferred vectors are those 1005373994 capable of autonomous replication and expression of the structural gene products present in the DNA segments to which they are operably joined.
- a coding sequence and regulatory sequences are said to be "operably” joined or linked when they are covalently linked in such a way as to place the expression or transcription of the coding sequence under the influence or control of the regulatory sequences.
- two DNA sequences are said to be operably joined or linked if induction of a promoter in the 5' regulatory sequences results in the transcription of the coding sequence and if the nature of the linkage between the two DNA sequences does not (1) result in the introduction of a frame-shift mutation, (2) interfere with the ability of the promoter region to direct the transcription of the coding sequences, or (3) interfere with the ability of the corresponding RNA transcript to be translated into a protein.
- a promoter region would be operably joined or linked to a coding sequence if the promoter region were capable of effecting transcription of that DNA sequence such that the resulting transcript can be translated into the desired protein or polypeptide.
- regulatory sequences needed for gene expression may vary between species or cell types, but shall in general include, as necessary, 5' non-transcribed and 5' non-translated sequences involved with the initiation of transcription and translation respectively, such as a TATA box, capping sequence, CAAT sequence, and the like.
- 5' non-transcribed regulatory sequences will include a promoter region which includes a promoter sequence for transcriptional control of the operably joined gene.
- Regulatory sequences may also include enhancer sequences or upstream activator sequences as desired.
- the vectors of the invention may optionally include 5' leader or signal sequences. The choice and design of an appropriate vector is within the ability and discretion of one of ordinary skill in the art.
- a “promoter” is a nucleotide sequence which initiates and regulates transcription of a polypeptide-encoding polynucleotide. Promoters can include inducible promoters (where expression of a polynucleotide sequence operably linked to the promoter is induced by an analyte, cofactor, regulatory protein, etc.), repressible promoters (where expression of a polynucleotide sequence operably linked to the promoter is repressed by an analyte, cofactor, regulatory protein, etc.), and constitutive promoters.
- promoter or “control element” includes full-length promoter regions and functional (e.g., controls transcription or translation) segments of these regions.
- the nucleic acids of the present invention are preferably operably linked to promoters such that the subject enzymes are expressed in the cell when cultured under suitable conditions for enabling consumption of hydrogen, as described herein.
- the promoters may be specific for individual bacterial cell species.
- the promoter may be a heterologous promoter which increases the expression of the gene above the typical expression level observed in the cell.
- the promoter may be an inducible promoter.
- a polynucleotide, expression cassette or vector according to the present invention may additionally comprise a signal peptide sequence.
- the signal peptide sequence is generally inserted in operable linkage with the promoter such that the signal peptide is expressed and facilitates secretion of a polypeptide encoded by coding sequence also in operable linkage with the promoter. It may further be understood that in any embodiment, any of the exemplary expression cassettes, vectors or sequences described herein may be further modified so as to not include a signal peptide sequence. [0132] Any appropriate expression vector (e.g., as described in Pouwels et al., Cloning Vectors: A Laboratory Manual (Elsevier, N.Y.: 1985)) and corresponding suitable host can be employed for production of recombinant polypeptides.
- Expression hosts include, but are not limited to, bacterial species within the genera Escherichia, Bacillus, Pseudomonas, Salmonella, host cell systems and the like. The skilled person is aware that the choice of expression host has ramifications for the type of polypeptide produced.
- the cell is engineered or selected (e.g., as described herein) to produce or have altered, optionally increased, production of a molecule of interest.
- the cell comprises a deletion or mutation of one or more genes (e.g., one or more regulatory or competing metabolic genes as described herein). In other examples, the one or more genes that are deleted or mutated are in a competing pathway.
- Mutations can be single or multiple point mutations, additions, partial internal deletions, N-terminal or C-terminal deletions (truncations), or complete deletions, all of which can affect amino acid sequence encoded the gene(s).
- Deletions or mutations can be made using standard methods in the art. Mutations can be non-random, partially random or random, or a combination of these mutations. 1005373994 For example, for a partially random mutation, the mutation(s) may be confined to a certain portion of the nucleic acid molecule encoding a polypeptide in which mutation(s) are to be made.
- Protein production and purification means the polypeptide that has been identified and separated and/or recovered from a component of its natural environment. Contaminant components of its natural environment are materials that would typically interfere with diagnostic or therapeutic uses for the polypeptide, and may include enzymes, hormones, and other proteinaceous or non-proteinaceous solutes.
- the polypeptide will be purified (1) to a degree sufficient to obtain at least 15 residues of N-terminal or internal amino acid sequence by use of a spinning cup sequenator, or (2) to homogeneity by SDS-PAGE under non-reducing or reducing conditions using Coomassie blue or, preferably, silver stain.
- Isolated protein includes polypeptide in situ within recombinant cells, since at least one component of the polypeptide natural environment will not be present. Ordinarily, however, isolated polypeptide will be prepared by at least one purification step.
- a "fragment” is a portion of a polypeptide of the present invention that retains substantially similar functional activity or substantially the same biological function or activity as the polypeptide, which can be determined using assays described herein.
- Percent (%) amino acid sequence identity or “percent (%) identical” with respect to a polypeptide sequence, i.e. a polypeptide of the invention defined herein, is defined as the percentage of amino acid residues in a candidate sequence that are identical with the amino acid residues in the specific polypeptide of the invention, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and not considering any conservative substitutions as part of the sequence identity.
- percent amino acid sequence identity X/Y100, where X is the number of amino acid residues scored as identical matches by the sequence alignment program's or algorithm's alignment of A and B and Y is the total number of amino acid residues in B.
- the percent amino acid sequence identity of A to B will not equal the percent amino acid sequence identity of B to A.
- the determination of percent identity between two sequences can be accomplished using a mathematical algorithm.
- a nonlimiting example of a mathematical algorithm utilized for the comparison of two sequences is the algorithm of Karlin and Altschul (1990) Proc. Natl. Acad. Sci. USA 87:2264, modified as in Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877. Such an algorithm is incorporated into the BLASTN and BLASTX programs of Altschul et al. (1990) J. MoI.
- Gapped BLAST in BLAST 2.0
- PSI-Blast can be used to perform an iterated search that detects distant relationships between molecules. See Altschul et al. (1997) supra.
- the default parameters of the respective programs e.g., BLASTX and BLASTN
- Alignment may also be performed manually by inspection.
- Another non- limiting example of a mathematical algorithm utilized for the comparison of sequences is the ClustalW algorithm (Higgins et al.
- ClustalW compares sequences and aligns the entirety of the amino acid or DNA sequence, and thus can provide data about the sequence conservation of the entire amino acid sequence.
- the ClustalW algorithm is used in several commercially available DNA/amino acid analysis software packages, such as the ALIGNX module of the Vector NTI Program Suite (Invitrogen Corporation, Carlsbad, CA). After alignment of amino acid sequences with ClustalW, the percent amino acid identity can be assessed.
- a non-limiting example of a software program useful for analysis of ClustalW alignments is GENEDOCTM or JalView (http://www.jalview.org/). GENEDOCTM allows assessment of amino acid (or DNA) similarity and identity between multiple proteins.
- the polypeptide desirably comprises an amino end and a carboxyl end.
- the polypeptide can comprise D-amino acids, L-amino acids or a mixture of D- and L-amino acids.
- the D-form of the amino acids is particularly preferred since a polypeptide comprised of D-amino acids is expected to have a greater retention of its biological activity in vivo.
- the polypeptide can be prepared by any of a number of conventional techniques.
- the polypeptide can be isolated or purified from a naturally occurring source or from a recombinant source. Recombinant production is preferred.
- a DNA fragment encoding a desired peptide can be subcloned into an appropriate vector using well-known molecular genetic techniques (see, e.g., Maniatis et al., Molecular Cloning: A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory, 1982); Sambrook et al., Molecular Cloning A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory, 1989).
- the fragment can be transcribed and the polypeptide subsequently translated in vitro.
- kits also can be employed (e.g., such as manufactured by Clontech, Palo Alto, Calif.; Amersham Pharmacia Biotech Inc., Piscataway, N.J.; InVitrogen, Carlsbad, Calif., and the like).
- the polymerase chain reaction optionally can be employed in the manipulation of nucleic acids.
- conservative substitution refers to the replacement of an amino acid present in the native sequence in the peptide with a naturally or non- naturally occurring amino acid or a peptidomimetic having similar steric properties.
- the conservative substitution should be with a naturally occurring amino acid, a non- naturally occurring amino acid or with a peptidomimetic moiety which is also polar or hydrophobic (in addition to having the same steric properties as the side-chain of the replaced amino acid).
- amino acids that may be considered to be conservative substitutions for one another: [0144] 1) Alanine (A), Serine (S), Threonine (T); [0145] 2) Aspartic acid (D), Glutamic acid (E); [0146] 3) Asparagine (N), Glutamine (Q); [0147] 4) Arginine (R), Lysine (K); [0148] 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); and [0149] 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W).
- non-conservative substitution or a “non-conservative residue” as used herein refers to replacement of the amino acid as present in the parent sequence by another naturally or non-naturally occurring amino acid, having different electrochemical and/or steric properties.
- the side chain of the substituting amino acid can be significantly larger (or smaller) than the side chain of the native amino acid being substituted and/or can have functional groups with significantly different electronic properties than the amino acid being substituted.
- non-conservative substitutions of this type include the substitution of phenylalanine or cycohexylmethyl glycine for alanine, isoleucine for glycine, or -NH-CH[(-CH2)5-COOH]-CO- for aspartic 1005373994 acid.
- Non-conservative substitution includes any mutation that is not considered conservative.
- a non-conservative amino acid substitution can result from changes in: (a) the structure of the amino acid backbone in the area of the substitution; (b) the charge or hydrophobicity of the amino acid; or (c) the bulk of an amino acid side chain.
- substitutions generally expected to produce the greatest changes in protein properties are those in which: (a) a hydrophilic residue is substituted for (or by) a hydrophobic residue; (b) a proline is substituted for (or by) any other residue; (c) a residue having a bulky side chain, e.g., phenylalanine, is substituted for (or by) one not having a side chain, e.g., glycine; or (d) a residue having an electropositive side chain, e.g., lysyl, arginyl, or histadyl, is substituted for (or by) an electronegative residue, e.g., glutamyl or aspartyl.
- a hydrophilic residue is substituted for (or by) a hydrophobic residue
- a proline is substituted for (or by) any other residue
- a residue having a bulky side chain e.g., phenylalanine
- an electropositive side chain e
- Alterations of the native amino acid sequence to produce mutant polypeptides can be done by a variety of means known to those skilled in the art.
- site-specific mutations can be introduced by ligating into an expression vector a synthesized oligonucleotide comprising the modified site.
- oligonucleotide-directed site-specific mutagenesis procedures can be used, such as disclosed in Walder et al., Gene 42: 133 (1986); Bauer et al., Gene 37: 73 (1985); Craik, Biotechniques, 12-19 (January 1995); and U.S. Pat. Nos.4,518,584 and 4,737,462.
- N-terminal and C-terminal are used herein to designate the relative position of any amino acid sequence or polypeptide domain or structure to which they are applied. The relative positioning will be apparent from the context. That is, an "N-terminal” feature will be located at least closer to the N-terminus of the polypeptide molecule than another feature discussed in the same context (the other feature possible referred to as "C-terminal” to the first feature). Similarly, the terms “5'-" and “3'-” can be used herein to designate relative positions of features of polynucleotides.
- a recombinant polypeptide made in accordance with the methods of the present invention may also be modified by, conjugated or fused to another moiety to facilitate purification of the polypeptides, or for use in enzymatic assays using methods known in the art.
- a polypeptide of the invention may be modified by glycosylation, 1005373994 acetylation, phosphorylation, amidation, derivatization by known protecting/blocking groups, proteolytic cleavage, etc.
- Modifications contemplated herein include, but are not limited to, modification to side chains, incorporating of unnatural amino acids and/or their derivatives during polypeptide synthesis and the use of crosslinkers and other methods which impose conformational constraints on the polypeptides of the invention.
- side chain modifications contemplated by the present invention include modifications of amino groups such as by reductive alkylation by reaction with an aldehyde followed by reduction with NaBH 4 ; amidination with methylacetimidate; acylation with acetic anhydride; carbamoylation of amino groups with cyanate; trinitrobenzylation of amino groups with 2, 4, 6-trinitrobenzene sulphonic acid (TNBS); acylation of amino groups with succinic anhydride and tetrahydrophthalic anhydride; and pyridoxylation of lysine with pyridoxal-5-phosphate followed by reduction with NaBH 4 .
- modifications of amino groups such as by reductive alkylation by reaction with an aldehyde followed by reduction with NaBH 4 ; amidination with methylacetimidate; acylation with acetic anhydride; carbamoylation of amino groups with cyanate; trinitrobenzylation of amino groups with 2, 4, 6-trinitrobenzene sulphonic acid (TNBS);
- the guanidine group of arginine residues may be modified by the formation of heterocyclic condensation products with reagents such as 2,3-butanedione, phenylglyoxal and glyoxal.
- the carboxyl group may be modified by carbodiimide activation via O- acylisourea formation followed by subsequent derivatisation, for example, to a corresponding amide.
- Sulphydryl groups may be modified by methods such as carboxymethylation with iodoacetic acid or iodoacetamide; performic acid oxidation to cysteic acid; formation of a mixed disulphides with other thiol compounds; reaction with maleimide, maleic anhydride or other substituted maleimide; formation of mercurial derivatives using 4- chloromercuribenzoate, 4-chloromercuriphenylsulphonic acid, phenylmercury chloride, 2- chloromercuri-4-nitrophenol and other mercurials; carbamoylation with cyanate at alkaline pH.
- Tryptophan residues may be modified by, for example, oxidation with N- bromosuccinimide or alkylation of the indole ring with 2-hydroxy-5-nitrobenzyl bromide or sulphenyl halides.
- Tyrosine residues on the other hand, may be altered by nitration with tetranitromethane to form a 3-nitrotyrosine derivative.
- Modification of the imidazole ring of a histidine residue may be accomplished by alkylation with iodoacetic acid derivatives or N-carboethoxylation with diethylpyrocarbonate.
- Examples of incorporating unnatural amino acids and derivatives during protein synthesis include, but are not limited to, use of norleucine, 4-amino butyric acid, 4-amino- 3-hydroxy-5-phenylpentanoic acid, 6-aminohexanoic acid, t-butylglycine, norvaline, phenylglycine, ornithine, sarcosine, 4-amino-3-hydroxy-6-methylheptanoic acid, 2-thienyl alanine and/or D-isomers of amino acids.
- a list of unnatural amino acids contemplated herein is shown in Table 6.
- Certain steps may be performed under aerobic conditions, such as initial culture of a host cell, and then certain steps may be performed under anaerobic conditions, such as harvesting the expressed hydrogenases to minimise or prevent inactivation by atmospheric oxygen.
- aerobic conditions such as initial culture of a host cell
- anaerobic conditions such as harvesting the expressed hydrogenases to minimise or prevent inactivation by atmospheric oxygen.
- culturing of recombinant host cells for production of recombinant proteins will be carried out at a temperature that is optimal for the growth and expression of proteins in the organism.
- the optimum temperature for growth of E. coli and related bacterial organisms is about 37 °C and the temperature for growth of yeasts for producing recombinant proteins is about 30-32°C.
- Genetically engineered or “genetically modified” refers to any cell modified by any recombinant DNA or RNA technology. In other words, the cell has been transfected, transformed, or transduced with a recombinant polynucleotide molecule, and thereby been altered so as to cause the cell to alter expression of a desired protein.
- Methods and vectors for genetically engineering host cells are well known in the art; for example, various techniques are illustrated in Current Protocols in Molecular Biology, Ausubel et al., eds. (Wiley & Sons, New York, 1988, and quarterly updates).
- exogenous polynucleotides is intended to mean polynucleotides that are not derived from naturally occurring polynucleotides in a given organism. Exogenous polynucleotides may be derived from polynucleotides present in a different organism. In accordance with the present invention, an E.
- coli cell may be genetically modified with a nucleic acid construct which contains one or more exogenous polynucleotides, encoding one or more enzymes which enable the cell to produce hydrogen.
- the exogenous polynucleotides may be heterologous or homologous.
- heterologous refers to a molecule or activity derived from a source other than the referenced species whereas "homologous” refers to a molecule or activity derived from the host microbial organism. Accordingly, exogenous expression of a nucleic acid molecule of the invention can be through the use of either or both a heterologous or homologous nucleic acid molecule.
- a particularly preferred heterologous microorganism for expression of a nucleic acid molecule or construct of the invention is E. coli.
- a strain of E. coli particularly preferred is DE3.
- DE3 indicates that the host is a lysogen of ⁇ DE3, and therefore carries a chromosomal copy of the T7 RNA polymerase gene under control of the lacUV5 promoter.
- a DE3 strain is suitable for production of protein from target genes cloned in pET vectors by induction with IPTG.
- coli strain is OrigamiTM B – with genotype F- ompT hsdS B (r B - m B -) gal dcm lacY1 ahpC (DE3) gor522:: Tn10 trxB (Kan R , Tet R ).
- the temperature of the culture is lowered to 20 o C to improve the yield of soluble protein.
- harvesting of the expressed polypeptide or hydrogenase and/or subsequent purification steps of the expressed polypeptide or hydrogenase are conducted in anaerobic conditions to prevent hydrogenase inactivation by atmospheric oxygen.
- Anaerobic conditions may be [O2] ⁇ 5 ppm.
- One or more other exemplary steps in the production and purification of a polypeptide of the invention are described in Example 1 below.
- the exogenous polynucleotides may be provided in one or more expression constructs (plasmid vectors).
- Methods of transforming microorganisms are well known in the art, and can include such non-limiting examples as electroporation, calcium chloride-, or lithium acetate-based methods. 1005373994
- the skilled person will be familiar with methods for confirming successful transformation of relevant constructs, as well as methods for determining whether the transformants possess the relevant enzyme activity provided by the encoded protein. For example, phosphofructokinase activity (and therefore inferring correct protein folding of the encoded protein) can be inferred using a commercially available enzyme assay kit.
- the skilled person will be familiar with standard techniques to confirm inhibition or deletion of the level of activity of a relevant protein or level of expression of the relevant gene.
- Selection marker genes refer to genetic material that encodes a protein necessary for the survival and/or growth of a host cell grown in a selective culture medium. Typical selection marker genes for use in microorganisms, including in E. coli are well known to the skilled person.
- the microorganism preferably an E.
- the microorganism of the invention or methods described herein may involve transformation of the microorganism with the required polynucleotides in order to generate a recombinant microorganism capable of oxidising hydrogen.
- the microorganism may then be harvested and stored under conditions suitable for storage of the microorganism (for example, at 4°C, -20°C, or -80°C in a suitable buffer) until required for hydrogen oxidation.
- the microorganism may be lyophilised until required for further use.
- the microorganism can be grown under conditions to enable expression of the required polynucleotides and then harvested, where necessary stored, and then resuspended in appropriate solutions to initiate bacterial oxidation of hydrogen.
- the bacteria may be immobilised or encapsulated. Methods for immobilisation or encapsulation of microorganisms are known, for example with the use of calcium alginate beads using standard techniques. The skilled person will be familiar with standard manual and mechanism techniques and equipment for bio- encapsulation, including by using a device such as the Inotech Encapsulator IE-50R (EncapBioSystems Inc), or Encapsulator B-390/B-395 pro (Buchi), or related systems.
- the recombinant microorganism does not need to be viable (i.e., capable of reproducing, “growing” or increasing in cell numbers) in order to be able to oxidise hydrogen in accordance with the present invention.
- the methods involve providing or generating a recombinant microorganism as herein described, culturing the microorganism under conditions and for a sufficient time to induce expression of the proteins required for oxidising hydrogen (e.g., the proteins forming the hydrogenase or anaerobic enzyme complex) and then inactivating the microorganism.
- the proteins required for oxidising hydrogen e.g., the proteins forming the hydrogenase or anaerobic enzyme complex
- the inactivated microorganisms remain intact, although it will be understood that this is not an essential requirement.
- Inactivated isolated, purified or recombinant microorganisms of the invention can be then be used to oxidise hydrogen, for example as described herein in the Examples.
- the skilled person will be familiar with methods for inactivating microorganisms so that the cells remain intact, but can still be utilised to oxidise hydrogen. Inactivation may be by gamma irradiation or by treatment with an antibiotic (such as mitomycin or similar).
- Systems and devices [0185] The present invention also provides systems and devices comprising the microorganisms of the invention, or reactor systems which include methods described herein for oxidising hydrogen and to preferably therefore generate energy from hydrogen. [0186] In certain embodiments, the systems and devices of the invention comprise systems and devices for use in protein film voltammetry.
- protein film voltammetry refers to a system for detecting/measuring electron transfer reactions that occur following oxidation of a substrate by an immobilised protein (typically adsorbed or covalently attached on an electrode). Further details relating to protein film voltammetry methods are described in the Examples herein, and also in Leger et al., Biochemistry, 42, 8653- 62, 2003.
- the present invention provides an electrode comprising an isolated, purified or recombinant polypeptide, enzyme or enzyme complex as described herein.
- the electrode comprises an electrode made of an electricity-conducting material (such as graphite), and an isolated, purified or recombinant polypeptide, enzyme or enzyme complex as described herein, attached thereto. Attachment of the enzyme to the electrode may be by any suitable method, including as described herein in the examples (eg, exposure of an abraded pyrolytic graphite edge electrode to a solution of enzyme in Tris buffer).
- an electricity-conducting material such as graphite
- Attachment of the enzyme to the electrode may be by any suitable method, including as described herein in the examples (eg, exposure of an abraded pyrolytic graphite edge electrode to a solution of enzyme in Tris buffer).
- conducting material refers to electricity-conducting materials that belong, for illustrative purposes and without limiting the scope of the invention, to the following group: metallic material (gold, copper, silver or platinum) or carbonaceous material (glassy carbon, “basal” pyrolytic carbon, “edge” pyrolytic carbon, carbon thread and carbon fabric, amongst other types of conducting carbon), optionally wherein on the surface thereof primary amines may be introduced by means of different techniques that allow for anchoring with the polypeptide, enzyme or enzyme complex.
- metallic material gold, copper, silver or platinum
- carbonaceous material glassy carbon, “basal” pyrolytic carbon, “edge” pyrolytic carbon, carbon thread and carbon fabric, amongst other types of conducting carbon
- the electrode material used is a carbonaceous material that belongs, for illustrative purposes and without limiting the scope of the invention, to the following group: glassy carbon, “basal” pyrolytic carbon, “edge” pyrolytic carbon, carbon thread and carbon fabric.
- Further methods for immobilising (or anchoring) hydrogenase enzymes to an electrode are described in the prior art, for example in US 20090142649, incorporated herein by reference.
- the invention further provides a fuel cell or an electrolytic cell comprising an electrode as described herein. 1005373994
- Fuel cells are electrochemical devices that convert the energy of a fuel directly into electrochemical and thermal energy.
- a fuel cell consists of an anode and a cathode, which are electrically connected via an electrolyte.
- a fuel such as, for example, hydrogen
- a fuel is fed to the anode where it is oxidized with the help of an electrocatalyst.
- an oxidant such as oxygen (or air)
- the electrochemical reactions which occur at the electrodes produce a current and thereby electrical energy.
- thermal energy is also produced which may be harnessed to provide additional electricity or for other purposes.
- the most common electrochemical reaction for use in a fuel cell is that between hydrogen and oxygen to produce water.
- the fuel cells of the present subject matter utilize hydrogen as a fuel wherein the source of hydrogen is provided from an exogenous source or simply atmospheric hydrogen, and the biocatalyst for conversion of the hydrogen to water is provided in the form of a purified or recombinant or isolated polypeptide, enzyme or enzyme complex of the present subject matter, or a recombinant or isolated microorganism of the invention.
- the source of hydrogen may be a waste gas stream, including syngas and biogas.
- the source of hydrogen may be ambient air.
- hydrogen is present in the fuel source in an amount of at least about 0.00005% by volume, or at least about 0.001% by volume, or at least about 0.01% by volume or at least about 0.1% by volume, preferably at least about 1% and more preferably at least about 5% by volume, for example about 10%, 20%, 30%, 40%, 50%, 75% or 90% by volume.
- an inert gas is used to form part of the fuel gas
- the inert gas is typically present in an amount of at least about 10%, such as at least about 25%, 50 % or 75% by volume, most preferably at least about 80% by volume.
- the fuel source is supplied from an optionally pressurized container of the fuel source in gaseous or liquid form.
- the fuel source is supplied to the electrode via an inlet, which can optionally comprise a valve.
- An outlet is also provided which enables used or waste fuel source to leave the fuel cell.
- the oxidant typically includes oxygen, although any other suitable oxidant can be used.
- the oxidant source typically provides the oxidant to the cathode in the form of a gas which includes the oxidant. In some embodiments, the oxidant can be provided in liquid form.
- the oxidant source also includes an inert gas, although the oxidant in its pure form can also be used.
- a mixture of oxygen with one or more gases such as nitrogen, helium, neon or argon can be used.
- the oxidant source can optionally comprise further components, for example alternative oxidants or other additives.
- An example of a suitable oxidant source is air.
- oxygen is present in the oxidant source in an amount of at least about 2% by volume, preferably at least about 5% and more preferably at least about 10% by volume or more.
- the oxidant source is supplied from an optionally pressurized container of. the oxidant source in gaseous or liquid form.
- the oxidant source is supplied to the electrode via an inlet, which optionally comprises a valve.
- the anode can be made of any conducting material for example stainless steel, brass or carbon, which can be graphite.
- the surface of the anode can, at least in part, be coated with a different material which facilitates adsorption of the catalyst.
- the surface onto which the catalyst is adsorbed is of a material which does not cause the hydrogenase to denature. Suitable surface materials include graphite, such as, for example, a polished graphite surface or a material having a high surface area such as carbon cloth or carbon sponge. Materials with a rough surface and/or with a high surface area are generally preferred.
- the cathode can be made of any suitable conducting material which will enable an oxidant to be reduced at its surface.
- materials used to form the cathode in conventional fuel cells can be used.
- An electrocatalyst (or bioelectrocatalyst) in the form of a polypeptide, enzyme or enzyme complex of the invention, is preferably present at the cathode.
- This electrocatalyst can, for example, be coated or adsorbed on the cathode itself, or it can be present in a solution surrounding the cathode.
- the fuel cell of the present subject matter is typically operated at a temperature of at least about 10°C, or at least about 20 °C or about 25°C, more preferably at least 1005373994 about 30°C. It is preferred that the fuel cell is operated at a temperature of from about 35 °C to about 65°C, such as from about 40°C to about 50°C.
- the present invention further contemplates the provision of a sensor comprising a polypeptide, enzyme or enzyme complex describe herein, wherein the sensor is useful for measuring/detecting hydrogen.
- the sensor device comprises a polypeptide, enzyme or enzyme complex immobilised on an electrode, such that upon oxidation of hydrogen by the polypeptide, enzyme or enzyme complex upon contact with hydrogen, and electrons are produced which enter an electrical circuit associated with the electrode.
- the magnitude of the electrical current generated could be correlated with the amount of hydrogen present.
- the system may also be calibrated using known amounts of hydrogen to enable the determination of unknown quantities of hydrogen in test samples.
- Kits [0204]
- the present invention also provides a kit comprising a polypeptide, enzyme or enzyme complex as described herein, or a device or sensor or electrode as described herein. [0205]
- the kit comprises written instructions for use in accordance with a method or system described herein.
- the kit may comprise buffers, co-factors and other components for enabling detection of hydrogen oxidation by the polypeptide, enzyme or enzyme complex of the invention and/or to enable detection and measurement of hydrogen in a test sample.
- buffers, co-factors and other components for enabling detection of hydrogen oxidation by the polypeptide, enzyme or enzyme complex of the invention and/or to enable detection and measurement of hydrogen in a test sample.
- Genome sources [0209] In this study, 130 archaeal genomes encoding [FeFe]-hydrogenases were analysed.40 genomes are novel metagenome-assembled genomes retrieved from our unpublished datasets at ggKbase (https://ggkbase-help.berkeley.edu/). 77 genomes were retrieved from the Genome Taxonomy Database (GTDB) R06-RS202 following a search of all 2,339 archaeal species representative genomes. GTDB was chosen as the data source as it offers a standardised taxonomy, a manageable search space due to pre-clustered sequences, and genomes pre-annotated with gene predictions.
- GTDB Genome Taxonomy Database
- MAGs came from previously reported study sites, namely Guaymas Basin hydrothermal vents (7 MAGs; Gulf of California, Mexico), Lac Lavin freshwater lake (6 MAGs; central France), Crystal Geyser (6 MAGs; Utah, USA), Chinese hot springs (5 MAGs; 2 from Yunnan Republic and 3 from Vietnamese Plateau, China), an aquifer (3 MAGs; Napa County, California, USA), Manure lagoon (2 MAGs, California, USA), an aquifer adjacent to the Colorado River (3 MAGs; Rifle, Colorado, USA), a wetland soil (2 MAGs; 1005373994 Napa County, California, USA), Alum Rock mineral spring (1 MAG; San Jose, California), Zodletone Spring (1 MAG; Anadarko, Oklahoma, USA), Azore Islands hot springs (1 MAGs, Portugal), a borehole (1 MAG, Muzunami, Japan).
- 1 MAG was obtained from samples from Corona Mine drainage at the Oat Hill Mine, Napa County, California, USA; in this case, water was filtered through 2.5 and 0.1 ⁇ m filters sequentially, DNA was extracted from each filter separately, and two runs of Illumina paired-end sequencing were performed at 150 bp and 250 bp read lengths.1 MAG were also obtained from hyporheic zone water sampled from beneath the riverbed of the East River, Gunnison County, Colorado, USA; in this case, water was filtered through a 0.1 ⁇ m filter, DNA was extracted from the filter, and Illumina sequencing was performed with a read length of 150 bp. For the East River and Oat Hill Mine samples, genome data were assembled using Metaspades v3.15.5 with default k-mer values.
- the analysis pipeline was written using the Nextflow pipeline framework, which allowed for the analysis to be run reproducibly in containers, which were executed in parallel across nodes of the computing cluster.
- the first process in the pipeline performed a multiple sequence alignment of training sequences, which was then used to build a hidden Markov model (HMM) using hmmer3 that modelled the primary sequence of the [FeFe]-hydrogenase catalytic domain.
- HMM hidden Markov model
- Candidate protein sequences with length greater than 100,000 amino 1005373994 acids were first excluded from the analysis. Each protein sequence from each representative genome was then matched against the HMM to obtain a bit score, which represented the similarity of the candidate sequence to the known enzyme profile.
- [FeFe]-hydrogenases are often multidomain proteins comprising a conserved H 2 - activating domain named the H-cluster and various accessory domains at the N- and C- terminals involved in electron transfer or H2 sensing.
- H-cluster a conserved H 2 - activating domain
- accessory domains at the N- and C- terminals involved in electron transfer or H2 sensing.
- To locate the positions of the H- cluster all retrieved archaeal [FeFe]-hydrogenase sequences were aligned against the trimmed reference [FeFe]-hydrogenase H-cluster sequences from HydDB using Clustal Omega v1.2.2 (default setting).
- N- and C-terminal regions outside of the H-cluster were extracted and annotated against CDD v3.19 using rpsblast (- evalue 0.01 -max_hsps 1 - max_target_seqs 10) in BLAST+ v2.9.0.
- the N-terminal fusion of the small subunit of [NiFe]-hydrogenase with certain archaeal [FeFe]-hydrogenases was confirmed by searching against Pfam protein family database v34.0 and protein structural modelling as detailed below.
- archaeal [FeFe]-hydrogenases were assigned into different subclasses. Analysis of [FeFe]-hydrogenase genetic organization [0213] To characterise the genetic context and potential interacting proteins of archaeal [FeFe]-hydrogenases, up to 10 genes upstream and downstream of the catalytic subunits were retrieved.
- flanking gene that encodes large subunit of [NiFe]-hydrogenase was classified using HydDB.
- all flanking genes were clustered at an identity threshold of 30% and a minimum coverage of 80% using MMseqs2 (--min-seq-id 0.3 -c 0.8 --cov-mode 1 -- cluster-mode 2 -s 7.5).
- the R package gggenes v0.4.1 https://github.com/wilkox/gggenes was used to construct gene arrangement diagrams.
- New and authentic divergent hits including from archaea were included in the final database, which was then used to screen for the presence of [FeFe]-hydrogenase maturases in archaeal genomes.
- a relaxed default setting of the DIAMOND v0.9.31 BLASTp algorithm was first applied. False positive hits were filtered by further searching against CDD v3.19 for the presence of rSAM_HydE (TIGR03956), rSAM_HydG (TIGR03955), and GTP_HydF (TIGR03918) domains.
- [FeFe]-hydrogenase models were generated using AlphaFold multimer v2.1.1 implemented on the Monash University MASSIVE M3 computing cluster.
- the amino acid sequences for HydA from each of the [FeFe]-hydrogenase groups (A1, A3, B, E, and F) shown in Table 7 were modelled both alone and with the sequences of putative complex partners present in the hydrogenase gene cluster. Modelling with putative complex partners was performed iteratively, with output models assessed to determine if a credible 1005373994 complex was generated. Predicted complexes were then modelled with higher stoichiometries, where possible given limitations of GPU RAM ( ⁇ 2,500 amino acids), to predict larger order structures.
- Models produced were validated based on confidence scores (pLDDT) with only regions with a confidence score of >85 utilised for analysis. Where complexes were predicted subunit interfaces were inspected manually for surface complementarity and the absence of clashing atoms. Interfaces were also analysed for stability using the program QT-PISA, with only interfaces predicted to be stable utilised for analysis.
- pLDDT confidence scores
- To assign cofactors to [FeFe]-hydrogenase models generated by AlphaFold the closest homologous structures or domains were identified by searching the PDB database using NCBI BLAST or the DALI server. The homologous structures were aligned with the AlphaFold models and cofactors were added in corresponding positions to that of the experimental structures, providing all conserved coordinating residues were present.
- the genes encoding group A, B, E, and F [FeFe]-hydrogenases were cloned in pET-11a(+) by Genscript using restriction sites NdeI and BamHI following codon optimization for expression in E. coli. Expression constructs were re-transformed in chemically competent E. coli BL21(DE3) cells to express the apo-forms of the hydrogenases lacking the diiron subsite of the H- cluster. Starter cultures were grown overnight in 5 mL LB medium containing 100 ⁇ g mL- 1 ampicillin at 37°C.
- Cells were thereafter harvested by centrifugation at 4,930 ⁇ g for 10 mins at 4°C. All subsequent operations were carried out under anaerobic conditions to prevent hydrogenase inactivation by atmospheric oxygen in an MBRAUN glovebox ([O2] ⁇ 5 ppm).
- the cell pellet was resuspended in a 0.5 mL lysis buffer (30 mM Tris-HCl pH 8.0, 0.2 % (v/v) Triton X-100, 0.6 mg mL -1 lysozyme, 0.1 mg mL -1 DNase, 0.1 mg mL -1 RNase).
- H2 production was determined by analysing the reaction headspace every 15 mins using a PerkinElmer Clarus 500 gas chromatograph (GC) equipped with a thermal conductivity detector (TCD) and a stainless-steel column packed with Molecular Sieve (60/80 mesh).
- GC PerkinElmer Clarus 500 gas chromatograph
- TCD thermal conductivity detector
- Molecular Sieve 60/80 mesh
- the operational temperatures of the injection port, oven, and detector were 100°C, 80°C, and 100°C, respectively.
- Argon was used as carrier gas at a flow rate of 35 mL min ⁇ 1 .
- the three biological replicates were run at varying times (1-4 hours) of incubating the cell lysates with the [2Fe] adt subsite mimic. Incubation time was not found to influence the observed H2 production.
- Induced cultures were incubated at 20°C and 150 rpm for approximately 16 h. Cells were thereafter harvested by centrifugation in a Beckman Coulter Avanti J-25 centrifuge (5,000 rpm/4,424 x g, 10 min). All subsequent operations were carried out under anaerobic conditions in the glovebox to prevent hydrogenase inactivation by atmospheric oxygen.
- the cell pellet was resuspended in 100 mM Tris-HCl pH 8.0, with NaCl (150 mM), MgCl 2 (10 mM), lysozyme from chicken egg white (1 mg/mL), DNAse I from bovine pancreas (0.05 mg/mL), RNase A from bovine pancreas (0.05 mg/mL), and a tablet of cOmpleteTM EDTA-free protease inhibitor cocktail, and was incubated inside the glovebox for 30 min. Cell lysis was performed by three cycles of freezing/thawing in liquid N 2 .
- the StrepTrapTM eluates were further separated using size exclusion chromatography via SuperdexTM 20010/300 GL, equilibrated in 100 mM Tris-HCl, 150 mM NaCl pH 8.0, to acquire purified Fm (Figure 11).
- Samples were stored anaerobically at ⁇ 80 °C.
- Coomassie-stained SDS-PAGE was used to assess purity. Protein estimations were performed via Bradford assay using bovine serum albumin as a standard. Quantification of Fe-content was performed using a previously reported assay (Fish 1988) using a commercially available Fe 2+ standard for AAS TraceCERT (Sigma Aldrich) for the calibration curve.
- the reconstituted enzyme 50 ⁇ M was mixed with sodium dithionite (1 mM, 20 ⁇ excess) in 100 mM phosphate buffer, pH 6.8 and incubated in room temperature for 10 minutes.
- Cofactor incorporation started with the addition of [2Fe] adt (600 ⁇ M, 12 ⁇ excess), and the reaction mixture was incubated for 1 h.
- the mixture was loaded onto a PD-10 desalting column (GE Healthcare) equilibrated with 10 mM Tris-HCl pH 8.0.
- Protein film electrochemistry experiments were carried out under anaerobic conditions at 20°C and pH 7.0.
- the three-electrode system was made up of (1) Ag/AgCl (4 M KCl) as reference electrode, (2) rotating disk 5 mm OD pyrolytic graphite edge (PGE) plane (epoxy encapsulated) as working electrode, and (3) graphite rod as the counter electrode.
- the gas-tight glass cell used featured a water jacket for temperature control and a cell gas inlet/outlet for hydrogen flow control.
- the buffer used was composed of 5 mM MES, 5 mM CHES, 5 mM HEPES, 5 mM TAPS, 5 mM sodium acetate (NaOAc), with 0.1 M Na2SO4 as carrying electrolyte titrated with H2SO4 to pH 7.0, and purged with N2 for 3 to 4 hours.
- the PGE working electrodes were polished with P1200 sandpaper and rinsed with purified water before they were brought into the glovebox.
- cyclic voltammograms were run at 100 mV/s from -100 to -600 mV (vs. standard hydrogen electrode (SHE)) for 40 scans.
- the cyclic voltammogram of the blank electrode (no enzyme immobilized) was then recorded at 10 mV/s with the working electrode rotated at 3 krpm.
- Polycationic polymyxin B sulfate (5 ⁇ L of 0.2 mg/mL) was added onto the deaerated PGE surface before adding 5 uL of 5 ⁇ M activated enzyme. The mixture was left for 10 min for maximal adsorption before the excess solution was removed by pipet.
- the ATR unit (BioRadII from Harrick) was sealed with a custom build PEEK cell that allowed for gas exchange and illumination mounted in a FTIR spectrometer (Vertex V70v, Bruker). The sample was dried under 100% nitrogen gas and rehydrated with a humidified aerosol (100 mM Tris-HCl, pH 8). Spectra were recorded with 2 cm -1 resolution, a scanner velocity of 80 Hz and averaged of varying number of scans (mostly 1000 Scans). All measurements were performed at ambient conditions (room temperature and pressure, hydrated enzyme films). Photochemical reduction was achieved through a previously established protocol (Lorenzi et al. 2022; Senger et al. 2019).
- the functional annotations were expanded to include marker genes for fermentation, fatty acid 1005373994 degradation, aromatic compound degradation, carbohydrate metabolism, and sulfur metabolism using curated HMMs from METABOLIC and KofamKOALA v1.3.0. Hits were further inspected through searching the NCBI CDD database.
- the final curated metabolic marker genes were visualized by a custom Python script to construct a heatmap that displayed the presence or absence of particular gene markers within the archaeal genomes. Counts of genes by genome were transformed into binary format, where "1" denotes the presence of at least one hit for a specific gene marker, while "0" signifies the absence of hits for the given marker.
- Genome-level phylogenetic analysis [0226] To construct the archaeal genome tree, 15 conserved syntenic ribosomal proteins were retrieved, aligned, and concatenated using GOOSOS (https://github.com/jwestrob/GOOSOS) using all archaeal genomes at least 60% complete and less than 5% contaminated (based on CheckM). Genomes containing at least 75% of the 15 syntenic proteins were retained. The concatenated ribosomal protein sequences were aligned using MAFFT, followed by trimming with trimAl using the -gt 0.1 option. The final length of the trimmed concatenated protein alignment was 3224 amino acids for 118 genomes.
- taxonomy was assigned to each sequence using ETE v3.0.0 and CD-HIT v4.6 was used to reduce the dataset at the 80% amino acid sequence identity level.
- the multiple sequences retrieved were aligned using MAFFT v7.304 (settings: -- localpair --maxiterate 1000 --reorder).
- the resulting alignment was trimmed using trimAl (settings: -gt 0.1) and manually inspected with Geneious to remove partial sequences.
- Maximum likelihood phylogenetic trees were constructed using IQ-TREE v1.6.1 to test various models and topologies, obtaining bootstrap values.
- the LG+F+G4 substitution model with 1,000 ultrafast bootstraps was selected for tree generation.
- Supplementary phylogenetic trees were constructed using two different models, LG+FO+R and mixed 1005373994 model LG+C60.
- the three models were applied to 3,677 amino acid sequences of the catalytic subunit (HydA) of [FeFe]-hydrogenases, including novel hybrid hydrogenases.
- NADH-quinone oxidoreductase subunit G (NuoG) and formate dehydrogenase subunit A (FdhA1) were used as outgroup sequences for rooting given they are reported to be related and ancestral to the [FeFe]-hydrogenase catalytic subunit (HydA).
- Example 2 Structurally and genetically diverse [FeFe]-hydrogenases are encoded by nine archaeal phyla [0228] The 2,339 archaeal species clusters of the Genome Taxonomy Database (GTDB) and the repository of novel archaeal metagenome-assembled genomes was searched for the gene encoding the catalytic subunit of [FeFe]-hydrogenases (HydA).
- GTDB Genome Taxonomy Database
- the archaeal group A1 and group E [FeFe]- hydrogenases are putative fermentative enzymes encoded by three DPANN phyla (Iainarchaeota, Micrarchaeota, Nanoarchaeota). With average sequences of just 363 and 286 residues respectively (after excluding any truncated sequences), these enzymes are much smaller than the most minimal hydrogenase previously characterised (for example Chlamydomonas reinhardtii HydA1; 457 residues). The most widespread hydrogenase, however, is the electron-bifurcating group A3 [FeFe]-hydrogenase.
- This hydrogenase together with its partner diaphorase (HydB) and thioredoxin (HydC) subunits, is encoded by at least six DPANN phyla, Thermoplasmatota (class E2), and some Asgard archaea (class Lokiarchaeia) ( Figure 1).
- the only cultured archaeon that encodes an [FeFe]-hydrogenase is ‘Candidatus Prometheoarchaeon syntrophicum’ ( Figure 1).
- Example 3 Diversity, distribution, conserved features, and classification of archaeal [FeFe]-hydrogenases
- a total of 136 [FeFe]-hydrogenases were identified from 130 genomes from nine archaeal phyla, including 3.3% of the 2,339 representative archaeal species genomes of the Genome Taxonomy Database (GTDB) R05-RS202.
- GTDB Genome Taxonomy Database
- the group A1 [FeFe]-hydrogenases are encoded by 26 genomes of two DPANN phyla, Iainarchaeota and Micrarchaeota.
- the hydrogenase is monomeric with the catalytic H- cluster as the signature domain and does not have additional iron-sulfur cluster binding domains at the N-terminal region (subclass M1) ( Figure 6).
- the intact archaeal group A1 [FeFe]-hydrogenase (324 - 386 residues) is much smaller than bacterial and eukaryotic group A1 enzymes from HydDB (average 564 residues).
- No known [FeFe]-hydrogenases structural subunits are present in 10 genes upstream and downstream of the hydrogenase.
- the group A3 [FeFe]-hydrogenases have the broadest distribution among archaea, encoded by 44 genomes from eight phyla (Aenigmatarchaeota / QMZS01, Altarchaeota, Asgardarchaeota, EX4484 ⁇ 52, Iainarchaeota, Micrarchaeota, Nanoarchaeota, Thermoplasmatota).
- subclass M2 a single 2[4Fe4S] cluster binding domain at N-terminus
- subclass M3 three signature iron-sulfur cluster binding domains at N-terminal from [2Fe2S], (Cys)3His-ligated [4Fe4S] to 2[4Fe4S]
- new subclass M3c two signature iron-sulfur cluster binding domains at N-terminus from [2Fe2S] to 2[4Fe4S]
- Figure 6 The average gene length of archaeal group A3 [FeFe]-hydrogenases is 560 residues, comparable with bacterial group A3 counterparts in HydDB (average 585 residues).
- the group E [FeFe]-hydrogenase is a novel monophyletic clade likely basal to the bacterial group C [FeFe]-hydrogenases. It consists of 29 sequences from two DPANN phyla, Iainarchaeota and Nanoarchaeota. The putative PAS sensory domain commonly present in bacterial group C [FeFe]-hydrogenases is not found in this archaeal clade. With complete sequence lengths between 267 - 316 residues, members of this group are the most minimalistic hydrogenases known. As a comparison, bacterial group C [FeFe]- hydrogenases in HydDB have an average length of 567 residues.
- group E [FeFe]-hydrogenases are monomeric with only the minimal catalytic H-cluster (new subclass M1a) ( Figure 6) and no other [FeFe]-hydrogenases structural subunits are identified in the flanking regions.
- a ribonuclease (elaC) and a small nuclear ribonucleoprotein (lsm) are sometimes found next to hydA, but their modest association is likely due to close relationships of some genomes analysed.
- the newly discovered archaeal group F and G [FeFe]-hydrogenases together form an early diverging branch sister to the group A [FeFe]-hydrogenases.
- Sequences from the two groups represent two deep-branching clades separating from each other early in evolution.
- the group F [FeFe]-hydrogenases (average 493 residues) were identified in 21 genomes from three phyla including Asgardarchaeota (1), Thermoplasmatota (6), and Thermoproteota (17).
- genes encoding structural subunits of [FeFe]-hydrogenases and [NiFe]-hydrogenases were present in the immediate vicinity of the fusion gene (Figure 7). They include the diaphorase (hydB) and thioredoxin (hydC) subunits of the electron- bifurcating group A3 [FeFe]-hydrogenase, and the group 3 [NiFe]-hydrogenase large subunit (hyhL). A neighbouring gene homologous to nuoG subunit of the NADH dehydrogenase was also found and is denoted hydD. As elaborated below, AlphaFold modelling suggests the five proteins associate into a stable complex.
- the short stretch of genetic region between the hyhS domain and H-cluster of the group F [FeFe]-hydrogenase also has a high degree of homology with nuoG.
- a possible scenario for the evolution of the group F [FeFe]- hydrogenases is homologous recombination between the nuoG-like region of group A3 and hydA of group G enzymes.
- nuoG (hydD) was split from hydA as a separate gene whereas hyhS was fused with hydA in group F [FeFe]-hydrogenases.
- Iainarchaeum andersonii using AlphaFold2. These enzymes have no obvious genetically associated interaction partners (Figure 7), and the structural modelling indicates they form a compact monomeric H- domain with a solvent exposed H-cluster Figure 2a; Figure 8). These group A1 enzymes do not feature additional [FeS] clusters, suggesting direct electron transfer to or from a soluble electron carrier. All five cysteines in the active site pocket previously identified as critical for H-cluster assembly and efficient catalysis in group A [FeFe]-hydrogenases are conserved in both Group A1 enzymes. This includes the four H-cluster coordinating cysteines (C 2 -C 5 ) and the proton transfer residue (C 1 ).
- the C 1 cysteine is strictly conserved in previously identified group A and B [FeFe]- hydrogenases, and has been identified as important for catalytic activity as it plays a key role in proton transfer to the active site. Sequence alignment further revealed that other active site residues that are generally well-conserved in bacterial and eukaryotic group A [FeFe]-hydrogenases were present also in Mu and Ia (for an extended comparison see Figure 9). In line with this, the structural modelling supports the notion that they form an H-cluster with canonical coordination ( Figure 2a, Figure 8). Models from the ultraminimal group E [FeFe]- hydrogenases from Ca.
- the HydA subunit consists of a H-domain fused to a 2x[4Fe-4S] ferredoxin-like domain that acts as an electron relay to/from the H-cluster.
- An additional 2x[4Fe-4S] ferredoxin-like domain is inserted in this domain (representing amino acids 122-182), forming a domain independent of the body of the protein.
- the predicted [FeS] clusters of this domain are 1005373994 outside of electron transfer distance (> 15 ⁇ ) with those associated with the H-cluster.
- the HydC subunit interacts with HydA so that its single predicted [2Fe- 2S] cluster is also not within electron transfer distance of the rest of the proteins in the structure ( Figure 2d).
- the lack of proximity of these FeS clusters suggests that additional unidentified subunits interact with this enzyme to complete the electron transfer relay. Alternatively, this enzyme may undergo extensive conformational changes during catalysis.
- the model of a group A3 [FeFe]-hydrogenase from DSAL01 sp011380095 (Altarchaeota) is composed of a complex of HydA, HydB and HydC subunits, which form a putative electron- bifurcating complex similar to that recently reported from Thermotoga maritima (Figure 2d).
- Example 5 Archaeal [FeFe]-hydrogenases are catalytically active and display H- cluster spectroscopic signals [0242] Group A1, B, and E [FeFe]-hydrogenases from archaea were tested to determine if they could bind the catalytic H-cluster and produce H2. To do so, five enzymes were heterologously expressed in Escherichia coli BL21(DE3), anaerobically matured using the synthetic mimic [2Fe] adt ([Fe2(azadithiolate)(CO) 4 (CN) 2 ] 2 ⁇ ), and H 2 production of whole-cell lysates was measured using gas chromatography (Figure 8).
- Example 6 ATR-FTIR analysis of archaeal [FeFe]-hydrogenases
- Attenuated Total Reflection Fourier transformed infrared (ATR-FTIR) spectroscopy confirmed successful assembly of the H-cluster of Ca.
- Sharp cofactor bands in the expected CO/CN ligand band region of the FTIR spectra were readily observed ( Figure 3b).
- the two CN (2105, 2097 cm -1 ) and four CO (2018, 2001, 1992, 1860 cm -1 ) cofactor ligand bands imply in total six ligands, indicative of a CO inhibited state; the formation of H ox -CO has been shown to occur during the artificial activation process.
- the observed peak positions are pronounced of the hydride state Hhyd or the inhibited Hinact and Htrans states, previously observed for bacterial [FeFe]-hydrogenases, suggesting a di-ferrous oxidation state of the [2Fe] H subsite.
- Example 7 - [FeFe]- and [NiFe]-hydrogenases associate into complexes in uncultivated archaea [0247]
- the group F [FeFe]-hydrogenases appear to form complexes with [NiFe]-hydrogenases.21 genomes encoded these complexes through five-gene clusters, including from the classes Bathyarchaeia, Brockarchaeia, Thermoplasmata, Thermoproteia, and Lokiarchaeia ( Figure 1).
- HydA C-terminal [FeFe]-hydrogenase catalytic domain
- HyhS N-terminal domain homologous to the group 3 [NiFe]-hydrogenase small subunit (HyhS) ( Figure 4a & Figure 8).
- the structural model suggests the complex receives electrons through the [FeFe]- hydrogenase arm or the [NiFe]-hydrogenase via a series of iron-sulfur clusters to a probable electron-converging [4Fe4S] cluster on the hybrid subunit. Thereafter electrons are predicted to be simultaneously transferred to the high-potential NAD + at the HydC subunit and an undetermined low-potential acceptor (likely ferredoxin) at the glutamate synthase (GltA) domain of the HydB subunit ( Figure 4b).
- These observations are remarkable given [NiFe]- and [FeFe]- hydrogenases are not known to associate.
- the hybrid HydA subunit When modelled alone using AlphaFold, the hybrid HydA subunit was predicted to contain a [FeFe]-hydrogenase catalytic domain and a [NiFe]- hydrogenase iron-sulfur cluster domain separated by a long flexible linker ( Figure 8).
- the cysteine residues required to ligate the [FeFe]-hydrogenase H-cluster of the HydA subunit (Cys357, Cys406, Cys536, Cys540) were present, and are overall well-conserved for all group F representatives, indicating that the HydA domain can bind an H-cluster.
- the proton transfer C 1 cysteine is absent.
- the active-site pocket also featured changes in other amino acids which are well-conserved in all other groups of [FeFe]-hydrogenases.
- the so-called APA (or APS) motif is replaced by a DPI motif in most identified group F enzymes ( Figure 9).
- the cysteine residues required to ligate the catalytic NiFe cofactor are present in HyhL (Cys63, Cys66, Cys418, Cys 421). Altogether, this indicates that both the HydA and HyhL subunits are likely to be active hydrogenases (Figure 4c).
- HydB from the archaeal group F hydrogenase this domain is substituted by a homologue of the GltA subunit from glutamate synthase, which contains 2x[4Fe4S] clusters (B3 and B4) and FAD, which likely serves as an electron donor to an external substrate. Electrons are likely transferred from C1 to B1, and then to the terminal FAD via the B3 and B4 clusters ( Figure 4c).
- A. mobile enzyme it is predicted that electron bifurcation is mediated by flexibility in the complex that changes the proximity of the [FeS]-clusters gating the flow of electrons through the complex.
- the GltA-like domain of the archaeal group F hydrogenase is attached to the remainder of HydB via a flexible linker, suggesting that electron bifurcation may be achieved by a similar mechanism in this complex.
- the species that accepts electrons from FAD in the GltA-like domain is unclear, although it may be soluble ferredoxin or another low potential electron acceptor.
- the archaeal group F hydrogenase functions in a hydrogen-evolving direction, then this process would occur in reverse with electrons entering the complex via the FMN and FAD cofactors and converging at cluster A2, before transfer to the [FeFe] and/or [NiFe] sites of HydA and HyhL respectively.
- Example 9 - [FeFe]-hydrogenases enable fermentation and electron-bifurcation in diverse archaea
- the role of the various [FeFe]-hydrogenases in the metabolism of archaea was studied. The high-quality archaeal genomes for genes associated with major energy conservation and carbon acquisition processes was annotated. All [FeFe]-hydrogenase- encoding archaea are predicted to be obligate anaerobes given they lacked terminal oxidases. This is consistent with the retrieval of the genomes from typically anoxic ecosystems, especially groundwater, anaerobic digesters, and sediments from hot springs, hydrothermal vents, and freshwater (Figure 1).
- DPANN archaea are likely to be symbiotic obligate fermenters dependent on host-derived organic compounds; consistently, they often encoded genes for the degradation and fermentation of carbohydrates (primarily starch) and aromatic compounds, but generally lacked respiratory reductases or carbon fixation pathways (Figure 1).
- the other archaeal phyla were predicted to be capable of a wider range of metabolic strategies, in line with their larger genome sizes (Figure 6), including beta oxidation, anaerobic respiration, and carbon fixation to varying extents ( Figure 1).
- 2-oxoacid-ferredoxin oxidoreductase genes are frequently adjacent to [FeFe]-hydrogenase genes in Nanoarchaeota and Micrarchaeota MAGs ( Figure 7).
- the trimeric group A3 [FeFe]-hydrogenases are predicted to simultaneously reoxidize ferredoxin (reduced primarily by pyruvate-ferredoxin oxidoreductase) and NADH (e.g. reduced during glycolysis) ( Figure 2), in line with their bacterial counterparts.
- the electron-bifurcating group A3 [FeFe]- hydrogenases are associated with those DPANN archaea harbouring relatively complex carbohydrate degradation pathways (i.e. Woesearchaeles, Altarchaeota, some Iainarchaeota) through which both NAD + and ferredoxin will be reduced (Figure 1).
- Example 10 - [FeFe]-hydrogenases have been acquired by archaea on multiple occasions and have an ancient association with [NiFe]-hydrogenases [0257]
- the evolutionary history of [FeFe]-hydrogenases was investigated through phylogenetic analysis of its catalytic subunit ( Figure 5; Figures 13 to 14) and three maturases ( Figure 15 to 17).
- the archaeal group A1, A3, and B [FeFe]-hydrogenases clustered with various bacterial and eukaryotic homologs; the group A enzymes form at least six radiations, suggesting they were laterally acquired from bacterial enzymes over several independent events, whereas the archaeal group B and E enzymes each formed a monophyletic clade ( Figure 5; Figures 13 to 14).
- Figure 5 Figures 13 to 14
- the ultraminimal fermentative group E enzymes of archaea are ancestral to the multidomain sensory group C hydrogenases of bacteria.
- [FeFe]- hydrogenases are absent from the Asgard genomes most related to eukaryotes ( Figure 1) and the alphaproteobacterial genomes most related to mitochondria, though this view could change with the addition of new genomes. Nevertheless, archaeal and eukaryotic [FeFe]-hydrogenases clustered together in three of these lineages (including Lokiarchaeia with Tritrichomonas), suggesting eukaryotes may have laterally acquired [FeFe]- hydrogenases from archaea during their diversification (Figure 5).
- the iron-sulfur (small) domain fused to the group F [FeFe]- hydrogenase and the iron-sulfur subunit downstream of the group G [FeFe]-hydrogenase each form distinct monophyletic subgroups within the group 3 [NiFe]-hydrogenases ( Figure 19).
- the iron-sulfur domain of hybrid complexes more strongly co-evolved with the [FeFe]- than [NiFe]-hydrogenase catalytic subunits.
- [FeFe]-hydrogenases may have first evolved in archaea in association with [NiFe]-hydrogenases; they potentially diversified into various monomeric and multimeric lineages, which were variably acquired by diverse bacteria, archaea, and eventually eukaryotes. It is proposed that the components of the hybrid hydrogenases are formally recognised as distinct lineages, namely the group F and G [FeFe]-hydrogenases and group 3f and 3g [NiFe]- hydrogenases, given their distinct phylogenies, structures, and potential physiological roles. [0260] Archaea have evolved remarkably disparate ways to use [FeFe]-hydrogenases to adapt to anaerobic environments.
- DPANN archaea have evolved ultraminimal enzymes to efficiently dispose of reductant derived from carbohydrate fermentation.
- lineages such as Brockarchaeia and Lokiarchaeia use unique hybrids of [NiFe]- and [FeFe]-hydrogenases – among the most complex hydrogenases described to date – to support their diverse redox biology.
- this 1005373994 approach also emphasises the potential of combining genome-resolved metagenomics with accurate protein structure prediction and heterologous production studies to discover new enzymes and functions in uncultured microorganisms.
- References [0262] Castelle, C.J., Brown, C.T., Anantharaman, K., Probst, A.J., Huang, R.H., and Banfield, J.F. (2016). Biosynthetic capacity, metabolic variety and unusual biology in the CPR and DPANN radiations. Nat. Rev. Microbiol.16, 629–645.
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Organic Chemistry (AREA)
- Genetics & Genomics (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Biochemistry (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Microbiology (AREA)
- Biomedical Technology (AREA)
- Biotechnology (AREA)
- Molecular Biology (AREA)
- Chemical Kinetics & Catalysis (AREA)
- General Chemical & Material Sciences (AREA)
- Sustainable Development (AREA)
- Manufacturing & Machinery (AREA)
- Sustainable Energy (AREA)
- Medicinal Chemistry (AREA)
- Electrochemistry (AREA)
- Physics & Mathematics (AREA)
- Biophysics (AREA)
- Plant Pathology (AREA)
- Enzymes And Modification Thereof (AREA)
Abstract
The invention relates to enzymes and polypeptide complexes for generating energy from hydrogen or for generating hydrogen, nucleic acid molecules encoding the same, devices and systems comprising the same, and uses thereof. An isolated, synthetic or purified nucleic acid molecule encoding a hydrogenase from archaea of a lineage Thermoplasmatota, Asgardarchaeota, Thermoproteota, EX4484-52, Aenigmarchaeota / QMZS01, Nanoarchaeota, Altarchaeota, lainarchaeota, or Micrarchaeota.
Description
Novel hydrogenases Field of the invention [0001] The invention relates to enzymes and polypeptide complexes for generating energy from hydrogen or for generating hydrogen, nucleic acid molecules encoding the same, devices and systems comprising the same, and uses thereof. Related application [0002] This application claims priority from Australian provisional application no. 2023902336, the entire contents of which are incorporated herein by reference. Sequence listing [0003] A sequence listing in ST.26 format is filed herewith, the entire contents of which are incorporated herein by reference. Background of the invention [0004] Molecular hydrogen (H2) is heralded as a future green energy carrier. In a biological context, this energy-rich gas already plays a central role in bioenergetics and evolution has driven an elaborate hydrogen economy. Numerous bacteria, and microbial eukaryotes consume and produce hydrogen gas (H2 ⇌ 2 H+ + 2 e-) using metalloenzymes called hydrogenases. Three hydrogenases have independently evolved in microorganisms, namely the [FeFe]-, [NiFe]-, and [Fe]-hydrogenases, which differ in their metal cofactors and catalytic mechanism. H2 serves multiple roles in microbial physiology. Microorganisms produce H2 to dispose of electrons during fermentation. [0005] Numerous bacteria and archaea also use electrons derived from H2 oxidation for respiration and carbon fixation. More recently, electron-bifurcating hydrogenase complexes have been discovered that are critical for energy conservation in obligate anaerobes. [0006] It is increasingly recognized that microbial H2 metabolism shapes global biogeochemical cycling, supports global biodiversity, and influences health and disease. In addition, these efficient enzymes have growing industrial applications in the developing H2 economy and serve as an inspiration for the design of synthetic catalysts. H2 was likely 1005373994
the primordial electron donor, but continues to have a central role in microbiology both as a desirable energy source and diffusible electron sink. Moreover, it is proposed that H2 exchange between bacteria and archaea underlies eukaryogenesis, as described in various syntrophy hypotheses. [0007] The three hydrogenase classes differ in their physiological roles and taxonomic distribution. [FeFe]-hydrogenases are typically fast-acting, but oxygen-sensitive, and are best known for their roles in obligate anaerobes. These enzymes currently comprise four phylogenetically distinct groups (groups A to D), which can be further subdivided through two different schemes based on domain architecture and genetic organisation. They include monomeric enzymes that couple ferredoxin oxidation to fermentative H2 production (group A1), trimeric enzymes that reversibly bifurcate electrons from H2 to NAD+ and ferredoxin (group A3), putative sensory hydrogenases in which the catalytic hydrogenase domain is fused with a PAS domain (group C), and several functionally undefined groups (e.g. groups B and D). Despite this diversity, they are all expected to rely on the same organometallic cofactor for catalysis, the “H-cluster”. To date, these enzymes have been exclusively characterised in anaerobic bacteria and eukaryotes, and appear to be absent in cultured archaea. [NiFe]-hydrogenases are extraordinarily structurally and functionally diverse enzymes encoded by bacteria and archaea across all ecosystems. They are presently subdivided into four major groups (groups 1 to 4) and 29 subgroups that each differ in their phylogeny, genetic organisation, and physiological roles. The catalytic (large) subunit and electron-relaying iron-sulfur (small) subunit of the [NiFe]-hydrogenase associate with other subunits depending on the subgroup; the different complexes formed can mediate respiration, fermentation, energy-conversion, electron-bifurcation, carbon fixation, and H2 sensing processes. In contrast, [Fe]- hydrogenases are a much narrower lineage that contribute to archaeal methanogenesis. The three hydrogenase classes are phylogenetically unrelated, despite having some similar structural features, and are not thought to genetically or structurally associate. [NiFe]-hydrogenases were present in the last universal common ancestor (LUCA), whereas [FeFe]-hydrogenases are proposed to have evolved later in fermentative bacteria. [0008] Several genomic studies have also suggested that [FeFe]-hydrogenases may be encoded by uncultivated DPANN archaea. However, this remains debatable given archaea seemingly lack the three maturation enzymes (HydEFG) required to synthesize 1005373994
the biologically unique catalytic H-cluster of the [FeFe]-hydrogenase, and to date there is no evidence for archaeal [FeFe]-hydrogenase activity. [0009] There is a need to identify alternative catalysts for the oxidation of hydrogen and/or the production of hydrogen which have a potential use in industrial applications, such as for energy production. [0010] Reference to any prior art in the specification is not an acknowledgment or suggestion that this prior art forms part of the common general knowledge in any jurisdiction or that this prior art could reasonably be expected to be understood, regarded as relevant, and/or combined with other pieces of prior art by a skilled person in the art. Summary of the invention [0011] The present invention is based on the identification by the inventors of previously unknown hydrogenases from archaea and the use thereof to oxidise or reduce hydrogen. [0012] Accordingly, in a first aspect, the present invention provides an isolated, synthetic or purified nucleic acid molecule encoding a hydrogenase from archaea of a lineage Thermoplasmatota, Asgardarchaeota, Thermoproteota, EX4484-52, Aenigmarchaeota / QMZS01, Nanoarchaeota, Altarchaeota, lainarchaeota, or Micrarchaeota. Preferably, the archaea is one shown in Figure 1. [0013] Typically, the nucleic acid molecule encodes a hydrogenase from group A1, A3, B, C1, E, F or G. [0014] Preferably, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 1 to 150, or a sequence at least about 80% identical thereto. [0015] Preferably, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 300, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 151 to 300. [0016] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 1 to 1005373994
150 (e.g. in Table 1), or a sequence at least about 80% identical thereto, and one or more HydB, HydC, HydD, HyhL or HyhS sequences as set forth in any one of SEQ ID Nos: 301 to 463, 627 to 629 (e.g. in Table 3), or a sequence at least about 80% identical thereto. Preferably the HydA, HydB, HydC, HydD, HyhL and/or HyhS combinations are any one shown in the row 1 to 73 of Table 5. [0017] Preferably, the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 300 (e.g. in Table 2), or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 151 to 300, and and one or more HydB, HydC, HydD, HyhL or HyhS sequences as set forth in any one of SEQ ID NOs: 464 to 626, 630 to 632 (e.g. in Table 4) or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 464 to 626. Preferably the HydA, HydB, HydC, HydD, HyhL and/or HyhS combinations are any one shown in rows 1 to 73 of Table 5. [0018] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 1 to 26, or a sequence at least about 80% identical thereto. [0019] In another embodiment, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 176, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 151 to 176. [0020] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 27 to 70, or a sequence at least about 80% identical thereto. [0021] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 27 to 70, or a sequence at least about 80% identical thereto, and one or both of a HydB and HydC sequence as set forth in any one of SEQ ID NOs: 301 to 334, 336 to 341, 343 to 386, 389 to 390, 393 to 396, 399 to 400, 404 to 407, 410 to 411, 415 to 416, 422 to 425, 1005373994
428 to 429, 432 to 433, 438 to 439, 443 to 444, 447 to 450, 452 to 453, and 457 to 458, or a sequence at least about 80% identical thereto. Preferably, the HydA, HydB and HydC combinations are any one shown in rows 2 to 39 in Table 5. In one embodiment, the combination shown in row 20 of Table 5 includes the HyhL. In one embodiment, the combination shown in row 22 of Table 5 includes the HyhL. In one embodiment, the combination shown in row 23 includes either HydB (SEQ ID NO: 343 or 344). [0022] In another embodiment, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 177 to 220, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 177 to 220. [0023] In another embodiment, the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 177 to 220, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 177 to 220, and and one or both of a HydB and HyC sequence as set forth in any one of SEQ ID NOs: 464 to 497, 499 to 504, 506 to 549, 552 to 553, 556 to 559, 562 to 563, 567 to 570, 573 to 574, 578 to 579, 585 to 588, 591 to 592, 595 to 596, 601 to 602, 606 to 607, 610 to 613, 615 to 616, and 620 to 621, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence a HydB and HyC sequence as set forth in any one of SEQ ID NOs: 464 to 497, 499 to 504, 506 to 549, 552 to 553, 556 to 559, 562 to 563, 567 to 570, 573 to 574, 578 to 579, 585 to 588, 591 to 592, 595 to 596, 601 to 602, 606 to 607, 610 to 613, 615 to 616, and 620 to 621. Preferably, the HydA, HydB and HydC combinations are any one shown in rows 2 to 39 in Table 5. In one embodiment, the combination shown in row 20 of Table 5 includes HyhL. In one embodiment, the combination shown in row 22 of Table 5 includes HyhL. In one embodiment, the combination shown in row 23 includes HydB as set forth in SEQ ID NO: 506 or 507. [0024] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 71 to 82, or a sequence at least about 80% identical thereto. 1005373994
[0025] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 71 to 82 (e.g. in Table 1), or a sequence at least about 80% identical thereto, and one HydC sequence as set forth in in Table 3, or a sequence at least about 80% identical thereto. Preferably, the HydA and HydC combinations are any one shown in rows 40 to 50 in Table 5. [0026] In another embodiment, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 221 to 232, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 221 to 232. [0027] In another embodiment, the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 221 to 232 (e.g. in Table 2), or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 221 to 232, and one HydC sequence as set forth in Table 4 or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the HydC sequence as set forth in Table 4. Preferably, the HydA and HydC combinations are any one shown in rows 40 to 50 in Table 5. [0028] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in SEQ ID NOs: 83, or a sequence at least about 80% identical thereto. [0029] In another embodiment, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in SEQ ID NOs: 233, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of SEQ ID NO: 233. [0030] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 84 to 112, or a sequence at least about 80% identical thereto. 1005373994
[0031] In another embodiment, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 234 to 262, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 234 to 262. [0032] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 113 to 133, or a sequence at least about 80% identical thereto. [0033] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 113 to 133, or a sequence at least about 80% identical thereto, and one or more HydB, HydC, HydD or HyhL sequences as set forth in SEQ ID NOs: 301 to 463, or a sequence at least about 80% identical thereto. Preferably, the HydA, HydB, HydC, HydD and HyhL combinations are any one shown in rows 51 to 70 in Table 5. In one embodiment, the HydA, HydB, HydC, HydD and HyhL are shown in row 69. In one embodiment, the HyhL shown in row 61 is SEQ ID NO: 419 or 420. [0034] In another embodiment, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 263 to 283, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 263 to 283. [0035] In another embodiment, the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 263 to 283, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 263 to 283, and at least one or more HydB, HydC, HydD or HyhL sequences as set forth in SEQ ID NOs: 464 to 626 or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the HydB, HydC, HydD or HyhL sequence as set forth in SEQ ID NOs: 464 to 626. Preferably, the HydA, HydB, HydC, HydD and HyhL combinations are any one shown in rows 51 to 70 in Table 5. In one embodiment, the amino acid 1005373994
sequence of the HydA, HydB, HydC, HydD and HyhL polypeptide is shown in row 69. In one embodiment, the HyhL shown in row 61 is SEQ ID NO: 582 or 583. [0036] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 134 to 136, or a sequence at least about 80% identical thereto. [0037] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of at least one HydA sequence as set forth in any one of SEQ ID NOs: 134 to 136, or a sequence at least about 80% identical thereto, and one or both of a HyhL and HyhS sequence as set forth in SEQ ID NOs: 335, 388, 392, 398, 402, 409, 413 to 414, 418 to 420, 427, 431, 435 to 436, 440 to 441, 445, 455 to 456, 460 to 463, and 627 to 629, or a sequence at least about 80% identical thereto. Preferably, the HydA, HyhL and HyhS combinations are any one shown in rows 71 to 73 in Table 5. In relation to HyhS each of rows 71 to 73 in Table 5 further include a HyhS. Row 71 further includes nucleotide SEQ ID NO: 628. Row 72 further includes nucleotide SEQ ID NO: 629. Row 73 further includes nucleotide SEQ ID NO: 627. [0038] In another embodiment, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 284 to 286, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 284 to 286. [0039] In another embodiment, the nucleic acid molecule encodes a polypeptide or polypeptide complex comprising, consisting essentially of or consisting of at least one HydA amino acid sequence as set forth in any one of SEQ ID NOs: 284 to 286, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NOs: 284 to 286, and at least one or both of a HyhL and HyhS sequence as set forth in SEQ ID NOs: 498, 551, 555, 561, 565, 572, 576 to 577, 581 to 583, 590, 594, 598 to 599, 603 to 604, 608, 618 to 619, 623 to 626, and 630 to 632, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to a HyhL sequence as set forth in SEQ ID NOs: 498, 551, 555, 561, 565, 572, 576 to 577, 581 to 583, 590, 594, 598 to 599, 603 to 604, 608, 618 to 619, and 623 to 626. Preferably, the HydA, HyhL and HyhS combinations are any one shown in rows 71 to 73 in Table 5. In relation to HyhS each of 1005373994
rows 71 to 73 in Table 5 further include a HyhS. Row 71 further includes nucleotide SEQ ID NO: 628 and amino acid SEQ ID NO: 631. Row 72 further includes nucleotide SEQ ID NO: 629 and amino acid SEQ ID NO: 632. Row 73 further includes nucleotide SEQ ID NO: 627 and amino acid SEQ ID NO: 630. [0040] In one embodiment, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 137 to 143, or a sequence at least about 80% identical thereto. [0041] In another embodiment, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 287 to 293, or a functionally equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 287 to 293. [0042] In any of the above embodiments, the nucleic acid molecule does not include any other nucleotide sequence or encode for any other protein that occurs in the native archea organism from which the HydA was derived. [0043] In a second aspect there is provided a nucleic acid construct comprising a nucleic acid molecule as described herein. Preferably, the construct is synthetic, recombinant or isolated. More preferably, the construct may comprise one or more heterologous promoters for enabling expression of the nucleic acid(s) comprised in the construct. [0044] In any embodiment of the second aspect of the invention, the construct comprises a heterologous sequence encoding a tag for enabling the purification of one or more proteins encoded by the nucleic acid sequences as described herein. [0045] In any embodiment of the second aspect of the invention, the construct is in the form of a vector or plasmid. [0046] In any embodiment of any aspect herein, a nucleic acid molecule or construct may comprise a codon optimised sequence for enabling expression of the nucleic acid in a heterologous host. For example, the nucleic acid molecule may comprise a codon optimised sequence for enabling expression of the nucleic acid(s) or operon in a cell that is not an archaea, or not an archaea of a lineage as shown in Figure 1. Preferably, the heterologous host is E. coli. 1005373994
[0047] In a third aspect there is provided a nucleic acid molecule comprising, consisting essentially of or consisting of a nucleotide sequence set forth in any one of SEQ ID Nos: 144 to 150. In one embodiment, the nucleic acid molecule comprises, consists essentially of or consists of a nucleotide sequence set forth in any one of SEQ ID Nos: 144 to 150 without a sequence encoding the C-terminal SSGWSHPQFEK. Preferably, the nucleic acid molecule is isolated, recombinant or synthetic. [0048] In another embodiment, the nucleic acid molecule is codon optimised for expression in a heterologous host, such as E. coli, and encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 286 or 294 to 300; and/or SEQ ID Nos: 464 to 626, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 151 to 286 or 294 to 300; and/or SEQ ID Nos: 464 to 626. Preferably, the nucleic acid molecule does not encode the C-terminal SSGWSHPQFEK as depicted in SEQ ID Nos: 294 to 300. [0049] In a fourth aspect there is provided a cell comprising a nucleic acid molecule as described herein, or comprising a nucleic acid construct as described herein. In a preferred embodiment, the cell is a microorganism. Preferably, the cell is E. coli. A particularly preferred strain of E. coli is DE3. In any embodiment, the cell is a recombinant cell. [0050] In a fifth aspect there is provided an isolated microorganism comprising a nucleic acid as described herein, or a construct as described herein, wherein the isolated microorganism is capable of oxidising hydrogen. [0051] Preferably, the isolated microorganism is any strain shown in Figure 1 or described herein. [0052] The invention also provides a method of producing a hydrogenase as described herein, the method including culturing a cell or microorganism comprising a nucleic acid molecule as described herein, or comprising a nucleic acid construct as described herein under conditions to allow expression of the hydrogenase. Preferable the expression is inducible expression. Preferably, the expressed protein is harvested from the cell under anaerobic conditions to minimise or prevent inactivation of the produced hydrogenase by atmospheric oxygen. 1005373994
[0053] In one embodiment, the cell or microorganism is cultured at a temperature of less than 37˚C after induction of expression, preferably the temperature is less than about 30°C, or less than about 20˚C or about 20˚C. [0054] In a further, fifth aspect there is provided an isolated, recombinant or purified polypeptide capable oxidising hydrogen, wherein the polypeptide comprises, consists essentially of or consists of an amino acid sequence selected from any one of SEQ ID NOs: 151 to 293, or a functional equivalent, homology or derivative thereof having a sequence at least 80% identical thereto. [0055] In one embodiment, thre is provided an isolated, recombinant or purified polypeptide complex capable oxidising hydrogen, wherein the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in any one of rows 1 to 73 of Table 5. In one embodiment, the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 20 without the HyhL shown. In one embodiment, the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 22 without the HyhL shown. In one embodiment, the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 23 with either HydB shown. In one embodiment, the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 22 without the HyhL shown. In one embodiment, the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 61 with either HyhL shown. In one embodiment, the polypeptide complex comprises, consists essentially of or consists of amino acid sequences referred to in row 69 with either HydB, HydC, HydD or HyhL shown. In relation to HyhS each of rows 71 to 73 in Table 5 further include a HyhS. Row 71 further includes amino acid SEQ ID NO: 631. Row 72 further includes amino acid SEQ ID NO: 632. Row 73 further includes amino acid SEQ ID NO: 630. [0056] In one embodiment there is provided an isolated, recombinant or purified polypeptide capable oxidising hydrogen, wherein the polypeptide comprises, consists essentially of or consists of an amino acid sequence selected from any one of SEQ ID NOs: 287 to 293, or a functional equivalent, homology or derivative thereof having a sequence at least 80% identical thereto. 1005373994
[0057] In one embodiment there is provided an isolated, recombinant or purified polypeptide capable oxidising hydrogen, wherein the polypeptide comprises, consists essentially of or consists of an amino acid sequence selected from any one of SEQ ID NOs: 294 to 300, or a functionally equivalent, homology or derivative thereof having a sequence at least 80% identical thereto. [0058] Preferably, the polypeptide is capable of H+ reduction catalysis and/or H2 gas oxidation. [0059] In a sixth aspect there is provided an anaerobically matured enzyme, or enzyme complex, for oxidising hydrogen, wherein the enzyme or enzyme complex comprises proteins comprising the amino acid sequences as set forth in any one or more of SEQ ID NOs: 151 to 300, or functional equivalent, homologs or derivatives having at least 80% sequence identity thereto. The enzyme complex may also be referred to herein as a multiprotein enzyme complex. In one embodiment, the enzyme or enzyme complex has been matured using a synthetic mimic. An example of a synthetic mimic is [2Fe]adt ([Fe2(azadithiolate)(CO)4(CN)2]2−). [0060] In one embodiment, the anaerobically matured enzyme, or enzyme complex, comprises, consists essentially of or consists of the subunits or amino acid sequences as referred to in any one of the rows 1 to 73 of Table 5. In one embodiment, the anaerobically matured enzyme, or enzyme complex, comprises, consists essentially of amino acid sequences referred to in row 20 without the HyhL shown. In one embodiment, the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 22 without the HyhL shown. In one embodiment, the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 23 with either HydB shown. In one embodiment, the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 22 without the HyhL shown. In one embodiment, the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 61 with either HyhL shown. In one embodiment, the anaerobically matured enzyme, or enzyme complex comprises, consists essentially of or consists of amino acid sequences referred to in row 69 with either HydB, HydC, HydD or HyhL shown. In relation to HyhS each of rows 71 to 73 in Table 5 further include a HyhS. Row 1005373994
71 further includes amino acid SEQ ID NO: 631. Row 72 further includes amino acid SEQ ID NO: 632. Row 73 further includes amino acid SEQ ID NO: 630. [0061] Preferably, the enzyme or enzyme complex is capable of H+ reduction catalysis and H2 gas oxidation. [0062] Preferably, the complex comprises a HydA subunit and a HydC subunit; or a HydA subunit, a HydB subunit and a HydC subunit; or a HydA subunit and HyhL subunit; a HydA subunit, a HydB subunit, a HydC subunit, a HydD subunit and a HyhL subunit; or a HydA subunit, a HyhL and a HyhS subunit. Preferably the complex is arranged substantially as depicted in the Examples and Figures herein. [0063] In another aspect, the present invention provides a method of activating a polypeptide, enzyme or enzyme complex as described herein. In one embodiment, activating a polypeptide, enzyme or enzyme complex includes contacting the polypeptide, enzyme or enzyme complex with iron and sulphur sources under anaerobic conditions. Exemplary iron and sulfur sources are ferrous ammonium sulfate and L-cysteine. Preferably, the iron and sulfur sources are both added in 1.5-fold, or about 1.5-fold, molar excess to the desired number of Fe-atoms to be added. [0064] In one embodiment, a method of activating a hydrogenase is outlined in Example 1 as described below. [0065] In any aspect of the invention, a functional equivalent, homolog or derivative of a HydA protein is a protein which retains the same or substantially the same function of a HydA protein as herein described; and/or a functional equivalent, homolog or derivative of an enzyme complex is a protein which retains the same or substantially the same function of an enzyme complex as herein described. [0066] In any aspect of the invention or embodiment described herein, the phrase “at least 80% sequence identity” should be understood to provide basis for at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity. [0067] Further it will be appreciated that a sequence having “at least 80% sequence identity” may consist of a sequence which is about 80%, about 81%, about 82%, about 1005373994
83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98% or about 99% sequence identity. [0068] In any embodiment of the invention, a functional equivalent of any amino acid sequence disclosed herein may include an equivalent amino acid sequence wherein the N-terminal methionine residue is cleaved in the final protein product. For example, any of the amino acid sequences set forth in Tables 2 and 4 should be understood to provide basis for the same amino acid sequences but which do not comprise an N terminal methionine. [0069] In any embodiment of the invention, an amino acid sequence disclosed herein will be understood to provide basis for an identical amino acid sequence wherein the N- terminal methionine residue is cleaved in the final protein product. For example, any of the amino acid sequences set forth in Tables 2 and 4 should be understood to provide basis for the same amino acid sequences but which do not comprise an N terminal methionine. [0070] In a further aspect there is provided a method for converting hydrogen to electrons, the method comprising contacting a source of hydrogen with an isolated or recombinant microorganism as described herein, or with an isolated, recombinant or purified polypeptide, enzyme, or enzyme complex as described herein. [0071] Further, there is provided a use of an isolated or recombinant microorganism as described herein, or an isolated, recombinant or purified polypeptide, enzyme or enzyme complex as described herein, for converting hydrogen to electrons. [0072] In a further aspect there is provided a method for producing hydrogen (H2), the method comprising contacting a source of protons with an isolated or recombinant microorganism as described herein, or with an isolated, recombinant or purified polypeptide, enzyme, or enzyme complex as described herein. [0073] Further, there is provided a use of an isolated or recombinant microorganism as described herein, or an isolated, recombinant or purified polypeptide, enzyme or enzyme complex as described herein, for producing hydrogen (H2). 1005373994
[0074] In a further aspect there is provided a method of generating energy from a source of hydrogen, the method comprising contacting a source of hydrogen with an isolated or recombinant microorganism as described herein, or with an isolated, recombinant or purified polypeptide, enzyme or enzyme complex as described herein. [0075] Further, there is provided a use of an isolated or recombinant microorganism as described herein, or an isolated, recombinant or purified polypeptide, enzyme or enzyme complex as described herein, for generating energy from a source of hydrogen. [0076] In any method of the invention, the isolated or recombinant microorganism, or the isolated, recombinant or purified polypeptide, enzyme or enzyme complex is immobilised or encapsulated. [0077] In any method or use described herein, the enzyme may be immobilised on an electrode. Methods for generating protein films for immobilising proteins on electrodes are well known in the art and are further described herein. In one example, the enzyme is covalently bound to the surface of an electrode. Preferably, the electrode is comprised of pyrolytic graphite edge (PGE), encapsulated in epoxy. [0078] In any embodiment, the microorganism may be encapsulated or immobilised. [0079] In a further aspect there is provided a device for converting hydrogen to electrons, the device comprising an immobilised polypeptide, enzyme or enzyme complex as described herein. The device may otherwise be referred to herein as an electrode. [0080] In accordance with this further aspect of the invention, the polypeptide, enzyme or enzyme complex is provided in the device/electrode covalently bound to the surface of an electricity-conducting material. Preferably, the material is comprised of pyrolytic graphite edge (PGE), encapsulated in epoxy. [0081] In a further aspect there is provided a system for oxidising hydrogen, the system comprising an isolated or recombinant microorganism of the invention or an isolated, purified or recombinant polypeptide, enzyme or enzyme complex of the invention, a source of hydrogen, and means for detecting the oxidation of hydrogen. Preferably, the system further comprises one or more cofactors for accepting the electrons produced by the oxidation of hydrogen. The co-factor may be a small molecule, or an enzyme that accepts the electrons and transfers to molecular oxygen. 1005373994
[0082] In one embodiment, the means for enabling detection of the oxidation of hydrogen comprises a means for direct measurement of electrical current (such as an amperometer). [0083] In a further aspect, the present invention further contemplates the provision of a hydrogen sensor comprising a polypeptide, enzyme or enzyme complex described herein, wherein the sensor is useful for measuring/detecting hydrogen. In certain embodiments, the sensor device comprises a polypeptide, enzyme or enzyme complex immobilised on an electrode, such that upon oxidation of hydrogen by the polypeptide, enzyme or enzyme complex upon contact with hydrogen, and electrons are produced which enter an electrical circuit associated with the electrode. The magnitude of the electrical current generated could be correlated with the amount of hydrogen present. The system may also be calibrated using known amounts of hydrogen to enable the determination of unknown quantities of hydrogen in test samples. [0084] In a further aspect there is provided a fuel cell comprising a device or electrode of the invention. [0085] In a further aspect there is provided an air-powered device comprising a fuel cell of the invention. In one embodiment, the fuel cell generates energy from hydrogen present in the ambient air thereby providing energy to the device. [0086] In still a further aspect there is provided a use of a system, an immobilised polypeptide, enzyme or enzyme complex, or electrode or device as described herein, for determining the concentration of hydrogen present in a test sample. Accordingly there is also provided a use of an isolated or recombinant microorganism as described herein, or an isolated, recombinant or purified polypeptide, enzyme or enzyme complex as described herein, for determining the concentration of hydrogen present in a test sample. [0087] Preferably the use comprises: - obtaining or having obtained a reference data set in the form of data comprising a measure of electron production by the polypeptide or enzyme present in the system, complex or device in response to known quantities of hydrogen, - contacting the polypeptide or enzyme in the system, complex or device with a test sample comprising an unknown quantity of hydrogen, 1005373994
- comparing the electron production resulting from the contacting, to the electron production in the reference data set to thereby determine the concentration of hydrogen present in the test sample. [0088] As used herein, except where the context requires otherwise, the term "comprise" and variations of the term, such as "comprising", "comprises" and "comprised", are not intended to exclude further additives, components, integers or steps. [0089] Further aspects of the present invention and further embodiments of the aspects described in the preceding paragraphs will become apparent from the following description, given by way of example and with reference to the accompanying drawings. Brief description of the drawings [0090] Figure 1: Phylogenetically and metabolically diverse archaea encode [FeFe]-hydrogenases. The left portion of the figure shows a maximum-likelihood phylogenomic tree (model LG+F+G4) based on the concatenated 15 ribosomal marker proteins of archaeal genomes that encode [FeFe]-hydrogenases. Results are shown for the 118 (out of 130) genomes that are at least 60% complete, less than 5% contaminated, and contain at least 75% of the 15 syntenic proteins. Branches are colour-coded encoding according to the respective phylum. Black circles indicate bootstrap support values over 80%. The middle portion shows the presence of key metabolic genes involved in different metabolic processes. Carbon fixation: ATP- citrate lyase beta-subunit (AclB), acetyl- CoA synthase beta subunit (AcsB), propionyl- CoA synthetase (PrpE), 4-hydroxybutyryl- CoA dehydratase / vinylacetyl-CoA-delta- isomerase (AbfD), CODH/ACS complex subunit delta (CdhD), CODH/ACS complex subunit gamma (CdhE), anaerobic carbon monoxide dehydrogenase catalytic subunit (CooS), type II/III ribulose-bisphosphate carboxylase (RbcL II/II), type III ribulose- bisphosphate carboxylase (RbcL III); respiration: reductive dehalogenase (RdhA), formaldehyde activating enzyme (Fae), glutathione-independent formaldehyde dehydrogenase (FdhA), the reversible succinate dehydrogenase and fumarate reductase flavoprotein (SdhA/FrdA); ATP synthesis: ATP synthase subunit alpha (AtpA), ATP synthase subunit beta (AtpB); fermentation: 2- oxoacid:ferredoxin or pyruvate:ferredoxin oxidoreductase alpha subunit (PorA/OorA), isocitrate dehydrogenase (Idh), ADP-forming acetyl-CoA synthetase (AcdA), acetate kinase (Ack), phosphate acetyltransferase (Pta), acetyl-CoA synthetase (Acs), formate C- acetyltransferase (PflD); fatty acid degradation: acyl-CoA dehydrogenase (ACAD); 1005373994
aromatics degradation: flavin prenyltransferase (UbiX); sulfur metabolism: sulfur dioxygenase (Sdo), sulfate adenylyltransferase (Sat), adenylylsulfate kinase (CysC), sulfate adenylyltransferase subunit 1 (CysN), anaerobic sulfite reductase subunit A (AsrA). The right portion shows the diverse environments the archaeal genomes were retrieved from. Phylum QMZS01 was classified as Aenigmatarchaeota in GTDB R06- RS207 while Thermoproteota class EX4484−205 was proposed as Brockarchaeia. [0091] Figure 2: Archaea encode genetically and structurally diverse [FeFe]- hydrogenases. Catalytic domain structure, genetic organisation, and AlphaFold2- based structural modelling of representative [FeFe]-hydrogenases encoded in archaeal genomes. (a) Group A1 [FeFe]-hydrogenase from UBA95 sp002499405 (Micrarchaeota). (b) Group E [FeFe]-hydrogenase from Ca. Forterrea multitransposorum. (c) Group B [FeFe]-hydrogenase complex from Ca. Prometheoarchaeum syntrophicum. (d) Group A3 [FeFe]-hydrogenase complex from DSAL01 sp011380095 (Altarchaeota). For each panel, the catalytic domain (H- cluster), iron-sulfur binding motifs, and amino acid sequence length are shown at the top. Genes encoding hydrogenase structural subunits are shown in their genetic context are shown beneath, labelled and coloured consistent with the corresponding subunit in the structural models. Predicted cofactors are positioned based on the structures of homologous proteins. A zoomed view of the H- cluster and conserved coordinating cysteine residues (C1 to C5) is shown for each group. For group B and A3 enzymes, FeS clusters within plausible electron transfer distance are connected by dashed lines. hydA, [FeFe]-hydrogenase; hydB, diaphorase; hydC, thioredoxin; hydD, nuoG-like conduit protein; hyd6TM, uncharacterised 4 to 6-helix transmembrane protein associated with group A [FeFe]-hydrogenases; (His)[4Fe4S], (Cys)3His-ligated [4Fe4S] cluster binding domain; [2Fe2S], [2Fe2S] cluster binding domain; [4Fe4S], [4Fe4S] cluster binding domain; 2[4Fe4S], bacterial ferredoxin-like 2[4Fe4S] cluster binding domain; 6Cys, putative iron-sulfur cluster binding domain. * HydC protein in group A1 gene cluster was not predicted to form a complex with HydA. Surface structures are used for the multisubunit group B and A3 [FeFe]-hydrogenases, with ribbon diagram versions provided in Figure 8. [0092] Figure 3: Three classes of [FeFe]-hydrogenases are catalytically active in archaea. (a) H2 gas production monitored from cell lysates in E. coli BL21(DE3) cells expressing group A1, B, and E [FeFe]-hydrogenases from archaea. The cell lysates were activated by addition of [2Fe]adt. H2 was measured by GC after addition of methyl viologen 1005373994
and dithionite to activated cell lysates, set to pH 6.8 with 100 mM KPi buffer. Activities are normalized for number of cells used (nmol H2 min-1 OD600-1) and error bars reflect standard deviation from biological triplicates. The strain expressing prototypical CrHydA1 was used as a positive control while “Blank” represents the same strain but containing an empty vector. (b) FTIR spectra of the group E [FeFe]-hydrogenase from Ca. Forterrea multitransposorum (Fm) after heterologous expression, semisynthetic maturation with [2Fe]adt, and purification. The absorbance spectrum (top) indicates a CO inhibited di- ferrous H-cluster state (Hsox-CO). The difference spectrum (bottom) illustrates the transitions of Fm into catalytically active states through photoreduction (illumination after the addition of eosin Y as a photosensitizer and triethanolamine as a sacrificial electron donor). During illumination bands associated with the highly oxidized CO-inhibited state decreased (grey bands), while new bands reflecting reduced and catalytically active H- cluster states appear, assigned to HoxH (cyan), Hox (blue) and Hred (red) (spectra arranged chronologically from top to bottom). (c) Cyclic voltammetry traces of immobilized Fm (orange) with H2 oxidation current densities at high potentials and H+ reduction currents at low potentials. The 2H+/H2 redox couple potential is indicated with a dashed line (E0′ 2H+/H2). Scan direction is indicated by black arrows. The “Blank” trace (grey) represents the electrode without an immobilized enzyme film. The experiments were performed on two independent films for each enzyme at pH 7.0 (5 mM MES, 5 mM CHES, 5 mM HEPES, 5 mM TAPS, 5 mM NaOAc, 0.1 M Na2SO4) and under 1 atm H2. [0093] Figure 4: [FeFe]- and [NiFe]-hydrogenases form unique complexes in archaea. (a) Catalytic domain structure and predicted operon encoding a putative complex of a group F [FeFe]-hydrogenase and group 3 [NiFe]-hydrogenase in Thermoplasmatota UBA147 sp002496385. hydA, [FeFe]-hydrogenase; hydB, diaphorase; hydC, thioredoxin; hydD, nuoG-like conduit protein; hyhL, group 3 [NiFe]- hydrogenase catalytic subunit; hyhS, group 3 [NiFe]-hydrogenase small subunit. (b) Predicted surface structure, cofactor composition, and electron flow through four potential arms in the hybrid hydrogenase complex. [FeS] clusters are numbered and labelled according to their subunit of origin (e.g. A1, A2, A3 originate from the HydA subunit). (c) Atomic structure of the predicted [FeFe]- and [NiFe]-hydrogenase active sites in the hybrid enzyme. Distances between catalytic cluster and coordinating residues of less than 2.5 Å are shown as blue dotted lines. (d) H2 gas production monitored from cell lysates in E. coli BL21(DE3) cells expressing the Th1 and Th2 [FeFe]- hydrogenases from archaea. The cell lysates were activated by addition of [2Fe]adt. H2 levels were 1005373994
measured every 15 mins for 2 hours by gas chromatography after addition of methyl viologen and dithionite to activated cell lysates, set to pH 6.8 with 100 mM KPi buffer. Activities are normalized for number of cells used (nmol H2 OD600 -1) and error bars reflect standard deviations from two biological triplicates. “Blank” represents the same strain but containing an empty vector. [0094] Figure 5: [FeFe]-hydrogenases are diverse, ancient, and potentially ancestral in archaea. Maximum-likelihood phylogenetic tree of the catalytic subunit (HydA) of [FeFe]-hydrogenases and novel hybrid hydrogenases. The tree was constructed based on 3,677 amino acid sequences using the LG+F+G4 model. The numbers at the branches indicate the ultrafast bootstrap support values. The tree was rooted using the NADH-quinone oxidoreductase subunit D (NuoD) and formate dehydrogenase alpha chain (FdhA) from Methylorubrum extorquens. [0095] Figure 6: Genome statistics, hydrogenase maturases, and hydrogenase domain structure of 130 [FeFe]-hydrogenase-encoding archaeal genomes. (a-b) Barcharts showing the size and completeness (CheckM) of [FeFe]-hydrogenase- encoding archaeal genomes from GTDB R05-RS202 (77 in total), newly assembled metagenomes (40 in total), and PATRIC (13 in total). (c) Heatmap showing the detection of known [FeFe]-hydrogenase maturase (HydEFG) homologs based on homology search. Navy shading denotes the presence of the homolog. (d) Domain and iron-sulfur cluster organization of archaeal [FeFe]-hydrogenase. Incomplete open reading frames are denoted by asterisks (*) next to arrows. Subclass (protein domain structure-based scheme) and subgroup (protein phylogeny-based scheme) classification of [FeFe]- hydrogenases are based on Land et al.2020 and Greening et al.2016, respectively, with proposed modifications. H−cluster, catalytic domain of [FeFe]-hydrogenase; hyhS, [NiFe]-hydrogenase small subunit / iron-sulfur domain; (His)[4Fe4S], (Cys)3His-ligated [4Fe4S] cluster binding domain; [2Fe2S], [2Fe2S] cluster binding domain; [4Fe4S], [4Fe4S] cluster binding domain; 2[4Fe4S], bacterial ferredoxin-like 2[4Fe4S] cluster binding domain; 6Cys, putative iron-sulfur cluster binding domain. Note that the phylum QMZS01 was classified as Aenigmatarchaeota in GTDB R06-RS207 while Thermoproteota class EX4484−205 was proposed as Brockarchaeia. [0096] Figure 7: Genetic organization of 136 archaeal [FeFe]-hydrogenases. Up to 10 genes upstream and downstream of the [FeFe]-hydrogenase (hydA) are shown. Gene length is shown to scale. hydA, [FeFe]-hydrogenase; hydS, [FeFe]-hydrogenase small 1005373994
subunit; hydB, [FeFe]-hydrogenase diaphorase subunit; hydC, [FeFe]-hydrogenase thioredoxin subunit; hydD, [FeFe]-hydrogenase nuoG-like conduit protein; hydF, [FeFe]- hydrogenase H-cluster maturation GTPase; hyd6TM, uncharacterised 4 to 6- helix transmembrane protein associated with group A [FeFe]-hydrogenases; hyhL / hoxH, group 3 [NiFe]-hydrogenase catalytic subunit; hyhS / hoxY, group 3 [NiFe]- hydrogenase small subunit; hyhD / hoxD, group 3 [NiFe]-hydrogenase iron-sulfur subunit D; hyhB, group 3 [NiFe]-hydrogenase diaphorase electron transfer subunit; hyhG, group 3 [NiFe]- hydrogenase diaphorase catalytic subunit; hyaD, [NiFe]- hydrogenase maturation protease; hypABCDEF, [NiFe]-hydrogenase maturation factors; hdrABC, CoB-CoM heterodisulfide reductase subunits; oorAB / porAB, 2- oxoacid:acceptor or pyruvate:ferredoxin oxidoreductase alpha and beta subunits. Nucleotide and amino acid sequences of each gene are available for example in Tables 1-4 and in the sequence listing, incorporated herein by reference. [0097] Figure 8: AlphaFold2 models of archaeal [FeFe]-hydrogenases. (a) Genetic organisation and model of the group A1 [FeFe]-hydrogenase from Ca. Iainarchaeum andersonii. (b) Genetic organisation and model of the group E [FeFe]-hydrogenase from CABMGN01 sp902385635 (Nanoarchaeota. (c) Model of the Group B [FeFe]- hydrogenase complex from Ca. Prometheoarchaeum syntrophicum. (d) Model of the group A3 [FeFe]-hydrogenase complex from DSAL01 sp011380095 (Altarchaeota). (e) Model of the complete group F [FeFe]-hydrogenase from Thermoplasmatota UBA147. (f) Model of the HydA-HyhS fusion from the group F [FeFe]-hydrogenase from Thermoplasmatota UBA147. [0098] Figure 9: Conservation of active site residues in different classes of [FeFe]- hydrogenases. (a) Structural view of the active site of Clostridium pasteurianum [FeFe]- hydrogenase (CpI, PDB ID: 4XDC) showing the H-cluster and interacting amino acid residues. (b) Normalized consensus logos of [FeFe]-hydrogenase groups A-F generated in Jalview using a ClustalΩ sequence alignment of sequences retrieved from Greening et al.2016 and current data. Coloring is based on the Clustal X color scheme. Numbering is based on CpI and black numbers are illustrated in the top panel. (c) Amino acid sequences of the archaeal [FeFe]-hydrogenases that were heterologously expressed. [0099] Figure 10: SDS-PAGE visualising the molecular weights of the heterologously expressed archaeal [FeFe]-hydrogenases. Expression constructs with verified sequences were transformed in chemically competent E. coli BL21(DE3). 1005373994
Protein bands are shown from before induction with IPTG (B), after induction (Name- Subclass and with the expected kDa size in parenthesis), and lysate or supernatant after cell lysis and centrifugation (L). The bands in each after-induction lane corresponded well with the expected molecular weights in kDa. Both group A1 [FeFe]-hydrogenases (Mu and Ia) had the highest expression and solubility levels. In contrast, the group F (Th1, Th2) and group B (Ps) enzymes exhibited moderate expression levels but poor solubilities in aqueous solutions. The group E (Na and Fm) enzymes exhibited high expression levels but poor solubility. [0100] Figure 11: Isolation and reconstitution of [4Fe-4S]+ cluster of Fm. (a) SDS- PAGE gel of purified Fm (33 kDa), obtained following expression in E. coli Origami™ B(DE3) and StrepTrap XT purification. Target protein indicated with horizontal arrow. (b) UV- visible spectra of Fm (205 µM) after reconstitution of [4Fe-4S]2+ cluster (blue spectrum), indicated by the absorbance at 405 nm. The Fe/protein content was 4.2 ± 0.4 after reconstitution, in agreement with the presence of a single [4Fe-4S] cluster. Upon addition of 20× excess sodium dithionite (NaDT, red spectrum), a decrease in absorbance at 405 nm is observable, indicating the reduction of [4Fe-4S]2+ cluster to [4Fe-4S]+. Spectra were collected in a 1 mm pathlength cuvette. [0101] Figure 12: FTIR difference spectra and redox state kinetics of Fm [FeFe]- hydrogenase. (a) The full set of difference spectra of the photoreduction experiment is shown in Figure 3b. The super-oxidised CO inhibited Hsox-CO species (grey bands) depopulates in favour of the one-electron-reduced oxidised states HoxH (cyan bands) and Hox (blue bands). During continuous photoreduction the oxidised species get further reduced to the [4Fe4S] cluster reduced state Hred’ (red bands). Illumination that facilitates photoreduction was applied for 88 seconds (compare b). (b) The summed delta peak area of each redox state of the difference spectra in (a) is plotted over the time course of the photoreduction experiment. The illumination period is indicated by the grey area. The depopulation of the super-oxidised CO inhibited Hsox-CO species (black) is mostly completed within 44 seconds. At the same time the oxidised states HoxH (cyan) and Hox (blue) reach their maximum population during photoreduction. Subsequently during the illumination period Hred’ accumulates at the expense of the oxidised species (until 88 seconds). After photoreduction (88 seconds) Hred’ converts back into the oxidised species. No re-population of Hsox-CO was detected. 1005373994
[0102] Figure 13: Phylogenetic tree of HydA sequences constructed with LG+FO+R model. The maximum-likelihood phylogenetic tree was constructed based on 3,677 amino acid sequences of the catalytic subunit (HydA) of [FeFe]-hydrogenases and novel hybrid hydrogenases. The model finder was used with default parameters and no mixed model testing; the best-fit model identified was LG+FO+R. Ultrafast bootstrap support values are denoted on the branches. The tree is rooted using the NADH-quinone oxidoreductase subunit D (NuoD) and formate dehydrogenase alpha chain (FdhA) derived from Methylorubrum extorquens. [0103] Figure 14: Phylogenetic tree of HydA sequences constructed with LG+C60 model. The maximum-likelihood phylogenetic tree was constructed based on 3,677 amino acid sequences of the catalytic subunit (HydA) of [FeFe]-hydrogenases and novel hybrid hydrogenases. The best model including the LG+C60 was utilised to test the amino acid replacement rate to vary across different sites. Ultrafast bootstrap support values are denoted on the branches. The tree is rooted using the NADH-quinone oxidoreductase subunit D (NuoD) and formate dehydrogenase alpha chain (FdhA) derived from Methylorubrum extorquens. [0104] Figure 15: Phylogenetic tree of the [FeFe]-hydrogenase maturase HydE. Different colours show archaeal (red), bacterial (black), and eukaryotic (green) sequences. Evolutionary history was inferred using the Maximum Likelihood method and JTT matrix-based model with 50 bootstrap replicates and midpoint rooting. [0105] Figure 16: Phylogenetic tree of the [FeFe]-hydrogenase maturase HydF. Different colours represent archaeal (red), bacterial (black), and eukaryotic (green) sequences. Evolutionary history was inferred using the Maximum Likelihood method and JTT matrix-based model with 50 bootstrap replicates and midpoint rooting. [0106] Figure 17. Phylogenetic tree of the [FeFe]-hydrogenase maturase HydG. Different colours represent archaeal (red), bacterial (black), and eukaryotic (green) sequences. Evolutionary history was inferred using the Maximum Likelihood method and JTT matrix-based model with 50 bootstrap replicates and midpoint rooting. [0107] Figure 18: Phylogenetic tree of the [NiFe]-hydrogenase large / catalytic subunit (HyhL) with focus on group 3 [NiFe]-hydrogenases. The subunits predicted to associate with [FeFe]-hydrogenases are shown in red (for group F [FeFe]- 1005373994
hydrogenases) and purple (for group G [FeFe]-hydrogenases). Evolutionary history was inferred using the Maximum Likelihood method and JTT matrix-based model with 50 bootstrap replicates and midpoint rooting. [0108] Figure 19. Phylogenetic tree of the [NiFe]-hydrogenase iron-sulfur domain / small subunit (HyhS) with focus on group 3 [NiFe]-hydrogenases. The domain that fuses with the group F [FeFe]-hydrogenases is shown in red. The unfused subunit encoded downstream of group G [FeFe]-hydrogenases is shown in purple. Evolutionary history was inferred using the Maximum Likelihood method and JTT matrix-based model with 50 bootstrap replicates and midpoint rooting. Sequence information [0109] Tables comprising sequence information Table 1: Nucleotide sequences of HydA SEQ Description Subgroup, Sequence ID subclass NO: 1 AQRS01000037.1 A1, M1 ATGGGTTCGATTGAGGACGTGAATGCCGCGCTCGCGGATGAAGGGA _10, nucleotide AAATGGTCATGGCGCAGGTAGCGCCCGCGGTTAGGGTTACTATCGG CGAGGAGTTCGGCCTTCCGGCGGGAACAATTGTGACGAAAAAGCTC GTGGGCGCGTTGAGGCAGGCCGGCTTTGAAAAGGTGTTTGACACCT CCGTTGCCGCGGATATTGTAACAATTGAGGAAGGAACGGAATTCCTG AACAGGCTCGAGGACCAGGAGGACCTTCCATTGCTGACTTCCTGCTG CCCTGCATCGGTTTTTTTTGTTGAGAACACTTTCCCGAAATTTTTGCAC CACTTCTGCACTGTTAAAAGCCCGCAGCAGGGCATGGGCTCGCTCAT AAAAACCTATTACGCGCGCAGGATGAAGATTGATGCCAAAAAAAATTT TGTTGTCGCTGTGATGCCGTGCATCGTGAAAAAAATGGAAGCGCGCC GCCCTGAAATGGAGTTCGACGGCGTGCATAACGTGGATGCGGTTCTT ACCACAAAAGAGGCAGCTGCTTTGCTCAAATCAAAGAAGGCCGATTT GAACGGCGCGAATGAAACGGGGTTCGACAGTTTGCTCGGAAAGGCT TCGGGAGCAGGCCAATTGTTCGGGCAAACTGGCGGGGTTTCTGAAG CTTTGCTGAGGTTTGTCGCGTGGAAGCTCGAGGGGAAAAAAGCGAG GGTTCTTTTCAAGGAGGTGCGCGGGAAAAAGGGTTTTCGCGAGGCC GAAGTGAAAATAGGCTCGAGGATGCTGAAGGTTGCGGTCATTGATGG GCTGAACAATTTGCGGGATTTGATGAGCAGCGAGGAAAAATTCCATT CGTATGATGTCGTGGAGATAATGACCTGCCCGGGTGGATGCATCGGC GGAGGCGGCCAGCCGGAATCAACGCCTGAAAAGCTGGAAGCGAGGA AAAAGGCGCTGCATAATGTGGATGCAATGGAAACCGTGAGGGTTGCC TCTGAAAACCCTGAAGTGAAGGAGCTTTACGGCTCATATCTCATTGAG 1005373994
CCGGGCTCGAAGGTCGCGCGCTCAATCCTGCACGTGAGCAGGATCT GCCTCAAGTGCGACTAG 2 CABMCJ0100000 A1, M1 ATGCGCAGGAGGGTTAAAAAAACTAACGAATTCGGGTTAGACAGGTG 04.1_889, GCGATATGGGAAATTTCTTAAAGTGTGCGCGAAAATCGTTATTATGCA nucleotide GGCAGACCCTATTGGAAAGTTTAACGAGATATGCGGATTGAAAAAGG AGCGCACGCTTTGCGTGCAGATTGCGCCTGCAGTAAGGGTTGCGCTT GGGGAAGAGTTCGGGATGCCCATTGGAACGGACGTTACCAAAAAGCT TGTGTGGGCAATGAGAAAAATTGGGGCGGAATATGTATTTGATACTCC TCTTGGAGCGGATATAATTGCAATGGAAGAGGCAAACGAGCTCAAGA AAATGCTTGAGGCTGGGGGACCGTTCCCTGTATTCACCTCTTGCTGT GCAGGCTGGATGCTTTTTTGGAAAAGGGCGCATCCGGAGCTGGAGA AGAATGTGTGCGAGCTTGTTGCGCCTCAAATGGCATTGGGCGCGCTC ATAAAAACATATTTTGCAAGAGTGAGCGGCATGAAGAAAGAGCCTTAT GTACTCTCCATAATGCCCTGCCTTTTGAAAGAGCAGGAGGCGCAGGG CACAATGAAAGACGGGAGAAAATATGTGGACTGCGTTTTCACTACAAA AGACATTGCAAAAGTGCTGAAGAGCGAGGGAATAGATTTGAAGAATG CGCCTGAAGAGGAGTTTGACAAAATTTCAGGCATGGGCTCGGGTGAA GGGGCTATTTTTGGGGCAAGCGGAGGAGTGATGGAGGCTGCACTGC AAAATTTGGGGAAAATGCTTGGGGAGAAAGTCGAGTTCAAGGAATTT CGAAATGAAGAGAACATGAAAAAGAGCACTGTGAAGATTGGGAAATA CACGCTGAACGTTGCTGCCGTGTGGGGCCTGCCGAATGTGGACAAA TTACTTGCGGAAATGACAGAGGGAAAAGTTTACCATTTTGTGGAGGTT ATGGCGTGCCCGGGAGGGTGCATTGGGGGAGGGGGACAGCCGCTT CCGCCTACGCCGGACGTTGTGAAGGCGCGCGCTGCTGCGATGCGAA AATATGCTGACATGCTGAAAATAAGCGCAAAAGAGAACCCGAAAGTG GATGAGATATACCGCGCATTTTTGAAAAAAGTGGGGAGCGAAATCGG GCACGAGCTTTTTAGTTCCAGAAAGAAAAATTAG 3 CABMDO010000 A1, M1 ATGGAAGAACTCGAGCGCCTGCTTTCATCGAAGAAAATCGTAGCGGT 013.1_20, CCAGACTGCGCCGTCGGTGCGCGTGACTCTCGGCGAGGAGTTCGGT nucleotide TTGAAGCCCGGCACCGATGTCACGGGGCGCGTCGTAGCCGCGTTGC GCATGCTCGGGTTCAAGTACGTTTTCGACACGGATTTCGGAGCCGAG GTTACTGTGCTCGAGGAGGCGCAGGAGCTTCTCGAACGCCTTGAACA GCAGGAGCGCCTGCCTCTTTTTACCTCGTGCTGCCCGGGCTGGACG CGCTTCTGCCTGAAACAGTTCCCGACGCTGGAGCAGAACCTTTCCGA GTGCAAGGCGCCGCAGCAGATTCTCGGCTCGCTCGCGAAGACGTAT TTCGCGAAGAAAATCCGCAAGCAGGCGAAGGACGTCGCCGTCGTAT CGATAATGCCCTGCTTCTCCAAGAAACTCGAAGCCAAGGAGCGCGAG CATGCCATCGCAGGGGCGTCGCAGGACGTCGATCTGGTCATAACGA CGCACGAGCTCGCGCAGATGCTCAAGGCCAAGGGCATAGACCTGCG CAGGCTTTCCGACGAGGAGTTCGACAATCCTCTCGGAGTAACGGGC GGGGCGGGTGCGCTGTTCGGCGCAACCGGCGGAATCGCCGAGGCG GTTATGCGCACTGCGTATCATGCAACAACTGGCGGCGCCCTTCCGCG CTTCGAATCGAACATGCGGCCAGCGCTCGGGCAAATCAAGGAGTATT CCATCATGATGGGAGACCGGGACGTGCATCTCGCGATAGTGCAGGG CTACCCCAACGCTTCCGTAGTGTGCAAGCGCGTGCTCGAGGAAAAGC 1005373994
GCAACGGAATGCGTTCGCTCGATTTCGTCGAGGTGCTCGGTTGCATC GGGGGCTGCGTCGGCGGTCCGGGGCAGCCCGATACGGGCGAGGAT GCGGTGCTGATGCGCGCGGCGGCACTGCACGAGCGCGATCGCCGC CTACCGCTCAGGAGCGCGCACGACAACCCCGCGGTGAAGAAGGTTT ACGCGTCGATGCTCGGGAAGCCCGGAAGCGCAAAGGCTAAGAAGTT CCTGCACCGGGGGCAGCGCGGTTATGCTGAGATGCTGAAGACGCTG TGA 4 CAIKOA0100000 A1, M1 ATGTCCACAGAGATTTACCTCAATAAGCCCGAGGAGCTTTTCGCAGC 26.1_15, GCTAGAAAAAGAGCTCGCAGACCCAGAAAATACCGTGGTTGTGCAGA nucleotide TCGCCCCCGCAGTGCGCGTCGCTATTGGCGAGGAATTCGGCTATCC CTCCGGAGAAGACCTCACATTCAAAACAATCGGGCTCCTGAATGCGC TCGGCTTCAAGCACGTAGTGGATACTCCTCTTGGCGCTGACATAAAC ATCTACGAGGAGGCTTACGAAATCCTCAACGCGCTCGAGCGCAAGGA CGACAATTACTTCCCGATTTTCAATTCCTGCTGCATCGGCTGGAGGTT ATATTGCTCCCGCGCACACCAGAAGGTCTGCCAGCACATCTCCCCGA TCGCTTCCCCGCACATGATTACGGGCAGCGTGATAAAGCATTATTTCT CCAAAAAACTGAACAAGCCCAAGGAGAAGATAATCTCCGTGAGCATA ATGCCCTGCGTCCTCAAGAAATATGAAAGCCTTGAGCGCTTCTCCGA CGGCTCCACCTATATTGATTACGTCGTCACAACCCGCGAGCTTGCGC AGTGGGCAAAGAAGAGGGGCCTGGACCTGCGCAAGGTGAATGAGGG GAAATTCTCAGAATTCCTTCCCAACTCCTCAAAGGATGGAGTAGGCTT CGGAGTCACCGGAGGCCTCGTTGAAGCCCTTCTCACTACAATGGCGC ACATTCTCGAAGTGAAGAAGGAAGCCGAATCCTTCCGCACCAACGAA CCCATAAAGGAGCGGCAGGTGAAAATAGGCGAATACACCCTGAATGT CGTTTCCATAAATGGCTTTGGGAACTTCGAGAAAGTCCTGAAGGAAAT CGAGTCCGGAGAGAGGAAATTCCACTTCGTAGAGGTAATGAACTGCC CATACGGCTGCGTTGGCGGCCCCGGCCAGCCCCTTCCAGTGAACGA TGAAATTCTTAAGGCAAGGGCGGCAGGGCTGCGCCTTGCAGCAGAC AAAAAACGCTCAATTGCAATCGTCCCGCAGGAAAACCCAACCGTGCA GATGCTCTACCGGGAGCTCCTGGGCCGCCCTGGCTCGGAGAAGGCG CGGGAGCTGCTCTATTTCCACAAAATAAAGCTCTAG 5 CAIWXP0100001 A1, M1 ATGGCAGAAGAAATTTACATGAATAAGCTAGATGACCTAATTAATGCC 75.1_2, nucleotide ATTGAAGCAGAAAAGAAGAAAGGTAAAGTATTGGTTGCTCAATTGGCT CCTGCTGTACGAATAACACTAGGAGAAGAGTTTGGCTATAATGTAGGT GAGGATTTAACCAAAAAATGTGTTGGTTTATTGAAAACTTTGGGTTTTG ATATAGTTATTGATACCCCTTTAGGTGCAGACATCGCAGTATATGAAG AAGTACATATTTTAAAAGAATTATTAGACTCAAAAGACCTAAGCTATTT CCCAATGTTTAACTCTTGCTGCATAGGCTGGAAAATGTATGCCAAAAG AATGCACCTAGACCTAATGCCCCATGTTTCTAGCATCGGTTCACCAAA TCAAATTGTTGGAAGCATAGCAAAGAATTACCTAGCATATCAACAAAA CAAATATCCTGATGATATTGTTGTAGTAGGAATAATGCCCTGCACTTTA AAGAAATTTGAGACTTTAAATACATTTGAACACAAAAAAGGAAATACCA TTGTTAAATTAAAATATGTTGACTATGTTGTTACAACAACAGAATTAGC AGAATGGTCAAGAAAGAAAAATATTGATTTTTCTCAAGTAGCAGAATAT GAAATGATTAATCCAGCATCAAAAGAAGGAACAATCTTTGGTGTTACT 1005373994
GGCGGCATAAGCGAAGCATTCATAAACGCATTTGCAAAATATATTGGT GAAGAAAAAGAAATTTTAGATTTTAGACAAGATGAAAAATCAAGAAAAT ATAAAGTTAAAATTGGCAACTATGAATTAAGTGTTGCAATTGTTTTTGG CGTAGGTCAATTAGAACATATTTTAGAAGACATAGAAAATGGAGAATT TTTTCATTTTGTTGAAGTAATGTATTGTAACTCTGGATGTGTAGGTGGT CCAGGACAACCTAGAGCCTCAGATACAACAATAGAAGAAAGAGCAAA AGCAATGAGAAATTGTTCAGATAAAATAAAAAAAACTACTTGTTTTGAT AATGGTGCATTATTAGAAATGTATAAAAATTTAAATATGAAACCTAATA ATGCAAAAGCAAAAGAAATGTTTTTCCTAAAAGAACAAAATTAA 6 CAIXOA0100001 A1, M1 ATGGAAGACGAAATTTACCTCAATAAGCCGGACGAGCTGTTCGCGGC 93.1_5, nucleotide CATAGAGAACGGGAAAAAGAACGGGAAGATAATGGTTGTGCAAATCG CCCCCGCGGTGCGCGTCTCCATCGGCGAGGAGTTCGGGCGCGCAC CTGGAGAGGACCTTACCCACAAAACCGTTGGGCTCCTGCAGGCGCT CGGCTTCGACCACGTCATGGACACCCCCCTCGGCGCTGATGTAAACA TTTACGAGGAGACGCTGGAAGTGCTGCACGCGCTCGAGCGCGACGA TGAGAAGTATTTCCCGGTGTTCAACTCCTGCTGCATCGGCTGGAGAT TGTATTGCAAGAACAAGCACCCGGAGCTCTACCGCCTGGTTTCGCCC ATCGCCTCGCCCCACATGATTGCGGGCAGCATCGGGAAGCACGTGC TCGCGAAAAAACTCGGGGTTCCGGTTGAAAAAATCTGCATGGTGAGC ATCATGCCCTGCGTGCTGAAGAAATACGAGACGCGCGAGCGGCTCC CGTCCGGAATAAAATACATAGATTACGTGCTGACGACGCACGAGCTC GGGATGTGGGCTAAGAAAAAAGGGCTGGACATAAACAGGGTGAGGG ACGGGAAGTTCACCGCGCTCCTGCCGGACAGCTCCAAGGACGGAGT CATATTCGGTGCGACCGGAGGAATCACCGAGGCGCTTTTGAGCACG CTCGCATGCATATGCGGGGAAAGCCCGGAGAAGGTGAGGTTCAGGG GCGACGAGCAGGTGAAGCACCTGTGCGTGCAGATAGGGAAGCACCA GCTCAACGTGGTTTCCATCTATGGGGTGACCAATCTGGACAAGGTGC TTGATGAGATAAAGCACGGGGTGAAATACCATTTCGTGGAAGTGATG AACTGCCCGTACGGGTGCGTGGGAGGGCCGGGGCAGCCCCTTCCC GCGAGCGAGGAAAAGTACAGGGCGAGGGCGCACGGGCTGAGGAAG GCCGCGGACAGGAAGCAGGGCAAGTGCCCGCTCGGGAAGATGGGC GTGCATTCCATCTACGAGGCGCTCGGCATAGGGCCCGGAAGCAGGG AGGCGCAGGAGCTTTTCTTCTTCCACAAAACGAACATCTGA 7 CG08_land_8_20 A1, M1 ATGAGATTTGAAATTTTGGAAAGAGCTGTAAGGGATAAGAGTATTCTA _14_0.20_scaffold AAGGTTGCACAATTAGCTCCAGCAGTAAGGGTAAGTTTAGGGGAAAT _32353_2, GTTTGGATTCGAAGCAGGAACAATTCTTACAAAAAAAATAGTTGGAGC nucleotide ACTTAAGGAATTAGGATTTGATTATATTTTTGATACGAGTTTTGGAGCA GATGTTGCAATTGTAGAGGAGAGCAAGGAACTGGGGGACAGACTAAA GAAAGGTGGAATATTTCCAATGATTAATTCCTGCTGTCCAGGGACAAT AAGCTTCCTGGAACATGCTTACCCAGATTTGGTTCCAAATATTGCTAC TGTAAAAAGCCCGATGGAAGTAACTGGAGTACTAATTAAAACGTATTT TGCAGAAAAGAAGAAAATTCCTCCAGAAAAAATACTCTCAGTTGCAGT TATGCCGTGCATAATCAAGAAGGCTGAGGCATTTAGGCCAGAATTGA GAATGAATGGGAAACTAGTAATAGACGGAGTTCTGACTACAGTTGAG CTTGGAGAATTACTAAAGGCAAAGAAAATTGACCTGAAAAATTGCAAG 1005373994
GAAAGGGAATTTGATTCCTTAATGGGCGTTGCTTCTGGTGCTGGGCA GATATTTGGCTCTACTGGAGGAGTAAGTGAGGCTGCAATTAGGAATTA TGCACACATGAATAATATTCCAATTGAGAAAATAGATACAAAACAACTG AGAAGTTTTGAGGGAGTAAGGGAAATGGAATTTTCTCTTGGTGAAAAA AAGATTAAGATTGCAATAATAAATTCCCTTAGAAATGCAAATCAAGTGT TAAATGATTCAGATAAAATGAAGGAGTATACGTTCATAGAAATTATGG CTTGCTTGGGTGGATGTGTTGGAGGAGCTGGACAGCCAACTTCAACT AAGGAAATTCTAGAGAAAAGAAGAGCTGGGCTGTATTCAATTGATGCA AAAACCAAAATAAAAATTTCTTCTGAAAATCCAGAAGTGAAAAAACTTT ATGAAGAATTCTTAGGGGCTCCAAGAAGTAAAAGAGCATTGAGAATAC TGCACACTGGATTTGTAAGAACATGTGTAGATTGCTTCTAA 8 CG10_big_fil_rev A1, M1 ATGGGTTTTGAGAAGGTTGTTGAGGAGCTTGAGGCAAAGGAAAAGTT _8_21_14_0.10_s CCTTGTCGCGCAGACAGCCCCCTCCATGAGGGTTTCGATAGGGGAG caffold_50598_2, GAGTTCGGCTACAGGCCGGGCGAGATAGTTACGGGGAAGCTTGCGG nucleotide GGGCTTTGAAGGAGCTGGGGTTTGACGCTGTTTTCGACACCTGCACA GGCGCGGACCTTGTGACCATGGAGGAAACCTACGAATTCCTCGAGA GGAAGAAAAAAGGGGAAAGGCTCCCGATAATGACGTCCTGCTGCCC GGGATTTGTGAGCTACATCGAGCATACGCATCCTGAATATGTCGATAA TCTCTGCAGCTGCCGAAGCCCCCAGGAGATTATGGGCGCGCTGATAA AGACATATTTTGCGCAGAAGAGGAAGCTTAAGCCCAAGGACATTTATG TTGTCTCAATAATGCCCTGCATAATCAAGAAAGCCGAGGCCCTGCGC CCGGAATTAAGGGTTAATGGAATGAAGAATGTCGACAAGGTGCTTAC TACAGTTGAGCTCGCGCAGCTCCTCAAGGCAAGGGGCATCGATTTGA AGAAAGTTAAGGAAGCTGATTTTGATTCATTGCTGGGGGAGTCGACC GGAAGCGCGAACATTTTCGGCGCGACCGGCGGCGTCCTTGAAACTG TGCTTAGGCTCGCCGCGAAAATAACCGATAAGAAAGTAGGGGTTATT GAATTCAGGGAGATAAGGGGCATGGAGGGGGTGAAGGAGGCGGAA GTGAGGATAGGCAAAAGCAAAGTTAAGGTTGCGGTGATAAATGGGTT GAGGTATGCCTCGCAGTTATTGAACGACAGGGAAAGAACCAAAAGCT TTGACATCATAGAGTGCATGGCGTGTTTCGGGGGCTGTGTTGGCGGG GCGGGCCAGCCGAGGACAACCCTTGACATAATCGACGCCAGGAAGG AGGCGTTGTATAGGATTGACAGGGGGAAGAAACAGAGAATAGCGTCT GAAAACCCCTCCGTGAAGAAGCTCTATAAGGACTTCCTGGGAAAGCC GGGCAGCGCAAAGGCGAAGAAGCTGCTGCACACCCACTACCACAAA TTCTTTGAATAG 9 CG10_big_fil_rev A1, M1 ATGGCTACAATTGAAGAGGTAAATAAGGCACTTGATGAGAATAAGGTA _8_21_14_0.10_s CAGGTAGTAGCTCAGGTGGCGCCCGCTCTTAGAGTTTCAATAGGGGA caffold_51986_4, GGAATTCGGTTTTGCACCTGGAAAAGTACTTACAAATGAATTTGTTGC nucleotide TGCACTCAAAAAAACTGGTTTTGACAAGGTTTTTGACACTTCAACCGC AGCCGACATTGTTACAATTGAAGAAGGAACAGAGCTACTTAAAAGACT CAAAACAAATGAATCTCTGCCAATGTTTACATCCTGTTGTCCAGGATC TGTTTTGTACATTGAAAATAATTACCCTGAATATGTGAACCATTTCTGT ACAGTAAAGTCGCCACAACAATCAATGGGTGCGCTCATAAAAACTTAT TATGCAAGGCAGATGAACTTACAGAGAAAGGATCTTTTTGTGGTTGCA ATAATGCCCTGTGTGGTGAAAAAAATGGAAGCAAAGCGACCTGAAAT 1005373994
GGAGTTTGATGGAATATCCCATGTTGATGCTGTGCTTACTACAAGAGA AATTGCTGAATTGCTCAAAGGACGAGGCATTTCAATTAAGGAAGGAGA AAAAGCAGATTTTGATAATTTATTGGGAGATGCATCAGGCGCGGGGC AGATTTTTGGCACAACAGGCGGAGTATCAGAAGCACTTATAAGATATG TTTCGGAAAAAACCGGTTCAAACCCAGGAAAAACTGAGTTTAAGGAAT TAAGAGGAACTGCAGGATTCAGAGAAGCTACAGTGAAAATTGCGGGC AAGAATGTTCATATTGCCGTTATTGATGGACTCAACAATCTAAAAAACC TTATAAGCGATCCTGAAAAATTCAATAGCTTTCAAGTAATCGAATTAAT GGTTTGCCCAAACGGGTGTATTGGTGGAGGAGGGCAGCCAAGAACA ACGCCTGAAGGATTGGCTGCAAGAAGGGCAGCATTAATGGGCATTGA TTCAGGCGAAAAAATCAGAGTAGCATCGGACAATGTTGAAGTACAAG AACTTTACAAAAAGTATTTATTAGAACCTGGTTCAAGAACAGCAAAAG CCGTATTGCACACAAACAGAATTTGTCTCAAATGCAATTGA 10 CG10_big_fil_rev A1, M1 ATGGAGGAACTCGAGCGCCTGCTCTCATCGAAGGAAATCGCGGTAGT _8_21_14_0.10_s GCAAACCGCGCCTTCGGTGCGCGTATCGCTCGGCGAGGAATTCGGC caffold_5639_c_1 CTGAAGCCGGGCGCGGACGTGACGGGGCGCGTCGTCGCCAGCCTG 5, nucleotide CGCATGCTCGGTTTCAAATACGTCTTTGACACCGACTTCGGAGCGGA GGTTACCGTGCTCGAGGAGGCGCAGGAGCTTCTCGAGAGGCTTGAA AAGCAGGAGCGGCTCCCGCTCTTTACGTCGTGCTGCCCCGGCTGGA CGCGCTTCTGTCTGAAGCAGTTCCCCTCGCTTGAAAACAATCTTTCGG AGTGCAAGGCGCCGCAGCAGATTCTCGCCTCGCTCGCGAAAACGTAT TTCGCGAAGAAAATCCGCAAGCCCGCAAAGGACGTCGTCGTCGTCTC CGTAATGCCCTGCTTCGCGAAGAAGCTCGAGGCGAAGGAGCGCGAG CATGCCATCGCAGGCGCGCCGCAGGACGTTGACTTGGTTATTACGAC GCACGAGCTCGCGCAGCTGCTCAAGGCAAGGGGTATTGATTTGCGC AGGCTTTCCGACGAGGAGTTCGACAACCCGCTCGGCGTTACTGGCG GCGCCGGCGCGCTTTTCGGGGCAACCGGAGGAATCGCGGAGGCGA TTATGCGCACCGCTTACCATGCCACGACGGGGGGCGCGCTCCCGCG TTTCGAATCCAACATGCGCCCCGCGCTCGGGCAGATCAAGGAGTATT CCATCGTGATGGGCGAGCGC 11 DALH01000010.1 A1, M1 ATGGACGATATTTCATTTGTCAAAGAAGTCCTAGCCGACAAATCCAAG _13, nucleotide ACTGTGATAGCGCAGACCGCGCCCTCGGTGCGCGTAACCGTTTGCG AGGAATTCGGTTCGGAGCCCTCAAGCGAGGAAACGGGCCGCGTGGT CGCCGCCCTCCGCAAACTCGGGTTCGACAACGTGTTCGACACGGATT TCGGCGCGGACGTCACAGTCGTAGAGGAATCCGCGGAACTCGTGCG CCGCCTGAGGGAGGGAGGCGCCCTGCCCGTGTTCACGTCGTGCTGC CCGGGCTGGACGCGCTTCTGCGTGAAGGCGTTCCCCGAACTCAACG AACACCTGTCCGGCGTGAAGTCGCCCCAGCAAATCGTGGGCGCGCT CACGAAAACCTATTTCGCGGAAAAAACCGGGAAGAAGGCGGGCGAA ATATTCGTGGTCGGCGTGATGCCCTGCTACTCCAAGAAACTGGAGTC CAGGGAAGAGGAAACCGAAATTCCGGGAGCAAAACAAGACGTGGAC TTCATTTTAACGACGAAGGACTTCGCCAAACTCATCAAACAGGAGGGA ATCGATTACCACTCCCTGGAACCGGATTCCTTCGACGAACTCCTCGG CACCGCGTCGGGCGCGAGCACCTTGTTCGCGGGCACTGGAGGCGTA ATGGAGTCGGTTGTCCGCGCCGCCTACTACGAGTTGATGGGCGGAA 1005373994
AGGAGTTGCTGCCCTTCGTGGAAGACCGCCGCCCCGGCTACGGCGG CACCAAGGAGTTCACAATTGAAGTGGGCCTGGAAAAACCGTTGCGCG TCGCGGTCGTCAACGGGTACGACGAGGCAAGAAAGATGTGCGAACG AGTGCTCAAGGAAAAAAAGGAGGGCGGGCGCACGATTGACTTCATTG AAGTGCTCGCGTGCCGCGGCGGCTGCGTGGGCGGACCCGGCCAGC CCGGCAACGACCCGCAACAAGTGGCGAAGCGCGCCTTCGGGTTACG CGACTTGGACGCCGAGAACACGCGCGTGCGCAACGCGCACCAGAAC CCGGACGTCAAAAAACTTTACCGCGAATTCCTCGGCGAAACCGGCGG GGAAAAAGCTCATAAACTCCTGCACCGAAAGGGAACCCAGGAGATAT TGGAGACGCTGAAATGA 12 DTGI01000011.1 A1, M1 ATGCTGTACGGCCAGGCTCTATTCTCGGCGATAGAAAAAGAGAAAAG _27, nucleotide AAGCGGAAAAACCCTTGTTGCCCAGCTTGCTCCAGCGGTGCGGGTTT CCATAGGAGAAGAATTCGGCCTTCAGCCAGGAACAGAGCTTACCGCA AAATCCATCTCACTCTTAAAATCCCTTGGCTTCGATGAAGTAATAGAC ACGCCGATTGGAGCCGATTTGATAACATTGGAAGAGGCGAAATTTTTC GTAAAAACAAAACTCGCCTTCAAGCCTCTCTTCAACTCCTGCTGTGTC GGATGGAGAGAGTACTGCAAATCAAACTACAAAAACCTCCTGGGCTA CATAAGCCACATAGTTTCGCCTATGATGGCAGCCGGCTTCATCACAAA AACCTATCTCGCAAACAGGATGGGAGTAGAACCAGAAAAGATCGTTT CTGTAGGTGTGATGCCATGCACAATAAAGAAGCTCGAGACAAAATAC AAAATGCGCTCCGGATTAAAGTATGTTGATTATGTTGTGACAACACAA GAGCTTGGTGAGTGGGCAAGGAGCAAGGGTATGGATATAAAGAACAT GAAGGATGGCGAGTTTTCAAAATTCCTCCCAACCAGTTCAAAGGATG GGACGATCTTCGGTGTGACAGGTGGGATAAGCGAGGCGTTCATAACC ACTGTAGCAAAACTTATGAATGAGGATAGCGAGATCCTTTTTTTCAGA AAAAACGAACCAGTGAGGGAATACGATTTTGCGATAGGAAAAATAAAG CTGAAAACCGCGGTCCTGCACGGAATCGCAAACTTCCCCTCCCTTCT TCCAAGAATAAAAGAATTCAATTTCATAGAGATAATGTTCTGCCCGTAT GGTTGCGTCGGCGGTCCAGGCCAGCCAGCACCGCCAACCGAAGAAA AGTTGAAAGCAAGGGCCGAGGCGTTAAGAAAATTTTCAGATGCAAAA AAACAGAGAACCCCGCTGGAAAACAAACCGCTCATGGAGTCGTTGGA CTGGTACGAGCAGAATAAGATTAATGTTTTTGAGGAGCTTACGTACTG GGGCGAATAA 13 DTNT01000026.1 A1, M1 ATGACTGGTAACAAACTTCTTTATGGTGACGAACTTTTTGCTGAAATC _4, nucleotide GAAAAGGAGATGAATTCCGGGAAGGTAGTTATTGCAGAAATCGCACC TGCGGTTCGCGTCTCTTTGGGCGAACTTTTCGGTTTGGAAGCAGGCA CGAACGTGCAGGGAAAGACGATCGCTCTTTTGCGAAGACTCGGTTTT CACGATGTTGTCGACACTCCGCTCGGCGCTGACATCGCGACCTATTA CGAGGCAGAGGACATAAAGAAAATGCTGGACAGCGGGACAGGCAAA TTCCCGATTTTTAATTCTTGTTGCATTGGTTGGAGGATGTACGCAGGG AAGATGCATCCTGAACTGCTCGATCATATCACTATCGTAGCTTCTCCT CAAATGACCACTGGTTCTGTTGCAAAATACTATTTTGCCGAAAAATTG AACAAGAAACCAGACGAAATCGTGATTGTCGGGATAATGCCGTGCGC TCTGAAAAAATACGAAACAATGGAAGTTATGCGAAACGGCAATCGTTA TGTTGATTATGTTGTTACCACCGTTGAACTCGCCCAGTGGGCAAAGAA 1005373994
GAAGAACATAGATTTTCTGCAACTGCCAGTTGAGTCCTTTTCATCGCT CTTGCCGACAAGTTCAAAGGATGGCATAATCTTCGGCGCGACTGGCG GAGTGACTGAAGCAGTGATAACAACGCTCGGTCGCTTGTATGGGCAG GAGATCACGATCAATGAGTTCCGGGACGATGCAGAGATAAAGAGAAA AACCGTGACTATCGGAAAGCATACTCTCAATATCGGCATCGTTCACG GTCTTCAAAATTTTGAGAAGCTTTATGAAGAAATAAAGGCAGGTAAGA CCTATCACTTGGTCGAAGTTATGATGTGCCCGTTCGGCTGCGTTGGC GGACCCGGGCAGCCCCAGGCATCAAGAGAAAAGATTCAGCAGCGCG CGAGATCCCTAAGACAATTTGCAGATTCGGCAAAGGAAAAGACACCA CTTGACAATCCAACAATGCAGATGCTCTTGCGAGATTTCTTTAGCAAA CTCCCGCGCGACAAATTGGAAGAGTTGATTTATTTTAACCGCTGA 14 DUGC01000052.1 A1, M1 GTGGGCATAGTTGATGACGTTAACGCAGCGCTTGATGACCGCAAGAA _8, nucleotide ATATGCCGTATGCCAGATAGCGCCAAGCGTGCGCGTTTCAATCGGCG AGGAATTCGGTTTTGCGCCGGGGGCTATTGTCACAAAAAAGCTCATC GGCGCCCTTAAAAATGCGGGTTTTAGAAAAGTCCTTGACACTTCGAGT GCGGCAGACATCGTGACAGTTGAGGAAGGGACAGAGCTGCTCGGAA GGCTGGCGGACCGCGAGCAATTGCCGCTCCTGACTTCATGCTGCAG CGCATCAGTGCTCTTCATAGAAAACAGCTACCCCAATTACCTGCCCCA TTTCTGCTCGGTAAAATCCCCGCAGCAATCCATGGGCGCGCTGATAA AGACCCATTATGCAAGGAGGATGAGGCTGCAGCGCAAAAACATCTAC TGCGTTTCCATCATGCCCTGCGTTGTAAAAAAGCTTGAGGCAAAGCG GCCTGAAATGGAATTCAACGGCGTGCCACACACCGATGCGGTGCTCA CGACACGCGAGGCCGCGCAGCTTTTGAAATTGCGCGGGATCGATTTA AAAAGCGCCAGGGGGCAGGAATTTGACCAGATACTGGGAAAGGGCT CGGGGGCAGGCCAGCTGTTTGGAACGACAGGCGGGGTAACGGAGG CGCTCTTGCGCTTTGTATCATGGAAGCTTGAGGGAAAGGGCGCACGC ACCGAATTTAAAGAAGTGCGCGGGGAGGAGGGCTTTCGCGAGGCTA AAGTTGCGATTGCAGGAAAGCCGATCAGGATTGCAATTGTTGACGGC CTGACAAACCTTCGCGACCTGCTCACAAACAGGGAAAAATTCTACAAT TATGACGTTATAGAGATAATGTCGTGCCCGGGAGGGTGCATCGGCGG CGCCGGCCAGCCGGCATCGACAAAAGAAAAGCTTTCAGCGCGAAGG AAGGCACTTTTTGAAATTGACTCGAAAGAGAAGGCTNNCTGA 15 JACCLI01000007 A1, M1 ATGGAAGATGAAATCTACCTCAACAAGCCGGACGAGCTTTTTGCTGC 7.1_25, nucleotide GCTGGAGAAGGAGAAGAAAGACGGAAAGGTGATGGTCGTGCAGATC GCTCCCGCGGTGCGCGTTTCCATCGGCGAGGAGTTCGGGCGCGCTC CTGGCGAGGACCTTACCTACCAGACGGTCGGATTGCTGCATGCGCTC GGGTTCGACCACGTAATGGACACGCCCCTTGGGGCGGATGTGAATAT TTATGAAGAGACTTTGGAGGTGCTGCACGCGCTCGAGCGCGGGGAT GAGAAATATTTCCCGGTGTTCAACTCCTGCTGCATCGGCTGGAGGCT GTACTGCAAGAACAAGCACCCGGAACTTTACCATCTGGTTTCCCCAAT CGGCTCTCCGCACATGGTTGCCGGAAGCCTGGGGAAGCACATACTC GCGAAAAAACTCGGGGTTCCGATAGAAAAAATCTGCATGGTGAGCGT GATGCCGTGCGTGCTCAAGAAATACGAGACGCGCGAGAGGCTTCCC TCCGGGATAAGATACATAGATTATGTGCTGACAACGCACGAGCTCGG GATTTGGGCGAAGAAAAAGGGGCTGGACATGAATAAAGTGAAGGAG 1005373994
GGTAAGTTTACGGAACTCCTGCCGGACAGCTCAAAGGACGGAGTCAT ATTCGGGGCGACCGGAGGAATCACGGAAGCGCTTTTGAGCACGCTC GCGTGCGTGTGCGGGGAGAGCCCTGAGAAGGTCAGGTTCAGGGGC GACGAACAGGTGAAGCACCTGTGCGTGCAGATAGGGAGGCACCGGC TGAATGTTGTTTCCATATATGGGGTAACAAACCTGGACAAGGTCCTTG ACGAGATAAAGCACGGGGTGAAATACCATTTCGTTGAAGTGATGAAC TGCCCTTACGGGTGCGTGGGCGGGCCGGGACAGCCGCTGCCTGCG AGCGAGGAAAAATACAGGGCGCGGGCAGCCGGGCTGAGGAAGGCT GCGGACAGGAAGCCCGGCAAGTGCCCGCTTGGGAAGATGGGGATCT GCGGGGTTTACGAGGCCCTCGGCATAGAGCCGGGAAGCAGGGAGG CGCAGGAGCTTTTCTTCTTCCACAAAACAAACATCTGA 16 JACCLJ01000003 A1, M1 ATGGACATGGAACTTCTTTCAGGAGACGCGCTTTTTGAAAAAATCGAA 7.1_10, nucleotide GGCGAGATGAAATCGGGCAAAATCGTCATTGCCGAAATCGCGCCTGC CGTGCGCGTAACCCTCGGCGAGCTTTTCGGATTTCCCGTCGGTGCGA ACGTGCTCGGGAAGATGACGGCACTCCTGAAAAAGCTCGGCTTTGCG CACGTGGTGGACACTCCGCTTGGAGCCGACATTGCGACATATTACGA AGCAGAGGACATCAAGAGGATGCTCGACAGCGGGAAAGGCAAATTC CCGATATTCAACTCATGCTGCATAGGGTGGAGGCTCTATGCTTCCCG GGCACATCCCGAACTGCTCGGCAACATAACCATAATCGCATCTCCGC AGATGACCATAGGCGCGGTTTCGAAATATTATATTGCCGATAAACTCA AGACCGACCCATCGAATGTCGTCGTCGTTGGGATAATGCCATGCGCG CTTAAGAAATACGAAAGTATGGAAGTCATGCGCAACGGCCACAAATA CATAGATTACGTGGTGACGACGATTGAGCTTGCGCAGTGGGCGAAGA AAAAAGGCCACGACCTGAAGAAGCTGAATGATGAGCCGTTATCCCAG CTCATGCCGCAAAGCTCAAAGGACGGGGTGATGTTCGGCGTAACGG GCGGGATGACAGAAGCGGTTATGACAACGCTTGCGGGGCTTTACGG GGAAAAAAAGGAAATCCTGGATTTCAGGGACGACCTTGAGATGCGGA AAAAGAAGGTCAGGATAGGGAAGCACGTGCTGAAAATAGCCGTGGTT TACGGCTTCCAGAATTTCGAAAAATTATACAAGGAAATAAAAGCTGGC GAGAAATACCATCTTGTCGAGGTCATGATGTGCCCGCTTGGCTGCGT CGGAGGGCCCGGCCAGCCTATTGCGCCAAAGGAAATAGTAACGGCA AGGGGAAATGCCCTAAGAACCGTGGCTGACGGGATAAAGGAAAGGA CCCCGATAGACAATCCGACCGTGCAGAAGCTGATTAAGGAATATCTC GGAAAACTTCCAAGGGAAAAACTCGAAGAGCTGATTTATTTCAACAGA TGA 17 LacPavin_0920_S A1, M1 ATGCCCGCGGAAATTTACCTCAATAAGCCCGAGGAGCTCTTCGCAGC ED3_scaffold_874 CCTTGAAAAGGAGCTTTCCAACCCGGAAAATACAGTGGTTGTGCAGA 915_3, nucleotide TTGCGCCCGCAGTGCGCGTCGCAATCGGAGAAGAATTCGGCTACCC TCCGGGAGTTGACCTCACCTTCAAGACAATCGGCCTCCTTAATGCGC TCGGCTTCAAGCACGTGGTGGACACCCCTCTCGGGGCAGACATAAAC ATCTACGAGGAGGCCTACGAAATCCTGAACGCGCTCGAGCGCAAGG ACGACGCGTATTTCCCGGTTTTCAACTCCTGCTGCATCGGCTGGAGG CTCTACTGCTCCCGCGCGCACCAGAAGGTCTGCCAGCACATCTCCCC GATCGCTTCCCCGCACATGATTACAGGCAGCGTCATAAAGCATTATTT TTCCAAGAAACTGGGTAAGCCCAAGGAGAAGATAATTTCCGTAAGCAT 1005373994
AATGCCCTGCGTGCTCAAGAAATATGAAAGCCTTGAGCGCTTCTCCG ACGGCTCCACCTATATTGATTACGTCGTCACAACCCATGAGCTTGCG CAATGGGCAAAGAAGAAAGGAATTGACCTCCGCAAGGTGAAGGAGG GGAAATTTTCCGAACTCCTCCCCAACTCCTCCAAGGACGGAGTGGTT TTCGGAGTCACCGGCGGCCTCATCGAAGCCCTTCTCACTACGATGGC GCACATTCTCGAAGTAAAGAAGGAAAACGAATCCTTCCGCACCAACG AACCCATAAAGGAGCGGCAGGTCAGAATAGGCGAATACACCCTGAAT GTGGTTTCCATAAACGGCTTTGGGAGCTTCGAGAAAGTCCTGAAGGA AATCGAGTCCGGAAAGAAGAAATTCCACTTCGTGGAGGTGATGAACT GCCCCTATGGGTGCGTTGGCGGCCCTGGGCAGCCCCTTCCGGTTAA CGACGCAATCCTGAAGGCGCGCGCAGAAGGGCTGCGCGCTGCCGC AGACAAAAAACGCTCAATCGCAATCGTCCCGCAGGAAAACCCGACCG TGCAGATGCTCTACCGGGAACTCCTCGAACGGCCAGGCTCGGAGAG GGCCCGCGAGCTGCTCTATTTCCACAAAATAAAAATCTGA 18 LacPavin_0920_S A1, M1 ATGGAAGATGAAATCTACCTCAACAGGCCGGACGAGCTTTTTGCTGC ED4_scaffold_300 GCTGGAAAAGGAGAAGAAATCCGGGAAAGTGATGGTCGTTCAGATCG 9403_2, CTCCCGCGGTCCGGGTTTCCATCGGGGAGGAGTTCGGGCGCGCCCC nucleotide TGGCGAGGACCTTACTTACAAGACGGTCGGGTTGCTGCAGGCGCTC GGGTTCGACCACGTGATGGACACGCCCCTTGGGGCAGACGTTAATAT TTATGAAGAGACTTTGGAGGTGCTGCACGCGCTCGAAAGGGGCGAT GAAAGCTATTTCCCGGTGTTCAACTCCTGCTGCATCGGATGGAGATT GTATTGCAAGAACAAGCACCCGGAACTTTACCATCTGGTTTCCCCAAT CGGCTCTCCGCACATGGTCGCGGGCAGCCTGGGGAAGCACATACTC GCGAAAAAACTCGGGGTTCCGATTGAAAAAATCTGCATGGTCAGCGT GATGCCGTGCGTGCTCAAGAAATACGAGACGCGCGAGATGCTTCCTT CCGGGATAAAATATATGGATTACGTGCTGACGACGCACGAGCTTGGG AATTGGGCGAAGAAAAAGGGGCTGGACATAAATAAAGTGAAGGGCG GGAAGTTTACGGAACTCCTGCCGGAAAGCTCAAAAGACGGAGTAATA TTCGGGGCGACCGGAGGAATCACCGAGGCGTTGCTGAGCACGCTCG CGTGCGTGTGCGGGGAGAGCCCGGAGAAAGTGAGGTTCAGGGGCA ACGAGCAGGTGAAACACCTCTGCGTGCAGATAGGCAGGCATCAGCT CAAAGTGGTTTCCATATATGGGGTAACAAACCTGGACAAGGTGCTCG ACGAGATAAAGCACGGGGTGAAATACCATTTCGTTGAAGTGATGAAC TGCCCCTACGGGTGCGTGGGCGGGCCGGGACAGCCGCTGCCTGCG AGCGAGGAAAGATACCGGGCGCGTGCAGCAGGGCTGAGGAAGGCG GCGGACAGGAGGCCGGGAAAGTGCCCCCTTGGAAAACTGGGCGTGC ACGGAGTGTACGGCGCACTCGGAATAGAGCCCGGGAGCAAGGAGGC GCAGGAGCTTTTCTTCTTCCACAAAACAAACATCTGA 19 Meg19_1012_Bin A1, M1 ATGGATGAGATTGAAAGAGTAAAAGAAGCCTTGGATAAAGAAGGAATA _505_scaffold_65 CAGGTTATTGCACAGGTTGCGCCTGCTTTAAGGGTTACAATTGGAGA 0_7, nucleotide ACTTTTTGGTTTTCCTCCCGGAACAGTTTTAACAAAAAAACTGGTTGG AGCATTAAAAGCTTTAGGAATAGAAAAAGTGTTTGACACTTCTTTAGCA GCAGATGTTGTTACAGTAGAAGAAGGAACTGAATTTATTCACAGACTT GAAAAAAACCAGAATCTGCCTTTATTTACTTCCTGCTGTCCTGCTTCA GTTGCTTTTGTGGAAAACAAGCACAAACAACATGTTAATCATTTTTGTG 1005373994
CAGTTAAATCTCCCCAGCAGACAATGGGTGCTTTGATTAAAAGTTATT ATTCAGAGAAAATGAATCTTTCTCATGAAAAATTTTTTGTTGTTTCAATA ATGCCTTGTGTGGTAAAAAAAATTGAAGCAAAAAGGCCTGAAATGGAT TTCAATGGAATCCATCATGTAGATGCAGTTTTAACTACAATTGAATTGG CAAAGCTTTTGAAAGAAAAAGGAATTGAATTAAAAGACGCAGAAGAAA AAGAATTTGACAGGCTGTTAGGAGACGCTGCAGGCGGTGGACAATTA TTTGGAGTTACAGGCGGAGTTTCTGAATCTCTTTTAAGGTTTGTTTCAA ACAAATTAGACCCGGAAAAAAAGAAGGTTGAATTCACTGAACTAAGAG GAAGTGAGGGAAGAAGAGAAATGGAATTAGAGATTGCAGGGAAAAAA CTAAAAATAGCAATAGTTCATGGCCTGCATAATTTGAATGATTTGATTG AAGATGAAAAAAAGTTTTCTTCTTTTCAGGTAATTGAAATGATGGCTTG TCCTGGCGGATGCATTGGCGGCGGAGGCCAGCCGGTTTCAACTCCA GAGATTAGAGAAAAAAGAATTCAGGGTTTAAGAAGTGTTGATGCAAAA GAAACTGTAAAAATTTCTTCTGACAGCAAAGAAGTCCAGGAATTATAT AAAACCTATTTAGATGAACCCGGTTCAAAAAAAGCCCGCGAATTGTTG CATACAGTTCACATTTGTTTAGAAAAATGCGATTAA 20 NZBD01000002.1 A1, M1 ATGGGAACAATAGAAGACGTAAACAAAGTACTGGACGAGGATAATGT _23, nucleotide TCAGGTTGTAGCACAGGTTGCTCCTGCATTGAGGGTAACAATCGGCG AGGAATTTGGTTACAAACCGGGAACTGTTCTTACAAAACAATTTATTG GATCACTAAAACAAGCGGGTTTTGATAAGGTGTTTGATACATCAACAT CCGCAGATGTTGTTACGGTAGAAGAAGGGACTGAATTTCTGAAAAGA CTTGAAACAGGGGAATGTCTGCCTCTTTTCACATCCTGCTGCCCAGC ATCAGTGCTGTTCATTGAAAGGCAGTTTCCTCAGTATGTCGATCATTT CTGCACTGTAAAATCCCCGCAACAAACAATGGGGGCATTGATCAAAA CATATTATGCTGAAAAAATGAAACTCAAACAGAAAGATCTTTTTGTGGT ATCAGTAATGCCTTGCGTTGTAAAAAAACTCGAAGCAAAAAGACCTGA AATGGAATTCAACAATATTCCGCACGTGGATGCGGTACTAACCACAAA GGAAATTGCTGAATTACTAAAAGGAAGAAGAATTGAACTTGAAAAAGT AAAAGAAGAAAATTTCGATAATTTGTTAGGAGATGCTTCAGGCGCAGG GCAATTGTTCGGAACAACCGGAGGAGTAGCAGAAGCTTTGTTAAGGT TTGTTTCAAATAAACTCGAACAGGATCTTGGAAGGGTTGAATTTAAAG AAGTGAGGGGCATGGAAGGTTTCAGGGAAGCAAATGTAAAAATTGCC GGGAAAAAATTCAAAGTAGCAATAGTTGATGGATTGCATCATTTAAAG GATTTACTAAGTAATGAAAAAAAATTCAAAAAATACCACGCAATAGAAC TAATGGTCTGCCCAAATGGTTGTATTGGCGGAGGAGGCCAGCCAATA TCAACTCCTGAAATAAGAGCAGCAAGAAGAAAAGCATTATTTAAAATC GATTCAAAAGAAAACGTGAGGATTTGTATGGATAATCCTGAAGTAAAA AAATTATACAACGACTATTTGGGGGAACCTGGTTCAAGAAAAGCCGTG TCTTTCTTACACACAAGCAAAGTATGCCTCAAATGTGACTAA 1005373994
21 PEXZ01000074.1 A1, M1 ATGAACGAACTCGAAAAACTGTTTGCCTCGGGGAAGACGGTCGTCGT _10, nucleotide CCAGACCGCGCCGTCGGTGCGCGTCTCGCTCGGCGAGGAATTCGGG CTGAAACCCGGCACCGACGTTACCGGGCGCACCGTAGCCGCGCTGC ACATGCTGGGCTTCAAGTACGTCTTCGACACCGACTTCGGAGCCGAG GCTACGGTGCTGGAGGAGTCGCAGGAGCTTCTGGAGCGGCTTGAGA AACAGGAACGCCTGCCGCTGCTGACGTCGTGCTGTCCCGGCTGGAC GCGCTACTGCCTCAAGATGTTTCCCGGCCTCGAGGGCAACCTTTCCG AAGCCAAGTCGCCGCAGCAGATGCTCGGCGCCATTGCGAAGACGTA TTTCGCGAAGAAAATCCGAAAATCCGCGGGGGACATCGCAGTTGTTT CCATAATGCCCTGCTTCGCGAAGAAGCTCGAAGCGAAGGAGCGCGA GCGCGAGATAGCGGGCGCGACTCAGGACGTGGACCGCGTCATAACG ACGCACGAACTGGCGCAGCTGCTCAAGAGCAAGGGCATAGACCTGC GGCGCCTTCCGGACGAGGCGTTTGACAATCCCCTCGGCGTCGCGGG CGGGGCGGGCGCGCTCTTCGGAGTCACCGGCGGCATAGCGGAGGC GGTGCTGCGCACGGCGTACTACACGGTGACCGAAGGCACGCTTCCC CGTTTCGAATCCAACATGCGCCCCTCCCTAGGGCAGATAAAGGAGCT CTCAATGATGCTCGGCGAGCGCGAGGTGCATCTCGCGATAGTGCAG GGCTACCGGGACGCTTCCGTTATATGCAAGCGCATCCTCGAGGAGA GGCGCAACGGCGTGCGCTCGCTCGATTTCGTCGAGGTGCTGGGATG CATCGGCGGCTGCGTCGGCGGACCCGGGCAGCCCGACACCGTCGA GGGGTCCGTGTTGTCGCGCGCCGCAGCGCTGCACGAGCACGACAG GAAGCTGCCGTTCCGAGACGCGCACAACAACCCCGCGGTAAAGGAG ACGTATGCCGCGTTGTTCGGGAAACCCGGAAGCGGAAAAGCGAAGA AGCTGCTGCATCGGAGCGAGCGCGACTACGAGGAGATAATGGGCAC GTTGACGTGA 22 PFAI01000011.1_ A1, M1 ATGGATGAAATTGAAAGAGTTCATGAAGTGCTGGAATATAAAGAAAAA 3, nucleotide ATTGTTGTCGCTCAGATTGCTCCTGCTTTAAGGGTTTCAATAGGAGAA GAATTTGGTTTTCCTGTGGGAAAAATTTTAACCAAAAAATTTGTTGGTG CATTAAAGCAGGCTGGATTCCAGAAAGTGTTTGACACTTCTGTTGCAG CTGATGTTGTTACAGTAGAAGAAGGAATTGAATTTATTTCCAGATTAGA AAAAAAAGAAAATCTCCCTTTGTTTACTTCTTGTTGTCCTGGTTCAATG GCTTTTATTGAAAACAAACACTCTGAATTGGTTTATCATTTTTGCACTG TTAAATCCCCTCAGCAGACAATGGGTGCATTAATTAAAACATATTATG CAAAAAAAATGAATATTCCTCAAGAAAAAATTTTTGTTGTTTCAATCAT GCCTTGTGCAGTAAAAAAGTTTGAATCCAGAAGACCTGAAATGGATTT TAATGGAATCCATCATGTTGATGCAGTTTTAACCACAAAAGATATTGCA GTATTGTTGAAGAACCTGAAAATTGATTTAAACAATGTAAAAGAAAATG ATTTTGACAGCCTGCTTGAAAATGCATCTGGGGCAGGGCAGATTTTTG GGGTTACAGGAGGAGTATCAGAATCTCTTTTAAGATTTGTTTCACATA AACTTGAACCAGAAAAAAAGAAGATTGAATTCAAAAAATTAAGGGGAA AAGAAGGAAGAAGAGAAATTGAATTAAGTATTGCTGGAAAAAAACTAA AGATTGCAATAGTTCACGGACTACAGAACCTGAATAATTTGATTTTAG ACAAAAACGAGTTTGATTCTTTTCAGGTAATTGAGATGATGGCTTGCC CTGGAGGCTGCATTGGAGGGGGAGGCCAGCCGGTTTCAACTCCTGA GATAAGAGAAAAAAGAATTCAGGCTTTAAGAAAAATTGATTCAAAAGA 1005373994
AAAAGTGAGAGTTTCTTCTGATAATCCTGAAGTAAAAGAACTGTATAAA TCTTTTTTGAAGAACCCTGGTTCAGAGACTGCAAAAGAACTGTTGCAT ACAAGAAGAATTTGTTTTAAATGCGATTAA 23 QZM_A1_scaffold A1, M1 ATGCTCTATGGCCAGGCTCTCTTTTCTGCAATAGAAAGTGAGAAAAAG _68_37, AGCGGGAAAATACTTGTTGCACAGATTGCACCTGCTGTGCGGGTTTC nucleotide CATCGGCGAAGAGTTTGGCCTTCCTCCTGGAACAGAGCTTACAGCAA AAACAATCTCCCTCCTAAAATGCTTAGGTTTTGATGAGGTAATAGATA CCCCGATCGGTGCAGACCTCATAACCCTGGAAGAAGCGAAATTTTTT GTAAAAACAAAACTCGCCTTCAAGCCCCTGCTCAACTCCTGCTGTGTC GGATGGCGGGAGTACTGCAAATCAAACCACAGAAAGCTTCTCAGTTA TATAAGCAACGTGGTTTCCCCAATGATGGCAACTGGCTTCCTCGCAAA AACCTATCTTGCAGGAAAGATGGGGGTTAAACCAGAAAATATACTCTC AGTTGGTGTGATGCCATGCACAATAAAAAAACTTGAGACAAGATACAG GATGCGCTCAGGCTTAAAATATGTTGATTATGTTGTAACCACGCAAGA GCTTGGCCAGTGGGCGAGAGGCAATAACTTGAACATAAAAGATATGG AGGGTAGAGGGTTCTCAAAATTCCTCCCTACAAGCTCAAAGAATGGA ACAATATTTGGTGTAACAGGTGGGATAAGCGAGGCATTCATAACAACC ATGGCAAGGCTCATGGGTGAGGAAAGCGAGCTTGTTTTTTTCAGAAA AAATGAGGCCGTGAGGGAGTATGACTTTGCAATTGGAAAAATAAAACT TAAAACAGCGGTTTTGCATGGGATTGTAAACTTCCCGTCCTTACTACC CAGAATAAATCAATTCAATTTTATAGAGATAATGTTCTGTCCTTATGGT TGTGTTGGCGGGCCCGGCCAGCCAACCCCCCAAACCGAAGAAAAAT TAAAGGCAAGGGCAGAAGCCTTAAGAAAATTTTCAGATGCAAAGAAG GAGAGAACCCCCCTCGACAATAAGCCGCTACTGGAAGCAGTTGGTTG GTTAGAGCAAAAGAATATTAATCTTTTTGATGAAATTACATATTGGGGC GAGTAA 24 QZM_B4_scaffold A1, M1 ATGGAGCTTCTCTCTGGTGACGCTCTTTTTTCAAAACTTGAAGAAGAG _1937_4, ATGAAATCAGGAAAGATTGTGGTCGCCGAGATCGCACCTGCTGTTCG nucleotide AGTAGCTTTAGGCGAGCTGTTCGATTTTCCAGTTGGCACAAATGCTCT TGGGAAAACAGTTTCCCTTCTTAAGAAGCTTGGCTTTCATCATGTTGT TGATACCCCCCTCGGCGCGGATATATCAACCTATTATGAAGCAGAAG ATCTTAAGAGAATGCTCGATACCGGAAAAGGTAAATTTCCGATTTTCA ATTCATGTTGTATTGGTTGGAGAATCTATGCCTCACGCAGCCATCCTG AACTGCTTAACAACATAACAATAATCGCATCACCGCAGATGACTATTG GTGCAGTTTCAAAATACTACCTCGCGCGCAAACTCAGTATCGATCCAG CAAACATAATCGTTGTCGGCATAATGCCCTGCGCACTGAAAAAATACG AGACACTTGAGGTGATGAAGAACGGCCATAGGTATATCGATTATGTCA TCACAACGCTCGAGCTTGCGCAATGGGCAAAGAAGTATGGCTACGAT TTGAAAGAGCTTAAGGACGAGCTGCTTGACCCGCTCATGCCAACAAG CTCAAAGGATGGTGTGATATTCGGCGTTACTGGCGGGATGACCGAAG CGGTTGTTACCACTCTTGCGCAGCTTTACGGAGAAAAAAAGGAAGTG CTCGATTTCAGGGAAGATGAAGGCATAAGGAAGAAGAAGGTTAGGAT AGGTAAGTATGTGCTGAACATTGCAGTGGTGCATGGCTTCCAGAATTT CGAAAAACTATACAGCGGAATAAATGCAGGCGAAAAATATCATCTTGT AGAGGTAATGATGTGCCCCCTCGGCTGCGTCGGTGGGCCTGGCCAG 1005373994
CCTGCGGCCTCGAAAGAGACAATCGCTGCCCGGGGAAGAGCCCTAA GGCAGCTTGCAGACAGCATAAGGGAAAGGACTCCGATAGACAATCCG ACACTGCAGAAGCTCGTGCGCGAATATCTTGGAAAGCTCCCAAGGGA AAAACTCGAAGAGCTGATTTATTTCAACAGATGA 25 S_p2_S4_170907 A1, M1 ATGGGAGCCGTTGAAGAGGTAATAGGGGCGCTGGAGGACGGCAGGA _scaffold_232397 AAGTCGCTGTTTGCCAGGTGGCGCCCTCCGTGAGGGTTTCCATAGGC 0_7, nucleotide GAGGAGTTCGGGATGCCTCCCGGAACCGTGGCCACGGAAAGGCTCG TCGGCGCCCTGAGGCACGCGGGATTCGGCAGGGTATTCGACACCTC GACCGCTGCAGACATAGTGACGATAGAGGAAGGCGCCGAGCTCCTG AGGAGGATGAAGGACAACGAAAGGCTTCCGCTCCTGACATCGTGCTG CAGCGCATCAGTGCTGTACGTGGAAAACAATTTCCCGGAACTCCTCG GCCACTTCTGCACCGTGAAGTCGCCGCAGCAGTCCATGGGGTCACT CGTGAAGACCTACTGCGCGAGGAAAATGCGCGTAAAAAGGCGGGAC ATCTACAGCGTTTCCATAATGCCGTGCATTGTCAAGAAGCTCGAGGC AAGGAGGCCGGAAATGGAGTTCAATGGCGTGAAGCACGTTGACGCC GTCCTCACCACGAAGGAAGCCGCGGAGCTGCTGAAGAGGCGCGGG CTGAGCCTCGCTGACGCCCCTGAATCCGGGTTTGACAGCCTCATGG GAAAGGCATCCGGCTCTGGCCAGCTCTTCGGCACCACCGGAGGAGT CACGGAGGCCCTGCTGAGGTTCGTCTCATGGAAGCTCGATGGCGGG AACGCAAGGATTGATTTTCCTGAGGTGCGCGGCTCTGGCGGATTGAG GGATGTGCGCGTGAAGGCAGGCGGAAGGGAAATAAGGATTGCGGTG GTGGACGGCCTCAACAACCTCAAGGACATACTGAGCAACGCGGACA GGTTCGGGTCATACGACATGATAGAGATAATGACGTGCCCCGGAGG GTGCATAGGCGGCTCGGGCCAGCCGGCATACACGCCACAGGGCCTG CTCGCGAGGAGGGACGCACTCTACGGGATTGACGCCACGGCAAAAG CAAGGACGGCGATGGACAACCCGGCGGTGCAGGCCGTGTACAGGAA TTACCTCCTCGAGCCCGGCTCCGGGGTATCGCAGTCCATACTGCACC TCAAGAGGATATGCCTCAAGTGCAAGTGA 26 SR- A1, M1 ATGAGGATGGAGTCCCTGGATGCCGCTCTCGCGCGGCCGGACGCGG VP_26_10_2020_ TGCTCGTGGCCCAGATGGCCCCGGCAGTCAGGGTCAGTTTGGGCGA 2_100CM_scaffol GATGTTCGGATACCTGCCAGGCACTGTGCTGACGAAAAAGATAGTGG d_1465248_544, GGGCGCTCAAGAAGCTCGGATTCAAGTACGTTTTCGACACCAGCTTC nucleotide GGCGCCGACGTCGCGGTGGTGGAAGAGAGCAAGGAGTTCCAGGAG CGGATAGAGAACGGCGGCGTCCTGCCGATGATAAACTCCTGCTGCC CTGGCACGGTCTCGTTCATGGAGCATTCGTACCCTGAACTGGTCCCG CACATCGGGACCGCGAAAAGCCCGATGGAAATAACGGGAGTCCTGA TAAAGACCTATTTCGCGCAGAAAAAGGGGATAGACCCCGAGAAAATA GTGTCAGTCGCCCTTATGCCGTGCGTGATAAAGAAGGCAGAGGCGCT GCGGCCTGAGCTACGCATGGACGGCAAGCTCGTGGTGGACAGCGTG ATGACCACGGTCGAGCTGGCCCAGGCGCTCAAGGCAAAGGGGATAG AGCTGGGCAAGGCCCCGGAAGCGGACTTCGACCAGCTTCTCGGCAC TGCCTCAGGAGGAGGGCAGATTTTCGGCTCGAGCGGCGGAGTATCG GAATCCGCGCTCAGGAATTTCGCGTTCAATTCCGGCGCGAGCTTTGA AAAGCTGGACACTGCGGCGCTAAGGGGCGCCGAGGGGCTGCGCGA GACCGTGTTCAGCATAGGGGGGAAGAAGCACAAGGTGCTGATAGTC 1005373994
AACTCATTTCGGAACGCCGGGGCAGTTCTCAACGACAAGGCGAAGAT GGGCAAATACTCGTTCATAGAGATTATGGCGTGCCCGGGCGGGTGC GTGGGCGGGGCAGGCCAGCCGCCGTCGAGCAAGGAAGCCATAGAA GCGCGGCGCAAGGGGCTCTATTCCATCGACAGCGGGATGAAAATAA AGGCCGCGGGCCAGAATCCTGACGTGAAACGGCTGTATGATGAATTC CTGGGCGAACCGGGCGGCGAAAAGGCGGTAAAACTGCTGCACACGT CTTTCGTGAAGAGCTGCGAGAACTGCTACTGA 71 Candidatus_Heim B, M3a' GAAAAATGTATACAATGTGGAAAATGTGAGCAAGTATGTCCATATAAT dallarchaeota_arc GCAATTGCTTATCGAGAACGACCGTGCGCGGCAGCATGCGGAGTAAA haeon_RS_5_1|J TGCGATTACATCAGATGCATTAGGATTTGCTGAAATTGATTATAATAAA AGLLF010000953 TGCACTTCGTGTGGTTTGTGTATTGTATCATGCCCCTTTGCCGCTATT _1, nucleotide GCTGAAAAATCCGAAATTGTACAAATTTTGCAAGCACTTCAATCAAAAA AACCCGTTTTTGCAGAGATTGCACCATCGTTCCCGAGTCAATTTGGAC CTTTGGTGAATCCCGAAGTAATTTTAAAGGCCATTACCCAATTAGGAT TTGTAGACGTTGCTGAAGTAGCCTATGGTGCCGATATCACTATCTTAA ATGAAGCTGAAGAACTAAAACAGTTAATTGAAAAATATAAAGGAGATC TCAGCGCTGTCAAAGCGGACAATCCCCAAACAGACCGAACTTTTGTT GGAACGAGTTGTTGCACCGCATGGAGCATAGCAGCTAAAAATAATTTT CCAAAAATTGCTGAGCAAAATATAAGTGAATCTTATACTCCCATGGTT GAAACAGCTAAAAAAATTAAGAAGATTAACCCAGATGCCATTGTGGTT TTTATCGGACCATGCATCGCCAAGAAAGAAGAATGTTTTGATCCATTC GTTCAACAATATGTTGATTTTGTTATGACTTATGAGGAATTGGCCGCG TTATTTCAAGCATATGAAATCGATCCTGCTGAAATTTCCGATTCTTTTC CCATTTCGGATGCTTCTGAATTGGGACGCGGATTTCCTGTTGCAGGT GGTGTTGCAAATGCAGTTGTTAGACAAACTCTCGCGGATCTAGGAAA AGAAGTATCTATTCCTTTAGAATCTGCGGAAACATTGAAAGATTGTAT GGGGATGTTAAAGAAAATCAAACAAAAGAAATATGATCCACTCCCTCT ACTTGTAGAAGGAATGGCCTGCCCTCATGGGTGCGTAGGGGGGCCA GGAACTCTCGCATCTCTTCGACGGGCTCAACGTTCAGCTAAAAAATTT GCCAAAGAAGCGGAATGGAAAAAACCCACCGATTATATTTCAGAAAAT TAG 72 Candidatus_Lokia B, M3a' ATGGAAAAAAAAGATTTGAATATATTTGAGAAAATGCGAGGGATATAT rchaeota_archaeo ACTCCTGTAACAGAAATAAGAAGAAAAGTGTTAGCAGCCGTCGCTAG n_bin106|JAGXO GATGGTTGTAGAAGATCAACCCCCACGATATATTGAATATATTCCATA A010000065_33, TCAAATAATAGATAAAAATATTCCCACCTATCGTGAATCTGTTTTTAAG nucleotide GAAAGAGCTATTGTCCGAGAGAGAATAAGACTCGCTTTTGGAATGGA GTTAAAAGAATTTGGAGCACATGGACCAATTCATGATGATGATGTAAT AAATACAATTACTGATTTAAAAGTGTTAAAGCGGCCTATAGTAAATGTA ATAAAAGCGGGATGTGAACGGTGTCCTGAAGATTCTTATATCGTAACA AATCTCTGTCAGGGTTGTATCGCCCACCCTTGCACGTCTGTTTGTCCA AAAAATGCCGTTTCTATTCAGAAAGGTAAATCATTAATAGACCAATTAA AATGCATTCGGTGTGGAAGATGTGCACAGGTATGTCCATATAATGCAA TAGCATATCGTGAACGTCCTTGTGCTACTGCATGTGGTGTAAAAGCAA TTTCATCTGATGAATATGGATTTGCTGATATCAATTATAATTTATGTGTT TCTTGTGGAATGTGTATCGTTTCATGTCCTTTTGGCGCAATAGGAGAA 1005373994
AAATCTGAAATTGTGCAAATAATTTCAGCAATTAAAAATGGTAAAAGAG TATATGCAGAAATTGCTCCTGCTTTTGTTAATCAATTTGGTCCATTAGC ATCTTCTGCAAAAATACGTAGCGCATTGAAGGAAATTGGTTTTCTTGA TATAAAAGAAGTCGCACTTGGAGCCGATAAAGTTATATTAAAAGAAGC TAAAGAGTTAGTCGCTATACTTGAAAATGACAATAAGAAAAATATCTCT AAGGATGAAAAACAATTTATAGGAACAAGCTGTTGTACATCATGGAAA ATGTGTGCAGATCGTCATTTTCCAGAACTTGCAAAAAATAATATTTCTG AATCATTTGCTCCTATGGTTGAAACTGCTCGAGTTATAAAAGAAAAAG ATCCTGATGCAATAGTTGTTTTTATTGGACCTTGTATTGCTAAAAAAGA AGAATGTTTTATTCCAGAAGTCACCGCCCTAGTGGATTTTGTTATGAC TTTTGAAGAATTAGTAGCAGTATTCCAAGCTTTCGAGATAGATCCAATA GAATTTGAAGAAGAAGAACCAATGCAAGATGCTTCATGTATTGGAAGA AATTTTCCTGTTGCAGGAGGTGTTGCTCAAGCCATAATTAAACAAACT AGGGCATTGCTTCCAGAGAATAAAAAAGATCGTGAGATCCCACATATA AATGCCGATACTCTTGCAAAATGTCTTACAATGCTAAAAAAACTAAAAT CAGGAAAATTTAATCCAAAACCATTGGTTGTTGAAGGAATGGCTTGTC CATTTGGATGTATTGGTGGGCCAGGATCTATATCATCACTTAATAGAG CCAAAAATGCCGTTAAAAAGTTTGCAGAAAAAGCGAAAAATGCTCTCC CTTCAGATTACTTAAAAAAGAAATAA 73 Candidatus_Lokia B, M3a' ATGCCTTTACCATCTTTGAATATTTTTGAAAAAATGCGGGGATTATATA rchaeota_archaeo CCCCTGTAATAAATATTAGGCGTCACGTTCTCTCTGAAGTAGGAAGAA n_FW102|JAIZW TGGTTGTCGAAGGAAGACCTCCACATTACATTGAAGAAATACCTTATC K010000001_123 GAGTAATTCCTAAATCAACACCAACTTATCGTGAATCTGTTTTCAAAGA 9, nucleotide AAGAGCAATTGTTAGAGAGCGGGTACGCTTAGCTTTTGGTATGGATTT AAAGGAATACGGGGCCCATGGACCAATTCATGATGATGAAGTTGTTG AATGCATGAAACCTATTAAAATTCTTAAGCGGCCTATTGTTAATGTGAT TAAAATCGGCTGCGAACGCTGTCCAGAAGAATCTTATTGGGTAACAAA TCTCTGTCGCGGTTGTATTGCTCATCCTTGTGTAACTGTTTGTCCTAA AAATGCAGTTTCAATTCAAAATGGGCGTTCAGTTATCAATCAAGATTTG TGTATTCAATGTGGGAGATGTGCACAAGTATGTCCATATAATGCTATT GCCTTCCGCGAACGACCTTGTGCTTCTGCTTGTGGTGTCAATGCTAT CGGTTCAGATGAAGAAGGCTATGCAAAAATTAATTACGAGAAATGTGT TTCATGCGGTCTTTGCATTGTTTCATGTCCATTTGGTGCAATAGCAGA AAAGTCCGAAATTGCCCAGATTATTTCATCCTTAAAAGCTGGAAAGAG GGTTTATGCTGAAATTGCTCCTGCTTTTGTTAATCAGTTTGGACCATTA GTCACTCCTCCTAAAATTGTTGCTTCTTTGAAAAAAATGGGTTTTAAAG ATGTAGTTGAAGTAGCTGCAGGTGCAGATAGAGTTGTTTTACAAGAAG CTCATGAATTAATAGAACTCATAAAGAAGCGTAAATCATACAGAATAG AAACAAGTAAAAAGAAAAATCAGAATCTACCAGATAATTTAGAAAGTAA TTTGAATATAGATTCGAAAGAAGTAGAAGATTCAGTGATAGATAGTAC TGTAAACAAGACTTTTGTTGGAACAAGCTGTTGCACTTCTTGGACATT GGCTGCAGAAAAGAATTTTCCTGAAATCACCGCTGCAAATATCAGTGA ATCATTTGCACCAATGGTAGAAGCGGCTAAAATAATCAAAGAAAAAGA CCATGAAGCCGTTGTAGTCTTTATCGGCCCTTGTATAGCAAAAAAAGA AGAGTGTTTTATTCCAAAAGTTGCTGAAGTTGTCGATTACGTAATGAC 1005373994
TTTTGAAGAATTAGCAGCAGTTTTTCAAGCTTTTCATATTGATCCCACC AAAGTACCTCCAGATAATGATCTTCAAGATGCATCTTCCCTTGGACGG GGATTTTCAGTTGCAGGAGGTGTTGCTGCAGCAGTAATTAAACAAACA AAAAAGATGCTTGAAGCGGATATTACCATTCCTCATGTATCAGTTGAT ACACTAAAAGAATGCATGGCATTGTTAAAAAAAGTAAAGCAGAATCGA CTTGATCCTGCTCCCCTTTTAGTGGAAGGTATGGCATGTCCTTATGGC TGTATAGGTGGACCTGGAACGTTAGCTCCCTTGAATCGGGCAAAACG CGCTGCCCAGAAATATGCAAAGGAAGCCATCCATGAATATCCATCCG ATTTTTCTTCAGCTGAATCTGATACTTGA 74 Candidatus_Lokia B, M3a' ATGCCGCGCATGTTCGAGCGCTTGAGGGGGCTCTACACGCCCGTCG rchaeota_archaeo TCGAGATAAGGCGCCGTGTTTTCGCGGAACTTGCCCGGCTGGTCGTC n_strain_AS27yjC GAGGGCCGCGCCACGCAAGAAAACATCGAGGCCATTCCCTACCAGG OA_147|JAAZNI0 TAATCAAGGATGATGTTCCTCATTACCGGTGCTGCGTCTTCAAGGAAC 10000050_6, GTGCTATCGTGCGGGAACGCGTGCGGTTAGCTTGTGGCCTCGACCT nucleotide CAAGGAGGGTGGGGAACACGAACCGAAGATCCCAGCCATTTCCAAC GTACTCACGGGCAAGAAGGTCATCTCGGACCCCCTCGTGAACGTGAT CAAGATCGGGTGCGAGCGTTGCCCGACCGATTCGGTCGTCGTCACG GACTTGTGCCGGAATTGCATGGCGCACCCGTGCATCATCGTGTGCCC GGTAAACGCGATTTCCATCGTGGAGGGACGGTCGCACGTGGACACG GGAAAATGCATCAAGTGCTTGCGCTGCACGCAGGTGTGCCCTTACGA GGCCATCGTCCGCCGCGGGAGGCCGTGTGCACAAGCTTGCGGCGTC GACGCTATCGGTTCGGATGCGCAAGGCTATGCTGAAATAGATCATGA CAAATGTGTTTCTTGCGGGCTCTGCACGGTGTCGTGCCCGTTCGGAG CGATCGCTGACAAGAGCGAGATCATGCAAGTGTTGCACGAGCTGCAC CAAAAGGAGCGCGTTCATGCCATCATCGCCCCATCCTTCGTCTCGCA GTTCGGGAACAAGGTATCCCCTCCCGCCATCTTCAGGGGCCTTCAAA AAGCGGGCTTCGCCGGGGTGATGGAGGTCGCCTATGGTGCCGATCA CGATATGATCATCGAGGCGGACAAGCTCGCGAAAATCATCGAGAAGA GCAAGCAAGCTTCAATGAATGCCATTCCGGGCAACCACGAGTTTCTC GGCACGTCATGCTGCCCTTCATGGGTCATGACTGCCCGCAAGTTCTT CCCTGAACTCGCATCGCACATATCCGATTCATACACGCCCATGGTGG TCACGGCCAAGAAGGTGAAGCAAGATGACCCGGGCGCGAAGGTCGT GTTCATCGGGCCATGCGTGGCGAAAAAACACGAGTCCTTGCTCTCGC CCATGAATGAATGGGTCGATCACGTGCTTACCTTCGAAGAATTGGCC GCGATTTTTGTCGCCATGCGTATCGACCTCATGGAGATTCAAGATGG CCTCCCGATCGCCGATGCAAGCCGGTACGGGCGGGGATACGCGGTG GCGGGTGGCGTGTGCAACGCGATCGCCACGTACACCCGTTCGTGTT ACGGATTCCCCGAGTTCGAGGTGAACCGCGCTGATACGCTCCGGGA CTGCCGGAAAATGCTTGCAAGCATCAAAAATGGTGAGATCATGCCTC GACTCGTTGAAGGCATGGCGTGCCCGGACGGATGCATCGGTGGACC AGGCACGCTGGCACCGCTCAAGCTCGCGAAGGTATCGGTAGAAAAAT TTTCCAAGGATGCTGGCAAGAAGGAGCCAAGCTGCGAATCACGTTGA 1005373994
75 Candidatus_Lokia B, M3a' ATGCCTCTGCCCCCTTTGAATATCTTTGAAAAAATGCGAGGACTTTAC rchaeota_archaeo ACCCCAGTGGTGAACATCCGGCGCCACGTGCTCTCAGAAGTTGGCC n_strain_B53_G9| GCATGGTTGTGGAAGAAAAACCTCCTCACTATATTGAAGAAATACCAT QMYW01000085_ ACCGAGTAATTCCAAAATCCACACCCACCTATCGAGAATCGGTTTTTC 2, nucleotide GAGAGCGAGCCATTGTCAGAGAACGGATTAGGCTGGCATTTGGAATG GAACTTAAGGAGTACGGTGCTCACGGACCGATTCATGATGACGAAGA GGCCGAATGCATGAAACCCGTAAAAATACTCACACGTCCCATCGTTAA CGTCATTAAAATCGGCTGTGAACGCTGTCCAGAAGTATCATATTGGGT AACTAATTTATGTCGCGGCTGTATTGCGCATCCATGTGTATCAGTTTG CCCCAAAAACGCTGTCTCTATTGTAAACGACCGCTCGGTCATCGATCA GAATCTTTGTATTCGATGCGGGAGATGCGCTCAAGTTTGCCCCTATAA TGCAATTTCTTATCGGGAAAGACCGTGTTCGGCTGCTTGCGGTGTAA AGGCAATAAGCTCGGATTCAGAGGGGTATGCTCAAATTAATTATGATC GGTGCGTATCATGTGGGCTTTGTGTTGTTTCTTGCCCCTTTGGAGCAA TTGCGGAGAAATCTGAAATTGCACAAATAATTTCCGCTTTGAAATCAA GTCAACATGTTTATGCCGAAATCGCACCTGCCTTTGTTAATCAATTCG GACCGTTAGTTTCACCCGCAAAATTAGTGGTTTCTTTAAAAAAGATGG GATTTCGAGATGTAATTGAGGTAGCTGCGGGAGCAGATCGGGTAGTT TTGCAAGAAGCTCGGGAATTATTGGAGTTAATTCAAGTCAAAAATGGA GAGAAAAACGAGGGTGAAAAAGGGAAATCCCCAAATATAATCAGGAA ATTTGTCGGCACAAGCTGCTGCACAGCGTGGACAATGGCGGCTGAAC GAAATTTCCCGGAATTAAAAAAGCTGAATATCAGCGAGTCTTTTGCTC CAATGGTCGAAGCTGCCAAGCTTATTAAAGCCACAGATCCGGAAGCT ATTGTAGTTTTTATTGGACCTTGCATAGCTAAAAAGGAAGAATGTTTTC AACCGGAAGTAGCGGAAGTGGTTAATTTTGTCATGACTTTTGAAGAAT TGGCAGCCATGTTTCAGTCCTTTTTAATTGATCCGACAAAAATAGAAC CGCAGAAAGATCTTCATGATGCTTCAACATTGGGGCGTGGTTTTCCG GTAGCTGGAGGAGTGGCTGCGGCAGTAATTAAACAGGCACAGAAGAT ATCTAGTGTTCCTCTTGATATCCCTTATGCTGCGGTTGATACCTTGAA GGATTGCATGAGTTTACTGAAGAAATTGGAAGGTAATAAGGTAGTTCC TGCCCCCTTATTAGTGGAAGGAATGGCATGTCCATTTGGATGTGTTG GAGGTCCAGGGACTCTCGCTCCAATAAACCGGGCTCGTCGTGCGGC TCAAAAATACGCTAAGGAAGCGGATCACCAGTATCCTTCAGATATTTT ACTCTAA 1005373994
76 Candidatus_Lokia B, M3a' ATGCCGCCCTTACCCGTGCTAAACATATTTGAAAAAATGCGAGGAATC rchaeota_archaeo TATAGTCCCGTCGTGAATATCCGTCGCCACGTTTTATCCGAGGTTGGA n_strain_HM1_B6 CGGATGGTGGTCGAAGGAAAACCTCCCAGTTATATCGAACAAATTCC _4|JABCSI010000 CTATCATGTTATTCCTTCATCCCTGCCCACATATCGAGAATCTGCGTT 124_6, nucleotide CAAAGAGCGCGCCATCGTTCGGGAACGGGTGCGATTGGCCTTTGGT ATGGATTTGAAGGAGTTTGGTGCCCATGGGCCGATCCACGATGAAGA TGTTGTTTTGTCCCTTCAACCAAATAAAATAATCACACGCCCCATTGTG AATGTAATTAAAATCGGGTGCGAACGATGCCCCGAAGATTCATATTGG GTGACAGATCTCTGCCGTGGCTGTATTGCGCATCCTTGCGTGTCTGT GTGCCCGAAAAACGCTGTTTCAATCATCGAAGGAAAATCGGTCATCG ATCAAGAAAAGTGCATCCTTTGCGGTAAATGTGCCCAGGTATGTCCGT ATAATGCCATTGCCTATCGGGAGCGACCTTGTGCGGTGGCGTGCGG CGTGAATGCCATTTCCTCCGATGCCGATGGATTTGCCGATATTGATTA CGAAAAATGTGTTTCCTGCGGTTTATGCATTGTATCATGCCCCTTTGG TGCGATTGCCGAAAAATCCGAAATTGCCCAAATTATTTCGGCGCTGAA GTCGGAAAAATCGGTGTATGCGGAGATTGCGCCGGCATTTGTGAATC AATTCGGCCCCCTTACCTCACCTGAAAAAATTGTCGCTTCCTTAAAAC AGATGGGATTTAAGGGTGTGCGGGAAGTCGCAGCAGGAGCGGATAA AGTGGTGCTCAATGAAGCGCAGGAACTATTGGAATTGATAAACCAAAA AAAATCCACAAAAAAAGGATCTCCAGCAAAAGGAGAGCGAACTTTTGT GGGGACTTCATGTTGTACAGCCTGGACCTTGGCAGCCCAAGATCATT ATCCCGATTTAGCCCGTTTAAATATCAGCGAATCCTTTGCACCCATGG TGGAAACCGCCCGTTTAATCAAGGCAAAGGATCCCGATGCGGTGGTT GTATTCATTGGGCCCTGTATCGCCAAAAAAGAAGAATGTTTCAGGGAA GAGGTTATTCCCTTCGTGGATTTTGTAATGACTTTCGAGGAGCTGGCT GCATTATTCCAGTCTTTCCTAATAGATCCTACGAAAATCGAACCCGAT GCGGATTTGAGCGATGCTTCTACTCTGGGGCGCGGATTTCCCGTTGC TGGGGGTGTGGCTGCAGCAGTCAAGGCACAAACTTGCAAACTTTTAG GCGATAACTGTGAAATACCTACCCAAGCTGCGGATACACTCAAGGAT TGCCTTGCCATGTTGAAGGATTTATCCCGCAATAAATTTGATCCAGCA CCTCTAATTTGTGAAGGAATGGCTTGTCCCTTTGGATGTATCGGAGGT CCCGGGTCGTTATCGTCCCTTCGAAGGGCTAAAAACGCTTCCAAGAA ATATGCCAAACAAGCGCCCTTTCAATTCCCTTCAGATCATGTGGCCGA TTTGGAGTGA 1005373994
77 Candidatus_Lokia B, M3a' ATGCCCCTACCCAGATTGAATATTTTTGAAAAAATGCGCGGTCTTTAT rchaeota_archaeo ACGCCTGTTATAGACATACGCAGGCGAGTTCTTTCAGCTGTTGCCCG n_strain_HM4_B4 AATGGTTGTTAAAGATTTATCACCAACACATATTGAGCATATTCCTTAT 8|JABCTF010000 CAAATTATAGATAAAGATACTCCAACTTATCGAGAATCAGTTTTTAAAG 130_2, nucleotide AGCGCGCTGTTGTAAGAGAAAGAGTCAGACTTGCTTTTGGAATGGAT TTGAAGGAATTCGGCGCGCATGGTCCAATAAATGACGATGATGTTATT TATTCATTAACCGATAAGAAAGTGATTAACAAACCTATTGTCAATGTTA TTAAAGTTGGATGTGAACGCTGTCCCGAGCATTCTTTTATCGTTACCG ATTTATGCAGGGGTTGTATTGCTCACCCGTGCACGATTGTCTGCCCAA AAAATGCTGTTTCAATAATTAATAATAGATCAGTTATAGATCAAGATCT ATGCATAAGATGTGGTAAATGTGAGCAAGTTTGTCCATATAACGCAAT TGCTTACCGAGAAAGACCATGTGCTGCAGCGTGCGGAGTAAAAGCAA TCAGCTCTGATGAAAACGGTTTTGCGGAAATTGATCAAGATAAATGTG TTTCTTGCGGGATGTGCATTGTATCTTGTCCGTTCGGAGCTATTGCAG AAAAATCAGAGATTGTACAAATAATAACAGCGCTAAAAGGGAAAAAAC CGGTTATAGCAGAAATTGCACCCGCTTTCGTAAGTCAATTTGGACCTT TAGTCACTCCAGGAAAATTAAAAGGGGCCTTAAAGGAGGTCGGTTTTA CTGATGTCCGAGAAGTTGCATATGGGGCTGATGTAGTTGTGATGAAC GAGACTGAAGAGTTGGTAGAACTAATTAAAAAACGGGAGGATTGTGA AGACAATGAGAAATCTGAAGCCAAATGTCGCACTTTCATAGGAACCAG TTGTTGTACCTCGTGGGCAATGGCAGCTGAAAAAAATTTTCCCGAATT ATCTAAGCTTAATATTTCAGAATCATTTGCTCCAATGGTGGAAATTGCT AAAAAGATAAAAAAAGAAACTCCGAATGCTATTGTAGTTTTCATTGGTC CTTGCATTTCTAAAAAAGAAGAATGTTTCATACCAGAAGTTGCTGAGG TTGTTGATTTTGTAATGACTTTTGAGGAATTGGTTGCAATATTCCAAGC TTTTACCATGGATCCTACAAAAATTCCTCCAGAAAAGGAGATGAAAGA TGCTTCTGCTCTAGGTAGGGGCTTTCCCGTTGCAGGTGGTGTGGCGA AAGCAGTTCTTGAACAAACCAAGGCTATAATTGGGGAAGATATTGATA TTCCAATTGTTTCCGCAGATACCTTGAAGGAATGTATGAGCTTACTAC GTCAAGTTAAAAACGGAAAACTCGACCCAAAACCATTGATCGTTGAAG GAATGGCCTGTCCTAATGGTTGTGTAGGGGGACCTGGAACTTTAGCT CCTATTAGAAGGGCTCAGAGAGAAGTAAAAAAATTTGCTAAAAAAGCG AAATGGGAAAAACCTAAGGATTATCTTTAG 78 Candidatus_Lokia B, M3a' ATGCCCGATCTACCGAAATTAACTATTTTTGAAAAGACTCGAGGTTTAT rchaeota_archaeo TCACCCCAGTAAAGGACGTACGACGACAAGTGCTAACCGAAGTAGCA n_strain_MAG_14 AAAATGATAGTGAAAAAAAAACCAGCGTCTTACATTGAAACCATCCCA |DXJH01000338_ TACAACATCATCGCAAAAGATACTCCAACCTACAGAGGATCCGTATTT 6, nucleotide AGAGAACGAGCGATTGTAAGAGAACGGGCTAGATTGGCATTTGGATT AGATTTAAAGGAATTTGGGGCACATGCTCCTATTATCGATGATGTATC TCCAGCGATGACAGATAAGAAATTCCTTGGATTGCCTATCATCAATAT AATTAAAGCAGGTTGCGAACGGTGTGAAACAGATTCATTCTGGGTAAC TAATAATTGCAGTAAATGTTTGGCACATCCATGTTTGATAGCCTGCCC AGTGGACGCTATTACCATCCAAGATCGGGCATTTATTGACCAAGAGAA GTGTATTAAATGCGGAAGGTGTGCACAAGTCTGCCCATATAACGCTAT CGTACATCGTGAACGTCCATGCGCACAAGCATGTGGCGTAGGAGCAA 1005373994
TTACTTCAGATGAAGATGGGTTTGCTGAAATTGATTATGAAAAATGTGT TTCGTGCGGGATGTGTATTATTGCATGCCCCTTTGGTGCAGCTGCCG AGAAATCAGAAATTGTACAAATTATCCACACACTCCGGGGAGATAAAC CAGTATATGCTGAAATTGCACCAAGTTTCGTTGGTCAGTTCGGTCCTT TAGTAAAAGCCCCAATGATCATTGAAGCAATTAAAAAACTGGGATTTA CAGGAGTGGTCGAAGTAGCGTATGGGGCTGATGTTGATATTCTCCTC GAAGCAAAAGAATTGGCAGAAATGATAACTAATAAGGGAACAACCCC GCGTGAATTTATTGGAACGAGTTGTTGCCCAGCGTGGGTACTTGCTG CAATATCCAATTTCCCAAAACAAGCAATCAATATTAGTGATTCATTTAC TCCAATGGTTGAAACGGCTAGAAAGATTAAGAAAATCGAACCAGATGC TCGTGTAGTCTTTATTGGACCCTGTATTGCTAAGAAAAATGAATGTTTT GCTCCCGAGGTTTCTGGGCTTGTGGATTTCGTAATGACGTTTGAGGA GTTAGGGGCACTTTTCCAGGCTTTTGGTGTTGATCCTTCCGAAATGAA AGCAACTGAAGACATAAAGGATGCATCCGAAGTAGGTCGTGGATTCG CTGTAGCAGGTGGAGTAGCAAATGCAATCATTGCTCAAACACAAGAG ATCTTAAAGAAAAAAGTAGATATTCCCCACACTTCTGCGGATACGCTT GCGGATTGTATGCAAATGCTAAAGGATATTCAAAATGGAAAGATAGAT CCAAAACCACTACTCGTAGAGGGAATGGCTTGTCCTTTTGGTTGTATC GGAGGCCCAGGTACCCTCGCACCCTTGCATAGAGCAAAACGGAAAG TGAAAACCTTTTCCAAGCGAGCTAAAGTGAAGTTACCTTCAGATCACT TGAAATAA 79 Candidatus_Lokia B, M3a' ATGCCCGATCTACCAAAATTAACGATTTTTGAAAAAAGTAGAGGTCTC rchaeota_archaeo CCAACTCCGATAAAGGGTATTCGAAGACGAGTCCTTTCCGAAGTCGC n_strain_YT2_012 TAAAATGATCGTAGAGAATAAACTTCCCTCGTATATTGAAACAATTCCC |JAEOTN0100000 TATATTATCATTTCAAAAGATACTCCAACATACCGAGATTCGGTTTTTA 30_11, nucleotide GAGAACGAGCTATCGTTCGTGAACGAGTTCGACTAGCTTTTGGATTA GATTTGCAAGAATTTGGAGCTCATGGACCTATTATTGATGATGTGACT CCCGCAATAACGGATAAAAAGTTCTTGGGATTACCTATAGTCAATGTT ATCAAAGCGGGATGTGAACGCTGTGAAACTGATAGCTTCTGGGTAAC TGATAATTGTAGACGTTGTTTAGCTCATCCATGCACGGTTGTATGTCC CGTAGACGCTGTCTCAATTCAAGAAAATAGAGGTTTAATTGACCAAGA AAAATGTGTAAGGTGTGGAAGATGCGCCCAAGTATGTCCATATAATGC AATTGTTCACCGAGAACGACCGTGTGCACAAGCTTGCGGAATTGGAG CGATTATCACCGACAAAGATGGATTTGCTGATATAGACTATGAAAAAT GTGTTAGTTGTGGAATGTGTTTGATATCATGTCCCTTCGGTGCTATTG GTGAAAAATCTGAAATTGTGCAGATTTTGCACGCTCTTAAAGGAAAGC ACCCAGTTTATGCTGAGATTGCTCCAAGTTTTGTAGGGCAATTTGGAC CCTTAGTTAAAGCATCCATGATTATTGAAGCAATAAAAAAGGTTGGATT TGCAGGAGTTGTAGAGGTAGCGTATGGTGCGGACGTTGGTACGTTAA AAGAGGCACAAGAACTGGTAGATATATGTTGCAATTCCGAAGATCATA AAGATTTCATTGCTACAAGTTGCTGCACGGCTTGGAAACAAGCAGCA GTGAACAATTTTCCAGATTTAGCGAGTAACATCTCTGAATCCTATACTC CCATGGTCGAAGCTGCTCGAAAAATAAAAGCTCAAGAACCTAAATCCC GAGTAGTATTCATAGGTCCTTGTATTGCAAAGAAAAATGAATGTTTTG CCCCTGAAGTCGCAGAACTTGTTGATTTTGTTATGACATATGAGGAAT 1005373994
TGGGAGCAATATTTCAAGCAACTGGAATAGATCCCTCAGAAATGAAAG CAACTAAGGATATTAAGGATGCTAGTGAAGCGGGGAGGGGGTATGCA GTAGCTGGTGGGGTTGCTAATGCGATTGTCTCTCAAACTCAAGAAATA CTAGGAAAAGAAGTTGACATCCCGTATACATCCGCTGATACACTTGCT GATTGTATGAAAATGCTGAACAATATTCAAAAGGGGAAACTTGATCCA AAACCGTTGCTTGTGGAAGGTATGGCTTGTCCTTTTGGTTGTATTGGA GGACCAGGAACACTAGCACCTTTACACAGGGCAAAACGGAAGGTAAA AACGTTTTCAAAACGAGCAAAGGTCAAATTACCTTCTGACTATTTGAAA AAATAA 80 Candidatus_Lokia B, M3a' ATGCCAGATTTACCCAAGCTCTCTATTTTTGAAAAAATGCGTGGAATTT rchaeota_archaeo TTACTCCGGTGAAGAATGTACGTCGCCAAGTTCTCACAGAGGTTGCA n_strain_Zod_Met AGAATGATCGTGGAGGGTCACCCTCCGTCATATATAGAAACAATCCCT abat.1044|JAFGB TATAATATCATATCTAAAGACACCCCAACCTATCGTGAATCGGTGTTTA S010000008_97, GAGAACGAGCAATCGTTCGAGAGAGAGCAAGATTAGCTTTTGGTCTA nucleotide GATTTGAAAGAATTTGGGGCACATGGTCCCATTATTGATGATGTAACT CCAGCTATGACGGATAAAAAATTCTTATCTCGTCCAATCGTCAATGTT ATTAAAGCGGGATGCGAACGATGCGAAACGGACAGTTTCTGGGTTAC AAATAACTGCAGAAAATGTATGGCACACCCCTGTAGTATTGTTTGCCC CGTCGAAGCAGTTACAATCGGAGAGAAGGCTGCGATTATCGATCAGG AGAAATGTATAAGGTGCGGTCGATGTGCACAAGCTTGCCCGTATAAT GCAATTGTCCATCGTGAACGACCTTGTGCTCAAGCATGCGGAGTAAA TGCGATAAGTTCTGATGAAAATGGTTTTGCCGAGATCGATTATGAAAA ATGCGTGAGCTGCGGATTATGCATAGTTAGTTGCCCGTTCGGGGCAA TTGGGGAGAAATCTGAGATTGTTCAGATCATCTATACTCTTAGAGGAG AGAAACCCGTGTACGTGGAGATCGCTCCAAGTTTTGTTGGCCAGTTT GGGCCGTTGGTAAAACCATCTATGATTATTGCAGCCTTAAAGCAAATT GGATTTGCGGGTGTCGTGGAGGTTGCATATGGAGCAGATGTTGATAC TCTTATCTTATCGAGGGAATTAGCAGATTTAGTATCTAAGAAGACTAAA GATAAGCGTAGGTATATAGGTACAAGCTGTTGTCCTGCATGGGTACT GGCAGCAATTACGAATTTTCCTAAACAAGCGATAAACATATCCGAGTC CTTTACTCCAATGGTGGAAGCTGCGCGCAAAATAAAAAAAATGGACC CTGAAGCCAGAGTTGTATTTCTTGGTCCGTGCATTGCAAAAAAGAATG AGTGTTTTACACCAGAAGTTGCTGAATTAGTAGATTTTGTAATGACTTT TGAAGAATTAGGTGCATTATTCCAAGCTTTTAGTATTGATCCCTCTGAA ATGAACGTAGAAGAAGAAATTAGCGATGCTTCAGAAGTAGGGAGGGC TTATGCTGTTGCGGGAGGAGTAGCCGATGCTATTCTTGCACAAACCA AAGAGATCTTAGGAAAAGAAATTGATATTCCTTTTACTTCCGCAGACA CCTTAGCAGATTGCATGATGATGTTAAAAAATATACAAAATAGTAAAAT TTCTCCCAGACCCCTTCTAGTTGAGGGTATGGCGTGCCCTTATGGAT GCATTGGAGGTCCCGGAACTCTATCCCCCCTACATAGGGCGAAACGT AAAGTGGAAACCTTTTCCAAGCGAGCTAAGGTTAAATTACCCTCCCAA CACCTTAAAGAATGA 1005373994
81 Candidatus_Lokia B, M3a' ATGGGAAAAAAGGATTTAAATATTTTCGAAAAGATGCGAGGGATTTAT rchaeota_archaeo ACTCCTGTTACTGAAATAAGGAGGAAGGTTTTGGCAGCAGTTGCCAG n_strain_Zod_Met AATGATAGTTGAAGATCAACCTCCGCGAAATATTGAATATATTCCATAT abat.578|JAFGOA CAAATAATCGATAAAGATATTCCTACATACAGAGAATCTGTTTTTAAAG 010000117_2, AAAGAGCAATTGTGAGAGAAAGAATAAGACTTGCATTTGGAATGGAAT nucleotide TAAAAGAATTTGGAGCACATGGACCAATTCATGATGATGATGTAATAA ATACAATTACAGACTCGAAGATATTAAAGCGACCTATTGTGAATGTGA TTAAAGCGGGATGTGAACGCTGTCCAGAACATTCTTTTATAGTTACTG ATCTCTGTAGAGGATGTATTGCTCATCCTTGCACATCAGTGTGCCCAA AAAATGCTGTTTCTATACAAAAAGGTAAATCATTCATTGACCAATCGAA ATGTATTCGGTGTGGGAGATGTGCTCAAGTTTGTCCATATAATGCAAT AGCCTATCGAGAAAGACCATGTGCTGCAGCATGTGGTGTAAAAGCAA TTACTTCTGATGACTATGGCTTTGCAGATATAAATTACAATTTATGTGT TTCATGTGGTATGTGCATTGTTTCATGTCCATTTGGAGCAATAGGGGA AAAATCTGAGATTGTCCAAATAATAACAGCCATTAAGAAGGGAAAAAG AGTATATGCTGAAATAGCTCCAGCTTTTATTAATCAATTTGGACCTTTA GCCACTCCAGCAAAAATACATAGTGCATTATTAAAAATGGGGTTTCTT GATATAAAAGAAGTCGCTTTAGGAGCAGATAAAGTTGTATTAAAGGAA GCTGAAGAATTAGTTAAACTAATTAAACAAAAAAAAGAACCTCATATCT CCAAAGATCAAAGAAACTTTATAGGTACCAGTTGTTGTACCTCTTGGA AAATGTGTGCTGATCATCATTTTCCTAATCTTGCGAAAAATAATATTTC TGAATCGTTTGCTCCAATGGTTGAGACAGCACGAGTTATAAAAAATAA AGATCCAGAGGGAATAGTTGTGTTTATAGGTCCTTGTATTGCGAAAAA AGAGGAATGTTTTATTCCAGAGGTTAAATCTTTAGTAGATTTTGTTATG ACTTTTGAAGAATTAGTTGCTGTATTTCAAGCTTTTGATATTGATCCTG TAGATTTAACTGAAGAAGAATCTATCAATGATGCTTCAAGCATAGGGA GAAATTTTCCGGTTGCAGGAGGAGTTGCCCAAGCGATTATTCAACAAA CACGTGCTTTACTTCCAGAAAGTGAAAAAAATTGTGATATCACTCATAT AAACGCAGATACTCTGGCAGAATGTCTTAAGATGTTAAAAAAACTAAA ATCAGAAAAATATGATCCTAAACCTTTAGTTGTTGAAGGTATGGCTTGT CCTTTTGGGTGTATAGGTGGACCTGGATCATTATCATCTCTTAATAGA GCAAAGATTGCTGTTAAAAAGTTTGCTGAAGAAGCAGAAAATACACTT CCTTCAGAGTATTTAAAAAAGAAATAG 1005373994
82 NZ_CP042905.1_ B, M3a' ATGCCGTTACCCAGATTAAATATATTTGAAAAAACGCGCGGTCTATAC 15, nucleotide ACGCCGGTAATTGATATTCGAAGGCGGGTTCTTTCAGCTGTTGCTCG AATGGTCGTAGAAAATAAACCACCGACATATATTGAGCATATTCCATA TCAAATTATTGATAAAGATACTCCAACATATCGAGAATCGGTATTCAAG GAACGTGCAGTTGTTAGAGAACGAGTCAGGCTTGCGTTCGGAATGGA TTTAAGGGAATTCGGCGCTCATGGACCGATTAATGATGATGATGTTAT ATATTCCTTAACCGATAAGAAAGTGATTAAAAAGCCTATTGTGAATGTC ATAAAAGTTGGATGTGAACGTTGCCCTGAACATTCTTATATTGTTACTG ATTTATGTAGAGGTTGTATTGCTCACCCTTGCACGATAGTTTGTCCAA AAAATGCGGTTTCGATAATTAATAACAGATCCCATATAGATCAAGATCT TTGTATAAGATGTGGTAAATGTGAGCAAGTGTGTCCATACAACGCAAT TGCATACCGAGAAAGACCATGTGCAGCGGCATGCGGGGTAAAAGCA ATCAGTTCAGATAAAGACGGTTTTGCGGATATCGACCAAAATAAATGT GTTTCTTGCGGAATGTGTATTGTGTCTTGTCCTTTTGGAGCAATTGCA GAAAAATCAGAAATCGTCCAAATAATTTCAGCGCTTAAAGGAAAAAAA CCTGTTGTTGCAGAAATTGCTCCTGCGTTTGTAAGTCAATTTGGACCT TTAGTCACCCCAGGCAAATTAAAAGGTGCTTTAAAAGAAGTCGGTTTT ATGGATGTTCGAGAAGTGGCCTATGGAGCTGATGTAGTTGTGATGAA TGAAACTGAAGAATTGGTCGAATTAATTAAAAAACGGGAAAATTGTAA AGAGAATGAAGAATCTGAAAGTAATTGTCGAACATTCATCGGAACCAG TTGTTGTACCTCATGGGCATTGGCTGCAGAGAAAAATTATCCTGAATT ATCTAAACTAAATATTTCTGAATCATTTGCTCCAATGGTAGAAATTGCA AAAAAGATCAAAAAAGATACTCCAGATGCTATTGTTGTATTCATTGGAC CATGTATATCTAAAAAGGAAGAGTGTTTCATCCCAAATGTGGCAGAAG TCGTGGATTTTGTAATGACTTTTGAAGAACTAGTCGCCATTTTTCAAGC TTTTACAATAGATCCTACAAAAATTCCCCCAGAAAAGGAGATGCAAGA TGCTTCTTCACTGGGAAGAGGTTTTCCCGTTGCTGGTGGAGTAGCAA ATGCAGTACTTGAGCAAACCAAAGCAATTATTGGACAAGATGTTGAGA TTCCAATTGTATCTGCAGATACTTTAAAAAACTGTATGAGTTTATTAAA TCAAATCAAAAAAGGAGCACTCGATCCAAAACCATTGATTGTTGAAGG AATGGCTTGCCCTAATGGGTGCGTGGGAGGACCGGGAACTTTAGCC CCTTTAAGAAGGGCACAACGAGAAGTAAAAAAATTTGCGAAGAAAGC ACAATGGGAAAAACCGAAGGATTATCTGTAG 84 Alum_Rock_MS4 E, M1a ATGTTAATCATGAAAAATGGTGAGAAGGGCATGGAAGAGCTGATTCCT _biofilm_2_scaffol GCGTTTGAGAAAAAAATAAAATTAGTCGCTATGCTCGCTCCTTCTTTT d_236_3, GTAGCGCATTTTGACCATCCTTCAATAATTCTTCAGTTGAAAAAGCTC nucleotide GGATTCGACAAAGTTGTTGAGCTTACATTCGGAGCTAAGATTGTCAAC AAGGAATATCACGAAATTCTAAAAAAGTCAAAAGGATTGGTAATATCTA CTGTCTGCCCAGGAGTTGTAGAAACTATAACCTCGCAATTCCCGCAAT ACAGGAAAAATTTTATAAAAATTGTCTCTCCTATGATAGCAACTGCTAT TATCTGCAGGAAGATATTTCCCAAGCATAAAACTGTTTTTATTTCTCCA TGCAATTTCAAAAGAATAGAAGCTTCCCGCTCTAAATATGTTGATTATG TGATTGGATGCGATGAACTCGATGCGTTATTTGAAAAATACGAGATAA AGCCAGTCAAAACCAAAAGAATGATTCATTTTGACAGGTTTTACAATG ATTACACCAAGATTTATCCTCTTGCAGGAGGGTTGAGCAAAACCGCAC 1005373994
ATGTGAGGCAAATCCTCAAGCCTGAAGAAACAAAAACAATTGACGGA ATTTCAGAACTGATTAGCTTTTTGAAGCATCCAGATAAAAAAGTAAGGT TTCTCGACGGAAACTTCTGCATTGGCGGCTGCATTGGCGGCCCTCTG CTGACAAAAAAGCGTTCGTTAAAAGAAAAGAAAGCCCGTGTCTTAAAA TATGTTCAATGGGCAGTTCATGAGAGAATTCCAAAATCATCAAAAGGG GTTTTCGAGCAGGCTAAAGGCTTGGATTTTTCAGTCAAAAGTTTTTAA 85 CABMEU0100000 E, M1a ATGGAATACGGTTCCCCAGACTTTTTTGGCTTTGCATATGCTGAAGAT 38.1_43, CAGAAAAGATTAGTCAGCCTTTTGCGTGAATCAAAGAAAAAGAATTCA nucleotide AAAAGAAAACTTTGCTTGATGACCGCCCCTTCCTTTGTTGTTGATTTTG ATTTTGTTGATTTTGTTCCAAAAATGAAGGGGCTTGGATTTGACAAGG TGACCGAGCTTACTTTCGGCGCAAAGATTGTAAACCAGCACTATCACA AACACATTAAGGAACACTTCTCTGATTATGCAATGAATGGGAGAAGTA TTGACAAAAAATTCCAGAACAAATTCATTTCTTCTGTTTGCCCCGCGA CCGTAGAACTTGTGAAAAACAGGCACCCTGAGCTTGTAAAATATTTAA TGCCTTTTGTTTCGCCGATGAGCGCGATGGCAAAAATAGTGAAAAAG AATTTTCCAAATCACAAAATAATTTTCCTCGCCCCGTGCAGTGCAAAA AAGTTTGAGGCGGCAAGGCTTCTTGACGGCAAAAATAAGAGAATTATT GATGTTGCGATTACTTTCGCAGAGATGAAACAGATTGTAGCAAAAGAA AATCCAAAAGCAAAAAACGGCTCAAGGAAATTTGATTCATTTTACAAT GATTATACAAAAGTTTATCCTCTTTCAGGGGGACTCTCAAGCACTCTT CACTGTAAGGGAATTCTTCACAAAAATGATGTGATTGTAAGAGATGGC TGCACCGAGCTTGAGAAATTGTTCTCAAAAACACCCGAAAGAGTTTTC TATGACGTTTTATTTTGCAAAGGCGGCTGTGTTGGCGGTCCGGGAAT CGCGTCCCGCGCACCTCTCTTTATTCGAAAACAAAGGGTGTTCAAATA CCGAAACTTCGCGAAAAAAGAAAAGATGGACGGAAAGCAGGGTGTTG ACAAATACACCAAAGGACTTGACTTTTCAAGAGAGTTTTAA 86 CABMGN010000 E, M1a ATGGTGAAGAAAAGTCCTAAATTGAAATTTCCTCTGAAGGGGAGAAAA 008.1_137, AAGTATGTTGTAATGCTTGCTCCAAGTTATATTGTAGACTTCTCTTACC nucleotide CCGAGATAATTTTTGCTTTGAGAAAATTAGGTTTTGATAAAGTTGTTGA GTTAACTTTTGGGGCGAAAATGGTTAATCGAGAGTATCATTCTATCCT GGAACATAACTTAAGTGCTCATGGTTTCTGGATTTCTAGTGTTTGTCC TGGCATTGTTGACCTTGTTTCGACTAGATTTCCACAATATAGAAAAAAT CTGATCCCAGTTGATAGCCCTATGATCGCTATGGCAAAGATTGTTAGA AAAACTTACTCTAAACACGGAATTGTTTTTATCTCTCCATGTAATTTTAA AAAGATTGAGGCTAAAGATAGTGGAGTTGTAGACTATGCTATTGATTA TTCTGAGCTTATGGAAATTTTTAGAAAGAAAAAAATCTCTTTAGAAAGT TTTTCTGATCATGAAAAAGCTCATTTTGATAAATTTTATAATGATTATAC AAAAGTTTACCCTCTTGCGGGAGGTTTATCTAAAACTGCGCGGTTAAA GGGACTTTTGAAAAGAAGAGAAATAAAAAAAATAGATGGTGCTGAGAA GGTAATTGAATTCTTGGAAAATCCTTCTATTAAGACTAAGTTTCTAGAT GCTAATTTCTGTGAAGGTGCTTGTATAGGCGGTCCGTGTATATACTCT AAAAAACTAAGCTTGAGAAAAAGAAGAAGAAAAGTTCTGAAGTATCTT AATCAGTCCAAACGAGAAGAGATCCCCAAAACAGATAAGGGTTTAGTT AAGTGCGCAGAGGGTATAAACTTTAGAAGGTACGATTTGTAG 1005373994
87 CABMGO010000 E, M1a ATGAAAAAAGAGGAAAAAGATCTAGATTCTTTCAAAAAAGCACTAGCT 005.1_287, AGAAAAGAAAAATTAGTTGCCCTACTTGCTCCAAGTTTCATCGCAGAC nucleotide TTTGATTATCCTTTAATAATTTACCAACTAAGAGCACGAGGATTTGACA AAATTGTAGAACTTACTTTTGGCGCAAAAATGATTAATAGAGAATATCA TAAAATTCTCAAAGAACAAAAAGACAAACTTTTCATTTCAAGCGTCTGC CCAGGAGTTGTAGAAACAATTCGAAATAAATACCCCCAATATAGAAAA AATCTTATACCAGTTTTAAGCCCAATGACAGCTACAGCAAAAATATGC AGAAAACTCTATCCTAAACATAAACTAGTTTTCATAAGCCCATGCCAAT TCAAAAAAATTGAAACAAATCAGTCAGAATATGTTGACTTTGCAATAGA CTACAATGAGTTGAGAAAATTATTTGAAGAAAATAAAGAAGAAATAAAC AAAAAAATAAAAATTCTAAAAAAGAAAAAAACAAATAAAATCGGATTTG ACAAATTTTACAATGATTACACAAAAATATATCCCCTAGCGGGAGGAT TAAGTAAAACTGCACACTTAAAGCAAGTTCTAAATCCTGGCGAAGAAG TAAAAATCGATGGAATACTAGAAGTTGAAAAATTCCTTCAAAATCCAGA CAAAAAAGTTAGGTTTATTGATGTCACATTCTGCAAAGGAGGATGTCT TGGCGGCCCTTGTGTATTAAACAAAAATCTTCCAGACAAAAGAAAAAG GCTAATGAAATACTTGAACTATTCCAAAAAAGAAAAAATAGAAGAGAA CAAAAAAGGATTAATAAGAAAAGCAGAAGGAATTAATTTTTCTAAAGTA TATTAA 88 CABMHR0100000 E, M1a ATGAAAAAAAGCAGTGATAAACTTGAATTTCCTCTTAAAGGCAAATATT 01.1_3, nucleotide TAGCAATGCTTGCTCCGAGCTTTGTTGCTGACTTTTCTTATCCCGCAA TAATCTCCCAGCTTAAAGATTTAGGTTTTGACAAAGTTGTTGAACTTAC TTTTGGAGCAAAAATGATTAACAGGGAATATCACAGGATATTGAAAAA TTCAAAGAAACTGATTATCTCAAGCGTCTGCCCTGGAATTGTTGAAAC AATAAAATCAAAAGCTTTAGAATTTAAGGAAAATCTTATTTCTGTTGAC AGCCCTATGACAGCAACTGCAAAAATCTGCAGAAAGATTTATCCGACT CACAAAATTGTCTTTATTTCTCCCTGCAATTTCAAGAAGATAGAGGCG CAACAATCAGAATATGTGGATTATACAATTGATTATCTTGAATTAAAAG AAATTTTTGATAAATTAAAATTAACTGATAAAAAATATTCTAAAAATACA TCTTCATTTGACAAATTTTACAACGATTCAACAAAAATTTATCCTCTTG CTGGAGGATTGTACAAGACTGCTAATCTGAACAATATACTTCAGCAGG ATGAAAGCATTATTATTGACGGGATTGAAGATGTTTTAAAATTCTTGAA CAAACCAAAAAAGAATATAAGATTTCTTGATGTTACTTTCTGCAAAGGC GGGTGCATTGGCGGGCCCTGCACAAATTCAAAATTATCTTTAATGGAA AAGAAAAAGAAAGTTTTGGATTATCTTAAAATAGCAGATAAAGAAAAAA TTCCAAACACAAAGAAAGGAAATATAAAACAAGCAGAAGGAATAAAAT TTAGTTCTTATTATCCAAACAGATAA 1005373994
89 CAIKMX0100000 E, M1a ATGAACAATAAAAACGACATAACAAAAATTGAAAAGGCGCTAAAGAAC 67.1_2, nucleotide AAAAAGCAAAAGATTGCTGCAATGCTTGCTCCGTCATTTGTTTCTGAG TTCAAATATCCTCAAATTATCTCTCAACTAAGAAAACTAGGATTTAACA ACCTCGTAGAGCTAACTTTCGGAGCGAAAATGATTAACCGAGAGTAC CACAAAATTCTTGAAAAATCAAATTCATTGGTAATAGCCTCTGTATGCC CTGGAATTGTCGAATCAATAAAAAACAATCCAGAATTAAAATCATACAA TAAAAACATAATACCGGTCAACTCTCCAATGATTGCAACAGCAAAGAT ATGCAAAAAGATTTATCCAAAACATAAAATCTGTTTTATATCTCCATGT CATTTCAAAAAGATTGAAGCGCAAAACTGCGAATGTGTAGACTATGTA ATAGACTACAATCAACTAAAACAGTTATTCACAAAATATAACATAAAAC CTTCAAAAGAAAAAGATCAGTTTGACAAGTTATACAACGATTACACAAA AATATATCCCTTATCAGGAGGTCTTTCCAAAACAGCCCATCTCAAGGG AATATTAAAAAAAGATCAAACAAAAACAATAGACCGCTGGAAAGATGT TGAAAAATTCCTAAAAGAATACAAAGCAGGAAAAAATAAATCCATCAG ATTCCTTGACGTTACATTCTGCAAAGGCGGTTGCATTGGCGGCCCAT GTACAAATCAAAAGCTCTCAATAGCAAAAAAGAAACAACTTGTTCTGA AATATCTAAAACAAGCCCGCCACGAAGATATTCCAGAATCAAAGAAAG GACTAATAAACAAGGCGAAGGGGATTAGTTTTAGAAACTAA 90 CAIKWG0100000 E, M1a ATGAAAAAAGAGATGAACTCGCTAAACTTCCCCCTAAGAGGAAAATAC 51.1_16, GTAGCCATGCTTGCTCCAAGTTTTATTGTGGATTTTTCATATCCAGAAA nucleotide TAATTCATATTCTTAAAAAATTAGGCTTTGATAAAGTCGTGGAATTAAC TTTCGGGGCTAAAATGATAAACAGAGATTATCACAAAATTCTTTCTAAT TCCAAGGAATTAAAGATTGCAACTGTTTGTCCGGGAGTTGTAGAATTA ATCAAAAATAAAGCTCCACAATATTCTAAGAACTTAATCCAAGTAGACA GCCCGATGATTGCAATGGCTAAAATTTGTAAGAAAGTTTATCCTAAGC ACAAAATAATATTTTTTTCTCCATGTCATTATAAAAAAGATGAATCCAAA AAATCAAAAATGATTTTTAAGGTAATAGATTACAAGGAACTGAAAGAAA TTATTGAAAAAAAGAAAATAAAGATAGATTCAAAAGAAAAAATTCACTT TGATAAATTTTACAATGATTACACAAAAATTTATCCTGTAACTGGAGGG TTATCCAAAACAGCTCACCTAAAAGGCGTAGTTAAGAAGAAAGAAGTT GAATCCATCGACGGAGTAAGCAATATAATCAAGTTTCTAGAGAAACCA AATAAAAGAATAAAATTTTTAGACTGCAATTTTTGTATTGGAGGATGCA TTGGGGGGCCATGCATAAATTCTAAAGAAAAATTAAGGAAAAAAAAGA GAAAAGTTATCAAATATCTTAACCAATCAAAAAAAGAAGACATCCCAG AAACTCGTAAAGGTCTAATTAGGAAAGCAGAAGGCATAACCTTTAAAA AATAA 1005373994
91 CAIPOI01000000 E, M1a ATGAAATCTCTAAACTCTCTCAAGCAAGCACTCGCAAAAAAACAAAAA 2.1_147, ATAGTAGTTATGCTTGCCCCCAGCTTTGTTGCTGATTTTGATTATCCTG nucleotide AAATTATCTATCAACTCAAAGCTTTGGATTTTGATAAAATTACTGAACT TACATTTGGAGCTAAGATGATTAATCGAGAGTATCACAGAATACTTGA AGAAAGTGAGAAGAGCGGTAAAAAAGAGTTATTCATAGCAACAGTTTG TCCGGGGGTTGTTAGTTTTATTAAAACTAAGTATCTTCAATATGCTAAG AATCTTATGAGTGTTGATAGTCCTATGATTGCAACAGCAAAAATATGC AAAAAAATCTATCCGAAGCATAAAGTTTGTTTTATCTCTCCTTGTGAAT TTAAAAAACAAGAATCCCAAGATTCAGAATATGTGGATTTTTGCATAGA TTATAATGAGCTTAGAAAATTAATTAAAGAAAATAAGAAACCTATAAAA AAATCAAAGACAGGATTTGATAAATTTTATAATGATTATACAAAAATATA CCCTCTTGCTGGGGGGCTAAGCAAAACAGCGCATCTTAAAGGAGTTC TAAAGCCAGGAGAAGAAATAAAAATAGATGGAATATTTGAAGTAGAAA AATTCTTGAAAAACCCTAAAAAAGAAGTTAGTTTTATGGATATTACTTT CTGCAAGGGTGGTTGTTTAGGCGGTCCTTGCATCTTAAATAATAATCT TAAGGACAAGAAAGAAAAACTTATGCACTATCTTGAAGTTGCTAAAGA AGAAAAAATACCTAAAGGCAGAGAAGGTTTGGTTGAGAAAGCTAAAG GGATAAGTTTTCTTAAGAGATATTAG 92 CG10_big_fil_rev E, M1a ATGAAATTTAAGTTTGATAAAAAAAAGAAGTATCTGGCAATGCTTGCTC _8_21_14_0.10_s CCAGTTTTGTTGTTGATTTTAATTATCCCGAAGTTATTTCTCAGCTCAG caffold_1732_4, GGAACTGGGATTTGATAAAGTAGTTGAGCTGACATTTGGAGCTAAGAT nucleotide GGTGAACAGGTGCTATCATGAGAAGTTGAAAAATTCTAAAGAGCTTGT TATTGCTAGTGTTTGTCCCGGGATTGTAGAAAGTGTTAGGGAAAGATT TCCAGAGTATGTTCGAAATTTAATTAAGGTTGACAGCCCGATGATTGC AATGGCTAAGATTTGCAGGAAGACATATCCCAATCATAAAATTGTTTTT ATTTCTCCGTGCAATTTTAAGAAAACTGAGGCTAGAAAATCAAAATACA TTGATTTTGTTATTGATTATAAAGAGCTTGCGGTTCTTTTGCAAGTTCA TAAGGTAAGGGGAAAGAAAAAAGATTGCTTTGATAAGTTCTATAATGA ATATACTCGAATTTATCCGATTGCAGGGGGGCTTTCAAAAACTGCGCA TCTTTTGGGGGTTTTGAATAAGAAAGAAGCGAGGGTGATTGATGGGA TATTGGATGTTGAGAAATTTTTAAAGAAGCCTAACAAAAAAATAAGGTT TTTAGATGCAACTTTTTGTAAGGGGGGCTGTATTGGTGGGCAGTGTG TTTCTTCTAAATTAAGCCTTGCTGGAAGAAAGAAAAAGGTTTTAGATTA TTTGAAGTTGGCTTTAAACGAGGAAATTCCGAGAGGCAAGAGGGGAA AGATTAATCGGGCAAAAGGGATTTGTTTTGAAATAAAAACCAGGGCTT CAAAAACTTCGATTTTTGATAGTCATCAGAAACCTAAGGTTTCTGAGAT TTTTGATCCTATAAAATGA 93 CG23_combo_of_ E, M1a ATGAAAAAGAGAGTTGAAAATTTAAAATTTCCTTTAAAGGGGAAGTATC CG06- TTGCAATGTTAGCTCCAAGTTTTGTAGTAGATTTTTCTTATCCTTCTGT 09_8_20_14_all_ AGTTTCACAATTAAAAGAACTTGGTTTTGAAAAAGTTGTTGAACTAACC 150_scaffold_495 TTTGGTGCTAAAATGATTAATAGAGAATATCATAAATTACTCGAAAATT 6_8, nucleotide CAGATAATCTTGTTATATCTTCTGTTTGTCCTGGAATTGTTGATTTTATA AAAAATAAATTTCCCAAATATAAGAAAAATTTAATTCTTGTAGATAGTC CAATGATCGCAATGGCTAAAATATGTAGAAAAACCTATTCTAAACATAA AATAGTTTTTATATCTCCTTGTAATTTTAAGAAGGAAGAAGCAAAAAAT 1005373994
TCTGGTATTGTCGATTATGTAATTGATTACAAGGAATTAAAAAATTTAT TTTTAAAATACAAAATAAATACCAAGAATAAGGAAGTTTGTTTTGACAA GTTTTATAATGATTATACTAAAATATACCCTCTTGCAGGGGGATTATCC AAAACTGCTCATTTAAAGGGAGTTGTAAAGAAAAGCGAAGTGAAAGTT ATAGATGGGATTAATGGAGTTGTTAAATTCCTTGAAAAACCAGATAAG AGTATAAAATTTTTGGATGTTAATCTTTGTGTTGGGGGTTGTATAGGG GGATCATGTATAAATTCTAAAGAAAGCATTCCAGAAAGGAAAAAAAGA GTATTGGGTTACCTTAGGCTTGCTAGAGAAGAGGATATTCCCAATGC GAGACTTGGTTTAATTGAAAAAGCCAAAGATATTAATTTTAATATTAAG AAGTTTTAA 94 CP045477.1_677, E, M1a ATGGAACTTAATACTGAACTTATAGATGTTTTAAAAACAATTAATAAAG nucleotide AAAAAGTTATTTGTCTTTTAGCACCAAGCTTTGTTGTAGATTTTAAATAT CCTAAAATAATTTTAGAACTTAGAAGAATTGGTTTTAATAAAATTGTTG AACTTACATATTCTGCAAAATTAATAAATAAAGAAATACACAAACAAAT ATTAGAAAATAAAAAAAAACAATATATTTGTGGAAACTGCCCTAGTGTT GTTAAATATATTGAAAATAAATATCCAGAACTTAAAAAAAATATTATGG ATATAGCTTCACCTATGGTTATTATGGCAAGATTTATGAAAGAAAAATA TCCTAGTCATATAATTGTTTTTGTAGGACCTTGTTTTTCAAAAAAACAA GAAGCAAAAGAAAACAAAGAAGTTGATTATGCACTTACCTTTAAAGAA ATAAATGATATGTTTTTGTATGCTAAAAAAAATGGTTTTTATAAAAAACA AGAAAATGAAAATAAGCATTTTGATAAATTTTATAATGATTATACAAAG ATATACCCTTTATCTGGTGCAGTTGCAGAAACCATGAATACTAGAGAA ATATTAAAACCAGAAGAAATGATTATTGCAGATGGAATTAAAGAAATTG ATAAGGCAATTATGAAATTTAAAGAAAATAAAAAAGTAAGATTTCTTGA CATGCTTTTCTGTACTGGAGGTTGTGTTGGAGGACCTGGCATAATTTC AACAGAAACTATTGAACAAAGAGAAGACAGAGTTATTCATTATAGAGA TAAATCAAAAAATGAAAAACTTGGAAAAAACTTTGGAAAATTTAAATAT GCAGAAAACCTTAGTTTAAAGAGGAAATAA 95 EastRiver_08_08_ E, M1a ATGATAAAAAACAACATCCCAAAAATAGAGAAAGTTCTTCAAAACAAG 2020_HR_Quigley AAAATAAAGAAGCTTGTAATGGTTGCACCTTCTTTTTTAACAGACTTCA _12455_length_7 ACTATCCTTCTTTGATTTCTCAGCTAAAAGAACTTGGATTCGACAAAGT 330_cov_52.4729 GGTTGAAGTTACATTCGGTGCAAAAATGGTAAACAGAGAGTATCATAG 90_7, nucleotide AATTTTATGCGAAGAAAATCAAAAACTCTGGATAGCAACAACCTGCCC AGGAATAACAGAAACAATAAAAAATAATCCAGAACTAAAAGTTTTTGAA GAAAATCTTATCCCTGTAGACTCTCCTATGGTTGCAATGGCAAAGATT TCAAGGAAAGCCTATCCTAAACACAAAATATTCTTTCTCTCTCCTTGCC ATATGAAGAAAATAGAAGCAGAAAAAACAAAACAAATAGATTTTGTAAT AGATTATCAACAACTCAGAGTCTTATTAGACAAATACAAGATTCTTTCC TCAAATCAACATGTTCAATTTGATAAGTTTTACAATGATTACACCAAAG TCTACCCCTTATCTGGAGGATTAACCAAAACCGCCAAAATAAAAAGCA TACTTAAATGGAGAGAATACAAAATAATAGACGGCTGGCAAAAAGTTG AAAAACTCCTACAAAATATTAAAGATAAACCAAAAAAATACAAAAATCA GAGGTTTTTAGACGTAACTTTTTGTGAAGGGGGGTGTATTGGAGGAC CTTGCACAAACAAAGAACTAAGCATAAGAAAAAAAAGAAAACTAGTCA 1005373994
TAAACTATTTAAAACAAGCAAAAAGAGAAGACATCCCCGAGTCTAAGA AAGGATTAATCAAAAAGGCAGAAGGAATAAGGTTTACCCAACAATAA 96 JAAZKV0100000 E, M1a ATGGTTGAATTGGATTCTTATGGTTATCCTGATTTTTTCGGATTTGCTT 01.1_4, nucleotide ATTCAGAAGATCAATTAAAAACTTTAAGACTTTTGAGAAACTCTTTGAG GTATAATGAGAGTAAAGTTATTTTGATGGTTGCTCCTTCTTTTGTAGTT GATTTTGATTTTAAAAAATTTGTTCCATTAATGAGAGGGCTTGGATTTG ATCTTATCACTGAATTAACTTTTGGGGCAAAAATAGTGAATAAAAATTA TCACAAATACATTAAAGAAAACAAGAAAACAAAAATAAAATTCATTTCA TCTGTTTGTCCATTAAGTGTTAATTTGTTGAAAGCAAAGTATCCTGATT TTTCTAGATTTCTTTTACCTTTTGATAGTCCAATGATTGCGATGGCAAA AGTTTTGAAAAAACACTATCCTAAACACAAAATAGTTTTTGTTTCCCCT TGCTCTGCAAAAAAAATTGAGTCAAAACAATTTAACGAAAAAAACAAAA AAACTCTTATTGATGTGGTAATAACTTTTAGTGAATTAAAACAAATTGT TGCTAAAGAAAAACCAAGATTAGTGGGTTCAAATATTTTTGATTCATTT TACAATGAATACACAAAAGTGTATCCTTTAAGTGGAGGGTTAACAGAA ACACTTCACACTAAACACATTTTGGAAGATGAAGAAATGATTTTTTCTG ACGGGTGTCAAAATTTACAAAAATTGTTTGGTTCAAACCCTGATAAAAT ATTTTACGATATTTTGTTTTGTGAAGGGGGATGCATTGGTGGAAATGG AATTGTTTCCAAAATGCCGATTGTTTGGAAAAAAAACAAAGTGCTCAA ATATAGAAAAGCCGCTTCTAGAATAAAAGAAGGAAAAAAAGTTGGTGT TACTAAATACTATAAAGGAATTAATTTTAGAAGAGAGTTTTGA 97 JAAZNJ01000000 E, M1a ATGGATAAAAAAAGAGATAATAATTTGAAATTTCCTTTAAAAGGAAAAT 4.1_18, nucleotide ACATTGCACTTGTTGCTCCTAGTTTTGTTGTTGATTTTCCTTATCCTAA AATCTTATCTCAATTAAAAGAACTCGGGTTTGATAAAGCTGTTGAATTA ACATTTGGGGCAAAAATTGTTAATAAAGAGTATTATGAGGAGATTAAA AATTCAAAAAAATTAATGATTTCAAGTGTTTGTCCTGGTGTTGTTGAAA CTATAAAAAATTCTTTTCCCGAATATAAAAATAATCTTTTGCTCGTTGAT AGCCCTATGGTAGCAACTGCTAAAATTTGTAAAAAAATATACCCTCAC CATAAAAGAGTTTTCATCTCACCTTGCAACTTTAAAAGATTAGAAGCTA AAAGAACTGGCTTTATTGATTATGTTATTGATTATAAAGAATTAAAAGA TCTTATTTTTAAACACTCTTTTGAATACACAAAAAATAAAAATAAAAAAT CTTATAAAAAAGAAATCTTATTTGATAAATTTTATAATGATTATACTAAG ATATATCCCATCTCCGGAGGATTATCTAAAACATTAAAAGTAAAAAAAC TTTTACAAAAAAATGAAATAAAAGAAATTGATGGAATAAAAAAAGTTAT TAAATTTCTCAAAAATCCTAATCCTAAAATAAAATTTTTAGATATAACTT TTTGTAAAGGCGGGTGTATTGGGAGTCAATTTATTAATTCTAAAATCC CAATATCTCTCAGAAAAATAAAAGTTTTAAAATATATAAGAAAAGCAAA TAAAGAAAGAATCCCTGAAAATAGAAAAGGGATATTTAAAAAAGCAGA AGGTTTATCCTTTAAATCAAATTATCCTATAAATATTTATAGTTTTTAA 1005373994
98 LacPavin_0920_S E, M1a ATGGTCGATCTAAATTTTAGTTATTCCACTGACCAGAAAAAAGTTCTGT ED2_scaffold_597 CGCTTTTGAATGAAAAGCAAAAAGTTTGCTTGATGGCGGCGCCCTCTT 388_197, TTGTTGTCGACTTCGATTATCTTTCTTTTGTTCCCTTGATGAAAGGGCT nucleotide CGGCTTCGACAGGATAACCGAGCTCACTTTCGGCGCAAAAATAGTAA ATGAACACTATCACAAATACATAAAAGAAAATAAAGGAAAGAAAGGCT ACGAGAAATTCATTTCCTCTGTCTGCCCTACTTCCGTTGAAATGGTGA AAAACCGTCACCCTGAATTAAAAAAATTCCTTTTGCCTTTTGATTCTCC AGTGATTTCAATGGCAAAGATACTTCATAAGGAATATCCGAAGCACAA AATCGTTTTTCTTGCCCCCTGCTCCGCGAAAAAAATTGAAGCAAAAAA TGGCAGGTTGATTTCAGCAGCGCTTACTTTTAGGGAAATGAAAGGAAT TATTGAAAAAGAAAAGCCAAAAAAGTCCGGTCGCTCGCATCTGTTCGA CCGCTTCTACAACGATTACACTAAAATTTACCCGCTCTCCGGGGGACT CGGTAAAACACTGCATTCAAAAGATATTTTGAAAGAGGGGGAGACTGT TTCCCGCGATGGTTGCGCAGATTTGCTAAAACTATTTGAAACGCATTC CGACAAAATTTTTTATGACATTTTGTTCTGCAAAGGCGGCTGCATCGG CGGCAACGGCGTCGCCTCAAAGCTTCCGCTCTTCCTTCGGAAGAAAA AAGTTTTAGATTACAAGAAATTCGCCGACAGGGAAAAAATTCCAGAAA AAATGGTTGGCTTGAATAGATATACCAAAGGATTAAGCTTTGAAACGG CATTTTAG 99 LacPavin_0920_S E, M1a ATGAAAAGAAAAAAAGAGAAGAAACTTGCTATGTTGGCTCCGAGTTTT ED3_scaffold_172 GCCAGCGAGTTTGACTATCCCGAGATTATTGGGATGTTGAAGAACCT 2076_9, AGGTTTTGATAAAGTAGTTGAATTGACATTTGGGGCCAAAATGGTTAA nucleotide TAGGGAGTATCATAAACTTCTTGAGAACTCAAAAGAGCTTGTTATTAC AAGTGTCTGCCCTGGAATTGTTTCTTTAATTGAAGGAAAGTTCTCCAA ATACAAGAAGAACCTAGCTAAAATAGATAGCCCAATGATTGCAACAGC AAAAATATGCAAAAAGGTTTTTCCTAAACATAGACTTATTTTTATTTCTC CTTGTAATTTTAAGAAAATTGAAGCTAAGAAAAGTAAACTTATTGATGG GGTTATTGATTACCAAGAATTAAAACAGATTTTTGATAAAAAAAGGATA AAACCCAAAAAAGGAATGTGGAAATTTGATAGGTTTTATAATGATTATA CCAAAGTTTATCCTTTACCCGGGGGTTTGTCCAGAACAGCCAACTTGA ATGGAATTGTGGGGTTTAATGAATGCATGATTATTGATGGAGCAAAAG AAGTTGAAAACTTTCTTAATAAGCCCGATAAATCAATAAAGTTCTTAGA TGTTACTTTTTGCAAAGGGGGGTGTATTGGCGGGCCTTTTTTATCTAA AACTAAAAGCTTGGGAGAAAAGCAAAGGGGGGTTTTGAAATATATTAA TCTTGCAAAAAAAGAAAGAATCTCAAGAGGAAGCAAGGGCGAAATTAA AGAAGCTAATGGTATAGATTTTAGGAAATAG 1005373994
100 LC_01_combined E, M1a ATGAAAAGGGATGATAAAAAATTGAAGTTTCCTCTTAAAGGGAAATAT _scaffold_270283 CTTGCCATGGTTGCTCCTAGTTTTGTAGTTGATTTTCCTTATTCTAAAA 4_31, nucleotide TTCTTTTCCAGCTTAAAGATTTGGGTTTTGATAAGACAGTTGAGTTGAC CTTTGGAGCGAAAATGGTTAATCGAGAATATCATAAAATATTAGAAAAT TCAAAAGGACTTGTTATTTCAAGTGTTTGCCCTGGAATTGTTGAGACA ATAAAATCAAAATATCCTCAATATAAAAATAATTTAATTCCTGTTGACAG CCCGATGACAGCGATGGCGAAAATATGCAAAAAAGTTTATCCTTCTTA CAAAATAGTTTTTATTGCTCCCTGCAATTTTAAAAAAATAGAAGCAAAG AGCTCAAAAAATATTGATTATGTCCTTGATTATTTTGAACTTAAAGAAA TCTTAGAAAAAAATAAATTAAATAAAAAGAAATATGGAAAAAAAGAAAT TTTCTTTGACAAATTTTACAATGATTATACTAAGATTTATCCGCTTTCG GGTGGATTGTCTAAAACTGCGAATTTGAAAAAAATTTTGAAATCTGAC GAGACAAAAATAATTGACGGAATTTTAGAAGTTGAGAAATTTTTGGAA AATCCTGATAAAAAAATAAGATTTCTTGATGTAACTTTCTGCGAAGGC GGATGCATTGGAGGACCTTTAGTGATTTCAAAACTGCCAATATTTTTA AGAAAAAGAAAAGTTTTGAATTATCTTAAAACCGCAGATAAAGAAAAAA TTCCACTTAAGAGAAAGGGTATAGTAAAAGAAGCGGAAGGCATTAGTT TTAGGTCTGAATACCCCAAACAAATAATTTATATATAA 101 Meg19_1012_Bin E, M1a ATGAAGAAAAATAATCTCCCTGAAATTGAGCGAGATTTGAAATCGAAA _278_scaffold_57 AATGTAAAATTTCTTGCGATGGTTGCGCCGAGTTTTGTTGCGGAATTT 52_32, nucleotide AATTATCCTTCAATTGTTTACAGATTAAAAGAACTGGGTTTTGATAAAG CAACTGAGCTGACTTTTGGAGCAAAAATGATTAATAGAGATTACCATA GAAAACTAAAAAATTCCAAGAAATTAGTGATTGCTTCGCCATGTCCTG GAATTGTCATGACGATAAAAAATAAGTATCCAAAATATTTTAAAAATTT AATTAGAACCGATAGTCCTGTTGTTGCTACTGGAAAAATTTGCAGAAA ACACTATCCACAGCACAAACTTGTCTTTATTTCTCCGTGCGATTTTAAA AAAATCGAAGCTGAAAACTCTGAGTATATTGATTATGTTATTGACTACA AACAATTAAGAAAATTATTCAGAAAATATAATGTTAAACCGAAAAAGTG CGAAATTTTGTTTGACAAATTTTACAATGATTATACAAAAATTTATCCCT TGTCTGGAGGGTTAGGTAAAACAGCACATTTGAAAGGAGTGGTTAAA GAAGAAGAGATTTTGGGTATTGACGGAATTAGAAAAGTTATGAAATTT TTAGATAATCCAGATCCTAAAATAAAATTTTTAGATGTTCTGTATTGTG TCGGCGGGTGTATTGGCGGGCAGCATACTTCAAAGAAACTGACTGTT GCTCAAAAAAGAAAAAAAGTTTTAGATTATCTGAATTTTTCAAAATCAG AGGATATTCCTGAAGACAGAAAGGGATTGATAAAAAAGGCAGAAGGG ATAAAATTTTCTGGAAAGTGTTGGTTTGATTAA 1005373994
102 Meg19_1012_Bin E, M1a ATGAATGTGTTTGATGGAATTGAAAAAAAAGTGGGTGTGCAGAATTTG _396_scaffold_36 AACTTTCCTTTAAAGGGTAAATATGTTGCAATGCTTGCTCCTAGTTTTG 153_8, nucleotide TTGTTGATTTTGAATATCCTTCAATTATTTCCAGATTAAAGGCTCTGGG ATTTGATAAAGTAGTTGAGTTGACTTTTGGAGCAAAAATGGTTAACCG AGAATATCAAAAACAACTTAAAAAATCTAAAAAATTATTGATTGCGAGC CCTTGTCCGGGAATTGTTGAAATAATAAAGCAAAAAGCACCAAAATAT GCAAAAAACCTTGCTCAAATAGATAGTCCTGTTACTGCAACAGGAAAA ATATGCAGAAAAATTTATCCTAATCACAAATTGGTTTTTATTTCTCCTTG TCATTATAAAAAATTAGAAGTTGCTAATTCAAAATATATTGATTATACAA TTGATTATAAACAATTAAATAAATTATTTGTGAAATTTAAAGTTCCTAAA TTTAAGACTAAATCTCATTTTGATAAGTTTTATAATGATTATACAAAAAT TTATCCTGTTTCAGGCGGGTTAGGAAAGACAGCTCACTTAAAGGGGG TTATTAAGCAAGATGAAATTTTAGTTATGGATGGATTAAATAATATTTT GAAATTTTTACAAAATCCAGATCAGAAAAAAAGATTTTTAGATATTTTAT TTTGTAAGGGGGGATGTATTGGGGGGCCTTGTATAAGCTGTAAATTG AGTATTCCTGCTCGTAGAAAAAAAGTCTTAGATTATTTAGAAAAATCAA AAGACGAAGATATCCCTGATGCTAAAAAAGGAGTTATTGATAAGGCAA AAGGAATTAATTTTTTGAGAAAAGTTTAA 103 MFWP01000006. E, M1a ATGAAAAAGAGTGCTGAAAATCTGAGTTTTCCTTTAAAAGAGAAAAGC 1_1, nucleotide GTAGCAATGCTTGCCCCAAGTTTTGTTGTTGATTTTTCTTACCCCAATA TAATTTCACAGCTTAAGGGTCTGGGTTTTGACAAAGTTGTTGAGCTTA CTTTTGGAGCCAAAATGATAAACAGAGAATATCATAAAATACTTGAACA TTCAAAAGAACTTGTTATCTCAAGCGTCTGCCCTGGGATTGTTGAAAC AATAAAATCAAAATATCCTCAATATAGAAAGAATCTTATTCCAATAGAC AGCCCCATGATAGCGATGGCAAAAATCTGCAGGAAAATTTATCCTAAG CATAAGGTTATTTTTATCTCTCCCTGCAATTTCAAGAAGATAGAAGCAG AAAATTCTGATTATGTTGATTATACTATTGATTACAAGGAACTCAAAAA TCTTTTAAAGAAACATAAAAAAAAGAAGAATAATTCAGAAACTTTTGAT AAATTTTATAATGATTATACAAAAATTTATCCTCTTGCAGGAGGATTAT ACAAAACAGCCCATCTGAAGAATATCCTAAAAGATGATGAAGCAATGG TTATTGACGGGATAGACAGGGTGATGGAATTTCTTGACAATCCTGACC CTAAAATAAAGTTTCTAGATGTTAATTTCTGCAAAGGGGGGTGCATTG GGGGACCGTGCATAAACTCAAAACTTCCATTAATACTAAAGAAAAAGA AAGTCCTTCATTATTTAAAATTGGCAGAAAAGGAAAAAATTCCCGAAG AAAGCAAGGGAATCATAAAAGAAGCAAAGGGCATTTCTTTTAAGTCTG ATTATTTAAATAAATAA 104 MWBC01000030. E, M1a ATGAGTTTTGGTTTACCCAATTTTTTTGGATTCGCTTATTCAGAGGATC 1_8, nucleotide AATTACTTGTTCTTAGATTACTAAAAGAATCAAAAAAAAGAGGAGGGA GTAAAGTTATTTTGATGAGTGCTCCTGCTTTTGTAGTTGATTTTGATTA CAAGGATTTTTGTCCATTAATGAAGGGGCTTGGTTTTGATAAAGTAAC TGAATTAACTTTTGGAGCAAAGATAGTGAACACTTGTTATAGAAAATAT ATTAAAGAAAACAAGGATAAGCAAGAAAAATTTATTGCAACTGTTTGC CCTTCTTCTGTGGAATTAATAAAAAATAGGTATCCTTATTTGAAAAGGT TCTTACTTCCTTTTGACTCTCCAATGGTTGCGATGGCCAAGGTTTTGA AAAAAAATTATCCTAAGCATAAAATAGTTTTTACTTCTCCTTGTAGTGC 1005373994
AAAAAAAATTGAGGCGAAAAAGGCTTTTTACAAAAAAAAGCAGTTGAT TGATGCGGTAATAACTTTTTCTGAATTGAAACAAATTATTGCTAAAGAG AGGCCAAAGAAGAAAAAAGTTTGTCATAAGTTTGATTCTTTTTATAATG ATTATACAAAAATTTATCCTCTTTCGGGGGGGCTGGGAGCAACTCTTA ACAAAAAAGGTATTTTGAAGGAGAGTGAAGTGGTTTCTACTGATGGAC ATAAGCGACTTTCTAGGATAATGGAAAAGAATTTGAACAAAACTTTTTT TGATGTGTTGTTTTGTGATGGGGGTTGTATTGGGGGTAATGGTGTTAG TTCTAAATTACCAATTGTTTTGAGAAAAAAAAGGGTTTTGGATTATAGG AATCTTTCAAAAAGAGAGAGTATGGATGGTAACCATGGTTTGAATAAG TATTTTAGAGGAATTGATTTTTCAAGAAAATTTGATTGA 105 NJDN01000007.1 E, M1a ATGAAAAAGAGTGGGCAGTTAGTTTTTCCATTGAGAGAAAAATTTGTT _7, nucleotide GCAATGGTTGCTCCAAGTTTTGTTGTGGATTTTCCATATCCTGGGATA ATTCACGGATTGAAAAAGTTGGGATTTGATAAAGTTGTGGAGCTTACA TTTGGAGCAAAACTTGTTAATCAAGAGTACCATAAAGAGTTAAAGAAG GAAGGTTTTTTCATATCTAGTGTTTGTCCGGGAATTGTAAATATTGTGC TTGAGAAGTTTCCAGAATATAAAGACAATCTGTTAAAGGTGGATAGCC CAATGGTTGCCATGGCAAAAATTGTCAGGAAGACTTATCCTTCTCACA AAGTGGTTTTTATCTCCCCTTGTTTTTATAAAAAAGAAGAAGCAAAAAA TTCCGGCCAGGTAGATTTTGTTATTGATTACAGGGAGTTGAAGGTTTT ATTTGATAATAAAAAAATAAATTTGAATTTAAAGAAAAAGATTCATTTTG ACAAATTTTATAATGACTATACAAAAATATATCCTATCGGAGGAGGCCT CTCAAAAACTGCTCACCTAAGAGGTGTTTTGAAGAAGGGGGAAGTAA AAGTTATAGATGGGATTTCAAAGGTGATAAAGTTCTTAGAAAATAGAG ATAAGAAAGTAAGGTTTCTTGATTGCAATTTTTGTGTTGGAGGGTGTA TTGGAGGGCCTTATATAAATTCAAAAGATAGTTTGGCTAAAAGGAAGA AGAGAGTCAGAGATTATCTTAAAAGATCATTAAAAGAAGATATTCCAG AGCCTAGAAAGGGCTTATCAAAGAGGGCGAAGGGCATCAATTTTACA ACCAATTTTTCCCATTAA 106 PNOQ01000025. E, M1a ATGCCAAAAAATAATATAGAACTAATATTGAAAGAATTGAAGCATAAAA 1_23, nucleotide AAATGGTTGCATTAGTTGCTCCGAGTTTTGTTGCGGATTTTGAATATC CTAAAATTATAACGCAATTAGAAATGTTAGGTTTTGATAAAGTTGTAGA ATTAACTTTTGGGGCTAAATTAGTTAATAGAGAGTATCATAGAATATTA AGAACGTCTAAAGAGTTAGTTATAGCTACAGTTTGTCCAGGTATTGTT GAAGTTGTTAATAAAAATTATCCAAAATATAAAAAGAATTTGATTAAAG TTGATAGCCCGATGATAGCAATGGCTAAGATTTGTAAAAAGATTTATC CTAAACATAAAACTGTTTTCTTATCTCCATGTGATTATAAAAAGATAGA AGCAAGTAAATCTAGTTATGTTGATTATGTTATAGATTATGAACAATTG AGAAAAATATTCAAAGATAAAAACTTAGAAAATGTGAAAGAGGGTAAG AAAATATTTGATAGGTTCTATAATGATTATACTAAAATTTATCCTTTAGC TGGTGGACTAAGTAAAACTGCTCATTTGAGTGGTATAATTAAACCTAG TGAGATTAAACATATTGATGGAATAACAAAAGTTATGCAATTTTTGAAA AAACCAGATAAGAAAATTAAATTTTTAGATGTAAATTTTTGTGAAGGCG GTTGTATTGGTGGGCTACATACTTGTAAAATACCAATAGCTCAGAAAA AGAAGAAAGTTATTGCTTACTTAAACAAAGCTAAAAGAGAGAAAATAC 1005373994
CTTTGCCTAGAAGAGGAACTTTTGAAAAGGCCGAAGGCTTGAAGTTTA ATTATTGA 107 PWLL01000017.1 E, M1a ATGGAAATATCAAAAGAACTTGTTGACTTGCTAAAAGTCCTACAAACC _4, nucleotide AAAAGATGTGTTTGTTTGTTAGCACCTAGCTTTGTTGTAGATTTTAAAT ATCCAAAAATAATTAAAACCCTTAAACAACTTGGGTTTTCAAAAGTCTC AGAACTTACTTTCGCTGCAAAAATAATAAATACAGAATACAAAAAACAA CTTAAAAAAACAAAAAAGCCAATCATTTGTACAAACTGTCCATCACTAG TAAAAACTATTGAAAATAAATATCCAGAATATAAAGAATATCTTGCAAA TATTGCATCACCAATGGTGGTTATGGGAAGATTTATTAAAAAACATTTT AAAGAAAAAAACACCTGTGTTTTTGTTGGACCTTGTATAACTAAAAAAA TAGAAGCACTAGAAAATAAAAAAGATATAGACTATGCCATAACCTTTAA AGAACTAAAACAAATGATAGACTATGCAAAAAAAAATAAACTACTTTTA GATACTAAAGAAAACCAAACAACTGATTTTGATAAATATTATAACGACT ATACAAAAATATATCCTTTAGCAGGAGCAGTAGCAGAAACAATGCACG CAAAAGAAATATTGTCTAAAAATCAAACCCTTTGTTGTCAAGGACCAG AAAATATAGAAAAAACTTTAAAAAAAATTAATAAAAACACAAAATTTGTA GACGCTTTGTTCTGTCATGGAGGATGTGTTGGTGGGCCAGGAATAAT TTCTAAAAAAAGTATTAAAAAAAAAGAAAATAAAGTAAGAAAATATAGA AAAGACTGTAAAAAAATAAAAATAGGACAAAATTTAGGTAAACAAAAAT ACGCAACAGAAATAAGTCTTAAAAGAAACTAA 108 RBG_13_scaffold E, M1a ATGAAAAAGAGTGCTGAAAATCTGAGTTTCCCCCTGAAAGAAAAATGC _108_63, GTAGCAATGCTTGCTCCAAGTTTTGTTGTTGATTTTTCTTATCCAAAAA nucleotide TAATTTCACAGCTTAAAAATTTGGGTTTTGATAAAGTTGTTGAATTAAC TTTCGGAGCAAAAATGATTAACAGGGAATACCATGAAGTACTGGAACA TTCAAAAAAACTGGTTATTTCAAGTGTCTGTCCCGGAATTGTTGAAACT ATAAAATCAAAATATCCGCAATATAAAAAGAATCTTATTTCAATAGACA GCCCGATGATAGCTATGGCAAAAATCTGCAGAAAGATTTACCCCCAG CATAAAATTATTTTTATTTCTCCCTGCAATTTCAAAAAAATAGAAGCAG AAAAATCCAATTATGTTGATTATGCTATTGATTATAAGGAACTCAAAAA TATTCTTAAGAAATATAAAAAATTGAACAATAATTCTCAGATTACTTTTG ATAAATTTTACAATGATTATACAAAAATTTATCCTATTGCAGGTGGATT ATACAAAACAGCCCGTCTGAAAAATATTCTGAAAGATGATGAAGCAAT AGTTATTGACGGAATAGACAGAGTCATGAAATTTCTTGACAATCCTGA TTCTAAAATAAAGTTTTTAGATGTTAATTTCTGCAAGGGAGGGTGCATT GGAGGGCCGTGCATAAACTCAAAACTTCCGTTAATACTAAGGAGAAA GAAGGTTCTTGATTATCTAAAATTGGCAGAAAAAGAAAAGATTCCAGA AGAAAGCAAAGGTGTTATCAAACAGGCATCAGGGATTTCATTCAAGTC TGATTATCTAAATAAATAA 1005373994
109 RBG_16_scaffold E, M1a ATGAAAAGAGGTGTGGAGAATCTGAGTTTTCCTCTGAAAGGGGAATAT _3952_10, GTAGCAATGCTTGCACCAAGTTTTGTTGTTGATTTTTCTTATCCAAAGA nucleotide TAATCCTTCGACTTAGGGATTTGGGCTTTGACAAGATTGTTGAGCTGA CTTTTGGAGCAAAGATGATTAACAGGTATTATCACAATAAACTGGAGA AAACAAAAGAATTAGTTATCTCCAGCGTGTGTCCAGGAGTTGTTGAGA CAATAAAAACAAAATTTCCACAATATAAAAAGAACCTGATACAAGTTGA TAGTCCGATGATTGCTACAGCAAAGATATGCAGAAAAATTTACCCACA ACATAAAATTGTTTTTATCTCTCCATGCAATTTCAAGAAAACAGAAGCA GAGAATTCAGAGTATATTGATTATGTCATAGATTACAGAGAACTGGGA GAATTGTTAAAGAGATTGGGGAGAAGACTAGACAAGAAAGATACAGA TTTGTTATTTGATAAATTCTATAATGATTATACTAAGATTTATCCTCTTT CAGGAGGGCTTTCCAAGACAGTCCATCTGAAAGGAATTTTGAGATCA GAAGAGACAAGGGCTATAGATGGAATGGGAGATGTAATTTTATTTTTG AATAACCCTGACAAAAAAATAAAATTTCTTGATGTAACTTTTTGCAAAG GGGGATGCATTGGAGGACCATGCATAAACTCAAAATTACCTTTACTAT TAAGAAAAAGGAAAGTTCTTGATTACATAAAAATAGCAGACAAAGAAG TTATTCCGCAAGGCAGGAAAGGACTTATGAAAGAGGCAAAAGGAATA TCGTTCAAATCTGAGAATGTGAATAAATAA 110 rifoxya1_full_scaff E, M1a ATGAAAAAAAATAATATTAATTTAATTTTAAAAGAATTAAAGCATAAGAA old_175_29, AATGGTTGCGCTAGTTGCTCCAAGTTTTGTTGCTGATTTTGAATATCCT nucleotide AAAATACTCTCTCAGTTAGAAAAATTAGGTTTTGACAAAATTGTTGAGT TAACTTTTGGTGCTAAATTAGTTAATAGAGAGTATCATAAAATATTAAA AAGTTCTAAAGGATTAGTTATAGCTACAGTTTGTCCAGGAATAGTTGA AGTTGTTAATAAAAATTATCCAAAATATAAAAAGAATTTGATTAGAGTT GACAGTCCTATGATTGCTATGGCTAAAATTTGTAAAAAAATTTATCCTA AACATAAAACTGTTTTCTTGTCTCCTTGCGATTATAAAAAAATAGAAGC AAATAAATCAAAATATGTTGATTATGTAATAGATTATGAACAATTAAGA GAAATTTTTAAAGAAAAAAAGTTAAATAATATTAAACCAAGTAAAAAAAT ATTTGATAAGTTCTATAATGATTATACTAAAATTTATCCTTTGGCTGGG GGATTAAGCAAAACTGCGCATTTAAATGATGTTATTAAAAAAAATGAAA TTAAACATATTGATGGAATAAAAAAAGTTATGCAGTTTTTAGATAAACC AGATAAAAAAATTAAATTTTTGGATGTGAATTTTTGTGTTGGAGGTTGT ATTGGTGGGTCACATACTTGTAATTTATCAATTGCTAAAAAGAAGAAG AGAATTATTGCTTATTTAAATAAAGCTAAAAGAGAAAGAATACCTTTAT CAAGAAGAGGAACTTTTGAAAGAGCTGAAGGTTTGAAGTTTACTTACT AA 1005373994
111 S2_GD2017_2_m E, M1a ATGGTTGAACTAAATTCAGAATTAAAAGCAGTGGTTGATGCTTTAAATA anure_scaffold_1 TTGAAAAGACAATTTGTCTTTTGGCACCAAGCTTTCCTGTAGATTTTGA 997_46, GTTCCCGGATATTATTCTAGATCTTAGAAGAATGGGTTTTACTAAAGTT nucleotide GTGGAATTGACTTATGCTGCTAAACTTATTAATTATAAATATATAAACA TAATAAATGAAAACCCCGAAAAACAATTTATTTGCGGTAATTGCCCAA CAATTGTTAAATTAATAGAAAATCAATATCCTGATCTAAAAGACAATAT TCTAGACGTTACTTCACCAATGGTGGTTATGGCGCGTTTTGTTAAAAG AGAGTTTGGGAATGATTATAAAACCATTTTTGTGGGTCCTTGCTTTGC TAAAAAAACAGAAGCCAAAGAAAATTCTGATTGCGTAGACTATGCTCT GACTTTTAAAGAATTAATAGAAATCTATGATTATTGCGAAGAAAAAAAT ATTTTAAAAAATATTGAAGAAAATAAAGCAAACAGGGAATTTGATAAAT TTTATAATGATGTGACAAAAATTTATCCACTTGGCGGCGGCGTTGCAA GTTCTATGATAACAAAAGACATACTTAATCTTGAACAAGTAATTGTATG CGATGGCCCAAAATGTATAGAAAAGGCAATGTTTGATTTTACTAATAAT ACAAAATATAGATTTGTAGATATACTTTTCTGCGAAGGCGGTTGCCTT GGAGGACCAGGCATTGTGTGTAAGGATGATCTTGAAACAAAAAAACA AAGACTATTTGATTACAAAGAAAAATCAAGATACTACGAGCCTAATGA AAACTACGGAAAGTTCGTGCATGCTTTTGGACTGGATATAAAACGAAA GAAATGA 112 SR- E, M1a ATGAAAAAAGAGATGTTAGAGTTTCCATTAAAAAAAGAAGAGAAGTAT VP_26_10_2020_ ATAGCAATGCTTGCTCCAAGTTTTGTAGTTGATTTTTCATATCCTGATA 1_100CM_scaffol TTATTTCTCAATTAAAAGGATTAGGTTTCGATAAGGTTGTTGAATTGAC d_4165154_7, TTTTGGGGCAAAAATGATTAATCGAGATTATCATAAAATACTTGGGAAA nucleotide AGCAAAGGACTGGTTATTTCAAGTGTTTGTCCTGGTGTGGTTGAAACG ATAAAATCAAAATTTCCAGAGTATAAAAATAATCTAATTGAAGTTGATA GTCCTATGATAGCAATGGCAAAGATATGCAAGAAAAACTACCCGCAAC ACAGGATTGTTTTTTTTGCTCCGTGTGATTTTAAGAAAATAGAAGCAG GAAAATCAAAGTATATAGACTATGTTTTTGATTTTACAGAATTGAAAGA GATACTGAAACAAAATAAAAAAAGAACAAATGAAAATTTGAAATTTGAC AGCTTTTATAATGATTATACAAAAATATATCCTGTTTCAGGAGGATTGT CAAAAACAGCAAATCTGAAAGATATAATAGAGTTAAAAGATGTCAAGA TTATTGATGGAATAGAAGAAGTTTCTGAGTTCCTGAAAAATCCAGAAA AAGATATAAAATTTCTTGATGTGACTTTTTGCAAGGGCGGATGCATTG GAGGGCCAAAGATAAATTCAAAACTGCCTATTGTTCTTAGAAAAACAA AAGTGATGAATTATATGAAAGTAGCAGATAAAGAGAGCATTCCTGAAA AAAGAAAAGGAGTAATAGAAAAAGCAAGAGGGATTTCTTTCAAGTCAG AATATCCGGCTTAA 137 AQRS01000037.1 A1, M1 ATGGGTTCGATTGAGGACGTGAATGCCGCGCTCGCGGATGAAGGGA _10, Ia, nucleotide AAATGGTCATGGCGCAGGTAGCGCCCGCGGTTAGGGTTACTATCGG CGAGGAGTTCGGCCTTCCGGCGGGAACAATTGTGACGAAAAAGCTC GTGGGCGCGTTGAGGCAGGCCGGCTTTGAAAAGGTGTTTGACACCT CCGTTGCCGCGGATATTGTAACAATTGAGGAAGGAACGGAATTCCTG AACAGGCTCGAGGACCAGGAGGACCTTCCATTGCTGACTTCCTGCTG CCCTGCATCGGTTTTTTTTGTTGAGAACACTTTCCCGAAATTTTTGCAC CACTTCTGCACTGTTAAAAGCCCGCAGCAGGGCATGGGCTCGCTCAT 1005373994
AAAAACCTATTACGCGCGCAGGATGAAGATTGATGCCAAAAAAAATTT TGTTGTCGCTGTGATGCCGTGCATCGTGAAAAAAATGGAAGCGCGCC GCCCTGAAATGGAGTTCGACGGCGTGCATAACGTGGATGCGGTTCTT ACCACAAAAGAGGCAGCTGCTTTGCTCAAATCAAAGAAGGCCGATTT GAACGGCGCGAATGAAACGGGGTTCGACAGTTTGCTCGGAAAGGCT TCGGGAGCAGGCCAATTGTTCGGGCAAACTGGCGGGGTTTCTGAAG CTTTGCTGAGGTTTGTCGCGTGGAAGCTCGAGGGGAAAAAAGCGAG GGTTCTTTTCAAGGAGGTGCGCGGGAAAAAGGGTTTTCGCGAGGCC GAAGTGAAAATAGGCTCGAGGATGCTGAAGGTTGCGGTCATTGATGG GCTGAACAATTTGCGGGATTTGATGAGCAGCGAGGAAAAATTCCATT CGTATGATGTCGTGGAGATAATGACCTGCCCGGGTGGATGCATCGGC GGAGGCGGCCAGCCGGAATCAACGCCTGAAAAGCTGGAAGCGAGGA AAAAGGCGCTGCATAATGTGGATGCAATGGAAACCGTGAGGGTTGCC TCTGAAAACCCTGAAGTGAAGGAGCTTTACGGCTCATATCTCATTGAG CCGGGCTCGAAGGTCGCGCGCTCAATCCTGCACGTGAGCAGGATCT GCCTCAAGTGCGACTAG 140 DALH01000010.1 A1, M1 ATGGACGATATTTCATTTGTCAAAGAAGTCCTAGCCGACAAATCCAAG _13, Mu, ACTGTGATAGCGCAGACCGCGCCCTCGGTGCGCGTAACCGTTTGCG nucleotide AGGAATTCGGTTCGGAGCCCTCAAGCGAGGAAACGGGCCGCGTGGT CGCCGCCCTCCGCAAACTCGGGTTCGACAACGTGTTCGACACGGATT TCGGCGCGGACGTCACAGTCGTAGAGGAATCCGCGGAACTCGTGCG CCGCCTGAGGGAGGGAGGCGCCCTGCCCGTGTTCACGTCGTGCTGC CCGGGCTGGACGCGCTTCTGCGTGAAGGCGTTCCCCGAACTCAACG AACACCTGTCCGGCGTGAAGTCGCCCCAGCAAATCGTGGGCGCGCT CACGAAAACCTATTTCGCGGAAAAAACCGGGAAGAAGGCGGGCGAA ATATTCGTGGTCGGCGTGATGCCCTGCTACTCCAAGAAACTGGAGTC CAGGGAAGAGGAAACCGAAATTCCGGGAGCAAAACAAGACGTGGAC TTCATTTTAACGACGAAGGACTTCGCCAAACTCATCAAACAGGAGGGA ATCGATTACCACTCCCTGGAACCGGATTCCTTCGACGAACTCCTCGG CACCGCGTCGGGCGCGAGCACCTTGTTCGCGGGCACTGGAGGCGTA ATGGAGTCGGTTGTCCGCGCCGCCTACTACGAGTTGATGGGCGGAA AGGAGTTGCTGCCCTTCGTGGAAGACCGCCGCCCCGGCTACGGCGG CACCAAGGAGTTCACAATTGAAGTGGGCCTGGAAAAACCGTTGCGCG TCGCGGTCGTCAACGGGTACGACGAGGCAAGAAAGATGTGCGAACG AGTGCTCAAGGAAAAAAAGGAGGGCGGGCGCACGATTGACTTCATTG AAGTGCTCGCGTGCCGCGGCGGCTGCGTGGGCGGACCCGGCCAGC CCGGCAACGACCCGCAACAAGTGGCGAAGCGCGCCTTCGGGTTACG CGACTTGGACGCCGAGAACACGCGCGTGCGCAACGCGCACCAGAAC CCGGACGTCAAAAAACTTTACCGCGAATTCCTCGGCGAAACCGGCGG GGAAAAAGCTCATAAACTCCTGCACCGAAAGGGAACCCAGGAGATAT TGGAGACGCTGAAATGA 141 NZ_CP042905.1_ B, M3a' ATGCCGTTACCCAGATTAAATATATTTGAAAAAACGCGCGGTCTATAC 15, Ps, nucleotide ACGCCGGTAATTGATATTCGAAGGCGGGTTCTTTCAGCTGTTGCTCG AATGGTCGTAGAAAATAAACCACCGACATATATTGAGCATATTCCATA TCAAATTATTGATAAAGATACTCCAACATATCGAGAATCGGTATTCAAG 1005373994
GAACGTGCAGTTGTTAGAGAACGAGTCAGGCTTGCGTTCGGAATGGA TTTAAGGGAATTCGGCGCTCATGGACCGATTAATGATGATGATGTTAT ATATTCCTTAACCGATAAGAAAGTGATTAAAAAGCCTATTGTGAATGTC ATAAAAGTTGGATGTGAACGTTGCCCTGAACATTCTTATATTGTTACTG ATTTATGTAGAGGTTGTATTGCTCACCCTTGCACGATAGTTTGTCCAA AAAATGCGGTTTCGATAATTAATAACAGATCCCATATAGATCAAGATCT TTGTATAAGATGTGGTAAATGTGAGCAAGTGTGTCCATACAACGCAAT TGCATACCGAGAAAGACCATGTGCAGCGGCATGCGGGGTAAAAGCA ATCAGTTCAGATAAAGACGGTTTTGCGGATATCGACCAAAATAAATGT GTTTCTTGCGGAATGTGTATTGTGTCTTGTCCTTTTGGAGCAATTGCA GAAAAATCAGAAATCGTCCAAATAATTTCAGCGCTTAAAGGAAAAAAA CCTGTTGTTGCAGAAATTGCTCCTGCGTTTGTAAGTCAATTTGGACCT TTAGTCACCCCAGGCAAATTAAAAGGTGCTTTAAAAGAAGTCGGTTTT ATGGATGTTCGAGAAGTGGCCTATGGAGCTGATGTAGTTGTGATGAA TGAAACTGAAGAATTGGTCGAATTAATTAAAAAACGGGAAAATTGTAA AGAGAATGAAGAATCTGAAAGTAATTGTCGAACATTCATCGGAACCAG TTGTTGTACCTCATGGGCATTGGCTGCAGAGAAAAATTATCCTGAATT ATCTAAACTAAATATTTCTGAATCATTTGCTCCAATGGTAGAAATTGCA AAAAAGATCAAAAAAGATACTCCAGATGCTATTGTTGTATTCATTGGAC CATGTATATCTAAAAAGGAAGAGTGTTTCATCCCAAATGTGGCAGAAG TCGTGGATTTTGTAATGACTTTTGAAGAACTAGTCGCCATTTTTCAAGC TTTTACAATAGATCCTACAAAAATTCCCCCAGAAAAGGAGATGCAAGA TGCTTCTTCACTGGGAAGAGGTTTTCCCGTTGCTGGTGGAGTAGCAA ATGCAGTACTTGAGCAAACCAAAGCAATTATTGGACAAGATGTTGAGA TTCCAATTGTATCTGCAGATACTTTAAAAAACTGTATGAGTTTATTAAA TCAAATCAAAAAAGGAGCACTCGATCCAAAACCATTGATTGTTGAAGG AATGGCTTGCCCTAATGGGTGCGTGGGAGGACCGGGAACTTTAGCC CCTTTAAGAAGGGCACAACGAGAAGTAAAAAAATTTGCGAAGAAAGC ACAATGGGAAAAACCGAAGGATTATCTGTAG 144 AQRS01000037.1 A1, M1 ATGGGCAGCATCGAAGACGTGAACGCGGCGCTGGCGGATGAGGGCA _10, la, nucleotide AGATGGTTATGGCGCAAGTGGCGCCGGCGGTGCGTGTTACCATCGG codon optimised CGAGGAATTCGGCCTGCCGGCGGGTACCATTGTTACCAAGAAACTGG TGGGCGCGCTGCGTCAAGCGGGTTTCGAAAAAGTTTTTGACACCAGC GTGGCGGCGGATATCGTTACCATTGAGGAAGGTACCGAGTTCCTGAA CCGTCTGGAGGACCAGGAAGATCTGCCGCTGCTGACCAGCTGCTGC CCGGCGAGCGTGTTCTTTGTTGAAAATACCTTCCCGAAGTTTCTGCAC CATTTTTGCACCGTTAAAAGCCCGCAGCAGGGTATGGGCAGCCTGAT CAAGACCTACTATGCGCGTCGCATGAAAATTGACGCGAAGAAAAACT TCGTGGTTGCGGTGATGCCGTGCATTGTTAAGAAAATGGAGGCGCGT CGCCCGGAGATGGAATTTGACGGCGTGCACAACGTTGATGCGGTGC TGACCACCAAGGAAGCGGCGGCGCTGCTGAAAAGCAAGAAAGCGGA CCTGAACGGCGCGAATGAGACCGGTTTCGATAGCCTGCTGGGTAAA GCGAGCGGTGCGGGTCAGCTGTTTGGTCAAACCGGTGGCGTGAGCG AAGCGCTGCTGCGTTTTGTTGCGTGGAAGCTGGAAGGTAAGAAAGCG CGCGTGCTGTTCAAAGAGGTTCGTGGTAAGAAAGGCTTTCGCGAGGC 1005373994
GGAAGTGAAGATCGGCAGCCGTATGCTGAAAGTGGCGGTTATTGACG GTCTGAACAATCTGCGCGATCTGATGAGCAGCGAGGAAAAGTTTCAC AGCTACGACGTGGTTGAAATCATGACCTGCCCGGGTGGCTGCATTGG TGGCGGTGGCCAACCGGAGAGCACCCCGGAGAAACTGGAAGCGCGT AAGAAAGCGCTGCATAACGTGGATGCGATGGAAACCGTGCGTGTTGC GAGCGAGAATCCGGAAGTTAAGGAGCTGTACGGCAGCTATCTGATCG AGCCGGGTAGCAAAGTGGCGCGTAGCATCCTGCACGTTAGCCGCAT TTGCCTGAAGTGCGATAGCAGCGGTTGGAGCCACCCGCAGTTCGAG AAATAA 145 CABMGN010000 E, M1a ATGGTTAAGAAAAGCCCGAAACTGAAGTTTCCGCTGAAAGGTCGTAA 008.1_137, Na, GAAATACGTGGTTATGCTGGCGCCGAGCTACATCGTGGACTTCAGCT nucleotide codon ATCCGGAGATCATTTTTGCGCTGCGTAAGCTGGGTTTCGATAAAGTG optimised GTTGAACTGACCTTCGGCGCGAAGATGGTGAACCGTGAGTATCACAG CATCCTGGAACACAACCTGAGCGCGCACGGTTTTTGGATCAGCAGCG TTTGCCCGGGCATTGTGGACCTGGTTAGCACCCGTTTCCCGCAGTAC CGTAAAAACCTGATTCCGGTGGATAGCCCGATGATCGCGATGGCGAA AATTGTGCGTAAGACCTATAGCAAACACGGTATCGTTTTCATTAGCCC GTGCAACTTTAAGAAAATCGAGGCGAAGGACAGCGGCGTGGTTGACT ACGCGATTGATTATAGCGAGCTGATGGAAATCTTCCGTAAGAAAAAGA TTAGCCTGGAGAGCTTCAGCGATCACGAAAAGGCGCACTTCGACAAG TTCTACAACGATTACACCAAGGTTTACCCGCTGGCGGGTGGCCTGAG CAAAACCGCGCGTCTGAAGGGTCTGCTGAAACGTCGTGAGATCAAAA AGATTGACGGCGCGGAAAAAGTGATCGAGTTTCTGGAAAACCCGAGC ATTAAAACCAAGTTCCTGGATGCGAACTTTTGCGAGGGTGCGTGCAT CGGTGGCCCGTGCATTTACAGCAAAAAGCTGAGCCTGCGTAAGCGTC GTCGTAAAGTGCTGAAGTATCTGAACCAAAGCAAACGTGAGGAAATC CCGAAAACCGACAAGGGTCTGGTTAAGTGCGCGGAAGGCATTAACTT CCGTCGTTATGATCTGAGCAGCGGTTGGAGCCACCCGCAGTTCGAGA AATAA 146 CP045477.1_677, E, M1a ATGGAGCTGAACACCGAACTGATTGACGTGCTGAAGACCATCAACAA Fm, nucleotide AGAGAAGGTTATTTGCCTGCTGGCGCCGAGCTTCGTGGTTGATTTTA codon optimised AATACCCGAAGATCATTCTGGAGCTGCGTCGTATTGGTTTCAACAAAA TCGTGGAACTGACCTATAGCGCGAAACTGATCAACAAGGAGATTCAC AAACAGATCCTGGAAAACAAGAAAAAGCAATACATCTGCGGCAACTG CCCGAGCGTGGTTAAGTACATTGAGAACAAATATCCGGAACTGAAAA AGAACATCATGGACATTGCGAGCCCGATGGTGATCATGGCGCGTTTC ATGAAAGAGAAGTACCCGAGCCACATCATTGTGTTCGTTGGTCCGTG CTTTAGCAAAAAGCAGGAAGCGAAAGAAAACAAAGAGGTGGACTATG CGCTGACCTTCAAGGAAATCAACGATATGTTTCTGTACGCGAAAAAGA ACGGCTTCTACAAGAAGCAAGAGAACGAAAACAAGCACTTCGACAAG TTCTACAACGATTACACCAAGATCTATCCGCTGAGCGGTGCGGTGGC GGAGACCATGAACACCCGTGAAATTCTGAAACCGGAGGAAATGATCA TTGCGGACGGCATCAAGGAGATCGATAAGGCGATCATGAAGTTCAAG GAAAACAAAAAGGTGCGTTTCCTGGATATGCTGTTTTGCACCGGTGG CTGCGTTGGTGGCCCGGGTATCATTAGCACCGAAACCATCGAGCAGC 1005373994
GTGAAGACCGTGTTATTCACTACCGTGATAAAAGCAAGAACGAGAAA CTGGGCAAGAACTTCGGCAAATTTAAGTATGCGGAAAACCTGAGCCT GAAACGTAAGAGCAGCGGTTGGAGCCACCCGCAGTTCGAGAAATAA 147 DALH01000010.1 A1, M1 ATGGACGATATCAGCTTTGTTAAAGAGGTGCTGGCGGATAAGAGCAA _13, Mu, AACCGTTATTGCGCAAACCGCGCCGAGCGTGCGTGTTACCGTGTGC nucleotide codon GAGGAATTCGGTAGCGAACCGAGCAGCGAAGAAACCGGTCGTGTGG optimised TTGCGGCGCTGCGTAAACTGGGTTTCGACAACGTTTTTGACACCGAT TTCGGCGCGGATGTGACCGTGGTTGAGGAAAGCGCGGAGCTGGTTC GTCGCCTGCGTGAAGGTGGCGCGCTGCCGGTGTTCACCAGCTGCTG CCCGGGTTGGACCCGCTTTTGCGTGAAGGCGTTCCCGGAACTGAAT GAGCACCTGAGCGGTGTTAAAAGCCCGCAGCAAATCGTGGGCGCGC TGACCAAGACCTACTTTGCGGAGAAAACCGGTAAGAAAGCGGGCGAA ATTTTCGTGGTTGGCGTTATGCCGTGCTATAGCAAGAAACTGGAGAG CCGTGAGGAAGAGACCGAAATCCCGGGTGCGAAGCAGGACGTGGAT TTTATTCTGACCACCAAAGACTTCGCGAAGCTGATCAAACAAGAGGG CATTGATTACCATAGCCTGGAACCGGACAGCTTTGATGAGCTGCTGG GTACCGCGAGCGGTGCGAGCACCCTGTTTGCGGGTACCGGTGGCGT TATGGAAAGCGTGGTTCGTGCGGCGTACTATGAACTGATGGGTGGCA AGGAGCTGCTGCCGTTTGTGGAAGATCGTCGCCCGGGTTACGGTGG CACCAAGGAATTCACCATCGAGGTTGGCCTGGAAAAACCGCTGCGTG TGGCGGTGGTTAACGGTTATGACGAGGCGCGTAAGATGTGCGAACG CGTTCTGAAAGAAAAGAAAGAGGGTGGCCGTACCATCGATTTTATTGA GGTTCTGGCGTGCCGCGGTGGCTGCGTGGGTGGCCCGGGTCAACC GGGTAACGACCCGCAGCAAGTGGCGAAGCGTGCGTTTGGCCTGCGC GACCTGGATGCGGAAAATACCCGTGTTCGCAACGCGCACCAAAATCC GGACGTGAAGAAACTGTATCGTGAGTTCCTGGGTGAAACCGGTGGC GAGAAGGCGCACAAACTGCTGCATCGTAAGGGCACCCAGGAAATTCT GGAGACCCTGAAAAGCAGCGGTTGGAGCCACCCGCAGTTCGAGAAA TAA 148 NZ_CP042905.1_ B, M3a’ ATGCCGCTGCCGCGTCTGAACATCTTTGAGAAAACCCGTGGTCTGTA 15, PS, TACCCCGGTGATTGACATCCGTCGCCGTGTGCTGAGCGCGGTTGCG nucleotide, codon CGTATGGTGGTTGAGAACAAGCCGCCGACCTACATCGAACACATTCC optimised GTATCAGATCATTGATAAGGACACCCCGACCTACCGTGAGAGCGTTT TCAAAGAACGTGCGGTGGTTCGTGAGCGTGTGCGTCTGGCGTTCGG CATGGATCTGCGTGAATTTGGTGCGCACGGCCCGATCAACGACGATG ACGTTATTTACAGCCTGACCGACAAGAAAGTGATCAAGAAACCGATTG TTAACGTGATCAAAGTTGGCTGCGAGCGTTGCCCGGAACACAGCTAT ATCGTGACCGATCTGTGCCGTGGTTGCATCGCGCACCCGTGCACCAT TGTGTGCCCGAAGAACGCGGTTAGCATCATTAACAACCGTAGCCACA TCGATCAGGACCTGTGCATTCGTTGCGGTAAATGCGAGCAAGTTTGC CCGTACAACGCGATTGCGTATCGTGAACGTCCGTGCGCGGCGGCGT GCGGTGTGAAGGCGATCAGCAGCGATAAAGACGGTTTCGCGGATATT GACCAAAACAAGTGCGTGAGCTGCGGCATGTGCATCGTTAGCTGCCC GTTTGGTGCGATCGCGGAGAAGAGCGAAATTGTTCAGATCATTAGCG CGCTGAAAGGCAAGAAACCGGTGGTTGCGGAGATTGCGCCGGCGTT 1005373994
CGTGAGCCAATTTGGTCCGCTGGTTACCCCGGGCAAGCTGAAAGGC GCGCTGAAAGAGGTGGGTTTCATGGATGTTCGTGAAGTGGCGTACG GTGCGGACGTGGTTGTGATGAACGAAACCGAGGAACTGGTGGAACT GATCAAGAAACGTGAGAACTGCAAAGAAAACGAGGAAAGCGAAAGCA ACTGCCGTACCTTCATTGGCACCAGCTGCTGCACCAGCTGGGCGCTG GCGGCGGAGAAGAACTATCCGGAACTGAGCAAACTGAACATCAGCG AGAGCTTTGCGCCGATGGTGGAAATTGCGAAGAAAATCAAGAAAGAT ACCCCGGACGCGATTGTTGTGTTCATCGGTCCGTGCATTAGCAAGAA AGAGGAATGCTTTATCCCGAACGTGGCGGAAGTGGTGGATTTCGTTA TGACCTTTGAGGAACTGGTGGCGATTTTCCAGGCGTTTACCATCGAT CCGACCAAGATTCCGCCGGAGAAAGAAATGCAAGACGCGAGCAGCC TGGGTCGTGGCTTTCCGGTTGCGGGTGGCGTGGCGAACGCGGTTCT GGAGCAGACCAAAGCGATCATTGGCCAAGATGTGGAAATCCCGATTG TTAGCGCGGACACCCTGAAGAACTGCATGAGCCTGCTGAACCAGATC AAGAAAGGTGCGCTGGACCCGAAACCGCTGATTGTTGAGGGTATGG CGTGCCCGAACGGCTGCGTGGGTGGCCCGGGTACCCTGGCGCCGC TGCGTCGTGCGCAGCGTGAAGTGAAAAAGTTTGCGAAGAAGGCGCA GTGGGAGAAGCCGAAAGACTACCTGAGCAGCGGTTGGAGCCACCCG CAGTTCGAGAAATAA Table 2: Amino acid sequences of HydA SEQ Description Subgroup, Sequence ID subclass NO: 151 AQRS01000037. A1, M1 MGSIEDVNAALADEGKMVMAQVAPAVRVTIGEEFGLPAGTIVTKKLVGAL 1_10, amino acid RQAGFEKVFDTSVAADIVTIEEGTEFLNRLEDQEDLPLLTSCCPASVFFVE NTFPKFLHHFCTVKSPQQGMGSLIKTYYARRMKIDAKKNFVVAVMPCIVK KMEARRPEMEFDGVHNVDAVLTTKEAAALLKSKKADLNGANETGFDSLL GKASGAGQLFGQTGGVSEALLRFVAWKLEGKKARVLFKEVRGKKGFRE AEVKIGSRMLKVAVIDGLNNLRDLMSSEEKFHSYDVVEIMTCPGGCIGGG GQPESTPEKLEARKKALHNVDAMETVRVASENPEVKELYGSYLIEPGSKV ARSILHVSRICLKCD* 152 CABMCJ0100000 A1, M1 MRRRVKKTNEFGLDRWRYGKFLKVCAKIVIMQADPIGKFNEICGLKKERT 04.1_889, amino LCVQIAPAVRVALGEEFGMPIGTDVTKKLVWAMRKIGAEYVFDTPLGADII acid AMEEANELKKMLEAGGPFPVFTSCCAGWMLFWKRAHPELEKNVCELVA PQMALGALIKTYFARVSGMKKEPYVLSIMPCLLKEQEAQGTMKDGRKYV DCVFTTKDIAKVLKSEGIDLKNAPEEEFDKISGMGSGEGAIFGASGGVME AALQNLGKMLGEKVEFKEFRNEENMKKSTVKIGKYTLNVAAVWGLPNVD KLLAEMTEGKVYHFVEVMACPGGCIGGGGQPLPPTPDVVKARAAAMRK YADMLKISAKENPKVDEIYRAFLKKVGSEIGHELFSSRKKN* 1005373994
153 CABMDO010000 A1, M1 MEELERLLSSKKIVAVQTAPSVRVTLGEEFGLKPGTDVTGRVVAALRMLG 013.1_20, amino FKYVFDTDFGAEVTVLEEAQELLERLEQQERLPLFTSCCPGWTRFCLKQ acid FPTLEQNLSECKAPQQILGSLAKTYFAKKIRKQAKDVAVVSIMPCFSKKLE AKEREHAIAGASQDVDLVITTHELAQMLKAKGIDLRRLSDEEFDNPLGVT GGAGALFGATGGIAEAVMRTAYHATTGGALPRFESNMRPALGQIKEYSI MMGDRDVHLAIVQGYPNASVVCKRVLEEKRNGMRSLDFVEVLGCIGGC VGGPGQPDTGEDAVLMRAAALHERDRRLPLRSAHDNPAVKKVYASMLG KPGSAKAKKFLHRGQRGYAEMLKTL* 154 CAIKOA0100000 A1, M1 MSTEIYLNKPEELFAALEKELADPENTVVVQIAPAVRVAIGEEFGYPSGED 26.1_15, amino LTFKTIGLLNALGFKHVVDTPLGADINIYEEAYEILNALERKDDNYFPIFNSC acid CIGWRLYCSRAHQKVCQHISPIASPHMITGSVIKHYFSKKLNKPKEKIISVSI MPCVLKKYESLERFSDGSTYIDYVVTTRELAQWAKKRGLDLRKVNEGKF SEFLPNSSKDGVGFGVTGGLVEALLTTMAHILEVKKEAESFRTNEPIKER QVKIGEYTLNVVSINGFGNFEKVLKEIESGERKFHFVEVMNCPYGCVGGP GQPLPVNDEILKARAAGLRLAADKKRSIAIVPQENPTVQMLYRELLGRPG SEKARELLYFHKIKL* 155 CAIWXP0100001 A1, M1 MAEEIYMNKLDDLINAIEAEKKKGKVLVAQLAPAVRITLGEEFGYNVGEDL 75.1_2, amino TKKCVGLLKTLGFDIVIDTPLGADIAVYEEVHILKELLDSKDLSYFPMFNSC acid CIGWKMYAKRMHLDLMPHVSSIGSPNQIVGSIAKNYLAYQQNKYPDDIVV VGIMPCTLKKFETLNTFEHKKGNTIVKLKYVDYVVTTTELAEWSRKKNIDF SQVAEYEMINPASKEGTIFGVTGGISEAFINAFAKYIGEEKEILDFRQDEKS RKYKVKIGNYELSVAIVFGVGQLEHILEDIENGEFFHFVEVMYCNSGCVG GPGQPRASDTTIEERAKAMRNCSDKIKKTTCFDNGALLEMYKNLNMKPN NAKAKEMFFLKEQN* 156 CAIXOA0100001 A1, M1 MEDEIYLNKPDELFAAIENGKKNGKIMVVQIAPAVRVSIGEEFGRAPGEDL 93.1_5, amino THKTVGLLQALGFDHVMDTPLGADVNIYEETLEVLHALERDDEKYFPVFN acid SCCIGWRLYCKNKHPELYRLVSPIASPHMIAGSIGKHVLAKKLGVPVEKIC MVSIMPCVLKKYETRERLPSGIKYIDYVLTTHELGMWAKKKGLDINRVRD GKFTALLPDSSKDGVIFGATGGITEALLSTLACICGESPEKVRFRGDEQVK HLCVQIGKHQLNVVSIYGVTNLDKVLDEIKHGVKYHFVEVMNCPYGCVGG PGQPLPASEEKYRARAHGLRKAADRKQGKCPLGKMGVHSIYEALGIGPG SREAQELFFFHKTNI* 157 CG08_land_8_20 A1, M1 MRFEILERAVRDKSILKVAQLAPAVRVSLGEMFGFEAGTILTKKIVGALKEL _14_0.20_scaffol GFDYIFDTSFGADVAIVEESKELGDRLKKGGIFPMINSCCPGTISFLEHAYP d_32353_2, DLVPNIATVKSPMEVTGVLIKTYFAEKKKIPPEKILSVAVMPCIIKKAEAFRP amino acid ELRMNGKLVIDGVLTTVELGELLKAKKIDLKNCKEREFDSLMGVASGAGQI FGSTGGVSEAAIRNYAHMNNIPIEKIDTKQLRSFEGVREMEFSLGEKKIKIA IINSLRNANQVLNDSDKMKEYTFIEIMACLGGCVGGAGQPTSTKEILEKRR AGLYSIDAKTKIKISSENPEVKKLYEEFLGAPRSKRALRILHTGFVRTCVDC F 158 CG10_big_fil_rev A1, M1 MGFEKVVEELEAKEKFLVAQTAPSMRVSIGEEFGYRPGEIVTGKLAGALK _8_21_14_0.10_s ELGFDAVFDTCTGADLVTMEETYEFLERKKKGERLPIMTSCCPGFVSYIE caffold_50598_2, HTHPEYVDNLCSCRSPQEIMGALIKTYFAQKRKLKPKDIYVVSIMPCIIKKA amino acid EALRPELRVNGMKNVDKVLTTVELAQLLKARGIDLKKVKEADFDSLLGES TGSANIFGATGGVLETVLRLAAKITDKKVGVIEFREIRGMEGVKEAEVRIG 1005373994
KSKVKVAVINGLRYASQLLNDRERTKSFDIIECMACFGGCVGGAGQPRTT LDIIDARKEALYRIDRGKKQRIASENPSVKKLYKDFLGKPGSAKAKKLLHT HYHKFFE 159 CG10_big_fil_rev A1, M1 MATIEEVNKALDENKVQVVAQVAPALRVSIGEEFGFAPGKVLTNEFVAAL _8_21_14_0.10_s KKTGFDKVFDTSTAADIVTIEEGTELLKRLKTNESLPMFTSCCPGSVLYIEN caffold_51986_4, NYPEYVNHFCTVKSPQQSMGALIKTYYARQMNLQRKDLFVVAIMPCVVK amino acid KMEAKRPEMEFDGISHVDAVLTTREIAELLKGRGISIKEGEKADFDNLLGD ASGAGQIFGTTGGVSEALIRYVSEKTGSNPGKTEFKELRGTAGFREATVK IAGKNVHIAVIDGLNNLKNLISDPEKFNSFQVIELMVCPNGCIGGGGQPRT TPEGLAARRAALMGIDSGEKIRVASDNVEVQELYKKYLLEPGSRTAKAVL HTNRICLKCN 160 CG10_big_fil_rev A1, M1 MEELERLLSSKEIAVVQTAPSVRVSLGEEFGLKPGADVTGRVVASLRMLG _8_21_14_0.10_s FKYVFDTDFGAEVTVLEEAQELLERLEKQERLPLFTSCCPGWTRFCLKQF caffold_5639_c_1 PSLENNLSECKAPQQILASLAKTYFAKKIRKPAKDVVVVSVMPCFAKKLEA 5, amino acid KEREHAIAGAPQDVDLVITTHELAQLLKARGIDLRRLSDEEFDNPLGVTGG AGALFGATGGIAEAIMRTAYHATTGGALPRFESNMRPALGQIKEYSIVMG E 161 DALH01000010.1 A1, M1 MDDISFVKEVLADKSKTVIAQTAPSVRVTVCEEFGSEPSSEETGRVVAAL _13, amino acid RKLGFDNVFDTDFGADVTVVEESAELVRRLREGGALPVFTSCCPGWTRF CVKAFPELNEHLSGVKSPQQIVGALTKTYFAEKTGKKAGEIFVVGVMPCY SKKLESREEETEIPGAKQDVDFILTTKDFAKLIKQEGIDYHSLEPDSFDELL GTASGASTLFAGTGGVMESVVRAAYYELMGGKELLPFVEDRRPGYGGT KEFTIEVGLEKPLRVAVVNGYDEARKMCERVLKEKKEGGRTIDFIEVLACR GGCVGGPGQPGNDPQQVAKRAFGLRDLDAENTRVRNAHQNPDVKKLY REFLGETGGEKAHKLLHRKGTQEILETLK* 162 DTGI01000011.1 A1, M1 MLYGQALFSAIEKEKRSGKTLVAQLAPAVRVSIGEEFGLQPGTELTAKSIS _27, amino acid LLKSLGFDEVIDTPIGADLITLEEAKFFVKTKLAFKPLFNSCCVGWREYCKS NYKNLLGYISHIVSPMMAAGFITKTYLANRMGVEPEKIVSVGVMPCTIKKL ETKYKMRSGLKYVDYVVTTQELGEWARSKGMDIKNMKDGEFSKFLPTSS KDGTIFGVTGGISEAFITTVAKLMNEDSEILFFRKNEPVREYDFAIGKIKLKT AVLHGIANFPSLLPRIKEFNFIEIMFCPYGCVGGPGQPAPPTEEKLKARAE ALRKFSDAKKQRTPLENKPLMESLDWYEQNKINVFEELTYWGE* 163 DTNT01000026.1 A1, M1 MTGNKLLYGDELFAEIEKEMNSGKVVIAEIAPAVRVSLGELFGLEAGTNVQ _4, amino acid GKTIALLRRLGFHDVVDTPLGADIATYYEAEDIKKMLDSGTGKFPIFNSCCI GWRMYAGKMHPELLDHITIVASPQMTTGSVAKYYFAEKLNKKPDEIVIVGI MPCALKKYETMEVMRNGNRYVDYVVTTVELAQWAKKKNIDFLQLPVESF SSLLPTSSKDGIIFGATGGVTEAVITTLGRLYGQEITINEFRDDAEIKRKTVT IGKHTLNIGIVHGLQNFEKLYEEIKAGKTYHLVEVMMCPFGCVGGPGQPQ ASREKIQQRARSLRQFADSAKEKTPLDNPTMQMLLRDFFSKLPRDKLEEL IYFNR* 164 DUGC01000052. A1, M1 MGIVDDVNAALDDRKKYAVCQIAPSVRVSIGEEFGFAPGAIVTKKLIGALK 1_8, amino acid NAGFRKVLDTSSAADIVTVEEGTELLGRLADREQLPLLTSCCSASVLFIEN SYPNYLPHFCSVKSPQQSMGALIKTHYARRMRLQRKNIYCVSIMPCVVKK LEAKRPEMEFNGVPHTDAVLTTREAAQLLKLRGIDLKSARGQEFDQILGK GSGAGQLFGTTGGVTEALLRFVSWKLEGKGARTEFKEVRGEEGFREAK 1005373994
VAIAGKPIRIAIVDGLTNLRDLLTNREKFYNYDVIEIMSCPGGCIGGAGQPA STKEKLSARRKALFEIDSKEKAX* 165 JACCLI01000007 A1, M1 MEDEIYLNKPDELFAALEKEKKDGKVMVVQIAPAVRVSIGEEFGRAPGED 7.1_25, amino LTYQTVGLLHALGFDHVMDTPLGADVNIYEETLEVLHALERGDEKYFPVF acid NSCCIGWRLYCKNKHPELYHLVSPIGSPHMVAGSLGKHILAKKLGVPIEKI CMVSVMPCVLKKYETRERLPSGIRYIDYVLTTHELGIWAKKKGLDMNKVK EGKFTELLPDSSKDGVIFGATGGITEALLSTLACVCGESPEKVRFRGDEQ VKHLCVQIGRHRLNVVSIYGVTNLDKVLDEIKHGVKYHFVEVMNCPYGCV GGPGQPLPASEEKYRARAAGLRKAADRKPGKCPLGKMGICGVYEALGIE PGSREAQELFFFHKTNI* 166 JACCLJ0100000 A1, M1 MDMELLSGDALFEKIEGEMKSGKIVIAEIAPAVRVTLGELFGFPVGANVLG 37.1_10, amino KMTALLKKLGFAHVVDTPLGADIATYYEAEDIKRMLDSGKGKFPIFNSCCI acid GWRLYASRAHPELLGNITIIASPQMTIGAVSKYYIADKLKTDPSNVVVVGIM PCALKKYESMEVMRNGHKYIDYVVTTIELAQWAKKKGHDLKKLNDEPLS QLMPQSSKDGVMFGVTGGMTEAVMTTLAGLYGEKKEILDFRDDLEMRK KKVRIGKHVLKIAVVYGFQNFEKLYKEIKAGEKYHLVEVMMCPLGCVGGP GQPIAPKEIVTARGNALRTVADGIKERTPIDNPTVQKLIKEYLGKLPREKLE ELIYFNR* 167 LacPavin_0920_ A1, M1 MPAEIYLNKPEELFAALEKELSNPENTVVVQIAPAVRVAIGEEFGYPPGVD SED3_scaffold_8 LTFKTIGLLNALGFKHVVDTPLGADINIYEEAYEILNALERKDDAYFPVFNS 74915_3, amino CCIGWRLYCSRAHQKVCQHISPIASPHMITGSVIKHYFSKKLGKPKEKIISV acid SIMPCVLKKYESLERFSDGSTYIDYVVTTHELAQWAKKKGIDLRKVKEGK FSELLPNSSKDGVVFGVTGGLIEALLTTMAHILEVKKENESFRTNEPIKER QVRIGEYTLNVVSINGFGSFEKVLKEIESGKKKFHFVEVMNCPYGCVGGP GQPLPVNDAILKARAEGLRAAADKKRSIAIVPQENPTVQMLYRELLERPG SERARELLYFHKIKI 168 LacPavin_0920_ A1, M1 MEDEIYLNRPDELFAALEKEKKSGKVMVVQIAPAVRVSIGEEFGRAPGED SED4_scaffold_3 LTYKTVGLLQALGFDHVMDTPLGADVNIYEETLEVLHALERGDESYFPVF 009403_2, amino NSCCIGWRLYCKNKHPELYHLVSPIGSPHMVAGSLGKHILAKKLGVPIEKI acid CMVSVMPCVLKKYETREMLPSGIKYMDYVLTTHELGNWAKKKGLDINKV KGGKFTELLPESSKDGVIFGATGGITEALLSTLACVCGESPEKVRFRGNE QVKHLCVQIGRHQLKVVSIYGVTNLDKVLDEIKHGVKYHFVEVMNCPYGC VGGPGQPLPASEERYRARAAGLRKAADRRPGKCPLGKLGVHGVYGALG IEPGSKEAQELFFFHKTNI 169 Meg19_1012_Bin A1, M1 MDEIERVKEALDKEGIQVIAQVAPALRVTIGELFGFPPGTVLTKKLVGALKA _505_scaffold_65 LGIEKVFDTSLAADVVTVEEGTEFIHRLEKNQNLPLFTSCCPASVAFVENK 0_7, amino acid HKQHVNHFCAVKSPQQTMGALIKSYYSEKMNLSHEKFFVVSIMPCVVKKI EAKRPEMDFNGIHHVDAVLTTIELAKLLKEKGIELKDAEEKEFDRLLGDAA GGGQLFGVTGGVSESLLRFVSNKLDPEKKKVEFTELRGSEGRREMELEI AGKKLKIAIVHGLHNLNDLIEDEKKFSSFQVIEMMACPGGCIGGGGQPVST PEIREKRIQGLRSVDAKETVKISSDSKEVQELYKTYLDEPGSKKARELLHT VHICLEKCD 1005373994
170 NZBD01000002.1 A1, M1 MGTIEDVNKVLDEDNVQVVAQVAPALRVTIGEEFGYKPGTVLTKQFIGSL _23, amino acid KQAGFDKVFDTSTSADVVTVEEGTEFLKRLETGECLPLFTSCCPASVLFIE RQFPQYVDHFCTVKSPQQTMGALIKTYYAEKMKLKQKDLFVVSVMPCVV KKLEAKRPEMEFNNIPHVDAVLTTKEIAELLKGRRIELEKVKEENFDNLLG DASGAGQLFGTTGGVAEALLRFVSNKLEQDLGRVEFKEVRGMEGFREA NVKIAGKKFKVAIVDGLHHLKDLLSNEKKFKKYHAIELMVCPNGCIGGGG QPISTPEIRAARRKALFKIDSKENVRICMDNPEVKKLYNDYLGEPGSRKAV SFLHTSKVCLKCD* 171 PEXZ01000074.1 A1, M1 MNELEKLFASGKTVVVQTAPSVRVSLGEEFGLKPGTDVTGRTVAALHML _10, amino acid GFKYVFDTDFGAEATVLEESQELLERLEKQERLPLLTSCCPGWTRYCLK MFPGLEGNLSEAKSPQQMLGAIAKTYFAKKIRKSAGDIAVVSIMPCFAKKL EAKEREREIAGATQDVDRVITTHELAQLLKSKGIDLRRLPDEAFDNPLGVA GGAGALFGVTGGIAEAVLRTAYYTVTEGTLPRFESNMRPSLGQIKELSMM LGEREVHLAIVQGYRDASVICKRILEERRNGVRSLDFVEVLGCIGGCVGG PGQPDTVEGSVLSRAAALHEHDRKLPFRDAHNNPAVKETYAALFGKPGS GKAKKLLHRSERDYEEIMGTLT* 172 PFAI01000011.1 A1, M1 MDEIERVHEVLEYKEKIVVAQIAPALRVSIGEEFGFPVGKILTKKFVGALKQ _3, amino acid AGFQKVFDTSVAADVVTVEEGIEFISRLEKKENLPLFTSCCPGSMAFIENK HSELVYHFCTVKSPQQTMGALIKTYYAKKMNIPQEKIFVVSIMPCAVKKFE SRRPEMDFNGIHHVDAVLTTKDIAVLLKNLKIDLNNVKENDFDSLLENASG AGQIFGVTGGVSESLLRFVSHKLEPEKKKIEFKKLRGKEGRREIELSIAGK KLKIAIVHGLQNLNNLILDKNEFDSFQVIEMMACPGGCIGGGGQPVSTPEI REKRIQALRKIDSKEKVRVSSDNPEVKELYKSFLKNPGSETAKELLHTRRI CFKCD* 173 QZM_A1_scaffold A1, M1 MLYGQALFSAIESEKKSGKILVAQIAPAVRVSIGEEFGLPPGTELTAKTISLL _68_37, amino KCLGFDEVIDTPIGADLITLEEAKFFVKTKLAFKPLLNSCCVGWREYCKSN acid HRKLLSYISNVVSPMMATGFLAKTYLAGKMGVKPENILSVGVMPCTIKKLE TRYRMRSGLKYVDYVVTTQELGQWARGNNLNIKDMEGRGFSKFLPTSS KNGTIFGVTGGISEAFITTMARLMGEESELVFFRKNEAVREYDFAIGKIKLK TAVLHGIVNFPSLLPRINQFNFIEIMFCPYGCVGGPGQPTPQTEEKLKARA EALRKFSDAKKERTPLDNKPLLEAVGWLEQKNINLFDEITYWGE 174 QZM_B4_scaffold A1, M1 MELLSGDALFSKLEEEMKSGKIVVAEIAPAVRVALGELFDFPVGTNALGKT _1937_4, amino VSLLKKLGFHHVVDTPLGADISTYYEAEDLKRMLDTGKGKFPIFNSCCIG acid WRIYASRSHPELLNNITIIASPQMTIGAVSKYYLARKLSIDPANIIVVGIMPC ALKKYETLEVMKNGHRYIDYVITTLELAQWAKKYGYDLKELKDELLDPLM PTSSKDGVIFGVTGGMTEAVVTTLAQLYGEKKEVLDFREDEGIRKKKVRI GKYVLNIAVVHGFQNFEKLYSGINAGEKYHLVEVMMCPLGCVGGPGQPA ASKETIAARGRALRQLADSIRERTPIDNPTLQKLVREYLGKLPREKLEELIY FNR 175 S_p2_S4_170907 A1, M1 MGAVEEVIGALEDGRKVAVCQVAPSVRVSIGEEFGMPPGTVATERLVGA _scaffold_232397 LRHAGFGRVFDTSTAADIVTIEEGAELLRRMKDNERLPLLTSCCSASVLYV 0_7, amino acid ENNFPELLGHFCTVKSPQQSMGSLVKTYCARKMRVKRRDIYSVSIMPCIV KKLEARRPEMEFNGVKHVDAVLTTKEAAELLKRRGLSLADAPESGFDSL MGKASGSGQLFGTTGGVTEALLRFVSWKLDGGNARIDFPEVRGSGGLR DVRVKAGGREIRIAVVDGLNNLKDILSNADRFGSYDMIEIMTCPGGCIGGS 1005373994
GQPAYTPQGLLARRDALYGIDATAKARTAMDNPAVQAVYRNYLLEPGSG VSQSILHLKRICLKCK 176 SR- A1, M1 MRMESLDAALARPDAVLVAQMAPAVRVSLGEMFGYLPGTVLTKKIVGAL VP_26_10_2020_ KKLGFKYVFDTSFGADVAVVEESKEFQERIENGGVLPMINSCCPGTVSFM 2_100CM_scaffol EHSYPELVPHIGTAKSPMEITGVLIKTYFAQKKGIDPEKIVSVALMPCVIKK d_1465248_544, AEALRPELRMDGKLVVDSVMTTVELAQALKAKGIELGKAPEADFDQLLGT amino acid ASGGGQIFGSSGGVSESALRNFAFNSGASFEKLDTAALRGAEGLRETVF SIGGKKHKVLIVNSFRNAGAVLNDKAKMGKYSFIEIMACPGGCVGGAGQP PSSKEAIEARRKGLYSIDSGMKIKAAGQNPDVKRLYDEFLGEPGGEKAVK LLHTSFVKSCENCY 221 Candidatus_Heim B, M3a' EKCIQCGKCEQVCPYNAIAYRERPCAAACGVNAITSDALGFAEIDYNKCT dallarchaeota_arc SCGLCIVSCPFAAIAEKSEIVQILQALQSKKPVFAEIAPSFPSQFGPLVNPE haeon_RS_5_1|J VILKAITQLGFVDVAEVAYGADITILNEAEELKQLIEKYKGDLSAVKADNPQ AGLLF01000095 TDRTFVGTSCCTAWSIAAKNNFPKIAEQNISESYTPMVETAKKIKKINPDAI 3_1, amino acid VVFIGPCIAKKEECFDPFVQQYVDFVMTYEELAALFQAYEIDPAEISDSFPI SDASELGRGFPVAGGVANAVVRQTLADLGKEVSIPLESAETLKDCMGML KKIKQKKYDPLPLLVEGMACPHGCVGGPGTLASLRRAQRSAKKFAKEAE WKKPTDYISEN* 222 Candidatus_Lokia B, M3a' MEKKDLNIFEKMRGIYTPVTEIRRKVLAAVARMVVEDQPPRYIEYIPYQIID rchaeota_archae KNIPTYRESVFKERAIVRERIRLAFGMELKEFGAHGPIHDDDVINTITDLKV on_bin106|JAGX LKRPIVNVIKAGCERCPEDSYIVTNLCQGCIAHPCTSVCPKNAVSIQKGKS OA010000065_3 LIDQLKCIRCGRCAQVCPYNAIAYRERPCATACGVKAISSDEYGFADINYN 3, amino acid LCVSCGMCIVSCPFGAIGEKSEIVQIISAIKNGKRVYAEIAPAFVNQFGPLA SSAKIRSALKEIGFLDIKEVALGADKVILKEAKELVAILENDNKKNISKDEKQ FIGTSCCTSWKMCADRHFPELAKNNISESFAPMVETARVIKEKDPDAIVVF IGPCIAKKEECFIPEVTALVDFVMTFEELVAVFQAFEIDPIEFEEEEPMQDA SCIGRNFPVAGGVAQAIIKQTRALLPENKKDREIPHINADTLAKCLTMLKKL KSGKFNPKPLVVEGMACPFGCIGGPGSISSLNRAKNAVKKFAEKAKNALP SDYLKKK* 223 Candidatus_Lokia B, M3a' MPLPSLNIFEKMRGLYTPVINIRRHVLSEVGRMVVEGRPPHYIEEIPYRVIP rchaeota_archae KSTPTYRESVFKERAIVRERVRLAFGMDLKEYGAHGPIHDDEVVECMKPI on_FW102|JAIZ KILKRPIVNVIKIGCERCPEESYWVTNLCRGCIAHPCVTVCPKNAVSIQNG WK010000001_1 RSVINQDLCIQCGRCAQVCPYNAIAFRERPCASACGVNAIGSDEEGYAKI 239, amino acid NYEKCVSCGLCIVSCPFGAIAEKSEIAQIISSLKAGKRVYAEIAPAFVNQFG PLVTPPKIVASLKKMGFKDVVEVAAGADRVVLQEAHELIELIKKRKSYRIET SKKKNQNLPDNLESNLNIDSKEVEDSVIDSTVNKTFVGTSCCTSWTLAAE KNFPEITAANISESFAPMVEAAKIIKEKDHEAVVVFIGPCIAKKEECFIPKVA EVVDYVMTFEELAAVFQAFHIDPTKVPPDNDLQDASSLGRGFSVAGGVA AAVIKQTKKMLEADITIPHVSVDTLKECMALLKKVKQNRLDPAPLLVEGMA CPYGCIGGPGTLAPLNRAKRAAQKYAKEAIHEYPSDFSSAESDT* 1005373994
224 Candidatus_Lokia B, M3a' MPRMFERLRGLYTPVVEIRRRVFAELARLVVEGRATQENIEAIPYQVIKDD rchaeota_archae VPHYRCCVFKERAIVRERVRLACGLDLKEGGEHEPKIPAISNVLTGKKVIS on_strain_AS27yj DPLVNVIKIGCERCPTDSVVVTDLCRNCMAHPCIIVCPVNAISIVEGRSHV COA_147|JAAZN DTGKCIKCLRCTQVCPYEAIVRRGRPCAQACGVDAIGSDAQGYAEIDHDK I010000050_6, CVSCGLCTVSCPFGAIADKSEIMQVLHELHQKERVHAIIAPSFVSQFGNKV amino acid SPPAIFRGLQKAGFAGVMEVAYGADHDMIIEADKLAKIIEKSKQASMNAIP GNHEFLGTSCCPSWVMTARKFFPELASHISDSYTPMVVTAKKVKQDDPG AKVVFIGPCVAKKHESLLSPMNEWVDHVLTFEELAAIFVAMRIDLMEIQDG LPIADASRYGRGYAVAGGVCNAIATYTRSCYGFPEFEVNRADTLRDCRK MLASIKNGEIMPRLVEGMACPDGCIGGPGTLAPLKLAKVSVEKFSKDAGK KEPSCESR* 225 Candidatus_Lokia B, M3a' MPLPPLNIFEKMRGLYTPVVNIRRHVLSEVGRMVVEEKPPHYIEEIPYRVI rchaeota_archae PKSTPTYRESVFRERAIVRERIRLAFGMELKEYGAHGPIHDDEEAECMKP on_strain_B53_G VKILTRPIVNVIKIGCERCPEVSYWVTNLCRGCIAHPCVSVCPKNAVSIVND 9|QMYW0100008 RSVIDQNLCIRCGRCAQVCPYNAISYRERPCSAACGVKAISSDSEGYAQI 5_2, amino acid NYDRCVSCGLCVVSCPFGAIAEKSEIAQIISALKSSQHVYAEIAPAFVNQF GPLVSPAKLVVSLKKMGFRDVIEVAAGADRVVLQEARELLELIQVKNGEK NEGEKGKSPNIIRKFVGTSCCTAWTMAAERNFPELKKLNISESFAPMVEA AKLIKATDPEAIVVFIGPCIAKKEECFQPEVAEVVNFVMTFEELAAMFQSFL IDPTKIEPQKDLHDASTLGRGFPVAGGVAAAVIKQAQKISSVPLDIPYAAV DTLKDCMSLLKKLEGNKVVPAPLLVEGMACPFGCVGGPGTLAPINRARR AAQKYAKEADHQYPSDILL* 226 Candidatus_Lokia B, M3a' MPPLPVLNIFEKMRGIYSPVVNIRRHVLSEVGRMVVEGKPPSYIEQIPYHVI rchaeota_archae PSSLPTYRESAFKERAIVRERVRLAFGMDLKEFGAHGPIHDEDVVLSLQP on_strain_HM1_B NKIITRPIVNVIKIGCERCPEDSYWVTDLCRGCIAHPCVSVCPKNAVSIIEG 6_4|JABCSI0100 KSVIDQEKCILCGKCAQVCPYNAIAYRERPCAVACGVNAISSDADGFADID 00124_6, amino YEKCVSCGLCIVSCPFGAIAEKSEIAQIISALKSEKSVYAEIAPAFVNQFGPL acid TSPEKIVASLKQMGFKGVREVAAGADKVVLNEAQELLELINQKKSTKKGS PAKGERTFVGTSCCTAWTLAAQDHYPDLARLNISESFAPMVETARLIKAK DPDAVVVFIGPCIAKKEECFREEVIPFVDFVMTFEELAALFQSFLIDPTKIEP DADLSDASTLGRGFPVAGGVAAAVKAQTCKLLGDNCEIPTQAADTLKDC LAMLKDLSRNKFDPAPLICEGMACPFGCIGGPGSLSSLRRAKNASKKYAK QAPFQFPSDHVADLE* 227 Candidatus_Lokia B, M3a' MPLPRLNIFEKMRGLYTPVIDIRRRVLSAVARMVVKDLSPTHIEHIPYQIIDK rchaeota_archae DTPTYRESVFKERAVVRERVRLAFGMDLKEFGAHGPINDDDVIYSLTDKK on_strain_HM4_B VINKPIVNVIKVGCERCPEHSFIVTDLCRGCIAHPCTIVCPKNAVSIINNRSV 48|JABCTF01000 IDQDLCIRCGKCEQVCPYNAIAYRERPCAAACGVKAISSDENGFAEIDQD 0130_2, amino KCVSCGMCIVSCPFGAIAEKSEIVQIITALKGKKPVIAEIAPAFVSQFGPLVT acid PGKLKGALKEVGFTDVREVAYGADVVVMNETEELVELIKKREDCEDNEK SEAKCRTFIGTSCCTSWAMAAEKNFPELSKLNISESFAPMVEIAKKIKKET PNAIVVFIGPCISKKEECFIPEVAEVVDFVMTFEELVAIFQAFTMDPTKIPPE KEMKDASALGRGFPVAGGVAKAVLEQTKAIIGEDIDIPIVSADTLKECMSLL RQVKNGKLDPKPLIVEGMACPNGCVGGPGTLAPIRRAQREVKKFAKKAK WEKPKDYL* 1005373994
228 Candidatus_Lokia B, M3a' MPDLPKLTIFEKTRGLFTPVKDVRRQVLTEVAKMIVKKKPASYIETIPYNIIA rchaeota_archae KDTPTYRGSVFRERAIVRERARLAFGLDLKEFGAHAPIIDDVSPAMTDKKF on_strain_MAG_ LGLPIINIIKAGCERCETDSFWVTNNCSKCLAHPCLIACPVDAITIQDRAFID 14|DXJH0100033 QEKCIKCGRCAQVCPYNAIVHRERPCAQACGVGAITSDEDGFAEIDYEKC 8_6, amino acid VSCGMCIIACPFGAAAEKSEIVQIIHTLRGDKPVYAEIAPSFVGQFGPLVKA PMIIEAIKKLGFTGVVEVAYGADVDILLEAKELAEMITNKGTTPREFIGTSC CPAWVLAAISNFPKQAINISDSFTPMVETARKIKKIEPDARVVFIGPCIAKK NECFAPEVSGLVDFVMTFEELGALFQAFGVDPSEMKATEDIKDASEVGR GFAVAGGVANAIIAQTQEILKKKVDIPHTSADTLADCMQMLKDIQNGKIDP KPLLVEGMACPFGCIGGPGTLAPLHRAKRKVKTFSKRAKVKLPSDHLK* 229 Candidatus_Lokia B, M3a' MPDLPKLTIFEKSRGLPTPIKGIRRRVLSEVAKMIVENKLPSYIETIPYIIISKD rchaeota_archae TPTYRDSVFRERAIVRERVRLAFGLDLQEFGAHGPIIDDVTPAITDKKFLGL on_strain_YT2_0 PIVNVIKAGCERCETDSFWVTDNCRRCLAHPCTVVCPVDAVSIQENRGLI 12|JAEOTN0100 DQEKCVRCGRCAQVCPYNAIVHRERPCAQACGIGAIITDKDGFADIDYEK 00030_11, amino CVSCGMCLISCPFGAIGEKSEIVQILHALKGKHPVYAEIAPSFVGQFGPLV acid KASMIIEAIKKVGFAGVVEVAYGADVGTLKEAQELVDICCNSEDHKDFIAT SCCTAWKQAAVNNFPDLASNISESYTPMVEAARKIKAQEPKSRVVFIGPC IAKKNECFAPEVAELVDFVMTYEELGAIFQATGIDPSEMKATKDIKDASEA GRGYAVAGGVANAIVSQTQEILGKEVDIPYTSADTLADCMKMLNNIQKGK LDPKPLLVEGMACPFGCIGGPGTLAPLHRAKRKVKTFSKRAKVKLPSDYL KK* 230 Candidatus_Lokia B, M3a' MPDLPKLSIFEKMRGIFTPVKNVRRQVLTEVARMIVEGHPPSYIETIPYNIIS rchaeota_archae KDTPTYRESVFRERAIVRERARLAFGLDLKEFGAHGPIIDDVTPAMTDKKF on_strain_Zod_M LSRPIVNVIKAGCERCETDSFWVTNNCRKCMAHPCSIVCPVEAVTIGEKA etabat.1044|JAF AIIDQEKCIRCGRCAQACPYNAIVHRERPCAQACGVNAISSDENGFAEIDY GBS010000008_ EKCVSCGLCIVSCPFGAIGEKSEIVQIIYTLRGEKPVYVEIAPSFVGQFGPL 97, amino acid VKPSMIIAALKQIGFAGVVEVAYGADVDTLILSRELADLVSKKTKDKRRYIG TSCCPAWVLAAITNFPKQAINISESFTPMVEAARKIKKMDPEARVVFLGPC IAKKNECFTPEVAELVDFVMTFEELGALFQAFSIDPSEMNVEEEISDASEV GRAYAVAGGVADAILAQTKEILGKEIDIPFTSADTLADCMMMLKNIQNSKIS PRPLLVEGMACPYGCIGGPGTLSPLHRAKRKVETFSKRAKVKLPSQHLK E* 231 Candidatus_Lokia B, M3a' MGKKDLNIFEKMRGIYTPVTEIRRKVLAAVARMIVEDQPPRNIEYIPYQIIDK rchaeota_archae DIPTYRESVFKERAIVRERIRLAFGMELKEFGAHGPIHDDDVINTITDSKILK on_strain_Zod_M RPIVNVIKAGCERCPEHSFIVTDLCRGCIAHPCTSVCPKNAVSIQKGKSFID etabat.578|JAFG QSKCIRCGRCAQVCPYNAIAYRERPCAAACGVKAITSDDYGFADINYNLC OA010000117_2, VSCGMCIVSCPFGAIGEKSEIVQIITAIKKGKRVYAEIAPAFINQFGPLATPA amino acid KIHSALLKMGFLDIKEVALGADKVVLKEAEELVKLIKQKKEPHISKDQRNFI GTSCCTSWKMCADHHFPNLAKNNISESFAPMVETARVIKNKDPEGIVVFI GPCIAKKEECFIPEVKSLVDFVMTFEELVAVFQAFDIDPVDLTEEESINDAS SIGRNFPVAGGVAQAIIQQTRALLPESEKNCDITHINADTLAECLKMLKKLK SEKYDPKPLVVEGMACPFGCIGGPGSLSSLNRAKIAVKKFAEEAENTLPS EYLKKK* 1005373994
232 NZ_CP042905.1_ B, M3a' MPLPRLNIFEKTRGLYTPVIDIRRRVLSAVARMVVENKPPTYIEHIPYQIIDK 15, amino acid DTPTYRESVFKERAVVRERVRLAFGMDLREFGAHGPINDDDVIYSLTDKK VIKKPIVNVIKVGCERCPEHSYIVTDLCRGCIAHPCTIVCPKNAVSIINNRSH IDQDLCIRCGKCEQVCPYNAIAYRERPCAAACGVKAISSDKDGFADIDQN KCVSCGMCIVSCPFGAIAEKSEIVQIISALKGKKPVVAEIAPAFVSQFGPLV TPGKLKGALKEVGFMDVREVAYGADVVVMNETEELVELIKKRENCKENE ESESNCRTFIGTSCCTSWALAAEKNYPELSKLNISESFAPMVEIAKKIKKD TPDAIVVFIGPCISKKEECFIPNVAEVVDFVMTFEELVAIFQAFTIDPTKIPPE KEMQDASSLGRGFPVAGGVANAVLEQTKAIIGQDVEIPIVSADTLKNCMS LLNQIKKGALDPKPLIVEGMACPNGCVGGPGTLAPLRRAQREVKKFAKKA QWEKPKDYL* 234 Alum_Rock_MS4 E, M1a MLIMKNGEKGMEELIPAFEKKIKLVAMLAPSFVAHFDHPSIILQLKKLGFDK _biofilm_2_scaffol VVELTFGAKIVNKEYHEILKKSKGLVISTVCPGVVETITSQFPQYRKNFIKIV d_236_3, amino SPMIATAIICRKIFPKHKTVFISPCNFKRIEASRSKYVDYVIGCDELDALFEK acid YEIKPVKTKRMIHFDRFYNDYTKIYPLAGGLSKTAHVRQILKPEETKTIDGI SELISFLKHPDKKVRFLDGNFCIGGCIGGPLLTKKRSLKEKKARVLKYVQW AVHERIPKSSKGVFEQAKGLDFSVKSF 235 CABMEU010000 E, M1a MEYGSPDFFGFAYAEDQKRLVSLLRESKKKNSKRKLCLMTAPSFVVDFD 038.1_43, amino FVDFVPKMKGLGFDKVTELTFGAKIVNQHYHKHIKEHFSDYAMNGRSIDK acid KFQNKFISSVCPATVELVKNRHPELVKYLMPFVSPMSAMAKIVKKNFPNH KIIFLAPCSAKKFEAARLLDGKNKRIIDVAITFAEMKQIVAKENPKAKNGSR KFDSFYNDYTKVYPLSGGLSSTLHCKGILHKNDVIVRDGCTELEKLFSKTP ERVFYDVLFCKGGCVGGPGIASRAPLFIRKQRVFKYRNFAKKEKMDGKQ GVDKYTKGLDFSREF* 236 CABMGN010000 E, M1a MVKKSPKLKFPLKGRKKYVVMLAPSYIVDFSYPEIIFALRKLGFDKVVELTF 008.1_137, amino GAKMVNREYHSILEHNLSAHGFWISSVCPGIVDLVSTRFPQYRKNLIPVD acid SPMIAMAKIVRKTYSKHGIVFISPCNFKKIEAKDSGVVDYAIDYSELMEIFR KKKISLESFSDHEKAHFDKFYNDYTKVYPLAGGLSKTARLKGLLKRREIKK IDGAEKVIEFLENPSIKTKFLDANFCEGACIGGPCIYSKKLSLRKRRRKVLK YLNQSKREEIPKTDKGLVKCAEGINFRRYDL* 237 CABMGO010000 E, M1a MKKEEKDLDSFKKALARKEKLVALLAPSFIADFDYPLIIYQLRARGFDKIVE 005.1_287, amino LTFGAKMINREYHKILKEQKDKLFISSVCPGVVETIRNKYPQYRKNLIPVLS acid PMTATAKICRKLYPKHKLVFISPCQFKKIETNQSEYVDFAIDYNELRKLFEE NKEEINKKIKILKKKKTNKIGFDKFYNDYTKIYPLAGGLSKTAHLKQVLNPG EEVKIDGILEVEKFLQNPDKKVRFIDVTFCKGGCLGGPCVLNKNLPDKRK RLMKYLNYSKKEKIEENKKGLIRKAEGINFSKVY* 238 CABMHR010000 E, M1a MKKSSDKLEFPLKGKYLAMLAPSFVADFSYPAIISQLKDLGFDKVVELTFG 001.1_3, amino AKMINREYHRILKNSKKLIISSVCPGIVETIKSKALEFKENLISVDSPMTATA acid KICRKIYPTHKIVFISPCNFKKIEAQQSEYVDYTIDYLELKEIFDKLKLTDKKY SKNTSSFDKFYNDSTKIYPLAGGLYKTANLNNILQQDESIIIDGIEDVLKFLN KPKKNIRFLDVTFCKGGCIGGPCTNSKLSLMEKKKKVLDYLKIADKEKIPN TKKGNIKQAEGIKFSSYYPNR* 239 CAIKMX0100000 E, M1a MNNKNDITKIEKALKNKKQKIAAMLAPSFVSEFKYPQIISQLRKLGFNNLVE 67.1_2, amino LTFGAKMINREYHKILEKSNSLVIASVCPGIVESIKNNPELKSYNKNIIPVNS acid PMIATAKICKKIYPKHKICFISPCHFKKIEAQNCECVDYVIDYNQLKQLFTKY 1005373994
NIKPSKEKDQFDKLYNDYTKIYPLSGGLSKTAHLKGILKKDQTKTIDRWKD VEKFLKEYKAGKNKSIRFLDVTFCKGGCIGGPCTNQKLSIAKKKQLVLKYL KQARHEDIPESKKGLINKAKGISFRN* 240 CAIKWG0100000 E, M1a MKKEMNSLNFPLRGKYVAMLAPSFIVDFSYPEIIHILKKLGFDKVVELTFGA 51.1_16, amino KMINRDYHKILSNSKELKIATVCPGVVELIKNKAPQYSKNLIQVDSPMIAMA acid KICKKVYPKHKIIFFSPCHYKKDESKKSKMIFKVIDYKELKEIIEKKKIKIDSK EKIHFDKFYNDYTKIYPVTGGLSKTAHLKGVVKKKEVESIDGVSNIIKFLEK PNKRIKFLDCNFCIGGCIGGPCINSKEKLRKKKRKVIKYLNQSKKEDIPETR KGLIRKAEGITFKK* 241 CAIPOI01000000 E, M1a MKSLNSLKQALAKKQKIVVMLAPSFVADFDYPEIIYQLKALDFDKITELTFG 2.1_147, amino AKMINREYHRILEESEKSGKKELFIATVCPGVVSFIKTKYLQYAKNLMSVD acid SPMIATAKICKKIYPKHKVCFISPCEFKKQESQDSEYVDFCIDYNELRKLIK ENKKPIKKSKTGFDKFYNDYTKIYPLAGGLSKTAHLKGVLKPGEEIKIDGIF EVEKFLKNPKKEVSFMDITFCKGGCLGGPCILNNNLKDKKEKLMHYLEVA KEEKIPKGREGLVEKAKGISFLKRY* 242 CG10_big_fil_rev E, M1a MKFKFDKKKKYLAMLAPSFVVDFNYPEVISQLRELGFDKVVELTFGAKMV _8_21_14_0.10_s NRCYHEKLKNSKELVIASVCPGIVESVRERFPEYVRNLIKVDSPMIAMAKI caffold_1732_4, CRKTYPNHKIVFISPCNFKKTEARKSKYIDFVIDYKELAVLLQVHKVRGKKK amino acid DCFDKFYNEYTRIYPIAGGLSKTAHLLGVLNKKEARVIDGILDVEKFLKKPN KKIRFLDATFCKGGCIGGQCVSSKLSLAGRKKKVLDYLKLALNEEIPRGKR GKINRAKGICFEIKTRASKTSIFDSHQKPKVSEIFDPIK 243 CG23_combo_of E, M1a MKKRVENLKFPLKGKYLAMLAPSFVVDFSYPSVVSQLKELGFEKVVELTF _CG06- GAKMINREYHKLLENSDNLVISSVCPGIVDFIKNKFPKYKKNLILVDSPMIA 09_8_20_14_all_ MAKICRKTYSKHKIVFISPCNFKKEEAKNSGIVDYVIDYKELKNLFLKYKINT 150_scaffold_495 KNKEVCFDKFYNDYTKIYPLAGGLSKTAHLKGVVKKSEVKVIDGINGVVKF 6_8, amino acid LEKPDKSIKFLDVNLCVGGCIGGSCINSKESIPERKKRVLGYLRLAREEDIP NARLGLIEKAKDINFNIKKF 244 CP045477.1_677 E, M1a MELNTELIDVLKTINKEKVICLLAPSFVVDFKYPKIILELRRIGFNKIVELTYS , amino acid AKLINKEIHKQILENKKKQYICGNCPSVVKYIENKYPELKKNIMDIASPMVIM ARFMKEKYPSHIIVFVGPCFSKKQEAKENKEVDYALTFKEINDMFLYAKKN GFYKKQENENKHFDKFYNDYTKIYPLSGAVAETMNTREILKPEEMIIADGI KEIDKAIMKFKENKKVRFLDMLFCTGGCVGGPGIISTETIEQREDRVIHYR DKSKNEKLGKNFGKFKYAENLSLKRK* 245 EastRiver_08_08 E, M1a MIKNNIPKIEKVLQNKKIKKLVMVAPSFLTDFNYPSLISQLKELGFDKVVEV _2020_HR_Quigl TFGAKMVNREYHRILCEENQKLWIATTCPGITETIKNNPELKVFEENLIPVD ey_12455_length SPMVAMAKISRKAYPKHKIFFLSPCHMKKIEAEKTKQIDFVIDYQQLRVLLD _7330_cov_52.47 KYKILSSNQHVQFDKFYNDYTKVYPLSGGLTKTAKIKSILKWREYKIIDGW 2990_7, amino QKVEKLLQNIKDKPKKYKNQRFLDVTFCEGGCIGGPCTNKELSIRKKRKL acid VINYLKQAKREDIPESKKGLIKKAEGIRFTQQ* 246 JAAZKV0100000 E, M1a MVELDSYGYPDFFGFAYSEDQLKTLRLLRNSLRYNESKVILMVAPSFVVD 01.1_4, amino FDFKKFVPLMRGLGFDLITELTFGAKIVNKNYHKYIKENKKTKIKFISSVCPL acid SVNLLKAKYPDFSRFLLPFDSPMIAMAKVLKKHYPKHKIVFVSPCSAKKIE SKQFNEKNKKTLIDVVITFSELKQIVAKEKPRLVGSNIFDSFYNEYTKVYPL SGGLTETLHTKHILEDEEMIFSDGCQNLQKLFGSNPDKIFYDILFCEGGCI GGNGIVSKMPIVWKKNKVLKYRKAASRIKEGKKVGVTKYYKGINFRREF* 1005373994
247 JAAZNJ0100000 E, M1a MDKKRDNNLKFPLKGKYIALVAPSFVVDFPYPKILSQLKELGFDKAVELTF 04.1_18, amino GAKIVNKEYYEEIKNSKKLMISSVCPGVVETIKNSFPEYKNNLLLVDSPMV acid ATAKICKKIYPHHKRVFISPCNFKRLEAKRTGFIDYVIDYKELKDLIFKHSFE YTKNKNKKSYKKEILFDKFYNDYTKIYPISGGLSKTLKVKKLLQKNEIKEID GIKKVIKFLKNPNPKIKFLDITFCKGGCIGSQFINSKIPISLRKIKVLKYIRKAN KERIPENRKGIFKKAEGLSFKSNYPINIYSF* 248 LacPavin_0920_ E, M1a MVDLNFSYSTDQKKVLSLLNEKQKVCLMAAPSFVVDFDYLSFVPLMKGL SED2_scaffold_5 GFDRITELTFGAKIVNEHYHKYIKENKGKKGYEKFISSVCPTSVEMVKNRH 97388_197, PELKKFLLPFDSPVISMAKILHKEYPKHKIVFLAPCSAKKIEAKNGRLISAAL amino acid TFREMKGIIEKEKPKKSGRSHLFDRFYNDYTKIYPLSGGLGKTLHSKDILK EGETVSRDGCADLLKLFETHSDKIFYDILFCKGGCIGGNGVASKLPLFLRK KKVLDYKKFADREKIPEKMVGLNRYTKGLSFETAF 249 LacPavin_0920_ E, M1a MKRKKEKKLAMLAPSFASEFDYPEIIGMLKNLGFDKVVELTFGAKMVNRE SED3_scaffold_1 YHKLLENSKELVITSVCPGIVSLIEGKFSKYKKNLAKIDSPMIATAKICKKVF 722076_9, amino PKHRLIFISPCNFKKIEAKKSKLIDGVIDYQELKQIFDKKRIKPKKGMWKFD acid RFYNDYTKVYPLPGGLSRTANLNGIVGFNECMIIDGAKEVENFLNKPDKSI KFLDVTFCKGGCIGGPFLSKTKSLGEKQRGVLKYINLAKKERISRGSKGEI KEANGIDFRK 250 LC_01_combined E, M1a MKRDDKKLKFPLKGKYLAMVAPSFVVDFPYSKILFQLKDLGFDKTVELTF _scaffold_270283 GAKMVNREYHKILENSKGLVISSVCPGIVETIKSKYPQYKNNLIPVDSPMT 4_31, amino acid AMAKICKKVYPSYKIVFIAPCNFKKIEAKSSKNIDYVLDYFELKEILEKNKLN KKKYGKKEIFFDKFYNDYTKIYPLSGGLSKTANLKKILKSDETKIIDGILEVE KFLENPDKKIRFLDVTFCEGGCIGGPLVISKLPIFLRKRKVLNYLKTADKEKI PLKRKGIVKEAEGISFRSEYPKQIIYI 251 Meg19_1012_Bin E, M1a MKKNNLPEIERDLKSKNVKFLAMVAPSFVAEFNYPSIVYRLKELGFDKATE _278_scaffold_57 LTFGAKMINRDYHRKLKNSKKLVIASPCPGIVMTIKNKYPKYFKNLIRTDSP 52_32, amino VVATGKICRKHYPQHKLVFISPCDFKKIEAENSEYIDYVIDYKQLRKLFRKY acid NVKPKKCEILFDKFYNDYTKIYPLSGGLGKTAHLKGVVKEEEILGIDGIRKV MKFLDNPDPKIKFLDVLYCVGGCIGGQHTSKKLTVAQKRKKVLDYLNFSK SEDIPEDRKGLIKKAEGIKFSGKCWFD 252 Meg19_1012_Bin E, M1a MNVFDGIEKKVGVQNLNFPLKGKYVAMLAPSFVVDFEYPSIISRLKALGFD _396_scaffold_36 KVVELTFGAKMVNREYQKQLKKSKKLLIASPCPGIVEIIKQKAPKYAKNLA 153_8, amino QIDSPVTATGKICRKIYPNHKLVFISPCHYKKLEVANSKYIDYTIDYKQLNKL acid FVKFKVPKFKTKSHFDKFYNDYTKIYPVSGGLGKTAHLKGVIKQDEILVMD GLNNILKFLQNPDQKKRFLDILFCKGGCIGGPCISCKLSIPARRKKVLDYLE KSKDEDIPDAKKGVIDKAKGINFLRKV 253 MFWP01000006. E, M1a MKKSAENLSFPLKEKSVAMLAPSFVVDFSYPNIISQLKGLGFDKVVELTFG 1_1, amino acid AKMINREYHKILEHSKELVISSVCPGIVETIKSKYPQYRKNLIPIDSPMIAMA KICRKIYPKHKVIFISPCNFKKIEAENSDYVDYTIDYKELKNLLKKHKKKKNN SETFDKFYNDYTKIYPLAGGLYKTAHLKNILKDDEAMVIDGIDRVMEFLDN PDPKIKFLDVNFCKGGCIGGPCINSKLPLILKKKKVLHYLKLAEKEKIPEES KGIIKEAKGISFKSDYLNK* 1005373994
254 MWBC01000030. E, M1a MSFGLPNFFGFAYSEDQLLVLRLLKESKKRGGSKVILMSAPAFVVDFDYK 1_8, amino acid DFCPLMKGLGFDKVTELTFGAKIVNTCYRKYIKENKDKQEKFIATVCPSSV ELIKNRYPYLKRFLLPFDSPMVAMAKVLKKNYPKHKIVFTSPCSAKKIEAK KAFYKKKQLIDAVITFSELKQIIAKERPKKKKVCHKFDSFYNDYTKIYPLSG GLGATLNKKGILKESEVVSTDGHKRLSRIMEKNLNKTFFDVLFCDGGCIG GNGVSSKLPIVLRKKRVLDYRNLSKRESMDGNHGLNKYFRGIDFSRKFD* 255 NJDN01000007.1 E, M1a MKKSGQLVFPLREKFVAMVAPSFVVDFPYPGIIHGLKKLGFDKVVELTFG _7, amino acid AKLVNQEYHKELKKEGFFISSVCPGIVNIVLEKFPEYKDNLLKVDSPMVAM AKIVRKTYPSHKVVFISPCFYKKEEAKNSGQVDFVIDYRELKVLFDNKKINL NLKKKIHFDKFYNDYTKIYPIGGGLSKTAHLRGVLKKGEVKVIDGISKVIKFL ENRDKKVRFLDCNFCVGGCIGGPYINSKDSLAKRKKRVRDYLKRSLKEDI PEPRKGLSKRAKGINFTTNFSH* 256 PNOQ01000025. E, M1a MPKNNIELILKELKHKKMVALVAPSFVADFEYPKIITQLEMLGFDKVVELTF 1_23, amino acid GAKLVNREYHRILRTSKELVIATVCPGIVEVVNKNYPKYKKNLIKVDSPMIA MAKICKKIYPKHKTVFLSPCDYKKIEASKSSYVDYVIDYEQLRKIFKDKNLE NVKEGKKIFDRFYNDYTKIYPLAGGLSKTAHLSGIIKPSEIKHIDGITKVMQF LKKPDKKIKFLDVNFCEGGCIGGLHTCKIPIAQKKKKVIAYLNKAKREKIPLP RRGTFEKAEGLKFNY* 257 PWLL01000017.1 E, M1a MEISKELVDLLKVLQTKRCVCLLAPSFVVDFKYPKIIKTLKQLGFSKVSELT _4, amino acid FAAKIINTEYKKQLKKTKKPIICTNCPSLVKTIENKYPEYKEYLANIASPMVV MGRFIKKHFKEKNTCVFVGPCITKKIEALENKKDIDYAITFKELKQMIDYAK KNKLLLDTKENQTTDFDKYYNDYTKIYPLAGAVAETMHAKEILSKNQTLCC QGPENIEKTLKKINKNTKFVDALFCHGGCVGGPGIISKKSIKKKENKVRKY RKDCKKIKIGQNLGKQKYATEISLKRN* 258 RBG_13_scaffold E, M1a MKKSAENLSFPLKEKCVAMLAPSFVVDFSYPKIISQLKNLGFDKVVELTFG _108_63, amino AKMINREYHEVLEHSKKLVISSVCPGIVETIKSKYPQYKKNLISIDSPMIAMA acid KICRKIYPQHKIIFISPCNFKKIEAEKSNYVDYAIDYKELKNILKKYKKLNNNS QITFDKFYNDYTKIYPIAGGLYKTARLKNILKDDEAIVIDGIDRVMKFLDNPD SKIKFLDVNFCKGGCIGGPCINSKLPLILRRKKVLDYLKLAEKEKIPEESKG VIKQASGISFKSDYLNK 259 RBG_16_scaffold E, M1a MKRGVENLSFPLKGEYVAMLAPSFVVDFSYPKIILRLRDLGFDKIVELTFG _3952_10, amino AKMINRYYHNKLEKTKELVISSVCPGVVETIKTKFPQYKKNLIQVDSPMIAT acid AKICRKIYPQHKIVFISPCNFKKTEAENSEYIDYVIDYRELGELLKRLGRRLD KKDTDLLFDKFYNDYTKIYPLSGGLSKTVHLKGILRSEETRAIDGMGDVILF LNNPDKKIKFLDVTFCKGGCIGGPCINSKLPLLLRKRKVLDYIKIADKEVIPQ GRKGLMKEAKGISFKSENVNK 260 rifoxya1_full_scaff E, M1a MKKNNINLILKELKHKKMVALVAPSFVADFEYPKILSQLEKLGFDKIVELTF old_175_29, GAKLVNREYHKILKSSKGLVIATVCPGIVEVVNKNYPKYKKNLIRVDSPMIA amino acid MAKICKKIYPKHKTVFLSPCDYKKIEANKSKYVDYVIDYEQLREIFKEKKLN NIKPSKKIFDKFYNDYTKIYPLAGGLSKTAHLNDVIKKNEIKHIDGIKKVMQF LDKPDKKIKFLDVNFCVGGCIGGSHTCNLSIAKKKKRIIAYLNKAKRERIPL SRRGTFERAEGLKFTY 261 S2_GD2017_2_m E, M1a MVELNSELKAVVDALNIEKTICLLAPSFPVDFEFPDIILDLRRMGFTKVVEL anure_scaffold_1 TYAAKLINYKYINIINENPEKQFICGNCPTIVKLIENQYPDLKDNILDVTSPM VVMARFVKREFGNDYKTIFVGPCFAKKTEAKENSDCVDYALTFKELIEIYD 1005373994
997_46, amino YCEEKNILKNIEENKANREFDKFYNDVTKIYPLGGGVASSMITKDILNLEQV acid IVCDGPKCIEKAMFDFTNNTKYRFVDILFCEGGCLGGPGIVCKDDLETKK QRLFDYKEKSRYYEPNENYGKFVHAFGLDIKRKK 262 SR- E, M1a MKKEMLEFPLKKEEKYIAMLAPSFVVDFSYPDIISQLKGLGFDKVVELTFG VP_26_10_2020_ AKMINRDYHKILGKSKGLVISSVCPGVVETIKSKFPEYKNNLIEVDSPMIAM 1_100CM_scaffol AKICKKNYPQHRIVFFAPCDFKKIEAGKSKYIDYVFDFTELKEILKQNKKRT d_4165154_7, NENLKFDSFYNDYTKIYPVSGGLSKTANLKDIIELKDVKIIDGIEEVSEFLKN amino acid PEKDIKFLDVTFCKGGCIGGPKINSKLPIVLRKTKVMNYMKVADKESIPEK RKGVIEKARGISFKSEYPA 287 AQRS01000037. A1, M1 MGSIEDVNAALADEGKMVMAQVAPAVRVTIGEEFGLPAGTIVTKKLVGAL 1_10, Ia, amino RQAGFEKVFDTSVAADIVTIEEGTEFLNRLEDQEDLPLLTSCCPASVFFVE acid, NTFPKFLHHFCTVKSPQQGMGSLIKTYYARRMKIDAKKNFVVAVMPCIVK KMEARRPEMEFDGVHNVDAVLTTKEAAALLKSKKADLNGANETGFDSLL GKASGAGQLFGQTGGVSEALLRFVAWKLEGKKARVLFKEVRGKKGFRE AEVKIGSRMLKVAVIDGLNNLRDLMSSEEKFHSYDVVEIMTCPGGCIGGG GQPESTPEKLEARKKALHNVDAMETVRVASENPEVKELYGSYLIEPGSKV ARSILHVSRICLKCD* 290 DALH01000010.1 A1, M1 MDDISFVKEVLADKSKTVIAQTAPSVRVTVCEEFGSEPSSEETGRVVAAL _13, Mu, amino RKLGFDNVFDTDFGADVTVVEESAELVRRLREGGALPVFTSCCPGWTRF acid CVKAFPELNEHLSGVKSPQQIVGALTKTYFAEKTGKKAGEIFVVGVMPCY SKKLESREEETEIPGAKQDVDFILTTKDFAKLIKQEGIDYHSLEPDSFDELL GTASGASTLFAGTGGVMESVVRAAYYELMGGKELLPFVEDRRPGYGGT KEFTIEVGLEKPLRVAVVNGYDEARKMCERVLKEKKEGGRTIDFIEVLACR GGCVGGPGQPGNDPQQVAKRAFGLRDLDAENTRVRNAHQNPDVKKLY REFLGETGGEKAHKLLHRKGTQEILETLK* 291 NZ_CP042905.1_ B, M3a' MPLPRLNIFEKTRGLYTPVIDIRRRVLSAVARMVVENKPPTYIEHIPYQIIDK 15, Ps, amino DTPTYRESVFKERAVVRERVRLAFGMDLREFGAHGPINDDDVIYSLTDKK acid VIKKPIVNVIKVGCERCPEHSYIVTDLCRGCIAHPCTIVCPKNAVSIINNRSH IDQDLCIRCGKCEQVCPYNAIAYRERPCAAACGVKAISSDKDGFADIDQN KCVSCGMCIVSCPFGAIAEKSEIVQIISALKGKKPVVAEIAPAFVSQFGPLV TPGKLKGALKEVGFMDVREVAYGADVVVMNETEELVELIKKRENCKENE ESESNCRTFIGTSCCTSWALAAEKNYPELSKLNISESFAPMVEIAKKIKKD TPDAIVVFIGPCISKKEECFIPNVAEVVDFVMTFEELVAIFQAFTIDPTKIPPE KEMQDASSLGRGFPVAGGVANAVLEQTKAIIGQDVEIPIVSADTLKNCMS LLNQIKKGALDPKPLIVEGMACPNGCVGGPGTLAPLRRAQREVKKFAKKA QWEKPKDYL* 294 AQRS01000037. A1, M1 MGSIEDVNAALADEGKMVMAQVAPAVRVTIGEEFGLPAGTIVTKKLVGAL 1_10, la, amino RQAGFEKVFDTSVAADIVTIEEGTEFLNRLEDQEDLPLLTSCCPASVFFVE acid, expressed NTFPKFLHHFCTVKSPQQGMGSLIKTYYARRMKIDAKKNFVVAVMPCIVK protein sequence KMEARRPEMEFDGVHNVDAVLTTKEAAALLKSKKADLNGANETGFDSLL which includes: GKASGAGQLFGQTGGVSEALLRFVAWKLEGKKARVLFKEVRGKKGFRE protein sequence AEVKIGSRMLKVAVIDGLNNLRDLMSSEEKFHSYDVVEIMTCPGGCIGGG +SSG+streptavidi GQPESTPEKLEARKKALHNVDAMETVRVASENPEVKELYGSYLIEPGSKV n tag ARSILHVSRICLKCDSSGWSHPQFEK* (WSHPQFEK) 1005373994
295 CABMGN010000 E, M1a MVKKSPKLKFPLKGRKKYVVMLAPSYIVDFSYPEIIFALRKLGFDKVVELTF 008.1_137, Na, GAKMVNREYHSILEHNLSAHGFWISSVCPGIVDLVSTRFPQYRKNLIPVD amino acid, SPMIAMAKIVRKTYSKHGIVFISPCNFKKIEAKDSGVVDYAIDYSELMEIFR expressed protein KKKISLESFSDHEKAHFDKFYNDYTKVYPLAGGLSKTARLKGLLKRREIKK sequence which IDGAEKVIEFLENPSIKTKFLDANFCEGACIGGPCIYSKKLSLRKRRRKVLK includes: protein YLNQSKREEIPKTDKGLVKCAEGINFRRYDLSSGWSHPQFEK* sequence +SSG+streptavidi n tag (WSHPQFEK) 296 CP045477.1_677 E, M1a MELNTELIDVLKTINKEKVICLLAPSFVVDFKYPKIILELRRIGFNKIVELTYS , Fm, amino acid, AKLINKEIHKQILENKKKQYICGNCPSVVKYIENKYPELKKNIMDIASPMVIM expressed protein ARFMKEKYPSHIIVFVGPCFSKKQEAKENKEVDYALTFKEINDMFLYAKKN sequence which GFYKKQENENKHFDKFYNDYTKIYPLSGAVAETMNTREILKPEEMIIADGI includes: protein KEIDKAIMKFKENKKVRFLDMLFCTGGCVGGPGIISTETIEQREDRVIHYR sequence DKSKNEKLGKNFGKFKYAENLSLKRKSSGWSHPQFEK* +SSG+streptavidi n tag (WSHPQFEK) 297 DALH01000010.1 A1, M1 MDDISFVKEVLADKSKTVIAQTAPSVRVTVCEEFGSEPSSEETGRVVAAL _13, Mu, amino RKLGFDNVFDTDFGADVTVVEESAELVRRLREGGALPVFTSCCPGWTRF acid, expressed CVKAFPELNEHLSGVKSPQQIVGALTKTYFAEKTGKKAGEIFVVGVMPCY protein sequence SKKLESREEETEIPGAKQDVDFILTTKDFAKLIKQEGIDYHSLEPDSFDELL which includes: GTASGASTLFAGTGGVMESVVRAAYYELMGGKELLPFVEDRRPGYGGT protein sequence KEFTIEVGLEKPLRVAVVNGYDEARKMCERVLKEKKEGGRTIDFIEVLACR +SSG+streptavidi GGCVGGPGQPGNDPQQVAKRAFGLRDLDAENTRVRNAHQNPDVKKLY n tag REFLGETGGEKAHKLLHRKGTQEILETLKSSGWSHPQFEK* (WSHPQFEK) 298 NZ_CP042905.1_ B, M3a’ MPLPRLNIFEKTRGLYTPVIDIRRRVLSAVARMVVENKPPTYIEHIPYQIIDK 15, Ps, amino DTPTYRESVFKERAVVRERVRLAFGMDLREFGAHGPINDDDVIYSLTDKK acid, expressed VIKKPIVNVIKVGCERCPEHSYIVTDLCRGCIAHPCTIVCPKNAVSIINNRSH protein sequence IDQDLCIRCGKCEQVCPYNAIAYRERPCAAACGVKAISSDKDGFADIDQN which includes: KCVSCGMCIVSCPFGAIAEKSEIVQIISALKGKKPVVAEIAPAFVSQFGPLV protein sequence TPGKLKGALKEVGFMDVREVAYGADVVVMNETEELVELIKKRENCKENE +SSG+streptavidi ESESNCRTFIGTSCCTSWALAAEKNYPELSKLNISESFAPMVEIAKKIKKD n tag TPDAIVVFIGPCISKKEECFIPNVAEVVDFVMTFEELVAIFQAFTIDPTKIPPE (WSHPQFEK) KEMQDASSLGRGFPVAGGVANAVLEQTKAIIGQDVEIPIVSADTLKNCMS LLNQIKKGALDPKPLIVEGMACPNGCVGGPGTLAPLRRAQREVKKFAKKA QWEKPKDYLSSGWSHPQFEK* Table 3: Nucleotide sequences of HydC 1005373994
SEQ Description Subunit Sequence ID NO: 301 AQRS01000037. HydC ATGTTTTCTTTCAAAAAGCACGTCCTCGTATGCACCTCGGAAAAGCCG 1_14, nucleotide (NuoE-like) GGGCATTGCGCCGAAAAAGGCGGCCCTGAATTGCTCGCTGCTTTCC GCGAGGAGGTTGCGAAGCGCGGCCTGCAGAACGAAATTTATGTTACC AAAACCGGCTGCACCTCGCAGCATCACTGCGGGCCTACAGTAATAAT TTACCCTGACGGCGTGTGGTACAAGTTGGTTACAAAAGAGGACATTC CTGAAATAATCGAGTCGCACCTGCTCGGCGGAAAAATTGTTGAAAGG ATCCTGAACAGGGAAATAGGCCTGTTCAGGAAACAGGCATAA 302 09mi20_z1_2019 HydC ATGCATTCAAAATGGTTAGAAAGCTATTTGAAAGAGAATGAGCACCAG _ig18392_10016_ (NuoE-like) AGCTTGCTTGCAGTACTGCTGAGGATACAGGAAAGGGAAGGCTTTCT 4, nucleotide ATCAAAAGAAAGCCTTGAGCATGTATCCAAAGGAATGAAGATCCCTTT AAGCAAGATTTATTCTGTTGCAACTTTTTATTCTGAATTCAAGCTGGAA AAGAGGGGAAAACACATAATTAAGCTCTGCTCAGGGACAGCTTGCCT TGTTAAGGGAAATAATGTGAATCTCAATTTCCTGAAAACTGTGCTGAA TCTAAAACCAGGGCAAACTACCCCGGATAAATTATTCACGCTTGAAAC TGTTAATTGCCTGGGGACATGCAGCTTAGCACCAGTAATCAATATTGA CGGAAAGATTTACCCGAATGTGACTATTGAAAAACTGAGTGAAATAAT TGAAAAGCTGAAAAGGTCGAAAAGATGA 305 AQSC01000060. HydC ATGGAAAAAGACATTGAATCCATTATCTCAAAGTATTCTGGTGCATCG 1_11, nucleotide (NuoE-like) GATCTAATCGACGCCCTGGAAGACGTTCAGGAAGCTTTCGGACACAT CTCTGAAGACAACATGCACAGCATAAACCAGGTACTGAAAATCCCGC TAGTCGACATTCTCGGAGTCGTCTCATTCTACTCAGCCTTCAAGACCA AGCCCCCCGGAAAGCACATAATAAGAATCTGCCGTGGAACAGCATGC CACATTAAAGGATCAACCATCCTTGAAGAGCACTTGGAAGACAAACTC GGCATAAAAGCGGGTGAGACCACTGAAGACGGGAAATTCACGCTCG AACCCGTAAACTGCATTGGAGCGTGCGCGAAAGCCCCCACCATGATG GTCGACGACATTGTGTACGGGGATCTGACAAAGGAGCGAATCGATGA AATACTGGGGGAGTATAAATGA 306 CAITEO0100000 HydC ATGGTGGAACTTAGAAGGGTAACTATTAATGGGAATCATTATTATTATT 66.1_1, (NuoE-like) TGTTCCATGAAATAAGAGAGAACGGACAGTTCAAAAAGTTCAGGCATT nucleotide ACATTGGCGCGCAAGAGCCAGATAATGCAATGCAGCAAAAACTTGAA CGAGATTTCATGGAGGATATCAAGAATAACCCTGATAAGTATAGCCCA AAAGAGAAGCAAAATATCATCGCCGTCCTTCAACAGATAATGGAGAAG GAGAATTATATTTCTGAAGAAAACTTCGTGCGTCTCTCATCAGAGCTT GATATTCCTTTGGTAAACCTTGTTGGTGTTGCGACATTCTATTCTCAGT TCAGACTCACAAAGCCAGGCAAGCACACAATTAAGATTTGTGATGGAA CTGCCTGTCATGTAAAAAACTCTGCTGCACTTCGAGTGTTCCTTGAAG AAACACTCGACATTAAATCTGGTCAAGTTACTAAGGATGGTAACTTCG GCCTGGAAGTTGTAAATTGCATTGGCGCGTGTGCTCGTGCTCCTTCA ATGATGATTGATGAAACAGTTTATGGTAAACTTGACAAGAAAAAAATCA AAGAAATAATTGGAACGTACAAATGA 310 Candidatus_Lokia HydC ATGCCTACAAAAATATCTGACATTTTATCCCATTTTACCCAAGGAAATT rchaeota_archae (NuoE-like) CTTCAGAACTTATTCCTATACTTCAAGCGGTGCAGGCAGAATATGGGT 1005373994
on_FW102|JAIZ TTATTTCTGAAGAATCAGTATATGAAATTTCGGAATATCTTAACATTCC WK010000001_1 CAGTTCAAAAATTTACGGAGTAGCAACATTCTATGCACAATTCCGCCT 633, nucleotide ACAACCTCCTGGACGACATGTCATCAATCTATGCACTGGAACTGCTTG TCACGTAAAGGGATCAGAAAAATTAATTCCAGTGTTTGAGCAGGAATT AAAGTGTAAGGCAGGAGAAACGACTAAAGATGGTCGTTTTACACTAAA TTTAGTTGCATGCTTAGGTGCTTGTGCGTTAAGTCCTGTTGTGAATAT CGATTCTGATTTTTACGGTAATTTGACTGCTGGAGAAATTAAGAAAATC TTGCGTAAATACAAATAG 312 Candidatus_Lokia HydC ATGTCGATTTTATCCTTTGAAAATCGAAATAAAAAACTTTCTGGAGGTT rchaeota_archae (NuoE-like) TAAAAAAACTAATGCCAGCAAAAATATCAGATGTCCTCTCCCGTTTCA on_strain_B53_G AGCAAGGAGATGCTTCAGAATTGATTCCAGTGCTCCAAGCCGTACAG 9|QMYW0100007 AGTGAATACGGATACATCTCTGAAGATAATACCTATGAAATAGCAGAA 8_8, nucleotide TACCTAAATCTGCCATCATCAAAAATCTATGGAGTTGCAACATTTTACG CTCAATTTCGACTCGAGCCATTAGGTCGCCATGTAATTAACTTGTGCA CCGGCACAGCATGCCATGTTAAAGGGTCGGAAAAATTGATCCCGGTT TTTGAACAAGAATTAAAGTGTAAAGCAGGAGAAACCACAAAGGATGGT CGATTTACTTTTAATTTGGTCGCTTGTTTGGGAGCTTGTGCATTAAGC CCAGTAGTAAATATTGATTCCGACTTTTATGGAAACGTTTCTCCAAGC GATGTCAAGAAAATCCTACGGAAATACAAATGA 315 Candidatus_Lokia HydC ATGCACTCAAAAACCATCACGGAGATTCTAGTTCCGTTTCCAAGAGAT rchaeota_archae (NuoE-like) GAACCTTCTGCTCTTATTCCAGTATTGCAAGCGGTTCAATCCGAATAT on_strain_HM1_B GGATATCTGTCTGAAAACAATATATATGCAATTTCCGATCATCTTAACG 6_4|JABCSI0100 TACCCAGTTCGAAGATTTATGGCGTTACAACCTTTTATGCGCAATTTC 00064_8, GACTTAAACCCCTTGGTCGTCATGTCATTAATCTTTGTCAAGGCACTG nucleotide CTTGTCATGTTAAAGGTTCTGAAAAATTAATTCCTGTATTTGAACAAGA ACTCAAATGCAAAGCAGGCGAAACAACCAAGGATGGAAAATTTACCTT CAATCTGGTAGCATGCCTTGGAGCTTGTGCTTTAAGCCCTGTAGTCAA TATTGATTCAGATTTCTATGGAAATATTAAAATAGCCGACATTAAAAAA ATTCTTCGGAAGTATGATTAA 317 DAWM01000035. HydC ATGAAAAGCAAAATCTTTGGGAATAATACCAAGATAATGAAAACTGAA 1_7, nucleotide (NuoE-like) AATACTATCATAATAATCTTATTATTCGCAATCGGAATAACGCTCTTTA CTCTCTCTGGATGTAATAAAGAACAGATTAATCCTGCCAGCGAAATAA ATAACTCATCCAACTTGTCTAGCAACCAATCCAACAACCAATCTTCTA GTGGCTCAACACAGCTCTTAGAGGTGTGTATTGAGCAGACCTGCTTT GAAGCGGAGATAGCGGATTCTCCAAAAGAGAGGGCGAAGGGCTTGA TGTTCAGGGAGAAGCTGGAAGAGGACAAAGGGATGCTCTTTGTTTAT CCTGAAGAGAGAACATATAACTTCTGGATGAAGAACACCTTAATTCCC TTGGATATTATCTGGATTAGCGCTGATAAGAGGATAGTGCATATTGAG GAGGCAGTGCCTTGCGAGGAAGAGCCATGCAAGATTTATAGCCCTGA GAAGGAAGCGCAATTCATTTTAGAGATTAAGGGAGGAATGGCTGAGA AGAAAGGTATTGAAATCGGAGATGAGGCAGGGTTCAACCTAAACATTT AA 319 DSAL01000038.1 HydC ATGATGGCGAAGATTGATGACATACTCGCGAAATACGATTCAGCGCA _4, nucleotide (NuoE-like) CGAGCTTATCGACATGCTCGAAGACATCCAGGCCGAGTACGGGTACA TTTCTGAAGAGAACATGCGCAAGGTCGAGCAGGACTTAAAGATCCCG 1005373994
CTAGTCGACATATACGGGGTCGTCACATTCTACTCCGCATTCAAGCTC AAGCCGTCGGGCAAGCACACCATCAAGGTTTGTACGGGAACGGCGT GTCACGTGAAAAAGTCCGATTCACTGAAGGAGCATTTGATGAAGGCG CTCTCAGTCAAGGAGGGTGAGACGACAAGCGACGGCAAGTTCACGC TTGAATTGGTCAACTGTATTGGAGCGTGCGCAAAATCCCCGGCCATG ATGATAGATGAAAAGGTTTACGGCGAGCTTACCGCGAAAAAGATCGA CTCGATATTGAAGGAATATTAG 321 DSBS01000042.1 HydC ATGGAAAAGGAAGGGAAAACTGTGATTGCTGGCGAGTCCAGGGTTCT _18, nucleotide (NuoE-like) TGAGATGCTCCAGGAGATTAACAAGAATGAGGGTTACATTTCCAGGG AAAGGCTTTCTTCGATAAGCAGGGAGCTTGGAATTCCGCTATCCCTG CTTTATGGCCTTGTTACATTCTACAATCAGTTTAACCTTGCAACAAGCG GGAGATATACTATTGAGGTTTGCGAGGGAACTGCCTGCCACATTAAC AAAAGCAGTGAGATAAAAAGAGCAATAAAGGATGCTGCAGGGATTAG TGTTGATGAGACTTCATCTGACGGCTTGTTTACTTTAAGGGATGTAAG GTGCCTTGGCGCCTGTGCGCTTGCCCCTGTGCTGAGGCTAAACAAAA AAATTTACAGCAAGATGACTTATGAAAAAACCAAAGAGCTAATTTTTAA GCTTAAGAAGGAAGCGGAGGGTGAAGCTAAGCTAAAATGA 323 DSVV01000048.1 HydC ATGAAAAAAGAGGGGTTCAATCTTATTGCTGAACTACAGGCTGTCCAG _6, nucleotide (NuoE-like) GACAGATATGGCTATCTTCCAATGGACGTCTTGAAAACCCTATCAAAA GAGCATAAAATCCCTGGAACTGAGATATTTGCAGTCGCCACATTCTAC AATCAGTTTAAGTTCGATAAGCCGGCAAAACACACCATCCAGGTATGC ACAGGGACTGCCTGCCATGTCAAAAGATCGGCTGACCTGTTGAGCCA GATACAAAAGACACTGAAGATTAGGCCAGGAGAAATCACAAAAGATG GACTGATAAAGCTTGAGACAGTAAACTGCATTGGTGCATGCGCTAAG GCTCCGGCTATGATGGTGGATGATAAGGTCTACGGGCTTGTTGACCA GGATAAGCTTAGGCAAATACTCGGTGCTCTCAGATGA 325 DUIB01000037.1 HydC ATGAATAATCATGAATCATTAAAACCAGATATTCATAAAGTCTTAAAGA _10, nucleotide (NuoE-like) AATATGATTCTAAGAAAGACATTATACCTGCATTACAGGATATCCAGG AAAATTTTGGATTTGTCTCAGAAGAGAATGCAGAGAGCCTAGCAAAAA AAATAAATTCTCCATTAGTTGATATAAGCGGCGTTGTAACATTCTATAA TATGTTCAGATTAAAGCCTGTCGGAAAATATCATATTGCTATCTGCAG AGGAACTGCCTGCCATGTACAGAATTCTGAGGAATTGTTGAAATATGT TGAAAAAAAGCTTAAGATAAAGACAGGAGAGATAACTCAAGATGGAAG ATTCAGCCTTGAAGCTGTGAATTGCATTGGTGCATGCGCAAAGGCTC CTGCAATGATGATACATGATAAGGTTTATGGGCAATTGACAGAAAAAA AGATAGATGCAATACTTGACTCAATGAAATGA 328 LacPavin_0920_ HydC ATGCTGATGTCTGAAAAAGCTGATGCAAAAAAAGAGGAACTTGCAAAG SED5_scaffold_1 (NuoE-like) AGCCCTTCATCCTATGAAATAGAGGAAAGTGTGCTTTCAATAATTGAC 418049_11, TCCCGCTCTGAGGAAAAGAGCCCGCTTCTGCCCATACTTCAGGATGT nucleotide GCAGAAAGCATTTGGCTTCATATCGCCGCAAGCCATGCTTGAGATAA GCGAAACCCTTGATATCCCTCTCTCGCAGGTTTATTCTGCAGTCACAT TCTACAATGAGTTCAGGACAAAGAAGCGCGGGAAGCATCTCTTCCGC GTCTGCATGGGAACTGCCTGCTGCATAAAGAAGGCAGATGCTGTCAT TCTTGAGCTTGAAAAGGAGCTCGGCATTCAGTGCGGCCAGACAGATG GGAATGGGCTTTTCACGCTGGAAACAGTCAACTGCTTCGGCGCCTGC 1005373994
GGCCTTGGCCCGATTGTTGAAGTAGATGGAAGGATTTTTTCATTGGTT GAGCCAAAAAAGGCAAAGGCTCTTGCAGAGGAAGTGAAGAAGAAAGA GAGGCTGCAGAAATGA 329 PEXD01000050.1 HydC ATGCTGATGTCTGGAAAAGCTGATGCAAAAAAAGAGGAAACTGCAAA _15, nucleotide (NuoE-like) GAGCCCTTCATCCTATGAAATTGAGGAAAGTGTGCTTTCAATAATTGA CTCCCACTCTGAGGAAAAGAGCCCGCTTCTGCCCATACTTCAGGATG TGCAGAAAGCATTTGGCTTCATATCGCCGCAAGCCATGCTTGAGATG AGCGAAACCCTTGATATCCCTCTCTCGCAGGTTTATTCTGCAGTCACA TTCTACAATGAGTTCAGGACAAAGAAGCGCGGGAAGCATCTCTTCCG CGTCTGCATTGGAACTGCCTGCTGCATAAAGAAGGCAGATGCTGTCA TTCTTGAGCTTGAAAAGGAGCTCGGCATTCAGTGCGGCCAGACAGAT GGGAAGGGGCTTTTCACGCTGGAGACAGTCAACTGCTTCGGCGCCT GCGGCCTTGGCCCGATTGTTGAAGTAGATGGAAGGATTTTTTCATTG GTTGAGCCAAAAAAGGCAAAGGCTCTTGCAGAGGAAGTGAAGAAGAA AGAGGGGCTGCAGAAATGA 331 PEXL01000024.1 HydC ATGTATTCAAAATGGTTAGAATCATGGCTAAGGGAAAATGACCACAAA _9, nucleotide (NuoE-like) GGTCTGCTTGAGATTCTGCTAGAGGTACAGCACAGGGAACAATTTTTA TCACGGGAAAATATTGAATTCGTTTCAAAGGCAAAAGAGATACCGCTA GCCAAAATTTATTCAGTAGCTACTTTTTACTCAGAATTTAGATTAGACC AGAGGGGAAAGCATGTGATCAGATTATGCGCTGGCACGGCTTGTTTG GTAAAAGGAAATAATGTTAACCTAAATTATTTAAAAACAGAATTAAATC TAATGCCGGGAAAAACAACTGCAGACAATTTATTTACATTAGAGGGGG TAAATTGCCTGGGCACATGCAGCCTTGCCCCTGTTGTGAGCATTGAT GGAAAAATTTACCCGAATGTTACAATAGAAAAACTTTCAGGGCTAATT GAAAAGATAAAAAAGGCGGAAAGATGA 333 PGXE01000078. HydC ATGGACAAGCTTGATGCGATAATAAAAGAACACGGTAAATCCCCTCTT 1_2, nucleotide (NuoE-like) CCGGTCCTAAAGGCCGCGAAGGCGGAGTACGGCCACTTATGCAAGG ATGTTCTCGAGGCTATTTCGGAGAAGATCGACGTGCCTGTCTCAAGG CTGCACGGCGTCGCGACCTTCTACTCGATGCTCGGAACCGAACAGAT TGGCGAGAACGTCATCTACGTATGCAACTCCCCGTCCTGTTACGTGA ACGGTTCCCTCAACGTCCTCGAGGAGTTCGAACGCCGCCTCGGCATA CGCTGCGGCGAGACCACGGCTGATGGCGGGGTAACGCTAGAAAAGA CCGCCTGCATCGGCTGCTGCGACATGGCGCCGGCGATCCTCCTCAA CGGCGAGCCGTGCGGGCCACTTGGCAAGAAGGACATCGCAAGGATC GTAAAGAGCATGAGGAAGTGA 337 PXDW01000025. HydC ATGGTTTTTTTAAAAGAGGTGCAATTAGGAAATAATAGCTATTTCTTTC 1_23, nucleotide (NuoE-like) TTTTCTATTCAATACAAAATGAAAAATTCAAGCCTTATATTAGATATATA GGGAAAAAAAAGCCGAATACAGAATATCTTAATTCCCTAAAAAAGAAA TTTCTTAAAGATGTCAAATCCAATCCAGAATTATTTGAAAAAAAACAGA AAAAAAATGTAATAATAATGCTCCAGGAAATACAGGAAAATGAAGGAT ACATATCAGAAGAAAATATAATAAGACTGTCAAAAGAGATTAATATACC TGCCACCCATATTTATGGTGTTCTGACATTCTATACATATTTTAGGTTT AATCCTCCGGGAAAATATAATATTGCTGTTTGCAACGGGACAGCATGC CACGTAAAAAACTCCGTGTCTTTAATAAGATATATAGAGAATATCCTTG ATATAAAAGTAAATGAAACAACCAAGGACAAAAAATTTAGCTTAGGAT 1005373994
CTGTCAATTGTATAGGGGCTTGCGCAAAAGCCCCGGCTATGATGATT AACAATACTGTATATGGTGATCTTGACGAAGAAAAGATTAAAAAAATTT TGGATGGTTTGGAATAA 338 SR- HydC ATGAATACAGCTAGTTCGGCGATTTTTATGGATGAAAGCAGAGAAGCA 2_scaffold_141_4 (NuoE-like) GACATCATAGGAAGATTGCTTGAGGTGCAGAAAAAAAACGGATGGCT 160621_8, TCCTAGAAAGGAAGTTGAAAGAATAGCAAAAGAAACAAGGACTCCGC nucleotide TTGCAAAAGTATATGGGATAGCGACGTTCTACGATTTTTTTTCCTTGAA TCTGCATAAAAAGAAGGAAGAGATCCGCAAATGCTTGAACCACGACT GTAGGCTGAAGGGAAGCGAAATAATTGAAAGCAGTTGA 340 AB_3033_bin_10 HydC ATGAGCTTTAGGGATATTGATGAGATCATAAAGGACAGAAAAGACAAT 1_scaffold_9312_ (NuoE-like) CTTCTTCTGCCGATGCTCGAGGCCATTCAGGCCAAGTTCGGCTGTGT 3, nucleotide TTCTGAGGAAAATGCACACTATTTAAGCAGAAAAACAGGAATTCCATT CTCAAAAATATATGGTGTTATAACTTTTTATGAAATGCTTTATACAGAG CCAAAAGGCAAGTACATTATAAGGATCTGCAACAGCCCTTCATGCTAT CTTAATGGTTCATTAAAATTAATTGAGTTTCTTGAATCATTATTGAAAAT AAAGTCAGGAGAAACAACAAAAGACAAAAAATTCTCTTTAGAGATTGT GTCATGCATTGGCTGCTGTGATAAAGCTCCGGCAATGATGATTAATGA TAAAGTATATGGAAATCTTGATGAGAAAAAAATAAGGAAAATAATCTCT AGCTTAAAATGA 345 AB_1215_Bin_11 HydC TGCAACAGCCCTTCATGCTATCTTAATGGCTCATTAAAATTAATTGAGT 1.fna_scaffold_95 (NuoE-like) TTCTTGAATCATTATTGAAAATAAAATCAGGAGAAACAACAAAAGACAA 187_7, nucleotide AAAATTTTCTTTAGAGATTGTTTCATGCATTGGCTGCTGTGACAAAGCT CCTGCTATGATAATTAATAATAAAGTTTATGGCAATCTTGATGAGAAAA AAATTAAGAAAATAATCTCTGGCTTAAAATGA 346 CABMGE010000 HydC ATGAAAAATATTCTAATAAATAGGCTTCGCGAAATCCAGAACAAAGAG 002.1_21, (NuoE-like) GGCTATGTCTCTGAAGAATCTTTAAAAAAACTAAGTCTTGAACTTAAAA nucleotide TTCCTATCTCCCAACTTTATGGGGTTGCAACATTCTATTCAATGATTTA TACAAAAAAACAAGGAAAATATGTAATTGAACTTTGCGCTTCTCCCTCA TGTTTTTTAAACGGATCCTGGAACCTTGAAGATTATTTAAAAAAAGAAT TAAAAATTGATATTGGAGAAACAACAAAAAACAAGAAGTTTAGTTTAAA GAAAACTTCCTGCATTGGTTGCTGTGACAAGCCACCTGCAATGCTACT TAATGGAAAAGTCTATACGAGTCTTACTGAAAAAAAATTAAAAGATATT TTAAAAAAATGCAAATAA 349 CAITKI01000006 HydC ATGGCAAAGGAAAAAGGCAAGAAGGACACCCGCCCCTTGATGAACAT 6.1_57, (NuoE-like) GCTCCATGAGGTGCAGGAAAGGCACGGCTACATCTCGGAGCACATG nucleotide CTCAAGCAGATTTCGGTGGATCAGGACATCCCCATTGCGCGGCTTTA CGGAGTGGTGAAGTTCTATACCATGTTCCACACCGAGCCGCAGGGCA AGTATGTGATTGAGATTTGCGGCTCGCCCTCCTGCGTGCTCAACAAC GGCGTCCGCTTGGAGAAGTTCCTGGAGAATGAAATCGGCGCAGGCA TAGGAGAGACGAGCAAGGACGGGCTGTTTTCCCTTTACAAGACGTCG TGCATAGGGTGCTGCGACGAGGCGCCCGCCATGCTGATAAACGGGG AGCCGCACACCAAGATGACAGTGGAGAGGCTCAAGCTCATCCTGAAA AAGCTTCGCGACACCGAGGCGGCGCAGGCCGGGGAGAAAAAGAAGT GA 1005373994
351 CAITNU0100000 HydC ATGGGCAAAAAGCTTAGCGAAAACCTCAGCGAAAAAACTTGCGCTGA 81.1_7, (NuoE-like) AGGCGTGAGCAAGGCGGACCAGCCGAAGGTGCTTGCGCTTTTGCGC nucleotide GAAGCTCAAGAGCGCGACGGGTACGTCACGCACCAGGCTGTCGAAC GCATTTCGAAGCAGACGGGCTTTACGGAGTCCGAAATCGACGGAGTC GCCTCTTTTTACGCGATGCTTTACTTGAAGCCGGTCGGCCGCTTCATC GTGCGCGTTTGCGCTTCGCCTTCCTGCGTAGTCAACGGCGGCGGCC GCGCGCTCGAGTGGGCGAGCGAGGTTTTGGGAGTGAGCGACGGCG AGACGACGAAAGACGGTTTGTTTACGCTTGAAGCCGTCAGCTGCTTC GGCAGGTGCGAAACCGCGCCGAACGTCATGATAAACGAAGAGAACT ACGGCGGAATCGATTCAAAAGAAAAAATGCGGAACCTGATAGAAAAA TTACGGCGCGAGGCGGCAGCGCGCGAGGCGACTAAATAA 352 JAACWB010000 HydC ATGGTGGAAGTTTTGATGAATAAACTTAGGGAAATTCAAGAAAAAGAG 015.1_3, (NuoE-like) GGTTTTCTTTCTGAAAAGTCCTTAAAAGAATTAAGCAATGAAATGAATA nucleotide TCCCAATTTCTAGACTTTATGGTATGGCAACATTTTATTCAATGTTTCA CACAAAAAAACCAGGAAAAAACATAATTGAAATTTGTGCTTCACCTTC GTGTTTTCTCAATGGCGGATTAACTCTTGAAAAATTTTTGATAAATGAA TTGAAAATTGATATCGGTGAAACTACAAAAGACAATAAATTCACTTTAC TAAAAACTTCTTGTATTGGTTGTTGTAATATTGCACCAGCGATGCTTCT AAATGGTAAACCAGTTGGAAATTTAACAGAAAAAAAATTAAAAAAGATT TTAAAAAAATGCAAATAA 355 JACCLF0100000 HydC ATGAAAGGGAAAAACACTCCTTTGCTCATTAACATCCTGCATGAGGTG 48.1_27, (NuoE-like) CAGGATAAGCAAGGCTATATCTCGGAACAGGCGCTCAAGAAAATTTC nucleotide TGTAGAGCAGAATATACCCATCTCGCGGCTTTTTGGGGTTGTCAAGTT CTATACAATGTTCCATACAGAGCCTCAGGGGAAATACGTTCTGGAAAT TTGTGGCTCGCCGTCATGCGTTCTGAATAACGGAATGAAGCTTGAGA AGTTCCTTGAAAATGAGATTGGCGTGAGAATAGGCGAAACCTCCAAG GACGGCATGTTCTCCCTTTACAAAACTTCCTGCATCGGGTGCTGCGA CGAGGCTCCTGCAATGCTCATAAACGGAAAGCCTTACACTAACATGA CTGTAGAGAGGGTGAAGCTGCTACTGAAGAAGCTGCGCAAATCAAAA AAGAAGTGA 356 JACCLG0100000 HydC ATGGAAAGGAAAAAAGGCACCAGCTTGCTTATGAATATCCTGCATGA 46.1_7, (NuoE-like) GGAGCAGGACAAGCACGGCTACATCTCCGAGCAGACTCTCAAGAAAA nucleotide TTTCAGTAGACGAAGGCATACCAATTTCACGGCTTTTCGGGGTTGTCA AGTTCTACACAATGTTCCATACAGAGCCACAGGGAAAGTATGTTGTTG AAATTTGCGGCTCCCCGTCATGCGTCCTGAATAACGGAGTGCAGCTC GAAAAGTTCCTTGAGAAGGAGATAGGCGTGAGAATAGGAGAGACTTC CAAAGACGGGATGTTCTCCCTCTACAAAACTTCCTGCATCGGCTGCT GCAACGAGGCTCCGGCAATGCTCATAAACGGCATGCCATACACCAAG ATGACTGTGGGGAGGCTGCAGCTGCTCCTGAAGAAGCTACGTGCAG CTTCAAAAAAGAAAAAGAAGTGA 359 Meg22_1012_Bin HydC ATGATTGCAAAAAAAGTTTTATTGAATGAATTGGGAAAAGAACAGAAA _224_scaffold_45 (NuoE-like) AAAAAGGGCTTTGTTTCCAAGAAAAAATTGAAAGAAATCGCCAAAAAT 14_12, nucleotide GTTTGCTTGCCGGAAAGCGAGGTTTTTTCCGCTGCAACTTTTTATTCC TTTTTATCATTGGAGAAAAGAGCAAAGCACATAATCCAGGTTTGCAAT TGTCCTTCTTCTCACTTGCATGGCTCTGATCGAATAATGAAATACTTG 1005373994
GAAAAAAAACTAAAAGTAAAAGCAGGTCATGCAACAAAAAACAAATTG TTTTTTTTAAGTGAAACTTCTTGTGTTGGATTGTGTGATAAAGCACCAG CAATAATTGTTGACGGAAAGCCGTATGTTAAAGTAAATGAAAAAAAAA TTGACAAAATTTTGAGGAAACTTAAATGA 362 MWBV01000007. HydC ATGGTGAAAAAACAAAACAAAAAAGCAAATAAAAAAGAAGATATTTTG 1_12, nucleotide (NuoE-like) CTTAATATCTTTCATGAAGAGCAAGACAAAAAAGGTTATATCTCAATAG ATTTCTTAAAAAAAATAAGCACAAAATATAATATTCCAATTTCAAGACTT TATGGTGTTGTTAAATTCTATACTATGTTAAGAACAGAACCTCAAGGAA AATACATAATAGAATTATGTGGTTCTCCAACTTGTGTTTTACATGAAAG CAGAGAAATAGAAAACTTCCTTAAAAAAGAACTCAAAATAGATATAGG AGATACAACAAAAGACAAAATGTTCTCTGTTTACAAAACATCTTGTATA GGATGTTGCGATGAACCTCCTGCAATGCTTCTCAATGGTAAACCAATA ACAAATTTGACTATTGAAAAAGTAAAGAAATTAATCAAGGAATTAAAAT CAAAGAAAAAATAA 364 NJBG01000001.1 HydC ATGAATTACTTGTTGCTTCCATTACTTAAGGATATACAGAAAAAAAAGA _784, nucleotide (NuoE-like) GATATATTTCTGAAAAGGATATGAAGAAGTTGAGTAAAAAAACAAGTAT TCCTATTGCTAAGATATATGCTACTGCTACATTCTATTCAATGTTGCAT ACAAAAAAACAAGGGAAATACATTATTGAAATTTGTGATTCTCCTTCTT GTTATGTTAATGGTTCTATTGATTTGATTAAATTTCTTGAGAAAAAATTG AAGATAAAGTCTGGTGAAACAACAAAGAATGGAAAATTTAGTCTGCAT ATTTGTTCTTGTATTGGTTGTTGTGATCAGGCGCCTGCTATGAAGATA AATGAAAGAGTTTATGGTAATTTAACAAAGAAAAAGATAGAGGAAATA CTAGATAAATGCAAATTTTAA 365 PCYE01000023.1 HydC ATGAGCTTTAGGGATATTGATGAGATTATAAAGGAGAGAAAGGATAAT _6, nucleotide (NuoE-like) TTTCTTCTGCCAATGCTTCAGGCCATCCAGGCCAAGTTTGGCTATGTT TCTGAGACAAATGCACACTACTTAAGCAGAAAAACAGGAATTCCATTC TCAAAAATATATGGTGTTATAACCTTCTACGAAATGCTTTATACCGAAA AAAAAGGCAAATACATTATAAGAATATGCAACAGTCCCTCATGCTATC TTAATTGCTCATTAAATTTAATTAAATTTCTTGAATCATCACTTAAAATA AAATCAGGAGAAACAACAAAAAACAAAAAATTTTCCTTGGAGATTGTTT CATGCATCGGCTGCTGCGACAAAGCACCGGCAATGATAATTAACAAT AAAGTTTACGGCAATCTTGATGAGAATAAGATTAAGAAAATAATATCTG GTTTAAAATGA 369 QMZS01000067. HydC ATGAAGAAACTTGACAACTTGATCTCAGAATACAAAAGAGGAGAATGC 1_2, nucleotide (NuoE-like) AAATTGATAACCTTGTTTAAAGAAATAGTTAAATCAAAGGGATTTCTCT CTTTTGAAAACCTGAATTATTTAAGCAAAAATCTGGATATTCCTCTTGC AAAATTATATACAACTGCAAGTTTTTACTCTTTTATTCCAACAGCAAAA AAAGGAAAATATATTATACGCGTCTGCAATAATCTATCATGCAATCTTA ATGGTTCGGAAAATATAATTGAAGTATTAAAAAAAGAACTAAAAATAAA CCTTGGGGAAACAACTGGAGATGGAAAATTCAGTCTTGAACTTACTTC ATGCATTGGGCAATGTGATTCCGCCCCTGCAATGATGATTAATAATAA AATATATACAAAACTTGACAAAACCAAAATCCGCAGAATACTGCGCGA ATTAAAATAA 372 QMZV01000029. HydC ATGAAGAAATTGGATTCATTAATTTTAAAATATAAAAAAAGAGAATTTAA 1_3, nucleotide (NuoE-like) CTTACTCACATTGCTTGAAGAAACTGTGAAAATAAAAGGATATCTCTCT 1005373994
TTTAAGACCTTAACTTATATCAGCGAAAATCTAAAAATCCCCCTCGCAA AATTATATGGCGTTGCAAGTTTTTATTCATTTCTTCCAACTGTGAAAAC AGGAAAATATATCATCAGGGTATGCAATGGACCTTCCTGCTATCTGAA TGGCTCGCGAGAAATTCTGAAAGTATTGAAAAAAGAACTTAAGATAGA TTTGGGCCAGACAACAAAAGATGGAAAATTTACTCTTGAATCAGCTTC CTGTATTGGTTGTTGCGATTCGCCTCCCGCAATTATGATTAATAACAA AGTATATAAAAATCTTGATAAAAATAAAATCAAAGACATAATAAAAAAAT TAAAATAA 373 Zodletone_Water HydC ATGACAAAAATTCTGATGAATAAATTAAGAGAAATTCAAGAAAGAGAT _assembly2018_ (NuoE-like) GGTTATCTTTCAGAAAATTCATTAAAAGAATTAAGTAAAGATATGGATG k141_1542890_2 TCCCGATTTCAAGACTTTATGGGATGGCAACTTTTTATTCAATGTTTCA 2, nucleotide TACAAAAAAAATGGGTAAAAATATAATTGAAATATGTGGTTCTCCGTCG TGTTTTTTGAATGGAGGTTTAACATTAGAGCAATTTTTAATAAAAGAAT TAAATATAGACATAGGAGAAACGACGAAAGATGGAAAATTTACTTTATT AAAAACATCTTGTATAGGTTGTTGCGATATCGCTCCAGCAATGCTTTTT AACGGGAGGCCTGTGGGTCATCTTACGGAGAATAAAATAAAAAAAATT TTTAAAAAATGCAAATCTTAA 375 Candidatus_Lokia HydC GTGATAAAAATAGAATTATGTTGTGGACTTAATTGTTTGGCTCATGGA rchaeota_archae (NuoE-like) GGACAAGAATTATTGGATATATTAGAGAATGATGAAAAATATAAAAAAA on_bin106|JAGX AATGTGAGATTGAATGTGTAAATTGCTTGGATGAATGCGGAAATACTG OA010000065_3 CTGAGAAAAGTCCTGTAATTAAAATCAACAATAAGATATACAGAAGAAT 4, nucleotide CACATCAGATTTTCTTATGGATTTATTAGATAGACTTATTTCTGAATAA 376 Candidatus_Lokia HydC ATGGCGAAAATTTCGTTAAAAATTTGCTGTGGAATGAATTGTTTGGCA rchaeota_archae (NuoE-like) CATGGGGGACAAGAACTTCTTGACTTAGTGGAAGATTCTCCTACGTAC on_FW102|JAIZ ACAGCATATGTTGAATTAGCATGTGTTGAATGTTTAAATACATGTGGA WK010000001_1 GATCGCGGCTATAATTCACCAGTTGTGGAATTAAATGGTAAAATCTAC 238, nucleotide TCAAATATGACTGCTGAAAAATTAATGGAGCTGCTTGATCAATTGATC CAAACTAATAATAATTAA 377 Candidatus_Lokia HydC ATGGCAAAGGTTCGCGTTCGCGTGTGCGTCGGGACCAATTGTGCATT rchaeota_archae (NuoE-like) CCACGGCGGGCAATCCATCAGCGACAAGCTCGATTCTGACCCTTTAT on_strain_AS27yj TCGAGGGAAAAGTTGACGTCGAGACCGTCAAGTGCTTCGACAAGTTG COA_147|JAAZN TGTGCTGATGGAAAAAATTCTCCCATCGTGGAAATCGACGGAAAGATT I010000050_7, TACAAGAAACTGTCGATGGAAAAGCTCTCCGAGATCGTGTTTTCCAAG nucleotide CTTTCAACCCTTCCCAAGCAGGTGAACTGA 378 Candidatus_Lokia HydC ATGGCGAAGATTGATCTGAAAATTTGCTGTGGGATGAATTGTCTCGCC rchaeota_archae (NuoE-like) CATGGTGGACAAGAATTACTCGATTTGGTGGAAGATAATCCCAAATAT on_strain_B53_G GAATCATTTATCGATCTTACTTGTGTGGAATGCCAAAATACTTGCGGA 9|QMYW0100008 GAACGGGGATACAATTCTCCCGTTGTAGTAATTAACAATCAAGTTTAT 5_3, nucleotide TCTAATATGACCGCACAACATTTGATGAAATTACTTGATAAATTAATTA ATGACTTAGTTGCAAAATAA 379 Candidatus_Lokia HydC ATGTTGATTTTTGAAGATAATAATTCTTTAGGTATACTTGGAAATCGAT rchaeota_archae (NuoE-like) TTTTCTGTAAGGTGATGATAGGTTTTTTTACTTCTTCGTGCAACTTAGT on_strain_HM1_B TACTAGCGCTATGGTGAAAAATACTAATAATACTAACAATGCTAAAGAT 6_4|JABCSI0100 GTTAAAAACACCAAAATCACTTTGAAGGTTTGTTGTGGGATGAACTGT CTGGCCCACGGTGGACAGGAGATTCTTGATTCGGTAGAGGATGAGA 1005373994
00124_7, GCCGTTTCACAAATCTGGTGGAGATTGAGGCTGTGGAATGCAGAGAT nucleotide ACATGCGGCGATGGAGGTTATCAATCTCCAGTTATCCAATTAAATGGT CGTATCTACCCCAAAATGACGGTGGAGCGACTGTGGGAATTGTTAGA TAAGGAAATTGACAATCTTTCTTGA 380 Candidatus_Lokia HydC AAACCTATTAGTTTAAAAATTTGCTGTGGTATGAATTGTTTGGTTCATG rchaeota_archae (NuoE-like) GCGGGCAAGAATTGCTTGATCTTGTTGAAAATGATTCAAAAATCTCAA on_strain_HM4_B ATAATTGTCAAATTGAGGGTGTTGAATGTCGTGAAACTTGCGGGGATT 48|JABCTF01000 TGGGAAAACAATCTCCAGTAGTGGAAATTAATGGTAAAATATATGCAA 0130_1, AAATGACAGCAGAACGGTTAATTGATATGTTATACAAAATGATTGATG nucleotide AGAAATAA 381 Candidatus_Lokia HydC ATGGGTAAAATTAGAGTAGAGATTTGCTGTGGATTACATTGTTCACTT rchaeota_archae (NuoE-like) CGAGGTGGTCAAGAAATCTTTGACGCGGTTGAATCGGAGGAACTCTT on_strain_MAG_ CGAAGGAAATAAATTCGATATTATACCTGTAAATTGTCTACAGTGTTGC 14|DXJH0100033 AATGATGGAGCATTGTCTCCAGTAGTCTCTATTAATGGGGAGTGTTAC 8_7, nucleotide ACTAAAATGACTACAGAGCGTCTTATTTCTACATTGCGTAATTTTATAA ATTAA 382 Candidatus_Lokia HydC ATGCAAGGAGGTCAAGAACTATATGATATGGTAGAATCCGACGAATTA rchaeota_archae (NuoE-like) CTTGCCGATAGCAAATTCGATATTATCCCTGTGAATTGCTTGCAGTGT on_strain_YT2_0 TGTAAAGATGGTCAATTTTCTCCAGTTGTAGCTATTAATGGAGAGTGC 12|JAEOTN0100 TACACTGAGATGAACGCTGCGCGTCTTATTTCCACATTGCGTAGTTTT 00030_10, TTAAATTAA nucleotide 383 Candidatus_Lokia HydC ATGCGAAAAGTATCTGTGGAAGTATGCTGTGGATTACATTGTAGTTTA rchaeota_archae (NuoE-like) AAGGGAGGTCAAGAATTACTGGATCTCATTGAATCTAATCCATTATTT on_strain_Zod_M CAAGGAGATCAATTCGATTTTATTCCCGTTAATTGTTTACAGTGTTGTG etabat.1044|JAF ATGATGGTGCGCTATCCCCTGTAGTTTCTATTAATGGAGAAAACTATA GBS010000008_ TGAAAATAACCCCAGAACGTCTAGTTTCCACATTGCGTACTTTTATAAA 98, nucleotide CTAA 384 Candidatus_Lokia HydC ATGATTAAAATAGAAATATGTTGTGGTTTAAATTGTTTAGCTCATGGAG rchaeota_archae (NuoE-like) GGCAAGAACTCTTGGATATTTTGGAGAATAATGAGAAGTATAAAAAAA on_strain_Zod_M ATTGTGAAATAAAATGTGTAAATTGTTTAAATGAATGTGGCGATACAGC etabat.578|JAFG AGAGAACAGCCCCGTGATTAAAATCAATGACAAGATTTATAAAAAGAT OA010000117_3, AACCGCAGATTGTATGATGAATTTGATAGAAAACTTTATTTCTAAATGA nucleotide 385 NZ_CP042905.1_ HydC ATGGAACCTATTAGTGTAAAAATTTGCTGTGGTATGAATTGTTTGGTTC 16, nucleotide (NuoE-like) ATGGTGGACAAGAATTGCTTGATCTTGTTGAAAATGATCCAAAGATTT CGAACAATTGTCAAATAGAAGGTGTAGAATGCCGTGAAACTTGCGGA GATTTAGGAAAACAATCTCCTGTAGTGGAAATTAATGGTAAGATATAT CCTAAAATGACTCCAGAGCGATTAATTGATATGTTATCTAAAATGATTG ATGGGACCAATAAAAAGTAA 389 DAEG01000100. HydC ATGAACGAGTCACAGAAAAAAGAGCTTGATGAGATCGTCGAGAAGTA 1_4, nucleotide (NuoE-like) CAGAAAGCTCCCAAACGCAAAGATCAAGATACTGCAGGAGACCCAGC TGAAGTTCGACTATCTAAGCGGTGAGAGCATGAGACACATAGCTGAG GCTCTCGGTGTGAGCTATACCGAGATATACGGCATAGCCACGTTCTA CTCGTTCTTCAACCTCAGCCCTCCTGGAGAGCACACAATAAATGTGTG 1005373994
CCTGGGCACGTCATGCCACGTGAAGGGCGGTGACGAAATCCTCGAA GAACTGCAGAAGACTCTAGGTGTGGACGTGGGAGAGACCACAAAAG ACGGTCTTTTCTCCATCAAGACCGTCCGCTGTCTGGGCTGCTGCGGT CTCTCCCCGGTCATCGACATAGACGGAAAGACCTTTGGCAGGGTTAA GAAGGCCAAAGTTAGTGGCCTGCTAGAGCCGTACAGGGAGGTGAAG GAATGA 393 DAJR01000051.1 HydC ATGGATGCAAAGGACATTTTCATCATCGACACCGTTCTTGATAAATAC _22, nucleotide (NuoE-like) AAGGACGTAAAGGATCCTGCTCTGATAATCCTCCAGAACGTACAGGC TGTGCTTGGATACGTCGGAAAAGAGTACCTCGAATACATTGCTGAGA AGACAGTTTATTCGCTCACCGAACTCTACGGTATCGTCACATTCTACC CATCATTCAGACTTGAGCCGCCCGGCAAACATATGATAAAGGTTTGC CAGGGCACAGCATGCCATGTCAGGGGCGGGGAGAAGGTTCTCAGGG AAATCCAAACGCAGTTGTGCATCAAACCAGGTGGCACAACCGAGGAT CGCCTTTTCTCGCTGGAATCTGTCAGATGTCTCGGGTGCTGCGGACT TTCTCCGGTCATAATGATCGACAACGAAACATATGGCAGGGTGAAGC CGGCCAAAATCCGTGAGATTCTGAATAAATACCGGGGGGGGAGCAAA TGA 395 DRVA01000011. HydC TTGTCGGAGAAACCTGAGGAGAAGCAGATAGACTTAGTTGAAAAGAT 1_2, nucleotide (NuoE-like) AATAGGGAATATAGCTGAAAGATATAATGAGCCTAGGAGAGCCTTAAT CTTAACTCTTCAGAAGATTCAAAATAGCCTAGGCTATCTTCCAAAGTG GAGTTTAGAACTTGCTTCAAACCATTTAAAAATACCACTTAGCACAATT TACGGAGTTGCTACATTCTACCATCAATTTAATCTCAACCCCCCTGGT AGAAACGTTATACAAGTCTGTATGGGTACGGCATGTCATATTAGAGGC AATTCTGAAAACTATGCTTTCCTGCTCAATCTACTTAACATAGATCCTG GTGAAAATACCTCTAAGGATGGGCAGTTCACAGTTCTTAAAGTCAGAT GCCTTGGATGTTGTAGTTTAGCACCTGTTATAAAACTGAACGACGACA TATATGGTAAAGTGGATTTTCCAACAATTCGTAGAATAATCTCTAAATA CCGCCTAGCTACTAAAGAGCTTCACGTGGTCTCTAGAGAGACTAGGT GA 399 DRYW01000042. HydC ATGTACTACGAGCAGTCAAAAACTTATTTGCGAGCGCCGCAATCGAG 1_34, nucleotide (NuoE-like) TGTCATGTTAGATTCTGGGAGGCGTCTTAGCGTCGTAGATAAGGTTAT TTATAGGTACGGTGCTGACAGGAGTAAGTTAATCAACATGCTTCAGGA AATACAGAAGGAGTTTCAGCATCTTCCTAGACAAGCTCTTGAGATGCT GTCTGTGAAGCTAGGTATCTCATTATCTGAAATCATTAATGTTGCTAC GTTCTATCATCAATTTAGGCTAGAACCCGTCGGGGTGTATGTTATTCA GGTATGCTTCGGCACAACCTGTTATCTTAAGGGATCTTCAGAAATCTA CGAATCTATGAGGAAGGTGCTTGGCTTGAGAGAAAGAGAGAACACAG CTAGAGACGGCACTATAACTGTGGAGAAAGCTAGGTGTTTTGGTTGTT GTAGCCTAGCTCCGGTCATCATGGTTACTTCATCGGATGGGAGTGAG AGGTATGTTCACGGCAGGCTTAACTCGTCTGAGAGTAAGAGGATAGT CCTAGATTACAGGTCGAAAGCTTTGATGAAATTAAAGGGGCCTAGAAA TGAGTGA 405 DTBU01000097.1 HydC TTGCAACAGCAAGTATCTAGAGATGTTGATGTAGCATTTATTGAGGAA _12, nucleotide (NuoE-like) GTTATAAACAATGTTATTGAAAGTAATGGTAGATCTAGAAAAACCCTCA TTCTGATCCTCCAGAAAATTCAGAATAGATTAGGTTATCTTCCAAAGTG 1005373994
GAGCTTAGAACTTGTTTCAAACAATTTAAAAGTACCACTAAGTAGCATC TATGGAGTCGCCACGTTTTACCATCAATTCAATATGGAGCCGCCCGG TGAGAATATAATACAAATATGTATGGGCACAGCGTGCCATATTAAAGG CAACTCTGAAAACTACAACTTCCTACTGAATTTACTTAATATTAATCCT GGCGAAAATACTTCTAAGGATAGGCTCTTCACTGTATTTAAAGTTAGA TGCCTAGGATGCTGCAGTCTAGCCCCAGTTATCAAAGTGAATGATGA AATATATGGTGAAGTTGATTTTAAGAAGCTCCGTAGGATAATTTCCAA GTATCGAACTAAAACTAAAGGAGAATCATATGGTGAAATATAG 406 DTCV01000006.1 HydC ATGTATTCAAATAGTAGAAGTGTGGATGAGGTTGTGGCAGGTATTTTA _4, nucleotide (NuoE-like) AGTAGGTATAGCGAACGCCAGCATTTAATGAGGATTCTCCAAGAAGTT CAGTCCACTTATGGTTACCTACCCCGTGAGGTCTTAGAGAAGATATCC CAGCATACGAAGATACCGTTGAGTGAGATAATCAACGTAGCCACTTTT TACCATCAATACAGGCTTGAAGTGCCTGGATACTACATATTCTTAGTTT GTATGGGTACTGCATGCCACTTAAGAGGTAATTATGATAATTACAATG CTTTAAGAAGTACTCTAGGGATTAAGAAGGGGAGTACCTCACCGGAT GGCCTGGTTAGTGTTGAGAAAGCTAGGTGTTTTGGTTGCTGCTCATTA GCACCCGTCATCATGGTTGTTAGTAGGGATGGTAATGAGAGGTATCT ACACGGTAATGTTGATATGAGGGAAGCTAAGAAATTAGCTCTTAACTA TAAGGCACTAGCTTCTAGGAAATTAAAAAGTGGTGGTCAACAATGA 410 DTEB01000060.1 HydC TTGTCAGAAAAACCTGAGAAGGAGAGGATAGACTTAGTTGAAAAGATA _3, nucleotide (NuoE-like) ATAGAGAATATAGCTGAAAGATATAATGAGCCTAGGAAAGCCCTAATC TTAACCCTCCAGAAGATTCAAAATAGTTTGGGTTATCTCCCAAGGTGG AGTTTAGAGCTTGTTTCAAACCATTTAAAAATGCCGCTTAGTACAATTT ACGGGGTTGCCACATTCTACCATCAATTTAATCTTAAACCCCCTGGTA GAAACATTATACAAGTCTGTATGGGTACGGCATGTCATATTAGAGGTA ATTCCGAAAACTATGCTTTCCTGCTCAATCTACTTAACATAGATCCCG GTGAAAATACCTCTAAGGATGGATTGTTCACAGTTCTTAAAGTCAGAT GCCTTGGATGCTGTAGTTTAGCACCCGTTGTAAAAGTGAATGACGAC ATATACGGTAAAGTGGATTTCTCGACAATTCGCAGAATAATCTCTAAAT ATCGCCTAGCTACTAGAGAGCTTCACGTAGTCTCTAGAGAGACTAGG TGA 415 DTOX01000047. HydC ATGCCCACGGTAAGTAACGTAGATAATGTTTTATCAGATATACTTGGT 1_15, nucleotide (NuoE-like) AGGTATAGTAGTAGGCAATATCTTATGAGGATACTTCAGGAGATACAG TCAGTTTATGGGTACCTACCTCGGGAAGCTCTTGAGAAACTCTCAAAC ATTATGAGGATACCCCTATCAGATATAATTAACGTAGCTACTTTCTACC ATCAATACAAGCTCGAACCCCCGGGGTACTACATATTTTTAGTCTGTA TGGGTACGGCATGTCATTTAAAAGGTAATCAAGATAATTATAAGGTTTT AAGAGATGCTTTAGGTATTAAGGGAGGTACTACATCTCCTGATGGTTT AGTTAGTGTTGAGAGGGCTAGATGCTTTGGTTGTTGTTCATTAGCACC AGTGATCATGGTTGTTAGTAGGGATGGTAGTGAAAGGTACCTGCATG GGTACGTCAATGTACAAGAAGCTAGGAAATTAGCTTTAAAGTATAGGG GTCTCACCTCCAGTAAGAGGAAATCAGGTGTTTAA 423 DTQO01000091. HydC ATGAGTCAAGACTTTAGATCTGTAGCAAGCGAGATAGTTTCAAAGCAC 1_20, nucleotide (NuoE-like) GGTGGATCTAGGAGTTCGTTGATATTAGTGCTTCAGGAGATACAGTCT AGATTCAACCACATCCCACTTGAGAGCATGCTGGAGGTTTCTGAGAA 1005373994
AATGGGAATCCCTCTCAGCGATATCTATGGTGTCGCCACATTCTATCA TCAATTCAGGTTGAAGCCGCAGGGAAGGCACCTCATTTCTGTATGCA TGGGTACAGCTTGCCATGTGAAGGGATCCCAGGCAATTTATGACTTG TTGAGGCAGAGGCTCGGTATAAAGGGAGATGAGGAGACAAGTCCGG ACGGAATGTTTACTGTGCAGAAAGTAAGGTGCATGGGCGCATGCAGC CTTTCTCCTGCCCTCAGAATTGATGGTGACATATATGGGAAAGTTAAA CCTGAGATTCTAGATACGATTCTCCATAAGTATTCGCGAGTCTCGGGG TGA 424 GMQ_scaffold_8 HydC TTGCAAAAGCTGCAGGAGAGATATGGGTTCTTGCCGAAGGACGCTCT _151, nucleotide (NuoE-like) AATGAAGCTTTCCAAGGATGTGGGAGCTCCTTTGAGCAGGATATTCG GCGTCGCGACTTTCTATCACCAGTTTAAGCTGGAGTCGCCTGGAAAA GTCCTGATCTCGGTTTGCATGGGAACAGCATGCCATCTTAGGGGCGA TGCCAACAACTACGAGTTCCTGAGGAGGTTTTTAAGAATAGCACCAG GCAAATCCGTCAGCGATGATGGCCTATTCTCCTTAGAGAAGGCTAGG TGCTTCGGCTGCTGTAGCCTGGCCCCAGTAGTGAGAGTTGGCGATAA GCTCATCGGCAAGGCCACGCCCAGGAGCTTCAGAGAATAA 428 GMQ_scaffold_9 HydC ATGCGTGCTGTGAGTGAACAAAAGACTCGTCTGTTATCACGCGCCGA 3_10, nucleotide (NuoE-like) CATTACGTTAGATTCTGAGAGGTACCTTAGTACAGTAGACAGAGTCAT TCATAAGTATGGCGTTGATAGAAGTAAGTTAATTAACATGCTTCAAGA AATACAGAGAGAATTTCATTATTTGCCTAGACAGTCTCTTGAAGCACT ATCTAACAAGCTAGGTATACCACTGTCTGAAATTCTTAACGTAGCTAC TTTCTATCATCAGTTTAGGCTAGAACCCCTCGGAACATACGTAATTCA CGTGTGTTTTGGCACAGCATGCTATCTTAAGGGGTCTCCGGAGATCT ACGAATCTATAAAGAAAGCACTCAACTTGAAAGAGAGAGAAAACACTA CAAGAGACGGCGCGATAACCGTAGAGAAAGCTAGGTGTTTCGGTTGT TGTAGTCTGGCTCCGGTAATGATGGTCACCTCGTCAGACGAGAGTGA GAGGTACGTGCACGGCAGGCTAAACATAGCTGAGAGCAAGAAAGTA GCTCTAGGTTATAGGACGAGAGCGCTGTCAAATATAGACAAGGCTTA G 432 NJDO01000061.1 HydC ATGTCCCTTCCGGAAGAAGAGAATCGGTTTCTTGATACGGTCCTCGA _2, nucleotide (NuoE-like) GAAAGCCTCCCGACTGGACTCCCCCATCCTTTTTATCCTGCATCAACT GCAGGAGAGATACGGCATGATCAAGCCGGAACATGCGGACCACGTT TCGGCGAGGCTCGGCATCCCCCTGGTTCGGCTCCATGCGGCGGCCG AGTTCTACGAGCATTTCACCCTGGAACGTAGAGGGAGATACATAATCA GATTGTGCAGGGGAATAGTCTGTCACGGGAAAGGTTCCCTTCCCATT CTTGAAGCCCTCAAGGAGAGGCTCGGAATAGAGGACGGTGAGACCA CGGAGGACGGTATGGTGACGCTTGAGACAGCATCATGCATCGGCCA GTGCGACGGAGCTCCCGCCATGATGATAAACGACGTTGTGTACAGG GACCTGACCGTGGAAGGTGCAATCGCCATCATCGATGACATCCTGAA GGAAGGGGGTGCGTGA 439 PKYH01000114.1 HydC TTGGGCGGTAGAAACATGAACGAGCAAGAGAAAATGGAAATCGAGGA _31, nucleotide (NuoE-like) GATAATTGCAGGGATAGAGGATGTGAAAACCGCTAAGATAAGGTTAC TGCAAGAGCTCCAGAAAAAATACGGATATCTTCGTGAAGAACACATGA GGTACGTTGCCAAAAGGATAGACGACAGCTACACCGACCTTTATGGC ATAGCAACTTTTTACGCCCAATTTAACCTCAACCCAGCTGGACGGCAT 1005373994
ACCATCTACGTCTGCGAAGGCACTTCTTGCCACGTGAAAGGCGGGAA AAAACTCCTCAGCAAAATCGGCGAACTCTTGAATGTCAAAGTTAAGGA AACCACTGAGGATAAACGCTTCACCTTAAAGGTTGTTCGTTGCCTAGG ATGTTGCGGTATCTCGCCCACTATAATGGTTGACAATGAGACTTTCGG CAGAGTAAGATTAACAGATCTTACAAGTATCTTTTGTAGGTTTGAATAA 444 SR- HydC ATGGACGCCAAGGACAGGAAGATGCTCGATGACATACTGAGCAGCCA 2_scaffold_141_2 (NuoE-like) CCGCGACACCAGGGACCCGGCCATGCTCATCCTGCAGAAGGTGCAG 318042_4, ACCTCCTTCGGCTACACGGGGCAGGACCACCTGACATACATCTCGGA nucleotide GAAGTCGTGCATCCCGCTTTCGAAGCTGTACGGGATAGTCACCTTCT ACCCGTCCTTCAGGCTCAGCCCGCCCGGCAAGAGCATCATCAGGATA TGCGAGGGGACGTCCTGCCACGTCAGGGGCGGGGCCAGGATCGTG AAGGAGCTCGAGCGGCAGCTCGGCATCGGGCCAGGCCAGACGACG AAGGACAGGAAGTTCTCCCTCGAGTCCGTGCGCTGCCTCGGGTGCT GCGGGATTTCCCCCGTCGTCATGATGGACGACAAGACCTACGGGCG CGTGAAGGCGACGAAGCTCGCGGAAATCCTCGCCACACATGGAGGG GAGTGA 448 SRVP18_trench_ HydC TTGAACGAACAAGAGAGAAGGAAGATTGACGAGCTAATTACGGGCAT 6_60cm_scaffold (NuoE-like) AGAAGATGTAAAGACTGCAAAGATAAGGTTGCTACAGGAACTTCAGAA _130_19, GAAGTACGGATACCTACGTGAGGATCACATGAGGTACGCCGCTGATA nucleotide AGATAGACGCCAGCTACACCGACCTCTACGGTATAGCAACCTTCTAC GCTCAGTTCAACCTCAGCCCAGCAGGGCGACATACCATCTATGTCTG CGAAGGTACATCATGCCATGTGAAAGGCGGTAAAAAACTCCTCAGCA AAACAAGCGAGCTCTTGAACGTCAAGGTGAAGGAAACCACTGAGGAT AAACGCTTCACCTTAAAAGTTGTCCGCTGCCTAGGCTGTTGCGGCAT CTCACCCACACTAATGGTGGATGCCGAAACCTTCGGCAGGGTTAGGT TAACCGATCTTAGAGGCATTTTTTCGAGGTTTGAGTGA 449 VMTD01000047. HydC ATGATACATCTAAAACTCAAGTGCCACCATTGTGGAGAAAGCCTGATG 1_33, nucleotide (NuoE-like) GACCCAGATTTCGAGATTGACGACCATCCCAGTGTGAGAATTATTGTT GCATCTAATGGAGAAAAAGGAATATTACGCTTGAGCTCACTCTATGGA AGTCACAAAAAAGATTCTGAATTAGATGTCCAAGAGGGAGAAATTGTT CGAATATTTTGTCCACATTGCAATGTCGATATGAAGAGTAGCAGGTTT TGTTATGAGTGCAAGGCACCCATGATCACATTTGAATCTCTTCTGGGT GGATATATTCGAGTATGCTCACGATGGGGTTGCAAGAAACAACTTGC AGAGTTTGAGAATTTGGAAACAGAGCTTAGGGCATTTCATGCGAAATA TTCTCTCTCTCCTCAAGGGAGGGGGGAAAAATGA 452 VMTD01000047. HydC GTGGACACCAAAGATAAGTTCATAATTGATACGATACTGGAAAAATTC 1_44, nucleotide (NuoE-like) ACCGACATGAAAGACCCTACTTTAATGATATTACAGAACATCCAGGAG ATACTGGGATGTGTGCGAGAGGAATATCTCACATACGTGGCCGAAAG CAACGGTTATTCTTTGACAGAGTTGTATGGTATAGTATCATTCTATCCG CAGTTCAGGTTGAGCCCGCCAGGTAAACACACGATCAAGATATGCAA GGGAACAGCGTGCCATGTTCGCGGCGGTGCCGCTATTCAGAAGAAC CTGCAGAACCTCCTTGGAATCAAACCAGGTGATACAACCGAGGACGG TGTTTTCACTCTGGAATCCGTGAGATGTCTGGGATGTTGTGGCCTTTC CCCGGTCATAATGATAGATAATGAAACCTATGGAAGGGTGAAAACATC 1005373994
AAAACTACAGGAAATTTTGGTCAAGTACAAGGGGGTGAGTGCTAATG AAAATCAGTAA 457 WOYO01000071. HydC ATGGAGGGGAAACATTCTTACGAAAGGGAAAAGCTGGAGCTGAAAAA 1_29, nucleotide (NuoE-like) TTCTGTCGATACAGCAATAGAAAATAAAACAAAAAATAGGGTTGAAGT AAAATACAAGGAGGAAAAAGGGGTATATGACGCTCTAAGCTTGGCCG ATGAGCTGAGAATTGATTATGACGTCACCGTTTTGACACCAAATCCGA TAAATGGCATCTATGTGTGTCAGTATGATCCCGTGTTTAACTATGCTG ACATCAAAATTTTGAAGGAGGCGTTATGA Table 4: Amino acid sequences of HydC SEQ Description Subunit Sequence ID NO: 464 AQRS01000037.1_14, HydC MFSFKKHVLVCTSEKPGHCAEKGGPELLAAFREEVAKRGLQNEI amino acid (NuoE-like) YVTKTGCTSQHHCGPTVIIYPDGVWYKLVTKEDIPEIIESHLLGGKI VERILNREIGLFRKQA* 465 09mi20_z1_2019_ig183 HydC MHSKWLESYLKENEHQSLLAVLLRIQEREGFLSKESLEHVSKGM 92_10016_4, amino acid (NuoE-like) KIPLSKIYSVATFYSEFKLEKRGKHIIKLCSGTACLVKGNNVNLNFL KTVLNLKPGQTTPDKLFTLETVNCLGTCSLAPVINIDGKIYPNVTIE KLSEIIEKLKRSKR* 468 AQSC01000060.1_11, HydC MEKDIESIISKYSGASDLIDALEDVQEAFGHISEDNMHSINQVLKIP amino acid (NuoE-like) LVDILGVVSFYSAFKTKPPGKHIIRICRGTACHIKGSTILEEHLEDKL GIKAGETTEDGKFTLEPVNCIGACAKAPTMMVDDIVYGDLTKERID EILGEYK* 469 CAITEO010000066.1_1, HydC MVELRRVTINGNHYYYLFHEIRENGQFKKFRHYIGAQEPDNAMQ amino acid (NuoE-like) QKLERDFMEDIKNNPDKYSPKEKQNIIAVLQQIMEKENYISEENFV RLSSELDIPLVNLVGVATFYSQFRLTKPGKHTIKICDGTACHVKNS AALRVFLEETLDIKSGQVTKDGNFGLEVVNCIGACARAPSMMIDE TVYGKLDKKKIKEIIGTYK* 473 Candidatus_Lokiarchaeo HydC MPTKISDILSHFTQGNSSELIPILQAVQAEYGFISEESVYEISEYLNI ta_archaeon_FW102|JAI (NuoE-like) PSSKIYGVATFYAQFRLQPPGRHVINLCTGTACHVKGSEKLIPVFE ZWK010000001_1633, QELKCKAGETTKDGRFTLNLVACLGACALSPVVNIDSDFYGNLTA amino acid GEIKKILRKYK* 475 Candidatus_Lokiarchaeo HydC MSILSFENRNKKLSGGLKKLMPAKISDVLSRFKQGDASELIPVLQA ta_archaeon_strain_B53 (NuoE-like) VQSEYGYISEDNTYEIAEYLNLPSSKIYGVATFYAQFRLEPLGRHVI _G9|QMYW01000078_8 NLCTGTACHVKGSEKLIPVFEQELKCKAGETTKDGRFTFNLVACL , amino acid GACALSPVVNIDSDFYGNVSPSDVKKILRKYK* 478 Candidatus_Lokiarchaeo HydC MHSKTITEILVPFPRDEPSALIPVLQAVQSEYGYLSENNIYAISDHL ta_archaeon_strain_HM (NuoE-like) NVPSSKIYGVTTFYAQFRLKPLGRHVINLCQGTACHVKGSEKLIPV 1_B6_4|JABCSI0100000 FEQELKCKAGETTKDGKFTFNLVACLGACALSPVVNIDSDFYGNI 64_8, amino acid KIADIKKILRKYD* 480 DAWM01000035.1_7, HydC MVQLSVESSISKQRNTESLIDGILENYNDYSRNIINILLRIQEKQGFV amino acid (NuoE-like) SEQDAIQLSDKTHIPLAKIFSILSFYNYFSFKRAGKNLLLLCDGTAC 1005373994
RVQGNRKLRQTLKDELDISPGETTKDNLFTLKEVRCLGACALAPV MMVNGKIYGNLDEKKVKEIISGLKQESGKKES* 482 DSAL01000038.1_4, HydC MMAKIDDILAKYDSAHELIDMLEDIQAEYGYISEENMRKVEQDLKI amino acid (NuoE-like) PLVDIYGVVTFYSAFKLKPSGKHTIKVCTGTACHVKKSDSLKEHLM KALSVKEGETTSDGKFTLELVNCIGACAKSPAMMIDEKVYGELTA KKIDSILKEY* 484 DSBS01000042.1_18, HydC MEKEGKTVIAGESRVLEMLQEINKNEGYISRERLSSISRELGIPLSL amino acid (NuoE-like) LYGLVTFYNQFNLATSGRYTIEVCEGTACHINKSSEIKRAIKDAAGI SVDETSSDGLFTLRDVRCLGACALAPVLRLNKKIYSKMTYEKTKE LIFKLKKEAEGEAKLK* 486 DSVV01000048.1_6, HydC MKKEGFNLIAELQAVQDRYGYLPMDVLKTLSKEHKIPGTEIFAVAT amino acid (NuoE-like) FYNQFKFDKPAKHTIQVCTGTACHVKRSADLLSQIQKTLKIRPGEI TKDGLIKLETVNCIGACAKAPAMMVDDKVYGLVDQDKLRQILGAL R* 488 DUIB01000037.1_10, HydC MNNHESLKPDIHKVLKKYDSKKDIIPALQDIQENFGFVSEENAESL amino acid (NuoE-like) AKKINSPLVDISGVVTFYNMFRLKPVGKYHIAICRGTACHVQNSEE LLKYVEKKLKIKTGEITQDGRFSLEAVNCIGACAKAPAMMIHDKVY GQLTEKKIDAILDSMK* 491 LacPavin_0920_SED5_s HydC MLMSEKADAKKEELAKSPSSYEIEESVLSIIDSRSEEKSPLLPILQD caffold_1418049_11, (NuoE-like) VQKAFGFISPQAMLEISETLDIPLSQVYSAVTFYNEFRTKKRGKHL amino acid FRVCMGTACCIKKADAVILELEKELGIQCGQTDGNGLFTLETVNC FGACGLGPIVEVDGRIFSLVEPKKAKALAEEVKKKERLQK* 492 PEXD01000050.1_15, HydC MLMSGKADAKKEETAKSPSSYEIEESVLSIIDSHSEEKSPLLPILQD amino acid (NuoE-like) VQKAFGFISPQAMLEMSETLDIPLSQVYSAVTFYNEFRTKKRGKH LFRVCIGTACCIKKADAVILELEKELGIQCGQTDGKGLFTLETVNCF GACGLGPIVEVDGRIFSLVEPKKAKALAEEVKKKEGLQK* 494 PEXL01000024.1_9, HydC MYSKWLESWLRENDHKGLLEILLEVQHREQFLSRENIEFVSKAKE amino acid (NuoE-like) IPLAKIYSVATFYSEFRLDQRGKHVIRLCAGTACLVKGNNVNLNYL KTELNLMPGKTTADNLFTLEGVNCLGTCSLAPVVSIDGKIYPNVTI EKLSGLIEKIKKAER* 496 PGXE01000078.1_2, HydC MDKLDAIIKEHGKSPLPVLKAAKAEYGHLCKDVLEAISEKIDVPVS amino acid (NuoE-like) RLHGVATFYSMLGTEQIGENVIYVCNSPSCYVNGSLNVLEEFERR LGIRCGETTADGGVTLEKTACIGCCDMAPAILLNGEPCGPLGKKDI ARIVKSMRK* 500 PXDW01000025.1_23, HydC MVFLKEVQLGNNSYFFLFYSIQNEKFKPYIRYIGKKKPNTEYLNSL amino acid (NuoE-like) KKKFLKDVKSNPELFEKKQKKNVIIMLQEIQENEGYISEENIIRLSK EINIPATHIYGVLTFYTYFRFNPPGKYNIAVCNGTACHVKNSVSLIR YIENILDIKVNETTKDKKFSLGSVNCIGACAKAPAMMINNTVYGDL DEEKIKKILDGLE* 501 SR- HydC MNTASSAIFMDESREADIIGRLLEVQKKNGWLPRKEVERIAKETRT 2_scaffold_141_416062 (NuoE-like) PLAKVYGIATFYDFFSLNLHKKKEEIRKCLNHDCRLKGSEIIESS* 1_8, amino acid 503 AB_3033_bin_101_scaff HydC MSFRDIDEIIKDRKDNLLLPMLEAIQAKFGCVSEENAHYLSRKTGIP old_9312_3, amino acid (NuoE-like) FSKIYGVITFYEMLYTEPKGKYIIRICNSPSCYLNGSLKLIEFLESLL 1005373994
KIKSGETTKDKKFSLEIVSCIGCCDKAPAMMINDKVYGNLDEKKIR KIISSLK* 508 AB_1215_Bin_111.fna_s HydC CNSPSCYLNGSLKLIEFLESLLKIKSGETTKDKKFSLEIVSCIGCCD caffold_95187_7, amino (NuoE-like) KAPAMIINNKVYGNLDEKKIKKIISGLK* acid 509 CABMGE010000002.1_ HydC MKNILINRLREIQNKEGYVSEESLKKLSLELKIPISQLYGVATFYSMI 21, amino acid (NuoE-like) YTKKQGKYVIELCASPSCFLNGSWNLEDYLKKELKIDIGETTKNKK FSLKKTSCIGCCDKPPAMLLNGKVYTSLTEKKLKDILKKCK* 512 CAITKI010000066.1_57, HydC MAKEKGKKDTRPLMNMLHEVQERHGYISEHMLKQISVDQDIPIAR amino acid (NuoE-like) LYGVVKFYTMFHTEPQGKYVIEICGSPSCVLNNGVRLEKFLENEIG AGIGETSKDGLFSLYKTSCIGCCDEAPAMLINGEPHTKMTVERLKL ILKKLRDTEAAQAGEKKK* 514 CAITNU010000081.1_7, HydC MGKKLSENLSEKTCAEGVSKADQPKVLALLREAQERDGYVTHQA amino acid (NuoE-like) VERISKQTGFTESEIDGVASFYAMLYLKPVGRFIVRVCASPSCVVN GGGRALEWASEVLGVSDGETTKDGLFTLEAVSCFGRCETAPNV MINEENYGGIDSKEKMRNLIEKLRREAAAREATK* 515 JAACWB010000015.1_3 HydC MVEVLMNKLREIQEKEGFLSEKSLKELSNEMNIPISRLYGMATFYS , amino acid (NuoE-like) MFHTKKPGKNIIEICASPSCFLNGGLTLEKFLINELKIDIGETTKDNK FTLLKTSCIGCCNIAPAMLLNGKPVGNLTEKKLKKILKKCK* 518 JACCLF010000048.1_2 HydC MKGKNTPLLINILHEVQDKQGYISEQALKKISVEQNIPISRLFGVVK 7, amino acid (NuoE-like) FYTMFHTEPQGKYVLEICGSPSCVLNNGMKLEKFLENEIGVRIGE TSKDGMFSLYKTSCIGCCDEAPAMLINGKPYTNMTVERVKLLLKK LRKSKKK* 519 JACCLG010000046.1_7, HydC MERKKGTSLLMNILHEEQDKHGYISEQTLKKISVDEGIPISRLFGV amino acid (NuoE-like) VKFYTMFHTEPQGKYVVEICGSPSCVLNNGVQLEKFLEKEIGVRI GETSKDGMFSLYKTSCIGCCNEAPAMLINGMPYTKMTVGRLQLLL KKLRAASKKKKK* 522 Meg22_1012_Bin_224_s HydC MIAKKVLLNELGKEQKKKGFVSKKKLKEIAKNVCLPESEVFSAATF caffold_4514_12, amino (NuoE-like) YSFLSLEKRAKHIIQVCNCPSSHLHGSDRIMKYLEKKLKVKAGHAT acid KNKLFFLSETSCVGLCDKAPAIIVDGKPYVKVNEKKIDKILRKLK* 525 MWBV01000007.1_12, HydC MVKKQNKKANKKEDILLNIFHEEQDKKGYISIDFLKKISTKYNIPISR amino acid (NuoE-like) LYGVVKFYTMLRTEPQGKYIIELCGSPTCVLHESREIENFLKKELKI DIGDTTKDKMFSVYKTSCIGCCDEPPAMLLNGKPITNLTIEKVKKLI KELKSKKK* 527 NJBG01000001.1_784, HydC MNYLLLPLLKDIQKKKRYISEKDMKKLSKKTSIPIAKIYATATFYSML amino acid (NuoE-like) HTKKQGKYIIEICDSPSCYVNGSIDLIKFLEKKLKIKSGETTKNGKFS LHICSCIGCCDQAPAMKINERVYGNLTKKKIEEILDKCKF* 528 PCYE01000023.1_6, HydC MSFRDIDEIIKERKDNFLLPMLQAIQAKFGYVSETNAHYLSRKTGIP amino acid (NuoE-like) FSKIYGVITFYEMLYTEKKGKYIIRICNSPSCYLNCSLNLIKFLESSL KIKSGETTKNKKFSLEIVSCIGCCDKAPAMIINNKVYGNLDENKIKKI ISGLK* 532 QMZS01000067.1_2, HydC MKKLDNLISEYKRGECKLITLFKEIVKSKGFLSFENLNYLSKNLDIP amino acid (NuoE-like) LAKLYTTASFYSFIPTAKKGKYIIRVCNNLSCNLNGSENIIEVLKKEL KINLGETTGDGKFSLELTSCIGQCDSAPAMMINNKIYTKLDKTKIR RILRELK* 1005373994
535 QMZV01000029.1_3, HydC MKKLDSLILKYKKREFNLLTLLEETVKIKGYLSFKTLTYISENLKIPL amino acid (NuoE-like) AKLYGVASFYSFLPTVKTGKYIIRVCNGPSCYLNGSREILKVLKKE LKIDLGQTTKDGKFTLESASCIGCCDSPPAIMINNKVYKNLDKNKIK DIIKKLK* 536 Zodletone_Water_assem HydC MTKILMNKLREIQERDGYLSENSLKELSKDMDVPISRLYGMATFY bly2018_k141_1542890 (NuoE-like) SMFHTKKMGKNIIEICGSPSCFLNGGLTLEQFLIKELNIDIGETTKD _22, amino acid GKFTLLKTSCIGCCDIAPAMLFNGRPVGHLTENKIKKIFKKCKS* 538 Candidatus_Lokiarchaeo HydC MIKIELCCGLNCLAHGGQELLDILENDEKYKKKCEIECVNCLDECG ta_archaeon_bin106|JA (NuoE-like) NTAEKSPVIKINNKIYRRITSDFLMDLLDRLISE* GXOA010000065_34, amino acid 539 Candidatus_Lokiarchaeo HydC MAKISLKICCGMNCLAHGGQELLDLVEDSPTYTAYVELACVECLN ta_archaeon_FW102|JAI (NuoE-like) TCGDRGYNSPVVELNGKIYSNMTAEKLMELLDQLIQTNNN* ZWK010000001_1238, amino acid 540 Candidatus_Lokiarchaeo HydC MAKVRVRVCVGTNCAFHGGQSISDKLDSDPLFEGKVDVETVKCF ta_archaeon_strain_AS2 (NuoE-like) DKLCADGKNSPIVEIDGKIYKKLSMEKLSEIVFSKLSTLPKQVN* 7yjCOA_147|JAAZNI010 000050_7, amino acid 541 Candidatus_Lokiarchaeo HydC MAKIDLKICCGMNCLAHGGQELLDLVEDNPKYESFIDLTCVECQN ta_archaeon_strain_B53 (NuoE-like) TCGERGYNSPVVVINNQVYSNMTAQHLMKLLDKLINDLVAK* _G9|QMYW01000085_3 , amino acid 542 Candidatus_Lokiarchaeo HydC MLIFEDNNSLGILGNRFFCKVMIGFFTSSCNLVTSAMVKNTNNTN ta_archaeon_strain_HM (NuoE-like) NAKDVKNTKITLKVCCGMNCLAHGGQEILDSVEDESRFTNLVEIE 1_B6_4|JABCSI0100001 AVECRDTCGDGGYQSPVIQLNGRIYPKMTVERLWELLDKEIDNLS 24_7, amino acid * 543 Candidatus_Lokiarchaeo HydC KPISLKICCGMNCLVHGGQELLDLVENDSKISNNCQIEGVECRET ta_archaeon_strain_HM (NuoE-like) CGDLGKQSPVVEINGKIYAKMTAERLIDMLYKMIDEK* 4_B48|JABCTF0100001 30_1, amino acid 544 Candidatus_Lokiarchaeo HydC MGKIRVEICCGLHCSLRGGQEIFDAVESEELFEGNKFDIIPVNCLQ ta_archaeon_strain_MA (NuoE-like) CCNDGALSPVVSINGECYTKMTTERLISTLRNFIN* G_14|DXJH01000338_7, amino acid 545 Candidatus_Lokiarchaeo HydC MQGGQELYDMVESDELLADSKFDIIPVNCLQCCKDGQFSPVVAIN ta_archaeon_strain_YT2 (NuoE-like) GECYTEMNAARLISTLRSFLN* _012|JAEOTN01000003 0_10, amino acid 546 Candidatus_Lokiarchaeo HydC MRKVSVEVCCGLHCSLKGGQELLDLIESNPLFQGDQFDFIPVNCL ta_archaeon_strain_Zod (NuoE-like) QCCDDGALSPVVSINGENYMKITPERLVSTLRTFIN* _Metabat.1044|JAFGBS 010000008_98, amino acid 1005373994
547 Candidatus_Lokiarchaeo HydC MIKIEICCGLNCLAHGGQELLDILENNEKYKKNCEIKCVNCLNECG ta_archaeon_strain_Zod (NuoE-like) DTAENSPVIKINDKIYKKITADCMMNLIENFISK* _Metabat.578|JAFGOA0 10000117_3, amino acid 548 NZ_CP042905.1_16, HydC MEPISVKICCGMNCLVHGGQELLDLVENDPKISNNCQIEGVECRE amino acid (NuoE-like) TCGDLGKQSPVVEINGKIYPKMTPERLIDMLSKMIDGTNKK* 552 DAEG01000100.1_4, HydC MNESQKKELDEIVEKYRKLPNAKIKILQETQLKFDYLSGESMRHIA amino acid (NuoE-like) EALGVSYTEIYGIATFYSFFNLSPPGEHTINVCLGTSCHVKGGDEIL EELQKTLGVDVGETTKDGLFSIKTVRCLGCCGLSPVIDIDGKTFGR VKKAKVSGLLEPYREVKE* 556 DAJR01000051.1_22, HydC MDAKDIFIIDTVLDKYKDVKDPALIILQNVQAVLGYVGKEYLEYIAEK amino acid (NuoE-like) TVYSLTELYGIVTFYPSFRLEPPGKHMIKVCQGTACHVRGGEKVL REIQTQLCIKPGGTTEDRLFSLESVRCLGCCGLSPVIMIDNETYGR VKPAKIREILNKYRGGSK* 558 DRVA01000011.1_2, HydC MSEKPEEKQIDLVEKIIGNIAERYNEPRRALILTLQKIQNSLGYLPK amino acid (NuoE-like) WSLELASNHLKIPLSTIYGVATFYHQFNLNPPGRNVIQVCMGTAC HIRGNSENYAFLLNLLNIDPGENTSKDGQFTVLKVRCLGCCSLAP VIKLNDDIYGKVDFPTIRRIISKYRLATKELHVVSRETR* 562 DRYW01000042.1_34, HydC MYYEQSKTYLRAPQSSVMLDSGRRLSVVDKVIYRYGADRSKLIN amino acid (NuoE-like) MLQEIQKEFQHLPRQALEMLSVKLGISLSEIINVATFYHQFRLEPV GVYVIQVCFGTTCYLKGSSEIYESMRKVLGLRERENTARDGTITV EKARCFGCCSLAPVIMVTSSDGSERYVHGRLNSSESKRIVLDYRS KALMKLKGPRNE* 568 DTBU01000097.1_12, HydC MQQQVSRDVDVAFIEEVINNVIESNGRSRKTLILILQKIQNRLGYLP amino acid (NuoE-like) KWSLELVSNNLKVPLSSIYGVATFYHQFNMEPPGENIIQICMGTAC HIKGNSENYNFLLNLLNINPGENTSKDRLFTVFKVRCLGCCSLAPV IKVNDEIYGEVDFKKLRRIISKYRTKTKGESYGEI* 569 DTCV01000006.1_4, HydC MYSNSRSVDEVVAGILSRYSERQHLMRILQEVQSTYGYLPREVLE amino acid (NuoE-like) KISQHTKIPLSEIINVATFYHQYRLEVPGYYIFLVCMGTACHLRGNY DNYNALRSTLGIKKGSTSPDGLVSVEKARCFGCCSLAPVIMVVSR DGNERYLHGNVDMREAKKLALNYKALASRKLKSGGQQ* 573 DTEB01000060.1_3, HydC MSEKPEKERIDLVEKIIENIAERYNEPRKALILTLQKIQNSLGYLPR amino acid (NuoE-like) WSLELVSNHLKMPLSTIYGVATFYHQFNLKPPGRNIIQVCMGTAC HIRGNSENYAFLLNLLNIDPGENTSKDGLFTVLKVRCLGCCSLAPV VKVNDDIYGKVDFSTIRRIISKYRLATRELHVVSRETR* 578 DTOX01000047.1_15, HydC MPTVSNVDNVLSDILGRYSSRQYLMRILQEIQSVYGYLPREALEKL amino acid (NuoE-like) SNIMRIPLSDIINVATFYHQYKLEPPGYYIFLVCMGTACHLKGNQD NYKVLRDALGIKGGTTSPDGLVSVERARCFGCCSLAPVIMVVSRD GSERYLHGYVNVQEARKLALKYRGLTSSKRKSGV* 586 DTQO01000091.1_20, HydC MSQDFRSVASEIVSKHGGSRSSLILVLQEIQSRFNHIPLESMLEVS amino acid (NuoE-like) EKMGIPLSDIYGVATFYHQFRLKPQGRHLISVCMGTACHVKGSQA IYDLLRQRLGIKGDEETSPDGMFTVQKVRCMGACSLSPALRIDGD IYGKVKPEILDTILHKYSRVSG* 1005373994
587 GMQ_scaffold_8_151, HydC MQKLQERYGFLPKDALMKLSKDVGAPLSRIFGVATFYHQFKLESP amino acid (NuoE-like) GKVLISVCMGTACHLRGDANNYEFLRRFLRIAPGKSVSDDGLFSL EKARCFGCCSLAPVVRVGDKLIGKATPRSFRE* 591 GMQ_scaffold_93_10, HydC MRAVSEQKTRLLSRADITLDSERYLSTVDRVIHKYGVDRSKLINML amino acid (NuoE-like) QEIQREFHYLPRQSLEALSNKLGIPLSEILNVATFYHQFRLEPLGT YVIHVCFGTACYLKGSPEIYESIKKALNLKERENTTRDGAITVEKAR CFGCCSLAPVMMVTSSDESERYVHGRLNIAESKKVALGYRTRAL SNIDKA* 595 NJDO01000061.1_2, HydC MSLPEEENRFLDTVLEKASRLDSPILFILHQLQERYGMIKPEHADH amino acid (NuoE-like) VSARLGIPLVRLHAAAEFYEHFTLERRGRYIIRLCRGIVCHGKGSL PILEALKERLGIEDGETTEDGMVTLETASCIGQCDGAPAMMINDV VYRDLTVEGAIAIIDDILKEGGA* 602 PKYH01000114.1_31, HydC MGGRNMNEQEKMEIEEIIAGIEDVKTAKIRLLQELQKKYGYLREEH amino acid (NuoE-like) MRYVAKRIDDSYTDLYGIATFYAQFNLNPAGRHTIYVCEGTSCHV KGGKKLLSKIGELLNVKVKETTEDKRFTLKVVRCLGCCGISPTIMV DNETFGRVRLTDLTSIFCRFE* 607 SR- HydC MDAKDRKMLDDILSSHRDTRDPAMLILQKVQTSFGYTGQDHLTYI 2_scaffold_141_231804 (NuoE-like) SEKSCIPLSKLYGIVTFYPSFRLSPPGKSIIRICEGTSCHVRGGARI 2_4, amino acid VKELERQLGIGPGQTTKDRKFSLESVRCLGCCGISPVVMMDDKT YGRVKATKLAEILATHGGE* 611 SRVP18_trench_6_60c HydC MNEQERRKIDELITGIEDVKTAKIRLLQELQKKYGYLREDHMRYAA m_scaffold_130_19, (NuoE-like) DKIDASYTDLYGIATFYAQFNLSPAGRHTIYVCEGTSCHVKGGKKL amino acid LSKTSELLNVKVKETTEDKRFTLKVVRCLGCCGISPTLMVDAETF GRVRLTDLRGIFSRFE* 612 VMTD01000047.1_33, HydC MIHLKLKCHHCGESLMDPDFEIDDHPSVRIIVASNGEKGILRLSSL amino acid (NuoE-like) YGSHKKDSELDVQEGEIVRIFCPHCNVDMKSSRFCYECKAPMITF ESLLGGYIRVCSRWGCKKQLAEFENLETELRAFHAKYSLSPQGR GEK* 615 VMTD01000047.1_44, HydC MDTKDKFIIDTILEKFTDMKDPTLMILQNIQEILGCVREEYLTYVAES amino acid (NuoE-like) NGYSLTELYGIVSFYPQFRLSPPGKHTIKICKGTACHVRGGAAIQK NLQNLLGIKPGDTTEDGVFTLESVRCLGCCGLSPVIMIDNETYGR VKTSKLQEILVKYKGVSANENQ* 620 WOYO01000071.1_29, HydC MDVKDRYIIDTIIEKCEGVRDPALIILQNVQEILGYVGKEYLEYVSEK amino acid (NuoE-like) SGYPLVELYGIITFYPQFKLNPPGKNTIKVCQGTACHVRGSDRILG TLEELLKISPGETTENRLFSLESVRCLGCCGLAPTIMINKKTYGRV KPSNLKNILAEYEEVDQYVVE* Table 5: Combinations of HydA, HydB, HydC, HydD, and HyhL Ro Nucleotide SEQ ID NOs: Amino Acid SEQ ID NOs: ws 1005373994
Description Subgroup, HydA HydB HydC HydD HyhL HydA HydB HydC HydD HyhL subclass 1 AQRS01000037.1 A1, M1 1, - 301 - - 151, - 464 - - 137, 287, 144 294 2 09mi20_z1_2019_ig18392 A3, M3 30 303 302 - - 180 466 465 - - _10016 3 AQSC01000060.1 A3, M3 32 304 305 - - 182 467 468 - - 4 CAITEO010000066.1 A3, M3 33 307 306 - - 183 470 469 - - 5 CAIYYO010000253.1 A3, M3 34 308 - - - 184 471 - - - 6 Candidatus_Lokiarchaeota A3, M3 35 309 310 - - 185 472 473 - - _archaeon_FW102|JAIZW K010000001 7 Candidatus_Lokiarchaeota A3, M3 37 311 312 - - 187 474 475 - - _archaeon_strain_B53_G9 |QMYW01000078 8 Candidatus_Lokiarchaeota A3, M3 38 313 - - - 188 476 - - - _archaeon_strain_CSSed 165cm_327R1|SLLX0100 0174 9 Candidatus_Lokiarchaeota A3, M3 39 314 315 - - 189 477 478 - - _archaeon_strain_HM1_B 6_4|JABCSI010000064 10 DAWM01000035.1 A3, M3 40 316 317 - - 190 479 480 - - 11 DSAL01000038.1 A3, M3 41 318 319 - - 191 481 482 - - 12 DSBS01000042.1 A3, M3 42 320 321 - - 192 483 484 - - 13 DSUA01000124.1 A3, M3 43 322 - - - 193 485 - - - 14 DSVV01000048.1 A3, M3 44 324 323 - - 194 487 486 - - 15 DUIB01000037.1 A3, M3 45 326 325 - - 195 489 488 - - 16 LacPavin_0920_SED5_sc A3, M3 46 327 328 - - 196 490 491 - - affold_1418049 17 PEXD01000050.1 A3, M3 47 330 329 - - 197 493 492 - - 18 PEXL01000024.1 A3, M3 48 332 331 - - 198 495 494 - - 19 PGXE01000078.1 A3, M3 49 334 333 - - 199 497 496 - - 20 PXDW01000025.1 A3, M3 50 336 337 - 335 200 499 500 - 498 21 SR- A3, M3 51 339 338 - - 201 502 501 - - 2_scaffold_141_4160621 22 AB_3033_bin_101_scaffol A3, M3' 52 341 340 342 - 202 504 503 505 - d_9312 23 AB_1215_Bin_111.fna_sc A3, M3c 53 343, 345 - - 203 506, 508 - - affold_95187 344 507 24 CABMGE010000002.1 A3, M3c 54 347 346 - - 204 510 509 - - 25 CAITKI010000066.1 A3, M3c 55 348 349 - - 205 511 512 - - 26 CAITNU010000081.1 A3, M3c 56 350 351 - - 206 513 514 - - 27 JAACWB010000015.1 A3, M3c 57 353 352 - - 207 516 515 - - 28 JACCLF010000048.1 A3, M3c 58 354 355 - - 208 517 518 - - 29 JACCLG010000046.1 A3, M3c 59 357 356 - - 209 520 519 - - 30 Meg22_1012_Bin_199_sc A3, M3c 61 358 - - - 211 521 - - - affold_273895 31 Meg22_1012_Bin_224_sc A3, M3c 62 360 359 - - 212 523 522 - - affold_4514 32 MWBV01000007.1 A3, M3c 63 361 362 - - 213 524 525 - - 33 NJBG01000001.1 A3, M3c 64 363 364 - - 214 526 527 - - 34 PCYE01000023.1 A3, M3c 65 366 365 - - 215 529 528 - - 1005373994
35 QMVA01000111.1 A3, M3c 66 367 - - - 216 530 - - - 36 QMVB01000024.1 A3, M3c 67 368 - - - 217 531 - - - 37 QMZS01000067.1 A3, M3c 68 370 369 - - 218 533 532 - - 38 QMZV01000029.1 A3, M3c 69 371 372 - - 219 534 535 - - 39 Zodletone_Water_assembl A3, M3c 70 374 373 - - 220 537 536 - - y2018_k141_1542890 40 Candidatus_Lokiarchaeota B, M3a' 72 - 375 - - 222 - 538 - - _archaeon_bin106|JAGXO A010000065 41 Candidatus_Lokiarchaeota B, M3a' 73 - 376 - - 223 - 539 - - _archaeon_FW102|JAIZW K010000001 42 Candidatus_Lokiarchaeota B, M3a' 74 - 377 - - 224 - 540 - - _archaeon_strain_AS27yj COA_147|JAAZNI010000 050 43 Candidatus_Lokiarchaeota B, M3a' 75 - 378 - - 225 - 541 - - _archaeon_strain_B53_G9 |QMYW01000085 44 Candidatus_Lokiarchaeota B, M3a' 76 - 379 - - 226 - 542 - - _archaeon_strain_HM1_B 6_4|JABCSI010000124 45 Candidatus_Lokiarchaeota B, M3a' 77 - 380 - - 227 - 543 - - _archaeon_strain_HM4_B 48|JABCTF010000130 46 Candidatus_Lokiarchaeota B, M3a' 78 - 381 - - 228 - 544 - - _archaeon_strain_MAG_1 4|DXJH01000338 47 Candidatus_Lokiarchaeota B, M3a' 79 - 382 - - 229 - 545 - - _archaeon_strain_YT2_01 2|JAEOTN010000030 48 Candidatus_Lokiarchaeota B, M3a' 80 - 383 - - 230 - 546 - - _archaeon_strain_Zod_Me tabat.1044|JAFGBS01000 0008 49 Candidatus_Lokiarchaeota B, M3a' 81 - 384 - - 231 - 547 - - _archaeon_strain_Zod_Me tabat.578|JAFGOA010000 117 50 NZ_CP042905.1 B, M3a' 82, - 385 - - 232, - 548 - - 141, 291, 148 298 51 Candidatus_Odinarchaeot F, M3d 113 386 - 387 388 263 549 - 550 551 a_archaeon_AUK265|JAG HBV010000041 52 DAEG01000100.1 F, M3d 114 390 389 391 392 264 553 552 554 555 53 DAJR01000051.1 F, M3d 115, 394 393 - - 265, 557 556 - - 142, 292, 149 299 54 DRVA01000011.1 F, M3d 116 396 395 397 398 266 559 558 560 561 55 DRYW01000042.1 F, M3d 117 400 399 401 402 267 563 562 564 565 56 DTBU01000097.1 F, M3d 119 404 405 403 - 269 567 568 566 - 57 DTCV01000006.1 F, M3d 120 407 406 408 409 270 570 569 571 572 58 DTEB01000060.1 F, M3d 121 411 410 412 413 271 574 573 575 576 59 DTEJ01000176.1 F, M3d 122 - - - 414 272 - - - 577 60 DTOX01000047.1 F, M3d 123 416 415 417 418 273 579 578 580 581 61 DTQO01000091.1 F, M3d 124 422 423 421 419, 274 585 586 584 582, 420 583 62 GMQ_scaffold_8 F, M3d 125 425 424 426 427 275 588 587 589 590 1005373994
63 GMQ_scaffold_93 F, M3d 126 429 428 430 431 276 592 591 593 594 64 NJDO01000061.1 F, M3d 127 433 432 434 435 277 596 595 597 598 65 PKYH01000114.1 F, M3d 128 438 439 437 436 278 601 602 600 599 66 QMRZ01000231.1 F, M3d 129 - - - 440 279 - - - 603 67 SR- F, M3d 130 443 444 442 441 280 606 607 605 604 2_scaffold_141_2318042 68 SRVP18_trench_6_60cm_ F, M3d 131 447 448 446 445 281 610 611 609 608 scaffold_130 69 VMTD01000047.1 F, M3d 132, 450, 449, 451, 455, 282, 613, 612, 614, 618, 143, 453 452 454 456 293, 616 615 617 619 150 300 70 WOYO01000071.1 F, M3d 133 458 457 459 460 283 621 620 622 623 71* 1_2_3_4_103298 G, M3 134 - - - 461 284 - - - 624 72* DRTY7_scaffold_969_cur G, M3 135 - - - 462 285 - - - 625 ated 73* JAAOZO010000036.1 G, M3 136 - - - 463 286 - - - 626 * each of rows 71 to 73 further include a HyhS. Row 71 further includes nucleotide SEQ ID NO: 628 and amino acid SEQ ID NO: 631. Row 72 further includes nucleotide SEQ ID NO: 629 and amino acid SEQ ID NO: 632. Row 73 further includes nucleotide SEQ ID NO: 627 and amino acid SEQ ID NO: 630 Detailed description of the embodiments [0110] It will be understood that the invention disclosed and defined in this specification extends to all alternative combinations of two or more of the individual features mentioned or evident from the text or drawings. All of these different combinations constitute various alternative aspects of the invention. [0111] Reference will now be made in detail to certain embodiments of the invention. While the invention will be described in conjunction with the embodiments, it will be understood that the intention is not to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the scope of the present invention as defined by the claims. [0112] One skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which could be used in the practice of the present invention. The present invention is in no way limited to the methods and materials described. It will be understood that the invention disclosed and defined in this specification extends to all alternative combinations of two or more of the individual features mentioned or evident from the text or drawings. All of these different combinations constitute various alternative aspects of the invention. 1005373994
[0113] All of the patents and publications referred to herein are incorporated by reference in their entirety. [0114] For purposes of interpreting this specification, terms used in the singular will also include the plural and vice versa. [0115] The inventors overturn three long-held assumptions featured in the introductions of almost every hydrogenase paper and book chapter. While thought to be restricted to bacteria and eukaryotes, here the inventors show that [FeFe]-hydrogenases are also widespread and active in archaea. Secondly, by purifying and characterising an ultraminimal hydrogenase from DPANN archaea, the inventors redefine the size requirements for hydrogen biocatalysis. Thirdly, while long thought that the two dominant hydrogenase classes (NiFe, FeFe) independently evolved, here the inventors show that these enzymes have a deep evolutionary association and even form hybrids in some complex archaea. In doing so, the inventord expand [FeFe]-hydrogenases from four to seven groups each with distinct phylogenies, structures, and functions. [0116] The inventors’ metabolic reconstructions suggest that [FeFe]-hydrogenases enable archaea to mediate fermentation and electron-bifurcation. This allows them to conserve energy and maintain redox balance in the diverse anoxic environments we sampled them from, such as groundwater and hot spring sediments. DPANN have notably evolved ultraminimal fermentative [FeFe]-hydrogenases, providing them with a genomically streamlined way to efficiently dispose excess electrons produced during carbohydrate fermentation. It is also remarkable that even the Asgard archaeon expresses a hitherto-overlooked [FeFe]-hydrogenase, which likely enables it to mediate syntrophic hydrogen exchange with its methanogenic partner. [0117] Here the inventors heterologously purified and biochemically, electrochemically, and spectroscopically characterised unique groups of [FeFe]-hydrogenases from archaea only known for their genomes. This allowed the inventors to prove that these enzymes are active, bind the hydrogen-converting metal centre, and are catalytically-biased towards hydrogen production. Nucleic acids [0118] As used herein an "isolated" nucleic acid molecule is a nucleic acid molecule that is identified and separated from at least one contaminant nucleic acid molecule with which 1005373994
it is ordinarily associated in the natural source of the polypeptide encoding nucleic acid. An isolated nucleic acid molecule is other than in the form or setting in which it is found in nature. Isolated nucleic acid molecules therefore are distinguished from the nucleic acid molecule as it exists in natural cells. However, an isolated nucleic acid molecule includes nucleic acid molecules contained in cells that ordinarily express the nucleic acid where, for example, the nucleic acid molecule is in a chromosomal location different from that of natural cells. [0119] The terms “nucleic acid molecule” and “polynucleotide” may be used interchangeably herein and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogues thereof. Non-limiting examples of polynucleotides include a gene, a gene fragment, messenger RNA (mRNA), cDNA, recombinant polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A nucleic acid sequence which “encodes” a selected polypeptide is a nucleic acid molecule which is transcribed (in the case of DNA) and translated into a polypeptide in vivo when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a start codon at the 5′ (amino) terminus and a translation stop codon at the 3′ (carboxy) terminus. A transcription termination sequence may be located 3′ to the coding sequence. [0120] Polynucleotides of the invention can be synthesised according to methods well known in the art, as described by way of example in Sambrook et al (1989, Molecular Cloning—a laboratory manual; Cold Spring Harbor Press). [0121] As used herein, “codon optimised” refers to optimisation of the DNA sequence to resemble the codon usage of genes in host microorganism. In preferred embodiments, the codon usage in the sequence is optimised to resemble that of highly expressed E. coli genes. [0122] The polynucleotide molecules of the present invention may be provided in the form of an expression cassette which includes control sequences operably linked to the inserted sequence, thus allowing for expression of the polypeptide. These expression cassettes, in turn, are typically provided within vectors (e.g., plasmids or recombinant vectors). A suitable vector may be any vector which is capable of carrying a sufficient amount of genetic information, and allowing expression of a polypeptide of the invention. 1005373994
[0123] The present invention thus includes expression vectors that comprise such polynucleotide sequences. Expression vectors are routinely constructed in the art of molecular biology and may for example involve the use of plasmid DNA and appropriate initiators, promoters, enhancers and other elements which may be necessary, and which are positioned in the correct orientation, in order to allow for expression of a desired polypeptide. Other suitable vectors would be apparent to persons skilled in the art. By way of further example in this regard we refer to Sambrook et al. [0124] Thus, a polypeptide of the invention may be provided by delivering such a vector to a cell and allowing transcription from the vector to occur. The skilled person will be familiar with standard techniques for delivery such expression vectors to a cell, including transformation techniques and the like. [0125] The vector may be a plasmid. In certain embodiments, the plasmid is a high copy number plasmid or a low copy number plasmid. Vectors are well known in the art and may include cloning vectors, expression vectors, etc. A cloning vector is a recombinant nucleic acid construct which is able to replicate autonomously or integrated in the genome in a host cell, and which is further characterized by one or more endonuclease restriction sites at which the vector may be cut in a determinable fashion and into which a desired DNA sequence may be ligated such that the new recombinant vector retains its ability to replicate in the host cell. In the case of plasmids, replication of the desired sequence may occur many times as the plasmid increases in copy number within the host bacterium or just a single time per host before the host reproduces by mitosis. In the case of phage, replication may occur actively during a lytic phase or passively during a lysogenic phase. An expression vector is a recombinant nucleic acid construct into which a desired DNA sequence may be inserted by restriction and ligation such that it is operably joined to regulatory sequences and may be expressed as an RNA transcript. Vectors may further contain one or more marker sequences suitable for use in the identification of cells which have or have not been transformed or transfected with the vector. Markers include, for example, genes encoding proteins which increase or decrease either resistance or sensitivity to antibiotics or other compounds, genes which encode polypeptides or enzymes whose activities are detectable by standard assays known in the art (e.g., β- galactosidase, luciferase or alkaline phosphatase), and genes which visibly affect the phenotype of transformed or transfected cells, hosts, colonies or plaques (e.g., fluorescent proteins such as green fluorescent protein). Preferred vectors are those 1005373994
capable of autonomous replication and expression of the structural gene products present in the DNA segments to which they are operably joined. [0126] As used herein, a coding sequence and regulatory sequences are said to be "operably" joined or linked when they are covalently linked in such a way as to place the expression or transcription of the coding sequence under the influence or control of the regulatory sequences. If it is desired that the coding sequences be translated into a functional protein, two DNA sequences are said to be operably joined or linked if induction of a promoter in the 5' regulatory sequences results in the transcription of the coding sequence and if the nature of the linkage between the two DNA sequences does not (1) result in the introduction of a frame-shift mutation, (2) interfere with the ability of the promoter region to direct the transcription of the coding sequences, or (3) interfere with the ability of the corresponding RNA transcript to be translated into a protein. Thus, a promoter region would be operably joined or linked to a coding sequence if the promoter region were capable of effecting transcription of that DNA sequence such that the resulting transcript can be translated into the desired protein or polypeptide. [0127] The precise nature of the regulatory sequences needed for gene expression may vary between species or cell types, but shall in general include, as necessary, 5' non- transcribed and 5' non-translated sequences involved with the initiation of transcription and translation respectively, such as a TATA box, capping sequence, CAAT sequence, and the like. In particular, such 5' non-transcribed regulatory sequences will include a promoter region which includes a promoter sequence for transcriptional control of the operably joined gene. [0128] Regulatory sequences may also include enhancer sequences or upstream activator sequences as desired. The vectors of the invention may optionally include 5' leader or signal sequences. The choice and design of an appropriate vector is within the ability and discretion of one of ordinary skill in the art. [0129] A “promoter” is a nucleotide sequence which initiates and regulates transcription of a polypeptide-encoding polynucleotide. Promoters can include inducible promoters (where expression of a polynucleotide sequence operably linked to the promoter is induced by an analyte, cofactor, regulatory protein, etc.), repressible promoters (where expression of a polynucleotide sequence operably linked to the promoter is repressed by an analyte, cofactor, regulatory protein, etc.), and constitutive promoters. It is intended 1005373994
that the term “promoter” or “control element” includes full-length promoter regions and functional (e.g., controls transcription or translation) segments of these regions. [0130] The nucleic acids of the present invention are preferably operably linked to promoters such that the subject enzymes are expressed in the cell when cultured under suitable conditions for enabling consumption of hydrogen, as described herein. The promoters may be specific for individual bacterial cell species. The promoter may be a heterologous promoter which increases the expression of the gene above the typical expression level observed in the cell. The promoter may be an inducible promoter. [0131] A polynucleotide, expression cassette or vector according to the present invention may additionally comprise a signal peptide sequence. The signal peptide sequence is generally inserted in operable linkage with the promoter such that the signal peptide is expressed and facilitates secretion of a polypeptide encoded by coding sequence also in operable linkage with the promoter. It may further be understood that in any embodiment, any of the exemplary expression cassettes, vectors or sequences described herein may be further modified so as to not include a signal peptide sequence. [0132] Any appropriate expression vector (e.g., as described in Pouwels et al., Cloning Vectors: A Laboratory Manual (Elsevier, N.Y.: 1985)) and corresponding suitable host can be employed for production of recombinant polypeptides. Expression hosts include, but are not limited to, bacterial species within the genera Escherichia, Bacillus, Pseudomonas, Salmonella, host cell systems and the like. The skilled person is aware that the choice of expression host has ramifications for the type of polypeptide produced. [0133] In some embodiments, the cell is engineered or selected (e.g., as described herein) to produce or have altered, optionally increased, production of a molecule of interest. In some embodiments, the cell comprises a deletion or mutation of one or more genes (e.g., one or more regulatory or competing metabolic genes as described herein). In other examples, the one or more genes that are deleted or mutated are in a competing pathway. Mutations can be single or multiple point mutations, additions, partial internal deletions, N-terminal or C-terminal deletions (truncations), or complete deletions, all of which can affect amino acid sequence encoded the gene(s). [0134] Deletions or mutations can be made using standard methods in the art. Mutations can be non-random, partially random or random, or a combination of these mutations. 1005373994
For example, for a partially random mutation, the mutation(s) may be confined to a certain portion of the nucleic acid molecule encoding a polypeptide in which mutation(s) are to be made. Protein production and purification [0135] The terms “isolated" or “purified” when used to describe the various polypeptides disclosed herein, mean the polypeptide that has been identified and separated and/or recovered from a component of its natural environment. Contaminant components of its natural environment are materials that would typically interfere with diagnostic or therapeutic uses for the polypeptide, and may include enzymes, hormones, and other proteinaceous or non-proteinaceous solutes. In preferred embodiments, the polypeptide will be purified (1) to a degree sufficient to obtain at least 15 residues of N-terminal or internal amino acid sequence by use of a spinning cup sequenator, or (2) to homogeneity by SDS-PAGE under non-reducing or reducing conditions using Coomassie blue or, preferably, silver stain. Isolated protein includes polypeptide in situ within recombinant cells, since at least one component of the polypeptide natural environment will not be present. Ordinarily, however, isolated polypeptide will be prepared by at least one purification step. [0136] A "fragment" is a portion of a polypeptide of the present invention that retains substantially similar functional activity or substantially the same biological function or activity as the polypeptide, which can be determined using assays described herein. [0137] “Percent (%) amino acid sequence identity” or “percent (%) identical” with respect to a polypeptide sequence, i.e. a polypeptide of the invention defined herein, is defined as the percentage of amino acid residues in a candidate sequence that are identical with the amino acid residues in the specific polypeptide of the invention, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and not considering any conservative substitutions as part of the sequence identity. [0138] Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms (non-limiting examples described below) needed to achieve maximal alignment over the full-length of the sequences being compared. When amino acid sequences are aligned, the percent amino acid sequence identity of a given 1005373994
amino acid sequence A to, with, or against a given amino acid sequence B (which can alternatively be phrased as a given amino acid sequence A that has or comprises a certain percent amino acid sequence identity to, with, or against a given amino acid sequence B) can be calculated as: percent amino acid sequence identity = X/Y100, where X is the number of amino acid residues scored as identical matches by the sequence alignment program's or algorithm's alignment of A and B and Y is the total number of amino acid residues in B. If the length of amino acid sequence A is not equal to the length of amino acid sequence B, the percent amino acid sequence identity of A to B will not equal the percent amino acid sequence identity of B to A. [0139] In calculating percent identity, typically exact matches are counted. The determination of percent identity between two sequences can be accomplished using a mathematical algorithm. A nonlimiting example of a mathematical algorithm utilized for the comparison of two sequences is the algorithm of Karlin and Altschul (1990) Proc. Natl. Acad. Sci. USA 87:2264, modified as in Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877. Such an algorithm is incorporated into the BLASTN and BLASTX programs of Altschul et al. (1990) J. MoI. Biol.215:403. To obtain gapped alignments for comparison purposes, Gapped BLAST (in BLAST 2.0) can be utilized as described in Altschul et al. (1997) Nucleic Acids Res.25:3389. Alternatively, PSI-Blast can be used to perform an iterated search that detects distant relationships between molecules. See Altschul et al. (1997) supra. When utilizing BLAST, Gapped BLAST, and PSI-Blast programs, the default parameters of the respective programs (e.g., BLASTX and BLASTN) can be used. Alignment may also be performed manually by inspection. Another non- limiting example of a mathematical algorithm utilized for the comparison of sequences is the ClustalW algorithm (Higgins et al. (1994) Nucleic Acids Res.22:4673- 4680). ClustalW compares sequences and aligns the entirety of the amino acid or DNA sequence, and thus can provide data about the sequence conservation of the entire amino acid sequence. The ClustalW algorithm is used in several commercially available DNA/amino acid analysis software packages, such as the ALIGNX module of the Vector NTI Program Suite (Invitrogen Corporation, Carlsbad, CA). After alignment of amino acid sequences with ClustalW, the percent amino acid identity can be assessed. A non-limiting example of a software program useful for analysis of ClustalW alignments is GENEDOC™ or JalView (http://www.jalview.org/). GENEDOC™ allows assessment of amino acid (or DNA) similarity and identity between multiple proteins. Another non- limiting example of a mathematical algorithm utilized for the comparison of sequences is 1005373994
the algorithm of Myers and Miller (1988) CABIOS 4:11-17. Such an algorithm is incorporated into the ALIGN program (version 2.0), which is part of the GCG Wisconsin Genetics Software Package, Version 10 (available from Accelrys, Inc., 9685 Scranton Rd., San Diego, CA, USA). When utilizing the ALIGN program for comparing amino acid sequences, a PAM 120 weight residue table, a gap length penalty of 12, and a gap penalty of 4 can be used. [0140] The polypeptide desirably comprises an amino end and a carboxyl end. The polypeptide can comprise D-amino acids, L-amino acids or a mixture of D- and L-amino acids. The D-form of the amino acids, however, is particularly preferred since a polypeptide comprised of D-amino acids is expected to have a greater retention of its biological activity in vivo. [0141] The polypeptide can be prepared by any of a number of conventional techniques. The polypeptide can be isolated or purified from a naturally occurring source or from a recombinant source. Recombinant production is preferred. For instance, in the case of recombinant polypeptides, a DNA fragment encoding a desired peptide can be subcloned into an appropriate vector using well-known molecular genetic techniques (see, e.g., Maniatis et al., Molecular Cloning: A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory, 1982); Sambrook et al., Molecular Cloning A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory, 1989). The fragment can be transcribed and the polypeptide subsequently translated in vitro. Commercially available kits also can be employed (e.g., such as manufactured by Clontech, Palo Alto, Calif.; Amersham Pharmacia Biotech Inc., Piscataway, N.J.; InVitrogen, Carlsbad, Calif., and the like). The polymerase chain reaction optionally can be employed in the manipulation of nucleic acids. [0142] The term "conservative substitution" as used herein, refers to the replacement of an amino acid present in the native sequence in the peptide with a naturally or non- naturally occurring amino acid or a peptidomimetic having similar steric properties. Where the side-chain of the native amino acid to be replaced is either polar or hydrophobic, the conservative substitution should be with a naturally occurring amino acid, a non- naturally occurring amino acid or with a peptidomimetic moiety which is also polar or hydrophobic (in addition to having the same steric properties as the side-chain of the replaced amino acid). 1005373994
[0143] Conservative amino acid substitution tables providing functionally similar amino acids are well known to one of ordinary skill in the art. The following six groups are examples of amino acids that may be considered to be conservative substitutions for one another: [0144] 1) Alanine (A), Serine (S), Threonine (T); [0145] 2) Aspartic acid (D), Glutamic acid (E); [0146] 3) Asparagine (N), Glutamine (Q); [0147] 4) Arginine (R), Lysine (K); [0148] 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); and [0149] 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W). [0150] As naturally occurring amino acids are typically grouped according to their properties, conservative substitutions by naturally occurring amino acids can be determined bearing in mind the fact that replacement of charged amino acids by sterically similar non-charged amino acids are considered as conservative substitutions. For producing conservative substitutions by non-naturally occurring amino acids it is also possible to use amino acid analogs (synthetic amino acids) well known in the art. A peptidomimetic of the naturally occurring amino acid is well documented in the literature known to the skilled person and non-natural or unnatural amino acids are described further below. When affecting conservative substitutions, the substituting amino acid should have the same or a similar functional group in the side chain as the original amino acid. [0151] The phrase "non-conservative substitution" or a “non-conservative residue” as used herein refers to replacement of the amino acid as present in the parent sequence by another naturally or non-naturally occurring amino acid, having different electrochemical and/or steric properties. Thus, the side chain of the substituting amino acid can be significantly larger (or smaller) than the side chain of the native amino acid being substituted and/or can have functional groups with significantly different electronic properties than the amino acid being substituted. Examples of non-conservative substitutions of this type include the substitution of phenylalanine or cycohexylmethyl glycine for alanine, isoleucine for glycine, or -NH-CH[(-CH2)5-COOH]-CO- for aspartic 1005373994
acid. Non-conservative substitution includes any mutation that is not considered conservative. [0152] A non-conservative amino acid substitution can result from changes in: (a) the structure of the amino acid backbone in the area of the substitution; (b) the charge or hydrophobicity of the amino acid; or (c) the bulk of an amino acid side chain. Substitutions generally expected to produce the greatest changes in protein properties are those in which: (a) a hydrophilic residue is substituted for (or by) a hydrophobic residue; (b) a proline is substituted for (or by) any other residue; (c) a residue having a bulky side chain, e.g., phenylalanine, is substituted for (or by) one not having a side chain, e.g., glycine; or (d) a residue having an electropositive side chain, e.g., lysyl, arginyl, or histadyl, is substituted for (or by) an electronegative residue, e.g., glutamyl or aspartyl. [0153] Alterations of the native amino acid sequence to produce mutant polypeptides, such as by insertion, deletion and/or substitution, can be done by a variety of means known to those skilled in the art. For instance, site-specific mutations can be introduced by ligating into an expression vector a synthesized oligonucleotide comprising the modified site. Alternately, oligonucleotide-directed site-specific mutagenesis procedures can be used, such as disclosed in Walder et al., Gene 42: 133 (1986); Bauer et al., Gene 37: 73 (1985); Craik, Biotechniques, 12-19 (January 1995); and U.S. Pat. Nos.4,518,584 and 4,737,462. A preferred means for introducing mutations is the QuikChange Site- Directed Mutagenesis Kit (Stratagene, LaJolla, Calif.). [0154] The terms "N-terminal" and "C-terminal" are used herein to designate the relative position of any amino acid sequence or polypeptide domain or structure to which they are applied. The relative positioning will be apparent from the context. That is, an "N-terminal" feature will be located at least closer to the N-terminus of the polypeptide molecule than another feature discussed in the same context (the other feature possible referred to as "C-terminal" to the first feature). Similarly, the terms "5'-" and "3'-" can be used herein to designate relative positions of features of polynucleotides. [0155] A recombinant polypeptide made in accordance with the methods of the present invention may also be modified by, conjugated or fused to another moiety to facilitate purification of the polypeptides, or for use in enzymatic assays using methods known in the art. For example, a polypeptide of the invention may be modified by glycosylation, 1005373994
acetylation, phosphorylation, amidation, derivatization by known protecting/blocking groups, proteolytic cleavage, etc. [0156] Modifications contemplated herein include, but are not limited to, modification to side chains, incorporating of unnatural amino acids and/or their derivatives during polypeptide synthesis and the use of crosslinkers and other methods which impose conformational constraints on the polypeptides of the invention. [0157] Examples of side chain modifications contemplated by the present invention include modifications of amino groups such as by reductive alkylation by reaction with an aldehyde followed by reduction with NaBH4; amidination with methylacetimidate; acylation with acetic anhydride; carbamoylation of amino groups with cyanate; trinitrobenzylation of amino groups with 2, 4, 6-trinitrobenzene sulphonic acid (TNBS); acylation of amino groups with succinic anhydride and tetrahydrophthalic anhydride; and pyridoxylation of lysine with pyridoxal-5-phosphate followed by reduction with NaBH4. [0158] The guanidine group of arginine residues may be modified by the formation of heterocyclic condensation products with reagents such as 2,3-butanedione, phenylglyoxal and glyoxal. [0159] The carboxyl group may be modified by carbodiimide activation via O- acylisourea formation followed by subsequent derivatisation, for example, to a corresponding amide. [0160] Sulphydryl groups may be modified by methods such as carboxymethylation with iodoacetic acid or iodoacetamide; performic acid oxidation to cysteic acid; formation of a mixed disulphides with other thiol compounds; reaction with maleimide, maleic anhydride or other substituted maleimide; formation of mercurial derivatives using 4- chloromercuribenzoate, 4-chloromercuriphenylsulphonic acid, phenylmercury chloride, 2- chloromercuri-4-nitrophenol and other mercurials; carbamoylation with cyanate at alkaline pH. [0161] Tryptophan residues may be modified by, for example, oxidation with N- bromosuccinimide or alkylation of the indole ring with 2-hydroxy-5-nitrobenzyl bromide or sulphenyl halides. Tyrosine residues on the other hand, may be altered by nitration with tetranitromethane to form a 3-nitrotyrosine derivative. 1005373994
[0162] Modification of the imidazole ring of a histidine residue may be accomplished by alkylation with iodoacetic acid derivatives or N-carboethoxylation with diethylpyrocarbonate. [0163] Examples of incorporating unnatural amino acids and derivatives during protein synthesis include, but are not limited to, use of norleucine, 4-amino butyric acid, 4-amino- 3-hydroxy-5-phenylpentanoic acid, 6-aminohexanoic acid, t-butylglycine, norvaline, phenylglycine, ornithine, sarcosine, 4-amino-3-hydroxy-6-methylheptanoic acid, 2-thienyl alanine and/or D-isomers of amino acids. A list of unnatural amino acids contemplated herein is shown in Table 6. [0164] Table 6 ______________________________________________________________________ Non-conventional Code Non-conventional Code amino acid amino acid ______________________________________________________________________ α-aminobutyric acid Abu L-N-methylalanine Nmala α-amino-α-methylbutyrate Mgabu L-N-methylarginine Nmarg aminocyclopropane- Cpro L-N-methylasparagine Nmasn carboxylate L-N-methylaspartic acid Nmasp aminoisobutyric acid Aib L-N-methylcysteine Nmcys aminonorbornyl- Norb L-N-methylglutamine Nmgln carboxylate L-N-methylglutamic acid Nmglu cyclohexylalanine Chexa L-N-methylhistidine Nmhis cyclopentylalanine Cpen L-N-methylisolleucine Nmile D-alanine Dal L-N-methylleucine Nmleu D-arginine Darg L-N-methyllysine Nmlys D-aspartic acid Dasp L-N-methylmethionine Nmmet D-cysteine Dcys L-N-methylnorleucine Nmnle D-glutamine Dgln L-N-methylnorvaline Nmnva D-glutamic acid Dglu L-N-methylornithine Nmorn D-histidine Dhis L-N-methylphenylalanine Nmphe D-isoleucine Dile L-N-methylproline Nmpro D-leucine Dleu L-N-methylserine Nmser D-lysine Dlys L-N-methylthreonine Nmthr 1005373994
D-methionine Dmet L-N-methyltryptophan Nmtrp D-ornithine Dorn L-N-methyltyrosine Nmtyr D-phenylalanine Dphe L-N-methylvaline Nmval D-proline Dpro L-N-methylethylglycine Nmetg D-serine Dser L-N-methyl-t-butylglycine Nmtbug D-threonine Dthr L-norleucine Nle D-tryptophan Dtrp L-norvaline Nva D-tyrosine Dtyr α-methyl-aminoisobutyrate Maib D-valine Dval α-methyl-γ-aminobutyrate Mgabu D-α-methylalanine Dmala α-methylcyclohexylalanine Mchexa D-α-methylarginine Dmarg α-methylcylcopentylalanine Mcpen D-α-methylasparagine Dmasn α-methyl-α-napthylalanine Manap D-α-methylaspartate Dmasp α-methylpenicillamine Mpen D-α-methylcysteine Dmcys N-(4-aminobutyl)glycine Nglu D-α-methylglutamine Dmgln N-(2-aminoethyl)glycine Naeg D-α-methylhistidine Dmhis N-(3-aminopropyl)glycine Norn D-α-methylisoleucine Dmile N-amino-α-methylbutyrate Nmaabu D-α-methylleucine Dmleu α-napthylalanine Anap D-α-methyllysine Dmlys N-benzylglycine Nphe D-α-methylmethionine Dmmet N-(2-carbamylethyl)glycine Ngln D-α-methylornithine Dmorn N-(carbamylmethyl)glycine Nasn D-α-methylphenylalanine Dmphe N-(2-carboxyethyl)glycine Nglu D-α-methylproline Dmpro N-(carboxymethyl)glycine Nasp D-α-methylserine Dmser N-cyclobutylglycine Ncbut D-α-methylthreonine Dmthr N-cycloheptylglycine Nchep D-α-methyltryptophan Dmtrp N-cyclohexylglycine Nchex D-α-methyltyrosine Dmty N-cyclodecylglycine Ncdec D-α-methylvaline Dmval N-cylcododecylglycine Ncdod D-N-methylalanine Dnmala N-cyclooctylglycine Ncoct D-N-methylarginine Dnmarg N-cyclopropylglycine Ncpro D-N-methylasparagine Dnmasn N-cycloundecylglycine Ncund D-N-methylaspartate Dnmasp N-(2,2-diphenylethyl)glycine Nbhm D-N-methylcysteine Dnmcys N-(3,3-diphenylpropyl)glycine Nbhe D-N-methylglutamine Dnmgln N-(3-guanidinopropyl)glycine Narg 1005373994
D-N-methylglutamate Dnmglu N-(1-hydroxyethyl)glycine Nthr D-N-methylhistidine Dnmhis N-(hydroxyethyl))glycine Nser D-N-methylisoleucine Dnmile N-(imidazolylethyl))glycine Nhis D-N-methylleucine Dnmleu N-(3-indolylyethyl)glycine Nhtrp D-N-methyllysine Dnmlys N-methyl-γ-aminobutyrate Nmgabu N-methylcyclohexylalanineNmchexa D-N-methylmethionine Dnmmet D-N-methylornithine Dnmorn N-methylcyclopentylalanine Nmcpen N-methylglycine Nala D-N-methylphenylalanine Dnmphe N-methylaminoisobutyrate Nmaib D-N-methylproline Dnmpro N-(1-methylpropyl)glycine Nile D-N-methylserine Dnmser N-(2-methylpropyl)glycine Nleu D-N-methylthreonine Dnmthr D-N-methyltryptophan Dnmtrp N-(1-methylethyl)glycine Nval D-N-methyltyrosine Dnmtyr N-methyla-napthylalanine Nmanap D-N-methylvaline Dnmval N-methylpenicillamine Nmpen γ-aminobutyric acid Gabu N-(p-hydroxyphenyl)glycine Nhtyr L-t-butylglycine Tbug N-(thiomethyl)glycine Ncys L-ethylglycine Etg penicillamine Pen L-homophenylalanine Hphe L-α-methylalanine Mala L-α-methylarginine Marg L-α-methylasparagine Masn L-α-methylaspartate Masp L-α-methyl-t-butylglycine Mtbug L-α-methylcysteine Mcys L-methylethylglycine Metg L-α-methylglutamine Mgln L-α-methylglutamate Mglu L-α-methylhistidine Mhis L-α-methylhomophenylalanine Mhphe L-α-methylisoleucine Mile N-(2-methylthioethyl)glycine Nmet L-α-methylleucine Mleu L-α-methyllysine Mlys L-α-methylmethionine Mmet L-α-methylnorleucine Mnle L-α-methylnorvaline Mnva L-α-methylornithine Morn L-α-methylphenylalanine Mphe L-α-methylproline Mpro L-α-methylserine Mser L-α-methylthreonine Mthr L-α-methyltryptophan Mtrp L-α-methyltyrosine Mtyr L-α-methylvaline Mval L-N-methylhomophenylalanine Nmhphe N-(N-(2,2-diphenylethyl) Nnbhm N-(N-(3,3-diphenylpropyl) Nnbhe carbamylmethyl)glycine carbamylmethyl)glycine 1-carboxy-1-(2,2-diphenyl-Nmbc 1005373994
ethylamino)cyclopropane Culturing and modification of microorganisms [0165] In particularly preferred embodiments, culturing of the microorganisms or host cells, as described herein, may be performed under aerobic and/or anaerobic conditions. Certain steps may be performed under aerobic conditions, such as initial culture of a host cell, and then certain steps may be performed under anaerobic conditions, such as harvesting the expressed hydrogenases to minimise or prevent inactivation by atmospheric oxygen. [0166] The skilled person will appreciate that culturing of recombinant host cells for production of recombinant proteins will be carried out at a temperature that is optimal for the growth and expression of proteins in the organism. For example, the optimum temperature for growth of E. coli and related bacterial organisms is about 37 °C and the temperature for growth of yeasts for producing recombinant proteins is about 30-32°C. [0167] “Genetically engineered” or “genetically modified” refers to any cell modified by any recombinant DNA or RNA technology. In other words, the cell has been transfected, transformed, or transduced with a recombinant polynucleotide molecule, and thereby been altered so as to cause the cell to alter expression of a desired protein. Methods and vectors for genetically engineering host cells are well known in the art; for example, various techniques are illustrated in Current Protocols in Molecular Biology, Ausubel et al., eds. (Wiley & Sons, New York, 1988, and quarterly updates). Genetic engineering techniques include but are not limited to expression vectors, targeted homologous recombination, and gene activation (see, for example, U.S. Pat. No. 5,272,071), and trans-activation by engineered transcription factors (see, for example, Segal et al., 1999, Proc Natl Acad Sci USA 96(6):2758-63). [0168] As used herein, the term "exogenous polynucleotides" is intended to mean polynucleotides that are not derived from naturally occurring polynucleotides in a given organism. Exogenous polynucleotides may be derived from polynucleotides present in a different organism. In accordance with the present invention, an E. coli cell may be genetically modified with a nucleic acid construct which contains one or more exogenous polynucleotides, encoding one or more enzymes which enable the cell to produce hydrogen. 1005373994
[0169] The exogenous polynucleotides may be heterologous or homologous. The term "heterologous" refers to a molecule or activity derived from a source other than the referenced species whereas "homologous" refers to a molecule or activity derived from the host microbial organism. Accordingly, exogenous expression of a nucleic acid molecule of the invention can be through the use of either or both a heterologous or homologous nucleic acid molecule. [0170] A particularly preferred heterologous microorganism for expression of a nucleic acid molecule or construct of the invention is E. coli. A strain of E. coli particularly preferred is DE3. DE3 indicates that the host is a lysogen of λDE3, and therefore carries a chromosomal copy of the T7 RNA polymerase gene under control of the lacUV5 promoter. A DE3 strain is suitable for production of protein from target genes cloned in pET vectors by induction with IPTG. An example of a DE3 E. coli strain is Origami™ B – with genotype F- ompT hsdSB(rB- mB-) gal dcm lacY1 ahpC (DE3) gor522:: Tn10 trxB (KanR, TetR). [0171] In a preferred embodiment, when expression of the nucleic acid molecule or construct of the invention occurs (for example when expression is induced, such as with IPTG) the temperature of the culture is lowered to 20oC to improve the yield of soluble protein. [0172] Further, in a preferred embodiment, harvesting of the expressed polypeptide or hydrogenase and/or subsequent purification steps of the expressed polypeptide or hydrogenase, are conducted in anaerobic conditions to prevent hydrogenase inactivation by atmospheric oxygen. Anaerobic conditions may be [O2] < 5 ppm. [0173] One or more other exemplary steps in the production and purification of a polypeptide of the invention are described in Example 1 below. [0174] The exogenous polynucleotides may be provided in one or more expression constructs (plasmid vectors). [0175] Methods of transforming microorganisms are well known in the art, and can include such non-limiting examples as electroporation, calcium chloride-, or lithium acetate-based methods. 1005373994
[0176] The skilled person will be familiar with methods for confirming successful transformation of relevant constructs, as well as methods for determining whether the transformants possess the relevant enzyme activity provided by the encoded protein. For example, phosphofructokinase activity (and therefore inferring correct protein folding of the encoded protein) can be inferred using a commercially available enzyme assay kit. [0177] Similarly, the skilled person will be familiar with standard techniques to confirm inhibition or deletion of the level of activity of a relevant protein or level of expression of the relevant gene. Successful gene modification, deletion or replacement can be confirmed using standard sequencing techniques. Successful inhibition of protein activity following contacting the cell with an inhibitor can be assessed by assessing for the activity of the relevant protein, for example using a commercially available enzyme assay kit. [0178] Successful transformation can also be determined by the inclusion of selection marker genes in the plasmid of vector to be transformed into the cell. As used herein, the term "selection marker genes" refer to genetic material that encodes a protein necessary for the survival and/or growth of a host cell grown in a selective culture medium. Typical selection marker genes for use in microorganisms, including in E. coli are well known to the skilled person. [0179] In any embodiment of the invention, the microorganism, preferably an E. coli microorganism, may be stored for a period of time prior to use in a method described herein (eg for oxidising hydrogen). For example, in certain embodiments, the microorganism of the invention or methods described herein may involve transformation of the microorganism with the required polynucleotides in order to generate a recombinant microorganism capable of oxidising hydrogen. The microorganism may then be harvested and stored under conditions suitable for storage of the microorganism (for example, at 4°C, -20°C, or -80°C in a suitable buffer) until required for hydrogen oxidation. [0180] It will also be appreciated that the microorganism may be lyophilised until required for further use. Further, it will be understood that the microorganism can be grown under conditions to enable expression of the required polynucleotides and then harvested, where necessary stored, and then resuspended in appropriate solutions to initiate bacterial oxidation of hydrogen. 1005373994
[0181] In certain embodiments, the bacteria may be immobilised or encapsulated. Methods for immobilisation or encapsulation of microorganisms are known, for example with the use of calcium alginate beads using standard techniques. The skilled person will be familiar with standard manual and mechanism techniques and equipment for bio- encapsulation, including by using a device such as the Inotech Encapsulator IE-50R (EncapBioSystems Inc), or Encapsulator B-390/B-395 pro (Buchi), or related systems. Other methods are described, for example in: Heidebach, et al., (2012) Critical Reviews in Food Science and Nutrition, 52: 291-311; Martín et al., (2015) Innovative Food Science & Emerging Technologies 27:15-25, the entire contents of which are hereby incorporated by reference. [0182] In other examples, the recombinant microorganism does not need to be viable (i.e., capable of reproducing, “growing” or increasing in cell numbers) in order to be able to oxidise hydrogen in accordance with the present invention. For example, in any embodiment, the methods involve providing or generating a recombinant microorganism as herein described, culturing the microorganism under conditions and for a sufficient time to induce expression of the proteins required for oxidising hydrogen (e.g., the proteins forming the hydrogenase or anaerobic enzyme complex) and then inactivating the microorganism. Preferably, the inactivated microorganisms remain intact, although it will be understood that this is not an essential requirement. [0183] Inactivated isolated, purified or recombinant microorganisms of the invention can be then be used to oxidise hydrogen, for example as described herein in the Examples. [0184] The skilled person will be familiar with methods for inactivating microorganisms so that the cells remain intact, but can still be utilised to oxidise hydrogen. Inactivation may be by gamma irradiation or by treatment with an antibiotic (such as mitomycin or similar). Systems and devices [0185] The present invention also provides systems and devices comprising the microorganisms of the invention, or reactor systems which include methods described herein for oxidising hydrogen and to preferably therefore generate energy from hydrogen. [0186] In certain embodiments, the systems and devices of the invention comprise systems and devices for use in protein film voltammetry. Also referred to as protein film 1005373994
electrochemistry or direct electrochemistry of proteins, protein film voltammetry refers to a system for detecting/measuring electron transfer reactions that occur following oxidation of a substrate by an immobilised protein (typically adsorbed or covalently attached on an electrode). Further details relating to protein film voltammetry methods are described in the Examples herein, and also in Leger et al., Biochemistry, 42, 8653- 62, 2003. [0187] The present invention provides an electrode comprising an isolated, purified or recombinant polypeptide, enzyme or enzyme complex as described herein. Preferably, the electrode comprises an electrode made of an electricity-conducting material (such as graphite), and an isolated, purified or recombinant polypeptide, enzyme or enzyme complex as described herein, attached thereto. Attachment of the enzyme to the electrode may be by any suitable method, including as described herein in the examples (eg, exposure of an abraded pyrolytic graphite edge electrode to a solution of enzyme in Tris buffer). [0188] The term “conducting material”, as used herein, refers to electricity-conducting materials that belong, for illustrative purposes and without limiting the scope of the invention, to the following group: metallic material (gold, copper, silver or platinum) or carbonaceous material (glassy carbon, “basal” pyrolytic carbon, “edge” pyrolytic carbon, carbon thread and carbon fabric, amongst other types of conducting carbon), optionally wherein on the surface thereof primary amines may be introduced by means of different techniques that allow for anchoring with the polypeptide, enzyme or enzyme complex. [0189] In any embodiment, the electrode material used is a carbonaceous material that belongs, for illustrative purposes and without limiting the scope of the invention, to the following group: glassy carbon, “basal” pyrolytic carbon, “edge” pyrolytic carbon, carbon thread and carbon fabric. [0190] Further methods for immobilising (or anchoring) hydrogenase enzymes to an electrode are described in the prior art, for example in US 20090142649, incorporated herein by reference. [0191] The invention further provides a fuel cell or an electrolytic cell comprising an electrode as described herein. 1005373994
[0192] Fuel cells are electrochemical devices that convert the energy of a fuel directly into electrochemical and thermal energy. Typically, a fuel cell consists of an anode and a cathode, which are electrically connected via an electrolyte. A fuel such as, for example, hydrogen, is fed to the anode where it is oxidized with the help of an electrocatalyst. At the cathode, the reduction of an oxidant such as oxygen (or air) takes place. The electrochemical reactions which occur at the electrodes produce a current and thereby electrical energy. Commonly, thermal energy is also produced which may be harnessed to provide additional electricity or for other purposes. Currently, the most common electrochemical reaction for use in a fuel cell is that between hydrogen and oxygen to produce water. Molecular hydrogen itself can be fed to the anode where it is oxidized, and the electrons produced are passed through an external circuit to the cathode where oxidant is reduced. Ion flow through an intermediate electrolyte maintains charge neutrality. [0193] The fuel cells of the present subject matter utilize hydrogen as a fuel wherein the source of hydrogen is provided from an exogenous source or simply atmospheric hydrogen, and the biocatalyst for conversion of the hydrogen to water is provided in the form of a purified or recombinant or isolated polypeptide, enzyme or enzyme complex of the present subject matter, or a recombinant or isolated microorganism of the invention. [0194] In any aspect or embodiment, the source of hydrogen may be a waste gas stream, including syngas and biogas. Therefore, also contemplated herein are waste gas-fed fuel cells. Further, the source of hydrogen may be ambient air. [0195] Typically, hydrogen is present in the fuel source in an amount of at least about 0.00005% by volume, or at least about 0.001% by volume, or at least about 0.01% by volume or at least about 0.1% by volume, preferably at least about 1% and more preferably at least about 5% by volume, for example about 10%, 20%, 30%, 40%, 50%, 75% or 90% by volume. Where an inert gas is used to form part of the fuel gas, the inert gas is typically present in an amount of at least about 10%, such as at least about 25%, 50 % or 75% by volume, most preferably at least about 80% by volume. [0196] Generally, the fuel source is supplied from an optionally pressurized container of the fuel source in gaseous or liquid form. The fuel source is supplied to the electrode via an inlet, which can optionally comprise a valve. An outlet is also provided which enables used or waste fuel source to leave the fuel cell. 1005373994
[0197] The oxidant typically includes oxygen, although any other suitable oxidant can be used. The oxidant source typically provides the oxidant to the cathode in the form of a gas which includes the oxidant. In some embodiments, the oxidant can be provided in liquid form. Generally, the oxidant source also includes an inert gas, although the oxidant in its pure form can also be used. For example, a mixture of oxygen with one or more gases such as nitrogen, helium, neon or argon can be used. The oxidant source can optionally comprise further components, for example alternative oxidants or other additives. An example of a suitable oxidant source is air. [0198] Typically, oxygen is present in the oxidant source in an amount of at least about 2% by volume, preferably at least about 5% and more preferably at least about 10% by volume or more. [0199] Generally, the oxidant source is supplied from an optionally pressurized container of. the oxidant source in gaseous or liquid form. The oxidant source is supplied to the electrode via an inlet, which optionally comprises a valve. An outlet is also provided which enables used or waste oxidant source to leave the fuel cell. [0200] The anode can be made of any conducting material for example stainless steel, brass or carbon, which can be graphite. The surface of the anode can, at least in part, be coated with a different material which facilitates adsorption of the catalyst. The surface onto which the catalyst is adsorbed is of a material which does not cause the hydrogenase to denature. Suitable surface materials include graphite, such as, for example, a polished graphite surface or a material having a high surface area such as carbon cloth or carbon sponge. Materials with a rough surface and/or with a high surface area are generally preferred. [0201] The cathode can be made of any suitable conducting material which will enable an oxidant to be reduced at its surface. For example materials used to form the cathode in conventional fuel cells can be used. An electrocatalyst (or bioelectrocatalyst) in the form of a polypeptide, enzyme or enzyme complex of the invention, is preferably present at the cathode. This electrocatalyst can, for example, be coated or adsorbed on the cathode itself, or it can be present in a solution surrounding the cathode. [0202] The fuel cell of the present subject matter is typically operated at a temperature of at least about 10°C, or at least about 20 °C or about 25°C, more preferably at least 1005373994
about 30°C. It is preferred that the fuel cell is operated at a temperature of from about 35 °C to about 65°C, such as from about 40°C to about 50°C. [0203] The present invention further contemplates the provision of a sensor comprising a polypeptide, enzyme or enzyme complex describe herein, wherein the sensor is useful for measuring/detecting hydrogen. In certain embodiments, the sensor device comprises a polypeptide, enzyme or enzyme complex immobilised on an electrode, such that upon oxidation of hydrogen by the polypeptide, enzyme or enzyme complex upon contact with hydrogen, and electrons are produced which enter an electrical circuit associated with the electrode. The magnitude of the electrical current generated could be correlated with the amount of hydrogen present. The system may also be calibrated using known amounts of hydrogen to enable the determination of unknown quantities of hydrogen in test samples. Kits [0204] The present invention also provides a kit comprising a polypeptide, enzyme or enzyme complex as described herein, or a device or sensor or electrode as described herein. [0205] Optionally the kit comprises written instructions for use in accordance with a method or system described herein. [0206] Further the kit may comprise buffers, co-factors and other components for enabling detection of hydrogen oxidation by the polypeptide, enzyme or enzyme complex of the invention and/or to enable detection and measurement of hydrogen in a test sample. [0207] It will be understood that the invention disclosed and defined in this specification extends to all alternative combinations of two or more of the individual features mentioned or evident from the text or drawings. All of these different combinations constitute various alternative aspects of the invention. Examples [0208] The following Examples describe the systematic analysis of [FeFe]- hydrogenases in the domain archaea by searching all publicly available species- representative genomes and MAGs, as well as multiple new MAGs. Two innovations were 1005373994
relied on to validate these findings. First, AlphaFold2-based structural modelling was used to test whether conserved gene clusters encode novel hydrogenase complexes. Critically, heterologous enzyme production was then combined with artificial maturation to confirm whether archaeal hydrogenases were catalytically active and displayed H- cluster spectroscopic signals. Through this integrated approach, it was demonstrated that archaea harbour active ultraminimal [FeFe]-hydrogenases. Novel complexes with distinct sequences, structures, and functions to previously described enzymes, including novel hybrid complexes with [NiFe]-hydrogenases were identified and characterised. Example 1 – materials and methods Genome sources [0209] In this study, 130 archaeal genomes encoding [FeFe]-hydrogenases were analysed.40 genomes are novel metagenome-assembled genomes retrieved from our unpublished datasets at ggKbase (https://ggkbase-help.berkeley.edu/). 77 genomes were retrieved from the Genome Taxonomy Database (GTDB) R06-RS202 following a search of all 2,339 archaeal species representative genomes. GTDB was chosen as the data source as it offers a standardised taxonomy, a manageable search space due to pre-clustered sequences, and genomes pre-annotated with gene predictions. In addition, 13 other Asgardarchaeota genomes were retrieved from PATRIC. All 130 genomes were examined for completeness and contamination with CheckM and manually curated through taxonomic profiling in ggKbase. Initial taxonomic classification of genome bins was performed with GTDB-Tk-R202 and subsequently confirmed by constructing phylogenetic trees. All archaeal genomes analysed were dereplicated at 95% average nucleotide identity using dRep v.3.3. Metagenomic assembly and binning [0210] The 40 newly reported genomes were retrieved from a range of anoxic ecosystems. Most MAGs came from previously reported study sites, namely Guaymas Basin hydrothermal vents (7 MAGs; Gulf of California, Mexico), Lac Lavin freshwater lake (6 MAGs; central France), Crystal Geyser (6 MAGs; Utah, USA), Chinese hot springs (5 MAGs; 2 from Yunnan Province and 3 from Tibetan Plateau, China), an aquifer (3 MAGs; Napa County, California, USA), Manure lagoon (2 MAGs, California, USA), an aquifer adjacent to the Colorado River (3 MAGs; Rifle, Colorado, USA), a wetland soil (2 MAGs; 1005373994
Napa County, California, USA), Alum Rock mineral spring (1 MAG; San Jose, California), Zodletone Spring (1 MAG; Anadarko, Oklahoma, USA), Azore Islands hot springs (1 MAGs, Portugal), a borehole (1 MAG, Muzunami, Japan). Sampling, DNA extraction, sequencing library preparation, and sequencing methods were previously described. Briefly, metagenomic sequencing reads assembled using IDBA-UD. Contigs larger than 2.5 kb were retained, and sequencing reads from all samples were mapped against each resulting assembly utilizing Bowtie2. Differential coverage profiles, filtered with a 95% read identity threshold, were then used for genome binning using a suite of binning tools (MetaBAT2, VAMB, MaxBin2, Abawaca (https://github.com/CK7/abawaca)), with the final bin choice determined by DAS Tool. Additionally, 1 MAG was obtained from samples from Corona Mine drainage at the Oat Hill Mine, Napa County, California, USA; in this case, water was filtered through 2.5 and 0.1 μm filters sequentially, DNA was extracted from each filter separately, and two runs of Illumina paired-end sequencing were performed at 150 bp and 250 bp read lengths.1 MAG were also obtained from hyporheic zone water sampled from beneath the riverbed of the East River, Gunnison County, Colorado, USA; in this case, water was filtered through a 0.1 μm filter, DNA was extracted from the filter, and Illumina sequencing was performed with a read length of 150 bp. For the East River and Oat Hill Mine samples, genome data were assembled using Metaspades v3.15.5 with default k-mer values. Draft genomes reconstructed using the MetaBAT2, VAMB, MaxBin2, and via ggKbase manual binning tools and the best genome selected using DAS Tool. Identification of archaeal [FeFe]-hydrogenase genes [0211] The process of identifying [FeFe]-hydrogenase enzymes in archaea involved matching a “training” profile of known [FeFe]-hydrogenases against a database of “candidate” protein sequences. Representative archaeal genomes from GTDB R06- RS202 were used as the data source for candidate hydrogenases. The training data was the catalytic domains of the [FeFe]-hydrogenases identified in the HydDB. The analysis pipeline was written using the Nextflow pipeline framework, which allowed for the analysis to be run reproducibly in containers, which were executed in parallel across nodes of the computing cluster. The first process in the pipeline performed a multiple sequence alignment of training sequences, which was then used to build a hidden Markov model (HMM) using hmmer3 that modelled the primary sequence of the [FeFe]-hydrogenase catalytic domain. Candidate protein sequences with length greater than 100,000 amino 1005373994
acids were first excluded from the analysis. Each protein sequence from each representative genome was then matched against the HMM to obtain a bit score, which represented the similarity of the candidate sequence to the known enzyme profile. Preliminary matches with a positive bit score were retained and then filtered by the presence of the CxxxC motif required to ligate the [FeFe]-hydrogenase catalytic centre. The matches passing this filter were then aggregated into a report, along with their taxonomy and bit score. Following this, the matches were manually inspected through a combination of Conserved Domain Database (CDD) annotations and phylogenetic analysis to derive a bit score cutoff. A final cutoff of 15.9 was chosen to include all true positives and exclude all other sequences. Analysis of [FeFe]-hydrogenase domain architecture [0212] [FeFe]-hydrogenases are often multidomain proteins comprising a conserved H2- activating domain named the H-cluster and various accessory domains at the N- and C- terminals involved in electron transfer or H2 sensing. To locate the positions of the H- cluster, all retrieved archaeal [FeFe]-hydrogenase sequences were aligned against the trimmed reference [FeFe]-hydrogenase H-cluster sequences from HydDB using Clustal Omega v1.2.2 (default setting). N- and C-terminal regions outside of the H-cluster were extracted and annotated against CDD v3.19 using rpsblast (- evalue 0.01 -max_hsps 1 - max_target_seqs 10) in BLAST+ v2.9.0. The N-terminal fusion of the small subunit of [NiFe]-hydrogenase with certain archaeal [FeFe]-hydrogenases was confirmed by searching against Pfam protein family database v34.0 and protein structural modelling as detailed below. The N- and C-terminal sequences of [FeFe]-hydrogenases were further searched for the presence of signature iron-sulfur cluster motifs or specific pattern of cysteine residues that are potentially involved in electron transfer: plant ferredoxin-like [2Fe2S] (Cx10-11Cx2Cx11-15C); bacterial ferredoxin-like 2[4Fe4S] (Cx2Cx2Cx3Cx23- 32Cx2Cx2Cx3C); (Cys)3His-ligated [4Fe4S] (Hx2-3Cx2Cx5C). Based on the domain organization and the classification scheme by Land et al., archaeal [FeFe]-hydrogenases were assigned into different subclasses. Analysis of [FeFe]-hydrogenase genetic organization [0213] To characterise the genetic context and potential interacting proteins of archaeal [FeFe]-hydrogenases, up to 10 genes upstream and downstream of the catalytic subunits were retrieved. These neighbouring genes were annotated against CDD v3.19 using 1005373994
rpsblast (-evalue 0.01 -max_hsps 1 -max_target_seqs 5) in BLAST+ v2.9.0, the Pfam protein family database v34.0 using PfamScan v1.6 (default setting), and the NCBI RefSeq protein database release 202 using DIAMOND v0.9.31 blastp algorithm (--max- hsps 1 --max-target-seqs 1). Protein subcellular localization and the presence of internal helices of the gene were predicted using PSORTb v3.0.3 (--archaea). The subgroup lineage of the flanking gene that encodes large subunit of [NiFe]-hydrogenase was classified using HydDB. To facilitate curation of the annotation results and identification of conserved neighbouring genes, all flanking genes were clustered at an identity threshold of 30% and a minimum coverage of 80% using MMseqs2 (--min-seq-id 0.3 -c 0.8 --cov-mode 1 -- cluster-mode 2 -s 7.5). The R package gggenes v0.4.1 (https://github.com/wilkox/gggenes) was used to construct gene arrangement diagrams. Identification of archaeal [FeFe]-hydrogenase maturases [0214] To probe the maturation pathway and evolution of [FeFe]-hydrogenases in Archaea, a genomic survey of the three conserved [FeFe]-hydrogenase maturases (HydE, HydF, HydG) in all representative archaeal species was performed in GTDB R06- RS202. First, a comprehensive database of known [FeFe]-hydrogenase maturase sequences from bacteria and eukaryotes was compiled, and the dataset was expanded based on BLAST searches against the NCBI non-redundant protein database (Nov 2021). New and authentic divergent hits including from archaea were included in the final database, which was then used to screen for the presence of [FeFe]-hydrogenase maturases in archaeal genomes. To enable the discovery of phylogenetically novel maturases, a relaxed default setting of the DIAMOND v0.9.31 BLASTp algorithm was first applied. False positive hits were filtered by further searching against CDD v3.19 for the presence of rSAM_HydE (TIGR03956), rSAM_HydG (TIGR03955), and GTP_HydF (TIGR03918) domains. AlphaFold2 structural modelling [0215] [FeFe]-hydrogenase models were generated using AlphaFold multimer v2.1.1 implemented on the Monash University MASSIVE M3 computing cluster. The amino acid sequences for HydA from each of the [FeFe]-hydrogenase groups (A1, A3, B, E, and F) shown in Table 7 were modelled both alone and with the sequences of putative complex partners present in the hydrogenase gene cluster. Modelling with putative complex partners was performed iteratively, with output models assessed to determine if a credible 1005373994
complex was generated. Predicted complexes were then modelled with higher stoichiometries, where possible given limitations of GPU RAM (~2,500 amino acids), to predict larger order structures. Models produced were validated based on confidence scores (pLDDT) with only regions with a confidence score of >85 utilised for analysis. Where complexes were predicted subunit interfaces were inspected manually for surface complementarity and the absence of clashing atoms. Interfaces were also analysed for stability using the program QT-PISA, with only interfaces predicted to be stable utilised for analysis. To assign cofactors to [FeFe]-hydrogenase models generated by AlphaFold the closest homologous structures or domains were identified by searching the PDB database using NCBI BLAST or the DALI server. The homologous structures were aligned with the AlphaFold models and cofactors were added in corresponding positions to that of the experimental structures, providing all conserved coordinating residues were present. Cofactor position was then manually adjusted to optimise coordination and to minimise clashes. [0216] Table 7: H2 production activities from archaeal [FeFe]-hydrogenases.
[0217] The PDB IDs for the structures were used to model cofactors for the archaeal hydrogenases are: • Group A1 – UBA95 sp002499405 (Micrarchaeota), Ca. Iainarchaeum andersonii: HydA = 6GM0 Chain A 1005373994
• Group A3 – DSAL01 sp011380095 (Altarchaeota): HydA = 6GM0 Chain A; HydB (NuoF-like) = 7E5Z chain B; HydC (NuoE-like) = 7E5Z chain A • Group B – Ca. Prometheoarchaeum syntrophicum: HydA = 6GM0 Chain A ([FeFe]- hydrogenase domain), 7BKD Chain A (ferredoxin domain); HydC (NuoE- like) = 8E9G Chain E • Group E – CABMGN01 sp902385635 (Nanoarchaeota), Ca. Forterrea multitransposorum: HydA = 6GM0 Chain A • Group F – UBA147 sp002496385 (Thermoplasmatota): HydA = 6GM0 chain A ([FeFe]-hydrogenase domain), 5ODC chain E ([NiFe]-hydrogenase small domain); HyhL = 5ODC chain L; HydD (NuoG-like) = 7T2R chain A, 5ODC chain E; HydB (NuoF-like) = 7E5Z chain B, 7MFM chain I; HydC (NuoE-like) = 7E5Z chain A Protein expression and characterization [0218] All chemicals used during the protein production and characterization were purchased from VWR and used as received unless otherwise stated. The genes encoding group A, B, E, and F [FeFe]-hydrogenases were cloned in pET-11a(+) by Genscript using restriction sites NdeI and BamHI following codon optimization for expression in E. coli. Expression constructs were re-transformed in chemically competent E. coli BL21(DE3) cells to express the apo-forms of the hydrogenases lacking the diiron subsite of the H- cluster. Starter cultures were grown overnight in 5 mL LB medium containing 100 µg mL- 1 ampicillin at 37°C. These cultures were subsequently used to inoculate 80 mL of M9 medium (22 mM Na2HPO4, 22 mM KH2PO4, 85 mM NaCl, 18 mM NH4Cl, 0.2 mM MgSO4, 0.1 mM CaCl2, 0.4% (v/v) glucose) containing 100 µg mL-1 ampicillin. Cultures were grown at 37°C and 150 rpm until an optical density (OD600) of approximately 0.4 to 0.6 was reached. Protein expression was induced by the addition of 0.1 mM FeSO4 and 1 mM IPTG. Induced cultures were incubated at 20°C and 150 rpm for approximately 16 h. Cells were thereafter harvested by centrifugation at 4,930 × g for 10 mins at 4°C. All subsequent operations were carried out under anaerobic conditions to prevent hydrogenase inactivation by atmospheric oxygen in an MBRAUN glovebox ([O2] < 5 ppm). The cell pellet was resuspended in a 0.5 mL lysis buffer (30 mM Tris-HCl pH 8.0, 0.2 % (v/v) Triton X-100, 0.6 mg mL-1 lysozyme, 0.1 mg mL-1 DNase, 0.1 mg mL-1 RNase). Cell 1005373994
lysis was performed by three cycles of freezing/thawing in liquid N2, and the supernatant was recovered by centrifugation (29,080 × g, 10 mins, 4°C). [0219] Hydrogenase semisynthetic maturation [0220] The subsite [2Fe]H subsite mimic, (Et4N)2[Fe2(SCH2NHCH2S)(CO)4(CN)2] ([2Fe]adt), was synthesized in accordance to literature protocols with minor modifications and verified by FTIR spectroscopy (Li and Rauchfuss 2002; Zaffaroni et al.2012; Cloirec et al.1999; Schmidt et al.1999; Lyon et al 1999). Incorporation of cofactor was performed by addition of 100 µg the [2Fe]adt subsite mimic (final concentration 80 μM) to 380 μL of the supernatant in potassium phosphate buffer (100 mM, pH 6.8) and 1 % (v/v) Triton X- 100. The reaction mixture was enclosed in an airtight vial and was anaerobically incubated at 20°C for 1-4 hr. The non-purified lysate containing the [2Fe]adt subsite mimic was mixed with 200 μL of potassium phosphate buffer (100 mM, pH 6.8) with 10 mM methyl viologen and 20 mM sodium dithionite. Reactions were incubated at 37°C up to 120 mins. H2 production was determined by analysing the reaction headspace every 15 mins using a PerkinElmer Clarus 500 gas chromatograph (GC) equipped with a thermal conductivity detector (TCD) and a stainless-steel column packed with Molecular Sieve (60/80 mesh). The operational temperatures of the injection port, oven, and detector were 100°C, 80°C, and 100°C, respectively. Argon was used as carrier gas at a flow rate of 35 mL min−1. The three biological replicates were run at varying times (1-4 hours) of incubating the cell lysates with the [2Fe]adt subsite mimic. Incubation time was not found to influence the observed H2 production. Thus, variation in H-cluster formation rates did not appear to have a substantial influence on the outcome of the screening process. Hydrogenase purification [0221] Expression construct with verified sequence of the group E [FeFe]-hydrogenase from Ca. Forterrea multitransposorum (Fm) was retransformed in chemically competent E. coli Origami™ B(DE3) cells since the strain yielded the highest expression levels from the small-scale expression tests. Starting cultures were grown overnight in 10 mL LB medium containing 100 µg/mL ampicillin and 15 µg/mL kanamycin at 37°C. These cultures were subsequently used to inoculate 1 L of M9 medium (22 mM Na2HPO4, 22 mM KH2PO4, 85 mM NaCl, 18 mM NH4Cl, 0.2 mM MgSO4, 0.1 mM CaCl2, 0.4% (v/v) glucose) containing 100 µg/mL ampicillin and 15 µg/mL kanamycin. Cultures were grown at 37°C and 150 rpm until an optical density (OD600) of approximately 0.4 was reached. 1005373994
Protein expression was induced by the addition of 0.1 mM FeSO4 and 1 mM IPTG. Induced cultures were incubated at 20°C and 150 rpm for approximately 16 h. Cells were thereafter harvested by centrifugation in a Beckman Coulter Avanti J-25 centrifuge (5,000 rpm/4,424 x g, 10 min). All subsequent operations were carried out under anaerobic conditions in the glovebox to prevent hydrogenase inactivation by atmospheric oxygen. The cell pellet was resuspended in 100 mM Tris-HCl pH 8.0, with NaCl (150 mM), MgCl2 (10 mM), lysozyme from chicken egg white (1 mg/mL), DNAse I from bovine pancreas (0.05 mg/mL), RNase A from bovine pancreas (0.05 mg/mL), and a tablet of cOmplete™ EDTA-free protease inhibitor cocktail, and was incubated inside the glovebox for 30 min. Cell lysis was performed by three cycles of freezing/thawing in liquid N2. Cell debris was removed by centrifugation in a Beckman Coulter Optima L-90K Ultracentrifuge (222,592 × g, 60 min). The supernatant was collected and filtered (0.45 µm syringe filter) before being loaded on a StrepTrapTM XT (Cytiva) affinity column using a BioLogic DuoFlow™ FPLC system (Bio-Rad) and purified according to the manufacturer´s instructions. Products eluted by 50 mM biotin were concentrated using Amicon®Ultra 30 kDa molecular weight cut-off (MWCO) centrifugal filters (Merck Millipore Ltd.). The StrepTrapTM eluates were further separated using size exclusion chromatography via Superdex™ 20010/300 GL, equilibrated in 100 mM Tris-HCl, 150 mM NaCl pH 8.0, to acquire purified Fm (Figure 11). Samples were stored anaerobically at −80 °C. Coomassie-stained SDS-PAGE was used to assess purity. Protein estimations were performed via Bradford assay using bovine serum albumin as a standard. Quantification of Fe-content was performed using a previously reported assay (Fish 1988) using a commercially available Fe2+ standard for AAS TraceCERT (Sigma Aldrich) for the calibration curve. Activation of purified hydrogenase [0222] To semi-enzymatically reconstitute the iron-sulfur clusters of Fm, a solution of 50 μM apoprotein in 100 mM Tris-HCl, 150 mM NaCl pH 8.0 was incubated with 500 μM dithiothreitol (DTT) under strictly anaerobic conditions for 10 min at room temperature. The iron and sulfur sources were ferrous ammonium sulfate and L-cysteine, respectively, both added in 1.5-fold molar excess to the desired number of Fe-atoms to be added. Reconstitution was initiated by adding a 1% molar equivalent of recombinant cysteine desulfurase (E. coli IscS), slowly releasing sulfide in situ from cysteine. Reaction mixtures were incubated at room temperature up to 2 hours. At the same time, the increase of 1005373994
absorbance around 405 nm was monitored by UV/Vis (Figure 11). The reconstitution process was stopped by running the reaction mixture through a PD-10 column (GE Healthcare), equilibrated in 100 mM Tris-HCl, 150 mM NaCl pH 8.0. To activate Fm with [2Fe]adt, under strictly anaerobic conditions, the reconstituted enzyme (50 µM) was mixed with sodium dithionite (1 mM, 20× excess) in 100 mM phosphate buffer, pH 6.8 and incubated in room temperature for 10 minutes. Cofactor incorporation started with the addition of [2Fe]adt (600 µM, 12× excess), and the reaction mixture was incubated for 1 h. The mixture was loaded onto a PD-10 desalting column (GE Healthcare) equilibrated with 10 mM Tris-HCl pH 8.0. The sample was concentrated using Amicon®Ultra 30 kDa MWCO centrifugal filters (Merck Millipore Ltd.), aliquoted into PCR tubes, and transferred into airtight serum vials (3-5 μL each) before they were flash-frozen in liquid N2 and stored at -80°C until further use. Protein film electrochemistry [0223] Protein film electrochemistry experiments were carried out under anaerobic conditions at 20°C and pH 7.0. The three-electrode system was made up of (1) Ag/AgCl (4 M KCl) as reference electrode, (2) rotating disk 5 mm OD pyrolytic graphite edge (PGE) plane (epoxy encapsulated) as working electrode, and (3) graphite rod as the counter electrode. The gas-tight glass cell used featured a water jacket for temperature control and a cell gas inlet/outlet for hydrogen flow control. The buffer used was composed of 5 mM MES, 5 mM CHES, 5 mM HEPES, 5 mM TAPS, 5 mM sodium acetate (NaOAc), with 0.1 M Na2SO4 as carrying electrolyte titrated with H2SO4 to pH 7.0, and purged with N2 for 3 to 4 hours. The PGE working electrodes were polished with P1200 sandpaper and rinsed with purified water before they were brought into the glovebox. To remove residual O2 in the PGE electrode, cyclic voltammograms were run at 100 mV/s from -100 to -600 mV (vs. standard hydrogen electrode (SHE)) for 40 scans. The cyclic voltammogram of the blank electrode (no enzyme immobilized) was then recorded at 10 mV/s with the working electrode rotated at 3 krpm. Polycationic polymyxin B sulfate (5 μL of 0.2 mg/mL) was added onto the deaerated PGE surface before adding 5 uL of 5 μM activated enzyme. The mixture was left for 10 min for maximal adsorption before the excess solution was removed by pipet. The cyclic voltammogram of the system with the immobilized enzyme was then recorded at 10 mV/s under 1 atm of H2. Electrochemical data was acquired using an Eco/Chemie PGSTAT10 and the GPES software (Metrohm/Autolab). Data were analyzed using Origin 8 software. All values are 1005373994
referenced versus SHE. Experiments were conducted on two independent enzyme films, with each film scanned at least three times. ATR-FTIR spectroscopy [0224] 2 µL enzyme solution (110 µM of Fm) in 10 mM Tris buffer (pH 8) was deposited on the ATR crystal. The ATR unit (BioRadII from Harrick) was sealed with a custom build PEEK cell that allowed for gas exchange and illumination mounted in a FTIR spectrometer (Vertex V70v, Bruker). The sample was dried under 100% nitrogen gas and rehydrated with a humidified aerosol (100 mM Tris-HCl, pH 8). Spectra were recorded with 2 cm-1 resolution, a scanner velocity of 80 Hz and averaged of varying number of scans (mostly 1000 Scans). All measurements were performed at ambient conditions (room temperature and pressure, hydrated enzyme films). Photochemical reduction was achieved through a previously established protocol (Lorenzi et al. 2022; Senger et al. 2019). In short, an enzyme film prepared as described above also including Eosin Y (6 mM, 0.5 µL) and triethanolamine (TEOA, 200 mM, 2 µL) was illuminated using a Schott KL2500LCD cold light source. Spectra shown in Figure 3 are representative examples of two technical replicates. Data were analysed using OPUS and Origin 2019 Software. Genome-wide metabolic annotations [0225] For all archaeal genomes encoding [FeFe]-hydrogenases, genes were predicted using Prodigal v2.6.3. Preliminary functional annotations were established and cross- referenced using KEGG HMMs, UniRef100 and UniProt, and collections of metabolic capacities in genome bins were reviewed using ggKbase genome summaries, as previously described (Castelle et al. 2015). Each open reading frame in the archaeal genomes was assigned a KEGG Orthology if the best-scoring KEGG HMM surpassed the bitscore cutoff. For a targeted profiling of the metabolic capacity of the archaeal genomes, a homology-based search was performed against a custom database (https://doi.org/10.26180/c.5230745) consisting metabolic marker genes involved in carbon fixation (RbcL, AcsB, AclB, Mcr, Hbs), alternative electron donors (FdhA, CoxL, CooS, McrA, MmoA, PmoA, IsoA, FCC, Sqr, Sor, SoxB, PsaA, PsbA, ARO), alternative electron acceptors (AsrA, DsrA, NarG, NapA, NirS, NirK, NrfA, NosZ, NorB, Nod, MtrB, OmcB, YgfK, RdhA), respiration (SdhA, FrdA, CoxA, CcoN, CyoA, CydA, AtpA) using DIAMOND v.2.0.11 with filtering cutoffs described previously (Morra 2022). The functional annotations were expanded to include marker genes for fermentation, fatty acid 1005373994
degradation, aromatic compound degradation, carbohydrate metabolism, and sulfur metabolism using curated HMMs from METABOLIC and KofamKOALA v1.3.0. Hits were further inspected through searching the NCBI CDD database. The final curated metabolic marker genes were visualized by a custom Python script to construct a heatmap that displayed the presence or absence of particular gene markers within the archaeal genomes. Counts of genes by genome were transformed into binary format, where "1" denotes the presence of at least one hit for a specific gene marker, while "0" signifies the absence of hits for the given marker. Genome-level phylogenetic analysis [0226] To construct the archaeal genome tree, 15 conserved syntenic ribosomal proteins were retrieved, aligned, and concatenated using GOOSOS (https://github.com/jwestrob/GOOSOS) using all archaeal genomes at least 60% complete and less than 5% contaminated (based on CheckM). Genomes containing at least 75% of the 15 syntenic proteins were retained. The concatenated ribosomal protein sequences were aligned using MAFFT, followed by trimming with trimAl using the -gt 0.1 option. The final length of the trimmed concatenated protein alignment was 3224 amino acids for 118 genomes. Branch support was obtained using the ultrafast bootstrap method implemented in IQ-TREE v1.6.12 and the phylogeny was estimated utilizing the following parameters -bb 1000 -m LG+F+G4. All trees were visualized using iTOL v6.3.2. Gene-level phylogenetic analysis [0227] The amino acid sequences of [FeFe]-hydrogenases catalytic subunits (HydA) were retrieved from three datasets: the archaeal genomes analysed in this study, all reference sequences from the hydrogenase database (HydDB), and additional eukaryotic genomes. For all datasets, taxonomy was assigned to each sequence using ETE v3.0.0 and CD-HIT v4.6 was used to reduce the dataset at the 80% amino acid sequence identity level. The multiple sequences retrieved were aligned using MAFFT v7.304 (settings: -- localpair --maxiterate 1000 --reorder). The resulting alignment was trimmed using trimAl (settings: -gt 0.1) and manually inspected with Geneious to remove partial sequences. Maximum likelihood phylogenetic trees were constructed using IQ-TREE v1.6.1 to test various models and topologies, obtaining bootstrap values. The LG+F+G4 substitution model with 1,000 ultrafast bootstraps was selected for tree generation. Supplementary phylogenetic trees were constructed using two different models, LG+FO+R and mixed 1005373994
model LG+C60. The three models were applied to 3,677 amino acid sequences of the catalytic subunit (HydA) of [FeFe]-hydrogenases, including novel hybrid hydrogenases. NADH-quinone oxidoreductase subunit G (NuoG) and formate dehydrogenase subunit A (FdhA1) were used as outgroup sequences for rooting given they are reported to be related and ancestral to the [FeFe]-hydrogenase catalytic subunit (HydA). The LG+FO+R model was chosen as the best fit for the data set using a standard model finder (ModelFinder) with default parameters, without testing protein mixture models. On the other hand, the protein mixture model LG+C60 was utilized to better account for the variation in amino acid replacement rates across different sites. Example 2 - Structurally and genetically diverse [FeFe]-hydrogenases are encoded by nine archaeal phyla [0228] The 2,339 archaeal species clusters of the Genome Taxonomy Database (GTDB) and the repository of novel archaeal metagenome-assembled genomes was searched for the gene encoding the catalytic subunit of [FeFe]-hydrogenases (HydA). In total, 130 archaeal genomes (90 previously reported, 40 novel) encoded [FeFe]- hydrogenases, spanning nine phyla and 17 classes (Figure 1). Except for some Asgard archaea, only one [FeFe]-hydrogenase was present in each genome. As detailed in Example 3 below, the enzymes were verified and classified based on analysis of domain structure (Figure 6), genetic organisation (Figure 8), maturases (Figure 1), and primary phylogeny (further details below). The enzymes fell into six distinct groups, namely the canonical groups A1 (n = 26), A3 (n = 44), and B (n = 12) and the novel groups E (n = 30), F (n = 21), and G (n = 3) (Figure 1). The archaeal group A1 and group E [FeFe]- hydrogenases are putative fermentative enzymes encoded by three DPANN phyla (Iainarchaeota, Micrarchaeota, Nanoarchaeota). With average sequences of just 363 and 286 residues respectively (after excluding any truncated sequences), these enzymes are much smaller than the most minimal hydrogenase previously characterised (for example Chlamydomonas reinhardtii HydA1; 457 residues). The most widespread hydrogenase, however, is the electron-bifurcating group A3 [FeFe]-hydrogenase. This hydrogenase, together with its partner diaphorase (HydB) and thioredoxin (HydC) subunits, is encoded by at least six DPANN phyla, Thermoplasmatota (class E2), and some Asgard archaea (class Lokiarchaeia) (Figure 1). [0229] The only cultured archaeon that encodes an [FeFe]-hydrogenase is ‘Candidatus Prometheoarchaeon syntrophicum’ (Figure 1). While this archaeon has previously been 1005373994
shown to fermentatively produce H2, this activity was assumed to originate from its [NiFe]- hydrogenase and its [FeFe]-hydrogenase was misannotated as an F420H2-dependent dehydrogenase subunit. The group B [FeFe]-hydrogenase gene is expressed at high levels (309 RPKM), comparable to its group 3c [NiFe]-hydrogenase genes (333 RPKM), suggesting it contributes to observed H2 production. It may even primarily account for H2 production in this culture given that [FeFe]-hydrogenases typically have higher activities than their [NiFe] counterparts and all biochemically characterised group 3c [NiFe]- hydrogenases oxidise H2 under cellular conditions. [FeFe]-hydrogenases are also encoded by MAGs of several other Lokiarchaeia and Heimdallarchaeia (Figure 1). [0230] To better understand the structure and function of the putative archaeal [FeFe]- hydrogenase, structural modelling was performed using AlphaFold2. This analysis suggested that the archaeal group A1 and E [FeFe]-hydrogenases are monomeric enzymes (HydA only) that each fold into a compact H-cluster domain with a solvent- exposed H-cluster and no additional iron-sulfur clusters (Figure 2a & 2b; Figure 8). As detailed below in Example 4, each of the modelled enzymes contained the cysteine residues required to ligate the H-cluster, though differed in their proton- transferring residues (Figure 9). Thus, despite their small size, these enzymes are theoretically capable of H2 catalysis. The group B [FeFe]-hydrogenase from Ca. P. syntrophicum modelled as a heterodimer between the HydA and HydC subunits encoded by the gene cluster; it contains two 2×[4Fe-4S] ferredoxin-like domains, one that acts as an electron relay from the H-cluster, the other of unknown function separate from the main body of the protein (Figure 2c, further details below). Structural modelling also supported that archaeal group A3 [FeFe]-hydrogenases form trimeric electron-bifurcating complexes (HydABC) similar to those recently structurally characterised in fermentative and acetogenic bacteria (Figure 2d). Example 3 - Diversity, distribution, conserved features, and classification of archaeal [FeFe]-hydrogenases [0231] A total of 136 [FeFe]-hydrogenases were identified from 130 genomes from nine archaeal phyla, including 3.3% of the 2,339 representative archaeal species genomes of the Genome Taxonomy Database (GTDB) R05-RS202. Based on sequence phylogeny using HydDB classification, 82 sequences cluster within the defined group A and B lineages.54 other sequences appear to form at least three major novel lineages, herein proposed as group E, F, and G. The linear phylogenetic tree of the [FeFe]-hydrogenase 1005373994
catalytic subunit (HydA) in archaea, bacteria, and eukaryotes is viewable at https://itol.embl.de/export/1283216342418841679681168 (a linearised version of the tree shown in Figure 5 with taxon labels). [0232] The distribution, domain features, and genetic organization defining each major hydrogenase group are summarized below. [0233] For group A [FeFe]-hydrogenases, two subgroup lineages were identified. The group A1 [FeFe]-hydrogenases are encoded by 26 genomes of two DPANN phyla, Iainarchaeota and Micrarchaeota. The hydrogenase is monomeric with the catalytic H- cluster as the signature domain and does not have additional iron-sulfur cluster binding domains at the N-terminal region (subclass M1) (Figure 6). The intact archaeal group A1 [FeFe]-hydrogenase (324 - 386 residues) is much smaller than bacterial and eukaryotic group A1 enzymes from HydDB (average 564 residues). No known [FeFe]-hydrogenases structural subunits are present in 10 genes upstream and downstream of the hydrogenase. [0234] The group A3 [FeFe]-hydrogenases have the broadest distribution among archaea, encoded by 44 genomes from eight phyla (Aenigmatarchaeota / QMZS01, Altarchaeota, Asgardarchaeota, EX4484−52, Iainarchaeota, Micrarchaeota, Nanoarchaeota, Thermoplasmatota). Three different domain organizations were identified: subclass M2 (a single 2[4Fe4S] cluster binding domain at N-terminus), subclass M3 (three signature iron-sulfur cluster binding domains at N-terminal from [2Fe2S], (Cys)3His-ligated [4Fe4S] to 2[4Fe4S]), and new subclass M3c (two signature iron-sulfur cluster binding domains at N-terminus from [2Fe2S] to 2[4Fe4S]) (Figure 6). The average gene length of archaeal group A3 [FeFe]-hydrogenases is 560 residues, comparable with bacterial group A3 counterparts in HydDB (average 585 residues). Genes encoding diaphorase (hydB) and thioredoxin (hydC) subunits necessary for the formation of the electron bifurcation complex (Figure 8; further details below) are highly conserved and typically present immediately upstream of hydA (Figure 7). Notably, hydB and hydC are fused in the novel phylum EX4484−52 (Figure 6). [0235] Interestingly, around half of the archaeal group A1 and A3 [FeFe]-hydrogenases are flanked by a conserved uncharacterised open-reading frame predicted to encode a transmembrane protein with four to six internal helices, herein denoted hyd6TM (Figure 7). Members of this uncharacterized gene have a gene length of approximately 200 1005373994
residues and all share at least 30% sequence identity. This gene cluster has no characterized homolog in the NCBI Conserved Domain Database and Pfam database. AlphaFold modelling of this gene with hydA and other structural subunits does not suggest a stable association with [FeFe]-hydrogenases. The role of this conserved gene in the functioning of archaeal group A1 and A3 hydrogenases remains to be investigated. [0236] Among archaea, the group B [FeFe]-hydrogenase is exclusively found in Asgardarchaeota, including the isolate Ca. P. syntrophicum. All variants share the same domain architecture (subclass M3a’) with a conserved putative iron-sulfur cluster binding motif containing six cysteine residues (6Cys; CxxCx10CxxCx4CxxxC) and two [4Fe4S] cluster binding domains at the N-terminal region (Figure 6). The archaeal group B [FeFe]- hydrogenases (average 522 residues) are typically larger than bacterial group B [FeFe]- hydrogenases in HydDB (average 475 residues). The gene encoding thioredoxin (hydC) is present immediately upstream of hydA, both of which are predicted to form a heterodimer (Figure 8; further details below). [0237] The group E [FeFe]-hydrogenase is a novel monophyletic clade likely basal to the bacterial group C [FeFe]-hydrogenases. It consists of 29 sequences from two DPANN phyla, Iainarchaeota and Nanoarchaeota. The putative PAS sensory domain commonly present in bacterial group C [FeFe]-hydrogenases is not found in this archaeal clade. With complete sequence lengths between 267 - 316 residues, members of this group are the most minimalistic hydrogenases known. As a comparison, bacterial group C [FeFe]- hydrogenases in HydDB have an average length of 567 residues. Like group A1 enzymes, group E [FeFe]-hydrogenases are monomeric with only the minimal catalytic H-cluster (new subclass M1a) (Figure 6) and no other [FeFe]-hydrogenases structural subunits are identified in the flanking regions. A ribonuclease (elaC) and a small nuclear ribonucleoprotein (lsm) are sometimes found next to hydA, but their modest association is likely due to close relationships of some genomes analysed. [0238] The newly discovered archaeal group F and G [FeFe]-hydrogenases together form an early diverging branch sister to the group A [FeFe]-hydrogenases. Sequences from the two groups represent two deep-branching clades separating from each other early in evolution. The group F [FeFe]-hydrogenases (average 493 residues) were identified in 21 genomes from three phyla including Asgardarchaeota (1), Thermoplasmatota (6), and Thermoproteota (17). The unique feature of this hydrogenase is the hybridization of a group 3 [NiFe]-hydrogenase small subunit (hyhS) at the N- 1005373994
terminal region and the [FeFe]-hydrogenase catalytic domain (hydA), interspaced by a 2[4Fe4S] ferredoxin-like iron-sulfur cluster binding domain (new subclass M3d) (Figure 6). This is the first observation that the genes encoding two phylogenetically distinct [NiFe]-hydrogenase and [FeFe]-hydrogenase associate. Consistently, genes encoding structural subunits of [FeFe]-hydrogenases and [NiFe]-hydrogenases were present in the immediate vicinity of the fusion gene (Figure 7). They include the diaphorase (hydB) and thioredoxin (hydC) subunits of the electron- bifurcating group A3 [FeFe]-hydrogenase, and the group 3 [NiFe]-hydrogenase large subunit (hyhL). A neighbouring gene homologous to nuoG subunit of the NADH dehydrogenase was also found and is denoted hydD. As elaborated below, AlphaFold modelling suggests the five proteins associate into a stable complex. In addition, genes encoding for maturation proteins of [NiFe]- hydrogenases (hypABCDEF, hyaD) and other [NiFe]-hydrogenase subunits were also often found in the gene cluster (Figure 7). [0239] The group G [FeFe]-hydrogenases are currently presented by three sequences from the Brockarchaeia genus JAAOZO01. This hydrogenase shares a similar domain architecture with group A3 (subclass M3) [FeFe]-hydrogenases, characterized by the presence of three signature iron-sulfur cluster binding motifs at N-terminus, for [2Fe2S], (Cys)3His-ligated [4Fe4S], and 2[4Fe4S] clusters (Figure 6). However, diaphorase (hydB) and thioredoxin (hydC) subunits of the electron-bifurcating group A3 [FeFe]-hydrogenase are absent from the gene cluster encoding the group G [FeFe]-hydrogenase. Instead, two individual genes for the small subunit (hyhS) and the large subunit (hyhL) of group 3 [NiFe]-hydrogenase were identified directly downstream of hydA (Figure 7), resembling partly the genetic organization of the group F. The genetic organisation of group G [FeFe]- hydrogenases thus is intermediate to the group A3 and group F [FeFe]-hydrogenases. It is worth noting that the N-terminal region of both group A3 and group G hydA are homologous to nuoG subunit of the NADH dehydrogenase or hydD subunit of the group E [FeFe]- hydrogenases. Notably, the short stretch of genetic region between the hyhS domain and H-cluster of the group F [FeFe]-hydrogenase also has a high degree of homology with nuoG. A possible scenario for the evolution of the group F [FeFe]- hydrogenases is homologous recombination between the nuoG-like region of group A3 and hydA of group G enzymes. As a result, nuoG (hydD) was split from hydA as a separate gene whereas hyhS was fused with hydA in group F [FeFe]-hydrogenases. The structural subunit genes of the group A3 [FeFe]-hydrogenases and group 3 [NiFe]- hydrogenases were thereby organized in the same gene cluster. An alternative scenario 1005373994
is the reverse process of the ancestral group F [FeFe]-hydrogenase diverging into group A3 and group G [FeFe]-hydrogenases lineages. Example 4 - AlphaFold2 modelling of archaeal [FeFe]-hydrogenases [0240] The structures of the group A1 [FeFe]-hydrogenases were modelled from UBA95 sp002499405 (Micrarchaeota) (Mu) and Ca. Iainarchaeum andersonii (Ia) using AlphaFold2. These enzymes have no obvious genetically associated interaction partners (Figure 7), and the structural modelling indicates they form a compact monomeric H- domain with a solvent exposed H-cluster Figure 2a; Figure 8). These group A1 enzymes do not feature additional [FeS] clusters, suggesting direct electron transfer to or from a soluble electron carrier. All five cysteines in the active site pocket previously identified as critical for H-cluster assembly and efficient catalysis in group A [FeFe]-hydrogenases are conserved in both Group A1 enzymes. This includes the four H-cluster coordinating cysteines (C2-C5) and the proton transfer residue (C1). The C1 cysteine is strictly conserved in previously identified group A and B [FeFe]- hydrogenases, and has been identified as important for catalytic activity as it plays a key role in proton transfer to the active site. Sequence alignment further revealed that other active site residues that are generally well-conserved in bacterial and eukaryotic group A [FeFe]-hydrogenases were present also in Mu and Ia (for an extended comparison see Figure 9). In line with this, the structural modelling supports the notion that they form an H-cluster with canonical coordination (Figure 2a, Figure 8). Models from the ultraminimal group E [FeFe]- hydrogenases from Ca. Forterrea multitransposorum (Fm) and CABMGN01 sp902385635 (Nanoarchaeota) (Na) were also monomeric with solvent exposed H- cluster, and no additional [FeS] clusters (Figure 2b; Figure 8). These hydrogenases lack the proton transfer cysteine C1 (Figure 9). However, this residue is not universally conserved, and noticeably absent also in previously characterized bacterial group C and D enzymes, indicating they use a different pathway for proton transfer. [0241] The Group B enzyme from Ca. P. syntrophicum (Ps) modelled as a heterodimer between the HydA and HydC subunits present in the gene cluster (Figure 2d). Modelling did not indicate that these subunits form a higher order oligomer. The HydA subunit consists of a H-domain fused to a 2x[4Fe-4S] ferredoxin-like domain that acts as an electron relay to/from the H-cluster. An additional 2x[4Fe-4S] ferredoxin-like domain is inserted in this domain (representing amino acids 122-182), forming a domain independent of the body of the protein. The predicted [FeS] clusters of this domain are 1005373994
outside of electron transfer distance (> 15 Å) with those associated with the H-cluster. The HydC subunit interacts with HydA so that its single predicted [2Fe- 2S] cluster is also not within electron transfer distance of the rest of the proteins in the structure (Figure 2d). The lack of proximity of these FeS clusters suggests that additional unidentified subunits interact with this enzyme to complete the electron transfer relay. Alternatively, this enzyme may undergo extensive conformational changes during catalysis. The model of a group A3 [FeFe]-hydrogenase from DSAL01 sp011380095 (Altarchaeota) is composed of a complex of HydA, HydB and HydC subunits, which form a putative electron- bifurcating complex similar to that recently reported from Thermotoga maritima (Figure 2d). Example 5 - Archaeal [FeFe]-hydrogenases are catalytically active and display H- cluster spectroscopic signals [0242] Group A1, B, and E [FeFe]-hydrogenases from archaea were tested to determine if they could bind the catalytic H-cluster and produce H2. To do so, five enzymes were heterologously expressed in Escherichia coli BL21(DE3), anaerobically matured using the synthetic mimic [2Fe]adt ([Fe2(azadithiolate)(CO)4(CN)2]2−), and H2 production of whole-cell lysates was measured using gas chromatography (Figure 8). H2 production was clearly discernible relative to negative controls in enzymes matured from all three groups (Figure 3a). The highest activity was observed from a Micrarchaeota group A1 enzyme (denoted Mu). Cell lysates containing Mu evolved H2 at half the rate of the well- known [FeFe]-hydrogenase HydA1 from the green alga C. reinhardtii, which was included in all assays as a positive control. The homologous enzyme from the groundwater archaeon ‘Ca. Iainarchaeum andersonii’ (denoted Ia) also produced H2, though at a hundred-fold lower rate. Substantial activity of the group B enzyme from the Asgard archaeon ‘Ca. P. syntrophicum’ (denoted Ps) and the group E enzyme from the groundwater archaeon ‘Ca. Forterrea multitransposorum’ (denoted Fm) was further observed. The observed H2 production activities should not be considered as specific activities given potential variations in the efficiency of heterologous expression, folding, and maturation between the enzymes; this is particularly illustrated by the differences in activity between the Mu and Ia hydrogenases despite their high degree of sequence homology. Nevertheless, these results validate that both DPANN and Asgard archaea encode functional [FeFe]-hydrogenases, and that the minimal group A1 and ultraminimal group E lineages are both active. 1005373994
[0243] A representative example of the group E [FeFe]-hydrogenase from ‘Ca. F. multitransposorum’ (Fm) was isolated to provide more detailed insight into the properties of these ultra-minimalistic enzymes. Following purification under strictly anaerobic conditions and semi-enzymatic reconstitution of the iron-sulfur clusters, the iron content of the enzyme was determined to 4.2 ± 0.4 per protein (Figure 9), in agreement with the structural modelling that this enzyme contains a single [4Fe4S] cluster (Figure 2). The reconstituted Fm hydrogenase was subsequently incubated with [2Fe]adt. Successful H- cluster assembly was verified through Attenuated Total Reflection Fourier transformed infrared (ATR-FTIR) spectroscopy, given sharp cofactor bands in the expected CO/CN ligand band region of the FTIR spectra were readily observed (Figure 3b). As elaborated below in Example 6, spectroscopic analysis suggested that the enzyme was isolated in an inhibited state. However, photochemical reduction resulted in the transition to catalytically active states (Figure 3b) reminiscent of the oxidised active ready states (Hox and HoxH) and a further reduced state (Hred’). Example 6 - ATR-FTIR analysis of archaeal [FeFe]-hydrogenases [0244] Attenuated Total Reflection Fourier transformed infrared (ATR-FTIR) spectroscopy confirmed successful assembly of the H-cluster of Ca. Forterrea multitransposorum (Fm) after heterologous expression, semisynthetic maturation with [2Fe]adt, and purification. Sharp cofactor bands in the expected CO/CN ligand band region of the FTIR spectra were readily observed (Figure 3b). However, the two CN (2105, 2097 cm-1) and four CO (2018, 2001, 1992, 1860 cm-1) cofactor ligand bands imply in total six ligands, indicative of a CO inhibited state; the formation of Hox-CO has been shown to occur during the artificial activation process. Regarding the apparent oxidation state of the H-cluster, the observed peak positions are reminiscent of the hydride state Hhyd or the inhibited Hinact and Htrans states, previously observed for bacterial [FeFe]-hydrogenases, suggesting a di-ferrous oxidation state of the [2Fe]H subsite. No detectable reactivity was observed when this form of Fm was exposed to H2, further supporting the notion that the enzyme is isolated in an inhibited state. [0245] Photochemical reduction resulted in the conversion of the redox state population into two new states (Figure 3b). Three new band patterns (2093, 2086, 1965, 1951, 1784 cm-1; 2081, 2068, 1965, 1953, 1776 cm-1; and 2075, 2061, 1944, 1933,1764 cm-1) are discernible, all of which most likely reflect H-cluster species with mixed valent FeIFeII [2Fe]H subsites. The three species are shifted to each other by 7-13 cm-1, indicative of 1005373994
differences in oxidation state of the [4Fe4S]H cluster or a protonation event close to the H-cluster. According to the increase of two of the band patterns at the beginning of the photoreduction experiment, these two were assigned to the catalytically active states HoxH and Hox. During extended photoreduction, the redox state population at lower wavenumbers accumulates and is consequently assigned to Hred’ (Hred’ is also referred to as Hred). The kinetics of the redox state specific bands supporting this assignment can be found in Figure 12. [0246] The catalytic properties of Fm were studied by protein film electrochemistry (PFE). Cyclic voltammetry traces of the enzyme recorded under a H2 atmosphere showed the typical bidirectional catalytic behaviour commonly associated with [FeFe]- hydrogenases (Figure 3c). A comparison of the reducing and oxidising currents observed at high driving force (± 300 mV vs reversible hydrogen electrode, RHE) indicated that the enzyme is clearly biased towards H+ reduction catalysis relative to H2 gas oxidation. This is in contrast to its closest characterized homolog, the group C [FeFe]-hydrogenase from the bacterium Thermotoga maritima (28% sequence identity), that showed a clear preference for H2 oxidation. Collectively, these findings suggest that the physiological role of the archaeal monomeric [FeFe]-hydrogenases is fermentative H2 production. Example 7 - [FeFe]- and [NiFe]-hydrogenases associate into complexes in uncultivated archaea [0247] Remarkably, the group F [FeFe]-hydrogenases appear to form complexes with [NiFe]-hydrogenases.21 genomes encoded these complexes through five-gene clusters, including from the classes Bathyarchaeia, Brockarchaeia, Thermoplasmata, Thermoproteia, and Lokiarchaeia (Figure 1). Their defining feature is the fusion of a C- terminal [FeFe]-hydrogenase catalytic domain (HydA) with an N-terminal domain homologous to the group 3 [NiFe]-hydrogenase small subunit (HyhS) (Figure 4a & Figure 8). Four other genes are contiguous with this fusion: the large subunit of the group 3 [NiFe]-hydrogenase (HyhL); the diaphorase (HydB) and thioredoxin (HydC) subunits of the electron-bifurcating group A3 [FeFe]-hydrogenase; and a conduit subunit containing four iron-sulfur clusters (herein HydD) (Figure 7). The thioredoxin, diaphorase, and conduit subunits are respectively homologous to NuoE, NuoF, and NuoG that together form the NADH dehydrogenase module of complex I. All four genes are also present in the recently identified and structurally characterised electron-bifurcating [NiFe]- hydrogenase from Acetomicrobium mobile; however, the archaeal complex is distinct in 1005373994
the presence of a true [FeFe]-hydrogenase catalytic subunit with a H-cluster domain and its fusion to the [NiFe]-hydrogenase small subunit. Of the archaeal [FeFe]-hydrogenases, the group F enzymes show the most conserved genetic organization in archaea, presumably due to their association with [NiFe]-hydrogenases (Figure 7). As elaborated above, the genetically and phylogenetically distinct group G [FeFe]- hydrogenases exclusive to the Brockarchaeia genus JAAOZO01 are also likely to form hybrid complexes (Figure 7). [0248] To confirm the catalytic activity of the group F [FeFe]-hydrogenase, two HydA- HyhS fusion proteins from the candidate lineage Thermoplasmata SG8-5 were recombinantly expressed and artificially matured. One of these enzymes rapidly produced H2 over the time course (Th1) with relative activities of 5.1% compared to the C. reinhardtii enzyme (Figure 4d; Table 7). In contrast, the other enzyme (Th2) showed only low levels of activity (Figure 4d). These measurements nevertheless confirm these are bona fide hydrogenases and that their activities are attributable to the H-cluster. [0249] Finally, AlphaFold2 modelling was performed to test whether the [FeFe]- and [NiFe]- hydrogenases associate (Figure 4b & 4c). When modelled alone, the hybrid subunit was predicted to contain a [FeFe]-hydrogenase catalytic domain and a [NiFe]- hydrogenase iron-sulfur cluster domain separated by a long flexible linker (Figure 8). Conserved cysteine ligands required to ligate both the H-cluster of the [FeFe]- hydrogenase and the three electron-relaying [4Fe4S] clusters of the [NiFe]- hydrogenase component were both observed (Figure 9; Figure 4c; further details below in Example 8). However, modelling of all five genetically contiguous subunits (HydA-HyhS, HydBCD, HyhL) indicated they form a stable electron-bifurcating complex. As elaborated below, the structural model suggests the complex receives electrons through the [FeFe]- hydrogenase arm or the [NiFe]-hydrogenase via a series of iron-sulfur clusters to a probable electron-converging [4Fe4S] cluster on the hybrid subunit. Thereafter electrons are predicted to be simultaneously transferred to the high-potential NAD+ at the HydC subunit and an undetermined low-potential acceptor (likely ferredoxin) at the glutamate synthase (GltA) domain of the HydB subunit (Figure 4b). These observations are remarkable given [NiFe]- and [FeFe]- hydrogenases are not known to associate. Moreover, they suggest surprising modularity of the [NiFe]-hydrogenase small subunit given it has evidently co-evolved with both [NiFe]- and [FeFe]- hydrogenases. Example 8 - AlphaFold modelling of hybrid complexes 1005373994
[0250] AlphaFold2 modelling was used to test whether the [FeFe]- and [NiFe]- hydrogenases associate, focusing on the more active hydrogenase from UBA147 sp002496385 (Thermoplasmatota). When modelled alone using AlphaFold, the hybrid HydA subunit was predicted to contain a [FeFe]-hydrogenase catalytic domain and a [NiFe]- hydrogenase iron-sulfur cluster domain separated by a long flexible linker (Figure 8). The cysteine residues required to ligate the [FeFe]-hydrogenase H-cluster of the HydA subunit (Cys357, Cys406, Cys536, Cys540) were present, and are overall well-conserved for all group F representatives, indicating that the HydA domain can bind an H-cluster. However, similarly to group C to E [FeFe]-hydrogenases, the proton transfer C1 cysteine is absent. More surprisingly, the active-site pocket also featured changes in other amino acids which are well-conserved in all other groups of [FeFe]-hydrogenases. For example, the so-called APA (or APS) motif is replaced by a DPI motif in most identified group F enzymes (Figure 9). Similar to the HydA subunit, the cysteine residues required to ligate the catalytic NiFe cofactor are present in HyhL (Cys63, Cys66, Cys418, Cys 421). Altogether, this indicates that both the HydA and HyhL subunits are likely to be active hydrogenases (Figure 4c). [0251] Modelling of all five genetically contiguous subunits (HydA-HyhS, HydBCD, HyhL) indicated they form a stable electron-bifurcating complex (Figure 8; Figure 4b). The closest characterized relative is the recently reported [NiFe]-hydrogenase from A. mobile; this enzyme contains homologs of HyhL, HyhS (as a separate subunit), HydB, HydC, and HydD (annotated as HydA despite lacking H-cluster), but lacks both the canonical catalytic HydA domain and the HyhS fusion characteristic of the archaeal enzyme. [0252] If the archaeal complex operates in a H2 oxidative direction, electrons are likely to be input into the complex through the [FeFe]-hydrogenase arm (through the H-cluster then FeS clusters A1 and A2) or the [NiFe]-hydrogenase arm (through the NiFe-centre then FeS clusters A3, D4, D3, D2, and A2). The relative proximity of the [FeS]-clusters indicates that A2 is the electron-converging site for the [FeFe]- and [NiFe]- hydrogenase arms. From A2, electrons likely transfer to FMN in the HydB subunit via the D1 and B1 [FeS]-clusters (Figure 4c). In the A. mobile enzyme, it has been proposed that this FMN is the site of electron-bifurcation with high-potential electrons transferred to bound NAD and low potential electrons transferred to [FeS] cluster C1, which is conserved in the archaeal group F [FeFe]-hydrogenase. In the A. mobile enzyme, it has been hypothesized that electrons are then transferred from C1 to soluble ferredoxin via iron-sulfur clusters in 1005373994
a flexible 2x[4Fe-4S] ferredoxin-like domain of the HydB subunit. In HydB from the archaeal group F hydrogenase, this domain is substituted by a homologue of the GltA subunit from glutamate synthase, which contains 2x[4Fe4S] clusters (B3 and B4) and FAD, which likely serves as an electron donor to an external substrate. Electrons are likely transferred from C1 to B1, and then to the terminal FAD via the B3 and B4 clusters (Figure 4c). In the A. mobile enzyme, it is predicted that electron bifurcation is mediated by flexibility in the complex that changes the proximity of the [FeS]-clusters gating the flow of electrons through the complex. The GltA-like domain of the archaeal group F hydrogenase is attached to the remainder of HydB via a flexible linker, suggesting that electron bifurcation may be achieved by a similar mechanism in this complex. The species that accepts electrons from FAD in the GltA-like domain is unclear, although it may be soluble ferredoxin or another low potential electron acceptor. [0253] If the archaeal group F hydrogenase functions in a hydrogen-evolving direction, then this process would occur in reverse with electrons entering the complex via the FMN and FAD cofactors and converging at cluster A2, before transfer to the [FeFe] and/or [NiFe] sites of HydA and HyhL respectively. Example 9 - [FeFe]-hydrogenases enable fermentation and electron-bifurcation in diverse archaea [0254] The role of the various [FeFe]-hydrogenases in the metabolism of archaea was studied. The high-quality archaeal genomes for genes associated with major energy conservation and carbon acquisition processes was annotated. All [FeFe]-hydrogenase- encoding archaea are predicted to be obligate anaerobes given they lacked terminal oxidases. This is consistent with the retrieval of the genomes from typically anoxic ecosystems, especially groundwater, anaerobic digesters, and sediments from hot springs, hydrothermal vents, and freshwater (Figure 1). Based on the retrieved genomes, DPANN archaea are likely to be symbiotic obligate fermenters dependent on host-derived organic compounds; consistently, they often encoded genes for the degradation and fermentation of carbohydrates (primarily starch) and aromatic compounds, but generally lacked respiratory reductases or carbon fixation pathways (Figure 1). The other archaeal phyla were predicted to be capable of a wider range of metabolic strategies, in line with their larger genome sizes (Figure 6), including beta oxidation, anaerobic respiration, and carbon fixation to varying extents (Figure 1). Notably, some Asgard archaea encoded fumarate reductases (Frd), reductive dehalogenases (Rdh), and anaerobic sulfite 1005373994
reductases (Asr), with the latter enzyme also encoded in certain MAGs from four other phyla. Several lineages were also predicted to be capable of autotrophy through the Wood-Ljungdahl (Lokiarchaeia, Thermoproteota) or reverse tricarboxylic acid (Lokiarchaeia, Heimdallarchaeia, Thermoplasmatota) pathways (Figure 1; Table 7). Most of the MAGs also encoded RuBisCO lineages known to function in nucleoside salvage. [0255] The group A, B, and E [FeFe]-hydrogenases likely facilitate cofactor regeneration during organic carbon fermentation in diverse archaea. Half of the archaea encode 2- oxoacid-ferredoxin oxidoreductases, such as pyruvate-ferredoxin oxidoreductases that couple the oxidation of the endproduct of glycolysis (pyruvate) to the reduction of ferredoxin, and most can gain ATP by converting the derived acetyl-CoA to the endproduct acetate via the acetyl-CoA synthetase or acetate kinase reaction (Figure 1). As supported by the structural modelling (Figure 2), the monomeric group A1, B, and E hydrogenases are predicted to couple ferredoxin reoxidation to H2 production. Also consistent with this role, 2-oxoacid-ferredoxin oxidoreductase genes are frequently adjacent to [FeFe]-hydrogenase genes in Nanoarchaeota and Micrarchaeota MAGs (Figure 7). By contrast, the trimeric group A3 [FeFe]-hydrogenases are predicted to simultaneously reoxidize ferredoxin (reduced primarily by pyruvate-ferredoxin oxidoreductase) and NADH (e.g. reduced during glycolysis) (Figure 2), in line with their bacterial counterparts. Congruently, the electron-bifurcating group A3 [FeFe]- hydrogenases are associated with those DPANN archaea harbouring relatively complex carbohydrate degradation pathways (i.e. Woesearchaeles, Altarchaeota, some Iainarchaeota) through which both NAD+ and ferredoxin will be reduced (Figure 1). Also notable is that most of the DPANN genomes encode the minimal group A1 and ultraminimal group E [FeFe]-hydrogenases as their sole H2-metabolising enzymes; together with the simple maturation pathway of [FeFe]-hydrogenases compared to [NiFe]- hydrogenases, it is probable that [FeFe]-hydrogenases contribute to the minimisation of the genetic and cellular requirements to metabolise H2 in these genome-reduced microorganisms (average completeness-normalized genome size: 1.2 Mbp). Further studies are needed to determine what biochemical features differentiate the monomeric group A1, B, and E [FeFe]-hydrogenases and whether they have distinct physiological roles. It cannot be ruled out that the group B enzymes instead consume H2 to support anaerobic respiration or carbon fixation in Asgard archaea; however, this seems unlikely given their reported H2-evolving activities (Figure 3), structural features (Figure 2), and the hydrogenogenic lifestyle of ‘Ca. P. syntrophicum’. 1005373994
[0256] The physiological role of the hybrid hydrogenases is unclear. These enzymes are exclusively encoded by more complex archaea, including those with the capacity for carbon fixation, beta oxidation, and energy conversion using group 4 [NiFe]- hydrogenases (Figure 1). The predicted structures suggest that that these enzymes contribute to electron bifurcation by transferring electrons from H2 to both NAD and likely ferredoxin, or vice versa (Figure 3). However, it is peculiar that a single complex contains two seemingly redundant hydrogenase modules. A potential explanation is that one the hydrogenase modules might act on a substrate other than H2, especially given group 3 [NiFe]-hydrogenases can behave as sulfhydrogenases in vitro (i.e. mediating reduction of elemental sulfur to hydrogen sulfide). Another possibility is that these complex act as a redox valve, transferring electrons either from NADH and reduced ferredoxin to a H2- evolving [FeFe]-hydrogenase when reductant accumulates, and from a H2-consuming [NiFe]-hydrogenase to NAD and oxidised ferredoxin otherwise. Other enzyme complexes regulating opposite reactions have recently been reported, namely between glutamate synthase (GltAB) and glutamate dehydrogenase (GudB), with the GltA domain of these enzymes also shared in the hybrid hydrogenase complex. A further rationale is that the two hydrogenase modules may differ in their affinities and/or oxygen tolerance, enabling efficient H2 oxidation across a wide range of environmental conditions. Example 10 - [FeFe]-hydrogenases have been acquired by archaea on multiple occasions and have an ancient association with [NiFe]-hydrogenases [0257] The evolutionary history of [FeFe]-hydrogenases was investigated through phylogenetic analysis of its catalytic subunit (Figure 5; Figures 13 to 14) and three maturases (Figure 15 to 17). In agreement with their classification into known groups, the archaeal group A1, A3, and B [FeFe]-hydrogenases clustered with various bacterial and eukaryotic homologs; the group A enzymes form at least six radiations, suggesting they were laterally acquired from bacterial enzymes over several independent events, whereas the archaeal group B and E enzymes each formed a monophyletic clade (Figure 5; Figures 13 to 14). Based on phylogenies constructed using the most robustly supported model (Figure 5) and their structural simplicity (Figure 2b), it’s possible that the ultraminimal fermentative group E enzymes of archaea are ancestral to the multidomain sensory group C hydrogenases of bacteria. Syntrophy hypotheses for eukaryogenesis were also reexamined based on the finding that [FeFe]-hydrogenases are present both in Asgard archaea and unicellular eukaryotes. The data do not support the hypothesis 1005373994
that all eukaryotic [FeFe]-hydrogenases were vertically acquired from an Asgard archaeal ancestor. The eukaryotic enzymes cluster into four disparate clades, spanning both group A and B, suggesting multiple horizontal acquisitions (Figure 5). In addition, [FeFe]- hydrogenases are absent from the Asgard genomes most related to eukaryotes (Figure 1) and the alphaproteobacterial genomes most related to mitochondria, though this view could change with the addition of new genomes. Nevertheless, archaeal and eukaryotic [FeFe]-hydrogenases clustered together in three of these lineages (including Lokiarchaeia with Tritrichomonas), suggesting eukaryotes may have laterally acquired [FeFe]- hydrogenases from archaea during their diversification (Figure 5). [0258] The distribution and phylogeny of the three maturases that synthesise the [FeFe]- hydrogenase cofactor, HydE, HydF, and HydG was also studied. Of the 130 archaeal genomes encoding [FeFe]-hydrogenases, just five encode a full set of maturases and three others encode an incomplete set. Though genome incompleteness means that the co-occurrence of hydrogenases and maturases will be underestimated, this does not explain the absence of maturases in most genomes. Indeed, maturase genes are even absent in the genome of the cultured hydrogenogenic archaeon ‘Ca. P. syntrophicum’ and the closed genome of ‘Ca. F. multitransposorum’ from which an active [FeFe]- hydrogenase was purified (Figure 3). In turn, it is likely that archaea synthesise [FeFe]- hydrogenases through an alternative pathway. The existence of an alternative pathway is consistent with reports that various eukaryotes lacking all (Giardia, Entamoeba) or some (Mastigamoeba) maturases still make catalytically active H2-producing hydrogenases, as well as recent reports of a cytosolic [FeFe]-hydrogenase in Trichomonas vaginalis and the heterologous production of active [FeFe]-hydrogenases in Synechocystis cells lacking maturases. In line with findings for the structural subunits, phylogenetic analysis of HydE, HydF, and HydG suggests that archaea acquired maturases on several occasions (Figures 15 to 17). Several archaeal maturases are closely related to those that were recently identified in Chlamydiae and eukaryotes. [0259] Phylogenetic analyses also suggest an ancient and potentially basal origin of the hybrid hydrogenases. In phylogenetic trees of the [FeFe]-hydrogenase catalytic subunit, sequences of group F and G hydrogenases formed long branches, as expected given their divergence from the group A to D hydrogenases (~25% sequence identity to the model hydrogenase C. reinhardtii HydA1). Their placement was highly unstable in unrooted trees: sequences clustered either within or basal to the group A [FeFe]- 1005373994
hydrogenases as either monophyletic or biphyletic lineages. In trees rooted with outgroups (FdhA and NuoG; thought to be ancestral to HydA), however, the hybrid enzymes clustered into well-supported monophyletic groups basal to the bacterial groups. Concordant findings were observed using three different phylogenetic models (LG+F+G4, LG+FO+R, LG+C60) (Figures 5, 13 to 14). In phylogenetic trees of the [NiFe]- hydrogenases, the catalytic (large) subunits of the fusion proteins formed multiple clusters (Figure 18). In contrast, the iron-sulfur (small) domain fused to the group F [FeFe]- hydrogenase and the iron-sulfur subunit downstream of the group G [FeFe]-hydrogenase each form distinct monophyletic subgroups within the group 3 [NiFe]-hydrogenases (Figure 19). Thus, as dictated by the gene fusion, the iron-sulfur domain of hybrid complexes more strongly co-evolved with the [FeFe]- than [NiFe]-hydrogenase catalytic subunits. Among several potential evolutionary scenarios, [FeFe]-hydrogenases may have first evolved in archaea in association with [NiFe]-hydrogenases; they potentially diversified into various monomeric and multimeric lineages, which were variably acquired by diverse bacteria, archaea, and eventually eukaryotes. It is proposed that the components of the hybrid hydrogenases are formally recognised as distinct lineages, namely the group F and G [FeFe]-hydrogenases and group 3f and 3g [NiFe]- hydrogenases, given their distinct phylogenies, structures, and potential physiological roles. [0260] Archaea have evolved remarkably disparate ways to use [FeFe]-hydrogenases to adapt to anaerobic environments. On one hand, DPANN archaea have evolved ultraminimal enzymes to efficiently dispose of reductant derived from carbohydrate fermentation. Conversely, lineages such as Brockarchaeia and Lokiarchaeia use unique hybrids of [NiFe]- and [FeFe]-hydrogenases – among the most complex hydrogenases described to date – to support their diverse redox biology. These discoveries expand the [FeFe]-hydrogenases from four to seven groups each with distinct phylogenies, structures, and functions. In addition to increasing understanding of archaeal biology, these findings also redefine existing understanding of hydrogen metabolism and hydrogenase biochemistry by showing: (i) [FeFe]-hydrogenases are active across all three domains of life, (ii) the two dominant hydrogenase classes have co-evolved, and (iii) providing a new size minimum for hydrogenases. The ultraminimal hydrogenases also provide ideal templates both to understand enzymatic H2 catalysis, for example determinants of rate, directionality, affinity, and oxygen sensitivity, as well as flexible scaffolds for directed evolution of efficient H2-converting biocatalysts. More broadly, this 1005373994
approach also emphasises the potential of combining genome-resolved metagenomics with accurate protein structure prediction and heterologous production studies to discover new enzymes and functions in uncultured microorganisms. [0261] References [0262] Castelle, C.J., Brown, C.T., Anantharaman, K., Probst, A.J., Huang, R.H., and Banfield, J.F. (2018). Biosynthetic capacity, metabolic variety and unusual biology in the CPR and DPANN radiations. Nat. Rev. Microbiol.16, 629–645. [0263] Cloirec, A., Davies, S., Evans, D., Hughes, D., Pickett, C., and Best, S. (1999). A di-iron dithiolate possessing structural elements of the carbonyl/cyanide sub- site of the H-centre of Fe-only hydrogenase. Chem. Commun., 2285–2286. [0264] Fish, W.W. (1988). Rapid colorimetric micromethod for the quantitation of complexed iron in biological samples in Methods in Enzymology (Riordan, JF and Vallee, BL, eds.) Vol.158A. [0265] Li, H., and Rauchfuss, T.B. (2002). Iron carbonyl sulfides, formaldehyde, and amines condense to give the proposed azadithiolate cofactor of the Fe-only hydrogenases. J. Am. Chem. Soc.124, 726–727. [0266] Lorenzi, M., Gamache, M.T., Redman, H.J., Land, H., Senger, M., and Berggren, G. (2022). Light-driven [FeFe] hydrogenase based H2 production in E. coli: A model reaction for exploring E. coli based semiartificial photosynthetic systems. ACS Sustain. Chem. Eng.10, 10760–10767. [0267] Lyon, E.J., Georgakaki, I.P., Reibenspies, J.H., and Darensbourg, M.Y. (1999). Carbon monoxide and cyanide ligands in a classical organometallic complex model for Fe‐only hydrogenase. Angew. Chemie Int. Ed.38, 3178–3180. [0268] Morra, S. (2022). Fantastic [FeFe]-hydrogenases and where to find them. Front. Microbiol.13. [0269] Schmidt, M., Contakes, S.M., and Rauchfuss, T.B. (1999). First generation analogues of the binuclear site in the Fe-only hydrogenases: Fe2(µ- SR)2(CO)4(CN)22-. J. Am. Chem. Soc.121, 9736–9737. 1005373994
[0270] Senger, M., Eichmann, V., Laun, K., Duan, J., Wittkamp, F., Knör, G., Apfel, U.- P., Happe, T., Winkler, M., and Heberle, J. (2019). How [FeFe]-hydrogenase facilitates bidirectional proton transfer. J. Am. Chem. Soc.141, 17394–17403. [0271] Zaffaroni, R., Rauchfuss, T.B., Gray, D.L., De Gioia, L., and Zampella, G. (2012). Terminal vs bridging hydrides of diiron dithiolates: protonation of Fe2(dithiolate)(CO)2(PMe3)4. J. Am. Chem. Soc.134, 19260–19269. 1005373994
Claims
CLAIMS 1. An isolated, synthetic or purified nucleic acid molecule encoding a hydrogenase from archaea of a lineage Thermoplasmatota, Asgardarchaeota, Thermoproteota, EX4484-52, Aenigmarchaeota / QMZS01, Nanoarchaeota, Altarchaeota, lainarchaeota, or Micrarchaeota.
2. The nucleic acid molecule of claim 1, wherein the archaea is one shown in Figure 1.
3. The nucleic acid molecule of claim 1, wherein the nucleic acid comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 1 to 150, or a sequence at least about 80% identical thereto.
4. The nucleic acid molecule of claim 1, wherein the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 300, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 151 to 300.
5. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 1 to 26, or a sequence at least about 80% identical thereto.
6. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 176, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 151 to 176.
7. The nucleic acid molecule of any one of claims 1 to 4, wherein the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 27 to 70, or a sequence at least about 80% identical thereto.
8. The nucleic acid molecule of any one of claims 1 to 4, wherein the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 177 to 220, or a functional 1005373994
equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 177 to 220.
9. The nucleic acid molecule of any one of claims 1 to 4, wherein the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 71 to 82, or a sequence at least about 80% identical thereto.
10. The nucleic acid molecule of any one of claims 1 to 4, wherein the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 221 to 232, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 221 to 232.
11. The nucleic acid molecule of any one of claims 1 to 4, wherein the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in SEQ ID NOs: 83, or a sequence at least about 80% identical thereto.
12. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in SEQ ID NOs: 233, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of SEQ ID NO: 233.
13. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 84 to 112, or a sequence at least about 80% identical thereto.
14. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 234 to 262, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 234 to 262.
15. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 113 to 133, or a sequence at least about 80% identical thereto. 1005373994
16. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 263 to 283, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 263 to 283.
17. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 134 to 136, or a sequence at least about 80% identical thereto.
18. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 284 to 286, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 284 to 286.
19. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule comprises, consists essentially of, or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 137 to 143, or a sequence at least about 80% identical thereto.
20. The nucleic acid molecule of any one of claims 1 to 4, the nucleic acid molecule encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 287 to 293, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 287 to 293.
21. A nucleic acid molecule comprising, consisting essentially of or consisting of a nucleotide sequence set forth in any one of SEQ ID Nos: 144 to 150.
22. A nucleic acid molecule comprising, consisting essentially of or consisting of a nucleotide sequence, wherein the nucleotide sequence is codon optimised for expression in a heterologous host, such as E. coli, and encodes a polypeptide comprising, consisting essentially of or consisting of an amino acid sequence as set forth in any one of SEQ ID NOs: 151 to 286 or 294 to 300, or a functional equivalent, homolog or derivative thereof comprising at least about 80% sequence identity to the sequence of any one of SEQ ID NO: 151 to 286 or 294 to 300. 1005373994
23. The nucleic acid molecule of claim 22, wherein nucleic acid molecule does not encode the C-terminal SSGWSHPQFEK as depicted in SEQ ID Nos: 294 to 300.
24. A nucleic acid construct comprising a nucleic acid molecule of any one of claims 1 to 23.
25. The nucleic acid construct of claim 24, the construct is synthetic, recombinant or isolated.
26. The nucleic acid construct of claim 24 or 25, wherein the construct comprises one or more heterologous promoters for enabling expression of the nucleic acid(s) comprised in the construct.
27. The nucleic acid construct of any one of claims 24 to 26, wherein the nucleic acid molecule or construct comprises a codon optimised sequence for enabling expression of the nucleic acid in a heterologous host. 29. A cell comprising a nucleic acid molecule of any one of claims 1 to 23 or nucleic acid construct of any one of claims 24 to 27. 30. The cell of claim 29, wherein the cell is a microorganism. 31. The cell of claim 30, wherein the cell is E. coli. 32. The cell of claim 31, wherein the E. coli is strain DE3. 33. The cell of any one of claims 29 to 32, wherein the cell is a recombinant cell. 34. An isolated or recombinant microorganism comprising a nucleic acid molecule of any one of claims 1 to 23 or nucleic acid construct of any one of claims 24 to 27. 35. An isolated, recombinant or purified polypeptide capable of oxidising hydrogen, wherein the polypeptide comprises, consists essentially of or consists of an amino acid sequence selected from any one of SEQ ID NOs: 151 to 293, or a functional equivalent, homology or derivative thereof having a sequence at least 80% identical thereto. 36. The isolated, recombinant or purified polypeptide of claim 35, wherein the polypeptide comprises, consists essentially of or consists of an amino acid sequence 1005373994
selected from any one of SEQ ID NOs: 287 to 293, or a functional equivalent, homology or derivative thereof having a sequence at least 80% identical thereto. 37. The isolated, recombinant or purified polypeptide of claim 36, wherein the polypeptide comprises, consists essentially of or consists of an amino acid sequence selected from any one of SEQ ID NOs: 294 to 300, or a functional equivalent, homology or derivative thereof having a sequence at least 80% identical thereto. 38. The isolated, recombinant or purified polypeptide of claim 35 to 37, wherein the polypeptide is capable of H+ reduction catalysis and/or H2 gas oxidation. 39. An anaerobically matured enzyme, or enzyme complex, for oxidising hydrogen, wherein the enzyme or enzyme complex comprises proteins comprising the amino acid sequences as set forth in any one or more of SEQ ID NOs: 151 to 300, or functional equivalent, homologs or derivatives having at least 80% sequence identity thereto. 40. The anaerobically matured enzyme, or enzyme complex of claim 39, wherein the enzyme or enzyme complex has been matured using a synthetic mimic. 41. The anaerobically matured enzyme, or enzyme complex of claim 40, wherein the synthetic mimic is [2Fe]adt ([Fe2(azadithiolate)(CO)4(CN)2]2−). 42. A method for converting hydrogen to electrons, the method comprising contacting a source of hydrogen with a cell of any one of claims 29 to 33, an isolated or recombinant microorganism of claim 34, or an isolated, recombinant or purified polypeptide, enzyme, or enzyme complex of claims 35 to 41. 43. A method for producing hydrogen (H2), the method comprising contacting a source of protons with a cell of any one of claims 29 to 33, an isolated or recombinant microorganism of claim 34, or an isolated, recombinant or purified polypeptide, enzyme, or enzyme complex of claims 35 to 41. 44. Use of a cell of any one of claims 29 to 33, an isolated or recombinant microorganism of claim 34, or an isolated, recombinant or purified polypeptide, enzyme, or enzyme complex of claims 35 to 41, for producing hydrogen (H2). 45. A method of generating energy from a source of hydrogen, the method comprising contacting a source of hydrogen with a cell of any one of claims 29 to 33, an isolated or 1005373994
recombinant microorganism of claim 34, or an isolated, recombinant or purified polypeptide, enzyme, or enzyme complex of claims 35 to 41. 46. Use of a cell of any one of claims 29 to 33, an isolated or recombinant microorganism of claim 34, or an isolated, recombinant or purified polypeptide, enzyme, or enzyme complex of claims 35 to 41, for generating energy from a source of hydrogen. 47. A device for converting hydrogen to electrons, the device comprising an immobilised polypeptide, enzyme, or enzyme complex of claims 35 to 41. 48. The device of claim 47, wherein the enzyme complex is provided in the device/ electrode immobilised or covalently bound to the surface of an electrically-conducting material. 49. A system for oxidising hydrogen, comprising an isolated, purified or recombinant polypeptide, enzyme, or enzyme complex of claims 35 to 41, a source of hydrogen, and means for detecting the oxidation of hydrogen. 50. The system of claim 49, wherein the system further comprises one or more co- factors for enabling the oxidation of hydrogen. 51. The system of claim 49 or 50, wherein the enzyme complex is immobilised, optionally, wherein the enzyme complex is covalently bound to the surface of an electrode. 52. A fuel cell comprising a device or system of any one of claims 47 to 51. 53. The fuel cell of claim 52, wherein the fuel cell is waste gas-fed from a source of hydrogen such as a waste gas stream, including syngas or biogas. 54. An air-powered device comprising a fuel cell of claim 52 or 53, preferably the fuel cell generates energy from hydrogen present in ambient air thereby providing energy to the device. 1005373994
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| AU2023902336A AU2023902336A0 (en) | 2023-07-21 | Novel hydrogenases | |
| AU2023902336 | 2023-07-21 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025019889A1 true WO2025019889A1 (en) | 2025-01-30 |
Family
ID=94373806
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/AU2024/050777 Pending WO2025019889A1 (en) | 2023-07-21 | 2024-07-19 | Novel hydrogenases |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025019889A1 (en) |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100311142A1 (en) * | 2008-09-05 | 2010-12-09 | Korea Ocean Research & Development Institute | Novel Hydrogenases Isolated from Thermococcus SPP., Genes Encoding the Same, and Methods for Producing Hydrogen Using Microorganisms Having the Genes |
-
2024
- 2024-07-19 WO PCT/AU2024/050777 patent/WO2025019889A1/en active Pending
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100311142A1 (en) * | 2008-09-05 | 2010-12-09 | Korea Ocean Research & Development Institute | Novel Hydrogenases Isolated from Thermococcus SPP., Genes Encoding the Same, and Methods for Producing Hydrogen Using Microorganisms Having the Genes |
Non-Patent Citations (15)
| Title |
|---|
| ARTERO VINCENT, BERGGREN GUSTAV, ATTA MOHAMED, CASERTA GIORGIO, ROY SOUVIK, PECQUEUR LUDOVIC, FONTECAVE MARC: "From Enzyme Maturation to Synthetic Chemistry: The Case of Hydrogenases", ACCOUNTS OF CHEMICAL RESEARCH, vol. 48, no. 8, 18 August 2015 (2015-08-18), US , pages 2380 - 2387, XP093271224, ISSN: 0001-4842, DOI: 10.1021/acs.accounts.5b00157 * |
| BENOIT STÉPHANE L, MAIER ROBERT J, SAWERS R GARY, GREENING CHRIS: "Molecular Hydrogen Metabolism: a Widespread Trait of Pathogenic Bacteria and Protists", MICROBIOLOGY AND MOLECULAR BIOLOGY REVIEWS, vol. 84, no. 1, 19 February 2020 (2020-02-19), US , pages 1 - 52, XP093271193, ISSN: 1092-2172, DOI: 10.1128/MMBR * |
| DATABASE Uniprot and GenBank 2016, "see Box V, Inventive Step", Database accession no. A1-A85 * |
| ESMIEU C., RALEIRAS P., BERGGREN G.: "From protein engineering to artificial enzymes – biological and biomimetic approaches towards sustainable hydrogen production", SUSTAINABLE ENERGY & FUELS, vol. 2, no. 4, 1 January 2018 (2018-01-01), UK , pages 724 - 750, XP093271199, ISSN: 2398-4902, DOI: 10.1039/C7SE00582B * |
| FIRRINCIELI ANDREA, NEGRONI ANDREA, ZANAROLI GIULIO, CAPPELLETTI MARTINA: "Unraveling the Metabolic Potential of Asgardarchaeota in a Sediment from the Mediterranean Hydrocarbon-Contaminated Water Basin Mar Piccolo (Taranto, Italy)", MICROORGANISMS, vol. 9, no. 4, CH, pages 859 - 859-15, XP093271175, ISSN: 2076-2607, DOI: 10.3390/microorganisms9040859 * |
| GREENING CHRIS, BISWAS AMBARISH, CARERE CARLO R, JACKSON COLIN J, TAYLOR MATTHEW C, STOTT MATTHEW B, COOK GREGORY M, MORALES SERGI: "Genomic and metagenomic surveys of hydrogenase distribution indicate H2 is a widely utilised energy source for microbial growth and survival", THE ISME JOURNAL, vol. 10, no. 3, 1 March 2016 (2016-03-01), UK, pages 761 - 777, XP093271222, ISSN: 1751-7362, DOI: 10.1038/ismej.2015.153 * |
| HUANG WEN-CONG, LIU YANG, ZHANG XINXU, ZHANG CUI-JING, ZOU DAYU, ZHENG SHILING, XU WEI, LUO ZHUHUA, LIU FANGHUA, LI MENG: "Comparative genomic analysis reveals metabolic flexibility of Woesearchaeota", NATURE COMMUNICATIONS, vol. 12, no. 1, UK, pages 1 - 14, XP093271184, ISSN: 2041-1723, DOI: 10.1038/s41467-021-25565-9 * |
| JAY ZJ ET AL.: "The distribution, diversity and function of predominant Thermoproteales in high-temperature environments of Yellowstone National Park", ENVIRONMENTAL MICROBIOLOGY, vol. 18, no. 12, 2016, pages 4755 - 4769, XP072197230, DOI: 10.1111/1462-2920.13366 * |
| LAND HENRIK, CECCALDI PIERRE, MÉSZÁROS LÍVIA S., LORENZI MARCO, REDMAN HOLLY J., SENGER MORITZ, STRIPP SVEN T., BERGGREN GUSTAV: "Discovery of novel [FeFe]-hydrogenases for biocatalytic H 2 -production", CHEMICAL SCIENCE, vol. 10, no. 43, 6 November 2019 (2019-11-06), UK, pages 9941 - 9948, XP093271217, ISSN: 2041-6520, DOI: 10.1039/C9SC03717A * |
| LAND HENRIK, SENGER MORITZ, BERGGREN GUSTAV, STRIPP SVEN T.: "Current State of [FeFe]-Hydrogenase Research: Biodiversity and Spectroscopic Investigations", ACS CATALYSIS, vol. 10, no. 13, 2 July 2020 (2020-07-02), US , pages 7069 - 7086, XP093271203, ISSN: 2155-5435, DOI: 10.1021/acscatal.0c01614 * |
| LEUNG POK MAN, GRINTER RHYS, TUDOR-MATTHEW EVE, JIMENEZ LUIS, LEE HAN, MILTON MICHAEL, HANCHAPOLA IRESHA, TANUWIDJAYA ERWIN, PEACH: "Atmospheric hydrogen oxidation extends to the domain archaea", BIORXIV, US, 13 December 2022 (2022-12-13), US, pages 1 - 37, XP093271181, Retrieved from the Internet <URL:https://www.biorxiv.org/content/10.1101/2022.12.13.520232v1.full.pdf> DOI: 10.1101/2022.12.13.520232 * |
| MÉSZÁROS LÍVIA S., LAND HENRIK, REDMAN HOLLY J., BERGGREN GUSTAV: "Semi-synthetic hydrogenases—in vitro and in vivo applications", CURRENT OPINION IN GREEN AND SUSTAINABLE CHEMISTRY, vol. 32, 1 December 2021 (2021-12-01), pages 100521 - 100521-8, XP093271208, ISSN: 2452-2236, DOI: 10.1016/j.cogsc.2021.100521 * |
| SIMMONS TREVOR R., BERGGREN GUSTAV, BACCHI MARINE, FONTECAVE MARC, ARTERO VINCENT: "Mimicking hydrogenases: From biomimetics to artificial enzymes", COORDINATION CHEMISTRY REVIEWS, vol. 270-271, 1 July 2014 (2014-07-01), NL , pages 127 - 150, XP093271214, ISSN: 0010-8545, DOI: 10.1016/j.ccr.2013.12.018 * |
| SPANG A ET AL.: "Proposal of the reverse flow model for the origin of the eukaryotic cell based on comparative analyses of Asgard archaeal metabolism", NATURE MICROBIOLOGY, vol. 4, no. 7, 2019, pages 1138 - 1148, XP036815857, DOI: 10.1038/s41564-019-0406-9 * |
| SUSANTI DWI, JOHNSON ERIC F., RODRIGUEZ JASON R., ANDERSON IAIN, PEREVALOVA ANNA A., KYRPIDES NIKOS, LUCAS SUSAN, HAN JAMES, LAPID: "Complete Genome Sequence of Desulfurococcus fermentans, a Hyperthermophilic Cellulolytic Crenarchaeon Isolated from a Freshwater Hot Spring in Kamchatka, Russia", JOURNAL OF BACTERIOLOGY, vol. 194, no. 20, 15 October 2012 (2012-10-15), US , pages 5703 - 5704, XP093271177, ISSN: 0021-9193, DOI: 10.1128/JB.01314-12 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Greening et al. | Minimal and hybrid hydrogenases are active from archaea | |
| Reeve et al. | Methanogenesis: genes, genomes, and who's on first? | |
| Ramazzina et al. | Completing the uric acid degradation pathway through phylogenetic comparison of whole genomes | |
| Kerscher et al. | A single external enzyme confers alternative NADH: ubiquinone oxidoreductase activity in Yarrowia lipolytica | |
| Kaster et al. | More than 200 genes required for methane formation from H2 and CO2 and energy conservation are present in Methanothermobacter marburgensis and Methanothermobacter thermautotrophicus | |
| Thorn et al. | Crystal structure ofEscherichia coliQOR quinone oxidoreductase complexed with NADPH | |
| AU2009296225B2 (en) | Identification and use of bacterial [2Fe-2S] dihydroxy-acid dehydratases | |
| EP3145947B1 (en) | Recombinant microorganisms capable of carbon fixation | |
| CN111019878B (en) | Recombinant Escherichia coli with improved L-threonine yield, construction method and application thereof | |
| US20140038263A1 (en) | Activity of Fe-S Cluster Requiring Proteins | |
| US20160108434A1 (en) | Photocatalytic hydrogen production and polypeptides capable of same | |
| CN103981158A (en) | Mutant enzyme and application thereof | |
| EP1272639A2 (en) | Genes encoding denitrification enzymes | |
| WO2021084526A1 (en) | Engineered autotrophic bacteria for co2 conversion to organic materials | |
| CN115838702B (en) | A 3′-O-methyltransferase mutant and its application | |
| CN119776314B (en) | Acyltransferase mutant and application thereof in synthesis of N-acetyl-trans-4-hydroxyproline | |
| Peng et al. | Structure-function analysis indicates that an active-site water molecule participates in dimethylsulfoniopropionate cleavage by DddK | |
| Frank et al. | Elucidation of substrate specificity in the cobalamin (vitamin B12) biosynthetic methyltransferases: structure and function of the C20 methyltransferase (CbiL) from Methanothermobacter thermautotrophicus | |
| Constantine et al. | Biochemical and structural studies of N 5-carboxyaminoimidazole ribonucleotide mutase from the acidophilic bacterium Acetobacter aceti | |
| WO2003095649A1 (en) | Nobel glyphosate-tolerant 5-enolpyruvylshikimate-3-phospha synthase and gene encoding it | |
| CN108998462B (en) | Escherichia coli expression system for recombinant protein containing manganese ions and its application method | |
| CN107267474B (en) | A kind of dihydrolipoamide dehydrogenase mutant protein and its preparation method and application | |
| WO2025019889A1 (en) | Novel hydrogenases | |
| US20160237442A1 (en) | Modified group i methanotrophic bacteria and uses thereof | |
| CN113999827A (en) | A kind of leucine dehydrogenase mutant and its preparation method and application |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24844102 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
