EP4732286A2 - Compositions and methods for encrypting, storing, and decrypting information in oligomers - Google Patents

Compositions and methods for encrypting, storing, and decrypting information in oligomers

Info

Publication number
EP4732286A2
EP4732286A2 EP24866009.4A EP24866009A EP4732286A2 EP 4732286 A2 EP4732286 A2 EP 4732286A2 EP 24866009 A EP24866009 A EP 24866009A EP 4732286 A2 EP4732286 A2 EP 4732286A2
Authority
EP
European Patent Office
Prior art keywords
monomers
bits
less
oligomer
unique
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24866009.4A
Other languages
German (de)
French (fr)
Inventor
Eric Anslyn
Livia SCHIAVINATO EBERLIN
Sarah MOOR
Samuel DAHLHAUSER
Mary King
Christopher Wight
James R. Howard
Julia SHULUK
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of Texas System
University of Texas at Austin
Original Assignee
University of Texas System
University of Texas at Austin
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of Texas System, University of Texas at Austin filed Critical University of Texas System
Publication of EP4732286A2 publication Critical patent/EP4732286A2/en
Pending legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA

Landscapes

  • Genetics & Genomics (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biotechnology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Organic Chemistry (AREA)
  • Zoology (AREA)
  • General Engineering & Computer Science (AREA)
  • Wood Science & Technology (AREA)
  • Biomedical Technology (AREA)
  • Microbiology (AREA)
  • Biophysics (AREA)
  • Physics & Mathematics (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Molecular Biology (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Plant Pathology (AREA)
  • Other Investigation Or Analysis Of Materials By Electrical Means (AREA)
  • Storage Device Security (AREA)
  • Compositions Of Macromolecular Compounds (AREA)

Abstract

Disclosed herein are compositions and methods for encrypting, storing, and decrypting information in oligomers.

Description

COMPOSITIONS AND METHODS FOR ENCRYPTING, STORING, AND DECRYPTING INFORMATION IN OLIGOMERS
CROSS-REFERENCE TO RELATED APPLICATIONS
This application claims the benefit of priority to U.S. Provisional Application No. 63/462.267 filed April 27, 2023. which is hereby incorporated herein by reference in its entirety.
STATEMENT OF GOVERNMENT SUPPORT
This invention was made with government support under Grant No. W911NF-17-1- 0522 awarded by The Army Research Office. The government has certain rights in the invention.
BACKGROUND
Synthetic sequence defined polymers have garnered significant interest, however information storage remains an underused function of these synthetic macromolecules. Reliable data storage is becoming an increasingly significant challenge. Encoding data at the molecular level could dramatically increase storage densities and overcome some of the significant drawbacks encountered with conventional silicon-based data storage, such as durability and longevity. However, high-throughput, simple, and facile means for writing and reading information in synthetic macromolecules are still needed. The compositions and methods discussed herein address these and other needs.
SUMMARY
In accordance with the purposes of the disclosed compositions and methods as embodied and broadly described herein, the disclosed subject matter relates to compositions and methods for encrypting, storing, and decrypting information in oligomers.
For example, disclosed herein are methods for decrypting information stored within a target oligomer, the target oligomer comprising a self-immolative oligourethane comprising a plurality of unique monomers in a pre-defined order and an endcap, wherein the target oligomer has a first end and a second end. the second end comprising the endcap; wherein each of the unique monomers has a unique mass spectrometry profile; the method comprising: creating a time-interval substrate, the time-interval substrate comprising a plurality of samples each collected at a different time-interval, the plurality of samples being disposed on a substrate in an ordered array; wherein each sample comprises a portion of a mixture formed by subjecting the target oligomer to 5-exo-trig cyclization and elimination; analyzing the time-interval substrate using desorption electrospray ionization (DESI) mass spectrometry, thereby generating a plurality of mass spectrometry profiles; and evaluating the plurality7 of mass spectrometry profiles to decrypt the information stored within the target oligomer.
In some examples, creating the time-interval substrate comprises: subjecting the target oligomer to 5-exo-trig cyclization and elimination, thereby forming a mixture; collecting a plurality of aliquots of the mixture over a plurality7 of time-intervals; placing each of the aliquots at a location on a substrate, such that the plurality of aliquots are disposed on the substrate in an ordered array, the location in the array corresponding to the time-interval at which the aliquot was collected.
In some examples, the substrate comprises a PTFE-coated slide, a multi-well glass slide, or a combination thereof.
In some examples, the plurality of samples further comprise a solvent. In some examples, the solvent comprises acetonitrile.
In some examples, the ordered array comprises a two-dimensional array.
In some examples, the plurality of samples comprises from 2 to 256 samples, such as from 8 to 32 samples.
In some examples, the time-intervals occur at regular intervals. In some examples, the time-intervals independently occur at an interval of from 15 minutes to 120 minutes.
In some examples, the plurality of unique monomers comprises from 2 to 256 unique monomers, such as from 8 to 32 unique monomers.
In some examples, the target oligomer has a total length of from 2 to 256 monomers, such as from 8 to 32 monomers.
In some examples, the self-immolative oligourethane is derived from P-amino alcohols.
In some examples, the mass spectrometry profiles are generated in positive-ion mode.
In some examples, the time-interval substrate is disposed on a movable stage, and analyzing the time-interval substrate using DESI-MS comprises translating the movable stage to sequentially subject each of the plurality of samples to the DESI-MS analysis. In some examples, the movable stage is translated using a constant velocity motion profile. In some examples, the constant velocity motion profile comprises a stage velocity of from 500 to 3000 pm/second.
In some examples, the endcap comprises rhodamine B or methyl tyrosine. In some examples, the information stored within the target oligomer is hexadecimal based.
In some examples, the information stored within the target oligomer comprises a cipher key.
In some examples, the information stored within the target oligomer comprises from 2 to 256 bits of information.
In some examples, the method further comprises steganography.
Also disclosed herein are methods for encrypting information within a target oligomer, the methods comprising: selecting a plurality of unique monomers, wherein each of the unique monomers has a unique mass spectrometry’ profile; assigning a unique value to each of the unique monomers within the plurality; synthesizing a target oligomer comprising a self-immolative oligourethane comprising the plurality' of unique monomers in a predefined order and an endcap, wherein the target oligomer has a first end and a second end, the second end comprising the endcap; wherein the endcap comprises rhodamine B or methyl ty rosine; wherein the assigned values and the pre-defined order encrypts pre-defined information into the target oligomer. In some examples, the plurality of unique monomers comprises from 2 to 256 unique monomers, such as from 8 to 32 unique monomers. In some examples, the target oligomer has a total length of from 2 to 256 monomers, such as from 8 to 32 monomers. In some examples, the self-immolative oligourethane is derived from f>-amino alcohols. In some examples, the information stored within the target oligomer is hexadecimal based. In some examples, the information stored within the target oligomer comprises a cipher key. In some examples, the information stored within the target oligomer comprises from 2 to 256 bits of information. In some examples, the method further comprises steganography.
Also disclosed herein are compositions for molecular cry ptography comprising: a target oligomer comprising a self-immolative oligourethane comprising a plurality of unique monomers in a pre-defined order and an endcap, wherein the target oligomer has a first end and a second end, the second end comprising the endcap; wherein each of the unique monomers has a unique mass spectrometry profile; wherein the endcap comprises rhodamine B or methyl tyrosine. In some examples, the composition further comprises a solvent. In some examples, the solvent comprises acetonitrile. In some examples, the composition further comprises: a truncated oligomer comprising a self-immolative oligourethane comprising at least a portion of the plurality7 of unique monomers and a truncated endcap, wherein the truncated oligomer has a leading end and a trailing end, the trailing end comprising the truncated endcap; wherein the number of monomers in the truncated oligomer is less than that of the target oligomer, the order of monomers in the truncated oligomer differs from the pre-defined order of the target oligomer by 1 monomer or more, or a combination thereof; and wherein the truncated cap comprises an anhydride. In some examples, the truncated cap comprises acetic anhydride. In some examples, the plurality of unique monomers comprises from 2 to 256 unique monomers, such as from 8 to 32 unique monomers. In some examples, the target oligomer has a total length of from 2 to 256 monomers, such as from 8 to 32 monomers. In some examples, the self-immolative oligourethane is derived from [3-amino alcohols. In some examples, the information stored within the target oligomer is hexadecimal based. In some examples, the information stored within the target oligomer comprises a cipher key. In some examples, the information stored within the target oligomer comprises from 2 to 256 bits of information.
Additional advantages of the disclosed compositions and methods will be set forth in part in the description which follows, and in part will be obvious from the description. The advantages of the disclosed compositions and methods will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory’ only and are not restrictive of the disclosed systems and methods, as claimed.
The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will be apparent from the description and drawings, and from the claims.
BRIEF DESCRIPTION OF THE FIGURES
The accompanying figures, which are incorporated in and constitute a part of this specification, illustrate several aspects of the disclosure, and together with the description, serve to explain the principles of the disclosure.
Figure 1. Sequencing of trimer 3 in 2.5: 1 Me0H:H20 with K3PO4. Reaction was heated to 70 °C in a microwave and sampled at the time intervals by LC/MS.
Figure 2. Sequencing of trimer 3 in 2.5: 1 MeOH:H2O with K3PO4. Reaction was heated to 70 °C in a microwave and sampled at the denoted time intervals by LC/MS.
Figure 3. An optimized Huffman tree generated from the exact frequencies of the text. Figure 4. Representation of how to convert between binary to octal and hexadecimal. Figure 5. The amino alcohols used to encode the information, with their octal notation.
Figure 6. Structures of Oligomers Al and A2.
Figure 7. Decoding scheme for Hello, World!
Figure 8. Reading the sequence of the oligomer “Al” by single quadrupole mass spectrometry after 300 minutes of immolation. Calculating the mass differences between parent oligomer masses (in blue, mass spectra masses circled in green) gives the molecular weight of the monomer (in red). The reading frame starting with the disappearance of phenylalaninol index as the 2-oxazolidinone. The octal symbol is then correlated to the monomer (in black).
Figure 9. Structure of the Al octamer, and corresponding chromatogram of the LC/MS at the 300 min time point, and low-resolution mass spectra used to calculate the mass differences.
Figure 10. Oligomer A2 is sequenced. LC/MS chromatograms of A2 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 11. Structure of the Al octamer, and corresponding chromatogram of the LC/MS at the 300 min time point, and low-resolution mass spectra used to calculate the mass differences.
Figure 12. Structure of the Al heptamer, and corresponding chromatogram of the LC/MS at the 300 min time point, and low-resolution mass spectra used to calculate the mass differences.
Figure 13. Structure of the Al hexamer, and corresponding chromatogram of the LC/MS at the 300 min time point, and low-resolution mass spectra used to calculate the mass differences.
Figure 14. Structure of the Al pentamer, and corresponding chromatogram of the LC/MS at the 300 min time point, and low-resolution mass spectra used to calculate the mass differences.
Figure 15. Structure of the Al tetramer, and corresponding chromatogram of the LC/MS at the 300 min time point, and low-resolution mass spectra used to calculate the mass differences.
Figure 16. Structure of the Al trimer, and corresponding chromatogram of the LC/MS at the 300 min time point, and low-resolution mass spectra used to calculate the mass differences.
Figure 17. Structure of the Al dimer, and corresponding chromatogram of the LC/MS at the 300 min time point, and low-resolution mass spectra used to calculate the mass differences.
Figure 18. Structure of the Al monomer, and corresponding chromatogram of the LC/MS at the 300 min time point, and low-resolution mass spectra used to calculate the mass differences.
Figure 19. Structures of the 16 monomers synthesized for solid phase synthesis with their respective labels in text, octal, and hexadecimal.
Figure 20. The information from Jane Austen’s Mansfield Park, given to Mary from Mrs. Grant on the topic of marriage, was converted to binary via an optimized Huffman Tree, then converted to hexadecimal by the standard ASCII conversion.
Figure 21. An overlay of the information in its various encodings. In red, the molecular form, in blue the hex string, in black the bit string, and in bold is written the English text.
Figure 22. LC/MS chromatograms of G1 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 23. LC/MS chromatograms of G2 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 24. LC/MS chromatograms of G3 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 25. LC/MS chromatograms of G4 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 26. LC/MS chromatograms of G5 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 27. LC/MS chromatograms of G6 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 28. LC/MS chromatograms of G7 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 29. LC/MS chromatograms of G8 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 30. LC/MS chromatograms of G9 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study. Figure 31. LC/MS chromatograms of GIO at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 32. LC/MS chromatograms of Gil at 470 nm, showing the gradual appearance and disappearance of oligomers throughout the time-course study.
Figure 33. LC/MS chromatograms of G12 at 470 nm, showing the gradual appearance and disappearance of oligomers throughout the time-course study.
Figure 34. LC/MS chromatograms of Hl at 470 nm, showing the gradual appearance and disappearance of oligomers throughout the time-course study.
Figure 35. LC/MS chromatograms of H2 at 470 nm, showing the gradual appearance and disappearance of oligomers throughout the time-course study.
Figure 36. LC/MS chromatograms of H3 at 470 nm, showing the gradual appearance and disappearance of oligomers throughout the time-course study.
Figure 37. LC/MS chromatograms of H4 at 470 nm, showing the gradual appearance and disappearance of oligomers throughout the time-course study.
Figure 38. LC/MS chromatograms of H5 at 470 nm, showing the gradual appearance and disappearance of oligomers throughout the time-course study.
Figure 39. LC/MS chromatograms of H6 at 470 nm, showing the gradual appearance and disappearance of immolating oligomers throughout the time-course study.
Figure 40. Instructions for blind participant.
Figure 41. Spectra of the oligomers in which the participant chose an incorrect mass. The "deletion" peak was the peak incorrectly chosen, and “8mer’' shows the peak that should have been chosen.
Figure 42. Amended instructions for the blind study participant.
Figure 43. Computationally generated isotope patterns.
Figure 44. Mass Spectra of monobromo, monochloro, dibromo, and di chloro monomers.
Figure 45. Stoichiometric mixing of isotopologue monomers (Ala and D2Ala) can be observed and the ratio calculated via high resolution mass spectrometry.
Figure 46. Stoichiometric mixing of isotopologue monomers (CHA and D2CHA) can be observed and the ratio calculated via high resolution mass spectrometry.
Figure 47. Stoichiometric mixing of isotopologue monomers (Vai and D2Val) can be observed and the ratio calculated via high resolution mass spectrometry.
Figure 48. Mass spectra of isotopologues Al, A2, and A3. The increase at the M+2 peak is directly correlated to the stoichiometry of non-deuterated to deuterated oligomer.
Figure 49. Mass spectra of monomers Fmoc-4ClPhe-ol and Fmoc-4BrPhe-ol, and oligomers Al and A2, as well as theoretical mass spectra of molecular formulas corresponding to oligomers Al and A2 if they had halogen isotope tags.
Figure 50. Isotope tags using isotopologues and halogen tags to provide specific “fingerprints” for each oligourethane via distinct and predictable isotope patterns.
Figure 51. Structure of decamer Bl, and corresponding LC trace at 470 nm. High resolution mass spectra Bl.
Figure 52. Structure of decamer B2, and corresponding LC trace at 470 nm. High resolution mass spectra B2.
Figure 53. Structure of decamer B3, and corresponding LC trace at 470 nm. High resolution mass spectra B3.
Figure 54. Structure of decamer B4, and corresponding LC trace at 470 nm. High resolution mass spectra B4.
Figure 55. Structure of decamer B5, and corresponding LC trace at 470 nm. High resolution mass spectra B5.
Figure 56. Structure of decamer B6, and corresponding LC trace at 470 nm. High resolution mass spectra B6.
Figure 57. Structure of decamer B7, and corresponding LC trace at 470 nm. High resolution mass spectra B7.
Figure 58. Structure of decamer B8, and corresponding LC trace at 470 nm. High resolution mass spectra B8.
Figure 59. 500 nmols of each oligomer (B1-B8) was added to a reaction vial. Before sequencing, the sample was analyzed as a 0-minute time point.
Figure 60 . Concurrent sequencing of oligomers Bl - B8 in DMSO with CS2CO3. Reaction was heated to 70 °C and sampled at designated time intervals by LC/MS.
Figure 61. LC/MS traces of three of the eight information containing oligomers. As observed, the stable isotope tag imparts a unique mass spectrum onto each ion, allowing for each mass to be sorted and assigned to the appropriate oligomer. The mass differences between ions are then calculated and correlated back to the monomer that was cyclized and cleaved, revealing the information stored within the macromolecule.
Figure 62. Comparison of the MS of two oligomers (B3-8mer and B5-6mer) with the same mass but different isotope pattern allowing for differentiation. Figure 63. Handwritten letter sent to collaborators containing oligourethane ink.
Figure 64. H1 spectra of Fmoc-3,5-DiBrTyr-ol.
Figure 65. C13 spectrum of Fmoc-3,5-DiBrTyr-ol.
Figure 66. H1 spectrum of Fmoc-3,5-DiBrTyrOMe-ol.
Figure 67. HRMS of Fmoc-3,5-DiBrTyrOMe-ol.
Figure 68. C13 spectrum of Fmoc-3,5-DiBrTyrOMe-ol.
Figure 69. Schematic of typical DESI setup.
Figure 70. Overview of the MultiPep 2 automated synthesizer. A) The monomer tube rack where 38 different monomers can be held and added to the resin. B) The single and multichannel needles used add reagents and solvent to the resins as well as rinse them. C) Where the respective solvent bottles are held, it is possible to easily modify which solvents are being added as well as the size of the bottle used. D) The reaction plate, where the solid phase synthesis occurs in either fritted wells or test tubes containing pre-weighed out resin.
Figure 71. Example of using UV-Vis trace at 470 nm used to assist in identifying truncated oligourethane sequences.
Figure 72. Spectrum of Gil taken at the 150-minute time point.
Figure 73. Oligomer G5 from Jane Austin encoding ionized using DESI-MS in a range of solvents in negative mode.
Figure 74. DESI instrument holding a PTFE slide containing oligo samples.
Figure 75. Effect of solvent on S/N ratio for each oligomer unit.
Figure 76. Analysis of the sample oligourethanes.
Figure 77. Replicate of runs shown in Figure 76 of the sample oligourethanes ensuring consistency across spectra.
Figure 78. Potential end caps for DESI analysis. Tyr: Tyrosine, Gly: Glycine.
Figure 79. Summary of Tyr, Gly, Mono-Cl, and Tri-Cl capping experiments, all masses are the potassiated adduct [M+K],
Figure 80. New Rhodamine based caps for oligomer N-terminus capping. TAMRA: 5(6)-carboxytetramethylrhodamine, RhoB: Rhodamine B.
Figure 81. HRMS of RhoB oligomer.
Figure 82. HRMS of TAMRA oligo.
Figure 83. DESI spectra comparing RhoB and TAMRA. All RhoB masses highlighted in purple are [M+K],
Figure 84. DESI spectrum of TAMRA-2 with 0.2% formic acid in ACN. Figure 85. DESI spectrum of Tyr-2 in ACN.
Figure 86. DESI spectrum of RhoB-2 in its lactam form in ACN.
Figure 87. DESI spectrum RhoB-2 its uncyclized form in ACN.
Figure 88. Raw file structure.
Figure 89. Visualization of scans and their corresponding average ionization in a DESI run. The X-axis corresponds to scan number while the Y-axis is the average ionization in that scan. While not scanning over samples the average is very7 low and increases as the DESI spray capillary and detector approaches the samples on the slide.
Figure 90. Software interpretation of DESI scans.
Figure 91. Adding an alanine. SMILES Code: CC(CO)NC(=O)OCCS
Figure 92. Adding a Leucine. SMILES Code: CC(COC(=O)NC(CO)C(C)C)NC(=O)OCCS
Figure 93. Adding a Phenylalanine. SMILES Code:
CC(COC(=O)NC(COC(=O)NC(CO)Cclcccccl)C(C)C)NC(=O)OCCS
Figure 94. Adding an Alanine. SMILES Code:
CC(CO)NC(=O)OCC(Cclcccccl)NC(=O)OCC(NC(=O)OCC(C)NC(=O)OCCS)C(C)C
Figure 95. Rhodamine B.
Figure 96A. Self-sequencing of beta-amino alcohol derived oligourethanes in base proceeds 0->N via a 5-exo-trig mechanism by successively removing the O-terminal monomer.
Figure 96B. Solid-phase synthesis of sequence-defined oligourethanes.
Figure 96C. General structure of sequence-defined oligourethanes used herein.
Figure 97A. Schematic of robotic liquid handler used for automated synthesis of sequence-defined oligourethanes.
Figure 97B. Schematic of robotic liquid handler used for automated sequencing of sequence-defined oligourethanes.
Figure 98. Workflow overview for the "reading’ and ‘writing’ of information
Figure 99. An example of MS data for a successful trial of automated sequencing.
Figure 100. Example input and output for the hexadecimal version.
Figure 101. Example input and output for the hexadecimal version.
Figure 102. Equilibria between the uncyclized rhodamine and its lactam.
DETAILED DESCRIPTION
The compositions, methods, and systems described herein may be understood more readily by reference to the following detailed description of specific aspects of the disclosed subject matter and the Examples included therein.
Before the present compositions, methods, and systems are disclosed and described, it is to be understood that the aspects described below are not limited to specific synthetic methods or specific reagents, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting.
Also, throughout this specification, various publications are referenced. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art to which the disclosed matter pertains. The references disclosed are also individually and specifically incorporated by reference herein for the material contained in them that is discussed in the sentence in which the reference is relied upon.
General Definitions
In this specification and in the claims that follow, reference will be made to a number of terms, which shall be defined to have the following meanings.
Throughout the description and claims of this specification, the word “comprise’' and other forms of the word, such as “comprising” and “comprises.” means including but not limited to, and is not intended to exclude, for example, other additives, components, integers, or steps.
As used in the description and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a composition” includes mixtures of two or more such compositions, reference to “an agent” includes mixtures of two or more such agents, reference to “the component” includes mixtures of two or more such components, and the like.
“Optional” or “optionally” means that the subsequently described event or circumstance can or cannot occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
Ranges can be expressed herein as from “about” one particular value, and/or to “about” another particular value. By “about” is meant within 5% of the value, e.g., within 4, 3, 2, or 1% of the value. When such a range is expressed, another aspect includes from the one particular value and/or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another aspect. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.
Values can be expressed herein as an ‘"average” value. ‘"Average” generally refers to the statistical mean value.
By "‘substantially” is meant within 5%, e.g., within 4%, 3%, 2%, or 1%.
“Exemplary” means “an example of’ and is not intended to convey an indication of a preferred or ideal embodiment. “Such as” is not used in a restrictive sense, but for explanatory’ purposes.
It is understood that throughout this specification the identifiers “first” and ‘‘second” are used solely to aid in distinguishing the various components and steps of the disclosed subject matter. The identifiers “first” and “second” are not intended to imply any particular order, amount, preference, or importance to the components or steps modified by these terms.
References in the specification and concluding claims to parts by weight of a particular element or component in a composition denotes the weight relationship between the element or component and any other elements or components in the composition or article for which a part by weight is expressed. Thus, in a compound containing 2 parts by weight of component X and 5 parts by weight component Y, X and Y are present at a weight ratio of 2:5, and are present in such ratio regardless of whether additional components are contained in the compound.
A weight percent (wt. %) of a component, unless specifically stated to the contrary, is based on the total weight of the formulation or composition in which the component is included.
The term “or combinations thereof’ as used herein refers to all permutations and combinations of the listed items preceding the term. For example, “A, B, C, or combinations thereof’ is intended to include at least one of: A, B, C, AB, AC, BC, or ABC, and if order is important in a particular context, also BA. CA, CB. CBA. BCA, ACB, BAC. or CAB. Continuing with this example, expressly included are combinations that contain repeats of one or more item or term, such as BB, AAA, AB, BBC, AAABCCCC, CBBAAA, CAB ABB, and so forth. The skilled artisan will understand that ty pically there is no limit on the number of items or terms in any combination, unless otherwise apparent from the context.
Chemical Definitions
Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
The organic moieties mentioned when defining variable positions within the general formulae described herein (e.g., the term ‘"halogen”) are collective terms for the individual substituents encompassed by the organic moiety. The prefix Cn-Cm preceding a group or moiety indicates, in each case, the possible number of carbon atoms in the group or moiety that follows.
The term “ion,” as used herein, refers to any molecule, portion of a molecule, cluster of molecules, molecular complex, moiety, or atom that contains a charge (positive, negative, or both at the same time within one molecule, cluster of molecules, molecular complex, or moiety (e.g., zwitterions)) or that can be made to contain a charge. Methods for producing a charge in a molecule, portion of a molecule, cluster of molecules, molecular complex, moiety, or atom are disclosed herein and can be accomplished by methods known in the art, e.g., protonation, deprotonation, oxidation, reduction, alkylation, acetylation, esterification, de-esterification, hydrolysis, etc.
The term “anion” is a type of ion and is included within the meaning of the term “ion.” An “anion” is any molecule, portion of a molecule (e.g., zwitterion), cluster of molecules, molecular complex, moiety, or atom that contains a net negative charge or that can be made to contain a net negative charge. The term “anion precursor” is used herein to specifically refer to a molecule that can be converted to an anion via a chemical reaction (e.g., deprotonation).
The term “cation” is a type of ion and is included within the meaning of the term “ion.” A “cation” is any molecule, portion of a molecule (e.g., zwitterion), cluster of molecules, molecular complex, moiety, or atom, that contains a net positive charge or that can be made to contain a net positive charge. The term “cation precursor” is used herein to specifically refer to a molecule that can be converted to a cation via a chemical reaction (e g., protonation or alkylation).
As used herein, the term “substituted” is contemplated to include all permissible substituents of organic compounds. In a broad aspect, the permissible substituents include acyclic and cyclic, branched and unbranched, carbocyclic and heterocyclic, and aromatic and nonaromatic substituents of organic compounds. Illustrative substituents include, for example, those described below. The permissible substituents can be one or more and the same or different for appropriate organic compounds. For purposes of this disclosure, the heteroatoms, such as nitrogen, can have hydrogen substituents and/or any permissible substituents of organic compounds described herein which satisfy the valencies of the heteroatoms. This disclosure is not intended to be limited in any manner by the permissible substituents of organic compounds. Also, the terms “substitution” or “substituted with” include the implicit proviso that such substitution is in accordance with permitted valence of the substituted atom and the substituent, and that the substitution results in a stable compound, e.g., a compound that does not spontaneously undergo transformation such as by rearrangement, cyclization, elimination, etc.
“Z1,” “Z2,” “Z3,” and “Z4” are used herein as generic symbols to represent various specific substituents. These symbols can be any substituent, not limited to those disclosed herein, and when they are defined to be certain substituents in one instance, they can, in another instance, be defined as some other substituents.
The term “aliphatic” as used herein refers to anon-aromatic hydrocarbon group and includes branched and unbranched, alkyl, alkenyl, or alkynyl groups.
As used herein, the term “alkyl” refers to saturated, straight-chained or branched saturated hydrocarbon moieties. Unless otherwise specified, C1-C24 (e.g., C1-C22, C1-C20, Ci- Cis, C1-C16, C1-C14, C1-C12, C1-C10, Ci-Cs, C1-C6, or C1-C4) alkyl groups are intended. Examples of alkyd groups include methyl, ethyl, propyl. 1-methyl-ethyl, butyl, 1 -methylpropyl, 2-methyl-propyl, 1,1-dimethyl-ethyl, pentyl, 1 -methyl-butyl, 2-methyl-butyl, 3- methyl-butyl, 2,2-dimethyl-propyl, 1-ethyl-propyl, hexyl, 1,1-dimethyl-propyl, 1 ,2-dimethyl- propyl, 1 -methyl-pentyl, 2-methyl-pentyl, 3-methyl-pentyl, 4-methyl-pentyl, 1,1 -dimethylbutyl, 1,2-dimethyl-butyl, 1,3-dimethyl-butyl. 2,2-dimethyl-butyl, 2,3-dimethyl-butyl, 3.3- dimethyl-butyl, 1-ethyl-butyl, 2-ethyl-butyl, 1,1,2-trimethyl-propyl, 1 ,2,2-trimethyl-propyl, 1- ethyl-l-methyl-propyl, l-ethyl-2-methyl-propyl, heptyl, octyl, nonyl, decyl, dodecyl, tetradecy l, hexadecyl, eicosyl, tetracosyl, and the like. Alky l substituents may be unsubstituted or substituted with one or more chemical moieties. The alky l group can be substituted with one or more groups including, but not limited to. hydroxyl, halogen, acyl, alkyl, alkoxy, alkenyl, alkynyl, aryl, heteroaryl, aldehyde, amino, cyano, carboxylic acid, ester, ether, ketone, nitro, phosphonyl, silyl, sulfo-oxo, sulfonyl, sulfone, sulfoxide, or thiol, as described below, provided that the substituents are sterically compatible and the rules of chemical bonding and strain energy are satisfied.
Throughout the specification “alkyd” is generally used to refer to both unsubstituted alkyl groups and substituted alkyl groups; however, substituted alkyl groups are also specifically referred to herein by identifying the specific substituent(s) on the alkyl group. For example, the term “halogenated alkyd” specifically refers to an alkyl group that is substituted with one or more halides (halogens; e.g., fluorine, chlorine, bromine, or iodine). The term “alkoxyalkyl” specifically refers to an alkyd group that is substituted with one or more alkoxygroups. as described below. The term “alkylamino” specifically refers to an alkyl group that is substituted with one or more amino groups, as described below, and the like. When “alkyl” is used in one instance and a specific term such as “alkylalcohol” is used in another, it is not meant to imply that the term “alkyl” does not also refer to specific terms such as “alkylalcohol” and the like.
This practice is also used for other groups described herein. That is, while a term such as “cycloalkyl” refers to both unsubstituted and substituted cycloalkyl moieties, the substituted moieties can, in addition, be specifically identified herein; for example, a particular substituted cycloalkyl can be referred to as, e.g., an “alkylcycloalkyl.” Similarly, a substituted alkoxy can be specifically referred to as. e.g, a “halogenated alkoxy.” a particular substituted alkenyl can be, e.g., an “alkenylalcohol,” and the like. Again, the practice of using a general term, such as “cycloalkyl,” and a specific term, such as “alkylcycloalkyl,” is not meant to imply that the general term does not also include the specific term.
As used herein, the term “alkenyl” refers to unsaturated, straight-chained, or branched hydrocarbon moieties containing a double bond. Unless otherwise specified, C2-C24 (e.g., C2- C22, C2-C20, C2-C18, C2-C16, C2-C14, C2-C12, C2-C10, C2-C8, C2-C6, or C2-C4) alkenyl groups are intended. Alkenyl groups may contain more than one unsaturated bond. Examples include ethenyl, 1-propenyl, 2-propenyl. 1 -methylethenyl. 1-butenyl, 2-butenyl, 3-butenyl, 1-methyl- 1 -propenyl, 2-methyl- 1-propenyl, 1 -methyl-2-propenyl, 2-methyl-2-propenyl, 1 -pentenyl, 2- pentenyl, 3-pentenyl, 4-pentenyl, 1 -methyl- 1-butenyl, 2-methyl- 1-butenyl, 3-methyl-l- butenyl, 1 -methyl-2-butenyl, 2-methyl-2-butenyl, 3-methyl-2-butenyl, l-methyl-3-butenyl, 2- methyl-3-butenyl, 3-methyl-3-butenyl, l,l-dimethyl-2-propenyl, 1 ,2-dimethyl- 1-propenyl, l,2-dimethyl-2-propenyl, 1 -ethyl- 1-propenyl, 1 -ethyl-2-propenyl. 1 -hexenyl. 2 -hexenyl, 3- hexenyl, 4-hexenyl, 5-hexenyl, 1 -methyl- 1 -pentenyl, 2-methyl- 1 -pentenyl, 3-methyl-l- pentenyl, 4-methyl-l -pentenyl, 1 -methyl-2-pentenyl, 2-methyl-2-pentenyl, 3-methyl-2- pentenyl, 4-methyl-2-pentenyl, l-methyl-3-pentenyl, 2-methyl-3-pentenyl, 3-methyl-3- pentenyl. 4-methyl-3-pentenyl, l-methyl-4-pentenyl, 2-methyl-4-pentenyl, 3-methyl-4- pentenyl, 4-methyl-4-pentenyl, l,l-dimethyl-2-butenyl, l,l-dimethyl-3-butenyl, 1,2- dimethyl- 1 -butenyl, 1 ,2-dimethyl-2-butenyl, 1 ,2-dimethyl-3-butenyl, 1 ,3-dimethyl- 1 -butenyl, l,3-dimethyl-2-butenyl, l,3-dimethyl-3-butenyl, 2,2-dimethyl-3-butenyL 2,3-dimethyl-l- butenyl, 2,3-dimethyl-2-butenyl, 2,3-dimethyl-3-butenyl, 3,3-dimethyl-l-butenyl, 3,3- dimethyl-2-butenyl, 1 -ethyl- 1-butenyl, l-ethyl-2-butenyl, l-ethyl-3-butenyl, 2-ethyl-l- butenyl, 2-ethyl-2-butenyl. 2-ethyl-3-butenyl, 1.1.2-trimethyl-2-propenyl, 1 -ethyl- 1-methyl-
2-propenyl, 1 -ethyl-2-methyl- 1 -propenyl, and l-ethyl-2-methyl-2-propenyl. The term "vinyl’’ refers to a group having the structure -CH=CH2i 1 -propenyl refers to a group with the structure -CH=CH-CH3; and 2-propenyl refers to a group with the structure -CH2-CH=CH2. Asymmetric structures such as (Z’Z2)C=C(Z?Z4) are intended to include both the E and Z isomers. This can be presumed in structural formulae herein wherein an asymmetric alkene is present, or it can be explicitly indicated by the bond symbol C=C. Alkenyl substituents may be unsubstituted or substituted with one or more chemical moieties. Examples of suitable substituents include, for example, alkyl, alkoxy, alkenyl, alkynyl, aryl, heteroaryl, acyl, aldehyde, amino, cyano, carboxylic acid, ester, ether, halide, hydroxyl, ketone, nitro, phosphonyl, silyl, sulfo-oxo. sulfonyl, sulfone, sulfoxide, or thiol, as described below, provided that the substituents are sterically compatible and the rules of chemical bonding and strain energy are satisfied.
As used herein, the term “alkynyl’' represents straight-chained or branched hydrocarbon moieties containing a triple bond. Unless otherwise specified. C2-C24 (e.g., C2- C24, C2-C20, C2-C18, C2-C16, C2-C14, C2-C12, C2-C10, C2-C8, C2-C6, or C2-C4) alkynyl groups are intended. Alkynyl groups may contain more than one unsaturated bond. Examples include C2-Ce-alkynyl, such as ethynyl, 1-propynyl, 2-propynyl (or propargyl), 1-butynyl, 2-butynyl,
3-butynyl, l-methyl-2-propynyl. 1 -pentynyl, 2-pentynyl, 3-pentynyl, 4-pentynyl, 3-methyl-l- butynyl, l-methyl-2-butynyl, l-methyl-3-butynyl, 2-methyl-3-butynyL l,l-dimethyl-2- propynyl, 1 -ethyl-2-propynyl, 1 -hexynyl, 2-hexynyl, 3-hexynyl, 4-hexynyl, 5-hexynyl, 3- methyl-1 -pentynyl, 4-methyl-l -pentynyl, 1 -methyl-2-pentynyl, 4-methyl-2-pentynyl, 1- methyl-3-pentynyl, 2-methyl-3-pentynyl, 1 -methyl-4-pentynyl, 2-methyl-4-pentynyl, 3- methyl-4-pentynyl. l,l-dimethyl-2-butynyl. l,l-dimethyl-3-butynyl, 1.2-dimethyl-3-butynyl, 2,2-dimethyl-3-butynyl, 3,3-dimethyl-l-butynyl, l-ethyl-2-butynyl, l-ethyl-3-butynyl, 2- ethyl-3-butynyl, and l-ethyl-l-methyl-2-propynyl. Alkynyl substituents may be unsubstituted or substituted with one or more chemical moieties. Examples of suitable substituents include, for example, alkyl, alkoxy, alkenyl, alkynyl. aryl, heteroaryl, acyl, aldehyde, amino, cyano, carboxylic acid, ester, ether, halide, hydroxyl, ketone, nitro, phosphonyl, silyl, sulfo-oxo, sulfonyl, sulfone, sulfoxide, or thiol, as described below. As used herein, the term '‘ary 1,” as well as derivative terms such as aryloxy, refers to groups that include a monovalent aromatic carbocyclic group of from 3 to 50 carbon atoms. Ary l groups can include a single ring or multiple condensed rings. In some examples, ary l groups include Ce-Cio aryl groups. Examples of aryl groups include, but are not limited to. benzene, phenyl, biphenyl, naphthyl, tetrahydronaphthyl, phenylcyclopropyl, phenoxybenzene, and indanyl. The term “aryf’ also includes “heteroaryl,” which is defined as a group that contains an aromatic group that has at least one heteroatom incorporated within the ring of the aromatic group. Examples of heteroatoms include, but are not limited to, nitrogen, oxygen, sulfur, and phosphorus. The term "non-heleroaryl." which is also included in the term '‘aryl,” defines a group that contains an aromatic group that does not contain a heteroatom. The aryl substituents may be unsubstituted or substituted with one or more chemical moieties. Examples of suitable substituents include, for example, alky l, alkoxy, alkenyl, alkynyl, aryl, heteroaryl, acyl, aldehyde, amino, cyano, carboxylic acid, ester, ether, halide, hydroxyl, ketone, nitro, phosphonyl. silyl, sulfo-oxo, sulfonyl, sulfone, sulfoxide, or thiol as described herein. The term “biaryl” is a specific type of aryl group and is included in the definition of aryl. Biaryl refers to two aryl groups that are bound together via a fused ring structure, as in naphthalene, or are attached via one or more carbon-carbon bonds, as in biphenyl.
The term '‘cycloalkyl” as used herein is a non-aromatic carbon-based ring composed of at least three carbon atoms. Examples of cycloalkyd groups include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, etc. The term “heterocycloalkyl” is a cycloalkyl group as defined above where at least one of the carbon atoms of the ring is substituted with a heteroatom such as, but not limited to, nitrogen, oxygen, sulfur, or phosphorus. The cycloalky 1 group and heterocycloalkyl group can be substituted or unsubstituted. The cycloalkyl group and heterocycloalky 1 group can be substituted with one or more groups including, but not limited to, alkyl, alkoxy, alkeny l, alkynyl, aryl, heteroary l. acyl, aldehyde, amino, cyano, carboxylic acid, ester, ether, halide, hydroxyl, ketone, nitro, phosphonyl, silyl, sulfo-oxo, sulfonyl, sulfone, sulfoxide, or thiol as described herein.
The term “cycloalkeny 1” as used herein is a non-aromatic carbon-based ring composed of at least three carbon atoms and containing at least one double bound, i.e., C=C. Examples of cycloalkenyl groups include, but are not limited to, cyclopropenyl, cyclobutenyl, cyclopentenyl, cyclopentadienyl, cyclohexenyl, cyclohexadienyl, and the like. The term “heterocycloalkenyl” is a type of cycloalkenyl group as defined above and is included within the meaning of the term “cycloalkenyl,” where at least one of the carbon atoms of the ring is substituted with a heteroatom such as, but not limited to, nitrogen, oxygen, sulfur, or phosphorus. The cycloalkenyl group and heterocycloalkenyl group can be substituted or unsubstituted. The cycloalkenyl group and heterocycloalkenyl group can be substituted with one or more groups including, but not limited to, alkyl, alkoxy, alkenyl, alkynyl, aryl, heteroaryl, acyl, aldehyde, amino, cyano, carboxylic acid, ester, ether, halide, hydroxyl, ketone, nitro, phosphonyl, silyl, sulfo-oxo, sulfonyl, sulfone, sulfoxide, or thiol as described herein.
The term “cyclic group" is used herein to refer to either aryl groups, non-aryl groups (i.e., cycloalkyl, heterocycloalkyl, cycloalkenyl, and heterocycloalkenyl groups), or both. Cyclic groups have one or more ring systems (e.g., monocyclic, bicyclic, tricyclic, polycyclic, etc.) that can be substituted or unsubstituted. A cyclic group can contain one or more aryl groups, one or more non-aryl groups, or one or more aryl groups and one or more non-aryl groups.
The term “acyl’’ as used herein is represented by the formula -C(O)Z1 where Z1 can be a hydrogen, hydroxyl, alkoxy, alkyd, alkenyl, alkynyl, aryl, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above. As used herein, the term “acyl” can be used interchangeably with “carbonyl.” Throughout this specification “C(O)” or “CO” is a shorthand notation for C=O.
The term “acetal” as used herein is represented by the formula (Z1Z2)C(=OZ3)(=OZ4), where Z1, Z2, Z3, and Z4 can be, independently, a hydrogen, halogen, hydroxyl, alky l, alkenyl, alkynyl, aryl, heteroaryl. cycloalkyl, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above.
The term “alkanol” as used herein is represented by the formula Z'OH. where Z1 can be an alkyl, alkenyl, alkynyl, aryl, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above.
As used herein, the term “alkoxy” as used herein is an alkyl group bound through a single, terminal ether linkage; that is, an “alkoxy” group can be defined as to a group of the formula Z'-O-. where Z1 is unsubstituted or substituted alk l as defined above. Unless otherwise specified, alkoxy groups wherein Z1 is a Ci-C24 (e.g., C1-C22. C1-C20, Ci-Cis, Ci- Ci6, C1-C14, C1-C12, C1-C10, Ci-Cs, C1-C6, or C1-C4) alkyl group are intended. Examples include methoxy, ethoxy, propoxy, 1 -methyl-ethoxy, butoxy, 1-methyl-propoxy, 2-methyl- propoxy, 1,1 -dimethyl-ethoxy, pentoxy, 1-methyl-butyloxy, 2-methyl -butoxy, 3-methyl- butoxy, 2,2-di-methyl-propoxy, 1 -ethyl-propoxy, hexoxy, 1,1-dimethyl-propoxy, 1,2- dimethyl-propoxy, 1-methyl-pentoxy, 2-methyl-pentoxy, 3-methyl-pentoxy, 4-methyl- penoxy, 1,1-dimethyl-butoxy, 1 ,2-dimethyl-butoxy, 1,3-dimethyl-butoxy, 2,2-dimethyl- butoxy, 2,3-dimethyl-butoxy, 3,3-dimethyl-butoxy, 1-ethyl-butoxy. 2-ethylbutoxy, 1,1,2- trimethyl-propoxy. 1 ,2,2-trimethyl-propoxy. 1 -ethyl- 1-methyl-propoxy, and l-ethyl-2- methyl -propoxy.
The term “aldehyde” as used herein is represented by the formula — C(O)H. Throughout this specification “C(O)” is a shorthand notation for C=O.
The term “amino” as used herein are represented by the formula — \Z'Z2Z3. where Z1, Z2, and Z3 can each be substitution group as described herein, such as hydrogen, an alkyl, alkenyl, alkynyl, aryl, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above.
The terms “amide” or “amido” as used herein are represented by the formula — C(O)NZ'Z2, where Z1 and Z2 can each be substitution group as described herein, such as hydrogen, an alkyl, alkenyl, alkynyl, aryl, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalkyd, or heterocycloalkenyl group described above.
The term “anhydride” as used herein is represented by the formula Z1C(O)OC(O)Z2 where Z1 and Z2. independently, can be an alkyl, alkenyl, alkynyl, ary 1. heteroaryl, cycloalky 1, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above.
The term “cyclic anhydride” as used herein is represented by the formula: where Z1 can be an alkyl, alkenyl, alkynyl, aryl, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above.
The term “azide” as used herein is represented by the formula -N=N=N.
The term “carboxylic acid” as used herein is represented by the formula — C(O)OH.
A “carboxylate” or “carboxyl” group as used herein is represented by the formula — C(O)O'
The term “cyano” as used herein is represented by the formula — CN.
The term “ester” as used herein is represented by the formula — OC(O)Z' or — C(O)OZ1, where Z1 can be an alkyl, alkenyl, alkynyl, aryl, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above. The term '‘ether” as used herein is represented by the formula Z'OZ2. where Z1 and Z2 can be, independently, an alkyd, alkenyl, alkynyl, aryl, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalk 1, or heterocycloalkenyl group described above.
The term “epoxy” or “epoxide” as used herein refers to a cyclic ether with a three atom ring and can represented by the formula: where Z1, Z2, Z3, and Z4 can be, independently, an alkyl, alkenyl, alkynyl, aryl, heteroaryl, cycloalky l, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above.
The term “ketone” as used herein is represented by the formula Z1C(O)Z2, where Z1 and Z2 can be, independently, an alkyl, alkenyl, alkynyl, aryl, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above.
The term “halide” or “halogen” or “halo” as used herein refers to fluorine, chlorine, bromine, and iodine.
The term “hydroxyl” as used herein is represented by the formula — OH.
The term '‘nitro” as used herein is represented by the formula — NO2.
The term “phosphonyl” is used herein to refer to the phospho-oxo group represented by the formula — P(O)(OZ1)2, where Z1 can be hy drogen, an alky l, alkenyl, alkyny l, ar l, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above.
The term “silyl” as used herein is represented by the formula — SiZ'Z2Z3. where Z1, Z2, and Z3 can be, independently, hydrogen, alkyd, alkoxy, alkenyl, alkynyl. aryl, heteroary l, cycloalky l, cycloalkeny l, heterocycloalkyl, or heterocycloalkenyl group described above.
The term “sulfonyl” or “sulfone” is used herein to refer to the sulfo-oxo group represented by the formula — S(O)2Z1, where Z1 can be hydrogen, an alkyl, alkenyl, alkynyl, aryl, heteroaryl, cycloalkyl, cycloalkenyl, heterocycloalkyl, or heterocycloalkenyl group described above.
The term “sulfide” as used herein comprises the formula — S — .
The term '‘thiol” as used herein is represented by the formula — SH.
“R1,” “R2,” “R3,” “Rn,” etc., where n is some integer, as used herein can, independently, possess one or more of the groups listed above. For example, if R1 is a straight chain alkyl group, one of the hydrogen atoms of the alkyl group can optionally be substituted with a hydroxyl group, an alkoxy group, an amino group, an alkyl group, a halide, and the like. Depending upon the groups that are selected, a first group can be incorporated within second group or, alternatively, the first group can be pendant (i.e., attached) to the second group. For example, with the phrase “an alkyl group comprising an amino group,” the amino group can be incorporated within the backbone of the alkyl group. Alternatively, the amino group can be attached to the backbone of the alkyl group. The nature of the group(s) that is (are) selected will determine if the first group is embedded or attached to the second group.
Unless stated to the contrary, a formula with chemical bonds shown only as solid lines and not as wedges or dashed lines contemplates each possible stereoisomer or mixture of stereoisomer (e.g., each enantiomer, each diastereomer, each meso compound, a racemic mixture, or scalemic mixture).
Compositions and Methods
Disclosed herein are compositions and methods for encrypting, storing, and decrypting information in oligomers.
Decryption Methods
For example, disclosed herein are methods for decrypting information stored within a target oligomer, the target oligomer comprising a self-immolative oligourethane comprising a plurality of unique monomers in a pre-defined order and an endcap, wherein the target oligomer has a first end and a second end. the second end comprising the endcap, and wherein each of the unique monomers has a unique mass spectrometry profile.
As used herein “a target oligomer” and “the target oligomer” can refer to one or more target oligomers. Accordingly, the methods disclosed herein can comprise decrypting information stored within one or more target oligomers, each of the one or more target oligomers independently comprising a self-immolative oligourethane comprising a plurality of unique monomers in a pre-defined order and an endcap, wherein each of the one or more target oligomers has a first end and a second end, the second end comprising the endcap, and wherein each of the unique monomers has a unique mass spectrometry profile.
The methods comprise: creating a time-interval substrate, the time-interval substrate comprising a plurality of samples each collected at a different time-interval, the plurality of samples being disposed on a substrate in an ordered array (e.g., a tw o-dimensional array); wherein each sample comprises a portion of a mixture formed by subjecting the target oligomer to 5-exo-trig cyclization and elimination; analyzing the time-interval substrate using desorption electrospray ionization (DESI) mass spectrometry, thereby generating a plurality of mass spectrometry profiles; and evaluating the plurality of mass spectrometry profiles to decrypt the information stored within the target oligomer.
Self-immolative oligourethanes, 5-exo-trig cyclization and elimination thereof, and mass spectrometry analysis and evaluation thereof are described, for example, by Dahlhauser et al. JACS, 2020, 142(6), 2744-2749; Dahlhauser et al. Call Reports Physical Science, 2021, 2, 100393; and Dahlhauser et al. ACS Central Science, 2022, 8. 1125-1133; each of which is incorporated herein by reference for its description thereof.
In some examples, the plurality of unique monomers can comprise 2 or more unique monomers (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more. 45 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, or 225 or more). In some examples, the plurality' of unique monomers can comprise 256 or less unique monomers (e.g.. 250 or less, 225 or less, 200 or less, 175 or less, 150 or less, 125 or less, 100 or less. 90 or less, 80 or less, 70 or less, 60 or less, 50 or less. 45 or less, 40 or less, 35 or less. 30 or less, 25 or less. 20 or less. 15 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, or 5 or less). The number of unique monomers in the plurality of unique monomers can range from any of the minimum values described above to any of the maximum values described above. For example, the plurality of unique monomers can comprise from 2 to 256 unique monomers (e.g., from 2 to 129, from 129 to 256, from 2 to 50, from 50 to 100, from 100 to 150, from 150 to 200, from 200 to 256, from 2 to 250, from 2 to 200, from 2 to 150, from 2 to 100, from 2 to 75, from 2 to 32, from 2 to 25, from 2 to 10, from 4 to 256, from 6 to 256, from 8 to 256, from 10 to 256, from 25 to 256, from 32 to 256, from 50 to 256, from 75 to 256, from 100 to 256, 6 to 250, from 8 to 248, from 8 to 200, from 8 to 100, from 8 to 32, from 8 to 20, or from 8 to 16).
In some examples, the target oligomer has a total length of 2 or more monomers (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 60 or more, 70 or more, 80 or more. 90 or more, 100 or more. 125 or more. 150 or more. 175 or more, 200 or more, or 225 or more). In some examples, the target oligomer has a total length of 256 or less monomers (e.g., 250 or less, 225 or less, 200 or less, 175 or less, 150 or less, 125 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 45 or less, 40 or less, 35 or less. 30 or less, 25 or less, 20 or less, 15 or less, 10 or less. 9 or less. 8 or less, 7 or less, 6 or less, or 5 or less). The number of monomers in the target oligomer can range from any of the minimum values described above to any of the maximum values described above. For example, the target oligomer can have a total length of from 2 to 256 unique monomers (e.g., from 2 to 129, from 129 to 256, from 2 to 50, from 50 to 100, from 100 to 150, from 150 to 200, from 200 to 256, from 2 to 250, from 2 to 200, from 2 to 150, from 2 to 100, from 2 to 75, from 2 to 32, from 2 to 25, from 2 to 10, from 4 to 256, from 6 to 256, from 8 to 256, from 10 to 256, from 25 to 256. from 32 to 256, from 50 to 256, from 75 to 256, from 100 to 256, 6 to 250, from 8 to 248, from 8 to 200, from 8 to 100, from 8 to 32, from 8 to 20, or from 8 to 16).
In some examples, the self-immolative oligourethane is derived from P-amino alcohols.
In some examples, the mass spectrometry profiles are generated in positive-ion mode.
As used herein, an ‘endcap’ is a single monomer added to the end of the oligomer that persists throughout the sequencing reaction. So, it is present on the parent oligomer and all sequenced products. End caps have two major functions: (1) increase the relative signal intensity of the desired parent and sequenced oligomers from background ions and undesired byproducts; and (2) favor the formation of a single adduct type in DESI (e.g., all ions are observed as singly potassiated adducts - [M+K]+). Seeing sequencing products as a single adduct greatly reduces the complexity of the signal deconvolution process, as the m/z difference between ions simply matches the mass of the monomer lost.
The endcap can comprise any suitable composition. Examples include, but are not limited to, 1-naphthol, glycine, TAMRA (5-Carboxytetramethylrhodamine), methyl tyrosine (Tyr(OMe)), tyrosine (Tyr(OH)), nitrobenzoxadiazole (NBD), and Rhodamine B. In some examples, the endcap comprises rhodamine B or methyl tyrosine.
In some examples, creating the time-interval substrate comprises: subjecting the target oligomer to 5-exo-trig cyclization and elimination, thereby forming a mixture; collecting a plurality7 of aliquots of the mixture over a plurality7 of time-intervals; placing each of the aliquots at a location on a substrate, such that the plurality of aliquots are disposed on the substrate in an ordered array, the location in the array corresponding to the time-interval at which the aliquot was collected.
In some examples, creating the time-interval substrate can further comprise further treating and/or processing each of the aliquots after collection and before placing on the substrate. For example, the solvent from the aliquot can be evaporated to provide a residue, and the residue can then be redissolved in a known volume of a second solvent before then being placed on the substrate. The substrate can comprise any suitable substrate. Examples of suitable substrates include, but are not limited to, polymers, glass, quartz, silicon, metals, ceramics, nitrides (e.g. silicon nitride), porous materials, and combinations thereof. In some examples, the substrate comprises a PTFE-coated slide, a multi-well glass slide, or a combination thereof.
In some examples, the plurality of samples further comprise a solvent. The solvent can comprise any suitable solvent. The solvent can, for example, comprise tetrahydrofuran (THF), N-methyl-2-pyrrolidone (NMP), dimethylformamide (DMF), N-methylformamide, formamide, dichloromethane (CH2CI2), ethylene glycol, polyethylene glycol, glycerol, alkane diol, ethanol, methanol, propanol, isopropanol, water, acetonitrile, chloroform, toluene, methyl acetate, ethyl acetate, acetone, hexane, heptane, tetraglyme, propylene carbonate, diglyme, dimethyl sulfoxide (DMSO), dimethoxy ethane, xylene, dimethylacetamide, methylene chloride, hexafluoro-2-propanol, or combinations thereof. In some examples, the solvent comprises acetonitrile.
In some examples, the plurality of samples can comprise 2 or more samples (e.g.. 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more. 150 or more, 175 or more, 200 or more, or 225 or more). In some examples, the plurality of samples can comprise 256 or less samples (e.g., 250 or less, 225 or less, 200 or less, 175 or less, 150 or less, 125 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 45 or less, 40 or less, 35 or less. 30 or less, 25 or less, 20 or less, 15 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, or 5 or less). The number of samples in the plurality of samples can range from any of the minimum values described above to any of the maximum values described above. For example, the plurality of samples can comprise from 2 to 256 samples (e.g., from 2 to 129, from 129 to 256, from 2 to 50, from 50 to 100, from 100 to 150, from 150 to 200. from 200 to 256. from 2 to 250, from 2 to 200, from 2 to 150, from 2 to 100, from 2 to 75. from 2 to 32, from 2 to 25, from 2 to 10, from 4 to 256. from 6 to 256, from 8 to 256, from 10 to 256, from 25 to 256, from 32 to 256, from 50 to 256, from 75 to 256, from 100 to 256, 6 to 250, from 8 to 248, from 8 to 200, from 8 to 100, from 8 to 32, from 8 to 20, or from 8 to 16).
The time-intervals can. for example, occur at regular intervals (e.g.. each interval being an equal amount of time) or irregular intervals. In some examples, the time-intervals are regular intervals. In some examples, the time-intervals independently occur at an interval of 15 minutes or more (e.g., 20 minutes or more, 25 minutes or more, 30 minutes or more, 35 minutes or more, 40 minutes or more, 45 minutes or more, 50 minutes or more, 55 minutes or more, 60 minutes or more, 65 minutes or more, 70 minutes or more, 75 minutes or more, 80 minutes or more. 85 minutes or more, 90 minutes or more, 95 minutes or more, 100 minutes or more, 105 minutes or more. 110 minutes or more, or 115 minutes or more). In some examples, the time-intervals independently occur at an interval of 120 minutes or less (e.g., 115 minutes or less, 110 minutes or less, 105 minutes or less, 100 minutes or less, 95 minutes or less, 90 minutes or less, 85 minutes or less, 80 minutes or less, 75 minutes or less, 70 minutes or less, 65 minutes or less, 60 minutes or less, 55 minutes or less, 50 minutes or less, 45 minutes or less, 40 minutes or less, 35 minutes or less, 30 minutes or less, 25 minutes or less, or 20 minutes or less). The interval at which the time-intervals occur can independently range from any of the minimum values described above to any of the maximum values described above. For example, the time-intervals can independently occur at an interval of from 15 minutes to 120 minutes (e.g., from 15 to 60 minutes, from 60 from 120 minutes, from 15 to 40 minutes, from 40 to 65 minutes, from 65 to 90 minutes, from 90 to 120 minutes, from 15 to 30 minutes, from 30 to 45 minutes, from 45 to 60 minutes, from 60 to 75 minutes, from 75 to 90 minutes, from 90 to 105 minutes, from 105 to 120 minutes, from 15 to 105 minutes, from 15 to 90 minutes, from 15 to 75 minutes, from 15 to 45 minutes, from 30 to 120 minutes, from 45 to 120 minutes, from 75 to 120 minutes, from 105 to 120 minutes, from 20 to 115 minutes, from 30 to 90 minutes, from 45 to 75 minutes, or from 55 to 65 minutes).
In some examples, the time-interval substrate is disposed on a movable stage, and analyzing the time-interval substrate using DESI-MS comprises translating the movable stage to sequentially subject each of the plurality of samples to the DESI-MS analysis.
In some examples, the movable stage is translated using a constant velocity motion profile. In some examples, the constant velocity motion profile comprises a stage velocity of 500 micrometers per second (pm/second) or more (e.g.. 525 pm/second or more. 550 pm/second or more, 575 pm/second or more, 600 pm/second or more, 625 pm/second or more, 650 pm/second or more, 675 pm/second or more, 700 pm/second or more, 725 pm/second or more, 750 pm/second or more, 800 pm/second or more, 850 pm/second or more, 900 pm/second or more, 950 pm/second or more, 1000 pm/second or more, 1100 pm/second or more, 1200 pm/second or more, 1300 pm/second or more, 1400 pm/second or more, 1500 pm/second or more, 1600 pm/second or more, 1700 pm/second or more, 1800 pm/second or more, 1900 pm/second or more, 2000 pm/second or more, 2250 pm/second or more, 2500 pm/second or more, or 2750 pm/second or more). In some examples, the constant velocity' motion profile comprises a stage velocity of 3000 pm/second or less (e.g., 2750 pm/second or less, 2500 pm/second or less, 2250 pm/second or less, 2000 pm/second or less, 1900 pm/second or less. 1800 pm/second or less, 1700 pm/second or less, 1600 pm/second or less, 1500 pm/second or less, 1400 pm/second or less, 1300 pm/second or less, 1200 pm/second or less, 1100 pm/second or less, 1000 pm/second or less, 950 pm/second or less, 900 pm/second or less, 850 pm/second or less, 800 pm/second or less, 750 pm/second or less, 725 pm/second or less, 700 pm/second or less. 675 pm/second or less, 650 pm/second or less, 625 pm/second or less, 600 pm/second or less, 575 pm/second or less, 550 pm/second or less, or 525 pm/second or less). The stage velocity can range from any of the minimum values described above to any of the maximum values described above. For example, the constant velocity motion profile can comprise a stage velocity’ of from 500 to 3000 pm/second (e.g.. from 500 to 1750 pm/second, from 1750 to 3000 pm/second. from 500 to 1000 pm/second, from 1000 to 1500 pm/second, from 1500 to 2000 pm/second, from 2000 to 2500 pm/second, from 2500 to 3000 pm/second, from 500 to 2750 pm/second, from 500 to 2500 pm/second, from 500 to 2250 pm/second, from 500 to 2000 pm/second. from 500 to 1500 pm/second, from 500 to 750 pm/second, from 750 to 3000 pm/second, from 1000 to 3000 pm/second, from 1250 to 3000 pm/second, from 1500 to 3000 pm/second, from 2000 to 3000 pm/second, from 2250 to 3000 pm/second, from 2750 to 3000 pm/second, from 750 to 2750 pm/second, from 1000 to 2500 pm/second, from 1000 to 2000 pm/second, from 1250 to 1750 pm/second, or from 1400 to 1600 pm/second).
In some examples, the information stored within the target oligomer comprises 2 or more bits of information (e.g., 3 bits or more, 4 bits or more, 5 bits or more, 6 bits or more, 7 bits or more, 8 bits or more, 9 bits or more, 10 bits or more, 15 bits or more, 20 bits or more, 25 bits or more, 30 bits or more, 35 bits or more, 40 bits or more, 45 bits or more, 50 bits or more, 60 bits or more, 70 bits or more, 80 bits or more, 90 bits or more, 100 bits or more, 125 bits or more, 150 bits or more, 175 bits or more, 200 bits or more, or 225 bits or more). In some examples, the information stored within the target oligomer can comprise 256 or less bits of information (e.g., 250 bits or less, 225 bits or less, 200 bits or less, 175 bits or less, 150 bits or less, 125 bits or less, 100 bits or less, 90 bits or less. 80 bits or less, 70 bits or less, 60 bits or less, 50 bits or less, 45 bits or less, 40 bits or less, 35 bits or less, 30 bits or less, 25 bits or less, 20 bits or less, 15 bits or less, 10 bits or less, 9 bits or less, 8 bits or less, 7 bits or less, 6 bits or less, or 5 bits or less). The bits of information stored within the target oligomer can range from any of the minimum values described above to any of the maximum values described above. For example, the information stored within the target oligomer can comprise from 2 to 256 bits of information (e.g., from 2 to 129 bits, from 129 to 256 bits, from 2 to 50 bits, from 50 to 100 bits, from 100 to 150 bits, from 150 to 200 bits, from 200 to 256 bits, from 2 to 250 bits, from 2 to 200 bits, from 2 to 150 bits, from 2 to 100 bits, from 2 to 75 bits, from 2 to 32 bits, from 2 to 25 bits, from 2 to 10 bits, from 4 to 256 bits, from 6 to 256 bits, from 8 to 256 bits, from 10 to 256 bits, from 25 to 256 bits, from 32 to 256 bits, from 50 to 256 bits, from 75 to 256 bits, from 100 to 256 bits, 6 to 250 bits, or from 8 to 248 bits).
In some examples, the information stored within the target oligomer is in binary (e.g., base 2), quaternary (e.g., base 4), octal (e.g., base 8), decimal (e.g., base 10), hexadecimal (e.g., base 16), duotrigesimal (e.g., base 32), tetrasexagesimal (e.g., base 64), or a combination thereof.
In some examples, the information stored within the target oligomer comprises a cipher key.
In some examples, the methods can further comprise steganography.
Encryption Methods
Also disclosed herein are methods for encrypting information within a target oligomer. The methods can, for example, comprise: selecting a plurality of unique monomers, wherein each of the unique monomers has a unique mass spectrometry profile; assigning a unique value to each of the unique monomers within the plurality; and synthesizing a target oligomer comprising a self-immolative oligourethane comprising the plurality of unique monomers in a pre-defined order and an endcap, wherein the target oligomer has a first end and a second end, the second end comprising the endcap; wherein the endcap comprises rhodamine B or methyl tyrosine; and wherein the assigned values and the pre-defined order encrypts pre-defined information into the target oligomer.
As used herein "‘a target oligomer” and “the target oligomer” can refer to one or more target oligomers. Accordingly, the methods disclosed herein can comprise encrypting information within one or more target oligomers, by: selecting a plurality of unique monomers, wherein each of the unique monomers has a unique mass spectrometry profile; assigning a unique value to each of the unique monomers within the plurality; and synthesizing one or more target oligomers, each of the one or more target oligomers independently comprising a self-immolative oligourethane comprising the plurality’ of unique monomers in a pre-defined order and an endcap, wherein each of the one or more target oligomers has a first end and a second end, the second end comprising the endcap; wherein the endcap comprises rhodamine B or methyl tyrosine; and wherein the assigned values and the pre-defined order encrypts pre-defined information into the target oligomer.
Synthesis of self-immolative oligourethanes, 5-exo-trig cyclization and elimination thereof, and mass spectrometry analysis and evaluation thereof are described, for example, by Dahlhauser et al. JACS, 2020, 142(6), 2744-2749; Dahlhauser et al. Call Reports Physical Science, 2021, 2, 100393; and Dahlhauser et al. ACS Central Science, 2022, 8, 1125-1133; each of which is incorporated herein by reference for its description thereof.
In some examples, the plurality of unique monomers can comprise 2 or more unique monomers (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more. 175 or more. 200 or more, or 225 or more). In some examples, the plurality of unique monomers can comprise 256 or less unique monomers (e.g., 250 or less, 225 or less, 200 or less, 175 or less, 150 or less, 125 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less. 45 or less, 40 or less, 35 or less, 30 or less, 25 or less. 20 or less, 15 or less, 10 or less. 9 or less. 8 or less. 7 or less. 6 or less, or 5 or less). The number of unique monomers in the plurality of unique monomers can range from any of the minimum values described above to any of the maximum values described above. For example, the plurality of unique monomers can comprise from 2 to 256 unique monomers (e.g., from 2 to 129, from 129 to 256. from 2 to 50, from 50 to 100, from 100 to 150, from 150 to 200, from 200 to 256, from 2 to 250, from 2 to 200, from 2 to 150, from 2 to 100, from 2 to 75, from 2 to 32, from 2 to 25, from 2 to 10, from 4 to 256, from 6 to 256, from 8 to 256, from 10 to 256, from 25 to 256, from 32 to 256, from 50 to 256, from 75 to 256, from 100 to 256, 6 to 250, from 8 to 248, from 8 to 200, from 8 to 100, from 8 to 32, from 8 to 20. or from 8 to 16).
In some examples, the target oligomer has a total length of 2 or more monomers (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more. 125 or more, 150 or more, 175 or more, 200 or more, or 225 or more). In some examples, the target oligomer has a total length of 256 or less monomers (e.g., 250 or less, 225 or less, 200 or less, 175 or less, 150 or less, 125 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less, 20 or less, 15 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, or 5 or less). The number of monomers in the target oligomer can range from any of the minimum values described above to any of the maximum values described above. For example, the target oligomer can have a total length of from 2 to 256 unique monomers (e.g., from 2 to 129. from 129 to 256, from 2 to 50, from 50 to 100, from 100 to 150, from 150 to 200, from 200 to 256, from 2 to 250, from 2 to 200, from 2 to 150, from 2 to 100, from 2 to 75, from 2 to 32, from 2 to 25, from 2 to 10, from 4 to 256, from 6 to 256, from 8 to 256, from 10 to 256, from 25 to 256, from 32 to 256, from 50 to 256, from 75 to 256, from 100 to 256, 6 to 250, from 8 to 248, from 8 to 200. from 8 to 100, from 8 to 32, from 8 to 20, or from 8 to 16).
In some examples, the self-immolative oligourethane is derived from P-amino alcohols.
In some examples, the information stored within the target oligomer comprises 2 or more bits of information (e.g.. 3 bits or more. 4 bits or more, 5 bits or more, 6 bits or more, 7 bits or more, 8 bits or more, 9 bits or more, 10 bits or more, 15 bits or more, 20 bits or more, 25 bits or more, 30 bits or more, 35 bits or more, 40 bits or more, 45 bits or more, 50 bits or more, 60 bits or more, 70 bits or more, 80 bits or more, 90 bits or more, 100 bits or more, 125 bits or more, 150 bits or more. 175 bits or more. 200 bits or more, or 225 bits or more). In some examples, the information stored within the target oligomer can comprise 256 or less bits of information (e.g., 250 bits or less, 225 bits or less, 200 bits or less, 175 bits or less, 150 bits or less, 125 bits or less, 100 bits or less, 90 bits or less, 80 bits or less, 70 bits or less, 60 bits or less, 50 bits or less, 45 bits or less, 40 bits or less. 35 bits or less, 30 bits or less, 25 bits or less, 20 bits or less, 15 bits or less, 10 bits or less, 9 bits or less, 8 bits or less, 7 bits or less, 6 bits or less, or 5 bits or less). The bits of information stored within the target oligomer can range from any of the minimum values described above to any of the maximum values described above. For example, the information stored within the target oligomer can comprise from 2 to 256 bits of information (e.g., from 2 to 129 bits, from 129 to 256 bits, from 2 to 50 bits, from 50 to 100 bits, from 100 to 150 bits, from 150 to 200 bits, from 200 to 256 bits, from 2 to 250 bits, from 2 to 200 bits, from 2 to 150 bits, from 2 to 100 bits, from 2 to 75 bits, from 2 to 32 bits, from 2 to 25 bits, from 2 to 10 bits, from 4 to 256 bits, from 6 to 256 bits, from 8 to 256 bits, from 10 to 256 bits, from 25 to 256 bits, from 32 to 256 bits, from 50 to 256 bits, from 75 to 256 bits, from 100 to 256 bits, 6 to 250 bits, or from 8 to 248 bits). In some examples, the information stored within the target oligomer is hexadecimal based. In some examples, the information stored within the target oligomer comprises a cipher key.
In some examples, the methods can further comprise steganography.
Compositions
Also disclosed herein are compositions for molecular cryptography. The compositions can, for example, comprise a target oligomer comprising a self-immolative oligourethane comprising a plurality of unique monomers in a pre-defined order and an endcap, wherein the target oligomer has a first end and a second end, the second end comprising the endcap; wherein each of the unique monomers has a unique mass spectrometry profile; and wherein the endcap comprises rhodamine B or methyl tyrosine.
As used herein “a target oligomer” and “the target oligomer” can refer to one or more target oligomers. Accordingly, the compositions disclosed herein can comprise one or more target oligomers, each of the one or more target oligomers independently comprising a self- immolative ohgourethane comprising a plurality of unique monomers in a pre-defined order and an endcap, wherein the target oligomer has a first end and a second end, the second end comprising the endcap; wherein each of the unique monomers has a unique mass spectrometry profile; and wherein the endcap comprises rhodamine B or methy l tyrosine.
In some examples, the compositions can further comprise a solvent. The solvent can comprise any suitable solvent. The solvent can, for example, comprise tetrahydrofuran (THF), N-methyl-2-pyrrolidone (NMP), dimethylformamide (DMF), N-methylformamide, formamide, dichloromethane (CH2CI2), ethylene glycol, polyethylene glycol, glycerol, alkane diol, ethanol, methanol, propanol, isopropanol, water, acetonitrile, chloroform, toluene, methyl acetate, ethyl acetate, acetone, hexane, heptane, tetraglyme, propylene carbonate, diglyme, dimethyl sulfoxide (DMSO), dimethoxy ethane, xylene, dimethylacetamide, methylene chloride, hexafluoro-2-propanol, or combinations thereof. In some examples, the solvent comprises acetonitrile.
In some examples, the compositions can further comprise a truncated oligomer comprising a self-immolative oligourethane comprising at least a portion of the plurality of unique monomers and a truncated endcap, wherein the truncated oligomer has a leading end and a trailing end, the trailing end comprising the truncated endcap; wherein the number of monomers in the truncated oligomer is less than that of the target oligomer, the order of monomers in the truncated oligomer differs from the pre-defined order of the target oligomer by 1 monomer or more, or a combination thereof; and wherein the truncated cap comprises an anhydride.
As used herein “a truncated oligomer” and “the truncated oligomer” can refer to one or more truncated oligomers. Accordingly, the compositions disclosed herein can further comprise one or more truncated oligomers, each of the one or more truncated oligomers independently comprising a self-immolative oligourethane comprising at least a portion of the plurality of unique monomers and a truncated endcap, wherein each of the one or more truncated oligomers has a leading end and a trailing end, the trailing end comprising the truncated endcap; wherein the number of monomers in each of the one or more truncated oligomers is less than that of the target oligomer, the order of monomers in each of the one or more truncated oligomer differs from the pre-defined order of the target oligomer by 1 monomer or more, or a combination thereof; and wherein the truncated cap comprises an anhydride.
As used herein a “truncated cap” is a chemical cap that is added to prevent sequences containing deletion products to proceed during the synthesis. During synthesis, monomers are added successively. If a 98% coupling efficiency in each step is assumed, that means 2% of the sequences will not have the proper monomer added in a given position. These 2% of sequences are said to have a deletion. If the synthesis continued without the use of a truncated cap, in the last step, all oligomers would be capped with an end cap, and so end cap labeled oligomers would have mixtures of desired products and deletion products. By using truncated caps, the number of theoretical synthetic by products observed during sequencing can be decreased, and propagation of deletion products can be avoided. Truncated caps can comprise any suitable compound or composition that reacts quantitatively and quickly with any unreacted amines on the growing oligomer end following coupling steps.
Examples of truncated caps include, but are not limited to, acetic anhydride, succinic anhydride, and trichloroacetic anhydride. In some examples, the truncated cap comprises acetic anhydride.
In some examples, the plurality of unique monomers can comprise 2 or more unique monomers (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, or 225 or more). In some examples, the plurality of unique monomers can comprise 256 or less unique monomers (e.g., 250 or less, 225 or less, 200 or less, 175 or less, 150 or less, 125 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less, 20 or less, 15 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, or 5 or less). The number of unique monomers in the plurality7 of unique monomers can range from any of the minimum values described above to any of the maximum values described above. For example, the plurality7 of unique monomers can comprise from 2 to 256 unique monomers (e.g., from 2 to 129, from 129 to 256, from 2 to 50, from 50 to 100, from 100 to 150, from 150 to 200, from 200 to 256, from 2 to 250, from 2 to 200, from 2 to 150, from 2 to 100, from 2 to 75, from 2 to 32, from 2 to 25, from 2 to 10, from 4 to 256, from 6 to 256, from 8 to 256, from 10 to 256, from 25 to 256, from 32 to 256, from 50 to 256, from 75 to 256, from 100 to 256, 6 to 250, from 8 to 248, from 8 to 200, from 8 to 100, from 8 to 32, from 8 to 20, or from 8 to 16).
In some examples, the target oligomer has a total length of 2 or more monomers (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 60 or more, 70 or more, 80 or more. 90 or more, 100 or more. 125 or more. 150 or more. 175 or more, 200 or more, or 225 or more). In some examples, the target oligomer has a total length of 256 or less monomers (e.g., 250 or less, 225 or less, 200 or less, 175 or less, 150 or less, 125 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 45 or less, 40 or less, 35 or less. 30 or less, 25 or less, 20 or less. 15 or less, 10 or less. 9 or less. 8 or less, 7 or less, 6 or less, or 5 or less). The number of monomers in the target oligomer can range from any of the minimum values described above to any of the maximum values described above. For example, the target oligomer can have a total length of from 2 to 256 unique monomers (e.g., from 2 to 129, from 129 to 256, from 2 to 50, from 50 to 100, from 100 to 150, from 150 to 200, from 200 to 256, from 2 to 250, from 2 to 200, from 2 to 150, from 2 to 100, from 2 to 75, from 2 to 32, from 2 to 25, from 2 to 10, from 4 to 256, from 6 to 256, from 8 to 256, from 10 to 256, from 25 to 256, from 32 to 256, from 50 to 256, from 75 to 256, from 100 to 256, 6 to 250, from 8 to 248, from 8 to 200. from 8 to 100, from 8 to 32, from 8 to 20. or from 8 to 16).
In some examples, the self-immolative oligourethane is derived from -amino alcohols.
In some examples, the information stored within the target oligomer comprises 2 or more bits of information (e.g., 3 bits or more, 4 bits or more, 5 bits or more, 6 bits or more, 7 bits or more, 8 bits or more, 9 bits or more, 10 bits or more, 15 bits or more, 20 bits or more, 25 bits or more, 30 bits or more, 35 bits or more, 40 bits or more, 45 bits or more, 50 bits or more, 60 bits or more, 70 bits or more, 80 bits or more, 90 bits or more, 100 bits or more, 125 bits or more, 150 bits or more, 175 bits or more, 200 bits or more, or 225 bits or more). In some examples, the information stored within the target oligomer can comprise 256 or less bits of information (e.g., 250 bits or less, 225 bits or less, 200 bits or less, 175 bits or less, 150 bits or less, 125 bits or less, 100 bits or less, 90 bits or less. 80 bits or less, 70 bits or less, 60 bits or less, 50 bits or less, 45 bits or less, 40 bits or less, 35 bits or less, 30 bits or less, 25 bits or less, 20 bits or less, 15 bits or less, 10 bits or less, 9 bits or less, 8 bits or less, 7 bits or less, 6 bits or less, or 5 bits or less). The bits of information stored within the target oligomer can range from any of the minimum values described above to any of the maximum values described above. For example, the information stored within the target oligomer can comprise from 2 to 256 bits of information (e.g., from 2 to 129 bits, from 129 to 256 bits, from 2 to 50 bits, from 50 to 100 bits, from 100 to 150 bits, from 150 to 200 bits, from 200 to 256 bits, from 2 to 250 bits, from 2 to 200 bits, from 2 to 150 bits, from 2 to 100 bits, from 2 to 75 bits, from 2 to 32 bits, from 2 to 25 bits, from 2 to 10 bits, from 4 to 256 bits, from 6 to 256 bits, from 8 to 256 bits, from 10 to 256 bits, from 25 to 256 bits, from 32 to 256 bits, from 50 to 256 bits, from 75 to 256 bits, from 100 to 256 bits, 6 to 250 bits, or from 8 to 248 bits). In some examples, the information stored within the target oligomer is hexadecimal based.
In some examples, the information stored within the target oligomer comprises a cipher key.
A number of embodiments of the invention have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the invention. Accordingly, other embodiments are within the scope of the following claims.
The examples below are intended to further illustrate certain aspects of the systems and methods described herein, and are not intended to limit the scope of the claims.
EXAMPLES
The following examples are set forth below to illustrate the methods and results according to the disclosed subject matter. These examples are not intended to be inclusive of all aspects of the subject matter disclosed herein, but rather to illustrate representative methods and results. These examples are not intended to exclude equivalents and variations of the present invention which are apparent to one skilled in the art.
Efforts have been made to ensure accuracy with respect to numbers (e.g., amounts, temperature, etc.) but some errors and deviations should be accounted for. Unless indicated otherwise, parts are parts by weight, temperature is in °C or is at ambient temperature, and pressure is at or near atmospheric. There are numerous variations and combinations of measurement conditions, e.g., component concentrations, temperatures, pressures and other measurement ranges and conditions that can be used to optimize the described process.
Example 1 - Sequence Defined Molecules and their Applications
Introduction. Sequence-defined macromolecules are a unique class of monodisperse oligomers or macromolecules that have a defined chain length, where monomer units are placed in a distinct order and position throughout the molecule [1]. Perhaps the most notable examples come from nature, comprising the majority of living systems in the form of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and proteins. These macromolecules have a precisely controlled structure that allowed the evolution of highly sophisticated machinery such a biocatalysis [2], information storage [3], and molecular recognition [4], Even slight variations in the amino acid sequence of proteins can drastically alter both their structure and function [5], With only four nucleotide-monomers, over 20.000 protein-coding genes have been identified and geneticists estimate many more remain unidentified [6], From this enormous pool of genes, over a million proteins are then produced, an absolutely staggering number that is responsible for complex life on Earth. In addition to the wide diversity of proteins that can be made from four simple nucleotide monomers, the speed at which these proteins are sequenced as well as the precise control is unfathomable and acts as a lofty goal for polymer chemistry.
Synthetic Sequence-Defined Macromolecules. With the complexity observed in biopolymers, one must assume that this can be increased with non-natural synthetic sequence-defined polymers. Their molecular composition is significantly broader, with a variety of backbones and pendent functionalities being explored to elicit structural properties [7], In the words of Professor Robert Breslow: "We also take inspiration, but not blueprints, from natural chemistry” [8], To that point, synthetic sequence defined polymers, e.g., ?- peptides, y-peptides. peptoids. polyureas, and poly carbamates [9], have garnered significant interest over recent decades, to the point that their size and structural complexity are nearing those of biopolymers [10], This has been accomplished by increasing the diversity of monomer pools as well as improving on the synthesis to create the sequence-defined polymers (SDPs). One consequence is this complexity’ renders sequence elucidation difficult, at times - impossible. When successful, analysis has relied upon an assortment of 1 and 2D NMR spectroscopy together with sophisticated mass spectrometry techniques [11], Molecular sequencing techniques such as Edman degradation for peptides and Sanger sequencing for DNA are among the most significant chemical achievements of the 20th century. Modem proteomic studies rely on comparisons to databases for protein identification [12], wherein many of the protein sequences were elucidated via Edman degradations. Notably, very few techniques analogous to Edman or Sanger sequencing exist for synthetic macromolecules, likely due to the fact that only recently has synthetic methodology been capable of creating monodisperse macromolecules as structurally complex as biopolymers [13], Peptoids are one exception, with Zuckerman realizing their stepwise chemical degradation on resin [14], As more examples of sequence-defined polymers with controlled primary [13, 15, 16] and secondary [17, 18] structures emerge, new advanced characterization techniques are necessary'.
Controlled chemical degradations, such as the classic Edman degradation, prevailed in the early days of polymer science but have fallen out of favor with the widespread adoption of NMR [19] and mass spectrometry [20, 21], However, if the sequence-defined polymer is designed with degradative sequencing in mind, controlled chain-end depolymerization still offers a powerful approach to primary' structure determination. A limit to this approach is that copolymers seldom depolymerize in a controlled manner [22, 23] as some homopolymers do [24. Triggered and complete depolymerizations, termed self-sequencing polymers, are one solution to this problem [25, 26], In general, removal of a triggering group from the chain end results in a spontaneous cascading elimination or cyclization event that releases all monomers, the kinetics for which depend upon the structure of the polymer backbone and its breakdown mechanism. The first example of such a system utilized a polyurethane backbone, where upon revealing the terminal amine caused a spontaneous 1,6-ehmination and subsequent decarboxylation, revealing the next amine [27], Typically, 1,6- and 1,4-quinone- methide eliminations [28] are much faster than alternating cyclization-eliminations [29] or intramolecular cyclization mechanisms [30], It was recognized that if the self-sequencing could be induced predictably, producing observable intermediates, it could act as a sequencing routine.
Self-Sequencing Oligourethanes. Herein, an example of a kinetically slow depolymerization designed for the characterization of the primary structure of a sequence- defined polymer is described. For the cyclization event that iteratively deconstructs the polymer, a urethane backbone was utilized, wherein a terminal /Talcohol facilitates a favorable 5-exo-trig cyclization that releases a 2-oxazolidinone and anew O-terminus, available for the next cyclization event (Scheme 1). Importantly, as described below, the sequence can be deciphered by a single LC/MS run. This work adds polyurethanes to the very short list of chemically sequenceable abiotic polymers.
Scheme 1. Self-Sequencing Mechanism of Urethanes
As precedent for this concept, random copoly(4-hydroxybutyrate) esters self-sequence via intramolecular cyclization to generate y-lactones, as monitored by NMR spectroscopy [30], It was postulated that urethanes should be capable of a similar intramolecular cyclization. It was reasoned that the use of P-amino alcohols as the monomers would allow the high effective molarity (EM) of the terminal alcohol to be exploited, promoting the sequencing in aqueous-alcohol media without hydrolysis or alcoholysis of the urethane linkages [31], Furthermore, /i-amino alcohols are an inexpensive and diverse pool of chiral monomers.
To promote chain-end depolymerization, a sufficiently basic media and appropriate temperature to induce the intramolecular cyclization via the terminal alcohol was needed. Potassium phosphate (K3PO4) was chosen, due to its low cost, solubility in aqueous/organic mixtures, and minimal footprint in spectroscopic techniques. Further, the p/G of the conjugate acid is 12.3 [32], making it sufficiently basic for reversible deprotonation of the tenninal alcohol, while lacking the basicity to cause an uncontrollable and rapid depolymerization.
To precisely control the temperature for predictable depolymerization, microwave irradiation was utilized over traditional convection heating. Trimer 3 was microwaved in the presence of K3PO4 in a MeOFFELO mixture at 70 °C and sampled for LC/MS at 20 minute intervals (Figure 1 -Figure 2). Each molecule: 3, 2, and 1, could be cleanly observed and quantified by absorbance at 280 nm (near the absorbance maximum of 2,4-dimethoxy benzyl amine) in the chromatogram. All three were defined by their distinctive retention time and absorbance, and characterized by their m/z. The gradual disappearance of trimer 3, appearance and subsequent disappearance of dimer 2, and final existence of only monomer 1, is indicative of a stepwise sequencing. First, upon deprotonation the terminal alcohol of 3 cyclizes intramolecularly with the proximal urethane to form 2-oxazolidinone 4. The newly released terminal alcohol of dimer 2 then cyclizes forming 5. Both 4 and 5 were observed by their characteristic m/z, but they did not have measurable absorbances due to their lack of a chromophore at the diode array detector (DAD) wavelengths (210-800 nm). The urea linkage in 1 was stable to hydrolysis and cyclization by the terminal alcohol.
Having achieved preliminary success, the importance of various reaction conditions were studied. Trimer 3 was subjected to the same basic conditions at 23 °C and monitored by LC/MS. A small amount of initial depolymerization was observed upon addition of the base (-10% of 3 is consumed over the first hour), but no further reaction is seen through 14 days. As another control, trimer 3 was subjected to the 70 °C immolation conditions, sans K3PO4, and monitored by LC/MS. No immolation, appreciable degradation, or hydrolysis was observed after 90 minutes. When the concentration of K3PO4 was increased dramatically to 50 equivalents, the rate of sequencing also increased dramatically. The immolation was complete by 60 minutes and was nearly finished at 40 minutes. This is in comparison to the 240 minutes required to achieve a similar immolation when 5-10 equivalents of K3PO4 is used. A variety of other bases and solvent conditions were tested however K3PO4 was ultimately chosen as it consistently performed well due to its solubility in both organic and aqueous solutions as well as the controlled rate of sequencing observed in the test runs.
After confirming success with controlled intramolecule self-sequence, a pool of unique monomers were synthesized by the reduction [33] and deutero-reduction of commercially available canonical and non-canonical amino acids (Scheme 3 and Scheme 2). By simply deuterating the amino alcohols at the a-methylene, the mass of each monomer increases by 2 AMU, effectively doubling the available monomer pool without changing the complexity of the synthesis or sequencing chemistry (a feature that was taken advantage of when encoding in hexadecimal, see below). The monomers were then converted to the activated carbonate by reaction with 4-nitrophenylchloroformate (Scheme 4). Having developed a facile workflow for rapidly creating a large pool of monomers, it was felt that this chemi st ry (both the synthesis and sequencing experiments) would be highly amenable to being conducted in parallel and applicable in the field of information storage.
Scheme 2. Generalized synthesis for the deuter-reduction of Fmoc-amino acids.
Scheme 3. Generalized synthesis for the reduction of commercial Fmoc-amino acids.
Scheme 4. Generalized synthesis for the activated carbonates.
Sequence-Defined Polymers for Information Storage. Sequence-defined polymers show promise for biomimetics, self-assembly, catalysis, and information storage, where in the primary structure begets complex chemical processes. Herein, the solution-phase and the high-yielding solid-phase syntheses of discrete oligourethanes and methods for their selfsequencing are reported, resulting in rapid and robust characterization of this class of oligomers and polymers, without the use of MS/MS. Crucial to the sequencing is the inherent reactivity of the terminal alcohol to ‘'unzip” the oligomers, in a controlled and iterative fashion, releasing each monomer as a 2-oxazolidinone. By monitoring the self-immolation reaction via LC/MS, an applied algorithm rapidly produces the sequence of the oligourethane. Not only does this process provide characterization of structurally complex molecules, but it also works as a reader of molecular information.
Sequence-defined polymers have shown potential as dense and durable information storage media [34], These macromolecules, often referred to as digital polymers, store information at the molecular level in the form of a defined, and absolute monomer sequence (i. e. , primary structure) [35-37], Encoding information at the molecular level can be used to surmount some drawbacks of conventional storage devices, such as durability, longevity, and excessive spatial occupation [34], Such macromolecular information storage is now well established with artificial DNA biopolymers, which have been shown to store and retrieve significant amounts of information [38-41], Two strengths of using DNA for information storage are the ability to replicate and retrieve data [42], as well as the ability to exploit rapid advances in sequencing, such as Next-Gen methods [39, 43] and nanopore technology [44, 45], Alongside DNA, advances in the synthesis and sequencing of abiotic sequence-defined polymers has improved the information storage capabilities of these macromolecular systems [46-60], with commensurate advances toward this goal seen with multicomponent reactions [1, 61, 62] and small molecule strategies [63-65], However, significant advances are still needed to rival the effective storage capacity of nucleic acids using sequence-defined polymers, let alone silicon-based data storage [55. 66],
Typically, only small pieces of digital information (e g., a few bytes, a word) can be stored in a single abiotic sequence-defined polymer due to limitations in the ability to synthesize long sequences and challenges in decoding the primary structure of these macromolecules. Frequently, the sequencing/read-out of the encoded information proves to be the most difficult aspect of molecular information storage [66], This challenge in sequence deconvolution is directly at odds with the parameters required to increase the storage capacity of digital polymers, namely: 1) chain length, and 2) number of unique monomers (bits/monomer, i.e., information density). As one increases the chain length and monomer pool, the capabilities of a given sequencing platform to correctly deconvolute the sequence information can be quickly reached. That being said, chain lengths and storage capacities in single molecules are continually increasing (128 bits in an aperiodic copolyester and 144 bits in a poly(phosphodiester)) [50, 57],
The most common sequencing methodologies for abiotic sequence-defined polymers utilize tandem mass spectrometry (MS/MS) [47-49, 52-54, 58, 59, 67], or more recently pseudo MS3 sequencing [57, 68], and mutated nanopores [69, 70], In each of these cases, subtle, or even dramatic, variations in the monomers become difficult to elucidate, and as a result, examples of multiply functionalized macromolecules for data storage are limited [1, 51, 53, 56, 57, 71],
Encoding Information Into Oligomers
Sanity Testing with “Hello, World!” To develop the self-sequencing urethanes as a medium for information storage, it was sought to first encode a small piece of information, in this case text, into a compressed bit-string that could then be translated into a molecular form, and ultimately read back. The first program written by aspiring programmers is often the well-established sanity' test “Hello, World!” [72], This is a historic program, dating back to its use in Bell Laboratories as w ell as The C Programming Language in 1978 [73], Sanity tests such as these are commonly used to test if systems are working before starting up more complicated processes, very analogous to scaling up chemical reactions and workflows. Another example of a commonly used sanity test in computer science from the textbook My computer likes me when I speak BASIC is '‘MY HUMAN UNDERSTANDS ME’’ [74], however it was felt that ‘'Hello, Word!” would sufficiently test this platform as it also contains punctuation.
Huffman Encoding and Decoding. Huffman coding [75], a form of lossless data compression, was utilized to convert “Hello, World!” to binary (Figure 3). This encryption method takes the phrase to be encoded (in this case '‘Hello, World!”) and first breaks it into its individual symbols (as seen in Figure 3). This is broken down regardless of what the symbol is, making it amenable for more than just phrases written in English. It then counts the occurrences in the selected phrases and creates a binary tree assigning the depth of the node containing the symbol based on the number of occurrences. This way the more commonly occurring symbols, in the case of “Hello, World!” the letters “1” and “o”, appear closer to the top of the binary tree than the others. This way Huffman Coding, like other entropy encoding methods, represents the more common symbols (in this case letters and punctuation) with fewer bits relative to the less common symbols. Thus, encoding “Hello, World!” required only 42 bits, as opposed to the 104 bits required when using a traditional ASCII table. From here, the bit string can be converted to any numerical base-system, such as octal or hexadecimal, as a means to further compact the information (Figure 4) [76], This is done by basic binary addition and then conversion to the desired base being used (base 8 for octal and base 16 for hexadecimal).
Encoding in the molecules. For this rather short bit string, it was chosen to convert the information to an octal notation, which will then be represented molecularly with eight unique monomers (Figure 5). Using the standard ASCII conversion, each octal character represents three bits (Figure 4). Thus, the 42-character bit string shrinks to a 14-character octal string “57451243036731'5 Encoding this octal string at a molecular level requires that each symbol be represented by a discrete monomer along the oligomer backbone. Thus, a pool of eight unique monomers was synthesized by the aforementioned reduction and activation steps. The amino alcohols used to encode the phrase are shown in Figure 5. Monomers were assigned their octal symbol according to their occurrence, such that cheaper monomers (e.g. alaninol) were assigned to the most frequently occurring symbols.
The fourteen-digit octal string was then written onto two oligomers (Al and A2, Figure 6), each carrying seven of the fourteen digits. Previous examples have used mass tags to indicate the position of letters in words [68], or the position of words in sentences [53] to allow reconstruction of information in the correct order. The physical location of the molecules in a 96-well plate was used to index the information, similar to how a mechanical hard disk drive uses a physical location and a directory to store a computer’s data [76],
Testing High-Throughput Oligomer-Synthesis. The two oligomers Al and A2 (read from N-0 termini: 4215475 -Pheindex, and 1376303-Pheindex, respectively) were successfully synthesized on the solid-phase, in parallel, on a fritted 96-well plate, using a previously described methodology, with slight modifications to adapt it a well plate format (Figure 5 and Scheme 5) [77], On each oligomer, the encoded information is preceded with a phenylalaninol as an indexing tool (Pheindex) to begin reading the mass spectra. This allows the same amino alcohol loaded resin to be utilized for every oligomer for consistent syntheses, while simultaneously providing a reading frame for the mass spectra. A key feature of this method is to use all of the oligomers directly from the solid-phase resin without any purification. As such, extraneous masses are present in the spectra and could possibly create reading errors. By implementing a reading frame with which to start calculating mass differences, the extraneous masses can be filtered out (Figure 8. for example). After the successful coupling of each monomer to the desired octamer, the N- terminus was “capped” 4-fluoro-7-nitrobenzofurazan (NBD-F) to act as a long-wav elength chromophore with which to monitor the self-immolation. From here, the oligomers are cleaved from the resin and directly sequenced.
Scheme 5. Exemplary scheme for each oligomer synthesized in parallel. Stepwise oligomerization enroute to octamer Al on a 2-chlorotrityl chloride polystyrene resin. Decoding the molecules. Sequencing of the oligomers is adapted from a previously established protocol [77], NBD-Labelled oligomers Al and A2 were sequenced in parallel in a well plate in DMSO with CS2CO3 at 70 °C and sampled for LC/MS at designated intervals (Figure 8). As Scheme 1 shows, immolation removes each monomer from the O-terminus in a controlled fashion, thus truncating the oligomers iteratively. The parent masses of each iteration (8mer, 7mer, 6mer, 5mer, 4mer, 3mer, 2mer, and Imer) were observed in the mass spectra (a spectra after 225 minutes of immolation is shown in Figure 8) and entered into a templated spreadsheet that is fed into the Mol.E-Coder decoding algorithm (https://github.com/PhysicalOrganic/Mol.E-Coder). The mass differences between subsequent peaks were calculated by the algorithm and correlated to each monomer, which are associated with a specific octal symbol (Figure 5 and Figure 7). The octal string is sorted according to the indices and converted back to binary by the same ASCII conversion. The bit string is finally decoded back through the Huffman algorithm to produce the original information wholly intact with no errors (Figure 7). The script is able to decode this information in less than a second.
It is important to note that from this sanity test, it was found that the workflow from information encoding, to synthesis, to sequencing, and finally decoding, integrated together seamlessly. It was observed that all eight monomers can be readily distinguished from one other, including the deuterated variant. Thus, mass differences between monomers can be as small as 2 AMU. The oligomers were sequenced without purification, meaning the noise accumulated after seven coupling and seven deprotection steps, one labelling step, and cleavage from the solid-phase, was not significant. Further, by utilizing the phenylalaninol reading frame, the 10% deletion observed in Al (presumed to be from incomplete deprotection after the first coupling step) was inconsequential. Likewise, the overlapping peaks in the chromatograms of A2 were easily deconvolved.
Encoding Jane Austen in Hexadecimal. With the success of this test, it was sought to scale the information storage to explore its capacity and efficacy to store more remarkable information. With seemingly infinite information to choose from, an apt but timeless quote from Jane Austen’s Mansfield Park was chosen: “If one scheme of happiness fails, human nature turns to another; if the first calculation is w rong, we make a second better: we find comfort somewhere.” The information, which is 153 characters long including spaces, was converted to a 632-character bit string by the optimized Huffman encoding, and converted to hexadecimal by the same standard ASCII conversion. Remarkably, the resulting hexadecimal code was 158 characters long, only 5 more than original information that contained 26 unique characters of the English language.
To write molecules in the hexadecimal numeral sy stem, the monomer pool was increased to sixteen by adding one more unique amino alcohol, and deuterating each one at the alpha methylene (Scheme 2, Figure 19). Having now sixteen unique monomers and the hex code, an encoding scheme w as devised. With molecular information density being the goal, it was chosen to extend the oligomers to ten in length. Thus, each oligomer would contain nine hex symbols and the index, resulting in seventeen 10-mers and a final 6-mer to encode all 158 characters. Figure 21 shows pictorially all of the information-strings for the concept: oligourethanes, hex strings, bit strings, and the English language translation.
The parallel synthesis was performed as described earlier, by iterative coupling and deprotection in a fritted 96-well plate (oligomers G1-G12, and H1-H6, named after the wells in the plate Figure 22-Figure 39). All 18 oligomers were synthesized in parallel. Through the course of the 18 simultaneous syntheses, 17 of the 18 were successful with only minor deletions (discovered to be inadequate mixing of the wells during deprotection). Oligomer H6 was found to have suffered a significant deletion, and as such, was resynthesized. With adequate mixing during the deprotection, no such deletions were again observed. All 18 oligomers were then labelled with NBD-F simultaneously, and cleaved from the solid-phase resin with 0.1% trifluoroacetic acid in dichloromethane for 5 minutes. The extended cleavage resulted in up to 50% trifluoroacetylation of the terminal alcohol. The trifluoroacetyl ester is readily hydrolyzed under the sequencing conditions, and as such, the material can be carried through in whole w ithout impacting the sequencing. As a control, but not for sequencing purposes, all 18 oligomers were confirmed by high resolution mass spectroscopy.
Oligomers G1-G12 and H1-H6 were sequenced concurrently via self-sequencing in a 96-well plate in a 2: 1 MeOH:Water mixture wdth K3PO4 at 70 °C in a heated shaker (Figure 22-Figure 39). The reactions were sampled for LC/MS every 30 minutes for 2.5 hours. Of the 176 masses (seventeen lOmers, one 6mer), 170 were observed clearly and distinctly. The parent 1 Omers for G8, Gil, H2. and H5 did not ionize in significant quantities under the generalized low-resolution LC/MS conditions. Considering these w ere indexing masses, no encoded information w as lost. Likewise, the 9mers for G8 and H2 did not ionize significantly under the generalized low-resolution LC/MS conditions. Due to the chromatographic traces at 470 nm, the presence of the 1 Omers and 9mers in the samples was clearly observed, making it trivial to pinpoint the specific information that was missing. Knowing what information was missing, this w as easily addressed by modifying the conditions and using a high-resolution instrument.
Considering future improvement of this sequencing methodology, some noteworthy trends regarding the effects of certain sidechains on the sequencing reaction should be mentioned. Firstly, oligomers in which valinol was the terminal monomer did not accumulate in significant quantities relative to other oligomers. Presumably, the isopropyl sidechain of the valine derived amino alcohol enforces greater steric compression by having a gem- dimethyl methine beta to the alcohol, rather than the methylene or methyl seen in each of the other sidechains. Steric compression is a significant driving force in the 5-exo-trig cyclization, resulting in an enhanced rate of cyclization [78], Thus, oligomers with terminal valinofs have a shorter lifetime in solution. Secondly, as expected, longer oligomers tend to have smaller changes in retention time per cyclization event, which can result in peak coelution. This is prominent when the terminal or cyclizing monomer is alaninol, likely due to its small sidechain having little effect on the overall polarity of the macromolecule. Since the masses are observed, this is not seen as a detriment to the sequencing, but worthy of note for future designs.
Each of the masses were entered into a templated spreadsheet and fed into the Mol.E-
Coder algorithm (described in greater detail below' and in Table 1).
Table 1. The recorded masses observed from oligomers G1-G12 and H1-H6 were entered into the templated spreadsheet.
Conducting A Blind Study. The algorithm assigned the hex symbol to the mass differences and sorted the hex string. The hex string was converted back to binary via ASCII, and decoded using the Huffman algorithm, returning the Jane Austen quote with no errors. To truly explore the robustness of the sequencing methodology, a blind study was performed, where a colleague unaffiliated with the project was given a set of instructions to “read” the molecular media (Figure 40). After one pass, the participant was able to correctly decipher 156 of the 158 information containing monomers. The two incorrect masses were attributed to a reading error when interpreting the mass spectra (a deletion was incorrectly chosen, see Figure 41). The participant was given a second set of instructions and was able to correctly decipher all 158 information encoding monomers (Figure 42). The success of the blind study lends promise to the robustness of the sequencing platform, as well as the plans to automate the decoding of mass spectrometry data.
Conclusions. Herein, a methodology for the encoding of information in sequence- defined macromolecular oligourethanes was described. The Mol.E-Coder algorithm is able to take presumably any information, regardless of the language or symbol table it is comprised of and compress it to a universally recognized hex string. The “writing” process is incredibly robust, efficient, and amenable to high throughput. 18 oligomers were synthesized in parallel, demonstrating the ease with which this could be scaled to full 96-well plates, specifically with increased automation. The encoding scheme utilizes a cheap feedstock of chemicals, being derived from commercially available amino acids. Further, the pool of monomers was effectively doubled by a simple deutero-enrichment. The information is read utilizing a very simple sequencing process, requiring only base, heat, and a single quadrupole LC/MS. The mass spectrometry data is fed into the Mol.E-Coder algorithm, which then assigns the hex symbol to each monomer, and decodes the information back to the original text.
Information storage aside, demonstrated herein is the high throughput synthesis of chiral abiotic oligomers, as well as the high throughput sequencing and characterization of complex macromolecules. Future research will look to scale these molecules for more advanced information storage systems, as well as utilize the rapid characterization techniques presented here to explore their applications as sequence-defined polymers; i.e. self-assembly, combinatorial chemistry, and catalysis.
Mol.E-Coder Program. Mol.E-Coder works to convert text into both octal and hexadecimal in the same way. It is a two-part program that encodes text input into monomer code and then takes LC/MS data and converts masses back to the original text that was encoded.
Both octal and hexadecimal work in the same way and are separated into an encode and decode portion.
Mol.E-Coder.py The first part of Mol.E-Coder is Mol.E-Coder{Hex/Octal].py. This program takes in one argument/input, a Text Document (.txt). This text document should contain the phrase the user would like encrypted. The program then parses the document and creates a. Dictionary data structure with the key being the symbol and the value being the frequency with which the symbol was used.
This dictionary is then passed to a function named encode that creates the binary' tree used to assign symbols their binary' representation. This function uses a Heap data structure where every parent node has a value less than or equal to any of its children. The dictionary values are initially converted into a 3D list, the list being comprised of [frequency, [symbol, binary representation]] which is subsequently converted into a heap using the python function heapify. The binary' code list space is left initially empty' to be populated by this function. After converting the information about the document into a heap, the data is then entered into a while loop function. While the heap is larger than 1. it will keep iterating through the heap. The heap is popped and the smallest value (lowest frequency) is returned as "Io" and the second smallest value (second lowest frequency) is returned as “hi”.
Following this visual representation of “Hello. World!”, the less frequently used symbols “H”, “e”, “W”, “r”. “d”, “!”, “,” and “ ” are all only used one time making their assignment completely arbitrary. They are assigned 0's and 1 ’s until their combined frequency reaches that of “1” (3) and “o” (2). Then “1” and “o” are assigned 0’s and l’s but fewer digits as it took longer for the binary' tree to reach them as they were higher up the tree due to their higher frequency. This is how more frequently used symbols are given shorter binary representation.
Because this simply returns a dictionary of the symbols and their binaryrepresentation ([T, '01'], ['o’, ’00'], ['r', TOO'], ['!', '1010'], [',', '1011'], ['H', '1100'], ['W', '1101'], ['d'. '1110'], ['e', T i l l']), it is necessary to parse through the data again to create the final bitstring.
The program then takes the binary string and breaks it up into units of 8 for hexadecimal and 4 for octal. Additionally, it is necessary to “pad” the binary code to ensure that the bitstring is broken up evenly. This means a certain number of 0’s are added to the end of the bitstring to ensure the string is evenly divisible by 8 (hex) or 4 (octal). This number will vary based on the length of the bitstring (e.g. for hexadecimal if the bitstring was already 7 digits long a single 0 would be added but if it were only 5 digits long, three 0’s would be added). This information is stored in the Huffman output file as well as the Huffman codes.
Lastly, the padded bitstring is then broken into sections of 8, assigned a hexadecimal value, and iteratively removed from the string until the bitstring is empty.
The program then outputs this bitstring as well as the Huffman tree into excel documents for the researcher. The output files are named CharactersToHuffmanCodes.xlsx and OutputCodes.xlsx. CharactersToHuffmanCodes.xlsx gives the binary representation of each character in the encoded text. OutputCodes.xlsx gives the text that is encoded in its hexadecimal, compressed representation.
To run this program the program takes a single input: a text document containing the text to be encoded.
An example input and output for the hexadecimal version is shown in Figure 100.
MoLE-Decoder Usage. The second part of Mol.E-Coder is Mol.E- Decoder{Hex/Octal}.py. This program takes in two inputs, the codes output by Mol.E- Coder.py in the Huffman algorithm document that are correlated to characters in the encoded text "CharactersToHuffmanCodes.xlsx" and the LC/MS masses entered into a template "DataTemplate.xlsx".
The program first takes in the masses entered into the template and matches the difference between the parent mass and the subsequent mass to a monomer.
After determined the parent mass (value 1) and next mass in the series (value 2), the values are then assigned a hexadecimal representation using another function.
If the mass is not found, the error message “NOT MATCHED” is printed. The monomer matching is given a range of +/- 1.5 amu’s to prevent overlap with the deuterated monomers. After a hexadecimal value is (hopefully) assigned to the difference between the two values, the hex value is then further examined.
If the value corresponds to a phenylalanine derivative and is the first monomer in the sequence then it is not added to the bitstring is this is simply the index monomer. Additionally, if the hex symbol is 0 it is necessary to indicate how many 0’s this corresponds to (4 for hex, 3 for octal). Next, the hex symbol is then converted back into binary.
This will then return the original binary string from the last program if the masses entered were correct. Lastly, the binary' string is converted back into the original phrase by matching this binary string with the Huffman codes (found in CharactersToHuffmanCodes.xlsx) and removing the added padding.
An example input and output for the hexadecimal version is shown in Figure 101. General Procedures for Jane Austen and Hello, World Encoding
Materials and Instrumentation. All materials used in the synthesis of each compound and related tests, were purchased from Sigma- Aldrich Chemical Co., Acros Organics, Tokyo Chemical Industry. Chem Impex International, etc. and used without further purification. Solvents (DCM, NMP, chloroform, DMSO. DMF, MeOH, MeCN. Isopropanol) were of reagent grade or HPLC grade quality and purchased from Fischer Scientific. NMR solvents (CDCls, DMSO-d6) were purchased from Cambridge Isotope Laboratories.
Column chromatography was performed using silica gel 60 (230 ± 400 mesh. 0.040 ± 0.063 mm) from Dynamic Adsorbents.
TLC analyses were carried out using Silica TLC Plates Aluminium Backing 20 by 20 cm sheet UV active at 254 nm.
'H and 13C spectra were recorded on Varian DirectDrive or Varian INOVA 400 MHz NMR spectrometers. The NMR spectra were referenced to solvent and the spectroscopic solvents were purchased from Cambridge Isotope Laboratories.
Liquid Chromatography /Mass spectra were recorded on an Agilent Technologies 6120 Single Quadropole or 6125B Single Quadrapole interfaced with an Agilent 1200 series liquid chromatography system equipped with a diode-array detector. Column: Agilent ZORBAX Eclipse Plus C18 narrow bore column; 2.1 mm internal diameter; 50 mm length; 5 micron particle size; P.N. 959746-902. Resulting spectra were analyzed using Agilent LC/MSD ChemStation. Separations achieved with MeOH/Water w/ or w/o 0.1% FA or MeCN/Water 5-95% w/ or w/o 0.1% FA, gradient elution. For oligomers MeCN/Water 5- 95% w/ 50mM ammonium acetate gradient elution provided the strongest ionization signals in the negative mode. High resolution mass spec was performed by the UT-Austin Mass Spectrometry Facility using Agilent Technologies 6530 Accurate Mass Q-TOF LC/MS system.
Parallel Synthesis was performed in Porvair Sciences Combinatorial Microlute 10 uM fritted deep-well 96-well plate catalogue number 240054.
Parallel Sequencing was performed in Nunc polypropylene DeepWell 96-well plates. Incubation at 70 °C was achieved using Ika KS 3000 i control incubator shaker.
Synthesis and Characterization
General procedure for reduction of commercial Fmoc Amino Acids. Procedure for the reduction of Fmoc amino acids was adapted from previous work [77],
To a stirred solution of Fmoc-L-amino acid 1.0 equivalent in anhydrous THF (3.3 mL) was added N,N-carbonyldiimidazole (1.33 equivalent) at room temperature. The reaction stirred for at least 10 minutes and was then cooled to 0 °C. Next, a solution of NaBEL (1.66 equivalent) in H2O (1.66 mL, or 0.6M) was added. The solution was stirred for at least 30 minutes, up to 1.5 hours. The reaction was quenched by addition of IM HC1 and extracted with EtOAc (3x). The combined organics were washed lx with brine, dried over Na2SO4, and concentrated under vacuum.
(9H-fluoren-9-yl)methyl(S)-(l-hydroxybutan-2-yl)carbamate (Fmoc-Abu-ol) .
According to the general procedure for the reduction of Fmoc-amino acids: L-2-(Fmoc- amino)butyric acid (2.0 g, 6.15 mmol) gave a crude product, identified by LC/MS, that was carried through to next step without prior purification. MS; ESI+ m/z 312.2 (M+H)+, calculated (C19H21NO3): 311.16
. (9H-fluoren-9-yl)methyl(S)-( 1 -hydroxy-4-phenylbutan-2-yl)carbamate (Fmoc- HoPhe-ol). According to the general procedure for the reduction of Fmoc-amino acids: Fmoc-L-homophenylalamne (2.0 g, 4.982 mmol) gave a crude product that w as then filtered through a plug of silica gel (2: 1 EtOAc:Hex) and concentrated as the crude product. The product was identified by LC/MS and carried through the next step without isolation. MS; ESI+ m/z 388.2 (M+H)+, calculated (C25H25NO3): 388.19.
(9H-fluoren-9-yl)methyl(S)-(l-(4-chlorophenyl)-3-hydroxypropan-2-yl)carbamate (Fmoc-4Cl-Phe-ol) . According to the general procedure for the reduction of Fmoc-amino acids: Fmoc-4-chloro-L-phenylalanine (2.0 g, 4.741 mmol) gave a crude product that w as then filtered through a plug of silica gel (EtOAc) and concentrated as the crude product. The product was identified by LC/MS and carried through the next step without isolation. MS; ESI+ m/z 408.2 (M+H)+, calculated (C24H22CINO3): 408.14.
(9H-fluoren-9-yl)methyl(S)-(l-cyclohexyl-3-hydroxypropan-2-yl)carbamate (Fmoc- CHA-ol). According to the general procedure for the reduction of Fmoc-amino acids: Fmoc- L-cyclohexylalanine (2.0 g, 5.082 mmol) gave a crude product that was then filtered through a plug of sihca gel (EtOAc) and concentrated as the crude product. The product was identified by LC/MS and carried through the next step without isolation. MS; ESI+ m/z 380.2 (M+H)+, calculated (C24H29NO3): 380.22.
Generalized synthesis for the deutero-reduction of Fmoc-amino acids. Procedure for the deuteron-reduction of Fmoc amino acids was adapted from previous work [77],
To a stirred solution of Fmoc-L-amino acid 1.0 equivalent in anhydrous THF (3.3 mL) was added N,N-carbonyldiimidazole (1.33 equivalent) at room temperature. The reaction stirred for at least 10 minutes and was then cooled to 0 °C. Next, a solution of NaBD4 (1.66 equivalent) in D2O (1.66 rnL, or 0.6M) was added. The solution was stirred for at least 30 minutes, up to 1.5 hours. The reaction was quenched by addition of IM HC1 and extracted with EtOAc (3x). The combined organics were washed lx with brine, dried over Na2SO4, and concentrated under vacuum.
(9H-fluoren-9-yl)methyl (S)-( l-hydroxypropan-2-yl-l, 1-d 2) carbamate (Fmoc-D Ala- ol). According to the general procedure for the deutero-reduction of Fmoc-amino acids: Fmoc-Alanine (2.0 g, 6.42 mmol) gave a crude product, identified by LC/MS, that was filtered through a silica plug with 1: 1 EtOAc:Hex, and carried through to the next step without isolation. MS; ESI+ m/z 300.2 (M+H)+, calculated (C18H17D2NO3): 300.16.
(9H-fluoren-9-yl)methyl(S)-(l-hydroxy-3-phenylpropan-2-yl-l,l-d2)carbamate (Fmoc-D2Phe-ol) . According to the general procedure for the deutero-reduction of Fmoc- amino acids: Fmoc-Z-Phenylalanine (2.0 g, 5.285 mmol) gave a crude product, identified by LC/MS, that was filtered through a silica plug with 1: 1 EtOAc:Hex, and carried through to the next step without isolation. MS; ESI+ m/z 376.2 (M+H)+, calculated (C24H21D2NO3): 376.19.
(9H-fluoren-9-yl)methyl(S)-( 1 -hydroxy-4-methylpentan-2-yl-l, l-d2) carbamate (Fmoc-D2Leu-ol) . According to the general procedure for the deutero-reduction of Fmoc- amino acids: Fmoc-L-Leucine (1.86 g, 5.285 mmol) gave a crude product, identified by LC/MS, that was filtered through a silica plug with 1 : 1 EtOAc:Hex, and carried through to the next step without isolation. MS; ESI+ m/z 342.2 (M+H)+, calculated (C21H23D2NO3): 342.21.
(9H-fluoren-9-yl)methyl (S)-(l-hydroxy-3-methylbutan-2-yl-l. l-d2)carbamate (Fmoc- D2Val-ol). According to the general procedure for the deutero-reduction of Fmoc-amino acids: Fmoc-L-Valine (1.8 g, 5.3 mmol) gave a crude product that was filtered through a silica plug in 2: 1 EtOAc:Hex and carried through to the next step without isolation. MS; ESI+ m/z 328.2 (M+H)+, calculated (C20H21D2NO3): 328.19.
(9H-fluoren-9-yl)methyl (S)-(l-hydroxybutan-2-yl-l, l-d2)carbamate (Fmoc-L PAbu- ol). According to the general procedure for the deutero-reduction of Fmoc-amino acids: L-2- (Fmoc-amino)butyric acid (1.73 g, 5.3 mmol) gave a crude product, identified by LC/MS, that was filtered through a silica plug with 1 : 1 EtOAc:Hex and carried through to the next step without isolation. MS; ESI+ m/z 314.2 (M+H)+, calculated (C19H19D2NO3): 314.17.
(9H-fluoren-9-yl)methyl (S)-(l -hydroxy-4-phenylbutan-2-yl-l, l-d2)carbamate (Fmoc- D2HoPhe-ol) . According to the general procedure for the deutero-reduction of Fmoc-amino acids: Fmoc-L-homophenylalanine (2.12 g, 5.3 mmol) gave a crude product, identified by LC/MS, that was filtered through a silica plug with 1 : 1 EtOAc:Hex, and carried through to the next step without isolation. MS; ESI+ m/z 390.2 (M+H)+, calculated (C25H23D2NO3): 390.21.
(9H-fluoren-9-yl)methyl (S)-( l-( 4-chlorophenyl)-3-hydroxypropan-2-yl-3.3- d2)carbamate (Fmoc-D2-4Cl-Phe-ol) . According to the general procedure for the deuteroreduction of Fmoc-amino acids: Fmoc-4-chloro-L-phenylalanine (2.23 g, 5.3 mmol) gave a crude product, identified by LC/MS. that was filtered through a silica plug with 2:3 EtOAc:Hex, and earned through to the next step without isolation. MS; ESI+ m/z 410. 1 (M+H)+, calculated (C24H20D2CINO3): 410.15.
(9H-fluoren-9-yl)methyl (S)-( 1 -cyclohexyl-3-hydroxypropan-2-yl-3, 3-d2) carbamate (Fmoc-D2CPL4-ol) . According to the general procedure for the deutero-reduction of Fmoc- amino acids: Fmoc-L-cyclohexylalanine (2.08 g, 5.3 mmol) gave a crude product, identified by LC/MS, that was filtered through a silica plug with 1 :2 EtOAc:Hex, and carried through to the next step without isolation. MS; ESI+ m/z 382.2 (M+H)+, calculated (C24H27D2NO3): 382.24.
Generalized synthesis for the activated carbonates. General procedure for the synthesis of activated carbonates was adapted from previous work [77],
To a stirring solution of Fmoc- Amino alcohol (1.0 equivalent) in anhydrous DCM (0.2M) was added pyridine (1.3 equivalents) dropwise. Next, 4-nitrophenyl chloroformate (1.5 equiv) was added, and the reaction left to stir overnight. Reaction was monitored by TLC and upon consumption of the starting material, was diluted excessively in DCM and transferred to a separatory funnel. The organic layer was washed with IM NaHSCL (3x), then IM Na2CC>3 (5x, or until it stopped turning bright yellow), and finally brine. The organic layer was dried over Na2SO4, filtered, and concentrated in vacuo. The product was purified by silica gel chromatography.
(9H-fluoren-9-yl)methyl (S)-( !-((( 4-nitrophenoxy)carbonyl)oxy)propan-2- yl)carbamate (Fmoc-Ala-PNOC) . According to the general procedure for synthesis of activated carbonates: Commercial Fmoc-L-Ala-ol (1.0g, 3.36 mmol) gave a crude product that was purified by silica gel chromatography (DCM then 2: 1 Hex:EtOAc) and isolated as a white solid (1.13 g, 72%), with matching XH NMR, 13C NMR, and mass spectrum as reported in the literature [37], (9H-fluoren-9-yl)methyl (S)-(4-methyl-l-(((4-nitrophenoxy)carbonyl)oxy)pentan-2- yl)carbamate (Fmoc-Leu-PNOC). According to the general procedure for synthesis of activated carbonates: Commercial Fmoc-Z-Leu-ol (1.0 g, 2.95 mmol) gave a crude product that was purified by silica gel chromatography (DCM then 2: 1 Hex:EtOAc) and isolated as a viscous glassy oil (1.24 g, 83% yield), with matching NMR. 13C NMR, and mass spectrum as reported in the literature [37],
(9H-fluoren-9-yl)methyl (S)-(l-(((4-nitrophenoxy)carbonyl)oxy)-3-phenylpropan-2- yl)carbamate (Fmoc-Phe-PNOC) . According to the general procedure for synthesis of activated carbonates: Commercial Fmoc-Z-Phe-ol (1.0 g, 2.68 mmol) gave a crude product that was purified by silica gel chromatography (DCM then 2: 1 Hex:EtOAc) and isolated as a white/yellow solid (1.03 g, 71% yield), with matching 'H NMR. 13C NMR, and mass spectrum as reported in the literature [37],
(9H-fluoren-9-yl)methyl (S)-(3-methyl-l-(((4-nitrophenoxy)carbonyl)oxy)butan-2- yl)carbamate (Fmoc-Val-PNOC). According to the general procedure for synthesis of activated carbonates: Commercial Fmoc-Z-Val-ol (2.0 g, 6. 146 mmol) gave a crude product that was purified by silica gel chromatography (DCM then 4: 1 Hex:EtOAc) and isolated as a white solid (2.5 g, 85% yield over 2 steps), with matching 'H NMR, 13C NMR. and mass spectrum as reported in the literature [37],
(9H-fluoren-9-yl)methyl (S)-(l-cyclohexyl-3-(((4-nitrophenoxy)carbonyl)oxy)propan- 2-yl)carbamate (Fmoc-CHA-PNOC) . According to the general procedure for synthesis of activated carbonates: Crude Fmoc-CHA-ol (~5 mmol) gave a crude product that was purified by silica gel chromatography (DCM then 9: 1 Hex:EtOAc) and isolated as a fluffy white solid (1.96g, 70% yield over 2 steps). 'H NMR (500 MHz, Chloroform- ) 5 8.24 (d, J= 8.8 Hz, 2H), 7.77 (d, J= 7.5 Hz, 2H), 7.59 (d, J= 7.4 Hz, 2H), 7.44 - 7.33 (m, 4H), 7.33 - 7.28 (m, 2H), 4.74 (d, J= 8.6 Hz, 1H), 4.50 - 4.39 (m, 2H), 4.37 - 4.30 (m, 1H), 4.23 (t, J= 7.0 Hz, 1H), 4.21 - 4.10 (m, 2H), 1.82 (d, J= 13.0 Hz, 1H). 1.76 - 1.64 (m, 4H), 1.49 - 1.34 (m, 3H), 1.31 - 1.10 (m, 3H), 1.03 - 0.85 (m, 2H). 13C NMR (126 MHz, Chloroform- ) 6 155.95, 155.44, 152.52, 145.43, 143.84, 143.80, 141.35, 127.75, 127.07, 127.05, 125.31, 124.97, 121.77, 120.04, 71.26, 66.75, 47.58, 47.27, 38.90, 34.09, 33.71, 32.68, 26.39, 26.21, 26.05. HRMS +ESI [M*Na]: 567.2107, calculated (C31H32N2O7): 567.2102.
(9H-fluoren-9-yl)methyl (S)-(l-(((4-nitrophenoxy)carbonyl)oxy)-4-phenylbutan-2- yl)carbamate (Fmoc-HoPhe-PNOC) . According to the general procedure for synthesis of activated carbonates: Crude Fmoc-HoPhe-ol (~5 mmol) gave a crude product that was purified by silica gel chromatography (DCM then 2: 1 Hex:EtOAc) and isolated as an off white solid (2.18 g, 79% yield over 2 steps). *H NMR (400 MHz, Chloroform-d) 5 8.25 (d, J = 9.1 Hz, 2H), 7.78 (d, J = 7.5 Hz, 2H), 7.61 (d, J = 7.5 Hz, 2H), 7.45 - 7.37 (m, 2H), 7.37 - 7.28 (m, 6H), 7.25 - 7.14 (m, 3H). 4.86 (d, J = 9.0 Hz, 1H), 4.55 - 4.41 (m, 2H), 4.37 - 4.16 (m. 3H), 4. 10 - 3.99 (m. 1H), 2.80 - 2.61 (m, 2H), 2.00 - 1.79 (m, 2H). 13C NMR (126 MHz, CDCh) 5 155.97, 155.37, 152.44, 145.46, 143.82, 143.74, 141.38, 140.75, 128.62, 128.35, 127.79, 127.10, 126.32, 125.33, 124.98, 124.94, 121.73, 120.06, 70.72, 66.70, 49.84, 47.32, 33.15, 32.14. HRMS +ESI [M+Na]: 575.1796, calculated (C32H28N2O7): 575.1789.
(9H-fluoren-9-yl)methyl (S)-(l-(((4-nitrophenoxy)carbonyl)oxy)butan-2-yl)carbamate (Fmoc-Abu-PNOC) . According to the general procedure for synthesis of activated carbonates: Crude Fmoc-Abu-ol (~6 mmol) gave a crude product that was purified by silica gel chromatography (CHCk) and isolated as an off white/clear viscous solid (1.3 g, 44% yield over 2 steps). Impure fractions were collected separately and not included. 'H NMR (400 MHz. Chloroform-ri) 5 8.25 (d. J = 9. 1 Hz. 2H). 7.77 (d. J = 7.6 Hz. 2H), 7.59 (d. J = 7.4 Hz, 2H), 7.44 - 7.34 (m, 4H), 7.33 - 7.28 (m, 2H), 4.80 (d, J= 9.0 Hz, 1H), 4.51 - 4.38 (m, 2H), 4.37 - 4.16 (m, 3H), 4.04 - 3.85 (m, 1H), 1.76 - 1.46 (mf, 2H), 1.01 (t, J= 7.4 Hz, 3H). 13C NMR (126 MHz, CDCh) 5 156.20. 155.53, 152.63, 145.57, 143.95, 143.91. 141.47, 127.87. 127.20, 125.44, 125.08, 121.87. 120.16. 70.54, 66.87, 51.66, 47.41, 24.59, 10.48. HRMS +ESI [M’Na]: 499.1486, calculated (C32H28N2O7): 499.1476.
(9H-fluoren-9-yl)methyl (S)-(l-(4-chlorophenyl)-3-(((4- nitrophenoxy)carbonyl)oxy)propan-2-yl)carbamate (Fmoc-4Cl-Phe-PNOC) . According to the general procedure for synthesis of activated carbonates: Crude Fmoc-4Cl-Phe-ol (4.7 mmol) gave a crude product that was purified by silica gel chromatography (3: 1 DCM:Hex then 4: 1, then 9: 1, then pure DCM) and isolated as an off white/yellow flaky solid (1.57 g, 58% yield over 2 steps). 'H NMR (400 MHz, Chloroform-J) 5 8.27 (d, J = 9.1 Hz, 2H), 7.77 (d, J= 7.5 Hz, 2H), 7.56 - 7.51 (m, 2H), 7.43 - 7.35 (m, 4H). 7.34 - 7.27 (m, 4H), 7.14 (d, J = 8.0 Hz. 2H), 4.86 (d. J= 8.1 Hz. 1H), 4.59 - 4.37 (m. 2H), 4.36 - 4.13 (mf 4H). 2.90 (d, J = 6.6 Hz, 2H). 13C NMR (126 MHz, CDCh) 5 155.71, 155.29, 152.34, 145.52, 143.70, 143.67, 141.36, 134.90, 132.99, 130.45, 128.99, 127.79, 127.07, 125.37, 124.87, 121.69, 120.05. 69.40, 66.77, 50.99, 47.21, 36.82. HRMS +ESI [M+Na]: 595.1260, calculated (C31H25CIN2O7): 595.1242.
(9H-fluoren-9-yl)methyl (S)-( !-((( 4-nitrophenoxy)carbonyl)oxy)propan-2-yl-l, 1- d2)carbamate (Fmoc-D2Ala-PNOC). According to the general procedure for synthesis of activated carbonates: Crude Fmoc-D2Ala-ol (4.7 mmol) gave a crude product that was purified by silica gel chromatography (3: 1 DCM:Hex then 4:1, then 9: 1, then pure DCM) and isolated as a fluffy white solid (2.2 g, 74% yield over 2 steps). 'H NMR (500 MHz, Chloroform-^ 5 8.26 (d. J = 9.1 Hz. 2H), 7.77 (d, J= 7.6 Hz. 2H), 7.59 (d, J= 7.5 Hz. 2H), 7.46 - 7.34 (m. 4H), 7.31 (t, J= 7.5 Hz. 2H), 4.85 (d. J= 8.3 Hz. 1H), 4.44 (d, J= 6.9 Hz. 2H), 4.22 (t, J= 6.8 Hz, 1H), 4.15 (t, = 7.5 Hz, 1H), 1.28 (d, = 6.9 Hz, 3H). 13C NMR (126 MHz, Chloroform-ti) 5 155.77, 155.40, 152.52, 145.46, 143.79, 1.43.76, 141.34, 127.77, 127.07. 125.34, 124.95, 121.75, 120.05, 66.84, 47.23, 45.74, 17.27. HRMS +ESI [M+Na]: 487.1456, calculated (C25H20D2N2O7): 487.1445.
(9H-fluoren-9-yl)methyl (S)-( !-((( 4-nitrophenoxy)carbonyl)oxy)-3-phenylpropan-2-yl-
1.1-d2)carbamate (Fmoc-D2Phe-PNOC) . According to the general procedure for synthesis of activated carbonates: Crude Fmoc-D2Phe-ol (4.7 mmol) gave a crude product that was purified by silica gel chromatography (3: 1 DCM:Hex then 4: 1, then 9: 1, then pure DCM) and isolated as an off white flaky solid. 'H NMR (400 MHz, Chloroform-c/) 6 8.27 (d, J = 9.1 Hz, 2H), 7.77 (d, J = 7.6 Hz, 2H), 7.56 - 7.51 (m, 2H), 7.44 - 7.26 (m, 9H), 7.22 (d, J= 7.4 Hz, 2H), 4.92 (d, J= 8.7 Hz, 1H), 4.51 - 4.33 (m, 2H), 4.29 (q, J= 8.0 Hz, 1H), 4.20 (t, J= 6.8 Hz. 1H), 2.93 (d, J= 7.3 Hz. 2H). 13C NMR (126 MHz, CDCh) 5 155.82, 155.36, 152.40, 145.48. 143.73, 141.34, 136.36, 129.16. 128.87. 127.77, 127.11, 127.07, 125.35. 124.99. 124.93, 121.72, 120.05, 66.86, 50.93, 47.21, 37.43. 21. HRMS +ESI [M+Na]: 563.1764, calculated (C31H24D2N2O7): 563.1758.
(9H-fluoren-9-yl)methyl (S)-(4-methyl-l-( ( 4-nitrophenoxy)carbonyl)oxy)pentan-2-yl-
1.1-d2) carbamate (Fmoc-DzLeu-PNOC). According to the general procedure for synthesis of activated carbonates: Crude Fmoc-D2Leu-ol (5.2 mmol) gave a crude product that was purified by silica gel chromatography (3: 1 DCM:Hex then 4:1, then 9: 1, then pure DCM) and isolated as a clear/ white/yellow glassy oil. 1 H NMR (500 MHz, Chloroform-<7) 5 8.24 (d, J = 8.8 Hz, 2H), 7.77 (d, J= 7.6 Hz, 2H), 7.59 (d, J= 7.5, 2H), 7.44 - 7.33 (m. 4H), 7.30 (t, J = 7.5 Hz. 2H), 4.74 (d. J= 9.0 Hz. 1H), 4.46 (qd, J= 10.7, 6.9 Hz, 2H), 4.22 (t, J= 6.9 Hz, 1H), 4.16 - 4.05 (m, 1H), 1.72 - 1.65 (m, 1H), 1.51 - 1.32 (m, 2H), 0.96 (t, J= 6.5 Hz, 6H). 13C NMR (126 MHz, Chloroform-J) 5 156.00, 155.42, 152.52, 145.43, 143.79 (d, J= 9.8 Hz), 141.36, 127.75, 127.07. 125.32, 124.94, 121.76, 120.03, 66.70, 48.12, 47.31, 40.28, 24.71, 23.02. 21.99. HRMS +ESI [M+Na]: 529.1924, calculated (C28H26D2N2O7): 529.1914.
(9H-fluoren-9-yl)methyl (S)-(3-methyl-l-(((4-nitrophenoxy)carbonyl)oxy)butan-2-yl-
1.1-d2)carbamate (Fmoc-D2Val-PNOC) . According to the general procedure for synthesis of activated carbonates: Crude Fmoc-D2Val-ol (4.9 mmol) gave a crude product that was purified by silica gel chromatography (DCM neat) and isolated as a fluffy white solid (2.34 g, 97% yield). *HNMR (500 MHz, Chloroform-J) 5 8.26 (d, J= 9.0 Hz, 2H), 7.79 (d, J= 7.5 Hz. 2H), 7.62 (d. J= 7.6, 2H), 7.42 (t, J= 7.5 Hz, 2H), 7.40 - 7.30 (m, 4H), 4.85 (d, J= 9.6 Hz. 1H), 4.55 - 4.42 (m. 2H), 4.25 (t, J= 6.8 Hz, 1H), 3.86 (d, J = 8.1 Hz, 1H), 1.92 (h, J =
6.9 Hz, 1H), 1.03 (dd, J= 16.9, 6.7 Hz, 6H). 13C NMR (126 MHz, Chloroform- ) 5 156.29, 155.41, 152.55, 145.43, 143.81, 143.77, 141.35, 127.75, 127.07, 125.31, 124.96, 121.74, 120.03. 66.76, 55.13, 47.32, 29.38, 19.44, 18.50. HRMS +ESI [M+Na]: 515.1769, calculated (C27H24D2N2O7): 515.1758.
(9H-fluoren-9-yl)methyl (S)-( !-((( 4-nitrophenoxy)carbonyl)oxy)butan-2-yl-l, 1- d2)carbamate (Fmoc-D2Abu-PNOC) . According to the general procedure for synthesis of activated carbonates: Crude Fmoc-D2Abu-ol (4.4 mmol) gave a crude product that was purified by silica gel chromatography (DCM then DCM w/ 1-5% EtOAC gradient elution) and isolated as a fluffy white solid (1.89 g, 90% yield). 'H NMR (400 MHz, Chloroform- ) 5 8.23 (d, J = 9.2 Hz, 2H), 7.75 (d, J= 7.6 Hz, 2H), 7.57 (d, J= 7.5 Hz, 2H), 7.43 - 7.26 (m, 6H), 4.78 (d, J= 9.0 Hz, 1H), 4.51 - 4.38 (m, 2H), 4.21 (t, J= 6.8 Hz, 1H), 3.92 (q, J= 8.3 Hz. 1H), 1.69 - 1.49 (m, 2H), 0.99 (t, J= 7.4 Hz, 3H). 13C NMR (126 MHz, Chloroform- ) 5 156.10. 155.41, 152.52, 145.44, 143.83. 143.78. 141.35, 127.75, 127.08, 125.32. 124.95. 121.74, 120.04, 66.76, 51.40, 47.28, 24.42, 10.35. HRMS +ESI [M+Na]: 501.1606, calculated (C26H22D2N2O7): 501.1601.
(9H-fluoren-9-yl)methyl (S)-(l -cyclohexyl-3-( ( 4-nitrophenoxy)carbonyl)oxy)propan- 2-yl-3,3-d2)carbamate (Fmoc-D2CHA-PNOC). According to the general procedure for synthesis of activated carbonates: Crude Fmoc-D2CHA-ol (4.7 mmol) gave a crude product that was purified by silica gel chromatography (DCM then DCM w/ 1-5% EtOAC gradient elution) and isolated as a fluffy white solid (2.08 g, 80% yield). 'H NMR (500 MHz, Chloroform-^ 5 8.24 (d. J= 9.0 Hz. 2H), 7.77 (d. J= 7.6 Hz. 2H), 7.63 - 7.56 (m, 2H), 7.44 - 7.28 (m, 6H), 4.77 (d, J= 8.9 Hz, 1H), 4.49 - 4.39 (m, 2H). 4.23 (t. J= 6.9 Hz. 1H). 4. 18 -
4.10 (m, 1H), 1.82 (d, J= 12.9 Hz, 1H), 1.76 - 1.62 (m, 3H), 1.50 - 1.32 (m, 3H), 1.32 - 1.08 (m, 4H), 1.08 - 0.83 (m, 2H). 13C NMR (126 MHz, CDCh) 5 155.98, 155.44, 152.53, 145.42,
143.84. 143.80, 141.35, 127.75, 127.08, 125.31. 124.97, 121.77, 120.04, 66.76, 47.44, 47.27,
38.84, 34.09. 33.71, 32.68, 26.40, 26.21, 26.06. HRMS +ESI [M+Na]: 569.2237. calculated (C31H30D2N2O7): 569.2227.
(9H-fluoren-9-yl)methyl(S)-( 1 -( 4-chlorophenyl)-3-( ((4- nitrophenoxy)carbonyl)oxy)propan-2-yl-3,3-d2)carbamate (Fmoc-D2-4Cl-Phe-PNOC) .
According to the general procedure for synthesis of activated carbonates: Crude Fmoc-D2- 4Cl-Phe-ol (5.3 mmol) gave a crude product that was purified by silica gel chromatography (3:2 Hex EtOAc then 1: 1 Hex:EtOAc) and isolated as an off white solid (1.99 g, 91% yield). 'H NMR (400 MHz. Chloroform^/) 5 8.27 (d. J = 9.2 Hz. 2H), 7.77 (d. J= 7.6 Hz. 2H), 7.59 - 7.49 (m, 2H), 7.45 - 7.27 (m, 8H), 7.14 (d, J= 8.0 Hz, 2H), 4.88 (d, J= 8.7 Hz, 1H), 4.53 - 4.34 (m, 2H), 4.32 - 4.15 (m, 2H), 2.89 (d, J= 7.3 Hz, 2H). 13C NMR (126 MHz, CDCls) 5 155.74. 155.30, 152.36, 145.51, 143.70, 143.66. 141.36, 134.91, 132.98, 130.45, 128.99, 127.79. 127.08, 125.37, 124.90, 124.87. 121.69. 120.05, 66.78, 50.86, 47.21. 36.77. HRMS +ESI [M+Na]: 597.1375, calculated (C31H23D2CIN2O7): 597.1368.
(9H-fluoren-9-yl)methyl (S)-(l-((( 4-nitrophenoxy)carbonyl)oxy)-4-phenylbutan-2-yl-
1 , l-d2)carbamate (Fmoc-D2HoPhe-PNOC) . According to the general procedure for synthesis of activated carbonates: Crude Fmoc-DzHoPhe-ol (5.3 mmol) gave a crude product that was purified by silica gel chromatography (3:2 Hex:EtOAc then 1: 1 Hex:EtOAc) and isolated as an off white solid (2.02 g, 68% yield). 'H NMR (400 MHz, Chloroform- ) 5 8.23 (d, J = 9.2 Hz, 2H), 7.76 (d, J= 7.5 Hz, 2H), 7.58 (d, J= 7.5 Hz, 2H), 7.43 - 7.25 (m, 8H), 7.23 - 7.10 (m, 3H), 4.79 (d, J= 9.1 Hz, 1H), 4.55 - 4.40 (m, 2H), 4.21 (t, J= 6.6 Hz, 1H), 4.10 - 3.98 (m. 1H), 2.69 (q. J= 8.7 Hz. 2H), 1.88 (p. J = 8.3 Hz. 2H). 13C NMR (126 MHz, CDCh) 5 155.95, 155.36, 152.45, 145.46, 143.80, 143.73, 141.37, 140.72, 128.62, 128.34, 127.78, 127.08, 126.32, 125.33, 124.97, 124.92, 121.72, 120.05, 66.69, 47.31, 33.11, 32.13. HRMS +ESI [M+Na]: 577.1915, calculated (C32H26D2N2O7): 577.1914.
Parallel Solid-Phase syntheses in a deep-well 96 well plate
General procedure for the solid-supported synthesis of oligourethanes in a plate.
Coupling: Phenyalaninol loaded (0.44 mmol/gram, 200-400 mesh) 2-chlorotrit l polystyrene resin (15 mg, 0.0075 mmol) was added to 18 wells (G1-G12, H1-H6) of a 10 pm fritted deepwell 96-well plate (combinatorial microlute by Porvair). A 96-well microplate deep well drain cap mat (cat. No. 219005, Porvair Sciences) was used to seal the base of the filtration plate. The resin was suspended in 0.2 rnL of anhydrous N-methyl-2-pyrrolidinone (NMP) and swollen. A stock solution of Hunig’s base (36 pL total, 2 pL or 0.01125 mmol per well) and hydroxybenzotriazole (110 mg total. 6.1 mg or 0.045 mmol per well) was made in NMP (1.8 mL total. 0. 1 mL/well). The plate was agitated to mix the contents of the wells. Finally, the activated amino alcohols Fmoc-XXX-PNOC (0.022 mmol per well) were dissolved in NMP (0.2 mL per well) in separate vials. Once dissolved, the monomers were transferred to their corresponding wells. The plate was sealed with Nunc aluminum sealing tape (no. 276014) and shaken overnight. Note: proper mixing of the wells is important for efficient coupling, and previous reports show the coupling to be complete in as little as 4 hours [37],
Resin in all 18 wells were washed with NMP (5x2 mL), then DCM (5x2 mL), and finally Et2O (3x2 mL) by vacuum filtration through the fritted plate (Supelco PlatePrep 96- well Vacuum Manifold kit). Resin was dried overnight under vacuum. Test cleavages were affected with 1% TFA in DCM (5x1 mL) for 10 seconds each. Coupling efficiency was checked by LC/MS. Note: depending on the cleavage times and amounts of TFA used, the trifluoroacetic ester was observed by LC/MS. This ester was readily hydrolysed dissolving the sample in DCM and shaking with saturated NaHCCL.
General procedure for labelling of the terminal amine with a NBD-Fluoride (on resin). In the fritted 96-well plate, to each of the 18 wells containing 0.0075 mmol oligomer loaded resin was added a solution of Hunig's base (13 uL/well, 0.075 mmol), 4-Fluoro-7- nitrobenzofurazan (6.9 mg or 0.0375 mmol per well) in anhydrous DMF (0.5 mL per well). The plate was sealed and the reactions were left to shake overnight. Resin was washed with DMF (5x2 mL), DCM (10x2 mL, or no more yellow/green was observed in the wash), and Et2O (3x5 mL).
Cleavage procedure: Cleavages were affected with 0.1% TFA in DCM (3x2 mL) for 5 minutes each. Resin was filtered off, and cleaved product was concentrated in a clean 96- well plate. Note: depending on the cleavage times and amounts of TFA used, the trifluoroacetic ester of the terminal alcohol was observed by LC/MS. This ester was readily hydrolysed in the proceeding sequencing methodology, and thus ignored, but is present in all initial LC/MS chromatograms.
Sequencing Experiments
Sequencing Al and A2. The oligomers (measured to be at a final concentration between 1-2 milimolar) were dissolved in DMSO and added to a 96-well plate. Next, cesium carbonate was added to the solution. The final concentration of base was approximately 5 milimolar. The reaction was heated to 70 °C, and held at the temperature for 45 minutes. Reaction was sampled for LC/MS by taking 50 pL of the reaction and diluting it into 75 pL methanol. This was repeated at 105, 165, 225, and 300 minutes.
Sequencing G1-G12 and H1-H6. The oligomers (measured to be at a final concentration between 1-2 milimolar) were dissolved in 2: 1 MeOH FLO and added to a 96- well plate. Next, potassium phosphate (K3PO4) was added to the solution. The final concentration of base was approximately 75 milimolar. The reaction was heated to 70 °C, and held at the temperature for 30 minutes. Reaction was sampled for LC/MS by taking 62.5 pL of the reaction and diluting it into 75 pL of 1 : 1 MeOftiffcO. This was repeated at 60, 90, 120, and 150 minutes.
Isotopically Tagged Sequence Defined Oligomers for Encryption
Nearly all abiotic sequence-defined polymers so-far applied to information storage purposes use short oligomeric structures, inherently limiting the data stored per molecule. While storing information in multiple oligomers offers the advantages of simplifying the synthesis and sequencing methodologies, it requires spatial organization of the oligomers for proper read-out. For example, self-assembled layers and mixtures [63, 79], crystals [61], or well plates [71, 80], have been used. Generally, every oligomer must be analyzed independently to prevent over-complication of the MS/MS spectra. Methodologies to analyze mixtures of information carrying macromolecules will effectively increase the information density of a given sample and are needed to achieve a truly dense storage medium.
To complement and improve upon many of the strategies discussed above, a chainend depolymerization sequencing methodology for sequence-defined oligourethanes (OUs) was recently described that utilizes a thermally induced intramolecular cyclization to iteratively remove the terminal monomer (Scheme 1). This allows each truncated oligourethane to be characterized by simple liquid chromatography -mass spectrometry (LC/MS), forgoing MS/MS protocols and greatly simplifying deconvolution [77], Here, the use of this chain-end depolymerization sequencing to deconvolute a complex mixture of eight different 10-mer oligourethanes as a means to increase the information density per sample is reported. The storage and retrieval of a 256-bit cipher key held within a single mixture of the oligourethanes is demonstrated, which is the most information stored in a single sample of abiotic sequence-defined polymers to date, and is applies to molecular cryptography and steganography.
Expanding Information Storage Capabilities in Digital Polymers. It was posited that a mixture containing X n-mers has the same effective storage capacity as a single X»n-mer. However, to achieve effective sequencing of such a complex mixture, it would be required that each oligomer be distinguishable from one another, and that no intermolecular crossreactivity occur during the sequencing event (Scheme 6). The oligomers need to be distinguished not only to correctly sequence, but also to know their proper ordering such that the stored information can be accurately reconstructed. One well-established methodology for distinguishing species in a complex mixture using mass spectrometry is isotopic labelling [81], Frequently in proteomics, stable isotope labels are utilized to identify otherwise equivalent peptide or peptide fragments. Isotope-ratio mass spectrometry can go so far as to measure the relative abundance of isotopes in a given sample [82], Given that one can measure the relative abundance of an isotopically-labelled and non-labelled sample, it was hypothesized that it would be possible to differentiate oligomers with unique isotopic labels and sort them according to their isotope tags. Scheme 6. Intramolecular, chain-end depolymerization for sequencing via a 5-exo-trig cyclization removes one monomer at a time, allow ing for controlled sequencing of multiple oligomers simultaneously without inter-oligomer cross-reactivity.
Isotopic Labeling and Isotopologue Design. To store 256 bits of information, it was chosen to encode a cipher key in hexadecimal (base- 16) in a mixture of eight 10-mer oligourethanes (8 of the 10 monomers encode information, see below). In base- 16, each monomer provides a storage density of 4 bits per monomer, thus 32 bits per 10-mer, and overall, 256 bits in the sample. Thus, sixteen monomers and eight unique isotope labels entirely distinct from one another, that are chemically persistent throughout the sequencing, were designed. The monomers w ere synthesized by protio-reduction [33] or deuteroreduction [71] of commercially available canonical and non-canonical amino acids (Figure 19 and Scheme 2-Scheme 4). Simple deuteration at the a-methylene of each monomer effectively doubles the monomer pool without changing the complexify of the synthesis or sequencing chemistries [71], Reaction with 4-nitrophenyl chloroformate converted each monomer to an activated carbonate in good yield. All sixteen monomers are shown in Figure 19.
Synthesis of Halogen Labels. Next, eight unique and differentiable isotope labels were sought. In this design criteria, each label would be placed on the N-terminal monomer, making it persistent and present on each precursor ion throughout the sequencing process (Scheme 6). Chlorine exists as two stable isotopes, 35C1 and 37C1, with relative abundances of 75.76% and 24.24%, respectively [83, 84], Thus, this natural abundance of the stable isotopes is observed in the mass spectra of monochlorinated organic molecules by an increase in intensity of the M+2 peak, relative to this natural abundance (an approximate 3: 1 ratio the M to M+2 peaks; Figure 44). The same is observed for bromine, which also exists as two stable isotopes, 79Br and 81Br (51% and 49% abundance, respectively) [83], By substituting a molecule with two chlorines, intensification of the M+2 and M+4 peaks was observed, generally resulting in a 9:6: 1 ratio relative to the molecular ion (Figure 44). Similarly, two bromines generally give a 1 :2:1 ratio relative to the molecular ion peak (Figure 44). To verify the prominence of the halogen tags on the end of the 1 Omers, an isotope distribution calculator available online (available from the UT Mass Spec Department DropBox) was first used to model how these tags would alter the isotope pattern of the molecule. In simulation, these molecules displayed patterns easily distinguishable from one another (Figure 43). To bring these molecules into reality, monomers with one chlorine, two chlorines, one bromine, and two bromines were then synthesized and isolated (Figure 44).
Synthesis of Isotopologue Labels. To generate four other isotope labels distinct from each halogen tag, the isotopic composition of the N-terminal urethane monomer was altered, creating isotopologues. Isotopologues are molecules with the same chemical formula and bonded arrangement of atoms, but one or more atoms have been replaced by an isotope with a different number of neutrons than the parent molecule (i.e., hydrogen vs. deuterium) [ 85 ] . By altering the deuteration level of the N-terminal amino alcohol, the relative proportions of the M+2 and M+3 peaks can be altered to give unique and identifiable mass spectra. Thus, protio-reduced and deutero-reduced monomers were combined at various stoichiometries to attain specific ratios of isotopologues. As a test, three tetramers were synthesized, Al, A2, and A3, using a previously reported methodology (Figure 48 and Scheme 7-Scheme 9) [71], The final coupling step for each oligomer used a stoichiometric mixture of deuterated and non-deuterated monomers (1:3, 1: 1, 3: 1 ratio, see Figure 45-Figure 47). By LC/MS, the unique isotopologue mixtures could be observed and easily distinguished (Figure 48).
Scheme 7. Synthesis of tetramer Al.
5 Scheme 8. Synthesis of tetramer A2.
Scheme 9. Synthesis of tetramer A3.
As may be expected, the high-resolution mass spectra reveals that the ratio of deuterated to non-deuterated oligomer is the same as the deuterated to non-deuterated monomer used during synthesis (Figure 45-Figure 48). Further, as may be predicted, applying ratios of 3: 1 and 1 :1 non-deuterated to deuterated monomer created isotope patterns mimicking the calculated chlorine and bromine isotope patterns (Figure 49). However, a 1 :3 ratio (A3) created an entirely new pattern. Thus, it was sought to create mixtures of isotopologues that would provide unique mass spectra distinguishable from one another, as well as being distinguishable from the halogen tags. In this regard, the non-deuterated to deuterated ratios implemented include 1:0, 5.6: 1, 1:2 and 1 :5.25 (ca., 0%, 15%, 67%, 84% deuterated, respectively). The entire suite of isotope labels includes one tag with no label, three isotopologues and four halogen-based labels (Figure 50).
Molecular Data Encryption. In order to synthesize and encode a 256-bit cipher key in oligourethanes, an encryption (Mol.Encrypter) and decry ption algorithm (Mol. Decry pier) capable of generating and utilizing the key was developed (https://github.com/PhysicalOrganic/Mol.E-Crypter). This encryption and decryption algorithm is based in the Advanced Encryption Standard (AES), the standard adopted by the US government, which is a publicly accessible and open cipher approved by the National Security' Agency (NSA) [86], AES utilizes a symmetric-key algorithm, meaning the key is used to both encrypt and decrypt the data, and is ideal for protecting data at rest. A 256-bit cipher key is considered impenetrable by brute force or exhaustive key searches. This encry ption process occurs through four matrix transformations: SubBytes, ShiftRows, Mix Columns, and Add Round Key. The length of the secret key indicates how many rounds of transformations the matrix will undergo, 10 rounds for 128-bit keys, 12 rounds for 192-bit keys, and 14 rounds for 256 bit keys. Therefore, the longer the secret key, the more secure the encryption is. The message or document being encrypted is broken it bits of data that would be easy for anyone to read (i.e., the phrase ‘“physical organic chemistry” is broken into “physi”, “calorg”, “aniche”, and “mistry” and then its hexadecimal representation.). The message is then mixed with the secret key and undergoes the matrix transformations previously listed to make it almost impossible to decipher without knowledge of the key. Thus, storing the cipher key within a molecular medium would offer a layer of security akin to the use of a hardware security module, as key storage is a considerably important part of cryptography.
A 256-bit string (the cipher key) was generated and used to encrypt a document containing the novel The Wonderful Wizard of Oz, by L. Frank Baum (1900). The 256-bit key was then converted to hexadecimal according to the Unicode standard [87], Within the encoding algorithm, the hexadecimal characters w ere arbitrarily assigned to each of the sixteen monomers (Figure 19). With each hex character representing four bits, the hexadecimal string w as 64 characters long. Thus, as mentioned previously, the 64 characters were encoded across eight 10-mer oligourethanes. In this encoding scheme, each oligomer contained 8 information containing characters, one indexing monomer (at the O-terminal). and the N-terminal isotope tag. The indexing monomer (phenylalaninol) provides a reading frame with which to begin sequence deconvolution, allowing the analyst to determine the isotope pattern present in the mass spectra between tw o precursor ions of the same oligomer.
Moreover, extremely important to the strategy of simultaneous sequencing is knowing the proper order of the eight 10-mers. Just as the 16-monomers were arbitrarily assigned hexadecimal characters, the mass tags of Figure 50 were arbitrarily assigned to the 1st, 2nd, 3rd .. . to 8th oligomer that would read out the 256-bit code in the proper order. In other words, identification of the isotope tag is required for accurate sorting of the eight hexadecimal strings to rebuild the original cipher key.
Synthesis of Eight Oligomers Isotopically Tagged Oligomers. The eight oligomers (B1-B8) w ere successfully synthesized on the solid-phase, in parallel, in a fritted 96-well plate (Scheme 10). The N-terminal amine was capped with a long-wavelength chromophore, 4-fluoro-7-nitrobenzofurazan to monitor the chain-end self-sequencing and aid in sequence deconvolution. B1-B8 were cleaved from the solid-phase, purified, and characterized by high-resolution MS (Figure 51 -Figure 58).
Scheme 10. Generic synthesis of the oligomers on the solid support (B1-B8).
Oligomer Simul-Sequence. For the simultaneous sequencing and decoding, 500 nmols of each oligomer (B1-B8) were combined in a single sample (Figure 59). The oligomers were dissolved in DMSO with 0.01 M cesium carbonate (CS2CO3) and heated to 70 °C in an incubated shaker and sampled for LC/MS at designated time points (Figure 60), monitoring the depolymerization over time. At the 0-minute time-point, only the eight starting oligomers were observed with each truncated iteration being formed as the reaction progressed until nearly full consolidation at the 1-mer for all eight sequence-defined polymers at 550 minutes. Figure 60 succinctly shows the complexity of the sequencing reaction. As the oligomers simultaneously and orthogonally depolymerize, the existence of all eighty discrete precursor ions were cleanly and clearly observed by negative mode ESIMS (Figure 59), as well as the sixteen unique cyclized oxazolidinones (observed in positive mode). It is worth noting that these oligomers were found to ionize strongly under an acetonitrile gradient with at least 10 - 50 mM ammonium acetate buffer, attributed to the suppression of undesired salt adducts in ammonium acetate buffers [88], Furthermore, during the time-point sampling of the sequencing reaction, a separate LC/MS sample was prepared wherein each time-point was combined. By high-resolution MS, the precursor ions of all 80 corresponding oligomers, as well as the sixteen oxazolidinones, were observed in a single high-resolution LC/MS sample (Figure 59). This highlights the power of using chain-end depolymerization rather than tandem MS for reading of the stored information.
Moreover, due to the chain-end depolymerization mechanism, simple visual inspection of the low-resolution LC/MS data was sufficient for complete sequence deconvolution. Prior to high-resolution confirmation, utilizing the isotopic fingerprints, two analysts were each able to correctly identify and sort the eighty precursor ion masses into a templated spreadsheet. Each analyst singlehandedly and independently analyzed and identified the eighty unique masses, with one analyst never having seen the structures of the oligomers prior to their analysis. To aid in mass sorting and sequence deconvolution, said analyst was provided a series of guidelines. Critical to rebuilding the cipher key, both analysts were able to successfully identify all eight of the stable isotope tags, which were used for the correct ordering of the hexadecimal character string. Figure 62 shows the low- resolution LC/MS data for three of the eight oligomers in detail, denoting how the isotope tags were capable of distinguishing oligomeric species (all 8 oligomers are shown in Figure 51 -Figure 58).
Observations. A few noteworthy observations were made during the data analysis and sequence convolution. First being that as a given cyclization event occurred, the resulting oligomer would lose one nitrogen atom from its molecular formula. Thus, the nitrogen rule [89] could be utilized as a tool to aid in mass sorting: as an oligomer sequences, its nominal mass will alternate from an even to odd number as the molecular ion alternates from an even to odd number of nitrogen atoms (as seen in Figure 61). Furthermore, when the oligomers are longer (i.e. 10-mer or 9-mer), the exact identity of a given isotope tag can obscured by the contributions of naturally-abundant isotopes (13C is 1.107%, 15N is 0.366%, and 18O is 0.204%) [90], Thus, the M+l, M+2, M+3, etc., peaks are attenuated with respect the amount of carbon, nitrogen, and oxygen present in the molecule, with this effect amplified as the number of atoms in each oligomer increases. Of course, this atomic composition varies between each oligomer (B1-B8), but also between a given oligomer as it depolymerizes. The isotope pattern of an oligomer will incrementally decrease in complexity with each cyclization event until arriving at a very simple and distinguishable pattern at the Imer (as seen in Figure 61). As such, if an isotope tag cannot be identified at the 10-mer, the mass spectra can be compared and matched (for example, comparing the similar patterns of 10-mer and 9-mer) until the isotope tag is identified at the smaller truncated oligomers. Likewise, knowing the indexing monomer prior to sequencing allows immediate identification of the 9- mer which, upon comparison with the 10-mer, provides a baseline MS with which to begin identifying and sorting the masses. Important to this encoding paradigm, the isotope pattern matching was capable of distinguishing two masses with the exact same base peak m/z due to the differences in isotope tag and resulting spectra (Figure 62).
Oligomer Decryption. Finally, the spreadsheet containing the precursor ion masses was fed into the decry ption algorithm, which then converted the sequencing information into the original 256-bit cipher key. This was accomplished by getting the difference in mass from the truncated oligomers. This mass was then correlated to a pre-assigned hex character which was then converted back to binary, thus recreating the 256-bit cipher key. In a similar manner to how the encryption process happens (described above), the document is decrypted using a series of matrix transformations along with the cipher key. This 256-bit cipher key was able to decrypt the document containing the novel The Wonderful Wizard of Oz. successfully revealing the encrypted information.
Independent Validation, Steganography. A medium of information storage is only as good as its ability to accurately and easily confer the information being stored. It w as believed the sequencing paradigm to be robust and simple, and therefore easily decodable by a third-party without any prior experience implementing these sequencing and deconvolution protocols. For this, a real-world demonstration of this platform in cryptography and steganography was performed using the molecular 256-bit cipher key. First, the molecular cipher (500 nmols of each oligomer, Bl - B8) was dissolved in isopropanol and combined with glycerol and soot, making a writable ink. The oligourethane “ink” mixture was placed in an emptied ballpoint pen, which was then used to handwrite a letter to a third- party collaborator (Figure 63). The molecular cipher key, embedded in a standard piece of printer paper w as mailed from Austin, Texas to Lowell, Massachusetts where it was extracted with dichloromethane and concentrated. The third-party collaborators were given a set of discreet instructions on the reaction set up and sequence deconvolution. In their very first attempt, the collaborators were able to correctly sequence the eight oligomers and identify the precursor ions, which were entered into the decryption algorithm, decry pting the file and revealing the document containing The Wonderful Wizard of Oz.
Conclusion. Herein, a robust method for the simultaneous synthesis and sequencing of multiple discreet macromolecules with the use of isotope tags to increase the information storage capacities of sequence-defined abiotic oligomers was described. The isotope tags made each oligomer distinguishable from one another, such that sequence deconvolution was possible even in a mixture containing up to 96 unique molecules. Important to note, all 96 molecules were characterized by high-resolution MS in a single LC/MS sample. Two programs, Mol.Encrypter and Mol.Decrypter were developed and utilized to encode and decode documents using a 256-bit cipher key. following the Advanced Encryption Standard. This molecular medium of information storage was able to store and deliver the 256-bit cipher key, which is the most information ever stored in single sample of abiotic sequence- defined macromolecules. The molecular cipher key and Mol.Decrypter were used successfully to decrypt a document containing The Wonderful Wizard of Oz. In an example of molecular steganography, a sample of discretely hidden oligourethanes embedded in '‘ink” was mailed to a third party, where the sequencing, structure deconvolution, and cipher key recovery was independently performed successfully. This paradigm has significant potential for widely accessible molecular information storage and encryption. Future iterations will look to robotically automate the writing and reading processes, furthering its accessibility and practicality7 for real-world applications.
Mol.E-Crypter
README. Mol.E-Crypter works to convert a plain text document into an encry pted file (extension .bin) using the Py cryptodome python library (https://github.com/Legrandin/pyct ptodome/blob/master/Doc/src/features.rst) and subsequently decode the encry pted file if the user enters in the correct secret key. The secret key is entered via sequencing oligourethanes and entering the masses into a spreadsheet, that is read by the decrypting program and if the key is correct, the unencrypted text file is output.
MoLEncrypter Usage. The first part of Mol.Encrypter. py uses a secure random number generator to be the secret key. A text document is the read in and undergoes AES- 256 encry ption and is then output as a .bin in addition to the secret key being output as a .txt file.
To run this program the program takes a single input: a text document containing the text to be encoded.
The program then outputs two files, encrypted.bin and secretkey.txt. The first file, encrypted.bin, contains an AES encrypted file that can only be deciphered using the secret key. The second file is the secret key, converted from binary into its hexadecimal representation to be encoded into oligourethanes.
Mol.Decrypter Usage. The second part of Mol.E-Crypter is Mol.Decoder.py. This program takes in three inputs, the encrypted file, the file containing the monomers correlation to hexadecimal, and the masses determining from sequencing the oligourethanes.
The program first reads in the monomers from the excel file containing the sequenced masses and generates the corresponding hexadecimal representation that is then converted to binary to decrypt the file. If it is the correct secret key, the file will be decrypted and a txt file containing the information will be created.
General Procedures for Isotope Encryption Work
Materials and Instrumentation. All materials used in the synthesis of each compound and related tests, were purchased from Sigma- Aldrich Chemical Co., Acros Organics, Tokyo Chemical Industry, Chem Impex International, etc. and used without further purification. Solvents (DCM, NMP, chloroform, DMSO, DMF, MeOH, MeCN, Isopropanol) were of reagent grade or HPLC grade uality and purchased from Fischer Scientific. NMR solvents (CDCh, DMSO-t*) were purchased from Cambridge Isotope Laboratories.
Column chromatography was performed using silica gel 60 (230 ± 400 mesh. 0.040 ± 0.063 mm) from Dynamic Adsorbents.
TLC analyses were carried out using Silica TLC Plates Aluminium Backing 20 by 20 cm sheet UV active at 254 nm.
Reverse phase column chromatography was done HPLC purifications were performed on Shimadzu Prominence HPLC system equipped with Zorbax SB-C 18 preparatory column (21.2 x 250 mm) with 7.0 pm packing material. Analytical HPLC traces were also carried out using a Zorbax SB-C18 analytical column (4.6 x 250 mm) with 5.0 pm packing material. 5- 95% gradient elution (MeCN/H2O with 0.1% formic acid). Hydrophobic urethanes utilized 30-95% gradient elution (MeCN/H2O with 0.1% formic acid).
'H and 13C spectra were recorded on Varian DirectDrive or Varian INOVA 400 MHz NMR spectrometers. The NMR spectra were referenced to solvent and the spectroscopic solvents were purchased from Cambridge Isotope Laboratories.
Liquid Chromatography /Mass spectra were recorded on an Agilent Technologies 6120 Single Quadropole or 6125B Single Quadrapole interfaced with an Agilent 1200 series liquid chromatography system equipped with a diode-array detector. Column: Agilent ZORBAX Eclipse Plus C18 narrow bore column; 2.1 mm internal diameter; 50 mm length; 5 micron particle size; P.N. 959746-902. Resulting spectra were analyzed using Agilent LC/MSD ChemStation. Separations achieved with or MeCN/Water 5-95% w/ 50 mM ammonium acetate gradient elution. High resolution mass spec was performed by the UT- Austin Mass Spectrometry Facility using Agilent Technologies 6530 Accurate Mass Q-TOF LC/MS system and Agilent Technologies 6546 Accurate-Mass Q-TOF LC/MS (Q-TOF G6546A), interfaced with Agilent Technologies 1260 Infinity II liquid chromatography system (G7112B), utilizing Agilent Technologies Dual Jet Stream ESI.
Parallel Synthesis was performed in Porvair Sciences Combinatorial Microlute 10 uM fritted deep-well 96-well plate catalogue number 240054.
Parallel Sequencing was performed in Nunc polypropylene DeepWell 96-well plates. Incubation at 70 °C was achieved using Ika KS 3000 i control incubator shaker.
Synthesis and Characterization. Preparations and monomers previously listed above are omitted.
(9H-fluoren-9-yl)methyl (R)-(l-hydroxy-3-methoxypropan-2-yl)carbamate (Fmoc- SerOMe-ol). According to the general procedure for the reduction of Fmoc-amino acids: Fmoc-O-methyl-L-serine (1.63 g, 4.79 mmol) gave a crude product that was then filtered through a plug of silica gel (1: 1 Hex:EtOAc) and concentrated as the semi-crude product. The product was identified by LC/MS and carried through the next step without isolation. MS; ESI+ m/z (M+H)+ 328.2, calculated (C25H25NO4): 328.15.
(9H-fluoren-9-yl)methyl (S)-(l -hydroxy-3-( 4-methoxyphenyl)propan-2-yl)carbamate (Fmoc- TyrOMe-ol) . According to the general procedure for the reduction of Fmoc-amino acids: Fmoc-O-methyl-L-tyrosine (2.0 g, 4.79 mmol) gave a crude product that was then filtered through a plug of silica gel (1 : 1 : 1 Hex:EtOAc:DCM) and concentrated as the semicrude product. The product was identified by LC/MS and carried through the next step without isolation. MS; ESI+ m/z (M+H)+ 404.2. calculated (C19H21NO4): 404.19
(9H-fluoren-9-yl)methyl (S)-( l-( 4-bromophenyl)-3-hydroxypropan-2-yl)carbamate (Fmoc-4BrPhe-ol) . According to the general procedure for the reduction of Fmoc-amino acids: Fmoc-4-bromo-/.-phenylalaninc (1.0 g, 2. 14 mmol) gave a crude product that was then filtered through a plug of silica gel (1 : 1 EtOAc:DCM) and concentrated as the semi-crude product. The product was identified by LC/MS and carried through the next step without isolation. MS; ESI+ m/z (M+H) 452.1, calculated (C^IfeBrNCF): 452.09
(9H-fluoren-9-yl)methyl (S)-(l-(3,5-dichlorophenyl)-3-hydroxypropan-2- yl)carbamate (Fmoc-3,5DiCl-Phe-ol) . According to the general procedure for the reduction of Fmoc-amino acids: Fmoc-3,5-Dichloro-L-phenylalanine (1.0 g, 2. 19 mmol) gave a crude product that was then filtered through a plug of silica gel (1 : 1 EtOAc:Hex) and concentrated as the semi-crude product. The product was identified by LC/MS and carried through the next step without isolation. MS; ESI+ m/z (M+H)+ 442.1, calculated (C24H21CI2NO3): 441.09 (9H-fluoren-9-yl)methyl (S)-(l -hydroxy-3-methoxypropan-2-yl-l, l-d2)carbamate (Fmoc-DiSerOMe-ol) . According to the general procedure for the deutero-reduction of Fmoc-amino acids: Fmoc-O-methyl-L-serine (1.63 g, 4.79 mmol) gave a crude product, identified by LC/MS, that was filtered through a silica plug with 1 : 1 EtOAc:Hex:, and carried through to the next step without isolation. MS; ESI+ m/z (M+H)+ 330.2, calculated (C19H19D2NO4): 330.17
(9H-fluoren-9-yl)methyl (S)-(l-hydroxy-3-(4-methoxyphenyl)propan-2-yl-l.l- d2)carbamate (Fmoc-D2TyrOMe-ol). According to the general procedure for the deuteroreduction of Fmoc-amino acids: Fmoc-O-methyl-£-tyrosine (2.0 g, 4.79 mmol) gave a crude product, identified by LC/MS, that was filtered through a silica plug with 1 : 1 : 1 EtOAc:Hex:DCM, and carried through to the next step without isolation. MS; ESI+ m/z (M+H)+ 406.2, calculated (C25H23D2NO4): 406.20
(9H-fluoren-9-yl)methyl (S)-( 1 -( 4-methoxyphenyl)-3-( ((4- nitrophenoxy)carbonyl)oxy)propan-2-yl)carbamate (Fmoc-TyrOMe-PNOC). According to the general procedure for synthesis of activated carbonates: Crude Fmoc-TyrOMe-ol (~4.8 mmol) gave a crude product that was purified by silica gel chromatography (DCM then DCM with 1-5% EtOAc) and isolated as flakey fluffy white solid (2.05 g. 75% yield over 2 steps). ‘H NMR (500 MHz, Chloroform-i/) 8 8.26 (d, J = 9.2 Hz, 2H), 7.77 (d, J= 7.5 Hz, 2H), 7.54 (t, J= 6.5 Hz, 2H), 7.42 - 7.36 (m, 4H), 7.30 (t, J= 7.5 Hz, 2H), 7.12 (d, J= 7.9 Hz, 2H), 6.86 (d, J= 8.0 Hz, 2H), 4.87 (d, J= 7.7 Hz, 1H), 4.50 - 4. 15 (m. 6H), 3.78 (s, 3H), 2.90- 2.80 (m, 2H). 13C NMR (126 MHz, CDCh) 8 158.66, 155.77, 155.38, 152.39, 145.49, 143.77, 143.75, 141.34, 130.15, 128.25, 127.75, 127.06, 125.34, 124.94, 121.71, 120.02, 114.26, 69.49, 66.82, 55.27, 51.15, 47.24, 36.56. HRMS +ESI [M+Na]: 591.1734, calculated (C32H28N2O8): 591.1738.
(9H-fluoren-9-yl)methyl (S)-(l-(4-bromophenyl)-3-(((4- nitrophenoxy)carbonyl)oxy)propan-2-yl)carbamate (Fmoc-4BrPhe-PNOC). According to the general procedure for synthesis of activated carbonates: Crude Fmoc-4Br-Phe-ol (2.14 mmol) gave a crude product that was purified by silica gel chromatography (9:1 DCM:Hex with 1% EtOAc) and isolated as a white flaky solid (0.85 g, 83% yield over 2 steps). ’H NMR (500 MHz, Chloroform-d) 8 8.26 (d, J= 8.8 Hz, 2H). 7.77 (d, J= 7.6 Hz, 2H). 7.54 (dd. J= 7.7, 2H), 7.49 - 7.33 (m, 6H), 7.31 (t, J= 7.6 Hz, 2H), 7.08 (d, J= 7.9 Hz, 2H), 4.89 (d, J= 8.1 Hz, 1H), 4.51- 4.37 (m, 2H), 4.37 - 4.05 (m, 4H), 2.88 (s, 2H). 13C NMR (126 MHz, CDCI3) 5 155.72, 155.30, 152.33, 145.52, 143.71, 143.67, 141.36, 135.44, 131.94, 130.82, 127.79, 127.08, 125.36, 124.87, 121.68, 121.03, 120.05, 69.41, 66.78, 50.93, 47.22, 36.89. HRMS +ESI [M+K]: 655.0473, calculated (CsiIfeBrlW?): 655.0477.
(9H-fluoren-9-yl)methyl (S)-(l-(3, 5 -dichlorophenyl)- 3 -( ((4- nitrophenoxy)carbonyl)oxy)propan-2-yl)carbamate (Fmoc-3, 5DiCl-Phe-PNOC). According to the general procedure for synthesis of activated carbonates: Crude Fmoc-3,5DiCl-Phe-ol (~2.2 mmol) gave a crude product that was purified by silica gel chromatography (3:1 DCM Hex, then 2: 1, then 1: 1, then DCM, then DCM w/ 3% EtOAc) and isolated as flaky fluffy white solid (0.91g. 68% yield over 2 steps). 'H NMR (500 MHz, Chloroform- ) 5 8.27 (d, J = 9.1 Hz, 2H), 7.77 (d, J= 7.5 Hz, 2H), 7.53 (d, J= 7.6 Hz, 2H), 7.45 - 7.33 (m, 5H), 7.32-7.27 (m, 2H), 7.23 - 7.13 (m, 2H), 4.97 (d, J= 8.2 Hz, 1H), 4.45 - 4.25 (m, 4H), 4.17 (t, J= 6.6 Hz, 1H), 3.11 - 2.97 (m, 2H). 13C NMR (126 MHz, Chloroform- ) 5 155.67, 155.31, 152.33, 145.53. 143.70, 141.35, 134.94, 133.67, 133.19. 131.96, 129.63, 127.78, 127.50, 127.06. 125.36. 124.89, 121.70, 120.04, 69.72, 66.83. 50.33, 47.18. 34.42. HRMS +ESI [M+K]: 645.0584, calculated (C31H24CI2N2O7): 645.0592.
(9H-fluoren-9-yl)methyl (S)-( 1 -( 3, 5-dibromo-4-methoxyphenyl)-3-( ((4- nitrophenoxy)carbonyl)oxy)propan-2-yl)carbamate (Fmoc-D2TyrOMe-PNOC). According to the general procedure for synthesis of activated carbonates: Crude Fmoc-D2TyrOMe-ol (4.8 mmol) gave a crude product that was purified by silica gel chromatography (DCM then DCM with 1-5% EtOAc) and isolated as flaky fluffy white solid (2.05 g, 66% yield over 2 steps). ’H NMR (500 MHz, Chloroform- ) 5 8.26 (d, J = 9.1, 2H), 7.79 (d, J= 7.5 Hz, 2H), 7.57 (t, J = 6.5 Hz. 2H), 7.45 - 7.36 (m, 4H), 7.32 (t, J= 7.5 Hz, 2H), 7. 14 (d, J= 8. 1 Hz, 2H), 6.88 (d, J= 8.0 Hz, 2H), 4.90 (d, J= 8.6 Hz, 1H), 4.44 (dt, J= 32.3, 10.2 Hz, 2H), 4.26 - 4.16 (m, 2H), 3.81 (s, 3H), 2.94 - 2.81 (m, 2H). 13C NMR (126 MHz, CDCI3) 5 158.66, 155.78, 155.38, 152.40, 145.48, 143.78, 143.76, 141.34, 130.15, 128.25, 127.75, 127.06, 125.33, 124.94. 121.71, 120.03, 114.26, 66.81, 55.27, 51.02, 47.24, 36.52. HRMS +ESI [M+Na]: 593.1863, calculated (C32H26D2N2O8): 593.1863
(9H-fluoren-9-yl)methyl (S)-(l-methoxy-3-(((4-nitrophenoxy)carbonyl)oxy)propan-2- yl-3, 3-d2)carbamate (Fmoc-PFSer-PNOC). According to the general procedure for synthesis of activated carbonates: Crude Fmoc-D2SerOMe-ol (4.8 mmol) ) gave a crude product that was purified by silica gel chromatography (3: 1 DCM:Hex) and isolated as an off white solid (2.08 g, 88% yield over 2 steps). 'H NMR (500 MHz, Chloroform^/) 5 8.25 (d, J= 9.1 Hz, 2H), 7.77 (d, J= 7.5 Hz, 2H), 7.59 (d, J= 7.4 Hz, 2H), 7.44 - 7.34 (m, 4H), 7.31 (td, J= 7.4, 1.0 Hz, 2H), 5.24 (d, J= 8.8 Hz, 1H), 4.52 - 4.36 (m, 2H), 4.26 - 4.16 (m, 2H), 3.62 - 3.46 (m, 2H), 3.39 (s, 3H). 13C NMR (126 MHz, CDCh) 5 156.04, 155.45, 152.38, 145.45, 143.78, 141.34, 127.76, 127.08, 125.30, 124.99, 121.76, 120.04, 70.95, 67.01, 59.28, 49.32, 47.22. HRMS +ESI [M~Na]: 517.1549. calculated (C26H22D2N2O8): 517.1550.
(9H-fluoren-9-yl)methyl (S)-( I -( 3, 5-dibromo-4-hydroxyphenyl)-3-hydroxypropan-2- yl)carbamate (Fmoc-3,5-DiBrTyr-ol). Synthesis of the dibromo isotope tag is shown in Scheme 11.
Scheme 11. Synthesis of Dibromo isotope tag
To a flask containing 250 mgs of Fmoc-3,5-Dibromotyrosine-OH (0.446 mmol), THF (1.5 mL) was added followed by CDI (187.5 mg, 0.867 mmol)) and allowed to stir for 1 hour. The reaction was cooled to 0 °C in an ice bath. NaBH4 (67 mg, 0.564 mmol) dissolved in water (1.0 mL) was added drop wise to the reaction and the reaction stirred stir for an additional hour. The reaction was quenched with 1 M HC1 and diluted with EtOAc. The reaction mixture was extracted with EtOAc (3 x 50 mL), brine (1 x 50 mL), and dried with anhydrous MgSO4and concentrated in vacuo. The crude mixture was purified by column chromatography (Silica, 1% MeOH in DCM, 212 mg 87% yield). 'H NMR (500 MHz. CDCh) 5 7.77 (d, J= 7.5 Hz, 2H), 7.56 (t, J= 6.8 Hz, 2H), 7.40 (t, J = 7.5 Hz, 2H), 7.35 - 7.29 (m, 4H), 5.80 (s, 1H), 4.97 (s, 1H), 4.46 - 4.39 (m, 2H), 4.21 (t, J= 6.5 Hz, 1H), 3.84 (s, 1H), 3.72 - 3.52 (s, 2H), 2.77 (s, 2H). 1?C NMR (126 MHz, CDCh) 5 148.06, 143.74, 143.68, 141.28. 132.51, 132.37, 127.63, 126.99, 124.87. 119.92, 109.75, 66.61. 63.15, 53.73. 47.21, 35.71. MS; ESI+ m/z (M+H)+ 546.0, calculated (C24H2iBr2NO4): 545.98
(9H-fluoren-9-yl)methyl (S)-( l-( 3, 5-dibromo-4-methoxyphenyl)-3-hydroxypropan-2- yl)carbamate (Fmoc-3,5-DiBrTyrOMe-ol). To a flame dried vial, Fmoc-3,5-DiBrTyr-ol (129 mgs, 0.23 mmol) and K2CO3 (98 mgs, 0.7 mmols) were added followed by acetone (1.7 mL) and lastly, Mel (73.4 pL, 1.17 mmols). The reaction was then heated to 35 °C for 30 minutes. The reaction was quenched with 0. 1 M NaS2Ch and extracted with 1 M HC1 (3 x 50 mL) followed by brine (1 x 50 mL). The reaction was carried forward without purification. MS; ESI+ m/z (M+H)+ 562.0, calculated (^sffeBnNCU): 562.01
(9H-fluoren-9-yl)methyl (S)-( 1 -( 3, 5-dibromo-4-methoxyphenyl)-3-( ((4- nitrophenoxy)carbonyl)oxy)propan-2-yl)carbamate (Fmoc-3,5-DiBrTyrOMe-PNOC) . To a vial of crude Fmoc-3,5-DiBrTyrOMe-ol (77 mg, 0.137 mmol), DCM (700 pL) was added, followed by pyridine (16.6 mg, 0.206 mmol) and then 4-nitrophenylchloroformate (41.6 mgs, 0.206 mmol). The reaction was stirred for 4 hours and quenched with 1 M sodium bisulfate and diluted with DCM. The reaction mixture was washed with 1 M sodium bisulfate (3 x 20 mL) followed by 1 M Na2CO3 (3 x 20 mL) and then brine (1 x 20 mL). The crude material was purified via silica column chromatography in pure DCM (34.8 mg, 35% yield over 2- steps). NMR (500 MHz, CDCh) 5 8.27 (d, J= 9.1 Hz, 2H), 7.77 (d, J= 7.6 Hz, 2H), 7.55 (t, J= 8.5 Hz, 2H), 7.44 - 7.26 (m, 6H), 7.36 - 7.27 (m, 2H), 4.99 (s, 1H), 4.52 - 4.45 (m, 1H), 4.44 - 4.36 (m, 1H), 4.35 - 4.30 (m, 1H), 4.28 - 4.18 (m, 3H), 3.85 (s, 3H). 2.91 - 2.81 (m. 2H). 13C NMR (126 MHZ. CDCh) 6 155.59, 155.13, 153.13, 152.18. 145.43. 143.56, 143.51, 141.23, 135.19, 133.07, 127.68, 126.99, 125.26, 124.81, 121.57, 119.94, 1 18.27, 69.11, 66.82, 60.50, 50.83, 47.09, 36.10, 29.57. HRMS (m/z): ESI calculated for C32H26Br2N2O8 [M+Na]+: 746.9948, obs. [M+Na]+: 746.9943.
Preparation of Isotopologue Monomer Ratios
To a vial was added Fmoc-D2Ala-PNOC (50.85 mg, 0.109 mmol) and Fmoc-Ala- PNOC (150.04 mg, 0.325 mmol). Next, 5.0 mL of dichloromethane was added, and the solution was mixed for one minute. The solvent was removed under reduced pressure and the sample dried overnight under vacuum. The isotopologue mixture was confirmed by high resolution mass spectrometry.
To a vial was added Fmoc-D2CHA-PNOC (100.95 mg, 0. 184 mmol) and Fmoc-CHA- PNOC (100.26 mg, 0.184 mmol). Next, 5.0 mL of dichloromethane was added, and the solution was mixed for one minute. The solvent was removed under reduced pressure and the sample dried overnight under vacuum. The isotopologue mixture was confirmed by high resolution mass spectrometry.
To a vial was added Fmoc-D2Val-PNOC (151.42 mg, 0.307 mmol) and Fmoc-Val- PNOC (50.35 mg, 0.103 mmol). Next. 5.0 mL of dichloromethane was added, and the solution was mixed for one minute. The solvent was removed under reduced pressure and the sample dried overnight under vacuum. The isotopologue mixture was confirmed by high resolution mass spectrometry.
CHA and D2CHA for synthesis of B2
To a vial was added Fmoc-D2CHA-PNOC (15. 15 mg. 0.0278 mmol) and Fmoc-CHA- PNOC 84.99 mg, 0.1567 mmol). Next, 5.0 mL of dichloromethane was added, and the solution was mixed for one minute. The solvent was removed under reduced pressure and the sample dried overnight under vacuum. The isotopologue mixture was confirmed by high resolution mass spectrometry’.
To a vial was added Fmoc-D2Val-PNOC (67. 16 mg, 0. 1364 mmol) and Fmoc-Val- PNOC (33.36 mg, 0.068 mmol). Next, 5.0 mL of dichloromethane was added, and the solution was mixed for one minute. The solvent was removed under reduced pressure and the sample dried overnight under vacuum. The isotopologue mixture was confirmed by high resolution mass spectrometry.
To a vial was added Fmoc-D2Ala-PNOC (84.16 mg, 0.1812 mmol) and Fmoc-Ala-
PNOC (16.31 mg, 0.0353 mmol). Next, 5.0 mL of dichloromethane was added, and the solution was mixed for one minute. The solvent was removed under reduced pressure and the sample dried overnight under vacuum. The isotopologue mixture was confirmed by high resolution mass spectrometry.
Parallel Solid-phase Synthesis in 96-well plate
General procedure for the solid-phase synthesis ofBl-B8
Coupling: Phenyalaninol loaded (0.44 mmol/gram, 200-400 mesh) 2-chlorotrityl polystyrene resin (20 mg, 0.088 mmol) was added to each of the 8 wells (B1-B8) of a 10 um fritted deep-well 96-well plate (combinatorial microlute by Porvair). A 96-well microplate deep well drain cap mat (cat. No. 219005, Porvair Sciences) was used to seal the base of the filtration plate. The resin was suspended in 0.2 mL of anhydrous N-methyl-2-pyrrolidinone (NMP) and swollen. A stock solution of Hunig’s base (21.2 pL total, 2.6 pL or 0.015 mmol per well) and hydroxybenzotriazole (64.8 mg total, 8.1 mg or 0.06 mmol per well) was made in NMP (0.8 mL total, 0.1 mL/well). Finally, the activated amino alcohols Fmoc-XXX- PNOC (0.030 mmol per well) were dissolved in NMP (0.2 mL per well) in separate vials. Once dissolved, the monomers were transferred to their corresponding wells. The plate was sealed with Nunc aluminum sealing tape (no. 276014) and shaken overnight. NOTE: proper mixing of the wells is important for efficient coupling, and previous reports show the coupling to be complete in as little as 4 hours.
Resin in all 8 wells were washed with NMP (5x2 mL), then DCM (5x2 mL), and finally Et20 (3x2 mL) by vacuum filtration through the fritted plate (Supelco PlatePrep 96- well Vacuum Manifold kit). Resin was dried overnight under vacuum. Test cleavages were affected with 1% TFA in DCM (5x1 mL) for 10 seconds each. Coupling efficiency was checked by LC/MS. NOTE: depending on the cleavage times and amounts of TFA used, the trifluoroacetic ester was observed by LC/MS. This ester was readily hydrolyzed dissolving the sample in DCM and shaking with saturated NaHCO
Deprotection: Resin loaded with terminal Fmoc-protected oligocarbamates (20 mg per well) was suspended in 20% piperidine in DMF (0.5 mL) and shaken for 2 hours. NOTE: Due to shorter reaction time, sufficient mixing of the well is required. Incomplete deprotections were observed with inadequate mixing. Cleavage of the dibenzofulvenepiperidine adduct can quantified at 301 nM using Beer’s Law. The resin was washed with DMF (5x2 mL), DCM (5x 2 mL), Et20 (3x2 mL). and dried overnight under vacuum.
General Procedure for labelling the terminal amine with NBD-Fluoride (on resin). In the fritted 96-well plate, to each of the 8 wells containing 0.01 mmol oligomer loaded resin was added a solution of Hunig’s base (17 uL/ well, 0.01 mmol), 4-Fluoro-7- nitrobenzofurazon (9.1 mg or 0.05 mmol per well) in anhydrous DMF (0.5 mL per well). The plate was sealed and the reactions were left to shake overnight. Resin was washed with DMF (5x2 mL), DCM (10x2 mL, or no more yellow/green was observed in the wash), and Et20 (3x5 mL).
Cleavage procedure: Cleavages were affected with 0.1% TFA in DCM (4x2 mL) for 15-30 seconds each (until no color observed upon addition of cleavage cocktail). Resin was filtered off. and cleaved product was concentrated in a 6-dram vial. Due to cleavage times and amounts of TFA used, the trifluoroacetic ester of the terminal alcohol was observed by LC/MS.
General Procedure for hydrolysis of TFA Ester: Prior to HPLC purification, each oligomer was worked up individually. Example work up: Oligomer Bl was dissolved in di chloromethane (20 mL) and transferred to a separatory funnel. To the separatory funnel was added 20 mL of sat. NaHCOs. The separatory funnel was adequately shaken for 3-5 minutes and the phases allowed to separate. The aqueous layer was extracted 3x with DCM (10 mL). The combined organics were washed lx with brine, dried over Na2SO4, filtered, and concentrated in vacuo.
General HPLC condition for B1-B8. Oligomers were dissolved in isopropanol (1-2 mL) and purified by reverse-phase HPLC chromatography (gradient elution, 30%-95% MeCN/HzO). Due to the hydrophobic character of the oligomers, the column was primed with at least 30% organic prior to injection.
Sequencing Experiments
Sequencing of Bl, B2, B3, B4, B5, B6, B7, B8 Simultaneously. The oligomers, B1-B8 (500 nmols of each) were dissolved in DMSO (1.2 mL) and added to a 2-dram vial. Next, cesium carbonate (4.08 mg) was added to the solution. The final concentration of base was approximately 0.01M. The reaction was heated to 70 °C, and held at the temperature for 45 minutes. Reaction was sampled for LC/MS by taking 50 pL of the reaction. This was repeated at 90, 150, 210, 270, 330, 420, and 550 minutes. The time points were placed into separate wells of a 96-well plate and dried under centrifugal vacuum evaporation (Genevac). Once dry, each sample was dissolved in 75 pL of MeOH and analyzed by LC/MS (5-95% MeCN/H2O with 50 mM ammonium acetate, negative mode).
Separately, 50 pL of the reaction was taken for each timepoint and combined. The combined sample was dried under centrifugal vacuum evaporation. The combined sample was dissolved in 400 pL of MeOH and analyzed by LC/MS (5-95% MeCN/I-bO with 50 mM ammonium acetate, negative mode).
References
1. Boukis AC et al. European Polymer Journal 2018, 104, 32-38.
2. Vk N et al. International Journal of Chemical Research 2009, 1 (2), 1-7.
3. Bornholt J et al. In A DNA-Based Archival Storage System, 2016; ACM.
4. Kaiser R et al. Nature 1991, 350 (6320), 656-657.
5. Schaefer C et al. BMC Genomics 2012, 13 (Suppl 4), S4.
6. Pertea M et al. Thousands of large-scale RNA sequencing experiments yield a comprehensive new human gene list and reveal extensive transcriptional noise. Cold Spring Harbor Laboratory: 2018.
7. Hill DJ et al. Chemical Reviews 2001, 101 (12), 3893-4012.
8. Breslow R. Journal of Biological Chemistry 2009, 284 (3), 1337-1342.
9. Solleder SC et al. Macromolecular Rapid Commun. 2017, 38 (9). 1600711.
10. Martinek TA et al. Chem. Soc. Rev. 2012, 41 (2), 687-702.
11. Mutlu H et al. Chem. , Int. Ed. 2014, 53 (48), 13010.
12. Cottrell JS. J. Proteomics 2011, 74 (10), 1842.
13. Solleder SC et al. Macromol. Rapid Commun. 2017, 38 (9), 1600711.
14. Proulx C et al. Biopolymers 2016, 106 (5), 726.
15. Zuckermann RN et al. J. Am. Chem. Soc. 1992, 114 (26), 10646.
16. Cho C et al. Science 1993, 261 (5126), 1303.
17. Hill DJ et al. Chem. Rev. 2001, 101 (12). 3893.
18. Martinek TA et al. Chem. Soc. Rev. 2012, 41 (2), 687.
19. Randall JC. Polymer Sequence Determination. 1977; p 1.
20. Gruendling T et al. Polym. Chem. 2010, 1 (5), 599.
21. Montaudo G et al. Prog. Polym. Sci. 2006, 31 (3), 277.
22. Pan G et al. J. Appl. Polym. Sci. 2004, 93 (2), 577.
23. Smith D A et al . Br. Polym. J. 1976, 8 (4), 101.
24. McCormick H. Journal of Chromatography A 1969, 40, 1.
25. Kim H et al. Angew. Chem., Int. Ed. 2015, 54 (44), 13063.
26. Peterson GI et al. Macromolecules 2012, 45 (18), 7317.
27. Sagi A et al. J. Am. Chem. Soc. 2008, 130 (16), 5434.
28. Erez R et al. Org. Biomol. Chem. 2008, 6 (15), 2669. 29. McBride RA et al. Macromolecules 2013, 46 ( 13), 5157.
30. Zhang LJ et al. Macromolecules 2013, 46 (24), 9554.
31. Kirby AJ et al. Advances in Physical Organic Chemistry. 1980; Vol. 17, p 183.
32. Beitia JLL. 5Weri 2011, 2011 (01), 139-140.
33. Sun-Hee Hwang et al. The Open Organic Chemistry Journal 2008, 2, 107- 109.
34. Rutten M et al. Nature Reviews Chemistry 2018, 2 (11), 365-381.
35. Colquhoun H et al. Nature Chemistry 2014, 6 (6), 455-456.
36. Lutz JF. Macromolecular Rapid Communications 2017, 38 (24), 1700582.
37. Lutz JF et al. Science 2013, 341 (6146), 1238149.
38. Ceze L et al. Nature Reviews Genetics 2019, 20 (8), 456-466.
39. Church GM et al. Science 2012, 337 (6102), 1628-1628.
40. Goldman N et al. Nature 2013, 494 (7435), 77-80.
41. Zhirnov V et al. Nature Materials 2016, 15 (4), 366-370.
42. Bornholt J et al. A DNA-Based Archival Storage System. In Proceedings of the Twenty -First International Conference on Architectural Support for Programming Languages and Operating Systems, Association for Computing Machinery: Atlanta, Georgia, USA, 2016; pp 637-649.
43. Behjati S et al. Arch Dis Child Educ Pract Ed 2013, 98 (6), 236-238.
44. Branton D et al. Nature Biotechnology 2008, 26 (10), 1146-1153.
45. Ying YL et al. Angewandte Chemie Int. Ed. 2013, 52 (50), 13154-13161.
46. Al Ouahabi A et al. JACS 2015, 137 (16), 5629-5635.
47. Cavallo G et al. JACS 2016, 138 (30), 9417-9420.
48. Ding K et al. European Polymer Journal 2019, 11 . 421-425.
49. Gunay Ufuk S et al. Chem 2016, 1 (1), 114-126.
50. Lee JM et al. Nature Communications 2020, 11 (1), 56.
51. Leguizamon SC et al. Nature Communications 2020, 11 (1), 784.
52. Liu B et al. Polymer Chemistry 2020, 11 (10), 1702-1707.
53. Martens S et al. Nature Communications 2018, 9 (1), 4451.
54. Roy RK et al. Nature Communications 2015, 6 (1), 7237.
55. Aksakal R et al. Advanced Science 2021, 8 (6), 2004038.
56. Frolich M et al. Communications Chemistry 2020, 3 (1), 184. 57. Laurent E et al. Macromolecules 2020, 53 (10), 4022-4029.
58. Soete M et al. Angewandte Chemie International Edition n/a (n/a), e202116718.
59. Song B et al. JACS 2022, 144 (4), 1672-1680.
60. Wetzel KS et al. Communications Chemistry 2020, 3 (1), 63.
61. Arcadia CE et al. Nature Communications 2020, 11 (1), 691.
62. Boukis AC et al. Nature Communications 2018, 9 (1), 1439.
63. Cafferty BJ et al. ACS Central Science 2019, 5 (5), 911-916.
64. La Clair JJ. Chemical Communications 2018, 54 (21), 2611-2614.
65. Sarkar T et al. Nature Communications 2016, 7 (1), 11374.
66. Mutlu H et al. Angewandte Chemie Int. Ed. 2014, 53 (48), 13010-13019.
67. Charles L et al. Macromolecules 2015, 48 (13), 4319-4328.
68. Al Ouahabi A et al. Nature Communications 2017, 8 ( 1 ), 967.
69. Boukhet M et al. Macromol. Rapid Commun. 2017, 38 (24). 1700680.
70. Cao C et al. Science Advances 2020, 6 (50), eabc2661 .
71. Dahlhauser SD et al. Cell Reports Physical Science 2021, 2 (4), 100393.
72. Langbridge JA. Professional Embedded ARM Development. 1st ed.: Wrox Press Ltd.: Birmingham. UK, 2014.
73. Kemighan BW et al. The C Programming Language. Prentice Hall Professional Technical Reference: 1988.
74. Albrecht B. My computer likes me when I speak in BASIC. Dilithium: Portland, 1972.
75. Huffman DA. Proceedings of the IRE 1952, 40 (9), 1098-1101.
76. M. Morris Mano, C. R. K., Logic and Computer Design Fundamentals, 4th Edition. Pearson: 2008.
77. Dahlhauser SD et al. JACS 2020, 142 (6), 2744-2749.
78. Kirby AJ. Effective Molarities for Intramolecular Reactions. In Advances in Physical Organic Chemistry, Gold, V.; Bethell, D., Eds. Academic Press: 1980; Vol. 17, pp 183-278.
79. Steinkoenig J et al. European Polymer Journal 2019, 120, 109260.
80. Nagy L et al. International Journal of Molecular Sciences 2020, 21 (4). 1318.
81. Chahrour O et al. J Pharm. and Biomed. Anal. 2015, 113, 2-20.
82. Paul D et al. Rapid Comm, in Mass Spectrometry 2007, 21 (18), 3006-3014. 83. Wieser ME. Pure and Applied Chemistry 2006, 78 (11), 2051-2066.
84. Kaufmann R et al. Nature 1984, 309 (5966), 338-340.
85. Seeman JI et al. J Chem. Soc., Chem. Comm.1992, (9), 713-714.
86. Miller FP et al. Advanced Encryption Standard. Alpha Press: 2009.
87. Staff CU. The Unicode Standard: Worldwide Character Encoding. Addison- Wesley Longman Publishing Co., Inc.: 1991.
88. Konermann L. J Am Soc Mass Spectrom 2017, 28 (9), 1827-1835.
89. Smith RM. Understanding Mass Spectra: A Basic Approach. 1998.
90. Rosman KJR et al. Pure and Applied Chemistry 1998, 70 (1), 217-235.
Example 2 - Real-Time Reading of Self-Sequencing Oligourethanes Using DESI
Introduction. Due to its speed, sensitivity, and specificity' [Al], mass spectrometry (MS) has found applications in an ever-expanding list of analytical techniques, ranging from medical diagnoses [A2], reaction screening [A3], sample characterization [A4], and many more. Both liquid chromatography/mass spectrometry (LC/MS) and high-resolution (Hi-Res) MS have been used to characterize the composition of complex reaction mixtures, including as the ability to distinguish between a host of isotopic patterns [A5, A6], However, these areas of MS analysis remain limited due to the required sample preparation. To be analyzed, the samples must be dissolved and diluted appropriately prior to analysis; especially if the solvents present in the sample are not compatible with the instrument’s mobile or solid phase. This is in addition to the sample processing time, which can range from about 25-45 minutes/sample [A7], Recent advances in MS have taken sample processing out of vials and columns and into ambient conditions.
Ambient Ionization Techniques. Desorption electrospray ionization (DESI) was developed by Professor Graham Cooks in 2004 as the first ambient ionization technique [A8, A9] . In this procedure, charged solvent droplets impact the surface where the analy te is mounted. Secondary charged droplets are then ejected from the surface and subsequently collected by a capillary inlet leading into the mass spectrometer (Figure 69) [A9], Since this breakthrough discovery in MS, over 30 methods have been developed to allow the analysis of samples at atmospheric pressure under ambient conditions [A10], One key feature of ambient ionization is the minimal processing that samples undergo prior to analysis [A8], No embedding in a matrix or dissolution of the analyte is necessary, other than the deposition or attachment of the sample onto a slide to be analyzed. With DESI-MS, the freely moving stage allows the solvent capillary and inlet to raster across the samples allowing specific areas of tissues [Al 1] or samples to be analyzed individually [A12], This precisely controlled movement allows the scanning speed to be optimized for different samples. In the case of screening alkenylation and azo-click reactions by the Cooks groups, speeds as fast as one second per reaction were achieved [A4], Additionally, it is possible to selectively desorb specific ions by varying the solvent spray composition allowing suppression of specific ions, creating a more easily interpretable spectrum. The ease of preparation, as well as strong ionization and rapid analysis speed, makes ambient ionization techniques such as DESI ideal for high-throughput analysis.
Optimizing Oligourethanes Synthesis for High-Throughput Workflows. Synthetic sequence-defined polymers are a promising means for molecular information storage. As discussed above, a workflow was developed and executed for the encoding and decoding of information in self-sequencing, macromolecular, sequence-defined oligourethanes. Subsequently, as also discussed above, the information density was increased by isotopically labelling eight oligomers to allow the precursor and truncated oligomers to be distinguished in one mixture (96 unique molecules). To continue to increase storage capacity and the complexity of these macromolecules, there remains two bottlenecks in the workflow: the oligourethane synthesis and subsequent LC/MS data acquisition and interpretation. Notwithstanding the robustness of the synthesis, preparing and individually distributing monomers by hand for multiple rows or even an entire well plate of reactions is a timeconsuming process. With successes having been made in reaction parallelization, the next was to automation. Through automation, it would be possible to dramatically increase the ease and throughput of synthesis, especially as more information is encoded into larger quantities of oligomers. Additionally, the LC/MS sample processing and data analysis of mass spectral information can take hours or days. This is not only because of acquisition time (as each LC/MS run takes ~25 minutes) but also analysis of the data due to each spectrum being parsed manually, with the researcher having to identify each precursor ion. To address these issues, the use of an automated peptide synthesizer (Figure 70) for the synthesis and a DESI-MS (Figure 69) for the analysis of the information-bearing oligourethanes were explored.
Optimization ofMultiPep 2 for Automated Oligourethane Synthesis. To tackle this first objective, clearing the synthetic bottleneck, the use of an automated peptide synthesizer was explored, CEM’s MultiPep 2 (Figure 70). Since its dissemination by Bruce Merrifield in 1965 [Al 3], solid phase synthesis has been optimized over decades with the advent of new resins [A14], protecting groups [A15], and instrumentation to facilitate efficient and rapid coupling reactions [Al 6] . Solid-phase peptide synthesis (SPPS) has advanced in leaps and bounds and laid the foundation for new sequence-defined oligomers and polymers to take form. The synthesis of a variety of oligomeric backbones using solid phase synthesis has been demonstrated [A6, A 17] and similar protecting groups are used across amino acid derivatives [A5, Al 8], After using such analogous techniques for the manual solid phase synthesis of the oligourethanes, an aim was to create the oligourethanes via automated synthesis as well. The MultiPep 2 was considered as a candidate for automated-oligourethane synthesis as it had been designed with all of the tools to execute SPPS (e.g. the capabilities to rinse resins, add solvent, extract solvent, agitate solutions, etc.), but it could also be adapted to other synthetic procedures.
Creating an Automated Workflow. Having been designed to create the oligourethanes in an automated manner with the MultiPep, as well as rapidly analyze samples with DES1-MS, this platform would be transformative for the use of sequence-defined polymers as a medium of information storage. Being able to synthesize sequence-defined oligourethanes via solid-phase synthesis on scales otherwise practically unachievable and able to analyze the readout in a few short minutes rather than hours brings this chemistry one step closer as a practical alternative to modem day hard drives. Here, the feasibility of DESIMS as an automated reading platform is demonstrated.
Testing DESI-MS with Oligourethanes
Combining Sequencing Time Points. To test the efficacy of using DESI-MS to sequence and analyze sequence-defined oligourethanes, we started with replicating results from our previous work on encoding a passage from Jane Austen into oligomers from above [ A5] . These 4-Fluoro-7-nitrobenzofurazan (NBD) labelled oligomers were tested initially in negative mode on the DESI-MS. Previously, time points would be analyzed individually, as the UV-Vis trace at 470 nm (the /.max of NBD) shows oligo species gradually accumulating and disappearing as the sequencing progresses. This visual analysis provided additional information in helping decipher the oligo sequence (Figure 71). Although this additional information proved useful, the ionization of the truncated species was often so prominent in the total ion chromatogram (TIC), the UV-Vis trace was not always necessary'. Using DESI- MS which has no chromatographic components precludes the need for keeping time points separate. As seen in Figure 72, at a single time point, four of the ten truncated species could be identified. How ever, in Figure 71, five spectra w ere acquired and used to identify the oligos in a reaction mixture despite many of the sequenced species being present across multiple runs. It was postulated that combining the time points would have an additive effect and would prevent redundant sample analysis.
Analyzing the Effect of Solvent Systems on Ionization. To begin the analysis, previously synthesized samples encoding a passage from Jane Austen, starting with sample G5 (Figure 73), were used. Glass slides were used to analyze the oligomer samples that had a polytetrafluoroethylene (PTFE) mat top with shallow circles cut out to ensure the contents of the individual wells stay isolated in one area on the slide (Figure 74). Samples were analyzed using the combined time points in a variety of solvents spray after being dissolved in methanol and deposited on the slide, then left to dry. As shown in Figure 73, while some solvent enhanced the signal/noise (S/N), as in the case with MeOH FhO 98:2 + 2mM ammonium acetate (A. A), IPA seemed to severely suppress the ionization of the sample (Figure 75). This cursory solvent analysis confirmed the importance of the solvent spray composition on the relative abundance of precursor ions observed in the analysis. Additionally, replicate spots of the same mixture were analyzed to ensure consistency across wells. Scanning replicate spots across plates showed nearly identical results, verifying the reliability of this ionization technique with the oligourethanes (Figure 76-Figure 77).
Testing New End Caps
Moving Away From 4-Fluoro-7-nitrobenzofurazan (NBD). Having verified that the oligomer masses could be observed and identified without the use of a UV-Vis trace, and that they ionized well under ambient conditions, next it was sought to optimize the N-terminal cap of the oligourethanes. This addition to the oligourethane chain served a variety of purposes. First, and perhaps most importantly, it prevented unforeseen side reactions (e.g. cross reactivity betw een molecules, cyclization to a urea, etc.) that would have resulted in an undecipherable series of truncated masses as the two ends sequenced at different rates. Second, it greatly enhanced the ionization of the oligourethanes in the negative mode, giving strong ionization w ith minimal adducts in the presence of ammonia acetate and acetonitrile. Finally, it provided a strong absorbance at 470 nm allowing the progress of the sequencing to be monitored visually (a necessary' feature w hen troubleshooting and developing this sequencing technology). While the first tw o features are important to this w orkflow', as stated above, the DESI has no chromatographic or UV/Vis component, rendering the chromophoricity of NBD unnecessary. Additionally, NBD-F is an expensive reagent (1 mg for $45), so replacing this N-terminal label with a more cost-effective cap that can accomplish equal, if not more efficient capping and ionization enhancement, would be preferred as synthesis is scaled up.
New End-Cap Designs. Caps were designed to append to oligomers synthesized by the Multipep that would aid the analysis beyond what NBD has provided. Initially, four caps were selected to test, each possessing a different property of interest: a unique isotope pattern or ionizable functionality. It was hoped that by favoring a negative charge on the oligomer, as in the case of using Gly and Tyr derivatives (Figure 78), the samples would be inherently biased towards the negative mode with enhanced signals on the mass spectrum. This would theoretically increase the limits of detection for the sample, requiring less material for successful analysis. Also synthesized were caps with halogens, inspired by previous work. With the halogen, particularly the Tri-Cl species, distinct isotope patterns would indicate an oligomer species was present, making identifying noise versus precursor ion peaks much simpler.
Optimization of Ambient Ionization. One obstacle encountered, which has been documented in the DESI community', is the affect ambient conditions have on this ambient ionization technique [Al 9], It was found that below a relative humidity of -35% , inconsistencies in data were found while using negative mode. It is speculated that this is due to the increased electrostatics found when water vapor content in the air is lower. This phenomenon could be remedied by ensuring DESI experiments are conducted in a humidity- controlled environment. Interestingly, these fluctuations were not found in similar tests while using the positive mode to analyze samples [Al 9], Having data acquisition be limited by weather or the time of year would greatly hinder its practical utility. In past work, samples were solely analyzed in negative mode due to the NBD-capped oligos favored ionization over what is observed in positive mode; as well as the decreased number of ion-paired adducts when ammonium acetate is present in the mobile phase. It was reasoned that with the development of new ionizable functionalities, the positive mode ionization results would vary less drastically, and the analysis would likely be more consistent, despite positive mode often creating a variety’ of adducts (e.g. M+K, M+Na). After several weeks of inconsistent results due to fluctuating negative mode conditions, test caps were analyzed solely in positive mode.
First Round of Positive-Mode Analysis. The subsequent analysis of the four unique N-terminal ionizable functionalities yielded very promising results. Figure 79 shows that it is possible to identify the majority of the species in solution regardless of the cap. However, Tyr gave the by far the cleanest spectra, with the truncated precursor ions observable well above the noise. Additionally, minimal adducts are found alongside the desired products. Gly gave the most adducts and was certainly the least intuitive spectra to deconvolute. Mono-Cl and Tri-Cl gave very distinct isotope patterns; however, it appeared as though minor degradation of the trichloro functional group occurred during the sequencing. Additionally, not all sequenced species could be identified in the spectrum, and many adducts were present in Mono-Cl and Tri-Cl. While Tyr was observed as the most promising cap, these results inspired several additional caps to be devised to further improve the DESI-MS spectra.
Capping with Rhodamine-Based Monomers. A series of rhodamine-based derivatives were synthesized or acquired. As stated previously, it was aimed to bias the samples towards the positive mode with a pre-installed positive charge on the molecule, as well take advantage of the added mass these larger caps would provide (to avoid the noisiest region of ambient ionization, below 400 m/z) [A20], Two rhodamine-based moi eties were chosen, 5(6)-carboxytetramethylrhodamine (TAMRA) and Rhodamine B (RhoB) (Figure 80). RhoB has many benefits for DESI analysis. This compound has been well documented for is use in labelling peptides and peptoids [A21] and has a noted strong ionization under ambient conditions [A22], Additionally, Rhodamine B is an inexpensive reagent (1 g for $0.25) and the activation to the NHS-ester is a single synthetic step where the resulting product can be appended to the oligomer on the solid-phase, without isolation (Scheme 12 and Figure 81). Rhodamine B has a well-documented propensity to cyclize to its lactone form in the presence of base [A23], In this cyclized form, the fluorescence of Rhodamine B is quenched and the molecule is no longer positively charged; however, this event is highly reversible if acid or metal ions (e.g. Hg2+, Fe3+, etc.) are added to the solution (Figure 102). If this cyclization did impact the studies however, the molecule could be easily uncyclized by adding acid in a protic solvent (Figure 102). Thus, analysis proceed for the RhoB oligo.
Scheme 12. Synthetic route to creation of a Rhodamine B capped oligourethane.
The second oligomer, TAMRA, was synthesized from the commercially available NHS-5(6)-carboxytetramethylrhodamine. This material exists as the 5,6 isomer and is significantly more costly than Rhodamine B (1 g for $800) but the NHS activation of the carboxylic acid at the 5 or 6 position prevents the cyclization as the lactam. Although the free carboxylic acid may still cyclize to for the lactone, this is much more reversible and would not occur in protic solvents [A24], This molecule was initially considered as an end-cap for previous oligourethane work; however, as it exists as the 5,6 isomer, it was found that this confounded the UV-Vis analysis as each isomer would have different retention times on the LC (data not included). Since DESI will be used to analyze these samples, and these isomers have the same exact mass, it was felt this end cap had potential to circumvent obstacles previously mentioned with the Rhodamine B cap. The 5(6)-NHS- carboxytetramethylrhodamine was appended to the oligomer on the solid-phase in the same manner as Rhodamine B (Scheme 13 and Figure 82) and was confirmed via HRMS.
TAMRA
Scheme 13. Synthetic route to TAMRA oligo using NHS-5(6)-carboxytetramethylrhodamine.
Only the 5-isomer is shown in this scheme.
Rhodamine-based Oligomer Analysis. RhoB and TAMRA were sequenced and analyzed using DESI-MS. RhoB was left in its cyclized form for the first analysis. As seen in Figure 83, RhoB displayed strong ionization with only the [M+K] adduct seen. A small deletion can be seen at 715.327 m/z; however, the precursor ion signals of the desired oligomer were consistently stronger than that of the deletion (the deletion being an artifact of an incomplete coupling step during synthesis). TAMRA, on the other hand, had many adducts and peaks corresponding to the [M+H], [M+Na], and [M+K] adducts were not found (Figure 83). Similar to Gly, it appears as though when a free carboxylic acid is present on the oligomer, the spectrum becomes much more convoluted with adducts (unlike what was observed with previous TAMRA oligomers using LC/MS, likely attributed to the acidic mobile phase and chromatography). Having concluded the initial investigation into rhodamine-based caps, these results with RhoB and TAMRA were repeated and subsequently compared to Tyr, the most promising of the previous cap study.
Tyrosine, Rhodamine B, and TAMRA Comparison. New oligomers were synthesized and sequenced using the MultiPep to repeat the analysis of these caps (Figure 84- Figure 87). The sequencing experiment did not go to completion; however, it was felt that the data would still be able to elucidate the quality of the different caps. TAMRA-2 was analyzed using acetonitrile spray containing 0.2% formic acid (Figure 84). All but the 3-mer were immediately visible from the spectra; however, the oligomers existed as [M+H] and in the case of the 3-mer the [M+K] species. Additionally, the spectrum was still very convoluted from adducts, as observed in the 600 m/z to 1200 m/z range (Figure 84). Next, Tyr-2 was analyzed in ACN. Although it may seem as though the signals were quite small compared to the peak at 414.216 m/z (a commonly seen ambient ion), the precursor ions were observed with very impressive relative abundances and most importantly, minimal noise. If the sample had sequenced to completion, it would be a simple matter to pick out each of the oligos after removing the background ambient ions (including the peak 414.215 m/z). The only drawback to this cap is the relatively small mass when compared to that of rhodamine. Looking at Figure 76-Figure 77 and Figure 85, there are significantly more ambient ions that can appear at a m/z less than 500. Using DESI-MS, it is possible to filter out masses that weigh less than a certain amount by varying a parameter on the instrument know n as the S-lens value. A higher S-lens value favors detection of larger ions, so lowering the value theoretically favors smaller ions. However, since the tyrosine-based cap itself is small (208.10 m/z + Imer), that filter would not be able to be used, unlike with the rhodamine caps (RhoB: 441.24 m/z + Imer and TAMRA: 429.17 m/z + Imer). Despite its low molecular weight, Tyr-2’s strong S/N ratio makes it a very promising endcap.
Last was the analysis of RhoB-2. In this experiment, the cyclized and uncyclized versions were analyzed to compare any differences in ionization and S/N ratio, if any (Figure 86 and Figure 87). Starting with the lactam, all sequenced oligomer peaks, as well as the potassiated adducts, were again immediately observed (Figure 86). The peaks ionized very well again, although it should be noticed that the S/N ratio was much higher for Tyr-2. Nonetheless, all precursor ions corresponding to the sequenced oligomer could be picked out from the DESI spectrum. The uncyclized version was not as clean to analyze (Figure 87) perhaps due to additional adducts created from the TFA added to open the ring (Figure 102). Formic acid was added to protonate the free carboxylic acid on RhoB-2 in an attempt create fewer ionizable adducts. However, this instead created a variety of [M+H] and other adducts. While having the RhoB-2 uncyclized does give the solution a bright pink color, making it easy to visualize the spots on the DESI slide, the analysis was so poor it was felt that this was not reason enough to continue with opening the lactam of RhoB prior to analysis.
Automated Software Read-out
Raw File Reader. As previously mentioned, one bottleneck of the workflow up until this point has been the sample processing/data acquisition times and data analysis. While DESI-MS will significantly decrease the acquisition time, if a user must manually enter MS data into a template, this poses a large impediment in the workflow. One benefit to using DESI-MS is the ease to which the output files can be formatted. Previously, sequencing data was solely analyzed on Agilent instruments. Unlike Agilent instruments which output *.D files (a propriety Agilent format), DESI creates *.RAW files, an extension commonly used by Thermo Fisher Scientific and Waters instruments. These files have a standard data format that can be accessed using a .NET application made by Thermo Fisher Scientific and is open source and accessible online (Figure 88). Although there are other tools for analyzing *.RAW files [A25, A26] (e.g. proteowizard), they lack a command line based interface that is important for total automation. Software such as Proteowizard. which readily converts * RAW files and other file formats into more usable forms, operates using a graphical user interface (GUI), meaning the user must drag and drop files to the program and specify the output. Although this is a simple process, if real-time readout is to be achieved, this input and output process must be automated. Using RawFileReader.c, a C# application, which readily converts *.RAW files to *.json files or any other specified output in an automated fashion, it is possible to interpret files as soon as they are created. This application can scan a specified directory as files are populated into it from the DESI-MS software. These files are then converted into .json files, which are easily interpretable via python libraries such as pandas or json. Python software can be developed to automatedly find all oligomers and their respective precursor ions in a DESI-MS spectrum and return the encoded information in real time.
Python Spectra Analysis. With any DESI-MS experiment, due to the constant input from the spray and inlet to the detector, a multitude of scans are created every run. While some of these correspond to sample spots, others can be the area between the spots on the slide, as seen in Figure 89. This scan information must be parsed by the program to identify sample spots, particularly where one sample starts and ends, and the next sample begins. This was accomplished similarly to how the DESI-MS software analyzes scans (as seen in Figure 89) by reading the scans and measuring the average intensity of the scans (Figure 90). As seen in the scans on the left, there is a marked relative ionization increase when scanning over actual oligomer samples versus an empty slide. This is summarized in the summary plot to the right (Figure 90) and a threshold can be set (see the red line) to discard scans that do not cross this threshold. Now that specific scans corresponding to oligomer samples can be identified and analyzed, the next aim w as to identify precursor ion peaks within each individual spectrum.
Creation of Oligos from MultiPep Files. As larger quantities of oligourethanes are created using the MultiPep, a concomitant method is needed for visualizing these peptides for both disseminating work to others as well as for the user to analyze. Typically, these structures are created by hand in ChemDraw, however as the oligos get longer and side chains get more complex, these structures become more tedious to create in ChemDraw. Below is an example of a MultiPep file that one would enter into the MultiPep software to create the desired sequence. d2Cha Ser Ser d2Cha Ser d2Ser d2Leu d2Phe Ala ;#Phe
Phe d2Abu d2Cha Leu d2Cha Phe Ser Phe d2Abu ;#Phe d2Cha HoPhe Phe Ala Ser Ser Ala d2Ala Cha ;#Phe
Ala Cha Ser d2Abu Tyr Ser Abu Phe Ser ;#Phe d2HoPhe HoPhe d2Abu Leu d2Cha Tyr ;#Phe
These codes (Ser = Fmoc-Serine-PNOC) were arbitrarily created for each monomer but used consistently throughout different runs. To expedite the visualization of these molecules, a python script was created using a library previously used in oligo work known as RDKit to create these molecules using SMILES code. By simply entering in the MultiPep run codes into the program, it will create the entire strand as a SMILES code which can simply be pasted into ChemDraw to create the oligos from the run.
An example of the creation of the sequence: “Ala Leu Phe Ala #Phe” is shown in Figure 91 to Figure 94.
Similar to how protecting group chemistry works in chemistry, the end group was initially started with an alkylthiol rather than the rhodamine or tyrosine caps as they have a free OH group which this program looks for as the next end. The software iterates through the string containing the monomer order and appends the monomers one by one to the growing strand of oligos. After the final monomer, the program will then add the specified end cap by replacing the thiol group.
General Procedure. All materials used in the synthesis of each compound and related tests, were purchased from Sigma-Aldrich Chemical Co., Acres Organics, Toky o Chemical Industry. Chem Impex International, etc. and used without further purification. Solvents (DCM, NMP, chloroform, DMSO. DMF, MeOH, MeCN, Isopropanol) were of reagent grade or HPLC grade quality and purchased from Fischer Scientific. NMR solvents (CDCh, DMSO-d6) were purchased from Cambridge Isotope Laboratories.
Column chromatography was performed using silica gel 60 (230 ± 400 mesh. 0.040 ± 0.063 mm) from Dynamic Adsorbents.
TLC analyses were carried out using Silica TLC Plates Aluminium Backing 20 by 20 cm sheet UV active at 254 nm.
XH and 13C spectra were recorded on Varian DirectDrive or Varian INOVA 400 MHz NMR spectrometers. The NMR spectra were referenced to solvent and the spectroscopic solvents were purchased from Cambridge Isotope Laboratories.
MultiPep 2 was from CEM.
Synthesis and Characterization. Preparations and monomers previously listed above are omitted.
General procedure for the reduction of commercial Fmoc Amino Acids. Procedure for the reduction of Fmoc amino acids was adapted from methods discussed above.
To a stirred solution of Fmoc-L-amino acid 1.0 equivalent in anhydrous THF (3.3 mL) was added N,N-carbonyldiimidazole (1.33 equivalent) at room temperature. The reaction stirred for at least 10 minutes and was then cooled to 0 °C. Next, a solution of NaBH4 (1.66 equivalent) in H2O (1.66 mL, or 0.6M) was added. The solution was stirred for at least 30 minutes, up to 1.5 hours. The reaction was quenched by addition of IM HC1 and extracted with EtOAc (3x). The combined organics were washed lx with brine, dried over Na2SOr, and concentrated under vacuum.
Generalized Synthesis of Activated Carbonates. General procedure for the synthesis of activated carbonates was adapted from methods discussed above.
To a stirring solution of Fmoc-Amino alcohol (1.0 equivalent) in anhydrous DCM (0.2M) was added pyridine (1.3 equiv) dropwise. Next, 4-nitrophenyl chloroformate (1.5 equiv) was added, and the reaction left to stir overnight. Reaction was monitored by TLC and upon consumption of the starting material, was diluted excessively in DCM and transferred to a separatory funnel. The organic layer was washed with IM NaHSCL (3x), then IM NazCCh (5x, or until it stopped turning bright yellow), and finally brine. The organic layer was dried over Na2SO4, filtered, and concentrated in vacuo. The product was purified by silica gel chromatography.
NHS-Rhodamine B. Rhodamine B (1.0 g, 2.087 mmol, 1 eq.) was dissolved in acetonitrile (60 mL) in a 500 mL round bottom flask. A solution of N-hydroxy-succinimide (288 mgs, 2.5044 mmol, 1.2 eq.) in acetonitrile (20 mL) was added and the solution was heated to 45 °C. After allowing the solution to heat up (about 10 minutes), a solution of 1- Ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) (388 mgs, 2.5044 mmol, 1.2 eq) in chloroform (40 mL) was added. The reaction was allowed to stir overnight at 45 °C. The solvent was removed by rotary evaporation and the product was carried forward without additional purification. Low-Res ESI calculated for C32H34N3O5+ [M]+: 520.25, obs. [M]+: 520.2. (Figure 95).
Oligo Cleavage. Cleavages were affected with 0.1% TFA in DCM (4x0.250 mL per well) for 15-30 seconds each (until no color was observed upon addition of cleavage cocktail). Resin was filtered using well-plate vacuum setup, and cleaved product was concentrated in a deep well-plate using a GeneVac, a vacuum centrifuge. Due to cleavage times and amounts of TFA used, the trifluoroacetic ester of the terminal alcohol was observed by LC/MS.
Sequencing Experiments. The oligomers (-500 nmols of each) were dissolved in methanol (0.66 mL) and added to a 2-dram vial. Next, potassium phosphate in water (10 mg/well in 0.33 mL of water) was added to the solution. The reaction was heated to 70 °C, and held at the temperature for 45 minutes. Reaction was sampled for DESI by taking 50 pL of the reaction for each timepoint and combined. This was repeated every’ 30 minutes for 6 hours. The time points were placed into wells of a 96-well plate and dried under centrifugal vacuum evaporation (Genevac). Once dry, each sample was dissolved in 200 pL of MeOH and analyzed by DESI.
References
AL Griffiths J. Analytical Chemistry 2008, 80 (15), 5678-5683.
A2. Zhang J et al. Cancer Research 2016, 76 (22), 6588-6597.
A3. Sun J et al. Mass Spectrometry Reviews 2022, 41 (1), 70-99.
A4. Huang Khet al. ChemPlusChem 2022, 87 (1).
A5. Dahlhauser SD et al. Cell Reports Physical Science 2021, 2 (4), 100393.
A6. Dahlhauser SD et al. Journal of the American Chemical Society 2020, 142 (6),
2744-2749.
A7. Sokolowska I et al. Analytical Chemistry 2020, 92 (3), 2369-2373.
A8. Takats Z et al. Science 2004, 306 (5695), 471-473.
A9. Wiseman JM et al. Nature Protocols 2008, 3 (3), 517-524.
A10. Feider CL et al. Analytical Chemistry 2019, 91 (7), 4266-4290.
Al 1. Sans M et al. J Am. Soc. for Mass Spec. 2020, 31 (2), 418-428.
A12. Wleklinski M et al. Chemical Science 2018, 9 (6), 1647-1653.
A13. Merrifield RB. Science 1965, 150 (3693), 178-185.
A14. Wang SS. JACS 1973, 95 (4), 1328-1333.
A15. Carpino LA et al. JACS 1970, 92 (19), 5748-5749. A16. Behrendt R et al. Journal of Peptide Science 2016, 22 (1), 4-27.
A17. Lone A et al. Frontiers in Chemistry 2020, 8.
Al 8. Soete M et al. Angewandte Chemie International Edition 2022, 61 (13).
A19. Feider CL et al. J Am. Soc. for Mass Spec. 2019, 30 (2), 376-380.
A20. Galhena AS et al. Analytical Chemistry 2010, 82 (22). 9159-9163.
A21. Birtalan E et al. Peptide Science 2011, 96 (5), 694-701.
A22. Venter A et al. Analytical Chemistry 2006, 78 (24), 8549-8555.
A23. Ramette RW et al. JACS 1956, 78 (19), 4872-4878.
A24. Anslyn E. Modern physical organic chemistry. University Science Books: Mill Valley, California, 2006.
A25. Chambers MC et al. Nature Biotechnology 2012, 30 (10), 918-920.
A26. Doran D et al. Cell Reports Physical Science 2021, 2 (12), 100685.
Example 3
Various modifications were employed to increase the intensity of the desired oligomers (parent oligomer and sequencing products), both in general, and, in relation to undesired synthetic by products.
For example, the target oligomers can comprise endcaps. An "End cap’ is a single monomer added to the end of the oligomer that persists throughout the sequencing reaction. So, it is present on the parent oligomer and all sequenced products. End caps have two major functions: (1) increase the relative signal intensity of the desired parent and sequenced oligomers from background ions and undesired byproducts; and (2) favor the formation of a single adduct type in DESI (e.g.. all ions are observed as singly potassiated adducts - [M+K]+).
Seeing sequencing products as a single adduct greatly reduces the complexity’ of the signal deconvolution process, as the m/z difference between ions simply matches the mass of the monomer lost.
End caps persist throughout the sequencing process, both on the parent oligomer and on the sequenced products:
Example: sequences observed during a sequencing reaction when end caps are employed. “A’", “B”, ... “H” represent information encoding monomers. Synthesis is done from ‘left-to-right’ where ’A' is the first monomer added. B’ is the second, etc. Sequencing is done from 'left-to-right’, where the first monomer added during sequencing also happens to be the first to be lost during separate sequencing experiments. o Desired Parent Oligomer: “A B C D E F G H EndCap” 8mer o Sequenced products: “B C D E F G H EndCap” 7mer
“C D E F G H EndCap” 6mer
D E F G H EndCap” 5mer
“E F G H EndCap” 4mer
“F G H EndCap” 3mer
“G H EndCap” 2mer
“H EndCap” Imer
Seven end caps were synthesized and appended to test oligomers that were synthesized, sequenced, and analyzed via DESI-MS to evaluate their signal. These included: 1 -naphthol, glycine, TAMRA (5-Carboxytetramethylrhodamine). methyl t rosine (Tyr(OMe)), tyrosine (Tyr(OH)), nitrobenzoxadiazole (NBD), and Rhodamine B. Four of these were subject to further testing. The four end caps most extensively screened were methyl ty rosine (Tyr(OMe)), ty rosine (Tyr(OH), nitrobenzoxadiazole (NBD), and Rhodamine B. This required the synthesis of numerous different oligomers, each independently synthesized and then labeled with each end cap.
NBD was initially screened as it was used as a chromophore (absorbent tag) in previously reported work. However, it was chosen to move away from NBD because of its expensive cost, and primarily because it has been seen to hydrolyze off during sequencing. Once hydrolyzed off, sequencing products change in both their absorbance and mass over time, dramatically complicating the sequence deconvolution process.
Rhodamine B (rhoB) was screened because it is a strong absorber (colored and easy to work with) and it is naturally cationic, which it was believed would enhance the signal of Rhodamine B labeled oligomers in positive mode. Rhodamine B is also relatively heavy (-443 amu), and so Imers labeled with rhodamine B would have a minimum mass of -550 amu. Many background peaks occur in the lower (100-350 m z) mass range for many solventbased ambient ionization mass spectrometry' methods such as DESI-MS, which have made the detection of shorter sequenced oligomers in that mass range challenging. Using Rhodamine B-labelled oligomers, the mass range to detect all species from the oligomers was increased to 500-2000 m/z, effectively enhancing ionization and thus detection of shorter sequenced oligomers.
Two tyrosine derivatives were chosen because of their propensity to form potassium adducts, likely driven by the strong cation-pi interaction, and the fact that they can be synthesized easily from commercial amino acids using the same methods as for information encoding monomers.
Methyl tyrosine was tested alongside ty rosine because it does not require special cleavage conditions when the synthesized oligomer is cleaved from the resin.
Tyrosine (Tyr(OH)) end caps are synthesized with a protected alcohol (Tyr(OlBu)) that must be deprotected when cleaved from the resin. This deprotection adds additional time, reagents, and complexity' when cleaving that would otherwise not be employed.
Both tyrosine derivates were tested in order to evaluate if there is a difference in the signal intensity and type of adduct formed. However, both behaved similarly, so Tyr(OMe) was chosen because it is easier to work with and requires no additional modifications in the cleavage protocol.
In some examples, the compositions can employ “Truncated Caps.” A “Truncated Cap” is a chemical cap that is added to prevent sequences containing deletion products to proceed during the synthesis.
During synthesis, monomers are added successively. If a 98% coupling efficiency in each step is assumed, that means at each step 2% of the sequences will not have the proper monomer added in a given position. These sequences are said to have a deletion.
If the synthesis continued without the use of a truncated cap, in the last step, all oligomers would be capped with an end cap, and so end cap labeled oligomers would have mixtures of desired products and deletion products. o Example: no truncated cap used: represent information encoding monomers.
Synthesis is done from ‘left-to-right’ where ’A’ is the first monomer attached, ‘B’ is the second, etc.
Desired 8mer Sequence: “A B C D E F G H EndCap”
Deletion in 1st synthesis step: “B C D E F G H EndCap”
Deletion in 3rd synthesis step: “A B D E F G H EndCap”
Deletion in 7th synthesis step: “A B C D E F H EndCap” o Example, truncated cap used’.
“A”-“H” represent information encoding monomers.
Synthesis is done from ‘left-to-right’ where ‘A’ is the first monomer attached, ‘B’ is the second, etc.
Desired 8mer Sequence: A B C D E F G H EndCap” Deletion in 1st synthesis step: ■‘A TruncatedCap”
Deletion in 3rd synthesis step: “A B C TruncatedCap”
Deletion in 7th synthesis step: “A B C D E F G TruncatedCap'
By using truncated caps, the number of theoretical synthetic by products observed during sequencing can be decreased, and propagation of deletion products can be avoided.
Truncated caps need to react quantitatively (100% yield) and quickly (~5 minutes) with any unreacted amines on the growing oligomer end following coupling steps in order to be useful. Three different truncated caps w ere tested alongside three different end caps (3x3 matrix) in order to determine the best end cap and truncated cap combination.
The following three anhydrides were evaluated as end caps: acetic anhydride, succinic anhydride, and trichloroacetic anhydride.
Acetic anhydride w as chosen because it is a commonly used end cap in peptide chemistry, also used to cap free amines and stop reactions of deletion products.
Succinic anhydride was chosen because once it reacts, it forms a negatively charged carboxylate on the truncated oligomer. It was hypothesized this negative charge could decrease the intensity of truncated ions observed, because imaging w as in positive mode.
Trichloroacetic anhydride was evaluated because chlorine exists as two stable isotopes (35C1 and 37C1), and so it was hypothesized that trichloroacetyl-truncated capped oligos would have a distinct isotope pattern that could be used as a selection criteria against the desired signal.
Acetic anhydride w as the most stable during sequencing runs and showed good results when paired with methyl tyrosine end caps.
The DESI signal was further improved, and analysis was expedited, by optimizing experimental parameters.
Various solvents were tested, including methanol (MeOH), acetonitrile (ACN), different ratios of these solvents with water, and a few additives (i.e., ammonium acetate, ammonium formate).
The molecular profde observ ed using ACN was substantially improved compared to using 98% MeOH, with only potassium-adducted species detected compared to both sodium- and potassium-adducts with 98% MeOH. Furthermore, the relative abundance of all oligomer species were higher using ACN, with less background noise observed.
Performance was evaluated in both positive and negative ion modes, determining that positive ion mode resulted in detection of higher relative abundances of species of interest compared to negative ion mode.
The stage velocity and motion profiles of the 2D stage were optimized to further accelerate the analysis without compromising data quality.
The point-to-point motion profiles “constant velocity” (CV) and “dwell” were evaluated, where after analysis of a certain area the stage is moved immediately to the next designated spot for analysis. The dwell motion profile allows for the stage, which holds the sample, to remain below the DESI spray in a specific area of interest for a certain amount of time before moving to the next spot, whereas the CV motion profile enables the stage to raster below the DESI spray at a user-specified velocity for certain intervals. A dwell time of two seconds and a CV of 800 pm/sec was used for the analysis of three sequential spots of G8, respectively. Using CV resulted in improved detection of oligomer species and lower background compared to using dwell settings. Thus, the CV motion profile was used for all subsequent DESI-MS analyses.
Four stage velocities were tested between 1000 and 5000 pm/sec using a rhoB-capped lOmer spotted in triplicate on a glass slide. All truncated oligomer species were observed with every stage rate. The median relative abundances of each oligomer species were highest using 1500 pm/sec, followed by 3000 pm/sec. Given the high quality of data obtained with 1500 pm/sec, this was chosen as the stage velocity moving forward coupled with a programmed delay between spots of two seconds.
Combinations of different time points were tested to yield a balanced molecular profile of all truncated oligomer species of interest and prevent redundant sample analysis.
By tuning DESI-MS parameters (i.e., solvent composition and stage velocity) to expedite analysis while still achieving high signal and sensitivity, it was shown that analysis of one 10-mer with all its iterative truncations can be achieved in approximately three seconds using the optimized stage rate of 1500 pm/sec (one spot, 4 mm in diameter). In contrast, Amalian et al. recently established the use of DESI in MS2 mode for reading molecularly encoded sequence-defined polymers (Amalian et al. Advanced Materials Technologies 2021, 6. 2001088). Eight spots (each 3 mm in diameter) of five oligomers were needed to encode the four-character word ”sty\" in binary, which took approximately 5.5 minutes for the analysis to be completed using a stage rate of 500 pm/sec. Remarkably, the methods described herein were able to encode a 42-character quote from Maya Angelou into five oligomers using hexadecimal coding, which altogether can be analyzed in under 30 seconds. Porous PTFE-coated slides were tested as substrates and found that this resulted in similar detection of Maya Angelou oligomers deposited on a multi -well glass slide as the surface substrate; work inspired by research from Cooks and coworkers where PTFE-coated slides were used as a multi-sample substrate in a study to screen thousands of organic reactions per hour (Morato et al. SLAS Technol. 2021, 26. 555-571). Furthermore, by using a PTFE-coated substrate instead of a glass slide with predefined wells, the storage capacity and analysis speed can be substantially increased by depositing oligomers more closely together as DESI-MS can be optimized to achieve spatial resolutions between 50-250 pm.
Example 4 - Molecular Record Player
Disclosed herein are methods and systems using desorption spray ionization (DESI) technology in tandem with oligourethanes to read stored information in molecules in realtime, e.g. methods and systems for a “molecular record player.” The systems and methods can further include deconvoluting the information stored in the molecules.
The methods and systems disclosed herein provide an alternative to existing modem data storage.
Storing information in carbon based molecules does not require expensive metals used in modem hard drives. Storing information in carbon based molecules also provides for extended storage capabilities due to the stability of the carbon based molecules under a wide variety of conditions.
Example 5 -Automated Synthesis and Sequencing of Sequence-Defined Oligourethanes using a commercial peptide synthesizer Analyzed via Desorption Electrospray Ionization (DESI) MS
Abstract: A single XYZ robotic liquid handler was used to synthesize and sequence sequence-defined oligourethanes (SDOUs). DESI-MS was used to automate and parallelize data acquisition. Sequences were decoded using a Python based program
Introduction. Abiotic Sequence-defined macromolecules have attracted increasing attention over the last decade, in large part due to wide ranging applications in life and material sciences [Bl], Sequence-defined oligomers (SDOs) and polymers (SDPs) have had a dramatic rise in the field of information storage, where they are seen as a complementary small molecule alternative to traditional silicon based storage devices [B2] . Sequence-defined macromolecules have diverse chemical structure, with examples including, but not limited to, DNA [B3, B4], peptides [B5, B6], abiotic phosphates and phosphodiesters [B7-B9], poly(alkoxyamine amide)s [B10-B12], esters [B13-B15], and urethanes [B16-B21], The use of sequence-defined macromolecules for information storage requires the ability to effectively ‘write’ and ‘read’ [B22] the information, where the ‘writing’ is often viewed as the synthesis, and the ‘reading’ as sequencing. Synthetically, chemically diverse sequence-defined oligomers can be achieved in high yields by employing solid-phase synthesis (SPS). Solid-phase synthesis methods often have been optimized for numerous backbones, and often have coupling efficiencies of >99% per step. At such coupling efficiencies, lOmers can be synthesized in 90% or greater yields, and depending on the application, may be used sans purification. For sequence-defined oligomers, solid-phase synthesis can often be applied to diverse monomers of a specific type (e.g. solid-phase synthesis of various amino acids) with protocols generally independent of length. In contrast, sequencing and sequence reconstruction protocols are often tailored to specific macromolecular structure.
Few examples of high-throughput sequencing methods for abiotic polymers can rival those developed for biopolymers. Recently, nanopores have been shown to be capable of sequencing abiotic oligomers and polymers [B23-B25], The use of nanopore for sequencing of abiotic macromolecules is still in its infancy, requiring monomers with significant differences and the addition of a large excess of non-information encoding monomers in order to create distinct signals. The majority of synthetic sequence-controlled polymers are characterized using tandem MS/MS [B6, Bl 1, B12, B15, B19, B20, B26-B30], A drawback seen in tandem MS analysis is the formation of adducts in a variety of charge states.
The synthesis and self-sequencing of sequence-defined oligourethanes (SDOUs), with applications in information storage and steganography, has recently been reported [B16, B17, B31 ] . This platform relies on a chain-end depolymerization sequencing methodology that utilizes a thermally induced intramolecular cyclization to iteratively remove the terminal monomer (Figure 96A). This allow s each truncated oligourethane to be characterized bysimple liquid chromatography-mass spectrometry (LC/MS), forgoing MS/MS protocols and greatly simplifying deconvolution [B31], The parent sequence can be reconstructed by looking at the difference in mass betw een two oligourethanes that differ by a single monomer, where the monomer lost from the sequencing O-terminus can be deciphered based on the mass difference of the two strands.
Mol.E-coder and Mol. E-decoder software were previously described [Bl 7], that in short, are capable of encoding information into sequence-defined oligourethanes based on the number of information encoding monomers used (e.g. hexadecimal encoding of bit strings encoded in base 16 with 16 unique monomers, 4 bits/monomer). The encoded hexadecimal string defines the sequence to be synthesized. Sequencing of each sequence-defined oligourethanes creates all possible chain-end depolymerized oligourethanes, and mass differences between these sequenced strands can be fed into the decoder to retrieve the original encoded information.
Recent work showed that eight different oligourethanes were mixed into a single sequencing vessel, and each sequence was able to be reconstructed based on the incorporation of unique isotope tags [Bl 6], While this is the most information stored in a single sample of abiotic sequence-defined polymers to date, the sequencing procedure required each time point to be analyzed using multiple 18 minute liquid chromatography-mass spectrometry (LC/MS) runs.
Desorption Electrospray Ionization (DESI) is a widely popular ambient ionization technique. One key feature of ambient ionization is the minimal processing that samples undergo prior to analysis [B32], No embedding in a matrix or dissolution of the analyte is necessary, other than the deposition or attachment of the sample onto a slide to be analyzed. With DESI-MS, the freely moving stage allows the solvent capillary and inlet to raster across the samples allowing specific areas of tissues [B33] or samples to be analyzed individually [B34] . This precisely controlled movement allows the scanning speed to be optimized for different samples. In the case of screening alkenylation and azo-click reactions by the Cooks groups, speeds as fast as one second per reaction were achieved [B35], Additionally, it is possible to selectively desorb specific ions by varying the solvent spray composition allowing suppression of specific ions, creating a more easily interpretable spectrum. The ease of preparation, as well as strong ionization and rapid analysis speed, makes ambient ionization techniques such as DESI ideal for high-throughput analysis.
Following previous work, a bottle necks in the use of oligourethanes for information storage is addressed by automating and increasing the speed for synthesis, sequencing, and sequence reconstruction. The development of separate automated synthesis and sequencing platforms by adapting a single commercially available synthesizer to perform the two different chemistries is described. Further, the development and implementation of DESI-MS based data acquisition with a new sequence reconstruction program is described. Thus, this report describes the automated ‘writing’ and ‘reading’ of data in the form of sequence- defined oligourethanes.
METHODS Workflow Overview. The creation of an automated platform for the '‘writing” and “reading” of information was envisioned, where the information is stored sequence-defined oligourethanes (SDOUs). Analogous to the storage of information in sequence-defined polymers, it was posited that the combined synthesis and sequencing of sequence-defined oligourethanes can be viewed as automated ‘‘writing” of information, while the DESI-MS analysis and subsequent sequence reconstruction, can be viewed as the '‘reading” of said information. Central to the goal of automating the “writing” and process was creating a unified workflow that utilized a single robotic liquid handler for both synthesis and sequencing.
To optimize the “writing” process, previously reported procedures for the solid-phase synthesis (SPS) of urethanes were adapted and optimized using a commercially available peptide synthesizer composed of an XYZ robot with a liquid handler. Optimization of synthesis protocols resulted in the fully automated synthesis of sequence-defined oligourethanes that could be sequenced sans purification. An automated sequencing protocol that employed the same robotic liquid handler was then able to be developed and optimized, simply by manipulating the programming of the robot. The “reading” process that followed sequencing involved optimizing DESI-MS experimental parameters to enhance the signal of the desired species (parent oligourethane and all truncated oligourethane strands) and attenuate signals of undesired species.
RESULTS AND DISCUSSION
Oligourethane Sequence Design. In order to optimize the automated workflow, automated synthesis of sequence-defined oligourethanes needed to (1) have near quantitative coupling steps, (2) prevent deletion products from continuing in synthesis, (3) bias sequenced products to form one type of adduct (e.g. [M+K]+) in DESI-MS, and (4) suppress ion formation of undesired byproducts (Figure 96C).
Several changes to monomers were made compared to previous reports to improve synthesis and sequence reconstruction. Monomers with protected side chains were removed in order to simplify and standardize cleavage from resin, and monomers were changed in order to keep the largest monomer mass less than twice the mass of the smallest monomer. While a subtle point, this change improved the ability to decipher the last information encoding monomer, as will be discussed later.
Automated Synthesis. Sequence-defined oligourethanes were synthesized by adapting an automated parallel peptide synthesizer (CEM MultiPep 2 Parallel Peptide Synthesizer, “MP”) for the coupling of PNOC activated amino alcohol monomers to an amino alcohol loaded resin (Figure 96B). A schematic showing the setup of the MP for synthesis is show n in Figure 97A. Optimization of synthesis methodology was aided by the development of a sequence deletion and MS adduct calculator.
Oligourethane End Caps. An ‘End cap’ is a single monomer added to the N-terminus of the oligourethane, thereby being present on the parent oligourethane and all sequenced products that form during the sequencing reaction. End caps have two major functions: 1) increase the relative signal intensity of the parent and sequenced oligourethanes from background ions and undesired byproducts and 2), end favor the formation of a single adduct type in DESI-MS (e.g. [M+K]+). By biasing sequencing products to form a single adduct, the sequence reconstruction process is greatly simplified, as the m/z difference between all observed ions simply matches the mass of the monomer lost. Six different end caps were evaluated. In the end, Tyr(OMe) was chosen as the end cap as (1) Fmoc-Tyr(OMe)-PNOC could be coupled to the growing strand using identical conditions as to information encoding monomers, (2) it could be synthesized and recrystallized on a multi-gram scale, and (3) it was observed to exclusively form singly potassiated adducts ([M+K]+). The exclusive formation of this adduct in MeCN is attributed to the strong intermolecular cation-pi interaction [B36, B37], and the presence of potassium cations (e.g. KOH) used during chain-end depolymerization sequencing.
Oligourethane Truncated Caps. “Truncated caps” w ere employed to block unreacted amines on growing oligourethane chain following an incomplete coupling step, as well as cap the N-terminus following completion of synthesis. Quantitative capping using acetic anhydride prevented deletion products from proceeding in synthesis, aiding sequence reconstruction. Further, the acetylation of free amines removed cationic ammonium salts, which were observed to have large ionization intensities when present.
Automated Sequencing. While the XYZ robot and liquid handler used for the project was commercially produced for the purpose of synthesis, it was envisioned that this instrument could be used to automate the oligourethane sequencing program (Figure 97B). In short, sequencing in the MP involved the automated addition of sequencing reagents (e.g. base) to a 96 well plate that contained the cleaved sequence-defined oligourethanes. Approximately every’ 60 minutes, an aliquot is pipetted from a sequencing well and dispensed into the corresponding w ell of a 'time points’ 96-well plate. Thus, at the end of sequencing, multiple time points are gathered in a single w ell. DESI-MS. Initial work was done to screen DESI-MS methods for the ohgourethanes. The goal was to develop a method suitable for screening end caps. Initial results indicated that all truncated species of test oligourethanes were easily observed in a MeCN solvent system in positive mode. Notably, screening various end caps revealed Tyr(OMe) as the optimal end cap as described above.
Sequence Reconstruction. With an efficient sequencing method in hand, a general protocol was then developed for extracting sequence information from the MS data. Using an inhouse Python script, the mass of the parent oligomer was first automatically identified. The program searches the window of m/z values between the masses of the n - 1 oligomer which would result from the loss of the heaviest monomer and the lightest monomer respectively. The largest signal in this window is selected as the sequence product (i.e., the n - 1 oligomer) and the difference between the parent oligomer and this new signal is then calculated. After comparing the mass loss to a dictionary of known monomer masses, the monomer which was closest in mass to the mass loss was selected. The process is then repeated with the n - 1 oligomer as the new parent oligomer to identify subsequent mass losses (n - 2, n - 3, etc.). The process is repeated until the endcap monomer mass is identified (in the present case Tyr(OMe)) and the resulting sequence data is obtained.
This strategy is robust and can be used to extract sequence information without error for 9-mers. However, because the sequence products are selected based on ion counts, noisy mass spectra and unanticipatedly large signals could be misinterpreted as sequence products. In this testing, this issue does not impact the sequencing of 9-mers or smaller. Because monomers are identified based on a ranked list, error rates could likely be further diminished by using monomers which have large mass differences.
An example of MS data for a successful trial of automated sequencing is shown in Figure 99.
Automated Synthesis and Sequencing. The combined and independently optimized synthesis and sequencing programs were tested to produce a continuous workflow for the automated encoding and decoding of information. To test this workflow, it was chosen to first encode the timeless quote by the writer and activist Maya Angelou: “When you learn, teach, when you get, give. ” To write this quote in base 16, six 9mers were synthesized and sequenced. Workflow overview for the 'reading' and ‘writing' of information is shown in
Figure 98.
References (Bl) Aksakal R et al. Applications of Discrete Synthetic Macromolecules in Life and Materials Science: Recent and Future Trends. Adv. Sci. 2021, n/a (n/a), 2004038.
(B2) Zhirnov V et al. Nucleic acid memory. Nat. Mater. 2016, 15 (4), 366-370.
(B3) Church GM et al. Next-Generation Digital Information Storage in DNA. Science 2012, 337 (6102), 1628.
(B4) Pan C et al. Rewritable two-dimensional DNA-based data storage with machine learning reconstruction. Nat. Commun. 2022, 73 (1), 2984.
(B5) Cafferty BJ et al. Storage of Information Using Small Organic Molecules. ACS Central Science 2019, 5 (5), 911-916.
(B6) Rossler SL et al. Abiotic peptides as carriers of information for the encoding of small-molecule library synthesis. Science 2023, 379 (6635), 939-945.
(B7) Al Ouahabi A et al. Synthesis of Non-Natural Sequence-Encoded Polymers Using Phosphoramidite Chemistry. J. Am. Chem. Soc. 2015, 737 (16), 5629-5635.
(B8) Al Ouahabi A et al. Synthesis of Monodisperse Sequence-Coded Polymers with Chain Lengths above DP100. ACS Macro Letters 2015, 4 (10), 1077-1080.
(B9) Konig NF et al. A Simple Post-Polymerization Modification Method for Controlling Side-Chain Information in Digital Polymers. Angew. Chem. Int. Ed. 2017, 56 (25), 7297-7301.
(B10) Charles L et al. Tandem mass spectrometry sequencing in the negative ion mode to read binary information encoded in sequence-defined poly (alkoxy amine amide)s. Rapid Commun. Mass Spectrom. 2016, 30 (1), 22-28.
(Bl 1) Roy RK et al. Design and synthesis of digitally encoded polymers that can be decoded and erased. Nat. Commun. 2015, 6 (1), 7237.
(Bl 2) Charles L et al. MS/MS Sequencing of Digitally Encoded Poly(alkoxy amine amide)s. Macromolecules 2015, 48 (13), 4319-4328.
(B13) Lee JM et al. Semiautomated synthesis of sequence-defined polymers for information storage. Sci. Adv. 2022. 8 (10), eabl8614.
(B14) Lee JM et al. Nondestructive Sequencing of Enantiopure Oligoesters by Nuclear Magnetic Resonance Spectroscopy. JACS Au 2022, 2 (9), 2108-2118.
(B15) Lee JM et al. High-density information storage in an absolutely defined aperiodic sequence of monodisperse copolyester. Nat. Commun. 2020, 77 (1), 56.
(B16) Dahlhauser SD et al. Molecular Encryption and Steganography Using Mixtures of Simultaneously Sequenced, Sequence-Defined Oligourethanes. ACS Central Science 2022, 8 (8), 1 125-1133.
(B17) Dahlhauser SD et al. Efficient molecular encoding in multifunctional self- immolative urethanes. Cell Reports Physical Science 2021, 2 (4), 100393.
(Bl 8) Soete M et al. Discrete, self-immolative N-substituted oligourethanes and their use as molecular tags. Polym. Chem. 2022, 13 (28). 4178-4185.
(Bl 9) Gunay Ufuk S et al. Chemoselective Synthesis of Uniform Sequence-Coded Polyurethanes and Their Use as Molecular Tags. Chem 2016, 1 (1), 114-126.
(B20) Martens S et al. Multifunctional sequence-defined macromolecules for chemical data storage. Nat. Commun. 2018, 9 (1), 4451.
(B21) Mondal T et al. Damage and Repair in Informational Poly (N-substituted urethane)s. Angew. Chem. Int. Ed. 2020, 59 (46), 20390-20393.
(B22) Soete M et al. Reading Information Stored in Synthetic Macromolecules. J. Am. Chem. Soc. 2022, 144 (49), 22378-22390.
(B23) Cao C et al. Aerolysin nanopores decode digital information stored in tailored macromolecular analytes. Sci. Adv. 2020, 6 (50), eabc2661.
(B24) Tabatabaei SK et al. Expanding the Molecular Alphabet of DNA-Based Data Storage Systems with Neural Network Nanopore Readout Processing. Nano Lett. 2022, 22 (5). 1905-1914.
(B25) Yan S et al. Non-binary Encoded Nucleic Acid Barcodes Directly Readable by a Nanopore. Angew. Chem. Int. Ed. 2022, 61 (20), e202116482.
(B26) Launay K et al. Precise Alkoxyamine Design to Enable Automated Tandem Mass Spectrometry’ Sequencing of Digital Poly(phosphodiester)s. Angew. Chem. Int. Ed. 2021, 60 (2), 917-926.
(B27) Al Ouahabi A et al. Mass spectrometry sequencing of long digital polymers facilitated by programmed inter-byte fragmentation. Nat. Commun. 2017, 8 (1), 967.
(B28) Cavallo G et al. Orthogonal Synthesis of '’Easy-lo-Read" Information- Containing Polymers Using Phosphoramidite and Radical Coupling Steps. J. Am. Chem. Soc. 1M6 38 (30), 9417-9420.
(B29) Liu B et al. Engineering digital polymer based on thiol-maleimide Michael coupling toward effective writing and reading. Polym. Chem. 2020, 11 (10), 1702-1707.
(B30) Laurent E et al. High-Capacity’ Digital Polymers: Storing Images in Single Molecules. Macromolecules 2020, 53 (10), 4022-4029.
(B31) Dahlhauser SD et al. Sequencing of Sequence-Defined Oligourethanes via Controlled Self-Immolation. J. Am. Chem. Soc. 2020, 142 (6), 2744-2749.
(B32) Takats Z et al. Mass Spectrometry Sampling Under Ambient Conditions with Desorption Electrospray Ionization. Science 2004, 306 (5695), 471-473.
(B33) Sans M et al. Spatially Controlled Molecular Analysis of Biological Samples Using Nanodroplet Arrays and Direct Droplet Aspiration. Journal of the American Society for Mass Spectrometry 2020, 31 (2), 418-428.
(B34) Wleklinski M et al. High throughput reaction screening using desorption electrospray ionization mass spectrometry. Chemical Science 2018. 9 (6), 1647-1653.
(B35) Huang KH et al. Late-Stage Functionalization and Characterization of Drugs by High-Throughput Desorption Electrospray Ionization Mass Spectrometry. ChemPlusChem 2022, <57 (1).
(B36) Dougherty DA. The Cation-n: Interaction. Acc. Chem. Res. 2013, 46 (4), 885- 893.
(B37) Mecozzi S et al. Cation-pi interactions in aromatics of biological and medicinal interest: electrostatic potential surfaces as a useful qualitative guide. PNAS 1996, 93 (20), 10566-10571
Other advantages which are obvious and which are inherent to the invention will be evident to one skilled in the art. It will be understood that certain features and subcombinations are of utility and may be employed without reference to other features and subcombinations. This is contemplated by and is within the scope of the claims. Since many possible embodiments may be made of the invention without departing from the scope thereof, it is to be understood that all matter herein set forth or shown in the accompanying drawings is to be interpreted as illustrative and not in a limiting sense.
The compositions and methods of the appended claims are not limited in scope by the specific compositions methods described herein, which are intended as illustrations of a few aspects of the claims and any compositions and methods that are functionally equivalent are intended to fall within the scope of the claims. Various modifications of the compositions and methods in addition to those shown and described herein are intended to fall within the scope of the appended claims. Further, w hile only certain representative method steps disclosed herein are specifically described, other combinations of the method steps also are intended to fall within the scope of the appended claims, even if not specifically recited. Thus, a combination of steps, elements, components, or constituents may be explicitly mentioned herein or less, however, other combinations of steps, elements, components, and constituents are included, even though not explicitly stated.

Claims

CLAIMS What is claimed is:
1. A method for decrypting information stored within a target oligomer, the target oligomer comprising a self-immolative oligourethane comprising a plurality of unique monomers in a pre-defined order and an endcap, wherein the target oligomer has a first end and a second end, the second end comprising the endcap; wherein each of the unique monomers has a unique mass spectrometry profile; the method comprising: creating a time-interval substrate, the time-interval substrate comprising a plurality of samples each collected at a different time-interval, the plurality of samples being disposed on a substrate in an ordered array; wherein each sample comprises a portion of a mixture formed by subjecting the target oligomer to 5-exo-trig cyclization and elimination; analyzing the time-interval substrate using desorption electrospray ionization (DESI) mass spectrometry, thereby generating a plurality of mass spectrometry profiles; and evaluating the plurality of mass spectrometry profiles to decrypt the information stored within the target oligomer.
2. The method of claim 1 , wherein creating the time-interval substrate comprises: subjecting the target oligomer to 5-exo-trig cyclization and elimination, thereby forming a mixture; collecting a plurality of aliquots of the mixture over a plurality of time-intervals; placing each of the aliquots at a location on a substrate, such that the plurality’ of aliquots are disposed on the substrate in an ordered array, the location in the array corresponding to the time-interval at which the aliquot was collected.
3. The method of claim 2, wherein the substrate comprises a PTFE-coated slide, a multiwell glass slide, or a combination thereof.
4. The method of any one of claims 1-3, wherein the plurality of samples further comprise a solvent.
5. The method of claim 4, wherein the solvent comprises acetonitrile.
6. The method of any one of claims 1-5, wherein the ordered array comprises a two- dimensional array.
7. The method of any one of claims 1-6, wherein the plurality of samples comprises from 2 to 256 samples, such as from 8 to 32 samples.
8. The method of any one of claims 1-7, wherein the time-intervals occur at regular intervals.
9. The method of any one of claims 1-8, wherein the time-intervals independently occur at an interval of from 15 minutes to 120 minutes.
10. The method of any one of claims 1-9, wherein the plurality of unique monomers comprises from 2 to 256 unique monomers, such as from 8 to 32 unique monomers.
11. The method of any one of claims 1-10, wherein the target oligomer has a total length of from 2 to 256 monomers, such as from 8 to 32 monomers.
12. The method of any one of claims 1-11, wherein the self-immolative oligourethane is derived from -amino alcohols.
13. The method of any one of claims 1-12, wherein the mass spectrometry7 profiles are generated in positive-ion mode.
14. The method of any one of claims 1-13, wherein the time-interval substrate is disposed on a movable stage, and analyzing the time-interval substrate using DESI-MS comprises translating the movable stage to sequentially subject each of the plurality of samples to the DESI-MS analysis.
15. The method of claim 14, wherein the movable stage is translated using a constant velocity motion profile.
16. The method of claim 15, wherein the constant velocity7 motion profile comprises a stage velocity7 of from 500 to 3000 pm/second.
17. The method of any one of claims 1-16, wherein the endcap comprises rhodamine B or methyl ty rosine.
18. The method of any one of claims 1-17, wherein the information stored within the target oligomer is hexadecimal based.
19. The method of any one of claims 1-18, wherein the information stored within the target oligomer comprises a cipher key.
20. The method of any one of claims 1-19, wherein the information stored within the target oligomer comprises from 2 to 256 bits of information.
21. The method of any one of claims 1-20, wherein the method further comprises steganography.
22. A method for encrypting information within a target oligomer, the method comprising: selecting a plurality of unique monomers, wherein each of the unique monomers has a unique mass spectrometry profile; assigning a unique value to each of the unique monomers within the plurality; synthesizing a target oligomer comprising a self-immolative oligourethane comprising the plurality of unique monomers in a pre-defined order and an endcap. wherein the target oligomer has a first end and a second end, the second end comprising the endcap; wherein the endcap comprises rhodamine B or methyl tyrosine; wherein the assigned values and the pre-defined order encrypts pre-defined information into the target oligomer.
23. The method of claim 22, wherein the plurality of unique monomers comprises from 2 to 256 unique monomers, such as from 8 to 32 unique monomers.
24. The method of claim 22 or claim 23. wherein the target oligomer has a total length of from 2 to 256 monomers, such as from 8 to 32 monomers.
25. The method of any one of claims 22-24, wherein the self-immolative oligourethane is derived from -amino alcohols.
26. The method of any one of claims 22-25, wherein the information stored within the target oligomer is hexadecimal based.
27. The method of any one of claims 22-26, wherein the information stored within the target oligomer comprises a cipher key.
28. The method of any one of claims 22-27, wherein the information stored within the target oligomer comprises from 2 to 256 bits of information.
I l l
29. The method of any one of claims 22-28, wherein the method further comprises steganography.
30. A composition for molecular cryptography comprising: a target oligomer comprising a self-immolative oligourethane comprising a plurality of unique monomers in a pre-defined order and an endcap, wherein the target oligomer has a first end and a second end, the second end comprising the endcap; wherein each of the unique monomers has a unique mass spectrometry profile; wherein the endcap comprises rhodamine B or methyl tyrosine.
31. The composition of claim 30, wherein the composition further comprises a solvent.
32. The composition of claim 31, wherein the solvent comprises acetonitrile.
33. The composition of any one of claims 30-32, wherein the composition further comprises: a truncated oligomer comprising a self-immolative oligourethane comprising at least a portion of the plurality of unique monomers and a truncated endcap, wherein the truncated oligomer has a leading end and a trailing end, the trailing end comprising the truncated endcap; wherein the number of monomers in the truncated oligomer is less than that of the target oligomer, the order of monomers in the truncated oligomer differs from the pre-defined order of the target oligomer by 1 monomer or more, or a combination thereof; and wherein the truncated cap comprises an anhydride.
34. The composition of claim 33, wherein the truncated cap comprises acetic anhydride.
35. The composition of any one of claims 30-34, wherein the plurality7 of unique monomers comprises from 2 to 256 unique monomers, such as from 8 to 32 unique monomers.
36. The composition of any one of claims 30-35, wherein the target oligomer has a total length of from 2 to 256 monomers, such as from 8 to 32 monomers.
37. The composition of any one of claims 30-36, yvherein the self-immolative oligourethane is derived from [3-amino alcohols.
38. The composition of any one of claims 30-37, yvherein the information stored yvithin the target oligomer is hexadecimal based.
39. The composition of any one of claims 30-38, wherein the information stored within the target oligomer comprises a cipher key.
40. The composition of any one of claims 30-39, wherein the information stored within the target oligomer comprises from 2 to 256 bits of information.
EP24866009.4A 2023-04-27 2024-04-26 Compositions and methods for encrypting, storing, and decrypting information in oligomers Pending EP4732286A2 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202363462267P 2023-04-27 2023-04-27
PCT/US2024/026494 WO2025058677A2 (en) 2023-04-27 2024-04-26 Compositions and methods for encrypting, storing, and decrypting information in oligomers

Publications (1)

Publication Number Publication Date
EP4732286A2 true EP4732286A2 (en) 2026-04-29

Family

ID=95022598

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24866009.4A Pending EP4732286A2 (en) 2023-04-27 2024-04-26 Compositions and methods for encrypting, storing, and decrypting information in oligomers

Country Status (2)

Country Link
EP (1) EP4732286A2 (en)
WO (1) WO2025058677A2 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP5997650B2 (en) * 2013-04-15 2016-09-28 株式会社日立ハイテクノロジーズ Analysis system
US10438662B2 (en) * 2016-02-29 2019-10-08 Iridia, Inc. Methods, compositions, and devices for information storage

Also Published As

Publication number Publication date
WO2025058677A2 (en) 2025-03-20
WO2025058677A3 (en) 2025-06-26

Similar Documents

Publication Publication Date Title
Boukis et al. Data storage in sequence-defined macromolecules via multicomponent reactions
Dahlhauser et al. Molecular encryption and steganography using mixtures of simultaneously sequenced, sequence-defined oligourethanes
Otto et al. Recent developments in dynamic combinatorial chemistry
Nanita et al. Serine octamers: cluster formation, reactions, and implications for biomolecule homochirality
US6475807B1 (en) Mass-based encoding and qualitative analysis of combinatorial libraries
DE60128900T2 (en) GROUND MARKER
Charles et al. MS/MS-assisted design of sequence-controlled synthetic polymers for improved reading of encoded information
Soete et al. Reading information stored in synthetic macromolecules
US20030100018A1 (en) Mass-based encoding and qualitative analysis of combinatorial libraries
Dailler et al. Divergent Synthesis of Aeruginosins Based on a C (sp3) H Activation Strategy
Von Eckardstein et al. Total synthesis and biological assessment of novel albicidins discovered by mass spectrometric networking
Mierke et al. Neuropeptide Y: Optimized solid‐phase synthesis and conformational analysis in trifluoroethanol
US6218551B1 (en) Combinatorial hydroxy-amino acid amide libraries
Sanchez-Martin et al. The impact of combinatorial methodologies on medicinal chemistry
Wu et al. Rapid Access to Multiple Classes of Peptidomimetics from Common γ‐AApeptide Building Blocks
Shuluk et al. A Workflow Enabling the Automated Synthesis, Chain-End Degradation, and Rapid Mass Spectrometry Analysis for Molecular Information Storage in Sequence-Defined Oligourethanes
Ervin et al. Proline Behavior in Model Prebiotic Peptides Formed by Wet–Dry Cycling
Zwillinger et al. Isotope ratio encoding of sequence-defined oligomers
EP4732286A2 (en) Compositions and methods for encrypting, storing, and decrypting information in oligomers
de Miguel et al. Generation and screening of synthetic receptor libraries
KR102806249B1 (en) Method of preparing sequence-defined polymer for information storage, information storage method, information decoding method and sequence-defined polymer for information storage
Peker et al. Analytical Tools for Dynamic Combinatorial Libraries of Cyclic Peptides
Peterse et al. Solid‐Phase Synthesis of Macrocyclic Peptides via Side‐Chain Anchoring of the Ornithine δ‐Amine
Arnusch et al. Solid phase synthesis of vancomycin mimics
Manku et al. Synthesis and high performance liquid chromatography/electrospray mass spectrometry single-bead decoding of split-pool structural libraries of polyamines supported on polystyrene and polystyrene/ethylene glycol resins

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251002

AK Designated contracting states

Kind code of ref document: A2

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR