EP4731787A1 - Operable random dna - Google Patents
Operable random dnaInfo
- Publication number
- EP4731787A1 EP4731787A1 EP23735638.1A EP23735638A EP4731787A1 EP 4731787 A1 EP4731787 A1 EP 4731787A1 EP 23735638 A EP23735638 A EP 23735638A EP 4731787 A1 EP4731787 A1 EP 4731787A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- sequence
- composition
- dsodns
- random
- sequencing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6876—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Organic Chemistry (AREA)
- Analytical Chemistry (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Microbiology (AREA)
- Immunology (AREA)
- Molecular Biology (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- Physics & Mathematics (AREA)
- Biochemistry (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Genetics & Genomics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
The invention is notably directed to a composition including a plurality of double-stranded oligodeoxynucleotides (dsODNs). The dsODNs of said plurality have a same length of between 47 and 300 bp (e.g., between 80 and 150 bp, or between 85 and 130 bp). The dsODNs are structured according to a same template structure, which consists of an orderly set of sequence portions having respective lengths that are constant across all the dsODNs of said plurality. The orderly set of sequence portions includes a first random segment, a first sequencing adapter, a second random segment, a second sequencing adapter, and a third random segment. Such segments are consecutively arranged to form a sequence. Each of the random segments of the dsODNs of said plurality consists of essentially random permutations of nucleotides, whereas the first and second sequencing adapters of the dsODNs of said plurality consist, independently of each other, of essentially a same sequence of nucleotides. The above template structure gives rise to partly random DNA, which can be operated for verification and/or authentication purposes. The above sequence structure allows the composition to be used as a mathematical one-way function and as a physical unclonable function (PUF). I.e., it can be used as a physical fingerprint, which can be challenged to verify or authenticate a product, an object, or any entity, with which the composition is associated. The same composition can be subjected to multiple challenges, hence providing higher certainty as to an associated entity. The underlying technology is scalable. Samples of the composition can be distributed to multiple users, unlike usual PUF objects. Accordingly, the proposed composition can adequately be used for securing objects or entities. The invention is further directed to methods of producing such a composition, the use of such a composition, a set comprising such a composition associated with an entity, and methods of verifying entities associated with such compositions.
Description
Operable Random DNA
TECHNICAL FIELD
The invention relates in general to compositions comprising a plurality of double-stranded oligodeoxynucleotides (dsODNs), methods of producing such compositions, the use of such compositions, a set comprising such a composition associated with an entity, and methods of verifying entities associated with such compositions. In particular, it is directed to dsODN compositions having a specific template structure, in which essentially random segments alternate with essentially identical segments, which results in partly random DNA, which can nevertheless be operated, e.g., for verification and/or authentication purposes.
BACKGROUND
Non-biological applications of DNA have gained importance due to the unique chemical properties of nucleic acids. Single-stranded DNA contains 455 exabytes of information per gram, and molecular tools exist to write, read, and copy this information chemically. The extraordinary storage density, synthetic accessibility, and structural properties of DNA open up new fields of application. Notably, synthetic DNA has been used for digital data storage, barcoding and tracing, and steganography. Techniques to encode information in synthetic DNA are known, wherein digital information is translated (using a given translation method) into a sequence combining the four natural deoxynucleotides (nucleobases: adenine, cytosine, guanine, and thymine). The sequence is then synthesized as DNA. In this form, the data can be stored in a highly compact way (with high storage density) and for long storage durations.
In addition, DNA computation has emerged as an interdisciplinary field that makes use of the tools provided by biology, which enable operations on a molecular level (e.g., copying, hybridization, extraction) and that can be exploited to perform calculations. Nucleic acids have been successfully used to solve combinatorial problems as well as computationally hard tasks and were implemented in logic gates and for random number generation.
As the cost of chemical synthesis and sequencing of DNA has dropped dramatically with the advent of the 21st century, research in DNA information technology has become increasingly accessible, opening the field towards more advanced applications.
In parallel to these developments in DNA research, digital transformation has led to the everyday use of cryptography in applications related to authentication and encryption,
electronic access control or payment. Such applications often rely on cryptographic hash functions to protect the authenticity of information. Cryptographic hash functions are "oneway" functions that calculate an output value from an input using mathematical operations that are relatively easy to perform in one direction, but very hard to invert. Moreover, such functions are designed in such a manner that "collisions" are very improbable. I.e., it is very unlikely to find two distinct inputs that produce the same output through such a function.
Although mathematical one-way functions are widely used, advancements in quantum computing and the lack of proof for the cryptographic security of such algorithms have led to the exploration of alternative methods, such as methods exploiting physical unclonable functions (PUFs), also called physical random functions. Such functions exploit random features occurring naturally, or according to a non-deterministic process. As a result, a PUF is a physical object, which, for a given input and conditions (i.e., the challenge), provides a given output (i.e., the response to the challenge). In other words, PUFs are characterized by their ability to translate an input (challenge) into an output (response) through a physical system that is unique and cannot be replicated, such that the challengeresponse pairs (CRPs) are very difficult or impossible to predict. The outputs of a PUF can accordingly serve as a "digital fingerprint". A fingerprint can notably serve as a unique identifier. PUFs are similar to cryptographic hash functions, except that they rely on a physical source of disorder instead of number theory. PUFs have been proposed for applications in intellectual property protection, public key cryptography, and anticounterfeiting of goods and services, amongst other examples.
The present inventors set themselves the challenge to design unclonable objects that are more versatile than usual PUF objects and allow new functionalities.
SUMMARY
According to a first aspect, the present invention is embodied as a composition including a plurality of double-stranded oligodeoxynucleotides (dsODNs). This composition is sometimes referred to as a dsODN composition in this document. The dsODNs of said plurality have a same length of between 47 and 300 bp, preferably between 80 and 150 bp, and more preferably between 85 and 130 bp. In addition, the dsODNs are structured according to a same template structure. The template structure consists of an orderly set of sequence portions having respective lengths that are constant across all the dsODNs of said plurality. The orderly set of sequence portions includes a first random segment, a first sequencing adapter, a second random segment, a second sequencing adapter, and a third random segment. Such segments are consecutively arranged to form
a sequence. Each of the first random segment, the second random segment, and the third random segment of the dsODNs of said plurality consists of essentially random permutations of nucleotides, whereas the first sequencing adapter and the second sequencing adapter of the dsODNs of said plurality consist, independently of each other, of essentially a same sequence of nucleotides.
A template structure as proposed herein gives rise to partly random DNA, which can be operated for verification and/or authentication purposes. For this reason, the underlying DNA sequence structure is sometimes referred to as relating to "operable random DNA" (or orDNA for short). This terminology may similarly refer to dsODNs having said template structure, and which form a DNA pool or a DNA composition.
The proposed template structure of the dsODNs results in the second random segments of the dsODNs of said plurality of dsODNs being simultaneously sequenceable, or co- sequenceable, thanks to the sequencing adapters. That is, the partly (if not fully) defined sequencing adapters flanking the second random segments ensure that the latter can be simultaneously sequenced (i.e., are co-sequenceable) across the plurality of dsODNs, using a same sequencing primer having a defined sequence (as opposed to a randomized primer).
The above sequence structure allows the composition to be used in a similar way as a mathematical one-way function and as a physical unclonable function (PUF). As a result, the composition can be used as a physical fingerprint, which can be challenged to verify or authenticate a product, an object, or any entity, with which the composition is paired. Interestingly, the random (or quasi-random) information it contains results in that the composition cannot be re-generated from scratch, nor be copied and distributed by a malicious actor. What is more, the same composition can be subjected to multiple challenges, hence providing higher certainty as to an associated item or entity. Moreover, the underlying technology is scalable, and samples of the composition can be distributed to multiple users, unlike usual PUF objects. Thus, the proposed composition is more versatile than usual PUF objects and enable more functionalities. Accordingly, the proposed composition can adequately be used for securing objects or entities, in particular objects having high stakes in authenticity.
In embodiments, in each of the dsODNs of said plurality, each of the first sequencing adapter and the second sequencing adapter has, independently of each other, a length of between 13 and 30 bp, preferably between 18 and 22 bp.
In embodiments, an average Levenshtein distance between first sequencing adapters of the dsODNs of said plurality is less than fa x Li, where Li is the length of the first sequencing adapters and fa is equal to 0.45. Similarly, an average Levenshtein distance
between second sequencing adapters of the dsODNs of said plurality is less than fa x L2, where L2 is the length of the second sequencing adapters. In variants, fa is equal to 0.30 or 0.10.
In embodiments, in each of the dsODNs of said plurality, the first random segment has a length /x of between 5 and 25 bp (preferably between 6 and 10 bp), the second random segment has a length /2 of between 11 and 200 bp (preferably between 15 and 50 bp, more preferably between 18 and 22 bp), and the third random segment has a length I3 of between 5 and 25 bp (preferably between 6 and 10 bp).
In embodiments, an average Levenshtein distance between first random segments of the dsODNs of said plurality is larger than ga x llr where /1 is the length of the first random segments. Similarly, an average Levenshtein distance between second random segments of the dsODNs of said plurality is larger than ga x /2, where I2 is the length of the second random segments, and an average Levenshtein distance between third random segments of the dsODNs of said plurality is larger than ga x /3, where I3 is the length of the third random segments. In that case, ga is equal to 0.55. In variants, ga is equal to 0.70, or 0.90.
In embodiments, the orderly set of sequence portions of each of the dsODNs of said plurality further comprises two outer handle sequences, these including a first handle sequence and a second handle sequence, whereby the first handle sequence, the first random segment, the first sequencing adapter, the second random segment, the second sequencing adapter, the third random segment, and the second handle sequence, are consecutively arranged in the sequence. In addition, each of the first handle sequence and the second handle sequence consist, independently of each other, of essentially a same sequence of nucleotides.
Preferably each of the first handle sequence and the second handle sequence has, independently of each other, a length of between 1 and 30 bp, preferably between 3 and 30 bp, and more preferably between 3 and 22 bp.
In embodiments, each of the first handle sequence and the second handle sequence has, independently of each other, a length of between 10 and 30 bp, preferably between 18 and 22 bp.
In embodiments, the first handle sequence and the second handle sequence have, independently of each other, a length of between 1 and 9 bp, preferably between 4 and 7 bp.
In embodiments, terminal base pairs (bp) of the dsODNs of said plurality are at least partly random, whereby no more than 55% of the dsODNs of said plurality have the same terminal base pairs.
The dsODNs may possibly be modified. In particular, in embodiments, terminal base pairs of the dsODNs are deoxynucleotide - dideoxynucleotide base pairs, preferably selected from a group consisting of ddA-dT base pairs, ddC-dG base pairs, ddG-dC base pairs, and ddT-dA base pairs. As usual in the art, ddA = 2',3'-dideoxyadenosine, ddC = 2', 3'- dideoxycytidine, ddG = 2',3'-dideoxyguanosine, and ddT = 2',3'-dideoxythymidine, while dA, dC, dG, and dT, refer to the respective 2'-deoxynucleotides.
In embodiments, the composition comprises a total amount of DNA in the range of 0.05 ng to 1000 ng, preferably 1 ng to 25 ng.
According to another aspect, the invention is embodied as a method of producing a dsODN composition as defined above. The method comprises a first step of (a) providing an initial composition including a plurality of single-stranded oligodeoxynucleotides (ssODNs). The initial composition is also referred to as an ssODN composition in this document. The ssODNs of said plurality have a same length of between 47 and 300 bp, preferably between 80 and 150 bp, and more preferably of between 85 and 130 bp. They are structured according to a same template structure, the latter consisting of an orderly set of sequence portions having respective lengths that are, by design, constant across all the ssODNs of said plurality. The orderly set of sequence portions includes a first handle sequence, a first random segment, a first sequencing adapter, a second random segment, a second sequencing adapter, a third random segment, and a second handle sequence, which are consecutively arranged to form a sequence. Each of the first random segment, the second random segment, and the third random segment of the ssODNs of said plurality consists of essentially random permutations of nucleotides. The first handle sequence, the second handle sequence, the first sequencing adapter and the second sequencing adapter of the ssODNs of said plurality consist, independently of each other, of essentially a same sequence of nucleotides. The method further comprises a step (b) of subjecting the initial composition to a polymerase chain reaction to obtain a dsODN composition (orDNA pool or orDNA composition) as described earlier. As the skilled person understands based on the present disclosure, the ssODN composition is amplified over its entire length using a pair of primers (forward and reverse primer) that anneal to the first handle sequence and the second handle sequence, respectively.
Preferably, the method further comprises a step (c) of at least partly cleaving the first handle sequence and the second handle sequence to obtain a dsODN composition as
described earlier, wherein residual portions of the first handle sequence and the second handle sequence, if any, have a length that is less than 10 bp, preferably less than 8 bp.
In embodiments, the first handle sequence and the second handle sequence are only partly cleaved, whereby each of the first handle sequence and the second handle sequence has, independently of each other, a length of between 1 and 9 bp, preferably between 4 and 7 bp.
According to a further aspect, the invention is embodied as a set comprising an entity and a dsODN composition as defined earlier, wherein the dsODN composition is associated with the entity. Preferably, the composition is attached to an object, or a packaging thereof, corresponding to that entity.
According to yet another aspect, the invention is embodied as a use of such a dsODN composition as a physical fingerprint of an entity of interest.
According to a final aspect, the invention is embodied as a method of verifying an entity of interest associated with the dsODN composition. This method comprises performing a verification procedure, wherein this verification procedure first comprises performing a polymerase chain reaction (PCR), on the composition. The PCR involves one or more pairs of PCR primers, wherein each of the one or more pairs includes a forward PCR primer and a reverse PCR primer, which are respectively adapted to bind the first random segment and the third random segment, respectively, of at least some of the dsODNs of said plurality, so as to amplify sequences of the dsODNs based on the one or more pairs of PCR primers. Note, it is not necessarily the entire first or second random segment that is bound by the primer. To that extent, the forward PCR primer and the reverse PCR primer may be regarded as being respectively adapted to bind, at least partially, the first random segment and the third random segment. The verification procedure further comprises sequencing the amplified sequences to obtain a sequencing dataset containing sequencing reads. The sequencing dataset is preferably obtained by filtering out sequencing reads that are inconsistent with said first sequencing adapter and said second sequencing adapter.
In embodiments, the verification procedure further comprises performing a k-mer extraction analysis of at least a subset of the sequencing reads, where 7 < k < 20, preferably 8 < k < 16. The k-mer analysis is preferably restricted to a subset of most- frequently occurring ones of the sequencing reads.
Preferably, the method further comprises: assigning a given portion of the composition to an entity; subsequently receiving the given portion assigned, or a part thereof, for verification purposes, whereby said verification procedure is performed based on the received portion, or the received part thereof, to obtain a test result; and verifying said
portion of the composition by comparing the test result with a reference result as obtained by performing the same verification procedure on a reference portion of the composition.
In preferred embodiments, comparing the test result with the reference result comprises measuring a statistical similarity between two outcomes of the -mer extraction analysis, as obtained by performing said verification procedure on the given portion and the reference portion of the composition. Said two outcomes preferably consist of two sets of extracted features, which are more preferably extracted, each, in a form of a onedimensional array of numbers.
Preferably, measuring the statistical similarity between said two outcomes comprises weighting -sequences reads obtained according to the k-mer analysis in accordance with respective frequencies of occurrence. The similarity coefficient is preferably measured as a Jaccard coefficient.
In embodiments, the method further comprises mapping a set of input numbers to unique pairs of the PCR primers and mapping the input numbers to output numbers generated based on the sequencing reads obtained.
BRIEF DESCRIPTION OF THE DRAWINGS
These and other objects, features, and advantages, of the present invention will become apparent from the following detailed description of illustrative embodiments thereof, which is to be read in connection with the accompanying drawings. The illustrations are for clarity in facilitating one skilled in the art in understanding the invention in conjunction with the detailed description. In the drawings:
FIG. 1 is a schematic representation of the template structure of dsODNs according to embodiments;
FIG. 2 is a schematic representation of an exemplary implementation of a chemical unclonable function using a template structure according to embodiments;
FIG. 3 shows the relative frequency (%) of the four nucleobases A (top left), C (top right), G (bottom left) and T (bottom right) across the 21 positions of the second random segment (output sequence) as per counts in Illumina sequencing results with two different inputs and their replicates, as obtained in embodiments;
FIG. 4 shows the relative counts (y-axis) of the 10 most frequent output sequences (x- axis numbered by rank) of individual executions of a procedure as involved in embodiments;
FIG. 5 shows the relative frequency (%) of the four nucleobases A (top left), C (top right), G (bottom left) and T(bottom right) across the 21 positions of the second random segment (output sequence) as per counts in Illumina sequencing results of an example according to embodiments;
FIG. 6 is a schematic of an example procedure for chemically synthesizing random DNA sequences;
FIG. 7 is an illustrative sketch of a principle of random DNA synthesis, schematically showing growing chains on solid support employing an equimolar mix of the four DNA nucleotides, as involved in embodiments;
FIG. 8 is a sketch summarizing procedures involved in generating dsODN compositions according to embodiments and operating them as chemical unclonable functions;
FIG. 9 shows the Ct number plotted against the number of arbitrary bases in the reverse primer with sigmoidal fit, as obtained in embodiments;
FIG. 10 is a histogram of output similarity scores of like and unlike inputs across all challenges of an example, as obtained in embodiments;
FIG. 11 illustrates an experimental procedure to amplify dsODN compositions according to embodiments, using handle primers by making copies of copies, leading to multiple generations, as in embodiments;
FIG. 12 is a sketch of a procedure to remove the handles and thereby obtain a dsODN composition than cannot be copied (unclonable), as in embodiments;
FIG. 13 shows an exemplary AGE photograph showing a purified dsODN band after PCR amplification of a 121 nt long ssODN library, as in embodiments;
FIG. 14 shows exemplary AGE photographs depicting purified dsODN bands after various stages of challenge-response-pair generation, as involved in embodiments;
FIG. 15 is a sketch illustrating the general principle of a -mer extraction from sequencing data, as involved in embodiments;
FIG. 16 is a sketch illustrating a comparison performed based on a weighted Jaccard similarity, as used in embodiments;
FIG. 17 is a schematic representation of an exemplary implementation of a chemical unclonable function, as in embodiments;
FIG. 18 shows AGE photographs depicting dsODN bands after restriction digest, as obtained in embodiments;
FIG. 19 shows a gel image showing that dsODNs comprising modified handle sequences containing terminal dideoxynucleotides, as obtained in embodiments, cannot be ligated;
FIG. 20 shows an expected number of sequences (x-axis) in a pool that perfectly matches with an input primer of a given length, as in embodiments.
Compositions, methods, uses, and sets, embodying the present invention will now be described, by way of non-limiting examples.
DETAILED DESCRIPTION OF EMBODIMENTS OF THE INVENTION
The following description is structured as follows. General embodiments and high-level variants are described in section 1. Section 2 addresses particularly preferred embodiments. Section 3 includes a detailed description of the appended drawings, including the numeral references therein.
Table of contents
1. General embodiments and high-level variants . 10
1.1. Compositions of double-stranded oligodeoxynucleotides (dsODNs) . 10
1.1.1. Main features . 10
1.1.2. Comments and definitions . 11
1.1.3. Advantages of the proposed dsODN compositions . 19
1.1.4. Preferred embodiments of the dsODN composition . 22
1.1.5. Exemplary procedure to determine whether a dsODN composition complies with a template structure as in section 1.1.1 . 24
1.2. Production of dsODN compositions . 28
1.3. Practical uses of the present compositions . 30
1.3.1. Associating a dsODN composition with an entity . 30
1.3.2. Verification methods . 31
2. Examples . 35
3. Detailed description of the drawings . 61
1. General embodiments and high-level variants
1.1. Compositions of double-stranded oligodeoxynucleotides (dsODNs)
A first aspect of the invention is now described, which concerns a composition of doublestranded oligodeoxynucleotides (dsODNs). This composition is also referred to as a dsODN composition in the following. Note, in embodiments, the dsODNs are unmodified (i.e., they consist of deoxynucleotide base pairs). However, the dsODNs may possibly be modified, as in other embodiments. In particular, terminal base pairs of the dsODNs can be deoxynucleotide - dideoxynucleotide base pairs, for reasons that will become apparent later.
1.1.1. Main features
The composition includes a plurality of dsODNs. The dsODNs of said plurality have a same length of between 47 and 300 base pairs (bp). Preferably, this length is of between 80 and 150 bp, and more preferably between 85 and 130 bp.
The dsODNs of said plurality are all structured according to a same template structure. This structure consists of an orderly set of sequence portions having respective lengths. The lengths of such sequence portions are, by design, constant across all the dsODNs of said plurality.
The orderly set of sequence portions includes five segments, i.e., a first random segment, a first sequencing adapter, a second random segment, a second sequencing adapter, and a third random segment. Such segments are consecutively arranged and accordingly form a sequence such as shown in FIG. 1. Note, the first, second, and third random segments can also be referred to as first input sequence, output sequence, and second input sequence, respectively, for reasons that will become apparent later.
As their names suggests, the random segments differ, structurally, from the sequencing adapters. Namely, each of the first random segment, the second random segment, and the third random segment of the dsODNs of said plurality consists of essentially random permutations of nucleotides. That is, any two random segments, whether belonging to a same dsODN of said plurality or not, essentially differ. I.e., when comparing two random segments, most of the nucleotides of the compared segments are likely to differ, as further explained below.
On the contrary, the first sequencing adapter and the second sequencing adapter of the dsODNs of said plurality consist, independently of each other, of essentially a same sequence of nucleotides. That is, any two first sequencing adapters (i.e., belonging to two different dsODNs) are essentially equal. I.e., when comparing two first sequencing
adapters of two distinct dsODNs of said plurality, most of the nucleotides of the compared segments are identical. The same holds for the second sequencing adapters of the dsODNs of said plurality.
1.1.2. Comments and definitions
1.1.2.1. dsODN structures across the composition
The above specifications concern dsODNs of the plurality of dsODNs of the composition. However, it should be kept in mind that the dsODN composition may possibly include additional dsODNs (i.e., in addition to the plurality of dsODNs referred to above) having different characteristics. In other words, the above specifications concern at least a given subset of dsODNs of the compositions. As a whole, the dsODNs can be subject to noise, meaning that distinct subsets of dsODNs may, in principle, be identified in the composition, where the dsODNs of such subsets have slightly different lengths, meaning also that lengths of the sequence portions may slightly vary across such subsets. However, it remains that a substantially large number of dsODNs can in principle be identified in the composition, which have a same length and are structured according to a same template structure as described above.
In the following, the dsODNs refer to sequences of a same subset (i.e., corresponding to said plurality); they obey a same template structure, such that they have all a same length, and the respective lengths of their sequence portions are identical across all such dsODNs. Nevertheless, the lengths of the individual segments within a given template structure, i.e., the lengths of the first random segment vs. the length of the first sequencing adapter vs. the length of the second random segment vs. the length of the second sequencing adapter vs. the length of the third random segment may, and typically do, differ, as exemplified later.
1.1.2.2. Template structure
A "sequence portion", as used above, is a segment of a sequence. So, each of the above sequence portions are segments of a sequence spanned by a dsODN. According to the present template structure, the random permutations occurring in the random segments cause to increase the entropy of the dsODNs, while the partial order in the interspersed sequencing adapters cause to decrease it.
1.1.2.3. Consequences in terms of co-sequenceability
The proposed template structure of the dsODNs results in that the second random segments (also referred to as output sequences) of the dsODNs of said plurality of dsODNs are simultaneously sequenceable, or co-sequenceable.
As the skilled person will appreciate, "simultaneously sequenceable", or "co- sequenceable", means that the dsODNs as a whole, or at least certain segments of the dsODNs, particularly the second random segments (output sequences), can be simultaneously sequenced using a same pair of sequencing adapter primers having a defined nucleotide sequence (as opposed to random primers, i.e., a mixture of primers essentially consisting of random permutations of nucleotides).
Now, the above definitions exclude the presence of dsODNs consisting only of essentially random segments as such dsODNs are not simultaneously sequenceable, or co- sequenceable (i.e., simultaneously sequenceable), using a same sequencing primer having a defined nucleotide sequence, as further explained below. In contrast, the partly (if not fully) defined segments (the essentially identical sequencing adapters) flanking the second random segments of said plurality of dsODNs ensure that the second random segments can be simultaneously sequenced (i.e., are co-sequenceable) using a same sequencing primer having a defined nucleotide sequence. The thereby obtained sequencing reads covering the second random segment (output sequence) may also be referred to as "output sequencing reads" or, in short, "outputs".
For instance, the first sequencing adapter and the second sequencing adapter of the dsODNs of said plurality are adapted to bind to a same pair of sequencing adapter primers, to allow sequencing of all dsODNs of said plurality using that same pair of sequencing adapter primers, as in embodiments.
Preferably, the first sequencing adapter and the second sequencing adapter are adapted for binding to each of the sequencing adapter primers with a melting temperature (Tm) of between 45°C to 60°C. In embodiments, the first sequencing adapter is configured to be bound by a first sequencing adapter primer with a temperature Tm of between 45°C to 60°C and the second sequencing adapter is configured to be bound by a second sequencing adapter primer with a temperature Tm of between 45°C to 60°C, each of the first and second sequencing adapter primer independently of each other having a defined nucleotide sequence. Measuring and/or calculating Tm is within the ordinary skill. For example, Tm may be calculated using the formula Tm = 64.9 + 41 x (c + d - 16.4)/(a + b + c + d), where a, b, c and d are the number of dA, dT, dG and dC nucleotides, respectively, in the primer.
By contrast, neither of the first random segment, the second random segment, and the third random segment of the dsODNs have a sufficient statistical identity (also in terms of average, relative Levenshtein distance, as explained below) across the plurality of dsODNs to be bound by a same oligonucleotide, such as a same primer. Accordingly, it is not possible to bind any one of the first random segments, the second random segments, and
the third random segments, of all the dsODNs, using a same primer or same oligonucleotide having a defined sequence. Since each of the first random segment, the second random segment, and the third random segment, of the dsODNs consists of essentially random permutations of nucleotides, a single primer having a defined nucleotide sequence (i.e., not consisting of random permutations of nucleotides) can only bind to a fraction of any one of the first random segments, the second random segments, and the third random segments, across all the dsODNs.
The skilled person understands that, statistically, a fraction of the first random segment, the second random segment, or the third random segment of the dsODNs will be bound by a primer or an oligonucleotide having a defined sequence, even though each of the first random segment, the second random segment, and the third random segment of the dsODNs, consists of essentially (i.e., substantially) random permutations of nucleotides.
Accordingly, the dsODNs of a certain subset of the composition may be amplified by PCR using a selected pair of PCR primers having a defined nucleotide sequence, also referred to as "input primers", as opposed to a pair of random primers, i.e., a mixture of primers essentially consisting of random permutations of nucleotides. That said, different subsets of dsODNs may be amplified, depending on the nucleotide sequence of the selected pair of defined PCR primers (input primers). In detail, the first random segment and the third random segment (i.e., the first input sequence and the second input sequence, respectively) of a certain subset of dsODNs can be at least partially bound by the selected pair of defined PCR primers (input primers), whereby said subset is amplifiable by PCR using the selected pair. I.e., any subset of dsODNs that happen to comprise a first random segment and a third random segment having sufficient complementarity to the selected pair of defined PCR primers can be amplified by PCR using that selected pair, thereby generating an amplified subset of dsODNs. This PCR is also referred to as "selection PCR" (since a subset of dsODNs is amplified and thereby "selected").
Binding of the selected pair of PCR primers (input primers) to the first random segment and to the third random segment (i.e., the first input sequence and the second input sequence, respectively), as opposed to binding to the second random segment (output sequence), the first sequencing adapter or the second sequencing adapter, can be improved by using a selected pair of PCR primers having partial complementarity to a constant segment neighbouring the respective random segment. The partial complementary of the selected pair directs the forward input primer towards binding to a sequence region partly spanning the first random segment and the first sequencing adapter, while directing the reverse input primer towards binding to a sequence region partly spanning the third random segment and the second sequencing adapter.
In particular, the first input primer may comprise a stretch of 1 to 20, e.g., 3 to 18, such as 4 to 15, 5 to 15 or 8 to 15, consecutive nucleotides that are complementary to the first sequencing adapter directly adjacent to a stretch of 3 to 20, e.g., 5 to 15, 5 to 12 or 5 to 10, consecutive nucleotides that are complementary to the first random segment of a subset of dsODNs. Further, the second input primer may comprise a stretch of 1 to 20, e.g., 3 to 18, such as 4 to 15, 5 to 15 or 8 to 15, consecutive nucleotides that are complementary to the second sequencing adapter directly adjacent to a stretch of 3 to 20, e.g., 5 to 15, 5 to 12 or 5 to 10, consecutive nucleotides that are complementary to the third random segment of a subset of dsODNs. In embodiments, the forward input primer comprises a stretch of 14 consecutive nucleotides that are complementary to the first sequencing adapter directly adjacent to a stretch of 6 consecutive nucleotides that are complementary to the first random segment of a subset of dsODNs. In embodiments, the reverse input primer comprises a stretch of 14 consecutive nucleotides that are complementary to the second sequencing adapter directly adjacent to a stretch of 7 consecutive nucleotides that are complementary to the third random segment of a subset of dsODNs.
As is apparent from the present disclosure, performing the PCR on the dsODN composition using a selected pair of PCR primers having partial complementarity to a constant segment neighbouring the respective random segment (e.g., first input primer partly spanning the first sequencing adapter and the first random segment and second input primer partly spanning the second sequencing adapter and the third random segment) thus improves selective amplification of the subset of dsODNs that happen to comprise a first random segment and a third random segment having sufficient complementarity to the selected pair of defined PCR primers.
Further, after amplification of said subset of dsODNs, the amplified subset can be sequenced using a same sequencing adapter primer as defined above, thereby generating sequencing reads covering the second random segment (also referred to as output sequence). Accordingly, a specific set of sequencing reads covering the second random segment (i.e., the output sequence) can be generated from the dsODN composition depending on the nucleotide sequence of the selected pair of PCR primers (input primers). As the skilled person understands, the pair of PCR primers comprises a forward primer (forward input primer) and a reverse primer (reverse input primer).
As is apparent from the present disclosure, said specific set of sequencing reads can be analysed to generate a specific sequencing fingerprint. Optionally, said specific fingerprint may be mapped to an output number, while said selected pair of PCR primers may be
mapped to an input number, thereby linking input numbers to output numbers. This, in turn, enables a verification and/or authentication, as discussed later in detail.
As the skilled person will appreciate, the dsODN compositions (orDNA compositions) described herein may be used to generate so-called challenge-response-pairs (CRPs). For example, a "challenge" may be a selected pair of input primers (defined PCR primers for selection PCR as described above) and the corresponding "response" being the sequencing dataset obtained after sequencing the amplified (selected) subset of dsODNs. Alternatively, provided the input primers are mapped to an input number and the sequencing dataset is further analysed and mapped to an output number, the combination of input numbers and output numbers may be regarded as CRP.
In embodiments, the orderly set of sequence portions of each of the dsODNs of said plurality of dsODNs further comprises two outer handle sequences as described above. In these embodiments, binding of the selected pair of PCR primers (input primers) to the first random segment and to the third random segment (i.e., the first input sequence and the second input sequence, respectively), thereby mitigating binding of input primers to the second random segment (output sequence), can be improved by using a selected pair of PCR primers having partial complementarity to the first handle sequence and the second handle sequence, respectively. Said partial complementarity of the selected pair directs the forward input primer towards binding a sequence region partly spanning the first handle sequence and the first random segment, while directing the reverse input primer towards binding to a sequence region partly spanning the second handle sequence and the third random segment. In particular, the first input primer may comprise a stretch of 1 to 20, e.g. 3 to 18, such as 4 to 15, 5 to 15 or 8 to 15, consecutive nucleotides that are complementary to the first handle sequence directly adjacent to a stretch of 3 to 20, e.g. 5 to 15, 5 to 12 or 5 to 10, consecutive nucleotides that are complementary to the first random segment of a subset of dsODNs. Further, the second input primer may comprise a stretch of 1 to 20, e.g., 3 to 18, such as 4 to 15, 5 to 15 or 8 to 15, consecutive nucleotides that are complementary to the second handle sequence directly adjacent to a stretch of 3 to 20, e.g. 5 to 15, 5 to 12 or 5 to 10, consecutive nucleotides that are complementary to the third random segment of a subset of dsODNs. In an embodiment, the forward input primer comprises a stretch of 14 consecutive nucleotides that are complementary to the first handle sequence directly adjacent to a stretch of 6 consecutive nucleotides that are complementary to the first random segment of a subset of dsODNs. In an embodiment, the reverse input primer comprises a stretch of 14 consecutive nucleotides that are complementary to the second handle sequence directly adjacent to a stretch of 7 consecutive nucleotides that are complementary to the third random segment of a subset of dsODNs. As is apparent from the present disclosure, performing the PCR on
the dsODN composition using a selected pair of PCR primers having partial complementarity to the first handle sequence and the second handle sequence (i.e., first input primer partly spanning the first handle sequence and the first random segment and second input primer partly spanning the second handle sequence and the third random segment) thus improves selective amplification of the subset of dsODNs that happen to comprise a first random segment and a third random segment having sufficient complementarity to the selected pair of defined PCR primers.
1.1.2.4. Definitions in terms of Levenshtein distances
In practice, the fact that each of the first sequencing adapter and the second sequencing adapter consists of essentially a same sequence of nucleotides (whereby that sequence can differ between the first and second sequencing adapter) means that their average, relative Levenshtein distance (ARLD) is strictly smaller than 0.50. That is, the average LD across the relevant sequence portions is less than half their respective lengths. In preferred embodiments, however, the ARLD is smaller than 0.45, and preferably smaller than 0.30 (e.g., smaller than 0.20, 0.15 or even 0.10, on average). The relative LD is the ratio of the LD to the length of the compared segments, which length is identical, by definition, for dsODNs of said plurality. Otherwise stated, the average LD across the first and second sequencing adapters of the dsODNs of said plurality can be formulated as being strictly less than fa x /, where fa is equal to 0.50. In preferred variants, however, this average LD is less than fa x /, where fa is equal to 0.45, 0.30, or even 0.10. Note, the ARLD is calculated as an arithmetic mean of the relative LDs computed over each pair of dsODNs in a set of interest.
Such Levenshtein distances translate into similarities, here expressed as identity percentages. I.e., in embodiments, each of the first sequencing adapter and the second sequencing adapter are, independently of each other, strictly more than 50% identical, preferably at least 55% identical, more preferably at least 70% identical, and even more preferably at least 90% identical, across all the dsODNs of said plurality, on average. For example, in embodiments, the first sequencing adapters are at least 55% identical across all dsODNs of said plurality, on average. And similarly, the second sequencing adapters are at least 55% identical across all the dsODNs of said plurality, on average. Note, the identity percentage figures given above are directly calculated from the ARLD figures provided above. Still, various algorithms can be used to compute such identity percentages, as exemplified later. In addition, the precision of the figures provided in this document is generally given by the last digit.
Such relative LD values of the first sequencing adapter and the second sequencing adapter allow pairs of dsODNs of said plurality to be simultaneously sequenceable, or co- sequenceable, as discussed above.
Note, in embodiments, the first sequencing adapter sequence and the second sequencing adapter sequence are meant to be identical by design, i.e., as per the fabrication method used (as discussed later in detail). However, they may, in practice, contain deviations and/or errors inherent to vagaries of the synthesis process used.
Conversely, the fact that each of the first random segment, the second random segment, and the third random segment of the dsODNs of said plurality consists of essentially random permutations of nucleotides typically means that their average, relative LD is strictly larger than 0.50, preferably larger than 0.55, and more preferably larger than 0.70 (e.g., larger than 0.80 or 0.90), on average. That is, the average LD across the relevant sequence portions is more than half their respective lengths /. I.e., the average LD across each of the first, second, or third random segments in the dsODNs of said plurality strictly larger than 0.50 x /. In preferred embodiments, this average LD is larger than ga x /, where ga is equal to 0.55, 0.70, or 0.90. Again, such figures can be translated in terms of maximal identity percentages. That is, the random segments have strictly less than 50% sequence identity across all the dsODNs of said plurality. Preferably, they have less than 45% sequence identity, more preferably less than 30%, and even more preferably less than 10% sequence identity.
As one understands, the first sequencing adapter sequence and the second sequencing adapter sequence are essentially (i.e., substantially) identical across the dsODNs of said plurality, whereas the random segments of the dsODNs of said plurality consist of essentially (i.e., substantially) random permutations of nucleotides.
Other distance metrics could be considered, starting with the Damerau-Levenshtein distance (DLD), the Smith-Waterman similarity (SWS), or the Needleman-Wunsch similarity (NWS), as usual in the art. Note, all such metrics lead to roughly similar results. For example, consider the two sequence portions pi = AGGTCCCAAAG (SEQ ID NO: 79) and P2 AGGTCCCAAAA (SEQ ID NO: 80). This example yields LD(pi, P2) = DLD(pi, P2) = 1, which, when normalized to the length of such sequences (the length is equal to 11 in this example), gives rise to an identity percentage of 1 - 1/11 = 10/11 « 90.91%. This number is identical to the relative similarity resulting from the SWS metric and close to the relative similarity resulting from the NWS metric, i.e., equal to 9/11 in the above example.
In this document, however, the Levenshtein distance is chosen for convenience, as the LD algorithm does not require any user input as to the alignment method, whereas the SWS
and NWS methods rely on distinct alignment methods and can be specified different gap penalty parameters. That is, the LD metric is clearer.
Moreover, various techniques for determining nucleic acid sequence identity are known in the art. Typically, such techniques revolve around determining the nucleotide sequences of the DNA (or, in the present context, the dsODNs) and comparing these sequences to a second DNA sequence.
In general, the concept of "identity" refers to an exact nucleotide-to-nucleotide correspondence of two DNA sequences, respectively. Accordingly, that a DNA sequence has a certain percent sequence identity to another DNA sequence means that, when aligned, the percentage of nucleobases are the same, and in the same relative position, when comparing the two sequences.
The percent identity of two DNA sequences is thus the number of exact matches between two aligned sequences divided by the length of the shorter sequences and multiplied by 100.
Programs for calculating the percent identity between sequences are generally known in the art.
In particular, "sequence identity" can generally be determined by alignment of two nucleic acid sequences using global or local alignment algorithms. As the skilled person understands, sequences of similar lengths are preferably aligned using a global alignment algorithm (e.g., Needleman Wunsch algorithm; cf. J. Mol. Biol. 48 (3) : 443-53) which aligns the sequences optimally over the entire length. Sequences of essentially (i.e., substantially) different lengths are preferably aligned using a local alignment algorithm (e.g., Smith Waterman algorithm; cf. J. Mol. Biol. 147 (1) : 195-197).
For example, global sequence alignments may be performed using the EMBOSS Needle sequence alignment tool [accessible via https://www.ebi.ac.uk/Tools/psa/emboss_needle/; Madeira et al., Nucl. Ac. Res., 2022, Vol. 50, Web Server issue] using default settings as indicated below:
OUTPUT FORMAT = pair; Matrix = DNAfull; GAP Open = 10; GAP EXTEND = 0.5; END GAP PENALTY = false; END GAP OPEN = 10; END GAP EXTEND = 0.5.
The percent identity can otherwise be computed directly from ARLD results, as exemplified above. That is, the percent identity can be approximated as being equal to 1 - the ARLD result, expressed as a percentage (i.e., multiplied by 100)
1.1.2.5. Maximal and minimal Levenshtein distances
In preferred embodiments, the maximal Levenshtein distance between any two sequencing adapters corresponding to the same dsODN segment is less that f x L, where L is the constant length of this sequencing adapter and f is a fraction, which is equal to 0.45, preferably equal to 0.30, and more preferably equal to 0.10. That is, the maximal LD between any pair of first sequencing adapters is less than f x Li, where Li is the constant length of this adapter, and the maximal LD between any pair of second sequencing adapters is similarly less than f x L2, where L2 is the constant length of the second adapter. In other words, the maximal LD between any pair of first sequencing adapters or second sequencing adapters is less than a fraction f of the respective lengths of the first sequencing adapter and the second sequencing adapter, across all the dsODNs of said plurality.
Such maximal distances of the first sequencing adapter and the second sequencing adapter allow any two dsODNs of said plurality to be simultaneously sequenceable, or co- sequenceable, notwithstanding the essentially random permutations occurring in the random segments.
Conversely, in embodiments, the minimal LD between any two random segments corresponding to the same dsODN segment is larger than g x I, where / is the constant length of this segment and g is a fraction, which is equal to 0.55, preferably equal to 0.70, and more preferably equal to 0.90. That is, the minimal LD between any pair of first random segments of length /1 is larger than g x llr the minimal LD between any pair of second random segments of length I2 is larger than g x /2, and the minimal LD between any pair of third random segments of length I3 is larger than g x /3. In other words, the minimal LD between any pair of first, second, or third random segments is larger than a fraction g of the respective lengths of such segments across the dsODNs of said plurality. And, again, such values may be translated into maximal identity percentages.
1.1.3. Advantages of the proposed dsODN compositions
As a result of the essentially random permutations, it is, in principle, always possible to identify, in the composition, a subset of n unique dsODNs, in which any two dsODNs differ by at least one permutation of the first random segment, the second random segment, and/or the third random segment.
A certain level of redundancy is expected, as in every DNA pool. I.e., in the present context, there are multiple copies of a same sequence in every orDNA pool. In addition, even if the first sequencing adapter sequence and the second sequencing adapter sequence will often be meant to be identical by design, they may, in practice, contain deviations and/or errors, which will impact the extent to which dsODNs differ.
Still, it remains that it is, in principle, possible to identify two dsODNs that differ by at least one permutation of nucleotides in any subset of dsODNs obeying the exact same template structure, notwithstanding the constant sequence portions. For example, the composition may include 106 to 1013, preferably 107 to 1011, more preferably 107 to 3 x IO10 unique (i.e., distinct) dsODNs, as well as several copies of each dsODN, preferably between 5 to 1000 copies, more preferably 8 to 100 copies (e.g., 80 copies).
A template structure as proposed herein gives rise to partly random DNA, which can be operated for verification and/or authentication purposes. For this reason, the underlying DNA sequence structure is sometimes referred to as "operable random DNA", or orDNA for short, in this document. As noted earlier, the terminology "operable random DNA" and the acronym orDNA similarly refer to dsODNs having said template structure, and which form a DNA pool or a DNA composition. In particular, the above sequence structure allows the composition to be used in a similar way as a mathematical one-way function (hereafter a "one-way function" for short) and as a physical unclonable function (PUF). As a result, the composition can be used as a physical fingerprint, which can be challenged to verify or authenticate a product, an object, or any entity, with which the composition is paired, as in embodiments discussed below.
Like PUFs, the proposed composition is unique, and subject to a very low, and therefore negligible, collision probability. In addition, such a composition contains sufficient random information, so that it is practically impossible to fully analyse and reproduce it. For instance, unlike one-way functions, the proposed sequence structure of the composition is unclonable.
In more detail, in the absence of outer handle sequences (i.e., a first handle sequence and a second handle sequence, see the discussion below), the composition cannot be copied, i.e., it is unclonable. Besides, even if the composition includes handle sequences, it can still not be copied. I.e., it is unclonable if the handle sequences are less than 10 bp long and are additionally modified, a concept referred to as "modified partial handle sequences", as discussed below in detail. Accordingly, this composition cannot be copied and distributed by a malicious actor.
Nevertheless, even if the composition includes handle sequences having a length of between 10 and 30 bp (i.e., full handle sequences as defined below), thereby making the composition copiable (or clonable; cf. example 3), the dsODNs are still operable. In particular, they can still be used as a one-way function. Further, a copiable (clonable) - but operable - dsODN composition may be transformed into an unclonable (and still operable) dsODN composition, provided the full handle sequences comprise a restriction
enzyme cut site, particularly a type IIS restriction enzyme cut site, e.g., a Pie I cut site, thereby rendering the full handle sequences cleavable (cf. example 5, FIG. 12).
What is more, the same composition (irrespective of the presence and length of outer handle sequences) can be subjected to multiple challenges, hence providing higher certainty as to an associated item or entity, as discussed later in reference to other aspects of the invention (cf. example 2). In that sense, the proposed composition is more versatile than usual PUF objects and enable more functionalities. For completeness, the technological barrier required to process DNA-based materials is considerably higher than that required to process numbers and functions in silico, making the proposed solution virtually unbreakable. Therefore, the present approach is suitable for securing objects or entities, in particular objects having high stakes in authenticity.
Despite clear conceptual resemblances, the present approach implies notable differences with one-way functions and PUFs, which led the present inventors to refer to it as a "chemical unclonable function", or CUF.
In particular, the proposed orDNA structure exploits a chemical process as a source of entropy, whereby the resulting orDNA can reliably be used to determine outputs from given inputs, as with one-way functions. Interestingly, the underlying CUF technology is scalable in terms of the potential input-output pairs it supports, and can be distributed to multiple users, unlike usual PUF objects. For this other reason, the proposed orDNA structure is more versatile than usual PUF objects.
Last but not least, the orDNA can be designed so that, e.g., a manufacturer may decide in advance how many times the function can be operated by limiting the amount of dsODN composition (i.e., total amount of dsDNA) that is deposited on an item, something that is not possible with conventional one-way functions, or PUFs. In essence, a relatively low amount of dsODN composition (e.g., a dsODN composition comprising a total amount of DNA in the range of 0.05 ng to 1000 ng, preferably 1 ng to 25 ng) may only be operated 5 to 10'000 times as approximately 0.01 to 10 ng of DNA are consumed during each round of operation.
Such a possibility opens the door to new cryptographic applications, i.e., new functionalities.
The present description discusses concrete use cases of verifications/authentications, which notably enable multiple simultaneous users to verify each other in a decentralized manner.
All this is now described in detail, in reference to particular embodiments of the invention.
1.1.4. Preferred embodiments of the dsODN composition
In embodiments, in each of the dsODNs of said plurality, each of the first sequencing adapter and the second sequencing adapter has, independently of each other, a length of between 13 and 30 bp. Preferably, this length is of between 18 and 22 bp. That is, the first sequencing adapter has a length Lx of between 13 and 30 bp (or preferably between 18 and 22 bp), and the second sequencing adapter similarly has a length L2 of between 13 and 30 bp (preferably between 18 and 22 bp), although Lx does not necessarily need to be equal to L2. It has been found that the above lengths of the first and second sequencing adapters (L and L2, respectively) are particularly suitable for selectively binding sequencing adapter primers.
As noted in section 1.1.2.4, the average LD between first sequencing adapters of the dsODNs of said plurality is typically less than fa x / . Similarly, the average LD between second sequencing adapters of the dsODNs of said plurality is typically less than fa x /2, where fa is equal to 0.45 (or preferably equal to 0.30 or 0.10).
In embodiments, in each of the dsODNs of said plurality, the first random segment has a length / of between 5 and 25 bp, preferably of between 6 and 10 bp. Meanwhile, the second random segment has a length /2 of between 11 and 200 bp, preferably between 15 and 50 bp, and more preferably between 18 and 22 bp. Finally, the third random segment has a length /3 of between 5 and 25 bp, preferably between 6 and 10 bp.
As noted in section 1.1.2.4, the average LD an average LD between first random segments of the dsODNs of said plurality is typically larger than ga x / . Similarly, the average LD between second random segments of the dsODNs of said plurality is typically larger than ga x /2, while the average LD between third random segments of the dsODNs of said plurality is typically larger than ga x /3, where ga is equal to 0.55, preferably equal to 0.70, or more preferably equal to 0.90.
As evoked in section 1.1.3, the orderly set of sequence portions of each of the dsODNs of said plurality may further comprises two outer handle sequences. The outer handle sequences include a first handle sequence and a second handle sequence. In such cases, the first handle sequence, the first random segment, the first sequencing adapter, the second random segment, the second sequencing adapter, the third random segment, and the second handle sequence, are consecutively arranged in the sequence. Moreover, each of the first handle sequence and the second handle sequence consist, independently of each other, of essentially a same sequence of nucleotides. Preferably each of the first handle sequence and the second handle sequence has, independently of each other, a length of between 1 and 30 bp. Preferably, this length is of between 3 and 30 bp, and more preferably of between 3 and 22 bp.
Each of the first handle sequence and the second handle sequence preferably consist, independently of each other, of essentially a same sequence of nucleotides. In particular, the orderly set of sequence portions may include a first handle sequence and a second handle sequence, where such sequence portions are at least 55% identical, preferably at least 70%, and more preferably at least 90% identical across all the dsODNs in said plurality, or even in the whole composition, as discussed earlier. Note, such percentages may again be calculated from relative LDs. If present, the first handle sequence is located upstream and directly adjacent to the first input sequence, while the second handle sequence is located downstream and directly adjacent to the second input sequence.
The present dsODN compositions can be formed as an orDNA pool containing full handle sequences. In variants, the dsODN composition contains partial handle sequences. In the present document, handle sequences having a length of between 10 and 30 bp (preferably between 18 and 22 bp) are referred to as "full handle sequences". Conversely, handle sequences having a length of between 1 and 9 bp (preferably between 4 and 7 bp) are referred to as "partial handle sequences", as now discussed in detail.
Partial handle sequences (cf. example 5). Each of the first handle sequence and the second handle sequence may possibly have, independently of each other, a length of between 1 and 9 bp. Preferably, this length is of between 4 and 7 bp. As defined herein, a handle sequence refers to any one of the two outer handle sequences that the present dsODNs may optionally include, as in preferred embodiments.
Modified partial handle sequences. Partial handle sequences may possibly be modified, a concept that is referred to as "modified partial handle sequences", which is now discussed in detail. Modified partial handle sequences contain a base pair formed by a dideoxynucleotide and a standard deoxynucleotide, where the dideoxynucleotide is present on the 3' end of each strand of the dsODN cf. example 5). Advantageously, such modification of the outer handle sequences prevent ligation, and consequently render the dsODNs containing such modified handle sequences unclonable. Note, the full handle sequences do not contain this modification. This modified base pair is formed by a dideoxynucleotide and a standard deoxynucleotide. It can be introduced by first digesting dsODNs having full handle sequences with a restriction enzyme (preferably a type IIS restriction enzyme, such as Piel), which produces a 5' overhang, and then using a DNA polymerase to blunt the sticky end with a dideoxynucleotide, e.g., a T7 polymerase without 3'-5' exonucleoase activity, which is preferably configured to incorporate ddATP, ddGTP, ddCTP, ddTTP with equal preference.
A commercially available example of such a modified T7 polymerase is the so-called Sequenase (commercially available from Thermo Fisher). As the skilled person knows, in
order to enable the cleavage of a full handle sequence to obtain a partial handle sequence, the full handle sequence must include a suitable restriction site, i.e., a restriction site that is recognized by the respective restriction enzyme, e.g., Piel.
In preferred embodiments, full handle sequences are initially employed to generate a modified partial handle sequence that contains a randomized base pair at the restriction enzyme cut site. This has the advantage that, after cleavage with a type IIS restriction enzyme, the 5' overhang resulting from the cleavage is randomized. Consequently, after the sticky end is converted into a blunt end with a dideoxynucleotide using a polymerase, the resulting base pair formed by the dideoxynucleotide and the standard deoxynucleotide is random, i.e., the 3' terminal position of each strand may consist of a matching base pair that is randomly selected from all four nucleotides, preferably a random selection from ddA-dT, ddC-dG, ddG-dC and ddT-dA base pairs, where the 3'-terminal position of each strand is a dideoxynucleotide. Cleaving the first and the second handle sequences (full handle sequences), if present, followed by enzymatic modification (e.g., using Sequenase) to obtain modified partial handle sequences, advantageously makes the composition unclonable.
In embodiments, terminal base pairs of the dsODNs of said plurality are at least partly random, whereby no more than 55% of the dsODNs of said plurality have the same terminal base pairs. For instance, if Ns (e.g., Ns > 1000) sequences are arbitrarily selected from the pool, a given sequence position is understood as being random if no more than 55% of the analysed sequences contain the same base at that position. This way, it is also possible to use only two of the four bases to generate a function (which would reduce entropy by half, but would otherwise still work the same), or of mixing randomly synthesized with non-randomly synthesized strands.
In embodiments, terminal base pairs of the dsODNs are deoxynucleotide - dideoxynucleotide base pairs, preferably selected from a group consisting of ddA-dT base pairs, ddC-dG base pairs, ddG-dC base pairs, and ddT-dA base pairs.
For completeness, the present compositions may advantageously comprise a total amount of DNA in the range of 0.05 ng to 1000 ng, preferably 1 ng to 25 ng. That is, the amount of DNA present in the composition limits the number of times the DNA pool can be analysed.
1.1.5. Exemplary procedure to determine whether a dsODN composition complies with a template structure as in section 1.1.1
The template structure described in section 1.1.1 may be verified by detecting the nucleotide composition using PCR and sequencing.
A composition as described in section 1.1.1 comprises a first sequencing adapter and a second sequencing adapter (which can also be referred to as 'forward' and 'reverse' sequencing adapters, respectively), a first and a second input sequence and an output sequence.
This verification requires to show that the first sequencing adapter and the second sequencing adapter consist, independently of each other, of essentially a same sequence of nucleotides (e.g., the average Levenshtein distance across the dsODN composition is small, as discussed earlier). This is the case for example, if a pair of primers exists that are complementary to the large majority of the adapter sequences across the dsODN composition and with which a PCR can be performed. Furthermore, it has to be shown that the first input sequence and the second input sequence (first and third random segments, respectively) consist of essentially random permutations of nucleotides (e.g., their average Levenshtein distance is large, as discussed earlier). This may be achieved by using modified primers that partially bind to either the first or the second sequencing adapter on the 3' end, but in addition contain arbitrarily defined nucleotides (dA, dT, dC, dG) at the 5' end for PCR and recording the Ct-value in dependence of the number of arbitrary nucleotides. This dependence is separately measured for the first and the second sequencing adapter. Lastly, it has to be shown that the output sequences (second random segments) consist of essentially random permutations of nucleotides (e.g., this corresponding to large average Levenshtein distances). This may be achieved by performing PCR with primers that bind to the first and second sequencing adapters and subjecting the thus amplified sequences to next generation sequencing.
Exemplary test protocol:
1. Subject 1 ng of the dsODN composition to PCR using sequencing adapter primers
2. Subject the amplified sequences to Illumina sequencing using at least 1'000'000 clusters. Filter the reads to only analyse the sequence portion located between the first and the second sequencing adapter portions (a Hamming distance of the sequencing adapter portions of up to 3 may be tolerated). The number of reads passing the filter should be at least 10'000. The respective sequence portions are subsequently analysed for their average relative Levenshtein distance.
3. Subject 1 ng of the dsODN composition to quantitative PCR (qPCR) using various primer sets (i.e., various pairs of primers) as described below. Each of the qPCR reactions is run under equal conditions, except that the sequence of the sequencing adapter primers varies for each qPCR (in one series of qPCR reactions, the first sequencing adapter primer remains constant, while the second sequencing adapter primer is varied, whereas in a second series of qPCR reactions, the second
sequencing adapter primer remains constant, while the first sequencing adapter primer is varied). A common three-step qPCR protocol can be used for each qPCR reaction, and the qPCR reactions can be performed on a commercially available qPCR platform (e.g., Roche Lightcycler 480 system). Each qPCR programme should include at least 35 cycles.
The following series of qPCR reactions may be performed:
3.1 In a first qPCR, primers with perfect complementarity to the first and second sequencing adapter portions are used (i.e. the primers only bind to the sequencing adapter portions and amplify the portion in between).
3.2 In a second qPCR, the first sequencing adapter primer (i.e., the forward sequencing adapter primer) remains unchanged, while the second sequencing adapter primer (i.e., the reverse sequencing adapter primer) is shortened by one nucleotide on the 3' end and lengthened by an arbitrary nucleotide on the 5' end compared to the pair of sequencing primers of the first qPCR (step 3.1), whereby the nucleotide arbitrarily added on the 5' end should not be the same one as the one removed from the 3' end.
3.3 In a third qPCR, the first sequencing adapter primer remains unchanged, while the second sequencing adapter primer is shortened by one nucleotide on the 3' end and lengthened by an arbitrary nucleotide on the 5' end, compared to the second qPCR (step 3.2), whereby the nucleotide arbitrarily added on the 5' end should not be the same one as the one removed from the 3' end.
3.4In a fourth qPCR, the first sequencing adapter primer remains unchanged, while the second sequencing adapter primer is shortened by one nucleotide on the 3' end and lengthened by an arbitrary nucleotide on the 5' end compared to the third qPCR (step 3.3), whereby the nucleotide arbitrarily added on the 5' end should not be the same one as the one removed from the 3' end.
3.5 Further qPCR reactions are performed, wherein for each further qPCR reaction, the second sequencing adapter primer is further varied in line with the first, second and third qPCR reaction, i.e. for each further qPCR, the first sequencing adapter primer remains unchanged, while the second sequencing adapter primer is shortened by one nucleotide on the 3' end and lengthened by an arbitrary nucleotide on the 5' end compared to the previous qPCR, whereby the nucleotide arbitrarily added on the 5'
end should not be the same one as the one removed from the 3' end. This procedure is followed at least until the second sequencing adapter primer fully consists of arbitrary nucleotides, or until the Ct-value reaches a plateau relative to the previous PCRs, or until no PCR-signal can be measured. A plateau is here defined as follows: The mean Ct value change over at least 3 consecutive PCRs as per the description of changing primers above is <0.5, after having been >0.5 for at least 3 consecutive PCRs as per the description of changing primers. Any adapted primer should not differ by more than 15% in GC-content relative to the original primer, i.e. the primer used for the first PCR. Optimally, the Tm (melting temperature) of all used primers is within a range of ± 3 °C.
3.6The measured Ct-values are plotted against the number of arbitrary nucleotides for each of the second sequencing adapter primers. The sole data points considered are those for which a Ct-value can be successfully determined.
3.7The data are fitted according to a formula describing a sigmoidal curve,
corresp ~onds to the maximum of the function, A2 to the minimum, loglO(xo) (i.e. the decade logarithm of xO) to the centre of the function, and p to the hill slope.
4. The procedure described in step 3 tests the fraction of random permutations of nucleotides in the third random segments (second input sequences) of dsODNs. To similarly test the first random segments (first input sequences) of the dsODNs, steps 3.2-3.5 are repeated analogously, wherein the second sequencing adapter primer remains constant in each qPCR reaction, whereas the first sequencing adapter primer is varied in each qPCR reaction as described above with respect to the second sequencing adapter primer.
For example, the tested dsODN composition complies with the template structure as defined in section 1.1.1 if the following criteria are fulfilled:
1. The average relative LDs of sequence reads for the portion amplified in between the sequencing adapters, as per step 1 of the exemplary test protocol, is above 0.45.
2. The p-parameter of the fitted curve as described in step 3.7 of the exemplary test protocol for sampling the first input is between 0.15 and 0.4.
3. The p-parameter of the fitted curve as described in step 3.7 of the exemplary test protocol for sampling the second input is between 0.15 and 0.4.
1.2. Production of dsODN compositions
Another aspect of the invention concerns a method of producing a dsODN composition as described in section 1.1 (cf. examples 1 and 4; step 202 in FIG. 2, step 802 in FIG. 8). Essentially, this method comprises two steps (a) and (b). First the method comprises (a) providing an initial composition, which includes a plurality of single-stranded oligodeoxynucleotides, or ssODNs for short (for example referred to as "Library 1" in example 2; cf., e.g., reference 201 in FIG. 2 and reference 801 in FIG. 8; synthesis of an ssODN composition is illustrated, e.g., in FIG. 6 and FIG. 7). Next, the method comprises (b) subjecting the initial composition to a PCR.
The ssODNs of said plurality have a same length of between 47 and 300 bp. Preferably, this length is of between 80 and 150 bp, and more preferably between 85 and 130 bp. Like the dsODNs, the ssODNs are structured according to a same template structure, the latter consisting of an orderly set of sequence portions having respective lengths that are constant across all the ssODNs of said plurality.
Consistently with the double-stranded sequences described earlier, the orderly set of sequence portions of the ssODNs includes a first handle sequence, a first random segment, a first sequencing adapter, a second random segment, a second sequencing adapter, a third random segment, and a second handle sequence. All such segments are consecutively arranged to form a sequence. Again, each of the first random segment, the second random segment, and the third random segment of the ssODNs of said plurality consists of essentially random permutations of nucleotides, while the first sequencing adapter and the second sequencing adapter of the ssODNs of said plurality consist, independently of each other, of essentially a same sequence of nucleotides. Moreover, the first and second handle sequences consist, independently of each other, of essentially a same sequence of nucleotides.
As said, the initial composition provided is then subjected to a PCR (step b) to obtain a dsODN composition as described earlier. As is apparent from the present disclosure, this PCR ("production PCR") is performed using a pair of primers that anneal to the first and second handle sequences, respectively (also referred to as handle primers; cf. example 1). Note, modification reactions may possibly be involved, as usual in the art. For example, such reactions may use handle primers having certain mismatches with the ssODNs to be
amplified, e.g., in order to introduce restriction enzyme cut sites that may not be present in the ssODN template.
As indicated above, the initial set of sequence portions defining the ssODN sequences and the set of sequence portions defining the dsODN sequence are structurally similar, subject to the lengths, possible modification, and possible absence of the outer handle sequences in the dsODNs, as in embodiments discussed earlier. Since the ssODNs refers to single strands, while the dsODNs refer to corresponding double strands, the dimensions and properties of the ssODNs remain consistent with the dsODNs'.
In particular, each of the ssODNs may have a length of between 47 and 300 bp and form an orderly set of sequence portions including a first random segment, a first sequencing adapter, a second random segment, a second sequencing adapter, and a third random segment, which are consecutively arranged to form a sequence. Additionally, each of the ssODNs contain outer handle sequences, i.e., a first handle sequence and a second handle sequence, as discussed above.
For example, the first random segment has a length of between 5 and 25 bp, preferably between 6 and 10 bp. The first sequencing adapter has a length of between 13 and 30 bp, preferably between 18 and 22 bp. The second random segment has a length of between 11 and 200 bp, preferably between 15 and 50 bp, more preferably between 18 and 22 bp. The second sequencing adapter has a length of between 13 and 30 bp, preferably between 18 and 22 bp, and the third random segment has a length of between 5 and 25 bp, preferably between 6 and 10 bp. For completeness, each of the first sequencing adapter and the second sequencing adapter will preferably be at least 55% identical, preferably at least 70% identical, more preferably at least 90% identical across all the ssODNs, on average, as discussed earlier. Again, such percentage values may be derived from average relative LDs.
So far, the dsODNs obtained still include handle sequences. However, the handle sequences may be cleaved, at least partly, as in preferred embodiments. That is, the method may further comprise a step (c) of at least partly cleaving the first handle sequence and the second handle sequence cf. example 5). In the resulting composition, the residual portions of the first handle sequence and the second handle sequence, if any, may for instance have a length that is less than 10 bp, preferably less than 8 bp.
One may, possibly, entirely cleave the first handle sequence and the second handle sequence to obtain dsODNs that are free of handles. I.e., the dsODNs do not comprise any handle sequence (see FIG. 1). As also noted earlier, another possibility is to only partly cleave the first handle sequence and the second handle sequence to obtain a composition,
in which each of the first handle sequence and the second handle sequence, as partly cleaved, has a length of between 1 and 9 bp, preferably between 4 and 7 bp.
To summarize, the main differences between the initial ssODN library and the dsODN composition (i.e., the orDNA composition) is the double- (instead of single-) stranded feature. Besides, the resulting dsODNs are, on average, present in multiple copies and are subject to an inevitable PCR bias. The outer handle sequences are present in the ssODN library but are only optional in the dsODN composition (i.e., the orDNA composition or orDNA pool). In addition, the length of the outer handle sequences may differ between the ssDNA library and the dsODN composition (i.e., the orDNA). As is apparent from the present disclosure, the lengths (or even presence/absence) of the handle sequences in the dsODN composition depend on the handle PCR primers used for amplifying the ssDNA library. For example, PCR amplification of the ssDNA library using a pair of handle PCR primers covering (i.e., binding to) only a part of the handle sequences that are proximal to the first and the third random segment, respectively (as opposed to covering the full handle sequences including the 5' and 3' terminal nucleotides of the ssDNA), yields a dsODN composition where each dsODN comprises partial handle sequences. As the skilled person understands, the lengths of the partial handle sequences depend on the lengths of the sequence regions of the handle sequences present in the ssDNA that is covered (bound) by the handle PCR primers.
Furthermore, if the outer handle sequences are present in the form of partial handle sequences as defined above, such partial handle sequences may contain a dideoxy nucleotide at the 3'-ends of both strands. This modification can for instance be introduced by digesting the dsODNs obtained after PCR amplification of the ssDNA library with a type IIS restriction enzyme, such as Piel, which produces a 5' overhang, and then using Sequenase to blunt the sticky end with a dideoxy nucleotide. Preferably, this position is also random, meaning that in the final dsODN composition, the terminal base pairs (both ends) of the dsODNs vary between dA-ddT, dC-ddG, dT-ddA and dG-ddC base pairs (the dideoxynucleotide being present on the 3'-end of each strand). Randomisation of the terminal position of each strand is achieved by designing the ssDNA library such that a random nucleotide is present at the restriction enzyme cut site.
1.3. Practical uses of the present compositions
1.3.1. Associating a dsODN composition with an entity
In embodiments, a dsODN composition as described in section 1.1. is associated (i.e., paired) with an object (e.g., a jewel, an item of goods, a consumer goods, such as a car, an electronic device), a person, or, more generally, an entity. As defined herein, an entity
can be any object, a human being, an animal, or an organization (e.g., a legal entity). When the entity is an abstract entity such as a company, the composition is typically associated with a physical medium used by this abstract entity, i.e., a physical carrier.
In that respect, a further aspect of the invention concerns a set including an entity (e.g., an object) and a dsODN composition as described earlier, wherein the composition is associated with the entity. For example, the composition can be attached to a physical object, or any item, or a packaging thereof. This way, the composition can be used as a physical fingerprint for the corresponding entity. In particular, it can be used as a physical fingerprint for physical objects, consumer goods, or other types of objects or items of goods. For instance, a sample of the composition may be packaged and attached (e.g., glued) to an object or its packaging.
Various types of containers may be used to house the composition in or on this object. For example, a sample of dsODN composition may be encapsulated in a matrix, made of, e.g., silica, calcium phosphate, or lipid nanoparticles. The encapsulated sample may subsequently be incorporated into (or coated onto) a product or packaging material, for example a paint or a dye, a solid, semi-solid or liquid pharmaceutical, a food product, a polymer-based product, a jewel, e.g., a ring or a gemstone.
In particular, a dsODN composition as described herein can be used as a physical fingerprint of an entity, i.e., a person, an object, or any other thing), according to another aspect of the invention. Note, the solutions disclosed herein are not meant to track humans. However, a dsODN composition could possibly be assigned to a human, or any legal entity willing to do so, to help this entity authenticate themselves, based on this composition. That said, in order to assign such a composition to an entity, the composition can first be associated, e.g., affixed to, a physical object, as noted above.
1.3.2. Verification methods
Another aspect of the invention concerns a method of verifying an entity of interest (e.g., a person or an object as defined above), where this entity is associated with the dsODN composition.
Basically, the verification method revolves around performing a verification procedure comprising: (i) performing a PCR on the composition; and (ii) sequencing the amplified sequences.
The PCR involves one or more pairs of PCR primers (input primers), where each pair of primers includes a forward PCR primer and a reverse PCR primer (each pair of input primers corresponds to a given challenge). The forward and reverse PCR primers are respectively adapted to at least partially bind the first random segment and the third
random segment of at least some of the dsODNs of said plurality. This, in turn, makes it possible to amplify sequences of the dsODNs based on said one or more pairs of PCR primers. That is, the PCR causes to amplify sequences of the dsODNs that are being bound by the forward PCR primer and the reverse PCR primer of each selected pair of PCR primers.
The amplified sequences are then sequenced to obtain a sequencing dataset, i.e., a set containing sequencing reads. Note, the sequencing dataset will advantageously be obtained by filtering out sequencing reads that are inconsistent with the first sequencing adapter and the second sequencing adapter.
So, performing the verification procedure leads to sequencing reads, which can subsequently be used to challenge and verify the physical fingerprint. Various methods can be contemplated to operationalize this, such as methods based on feature extractions (i.e., feature extractors in a machine learning sense) and/or k-mer analyses. In fact, feature extraction can be applied to the sequencing reads or, even, to the results of the k-mer analysis. More generally, any mapping function can be used to map the sequencing reads onto numbers (or arrays of numbers), something that enables easy comparisons.
Preferred embodiments of the verification procedure rely on performing a k-mer extraction analysis, where, typically, 7 < k < 20. Preferably, however, the value of k is between 8 and 16. For example, using k = 16 typically provides satisfactory results. This analysis is performed with respect to at least a subset of the sequencing reads. In particular, the k- mer analysis can advantageously be restricted to a subset of most-frequently occurring sequencing reads. That is, the sequencing dataset considered in the analysis may be restricted to the top-k sequences, e.g., the K sequences that are the most frequently occurring, where K is typically between 1 and 1000. In practice, it is often sufficient to restrict the value of K to the 10 to 150 most frequently occurring sequencing reads, e.g. to the 50 most frequently occurring sequencing reads. This will normally be sufficient to suitably authenticate a composition sample.
In practice, the verification procedure is normally performed at least twice, i.e., to identify reference results and to verify entities to which the same composition is assigned, as further discussed below. The reference results can for instance be identified in the form of output numbers (e.g., arrays of numbers, such as vectors and matrices), with which respective pairs of PCR primers are associated. This verification may notably be used for authentication purposes, as noted earlier. This way, the above composition can be used as a PUF and play the role of a physical fingerprint.
For example, embodiments of the verification method involve assigning a given portion of the dsODN composition to an entity of interest, which typically is the entity one wants to
verify, though not necessarily. I.e., an intermediate carrier may be used, as noted earlier. At a later time, a verifier receives the assigned portion of the dsODN, or a part thereof, for verification purposes. From this moment on, the verifier may perform the verification procedure based on the received portion (or part thereof). This yields a test result, which can then be compared with a reference result. The reference result is obtained by performing the same verification procedure as described above, albeit on a reference portion of the composition. Importantly, the reference result can be obtained before (i.e., ex-ante) or after receiving the assigned composition (or part thereof) for verification purposes. In both cases, the comparison makes it possible to verify the (portion of the) composition as initially assigned to the entity of interest.
For example, the verification procedure may be performed by directly comparing given output numbers (as obtained for the test composition) with reference output numbers, as obtained by performing the same procedure on the reference composition (or sample thereof). In that case, the comparison is performed so as to verify that the given output numbers match the reference output numbers. Note, the reference composition is assumed to be statistically equivalent to the assigned portion. E.g., the assigned portion is a sufficiently large sample of a reference composition or, conversely, both the assigned portion and the reference composition are sufficiently samples of an initial composition.
The comparison can be eased by mapping given input numbers to both the input primers and corresponding outputs. That is, in embodiments, the verification method further comprises mapping a set of input numbers to unique pairs of the PCR primers. Moreover, the same set of input numbers is mapped to output numbers generated based on the sequencing reads obtained thanks to such primers. In that case, verifying that the given output numbers match the reference output numbers may be performed by: selecting one or more input numbers of the input set; identifying the corresponding one or more unique pairs of primers; and performing the verification PCR reaction based on the identified pairs of PCR primers. Eventually, one verifies whether the output numbers obtained in accordance with the selected input numbers match the reference output numbers.
As noted above, the verification procedure can possibly be performed based on reference results obtained ex-ante for the reference composition. Once the reference results have been obtained, one can assign the residual portion of the composition to the target entity, or only a sample thereof and, in that case, eliminate the unused part of the composition, for more security. I.e., the composition can no longer be "hacked", at least not at the manufacturer site. Only the target entity is in possession of the composition. Alternatively, the initial composition can be safely stored and then used to perform verifications a posteriori, when and if necessary. In that case, there is no need to map input numbers to
output numbers a priori. The advantage of this variant is that it does not require systematically obtaining the reference results in the first place. This may be advantageous in scenarios where entities do not systematically need to be verified.
In embodiments, the comparison between the test result and the reference involves a statistical similarity metric. That is, one measures a statistical similarity between two outcomes of the k-mer extraction analysis, as obtained by performing said verification procedure on the given portion (i.e., as corresponding to the position assigned to the entity of interest) and the reference portion of the composition.
Note, the two outcomes at issue may for instance consist of two sets of extracted features. The two sets of features can notably be extracted, each, in the form of a one-dimensional array of numbers, i.e., a vector. Various statistical analysis methods can be contemplated. In general, one will seek to favour methods that allow stable arrays (i.e., keys) to be extracted from noisy data. A convenient option is to use fuzzy extractors, as these have a tolerance for noise. In variants, autoencoders can be used, too, which can be trained to filter out noise and extract noise-free features. Noise-free features can then safely be compared to authenticate a composition sample.
In embodiments, the statistical similarity between said two outcomes is measured by weighting k-sequence reads obtained according to the k-mer analysis in accordance with respective frequencies of occurrence. The similarity coefficient is preferably measured as a Jaccard coefficient J, as assumed in FIG. 16, which can be weighted or unweighted. PCR and sequencing are noisy channels with error sources, in particular due to off-target amplification caused by similarity with other sequences, or low template concentration, relative to background.
Both factors may adversely impact the orDNA pool, where sequences can differ by a single base and only occur a few times within a large background. Such error sources may lead to similar off-target amplifications in each PCR performed with the same input primers, as the sequences similar to the target are constant within a given pool, hence the benefit of a statistical similarity analysis.
A particularly suitable procedure is one that computes a similarity score between 0 and 1 for the sequence sets obtained from individual experiments, which considers the presence or absence of given sequences in the compared sets, as well as their frequency. To calculate the score, the reads are filtered for constant adapter regions, which are expected to be present in all correctly amplified sequences. I.e., only those sequences exhibiting the expected sequencing adapter regions are considered for further analysis. Among the sequences passing the filter, only the output sections (e.g., 21 bp) covering at least 0.1% of the total read count are kept. This filter procedure aims at removing the sequence
information that is either artefactual or belongs to constant adapter regions. The remaining 21-mers can be analysed using the -mer method with k = 16, to identify all 16-mers occurring in a dataset. These signature sets of -mers are then compared between experiments. The similarity score of two sets corresponds to the Jaccard coefficient weighted by the frequency of occurrence of each k-mer. This method leads to well separated similarity distributions of identical and different keys, without any false positives or false negatives identified within the available data set.
The above embodiments have been succinctly described in reference to the accompanying drawings and may accommodate a number of variants. Several combinations of the above features may be contemplated. Examples are given in the next section.
2. Examples
Example 1: Synthesis of dsODN composition (orDNa composition, orDNA pool)
Description :
Example 1 comprised generating orDNA compositions comprising random and constant portions. First, a single-stranded library (Library 1, SEQ ID NO 1) was ordered from an external supplier. The library layout is shown in FIG. 2 (A-G and 201). Synthesis was performed on solid support using standard phosphoramidite chemistry, which is schematically described in FIG. 6. To generate the random segments (first random segment, second random segment and third random segment), the four nucleobases were added in equimolar amount, leaving it to stochastics which base is incorporated. This procedure is schematically described in FIG. 7 and results in a combinatorial library. From this ssODN library, approx. 108 sequences were arbitrarily extracted to generate a doublestranded orDNA pool via PCR (Library 1 orDNA pool, SEQ ID NO 14).
Methods:
An ssODN library (Library 1) of the composition described in FIG. 2 (A-G, 201), and of the sequence structure
ATGCGATGCAGTAAGCACTCN N N N N N N N N ACACGACGCTCTTCCGATCTN NNNNNNNNNNNNN N N N N N N NGCTCAGG ATACCAAGCTGTCCN N N N N N N N N NGATATCTGCTCGG ACCGCTA with SEQ ID NO 1 was ordered from Microsynth AG (Balgach, Switzerland). The synthesis yielded around 4 mg of DNA, equivalent to ca. 6 1016 sequences. A 5 nmol aliquot of the dried library as received from the supplier was dissolved in PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). Dilution series were performed to achieve the concentration needed to pipette the amount corresponding to the desired pool size of approx. 108 sequences into a PCR reaction. Aside from the template, the final
PCR mix contained lx KAPA SYBR FAST qPCR master mix (KAPA Biosystems, Wilmington, USA) and 0.5 pM of each primer. The primers were Handle 1.1 (SEQ ID NO 45, ATGCGATGCAGTAAGCACTC) and Handle II. I (SEQ ID NO 46, TAGCGGTCCGAGCAGATATC) primers (Microsynth AG, Balgach, Switzerland). The final reaction volume comprised 20 pl. Dilutions and reaction mixes were prepared under laminar flow. Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 30 cycles of 15 s denaturing at 95 °C, 30 s annealing at 56 °C and 30 s elongation at 72 °C, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland). The reaction was completed with 180 s of final elongation at 72 °.
DNA purification and work-up was performed using the DNA Clean & Concentrator kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's protocol. The purified product was eluted in 50 pl PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q® ; Merck, Darmstadt, Germany). Analytical agarose gel electrophoresis (AGE) was performed to confirm product size, using E-Gel EX gels (2%, Invitrogen, Thermo Fisher Scientific, Waltham, MA, USA) on a Power Snap Electrophoresis Device (Thermo Fisher Scientific). For size comparison, a 50 bp ladder (GeneRuler 50 bp DNA Ladder, ready-to-use, Thermo Fisher Scientific, Waltham, MA, USA) was used. Concentration measurements were performed using Qubit fluorometric quantification (Thermo Fisher Scientific, Waltham, MA, USA).
Results:
Approx. 4 mg of single-stranded DNA were received from the supplier. After PCR amplification, Qubit measurement of the purified product yielded a total amount of 195 ng of dsDNA. AGE quality control of the purified showed a single band in the expected length range for 121 bp without visible side products or impurities, as displayed in FIG. 13.
Interpretation of the results:
As Library 1 contains a total of 40 random positions in the input and output segments (first random segment, second random segment and third random segment), there are 4n = 440 = ca. 1024 possible sequences that arise from a combinatorial synthesis, where n denotes the number of synthetic cycles performed with the nucleotide mixture. To produce all possible combinations once on average, approx. 66 kg of single-stranded DNA would have to be produced. This estimate is based on a molecular weight of ca. 40'000 g/mol, given by an average molecular weight of 330 g/mol/nt x 121 nt. It follows that 4 mg of ssDNA are equivalent to ca. 6 x 1016 single-stranded sequences. Before PCR amplification, any duplicates only exist by chance and not by design. As approx. 0.006 ng were used in the subsequent PCR reaction, it can be derived that 195 ng PCR yield corresponds to approx. 16'000 double-stranded copies per sequence. Example 1 thus shows that a subset with a
size of approx. 108 sequences of the single-stranded Library 1 composition can be PCR- amplified using handle 1.1 and handle II. I primers to produce double-stranded orDNA of the structure described in FIG. 2 (step 203). This double-stranded composition was named Library 1 orDNA pool
( ATGCGATGCAGTAAGCACTCN N N N N N N N N ACACG ACGCTCTTCCGATCTN NNNNNNNNNNNN N N N N N N N NGCTCAGGATACCAAGCTGTCCN N N N N N N N N NG ATATCTGCTCGGACCGCTA, SEQ ID NO 14).
Example 2: Challenge response pair generation and evaluation (operation of orDNA composition; verification method)
Description :
Example 2 was performed to show that an orDNA composition such as the one generated in example 1 can be used to generate challenge-response-pairs (CRPs). Here, a challenge- response-pair refers to the use of a pair of primers (challenge) to selectively amplify sequences from the orDNA composition, the output portions of which are sequenced to give an overall output (the response). The response is unique to a given challenge, and different challenges consistently produce different responses. To generate different CRPs, different primer pairs were used, each primer pair binding to and amplifying a different subset of sequences out of the Library 1 orDNA pool in a PCR reaction. Library 1 orDNA pool (SEQ ID NO 14) as produced in example 1 comprises approx. 108 unique sequences. In order to achieve an expectation value of above 1 for the number perfectly matching sequences in the pool, the number of selective nucleotides is given by the equation 4X < 108, meaning x < 8/log(4). This relationship is shown in FIG. 20. Based on the aforementioned equation, 13 selective nucleotides (complementary to the first and second random segments of a subset of dsODNs) distributed over the two primers (6 and 7 nt, respectively) were used. The remaining nucleotides of both primers were chosen to overlap with the adjacent constant handle segments in order to stabilize the primer-template pair during PCR and to guide the primers to the correct binding position. The sequences amplified in this reaction ('selection PCR'), are termed Library 1 orDNA selection (SEQ ID NO 15). They were then processed further and prepared for sequencing. This comprises 3 consecutive PCR reactions ('trimming PCR', 'Illumina preparation T and 'Illumina preparation IT). These three steps serve the sole purpose of making the orDNA sequences emerging from selection PCR compatible with the Illumina iSeq 100 platform for Next Generation Sequencing, which was chosen as an exemplary sequencing platform for this example. The procedure is further explained in a separate section ("note to sequencing preparation" further below).
The procedure for example 2 starting from selection PCR and ending at the composition for sequencing is schematically shown in FIG. 2 (steps 204-209). The output segments (second random segment) of the amplified sequences were subsequently sequenced using next generation sequencing (NGS), and the data processed using a similarity score based on -mer analysis. The general principle of -mer extraction is illustrated in FIG. 15. The different responses in terms of generated k-mers were recorded and cross-compared for similarity via weighted Jaccard coefficient. This principle is shown in FIG. 16. FIG. 8 provides a further overview of the procedure, with steps 803-812 describing the process of challenge-response pair generation starting with the ds orDNA pool and ending with the processed sequencing data. The steps implemented in example 2 show that repetitions with the same input primers lead to similar challenge response pairs and changing the input primer distinguishably changes the output (response).
Note concerning sequencing preparation:
The Illumina NGS system requires the presence of specific DNA segments adjacent to the segment of interest that is to be sequenced. These segments serve as the starting points for the sequence reads. Furthermore, the Illumina system works with indices, which are 6-nt long sequence segments to identify which sequences belong to which sample. In case of example 2, a different index was used for the processing of each challenge as a unique identifier, which allows to run several challenges in parallel and still distinguish which response belongs to which challenge. These Illumina sequence segments, summarized under the term 'Illumina adapters', have to be added to the orDNA sequences after the selection PCR. The first step in this procedure (the 'trimming PCR'), amplifies the orDNA selection further and ensures that only sequences of the correct format, i.e., containing the second random segment (output sequence) flanked by the first sequencing adapter and the second sequencing adapter (essentially constant segments), are processed further. They are then subjected to Illumina preparation I, which is an overhang PCR. Overhang PCR uses primers that are longer than necessary and lead to a template-primer duplex with a single-stranded overhang formed by the primer. The polymerase then extends in both directions (the primer overhang and the template), leading to a product that is longer than the original PCR template by the number of base pairs corresponding to the length of the primer overhang). Thus, new short sequence portions can be introduced at will, at either end of the DNA double strand. Illumina preparation II is again an overhang PCR and adds another short portion to both ends.
Methods:
Selection PCR: All dilutions and reaction mixes were prepared under laminar flow. PCR reaction mixes contained 10 pl KAPA SYBR FAST qPCR master mix (KAPA Biosystems,
Wilmington, USA), 1 pl of the desired input I (forward) and input II (reverse) primer (10 pM), respectively, 1 pl template (1 ng/pl Library 1 orDNA from example 1) and 7 pl PCR- grade water. The challenge-response pairs differ from each other by the primer combinations that were used. The different input I and input II primer combinations are listed in Table 1 below. Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 30 cycles of 15 s denaturing at 95 °C, 30 s annealing at 62 °C and 30 s elongation at 72 °C with 180 s final elongation, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland). DNA purification and work-up was performed using the DNA Clean & Concentrator kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's protocol. The purified product was eluted in 50 pl PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). Analytical agarose gel electrophoresis (AGE) was performed to confirm product size, using E-Gel EX gels (2%, Invitrogen, Thermo Fisher Scientific, Waltham, MA, USA) on a Power Snap Electrophoresis Device (Thermo Fisher Scientific). For size comparison, a 50 bp ladder (GeneRuler 50 bp DNA Ladder, ready-to-use, Thermo Fisher Scientific, Waltham, MA, USA) was used. Concentration measurements were performed using Qubit fluorometric quantification (Thermo Fisher Scientific, Waltham, MA, USA). The purified product was designated Library 1 orDNA selection
( ATGCAGTAAGCACTCN N N N N N N N N ACACGACGCTCTTCCGATCTN NNNNNNNNNNNNNNNNN NNNGCTCAGGATACCAAGCTGTCCNNNNNNNNNNGATATCTGCTCGGA, SEQ ID NO 15).
Table 1 : List of challenges with input I and input II primers used within example 2. Sequence segments in italic overlap with the forward strand of the first handle segment (SEQ ID NO 2) of the Library 1 orDNA pool, underlined sequence segments overlap with the reverse strand of the second handle segment (SEQ ID NO 9) of the Library 1 orDNA pool. The nucleotides in bold characters bind to the complementary random input I and input II segments (first and third random segment) adjacent to the first and second handle segments, respectively. These nucleotides are termed the 'input sequences' and are also separately listed in the fourth column.
Challenge Forward primer (input I Reverse primer (input II Input number primer) primer) sequences
(input I/input II, 5'-3' direction)
1.1 infw4, SEQ ID NO: 49 inrv5, SEQ ID NO: 50 TACGAC/GGCAACG
TGCAGTAAGCACTCTACGAC TCCGAGCAGATATCGGCAACG
1.2 infw4, SEQ ID NO: 49 inrv5, SEQ ID NO: 50 TACGAC/GGCAACG
TGCAGTAAGCACTCTACGAC TCCGAGCAGATATCGGCAACG
2.1 infw4.1, SEQ ID NO 51 inrv5, SEQ ID NO 50 TACGAT/GGCAACG
TGCAGTAAGCACTCTACGAT TCCGAGCAGATATCGGCAACG
2.2 infw4.1, SEQ ID NO 51 inrv5, SEQ ID NO 50 TACGAT/GGCAACG
TGCAGTAAGCACTCTACGAT TCCGAGCAGATATCGGCAACG
3.1 infw4, SEQ ID NO 49 inrv5.2, SEQ ID NO 52 TACGAC/AGCAACG
TGCAGTAAGCACTCTACGAC TCCGAGCAGATATCAGCAACG
3.2 infw4, SEQ ID NO 49 inrv5.2, SEQ ID NO 52 TACGAC/AGCAACG
TGCAGTAAGCACTCTACGAC TCCGAGCAGATATCAGCAACG
4.1 infw4.3, SEQ ID NO 53 inrv5.3, SEQ ID NO 54 AGGTCG/ ATTCTTC
TGCAGTAAGCACTCAGGTCG TCCGAGCAGATATCATTCTTC
4.2 infw4.3, SEQ ID NO 53 inrv5.3, SEQ ID NO 54 AGGTCG/ ATTCTTC
TGCAGTAAGCACTCAGGTCG TCCGAGCAGATATCATTCTTC
Trimming PCR: Reaction mixes contained 10 pl KAPA SYBR FAST qPCR master mix (KAPA Biosystems, Wilmington, USA), 1 pl of forward and reverse primers (both 10 pM), 1 pl of a 1 ng/pl template solution (DNA purified from selection PCR) and 7 pl PCR-grade water. The primers used were sequencing adapter I primer (SEQ ID NO 59, ACACGACGCTCTTCCGATCT; anneals to first sequencing adapter), and sequencing adapter II primer (SEQ ID NO 60, GGACAGCTTGGTATCCTGAGC; anneals to second sequencing adapter). Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 12 cycles of 15 s denaturing at 95 °C, 30 s annealing at 56 °C and 30 s elongation at 72 °C with 180 s final elongation, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland). DNA purification and work-up was performed using the DNA Clean & Concentrator kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's protocol. The purified product was eluted in 50 pl PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). Analytical agarose gel electrophoresis (AGE) was performed to confirm product size, using E-Gel EX gels (2%, Invitrogen, Thermo Fisher Scientific, Waltham, MA, USA) on a Power Snap Electrophoresis Device (Thermo Fisher Scientific). For size comparison, a 50 bp ladder (GeneRuler 50 bp DNA
Ladder, ready-to-use, Thermo Fisher Scientific, Waltham, MA, USA) was used. Concentration measurements were performed using Qubit fluorometric quantification (Thermo Fisher Scientific, Waltham, MA, USA).
Illumina preparation I: Illumina preparation I reaction mix contained 10 pl KAPA SYBR FAST qPCR master mix (KAPA Biosystems, Wilmington, USA), 1 pl each of forward and reverse primer (both 10 pM), 1 pl of a 1 ng/pl template solution (purified DNA from previous step) and 7 pl PCR-grade water. The primers used were Illumina primer I forward (SEQ ID NO 61, ACACTCTTTCCCTACACGACGCTCTTCCGATCT) and Illumina primer I reverse (SEQ ID NO 62,
GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGGACAGCTTGGTATCCTGAGC) . Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 12 cycles of 15 s denaturing at 95 °C, 30 s annealing at 56 °C and 30 s elongation at 72 °C with 180 s final elongation, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland). Preparative agarose gel electrophoresis (AGE) was performed using E-Gel EX gels (2%, Invitrogen, Thermo Fisher Scientific, Waltham, MA, USA) on a Power Snap Electrophoresis Device (Thermo Fisher Scientific). For size comparison, a 50 bp ladder (GeneRuler 50 bp DNA Ladder, ready-to-use, Thermo Fisher Scientific, Waltham, MA, USA) was used. The desired product band was extracted and purified using the Zymoclean Gel DNA recovery kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's instructions, eluting the purified product in 50 pl PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). Concentration measurements were performed using Qubit fluorometric quantification (Thermo Fisher Scientific, Waltham, MA, USA).
Illumina preparation II: Illumina preparation II reaction mix contained 10 pl KAPA SYBR FAST qPCR master mix (KAPA Biosystems, Wilmington, USA), 1 pl each of forward and reverse primer (both 10 pM), 1 pl of a 1 ng/pl template solution (purified DNA from previous step) and 7 pl PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). The primers used were Illumina primer II forward (SEQ ID NO 63, AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGC) and Illumina primer II reverse (SEQ ID NO 64-71, depending on the sample). As explained in the section 'note to sequencing preparation', Illumina primer II reverse exists in several variants, each with a different 6-base index segment. Which Illumina primer II reverse was used for which challenge is indicated in Table 2 below. Illumina primer II forward is identical for all challenges. Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 12 cycles of 15 s denaturing at 95 °C, 30 s annealing at 56 °C and 30 s elongation at 72 °C with 180 s final elongation, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland). Preparative agarose gel electrophoresis (AGE) was performed using E-Gel EX gels (2%, Invitrogen, Thermo Fisher Scientific, Waltham, MA, USA) on a
Power Snap Electrophoresis Device (Thermo Fisher Scientific). For size comparison, a 50 bp ladder (GeneRuler 50 bp DNA Ladder, ready-to-use, Thermo Fisher Scientific, Waltham, MA, USA) was used. The desired product band was extracted and purified using the Zymoclean Gel DNA recovery kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's instruction, eluting the purified product in 50 pl water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). Concentration measurements were performed using Qubit fluorometric quantification (Thermo Fisher Scientific, Waltham, MA, USA).
Table 2: List of challenges with their corresponding Illumina primer II reverse sequences that were used for Illumina preparation. Nucleotides in italic partially overlap with Illumina primer I reverse. The nucleotides in bold are the Illumina indices, which are used to assign each read to the correct sample when several samples are run in parallel.
Challenge number Illumina II primer reverse
CAAGCAGAAGACGGCATACGAGATTCAAGTGTGACTGGAGTTCAGACGTGT
Sequencing: Library 1 Illumina orDNA pools (SEQ ID NO 18-25) carrying the Illumina adapters as described above were sequenced on an iSeq 100 system (Illumina, San Diego, CA, USA) at a final total DNA concentration of 50 pM containing 2% PhiX reference (Illumina, San Diego, CA, USA).
Data processing, read frequency analysis: A python pipeline was used to filter the sequencing reads, including only the reads containing the sequencing adapter I portion at the expected position. The output portions were then analysed by their read count for each challenge using MS Excel. Relative counts of the 10 most frequently read output portions for challenges 1.1, 1.2, 4.1 and 4.2 were plotted using the Origin 2021b software.
Data processing, analysis of base content by position: A python pipeline was used to filter the sequencing reads, including only the reads containing the first sequencing adapter segment. The output segments (second random segments) were then analysed by the relative frequencies of the four nucleobases A, C, G and T across the 21 positions of the output portion over all included reads using a MATLAB pipeline, and plotted using the Origin 2021b software.
Data processing, -mer extraction/similaritv score calculation: The FASTQ files returned by the sequencing platforms were processed with a python pipeline. This pipeline first filtered the reads, only including the sequences containing the sequencing adapter segments at the expected position, with a maximum allowed hamming distance of 3. In a secondary filtering step, only sequences were included that had a 21-mer insert (comprising the 'output' portion, or second random segment) between the sequencing adapter segments. In a third filtering step, only sequences with read counts comprising at least 0.1% of the overall index reads were included in further analysis. The 'output' portions (i.e., the second random segments of the dsODNs) of the thus filtered sequences were subsequently subjected to circular -mer extraction with k = 16. Thus, each challenge resulted in a 16-mer set. These sets were then cross-compared and for each comparison a logarithmically weighted Jaccard similarity score was calculated. If the similarity was higher than 0.37, the two sets were assigned as equal, if the similarity was below 0.37, the two sets were assigned as different.
Results:
Representative AGE quality control gel images for all the PCR stages are shown in FIG. 14. For each stage, the expected band sizes of 110 bp, 62 bp, 109 bp, and 164 bp are present. An example of read frequency analysis is displayed in FIG. 4 and an example for analysis of base content by position is given in FIG. 3. Table 3 shows the similarity scores assigned
to the responses generated from the different challenges in the form of a matrix. The table indicates the log-weighted Jaccard similarity score (rounded to the first decimal position) as calculated after -mer extraction. FIG. 10 shows a histogram of the same data, indicating the normalized counts of like and unlike output comparisons given similarity scores.
Table 3: Similarity matrix between responses to challenges performed in example 2.
Challenge 1.1 1.2 2.1 2.2 3.1 3.2 4.1 4.2
Interpretation of the results:
All three of the performed analyses demonstrate the reproducibility of the challenge- response-pair generation. Further, the results show that each amplified subset is specific to the primer pair used in selection PCR. Different input primer combinations can reliably be distinguished from each other, even if the input primer sequences only differ by the minimal distance of a single nucleotide. This is best reflected in the strong separation of similarity scores when comparing like vs. comparing unlike inputs as shown in FIG. 10. Moreover, it was demonstrated that challenging an orDNA pool using the aforementioned procedure produces a fuzzy dataset that can be converted to similarity scores using k-mer extraction and log-weighted Jaccard similarity. Using a fuzzy vault system, all of the data stemming from the same input map to the same output, while the data stemming from different inputs map to different inputs.
Example 3: Amplification of orDNA compositions fdsODN composition or orDNA
Description:
Example 3 was performed to demonstrate that an orDNA pool with full handles can be copied multiple times and retain its properties regarding challenge-response-pair (CRP) generation. An orDNA pool comprising 108 dsDNA sequences was copied by PCR, whereby the copy was again used as a template to produce more copies for a total of 5 'generations'. For each generation, a selection PCR using the same input primers was performed to show that the resulting outputs matched across all generations.
Methods:
Copying of orDNA: Each generation was synthesized by using 1 ng (approx. 80 copies of the Library 1 orDNA pool, SEQ ID NO 14) of the purified previous generation as a PCR template, as illustrated in FIG. 11. Aside from the template, each PCR mix contained lx KAPA SYBR FAST qPCR master mix (KAPA Biosystems, Wilmington, USA) and 0.5 pM of each primer (Handle 1.1 primer, SEQ ID NO 45, ATGCGATGCAGTAAGCACTC and Handle II. I primer, SEQ ID NO 46, TAGCGGTCCGAGCAGATATC, Microsynth AG, Balgach, Switzerland). The final reaction volume comprised 20 pl. Dilutions and reaction mixes were prepared under laminar flow. Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 10 cycles of 15 s denaturing at 95 °C, 30 s annealing at 56 °C and 30 s elongation at 72 °C, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland). The reaction was completed with 180 s of final elongation at 72 °C. DNA purification and work-up was performed using the DNA Clean & Concentrator kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's protocol. Concentration measurements were performed using Qubit fluorometric quantification (Thermo Fisher Scientific, Waltham, MA, USA). The purified product was eluted in 50 pl PCR-grade water (type 1, 18.2 MQ x cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). Analytical agarose gel electrophoresis (AGE) was performed to confirm product size, using E-Gel EX gels (2%, Invitrogen, Thermo Fisher Scientific, Waltham, MA, USA) on a Power Snap Electrophoresis Device (Thermo Fisher Scientific). For size comparison, a 50 bp ladder (GeneRuler 50 bp DNA Ladder, ready-to-use, Thermo Fisher Scientific, Waltham, MA, USA) was used.
Generating challenqe-response-pairs: The procedure to generate and compare challenge- response-pairs was identical to the procedure described in example 2. The entire procedure from steps 'selection PCR' to 'Illumina preparation IT was performed separately for generations 0-5 of Library 1 orDNA pool. For selection PCR, 1 ng of the respective Library 1 orDNA pool generation was used, with input I primer infw4 (TGCAGTAAGCACTCTACGAC, SEQ ID NO 49) and input II primer
inrv5TCCGAGCAGATATCGGCAACG, SEQ ID NO 50). This was followed by the consecutive steps of trimming PCR, Illumina preparation I, Illumina preparation II, and sequencing. Afterwards, all sequencing data were analysed with the data processing steps 'base content by position' and ' -mer extraction/similarity score calculation' as described in the methods of example 2. The Illumina primer II reverse sequences used for Illumina preparation II are listed in Table 4 below:
Table 4: List of challenges with their corresponding Illumina primer II reverse sequences that were used for Illumina preparation. Nucleotides in italic partially overlap with Illumina primer I reverse. The nucleotides in bold are the Illumina indices, which are used to assign each read to the correct sample when several samples are run in parallel.
Generation Used Illumina II primer reverse
0 Illumina primer II reverse index 9, SEQ ID NO 72
CAAGCAGAAGACGGCATACGAGATCTGATCGTGACTGGAGTTCAGACGTGT
1 Illumina primer II reverse index 10, SEQ ID NO 73
CAAGCAGAAGACGGCATACGAGATAAGCTAGTGACTGGAGTTCAGACGTGT
2 Illumina primer II reverse index 11, SEQ ID NO 74
CAAGCAGAAGACGGCATACGAGATGTAGCCGTGACTGGAGTTCAGACGTGT
3 Illumina primer II reverse index 12, SEQ ID NO 75
CAAGCAGAAGACGGCATACGAGATTACAAGGTGACTGGAGTTCAGACGTGT
4 Illumina primer II reverse index 13, SEQ ID NO 76
CAAGCAGAAGACGGCATACGAGATTTGACTGTGACTGGAGTTCAGACGTGT
5 Illumina primer II reverse index 14, SEQ ID NO 77
CAAGCAGAAGACGGCATACGAGATGGAACTGTGACTGGAGTTCAGACGTGT
Results:
PCR on average yielded approx. 16'000 copies per generation, as estimated by measuring the concentration of the eluted product after purification. FIG. 5 shows the base content by position across the output sequencing data resulting from example 3. Table 5 shows the similarity scores assigned to the outputs generated from a constant input across the
Library 1 orDNA pool generations in the form of a matrix. Shown is the log-weighted Jaccard similarity score as calculated after -mer extraction.
Table 5: Similarity matrix between responses given by an orDNA pool and five of its copy generations to the same challenge.
Generation 0 1 2 3 4 5
0 1.0 0.8 0.6 0.6 0.5 0.5
1 0.8 1.0 0.6 0.6 0.5 0.5
2 0.6 0.6 1.0 0.8 0.8 0.8
3 0.6 0.6 0.8 1.0 0.8 0.9
4 0.5 0.5 0.8 0.8 1.0 0.8
5 0.5 0.5 0.8 0.9 0.8 1.0
Interpretation of the Results:
As described in example 2, two compared responses yielding a similarity score above 0.37 are assigned as belonging to the same input. The results show that the responses measured across all generations fulfil that criterion, meaning the function readout is robust against copying by PCR with the handle primers. An average of 16'000 copies of the Library 1 orDNA pool means each generation comprises enough material to perform 200 PCR amplifications according to the described procedure. In theory, this translates to at least 2005 = 3.2 x 1011 challenge-response-pair executions over which the output to a given input remains constant.
Example 4: Upscaling of orDNA compositions (orDNA pools, dsODN compositions)
Description :
Example 4 comprised generating larger orDNA pools and subjecting them to challenges. The orDNA pools were synthesized analogous to the procedure described in example 1, using a composition extracted from Library 1 as a template. Approx. 1.6 x 109 and 2.6 x IO10 sequences were arbitrarily extracted from ssODN Library 1 as described in example 1 to generate two new double-stranded orDNA pools via PCR. These pools are identically structured as the Library 1 orDNA pool but contain a 16-fold and 256-fold higher amount of unique sequences. Both pools were afterwards challenged with two inputs each, which were evaluated to show that orDNA pools can be scaled up and still successfully and reproducibly generate challenge-response-pairs that can be separated from each other.
Methods:
Generating orDNA pools: A 5 nmol aliquot of the dried ssODN Library 1 as received from the supplier (SEQ ID NO 1,
ATGCGATGCAGTAAGCACTCN N N N N N N N N ACACGACGCTCTTCCGATCTN NNNNNNNNNNNNN N N N N N N NGCTCAGG ATACCAAGCTGTCCN N N N N N N N N NGATATCTGCTCGG ACCGCTA) was dissolved in PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q® ; Merck, Darmstadt, Germany). Dilution series were performed to achieve the concentration needed to pipette the amount corresponding to the desired pool size of approx. 1.6 x 109 and 2.6 x IO10 sequences, respectively, into a PCR reaction. Aside from the template, both of the final PCR mixes contained lx KAPA SYBR FAST qPCR master mix (KAPA Biosystems, Wilmington, USA) and 0.5 pM of each primer (Handle 1.1 primer, SEQ ID NO 45, ATGCGATGCAGTAAGCACTC, and Handle II. I primer, SEQ ID NO 46, TAGCGGTCCGAGCAGATATC, Microsynth AG, Balgach, Switzerland). The final reaction volume comprised 20 pl for the 1.6 x 109 pool size and 10x50 pl for the 2.6 x IO10 pool size. Dilutions and reaction mixes were prepared under laminar flow. Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 21 cycles of 15 s denaturing at 95 °C, 30 s annealing at 56 °C and 30 s elongation at 72 °C, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland). The reaction was completed with 180 s of final elongation at 72 °C.
DNA purification and work-up was performed using the DNA Clean & Concentrator kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's protocol. Analytical agarose gel electrophoresis (AGE) was performed to confirm product size, using E-Gel EX gels (2%, Invitrogen, Thermo Fisher Scientific, Waltham, MA, USA) on a Power Snap Electrophoresis Device (Thermo Fisher Scientific). For size comparison, a 50 bp ladder
(GeneRuler 50 bp DNA Ladder, ready-to-use, Thermo Fisher Scientific, Waltham, MA, USA) was used. Concentration measurements were performed using Qubit fluorometric quantification (Thermo Fisher Scientific, Waltham, MA, USA). The resulting Library 1 orDNA pools are of the structure
ATGCGATGCAGTAAGCACTCN N N N N N N N N ACACGACGCTCTTCCGATCTN NNNNNNNNNNNNN N N N N N N NGCTCAGG ATACCAAGCTGTCCN N N N N N N N N NGATATCTGCTCGG ACCGCTA, SEQ ID NO 14.
Generating challenqe-response-pairs: The procedure to generate and compare challenge- response-pairs was analogous to the procedure described in example 2. Two inputs per Library 1 orDNA pool size were tested, each in duplicates. For selection PCR with the 1.6 x 109 pool size, 1 ng of the respective orDNA pool (SEQ ID NO 14) was used. For selection PCR with the 2.6 x 1010 pool size, 23 ng of the respective orDNA pool (SEQ ID NO 14) was used. The concentration of the primers and mastermix were identical in both cases. However, the total reaction volume was 20 pl for the 1.6 x 109 pool size and 50 pl for the
2.6 x IO10 pool size. Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 25 cycles of 15 s denaturing at 95 °C, 30 s annealing at 62 °C and 30 s elongation at 72 °C, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland). The primers used for the four selection PCRs of the 1.6 x 109 orDNA pool are shown in Table 6 below.
Table 6: List of challenges with primers and corresponding input sequences used for the
1.6 x 109 orDNA pool within example 4.
Challenge Forward primer Reverse primer Input sequence number designation (input designation (input primer (input I/input II, primer I) II) 5'-3' direction)
5.1 infw5, SEQ ID NO 55 inrv5, SEQ ID NO 50 TTACGAC/GGCAACG
TGCAGTAAGCACTTTACGAC TCCGAGCAGATATCGGCAACG
5.2 infw5, SEQ ID NO 55 inrv5, SEQ ID NO 50 TTACGAC/GGCAACG
TGCAGTAAGCACTTTACGAC TCCGAGCAGATATCGGCAACG
6.1 infw5 SEQ ID NO 55 inrv6, SEQ ID NO 56 TTACGAC/TGGCAACG
TGCAGTAAGCACTTTACGAC CCGAGCAGATATCTGGCAACG
6.2 infw5, SEQ ID NO 55 inrv6, SEQ ID NO 56 TTACGAC/TGGCAACG
TGCAGTAAGCACTTTACGAC CCGAGCAGATATCTGGCAACG
The primers used for the four selection PCRs of the 2.6 x IO10 orDNA pool are shown in Table 7 below.
Table 7: List of challenges with primers and corresponding input sequences used for the 2.6 x 1O10 orDNA pool within example 4.
Challenge Forward primer Reverse primer Input sequence number designation (input designation (input (input I/input II, 5'-3'
The consecutive steps of trimming PCR, Illumina preparation I and II, sequencing and data processing were analogous to the procedure described in example 2, without pool sizespecific deviations.
The Illumina primer II reverse sequences used for Illumina preparation II for each challenge are listed in Table 8 below.
Table 8/ List of challenges with their corresponding Illumina primer II reverse sequences that were used for Illumina preparation. Nucleotides in italic partially overlap with Illumina
primer I reverse. The nucleotides in bold are the Illumina indices, which are used to assign each read to the correct sample when several samples are run in parallel.
Challenge Illumina II primer reverse sequence number
5.1 Illumina primer II reverse index 1, SEQ ID NO 64
CAAGCAGAAGACGGCATACGAGATCGTGATGTGACTGGAGTTCAGACGTGT
5.2 Illumina primer II reverse index 2, SEQ ID NO 65
CAAGCAGAAGACGGCATACGAGATACATCGGTGACTGGAGTTCAGACGTGT
6.1 Illumina primer II reverse index 3, SEQ ID NO 66
CAAGCAGAAGACGGCATACGAGATGCCTAAGTGACTGGAGTTCAGACGTGT
6.2 Illumina primer II reverse index 4, SEQ ID NO 67
CAAGCAGAAGACGGCATACGAGATTGGTCAGTGACTGGAGTTCAGACGTGT
7.1 Illumina primer II reverse index 5, SEQ ID NO 68
CAAGCAGAAGACGGCATACGAGATCACTGTGTGACTGGAGTTCAGACGTGT
7.2 Illumina primer II reverse index 6, SEQ ID NO 69
CAAGCAGAAGACGGCATACGAGATATTGGCGTGACTGGAGTTCAGACGTGT
8.1 Illumina primer II reverse index 7, SEQ ID NO 70
CAAGCAGAAGACGGCATACGAGATGATCTGGTGACTGGAGTTCAGACGTGT
8.2 Illumina primer II reverse index 8, SEQ ID NO 71
CAAGCAGAAGACGGCATACGAGATTCAAGTGTGACTGGAGTTCAGACGTGT
Results:
Table 9 below shows the similarity scores assigned to the outputs generated from the four challenges by the 1.6 x 109 orDNA pool in the form of a matrix. Shown is the weighted Jaccard similarity score as calculated after -mer extraction.
Table 9: Similarity matrix between responses to challenges performed in example 4 with the 1.6 x 109 orDNA pool.
5.2 0.9 1.0 0.2 0.2
6.1 0.2 0.2 1.0 0.7
6.2 0.2 0.2 0.7 1.0
Table 10 below shows the similarity scores assigned to the outputs generated from the four challenges by the 2.6 x IO10 orDNA pool in the form of a matrix. Shown is the weighted Jaccard similarity score as calculated after -mer extraction.
Table 10: Similarity matrix between responses to challenges performed in example 4 with the 2.6 x 1O10 orDNA pool.
7.2 0.8 1.0 0.3 0.3
8.1 0.3 0.3 1.0 0.8
8.2 0.3 0.3 0.8 1.0
Interpretation of the results:
For both Library 1 orDNA pools tested in example 4, the challenges with the same input primers show a high similarity while the challenges with different input primers show a low similarity. Like inputs can be well distinguished from unlike inputs, even though their primer sequences only differ minimally, meaning one of the hardest possible distinctions can still be made with both orDNA pools. This demonstrates the scalability and functionality of orDNA pools.
Example 5: Generating unclonable functions
Description :
Example 5 was performed to show that orDNA pools can be switched from a clonable, i.e., copiable, to an unclonable, i.e., uncopiable, state. The procedure is schematically described in FIG. 12. First, an ssODN composition with cleavable handles as shown in Figure 17 (steps 1701-1708) was synthesized using a procedure equivalent to the chemical DNA synthesis described in example 1. From this ssODN library (Library 2, SEQ ID NO 32), approx. 108 sequences were arbitrarily extracted to generate a double-stranded orDNA pool via PCR (Library 2 orDNA pool, SEQ ID NO 41). The thus generated copies all contained first handle and second handle segments with Piel restriction enzyme recognition sites at both ends. Following treatment with Piel, a large portion of both handles was removed, leaving only 5 and 6 constant nucleotides on each side, respectively, and a 5'N-overhang (N being dA, dC, dG or dT, determined by the random synthesis process for each individual sequence). This composition was termed 'Library 2 orDNA digested' (SEQ ID NOs 42 and 43, respectively). This procedure disables PCR amplification of the composition in its entirety, as the remaining constant nucleotides are insufficient for successful primer annealing due to the low melting temperature. Additionally, the overhang ends were blunted with 2',3'-dideoxy nucleotides (ddNTPs), leading to a composition of the type 'Library 2 orDNA blunted' (SEQ ID NO 44). As a consequence, the composition no longer contains any 3'OH groups, which disables ligation reactions. To what degree ligation of the entire composition is impaired was investigated experimentally. This was achieved by subjecting the composition to a ligation protocol commonly used for whole genome sequencing preparation. As ligation controls, a nondigested composition and a composition blunted with dNTPs instead of ddNTPs were used.
Methods:
OrDNA synthesis: First, an ssODN composition (Library 2, SEQ ID NO 32, ATGCGAGTCAG ATNGCACTCN N N N N N N N N ACACGACGCTCTTCCGATCTN NNNNNNNNNNNNN N N N N N N N GOTO AGG ATACC AAGCTGTCC N NNNNNNNNN G AC ATNGG ACG ACTCAGCTA) was ordered from an external supplier (Microsynth AG, Balgach, Switzerland). The library layout is shown in FIG. 17 (1701-1708). The composition differs from the one synthesized in example 1 in that it contains Piel restriction sites in both handle segments, with the cut site being located next to a randomly synthesized position. The synthetic approach was equivalent to the procedure described in example 1. A 5 nmol aliquot of the dried library as received from the supplier was dissolved in PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). Dilution series were performed to achieve the concentration needed to pipette the amount corresponding to the desired pool size of
approx. 108 sequences into a PCR reaction. Aside from the template, the final PCR mix contained lx KAPA SYBR FAST qPCR master mix (KAPA Biosystems, Wilmington, USA) and 0.5 pM of each primer (handle I. II primer, SEQ ID NO 47, ATGCGAGTCAGATNGCACTC, and handle II. II primer, SEQ ID NO 48, TAGCTGAGTCGTCCNATGTC, Microsynth AG, Balgach, Switzerland). The final reaction volume comprised 20 pl. Dilutions and reaction mixes were prepared under laminar flow. Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 30 cycles of 15 s denaturing at 95 °C, 30 s annealing at 56 °C and 30 s elongation at 72 °C, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland). The reaction was completed with 180 s of final elongation at 72 °C. DNA purification and work-up was performed using the DNA Clean & Concentrator kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's protocol. The purified product was eluted in 50 pl PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q® ; Merck, Darmstadt, Germany). Concentration measurements were performed using Qubit fluorometric quantification (Thermo Fisher Scientific, Waltham, MA, USA). Analytical agarose gel electrophoresis (AGE) was performed to confirm product size, using E-Gel EX gels (2%, Invitrogen, Thermo Fisher Scientific, Waltham, MA, USA) on a Power Snap Electrophoresis Device (Thermo Fisher Scientific). For size comparison, a 50 bp ladder (GeneRuler 50 bp DNA Ladder, ready-to-use, Thermo Fisher Scientific, Waltham, MA, USA) was used.
The composition was thereafter further amplified, using 40 x 20 pl reaction volume, each containing 1 ng of the purified composition, lx KAPA SYBR FAST qPCR master mix (KAPA Biosystems, Wilmington, USA) and 0.5 pM of each primer (handle I. II primer, SEQ ID NO 47, and handle II. II primer, SEQ ID NO 48, as above). Again, the product was purified and subjected to concentration measurement and AGE quality control. Amplification and purification yielded approx. 5 pg of DNA, termed Library 2 orDNA pool (SEQ ID NO 41, ATGCGAGTCAG ATNGCACTCN N N N N N N N N ACACGACGCTCTTCCGATCTN NNNNNNNNNNNNN N N N N N N N GOTO AGG ATACC AAGCTGTCC N NNNNNNNNN G AC ATNGG ACG ACTCAGCTA) .
Restriction digestion : From this Library 2 orDNA pool, 2 x 1 pg of the composition were subjected to a restriction digest, each of the two assays consisting of 5 pl rCutSmart buffer, 10 pl Piel enzyme (concentrated at 5U/ pl) and 18 pl PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany), combining to a total reaction volume of 50 pl per assay. The buffer and enzyme were obtained from New England Biolabs (Ipswich, MA, USA). All components were mixed on ice with the enzyme added last, then incubated at 37°C for 90 minutes. Preparative agarose gel electrophoresis was performed using a 2% agarose (Ultrapure, Thermo Fisher Scientific, Waltham, MA, USA) gel stained with gel red nucleic acid stain (Biotium, Fremont, CA, USA). The AGE was run at 100 mA/130 V for 90 min on a PowerEase 90 W device (Life Technologies, Thermo Fisher Scientific, Waltham,
MA, USA). For size comparison, a 50 bp ladder (GeneRuler 50 bp DNA Ladder, ready-to- use, Thermo Fisher Scientific, Waltham, MA, USA) was used. As a negative control, an undigested sample of Library 2 orDNA was run. The product bands were excised and purified using the Zymoclean Gel DNA recovery kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's instruction. Concentration measurements were performed using Qubit fluorometric quantification (Thermo Fisher Scientific, Waltham, MA, USA). The purified digested product was termed 'Library 2 orDNA digested'. The forward strand of the composition (SEQ ID NO 42,
NGCACTCNNNNNNNNNACACGACGCTCTTCCGATCTNNNNNNNNNNNNNNNNNNNNNGCTCA GGATACCAAGCTGTCCNNNNNNNNNNGACAT) and the reverse strand (SEQ ID NO 43, NATGTCNNNNNNNNNNGGACAGCTTGGTATCCTGAGCNNNNNNNNNNNNNNNNNNNNNAGAT CGGAAGAGCGTCGTGTNNNNNNNNNGAGTGC) each contain a 5'N-overhang.
Blunting with ddNTPs: Library 2 orDNA digested (having a 5' single nucleotide overhang on both ends) was treated with Sequenase (Thermo Fisher Scientific, Waltham, MA, USA) and equimolar ddNTPs (New England Biolabs Ipswich, MA, USA) for 3' recessed end fill-in. The reaction containing 300 ng of the restriction product was performed in IX reaction buffer provided with the sequenase, at a 3 mM total ddNTP concentration (distributed equally between ddATP, ddCTP, ddGTP, ddTTP) in a total volume of 100 pl for 1 minute at 37 °C. The product was purified using the DNA Clean & Concentrator kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's protocol and eluted in 10 pl PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). The blunted product was termed Library 2 orDNA blunted (SEQ ID NO 44).
Blunting with dNTPs: Library 2 orDNA digested (having a 5' single nucleotide overhang on both ends) was treated with Sequenase (Thermo Fisher Scientific, Waltham, MA, USA) and equimolar dNTPs (New England Biolabs Ipswich, MA, USA) for 3' recessed end fill-in. The reaction containing 300 ng of the restriction product (Library 2 orDNA digested) was performed in IX reaction buffer provided with the sequenase, at a 3 mM total dNTP concentration (distributed equally between dATP, dCTP, dGTP, dTTP) in a total volume of 100 pl for 1 minute at 37 °C. The product was purified using the DNA Clean & Concentrator kit (Zymo Research, Irvine, CA, USA) according to the manufacturer's protocol and eluted in 10 pl PCR-grade water (type 1, 18.2 MQ-cm at 24°C, Milli-Q®; Merck, Darmstadt, Germany). The blunted product was termed Library 2 orDNA dNTP blunted (SEQ ID NO 78).
Ligation test: The digested and ddNTP blunted product (Library 2 orDNA blunted, SEQ ID NO 44) was tested for ligation using standard analysis and quality controls. As controls, a composition blunted with dNTPs (Library 2 orDNA dNTP blunted, SEQ ID NO 78,
NGCACTCNNNNNNNNNACACGACGCTCTTCCGATCTNNNNNNNNNNNNNNNNNNNNNGCTCA GGATACCAAGCTGTCCNNNNNNNNNNGACATN) and the undigested product (Library 2 orDNA pool, SEQ ID NO 41,
ATGCGAGTCAG ATNGCACTCN N N N N N N N N ACACGACGCTCTTCCGATCTN NNNNNNNNNNNNN N N N N N N N GOTO AGG ATACC AAGCTGTCC N NNNNNNNNN G AC ATNGG ACG ACTCAGCTA) we re subjected to the same ligation protocol.
Results:
Piel restriction resulted in clean bands of 93 bp without residual undigested product visible on the quality control gel, displayed in FIG. 18. The data showed that ligation was unsuccessful for the composition with cut handles containing a dideoxy nucleotide at both ends (Library 2 orDNA blunted, SEQ ID NO 44), while the composition with partially removed handles blunted with regular dNTPs (Library 2 orDNA dNTP blunted, SEQ ID NO 78) and the untreated composition with fully intact handles (Library 2 orDNA, SEQ ID NO 41) were successfully ligated. This is evidenced by the results provided in FIG. 19.
Interpretation of the results:
The results obtained above show that the handle sequences can be successfully shortened or removed using a restriction enzyme, and that blunting of the produced overhang sequences with ddNTPs impairs ligation. It can therefore be concluded that thus treated orDNA is unclonable, as direct end-to-end PCR or ligation of a constant segment to both ends with subsequent PCR are the only known ways to copy a diverse DNA sequence composition. However, the composition can still be challenged with inputs, as the portions of the composition enabling the generation of challenge-response pairs are not affected by the restriction and blunting procedures. Thus, the composition can be used as a physical unclonable function. As in its clonable state, the orDNA pool can be copied extensively as shown in example 3, the issuer is able to decide how many copies to generate before switching the orDNA to its unclonable state.
Example 6: Execution of an exemplary test procedure to determine that a dsODN composition complies with a template structure as in section 1.1.1
Description :
In example 6, an experimental test as laid out in section 1.1.5 was performed to demonstrate that a template structure as described in section 1.1.1 can be experimentally verified. More specifically, it was shown that the composition of interest contained a first and a second sequencing adapter with an essentially same sequence of nucleotides, as well as a first and a second input sequence (first and third random segment) consisting of essentially random permutations of nucleotides.
The composition used for example 6 was Library 3 orDNA pool (SEQ ID NO 79, ATGCGATGCAGTAAGCACTCN N NNNNNNNNNNNNNNNNNN N ACACGACGCTCTTCCGATCTN N NNNNNNNNNNNNNNNNNNNGCTCAGGATACCAAGCTGTCCNNNNNNNNNNNNNNNNNNNNN GATATCTGCTCGGACCGCTA) .
The performed experiment comprises 20 PCR reactions using forward primers that partially bind to the first sequencing adapter and the first random segment, whereby for each PCR a different ratio of the primer nucleotides binding to either of the two template portions was chosen. The reverse primer was sequencing adapter II primer (SEQ ID NO 60, GGACAGCTTGGTATCCTGAGC) and constant over all reactions.
More specifically, the number of arbitrary nucleotides (which bind to the first input sequence) in the forward primer is increased from one PCR to the next and,s concurrently, the number of constant nucleotides (which bind to the first sequence adapter) is decreased. This corresponds to the pattern described in the test procedure in section 1.1.5, step 3. Provided the first input sequence consists of essentially random permutations of nucleotides, increasing the number of arbitrary nucleotides in one of the primers decreases the number of sequences in the composition that can be bound by the primer at a constant annealing temperature, thus increasing the Ct-value of the quantitative PCR readout. This dependence would not occur if the thus probed sequence region consisted of essentially a same sequence of nucleotides across the composition. Hence, this experiment is suitable to verify that a randomly synthesized sequence portion is adjacent to a constantly synthesized sequence portion.
Methods:
PCR with different primer pairs: All 20 PCR reactions were conducted under equal conditions, with the single exception of the varying forward primer. The reverse primer was always sequencing adapter II primer (SEQ ID NO 60, GGACAGCTTGGTATCCTGAGC) and constant over all reactions, while the forward primer changed with each PCR as per Table 11. Each PCR mix contained 1 ng of the orDNA composition (Library 3 orDNA pool, SEQ ID NO 79,
ATGCGATGCAGTAAGCACTCN N NNNNNNNNNNNNNNNNNN N ACACGACGCTCTTCCGATCTN N NNNNNNNNNNNNNNNNNNNGCTCAGGATACCAAGCTGTCCNNNNNNNNNNNNNNNNNNNNN GATATCTGCTCGGACCGCTA), 0.5 pM of each primer (Microsynth AG, Balgach, Switzerland) and lx KAPA SYBR FAST qPCR master mix (KAPA Biosystems, Wilmington, USA) in a total volume of 20 pl. Thermal cycling consisted of 180 s pre-incubation at 95 °C, followed by 30 cycles of 15 s denaturing at 95 °C, 30 s annealing at 62 °C and 30 s elongation at 72 °C, performed on a LightCycler 96 platform (Roche Diagnostics, Rotkreuz, Switzerland).
Data analysis: Ct values were determined automatically by the LightCycler 96 software (Roche Diagnostics, Rotkreuz, Switzerland). In cases where the programme failed to call a Ct value because the initial slope was too high, a Ct value of 1 was assumed per default. The Ct values were then plotted against the number of arbitrary nucleotides in the respective forward primer using the OriginPro 2021b software, and a sigmoidal fit using the dose response equation y = Al + A2 - Al )/(l + 10 ((log(x0) - x) x p)) as provided by the software was calculated.
Table 11 : List of forward primers used in example 6. Arbitrary nucleotides binding to the first input sequence are marked in bold, while nucleotides binding to the first sequencing adapter are underlined.
PCR number Forward primer sequence Number of arbitrary nt in primer
1 Test procedure primer 1, SEQ ID NO 80 1
TACACGACGCTCTTCCGATC
2 Test procedure primer 2, SEQ ID NO 81 2
CTACACGACGC 1 C 1 1 CCGAT
3 Test procedure primer 3, SEQ ID NO 82 3
CCTACACGACGCTCTTCCGA
4 Test procedure primer 4, SEQ ID NO 83 4
GCCTACACGACGC 1 C 1 1 CCG
5 Test procedure primer 5, SEQ ID NO 84 5
AGCCTACACGACGCTCTTCC
6 Test procedure primer 6, SEQ ID NO 85 6
TAGCCTACACGACGCTCTTC
7 Test procedure primer 7, SEQ ID NO 86 7
TTAGCCTACACGACGCTCTT
8 Test procedure primer 8, SEQ ID NO 87 8
CTTAGCCTACACGACGCTCT Test procedure primer 9, SEQ ID NO 88 9
CCTTAGCCTACACGACGCTC Test procedure primer 10, SEQ ID NO 89 10
TCCTTAGCCTACACGACGCT Test procedure primer 11, SEQ ID NO 90 11
CTCCTTAGCCTACACGACGC Test procedure primer 12, SEQ ID NO 91 12
TCTCCTTAGCCTACACGACG Test procedure primer 13, SEQ ID NO 92 13
ATCTCCTTAGCCTACACGAC Test procedure primer 14, SEQ ID NO 93 14
TATCTCCTTAGCCTACACGA Test procedure primer 15, SEQ ID NO 94 15
GTATCTCCTTAGCCTACACG Test procedure primer 16, SEQ ID NO 95 16
AGTATCTCCTTAGCCTACAC Test procedure primer 17, SEQ ID NO 96 17
TAGTATCTCCTTAGCCTACA Test procedure primer 18, SEQ ID NO 97 18
GTAGTATCTCCTTAGCCTAC Test procedure primer 19, SEQ ID NO 98 19
TGTAGTATCTCCTTAGCCTA Test procedure primer 20, SEQ ID NO 99 20
GTGTAGTATCTCCTTAGCCT
Results:
The results are shown in FIG 9. The dependence between the Ct value and the number of arbitrary nucleotides in the forward primer approximately follows a sigmoidal curve, with very low Ct values for a low number of arbitrary nucleotides, followed by a steep increase for a medium number of arbitrary nucleotides, before the curve starts to flatten and plateaus at a Ct value of around 19 for 16 arbitrary nucleotides and above. This translates into sigmoidal fit parameters of approx. 0.60 for Al, 18.57 for A2, 11.28 for log(xO) and 0.26 for p.
Interpretation of results:
The probed orDNA composition is used in a high initial concentration (0.05 ng/pl in the reaction) and is highly diverse, i.e. contains approx. 108 unique sequences. As all the sequences in the composition are expected to contain the first sequencing adapter portion and the second sequencing adapter portion, primer pairings where the forward primer largely overlaps with the first sequencing adapter (i.e. most primer nucleotides bind to said adapter) will bind and amplify nearly all sequences in the pool. This leads to an immediate and steeply increasing fluorescence signal, which registers as a Ct-value of 1. When the number of arbitrary nucleotides are increased, the amount of amplifiable template sequences in the composition decreases, as the forward primer can only bind and amplify the sequences with full complementarity to both primers. The more arbitrary nucleotides the forward primer contains, the less sequences have the necessary complementary nucleotides in their respective first random segments (first input segments). However, this effect is not immediately reflected in the Ct-value, as the Ct- value has a lower limit of 1, and forward primers with only a low number of arbitrary nucleotides still find sufficient template sequences to cause a steep fluorescence increase from the start. But, once the number of arbitrary nucleotides increases further, less and less matching templates are amplified and the Ct-value starts to increase measurably, as reflected in the increased slope for 7-15 arbitrary nucleotides in the primer (shown in FIG 9). The plateau that is reached with forward primers containing 16 or more arbitrary nucleotides can be explained by unspecific amplification. Due to the presence of a high background concentration of DNA, combined with the fact that the reverse primer can bind all present DNA molecules with full complementarity, a subset of the background always amplifies, corresponding to a Ct-value of approx. 19.
Such a dependence is proof that the sequence portion adjacent to the first sequencing adapter consists of essentially random permutations of nucleotides.
3. Detailed description of the drawings
FIG. 1 : Schematic representation of the template structure of dsODNs according to embodiments.
101 : First random segment (first input sequence), e.g., comprising between 5-25 essentially randomly permutated base pairs
102: First sequencing adapter, e.g., comprising between 13 - 30 bp, e.g., with a maximal average Levenshtein distance of 0.45 x Li (= length of the first sequencing adapter) across the plurality of dsODNs
103: Second random segment (output sequence), e.g., comprising between 11-200 essentially randomly permutated base pairs
104: Second sequencing adapter, e.g., comprising between 13 - 30 bp with a maximal average Levenshtein distance of 0.45 x L2 (= length of the second sequencing adapter) across the plurality of dsODNs
105: Third random segment (second input sequence), e.g., comprising between 5-25 essentially randomly permutated base pairs across the plurality of dsODNs
FIG. 2: Schematic representation of an exemplary implementation of a Chemical Unclonable Function using a template structure according to embodiments. Shown are the template structure, a method for producing a dsODN composition (orDNA composition), and an exemplary implementation of its use as a Chemical Unclonable Function (method of verification I authentication; i.e., operation of the dsODN composition) as performed in examples 1-4.
201 : Template structure of an ssODN composition as synthesized and described, e.g., in example 1 (Library 1), e.g., having a length of 121 nt.
202: Amplification with handle 1.1 and II. I primers, as described in experiment 1.
203: Template structure of dsODNs (orDNA) amplified from ssODN composition (library I) as obtained from example 1.
204: Selection PCR with input primers, as described in example 2.
205: Amplified subset of dsODN composition, e.g., comprising 110 bp. Subset as amplified by selection PCR using input primer I and input primer II, as described in example 2.
206: Trimming PCR with primers annealing to the first sequencing adapter and the second sequencing adapter (sequencing adapter primer I and sequencing adapter primer II).
207: Amplified ("selected") and trimmed dsODN obtained by PCR with sequencing adapter primer I and sequencing adapter primer II, as described in example 2.
208: Preparation for Illumina sequencing including two PCR reactions (Illumina preparation I and II) using Illumina primer I forward with Illumina primer I reverse in the first reaction and Illumina primer II forward with Illumina primer II reverse in the second reaction, as described in example 2.
209: Amplified and trimmed dsODN composition ready for sequencing (referred to Library 1 Illumina orDNA in example 2).
A: First handle sequence (handle I), e.g., comprising 20 nt, e.g., having the sequence 5'ATGCGATGCAGTAAGACTC3' (SEQ ID NO: 2)
B: First random segment (first input sequence; input I), e.g., comprising 9 nt, e.g., having the sequence 5'NNNNNNNNN3'
C: First sequencing adapter (sequencing adapter I), e.g., comprising 20 nt, e.g., having the sequence 5'ACACGACGTCTTCCGATCT3' (SEQ ID NO: 3)
D: Second random segment (output sequence), e.g., comprising 21 nt, e.g., having the sequence 5'NNNNNNNNNNNNNNNNNNNNN3'
E: Second sequencing adapter (sequencing adapter II), e.g., comprising 21 nt, e.g., having the sequence 5'GCTCAGGATACCAAGCTGTCC3' (SEQ ID NO: 4)
F: Second random segment (second input sequence; Input II), e.g., comprising 10 nt, e.g., having the sequence 5'NNNNNNNNNN3'
G: Second handle sequence (handle II), e.g., comprising 20 nt, e.g., having the sequence 5'GATATCTGCTCGGACCGCTA3' (SEQ ID NO: 5)
H-N denote the reverse complementary segments corresponding to A-G:
H: Handle I reverse, e.g., comprising 20 nt, e.g., having the sequence 3TACGCTACGTCATTCGTGAG5' (SEQ ID NO: 6)
I: Input I reverse, e.g., comprising 9nt, e.g., having the sequence 3'NNNNNNNNN5'
J: Sequencing Adapter I reverse, e.g., comprising 20 nt, e.g., having the sequence 3TGTGCTGCGAGAAGGCTAGA5' (SEQ ID NO: 7)
K: Output sequence reverse, e.g., comprising 21 nt, e.g., having the sequence 3'NNNNNNNNNNNNNNNNNNNNN5'
L: Sequencing Adapter II reverse, e.g., comprising 21 nt, e.g., having the sequence 3'CGAGTCCTATGGTTCGACAGG5' (SEQ ID NO: 8)
M: Input II reverse, e.g., comprising 10 nt, e.g., having the sequence 3'NNNNNNNNNN5'
N: Handle II reverse, e.g., comprising 20 nt, e.g., having the sequence 3'CTATAGACGAGCCTGGCGAT5' (SEQ ID NO: 9)
O: Handle I after selection PCR, e.g., comprising 15 nt, e.g., having the sequence 5'ATGCAGTAAGCACTC3' (SEQ ID NO: 10)
P: Handle II after selection PCR, e.g., comprising 14 nt, e.g., having the sequence 5'GATATCTGCTCGGA3' (SEQ ID NO: 11)
Q: Handle I reverse after selection PCR, e.g., comprising 15 nt, e.g., having the sequence 3 ACGTCATTCGTGAG5' (SEQ ID NO: 12)
R: Handle II reverse after selection PCR, e.g., comprising 14 nt, e.g., having the sequence 3'CTATAGACGAGCCT5' (SEQ ID NO: 13)
S: Illumina adapter I forward; T: Illumina adapter II forward; U: Illumina adapter I reverse; V: Illumina adapter II reverse. S, T, U and V are the sequences arising from preparation for Illumina sequencing.
FIG. 3: Relative frequency (%) of the four nucleobases A (top left), C (top right), G (bottom left) and T (bottom right) across the 21 positions of the second random segment (output sequence) as per counts in Illumina sequencing results with two different inputs (challenge 1.1 and challenge 4.1) and their replicates (challenge 1.2 and challenge 4.2), as described in example 2. The position-dependent frequencies show high similarity between replicates but a clear difference between the two inputs.
31 : Challenge 1.1; 32: Challenge 1.2; 33: Challenge 4.1; 34: Challenge 4.2
FIG. 4: Relative counts (y-axis) of the 10 most frequent output sequences (x-axis numbered by rank) of individual executions of the procedure as described in example 2, with challenges 1.1, 1.2, 4.1 and 4.2. There is a clear similarity between repetitions with the same input, and a clear difference between different inputs.
41 : Challenge 1.1, input TACGAC/GGCAACG
42: Challenge 1.2, input TACGAC/GGCAACG
43: Challenge 4.1, input AGGTCG/AATCATG
44: Challenge 4.2, input AGGTCG/ATTCATG
FIG. 5: Relative frequency (%) of the four nucleobases A (top left), C (top right), G (bottom left) and T(bottom right) across the 21 positions of the second random segment (output sequence) as per counts in Illumina sequencing results of example 3. The position-
dependent frequencies show high similarity between all generations of the Library 1 orDNA pool (cf. reference 209 in FIG. 2).
51 : Generation 0; 52: Generation 1; 53: Generation 2; 54: Generation 3; 55: Generation 4; 56: Generation 5
FIG. 6: Schematic of an example procedure for chemically synthesizing random DNA sequences as supplied by several commercial suppliers. Shown is the synthetic process for a single growing chain, whereby each cycle as illustrated by the arrows adds a single nucleotide, which is incorporated at random by the entropy of a mix of nucleotides.
61 : Start of the cycle; 62: Detritylation; 63: Incoming nucleotide, from a mix of nucleotides; 64: Coupling; 65: Oxidation; 66: Cleavage from solid support and deprotection
FIG. 7: Illustrative sketch of the principle of random DNA synthesis, schematically showing growing chains on solid support employing an equimolar mix of the four DNA nucleotides.
71 : Synthesis chamber; 72: Solid support; 73: First synthetic cycle adding a nucleotide at random; 74: Second synthetic cycle adding the next nucleotide at random; 75: Subsequent synthetic cycles 3 to n adding the next n - 2 nucleotides at random; 76: 4n possible sequences arising from the combinatorial possibilities of n random nucleotides
FIG. 8: Summarizing sketch of the procedures involved in generating orDNA pools (dsODN compositions) and operating them as chemical unclonable functions. Shown are the various steps from generating a dsODN composition (orDNA composition) to receiving responses (outputs) to challenges (inputs).
801 : ssODN composition comprising constant and random portions, with handles on both ends
802: Generating a dsODN composition (orDNA composition) from ssODN composition (also referred to as "function generation"
803: dsODN composition (orDNA composition) serving as chemical unclonable function (optionally comprising first and second handle sequences)
804: Numeric input to the function (optional); 805: Mapping of numeric input to primer sequence (optional); 806: Input primers
807: Amplification of a subset of the dsODN composition (orDNA composition) using a pair of input primers ("selection PCR"), followed by sequencing of the amplified subset, thereby obtaining sequencing reads covering the second random segments of the amplified subset (output sequences)
808: Sequencing reads
809: Data processing using k-mer extraction (optional); 810: Set of k-mers with frequency of occurrence (optional)
811 : Calculation of similarity between k-mer sets, optional, for example using a weighted Jaccard similarity
812: Jaccard similarity score
813: Optional mapping to a numeric output, for example using a fuzzy vault
814: Optional numeric output
FIG. 9: Ct number plotted against the number of arbitrary bases in the reverse primer with sigmoidal fit, as described in test procedure steps 3.6-3.7. The equation for the curve fit is y = Al + (A2 - Al)/(1 + 10/K((log(xo) - x) x p)). Al is equal to 0.5956 ± 0.32332. A2 is equal to 18.57025 ± 0.38632. Log(xo) is equal to 11.28173 ± 0.17121, and p is equal to 0.25948 ± 0.0245. The reduced chi-square obtained is 0.46704, the reduced R- square is 0.99305, and the adjusted R-square is 0.99175. Parameters are unrounded as returned for the fit by the Origin2021b software.
FIG. 10: Histogram of output similarity scores of like and unlike inputs across all challenges of example 2. White bars represent normalized comparisons between challenges conducted with unlike inputs, black bars represent normalized comparisons between challenges with like inputs. The similarity score is the number calculated by weighted k-mer extraction (k = 16, weight = logarithmic, including all sequences contributing at least 0.1% of total sequence counts).
1001 : Range of similarity scores assigned to unlike inputs; 1002: Range of similarity scores assigned to like inputs; 1003: Threshold between similarity scores assigned to unlike/like inputs
FIG. 11: Experimental procedure to amplify dsODN compositions using handle primers by making copies of copies, leading to multiple generations, as described in example 3. Starting with approx. 16'000 copies of a composition comprising approx. 108 sequences (generation 0), the first generation was obtained by performing PCR of approx. 0.5% of generation 0 with the handle primers, yielding again approx. 16'000 copies on average. This procedure was repeated 5 times in total, each time using the previous generation as a template.
FIG. 12: Sketch of the procedure as described in example 5 to remove the handles, to thereby obtain a dsODN composition than cannot be copied (unclonable); i.e., no PCR amplification or ligation being possible.
1201 : Handle I; 1202: Input I; 1203: Sequencing Adapter I; 1204: Output; 1205: Sequencing Adapter II; 1206: Input II; 1207: Handle II; 1208: Restriction site; 1209: Restriction digest using restriction enzyme; 1210: End fill-in with dideoxy nucleotides; 1211 : Incorporated dideoxy nucleotides
FIG. 13: Exemplary AGE photograph showing the purified dsODN band (Library 1 orDNA pool) after PCR amplification of a 121 nt long ssODN library in example 1. For size comparison a 50 bp ladder is shown.
131 : 50 bp marker; 132: 100 bp marker; 133: 150 bp marker; 134: Sample band, 121 bp; 135: Ladder lane; 136: Sample lane
FIG. 14: Exemplary AGE photographs showing the purified dsODN bands after the various stages of challenge-response-pair generation as described in example 2. For size comparison a 50 bp ladder is shown on all gels.
1401 : AGE gel after 'selection PCR'; 1402: Ladder lane; 1403: Sample lane; 1404: 50 bp marker; 1405: 100 bp marker; 1406: 150 bp marker; 1407: Sample band, 110 bp; 1408: AGE gel after 'trimming PCR'; 1409: Ladder lane; 1410: Sample lane; 1411 : 150 bp marker; 1412: 100 bp marker; 1413: 50 bp marker; 1414: Sample band, 62 bp; 1415: AGE gel after 'sequencing preparation step T; 1416: Ladder lane; 1417: Sample lane; 1418: 150 bp marker; 1419: 100 bp marker; 1420: 50 bp marker; 1421 : Sample band, 109 bp; 1422: AGE gel after 'sequencing preparation step II'; 1423: Ladder lane; 1424: Sample lane; 1425: 200 bp marker; 1426: 150 bp marker; 1427: 100 bp marker; 1428: 50 bp marker; 1429: Sample band, 164 bp
FIG. 15: Schematic sketch illustrating the general principle of k-mer extraction from sequencing data.
1501 : NGS sequence reads; 1502: k-mer extraction process; 1503: Unordered set of k- mers with counts per k-mer; 1504: Example sequence, 21-mer; 1505: k-mers for k = 10 that can be extracted from the 21-mer shown above. For circular k-mer extraction of an n-mer, n k-mers are obtained.
FIG. 16: Sketch showing the comparison via weighted Jaccard similarity. Two k-mer sets as extracted from the filtered sequence reads stemming from two challenges are compared by calculating their weighted Jaccard similarity.
1601 : k-mer set 1; 1602: k-mer set 2; 1603: Formula for calculating the Jaccard coefficient
FIG. 17: Schematic representation of an exemplary implementation of a Chemical Unclonable Function using a sequence composition structure as described in example 5.
1701 : Template structure of the ssODN composition as synthesized and described in example 5 (121 nt) .
5'ATGCGAGTCAGATNGCACTCN N N N N N N N N ACACGACGCTCTTCCGATCTN NNNNNNNNNNNN N N N N N N N NGCTCAGGATACCAAGCTGTCCN N N N N N N N N NG ACATNGG ACGACTCAGCTA3' (SEQ ID NO: 32).
1702 : Handle I portion, 5'ATGCGAGTCAGATNGCACTC3' (SEQ ID NO: 33)
1703 : Input I, 9 nt, 5'NNNNNNNNN3'
1704: Sequencing adapter I, 5'ACACCACGCTCTTCCGATCT3' (SEQ ID NO: 34)
1705 : Output, 21 nt, 5'NNNNNNNNNNNNNNNNNNNNN3'
1706: Sequencing Adapter II, 21 nt, 5'GCTCAGGATACCAAGCTGTCC3' (SEQ ID NO: 35)
1707 : Input II, 10 nt, 5'NNNNNNNNNN3'
1708 : Handle II, 5'GACATNGGACGACTCAGCTA3' (SEQ ID NO: 36)
1709 : Amplification with handle I and II primers
1710 : dsODN composition (orDNA pool) according to example 5 (121 bp), 5'ATGCGAGTCAGATNGCACTCN N N N N N N N N ACACGACGCTCTTCCGATCTN NNNNNNNNNNNN N N N N N N N NGCTCAGGATACCAAGCTGTCCN N N N N N N N N NG ACATNGG ACGACTCAGCTA3' (SEQ ID NO: 41)
1711 : 5'AT3'; 1712 : Piel recognition site, 5'GCGAG3'; 1713: 5TCAGAT3'; 1714: Piel cut site; 1715 : 5'NGCACTC3'; 1716: 5'GACAT3'; 1717 : Piel cut site; 1718 : 5'NGGAC3'; 1719 : Piel recognition site, 5'GACTC3'; 1720 : STAS'; 1721 : 3'AT5'; 1722 : 3'CGCTC5'; 1723: 3'AGTCTAN5'; 1724: Piel cut site; 1725 : 3'CGTGAG5'; 1726: Input I reverse, 9nt, 3'NNNNNNNNN5';
1727 : Sequencing adapter I, reverse, 20 nt, 3TGTGCTGCGAGAAGGCTAGA5' (SEQ ID NO: 39)
1728 : Output reverse, 21 nt, 3'NNNNNNNNNNNNNNNNNNNNN5'
1729 : Sequencing Adapter II reverse, 21 nt, 3'CGAGTCCTATGGTTCGACAGG5' (SEQ ID NO: 40)
1730 : Input II reverse, 10 nt, 3'NNNNNNNNNN5'
1731 : 3'CTGTAN5'; 1732 : Piel cut site; 1733 : 3'CCTG5'; 1734: Piel recognition site, 3'CTGAG5'; 1735 : 3TCGAT5'; 1736: Restriction digestion with Piel
1737 : Digested dsODN composition with 5'N-overhangs, forward (SEQ ID NO: 42) and reverse strand (SEQ ID NO: 43)
1738: 5'N-overhang; 1739: 5'NGCACT3'; 1740: 5'GACAT3'; 1741 : 3'CGTGAG5'; 1742: 3'CTGTAN5'; 1743: 5'N-overhang
1744: Blunting by end fill-in with ddNTP
1745: Digested dsODN composition with blunted ends, 94 bp
(NGCACTCNNNNNNNNNACACGACGCTCTTCCGATCTNNNNNNNNNNNNNNNNNNNNNGCTCA GGATACCAAGCTGTCCNNNNNNNNNNGACATN, SEQ ID NO: 44). Composition that can be used for challenge-response pair generation.
1746: Filled-in ddNTP; 1747: Filled- in ddNTP
1748: Subsequent steps for challenge-response pair generation in analogy to the description in example 2 and FIG. 2.
FIG. 18: AGE photographs showing the dsODN bands after restriction digest. For size comparison a 50 bp ladder is shown.
1801 : Ladder lane; 1802: Sample lane; 1803: 150 bp marker; 1804: 100 bp marker; 1805: 50 bp marker; 1806: Sample band digested product, 92 bp double-stranded with 5'N-overhangs (forward strand with SEQ ID NO: 42, reverse strand with SEQ ID NO: 43)
1807: Sample band undigested product (negative control, Library 2 orDNA pool, SEQ ID NO: 41), 121 bp
FIG. 19: Gel image showing samples 1-3 of example 5 after ligation.
1901 : Sample 1, composition Piel digested and blunted with ddNTPs (SEQ ID NO: 44)
1902: Sample 2, composition Piel digested and blunted with dNTPs (SEQ ID NO: 78)
1903: Sample 3, undigested composition (Library 2 orDNA pool, SEQ ID NO: 41)
FIG. 20: Expected number of sequences (x-axis) in a pool that perfectly match to an input primer of a given length, measured by the number of variable base pairs within a primer pair (y-axis), for the pool sizes as indicated by the three curves.
While the present invention has been described with reference to a limited number of embodiments, variants, and the accompanying drawings, it will be understood by those skilled in the art that various changes may be made, and equivalents may be substituted without departing from the scope of the present invention. In particular, a feature (devicelike or method-like) recited in a given embodiment, variant or shown in a drawing may be combined with or replace another feature in another embodiment, variant or drawing, without departing from the scope of the present invention. Various combinations of the features described in respect of any of the above embodiments or variants may accordingly be contemplated, that remain within the scope of the appended claims. In addition, many
minor modifications may be made to adapt a particular situation or material to the teachings of the present invention without departing from its scope. Therefore, it is intended that the present invention is not limited to the particular embodiments disclosed, but that the present invention will include all embodiments falling within the scope of the appended claims. In addition, many other variants than explicitly touched above can be contemplated. For example, other verification procedures can be contemplated.
Claims
1. A composition including a plurality of double-stranded oligodeoxynucleotides, or dsODNs, wherein: the dsODNs of said plurality have a same length of between 47 and 300 bp, preferably between 80 and 150 bp, more preferably between 85 and 130 bp, and are structured according to a same template structure, the latter consisting of an orderly set of sequence portions having respective lengths that are constant across all the dsODNs of said plurality; the orderly set of sequence portions includes a first random segment, a first sequencing adapter, a second random segment, a second sequencing adapter, and a third random segment, which are consecutively arranged to form a sequence; each of the first random segment, the second random segment, and the third random segment of the dsODNs of said plurality consists of essentially random permutations of nucleotides; and the first sequencing adapter and the second sequencing adapter of the dsODNs of said plurality consist, independently of each other, of essentially a same sequence of nucleotides.
2. The composition according to claim 1, wherein, in each of the dsODNs of said plurality, each of the first sequencing adapter and the second sequencing adapter has, independently of each other, a length of between 13 and 30 bp, preferably between 18 and 22 bp.
3. The composition according to claim 1 or 2, wherein an average Levenshtein distance between first sequencing adapters of the dsODNs of said plurality is less than fa x Li, where Li is the length of the first sequencing adapters, an average Levenshtein distance between second sequencing adapters of the dsODNs of said plurality is less than fa x L2, where l_2 is the length of the second sequencing adapters, and fa is equal to 0.45, preferably equal to 0.30, and more preferably equal to 0.10.
4. The composition according to any one of claims 1 to 3, wherein, in each of the dsODNs of said plurality, the first random segment has a length /1 of between 5 and 25 bp, preferably between 6 and 10 bp,
the second random segment has a length I2 of between 11 and 200 bp, preferably between 15 and 50 bp, more preferably between 18 and 22 bp, and the third random segment has a length I3 of between 5 and 25 bp, preferably between 6 and 10 bp.
5. The composition according to any one of claims 1 to 4, wherein an average Levenshtein distance between first random segments of the dsODNs of said plurality is larger than ga x llr where /1 is the length of the first random segments, an average Levenshtein distance between second random segments of the dsODNs of said plurality is larger than ga x /2, where I2 is the length of the second random segments, an average Levenshtein distance between third random segments of the dsODNs of said plurality is larger than ga x /3, where I3 is the length of the third random segments, and ga is equal to 0.55, preferably equal to 0.70, and more preferably equal to 0.90.
6. The composition according to any one of claims 1 to 5, wherein the orderly set of sequence portions of each of the dsODNs of said plurality further comprises two outer handle sequences, these including a first handle sequence and a second handle sequence, whereby the first handle sequence, the first random segment, the first sequencing adapter, the second random segment, the second sequencing adapter, the third random segment, and the second handle sequence, are consecutively arranged in the sequence, each of the first handle sequence and the second handle sequence consist, independently of each other, of essentially a same sequence of nucleotides, and, preferably, each of the first handle sequence and the second handle sequence has, independently of each other, a length of between 1 and 30 bp, preferably between 3 and 30 bp, more preferably between 3 and 22 bp.
7. The composition according to claim 6, wherein each of the first handle sequence and the second handle sequence has, independently of each other, a length of between 10 and 30 bp, preferably between 18 and 22 bp.
8. The composition according to claim 6, wherein
each of the first handle sequence and the second handle sequence has, independently of each other, a length of between 1 and 9 bp, preferably between 4 and 7 bp.
9. The composition according to claim 6 or 8, wherein terminal base pairs of the dsODNs of said plurality are at least partly random, whereby no more than 55% of the dsODNs of said plurality have the same terminal base pairs.
10. The composition according to any one of claims 1 to 9, wherein terminal base pairs of the dsODNs are deoxynucleotide - dideoxynucleotide base pairs, preferably selected from a group consisting of ddA-dT base pairs, ddC-dG base pairs, ddG-dC base pairs, and ddT-dA base pairs.
11. The composition according to any one of claims 1 to 10, wherein the composition comprises a total amount of DNA in the range of 0.05 ng to 1000 ng, preferably 1 ng to 25 ng.
12. A method of producing a composition as defined in any one of claims 1 to 11, the method comprising the steps of: a) providing an initial composition including a plurality of single-stranded oligodeoxynucleotides, or ssODNs, wherein the ssODNs of said plurality have a same length of between 47 and 300 bp, preferably between 80 and 150 bp, more preferably between 85 and 130 bp, and are structured according to a same template structure, the latter consisting of an orderly set of sequence portions having respective lengths that are constant across all the ssODNs of said plurality; the orderly set of sequence portions includes a first handle sequence, a first random segment, a first sequencing adapter, a second random segment, a second sequencing adapter, a third random segment, and a second handle sequence which are consecutively arranged to form a sequence; each of the first random segment, the second random segment, and the third random segment of the ssODNs of said plurality consists of essentially random permutations of nucleotides; and the first handle sequence, the second handle sequence, the first sequencing adapter and the second sequencing adapter of the ssODNs of said plurality consist, independently of each other, of essentially a same sequence of nucleotides, and
b) subjecting the initial composition to a polymerase chain reaction to obtain a composition according to any one of claims 1 to 11.
13. The method according to claim 12, wherein the method further comprises a step of: c) at least partly cleaving the first handle sequence and the second handle sequence to obtain a composition according to any one of claims 1 to 11, wherein residual portions of the first handle sequence and the second handle sequence, if any, have a length that is less than 10 bp, preferably less than 8 bp.
14. The method according to claim 13, wherein the first handle sequence and the second handle sequence are only partly cleaved, whereby each of the first handle sequence and the second handle sequence has, independently of each other, a length of between 1 and 9 bp, preferably between 4 and 7 bp.
15. A set comprising an entity and the composition according to any one of claims 1 to 11, wherein the composition is associated with the entity, and, preferably, the composition is attached to an object, or a packaging thereof, corresponding to that entity.
16. Use of the composition according to any one of claims 1 to 11 as a physical fingerprint of an entity.
17. A method of verifying an entity of interest associated with the composition according to any one of claims 1 to 11, wherein the method comprises performing a verification procedure comprising: performing a polymerase chain reaction, or PCR, on the composition, the PCR involving one or more pairs of PCR primers, wherein each of the one or more pairs includes a forward PCR primer and a reverse PCR primer, which are respectively adapted to bind the first random segment and the third random segment, respectively, of at least some of the dsODNs of said plurality, so as to amplify sequences of the dsODNs based on the one or more pairs of PCR primers; and sequencing the amplified sequences to obtain a sequencing dataset containing sequencing reads, wherein the sequencing dataset is preferably obtained by filtering out sequencing reads that are inconsistent with said first sequencing adapter and said second sequencing adapter.
18. The method according to claim 17, wherein the verification procedure further comprises: performing a -mer extraction analysis of at least a subset of the sequencing reads, where 7 < k < 20, preferably 8 < k < 16, wherein the -mer analysis is preferably restricted to a subset of most-frequently occurring ones of the sequencing reads.
19. The method according to claim 18, wherein the method further comprises: assigning a given portion of the composition to an entity; subsequently receiving the given portion assigned, or a part thereof, for verification purposes, whereby said verification procedure is performed based on the received portion, or the received part thereof, to obtain a test result; and verifying said portion of the composition by comparing the test result with a reference result as obtained by performing the same verification procedure on a reference portion of the composition.
20. The method according to claim 19, wherein comparing the test result with the reference result comprises measuring a statistical similarity between two outcomes of the -mer extraction analysis, as obtained by performing said verification procedure on the given portion and the reference portion of the composition, and said two outcomes preferably consist of two sets of extracted features, which are more preferably extracted, each, in a form of a one-dimensional array of numbers.
21. The method according to claim 20, wherein measuring the statistical similarity between said two outcomes comprises weighting -sequences reads obtained according to the -mer analysis in accordance with respective frequencies of occurrence, and the similarity coefficient is preferably measured as a Jaccard coefficient.
22. The method according to any one of claims 17 to 21, wherein the method further comprises: mapping a set of input numbers to unique pairs of the PCR primers, and mapping the input numbers to output numbers generated based on the sequencing reads obtained.
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/EP2023/066962 WO2024260562A1 (en) | 2023-06-22 | 2023-06-22 | Operable random dna |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4731787A1 true EP4731787A1 (en) | 2026-04-29 |
Family
ID=87060214
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23735638.1A Pending EP4731787A1 (en) | 2023-06-22 | 2023-06-22 | Operable random dna |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4731787A1 (en) |
| WO (1) | WO2024260562A1 (en) |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9790538B2 (en) * | 2013-03-07 | 2017-10-17 | Apdn (B.V.I.) Inc. | Alkaline activation for immobilization of DNA taggants |
| GB2439960B (en) * | 2006-07-08 | 2011-11-16 | Redweb Security | Material for marking an article using DNA |
| JPWO2010029629A1 (en) * | 2008-09-11 | 2012-02-02 | 長浜バイオラボラトリー株式会社 | DNA-containing ink composition |
| US9963740B2 (en) * | 2013-03-07 | 2018-05-08 | APDN (B.V.I.), Inc. | Method and device for marking articles |
| US11473140B2 (en) * | 2013-11-26 | 2022-10-18 | Lc Sciences Lc | Highly selective omega primer amplification of nucleic acid sequences |
| US20190144958A1 (en) * | 2016-05-06 | 2019-05-16 | Provenance Biofabrics, Inc. | Cultured leather and products made therefrom |
| WO2018045109A1 (en) * | 2016-08-30 | 2018-03-08 | Metabiotech Corporation | Methods and compositions for phased sequencing |
| WO2019236787A1 (en) * | 2018-06-07 | 2019-12-12 | Videojet Technologies Inc. | Dna-tagged inks and systems and methods of use |
-
2023
- 2023-06-22 WO PCT/EP2023/066962 patent/WO2024260562A1/en not_active Ceased
- 2023-06-22 EP EP23735638.1A patent/EP4731787A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024260562A1 (en) | 2024-12-26 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Lopez et al. | DNA assembly for nanopore data storage readout | |
| Carøe et al. | Single‐tube library preparation for degraded DNA | |
| DK2245187T3 (en) | Methods for accurate sequence data and modified due to localization | |
| CN104272311B (en) | The data analysis of DNA sequence dna | |
| EP2718866B1 (en) | Providing nucleotide sequence data | |
| ES2403312T3 (en) | New strategies for genome sequencing | |
| US9334532B2 (en) | Complexity reduction method | |
| US20220389416A1 (en) | COMPOSITIONS AND METHODS FOR CONSTRUCTING STRAND SPECIFIC cDNA LIBRARIES | |
| JP2018521675A (en) | Target enrichment by single probe primer extension | |
| CA3187549A1 (en) | Compositions and methods for nucleic acid analysis | |
| WO2017037657A1 (en) | Method of identifying sequence variants using concatenation | |
| KR20240113772A (en) | Nucleic acid storage for blockchain and non-fungible tokens | |
| CN114555821A (en) | Detecting sequences uniquely associated with target regions of DNA | |
| CN109825552B (en) | Primer and method for enriching target region | |
| JP5926189B2 (en) | RNA analysis method | |
| CN112634984B (en) | Method, device and storage medium for simultaneously detecting DNA methylation and genome variation | |
| Park et al. | Selection of self-priming molecular replicators | |
| JP7104770B2 (en) | Methods for Amplifying and Determining Target Nucleotide Sequences | |
| EP4731787A1 (en) | Operable random dna | |
| AU2021202166A1 (en) | Composition for improving molecular barcoding efficiency and use thereof | |
| CA3200114C (en) | Rna probe for mutation profiling and use thereof | |
| EP3237635A1 (en) | Bubble primers | |
| Liu et al. | High‐Fidelity Data Retrieval from Synthetic DNA Pools via Machine Learning Model | |
| EP4494143A1 (en) | Combinatorial enumeration and search for nucleic acid-based data storage | |
| Kumar Sahu et al. | DNA Sequencing by Hybridization and Shotgun Sequencing Technique. |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20260113 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |