EP3589753A1 - Verfahren zur detektion von bekannten nukleotid-modifikationen in einer rna - Google Patents
Verfahren zur detektion von bekannten nukleotid-modifikationen in einer rnaInfo
- Publication number
- EP3589753A1 EP3589753A1 EP18710996.2A EP18710996A EP3589753A1 EP 3589753 A1 EP3589753 A1 EP 3589753A1 EP 18710996 A EP18710996 A EP 18710996A EP 3589753 A1 EP3589753 A1 EP 3589753A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- signature
- phase
- nucleotide
- different
- rna
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6869—Methods for sequencing
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B35/00—ICT specially adapted for in silico combinatorial libraries of nucleic acids, proteins or peptides
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B35/00—ICT specially adapted for in silico combinatorial libraries of nucleic acids, proteins or peptides
- G16B35/10—Design of libraries
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B35/00—ICT specially adapted for in silico combinatorial libraries of nucleic acids, proteins or peptides
- G16B35/20—Screening of libraries
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2521/00—Reaction characterised by the enzymatic activity
- C12Q2521/10—Nucleotidyl transfering
- C12Q2521/107—RNA dependent DNA polymerase,(i.e. reverse transcriptase)
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2525/00—Reactions involving modified oligonucleotides, nucleic acids, or nucleotides
- C12Q2525/10—Modifications characterised by
- C12Q2525/117—Modifications characterised by incorporating modified base
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Y—ENZYMES
- C12Y207/00—Transferases transferring phosphorus-containing groups (2.7)
- C12Y207/07—Nucleotidyltransferases (2.7.7)
- C12Y207/07049—RNA-directed DNA polymerase (2.7.7.49), i.e. telomerase or reverse-transcriptase
Definitions
- the present invention relates to a method of detection, i. Determination of number and position (locus) of a selected known nucleotide modification in one or more RNAs (including transcriptome).
- RNA ribonucleic acid (ribonucleic acid)
- mRNA messenger RNA (messenger RNA)
- tRNA transfer RNA
- rRNA ribosomal RNA
- ml A Nl-methyladenosine (m'A)
- mlG Nl-methylguanosine (m'G)
- dNTP deoxynucleotide triphosphate
- the transcriptome that is to say the entirety of the RNA transcripts (ie the RNA polymerase read or transcribed genes) of a genome of a cell or of a cell type or of an organism, in particular mRNAs, tRNAs and rRNAs, but also others encoding RNAs plays a crucial role in various aspects of gene expression, cell development and cell function. Errors in the transcriptome, for example due to modified nucleotides in an mRNA or tRNA or rRNA, can lead to diseases. The identification and characterization of various types of RNA base modifications in different types of RNA have become increasingly important in recent years. Interest in current research is growing, and this field continues to grow in importance.
- RNAs reverse transcriptases
- RTases reverse transcriptases
- RNA selected as template which are obtained in a reverse transcription (RT) for short with a specific RTase, are first amplified and then sequenced.
- RT reverse transcription
- the resulting sequencing data of the cDNAs are compared with the genomic reference sequence, and in the course of the so-called “mapping" the sequenced cDNAs / reads are assigned to the reference genome or reference transcriptome.
- RTase reverse transcriptase
- RT signature The type and number of different RT events form a characteristic event pattern, the so-called reverse transcription signature (in the following: RT signature) at each individual nucleotide position.
- the RT signature for the nucleotide positions of an RNA under study is in principle based on the characteristic features abort events and mismatch read-through events (ie read-through events with missinkorpor convinced or mismatched cDNA building blocks).
- mapping results obtained after reverse transcription, amplification, sequencing and mapping for this template RNA are examined and evaluated as to whether and, if so, at which nucleotide position which RT Events occur at what frequency and hence what the RT signatures look like for each nucleotide position.
- RT signature for the relevant template RNA can be closed to existing nucleotide modifications. If a particular characteristic RT signature could be determined for a specific nucleotide modification, as has been done in the prior art for mlA, the presence of the relevant nucleotide modification in the template can be determined by comparison with this known and modification-specific RT signature Be closed.
- the template RNA may be a particular RNA species, as well as a group of different RNA species.
- Amplification and sequencing of the cDNAs are usually carried out in the prior art with sequencing methods based on high throughput methods in the form of massive parallel sequencing, the so-called “Next-Generation Sequencing” (NGS), and in which the acquired sequence data are output in digital form ,
- NGS Next-Generation Sequencing
- next-generation sequencing is so-called "bridge amplification sequencing”.
- a different adapter DNA sequence is introduced at each end of the (double-stranded) DNA to be sequenced.
- the DNA is denatured, after dilution (single-stranded) hybridized to a support plate and amplified by bridge amplification.
- the carrier plate On the carrier plate, individual regions (clusters) of amplified DNA are formed which have the same sequence within a cluster.
- a sequencing-by-synthesis-related PCR reaction ie, a PCR reaction in which synthesis is sequenced
- modified nucleotides coupled to a reversible 3 'blocker and a fluorescent label each of the four nucleotides with a different colored fluorescent label coupled
- the built-in nucleotide per cycle in a cluster is detected.
- mapping is preferably carried out by means of computer-aided alignment methods known in the art, and the evaluation (analysis) of the mapping results with regard to the reverse transcription event pattern (the RT signature) is usually computer-assisted.
- the characteristic RT signature described in the prior art for mlA was determined using computer-aided, automatic and supervised machine learning-based classification techniques known and well known in the art.
- the determined RT signature for mlA (ie, at mlA sites) was used for the review and confirmation of suspect objects. Suspected positions of mlA in the Sequences of several human R As could be confirmed, and in trypanosoma brucei tRNA previously unknown m lA positions were determined by signature matching and sequence homology.
- RTase reverse transcriptase
- a solution to this problem is to provide a method for determining the number and position (locus) of a selected (predetermined), known Nucleotide modification in one or more RNAs (including transcriptome) of the so-called template RNA (s), comprising the following steps in the order named:
- a first phase (I) of the method the calibration phase, steps (1) to (5) with one or more different, known and with respect to nucleotide sequence and optionally present nucleotide modification (s) identified and annotated RNAs as template RNAs and in step (5) determined RT signatures of nucleotide positions with the known nucleotide modification and detected RT signatures of nucleotide positions of the same nucleoside without nucleotide modification are fed into the classification system, and the classification system during training - and self-test (classification) run implicitly (“learns") the (characteristic) profile of the RT signature (ie the characteristic quantitative expression of the RT signature features) at the nucleotide modification having nucleotide position, and (as a result) as a classification result those positions on the (each ) Detects and reports TempIate RNA (s) that have an RT signature that approximately or completely coincides with, ie, is similar to, or with, this (
- steps (1) to (5) are carried out with one or more unknown test RNA (s) to be examined as template RNA (s) , and steps (1) to (4) are performed under the same conditions as in phase (I), and RT signatures of (preferably all or nearly all) nucleotide positions of the test template RNA determined in step (5) (s) are fed into the classification system, and the classification system classifies the (and preferably each of) the entered RT signatures based on the (implicitly) profile implicitly learned in step (I) step (5) (ie, criterion) to what extent (ie to what degree or degree) they are similar or coincident to this profile, and where classification results have the meaning of "similar” or “approximately coincidental” or “consistent” (ie classification results, which correspond to the statement "similar” or “approximately coincident” or “consistent”) indicate the presence of the relevant nucleotide
- step (1) of phase (I) and phase (II) of the method the reverse transcription of the template RNAs in two or more reaction mixtures and run through with different RTases under the same reaction conditions and / or with the same RTase (s) is carried out under different reaction conditions per batch, with / from each batch a cDNA library is obtained,
- phase (I) and phase (II) of the method for evaluating the mapping results with respect to the RT signature the event (s) 'abort' and / or 'read-through' Mismatch 'and / or the additional event' read-through with sequence gap (s) (jump / jumps) 'determined and evaluated as RT signature feature (s),
- phase (I) and phase (II) of the method data sets of RT signatures of (all or almost all or at least nucleotide positions with the base type of the respective nucleotide modification determined) the cDNA libraries obtained in step (1) which interact with the different RTases under the same reaction conditions and / or with the same RTase (s) were received into the classification system.
- RNAs to be investigated including transcriptome, the template RNA (s), consists of two phases ( I) and (II):
- RNAs identified and annotated with regard to their nucleotide sequence and optionally the selected selected nucleotide modifications as template RNA (s)
- RNAs preferred are synthetic RNAs or RNAs isolated from natural sources according to database Information from eg MODOMICS according to Machnicka et al., 2013
- RTases including or modified by mutations specifically for this purpose
- RTases including or modified by mutations specifically for this purpose
- the same reaction conditions and / or with (the) same RTase (s) under different reaction conditions per batch and creating cDNA libraries, one each per reaction run, wherein each created cDNA library, the reverse transcription products (cDNAs) of the RTase used in the relevant reaction run from the one or the ver used template RNA (s).
- step 2 For each cDNA library (obtained in step 1), amplification of the cDNAs and sequencing of the amplified cDNAs are carried out by a high-throughput sequencing method, whereby the obtained sequence data, i. the sequence information of the individual cDNAs (synonym: reads), are output in digital form. Preference is given here to "sequencing with bridge amplification", e.g. the Illumina sequencing method.
- adapter trimming removal of the adapter sequences
- Preferred here is the use of a computer-based method for sequence alignment and sequence analysis, for example the Bowtie 2 software.
- step (4) feeding the (digitized) data (data sets) of the RT-signatures of all or almost all or at least those RT-signatures determined at the nucleotide positions with the base type of the respective nucleotide position into a computer-based, automatic one , machine-based and supervised learning-based classification system (synonyms: classification method, classifier), eg into a Random Forest classifier, and train this (learning) classification system on the particular (characteristic, typical) profile of the RT signatures obtained in step (4) (ie on the characteristic quantitative expression of the RT signature features) for the resp at the nucleotide position (s) with the relevant nucleotide modification (ie having the relevant nucleotide modification), such that it determines as a classification result those positions on the template RNA (s) and indicating having an RT signature that approximately or completely matches that profile, ie which are similar or in agreement with it, and thus indicate the presence of the relevant nucleotide modification at these positions.
- phase II the analysis or investigation phase with at least one test RNA, the following steps are carried out in the order named:
- phase I step (1) reverse transcription of the test RNA (s) to be tested as template RNA (s) under the same conditions as in phase I step (1), i. with the RTase (s) and reaction conditions used in phase I step (1) and preparation of cDNA libraries (one per batch) comprising the reverse transcription products of the particular RTase (s) used for this test template RNA (s) included.
- step (2) Amplifying the cDNAs recovered in step (1) and sequencing the amplified cDNAs using the method used in step I step (2), wherein the obtained sequence data (reads) are output in digital form.
- each entered RT signature is classified according to how similar it is to this profile. (That is, each entered RT signature is classified in terms of the criterion of how much, or to what degree or degree, that it resembles or conforms to that profile.) Classification results corresponding to the statement “similar” or “approximately consistent” or “consistent "indicate the presence of the subject nucleotide modification in the test template RNA (s) at the nucleotide position with this RT signature.
- the core result of the classification is the indication of the determined positions on the test template RNA (s) that have an RT signature that approximately or fully matches this (particular) profile, i. which is similar or in agreement with it, and thus indicates the presence of the subject nucleotide modification at these positions.
- a numerical score is given on a one-dimensional numeric rating scale as a measure of the quality of the match.
- the method according to the invention is based on the surprising findings:
- the RT signature at a nucleotide modification site depends not only on the type of nucleotide modification, but also on the RTase type (the RTase species). Because of their very specific and characteristic behavior at the site of a nucleotide modification, an RTase type-specific RT signature is obtained at this nucleotide position. (ii) By combining at least two RTases of different types in reverse transcription, surprisingly large improvements in predictive performance are obtained by means of classifiers.
- the RT signature is characterized not only by the special features (special RT events) abort and mismatch read-through (broken down into overall rate and single rates of the various mismatches), but also by the feature "read-through events with sequence gap (s) (Synonyms: jump / jumps; Jump (s)) "short” Jump-Read-Through ", ie by events in which the RTase skips the site of nucleotide modification.
- This feature "jump-read-through” can also be further broken down into: total jump rate, rate of direct single jumps, rate of delayed single jumps and rate of double jumps.
- the reverse transcription event of RTases at a nucleotide modification site can not only consist of transcription termination or read-through with mismatch / mismatch, but also in jumps of the RTase in question across the position of the modified nucleotide, resulting in characteristic gaps in the sequence reading.
- Such jumps were found especially in RTases with a high coverage / coverage rate, ie with a strong read-through capability.
- Single and double jumps can be distinguished, and the single jumps can be either direct or delayed single jumps, that is, the (skipped) gap is either directly at the mlA site or at the location of its 5 'adjacent neighbor when - 1 position. Double jumps lead to appear as gaps at the two positions mlA- and-1.
- step (4) of phase (I) and phase (II) of the method for the evaluation of the mapping results with respect to the RT signature all three events' abort 'and' read-through with mismatch 1 and 'read-through with sequence gap (s) (jump (s), jump / jumps)' qualitatively and quantitatively determined and evaluated as forming characteristics. This enhances the conciseness and uniqueness of each RT signature.
- the analogous reaction mixtures and runs are carried out with at least two reverse transcriptases ("RTases") whose RT signatures at or for the nucleotide modification site in question the weighting (synonyms: importance, importance) of their RT signature features have a different pattern.
- the patterns differ in the weighting of at least one of the features such that this feature is pronounced in the RT signature of one RTase and weak in the RT signature of the other RTase or at least significantly less pronounced.
- Particularly preferred different patterns are those having a mutually opposite pattern in at least two RT signature features (M1 and M2, e.g., the arrest rate and the Mismach rate). That is, of the characteristics involved, e.g. Ml and M2 are strong in one RTase (A) feature Ml, and feature M2 is weak, while in the other RTase (B) the ratios are reversed, namely, Ml is weak and M2 is strong.
- M1 and M2 e.g., the arrest rate and the Mismach rate
- RTases used in step (1) of the calibration phase (phase I) and the application or examination phase (phase II) may also be well according to the invention which have been generated for this purpose by mutations.
- step (1) of phase I and phase II it is possible, in particular, for (a) different concentrations of dNTPs, and / or (b) different divalent cations, in particular Mg 2+ and Mn 2+, and / or (c) different concentrations of divalent cations and / or (d) different pH values and / or (e) different temperatures and / or (f) different concentrations of polyethylene glycol (PEG).
- dNTPs different concentrations of dNTPs
- divalent cations in particular Mg 2+ and Mn 2+
- PEG polyethylene glycol
- RNA modification mlA For detection of other RNA modifications, especially those for which sequencing data analysis provides a typical profile of the RT signature, e.g. Guanosine derivatives Nl -Methyl guanosine (mlG) and N2, N2-dimethylguanosine (m2.2G), it is also suitable and intended.
- An embodiment of the method according to the invention is therefore in particular that the nucleotide modification is a nucleoside methylation, in particular a Nl methylation of adenosine or guanosine.
- step (2) of phase I and phase II of the method according to the invention has a sequencing with bridge amplification, in particular an illumina sequencing method, proved to be well suited.
- mapping i. the assignment of the sequenced cDNAs / reads to the reference genome or reference transcriptome by computer-aided AI ignment- method in step (3) of Phase I and Phase II of the method has in practice a computer-based method for sequence AI ignment and sequence analysis, such as eg the Bowtie 2 software proved to be well suited.
- a Random Forest classifier As a computer-based automatic machine learning-based classification system in step (5) of phase I and phase II of the method of the invention, a Random Forest classifier has been found to work well.
- the sequence data obtained in step (2) for the implementation of steps (3) to (5) are fed into a bioinformatics pipeline which controls the combination of steps (3) to (5) .
- a bioinformatics pipeline ie the software program that completes steps (3) to (3) (5) combined in the prescribed order or coupled to each other, can be created according to the invention with the programming language Python (Version v2.7.6).
- the known RNAs used in the calibration phase (phase I) step (1) according to the invention are preferably synthetic RNAs of known sequence including known positions of the relevant (selected) nucleotide modification or natural RNAs isolated on the basis of database information, the sequence of which, including the positions of the relevant (selected) nucleotide modification according to the relevant database entries, is well understood.
- step (5) of phase II and optionally also of phase I for each classification result a numerical score is given on a one-dimensional numerical rating scale as a measure of the quality of the match.
- the present invention also provides a kit for carrying out the method according to the invention, which comprises at least two RTases whose RT signatures at the relevant nucleotide modification point have an effect on the weighting of the RT signature features (termination rate, overall mismatch rate, Single mismatch rates of the respective mismatched nucleotides, total hopping rate, rate of direct hops, rate of delayed hops, double hopping rate) have a different, preferably opposite, pattern in at least one of the RT signature features.
- Opposite pattern in at least one RT signature feature means here that, for example, the feature Ml is pronounced in RTase A and only weakly in RTase B.
- the kit comprises at least two different premixed reaction mixtures (synonyms: reaction mixtures, buffer mixtures) which are preferably in the concentration of dNTPs and / or divalent cations and / or polyethylene glycol (PEG) and / or in the nature of the divalent compounds present Cations (especially Mg2 + and Mn2 +) and / or in the pH, differ.
- reaction mixtures buffer mixtures
- PEG polyethylene glycol
- the template RNA (s) is required for carrying out the method according to the invention with such a kit.
- the method according to the invention is a powerful tool for the detection of modified nucleotides in RNA on the basis of the RT signature at the odhuisstelle, ie based on the analysis of the modification-specific behavior of RTase during the reverse transcription of the RNA to cDNA. It allows for accurate localization of RNA modifications in a single nucleotide resolution, and thus, for example, a much more accurate identification and prediction of mlA sites than conventional methods, and it may in principle be just as useful for analyzing other modifications such as m lG or m2.2G are used.
- step (1) of phase (I) and phase (II) of the method Carrying out the reverse transcription of the template RNA (s) in step (1) of phase (I) and phase (II) of the method in two or more parallel (analogous) reaction batches and reaction runs with RTases different from one another and / or with under different reaction conditions per batch, and the comparison of the thus obtained and usually not quite identical RT signatures for the same nucleotide modification site makes it possible to clarify the characteristics of the RT signature for the relevant nucleotide modification site in more detail and further specify.
- the more succinctly and more specifically the characteristic features for the RT signature can be given at a particular nucleotide modification site the more accurate can be found for the RT signature of the reverse transcription of a test template RNA (eg from a patient sample), whether or not it represents an embodiment of this known RT signature, ie whether or not the nucleotide modification in question is present in the test template RNA (s).
- a test template RNA eg from a patient sample
- the method according to the invention is a universal method for the transcriptome-wide detection of RNA modifications involving very specific properties of the reverse transcription or the RTase, which allows detection of individual, modified nucleotides within the sequence solely on the basis of their characteristic RT signature.
- An accumulation of sequence regions obtained by immunoprecipitation, which contain (presumably) the nucleotide modification, can be completely dispensed with here.
- the method according to the invention can be used in the field of clinical diagnostics by analytical service providers or medical-diagnostic laboratories and can be used to customize the personalized medicine with regard to patient-specific Further develop diagnostics.
- analytical service providers or medical-diagnostic laboratories can be used to customize the personalized medicine with regard to patient-specific Further develop diagnostics.
- many new insights can be expected in the field in the coming years, which makes the precise determination of modified sequence positions or nucleotide positions all the more important.
- the method of the invention makes it possible to make serious statements about their effect and function and a routine application in an economical manner.
- Analytical service providers or clinical diagnostic laboratories can analyze patient-derived RNA samples and generate a report of classified modification candidates, providing additional information for the patient's diagnostic workup.
- Figure 1 The inventive principle of generation (generation) and analysis (evaluation) of RNA sequencing data for the detection of m lA residues
- Figure 2 A) The RT-signature of a MIA point obtained by a conventional method using a single RTase ( "single RT-signature"), here the RTase 5 (SuperScript ® III), that is, with use of the RT-signature an RT approach using only the RTase 5 (SuperScript ® III.) According to Table 1.
- RT signature of an IA site obtained by the method according to the invention in which the information of the RT signature is combined from two different RT approaches which differ in the RTase used.
- the ienexen RTases were (i) 12 RTase (SuperScript ® IV) and (ii) RTase 4 (GoScript TM) according to Table 1 below.
- FIG. 3 mlA signatures of 13 RTases at 26 mlA sites in the cytosolic tRNA from yeast.
- the size of the pie charts represents the overall hop rate, i. the
- Single jump direct 1 nucleotide was omitted and skipped at m lA itself.
- Single-bound delay 1 nucleotide was omitted at the 5 'adjacent position of mlA and skipped.
- Double jump 2 nucleotides were skipped and skipped, at the mlA site and at the -1 position.
- the percentages of the abort rate refer to the reads comprising the 3 'adjacent position of m lA (+1).
- the percentages of mismatch rate and hopping rate refer to the reads comprising the m l A position.
- Figure 4 Random Forest Execution and weighting of RT signature features for 13 different RTases.
- Classification power is represented as the Area Under Curve (AUC) of the Receiver Operating Characteristic (ROC).
- AUC Area Under Curve
- ROC Receiver Operating Characteristic
- Weighting average loss of classification accuracy as the values of the respective features between the training instances (m lA instances and non m lA instances) permute, i. E. be replaced.
- Figure 5 Random Forest implementation for determining the prediction performance under
- FIG. 6 Box plot (box whisker plot, box graphic) for the m lA prediction performance of the Random Forest classification, which is combined with the information or data of the RT signature features of an RT signature (from one of the 13 different RTases ), two RT signatures (from one of the 156 RTase pairs), and three RT signatures (from one of the 1716 RTase triplets).
- AUC Area Under Curve
- the boxes indicate the area in which the middle 50% of each data population is located.
- the whiskers mark the percentile values 5% and 95%, i. the values that form the boundary to the lower 5% and the upper 5% of the data, respectively.
- Figure 7 An example of a profile file.
- Mismatch type 1, type 2 and type 3 are synonymous with the three concrete mismatches with the three bases that naturally occur next to the reference base (and its modification) in the genome; ie in the case of a modification of A, the mismatch types G, T and C.
- Example 1 Recovery of Template RNAs for the Calibration Phase (Phase I)
- RNA (s) used in the calibration phase and identified with respect to their nucleotide sequence and optionally the selected selected nucleotide modifications were either synthetic RNAs (commercially available eg from IBA, Göttingen, Germany) or RNAs derived from natural sources For example, yeast RNAs whose sequence information from databases such as MODOMICS are known.
- yeast rRNA and yeast tRNA were recovered by known and well known methods, e.g. as in Tserovski et al. (2016).
- RNA 0.5 ⁇ g was used per sample / batch for a reverse transcription (reaction).
- the protocol corresponds in principle to that in Tserovski et al. (2016).
- EDTA ethylenediaminetetraacetic acid
- the template RNA (s) (about 0.5 ⁇ g per sample / batch) was dephosphorylated at both endpoints.
- the dephosphorylation mixture (total 10 ⁇ ) consisted of 100 mM Tris-HCl, pH 7.4, 20 mM MgCl 2 , 0.1 mg / ml BSA, 100 mM 2-mercaptoethanol and 0.5 U FastAP Alkaline Phosphatase (Thermo Scientific, # EF0651) at 37 ° C for 30 min. Before the addition of the enzyme, the RNA was denatured at 90 ° C for 30 sec and then cooled on ice (hereinafter, this treatment is called "heat denaturation").
- the RNA was heat denatured again for 30 sec and then the described dephosphorylation step was performed a second time.
- an adapter was ligated (ligated) to the 3 'end of the dephosphorylated RNA.
- the ligation (attachment) of the preadeylated 3 'RNA adapter (whose 5' end was blocked by a C6 body) to the 3 'end of the RNA was as described in Tserovski et al.
- ligases in this case T4 RNA Ligase 2 truncated, New England Biolabs, # M0242L, and T4 RNA Ligase, Thermo Scientific, # EL0021
- T4 RNA Ligase 2 truncated New England Biolabs, # M0242L
- T4 RNA Ligase Thermo Scientific, # EL0021
- the ligation reaction was carried out at 4 ° C overnight. Subsequently, the enzymes were inactivated at 75 ° C for 15 min.
- RNA adapter Prior to the step of reverse transcription, the excess of RNA adapter was removed using the enzymes deadenylases and exonucleases (here 5'-deadenylase, New England Biolabs, # M0331S and Lambda Exonuclease, Thermo Scientific, # EN0561).
- deadenylases and exonucleases here 5'-deadenylase, New England Biolabs, # M0331S and Lambda Exonuclease, Thermo Scientific, # EN0561.
- an amount of 20 U of 5'-deadenylase e.g., from New England Biolabs, Frankfurt, Germany
- RNA was precipitated, here with the addition of initially 1 ⁇ glycogen (Thermo Scientific, Dreieich, Germany, # R0561) and NH4AC ammonium acetate (final concentration: 0.5 M) to a total volume of 50.0 ⁇ and subsequent Add 150 ⁇ ethanol per sample.
- the composition of the reverse transcription mixture was as described in Tserovski et al. (2016).
- FS First beach
- the degradation of the RNA was then carried out by addition of NaOH (final concentration: 0.15 M), heating to 55 ° C. for 25 minutes and subsequent cooling on ice for 2 minutes.
- the reaction was carried out by neutralization with an equal amount of acetic acid ( Final concentration: 0.15 M) stopped.
- the cDNA pellet obtained in (F) in the reaction mixture was extracted from 1x TdT buffer, 1.25mM rCTP and 1 U / ⁇ TdT taken and resuspended.
- the mixture was incubated at 37 ° C for 30 min. This is followed by a heat treatment at 70 ° C for 10 min. To inactivate the enzyme. - total volume: 10.0 ⁇ . Subsequently, the ligation reaction was carried out with the mixture obtained, here for example with the aid of T4 DNA ligase (Thermo Scientific, # EL0013).
- the extraction of the cDNA ligation products from the mixture obtained was carried out by ethanol precipitation with the addition of initially 1 ⁇ glycogen (Thermo Scientific, # R0561) and NH4AC (final concentration: 0.5 M) to this mixture (total volume: 50.0 ⁇ ) and final addition of 150 ⁇ ethanol.
- Polyacrylamide gel electrophoresis was performed to remove excess DNA adapter. For this purpose, the last pellet obtained with the ligation products in ⁇ H 2 0 were added and resuspended. This resuspended ligation product mixture was applied to a denaturing 10% polyacrylamide gel. After electrophoresis, the areas of the size range between 40 nt and 150 nt were excised from the gel and eluted overnight with 300 ⁇ M 0.5 M NH4AC.
- the recovery of the cDNA ligation products from the recovered eluate was carried out by ethanol precipitation with the addition of initially 1 ⁇ glycogen (Thermo Scientific, # R0561) (total volume: 301.0 ⁇ ) and final addition of 750 ⁇ ethanol.
- the cDNAs obtained from (G) were analyzed by means of the polymerase chain reaction (PCR) using a Taq polymerase, here for example the Taq polymerase from Rapidozym (# Gen-003-1000), and corresponding barcoded P5 and P7 primers, here eg each with 8 nt barcodes, amplified.
- PCR polymerase chain reaction
- the last pellet obtained in (G) with the ligation products of size 40 nt to 150 nt was taken up in 20 ⁇ l PCR reaction mixture and resuspended.
- the PCR reaction mixture consisted per 20 ⁇ of lx Taq polymerase buffer, 3 mM MgCl 2> 5 ⁇ P5 primer, 5 ⁇ P7 primer, 0.5 mM dNTP mix and 0.25 U / ⁇ Taq polymerase.
- the recovered resuspension with the adapter ligated cDNAs contained therein, the P5 and P7 primers, the Taq polymerase and the dNTPs was subjected to 12 PCR cycles.
- the PCR started with a denaturation step (of DNA double strands in single strands) at 95 ° C for 5 min. Subsequently, 12 cycles consisting of denaturation at 95 ° C for 1 min, annealing (hybridization) at 65 ° C for 1 min and elongation at 72 ° C for 1 min.
- the PCR was terminated with a final elongation step at 72 ° C for 5 min.
- the recovery of the PCR products, i. of the amplified cDNAs was carried out by means of ethanol precipitation with the addition of initially 1 ⁇ glycogen (Thermo Scientific, # R0561) and NH4Ac (final concentration: 0.5 M) to this mixture (total volume: 50.0 ⁇ ) and final addition of 150 ⁇ ethanol ,
- PCR products (amplified cDNAs) were size separated by 10% denaturing polyacrylamide gel electrophoresis (PAGE).
- the last pellet obtained with the amplified cDNAs in 10 ⁇ 1 ⁇ 2 ⁇ was added and resuspended. This resuspension was applied to a denaturing 10% polyacrylamide gel. After electrophoresis, the gel sized areas were cut out between 150 nt (the size of adapter dimers) and 300 nt (the maximum size of PCR amplification products) and eluted overnight with 300 ⁇ M 0.5 M NR, Ac.
- the recovery of the amplified cDNAs from the eluate was carried out by ethanol precipitation with the addition of initially 1 ⁇ glycogen (Thermo Scientific, # R0561) (total volume: 301.0 ⁇ ) and finally adding 750 ⁇ ethanol.
- the recovered pellet containing the amplified cDNAs of size 150-300 nt was taken up in 10 ⁇ H 2 O and resuspended.
- the cDNAs contained in this suspension were ready for sequencing, in particular also for high-throughput sequencing with NGS methods, eg "sequencing with bridge amplification".
- the aliquots were diluted (5-500 pg / ⁇ ) and loaded on an Agilent High Sensitivity DNA chip.
- the thus loaded chip was introduced into the analyzer.
- the sample components DNA molecules
- the sample components were electrophoretically separated, detected and translated into gel-like images (bands) and / or electropherograms (peaks).
- the data was generated in digital form and automatically analyzed in real time. If the quality of the aliquot examined was satisfactory, the corresponding sample was used further.
- Example 2 cDNA library For this purpose, the examined in terms of quality and quantity and found suitable samples of (possibly several parallel) prepared according to Example 2 cDNA library (s) were combined, denatured with 2 N NaOH and diluted (10 pM), and on the support plate, the so-called "Flow Cell", applied.
- the determined sequence information per cDNA molecule, the "Reads”, were generated and output in digital form and were ready for injection and further processing in a bioinformatics pipeline.
- the obtained sequencing data were checked for quality and adapter contamination. For this they were (here and preferably) examined via a bioinformatics or high-throughput sequencing pipeline with the well-known in the art software program FastQC.
- the FastQC program created a quality control (QC) report of the detected issues that had arisen either in the sequencer or in the source library material.
- QC quality control
- FastQC could be run in one of two modes. It could either run as a stand-alone interactive application for instant analysis of small numbers of FastQ files, or it could run in a non-interactive mode suitable for systematically processing a large number of FastQ files. In this non-interactive mode, it was well integrated into a larger analysis pipeline.
- the examination was carried out by means of FastQC within the MiSeq RTA software.
- the barcode sequences from the barcoding PCR step were first identified (no fault tolerance - 0 mismatch) and then the reads (sequencing data) into individual FastQ files (one FastQ file per sample or per original cDNA). Library).
- Adapter sequences in particular the sequences of the adapters P5 and P7 from the PCR reaction (see Example 3 (H)) and also random 10 nt sequences of the 3'-RNA adapter at the 3'-end of the RNA (see Example 2 (cf. C)) and variable number of 5'-G RNA nucleotides from the CTP cDNA tailing step (see Example 3 (G)) were removed.
- Trimming was computer-based (here and preferably) with bioinformatics software for adapter trimming commonly used in the art, in this example the Cutadapt vi .8.1 software.
- the mapping ie the assignment to the reference genome, was performed using the Bowtie 2 software.
- RT signature diagnostics i. the identification and quantitative measurement of the reverse transcription event pattern for the template RNA (s) tested was done at each individual nucleotide position of the RNA of interest using software programs well known and well known in the art, e.g. the SAMtools software (version 1.2).
- the SAM files from the mapping step were first converted into BAM files. Then the steps followed: (i) sorting and indexing the BAM files, (ii) converting the BAM files to the Pileup format, and (iii) converting the Pileup format into a custom tab-delimited text file (so-called " Profile File ").
- Pileup format Every line of the Pileup format accurately reflects the coverage of a reference position.
- Bases that differ from the reference base (in the template RNA) appear in Pileup format in the form of the usual letters A, G, T or C (if the respective Read in "sense" direction, ie as it is has been aligned) or a, g, t, or c (if the respective read has been aligned in anti-sense direction, ie as its reverse complement).
- a jump over the corresponding position is displayed as an asterisk.
- CSA context-sensitive termination rate
- CSA is defined as the ratio of position-specific RT termination (arrest) a t at a position i to the RT abort observed in the local environment, ie in the adjacent sequences
- the pileup Format the 5 digits before and 5 digits after the respective nucleotide position i (ie, the neighboring sequences five bases upstream (+ 5 bp) and five bases downstream (- 5 bp)) and the arrest rate at position i by the median of the Divided arrest rates of all eleven positions in this window.
- the window size may be increased or decreased depending on knowledge of the nature of the RNA present, to improve, if appropriate, the predictive power in the cross-validation.
- the data thus obtained were taken from the pileup format e.g. (as here and preferably) transformed into the profile format and stored there and displayed as needed.
- FIG. 7 shows an example of such a profile file.
- reduced-data profile files can be created, e.g. only with the data of positions corresponding to reference base A of the considered modification mlA.
- Example 6 Computer-based and machine-learning-based supervised prediction of mlA sites
- a computer-based, machine-based classification method e.g. B. (here and preferably) in the known and used in the art, based on decision trees Random Forest classification system (R Version v3.3.1) fed.
- the classifier was given at least the attributes: termination rate a, total mismatch (mismatch) rate m, m / a ratio, relative mismatch composition (fraction content of G, T and C), and the context-sensitive abort rate CSA entered, and preferably also the jump rate.
- RT mappings from known m lA sites (identified from known template RNAs) mlA sites) and to RT signatures of known unmodified (or not identifiable modified) A sites (from known template RNAs without mA sites).
- profile files with a reduced data record created in accordance with example 5 were preferably used, ie profile files which contain only the RT signature data of the positions indicated for the relevant reference base (in this example for A) (in this example, mA or non-MLA positions).
- mA or non-MLA positions Preferably, as described in Hauenschild et al., 2015, mlA-like non-ml A sites were included in the training to prepare the classifier for recognition of difficult cases in unknown template RNAs as well.
- the classifiers Based on the RT signatures of the known positive ml A sites ("learned") and adapted (corrected and optimized) of the classifiers (here, for example, and preferably the Random Forest classifier, R version v3.3.1) created the special and typical (characteristic) profile for the mlA site (see Fig. 2A). In other words, the classifier implicitly learned the typical m lA-RT signature profile during the training and (self-) testing / verification runs.
- the classifier implicitly learned the typical m lA-RT signature profile during the training and (self-) testing / verification runs.
- the classification result consisted of a listing of all tested positions ("instances") on the template RNA (s) with a rating per position ("instance”) with respect to the decision or question as to whether the respective RT Signature with the learned typical mlA RT signature profile (see Fig. 2A) approximately or completely agree, ie that was similar to him or agreed with him.
- the evaluation was carried out by specifying a numerical value between 0 and 1, where the value "0" corresponds to a clear “no” and the value "1" to a clear “yes".
- Intermediate values represent a corresponding probability for a "yes” or a "no” (for example, the value 0.99 stands for a relatively very certain "yes” and the value 0.4 stands for a weak "no”).
- the mean predictive power (sensitivity, specificity) was calculated on the basis of training and testing / verification using cross-validation.
- the unknown template RNA (s) to be examined was prepared according to Examples 2 and 3.
- the retrieved sequence information in the form of reads was trimmed according to Example 4 and mapped to a reference genome and examined for RT signatures according to Example 5, ie the RT signature features coverage, arrest rate ), Total mismatch rate, single mismatch rates (of the corresponding mismatched nucleotides) rate of direct single jumps (single jump Rate Direct), Single Jump Rate Delayed, Double Jump Rate were checked for presence at each nucleotide position and quantified if necessary.
- nucleoside in question For the purpose of minimizing the expenditure, only those nucleotide positions in the profile format were stored according to example 5 and fed into the classifier having the reference base in question (the nucleoside in question).
- a shortened runtime of the prediction procedure could be achieved by using in the profile file apparently unnatural lines (ie nucleotide positions that had been correctly transcribed by the RTase and thus where none of the characteristic RT signature features are present) by means of a simple Filters (requesting user-selected minimum values for signature characteristics) were removed in a controlled manner.
- the classifier eg the Random Forest model
- the classifier provided an estimate between "yes” and “no” for each nucleotide position in the transcriptome being examined and thus a listing of those positions on the unknown template RNA (s), which had an RT signature which corresponded to the particular and typical (characteristic) profile (implicitly learned by the classifier) for the relevant nucleotide modification site, here in the example for the m lA site, approximately or completely.
- the so-called “score” a numerical score on a one-dimensional rating scale, was also indicated here (for example and preferably).
- the creation of the score is a (possible) component of the Random Forest classifier.
- RTases reverse transcriptases
- RNA used for each of the RTases was the well-analyzed (annealed) total tRNA of Saccharomyces cerevisiae (available, for example, from Roche Diagnostics: Ref 10109525001 / lot 13407921), which contains sufficiently many known m lA sites.
- the reference sequences used were a set of 43 tRNAs, compiled from the databases MODOMICS (Machnicka et al., 2013) and SRocl (Jühling et al., 2009). Among these 43 tRNAs were 26 tRNAs carrying one ml A in their sequence.
- RTase showed large variations in their termination and mismatch rates, ie the termination and mismatch rates of the individual RTases were very different in comparison with one another (see FIG. 3).
- RTase showed 10 (Monster Script TM) and RTase 4 (GoScript TM) demolition high and low mismatch (mismatch) -Raten during 12 RTase (SuperScript ® IV) compared to the opposite behavior exhibited.
- a hitherto unknown phenomenon was surprisingly found to be of varying severity in the individual RTases: in the sequencing data sets of some RTases, characteristic sequence gaps were recognizable at mlA sites, which indicate jumps in the depreciation of the RNA in cDNA. In other words, in addition to mismatch / mismatching and termination, the RTases in question also showed characteristic gaps in the sequence reading resulting from jumps in the RTase in question over the mlA position as hitherto unknown phenomenon (see FIGS. 1 and Fig. 2).
- the single jumps may be direct or delayed single jumps, that is, the (skipped) gap lies either directly at the mlA site or at the location of its 5 'adjacent neighbor, known as the -1 position. Double jumps lead to appear as gaps at the two positions ml A- and -1.
- the scatter plot in Figure 3 shows for the 13 different RTases examined the great variety of RT signatures at m lA sites, taking into account this newly discovered third core feature "jumps".
- the abort rate, mismatch rate, and total hopping rate values shown represent the average of the three individual values of each technical triplicate; the error bars show the standard deviations of the truncation and misincorporation rates of these triplicates.
- Example 7 To evaluate the predictive power of the RT signature of an individual RTase, in each case the characteristic features of the RT signature with which weighting were determined in accordance with Example 6 (A) for the RT signatures of 13 RTases obtained in Example (7) characterizes or co-imprints the relevant RTase-specific RT signature.
- the six RT signature features were examined (here, for example and preferably): abort rate, overall hopping rate and mismatch rate and, with regard to the mismatch events, also the relative contents of the mismatch components G, T and C.
- the RT signatures of 26 mlA cases were mated with an equal number of non-m lA signatures randomly drawn from the surrounding sequence pool. These pairs were mixed and divided into three groups of equal frequency (so-called "folds").
- Each signature data point contained the RT signature characteristics abort rate, relative mismatch rate, relative mismatch components (G, T and C), and the overall hopping rate.
- a cross validation a random forest model (as described in Liaw and Wiener, 2002) was trained (trained) on two of these groups ("folds") and tested on (the third). Shuffling (i.e., group composition blending) and cross-validation were repeated 100 times to account for statistical variance.
- each RTase provided the results shown in Figure 4: for each of the 13 RTases (# 1 to # 13), for each of the six RT signature features, the averaged ranking of their performance for a m 1 A prediction and thus their ""Classificationpower" as specified in the "Area Under Curve (AUC)" of the "Receiver Operating Characteristic (ROC)".
- the data of the RT signature triplicates were averaged; the black vertical lines show standard deviations over 3 sequence runs). For each RTase, 100 repetitions of a 3-fold cross-validation were performed in each classification run.
- RT-signatures of some RTases lead to better predictive power results than those of others.
- the predictive power results differed by several percentage points.
- Targeted selection of the presumptive most appropriate RTase (s) for a planned sequencing of a test template RNA for a given nucleotide modification can improve workflow and significantly reduce residual errors.
- a comparison of the weighting values obtained for the RT signature characteristics of different RTases (at mlA sites) with machine learning models designed for them indicates an individual weighting pattern of the features in the RTS signature of each RTase that contributes to the decision making.
- the weighting (synonyms: significance, importance) of an RT signature feature was determined by permuting the (all) values obtained for this feature with all 13 RTases examined (ie, commutation of values, also of Negative instances, namely nucleotide positions with potentially weak expression of this feature, with corresponding values of positive instances, namely nucleotide positions with mostly stronger expression of this feature), and measurement of the corresponding decrease in the classification accuracy. This decrease in classification accuracy is tends to be higher, the more important (defining) this feature is for the RT signature.
- RTase 3 (ProtoScript ® II) for the RT-signature of RTase has the permutation of the pronounced drop-out rate large impact, while this feature for RTase 12 (superscript ® IV), where it is only slightly pronounced only of secondary importance. Although they are different in the patterns of feature weighting, RTase are 3 (ProtoScript® II) and RTase 12 (superscript ® IV) to the highest AUC ranks, ie their RT-signatures have the strongest classification capability and allow the best machine learning services ,
- RTase pairs a significantly optimized prediction performance for the obtained RT signatures (the RTase pairs) can consequently be achieved.
- two (or correspondingly more) RT signatures are obtained, the combined application of which in the supervised machine learning experiment according to Example 8 has a significantly improved classification performance, ie Performance for m 1 A prediction causes (results). Residual errors are significantly reduced.
- Example 9 The studies described in Example 9 were carried out analogously with RTase triplets, i. with triple combinations of different RTases, performed.
- results obtained were expected to improve that the predictive performance (detection performance) for a m lA site can be improved by combining the RT signature data from three different RTases.
- Numerous RTase triplets delivered AUC values of 1.000, demonstrating ideal classification (classification) performance and m / L predictive performance.
- a detailed comparison between RT signature pairs and RT signature triplets shows however, the RT signatures of some RTase pairs already provide a quasi-best m lA predictive power, ie, their AUC values are in a range that overlaps with that of the RT signature triplets.
- Fig. 6 is a comparison of the m lA predictive powers in a Random Forest classification for the three alternative training and application modes (states) of the classification method, namely training (according to Example 6 A) and application according to Example 6 B) with the Information or data of the RT signature characteristics of (i) an RT signature (of the 13 different RTases), (ii) two RT signatures (of one of the 156 RTase pairs) and (iii) three RT signatures ( from one of the 1716 RTase triplets) using a box plot.
- the method for detecting nucleotide modifications other than m lA, eg of mlG or m2.2G, is carried out as in Examples 1 to 9 or 1 to 10, with the modification that as template RNAs in Step 1 of Phase I, and thus as training RNA set for the classifier such RNA sequences are used, which are known and proven to contain the desired Nuleotid modification, eg mlG or m2.2G.
- Example 12 Kit for carrying out the process according to the invention
- the kit comprises (a) at least two reverse transcriptases RTase X and RTase Y, whose RT signatures at the respective nucleotide modification site have respect to the weighting of the RT signature characteristics (termination rate, total mismatch rate, single mismatch rates of the respective mismatched nucleotides, total hopping rate, direct single hopping rate, delayed single hopping rate, double hopping rate) have a different pattern in at least one of the RT signature features, and or (b) at least two different premixed reaction batches (synonyms: reaction mixtures; Buffer mixtures) A and B, which embody different (ie divergent) reaction conditions, for example by different concentrations of dNTPs, and / or different divalent cations, especially Mg2 + and Mn2 +, and / or different concentrations of divalent cations and / or different pHs, and / or different concentrations of polyethylene glycol (PEG).
- step (1) in phase (I) and in phase (II) of the process, it is only necessary to mix the RTase (s) with the reaction mixture or the reaction mixtures and the relevant template RNA (s) and to incubate.
- MODOMICS a database of RNA modification pathways-2013 update. Nucleic Acids Res., 41, D262-D267.
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Engineering & Computer Science (AREA)
- Chemical & Material Sciences (AREA)
- Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Organic Chemistry (AREA)
- Biotechnology (AREA)
- General Health & Medical Sciences (AREA)
- Biophysics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Biochemistry (AREA)
- Molecular Biology (AREA)
- Library & Information Science (AREA)
- Wood Science & Technology (AREA)
- Zoology (AREA)
- Evolutionary Biology (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Medical Informatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Analytical Chemistry (AREA)
- Genetics & Genomics (AREA)
- General Engineering & Computer Science (AREA)
- Microbiology (AREA)
- Immunology (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| DE102017002092.2A DE102017002092B4 (de) | 2017-03-04 | 2017-03-04 | Verfahren zur Detektion von bekannten Nukleotid-Modifikationen in einer RNA |
| PCT/DE2018/000044 WO2018161981A1 (de) | 2017-03-04 | 2018-02-21 | Verfahren zur detektion von bekannten nukleotid-modifikationen in einer rna |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP3589753A1 true EP3589753A1 (de) | 2020-01-08 |
Family
ID=61628084
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP18710996.2A Withdrawn EP3589753A1 (de) | 2017-03-04 | 2018-02-21 | Verfahren zur detektion von bekannten nukleotid-modifikationen in einer rna |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20190390269A1 (de) |
| EP (1) | EP3589753A1 (de) |
| DE (1) | DE102017002092B4 (de) |
| WO (1) | WO2018161981A1 (de) |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110379464B (zh) * | 2019-07-29 | 2023-05-12 | 桂林电子科技大学 | 一种细菌中dna转录终止子的预测方法 |
| CN111951889B (zh) * | 2020-08-18 | 2023-12-22 | 安徽农业大学 | 一种rna序列中m5c位点的识别预测方法及系统 |
| CN113257354B (zh) * | 2021-05-12 | 2022-03-11 | 广州万德基因医学科技有限公司 | 基于高通量实验数据挖掘进行关键rna功能挖掘的方法 |
| WO2024073730A2 (en) * | 2022-09-29 | 2024-04-04 | The University Of Chicago | Methods and systems for rna sequencing and analysis |
| CN116926039A (zh) * | 2023-09-19 | 2023-10-24 | 魔因生物科技(北京)有限公司 | 反转录酶HIV p66突变体及其应用 |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2013123481A1 (en) * | 2012-02-16 | 2013-08-22 | Cornell University | Methods and kit for characterizing the modified base status of a transcriptome |
-
2017
- 2017-03-04 DE DE102017002092.2A patent/DE102017002092B4/de active Active
-
2018
- 2018-02-21 US US16/483,896 patent/US20190390269A1/en not_active Abandoned
- 2018-02-21 WO PCT/DE2018/000044 patent/WO2018161981A1/de not_active Ceased
- 2018-02-21 EP EP18710996.2A patent/EP3589753A1/de not_active Withdrawn
Non-Patent Citations (3)
| Title |
|---|
| MOTORIN YURI ET AL: "Identification of modified residues in RNAs by reverse transcription-based methods", RNA MODIFICATION ELSEVIER ACADEMIC PRESS INC, 525 B STREET, SUITE 1900, SAN DIEGO, CA 92101-4495 USA SERIES : METHODS IN ENZYMOLOGY (0076-6879(PRINT)), 2007, pages 21 - 53, XP009523854 * |
| SCHRAGA SCHWARTZ ET AL: "Next-generation sequencing technologies for detection of modified nucleotides in RNAs", RNA BIOLOGY, vol. 14, no. 9, 5 December 2016 (2016-12-05), pages 1124 - 1137, XP055745529, ISSN: 1547-6286, DOI: 10.1080/15476286.2016.1251543 * |
| See also references of WO2018161981A1 * |
Also Published As
| Publication number | Publication date |
|---|---|
| US20190390269A1 (en) | 2019-12-26 |
| WO2018161981A1 (de) | 2018-09-13 |
| DE102017002092A1 (de) | 2018-09-06 |
| DE102017002092B4 (de) | 2018-11-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| DE102017002092B4 (de) | Verfahren zur Detektion von bekannten Nukleotid-Modifikationen in einer RNA | |
| DE69225333T2 (de) | Verfahren für den Nachweis von Mikroorganismen unter verwendung von direkter undwillkürlicher DNA Amplifikation. | |
| Ding et al. | In vivo genome-wide profiling of RNA secondary structure reveals novel regulatory features | |
| DE69230873T3 (de) | Selektive Restriktionsfragmentenamplifikation: generelles Verfahren für DNS-Fingerprinting | |
| DE3855064T2 (de) | Selektive Amplifikation von Oligonukleotiden-Zielsequenzen | |
| DE69713599T2 (de) | Verfahren zur bestimmung von nukleinsäure-sequenzen und diagnostische anwendungen derselben | |
| EP0438512B1 (de) | Verfahren zur analyse von längenpolymorphismen in dna-bereichen | |
| WO2001007648A1 (de) | Verfahren zum speziesspezifischen nachweis von organismen | |
| DE102008025656A1 (de) | Verfahren zur quantitativen Analyse von Nikleinsäuren, Marker dafür und deren Verwendung | |
| DE60030811T2 (de) | Verfahren zur Ampifizierung von RNA | |
| Holland et al. | MPS analysis of the mtDNA hypervariable regions on the MiSeq with improved enrichment | |
| CN109559780A (zh) | 一种高通量测序的rna数据处理方法 | |
| DE60133321T2 (de) | Methoden zur Detektion des mecA Gens beim methicillin-resistenten Staphylococcus Aureus | |
| WO2018019610A1 (de) | Dna-sonden für eine in-situ hybridisierung an chromosomen | |
| DE60311263T2 (de) | Verfahren zur bestimmung der kopienzahl einer nukleotidsequenz | |
| DE102023105888A1 (de) | Verfahren zur Identifizierung eines Kandidaten, nämlich eines Genlocus und/oder einer Sequenzvariante, der für mindestens ein (phänotypisches) Merkmal indikativ ist | |
| DE69834422T2 (de) | Klonierungsverfahren durch multiple verdauung | |
| EP0698122A1 (de) | Mittel zur komplexen diagnostik der genexpression und verfahren zur anwendung für die medizinische diagnostik und die genisolierung | |
| DE102007010311A1 (de) | Organismusspezifisches hybridisierbares Nucleinsäuremolekül | |
| WO2007068305A1 (de) | Verfahren zur bestimmung des genotyps aus einer biologischen probe enthaltend nukleinsäuren unterschiedlicher individuen | |
| DE60109002T2 (de) | Methode zum Nachweis von transkribierten genomischen DNA-Sequenzen | |
| Liao et al. | Nanopore sequencing and haplotyping of mitochondrial DNA hypervariable regions and its application on mixed stain | |
| Lee et al. | Unlocking the Potential of Low Quality Total RNA-seq Data: A Stepwise Mapping Approach for Improved Quantitative Analyses | |
| Calhoun | INVESTIGATION INTO THE GENETIC BASIS OF CAPSAICIN PRODUCTION IN PEPPERS USING NEXT GENERATION RNA SEQUENCING AND SYNTHETIC BIOLOGY APPROACHES | |
| Calhoun | Investigation in to the Genetic Basis of Capsaicin Production in Peppers Using Next Generation RNA Sequencing and Synthetic Biology Approaches |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20190831 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| AX | Request for extension of the european patent |
Extension state: BA ME |
|
| RIN1 | Information on inventor provided before grant (corrected) |
Inventor name: KEMMER, THOMAS Inventor name: TSEROVSKI, LYUDMIL Inventor name: HAUENSCHILD, RALF Inventor name: HELM, MARK Inventor name: HILDEBRANDT, ANDREAS Inventor name: WERNER, STEPHAN Inventor name: LECLAIRE, JENNIFER |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| 17Q | First examination report despatched |
Effective date: 20201105 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20210518 |