GB2579110A - Method for determining a consensus sequence of a target polymer - Google Patents

Method for determining a consensus sequence of a target polymer Download PDF

Info

Publication number
GB2579110A
GB2579110A GB1907165.3A GB201907165A GB2579110A GB 2579110 A GB2579110 A GB 2579110A GB 201907165 A GB201907165 A GB 201907165A GB 2579110 A GB2579110 A GB 2579110A
Authority
GB
United Kingdom
Prior art keywords
sequence
graph
consensus
reads
polymer
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
GB1907165.3A
Other versions
GB201907165D0 (en
Inventor
James Wright Christopher
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Oxford Nanopore Technologies PLC
Original Assignee
Oxford Nanopore Technologies PLC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Oxford Nanopore Technologies PLC filed Critical Oxford Nanopore Technologies PLC
Priority to GB1907165.3A priority Critical patent/GB2579110A/en
Publication of GB201907165D0 publication Critical patent/GB201907165D0/en
Publication of GB2579110A publication Critical patent/GB2579110A/en
Withdrawn legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • G16B30/10Sequence alignment; Homology search
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • G06N3/0455Auto-encoder networks; Encoder-decoder networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/042Knowledge-based neural networks; Logical representations of neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/044Recurrent networks, e.g. Hopfield networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/088Non-supervised learning, e.g. competitive learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/02Knowledge representation; Symbolic representation
    • G06N5/022Knowledge engineering; Knowledge acquisition
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • G16B30/20Sequence assembly
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B40/00ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
    • G16B40/20Supervised data analysis
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6869Methods for sequencing

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Biophysics (AREA)
  • Software Systems (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Computing Systems (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Molecular Biology (AREA)
  • Computational Linguistics (AREA)
  • Chemical & Material Sciences (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Biomedical Technology (AREA)
  • Medical Informatics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Biotechnology (AREA)
  • Analytical Chemistry (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Organic Chemistry (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Wood Science & Technology (AREA)
  • Zoology (AREA)
  • Databases & Information Systems (AREA)
  • Epidemiology (AREA)
  • Public Health (AREA)
  • Immunology (AREA)
  • Microbiology (AREA)
  • Bioethics (AREA)
  • Biochemistry (AREA)
  • Genetics & Genomics (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

The invention relates to a method of determining a consensus sequence of a target polymer comprising polymer units, the target polymer can be a strand of DNA. The method includes obtaining a first sequence taken from a signal relating to the target polymer and obtaining a second sequence taken from a second signal relating to the target polymer. A multiple sequence alignment is constructed from the plurality of sequences. A machine learning technique is used to determine the consensus sequence of the target polymer exclusively from the multiple sequence alignment and the differences between the sequences derivable therefrom. Optionally, the signals can be obtained during translocation of a polymer through a nanopore and measurements are dependent upon the k-mers in the nanopore, a k-mer being k polymer units of the polymer.

Description

METHOD FOR DETERMINING
A CONSENSUS SEQUENCE OF A TARGET POLYMER
Background
The present invention relates to methods of determining a consensus sequence of a target polymer, for example but without limitation a consensus sequence of polynucleotides, obtained from signals obtained during translocation of the polymer with respect to a nanopore. Aspects of the invention relate to the preparation of a polymer for use in such methods, and the determination of a consensus sequence.
A type of measurement system for estimating a target sequence of polymer units in a polymer uses a nanopore, and the polymer is translocated with respect to the nanopore. Some property of the system depends on the polymer units in the nanopore, and measurements of that property are taken. This type of measurement system using a nanopore has been shown to be highly effective, particularly in the field of sequencing a polynucl eoti de such as DNA or RNA and has been the subject of much recent development.
Such nanopore measurement systems can provide long continuous reads of polynucleotides ranging from hundreds to millions of nucleotides. The data gathered in this way comprise measurements, such as measurements of ion current, where each translocation of the sequence with respect to the sensitive part of the nanopore can result in a change in the measured property.
The signal measured during movement of a polynucl eoti de with respect to a nanopore, such as for example translocation of the polymer through a nanopore, has been shown to be dependent upon plural nucleotides and is complex. Analytical techniques of estimating a polymer sequence from measurements taken during interaction of the polynucleotide with a nanopore include the use of machine learning techniques such as a recurrent neural network, such a technique is disclosed in PCT/GB2018/051208, hereby incorporated by reference in its entirety.
Examining multiple sequence reads of an analyte to determine a consensus sequence has at least two types of variation, namely 0 errors with respect to the true identity of the analyte, and (ii) differences due to genuine population differences within a sample of the analyte.
When examining multiple sequence reads of an analyte, whether the term sequence read is understood to mean a raw measurements from a sequencing device, or derived nucleotide sequences, produced by a basecaller independently for each raw measurement read, then at least two types of variation can be present: errors with respect to the true identity of the analyte, and differences due to genuine population differences within a sample of the analyte. The invention seeks to inhibit the former variation while enhance the latter.
Consensus sequencing is understood to mean deriving a single or small set of sequences from the multiple reads which are determined to be the correct identity, or identities, of the analyte from which reads have been taken. Naively the consensus sequence is an "averaged" sequence.
Classical techniques for constructing a consensus sequence are centred on first constructing a multiple sequence alignment (NBA) [1, 2]. A multiple sequence alignment encodes, in either a graph or matrix form, the simultaneous alignment of all sequence reads to each other. There are known methods (naive or otherwise) for extracting a consensus sequence from an A/ISA graph -commonly dynamic programming is used [3]. The improved methods herein eschew the dynamic programming on IVISA graphs and utilise machine learning techniques, such as the use of neural networks, to extract a consensus sequence.
In addition to derivation of a single or small set of sequences representative of the analyte, classical techniques also enable variant calling. In addition to establishing a sequence a comparison to the reference sequence can be made. There are various traditional techniques to perform variant calling, in outline [4] these stages are: filtering the set of reads to remove sources of error or bias from the sequencing device; aligning the reads to a reference genome; using an algorithm, either based on a statistical model or some heuristics, to predict the likelihood of variation at each locus, based on the quality scores and allele counts of the aligned reads at that locus; filtering the predicted results, often based on metrics relevant to the application; and using a single nucleotide polymorphism (SNP) annotation to predict the functional effect of each variation.
There is some similarity between consensus sequencing and variant calling, and one way to achieve consensus sequencing is to perform variant calling, in some manner, and mutate the original reference sequence with the variation found.
Research has led to the application neural networks (NNs) to the task of variant calling. The two most prominent examples are the methods of DeepVariant [5] and Clairvoyant [6], and each relies on independent alignment of basecalled reads to a reference sequence and subsequent analysis of the collection of alignments, the so-called pileup, by a neural network to produce a list of variants. In both cases this analysis is performed using a convolution neural network (CNN) to examine the pileup in the form of a set of matrices. In some sense these matrices can be seen as a set of images, indeed the description of DeepVariant draws heavily on images classification literature. There are minimal differences between these two methods. It is at least one aim of the invention to inhibit errors with respect to the true identity of the analyte, whilst retaining an understanding of the variations due to genuine population differences within a sample of the analyte
Summary of Invention
According to a first aspect, the invention resides in a method of determining a consensus sequence of a target polymer comprising polymer units, the method comprising: obtaining a first sequence taken from a first signal relating to the target polymer, and obtaining a second sequence taken from a second signal relating to the target polymer; constructing a multiple sequence alignment from the plurality of sequences; determining the differences between the sequences in the multiple sequence alignment; and determining using a machine learning technique the consensus sequence of the target polymer exclusively from the multiple sequence alignment and the differences between the sequences derivable therefrom. The method can obtain three or more sequences from respective signals relating to the target polymer.
Each signal can be a series of measurements taken from a nanopore. Nanopore sensing is known and, in particular, known for measuring fluctuations in voltage or current as a polymer passes via or through a nanopore. Such measurements produce a signal than can be analysed to produce a sequence, which are often referred to as "reads". It is to be noted that multiple signals can be recorded for the same polymer or element thereof The differences between two or more sequences in the multiple sequence alignment can be used to derive feature vectors, such as weighting or distribution values.
For the avoidance of doubt, the machine learning technique determines a consensus from sequence information alone, during the determination of the consensus. The invention does not use a reference sequence or draft sequence obtained from an external source. The invention determines a consensus without a reference sequence or draft sequence.
The multiple sequence alignment can be configured by at least one of an external reference sequence, used to order the sequences; comparing short sequences and producing a longer continuous sequence therefrom; and selecting one of the plurality of sequences as a reference.
The machine learning technique can determine a single continuous consensus sequence. The method can produce a single continuous consensus sequence for different regions of the target polymer.
The machine learning technique can be applied directly to the multiple sequence alignment The method can further include performing analysis on the multiple sequence alignment to identify feature vectors, and analysing said feature vectors using the machine learning technique.
Information, such as feature vectors, can be extracted from the multiple sequence alignment using an algorithm and entered in to a machine learning technique to produce a consensus sequence. The information can be obtained using a first machine learning technique and the consensus sequence can be determined using a second machine learning technique. The first and second machine learning techniques can be combined.
The analysis on the multiple sequence alignment can be performed by a preliminary machine learning technique, such as a neural network.
The machine learning technique can include a recurrent neural network for determining the consensus.
The preliminary machine learning technique, such as a neural network, can produce a graph having nodes, and edges connecting nodes, wherein the nodes represent at least one polymer unit or absence thereof and edges represent a transition from one polymer unit to another. The graph can be a partial order alignment graph.
The edges can be directed and the graph can be a directed graph without cycles. The edges can be undirected and the graph can be a directed acyclic graph.
A plurality of multiple sequence alignments can be constructed. Each of the plurality can be a group of sequences, aligned, representing different sub-regions of the target polymer.
The machine learning technique can retain knowledge of a probability of each edge between adjacent nodes. The machine learning technique can assimilate the probability of each edge within a window of the multiple sequence alignment such that 3 or more transitions are taken into account when determining the consensus.
Obtaining at least one of the first and second sequence can include taking a series of measurements of a signal relating to the target polymer, wherein a measurement of the first signal can be dependent upon a plurality of polymer units, and wherein the polymer units of the target polymer modulate the signal.
The signal can be obtained during translocation of a polymer through a nanopore and measurements can be dependent on the k-mers in the nanopore, a k-mer being k polymer units of the polymer, where k is a positive integer, wherein the measurements can be made (i) in respect of individual k-mers, or (ii) in respect of analysis performed on the input signal itself According to a further aspect, the invention resides a computer program executable in a computer apparatus or analysis apparatus having instructions which, when executed, cause the computer to implement a method claimed herein. According to a further aspect, the invention resides in a data processing apparatus, device or system comprising a processor adapted to perform a method claimed herein. According to a further aspect, the invention resides in a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out a method claimed herein.
Overall, the advantage of the improved methods of this invention is that the sequence is determined independently of any reference or draft sequence normally required to provide a 'yardstick' i.e. without a reference sequence that functions as a 'yardstick'. These reference sequences were found to impose a bias on the results. Analysing reads in absence of the reference removes the bias because the reads are being aligned and analysed with respect to each other, for example in a sequential manner (rather than being vertically aligned to a yardstick'). This technique has been shown to improve the outcome from Q27 to Q30.
Figures The invention is discussed below, by way of example only, with reference to the following figures in which: Figure 1 is a flow diagram of the steps taken in a known method for consensus formation using a plurality of reads and relying upon a reference sequence; Figure 2 is a flow diagram of the steps taken in an improved method of forming a consensus, wherein the consensus is formed independently of a reference sequence; Figure 3 contains a plurality of graphs showing a comparison of the performance of the improved method versus a known consensus tool; Figure 4 is an excerpt taken from the improved consensus method showing an alignment of reads and the contains a plurality of graphs showing a comparison of the performance of the improved method versus a known consensus tool; Figure 5 is a flow diagram of the steps taken in another improved method of forming a consensus, wherein the consensus is formed independently of a reference sequence; and Figure 6 is a section of a partial order alignment graph of the improved method.
Description
There is no known method that can form a consensus sequence using a neural network without direct aid of a reference sequence or perform variant calling without first comparing reads independently to a reference sequence and subsequent analysis of the pileup. Known use of neural networks to analyse multiple reads has relied on the simplification of the task at hand afforded by the use of a reference sequence. The improved method herein moves away from such conventionality.
The invention builds upon known methods and analysis of the collection of alignments can benefit from at least one of an external reference sequence, possibly from a public database, relating to the analyte under analysis that is used restrictively for the purposes of forming a multiple sequence alignment and not for the determination of the consensus; an assembler, which is a process where short sequencing reads are compared, and a longer continuous sequence is produced; and a selection of one of the sequencing reads.
The improved method draws on the use of neural networks to average basecalled sequencing reads from known methods, with the principle differences being: it produces a single continuous consensus sequence; it performs transformations on the pileup to extract features which are provided to a neural network, not simply the pileup; and it predominantly uses a recurrent neural network (RNN) not an CNN.
The improved method performs the tasks of consensus sequencing and variant calling without comparison of sequence reads to a reference sequence using construction of a multiple sequence alignment and subsequent analysis by a neural network.
The notion of creating variants without referring reads to a reference naturally sounds curious, however it is possible: one can derive variants by first constructing a consensus sequence (or sequences, or variation graph [9]) and compare these to the reference to produce the final desired output.
By analysing the sequences independently of a draft sequence or reference sequence i.e. without such a reference sequence, "reference bias" is removed. The comparison, i.e. the sequence alignment, of a read to the reference is by definition asking the question: "how is this read similar and dissimilar to the reference." The resultant pileup of a method using a reference has a skewed view of how individual reads are similar and dissimilar to each other. This bias is not desirable, particularly in the case of consensus sequencing. By contrast an MSA does not contain this reference bias and is more naturally comparing reads to each in and of themselves.
Further, the improved method supports the determination of a consensus sequence where it is not useful or desirable to first construct an assembly of the reads, for example when the reads fully constitute the analyte under study. It is more natural in this instance to construct and examine an MSA, Moreover, the improved method avoids problems associated with hand crafted and optimised heuristic algorithms [3] when extracting multiple consensus sequences from an MSA The improved method uses a machine learning technique such as a neural network (NN), to perform the extraction of consensus sequences.
The improved method is used to extract a consensus from an MSA graph and can be independent of the graph structure and locality. An NN approach can model time-dependent features of the graph, it can extract information from the large-scale structure of the graph.
Core method The method of constructing an MSA graph from multiple sequence reads is generally known, though we note [2, 1] there are a variety of methods to perform this step and there can be benefits to choosing a particular method contingent on subsequent analysis in the improved method. For concreteness and in the improved method' s implementation, partial order alignment (POA) [10, 11] is used. In order to construct a consensus sequence using an NN to process an MSA graph, three alternatives are proposed: using the graph as-is, perhaps represented as an adjacency matrix or otherwise, as input to an INN; (ii) perform algorithmic feature extraction based on expert knowledge from the graph and use these features as input to an MN, and (iii) extract features from the graph as-is using a preliminary neural network -the outputs of which are processed and given to a second network, while overcoming human bias and outperforming curated algorithms.
The third alternative can be regarded a subclass of the first if we regard the two networks as a single "black-box" process. It is to be noted that while the improved method is explained herein as deriving a consensus sequence from basecalled reads, it can be performed using raw sequencing measurements. For example, one can create an MSA of a nanopore ionic current measurements based around dynamic time warping with barycentre averaging [17].
Method -2nd alternative The first alternative is considered self-explanatory in light of the teaching herein, thus the improved method is explained using the second alternative, wherein the method performs algorithmic feature extraction from the graph and uses these features as input to an NN.
This second form removes reference-based issues discussed previously and will be shown to represent an improvement over the classical dynamic programming algorithms used on MSA graphs.
A known method is implemented using a "medaka" computer program as outlined in Figure 1, wherein a reference sequence is combined with a plurality of reads and aligned. The read alignment allows the 'pileup' to be read before feature vectors are identified and provided with labels that are scored before a consensus is determined. The program performs consensus by first aligning reads to the reference sequence independently to create a pileup resulting in a representation such as: * 1 2 3 4 5 6E7 9 10 11 12 13 14 15 16 17 18 19 Reference A A GC A CG T - A C A -:A C A Read! A GC A CG TIAIA C A A A Read 2 AA GC A CC- A A A A Read 3 AA T CA C A C T A A A Where part of a sequence differs from the reference the "basecall" is enclosed in a solid-lined box, as shown. In the common use case of creating a consensus sequence, the reference here is ordinarily a sequence derived from a prior assembly of the reads. ;The second step is to derive "feature vectors" for every column of the above pileup: in the implementation this is to count the occurrence of the five symbols, the four types of base and deletion "-" characters. Note the asymmetry between the two types of "-"character: ones indicating a deletion in a read with respect to the reference, and a second type indicating a deletion in a read with respect to other reads. The series of feature vectors constitutes the input to the RNN. The output of the RNN is a score for each of the five symbols (four bases and deletion) and each pileup column. A consensus sequence can be derived by choosing the highest scoring symbol at each column. In order to call secondary consensus sequences, and variants, lower scoring symbols can be extracted. ;To train the NN the truth sequence is also aligned to the reference sequence such that for every column of the pileup a truth symbol can be obtained. Training occurs using standard methods optimizing the categorical cross entropy of symbol scores. ;Figure 2 shows an MSA-NN approach of the improved method, which creates a similar pileup albeit in the absence of a reference sequence. This difference is significant in that the pileup is not informed by the local structure (the base and short "word" content) of the reference sequence. For the above reads we obtain: Read3 AA'TCACGT In similarity to the known method one 11 12 13 14 15 16 1. 17 1. 18 19 ;AC TAG A C TA C ;AC TA T AC TIAC ;A C T A A C TA C. example of the MSA-NN implementation 1 2 3 4 5 6:78 Readl A -GCAC Read2 AAGCACG TA creates a feature vector by counting bases down the columns of the pileup; and the feature vector is passed to the NN. Again, training is affected by assigning a truth label to every column of the A/ISA pileup and optimization of the categorical cross entropy. ;In the example presented above the structure of the pileup is the same in both cases. In such simple cases we naturally do not expect one method to be superior to the other. Examples where the structure of the pileup can vary arise when multiple differences between reads are coincident, that is to say close to each other in the multiple alignment. ;The improved method uses a partial order alignment (POA) graph having, by way of example, three types of score: 1. match score: a reward for aligning bases in sequences/graphs that are identical; 2. mismatch score: a penalty for aligning non-identical bases; and 3. gap penalty: a penalty for introducing gap symbols. It is the interplay between these factors, particularly with coincident variants, that can lead to different pileups either between the two approaches, or with a single approach when the relative values of the scores are changed. ;Method -rd alternative results The improved method -2nd alternative -is an improved method that is an adaptation of consensus tool called "medakaT [18] issued by Oxford nanopore Technologies. This tool can create a consensus sequence from nanopore sequencing data, which is performed using neural networks applied from a pileup of individual sequencing reads against a draft assembly. It outperforms graph-based methods operating on basecalled data and can be competitive with state-of-the-art signal-based methods, whilst being much faster. ;In the improved method the POA graph and MSA were constmcted using the spoa library [12, 13]. The genomes of seven organisms were considered, five of these were used for training the NN while two were held back for testing the veracity of the trained network. Since these genomes are longer than the typical read lengths we perform assembly of the reads using the 'cam' assembler [14]. ;The assembly sequence (the reference) is used only to scaffold the MSA construction for computational performance reasons. It is to be noted that this use of the assembly is not to be confused with how it is used in the previous medaka method. To be clear, information is not passed from the reference scaffold sequence to the MSA graph or the process of determining the consensus sequence. Reads are simply aligned to the assembly in order to find clusters of reads overlapping relatively long segments of the assembly sequence. This is the only information obtained from the reference, from which POA graphs are built. Separate MSA graphs are built for each long segment of the reference. In practice, segments of length 1500 bases were used in this example. ;As is typical in the assessment of consensus algorithms we examine the errors in the consensus sequence as a function of the "depth" of reads: the number of reads spanning a genomic loci (on average). Error is expressed on a log scale through use of Qscores: Perrot) Q = -10 log] 0(1 -P 1 wherein P CITOr is the proportion of errors (incorrect bases, additional bases, missing bases the consensus sequence. ;The process naturally has some variation with the precise set of sequencing reads used and the structural accuracy of the assembly sequence. To aid comparison the consensus is broken into 100 kilobase chunks whose accuracy is determined independently and a median Qscorc is calculated for each full consensus. A second median is performed across multiple independent consensus sequences calculated from distinct sets of reads. ;Three methods for calculating the consensus sequence were used: A. "racon": 'cam' assembly followed by application of the racon program [12], which uses a traditional dynamic programming approach to extract a consensus sequence from the POA graph. Here the program has been applied iteratively to its own output until no further improvement in the accuracy of the result is obtained. ;B. "moussaka": the racon approach of method A, followed by application of the new MSA-NN approach. ;C. "moussaka canu": 'came assembly followed by application of the new MSA-NN without intermediate use of racon. ;Figure 3 compares these methods and shows graphs showing the performance of each method during training and during test, except the known "racon" too which was not trained in this evaluation. As can be readily seen the MSA-NN provides routes to higher Qscores, thus indicating higher consensus accuracies. Figure 3 illustrates via method C that the MSA-NN approach is able to extract a more accurate consensus from a POA graph than static dynamic programming methods. ;Figure 4 is an illustration of an isolated example of an improvement to a region of a consensus sequence derived from the "racon" and "moussaka canu" approaches. The top section ("reads.bam") graphically represents the read sequences. Grey blocks (shaded horizontal strips) indicate regions where no errors are present in the reads. Purple marks (shaped like an I-beam cross-section) indicate incorrect additional bases are present in the reads, while thick black horizontal lines indicate bases which are missing. Bases which are incorrect (substituted) are shown by a letter. ;It was observed that down the central portion of the image many reads contain a deleted base and a substituted T (thymine) base. The second section of the figure ("racon.bam") illustrates the result of dynamic programming to extract a consensus sequence from the POA graph, whilst the third section shows the consensus from the MSA-NN approach. The final bottom panel of the figure records the correct (truth) sequence of the analyte studied. It can be seen that the MSA-NN approach makes a single deletion error in this region while the dynamic programming approach makes a deletion and a substitution error. ;As an additional point it is to be noted that the differences between the "moussaka" and ''moussaka canu" methods illustrate an independence on the quality of the scaffold sequence (the canu assemblies themselves achieve accuracies of Qscore=26), this can be taken as an example of a benefit. ;Overall, the MSA-NN has been shown to improve performance, which is considered to have been achieved, at least in part, by the inhibition of reference-bias via removal of the reference (scaffold) sequence from the pileup construction. ;Method -St/3'd alternative Figure 5 shows the improved method in its simplest form -that is obtaining reads, creating an MSA graph and applying neural network analysis to extract a consensus sequence. This direct application of neural networks to an MSA graph to form a consensus sequence without the use of an intermediate pileup representation (or predetermined algorithmic extraction of features) is referred to as an MSA-graph neural network (MSA-GNN) ;U ;By way of background, a graph G = (V, E) consists a set of vertices (or nodes) V and edges E. Nodes may have one or more attributes. The edges connect nodes. Edges can be either directed or undirected: a POA graph is a directed graph without cycles (paths which loop-back on themselves), a directed acyclic graph (DAG). An example POA graph is shown in Figure 6. ;The nodes of such graphs are represented as ovals, the letters within these are bases. The graph edges are represented as arrows, the numbers alongside the arrows indicate the number of reads which traverse the edge, i.e. contain the bases within the nodes connected by the edge. Dashed lines indicate nodes which are aligned but with mismatched bases. No two nodes containing an identical base are connected by a dashed line as this would constitute a match and the two nodes collapse to a single node. The orange nodes (shaded, when black and white) and edges trace the path of the consensus sequence extracted from the graph via a dynamic programming method. ;The first step of the MSA-GNN approach is a to build a POA graph whilst retaining 15 knowledge of which of the two complementary DNA strands was sequenced. One method to accomplish this is to record simply two weights for every edge, one for each strand. ;A second method is to record this information as attributes of the nodes, a count of bases from each strand independently. These two approaches are in essence the same, the former has the benefit of encoding directly counts of pairs of bases occurring in reads. These "transition counts" form a useful input to a neural network. ;An alternative more distinct method is to create two subgraphs, one for each strand. In this second approach a graph-to-graph alignment can be made to find a correspondence between the nodes of the two subgraphs if required. This alternative has the benefit that the graph construction can take advantage of possible differences in error modes of reads of the two strands, i.e. the relative rates of substitutions, deletions, and insertions. ;Having obtained a POA graph (or graphs in the strand specific graph case) neural networks can be leveraged to extract a consensus sequence. Below we outline how various GNN structures can be brought to bear on this task. ;Graph convolutional networks Convolutional neural networks (CNNs) have been leveraged on a plethora of tasks in recent years. A convolution is defined succinctly as the application of a filter function to a neighbourhood of points around a centre point; the output signal is the result of this application in turn to all input points considered the centre point. Mostly CNNs have been limited to data I 3' with a well-defined structure, two chief examples axe: I. Audio data: an ordered 1-dimensional series of values; and 2. Image data: a 2-dimension array of values where only nearest neighbours are considered connected. ;The notion of a convolution extends quite naturally to arbitrary graphs: it is not limited to linear (audio) or lattice (image) structures. In a graph convolution an output graph is constructed with the same connectivity (same node and edge structure) but with new node attributes. The convolutional filter is again applied to a neighbourhood of nodes around a central node, the output of which is stored on the equivalent node in the output graph. In similarity to the structured cases, the convolutional filters are learnt during a training process such that the node attributes of the output graph contain a desired response [15]. A graph convolution network (GCN) may comprise multiple convolutional stages, or layers, and additional transforms such as non-linear filters and graph simplifications (pooling). ;Within the consensus problem and for the purpose of the convolution, the graph edge weights should be considered as part of a node's attributes (along with counts of bases, possibly segregated by strand). It can also be advantageous to either: a) form node attributes corresponding to counts of >1 base substrings from the reads, or b) directly form the POA from overlapping substrings of the reads (together with an expanded substitution matrix). These methods allow direct flow of information concerning bases neighbouring each other within reads, a task which otherwise would have to be recapitulated within the actions of the neural network. ;In order to train a GCN to extract a consensus sequence we train it to trace the correct path through an input graph. This can be achieved by aligning the truth sequence to the graph to determine which edges should be followed. To achieve maximal performance, this method can use the prerequisite nodes present in the graph. Cases where this criterion is not met are useful to include in training as they also can arise during inference. For training one can choose a path with minimal edit-distance to the truth sequence. ;Nodes in the output graph are assigned one of two labels: whether or not they belong to the truth path. The training objective is to score the likelihood of these labels and maximise some function of the scores. The process is analogous to image segmentation tasks using 30 CNNs. ;During prediction the GCN is applied to a POA graph to produce a graph with the same structure but with node attributes indicating the scores of the two labels. For extracting a consensus sequence the simplest case is that all label scores are identically 0 or 1 and a path exists with no branches or dead ends. In practice this does not occur, and node labels lie in the ;N ;range (0,1); a supplementary method must be used to extract a single best sequence with maximal score. Methods are known to accomplish this task. At first sight it appears the classical POA dynamic programming solution has simply been swapped for a second dynamic programming problem. However, we note that the use of a GCN can output a graph which is more amenable to decoding with a static (time-independent) algorithm; the GCN is responsible for transforming a graph with complex long-range interactions into a graph with a simpler interpretation. ;Graph auto-encoder The transform of the complex POA graph to the linear chain of the consensus sequence can be viewed as a form of lossy compression. In the parlance of neural networks this is a task performed by lossy "autoencoders". An autoencoder consists of two components, an encoder and a decoder. For a non-lossy encoder/decoder pair the effect of applying the encoder to an input followed by subsequent output of the decoder is to regenerate the input. A lossy pair however can perform a non-trivial transform of the input, to the extent that the output can constitute the adjacency matrix of a new graph. Both the encoder and decoder can comprise a GCN, in which case the resultant model is conceptually simply a GCN. ;Alternatively, the encoder can be a GCN, the output of which is fed to a generative model. Examples where the generative model is an inner product of the GCN output are known [16]. ;Convolution on graph adjacency matrix The extraction of a consensus sequence from a POA graph is not limited to use of graph convolutional networks: a more typical lattice convolution can also be employed. To achieve this the POA graph adjacency matrix is first constructed. Here it is useful to topologically sort the nodes with respect to the matrix indices to maintain notions of spatial locality. This adjacency matrix contains an additional dimension (or dimensions) corresponding to the multiplicity of the graph node attributes and can be used as input to a non-graph CNN. Alternatively, as the matrix is index sorted it may be natural to utilise a recurrent neural network, treating the first dimension of the matrix as a series of time-points. ;The output of these networks can be trained to be one of the two we have seen previously above: 1. the adjacency matrix corresponding to a new simplified graph (akin to a lossy autoencoder); 2. a labelling (truth sequence scoring) of the nodes of the original graph (as in the GCN approach). A third, hybrid, option is possible also: 3. an adjacency matrix of a simplified graph but with node labelling. This third option combines hard choices of removal of some information from the input graph with a representation of uncertainty when required. One way to achieve this is as a subclass of 2.: one can add in an additional score indicating that that node should be interpreted as removed entirely from the graph. For effective training it can be useful to modify the function being optimized to include a penalty term for the removal of a node. In common to the GCN approach to the consensus task, the output of this graph adjacency matrix approach can require decoding via supplementary methods. ;The features in Figures of the invention are interchangeable and compatible in light of the teaching herein. The present invention has been described above purely by way of example, and modifications can be made within the spirit and scope of the invention, which extends to equivalents of the features described and combinations of one or more features described herein. The invention also consists in any individual features described or implicit herein. ;Bibliography [1] "Multiple sequence alignment," [Online]. Available: httpsfien.wikipedia.org/wiki/Multiple sequence alignment. ;[2] Durbin, "Multiple sequence alignment methods," in Biological sequence analysis. ;[3] Lee, "Generating consensus sequences from partial order multiple sequence alignment graphs," Bioinformatics, vol. 19, no. 8, pp. 999-1008, 22 May 2003. ;[4] "SNV calling from NGS data," [Online]. Available: https://en.wiki pedi a. org/wi ki/S N V_cal 1 i ng_from_NG S_data [5] Poplin, "A universal SNP and small-indel variant caller using deep neural networks," Nature Biotechnology, vol. 36, pp. 983-987, September 2018. ;[6] Luo, "Clairvoyante: a multi-task convolutional deep neural network for variant calling in Single Molecule Sequencing," 2018. [Online]. Available: https://doi.org/10.1101/310458. ;[7] ONT, "medaka," 2016-2019. [Online]. Available: https.figithub.com/nanoporetech/medaka. ;[8] Chin, "DCNet Denoi sing (DNA) Sequence With a LSTM-RNN and PyTorch," 28 March 2018. [Online]. Available: https://medium. com/ainfoecho/dcnet-denoising-dna-sequence-with-a-lstm-rnn-and-pytorch-3b4 54f1727e7. ;[9] Garrison, "Variation graph toolkit improves read mapping by representing genetic variation in the reference," Nature Biotechnology, vol. 36, pp. 875-879, 2018. ;[10] Lee, "Multiple sequence alignment using partial order graphs," Bioinformatics, vol. 18, no. 3, pp. 452-464, March 2002. ;[11] Dursi, "Understanding Partial Order Alignment for Multiple Sequence Alignment," May 2015. [Online]. Available: http://si mp sonl ab.gi thub o/2015/05/01/understanding-poa/. ;[12] Vaser, "Fast and accurate de novo genome assembly from long uncorrected reads," Genome Research, vol. 27, pp. 737-746, 18 January 2017. ;[13] Vaser, "spoa," [Online]. Available: https://github.com/rvaser/spoa [Accessed March 2019]. ;[14] Koren, "Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation," Genome Research, vol. 27, pp. 722-736, 2017. ;[15] Z. Wu, "A Comprehensive Survey on Graph Neural Networks," 10 03 2019. [Online]. Available: https://arxiv.org/abs/1901.00596. [Accessed 04 2019]. ;[16] T. Kipf, "arXiv," 11 21 2016. [Online]. Available: https://arxiv.org/abs/1611.07308. ;[Accessed 04 2019]. ;[17] Wright, "align dtw.py," 9 October 2014. [Online]. Available ts* wthianolabsiocaliresearaii 0,n) /97e0,3:4 e get, c f 181 havs./A4ithib comina no' F'
GB1907165.3A 2019-05-21 2019-05-21 Method for determining a consensus sequence of a target polymer Withdrawn GB2579110A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
GB1907165.3A GB2579110A (en) 2019-05-21 2019-05-21 Method for determining a consensus sequence of a target polymer

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
GB1907165.3A GB2579110A (en) 2019-05-21 2019-05-21 Method for determining a consensus sequence of a target polymer

Publications (2)

Publication Number Publication Date
GB201907165D0 GB201907165D0 (en) 2019-07-03
GB2579110A true GB2579110A (en) 2020-06-10

Family

ID=67385218

Family Applications (1)

Application Number Title Priority Date Filing Date
GB1907165.3A Withdrawn GB2579110A (en) 2019-05-21 2019-05-21 Method for determining a consensus sequence of a target polymer

Country Status (1)

Country Link
GB (1) GB2579110A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
TWI801287B (en) * 2021-07-20 2023-05-01 奧義智慧科技股份有限公司 Event visualization device and related computer program product for generating hierarchical directed acyclic graph

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP4211691A1 (en) * 2020-09-11 2023-07-19 F. Hoffmann-La Roche AG Deep-learning-based techniques for generating a consensus sequence from multiple noisy sequences
CN114237154B (en) * 2021-11-26 2023-06-30 重庆大学 Thermal error prediction model based on ONT-GCN space-time model, modeling method and haze-edge-fog-cloud error compensation system

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
arXiv preprint arXiv, 2019, MICULINIC N. ET AL., "MinCall MinION end2end convolutional deep learning basecaller", 1904.10337 *
Bioinformatics, 1999, Vol 15 No 9, ALLEX C.F. ET AL., "Neural network input representations that produce accurate consensus sequences from DNA fragment assemblies", pages 723-728 *
GigaScience, 2018, Vol 7 No 5, TENG H. ET AL., "Chiron: translating nanopore raw signal directly into nucleotide sequence using deep learning", Article giy037 *

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
TWI801287B (en) * 2021-07-20 2023-05-01 奧義智慧科技股份有限公司 Event visualization device and related computer program product for generating hierarchical directed acyclic graph

Also Published As

Publication number Publication date
GB201907165D0 (en) 2019-07-03

Similar Documents

Publication Publication Date Title
KR20240124392A (en) Variant calling without using target reference genome
US6681186B1 (en) System and method for improving the accuracy of DNA sequencing and error probability estimation through application of a mathematical model to the analysis of electropherograms
Anisimova et al. State-of the art methodologies dictate new standards for phylogenetic analysis
JP6066924B2 (en) DNA sequence data analysis method
Redelings Erasing errors due to alignment ambiguity when estimating positive selection
Bernardes et al. A multi-objective optimization approach accurately resolves protein domain architectures
Bai et al. Hierarchical clustering split for low-bias evaluation of drug-target interaction prediction
CN113555062B (en) Data analysis system and analysis method for genome base variation detection
GB2579110A (en) Method for determining a consensus sequence of a target polymer
CN119418762A (en) A genetic testing data cleaning method and system based on artificial intelligence
WO2022123067A2 (en) Method and apparatus for classification and/or prioritization of genetic variants
Papetti et al. Barcode demultiplexing of nanopore sequencing raw signals by unsupervised machine learning
Černý et al. Statistical evaluation of character support reveals the instability of higher-level dinosaur phylogeny
Pagni et al. Density-based hierarchical clustering of pyro-sequences on a large scale—the case of fungal ITS1
Liu et al. A New Method for Constructing High‐Resolution Phylogenomic Topologies Using Core Gene‐Associated MNP Markers: A Case Study From Agaricus bisporus
CN117275577A (en) Algorithm for detecting human mitochondrial genetic mutation sites based on second-generation sequencing technology
CN119832980B (en) Gene mutation detection method, device, electronic equipment and storage medium
Pinto et al. Ranked Adjusted Rand: integrating distance and partition information in a measure of clustering agreement
KR101629178B1 (en) Apparatus for technology life analysis using multiple patent indicators
KR102405732B1 (en) Method and appratus for cell clustering
Nyström-Persson et al. Precise and scalable metagenomic profiling with sample-tailored minimizer libraries
Zachariasen et al. Identification of representative species-specific genes for abundance measurements
WO2023181370A1 (en) Information processing device, information processing method, and information processing program
David et al. Targeted domain assembly for fast functional profiling of metagenomic datasets with S3A
JP4894860B2 (en) Nucleotide sequence reliability calculation method

Legal Events

Date Code Title Description
WAP Application withdrawn, taken to be withdrawn or refused ** after publication under section 16(1)