EP4374378A1 - Specialist signal profilers for base calling - Google Patents
Specialist signal profilers for base callingInfo
- Publication number
- EP4374378A1 EP4374378A1 EP22751936.0A EP22751936A EP4374378A1 EP 4374378 A1 EP4374378 A1 EP 4374378A1 EP 22751936 A EP22751936 A EP 22751936A EP 4374378 A1 EP4374378 A1 EP 4374378A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- signal
- specialist
- profilers
- different
- analytes
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
- 238000012549 training Methods 0.000 claims abstract description 82
- 230000015654 memory Effects 0.000 claims abstract description 44
- 239000012491 analyte Substances 0.000 claims abstract description 26
- 238000012163 sequencing technique Methods 0.000 claims description 272
- 238000009826 distribution Methods 0.000 claims description 67
- 230000002123 temporal effect Effects 0.000 claims description 42
- 238000000034 method Methods 0.000 claims description 36
- 230000002093 peripheral effect Effects 0.000 claims description 18
- 230000008569 process Effects 0.000 claims description 16
- 230000004044 response Effects 0.000 claims description 7
- 238000013507 mapping Methods 0.000 claims description 5
- 239000000523 sample Substances 0.000 description 45
- 238000013528 artificial neural network Methods 0.000 description 35
- 238000003384 imaging method Methods 0.000 description 33
- 238000012545 processing Methods 0.000 description 33
- 230000003287 optical effect Effects 0.000 description 17
- 238000005516 engineering process Methods 0.000 description 16
- 230000006870 function Effects 0.000 description 15
- 230000011218 segmentation Effects 0.000 description 12
- 239000000758 substrate Substances 0.000 description 11
- 238000004458 analytical method Methods 0.000 description 10
- 238000003860 storage Methods 0.000 description 10
- 238000012937 correction Methods 0.000 description 8
- 150000007523 nucleic acids Chemical class 0.000 description 8
- 239000012472 biological sample Substances 0.000 description 7
- 238000005286 illumination Methods 0.000 description 7
- 102000039446 nucleic acids Human genes 0.000 description 7
- 108020004707 nucleic acids Proteins 0.000 description 7
- 238000010223 real-time analysis Methods 0.000 description 7
- 238000005192 partition Methods 0.000 description 6
- 238000007493 shaping process Methods 0.000 description 6
- 230000003321 amplification Effects 0.000 description 5
- 230000006872 improvement Effects 0.000 description 5
- 230000001537 neural effect Effects 0.000 description 5
- 238000003199 nucleic acid amplification method Methods 0.000 description 5
- 239000002773 nucleotide Substances 0.000 description 5
- 125000003729 nucleotide group Chemical group 0.000 description 5
- 239000000203 mixture Substances 0.000 description 4
- 230000009466 transformation Effects 0.000 description 4
- 230000003044 adaptive effect Effects 0.000 description 3
- 238000013459 approach Methods 0.000 description 3
- 230000015572 biosynthetic process Effects 0.000 description 3
- 239000003153 chemical reaction reagent Substances 0.000 description 3
- 239000007788 liquid Substances 0.000 description 3
- 230000007246 mechanism Effects 0.000 description 3
- 238000011176 pooling Methods 0.000 description 3
- 230000000306 recurrent effect Effects 0.000 description 3
- 230000002441 reversible effect Effects 0.000 description 3
- 239000000126 substance Substances 0.000 description 3
- 230000000007 visual effect Effects 0.000 description 3
- 230000008901 benefit Effects 0.000 description 2
- 230000001427 coherent effect Effects 0.000 description 2
- OPTASPLRGRRNAP-UHFFFAOYSA-N cytosine Chemical compound NC=1C=CNC(=O)N=1 OPTASPLRGRRNAP-UHFFFAOYSA-N 0.000 description 2
- 238000001514 detection method Methods 0.000 description 2
- 238000010586 diagram Methods 0.000 description 2
- 230000000694 effects Effects 0.000 description 2
- 239000000835 fiber Substances 0.000 description 2
- 239000012530 fluid Substances 0.000 description 2
- 239000007850 fluorescent dye Substances 0.000 description 2
- UYTPUPDQBNUYGX-UHFFFAOYSA-N guanine Chemical compound O=C1NC(N)=NC2=C1N=CN2 UYTPUPDQBNUYGX-UHFFFAOYSA-N 0.000 description 2
- 238000010191 image analysis Methods 0.000 description 2
- 238000004519 manufacturing process Methods 0.000 description 2
- 238000013178 mathematical model Methods 0.000 description 2
- 239000011159 matrix material Substances 0.000 description 2
- 238000012634 optical imaging Methods 0.000 description 2
- 238000005457 optimization Methods 0.000 description 2
- XEBWQGVWTUSTLN-UHFFFAOYSA-M phenylmercury acetate Chemical compound CC(=O)O[Hg]C1=CC=CC=C1 XEBWQGVWTUSTLN-UHFFFAOYSA-M 0.000 description 2
- 238000001228 spectrum Methods 0.000 description 2
- 238000003786 synthesis reaction Methods 0.000 description 2
- 238000012360 testing method Methods 0.000 description 2
- RWQNBRDOKXIBIV-UHFFFAOYSA-N thymine Chemical compound CC1=CNC(=O)NC1=O RWQNBRDOKXIBIV-UHFFFAOYSA-N 0.000 description 2
- PXFBZOLANLWPMH-UHFFFAOYSA-N 16-Epiaffinine Natural products C1C(C2=CC=CC=C2N2)=C2C(=O)CC2C(=CC)CN(C)C1C2CO PXFBZOLANLWPMH-UHFFFAOYSA-N 0.000 description 1
- ORILYTVJVMAKLC-UHFFFAOYSA-N Adamantane Natural products C1C(C2)CC3CC1CC2C3 ORILYTVJVMAKLC-UHFFFAOYSA-N 0.000 description 1
- 229930024421 Adenine Natural products 0.000 description 1
- GFFGJBXGBJISGV-UHFFFAOYSA-N Adenine Chemical compound NC1=NC=NC2=C1N=CN2 GFFGJBXGBJISGV-UHFFFAOYSA-N 0.000 description 1
- 240000001436 Antirrhinum majus Species 0.000 description 1
- 241000894006 Bacteria Species 0.000 description 1
- 208000037170 Delayed Emergence from Anesthesia Diseases 0.000 description 1
- 229920002307 Dextran Polymers 0.000 description 1
- 102000004190 Enzymes Human genes 0.000 description 1
- 108090000790 Enzymes Proteins 0.000 description 1
- 239000004743 Polypropylene Substances 0.000 description 1
- 239000004793 Polystyrene Substances 0.000 description 1
- 240000000136 Scabiosa atropurpurea Species 0.000 description 1
- 241001290864 Schoenoplectus Species 0.000 description 1
- XUIMIQQOPSSXEZ-UHFFFAOYSA-N Silicon Chemical compound [Si] XUIMIQQOPSSXEZ-UHFFFAOYSA-N 0.000 description 1
- 230000004913 activation Effects 0.000 description 1
- 229960000643 adenine Drugs 0.000 description 1
- 238000003491 array Methods 0.000 description 1
- 238000003556 assay Methods 0.000 description 1
- 239000012620 biological material Substances 0.000 description 1
- 239000000872 buffer Substances 0.000 description 1
- 230000015556 catabolic process Effects 0.000 description 1
- 238000006243 chemical reaction Methods 0.000 description 1
- 238000003776 cleavage reaction Methods 0.000 description 1
- 230000006835 compression Effects 0.000 description 1
- 238000007906 compression Methods 0.000 description 1
- 239000000470 constituent Substances 0.000 description 1
- 229940104302 cytosine Drugs 0.000 description 1
- 238000013480 data collection Methods 0.000 description 1
- 230000007423 decrease Effects 0.000 description 1
- 238000013135 deep learning Methods 0.000 description 1
- 238000006731 degradation reaction Methods 0.000 description 1
- 230000001419 dependent effect Effects 0.000 description 1
- 238000013461 design Methods 0.000 description 1
- 230000001627 detrimental effect Effects 0.000 description 1
- 230000002708 enhancing effect Effects 0.000 description 1
- 230000005284 excitation Effects 0.000 description 1
- 230000007717 exclusion Effects 0.000 description 1
- 230000001747 exhibiting effect Effects 0.000 description 1
- 238000000605 extraction Methods 0.000 description 1
- 238000005562 fading Methods 0.000 description 1
- 238000001914 filtration Methods 0.000 description 1
- 238000000799 fluorescence microscopy Methods 0.000 description 1
- 239000012634 fragment Substances 0.000 description 1
- 239000000499 gel Substances 0.000 description 1
- 230000002068 genetic effect Effects 0.000 description 1
- 239000011521 glass Substances 0.000 description 1
- PCHJSUWPFVWCPO-UHFFFAOYSA-N gold Chemical compound [Au] PCHJSUWPFVWCPO-UHFFFAOYSA-N 0.000 description 1
- 239000010931 gold Substances 0.000 description 1
- 229910052737 gold Inorganic materials 0.000 description 1
- 238000012165 high-throughput sequencing Methods 0.000 description 1
- 239000012535 impurity Substances 0.000 description 1
- 230000010354 integration Effects 0.000 description 1
- 230000003993 interaction Effects 0.000 description 1
- 238000003064 k means clustering Methods 0.000 description 1
- 239000004816 latex Substances 0.000 description 1
- 229920000126 latex Polymers 0.000 description 1
- 239000004973 liquid crystal related substance Substances 0.000 description 1
- 238000007477 logistic regression Methods 0.000 description 1
- 239000000463 material Substances 0.000 description 1
- 238000002156 mixing Methods 0.000 description 1
- 238000012986 modification Methods 0.000 description 1
- 230000004048 modification Effects 0.000 description 1
- 230000008450 motivation Effects 0.000 description 1
- 238000010606 normalization Methods 0.000 description 1
- 230000005693 optoelectronics Effects 0.000 description 1
- 230000002085 persistent effect Effects 0.000 description 1
- 230000000704 physical effect Effects 0.000 description 1
- 239000004033 plastic Substances 0.000 description 1
- 229920003023 plastic Polymers 0.000 description 1
- 229920002401 polyacrylamide Polymers 0.000 description 1
- 102000040430 polynucleotide Human genes 0.000 description 1
- 108091033319 polynucleotide Proteins 0.000 description 1
- 239000002157 polynucleotide Substances 0.000 description 1
- -1 polypropylene Polymers 0.000 description 1
- 229920001155 polypropylene Polymers 0.000 description 1
- 229920002223 polystyrene Polymers 0.000 description 1
- 102000004169 proteins and genes Human genes 0.000 description 1
- 108090000623 proteins and genes Proteins 0.000 description 1
- 230000007017 scission Effects 0.000 description 1
- 238000005204 segregation Methods 0.000 description 1
- 230000006403 short-term memory Effects 0.000 description 1
- 229910052710 silicon Inorganic materials 0.000 description 1
- 239000010703 silicon Substances 0.000 description 1
- 230000001360 synchronised effect Effects 0.000 description 1
- 229940113082 thymine Drugs 0.000 description 1
- 238000000844 transformation Methods 0.000 description 1
- 238000013519 translation Methods 0.000 description 1
- 238000013024 troubleshooting Methods 0.000 description 1
- 235000012431 wafers Nutrition 0.000 description 1
- 238000005406 washing Methods 0.000 description 1
- 239000002699 waste material Substances 0.000 description 1
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/10—Signal processing, e.g. from mass spectrometry [MS] or from PCR
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/20—Supervised data analysis
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6869—Methods for sequencing
Definitions
- the technology disclosed relates to apparatus and corresponding methods for the automated analysis of an image or recognition of a pattern. Included herein are systems that transform an image for the purpose of (a) enhancing its visual quality prior to recognition, (b) locating and registering the image relative to a sensor or stored prototype, or reducing the amount of image data by discarding irrelevant data, and (c) measuring significant characteristics of the image.
- the technology disclosed relates to removing spatial crosstalk from sensor pixels using equalization-based image processing techniques.
- Base calling accuracy is crucial for high-throughput sequencing and downstream analysis such as read mapping and genome assembly.
- This disclosure relates to optimizing image data to accurately base call clusters during a sequencing run.
- One challenge with the optimization of image data is variation in intensity profiles (or intensity distributions) of clusters in a cluster population being base called. This is particularly detrimental for multi-cycle imaging of substrates ( e.g ., flow cells) having a large number (e.g, thousands, millions, billions, etc.) of clusters, as it makes the scale of variation unmanageable, thereby causing a drop in data throughput and an increase in error rate.
- Intensity profiles of millions of clusters on a flow cell can vary between respective clusters or between subpopulations of clusters. There are many potential reasons for this variation. It may result from differences in cluster brightness, caused by fragment length distribution in the cluster population or unwanted light emissions from adjacent clusters (spatial crosstalk). It may result from phase error, which occurs when a molecule in a cluster does not incorporate a nucleotide in some sequencing cycle and lags behind other molecules, or when a molecule incorporates more than one nucleotide in a single sequencing cycle. It may result from fading, i.e., an exponential decay in signal intensity of clusters as a function of sequencing cycle number due to excessive washing and laser exposure as the sequencing run progresses.
- underdeveloped cluster colonies i.e., small cluster sizes that produce empty or partially filled wells on a patterned flow cell. It may result from overlapping cluster colonies caused by unexclusive amplification. It may result from under-illumination or uneven- illumination, for example, due to clusters being located on edges of a flow cell. It may result from impurities (e.g ., bubbles) on a flow cell that obfuscate emitted signal. It may result from polyclonal clusters, i.e., when multiple clusters are deposited in the same well. It may result from different types of distortion in the image induced by the geometry of the optical lens. Such distortions may include, for example, magnification distortion, skew distortion, translation distortion, and nonlinear distortions such as barrel distortion and pincushion distortion.
- This variation can be corrected in a coarse way by training an intensity corrector for the entire cluster population. Different than this would be to train respective intensity correctors for respective subpopulations of clusters, where the subpopulations are segmented in a way that minimizes sequencing errors and maximizes base calling accuracy within the bounds of available compute. This disclosure relates to the latter more granular approach. More details follow.
- Figure 1 shows an example sequencing environment with an imaging system.
- Figure 2 is a block diagram illustrating an example two-channel, line-scanning modular optical imaging system that can be implemented in particular implementations.
- Figure 3 shows one implementation of respective signal profilers 1 to N trained to maximize signal-to-noise ratio of respective image data subsets 1 to N generated for respective classes 1 to N of clusters located on a flow cell in respective spatial configurations 1 to N.
- Figure 4 illustrates an example configuration of a flow cell that can be imaged in accordance with implementations disclosed herein.
- Figure 5 A depicts lanes of a top surface of a flow cell.
- Figure 5B shows swathes of tiles in a lane of a top surface of a flow cell.
- Figure 5C illustrates a tile in a swath of a lane of a top surface of a flow cell.
- Figure 5D portrays sub-tiles in a tile in a swath of a lane of a top surface of a flow cell.
- Figure 6A shows one implementation of training respective surface-specific specialist signal profilers for respective cluster classes during a sequencing run 600
- Figure 6B shows one implementation of applying the trained surface-specific specialist signal profilers to image data subsets corresponding to the respective cluster classes.
- Figure 7 A shows one implementation of training respective lane group-specific specialist signal profilers for respective cluster classes.
- Figure 7B shows one implementation of training respective lane-specific specialist signal profilers for respective cluster classes.
- Figure 7C shows one implementation of training respective swath-specific specialist signal profilers for respective cluster classes.
- Figure 7D shows one implementation of training respective tile-specific specialist signal profilers for respective cluster classes.
- Figure 7E shows one implementation of training respective sub-tile-specific specialist signal profilers for respective cluster classes.
- Figure 8 shows one implementation of applying the trained lane group-specific specialist signal profilers to image data subsets corresponding to the respective cluster classes during a sequencing run.
- Figure 9 shows one implementation of applying the trained lane-specific specialist signal profilers to image data subsets corresponding to the respective cluster classes during a sequencing run.
- Figure 10 shows one implementation of applying the trained swath-specific specialist signal profilers to image data subsets corresponding to the respective cluster classes during a sequencing run.
- Figure 11 shows one implementation of applying the trained tile-specific specialist signal profilers to image data subsets corresponding to the respective cluster classes during a sequencing run.
- Figure 12 shows one implementation of applying the trained sub-tile-specific specialist signal profilers to image data subsets corresponding to the respective cluster classes during a sequencing run.
- Figure 13 shows one implementation of respective/separate/different/independent specialist signal profilers for respective sub-series of sequencing cycles of a sequencing run with a total ofN sequencing cycles.
- Figure 14 shows one implementation of respective/separate/different/independent specialist signal profilers for a combination of different spatial configurations (e.g, different sub tiles) and different temporal configurations (e.g, different sub-series of sequencing cycles).
- Figure 15 shows one implementation of respective/separate/different/independent specialist signal profilers for each cluster/well sequenced during a sequencing run.
- Figure 16 shows one implementation of offline training of specialist signal profilers on sequenced data from one or more completed/already-executed sequencing runs, and application of the trained specialist signal profilers on sequenced data from an ongoing sequencing run.
- Figure 17 shows an example tile image partitioned into sub-tiles images.
- Figure 18 shows one implementation of online training of specialist signal profilers on sequenced data from earlier sequencing cycles of an ongoing sequencing run, and application of the trained specialist signal profilers on sequenced data from later sequencing cycles of the ongoing sequencing run.
- Figure 19 shows one implementation of training respective/separate/different/independent specialist signal profilers for respective signal distributions observed in sequenced data.
- Figure 20 shows one example of a signal distribution/signal profile/cluster intensity profile.
- Figure 21 shows one implementation of a processing pipeline that implements the technology disclosed.
- Figures 22A-22E show equalizer coefficient sets of an example spatial equalizer.
- Figures 23-31 show one implementation of training an equalizer.
- Figure 32A shows base-wise signal distributions of a cluster population without use of an equalizer, and with a signal -to-noise ratio of 11.96 decibels (dBs).
- Figure 32B shows the base-wise signal distributions of the same cluster population with the use of an equalizer, and with an improvement in the signal -to-noise ratio to 13.13 dBs.
- Figure 33 shows how a cost function for a specialist signal profiler improves with each iteration of gradient descent.
- Figure 34 is a plot that shows initial and final values of the cost function of Figure 33 when the specialist signal profiler is adapted/trained/configured/updated at each sequencing cycle.
- Figures 35 A and 35B show the improvement in primary analysis metrics for a sequencing run when we adapt/train/configure/update the specialist signal profiler.
- Figures 36A and 36B show two plots that assess a number of sub-tiles into which a sequencing tile can be partitioned for adaptive equalization of respective specialist signal profilers.
- Figure 37 shows an example computer system that can be used to implement the technology disclosed. PET ATT, ED DESCRIPTION
- a “signal profiler” maximizes the signal-to-noise ratio of a signal that is disturbed by noise.
- a signal profiler can be a value or function that is applied to data to modify the data in a desired way. For example, the data can be modified to increase its accuracy, relevance, or applicability with regard to a particular situation.
- the signal profiler can be applied to the data by any of a variety of mathematical manipulations including, but not limited to addition, subtraction, division, multiplication, or a combination thereof.
- the signal profiler can be a mathematical formula, logic function, computer implemented algorithm, or the like.
- the data can be image data, electrical data, or a combination thereof.
- the signal profiler is an equalizer (e.g ., a spatial equalizer).
- the equalizer can be trained (e.g., using least square estimation, adaptive equalization algorithm) to maximize the signal-to-noise ratio of cluster intensity data in sequencing images.
- the equalizer is a lookup table (LUT) bank with a plurality of LUTs with subpixel resolution, also referred to as “equalizer filters” or “convolution kernels ”
- the number of LUTs in the equalizer depends on the number of subpixels into which pixels of the sequencing images can be divided. For example, if the pixels are divisible into n by n subpixels (e.g, 5 x 5 subpixels), then the equalizer generates n 2 LUTs (e.g, 25 LUTs).
- data from the sequencing images is binned by well subpixel location. For example, for a 5 x 5 LUT, approximately l/25th of the wells have a center that is in bin (1,1) (e.g, the upper left comer of a sensor pixel), l/25th of the wells are in bin (1,2), and so on.
- the equalizer coefficients for each bin are determined using least squares estimation on the subset of data from the wells corresponding to the respective bins. This way the resulting estimated equalizer coefficients are different for each bin.
- Each LUT/equalizer filter/convolution kernel has a plurality of coefficients that are learned from the training.
- the number of coefficients in a LUT corresponds to the number of pixels that are used for base calling a cluster. For example, if a local grid of pixels (image or pixel patch) that is used to base call a cluster is of size p x p ( e.g ., 9 x 9 pixel patch), then each LUT has p 2 coefficients (e.g., 81 coefficients).
- the training produces equalizer coefficients that are configured to mix/combine intensity values of pixels that depict intensity emissions from a target cluster being base called and intensity emissions from one or more adjacent clusters in a manner that maximizes the signal-to-noise ratio.
- the signal maximized in the signal-to-noise ratio is the intensity emissions from the target cluster
- the noise minimized in the signal-to-noise ratio is the intensity emissions from the adjacent clusters, i.e., spatial crosstalk, plus some random noise (e.g, to account for background intensity emissions).
- the equalizer coefficients are used as weights and the mixing/combining includes executing element-wise multiplication between the equalizer coefficients and the intensity values of the pixels to calculate a weighted sum of the intensity values of the pixels, i.e., a convolution operation. Furthermore, in cases the image data spans across multiple color channels, a set of equalizer coefficients is generated for each color channel (e.g, one channel, three channels, four channels, etc.).
- Figures 22A-22E show equalizer coefficient sets of an example spatial equalizer. As indicated by the heat maps, the different equalizer coefficient sets are configured to differently attenuate and augment signals of pixels depending on locations of the pixels.
- Figure 23 shows one implementation of training an equalizer.
- the equalizer in Figure 23 has a first set of equalizer coefficients 2302 for the green color channel, and a second set of equalizer coefficients 2304 for the blue color channel.
- a first cluster has input image pixels 2306 for the green color channel and input image pixels 2308 for the blue color channel.
- the first set of equalizer coefficients 2302 are element-wise multiplied 2402 with the input image pixels 2306 to generate a weighted sum 2316 for the green color channel.
- the second set of equalizer coefficients 2304 are element-wise multiplied 2502 with the input image pixels 2308 to generate a weighted sum 2318 for the blue color channel.
- a base calling logic 2322 uses the expectation maximization (EM) algorithm discussed above to predict a base call 2324 based on the weighted sums 2316, 2318.
- EM expectation maximization
- the weighted sum 2316 for the green color channel is compared against the centroid value 2612 of the called base for the green color channel with the base calling logic 2322.
- the comparison yields a base calling error 2336 for the green color channel.
- the weighted sum 2318 for the blue color channel is compared against the centroid value 2712 of the called base for the blue color channel with the base calling logic 2322.
- the comparison yields a base calling error 2338 for the blue color channel.
- the base calling errors 2336, 2338 are used by an update logic 2342 to produce an updated first set of equalizer coefficients 2356 for the green color channel, and an updated second set of equalizer coefficients 2358 for the blue color channel.
- the above steps are executed for a plurality of clusters.
- three updated versions of the first set of equalizer coefficients for the green color channel, and three updated versions of the second set of equalizer coefficients for the blue color channel are generated.
- the three updated versions are used to calculate, for a second sequencing cycle (cycle 2), a first set of equalizer coefficients 2362 for the green color channel, and a second set of equalizer coefficients 2354 for the blue color channel.
- Figure 32A shows base-wise signal distributions of a cluster population without use of an equalizer, and with a signal -to-noise ratio of 11.96 decibels (dBs).
- Figure 32B shows the base-wise signal distributions of the same cluster population with the use of an equalizer, and with an improvement in the signal-to-noise ratio to 13.13 dBs. The improvement in the signal-to- noise ratio is also visually observable by tighter/more discrete base-wise clouds/distributions in Figure 32B compared to the base-wise clouds in Figure 32A.
- a global intensity correction applied at a pan-flow cell level or a pan-sequencing run level fails to consider a variety of noises in the image data. For example, non-linear distortion and noise can be induced by the shape of the optical lens that captures the image data.
- the imaged flow cell can also introduce distortion in the well pattern due to the manufacturing process, e.g ., a 3D bathtub effect introduced by bonding or movement of the wells due to non-rigidity of the substrate.
- the tilt of the flow cell within the holder is not accounted for by the global intensity correction.
- a “specialist signal profiler” is a signal profiler that is configured to/trained to maximize the signal-to-noise ratio of a particular category/type/configuration/characteristic/class/bin of data.
- a “surface-specific specialist signal profiler” is configured to/trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular surface or a particular surf ace-type/ category/cl ass ( e.g ., top surfaces or bottom surfaces or surfaces 1 to N of a flow cell).
- a “lane-specific specialist signal profiler” is configured to/trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular lane or a particular lane-type/category/class (e.g., central lanes or peripheral lanes or lanes 1 to N of a flow cell).
- a “tile-specific specialist signal profiler” is configured to/trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular tile or a particular tile-type/category/class (e.g, central tiles or peripheral tiles or tiles 1 to N of a flow cell).
- a “sub-tile-specific specialist signal profiler” is configured to/trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular sub tile or a particular sub-tile-type/category/class (e.g., central sub-tiles or peripheral sub-tiles or sub-tiles 1 to N of a flow cell). More examples and details of the disclosed specialist signal profilers follow.
- a single signal profiler can comprise a plurality of specialist coefficient sets, such that each specialist coefficient set is configured to/trained to maximize the signal-to-noise ratio of a particular category/type/configuration/characteristic/class/bin of data.
- the single signal profiler can comprise a variety of specialist coefficient sets.
- a “surface- specific specialist coefficient set” is configured to/trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular surface or a particular surface- type/category/class (e.g, top surfaces or bottom surfaces or surfaces 1 to N of a flow cell).
- a “lane-specific specialist coefficient set” is configured to/trained to maximize the signal-to-noise ratio of sequencing data of clusters located on a particular lane or a particular lane-type/category/class (e.g, central lanes or peripheral lanes or lanes 1 to N of a flow cell).
- a “tile-specific specialist coefficient set” is configured to/trained to maximize the signal-to- noise ratio of sequencing data of clusters located on a particular tile or a particular tile- type/ category/ cl ass (e.g, central tiles or peripheral tiles or tiles 1 to N of a flow cell).
- a “sub-tile-specific specialist coefficient set” is configured to/trained to maximize the signal-to- noise ratio of sequencing data of clusters located on a particular sub-tile or a particular sub-tile- type/category/class (e.g ., central sub-tiles or peripheral sub-tiles or sub-tiles 1 to N of a flow cell). More examples and details of the disclosed specialist coefficient sets follow.
- the disclosed specialist signal profilers are applicable to clusters located on both patterned and unpatterned surfaces of a flow cell. With unpatterned surfaces, the clusters are randomly distributed on the flow cell. The randomly distributed clusters and data therefor (e.g., images) can be binned spatially, temporally, signal-wise, or by any combination thereof. Accordingly, the specialist signal profilers can be configured and trained for different configurations of the differently binned randomly distributed clusters. With patterned surfaces, the clusters are located on patterned wells with fixed locations. The patterned wells and the constituent clusters can be binned spatially, temporally, signal-wise, or by any combination thereof. Accordingly, the specialist signal profilers can be configured and trained for different configurations of the differently binned patterned clusters.
- the disclosed specialist signal profilers are configuration-specific signal profilers that are trained to maximize the signal -to-noise ratio of image data generated for different configurations of a sequencing run. These configurations can be spatial configurations relating to different regions on a flow cell, temporal configurations relating to different sequencing/imaging cycles of the sequencing run, signal distribution configurations relating to different distributions/pattems of signal profiles observed/encoded in the imaged data, or a combination thereof.
- Figure 1 shows an example sequencing environment with an imaging system 100.
- the example imaging system 100 can include a device for obtaining or producing an image of a sample.
- the example outlined in Figure 1 shows an example imaging configuration of a backlight design implementation. It should be noted that although systems and methods can be described herein from time to time in the context of example imaging system 100, these are only examples with which implementations of the specialist signal profilers disclosed herein can be implemented.
- sample container 110 e.g., a flow cell as described herein
- sample stage 170 under an objective lens 142.
- Light source 160 and associated optics direct a beam of light, such as laser light, to a chosen sample location on the sample container 110.
- the sample fluoresces and the resultant light is collected by the objective lens 142 and directed to an image sensor of camera system 140 to detect the florescence.
- Sample stage 170 is moved relative to objective lens 142 to position the next sample location on sample container 110 at the focal point of the objective lens 142. Movement of sample stage 110 relative to objective lens 142 can be achieved by moving the sample stage itself, the objective lens, some other component of the imaging system 100, or any combination of the foregoing. Further implementations can also include moving the entire imaging system 100 over a stationary sample.
- Fluid delivery module or device 100 directs the flow of reagents (e.g ., fluorescently labeled nucleotides, buffers, enzymes, cleavage reagents, etc.) to (and through) sample container 110 and waste valve 120.
- Sample container 110 can include one or more substrates upon which the samples are provided.
- sample container 110 can include one or more substrates on which nucleic acids to be sequenced are bound, attached, or associated.
- the substrate can include any inert substrate or matrix to which nucleic acids can be attached, such as for example glass surfaces, plastic surfaces, latex, dextran, polystyrene surfaces, polypropylene surfaces, polyacrylamide gels, gold surfaces, and silicon wafers.
- the substrate is within a channel or other area at a plurality of locations formed in a matrix or array across the sample container 110.
- the sample container 110 can include a biological sample that is imaged using one or more fluorescent dyes.
- the sample container 110 can be implemented as a patterned flow cell including a translucent cover plate, a substrate, and a liquid sandwiched therebetween, and a biological sample can be located at an inside surface of the translucent cover plate or an inside surface of the substrate.
- the flow cell can include a large number (e.g., thousands, millions, or billions) of wells or regions that are patterned into a defined array (e.g, a hexagonal array, rectangular array, etc.) into the substrate. Each region can form a cluster (e.g, a monoclonal cluster) of a biological sample such as DNA, RNA, or another genomic material which can be sequenced, for example, using sequencing by synthesis.
- the flow cell can be further divided into a number of spaced apart lanes (e.g, eight lanes), each lane including a hexagonal array of clusters.
- Example flow cells that can be used in implementations disclosed herein are described in U.S. Pat. No. 8,778,848.
- the system also comprises temperature station actuator 130 and heater/cooler 135 that can optionally regulate the temperature of conditions of the fluids within the sample container 110.
- Camera system 140 can be included to monitor and track the sequencing of sample container 110.
- Camera system 140 can be implemented, for example, as a charge-coupled device (CCD) camera (e.g, a time delay integration (TDI) CCD camera), which can interact with various filters within filter switching assembly 145, objective lens 142, and focusing laser/focusing laser assembly 150.
- CCD charge-coupled device
- TDI time delay integration
- Camera system 140 is not limited to a CCD camera and other cameras and image sensor technologies can be used. In particular implementations, the camera sensor can have a pixel size between about 5 and about 15 pm.
- Output data from the sensors of camera system 140 can be communicated to a real time analysis module (not shown) that can be implemented as a software application that analyzes the image data (e.g ., image quality scoring), reports or displays the characteristics of the laser beam (e.g., focus, shape, intensity, power, brightness, position) to a graphical user interface (GUI), and, as further described below, dynamically corrects intensity noise in the image data.
- a real time analysis module e.g., an excitation laser within an assembly optionally comprising multiple lasers
- other light source can be included to illuminate fluorescent sequencing reactions within the samples via illumination through a fiber optic interface (which can optionally comprise one or more re-imaging lenses, a fiber optic mounting, etc.).
- Low watt lamp 165, focusing laser 150, and reverse dichroic 185 are also presented in the example shown.
- focusing laser 150 can be turned off during imaging.
- an alternative focus configuration can include a second focusing camera (not shown), which can be a quadrant detector, a Position Sensitive Detector (PSD), or similar detector to measure the location of the scattered beam reflected from the surface concurrent with data collection.
- PSD Position Sensitive Detector
- sample container 110 can be ultimately mounted on a sample stage 170 to provide movement and alignment of the sample container 110 relative to the objective lens 142.
- the sample stage can have one or more actuators to allow it to move in any of three dimensions.
- actuators can be provided to allow the stage to move in the X, Y, and Z directions relative to the objective lens. This can allow one or more sample locations on sample container 110 to be positioned in optical alignment with objective lens 142.
- Focus component 175 is shown in this example as being included to control positioning of the optical components relative to the sample container 110 in the focus direction (typically referred to as the z axis, or z direction).
- Focus component 175 can include one or more actuators physically coupled to the optical stage or the sample stage, or both, to move sample container 110 on sample stage 170 relative to the optical components (e.g, the objective lens 142) to provide proper focusing for the imaging operation.
- the actuator can be physically coupled to the respective stage such as, for example, by mechanical, magnetic, fluidic, or other attachment or contact directly or indirectly to or with the stage.
- the one or more actuators can be configured to move the stage in the z-direction while maintaining the sample stage in the same plane (e.g ., maintaining a level or horizontal attitude, perpendicular to the optical axis).
- the one or more actuators can also be configured to tilt the stage. This can be done, for example, so that sample container 110 can be leveled dynamically to account for any slope in its surfaces.
- Focusing of the system generally refers to aligning the focal plane of the objective lens with the sample to be imaged at the chosen sample location. However, focusing can also refer to adjustments to the system to obtain a desired characteristic for a representation of the sample such as, for example, a desired level of sharpness or contrast for an image of a test sample. Because the usable depth of field of the focal plane of the objective lens can be small (sometimes on the order of 1 pm or less), focus component 175 closely follows the surface being imaged. Because the sample container is not perfectly flat as fixtured in the instrument, focus component 175 can be set up to follow this profile while moving along in the scanning direction (herein referred to as the y-axis).
- the light emanating from a test sample at a sample location being imaged can be directed to one or more detectors of camera system 140.
- An aperture can be included and positioned to allow only light emanating from the focus area to pass to the detector.
- the aperture can be included to improve image quality by filtering out components of the light that emanate from areas that are outside of the focus area.
- Emission filters can be included in filter switching assembly 145, which can be selected to record a determined emission wavelength and to cut out any stray laser light.
- a controller can be provided to control the operation of the scanning system.
- the controller can be implemented to control aspects of system operation such as, for example, focusing, stage movement, and imaging operations.
- the controller can be implemented using hardware, algorithms (e.g., machine executable instructions), or a combination of the foregoing.
- the controller can include one or more CPUs or processors with associated memory.
- the controller can comprise hardware or other circuitry to control the operation, such as a computer processor and a non-transitory computer readable medium with machine-readable instructions stored thereon.
- this circuitry can include one or more of the following: field programmable gate array (FPGA), application specific integrated circuit (ASIC), programmable logic device (PLD), complex programmable logic device (CPLD), a programmable logic array (PLA), programmable array logic (PAL) or other similar processing device or circuitry.
- FPGA field programmable gate array
- ASIC application specific integrated circuit
- PLD programmable logic device
- CPLD complex programmable logic device
- PLA programmable logic array
- PAL programmable array logic
- the controller can comprise a combination of this circuitry with one or more processors.
- Figure 2 is a block diagram illustrating an example two-channel, line-scanning modular optical imaging system 200 that can be implemented in particular implementations. It should be noted that although systems and methods can be described herein from time to time in the context of example imaging system 200, these are only examples with which implementations of the technology disclosed herein can be implemented.
- system 200 can be used for the sequencing of nucleic acids. Applicable techniques include those where nucleic acids are attached at fixed locations in an array (e.g ., the wells of a flow cell) and the array is imaged repeatedly. In such implementations, system 200 can obtain images in two different color channels, which can be used to distinguish a particular nucleotide base type from another. More particularly, system 200 can implement a process referred to as “base calling,” which generally refers to a process of a determining a base call (e.g., adenine (A), cytosine (C), guanine (G), or thymine (T)) for a given spot location of an image at an imaging cycle.
- base calling generally refers to a process of a determining a base call (e.g., adenine (A), cytosine (C), guanine (G), or thymine (T)) for a given spot location of an image at an imaging cycle.
- image data extracted from two images can be used to determine the presence of one of four base types by encoding base identity as a combination of the intensities of the two images. For a given spot or location in each of the two images, base identity can be determined based on whether the combination of signal identities is [on, on], [on, off], [off, on], or [off, off]
- the system includes a line generation module (LGM) 210 with two light sources, 211 and 212, disposed therein.
- Light sources 211 and 212 can be coherent light sources such as laser diodes which output laser beams.
- Light source 211 can emit light in a first wavelength (e.g, a red color wavelength), and light source 212 can emit light in a second wavelength (e.g, a green color wavelength).
- the light beams output from laser sources 211 and 212 can be directed through a beam shaping lens or lenses 213.
- a single light shaping lens can be used to shape the light beams output from both light sources.
- a separate beam shaping lens can be used for each light beam.
- the beam shaping lens is a Powell lens, such that the light beams are shaped into line patterns.
- the beam shaping lenses of LGM 210 or other optical components imaging system 200 be configured to shape the light emitted by light sources 211 and 212 into a line patterns (e.g, by using one or more Powel lenses, or other beam shaping lenses, diffractive or scattering components).
- LGM 210 can further include mirror 214 and semi-reflective mirror 215 configured to direct the light beams through a single interface port to an emission optics module (EOM) 230.
- EOM emission optics module
- the light beams can pass through a shutter element 216.
- EOM 230 can include objective 235 and a z-stage 236 which moves objective 235 longitudinally closer to or further away from a target 250.
- target 250 can include a liquid layer 252 and a translucent cover plate 251, and a biological sample can be located at an inside surface of the translucent cover plate as well an inside surface of the substrate layer located below the liquid layer.
- the z-stage can then move the objective as to focus the light beams onto either inside surface of the flow cell ( e.g ., focused on the biological sample).
- the biological sample can be DNA, RNA, proteins, or other biological materials responsive to optical sequencing as known in the art.
- EOM 230 can include semi-reflective mirror 233 to reflect a focus tracking light beam emitted from a focus tracking module (FTM) 240 onto target 250, and then to reflect light returned from target 250 back into FTM 240.
- FTM 240 can include a focus tracking optical sensor to detect characteristics of the returned focus tracking light beam and generate a feedback signal to optimize focus of objective 235 on target 250.
- EOM 230 can also include semi -reflective mirror 234 to direct light through objective 235, while allowing light returned from target 250 to pass through.
- EOM 230 can include a tube lens 232.
- Light transmitted through tube lens 232 can pass through filter element 231 and into camera module (CAM) 220.
- CAM 220 can include one or more optical sensors 221 to detect light emitted from the biological sample in response to the incident light beams (e.g., fluorescence in response to red and green light received from light sources 211 and 212).
- Output data from the sensors of CAM 220 can be communicated to a real time analysis module 225.
- Real time analysis module executes computer readable instructions for analyzing the image data (e.g, image quality scoring, base calling, etc.), reporting or displaying the characteristics of the beam (e.g, focus, shape, intensity, power, brightness, position) to a graphical user interface (GUI), etc. These operations can be performed in real-time during imaging cycles to minimize downstream analysis time and provide real time feedback and troubleshooting during an imaging run.
- real time analysis module can be a computing device (e.g, computing device 1000) that is communicatively coupled to and controls imaging system 200.
- real time analysis module 225 can additionally execute computer readable instructions for maximizing the signal-to-noise ratio of the output image data received from CAM 220.
- each of the sequencing images has one or more image (or intensity) channels (analogous to the red, green, blue (RGB) channels of a color image).
- each image channel corresponds to one of a plurality of filter wavelength bands.
- each image channel corresponds to one of a plurality of imaging events at a sequencing cycle.
- each image channel corresponds to a combination of illumination with a specific laser and imaging through a specific optical filter.
- the image patches are tiled (or accessed) from each of the m image channels for a particular sequencing cycle.
- m is 4 or 2.
- m is 1, 3, or greater than 4.
- the images can be in blue and violet color channels instead of or in addition to the red and green channels.
- Figure 3 shows one implementation of respective/separate/different/independent specialist signal profilers 1 to N trained to maximize signal-to-noise ratio of respective image data subsets 1 to N generated for respective classes 1 to N of clusters located on a flow cell in respective spatial configurations 1 to N.
- Sequencing run 300 sequences a population of clusters over a plurality of sequencing cycles/imaging cycles 1 to K and generates image data that depicts intensity emissions of the population of clusters.
- imaging cycle of the flow cell is performed such that the image data can be collected for the entire flow cell by scanning the flow cell area (e.g ., using a line scanner), with one or more coherent sources of light.
- the imaging system 200 can use LGM 210 in coordination with the optics of the system to line scan the flow cell with light having wavelengths within the red color spectrum and to line scan the sample with light having wavelengths within the green color spectrum.
- fluorescent dyes situated at the different clusters of the flow cell fluoresce and the resultant light can be collected by the objective lens 235 and directed to an image sensor of CAM 220 to detect the florescence.
- fluorescence of each cluster can be detected by a few pixels of CAM 220.
- Image data output from CAM 220 can then be communicated to the real time analysis module 225 for image noise correction and base calling.
- the population of clusters is spatially distributed across the flow cell.
- the flow cell and therefore the population of clusters and the image data, is partitionable into different spatial configurations defined by different regions of the flow cell. So, for example, if the flow cell is partitionable into three rectangular regions, this would result in three spatial configurations of the flow cell, three sub-populations or classes of the clusters, three subsets of the image data, and three specialist signal profilers.
- each of the three rectangular regions of the flow cell can be further divided into three squares, resulting in a total of nine squares.
- the image data is divided into a plurality of image data subsets corresponding to a respective region of the flow cell.
- the size of the image data subsets can be determined using the placement of fiducial markers or fiducials in the field of view of the imaging system 200, in the flow cell, or on the flow cell.
- the image data subsets can be divided such that the pixels of each image data subset have a predetermined number of fiducials (e.g ., at least three fiducials, four fiducials, six fiducials, eight fiducials, etc.).
- the total number of pixels of the image data subset may be predetermined based on predetermined pixel distances between the boundaries of the image data subset and the fiducials.
- intensity data in each subset of image data can span across multiple color channels (e.g., red, blue, green, and/or violet image channels). Accordingly, each specialist signal profiler has a respective set of coefficients trained to maximize the signal-to-noise ratio of intensity data in a corresponding color channel.
- new/unseen/wild image data is segmented on the same basis as the spatial configurations defined and used to generate the respective/separate/different/independent specialist signal profilers at training.
- the respective/separate/different/independent specialist signal profilers are applied to segmented new image data subsets in a way that the application of a particular specialist signal profiler is confined only to a corresponding new image data subset.
- the application of the specialist signal profilers 1 to N to the image data subsets 1 to N generates signal-to-noise ratio maximized image data subsets 1 to N, additional details about which can be found in U.S.
- Other corrections for channel, spatial, or phasing noise can be applied in advance to the unprocessed image data, or subsequently to the signal-to-noise ratio maximized image data subsets 1 to N.
- the output of the specialist signal profilers 1 to N i.e., the signal-to-noise ratio maximized image data subsets 1 to N, are provided as input to a base caller 332 to produce base calls 1 to N for the population of clusters.
- Base calling may be performed by fitting a mathematical model to the intensity data. Suitable mathematical models that can be used include, for example, a k-means clustering algorithm, a k-means-like clustering algorithm, expectation maximization clustering algorithm, a histogram based method, and the like.
- Four Gaussian distributions may be fit to the set of two-channel intensity data such that one distribution is applied for each of the four nucleotides represented in the data set.
- an expectation maximization (EM) algorithm may be applied.
- EM expectation maximization
- a value can be generated which represents the likelihood that a certain X, Y intensity value belongs to one of four Gaussian distributions to which the data is fitted.
- each X, Y intensity value will also have four associated likelihood values, one for each of the four bases.
- the maximum of the four likelihood values indicates the base call. For example, if a cluster is “off’ in both channels, the base call is G. If the cluster is “off’ in one channel and “on” in another channel the base call is either C or T (depending on which channel is on), and if the cluster is “on” in both channels the base call is A.
- Figure 4 illustrates an example configuration of a flow cell 400 that can be imaged in accordance with implementations disclosed herein.
- the flow cell 400 is patterned with a hexagonal array of ordered clusters or spots or features that can be simultaneously imaged during an imaging run.
- the flow cell 400 can be patterned using a rectilinear array, a circular array, an octagonal array, or some other array pattern.
- the flow cell 400 can have tens, hundreds, thousands, millions, or billions of clusters that are imaged.
- the flow cell 400 can be patterned with millions or billions of wells that are divided into lanes.
- each well of the flow cell 400 can contain at least one cluster.
- the flow cell 400 can be a multi-plane sample comprising multiple planes of clusters that are sampled during an imaging run.
- the multiple planes of clusters include a top surface 402 and a bottom surface 412.
- the flow cell 400 can have two spatial configurations corresponding to the top and bottom surfaces 402 and 412, which in turn result in two sub-populations or classes of the clusters, two subsets of the image data, and two specialist signal profilers.
- Figure 6A shows one implementation of training respective surface-specific specialist signal profilers 604 for respective cluster classes 602.
- the respective cluster classes 602 include groups of clusters respectively located on the top surface 402, and clusters located on the bottom surface 412. What results is a first specialist signal profiler 604a configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the top surface 402, and a second specialist signal profiler 604b configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the bottom surface 412.
- Figure 6B shows one implementation of applying the trained surface-specific specialist signal profilers 604 to image data subsets 632, 642 corresponding to the respective cluster classes 602 during a sequencing run 600.
- the flow cell 400 is imaged at the tile-level. So, assuming that the top surface 402 has a first set of 1600 tiles and the bottom surface 412 has a second set of 1600 tiles, a first image data subset 632 includes a first set of 1600 tile images of the first set of 1600 tiles, and a second image data subset 642 includes a second set of 1600 tile images of the second set of 1600 tiles.
- the first specialist signal profiler 604a is configured to maximize the signal-to-noise ratio of intensity data in the first image data subset 632 to generate a signal-to-noise ratio maximized version 634 of the first image data subset 632.
- the second specialist signal profiler 604b is configured to maximize the signal-to-noise ratio of intensity data in the second image data subset 642 to generate a signal-to-noise ratio maximized version 644 of the second image data subset 642.
- the base caller 332 processes the signal-to-noise ratio maximized versions 634, 644 and generates base calls 638, 648.
- the term “maximized version” refers to data produced as output by a specialist signal profiler.
- the maximized version of an input is an output produced by a specialist signal profiler in response to processing a corresponding input.
- the maximized version of the input has a signal-to-noise ratio (SNR, S/R) that is greater than the corresponding input processed by the specialist signal profiler.
- SNR, S/R signal-to-noise ratio
- the maximized version of the input (i.e., the output) produced by the specialist signal profiler is corrected by the specialist signal profiler to have less spatial crosstalk and background noise compared to the corresponding input.
- the maximized version of the input (i.e., the output) produced by the specialist signal profiler is corrected by the specialist signal profiler to have less phasing and pre-phasing noise compared to the corresponding input.
- the top surface 402 can be divided or partitioned into a plurality of lanes 508a, 508b, ..., 5081.
- the top surface 402 has eight lanes, although the number of lanes is implementation specific.
- the flow cell 400 can have sixteen spatial configurations corresponding to eight lanes of the top surface 402 and eight lanes of the bottom surface 412, which in turn result in sixteen sub-populations or classes of the clusters, sixteen subsets of the image data, and sixteen specialist signal profilers.
- Figure 7B shows one implementation of training respective lane-specific specialist signal profilers 714 for respective cluster classes 712.
- the respective cluster classes 712 include groups of clusters respectively located on the lanes 508a, 508b, ..., 5081.
- a first specialist signal profiler 714a configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the first lane 508a
- a second specialist signal profiler 714b configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the second lane 508b
- so on continuing to the 1 th specialist signal profiler 7141 configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the 1 th lane 5081).
- Figure 9 shows one implementation of applying the trained lane-specific specialist signal profilers 714 to image data subsets 902, 912, ..., 922 corresponding to the respective cluster classes 712 during a sequencing run 900.
- the flow cell 400 is imaged at the tile-level. So, assuming that each lane has 200 tiles, a first image data subset 902 includes a first set of 200 tile images of the 200 tiles on the first lane 508a, a second image data subset 912 includes a second set of 200 tile images of the 200 tiles on the second lane 508b, and so on (continuing to the 1 th image data subset 922).
- the first specialist signal profiler 714a is configured to maximize the signal-to-noise ratio of intensity data in the first image data subset 902 to generate a signal-to-noise ratio maximized version 904 of the first image data subset 902.
- the second specialist signal profiler 714b is configured to maximize the signal-to-noise ratio of intensity data in the second image data subset 912 to generate a signal-to-noise ratio maximized version 914 of the second image data subset 912, and so on (continuing to signal-to-noise ratio maximized version 924 of the 1 th image data subset 922).
- the base caller 332 processes the signal-to-noise ratio maximized versions 904, 914, ..., 924 and generates base calls 908, 918, ..., 928.
- the lanes can be grouped into lane groups 502a, 502b, and 502c.
- the lane groups include top peripheral lanes, central lanes, and bottom peripheral lanes.
- the flow cell 400 can have six spatial configurations corresponding to three lane groups of the top surface 402 and three lane groups of the bottom surface 412, which in turn result in six sub-populations or classes of the clusters, six subsets of the image data, and six specialist signal profilers.
- Figure 7A shows one implementation of training respective lane group-specific specialist signal profilers 704 for respective cluster classes 702.
- the respective cluster classes 702 include groups of clusters respectively located on the lane groups 502a, 502b, 502c. What results is a first specialist signal profiler 704a configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the first lane group 502a, a second specialist signal profiler 704b configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the second lane group 502b, and a third specialist signal profiler 704c configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the third lane group 502c.
- Figure 8 shows one implementation of applying the trained lane group-specific specialist signal profilers 704 to image data subsets 802, 812, 822 corresponding to the respective cluster classes 702 during a sequencing run 800.
- the flow cell 400 is imaged at the tile-level.
- a first image data subset 802 includes a first set of 600 tile images of the 600 tiles on the first lane group 502a
- a second image data subset 812 includes a second set of 600 tile images of the 600 tiles on the second lane group 502b
- a third image data subset 822 includes a third set of 400 tile images of the 400 tiles on the third lane group 502c.
- the first specialist signal profiler 704a is configured to maximize the signal-to-noise ratio of intensity data in the first image data subset 802 to generate a signal-to-noise ratio maximized version 804 of the first image data subset 802.
- the second specialist signal profiler 704b is configured to maximize the signal-to-noise ratio of intensity data in the second image data subset 812 to generate a signal-to-noise ratio maximized version 814 of the second image data subset 812.
- the third specialist signal profiler 704c is configured to maximize the signal-to- noise ratio of intensity data in the third image data subset 822 to generate a signal-to-noise ratio maximized version 824 of the third image data subset 822.
- the base caller 332 processes the signal-to-noise ratio maximized versions 804, 814, 824 and generates base calls 808, 818, 828.
- each lane comprises one or more columns/ swathes 506a, 506b of tiles, as shown in Figure 5B.
- the flow cell 400 can have thirty -two spatial configurations corresponding to thirty -two swathes of tiles of the top surface 402 and thirty -two swathes of tiles of the bottom surface 412, which in turn result in thirty-two sub-populations or classes of the clusters, thirty-two subsets of the image data, and thirty-two specialist signal profilers.
- Figure 7C shows one implementation of training respective swath-specific specialist signal profilers 724 for respective cluster classes 722.
- the respective cluster classes 722 include groups of clusters respectively located on the swathes 506a, 506b. What results is a first specialist signal profiler 724a configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the first swath 506a, and a second specialist signal profiler 724b configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the second swath 506b.
- Figure 10 shows one implementation of applying the trained swath-specific specialist signal profilers 724 to image data subsets 1002, 1012 corresponding to the respective cluster classes 722 during a sequencing run 1000.
- the flow cell 400 is imaged at the tile-level. So, assuming that the first swath 506a has 100 tiles, and the second swath 506b has 100 tiles, a first image data subset 1002 includes a first set of 100 tile images of the 100 tiles on the first swath 506a, and a second image data subset 1012 includes a second set of 100 tile images of the 100 tiles on the second swath 506b.
- the first specialist signal profiler 724a is configured to maximize the signal-to-noise ratio of intensity data in the first image data subset 1002 to generate a signal-to-noise ratio maximized version 1004 of the first image data subset 1002.
- the second specialist signal profiler 724b is configured to maximize the signal-to-noise ratio of intensity data in the second image data subset 1012 to generate a signal-to-noise ratio maximized version 1014 of the second image data subset 1012.
- the base caller 332 processes the signal-to-noise ratio maximized versions 1004, 1014, 1024 and generates base calls 1008, 1018.
- Each swath comprises a plurality of tiles 512a, 512b, ..., 512t, as shown in Figure 5B.
- the number of tiles within each swath is implementation specific, and in different examples, there can be 50 tiles, 60 tiles, 80 tiles, and so on.
- each swath has 100 tiles.
- the flow cell 400 will have 200 tiles per lane, resulting in 1600 tiles for the top surface 402 and another 1600 tiles for the bottom surface 412, /. e. , a total of 3200 tiles.
- the flow cell 400 can have 3200 spatial configurations corresponding to the 3200 tiles, which in turn result in 3200 sub-populations or classes of the clusters, 3200 subsets of the image data, and 3200 specialist signal profilers.
- Figure 7D shows one implementation of training respective tile-specific specialist signal profilers 734 for respective cluster classes 732.
- the respective cluster classes 732 include groups of clusters respectively located on the tiles 512a, 512b, ..., 512t.
- a first specialist signal profiler 734a configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the first tile 512a
- a second specialist signal profiler 734b configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the second tile 512b
- so on continuing to the t th specialist signal profiler 734t configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the t th tile 512t).
- Figure 11 shows one implementation of applying the trained tile-specific specialist signal profilers 734 to image data subsets 1102, 1112, ... , 1122 corresponding to the respective cluster classes 732 during a sequencing run 1100.
- the flow cell 400 is imaged at the tile-level. So, for the tiles 512a, 512b, ..., 512t, a first image data subset 1102 includes a first tile image of the first tile 512a, a second image data subset 1112 includes a second tile image of the second tile 512b, and so on (continuing to the t* 11 image data subset 1122 comprising the t th tile image of the t th tile 512t (as shown in Figure 5C)).
- the first specialist signal profiler 734a is configured to maximize the signal-to-noise ratio of intensity data in the first image data subset 1102 to generate a signal-to-noise ratio maximized version 1104 of the first image data subset 1102.
- the second specialist signal profiler 734b is configured to maximize the signal-to-noise ratio of intensity data in the second image data subset 1112 to generate a signal-to-noise ratio maximized version 1114 of the second image data subset 1112, and so on (continuing to signal-to-noise ratio maximized version 1124 of the t th image data subset 1122).
- the base caller 332 processes the signal-to-noise ratio maximized versions 1104, 1114, ..., 1124 and generates base calls 1108, 1118, ..., 1128.
- Each tile can be divided into a plurality of sub-tiles 518a, 518b, ... , 518s, as shown in Figure 5D.
- the number of sub-tiles into which a tile can be divided is implementation specific, and in different examples, there can be 10 sub-tiles, 30 sub-tiles, 50 sub-tiles, and so on.
- each tile is divided into 9 sub-tiles.
- the flow cell 400 will have a total of 28,800 sub-tiles for 3200 tiles. Accordingly, in one implementation, the flow cell 400 can have 28,800 spatial configurations corresponding to the 28,800 tiles, which in turn result in 28,800 sub-populations or classes of the clusters, 28,800 subsets of the image data, and 28,800 specialist signal profilers.
- Figure 7E shows one implementation of training respective sub-tile-specific specialist signal profilers 744 for respective cluster classes 742.
- the respective cluster classes 742 include groups of clusters respectively located on the sub-tiles 518a, 518b, 518c, 518d, ..., 518s.
- first specialist signal profiler 744a configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the first sub-tile 518a
- second specialist signal profiler 744b configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the second sub-tile 518b
- third specialist signal profiler 744c configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the third sub-tile 518c
- a fourth specialist signal profiler 744d configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the fourth sub-tile 518d, and so on (continuing to the s th specialist signal profiler 744s configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the s th sub-tile 518s).
- Figure 12 shows one implementation of applying the trained sub-tile-specific specialist signal profilers 744 to image data subsets 1202, 1212, 1222, 1232, ..., 1242 corresponding to the respective cluster classes 742 during a sequencing run 1200.
- the flow cell 400 is imaged at the sub-tile-level.
- a first image data subset 1202 includes a first sub-tile image patch of the first sub-tile 518a
- a second image data subset 1212 includes a second sub-tile image patch of the second sub-tile 518b
- a third image data subset 1222 includes a third sub-tile image patch of the third sub-tile 518c
- a fourth image data subset 1232 includes a fourth sub-tile image patch of the fourth sub-tile 518d, and so on (continuing to the s th image data subset 1242 comprising the s th sub-tile image patch of the s th sub-tile 518s).
- the first specialist signal profiler 744a is configured to maximize the signal-to-noise ratio of intensity data in the first image data subset 1202 to generate a signal-to-noise ratio maximized version 1204 of the first image data subset 1202.
- the second specialist signal profiler 744b is configured to maximize the signal-to-noise ratio of intensity data in the second image data subset 1212 to generate a signal-to-noise ratio maximized version 1214 of the second image data subset 1212.
- the third specialist signal profiler 744c is configured to maximize the signal- to-noise ratio of intensity data in the third image data subset 1222 to generate a signal-to-noise ratio maximized version 1224 of the third image data subset 1222.
- the fourth specialist signal profiler 744d is configured to maximize the signal-to-noise ratio of intensity data in the fourth image data subset 1232 to generate a signal-to-noise ratio maximized version 1234 of the fourth image data subset 1232, and so on (continuing to signal-to-noise ratio maximized version 1244 of the s * image data subset 1242).
- the base caller 332 processes the signal-to-noise ratio maximized versions 1204, 1214, 1224, 1234, ..., 1244 and generates base calls 1208, 1218,
- Figure 13 shows one implementation of respective/separate/different/independent specialist signal profilers for respective sub-series of sequencing cycles of a sequencing run 1300 with a total of N sequencing cycles.
- the first specialist signal profiler 1312 is configured to maximize the signal-to-noise ratio of intensity data of clusters located on sub-tile M and generated during sequencing cycles 1 to Nl.
- the second specialist signal profiler 1314 is configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the sub-tile M and generated during sequencing cycles Nl + 1 to N2.
- the third specialist signal profiler 1318 is configured to maximize the signal-to-noise ratio of intensity data of the clusters located on the sub-tile M and generated during sequencing cycles N2 + 1 to N.
- Figure 14 shows one implementation of respective/separate/different/independent specialist signal profilers for a combination of different spatial configurations (e.g, different sub tiles) and different temporal configurations (e.g, different sub-series of sequencing cycles).
- the cluster classes 1410 are defined by the different spatial configurations (e.g, different sub-tiles).
- the cluster sub-classes 1412, 1414, and 14 are defined by different temporal configurations (e.g, different sub-series of sequencing cycles).
- the first specialist signal profiler 1422 is configured to maximize the signal -to-noise ratio of intensity data of the clusters located on the sub-tile 518a and generated during the sequencing cycles 1 to Nl.
- the second specialist signal profiler 1424 is configured to maximize the signal -to-noise ratio of intensity data of the clusters located on the sub-tile 518a and generated during the sequencing cycles Nl + 1 to N2.
- the third specialist signal profiler 1428 is configured to maximize the signal -to-noise ratio of intensity data of the clusters located on the sub-tile 518a and generated during the sequencing cycles N2 + 1 to N.
- the fourth specialist signal profiler 1432 is configured to maximize the signal-to- noise ratio of intensity data of the clusters located on the sub-tile 518b and generated during the sequencing cycles 1 to Nl.
- the fifth specialist signal profiler 1434 is configured to maximize the signal -to-noise ratio of intensity data of the clusters located on the sub-tile 518b and generated during the sequencing cycles Nl + 1 to N2.
- the sixth specialist signal profiler 1438 is configured to maximize the signal -to-noise ratio of intensity data of the clusters located on the sub-tile 518b and generated during the sequencing cycles N2 + 1 to N.
- Figure 15 shows one implementation of respective/separate/different/independent specialist signal profilers for each cluster/well sequenced during a sequencing run.
- the clusters/wells on a flow cell can be pre-identified by their location coordinates. These location coordinates of the clusters/wells can be used to train per-cluster/well specialist signal profilers during training on training per-cluster/well intensity data, and to apply the trained per-cluster specialist signal profilers during inference on inference per-cluster/well intensity data based on the location coordinates of the clusters/wells.
- the cluster/well population 1502 has N clusters/wells, and therefore the specialist signal profilers 1508 comprise respective N per- cluster/well specialist signal profilers.
- different per-cluster/well signal profilers can be trained for different temporal configurations as well, as discussed above. Offline Training
- Figure 16 shows one implementation of offline training of specialist signal profilers on sequenced data from one or more completed/already-executed sequencing runs, and application of the trained specialist signal profilers on sequenced data from an ongoing sequencing run.
- Training data 1612 is generated during a training stage 1602.
- the training data 1612 includes the sequenced data from the one or more completed/already-executed sequencing runs.
- Segmentation logic 1622 segments the training data 1612 based on one or more configurations selected from different spatial configurations, temporal configurations, signal distribution configurations, or any combination thereof. What results is segmented training data 1632 with configuration-specific training data subsets 1 to N.
- the training data 1612 can include K images of a tile from K imaging cycles of a completed/already-executed sequencing run, with each tile image having multiple color channels.
- Figure 17 shows an example of a tile image with C color channels.
- the segmentation logic 1622 logically partitions each tile image in the training data 1612 into sub-tile images by specifying pixel ranges. For example, a first sub-tile image of a tile can range from pixels 1 to 500, a second sub tile image of the tile can range from pixels 501 to 1000, and so on.
- the pixel ranges can be defined using fiducial markers, as discussed above.
- Offline training logic 1642 trains respective/separate/different/independent specialist signal profilers 1 to N on the respective configuration-specific training data subsets 1 to N. What results is trained specialist signal profilers 1 to N.
- a corresponding specialist signal profiler is trained to maximize the signal-to-noise ratio of each sub-tile image in the training data 1612.
- Inference data 1618 is generated during an inference stage 1608.
- the inference data 1618 includes the sequenced data from the ongoing sequencing run (e.g ., the first i cycles of the ongoing sequencing run).
- the segmentation logic 1622 segments the inference data 1618 based on the same one or more configurations that were used to segment the training data 1612 during the training stage 1602. What results is segmented inference data 1638 with configuration-specific inference data subsets 1 to N.
- the inference data 1618 can include K images of a tile from K imaging cycles of the ongoing sequencing run, with each tile image having multiple color channels.
- the segmentation logic 1622 logically partitions each tile image in the inference data 1618 into sub-tile images by specifying the same pixel ranges used to partition the training data 1612.
- Runtime logic 1648 applies the respective trained specialist signal profilers 1 to N 1658 on the respective configuration-specific inference data subsets 1 to N.
- a corresponding trained specialist signal profiler is applied to maximize the signal -to-noise ratio of each sub-tile image in the segmented inference data 1638.
- Figure 18 shows one implementation of online training of specialist signal profilers on sequenced data from earlier sequencing cycles of an ongoing sequencing run, and application of the trained specialist signal profilers on sequenced data from later sequencing cycles of the ongoing sequencing run.
- Inference data 1802 is generated during an inference stage 1802.
- the inference data 1812 includes the sequenced data from the earlier sequencing cycles ( e.g ., cycles 1 to Nl) of the ongoing sequencing run.
- Segmentation logic 1622 segments the inference data 1812 based on one or more configurations selected from different spatial configurations, temporal configurations, signal distribution configurations, or any combination thereof. What results is segmented inference data 1832 with configuration-specific training data subsets 1 to N.
- the inference data 1812 can include Nl images of a tile from the Nl earlier sequencing cycles of the ongoing sequencing run, with each tile image having multiple color channels.
- Figure 17 shows an example of a tile image with C color channels.
- the segmentation logic 1622 logically partitions each tile image in the inference data 1812 into sub-tile images by specifying pixel ranges. For example, a first sub-tile image of a tile can range from pixels 1 to 500, a second sub tile image of the tile can range from pixels 501 to 1000, and so on.
- the pixel ranges can be defined using fiducial markers, as discussed above.
- Offline training logic 1842 trains respective/separate/different/independent specialist signal profilers 1 to N on the respective configuration-specific training data subsets 1 to N. What results is trained specialist signal profilers 1 to N.
- a corresponding specialist signal profiler is trained to maximize the signal-to-noise ratio of each sub-tile image in the inference data 1812.
- Inference data 1818 is also generated during the inference stage 1802.
- the inference data 1818 includes the sequenced data from the later sequencing cycles (e.g., cycles Nl+1 to N2) of the ongoing sequencing run.
- the segmentation logic 1622 segments the inference data 1818 based on the same one or more configurations that were used to segment the inference data 1812 during the earlier sequencing cycles (e.g, cycles 1 to Nl) of the ongoing sequencing run. What results is segmented inference data 1838 with configuration-specific inference data subsets 1 toN.
- the inference data 1818 can include N2 images of a tile from the N2 later sequencing cycles of the ongoing sequencing run, with each tile image having multiple color channels.
- the segmentation logic 1622 logically partitions each tile image in the inference data 1818 into sub-tile images by specifying the same pixel ranges used to partition the inference data 1812.
- Runtime logic 1648 applies the respective trained specialist signal profilers 1 to N 1858 on the respective configuration-specific inference data subsets 1 to N.
- a corresponding trained specialist signal profiler is applied to maximize the signal -to-noise ratio of each sub-tile image in the segmented inference data 1838.
- the training process is iteratively repeated where the respective trained specialist signal profilers 1 to N 1858 are retrained/further trained on the segmented inference data 1838 from the later sequencing cycles (e.g ., cycles Nl+1 to N2) of the ongoing sequencing run, and applied on segmented inference data from even later sequencing cycles (e.g., cycles N2+1 to N3) of the ongoing sequencing run.
- the later sequencing cycles e.g ., cycles Nl+1 to N2
- cycles N2+1 to N3 e.g., cycles N2+1 to N3
- a control logic can iterate, at each successive sequencing cycle of the ongoing sequencing run, or at successive subseries of sequencing cycles of the ongoing sequencing run (e.g, every ten or twenty sequencing cycles) — (i) the segmentation of a current batch of image data based on one or more configurations selected from different spatial configurations, temporal configurations, signal distribution configurations, or any combination thereof, (ii) the retraining of the respective trained specialist signal profilers 1 to N on the segmented current batch of image data, (iii) segmentation of a next batch of image data on the same basis as the current batch of image data, and (iv) the application of the respective retrained specialist signal profilers 1 to N to the segmented next batch of image data.
- Figure 19 shows one implementation of training respective/separate/different/independent specialist signal profilers for respective signal distributions observed in sequenced data.
- the respective signal distributions can be observed in sequenced data from an offline/already-executed sequencing run.
- the respective signal distributions can be also observed in sequenced data from an online/ongoing sequencing run (e.g, observed in the first ten sequencing cycles of the ongoing sequencing run).
- Figure 20 shows one example of a signal distribution/signal profile/cluster intensity profile.
- the cluster intensity profile depicted in Figure 20 follows an attenuation pattern in which the cluster signal is strongest at a cluster center and attenuates as it propagates away from the cluster center.
- Sub-populations/groups/sets of clusters 1904 in a cluster population can have similar signal distributions.
- Those clusters that share same or similar signal distributions can be bucketed together ( e.g ., by grouping/addressing clusters by their location coordinates), such that a specialist signal profiler can be trained for a corresponding cluster group/set exhibiting a corresponding signal distribution.
- signal distribution-based grouping can group non-contiguous clusters. For example, edge clusters on opposite edges of a tile can have similar signal distributions, and can be grouped in a what that their intensity data is corrected by a same specialist signal profiler (e.g., by grouping/addressing clusters by their location coordinates).
- similar signal distribution refers to those signal distributions that share substantially overlapping signal patterns.
- two signal patters of similar shape e.g, trapezoids
- different shape sizes e.g. one bigger trapezoid and one smaller trapezoid
- two signal patterns with a common centroid or centroids within a range of, for example, one to five units in each dimension can be considered to have similar signal distributions.
- respective/separate/different/independent specialist signal profilers 1 to N 1908 are trained to maximize the signal -to-noise ratio of respective signal distributions 1 to N 1902 corresponding to respective cluster sets 1 to N 1904.
- different signal distributions can be observed at different sequencing cycles, and therefore different specialist signal profilers can be trained for and configured for application at different temporal stages of an ongoing sequencing run.
- the phrase “different temporal stages of an ongoing sequencing run” refers to different sequencing cycles or different sequencing cycle groups of the ongoing sequencing run. For example, if the sequencing run has 150 sequencing cycles, then each successive sequencing cycle can be considered a different temporal stage, or groups of sequencing cycles like cycles 1-20, cycles 20-40, cycles 40-70, and so on can be considered different temporal stages.
- Figure 21 shows one implementation of a processing pipeline that implements the technology disclosed.
- the processing pipeline can be implemented by the real time analysis module 225.
- the processing pipeline is executed on a cycle-by-cycle basis 2100, and iterated 2102 for each new cycle, according to one implementation.
- the input to the processing pipeline is tile images having a first (green color) channel and a second (blue color) channel.
- a template image is produced that identifies locations of clusters on a tile using sequencing images from some number of initial sequencing cycles called template cycles.
- the template image is used as a reference for subsequent registration and intensity extraction steps.
- the template image is generated by detecting and merging bright spots in each sequencing image of the template cycles, which in turn involves sharpening a sequencing image 2114 ( e.g ., using the Laplacian convolution), determining an “on” threshold by a spatially segregated Otsu approach, and subsequent five-pixel local maximum detection with subpixel location interpolation.
- the phrase “on threshold” can refer to an intensity value that exceeds a preset value, for example, 200 or 320, such that the intensity value is detected to be more than a background intensity value or a noise intensity value.
- the processing pipeline then registers a current sequencing image against the template image. This is achieved by using image correlation to align the current sequencing image to the template image on a sub-region, or by using non-linear transformations (e.g., a full six-parameter linear affine transformation).
- the processing pipeline applies a non-linear distortion to each spot to account for, for example, optical distortions caused by the geometry of the optical lens.
- the non-linear distortion can be applied as channel-dependent coefficients of third order polynomials.
- the processing pipeline segments the tile images into sub-tile images based on one or more configurations selected from different spatial configurations, temporal configurations, signal distribution configurations, or any combination thereof.
- intensity is extracted from the segmented sub-tile images using the corresponding specialist signal profilers 1 to N.
- the sub-tile intensities are spatially normalized, for example, by making the 90 th percentiles of their extracted intensities are equal.
- the processing pipeline applies empirical phasing correction to compensate noise in the image data caused by phasing and pre-phasing errors.
- the processing pipeline spatially normalizes the extracted signal intensities to account for variation in illumination across the sampled imaged. For example, intensity values can be normalized such that a 5th and 95th percentiles have values of 0 and 1, respectively.
- the normalized signal intensities for the image e.g, normalized intensities for each channel
- the processing pipeline scale the intensities on a cluster-by-cluster basis to account for variation in brightness of clusters.
- the processing pipeline uses the expectation maximization (EM) algorithm to produce base calls, as discussed above.
- EM expectation maximization
- the processing pipeline assigns quality scores to the called bases using quality tables (Q-Tables) 2152.
- the processing pipeline aligns the called bases to a reference genome (e.g ., of the PhiX bacteria) to compute a mismatch rate.
- a reference genome e.g ., of the PhiX bacteria
- the processing pipeline generates certain outputs, such as base call and quality score 2128, inter-operation (InterOp) files 2138 (binary reporting files for sequencing analysis view), and logs 2148 (e.g., error logs, general event logs, processing event logs, warning event logs).
- base call and quality score 2128 inter-operation (InterOp) files 2138 (binary reporting files for sequencing analysis view)
- InterOp inter-operation
- logs 2148 e.g., error logs, general event logs, processing event logs, warning event logs.
- Other examples of configurations covered by this disclosure include segmenting sequencing data and training corresponding specialist signal profilers by library type, sample type, indexing type (first index read v/s second index read), read type (forward read v/s reverse read), physical properties of the sample, noise type (e.g, bubble), and reagent type.
- Figure 33 shows how a cost function for a specialist signal profiler improves with each iteration of gradient descent.
- a cost function (or loss function) measures the performance of a model on a given data.
- the different colored lines correspond to a number of sub tiles into which a tile is partitioned, such that for each sub-tile, the specialist signal profiler is adapted/trained/configured/updated independently.
- the cost function is lower.
- the cost function is the sum of square of Euclidean distance between well intensity and base call centroid for a sample of wells (clusters). Improving this cost function improves base calling accuracy as well.
- Figure 34 is a plot that shows initial and final values of the cost function of Figure 33 when the specialist signal profiler is adapted/trained/configured/updated at each sequencing cycle. At each successive sequencing cycle, we started with a same initial specialist signal profiler and adapted the specialist signal profiler using gradient descent. The plot shows that it is possible to adapt/train/configure/update the specialist signal profiler at any sequencing cycle.
- Figures 35 A and 35B show the improvement in primary analysis metrics for a sequencing run when we adapt/train/configure/update the specialist signal profiler. In this case, mean PhiX error rate improved from 0.3520% to 0.3316%.
- Figures 36A and 36B show two plots that assess a number of sub-tiles into which a sequencing tile can be partitioned for adaptive equalization of respective specialist signal profilers.
- a separate specialist signal profiler is adapted/trained/configured/updated using wells from a corresponding sub-tile.
- Using more sub tiles allows us to model spatially varying phenomena within a tile more accurately which is why the error rate and Q30 improve significantly as we increase the number of sub-tiles from 1 to 9.
- the number of wells available to adapt the model decreases as the sub-tile gets smaller. That is why there is a tradeoff in picking the right number of sub-tiles. In this particular case, increasing the number of sub-tiles from 9 to 16 degrades error rate and Q30 slightly.
- Figure 37 shows an example computer system 3700 that can be used to implement the technology disclosed.
- Computer system 3700 includes at least one central processing unit (CPU) 3772 that communicates with a number of peripheral devices via bus subsystem 3755.
- peripheral devices can include a storage subsystem 3710 including, for example, memory devices and a file storage subsystem 3736, user interface input devices 3738, user interface output devices 3776, and a network interface subsystem 3774.
- the input and output devices allow user interaction with computer system 3700.
- Network interface subsystem 3774 provides an interface to outside networks, including an interface to corresponding interface devices in other computer systems.
- specialist signal profilers 3718 are communicably linked to the storage subsystem 3710 and the user interface input devices 3738.
- User interface input devices 3738 can include a keyboard; pointing devices such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touch screen incorporated into the display; audio input devices such as voice recognition systems and microphones; and other types of input devices.
- pointing devices such as a mouse, trackball, touchpad, or graphics tablet
- audio input devices such as voice recognition systems and microphones
- use of the term “input device” is intended to include all possible types of devices and ways to input information into computer system 3700.
- User interface output devices 3776 can include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices.
- the display subsystem can include an LED display, a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image.
- the display subsystem can also provide a non-visual display such as audio output devices.
- output device is intended to include all possible types of devices and ways to output information from computer system 3700 to the user or to another machine or computer system.
- Storage subsystem 3710 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by processors 3778.
- Processors 3778 can be graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and/or coarse-grained reconfigurable architectures (CGRAs).
- GPUs graphics processing units
- FPGAs field-programmable gate arrays
- ASICs application-specific integrated circuits
- CGRAs coarse-grained reconfigurable architectures
- Processors 3778 can be hosted by a deep learning cloud platform such as Google Cloud PlatformTM, XilinxTM, and CirrascaleTM.
- processors 3778 include Google’s Tensor Processing Unit (TPU)TM, rackmount solutions like GX4 Rackmount SeriesTM, GX37 Rackmount SeriesTM, NVIDIA DGX-1TM, Microsoft' Stratix V FPGATM, Graphcore's Intelligent Processor Unit (IPU)TM, Qualcomm’s Zeroth PlatformTM with Snapdragon processorsTM, NVIDIA’ s VoltaTM, NVIDIA’ s DRIVE PXTM, NVIDIA’ s JETSON TX1/TX2 MODULETM, Intel’s NirvanaTM, Movidius VPUTM, Fujitsu DPITM, ARM’s DynamicIQTM, IBM TrueNorthTM, Lambda GPU Server with Testa VI 00sTM, and others.
- TPU Tensor Processing Unit
- rackmount solutions like GX4 Rackmount SeriesTM, GX37 Rackmount SeriesTM, NVIDIA DGX-1TM, Microsoft' Stratix V FPGATM, Graphcore's Intelligent Processor Unit (IPU)TM
- Memory subsystem 3722 used in the storage subsystem 3710 can include a number of memories including a main random access memory (RAM) 3732 for storage of instructions and data during program execution and a read only memory (ROM) 3734 in which fixed instructions are stored.
- a file storage subsystem 3736 can provide persistent storage for program and data files, and can include a hard disk drive, a floppy disk drive along with associated removable media, a CD-ROM drive, an optical drive, or removable media cartridges.
- the modules implementing the functionality of certain implementations can be stored by file storage subsystem 3736 in the storage subsystem 3710, or in other machines accessible by the processor.
- Bus subsystem 3755 provides a mechanism for letting the various components and subsystems of computer system 3700 communicate with each other as intended. Although bus subsystem 3755 is shown schematically as a single bus, alternative implementations of the bus subsystem can use multiple busses.
- Computer system 3700 itself can be of varying types including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a widely -distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 3700 depicted in Figure 37 is intended only as a specific example for purposes of illustrating the preferred implementations of the present invention. Many other configurations of computer system 3700 are possible having more or less components than the computer system depicted in Figure 37. Neural Network-Based Base Caller
- the following discussion focuses on a neural network-based base caller described herein that can be used in conjunction with the specialist signal profilers.
- the input to the neural network-based base caller is described, in accordance with one implementation.
- examples of the structure and form of the neural network-based base caller are provided.
- the output of the neural network-based base caller is described, in accordance with one implementation.
- a data flow logic provides the sequencing images to the neural network-based base caller for base calling.
- the neural network-based base caller accesses the sequencing images on a patch-by-patch basis (or a tile-by-tile basis).
- Each of the patches is a sub-grid (or sub-array) of pixelated units in the grid of pixelated units that forms the sequencing images.
- the patches have dimensions q x r of the sub-grid of pixelated units, where q (width) and r (height) are any numbers ranging from 1 and 10000 ( e.g ., 3 x 3, 5 x 5, 7 x 7, 10 x 10, 15 x 15, 25 x 25, 64 x 64,
- the patches extracted from a sequencing image are of the same size. In other implementations, the patches are of different sizes. In some implementations, the patches can have overlapping pixelated units (e.g., on the edges).
- each of the sequencing images has one or more image (or intensity) channels (analogous to the red, green, blue (RGB) channels of a color image).
- each image channel corresponds to one of a plurality of filter wavelength bands.
- each image channel corresponds to one of a plurality of imaging events at a sequencing cycle.
- each image channel corresponds to a combination of illumination with a specific laser and imaging through a specific optical filter.
- the image patches are tiled (or accessed) from each of the m image channels for a particular sequencing cycle.
- m is 4 or 2.
- m is 1, 3, or greater than 4.
- the images can be in blue and violet color channels instead of or in addition to the red and green channels.
- the input image data to the neural network-based base caller for a single iteration of base calling comprises data for a sliding window of multiple sequencing cycles.
- the sliding window can include, for example, a current sequencing cycle, one or more preceding sequencing cycles, and one or more successive sequencing cycles.
- the input image data comprises data for three sequencing cycles, such that data for a current (time t) sequencing cycle to be base called is accompanied with (i) data for a left flanking/context/previous/preceding/prior (time t- 1) sequencing cycle and (ii) data for a right flanking/context/next/successive/subsequent (time /+1) sequencing cycle.
- the input image data comprises data for five sequencing cycles, such that data for a current (time t) sequencing cycle to be base called is accompanied with (i) data for a first left flanking/context/previous/preceding/prior (time t- 1) sequencing cycle,
- the input image data comprises data for seven sequencing cycles, such that data for a current (time t) sequencing cycle to be base called is accompanied with (i) data for a first left flanking/context/previous/preceding/prior (time t- 1) sequencing cycle, (ii) data for a second left flanking/context/previous/preceding/prior (time 1-2) sequencing cycle, (iii) data for a third left flanking/context/previous/preceding/prior (time t- 3) sequencing cycle, (iv) data for a first right flanking/context/next/successive/subsequent (time /+1), (v) data for a second right flanking/context/next/successive/subsequent (time 1+2) sequencing cycle, and (vi) data for a third right flanking/context/next/successive/subsequent (time /+) sequencing cycle, and (vi)
- the input image data comprises data for a single sequencing cycle. In yet other implementations, the input image data comprises data for 10, 15, 20, 30, 58, 75, 92, 130, 168, 175, 209, 225, 230, 275, 318, 325, 330, 525, or 625 sequencing cycles.
- the neural network-based base caller processes the image patches through its convolution layers and produces an alternative representation, according to one implementation.
- the alternative representation is then used by an output layer (e.g ., a softmax layer) for generating a base call for either just the current (time t) sequencing cycle or each of the sequencing cycles, i.e., the current (time I) sequencing cycle, the first and second preceding (time t- 1, time 1-2) sequencing cycles, and the first and second succeeding (time /+ 1, time 1+2) sequencing cycles.
- the resulting base calls form the sequencing reads.
- the neural network-based base caller outputs a base call for a single target cluster for a particular sequencing cycle.
- the neural network-based base caller outputs a base call for each target cluster in a plurality of target clusters for the particular sequencing cycle. In yet another implementation, the neural network-based base caller outputs a base call for each target cluster in a plurality of target clusters for each sequencing cycle in a plurality of sequencing cycles, thereby producing a base call sequence for each target cluster.
- the neural network-based base caller is a multilayer perceptron (MLP). In another implementation, the neural network-based base caller is a feedforward neural network. In yet another implementation, the neural network-based base caller is a fully-connected neural network. In a further implementation, the neural network-based base caller is a fully convolution neural network. In yet further implementation, the neural network-based base caller is a semantic segmentation neural network. In yet another further implementation, the neural network-based base caller is a generative adversarial network (GAN). [00209] In one implementation, the neural network-based base caller is a convolution neural network (CNN) with a plurality of convolution layers.
- CNN convolution neural network
- the neural network-based base caller is a recurrent neural network (RNN) such as a long short-term memory network (LSTM), bi-directional LSTM (Bi-LSTM), or a gated recurrent unit (GRU).
- RNN recurrent neural network
- LSTM long short-term memory network
- Bi-LSTM bi-directional LSTM
- GRU gated recurrent unit
- the neural network-based base caller includes both a CNN and an RNN.
- the neural network-based base caller can use ID convolutions, 2D convolutions, 3D convolutions, 4D convolutions, 5D convolutions, dilated or atrous convolutions, transpose convolutions, depthwise separable convolutions, pointwise convolutions, l x l convolutions, group convolutions, flattened convolutions, spatial and cross channel convolutions, shuffled grouped convolutions, spatial separable convolutions, and deconvolutions.
- the neural network-based base caller can use one or more loss functions such as logistic regression/log loss, multi-class cross-entropy/softmax loss, binary cross-entropy loss, mean-squared error loss, LI loss, L2 loss, smooth LI loss, and Huber loss.
- the neural network- based base caller can use any parallelism, efficiency, and compression schemes such TFRecords, compressed encoding (e.g ., PNG), sharding, parallel calls for map transformation, batching, prefetching, model parallelism, data parallelism, and synchronous/asynchronous stochastic gradient descent (SGD).
- loss functions such as logistic regression/log loss, multi-class cross-entropy/softmax loss, binary cross-entropy loss, mean-squared error loss, LI loss, L2 loss, smooth LI loss, and Huber loss.
- the neural network- based base caller can use any parallelism, efficiency, and compression schemes such TFRecords, compressed encoding (e.g ., PNG
- the neural network-based base caller can include upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (like an LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., non-linear transformation functions like rectifying linear unit (ReLU), leaky ReLU, exponential liner unit (ELU), sigmoid and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g ., max or average pooling), global average pooling layers, and attention mechanisms.
- ReLU rectifying linear unit
- ELU exponential liner unit
- sigmoid and hyperbolic tangent sigmoid and hyperbolic tangent
- the neural network-based base caller is trained using backpropagati on-based gradient update techniques.
- Example gradient descent techniques that can be used for training the neural network-based base caller include stochastic gradient descent, batch gradient descent, and mini batch gradient descent.
- Some examples of gradient descent optimization algorithms that can be used to train the neural network-based base caller are Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.
- the neural network-based base caller uses a specialized architecture to segregate processing of data for different sequencing cycles. The motivation for using the specialized architecture is described first. As discussed above, the neural network-based base caller processes image patches for a current sequencing cycle, one or more preceding sequencing cycles, and one or more successive sequencing cycles. Data for additional sequencing cycles provides sequence-specific context. The neural network-based base caller learns the sequence-specific context during training and base calls them. Furthermore, data for pre and post sequencing cycles provides second order contribution of pre-phasing and phasing signals to the current sequencing cycle.
- the specialized architecture comprises spatial convolution layers that do not mix information between sequencing cycles and only mix information within a sequencing cycle.
- Spatial convolution layers use so-called “segregated convolutions” that operationalize the segregation by independently processing data for each of a plurality of sequencing cycles through a “dedicated, non-shared” sequence of convolutions.
- the segregated convolutions convolve over data and resulting feature maps of only a given sequencing cycle, i.e ., intra-cycle, without convolving over data and resulting feature maps of any other sequencing cycle.
- the input image data comprises (i) current image patch for a current (time t) sequencing cycle to be base called, (ii) previous image patch for a previous (time t- 1) sequencing cycle, and (iii) next image patch for a next (time /+ 1 ) sequencing cycle.
- the specialized architecture then initiates three separate convolution pipelines, namely, a current convolution pipeline, a previous convolution pipeline, and a next convolution pipeline.
- the current data processing pipeline receives as input the current image patch for the current (time t) sequencing cycle and independently processes it through a plurality of spatial convolution layers to produce a so-called “current spatially convolved representation” as the output of a final spatial convolution layer.
- the previous convolution pipeline receives as input the previous image patch for the previous (time t- 1) sequencing cycle and independently processes it through the plurality of spatial convolution layers to produce a so-called “previous spatially convolved representation” as the output of the final spatial convolution layer.
- the next convolution pipeline receives as input the next image patch for the next (time /+1) sequencing cycle and independently processes it through the plurality of spatial convolution layers to produce a so-called “next spatially convolved representation” as the output of the final spatial convolution layer.
- the current, previous, and next convolution pipelines are executed in parallel.
- the spatial convolution layers are part of a spatial convolution network (or subnetwork) within the specialized architecture.
- the neural network-based base caller further comprises temporal convolution layers (or temporal logic) that mix information between sequencing cycles, i.e., inter-cycles.
- the temporal convolution layers receive their inputs from the spatial convolution network and operate on the spatially convolved representations produced by the final spatial convolution layer for the respective data processing pipelines.
- the inter-cycle operability freedom of the temporal convolution layers emanates from the fact that the misalignment property, which exists in the image data fed as input to the spatial convolution network, is purged out from the spatially convolved representations by the stack, or cascade, of segregated convolutions performed by the sequence of spatial convolution layers.
- Temporal convolution layers use so-called “combinatory convolutions” that groupwise convolve over input channels in successive inputs on a sliding window basis.
- the successive inputs are successive outputs produced by a previous spatial convolution layer or a previous temporal convolution layer.
- the temporal convolution layers are part of a temporal convolution network (or subnetwork) within the specialized architecture.
- the temporal convolution network receives its inputs from the spatial convolution network.
- a first temporal convolution layer of the temporal convolution network groupwise combines the spatially convolved representations between the sequencing cycles.
- subsequent temporal convolution layers of the temporal convolution network combine successive outputs of previous temporal convolution layers.
- the output of the final temporal convolution layer is fed to an output layer that produces an output.
- the output is used to base call one or more clusters at one or more sequencing cycles.
- a system comprising: memory storing a plurality of specialist signal profilers, wherein each specialist signal profiler in the plurality of specialist signal profilers is trained to maximize signal-to-noise ratio of sequenced signals in a particular signal profile detected for analytes in a particular analyte class and characterized in a particular training data set; and runtime logic, having access to the memory, configured to execute a base calling operation by applying respective specialist signal profilers in the plurality of specialist signal profilers to sequenced signals in respective signal profiles detected for analytes in respective analyte classes during the base calling operation.
- each specialist signal profiler is further trained to maximize signal-to-noise ratio of sequenced signals in a particular signal profile detected for analytes in a particular analyte sub-class and characterized in a particular training data sub-set
- the runtime logic is further configured to execute the base calling operation by applying the respective specialist signal profilers to sequenced signals in respective signal profiles detected for analytes in respective analyte sub-classes during the base calling operation.
- each specialist signal profiler is configured with channel-specific equalizers, wherein each channel-specific equalizer has a plurality of convolution kernels.
- the runtime logic is further configured to implement expectation maximization that iteratively maximizes a likelihood of channel-wise observing base-wise signal centroids and signal distributions that best fit sequenced signals detected so far during the base calling operation, to channel-wise determine signal-to-noise ratio-maximized sequenced signals in response to applying the respective specialist signal profilers to the sequenced signals, to call bases based on the signal-to-noise ratio-maximized sequenced signals, to channel-wise determine base calling errors based on comparing the signal-to-noise ratio-maximized sequenced signals against signal centroids of the called bases, and to channel-wise update coefficients of convolution kernels of the respective specialist signal profilers based on the base calling errors.
- a system comprising: memory storing initially sequenced signals detected during initial sequencing cycles of a sequencing run; fitting logic, having access to the memory, configured to fit a plurality of signal distributions on the initially sequenced signals, and to store the plurality of signal distributions in the memory; online training logic, having access to the memory, configured to train respective specialist signal profilers in a plurality of specialist signal profilers to maximize signal-to-noise ratio of respective signal distributions in the plurality of signal distributions, and to store the trained respective specialist signal profilers in the memory; and runtime logic, having access to the memory, configured to uniquely map subsequently sequenced signals detected during subsequent sequencing cycles of the sequence run to the respective signal distributions, and to apply the trained respective specialist signal profilers to the subsequently sequenced signals based on the unique mapping to the respective signal distributions to generate base calls for the subsequent sequencing cycles.
- each set of variation coefficients includes channel- specific amplification coefficients that correct scale variations in a corresponding signal distribution.
- each set of variation coefficients includes channel-specific offset coefficients that correct shift variations in a corresponding signal distribution.
- control logic configured to iterate execution of the fitting logic, the online training logic, and the runtime logic at each sequencing cycle of the sequencing run.
- control logic configured to iterate execution of the fitting logic, the online training logic, and the runtime logic after certain number of sequencing cycles of the sequencing run.
- control logic configured to iterate execution of the fitting logic, the online training logic, and the runtime logic at certain predetermined sequencing cycles of the sequencing run.
- a system comprising: memory storing a plurality of specialist signal profilers configured for use in a base calling operation, respective specialist signal profilers in the plurality of specialist signal profilers trained to maximize signal-to-noise ratio of sensor data in respective signal distributions observed at respective sequencing events of the base calling operation and characterized in respective training data sets; and runtime logic, having access to the memory, configured to select a specialist signal profiler from the plurality of specialist signal profilers based on a subject sequencing event that generated subject sensor data, and apply the selected specialist signal profiler on the subject sensor data to produce base call classification data for the subject sequencing event.
- runtime logic is further configured to iteratively train the respective specialist signal profilers during the sequencing run at the respective sequencing events.
- the runtime logic is further configured to implement expectation maximization that iteratively maximizes a likelihood of channel-wise observing base-wise signal centroids and signal distributions that best fit the sensor data, to channel-wise determine signal-to-noise ratio-maximized sensor data in response to applying the respective specialist signal profilers to the sensor data, to call bases based on the signal-to-noise ratio-maximized sensor data, to channel-wise determine base calling errors based on comparing the signal-to-noise ratio-maximized sensor data against signal centroids of the called bases, and to channel-wise update coefficients of convolution kernels of the respective specialist signal profilers based on the base calling errors.
- a system comprising: memory storing initially sequenced signals detected during initial sequencing cycles of a sequencing run for a population of analytes; fitting logic, having access to the memory, configured to process the initially sequenced signals on an analyte-by-analyte basis, to fit respective signal profiles for respective analytes in the population of analytes, and to store the respective signal profiles in the memory; online training logic, having access to the memory, configured to train respective specialist signal profilers in a plurality of specialist signal profilers to maximize signal-to-noise ratio of the respective signal profiles fitted for the respective analytes, and to store the trained respective specialist signal profilers in the memory; and runtime logic, having access to the memory, configured to uniquely map subsequently sequenced signals detected during subsequent sequencing cycles of the sequence run to the respective signal profiles on the analyte-by-analyte basis, and to apply the trained respective specialist signal profilers to the subsequently sequenced signals based on the unique mapping to the respective signal profiles to generate base calls for the subsequent sequencing cycles on
- each set of variation coefficients includes channel- specific amplification coefficients that correct scale variations in a corresponding signal profile.
- each set of variation coefficients includes channel- specific offset coefficients that correct shift variations in a corresponding signal profile.
- control logic configured to iterate execution of the fitting logic, the online training logic, and the runtime logic at each sequencing cycle of the sequencing run.
- control logic configured to iterate execution of the fitting logic, the online training logic, and the runtime logic after certain number of sequencing cycles of the sequencing run.
- control logic configured to iterate execution of the fitting logic, the online training logic, and the runtime logic at certain predetermined sequencing cycles of the sequencing run.
- a system comprising: memory storing a plurality of specialist signal profilers, wherein each specialist signal profiler in the plurality of specialist signal profilers is trained to maximize signal-to-noise ratio of sequenced signals in a particular signal profile detected for a particular analyte and characterized in a particular training data set; and runtime logic, having access to the memory, configured to execute a base calling operation by applying respective specialist signal profilers in the plurality of specialist signal profilers to sequenced signals in respective signal profiles detected for respective analytes during the base calling operation.
- memory storing a plurality of specialist signal profilers, wherein each specialist signal profiler in the plurality of specialist signal profilers is trained to maximize signal-to-noise ratio of sequenced signals in a particular signal profile detected for a particular analyte and characterized in a particular training data set
- runtime logic having access to the memory, configured to execute a base calling operation by applying respective specialist signal profilers in the plurality of specialist signal profilers to sequenced signals in respective signal profiles detected for respective ana
- a system comprising: spatial classification logic configured to use analyte location to segment a population of analytes into spatial classes, wherein each spatial class includes a nonoverlapping set of analytes culled from the population of analytes; signal profiling logic configured to use sequenced signals detected for the population of analytes to estimate one or more signal profiles for each spatial class, wherein each signal profile includes a nonoverlapping subset of analytes culled from a nonoverlapping set of analytes; online training logic configured to train at least one specialist signal profiler to maximize a signal-to-noise ratio of each signal profile estimated for each spatial class; and runtime logic configured to base call analytes in respective nonoverlapping subsets of analytes using respective trained specialist signal profilers.
- control logic configured to iterate execution of the signal profiling logic, the online training logic, and the runtime logic at certain predetermined sequencing cycles of the sequencing run.
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Biophysics (AREA)
- General Health & Medical Sciences (AREA)
- Biotechnology (AREA)
- Chemical & Material Sciences (AREA)
- Bioinformatics & Computational Biology (AREA)
- Data Mining & Analysis (AREA)
- Theoretical Computer Science (AREA)
- Evolutionary Biology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Evolutionary Computation (AREA)
- Bioethics (AREA)
- Software Systems (AREA)
- Public Health (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Artificial Intelligence (AREA)
- Epidemiology (AREA)
- Databases & Information Systems (AREA)
- Organic Chemistry (AREA)
- Molecular Biology (AREA)
- Analytical Chemistry (AREA)
- Wood Science & Technology (AREA)
- Zoology (AREA)
- Signal Processing (AREA)
- Immunology (AREA)
- Microbiology (AREA)
- Biochemistry (AREA)
- General Engineering & Computer Science (AREA)
- Genetics & Genomics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
Description
Claims
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202163223408P | 2021-07-19 | 2021-07-19 | |
| US17/839,353 US20230018469A1 (en) | 2021-07-19 | 2022-06-13 | Specialist signal profilers for base calling |
| PCT/US2022/037226 WO2023003758A1 (en) | 2021-07-19 | 2022-07-14 | Specialist signal profilers for base calling |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4374378A1 true EP4374378A1 (en) | 2024-05-29 |
Family
ID=82846128
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP22751936.0A Pending EP4374378A1 (en) | 2021-07-19 | 2022-07-14 | Specialist signal profilers for base calling |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4374378A1 (en) |
| JP (1) | JP2024528511A (en) |
| KR (1) | KR20240026931A (en) |
| WO (1) | WO2023003758A1 (en) |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US2073908A (en) | 1930-12-29 | 1937-03-16 | Floyd L Kallam | Method of and apparatus for controlling rectification |
| US8778848B2 (en) | 2011-06-09 | 2014-07-15 | Illumina, Inc. | Patterned flow-cells useful for nucleic acid analysis |
| CN112203648A (en) * | 2018-03-30 | 2021-01-08 | 朱诺诊断学公司 | Deep learning based methods, devices and systems for prenatal care |
| US11783917B2 (en) * | 2019-03-21 | 2023-10-10 | Illumina, Inc. | Artificial intelligence-based base calling |
| US11593649B2 (en) * | 2019-05-16 | 2023-02-28 | Illumina, Inc. | Base calling using convolutions |
-
2022
- 2022-07-14 KR KR1020237043986A patent/KR20240026931A/en active Pending
- 2022-07-14 EP EP22751936.0A patent/EP4374378A1/en active Pending
- 2022-07-14 JP JP2023579788A patent/JP2024528511A/en active Pending
- 2022-07-14 WO PCT/US2022/037226 patent/WO2023003758A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2023003758A1 (en) | 2023-01-26 |
| JP2024528511A (en) | 2024-07-30 |
| KR20240026931A (en) | 2024-02-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12100125B2 (en) | Optical distortion correction for imaged samples | |
| JP7566638B2 (en) | Artificial Intelligence-Based Sequencing | |
| NL2023316B1 (en) | Artificial intelligence-based sequencing | |
| NL2023310B1 (en) | Training data generation for artificial intelligence-based sequencing | |
| US20220067489A1 (en) | Detecting and Filtering Clusters Based on Artificial Intelligence-Predicted Base Calls | |
| US20230018469A1 (en) | Specialist signal profilers for base calling | |
| US20220319639A1 (en) | Artificial intelligence-based base caller with contextual awareness | |
| WO2023003758A1 (en) | Specialist signal profilers for base calling | |
| US20240257320A1 (en) | Jitter correction image analysis | |
| CN117546247A (en) | Specialized signal analyzer for base detection | |
| WO2022212180A1 (en) | Artificial intelligence-based base caller with contextual awareness | |
| US20260091385A1 (en) | Apparatus and methods for sample loading | |
| WO2026080856A1 (en) | Substrates and related methods and systems | |
| WO2026096670A1 (en) | Precision, recall, and f-score for continuous distributions | |
| WO2023014741A1 (en) | Base calling using multiple base caller models | |
| WO2023049215A1 (en) | Compressed state-based base calling | |
| NZ752391B2 (en) | Optical distortion correction for imaged samples |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20231220 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| REG | Reference to a national code |
Ref country code: HK Ref legal event code: DE Ref document number: 40104565 Country of ref document: HK |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |