WO2015154137A1 - A method and a system for identifying molecules - Google Patents
A method and a system for identifying molecules Download PDFInfo
- Publication number
- WO2015154137A1 WO2015154137A1 PCT/AU2015/000216 AU2015000216W WO2015154137A1 WO 2015154137 A1 WO2015154137 A1 WO 2015154137A1 AU 2015000216 W AU2015000216 W AU 2015000216W WO 2015154137 A1 WO2015154137 A1 WO 2015154137A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- fluorescent
- detection
- events
- detection data
- molecules
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N15/00—Investigating characteristics of particles; Investigating permeability, pore-volume or surface-area of porous materials
- G01N15/10—Investigating individual particles
- G01N15/14—Optical investigation techniques, e.g. flow cytometry
- G01N15/1429—Signal processing
- G01N15/1433—Signal processing using image recognition
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N15/00—Investigating characteristics of particles; Investigating permeability, pore-volume or surface-area of porous materials
- G01N2015/0042—Investigating dispersion of solids
- G01N2015/0053—Investigating dispersion of solids in liquids, e.g. trouble
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N15/00—Investigating characteristics of particles; Investigating permeability, pore-volume or surface-area of porous materials
- G01N15/10—Investigating individual particles
- G01N2015/1006—Investigating individual particles for cytology
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N15/00—Investigating characteristics of particles; Investigating permeability, pore-volume or surface-area of porous materials
- G01N15/10—Investigating individual particles
- G01N15/14—Optical investigation techniques, e.g. flow cytometry
- G01N2015/1486—Counting the particles
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N15/00—Investigating characteristics of particles; Investigating permeability, pore-volume or surface-area of porous materials
- G01N15/10—Investigating individual particles
- G01N15/14—Optical investigation techniques, e.g. flow cytometry
- G01N2015/1488—Methods for deciding
Definitions
- the present invention relates to microscopy methods and systems for identifying molecules, and particularly, but not exclusively to a system and a method for locating and counting molecules.
- Microscopy techniques are generally used to investigate macromolecular structures at intermediate and high
- EM electron microscopy
- cryo-EM and crystallography are used to examine the molecular structure of proteins, but these techniques often involve elaborate sample preparation or require isolating the protein complexes from their natural environment.
- Optical microscopy techniques with fluorescent molecules are used as a less invasive alternative.
- the resolution of conventional light microscopy is limited by the diffraction limit of light to approximately 250 nm. This is generally larger than the size of protein
- Photo-activated localization microscopy and related techniques aim to enable imaging and counting fluorescent molecules spatially closer than the diffraction limit.
- errors related to the fluorescent blinking, photo-physics and photo-bleaching behaviour make this technique unreliable when quantitative data are required.
- microscopy techniques for imaging and counting fluorescent molecules.
- the present invention provides a method for counting fluorescent molecules in a sample comprising the steps of:
- the method further comprises the step of associating detection paths corresponding to fluorescent events of existing fluorescent molecules into detection trajectories; each trajectory representing an exclusive association of detection events of a single fluorescent molecule.
- the step of dividing detection data into spatial regions of interest comprises calculating a spatial distance between two fluorescent events and comparing the calculated distance with a predetermined distance.
- a modified version of a DBSCAN algorithm may be used to perform the step of dividing detection data into spatial regions of interest.
- a plurality of spatial distances are calculated to compose at least one region of interest comprising at least a portion of the sample surface.
- the step of linking detection data in a plurality of detection paths further comprises the step of setting a variable metric to measure a distance between fluorescent events occurring in different frames.
- the variable metric may be a mathematical function of the Hellinger Distance.
- the variable metric may also be a probabilistic metric related to one or more probabilistic distributions of the fluorescent events. Further, the different frames may be consecutive frames. Each of the one or more probabilistic distributions may be used to define a position and a localisation precision of a fluorescent event.
- the method further comprises the step of comparing the metric to a score threshold to define if two fluorescent events are linked.
- the score threshold may vary for different samples and may be related to dye photo-physical kinetics.
- the method further comprises the step of organising linked fluorescent events in short tracks of fluorescent events. In an embodiment, the method comprises the step of
- the method further comprises the step of calculating a number of new particles born in each frame based on the histogram of the frame.
- the method further comprises the step of calculating the probability that a new fluorescent
- the method further comprises the step of calculating a likelihood of a short track being the result of a new molecule arising at a given time in a given location .
- the method further comprises the step of calculating a likelihood of a short track being the result of the same molecule.
- the likelihood may be calculated simultaneously for all the short tracks in a given region of interest.
- a Munkres algorithm may be used to associate short tracks to existing molecules of new molecules arising at a given time in a given location.
- the method further comprises the step of calculating a plurality of trajectories related to a plurality of fluorescent molecules by pairing short tracks. In an embodiment, the method further comprises the step of simulating a kinetic behaviour of a molecular cluster compatible with the detection data extracted from at least a portion of the acquired frames and comparing the
- the method further comprises the step of calculating a kinetic behaviour of a molecular cluster compatible with the detection data extracted from at least a portion of the acquired frames based on a theoretical model and comparing the calculated plurality of
- a data processing apparatus comprising a plurality of operating modules arranged for carrying out a method in accordance with the first aspect.
- the present invention provides a computer program, comprising instructions for controlling a computer to perform a method in accordance with the first aspect.
- the present invention provides a computer readable medium, providing the computer program of the third aspect.
- the present invention provides a data signal, providing a computer program in accordance with the third aspect.
- the present invention provides an apparatus for counting fluorescent molecules in a sample, from fluorescence detection data regarding timing and location of fluorescent events obtained from a plurality of image frames of the sample, the apparatus comprising;
- a splitting module for dividing the detection data into spatial regions of interest based on location of fluorescent events
- a linking module for linking the split detection data in a plurality of detection paths in a region of interest based on location and timing of fluorescent events in the region of interest;
- a processing module for analysing the detection paths and determine whether each detection path is related to a new fluorescent molecule or a new fluorescent event of an existing fluorescent molecule.
- the apparatus further comprises an associator for associating detection paths corresponding to fluorescent events of existing fluorescent molecules into detection trajectories; each trajectory representing an exclusive association of detection events of a single fluorescent molecule.
- the apparatus further comprises a microscopy unit arranged to acquire image frames of a biological sample and an analyser unit arranged to extract detection data related to timing and location of
- invention provides a method for counting fluorescent molecules in a sample comprising the steps of:
- invention provides an apparatus for counting fluorescent molecules in a biological sample, the apparatus
- a microscopy unit arranged to acquire image frames of a biological sample
- an image analyser unit arranged to extract detection data related to timing and location of fluorescent events; a comparator arranged to compare the timing detection data with the location detection data to identify a number and location of the fluorescent molecules.
- Advantageous embodiments provide microscopy methods and systems for counting molecules based on spatial and temporal information. Spatial and temporal information are processed using mathematical models and variable thresholds to extract a reliable number of fluorophores and their trajectories in a cluster in a regime where the activation rate of the clusters is not constant in time.
- Advantageous embodiments may be used to identify the presence and number of target molecules in situ, within a cell for example.
- the system and methods disclosed herein are non-destructive for the sample as they are based on the acquisition of a sequence of optical images.
- Figure 1 shows a common photo-conversion scheme of photo- activatable fluorescent proteins
- Figures 2 and 3 shows flow-charts with steps for counting fluorescent molecules in a sample in accordance with embodiments ;
- Figure 4 schematically illustrates the linking of short tracks into trajectories in accordance with embodiments
- Figures 5 to 7 show maps of a simulated data set, a calculated data set and a comparison of true
- Figure 8 shows a fit of the calculated cluster data to a probability distribution and the variance of cluster data as a function of sampling fraction in accordance with embodiments ;
- Figure 9 shows a plot with experimental data and simulated data in accordance with embodiments
- Figure 10 is a schematic illustration of a computer system suitable to implement a method in accordance with
- Embodiments of the present invention relate to microscopy methods and systems for counting molecules.
- microscopy domain The approach used allows tracking and deriving
- the process is based and operates on detection data of fluorescent events extracted from the high
- the photo-activatable fluorescent proteins may be in a x pre-activation' state 102, an x active' state 104, a x bleached state' 106 or two
- transition rates k a and k b irreversible and have respective transition rates k a and k b .
- the transitions from active 104 to dark 1 108 or dark 2 110 are considered to be reversible and have respective transition rates k d and ak d .
- a is a scaling factor which biases the transition probability towards one dark state over the other.
- the transition rates from dark 1 108 and dark 2 110 to active 104 are instead different and defined respectively as k rl and k r2 .
- values of k b , k d , a, k r i and k r 2 estimated by measurements of a known sample of monomeric fluorophores are used. The measurements are performed in the same conditions as per the molecule counting process. In particular, the same incident laser intensity is used.
- a flow-chart 200 with steps for counting fluorescent molecules in a sample in accordance with embodiments.
- a plurality of image frames of a sample are acquired using an image acquisition system, such as a fluorescence optical
- the data are processed at step 210 to extract information related to timing and spatial location of fluorescent events. In some embodiments these information are
- the detection data are then divided into spatial regions of interest at step 220 based on the location of fluorescent events.
- the splitting of the data in regions of interest (ROIs) is performed to optimise the analysis of the data using performing fitting algorithms and ultimately improve molecule detection.
- ROIs regions of interest
- the splitting process allows reducing the number of detection events that must be considered simultaneously.
- the division is performed using a recursive approach .
- each ROI is considered to be exclusive from all other regions of interest and events from one region cannot be linked to events in another region.
- the detection events in each ROI are linked in a plurality of detection paths in a ROI at step 225.
- the detection events are linked based on both their location and timing.
- the data related to detection paths obtained by linking detection events in each ROI are analysed at step 230 to determine whether each detection path is related to a new fluorescent molecule or a new fluorescent event of an existing fluorescent molecule.
- the number of molecules in each ROI is added to obtain the number of molecules in the sample .
- FIG 3 there is shown a flow-chart 300 listing method steps to perform the counting algorithm in accordance with embodiments.
- the flow-chart 300 is based on the chart 200 of figure 2 and adds further details.
- Calibration data 305 is collected by analysis of data acquired in a similar manner as the image data from a sample of known form of isolated molecules. Analysis of the signal intermittency and total length for individual molecules the kinetics constants k b , k d , k rl , k r2 , and a are determined .
- the detection events in each ROI are first linked into short tracks of events occurring in consecutive frames at step 310.
- the algorithm monitors detection events occurring in the next frame. In case an event is found, Gaussian distributions related to the two detection events are calculated together with a mathematical
- the distance 315 which separates the two distributions.
- the position of each detection event is the mean of the respective Gaussian distribution and the localization precision is the covariance of the distributions.
- the distance metric at 315 is compared against a score
- threshold (ST) 320 which defines a threshold for detection events in consecutive frames to be linked.
- the value of ST may be varied in accordance with the requirements in specific experiments. In advantageous embodiments the value of ST is based on the dye kinetics yields.
- the kinetics-based value of ST is given by the probability that a molecule transition away from an active state in a single frame.
- t f is the time between consecutive frames. If the distance metric exceeds the ST value, the two detection events are linked into a short track. In the case where more than one detection event may occur in either the current or next frame then the distance metric and score threshold values are concatenated into an appropriate matrix and a global minimization assignment algorithm executed to pair current detection events with either future detection events or the ST. Those
- the distance metric 315 is the square root of 1 minus the Hellinger Distance ((1- D H ) 1/2 ) and can be calculated according to the equation below based on the multivariate Bhattacharyya Coefficient.
- positions of the new aggregated events and their accuracy are calculated by weighted averages accounting for the number of photons detected for each short track and the number of events in the short track.
- x and y ⁇ are the positions of the detection events and N ⁇ the number of photons detected for each event i in the short track.
- the spatial uncertainty of the localisation is given by ⁇ .
- ⁇ and N Events correspond to the uncertainty of localization i in the short track and the number of events in the track, respectively.
- the overall uncertainty, ⁇ is taken as the greater of ⁇ ⁇ and Oy (equivalent to ⁇ ⁇ given above) .
- the greater of the two uncertainties is carried forward.
- a bin width of 1000-2000 frames may be appropriate for a dataset of 25000 frames.
- Scaling this histogram provides a smoothed approximation of the number of new particles born in a given frame, 330. This approximation may be a slight over- estimate of the actual probability as some of the tracks in some frames are continuations of other short tracks rather than true new molecules.
- the accuracy of the estimate is related to the kinetics of the fluorophores , namely the number of occurrences of a molecule returning from a dark state 108 or 110 to an on state 104 before transitioning to a bleach state 106.
- P Wew The probability of a new molecule arriving in a given position in a given frame
- fractional area for randomly-distributed molecules is taken as the area occupied by a circle with a radius equal to the sum of the objective PSFs of the two molecules that could be linked, divided by the area of the entire
- the fractional area is taken as a fraction of the intensity of the image generated from the sum of all frames in the data set that is present within a given distance from the detection event.
- this distance is a circle with the radius of one objective PSF.
- the product of the area fraction PSF Area and P New yields a score S New which is a measure of the likelihood of a short track being the result of a new molecule arising at a given time in a given location.
- the complementary score S L i nk is a measure of the likelihood of a short track being related to the same molecule and is calculated at 335.
- S L i nk is the product of three values: the overlap score between short track positions, ⁇ 1-D H ) 1/2 , the off time score S off , and the dark state probability, P d ar k -
- S 0ff is proportional to the probability that a molecule in the dark state will return to the active state after a dark state dwell time of At. S 0ff is calculated in
- the dark state probability, Pdar k r is the probability that molecule in the active state will transition to a dark state rather than to the bleach state.
- Pdar k is calculated according to the following equation. The values for k ⁇ , and are obtained from the results of calibration experiments with the same fluorescent molecule as described previously.
- the product of (1-D H ) 1 2 , S 0ff and Pdar k generated at 335 is compared in an inequality with the product of PSF Area and P Neu of 330 to ascertain whether a short track originating in the spatial and temporal proximity of the end of another short track is more likely due to a molecule returning from a dark state to the active state or from the activation of a new molecule.
- FIG. 4 there is shown a line diagram 400 that schematically illustrates the linking of short tracks (solid lines) into trajectories by bridging off gaps between valid short track termini pairs (dashed and dotted lines) .
- Scores for each valid gap are assembled into a matrix to be passed to an optimal assignment algorithm. Those associations found to be valid are used to assemble trajectories from short tracks and gaps, here shown as dashed lines.
- the structure of a possible score matrix is provided below. Tuck J
- Scores are indicated by S ⁇ , j and are generated for all valid parings (or where a short track start is followed by a short track end) taking into account spatial and
- S New values are considered for the later short tracks in a further matrix appended to the matrix of S L i nk values.
- the optimal assignment is determined via Munkres algorithm which enforces the constraint that each short track origination can be paired with either a single short track end or the null S New option, and each short track end can be paired with at most a single short track start .
- the elements indicated by S ⁇ fj designate a valid potential pairing between the end of short track i and the start of short track j (this element will have the value of S ⁇ ink for this pairing) .
- the elements indicated by S NeH/ j designate the value of S New for that particular short track.
- the trajectories of each molecule are determined at 340 by pairing short tracks.
- the kinetic behaviour of the analysed individual molecule trajectory data from at least a portion of the acquired frames and at least a portion of the analysed trajectories are compared to the calculated kinetic behaviour of a similar molecular cluster based on a theoretical model, in a manner similar to the isolated fluorophore calibration data. Those trajectories whose kinetic behaviours are found to lead to the greatest deviations from the theoretical model are then re ⁇ evaluated to result in a smaller deviation from the model. This process is repeated recursively until a satisfactory agreement or convergence is reached.
- individual molecule trajectories are used for simulating a kinetic behaviour of a molecular cluster compatible with the detection data extracted from at least a portion of the acquired frames.
- the simulated data are then compared to the analysed detection paths with the simulated kinetic behaviour of a molecular cluster to correct errors in the analysed detection paths.
- This comparative step may lead to detection of counting errors due to over-joining of very short tracks
- performing calibration routines on known molecular systems enables optimising detection thresholds for investigating unknown molecular systems, for which calibrated data may not be available of may have a poor accuracy.
- a Euclidean distance may be imposed between molecules for consideration as a multimer molecule and the stoichiometry of molecular clusters may be determined.
- Other techniques such as DBSCAN, may be used to segment the trajectories into clusters. Determining the overall underlying stoichiometry in a sample is a difficult task for expected clusters of only a few molecules.
- V (N sub ) The behaviour of V (N sub ) as N sub is incremented from 1 to N is characteristic of the underlying stoichiometry of the sample. The position and amplitude of a peak in this function may be used to determine the underlying
- Samples with mixed cluster stoichiometry may also have characteristic V (N sub ) curves. However, these distributions can be very similar to the curves of a different number of binding sites of a stoichiometry between the two in the mixed sample. In these situations the variance may not be sufficient to adequately describe the clustering
- a third order moment (the skew) is used as a further indicator of underlying stoichiometry.
- FIG. 5 there is shown a map 500 of a molecular data set simulated in accordance with
- the data set shows the simulated positions and stoichiometry of clusters of molecules across the area of a sample. Although the simulated data set in this figure is of dimers a
- FIG. 6 shows a map 600 similar to map 500 calculated in accordance with an embodiment of the method described above for the same data set as for map 500.
- the calculated position and number of clusters generated by the method can be related to the simulated data 500 for optimisation of calculation parameters.
- Figure 6 shows the output of an embodiment of the method obtained with an input similar to the input used for the simulation of figure 5.
- Figure 6 shows the apparent stoichiometry of each cluster in the data set.
- circles represent clusters in which one fluorophore trajectory was assigned, crosses where two trajectories were assigned, squares where three clusters assigned, and diamonds where four clusters were assigned.
- Map 700 is comparison of the ground truth shown in figure 5 and the output of the embodiment of the method shown in figure 6. Circles indicate clusters where the assignment of trajectories is correct and the apparent cluster stoichiometry in figure 5 matches that of
- FIG. 8 (a) there is shown a histogram 800 generated from the true 802 and calculated 804 apparent cluster stoichiometries for the data
- Figure 8 (b) is a plot 850 showing the standard deviation of cluster data as a function of sampling fraction.
- Plot 850 shows the curves of the standard deviation of the apparent cluster stoichiometry distribution (which follows the zero-truncated binomial distribution) as a function of sampling (detection) fraction for underlying cluster stoichiometries of monomers through to decamers (10-mers).
- the curve for monomers is a straight line at a value of 0, with increasing values at a given sampling fraction for increasing underlying cluster stoichiometry. These curves do not intersect at points greater than 0 and less than 1, exclusive, and are independent of the number of clusters or molecules. As such a curve generated from experimental data as described above is expected to fall along one of these curves which can be used for visualization and identification of the underlying cluster stoichiometry at a range of sampling fractions and when the sampling fraction is unknown a priori.
- Figure 9 shows a plot 900 with experimental data (solid lines) and fits (dashed lines) for simulated data
- Embodiments of the methods disclosed herein can be carried out on a computer system.
- the computer system may be interconnected with a high
- the computer system 901 may be a high performance machine, a desktop workstation or a personal computer, or may be a portable computer such as a laptop or a notebook or may be a distributed computing array or a computer cluster or a networked cluster of computers.
- the computer system 901 also comprises a suitable
- the computer system 901 comprises one or more data
- CPUs 902 processing units 902
- memory 904 which may include volatile or non volatile memory, such as various types of RAM memories, magnetic discs, optical disks and solid state memories
- user interface 906 which may comprise a monitor, keyboard, mouse and/or touch-screen display
- network or other communication interface 908 for example
- Memory 904 comprises program code to execute computer programs in accordance with embodiments of the methods described herein.
- the computer system 901 may also be connected directly to microscope 912 to download data.
- the computer system 901 may also access data stored in a remote database 914 via network interface 908.
- Remote database 914 may be a distributed database.
- a computer system for implementing embodiments of the invention is not limited to the computer system described in the preceding paragraphs. Any computer system
- the architecture may comprise client/server architecture, or any other architecture.
- the software for implementing embodiments of the invention may be processed by "cloud" computing architecture .
- the method, system and apparatus described herein are not limited to the detection of fluorophores in a luminescence experiment, but may be used in any circumstance where individual members of a population can be assembled into a list of clusters for sub-sampling.
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Signal Processing (AREA)
- Dispersion Chemistry (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Analytical Chemistry (AREA)
- Biochemistry (AREA)
- General Health & Medical Sciences (AREA)
- General Physics & Mathematics (AREA)
- Immunology (AREA)
- Pathology (AREA)
- Investigating, Analyzing Materials By Fluorescence Or Luminescence (AREA)
Abstract
A method and an apparatus for counting fluorescent molecules in a sample are disclosed in this specification. The disclosed method and apparatus allows counting fluorescent molecules and derive molecules trajectories based on spatial and temporal information. Spatial and temporal information are processed using mathematical models and variable thresholds to extract a reliable number of fluorescent molecules and their trajectories in a cluster in a regime where the activation rate is not constant in time. The system and methods disclosed herein are non-destructive for the sample as they are based on the acquisition of a sequence of optical images. They may be used to identify the presence and number of target molecules in situ, within a cell for example.
Description
A METHOD AND A SYSTEM FOR IDENTIFYING MOLECULES
FIELD OF THE INVENTION
The present invention relates to microscopy methods and systems for identifying molecules, and particularly, but not exclusively to a system and a method for locating and counting molecules.
BACKGROUND OF THE INVENTION
Microscopy techniques are generally used to investigate macromolecular structures at intermediate and high
resolution. For example, electron microscopy (EM), cryo-EM and crystallography are used to examine the molecular structure of proteins, but these techniques often involve elaborate sample preparation or require isolating the protein complexes from their natural environment.
Optical microscopy techniques with fluorescent molecules are used as a less invasive alternative. However, the resolution of conventional light microscopy is limited by the diffraction limit of light to approximately 250 nm. This is generally larger than the size of protein
complexes .
Photo-activated localization microscopy (PALM) and related techniques aim to enable imaging and counting fluorescent molecules spatially closer than the diffraction limit. However, errors related to the fluorescent blinking, photo-physics and photo-bleaching behaviour make this technique unreliable when quantitative data are required.
There is a need for improved microscopy techniques for imaging and counting fluorescent molecules.
SUMMARY OF THE INVENTION
In accordance with the first aspect, the present invention provides a method for counting fluorescent molecules in a sample comprising the steps of:
obtaining detection data from a plurality of image frames of a sample, the detection data regarding timing and location of fluorescent events;
dividing detection data into spatial regions of interest based on location of fluorescent events;
linking detection data in a plurality of detection paths in a region of interest based on location and timing of fluorescent events in the region of interest;
analysing the detection paths to determine whether each detection path is related to a new fluorescent molecule or a new fluorescent event of an existing
fluorescent molecule. In an embodiment, the method further comprises the step of associating detection paths corresponding to fluorescent events of existing fluorescent molecules into detection trajectories; each trajectory representing an exclusive association of detection events of a single fluorescent molecule.
In an embodiment, the step of dividing detection data into spatial regions of interest comprises calculating a spatial distance between two fluorescent events and comparing the calculated distance with a predetermined
distance. A modified version of a DBSCAN algorithm may be used to perform the step of dividing detection data into spatial regions of interest. In embodiments, a plurality of spatial distances are calculated to compose at least one region of interest comprising at least a portion of the sample surface.
In embodiments, the step of linking detection data in a plurality of detection paths further comprises the step of setting a variable metric to measure a distance between fluorescent events occurring in different frames. The variable metric may be a mathematical function of the Hellinger Distance. The variable metric may also be a probabilistic metric related to one or more probabilistic distributions of the fluorescent events. Further, the different frames may be consecutive frames. Each of the one or more probabilistic distributions may be used to define a position and a localisation precision of a fluorescent event.
In some embodiments, the method further comprises the step of comparing the metric to a score threshold to define if two fluorescent events are linked. The score threshold may vary for different samples and may be related to dye photo-physical kinetics.
In an embodiment, the method further comprises the step of organising linked fluorescent events in short tracks of fluorescent events.
In an embodiment, the method comprises the step of
calculating a histogram of the number of short tracks for each frame. In an embodiment, the method further comprises the step of calculating a number of new particles born in each frame based on the histogram of the frame.
In an embodiment, the method further comprises the step of calculating the probability that a new fluorescent
molecule would stochastically be born in an area proximal to an existing fluorescent molecule.
In an embodiment, the method further comprises the step of calculating a likelihood of a short track being the result of a new molecule arising at a given time in a given location .
In embodiments, the method further comprises the step of calculating a likelihood of a short track being the result of the same molecule. The likelihood may be calculated simultaneously for all the short tracks in a given region of interest. A Munkres algorithm may be used to associate short tracks to existing molecules of new molecules arising at a given time in a given location.
In an embodiment, the method further comprises the step of calculating a plurality of trajectories related to a plurality of fluorescent molecules by pairing short tracks.
In an embodiment, the method further comprises the step of simulating a kinetic behaviour of a molecular cluster compatible with the detection data extracted from at least a portion of the acquired frames and comparing the
calculated plurality of trajectories with the simulated kinetic behaviour of a molecular cluster to correct errors .
In an embodiment, the method further comprises the step of calculating a kinetic behaviour of a molecular cluster compatible with the detection data extracted from at least a portion of the acquired frames based on a theoretical model and comparing the calculated plurality of
trajectories with the calculated kinetic behaviour of a molecular cluster to correct errors.
In accordance with the second aspect, the present
invention provides a data processing apparatus comprising a plurality of operating modules arranged for carrying out a method in accordance with the first aspect.
In accordance with the third aspect, the present invention provides a computer program, comprising instructions for controlling a computer to perform a method in accordance with the first aspect.
In accordance with the fourth aspect, the present
invention provides a computer readable medium, providing the computer program of the third aspect.
In accordance with the fifth aspect, the present invention provides a data signal, providing a computer program in accordance with the third aspect. In accordance with the sixth aspect, the present invention provides an apparatus for counting fluorescent molecules in a sample, from fluorescence detection data regarding timing and location of fluorescent events obtained from a plurality of image frames of the sample, the apparatus comprising;
a splitting module for dividing the detection data into spatial regions of interest based on location of fluorescent events;
a linking module for linking the split detection data in a plurality of detection paths in a region of interest based on location and timing of fluorescent events in the region of interest;
a processing module for analysing the detection paths and determine whether each detection path is related to a new fluorescent molecule or a new fluorescent event of an existing fluorescent molecule.
In an embodiment, the apparatus further comprises an associator for associating detection paths corresponding to fluorescent events of existing fluorescent molecules into detection trajectories; each trajectory representing an exclusive association of detection events of a single fluorescent molecule. In an embodiment, the apparatus further comprises a microscopy unit arranged to acquire image frames of a
biological sample and an analyser unit arranged to extract detection data related to timing and location of
fluorescent events. In accordance with the seventh aspect, the present
invention provides a method for counting fluorescent molecules in a sample comprising the steps of:
acquiring images of the sample showing the
fluorescent molecules;
obtaining from images of the sample showing the fluorescent molecules, detection data regarding timing of fluorescence and location of fluorescence;
comparing the timing detection data with the location detection data to identify a number and location of the fluorescent molecules.
In accordance with the eighth aspect, the present
invention provides an apparatus for counting fluorescent molecules in a biological sample, the apparatus
comprising:
a microscopy unit arranged to acquire image frames of a biological sample;
an image analyser unit arranged to extract detection data related to timing and location of fluorescent events; a comparator arranged to compare the timing detection data with the location detection data to identify a number and location of the fluorescent molecules.
Advantageous embodiments provide microscopy methods and systems for counting molecules based on spatial and temporal information. Spatial and temporal information are processed using mathematical models and variable
thresholds to extract a reliable number of fluorophores and their trajectories in a cluster in a regime where the activation rate of the clusters is not constant in time.
Advantageous embodiments may be used to identify the presence and number of target molecules in situ, within a cell for example. The system and methods disclosed herein are non-destructive for the sample as they are based on the acquisition of a sequence of optical images.
BRIEF DESCRIPTION OF THE DRAWINGS In order to fully understand the present invention, embodiments of the present invention will now be
described, by way of example only, with reference to the accompanying drawings, in which:
Figure 1 shows a common photo-conversion scheme of photo- activatable fluorescent proteins;
Figures 2 and 3 shows flow-charts with steps for counting fluorescent molecules in a sample in accordance with embodiments ;
Figure 4 schematically illustrates the linking of short tracks into trajectories in accordance with embodiments;
Figures 5 to 7 show maps of a simulated data set, a calculated data set and a comparison of true and
calculated cluster stoichiometry respectively in
accordance with embodiments; Figure 8 shows a fit of the calculated cluster data to a probability distribution and the variance of cluster data
as a function of sampling fraction in accordance with embodiments ;
Figure 9 shows a plot with experimental data and simulated data in accordance with embodiments; Figure 10 is a schematic illustration of a computer system suitable to implement a method in accordance with
embodiments .
DETAILED DESCRIPTION OF EMBODIMENTS
Embodiments of the present invention relate to microscopy methods and systems for counting molecules. Some
embodiments provide techniques to utilize fluorophore photophysics and the spatiotemporal localization to quantify cluster numbers in a serial localization
microscopy domain. The approach used allows tracking and deriving
trajectories of fluorescent molecule using spatial and temporal information extracted from high resolution images of samples. The process is based and operates on detection data of fluorescent events extracted from the high
resolution images. The process does not operate on image data as such.
Referring now to figure 1, there is shown a common photo- conversion scheme 100 of photo-activatable fluorescent proteins, such as mEos. The photo-activatable fluorescent proteins may be in a xpre-activation' state 102, an xactive' state 104, a xbleached state' 106 or two
different dark states, Mark 1' 108 and Mark 2' 110.
The transitions from pre-activation 102 to active 104, and active 104 to bleached 106 are considered to be
irreversible and have respective transition rates ka and kb. The transitions from active 104 to dark 1 108 or dark 2 110 are considered to be reversible and have respective transition rates kd and akd. a is a scaling factor which biases the transition probability towards one dark state over the other. The transition rates from dark 1 108 and dark 2 110 to active 104 are instead different and defined respectively as krl and kr2.
In embodiments, values of kb, kd, a, kri and kr2 estimated by measurements of a known sample of monomeric fluorophores are used. The measurements are performed in the same conditions as per the molecule counting process. In particular, the same incident laser intensity is used.
Referring now to figure 2, there is shown a flow-chart 200 with steps for counting fluorescent molecules in a sample in accordance with embodiments. At step 205 a plurality of image frames of a sample are acquired using an image acquisition system, such as a fluorescence optical
microscope with an objective lens of high numerical aperture and a camera detector.
The data are processed at step 210 to extract information related to timing and spatial location of fluorescent events. In some embodiments these information are
collected in a data file which can be processed by a computing system.
The detection data are then divided into spatial regions of interest at step 220 based on the location of
fluorescent events. The splitting of the data in regions of interest (ROIs) is performed to optimise the analysis of the data using performing fitting algorithms and ultimately improve molecule detection. The splitting process allows reducing the number of detection events that must be considered simultaneously. In advantageous embodiments the division is performed using a recursive approach .
Starting from the first detection event in the data file, other detection events within a given distance from the first detection event location are added to a list of localizations for a given detection region. The survey region is then expanded to that given distance from the new additions to the detection region list, capturing more detection events. The process is repeated until there are no detection events within the given distance to add to the list in an iteration. The ROI created is stored and the related detection events are removed from the data file. The process restarts from the new first detection event in the data file to create another ROI.
After the detection data are divided into spatial ROIs at step 220, each ROI is considered to be exclusive from all other regions of interest and events from one region cannot be linked to events in another region. Once the ROIs are identified, the detection events in each ROI are linked in a plurality of detection paths in a ROI at step 225. The detection events are linked based on both their location and timing.
The data related to detection paths obtained by linking detection events in each ROI are analysed at step 230 to determine whether each detection path is related to a new fluorescent molecule or a new fluorescent event of an existing fluorescent molecule. The number of molecules in each ROI is added to obtain the number of molecules in the sample .
Referring now to figure 3, there is shown a flow-chart 300 listing method steps to perform the counting algorithm in accordance with embodiments. The flow-chart 300 is based on the chart 200 of figure 2 and adds further details.
Calibration data 305 is collected by analysis of data acquired in a similar manner as the image data from a sample of known form of isolated molecules. Analysis of the signal intermittency and total length for individual molecules the kinetics constants kb, kd, krl, kr2, and a are determined .
After the detection data are divided into spatial ROIs at step 220, as explained above with reference to figure 2, the detection events in each ROI are first linked into short tracks of events occurring in consecutive frames at step 310. Starting with the earliest detection event in a region, at step 310, the algorithm monitors detection events occurring in the next frame. In case an event is found, Gaussian distributions related to the two detection events are calculated together with a mathematical
distance 315 which separates the two distributions. The position of each detection event is the mean of the respective Gaussian distribution and the localization precision is the covariance of the distributions. The
distance metric at 315 is compared against a score
threshold (ST) 320 which defines a threshold for detection events in consecutive frames to be linked. The value of ST may be varied in accordance with the requirements in specific experiments. In advantageous embodiments the value of ST is based on the dye kinetics yields.
The kinetics-based value of ST is given by the probability that a molecule transition away from an active state in a single frame.
In the equation tf is the time between consecutive frames. If the distance metric exceeds the ST value, the two detection events are linked into a short track. In the case where more than one detection event may occur in either the current or next frame then the distance metric and score threshold values are concatenated into an appropriate matrix and a global minimization assignment algorithm executed to pair current detection events with either future detection events or the ST. Those
assignments paired with future detection events are associated together as a part of the respective short track; those assignments paired with the ST are taken as terminal events in the track and not associated with later events . The status of new active molecule or part of a short track is assigned to each detection event in the consecutive frame, then the algorithm moves to the next frame and the process is repeated until all detection events in a given region are assigned to short tracks or as a track end.
In advantageous embodiments, the distance metric 315 is the square root of 1 minus the Hellinger Distance ((1- DH)1/2) and can be calculated according to the equation below based on the multivariate Bhattacharyya Coefficient.
Once all the short tracks are calculated for each ROI, detection events within a single short track are
aggregated into a single detection event. The new
positions of the new aggregated events and their accuracy are calculated by weighted averages accounting for the number of photons detected for each short track and the number of events in the short track.
' V - ..„^ * V
In the equations x and y± are the positions of the detection events and N± the number of photons detected for each event i in the short track. The spatial uncertainty of the localisation is given by σ. σ± and NEvents correspond to the uncertainty of localization i in the short track and the number of events in the track, respectively. The overall uncertainty, σ, is taken as the greater of σχ and Oy (equivalent to σχ given above) . The greater of the two uncertainties is carried forward.
After a list of short tracks, each with a start frame, end frame, position, and uncertainty has been collected at 310, 315, 320, a histogram of the number of short tracks starting from each frame is generated. The bin width is specified as an algorithm input. For example, a bin width of 1000-2000 frames may be appropriate for a dataset of 25000 frames. Scaling this histogram provides a smoothed approximation of the number of new particles born in a given frame, 330. This approximation may be a slight over- estimate of the actual probability as some of the tracks in some frames are continuations of other short tracks rather than true new molecules. The accuracy of the estimate is related to the kinetics of the fluorophores , namely the number of occurrences of a molecule returning from a dark state 108 or 110 to an on state 104 before transitioning to a bleach state 106.
The probability of a new molecule arriving in a given position in a given frame (PWew) is the product of the probability of a new molecule arising in that frame and the fraction of the total experimental area under
consideration (PSFArea) .
The fractional area for randomly-distributed molecules is taken as the area occupied by a circle with a radius equal to the sum of the objective PSFs of the two molecules that could be linked, divided by the area of the entire
experimental field of view. For non-randomly-distributed molecules the fractional area is taken as a fraction of the intensity of the image generated from the sum of all frames in the data set that is present within a given distance from the detection event. In advantageous
embodiments this distance is a circle with the radius of one objective PSF.
The product of the area fraction PSFArea and PNew yields a score SNew which is a measure of the likelihood of a short track being the result of a new molecule arising at a given time in a given location.
The complementary score SLink is a measure of the likelihood of a short track being related to the same molecule and is calculated at 335. SLink is the product of three values: the overlap score between short track positions, {1-DH)1/2 , the off time score Soff, and the dark state probability, Pdark-
S0ff is proportional to the probability that a molecule in the dark state will return to the active state after a dark state dwell time of At. S0ff is calculated in
accordance with the equations below. The values for krl, kr2r and a are obtained from the results of calibration experiments with the same fluorescent molecule as
J
The dark state probability, Pdarkr is the probability that molecule in the active state will transition to a dark state rather than to the bleach state. Pdark is calculated according to the following equation.
The values for k^, and are obtained from the results of calibration experiments with the same fluorescent molecule as described previously. The product of (1-DH)1 2, S0ff and Pdark generated at 335 is compared in an inequality with the product of PSFArea and PNeu of 330 to ascertain whether a short track originating in the spatial and temporal proximity of the end of another short track is more likely due to a molecule returning from a dark state to the active state or from the activation of a new molecule. In practice there may be many molecules in a single region activating, blinking, and bleaching simultaneously. This problem may be addressed by considering all short tracks built from detection events in a single ROI . A matrix of SLink values is determined for each possible pairings between a short track and all other tracks sharing as origination frame subsequent to the final frame the short track .
Referring now to figure 4 there is shown a line diagram 400 that schematically illustrates the linking of short tracks (solid lines) into trajectories by bridging off gaps between valid short track termini pairs (dashed and dotted lines) . Scores for each valid gap are assembled into a matrix to be passed to an optimal assignment algorithm. Those associations found to be valid are used to assemble trajectories from short tracks and gaps, here shown as dashed lines. The structure of a possible score matrix is provided below.
Tuck J
Scores are indicated by S±,j and are generated for all valid parings (or where a short track start is followed by a short track end) taking into account spatial and
temporal information. For each track a score SNew is also calculated as a metric for that track not linking to a prior track. This matrix is entered into an optimal assignment algorithm and the resulting assignments used to create trajectories from short tracks. The associations of figure 4 show termini of short tracks (solid lines) in which a short track start comes
subsequent to another short track end. These are indicated as valid potential associations by a dashed or dotted line. The SLink values for each of these potential
associations is used to generate the matrix.
SNew values are considered for the later short tracks in a further matrix appended to the matrix of SLink values. In embodiments, the optimal assignment is determined via Munkres algorithm which enforces the constraint that each short track origination can be paired with either a single short track end or the null SNew option, and each short track end can be paired with at most a single short track start .
In the matrix shown above, the elements indicated by S±fj designate a valid potential pairing between the end of
short track i and the start of short track j (this element will have the value of S∑ink for this pairing) . The elements indicated by SNeH/j designate the value of SNew for that particular short track. The trajectories of each molecule are determined at 340 by pairing short tracks.
In embodiments, the kinetic behaviour of the analysed individual molecule trajectory data from at least a portion of the acquired frames and at least a portion of the analysed trajectories are compared to the calculated kinetic behaviour of a similar molecular cluster based on a theoretical model, in a manner similar to the isolated fluorophore calibration data. Those trajectories whose kinetic behaviours are found to lead to the greatest deviations from the theoretical model are then re¬ evaluated to result in a smaller deviation from the model. This process is repeated recursively until a satisfactory agreement or convergence is reached.
In other embodiments, individual molecule trajectories are used for simulating a kinetic behaviour of a molecular cluster compatible with the detection data extracted from at least a portion of the acquired frames. The simulated data are then compared to the analysed detection paths with the simulated kinetic behaviour of a molecular cluster to correct errors in the analysed detection paths.
This comparative step may lead to detection of counting errors due to over-joining of very short tracks,
especially those of a single frame in length, or under- joining. In some embodiments, performing calibration
routines on known molecular systems enables optimising detection thresholds for investigating unknown molecular systems, for which calibrated data may not be available of may have a poor accuracy. A Euclidean distance may be imposed between molecules for consideration as a multimer molecule and the stoichiometry of molecular clusters may be determined. Other techniques such as DBSCAN, may be used to segment the trajectories into clusters. Determining the overall underlying stoichiometry in a sample is a difficult task for expected clusters of only a few molecules. It is typical for a significant portion of the molecules to go undetected during the course of an experiment; this portion is unknown, increases during the course of an experiment, and extremely cumbersome to control. Here we contrast the underlying stoichiometry, or the true stoichiometry of the sample of interest, against the apparent stoichiometry, or the stoichiometry of the sample as has been measured in an experiment. The complexity of using the mean value to determine stoichiometry of molecules in clusters and complexes increases linearly as more molecules are detected. In advantageous embodiments the standard deviation (or variance) of the detected stoichiometry distribution is used.
If for example a perfectly labelled sample is considered, where all fluorophores are correctly detected. The
molecules represent all potential binding sites (of number N) and can be assembled into M clusters based on spatial
criteria as described above. In each step of the analysis a subset of Nsub molecules is randomly chosen from the master list of molecules. The Nsub molecules are assembled into Msub clusters as before and the distribution of apparent stoichiometries in those Msub clusters determined and the variance of this distribution is given as V (Nsub) . The behaviour of V (Nsub) as Nsub is incremented from 1 to N is characteristic of the underlying stoichiometry of the sample. The position and amplitude of a peak in this function may be used to determine the underlying
stoichiometry of the sample as well as the number of potential binding sites within that sample.
For example, for a sample of dimers, the maximum amplitude of V(Nsub) = 0,25 at a position of Npeak/N = 2/3 . With increasing underlying stoichiometry the peak amplitude increases and Npeak/N moves towards 0.5.
Samples with mixed cluster stoichiometry may also have characteristic V (Nsub) curves. However, these distributions can be very similar to the curves of a different number of binding sites of a stoichiometry between the two in the mixed sample. In these situations the variance may not be sufficient to adequately describe the clustering
distribution behaviour.
In embodiments, a third order moment (the skew) is used as a further indicator of underlying stoichiometry. The distributions consisting of a single stoichiometry
generally have a skew function Sk (Nsub) that decreases with a sharply decrease in slope as Nsub → N. In a mixed
stoichiometry distribution Sk(Nsub) generally does not show this sharp decrease or may increase, depending on the
nature of the stoichiometry. With this additional
information it is sometimes possible to separate, for example, a mixture of dimers and tetramers from that of a sample of trimers . Referring now to figure 5 there is shown a map 500 of a molecular data set simulated in accordance with
theoretical molecular models. The data set shows the simulated positions and stoichiometry of clusters of molecules across the area of a sample. Although the simulated data set in this figure is of dimers a
significant portion of the fluorophores will not be activated during the course of the experiment and as such the apparent stoichiometry of the resulting clusters will vary according a zero-truncated binomial distribution with variables of underlying stoichiometry and detection fraction (Nsub above) . In figure 5 the crosses represent clusters where both fluorophores were activated in the simulation while circles represent clusters in which only one of the two fluorophores was activated in the
simulation. The analysis method facilitates taking the simulated detection event output of the simulation and reassembles those events into trajectories with the same number of apparent fluorophores in each corresponding cluster as shown in figure 5. Figure 6 shows a map 600 similar to map 500 calculated in accordance with an embodiment of the method described above for the same data set as for map 500. The calculated position and number of clusters generated by the method can be related to the simulated data 500 for optimisation of calculation parameters.
Figure 6 shows the output of an embodiment of the method obtained with an input similar to the input used for the simulation of figure 5. Figure 6 shows the apparent stoichiometry of each cluster in the data set. Here circles represent clusters in which one fluorophore trajectory was assigned, crosses where two trajectories were assigned, squares where three clusters assigned, and diamonds where four clusters were assigned.
Referring now to figure 7, there is shown a comparison map 700 of simulated data and calculated cluster
stoichiometry. Markers show under-counting and overcounting errors. Map 700 is comparison of the ground truth shown in figure 5 and the output of the embodiment of the method shown in figure 6. Circles indicate clusters where the assignment of trajectories is correct and the apparent cluster stoichiometry in figure 5 matches that of
corresponding cluster in figure 6. Diamonds and squares in figure 7 indicate undercounting and overcounting errors, respectively. An undercounting error corresponds to a cluster where the number of assigned trajectories in a cluster is less than the true apparent cluster
stoichiometry; an overcounting error corresponds to a cluster in which the number of assigned trajectories is greater than the true apparent cluster stoichiometry. Referring now to figure 8 (a) , there is shown a histogram 800 generated from the true 802 and calculated 804 apparent cluster stoichiometries for the data
corresponding to figures 5 and 6. In addition a linear fit (black line and circles) of the calculated
stoichiometry to a zero-truncated binomial distribution
with a set stoichiometry of dimers is shown. This figure shows that the assigned cluster stoichiometries follow well the distribution as expected for dimers and with a fraction of detected molecules close to that as produced by the simulation.
Figure 8 (b) is a plot 850 showing the standard deviation of cluster data as a function of sampling fraction. Plot 850 shows the curves of the standard deviation of the apparent cluster stoichiometry distribution (which follows the zero-truncated binomial distribution) as a function of sampling (detection) fraction for underlying cluster stoichiometries of monomers through to decamers (10-mers). The curve for monomers is a straight line at a value of 0, with increasing values at a given sampling fraction for increasing underlying cluster stoichiometry. These curves do not intersect at points greater than 0 and less than 1, exclusive, and are independent of the number of clusters or molecules. As such a curve generated from experimental data as described above is expected to fall along one of these curves which can be used for visualization and identification of the underlying cluster stoichiometry at a range of sampling fractions and when the sampling fraction is unknown a priori.
In some embodiments, a combination of multiple
distributions can be used to fit to the experimental data. Figure 9 shows a plot 900 with experimental data (solid lines) and fits (dashed lines) for simulated data
generated in the same manner as for figures 5 to 7 but with underlying cluster stoichiometries of monomers (902), dimers (904), trimers (906) and tetramers (908). Results
from three data sets at each condition are shown. Here the fit is taken to be a combination of two standard deviation vs sampling fraction curves. In each of the curves the data points related to a given stoichiometry condition are very close to one another and separated from the other curves, providing a manner in which to identify the underlying cluster stoichiometry in a rapid and accurate manner. As the number of molecules detected is significantly smaller than the underlying number of molecules the experimental curves stop at a position characteristic of a sampling fraction of approximately 2/3. This information can furthermore be used to
calculate the number of underlying molecules, even though a significant portion is not detected during the
experiment.
Embodiments of the methods disclosed herein can be carried out on a computer system. In advantageous embodiments the computer system may be interconnected with a high
resolution microscopy unit. An example of computer system which may be used to implement the methods is schematised in figure 10. The computer system 901 may be a high performance machine, a desktop workstation or a personal computer, or may be a portable computer such as a laptop or a notebook or may be a distributed computing array or a computer cluster or a networked cluster of computers.
The computer system 901 also comprises a suitable
operating system and appropriate software for
implementation of embodiments of the present invention. The computer system 901 comprises one or more data
processing units (CPUs) 902; memory 904, which may include volatile or non volatile memory, such as various types of
RAM memories, magnetic discs, optical disks and solid state memories; a user interface 906, which may comprise a monitor, keyboard, mouse and/or touch-screen display; a network or other communication interface 908 for
communicating with other computers as well as other devices; and one or more communication busses 910 for interconnecting the different parts of the system 901. Memory 904 comprises program code to execute computer programs in accordance with embodiments of the methods described herein.
The computer system 901 may also be connected directly to microscope 912 to download data. The computer system 901 may also access data stored in a remote database 914 via network interface 908. Remote database 914 may be a distributed database.
A computer system for implementing embodiments of the invention is not limited to the computer system described in the preceding paragraphs. Any computer system
architecture may be utilised, such as standalone
computers, networked computers, dedicated computing devices, handheld devices or any device capable of receiving processing information in accordance with embodiments of the present invention. The architecture may comprise client/server architecture, or any other architecture. The software for implementing embodiments of the invention may be processed by "cloud" computing architecture .
The method, system and apparatus described herein are not limited to the detection of fluorophores in a luminescence experiment, but may be used in any circumstance where
individual members of a population can be assembled into a list of clusters for sub-sampling.
It is to be understood that, if any prior art publication is referred to herein, such reference does not constitute an admission that the publication forms a part of the common general knowledge in art, in Australia or any other country .
In the claims which follow and in the preceding
description of the invention, except where the context requires otherwise due to express language or necessary implication the word "comprise" or variations such as "comprises" or "comprising" is used in an inclusive sense i.e. to specify the presence of the stated features but not to preclude the presence or addition of further features in various embodiments of the invention.
It will be appreciated by persons skilled in the art that numerous variations and/or modifications may be made to the invention as shown in the specific embodiments without departing from the spirit or scope of the invention as broadly described. The present embodiments are,
therefore, to be considered in all respects as
illustrative and not restrictive.
Claims
1. A method for counting fluorescent molecules in a sample comprising the steps of:
obtaining detection data from a plurality of image frames of a sample, the detection data regarding timing and location of fluorescent events;
dividing detection data into spatial regions of interest based on location of fluorescent events;
linking detection data in a plurality of detection paths in a region of interest based on location and timing of fluorescent events in the region of interest;
analysing the detection paths to determine whether each detection path is related to a new fluorescent molecule or a new fluorescent event of an existing
fluorescent molecule.
2. The method of claim 1 further comprising the step of associating detection paths corresponding to fluorescent events of existing fluorescent molecules into detection trajectories; each trajectory representing an exclusive association of detection events of a single fluorescent molecule .
3. The method of claim 1 or claim 2 wherein the step of dividing detection data into spatial regions of interest comprises calculating a spatial distance between two fluorescent events and comparing the calculated distance with a predetermined distance.
4. The method of claim 3 further comprising the step of using a modified version of a DBSCAN algorithm is used to
perform the step of dividing detection data into spatial regions of interest.
5. The method of claim 3 or claim 4 wherein a plurality of spatial distances are calculated to compose at least one region of interest comprising at least a portion of the sample surface.
6. The method of any one of the preceding claims wherein the step of linking detection data in a plurality of detection paths further comprises the step of setting a variable metric to measure a distance between fluorescent events occurring in different frames.
7. The method of claim 6 wherein the variable metric is a mathematical function of the Hellinger Distance.
8. The method of claim 6 or 7 wherein the different frames are consecutive frames.
9. The method of any one of claims 6 to 8 wherein the variable metric is a probabilistic metric and is related to one or more probabilistic distributions of the
fluorescent events.
10. The method of claim 9 wherein each of the one or more probabilistic distributions defines a position and a localisation precision of a fluorescent event.
- so
il. The method of any one of claims 6 to 10 further comprising the step of comparing the metric to a score threshold to define if two fluorescent events are linked.
12. The method of claim 11 wherein the score threshold varies for different samples and is related to dye photo- physical kinetics.
13. The method of claim 11 or claim 12 further comprising the step of organising linked fluorescent events in short tracks of fluorescent events.
14. The method of claim 13 further comprising the step of calculating a histogram of the number of short tracks for each frame.
15. The method of claim 14 further comprising the step of calculating a number of new particles born in each frame based on the histogram of the frame.
16. The method of any one of claims 13 to 15 further comprising the step of calculating the probability that a new fluorescent molecule would stochastically be born in an area proximal to an existing fluorescent molecule.
17. The method of any one claims 13 to 16 further
comprising the step of calculating a likelihood of a short track being the result of a new molecule arising at a given time in a given location.
18. The method of any one claims 13 to 16 further
comprising the step of calculating a likelihood of a short track being the result of the same molecule.
19. The method of claim 18 wherein the likelihood is calculated simultaneously for all the short tracks in a given region of interest.
20. The method of any one of claims 17 to 19 wherein a Munkres algorithm is used to associate short tracks to existing molecules of new molecules arising at a given time in a given location.
21. The method of any one of claims 17 to 20 further comprising the step of calculating a plurality of
trajectories related to a plurality of fluorescent
molecules by pairing short tracks.
22. The method of claim 21 further comprising the step of simulating a kinetic behaviour of a molecular cluster; the molecular cluster being compatible with the detection data extracted from at least a portion of the acquired frames; and
comparing the calculated plurality of trajectories with the simulated kinetic behaviour of a molecular cluster to correct errors.
23. The method of claim 21 further comprising the step of calculating a kinetic behaviour of a molecular cluster based on a theoretical model; the molecular cluster being
compatible with the detection data extracted from at least a portion of the acquired frames; and comparing the calculated plurality of trajectories with the calculated kinetic behaviour of a molecular cluster to correct errors.
24. A data processing apparatus comprising a plurality of operating modules arranged for carrying out the method of any one of claims 1 to 23.
25. A computer program, comprising instructions for controlling a computer to perform the method of any one of claims 1 to 23.
26. A computer readable medium, providing the computer program of claim 25.
27. A data signal, providing a computer program in accordance with claim 25.
28. An apparatus for counting fluorescent molecules in a sample, from fluorescence detection data regarding timing and location of fluorescent events obtained from a
plurality of image frames of the sample, the apparatus comprising;
a splitting module for dividing the detection data into spatial regions of interest based on location of fluorescent events;
a linking module for linking the split detection data in a plurality of detection paths in a region of interest
based on location and timing of fluorescent events in the region of interest;
a processing module for analysing the detection paths and determine whether each detection path is related to a new fluorescent molecule or a new fluorescent event of an existing fluorescent molecule.
29. The apparatus of claim 28 further comprising an associator for associating detection paths corresponding to fluorescent events of existing fluorescent molecules into detection trajectories; each trajectory representing an exclusive association of detection events of a single fluorescent molecule.
30. The apparatus of claim 29 further comprising a microscopy unit arranged to acquire image frames of a biological sample and an analyser unit arranged to extract detection data related to timing and location of
fluorescent events.
31. A method for counting fluorescent molecules in a sample comprising the steps of:
acquiring images of the sample showing the
fluorescent molecules;
obtaining from images of the sample showing the fluorescent molecules, detection data regarding timing of fluorescence and location of fluorescence;
comparing the timing detection data with the location detection data to identify a number and location of the fluorescent molecules.
32. An apparatus for counting fluorescent molecules in a biological sample, the apparatus comprising:
a microscopy unit arranged to acquire image frames of a biological sample;
an image analyser unit arranged to extract detection data related to timing and location of fluorescent events; a comparator arranged to compare the timing detection data with the location detection data to identify a number and location of the fluorescent molecules.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| AU2014901336 | 2014-04-11 | ||
| AU2014901336A AU2014901336A0 (en) | 2014-04-11 | A Microscopy Technique |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2015154137A1 true WO2015154137A1 (en) | 2015-10-15 |
Family
ID=54287018
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/AU2015/000216 Ceased WO2015154137A1 (en) | 2014-04-11 | 2015-04-10 | A method and a system for identifying molecules |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2015154137A1 (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113223632A (en) * | 2021-05-12 | 2021-08-06 | 北京望石智慧科技有限公司 | Molecular fragment library determination method, molecular segmentation method and device |
-
2015
- 2015-04-10 WO PCT/AU2015/000216 patent/WO2015154137A1/en not_active Ceased
Non-Patent Citations (2)
| Title |
|---|
| CHENOUARD, N. ET AL.: "Objective Comparison of Particle Tracking Methods", NATURE METHODS, vol. 11, no. 3, March 2014 (2014-03-01), pages 281 - 290, XP055230226, [retrieved on 20140119] * |
| SBALZARINI, I. F. ET AL.: "Feature Point Tracking and Trajectory Analysis for Video Imaging in Cell Biology", JOURNAL OF STRUCTURAL BIOLOGY, vol. 151, 2005, pages 182 - 195, XP027216379 * |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113223632A (en) * | 2021-05-12 | 2021-08-06 | 北京望石智慧科技有限公司 | Molecular fragment library determination method, molecular segmentation method and device |
| CN113223632B (en) * | 2021-05-12 | 2024-02-13 | 北京望石智慧科技有限公司 | A method for determining a molecular fragment library, a molecular segmentation method and a device |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Béthermin et al. | The impact of clustering and angular resolution on far-infrared and millimeter continuum observations | |
| Krone-Martins et al. | UPMASK: unsupervised photometric membership assignment in stellar clusters | |
| JP6525864B2 (en) | Classification method of samples based on spectral data, method of creating database and method of using the database, and corresponding computer program, data storage medium and system | |
| Karpievitch et al. | Normalization and missing value imputation for label-free LC-MS analysis | |
| Bodensteiner et al. | The young massive SMC cluster NGC 330 seen by MUSE-II. Multiplicity properties of the massive-star population | |
| Kenar et al. | Automated label-free quantification of metabolites from liquid chromatography–mass spectrometry data | |
| US20180300609A1 (en) | Facilitating machine-learning and data analysis by computing user-session representation vectors | |
| Ivanov et al. | Empirical multidimensional space for scoring peptide spectrum matches in shotgun proteomics | |
| EP3611695A1 (en) | Generating annotation data of tissue images | |
| JP2016200435A (en) | Mass spectrum analysis system, method, and program | |
| Feher et al. | Cell population identification using fluorescence-minus-one controls with a one-class classifying algorithm | |
| Mantini et al. | Independent component analysis for the extraction of reliable protein signal profiles from MALDI-TOF mass spectra | |
| Jaffeux et al. | Ice crystals images from Optical Array Probes: classification with Convolutional Neural Networks | |
| Li et al. | The effect of ground truth on performance evaluation of hyperspectral image classification | |
| Harada et al. | Enhanced conformational sampling method for proteins based on the T a B oo S e A rch algorithm: Application to the folding of a mini‐protein, chignolin | |
| Ahmed et al. | Investigation on the use of ensemble learning and big data in crop identification | |
| Gupta et al. | Deep learning for morphological identification of extended radio galaxies using weak labels | |
| Chen et al. | Towards biologically plausible and private gene expression data generation | |
| Pino et al. | Semantic segmentation of radio-astronomical images | |
| Murata et al. | Three-dimensional leaf edge reconstruction combining two-and three-dimensional approaches | |
| Penzel et al. | Analyzing the behavior of cauliflower harvest-readiness models by investigating feature relevances | |
| WO2015154137A1 (en) | A method and a system for identifying molecules | |
| Tang et al. | Cluster population demographics in NGC 628 derived from stochastic population synthesis models | |
| Tempel et al. | TOPz: Photometric redshifts using template fitting applied to the GAMA survey | |
| Martins et al. | Evaluating unified training optimisations for MobileNetV2: Efficiency-accuracy trade-offs in fine-grained dog breed classification |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 15776139 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 15776139 Country of ref document: EP Kind code of ref document: A1 |


