WO2013144651A2 - A method for the investigation of differences in analytical data and an apparatus adapted to perform such a method - Google Patents

A method for the investigation of differences in analytical data and an apparatus adapted to perform such a method Download PDF

Info

Publication number
WO2013144651A2
WO2013144651A2 PCT/GB2013/050840 GB2013050840W WO2013144651A2 WO 2013144651 A2 WO2013144651 A2 WO 2013144651A2 GB 2013050840 W GB2013050840 W GB 2013050840W WO 2013144651 A2 WO2013144651 A2 WO 2013144651A2
Authority
WO
WIPO (PCT)
Prior art keywords
data
bins
probability distribution
posterior probability
count rate
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/GB2013/050840
Other languages
French (fr)
Other versions
WO2013144651A3 (en
Inventor
Richard Denny
Kieran NEESON
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Micromass UK Ltd
Original Assignee
Micromass UK Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Micromass UK Ltd filed Critical Micromass UK Ltd
Priority to US14/388,911 priority Critical patent/US20150120212A1/en
Priority to GB1414865.4A priority patent/GB2514942B/en
Publication of WO2013144651A2 publication Critical patent/WO2013144651A2/en
Publication of WO2013144651A3 publication Critical patent/WO2013144651A3/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/0004Gaseous mixtures, e.g. polluted air
    • HELECTRICITY
    • H01ELECTRIC ELEMENTS
    • H01JELECTRIC DISCHARGE TUBES OR DISCHARGE LAMPS
    • H01J49/00Particle spectrometers or separator tubes
    • H01J49/0027Methods for using particle spectrometers
    • H01J49/0036Step by step routines describing the handling of the data generated during a measurement
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N30/00Investigating or analysing materials by separation into components using adsorption, absorption or similar phenomena or using ion-exchange, e.g. chromatography or field flow fractionation

Definitions

  • the present invention relates to method of investigating differences in data produced by at least one analytical instrument and apparatus adapted to perform the method. More particularly, but not exclusively, the present invention relates to the identification of differences in data sets using a data model for the data sets wherein the model ties the data in one bin in one data set to the data in the corresponding bin of the other data set.
  • the user of analytical instruments may wish to identify changes within two or more samples in various different applications. These applications may include, for example, reaction monitoring over time, quality control monitoring, patient diagnosis, and petrochemical investigative analysis.
  • concentration or expression level of one or more components, molecules or analytes in a first sample is quantitated relative to the intensity, concentration or expression level of one or more components, molecules or analytes in a second sample.
  • the method and apparatus according to the present invention seeks to overcome the problems of the prior art.
  • the present invention provides a method of investigating differences in data produced by at least one analytical instrument comprising providing a first data set from a first sample in a plurality of data bins; providing a second data set from a second sample in a plurality of data bins; providing a data model of said data sets in which the data in a plurality of data bins from the first data set is linked to the data in the corresponding bins of the second data set, each linked pair having an associated switch parameter linking the two together; and, exploring the posterior probability distribution for the data model as a function of the switch parameters to produce a posterior probability distribution map.
  • the method according to the invention has the advantage of a considerable increase in speed and accuracy of processing the data.
  • a further advantage of the method according to the invention is that not only are differences in the quantity of components within a sample identified, but components only present in one sample can also be identified.
  • At least one of the first and second data sets is a raw data set.
  • each data bin represents a region of mass to charge ratio against mobility cell drift time.
  • the data held by each bin preferably relates to ion arrival count rate.
  • the switch parameter for each pair of linked data bins relates to the difference in the ion arrival count rate between the linked bins.
  • the data model models the ion arrival count rate for each bin as a product of a normalised count rate for that bin multiplied by a count rate scale factor for the data set to which the bin belongs, the switch parameter for each pair of linked bins relating to the difference in normalised ion count rate between the two bins.
  • the step of exploring the posterior probability distribution further comprises exploring the posterior probability distribution as a function of the count rate scale factors of the data sets.
  • the switch parameter for each pair of linked data bins is a boolean parameter with one value corresponding to the same ion arrival count rate between the two linked bins and the other value corresponding to a different ion arrival count rate between the two bins.
  • the data model further includes a parameter relating to the gain factor of the analytical instrument and the step of exploring the posterior probability distribution further comprises exploring the posterior probability distribution as a function of gain factor.
  • the data model associates at least one shift correction with each bin, the step of exploring the posterior probability distribution further comprising exploring the posterior probability distribution as a function of the at least one shift correction.
  • the at least one shift correction is at least one of drift time, retention time, mass, precursor ion mass and product ion mass.
  • multiple data sets are provided from at least one of the first and second data samples.
  • the method further comprises the step of identifying differences in data between the multiple data sets from the same sample to produce an estimate of variation in data produced by the at least one analytical instrument.
  • the posterior probability distribution is explored by a Monte Carlo algorithm.
  • the Monte Carlo algorithm may be a Markov Chain algorithm.
  • the Monte Carlo algorithm further comprises at least one sampling technique from the list comprising Gibbs sampling and Slice sampling.
  • the Monte Carlo algorithm further comprises at least one sampling technique from the list comprising Metropolis Hastings sampling and Nested sampling.
  • the method further comprises the step of analysing at least a portion of the explored region of the posterior probability distribution map to produce a result for at least one of the parameters of the data model.
  • the method further comprises the step of analysing at least a portion of the posterior probability distribution map to produce a map indicative of the differences between the data produced by the at least one analytical instrument from the first sample and the second sample.
  • the method further comprises the step of further investigating the map indicative of the differences between the data produced by the at least one analytical instrument from the first sample and the second sample to determine differences in composition between the first and second samples.
  • the first data set and the second data set includes data produced by hydrogen deuterium exchange.
  • a computer program element comprising computer readable code means for causing a processor to implement the method of any of claims 1-21.
  • the computer program element is embodied on a computer readable medium.
  • a computer readable medium having a program stored thereon, wherein the program is adapted to make a computer execute a procedure to implement the method of any of claims 1-21.
  • Figure 1 shows raw data produced by an analytical instrument.
  • Shown in figure 1 is raw data produced by an analytical instrument, in this case a mass spectrometer.
  • the data may be viewed as a data set comprising a plurality of bins which in this embodiment are arranged in a rectangular array.
  • the x axis of the array is mass to charge ratio (m/z).
  • the y axis is mobility cell drift time.
  • Each bin therefore represents a small area of mass to charge ratio and mobility cell drift time.
  • Contained within each bin is a count of the number of ion arrivals within a predetermined time (the ion arrival count rate).
  • a detector response is recorded for each bin which may be taken to be proportional to the count.
  • the production of such data sets from analytical instruments is known and will not be discussed in further detail.
  • the purpose of the method according to the invention is to is to compare two similar data sets to obtain a bin by bin probability map that gives the probability that the counts ascribed to corresponding bins arise from the same Poisson source.
  • a first data set from a first sample is provided.
  • the first data set (image) is arranged in a plurality of data bins as described above.
  • a second data set from a second sample is provided.
  • the second data set (image) is also arranged in a plurality of bins as described above.
  • the data model comprises two data arrays of bins corresponding to the bins of the two data sets.
  • ⁇ - ⁇ is exponentially distributed with unit mean, with an identically distributed rate ⁇ 2 , in the second image.
  • Scaling by the count rate scale factor ⁇ 1 (or ⁇ 2 in the second image) allows patterns of rates ⁇ 1 and p 2 to be established independent of the scale factors required to achieve agreement with the data.
  • the basic scheme is to explore the parameters of the data model by constructing an ergodic Markov chain whose stationary distribution is the joint probability distribution of data and model parameters.
  • the chain is constructed using transitions for each parameter which leaves this desired distribution invariant.
  • the chain will be ergodic if, for each transition, the probability distribution is greater than zero. This ensures that there is a non-zero probability of accessing any state, after iterating over each parameter starting from any initial state.
  • p. be a switch state linking a bin in one data set to a corresponding bin in the other data set.
  • Pr(/3 ⁇ 4 false
  • ⁇ , c3 ⁇ 4) 1 - r and r is a random number drawn from (0, 1).
  • a posterior probability map Once such a posterior probability map has been determined it can be analysed to produce a variety of results. By sampling appropriate states in the map one can derive an average value of the switch state for one or more bins. For bins where the average value is close to false this indicates likelihood of a significant difference in data from the two samples in that bin, so indicating a difference in composition of the two samples. For bins where the average value of the switch state is close to false this suggests no difference in composition.
  • the posterior probability distribution is explored as a function of the switch state ⁇ .
  • the posterior probability distribution can be explored as a function of further variables.
  • the posterior probability distribution map as a function of ⁇ one can analyse it by sampling appropriate states to determine the likely values of ⁇ for the two data sets.
  • a further suitable variable is the gain factor.
  • the response of the detector x of the analytical instrument may be proportional to the ion arrival count n rather than identical to it and one may be uncertain about the constant of proportionality or gain factor Y.
  • the factor can be included in the above analysis by scaling down ⁇ and S by Y so that the prior on ⁇ becomes
  • each image may be shifted and re-sampled onto a common drift time axis, from where the likelihood can be re-computed and explored with slice sampling.
  • the first and last points on the drift time axis are not shifted so that the total number of counts is conserved when data are resampled. For safety the extremities are placed away from the interior values by a large margin.
  • the common axis may be shifted to relax the image shifts against their combined prior probability.
  • a Gaussian prior with a standard deviation of one or two bins typically reflects the drift time variability adequately.
  • Slice sampling may again be employed to generate transitions.
  • subscripts c,a and r indicate members of the control group, analyte group and entire group respectively.
  • the probability ratios for the switch states ⁇ and scale factors ⁇ are easily modified as in the above to accommodate the control and analyte groupings.
  • the first data set and second data set can include data produced by hydrogen deuterium exchange.
  • the method and apparatus of the invention may be used to monitor samples in a batch control process, wherein samples may be compared to a predetermined standard, to ascertain whether, and by how much, and in what components, the sample deviates from the standard. This is of use in, for example, the assessment of petroleum and biofuel samples, which must adhere to strict standards.
  • the method and apparatus of the invention may further be used in a sequential process. Rather than comparing a sample against a fixed standard, the samples may be compared against an earlier sample. This is of use in, for example, monitoring drug metabolism over time, or the degradation of petroleum over time. Also for the relative comparison of materials in the chemical industry such as polymers and formulated blends such as paints, coatings, sealants, cosmetics and agrochemicals.
  • the method and apparatus of invention are used to detect and characterise composition changes resulting from varying reaction conditions, errors in formulation make-up, degradation and ageing of materials as a result of environmental conditions and/or mechanical use.

Landscapes

  • Chemical & Material Sciences (AREA)
  • Analytical Chemistry (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • General Physics & Mathematics (AREA)
  • Immunology (AREA)
  • Pathology (AREA)
  • Combustion & Propulsion (AREA)
  • Food Science & Technology (AREA)
  • Medicinal Chemistry (AREA)
  • Analysing Materials By The Use Of Radiation (AREA)
  • Other Investigation Or Analysis Of Materials By Electrical Means (AREA)

Description

A method for the investigation of differences in analytical data and an apparatus adapted to perform such a method
The present invention relates to method of investigating differences in data produced by at least one analytical instrument and apparatus adapted to perform the method. More particularly, but not exclusively, the present invention relates to the identification of differences in data sets using a data model for the data sets wherein the model ties the data in one bin in one data set to the data in the corresponding bin of the other data set.
The user of analytical instruments may wish to identify changes within two or more samples in various different applications. These applications may include, for example, reaction monitoring over time, quality control monitoring, patient diagnosis, and petrochemical investigative analysis.
It may be possible, particularly with relatively simple data, to identify changes by visibly looking at the data sets, and looking for differences. However, doing this can lead to incorrect identification of differences, due to various different factors. For example, these factors may include changes in conditions in which the analysis may be performed; and changes in concentration of the samples. These may lead to changes in the data produced by the instrument, which are not reflective of a change in the sample.
US8012764, and Proteinlynx global server (Micromass UK Limited) discloses a method where two separate samples are mass analysed and then the relative intensity,
concentration or expression level of one or more components, molecules or analytes in a first sample is quantitated relative to the intensity, concentration or expression level of one or more components, molecules or analytes in a second sample.
This is performed by the identification of peaks in the data resulting from components within the first sample, and identification of the corresponding peaks from data relating to those same components in the second sample. The intensities of corresponding peaks are then interrogated to provide information relating to differences in the quantity of the components in both samples. However, the quantity of data produced by modern analytical instruments may be extremely large. This may result in the time taken for analysis of the data to be extremely long.
The method and apparatus according to the present invention seeks to overcome the problems of the prior art.
Accordingly in a first aspect the present invention provides a method of investigating differences in data produced by at least one analytical instrument comprising providing a first data set from a first sample in a plurality of data bins; providing a second data set from a second sample in a plurality of data bins; providing a data model of said data sets in which the data in a plurality of data bins from the first data set is linked to the data in the corresponding bins of the second data set, each linked pair having an associated switch parameter linking the two together; and, exploring the posterior probability distribution for the data model as a function of the switch parameters to produce a posterior probability distribution map.
The method according to the invention has the advantage of a considerable increase in speed and accuracy of processing the data. A further advantage of the method according to the invention is that not only are differences in the quantity of components within a sample identified, but components only present in one sample can also be identified.
Preferably at least one of the first and second data sets is a raw data set.
Preferably, each data bin represents a region of mass to charge ratio against mobility cell drift time.
The data held by each bin preferably relates to ion arrival count rate. Preferably the switch parameter for each pair of linked data bins relates to the difference in the ion arrival count rate between the linked bins.
Preferably the data model models the ion arrival count rate for each bin as a product of a normalised count rate for that bin multiplied by a count rate scale factor for the data set to which the bin belongs, the switch parameter for each pair of linked bins relating to the difference in normalised ion count rate between the two bins.
Preferably, the step of exploring the posterior probability distribution further comprises exploring the posterior probability distribution as a function of the count rate scale factors of the data sets.
Preferably the switch parameter for each pair of linked data bins is a boolean parameter with one value corresponding to the same ion arrival count rate between the two linked bins and the other value corresponding to a different ion arrival count rate between the two bins.
Preferably, the data model further includes a parameter relating to the gain factor of the analytical instrument and the step of exploring the posterior probability distribution further comprises exploring the posterior probability distribution as a function of gain factor.
Preferably, the data model associates at least one shift correction with each bin, the step of exploring the posterior probability distribution further comprising exploring the posterior probability distribution as a function of the at least one shift correction.
Preferably the at least one shift correction is at least one of drift time, retention time, mass, precursor ion mass and product ion mass. Optionally multiple data sets are provided from at least one of the first and second data samples.
Preferably, the method further comprises the step of identifying differences in data between the multiple data sets from the same sample to produce an estimate of variation in data produced by the at least one analytical instrument.
Optionally the posterior probability distribution is explored by a Monte Carlo algorithm.
The Monte Carlo algorithm may be a Markov Chain algorithm.
Preferably the Monte Carlo algorithm further comprises at least one sampling technique from the list comprising Gibbs sampling and Slice sampling.
Preferably the Monte Carlo algorithm further comprises at least one sampling technique from the list comprising Metropolis Hastings sampling and Nested sampling.
Optionally the method further comprises the step of analysing at least a portion of the explored region of the posterior probability distribution map to produce a result for at least one of the parameters of the data model.
Optionally the method further comprises the step of analysing at least a portion of the posterior probability distribution map to produce a map indicative of the differences between the data produced by the at least one analytical instrument from the first sample and the second sample.
Preferably, the method further comprises the step of further investigating the map indicative of the differences between the data produced by the at least one analytical instrument from the first sample and the second sample to determine differences in composition between the first and second samples.
Optionally the first data set and the second data set includes data produced by hydrogen deuterium exchange.
In a further aspect of the invention there is provided a computer program element comprising computer readable code means for causing a processor to implement the method of any of claims 1-21.
Preferably the computer program element is embodied on a computer readable medium.
In a further aspect of the invention there is provided a computer readable medium having a program stored thereon, wherein the program is adapted to make a computer execute a procedure to implement the method of any of claims 1-21.
In a further aspect of the invention there is provided an analytical instrument, adapted to perform the method as claimed in any one of claims 1-21.
The present invention will now be described by way of example only, and not in any limitative sense with reference to the accompanying drawings in which
Figure 1 shows raw data produced by an analytical instrument.
Shown in figure 1 is raw data produced by an analytical instrument, in this case a mass spectrometer. The data may be viewed as a data set comprising a plurality of bins which in this embodiment are arranged in a rectangular array. The x axis of the array is mass to charge ratio (m/z). The y axis is mobility cell drift time. Each bin therefore represents a small area of mass to charge ratio and mobility cell drift time. Contained within each bin is a count of the number of ion arrivals within a predetermined time (the ion arrival count rate). In practice a detector response is recorded for each bin which may be taken to be proportional to the count. The production of such data sets from analytical instruments is known and will not be discussed in further detail.
In non-limiting outline, the purpose of the method according to the invention is to is to compare two similar data sets to obtain a bin by bin probability map that gives the probability that the counts ascribed to corresponding bins arise from the same Poisson source.
In a first step of an embodiment of a method according to the invention a first data set from a first sample is provided. The first data set (image) is arranged in a plurality of data bins as described above.
In a second step a second data set from a second sample is provided. The second data set (image) is also arranged in a plurality of bins as described above.
In order to compare the two data sets and identify differences between them a data model is provided. The data model comprises two data arrays of bins corresponding to the bins of the two data sets.
The i'th bin (where i indexes the corresponding bins in the two images) is assigned a Poisson ion arrival rate hri in the first image, which is split into two factors, λ·π1 ίσ1. Here, μ-ι, is exponentially distributed with unit mean,
Figure imgf000007_0001
with an identically distributed rate μ2, in the second image. Scaling by the count rate scale factor σ1 (or σ2 in the second image) allows patterns of rates μ1 and p2to be established independent of the scale factors required to achieve agreement with the data. The ion count Hii in the bin image 1 has probability Pr(nH I μα, σι) = —
and the joint probability is
Figure imgf000008_0001
can marginalise over (integrate out) to give
Figure imgf000008_0002
If the corresponding bins in the two images are assumed to have the same Poisson rate so that μΐ ί then μ, may again be marginalised away to give
Pr{nlh n2i σ1, σ2 μα = μ·&} - ι ; , Ί \Βι .+¾¾ .+ι i Γ—
Conversely, if they are assumed to have different rates μ^ and μ2, are marginalised separately to give
Figure imgf000008_0003
From these equations the posterior probability distribution for the data model can be derived.
The basic scheme is to explore the parameters of the data model by constructing an ergodic Markov chain whose stationary distribution is the joint probability distribution of data and model parameters. The chain is constructed using transitions for each parameter which leaves this desired distribution invariant. The chain will be ergodic if, for each transition, the probability distribution is greater than zero. This ensures that there is a non-zero probability of accessing any state, after iterating over each parameter starting from any initial state.
One wishes to obtain a posterior probability distribution map of Pr^p^paJ!image 1 = {r\Y , image 2 ={η}) and so must explore the space of same/different assignment of rates for each bin along with the two scale factors σ-ι and σ2 for each image.
Let p. be a switch state linking a bin in one data set to a corresponding bin in the other data set. Let βί≡(μ·ϋ = μ2,) be a Boolean variable. A transition from β, = false to β, = true has a probability ratio
Pr(A = true | σ1, σ2) = Pr(/¾ = true) {σ1 + l)""+1 (g2 + (nu + nu) \
Pr(/¾ = false j σι, σ2) Pr(/¾ = false) (σχ + σ2 + l) »"+*2i+i '
These transitions can be explored by a Gibbs sampling scheme whereby a transition to ,=b is made where h = Pr(A = true | a a2)
Pr(/¾ = false | σι , c¾) 1 - r and r is a random number drawn from (0, 1).
Once such a posterior probability map has been determined it can be analysed to produce a variety of results. By sampling appropriate states in the map one can derive an average value of the switch state for one or more bins. For bins where the average value is close to false this indicates likelihood of a significant difference in data from the two samples in that bin, so indicating a difference in composition of the two samples. For bins where the average value of the switch state is close to false this suggests no difference in composition. One can perform this analysis for each bin to produce a map which is indicative of the differences between the data produced by the at least one analytical instrument from the first sample and the second sample. This in turn can be used to determine differences in composition between the first sample and the second sample. Typically when sampling states in the posterior probability map one tends not to take into consideration the first few determined states in the chain. This is because the initial starting state of the model may be a highly improbable state and it takes a few iterations to approach the region of probable states.
In the above embodiment the posterior probability distribution is explored as a function of the switch state β. In more complex models the posterior probability distribution can be explored as a function of further variables.
One particular further variable is the count rate scale factor. The likelihood for scale factors σ1 or o2 is
Pr({nii}, {n2i} | σ1 , σ2, {βί}) oc ·
Figure imgf000010_0001
(σι + σ2 + l) =trae (σι + (σ2 + ij^-w-
Probable values for σ1 and σ2 may easily be explored by slice sampling, once a suitable prior has been established. It is straightforward to decouple slice sampling from the details of a prior probability distribution for & for instance by driving it through a controlling variable u uniformly distributed on (0,1) such that
Θ
Figure imgf000010_0002
A convenient and unrestrictive prior for σ is
Figure imgf000010_0003
so that
Figure imgf000011_0001
and σ(ν Su
1-u where S is the median value.
Once one has determined the posterior probability distribution map as a function of σ one can analyse it by sampling appropriate states to determine the likely values of σ for the two data sets.
A further suitable variable is the gain factor. The response of the detector x of the analytical instrument may be proportional to the ion arrival count n rather than identical to it and one may be uncertain about the constant of proportionality or gain factor Y. The factor can be included in the above analysis by scaling down σ and S by Y so that the prior on σ becomes
Figure imgf000011_0002
and by modifying the Poisson likelihood to the form
Figure imgf000011_0003
where [x Y] = n. The complete likelihood over all pixels (bins) is P n2i
Figure imgf000012_0001
£fc=ialse where I is the total number of pixels (bins) in an image. A similar prior for Y may be employed to that used for the scale factors σ but offset by one as the lowest attainable gain corresponds to the counting of single ions. Slice sampling may again be employed to generate transitions.
Often there is some shift in the drift time calibration between acquisitions, the shift being of the order of one bin (of typically 200). Each image may be shifted and re-sampled onto a common drift time axis, from where the likelihood can be re-computed and explored with slice sampling. There are a couple of technicalities involved in this procedure. Firstly, the first and last points on the drift time axis are not shifted so that the total number of counts is conserved when data are resampled. For safety the extremities are placed away from the interior values by a large margin. Secondly, once all images have been shifted the common axis may be shifted to relax the image shifts against their combined prior probability. A Gaussian prior with a standard deviation of one or two bins typically reflects the drift time variability adequately. Slice sampling may again be employed to generate transitions.
There is some variation between acquisitions of nominally equivalent samples beyond strict application of Poisson statistics. These replicate acquisitions may be used to accommodate extra variation between nominally non-equivalent samples that may not be significant. Consider C replicate injections of sample 1 (the control) and A replicate injections of sample 2 (the analyte) giving a total of C+A = R images. The complete likelihood becomes P
Figure imgf000013_0001
Where subscripts c,a and r indicate members of the control group, analyte group and entire group respectively.
The probability ratios for the switch states β and scale factors σ are easily modified as in the above to accommodate the control and analyte groupings.
The above analysis depends on the image pixel (bin) size. At both extremes (single pixel per image and single count or zero counts per pixel) the data becomes uninformative with regard to the model being applied. Choice of pixel size is therefore an important consideration.
The method according to the invention has been described in various embodiments above with reference to Slice sampling and Gibbs sampling. Other sampling methods may be employed more particularly but not limited to Metropolis Hastings sampling and Nested sampling.
The first data set and second data set can include data produced by hydrogen deuterium exchange.
The method and apparatus of the invention may be used to monitor samples in a batch control process, wherein samples may be compared to a predetermined standard, to ascertain whether, and by how much, and in what components, the sample deviates from the standard. This is of use in, for example, the assessment of petroleum and biofuel samples, which must adhere to strict standards.
The method and apparatus of the invention may further be used in a sequential process. Rather than comparing a sample against a fixed standard, the samples may be compared against an earlier sample. This is of use in, for example, monitoring drug metabolism over time, or the degradation of petroleum over time. Also for the relative comparison of materials in the chemical industry such as polymers and formulated blends such as paints, coatings, sealants, cosmetics and agrochemicals. Here the method and apparatus of invention are used to detect and characterise composition changes resulting from varying reaction conditions, errors in formulation make-up, degradation and ageing of materials as a result of environmental conditions and/or mechanical use.
When used in this specification and claims, the terms "comprises" and "comprising" and variations thereof mean that the specified features, steps or integers are included. The terms are not to be interpreted to exclude the presence of other features, steps or components.
The features disclosed in the foregoing description, or the following claims, or the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for attaining the disclosed result, as appropriate, may, separately, or in any combination of such features, be utilised for realising the invention in diverse forms thereof.

Claims

A method of investigating differences in data produced by at least one analytical instrument comprising providing a first data set from a first sample in a plurality of data bins; providing a second data set from a second sample in a plurality of data bins; providing a data model of said data sets in which the data in a plurality of data bins from the first data set is linked to the data in the corresponding bins of the second data set, each linked pair having an associated switch parameter linking the two together; and, exploring the posterior probability distribution for the data model as a function of the switch parameters to produce a posterior probability distribution map.
A method as claimed in claim 1 , wherein at least one of the first and second data sets is a raw data set.
A method as claimed in either of claims 1 or 2, wherein each bin represents a region of mass to charge ratio against mobility cell drift time.
A method as clamed in any one of claims 1 to 3, wherein the data held by each bin relates to ion arrival count rate.
A method as claimed in claim 4, wherein the switch parameter for each pair of linked data bins relates to the difference in the ion arrival count rate between the linked bins.
6. A method as claimed in claim 5, wherein the data model models the ion arrival count rate for each bin as a product of normalised count rate for that bin multiplied by a count rate scale factor for the data set to which the bin belongs, the switch parameter for each pair of linked bins relating to the difference in normalised ion count rate count rate between the linked bins.
7. A method as claimed in claim 6, wherein the step of exploring the posterior probability distribution further comprises exploring the posterior probability distribution as a function of the count rate scale factors of the data sets.
8. A method as claimed in any one of claims 5 to 7, wherein the switch parameter for each pair of linked data bins is a boolean parameter with one value corresponding to the same ion arrival count rate between the two linked bins and the other value corresponding to a different ion arrival count rate between the two bins.
9. A method as claimed in any one of claims 1 to 8, wherein the data model further includes a parameter relating to the gain factor of the analytical instrument and the step of exploring the posterior probability distribution further comprises exploring the posterior probability distribution as a function of gain factor.
10. A method as claimed in any one of claims 1 to 9, wherein the data model associates at least one shift correction with each bin, the step of exploring the posterior probability distribution further comprising exploring the posterior probability distribution as a function of the at least one shift correction.
11. A method as claimed in claim 10, where in the at least one shift correction is at least one of drift time, retention time, mass, precursor ion mass and product ion mass.
A method as claimed in any one of claims 1 to 1 1 , wherein multiple data sets are provided from at least one of the first and second data samples.
13. A method as claimed in claim 12 further comprising the step of identifying differences in data between the multiple data sets from the same sample to produce an estimate of the variation in data produced by the at least one analytical instrument.
14. A method as claimed in any one of claims 1 to 13, wherein the posterior probability distribution is explored by a Monte Carlo algorithm.
15. A method as clamed in claim 14, wherein the Monte Carlo algorithm is a Markov Chain algorithm.
16. A method as claimed in either of claims 14 or 15, wherein the Monte Carlo algorithm further comprises at least one sampling technique from the list comprising Gibbs sampling and Slice sampling.
17. A method as claimed in any one of claims 14 to 16, wherein the Monte Carlo algorithm further comprises at least one sampling technique from the list comprising Metropolis Hastings sampling and Nested sampling.
18. A method as claimed in any one of claims 1 to 17, further comprising the step of analysing at least a portion of the posterior probability distribution map to produce a result for at least one of the parameters of the data model.
19. A method as claimed in any one of claims 1 to 18, further comprising the step of analysing at least a portion of the posterior probability distribution map to produce a map indicative of the differences between the data produced by the at least one analytical instrument from the first sample and the second sample.
20. A method as claimed in claim 19, further comprising further investigating the map indicative of differences between the data produced by the at least one analytical instrument from the first sample and the second sample to determine differences in composition between the first and the second samples.
21. A method as claimed in any one of claims 1 to 20, wherein the first data set and the second data set includes data produced by hydrogen deuterium exchange.
22. A computer program element comprising computer readable code means for causing a processor to implement the method of any of claims 1-21.
23. A computer program element according to claim 22 embodied on a computer readable medium.
24. A computer readable medium having a program stored thereon, where the program is adapted to make a computer execute a procedure to implement the method of any of claims 1-21.
25. An analytical instrument, adapted to perform the method as claimed in any one of claims 1-21
PCT/GB2013/050840 2012-03-30 2013-03-28 A method for the investigation of differences in analytical data and an apparatus adapted to perform such a method Ceased WO2013144651A2 (en)

Priority Applications (2)

Application Number Priority Date Filing Date Title
US14/388,911 US20150120212A1 (en) 2012-03-30 2013-03-28 Method for the investigation of differences in analytical data and an apparatus adapted to perform such a method
GB1414865.4A GB2514942B (en) 2012-03-30 2013-03-28 A method for the investigation of differences in analytical data and an apparatus adapted to perform such a method

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
GB201205720A GB201205720D0 (en) 2012-03-30 2012-03-30 A method for the investigation of differences in analytical data and an apparatus adapted to perform such a method
GB1205720.4 2012-03-30

Publications (2)

Publication Number Publication Date
WO2013144651A2 true WO2013144651A2 (en) 2013-10-03
WO2013144651A3 WO2013144651A3 (en) 2014-02-27

Family

ID=46160059

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/GB2013/050840 Ceased WO2013144651A2 (en) 2012-03-30 2013-03-28 A method for the investigation of differences in analytical data and an apparatus adapted to perform such a method

Country Status (3)

Country Link
US (1) US20150120212A1 (en)
GB (2) GB201205720D0 (en)
WO (1) WO2013144651A2 (en)

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8012764B2 (en) 2004-04-30 2011-09-06 Micromass Uk Limited Mass spectrometer

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
GB2308917B (en) * 1996-01-05 2000-04-12 Maxent Solutions Ltd Reducing interferences in elemental mass spectrometers
CA2501003C (en) * 2004-04-23 2009-05-19 F. Hoffmann-La Roche Ag Sample analysis to provide characterization data
FR2920235B1 (en) * 2007-08-22 2009-12-25 Commissariat Energie Atomique METHOD FOR ESTIMATING MOLECULE CONCENTRATIONS IN A SAMPLE STATE AND APPARATUS
GB201019337D0 (en) * 2010-11-16 2010-12-29 Micromass Ltd Controlling hydrogen-deuterium exchange on a spectrum by spectrum basis
FR2979705B1 (en) * 2011-09-05 2014-05-09 Commissariat Energie Atomique METHOD AND DEVICE FOR ESTIMATING A MOLECULAR MASS PARAMETER IN A SAMPLE

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8012764B2 (en) 2004-04-30 2011-09-06 Micromass Uk Limited Mass spectrometer

Also Published As

Publication number Publication date
GB201414865D0 (en) 2014-10-08
US20150120212A1 (en) 2015-04-30
GB2514942B (en) 2018-07-18
WO2013144651A3 (en) 2014-02-27
GB201205720D0 (en) 2012-05-16
GB2514942A (en) 2014-12-10

Similar Documents

Publication Publication Date Title
Aguilan et al. Guide for protein fold change and p-value calculation for non-experts in proteomics
CA2641025C (en) Overlap density (od) heatmaps and consensus data displays
US20250182855A1 (en) Methods and systems for visualizing and evaluating data
US10957523B2 (en) 3D mass spectrometry predictive classification
WO2011127544A1 (en) Intensity normalization in imaging mass spectrometry
La Ferlita et al. RNAdetector: a free user-friendly stand-alone and cloud-based system for RNA-Seq data analysis
WO2015097217A1 (en) Method and system for preparing synthetic multicomponent biotechnological and chemical process samples
JP5945365B2 (en) Method for identifying substances from NMR spectra
EP1238359A2 (en) Methods for normalization of experimental data
Dutta et al. Data-driven equation for drug–membrane permeability across drugs and membranes
Szymańska et al. Increasing conclusiveness of clinical breath analysis by improved baseline correction of multi capillary column–ion mobility spectrometry (MCC-IMS) data
CN107664655B (en) Method and apparatus for characterizing analytes
EP3724654B1 (en) Method for analyzing small molecule components of a complex mixture, and associated apparatus and computer program product
CN102906851A (en) Method, computer program, and system to analyze mass spectra
WO2013144651A2 (en) A method for the investigation of differences in analytical data and an apparatus adapted to perform such a method
EP2646811B1 (en) Method for automatic peak finding in calorimetric data
Kumar Application of Akaike information criterion assisted probabilistic latent semantic analysis on non-trilinear total synchronous fluorescence spectroscopic data sets: Automatizing fluorescence based multicomponent mixture analysis
Jin et al. Robust discriminant analysis and its application to identify protein coding regions of rice genes
Hu et al. Joint precursor elution profile inference via regression for peptide detection in data-independent acquisition mass spectra
JP5866287B2 (en) Apparatus and related methods for small molecule component analysis in complex mixtures
EP4361624B1 (en) Method for estimating content ratio of components contained in sample, composition estimating device, and program
Kirchner et al. Non-linear classification for on-the-fly fractional mass filtering and targeted precursor fragmentation in mass spectrometry experiments
JP4835695B2 (en) Chromatograph mass spectrometer
Chung et al. Non-parametric Bayesian approach to post-translational modification refinement of predictions from tandem mass spectrometry
Kumar et al. Constraint randomised non-negative factor analysis (CRNNFA): an alternate chemometrics approach for analysing the biochemical data sets

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13716822

Country of ref document: EP

Kind code of ref document: A2

ENP Entry into the national phase

Ref document number: 1414865

Country of ref document: GB

Kind code of ref document: A

Free format text: PCT FILING DATE = 20130328

WWE Wipo information: entry into national phase

Ref document number: 1414865.4

Country of ref document: GB

WWE Wipo information: entry into national phase

Ref document number: 14388911

Country of ref document: US

122 Ep: pct application non-entry in european phase

Ref document number: 13716822

Country of ref document: EP

Kind code of ref document: A2