WO2025035386A1 - 一种高灵敏、高通量的化学物质注释方法与系统 - Google Patents
一种高灵敏、高通量的化学物质注释方法与系统 Download PDFInfo
- Publication number
- WO2025035386A1 WO2025035386A1 PCT/CN2023/113115 CN2023113115W WO2025035386A1 WO 2025035386 A1 WO2025035386 A1 WO 2025035386A1 CN 2023113115 W CN2023113115 W CN 2023113115W WO 2025035386 A1 WO2025035386 A1 WO 2025035386A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- candidate
- peak
- chromatographic
- chemical
- target
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N27/00—Investigating or analysing materials by the use of electric, electrochemical, or magnetic means
- G01N27/62—Investigating or analysing materials by the use of electric, electrochemical, or magnetic means by investigating the ionisation of gases, e.g. aerosols; by investigating electric discharges, e.g. emission of cathode
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N30/00—Investigating or analysing materials by separation into components using adsorption, absorption or similar phenomena or using ion-exchange, e.g. chromatography or field flow fractionation
- G01N30/02—Column chromatography
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16C—COMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
- G16C20/00—Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
- G16C20/20—Identification of molecular entities, parts thereof or of chemical compositions
Definitions
- the present invention relates to the technical field of chemical substance identification, and more specifically, to a highly sensitive and high-throughput chemical substance annotation method and system.
- Chemical annotation is an important part of the exposome research.
- the concept of exposure was originally used to explain the environmental drivers of health and disease. Now, it represents all exogenous and endogenous environmental exposures and related biological effects throughout the life cycle.
- the concentration range of xenobiotic chemicals in biological and environmental samples is very wide, with low concentrations reaching ppb or nanomolar levels. Therefore, there is an urgent need for a highly sensitive chemical exposure localization technology with a wide range of chemical coverage to promote the development of exposure type research.
- Traditional chemical annotation requires the use of a tandem mass spectrometer to obtain parent ion mass spectrometry (MS) data and fragment ion mass spectrometry data respectively.
- Parent ions also known as precursor ions, are usually generated in the ion source.
- fragment ions usually requires the combination of an electrospray ion source (ESI) and a tandem mass spectrometer.
- EESI electrospray ion source
- tandem mass spectrometer When the parent ion enters the collision chamber, it reacts with inert gas molecules under the action of energy to produce fragment ions.
- Targeted detection based on triple quadrupole mass spectrometry is generally considered to be the most sensitive technology for monitoring biomarkers of exposure. However, this method can only monitor a limited number of chemicals, limiting its application in exposure studies.
- Non-targeted analysis methods based on high resolution mass spectrometry (HRMS) can detect a wide range of chemicals.
- Full scan acquisition methods are usually used to obtain the m/z (mass-to-charge ratio) and retention time (RT) information of the parent ion, and then the mass spectrum (MS/MS) is generated using data-dependent acquisition (DDA) mode or data-independent acquisition (DIA) mode.
- DDA data-dependent acquisition
- DIA data-independent acquisition
- DDA mode the mass spectrometer automatically performs MS/MS analysis on the parent ion list selected from the full scan spectrum when performing an MS full scan.
- DDA mode is often used to collect fragment spectra of exposure features of interest. The selection of parent ions depends on the ion intensity, so mass spectrometric features of low-abundance exposures may never be selected to produce fragments in the collision cell.
- the advantage of DDA mode is high selectivity, but the throughput is low and it is difficult to achieve batch acquisition of chemical fragment ions;
- Data Independent Acquisition (DIA) mode is a mode for batch acquisition of chemical fragment ions in mass spectrometry. In this mode, all ions in the selected m/z range will be fragmented and analyzed in the second stage of the tandem mass spectrometer.
- tandem mass spectrometry data can be obtained by sequentially separating and fragmenting m/z ranges.
- DIA mode can achieve batch acquisition of chemical fragment ions, it has low selectivity.
- Informatics algorithms have certain limitations in the annotation of chemicals at lower levels.
- data processing informatics algorithms used for autonomous chemical identification and annotation in exposure studies mainly draw on algorithms from the field of metabolomics.
- concentrations of endogenous metabolites in biological samples are usually several orders of magnitude higher than those of exogenous chemicals in the same sample.
- concentrations of chemical pollutants are low, usually around ppb.
- Early studies have shown that traditional metabolomics data preprocessing software such as XCMS will miss a large number of metabolic features, i.e., features with low abundance and/or poor shape. Therefore, traditional peak extraction algorithms will lead to feature loss and affect downstream substance identification.
- EISA highly sensitive chemical analysis technology
- ESI electrospray ionization
- ESI electrospray ionization
- the source fragment generation uses the EISA technology full scan mode HRMS to simulate the endogenous metabolites and peptides of the medium and high energy MS/MS fragment spectrum through the collision induced dissociation (CID) technology in the collision cell without affecting the intensity of the precursor ion, thereby effectively realizing the identification of low-concentration chemical substances.
- CID collision induced dissociation
- traditional mass spectrometry data analysis software such as XCMS cannot analyze EISA data. Therefore, it is urgent to propose a highly sensitive and high-throughput chemical substance annotation method.
- the prior art provides a compound fragment ion prediction method and application, including: 1) determining the compound used for the experiment, the parent ion molecular formula and the fragment ion molecular formula of the compound; 2) selecting stable isotopes according to the elemental composition of the parent ion and the fragment ion; 3) calculating the stable isotope labeling of the parent ion and the fragment ion according to the determined compound, the parent ion molecular formula and the fragment ion molecular formula, and the number of stable isotope elements in the parent ion and the fragment ion; 4) calculating the number of stable isotope labeling of all parent ions and all fragment ions; 5) calculating the accurate mass number of the parent ion and the fragment ion in all labeling situations according to the number of labeling situations obtained in step 4), and calculating the mass-to-charge ratio; 6) establishing a mass-to-charge ratio database of the fragment ions of the compound.
- the prior art needs to clarify the relationship between the parent ion and the fragment ion in the compound used for the experiment, while the EISA technology collects all parent ions and fragment ions in a full scan, and the relationship between the parent ion and the fragment ion is unclear, which is not suitable for EISA data.
- the present invention provides a highly sensitive and high-throughput chemical substance annotation method to overcome the defects of the above-mentioned prior art that the EISA data cannot be accurately analyzed, and thus the high-sensitivity and high-throughput chemical substance detection cannot be achieved.
- the method and system significantly increase the number of low-concentration chemicals detected with high sensitivity and high throughput.
- the present invention provides a highly sensitive and high-throughput chemical substance annotation method, comprising:
- S8 Output the first several candidate chromatographic peaks in the sorting results as the final screening results of the chemical substances in the sample to be detected.
- a chemical exposure group database is constructed based on a local database of chemical standards, an open source mass spectrometry database, and mass spectrometry data from public literature;
- the chemical exposure group database includes several chemical substances and their mass spectrometry data, and the mass spectrometry data of each chemical substance includes the name of the chemical substance, the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of the fragment ion, and the intensity.
- step S3 a mass accuracy standard is set, and the chemical substances in the chemical exposure group database are sequentially used as target chemical substances, and the specific method for extracting the parent ion chromatogram of the sample to be detected is:
- the chemical substances in the chemical exposure group database are taken as target chemical substances in turn, and the mass-to-charge ratio range of the parent ion that needs to be extracted is obtained according to the mass-to-charge ratio of the parent ion of the target chemical substance and the set mass accuracy standard; in the sample to be tested, for the ions whose parent ion mass-to-charge ratio is within the range of the parent ion mass-to-charge ratio, the operation of extracting and constructing the chromatogram is performed.
- step S4 the parent ion chromatogram of the sample to be detected is corrected, and the specific method for obtaining the target chromatogram is:
- the chromatogram For the chromatogram of the extracted sample to be tested, the chromatogram is taken before the start and the end of the chromatographic elution. The corresponding part of the chromatogram is deleted, and the remaining part of the chromatogram is used as the target chromatogram for subsequent analysis.
- the chromatographic elution is set according to actual conditions, and the corresponding chromatogram parts before and after the elution are deleted, and only the chromatogram part during the elution process is retained, thereby reducing the occurrence of false positive features and making the detection result more accurate.
- step S5 is:
- the peak type of the corresponding candidate chromatographic peak is classified as the first type; for candidate chromatographic peaks that cannot be detected by the existing mass spectrometry data analysis algorithm, if their sawtooth index is less than the preset sawtooth index threshold, the peak type of the corresponding candidate chromatographic peak is classified as the second type; otherwise, the peak type of the corresponding candidate chromatographic peak is classified as the third type.
- Filtering chromatographic peaks with peak heights lower than the preset peak height threshold can reduce the number of candidate chromatographic peaks without affecting the accuracy of the test results, improve the detection speed and reduce the occurrence of false positive features.
- classifying the peak types of candidate chromatographic peaks not only whether they can be detected by the existing mass spectrometry data analysis algorithm is considered, but also the sawtooth index is considered, retaining features that cannot be detected by the existing mass spectrometry data analysis algorithm, and expanding the scope of chemical substance identification.
- the third type of candidate chromatographic peak is used as a reference type, and whether to filter further is selected according to actual needs; if filtered, there are two types of peaks used for subsequent analysis, namely the first type and the second type; if not filtered, there are three types of peaks used for subsequent analysis.
- step S5.1 if the target chemical substance in the chemical exposure group database has a retention time, only the peak height and sawtooth index of the chromatographic peak appearing in the retention time window are calculated.
- the retention time window By using the retention time window, the range and number of chromatographic peaks that need to be screened are narrowed, which can effectively improve the subsequent matching speed and accuracy.
- step S6 is:
- S6.2 Calculate a first score and a second score based on the matched fragment ions at the top of the candidate chromatographic peak and the fragment ions of the target chemical in the chemical exposure group database;
- step S6.2 the specific method for calculating the first score is:
- MFR i represents the first score of the fragment ion matched at the top of the i-th candidate chromatographic peak of the target chemical
- Ni represents the number of fragment ions matched at the top of the i-th candidate chromatographic peak
- NT represents the number of fragment ions of the target chemical in the chemical exposure group database.
- step S6.2 the specific method for calculating the second score is:
- the second score is calculated based on the intensity of the fragment ion matched at the top of the candidate chromatographic peak and the intensity of the fragment ion of the target chemical in the chemical exposure group database, and the calculation formula is:
- SSM i represents the second score of the fragment ion matched at the top of the i-th candidate chromatographic peak of the target chemical
- W Qi represents the intensity of the fragment ion matched at the top of the i-th candidate chromatographic peak
- W Ri represents the intensity of the i-th fragment ion of the target chemical in the chemical exposure group database.
- step S6.3 based on the first score and the second score, a specific method for calculating the relevant characteristic score of the candidate chromatographic peak corresponding to the target chemical substance is:
- Score i ⁇ MFR i + ⁇ SSM i
- Score i represents the relevant feature score of the candidate feature of the i-th candidate chromatographic peak corresponding to the target chemical substance
- ⁇ and ⁇ represent the first and second weight coefficients, respectively.
- step S7 is:
- S7.1 Set the sorting priority of the peak types of candidate chromatographic peaks, from high to low, first type, second type, and third type;
- S7.3 Concatenate the candidate features of the target chemical substance within each peak type according to the sorting priority of the peak type to obtain a sorting result.
- the present invention also provides a highly sensitive and high-throughput chemical substance annotation system for implementing the above-mentioned annotation method, comprising:
- a data acquisition module used to collect parent ions and fragment ions in the sample to be detected based on EISA technology
- a database construction module is used to construct a chemical exposure group database, including several chemical substances and their mass spectrometry data;
- a chromatogram generation module is used to set a mass accuracy standard, take the chemical substances in the chemical exposure group database as target chemical substances in turn, and extract the parent ion chromatogram of the sample to be tested accordingly;
- a chromatogram correction module is used to correct the parent ion chromatogram of the sample to be detected to obtain a target chromatogram; the target chromatogram includes a plurality of chromatographic peaks;
- a chromatographic peak screening module is used to screen all chromatographic peaks to obtain candidate chromatographic peaks and their peak types
- a data matching module is used to traverse and match target chemical substances for each candidate chromatographic peak, and calculate the relevant characteristic score of each candidate chromatographic peak corresponding to the target chemical substance;
- a sorting module used to sort each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and related characteristic score of the candidate chromatographic peak to obtain a sorting result
- the chemical substance detection module is used to output the first several candidate chromatographic peaks in the sorting results as the final screening results of the chemical substances in the sample to be detected.
- the present invention first uses EISA technology to collect all parent ions and fragment ions of the sample to be detected; then constructs a chemical exposure group database including several chemical substances and their mass spectrometry data, and extracts parent ion chromatograms of the sample to be detected by taking the chemical substances as target chemical substances in turn; corrects the parent ion chromatogram to obtain a target chromatogram containing several chromatographic peaks; then screens all chromatographic peaks to obtain candidate chromatographic peaks and their peak types; for each candidate chromatographic peak, traverses and matches the target chemical substance to obtain the target chemical substance
- the method can detect the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance, and finally sort each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and the relevant characteristic scores of the candidate chromatographic peaks, and output the first several candidate chromatographic peaks in the sorting results as the final screening results of the chemical substances in the sample to be detected.
- the present invention can significantly increase the number of low-concentr
- FIG1 is a flow chart of a highly sensitive, high-throughput chemical substance annotation method described in Example 1.
- FIG. 2 is a schematic diagram of the chemical exposure group database described in Example 2.
- FIG3 is a chromatogram of the sample to be detected extracted based on the preset mass accuracy standard described in Example 2.
- FIG. 4 is a target chromatogram as described in Example 2.
- FIG. 5 is a chromatographic peak as a processing target when the retention time described in Example 2 exists.
- FIG. 6 is a chromatographic peak as a processing target when there is no retention time as described in Example 2.
- FIG. 7 is a schematic diagram comparing the detection results of the method provided in this embodiment described in Example 2 and the traditional peak extraction algorithm at different concentrations.
- FIG8 is a schematic diagram showing a comparison of the detection results of the method provided in this embodiment described in Example 2 and the traditional TMM collection method at different concentrations.
- FIG9 is a schematic diagram of the structure of a highly sensitive, high-throughput chemical substance annotation system described in Example 3.
- This embodiment provides a highly sensitive and high-throughput chemical substance annotation method, as shown in FIG1 , comprising:
- S8 Output the first several candidate chromatographic peaks in the sorting results as the final screening results of the chemical substances in the sample to be detected.
- this embodiment first uses EISA technology to collect all parent ions and fragment ions of the sample to be detected; then constructs a chemical exposure group database including several chemical substances and their mass spectrometry data, and extracts the parent ion chromatogram of the sample to be detected as the target chemical substance in turn; corrects the parent ion chromatogram to obtain a target chromatogram containing several chromatographic peaks; then screens all chromatographic peaks to obtain candidate chromatographic peaks and their peak types; for each candidate chromatographic peak, traverses and matches the target chemical substance to obtain the relevant characteristic score of each candidate chromatographic peak corresponding to the target chemical substance; finally, the peak type and relevant characteristic score of the candidate chromatographic peak are used to sort each candidate chromatographic peak corresponding to the target chemical substance, and the first several candidate chromatographic peaks in the sorting result are output as the final screening result of the chemical substance in the sample to be detected.
- This embodiment can significantly increase the number of low-concentration chemical substances detected, with high sensitivity and high throughput.
- This embodiment provides a highly sensitive and high-throughput chemical substance annotation method, comprising:
- S1 is based on EISA technology, which collects parent ions and fragment ions in the sample to be detected.
- EISA technology collects all parent ions and fragment ions in the sample to be detected in one full scan, and the correspondence between fragment ions and parent ions cannot be determined.
- a chemical exposure group database is constructed based on the local database of chemical standards, the open source mass spectrometry database, and the mass spectrometry data of public literature;
- the chemical exposure group database includes several chemical substances and their mass spectrometry data, and the mass spectrometry data of each chemical substance includes the name of the chemical substance, the mass-to-charge ratio of the parent ion, the mass-to-charge ratio of the fragment ion, and the intensity;
- the chemical substances in the chemical exposure group database are taken as target chemical substances in turn, and the mass-to-charge ratio range of the parent ion that needs to be extracted is obtained according to the mass-to-charge ratio of the parent ion of the target chemical substance and the set mass accuracy standard; in the sample to be detected, for the ions whose parent ion mass-to-charge ratio is within the range of the parent ion mass-to-charge ratio, the operation of extracting and constructing the chromatogram is performed.
- the parent ion mass-to-charge ratio of the target chemical substance in the chemical exposure group database is 183.0991Da, and the preset mass accuracy standard is ⁇ 0.01Da
- the parent ion mass-to-charge ratio range in the sample to be detected that needs to be extracted is [183.0891, 183.1091]; as shown in Figure 3, in the sample to be detected, all ions whose parent ion mass-to-charge ratio is located in [183.0891, 183.1091] are extracted and the chromatogram is constructed;
- the chromatographic elution is set according to the actual situation, and the corresponding chromatogram parts before and after the elution are deleted, and only the chromatogram part during the elution process is retained, thereby reducing the occurrence of false positive features and making the detection result more accurate.
- the chromatographic system starts elution at 90s and ends elution at 900s, then the parts before 90s and after 900s of the entire chromatogram are deleted, and the part between 90s-900s is retained to obtain the target chromatogram.
- the peak type of the corresponding candidate chromatographic peak is classified as the first type; for candidate chromatographic peaks that cannot be detected by the existing mass spectrometry data analysis algorithm, if the sawtooth index is less than the preset sawtooth index threshold, the corresponding candidate The peak type of the chromatographic peak is classified as the second type; otherwise, the peak type of the corresponding candidate chromatographic peak is classified as the third type.
- the peak height threshold is 1000
- the sawtooth index is 0.2
- the existing mass spectrometry data analysis algorithm is the XCMS algorithm.
- the range and number of chromatographic peaks that need to be screened can be narrowed, which can effectively improve the subsequent matching speed and accuracy.
- Filtering chromatographic peaks with peak heights lower than a preset peak height threshold or sawtooth indexes greater than a preset sawtooth index threshold can reduce the number of candidate chromatographic peaks without affecting the accuracy of the test results, thereby improving the detection speed and reducing the false positive rate.
- the sawtooth index is considered, retaining features that cannot be detected by existing mass spectrometry data analysis algorithms, and expanding the scope of chemical substance identification.
- the third type of candidate chromatographic peaks are used as reference types, and whether to filter further is selected according to actual needs; if filtered, there are two types of peaks used for subsequent analysis, namely the first type and the second type; if not filtered, there are three types of peaks used for subsequent analysis.
- MFR i represents the first score of the fragment ion matched at the top of the i-th candidate chromatographic peak of the target chemical
- Ni represents the number of fragment ions matched at the top of the i-th candidate chromatographic peak
- NT represents the number of fragment ions of the target chemical in the chemical exposure group database
- the second score is calculated based on the intensity of the fragment ion matched at the top of the candidate chromatographic peak and the intensity of the fragment ion of the target chemical in the chemical exposure group database, and the calculation formula is:
- SSM i represents the second score of the fragment ion matched at the top of the i-th candidate chromatographic peak of the target chemical
- W Qi represents the intensity of the fragment ion matched at the top of the i-th candidate chromatographic peak
- W Ri represents the intensity of the i-th fragment ion of the target chemical in the chemical exposure group database
- Score i ⁇ MFR i + ⁇ SSM i
- Score i represents the relevant characteristic score of the candidate characteristic of the i-th candidate chromatographic peak corresponding to the target chemical substance
- S7 Sort each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and related characteristic score of the candidate chromatographic peak to obtain a sorting result; specifically:
- S7.1 Set the sorting priority of the peak types of candidate chromatographic peaks, from high to low, first type, second type, and third type;
- S7.3 Concatenate the candidate features of the target chemical substance within each peak type according to the sorting priority of the peak type to obtain a sorting result.
- S8 Output the first several candidate chromatographic peaks in the sorting results as the final screening results of the chemical substances in the sample to be detected.
- the method provided in this embodiment is compared with traditional peak extraction algorithms, such as the XCMS algorithm, the MZmine3 algorithm, and the MSDIAL algorithm, to perform characteristic detection of chemical substances at concentrations of 500 ppb and 0.8 ppb, respectively, and compared with the manual inspection results.
- the detection result comparison diagram is shown in Figure 7; it can be seen that the method of this embodiment detects the largest number of chemical substance characteristics at 500 ppb and 0.8 ppb, and has a higher sensitivity.
- the method provided in this embodiment is compared with the traditional TMM acquisition method.
- the object is a mixed standard containing 50 pollutants, and the characteristic number of pollutants at concentrations of 20 ppb, 4 ppb and 0.8 ppb is obtained.
- FIG8 it can be seen that the method provided in this embodiment can detect the pollutants at a lower concentration.
- the chemical characteristics of the samples vary little and far exceed the number of chemical characteristics detected in the traditional TMM acquisition mode.
- the method provided in this embodiment can still accurately detect chemical substances at a low concentration of 0.8 ppb, with high sensitivity and high throughput.
- the chemical exposure group database constructed includes 200 pesticides. It can be seen that the method of this embodiment identified 25 pesticides; and the TTM acquisition method identified 13 pesticides through manual inspection.
- the chemical substance with MFR equal to 0 indicates that no parent ion is matched; and the intensity of the parent ion identified by the method of this embodiment is much higher than the intensity of the parent ion identified by the TTM acquisition method, that is, the method provided in this embodiment can collect the characteristics of fragment ions without affecting the abundance of parent ions, thereby realizing high-sensitivity and high-throughput detection.
- This embodiment also provides a highly sensitive and high-throughput chemical substance annotation system, which is used to implement the annotation method described in Embodiment 1 or 2, as shown in FIG9 , comprising:
- a data acquisition module used to collect parent ions and fragment ions in the sample to be detected based on EISA technology
- a database construction module is used to construct a chemical exposure group database, including several chemical substances and their mass spectrometry data;
- a chromatogram generation module is used to set a mass accuracy standard, take the chemical substances in the chemical exposure group database as target chemical substances in turn, and extract the parent ion chromatogram of the sample to be tested accordingly;
- a chromatogram correction module is used to correct the parent ion chromatogram of the sample to be detected to obtain a target chromatogram; the target chromatogram includes a plurality of chromatographic peaks;
- a chromatographic peak screening module is used to screen all chromatographic peaks to obtain candidate chromatographic peaks and their peak types
- a data matching module is used to traverse and match target chemical substances for each candidate chromatographic peak, and calculate the relevant characteristic score of each candidate chromatographic peak corresponding to the target chemical substance;
- a sorting module used to sort each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and related characteristic score of the candidate chromatographic peak to obtain a sorting result
- the chemical substance detection module is used to output the first several candidate chromatographic peaks in the sorting results as the final screening results of the chemical substances in the sample to be detected.
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Pathology (AREA)
- Immunology (AREA)
- General Physics & Mathematics (AREA)
- General Health & Medical Sciences (AREA)
- Biochemistry (AREA)
- Analytical Chemistry (AREA)
- Engineering & Computer Science (AREA)
- Electrochemistry (AREA)
- Chemical Kinetics & Catalysis (AREA)
- Theoretical Computer Science (AREA)
- Computing Systems (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Crystallography & Structural Chemistry (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Other Investigation Or Analysis Of Materials By Electrical Means (AREA)
Abstract
一种高灵敏、高通量的化学物质注释方法与系统,涉及化学物质识别的技术领域,包括利用EISA技术采集待检测样品的母离子和碎片离子;构建化学暴露组数据库,依次作为目标化学物质提取待检测样品的母离子色谱图;校正后获得目标色谱图,对其中所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;对每个候选色谱峰均遍历匹配目标化学物质,获得目标化学物质对应的每个候选色谱峰的相关特征分值;基于峰值类型和相关特征分值对目标化学物质对应的每个候选色谱峰进行排序,将排序结果中前若干位的候选色谱峰作为待检测样品中化学物质的最终筛查结果输出。能够显著提高检测到的低浓度化学物质的数量,具有高灵敏度和高通量的优势。
Description
本发明涉及化学物质识别的技术领域,更具体地,涉及一种高灵敏、高通量的化学物质注释方法与系统。
化学物质注释是暴露组学研究的重要组成部分,暴露的概念最初是用来解释健康和疾病的环境驱动因素的。现在,其代表了整个生命周期中所有的外源性和内源性环境暴露以及相关的生物学效应。生物和环境样品中的异生物化学物质浓度范围很广,低浓度达到ppb或纳摩尔水平。因此,迫切需要一种具有广泛化学覆盖范围的高敏感度化学暴露定位技术来推动暴露类型研究的向前发展。传统化学物质注释需要通过串联质谱仪,分别获取母离子质谱(mass spectrometry,MS)数据和碎片离子的质谱数据。母离子又称作前体离子,通常在离子源内生成。碎片离子的生成通常需要将电喷雾离子源(ESI)与串联质谱相结合,当母离子进入碰撞室内后在能量作用下与惰性气体分子发生反应产生碎片离子。
基于三重四极杆质谱的靶向检测通常被认为是暴露体生物标志物监测中最敏感的技术。然而,这种方法只能监测有限数量的化学物质,限制了其在暴露研究中的应用。基于高分辨率质谱(high resolution mass spectrum,HRMS)的非靶向分析方法可以检测到广泛的化学物质,通常采用全扫描采集方法获取母离子的m/z(质荷比)和保留时间(Retention Time,RT)信息,然后使用数据依赖采集(DDA)模式或数据独立采集(DIA)模式生成质谱图(MS/MS)。数据依赖采集(DDA)模式是串联质谱中数据采集的一种主要模式。在DDA模式下,质谱仪器在执行MS全扫描时会自动对从全扫描谱图中选择的母离子列表进行MS/MS分析。DDA模式通常用于收集感兴趣的暴露特征的碎片光谱,母离子的选择取决于离子强度,那么低丰度暴露的质谱特征可能永远不会被选择在碰撞细胞中产生碎片DDA模式优点是选择性高,但通量较低、难以实现化学物质碎片离子的批量采集;数据独立采集(DIA)模式是质谱分析中一种批量采集化学物质碎片离子的模式。在该模式中,选定m/z范围内的所有离子都将在串联质谱的第二阶段进行碎片化和分析。该模式通过裂解在给定时间进入质谱仪的所有离子
或通过顺序分离和裂解m/z范围来获得串联质谱数据。DIA模式尽管可以实现化学物质碎片离子的批量采集,但选择性较低。
信息学算法在较低水平的化学物质注释中具有一定局限性。目前,在暴露研究中用于自主化学识别和注释的数据处理信息学算法主要借鉴了代谢组学领域的算法。然而,生物样品中内源性代谢物的浓度通常比同一样品中的外源性化学物质高几个数量级。即使在环境样本中,化学污染物的浓度也较低,通常在ppb左右。早期的研究表明,传统的代谢组学数据预处理软件如XCMS会遗漏大量的代谢特征,即低丰度和/或形状不良的特征,所以,传统的峰提取算法会导致特征丢失,影响下游的物质鉴定。
申请人前期提出了基于质谱离子源内高能裂解现象的化学物质高灵敏分析技术(EISA)。EISA作为一种简单有效的在电喷雾电离(ESI)源中生成源内片段的方法,改变了传统的串联质谱片段生成策略,通过优化源内碎片条件,源内片段生成使用EISA技术全扫描模式HRMS可以模拟中高能MS/MS片段谱的内生代谢产物和肽通过碰撞诱导解离(CID)技术在碰撞细胞,而不影响前体离子的强度,从而能够有效实现低浓度化学物质的鉴定。但是,传统的质谱数据分析软件如XCMS,却无法分析EISA数据。因此,亟需提出一种高灵敏、高通量的化学物质注释方法。
现有技术提供了一种化合物的碎片离子预测方法及应用,包括:1)确定用于实验的化合物,化合物的母离子分子式和碎片离子分子式;2)根据母离子和碎片离子的元素组成,选取稳定同位素;3)根据确定的化合物、母离子分子式和碎片离子分子式、母离子和碎片离子中稳定同位素元素个数,计算所述母离子和所述碎片离子被稳定同位素标记的情况;4)计算全部母离子和全部碎片离子被所述稳定同位素标记情况的个数;5)根据步骤4)中得到的标记情况的数量,计算全部标记情况中母离子和碎片离子的精确质量数,并计算质荷比;6)建立化合物的碎片离子的质荷比数据库。该现有技术需要明确用于实验的化合物中的母离子和碎片离子的关系,而EISA技术是在一次全扫描中采集所有的母离子与碎片离子,母离子与碎片离子的关系是不明确的,不适用于EISA数据。
发明内容
本发明为克服上述现有技术无法准确分析EISA数据,从而无法实现高灵敏度、高通量的化学物质检测的缺陷,提供一种高灵敏、高通量的化学物质注释方
法与系统,显著提高检测到的低浓度化学物质的数量,灵敏度高,通量高。
为解决上述技术问题,本发明的技术方案如下:
本发明提供了一种高灵敏、高通量的化学物质注释方法,包括:
S1:基于EISA技术,采集待检测样品中的母离子和碎片离子;
S2:构建化学暴露组数据库,包括若干种化学物质及其质谱数据;
S3:设置质量精度标准,将所述化学暴露组数据库中的化学物质依次作为目标化学物质,对应提取待检测样品的母离子色谱图;
S4:对待检测样品的母离子色谱图进行校正,获得目标色谱图;所述目标色谱图包括若干个色谱峰;
S5:对所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;
S6:对于每个候选色谱峰,均遍历匹配目标化学物质,并计算目标化学物质对应的每个候选色谱峰的相关特征分值;
S7:基于候选色谱峰的峰值类型和相关特征分值,对目标化学物质对应的每个候选色谱峰进行排序,获得排序结果;
S8:将排序结果中前若干位的候选色谱峰,作为待检测样品中化学物质的最终筛查结果输出。
优选地,所述步骤S2中,基于化学标准品本地数据库、开源质谱数据库和公开文献的质谱数据,构建化学暴露组数据库;所述化学暴露组数据库包括若干种化学物质及其质谱数据,每种化学物质的质谱数据包括化学物质名称、母离子质荷比、碎片离子质荷比和强度。
优选地,所述步骤S3中,设置质量精度标准,将所述化学暴露组数据库中的化学物质依次作为目标化学物质,对应提取待检测样品的母离子色谱图的具体方法为:
将所述化学暴露组数据库中的化学物质依次作为目标化学物质,根据目标化学物质的母离子质荷比,结合设置的质量精度标准,获得需要提取的母离子质荷比范围;在待检测样品中,对于母离子质荷比位于母离子质荷比范围内的离子,进行提取构建色谱图的操作。
优选地,所述步骤S4中,对待检测样品的母离子色谱图进行校正,获得目标色谱图的具体方法为:
对于提取到的待检测样品的色谱图,将其在色谱洗脱开始前和色谱洗脱结束
后对应的色谱图部分进行删除,将剩余部分的色谱图作为目标色谱图,用于后续分析。
所述色谱洗脱根据实际情况设定,删除洗脱开始前和洗脱结束后对应的色谱图部分,仅保留洗脱过程中的色谱图部分,减少了假阳性特征的出现,使检测结果更加准确。
优选地,所述步骤S5的具体方法为:
S5.1:计算所有色谱峰的峰高和锯齿形指数;
S5.2:将峰高低于预设峰高阈值的色谱峰过滤,将剩余色谱峰作为候选色谱峰;
S5.3:利用现有的质谱数据分析算法对候选色谱峰进行检测;
S5.4:对于能够被现有的质谱数据分析算法检测到的候选色谱峰,将对应的候选色谱峰的峰值类型分类为第一类型;对于无法被现有的质谱数据分析算法检测到的候选色谱峰,若其锯齿形指数小于预设锯齿形指数阈值,则将对应的候选色谱峰的峰值类型分类为第二类型;否则,将对应的候选色谱峰的峰值类型分类为第三类型。
过滤峰高低于预设峰高阈值的色谱峰,能够在不影响检测结果准确性的前提下减少候选色谱峰的数量,提高了检测速度并减少了假阳性特征的出现。其次,对候选色谱峰的峰值类型进行分类时,不仅考虑了能否被现有的质谱数据分析算法检测到,还考虑了锯齿形指数,保留了不能被现有的质谱数据分析算法检测到的特征,扩大了化学物质鉴定的范围。将第三类型的候选色谱峰作为参考类型,根据实际需求选择是否进一步过滤;若过滤,则用于后续分析的峰值类型为两种,即第一类型和第二类应;若不过滤,则用于后续分析的峰值类型为三种。
优选地,在步骤S5.1中,若化学暴露组数据库中目标化学物质存在保留时间,则只计算保留时间窗口中出现的色谱峰的峰高和锯齿形指数。
通过使用保留时间窗口,缩小了需要筛查的色谱峰的范围和数量,能够有效提高后续的匹配速度和准确性。
优选地,所述步骤S6的具体方法为:
S6.1:对于每个候选色谱峰,将候选色谱峰顶端对应的碎片离子与化学暴露组数据库中目标化学物质的碎片离子依次进行匹配,获得每个候选色谱峰顶端处匹配的碎片离子及其数量、质荷比和强度;
S6.2:基于候选色谱峰顶端处匹配的碎片离子和化学暴露组数据库中目标化学物质的碎片离子,计算第一分值和第二分值;
S6.3:基于第一分值和第二分值,计算目标化学物质对应的候选色谱峰的相关特征分值;
S6.4:重复步骤S6.1-S6.3,获得目标化学物质对应的每个候选色谱峰的相关特征分值。
优选地,所述步骤S6.2中,计算第一分值的具体方法为:
基于候选色谱峰顶端处匹配的碎片离子的数量和化学暴露组数据库中目标化学物质的碎片离子的数量,计算第一分值,计算公式为:
式中,MFRi表示目标化学物质的第i个候选色谱峰顶端处匹配的碎片离子的第一分值,Ni表示第i个候选色谱峰顶端处匹配的碎片离子的数量,NT表示化学暴露组数据库中目标化学物质的碎片离子的数量。
优选地,所述步骤S6.2中,计算第二分值的具体方法为:
基于候选色谱峰顶端处匹配的碎片离子的强度与化学暴露组数据库中目标化学物质的碎片离子的强度,计算第二分值,计算公式为:
式中,SSMi表示目标化学物质的第i个候选色谱峰顶端处匹配的碎片离子的第二分值,WQi表示第i个候选色谱峰顶端处匹配的碎片离子的强度,WRi表示化学暴露组数据库中目标化学物质的第i个碎片离子的强度。
优选地,所述步骤S6.3中,基于第一分值和第二分值,计算目标化学物质对应的候选色谱峰的相关特征分值的具体方法为:
对第一分值和第二分值分别赋予权重系数,计算目标化学物质对应的候选特征的相关特征分值:
Scorei=αMFRi+βSSMi
Scorei=αMFRi+βSSMi
式中,Scorei表示目标化学物质对应的第i个候选色谱峰的候选特征的相关特征分值,α,β分别表示第一、二权重系数。
优选地,所述步骤S7的具体方法为:
S7.1:设置候选色谱峰的峰值类型的排序优先级,从高到低依次为第一类型、第二类型、第三类型;
S7.2:对于属于同一峰值类型的所有候选色谱峰,将其相关特征分值按照从大到小的顺序进行排列,获得每种峰值类型内的目标化学物质候选特征;
S7.3:将每种峰值类型内的目标化学物质候选特征按照峰值类型的排序优先级进行拼接,获得排序结果。
本发明还提供了一种高灵敏、高通量的化学物质注释系统,用于实现上述的注释方法,包括:
数据采集模块,用于基于EISA技术,采集待检测样品中的母离子和碎片离子;
数据库构建模块,用于构建化学暴露组数据库,包括若干种化学物质及其质谱数据;
色谱图生成模块,用于设置质量精度标准,将所述化学暴露组数据库中的化学物质依次作为目标化学物质,对应提取待检测样品的母离子色谱图;
色谱图校正模块,用于对待检测样品的母离子色谱图进行校正,获得目标色谱图;所述目标色谱图包括若干个色谱峰;
色谱峰筛选模块,用于对所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;
数据匹配模块,用于对于每个候选色谱峰,均遍历匹配目标化学物质,并计算目标化学物质对应的每个候选色谱峰的相关特征分值;
排序模块,用于基于候选色谱峰的峰值类型和相关特征分值对目标化学物质对应的每个候选色谱峰进行排序,获得排序结果;
化学物质检测模块,用于将排序结果中前若干位的候选色谱峰,作为待检测样品中化学物质的最终筛查结果输出。
与现有技术相比,本发明技术方案的有益效果是:
本发明首先利用EISA技术采集待检测样品的所有母离子和碎片离子;之后构建包括若干种化学物质及其质谱数据的化学暴露组数据库,将化学物质依次作为目标化学物质提取待检测样品的母离子色谱图;对母离子色谱图进行校正,获得包含若干色谱峰的目标色谱图;然后对所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;对每个候选色谱峰,均遍历匹配目标化学物质,获得目标化学物
质对应的每个候选色谱峰的相关特征分值;最后基于候选色谱峰的峰值类型和相关特征分值对目标化学物质对应的每个候选色谱峰进行排序,将排序结果中前若干位的候选色谱峰,作为待检测样品中化学物质的最终筛查结果输出。本发明能够显著提高检测到的低浓度化学物质的数量,灵敏度高,通量高。
图1为实施例1所述的一种高灵敏、高通量的化学物质注释方法的流程图。
图2为实施例2所述的化学暴露组数据库的示意图。
图3为实施例2所述的基于预设的质量精度标准对待检测样品提取的色谱图。
图4为实施例2所述的目标色谱图。
图5为实施例2所述的存在保留时间时,作为处理对象的色谱峰。
图6为实施例2所述的不存在保留时间时,作为处理对象的色谱峰。
图7为实施例2所述的本实施例提供的方法与传统的峰提取算法在不同浓度下的检测结果对比示意图。
图8为实施例2所述的本实施例提供的方法与传统的TMM采集方法在不同浓度下的检测结果对比示意图。
图9为实施例3所述的一种高灵敏、高通量的化学物质注释系统的结构示意图。
附图仅用于示例性说明,不能理解为对本专利的限制;
为了更好说明本实施例,附图某些部件会有省略、放大或缩小,并不代表实际产品的尺寸;
对于本领域技术人员来说,附图中某些公知结构及其说明可能省略是可以理解的。
下面结合附图和实施例对本发明的技术方案做进一步的说明。
实施例1
本实施例提供了一种高灵敏、高通量的化学物质注释方法,如图1所示,包括:
S1:基于EISA技术,采集待检测样品中的母离子和碎片离子;
S2:构建化学暴露组数据库,包括若干种化学物质及其质谱数据;
S3:设置质量精度标准,将所述化学暴露组数据库中的化学物质依次作为目标化学物质,对应提取待检测样品的母离子色谱图;
S4:对待检测样品的母离子色谱图进行校正,获得目标色谱图;所述目标色谱图包括若干个色谱峰;
S5:对所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;
S6:对于每个候选色谱峰,均遍历匹配目标化学物质,并计算目标化学物质对应的每个候选色谱峰的相关特征分值;
S7:基于候选色谱峰的峰值类型和相关特征分值对目标化学物质对应的每个候选色谱峰进行排序,获得排序结果;
S8:将排序结果中前若干位的候选色谱峰,作为待检测样品中化学物质的最终筛查结果输出。
在具体实施过程中,本实施例首先利用EISA技术采集待检测样品的所有母离子和碎片离子;之后构建包括若干种化学物质及其质谱数据的化学暴露组数据库,依次作为目标化学物质提取待检测样品的母离子色谱图;对母离子色谱图进行校正,获得包含若干色谱峰的目标色谱图;然后对所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;对每个候选色谱峰,均遍历匹配目标化学物质,获得目标化学物质对应的每个候选色谱峰的相关特征分值;最后候选色谱峰的峰值类型和相关特征分值对目标化学物质对应的每个候选色谱峰进行排序,将排序结果中前若干位的候选色谱峰,作为待检测样品中化学物质的最终筛查结果输出。本实施例能够显著提高检测到的低浓度化学物质的数量,灵敏度高,通量高。
实施例2
本实施例提供了一种高灵敏、高通量的化学物质注释方法,包括:
S1基于EISA技术,采集待检测样品中的母离子和碎片离子;EISA技术是在一次全扫描中采集待检测样品中的全部母离子和碎片离子,无法确定碎片离子与母离子的对应关系;
S2:构建化学暴露组数据库,包括若干种化学物质及其质谱数据;
如图2所示,基于化学标准品本地数据库、开源质谱数据库和公开文献的质谱数据,构建化学暴露组数据库;所述化学暴露组数据库包括若干种化学物质及其质谱数据,每种化学物质的质谱数据包括化学物质名称、母离子质荷比、碎片离子质荷比和强度;
S3:设置质量精度标准,将所述化学暴露组数据库中的化学物质依次作为目标化学物质,对应提取待检测样品的母离子色谱图;
具体的,将所述化学暴露组数据库中的化学物质依次作为目标化学物质,根据目标化学物质的母离子质荷比,结合设置的质量精度标准,获得需要提取的母离子质荷比范围;在待检测样品中,对于母离子质荷比位于母离子质荷比范围内的离子,进行提取构建色谱图的操作。如化学暴露组数据库中目标化学物质的母离子质荷比为183.0991Da,预设的质量精度标准为±0.01Da,则需要提取的待检测样品中的母离子质荷比范围[183.0891,183.1091];如图3所示,在待检测样品中,对所有母离子质荷比位于[183.0891,183.1091]的离子,均提取并构建色谱图;
S4:对于提取到的待检测样品的色谱图,将其在色谱洗脱开始前和色谱洗脱结束后对应的色谱图部分进行删除,将剩余部分的色谱图作为目标色谱图,用于后续分析;每张所述目标色谱图包括若干个色谱峰;
如图4所示,所述色谱洗脱根据实际情况设定,删除洗脱开始前和洗脱结束后对应的色谱图部分,仅保留洗脱过程中的色谱图部分,减少了假阳性特征的出现,使检测结果更加准确;本实施例中,色谱系统在90s时开始洗脱,900s时结束洗脱,则将整个色谱图中90s前和900s后的部分删除,90s-900s之间的部分保留,获得目标色谱图。
S5:对所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;具体的:
S5.1:若化学暴露组数据库中目标化学物质存在保留时间,则只计算保留时间窗口中出现的色谱峰的峰高和锯齿形指数;如图5所示,若目标化学物质在化学暴露组数据库中存在保留时间为4.822min,转化为289.32s;则根据预设的时间窗口在色谱图中提取色谱峰并进行后续分析,本实施例中,时间窗口为±90s,289.32±90s时间段内出现的色谱峰的峰高和锯齿形指数;
如图6所述,若化学暴露组数据库中目标化学物质不存在保留时间,则计算所有出现的色谱峰的峰高和锯齿形指数;
S5.2:将峰高低于预设峰高阈值的色谱峰过滤,将剩余色谱峰作为候选色谱峰;
S5.3:利用现有的质谱数据分析算法对候选色谱峰进行检测;
S5.4:对于能够被现有的质谱数据分析算法检测到的候选色谱峰,将对应的候选色谱峰的峰值类型分类为第一类型;对于无法被现有的质谱数据分析算法检测到的候选色谱峰,若其锯齿形指数小于预设锯齿形指数阈值,则将对应的候选
色谱峰的峰值类型分类为第二类型;否则,将对应的候选色谱峰的峰值类型分类为第三类型。
本实施例中,峰高阈值为1000,锯齿形指数为0.2,现有的质谱数据分析算法为XCMS算法。
化学暴露组数据库中目标化学物质存在保留时间的话,可以缩小需要筛查的色谱峰的范围和数量,能够有效提高后续的匹配速度和准确性。
过滤峰高低于预设峰高阈值或锯齿形指数大于预设锯齿形指数阈值的色谱峰,能够在不影响检测结果准确性的前提下减少候选色谱峰的数量,提高了检测速度和减少了假阳性率。并且,对候选色谱峰的峰值类型进行分类时,不仅考虑了能否被现有的质谱数据分析算法检测到,还考虑了锯齿形指数,保留了不能被现有的质谱数据分析算法检测到的特征,扩大了化学物质鉴定的范围。将第三类型的候选色谱峰作为参考类型,根据实际需求选择是否进一步过滤;若过滤,则用于后续分析的峰值类型为两种,即第一类型和第二类应;若不过滤,则用于后续分析的峰值类型为三种。
S6:对于每个候选色谱峰,均遍历匹配目标化学物质,并计算目标化学物质对应的每个候选色谱峰的相关特征分值;具体的:
S6.1:对于每个候选色谱峰,将候选色谱峰顶端对应的碎片离子与化学暴露组数据库中目标化学物质的碎片离子依次进行匹配,获得每个候选色谱峰顶端处匹配的碎片离子及其数量、质荷比和强度;
S6.2:基于候选色谱峰顶端处匹配的碎片离子的数量和化学暴露组数据库中目标化学物质的碎片离子的数量,计算第一分值,计算公式为:
式中,MFRi表示目标化学物质的第i个候选色谱峰顶端处匹配的碎片离子的第一分值,Ni表示第i个候选色谱峰顶端处匹配的碎片离子的数量,NT表示化学暴露组数据库中目标化学物质的碎片离子的数量;
基于候选色谱峰顶端处匹配的碎片离子的强度与化学暴露组数据库中目标化学物质的碎片离子的强度,计算第二分值,计算公式为:
式中,SSMi表示目标化学物质的第i个候选色谱峰顶端处匹配的碎片离子的第二分值,WQi表示第i个候选色谱峰顶端处匹配的碎片离子的强度,WRi表示化学暴露组数据库中目标化学物质的第i个碎片离子的强度;
S6.3:基于第一分值和第二分值,计算目标化学物质对应的候选色谱峰的相关特征分值的具体方法为:
对第一分值和第二分值分别赋予权重系数,计算目标化学物质对应的候选特征的相关特征分值:
Scorei=αMFRi+βSSMi
Scorei=αMFRi+βSSMi
式中,Scorei表示目标化学物质对应的第i个候选色谱峰的候选特征的相关特征分值,α,β分别表示第一、二权重系数;本实施例中,α=0.7,β=0.3;
S6.4:重复步骤S6.1-S6.3,获得目标化学物质对应的每个候选色谱峰的相关特征分值。
S7:基于候选色谱峰的峰值类型和相关特征分值对目标化学物质对应的每个候选色谱峰进行排序,获得排序结果;具体的:
S7.1:设置候选色谱峰的峰值类型的排序优先级,从高到低依次为第一类型、第二类型、第三类型;
S7.2:对于属于同一峰值类型的所有候选色谱峰,将其相关特征分值按照从大到小的顺序进行排列,获得每种峰值类型内的目标化学物质候选特征;
S7.3:将每种峰值类型内的目标化学物质候选特征按照峰值类型的排序优先级进行拼接,获得排序结果。
S8:将排序结果中前若干位的候选色谱峰,作为待检测样品中化学物质的最终筛查结果输出。
在具体实施过程中,将本实施例提供的方法与传统的峰提取算法,如XCMS算法、MZmine3算法、MSDIAL算法,分别在浓度为500ppb和0.8ppb进行化学物质的特征检测,并且与手动检查结果进行了比对,检测结果对比示意图如图7所示;可以看出,本实施例的方法在500ppb和0.8ppb下检测出的化学物质特征的数量均为最多的,灵敏度更高。
将本实施例提供的方法与传统的TMM采集方法进行对比,对象为包含50个污染物的混标,获得在浓度为20ppb、4ppb和0.8ppb下的污染物特征数量;如图8所示,可以看出,本实施例提供的方法可以在浓度降低的情况下,检测出
的化学物质特征变化很小,并且远远超过传统的TMM采集模式下检测到的化学物质特征数量。本实施例提供的方法在低浓度0.8ppb下仍能准确检测到化学物质,灵敏度高,通量高。
进一步设置实验进行验证,将灰尘样品作为待检测样品,分别利用本实施例的方法和传统的TMM采集方法进行检测,检测结果如下表所示:
本实验中,构建的化学暴露组数据库包括200种农药,可以看出,本实施例的方法鉴定出了25种农药;而TTM采集方法通过人工检查识别出13中农药,MFR等于0的化学物质说明没有匹配到母离子;并且,利用本实施例方法鉴定的母离子的强度远高于用TTM采集方法识别的母离子强度,即本实施例提供的方法可以在不影响母离子丰度的前提下,采集碎片离子的特征,实现高灵敏度、高通量的检测。
实施例3
本实施例还提供了一种高灵敏、高通量的化学物质注释系统,用于实现实施例1或2所述的注释方法,如图9所示,包括:
数据采集模块,用于基于EISA技术,采集待检测样品中的母离子和碎片离子;
数据库构建模块,用于构建化学暴露组数据库,包括若干种化学物质及其质谱数据;
色谱图生成模块,用于设置质量精度标准,将所述化学暴露组数据库中的化学物质依次作为目标化学物质,对应提取待检测样品的母离子色谱图;
色谱图校正模块,用于对待检测样品的母离子色谱图进行校正,获得目标色谱图;所述目标色谱图包括若干个色谱峰;
色谱峰筛选模块,用于对所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;
数据匹配模块,用于对于每个候选色谱峰,均遍历匹配目标化学物质,并计算目标化学物质对应的每个候选色谱峰的相关特征分值;
排序模块,用于基于候选色谱峰的峰值类型和相关特征分值对目标化学物质对应的每个候选色谱峰进行排序,获得排序结果;
化学物质检测模块,用于将排序结果中前若干位的候选色谱峰,作为待检测样品中化学物质的最终筛查结果输出。
相同或相似的标号对应相同或相似的部件;
附图中描述位置关系的用语仅用于示例性说明,不能理解为对本专利的限制;
显然,本发明的上述实施例仅仅是为清楚地说明本发明所作的举例,而并非是对本发明的实施方式的限定。对于所属领域的普通技术人员来说,在上述说明的基础上还可以做出其它不同形式的变化或变动。这里无需也无法对所有的实施方式予以穷举。凡在本发明的精神和原则之内所作的任何修改、等同替换和改进等,均应包含在本发明权利要求的保护范围之内。
Claims (10)
- 一种高灵敏、高通量的化学物质注释方法,其特征在于,包括:S1:基于EISA技术,采集待检测样品中的母离子和碎片离子;S2:构建化学暴露组数据库,包括若干种化学物质及其质谱数据;S3:设置质量精度标准,将所述化学暴露组数据库中的化学物质依次作为目标化学物质,对应提取待检测样品的母离子色谱图;S4:对待检测样品的母离子色谱图进行校正,获得目标色谱图;所述目标色谱图包括若干个色谱峰;S5:对所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;S6:对于每个候选色谱峰,均遍历匹配目标化学物质,并计算目标化学物质对应的每个候选色谱峰的相关特征分值;S7:基于候选色谱峰的峰值类型和相关特征分值,对目标化学物质对应的每个候选色谱峰进行排序,获得排序结果;S8:将排序结果中前若干位的候选色谱峰,作为待检测样品中化学物质的最终筛查结果输出。
- 根据权利要求1所述的高灵敏、高通量的化学物质注释方法,其特征在于,所述步骤S2中,基于化学标准品本地数据库、开源质谱数据库和公开文献的质谱数据,构建化学暴露组数据库;所述化学暴露组数据库包括若干种化学物质及其质谱数据,每种化学物质的质谱数据包括化学物质名称、母离子质荷比、碎片离子质荷比和强度。
- 根据权利要求1所述的高灵敏、高通量的化学物质注释方法,其特征在于,所述步骤S4中,对待检测样品的母离子色谱图进行校正,获得目标色谱图的具体方法为:对于提取到的待检测样品的色谱图,将其在色谱洗脱开始前和色谱洗脱结束后对应的色谱图部分进行删除,将剩余部分的色谱图作为目标色谱图,用于后续分析。
- 根据权利要求2或3所述的高灵敏、高通量的化学物质注释方法,其特征在于,所述步骤S5的具体方法为:S5.1:计算所有色谱峰的峰高和锯齿形指数;S5.2:将峰高低于预设峰高阈值的色谱峰过滤,将剩余色谱峰作为候选色谱峰;S5.3:利用现有的质谱数据分析算法对候选色谱峰进行检测;S5.4:对于能够被现有的质谱数据分析算法检测到的候选色谱峰,将对应的候选色谱峰的峰值类型分类为第一类型;对于无法被现有的质谱数据分析算法检测到的候选色谱峰,若其锯齿形指数小于预设锯齿形指数阈值,则将对应的候选色谱峰的峰值类型分类为第二类型;否则,将对应的候选色谱峰的峰值类型分类为第三类型。
- 根据权利要求4所述的高灵敏、高通量的化学物质注释方法,其特征在于,所述步骤S6的具体方法为:S6.1:对于每个候选色谱峰,将候选色谱峰顶端对应的碎片离子与化学暴露组数据库中目标化学物质的碎片离子依次进行匹配,获得每个候选色谱峰顶端处匹配的碎片离子及其数量、质荷比和强度;S6.2:基于候选色谱峰顶端处匹配的碎片离子和化学暴露组数据库中目标化学物质的碎片离子,计算第一分值和第二分值;S6.3:基于第一分值和第二分值,计算目标化学物质对应的候选色谱峰的相关特征分值;S6.4:重复步骤S6.1-S6.3,获得目标化学物质对应的每个候选色谱峰的相关特征分值。
- 根据权利要求5所述的高灵敏、高通量的化学物质注释方法,其特征在于,所述步骤S6.2中,计算第一分值的具体方法为:基于候选色谱峰顶端处匹配的碎片离子的数量和化学暴露组数据库中目标化学物质的碎片离子的数量,计算第一分值,计算公式为:
式中,MFRi表示目标化学物质的第i个候选色谱峰顶端处匹配的碎片离子的第一分值,Ni表示第i个候选色谱峰顶端处匹配的碎片离子的数量,NT表示化学暴露组数据库中目标化学物质的碎片离子的数量。 - 根据权利要求5所述的高灵敏、高通量的化学物质注释方法,其特征在于,所述步骤S6.2中,计算第二分值的具体方法为:基于候选色谱峰顶端处匹配的碎片离子的强度与化学暴露组数据库中目标化学物质的碎片离子的强度,计算第二分值,计算公式为:
式中,SSMi表示目标化学物质的第i个候选色谱峰顶端处匹配的碎片离子的第二分值,WQi表示第i个候选色谱峰顶端处匹配的碎片离子的强度,WRi表示化学暴露组数据库中目标化学物质的第i个碎片离子的强度。 - 根据权利要求6或7所述的高灵敏、高通量的化学物质注释方法,其特征在于,所述步骤S6.3中,基于第一分值和第二分值,计算目标化学物质对应的候选色谱峰的相关特征分值的具体方法为:对第一分值和第二分值分别赋予权重系数,计算目标化学物质对应的候选特征的相关特征分值:
Scorei=αMFRi+βSSMi式中,Scorei表示目标化学物质对应的第i个候选色谱峰的候选特征的相关特征分值,α,β分别表示第一、二权重系数。 - 根据权利要求8所述的高灵敏、高通量的化学物质注释方法,其特征在于,所述步骤S7的具体方法为:S7.1:设置候选色谱峰的峰值类型的排序优先级,从高到低依次为第一类型、第二类型、第三类型;S7.2:对于属于同一峰值类型的所有候选色谱峰,将其相关特征分值按照从大到小的顺序进行排列,获得每种峰值类型内的目标化学物质候选特征;S7.3:将每种峰值类型内的目标化学物质候选特征按照峰值类型的排序优先级进行拼接,获得排序结果。
- 一种高灵敏、高通量的化学物质注释系统,用于实现权利要求1-9任一项所述的注释方法,其特征在于,包括:数据采集模块,用于基于EISA技术,采集待检测样品中的母离子和碎片离子;数据库构建模块,用于构建化学暴露组数据库,包括若干种化学物质及其质谱数据;色谱图生成模块,用于设置质量精度标准,将所述化学暴露组数据库中的化 学物质依次作为目标化学物质,对应提取待检测样品的母离子色谱图;色谱图校正模块,用于对待检测样品的母离子色谱图进行校正,获得目标色谱图;所述目标色谱图包括若干个色谱峰;色谱峰筛选模块,用于对所有色谱峰进行筛选,获得候选色谱峰及其峰值类型;数据匹配模块,用于对于每个候选色谱峰,均遍历匹配目标化学物质,并计算目标化学物质对应的每个候选色谱峰的相关特征分值;排序模块,用于基于候选色谱峰的峰值类型和相关特征分值对目标化学物质对应的每个候选色谱峰进行排序,获得排序结果;化学物质检测模块,用于将排序结果中前若干位的候选色谱峰,作为待检测样品中化学物质的最终筛查结果输出。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/113115 WO2025035386A1 (zh) | 2023-08-15 | 2023-08-15 | 一种高灵敏、高通量的化学物质注释方法与系统 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/113115 WO2025035386A1 (zh) | 2023-08-15 | 2023-08-15 | 一种高灵敏、高通量的化学物质注释方法与系统 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025035386A1 true WO2025035386A1 (zh) | 2025-02-20 |
Family
ID=94631885
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/113115 Pending WO2025035386A1 (zh) | 2023-08-15 | 2023-08-15 | 一种高灵敏、高通量的化学物质注释方法与系统 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025035386A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN120823904A (zh) * | 2025-07-16 | 2025-10-21 | 西安理工大学 | 基于列表依赖模式的植物油氧化甘油三酯高覆盖分析方法 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20060172430A1 (en) * | 2005-01-31 | 2006-08-03 | Kelleher Neil L | Identification and characterization of protein fragments |
| CN104215729A (zh) * | 2014-08-18 | 2014-12-17 | 中国科学院计算技术研究所 | 串联质谱数据母离子检测模型训练方法及母离子检测方法 |
| US20150160231A1 (en) * | 2013-12-06 | 2015-06-11 | Premier Biosoft | Identification of metabolites from tandem mass spectrometry data using databases of precursor and product ion data |
| CN113758989A (zh) * | 2021-08-26 | 2021-12-07 | 清华大学深圳国际研究生院 | 基于碎片树的现场质谱目标物识别以及衍生物预测方法 |
| CN114965662A (zh) * | 2022-07-25 | 2022-08-30 | 广东工业大学 | 一种化学物质注释方法 |
| CN115453009A (zh) * | 2022-10-14 | 2022-12-09 | 广东工业大学 | 一种不依赖保留时间的化学物质注释方法 |
-
2023
- 2023-08-15 WO PCT/CN2023/113115 patent/WO2025035386A1/zh active Pending
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20060172430A1 (en) * | 2005-01-31 | 2006-08-03 | Kelleher Neil L | Identification and characterization of protein fragments |
| US20150160231A1 (en) * | 2013-12-06 | 2015-06-11 | Premier Biosoft | Identification of metabolites from tandem mass spectrometry data using databases of precursor and product ion data |
| CN104215729A (zh) * | 2014-08-18 | 2014-12-17 | 中国科学院计算技术研究所 | 串联质谱数据母离子检测模型训练方法及母离子检测方法 |
| CN113758989A (zh) * | 2021-08-26 | 2021-12-07 | 清华大学深圳国际研究生院 | 基于碎片树的现场质谱目标物识别以及衍生物预测方法 |
| CN114965662A (zh) * | 2022-07-25 | 2022-08-30 | 广东工业大学 | 一种化学物质注释方法 |
| CN115453009A (zh) * | 2022-10-14 | 2022-12-09 | 广东工业大学 | 一种不依赖保留时间的化学物质注释方法 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN120823904A (zh) * | 2025-07-16 | 2025-10-21 | 西安理工大学 | 基于列表依赖模式的植物油氧化甘油三酯高覆盖分析方法 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN109828068B (zh) | 质谱数据采集及分析方法 | |
| WO2022262132A1 (zh) | 一种样品未知成分的液质联用非靶向分析方法 | |
| US7538321B2 (en) | Method of identifying substances using mass spectrometry | |
| Wallace et al. | High-resolution mass spectrometry | |
| CN117250267A (zh) | 一种新污染物非靶向筛查的高分辨质谱数据处理方法 | |
| EP4078600B1 (en) | Method and system for the identification of compounds in complex biological or environmental samples | |
| CN115453009B (zh) | 一种不依赖保留时间的化学物质注释方法 | |
| WO2026051643A1 (zh) | 一种基于数据分析的质谱仪分辨率提升方法及系统 | |
| GB2575168A (en) | Precursor selection for data-dependent tandem mass spectrometry | |
| CN118566391B (zh) | 一种基于液质联用的溴代污染物及其代谢物多策略注释方法 | |
| CN117789848A (zh) | 一种应用特征碎片及特征碎片组辅助非靶向筛查的方法 | |
| WO2025035386A1 (zh) | 一种高灵敏、高通量的化学物质注释方法与系统 | |
| JP4929149B2 (ja) | 質量分析スペクトル分析方法 | |
| CN111551626A (zh) | 基于分子组成和结构指纹识别的串级质谱解析方法 | |
| CN114965662A (zh) | 一种化学物质注释方法 | |
| US20240369516A1 (en) | Method of bioanalytical analysis utilizing ion spectrometry, including mass analysis | |
| CN106908527B (zh) | 一种鉴别荔枝蜜产地的方法 | |
| CN117110466B (zh) | 一种高灵敏、高通量的化学物质注释方法与系统 | |
| WO2024223367A1 (en) | A method for predicting one or more sample metric values by machine learning | |
| CN113740408B (zh) | 用于确定ms扫描数据中的干扰、过滤离子并对样品进行质谱分析的方法和装置 | |
| CN118534008A (zh) | 一种基于非靶向代谢组学技术的酱香型白酒组分分析方法及应用 | |
| CN115932142A (zh) | 一种谱图的解析方法和装置 | |
| CN116298036A (zh) | 一种四维代谢组学数据处理方法 | |
| JP2007121134A (ja) | タンデム質量分析システム | |
| CN115236165A (zh) | 基于直接电离质谱的爆炸物检测方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23948818 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |