EP4458989A1 - Kit for identifying tumor tissue-of-origin and data analysis method - Google Patents

Kit for identifying tumor tissue-of-origin and data analysis method Download PDF

Info

Publication number
EP4458989A1
EP4458989A1 EP23189193.8A EP23189193A EP4458989A1 EP 4458989 A1 EP4458989 A1 EP 4458989A1 EP 23189193 A EP23189193 A EP 23189193A EP 4458989 A1 EP4458989 A1 EP 4458989A1
Authority
EP
European Patent Office
Prior art keywords
kit
tumor tissue
identifying tumor
origin
cancer
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
EP23189193.8A
Other languages
German (de)
French (fr)
Other versions
EP4458989B1 (en
Inventor
Hongcang Gu
Yunfei Wang
Xianrong CHE
Xianghe Meng
Changxiao Xie
Qiongwan XIE
Wenjun Wang
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hangzhou Shengting Medical Technology Ltd
Original Assignee
Hangzhou Shengting Medical Technology Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hangzhou Shengting Medical Technology Ltd filed Critical Hangzhou Shengting Medical Technology Ltd
Publication of EP4458989A1 publication Critical patent/EP4458989A1/en
Application granted granted Critical
Publication of EP4458989B1 publication Critical patent/EP4458989B1/en
Active legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Images

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • C12Q1/6886Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • G16B20/20Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • G16B30/10Sequence alignment; Homology search
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B40/00ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/154Methylation markers

Definitions

  • the present disclosure relates to the technical field of biomedicine, in particular to a kit for identifying tumor tissue-of-origin and a data analysis method for a kit for identifying tumor tissue-of-origin.
  • Cancer metastasis refers to a process in which tumor cells fall off from a primary lesion, enter a circulatory system, transfer to other parts of the body, and continue to grow.
  • the original tumor is called “primary lesion”
  • the newly formed tumor is called “secondary lesion” or “metastatic lesion”.
  • cancer metastasis can be divided into early dissemination model and late dissemination model. A difference between the two models lies primarily in an environment in which metastatic cancer cells mutate. In the late dissemination model, cancer cells mutate and proliferate at the primary lesion, and then metastasize through the circulatory system.
  • the metastatic tumor has a high degree of genetic similarity with the primary tumor, and can be treated with same treatment methods as the primary tumor.
  • cancer cells mutate after passing through the circulatory system to reach their destination.
  • the metastatic tumor shows a low genetic similarity with that of the cancer cells in the primary tumor, making it difficult to determine the location of the primary tumor.
  • the metastasis advantage of cancer cells is the main cause of cancer death.
  • Tumor metastasis refers to a process in which tumor cells migrate from the original site to other body parts and continue to grow by invading the circulatory system.
  • Cancer of unknown primary is a type of tumor that is diagnosed as metastatic, but has a primary location undetermined even after a comprehensive diagnostic workup.
  • the CUP accounts for about 3% to 9% of the total number of malignant tumors, and shows a mortality rate ranking fourth or fifth among all cancers.
  • the reason why a primary tumor cannot be determined may be that the primary tumor is too small, has been destroyed by the human immune system, or has been surgically removed.
  • Patients with CUP generally have a poor prognosis.
  • a median survival time after chemotherapy is 4 to 12 months, a 1-year survival rate is about 50%, and a 5-year survival rate is about 10%.
  • Approximately 15-20% of CUP patients who receive site-specific chemotherapy have improved overall survival compared to patients treated with empiric chemotherapy.
  • PET/CT Positron emission tomography/computed tomography
  • Immunohistochemistry with antibodies against tumor antigens has been the "gold standard" for the past two decades. Yet, the challenges remain: hand-picked antibody panels are primarily subjective, and the IHC analysis can identify primary sites in only 50-65% of patients with metastases and an even lower rate of 20-25% in CUP patients.
  • researchers have been able to detect the gene expression, miRNA, lncRNA, and methylation levels of tumor tissue cells, and obtain tissue-specific tumor characteristics through analysis.
  • the molecular expression profiles of tumor cells in metastases are more similar to those in the primary tumor, but different from those in the metastatic site. This suggests that tumor origin can be traced based on the molecular expression profiles of metastatic tumor cells.
  • DNA methylation the addition of a methyl group to the cytosine almost exclusively in the context of CpG dinucleotides, shows both cell- and tissue-specific patterns in the human genome, which can be used to determine the primary location of the tumor.
  • Machine learning algorithms can discover patterns from a large amount of methylation data, and classify the tumor tissue origin accordingly. As a result, the machine learning algorithms are suitable for the methylation-based tumor tracing.
  • the present disclosure designs a novel DNA methylation-based technology, including methylated adapters with inline barcodes and matched PCR amplification primers.
  • the technology has higher accuracy and effectiveness.
  • the present disclosure provides a kit for identifying tumor tissue-of-origin, including a serial methylated adapter that have nucleotide sequences of A01-T, A01-B, A02-T, A02-B, A03-T, A03-B, A04-T, A04-B, A05-T, A05-B, A06-T, and A06-B.
  • the kit for identifying tumor tissue-of-origin further includes PCR amplification primers and a PCR amplification reagent, where the PCR amplification primers include nucleotide sequences of R01-F, R01-R, R02-F, and R02-R.
  • the kit for identifying tumor tissue-of-origin further includes a PCR amplification reagent, where the PCR amplification reagent includes a polymerase, dNTP, MgCl 2 , and Tris-HCl.
  • the PCR amplification reagent includes a polymerase, dNTP, MgCl 2 , and Tris-HCl.
  • the kit for identifying tumor tissue-of-origin further includes a dephosphorylase and a 10 ⁇ buffer selected from the group consisting of KAc, Tris-Ac, and Mg(Ac) 2 .
  • the buffer works well for multiple enzymatic reactions employed in this kit.
  • the kit for identifying tumor tissue-of-origin further includes a dNTP mixture, where the dNTP mixture includes dATP, dCTP, and dGTP and excludes dTTP.
  • the kit for identifying tumor tissue-of-origin further includes an end repair enzyme, a ligase, and ATP.
  • the kit for identifying tumor tissue-of-origin further includes a negative control and a positive control, where the negative control is a healthy human blood leukocyte DNA, and the positive control is a cancer tissue sample DNA.
  • the present disclosure further provides a data analysis method for a kit for identifying tumor tissue-of-origin, including the following steps:
  • the quality control is conducted on the raw off-machine data by a FastQC software component in step 1).
  • the adapter sequence and the inline barcode sequence of the raw data are removed by Trim Galore software in step 1).
  • the beneficial effects of the present disclosure are: (1) The present disclosure is based on two methodological patents including " CN113604540A , Method for Quickly Constructing RRBS Sequencing Library Using Circulating Tumor DNA” and " CN113550013A , Method for Rapidly Constructing RRBS Sequencing Library Using Formalin-fixed Paraffin-Embedded Sample".
  • a kit is developed for qualitatively detecting a methylation profile of a human genome in formalin-fixed paraffin-embedded (FFPE) tumor tissues or blood samples of patients initially diagnosed as cancer by clinicians. The kit has higher accuracy and effectiveness.
  • the present disclosure solves the current problem that the primary tumor of some tumor patients cannot be determined by IHC.
  • the present disclosure further improves the prediction accuracy of tumor primary site through a matching analysis system.
  • the present disclosure provides a kit for identifying tumor tissue-of-origin, including an adapter, PCR amplification primers, PCR amplification reagents, a dephosphorylase, a 10 ⁇ buffer, a dNTP mixture, an end repair enzyme, a ligase, ATP, a negative control, and a positive control, where the adapter includes nucleotide sequences of A01-T, A01-B, A02-T, A02-B, A03-T, A03-B, A04-T, A04-B, A05-T, A05-B, A06-T, and A06-B.
  • the PCR amplification reagent includes a polymerase, dNTP, MgCl 2 , and Tris-HCl.
  • the 10 ⁇ buffer is selected from the group consisting of KAc, Tris-Ac, and Mg(Ac) 2 .
  • Adapter nucleotide sequences are specifically as follows: A01-T mCTmCAmCGAmCGmCTmCTmCmCGTmCTAmCAAmCmC-S-T A01-B Pho-GGTTGTAGATmCGGAAGAGmCAmCAmCGTmCTAAmC A02-T mCTAmCAmCGmCGmCTmCTmCmCGATmCTAmCTmCAmC-S-T A02-B Pho-GTGAGTAGATmCGGAAGAGmCAmCAmCGTmCTGAAmC A03-T mCTAmCAmCGmCGmCTmCTmCmCGATmCTAGGATG-S-T A03-B Pho-mCATmCmCTAGATmCGGAAGAGmCAmCAmCGTmCTGAAmC A04-T mCTAmCAmCAmCGmCTmCTTmCmCGATmCTATmCGAmC-S-T A04-B Pho-GTmCGATAGATmCGGA
  • the PCR amplification primers include nucleotide sequences of R01-F, R01-R, R02-F, and R02-R.
  • the primer nucleotide sequences of the PCR amplification primers are specifically as follows: R01-F AATGATACGGCGACCACCGCACTCTTTCCCTACACGACGCTCTTCCGATCT R01-R R02-F AATGATACGGCGACCACCGATACACTCTTTCCCTACACGACGCTCTTCCGATCT R02-R
  • the present disclosure provides a data analysis method for a kit for identifying tumor tissue-of-origin, including the following steps:
  • Example 1 a detection process of this product was as follows:
  • DNA dephosphorylation treatment A DNA dephosphorylation system included: 5 ng to 500 ng of DNA, 0.1 ⁇ L to 1.0 ⁇ L of dephosphorylase, 0.3 ⁇ L to 3.0 ⁇ L of buffer, and adding water to a total volume of 10 ⁇ L to 50 ⁇ L.
  • the DNA dephosphorylation included: treatment at 37°C for 5 min to 50 min, and treatment at 75°C for 5 min to 30 min.
  • DNA digestion A DNA digestion system included: 0.1 ⁇ L to 1.0 ⁇ L of endonuclease and 0.1 ⁇ L to 1.0 ⁇ L of buffer were added to the DNA dephosphorylation treatment system.
  • DNA digestion included: treatment at 37°C for 10 min to 100 min, and treatment at 70°C for 5 min to 30 min. 3)
  • End repair of the library An end repair system included: 0.1 ⁇ L to 1.0 ⁇ L of end repair enzyme, 0.1 ⁇ L to 1.0 ⁇ L of buffer, and 0.1 ⁇ L to 1.0 ⁇ L of dNTP mixture were added to the DNA digestion system.
  • End repair reaction included: treatment at 30°C for 10 min to 30 min, treatment at 37°C for 10 min to 30 min, treatment at 70°C for 5 min to 30 min.
  • Adapter ligation An adapter ligation system included: 0.1 ⁇ L to 1.0 ⁇ L of of ATP, 0.4 ⁇ L to 4.0 ⁇ L of ligase, 0.5 ⁇ L to 5.0 ⁇ L of buffer, and 0.3 ⁇ L to 3.0 ⁇ L of adapter were added to the library end repair system.
  • An adapter ligation procedures included: treatment at 16°C for 1 h to 4 h, treatment at 70°C for 10 min to 30 min. 5) Nucleic acid purification: The system was purified with a nucleic acid purification reagent, and 40 ⁇ L to 60 ⁇ L of a DNA solution was taken into a transformation reaction.
  • PCR amplification A PCR amplification system included: 20 ⁇ L to 30 ⁇ L of transformed DNA, 10 ⁇ L to 25 ⁇ L of PCR amplification reagent, and 1 ⁇ L to 5 ⁇ L of PCR amplification primers.
  • PCR amplification program (as shown in the table below): Temperature Time Number of cycles 98°C 45 sec 1 98°C 5 sec 6 58°C 5 sec 72°C 10 sec 98°C 5 sec 14 65°C 5 sec 72°C 10 sec 72°C 5 min 1 4°C hold 8)
  • Library purification The library was purified with a nucleic acid purification reagent, and 10 ⁇ L to 50 ⁇ L of the DNA solution was stored in a 1.5 mL EP tube. 9)
  • Library quality control Quality control of library concentration: Qubit ® dsDNA HS Assay Kit was used for quantifying a library concentration, and the library concentration should be greater than 2.5 ng/ ⁇ L.
  • the LGR and LinearSVC algorithms could also be used to analyze the Beta value-based methylation index, showing the maximum AUC.
  • the produced human genome methylation detection kit was tested, the clinical samples from different hospitals were detected, and the clinicopathologic diagnosis method was used as a comparative verification method.
  • Comparative verification method - the clinicopathologic diagnosis method is based on the current medical level and is the result of a comprehensive determination of information by professional doctors, where the information included blood and biochemical tests, urinalysis, fecal occult blood test, radiological examination of the suspected tumor primary, pathological immunohistochemistry, and medical history.
  • a testing method of this kit included the following steps: Paraffin-embedded pathological tissues: a sample was collected from the lesion tissues to confirm that there were at least 30% tumor-related lesion tissues. Paraffin-embedded pathological tissues or section samples were confirmed to carry neoplastic lesion cells. The samples that had been stored for not more than 1 year were selected, and DNA extraction was conducted with a commercial kit using not less than 8 pieces of 5 ⁇ m sections or no less than 5 pieces of 10 ⁇ m sections.
  • DNA dephosphorylation treatment A DNA dephosphorylation system included: 50 ng of DNA, 1.0 ⁇ L of dephosphorylase, 10 ⁇ L of buffer, and adding water to a total volume of 30 ⁇ L. DNA dephosphorylation included: treatment at 37°C for 50 min, and treatment at 75°C for 20 min.
  • DNA digestion A DNA digestion system included: 1.0 ⁇ L of endonuclease and 0.2 ⁇ L of buffer were added to the DNA dephosphorylation treatment system.
  • DNA digestion included: treatment at 37°C for 90 min, and treatment at 70°C for 10 min.
  • End repair of the library An end repair system included: 1.0 ⁇ L of end repair enzyme, 0.4 ⁇ L of buffer, and 0.8 ⁇ L of dNTP mixture were added to the DNA digestion system.
  • End repair reaction included: treatment at 30°C for 25 min, treatment at 37°C for 25 min, treatment at 70°C for 10 min.
  • Adapter ligation An adapter ligation system included: 0.2 ⁇ L of ATP, 0.4 ⁇ L of ligase, 0.6 ⁇ L of buffer, and 3.0 ⁇ L of adapter were added to the library end repair system.
  • An adapter ligation procedures included: treatment at 16°C for 4 h, treatment at 70°C for 15 min. 5) Nucleic acid purification: The system was purified with a nucleic acid purification reagent, and 40 ⁇ L of a DNA solution was taken into a transformation reaction. 6) Bisulfite conversion: The purified nucleic acid was subjected to sodium bisulfite conversion with a conversion kit. 7) PCR amplification: A PCR amplification system included: 20 ⁇ L of transformed DNA, 25 ⁇ L of PCR amplification reagent, and 5 ⁇ L of PCR primers.
  • PCR amplification program (as shown in the table below): Temperature Time Number of cycles 98°C 45 sec 1 98°C 5 sec 6 58°C 5 sec 72°C 10 sec 98°C 5 sec 14 65°C 5 sec 72°C 10 sec 72°C 5 min 1 4°C hold 8)
  • Library purification The library was purified with a nucleic acid purification reagent, and 35 ⁇ L of the DNA solution was stored in a 1.5 mL EP tube. 9)
  • Library quality control Quality control of library concentration: Qubit ® dsDNA HS Assay Kit was used for quality control of a library concentration, and the library concentration was greater than 2.5 ng/ ⁇ L.
  • the present disclosure further provided a data analysis method for a kit for identifying tumor tissue-of-origin, including the following steps:
  • Table 2 Consistency analysis results (compared with results ranked TOP2 in probability value) SN Cancer name Sample size TOP2 prediction samples Accuracy rate 1 Breast cancer 27 25 92.59% 2 Lung cancer group 27 26 96.30% 3 Gastric cancer 30 22 73.33% 4 Intestinal cancer 31 30 96.77% 5 Head and neck cancer 22 20 90.91% 6 Thyroid cancer 25 25 100.00% 7 Cervical cancer 16 12 75.00% 8 ovarian cancer 26 22 84.62% 9 Liver cancer 23 20 86.96% 10 Esophagus cancer 13 10 76.92% Total Total number of samples 240 212 88.33%
  • the determination result of this product had high total accordance rate and consistency with those of the clinical diagnosis results.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Analytical Chemistry (AREA)
  • Organic Chemistry (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Biophysics (AREA)
  • Biotechnology (AREA)
  • Genetics & Genomics (AREA)
  • Medical Informatics (AREA)
  • Wood Science & Technology (AREA)
  • Zoology (AREA)
  • Pathology (AREA)
  • Immunology (AREA)
  • Evolutionary Biology (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Theoretical Computer Science (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Molecular Biology (AREA)
  • Hospice & Palliative Care (AREA)
  • Oncology (AREA)
  • General Engineering & Computer Science (AREA)
  • Biochemistry (AREA)
  • Microbiology (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Evolutionary Computation (AREA)
  • Epidemiology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Bioethics (AREA)
  • Artificial Intelligence (AREA)
  • Software Systems (AREA)
  • Public Health (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

The present disclosure provides a kit for identifying tumor tissue-of-origin, including an adapter and PCR amplification primers, where the adapter includes nucleotide sequences of A01-T, A01-B, A02-T, A02-B, A03-T, A03-B, A04-T, A04-B, A05-T, A05-B, A06-T, and A06-B; and the PCR amplification primers include nucleotide sequences of RO1-F, R01-R, R02-F, and R02-R. In the present disclosure, new adapter nucleotide sequences and PCR amplification primers are designed, with higher accuracy and effectiveness. The present disclosure further provides a data analysis method for a kit for identifying tumor tissue-of-origin, including data preprocessing, alignment, methylation information statistics, quality control, and analysis. In the present disclosure, the analysis of a Beta value-based methylation index further improves a recognition ratio.

Description

    TECHNICAL FIELD
  • The present disclosure relates to the technical field of biomedicine, in particular to a kit for identifying tumor tissue-of-origin and a data analysis method for a kit for identifying tumor tissue-of-origin.
  • BACKGROUND
  • Cancer metastasis refers to a process in which tumor cells fall off from a primary lesion, enter a circulatory system, transfer to other parts of the body, and continue to grow. The original tumor is called "primary lesion", and the newly formed tumor is called "secondary lesion" or "metastatic lesion". According to the stage of tumor cells entering the circulatory system, cancer metastasis can be divided into early dissemination model and late dissemination model. A difference between the two models lies primarily in an environment in which metastatic cancer cells mutate. In the late dissemination model, cancer cells mutate and proliferate at the primary lesion, and then metastasize through the circulatory system. At this time, the metastatic tumor has a high degree of genetic similarity with the primary tumor, and can be treated with same treatment methods as the primary tumor. In the early dissemination model, cancer cells mutate after passing through the circulatory system to reach their destination. In this model, the metastatic tumor shows a low genetic similarity with that of the cancer cells in the primary tumor, making it difficult to determine the location of the primary tumor.
  • When the tumor has developed into a real cancer, namely invasive cancer, it means that the cancer cells have invaded and infiltrated deeper from the place where they occurred. For invasive cancer, the metastasis advantage of cancer cells is the main cause of cancer death.
  • According to statistics, 90% of cancer deaths are caused by metastasis. Metastatic cancer is extremely difficult to treat. The attachment ability of cancer cells in patients is reduced or even completely lost, and the migration ability of cancer cells is enhanced, leading to the transfer of cancer cells from the original site to other body parts along with the blood or lymphatic system, thus finally forming new tumors. Tumor metastasis refers to a process in which tumor cells migrate from the original site to other body parts and continue to grow by invading the circulatory system.
  • Cancer of unknown primary (CUP) is a type of tumor that is diagnosed as metastatic, but has a primary location undetermined even after a comprehensive diagnostic workup. The CUP accounts for about 3% to 9% of the total number of malignant tumors, and shows a mortality rate ranking fourth or fifth among all cancers. The reason why a primary tumor cannot be determined may be that the primary tumor is too small, has been destroyed by the human immune system, or has been surgically removed. Patients with CUP generally have a poor prognosis. A median survival time after chemotherapy is 4 to 12 months, a 1-year survival rate is about 50%, and a 5-year survival rate is about 10%. Approximately 15-20% of CUP patients who receive site-specific chemotherapy have improved overall survival compared to patients treated with empiric chemotherapy.
  • Positron emission tomography/computed tomography (PET/CT), as once the most effective medical imaging tool to identify the primary location of CUP, was capable of diagnosing about 30% of CUP primary locations. Moller et al. detected the primary location of CUP using FDG-PET-CT. However, when the primary tumor has a small size, the PET/CT plays a limited role in tracing the origin of cancer.
  • Immunohistochemistry (IHC) with antibodies against tumor antigens has been the "gold standard" for the past two decades. Yet, the challenges remain: hand-picked antibody panels are primarily subjective, and the IHC analysis can identify primary sites in only 50-65% of patients with metastases and an even lower rate of 20-25% in CUP patients. In recent years, with the development of biotechnology, researchers have been able to detect the gene expression, miRNA, lncRNA, and methylation levels of tumor tissue cells, and obtain tissue-specific tumor characteristics through analysis. The molecular expression profiles of tumor cells in metastases are more similar to those in the primary tumor, but different from those in the metastatic site. This suggests that tumor origin can be traced based on the molecular expression profiles of metastatic tumor cells. At present, studies have constructed tumor traceability models, and achieved high accuracy by analyzing mRNA, miRNA, or lncRNA levels. Compared to the single-stranded RNAs, the double-stranded nature and the absence of a reactive 2'-hydroxyl group on the pentose ring make DNA more attractive for genetic testing. DNA methylation, the addition of a methyl group to the cytosine almost exclusively in the context of CpG dinucleotides, shows both cell- and tissue-specific patterns in the human genome, which can be used to determine the primary location of the tumor. Machine learning algorithms can discover patterns from a large amount of methylation data, and classify the tumor tissue origin accordingly. As a result, the machine learning algorithms are suitable for the methylation-based tumor tracing.
  • SUMMARY
  • In order to overcome the low accuracy and effectiveness of IHC marker recognition in the prior part, the present disclosure designs a novel DNA methylation-based technology, including methylated adapters with inline barcodes and matched PCR amplification primers. The technology has higher accuracy and effectiveness.
  • To achieve the above objective, the present disclosure adopts the following technical solutions: the present disclosure provides a kit for identifying tumor tissue-of-origin, including a serial methylated adapter that have nucleotide sequences of A01-T, A01-B, A02-T, A02-B, A03-T, A03-B, A04-T, A04-B, A05-T, A05-B, A06-T, and A06-B.
  • Preferably, the kit for identifying tumor tissue-of-origin further includes PCR amplification primers and a PCR amplification reagent, where the PCR amplification primers include nucleotide sequences of R01-F, R01-R, R02-F, and R02-R.
  • Preferably, the kit for identifying tumor tissue-of-origin further includes a PCR amplification reagent, where the PCR amplification reagent includes a polymerase, dNTP, MgCl2, and Tris-HCl.
  • Preferably, the kit for identifying tumor tissue-of-origin further includes a dephosphorylase and a 10× buffer selected from the group consisting of KAc, Tris-Ac, and Mg(Ac)2. The buffer works well for multiple enzymatic reactions employed in this kit.
  • Preferably, the kit for identifying tumor tissue-of-origin further includes a dNTP mixture, where the dNTP mixture includes dATP, dCTP, and dGTP and excludes dTTP.
  • Preferably, the kit for identifying tumor tissue-of-origin further includes an end repair enzyme, a ligase, and ATP.
  • Preferably, the kit for identifying tumor tissue-of-origin further includes a negative control and a positive control, where the negative control is a healthy human blood leukocyte DNA, and the positive control is a cancer tissue sample DNA.
  • The present disclosure further provides a data analysis method for a kit for identifying tumor tissue-of-origin, including the following steps:
    1. 1) data preprocessing: conducting quality control on a raw off-machine data, and removing an adapter sequence and an inline barcode sequence of a raw data to obtain clean data;
    2. 2) alignment: allowing the clean data to be aligned to the human reference genome, and converting a resulting bam file generated by the alignment into an mHap file;
    3. 3) CpG methylation information statistics: extracting a methylation information of each CpG site;
    4. 4) quality control: removing a sample with less than 800,000 CpG sites and having a bisulfite conversion rate of less than 99%, an alignment rate of less than 50%, and a coverage of not less than 10×; and
    5. 5) analysis: analyzing a Beta value-based DNA methyaltion index, predicting all samples in sequence with a training set model, and outputting a probability value of each sample on 10 cancer types.
  • Preferably, the quality control is conducted on the raw off-machine data by a FastQC software component in step 1).
  • Preferably, the adapter sequence and the inline barcode sequence of the raw data are removed by Trim Galore software in step 1).
  • The beneficial effects of the present disclosure are: (1) The present disclosure is based on two methodological patents including " CN113604540A , Method for Quickly Constructing RRBS Sequencing Library Using Circulating Tumor DNA" and " CN113550013A , Method for Rapidly Constructing RRBS Sequencing Library Using Formalin-fixed Paraffin-Embedded Sample". In the present disclosure, a kit is developed for qualitatively detecting a methylation profile of a human genome in formalin-fixed paraffin-embedded (FFPE) tumor tissues or blood samples of patients initially diagnosed as cancer by clinicians. The kit has higher accuracy and effectiveness. The present disclosure solves the current problem that the primary tumor of some tumor patients cannot be determined by IHC. (2) The present disclosure further improves the prediction accuracy of tumor primary site through a matching analysis system.
  • BRIEF DESCRIPTION OF THE DRAWINGS
    • FIG. 1 shows a flowchart of the methylation detection of the present disclosure;
    • FIG. 2 shows a flowchart of constructing an analysis model of the present disclosure; and
    • FIG. 3 shows a flowchart of the methylation analysis of the present disclosure.
    DETAILED DESCRIPTION OF THE EMBODIMENTS
  • The following further describes the present disclosure in detail with reference to the accompanying drawings and examples.
  • In one example, the present disclosure provides a kit for identifying tumor tissue-of-origin, including an adapter, PCR amplification primers, PCR amplification reagents, a dephosphorylase, a 10× buffer, a dNTP mixture, an end repair enzyme, a ligase, ATP, a negative control, and a positive control, where the adapter includes nucleotide sequences of A01-T, A01-B, A02-T, A02-B, A03-T, A03-B, A04-T, A04-B, A05-T, A05-B, A06-T, and A06-B. The PCR amplification reagent includes a polymerase, dNTP, MgCl2, and Tris-HCl. The 10× buffer is selected from the group consisting of KAc, Tris-Ac, and Mg(Ac)2.
  • Adapter nucleotide sequences are specifically as follows:
    A01-T mCTmCAmCGAmCGmCTmCTmCmCGTmCTAmCAAmCmC-S-T
    A01-B Pho-GGTTGTAGATmCGGAAGAGmCAmCAmCGTmCTAAmC
    A02-T mCTAmCAmCGmCGmCTmCTmCmCGATmCTAmCTmCAmC-S-T
    A02-B Pho-GTGAGTAGATmCGGAAGAGmCAmCAmCGTmCTGAAmC
    A03-T mCTAmCAmCGmCGmCTmCTmCmCGATmCTAGGATG-S-T
    A03-B Pho-mCATmCmCTAGATmCGGAAGAGmCAmCAmCGTmCTGAAmC
    A04-T mCTAmCAmCAmCGmCTmCTTmCmCGATmCTATmCGAmC-S-T
    A04-B Pho-GTmCGATAGATmCGGAAGAGmCAmCAmCGTmCTGAAmC
    A05-T mCTAmCAmCGAmCGmCTmCTTmCmCGTmCTmCAAGAG-S-T
    A05-B Pho-mCTmCTTGAGATmCGGAAGAGmCAmCAmCGTmCTGAAmC
    A06-T mCTAmCAmCGAmCGmCTmCTTmCmCGATmCTmCATAmC-S-T
    A06-B Pho-GTmCATGAGATmCGGAAGAGmCAmCAmCmCTGAAmC
  • The PCR amplification primers include nucleotide sequences of R01-F, R01-R, R02-F, and R02-R. The primer nucleotide sequences of the PCR amplification primers are specifically as follows:
    R01-F AATGATACGGCGACCACCGCACTCTTTCCCTACACGACGCTCTTCCGATCT
    R01-R
    Figure imgb0001
    R02-F AATGATACGGCGACCACCGATACACTCTTTCCCTACACGACGCTCTTCCGATCT
    R02-R
    Figure imgb0002
  • In another example, the present disclosure provides a data analysis method for a kit for identifying tumor tissue-of-origin, including the following steps:
    1. 1) data preprocessing: conducting quality control on a raw off-machine data, and removing an adapter sequence and an inline barcode sequence of a raw data to obtain clean data; the quality control is conducted on the raw off-machine data by a FastQC software component in step 1); the adapter sequence and the inline barcode sequence of the raw off-machine data are removed by Trim Galore software in step 1);
    2. 2) alignment: allowing the clean data to be aligned to the human reference genome, and converting a resulting bam file generated by the alignment into an mHap file;
    3. 3) CpG methylation information statistics: extracting a methylation information of each CpG site;
    4. 4) quality control: removing a sample with less than 800,000 CpG sites and having a bisulfite conversion rate of less than 99%, an alignment rate of less than 50%, and a coverage of not less than 10×; and
    5. 5) analysis: analyzing a Beta value-based methylation index, predicting all samples in sequence with a training set model, and outputting a probability value of each sample on 10 cancer types.
    Example 1: a detection process of this product was as follows:
  • 1) DNA dephosphorylation treatment:
    A DNA dephosphorylation system included: 5 ng to 500 ng of DNA, 0.1 µL to 1.0 µL of dephosphorylase, 0.3 µL to 3.0 µL of buffer, and adding water to a total volume of 10 µL to 50 µL. The DNA dephosphorylation included: treatment at 37°C for 5 min to 50 min, and treatment at 75°C for 5 min to 30 min.
    2) DNA digestion:
    A DNA digestion system included: 0.1 µL to 1.0 µL of endonuclease and 0.1 µL to 1.0 µL of buffer were added to the DNA dephosphorylation treatment system. DNA digestion included: treatment at 37°C for 10 min to 100 min, and treatment at 70°C for 5 min to 30 min.
    3) End repair of the library:
    An end repair system included: 0.1 µL to 1.0 µL of end repair enzyme, 0.1 µL to 1.0 µL of buffer, and 0.1 µL to 1.0 µL of dNTP mixture were added to the DNA digestion system. End repair reaction included: treatment at 30°C for 10 min to 30 min, treatment at 37°C for 10 min to 30 min, treatment at 70°C for 5 min to 30 min.
    4) Adapter ligation:
    An adapter ligation system included: 0.1 µL to 1.0 µL of of ATP, 0.4 µL to 4.0 µL of ligase, 0.5 µL to 5.0 µL of buffer, and 0.3 µL to 3.0 µL of adapter were added to the library end repair system. An adapter ligation procedures included: treatment at 16°C for 1 h to 4 h, treatment at 70°C for 10 min to 30 min.
    5) Nucleic acid purification:
    The system was purified with a nucleic acid purification reagent, and 40 µL to 60 µL of a DNA solution was taken into a transformation reaction.
    6) Bisulfite conversion:
    The purified nucleic acid was subjected to sodium bisulfite conversion with a conversion kit.
    7) PCR amplification:
    A PCR amplification system included: 20 µL to 30 µL of transformed DNA, 10 µL to 25 µL of PCR amplification reagent, and 1 µL to 5 µL of PCR amplification primers.
    PCR amplification program (as shown in the table below):
    Temperature Time Number of cycles
    98°C 45 sec 1
    98°C 5 sec 6
    58°C 5 sec
    72°C 10 sec
    98°C 5 sec 14
    65°C 5 sec
    72°C 10 sec
    72°C 5 min 1
    4°C hold
    8) Library purification:
    The library was purified with a nucleic acid purification reagent, and 10 µL to 50 µL of the DNA solution was stored in a 1.5 mL EP tube.
    9) Library quality control:
    Quality control of library concentration: Qubit® dsDNA HS Assay Kit was used for quantifying a library concentration, and the library concentration should be greater than 2.5 ng/µL. Quality control of library fragment length: Agilent High Sensitivity DNA Kit was used for quality inspection of library fragments. If the size distribution of the library DNA was between 200 bp to 400 bp, the library was considered as qualified.
    10) Convention of general library into MGI library (Illumina platform skipped this step):
    A universal library was converted into an MGI library with an MGIEasy Universal Library Conversion Kit.
    11) Sequencing:
    Sequencing was conducted with the matching sequencing reagents of a gene sequencer MGISEQ-2000 or the matching sequencing reagents of an Illumina sequencer.
  • A data analysis method for the kit for identifying tumor tissue-of-origin included:
    1. 1) Data preprocessing: quality control was conducted on a raw off-machine data with a FastQC software component, and an adapter sequence and an inline barcode sequence of a raw data were removed with Trim Galore software to obtain clean data.
    2. 2) Alignment: the clean data was aligned to the human reference genome (version hg19), and a resulting bam file generated by the alignment was converted into an mHap file.
    3. 3) CpG methylation information statistics: a methylation information of each CpG site was extracted with methylation data analysis software MethylDackel.
    4. 4) Quality control: to obtain high-quality samples, a sample with less than 800,000 CpG sites and having a bisulfite conversion rate of less than 99%, an alignment rate of less than 50%, and a coverage of not less than 10× was removed.
    5. 5) Feature selection: CpG islands that met the conditions for subsequent analysis were selected. A specific screening method was as follows: the coverage of CpG islands should be greater than 100×. Moreover, potential CpG islands that could be used as biomarkers by t-test were selected from the DNA methylation data set of 498 primary tumors and 57 tumor-adjacent normal tissues. The selected CpG islands should exhibit methylation levels in one cancer type that were significantly different (FDR≤0.01) from those in other cancer types. Moreover, the methylation level should be significantly higher or lower than 57 tumor-adjacent normal tissues (FDR≤0.01).
      In this example, 4 different methods were used to characterize the methylation level of CpG islands. The 4 methods were: mean methylation (Beta value), Proportion of Discordant Reads (PDR), Cell Heterogeneity-Adjusted cLonal Methylation (CHALM), and Methylated Haplotype Load (MHL). Each method could obtain a set of CpG islands, which were used for subsequent classifier construction.
    6. 6) Classifier construction: 28 classifiers were constructed using 7 machine learning algorithms (adaboost, KNN, LGR, LinearSVC, NB, RF, and SVM) and the above 4 sets of CpG islands, and then a validation set was used to evaluate these classifiers. The main indicators of evaluation included Precision, Recall, F1 score, and Accuracy. The calculation formulas were as follows: Recall = TP / TP + FN
      Figure imgb0003
      Precision = TP / TP + FP
      Figure imgb0004
      F 1 = 2 recall precision / recall + precision
      Figure imgb0005
      Accuracy = TP + TN / TP + FP + TN + FN
      Figure imgb0006
  • The predicted AUC of the classifier constructed based on the above method (the table below)
    Methods Beta CHALM MHL PDR
    adaboost 0.88 0.56 0.88 0.52
    KNN 0.81 0.58 0.80 0.54
    LGR 0.95 0.58 0.92 0.60
    LinearSVC 0.95 0.54 0.92 0.57
    NB 0.58 0.53 0.55 0.54
    RF 0.88 0.57 0.86 0.58
    SVM 0.90 0.61 0.85 0.59
  • In this example, the LGR and LinearSVC algorithms could also be used to analyze the Beta value-based methylation index, showing the maximum AUC.
  • Example 2:
  • The produced human genome methylation detection kit was tested, the clinical samples from different hospitals were detected, and the clinicopathologic diagnosis method was used as a comparative verification method.
  • Comparative verification method - the clinicopathologic diagnosis method is based on the current medical level and is the result of a comprehensive determination of information by professional doctors, where the information included blood and biochemical tests, urinalysis, fecal occult blood test, radiological examination of the suspected tumor primary, pathological immunohistochemistry, and medical history.
  • A testing method of this kit included the following steps:
    Paraffin-embedded pathological tissues: a sample was collected from the lesion tissues to confirm that there were at least 30% tumor-related lesion tissues. Paraffin-embedded pathological tissues or section samples were confirmed to carry neoplastic lesion cells. The samples that had been stored for not more than 1 year were selected, and DNA extraction was conducted with a commercial kit using not less than 8 pieces of 5 µm sections or no less than 5 pieces of 10 µm sections.
  • An effective DNA concentration was determined by Qubit® dsDNA HS Assay Kit. If the DNA concentration was not less than 2 ng/µL, and a total amount of DNA was not less than 50 ng, it was determined to be qualified.
    1) DNA dephosphorylation treatment:
    A DNA dephosphorylation system included: 50 ng of DNA, 1.0 µL of dephosphorylase, 10 µL of buffer, and adding water to a total volume of 30 µL. DNA dephosphorylation included: treatment at 37°C for 50 min, and treatment at 75°C for 20 min.
    2) DNA digestion:
    A DNA digestion system included: 1.0 µL of endonuclease and 0.2 µL of buffer were added to the DNA dephosphorylation treatment system. DNA digestion included: treatment at 37°C for 90 min, and treatment at 70°C for 10 min.
    3) End repair of the library:
    An end repair system included: 1.0 µL of end repair enzyme, 0.4 µL of buffer, and 0.8 µL of dNTP mixture were added to the DNA digestion system. End repair reaction included: treatment at 30°C for 25 min, treatment at 37°C for 25 min, treatment at 70°C for 10 min.
    4) Adapter ligation:
    An adapter ligation system included: 0.2 µL of ATP, 0.4 µL of ligase, 0.6 µL of buffer, and 3.0 µL of adapter were added to the library end repair system. An adapter ligation procedures included: treatment at 16°C for 4 h, treatment at 70°C for 15 min.
    5) Nucleic acid purification:
    The system was purified with a nucleic acid purification reagent, and 40 µL of a DNA solution was taken into a transformation reaction.
    6) Bisulfite conversion:
    The purified nucleic acid was subjected to sodium bisulfite conversion with a conversion kit.
    7) PCR amplification:
    A PCR amplification system included: 20 µL of transformed DNA, 25 µL of PCR amplification reagent, and 5 µL of PCR primers.
    PCR amplification program (as shown in the table below):
    Temperature Time Number of cycles
    98°C 45 sec 1
    98°C 5 sec 6
    58°C 5 sec
    72°C 10 sec
    98°C 5 sec 14
    65°C 5 sec
    72°C 10 sec
    72°C 5 min 1
    4°C hold
    8) Library purification:
    The library was purified with a nucleic acid purification reagent, and 35 µL of the DNA solution was stored in a 1.5 mL EP tube.
    9) Library quality control:
    Quality control of library concentration: Qubit® dsDNA HS Assay Kit was used for quality control of a library concentration, and the library concentration was greater than 2.5 ng/µL. Quality control of library fragment length: Agilent High Sensitivity DNA Kit was used for quality inspection of library fragments. If the size of the main band of the fragment was 200 bp to 400 bp, the library fragment was qualified.
    10) Convention of general library into MGI library:
    A universal library was converted into an MGI library with an MGIEasy Universal Library Conversion Kit.
    11) Sequencing:
    Sequencing was conducted with the matching sequencing reagents of a gene sequencer MGISEQ-2000.
  • Example 3
  • The present disclosure further provided a data analysis method for a kit for identifying tumor tissue-of-origin, including the following steps:
    1. 1) Data preprocessing: quality control was conducted on a raw off-machine data with a FastQC software component, and an adapter sequence and an inline barcode sequence of a raw data were removed with Trim Galore software to obtain clean data.
    2. 2) Alignment: the clean data was aligned to the human reference genome (version hg19), and a resulting bam file generated by the alignment was converted into an mHap file.
    3. 3) CpG methylation information statistics: a methylation information of each CpG site was extracted with methylation data analysis software MethylDackel.
    4. 4) Quality control: a sample with less than 800,000 CpG sites and having a bisulfite conversion rate of less than 99%, an alignment rate of less than 50%, and a coverage of not less than 10× was removed.
    5. 5) Analysis: a Beta value-based methylation index was analyzed with LGR and LinearSVC algorithms, all samples were predicted in sequence with a training set model, and a probability value of each sample on 10 cancer types was output.
    6. 6) The consistency analysis results were shown in summary Table 1. The analysis showed the prediction results of the human genome methylation detection kit (combined with a probe-anchored polymerization sequencing method) on the primary site of cancer (according to the probability value, a first-ranked result was selected to be compared with a pathological diagnosis result), with an overall accordance rate of 78.75%. The above results had high total accordance rate and consistency with the clinical diagnosis results, thus preliminarily verifying that the kit showed high accuracy and effectiveness in detecting the primary site of cancer.
    Table 1: Consistency analysis results (compared with results ranked first in probability value)
    SN Cancer name Sample size Samples with accurate prediction Accordance rate
    1 Breast cancer 27 24 88.89%
    2 Lung cancer group 27 23 85.19%
    3 Gastric cancer 30 20 66.67%
    4 Intestinal cancer 31 30 96.77%
    5 Head and neck cancer 22 15 68.18%
    6 Thyroid cancer 25 23 92.00%
    7 Cervical cancer 16 9 56.25%
    8 ovarian cancer 26 21 80.77%
    9 Liver cancer 23 18 78.26%
    10 Esophagus cancer 13 6 46.15%
    Total Total number of samples 240 189 78.75%
  • Considering that some cancers such as cervical cancer only occur in women, and when a difference between the first and second predicted probability values is small, it is possible that a corresponding part of the predicted probability value ranked second is the real primary cancer lesion. Therefore, the top two results of the predicted probability values were analyzed, and one of the top two consistent with the clinical diagnosis result was regarded to be consistent. The results were shown in the table. The accordance rate of cancer types including gastric cancer, head and neck cancer, cervical cancer, and esophagus cancer had been significantly improved, with an overall accuracy increased to 88.33% (Table 2). Table 2: Consistency analysis results (compared with results ranked TOP2 in probability value)
    SN Cancer name Sample size TOP2 prediction samples Accuracy rate
    1 Breast cancer 27 25 92.59%
    2 Lung cancer group 27 26 96.30%
    3 Gastric cancer 30 22 73.33%
    4 Intestinal cancer 31 30 96.77%
    5 Head and neck cancer 22 20 90.91%
    6 Thyroid cancer 25 25 100.00%
    7 Cervical cancer 16 12 75.00%
    8 ovarian cancer 26 22 84.62%
    9 Liver cancer 23 20 86.96%
    10 Esophagus cancer 13 10 76.92%
    Total Total number of samples 240 212 88.33%
  • Among the 240 samples tested in this study, 89 carcinomas were known to be poorly or moderately differentiated. The results were shown in Table 3. Calculated according to the highest predicted probability value (TOP1), the accuracy could reach 75.28%. The accuracy was increased to 85.39% by including the top two predicted probability values (TOP2) in the calculation. This indicated that the ability of this product to determine the primary site of poorly differentiated cancer was better than that of immunohistochemistry reported by the current technical level. Table 3 Consistency analysis results of moderately and poorly differentiated cancer tissues
    SN Cancer name Sample size TOP1 prediction samples Accuracy rate TOP2 prediction samples Accuracy rate
    1 Breast cancer N/A N/A N/A N/A N/A
    2 Lung cancer group 14 11 78.57% 13 92.86%
    3 Gastric cancer 21 15 71.43% 16 76.19%
    4 Intestinal cancer 17 16 94.12% 16 94.12%
    5 Head and neck cancer 12 8 66.67% 11 91.67%
    6 Thyroid cancer N/A N/A N/A N/A N/A
    7 Cervical cancer 6 3 50.00% 3 50.00%
    8 ovarian cancer 9 9 100.00% 9 100.00%
    9 Liver cancer 5 4 80.00% 4 80.00%
    10 Esophagus cancer 5 1 20.00% 4 80.00%
    Total 89 67 75.28% 76 85.39%
  • In general, the determination result of this product had high total accordance rate and consistency with those of the clinical diagnosis results. This preliminarily verified that the human genome methylation detection kit had a clinical value for the detection of the primary site of cancer, and showed high accuracy and effectiveness.
  • The foregoing is merely a preferable example of the present disclosure without limitation on the scope of the present disclosure. Any equivalent structure change made by using the description and the accompanying drawings of the present disclosure, or direct or indirect application thereof in other related technical fields, shall still fall in the protection scope of the patent of the present disclosure.

Claims (10)

  1. A kit for identifying tumor tissue-of-origin, comprising an adapter, wherein the adapter comprises nucleotide sequences of A01-T, A01-B, A02-T, A02-B, A03-T, A03-B, A04-T, A04-B, A05-T, A05-B, A06-T, and A06-B.
  2. The kit for identifying tumor tissue-of-origin according to claim 1, further comprising PCR amplification primers and a PCR amplification reagent, wherein the PCR amplification primers comprise nucleotide sequences of R01-F, R01-R, R02-F, and R02-R.
  3. The kit for identifying tumor tissue-of-origin according to claim 1, further comprising a PCR amplification reagent, wherein the PCR amplification reagent comprises a polymerase, dNTP, MgCl2, and Tris-HCl.
  4. The kit for identifying tumor tissue-of-origin according to claim 1, further comprising a dephosphorylase and a 10× buffer, wherein the 10× buffer is selected from the group consisting of KAc, Tris-Ac, and Mg(Ac)2.
  5. The kit for identifying tumor tissue-of-origin according to claim 1, further comprising a dNTP mixture, wherein the dNTP mixture comprises dATP, dCTP, and dGTP.
  6. The kit for identifying tumor tissue-of-origin according to claim 1 or 2 or 3 or 4 or 5, further comprising an end repair enzyme, a ligase, and ATP.
  7. The kit for identifying tumor tissue-of-origin according to claim 1 or 2 or 3 or 4 or 5, further comprising a negative control and a positive control, wherein the negative control is a healthy human blood leukocyte DNA, and the positive control is a cancer tissue sample DNA.
  8. A data analysis method for a kit for identifying tumor tissue-of-origin, comprising the following steps:
    1) data preprocessing: conducting quality control on a raw off-machine data, and removing an adapter sequence and an inline barcode sequence of a raw data to obtain clean data;
    2) alignment: allowing the clean data aligned to the human reference genome, and converting a resulting bam file generated by the alignment into an mHap file;
    3) CpG methylation information statistics: extracting a methylation information of each CpG site;
    4) quality control: removing a sample with less than 800,000 CpG sites and having a bisulfite conversion rate of less than 99%, an alignment rate of less than 50%, and a coverage of not less than 10×; and
    5) analysis: analyzing a Beta value-based methylation index, predicting all samples in sequence with a training set model, and outputting a probability value of each sample on 10 cancer types.
  9. The data analysis method for a kit for identifying tumor tissue-of-origin according to claim 8, wherein the quality control is conducted on the raw off-machine data by a FastQC software component in step 1).
  10. The data analysis method for a kit for identifying tumor tissue-of-origin according to claim 9, wherein the adapter sequence and the inline barcode sequence of the raw data are removed by Trim Galore software in step 1).
EP23189193.8A 2023-05-04 2023-08-02 Kit for identifying tumor tissue-of-origin and data analysis method Active EP4458989B1 (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
CN202310501572.7A CN116555426B (en) 2023-05-04 2023-05-04 A kit for identifying the origin of tumor tissue and a data analysis method

Publications (2)

Publication Number Publication Date
EP4458989A1 true EP4458989A1 (en) 2024-11-06
EP4458989B1 EP4458989B1 (en) 2026-02-18

Family

ID=87502997

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23189193.8A Active EP4458989B1 (en) 2023-05-04 2023-08-02 Kit for identifying tumor tissue-of-origin and data analysis method

Country Status (3)

Country Link
US (1) US20240368699A1 (en)
EP (1) EP4458989B1 (en)
CN (2) CN116555426B (en)

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008015396A2 (en) * 2006-07-31 2008-02-07 Solexa Limited Method of library preparation avoiding the formation of adaptor dimers
EP2898100A1 (en) * 2012-09-20 2015-07-29 The Chinese University Of Hong Kong Non-invasive determination of methylome of fetus or tumor from plasma
CN113249439A (en) * 2021-05-11 2021-08-13 杭州圣庭医疗科技有限公司 Construction method of simplified DNA methylation library and transcriptome co-sequencing library
CN113550013A (en) 2021-07-23 2021-10-26 杭州圣庭医疗科技有限公司 Method for rapidly constructing RRBS sequencing library by using formalin-fixed paraffin embedded sample
CN113604540A (en) 2021-07-23 2021-11-05 杭州圣庭医疗科技有限公司 Method for rapidly constructing RRBS sequencing library by using blood circulation tumor DNA
WO2022159035A1 (en) * 2021-01-20 2022-07-28 National University Of Singapore Heatrich-bs: heat enrichment of cpg-rich regions for bisulfite sequencing
US20230126920A1 (en) * 2019-11-08 2023-04-27 Beijing Institute of Genomics, Chinese Academy of Sciences (China National Center for Bioinformation Method and device for classification of urine sediment genomic dna, and use of urine sediment genomic dna

Family Cites Families (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104250663B (en) * 2013-06-27 2017-09-15 北京大学 The high-flux sequence detection method on methylated CpG island
EP3645718B1 (en) * 2017-06-30 2025-12-17 The Regents of the University of California Methods and systems for evaluating dna methylation in cell-free dna
CN107541791A (en) * 2017-10-26 2018-01-05 中国科学院北京基因组研究所 Construction method, kit and the application in plasma DNA DNA methylation assay library
KR20220015367A (en) * 2019-05-31 2022-02-08 프리놈 홀딩스, 인크. Methods and Systems for Deep Sequencing of Methylated Nucleic Acids
CN110379465A (en) * 2019-07-19 2019-10-25 元码基因科技(北京)股份有限公司 Based on RNA target to sequencing and machine learning cancerous tissue source tracing method
CN112795620B (en) * 2019-11-13 2024-08-13 深圳华大基因股份有限公司 Double-stranded nucleic acid cyclization method, methylation sequencing library construction method and kit
CN110938674B (en) * 2019-12-05 2024-03-19 广州金域医学检验集团股份有限公司 Construction method and application of methylation sequencing DNA library
CN112708622A (en) * 2021-02-01 2021-04-27 深圳裕康医学检验实验室 Joint primer combination for library construction and kit thereof
CN112941180A (en) * 2021-02-25 2021-06-11 浙江大学医学院附属妇产科医院 Group of lung cancer DNA methylation molecular markers and application thereof in preparation of lung cancer early diagnosis kit
CN113539355B (en) * 2021-07-15 2022-11-25 云康信息科技(上海)有限公司 Tissue-specific source for predicting cfDNA (deoxyribonucleic acid), related disease probability evaluation system and application
CN114045342A (en) * 2021-12-01 2022-02-15 大连晶泰生物技术有限公司 Detection method and kit for methylation mutation of free DNA (cfDNA)
CN114214408B (en) * 2021-12-22 2023-02-17 重庆大学附属肿瘤医院 A method, probe library and kit for detecting tumor ctDNA methylation with high throughput and high sensitivity
CN115896027A (en) * 2022-04-07 2023-04-04 广州燃石医学检验所有限公司 A kind of biological composition, its preparation method and application
CN115132273B (en) * 2022-08-01 2023-07-28 广州燃石医学检验所有限公司 Method and system for evaluating tumor formation risk and tumor tissue source

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008015396A2 (en) * 2006-07-31 2008-02-07 Solexa Limited Method of library preparation avoiding the formation of adaptor dimers
EP2898100A1 (en) * 2012-09-20 2015-07-29 The Chinese University Of Hong Kong Non-invasive determination of methylome of fetus or tumor from plasma
US20230126920A1 (en) * 2019-11-08 2023-04-27 Beijing Institute of Genomics, Chinese Academy of Sciences (China National Center for Bioinformation Method and device for classification of urine sediment genomic dna, and use of urine sediment genomic dna
WO2022159035A1 (en) * 2021-01-20 2022-07-28 National University Of Singapore Heatrich-bs: heat enrichment of cpg-rich regions for bisulfite sequencing
CN113249439A (en) * 2021-05-11 2021-08-13 杭州圣庭医疗科技有限公司 Construction method of simplified DNA methylation library and transcriptome co-sequencing library
CN113550013A (en) 2021-07-23 2021-10-26 杭州圣庭医疗科技有限公司 Method for rapidly constructing RRBS sequencing library by using formalin-fixed paraffin embedded sample
CN113604540A (en) 2021-07-23 2021-11-05 杭州圣庭医疗科技有限公司 Method for rapidly constructing RRBS sequencing library by using blood circulation tumor DNA

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
GUO SHICHENG ET AL: "Identification of methylation haplotype blocks aids in deconvolution of heterogeneous tissue samples and tumor tissue-of-origin mapping from plasma DNA", vol. 49, no. 4, 6 March 2017 (2017-03-06), New York, pages 635 - 642, XP093043427, ISSN: 1061-4036, Retrieved from the Internet <URL:http://www.nature.com/articles/ng.3805> DOI: 10.1038/ng.3805 *

Also Published As

Publication number Publication date
CN116555426A (en) 2023-08-08
EP4458989B1 (en) 2026-02-18
US20240368699A1 (en) 2024-11-07
CN118737268A (en) 2024-10-01
CN116555426B (en) 2024-07-12

Similar Documents

Publication Publication Date Title
KR102930361B1 (en) Cell-free DNA for the evaluation and/or treatment of cancer
US20230366034A1 (en) Compositions and methods for diagnosing lung cancers using gene expression profiles
CN110760579B (en) Reagent for amplifying free DNA and amplification method
US20080108071A1 (en) Methods and Systems to Determine Fetal Sex and Detect Fetal Abnormalities
CN110387421A (en) DNA methylation qPCR kit and application method for lung cancer detection
CN109825586A (en) DNA methylation qPCR kit and application method for lung cancer detection
US12297504B2 (en) Chromosomal assessment to diagnose urogenital malignancy in dogs
WO2018166476A1 (en) Method for detecting mutation site in sample
CN114717311A (en) Marker, kit and device for detecting urothelial cancer
CN109504780A (en) DNA methylation qPCR kit and application method for lung cancer detection
CN113337608A (en) Combined marker for early diagnosis of liver cancer and application thereof
CN117165688A (en) Marker for urothelial cancer and application thereof
CN116804218A (en) Methylation marker for detecting benign and malignant lung nodules and application thereof
CN112951325A (en) Design method and application of probe combination for cancer detection
CN115896281B (en) Methylation biomarker, kit and application
EP4458989A1 (en) Kit for identifying tumor tissue-of-origin and data analysis method
CN116987791B (en) Application of plasma markers in identification of benign and malignant thyroid nodule
US20240136022A1 (en) Methods and compositions for detecting cancer using fragmentomics
CN118621015A (en) Methylation markers for liver cancer and their uses and products
CN114480636B (en) Application of bile bacteria as diagnosis and prognosis marker of hepatic portal bile duct cancer
CN118326045A (en) Primer pair and kit for detecting methylation of ovarian cancer genes
CN117106918A (en) Method for differential diagnosis of benign lung nodules and malignant tumors by gene methylation and kit thereof
CN108588218A (en) A kind of minimally invasive detection kit of serum miRNA combination
CN118006781B (en) Markers, primer sets, high-sensitivity and high-specificity kits and detection methods for detecting urothelial carcinoma
CN115961048B (en) A gene methylation detection primer combination, reagent and application thereof

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20230802

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: EXAMINATION IS IN PROGRESS

17Q First examination report despatched

Effective date: 20250509

GRAP Despatch of communication of intention to grant a patent

Free format text: ORIGINAL CODE: EPIDOSNIGR1

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: GRANT OF PATENT IS INTENDED

INTG Intention to grant announced

Effective date: 20251007

GRAS Grant fee paid

Free format text: ORIGINAL CODE: EPIDOSNIGR3

GRAA (expected) grant

Free format text: ORIGINAL CODE: 0009210

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE PATENT HAS BEEN GRANTED

AK Designated contracting states

Kind code of ref document: B1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

REG Reference to a national code

Ref country code: CH

Ref legal event code: F10

Free format text: ST27 STATUS EVENT CODE: U-0-0-F10-F00 (AS PROVIDED BY THE NATIONAL OFFICE)

Effective date: 20260218

Ref country code: GB

Ref legal event code: FG4D

REG Reference to a national code

Ref country code: IE

Ref legal event code: FG4D

REG Reference to a national code

Ref country code: DE

Ref legal event code: R096

Ref document number: 602023012108

Country of ref document: DE