WO2024254482A2 - Cell-free dna biomarker for diagnosis and prognosis of diseases with degenerative processes - Google Patents
Cell-free dna biomarker for diagnosis and prognosis of diseases with degenerative processes Download PDFInfo
- Publication number
- WO2024254482A2 WO2024254482A2 PCT/US2024/033056 US2024033056W WO2024254482A2 WO 2024254482 A2 WO2024254482 A2 WO 2024254482A2 US 2024033056 W US2024033056 W US 2024033056W WO 2024254482 A2 WO2024254482 A2 WO 2024254482A2
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- cfdna
- methylation
- als
- cell
- estimator
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6876—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
- C12Q1/6883—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6876—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
- C12Q1/6883—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
- C12Q1/6886—Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material for cancer
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2600/00—Oligonucleotides characterized by their use
- C12Q2600/112—Disease subtyping, staging or classification
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2600/00—Oligonucleotides characterized by their use
- C12Q2600/154—Methylation markers
Definitions
- cfDNA Cell-free DNA
- cfDNA is an emerging biomarker or biomarker candidate for multiple diseases, as it originates from dying tissues and can be non-invasively measured through a blood draw.
- CfDNA has been used in the detection of cancer (1–3), to identify fetal genetic abnormalities (4,5), to screen for infectious diseases (6,7), and to predict pregnancy complications (8).
- One underexplored domain for cfDNA is neurodegenerative disease.
- Biomarkers for neurodegenerative diseases are critically needed for improving patient care and evaluating the efficiency of clinical trials (9). While the application of cfDNA to neurodegeneration is nascent, previous work (10–13) has shown alterations in the cell- free DNA and RNA of patients with neurodegeneration relative to healthy controls.
- ALS Amyotrophic Lateral Sclerosis
- Our previous work focused on analyzing the cell-free DNA of ALS patients using whole-genome bisulfite sequencing (WGBS). However, about 80% of CpG sites are not variable between tissues, limiting their utility in characterizing tissue-specific cell death and its relevance to disease.
- WGBS can be prohibitively costly, especially for clinical applications.
- a biomarker as well as practical methods, to expedite diagnosis, monitor disease progression, or prioritize patients for clinical trials for ALS and other degenerative conditions.
- CfDNA cell-free DNA
- the methods described herein provide cell-free DNA (cfDNA) as a non-invasive biomarker for ALS, ALS progression, detection and monitoring of other neurodegenerative diseases, and other conditions with degenerative processes.
- CfDNA is an ideal candidate because it is enriched in ALS patients, is informative about tissue-specific cell death, and can be extracted from a standard blood draw.
- the method comprises: (a) contacting circulating extracellular DNA extracted from the sample with a set of capture probes, wherein the capture probes are less than 50,000 in number, and wherein the capture probes specifically bind within 100 base pairs (bp) of methylated and unmethylated CpG sites on cell-free DNA (cfDNA), and wherein the methylation status of CpG sites is specific to a cell type of interest.
- the method further comprises: (b) performing methylation sequencing of cfDNA captured in step (a); (c) creating a methylation status dataset from the sequencing of step (b), wherein the methylation status dataset contains the methylation status of the cell type specific CpG sites (CTSCS); and (d) detecting cell type specific degeneration when the methylation status of the CTSCS is distinct from a reference methylation status dataset, wherein the reference methylation status dataset contains the methylation status for the CTSCS of all cell types.
- the method further comprises inputting the methylation status dataset into a statistical model. Alternatively, the methylation status dataset has previously been inputted into the statistical model.
- the statistical model comprises a penalized regression model.
- the biological sample is a blood sample.
- the biological sample is extracellular DNA that has been extracted from a blood sample, or other biological specimen.
- the cfDNA is extracted from the biological sample as a separate step prior to contact with the capture probes.
- the extracellular DNA extracted from the sample is treated with bisulfite, which introduces methylation-dependent sequence changes through selective chemical conversion of non-methylated cytosine to uracil. After treatment, all non-methylated cytosine bases are converted to uracil but all methylated cytosine bases remain cytosine.
- the cell type specific degeneration is amyotrophic lateral sclerosis (ALS).
- the cell type specific degeneration is traumatic brain injury (TBI), cancer, Alzheimer’s disease, Parkinson’s disease, or a pregnancy-related phenotype. Examples of a pregnancy-related phenotype include, but are not limited to, gestational diabetes, pre-term birth, and preeclampsia.
- the methylation sequencing comprises high throughput sequencing.
- the CTSCS are specific to skeletal muscle, fibroblasts, neurons, and/or hematopoietic cells.
- the method comprises: (a) storing a set of data comprising a plurality of prospective patient records from more than 100 patients, each prospective patient record including a plurality of parameters and corresponding values for each patient included in the patient records, and a diagnostic indicator indicating whether or not the patient included in the patient records has been diagnosed with a specific condition associated with a cell type specific degeneration; (b) selecting a subset of the plurality of parameters for inputs into the machine learning system, wherein the subset comprises CTSCS, the subset further including at least one clinical parameter selected from age, gender, and smoking status; (c) randomly partitioning the set of data into training data and validation data; and (d) generating the estimator wherein the machine learning system is trained based on the training data and the subset of inputs.
- the machine learning system trained on the subset of inputs comprises a supervised machine learning model.
- the estimator is trained with an expectation maximization (EM) algorithm that estimates the proportion of the reference cell types for the subject and that estimates a contribution of unknown cell types not included in the reference methylation status dataset, whereby the machine learning system is trained to generate the estimator.
- the estimator when used with individual patient data, generates a composite algorithm value that is converted to a predictive score relative to a reference population.
- the set of data comprising a plurality of prospective patient records includes records from more than 1,000 patients.
- the computer implemented method further comprises iteratively regenerating the estimator when the estimator does not meet a predetermined receiver operating curve (ROC) statistic, by using a different subset of inputs and/or by adjusting the associated weights of the inputs until the regenerated estimator meets the predetermined ROC statistic.
- the computer implemented method further comprises generating a static configuration of the estimator when the machine learning system meets a predetermined ROC statistic.
- the computer implemented method further comprises configuring a computing device accessible by a user with the static estimator; entering values for the subset of the plurality of parameters corresponding to the patient into the computing device; and estimating, using the static estimator, the patient into a category indicative of a likelihood of having the cell-type specific degeneration or into another category indicative of a likelihood of not having the cell-type specific degeneration.
- the computer implemented method further comprises obtaining test results from a diagnostic test that confirms or denies the presence of the cell- type specific degeneration; incorporating the test results into the training data for further training of the machine learning system; and generating an improved estimator by the machine learning system.
- the estimator functions as a classifier (e.g., when predicting ALS vs healthy). In some embodiments, it functions as a regression (e.g., when estimating a score that measures disease progress).
- FIGS.1A-1D provide an overview of epigenetic cfDNA biomarker development approach.1A: Firstly, tissue informative markers (TIMs) were selected using WGBS data to capture CpG sites that were hypermethylated or hypomethylated in a tissue of interest (SEQ ID NO: 1).1B: Next, cfDNA was extracted from the blood plasma of ALS cases and controls.
- FIGS.3A-3E illustrate the capture panel design.
- the panel was designed to capture both hypomethylated TIMs, which were CpG sites that were less methylated in a tissue of interest relative to other tissues, and hypermethylated TIMs, which were designed to capture sites more methylated in a tissue of interest than other tissues (SEQ ID NO: 1).
- FIGS.4A-4D demonstrate the capture panel performance on cfDNA data.
- FIGS.5A-5D show ALS disease classification with cfDNA epigenetic features.
- FIGS.6A-6G show the features selected by the elastic net algorithm. (6A) For each tissue the TIMs were selected for, and for the type of TIM, the total absolute ⁇ value. A larger absolute ⁇ sum indicated that the feature type contributed more to model predictions.
- FIGS.7A-7C demonstrate predictive performance of cfDNA epigenetic features for ALS phenotypes. For a tenfold cross-validated model trained using cfDNA methylation proportion and coverage features the predicted versus true (7A) ALSFRS-R, (7B) FVC, and (7C) ALSFRS-R slope. Each point represents one ALS case.
- FIGS.8A-8D illustrate cohort demographic characteristics. For the UQ and UCSF cohorts, (8A) the distribution of the age of the cases and controls, (8B) the percentage of the cohorts that are female, and the percentage of the (8C) ALS cases and (8D) controls that identify as five different racial/ethnic categories.
- FIGS.9A-9B show properties of captured TIMs.
- FIGS.10A-10C show deconvolution of validation data.
- 10C For cfDNA taken from one individual before and after exercise, the proportion of cfDNA estimated to be originating from neutrophils.
- FIGS.11A-11E illustrate the on target percentage.
- FIGS.12A-12D show the cell-type decomposition estimates. The proportion of cfDNA estimated by CelFiE to originate from each tissue for each sample type in the (12A) UCSF cohort and (12B) UQ cohort.
- FIGS.13A-13D show ALS classification using CpG coverage.
- FIGS.14A-14D show ALS disease classification using CpG methylation.
- FIGS.15A-15D show ALS disease classification using only covariate information.
- FIG.16 demonstrates the relationship between read coverage and predictive performance. For UCSF cfDNA samples, the total number of reads was randomly downsampled to reduce overall on-target CpG coverage relative to the actual UCSF read coverage.
- FIGS.17A-17D show ALS disease classification without skeletal muscle TIMS.
- FIGS.18A-18B show ALS disease classification with off-target CpGs.
- ALS amyotrophic lateral sclerosis
- cfDNA-based detection applications such as cancer and pregnancy-related conditions.
- a capture panel technology that has a projected cost of under $300 per sample. The panel was designed to capture methylation sites that are informative for tissue status. Using this technology, we captured over four thousand regions and performed next-generation sequencing of cases and controls. Sequencing data was combined with supervised machine learning to predict ALS disease status using the methylation status of the reads and clinical covariates.
- control or reference sample means a sample that is representative of normal measures of the respective marker, such as would be obtained from normal, healthy control subjects, or a baseline marker to be used for comparison. Typically, a baseline will be a measurement taken from the same subject or patient.
- the sample can be an actual sample used for testing, or a reference level or range, based on known normal measurements of the corresponding marker or based on sampling a broad group of subjects and/or cell types.
- an “increase”, “decrease”, or “distinct from” are terms used to indicate an observable difference relative to a comparison or reference value, as would be understood by a person skilled in the art.
- the observable difference is a statistically significant difference.
- the term “probe” refers to an oligonucleotide, including DNA and RNA, naturally or synthetically produced, via recombinant methods or by PCR amplification, that hybridizes to at least part of another oligonucleotide of interest.
- a probe can be single- stranded or double-stranded.
- the probe is labeled with a detectable marker, such as biotin.
- active fragment refers to a substantial portion of an oligonucleotide that is capable of performing the same function of specifically hybridizing to a target polynucleotide.
- hybridizes means that the oligonucleotide forms a noncovalent interaction with the target DNA molecule under standard conditions.
- Standard hybridizing conditions are those conditions that allow an oligonucleotide probe or primer to hybridize to a target DNA molecule. Such conditions are readily determined for an oligonucleotide probe or primer and the target DNA molecule using techniques well known to those skilled in the art.
- the nucleotide sequence of a target polynucleotide is generally a sequence complementary to the oligonucleotide primer or probe.
- the hybridizing oligonucleotide may contain nonhybridizing nucleotides that do not interfere with forming the noncovalent interaction.
- the nonhybridizing nucleotides of an oligonucleotide primer or probe may be located at an end of the hybridizing oligonucleotide or within the hybridizing oligonucleotide.
- an oligonucleotide probe or primer does not have to be complementary to all the nucleotides of the target sequence as long as there is hybridization under standard hybridization conditions.
- non-human animal includes all vertebrates, e.g., mammals and non-mammals, such as non-human primates, horses, sheep, dogs, cows, pigs, chickens, and other veterinary subjects. In a typical embodiment, the subject is a human.
- “a” or “an” means at least one, unless clearly indicated otherwise.
- “Cell Free DNA Estimation via expectation-maximization” or “CelFiE” estimates the contribution of various cell types to the cfDNA of an individual via an EM optimization algorithm.
- the input to CelFiE is WGBS reference data consisting of T total cell types and WGBS cfDNA samples for N total individuals.
- Its output is the proportion of the reference cell types that make up each individual’s cfDNA, such that the proportion of all T cell types sums to one for each individual.
- an arbitrary number of cell types can be missing, which addresses potential biases arising from estimating the proportions of cell types from a restricted reference panel.
- CelFiE also estimates the methylation values for each of the cell types included in the reference, which accommodates the currently noisy and low-coverage reference data sets. These developments are facilitated by CelFiE’s EM algorithm, which is a flexible framework for parameter estimation, even when there is missing data.
- CelFiE is detailed in Caggiano, C., et al. Comprehensive cell type decomposition of circulating cell-free DNA with CelFiE. Nat Commun 12, 2717 (2021).
- the methods described herein provide a means to use cell-free DNA (cfDNA) as a non-invasive biomarker for ALS, ALS progression, detection and monitoring of other neurodegenerative diseases, and other conditions with degenerative processes.
- CfDNA is an ideal candidate because it is enriched in ALS patients, is informative about tissue-specific cell death, and can be extracted from a standard blood draw.
- the methods described herein can be used to distinguish between ALS patients and healthy controls or other neurodegenerative diseases, such as, for example, Alzheimer’s disease.
- the method can also be used to distinguish different subtypes of ALS, such as, for example, primary lateral sclerosis (PLS), which has a much different time-course and disease progression relative to other forms of ALS.
- PLS primary lateral sclerosis
- the method comprises: (a) contacting circulating extracellular DNA extracted from a biological sample with a set of capture probes, wherein the capture probes are less than 50,000 in number, and wherein the capture probes specifically bind within 100 base pairs (bp) of methylated and unmethylated CpG sites on cell-free DNA (cfDNA), and wherein the methylation status of CpG sites is specific to a cell type of interest.
- the method further comprises: (b) performing methylation sequencing of cfDNA captured in step (a); and (c) creating a methylation status dataset from the sequencing of step (b).
- the methylation status dataset contains the methylation status of the cell type specific CpG sites (CTSCS).
- CTSCS cell type specific CpG sites
- the method further comprises (d) detecting cell type specific degeneration when the methylation status of the CTSCS is distinct from a reference methylation status dataset, wherein the reference methylation status dataset contains the methylation status for the CTSCS of all cell types.
- the method further comprises inputting the methylation status dataset into a statistical model.
- the methylation status dataset has previously been inputted into the statistical model.
- the model maps the methylation status dataset to a reference methylation status dataset.
- the reference methylation status dataset contains the methylation status for the CTSCS of all cell types.
- the CTSCS are identified using a supervised machine learning model and expectation maximization (EM) algorithm that estimates the proportion of the reference cell types for the subject and that estimates a contribution of unknown cell types not included in the reference methylation status dataset.
- EM expectation maximization
- the sites selected to capture and employ as input to the machine learning model are assigned by the model a weight for each of the sites. Some sites are assigned 0 weight because the model determined the site was not important. The data generated in this was informs which sites the model is picking up as important, whereby a higher weight is indicative of a site that is more important for discriminating between ALS and controls.
- the sites found to be informative of regions of the genome can be used to identify the tissues that are important in the cfDNA of ALS patients.
- Fig.6A shows the average weight per tissue category for various tissue types. This can be used as a guide for identifying informative tissues, e.g., for discriminating between ALS and healthy patients.
- the set of capture probes is between 20 and 20,000 in number. In some embodiments, the set of capture probes is between 500 and 15,000 in number. In some embodiments, the set of capture probes is between 1,000 and 10,000 in number. In some embodiments, about 5,000 capture probes are included in the set. In some embodiments, the statistical model comprises a penalized regression model. [0054] Also provided is a method of monitoring progression of a degenerative disease in a subject.
- the method comprises performing the method of described above at a first time point on a biological sample obtained from the subject, and repeating the method at a subsequent time point, wherein an increase or decrease in the methylation status of the CTSCS is indicative of disease progression.
- a separate algorithm is generated using training data in the form of progression as the input. The lasso approach described herein is employed to learn features of progression. One can assess changes in predicted progression by running the algorithm on subsequent samples.
- the biological sample is a blood sample.
- the biological sample is extracellular DNA that has been extracted from a blood sample, or other biological specimen.
- the cfDNA is extracted from the biological sample as a separate step prior to contact with the capture probes.
- the extracellular DNA extracted from the sample is treated with bisulfite, which introduces methylation-dependent sequence changes through selective chemical conversion of non-methylated cytosine to uracil. After treatment, all non-methylated cytosine bases are converted to uracil, but all methylated cytosine bases remain cytosine. These methylation dependent C-to-T changes can subsequently be studied using conventional DNA analysis technologies.
- the cell type specific degeneration is amyotrophic lateral sclerosis (ALS).
- the cell type specific degeneration is traumatic brain injury (TBI), cancer, Alzheimer’s disease, Parkinson’s disease, or a pregnancy-related phenotype.
- a pregnancy-related phenotype examples include, but are not limited to, gestational diabetes, pre-term birth, and preeclampsia.
- the methylation sequencing comprises high throughput sequencing.
- the CTSCS are specific to skeletal muscle, fibroblasts, neurons, and/or hematopoietic cells.
- Computer Implementation & System Also described is a computer implemented method of training a machine learning system to generate an estimator for identifying a cell type specific degeneration in a blood sample obtained from a subject.
- the method comprises: (a) storing a set of data comprising a plurality of prospective patient records from more than 100 patients, each prospective patient record including a plurality of parameters and corresponding values for each patient included in the patient records, and a diagnostic indicator indicating whether or not the patient included in the patient records has been diagnosed with a specific condition associated with a cell type specific degeneration; (b) selecting a subset of the plurality of parameters for inputs into the machine learning system.
- the subset comprises CTSCS, and the subset further includes at least one clinical parameter selected from age, gender, and smoking status.
- the method further comprises (c) randomly partitioning the set of data into training data and validation data; and (d) generating the estimator.
- the machine learning system is trained based on the training data and the subset of inputs.
- the estimator is trained with a supervised machine learning model and expectation maximization (EM) algorithm that estimates the proportion of the reference cell types for the subject and that estimates a contribution of unknown cell types not included in the reference methylation status dataset, whereby the machine learning system is trained to generate the estimator.
- the estimator when used with individual patient data, generates a composite algorithm value that is converted to a predictive score relative to a reference population.
- the set of data comprising a plurality of prospective patient records includes records from more than 1,000 patients.
- the computer implemented method further comprises iteratively regenerating the estimator when the estimator does not meet a predetermined receiver operating curve (ROC) statistic, by using a different subset of inputs and/or by adjusting the associated weights of the inputs until the regenerated estimator meets the predetermined ROC statistic.
- the computer implemented method further comprises generating a static configuration of the estimator when the machine learning system meets a predetermined ROC statistic.
- the computer implemented method further comprises configuring a computing device accessible by a user with the static estimator; entering values for the subset of the plurality of parameters corresponding to the patient into the computing device; and estimating, using the static estimator, the patient into a category indicative of a likelihood of having the cell-type specific degeneration or into another category indicative of a likelihood of not having the cell-type specific degeneration.
- a system comprising a computing device programmed to perform the methods described herein. In some embodiments, the system interfaces with tubes for receiving drawn blood, machinery that extracts cell free DNA, kits for processing DNA, sequencers, and computers for prediction.
- Such interface can include, for example, communicating instructions to initiate these tasks, and/or receiving input from other elements of the system.
- the computer implemented method further comprises obtaining test results from a diagnostic test that confirms or denies the presence of the cell- type specific degeneration; incorporating the test results into the training data for further training of the machine learning system; and generating an improved estimator by the machine learning system.
- the estimator functions as a classifier (e.g., when predicting ALS vs healthy). In some embodiments, it functions as a regression (e.g., when estimating a score that measures disease progress).
- Example 1 Development of a non-invasive biomarker discovery in amyotrophic lateral sclerosis
- This Example demonstrates the development of cfDNA as a biomarker for ALS, as well as for its use in identifying the tissue of origin, as well as providing a biomarker for other degenerative conditions. Circulating cell-free DNA (cfDNA) in the bloodstream comes from dying cells and can be used to learn about tissue health. However, cfDNA sequences themselves provide very little information about the tissue of origin.
- Probe capture can pull down a select number of chosen fragments in both their methylated and unmethylated states.
- PCA principle component analysis
- a limitation of whole genome epigenetic approaches is that the cost to achieve high sequencing coverage(14) is currently too expensive to be routinely applicable in clinical settings.(2,15) High sequencing coverage, however, is needed since certain cfDNA fragments may only be present in low quantities, which could be missed by shallow sequencing.(16) Furthermore, many methylation sites are not variable,(17) limiting their value in biomarker development. [0077] To address these limitations, previous work has successfully used DNA methylation capture(18) to enrich for only relevant genomic regions, which can reduce sequencing costs while maintaining high coverage. Examples of DNA methylation capture in cfDNA applications include an approach to classify cancer types and to predict whether a patient develops preeclampsia.
- ALS cases at the time of visit FVC and ALSFRS-R were taken, and ALSFRS- R slope and FVC slope relative to the previous visit were calculated. The symptom onset site and date of first symptoms were also recorded.
- all blood samples were collected in the PAXgene Blood ccfDNA Tubes following a clinic appointment. To ensure enough cfDNA was available for downstream applications 20 mL of whole blood from controls/OND and 10 mL of whole blood from cases were collected. Following laboratory receipt (typically within 24-48hrs of collection) blood was spun with the brake off (10mins, 1900g) before plasma was aliquoted and spun twice (10mins, 16000g) to remove any further debris.
- Library Preparation and Sequencing [0087] Using a harmonized protocol across two sites (UCSF and UQ) cfDNA was extracted and prepared for sequencing. Briefly, plasma was thawed at room temperature and cfDNA was extracted from all available plasma (range 2-8 ml) using the QIAGEN Circulating Nucleic Acid kit (Cat No: 55114) according to the manufacturer’s recommendations. Extracted cfDNA was quantified using Qubit dsDNA HS Assay and visualized using the cfDNA assay (Agilent - TapeStation 4200 (UCSF) and Agilent Bioanalyzer 2100 (HS kit) (UQ)).
- cfDNA was bisulfite converted using the Zymo Lightning kit (Zymo Research) and underwent library preparation using the Accel-NGS Methyl-Seq (Swift Biosciences) according to the manufacturer’s instructions with a major modification. Briefly, the denatured BS-converted cfDNA was subject to the adaptase, extension, and ligation reaction. Following the ligation purification, the DNA underwent primer extension (98C for 1 minute; 70C for 2 minutes; 65C for 5 minutes; 4C hold) using oligos containing random UMI and i5 barcodes.
- the resulting unique-dual indexed libraries were then purified, quantified using the Qubit HS-dsDNA assay, the quality checked using the D1000-HS assay (Agilent - TapeStation 4200), and grouped as 12-plex pools. Each pool was then subject to hybridization capture using the xGen Hybridization Capture Kit (IDT) using custom probes designed on approximately 5000 pre-selected regions. [0089] For each top and bottom strands of the regions of interest, two probes were designed: one “unmethyl” probe with all G bases converted to A, and one “methyl” probe with all non-CpG G bases converted to A.
- Tissue informative marker selection [0092] Tissue informative marker selection
- TIMs were selected for 19 tissues and cell types: dendritic cells, endothelial cells, eosinophils, erythroblasts, macrophages, monocytes, neutrophils, T-cells, adipose, brain, fibroblast, heart, hepatocytes, lung, megakaryocytes, skeletal muscle, small intestine, placenta, and mammary epithelial cells.
- tissues were determined based on our previous work to be relevant to ALS, or selected based on previous publications to be the primary contributors to cfDNA. At least two WGBS samples per reference dataset were obtained. The average methylation per CpG for the reference tissue replicates was calculated.
- Probe design For each of the 4,994 TIMs, both a methylated and unmethylated probe were designed to bind to and capture both possible states of the targeted CpG. To increase the efficiency of the capture, 120 base pair probes were designed to target a window around the TIM.
- any cytosine base not protected by a methyl group in position 5 is converted into thymine.(60) Since methylation in humans primarily occurs at CpG sites, this means that all cytosines on the forward strand would be converted to thymine. Thus, to capture the unmethylated CpG state, the unmethylated probe was designed with all guanine bases converted to adenine. For the methylated state, where only cytosines in a CpG dinucleotide would be protected from the bisulfite treatment, only non- CpG guanine bases were converted to an adenine.
- cfDNA deconvolution was performed using CelFiE, which is a supervised deconvolution algorithm that is designed for noisy read count data and missing reference tissues. Input sites for CelFiE were the on-target TIMs selected for capture, As demonstrated in the CelFiE publication, summing reads from adjacent CpGs can improve deconvolution performance by decreasing sampling noise. As such, reads were summed +/-250bp around the target CpG. Sites with no reads covering the CpG were set to have a read depth of zero.
- any site that had a median read coverage of 1 read or less was also removed.
- the input matrix was made by dividing the number of methylated reads by the total number of reads. Imputation was performed per cohort over the methylation proportion matrix using SoftImpute, implemented in the Python package fancyImpute. For methylation coverage features, the coverage was normalized per sample by dividing the number of reads at a CpG by the total number of sequencing reads per individual. [0111] Sex and SIRE were one-hot encoded and added as columns in the input matrix. Age, cfDNA starting concentration, and total cfDNA input were included as continuous covariates.
- the alpha parameter which controls model sparsity was selected by performing ten- fold cross-validation on the training cohort and picking the optimal value.
- the BigStatsR package removes the manual selection of an optimal lambda value by introducing the Cross- Model Selection and Averaging (CMSA) procedure.(42,66)
- CMSA separates the training set into K folds and then performs cross-validation within the training set to obtain a set of vectors of predictions. This set of coefficients is averaged to produce the final coefficient value.
- CMSA Cross- Model Selection and Averaging
- CMSA separates the training set into K folds and then performs cross-validation within the training set to obtain a set of vectors of predictions. This set of coefficients is averaged to produce the final coefficient value.
- To standardize the weights produced per CpG site in each model we scaled input value parameters to have mean zero and variance one. We scaled the test and training data separately.
- ALS disease phenotype prediction ALS disease prediction models were trained for ALSFRS, ALSFRS Slope, and FVC.
- the top 1000 methylation features and top 1000 coverage features from the combined case- control prediction model were used as input to the model along with age, sex, SIRE, input cfDNA concentration, and total cfDNA input as non-penalized covariates. Due to low sample sizes for the case-only analysis, we meta-analyzed the two cohorts and additionally added cohort as a non-penalized covariate. We trained the elastic net model using the BigStatsR package with the big_spLinReg command. Each of the three models were evaluated against an elastic net model trained on only the covariates. [0121] Off target prediction models [0122] Off target prediction models incorporated information for all CpGs obtained from high throughput sequencing.
- cfDNA WGBS data from diverse disease contexts was used to screen candidate TIMs for those actually observed in cfDNA.
- cfDNA from our two cohorts was extracted (Fig.1B) and underwent methylation profiling on the TIM-enriched cfDNA (Fig.1C).
- Fig.1D we analyzed the methylation status of the targeted regions and developed statistical and machine learning approaches to learn about the disease status of the ALS patients and controls.
- the UCSF cohort comprised 42 ALS cases, 9 PLS cases, and 45 healthy age- matched controls consisting of unrelated partners or carers.
- the UQ OND samples included a cross-section of neurological conditions, including diseases that share pathophysiology with ALS, like frontotemporal degeneration,(23) and other neurodegenerative diseases like Alzheimer’s disease (Table 2). Therefore, the UQ cohort represented a challenging real-world scenario for ALS biomarker development.
- Table 2 Other neurological disease patients.
- Table 2 Other neurological disease patients.
- ALSFRS-R ALS Functional Rating Scale-Revised
- ALSFRS-R slope The change in ALSFRS-R between visits, referred to as ALSFRS-R slope, was also calculated as a metric of disease progression.
- Fig 2B The change in ALSFRS-R between visits, referred to as ALSFRS-R slope, was also calculated as a metric of disease progression.
- FVC forced vital capacity
- Fig.2C The two cohorts were also similar in the distribution of days between cfDNA collection and symptom onset (Fig.2D).
- TIMs tissue informative markers
- a TIM is a site that is either hyper- or hypo-methylated relative to the average methylation proportion of all other tissues at that site (Fig.3A).
- WGBS methylomes that were obtained from two public reference consortiums, ENCODE(25) and Blueprint.(26)
- CpG sites were obtained from two public reference consortiums, ENCODE(25) and Blueprint.(26)
- CpG sites we focused on CpG sites as candidate TIMs, as most non-CpG sites are not methylated in adult tissues.27
- These tissues included several hematopoietic cell types, organs, epithelium, and brain (Table 3).
- TIM selection design Per tissue selected for capture, the number of hypermethylated TIMs selected, the number of hypomethylated TIMs selected, and the total number of final TIMs selected for capture.
- An important property of cfDNA is that their fragmentation patterns are non- random.(30–32) cfDNA observed in blood generally are fragments approximately 160 base pairs long,(33) suggesting that cfDNA fragments are protected from degradation in the blood by the presence of tightly associated histone proteins.
- Cell-type decomposition [0151] Since TIMs were designed to be specific to a given tissue type, they can be used to estimate what tissues are contributing to the cfDNA in the context of neurodegeneration. To do this, we performed cfDNA cell-type decomposition with CelFiE.(10) CelFiE is a supervised decomposition algorithm that is designed to work with methylation read count data and missing or noisy reference data.
- CelFiE takes the TIM read count data for each cfDNA sample and estimates the proportion of the cfDNA mixture originating from the tissues in the reference dataset, along with a specified number of unknown tissues.
- Fig.12A-12B we ran CelFiE with two unknown components using the methylation proportion of the captured sites as input.
- Fig.5D we observed elevated skeletal muscle in ALS patients in both cohorts relative to the healthy control samples (t-test p-value UCSF: 1.1 ⁇ 10 -3 , UQ: 4.7 ⁇ 10 -2 ) (Fig.5D). This is consistent with muscle atrophy that occurs as part of their disease.
- Model parameters including the elastic net mixing parameter, were selected by using a cross-model selection and averaging procedure within the training set.
- Non-penalized covariates included age at the time of cfDNA sampling, sex, SIRE, cfDNA concentration, and total cfDNA input.
- AUC receiver operating characteristic curve
- TIMs can provide two classes of features for the prediction model, the methylation proportion and the coverage of the TIMs. Coverage was included because cfDNA fragmentation is non-random; we therefore reasoned that CpG coverage may also be informative of disease status. In total, we trained models using CpG coverage only, CpG methylation proportion only, and a combination of both as input features. [0158] Overall, we found that tissue informative epigenetic features could significantly predict ALS case-control status in both cohorts (Fig.5, Fig.13-14, Table 4). The best- performing model incorporated both TIM coverage and methylation features (Fig.5).
- An advantage of using a regularized regression model like an elastic net is that the model performs feature selection and assigns a higher weight, or absolute ⁇ value, to features that contribute more to accurately predicting the outcome. Features that do not contribute to the prediction will have an absolute ⁇ value near zero.
- the absolute ⁇ value for each TIM from an elastic net model trained on the entire UQ and UCSF cohorts (Fig.5C and Fig.5D). Then we examined how these values related to different characteristics of the TIMs. [0167]
- skeletal muscle TIMs were highly important in making model predictions, especially for TIMs that were hypermethylated in skeletal muscle (Fig. 6A).
- TIMs for every tissue type contributed to the model predictions (Fig.6).
- T-cell TIMs were highly important (Fig.6A), indicating that cfDNA originating from immune cell types may be relevant in ALS disease.
- Hypermethylated TIMs generally had higher absolute ⁇ values than hypomethylated TIMs (Fig.6B-6E), which could be related to our previous observation that hypermethylated TIMs were more likely to be in promoter or genic regions (Fig.3).
- Fig.6B-6E hypomethylated TIMs
- Fig.3 hypermethylated TIMs were more likely to be in promoter or genic regions
- Fig.6b-c absolute ⁇ values
- TIMs with a non-zero absolute ⁇ value were chosen for association with ALS case-control status, along with covariates and correcting for cohort. Multiple test correction was employed using false discovery rate at 10%.
- One of the most important methylation proportion features was a hypermethylated TIM selected for epithelium.
- TIM was selected for hepatocytes, it is in the XRCC6 gene, which was highly expressed in many tissues in bulk RNA-seq from the Genotype-Tissue Expression (GTEx) Project.(45)
- GTEx Genotype-Tissue Expression
- the UCSF model outperformed the UQ model. Furthermore, the transferability was better when the UQ model was applied to the UCSF cohort. While it is likely a combination of factors, one explanation may be attributed to differences in sequencing depth.
- the UCSF cohort had higher on-target CpG coverage. Additional coverage may reduce noise, especially in analyses utilizing methylation proportion. In some cases, the overall coverage is limited by the total amount of cfDNA available as input to the sequencing assay. This could be improved by recent high-throughput extraction technologies with the ability to increase cfDNA yield from a plasma sample.(53,54) [0180] Model performance also may be affected by the slight differences in ALS patient characteristics between the cohorts.
- the UCSF cohort had patients with lower ALSFRS-R scores and whose advanced condition may be easier to detect in cfDNA.
- ALS is also an extremely heterogeneous disease,(55) which can make designing biomarkers that generalize across patient populations difficult. It is also important to note that both cohorts were of majority European ancestry. Further exploration of how epigenetic cfDNA profiles differ between diverse subtypes of patients or change longitudinally as patients progress is now needed. [0181] This Example only examined the performance of tissue informative markers in characterizing ALS.
- methylation capture arrays allow for a more cost-effective and focused analysis over relevant CpG sites, targeted capture also limits the coverage of the genome. This has the potential to miss important methylation changes occurring outside the targeted regions. Additionally, since we relied on published tissue methylation data sets that are low coverage and inherently noisy, TIM selection might be affected. Marker selection and overall algorithm performance might be improved by better, high-coverage reference data. Reference panel design for cfDNA applications is a robust area of current research, and incorporating new samples or biobanks into ALS disease prediction could be an area for future research.
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Organic Chemistry (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Engineering & Computer Science (AREA)
- Analytical Chemistry (AREA)
- Zoology (AREA)
- Genetics & Genomics (AREA)
- Wood Science & Technology (AREA)
- Immunology (AREA)
- Pathology (AREA)
- Physics & Mathematics (AREA)
- Biotechnology (AREA)
- Microbiology (AREA)
- Molecular Biology (AREA)
- Biophysics (AREA)
- Biochemistry (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Hospice & Palliative Care (AREA)
- Oncology (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Investigating Or Analysing Biological Materials (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24820151.9A EP4724607A2 (en) | 2023-06-08 | 2024-06-07 | Cell-free dna biomarker for diagnosis and prognosis of diseases with degenerative processes |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363506899P | 2023-06-08 | 2023-06-08 | |
| US63/506,899 | 2023-06-08 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| WO2024254482A2 true WO2024254482A2 (en) | 2024-12-12 |
| WO2024254482A3 WO2024254482A3 (en) | 2025-03-13 |
Family
ID=93794531
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/033056 Ceased WO2024254482A2 (en) | 2023-06-08 | 2024-06-07 | Cell-free dna biomarker for diagnosis and prognosis of diseases with degenerative processes |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4724607A2 (en) |
| WO (1) | WO2024254482A2 (en) |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113151458A (en) * | 2016-07-06 | 2021-07-23 | 优美佳肿瘤技术有限公司 | Solid tumor methylation marker and application thereof |
-
2024
- 2024-06-07 WO PCT/US2024/033056 patent/WO2024254482A2/en not_active Ceased
- 2024-06-07 EP EP24820151.9A patent/EP4724607A2/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024254482A3 (en) | 2025-03-13 |
| EP4724607A2 (en) | 2026-04-15 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP2023521308A (en) | Cancer classification with synthetic training samples | |
| US12258634B2 (en) | Fragmentation for measuring methylation and disease | |
| WO2019025004A1 (en) | A method for non-invasive prenatal detection of fetal sex chromosomal abnormalities and fetal sex determination for singleton and twin pregnancies | |
| KR20250041134A (en) | Epigenetic analysis of cell-free DNA | |
| Caggiano et al. | Tissue informative cell-free DNA methylation sites in amyotrophic lateral sclerosis | |
| WO2024007971A1 (en) | Analysis of microbial fragments in plasma | |
| US20240170099A1 (en) | Methylation-based age prediction as feature for cancer classification | |
| US20240412821A1 (en) | Methylation-based biological sex prediction | |
| AU2024210217A1 (en) | Methods and systems for detecting and assessing liver conditions | |
| WO2024254482A2 (en) | Cell-free dna biomarker for diagnosis and prognosis of diseases with degenerative processes | |
| Caggiano et al. | Epigenetic profiles of tissue informative CpGs inform ALS disease status and progression | |
| US20240182982A1 (en) | Fragmentomics in urine and plasma | |
| US20230272477A1 (en) | Sample contamination detection of contaminated fragments for cancer classification | |
| US20250171858A1 (en) | Enrichment of clinically-relevant nucleic acids | |
| JP7138074B2 (en) | Method for determining the risk of hepatitis B and/or hepatitis C | |
| CA3270171A1 (en) | Fragmentomics in urine and plasma | |
| HK40129203A (en) | Fragmentomics in urine and plasma | |
| WO2026015665A1 (en) | Determining methylation status of biological samples | |
| JP2026509734A (en) | Sample barcodes in multiplex sample sequencing | |
| KR20260011831A (en) | A method for diagnosing anxiety disorders by detection of methylation markers |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024820151 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2024820151 Country of ref document: EP Effective date: 20260108 |
|
| ENP | Entry into the national phase |
Ref document number: 2024820151 Country of ref document: EP Effective date: 20260108 |
|
| ENP | Entry into the national phase |
Ref document number: 2024820151 Country of ref document: EP Effective date: 20260108 |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24820151 Country of ref document: EP Kind code of ref document: A2 |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024820151 Country of ref document: EP |



