WO2025006579A1 - Classification of single cells as tumor or normal from single cell sequences - Google Patents
Classification of single cells as tumor or normal from single cell sequences Download PDFInfo
- Publication number
- WO2025006579A1 WO2025006579A1 PCT/US2024/035582 US2024035582W WO2025006579A1 WO 2025006579 A1 WO2025006579 A1 WO 2025006579A1 US 2024035582 W US2024035582 W US 2024035582W WO 2025006579 A1 WO2025006579 A1 WO 2025006579A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- reads
- sequence
- entity
- reference positions
- tumor
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
- G16B30/10—Sequence alignment; Homology search
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/20—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/30—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
Definitions
- Tumors can include tumor and non-tumor cells. Identification of single cells as tumor or non-tumor can be a diagnostic and therapeutic tool used in treatment of diseases such as cancer.
- a computer- implemented method for identifying one or more single cells as tumor or normal in a biological sample is disclosed.
- a method can include actions of [0004]
- Other versions include corresponding systems, apparatus, and computer programs to perform the actions of methods defined by instructions encoded on computer readable storage devices. These and other versions may optionally include one or more of the following features.
- a system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions.
- the obtaining also includes classifying, by one or more 1 ATTORNEY DOCKET NO.: 35629 ⁇ 0043WO1 / IP ⁇ 2544 ⁇ PCT computers, the single cell as a tumor cell or normal cell based on an aggregation of the score determined for the respective reads of the obtained plurality of reads.
- Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
- Implementations may include one or more of the following features.
- the method also includes obtaining data indicating a plurality of reference positions where a known variant sequence exists for the entity in respective reference positions of the plurality of reference positions; obtaining, by one or more computers, a plurality of reads for the single cell from the biological sample of the entity; determining, by one or more computers and for respective reads of the obtained plurality of reads, a first score indicating whether a variant sequence in the respective reads of the biological sample of the entity matches the plurality of reference positions where the known variant sequence exists; determining, by one or more computers and for respective reads of the obtained plurality of reads, a second score corresponding to respective base calls of the respective reads that match the plurality of reference positions where the known variant sequence exists.
- FIG.5 is a block diagram of system components that can be used to implement a system for classification of single cells as tumor or normal from single cell sequences.
- DETAILED DESCRIPTION [0016] The present disclosure is directed to systems, methods, apparatuses, computer programs, or any combination thereof, for classification of single cells as tumor or normal based on single cell sequence reads.
- Tumor cells are known to exhibit genetic, epigenetic, and phenotypic heterogeneity.
- the accurate identification of a single cell as tumor or normal can be important to the identification of disease, research, and treatment selection. For example, accurate identification of an individual cell as tumor or normal can be important to the understanding of tumor heterogeneity.
- Accurate identification at the single-cell level can provide a more complete understanding of tumor heterogeneity, enabling researchers and health care providers to identify different subpopulations of cells with distinct properties, such as drug resistance or metastatic potential. Accordingly, the accurate classification of single cells can guide treatment decisions for a subject affected by tumors (e.g., malignant or benign) and increase the understanding of genetic, epigenetic, and phenotypic diversity of tumor cells within a tumor or across different tumors.
- tumors e.g., malignant or benign
- tens of thousands of reads can be analyzed for a respective single cell.
- the analysis of each respective read can include determining one or more scores for each of the respective tens of thousands of reads.
- the aggregate of the determined scores for each of the respective tens of thousands of reads can classify a single cell as tumor or normal.
- the score can be based on one or more variables that are used to determine a classification of the single cell as tumor or normal.
- the variables can include a likelihood that a respective read includes one or more variant sequences (e.g., a single nucleotide variant (SNV) also called a TN somatic variant an alteration in gene expression) and/or a base call quality score corresponding to each base call of the respective read.
- SNV single nucleotide variant
- the classification of a single cell as tumor or normal can include an aggregate of more than one scored variable determined of each of the respective tens of thousands of reads. For example, first, the respective (e.g., tens of thousands) reads for the single cell are scored using a first score to indicate a likelihood the read includes one or more variant sequences (e.g., SNV or alterations in gene expression). Second, the respective reads are scored using a second score that is based on the base call quality score corresponding to each base call of the respective read. The present disclosure then classifies the single cell as a normal cell or a tumor cell based on the aggregated first score and second score determined for the respective reads of the tens of thousands of reads.
- the classification of a single cell as a tumor cell or normal cells is a technological improvement in the field of biological classification.
- the accurate identification of a single cell as tumor or normal can be important to the identification of disease, research, and treatment selection.
- accurate identification of an individual cell as tumor or normal can be important to the understanding of tumor heterogeneity.
- Tumor cells are known to exhibit genetic, epigenetic, and phenotypic heterogeneity.
- Accurate identification at the single-cell level can provide a more complete understanding of tumor heterogeneity, enabling researchers and health care providers to identify different subpopulations of cells with distinct properties, such as drug resistance or metastatic potential.
- Prior methods to classify biological samples as tumor or normal at the granularity of a single cell have failed.
- FIG.1 is a block diagram of an example of a system 100 for classification of a single cell as tumor or normal from single cell sequence reads.
- the system 100 can include a nucleotide sequencing device 110, a memory 120, a secondary analysis unit 130, variant detection engine 140, confidence score engine 150, and a classification engine 160, an output application program interface (API) engine 190, and an output display 195.
- API application program interface
- each of these components is described as being implemented within the nucleic acid sequencing device 110.
- the present disclosure is not limited to such embodiments.
- one or more of the “units” or “engines” described in FIG.1 can be executed on a computer outside the nucleic acid sequencing device 110.
- the secondary analysis unit 130 may be implemented within the nucleic acid sequencing device 110 and the variant detection engine 140, a confidence score engine 150, a classification engine 160, an output application program interface (API) engine 190 can be implemented in one or more different computers outside of the sequencing device 110.
- the one or more different computers and the nucleic acid sequencing device 110 can be communicatively coupled using one or more wired networks, one or more wireless networks, or a combination thereof.
- the network may be one or more of a wired Ethernet, a wired optical network, a LAN, a WAN, a cellular network, the Internet, or a combination thereof.
- one or more of the computers communicatively coupled to the nucleic acid sequencing device 110 can be a remote cloud server, the present disclosure is not so limited. Instead, in other implementations, the one or more computers can connected to the sequencing device 110 via a direct connection such as a direct Ethernet connection, a USB-C connection, or the like.
- the term “engine” includes one or more software components, one or more hardware components, or any combination thereof, which can be used 8 ATTORNEY DOCKET NO.: 35629 ⁇ 0043WO1 / IP ⁇ 2544 ⁇ PCT to realize the functionality attributed to a respective engine by this specification.
- an “engine,” as described herein, uses one or more processors to execute software instructions to realize the functionality of the engine described herein.
- a processor can include a central processing unit (CPU), graphics processing unit (GPU), or the like.
- the term “unit” as used in this specification includes one or more software components, one or more hardware components, or any combination thereof, which can be used to realize the functionality attributed to a respective unit by this specification.
- a “unit,” as described herein uses one or more hardware components such as hardwired digital logic gates or hardwired digital logic blocks arranged as processing engines to perform operations that realize the functionality of the unit described herein.
- the nucleic acid sequencing device 110 (also referred to herein as sequencing device 110) is configured to perform primary nucleic acid sequence analysis.
- the sequencing device 110 is configured to perform single cell sequencing.
- the biological sample 105 sequenced by the sequencing device 110 can be comprised of a single cell.
- the single cell is isolated from tissue.
- tissue include whole blood, peripheral blood mononuclear cells (PBMCs), saliva, tumor tissue, non-tumor tissue, urine, sweat, cerebral spinal fluid, etc.
- individual cells can be isolated from a tissue sample using a variety of techniques, such as fluorescence-activated cell sorting (FACS), micromanipulation, or laser capture microdissection.
- FACS fluorescence-activated cell sorting
- isolated cells are then lysed to release their DNA or RNA, which is amplified using various methods to generate sufficient material for sequencing. Different amplification methods can be used depending on whether DNA or RNA is being sequenced.
- FACS fluorescence-activated cell sorting
- RNA DNA or RNA
- it can be prepared for sequencing using a library preparation method that adds adapter sequences to the ends of the amplified fragments. These adapters allow the fragments to be attached to a sequencing flow cell and amplified further using bridge amplification or clonal amplification methods.
- the sequencing device 110 is configured to generate ordered sequences of nucleotides, respectively referred to herein as “reads” or “sequence reads.”
- the nucleic acid sequencer 110 can be used to produce RNA reads of a biological sample 105. In such implementations, this can occur using RNA-seq protocols.
- a biological sample can be preprocessed using reverse-transcription to form complementary DNA (cDNA) using a reverse transcriptase enzyme.
- the nucleic acid sequencer 110 can include an RNA sequencer
- the biological sample 105 can include an RNA sample.
- RNA reads produced using cDNA or via an RNA sequencer can be comprised of C, G, A, and Uracil (U).
- U Uracil
- the sequencing device 110 can sequence the biological sample 105 (e.g., a single cell) and generate a corresponding set of RNA reads (e.g., tens of thousands of reads) represented using base calls corresponding to nucleotides of A, C, U, and G.
- the RNA sequence reads 112-1, 112-2, 112-n are output by the sequencing device 110 and stored in the memory device 120.
- the memory device 120 can be accessible by each of the components of FIG.1 including the secondary analysis unit 130, variant detection engine 140, confidence score engine 150, the classification engine 160, and the output API engine 190.
- the secondary analysis unit 130 can access the reads 112-1, 112-2, 112-n stored in the memory device 120 and perform one or more secondary analysis operations on the reads 112-1, 112-2, 112-n.
- the reads 112-1, 112-2, 112-n may be stored in the memory device 120 in compressed data records.
- the secondary analysis unit 130 can perform decompression operations on the compressed read records prior to performing secondary analysis operations on the read records.
- Secondary analysis operations can include mapping one or more reads to a reference sequence stored in memory device 120, aligning one or more reads to the reference sequence, or both.
- the secondary analysis unit 130 can also be configured to perform 10 ATTORNEY DOCKET NO.: 35629 ⁇ 0043WO1 / IP ⁇ 2544 ⁇ PCT sorting operations. Sorting operations can include, for example, ordering reads that have been aligned by the secondary analysis unit 130 based on the position in the reference genome to which the aligned reads were mapped.
- the functionality of the read alignment unit 136 can include obtaining data indicating a plurality of reference positions where a known variant sequence exists in respective reference positions of the plurality of reference positions.
- obtaining data indicating a plurality of reference positions can include obtaining a reference sequence.
- a reference sequence includes a sequence (e.g., nucleic acid, amino acid, peptide, or chromosome) that has known characteristics and can serve as a template for comparisons with other sequences.
- a reference sequence can be a high-quality, annotated, and well-characterized sequence that represents the consensus sequence of a particular species, organism, or biological sample.
- a reference sequence can provide a framework for the study of genetic variation, gene expression, and functional genomics.
- a reference sequence can be used as a basis for comparing and analyzing genetic variations in different populations, individuals, or tissues from individuals.
- a reference sequence is a sequence that includes one or more known variant sequences in the respective reference positions. In some example embodiments, a reference sequence is a sequence that includes one or more known non-variant reference sequences in the respective reference positions. In some example embodiments, a reference sequence is a sequence that includes one or more known variant sequences in respective reference positions and one or more known non-variant reference sequences in the respective reference positions.
- the functionality of the read alignment unit 136 can also include obtaining one or more reads such as RNA reads 112-1, 112-2, 112-n that were stored in memory 120 by the sequencing device 110, mapping the obtained reads 112-1, 112-2, 112-n to one or more reference sequence locations of a reference sequence, and then aligning the mapped reads 112-1, 112-2, 112-n to the reference sequence.
- sequence reads 112-1, 112-2, 112-n are compared to a known reference sequence using read alignment unit 136.
- the reference sequence is a sequence generated by sequencing an initial tissue sample of the same entity from which the single cell biological sample 105 was obtained.
- the initial 11 ATTORNEY DOCKET NO.: 35629 ⁇ 0043WO1 / IP ⁇ 2544 ⁇ PCT tissue sample may be a tumor that formed a portion of the entity’s body (e.g., lung, pancreas, stomach, etc.) and the single cell may be obtained from the same portion of the entity’s body on which the tumor formed.
- the single cell may be obtained from a new tumor that has formed after the tumor (which yielded the initial tissue sample) has been removed. Since tissue samples such as tumor tissue samples can comprise both tumor cells and normal cells, the reference sequence in this implementation was analyzed to identify normal sequences (e.g., known reference sequence 115) and tumor supporting sequences (known variant sequences 113).
- a non-single cell biological sample from the tissue of an entity can be sequenced to perform tumor normal (SNV calling). This process can identify variants that are present in tumor samples but not present in non-tumor samples.
- the sequencing method can be whole genome sequencing (WGS) or whole exome sequencing (WES), or any technology that generates a fingerprint of tumor specific SNVs (also called a TN somatic variant).
- the reference sequence can include a known tumor genomic library with a plurality of known variant sequences or a known tumor gene expression library with a plurality of known variant sequences.
- the secondary analysis unit 130 can access the known reference sequence 115, the known variant sequence 113, or both, stored in the memory device 120 and perform one or more secondary analysis operations on the reads the known reference sequence 115, the known variant sequence 113, or both.
- the known reference sequence 115, the known variant sequence 113, or both may be stored in the memory device 120 in compressed data records.
- the secondary analysis unit 130 can perform decompression operations on the compressed read records prior to performing secondary analysis operations on the read records.
- the known variant sequence 113 can include a combination of TN somatic variants.
- a single TN somatic variant or a combination of TN somatic variants in the known variant sequence 113 can be indicative of a particular tumor or biological sample.
- the obtained reads 112-1, 112-2, 112-n can be mapped by the 12 ATTORNEY DOCKET NO.: 35629 ⁇ 0043WO1 / IP ⁇ 2544 ⁇ PCT read alignment unit 136 to the known reference sequence such as known variant sequence 113.
- the reference sequence such as reference sequence 115 does not include the TN somatic variants.
- the read alignment unit 136 can align reads that match the known reference sequence when the reads do not contain a TN somatic mutation.
- the read alignment unit can align read 112-1 with the reference sequence 113.
- an eight base call portion 114 of the known variant sequence 113 is shown with the sequence AUCUUCGA which represents a TN somatic variant.
- the read 112-1 is aligned with the known variant sequence 113 because nucleotide portion 114 of the known variant sequence 113 matches the read 112-1.
- an eight nucleotide portion 116 of the known reference sequence 115 is shown with the sequence AUCUUCAA.
- the read 112-1 is not aligned with the known variant sequence 115 because nucleotide portion 116 of the known reference sequence 115 does not match.
- Read records describing the aligned reads can be output by the secondary analysis unit 130 and stored in the memory for later access by one or more other engines of system 100 such as the variant detection engine 140.
- a read record can be stored for each single-cell read 112 indicating whether or not the single cell read such as 112-1 includes a known variant sequence.
- the reference sequence can be autogenous.
- the single cell biological sample 105 from which the sequencing device 110 generates reads 112-1, 112-2, 112-n is a single cell that was isolated from the same biological sample from which the reference sequence was obtained.
- the single cell biological sample 105 from which the sequencing device 110 generates reads 112-1, 112-2, 112-n is a single cell that was isolated from a biological sample that was adjacent to a biological sample from which the reference sequence was obtained.
- the single cell biological sample 105 could be isolated from tissue that is adjacent to a location where a tumor was removed from the entity.
- the reference sequence could be generated from the tissue of the removed tumor.
- the variant detection engine 140 can obtain read records corresponding to a batch of aligned and sorted reads that were aligned by the read alignment unit 136 and determine if each read records corresponds to a single cell read sequence that includes a known variant sequence. In some implementations, this can be achieved by determining whether the obtained read record corresponds to a read such as 112-1 that aligns with the known variant sequence 113 or the known reference sequence 115. In this example, the variant detection engine 140 would determine that the read 112-1 includes a variant sequence (e.g., a TN somatic mutation). However, the same result can be determined in different ways.
- a variant sequence e.g., a TN somatic mutation
- the confidence score engine 150 assigns a second score to each single-cell read such as read 112-1 based on a base call quality score of one or more base calls of the read 112-1.
- the classification engine 160 is configured to determine, based on an aggregation of the first score and the second score for each of the plurality of single-cell reads, a classification of the single cell as a tumor cell or a normal cell.
- the classification engine 160 can receive as an input, multiple different parameters. These parameters, as will be discussed in more detail below, include a number of alt-supporting reads, a number of ref-supporting reads, and a base call error rate.
- the process 300 includes determining, by one or more computers and for respective reads of the obtained plurality of reads, a score corresponding to respective base calls of the respective reads that match the plurality of reference positions where the known variant sequence exists (330).
- the classification engine 160 can determine a base call error rate based on the score for each single-cell read. For example, the classification engine can determine that any single-cell read having a score that satisfies a predetermined threshold has a sufficient base call quality and those below it have insufficient base call quality. Then, the base call error rate can be determined as a ration of the single-cell reads having, e.g., insufficient base call quality over the total number of single-cell reads.
- FIG.4 is a flowchart of an example of a process 400 for performing classification of single cells as tumor or normal from single cell sequences.
- the process 400 may be performed by one or more electronic systems, for example, the system 100 of FIG.1.
- the process 400 includes obtaining data indicating a plurality of reference positions where a known variant sequence exists for an entity in respective reference positions of the plurality of reference positions (410).
- functionality of the read alignment unit 136 obtaining data indicating a plurality of reference positions can include obtaining a reference sequence.
- a reference sequence is a sequence that includes one or more known non-variant reference sequences in the respective reference positions.
- a reference sequence is a sequence that includes one or more known variant sequences in respective reference positions and one or more known non-variant reference sequences in the respective reference positions.
- obtaining data indicating a plurality of reference positions includes obtaining the reference sequence.
- the process 400 includes obtaining, by one or more computers, a plurality of reads for the single cell from a biological sample of the entity (420).
- the classification engine 160 can use the first score to provide an indication of (i) a number of single-cell reads that support a known variant sequence 113 and (ii) a number of single-cell reads that support a known reference sequence 115.
- the number of single-cell reads supporting a known variant sequence can be a sum of the number of single-cell reads that have a “1” as their score and the number of single-cell reads supporting a known reference can be a sum of the number of reads having a “0” as their score.
- the process 400 includes determining, by one or more computers and for respective reads of the obtained plurality of reads, a second score corresponding to respective base calls of the respective reads that match the plurality of reference positions where the known variant sequence exists (440).
- the classification engine 160 can determine a base call error rate based on the second score for each single-cell read.
- the classification engine can determine that any single-cell read having a second score that satisfies a predetermined threshold has a sufficient base call quality and those below it have insufficient base call quality.
- FIG.5 is a block diagram of system components that can be used to implement a system for classification of single cells as tumor or normal from single cell sequences.
- Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers.
- Computing device 550 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, and other similar computing devices. Additionally, computing device 500 or 550 can include Universal Serial Bus (USB) flash drives. The USB flash drives can store operating systems and other applications. The USB flash drives can include input/output components, such as a wireless transmitter or USB connector that can be inserted into a USB port of another 24 ATTORNEY DOCKET NO.: 35629 ⁇ 0043WO1 / IP ⁇ 2544 ⁇ PCT computing device.
- the components shown here, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the inventions described and/or claimed in this document.
- Computing device 500 includes a processor 502, memory 504, a storage device 506, a high-speed interface 508 connecting to memory 504 and high-speed expansion ports 510, and a low speed interface 512 connecting to low speed bus 514 and storage device 506.
- Each of the components 502, 504, 506, 508, 510, and 512 are interconnected using various busses, and can be mounted on a common motherboard or in other manners as appropriate.
- the processor 502 can process instructions for execution within the computing device 500, including instructions stored in the memory 504 or on the storage device 506 to display graphical information for a GUI on an external input/output device, such as display 516 coupled to high speed interface 508.
- multiple processors and/or multiple buses can be used, as appropriate, along with multiple memories and types of memory.
- multiple computing devices 500 can be connected, with each device providing portions of the necessary operations, e.g., as a server bank, a group of blade servers, or a multi-processor system.
- the memory 504 stores information within the computing device 500.
- the memory 504 is a volatile memory unit or units.
- the memory 504 is a non-volatile memory unit or units.
- the memory 504 can also be another form of computer-readable medium, such as a magnetic or optical disk.
- the storage device 506 is capable of providing mass storage for the computing device 500.
- the storage device 506 can be or contain a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations.
- a computer program product can be tangibly embodied in an information carrier.
- the computer program product can also contain instructions that, when executed, perform one or more methods, such as those described above.
- the information carrier is a computer- or machine-readable medium, such as the memory 504, the storage device 506, or memory on processor 502. [0082]
- the high speed controller 508 manages bandwidth-intensive operations for the computing device 500, while the low speed controller 512 manages lower bandwidth intensive operations.
- the high- 25 ATTORNEY DOCKET NO.: 35629 ⁇ 0043WO1 / IP ⁇ 2544 ⁇ PCT speed controller 508 is coupled to memory 504, display 516, e.g., through a graphics processor or accelerator, and to high-speed expansion ports 510, which can accept various expansion cards (not shown).
- low-speed controller 512 is coupled to storage device 506 and low-speed expansion port 514.
- the low-speed expansion port which can include various communication ports, e.g., USB, Bluetooth, Ethernet, wireless Ethernet can be coupled to one or more input/output devices, such as a keyboard, a pointing device, microphone/speaker pair, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.
- the computing device 500 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a standard server 520, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 524. In addition, it can be implemented in a personal computer such as a laptop computer 522.
- components from computing device 500 can be combined with other components in a mobile device (not shown), such as device 550.
- Each of such devices can contain one or more of computing device 500, 550, and an entire system can be made up of multiple computing devices 500, 550 communicating with each other.
- the computing device 500 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a standard server 520, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 524. In addition, it can be implemented in a personal computer such as a laptop computer 522.
- components from computing device 500 can be combined with other components in a mobile device (not shown), such as device 550.
- Computing device 550 includes a processor 552, memory 564, and an input/output device such as a display 554, a communication interface 566, and a transceiver 568, among other components.
- the device 550 can also be provided with a storage device, such as a micro-drive or other device, to provide additional storage.
- a storage device such as a micro-drive or other device, to provide additional storage.
- Each of the components 550, 552, 564, 554, 566, and 568 are interconnected using various buses, and several of the components can be mounted on a common motherboard or in other manners as appropriate.
- the processor 552 can execute instructions within the computing device 550, including instructions stored in the memory 564.
- the processor can be implemented as a chipset 26 ATTORNEY DOCKET NO.: 35629 ⁇ 0043WO1 / IP ⁇ 2544 ⁇ PCT of chips that include separate and multiple analog and digital processors. Additionally, the processor can be implemented using any of a number of architectures.
- the processor 510 can be a CISC (Complex Instruction Set Computers) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimal Instruction Set Computer) processor.
- the processor can provide, for example, for coordination of the other components of the device 550, such as control of user interfaces, applications run by device 550, and wireless communication by device 550.
- Processor 552 can communicate with a user through control interface 558 and display interface 556 coupled to a display 554.
- the display 554 can be, for example, a TFT (Thin-Film- Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology.
- the display interface 556 can comprise appropriate circuitry for driving the display 554 to present graphical and other information to a user.
- the control interface 558 can receive commands from a user and convert them for submission to the processor 552.
- an external interface 562 can be provide in communication with processor 552, so as to enable near area communication of device 550 with other devices.
- External interface 562 can provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces can also be used.
- the memory 564 stores information within the computing device 550.
- the memory 564 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units.
- Expansion memory 574 can also be provided and connected to device 550 through expansion interface 572, which can include, for example, a SIMM (Single In Line Memory Module) card interface.
- SIMM Single In Line Memory Module
- expansion memory 574 can provide extra storage space for device 550, or can also store applications or other information for device 550.
- expansion memory 574 can include instructions to carry out or supplement the processes described above, and can include secure information also.
- expansion memory 574 can be provide as a security module for device 550, and can be programmed with instructions that permit secure use of device 550.
- secure applications can be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non-hackable manner.
- the memory can include, for example, flash memory and/or NVRAM memory, as discussed below.
- a computer program product is tangibly embodied in an information carrier.
- the computer program product contains instructions that, when executed, perform one or more methods, such as those described above.
- the information carrier is a computer- or machine-readable medium, such as the memory 564, expansion memory 574, or memory on processor 552 that can be received, for example, over transceiver 568 or external interface 562.
- Device 550 can communicate wirelessly through communication interface 566, which can include digital signal processing circuitry where necessary. Communication interface 566 can provide for communications under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others.
- Such communication can occur, for example, through radio-frequency transceiver 568.
- short-range communication can occur, such as using a Bluetooth, Wi-Fi, or other such transceiver (not shown).
- GPS (Global Positioning System) receiver module 570 can provide additional navigation- and location-related wireless data to device 550, which can be used as appropriate by applications running on device 550.
- Device 550 can also communicate audibly using audio codec 560, which can receive spoken information from a user and convert it to usable digital information. Audio codec 560 can likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of device 550.
- Such sound can include sound from voice telephone calls, can include recorded sound, e.g., voice messages, music files, etc. and can also include sound generated by applications operating on device 550.
- the computing device 550 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a cellular telephone 580. It can also be implemented as part of a smartphone 582, personal digital assistant, or other similar mobile device. [0092] Various implementations of the systems and methods described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and/or combinations of such implementations.
- ASICs application specific integrated circuits
- These various implementations can include implementation in one or more computer programs that are executable and/or interpretable on a programmable system including 28 ATTORNEY DOCKET NO.: 35629 ⁇ 0043WO1 / IP ⁇ 2544 ⁇ PCT at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
- These computer programs also known as programs, software, software applications or code
- machine-readable medium As used herein, the terms “machine-readable medium” "computer- readable medium” refers to any computer program product, apparatus and/or device, e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs), used to provide machine instructions and/or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal.
- machine- readable signal refers to any signal used to provide machine instructions and/or data to a programmable processor.
- the systems and techniques described here can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball by which the user can provide input to the computer.
- a display device e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball by which the user can provide input to the computer.
- Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.
- the systems and techniques described here can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here, or any combination of such back end, middleware, or front end components.
- the components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network ("LAN”), a wide area network (“WAN”), and the Internet.
- LAN local area network
- WAN wide area network
- the Internet the global information network
- the computing system can include clients and servers.
- a client and server are generally remote from each other and typically interact through a communication network.
- the relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
- OTHER EMBODIMENTS [0097]
- a number of embodiments have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the invention.
- the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results.
- other steps can be provided, or steps can be eliminated, from the described flows, and other components can be added to, or removed from, the described systems. Accordingly, other embodiments are within the scope of the following claims.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- General Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Public Health (AREA)
- Physics & Mathematics (AREA)
- Biomedical Technology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Pathology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Epidemiology (AREA)
- Databases & Information Systems (AREA)
- Chemical & Material Sciences (AREA)
- Analytical Chemistry (AREA)
- Biophysics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Data Mining & Analysis (AREA)
- Primary Health Care (AREA)
- Biotechnology (AREA)
- Evolutionary Biology (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Genetics & Genomics (AREA)
- Molecular Biology (AREA)
- Investigating Or Analysing Biological Materials (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24743191.9A EP4732285A1 (en) | 2023-06-26 | 2024-06-26 | Classification of single cells as tumor or normal from single cell sequences |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363523286P | 2023-06-26 | 2023-06-26 | |
| US63/523,286 | 2023-06-26 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025006579A1 true WO2025006579A1 (en) | 2025-01-02 |
Family
ID=91950394
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/035582 Ceased WO2025006579A1 (en) | 2023-06-26 | 2024-06-26 | Classification of single cells as tumor or normal from single cell sequences |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4732285A1 (en) |
| WO (1) | WO2025006579A1 (en) |
-
2024
- 2024-06-26 EP EP24743191.9A patent/EP4732285A1/en active Pending
- 2024-06-26 WO PCT/US2024/035582 patent/WO2025006579A1/en not_active Ceased
Non-Patent Citations (4)
| Title |
|---|
| FASTERIUS ERIK ET AL: "Single-cell RNA-seq variant analysis for exploration of genetic heterogeneity in cancer", SCIENTIFIC REPORTS, vol. 9, no. 1, 2 July 2019 (2019-07-02), US, XP093206833, ISSN: 2045-2322, DOI: 10.1038/s41598-019-45934-1 * |
| GASPER WILLIAM ET AL: "Variant calling enhances the identification of cancer cells in single-cell RNA sequencing data", PLOS COMPUTATIONAL BIOLOGY, vol. 18, no. 10, 3 October 2022 (2022-10-03), pages e1010576, XP093096686, DOI: 10.1371/journal.pcbi.1010576 * |
| MYERS MATTHEW A ET AL: "Identifying tumor clones in sparse single-cell mutation data", BIOINFORMATICS, vol. 36, no. Supplement_1, 1 July 2020 (2020-07-01), GB, pages i186 - i193, XP093207188, ISSN: 1367-4803, DOI: 10.1093/bioinformatics/btaa449 * |
| ROZHONOVÁ HANA ET AL: "SECEDO: SNV-based subclone detection using ultra-low coverage single-cell DNA sequencing", BIOINFORMATICS, vol. 38, no. 18, 15 September 2022 (2022-09-15), GB, pages 4293 - 4300, XP093206839, ISSN: 1367-4803, DOI: 10.1093/bioinformatics/btac510 * |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4732285A1 (en) | 2026-04-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7689557B2 (en) | An integrated machine learning framework for inferring homologous recombination defects | |
| US20240321389A1 (en) | Models for Targeted Sequencing | |
| US20210358626A1 (en) | Systems and methods for cancer condition determination using autoencoders | |
| US20260120798A1 (en) | Systems and methods for classifying patients with respect to multiple cancer classes | |
| US20210104297A1 (en) | Systems and methods for determining tumor fraction in cell-free nucleic acid | |
| US11830579B2 (en) | Methods for detecting biallelic loss of function in next-generation sequencing genomic data | |
| US20200385813A1 (en) | Systems and methods for estimating cell source fractions using methylation information | |
| CA3119328C (en) | Cancer tissue source of origin prediction with multi-tier analysis of small variants in cell-free dna samples | |
| US12236346B2 (en) | Systems and methods for using a convolutional neural network to detect contamination | |
| JP2023540257A (en) | Validation of samples to classify cancer | |
| EP3753021B1 (en) | Systems and methods for correlated error event mitigation for variant calling | |
| US20240347132A1 (en) | Classification of single cells as tumor or normal from single cell sequences | |
| WO2025006579A1 (en) | Classification of single cells as tumor or normal from single cell sequences | |
| US20210295948A1 (en) | Systems and methods for estimating cell source fractions using methylation information | |
| US20240312564A1 (en) | White blood cell contamination detection | |
| KR20250158791A (en) | Optimizing sequencing panel allocation | |
| EP3588506A1 (en) | Systems and methods for genomic and genetic analysis | |
| US20220301654A1 (en) | Systems and methods for predicting and monitoring treatment response from cell-free nucleic acids | |
| CA3080170C (en) | Models for targeted sequencing | |
| HK40087494A (en) | Systems and methods for cancer condition determination using autoencoders |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24743191 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2024743191 Country of ref document: EP Effective date: 20260126 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024743191 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2024743191 Country of ref document: EP Effective date: 20260126 |
|
| ENP | Entry into the national phase |
Ref document number: 2024743191 Country of ref document: EP Effective date: 20260126 |



