WO2026012003A1 - 图像处理方法、癌症的计算机辅助诊断方法 - Google Patents
图像处理方法、癌症的计算机辅助诊断方法Info
- Publication number
- WO2026012003A1 WO2026012003A1 PCT/CN2025/098321 CN2025098321W WO2026012003A1 WO 2026012003 A1 WO2026012003 A1 WO 2026012003A1 CN 2025098321 W CN2025098321 W CN 2025098321W WO 2026012003 A1 WO2026012003 A1 WO 2026012003A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- lesion
- feature information
- information
- image
- image processing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/0002—Inspection of images, e.g. flaw detection
- G06T7/0012—Biomedical image inspection
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
- G06N3/0455—Auto-encoder networks; Encoder-decoder networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/084—Backpropagation, e.g. using gradient descent
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
- G06V10/765—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects using rules for classification or partitioning the feature space
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
- G06V10/806—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/20—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10072—Tomographic images
- G06T2207/10081—Computed x-ray tomography [CT]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20081—Training; Learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20084—Artificial neural networks [ANN]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30004—Biomedical image processing
- G06T2207/30092—Stomach; Gastric
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30004—Biomedical image processing
- G06T2207/30096—Tumor; Lesion
Definitions
- This disclosure relates to the field of computer technology, and in particular to an image processing method.
- contrast-enhanced CT offers high accuracy, it exposes patients to contrast agents, which may cause allergic reactions or organ failure.
- limitations in technology and equipment prevent 24-hour availability in all regions, impacting its convenience.
- Plain CT due to its lower resolution, makes lesions visually difficult to identify, resulting in a slightly lower detection rate compared to contrast-enhanced CT. Marking lesions on plain CT also presents challenges. Therefore, it is crucial to explore how to utilize plain CT for tumor screening and improve its accuracy.
- this disclosure provides an image processing method.
- One or more embodiments of this specification also relate to a CT image processing method, a computer-aided diagnosis method for cancer, a computer-aided diagnosis method for liver cancer, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
- an image processing method comprising:
- the image processing task is received, wherein the image processing task carries multiple target images corresponding to the target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area;
- the multiple target images are input into an image processing model to obtain the detection results corresponding to the target detection regions.
- the image processing model generates object feature information and lesion feature information based on each target image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates detection results based on the object classification information, the lesion classification information, and the lesion segmentation information.
- a CT image processing method comprising:
- the CT image processing task carries multiple plain CT images corresponding to the target detection area, and the CT image processing task is used to detect whether there is an abnormal object in the target detection area;
- the multiple plain CT images are input into an image processing model to obtain the detection results corresponding to the target detection area.
- the image processing model generates object feature information and lesion feature information based on each plain CT image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates detection results based on the object classification information, the lesion classification information, and the lesion segmentation information.
- a computer-aided diagnostic method for cancer comprising:
- the CT image processing task carries multiple plain CT images corresponding to the target detection area, and the CT image processing task is used to detect whether there is a tumor in the target detection area;
- the multiple plain CT images are input into an image processing model to obtain the detection results corresponding to the target detection area.
- the image processing model generates object feature information and lesion feature information based on each plain CT image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates detection results based on the object classification information, the lesion classification information, and the lesion segmentation information.
- a computer-aided diagnostic method for liver cancer comprising:
- the CT image processing task carries multiple plain CT images corresponding to the liver region, and the CT image processing task is used to detect whether there is a tumor in the liver region;
- the multiple plain CT images are input into an image processing model to obtain the detection results corresponding to the liver region.
- the image processing model generates object feature information and lesion feature information based on each plain CT image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates detection results based on the object classification information, the lesion classification information, and the lesion segmentation information.
- a computer-aided diagnostic system for cancer comprising a client and a server, wherein:
- the client is used to send a CT image processing task to the server, wherein the CT image processing task carries multiple plain CT images corresponding to the target detection area, and the CT image processing task is used to detect whether there is a tumor in the target detection area;
- the server is used to input the multiple plain CT images into an image processing model to obtain the detection result corresponding to the target detection area.
- the image processing model generates object feature information and lesion feature information based on each plain CT image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates the detection result based on the object classification information, the lesion classification information, and the lesion segmentation information.
- a computing device comprising:
- the memory is used to store computer programs/instructions
- the processor is used to execute the computer programs/instructions, which, when executed by the processor, implement the steps of the above method.
- a computer-readable storage medium stores a computer program/instructions that, when executed by a processor, implement the steps of the method described above.
- a computer program product including a computer program/instructions that, when executed by a processor, implement the steps of the above-described method.
- One embodiment of this specification provides an image processing method, including receiving an image processing task, wherein the image processing task carries multiple target images corresponding to a target detection region, and the image processing task is used to detect whether there is an abnormal object in the target detection region; inputting the multiple target images into an image processing model to obtain a detection result corresponding to the target detection region, wherein the image processing model generates object feature information and lesion feature information based on each target image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates a detection result based on the object classification information, the lesion classification information, and the lesion segmentation information.
- the image processing method provided in this disclosure inputs multiple target images into an image processing model.
- object feature information and lesion feature information are generated based on each target image.
- the object feature information and lesion feature information are fused to generate object fusion feature information. This allows the model to use not only object feature information but also lesion feature information when predicting the final object classification information, thus fusing global and local information and improving the detection capability of small lesions.
- FIG. 1 is a flowchart of an image processing method provided in one embodiment of this specification
- Figure 2 is a schematic diagram of the model structure of an image processing model provided in one embodiment of this specification
- Figure 3 is a schematic diagram of data processing of an object branch network provided in an embodiment of this specification.
- Figure 4 is a schematic diagram of data processing of a lesion branch network provided in one embodiment of this specification.
- Figure 5 is a schematic diagram of data processing of a fusion network provided in one embodiment of this specification.
- Figure 6 is a flowchart of an image processing model training method provided in one embodiment of this specification.
- FIG. 7 is a flowchart of a CT image processing method provided in one embodiment of this specification.
- Figure 8 is a flowchart illustrating a computer-aided diagnosis method for cancer according to an embodiment of this specification
- Figure 9 is a flowchart illustrating a computer-aided diagnostic method for liver cancer according to an embodiment of this specification.
- Figure 10 is an architecture diagram of a computer-aided diagnosis system for cancer provided in one embodiment of this specification.
- FIG 11 is a schematic diagram of an image processing apparatus provided in one embodiment of this specification.
- Figure 12 is a structural block diagram of a computing device provided in one embodiment of this specification.
- first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word “if” as used herein may be interpreted as "when,” “when,” or "in response to a determination.”
- the user information including but not limited to user device information, user personal information, etc.
- data including but not limited to data used for analysis, data stored, data displayed, etc.
- the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
- Computed tomography is a scanning technology that uses a precisely collimated X-ray beam and a highly sensitive detector to scan a specific part of the human body one section after another. It features fast scanning time and clear images and can be used to examine a variety of diseases.
- Plain CT scan also known as a regular scan, refers to a scan performed intravenously without the administration of contrast agents.
- Enhanced CT This refers to a scanning method that involves injecting a contrast agent into a blood vessel before scanning. The purpose is to increase the density difference between the diseased tissue and normal tissue, so as to show lesions that are not shown or are not clearly shown on plain CT. The presence or absence of enhancement and the type of enhancement help to characterize the lesion.
- Computer-aided diagnosis refers to the use of imaging, medical image processing technology, and other possible physiological and biochemical methods, combined with computer analysis and calculation, to assist in the detection of lesions and improve the accuracy of diagnosis.
- contrast-enhanced CT offers high accuracy, it exposes patients to contrast agents, which can potentially cause allergic reactions or organ failure.
- limitations in technology and equipment restrict the availability of contrast-enhanced CT scans to 24/7 operation in all regions, impacting their convenience.
- Some screening techniques are based on ultrasound, but their sensitivity is low, leading to missed diagnoses.
- Other techniques are based on plain CT scans, but these have lower resolution, making lesions on organs difficult to visually identify, resulting in a slightly lower detection rate compared to contrast-enhanced CT. Therefore, improving tumor detection remains a pressing issue for researchers.
- One or more embodiments of this specification also relate to a CT image processing method, a computer-aided diagnosis method for cancer, a computer-aided diagnosis method for liver cancer, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
- Figure 1 shows a flowchart of an image processing method provided according to an embodiment of this specification, which specifically includes the following steps.
- Step 102 Receive an image processing task, wherein the image processing task carries multiple target images corresponding to the target detection area, and the image processing task is used to detect whether there are abnormal objects in the target detection area.
- image processing tasks sent by users can be received through either a server or a client.
- an image processing task can be understood as a task to detect whether there are abnormal objects within a target detection region.
- the image processing task carries multiple target images corresponding to the target detection region.
- the target detection region can be understood as a partition used to detect the presence of abnormal objects, and the abnormal object can be understood as a foreign object within the target detection region.
- an abnormal object could be a tumor within the human body.
- the target detection area can be any organ within the human body, such as the liver, lungs, stomach, or esophagus.
- the state of the object to be detected can be further determined based on the prediction results, thereby aiding in the localization and precise treatment of abnormal objects.
- the object to be detected can be understood as the object described in the target detection area.
- the target image is a CT image of Zhang San's stomach
- the CT image is the target image
- the object to be detected is Zhang San.
- This image processing task is used to detect whether there is a tumor in the stomach area.
- the object to be detected can be a person or other living organisms, and this is not limited in one or more specific embodiments provided in this specification.
- the image processing task can be applied to the recognition of various types of medical images, and to determine whether there are abnormal objects in the target detection area of the medical image based on image features.
- the presence of a tumor in the stomach region can be predicted based on a medical image of the stomach region, thereby helping doctors to accurately locate abnormal areas
- the presence of a tumor in the liver can be predicted based on a medical image of the liver region, thereby helping doctors to accurately locate abnormal areas, facilitating subsequent treatment.
- the target image acquired is an image of the liver.
- multiple target images acquired are plain CT images of the liver, which can be combined to form a 3D image of the liver region.
- Anomalies can be understood as tumors within the liver.
- the abnormal object can be a certain type of cell, a certain type of tissue structure, etc., such as a malignant tumor, a benign tumor, proliferating tissue, etc. This is not limited in the one or more embodiments provided in this specification.
- multiple target images corresponding to the target detection area carried in the image processing task can be used as input to detect whether there are abnormal objects within the target detection area.
- the method before receiving the image processing task, the method further includes:
- Image segmentation task carries multiple initial images corresponding to a target detection region, and the image segmentation task is used to extract the target image corresponding to the target detection region;
- Each initial image is input into a pre-trained image segmentation model to obtain the target image corresponding to each initial image output by the image segmentation model.
- the target image can be understood as a close-up image corresponding to the target detection region.
- multiple received images may contain other regions besides the target detection region, and these other regions can affect the target detection region. Therefore, the method provided in this disclosure also performs image segmentation processing.
- the first step is to obtain an image segmentation task from multiple initial images. These initial images include both target detection regions and regions that influence the target detection regions. This image segmentation task is used to extract the target images corresponding to the target detection regions from each initial image.
- the image segmentation model is trained to identify target detection regions in the initial images, extract the target detection regions from the initial images, and generate target images corresponding to the target detection regions.
- initial plain CT images are those that meet the image quality requirements.
- the sources of these initial plain CT images can be multiple CT scanners or the same CT scanner; this manual does not impose any limitations on this.
- the format of each initial plain CT image must be standardized.
- the image segmentation model can be 3DUNet.
- 3DUNet is a deep learning architecture for image segmentation in three-dimensional space. It is an extended version of the U-Net model, which was originally designed for semantic segmentation of two-dimensional biomedical images and has gained widespread recognition for its excellent performance and high accuracy in segmenting small objects. 3DUNet applies this idea to three-dimensional datasets, such as medical images (e.g., CT, MRI scans), which is very useful in many medical fields because the data in these fields often have rich three-dimensional structural information.
- medical images e.g., CT, MRI scans
- 3DUNet includes an encoder-decoder structure to capture global contextual information, recover lost spatial details, and generate accurate pixel-level segmentation labels.
- 3D convolutional kernels replace 2D convolutional kernels, allowing simultaneous processing of input data features in three dimensions (length, width, and height).
- 3DUNet effectively fuses 3D features at different levels, facilitating the extraction of complex shape and structural information. This framework demonstrates good performance in medical image segmentation.
- the initial CT image is segmented using a preprocessing strategy based on the 3DUNet structure to obtain the target image corresponding to the target detection region for subsequent processing.
- the method for detecting the liver region is used as an example for explanation.
- the initial plain CT image is resampled to a size of 0.7*0.7*5mm in the x, y, and z directions.
- the liver and several surrounding organs are masked on the initial plain CT image.
- the surrounding organs include the gallbladder, intrahepatic vessels, spleen, stomach, pancreas, right kidney, inferior vena cava, etc.
- the bounding boxes are expanded outward by a preset range of pixel distance to obtain the target image.
- the purpose of obtaining the regions of the liver and spleen is to use the spleen as a reference organ for the liver, so that when the subsequent model processes the image, it can learn not only the features of the target detection region (liver) but also the features of the reference detection region (spleen).
- the features of the liver need to be considered, while the features of the spleen do not need to be considered.
- Step 104 Input the multiple target images into the image processing model to obtain the detection results corresponding to the target detection regions.
- the image processing model generates object feature information and lesion feature information based on each target image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates detection results based on the object classification information, the lesion classification information, and the lesion segmentation information.
- multiple target images carried by the task can be obtained.
- the detection results corresponding to the target detection area output by the image processing model can be obtained.
- the detection results specifically include the object classification information of the object to be detected, the lesion segmentation information and lesion classification information on the target detection area.
- the object classification information of the object to be detected can be understood as classification information at the object level, which includes whether there are abnormal objects, the type of abnormal objects, and the probability of abnormal objects corresponding to each category.
- Lesion segmentation information can be understood as mask information for abnormal objects on the image to be detected.
- lesion segmentation information includes global mask information and local mask information.
- Global mask information can be understood as the mask information of all abnormal objects in the target detection region
- local mask information can be understood as the mask information for each type of abnormal object. For example, if there are 3 abnormal objects in a certain target detection region, the lesion segmentation information includes 4 parts: the complete mask information of the three abnormal objects, and the mask information corresponding to each of the three abnormal objects.
- Lesion classification information can be understood as determining the type of lesion in a single step after image processing model recognition, provided lesion segmentation information exists. Furthermore, lesion classification information includes lesion category information and organ category information. In the liver tumor detection provided in this disclosure, lesions are divided into 9 categories (hepatocellular carcinoma, cholangiocarcinoma, metastatic tumors, hemangiomas, focal nodular hyperplasia, cysts, calcifications, other malignant tumors, and other benign tumors), and organs are divided into 8 categories (liver, gallbladder, intrahepatic vessels, spleen, stomach, pancreas, right kidney, and inferior vena cava).
- the image processing model processes the multiple target images to generate object feature information and lesion feature information.
- the object feature information can be understood as feature information used to generate object classification information
- the lesion feature information can be understood as feature information used to generate lesion classification information and lesion segmentation information.
- the lesion feature information includes lesion sub-feature information and organ sub-feature information.
- the lesion sub-feature information is used to determine the lesion type
- the organ sub-feature information is used to determine the organ type.
- object feature information can be used to generate object classification information
- lesion feature information is fused with object feature information to improve prediction accuracy.
- object feature information and lesion feature information can be fused to obtain object fused feature information.
- the object fused feature information is still used for predicting object classification information; compared to object feature information, it includes relevant content related to lesion feature information, making subsequent predictions more accurate.
- the image processing model determines object classification information based on object fusion feature information, and determines lesion classification information and lesion segmentation information based on lesion feature information, thus forming the final detection result.
- the method provided in this disclosure further explains the model structure of the image processing model.
- the image processing model includes an image processing backbone network, an object branch network, a lesion branch network, and a fusion network.
- the image processing model includes an image processing backbone network, an object branch network, a lesion branch network, and a fusion network.
- the plurality of target images are input into an image processing model to obtain the detection results corresponding to the target detection regions, including S1042-S1048:
- S1042. Input the multiple target images into the image processing backbone network, and extract the feature information of the object to be processed and the feature information of the lesion to be processed from the image processing backbone network.
- Multiple target images can be combined to form a 3D detection region for the target detection area.
- the feature information of the object to be processed and the feature information of the lesion to be processed generated by the image processing backbone network during the processing of multiple target images can be extracted.
- both the object feature information and the lesion feature information are intermediate parameter features generated by the image processing model during the processing of multiple target images. Based on their respective uses, they are categorized as object feature information and lesion feature information. If a parameter feature is used to generate both object feature information and lesion feature information, then that parameter feature belongs to both categories.
- the image processing backbone network includes an image encoder and an image decoder
- the multiple target images are input into the image processing backbone network, and the feature information of the object to be processed and the feature information of the lesion to be processed are extracted from the image processing backbone network, including:
- the plurality of target images are input into the image encoder to obtain target image coding features and at least one multi-scale image coding feature;
- the target image encoding features are input into the image decoder to obtain at least one first multi-scale image decoding feature, at least one second multi-scale image decoding feature, and the target image decoding feature;
- the at least one multi-scale image coding feature and the at least one first multi-scale image decoding feature are determined as the feature information of the object to be processed, and the at least one second multi-scale image decoding feature and the target image decoding feature are determined as the feature information of the lesion to be processed.
- the method provided in this disclosure uses the nnU-Net framework to build an image processing model.
- the image processing backbone network includes an image encoder and an image decoder.
- the U-Net Pixel Encoder is used as the image encoder
- the feature pyramid network is used as the image decoder.
- target images When multiple target images are input into an image processing model, they are first encoded by an image encoder to obtain target image encoding features and at least one multi-scale image encoding feature.
- the target image encoding features are the encoding features output by the image encoder, and the at least one multi-scale image encoding feature is the encoding parameter feature generated by the image encoder during the processing of multiple target images.
- the target image encoding features are input into the image decoder to obtain at least one first multi-scale image decoding feature, at least one second multi-scale image decoding feature, and the final target image decoding feature generated by the image decoder during the decoding process.
- the multi-scale image coding features and the first multi-scale image decoding features are used as the feature information of the object to be processed, and the second multi-scale image decoding features and the target image decoding features are used as the feature information of the lesion to be processed.
- the image encoder includes multiple sequentially connected image coding layers
- the plurality of target images are input into the image encoder to obtain target image coding features and at least one multi-scale image coding feature, including:
- the plurality of target images are input into the image encoder to obtain the target image encoding features output by the image encoder;
- At least one target image coding layer is determined among the plurality of image coding layers, and the multi-scale image coding features output by each target image coding layer are obtained.
- the target image coding layer doesn't refer to a specific image coding layer, but rather to the image coding layer from which image coding features need to be extracted. For example, if there are six image coding layers in total, and the multi-scale image coding features output from the 1st, 3rd, and 5th image coding layers are needed, then the 1st, 3rd, and 5th image coding layers are the target image coding layers. Similarly, if there are twelve image coding layers in total, and the multi-scale image coding features output from the 9th, 10th, 11th, and 12th image coding layers are needed, then the 9th, 10th, 11th, and 12th image coding layers are the target image coding layers.
- the image encoder includes six image coding layers. Multiple target images are input into the image encoder for processing.
- the 5th and 6th image coding layers encode the target images, thus obtaining the target image coding features output by the image encoder, and the multi-scale image coding features output by the 5th and 6th image coding layers.
- the image decoder includes a plurality of sequentially connected image decoding layers
- the target image encoding features are input into the image decoder to obtain at least one first multi-scale image decoding feature and at least one second multi-scale image decoding feature, including:
- At least one first image decoding layer and at least one second image decoding layer are determined among the plurality of image decoding layers;
- the first image decoding layer and the second image decoding layer in the image decoding layer do not refer to a specific image decoding layer, but rather to the image decoding layer used to determine the first multi-scale image decoding features and the second multi-scale image decoding features.
- the image decoder includes 6 image decoding layers.
- the target image encoded features are input into the image decoder. It is determined that the output features of the 5th and 6th image decoding layers are the first multi-scale image decoding features, and the output features of the 2nd, 3rd, and 4th image decoding layers are the second multi-scale image decoding features. Therefore, the first multi-scale image decoding features output by the 5th and 6th image decoding layers are obtained. and and the second multi-scale image decoding features output from the 2nd, 3rd, and 4th image decoding layers. and and the target image decoding features output by the final image decoder
- multi-scale image coding features and First multi-scale image decoding features and The feature information of the object to be processed, the second multi-scale image decoding features and and target image decoding features The information constitutes the characteristic features of the lesion to be treated.
- S104 Input the feature information of the object to be processed into the object branch network to obtain object feature information, and input the feature information of the lesion to be processed into the lesion branch network to obtain lesion feature information and image feature information to be segmented.
- the object branch network is used to generate object classification information for the object to be processed. After obtaining the feature information of the object to be processed in the above steps, the feature information of the object to be processed is input into the object branch network and processed in the object branch network to obtain the object feature information.
- FIG. 3 illustrates a data processing diagram of an object branching network provided in one embodiment of this specification.
- the feature information of the object to be processed is iteratively input into a 4-layer Dual-path Transformer Block (DPB) for processing to obtain the object feature information.
- the Dual-path Transformer Block is a special Transformer structure that optimizes performance and reduces computational complexity by decomposing the processing into two parallel paths (usually called local paths and global paths), thereby improving processing efficiency and effectiveness. These two paths typically focus on different types of features; the local path focuses more on fine-grained local features, while the global path focuses more on scene-level global features.
- the lesion branch network is used to generate segmentation and classification information for lesions. After obtaining the lesion feature information in the above steps, the lesion feature information is input into the lesion branch network for processing to obtain lesion feature information and image feature information to be segmented.
- the lesion feature information to be processed includes at least one second multi-scale image decoding feature and a target image decoding feature;
- the lesion feature information to be processed is input into the lesion branch network to obtain lesion feature information and image feature information to be segmented, including:
- Feature decoding is performed on the decoding features of each second multi-scale image to obtain lesion feature information
- the lesion feature information and the target image decoding features are fused to generate image feature information to be segmented.
- lesion feature information and image feature information to be segmented are generated simultaneously.
- the lesion feature information is used to determine lesion classification information
- the image feature information to be segmented is used to determine lesion segmentation information.
- Figure 4 illustrates a data processing diagram of a lesion branch network provided in an embodiment of this specification.
- the lesion branch network includes a Transformer Decoder structure. Fifty randomly initialized queries are input, and the Transformer Decoder performs decoding processing on the second multi-scale image decoding features.
- the second multi-scale image decoding features ... and The 50 queries are input into the Transformer Decoder for processing, obtaining the lesion feature information output by the Transformer Decoder. Simultaneously, to better segment abnormal objects at the pixel level, the lesion feature information and the target image decoding features output by the image decoder are combined. The images are fused to obtain feature information of the image to be segmented, which is then used to mask abnormal objects in the detection area at the pixel level.
- the object feature information generated in the object branch network and the lesion feature information and image feature information to be segmented generated in the lesion branch network have been obtained.
- a fusion network can be understood as a network used to combine object feature information and lesion feature information to generate object fusion feature information.
- Object fusion feature information enables image processing models to combine global information of objects and local information of lesions, which can improve the detection capability of smaller abnormal objects.
- the object feature information and the lesion feature information are input into the fusion network to generate object fusion feature information, including:
- the object feature information and the lesion feature information are input into the fusion network.
- the lesion sub-feature information in the lesion feature information is extracted in the fusion network, and the object feature information and the lesion sub-feature information are spliced together to generate object fusion feature information.
- lesion feature information includes lesion sub-feature information and organ sub-feature information.
- the lesion sub-feature information and the object feature information can be fused to generate object fused feature information.
- FIG. 5 illustrates a data processing diagram of a fusion network provided in an embodiment of this specification
- lesion sub-feature information for determining the lesion type is extracted from the lesion feature information.
- object fusion feature information is obtained.
- Object fusion feature information is the feature information used to finally determine the object classification information. It includes both object feature information and lesion feature information, making the final prediction result more accurate.
- lesion classification information image feature information to be segmented, and object fusion feature information are obtained. Then, corresponding lesion classification information, lesion segmentation information, and object classification information are generated based on these three pieces of information. Finally, the final detection result is composed of the lesion classification information, lesion segmentation information, and object classification information.
- lesion classification information is generated based on the lesion feature information
- lesion segmentation information is generated based on the image feature information to be segmented, including:
- the lesion feature information is input into the lesion classifier to generate lesion classification information
- the feature information of the image to be segmented is input into the segmenter to generate lesion segmentation information.
- object classification information is generated, including:
- the object fusion feature information is input into the object classifier to generate object classification information.
- lesion feature information is input into a lesion classifier for processing to obtain lesion classification information output by the lesion classifier;
- lesion segmentation information is input into a segmenter for processing to obtain lesion segmentation information output by the segmenter;
- object fusion feature information is input into an object classifier for processing to obtain object classification information output by the object classifier.
- the image processing method provided in this disclosure inputs multiple target images into an image processing model.
- object feature information and lesion feature information are generated based on each target image.
- the object feature information and lesion feature information are fused to generate object fusion feature information. This allows the model to use not only object feature information but also lesion feature information when predicting the final object classification information, thus fusing global and local information and improving the detection capability of small lesions.
- the image processing model provided in this disclosure is a pre-trained image processing model.
- Figure 6 shows a flowchart of an image processing model training method provided in an embodiment of this specification. As shown in Figure 6, the image processing model is trained and generated through the following steps:
- Step 602 Obtain multiple sample target images and the corresponding sample detection results, wherein the sample detection results include sample object classification information, sample lesion classification information and sample lesion segmentation information.
- the image processing model training method uses supervised training, which includes training sample pairs.
- Each training sample pair comprises multiple sample target images targeting the target detection region, and corresponding sample detection results for these multiple sample target images. These multiple sample target images can be combined to form a three-dimensional image of the target detection region.
- the sample detection results specifically include sample object classification information, sample lesion classification information, and sample lesion segmentation information.
- Sample object classification information includes whether there are abnormal objects in the target detection area of the object, the type of abnormal object, and the probability of the abnormal object corresponding to each category.
- Sample lesion classification information can be understood as the lesion classification for the lesion corresponding to the target detection area, and the organ classification of the organ where the lesion is located.
- Sample lesion segmentation information can be understood as the mask information for abnormal objects in the target detection area.
- the image processing model is trained in a supervised manner using multiple sample target images and the sample detection results corresponding to the multiple sample target images.
- acquiring multiple sample target images and corresponding sample detection results includes:
- sample detection object Identify the sample detection object and obtain the sample enhancement image and sample target image corresponding to the sample detection object, wherein the sample enhancement image and sample target image corresponding to the sample detection object correspond to the same target detection region of the sample detection object;
- a reference sample detection result is generated on the enhanced sample image for the target detection region.
- the sample detection results are labeled on the sample target image.
- a sample-enhanced image refers to an image whose contrast has been enhanced for the target detection region
- a sample target image refers to a regular image of the target detection region.
- the sample-enhanced image and the sample target image correspond to the same target detection region of the same sample detection object.
- the corresponding sample enhancement image and sample target image are obtained based on the target detection area of the sample detection object.
- a sample-enhanced image is an image whose contrast has been enhanced for the target detection region. This allows for better acquisition of the detection results for the target detection region within the sample-enhanced image.
- the detection results for the target detection region in the sample-enhanced image serve as the reference sample detection results.
- the reference sample detection results are then mapped to the target image through registration to obtain the corresponding sample detection results for the target image.
- the process of generating reference sample detection results for target detection regions on augmented images is usually done through manual annotation.
- manual annotation requires experienced technicians, which is time-consuming, labor-intensive, and costly.
- technicians can annotate a portion of the augmented image. Training sample pairs of augmented images and reference sample detection results are formed. A augmented image annotation model is then trained using these training sample pairs. After the model is trained, unannotated augmented images are input into the model for annotation, thereby generating reference sample detection results for the augmented image.
- paired enhanced CT images and plain CT images of acceptable quality are obtained from the radiology department.
- the paired enhanced CT images and plain CT images are CT images of the same sample subject taken at the same time and from the same angle.
- Enhanced CT images include arterial phase enhanced CT, venous phase enhanced CT, and delayed phase enhanced CT, etc.
- An enhanced CT model is trained using labeled enhanced CT images and reference sample detection results, enabling the model to predict and annotate detection results based on enhanced CT images. Unlabeled enhanced CT images are then input into this model for processing. The model processes the unlabeled enhanced CT images to generate reference sample detection results corresponding to the unlabeled enhanced CT images.
- the detection results of the reference sample on the enhanced CT image can be transferred to the plain CT image as the sample detection results of the plain CT image.
- Step 604 Input the multiple sample target images into the image processing model to obtain the prediction detection results and training contrast loss values output by the image processing model.
- the image processing model outputs predicted detection results and training comparison loss values based on the sample target images.
- the concept of comparison loss value is introduced, that is, a two-stage class-balanced comparison loss function is constructed, and the comparison loss function is processed in the object branch network and the lesion branch network of the image processing model respectively, thereby obtaining the training comparison loss value during the training process.
- the multiple sample target images are input into an image processing model to obtain the predicted detection results and training contrastive loss values output by the image processing model, including:
- the multiple sample target images are input into the image processing backbone network to extract the feature information of the sample object to be processed and the feature information of the sample lesion to be processed from the image processing backbone network;
- the feature information of the sample object to be processed is input into the object branch network to obtain the predicted object feature information
- the feature information of the sample lesion to be processed is input into the lesion branch network to obtain the predicted lesion feature information and the predicted image feature information to be segmented.
- the object comparison loss value is obtained based on the sample object classification information corresponding to the sample target image; in the lesion branch network, the lesion comparison loss value is obtained based on the sample lesion classification information corresponding to the sample target image.
- the predicted object feature information and the predicted lesion feature information are input into the fusion network to generate sample object fusion feature information;
- First object classification information is generated based on the predicted object feature information; second object classification information is generated based on the sample object fusion feature information; predicted lesion classification information is generated based on the predicted lesion feature information; and predicted lesion segmentation information is generated based on the predicted image to be segmented feature information.
- a prediction detection result is generated based on the first object classification information, the second object classification information, the predicted lesion classification information, and the predicted lesion segmentation information.
- a training comparison loss value is generated based on the object comparison loss value and the lesion comparison loss value.
- model structure of the image processing model is the same as that of the model application phase described above.
- model structure and data processing flow of the image processing model please refer to the relevant sections of the model application phase described above, which will not be repeated here.
- the image processing model first processes multiple sample target images in the image processing backbone network. It then extracts the feature information of the sample objects to be processed and the feature information of the sample lesions to be processed generated in the image processing backbone network.
- the feature information of the sample objects to be processed is input into the object branch network to obtain predicted object feature information; the feature information of the sample lesions to be processed is input into the lesion branch network to obtain predicted lesion feature information and predicted image segmentation feature information.
- the data processing procedures of the image processing backbone network, object branch network, and lesion branch network during the model training phase are the same as those in the application phase described above, and will not be repeated here.
- a two-stage class-balanced contrastive loss function is introduced during model training.
- a two-class contrastive learning loss function is introduced to determine whether there are abnormal objects in the sample object.
- a contrastive learning loss function for several abnormal objects is introduced.
- the image processing model learns the differences in the feature information of the sample objects to be processed between different training batches, and the differences in the feature information of the sample lesions to be processed between different training batches.
- the contrastive loss function yields the object contrastive loss value and the lesion contrastive loss value.
- the training method of this disclosure uses a method that references historical sample object feature information.
- the sample object feature information is input into the object branch network to obtain predicted object feature information, including:
- the historical object feature information and the sample object feature information to be processed are input into the object branch network to obtain the predicted object feature information;
- the lesion feature information of the sample to be processed is input into the lesion branch network to obtain predicted lesion feature information and predicted image feature information to be segmented, including:
- the lesion feature information of the historical samples and the lesion feature information of the samples to be processed are input into the lesion branch network to obtain the predicted lesion feature information and the predicted image feature information to be segmented.
- a memory bank strategy is introduced, adding object feature storage blocks and lesion feature storage blocks to the object branch network and lesion branch network, respectively.
- the object feature storage block stores historical sample object feature information
- the lesion feature storage block stores historical sample lesion feature information. Both feature storage blocks have the same capacity.
- the number of features stored in each category in both feature storage blocks is balanced.
- the number of features already stored in each category is also considered to ensure that the number of features in each category remains balanced.
- the training method provided in this disclosure also generates first object classification information based on the predicted object feature information, and generates second object classification information based on the sample object fusion feature information. That is, during the model training process, two object classification information will be generated. One is the first object classification information generated based on the predicted object feature information, and the other is the second object classification information generated based on the predicted object feature information and the predicted lesion feature information. Both will be used as the prediction results of the model during the model training stage and used to calculate the model loss value in the subsequent calculation.
- predicted lesion classification information will be generated based on the predicted lesion feature information
- predicted lesion segmentation information will be generated based on the predicted image feature information
- the first object classification information, the second object classification information, the predicted lesion classification information, and the predicted lesion segmentation information are determined to generate the predicted detection result.
- the object alignment loss value and the lesion alignment loss value are determined to be the training contrastive loss value.
- Step 606 Calculate the predicted loss value based on the predicted detection results and the sample detection results.
- the predicted detection results are compared with the sample detection results to calculate the predicted loss value.
- the image processing model is not yet trained. By calculating the predicted loss value, the difference between the predicted results and the sample results is determined, thereby further adjusting the parameters of the image processing model and achieving the training of the image processing model.
- calculating the predicted loss value based on the predicted detection result and the sample detection result includes:
- the lesion segmentation loss value is calculated based on the predicted lesion segmentation information and the sample lesion segmentation information.
- the predicted detection results include first object classification information and second object classification information. Both are used to calculate loss values along with the sample object classification information. Specifically, the first object loss value is calculated using the first object classification information and the sample object classification information, and the second object loss value is calculated using the second object classification information and the sample object classification information. Similarly, the lesion classification loss value is calculated based on the predicted lesion classification information and the sample lesion classification information, and the lesion segmentation loss value is calculated based on the predicted lesion segmentation information and the predicted lesion classification information.
- model loss values such as the cross-entropy loss function, the maximum loss function, the average loss function, etc. This specification does not limit the specific method of the loss function; the actual application shall prevail.
- Step 608 Adjust the model parameters of the image processing model according to the predicted loss value and the training contrast loss value, and continue to train the image processing model until the model training stops.
- the model parameters of the image processing model can be adjusted by backpropagation based on these two types of loss values.
- adjusting the model parameters of the image processing model based on the predicted loss value and the training contrastive loss value includes:
- the model loss value is calculated based on the first object loss value, the second object loss value, the lesion classification loss value, the lesion segmentation loss value, the object comparison loss value, and the lesion comparison loss value;
- the predicted loss value and the training comparison loss value are fused to obtain a new loss value. Furthermore, according to the preset loss weights, the loss values of the first object, the second object, the lesion classification, the lesion segmentation, the object comparison, and the lesion comparison can be fused to obtain a new model loss value.
- the model parameters of the image processing model are then adjusted based on the model loss value until the model training stopping condition is met, and a trained image processing model is obtained.
- the image processing model training method disclosed herein proposes to train a sample augmentation image annotation model using sample augmentation images with partial manual annotation. This allows the sample augmentation image annotation model to annotate sample augmentation images that have not been manually annotated, saving annotation time. Furthermore, by registering the sample augmentation image and the target image, the detection results of reference samples on the sample augmentation image are transferred to the target image to obtain the sample detection results of the target image. This solves the problems of difficult and missing annotations on the target image, reducing annotation costs.
- FIG. 7 shows a flowchart of a CT image processing method according to an embodiment of this specification, specifically including the following steps.
- Step 702 Receive a CT image processing task, wherein the CT image processing task carries multiple plain CT images corresponding to the target detection area, and the CT image processing task is used to detect whether there is an abnormal object in the target detection area.
- Step 704 Input the multiple plain CT images into the image processing model to obtain the detection result corresponding to the target detection area.
- the image processing model generates object feature information and lesion feature information based on each plain CT image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates the detection result based on the object classification information, the lesion classification information, and the lesion segmentation information.
- steps 702 to 704 are the same as those of steps 102 to 104 above, and will not be repeated here.
- the method is further explained using a CT image processing task to detect the presence of abnormal objects within a target detection area.
- the CT image processing task includes taking a plain CT image of the target detection area and inputting it into an image processing model for identification.
- the image processing model can identify whether abnormal objects exist in the plain CT image, and if so, the mask information of the abnormal objects. It can also predict the object classification information of the object to be detected, i.e., whether an abnormal object exists in the target detection area of the object to be detected, and the type of the abnormal object. This method solves the problem of poor detection effect and low accuracy in current applications where plain CT images cannot be used to detect target detection areas.
- Figure 8 shows a flowchart of a computer-aided diagnosis method for cancer according to an embodiment of this specification, specifically including:
- Step 802 Receive a CT image processing task, wherein the CT image processing task carries multiple plain CT images corresponding to the target detection area, and the CT image processing task is used to detect whether there is a tumor in the target detection area.
- Step 804 Input the multiple plain CT images into the image processing model to obtain the detection result corresponding to the target detection area.
- the image processing model generates object feature information and lesion feature information based on each plain CT image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates the detection result based on the object classification information, the lesion classification information, and the lesion segmentation information.
- the computer-aided cancer diagnosis method provided in this embodiment is applicable to screening for the presence of tumors in various target detection areas, including but not limited to the presence of tumors in organs throughout the body such as the pancreas, esophagus, liver, lungs, breast, intestines, stomach, and lymph nodes.
- the detection results include user classification information (i.e., whether a tumor exists, and whether the tumor is malignant), lesion classification information (what type of tumor it is), and lesion segmentation information (mask information marking the tumor on the CT image).
- the method provided in this disclosure can offer guidance to doctors, helping to improve their diagnostic accuracy and providing data support for their diagnostic results.
- FIG. 9 shows a flowchart of a computer-aided diagnostic method for liver cancer according to an embodiment of this specification, specifically including:
- Step 902 Receive a CT image processing task, wherein the CT image processing task carries multiple plain CT images corresponding to the liver region, and the CT image processing task is used to detect whether there is a tumor in the liver region.
- Step 904 Input the multiple plain CT images into the image processing model to obtain the detection results corresponding to the liver region.
- the image processing model generates object feature information and lesion feature information based on each plain CT image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates detection results based on the object classification information, the lesion classification information, and the lesion segmentation information.
- the computer-aided diagnostic method for liver cancer provided in this embodiment can be used to screen for diseases such as hepatocellular carcinoma, cholangiocarcinoma, metastatic tumors, hemangiomas, focal nodular hyperplasia, cysts, and calcifications. It provides guidance to doctors, helps improve their diagnostic accuracy, and provides data support for doctors to give diagnostic results.
- diseases such as hepatocellular carcinoma, cholangiocarcinoma, metastatic tumors, hemangiomas, focal nodular hyperplasia, cysts, and calcifications. It provides guidance to doctors, helps improve their diagnostic accuracy, and provides data support for doctors to give diagnostic results.
- the computer-aided cancer diagnosis system may include a client 100 and a server 200.
- Client 100 is used to send a CT image processing task to server 200, wherein the CT image processing task carries multiple plain CT images corresponding to the target detection area, and the CT image processing task is used to detect whether there is a tumor in the target detection area;
- the image processing model generates object feature information and lesion feature information based on each plain CT image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, generates the detection result based on the object classification information, the lesion classification information, and the lesion segmentation information, and sends the detection result corresponding to the target detection area to client 100.
- Client 100 is also used to receive the detection results sent by server 200.
- a computer-aided diagnosis system for cancer may include multiple clients 100 and a server 200.
- Clients 100 can be referred to as edge devices, and the server 200 as cloud devices. Multiple clients 100 can establish communication connections through the server 200.
- the server 200 is used to provide computer-aided diagnosis services for cancer among the multiple clients 100.
- Each client 100 can act as either a sender or a receiver, communicating through the server 200.
- Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc.
- users can publish data streams to server 200 through client 100, server 200 can generate test results based on the data stream, and push the test results to other clients that have established communication.
- client 100 and server 200 establish a connection via a network.
- the network provides the medium for communication between client 100 and server 200.
- the network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.
- Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application.
- Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK.
- SDK software development kit
- Client 100 can be deployed on electronic devices and depends on the device or certain apps on the device to run. Electronic devices may have displays and support information browsing, such as personal mobile terminals like mobile phones, tablets, and personal computers.
- Various other types of applications can also be configured on electronic devices, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.
- Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server.
- the server can also be a server in a distributed system, or a server integrated with blockchain.
- the server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
- the computer-aided diagnosis method for cancer provided in this disclosure is generally executed by the server.
- the client may also have similar functions to the server, thereby executing the computer-aided diagnosis method for cancer provided in this disclosure.
- the computer-aided diagnosis method for cancer provided in this disclosure may also be executed jointly by the client and the server.
- Figure 11 shows a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification. As shown in Figure 11, the apparatus includes:
- the receiving module 1102 is configured to receive an image processing task, wherein the image processing task carries multiple target images corresponding to the target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area.
- the detection module 1104 is configured to input the plurality of target images into an image processing model to obtain detection results corresponding to the target detection regions.
- the image processing model generates object feature information and lesion feature information based on each target image, generates object fusion feature information based on the object feature information and the lesion feature information, determines object classification information based on the object fusion feature information, determines lesion classification information and lesion segmentation information based on the lesion feature information, and generates detection results based on the object classification information, the lesion classification information, and the lesion segmentation information.
- the image processing model includes an image processing backbone network, an object branch network, a lesion branch network, and a fusion network;
- the detection module 1104 is further configured as follows:
- the multiple target images are input into the image processing backbone network to extract the feature information of the object to be processed and the feature information of the lesion to be processed from the image processing backbone network;
- the feature information of the object to be processed is input into the object branch network to obtain object feature information
- the feature information of the lesion to be processed is input into the lesion branch network to obtain lesion feature information and image feature information to be segmented
- the object feature information and the lesion feature information are input into the fusion network to generate object fusion feature information;
- lesion classification information is generated; based on the image feature information to be segmented, lesion segmentation information is generated; based on the object fusion feature information, object classification information is generated; and based on the object classification information, the lesion classification information, and the lesion segmentation information, a detection result is generated.
- the image processing backbone network includes an image encoder and an image decoder
- the detection module 1104 is further configured as follows:
- the plurality of target images are input into the image encoder to obtain target image coding features and at least one multi-scale image coding feature;
- the target image encoding features are input into the image decoder to obtain at least one first multi-scale image decoding feature, at least one second multi-scale image decoding feature, and the target image decoding feature;
- the at least one multi-scale image coding feature and the at least one first multi-scale image decoding feature are determined as the feature information of the object to be processed, and the at least one second multi-scale image decoding feature and the target image decoding feature are determined as the feature information of the lesion to be processed.
- the image encoder includes multiple sequentially connected image coding layers
- the detection module 1104 is further configured as follows:
- the multiple target images are input into the image encoder to obtain the target image encoding features output by the image encoder;
- At least one target image coding layer is determined among the plurality of image coding layers, and the multi-scale image coding features output by each target image coding layer are obtained.
- the image decoder includes multiple sequentially connected image decoding layers
- the detection module 1104 is further configured as follows:
- At least one first image decoding layer and at least one second image decoding layer are determined among the plurality of image decoding layers;
- the lesion feature information to be processed includes at least one second multi-scale image decoding feature and a target image decoding feature;
- the detection module 1104 is further configured as follows:
- Feature decoding is performed on the decoding features of each second multi-scale image to obtain lesion feature information
- the detection module 1104 is further configured to:
- the lesion feature information is input into the lesion classifier to generate lesion classification information
- the feature information of the image to be segmented is input into the segmenter to generate lesion segmentation information.
- the detection module 1104 is further configured to:
- the object feature information and the lesion feature information are input into the fusion network.
- the lesion sub-feature information in the lesion feature information is extracted in the fusion network, and the object feature information and the lesion sub-feature information are spliced together to generate object fusion feature information.
- the detection module 1104 is further configured to:
- the object fusion feature information is input into the object classifier to generate object classification information.
- the device further includes a training module configured to:
- sample detection results include sample object classification information, sample lesion classification information, and sample lesion segmentation information
- the multiple sample target images are input into the image processing model to obtain the prediction detection results and training contrast loss values output by the image processing model.
- the model parameters of the image processing model are adjusted based on the predicted loss value and the training contrast loss value, and the image processing model is trained until the model training stops.
- the training module is further configured to:
- sample detection object Identify the sample detection object and obtain the sample enhancement image and sample target image corresponding to the sample detection object, wherein the sample enhancement image and sample target image corresponding to the sample detection object correspond to the same target detection region of the sample detection object;
- a reference sample detection result is generated on the enhanced sample image for the target detection region.
- the sample detection results are labeled on the sample target image.
- the image processing model includes an image processing backbone network, an object branch network, a lesion branch network, and a fusion network;
- the training module is further configured as follows:
- the multiple sample target images are input into an image processing model to obtain the predicted detection results and training contrastive loss values output by the image processing model, including:
- the multiple sample target images are input into the image processing backbone network to extract the feature information of the sample object to be processed and the feature information of the sample lesion to be processed from the image processing backbone network;
- the feature information of the sample object to be processed is input into the object branch network to obtain the predicted object feature information
- the feature information of the sample lesion to be processed is input into the lesion branch network to obtain the predicted lesion feature information and the predicted image feature information to be segmented.
- the object comparison loss value is obtained based on the sample object classification information corresponding to the sample target image; in the lesion branch network, the lesion comparison loss value is obtained based on the sample lesion classification information corresponding to the sample target image.
- the predicted object feature information and the predicted lesion feature information are input into the fusion network to generate sample object fusion feature information;
- First object classification information is generated based on the predicted object feature information; second object classification information is generated based on the sample object fusion feature information; predicted lesion classification information is generated based on the predicted lesion feature information; and predicted lesion segmentation information is generated based on the predicted image to be segmented feature information.
- a prediction detection result is generated based on the first object classification information, the second object classification information, the predicted lesion classification information, and the predicted lesion segmentation information.
- a training comparison loss value is generated based on the object comparison loss value and the lesion comparison loss value.
- the training module is further configured to:
- the lesion segmentation loss value is calculated based on the predicted lesion segmentation information and the sample lesion segmentation information.
- the training module is further configured to:
- the historical object feature information and the sample object feature information to be processed are input into the object branch network to obtain the predicted object feature information;
- the lesion feature information of the historical samples and the lesion feature information of the samples to be processed are input into the lesion branch network to obtain the predicted lesion feature information and the predicted image feature information to be segmented.
- the training module is further configured to:
- the model loss value is calculated based on the first object loss value, the second object loss value, the lesion classification loss value, the lesion segmentation loss value, the object comparison loss value, and the lesion comparison loss value.
- the image processing apparatus inputs multiple target images into an image processing model.
- object feature information and lesion feature information are generated based on each target image.
- the object feature information and lesion feature information are fused to generate object fusion feature information. This allows the model to use not only object feature information but also lesion feature information when predicting the final object classification information, thus fusing global and local information and improving the detection capability of small lesions.
- the training module provided in this disclosure proposes training a sample augmentation image annotation model using partially manual annotation of the augmented images. This allows the model to annotate sample augmentation images that have not been manually annotated, saving annotation time. Additionally, by registering the augmented images with the target images, the detection results of reference samples from the augmented images are transferred to the target images to obtain the target image's detection results. This solves the problems of difficult and missing annotations on target images, reducing annotation costs.
- Figure 12 shows a structural block diagram of a computing device 1200 according to an embodiment of this specification.
- the components of the computing device 1200 include, but are not limited to, a memory 1210 and a processor 1220.
- the processor 1220 is connected to the memory 1210 via a bus 1230, and a database 1250 is used to store data.
- the computing device 1200 also includes an access device 1240, which enables the computing device 1200 to communicate via one or more networks 1260.
- networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet.
- Access device 1240 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
- NIC network interface card
- the aforementioned components of the computing device 1200 may be interconnected, for example, via a bus. It should be understood that the block diagram of the computing device shown in FIG. 12 is merely illustrative and not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
- the computing device 1200 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs).
- the computing device 1200 can also be a mobile or stationary server.
- the processor 1220 is configured to execute the following computer program/instructions, which, when executed by the processor, implement the steps of the above-mentioned image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer.
- the various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments.
- the computing device embodiments are basically similar to the image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer embodiments, so the description is relatively simple. Relevant parts can be referred to the descriptions of the image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer embodiments.
- An embodiment of this specification also provides a computer-readable storage medium storing a computer program/instructions that, when executed by a processor, implement the steps of the above-described image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer.
- the various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments.
- the computer-readable storage medium embodiments are relatively simple in description because they are substantially similar to the image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer embodiments. Relevant parts can be referred to the descriptions of the image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer embodiments.
- An embodiment of this specification also provides a computer program product, including a computer program/instructions that, when executed by a processor, implement the steps of the above-described image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer.
- the above is an illustrative scheme of a computer program product according to this embodiment.
- the technical solution of this computer program product belongs to the same concept as the technical solutions of the image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer described above.
- the technical solutions of the image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer described above please refer to the descriptions of the technical solutions of the image processing method, CT image processing method, computer-aided diagnosis method for cancer, or computer-aided diagnosis method for liver cancer described above.
- the computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms.
- the computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Evolutionary Computation (AREA)
- General Physics & Mathematics (AREA)
- Medical Informatics (AREA)
- Software Systems (AREA)
- Artificial Intelligence (AREA)
- Computing Systems (AREA)
- Biomedical Technology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Computational Linguistics (AREA)
- Multimedia (AREA)
- General Engineering & Computer Science (AREA)
- Molecular Biology (AREA)
- Biophysics (AREA)
- Mathematical Physics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Public Health (AREA)
- Epidemiology (AREA)
- Primary Health Care (AREA)
- Pathology (AREA)
- Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
- Radiology & Medical Imaging (AREA)
- Quality & Reliability (AREA)
- Apparatus For Radiation Diagnosis (AREA)
- Image Processing (AREA)
Abstract
本公开提供图像处理方法、癌症的计算机辅助诊断方法,其中图像处理方法包括:接收图像处理任务,其中,图像处理任务携带目标检测区域对应的多个目标图像,图像处理任务用于检测目标检测区域内是否存在异常对象;将多个目标图像输入至图像处理模型,获得目标检测区域对应的检测结果,其中,图像处理模型基于各目标图像生成对象特征信息和病灶特征信息,根据对象特征信息和病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据对象分类信息、病灶分类信息和病灶分割信息生成检测结果。通过本方法,提升了图像检测的准确率。
Description
本公开要求申请号为202410927473.X的中国专利申请的优先权,该中国专利申请于2024年07月10日提交中国专利局,申请名称为“图像处理方法、癌症的计算机辅助诊断方法”,其全部内容通过引用结合在本公开中。
本公开涉及计算机技术领域,特别涉及一种图像处理方法。
随着人们生活水平的提高,越来越多的人开始注重自己的健康问题,癌症作为影响人们身体健康的重要因素,早期的诊断对癌症的治疗有积极的作用。随着计算机技术的不断发展,各种学习模型逐渐被应用于各种应用场景的预测,图像处理模型在医学影像计算机辅助诊断(computer aided diagnosis,CAD)任务中也取得了显著成功,从医学图像中进行病理分析和分类是计算机辅助诊断中的一个重要课题。
目前肿瘤的筛查技术主要是基于增强CT的,增强CT虽然精度高,但是增强CT会使患者暴露于造影剂,造影剂可能会引起过敏反应或器官衰竭。另外增强CT检查受到技术和设备的限制,无法在所有地区进行24小时使用,其便捷性受到影响。平扫CT由于分辨率不高,在平扫CT上的病灶视觉上难以辨认,检出率会比增强CT稍差,在平扫CT上为病灶进行标注也存在一定难度。因此,如何使用平扫CT进行肿瘤筛查,提升利用平扫CT进行肿瘤筛查的准确率。
有鉴于此,本公开提供了一种图像处理方法。本说明书一个或者多个实施例同时涉及一种CT图像处理方法,一种癌症的计算机辅助诊断方法,一种肝癌的计算机辅助诊断方法,一种计算设备,一种计算机可读存储介质以及一种计算机程序产品,以解决现有技术中存在的技术缺陷。
根据本公开的第一方面,提供了一种图像处理方法,包括:
接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象;
将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各目标图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
根据本公开的第二方面,提供了一种CT图像处理方法,包括:
接收CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在异常对象;
将所述多个平扫CT图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
根据本公开的第三方面,提供了一种癌症的计算机辅助诊断方法,包括:
接收CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在肿瘤;
将所述多个平扫CT图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
根据本公开的第四方面,提供了一种肝癌的计算机辅助诊断方法,包括:
接收CT图像处理任务,其中,所述CT图像处理任务携带肝脏区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述肝脏区域内是否存在肿瘤;
将所述多个平扫CT图像输入至图像处理模型,获得所述肝脏区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
根据本公开的第五方面,提供了一种癌症的计算机辅助诊断系统,包括客户端和服务端,其中:
所述客户端,用于向所述服务端发送CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在肿瘤;
所述服务端,用于将所述多个平扫CT图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
根据本公开的第六方面,提供了一种计算设备,包括:
存储器和处理器;
所述存储器用于存储计算机程序/指令,所述处理器用于执行所述计算机程序/指令,该计算机程序/指令被处理器执行时实现上述方法的步骤。
根据本公开的第七方面,提供了一种计算机可读存储介质,其存储有计算机程序/指令,该计算机程序/指令被处理器执行时实现上述方法的步骤。
根据本公开的第八方面,提供了一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现上述方法的步骤。
本说明书一个实施例提供了一种图像处理方法,包括接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象;将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各目标图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
通过本公开提供的图像处理方法,将多个目标图像输入至图像处理模型,在图像处理模型中,基于各目标图像生成对象特征信息和病灶特征信息,同时将对象特征信息和病灶特征信息融合,生成对象融合特征信息,使得模型在预测最终的对象分类信息时,除了使用对象特征信息,还会参考病灶特征信息,对全局信息和局部信息进行融合,提升小病灶的检出能力。
图1是本说明书一个实施例提供的一种图像处理方法的流程图;
图2是本说明书一个实施例提供的图像处理模型的模型结构示意图;
图3是本说明书一个实施例提供的对象分支网络的数据处理示意图;
图4是本说明书一个实施例提供的病灶分支网络的数据处理示意图;
图5是本说明书一个实施例提供的融合网络的数据处理示意图;
图6是本说明书一个实施例提供的图像处理模型训练方法的流程图;
图7是本说明书一个实施例提供的一种CT图像处理方法的处理过程流程图;
图8是本说明书一个实施例提供的一种癌症的计算机辅助诊断方法的流程示意图;
图9是本说明书一个实施例提供的一种肝癌的计算机辅助诊断方法的流程示意图;
图10是本说明书一个实施例提供的一种癌症的计算机辅助诊断系统的架构图;
图11是本说明书一个实施例提供的一种图像处理装置的结构示意图;
图12是本说明书一个实施例提供的一种计算设备的结构框图。
在下面的描述中阐述了很多具体细节以便于充分理解本说明书。但是本说明书能够以很多不同于在此描述的其它方式来实施,本领域技术人员可以在不违背本说明书内涵的情况下做类似推广,因此本说明书不受下面公开的具体实施的限制。
在本说明书一个或多个实施例中使用的术语是仅仅出于描述特定实施例的目的,而非旨在限制本说明书一个或多个实施例。在本说明书一个或多个实施例和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。还应当理解,本说明书一个或多个实施例中使用的术语“和/或”是指并包含一个或多个相关联的列出项目的任何或所有可能组合。
应当理解,尽管在本说明书一个或多个实施例中可能采用术语第一、第二等来描述各种信息,但这些信息不应限于这些术语。这些术语仅用来将同一类型的信息彼此区分开。例如,在不脱离本说明书一个或多个实施例范围的情况下,第一也可以被称为第二,类似地,第二也可以被称为第一。取决于语境,如在此所使用的词语“如果”可以被解释成为“在……时”或“当……时”或“响应于确定”。
需要说明的是,本说明书所涉及的用户信息(包括但不限于用户设备信息、用户个人信息等)和数据(包括但不限于用于分析的数据、存储的数据、展示的数据等),均为经用户授权或者经过各方充分授权的信息和数据,并且相关数据的收集、使用和处理需要遵守相关地区的相关法律法规和标准,并提供有相应的操作入口,供用户选择授权或者拒绝。
首先,对本说明书一个或多个实施例涉及的名词术语进行解释。
CT(Computed Tomography):电子计算机断层扫描,它是利用精确准直的X线束,与灵敏度极高的探测器一同围绕人体的某一部位作一个接一个的断面扫描,具有扫描时间快,图像清晰等特点,可用于多种疾病的检查。
平扫CT:又称普通扫描,是指静脉内不给含造影剂的扫描。
增强CT:是指血管内注射对比剂后再行扫描的方法,目的是提高病变组织同正常组织的密度差,以显示平扫CT上未被显示或显示不清楚的病变,通过有无强化及强化类型,有助于病变的定性。
CAD(computer aided diagnosis):计算机辅助诊断,是指通过影像学、医学图像处理技术以及其他可能的生理、生化手段,结合计算机的分析计算,辅助发现病灶,提高诊断的准确率。
随着人民生活水平提高,越来越多的人重视自身的健康,肿瘤是影响人健康的重大因素之一,在医学图像中识别肿瘤是需要专业医生根据经验进行识别,受限于医生的经验,借助医学图像的进行图像识别分析成为一个重要课题。
现在的癌症筛查技术主要是基于增强CT的,增强CT虽然精度高,但是增强CT会使患者暴露于造影剂,造影剂可能会引起过敏反应或器官衰竭。另外增强CT检查受到技术和设备的限制,无法在所有地区进行24小时使用,其便捷性受到影响。还有一些筛查技术是基于超声诊断的,但是灵敏度不高,容易漏诊。目前还有一些技术是基于平扫CT的,但是平扫CT由于分辨率不高,器官上的病灶视觉上难以辨认,因此检出率比增强CT稍差。因此,如何能更好的进行肿瘤检测就成为技术人员亟待解决的问题。
基于此,在本说明书中,提供了一种图像处理方法。本说明书一个或者多个实施例同时涉及一种CT图像处理方法,一种癌症的计算机辅助诊断方法,一种肝癌的计算机辅助诊断方法,一种计算设备,一种计算机可读存储介质以及一种计算机程序产品,在下面的实施例中逐一进行详细说明。
参见图1,图1示出了根据本说明书一个实施例提供的一种图像处理方法的流程图,具体包括以下步骤。
步骤102:接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象。
在实际应用中,可以通过服务端,也可以通过客户端接收用户发送的图像处理任务。
具体的,图像处理任务可以理解为用于检测目标检测区域内是否存在异常对象的任务。图像处理任务重携带有目标检测区域对应的多个目标图像。更进一步的,目标检测区域可以理解为用于检测是否存在异常对象的分区,异常对象可以理解为目标检测区域内的异物。例如异常对象可以是人体内的肿瘤。
目标检测区域可以是人体内的任一器官,如肝脏、肺、胃、食管等。通过对目标检测区域是否存在异常对象进行预测,可以根据预测结果进一步判断待检测对象的状态,从而对异常对象的定位和精准治疗提供帮助。
待检测对象可以理解为目标检测区域所述的对象,例如,目标图像是张三胃部的CT图像,则目标检测区域为胃部,CT图像即为目标图像,待检测对象即为张三,该图像处理任务用于检测胃部区域是否存在有肿瘤。在实际应用中,待检测对象可以是人,也可以是其他生命体,在本说明书提供的一个或多个具体实施方式中,对此不做限定。
需要说明的是,本说明书一个或多个实施例中,图像处理任务可以应用于对各类医学图像的识别,并根据图像特征判断医学图像中的目标检测区域中是否存在异常对象。示例性的,在胃癌检测的应用场景中,可以根据胃部区域的医学图像,预测在胃部区域中是否存在肿瘤,从而帮助医生对出现异常的部位进行精准定位;在肝癌检测的应用场景中,可以根据肝脏区域的医学图像,预测在肝脏中是否存在肿瘤,从而帮助医生对出现异常的部位进行精准定位,便于后续的治疗。
示例性的,在肝癌检测的场景中,获取的目标图像即为肝脏的图像,具体的,获取的多个目标图像即为肝脏的平扫CT图像,多个目标图像可以组成肝脏区域的3D图像。异常对象可以理解为肝脏中的肿瘤。获取肝脏区域对应的多个平扫CT图像,通过对多个平扫CT图像进行图像检测处理,检测在肝脏中是否存在肿瘤。
在实际应用中,异常对象可以为某一种细胞,某一种组织结构等等,例如可以为恶性肿瘤、良性肿瘤、增生组织等等。在本说明书提供的一个或多个实施例中,对此不做限定。
通过接收图像处理任务,可以将图像处理任务中携带的目标检测区域对应的多个目标图像作为输入,用于检测目标检测区域内是否存在异常对象。
在本说明书提供的一具体实施方式中,在接收图像处理任务之前,还包括:
接收图像分割任务,其中,所述图像分割任务携带目标检测区域对应的多个初始图像,所述图像分割任务用于提取所述目标检测区域对应的目标图像;
将各初始图像输入至预先训练的图像分割模型,获得所述图像分割模型输出的各初始图像对应的目标图像。
在本公开提供的实施例中,目标图像可以理解为目标检测区域对应的特写图像。而在实际应用中,接收到的多个图像可能会存在除了目标检测区域外,还包括了其他的区域,而其他的区域会对目标检测区域造成影响。因此,在本公开提供的方法中,还对图像做了分割处理。
具体的,首先获取针对多个初始图像的图像分割任务,初始图像中即包括有目标检测区域,又包括对目标检测区域带来影响的区域。该图像分割任务用于从各初始图像中提取出目标检测区域对应的目标图像。
将多个初始图像输入至预先训练的图像分割模型进行处理。图像分割模型被训练于识别到初始图像中的目标检测区域,并将目标检测区域从初始图像中截取出来,生成目标检测区域对应的目标图像。
以平扫CT图像为例,在本说明书提供的方法中,首先获取初始平扫CT图像,初始平扫CT图像为符合图像质量要求的平扫CT图像,初始平扫CT图像的来源可以是多个CT扫描机,也可以是同一个CT扫描机,在本说明书中对此不做限定。在获得初始平扫CT图像之后,要统一各初始平扫CT图像的格式。图像分割模型可以是3DUNet。
3DUNet是一种在三维空间中进行图像分割的深度学习架构,是U-Net模型的扩展版本,U-Net最初被设计用于二维生物医学图像的语义分割,并因其优异的表现和对小目标分割的高精度而受到广泛认可。3DUNet将这一思想应用于三维数据集,例如医学影像(如CT、MRI扫描等),这在许多医疗领域非常有用,因为这些领域的数据通常具有丰富的三维结构信息。
在3DUNet中包括有编码器-解码器结构,用于捕获全局上下文信息,以及恢复丢失的空间细节并生成精确的像素级分割标签。3DUNet结构中的3D卷积核在网络中替代了2D卷积核,可以同时处理输入数据在三个维度(长、宽、高)上的特征。同时,3DUNet还能有效地融合不同层次的三维特征,有助于提取复杂的形状和结构信息。该框架在医学图像分割方法有较好的效果。在本说明书提供的方法中,对初始CT图像使用3DUNet结构的预处理策略进行切割处理,获得目标检测区域对应的目标图像,用于后续的处理。
在本公开提供的方法中,以检测肝部区域为例进行解释说明。将初始平扫CT图像重采样到x,y,z三个方向的尺寸为0.7*0.7*5mm,通过3DUNet分割模型,在初始平扫CT图像上分割出肝脏及其周围若干个器官的mask,周围的器官包括胆囊、肝内血管、脾脏、胃、胰腺、右肾、下腔静脉等等。
再根据肝脏和脾脏的mask,确定各自对应的标记框,再在标记框的基础上向外扩充预设范围的像素距离,得到目标图像。获取肝脏和脾脏的区域,目的是将脾脏作为肝脏的参考器官,使得后续模型在处理图像时,不仅能学习到目标检测区域(肝脏)的特征,还可以学习到参考检测区域(脾脏)的特征,从而学习到肝脏的特征是需要考虑的,脾脏的特征是无需考虑的。
步骤104:将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各目标图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
在接收到图像处理任务后,可以从图像处理任务中获取其携带的多个目标图像,将多个目标图像输入至图像处理模型,可以获得图像处理模型输出的目标检测区域对应的检测结果,检测结果具体包括待检测对象的对象分类信息,目标检测区域上的病灶分割信息和病灶分类信息。
待检测对象的对象分类信息可以理解为在对象级别上的分类信息,其包括有是否存在异常对象,异常对象类型,以及异常对象对应各类别的概率。
病灶分割信息可以理解为在待检测图像上针对异常对象的mask信息,在本公开提供的方法中,病灶分割信息包括全局mask信息和局部mask信息。全局mask信息可以理解为目标检测区域上所有异常对象的mask信息,局部mask信息可以理解为针对每个类型的异常对象的mask信息。例如,在某个目标检测区域中包括有3个异常对象,则病灶分割信息包括4部分,分别为三个异常对象的全部mask信息,以及三个异常对象各自分别对应的mask信息。
病灶分类信息可以理解为在经过图像处理模型识别之后,如果存在病灶分割信息的情况下,仅一步确定病灶的类型。更进一步的,病灶分类信息包括病灶类别信息和器官类别信息,在本公开提供的肝脏肿瘤的检测中,病灶分为9个类别(分别为肝细胞肝癌、胆管细胞癌、转移瘤、血管瘤、局灶性结节增生、囊肿、钙化、其他恶性肿瘤、其他良性肿瘤),器官分为8个类别(分别为肝脏、胆囊、肝内血管、脾脏、胃、胰腺、右肾、下腔静脉)。
在本公开提供的方法中,图像处理模型接收到多个目标图像后,对多个目标图像进行处理,生成对象特征信息和病灶特征信息,其中,对象特征信息可以理解为用于生成对象分类信息的特征信息,病灶特征信息可以理解为用于生成病灶分类信息和病灶分割信息的特征信息,具体的,病灶特征信息包括病灶子特征信息和器官子特征信息,病灶子特征信息用于确定病灶类型,器官子特征信息用于确定器官类型。
在本公开提供的方法中,虽然可以用对象特征信息生成对象分类信息,但是为了提升预测准确率,将病灶特征信息与对象特征信息进行融合,使其预测能更加准确。基于此,可以将对象特征信息和病灶特征信息进行融合,获得对象融合特征信息。对象融合特征信息依然是用于预测对象分类信息的特征信息,其相比于对象特征信息而言,增加了病灶特征信息的相关内容,使得在后续的预测过程中,可以更加准确。
最后图像处理模型根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,从而组成最终的检测结果。
本公开提供的方法中,进一步解释说明了图像处理模型的模型结构,具体的,图像处理模型包括图像处理主干网络、对象分支网络、病灶分支网络和融合网络;
参见图2,图2示出了本说明书一实施例提供的图像处理模型的模型结构示意图,如图2所示,图像处理模型包括有图像处理主干网络、对象分支网络、病灶分支网络和融合网络。
将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,包括S1042-S1048:
S1042、将所述多个目标图像输入至所述图像处理主干网络,提取所述图像处理主干网络中的待处理对象特征信息和待处理病灶特征信息。
多个目标图像可以组成针对目标检测区域的3D检测区域,将多个目标图像输入至图像处理主干网络进行处理,可以提取图像处理主干网络在对多个目标图像处理的过程中生成的待处理对象特征信息和待处理病灶特征信息。
需要注意的是,待处理对象特征信息和待处理病灶特征信息均为图像处理模型在对多个目标图像进行处理的过程中生成的中间参数特征。依据各自的用于分为了待处理对象特征信息和待处理病灶特征信息,若其中的某个参数特征即用于生成对象特征信息,又用于生成病灶特征信息,则该参数特征即属于待处理对象特征信息,又属于待处理病灶特征信息。
具体的,所述图像处理主干网络包括图像编码器和图像解码器;
将所述多个目标图像输入至所述图像处理主干网络,提取所述图像处理主干网络中的待处理对象特征信息和待处理病灶特征信息,包括:
将所述多个目标图像输入至所述图像编码器,获得目标图像编码特征和至少一个多尺度图像编码特征;
将所述目标图像编码特征输入至所述图像解码器,获得至少一个第一多尺度图像解码特征、至少一个第二多尺度图像解码特征和目标图像解码特征;
将所述至少一个多尺度图像编码特征和至少一个第一多尺度图像解码特征确定为待处理对象特征信息,将所述至少一个第二多尺度图像解码特征和所述目标图像解码特征确定为待处理病灶特征信息。
在本公开提供的方法横纵,使用nnU-Net框架搭建图像处理模型,图像处理主干网络包括图像编码器和图像解码器,在本实施方式中,使用U-Net的Pixel Encoder作为图像编码器,使用特征金字塔网络作为图像解码器。
将多个目标图像输入至图像处理模型处理的过程中,先经过图像编码器对多个目标图像进行编码,获得目标图像编码特征和至少一个多尺度图像编码特征。目标图像编码特征为图像编码器输出的编码特征,至少一个多尺度图像编码特征为图像编码器在对多个目标图像处理过程中生成的编码参数特征。
将目标图像编码特征输入至图像解码器,获得图像解码器在解码过程中生成的至少一个第一多尺度图像解码特征、至少一个第二多尺度图像解码特征和最终的目标图像解码特征。
此时,将多尺度图像编码特征和第一多尺度图像解码特征作为待处理对象特征信息,将第二多尺度图像解码特征和目标图像解码特征作为待处理病灶特征信息。
其中,所述图像编码器包括多个顺次连接的图像编码层;
将所述多个目标图像输入至所述图像编码器,获得目标图像编码特征和至少一个多尺度图像编码特征,包括:
将所述多个目标图像输入至所述图像编码器,获得所述图像编码器输出的目标图像编码特征;
在所述多个图像编码层中确定至少一个目标图像编码层,获得各目标图像编码层输出的多尺度图像编码特征。
需要注意的是,目标图像编码层并不特指某一个图像编码层,而是指需要提取图像编码特征的图像编码层,例如,图像编码层一共有6个,其中,需要用到第1、3、5个图像编码层输出的多尺度图像编码特征,则第1、3、5个图像编码层为目标图像编码层。又例如,图像编码层一共有12个,需要用到第9、10、11、12个图像编码层输出的多尺度图像编码特征,则第9、10、11、12个图像编码层为目标图像编码层。
参见图2,以图2所示的图像处理主干网络为例进行解释说明,在图像编码器中包括有6个图像编码层。将多个目标图像输入至图像编码器处理,第5、6个图像编码层为目标图像编码成,则获得图像编码器输出的目标图像编码特征,和第5、第6个图像编码层输出的多尺度图像编码特征和
在本公开提供的另一具体实施方式中,所述图像解码器包括多个顺次连接的图像解码层;
将所述目标图像编码特征输入至所述图像解码器,获得至少一个第一多尺度图像解码特征和至少一个第二多尺度图像解码特征,包括:
在所述多个图像解码层中确定至少一个第一图像解码层和至少一个第二图像解码层;
获得第一图像解码层输出的第一多尺度图像解码特征,获得第二图像解码层输出的第二多尺度图像解码特征。
如上述的目标图像编码层类似,图像解码层中的第一图像解码层和第二图像解码层业并不特指某个图像解码层,而是指用于确定第一多尺度图像解码特征和第二多尺度图像解码特征的图像解码层。
参见图2,以图2所示的图像处理主干网络为例进行解释说明,在图像解码器中包括有6个图像解码层,将目标图像编码特征输入至图像解码器,确定,第5、6个图像解码层的输出特征为第一多尺度图像解码特征,第2、3、4个图像解码层的输出特征为第二多尺度图像解码特征,则获得第5、第6个图像解码层输出的第一多尺度图像解码特征和和第2、3、4个图像解码层输出的第二多尺度图像解码特征和以及最终图像解码器输出的目标图像解码特征
以图2所示为例,多尺度图像编码特征和第一多尺度图像解码特征和组成待处理对象特征信息,第二多尺度图像解码特征和和目标图像解码特征组成待处理病灶特征信息。
S1044、将所述待处理对象特征信息输入至所述对象分支网络,获得对象特征信息,将所述待处理病灶特征信息输入至所述病灶分支网络,获得病灶特征信息和待分割图像特征信息。
对象分支网络用于生成针对待处理对象的对象分类信息,在上述步骤中获得待处理对象特征信息后,将待处理对象特征信息输入至对象分支网络,在对象分支网络中进行处理,获得对象特征信息。
参见图3,图3示出了本说明书一实施例提供的对象分支网络的数据处理示意图。如图3所示,将待处理对象特征信息迭代地输入到4层Dual-path Transformer Block(DPB)中进行处理,得到对象特征信息。Dual-path Transformer Block是一种特殊的Transformer结构,它通过将处理过程分解为两个并行路径(通常称为局部路径和全局路径)来优化性能并降低计算复杂度,从而提升处理效率和处理效果。这两个路径通常关注不同类型的特征,局部路径更关注细粒度的局部特征,全局路径更关注场景级的全局特征。
病灶分支网络用于生成针对病灶的分割信息和分类信息,在上述步骤中获得处理病灶特征信息后,将处理病灶特征信息输入至病灶分支网络中,在病灶分支网络中进行处理,获得病灶特征信息和待分割图像特征信息。
具体的,所述待处理病灶特征信息包括至少一个第二多尺度图像解码特征和目标图像解码特征;
将所述待处理病灶特征信息输入至所述病灶分支网络,获得病灶特征信息和待分割图像特征信息,包括:
对各第二多尺度图像解码特征进行特征解码,获得病灶特征信息;
融合所述病灶特征信息和所述目标图像解码特征,生成待分割图像特征信息。
在病灶分支网络中,会同时生成病灶特征信息和待分割图像特征信息,病灶特征信息用于确定病灶分类信息,待分割图像特征信息用于确定病灶分割信息。
参见图4,图4示出了本说明书一实施例提供的病灶分支网络的数据处理示意图。如图4所示,在病灶分支网络中,包括一个Transformer Decoder结构,输入随机初始化的50个Query,在Transformer Decoder与第二多尺度图像解码特征进行解码处理。如图4所示,将第二多尺度图像解码特征和和50个Query输入到Transformer Decoder中处理,获得Transformer Decoder输出的病灶特征信息。同时,为了更好的对异常对象进行像素级的分割,将病灶特征信息和图像解码器输出的目标图像解码特征进行融合,获得待分割图像特征信息,用于对待检测区域中的异常对象进行像素级别的Mask。
至此,获得了对象分支网络中生成的对象特征信息,同时获得了病灶分支网络中生成的病灶特征信息和待分割图像特征信息。
S1046、将所述对象特征信息和所述病灶特征信息输入至所述融合网络,生成对象融合特征信息。
融合网络可以理解为用于将对象特征信息和病灶特征信息融合到一起的网络,生成对象融合特征信息。对象融合特征信息使得图像处理模型可以融合对象的全局信息和病灶的局部信息,可以提升较小异常对象的检测能力。
在本公开提供的一具体实施方式中,将所述对象特征信息和所述病灶特征信息输入至所述融合网络,生成对象融合特征信息,包括:
将所述对象特征信息和所述病灶特征信息输入至所述融合网络,在所述融合网络中提取所述病灶特征信息中的病灶子特征信息,并拼接所述对象特征信息和所述病灶子特征信息,生成对象融合特征信息。
在实际应用中,病灶特征信息中包括了病灶子特征信息和器官子特征信息,为了使得对象特征中可以参考到病灶特征,可以将病灶子特征信息与对象特征信息进行融合,生成对象融合特征信息。
参见图5,图5示出了本说明书一实施例提供的融合网络的数据处理示意图。如图5所示,在病灶特征信息中提取出用于确定病灶类型的病灶子特征信息,将病灶子特征信息与对象特征信息进行拼接后,获得对象融合特征信息。
对象融合特征信息是用于最终确定对象分类信息的特征信息,其即包括了对象特征信息,有包括了病灶子特征信息,使得最终的预测结果更加准确。
S1048、根据所述病灶特征信息生成病灶分类信息,根据所述待分割图像特征信息生成病灶分割信息,根据所述对象融合特征信息生成对象分类信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
在经过上述步骤处理后,即可获得了病灶分类信息、待分割图像特征信息和对象融合特征信息。之后即可分别根据这三个信息生成对应的病灶分类信息、病灶分割信息和对象分类信息。最后根据病灶分类信息、病灶分割信息和对象分类信息组成最终的检测结果。
在实际应用中,分别用三个不同的分类器对不同的信息进行处理。具体的,在本说明书提供的一具体实施方式中,根据所述病灶特征信息生成病灶分类信息,根据所述待分割图像特征信息生成病灶分割信息,包括:
将所述病灶特征信息输入至病灶分类器,生成病灶分类信息;
将所述待分割图像特征信息输入至分割器,生成病灶分割信息。
根据所述对象融合特征信息生成对象分类信息,包括:
将所述对象融合特征信息输入至对象分类器,生成对象分类信息。
在实际应用中,将病灶特征信息输入至病灶分类器中进行处理,可以获得病灶分类器输出的病灶分类信息;将待分割图像特征信息输入至分割器处理,可以获得分割器输出的病灶分割信息;将对象融合特征信息输入至对象分类器处理,可以获得对象分类器输出的对象分类信息。
通过本公开提供的图像处理方法,将多个目标图像输入至图像处理模型,在图像处理模型中,基于各目标图像生成对象特征信息和病灶特征信息,同时将对象特征信息和病灶特征信息融合,生成对象融合特征信息,使得模型在预测最终的对象分类信息时,除了使用对象特征信息,还会参考病灶特征信息,对全局信息和局部信息进行融合,提升小病灶的检出能力。
本公开提供的图像处理模型是预先训练好的图像处理模型,参见图6,图6示出了本说明书一实施例提供的图像处理模型训练方法的流程图,如图6所示,图像处理模型通过下述步骤训练生成:
步骤602:获取多个样本目标图像和多个样本目标图像对应的样本检测结果,其中,所述样本检测结果包括样本对象分类信息、样本病灶分类信息和样本病灶分割信息。
具体的,本说明书提供的图像处理模型的训练方法使用的时有监督训练,其包括有训练样本对,训练样本对具体包括针对目标检测区域的多个样本目标图像,多个样本目标图像对应的样本检测结果。其中,多个样本目标图像可以组成针对目标检测区域的三维图像。样本检测结果具体包括了样本对象分类信息,样本病灶分类信息和样本病灶分割信息。
样本对象分类信息包括该检测对象的目标检测区域是否存在异常对象,异常对象类型,异常对象对应各类别的概率。样本病灶分类信息可以理解为针对目标检测区域对应病灶的病灶分类,和病灶所在器官的器官分类。样本病灶分割信息可以理解为目标检测区域中针对异常对象的mask信息。
在本公开提供的方法中,通过多个样本目标图像和多个样本目标图像对应的样本检测结果来对图像处理模型进行有监督训练。
在本说明书提供的一具体实施方式中,获取多个样本目标图像和多个样本目标图像对应的样本检测结果,包括:
确定样本检测对象,并获取样本检测对象对应的样本增强图像和样本目标图像,其中,样本检测对象对应的样本增强图像和样本目标图像对应样本检测对象的同一目标检测区域;
在所述样本增强图像上针对目标检测区域生成参考样本检测结果;
基于所述参考样本检测结果在所述样本目标图像上标注样本检测结果。
其中,样本增强图像具体是指针对目标检测区域的对比度经过增强的图像,样本目标图像具体是指针对目标检测区域的普通图像。样本增强图像和样本目标图像对应的是同一个样本检测对象的同一个目标检测区域。
例如,以样本检测对象为张三,目标检测区域为肝脏为例,在确定了样本检测对象之后,基于该样本检测对象的目标检测区域获取到其对应的样本增强图像和样本目标图像。
在实际应用中,样本增强图像是针对目标检测区域的对比度经过增强的图像,可以更好的在样本增强图像中获取到针对目标检测区域的检测结果,在本公开中,针对样本增强图像的目标检测区域的检测结果成为参考样本检测结果。再通过配准的方式将参考样本检测结果对应到样本目标图像中,获得样本目标图像对应的样本检测结果。
在样本增强图像上针对目标检测区域生成参考样本检测结果的过程通常是通过人工标注的形式,而人工标注需要经验丰富的技术人员进行手动标注,费事费力,成本较高。在本公开提供的一可选实施方式中,可以由技术人员标注部分样本增强图像。形成样本增强图像和参考样本检测结果的训练样本对,再通过该训练样本对训练一个样本增强图像标注模型,在样本增强图像标注模型训练完成后,将未经过人工标注的样本增强图像输入至该样本增强图像标注模型中进行标注,进而生成样本增强图像的参考样本检测结果。
例如,在辅助医疗场景中,在经过用户许可的情况下,从影像科获取到图像质量符合要求的、成对的增强CT图像和平扫CT图像。成对的增强CT图像和平扫CT图像为同一个样本检测对象在同一个时间点拍摄的同一个角度的CT图像。增强CT图像包括动脉期增强CT、静脉期增强CT、延迟期增强CT等。
邀请经验丰富的影像科医生和专家,对部分增强CT图像上标注出针对目标检测区域中异常对象的mask信息,并参考样本检测对象的病理金标准给出异常对象的病灶分类信息和对象分类信息。mask信息、病灶分类信息和对象分类信息组成了参考样本检测结果。
根据标注的增强CT图像和参考样本检测结果训练一个增强CT模型,使得增强CT模型具备根据增强CT图像进行检测结果预测标注的能力。将未标注的增强CT图像输入至该增强CT模型中进行处理,通过增强CT模型对未标注的增强CT图像进行处理,生成未标注增强CT图像对应的参考样本检测结果。
在前述步骤中,由于增强CT图像和平扫CT图像是通过配准的,因此,增强CT图像上的参考样本检测结果可以平移到平扫CT图像中,作为平扫CT图像的样本检测结果。
经过上述处理,可以获得样本目标图像和样本图像对应的样本检测结果,需要注意的是,其中,样本目标图像和样本图像对应的样本检测结果中,第一部分是通过人工标注的样本增强图像和样本参考检测结果配准获得的,第二部分是通过模型标注的参考检测结果配准获得的。在后续的处理中,将样本目标图像和样本图像对应的样本检测结果划分为训练集和测试集的过程中,为了保证训练数据的准确性,在第一部分人工标注的数据中选取测试集,用于后续对模型性能的测试。
步骤604:将所述多个样本目标图像输入至图像处理模型,获得所述图像处理模型输出的预测检测结果和训练对比损失值。
在获得了样本目标图像和样本目标图像对应的样本检测结果后,即可对图像处理模型进行模型训练。此时的图像处理模型为还未被训练好的图像处理模型,需要通过样本目标图像和样本目标图像对应的样本检测结果对其进行模型训练。
在训练过程中,图像处理模型会输出基于样本目标图像生成的预测检测结果和训练比对损失值。在本说明书提供的具体实施方式中,引入了对比损失值的概念,即构建2阶段类别平衡的对比损失函数,分别在图像处理模型的对象分支网络和病灶分支网络中执行比对损失函数的处理,同时获得在训练过程中的训练比对损失值。
在本公开提供的一具体实施方式中,所述图像处理模型包括图像处理主干网络、对象分支网络、病灶分支网络和融合网络;
将所述多个样本目标图像输入至图像处理模型,获得所述图像处理模型输出的预测检测结果和训练对比损失值,包括:
将所述多个样本目标图像输入至图像处理主干网络,提取所述图像处理主干网络中的样本待处理对象特征信息和样本待处理病灶特征信息;
将所述样本待处理对象特征信息输入至所述对象分支网络,获得预测对象特征信息,将所述样本待处理病灶特征信息输入至所述病灶分支网络,获得预测病灶特征信息和预测待分割图像特征信息;
在所述对象分支网络中,根据样本目标图像对应的样本对象分类信息获取对象比对损失值,在所述病灶分支网络中,根据样本目标图像对应的样本病灶分类信息获取病灶比对损失值;
将所述预测对象特征信息和所述预测病灶特征信息输入至所述融合网络,生成样本对象融合特征信息;
根据所述预测对象特征信息生成第一对象分类信息,根据所述样本对象融合特征信息生成第二对象分类信息,根据所述预测病灶特征信息生成预测病灶分类信息,根据所述预测待分割图像特征信息生成预测病灶分割信息;
根据所述第一对象分类信息、所述第二对象分类信息、所述预测病灶分类信息和所述预测病灶分割信息生成预测检测结果,根据所述对象比对损失值和所述病灶比对损失值生成训练对比损失值。
在模型训练阶段,图像处理模型的模型结构与上述模型应用阶段的模型结构相同,关于图像处理模型的模型结构和数据处理流程的相关内容,参见上述模型应用阶段的相关部分,在此不再赘述。
图像处理模型在模型训练阶段,先在图像处理主干网络中处理多个样本目标图像。提取图像处理主干网络中生成的样本待处理对象特征信息和样本待处理病灶特征信息。将样本待处理对象特征信息输入到对象分支网络,获得预测对象特征信息;将样本待处理病灶特征信息输入至病灶分支网络,获得预测病灶特征信息和预测待分割图像特征信息。模型训练阶段中图像处理主干网络、对象分支网络和病灶分支网络的数据处理过程与上述应用阶段相同,在此不再赘述。
需要注意的是,在模型训练阶段,引入了2阶段类别平衡的对比损失函数,在对象分支网络中,在得到样本待处理对象特征信息的基础上,引入该待检测对象是否有异常对象得到2类的对比学习损失函数。在病灶分支网络中,在得到样本待处理病灶特征信息的基础上,引入几种异常对象的比对学习损失函数。通过比对学习的方式,使得图像处理模型学习到各训练批次之间样本待处理对象特征信息之间的差异,和各训练批次之间样本待处理病灶特征信息之间的差异。通过对比损失函数,可以获得对象比对损失值和病灶比对损失值。
在本公开提供的另一具体实施方式中,本公开提供的图像处理模型的训练过程中,如果batch size过小,会容易造成模型坍塌的问题,在此基础上,本公开的训练方法使用了参考历史样本对象特征信息的方式。具体的,将所述样本待处理对象特征信息输入至所述对象分支网络,获得预测对象特征信息,包括:
获取历史样本对象特征信息;
将所述历史对象特征信息和所述样本待处理对象特征信息输入至所述对象分支网络,获得预测对象特征信息;
相应的,将所述样本待处理病灶特征信息输入至所述病灶分支网络,获得预测病灶特征信息和预测待分割图像特征信息,包括:
获取历史样本病灶特征信息;
将所述历史样本病灶特征信息和所述样本待处理病灶特征信息输入至所述病灶分支网络,获得预测病灶特征信息和预测待分割图像特征信息。
在本实施方式中,引入了memory bank的策略,在对象分支网络和病灶分支网络中分别加入了对象特征存储块和病灶特征存储块。其中,对象特征存储块用于存储历史样本对象特征信息,病灶特征存储块用于存储历史样本病灶特征信息。两个特征存储块的容量相同,同时考虑的分类问题中由于数据类型的不同,会有某些信息分类的长尾问题,即某些分类信息的数据量少的问题。为了解决分类的长尾问题,在两个特征存储块中存储的各个类别的特征数量是均衡的,在更新各特征存储块的过程中,也要考虑各个类别已经存储的特征的数量,保证各类别的特征数量保持均衡。
除了引入了对比损失函数之外,本公开提供的训练方法中,还会根据预测对象特征信息生成第一对象分类信息,同时根据样本对象融合特征信息生成第二对象分类信息,即在模型训练过程中,会生成两个对象分类信息,其中一个是基于预测对象特征信息生成的第一对象分类信息,另外一个是基于预测对象特征信息和预测病灶特征信息生成第二对象分类信息,两者在模型训练阶段都会作为模型的预测结果,用于后续计算模型损失值。
同时,与上述应用阶段相同的,在模型训练阶段,还会根据预测病灶特征信息生成预测病灶分类信息,根据预测待分割图像特征信息生成预测病灶分割信息。
最终,确定第一对象分类信息、第二对象分类信息、预测病灶分类信息和所述预测病灶分割信息生成预测检测结果。确定对象比对损失值和病灶比对损失值为训练对比损失值
步骤606:根据所述预测检测结果和所述样本检测结果计算预测损失值。
在获得预测检测结果后,根据预测检测结果和样本检测结果进行比对,计算预测损失值。此时,图像处理模型是还未训练好的模型,通过计算预测损失值的方式,确定预测结果和样本结果之间的差异,从而进一步对图像处理模型的参数进行调整,实现对图像处理模型的训练。
具体的,在本说明书提供的一具体实施方式中,根据所述预测检测结果和所述样本检测结果计算预测损失值,包括:
根据所述第一对象分类信息和所述样本对象分类信息计算第一对象损失值;
根据所述第二对象分类信息和所述样本对象分类信息计算第二对象损失值;
根据所述预测病灶分类信息和所述样本病灶分类信息计算病灶分类损失值;
根据所述预测病灶分割信息和所述样本病灶分割信息计算病灶分割损失值。
在实际应用中,预测检测结果中包括了第一对象分类信息和第二对象分类信息,两者均与样本对象分类信息计算损失值,具体的,由第一对象分类信息和样本对象分类信息计算第一对象损失值,由第二对象分类信息和样本对象分类信息计算第二对象损失值。同样的,根据预测病灶分类信息和样本病灶分类信息计算病灶分类损失值,根据预测病灶分割信息和预测病灶分类信息计算病灶分割损失值。计算模型损失值的方法有很多,例如交叉熵损失函数、最大损失函数、平均值损失函数等等,在本说明书中,对损失函数的具体方式不做限定,以实际应用为准。
步骤608:根据所述预测损失值和所述训练对比损失值调整图像处理模型的模型参数,并继续训练所述图像处理模型,直至达到模型训练停止条件。
在获得预测损失值和训练比对损失值之后,即可根据这两类损失值反向传播,调整图像处理模型的模型参数。
具体的,在本公开提供的一具体实施方式中,根据所述预测损失值和所述训练对比损失值调整图像处理模型的模型参数,包括:
根据所述第一对象损失值、所述第二对象损失值、所述病灶分类损失值、所述病灶分割损失值、所述对象比对损失值和所述病灶比对损失值计算模型损失值;
根据所述模型损失值调整图像处理模型的模型参数。
在实际应用中,是将预测损失值和训练比对损失值进行融合后获得新的损失值,更进一步的,可以根据预设的损失权重,将第一对象损失值、第二对象损失值、病灶分类损失值、病灶分割损失值、对象比对损失值和病灶比对损失值进行融合,获得新的模型损失值,再进一步通过模型损失值调整图像处理模型的模型参数,直至达到模型训练停止条件,获得训练好的图像处理模型。
本公开提供的图像处理模型的训练方法,提出了使用样本增强图像通过部分手动标注的形式,训练一个样本增强图像标注模型,使得样本增强图像标注模型可以对未进行人工标注的样本增强图像进行标注,节省了标注时间。另外,通过样本增强图像和样本目标图像配准的方式,将样本增强图像上的参考样本检测结果迁移到样本目标图像中,获得样本目标图像的样本检测结果,解决了样本目标图像上标注困难、标注缺失的问题,降低了标注成本。
另外,在模型训练过程中,引入了两阶段类别平衡的对比学习损失函数,增强了图像处理模型的鉴别诊断能力,在对象分支网络和病灶分支网络中都能提升分类的精度,尤其是对于某些特征数据较少的长尾类别,具有更好的分类识别精度。
下述结合附图7,以本说明书提供的图像处理方法在CT图像的处理应用为例,对图像处理方法进行进一步说明。其中,图7示出了本说明书一个实施例提供的一种CT图像处理方法的处理过程流程图,具体包括以下步骤。
步骤702:接收CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在异常对象。
步骤704:将所述多个平扫CT图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
需要说明的是,步骤702至步骤704的实现方式,与上述步骤102至步骤104的实现方式相同,本公开便不再赘述。
具体的,在本实施例提供的方法中,以CT图像处理任务为目标检测区域内是否存在异常对象进行检测为例进行进一步解释说明,基于此,CT图像处理任务包括目标检测区域的平扫CT图喜爱那个,将平扫CT图像输入至图像处理模型中进行识别。图像处理模型可以识别出平扫CT图像中的是否存在异常对象,以及如果存在异常对象时异常对象的Mask信息,同时还可以预测出待检测对象的对象分类信息,即该待检测对象的目标检测区域中是否存在异常对象,异常对象的类型等信息。通过本方法,解决了目前应用中无法根据平扫CT图像在对目标检测区域进行检测时,检测效果差,准确率较低的问题。
参见图8,图8示出了本说明书一实施例提供的一种癌症的计算机辅助诊断方法的流程示意图,具体包括:
步骤802:接收CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在肿瘤。
步骤804:将所述多个平扫CT图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
本实施例提供的癌症的计算机辅助诊断方法,可以适用于筛查各目标检测区域内是否存在肿瘤,目标检测区域包括但不限于胰腺、食道、肝脏、肺、乳腺、肠、胃、全身淋巴结等全身各器官中是否存在肿瘤的场景。在检测结果中包括了待检测用户的用户分类信息(即是否存在肿瘤,肿瘤是否为恶性肿瘤等)、病灶分类信息(肿瘤属于什么类型的肿瘤)和病灶分割信息(在CT图像上标记出肿瘤的mask信息)。通过本公开提供的方法,可以为医生提供指导性建议,辅助提升医生的诊断准确率,为医生给出诊断结果提供数据支持。
参见图9,图9示出了本说明书一实施例提供的一种肝癌的计算机辅助诊断方法的流程示意图,具体包括:
步骤902:接收CT图像处理任务,其中,所述CT图像处理任务携带肝脏区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述肝脏区域内是否存在肿瘤。
步骤904:将所述多个平扫CT图像输入至图像处理模型,获得所述肝脏区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
本实施例提供的肝癌的计算机辅助诊断方法,可以使用于筛查肝细胞肝癌、胆管细胞癌、转移瘤、血管瘤、局灶性结节增生、囊肿、钙化等疾病,为医生提供指导性建议,辅助提升医生的诊断准确率,为医生给出诊断结果提供数据支持。
参见图10,图10示出了本说明书一个实施例提供的一种癌症的计算机辅助诊断系统的架构图,癌症的计算机辅助诊断系统可以包括客户端100和服务端200;
客户端100,用于向服务端200发送CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在肿瘤;
服务端200,用于将所述多个平扫CT图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果;向客户端100发送目标检测区域对应的检测结果;
客户端100,还用于接收服务端200发送的检测结果。
癌症的计算机辅助诊断系统可以包括多个客户端100以及服务端200,其中,客户端100可以称为端侧设备,服务端200可以称为云侧设备。多个客户端100之间通过服务端200可以建立通信连接,在癌症的计算机辅助诊断场景中,服务端200即用来在多个客户端100之间提供癌症的计算机辅助诊断服务,多个客户端100可以分别作为发送端或接收端,通过服务端200实现通信。
用户通过客户端100可与服务端200进行交互以接收其它客户端100发送的数据,或将数据发送至其它客户端100等。在癌症的计算机辅助诊断场景中,可以是用户通过客户端100向服务端200发布数据流,服务端200根据该数据流生成检测结果,并将检测结果推送至其他建立通信的客户端中。
其中,客户端100与服务端200之间通过网络建立连接。网络为客户端100与服务端200之间提供了通信链路的介质。网络可以包括各种连接类型,例如有线、无线通信链路或者光纤电缆等等。客户端100所传输的数据可能需要经过编码、转码、压缩等处理之后才发布至服务端200。
客户端100可以为浏览器、APP(Application,应用程序)、或网页应用如H5(HyperText Markup Language5,超文本标记语言第5版)应用、或轻应用(也被称为小程序,一种轻量级应用程序)或云应用等,客户端100可以基于服务端200提供的相应服务的软件开发工具包(SDK,Software Development Kit),如基于实时通信(RTC,Real Time Communication)SDK开发获得等。客户端100可以部署在电子设备中,需要依赖设备运行或者设备中的某些APP而运行等。电子设备例如可以具有显示屏并支持信息浏览等,如可以是个人移动终端如手机、平板电脑、个人计算机等。在电子设备中通常还可以配置各种其它类应用,例如人机对话类应用、模型训练类应用、文本处理类应用、网页浏览器应用、购物类应用、搜索类应用、即时通信工具、邮箱客户端、社交平台软件等。
服务端200可以包括提供各种服务的服务器,例如为多个客户端提供通信服务的服务器,又如为客户端上使用的模型提供支持的用于后台训练的服务器,又如对客户端发送的数据进行处理的服务器等。需要说明的是,服务端200可以实现成多个服务器组成的分布式服务器集群,也可以实现成单个服务器。服务器也可以为分布式系统的服务器,或者是结合了区块链的服务器。服务器也可以是云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络(CDN,Content Delivery Network)以及大数据和人工智能平台等基础云计算服务的云服务器,或者是带人工智能技术的智能云计算服务器或智能云主机。
值得说明的是,本公开中提供的癌症的计算机辅助诊断方法一般由服务端执行,但是,在本说明书的其它实施例中,客户端也可以与服务端具有相似的功能,从而执行本公开所提供的癌症的计算机辅助诊断方法。在其它实施例中,本公开所提供的癌症的计算机辅助诊断方法还可以是由客户端与服务端共同执行。
与上述方法实施例相对应,本说明书还提供了图像处理装置实施例,图11示出了本说明书一个实施例提供的一种图像处理装置的结构示意图。如图11所示,该装置包括:
接收模块1102,被配置为接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象。
检测模块1104,被配置为将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各目标图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
可选的,所述图像处理模型包括图像处理主干网络、对象分支网络、病灶分支网络和融合网络;
所述检测模块1104,进一步被配置为:
将所述多个目标图像输入至所述图像处理主干网络,提取所述图像处理主干网络中的待处理对象特征信息和待处理病灶特征信息;
将所述待处理对象特征信息输入至所述对象分支网络,获得对象特征信息,将所述待处理病灶特征信息输入至所述病灶分支网络,获得病灶特征信息和待分割图像特征信息;
将所述对象特征信息和所述病灶特征信息输入至所述融合网络,生成对象融合特征信息;
根据所述病灶特征信息生成病灶分类信息,根据所述待分割图像特征信息生成病灶分割信息,根据所述对象融合特征信息生成对象分类信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
可选的,所述图像处理主干网络包括图像编码器和图像解码器;
所述检测模块1104,进一步被配置为:
将所述多个目标图像输入至所述图像编码器,获得目标图像编码特征和至少一个多尺度图像编码特征;
将所述目标图像编码特征输入至所述图像解码器,获得至少一个第一多尺度图像解码特征、至少一个第二多尺度图像解码特征和目标图像解码特征;
将所述至少一个多尺度图像编码特征和至少一个第一多尺度图像解码特征确定为待处理对象特征信息,将所述至少一个第二多尺度图像解码特征和所述目标图像解码特征确定为待处理病灶特征信息。
可选的,所述图像编码器包括多个顺次连接的图像编码层;
所述检测模块1104,进一步被配置为:
将所述多个目标图像输入至所述图像编码器,获得所述图像编码器输出的目标图像编码特征;
在所述多个图像编码层中确定至少一个目标图像编码层,获得各目标图像编码层输出的多尺度图像编码特征。
可选的,所述图像解码器包括多个顺次连接的图像解码层;
所述检测模块1104,进一步被配置为:
在所述多个图像解码层中确定至少一个第一图像解码层和至少一个第二图像解码层;
获得第一图像解码层输出的第一多尺度图像解码特征,获得第二图像解码层输出的第二多尺度图像解码特征。
可选的,所述待处理病灶特征信息包括至少一个第二多尺度图像解码特征和目标图像解码特征;
所述检测模块1104,进一步被配置为:
对各第二多尺度图像解码特征进行特征解码,获得病灶特征信息;
融合所述病灶特征信息和所述目标图像解码特征,生成待分割图像特征信息。
可选的,所述检测模块1104,进一步被配置为:
将所述病灶特征信息输入至病灶分类器,生成病灶分类信息;
将所述待分割图像特征信息输入至分割器,生成病灶分割信息。
可选的,所述检测模块1104,进一步被配置为:
将所述对象特征信息和所述病灶特征信息输入至所述融合网络,在所述融合网络中提取所述病灶特征信息中的病灶子特征信息,并拼接所述对象特征信息和所述病灶子特征信息,生成对象融合特征信息。
可选的,所述检测模块1104,进一步被配置为:
将所述对象融合特征信息输入至对象分类器,生成对象分类信息。
可选的,所述装置还包括训练模块,被配置为:
获取多个样本目标图像和多个样本目标图像对应的样本检测结果,其中,所述样本检测结果包括样本对象分类信息、样本病灶分类信息和样本病灶分割信息;
将所述多个样本目标图像输入至图像处理模型,获得所述图像处理模型输出的预测检测结果和训练对比损失值;
根据所述预测检测结果和所述样本检测结果计算预测损失值;
根据所述预测损失值和所述训练对比损失值调整图像处理模型的模型参数,并继续训练所述图像处理模型,直至达到模型训练停止条件。
可选的,所述训练模块,进一步被配置为:
确定样本检测对象,并获取样本检测对象对应的样本增强图像和样本目标图像,其中,样本检测对象对应的样本增强图像和样本目标图像对应样本检测对象的同一目标检测区域;
在所述样本增强图像上针对目标检测区域生成参考样本检测结果;
基于所述参考样本检测结果在所述样本目标图像上标注样本检测结果。
可选的,所述图像处理模型包括图像处理主干网络、对象分支网络、病灶分支网络和融合网络;
所述训练模块,进一步被配置为:
将所述多个样本目标图像输入至图像处理模型,获得所述图像处理模型输出的预测检测结果和训练对比损失值,包括:
将所述多个样本目标图像输入至图像处理主干网络,提取所述图像处理主干网络中的样本待处理对象特征信息和样本待处理病灶特征信息;
将所述样本待处理对象特征信息输入至所述对象分支网络,获得预测对象特征信息,将所述样本待处理病灶特征信息输入至所述病灶分支网络,获得预测病灶特征信息和预测待分割图像特征信息;
在所述对象分支网络中,根据样本目标图像对应的样本对象分类信息获取对象比对损失值,在所述病灶分支网络中,根据样本目标图像对应的样本病灶分类信息获取病灶比对损失值;
将所述预测对象特征信息和所述预测病灶特征信息输入至所述融合网络,生成样本对象融合特征信息;
根据所述预测对象特征信息生成第一对象分类信息,根据所述样本对象融合特征信息生成第二对象分类信息,根据所述预测病灶特征信息生成预测病灶分类信息,根据所述预测待分割图像特征信息生成预测病灶分割信息;
根据所述第一对象分类信息、所述第二对象分类信息、所述预测病灶分类信息和所述预测病灶分割信息生成预测检测结果,根据所述对象比对损失值和所述病灶比对损失值生成训练对比损失值。
可选的,所述训练模块,进一步被配置为:
根据所述第一对象分类信息和所述样本对象分类信息计算第一对象损失值;
根据所述第二对象分类信息和所述样本对象分类信息计算第二对象损失值;
根据所述预测病灶分类信息和所述样本病灶分类信息计算病灶分类损失值;
根据所述预测病灶分割信息和所述样本病灶分割信息计算病灶分割损失值。
可选的,所述训练模块,进一步被配置为:
获取历史样本对象特征信息;
将所述历史对象特征信息和所述样本待处理对象特征信息输入至所述对象分支网络,获得预测对象特征信息;
获取历史样本病灶特征信息;
将所述历史样本病灶特征信息和所述样本待处理病灶特征信息输入至所述病灶分支网络,获得预测病灶特征信息和预测待分割图像特征信息。
可选的,所述训练模块,进一步被配置为:
根据所述第一对象损失值、所述第二对象损失值、所述病灶分类损失值、所述病灶分割损失值、所述对象比对损失值和所述病灶比对损失值计算模型损失值;
根据所述模型损失值调整图像处理模型的模型参数。
通过本公开提供的图像处理装置,将多个目标图像输入至图像处理模型,在图像处理模型中,基于各目标图像生成对象特征信息和病灶特征信息,同时将对象特征信息和病灶特征信息融合,生成对象融合特征信息,使得模型在预测最终的对象分类信息时,除了使用对象特征信息,还会参考病灶特征信息,对全局信息和局部信息进行融合,提升小病灶的检出能力。
另外,本公开提供的训练模块,提出了使用样本增强图像通过部分手动标注的形式,训练一个样本增强图像标注模型,使得样本增强图像标注模型可以对未进行人工标注的样本增强图像进行标注,节省了标注时间。另外,通过样本增强图像和样本目标图像配准的方式,将样本增强图像上的参考样本检测结果迁移到样本目标图像中,获得样本目标图像的样本检测结果,解决了样本目标图像上标注困难、标注缺失的问题,降低了标注成本。
同时,在模型训练过程中,引入了两阶段类别平衡的对比学习损失函数,增强了图像处理模型的鉴别诊断能力,在对象分支网络和病灶分支网络中都能提升分类的精度,尤其是对于某些特征数据较少的长尾类别,具有更好的分类识别精度。
上述为本实施例的一种图像处理装置的示意性方案。需要说明的是,该图像处理装置的技术方案与上述的图像处理方法的技术方案属于同一构思,图像处理装置的技术方案未详细描述的细节内容,均可以参见上述图像处理方法的技术方案的描述。
图12示出了根据本说明书一个实施例提供的一种计算设备1200的结构框图。该计算设备1200的部件包括但不限于存储器1210和处理器1220。处理器1220与存储器1210通过总线1230相连接,数据库1250用于保存数据。
计算设备1200还包括接入设备1240,接入设备1240使得计算设备1200能够经由一个或多个网络1260通信。这些网络的示例包括公用交换电话网(PSTN,Public Switched Telephone Network)、局域网(LAN,Local Area Network)、广域网(WAN,Wide Area Network)、个域网(PAN,Personal Area Network)或诸如因特网的通信网络的组合。接入设备1240可以包括有线或无线的任何类型的网络接口(例如,网络接口卡(NIC,network interface controller))中的一个或多个,诸如IEEE802.11无线局域网(WLAN,Wireless Local Area Network)无线接口、全球微波互联接入(Wi-MAX,Worldwide Interoperability for Microwave Access)接口、以太网接口、通用串行总线(USB,Universal Serial Bus)接口、蜂窝网络接口、蓝牙接口、近场通信(NFC,Near Field Communication)。
在本说明书的一个实施例中,计算设备1200的上述部件以及图12中未示出的其他部件也可以彼此相连接,例如通过总线。应当理解,图12所示的计算设备结构框图仅仅是出于示例的目的,而不是对本说明书范围的限制。本领域技术人员可以根据需要,增添或替换其他部件。
计算设备1200可以是任何类型的静止或移动计算设备,包括移动计算机或移动计算设备(例如,平板计算机、个人数字助理、膝上型计算机、笔记本计算机、上网本等)、移动电话(例如,智能手机)、可佩戴的计算设备(例如,智能手表、智能眼镜等)或其他类型的移动设备,或者诸如台式计算机或个人计算机(PC,Personal Computer)的静止计算设备。计算设备1200还可以是移动式或静止式的服务器。
其中,处理器1220用于执行如下计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法、CT图像处理方法、癌症的计算机辅助诊断方法或肝癌的计算机辅助诊断方法的步骤。
本说明书中的各个实施例均采用递进的方式描述,各个实施例之间相同相似的部分互相参见即可,每个实施例重点说明的都是与其他实施例的不同之处。尤其,对于计算设备实施例而言,由于其基本相似于图像处理方法、CT图像处理方法、癌症的计算机辅助诊断方法或肝癌的计算机辅助诊断方法实施例,所以描述的比较简单,相关之处参见图像处理方法、CT图像处理方法、癌症的计算机辅助诊断方法或肝癌的计算机辅助诊断方法实施例的部分说明即可。
本说明书一实施例还提供一种计算机可读存储介质,其存储有计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法、CT图像处理方法、癌症的计算机辅助诊断方法或肝癌的计算机辅助诊断方法的步骤。
本说明书中的各个实施例均采用递进的方式描述,各个实施例之间相同相似的部分互相参见即可,每个实施例重点说明的都是与其他实施例的不同之处。尤其,对于计算机可读存储介质实施例而言,由于其基本相似于图像处理方法、CT图像处理方法、癌症的计算机辅助诊断方法或肝癌的计算机辅助诊断方法实施例,所以描述的比较简单,相关之处参见图像处理方法、CT图像处理方法、癌症的计算机辅助诊断方法或肝癌的计算机辅助诊断方法实施例的部分说明即可。
本说明书一实施例还提供一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法、CT图像处理方法、癌症的计算机辅助诊断方法或肝癌的计算机辅助诊断方法的步骤。
上述为本实施例的一种计算机程序产品的示意性方案。需要说明的是,该计算机程序产品的技术方案与上述的图像处理方法、CT图像处理方法、癌症的计算机辅助诊断方法或肝癌的计算机辅助诊断方法的技术方案属于同一构思,计算机程序产品的技术方案未详细描述的细节内容,均可以参见上述图像处理方法、CT图像处理方法、癌症的计算机辅助诊断方法或肝癌的计算机辅助诊断方法的技术方案的描述。
上述对本说明书特定实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。
所述计算机指令包括计算机程序代码,所述计算机程序代码可以为源代码形式、对象代码形式、可执行文件或某些中间形式等。所述计算机可读介质可以包括:能够携带所述计算机程序代码的任何实体或装置、记录介质、U盘、移动硬盘、磁碟、光盘、计算机存储器、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、电载波信号、电信信号以及软件分发介质等。需要说明的是,所述计算机可读介质包含的内容可以根据专利实践的要求进行适当的增减,例如在某些地区,根据专利实践,计算机可读介质不包括电载波信号和电信信号。
需要说明的是,上述对本说明书特定实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。其次,本领域技术人员也应该知悉,说明书中所描述的实施例均属于优选实施例,所涉及的动作和模块并不一定都是本公开所必须的。
在上述实施例中,对各个实施例的描述都各有侧重,某个实施例中没有详述的部分,可以参见其它实施例的相关描述。
以上公开的本说明书优选实施例只是用于帮助阐述本说明书。可选实施例并没有详尽叙述所有的细节,也不限制该发明仅为所述的具体实施方式。显然,根据本公开的内容,可作很多的修改和变化。本说明书选取并具体描述这些实施例,是为了更好地解释本公开的原理和实际应用,从而使所属技术领域技术人员能很好地理解和利用本说明书。本说明书仅受权利要求书及其全部范围和等效物的限制。
Claims (22)
- 一种图像处理方法,包括:接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象;将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各目标图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
- 如权利要求1所述的方法,所述图像处理模型包括图像处理主干网络、对象分支网络、病灶分支网络和融合网络;将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,包括:将所述多个目标图像输入至所述图像处理主干网络,提取所述图像处理主干网络中的待处理对象特征信息和待处理病灶特征信息;将所述待处理对象特征信息输入至所述对象分支网络,获得对象特征信息,将所述待处理病灶特征信息输入至所述病灶分支网络,获得病灶特征信息和待分割图像特征信息;将所述对象特征信息和所述病灶特征信息输入至所述融合网络,生成对象融合特征信息;根据所述病灶特征信息生成病灶分类信息,根据所述待分割图像特征信息生成病灶分割信息,根据所述对象融合特征信息生成对象分类信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
- 如权利要求2所述的方法,所述图像处理主干网络包括图像编码器和图像解码器;将所述多个目标图像输入至所述图像处理主干网络,提取所述图像处理主干网络中的待处理对象特征信息和待处理病灶特征信息,包括:将所述多个目标图像输入至所述图像编码器,获得目标图像编码特征和至少一个多尺度图像编码特征;将所述目标图像编码特征输入至所述图像解码器,获得至少一个第一多尺度图像解码特征、至少一个第二多尺度图像解码特征和目标图像解码特征;将所述至少一个多尺度图像编码特征和至少一个第一多尺度图像解码特征确定为待处理对象特征信息,将所述至少一个第二多尺度图像解码特征和所述目标图像解码特征确定为待处理病灶特征信息。
- 如权利要求3所述的方法,所述图像编码器包括多个顺次连接的图像编码层;将所述多个目标图像输入至所述图像编码器,获得目标图像编码特征和至少一个多尺度图像编码特征,包括:将所述多个目标图像输入至所述图像编码器,获得所述图像编码器输出的目标图像编码特征;在所述多个图像编码层中确定至少一个目标图像编码层,获得各目标图像编码层输出的多尺度图像编码特征。
- 如权利要求3或4所述的方法,所述图像解码器包括多个顺次连接的图像解码层;将所述目标图像编码特征输入至所述图像解码器,获得至少一个第一多尺度图像解码特征和至少一个第二多尺度图像解码特征,包括:在所述多个图像解码层中确定至少一个第一图像解码层和至少一个第二图像解码层;获得第一图像解码层输出的第一多尺度图像解码特征,获得第二图像解码层输出的第二多尺度图像解码特征。
- 如权利要求2至5任意一项所述的方法,所述待处理病灶特征信息包括至少一个第二多尺度图像解码特征和目标图像解码特征;将所述待处理病灶特征信息输入至所述病灶分支网络,获得病灶特征信息和待分割图像特征信息,包括:对各第二多尺度图像解码特征进行特征解码,获得病灶特征信息;融合所述病灶特征信息和所述目标图像解码特征,生成待分割图像特征信息。
- 如权利要求2至6任意一项所述的方法,根据所述病灶特征信息生成病灶分类信息,根据所述待分割图像特征信息生成病灶分割信息,包括:将所述病灶特征信息输入至病灶分类器,生成病灶分类信息;将所述待分割图像特征信息输入至分割器,生成病灶分割信息。
- 如权利要求2至7任意一项所述的方法,将所述对象特征信息和所述病灶特征信息输入至所述融合网络,生成对象融合特征信息,包括:将所述对象特征信息和所述病灶特征信息输入至所述融合网络,在所述融合网络中提取所述病灶特征信息中的病灶子特征信息,并拼接所述对象特征信息和所述病灶子特征信息,生成对象融合特征信息。
- 如权利要求2至8任意一项所述的方法,根据所述对象融合特征信息生成对象分类信息,包括:将所述对象融合特征信息输入至对象分类器,生成对象分类信息。
- 如权利要求1至9任意一项所述的方法,所述图像处理模型通过下述步骤训练生成:获取多个样本目标图像和多个样本目标图像对应的样本检测结果,其中,所述样本检测结果包括样本对象分类信息、样本病灶分类信息和样本病灶分割信息;将所述多个样本目标图像输入至图像处理模型,获得所述图像处理模型输出的预测检测结果和训练对比损失值;根据所述预测检测结果和所述样本检测结果计算预测损失值;根据所述预测损失值和所述训练对比损失值调整图像处理模型的模型参数,并继续训练所述图像处理模型,直至达到模型训练停止条件。
- 如权利要求10所述的方法,获取多个样本目标图像和多个样本目标图像对应的样本检测结果,包括:确定样本检测对象,并获取样本检测对象对应的样本增强图像和样本目标图像,其中,样本检测对象对应的样本增强图像和样本目标图像对应样本检测对象的同一目标检测区域;在所述样本增强图像上针对目标检测区域生成参考样本检测结果;基于所述参考样本检测结果在所述样本目标图像上标注样本检测结果。
- 如权利要求10或11所述的方法,所述图像处理模型包括图像处理主干网络、对象分支网络、病灶分支网络和融合网络;将所述多个样本目标图像输入至图像处理模型,获得所述图像处理模型输出的预测检测结果和训练对比损失值,包括:将所述多个样本目标图像输入至图像处理主干网络,提取所述图像处理主干网络中的样本待处理对象特征信息和样本待处理病灶特征信息;将所述样本待处理对象特征信息输入至所述对象分支网络,获得预测对象特征信息,将所述样本待处理病灶特征信息输入至所述病灶分支网络,获得预测病灶特征信息和预测待分割图像特征信息;在所述对象分支网络中,根据样本目标图像对应的样本对象分类信息获取对象比对损失值,在所述病灶分支网络中,根据样本目标图像对应的样本病灶分类信息获取病灶比对损失值;将所述预测对象特征信息和所述预测病灶特征信息输入至所述融合网络,生成样本对象融合特征信息;根据所述预测对象特征信息生成第一对象分类信息,根据所述样本对象融合特征信息生成第二对象分类信息,根据所述预测病灶特征信息生成预测病灶分类信息,根据所述预测待分割图像特征信息生成预测病灶分割信息;根据所述第一对象分类信息、所述第二对象分类信息、所述预测病灶分类信息和所述预测病灶分割信息生成预测检测结果,根据所述对象比对损失值和所述病灶比对损失值生成训练对比损失值。
- 如权利要求12所述的方法,根据所述预测检测结果和所述样本检测结果计算预测损失值,包括:根据所述第一对象分类信息和所述样本对象分类信息计算第一对象损失值;根据所述第二对象分类信息和所述样本对象分类信息计算第二对象损失值;根据所述预测病灶分类信息和所述样本病灶分类信息计算病灶分类损失值;根据所述预测病灶分割信息和所述样本病灶分割信息计算病灶分割损失值。
- 如权利要求12或13所述的方法,将所述样本待处理对象特征信息输入至所述对象分支网络,获得预测对象特征信息,包括:获取历史样本对象特征信息;将所述历史对象特征信息和所述样本待处理对象特征信息输入至所述对象分支网络,获得预测对象特征信息;相应的,将所述样本待处理病灶特征信息输入至所述病灶分支网络,获得预测病灶特征信息和预测待分割图像特征信息,包括:获取历史样本病灶特征信息;将所述历史样本病灶特征信息和所述样本待处理病灶特征信息输入至所述病灶分支网络,获得预测病灶特征信息和预测待分割图像特征信息。
- 如权利要求12至14任意一项所述的方法,根据所述预测损失值和所述训练对比损失值调整图像处理模型的模型参数,包括:根据所述第一对象损失值、所述第二对象损失值、所述病灶分类损失值、所述病灶分割损失值、所述对象比对损失值和所述病灶比对损失值计算模型损失值;根据所述模型损失值调整图像处理模型的模型参数。
- 一种CT图像处理方法,包括:接收CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在异常对象;将所述多个平扫CT图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
- 一种癌症的计算机辅助诊断方法,包括:接收CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在肿瘤;将所述多个平扫CT图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
- 一种肝癌的计算机辅助诊断方法,包括接收CT图像处理任务,其中,所述CT图像处理任务携带肝脏区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述肝脏区域内是否存在肿瘤;将所述多个平扫CT图像输入至图像处理模型,获得所述肝脏区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
- 一种癌症的计算机辅助诊断系统,包括客户端和服务端,其中,所述客户端,用于向所述服务端发送CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个平扫CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在肿瘤;所述服务端,用于将所述多个平扫CT图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于各平扫CT图像生成对象特征信息和病灶特征信息,根据所述对象特征信息和所述病灶特征信息生成对象融合特征信息,根据对象融合特征信息确定对象分类信息,根据病灶特征信息确定病灶分类信息和病灶分割信息,根据所述对象分类信息、所述病灶分类信息和所述病灶分割信息生成检测结果。
- 一种计算设备,包括:存储器和处理器;所述存储器用于存储计算机程序/指令,所述处理器用于执行所述计算机程序/指令,该计算机程序/指令被处理器执行时实现权利要求1至18任意一项所述方法的步骤。
- 一种计算机可读存储介质,其存储有计算机程序/指令,该计算机程序/指令被处理器执行时实现权利要求1至18任意一项所述方法的步骤。
- 一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现权利要求1至18任意一项所述方法的步骤。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410927473.X | 2024-07-10 | ||
| CN202410927473.XA CN119048419A (zh) | 2024-07-10 | 2024-07-10 | 图像处理方法、癌症的计算机辅助诊断方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026012003A1 true WO2026012003A1 (zh) | 2026-01-15 |
Family
ID=93569650
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/098321 Pending WO2026012003A1 (zh) | 2024-07-10 | 2025-05-30 | 图像处理方法、癌症的计算机辅助诊断方法 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN119048419A (zh) |
| WO (1) | WO2026012003A1 (zh) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119048419A (zh) * | 2024-07-10 | 2024-11-29 | 阿里巴巴(中国)有限公司 | 图像处理方法、癌症的计算机辅助诊断方法 |
| TWI901468B (zh) * | 2024-12-04 | 2025-10-11 | 智德萬生醫科技股份有限公司 | 醫學影像處理方法及其系統 |
| CN119251222B (zh) * | 2024-12-04 | 2025-04-01 | 南方医科大学 | 基于反投影数据的平扫ct图像病变检测方法及相关设备 |
| CN120746925A (zh) * | 2025-05-15 | 2025-10-03 | 阿里巴巴达摩院(杭州)科技有限公司 | 图像处理模型训练方法、图像处理方法 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117854705A (zh) * | 2023-12-28 | 2024-04-09 | 澳门大学 | 上消化道多病变多任务智能诊断方法、装置、设备及介质 |
| CN117853490A (zh) * | 2024-03-06 | 2024-04-09 | 阿里巴巴达摩院(杭州)科技有限公司 | 图像处理方法、图像处理模型的训练方法 |
| CN117911307A (zh) * | 2022-10-09 | 2024-04-19 | 上海微创卜算子医疗科技有限公司 | 医学图像分类方法、系统、电子设备和存储介质 |
| CN118154965A (zh) * | 2024-03-18 | 2024-06-07 | 紫桐智能(厦门)信息科技有限公司 | 一种病灶的检测方法和装置 |
| CN119048419A (zh) * | 2024-07-10 | 2024-11-29 | 阿里巴巴(中国)有限公司 | 图像处理方法、癌症的计算机辅助诊断方法 |
-
2024
- 2024-07-10 CN CN202410927473.XA patent/CN119048419A/zh active Pending
-
2025
- 2025-05-30 WO PCT/CN2025/098321 patent/WO2026012003A1/zh active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117911307A (zh) * | 2022-10-09 | 2024-04-19 | 上海微创卜算子医疗科技有限公司 | 医学图像分类方法、系统、电子设备和存储介质 |
| CN117854705A (zh) * | 2023-12-28 | 2024-04-09 | 澳门大学 | 上消化道多病变多任务智能诊断方法、装置、设备及介质 |
| CN117853490A (zh) * | 2024-03-06 | 2024-04-09 | 阿里巴巴达摩院(杭州)科技有限公司 | 图像处理方法、图像处理模型的训练方法 |
| CN118154965A (zh) * | 2024-03-18 | 2024-06-07 | 紫桐智能(厦门)信息科技有限公司 | 一种病灶的检测方法和装置 |
| CN119048419A (zh) * | 2024-07-10 | 2024-11-29 | 阿里巴巴(中国)有限公司 | 图像处理方法、癌症的计算机辅助诊断方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN119048419A (zh) | 2024-11-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Jin et al. | AI-assisted CT imaging analysis for COVID-19 screening: Building and deploying a medical AI system in four weeks | |
| WO2026012003A1 (zh) | 图像处理方法、癌症的计算机辅助诊断方法 | |
| AlGhamdi et al. | DU-Net: Convolutional network for the detection of arterial calcifications in mammograms | |
| Singh et al. | Pneumonia detection with QCSA network on chest X-ray | |
| CN116797554B (zh) | 图像处理方法以及装置 | |
| CN117853490B (zh) | 图像处理方法、图像处理模型的训练方法 | |
| WO2025246480A1 (zh) | 图像处理模型的训练方法、图像处理方法 | |
| TW202346826A (zh) | 影像處理方法 | |
| WO2025180099A1 (zh) | 目标对象识别方法、对象识别模型训练方法、ct图像中的可见淋巴结检测方法、计算机辅助诊断方法、电子设备、存储介质及程序产品 | |
| WO2026001027A1 (zh) | 目标图像处理模型训练方法、图像处理方法 | |
| CN111445457A (zh) | 网络模型的训练方法及装置、识别方法及装置、电子设备 | |
| Sui et al. | Deep learning-based channel squeeze U-structure for lung nodule detection and segmentation | |
| WO2025007942A1 (zh) | 图像处理方法、图像处理模型的训练方法 | |
| CN117408948A (zh) | 图像处理方法、图像分类分割模型的训练方法 | |
| CN116416239B (zh) | 胰腺ct图像分类方法、图像分类模型、电子设备及介质 | |
| Gao et al. | A lung CT vision foundation model facilitating disease diagnosis and medical imaging | |
| CN117408946A (zh) | 图像处理模型的训练方法、图像处理方法 | |
| Liu et al. | An end to end thyroid nodule segmentation model based on optimized U-net convolutional neural network | |
| Kabir et al. | TriGWONet a lightweight multibranch convolutional neural network using gray wolf optimization for accurate oral cancer image classification | |
| Yang et al. | ESFCU‐Net: A Lightweight Hybrid Architecture Incorporating Self‐Attention and Edge Enhancement Mechanisms for Enhanced Polyp Image Segmentation | |
| Wang et al. | Efficient hierarchical multiscale convolutional attention for accurate medical image segmentation: B. Wang et al. | |
| CN121281731A (zh) | 基于全局依赖学习和多模态对齐网络的放射学报告生成方法及系统 | |
| CN116993663B (zh) | 图像处理方法、图像处理模型的训练方法 | |
| CN120782704A (zh) | 图像处理方法及图像处理模型训练方法 | |
| Prazuch et al. | CIRCA: comprehensible online system in support of chest X-rays-based screening by COVID-19 example |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25835906 Country of ref document: EP Kind code of ref document: A1 |