WO2026001027A1 - 目标图像处理模型训练方法、图像处理方法 - Google Patents

目标图像处理模型训练方法、图像处理方法

Info

Publication number
WO2026001027A1
WO2026001027A1 PCT/CN2025/078597 CN2025078597W WO2026001027A1 WO 2026001027 A1 WO2026001027 A1 WO 2026001027A1 CN 2025078597 W CN2025078597 W CN 2025078597W WO 2026001027 A1 WO2026001027 A1 WO 2026001027A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
target
features
image processing
processing model
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/078597
Other languages
English (en)
French (fr)
Inventor
唐禹行
张灵
方伟
曹维维
莫志榮
夏英达
张建鹏
许敏丰
吕乐
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba China Co Ltd
Original Assignee
Alibaba China Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba China Co Ltd filed Critical Alibaba China Co Ltd
Publication of WO2026001027A1 publication Critical patent/WO2026001027A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/0002Inspection of images, e.g. flaw detection
    • G06T7/0012Biomedical image inspection
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/25Determination of region of interest [ROI] or a volume of interest [VOI]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/26Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/44Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
    • G06V10/443Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components by matching or filtering
    • G06V10/449Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters
    • G06V10/451Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters with interaction between the filter responses, e.g. cortical complex cells
    • G06V10/454Integrating the filters into a hierarchical structure, e.g. convolutional neural networks [CNN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/774Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/80Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
    • G06V10/803Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of input or preprocessed data
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10072Tomographic images
    • G06T2207/10081Computed x-ray tomography [CT]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30004Biomedical image processing
    • G06T2207/30096Tumor; Lesion
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V2201/00Indexing scheme relating to image or video recognition or understanding
    • G06V2201/03Recognition of patterns in medical or anatomical images
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V2201/00Indexing scheme relating to image or video recognition or understanding
    • G06V2201/07Target detection

Definitions

  • This specification relates to the field of computer technology, and in particular to a target image processing model training method and an image processing method; one or more embodiments of this specification also relate to a target image processing model training device, an image processing method, a computing device, a computer-readable storage medium, and a computer program product.
  • NCCT non-contrast computed tomography
  • embodiments of this specification provide a method for training a target image processing model.
  • One or more embodiments of this specification also relate to a target image processing model training apparatus, an image processing method, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies in the prior art where examination results obtained using plain CT scans are inaccurate.
  • a method for training a target image processing model comprising:
  • the initial image and the enhanced image are fused to obtain a fused image, and the initial image is masked to obtain a masked image;
  • the fused image is input into a reference image processing model, and the reference image processing model is used to encode and decode the fused image to obtain initial decoded image features and initial decoded classification features.
  • the masked image is input into the target image processing model, and the target image processing model is used to encode and decode the masked image to obtain target decoded image features and target decoded classification features;
  • the target image processing model is trained based on the initial decoded image features, the initial decoded classification features, the target decoded image features, and the target decoded classification features.
  • an image processing method comprising:
  • the target image is encoded and decoded using the target image processing model to obtain the target image features and target classification features.
  • the classification result corresponding to the target image is obtained.
  • a computer-aided diagnosis method for cancer comprising:
  • the CT image is input into the CT image processing model, and the CT image is processed by the CT image processing model to obtain the CT segmentation image and CT classification result corresponding to the CT image.
  • the CT image processing model is trained by the above-mentioned target image processing model training method.
  • a detection result is obtained to determine whether a tumor exists in the target detection area.
  • a computer-aided diagnostic method for breast cancer comprising:
  • the CT image is input into the CT image processing model, and the CT image is processed by the CT image processing model to obtain the CT segmentation image and CT classification result corresponding to the CT image.
  • the CT image processing model is trained by the above-mentioned target image processing model training method.
  • the detection result of whether there is a tumor in the breast region is obtained.
  • another image processing method is provided, applied to a client of a medical system, comprising:
  • a medical image is determined
  • the medical image is sent to the server of the medical system, and the segmented image and classification result corresponding to the medical image are received from the server.
  • the segmented image and classification result corresponding to the medical image are obtained by processing the medical image according to the target image processing model.
  • the target image processing model is trained by the above-mentioned target image processing model training method.
  • the segmented image and classification results corresponding to the medical image are displayed to the user through the user interface.
  • a computer-aided diagnosis system for cancer comprising a client and a server, wherein,
  • the client is used to send CT images of the target detection area to the server;
  • the server is used to input the CT image into a CT image processing model, process the CT image using the CT image processing model, obtain the segmentation image and classification result corresponding to the CT image, and obtain the detection result of whether there is a tumor in the target detection area based on the CT segmentation image and the CT classification result, and return the detection result to the client.
  • the CT image processing model is trained using the above-mentioned target image processing model training method.
  • a target image processing model training apparatus comprising:
  • the image determination module is configured to determine the initial image and the enhanced image of the target object
  • the image acquisition module is configured to fuse the initial image and the enhanced image to obtain a fused image, and to perform masking processing on the initial image to obtain a masked image;
  • the initial feature acquisition module is configured to input the fused image into a reference image processing model, and use the reference image processing model to perform encoding and decoding processing on the fused image to obtain initial decoded image features and initial decoded classification features.
  • the target feature acquisition module is configured to input the mask image into the target image processing model, and use the target image processing model to perform encoding and decoding processing on the mask image to obtain target decoded image features and target decoded classification features;
  • the model training module is configured to train the target image processing model based on the initial decoded image features, the initial decoded classification features, the target decoded image features, and the target decoded classification features.
  • an image processing apparatus comprising:
  • the determination module is configured to determine the target image and input the target image into the target image processing model
  • the feature acquisition module is configured to use the target image processing model to perform encoding and decoding processing on the target image to obtain the target image features and target classification features of the target image.
  • the image acquisition module is configured to obtain a segmented image corresponding to the target image using the features of the target image;
  • the result acquisition module is configured to use the target classification features to obtain the classification result corresponding to the target image.
  • a computing device comprising:
  • the memory is used to store computer programs/instructions
  • the processor is used to execute the computer programs/instructions.
  • the computer programs/instructions are executed by the processor, they implement the steps of the above-mentioned target image processing model training method and image processing method.
  • a computer-readable storage medium stores a computer program/instructions, which, when executed by a processor, implement the steps of the above-described target image processing model training method and image processing method.
  • a computer program product including a computer program/instructions, which, when executed by a processor, implement the steps of the above-described target image processing model training method and image processing method.
  • This specification provides one or more embodiments of a target image processing model training method.
  • the method utilizes the initial decoded image features and initial decoded classification features, which contain rich information about the target object and can be obtained from the reference image processing model.
  • the method can refer to the initial decoded image features and initial decoded classification features when processing the mask image to recover its features.
  • the method can improve the accuracy of the target image processing model in detecting and segmenting the target object using the initial image.
  • Figure 1 is a schematic diagram of a scene of an image processing method provided in one embodiment of this specification
  • Figure 2 is a flowchart of a target image processing model training method provided in one embodiment of this specification
  • FIG. 3 is a flowchart of an image processing method provided in one embodiment of this specification.
  • Figure 4 is a flowchart of a computer-aided diagnostic method for breast cancer provided in one embodiment of this specification
  • Figure 5 is a flowchart of an image processing method for a client application in a medical system, provided in one embodiment of this specification.
  • Figure 6 is a structural framework diagram of a target image processing model training method provided in an embodiment of this specification, applied to a medical image diagnosis scenario.
  • Figure 7 is a schematic diagram of the structure of a target image processing model training device provided in one embodiment of this specification.
  • Figure 8 is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification.
  • Figure 9 is a structural block diagram of a computing device provided in one embodiment of this specification.
  • first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word “if” as used herein may be interpreted as "when,” “when,” or "in response to a determination.”
  • This specification provides a method for training a target image processing model. It also relates to a target image processing model training device, an image processing method, an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
  • the image processing methods provided in the embodiments of this specification can be applied to different scenarios.
  • the image processing method when the image processing method is applied to an autonomous driving scenario, it can detect and identify pedestrians, vehicles, or traffic signs in traffic images.
  • the image processing method when the image processing method is applied to a medical scenario, specifically for medical image analysis, it can detect and identify organs, tumors, or lesion areas in medical images.
  • a target image processing model is trained in server 104, as shown in Figure 1.
  • This target image processing model includes an encoding layer, a decoding layer, and a fully connected layer.
  • server 104 receives a plain CT image sent by end device 102, it inputs the plain CT image into the target image processing model.
  • the encoding and decoding layers of the target image processing model encode and decode the plain CT image to obtain target image features and classification features corresponding to each decoding layer.
  • the segmentation image corresponding to the plain CT image is obtained through the image features.
  • This segmentation image is an image of the same size as the plain CT image, using different gray values to indicate different regions. For example, 0 represents the background, 1 represents an organ, and 2 represents a tumor.
  • the training steps of the target image processing model are as follows: Determine the initial image and enhanced image of the target object; fuse the initial image and the enhanced image to obtain a fused image, and perform masking processing on the initial image to obtain a masked image; input the fused image into a reference image processing model, and use the reference image processing model to perform encoding and decoding processing on the fused image to obtain initial decoded image features and initial decoded classification features; input the masked image into the target image processing model, and use the target image processing model to perform encoding and decoding processing on the masked image to obtain target decoded image features and target decoded classification features; train the target image processing model based on the initial decoded image features, the initial decoded classification features, the target decoded image features, and the target decoded classification features.
  • the edge device 102 may include a browser, an app (application), or a web application such as an H5 (Hypertext Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application.
  • the edge device can be developed based on a software development kit (SDK) provided by the server, such as a Real-Time Communication (RTC) SDK.
  • SDK software development kit
  • RTC Real-Time Communication
  • the edge device can be deployed in an electronic device and depends on the device's operation or certain apps within the device to run.
  • the electronic device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer.
  • Various other types of applications can also be configured in the electronic device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.
  • Server 104 can be understood as a server providing various services, including physical servers and cloud servers. For example, it could be a server providing communication services to multiple clients, a server supporting backend training of models used on clients, or a server processing data sent by clients. It's important to note that Server 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. Server 104 can also be a server in a distributed system, or a server integrated with blockchain.
  • Server 104 can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
  • basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
  • the image processing method provided in the embodiments of this specification can be executed by the server 104.
  • the target image processing model can be deployed in the edge device 102, so that the edge device 102 can also have similar functions to the server 104, thereby executing the image processing method provided in the embodiments of this specification.
  • the image processing method provided in the embodiments of this specification can also be jointly executed by the edge device 102 and the server 104.
  • the image processing method provided in the embodiments of this specification inputs an image into a target image processing model, and uses the target image processing model to obtain the target image features and target classification features of the image.
  • the target image processing model is trained with features containing enhanced image information, the target image processing model can better capture detailed feature-level information of the target object in the image, thereby obtaining more accurate image segmentation and classification results.
  • the method includes the following steps.
  • Step 202 Determine the initial image and enhanced image of the target object.
  • the target object can be understood as a region, object, or feature in an image that has specific meaning or interest. These are key elements in tasks such as image analysis, image processing, and image understanding.
  • the target object can be understood differently in different application scenarios. For example, in a traffic scenario, the target object can be understood as pedestrians, vehicles, or traffic signs; in a logistics scenario, the target object can be understood as shelves, goods, or electronic tags; and in a medical scenario, the target object can be understood as organs, tumors, etc.
  • An initial image can be understood as the original image obtained by capturing (such as taking a picture or scanning) a target object.
  • An initial image can be obtained directly through a camera, scanner, or other image capture device.
  • An enhanced image can be understood as an image that highlights the target object.
  • An enhanced image can be obtained by adjusting the initial image.
  • the initial image might be a plain CT scan.
  • the enhanced image could be a contrast-enhanced CT scan.
  • Contrast-enhanced CT scans through the use of contrast agents, can more clearly display vascular structures, tumors, inflammatory areas, or other pathological changes compared to plain CT scans because the contrast between them and surrounding normal tissue is increased.
  • the initial image might be a shelf image captured by an inspection robot. If the target is goods, the enhanced image could be a sharpened version of the initial image, making the goods on the shelf more clearly visible.
  • the initial image and the enhanced image of the target object are determined.
  • the enhanced image contains richer privileged information related to the target object than the initial image.
  • a fused image can be understood as an image obtained by stitching and fusing the initial image and the enhanced image, weighted fusion, Poisson fusion, etc.
  • a masked image can be understood as an image obtained by masking the initial image, randomly masking some of the image pixels in the initial image.
  • the reference image processing model can better understand the complex features in the plain CT images. This allows the reference image processing model (as the teacher model) to transfer feature-level information to the target image processing model (as the student model). Subsequently, when the input data of the target image processing model is a plain CT image, it can also perform reasonable processing on the plain CT image, thereby improving the accuracy and effectiveness of detection and segmentation based on plain CT images.
  • the initial image sample and the enhanced image sample are fused to obtain a fused image sample, and the fused image sample is input into a reference image processing model;
  • the reference image processing model is trained based on the predicted segmented image, the image label, the predicted classification result, and the classification label.
  • obtaining the predicted segmentation image and the predicted classification result corresponding to the initial image sample using the reference image processing model includes:
  • the fused image samples are encoded and decoded to obtain target decoded image sample features and target decoded classification sample features;
  • the predicted segmented image corresponding to the fused image sample is obtained;
  • the predicted classification result corresponding to the fused image sample is obtained.
  • the reference image processing model includes an encoder and a decoder, wherein the encoder is composed of multiple encoding layers and the decoder is composed of multiple decoding layers;
  • the step of using the reference image processing model to encode and decode the fused image samples to obtain target decoded image sample features and target decoded classification sample features includes:
  • the fused image samples are encoded using the multiple coding layers to obtain initial encoded image sample features
  • the initial encoded image sample features are decoded using the multiple decoding layers to obtain the initial decoded image sample features corresponding to each of the multiple decoding layers.
  • the initial decoded image sample features corresponding to the final decoded layer in the plurality of decoding layers are determined as target decoding classification sample features; or, according to the image classification task, the initial decoded image sample features corresponding to each of the plurality of decoding layers are subjected to convolution and pooling processing to obtain key image sample features corresponding to each decoding layer, and the key image sample features corresponding to each decoding layer are fused to obtain target decoding classification sample features.
  • the training process of the reference image processing model is explained in detail:
  • the initial image sample and the enhanced image sample of the target object are fused to obtain a fused image sample, which is then used as input to the reference image processing model.
  • the network architecture of the reference image processing model is a multi-task network architecture, including an encoder-decoder structure with skip connections.
  • the reference image processing model contains two task-specific branches: a segmentation branch for obtaining segmented images and a classification branch for obtaining classification results.
  • the encoder and decoder of the reference image processing model are used to encode and decode the fused image samples. That is, the fused image samples are encoded through multiple encoding layers in the encoder to obtain the initial encoded image sample features, and the initial encoded image sample features are decoded through multiple decoding layers in the decoder to obtain the initial decoded image sample features corresponding to each decoding layer.
  • the initial decoded image sample features corresponding to the last decoding layer among multiple decoding layers can be determined as the target decoding classification sample features.
  • the corresponding initial decoded image sample features can be obtained from each decoding layer, and convolution processing can be performed on these features.
  • Global max pooling can be applied to extract representative key image sample features from the convolutional features, and the key image sample features corresponding to each decoding layer can be concatenated and fused to obtain the target decoding classification sample features.
  • the encoder captures high-level features of the fused image samples to obtain a highly abstract representation of the content of the fused image samples.
  • the decoder is responsible for restoring the abstract feature map generated by the encoder to the spatial size of the original image.
  • each layer of the encoder is directly connected to the corresponding decoding layer (usually at the same or similar scale), directly passing the feature map from the encoding stage, thus preserving the detailed information of the original image and helping to reconstruct finer boundaries and textures in the final output.
  • the features finally output by the decoder after processing by multiple decoding layers are determined as the features of the target decoded image sample.
  • the predicted segmentation image corresponding to the initial image sample is obtained by using the features of the target decoded image sample.
  • This predicted segmentation image uses different gray values to represent different objects to achieve the segmentation of different objects.
  • the segmentation loss function of the segmentation branch is calculated using the predicted segmented image and image label
  • the classification loss function of the classification branch is calculated using the predicted classification result and classification label.
  • a reference image processing model is trained, which is a more accurate reference image processing model trained using paired initial image samples and augmented image samples. This reference image processing model benefits from a larger training dataset and is able to capture the rich information present in the augmented image.
  • a knowledge distillation technique is used to transfer the knowledge learned from the reference image processing model (teacher model) to the target image processing model (student model), which has lower accuracy and is trained using the initial image samples.
  • the target image processing model has the same model structure as the reference image processing model, thereby improving the accuracy of segmentation images and classification results obtained by the target image processing model based on the initial image.
  • Step 206 Input the fused image into the reference image processing model, and use the reference image processing model to perform encoding and decoding processing on the fused image to obtain initial decoded image features and initial decoded classification features.
  • the initial decoded image features can be understood as the features obtained using the decoding layers of the reference image processing model, used to generate the segmented image corresponding to the initial image. These initial decoded image features can be features corresponding to each decoding layer used to generate the segmented image corresponding to the initial image, or features corresponding to the last decoding layer after multiple decoding layers used to generate the segmented image corresponding to the initial image.
  • the initial decoded classification features can be understood as the features obtained using the decoding layers of the reference image processing model used to classify the initial image. These initial decoded classification features can be features corresponding to each decoding layer used to classify the initial image, or features corresponding to the last decoding layer after multiple decoding layers used to classify the initial image.
  • initial decoded image features and initial decoded classification features corresponding to the fused image are obtained.
  • These initial decoded image features are suitable for generating or understanding image segmentation.
  • Image segmentation refers to dividing an image into multiple regions, each region corresponding to different objects or object categories in the image. Therefore, these features help to identify and separate different objects or regions in the image.
  • the initial decoded classification features focus on classifying the entire image or the main body in the image, that is, identifying which category the image belongs to or what attributes it has. They are key to the model's judgment of image content, such as determining whether an image is an organ, tumor, etc.
  • the reference image processing model includes an encoder and a decoder.
  • the encoder is composed of multiple encoding layers
  • the decoder is composed of multiple decoding layers. Through processing by multiple decoding layers, initial decoded image features corresponding to multiple decoding layers can be obtained, thereby determining initial decoded classification features based on the initial decoded image features corresponding to multiple decoding layers.
  • the specific implementation is as follows:
  • the step of using the reference image processing model to encode and decode the fused image to obtain initial decoded image features and initial decoded classification features includes:
  • the fused image is encoded using the multiple coding layers to obtain initial coded image features
  • the initial encoded image features are decoded using the multiple decoding layers to obtain the initial decoded image features corresponding to the multiple decoding layers, and the initial decoded classification features are determined based on the initial decoded image features corresponding to the multiple decoding layers.
  • the initial encoded image features can be understood as the compact and information-rich feature representation obtained by encoding the fused image through multiple encoding layers; the initial decoded image features can be understood as the image representations at different levels of abstraction output by the decoding layer; and the initial decoded classification features can be understood as the features related to image classification determined based on the initial decoded image features corresponding to the decoding layer.
  • multiple coding layers are used to process the fused image.
  • the coding layers typically involve progressively reducing the spatial resolution of the image while increasing the abstraction level of the feature representation. This process can be viewed as extracting features from the fused image.
  • a series of dimensionality reduction operations such as convolution and pooling, the fused image is transformed into a set of compact and information-rich feature representations, i.e., the initial encoded image features.
  • decoding layers are used to decode the initial encoded image features obtained by the encoding layer.
  • the role of the decoding layer is to inversely decode and upsample the highly compressed feature information step by step to restore the spatial structure close to the initial image.
  • partial or complete structure of the image is reconstructed through operations such as deconvolution and upsampling, generating initial decoded image features corresponding to multiple decoding layers.
  • the output of each decoding layer can be regarded as an image representation at different levels of abstraction. The closer the decoding layer is to the output end, the closer its features are to the original spatial structure of the image.
  • the target image processing model training method uses multiple encoding layers and multiple decoding layers to encode and decode the fused image, which can obtain more refined initial decoded image features corresponding to each decoding layer. This allows for the subsequent determination of initial decoded classification features using the initial decoded image features, while also obtaining initial decoded classification features containing rich information. This facilitates the improvement of the performance of the trained target image processing model and the accuracy of image segmentation and image classification.
  • the step of determining the initial decoding classification features based on the initial decoded image features corresponding to the multiple decoding layers includes:
  • the initial decoded image features corresponding to the final decoding layer among the multiple decoding layers are determined as the initial decoding classification features; or,
  • the initial decoded image features corresponding to each of the multiple decoding layers are subjected to convolution and pooling to obtain the first key decoded image features corresponding to each decoding layer.
  • the first key decoded image features corresponding to each decoding layer are then fused to obtain the initial decoding classification features.
  • the final decoding layer can be understood as the last decoding layer among multiple decoding layers. Since each decoding layer processes features based on the output of the previous decoding layer, the features output by this final decoding layer are the final product of the entire decoding process, containing the key information required for the model's in-depth understanding and prediction of the fused image.
  • the first key decoded image feature can be understood as the representative features for image classification among the initial decoded image features corresponding to each decoding layer.
  • Image classification tasks can be understood as classifying images into different types based on the target objects within them.
  • the goal of an image classification task could be to classify images as containing tumors or not, based on the target objects within them.
  • the goal of an image classification task could be to classify images as having normal or abnormal cargo displays, based on the target objects within them.
  • the features that the decoder finally outputs after processing through multiple decoding layers i.e., the initial decoded image features output by the last decoding layer, can be determined as the initial decoding classification features.
  • initial decoded image features can be extracted from each decoding layer according to the image classification task. This involves obtaining initial decoded image features corresponding to each decoding layer, containing different hierarchical scales. Convolutional processing is then used to further refine these initial decoded image features, obtaining features meaningful for the image classification task. Pooling operations (global large pooling is used in this embodiment) reduce spatial dimensions while retaining the most important information in each initial decoded image feature after convolution. Pooling also reduces the number of parameters in subsequent fully connected layers, thereby reducing the risk of overfitting and computational burden. The first key decoded image features corresponding to each decoding layer are then obtained. By concatenating and fusing these first key decoded image features, initial decoded classification features are obtained.
  • the target image processing model training method provided in the embodiments of this specification can obtain initial decoding classification features that meet different needs according to actual requirements. By fusing the first key decoding image features corresponding to each decoding layer to obtain initial decoding classification features, it can integrate the overall contextual information while capturing details at the local level, thereby obtaining more accurate results.
  • an initial segmentation image corresponding to the image segmentation task is obtained, and an initial classification result corresponding to the image classification task is obtained.
  • the specific implementation method is as follows:
  • the process further includes:
  • the initial segmented image corresponding to the initial image is obtained using the features of the initial decoded image
  • the initial classification result corresponding to the initial image is obtained by utilizing the initial decoding classification features.
  • Image segmentation can be understood as the task of segmenting different objects or regions in an image.
  • the process of image segmentation can be seen as classifying each pixel in the image and assigning a label to each pixel to indicate the category or object to which it belongs.
  • the background in the image is labeled as 0
  • the organ is labeled as 1
  • the tumor is labeled as 2.
  • the initial segmentation image can be understood as the image obtained by segmenting different objects in the initial image; the initial classification result can be understood as the classification result determined based on the target object in the initial image; for example, in the medical field, the initial segmentation image is an image with different gray values using 0 to represent the background, 1 to represent the organ, and 2 to represent the tumor, and the initial classification result includes two types, one for cancer and one for non-cancer.
  • the initial segmented image corresponding to the initial image can be obtained by using the initial decoded image features; and for image classification, the initial classification result corresponding to the initial image can be obtained by using the initial decoded classification features.
  • the target image processing model training method provided in the embodiments of this specification can not only achieve pixel-level image segmentation tasks using the features of the initial decoded image, but also perform image-level image classification tasks using the features of the initial decoded classification, thereby obtaining diverse results related to the initial image.
  • Step 208 Input the mask image into the target image processing model, and use the target image processing model to encode and decode the mask image to obtain target decoded image features and target decoded classification features.
  • the target image processing model has the same network architecture as the reference image processing model. That is, the target image processing model is also a multi-task network architecture, including an encoder-decoder structure with skip connections.
  • the target image processing model obtains target decoded image features and target decoded classification features.
  • the target decoded image features can be understood as the decoded features obtained using the decoding layer of the target image processing model, which are used to generate the segmented image corresponding to the mask image;
  • the target decoded classification features can be understood as the features obtained using the decoding layer of the target image processing model, which are used to classify the mask image.
  • the target image processing model includes an encoder and a decoder, wherein the encoder is composed of multiple encoding layers and the decoder is composed of multiple decoding layers;
  • the step of using the target image processing model to encode and decode the mask image to obtain target decoded image features and target decoded classification features includes:
  • the mask image is encoded using the multiple encoding layers to obtain the target encoded image features
  • the target encoded image features are decoded using the multiple decoding layers to obtain the target decoded image features corresponding to the multiple decoding layers, and the target decoding classification features are determined based on the target decoded image features corresponding to the multiple decoding layers.
  • determining the target decoding classification features based on the target decoded image features corresponding to the plurality of decoding layers includes:
  • the target decoded image features corresponding to the final decoding layer among the multiple decoding layers are determined as the target decoding classification features; or,
  • the target decoded image features corresponding to each of the multiple decoding layers are subjected to convolution and pooling processing to obtain the second key decoded image features corresponding to each decoding layer.
  • the second key decoded image features corresponding to each decoding layer are then fused to obtain the target decoding classification features.
  • the method further includes:
  • the target segmentation result corresponding to the initial image is obtained by utilizing the target decoded image features
  • the target classification result corresponding to the initial image is obtained by using the target decoding classification features.
  • the specific implementation method is similar to the process of processing fused images in the reference image processing model mentioned above, and will not be repeated here.
  • Step 210 Train the target image processing model based on the initial decoded image features, the initial decoded classification features, the target decoded image features, and the target decoded classification features.
  • the intermediate features between the reference image processing model and the target image processing model are minimized through Feature-Level Knowledge Distillation (FKD) loss. That is, the difference between the initial decoded image features and the target decoded image features, as well as the difference between the initial decoded classification features and the target decoded classification features, are minimized.
  • FKD Feature-Level Knowledge Distillation
  • a loss function between the initial decoded image features and the target decoded image features is calculated through feature-level knowledge distillation, and a loss function between the initial decoded classification features and the target decoded classification features is also calculated to train the target image processing model.
  • the specific implementation is as follows:
  • the step of training the target image processing model based on the initial decoded image features, the initial decoded classification features, the target decoded image features, and the target decoded classification features includes:
  • the target image processing model is trained based on the segmentation loss function and the classification loss function.
  • the segmentation loss function and the classification loss function can be measured using feature similarity, such as by using cosine similarity or mean square error.
  • the similarity between the initial decoded image features corresponding to each decoding layer of the reference image processing model and the target decoded image features corresponding to each decoding layer of the target image processing model is compared to obtain the segmentation loss function of each decoding layer. This can better guide the target image processing model (student model) to learn information at different scales and details.
  • the features output by the decoder in the reference image processing model, which uses multiple decoding layers can be used as the initial decoding classification features.
  • the features output by the decoder in the target image processing model, which uses multiple decoding layers can be used as the target decoding classification features.
  • the classification loss function can be obtained by comparing the similarity between the final feature representations (the features output by the decoder).
  • the target image processing model is trained using the segmentation loss function for image segmentation tasks and the classification loss function for image classification tasks.
  • the target image processing model training method provided in this specification, in the case that image classification tasks require layer-by-layer guidance at each decoding layer to obtain multi-scale information and ensure the learning of detailed information, obtains the segmentation loss function of the initial decoded image features and the target decoded image features at each decoding layer; and in the case that image classification tasks can effectively express global semantic information through the final feature representation, obtaining the classification loss function through the final feature representation can simplify calculation and improve efficiency.
  • the enhanced image is reconstructed using the target decoded image features corresponding to the decoding layer, further enhancing the distillation process, thereby obtaining the enhanced image information missing from the input of the target image processing model.
  • the method Before training the target image processing model based on the initial decoded image features, the initial decoded classification features, the target decoded image features, and the target decoded classification features, the method further includes:
  • the target decoded image features are convolved and upsampled to obtain the predicted decoded image features, and the predicted enhanced image is obtained based on the predicted decoded image features.
  • the step of training the target image processing model based on the initial decoded image features, the initial decoded classification features, the target decoded image features, and the target decoded classification features includes:
  • the target image processing model is trained based on the initial decoded image features, the initial decoded classification features, the target decoded image features, the target decoded classification features, the predicted enhanced image, and the enhanced image.
  • predicted decoded image features can be understood as features that contain the specific details and structural information required to generate the predicted enhanced image.
  • the target decoded image features are further refined and combined. Upsampling operations are used to increase the resolution of the feature map and enlarge its size, making it close to or restored to the size of the original image, in preparation for generating a high-resolution prediction image. This process can also restore the spatial details of the image and obtain the predicted decoded image features. Using the predicted decoded image features, a predicted enhanced image is generated. This predicted enhanced image should contain richer information related to the target object than the initial image.
  • the model is trained using the initial decoded image features, initial decoded classification features, target decoded image features, target decoded classification features, predicted augmented image, and augmented image.
  • the target image processing model training method integrates an additional reconstruction branch into the target image processing model, uses the additional reconstruction branch to obtain a predicted enhanced image, and then further enhances the distillation process through the predicted enhanced image, so that the target image processing model can implicitly acquire knowledge from the enhanced image.
  • the step of training the target image processing model based on the initial decoded image features, the initial decoded classification features, the target decoded image features, the target decoded classification features, the predicted enhanced image, and the enhanced image includes:
  • the target image processing model is trained based on the segmentation loss function, the classification loss function, and the reconstruction loss function.
  • the similarity between the obtained predicted enhanced image and the enhanced image is calculated to obtain the reconstruction loss function corresponding to the image prediction task.
  • a reconstruction loss function is obtained, and then the segmentation loss function, the classification loss function, and the reconstruction loss function are used to train the target image processing model.
  • the target image processing model provided in the embodiments of this specification, by reconstructing and enhancing the image, and utilizing the predicted enhanced image obtained from the reconstruction and the reconstruction loss function of the enhanced image, enables the target image processing model to learn and capture more fine-grained features that may not be obvious in the initial image. In this way, the target image processing model can gain a deeper understanding of the complex features in the initial image.
  • the reconstruction branch provides a strong additional constraint, prompting the target image processing model not only to match the output or feature representation of the reference image processing model, but also to actually generate a high-quality enhanced image. This multi-task learning strategy can improve the generalization ability and prediction performance of the target image processing model.
  • the initial image and enhanced image of the target object can be obtained through a client.
  • the target image processing model can be deployed on the client according to the client's actual needs, or model interface information for interaction between the client and the target image processing model can be provided to the client.
  • the specific implementation is as follows:
  • the determination of the initial image and enhanced image of the target object includes:
  • the method further includes:
  • the target image processing model or the model interface information corresponding to the target image processing model is sent to the client.
  • the model interface information can be understood as containing detailed information on how to interact with the target image processing model, including data format, calling method, configuration options, and error handling, to ensure that the target image processing model can be effectively and accurately integrated into various applications and services.
  • the client can send the initial image and the enhanced image of the target object to the server.
  • the server Upon receiving the initial image and the enhanced image of the target object, the server processes them using the aforementioned target image processing model training method to train and obtain the target image processing model.
  • the trained target image processing model can be sent to the client for local deployment on the client.
  • the target image processing model can be deployed on the server (which can also be in the cloud), and the corresponding model interface information can be sent to the client. This allows the client to interact with the target image processing model through the model interface information. For example, when the client sends a plain CT image to the server, the server returns the segmentation image and classification result corresponding to the plain CT image to the client.
  • the target image processing model training method provided in the embodiments of this specification allows the target image processing model to be deployed on the client or the server as a model providing platform to provide model services to the client, thus saving the client's computing resources.
  • the target image processing model training method uses paired initial and enhanced images to train a higher-accuracy reference image processing model.
  • This reference image processing model captures rich information present in the enhanced image.
  • the knowledge learned from the reference image processing model is transferred to the lower-accuracy target image processing model trained using the mask image corresponding to the initial image. This achieves the transfer of privileged information from the reference image processing model to the target image processing model.
  • the distillation constraints can be enhanced and the robustness of the target image processing model can be improved, thereby increasing the accuracy of detecting and segmenting target objects using the initial image.
  • FIG. 3 shows a flowchart of an image processing method provided in one embodiment of this specification, the method specifically includes the following steps.
  • Step 302 Determine the target image and input the target image into the target image processing model.
  • the target image processing model is obtained by training using the target image processing model training method described above.
  • Step 304 Use the target image processing model to encode and decode the target image to obtain the target image features and target classification features.
  • the target image can be understood as an initial image containing the target object (i.e., the initial image in the above embodiments).
  • the target image can be understood as a plain CT image.
  • the initial image containing the target object which is unprocessed, is input into the target image processing model.
  • the target image processing model contains an encoder and a decoder, the target image is encoded and decoded to obtain target image features for image segmentation and target classification features for image classification.
  • the target image processing model includes an encoder and a decoder.
  • the encoder is composed of multiple encoding layers
  • the decoder is composed of multiple decoding layers.
  • the encoded image features obtained through encoding are decoded using multiple decoding layers to obtain the target image features corresponding to each decoding layer, thereby further obtaining target classification features.
  • the step of using the target image processing model to encode and decode the target image to obtain the target image features and target classification features includes:
  • the target image is encoded using the multiple encoding layers to obtain encoded image features
  • the encoded image features are decoded using the multiple decoding layers to obtain the target image features corresponding to the multiple decoding layers, and the target classification features are determined based on the target image features corresponding to the multiple decoding layers.
  • encoded image features can be understood as compact and information-rich feature representations obtained by encoding the target image through multiple encoding layers; target image features can be understood as image representations at different levels of abstraction output by the decoding layer; and target classification features can be understood as features related to image classification determined based on the target image features corresponding to the decoding layer.
  • each decoding layer corresponds to a target image feature, and the output of each decoding layer can be regarded as an image representation at different levels of abstraction. The closer the decoding layer is to the output, the closer its features are to the original spatial structure of the image.
  • the features of the decoded target image are analyzed to extract the parts most relevant to image classification. This can include a combination of global or local features for classifying the initial image.
  • the image processing method provided in the embodiments of this specification encodes and decodes the target image through multiple encoding layers and multiple decoding layers, which can obtain more granular target image features corresponding to each decoding layer.
  • the target classification features can also contain rich information about the target object.
  • the step of determining the target classification features based on the target image features corresponding to the multiple decoding layers includes:
  • the target image features corresponding to the final decoding layer among the multiple decoding layers are determined as the target classification features; or,
  • the target image features corresponding to each decoding layer in the multiple decoding layers are subjected to convolution and pooling processing to obtain the key image features corresponding to each decoding layer.
  • the key image features corresponding to each decoding layer are then fused to obtain the target classification features.
  • features are extracted from each decoding layer, convolution is performed on the extracted features, and global max pooling is applied to extract representative features (i.e., key image features) from the convolutional features. These representative features are then concatenated, and the concatenated features (i.e., target classification features) are input into the fully connected layer to obtain the classification result.
  • the target image processing model training method provided in the embodiments of this specification when extracting features in each decoding layer, is beneficial for fine-grained segmentation and also helps with classification. Through this stitching method, it is possible to capture details at the local level and integrate contextual information at the overall level, thereby improving the overall performance and interpretability of the target image processing model.
  • Step 306 Using the features of the target image, obtain the segmented image corresponding to the target image.
  • Step 308 Using the target classification features, obtain the classification result corresponding to the target image.
  • the image processing method provided in the embodiments of this specification when the target image processing model is obtained through the above-described target image processing model training method, enables the target image processing model to acquire target image features and target classification features containing rich information about the target image, thereby improving the accuracy of the segmented image and classification results corresponding to the target image through the target image processing model.
  • This specification also provides a computer-aided diagnosis method for cancer, specifically including:
  • the CT image is input into the CT image processing model, and the CT image is processed by the CT image processing model to obtain the CT segmentation image and CT classification result corresponding to the CT image.
  • the CT image processing model is trained by the above-mentioned target image processing model training method.
  • a detection result is obtained to determine whether a tumor exists in the target detection area.
  • the target detection area can be the area to be detected, such as a certain organ area of the human body
  • the CT image can be the CT image corresponding to the target detection area, including plain CT images and enhanced CT images.
  • the CT image processing model can obtain the CT segmentation image and CT classification result corresponding to the CT image.
  • the detection result of whether there is a tumor in the target detection area can be obtained.
  • computer-aided diagnosis of cancer which is implemented by a computer
  • the purpose is to improve the accuracy of image processing and facilitate the recognition, storage, and transmission of images.
  • the detection results provided by the computer are probability values, which can usually provide a reference for medical staff to accurately diagnose diseases and formulate treatment plans.
  • the detection results obtained from CT segmentation images and CT classification results can be understood as the probability value of whether a tumor exists in the target detection area.
  • Medical staff can use this probability value as a reference for diagnosing diseases. For example, if the probability value is greater than a preset threshold (e.g., the preset threshold is 80%), medical staff can increase the probability of diagnosing a tumor in the target detection area to reduce the risk of misdiagnosis.
  • a preset threshold e.g., the preset threshold is 80%
  • the computer-aided diagnosis method for cancer when the CT image processing model is obtained through the above-described target image processing model training method, can use the CT image processing model to obtain accurate CT segmentation images and CT classification results, thereby improving the accuracy of the detection results of whether there is a tumor in the target detection area based on the CT segmentation images and CT classification results.
  • Figure 4 shows a flowchart of a computer-aided diagnostic method for breast cancer provided in one embodiment of this specification, which specifically includes the following steps.
  • Step 402 Determine the CT image of the breast region
  • Step 404 Input the CT image into the CT image processing model, and use the CT image processing model to process the CT image to obtain the CT segmentation image and CT classification result corresponding to the CT image.
  • the CT image processing model is trained by the above-mentioned target image processing model training method.
  • Step 406 Based on the CT segmentation image and the CT classification result, obtain the detection result of whether there is a tumor in the breast region.
  • the target detection area can be the breast region.
  • the CT image can be a plain CT image or an enhanced image containing organs or tumors.
  • the CT segmentation image can be a CT segmentation image that uses 0 to represent the background, 1 to represent the organ, and 2 to represent the tumor. This CT segmentation image uses different gray values to represent different objects.
  • the CT classification result can be a CT classification result that uses 0 to represent non-cancer and 1 to represent cancer.
  • the results of these two parts can be combined to obtain the detection results for the breast region. These detection results are used to detect whether there is a tumor in the breast region.
  • the computer-aided diagnostic method for breast cancer can process CT images of the breast region using a CT image processing model.
  • plain CT scans used for tumor screening and opportunistic detection it can improve the accuracy of plain CT scans, thereby improving the diagnostic effect for breast cancer.
  • Figure 5 shows a flowchart of an image processing method for a client application in a medical system according to an embodiment of this specification, specifically including the following steps.
  • Step 502 In response to the user's selection operation on the user interface of the client, determine the medical image
  • the medical image can be a plain CT scan of the breast.
  • Step 504 Send the medical image to the server of the medical system, and receive the segmented image and classification result corresponding to the medical image returned by the server.
  • the segmented image and classification result corresponding to the medical image are obtained by processing the medical image according to the target image processing model.
  • the target image processing model is trained by the above-mentioned target image processing model training method.
  • the medical images are sent to the server of the medical system.
  • the server of the medical system has a target image processing model.
  • the target image processing model is used to process the medical images to obtain the corresponding segmented images and classification results.
  • the segmented images and classification results are then returned to the client.
  • Step 506 Display the segmented image and classification results corresponding to the medical image to the user through the user interface.
  • the segmented image and classification results are displayed to the user through the user interface, so that the user can determine the information processing result of the medical image through the segmented image and classification results.
  • the image processing method provided in this specification through the interaction between the client and the server, enables automated processing of medical images.
  • the automated processing and real-time feedback can significantly shorten the time from image upload to result analysis, accelerating the entire diagnostic process.
  • the use of a target image processing model can improve diagnostic accuracy. In regions where plain CT scans are the preferred method for tumor screening, this method can improve the accuracy of plain CT scans, thereby better serving public health.
  • This specification also provides a computer-aided cancer diagnosis system in its embodiments, including a client and a server, wherein,
  • the client is used to send CT images of the target detection area to the server;
  • the server is used to input the CT image into a CT image processing model, process the CT image using the CT image processing model, obtain the segmentation image and classification result corresponding to the CT image, and obtain the detection result of whether there is a tumor in the target detection area based on the CT segmentation image and the CT classification result, and return the detection result to the client.
  • the CT image processing model is trained using the above-mentioned target image processing model training method.
  • the computer-aided cancer diagnosis system provided in the embodiments of this specification achieves automated processing of CT images of the target detection area through interaction between the client and the server, thereby accelerating the entire diagnosis process and improving the accuracy of the detection results by utilizing CT image processing models.
  • Figure 6 shows a structural framework diagram of a target image processing model training method provided in one embodiment of this specification, applied to a medical image diagnosis scenario.
  • enhanced CT scans incorporate additional information that can improve diagnostic performance.
  • AI Artificial Intelligence
  • plain CT scans to improve and enhance the diagnostic performance of plain CT images, which are widely used in routine clinical practice. This requires transferring knowledge gained from enhanced CT scans to plain CT images.
  • a higher-accuracy teacher model i.e., the reference image processing model in the above embodiment
  • This teacher model benefited from a larger training dataset and captured the rich information present in the enhanced CT images.
  • a knowledge distillation technique was employed to transfer the knowledge learned from the teacher model to a lower-accuracy student model trained using plain CT images (i.e., the target image processing model in the above embodiment).
  • the aim is to improve the diagnostic performance of the student models corresponding to plain CT scans.
  • masked image modeling is integrated into the knowledge distillation process. That is, by masking the image blocks of the input student model and recovering features from the teacher model, the distillation constraints can be strengthened.
  • the distillation process is further enhanced, thus enabling the student model to supplement the missing enhanced CT information in the input to some extent.
  • the architecture diagram shown in Figure 6 includes a teacher model and a student model. The following sections will provide a detailed explanation of this architecture diagram in stages.
  • a multi-phase CT image teacher model (the input of the teacher network is both enhanced CT and plain CT, so it is a multi-phase CT image) is constructed to model knowledge of plain CT and enhanced CT, and to assist the learning of the single-phase CT image student model (the input of the student model is plain CT).
  • the teacher model is a U-net structure containing two task-specific branches for image segmentation and image classification tasks.
  • the encoder extracts features from the input images (stitched plain CT and enhanced CT scans), and the encoder and decoder are connected using skip connections. This means that the features of the encoding layer are combined with the upsampled features of the corresponding decoding layer, which can inject the fine spatial information retained in the encoder into the decoder to generate accurate segmented images.
  • Features are extracted from each decoding layer, and convolutional processing is performed on the extracted features.
  • Global max pooling is then applied to extract representative features from the convolutional features, and these representative features are concatenated and input into a fully connected layer to obtain the classification result.
  • the segmentation loss function of the generated segmented image and its label is calculated, and the classification loss function of the obtained classification result and its label is calculated.
  • a teacher network is trained to obtain the network.
  • the goal of multi-task learning for the teacher model is to minimize the following loss function:
  • ⁇ 1 and ⁇ 2 represent the segmentation loss function and the classification loss function, respectively, and are hyperparameters used to balance the importance of segmentation and classification tasks.
  • the labels represent the segmented images corresponding to enhanced CT and plain CT scans; Represents the category label of the entire input; and These are the predicted segmented image and the classification result, respectively; it should be noted that for paired inputs, They share the same values because they represent global labels for the same patient (e.g., the same breast of the same patient), and similarly, They also share the same values.
  • Student model training phase In order to extract privileged information (information included in enhanced CT scans but not in plain CT scans) from the teacher network, feature-level KD loss (represented as(7) is used. Constraints are imposed on intermediate features and the student model learns from the teacher model to guide the student model to inherit contrast enhancement knowledge from the teacher model.
  • the similarity between the features extracted by the first layer of decoder 1 and the features extracted by the first layer of decoder 2 is calculated. Based on the features extracted by the first layer of decoder 2, the features extracted by the second layer of decoder 2 are obtained. The similarity between the features extracted by the second layer of decoder 1 and the features extracted by the second layer of decoder 2 is then calculated. This process continues until the similarity between the features extracted by the fifth layer of decoder 1 and the features extracted by the fifth layer of decoder 2 is calculated.
  • the student model can be better guided to learn different scales and details, capture fine-grained local features, and generate more accurate segmented images.
  • the features that are processed by multiple decoder layers and are finally output by the decoder are called final features.
  • the similarity between the final feature representation of the classification branch extracted by the teacher network and the final feature representation of the classification branch extracted by the teacher network is calculated to obtain a classification result consistent with the classification result output by the teacher network.
  • feature-level knowledge distillation is achieved between the teacher network and the student model.
  • An additional reconstruction branch (containing convolution operations) is integrated into the student model, performing convolution and upsampling on features of each decoder layer of the student model to obtain reconstructed enhanced CT images.
  • the loss function of the reconstructed enhanced CT image and the corresponding enhanced CT image is calculated, and the student model is trained by parameter tuning to obtain the student model.
  • T and S represent the feature representations extracted by the i-th decoding layer (there are n decoding layers in total) of the teacher (T) and student (S) models, respectively, for the segmentation task; and These represent the final feature representations of the classification task extracted by the teacher (T) and student (S) models, respectively (before the fully connected layer); It is a measure of the similarity between two features, such as cosine similarity or mean squared error (MSE).
  • MSE mean squared error
  • Student model inference phase During the inference phase, the student network can make predictions using plain CT images (e.g., cancer detection [cancer vs. non-cancer], tumor segmentation).
  • plain CT images e.g., cancer detection [cancer vs. non-cancer], tumor segmentation.
  • the plain CT scan is input into the student model.
  • the encoder-decoder structure of the student model is used to obtain the segmented image.
  • the segmented image is a grayscale image with the same size as the plain CT scan (the segmented image has 2 or 3 values: 0 represents the background, 1 represents the organ, and 2 represents the malignant tumor; if no tumor is detected/segmented, then there are only two values, 0 and 1).
  • FIG. 7 shows a schematic diagram of the structure of a target image processing model training device provided in one embodiment of this specification. As shown in Figure 7, the device includes:
  • Image determination module 702 is configured to determine the initial image and the enhanced image of the target object
  • the image acquisition module 704 is configured to fuse the initial image and the enhanced image to obtain a fused image, and to perform masking processing on the initial image to obtain a masked image;
  • the initial feature acquisition module 706 is configured to input the fused image into a reference image processing model, and use the reference image processing model to perform encoding and decoding processing on the fused image to obtain initial decoded image features and initial decoded classification features.
  • the target feature acquisition module 708 is configured to input the mask image into the target image processing model, and use the target image processing model to perform encoding and decoding processing on the mask image to obtain target decoded image features and target decoded classification features;
  • the model training acquisition module 710 is configured to train the target image processing model based on the initial decoded image features, the initial decoded classification features, the target decoded image features, and the target decoded classification features.
  • the initial feature acquisition module 706 is further configured to:
  • the fused image is encoded using the multiple coding layers to obtain initial coded image features
  • the initial encoded image features are decoded using the multiple decoding layers to obtain the initial decoded image features corresponding to the multiple decoding layers, and the initial decoded classification features are determined based on the initial decoded image features corresponding to the multiple decoding layers.
  • the initial feature acquisition module 706 is further configured to:
  • the initial decoded image features corresponding to the final decoding layer among the multiple decoding layers are determined as the initial decoding classification features; or,
  • the initial decoded image features corresponding to each of the multiple decoding layers are subjected to convolution and pooling to obtain the first key decoded image features corresponding to each decoding layer.
  • the first key decoded image features corresponding to each decoding layer are then fused to obtain the initial decoding classification features.
  • the target feature acquisition module 708 is further configured to:
  • the mask image is encoded using the multiple encoding layers to obtain the target encoded image features
  • the target encoded image features are decoded using the multiple decoding layers to obtain the target decoded image features corresponding to the multiple decoding layers, and the target decoding classification features are determined based on the target decoded image features corresponding to the multiple decoding layers.
  • the target feature acquisition module 708 is further configured to:
  • the target decoded image features corresponding to the final decoding layer among the multiple decoding layers are determined as the target decoding classification features; or,
  • the target decoded image features corresponding to each of the multiple decoding layers are subjected to convolution and pooling processing to obtain the second key decoded image features corresponding to each decoding layer.
  • the second key decoded image features corresponding to each decoding layer are then fused to obtain the target decoding classification features.
  • model training acquisition module 710 is further configured to:
  • the target image processing model is trained based on the segmentation loss function and the classification loss function.
  • the device further includes:
  • the reconstruction module is configured to perform convolution and upsampling processing on the target decoded image features according to the image prediction task to obtain predicted decoded image features, and obtain a predicted enhanced image based on the predicted decoded image features.
  • model training acquisition module 710 is further configured to:
  • the target image processing model is trained based on the initial decoded image features, the initial decoded classification features, the target decoded image features, the target decoded classification features, the predicted enhanced image, and the enhanced image.
  • model training acquisition module 710 is further configured to:
  • the target image processing model is trained based on the segmentation loss function, the classification loss function, and the reconstruction loss function.
  • the device further includes:
  • the initial result acquisition module is configured to obtain an initial segmented image corresponding to the initial image based on the initial decoded image features according to the image segmentation task; and to obtain an initial classification result corresponding to the initial image based on the initial decoded classification features according to the image classification task.
  • the device further includes:
  • the target result acquisition module is configured to obtain the target segmentation result corresponding to the initial image by utilizing the target decoded image features according to the image segmentation task; and to obtain the target classification result corresponding to the initial image by utilizing the target decoded classification features according to the image classification task.
  • the device further includes:
  • the reference model training module is configured to determine the initial image sample, enhanced image sample, image label, and classification label of the target object sample; fuse the initial image sample and the enhanced image sample to obtain a fused image sample, and input the fused image sample into the reference image processing model; use the reference image processing model to obtain the predicted segmentation image and predicted classification result corresponding to the initial image sample; and train the reference image processing model based on the predicted segmentation image, the image label, the predicted classification result, and the classification label.
  • the image determination module 702 is further configured to:
  • the device further includes:
  • the sending module is configured to send the target image processing model or the model interface information corresponding to the target image processing model to the client.
  • One embodiment of this specification provides a target image processing model training apparatus.
  • the reference image processing model can obtain initial decoded image features and initial decoded classification features containing rich information about the target object.
  • a mask image corresponding to the initial image is input into the target image processing model to obtain target decoded image features and target decoded classification features, and the target image processing model is trained using the initial decoded image features, initial decoded classification features, target decoded image features, and target decoded classification features
  • the target image processing model can process the mask image and recover its features by referring to the initial decoded image features and initial decoded classification features.
  • the above is a schematic scheme of a target image processing model training device according to this embodiment. It should be noted that the technical solution of this target image processing model training device and the technical solution of the target image processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the target image processing model training device, please refer to the description of the technical solution of the target image processing model training method described above.
  • Figure 8 shows a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification. As shown in Figure 8, the apparatus includes:
  • the determination module 802 is configured to determine the target image and input the target image into the target image processing model
  • the feature acquisition module 804 is configured to use the target image processing model to perform encoding and decoding processing on the target image to obtain the target image features and target classification features of the target image.
  • the image acquisition module 806 is configured to obtain a segmented image corresponding to the target image using the features of the target image;
  • the result acquisition module 808 is configured to use the target classification features to obtain the classification result corresponding to the target image.
  • the feature acquisition module 804 is further configured to:
  • the target image is encoded using the multiple encoding layers to obtain encoded image features
  • the encoded image features are decoded using the multiple decoding layers to obtain the target image features corresponding to the multiple decoding layers, and the target classification features are determined based on the target image features corresponding to the multiple decoding layers.
  • the feature acquisition module 804 is further configured to:
  • the target image features corresponding to the final decoding layer among the multiple decoding layers are determined as the target classification features; or,
  • the target image features corresponding to each decoding layer in the multiple decoding layers are subjected to convolution and pooling processing to obtain the key image features corresponding to each decoding layer.
  • the key image features corresponding to each decoding layer are then fused to obtain the target classification features.
  • the image processing apparatus when the target image processing model is obtained through the above-described target image processing model training method, can acquire target image features and target classification features containing rich information about the target image, thereby improving the accuracy of the segmented image and classification results corresponding to the target image through the target image processing model.
  • Figure 9 shows a structural block diagram of a computing device 900 according to one embodiment of this specification.
  • the components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920.
  • the processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.
  • the computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960.
  • networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet.
  • Access device 940 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
  • NIC network interface card
  • the aforementioned components of the computing device 900, as well as other components not shown in FIG. 9, may be connected to each other, for example, via a bus. It should be understood that the block diagram of the computing device shown in FIG. 9 is merely for illustrative purposes and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
  • the computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs).
  • the computing device 900 can also be a mobile or stationary server.
  • the processor 920 is used to execute the following computer program/instructions, which, when executed by the processor, implement the steps of the above-described target image processing model training method.
  • the various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments.
  • the computing device embodiments are basically similar to the target image processing model training method and various image processing method embodiments, so the description is relatively simple. Relevant parts can be referred to the descriptions of the target image processing model training method and various image processing method embodiments.
  • An embodiment of this specification also provides a computer-readable storage medium storing a computer program/instructions that, when executed by a processor, implement the steps of the above-described target image processing model training method and various image processing methods.
  • An embodiment of this specification also provides a computer program product, including a computer program/instructions that, when executed by a processor, implement the steps of the above-described target image processing model training method and multiple image processing methods.
  • the computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms.
  • the computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Medical Informatics (AREA)
  • Evolutionary Computation (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Databases & Information Systems (AREA)
  • Quality & Reliability (AREA)
  • Radiology & Medical Imaging (AREA)
  • Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biodiversity & Conservation Biology (AREA)
  • Biomedical Technology (AREA)
  • Molecular Biology (AREA)
  • Image Analysis (AREA)

Abstract

本说明书实施例提供目标图像处理模型训练方法、图像处理方法,该目标图像处理模型训练方法包括,确定目标对象的初始图像、增强图像;将初始图像以及增强图像进行融合获得融合图像以及对初始图像进行掩码处理获得掩码图像;将融合图像输入参考图像处理模型,利用参考图像处理模型对融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征;将掩码图像输入目标图像处理模型,利用目标图像处理模型对掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征;根据初始解码图像特征、初始解码分类特征、目标解码图像特征、目标解码分类特征,训练目标图像处理模型;提升利用初始图像对目标对象进行检测和分割的准确性。

Description

目标图像处理模型训练方法、图像处理方法
本申请要求于2024年6月26日提交中国专利局、申请号为202410843483.5、发明名称为“目标图像处理模型训练方法、图像处理方法”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本说明书实施例涉及计算机技术领域,特别涉及目标图像处理模型训练方法、图像处理方法;本说明书一个或者多个实施例同时涉及一种目标图像处理模型训练装置,图像处理方法,一种计算设备,一种计算机可读存储介质以及一种计算机程序产品。
背景技术
随着乳腺癌筛查的普及,越来越多的人接受平扫CT(NCCT,Non-Contrast Computed Tomography,非对比度计算机断层扫描)检查。然而,NCCT检查的结果可能不如增强CT(CECT,Contrast-Enhanced Computed Tomography,对比增强计算机断层扫描)检查准确,即利用平扫CT进行检查的情况下可能会影响早期诊断和治疗。
因此,亟需一种方法来提高平扫CT检查的诊断效果,提高平扫CT数据中肿瘤的检测敏感性和特异性。
发明内容
有鉴于此,本说明书实施例提供了一种目标图像处理模型训练方法。本说明书一个或者多个实施例同时涉及一种目标图像处理模型训练装置,图像处理方法,一种计算设备,一种计算机可读存储介质以及一种计算机程序产品,以解决现有技术中存在的、利用平扫CT获得的检查结果存在不准确的技术缺陷。
根据本说明书实施例的第一方面,提供了一种目标图像处理模型训练方法,包括:
确定目标对象的初始图像、增强图像;
将所述初始图像以及所述增强图像进行融合,获得融合图像,以及对所述初始图像进行掩码处理,获得掩码图像;
将所述融合图像输入参考图像处理模型,利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征;
将所述掩码图像输入目标图像处理模型,利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征;
根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型。
根据本说明书实施例的第二方面,提供了一种图像处理方法,包括:
确定目标图像,将所述目标图像输入目标图像处理模型;
利用所述目标图像处理模型对所述目标图像进行编码解码处理,获得目标图像的目标图像特征以及目标分类特征;
利用所述目标图像特征,获得所述目标图像对应的分割图像;
利用所述目标分类特征,获得所述目标图像对应的分类结果。
根据本说明书实施例的第三方面,提供了一种癌症的计算机辅助诊断方法,包括:
确定目标检测区域的CT图像;
将所述CT图像输入CT图像处理模型,利用所述CT图像处理模型对所述CT图像进行处理,获得所述CT图像对应的CT分割图像以及CT分类结果,其中,所述CT图像处理模型通过上述目标图像处理模型训练方法训练获得;
根据所述CT分割图像以及所述CT分类结果,获得所述目标检测区域是否存在肿瘤的检测结果。
根据本说明书实施例的第四方面,提供了一种乳腺癌的计算机辅助诊断方法,包括:
确定乳腺区域的CT图像;
将所述CT图像输入CT图像处理模型,利用所述CT图像处理模型对所述CT图像进行处理,获得所述CT图像对应的CT分割图像以及CT分类结果,其中,所述CT图像处理模型通过上述目标图像处理模型训练方法训练获得;
根据所述CT分割图像以及所述CT分类结果,获得所述乳腺区域是否存在肿瘤的检测结果。
根据本说明书实施例的第五方面,提供了另一种图像处理方法,应用于医疗系统的客户端,包括:
响应于用户针对所述客户端的用户交互界面的点选操作,确定医疗图像;
将所述医疗图像发送至所述医疗系统的服务端,接收所述服务端返回的所述医疗图像对应的分割图像以及分类结果,其中,所述医疗图像对应的分割图像以及分类结果,根据目标图像处理模型对所述医疗图像进行处理获得,所述目标图像处理模型通过上述目标图像处理模型训练方法训练获得;
将所述医疗图像对应的分割图像以及分类结果通过所述用户交互界面展示给所述用户。
根据本说明书实施例的第六方面,提供了一种癌症的计算机辅助诊断系统,包括客户端和服务端,其中,
所述客户端,用于向所述服务端发送目标检测区域的CT图像;
所述服务端,用于将所述CT图像输入CT图像处理模型,利用所述CT图像处理模型对所述CT图像进行处理,获得所述CT图像对应的分割图像以及分类结果,并根据所述CT分割图像以及所述CT分类结果,获得所述目标检测区域是否存在肿瘤的检测结果,并将所述检测结果返回至所述客户端,其中,CT图像处理模型通过上述目标图像处理模型训练方法训练获得。
根据本说明书实施例的第七方面,提供了一种目标图像处理模型训练装置,包括:
图像确定模块,被配置为确定目标对象的初始图像、增强图像;
图像获得模块,被配置为将所述初始图像以及所述增强图像进行融合,获得融合图像,以及对所述初始图像进行掩码处理,获得掩码图像;
初始特征获得模块,被配置为将所述融合图像输入参考图像处理模型,利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征;
目标特征获得模块,被配置为将所述掩码图像输入目标图像处理模型,利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征;
模型训练获得模块,被配置为根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型。
根据本说明书实施例的第八方面,提供了一种图像处理装置,包括:
确定模块,被配置为确定目标图像,将所述目标图像输入目标图像处理模型;
特征获得模块,被配置为利用所述目标图像处理模型对所述目标图像进行编码解码处理,获得目标图像的目标图像特征以及目标分类特征;
图像获得模块,被配置为利用所述目标图像特征,获得所述目标图像对应的分割图像;
结果获得模块,被配置为利用所述目标分类特征,获得所述目标图像对应的分类结果。
根据本说明书实施例的第九方面,提供了一种计算设备,包括:
存储器和处理器;
所述存储器用于存储计算机程序/指令,所述处理器用于执行所述计算机程序/指令,该计算机程序/指令被处理器执行时实现上述目标图像处理模型训练方法、图像处理方法的步骤。
根据本说明书实施例的第十方面,提供了一种计算机可读存储介质,其存储有计算机程序/指令,该计算机程序/指令被处理器执行时实现上述目标图像处理模型训练方法、图像处理方法的步骤。
根据本说明书实施例的第十一方面,提供了一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现上述目标图像处理模型训练方法、图像处理方法的步骤。
本说明书一个或多个实施例提供了目标图像处理模型训练方法,在将初始图像以及增强图像融合输入参考图像处理模型的情况下,利用参考图像处理模型能够获得、包含目标对象的丰富信息的初始解码图像特征以及初始解码分类特征,在将初始图像对应的掩码图像输入到目标图像处理模型,获得目标解码图像特征以及目标解码分类特征,并利用初始解码图像特征、初始解码分类特征、目标解码图像特征、目标解码分类特征,训练目标图像处理模型的情况下,能够在目标图像处理模型针对掩码图像进行处理,使掩码图像恢复特征时,能够参考初始解码图像特征以及初始解码分类特征进行恢复,从而更好地捕捉目标对象详细的特征级信息(即上述的初始解码图像特征以及初始解码分类特征),实现将从参考图像处理模型学到的知识转移到目标图像处理模型中,后续在将目标对象的初始图像输入目标图像处理模型的情况下,能够提升目标图像处理模型利用初始图像针对目标对象进行检测和分割时的准确性。
附图说明
图1是本说明书一个实施例提供的一种图像处理方法的场景示意图;
图2是本说明书一个实施例提供的一种目标图像处理模型训练方法的流程图;
图3是本说明书一个实施例提供的一种图像处理方法的流程图;
图4是本说明书一个实施例提供的一种乳腺癌的计算机辅助诊断方法的流程图;
图5是本说明书一个实施例提供的一种应用于医疗系统的客户端的、一种图像处理方法的流程图;
图6是本说明书一个实施例提供的一种目标图像处理模型训练方法、应用于医学影像诊断场景的结构框架图;
图7是本说明书一个实施例提供的一种目标图像处理模型训练装置的结构示意图;
图8是本说明书一个实施例提供的一种图像处理装置的结构示意图;
图9是本说明书一个实施例提供的一种计算设备的结构框图。
具体实施方式
在下面的描述中阐述了很多具体细节以便于充分理解本说明书。但是本说明书能够以很多不同于在此描述的其它方式来实施,本领域技术人员可以在不违背本说明书内涵的情况下做类似推广,因此本说明书不受下面公开的具体实施的限制。
在本说明书一个或多个实施例中使用的术语是仅仅出于描述特定实施例的目的,而非旨在限制本说明书一个或多个实施例。在本说明书一个或多个实施例和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。还应当理解,本说明书一个或多个实施例中使用的术语“和/或”是指并包含一个或多个相关联的列出项目的任何或所有可能组合。
应当理解,尽管在本说明书一个或多个实施例中可能采用术语第一、第二等来描述各种信息,但这些信息不应限于这些术语。这些术语仅用来将同一类型的信息彼此区分开。例如,在不脱离本说明书一个或多个实施例范围的情况下,第一也可以被称为第二,类似地,第二也可以被称为第一。取决于语境,如在此所使用的词语“如果”可以被解释成为“在……时”或“当……时”或“响应于确定”。
此外,需要说明的是,本说明书一个或多个实施例所涉及的用户信息(包括但不限于用户设备信息、用户个人信息等)和数据(包括但不限于用于分析的数据、存储的数据、展示的数据等),均为经用户授权或者经过各方充分授权的信息和数据,并且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准,并提供有相应的操作入口,供用户选择授权或者拒绝。
首先,对本说明书一个或多个实施例涉及的名词术语进行解释。
知识蒸馏:一种模型压缩技术,旨在将一个复杂模型(教师模型)学到的知识迁移到一个轻量级模型(学生模型)上,以实现模型压缩和加速,这种方法通过较小化教师模型和学生模型输出之间的差异来优化学生模型。
FKD:Feature-level Knowledge Distillation,特征级知识蒸馏,一种在深度学习领域中用于模型压缩和知识转移的技术。它通过使学生模型的中间层特征向教师模型对应层的特征靠近,以此来传递教师模型的隐含知识。
masked image modeling:遮掩图像建模,一种自监督学习方法,具体的,算法会随机地“遮盖”(mask)输入图像的部分区域,然后让模型去预测这些被遮盖部分的内容。这种方法鼓励模型理解和学习图像的全局上下文和局部细节,从而提升其对图像内容的综合理解能力。
在本说明书中,提供了一种目标图像处理模型训练方法,本说明书同时涉及一种目标图像处理模型训练装置,图像处理方法,图像处理装置,一种计算设备,一种计算机可读存储介质以及一种计算机程序产品,在下面的实施例中逐一进行详细说明。
参见图1,图1示出了根据本说明书一个实施例提供的一种图像处理方法的场景示意图。
本说明书实施例提供的图像处理方法可以应用于不同场景,例如在图像处理方法应用于自动驾驶场景的情况下,可以实现针对交通图像中的行人、车辆或交通标志进行检测识别;在图像处理方法应用于医疗场景、具体针对医学影像分析的情况下,可以实现针对医学影像中的器官、肿瘤或病变区域进行检测识别。
对本说明书实施例提供的图像处理方法应用于医疗场景中为例,具体对平扫CT图像中的肿瘤进行检测的场景进行详细说明。
具体的,该图像处理方法通过应用端侧设备102以及服务器104实现,端侧设备102用于向服务器104发送平扫CT图像,如平扫CT图像为一张乳腺平扫CT图像。
在服务器104中训练获得目标图像处理模型,如图1所示,该目标图像处理模型包括编码层、解码层以及全连接层;从而在服务器104接收到端侧设备102发送的平扫CT图像的情况下,将平扫CT图像输入目标图像处理模型,利用目标图像处理模型的编码层、解码层针对平扫CT图像进行编码以及解码处理,获得目标图像特征以及各解码层对应的分类特征,通过图像特征获得平扫CT图像对应的分割图像,该分割图像为与平扫CT图像大小一致的、利用不同灰度值表明不同区域的图像,例如,0代表背景,1代表器官,2代表肿瘤;通过对各解码层对应的分类特征进行卷积以及池化处理,获得各解码层对应的关键分类特征,将这些关键分类特征进行拼接后获得的、目标分类特征输入全连接层,从而获得平扫CT图像对应的分类结果,如0代表非癌症、1代表癌症;将获得的分割图像以及分类结果返回至端侧设备102。
具体的,目标图像处理模型的训练步骤如下所述:确定目标对象的初始图像、增强图像;将所述初始图像以及所述增强图像进行融合,获得融合图像,以及对所述初始图像进行掩码处理,获得掩码图像;将所述融合图像输入参考图像处理模型,利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征;将所述掩码图像输入目标图像处理模型,利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征;根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型。
端侧设备102可以包括浏览器、APP(Application,应用程序)、或网页应用如H5(Hyper Text Markup Language5,超文本标记语言第5版)应用、或轻应用(也被称为小程序,一种轻量级应用程序)或云应用等,端侧设备可以基于服务端提供的相应服务的软件开发工具包(SDK,Software Development Kit),如基于实时通信(RTC,Real Time Communication)SDK开发获得等。端侧设备可以部署在电子设备中,需要依赖设备运行或者设备中的某些APP而运行等。电子设备可以具有显示屏并支持信息浏览等,如可以是个人移动终端如手机、平板电脑、个人计算机等。在电子设备中通常还可以配置各种其它类应用,例如人机对话类应用、模型训练类应用、文本处理类应用、网页浏览器应用、购物类应用、搜索类应用、即时通信工具、邮箱客户端、社交平台软件等。
服务器104可以理解为提供各种服务的服务器,包括物理服务器、云服务器,例如为多个客户端提供通信服务的服务器,又如为客户端上使用的模型提供支持的用于后台训练的服务器,又如对客户端发送的数据进行处理的服务器等。需要说明的是,服务器104可以实现成多个服务器组成的分布式服务器集群,也可以实现成单个服务器。服务器104也可以为分布式系统的服务器,或者是结合了区块链的服务器。服务器104也可以是云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络(CDN,Content Delivery Network)、以及大数据和人工智能平台等基础云计算服务的云服务器,或者是带人工智能技术的智能云计算服务器或智能云主机。
值得说明的是,本说明书实施例中提供的图像处理方法可以由服务器104执行,在本说明书的其它实施例中,可以将目标图像处理模型部署在端侧设备102中,使得端侧设备102也可以与服务器104具有相似的功能,从而执行本说明书实施例所提供的图像处理方法;在其它实施例中,本说明书实施例所提供的图像处理方法还可以是由端侧设备102与服务器104共同执行。
本说明书实施例提供的图像处理方法,通过将图像输入目标图像处理模型,利用目标图像处理模型获得图像的目标图像特征以及目标分类特征,在目标图像处理模型利用包含增强图像信息的特征进行训练的情况下,目标图像处理模型能够更好地捕捉图像中目标对象详细的特征级信息,从而获得更准确的分割图像以及分类结果。
参见图2,图2示出了本说明书一个实施例提供的一种目标图像处理模型训练方法的流程图,具体包括以下步骤。
步骤202:确定目标对象的初始图像、增强图像。
其中,目标对象可以理解为图像中具有特定意义或兴趣的区域、物体或特征,它们是图像分析、图像处理和图像理解等任务中的关键要素;目标对象在不同的应用场景理解不同,比如在交通场景中,目标对象可以理解为行人、车辆或交通标志;在物流场景中,目标对象可以理解为货架、货物或电子标签;在医疗场景中,目标对象可以理解为器官、肿瘤等。
初始图像可以理解为,针对目标对象进行采集(比如拍摄或者扫描)获得的原始图像,初始图像可以通过相机、扫描仪或其他图像捕获设备直接获得;增强图像可以理解为,能够重点突出目标对象的图像,增强图像可以通过针对初始图像进行调整获得。
例如,在医疗场景中,初始图像可以为平扫CT图像,在目标对象为肿瘤的情况下,增强图像可以为增强CT图像,增强CT图像通过对比剂的作用,能够相比平扫CT图像更清晰地显示血管结构、肿瘤、炎症区域或其他病理性改变,因为它们与周围正常组织之间的对比度提高了。在物流场景中,初始图像为巡检机器人采集到的一张货架图像,在目标对象为货物的情况下,增强图像可以为,通过针对初始图像进行锐化,使得该货架图像中、货架上的货物能够更清晰的进行展示的图像。
具体的,确定目标对象的初始图像以及增强图像,增强图像中包含相比初始图像更丰富的、与目标对象相关的特权信息。
步骤204:将所述初始图像以及所述增强图像进行融合,获得融合图像,以及对所述初始图像进行掩码处理,获得掩码图像。
其中,融合图像可以理解为,将初始图像以及增强图像拼接融合、加权融合、泊松融合等方式获得的图像;掩码图像可以理解为,通过对初始图像进行掩码处理、将初始图像中的部分图像像素随机掩蔽起来、获得的图像。
具体的,通过将初始图像以及增强图像进行拼接融合,获得拼接融合后的融合图像,使得后续参考图像处理模型针对融合图像进行处理的情况下,能够利用多模态图像的互补信息,丰富模型的特征学习,增强泛化能力,为后续的知识蒸馏过程提供有力的支持。
例如,将平扫CT图像以及增强CT图像进行通道级的拼接融合,输入参考图像处理模型,能够帮助参考图像处理模型更好地理解平扫CT图像中的复杂特征,以在参考图像处理模型(作为教师模型)将特征级信息转移给目标图像处理模型(作为学生模型)的情况下,后续在目标图像处理模型的输入数据为平扫CT图像时,也能针对平扫CT图像进行合理的处理,从而提高基于平扫CT图像的检测和分割的准确性和有效性。
通过将掩码图像输入目标图像处理模型,能够在基于掩码图像恢复特征时,更多的依据参考图像处理模型对应的特征进行恢复,从而从参考图像处理模型中获得、包含目标对象详细的特征级信息,加强了蒸馏约束。
在本说明书一个实施例中,在目标图像处理模型需要向参考图像处理模型学习,从参考图像处理模型中获得、包含目标对象详细的特征级信息的情况下,参考图像处理模型能够针对包含目标对象的图像进行处理时,获得包含目标对象详细的特征级信息。所述参考图像处理模型的训练步骤如下所述:
确定目标对象样本的初始图像样本、增强图像样本、图像标签、分类标签;
将所述初始图像样本以及所述增强图像样本进行融合,获得融合图像样本,并将所述融合图像样本输入参考图像处理模型;
利用所述参考图像处理模型,获得所述初始图像样本对应的预测分割图像以及预测分类结果;
根据所述预测分割图像、所述图像标签、所述预测分类结果、所述分类标签,训练所述参考图像处理模型。
在本说明书一个实施例中,所述利用所述参考图像处理模型,获得所述初始图像样本对应的预测分割图像以及预测分类结果,包括:
利用所述参考图像处理模型,对所述融合图像样本进行编码解码处理,获得目标解码图像样本特征以及目标解码分类样本特征;
利用所述目标解码图像样本特征,获得所述融合图像样本对应的预测分割图像;
利用所述目标解码分类样本特征,获得所述融合图像样本对应的预测分类结果。
在本说明书一个实施例中,所述参考图像处理模型包括编码器以及解码器,所述编码器通过多个编码层构成,所述解码器通过多个解码层构成;
所述利用所述参考图像处理模型对所述融合图像样本进行编码解码处理,获得目标解码图像样本特征以及目标解码分类样本特征,包括:
利用所述多个编码层对所述融合图像样本进行编码处理,获得初始编码图像样本特征;
利用所述多个解码层对所述初始编码图像样本特征进行解码处理,获得所述多个解码层中各解码层对应的初始解码图像样本特征;
将所述多个解码层中、最终解码层对应的初始解码图像样本特征,确定为目标解码分类样本特征;或者,根据图像分类任务,对所述多个解码层中各解码层对应的初始解码图像样本特征,进行卷积以及池化处理,获得所述各解码层对应的关键图像样本特征,并将所述各解码层对应的关键图像样本特征进行融合,获得目标解码分类样本特征。
具体的,针对参考图像处理模型的训练过程进行详细说明:在对参考图像处理模型进行训练时,将目标对象样本的初始图像样本、增强图像样本进行融合,获得融合图像样本,并将融合图像样本作为输入,输入到参考图像处理模型中;参考图像处理模型的网络架构为一个多任务网络架构,包括一个带有跳跃连接的编码器-解码器结构;参考图像处理模型包含两个特定任务的分支,一个为分割分支,用于获得分割图像,一个为分类分支,用于获得分类结果。
利用参考图像处理模型的编码器、解码器,对融合图像样本进行编码解码处理,即通过编码器中的多个编码层对融合图像样本进行编码处理,获得初始编码图像样本特征,通过解码器中的多个解码层对初始编码图像样本特征进行解码处理,获得多个解码层中各解码层对应的初始解码图像样本特征。
实际应用中,针对分类分支,可以将多个解码层中、最后一个解码层对应的初始解码图像样本特征,确定为目标解码分类样本特征;也可以为了更准确的获得分类结果,从每个解码层中获得对应的初始解码图像样本特征,并针对这些特征进行卷积处理,应用全局较大池化从经过卷积处理后的特征中提取具有代表性的关键图像样本特征,并将各解码层对应的关键图像样本特征进行拼接融合,获得目标解码分类样本特征;通过将目标解码分类样本特征输入全连接层,利用全连接层获得初始图像样本的预测分类结果。
针对分割分支,在编码器中捕捉融合图像样本的高层次特征,得到对融合图像样本内容的高度抽象表示;解码器负责将编码器产生的抽象特征图恢复到原始图像的空间尺寸,在参考图像处理模型为跳跃连接的编码器、解码器结构的情况下,编码器的各层直接连接到相应解码层(通常是同尺度或接近的尺度),直接传递编码阶段的特征图,从而保留了原始图像的细节信息,有助于在最终输出中重建更精细的边界和纹理,通过编码器、解码器的处理,将经过多个解码层处理,解码器最终输出的特征确定为目标解码图像样本特征,从而利用目标解码图像样本特征获得初始图像样本对应的预测分割图像,该预测分割图像利用不同的灰度值对不同对象进行表示,以实现不同对象的分割。
在获得预测分割图像以及预测分类结果的情况下,通过预测分割图像以及图像标签计算分割分支的分割损失函数,通过预测分类结果以及分类标签计算分类分支的分类损失函数;根据该分割损失函数以及分类损失函数,训练参考图像处理模型,使得利用成对的初始图像样本以及增强图像样本训练获得、一个准确度更高的参考图像处理模型,该参考图像处理模型受益于更大的训练数据集,能够捕获增强图像中存在的丰富信息。
在要提高初始图像对应的分割图像、分类结果准确度的情况下,采用知识蒸馏技术,将从参考图像处理模型(教师模型)学到的知识转移到,使用初始图像样本训练的较低准确度的目标图像处理模型(学生模型),目标图像处理模型与参考图像处理模型的模型结构一致,从而提高目标图像处理模型基于初始图像获得的分割图像、分类结果的准确度。
步骤206:将所述融合图像输入参考图像处理模型,利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征。
其中,初始解码图像特征可以理解为,利用参考图像处理模型的解码层获得的解码特征,用于生成初始图像对应的分割图像的特征;该初始解码图像特征可以为每个解码层对应的、用于生成初始图像对应的分割图像的特征,也可以为经过多个解码层、最后一个解码层对应的、用于生成初始图像对应的分割图像的特征;初始解码分类特征可以理解为,利用参考图像处理模型的解码层获得的、用于对初始图像进行分类的特征;该初始解码分类特征可以为每个解码层对应的、用于对初始图像进行分类的特征,也可以为经过多个解码层、最后一个解码层对应的、用于对初始图像进行分类的特征。
具体的,通过将融合图像输入参考图像处理模型,通过参考图像处理模型的编码解码处理,获得融合图像对应的初始解码图像特征以及初始解码分类特征,该初始解码图像特征适用于生成或理解图像的分割,图像分割是指将图像划分为多个区域,每个区域对应于图像中不同的物体或对象类别,因此这类特征有助于识别和分离出图像中的不同对象或区域;该初始解码分类特征则专注于对整个图像或图像中的主体进行分类,即识别图像属于哪一类或具有什么属性;它们是模型用来判断图像内容的关键,比如确定一张图像是否为器官、肿瘤等。
在本说明书一个或多个实施例中,所述参考图像处理模型包括编码器以及解码器,所述编码器通过多个编码层构成,所述解码器通过多个解码层构成;通过多个解码层的处理,能够获得多个解码层对应的初始解码图像特征,从而根据多个解码层对应的初始解码图像特征确定初始解码分类特征。具体实现方式如下所述:
所述利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征,包括:
利用所述多个编码层对所述融合图像进行编码处理,获得初始编码图像特征;
利用所述多个解码层对所述初始编码图像特征进行解码处理,获得所述多个解码层对应的初始解码图像特征,并根据所述多个解码层对应的初始解码图像特征,确定初始解码分类特征。
其中,初始编码图像特征可以理解为,经过多个编码层对融合图像进行编码处理、获得的紧凑且富含信息的特征表示;初始解码图像特征可以理解为,解码层输出的、不同抽象层次的图像表示;初始解码分类特征可以理解为,根据解码层对应的初始解码图像特征、确定的与图像分类相关的特征。
具体的,使用多个编码层对融合图像进行处理,编码层通常涉及逐步降低图像的空间分辨率,同时增加特征表达的抽象层次;这个过程可以看作是对融合图像进行特征提取,通过一系列降维操作,如卷积、池化等,将融合图像转换为一组紧凑且富含信息的特征表示,即初始编码图像特征。
利用多个解码层对编码层获得的初始编码图像特征进行解码,解码层的作用是逆向地将高度压缩的特征信息逐步解码、上采样,恢复到接近初始图像的空间结构;在这个过程中,根据编码得到的初始编码图像特征,通过反卷积、上采样等操作重建图像的部分或完整结构,产生多个解码层对应的初始解码图像特征;每一个解码层的输出都可以视作是不同抽象层次的图像表示,越靠近输出端的解码层,其特征越接近于图像的原始空间结构。
针对多个解码层对应的多个初始解码图像特征进一步分析,实际应用中,对解码后的初始解码图像特征进行分析,从中提取与图像分类最相关的部分,可以包括全局或局部特征的组合,用于针对初始图像的类别判断。
本说明书实施例提供的目标图像处理模型训练方法,通过多个编码层、多个解码层对融合图像进行编码解码处理,能够获得每个解码层对应的、更精细的初始解码图像特征,从而后续利用初始解码图像特征确定初始解码分类特征的情况下,也能获得包含丰富信息的初始解码分类特征,便于提高训练后的目标图像处理模型的性能,提高图像分割、以及图像分类的准确性。
在本说明书一个或多个实施例中,可以根据实际需求,针对初始解码图像特征进行不同的处理,从而获得符合不同需求的初始解码分类特征。具体实现方式如下所述:
所述根据所述多个解码层对应的初始解码图像特征,确定初始解码分类特征,包括:
将所述多个解码层中最终解码层对应的初始解码图像特征,确定为初始解码分类特征;或者,
根据图像分类任务,对所述多个解码层中各解码层对应的初始解码图像特征,进行卷积以及池化处理,获得所述各解码层对应的第一关键解码图像特征,并将所述各解码层对应的第一关键解码图像特征进行融合,获得初始解码分类特征。
其中,最终解码层可以理解为,多个解码层中的最后一个解码层,在每个解码层均基于上一个解码层输出的特征进行处理的情况下,该最终解码层输出的特征为整个解码过程的最终产物,包含了模型对融合图像的深入理解和预测所需的关键信息。第一关键解码图像特征,可以理解为各解码层对应的初始解码图像特征中、针对图像分类具有代表性的特征。
图像分类任务可以理解为,根据图像中的目标对象,将图像进行分类,划分为不同类型的图像的任务;例如在医疗场景中,图像分类任务的目标可以是根据图像中的目标对象、将图像分为包含肿瘤、不包含肿瘤的图像;在物流场景中,图像分类任务的目标可以是根据图像中的目标对象、将图像分为货物陈列正常、货物陈列异常的图像。
具体的,在想要简化计算、提高效率的情况下,可以将经过多个解码层处理,解码器最终输出的特征,即最后一个解码层输出的初始解码图像特征,确定为初始解码分类特征。
在想要获得更精准的分类结果的情况下,可以根据图像分类任务,从每个解码层中提取对应的初始解码图像特征,即获得各解码层对应的、包含不同层级尺度的初始解码图像特征,并通过卷积处理,进一步提炼各解码层对应的初始解码图像特征,获得对图像分类任务有意义的特征,通过池化操作(本说明书实施例应用全局较大池化),减少空间维度,并保留卷积后的每个初始解码图像特征中最重要的信息,通过池化操作也能减少后续全连接层的参数数量,从而降低过拟合风险,并减少计算负担;获得各解码层对应的第一关键解码图像特征,通过将各解码层对应的第一关键解码图像特征进行拼接融合,获得初始解码分类特征。
本说明书实施例提供的目标图像处理模型训练方法,可以根据实际需求,获得符合不同需求的初始解码分类特征,并在将各解码层对应的第一关键解码图像特征进行融合,获得初始解码分类特征的情况下,能在捕捉局部层面的细节的情况下,整合整体层面的上下文信息,从而获得更准确的结果。
在本说明书一个或多个实施例中,在获得初始解码图像特征以及初始解码分类特征的情况下,获得图像分割任务对应的初始分割图像,获得图像分类任务对应的初始分类结果。具体实现方式如下所述:
所述将所述融合图像输入参考图像处理模型,利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征之后,还包括:
根据图像分割任务,利用所述初始解码图像特征,获得所述初始图像对应的初始分割图像;
根据图像分类任务,利用所述初始解码分类特征,获得所述初始图像对应的初始分类结果。
其中,图像分割任务可以理解为,将图像中的不同对象或不同区域进行分割的任务,图像分割的过程可以看作是对图像中的每个像素进行分类,为每个像素分配一个标签,以表示它所属的类别或对象;例如在本说明书实施例中,目标图像处理模型应用于医疗领域的情况下,图像中的背景处标记为0,器官处标记为1,肿瘤处标记为2。
初始分割图像可以理解为,对初始图像中不同的对象进行分割获得的图像;初始分类结果可以理解为,根据初始图像中的目标对象、确定的分类结果;例如在医疗领域,初始分割图像为利用0表示背景,1表示器官,2表示肿瘤的不同灰度值的图像,初始分类结果包含两种,一种为癌症、一种为非癌症。
实际应用中,根据具体的下游任务,如根据图像分割任务,利用初始解码图像特征即可获得与图像分割任务对应的、初始图像的初始分割图像;根据图像分类任务,利用初始解码分类特征即可获得与图像分类任务对应的、初始图像的初始分类结果。
本说明书实施例提供的目标图像处理模型训练方法,通过参考图像处理模型不仅能够利用初始解码图像特征实现像素级的图像分割任务,又能利用初始解码分类特征进行图像级的图像分类任务,获得多样的、与初始图像相关的结果。
步骤208:将所述掩码图像输入目标图像处理模型,利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征。
其中,目标图像处理模型与参考图像处理模型具有相同的网络架构,即目标图像处理模型同样为一个多任务网络架构,包括一个带有跳跃连接的编码器-解码器结构。
利用与参考图像处理模型获得初始解码图像特征以及初始解码分类特征、类似的方式,目标图像处理模型获得目标解码图像特征以及目标解码分类特征。
目标解码图像特征可以理解为,利用目标图像处理模型的解码层获得的解码特征,用于生成掩码图像对应的分割图像的特征;目标解码分类特征可以理解为,利用目标图像处理模型的解码层获得的、用于对掩码图像进行分类的特征。
在本说明书一个或多个实施例中,所述目标图像处理模型包括编码器以及解码器,所述编码器通过多个编码层构成,所述解码器通过多个解码层构成;
所述利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征,包括:
利用所述多个编码层对所述掩码图像进行编码处理,获得目标编码图像特征;
利用所述多个解码层对所述目标编码图像特征进行解码处理,获得所述多个解码层对应的目标解码图像特征,并根据所述多个解码层对应的目标解码图像特征,确定目标解码分类特征。
在本说明书一个或多个实施例中,所述根据所述多个解码层对应的目标解码图像特征,确定目标解码分类特征,包括:
将所述多个解码层中最终解码层对应的目标解码图像特征,确定为目标解码分类特征;或者,
根据图像分类任务,对所述多个解码层中各解码层对应的目标解码图像特征,进行卷积以及池化处理,获得所述各解码层对应的第二关键解码图像特征,并将所述各解码层对应的第二关键解码图像特征进行融合,获得目标解码分类特征。
在本说明书一个或多个实施例中,所述将所述掩码图像输入目标图像处理模型,利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征之后,还包括:
根据图像分割任务,利用所述目标解码图像特征,获得所述初始图像对应的目标分割结果;
根据图像分类任务,利用所述目标解码分类特征,获得所述初始图像对应的目标分类结果。
具体实现方式与上述参考图像处理模型针对融合图像进行处理的过程类似,在此不再赘述。
步骤210:根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型。
具体的,针对目标图像处理模型进行训练时,为了从参考图像处理模型中提取特权信息,通过特征级知识蒸馏(FKD)损失,较小化参考图像处理模型与目标图像处理模型的中间特征,即较小化初始解码图像特征与目标解码图像特征之间的差异,以及较小化初始解码分类特征与目标解码分类特征之间的差异,以指导目标图像处理模型从参考图像处理模型中继承、增强图像中的知识。
在本说明书一个或多个实施例中,通过特征级的知识蒸馏,计算初始解码图像特征以及目标解码图像特征之间的损失函数,计算初始解码分类特征以及目标解码分类特征之间的损失函数,训练目标图像处理模型。具体实现方式如下所述:
所述根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型,包括:
根据所述初始解码图像特征以及所述目标解码图像特征,获得分割损失函数;
根据所述初始解码分类特征以及所述目标解码分类特征,获得分类损失函数;
根据所述分割损失函数以及所述分类损失函数,训练所述目标图像处理模型。
其中,分割损失函数以及分类损失函数可以利用特征相似度衡量,比如利用余弦相似度或均方误差获得。
实际应用中,由于图像分割任务需要捕捉细粒度的局部特征,因此对参考图像处理模型每个解码层对应的初始解码图像特征、以及目标图像处理模型每个解码层对应的目标解码图像特征进行相似度比较,获得每个解码层的分割损失函数;从而可以更好地指导目标图像处理模型(学生模型)学习不同尺度和细节信息。
在图像分类任务更关注全局特征的情况下,可以将参考图像处理模型中利用多个解码层处理,解码器最终输出的特征作为初始解码分类特征、将目标图像处理模型中利用多个解码层处理,解码器最终输出的特征作为目标解码分类特征,通过在最终特征表示(解码器最终输出的特征)上进行相似度比较即可,获得分类损失函数。
通过图像分割任务的分割损失函数,以及图像分类任务的分类损失函数训练目标图像处理模型。
本说明书实施例提供的目标图像处理模型训练方法,在图像分类任务需要在每个解码层逐层指导以获取多尺度信息、确保细节信息的学习的情况下,获得每个解码层上初始解码图像特征与目标解码图像特征的分割损失函数;而在图像分类任务通过最终特征表示可以有效表达全局语义信息的基础上,通过最终特征表示获得分类损失函数,可以简化计算、提升效率。
在本说明书一个或多个实施例中,通过在目标图像处理模型中引入额外的重建分支,利用解码层对应的目标解码图像特征,重建增强图像,进一步加强蒸馏过程,从而获得目标图像处理模型的输入中缺失的增强图像信息。具体实现方式如下所述:
所述根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型之前,还包括:
根据图像预测任务,对所述目标解码图像特征进行卷积以及上采样处理,获得预测解码图像特征,并根据所述预测解码图像特征,获得预测增强图像;
所述根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型,包括:
根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征、所述预测增强图像、所述增强图像,训练所述目标图像处理模型。
其中,预测解码图像特征可以理解为,包含生成预测增强图像所需的具体细节和结构信息的特征。
具体的,通过对各解码层对应的目标解码图像特征进行卷积处理,进一步提炼和组合目标解码图像特征,并利用上采样操作增加特征图分辨率、放大特征图的尺寸,使其接近或恢复到原始图像的尺寸,为生成高分辨率的预测图像做准备,且能够恢复图像的空间细节,获得预测解码图像特征;利用预测解码图像特征,生成预测增强图像,该预测增强图像应相比初始图像包含更丰富的、与目标对象相关的信息。
在获得了预测增强图像的情况下,训练目标图像处理模型时,利用初始解码图像特征、初始解码分类特征、目标解码图像特征、目标解码分类特征、预测增强图像、增强图像对目标图像处理模型进行训练。
即在额外增加了重建分支的情况下,在上述训练目标图像处理模型的基础上,增加预测增强图像、增强图像对目标图像处理模型的训练。
本说明书实施例提供的目标图像处理模型训练方法,将额外的重建分支集成到目标图像处理模型中,利用该额外的重建分支获得预测增强图像,从而通过预测增强图像以及增强图像进一步增强蒸馏过程,使得目标图像处理模型可以从增强图像中隐形的获取知识。
在本说明书一个或多个实施例中,在获得了预测增强图像的情况下,针对目标图像处理模型进行训练时,需要在获得分割损失函数、分类损失函数的基础上,获得重建损失函数,从而根据这三个损失函数,训练获得目标图像处理模型。具体实现方式如下所述:
所述根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征、所述预测增强图像、所述增强图像,训练所述目标图像处理模型,包括:
根据所述初始解码图像特征、所述目标解码图像特征,获得分割损失函数;
根据所述初始解码分类特征、所述目标解码分类特征,获得分割损失函数;
根据所述预测增强图像、所述增强图像,获得重建损失函数;
根据所述分割损失函数、所述分类损失函数、所述重建损失函数,训练所述目标图像处理模型。
具体的,因为针对同一目标图像存在成对的初始图像、增强图像,在将初始图像对应的掩码图像输入目标图像处理模型的情况下,计算获得的预测增强图像、与增强图像(与初始图像配对的增强图像)之间的相似度,获得图像预测任务对应的重建损失函数。
实际应用中,在获得分割损失函数、分类损失函数的基础上,获得重建损失函数,从而利用分割损失函数、分类损失函数以及重建损失函数训练目标图像处理模型。
本说明书实施例提供的目标图像处理模型,通过重建增强图像,并利用重建获得的预测增强图像、与增强图像的重建损失函数,可以使目标图像处理模型学习和捕捉更多细粒度的特征,这些特征在初始图像中可能不明显,且通过这种方式,目标图像处理模型可以更深入地理解初始图像中的复杂特征;另外,重建分支提供了一个强有力的额外约束,促使目标图像处理模型不仅要匹配参考图像处理模型的输出或特征表示,还需要实际生成高质量的增强图像,通过这种多任务学习策略可以提高目标图像处理模型的泛化能力和预测性能。
在本说明书一个或多个实施例中,目标对象的初始图像、增强图像可以通过客户端获得,并在训练获得目标图像处理模型的情况下,根据客户端的实际需要,将目标图像处理模型部署在客户端,或者提供客户端与目标图像处理模型进行交互的模型接口信息给客户端。具体实现方式如下所述:
所述确定目标对象的初始图像、增强图像,包括:
接收客户端发送的目标对象的初始图像、增强图像;
所述训练所述目标图像处理模型之后,还包括:
将所述目标图像处理模型或者、所述目标图像处理模型对应的模型接口信息发送至所述客户端。
其中,模型接口信息可以理解为,包含如何与目标图像处理模型进行交互的详细信息,包括数据格式、调用方式、配置选项以及错误处理等,以确保目标图像处理模型能被有效且准确地集成到各种应用和服务中。
具体的,客户端可以向服务端发送目标对象的初始图像以及增强图像,服务端在接收到目标对象的初始图像以及增强图像的情况下,通过上述目标图像处理模型训练方法进行处理,训练获得目标图像处理模型。
在服务端获得目标图像处理模型的情况下,可以将训练完成的目标图像处理模型发送至客户端,以便在客户端本地部署目标图像处理模型;或者为节省客户端的计算资源,将目标图像处理模型部署在服务端(该服务端也可以为云端),并将目标图像处理模型对应的模型接口信息发送至客户端,方便客户端通过该模型接口信息实现与目标图像处理模型的交互,例如,在客户端向服务端发送平扫CT图像的情况下,服务端向客户端返回该平扫CT图像对应的分割图像以及分类结果。
本说明书实施例提供的目标图像处理模型训练方法,在训练获得目标图像处理模型的情况下,可以将目标图像处理模型部署在客户端,或者服务端作为模型提供平台,为客户端提供模型服务,节省客户端的计算资源。
本说明书实施例提供的目标图像处理模型训练方法,使用成对的初始图像和增强图像训练了一个更高准确度的参考图像处理模型,该参考图像处理模型捕获了增强图像中存在的丰富信息,采用知识蒸馏技术,将从参考图像处理模型学到的知识转移到、使用初始图像对应的掩码图像训练的较低准确度的目标图像处理模型,实现将参考图像处理模型的特权信息传递给目标图像处理模型,且在利用掩码图像进行训练的基础上,可以增强蒸馏约束并提高目标图像处理模型的鲁棒性,提升利用初始图像针对目标对象进行检测和分割时的准确性。
参见图3,图3示出了本说明书一个实施例提供的一种图像处理方法的流程图,具体包括以下步骤。
步骤302:确定目标图像,将所述目标图像输入目标图像处理模型。
具体的,该目标图像处理模型通过上述目标图像处理模型训练方法训练获得。
步骤304:利用所述目标图像处理模型对所述目标图像进行编码解码处理,获得目标图像的目标图像特征以及目标分类特征。
其中,目标图像可以理解为,包含目标对象的初始图像(即上述实施例中的初始图像),在医疗领域,目标图像可以理解为平扫CT图像。
实际应用中,将未经处理的、包含目标对象的初始图像输入到目标图像处理模型中,并在目标图像处理模型包含编码器、解码器的情况下,针对目标图像进行编码解码处理,获得用于图像分割的目标图像特征,以及用于图像分类的目标分类特征。
在本说明一个或多个实施例中,所述目标图像处理模型包括编码器以及解码器,所述编码器通过多个编码层构成,所述解码器通过多个解码层构成;利用多个解码层对编码获得的编码图像特征进行解码处理,获得各解码层对应的目标图像特征,从而进一步的获得目标分类特征。具体实现方式如下所述:
所述利用所述目标图像处理模型对所述目标图像进行编码解码处理,获得目标图像的目标图像特征以及目标分类特征,包括:
利用所述多个编码层对所述目标图像进行编码处理,获得编码图像特征;
利用所述多个解码层对所述编码图像特征进行解码处理,获得所述多个解码层对应的目标图像特征,并根据所述多个解码层对应的目标图像特征,确定目标分类特征。
其中,编码图像特征可以理解为,经过多个编码层对目标图像进行编码处理、获得的紧凑且富含信息的特征表示;目标图像特征可以理解为,解码层输出的、不同抽象层次的图像表示;目标分类特征可以理解为,根据解码层对应的目标图像特征、确定的与图像分类相关的特征。
具体的,使用多个编码层对目标图像进行编码处理,将目标图像转换为一组紧凑且富含信息的特征表示,即编码图像特征;利用多个解码层对编码层获得的编码图像特征进行解码,恢复到接近目标图像的空间结构;在这个过程中,每个解码层均对应一个目标图像特征,每一个解码层的输出都可以视作是不同抽象层次的图像表示,越靠近输出端的解码层,其特征越接近于图像的原始空间结构。
实际应用中,对解码后的目标图像特征进行分析,从中提取与图像分类最相关的部分,可以包括全局或局部特征的组合,用于针对初始图像的类别判断。
本说明书实施例提供的图像处理方法,通过多个编码层、多个解码层对目标图像进行编码解码处理,能够获得每个解码层对应的、更细粒度的目标图像特征,从而后续利用目标解码图像特征确定目标分类特征的情况下,目标分类特征中也能包含针对目标对象的丰富信息。
在本说明书一个或多个实施例中,可以根据实际需求,针对目标图像特征进行不同的处理,从而获得符合不同需求的目标分类特征。具体实现方式如下所述:
所述根据所述多个解码层对应的目标图像特征,确定目标分类特征,包括:
将所述多个解码层中最终解码层对应的目标图像特征,确定为目标分类特征;或者,
根据图像分类任务,对所述多个解码层中各解码层对应的目标图像特征,进行卷积以及池化处理,获得所述各解码层对应的关键图像特征,并将所述各解码层对应的关键图像特征进行融合,获得目标分类特征。
具体实现可参见上述实施例,在此不再赘述。
实际应用中,对每个解码层中的特征进行提取,针对提取出的特征作卷积处理,并应用全局较大池化从经过卷积处理后的特征中提取具有代表性的特征(即关键图像特征),将这些具备代表性的特征进行拼接,将拼接后的特征(即目标分类特征)输入全连接层,获得分类结果。
本说明书实施例提供的目标图像处理模型训练方法,在对每个解码层中的特征进行提取的情况下,有利于细粒度的分割且对分类也有帮助,通过这种拼接的方式,能够捕捉局部层面的细节和整合整体层面的上下文信息,提高了目标图像处理模型的整体性能和可解释性。
步骤306:利用所述目标图像特征,获得所述目标图像对应的分割图像。
步骤308:利用所述目标分类特征,获得所述目标图像对应的分类结果。
本说明书实施例提供的图像处理方法,在目标图像处理模型通过上述目标图像处理模型训练方法获得的情况下,目标图像处理模型能够获取到包含目标图像丰富信息的目标图像特征以及目标分类特征,进而通过目标图像处理模型能够提高目标图像对应的分割图像、分类结果的准确性。
本说明书实施例还提供一种癌症的计算机辅助诊断方法,具体包括:
确定目标检测区域的CT图像;
将所述CT图像输入CT图像处理模型,利用所述CT图像处理模型对所述CT图像进行处理,获得所述CT图像对应的CT分割图像以及CT分类结果,其中,所述CT图像处理模型通过上述目标图像处理模型训练方法训练获得;
根据所述CT分割图像以及所述CT分类结果,获得所述目标检测区域是否存在肿瘤的检测结果。
具体的,目标检测区域可以为待检测的区域,比如人体的某个器官区域,CT图像可以为目标检测区域对应的CT图像,包括平扫CT图像、增强CT图像。
将CT图像输入CT图像处理模型,在CT图像处理模型通过上述目标图像处理模型训练方法训练获得的情况下,能够通过该CT图像处理模型获得CT图像对应的CT分割图像以及CT分类结果,从而根据CT分割图像以及CT分类结果获得目标检测区域是否存在肿瘤的检测结果。
实际应用中,由计算机实现该癌症的计算机辅助诊断方法,由计算机等具有图像处理能力的装置实施的涉及诊断的图像处理方法,是为了提高图像处理的准确率,方便图像的识别、存储和传输,由计算机提供的检测结果为概率值,通常能为医护人员准确诊断疾病和制定治疗方案提供参考。
即通过上述内容可知,根据CT分割图像以及CT分类结果获得的检测结果可以理解为,目标检测区域是否存在肿瘤的概率值,医护人员可以将该概率值作为诊断疾病的参考,例如在该概率值大于预设阈值(例如预设阈值为80%)的情况下,医护人员可以增加确诊目标检测区域存在肿瘤的概率,以减少误诊的风险。
本说明书实施例提供的癌症的计算机辅助诊断方法,在CT图像处理模型通过上述目标图像处理模型训练方法获得的情况下,能够利用该CT图像处理模型获得准确的、CT分割图像以及CT分类结果,从而根据CT分割图像以及CT分类结果,提高目标检测区域是否存在肿瘤的检测结果的准确性。
参见图4,图4示出了本说明书一个实施例提供的一种乳腺癌的计算机辅助诊断方法的流程图,具体包括以下步骤。
步骤402:确定乳腺区域的CT图像;
步骤404:将所述CT图像输入CT图像处理模型,利用所述CT图像处理模型对所述CT图像进行处理,获得所述CT图像对应的CT分割图像以及CT分类结果,其中,所述CT图像处理模型通过上述目标图像处理模型训练方法训练获得。
步骤406:根据所述CT分割图像以及所述CT分类结果,获得所述乳腺区域是否存在肿瘤的检测结果。
以目标检测区域为乳腺区域为例,对乳腺癌的计算机辅助诊断方法进行说明,具体的,在应用于医学影像诊断场景的情况下,目标检测区域可以为乳腺区域,此时CT图像可以为包含器官或肿瘤的平扫CT图像或者增强图像,对应地,CT分割图像可以为,利用0代表背景、1代表器官、2代表肿瘤的CT分割图像,该CT分割图像通过不同的灰度值对不同的对象进行表示;CT分类结果可以为,利用0代表非癌症、1代表癌症的CT分类结果。
实际应用中,在获得CT分割图像以及CT分类结果的情况下,即可综合这两部分的结果,获得针对该乳腺区域的检测结果,该检测结果用于检测该乳腺区域是否存在肿瘤。
具体实现可参见上述实施例,在此不再赘述。
本说明书实施例提供的乳腺癌的计算机辅助诊断方法,利用CT图像处理模型可以针对乳腺区域的CT图像进行处理,在用于针对肿瘤筛查和机会性发现的平扫CT检查中,可以提高平扫CT检查的准确性,从而提高针对乳腺癌的诊断效果。
参见图5,图5示出了本说明书一个实施例提供的一种应用于医疗系统的客户端的、一种图像处理方法的流程图,具体包括以下步骤。
步骤502:响应于用户针对所述客户端的用户交互界面的点选操作,确定医疗图像;
实际应用中,用户(通常是医护人员)通过客户端在用户交互界面的点选操作,比如点击用户交互界面上的上传按钮,上传医疗图像,或者直接在用户交互界面中选择医疗图像;该医疗图像可以为乳腺平扫CT图像。
步骤504:将所述医疗图像发送至所述医疗系统的服务端,接收所述服务端返回的所述医疗图像对应的分割图像以及分类结果,其中,所述医疗图像对应的分割图像以及分类结果,根据目标图像处理模型对所述医疗图像进行处理获得,所述目标图像处理模型通过上述目标图像处理模型训练方法训练获得。
具体的,将医疗图像发送至医疗系统的服务端,在医疗系统的服务端部署有目标图像处理模型,利用目标图像处理模型对医疗图像进行处理,获得医疗图像对应的分割图像以及分类结果,并将医疗图像对应的分割图像以及分类结果返回至客户端。
步骤506:将所述医疗图像对应的分割图像以及分类结果通过所述用户交互界面展示给所述用户。
具体的,在接收到服务端发送的、医疗图像对应的分割图像以及分类结果的情况下,通过用户交互界面将分割图像以及分类结果展示给用户,以使用户通过该分割图像以及分类结果确定该医疗图像的信息处理结果。
本说明书实施例提供的图像处理方法,通过客户端与服务端的交互,实现针对医疗图像的自动化处理,并且自动化处理与即时反馈能够大幅缩短从图像上传到结果分析的时间,加快了整个诊断流程,且利用目标图像处理模型能够提升诊断精确度;在将平扫CT检查作为肿瘤筛选优选方案的地区,该方法可以提高平扫CT检查的准确性,从而更好地服务于公众健康。
本说明书实施例中还提供一种癌症的计算机辅助诊断系统,包括客户端和服务端,其中,
所述客户端,用于向所述服务端发送目标检测区域的CT图像;
所述服务端,用于将所述CT图像输入CT图像处理模型,利用所述CT图像处理模型对所述CT图像进行处理,获得所述CT图像对应的分割图像以及分类结果,并根据所述CT分割图像以及所述CT分类结果,获得所述目标检测区域是否存在肿瘤的检测结果,并将所述检测结果返回至所述客户端,其中,CT图像处理模型通过上述目标图像处理模型训练方法训练获得。
以上述的医疗图像为目标检测区域的CT图像为例,具体实现可参见上述实施例,在此不再赘述。
本说明书实施例提供的癌症的计算机辅助诊断系统,通过客户端与服务端的交互,实现针对目标检测区域CT图像的自动化处理,加快了整个诊断流程,并利用CT图像处理模型提高检测结果的准确性。
参见图6,图6示出了本说明书一个实施例提供的一种目标图像处理模型训练方法、应用于医学影像诊断场景的结构框架图。
实际应用中,增强CT包含了能够增强诊断性能的附加信息的认识。然而,仅仅依赖于增强CT学习到的AI(Artificial Intelligence,人工智能)模型并不能推广和泛化到平扫CT,为针对常规临床实践中大量使用的平扫CT图像的诊断性能进行改进和提升;将从增强CT获得的知识转移至平扫CT图像。
具体的,使用成对的平扫CT和增强CT图像训练了一个更高准确度的教师模型(即上述实施例中的参考图像处理模型),这个教师模型受益于更大的训练数据集,并捕获了增强CT图像中存在的丰富信息。然而,由于两种CT期像之间的固有差异,它不能直接推广到平扫CT图像。因此采用知识蒸馏技术,将从教师模型学到的知识转移到、使用平扫CT图像训练的较低准确度的学生模型(即上述实施例中的目标图像处理模型)。
通过从增强CT以及平扫CT对应的教师模型中提取知识,旨在提高平扫CT对应的学生模型的诊断性能,为了通过蒸馏过程更好地捕捉详细的特征级信息,将遮掩图像建模(masked image modeling)整合到知识蒸馏过程中,即通过对输入学生模型的图像块进行遮掩处理,并从教师模型中恢复特征,能够加强蒸馏约束。
此外,通过在学生模型中引入额外的重建分支,利用解码器获得的特征重建增强CT图像,进一步加强了蒸馏过程,从而学生模型能够一定程度上补充输入中缺失的增强CT信息。
具体的在如图6所示的架构图中,包括教师模型以及学生模型,下面分阶段的对该架构图进行详细说明。
教师模型训练阶段:实际应用中,构建一个多期像CT(教师网络的输入是既有增强CT,又有平扫CT,所以是多期像CT)教师模型,用于平扫CT以及增强CT知识建模,辅助单期像CT(学生模型的输入为平扫CT)学生模型的学习。
获得成对的平扫CT以及增强CT,将成对的平扫CT以及增强CT进行拼接后、输入教师模型中,目标是使教师模型能够指导鲁棒学生模型的学习过程,改进学生模型对平扫CT图像的预测。
教师模型为U-net结构,包含两个任务特定的分支,用于图像分割任务以及图像分类任务,具体的,利用编码器提取输入图像(拼接后的平扫CT以及增强CT)中的特征,并且编码器与解码器使用跳跃连接(skip connections)进行连接,即将编码层的特征与相应解码层的上采样特征相结合,可以将编码器中保留的精细空间信息注入到解码器中,生成准确的分割图像。
对每个解码层中的特征进行提取,针对提取出的特征作卷积处理,应用全局较大池化从经过卷积处理后的特征中提取具有代表性的特征,并将这些具备代表性的特征拼接输入全连接层,获得分类结果。计算生成的分割图像、与分割图像标签的分割损失函数,计算获得的分类结果与分类标签的分类损失函数,从而根据图像损失函数以及分类损失函数,训练获得教师网络。
对教师模型的多任务学习的目标,是较小化以下损失函数:
其中,分别代表分割损失函数和分类损失函数,λ1和λ2是用来平衡分割和分类任务重要性的超参数;表示增强CT、平扫CT对应的分割图像标签;表示整个输入的分类标签;分别是预测的分割图像和分类结果;需要说明的是,针对成对输入,共享相同的值,因为它们代表同一患者(比如患者的同一乳房)的全局标签,同样地,也共享相同的值。
学生模型训练阶段:为了从教师网络中提取特权信息(增强CT包含的、平扫CT不包含的信息),通过使用特征级KD损失(表示为)对中间特征施加、学生模型向教师模型学习的约束,以指导学生模型从教师模型中继承对比度增强的知识。
从训练教师网络的训练集中,获得平扫CT,并进行掩码处理,将掩码后的掩码CT输入学生模型;学生模型与教师网络是相同的编码器-解码器结构;学生模型的解码器针对编码器提取的特征进行上采样时,利用教师网络中解码器的特征进行指导,即在教师网络的解码器为解码器1、学生模型的解码器为解码器2的情况下,针对分割分支,计算解码器1中每层提取出的特征、与解码器2中每层提取出的特征的相似度。
具体的,以解码器包含5层为例,计算解码器1中第1层提取出的特征、与解码器2中第1层提取出的特征的相似度,并依据解码器2中第1层提取出的特征、获得解码器2中第2层提取出的特征,计算解码器1中第2层提取出的特征、与解码器2中第2层提取出的特征的相似度,以此类推,直至计算解码器1中第5层提取出的特征、与解码器2中第5层提取出的特征的相似度,实现通过在每个解码器层进行相似度比较,更好地指导学生模型学习不同尺度和细节信息,捕捉细粒度的局部特征,生成更准确的分割图像。
将经过多个解码器层处理,解码器最终输出的特征称为最终特征,计算教师网络提取出的分类分支的最终特征表示、与教师网络提取出的分类分支的最终特征表示之间的相似度,从而获得与教师网络输出的分类结果一致的分类结果。
通过分割分支针对每个解码器层的特征进行相似度比较,以及分类分支针对解码器的最终特征进行相似度比较,实现教师网络、学生模型之间的特征级知识蒸馏;并将一个额外的重建分支(包含卷积操作)集成到学生模型中,对学生模型的每一个解码器层的特征均进行卷积以及上采样,获得重建增强CT图像;计算该重建增强CT图像、与对应的增强CT图像的损失函数,对学生模型进行调参,训练获得学生模型。
实际应用中,在训练学生模型时,可以将平扫CT输入学生模型中,不过在平扫CT与增强CT表现出相似的外观时,计算损失函数时,从教师模型提取知识到学生模型方面可能仍然缺乏有效性,因此将掩码后的掩码CT输入学生模型。
在将掩码后的掩码CT输入学生模型的情况下,从教师模型中恢复这些掩码的特征,强制知识提取损失从教师模型中提取有价值的信息,从而为知识提取制定了一个强约束。
因此,具体的训练目标可以表示为:
其中,XNC表示平扫图像,表示掩码后的平扫图像,ΘT表示教师模型对应的参数,ΘS表示学生模型对应的参数;分别表示教师(T)和学生(S)模型的第i个解码层(总共有n个解码层)提取的特征表示,用于分割任务;分别表示教师(T)和学生(S)模型的提取的分类任务的最终特征表示(在全连接层前);是衡量两个特征相似性的度量,例如余弦相似度或均方误差(MSE)。
学生模型推理阶段:在推理阶段,学生网络可以使用平扫CT图像进行预测(例如,癌症检测[癌症与非癌症],肿瘤分割)。
将平扫CT输入学生模型,利用学生模型的编码器-解码器结构,获得分割图像,其中分割图像为与平扫CT大小一致的灰度图(分割图像有2个值或者3个值:0代表背景,1器官,2代表恶性肿瘤;如果没有检测/分割到肿瘤的话,那就是只有0和1两个值)。
通过对每个解码层中的特征进行提取,针对提取出的特征作卷积处理,应用全局最大池化从经过卷积处理后的特征中提取具有代表性的特征,并将这些具备代表性的特征拼接输入全连接层,获得分类结果。
本说明书实施例提供的图像处理方法,提出了一种新颖的知识蒸馏框架,将来自经过成对增强CT和平扫CT图像训练的高准确度教师模型的知识转移到、使用平扫CT图像训练的低准确度学生模型;学生模型被鼓励模仿教师网络的行为,以较小化它们输出之间的差异;通过将教师网络的特权知识传递给学生模型,从而提高诊断效果;将掩码图像建模纳入到知识蒸馏中,着重于从未遮掩的CT图像块中蒸馏知识,以加强约束并提高鲁棒性;显著提高了基于平扫CT的乳腺癌检测和分割的性能。
与上述方法实施例相对应,本说明书还提供了目标图像处理模型训练装置实施例,图7示出了本说明书一个实施例提供的一种目标图像处理模型训练装置的结构示意图。如图7所示,该装置包括:
图像确定模块702,被配置为确定目标对象的初始图像、增强图像;
图像获得模块704,被配置为将所述初始图像以及所述增强图像进行融合,获得融合图像,以及对所述初始图像进行掩码处理,获得掩码图像;
初始特征获得模块706,被配置为将所述融合图像输入参考图像处理模型,利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征;
目标特征获得模块708,被配置为将所述掩码图像输入目标图像处理模型,利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征;
模型训练获得模块710,被配置为根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型。
可选地,所述初始特征获得模块706,进一步被配置为:
利用所述多个编码层对所述融合图像进行编码处理,获得初始编码图像特征;
利用所述多个解码层对所述初始编码图像特征进行解码处理,获得所述多个解码层对应的初始解码图像特征,并根据所述多个解码层对应的初始解码图像特征,确定初始解码分类特征。
可选地,所述初始特征获得模块706,进一步被配置为:
将所述多个解码层中最终解码层对应的初始解码图像特征,确定为初始解码分类特征;或者,
根据图像分类任务,对所述多个解码层中各解码层对应的初始解码图像特征,进行卷积以及池化处理,获得所述各解码层对应的第一关键解码图像特征,并将所述各解码层对应的第一关键解码图像特征进行融合,获得初始解码分类特征。
可选地,所述目标特征获得模块708,进一步被配置为:
利用所述多个编码层对所述掩码图像进行编码处理,获得目标编码图像特征;
利用所述多个解码层对所述目标编码图像特征进行解码处理,获得所述多个解码层对应的目标解码图像特征,并根据所述多个解码层对应的目标解码图像特征,确定目标解码分类特征。
可选地,所述目标特征获得模块708,进一步被配置为:
将所述多个解码层中最终解码层对应的目标解码图像特征,确定为目标解码分类特征;或者,
根据图像分类任务,对所述多个解码层中各解码层对应的目标解码图像特征,进行卷积以及池化处理,获得所述各解码层对应的第二关键解码图像特征,并将所述各解码层对应的第二关键解码图像特征进行融合,获得目标解码分类特征。
可选地,所述模型训练获得模块710,进一步被配置为:
根据所述初始解码图像特征以及所述目标解码图像特征,获得分割损失函数;
根据所述初始解码分类特征以及所述目标解码分类特征,获得分类损失函数;
根据所述分割损失函数以及所述分类损失函数,训练所述目标图像处理模型。
所述装置,还包括:
重建模块,被配置为根据图像预测任务,对所述目标解码图像特征进行卷积以及上采样处理,获得预测解码图像特征,并根据所述预测解码图像特征,获得预测增强图像。
可选地,所述模型训练获得模块710,进一步被配置为:
根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征、所述预测增强图像、所述增强图像,训练所述目标图像处理模型。
可选地,所述模型训练获得模块710,进一步被配置为:
根据所述初始解码图像特征、所述目标解码图像特征,获得分割损失函数;
根据所述初始解码分类特征、所述目标解码分类特征,获得分类损失函数;
根据所述预测增强图像、所述增强图像,获得重建损失函数;
根据所述分割损失函数、所述分类损失函数、所述重建损失函数,训练所述目标图像处理模型。
所述装置,还包括:
初始结果获得模块,被配置为根据图像分割任务,利用所述初始解码图像特征,获得所述初始图像对应的初始分割图像;根据图像分类任务,利用所述初始解码分类特征,获得所述初始图像对应的初始分类结果。
所述装置,还包括:
目标结果获得模块,被配置为根据图像分割任务,利用所述目标解码图像特征,获得所述初始图像对应的目标分割结果;根据图像分类任务,利用所述目标解码分类特征,获得所述初始图像对应的目标分类结果。
所述装置,还包括:
参考模型训练模块,被配置为确定目标对象样本的初始图像样本、增强图像样本、图像标签、分类标签;将所述初始图像样本以及所述增强图像样本进行融合,获得融合图像样本,并将所述融合图像样本输入参考图像处理模型;利用所述参考图像处理模型,获得所述初始图像样本对应的预测分割图像以及预测分类结果;根据所述预测分割图像、所述图像标签、所述预测分类结果、所述分类标签,训练所述参考图像处理模型。
可选地,所述图像确定模块702,进一步被配置为:
接收客户端发送的目标对象的初始图像、增强图像。
所述装置,还包括:
发送模块,被配置为将所述目标图像处理模型或者、所述目标图像处理模型对应的模型接口信息发送至所述客户端。
本说明书一个实施例提供了目标图像处理模型训练装置,在将初始图像以及增强图像融合输入参考图像处理模型的情况下,利用参考图像处理模型能够获得、包含目标对象的丰富信息的初始解码图像特征以及初始解码分类特征,在将初始图像对应的掩码图像输入到目标图像处理模型,获得目标解码图像特征以及目标解码分类特征,并利用初始解码图像特征、初始解码分类特征、目标解码图像特征、目标解码分类特征,训练目标图像处理模型的情况下,能够在目标图像处理模型针对掩码图像进行处理,使掩码图像恢复特征时,能够参考初始解码图像特征以及初始解码分类特征进行恢复,从而更好地捕捉目标对象详细的特征级信息(即上述的初始解码图像特征以及初始解码分类特征),实现将从参考图像处理模型学到的知识转移到目标图像处理模型中,后续在将目标对象的初始图像输入目标图像处理模型的情况下,能够提升目标图像处理模型利用初始图像针对目标对象进行检测和分割时的准确性。
上述为本实施例的一种目标图像处理模型训练装置的示意性方案。需要说明的是,该目标图像处理模型训练装置的技术方案与上述的目标图像处理模型训练方法的技术方案属于同一构思,目标图像处理模型训练装置的技术方案未详细描述的细节内容,均可以参见上述目标图像处理模型训练方法的技术方案的描述。
与上述方法实施例相对应,本说明书还提供了图像处理装置实施例,图8示出了本说明书一个实施例提供的一种图像处理装置的结构示意图。如图8所示,该装置包括:
确定模块802,被配置为确定目标图像,将所述目标图像输入目标图像处理模型;
特征获得模块804,被配置为利用所述目标图像处理模型对所述目标图像进行编码解码处理,获得目标图像的目标图像特征以及目标分类特征;
图像获得模块806,被配置为利用所述目标图像特征,获得所述目标图像对应的分割图像;
结果获得模块808,被配置为利用所述目标分类特征,获得所述目标图像对应的分类结果。
可选地,所述特征获得模块804,进一步被配置为:
利用所述多个编码层对所述目标图像进行编码处理,获得编码图像特征;
利用所述多个解码层对所述编码图像特征进行解码处理,获得所述多个解码层对应的目标图像特征,并根据所述多个解码层对应的目标图像特征,确定目标分类特征。
可选地,所述特征获得模块804,进一步被配置为:
将所述多个解码层中最终解码层对应的目标图像特征,确定为目标分类特征;或者,
根据图像分类任务,对所述多个解码层中各解码层对应的目标图像特征,进行卷积以及池化处理,获得所述各解码层对应的关键图像特征,并将所述各解码层对应的关键图像特征进行融合,获得目标分类特征。
本说明书实施例提供的图像处理装置,在目标图像处理模型通过上述目标图像处理模型训练方法获得的情况下,目标图像处理模型能够获取到包含目标图像丰富信息的目标图像特征以及目标分类特征,进而通过目标图像处理模型能够提高目标图像对应的分割图像、分类结果的准确性。
上述为本实施例的一种图像处理装置的示意性方案。需要说明的是,该图像处理装置的技术方案与上述的图像处理方法的技术方案属于同一构思,图像处理装置的技术方案未详细描述的细节内容,均可以参见上述图像处理方法的技术方案的描述。
图9示出了根据本说明书一个实施例提供的一种计算设备900的结构框图。该计算设备900的部件包括但不限于存储器910和处理器920。处理器920与存储器910通过总线930相连接,数据库950用于保存数据。
计算设备900还包括接入设备940,接入设备940使得计算设备900能够经由一个或多个网络960通信。这些网络的示例包括公用交换电话网(PSTN,Public Switched Telephone Network)、局域网(LAN,Local Area Network)、广域网(WAN,Wide Area Network)、个域网(PAN,Personal Area Network)或诸如因特网的通信网络的组合。接入设备940可以包括有线或无线的任何类型的网络接口(例如,网络接口卡(NIC,network interface controller))中的一个或多个,诸如IEEE802.11无线局域网(WLAN,Wireless Local Area Network)无线接口、全球微波互联接入(Wi-MAX,Worldwide Interoperability for Microwave Access)接口、以太网接口、通用串行总线(USB,Universal Serial Bus)接口、蜂窝网络接口、蓝牙接口、近场通信(NFC,Near Field Communication)。
在本说明书的一个实施例中,计算设备900的上述部件以及图9中未示出的其他部件也可以彼此相连接,例如通过总线。应当理解,图9所示的计算设备结构框图仅仅是出于示例的目的,而不是对本说明书范围的限制。本领域技术人员可以根据需要,增添或替换其他部件。
计算设备900可以是任何类型的静止或移动计算设备,包括移动计算机或移动计算设备(例如,平板计算机、个人数字助理、膝上型计算机、笔记本计算机、上网本等)、移动电话(例如,智能手机)、可佩戴的计算设备(例如,智能手表、智能眼镜等)或其他类型的移动设备,或者诸如台式计算机或个人计算机(PC,Personal Computer)的静止计算设备。计算设备900还可以是移动式或静止式的服务器。
其中,处理器920用于执行如下计算机程序/指令,该计算机程序/指令被处理器执行时实现上述目标图像处理模型训练方法的步骤。
本说明书中的各个实施例均采用递进的方式描述,各个实施例之间相同相似的部分互相参见即可,每个实施例重点说明的都是与其他实施例的不同之处。尤其,对于计算设备实施例而言,由于其基本相似于目标图像处理模型训练方法、多种图像处理方法实施例,所以描述的比较简单,相关之处参见目标图像处理模型训练方法、多种图像处理方法实施例的部分说明即可。
本说明书一实施例还提供一种计算机可读存储介质,其存储有计算机程序/指令,该计算机程序/指令被处理器执行时实现上述目标图像处理模型训练方法、多种图像处理方法的步骤。
本说明书中的各个实施例均采用递进的方式描述,各个实施例之间相同相似的部分互相参见即可,每个实施例重点说明的都是与其他实施例的不同之处。尤其,对于计算机可读存储介质实施例而言,由于其基本相似于目标图像处理模型训练方法、多种图像处理方法实施例,所以描述的比较简单,相关之处参见目标图像处理模型训练方法、多种图像处理方法实施例的部分说明即可。
本说明书一实施例还提供一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现上述目标图像处理模型训练方法、多种图像处理方法的步骤。
上述为本实施例的一种计算机程序产品的示意性方案。需要说明的是,该计算机程序产品的技术方案与上述的目标图像处理模型训练方法、多种图像处理方法的技术方案属于同一构思,计算机程序产品的技术方案未详细描述的细节内容,均可以参见上述目标图像处理模型训练方法、多种图像处理方法的技术方案的描述。
上述对本说明书特定实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。
所述计算机指令包括计算机程序代码,所述计算机程序代码可以为源代码形式、对象代码形式、可执行文件或某些中间形式等。所述计算机可读介质可以包括:能够携带所述计算机程序代码的任何实体或装置、记录介质、U盘、移动硬盘、磁碟、光盘、计算机存储器、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、电载波信号、电信信号以及软件分发介质等。需要说明的是,所述计算机可读介质包含的内容可以根据专利实践的要求进行适当的增减,例如在某些地区,根据专利实践,计算机可读介质不包括电载波信号和电信信号。
需要说明的是,对于前述的各方法实施例,为了简便描述,故将其都表述为一系列的动作组合,但是本领域技术人员应该知悉,本说明书实施例并不受所描述的动作顺序的限制,因为依据本说明书实施例,某些步骤可以采用其它顺序或者同时进行。其次,本领域技术人员也应该知悉,说明书中所描述的实施例均属于优选实施例,所涉及的动作和模块并不一定都是本说明书实施例所必须的。
在上述实施例中,对各个实施例的描述都各有侧重,某个实施例中没有详述的部分,可以参见其它实施例的相关描述。
以上公开的本说明书优选实施例只是用于帮助阐述本说明书。可选实施例并没有详尽叙述所有的细节,也不限制该发明仅为所述的具体实施方式。显然,根据本说明书实施例的内容,可作很多的修改和变化。本说明书选取并具体描述这些实施例,是为了更好地解释本说明书实施例的原理和实际应用,从而使所属技术领域技术人员能很好地理解和利用本说明书。本说明书仅受权利要求书及其全部范围和等效物的限制。

Claims (19)

  1. 一种目标图像处理模型训练方法,包括:
    确定目标对象的初始图像、增强图像;
    将所述初始图像以及所述增强图像进行融合,获得融合图像,以及对所述初始图像进行掩码处理,获得掩码图像;
    将所述融合图像输入参考图像处理模型,利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征;
    将所述掩码图像输入目标图像处理模型,利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征;
    根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型。
  2. 根据权利要求1所述的目标图像处理模型训练方法,所述参考图像处理模型包括编码器以及解码器,所述编码器通过多个编码层构成,所述解码器通过多个解码层构成;
    所述利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征,包括:
    利用所述多个编码层对所述融合图像进行编码处理,获得初始编码图像特征;
    利用所述多个解码层对所述初始编码图像特征进行解码处理,获得所述多个解码层对应的初始解码图像特征,并根据所述多个解码层对应的初始解码图像特征,确定初始解码分类特征。
  3. 根据权利要求2所述的目标图像处理模型训练方法,所述根据所述多个解码层对应的初始解码图像特征,确定初始解码分类特征,包括:
    将所述多个解码层中最终解码层对应的初始解码图像特征,确定为初始解码分类特征;或者,
    根据图像分类任务,对所述多个解码层中各解码层对应的初始解码图像特征,进行卷积以及池化处理,获得所述各解码层对应的第一关键解码图像特征,并将所述各解码层对应的第一关键解码图像特征进行融合,获得初始解码分类特征。
  4. 根据权利要求1所述的目标图像处理模型训练方法,所述目标图像处理模型包括编码器以及解码器,所述编码器通过多个编码层构成,所述解码器通过多个解码层构成;
    所述利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征,包括:
    利用所述多个编码层对所述掩码图像进行编码处理,获得目标编码图像特征;
    利用所述多个解码层对所述目标编码图像特征进行解码处理,获得所述多个解码层对应的目标解码图像特征,并根据所述多个解码层对应的目标解码图像特征,确定目标解码分类特征。
  5. 根据权利要求4所述的目标图像处理模型训练方法,所述根据所述多个解码层对应的目标解码图像特征,确定目标解码分类特征,包括:
    将所述多个解码层中最终解码层对应的目标解码图像特征,确定为目标解码分类特征;或者,
    根据图像分类任务,对所述多个解码层中各解码层对应的目标解码图像特征,进行卷积以及池化处理,获得所述各解码层对应的第二关键解码图像特征,并将所述各解码层对应的第二关键解码图像特征进行融合,获得目标解码分类特征。
  6. 根据权利要求1-5任意一项所述的目标图像处理模型训练方法,所述根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型,包括:
    根据所述初始解码图像特征以及所述目标解码图像特征,获得分割损失函数;
    根据所述初始解码分类特征以及所述目标解码分类特征,获得分类损失函数;
    根据所述分割损失函数以及所述分类损失函数,训练所述目标图像处理模型。
  7. 根据权利要求1-5任意一项所述的目标图像处理模型训练方法,所述根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型之前,还包括:
    根据图像预测任务,对所述目标解码图像特征进行卷积以及上采样处理,获得预测解码图像特征,并根据所述预测解码图像特征,获得预测增强图像;
    所述根据所述初始解码图像特征、所述初始解码分类特征、所述目标解码图像特征、所述目标解码分类特征,训练所述目标图像处理模型,包括:
    根据所述初始解码图像特征、所述目标解码图像特征,获得分割损失函数;
    根据所述初始解码分类特征、所述目标解码分类特征,获得分类损失函数;
    根据所述预测增强图像、所述增强图像,获得重建损失函数;
    根据所述分割损失函数、所述分类损失函数、所述重建损失函数,训练所述目标图像处理模型。
  8. 根据权利要求1所述的目标图像处理模型训练方法,所述将所述融合图像输入参考图像处理模型,利用所述参考图像处理模型对所述融合图像进行编码解码处理,获得初始解码图像特征以及初始解码分类特征之后,还包括:
    根据图像分割任务,利用所述初始解码图像特征,获得所述初始图像对应的初始分割图像;
    根据图像分类任务,利用所述初始解码分类特征,获得所述初始图像对应的初始分类结果。
  9. 根据权利要求1所述的目标图像处理模型训练方法,所述将所述掩码图像输入目标图像处理模型,利用所述目标图像处理模型对所述掩码图像进行编码解码处理,获得目标解码图像特征以及目标解码分类特征之后,还包括:
    根据图像分割任务,利用所述目标解码图像特征,获得所述初始图像对应的目标分割结果;
    根据图像分类任务,利用所述目标解码分类特征,获得所述初始图像对应的目标分类结果。
  10. 根据权利要求1所述的目标图像处理模型训练方法,所述参考图像处理模型的训练步骤如下所述:
    确定目标对象样本的初始图像样本、增强图像样本、图像标签、分类标签;
    将所述初始图像样本以及所述增强图像样本进行融合,获得融合图像样本,并将所述融合图像样本输入参考图像处理模型;
    利用所述参考图像处理模型,获得所述初始图像样本对应的预测分割图像以及预测分类结果;
    根据所述预测分割图像、所述图像标签、所述预测分类结果、所述分类标签,训练所述参考图像处理模型。
  11. 根据权利要求1所述的目标图像处理模型训练方法,所述确定目标对象的初始图像、增强图像,包括:
    接收客户端发送的目标对象的初始图像、增强图像;
    所述训练所述目标图像处理模型之后,还包括:
    将所述目标图像处理模型或者、所述目标图像处理模型对应的模型接口信息发送至所述客户端。
  12. 一种图像处理方法,包括:
    确定目标图像,将所述目标图像输入目标图像处理模型;
    利用所述目标图像处理模型对所述目标图像进行编码解码处理,获得目标图像的目标图像特征以及目标分类特征;
    利用所述目标图像特征,获得所述目标图像对应的分割图像;
    利用所述目标分类特征,获得所述目标图像对应的分类结果。
  13. 根据权利要求12所述的图像处理方法,所述目标图像处理模型包括编码器以及解码器,所述编码器通过多个编码层构成,所述解码器通过多个解码层构成;
    所述利用所述目标图像处理模型对所述目标图像进行编码解码处理,获得目标图像的目标图像特征以及目标分类特征,包括:
    利用所述多个编码层对所述目标图像进行编码处理,获得编码图像特征;
    利用所述多个解码层对所述编码图像特征进行解码处理,获得所述多个解码层对应的目标图像特征,并根据所述多个解码层对应的目标图像特征,确定目标分类特征。
  14. 根据权利要求13所述的图像处理方法,所述根据所述多个解码层对应的目标图像特征,确定目标分类特征,包括:
    将所述多个解码层中最终解码层对应的目标图像特征,确定为目标分类特征;或者,
    根据图像分类任务,对所述多个解码层中各解码层对应的目标图像特征,进行卷积以及池化处理,获得所述各解码层对应的关键图像特征,并将所述各解码层对应的关键图像特征进行融合,获得目标分类特征。
  15. 一种癌症的计算机辅助诊断方法,包括:
    确定目标检测区域的CT图像;
    将所述CT图像输入CT图像处理模型,利用所述CT图像处理模型对所述CT图像进行处理,获得所述CT图像对应的CT分割图像以及CT分类结果,其中,所述CT图像处理模型通过权利要求1-10任意一项目标图像处理模型训练方法训练获得;
    根据所述CT分割图像以及所述CT分类结果,获得所述目标检测区域是否存在肿瘤的检测结果。
  16. 一种乳腺癌的计算机辅助诊断方法,包括:
    确定乳腺区域的CT图像;
    将所述CT图像输入CT图像处理模型,利用所述CT图像处理模型对所述CT图像进行处理,获得所述CT图像对应的CT分割图像以及CT分类结果,其中,所述CT图像处理模型通过权利要求1-10任意一项目标图像处理模型训练方法训练获得;
    根据所述CT分割图像以及所述CT分类结果,获得所述乳腺区域是否存在肿瘤的检测结果。
  17. 一种癌症的计算机辅助诊断系统,包括客户端和服务端,其中,
    所述客户端,用于向所述服务端发送目标检测区域的CT图像;
    所述服务端,用于将所述CT图像输入CT图像处理模型,利用所述CT图像处理模型对所述CT图像进行处理,获得所述CT图像对应的分割图像以及分类结果,并根据所述CT分割图像以及所述CT分类结果,获得所述目标检测区域是否存在肿瘤的检测结果,并将所述检测结果返回至所述客户端,其中,CT图像处理模型通过权利要求1-12任意一项目标图像处理模型训练方法训练获得。
  18. 一种计算设备,包括:
    存储器和处理器;
    所述存储器用于存储计算机程序/指令,所述处理器用于执行所述计算机程序/指令,该计算机程序/指令被处理器执行时实现权利要求1-16任意一项所述方法的步骤。
  19. 一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现权利要求1-16任意一项所述方法的步骤。
PCT/CN2025/078597 2024-06-26 2025-02-21 目标图像处理模型训练方法、图像处理方法 Pending WO2026001027A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410843483.5 2024-06-26
CN202410843483.5A CN118397377B (zh) 2024-06-26 2024-06-26 目标图像处理模型训练方法、图像处理方法

Publications (1)

Publication Number Publication Date
WO2026001027A1 true WO2026001027A1 (zh) 2026-01-02

Family

ID=92006011

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/078597 Pending WO2026001027A1 (zh) 2024-06-26 2025-02-21 目标图像处理模型训练方法、图像处理方法

Country Status (2)

Country Link
CN (1) CN118397377B (zh)
WO (1) WO2026001027A1 (zh)

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118397377B (zh) * 2024-06-26 2024-09-13 阿里巴巴(中国)有限公司 目标图像处理模型训练方法、图像处理方法
CN119107223B (zh) * 2024-08-01 2025-09-30 北京达佳互联信息技术有限公司 图像处理模型训练方法、图像处理方法及相关设备
CN118674725B (zh) * 2024-08-22 2024-11-05 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法、计算机辅助癌症预后方法及系统
CN119760627B (zh) * 2024-12-06 2025-11-25 中国科学院信息工程研究所 基于流量多模态特征融合的社交机器人发现方法及系统

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210327054A1 (en) * 2020-04-15 2021-10-21 Siemens Healthcare Gmbh Medical image synthesis of abnormality patterns associated with covid-19
CN116935166A (zh) * 2023-08-10 2023-10-24 Oppo广东移动通信有限公司 模型训练方法、图像处理方法及装置、介质、设备
WO2023241410A1 (zh) * 2022-06-14 2023-12-21 北京有竹居网络技术有限公司 数据处理方法、装置、设备及计算机介质
CN117408948A (zh) * 2023-09-11 2024-01-16 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法、图像分类分割模型的训练方法
CN118247284A (zh) * 2024-05-28 2024-06-25 阿里巴巴达摩院(杭州)科技有限公司 图像处理模型的训练方法、图像处理方法
CN118397377A (zh) * 2024-06-26 2024-07-26 阿里巴巴(中国)有限公司 目标图像处理模型训练方法、图像处理方法

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111583165B (zh) * 2019-02-19 2023-08-08 京东方科技集团股份有限公司 图像处理方法、装置、设备及存储介质
CN110363776B (zh) * 2019-06-28 2021-10-22 联想(北京)有限公司 图像处理方法及电子设备
CN113112559A (zh) * 2021-04-07 2021-07-13 中国科学院深圳先进技术研究院 一种超声图像的分割方法、装置、终端设备和存储介质
WO2022246677A1 (zh) * 2021-05-26 2022-12-01 深圳高性能医疗器械国家研究院有限公司 一种增强ct图像的重建方法
CN115482221A (zh) * 2022-09-22 2022-12-16 深圳先进技术研究院 一种病理图像的端到端弱监督语义分割标注方法
CN117115833A (zh) * 2023-08-29 2023-11-24 平安健康保险股份有限公司 一种证件分类方法、装置、设备及存储介质
CN117237756A (zh) * 2023-09-14 2023-12-15 深圳数联康健智能科技有限公司 一种训练目标分割模型的方法、目标分割方法及相关装置
CN117952920A (zh) * 2024-01-08 2024-04-30 南方医科大学 一种基于深度学习的增强-平扫ct图像合成方法和系统

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210327054A1 (en) * 2020-04-15 2021-10-21 Siemens Healthcare Gmbh Medical image synthesis of abnormality patterns associated with covid-19
WO2023241410A1 (zh) * 2022-06-14 2023-12-21 北京有竹居网络技术有限公司 数据处理方法、装置、设备及计算机介质
CN116935166A (zh) * 2023-08-10 2023-10-24 Oppo广东移动通信有限公司 模型训练方法、图像处理方法及装置、介质、设备
CN117408948A (zh) * 2023-09-11 2024-01-16 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法、图像分类分割模型的训练方法
CN118247284A (zh) * 2024-05-28 2024-06-25 阿里巴巴达摩院(杭州)科技有限公司 图像处理模型的训练方法、图像处理方法
CN118397377A (zh) * 2024-06-26 2024-07-26 阿里巴巴(中国)有限公司 目标图像处理模型训练方法、图像处理方法

Also Published As

Publication number Publication date
CN118397377B (zh) 2024-09-13
CN118397377A (zh) 2024-07-26

Similar Documents

Publication Publication Date Title
WO2026001027A1 (zh) 目标图像处理模型训练方法、图像处理方法
CN115861616B (zh) 面向医学图像序列的语义分割系统
CN113569840A (zh) 基于自注意力机制的表单识别方法、装置及存储介质
Holste et al. Efficient deep learning-based automated diagnosis from echocardiography with contrastive self-supervised learning
Liu et al. Skin lesion segmentation with a multiscale input fusion U-Net incorporating Res2-SE and pyramid dilated convolution
CN116912268B (zh) 一种皮肤病变图像分割方法、装置、设备及存储介质
Sun et al. Automated classification and segmentation and feature extraction from breast imaging data
Liang et al. A spatiotemporal network using a local spatial difference stack block for facial micro-expression recognition
Ma Deep learning‐based image processing for financial audit risk quantification in healthcare
DS et al. G-Net: Implementing an enhanced brain tumor segmentation framework using semantic segmentation design
Martins et al. A review toward deep learning for high dynamic range reconstruction
Singh et al. A Multi-Model Image Enhancement and Tailored U-Net Architecture for Robust Diabetic Retinopathy Grading
Hong et al. Detection of coal gangue based on MSRCR algorithm and improved lightweight YOLOv8n
Sarica et al. Comparative assessment of CNN and transformer U‐Nets in multiple sclerosis lesion segmentation
Kim et al. Concept graph embedding models for enhanced accuracy and interpretability
CN117036750B (zh) 膝关节病灶检测方法及装置、电子设备和存储介质
Ahmed et al. Melanoma Detection Using Augmented ResNet34 for High-Precision Dermoscopic Image Classification
Cahan et al. X-ray2CTPA: leveraging diffusion models to enhance pulmonary embolism classification
WO2024251271A1 (zh) 图像分类方法和系统
Hemalatha et al. Masked and Noise-Masked Multimodal Brain Tumor Image Segmentation Using SegFormer and Shared Encoder Framework
Zhao et al. Effective Algorithm for Biomedical Image Segmentation
CN115631196B (zh) 图像分割方法、模型的训练方法、装置、设备和存储介质
Okila et al. Automated AI‐Based Lung Disease Classification Using Point‐of‐Care Ultrasound
CN118570223B (zh) 多模态核医学图像分割方法、系统、装置及可读存储介质
Su et al. GSCCANet: Dual Decoder Network Fusing Edge Focus and Global Channel Attention for Precise Segmentation of Colonic Polyps

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25824422

Country of ref document: EP

Kind code of ref document: A1