WO2025185337A1 - 图像处理方法、计算机辅助诊断方法、图像处理模型的训练方法 - Google Patents

图像处理方法、计算机辅助诊断方法、图像处理模型的训练方法

Info

Publication number
WO2025185337A1
WO2025185337A1 PCT/CN2025/070596 CN2025070596W WO2025185337A1 WO 2025185337 A1 WO2025185337 A1 WO 2025185337A1 CN 2025070596 W CN2025070596 W CN 2025070596W WO 2025185337 A1 WO2025185337 A1 WO 2025185337A1
Authority
WO
WIPO (PCT)
Prior art keywords
information
sample
image processing
text
image
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/070596
Other languages
English (en)
French (fr)
Other versions
WO2025185337A8 (zh
Inventor
姚佳文
郭广宇
夏英达
莫志榮
郑智琳
吕乐
张灵
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba China Co Ltd
Original Assignee
Alibaba China Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba China Co Ltd filed Critical Alibaba China Co Ltd
Publication of WO2025185337A1 publication Critical patent/WO2025185337A1/zh
Publication of WO2025185337A8 publication Critical patent/WO2025185337A8/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/0002Inspection of images, e.g. flaw detection
    • G06T7/0012Biomedical image inspection
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • G06N3/0455Auto-encoder networks; Encoder-decoder networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/096Transfer learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/26Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/44Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
    • G06V10/443Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components by matching or filtering
    • G06V10/449Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters
    • G06V10/451Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters with interaction between the filter responses, e.g. cortical complex cells
    • G06V10/454Integrating the filters into a hierarchical structure, e.g. convolutional neural networks [CNN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/80Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
    • G06V10/806Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/20ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/70ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10072Tomographic images
    • G06T2207/10081Computed x-ray tomography [CT]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30004Biomedical image processing
    • G06T2207/30092Stomach; Gastric
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30004Biomedical image processing
    • G06T2207/30096Tumor; Lesion
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02DCLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • the present disclosure relates to the field of computer technology, and in particular to an image processing method.
  • Tumors are one of the major factors affecting human health. Identifying tumors in medical images requires professional doctors to identify them based on their experience. Due to the limitations of doctors' experience, image recognition and analysis with the help of medical images has become an important topic.
  • embodiments of the present disclosure provide an image processing method.
  • One or more embodiments of the present disclosure also provide an image processing method, a computer-aided diagnosis method for cancer, a training method for an image processing model, a computer-aided diagnosis method, an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
  • an image processing method including:
  • the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area;
  • the multiple target images are input into an image processing model to obtain detection results corresponding to the target detection area, wherein the image processing model generates detection results corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, and the detection results include detection annotation information, detection category information and detection guidance text.
  • a computer-aided diagnosis method for cancer comprising:
  • CT image processing task carries a plurality of CT images corresponding to a target detection area, and the CT image processing task is used to detect whether a tumor exists in the target detection area;
  • the multiple CT images are input into a CT image processing model to obtain detection results corresponding to the target detection area, wherein the CT image processing model generates detection results corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection results include detection annotation information, detection category information and detection guidance text.
  • a method for training an image processing model is provided, which is applied to a cloud-side device and includes:
  • sample image as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image;
  • the model parameters of the image processing model are sent to the end-side device.
  • a computer-aided diagnosis method comprising:
  • the CT image processing task carries multiple CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormal object in the target detection area;
  • the CT image processing model Inputting the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, wherein the CT image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information, and detection guidance text;
  • a computing device including:
  • the memory is used to store computer-executable instructions
  • the processor is used to execute the computer-executable instructions.
  • the steps of the above-mentioned image processing method, CT image processing method or image processing model training method are implemented.
  • a computer-readable storage medium which stores computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.
  • a computer program product comprising a computer program/instruction, which, when executed by a processor, implements the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.
  • An image processing method provided by one embodiment of the present disclosure uses an image processing model to obtain multiple scale feature information corresponding to a target image, thereby improving the accuracy of subsequent detection results.
  • the generated detection results include the location of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results and providing users with multi-dimensional detection information, thereby enhancing the user experience.
  • FIG1 is an architecture diagram of an image processing system provided by one embodiment of the present disclosure.
  • FIG2 is a flowchart of an image processing method provided by one embodiment of the present disclosure.
  • FIG3 is a schematic diagram of the structure of an image processing model provided by an embodiment of the present disclosure.
  • FIG4 is a flowchart of a CT image processing method provided by one embodiment of the present disclosure.
  • FIG5 is a flowchart of a method for training an image processing model provided by one embodiment of the present disclosure
  • FIG6 is a flowchart of another image processing method provided by an embodiment of the present disclosure.
  • FIG7 is a flowchart of an image processing method applied to an esophageal cancer detection scenario provided by one embodiment of the present disclosure
  • FIG8 is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure.
  • FIG9 is a schematic structural diagram of a CT image processing device provided by one embodiment of the present disclosure.
  • FIG10 is a schematic structural diagram of another image processing device provided by an embodiment of the present disclosure.
  • FIG11 is a structural block diagram of a computing device provided by an embodiment of the present disclosure.
  • first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other.
  • the first may also be referred to as the second, and similarly, the second may also be referred to as the first.
  • word "if” as used herein may be interpreted as "at the time of” or "when” or "in response to determining”.
  • the user information including but not limited to user device information, user personal information, etc.
  • data including but not limited to data used for analysis, stored data, displayed data, etc.
  • the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
  • CT Computer tomography: Computer tomography uses precisely collimated X-ray beams, gamma rays, ultrasound waves, etc., together with highly sensitive detectors to perform cross-sectional scans around a certain part of the human body one after another. It has the characteristics of fast scanning time and clear images, and can be used to detect a variety of diseases.
  • Computer-aided diagnosis refers to the use of imaging, medical image processing technology and other possible physiological and biochemical means, combined with computer analysis and calculation, to assist in the discovery of lesions and improve the accuracy of diagnosis.
  • Esophageal cancer is a highly lethal cancer with a low 5-year survival rate according to incomplete statistics. However, early detection of resectable/curable esophageal cancer can greatly reduce the mortality rate. Among them, lymph node metastasis is a relatively common and typical disease.
  • Tumors are one of the major factors affecting human health. Identifying tumors in medical images requires professional doctors to identify them based on their experience. Due to the limitations of doctors' experience, image recognition and analysis with the help of medical images has become an important topic.
  • an image processing method is provided in the present disclosure.
  • the present disclosure also involves a CT image processing method, a training method for an image processing model, an image processing device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
  • FIG1 shows an architecture diagram of an image processing system provided by an embodiment of the present disclosure.
  • the image processing system may include a client 100 and a server 200;
  • the client 100 is configured to send an image processing task to the server 200, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is configured to detect whether there is an abnormal object in the target detection area;
  • the server 200 is configured to input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection area, wherein the image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, the detection result including detection annotation information, detection category information, and detection guidance text; and send the detection result to the client 100;
  • the client 100 is also used to receive the detection results sent by the server 200.
  • an image processing task is received, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is used to detect whether there are abnormal objects in the target detection area; the multiple target images are input into an image processing model to obtain a detection result corresponding to the target detection area, wherein the image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, and the detection result includes detection annotation information, detection category information and detection guidance text.
  • multi-scale feature information of multiple target images is extracted, and detection annotation information, detection category information and detection guidance text are generated based on the multi-scale feature information.
  • the image processing system may include multiple clients 100 and a server 200.
  • the clients 100 may be referred to as end-side devices, and the server 200 may be referred to as cloud-side devices. Multiple clients 100 may establish communication connections through the server 200.
  • the server 200 provides image processing services to multiple clients 100. Multiple clients 100 may function as either senders or receivers, communicating through the server 200.
  • Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100.
  • a user can publish a data stream to the server 200 through the client 100, and the server 200 generates detection results based on the data stream and pushes the detection results to other clients with which communication has been established.
  • the client 100 and the server 200 are connected via a network.
  • the network provides a medium for the communication link between the client 100 and the server 200.
  • the network can include various connection types, such as wired or wireless communication links or fiber optic cables.
  • the data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.
  • the client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program, a lightweight application), or a cloud application.
  • the client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by the server 200, such as a real-time communication (RTC) SDK.
  • SDK software development kit
  • RTC real-time communication
  • the client 100 can be deployed in an electronic device and rely on the device or certain apps in the device to run.
  • the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, tablet computer, or personal computer.
  • Various other types of applications can also be configured in the electronic device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
  • the server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers for background training that provide support for models used on the clients, and servers that process data sent by the clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server.
  • the server can also be a server of a distributed system, or a server combined with a blockchain.
  • the server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
  • the image processing methods provided in the embodiments of the present disclosure are generally executed by the server.
  • the client may also have similar functions to the server to execute the image processing methods provided in the embodiments of the present disclosure.
  • the image processing methods provided in the embodiments of the present disclosure may also be executed jointly by the client and the server.
  • FIG. 2 shows a flow chart of an image processing method provided by an embodiment of the present disclosure, which specifically includes the following steps:
  • Step 202 Receive an image processing task, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area.
  • the image processing tasks sent by users can be received through the server or the client.
  • an image processing task is a task for detecting the presence of abnormal objects within a target detection area.
  • the image processing task carries multiple target images corresponding to the target detection area.
  • the target detection area can be understood as a subarea used to detect the presence of abnormal objects, and the abnormal object specifically refers to a foreign object within the target detection area.
  • the abnormal object can be a tumor in the human body.
  • the target detection area can be any organ in the human body, such as the liver, lungs, stomach, esophagus, etc.
  • the status of the object to be detected can be further determined based on the prediction results, thereby providing assistance in locating the abnormal object and providing precise treatment.
  • the object to be detected can be understood as the object described in the target detection area.
  • the target image is a CT image of Zhang San's stomach
  • the CT image is the target image
  • the object to be detected is Zhang San.
  • the image processing task is used to detect whether there is a tumor in the stomach area.
  • the object to be detected can be a human or other living organism, and this is not limited to this in one or more specific embodiments provided in this disclosure.
  • the image processing task can be applied to the recognition of various types of medical images, and the determination of whether there are abnormal objects in the target detection area of the medical image based on the image features.
  • the presence of a tumor in the stomach area can be predicted based on the medical image of the stomach area, thereby helping doctors to accurately locate the abnormal area
  • the presence of a tumor in the esophagus can be predicted based on the medical image of the esophagus area, thereby helping doctors to accurately locate the abnormal area and facilitate subsequent treatment.
  • the target image acquired is an image of the esophagus.
  • the multiple target images acquired are CT images of the esophagus. These multiple target images can form a 3D image of the esophageal region.
  • the abnormal object can be understood as a tumor in the esophagus.
  • Multiple CT images corresponding to the esophageal region are acquired and image detection processing is performed on the multiple CT images to detect whether a malignant tumor is present in the esophagus.
  • the abnormal object may be a certain type of cell, a certain type of tissue structure, etc., for example, a malignant tumor, a benign tumor, a hyperplastic tissue, etc. This is not limited in one or more embodiments provided in the present disclosure.
  • multiple target images corresponding to the target detection area carried in the image processing task can be used as input to detect whether there is an abnormal object in the target detection area.
  • the method before receiving the image processing task, the method further includes:
  • the image segmentation task carries a plurality of initial images corresponding to a target detection area, and the image segmentation task is used to extract a target image corresponding to the target detection area;
  • Each initial image is input into a pre-trained image segmentation model to obtain a target image corresponding to each initial image output by the image segmentation model.
  • the target image can be understood as a close-up image corresponding to the target detection area.
  • the multiple images received may include other areas in addition to the target detection area, and these other areas may affect the target detection area. Therefore, the method provided in this disclosure also processes the image.
  • an image segmentation task is first obtained for multiple initial images, where the initial images include both target detection regions and regions that affect the target detection regions.
  • the image segmentation task is used to extract the target images corresponding to the target detection regions from each of the initial images.
  • the image segmentation model is trained to identify the target detection region in the initial image, extract the target detection region from the initial image, and generate a target image corresponding to the target detection region.
  • the method provided herein first acquires an initial CT image.
  • This initial CT image is a plain scan CT image that meets image quality requirements.
  • the initial CT image can be obtained from multiple CT scanners or from a single CT scanner, a fact not limited in this disclosure.
  • the format of each initial CT image must be standardized.
  • the image segmentation model can be 3DUNet.
  • 3DUNet is a deep learning architecture for image segmentation in three-dimensional space. It is an extension of the U-Net model, which was originally designed for semantic segmentation of two-dimensional biomedical images and is widely recognized for its excellent performance and high accuracy in segmenting small objects. 3DUNet applies this concept to three-dimensional datasets, such as medical imaging (CT and MRI scans), which is useful in many medical fields because these data are often rich in three-dimensional structural information.
  • CT and MRI scans three-dimensional datasets
  • 3DUNet includes an encoder-decoder structure for capturing global context information, restoring lost spatial details and generating accurate pixel-level segmentation labels.
  • the 3D convolution kernel in the 3DUNet structure replaces the 2D convolution kernel in the network, and can simultaneously process the features of the input data in three dimensions (length, width, and height).
  • 3DUNet can also effectively fuse three-dimensional features at different levels, which helps to extract complex shape and structure information. This framework has a good effect in medical image segmentation methods.
  • the initial CT image is cut using the preprocessing strategy of the 3DUNet structure to obtain the target image corresponding to the target detection area for subsequent processing.
  • Step 204 Input the multiple target images into the image processing model to obtain the detection results corresponding to the target detection area, wherein the image processing model generates the detection results corresponding to the target detection area based on the multi-scale feature information corresponding to the multiple target images, and the detection results include detection annotation information, detection category information and detection guidance text.
  • multiple target images carried by the image processing task can be obtained from the image processing task, and the multiple target images can be input into the image processing model to obtain the detection results corresponding to the target detection area output by the image processing model.
  • the detection results specifically include detection annotation information, detection category information and detection guidance text.
  • the detection annotation information specifically refers to the area where the abnormal object is marked in the target image. If there is no abnormal object in the target image, the detection annotation information is empty.
  • Detection category information specifically refers to the category information of the object to be detected corresponding to the target detection area for detection of the target image. If the target detection area contains an abnormal object, the detection category information is abnormal; if the target detection area does not contain an abnormal object, the detection category information is normal.
  • the detection guidance text specifically refers to the guidance text for the object to be detected, which is output based on the detection annotation information, detection category information, etc. of the target image, regarding the position, status, and other information of the abnormal object.
  • the detection guidance text provided by the embodiment of the present disclosure is a guidance text generated by adding the detected abnormal object position, status, and other information to the guidance text template. If the detection annotation information determines that there is an abnormal object, the specific location of the abnormal object and the category information of the object to be detected will be given in the detection annotation information.
  • the detection guidance text can be "The patient has a tumor in the upper part of the esophagus, and the patient has esophageal cancer.”
  • the detection guidance text can be "No tumor was detected in the patient's stomach, and the patient was not detected with gastric cancer," and so on.
  • the multi-scale feature information of each target image is extracted within the image processing model, and then the detection results corresponding to the target detection area are extracted based on the multi-scale feature information.
  • the image processing model includes a multi-scale feature extraction module, a feature fusion module, and a feature processing module;
  • Inputting the multiple target images into an image processing model to obtain detection results corresponding to the target detection area includes S2042-S2046:
  • S2042 Input the multiple target images into the multi-scale feature extraction module to obtain at least one scale feature information.
  • the multi-scale feature extraction module specifically refers to using different scale information to extract target feature information of a 3D model composed of multiple target images. For example, for a 3D model with a size of W*H*D, it is input into the multi-scale feature extraction module.
  • Figure 3 shows a structural schematic diagram of an image processing model provided by an embodiment of the present disclosure.
  • multiple target images are input into the multi-scale feature extraction module of the image processing model.
  • multiple initial feature information of different scales is obtained.
  • the initial feature information of multiple different scales is convolved or deconvolved-convolved to obtain scale feature maps corresponding to each scale, thereby obtaining 6 multi-scale feature maps.
  • S2044 Input the feature information of each scale into the feature fusion module to obtain feature fusion information.
  • the multiple scale feature information is input into the feature fusion layer for feature fusion, and the multiple scale feature information is unified into the same scale and fused to obtain feature fusion information.
  • multiple scale feature maps ⁇ F0, F1, F2, F3, F4, F5 ⁇ are input into the feature fusion layer.
  • the multiple scale feature maps are unified into the same scale for feature fusion to obtain fused feature information Fa.
  • S2046 Input the feature fusion information into the feature processing module to obtain a detection result corresponding to the target detection area.
  • the feature fusion information can be input into the feature processing module, where the feature fusion information is processed to generate a detection result corresponding to the target detection area.
  • the detection result specifically includes detection annotation information, detection category information, and detection guidance text.
  • the feature processing module includes an abnormal object segmentation unit, an abnormal object classification unit and a guidance text generation unit;
  • Inputting the feature fusion information into the feature processing module to obtain the detection result corresponding to the target detection area includes:
  • a detection result corresponding to the target detection area is generated according to the detection label information, detection category information, and detection guidance text.
  • the feature processing module is used to process feature fusion information, extract features from the feature fusion information, and decode the extracted features to generate corresponding detection results.
  • the feature processing module specifically includes an abnormal object segmentation unit, an abnormal object classification unit, and a guidance text generation unit.
  • the abnormal object segmentation unit is used to mark the area of the abnormal object in the target image according to the feature fusion information, and segment the abnormal object corresponding to the target detection area in the target image;
  • the abnormal object classification unit is used to determine whether there is an abnormal object in the target detection area according to the feature fusion information, and determine the classification information for the object to be detected;
  • the guidance text generation unit is used to generate guidance text for indicating information such as the position and status of the abnormal object to be detected according to the feature fusion information.
  • the final detection result is generated according to the detection annotation information, detection category information and detection guidance text output by the abnormal object segmentation unit, the abnormal object classification unit and the guidance text generation unit respectively.
  • the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit;
  • Inputting the feature fusion information into the guidance text generation unit to obtain the detection guidance text includes:
  • the detection guidance text specifically includes guidance information on the abnormal location and the status of the object to be detected. Therefore, the guidance text generation unit includes an abnormal location text subunit and an abnormal result text subunit.
  • the abnormal location text subunit is used to generate the location information of the abnormal object in the target detection area based on the feature fusion information, and the abnormal result text subunit is used to generate the status information of the object to be detected.
  • the abnormal location text subunit processes the feature fusion information and determines the location guidance information as "the middle of the stomach.”
  • the abnormal result text subunit also processes the feature fusion information and determines the result guidance information as "the user has stomach cancer.” Based on the location guidance information and the result guidance information, the final detection guidance text is generated: "The user has a tumor in the middle of the stomach and has stomach cancer.”
  • the image processing method provided by the embodiment of the present disclosure inputs multiple target images of the target detection area of the object to be detected into the image processing model.
  • the image processing model multiple scale features are extracted from the 3D images corresponding to the multiple target images. After the multiple scale features are fused, annotation information, category information, and guidance text are generated respectively, thereby finally generating a detection result.
  • the accuracy of the subsequent generation of detection results is improved by the multiple scale feature information.
  • the generated detection results include the location information of the abnormal object, the information of the object to be detected, and the guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.
  • the training data for image processing models is images that include abnormal objects in the target detection area.
  • the processing effect of images that do not include abnormal objects in the target detection area is poor.
  • sample image as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image;
  • the training method of the image processing model provided in the present invention uses the ideas of supervised training and transfer training. First, an initial model is trained using labeled training data, and then the backbone network in the initial model is migrated to the image processing model, and further training is performed with another batch of training data to obtain the final image processing model.
  • Sample images specifically refer to images used for model training. In practical applications, sample images include both images with and without abnormal objects.
  • Sample annotation information corresponding to sample images specifically refers to the annotation information for abnormal objects in the sample images;
  • sample category information specifically refers to the status information of the object to be detected corresponding to the sample images;
  • sample guidance text specifically refers to the text that indicates the location of abnormal objects and the status of the object to be detected in the sample images.
  • Sample category information can be obtained from the sample guidance text or can be separate sample category information.
  • the image processing model After obtaining the sample images, they are fed into the image processing model (at this point, the model is still untrained).
  • the image processing model generates prediction detection results based on each sample image.
  • the prediction detection results include predicted label information, predicted category information, and predicted guidance text.
  • the model structure of the image processing model such as the image processing model in the above steps, also includes a multi-scale feature extraction module, a feature fusion module, and a feature processing module.
  • the data processing process of the sample image in the untrained image processing model is the same as the data processing process of the image processing model in the above embodiment.
  • the data processing process of the sample image in the untrained image processing model please refer to the data processing process of the target image in the image processing model above, which will not be repeated here.
  • the model loss value can be calculated based on the predicted detection results and the sample detection results.
  • the method provided in the present disclosure there are many methods for calculating the model loss value, such as cross entropy loss function, maximum loss function, average loss function, etc.
  • the specific method of the loss function is not limited, and it is subject to actual application.
  • the sample detection result includes sample labeling information and sample category information
  • the predicted detection result includes predicted labeling information and predicted category information. Calculating the model loss value based on the sample labeling information, the sample category information, the predicted labeling information, and the predicted category information includes:
  • a model loss value is calculated based on the first loss value and the second loss value.
  • sample detection results include sample labeling information and sample category information
  • predicted detection results include predicted labeling information and predicted category information.
  • Technicians hope that the predicted detection results are consistent with the sample detection results, thereby improving the accuracy of image processing model predictions.
  • the first loss value is calculated using sample labeling information and predicted labeling information
  • the second loss value is calculated using sample category information and predicted category information.
  • the predicted guidance text and predicted labeling information can also be verified with each other, further improving the accuracy of model prediction.
  • the first loss value and the second loss value are then fused to obtain a model loss value. Specifically, the first loss value and the second loss value are added to obtain the model loss value.
  • the sample guidance text is processed within the image processing model to obtain the text loss value. That is, the sample guidance text needs to be input into the image processing model.
  • the image processing model includes a multi-scale feature extraction module, a feature fusion module, a feature processing module, and a text feature extraction module;
  • the feature fusion information and the sample text feature information are input into the feature processing module to obtain the predicted labeling information, predicted category information and text loss value corresponding to the sample image.
  • the model also includes a text feature extraction module.
  • the sample guidance text is input into the text feature extraction module to obtain sample text feature information.
  • the feature fusion information and sample text feature information are input into the feature processing module. After processing in the feature processing module, predicted label information, predicted category information, and text loss value are obtained.
  • the model parameters of the image processing model are adjusted according to the text loss value and the model loss value. Specifically, the text loss value and the model loss value are added to obtain a new loss value, and the model parameters of the image processing model are adjusted by backpropagation according to the new loss value.
  • the feature processing module includes an abnormal object segmentation unit, an abnormal object classification unit, and a guidance text generation unit;
  • a text loss value is obtained by calculation based on the predicted text feature information and the sample text feature information.
  • the feature processing module in the training phase includes an abnormal object segmentation unit, an abnormal object classification unit, and a guidance text generation unit.
  • the abnormal object segmentation unit is used to generate detection annotation information based on feature fusion information;
  • the abnormal object classification unit is used to generate detection category information based on feature fusion information;
  • the guidance text generation unit is used to generate predicted text feature information based on feature fusion information, specifically generating a predicted text feature vector, and then calculating the text loss value using the predicted text feature information and the sample text feature information.
  • the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit
  • the sample text feature information includes sample position feature information and sample result feature information
  • Inputting the feature fusion information into the guidance text generation unit to obtain predicted text feature information includes:
  • calculating and obtaining a text loss value according to the predicted text feature information and the sample text feature information includes:
  • the text loss value is calculated based on the sample position feature information, the sample result feature information, the predicted position guidance information feature and the predicted result guidance information feature.
  • the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit.
  • the guidance text generation unit first generates predicted text feature information.
  • the predicted text feature information is composed of predicted position guidance information features and predicted result guidance information features.
  • the sample text feature information includes sample position feature information and sample result feature information.
  • a text loss value is calculated based on the sample position feature information, sample result feature information, predicted position guidance information features, and predicted result guidance information features.
  • sample guidance text specifically includes two parts: sample location guidance information and sample result guidance information.
  • Sample location guidance information specifically refers to the location information of abnormal objects within the target detection area, while sample result guidance information specifically refers to the status information of the object to be detected.
  • sample location guidance information may include "a tumor is found in the middle of the esophagus," while sample result guidance information may include "the patient suffers from esophageal cancer.”
  • Sample location feature information specifically refers to the feature vector corresponding to the sample location guidance information
  • sample result feature information specifically refers to the feature vector corresponding to the sample result guidance information.
  • the sample guidance text is input into a text feature extraction module.
  • This text feature extraction module specifically refers to a module that can convert the sample guidance text into corresponding text feature information, such as the encoder of the Transformer model, the text encoder of the CLIP model, etc.
  • the sample position guidance information and sample result guidance information in the sample guidance text are input into the feature extraction module to obtain the sample position guidance information features corresponding to the sample position guidance information and the sample result guidance information features corresponding to the sample result guidance information.
  • the image processing model includes a guidance text generation unit, which includes an abnormal position text sub-unit and an abnormal result text sub-unit. During the processing process, the image processing model will obtain the predicted position guidance information features output by the abnormal position text sub-unit and the predicted result guidance information features output by the abnormal result text sub-unit.
  • the position loss value is calculated based on the sample position guidance information feature and the predicted position guidance information feature, the result loss value is calculated based on the sample result guidance information feature and the predicted result guidance information feature, and then the text loss value is determined based on the position loss value and the result loss value.
  • the sample guidance text is processed by a text feature extraction module.
  • the text feature extraction module can use the sample position guidance information and sample result guidance information of the object to be detected in the sample guidance text to supervise the image processing model, making full use of the existing sample guidance text and improving the training speed and prediction accuracy of the image processing model.
  • the sample image includes a first sample image and a second sample image, wherein the first sample image is marked with sample annotation information, and the second sample image is not marked with sample annotation information;
  • sample image as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image, including:
  • the second sample is input into the image annotation model to obtain predicted annotation information output by the image annotation model, and the predicted annotation information is used as sample annotation information of the second sample image.
  • the first sample image is a sample image with sample annotation information
  • the second sample image is a sample image without sample annotation information. Both the first sample image and the second sample image have corresponding sample guidance texts.
  • the sample category information corresponding to the object to be detected can be identified and extracted in each sample guidance text. After obtaining the sample category information corresponding to each sample image, an image annotation model can be trained using the first sample image with sample annotation information to generate sample annotation information for the second sample image.
  • each sample image includes corresponding sample guidance text
  • the first sample image has sample annotation information
  • the second sample image does not.
  • An image annotation model is pre-trained using the first sample image and the sample annotation information and sample guidance text corresponding to the first sample image.
  • the image annotation model includes an abnormal object segmentation unit and a guidance text generation unit.
  • the first sample image is input into the image annotation model to obtain predicted annotation information output by the abnormal object segmentation unit and predicted guidance text output by the guidance text generation unit.
  • the predicted annotation information, sample annotation information, predicted guidance text, and sample guidance text are used to calculate a model loss value for the image annotation model.
  • the model parameters of the image annotation model are adjusted based on the model loss value until a model training stopping condition is met, thereby obtaining a trained image annotation model that can annotate images that are not labeled with image annotation information.
  • the second sample image is input into the trained image annotation model.
  • the image annotation model can generate predicted annotation information for the second sample image, and use the predicted annotation information as the sample annotation information corresponding to the second sample image.
  • the image annotation model includes an abnormal object segmentation unit and a guidance text generation unit.
  • the image processing model in the above steps also includes an abnormal object segmentation unit and a guidance text generation unit.
  • the method further includes:
  • the abnormal object segmentation unit and the guidance text generation unit in the image annotation model are used as the abnormal object segmentation unit and the guidance text generation unit of the image processing model.
  • the image annotation model is used as a teacher model.
  • the first sample image, the sample annotation information corresponding to the first sample image, and the sample guidance text are utilized.
  • the abnormal object segmentation unit and guidance text generation unit in the obtained image annotation model have been trained and parameterized, and have corresponding data processing capabilities.
  • the abnormal object segmentation unit and guidance text generation unit can be directly migrated to the image processing model, and the abnormal object segmentation unit and guidance text generation unit in the image annotation model can be used as the abnormal object segmentation unit and guidance text generation unit of the image processing model.
  • the image annotation model can be set to also include a multi-scale feature extraction module and a feature fusion module, and continue to be trained during the training process of the image annotation model. Then it is migrated to the image processing model, so that the multi-scale feature extraction module, feature fusion module, abnormal object segmentation unit and guidance text generation unit in the feature processing module in the image processing model are preliminarily trained during the training process of the image annotation model. When the image processing model is trained, it is trained again, or no longer trained, thereby improving the model training efficiency of the image processing model.
  • the image processing model includes a text feature extraction module in the model training stage, and may not include the text feature extraction module in the model application stage of the image processing model.
  • the training method of the image processing model utilizes a text feature extraction module to obtain corresponding sample text feature information from the sample guidance text through the text feature extraction module, and uses the sample text feature information to supervise the training of the image processing model, thereby improving the efficiency and accuracy of the image processing model training.
  • a semi-supervised image labeling model training method is adopted. An image labeling model is trained using some of the labeled sample images, and the image labeling model is used to label the unlabeled sample images, thereby enriching the number of sample images with labeled information.
  • model structure in the image annotation model is migrated to the image processing model, and the image processing model is further trained using the model structure in the trained image annotation model, which improves the training speed of the image processing model and avoids the waste of computing resources.
  • FIG. 4 shows a flow chart of a CT image processing method provided by an embodiment of the present disclosure, specifically comprising the following steps:
  • Step 402 Receive a CT image processing task, wherein the CT image processing task carries multiple CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormal object in the target detection area.
  • the abnormal object may be a tumor.
  • the CT image processing method provided in this embodiment is a computer-aided diagnosis method for cancer, which is used to detect whether a tumor exists in a CT image corresponding to a target detection area.
  • Step 404 Input the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, wherein the CT image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information and detection guidance text.
  • steps 402 to 404 is the same as the implementation of steps 202 to 204 described above, and will not be described in detail in this embodiment of the present disclosure.
  • a CT image processing task is received.
  • the CT image processing task includes multiple CT images corresponding to the target user's stomach.
  • the multiple CT images can form a 3D image of the target user's stomach.
  • the CT image processing task is used to detect whether there is a malignant tumor in the target user's stomach.
  • the CT image processing model is the image processing model in the above embodiment.
  • the model structure of the CT image processing model is the same as the structure of the image processing model in the above embodiment, and will not be repeated here.
  • the CT image processing model generates multiple scale feature information based on the multiple CT images, and then generates tumor annotation information corresponding to the stomach, the user's disease status (suffering from gastric cancer or not), user detection text (tumor location in the user's stomach, whether the user is sick, etc.) and other information based on the multiple scale feature information.
  • the CT image processing method extracts multiple scale features from 3D images corresponding to multiple CT images within an image processing model. After fusing these multiple scale features, the method then generates annotation information, category information, and guidance text, ultimately generating a detection result. This multiple scale feature information improves the accuracy of subsequent detection results.
  • the generated detection results include location information of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.
  • FIG. 5 shows a flow chart of a method for training an image processing model according to an embodiment of the present disclosure, which is applied to a cloud-side device and specifically includes the following steps:
  • Step 502 Obtain a sample image, as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image.
  • Step 504 Input the sample image and the sample guidance text into the image processing model to obtain the predicted labeling information, predicted category information and text loss value output by the image processing model.
  • Step 506 Calculate a model loss value based on the sample labeling information, the sample category information, the predicted labeling information, and the predicted category information.
  • Step 508 Adjust the model parameters of the image processing model according to the model loss value and the text loss value, and continue to train the image processing model until the model training stop condition is reached, thereby obtaining the model parameters of the image processing model.
  • Step 510 Send the model parameters of the image processing model to the terminal device.
  • steps 502 to 508 are implemented in the same manner as the training method of the above-mentioned image processing model, and will not be described in detail in this embodiment of the present disclosure.
  • model training requires large amounts of data and significant computing resources, which edge devices may not have the necessary processing capabilities. Therefore, model training can be performed on cloud-side devices.
  • the cloud-side device can also send the model parameters to the edge device.
  • the edge device can then locally construct an image processing model based on the model parameters and further use the image processing model to perform image processing.
  • the method provided by the embodiment of the present disclosure utilizes a text feature extraction module during the training of an image processing model, obtains corresponding sample text feature information from the sample guidance text through the text feature extraction module, and uses the sample text feature information to supervise the training of the image processing model, thereby improving the efficiency and accuracy of the image processing model training.
  • a semi-supervised image labeling model training method is adopted. An image labeling model is trained using some of the labeled sample images, and the image labeling model is used to label the unlabeled sample images, thereby enriching the number of sample images with labeled information.
  • model structure in the image annotation model is migrated to the image processing model, and the image processing model is further trained using the model structure in the trained image annotation model, which improves the training speed of the image processing model and avoids the waste of computing resources.
  • FIG. 6 shows a flow chart of an image processing method provided by an embodiment of the present disclosure, specifically comprising the following steps:
  • Step 602 Receive an image processing task sent by a user, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area.
  • Step 604 Input the multiple target images into the image processing model to obtain the detection results corresponding to the target detection area, wherein the image processing model generates the detection results corresponding to the target detection area based on the multi-scale feature information corresponding to the multiple target images, and the detection results include detection annotation information, detection category information and detection guidance text.
  • Step 606 Send the detection result corresponding to the target detection area to the user.
  • steps 602 to 604 is the same as the implementation of steps 202 to 204 described above, and will not be described in detail in the embodiment of the present disclosure.
  • an image processing request sent by a user is received, and the image processing request includes an image processing task.
  • the detection result needs to be returned to the user so that the user can perform corresponding subsequent processing based on the detection result.
  • the image processing method provided by one or more embodiments of the present disclosure uses an image processing model to obtain multiple scale feature information corresponding to a target image, thereby improving the accuracy of subsequent detection results.
  • the generated detection results include the location information of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results, providing users with multi-dimensional detection information, and enhancing the user experience.
  • One embodiment of the present disclosure also provides a computer-aided diagnosis method, comprising: receiving a CT image processing task sent by a user, wherein the CT image processing task carries multiple CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormal object within the target detection area; inputting the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, wherein the CT image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information, and detection guidance text; and sending the detection result corresponding to the target detection area to the user.
  • the specific implementation of this method is the same as the implementation of steps 202-204 above, and will not be repeated in this embodiment of the present disclosure.
  • FIG7 shows a flowchart of an image processing method provided by one embodiment of the present disclosure for the esophageal cancer detection scenario, specifically comprising the following steps:
  • Step 702 Receive a CT image processing task, wherein the CT image processing task carries multiple CT images corresponding to the esophagus, and the CT image processing task is used to detect whether there is a tumor in the esophagus.
  • Step 704 Input the multiple CT images into a CT image processing model to obtain a detection result corresponding to the esophagus, wherein the CT image processing model generates a detection result corresponding to the esophagus based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information and detection guidance text.
  • the collected dataset includes esophageal cancer screening data from 1,617 patients, including their corresponding CT images and test reports. Of these, 946 patients had esophageal cancer, while 671 did not. Experienced physicians were invited to annotate the tumors in the CT images of the 30% of patients diagnosed with esophageal cancer and make final decisions for each patient based on the test reports.
  • the image annotation model includes a multi-scale feature extraction module, a feature fusion module, and a first feature processing module, among which the first feature processing module includes an abnormal object segmentation unit and a guidance text generation unit.
  • the image annotation model After obtaining the image annotation model, the image annotation model is used to annotate the remaining CT images.
  • the annotation information corresponding to the CT images of users who do not suffer from esophageal cancer is "empty".
  • the multi-scale feature extraction module, feature fusion module, and first feature processing module in the image annotation model are extracted, and a new abnormal object classification unit is introduced.
  • the abnormal object classification unit, the abnormal object segmentation unit, and the guidance text generation unit are combined to form a second feature processing module, and the CT image processing model is constructed by the multi-scale feature extraction module, the feature fusion module, and the second feature extraction module.
  • the patient status corresponding to each patient i.e., whether the patient has esophageal cancer
  • the image annotation information is used as sample annotation information
  • the location guidance information and result guidance information are extracted from the test report as sample guidance text.
  • sample category information, sample annotation information, sample guidance text and CT images are used as training samples to train the CT image processing model until the model training stopping condition of the CT image processing model is reached.
  • the trained CT image processing model can be used to detect esophageal cancer.
  • the corresponding detection annotation information, detection category information and detection guidance text of the patient can be obtained.
  • FIG8 shows a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure. As shown in FIG8 , the device includes:
  • a receiving module 802 is configured to receive an image processing task, wherein the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area;
  • the detection module 804 is configured to input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection area, wherein the image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, and the detection result includes detection annotation information, detection category information and detection guidance text.
  • the image processing model includes a multi-scale feature extraction module, a feature fusion module, and a feature processing module;
  • the detection module 804 is further configured to:
  • the feature fusion information is input into the feature processing module to obtain the detection result corresponding to the target detection area.
  • the feature processing module includes an abnormal object segmentation unit, an abnormal object classification unit and a guidance text generation unit;
  • the detection module 804 is further configured to:
  • the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit;
  • the detection module 804 is further configured to:
  • the device further includes a segmentation module configured to:
  • the image segmentation task carries a plurality of initial images corresponding to a target detection area, and the image segmentation task is used to extract a target image corresponding to the target detection area;
  • Each initial image is input into a pre-trained image segmentation model to obtain a target image corresponding to each initial image output by the image segmentation model.
  • the device further includes a training module configured to:
  • sample image as well as sample annotation information, sample category information, and sample guidance text corresponding to the sample image;
  • the training module is further configured to:
  • a model loss value is calculated based on the first loss value and the second loss value.
  • the image processing model includes a multi-scale feature extraction module, a feature fusion module, a feature processing module, and a text feature extraction module;
  • the training module is further configured to:
  • the feature fusion information and the sample text feature information are input into the feature processing module to obtain the predicted labeling information, predicted category information and text loss value corresponding to the sample image.
  • the feature processing module includes an abnormal object segmentation unit, an abnormal object classification unit, and a guidance text generation unit;
  • the training module is further configured to:
  • a text loss value is obtained by calculation based on the predicted text feature information and the sample text feature information.
  • the guidance text generation unit includes an abnormal position text subunit and an abnormal result text subunit, and the sample text feature information includes sample position feature information and sample result feature information;
  • the training module is further configured to:
  • the text loss value is calculated based on the sample position feature information, the sample result feature information, the predicted position guidance information feature and the predicted result guidance information feature.
  • the sample image includes a first sample image and a second sample image, wherein the first sample image is marked with sample annotation information, and the second sample image is not marked with sample annotation information;
  • the training module is further configured to:
  • the second sample is input into the image annotation model to obtain predicted annotation information output by the image annotation model, and the predicted annotation information is used as sample annotation information of the second sample image.
  • the image annotation model includes an abnormal object segmentation unit and a guidance text generation unit;
  • the training module is further configured to:
  • the abnormal object segmentation unit and the guidance text generation unit in the image annotation model are used as the abnormal object segmentation unit and the guidance text generation unit of the image processing model.
  • the image processing device inputs multiple target images of the target detection area of the object to be detected into the image processing model.
  • multiple scale features are extracted from the 3D images corresponding to the multiple target images.
  • annotation information, category information, and guidance text are generated respectively, thereby finally generating a detection result.
  • the accuracy of the subsequent generation of detection results is improved by the multiple scale feature information.
  • the generated detection results include the location information of the abnormal object, the information of the object to be detected, and the guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.
  • the text feature extraction module is used to obtain the corresponding sample text feature information of the sample guidance text through the text feature extraction module, and the sample text feature information is used to supervise the training of the image processing model, thereby improving the efficiency and accuracy of the image processing model training.
  • a semi-supervised image labeling model training method is adopted. An image labeling model is trained using some of the labeled sample images, and the image labeling model is used to label the unlabeled sample images, thereby enriching the number of sample images with labeled information.
  • model structure in the image annotation model is migrated to the image processing model, and the image processing model is further trained using the model structure in the trained image annotation model, which improves the training speed of the image processing model and avoids the waste of computing resources.
  • the above is a schematic diagram of an image processing device according to this embodiment. It should be noted that the technical solution of the image processing device and the technical solution of the above-mentioned image processing method are based on the same concept. For details not described in detail in the technical solution of the image processing device, please refer to the description of the technical solution of the above-mentioned image processing method.
  • FIG9 shows a schematic structural diagram of a CT image processing device provided by an embodiment of the present disclosure. As shown in FIG9 , the device includes:
  • a receiving module 902 is configured to receive a CT image processing task, wherein the CT image processing task carries a plurality of CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormal object in the target detection area;
  • the detection module 904 is configured to input the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, wherein the CT image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple CT images, and the detection result includes detection annotation information, detection category information and detection guidance text.
  • the CT image processing device provided by one or more embodiments of the present disclosure extracts multiple scale features from 3D images corresponding to multiple CT images within an image processing model. After fusing these multiple scale features, the device generates annotation information, category information, and guidance text, ultimately generating a detection result.
  • the use of multiple scale feature information improves the accuracy of subsequent detection results.
  • the generated detection results include the location information of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.
  • the above is a schematic diagram of a CT image processing device according to this embodiment. It should be noted that the technical solution of the CT image processing device and the technical solution of the aforementioned CT image processing method are based on the same concept. For details not described in detail in the technical solution of the CT image processing device, please refer to the description of the technical solution of the aforementioned CT image processing method.
  • FIG10 shows a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure. As shown in FIG10 , the device includes:
  • the receiving module 1002 is configured to receive an image processing task sent by a user, wherein the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area.
  • the detection module 1004 is configured to input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection area, wherein the image processing model generates a detection result corresponding to the target detection area based on multi-scale feature information corresponding to the multiple target images, and the detection result includes detection annotation information, detection category information and detection guidance text.
  • the sending module 1006 is configured to send the detection result corresponding to the target detection area to the user.
  • the image processing device provided by one or more embodiments of the present disclosure extracts multiple scale features from 3D images corresponding to multiple images within an image processing model. After fusing these multiple scale features, the device then generates annotation information, category information, and guidance text, ultimately generating a detection result. This multiple scale feature information improves the accuracy of the subsequent detection results.
  • the generated detection results include the location information of abnormal objects, information about the object to be detected, and guidance text, enriching the detection results, providing users with multi-dimensional detection information, and improving the user experience.
  • the above is a schematic diagram of an image processing device according to this embodiment. It should be noted that the technical solution of the image processing device and the technical solution of the above-mentioned image processing method are based on the same concept. For details not described in detail in the technical solution of the image processing device, please refer to the description of the technical solution of the above-mentioned image processing method.
  • Figure 11 shows a block diagram of a computing device 1100 according to one embodiment of the present disclosure.
  • Components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120.
  • the processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.
  • the computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160.
  • networks 1160 include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet.
  • PSTN public switched telephone network
  • LAN local area network
  • WAN wide area network
  • PAN personal area network
  • Internet a combination of communication networks such as the Internet.
  • the access device 1140 may include one or more of any type of network interface, wired or wireless (e.g., a network interface card (NIC)), such as an IEEE802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, and a Near Field Communication (NFC).
  • NIC network interface card
  • the aforementioned components of the computing device 1100 and other components not shown in FIG11 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG11 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.
  • Computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC).
  • Computing device 1100 may also be a mobile or stationary server.
  • the processor 1120 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.
  • the above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method.
  • An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.
  • the above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method.
  • An embodiment of the present disclosure further provides a computer program product, comprising a computer program/instruction, which, when executed by a processor, implements the steps of the above-mentioned image processing method, CT image processing method or image processing model training method.
  • the above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product is based on the same concept as the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the aforementioned image processing method, CT image processing method, or image processing model training method.
  • the computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form.
  • the computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • General Physics & Mathematics (AREA)
  • Biomedical Technology (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Data Mining & Analysis (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Medical Informatics (AREA)
  • Molecular Biology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Biophysics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Public Health (AREA)
  • Databases & Information Systems (AREA)
  • Multimedia (AREA)
  • Pathology (AREA)
  • Epidemiology (AREA)
  • Primary Health Care (AREA)
  • Biodiversity & Conservation Biology (AREA)
  • Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
  • Radiology & Medical Imaging (AREA)
  • Quality & Reliability (AREA)
  • Apparatus For Radiation Diagnosis (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)

Abstract

本公开实施例提供图像处理方法、图像处理模型的训练方法,其中图像处理方法包括:接收图像处理任务,其中,图像处理任务携带目标检测区域对应的多个目标图像,图像处理任务用于检测目标检测区域内是否存在异常对象;将多个目标图像输入至图像处理模型,获得目标检测区域对应的检测结果,其中,基于多个目标图像对应的多尺度特征信息生成目标检测区域对应的检测结果,检测结果包括检测标注信息、检测类别信息和检测指导文本。通过获得目标图像对应的多个尺度特征信息提高了后续生成检测结果的准确率。检测结果包括了异常对象的位置信息、待检测对象的信息和指导文本,丰富了检测结果,为用户提供多维度的检测信息,提升了用户的使用体验。

Description

图像处理方法、计算机辅助诊断方法、图像处理模型的训练方法
本公开要求于2024年03月06日提交中国专利局、申请号为202410257868.3、申请名称为“图像处理方法、图像处理模型的训练方法”的中国专利申请的优先权,其全部内容通过引用结合在本公开中。
技术领域
本公开实施例涉及计算机技术领域,特别涉及一种图像处理方法。
背景技术
随着人民生活水平提高,越来越多的人重视自身的健康,肿瘤是影响人健康的重大因素之一,在医学图像中识别肿瘤是需要专业医生根据经验进行识别,受限于医生的经验,借助医学图像的进行图像识别分析成为一个重要课题。
目前人工智能系统已经展现出巨大的潜力,利用大模型对医学图像进行识别在医学影像计算机辅助诊断(computer aided diagnosis,CAD)任务中也取得了长足的进步,但是目前对医学图像进行识别分析的模型,识别准确度较低,而且模型在训练时,需要用到大量的标注数据,而标注数据又需要经验丰富的专业医生经验,因此,模型训练的效果也较差。因此,如何提升图像识别模型的图像识别准确度,就成为技术人员亟待解决的问题。
发明内容
有鉴于此,本公开实施例提供了一种图像处理方法。本公开一个或者多个实施例同时涉及提供了一种图像处理方法,本公开同时涉及癌症的计算机辅助诊断方法、图像处理模型的训练方法、计算机辅助诊断方法、图像处理装置,一种计算设备、一种计算机可读存储介质,以及一种计算机程序产品,以解决现有技术中存在的技术缺陷。
根据本公开实施例的第一方面,提供了一种图像处理方法,包括:
接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象;
将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于多个目标图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
根据本公开实施例的第二方面,提供了一种癌症的计算机辅助诊断方法,包括:
接收计算机断层扫描CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在肿瘤;
将所述多个CT图像输入至CT图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述CT图像处理模型基于多个CT图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
根据本公开实施例的第三方面,提供了一种图像处理模型的训练方法,应用于云侧设备,包括:
获取样本图像,以及所述样本图像对应的样本标注信息、样本类别信息和样本指导文本;
将所述样本图像和所述样本指导文本输入至图像处理模型,获得所述图像处理模型输出的预测标注信息、预测类别信息和文本损失值;
根据所述样本标注信息、所述样本类别信息和所述预测标注信息、所述预测类别信息计算模型损失值;
根据所述模型损失值和所述文本损失值调整所述图像处理模型的模型参数,并继续训练所述图像处理模型,直至达到模型训练停止条件,获得图像处理模型的模型参数;
向端侧设备发送所述图像处理模型的模型参数。
根据本公开实施例的第四方面,提供了一种计算机辅助诊断方法,包括:
接收用户发送的CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在异常对象;
将所述多个CT图像输入至CT图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述CT图像处理模型基于多个CT图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本;
向用户发送所述目标检测区域对应的检测结果。
根据本公开实施例的第五方面,提供了一种计算设备,包括:
存储器和处理器;
所述存储器用于存储计算机可执行指令,所述处理器用于执行所述计算机可执行指令,该计算机可执行指令被处理器执行时实现上述图像处理方法、CT图像处理方法或图像处理模型的训练方法的步骤。
根据本公开实施例的第六方面,提供了一种计算机可读存储介质,其存储有计算机可执行指令,该指令被处理器执行时实现上述图像处理方法、CT图像处理方法或图像处理模型的训练方法的步骤。
根据本公开实施例的第七方面,提供了一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法、CT图像处理方法或图像处理模型的训练方法的步骤。
本公开一个实施例提供的图像处理方法,通过图像处理模型获得目标图像对应的多个尺度特征信息提高了后续生成检测结果的准确率。在生成的检测结果中,包括了异常对象的位置信息、待检测对象的信息和指导文本,丰富了检测结果,为用户提供多维度的检测信息,提升了用户的使用体验。
附图说明
图1是本公开一个实施例提供的一种图像处理系统的架构图;
图2是本公开一个实施例提供的一种图像处理方法的流程图;
图3是本公开一个实施例提供的一种图像处理模型的结构示意图;
图4是本公开一个实施例提供的一种CT图像处理方法的流程图;
图5是本公开一个实施例提供的一种图像处理模型的训练方法的流程图;
图6是本公开一个实施例提供的另一种图像处理方法的流程图;
图7是本公开一个实施例提供的应用于食管癌检测场景的图像处理方法的流程图;
图8是本公开一个实施例提供的一种图像处理装置的结构示意图;
图9是本公开一个实施例提供的一种CT图像处理装置的结构示意图;
图10是本公开一个实施例提供的另一种图像处理装置的结构示意图;
图11是本公开一个实施例提供的一种计算设备的结构框图。
具体实施方式
在下面的描述中阐述了很多具体细节以便于充分理解本公开。但是本公开能够以很多不同于在此描述的其它方式来实施,本领域技术人员可以在不违背本公开内涵的情况下做类似推广,因此本公开不受下面公开的具体实施的限制。
在本公开一个或多个实施例中使用的术语是仅仅出于描述特定实施例的目的,而非旨在限制本公开一个或多个实施例。在本公开一个或多个实施例和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。还应当理解,本公开一个或多个实施例中使用的术语“和/或”是指并包含一个或多个相关联的列出项目的任何或所有可能组合。
应当理解,尽管在本公开一个或多个实施例中可能采用术语第一、第二等来描述各种信息,但这些信息不应限于这些术语。这些术语仅用来将同一类型的信息彼此区分开。例如,在不脱离本公开一个或多个实施例范围的情况下,第一也可以被称为第二,类似地,第二也可以被称为第一。取决于语境,如在此所使用的词语“如果”可以被解释成为“在……时”或“当……时”或“响应于确定”。
此外,需要说明的是,本公开一个或多个实施例所涉及的用户信息(包括但不限于用户设备信息、用户个人信息等)和数据(包括但不限于用于分析的数据、存储的数据、展示的数据等),均为经用户授权或者经过各方充分授权的信息和数据,并且相关数据的收集、使用和处理需要遵守相关地区的相关法律法规和标准,并提供有相应的操作入口,供用户选择授权或者拒绝。
首先,对本公开一个或多个实施例涉及的名词术语进行解释。
CT(Computed Tomography):电子计算机断层扫描,它是利用精确准直的X线束、γ射线、超声波等,与灵敏度极高的探测器一同围绕人体的某一部位作一个接一个的断面扫描,具有扫描时间快,图像清晰等特点,可用于多种疾病的检查。
CAD(computer aided diagnosis):计算机辅助诊断,是指通过影像学、医学图像处理技术以及其他可能的生理、生化手段,结合计算机的分析计算,辅助发现病灶,提高诊断的准确率。
EC(esophageal cancer):食管癌,是致命率极高的癌症,具不完全统计5年生存率较低,然而若早期发现可切除/可治愈的食管癌会很大程度上降低死亡率,其中,淋巴结转移是较为常见,具有典型性的一类病症。
随着人民生活水平提高,越来越多的人重视自身的健康,肿瘤是影响人健康的重大因素之一,在医学图像中识别肿瘤是需要专业医生根据经验进行识别,受限于医生的经验,借助医学图像的进行图像识别分析成为一个重要课题。
目前人工智能系统已经展现出巨大的潜力,利用大模型对医学图像进行识别在医学影像计算机辅助诊断(computer aided diagnosis,CAD)任务中也取得了长足的进步,但是目前对医学图像进行识别分析的模型,识别准确度较低。大多数人工智能系统严重依赖肿瘤级别的标注信息,这需要经验丰富的放射科医生的标注。另一方面,临床报告中同时会包含有丰富的描述性信息,而目前的计算机辅助诊断系统还无法有效的利用临床报告中的信息。
基于此,在本公开中,提供了一种图像处理方法,本公开同时涉及CT图像处理方法、图像处理模型的训练方法、图像处理装置,一种计算设备、一种计算机可读存储介质,以及一种计算机程序产品,在下面的实施例中逐一进行详细说明。
参见图1,图1示出了本公开一个实施例提供的一种图像处理系统的架构图,图像处理系统可以包括客户端100和服务端200;
客户端100,用于向服务端200发送图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象;
服务端200,用于将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于多个目标图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本;向客户端100发送检测结果;
客户端100,还用于接收服务端200发送的检测结果。
应用本公开实施例的方案,接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象;将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于多个目标图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
如此,在利用图像处理模型在对多个目标图像进行处理的过程中,提取了多个目标图像的多尺度特征信息,并基于多尺度特征信息生成了检测标注信息、检测类别信息和检测指导文本,
图像处理系统可以包括多个客户端100以及服务端200,其中,客户端100可以称为端侧设备,服务端200可以称为云侧设备。多个客户端100之间通过服务端200可以建立通信连接,在图像处理场景中,服务端200即用来在多个客户端100之间提供图像处理服务,多个客户端100可以分别作为发送端或接收端,通过服务端200实现通信。
用户通过客户端100可与服务端200进行交互以接收其它客户端100发送的数据,或将数据发送至其它客户端100等。在图像处理场景中,可以是用户通过客户端100向服务端200发布数据流,服务端200根据该数据流生成检测结果,并将检测结果推送至其他建立通信的客户端中。
其中,客户端100与服务端200之间通过网络建立连接。网络为客户端100与服务端200之间提供了通信链路的介质。网络可以包括各种连接类型,例如有线、无线通信链路或者光纤电缆等等。客户端100所传输的数据可能需要经过编码、转码、压缩等处理之后才发布至服务端200。
客户端100可以为浏览器、APP(Application,应用程序)、或网页应用如H5(HyperText Markup Language5,超文本标记语言第5版)应用、或轻应用(也被称为小程序,一种轻量级应用程序)或云应用等,客户端100可以基于服务端200提供的相应服务的软件开发工具包(SDK,Software Development Kit),如基于实时通信(RTC,Real Time Communication)SDK开发获得等。客户端100可以部署在电子设备中,需要依赖设备运行或者设备中的某些APP而运行等。电子设备例如可以具有显示屏并支持信息浏览等,如可以是个人移动终端如手机、平板电脑、个人计算机等。在电子设备中通常还可以配置各种其它类应用,例如人机对话类应用、模型训练类应用、文本处理类应用、网页浏览器应用、购物类应用、搜索类应用、即时通信工具、邮箱客户端、社交平台软件等。
服务端200可以包括提供各种服务的服务器,例如为多个客户端提供通信服务的服务器,又如为客户端上使用的模型提供支持的用于后台训练的服务器,又如对客户端发送的数据进行处理的服务器等。需要说明的是,服务端200可以实现成多个服务器组成的分布式服务器集群,也可以实现成单个服务器。服务器也可以为分布式系统的服务器,或者是结合了区块链的服务器。服务器也可以是云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络(CDN,Content Delivery Network)以及大数据和人工智能平台等基础云计算服务的云服务器,或者是带人工智能技术的智能云计算服务器或智能云主机。
值得说明的是,本公开实施例中提供的图像处理方法一般由服务端执行,但是,在本公开的其它实施例中,客户端也可以与服务端具有相似的功能,从而执行本公开实施例所提供的图像处理方法。在其它实施例中,本公开实施例所提供的图像处理方法还可以是由客户端与服务端共同执行。
参见图2,图2示出了本公开一个实施例提供的一种图像处理方法的流程图,具体包括以下步骤:
步骤202:接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象。
在实际应用中,可以通过服务端,也可以通过客户端接收用户发送的图像处理任务。
具体的,图像处理任务具体是指用于检测目标检测区域内是否存在异常对象的任务。图像处理任务中携带目标检测区域对应的多个目标图像。进一步的,目标检测区域可以理解为用于检测是否存在异常对象的分区,异常对象具体是指目标检测区域内的异物。例如,异常对象可以是人体内的肿瘤。
目标检测区域可以是人体内的任一器官,如肝脏、肺、胃、食管等。通过对目标检测区域是否存在异常对象进行预测,可以根据预测结果进一步判断待检测对象的状态,从而对异常对象的定位和精准治疗提供帮助。
待检测对象可以理解为目标检测区域所述的对象,例如,目标图像是张三胃部的CT图像,则目标检测区域为胃部,CT图像即为目标图像,待检测对象即为张三,该图像处理任务用于检测胃部区域是否存在有肿瘤。在实际应用中,待检测对象可以是人,也可以是其他生命体,在本公开提供的一个或多个具体实施方式中,对此不做限定。
需要说明的是,本公开一个或多个实施例中,图像处理任务可以应用于对各类医学图像的识别,并根据图像特征判断医学图像中的目标检测区域中是否存在异常对象。示例性的,在胃癌检测的应用场景中,可以根据胃部区域的医学图像,预测在胃部区域中是否存在肿瘤,从而帮助医生对出现异常的部位进行精准定位;在食道癌检测的应用场景中,可以根据食道区域的医学图像,预测在食道中是否存在肿瘤,从而帮助医生对出现异常的部位进行精准定位,便于后续的治疗。
示例性的,在食道癌检测的场景中,获取的目标图像即为食道的图像,具体的,获取的多个目标图像即为食道的CT图像,多个目标图像可以组成食道区域的3D图像。异常对象可以理解为食道中的肿瘤。获取食道区域对应的多个CT图像,通过对多个CT图像进行图像检测处理,检测在食道中是否存在恶性肿瘤。
在实际应用中,异常对象可以为某一种细胞,某一种组织结构等等,例如可以为恶性肿瘤、良性肿瘤、增生组织等等。在本公开提供的一个或多个实施例中,对此不做限定。
通过接收图像处理任务,可以将图像处理任务中携带的目标检测区域对应的多个目标图像作为输入,用于检测目标检测区域内是否存在异常对象。
在本公开提供的一具体实施方式中,在接收图像处理任务之前,还包括:
接收图像分割任务,其中,所述图像分割任务携带目标检测区域对应的多个初始图像,所述图像分割任务用于提取所述目标检测区域对应的目标图像;
将各初始图像输入至预先训练的图像分割模型,获得所述图像分割模型输出的各初始图像对应的目标图像。
在本公开实施例提供的实施例中,目标图像可以理解为目标检测区域对应的特写图像。而在实际应用中,接收到的多个图像可能会存在除了目标检测区域外,还包括了其他的区域,而其他的区域会对目标检测区域造成影响。因此,在本公开实施例提供的方法中,还对图像做了处理。
具体的,首先获取针对多个初始图像的图像分割任务,初始图像中即包括有目标检测区域,又包括对目标检测区域带来影响的区域。该图像分割任务用于从各初始图像中提取出目标检测区域对应的目标图像。
将多个初始图像输入至预先训练的图像分割模型进行处理。图像分割模型被训练于识别到初始图像中的目标检测区域,并将目标检测区域从初始图像中截取出来,生成目标检测区域对应的目标图像。
以CT图像为例,在本公开提供的方法中,首先获取初始CT图像,初始CT图像为图像质量要求的平扫CT图像,初始CT图像的来源可以是多个CT扫描机,也可以是同一个CT扫描机,在本公开中对此不做限定。在获得初始CT图像之后,要统一各初始CT图像的格式。图像分割模型可以是3DUNet。
3DUNet是一种在三维空间中进行图像分割的深度学习架构,是U-Net模型的扩展版本,U-Net最初被设计用于二维生物医学图像的语义分割,并因其优异的表现和对小目标分割的高精度而受到广泛认可。3DUNet将这一思想应用于三维数据集,例如医学影像(如CT、MRI(Magnetic Resonance Imaging,磁共振成像)扫描等),这在许多医疗领域非常有用,因为这些领域的数据通常具有丰富的三维结构信息。
在3DUNet中包括有编码器-解码器结构,用于捕获全局上下文信息,以及恢复丢失的空间细节并生成精确的像素级分割标签。3DUNet结构中的3D卷积核在网络中替代了2D卷积核,可以同时处理输入数据在三个维度(长、宽、高)上的特征。同时,3DUNet还能有效地融合不同层次的三维特征,有助于提取复杂的形状和结构信息。该框架在医学图像分割方法有较好的效果。在本公开提供的方法中,对初始CT图像使用3DUNet结构的预处理策略进行切割处理,获得目标检测区域对应的目标图像,用于后续的处理。
步骤204:将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于多个目标图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
实际应用中,在接收到图像处理任务后,可以从图像处理任务中获取其携带的多个目标图像,将多个目标图像输入至图像处理模型,可以获得图像处理模型输出的目标检测区域对应的检测结果,检测结果具体包括检测标注信息、检测类别信息和检测指导文本。
其中,检测标注信息具体是指在目标图像中标记出异常对象的区域,若在目标图像中没有异常对象,则检测标注信息为空。
检测类别信息具体是指针对目标图像的检测,目标检测区域对应的待检测对象所述类别信息。若目标检测区域中包括异常对象,则检测类别信息为异常;若目标检测区域中未包括异常对象,则检测类别信息为正常。
检测指导文本具体是指根据目标图像的检测标注信息、检测类别信息等,输出的针对待检测对象关于异常对象位置、状态等信息的指导性文本。需要注意的是,本公开实施例提供的检测指导文本是将检测到的异常对象位置、状态等信息,添加到指导文本模版中生成的指导文本。如果在检测标注信息确定有异常对象的情况下,在检测标注信息中会给出异常对象的具体位置、以及待检测对象的类别信息。例如检测指导文本可以是“患者在食管的上部有一个肿瘤,患者患有食管癌”,又例如检测指导文本可以是“患者胃部未检测到肿瘤,患者未检测到胃癌”等等。
在图像处理模型内部会提取各目标图像的多尺度特征信息,再根据多尺度特征信息提取出目标检测区域对应的检测结果。
具体的,所述图像处理模型包括多尺度特征提取模块、特征融合模块、特征处理模块;
将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,包括S2042-S2046:
S2042、将所述多个目标图像输入至所述多尺度特征提取模块,获得至少一个尺度特征信息。
其中,多尺度特征提取模块具体是指用不同的尺度信息,提取多个目标图像组成的3D模型的目标特征信息。例如,对于给尺寸为W*H*D的3D模型,将其输入到多尺度特征提取模块,该多尺度特征提取模块可以理解为3DUNet的主干特征提取网络,得到多尺度特征图F={F0,F1,F2,……FS},其中,S为尺度层数,在本公开提供的一个实施方式中,S=5,即多尺度特征图F={F0,F1,F2,F3,F4,F5}。对于第i个尺度层数对应的特征图
参见图3,图3示出了本公开一个实施例提供的图像处理模型的结构示意图,如图3所示,将多个目标图像输入到图像处理模型的多尺度特征提取模块中,经过5个stage的降采样处理,获得多个不同尺度特初始特征信息,再对多个不同尺度的初始特征信息进行卷积处理或反卷积-卷积处理,获得各尺度对应的尺度特征图,从而获得6个多尺度特征图。
S2044、将各尺度特征信息输入至所述特征融合模块,获得特征融合信息。
在获得了多个尺度特征信息之后,将多个尺度特征信息输入到特征融合层中进行特征融合,将多个尺度特征信息统一到同一个尺度内,进行融合,获得了特征融合信息。
参见图3,将多个尺度特征图{F0,F1,F2,F3,F4,F5}输入到特征融合层,在特征融合层中,将多个尺度特征图统一到同一个尺度内进行特征融合,获得融合特征信息Fa。
S2046、将所述特征融合信息输入至所述特征处理模块,获得所述目标检测区域对应的检测结果。
在获得了特征融合信息之后,即可将特征融合信息输入到特征处理模块,在特征处理模块中对特征融合信息进行处理,从而生成针对目标检测区域对应的检测结果。在本公开提供的实施方式中,检测结果具体包括检测标注信息、检测类别信息和检测指导文本。
在本公开提供的一具体实施方式中,所述特征处理模块包括异常对象分割单元、异常对象分类单元和指导文本生成单元;
将所述特征融合信息输入至所述特征处理模块,获得所述目标检测区域对应的检测结果,包括:
将所述特征融合信息输入至所述异常对象分割单元,获得检测标注信息;
将所述特征融合信息输入至所述异常对象分类单元,获得检测类别信息;
将所述特征融合信息输入至所述指导文本生成单元,获得检测指导文本;
根据所述检测标注信息、检测类别信息、检测指导文本生成所述目标检测区域对应的检测结果。
在实际应用中,特征处理模块用于对特征融合信息进行处理,提取特征融合信息中的特征,并对提取到的特征进行解码,从而生成对应的检测结果,在特征处理模块中具体包括有异常对象分割单元、异常对象分类单元和指导文本生成单元。
其中,异常对象分割单元用于根据特征融合信息在目标图像中标记出异常对象的区域,在目标图像中分割出目标检测区域对应的异常对象;异常对象分类单元用于根据特征融合信息确定目标检测区域中是否有异常对象,并确定针对待检测对象的分类信息;指导文本生成单元用于根据特征融合信息生成用于指示待检测对象关于异常对象位置、状态等信息的指导性文本。
根据异常对象分割单元、异常对象分类单元和指导文本生成单元分别输出的检测标注信息、检测类别信息和检测指导文本,生成最终的检测结果。
更进一步的,所述指导文本生成单元包括异常位置文本子单元和异常结果文本子单元;
将所述特征融合信息输入至所述指导文本生成单元,获得检测指导文本,包括:
将所述特征融合信息输入至所述异常位置文本子单元,获得位置指导信息;
将所述特征融合信息输入至所述异常结果文本子单元,获得结果指导信息;
根据所述位置指导信息和所述结果指导信息,生成检测指导文本。
在实际应用中,检测指导文本中具体包括有异常位置指导信息和待检测对象的状态指导信息。因此,在指导文本生成单元中包括有异常位置文本子单元和异常结果文本子单元。其中,异常位置文本子单元用于根据特征融合信息生成异常对象在目标检测区域中的位置信息,异常结果文本子单元,用于生成待检测对象的状态信息。
例如,在对某个用户的胃部CT图像进行检测的过程,用于检测胃部的异常对象,生成检测指导文本,异常位置文本子单元对特征融合信息进行处理后,确定其位置指导信息为“胃的中部”,异常结果文本子单元对特征融合信息进行处理后,确定其结果指导信息为“该用户患有胃癌”。根据位置指导信息和结果指导信息生成最终的检测指导文本“该用户在胃的中部有肿瘤,患有胃癌”。
通过本公开实施例提供的图像处理方法,将待检测对象针对目标检测区域的多个目标图像输入至图像处理模型,在图像处理模型中对多个目标图像对应的3D图像进行多个尺度特征提取,并对多个尺度特征进行融合后,分别进行标注信息、类别信息和指导文本的生成,从而最终生成检测结果。通过多个尺度特征信息提高了后续生成检测结果的准确率。在生成的检测结果中,包括了异常对象的位置信息、待检测对象的信息和指导文本,丰富了检测结果,为用户提供多维度的检测信息,提升了用户的使用体验。
随着计算机技术的不断发展,深度学习逐渐能应用于各种医学影响计算机辅助诊断任务中,深度学习模型依赖大规模精准标注的训练样本和样本标签作为训练数据进行模型训练,目前对图像处理模型在模型训练过程中,其训练数据为目标检测区域内包括异常对象的图像,其对于目标检测区域内没有包括异常对象的处理效果较差。基于此,在本公开提供的一具体实施方式中,所述图像处理模型通过下述步骤训练获得:
获取样本图像,以及所述样本图像对应的样本标注信息、样本类别信息和样本指导文本;
将所述样本图像和所述样本指导文本输入至图像处理模型,获得所述图像处理模型输出的预测标注信息、预测类别信息和文本损失值;
根据所述样本标注信息、所述样本类别信息和所述预测标注信息、所述预测类别信息计算模型损失值;
根据所述模型损失值和所述文本损失值调整所述图像处理模型的模型参数,并继续训练所述图像处理模型,直至达到模型训练停止条件。
需要说明的是,在本公开提供的图像处理模型的训练方法使用了有监督训练和迁移训练的思想,先使用有标注的训练数据训练一个初始模型,再将初始模型中的骨干网络迁移到图像处理模型中,用另一批训练数据进一步进行训练,获得最终的图像处理模型。
样本图像具体是指用于进行模型训练的图像,在实际应用中,样本图像中不仅包含有异常对象的图像,还包含有不存在异常对象的图像。样本图像对应的样本标注信息具体是指在样本图像中对异常对象的标注信息;样本类别信息具体是指样本图像对应的待检测对象的状态信息;样本指导文本具体是指样针对样本图像中异常对象的位置、待检测对象的状态的文本。样本类别信息可以从样本指导文本中获取,也可以是单独的样本类别信息。
以训练计算机辅助诊断领域的辅助诊断图像处理模型为例,在训练过程中,引入了健康者的图像,而不是仅使用肿瘤患者的图像,从而避免了在图像处理模型在真实应用时,面对多样化的无肿瘤图像时,预测出错误的结果。
在获得了样本图像之后,将样本图像输入至图像处理模型中,此时的图像处理模型还是未训练好的图像处理模型。在图像处理模型中,根据各样本图像生成预测检测结果,预测检测结果包括预测标注信息、预测类别信息、预测指导文本。
在本公开提供的方法中,图像处理模型如上述步骤中的图像处理模型的模型结构,也包括有多尺度特征提取模块、特征融合模块、特征处理模块,样本图像在未训练好的图像处理模型中的数据处理过程,与上述实施例中图像处理模型的数据处理过程相同,关于样本图像在未训练好的图像处理模型中的数据处理过程,参见上述目标图像在图像处理模型中的数据处理过程,在此不在赘述。
在获得了样本图像的预测检测结果之后,可根据预测检测结果和样本检测结果计算模型损失值,在本公开提供的方法中,计算模型损失值的方法有很多,例如交叉熵损失函数、最大损失函数、平均值损失函数等等,在本公开中,对损失函数的具体方式不做限定,以实际应用为准。
在本公开一个或多个具体实施方式中,样本检测结果包括样本标注信息、样本类别信息,预测检测结果包括预测标注信息、预测类别信息,根据所述样本标注信息、所述样本类别信息和所述预测标注信息、所述预测类别信息计算模型损失值,包括:
根据所述样本标注信息和所述预测标注信息计算第一损失值;
根据所述样本类别信息和所述预测类别信息计算第二损失值;
根据所述第一损失值和所述第二损失值计算模型损失值。
具体的,在本公开一个或多个具体实施方式中,样本检测结果包括样本标注信息、样本类别信息,预测检测结果包括预测标注信息、预测类别信息。技术人员希望预测检测结果和样本检测结果一致,从而提升图像处理模型预测的准确性。
用样本标注信息和预测标注信息计算第一损失值,用样本类别信息和预测类别信息计算第二损失值,同时预测指导文本、预测标注信息之间还可以相互验证,进一步提升模型预测的准确性。
再将第一损失值和第二损失值融合,获得模型损失值,具体的,是将第一损失值和第二损失值相加,获得模型损失值。
除此之外,需要注意的是,在本公开提供的图像处理模型的训练方法中,样本指导文本是在图像处理模型的内部进行处理,获得文本损失值。即需要将样本指导文本也输入到图像处理模型中。图像处理模型包括多尺度特征提取模块、特征融合模块、特征处理模块、文本特征提取模块;
将所述样本图像和所述样本指导文本输入至图像处理模型,获得所述图像处理模型输出的预测标注信息、预测类别信息和文本损失值,包括:
将所述样本图像输入至所述多尺度特征提取模块,获得至少一个尺度特征信息;
将各尺度特征信息输入至所述特征融合模块,获得特征融合信息;
将所述样本指导文本输入至所述文本特征提取模块,获得样本文本特征信息;
将所述特征融合信息和所述样本文本特征信息输入至所述特征处理模块,获得所述样本图像对应的预测标注信息、预测类别信息和文本损失值。
图像处理模型中多尺度特征提取模块、特征融合模块的处理方式参见上述步骤中的内容,在此不再赘述。
在图像处理模型的模型训练过程中,图像处理模型中还包括有文本特征提取模块,在训练阶段,将所述样本指导文本输入至所述文本特征提取模块,获得样本文本特征信息。同时,在训练阶段,将特征融合信息和样本文本特征信息输入到特征处理模块,在特征处理模块中进行处理后,获得预测标注信息、预测类别信息和文本损失值。
在获得文本损失值和模型损失值后,根据文本损失值和模型损失值一起对图像处理模型的模型参数进行调整,具体的,是将文本损失值和模型损失值相加,获得新的损失值,并根据新的损失值反向传播,调整图像处理模型的模型参数。
在本公开提供的一具体实施方式中,所述特征处理模块包括异常对象分割单元和异常对象分类单元、指导文本生成单元;
将所述特征融合信息和所述样本文本特征信息输入至所述特征处理模块,获得所述样本图像对应的预测标注信息、预测类别信息和文本损失值,包括:
将所述特征融合信息输入至所述异常对象分割单元,获得检测标注信息;
将所述特征融合信息输入至所述异常对象分类单元,获得检测类别信息;
将所述特征融合信息输入至所述指导文本生成单元,获得预测文本特征信息;
根据所述预测文本特征信息和所述样本文本特征信息计算获得文本损失值。
具体的,在训练阶段的特征处理模块中包括有异常对象分割单元、异常对象分类单元和指导文本生成单元。其中,异常对象分割单元用于根据特征融合信息生成检测标注信息;异常对象分类单元用于根据特征融合信息生成检测类别信息;指导文本生成单元用于根据特征融合信息生成预测文本特征信息,具体是指生成预测文本特征向量,再通过预测文本特征信息和所述样本文本特征信息计算获得文本损失值。
更进一步的,所述指导文本生成单元包括异常位置文本子单元和异常结果文本子单元,所述样本文本特征信息包括样本位置特征信息和样本结果特征信息;
将所述特征融合信息输入至所述指导文本生成单元,获得预测文本特征信息,包括:
将所述特征融合信息输入至所述异常位置文本子单元,获得预测位置指导信息特征;
将所述特征融合信息输入至所述异常结果文本子单元,获得预测结果指导信息特征;
相应的,根据所述预测文本特征信息和所述样本文本特征信息计算获得文本损失值,包括:
根据所述样本位置特征信息、样本结果特征信息、预测位置指导信息特征和预测结果指导信息特征计算获得文本损失值。
在实际应用中,指导文本生成单元包括异常位置文本子单元和异常结果文本子单元,指导文本生成单元在根据特征融合信息生成预测指导文本的过程中,先生成预测文本特征信息,更进一步的,预测文本特征信息由预测位置指导信息特征和预测结果指导信息特征组成,样本文本特征信息包括样本位置特征信息和样本结果特征信息。根据所述样本位置特征信息、样本结果特征信息、预测位置指导信息特征和预测结果指导信息特征计算获得文本损失值。
在实际应用中,在样本指导文本中具体包括有两部分内容,即样本位置指导信息和样本结果指导信息。其中,样本位置指导信息具体是指目标检测区域内异常对象的位置信息,样本结果指导信息具体是指待检测对象的状态信息。例如样本位置指导信息包括“食管的中部有肿瘤”,样本结果指导信息包括“患者患有食管癌”。样本位置特征信息具体是指样本位置指导信息对应的特征向量,样本结果特征信息具体是指样本结果指导信息对应的特征向量。
将样本指导文本输入到文本特征提取模块中,该文本特征提取模块具体是指可以将样本指导文本转换为对应的文本特征信息的模块,例如Transformer模型的编码器、CLIP模型的文本编码器等等。更进一步的,是将样本指导文本中的样本位置指导信息和样本结果指导信息输入到特征提取模块中,获得样本位置指导信息对应的样本位置指导信息特征、样本结果指导信息对应的样本结果指导信息特征。
在图像处理模型中,包括有指导文本生成单元,在指导文本生成单元中包括异常位置文本子单元和异常结果文本子单元,图像处理模型在处理过程中,会获得异常位置文本子单元输出的预测位置指导信息特征、异常结果文本子单元输出的预测结果指导信息特征。
根据所述样本位置指导信息特征和所述预测位置指导信息特征计算位置损失值,根据所述样本结果指导信息特征和所述预测结果指导信息特征计算结果损失值,再根据位置损失值和所述结果损失值确定文本损失值。
通过文本特征提取模块对样本指导文本进行处理,该文本特征提取模块可以利用样本指导文本中待检测对象的样本位置指导信息和样本结果指导信息对图像处理模型进行监督,充分利用了目前已有的样本指导文本,提升了图像处理模型的训练速度和图像处理模型的预测准确率。
在实际应用中,还存在有样本图像标记信息不全的问题。在实际应用中,只有部分图像中设置有标注信息和样本指导文本,大部分图像只有样本指导文本而没有标注信息。例如对于CT图像,只有部分CT图像由有经验的医生对肿瘤区域进行了标注,而大部分的CT图像没有对肿瘤区域进行标注,在实际应用中,有肿瘤区域标注的CT图像占比可能只有20%-30%,而没有肿瘤区域标注的CT图像占比有70%-80%。样本指导文本可以理解为CT图像对应的临床报告。
为了充分利用样本文本,在本公开提供的另一具体实施方式中,所述样本图像包括第一样本图像和第二样本图像,其中,所述第一样本图像中标记有样本标注信息,第二样本图像中未标记样本标注信息;
获取样本图像,以及所述样本图像对应的样本标注信息、样本类别信息和样本指导文本,包括:
获取样本图像和样本图像对应的样本指导文本,并从所述样本指导文本中提取样本类别信息;
根据第一样本图像,以及所述第一样本图像对应的样本标注信息和样本指导文本训练图像标注模型,获得用于生成标注信息的图像标注模型;
将所述第二样本输入至所述图像标注模型,获得所述图像标注模型输出的预测标注信息,将所述预测标注信息作为所述第二样本图像的样本标注信息。
其中,第一样本图像即为有样本标注信息的样本图像,第二样本图像即为没有样本标注信息的样本图像。第一样本图像和第二样本图像均有对应的样本指导文本。
在各样本指导文本中可以识别并提取出待检测对象对应的样本类别信息,在获得各样本图像对应的样本类别信息之后,即可先利用有样本标注信息的第一样本图像训练一个图像标注模型,用于为第二样本图像生成样本标注信息。
具体的,由于各样本图像均包括有对应的样本指导文本,第一样本图像有样本标注信息,第二样本图像没有样本标注信息。则先利用第一样本图像和第一样本图像对应的样本标注信息和样本指导文本,预先训练一个图像标注模型,该图像标注模型中包括有异常对象分割单元和指导文本生成单元。
将第一样本图像输入到图像标注模型中,获得异常对象分割单元输出的预测标注信息,获得指导文本生成单元输出的预测指导文本。用预测标注信息、样本标注信息、预测指导文本和样本指导文本计算图像标注模型的模型损失值,基于图像标注模型的模型损失值调整图像标注模型的模型参数,直至达到模型训练停止条件,获得训练好的图像标注模型,该图像标注模型可以对未标记有图像标注信息的图像进行图像标注。
将第二样本图像输入至训练好的图像标注模型,图像标注模型可以为第二样本图像生成预测标注信息,并将该预测标注信息作为第二样本图像对应的样本标注信息。
在训练图像标注模型的过程中,图像标注模型中包括有异常对象分割单元和指导文本生成单元。而在上述步骤中的图像处理模型也包括有异常对象分割单元和指导文本生成单元,为了进一步节省计算资源,所述方法还包括:
将所述图像标注模型中的异常对象分割单元和指导文本生成单元,作为所述图像处理模型的异常对象分割单元和指导文本生成单元。
即将图像标注模型作为教师模型,在对图像标注模型的训练过程中,利用了第一样本图像、第一样本图像对应的样本标注信息和样本指导文本,获得的图像标注模型中的异常对象分割单元和指导文本生成单元已经经过训练调参,具备相应的数据处理能力。可以将异常对象分割单元和指导文本生成单元直接迁移至图像处理模型中,将图像标注模型中的异常对象分割单元和指导文本生成单元,作为图像处理模型的异常对象分割单元和指导文本生成单元。
更进一步的,在训练图像标注模型的过程中,可以设置图像标注模型同样包括多尺度特征提取模块、特征融合模块,在图像标注模型的训练过程中继续进行训练。再将其迁移到图像处理模型中,使得图像处理模型中的多尺度特征提取模块、特征融合模块、特征处理模块中的异常对象分割单元和指导文本生成单元,在图像标注模型的训练过程中进行初步训练。在图像处理模型训练时,再次进行训练,或不再训练,从而提升图像处理模型的模型训练效率。在实际应用中,图像处理模型在模型训练阶段包括文本特征提取模块,在图像处理模型的模型应用阶段,可以不包括文本特征提取模块。
本公开实施例提供的图像处理模型的训练方法,利用文本特征提取模块,将样本指导文本通过文本特征提取模块获得对应的样本文本特征信息,将样本文本特征信息对图像处理模型的训练进行监督,提升了图像处理模型训练的效率和准确度。
另外,针对样本图像的标注信息不全的情况,采用半监督的图像标注模型的训练方法,利用部分由标注的样本图像,训练一个图像标注模型,用该图像标注模型对未标注的样本图像进行标注,从而丰富了有标注信息的样本图像的数量。
最后,将图像标注模型中的模型结构迁移到图像处理模型中,利用已经训练的图像标注模型中的模型结构继续训练图像处理模型,提升了图像处理模型的训练速度,也避免了计算资源的浪费。
参见图4,图4示出了本公开一个实施例提供的一种CT图像处理方法的流程图,具体包括以下步骤:
步骤402:接收CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在异常对象。
其中,异常对象可以是肿瘤。本实施例提供的CT图像处理方法属于一种癌症的计算机辅助诊断方法,用于检测目标检测区域对应的CT图像内是否存在肿瘤。
步骤404:将所述多个CT图像输入至CT图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述CT图像处理模型基于多个CT图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
需要说明的是,步骤402-步骤404的实现方式,与上述步骤202-步骤204的实现方式相同,本公开实施例便不再进行赘述。
示例性的,以目标检测区域为胃部为例,接收CT图像处理任务,在CT图像处理任务中包括目标用户的胃部对应的多个CT图像,多个CT图像可以组成目标用户的胃部3D图,该CT图像处理任务用于检测目标用户的胃部中是否存在恶性肿瘤。
应用本公开实施例的方法,CT图像处理模型为上述实施例中的图像处理模型,CT图像处理模型的模型结构与上述实施例中图像处理模型的结构相同,在此不再赘述。通过将胃部对应的多个CT图像输入至CT图像处理模型,能获得图像处理模型输出的胃部对应的检测结果,从而实现对胃部中是否存在恶性肿瘤的自动检测,具体的,CT图像处理模型根据多个CT图像生成多个尺度特征信息,再根据多个尺度特征信息生成胃部对应的肿瘤标注信息、用户患病状态(患有胃癌或未患病)、用户检测文本(用户胃部的肿瘤部位、用户是否患病等)等信息。
本公开一个或多个实施例提供的CT图像处理方法,在图像处理模型中对多个CT图像对应的3D图像进行多个尺度特征提取,并对多个尺度特征进行融合后,分别进行标注信息、类别信息和指导文本的生成,从而最终生成检测结果。通过多个尺度特征信息提高了后续生成检测结果的准确率。在生成的检测结果中,包括了异常对象的位置信息、待检测对象的信息和指导文本,丰富了检测结果,为用户提供多维度的检测信息,提升了用户的使用体验。
参见图5,图5示出了本公开一个实施例提供的一种图像处理模型的训练方法的流程图,应用于云侧设备,具体包括以下步骤:
步骤502:获取样本图像,以及所述样本图像对应的样本标注信息、样本类别信息和样本指导文本。
步骤504:将所述样本图像和所述样本指导文本输入至图像处理模型,获得所述图像处理模型输出的预测标注信息、预测类别信息和文本损失值。
步骤506:根据所述样本标注信息、所述样本类别信息和所述预测标注信息、所述预测类别信息计算模型损失值。
步骤508:根据所述模型损失值和所述文本损失值调整所述图像处理模型的模型参数,并继续训练所述图像处理模型,直至达到模型训练停止条件,获得图像处理模型的模型参数。
步骤510:向端侧设备发送所述图像处理模型的模型参数。
需要说明的是,步骤502-步骤508与上述图像处理模型的训练方法的实现方式相同,本公开实施例不在进行赘述。
在实际应用中,由于对模型进行训练需要大量的数据和较好的计算资源,端侧设备可能不具备相应的处理能力,因此,模型训练的过程可以在云侧设备实现,云侧设备在获得了图像处理模型的模型参数后,还可以将模型参数发送至端侧设备。端侧设备可以根据图像处理模型的模型参数在本地构建图像处理模型,进一步利用图像处理模型进行图像处理。
本公开实施例提供的方法,在对图像处理模型进行训练的过程中,利用文本特征提取模块,将样本指导文本通过文本特征提取模块获得对应的样本文本特征信息,将样本文本特征信息对图像处理模型的训练进行监督,提升了图像处理模型训练的效率和准确度。
另外,针对样本图像的标注信息不全的情况,采用半监督的图像标注模型的训练方法,利用部分由标注的样本图像,训练一个图像标注模型,用该图像标注模型对未标注的样本图像进行标注,从而丰富了有标注信息的样本图像的数量。
最后,将图像标注模型中的模型结构迁移到图像处理模型中,利用已经训练的图像标注模型中的模型结构继续训练图像处理模型,提升了图像处理模型的训练速度,也避免了计算资源的浪费。
参见图6,图6示出了本公开一个实施例提供的一种图像处理方法的流程图,具体包括以下步骤:
步骤602:接收用户发送的图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象。
步骤604:将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于多个目标图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
步骤606:向用户发送所述目标检测区域对应的检测结果。
需要说明的是,步骤602-步骤604的具体实现方式与上述步骤202-步骤204的实现方式相同,在本公开实施例中不再进行赘述。
在本实施方式中,是接收到用户发送的图像处理请求,在该图像处理请求中包括有图像处理任务,并且在通过上述实施例的图像处理方法处理完成、获得检测结果之后,还需要将检测结果返回给用户,以使用户根据检测结果进行相应的后续处理。
本公开一个或多个实施例提供的图像处理方法,通过图像处理模型获得目标图像对应的多个尺度特征信息提高了后续生成检测结果的准确率。在生成的检测结果中,包括了异常对象的位置信息、待检测对象的信息和指导文本,丰富了检测结果,为用户提供多维度的检测信息,提升了用户的使用体验。
本公开一实施例还提供一种计算机辅助诊断方法,该方法包括:接收用户发送的CT图像处理任务,其中,CT图像处理任务携带目标检测区域对应的多个CT图像,CT图像处理任务用于检测目标检测区域内是否存在异常对象;将多个CT图像输入至CT图像处理模型,获得目标检测区域对应的检测结果,其中,CT图像处理模型基于多个CT图像对应的多尺度特征信息生成目标检测区域对应的检测结果,检测结果包括检测标注信息、检测类别信息和检测指导文本;向用户发送目标检测区域对应的检测结果。具体地,该方法的具体实现方式于上述步骤202-步骤204的实现方式相同,在本公开实施例中不再进行赘述。
下述结合附图7,以本公开提供的图像处理方法在食管癌检测场景的应用为例,对所述图像处理方法进行进一步说明。其中,图7示出了本公开一个实施例提供的一种应用于食管癌检测场景的图像处理方法的流程图,具体包括以下步骤:
步骤702:接收CT图像处理任务,其中,所述CT图像处理任务携带食管对应的多个CT图像,所述CT图像处理任务用于检测所述食管内是否存在肿瘤。
步骤704:将所述多个CT图像输入至CT图像处理模型,获得食管对应的检测结果,其中,所述CT图像处理模型基于多个CT图像对应的多尺度特征信息生成食管对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
在本实施例中,以对食管中是否存在肿瘤为例进行解释说明,预先对CT图像处理模型进行训练。
具体的,采集的数据集中包括有1617例患者的食管癌筛查数据,数据集包括患者对应的CT图像和检测报告。其中,946名患者患有食管癌,671名用户未患有食管癌。邀请资深医生对30%患有食管癌的患者对应的CT图像中的肿瘤进行标注,并根据检测报告为各患者进行最终决定。
先根据30%标注有肿瘤位置的CT图像和其对应的检测报告,训练图像标注模型,图像标注模型中包括有多尺度特征提取模块、特征融合模块、第一特征处理模块,其中第一特征处理模块包括有异常对象分割单元和指导文本生成单元。
在获得图像标注模型后,用图像标注模型为剩余的CT图像进行标注,未患有食管癌的用户对应的CT图像对应的标注信息为“空”。
在标注完成后,提取图像标注模型中的多尺度特征提取模块、特征融合模块、第一特征处理模块,并引入新的异常对象分类单元,将异常对象分类单元与异常对象分割单元和指导文本生成单元组成第二特征处理模块,并由多尺度特征提取模块、特征融合模块和第二特征提取模块构建CT图像处理模型。
从检测报告中获取到各患者对应的患者状态(即患者是否患有食管癌)作为样本类别信息,将图像标注信息作为样本标注信息,从检测报告中提取位置指导信息和结果指导信息作为样本指导文本。
将样本类别信息、样本标注信息、样本指导文本和CT图像作为训练样本,对CT图像处理模型,直至达到CT图像处理模型的模型训练停止条件。
将训练好的CT图像处理模型即可用于食管癌的检测,将新患者的CT图像输入至该CT图像处理模型中进行预测,即可获得该患者对应的检测标注信息、检测类别信息和检测指导文本。
与上述方法实施例相对应,本公开还提供了图像处理装置实施例,图8示出了本公开一个实施例提供的一种图像处理装置的结构示意图。如图8所示,该装置包括:
接收模块802,被配置为接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象;
检测模块804,被配置为将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于多个目标图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
可选的,所述图像处理模型包括多尺度特征提取模块、特征融合模块、特征处理模块;
相应的,所述检测模块804,进一步被配置为:
将所述多个目标图像输入至所述多尺度特征提取模块,获得至少一个尺度特征信息;
将各尺度特征信息输入至所述特征融合模块,获得特征融合信息;
将所述特征融合信息输入至所述特征处理模块,获得所述目标检测区域对应的检测结果。
可选的,所述特征处理模块包括异常对象分割单元、异常对象分类单元和指导文本生成单元;
相应的,所述检测模块804,进一步被配置为:
将所述特征融合信息输入至所述异常对象分割单元,获得检测标注信息;
将所述特征融合信息输入至所述异常对象分类单元,获得检测类别信息;
将所述特征融合信息输入至所述指导文本生成单元,获得检测指导文本;
根据所述检测标注信息、检测类别信息、检测指导文本生成所述目标检测区域对应的检测结果。
可选的,所述指导文本生成单元包括异常位置文本子单元和异常结果文本子单元;
相应的,所述检测模块804,进一步被配置为:
将所述特征融合信息输入至所述异常位置文本子单元,获得位置指导信息;
将所述特征融合信息输入至所述异常结果文本子单元,获得结果指导信息;
根据所述位置指导信息和所述结果指导信息,生成检测指导文本。
可选的,所述装置还包括分割模块,被配置为:
接收图像分割任务,其中,所述图像分割任务携带目标检测区域对应的多个初始图像,所述图像分割任务用于提取所述目标检测区域对应的目标图像;
将各初始图像输入至预先训练的图像分割模型,获得所述图像分割模型输出的各初始图像对应的目标图像。
可选的,所述装置还包括训练模块,被配置为:
获取样本图像,以及所述样本图像对应的样本标注信息、样本类别信息和样本指导文本;
将所述样本图像和所述样本指导文本输入至图像处理模型,获得所述图像处理模型输出的预测标注信息、预测类别信息和文本损失值;
根据所述样本标注信息、所述样本类别信息和所述预测标注信息、所述预测类别信息计算模型损失值;
根据所述模型损失值和所述文本损失值调整所述图像处理模型的模型参数,并继续训练所述图像处理模型,直至达到模型训练停止条件。
可选的,所述训练模块,进一步被配置为:
根据所述样本标注信息和所述预测标注信息计算第一损失值;
根据所述样本类别信息和所述预测类别信息计算第二损失值;
根据所述第一损失值和所述第二损失值计算模型损失值。
可选的,所述图像处理模型包括多尺度特征提取模块、特征融合模块、特征处理模块、文本特征提取模块;
所述训练模块,进一步被配置为:
将所述样本图像和所述样本指导文本输入至图像处理模型,获得所述图像处理模型输出的预测标注信息、预测类别信息和文本损失值,包括:
将所述样本图像输入至所述多尺度特征提取模块,获得至少一个尺度特征信息;
将各尺度特征信息输入至所述特征融合模块,获得特征融合信息;
将所述样本指导文本输入至所述文本特征提取模块,获得样本文本特征信息;
将所述特征融合信息和所述样本文本特征信息输入至所述特征处理模块,获得所述样本图像对应的预测标注信息、预测类别信息和文本损失值。
可选的,所述特征处理模块包括异常对象分割单元和异常对象分类单元、指导文本生成单元;
所述训练模块,进一步被配置为:
将所述特征融合信息输入至所述异常对象分割单元,获得检测标注信息;
将所述特征融合信息输入至所述异常对象分类单元,获得检测类别信息;
将所述特征融合信息输入至所述指导文本生成单元,获得预测文本特征信息;
根据所述预测文本特征信息和所述样本文本特征信息计算获得文本损失值。
可选的,所述指导文本生成单元包括异常位置文本子单元和异常结果文本子单元,所述样本文本特征信息包括样本位置特征信息和样本结果特征信息;
所述训练模块,进一步被配置为:
将所述特征融合信息输入至所述异常位置文本子单元,获得预测位置指导信息特征;
将所述特征融合信息输入至所述异常结果文本子单元,获得预测结果指导信息特征;
根据所述样本位置特征信息、样本结果特征信息、预测位置指导信息特征和预测结果指导信息特征计算获得文本损失值。
可选的,所述样本图像包括第一样本图像和第二样本图像,其中,所述第一样本图像中标记有样本标注信息,第二样本图像中未标记样本标注信息;
所述训练模块,进一步被配置为:
获取样本图像和样本图像对应的样本指导文本,并从所述样本指导文本中提取样本类别信息;
根据第一样本图像,以及所述第一样本图像对应的样本标注信息和样本指导文本训练图像标注模型,获得用于生成标注信息的图像标注模型;
将所述第二样本输入至所述图像标注模型,获得所述图像标注模型输出的预测标注信息,将所述预测标注信息作为所述第二样本图像的样本标注信息。
可选的,所述图像标注模型中包括异常对象分割单元和指导文本生成单元;
所述训练模块,进一步被配置为:
将所述图像标注模型中的异常对象分割单元和指导文本生成单元,作为所述图像处理模型的异常对象分割单元和指导文本生成单元。
通过本公开实施例提供的图像处理装置,将待检测对象针对目标检测区域的多个目标图像输入至图像处理模型,在图像处理模型中对多个目标图像对应的3D图像进行多个尺度特征提取,并对多个尺度特征进行融合后,分别进行标注信息、类别信息和指导文本的生成,从而最终生成检测结果。通过多个尺度特征信息提高了后续生成检测结果的准确率。在生成的检测结果中,包括了异常对象的位置信息、待检测对象的信息和指导文本,丰富了检测结果,为用户提供多维度的检测信息,提升了用户的使用体验。
在对图像处理模型进行训练的过程中,利用文本特征提取模块,将样本指导文本通过文本特征提取模块获得对应的样本文本特征信息,将样本文本特征信息对图像处理模型的训练进行监督,提升了图像处理模型训练的效率和准确度。
另外,针对样本图像的标注信息不全的情况,采用半监督的图像标注模型的训练方法,利用部分由标注的样本图像,训练一个图像标注模型,用该图像标注模型对未标注的样本图像进行标注,从而丰富了有标注信息的样本图像的数量。
最后,将图像标注模型中的模型结构迁移到图像处理模型中,利用已经训练的图像标注模型中的模型结构继续训练图像处理模型,提升了图像处理模型的训练速度,也避免了计算资源的浪费。
上述为本实施例的一种图像处理装置的示意性方案。需要说明的是,该图像处理装置的技术方案与上述的图像处理方法的技术方案属于同一构思,图像处理装置的技术方案未详细描述的细节内容,均可以参见上述图像处理方法的技术方案的描述。
与上述方法实施例相对应,本公开还提供了CT图像处理装置实施例,图9示出了本公开一个实施例提供的一种CT图像处理装置的结构示意图。如图9所示,该装置包括:
接收模块902,被配置为接收CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在异常对象;
检测模块904,被配置为将所述多个CT图像输入至CT图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述CT图像处理模型基于多个CT图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
本公开一个或多个实施例提供的CT图像处理装置,在图像处理模型中对多个CT图像对应的3D图像进行多个尺度特征提取,并对多个尺度特征进行融合后,分别进行标注信息、类别信息和指导文本的生成,从而最终生成检测结果。通过多个尺度特征信息提高了后续生成检测结果的准确率。在生成的检测结果中,包括了异常对象的位置信息、待检测对象的信息和指导文本,丰富了检测结果,为用户提供多维度的检测信息,提升了用户的使用体验。
上述为本实施例的一种CT图像处理装置的示意性方案。需要说明的是,该CT图像处理装置的技术方案与上述的CT图像处理方法的技术方案属于同一构思,CT图像处理装置的技术方案未详细描述的细节内容,均可以参见上述CT图像处理方法的技术方案的描述。
与上述方法实施例相对应,本公开还提供了图像处理装置实施例,图10示出了本公开一个实施例提供的一种图像处理装置的结构示意图。如图10所示,该装置包括:
接收模块1002,被配置为接收用户发送的图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象。
检测模块1004,被配置为将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于多个目标图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
发送模块1006,被配置为向用户发送所述目标检测区域对应的检测结果。
本公开一个或多个实施例提供的图像处理装置,在图像处理模型中对多个图像对应的3D图像进行多个尺度特征提取,并对多个尺度特征进行融合后,分别进行标注信息、类别信息和指导文本的生成,从而最终生成检测结果。通过多个尺度特征信息提高了后续生成检测结果的准确率。在生成的检测结果中,包括了异常对象的位置信息、待检测对象的信息和指导文本,丰富了检测结果,为用户提供多维度的检测信息,提升了用户的使用体验。
上述为本实施例的一种图像处理装置的示意性方案。需要说明的是,该图像处理装置的技术方案与上述的图像处理方法的技术方案属于同一构思,图像处理装置的技术方案未详细描述的细节内容,均可以参见上述图像处理方法的技术方案的描述。
图11示出了根据本公开一个实施例提供的一种计算设备1100的结构框图。该计算设备1100的部件包括但不限于存储器1110和处理器1120。处理器1120与存储器1110通过总线1130相连接,数据库1150用于保存数据。
计算设备1100还包括接入设备1140,接入设备1140使得计算设备1100能够经由一个或多个网络1160通信。这些网络的示例包括公用交换电话网(PSTN,Public Switched Telephone Network)、局域网(LAN,Local Area Network)、广域网(WAN,Wide Area Network)、个域网(PAN,Personal Area Network)或诸如因特网的通信网络的组合。接入设备1140可以包括有线或无线的任何类型的网络接口(例如,网络接口卡(NIC,network interface controller))中的一个或多个,诸如IEEE802.11无线局域网(WLAN,Wireless Local Area Network)无线接口、全球微波互联接入(Wi-MAX,Worldwide Interoperability for Microwave Access)接口、以太网接口、通用串行总线(USB,Universal Serial Bus)接口、蜂窝网络接口、蓝牙接口、近场通信(NFC,Near Field Communication)。
在本公开的一个实施例中,计算设备1100的上述部件以及图11中未示出的其他部件也可以彼此相连接,例如通过总线。应当理解,图11所示的计算设备结构框图仅仅是出于示例的目的,而不是对本公开范围的限制。本领域技术人员可以根据需要,增添或替换其他部件。
计算设备1100可以是任何类型的静止或移动计算设备,包括移动计算机或移动计算设备(例如,平板计算机、个人数字助理、膝上型计算机、笔记本计算机、上网本等)、移动电话(例如,智能手机)、可佩戴的计算设备(例如,智能手表、智能眼镜等)或其他类型的移动设备,或者诸如台式计算机或个人计算机(PC,Personal Computer)的静止计算设备。计算设备1100还可以是移动式或静止式的服务器。
其中,处理器1120用于执行如下计算机可执行指令,该计算机可执行指令被处理器执行时实现上述图像处理方法、CT图像处理方法或图像处理模型的训练方法的步骤。
上述为本实施例的一种计算设备的示意性方案。需要说明的是,该计算设备的技术方案与上述的图像处理方法、CT图像处理方法或图像处理模型的训练方法的技术方案属于同一构思,计算设备的技术方案未详细描述的细节内容,均可以参见上述图像处理方法、CT图像处理方法或图像处理模型的训练方法的技术方案的描述。
本公开一实施例还提供一种计算机可读存储介质,其存储有计算机可执行指令,该计算机可执行指令被处理器执行时实现上述图像处理方法、CT图像处理方法或图像处理模型的训练方法的步骤。
上述为本实施例的一种计算机可读存储介质的示意性方案。需要说明的是,该存储介质的技术方案与上述的图像处理方法、CT图像处理方法或图像处理模型的训练方法的技术方案属于同一构思,存储介质的技术方案未详细描述的细节内容,均可以参见上述图像处理方法、CT图像处理方法或图像处理模型的训练方法的技术方案的描述。
本公开一实施例还提供一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法、CT图像处理方法或图像处理模型的训练方法的步骤。
上述为本实施例的一种计算机程序产品的示意性方案。需要说明的是,该计算机程序产品的技术方案与上述的图像处理方法、CT图像处理方法或图像处理模型的训练方法的技术方案属于同一构思,计算机程序产品的技术方案未详细描述的细节内容,均可以参见上述图像处理方法、CT图像处理方法或图像处理模型的训练方法的技术方案的描述。
上述对本公开特定实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。
所述计算机指令包括计算机程序代码,所述计算机程序代码可以为源代码形式、对象代码形式、可执行文件或某些中间形式等。所述计算机可读介质可以包括:能够携带所述计算机程序代码的任何实体或装置、记录介质、U盘、移动硬盘、磁碟、光盘、计算机存储器、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、电载波信号、电信信号以及软件分发介质等。
需要说明的是,对于前述的各方法实施例,为了简便描述,故将其都表述为一系列的动作组合,但是本领域技术人员应该知悉,本公开实施例并不受所描述的动作顺序的限制,因为依据本公开实施例,某些步骤可以采用其它顺序或者同时进行。其次,本领域技术人员也应该知悉,说明书中所描述的实施例均属于优选实施例,所涉及的动作和模块并不一定都是本公开实施例所必须的。
在上述实施例中,对各个实施例的描述都各有侧重,某个实施例中没有详述的部分,可以参见其它实施例的相关描述。
以上公开的本公开优选实施例只是用于帮助阐述本公开。可选实施例并没有详尽叙述所有的细节,也不限制该发明仅为所述的具体实施方式。显然,根据本公开实施例的内容,可作很多的修改和变化。本公开选取并具体描述这些实施例,是为了更好地解释本公开实施例的原理和实际应用,从而使所属技术领域技术人员能很好地理解和利用本公开。本公开仅受权利要求书及其全部范围和等效物的限制。

Claims (18)

  1. 一种图像处理方法,包括:
    接收图像处理任务,其中,所述图像处理任务携带目标检测区域对应的多个目标图像,所述图像处理任务用于检测所述目标检测区域内是否存在异常对象;
    将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述图像处理模型基于多个目标图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
  2. 如权利要求1所述的方法,所述图像处理模型包括多尺度特征提取模块、特征融合模块、特征处理模块;
    将所述多个目标图像输入至图像处理模型,获得所述目标检测区域对应的检测结果,包括:
    将所述多个目标图像输入至所述多尺度特征提取模块,获得至少一个尺度特征信息;
    将各尺度特征信息输入至所述特征融合模块,获得特征融合信息;
    将所述特征融合信息输入至所述特征处理模块,获得所述目标检测区域对应的检测结果。
  3. 如权利要求2所述的方法,所述特征处理模块包括异常对象分割单元、异常对象分类单元和指导文本生成单元;
    将所述特征融合信息输入至所述特征处理模块,获得所述目标检测区域对应的检测结果,包括:
    将所述特征融合信息输入至所述异常对象分割单元,获得检测标注信息;
    将所述特征融合信息输入至所述异常对象分类单元,获得检测类别信息;
    将所述特征融合信息输入至所述指导文本生成单元,获得检测指导文本;
    根据所述检测标注信息、检测类别信息、检测指导文本生成所述目标检测区域对应的检测结果。
  4. 如权利要求3所述的方法,所述指导文本生成单元包括异常位置文本子单元和异常结果文本子单元;
    将所述特征融合信息输入至所述指导文本生成单元,获得检测指导文本,包括:
    将所述特征融合信息输入至所述异常位置文本子单元,获得位置指导信息;
    将所述特征融合信息输入至所述异常结果文本子单元,获得结果指导信息;
    根据所述位置指导信息和所述结果指导信息,生成检测指导文本。
  5. 如权利要求1至4任意一项所述的方法,在接收图像处理任务之前,还包括:
    接收图像分割任务,其中,所述图像分割任务携带目标检测区域对应的多个初始图像,所述图像分割任务用于提取所述目标检测区域对应的目标图像;
    将各初始图像输入至预先训练的图像分割模型,获得所述图像分割模型输出的各初始图像对应的目标图像。
  6. 如权利要求1至5任意一项所述的方法,所述图像处理模型通过下述步骤训练获得:
    获取样本图像,以及所述样本图像对应的样本标注信息、样本类别信息和样本指导文本;
    将所述样本图像和所述样本指导文本输入至图像处理模型,获得所述图像处理模型输出的预测标注信息、预测类别信息和文本损失值;
    根据所述样本标注信息、所述样本类别信息和所述预测标注信息、所述预测类别信息计算模型损失值;
    根据所述模型损失值和所述文本损失值调整所述图像处理模型的模型参数,并继续训练所述图像处理模型,直至达到模型训练停止条件。
  7. 如权利要求6所述的方法,根据所述样本标注信息、所述样本类别信息和所述预测标注信息、所述预测类别信息计算模型损失值,包括:
    根据所述样本标注信息和所述预测标注信息计算第一损失值;
    根据所述样本类别信息和所述预测类别信息计算第二损失值;
    根据所述第一损失值和所述第二损失值计算模型损失值。
  8. 如权利要求6所述的方法,所述图像处理模型包括多尺度特征提取模块、特征融合模块、特征处理模块、文本特征提取模块;
    将所述样本图像和所述样本指导文本输入至图像处理模型,获得所述图像处理模型输出的预测标注信息、预测类别信息和文本损失值,包括:
    将所述样本图像输入至所述多尺度特征提取模块,获得至少一个尺度特征信息;
    将各尺度特征信息输入至所述特征融合模块,获得特征融合信息;
    将所述样本指导文本输入至所述文本特征提取模块,获得样本文本特征信息;
    将所述特征融合信息和所述样本文本特征信息输入至所述特征处理模块,获得所述样本图像对应的预测标注信息、预测类别信息和文本损失值。
  9. 如权利要求8所述的方法,所述特征处理模块包括异常对象分割单元和异常对象分类单元、指导文本生成单元;
    将所述特征融合信息和所述样本文本特征信息输入至所述特征处理模块,获得所述样本图像对应的预测标注信息、预测类别信息和文本损失值,包括:
    将所述特征融合信息输入至所述异常对象分割单元,获得检测标注信息;
    将所述特征融合信息输入至所述异常对象分类单元,获得检测类别信息;
    将所述特征融合信息输入至所述指导文本生成单元,获得预测文本特征信息;
    根据所述预测文本特征信息和所述样本文本特征信息计算获得文本损失值。
  10. 如权利要求9所述的方法,所述指导文本生成单元包括异常位置文本子单元和异常结果文本子单元,所述样本文本特征信息包括样本位置特征信息和样本结果特征信息;
    将所述特征融合信息输入至所述指导文本生成单元,获得预测文本特征信息,包括:
    将所述特征融合信息输入至所述异常位置文本子单元,获得预测位置指导信息特征;
    将所述特征融合信息输入至所述异常结果文本子单元,获得预测结果指导信息特征;
    相应的,根据所述预测文本特征信息和所述样本文本特征信息计算获得文本损失值,包括:
    根据所述样本位置特征信息、样本结果特征信息、预测位置指导信息特征和预测结果指导信息特征计算获得文本损失值。
  11. 如权利要求6所述的方法,所述样本图像包括第一样本图像和第二样本图像,其中,所述第一样本图像中标记有样本标注信息,第二样本图像中未标记样本标注信息;
    获取样本图像,以及所述样本图像对应的样本标注信息、样本类别信息和样本指导文本,包括:
    获取样本图像和样本图像对应的样本指导文本,并从所述样本指导文本中提取样本类别信息;
    根据第一样本图像,以及所述第一样本图像对应的样本标注信息和样本指导文本训练图像标注模型,获得用于生成标注信息的图像标注模型;
    将所述第二样本输入至所述图像标注模型,获得所述图像标注模型输出的预测标注信息,将所述预测标注信息作为所述第二样本图像的样本标注信息。
  12. 如权利要求11所述的方法,所述图像标注模型中包括异常对象分割单元和指导文本生成单元,所述方法还包括:
    将所述图像标注模型中的异常对象分割单元和指导文本生成单元,作为所述图像处理模型的异常对象分割单元和指导文本生成单元。
  13. 一种癌症的计算机辅助诊断方法,包括:
    接收计算机断层扫描CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在肿瘤;
    将所述多个CT图像输入至CT图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述CT图像处理模型基于多个CT图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本。
  14. 一种图像处理模型的训练方法,应用于云侧设备,包括:
    获取样本图像,以及所述样本图像对应的样本标注信息、样本类别信息和样本指导文本;
    将所述样本图像和所述样本指导文本输入至图像处理模型,获得所述图像处理模型输出的预测标注信息、预测类别信息和文本损失值;
    根据所述样本标注信息、所述样本类别信息和所述预测标注信息、所述预测类别信息计算模型损失值;
    根据所述模型损失值和所述文本损失值调整所述图像处理模型的模型参数,并继续训练所述图像处理模型,直至达到模型训练停止条件,获得图像处理模型的模型参数;
    向端侧设备发送所述图像处理模型的模型参数。
  15. 一种计算机辅助诊断方法,包括:
    接收用户发送的CT图像处理任务,其中,所述CT图像处理任务携带目标检测区域对应的多个CT图像,所述CT图像处理任务用于检测所述目标检测区域内是否存在异常对象;
    将所述多个CT图像输入至CT图像处理模型,获得所述目标检测区域对应的检测结果,其中,所述CT图像处理模型基于多个CT图像对应的多尺度特征信息生成目标检测区域对应的检测结果,所述检测结果包括检测标注信息、检测类别信息和检测指导文本;
    向用户发送所述目标检测区域对应的检测结果。
  16. 一种计算设备,包括:
    存储器和处理器;
    所述存储器用于存储计算机可执行指令,所述处理器用于执行所述计算机可执行指令,该计算机可执行指令被处理器执行时实现权利要求1至15任意一项所述方法的步骤。
  17. 一种计算机可读存储介质,其存储有计算机可执行指令,该计算机可执行指令被处理器执行时实现权利要求1至15任意一项所述方法的步骤。
  18. 一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现权利要求1至15任意一项所述方法的步骤。
PCT/CN2025/070596 2024-03-06 2025-01-03 图像处理方法、计算机辅助诊断方法、图像处理模型的训练方法 Pending WO2025185337A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410257868.3A CN117853490B (zh) 2024-03-06 2024-03-06 图像处理方法、图像处理模型的训练方法
CN202410257868.3 2024-03-06

Publications (2)

Publication Number Publication Date
WO2025185337A1 true WO2025185337A1 (zh) 2025-09-12
WO2025185337A8 WO2025185337A8 (zh) 2025-10-02

Family

ID=90533065

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/070596 Pending WO2025185337A1 (zh) 2024-03-06 2025-01-03 图像处理方法、计算机辅助诊断方法、图像处理模型的训练方法

Country Status (2)

Country Link
CN (2) CN120219278A (zh)
WO (1) WO2025185337A1 (zh)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120219278A (zh) * 2024-03-06 2025-06-27 阿里巴巴达摩院(杭州)科技有限公司 癌症的计算机辅助诊断方法及系统、食管癌的计算机辅助诊断方法、胃癌的计算机辅助诊断方法
CN118247284B (zh) * 2024-05-28 2024-09-13 阿里巴巴达摩院(杭州)科技有限公司 图像处理模型的训练方法、图像处理方法
CN119048419A (zh) * 2024-07-10 2024-11-29 阿里巴巴(中国)有限公司 图像处理方法、癌症的计算机辅助诊断方法

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2020088328A1 (zh) * 2018-10-31 2020-05-07 腾讯科技(深圳)有限公司 一种结肠息肉图像的处理方法和装置及系统
CN114972368A (zh) * 2021-02-26 2022-08-30 阿里巴巴集团控股有限公司 图像分割处理方法及装置、存储介质和电子设备
CN115115576A (zh) * 2022-05-18 2022-09-27 中国人民解放军北部战区总医院 一种医学图像异常信号强度的检测方法
CN116206331A (zh) * 2023-01-29 2023-06-02 阿里巴巴(中国)有限公司 图像处理方法、计算机可读存储介质以及计算机设备
CN116993663A (zh) * 2023-06-12 2023-11-03 阿里巴巴(中国)有限公司 图像处理方法、图像处理模型的训练方法
WO2024040576A1 (zh) * 2022-08-26 2024-02-29 京东方科技集团股份有限公司 目标检测方法、深度学习的训练方法、电子设备以及介质
CN117853490A (zh) * 2024-03-06 2024-04-09 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法、图像处理模型的训练方法

Family Cites Families (16)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106156766B (zh) * 2015-03-25 2020-02-18 阿里巴巴集团控股有限公司 文本行分类器的生成方法及装置
CN112292691B (zh) * 2018-06-18 2024-06-04 谷歌有限责任公司 用于使用深度学习提高癌症检测的方法与系统
CN110752028B (zh) * 2019-10-21 2026-02-03 腾讯医疗健康(深圳)有限公司 一种图像处理方法、装置、设备以及存储介质
CN113269257A (zh) * 2021-05-27 2021-08-17 中山大学孙逸仙纪念医院 一种图像分类方法、装置、终端设备及存储介质
CN113435522A (zh) * 2021-06-30 2021-09-24 平安科技(深圳)有限公司 图像分类方法、装置、设备及存储介质
CN113962274B (zh) * 2021-11-18 2022-03-08 腾讯科技(深圳)有限公司 一种异常识别方法、装置、电子设备及存储介质
CN114495058A (zh) * 2022-01-24 2022-05-13 京东鲲鹏(江苏)科技有限公司 交通标志检测方法和装置
CN114898150A (zh) * 2022-05-12 2022-08-12 山东大学 一种宫腔镜下子宫内膜图像处理系统及图像处理方法
CN115249304A (zh) * 2022-08-05 2022-10-28 腾讯科技(深圳)有限公司 检测分割模型的训练方法、装置、电子设备和存储介质
CN115953394B (zh) * 2023-03-10 2023-06-23 中国石油大学(华东) 基于目标分割的海洋中尺度涡检测方法及系统
CN116797554B (zh) * 2023-05-31 2025-05-09 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法以及装置
CN117036855A (zh) * 2023-08-02 2023-11-10 深圳市微埃智能科技有限公司 目标检测模型训练方法、装置、计算机设备和存储介质
CN117408946A (zh) * 2023-09-11 2024-01-16 阿里巴巴达摩院(杭州)科技有限公司 图像处理模型的训练方法、图像处理方法
CN117408948A (zh) * 2023-09-11 2024-01-16 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法、图像分类分割模型的训练方法
CN117475253A (zh) * 2023-10-08 2024-01-30 英特灵达信息技术(深圳)有限公司 一种模型训练方法、装置、电子设备及存储介质
CN117474918B (zh) * 2023-12-27 2024-04-16 苏州镁伽科技有限公司 异常检测方法和装置、电子设备以及存储介质

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2020088328A1 (zh) * 2018-10-31 2020-05-07 腾讯科技(深圳)有限公司 一种结肠息肉图像的处理方法和装置及系统
CN114972368A (zh) * 2021-02-26 2022-08-30 阿里巴巴集团控股有限公司 图像分割处理方法及装置、存储介质和电子设备
CN115115576A (zh) * 2022-05-18 2022-09-27 中国人民解放军北部战区总医院 一种医学图像异常信号强度的检测方法
WO2024040576A1 (zh) * 2022-08-26 2024-02-29 京东方科技集团股份有限公司 目标检测方法、深度学习的训练方法、电子设备以及介质
CN116206331A (zh) * 2023-01-29 2023-06-02 阿里巴巴(中国)有限公司 图像处理方法、计算机可读存储介质以及计算机设备
CN116993663A (zh) * 2023-06-12 2023-11-03 阿里巴巴(中国)有限公司 图像处理方法、图像处理模型的训练方法
CN117853490A (zh) * 2024-03-06 2024-04-09 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法、图像处理模型的训练方法

Also Published As

Publication number Publication date
CN120219278A (zh) 2025-06-27
CN117853490A (zh) 2024-04-09
CN117853490B (zh) 2024-05-24
WO2025185337A8 (zh) 2025-10-02

Similar Documents

Publication Publication Date Title
US11861829B2 (en) Deep learning based medical image detection method and related device
US10755411B2 (en) Method and apparatus for annotating medical image
WO2025185337A1 (zh) 图像处理方法、计算机辅助诊断方法、图像处理模型的训练方法
EP4030381B1 (en) Artificial-intelligence-based image processing method and apparatus, and device and storage medium
CN113284572B (zh) 多模态异构的医学数据处理方法及相关装置
CN107680088A (zh) 用于分析医学影像的方法和装置
CN116797554B (zh) 图像处理方法以及装置
WO2026012003A1 (zh) 图像处理方法、癌症的计算机辅助诊断方法
WO2025007942A1 (zh) 图像处理方法、图像处理模型的训练方法
WO2025246480A1 (zh) 图像处理模型的训练方法、图像处理方法
WO2025180099A1 (zh) 目标对象识别方法、对象识别模型训练方法、ct图像中的可见淋巴结检测方法、计算机辅助诊断方法、电子设备、存储介质及程序产品
WO2025016322A1 (zh) 目标检测方法
WO2024074921A1 (en) Distinguishing a disease state from a non-disease state in an image
CN118351092A (zh) 一种牙齿图像处理分析方法及其相关设备
CN117408948A (zh) 图像处理方法、图像分类分割模型的训练方法
Moosavi et al. Segmentation and classification of lungs CT-scan for detecting COVID-19 abnormalities by deep learning technique: U-Net model
Agarwal et al. Weakly-supervised lesion segmentation on CT scans using co-segmentation
CN117408946A (zh) 图像处理模型的训练方法、图像处理方法
CN115762763A (zh) 医学影像报告多标签分类方法和装置
CN120782704A (zh) 图像处理方法及图像处理模型训练方法
CN112397194B (zh) 用于生成患者病情归因解释模型的方法、装置和电子设备
CN116993663B (zh) 图像处理方法、图像处理模型的训练方法
CN117726822B (zh) 基于双分支特征融合的三维医学图像分类分割系统及方法
CN119107322A (zh) 一种轻量级图像分割方法、装置、计算机设备及存储介质
CN114863521A (zh) 表情识别方法、表情识别装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25767082

Country of ref document: EP

Kind code of ref document: A1