WO2025007940A1 - 图像处理方法、电子设备及存储介质 - Google Patents

图像处理方法、电子设备及存储介质 Download PDF

Info

Publication number
WO2025007940A1
WO2025007940A1 PCT/CN2024/103723 CN2024103723W WO2025007940A1 WO 2025007940 A1 WO2025007940 A1 WO 2025007940A1 CN 2024103723 W CN2024103723 W CN 2024103723W WO 2025007940 A1 WO2025007940 A1 WO 2025007940A1
Authority
WO
WIPO (PCT)
Prior art keywords
feature
image
area
images
monitoring
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2024/103723
Other languages
English (en)
French (fr)
Inventor
张建鹏
张灵
吕乐
张剑锋
唐禹行
许敏丰
郭建飞
周靖人
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba China Co Ltd
Original Assignee
Alibaba China Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba China Co Ltd filed Critical Alibaba China Co Ltd
Publication of WO2025007940A1 publication Critical patent/WO2025007940A1/zh
Priority to US19/419,601 priority Critical patent/US20260105717A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/20Scenes; Scene-specific elements in augmented reality scenes
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/25Determination of region of interest [ROI] or a volume of interest [VOI]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/26Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/42Global feature extraction by analysis of the whole pattern, e.g. using frequency domain transformations or autocorrelation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/44Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/762Arrangements for image or video recognition or understanding using pattern recognition or machine learning using clustering, e.g. of similar faces in social networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/80Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
    • G06V10/806Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/50Context or environment of the image
    • G06V20/52Surveillance or monitoring of activities, e.g. for recognising suspicious objects

Definitions

  • the present application relates to the field of image processing, and in particular to an image processing method, an electronic device and a storage medium.
  • image processing technology has been increasingly applied to life, such as medicine and teaching.
  • the image processing method in the prior art when processing the target part of the image, only extracts the features of the target part of the image, and then recognizes the extracted features to obtain the recognition result.
  • the image clarity is low, the accuracy of the extracted features will be low, which will lead to a low recognition accuracy rate for the target part in the image.
  • the embodiments of the present application provide an image processing method, an electronic device, and a storage medium to at least solve the technical problem of low recognition accuracy of images to be monitored in the related art.
  • an image processing method comprising: acquiring multiple images, wherein the displayed content of the images at least includes a monitoring area of a target part of an object to be monitored; performing semantic segmentation on the images to obtain a first area feature of the monitoring area in the images; determining a second area feature of the monitoring area based on a dependency relationship between multiple pre-constructed prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; identifying feature information of the monitoring area based on the first area feature and the second area feature, and determining an identification result of the monitoring area.
  • an image processing method including: responding to an input instruction acting on an operation interface, displaying multiple images on the operation interface, wherein the displayed content of the image at least includes a monitoring area of a target part of the object to be monitored; responding to an image processing instruction acting on the operation interface, displaying an identification result of the monitoring area on the operation interface, wherein the identification result is obtained by identifying feature information of the monitoring area based on a first area feature and a second area feature, the second area feature is determined based on a dependency relationship between multiple pre-constructed prototypes and the first area feature, and the first area feature is obtained by semantic segmentation of the medical image.
  • an image processing method comprising: displaying a plurality of images on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the displayed content of the image at least includes A monitoring area for a target part of a monitoring object; performing semantic segmentation on an image to obtain a first area feature of the monitoring area in the image; determining a second area feature of the monitoring area based on a dependency relationship between a plurality of pre-built prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; identifying feature information of the monitoring area based on the first area feature and the second area feature, and determining an identification result of the monitoring area; and driving a VR device or an AR device to render and display the identification result.
  • VR virtual reality
  • AR augmented reality
  • an image processing method including: acquiring multiple images by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter is multiple images, and the display content of the image at least includes a monitoring area of a target part of the object to be monitored; performing semantic segmentation on the image to obtain a first area feature of the monitoring area in the image; determining a second area feature of the monitoring area based on a dependency relationship between multiple pre-constructed prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; identifying feature information of the monitoring area based on the first area feature and the second area feature, and determining an identification result of the monitoring area; outputting the identification result by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the identification result.
  • an image processing device including: an acquisition module, used to acquire multiple images, wherein the display content of the image at least includes a monitoring area of a target part of the object to be monitored; a segmentation module, used to perform semantic segmentation on the image to obtain a first area feature of the monitoring area in the image; a first determination module, used to determine a second area feature of the monitoring area based on a dependency relationship between multiple pre-constructed prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; a second determination module, used to identify feature information of the monitoring area based on the first area feature and the second area feature, and determine an identification result of the monitoring area.
  • an acquisition module used to acquire multiple images, wherein the display content of the image at least includes a monitoring area of a target part of the object to be monitored
  • a segmentation module used to perform semantic segmentation on the image to obtain a first area feature of the monitoring area in the image
  • a first determination module used to determine a second area feature of the monitoring area based on
  • an image processing device including: a first display module, used to respond to an input instruction acting on an operation interface, and display multiple images on the operation interface, wherein the displayed content of the image at least includes a monitoring area of a target part of the object to be monitored; a second display module, used to respond to an image processing instruction acting on the operation interface, and display an identification result of the monitoring area on the operation interface, wherein the identification result is obtained by identifying feature information of the monitoring area based on a first area feature and a second area feature, the second area feature is determined based on a dependency relationship between multiple pre-constructed prototypes and the first area feature, and the first area feature is obtained by semantic segmentation of the medical image.
  • an image processing device including: a display module, used to display multiple images on a presentation screen of a virtual reality VR device or an augmented reality AR device, wherein the display content of the image at least includes a monitoring area of a target part of an object to be monitored; a segmentation module, used to perform semantic segmentation on the image to obtain a first area feature of the monitoring area in the image; a first determination module, used to determine a second area feature of the monitoring area based on a dependency relationship between multiple pre-built prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; a second determination module, used to identify feature information of the monitoring area based on the first area feature and the second area feature, and determine an identification result of the monitoring area; and a driving module, used to drive the VR device or AR device to render and display the identification result.
  • a display module used to display multiple images on a presentation screen of a virtual reality VR device or an augmented reality AR device, wherein the display content of the image at least includes a monitoring area
  • an image processing device comprising: an acquisition module, configured to acquire multiple images by calling a first interface, wherein the first interface comprises a first parameter, and a parameter value of the first parameter is multiple An image, wherein the displayed content of the image at least includes a monitoring area of a target part of an object to be monitored; a segmentation module, used for performing semantic segmentation on the image to obtain a first area feature of the monitoring area in the image; a first determination module, used for determining a second area feature of the monitoring area based on a dependency relationship between a plurality of pre-constructed prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; a second determination module, used for identifying feature information of the monitoring area based on the first area feature and the second area feature, and determining an identification result of the monitoring area; an output module, used for outputting an identification result by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of
  • an electronic device including: a memory storing an executable program; and a processor for running the program, wherein any one of the above methods is executed when the program is running.
  • a computer-readable storage medium including a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above methods.
  • a computer-aided diagnosis method for cancer comprising: acquiring multiple medical images, wherein the medical images contain a target monitoring part of an object to be monitored; performing semantic segmentation on the medical images to obtain a first part feature of the monitoring part in the medical image; determining a second part feature of the monitoring part based on a dependency relationship between multiple pre-constructed prototypes and the first part feature, wherein different prototypes are used to characterize different types of monitoring parts; diagnosing the monitoring part based on the first part feature and the second part feature to obtain a diagnosis result of the monitoring part, wherein the diagnosis result is used to characterize the presence of a malignant condition or a benign condition in the monitoring part.
  • a computer-aided diagnosis system for cancer comprising a memory, a processor, and a computer program stored on the memory and running on the processor, wherein the processor executes the computer program to perform a computer-aided diagnosis method for cancer, the method comprising: acquiring multiple medical images, wherein the medical images contain a target monitoring part of an object to be monitored; performing semantic segmentation on the medical images to obtain a first part feature of the monitoring part in the medical images; determining a second part feature of the monitoring part based on a dependency relationship between multiple pre-constructed prototypes and the first part feature, wherein different prototypes are used to characterize different types of monitoring parts; diagnosing the monitoring part based on the first part feature and the second part feature to obtain a diagnosis result of the monitoring part, wherein the diagnosis result is used to characterize whether the monitoring part is a malignant disease or a benign disease.
  • a computer-aided diagnosis method for lung cancer comprising: acquiring multiple medical images, wherein the medical images contain lung lesions; performing semantic segmentation on the medical images to obtain a first lesion feature of the lung lesions in the medical images; determining a second lesion feature of the lung lesions based on a dependency relationship between multiple pre-constructed prototypes and the first nodule feature, wherein different prototypes are used to characterize different types of lung lesions; diagnosing the lung nodules based on the first lesion feature and the second lesion feature to obtain a diagnosis result of the lung lesions, wherein the diagnosis result is used to characterize whether the lung lesions are benign or malignant.
  • a method for diagnosing lung nodules comprising: obtaining a plurality of medical images; An image is provided, wherein multiple medical images contain lung nodules; semantic segmentation is performed on the medical images to obtain a first nodule feature of the lung nodules in the medical images; a second nodule feature of the lung nodules is determined based on a dependency relationship between multiple pre-constructed prototypes and the first nodule feature, wherein different prototypes are used to characterize different types of lung nodules; the lung nodules are diagnosed based on the first nodule feature and the second nodule feature to obtain a diagnosis result of the lung nodules, wherein the diagnosis result is used to characterize whether the lung nodule is a benign nodule or a malignant nodule.
  • a lung nodule diagnostic device including: an acquisition module, used to acquire multiple medical images, wherein the multiple medical images contain lung nodules; a segmentation module, used to perform semantic segmentation on the medical images to obtain a first nodule feature of the lung nodules in the medical images; a determination module, used to determine a second nodule feature of the lung nodule based on a dependency relationship between multiple pre-constructed prototypes and the first nodule feature, wherein different prototypes are used to characterize different types of lung nodules; a diagnosis module, used to diagnose the lung nodules based on the first nodule feature and the second nodule feature to obtain a diagnosis result of the lung nodule, wherein the diagnosis result is used to characterize whether the lung nodule is a benign nodule or a malignant nodule.
  • a computer program which stores computer executable instructions, and when the instructions are executed by a processor, any one of the above methods is implemented.
  • a computer program product including a computer program, which, when executed in a computer, causes the computer to execute any one of the above methods.
  • a method is adopted in which a plurality of images are acquired; the images are semantically segmented to obtain the first regional features of the monitoring area in the image; the second regional features of the monitoring area are determined based on the dependency relationship between the plurality of pre-built prototypes and the first regional features; the feature information of the monitoring area is identified based on the first regional features and the second regional features, and the recognition result of the monitoring area is determined.
  • the present application not only semantically segments the image to obtain the regional features, but also combines the dependency relationship between the prototype and the regional features to ensure that the final regional features are more consistent with the attributes of the target part itself, and the feature extraction accuracy is higher, thereby achieving the purpose of more accurately identifying the target part of the monitored object, thereby achieving the technical effect of improving the recognition accuracy of the target part of the monitored object, thereby solving the technical problem of low recognition accuracy of the image to be monitored in the related technology.
  • FIG1 is a schematic diagram of a hardware environment of a virtual reality device according to an image processing method according to an embodiment of the present application
  • FIG2 is a structural block diagram of a computing environment of an image processing method according to an embodiment of the present application.
  • FIG3 is a flow chart of an image processing method according to Embodiment 1 of the present application.
  • FIG4 is a schematic diagram of an optional image processing method according to Embodiment 1 of the present application.
  • FIG5 is a flow chart of an image processing method according to Embodiment 2 of the present application.
  • FIG6 is a schematic diagram of an optional operation interface according to Embodiment 2 of the present application.
  • FIG7 is a flow chart of an image processing method according to Embodiment 3 of the present application.
  • FIG8 is a flow chart of an image processing method according to Embodiment 4 of the present application.
  • FIG9 is a flow chart of a method for diagnosing lung nodules according to Example 5 of the present application.
  • FIG10 is a schematic diagram of an optional comparison between reader research and artificial intelligence according to Example 5 of the present application.
  • FIG11 is a schematic diagram of the structure of an image processing device according to Embodiment 6 of the present application.
  • FIG12 is a schematic diagram of the structure of an image processing device according to Embodiment 7 of the present application.
  • FIG13 is a schematic diagram of the structure of an image processing device according to Embodiment 8 of the present application.
  • FIG14 is a schematic diagram of the structure of an image processing device according to Embodiment 9 of the present application.
  • FIG15 is a schematic diagram of the structure of a lung nodule diagnosis device according to Example 10 of the present application.
  • FIG16 is a structural block diagram of a computer terminal according to an embodiment of the present application.
  • U-shaped neural network includes a feature extraction network (encoder), which is used to extract features and obtain abstract semantic features, and a feature fusion network (decoder), which uses the abstract features encoded previously to restore the original image size, and finally obtains the segmentation result (masked image).
  • the feature extraction network and the feature fusion network can be connected to obtain a U-shaped U-shaped neural network.
  • Self-Attention Model An attention model where queries, keys, and values come from the same set of inputs, better able to understand contextual information when processing sequences.
  • Cross-Attention Model An attention model where the key and value are the same but different from the query.
  • Prototype Images with similar features.
  • the learned images are clustered in the representation space through a clustering algorithm, and the obtained class center is used as the prototype of the category.
  • an image processing method is provided. It should be noted that the flowchart in the accompanying drawing shows The steps may be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flow chart, in some cases the steps shown or described may be performed in an order different from that shown here.
  • Fig. 1 is a schematic diagram of the hardware environment of a virtual reality device according to an image processing method of an embodiment of the present application.
  • a virtual reality device 104 is connected to a terminal 106, and the terminal 106 is connected to a server 102 through a network.
  • the virtual reality device 104 is not limited to: a virtual reality helmet, virtual reality glasses, a virtual reality all-in-one machine, etc.
  • the terminal 104 is not limited to a PC, a mobile phone, a tablet computer, etc.
  • the server 102 can be a server corresponding to a media file operator, and the network includes but is not limited to: a wide area network, a metropolitan area network or a local area network.
  • the virtual reality device 104 of this embodiment includes: a memory, a processor, and a transmission device.
  • the memory is used to store an application program, which can be used to execute: obtaining multiple images, wherein the display content of the image at least includes the monitoring area of the target part of the object to be monitored; performing semantic segmentation on the image to obtain the first area feature of the monitoring area in the image; determining the second area feature of the monitoring area based on the dependency relationship between the pre-built multiple prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; identifying the feature information of the monitoring area based on the first area feature and the second area feature, and determining the recognition result of the monitoring area, thereby solving the technical problem of low recognition accuracy of the image to be monitored in the related art, and achieving the purpose of accurately identifying the target part of the object to be monitored.
  • the terminal of this embodiment can be used to display multiple images on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the displayed content of the image at least includes a monitoring area of a target part of the object to be monitored; perform semantic segmentation on the image to obtain a first area feature of the monitoring area in the image; determine a second area feature of the monitoring area based on a dependency relationship between a plurality of pre-constructed prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; identify feature information of the monitoring area based on the first area feature and the second area feature, and determine a recognition result of the monitoring area; and drive the VR device or AR device to render and display the recognition result.
  • VR virtual reality
  • AR augmented reality
  • the virtual reality device 104 of this embodiment has an eye-tracking HMD (Head Mount Display) head display and an eye-tracking module that have the same functions as those in the above embodiment, that is, the screen in the HMD head display is used to display real-time images, and the eye-tracking module in the HMD is used to obtain the real-time movement trajectory of the user's eyeballs.
  • the terminal of this embodiment obtains the user's position information and movement information in the real three-dimensional space through the tracking system, and calculates the three-dimensional coordinates of the user's head in the virtual three-dimensional space, as well as the user's field of view direction in the virtual three-dimensional space.
  • FIG2 shows an embodiment of using the AR/VR device (or mobile device) shown in FIG1 as a computing node in the computing environment 201 in a block diagram.
  • FIG2 is a structural block diagram of a computing environment of an image processing method according to an embodiment of the present application. As shown in FIG2, the computing environment 201 includes multiple (210-1, 210-2, ..., shown in the figure) computing nodes (such as servers) running on a distributed network.
  • Different computing nodes contain local processing and memory resources, and the terminal user 202 can remotely run applications or store data in the computing environment 201.
  • the application can be provided as multiple services 220-1, 220-2, 220-3 and 220-4 in the computing environment 201, representing services "A”, “D”, “E” and "H” respectively.
  • the end user 202 can provide and access services through a web browser or other software application on the client, and in some embodiments, the end user 202's provision and/or request can be provided to the entry gateway 230.
  • the entry gateway 230 can include a corresponding agent to handle the provision and/or request for the service (one or more services provided in the computing environment 201).
  • Services are provided or deployed based on various virtualization technologies supported by the computing environment 201.
  • services can be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and/or similar methods.
  • Virtual machine-based virtualization can be to simulate a real computer by initializing a virtual machine to execute programs and applications without directly contacting any actual hardware resources. While the virtual machine virtualizes the machine, according to container-based virtualization, a container can be started to virtualize the entire operating system (OS) so that multiple workloads can run on a single operating system instance.
  • OS operating system
  • a Pod e.g., a Kubernetes Pod
  • service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods).
  • the Pod may include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers).
  • One or more containers in the Pod process requests related to one or more corresponding functions of the service, and the proxy 245 generally controls network functions related to the service, such as routing, load balancing, etc.
  • Other services may also be equipped with similar Pods.
  • executing a user request from the end user 202 may require invoking one or more services in the computing environment 201, and executing one or more functions of one service requires invoking one or more functions of another service.
  • service “A” 220-1 receives a user request from the end user 202 from the ingress gateway 230, service “A” 220-1 may call service “D” 220-2, and service “D” 220-2 may request service “E” 220-3 to execute one or more functions.
  • the computing environment described above can be a cloud computing environment, where the allocation of resources is managed by the cloud service provider, allowing the development of functions without considering the implementation, adjustment or expansion of servers.
  • the computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be divided into a set of functions that can be automatically and independently scaled, rather than expanding a single hardware device to handle potential loads.
  • FIG3 is a flow chart of an image processing method according to Embodiment 1 of the present application. As shown in FIG3, the method may include the following steps:
  • Step S302 acquiring a plurality of images, wherein the display content of the images at least includes a monitoring area of a target part of the object to be monitored.
  • the above-mentioned object to be monitored may be a part of the human body, but is not limited thereto, and may also be a part of a building, etc.
  • the above-mentioned target part may be a part of the object to be monitored that needs to be specifically monitored.
  • the target part may be a nodule, blood vessel, trachea, etc. in the lungs, but is not limited thereto.
  • the target part may be a window handle, a frame corner, a middle part, etc., but is not limited thereto.
  • the above-mentioned monitoring area may be an area containing the target part of the object to be monitored, which may be referred to as a region of interest (ROI).
  • ROI region of interest
  • an original image containing the target part of the monitored object can be obtained, and secondly, the original image can be cropped based on the monitoring area to obtain an image of the target part, that is, the above-mentioned multiple images, wherein the multiple images at least contain the target part of the monitored object.
  • a CT image taken of the lungs can be obtained, wherein the CT image is a 3D image, and secondly, the CT image can be cropped based on the lungs to obtain a 3D cropped image, and then in order to be able to extract features from the 3D cropped image, the 3D cropped image can be converted into multiple 2D images (i.e., multiple images).
  • multiple original images taken of the windows can be obtained, and secondly, the multiple original images can be cropped based on the windows, that is, the above-mentioned multiple images can be obtained.
  • a video obtained by shooting the monitored object can be obtained, and then the video can be extracted multiple times to obtain multiple original images, and then the multiple original images can be cropped based on the monitoring area to obtain multiple images of the target part, that is, the above-mentioned multiple images, wherein the multiple images at least contain the target part of the monitored object.
  • a video obtained by shooting the windows can be obtained, and then the video can be extracted multiple times to obtain multiple original images, and then the multiple original images can be cropped based on the windows, that is, the above-mentioned multiple images can be obtained.
  • Step S304 semantically segment the image to obtain a first region feature of the monitoring region in the image.
  • the monitoring area image after acquiring the image, can be semantically segmented through a semantic segmentation model, that is, the first regional feature of the monitoring area can be obtained.
  • the monitoring area image can be contextually semantically segmented through a semantic segmentation model, that is, the first regional feature of the monitoring area can be obtained.
  • the monitoring area image can be contextually semantically segmented through a semantic segmentation model to obtain the preset regional features of the monitoring area, and secondly, the preset regional features can be contextually parsed through the semantic segmentation model to obtain the first regional features of the monitoring area, but it is not limited to this.
  • semantic segmentation model can be any one or more models in the related technology that can perform semantic segmentation on the monitoring area image to obtain the first area feature, and is not specifically limited in this embodiment.
  • Step S306 determining a second area feature of the monitoring area based on a dependency relationship between the pre-built multiple prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas.
  • the above-mentioned prototypes may be monitoring areas of confirmed types, and different prototypes may correspond to different types.
  • the above-mentioned types may be types of target parts.
  • the type of the monitoring area may be joints
  • the type of the monitoring area may be windows, but is not limited thereto.
  • a dependency relationship between different prototypes and the first area feature can be constructed. Secondly, after the first area feature is obtained, the prototype corresponding to the first area feature can be determined based on the dependency relationship. Then, the first area feature and the prototype corresponding to the first area feature can be processed to obtain the second area feature of the monitoring area.
  • Step S308 identifying feature information of the monitoring area based on the first area feature and the second area feature, and determining the monitoring area Domain identification results.
  • the above-mentioned recognition result may be a result obtained after recognizing the target part in the monitoring area.
  • the recognition result may be that the joint is in good condition or that the joint is in poor condition.
  • the recognition result may be that the window meets the requirements or that the window does not meet the requirements, but is not limited thereto.
  • the feature information of the monitoring area can be identified based on the first regional feature and the second regional feature to obtain an identification result of the monitoring area.
  • the feature information of the monitoring area can be identified based on the first regional feature and the second regional feature, respectively, to obtain a first identification result and a second identification result, and then the first identification result and the second identification result are compared, and the identification result with higher accuracy is selected as the final comparison result.
  • the feature information of the monitoring area can be identified based on the first regional feature and the second regional feature, respectively, to obtain a first identification result and a second identification result, and then the average of the first identification result and the second identification result is taken to obtain the final identification result.
  • the first regional feature and the second regional feature can be firstly feature-fused, and then the feature information of the monitoring area can be identified based on the fused regional feature to obtain an identification result of the monitoring area, but it is not limited to this.
  • a CT image containing human lungs can be obtained, and then the CT image can be cropped based on the lung region to obtain a 3D cropped image. Then, in order to be able to extract features from the 3D cropped image, the 3D cropped image can be converted into multiple 2D images (i.e., multiple images), wherein the multiple images all contain lung nodules, and the lung image is contextually semantically segmented through a semantic segmentation model to obtain preset regional features of the lung image. Then, the preset regional features can be contextually parsed through the semantic segmentation model to obtain a first regional feature of the lung image.
  • the prototype corresponding to the first regional feature can be determined based on a pre-constructed dependency relationship, and the first regional feature and the prototype corresponding to the first regional feature can be processed to obtain a second regional feature.
  • the feature information of the lung image can be identified based on the first regional feature and the second regional feature, respectively, to obtain a first recognition result and a second recognition result.
  • the average of the first recognition result and the second recognition result can be obtained to obtain a final recognition result.
  • the method of acquiring multiple images; semantically segmenting the images to obtain the first regional features of the monitoring area in the images; determining the second regional features of the monitoring area based on the dependency relationship between the pre-built multiple prototypes and the first regional features; identifying the feature information of the monitoring area based on the first regional features and the second regional features, and determining the recognition result of the monitoring area is adopted.
  • the present application not only semantically segments the images to obtain regional features, but also combines the dependency relationship between the prototypes and the regional features to ensure that the final regional features are more consistent with the properties of the target part itself, and the feature extraction accuracy is higher, thereby achieving the purpose of more accurately identifying the target part of the monitored object, thereby achieving the technical effect of improving the recognition accuracy of the target part of the monitored object, and thereby solving the technical problem of low recognition accuracy of the image to be monitored in the related technology.
  • semantic segmentation is performed on the image to obtain the first regional features of the monitoring area in the image, including: semantic segmentation is performed on the image to obtain the semantic segmentation result and global features of the image; feature fusion is performed on the semantic segmentation result and the image to obtain fusion features; attention processing is performed on the global features and the fusion features to obtain the first regional features feature.
  • the above-mentioned semantic segmentation result can be a semantic mask M, which can reflect whether the pixel in the image belongs to the pixel of the target area.
  • the semantic segmentation result of the pixel can be 1, but is not limited to this, and can also be 0.
  • the semantic segmentation result M contains different voxels (for example: nodules, blood vessels, etc.), belonging to the set ⁇ 0: background, 1: lungs, 2: nodules, 3: blood vessels, 4: trachea ⁇ .
  • the image can be semantically segmented by a semantic segmentation module to obtain semantic segmentation results and global features of the image.
  • the semantic segmentation results and the image can be segmented into small blocks based on the target part.
  • the semantic segmentation results and the image segmented into small blocks can be feature fused. That is, the semantic segmentation results and the image segmented into small blocks containing the same target part can be feature fused to obtain fused features.
  • the global features and the fused features can be attention processed to obtain the first region features.
  • semantic segmentation is performed on the image to obtain a semantic segmentation result and a global feature of the image, including: using an encoder module of a U-type neural network model to extract features of the image to obtain a first image feature of the image; extracting global features from the bottleneck layer of the U-type neural network model; and using the encoder module of the U-type neural network model to decode the first image feature to obtain a semantic segmentation result.
  • the encoder module of the U-type neural network model can be used to extract features of the image to obtain the first image features of the image.
  • the first image features can be decoded by the decoder module of the U-type neural network model to obtain semantic segmentation results.
  • the global features of the image can be extracted by the bottleneck layer of the U-type neural network model, wherein the bottleneck layer is located in the middle layer of the U-type neural network model.
  • features of the semantic segmentation results and the images are fused to obtain fused features, including: dividing the semantic segmentation results and the images respectively to obtain multiple sub-segmentation results and multiple sub-images; extracting features from the multiple sub-segmentation results and the multiple sub-images respectively to obtain sub-segmentation features of the multiple sub-segmentation results and sub-image features of the multiple sub-images; fusing the sub-segmentation features and the sub-image features to obtain fused features.
  • the semantic segmentation result and the image can be firstly divided based on the target part to obtain multiple sub-segmentation results and multiple sub-images, wherein the multiple sub-segmentation results and the corresponding multiple sub-images contain the same target part.
  • feature extraction can be performed on the multiple sub-segmentation results and the multiple sub-images respectively to obtain sub-segmentation features of the multiple sub-segmentation results and sub-image features of the multiple sub-images.
  • the sub-segmentation features and the sub-image features can be fused to obtain fused features.
  • attention processing is performed on the global feature and the fused feature to obtain the first region feature, including: splicing the global feature and the fused feature to obtain the first spliced feature; and performing self-attention processing on the first spliced feature using a self-attention model to obtain the first region feature.
  • the global feature and the fused feature can first be inserted into the fragment position (i.e., spliced), so that the first spliced feature token [q; t 1 , ⁇ ,t g ] ⁇ R (g+1)D can be obtained, wherein q is the global feature, t is the fused feature, R is a real number set of dimension (g+1)D, D represents the embedding dimension, and g represents the number of fused features.
  • the semantic segmentation result is firstly divided into small image blocks and concatenated with the corresponding regions of the original image. Then, a sequence is generated through image block encoding and position encoding. High-level semantic features are extracted from the neural network as the global features of the nodules.
  • the first splicing feature can be processed by self-attention through a self-attention model.
  • the first splicing feature can be processed by self-attention through a normalization function (Norm), self-attention modeling (Service Component Architecture, SCA), and a feedforward neural network (Feed Forward Network, FFN) in the self-attention model, so as to obtain the first region feature.
  • Norm normalization function
  • SCA Service Component Architecture
  • FFN feedforward neural network
  • the second area characteristics of the monitoring area are determined, including: using a cross-attention model to perform attention processing on the first area characteristics and multiple prototypes to obtain the second area characteristics.
  • the first region feature and multiple prototypes can be processed by attention through the Norm, cross-prototype attention module (CPA) and FFN in the cross-attention model, that is, the second region feature can be obtained.
  • CPA cross-prototype attention module
  • FFN cross-attention model
  • the method also includes: obtaining global features of different monitoring areas; clustering the global features of different monitoring areas to obtain multiple feature sets; and constructing multiple prototypes based on the central features of the multiple feature sets.
  • global features of different monitoring areas may be acquired, and then the global features may be clustered to obtain multiple feature sets ⁇ C 1 , . . . , C N ⁇ , where C represents the clustered features and N represents the number of features.
  • the objective function can be minimized by And the central features of multiple feature sets get multiple prototypes
  • d is the Euclidean function
  • p represents the global feature.
  • the first prototype can be expressed as P B ⁇ RN/2 ⁇ D
  • the second prototype can be expressed as P M ⁇ RN/2 ⁇ D .
  • the method after determining the second regional feature of the monitoring area based on the dependency relationship between the pre-constructed multiple prototypes and the first regional feature, the method also includes: determining a target prototype that successfully matches the second regional feature from the pre-constructed multiple prototypes; performing momentum update on the target prototype to obtain an updated regional feature; and updating the multiple prototypes based on the updated regional feature.
  • the momentum of the target prototype may be updated by the following formula:
  • the updated regional features can be obtained based on the updated target prototype, and then the multiple prototypes can be updated based on the updated regional features.
  • the characteristic information of the monitoring area is identified based on the first area feature and the second area feature to determine the identification result of the monitoring area, including: based on the global feature, the characteristic information of the monitoring area is identified to obtain the first sub-identification result; based on the first area feature, the characteristic information of the monitoring area is identified to obtain the second sub-identification result; based on the second area feature, the characteristic information of the monitoring area is identified to obtain the third sub-identification result; the first sub-identification result, the second sub-identification result and the third sub-identification result are summarized to obtain the identification result.
  • the feature information of the detection area can be identified based on the global feature, the first area feature and the second area feature by a multi-layer perceptron (MLP), and the first recognition result, the second recognition result and the third recognition result can be obtained respectively.
  • MLP multi-layer perceptron
  • the first recognition result, the second recognition result and the third recognition result can be summarized to obtain the recognition result.
  • the recognition result can be obtained by obtaining the average value of the first recognition result, the second recognition result and the third recognition result, and the more accurate recognition result among the first recognition result, the second recognition result and the third recognition result can be obtained as the recognition result, but it is not limited to this.
  • FIG4 is a schematic diagram of an optional image processing method according to Example 1 of the present application.
  • multiple images are input into a U-shaped neural network model.
  • the U-shaped neural network can decode and encode multiple images to obtain semantic segmentation results.
  • the U-shaped neural network model can output global features of multiple images through a bottleneck layer.
  • the semantic segmentation results and multiple images can be segmented to obtain multiple sub-segmentation results and multiple sub-images, and feature extraction is performed on the multiple sub-segmentation results and multiple sub-images to obtain multiple sub-segmentation features and sub-image features.
  • the multiple sub-segmentation features and sub-image features can be fused to obtain fused features, such as the white small square in FIG4.
  • the global features and The fused features are spliced (i.e., block position embedding), and the first spliced feature can be obtained, such as the white rectangular block and the oblique shaded rectangular block in Figure 4.
  • the first spliced feature can be input into the self-attention model, and the first regional feature can be obtained through normalization function, self-attention modeling and feedforward neural network.
  • the first regional feature can be input into the cross-attention model, and the second regional feature can be obtained through normalization function, cross-prototype attention and feedforward neural network.
  • the representation space of the global feature, the first regional feature and the second regional feature can be mapped to the category space through MLP to obtain the recognition results of the three MLPs, and finally the average of the recognition results of the three MLPs is obtained, that is, the final recognition result can be obtained.
  • the query (Query), key (Key) and value (Value) in the self-attention model are the same and come from the same set of inputs.
  • the query, key and value in the cross-attention model are different.
  • the target prototype that successfully matches the second regional feature can be determined, and then the target prototype is momentum updated to obtain the updated regional feature; multiple prototypes can be updated based on the updated regional feature.
  • user information including but not limited to user device information, user personal information, etc.
  • data including but not limited to data used for analysis, stored data, displayed data, etc.
  • user information including but not limited to user device information, user personal information, etc.
  • data including but not limited to data used for analysis, stored data, displayed data, etc.
  • the computer software product is stored in a storage medium (such as ROM/RAM, disk, or CD), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
  • a storage medium such as ROM/RAM, disk, or CD
  • an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
  • FIG5 is a flow chart of an image processing method according to Embodiment 2 of the present application. As shown in FIG5 , the method may include the following steps:
  • Step S502 in response to an input instruction on the operation interface, displaying a plurality of images on the operation interface, wherein the displayed content of the images at least includes a monitoring area of a target part of the object to be monitored;
  • Step S504 in response to the image processing instruction acting on the operation interface, the recognition result of the monitoring area is displayed on the operation interface, wherein the recognition result is obtained by identifying the feature information of the monitoring area based on the first region feature and the second region feature, the second region feature is determined based on the dependency relationship between multiple pre-constructed prototypes and the first region feature, and the first region feature is obtained by semantic segmentation of the medical image.
  • Figure 6 is a schematic diagram of an optional operation interface according to Example 2 of the present application.
  • the operation interface includes: an input instruction input area, a processing instruction input area and a display area.
  • the display instruction can be first input in the input instruction input area of the operation interface, and the operation interface can display multiple display images in the display area.
  • the processing instruction can be input in the processing instruction input area, and the operation interface can display the recognition result of the monitoring area in the display area, wherein the recognition result is obtained by identifying the feature information of the monitoring area based on the first area feature and the second area feature, the second area feature is determined based on the dependency relationship between the pre-constructed multiple prototypes and the first area feature, and the first area feature is obtained by semantic segmentation of the medical image.
  • an image processing method that can be applied to virtual reality scenarios such as virtual reality VR devices and augmented reality AR devices.
  • steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
  • FIG7 is a flow chart of an image processing method according to Embodiment 3 of the present application. As shown in FIG7 , the method may include the following steps:
  • Step S702 displaying a plurality of images on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the displayed content of the images at least includes a monitoring area of a target part of the object to be monitored;
  • VR virtual reality
  • AR augmented reality
  • Step S704 performing semantic segmentation on the image to obtain a first region feature of a monitoring region in the image;
  • Step S706 determining a second area feature of the monitoring area based on a dependency relationship between the pre-built multiple prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas;
  • Step S708 identifying feature information of the monitoring area based on the first area feature and the second area feature, and determining an identification result of the monitoring area;
  • Step S7010 drive the VR device or AR device to render and display the recognition result.
  • multiple images can be displayed on the presentation screen of the virtual reality VR device or the augmented reality AR device, wherein the displayed content of the image at least includes the monitoring area of the target part of the monitored object;
  • the image can be semantically segmented to obtain the first area feature of the monitoring area in the image; then, based on the dependency relationship between the pre-built multiple prototypes and the first area feature, the second area feature of the monitoring area can be determined, wherein different prototypes are used to characterize different types of monitoring areas; then, based on the first area feature and the second area feature, the feature information of the monitoring area can be identified, and the recognition result of the monitoring area can be determined; finally, the VR device or AR device can be driven to render and display the recognition result.
  • the image processing method can be applied to a hardware environment composed of a server and a virtual reality device.
  • the recognition result is displayed on the presentation screen of the virtual reality VR device or the augmented reality AR device.
  • the server can be a server corresponding to the media file operator.
  • the network includes but is not limited to: a wide area network, a metropolitan area network or a local area network.
  • the virtual reality device is not limited to: a virtual reality helmet, virtual reality glasses, a virtual reality all-in-one machine, etc.
  • the virtual reality device includes: a memory, a processor, and a transmission device.
  • the memory is used to store an application, which can be used to execute: display multiple images on a presentation screen of a virtual reality VR device or an augmented reality AR device, wherein the displayed content of the image at least includes a monitoring area of a target part of the object to be monitored; semantically segment the image to obtain a first area feature of the monitoring area in the image; determine a second area feature of the monitoring area based on a dependency relationship between multiple pre-built prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; identify feature information of the monitoring area based on the first area feature and the second area feature, and determine a recognition result of the monitoring area; drive the VR device or AR device to render and display the recognition result.
  • the above-mentioned image processing method applied in the VR device or AR device of this embodiment may include the method of the embodiment shown in Figure 3, so as to achieve the purpose of driving the VR device or AR device to display the recognition result.
  • the processor of this embodiment can call the application stored in the memory to execute the above steps through a transmission device.
  • the transmission device can receive media files sent by the server through the network, and can also be used for data transmission between the processor and the memory.
  • a head-mounted display with eye tracking is provided, wherein the screen in the HMD is used to display the displayed video images, the eye tracking module in the HMD is used to obtain the real-time movement trajectory of the user's eyes, the tracking system is used to track the user's position information and motion information in the real three-dimensional space, and the computing processing unit is used to obtain the user's real-time position and motion information from the tracking system, and calculate the three-dimensional coordinates of the user's head in the virtual three-dimensional space, as well as the user's field of view direction in the virtual three-dimensional space, etc.
  • a virtual reality device may be connected to a terminal, and the terminal and the server are connected through a network.
  • the virtual reality device is not limited to: a virtual reality helmet, virtual reality glasses, a virtual reality all-in-one machine, etc.
  • the terminal is not limited to a PC, a mobile phone, a tablet computer, etc.
  • the server may be a server corresponding to a media file operator, and the network includes but is not limited to: a wide area network, a metropolitan area network or a local area network.
  • an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
  • FIG8 is a flow chart of an image processing method according to Embodiment 4 of the present application. As shown in FIG8 , the method may include the following steps:
  • Step S802 acquiring multiple images by calling a first interface, wherein the first interface includes a first parameter, a parameter value of the first parameter is multiple images, and display content of the image at least includes a monitoring area of a target part of the object to be monitored;
  • Step S804 performing semantic segmentation on the image to obtain a first region feature of a monitoring region in the image
  • Step S806 determining a second area feature of the monitoring area based on a dependency relationship between the pre-built multiple prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas;
  • Step S808 identifying feature information of the monitoring area based on the first area feature and the second area feature, and determining an identification result of the monitoring area;
  • Step S8010 output the recognition result by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the recognition result.
  • the first interface may be an interface for acquiring multiple images from a server
  • the second interface may be an interface for sending recognition results to the server.
  • multiple images can be acquired by calling the first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter is multiple images, and the display content of the image at least includes the monitoring area of the target part of the monitored object;
  • the image can be semantically segmented to obtain the first area feature of the monitoring area in the image; then, based on the dependency relationship between the pre-constructed multiple prototypes and the first area feature, the second area feature of the monitoring area can be determined, wherein different prototypes are used to characterize different types of monitoring areas; then, based on the first area feature and the second area feature, the feature information of the monitoring area can be identified to determine the recognition result of the monitoring area; finally, the recognition result can be output by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the recognition result.
  • a method for diagnosing lung nodules is also provided. It should be noted that in the flowchart of the attached drawing The steps shown may be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flow chart, in some cases the steps shown or described may be performed in an order different from that shown.
  • FIG9 is a flow chart of a method for diagnosing lung nodules according to Example 5 of the present application. As shown in FIG9 , the method may include the following steps:
  • Step S902 Acquire multiple medical images, wherein the multiple medical images contain lung nodules.
  • the above-mentioned multiple medical images may be multiple 2D medical images, which are multiple 2D images obtained by cropping and converting the ROI area (for example, the lung area) of a computer tomography (CT) image of the human body, wherein the multiple medical images contain lung nodules.
  • CT computer tomography
  • a medical image containing lung nodules can be first obtained, then the medical image can be cropped based on the lung area to obtain a lung image, and finally the lung image can be converted to obtain multiple 2D medical images, wherein the multiple medical images contain lung nodules.
  • Step S904 performing semantic segmentation on the medical image to obtain a first nodule feature of the lung nodule in the medical image;
  • semantic segmentation can be performed on the lung image through a semantic segmentation model, that is, a first nodule feature of the lung nodule can be obtained.
  • semantic segmentation model that is, a first nodule feature of the lung nodule in the lung image can be obtained.
  • contextual semantic segmentation can be performed on the lung image through a semantic segmentation model to obtain a first preset nodule feature of the lung nodule in the lung image
  • first preset nodule feature can be contextually parsed through the semantic segmentation model to obtain the first nodule feature of the lung nodule in the lung image, but it is not limited to this.
  • nodule context information has an important impact on benign and malignant diagnosis.
  • a nodule associated with a blood transfusion vessel is more likely to be malignant than an isolated nodule. Therefore, a semantic mask m (i.e., a contextual semantic segmentation image) of an input image (i.e., a medical image) can be parsed using a U-shaped neural network (UNet), where the input image is an image obtained by cropping the pulmonary nodule ROI area in the original CT image, and the input is a three-dimensional volume composed of multiple 2D images (slices). This allows subsequent context modeling of both the nodule and its surrounding structures.
  • U-shaped neural network U-shaped neural network
  • each voxel of m belongs to ⁇ 0: background, 1: lung, 2: nodule, 3: blood vessel, 4: trachea ⁇ .
  • This segmentation process can collect comprehensive contextual information that is crucial for accurate diagnosis.
  • global features can be extracted from the bottleneck of UNet as nodule embedding q, which will be used in the subsequent diagnosis stage.
  • the context content required for distinguishing benign and malignant lung nodules includes normal lung tissue, nodules, blood vessels and trachea. This information not only reflects the shape, position, and size of the nodule itself, but also reflects the structural relationship between the nodule and the surrounding tissue.
  • This application uses a U-shaped convolutional neural network to perform pixel-level recognition on the above-mentioned context semantic information, and finally obtains a context semantic segmentation map for each nodule.
  • Step S906 determining a second nodule feature of the pulmonary nodule based on the dependency relationship between the pre-constructed multiple prototypes and the first nodule feature, wherein different prototypes are used to characterize different types of pulmonary nodules.
  • a dependency relationship between different prototypes and the first nodule feature can be constructed. Then, after obtaining the first nodule feature, the prototype corresponding to the first nodule feature can be determined based on the dependency relationship. The nodule feature and the prototype corresponding to the first nodule feature are processed to obtain a second nodule feature of the lung nodule in the lung image.
  • the present application designs a context parsing module based on an attention mechanism to deeply analyze nodules, integrate their context information, and improve the ability to distinguish benign and malignant nodules.
  • their context semantic segmentation images are divided into small image blocks, and the areas corresponding to the original image (i.e., the input image) are spliced together to obtain multiple overlapping blocks, and a string of context features (tokens) are generated through an image block encoding and position encoding.
  • high-level semantic features are extracted from the convolutional neural network as the global representation of the nodule, also known as the nodule token.
  • the long-distance dependency relationship between the nodule token and the context token is modeled, and the relevant basis for benign and malignant identification is extracted from the context information.
  • the nodule token output by the self-attention module is used as a new representation for distinguishing benign and malignant nodules.
  • the discriminative representation of nodules can be enhanced by aggregating the contextual information generated by the segmentation model.
  • the context mask can be symbolized into a set of sequences through overlapping block embedding.
  • the input image is also segmented into small blocks and embedded in context tokens to maintain the original image information.
  • position encoding is added in a learnable way to maintain position information.
  • the nodule embedding token can be pre-appended to the context sequence, represented as [q; t 1 , ⁇ ,t g ] ⁇ R (g+1)D . Where g is the number of context tokens and D represents the embedding dimension.
  • nodule embedding token at the output of the last SCA block is used as an updated nodule representation.
  • Explicitly modeling the dependency between nodule embeddings and their background structures can lead to the evolution of more discriminative representations, thereby improving the distinction between benign and malignant nodules.
  • the present application designs a nodule diagnostic knowledge prototype review module.
  • the prototype of the nodule is defined as a nodule representative with similar features.
  • the learned lung nodules are clustered in the representation space through a clustering algorithm, and the obtained class center is used as the prototype of the category.
  • the prototype has the distinction between benign and malignant.
  • the benign prototype is calculated through benign nodules with similar features, while the malignant prototype comes from malignant nodules.
  • the present application designs a cross-prototype attention module to construct the relationship between the current nodule and other prototypes, where the query is the representation from the current nodule, and the key and value are the representations from the prototype respectively.
  • the query token output by the cross-attention module serves as the final benign and malignant discrimination representation.
  • nodules can be condensed into prototypes.
  • a group of nodules i.e., multiple nodule images
  • they can be clustered into N groups ⁇ C 1 , ⁇ ,C N ⁇
  • the objective function Where d is the Euclidean distance function, p represents the nodule embedding, and the center of each cluster is taken as the prototype.
  • the prototypes can be divided into benign and malignant groups, represented by PB ⁇ RN /2 ⁇ D and PM ⁇ RN /2 ⁇ D .
  • the inter-layer dependencies between nodules and external prototypes can also be captured.
  • This enables PARE to explore relevant identification bases beyond a single nodule.
  • the present application designs a cross-prototype attention (CPA) module, which utilizes nodule embeddings as queries and prototypes as keywords and values. It allows nodule embeddings to selectively attend to the most relevant parts of the prototype sequence.
  • PARE is a model for diagnosing lung nodules proposed in this application, which includes three parts: context segmentation, nodule internal context analysis, and nodule prototype recall learning.
  • the present application can also update the prototype online, thereby allowing the prototype to quickly adjust to changes in nodule embeddings.
  • a nodule embedding q with data (x, y) select its nearest prototype and then update it using the following momentum rule:
  • is the momentum factor, which is generally set to 0.95, but not limited to this.
  • Momentum update can help accelerate convergence and improve generalization ability.
  • Step S908 diagnose the lung nodule based on the first nodule feature and the second nodule feature to obtain a diagnosis result of the lung nodule, wherein the diagnosis result is used to characterize whether the lung nodule is a benign nodule or a malignant nodule.
  • the lung nodule after obtaining the first nodule feature and the second nodule feature, can be diagnosed based on the first nodule feature and the second nodule feature to obtain a diagnosis result of the lung nodule.
  • the lung nodule can be diagnosed based on the first nodule feature and the second nodule feature, respectively, to obtain a first diagnosis result and a second diagnosis result, and then the first diagnosis result and the second diagnosis result are compared to select the diagnosis result with higher accuracy as the final diagnosis result.
  • the lung nodule can be diagnosed based on the first nodule feature and the second nodule feature, respectively, to obtain a first diagnosis result and a second diagnosis result, and then the average of the first diagnosis result and the second diagnosis result is taken to obtain the final diagnosis result.
  • the first nodule feature and the second nodule feature can be firstly fused, and then the lung nodule can be diagnosed based on the fused nodule feature to obtain a diagnosis result of the lung nodule, but is not limited to this.
  • the deep supervision signals are respectively in the nodule global representation output by the convolutional neural network, the nodule representation output by the self-context attention module, and the nodule representation output by the cross-prototype attention module.
  • MLP Multi-Layer Perception
  • This application proposes a radiologist motivation method that simulates the diagnostic process of radiologists, which consists of context parsing and prototype review modules.
  • the context parsing module first segments the context structure of the nodule, and then aggregates the context information to understand the nodule more comprehensively.
  • the prototype review module uses prototype-based learning to compress previously learned situations into prototypes for comparative analysis, which are updated online in a momentum manner during training. Based on these two modules, the method of this application utilizes both the inherent characteristics of nodules and the external knowledge accumulated from other nodules to achieve a reasonable diagnosis.
  • Table 1 is an ablation comparison of optional hyperparameters according to Example 5 of the present application.
  • the present application investigates the impact of different configurations on the PARE performance on the validation set, including transformer layers, number of prototypes, embedding dimension, and deep supervision.
  • higher AUC scores can be obtained by increasing the number of transformer layers, increasing the number of prototypes, doubling the channel size of token embeddings, or using deep classification supervision.
  • the hyperparameters include: transformer layer (L), number of prototypes (N), embedding dimension (D), and deep supervision (DS).
  • Table 2 is the effectiveness of different optional modules according to Example 5 of the present application.
  • the present application investigates the ablation study of different methods/modules on the validation set and observes the following results: (1) Pure segmentation methods perform better than pure classification methods, mainly because it enables greater supervision at the pixel level, (2) Joint segmentation and classification outperform any single method, indicating the complementary effect of the two tasks, (3) Both context parsing and prototype comparison help improve performance over a strong baseline, thereby proving the effectiveness of the two modules, and (4) Segmenting more context structures (such as blood vessels, lungs, and trachea) provides a slight improvement compared to just segmenting nodules.
  • MT stands for multi-task learning.
  • Context intra-frame context parsing.
  • Prototype review between prototypes. * Indicates that only nodule masks are used in the segmentation task.
  • Table 3 is a comparison of different methods on an optional NNLST according to Example 5 of the present application and an internal test set, including a pure classification-based method, a pure segmentation-based method, and a multi-task-based method.
  • a hierarchical evaluation was performed in two test groups based on nodule size distribution. These results indicate that the segmentation-based method outperforms the pure classification method, mainly due to its superior ability to segment contextual structures.
  • the multi-task based CA-Net outperforms any single-task method.
  • the PARE method of this application surpasses most other methods.
  • the overall AUC is further improved to 0.931 for both datasets.
  • LUNGx External evaluation of LUNGx: This application uses LUNGx as an external test to evaluate the generalization of PARE. It is worth noting that these compared methods have never been trained on LUNGx.
  • Table 4 is an optional comparison with other methods on LUNGx according to Example 5 of the present application. It can be seen from Table 4 that the AUC of the PARE model of the present application is as high as 0.801, which is 2% higher than the method DAR.
  • the present application also conducted a reader study to compare PARE with two experienced radiologists with 8 and 13 years of experience in lung nodule diagnosis, respectively.
  • Figure 10 is a schematic diagram of an optional reader study compared with artificial intelligence according to Example 5 of the present application. The results of Figure 3 show that the method of the present application has achieved performance comparable to that of radiologists.
  • user information including but not limited to user device information, user personal information, etc.
  • data including but not limited to data used for analysis, stored data, displayed data, etc.
  • user information including but not limited to user device information, user personal information, etc.
  • data including but not limited to data used for analysis, stored data, displayed data, etc.
  • LDCT and NCCT The model of this application is trained on a mixture of LDCT and NCCT datasets, and can perform well in low-dose and regular-dose applications.
  • Table 5 shows the generalization performance comparison of the models obtained under the three training configurations. It can be seen from Table 5 that the models trained on the LDCT or NCCT dataset alone cannot be generalized well to other modalities, with at least a 6% drop in AUC. However, the hybrid training method of this application performs well on LDCT and NCCT. The best performance was achieved with almost no performance degradation.
  • a method for computer-aided diagnosis of cancer comprising:
  • the monitored part is diagnosed based on the first part feature and the second part feature to obtain a diagnosis result of the monitored part, wherein the diagnosis result is used to characterize whether the monitored part is a malignant disease or a benign disease.
  • a computer-aided diagnosis system for cancer comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to perform a computer-aided method for cancer, the method comprising:
  • the monitored part is diagnosed based on the first part feature and the second part feature to obtain a diagnosis result of the monitored part, wherein the diagnosis result is used to characterize whether the monitored part is a malignant disease or a benign disease.
  • a method for computer-aided diagnosis of lung cancer comprising:
  • the lung nodule is diagnosed based on the first lesion feature and the second lesion feature to obtain a diagnosis result of the lung lesion, wherein the diagnosis result is used to characterize whether the lung lesion is a benign disease or a malignant disease.
  • the lung lesions may include lung nodules. If the diagnosis result is that the lung lesions are malignant diseases, the lung lesions may be caused by lung cancer, so that the doctor or the diagnostic device can determine the treatment plan according to the diagnosis result.
  • FIG11 is a structural schematic diagram of an image processing device according to Embodiment 6 of the present application. As shown in FIG11 , the device includes: an acquisition module 1102, a segmentation module 1104, a first determination module 1106, and a second determination module 1108.
  • the acquisition module is used to acquire multiple images, wherein the display content of the image at least includes the monitoring area of the target part of the object to be monitored; the segmentation module is used to perform semantic segmentation on the image to obtain the first area feature of the monitoring area in the image; the first determination module is used to determine the second area feature of the monitoring area based on the dependency relationship between the pre-built multiple prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; the second determination module is used to identify the feature information of the monitoring area based on the first area feature and the second area feature, and determine the recognition result of the monitoring area.
  • the acquisition module, segmentation module, first determination module and second determination module correspond to steps S302 to S308 in Example 1, and the four modules are the same as the instances and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above-mentioned Example 1.
  • the above-mentioned modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above-mentioned modules may also be part of the device and run in the AR/VR device provided in Example 1.
  • the segmentation module includes: a segmentation unit, a fusion unit and a first processing unit.
  • the segmentation unit is used to perform semantic segmentation on the image to obtain the semantic segmentation result and global features of the image;
  • the fusion unit is used to perform feature fusion on the semantic segmentation result and the image to obtain fusion features;
  • the first processing unit is used to perform attention processing on the global features and the fusion features to obtain the first regional features.
  • the segmentation unit includes: a first extraction subunit, a second extraction subunit and a decoding subunit.
  • the first extraction subunit is used to use the encoder module of the U-type neural network model to extract features of the image to obtain the first image features of the image;
  • the second extraction subunit is used to extract global features from the bottleneck layer of the U-type neural network model;
  • the decoding subunit is used to use the encoder module of the U-type neural network model to decode the first image features to obtain semantic segmentation results.
  • the fusion unit includes: a cutting subunit, a third extraction subunit and a fusion subunit.
  • the segmentation sub-unit is used to segment the semantic segmentation results and the image respectively, so as to obtain multiple sub-segmentation results and multiple sub-images;
  • the third extraction sub-unit is used to extract features from the multiple sub-segmentation results and the multiple sub-images respectively, so as to obtain sub-segmentation features of the multiple sub-segmentation results and sub-image features of the multiple sub-images;
  • the fusion sub-unit is used to fuse the sub-segmentation features and the sub-image features to obtain the fusion features.
  • the first processing unit includes: a splicing subunit and a processing subunit.
  • the splicing subunit is used to splice the global feature and the fusion feature to obtain a first splicing feature; the processing subunit is used to perform self-attention processing on the first splicing feature using a self-attention model to obtain a first regional feature.
  • the first determination module includes: a second processing unit.
  • the second processing unit is used to use the cross attention model to perform attention processing on the first region feature and multiple prototypes to obtain the second region feature.
  • the first determination module further includes: an acquisition unit, a clustering unit and a construction unit.
  • the acquisition unit is used to obtain the global features of different monitoring areas;
  • the clustering unit is used to cluster the global features of different monitoring areas to obtain multiple feature sets; and
  • the construction unit is used to construct multiple prototypes based on the central features of multiple feature sets.
  • the device after determining the second area feature of the monitoring area based on the dependency relationship between the pre-built multiple prototypes and the first area feature, the device also includes: a third determination module, a first update module and a second update module.
  • the third determination module is used to determine the target prototype that successfully matches the second regional feature from multiple pre-built prototypes; the first update module is used to perform momentum update on the target prototype to obtain updated regional features; the second update module is used to update multiple prototypes based on the updated regional features.
  • the second determination module includes: a third processing unit, a fourth processing unit, a fifth processing unit and a summarizing unit.
  • the third processing unit is used to identify the characteristic information of the monitoring area based on the global feature to obtain a first sub-identification result;
  • the fourth processing unit is used to identify the characteristic information of the monitoring area based on the first area feature to obtain a second sub-identification result;
  • the fifth processing unit is used to identify the characteristic information of the monitoring area based on the second area feature to obtain a third sub-identification result;
  • the summary unit is used to summarize the first sub-identification result, the second sub-identification result and the third sub-identification result to obtain an identification result.
  • an image processing device for implementing the above-mentioned image processing method is also provided.
  • FIG12 is a schematic diagram of the structure of an image processing device according to Embodiment 7 of the present application. As shown in FIG12 , the device includes: a first display module 1202 and a second display module 1204 .
  • the first display module is used to respond to the input command on the operation interface and display multiple pictures on the operation interface.
  • An image wherein the displayed content of the image at least includes a monitoring area of a target part of the object to be monitored;
  • the second display module is used to respond to an image processing instruction acting on an operation interface, and display a recognition result of the monitoring area on the operation interface, wherein the recognition result is obtained by identifying feature information of the monitoring area based on a first area feature and a second area feature, the second area feature is determined based on a dependency relationship between a plurality of pre-constructed prototypes and the first area feature, and the first area feature is obtained by semantic segmentation of the medical image.
  • first display module and the second display module correspond to steps S502 to S504 in Example 2, and the two modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 2.
  • the modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the modules may also be part of a device and run in the AR/VR device provided in Example 1.
  • FIG13 is a structural schematic diagram of an image processing device according to Embodiment 8 of the present application. As shown in FIG13 , the device includes: a display module 1302, a segmentation module 1304, a first determination module 1306, a second determination module 1308 and a driving module 13010.
  • the display module is used to display multiple images on the presentation screen of a virtual reality VR device or an augmented reality AR device, wherein the display content of the image at least includes the monitoring area of the target part of the object to be monitored;
  • the segmentation module is used to perform semantic segmentation on the image to obtain the first area feature of the monitoring area in the image;
  • the first determination module is used to determine the second area feature of the monitoring area based on the dependency relationship between multiple pre-built prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas;
  • the second determination module is used to identify the feature information of the monitoring area based on the first area feature and the second area feature, and determine the recognition result of the monitoring area;
  • the driving module is used to drive the VR device or AR device to render and display the recognition result.
  • the above display module, segmentation module, first determination module, second determination module and driving module correspond to steps S702 to S7010 in Example 3, and the five modules are the same as the examples and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above Example 3.
  • the above modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the above modules can also be run in the AR/VR device provided in Example 1 as part of the device.
  • FIG14 is a structural schematic diagram of an image processing device according to Embodiment 9 of the present application. As shown in FIG14 , the device includes: an acquisition module 1402, a segmentation module 1404, a first determination module 1406, a second determination module 1408 and an output module 14010.
  • the acquisition module is used to acquire multiple images by calling the first interface, wherein the first interface includes a first parameter A number of images, the parameter value of the first parameter is a plurality of images, and the display content of the image at least includes the monitoring area of the target part of the object to be monitored; a segmentation module, used to perform semantic segmentation on the image to obtain a first area feature of the monitoring area in the image; a first determination module, used to determine a second area feature of the monitoring area based on a dependency relationship between a plurality of pre-built prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; a second determination module, used to identify feature information of the monitoring area based on the first area feature and the second area feature, and determine an identification result of the monitoring area; an output module, used to output the identification result by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the identification result.
  • the acquisition module, segmentation module, first determination module, second determination module and output module correspond to steps S802 to S8010 in Example 4, and the five modules are the same as the examples and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above-mentioned Example 4.
  • the above-mentioned modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above-mentioned modules may also be part of the device and run in the AR/VR device provided in Example 1.
  • FIG15 is a schematic diagram of the structure of a pulmonary nodule diagnosis device according to Embodiment 10 of the present application. As shown in FIG15 , the device includes: an acquisition module 1502, a segmentation module 1504, a determination module 1506, and a diagnosis module 1508.
  • the acquisition module is used to acquire multiple medical images, wherein the multiple medical images contain lung nodules;
  • the segmentation module is used to perform semantic segmentation on the medical images to obtain the first nodule feature of the lung nodules in the medical images;
  • the determination module is used to determine the second nodule feature of the lung nodules based on the dependency relationship between the pre-built multiple prototypes and the first nodule feature, wherein different prototypes are used to characterize different types of lung nodules;
  • the diagnosis module is used to diagnose the lung nodules based on the first nodule feature and the second nodule feature to obtain the diagnosis result of the lung nodules, wherein the diagnosis result is used to characterize whether the lung nodule is a benign nodule or a malignant nodule.
  • the acquisition module, segmentation module, determination module, and diagnosis module described above correspond to steps S902 to S908 in Example 5, and the five modules and corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 5.
  • the modules or units described above may be hardware components or software components stored in a memory and processed by one or more processors, and the modules described above may also be run in the AR/VR device provided in Example 1 as part of the device.
  • the embodiment of the present application may provide an AR/VR device, which may be any AR/VR device in an AR/VR device group.
  • the AR/VR device may also be replaced by a terminal device such as a mobile terminal.
  • the AR/VR device may be located in at least one network device among a plurality of network devices of a computer network.
  • the above-mentioned AR/VR device can execute the program code of the following steps in the image processing method: display multiple images on the presentation screen of the virtual reality VR device or the augmented reality AR device, wherein the display content of the image at least includes the monitoring area of the target part of the object to be monitored; perform semantic segmentation on the image to obtain the first area feature of the monitoring area in the image; determine the second area feature of the monitoring area based on the dependency relationship between the pre-built multiple prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; identify the feature information of the monitoring area based on the first area feature and the second area feature, and determine the recognition result of the monitoring area; drive the VR device or AR device to render and display the recognition result.
  • Figure 16 is a block diagram of a computer terminal according to an embodiment of the present application.
  • the computer terminal A may include: one or more (only one is shown in the figure) processors 1602, a memory 1604, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display.
  • the memory can be used to store software programs and modules, such as program instructions/modules corresponding to the image processing method and device in the embodiment of the present application.
  • the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned image processing method.
  • the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory.
  • the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
  • the processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: obtain multiple images, wherein the display content of the image at least includes the monitoring area of the target part of the object to be monitored; perform semantic segmentation on the image to obtain the first area feature of the monitoring area in the image; determine the second area feature of the monitoring area based on the dependency relationship between the pre-constructed multiple prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; identify the feature information of the monitoring area based on the first area feature and the second area feature, and determine the recognition result of the monitoring area.
  • the processor may also execute the program code of the following steps: performing semantic segmentation on the image to obtain the semantic segmentation result and global features of the image; performing feature fusion on the semantic segmentation result and the image to obtain fused features; performing attention processing on the global features and the fused features to obtain first region features.
  • the processor may also execute the program code of the following steps: using the encoder module of the U-type neural network model to extract features of the image to obtain the first image features of the image; extracting global features from the bottleneck layer of the U-type neural network model; and using the encoder module of the U-type neural network model to decode the first image features to obtain semantic segmentation results.
  • the processor may also execute the following steps of program code: segmenting the semantic segmentation result and the image respectively to obtain a plurality of sub-segmentation results and a plurality of sub-images; extracting features from the plurality of sub-segmentation results and the plurality of sub-images respectively to obtain sub-segmentation features of the plurality of sub-segmentation results and sub-image features of the plurality of sub-images; extracting features from the ...; extracting features from the sub-segmentation results and the plurality of sub-images respectively; extracting features from the sub-segmentation results and the plurality of sub-images respectively; extracting features from the sub-segmentation results and the plurality of sub-images respectively; extracting features from the sub-segmentation results and the plurality of sub-images respectively; extracting features from the sub-s
  • the feature of the image is fused with the sub-image feature to obtain the fused feature.
  • the processor may also execute the following program code: concatenating the global feature and the fused feature to obtain a first concatenated feature; and performing self-attention processing on the first concatenated feature using a self-attention model to obtain a first regional feature.
  • the processor may also execute the program code of the following steps: using a cross attention model to perform attention processing on the first region feature and multiple prototypes to obtain a second region feature.
  • the processor may also execute program codes of the following steps: obtaining global features of different monitoring areas; clustering global features of different monitoring areas to obtain multiple feature sets; and constructing multiple prototypes based on central features of multiple feature sets.
  • the processor may also execute program code of the following steps: determining a target prototype that successfully matches the second regional feature from multiple pre-built prototypes; performing momentum update on the target prototype to obtain an updated regional feature; and updating multiple prototypes based on the updated regional feature.
  • the processor may also execute the program code of the following steps: obtaining a first sub-identification result based on the feature information of the monitoring area identified by the global feature; obtaining a second sub-identification result based on the feature information of the monitoring area identified by the first area feature; obtaining a third sub-identification result based on the feature information of the monitoring area identified by the second area feature; and summarizing the first sub-identification result, the second sub-identification result and the third sub-identification result to obtain an identification result.
  • a method for acquiring multiple images; performing semantic segmentation on the images to obtain the first regional features of the monitoring area in the images; determining the second regional features of the monitoring area based on the dependency relationship between the pre-built multiple prototypes and the first regional features; identifying the feature information of the monitoring area based on the first regional features and the second regional features, and determining the recognition result of the monitoring area.
  • the present application not only performs semantic segmentation on the images to obtain regional features, but also combines the dependency relationship between the prototypes and the regional features to ensure that the final regional features are more consistent with the properties of the target part itself, and the feature extraction accuracy is higher, thereby achieving the purpose of more accurately identifying the target part of the monitored object, thereby achieving the technical effect of improving the recognition accuracy of the target part of the monitored object, and thereby solving the technical problem of low recognition accuracy of the image to be monitored in the related technology.
  • the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, and other terminal devices.
  • FIG. 16 does not limit the structure of the above-mentioned electronic device.
  • the computer terminal A may also include more or fewer components (such as a network interface, a display device, etc.) than those shown in FIG. 16, or have a configuration different from that shown in FIG. 16.
  • a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
  • the embodiment of the present application also provides a computer readable storage medium.
  • the computer can be used to store the program code executed by the image processing method provided in the above-mentioned embodiment 1.
  • the above-mentioned computer-readable storage medium can be located in any computer terminal in the AR/VR device terminal group in the AR/VR device network, or in any mobile terminal in the mobile terminal group.
  • the computer-readable storage medium is configured to store program codes for executing the following steps: acquiring multiple images, wherein the displayed content of the images at least includes a monitoring area of a target part of the object to be monitored; performing semantic segmentation on the images to obtain a first area feature of the monitoring area in the images; determining a second area feature of the monitoring area based on a dependency relationship between multiple pre-constructed prototypes and the first area feature, wherein different prototypes are used to characterize different types of monitoring areas; identifying feature information of the monitoring area based on the first area feature and the second area feature, and determining an identification result of the monitoring area.
  • the computer-readable storage medium is also configured to store program codes for executing the following steps: performing semantic segmentation on the image to obtain semantic segmentation results and global features of the image; performing feature fusion on the semantic segmentation results and the image to obtain fused features; performing attention processing on the global features and the fused features to obtain first region features.
  • the computer-readable storage medium is also configured to store program codes for executing the following steps: using the encoder module of the U-type neural network model to extract features of the image to obtain the first image features of the image; extracting global features from the bottleneck layer of the U-type neural network model; using the encoder module of the U-type neural network model to decode the first image features to obtain semantic segmentation results.
  • the computer-readable storage medium is also configured to store program codes for executing the following steps: segmenting the semantic segmentation results and the image respectively to obtain multiple sub-segmentation results and multiple sub-images; extracting features from the multiple sub-segmentation results and the multiple sub-images respectively to obtain sub-segmentation features of the multiple sub-segmentation results and sub-image features of the multiple sub-images; and fusing the sub-segmentation features and the sub-image features to obtain fused features.
  • the computer-readable storage medium is also configured to store program codes for executing the following steps: splicing the global feature and the fusion feature to obtain a first spliced feature; performing self-attention processing on the first spliced feature using a self-attention model to obtain a first regional feature.
  • the computer-readable storage medium is also configured to store program code for executing the following steps: using a cross-attention model to perform attention processing on the first region feature and multiple prototypes to obtain a second region feature.
  • the computer-readable storage medium is also configured to store program codes for executing the following steps: obtaining global features of different monitoring areas; clustering the global features of different monitoring areas to obtain multiple feature sets; and constructing multiple prototypes based on the central features of the multiple feature sets.
  • the computer-readable storage medium is also configured to store program code for performing the following steps: determining a target prototype that successfully matches the second regional feature from multiple pre-built prototypes; performing momentum update on the target prototype to obtain an updated regional feature; and updating multiple prototypes based on the updated regional feature.
  • the computer-readable storage medium is further configured to store program codes for executing the following steps: identifying feature information of the monitoring area based on the global feature to obtain a first sub-recognition result; identifying feature information of the monitoring area based on the first feature to obtain a first sub-recognition result; The feature information of the monitoring area is identified based on the feature of the second area to obtain a second sub-identification result; the feature information of the monitoring area is identified based on the feature of the second area to obtain a third sub-identification result; the first sub-identification result, the second sub-identification result and the third sub-identification result are summarized to obtain an identification result.
  • An embodiment of the present application further provides a computer program product, including a computer program, which, when executed in a computer, enables the computer to execute the method provided in the embodiment of the present application.
  • the embodiment of the present application also provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the method provided in the embodiment of the present application.
  • the serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
  • the disclosed technical content can be implemented in other ways.
  • the device embodiments described above are only schematic.
  • the division of the units is only a logical function division.
  • multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
  • Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
  • the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
  • each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
  • the above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
  • the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
  • the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.
  • the aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc., which can store program code.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Evolutionary Computation (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Computing Systems (AREA)
  • General Health & Medical Sciences (AREA)
  • Software Systems (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • Computational Linguistics (AREA)
  • Data Mining & Analysis (AREA)
  • Molecular Biology (AREA)
  • General Engineering & Computer Science (AREA)
  • Mathematical Physics (AREA)
  • Image Processing (AREA)
  • Alarm Systems (AREA)

Abstract

本申请公开了一种图像处理方法、电子设备及存储介质。其中,该方法应用于图像处理领域,包括:获取多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果。本申请解决了相关技术中对待监测图像进行识别的识别准确率低的技术问题。

Description

图像处理方法、电子设备及存储介质
本申请要求于2023年07月04日提交中国专利局、申请号为202310814294.0、申请名称为“图像处理方法、电子设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及图像处理领域,具体而言,涉及一种图像处理方法、电子设备及存储介质。
背景技术
随着科技的飞速发展,针对图像的处理技术也越来越多地应用到了生活领域中,例如医学领域和教学领域。现有技术中的图像处理方法,在对图像的目标部位进行处理时,仅仅只是对图像的目标部位图像进行特征提取,然后对提取到的特征进行识别,得到识别结果,而当图像的清晰度较低时,提取到的特征的准确度就会较低,进而会导致对图像中的目标部位进行识别的识别准确率低。
针对上述的问题,目前尚未提出有效的解决方案。
发明内容
本申请实施例提供了一种图像处理方法、电子设备及存储介质,以至少解决相关技术中对待监测图像进行识别的识别准确率低的技术问题。
根据本申请实施例的一个方面,提供了一种图像处理方法,包括:获取多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果。
根据本申请实施例的另一方面,还提供了一种图像处理方法,包括:响应作用于操作界面上的输入指令,在操作界面上显示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;响应作用于操作界面上的图像处理指令,在操作界面上显示监测区域的识别结果,其中,识别结果是基于第一区域特征和第二区域特征识别监测区域的特征信息得到的,第二区域特征是基于预先构建的多个原型和第一区域特征之间的依赖关系确定的,第一区域特征是对医学图像进行语义分割得到的。
根据本申请实施例的另一方面,还提供了一种图像处理方法,包括:在虚拟现实VR设备或增强现实AR设备的呈现画面上展示多张图像,其中,图像的显示内容至少包含待 监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;驱动VR设备或AR设备渲染展示识别结果。
根据本申请实施例的另一方面,还提供了一种图像处理方法,包括:通过调用第一接口获取多张图像,其中,第一接口包括第一参数,第一参数的参数值为多张图像,图像的显示内容至少包含待监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;通过调用第二接口输出识别结果,其中,第二接口包括第二参数,第二参数的参数值为识别结果。
根据本申请实施例的另一方面,还提供了一种图像处理装置,包括:获取模块,用于获取多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;分割模块,用于对图像进行语义分割,得到图像中监测区域的第一区域特征;第一确定模块,用于基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;第二确定模块,用于基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果。
根据本申请实施例的另一方面,还提供了一种图像处理装置,包括:第一显示模块,用于响应作用于操作界面上的输入指令,在操作界面上显示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;第二显示模块,用于响应作用于操作界面上的图像处理指令,在操作界面上显示监测区域的识别结果,其中,识别结果是基于第一区域特征和第二区域特征识别监测区域的特征信息得到的,第二区域特征是基于预先构建的多个原型和第一区域特征之间的依赖关系确定的,第一区域特征是对医学图像进行语义分割得到的。
根据本申请实施例的另一方面,还提供了一种图像处理装置,包括:显示模块,用于在虚拟现实VR设备或增强现实AR设备的呈现画面上展示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;分割模块,用于对图像进行语义分割,得到图像中监测区域的第一区域特征;第一确定模块,用于基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;第二确定模块,用于基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;驱动模块,用于驱动VR设备或AR设备渲染展示识别结果。
根据本申请实施例的另一方面,还提供了一种图像处理装置,包括:获取模块,用于通过调用第一接口获取多张图像,其中,第一接口包括第一参数,第一参数的参数值为多 张图像,图像的显示内容至少包含待监测对象的目标部位的监测区域;分割模块,用于对图像进行语义分割,得到图像中监测区域的第一区域特征;第一确定模块,用于基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;第二确定模块,用于基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;输出模块,用于通过调用第二接口输出识别结果,其中,第二接口包括第二参数,第二参数的参数值为识别结果。
根据本申请实施例的另一方面,还提供了一种电子设备,包括:存储器,存储有可执行程序;处理器,用于运行程序,其中,程序运行时执行上述任意一项的方法。
根据本申请实施例的另一方面,还提供了一种计算机可读存储介质,计算机可读存储介质包括存储的可执行程序,其中,在可执行程序运行时控制计算机可读存储介质所在设备执行上述任意一项的方法。
根据本申请实施例的另一方面,还提供了一种癌症计算机辅助诊断方法,包括:获取多张医学图像,其中,所述医学图像包含待监测对象的目标监测部位;对所述医学图像进行语义分割,得到所述医学图像中所述监测部位的第一部位特征;基于预先构建的多个原型和所述第一部位特征之间的依赖关系,确定所述监测部位的第二部位特征,其中,不同原型用于表征不同类型的监测部位;基于所述第一部位特征和所述第二部位特征对所述监测部位进行诊断,得到所述监测部位的诊断结果,其中,所述诊断结果用于表征所述监测部位存在恶性病症或良性病症。
根据本申请实施例的另一方面,还提供了一种癌症的计算机辅助诊断系统,包括存储器、处理器以及存储在所述存储器上并在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序可用于执行一种癌症计算机辅助诊断方法,所述方法包括:获取多张医学图像,其中,所述医学图像包含待监测对象的目标监测部位;对所述医学图像进行语义分割,得到所述医学图像中所述监测部位的第一部位特征;基于预先构建的多个原型和所述第一部位特征之间的依赖关系,确定所述监测部位的第二部位特征,其中,不同原型用于表征不同类型的监测部位;基于所述第一部位特征和所述第二部位特征对所述监测部位进行诊断,得到所述监测部位的诊断结果,其中,所述诊断结果用于表征所述监测部位是恶性病症或良性病症。
根据本申请实施例的另一方面,还提供了一种肺癌计算机辅助诊断方法,包括:获取多张医学图像,其中,所述医学图像包含肺部病变;对所述医学图像进行语义分割,得到所述医学图像中肺部病变的第一病变特征;基于预先构建的多个原型和所述第一结节特征之间的依赖关系,确定所述肺部病变的第二病变特征,其中,不同原型用于表征不同类型的肺部病变;基于所述第一病变特征和所述第二病变特征对所述肺结节进行诊断,得到所述肺部病变的诊断结果,其中,所述诊断结果用于表征所述肺部病变是良性病症或恶性病症。
根据本申请实施例的另一方面,还提供了一种肺结节诊断方法,包括:获取多张医学 图像,其中,多张医学图像包含肺结节;对医学图像进行语义分割,得到医学图像中肺结节的第一结节特征;基于预先构建的多个原型和第一结节特征之间的依赖关系,确定肺结节的第二结节特征,其中,不同原型用于表征不同类型的肺结节;基于第一结节特征第二结节特征对肺结节进行诊断,得到肺结节的诊断结果,其中,诊断结果用于表征肺结节是良性结节或恶性结节。
根据本申请实施例的另一方面,还提供了一种肺结节诊断装置,包括:获取模块,用于获取多张医学图像,其中,多张医学图像包含肺结节;分割模块,用于对医学图像进行语义分割,得到医学图像中肺结节的第一结节特征;确定模块,用于基于预先构建的多个原型和第一结节特征之间的依赖关系,确定肺结节的第二结节特征,其中,不同原型用于表征不同类型的肺结节;诊断模块,用于基于第一结节特征第二结节特征对肺结节进行诊断,得到肺结节的诊断结果,其中,诊断结果用于表征肺结节是良性结节或恶性结节。
根据本申请实施例的另一方面,还提供了一种计算机程序,其存储有计算机可执行指令,该指令被处理器执行时实现上述任意一项的方法。
根据本申请实施例的另一方面,还提供了一种计算机程序产品,包括计算机程序,当所述计算机程序在计算机中执行时,令计算机执行上述任意一项的方法。在本申请实施例中,采用获取多张图像;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果的方式。容易注意到的是,本申请不仅对图像进行了语义分割,得到区域特征,还结合了原型与区域特征之间的依赖关系,确保最终的区域特征更符合目标部位本身的属性,特征提取准确度更高,进而达到了更准确的对待监测对象的目标部位进行识别的目的,从而实现了提高待监测对象的目标部位的识别准确率的技术效果,进而解决了相关技术中对待监测图像进行识别的识别准确率低的技术问题。
容易注意到的是,上面的通用描述和后面的详细描述仅仅是为了对本申请进行举例和解释,并不构成对本申请的限定。
附图说明
此处所说明的附图用来提供对本申请的进一步理解,构成本申请的一部分,本申请的示意性实施例及其说明用于解释本申请,并不构成对本申请的不当限定。在附图中:
图1是根据本申请实施例的一种图像处理方法的虚拟现实设备的硬件环境的示意图;
图2是根据本申请实施例的一种图像处理方法的计算环境的结构框图;
图3是根据本申请实施例1的一种图像处理方法的流程图;
图4是根据本申请实施例1的一种可选的图像处理方法的示意图;
图5是根据本申请实施例2的一种图像处理方法的流程图;
图6是根据本申请实施例2的一种可选的操作界面的示意图;
图7是根据本申请实施例3的一种图像处理方法的流程图;
图8是根据本申请实施例4的一种图像处理方法的流程图;
图9是根据本申请实施例5的一种肺结节诊断方法的流程图;
图10是根据本申请实施例5的一种可选的读者研究与人工智能对比的示意图;
图11是根据本申请实施例6的一种图像处理装置的结构示意图;
图12是根据本申请实施例7的一种图像处理装置的结构示意图;
图13是根据本申请实施例8的一种图像处理装置的结构示意图;
图14是根据本申请实施例9的一种图像处理装置的结构示意图;
图15是根据本申请实施例10的一种肺结节诊断装置的结构示意图;
图16是根据本申请实施例的一种计算机终端的结构框图。
具体实施方式
为了使本技术领域的人员更好地理解本申请方案,下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本申请一部分的实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都应当属于本申请保护的范围。
需要说明的是,本申请的说明书和权利要求书及上述附图中的术语“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便这里描述的本申请的实施例能够以除了在这里图示或描述的那些以外的顺序实施。此外,术语“包括”和“具有”以及他们的任何变形,意图在于覆盖不排他的包含,例如,包含了一系列步骤或单元的过程、方法、系统、产品或设备不必限于清楚地列出的那些步骤或单元,而是可包括没有清楚地列出的或对于这些过程、方法、产品或设备固有的其它步骤或单元。
首先,在对本申请实施例进行描述的过程中出现的部分名词或术语适用于如下解释:
U型神经网络:包括特征提取网络(编码器),用于特征提取,得到抽象语义特征,以及特征融合网络(解码器),利用前面编码的抽象特征来恢复到原图尺寸的过程,最终得到分割结果(掩码图片),其中,特征提取网络和特征融合网络连接后可以得到形状为U型的U型神经网络。
自注意力模型:查询、键和值来自同一组输入的注意力模型,在处理序列时能更好地理解上下文信息。
交叉注意力模型:键和值相同但与查询不同的注意力模型。
原型:具有相似特征的图像,通过聚类算法将学习过的图像在表征空间上进行聚类,得到的类中心作为该类别的原型。
实施例1
根据本申请实施例,提供了一种图像处理方法,需要说明的是,在附图的流程图示出 的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
图1是根据本申请实施例的一种图像处理方法的虚拟现实设备的硬件环境的示意图。如图1所示,虚拟现实设备104与终端106相连接,终端106与服务器102通过网络进行连接,上述虚拟现实设备104并不限定于:虚拟现实头盔、虚拟现实眼镜、虚拟现实一体机等,上述终端104并不限定于PC、手机、平板电脑等,服务器102可以为媒体文件运营商对应的服务器,上述网络包括但不限于:广域网、城域网或局域网。
可选地,该实施例的虚拟现实设备104包括:存储器、处理器和传输装置。存储器用于存储应用程序,该应用程序可以用于执行:获取多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果,从而解决了相关技术中对待监测图像进行识别的识别准确率低的技术问题,达到了准确对待监测对象的目标部位进行识别的目的。
该实施例的终端可以用于执行在虚拟现实(Virtual Reality,简称为VR)设备或增强现实(Augmented Reality,简称为AR)设备的呈现画面上展示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;驱动VR设备或AR设备渲染展示识别结果。
可选地,该实施例的虚拟现实设备104带有的眼球追踪的HMD(Head Mount Display,头戴式显示器)头显与眼球追踪模块与上述实施例中的作用相同,也即,HMD头显中的屏幕,用于显示实时的画面,HMD中的眼球追踪模块,用于获取用户眼球的实时运动轨迹。该实施例的终端通过跟踪系统获取用户在真实三维空间的位置信息与运动信息,并计算出用户头部在虚拟三维空间中的三维坐标,以及用户在虚拟三维空间中的视野朝向。
图1示出的硬件结构框图,不仅可以作为上述AR/VR设备(或移动设备)的示例性框图,还可以作为上述服务器的示例性框图,一种可选实施例中,图2以框图示出了使用上述图1所示的AR/VR设备(或移动设备)作为计算环境201中计算节点的一种实施例。图2是根据本申请实施例的一种图像处理方法的计算环境的结构框图,如图2所示,计算环境201包括运行在分布式网络上的多个(图中采用210-1,210-2,…,来示出)计算节点(如服务器)。不同计算节点都包含本地处理和内存资源,终端用户202可以在计算环境201中远程运行应用程序或存储数据。应用程序可以作为计算环境201中的多个服务220-1,220-2,220-3和220-4进行提供,分别代表服务“A”,“D”,“E”和“H”。
终端用户202可以通过客户端上的web浏览器或其他软件应用程序提供和访问服务,在一些实施例中,可以将终端用户202的供应和/或请求提供给入口网关230。入口网关230可以包括一个相应的代理来处理针对服务(计算环境201中提供的一个或多个服务)的供应和/或请求。
服务是根据计算环境201支持的各种虚拟化技术来提供或部署的。在一些实施例中,可以根据基于虚拟机(Virtual Machine,VM)的虚拟化、基于容器的虚拟化和/或类似的方式提供服务。基于虚拟机的虚拟化可以是通过初始化虚拟机来模拟真实的计算机,在不直接接触任何实际硬件资源的情况下执行程序和应用程序。在虚拟机虚拟化机器的同时,根据基于容器的虚拟化,可以启动容器来虚拟化整个操作系统(Operating System,OS),以便多个工作负载可以在单个操作系统实例上运行。
在基于容器虚拟化的一个实施例中,服务的若干容器可以被组装成一个Pod(例如,Kubernetes Pod)。举例来说,如图2所示,服务220-2可以配备一个或多个Pod 240-1,240-2,…,240-N(统称为Pod)。Pod可以包括代理245和一个或多个容器242-1,242-2,…,242-M(统称为容器)。Pod中一个或多个容器处理与服务的一个或多个相应功能相关的请求,代理245通常控制与服务相关的网络功能,如路由、负载均衡等。其他服务也可以配备类似的Pod。
在操作过程中,执行来自终端用户202的用户请求可能需要调用计算环境201中的一个或多个服务,执行一个服务的一个或多个功能需要调用另一个服务的一个或多个功能。如图2所示,服务“A”220-1从入口网关230接收终端用户202的用户请求,服务“A”220-1可以调用服务“D”220-2,服务“D”220-2可以请求服务“E”220-3执行一个或多个功能。
上述的计算环境可以是云计算环境,资源的分配由云服务提供上管理,允许功能的开发无需考虑实现、调整或扩展服务器。该计算环境允许开发人员在不构建或维护复杂基础设施的情况下执行响应事件的代码。服务可以被分割完成一组可以自动独立伸缩的功能,而不是扩展单个硬件设备来处理潜在的负载。
在上述运行环境下,本申请提供了如图3所示的图像处理方法。需要说明的是,该实施例的图像处理方法可以由图1所示实施例的移动终端执行。图3是根据本申请实施例1的一种图像处理方法的流程图。如图3所示,该方法可以包括如下步骤:
步骤S302,获取多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域。
上述的待监测对象可以为人体的部位,但不仅限于此,还可以是建筑物的局部等。上述的目标部位可以是待监测对象的需要具体监测的部位。例如,当待监测对象为人体的肺部时,目标部位可以是肺部的结节、血管、气管等,但不仅限于此。当待监测对象为建筑物墙上的窗户时,目标部位可以是窗户的把手、边框角落、中间部位等,但不仅限于此。上述的监测区域可以是包含有待监测对象的目标部位的区域,可以称为感兴趣区域(Region Of Interest,ROI)。
在一种可选的实施例中,当需要对待监测对象的目标部位进行监测时,首先可以获取包含待监测对象的目标部位的原始图像,其次可以基于监测区域对原始图像进行裁剪,得到目标部位的图像,即上述的多张图像,其中,多张图像中至少包含有待监测对象的目标部位。例如,当需要对人体肺部的结节进行监测时,首先可以获取对肺部拍摄得到的CT图像,其中,CT图像为3维(Dimensional)图像,其次可以基于肺部对CT图像进行裁剪得到3D裁剪图像,然后为了能够对3D裁剪图像进行特征提取,可以将3D裁剪图像转换为多张2D图像(即多张图像)。又例如,当需要对建筑物墙上的窗户进行监测时,首先可以获取对窗户进行拍摄得到的多张原始图像,其次可以基于窗户对多张原始图像进行裁剪,即可以得到上述的多张图像。
在另一种可选的实施例中,当需要对待监测对象的目标部位进行监测时,首先可以获取对待监测对象拍摄后得到的视频,其次可以对该视频进行多次抽取得到多张原始图像,然后可以基于监测区域对多张原始图像进行裁剪,得到多张目标部位的图像,即上述的多张图像,其中,多张图像中至少包含由待监测对象的目标部位。例如,当需要对建筑物墙上的窗户进行监测时,首先可以获取对窗户进行拍摄得到视频,其次可以对该视频进行多次抽取得到多张原始图像,然后可以基于窗户对多张原始图像进行裁剪,即可以得到上述的多张图像。
步骤S304,对图像进行语义分割,得到图像中监测区域的第一区域特征。
在一种可选的实施例中,当获取到图像后,可以通过语义分割模型对监测区域图像进行语义分割,即可以得到监测区域的第一区域特征。例如,可以通过语义分割模型对监测区域图像进行上下文语义分割,即可以得到监测区域的第一区域特征。又例如,首先可以通过语义分割模型对监测区域图像进行上下文语义分割,得到监测区域的预设区域特征,其次还可以通过语义分割模型对预设区域特征进行上下文解析,进而可以得到监测区域的第一区域特征,但不仅限于此。
需要说明的是,上述的语义分割模型可以是相关技术中,任意一种或多种能够对监测区域图像进行语义分割,以得到第一区域特征的模型,在本实施例中不做具体限定。
步骤S306,基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域。
上述的原型可以是已经确认类型的监测区域,不同原型对应的类型也不同。上述的类型可以是目标部位的类型。例如,当监测区域包含的目标部位为人类的关节时,该监测区域的类型可以是关节,当监测区域包含的目标部位为建筑物的窗户时,该监测区域的类型可以是窗户,但不仅限于此。
在一种可选的实施例中,首先可以构建不同原型与第一区域特征之间的依赖关系,其次当得到第一区域特征后,可以基于依赖关系确定第一区域特征对应的原型,然后可以对第一区域特征以及第一区域特征对应的原型进行处理,得到监测区域的第二区域特征。
步骤S308,基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区 域的识别结果。
上述的识别结果可以是对监测区域中的目标部位进行识别后得到的结果。例如,当监测区域中的目标部位为关节时,识别结果可以是关节状态好,也可以是关节状态差,当监测区域中的目标部位为窗户时,识别结果可以是窗户符合要求,也可以是窗户不符合要求,但不仅限于此。
在一种可选的实施例中,当得到第一区域特征和第二区域特征后,可以基于第一区域特征和第二区域特征识别监测区域的特征信息,得到监测区域的识别结果。例如,可以分别基于第一区域特征和第二区域特征对监测区域的特征信息进行识别,得到第一识别结果和第二识别结果,然后将第一识别结果和第二识别结果进行对比,选取准确度较高的识别结果作为最终的对比结果。又例如,可以分别基于第一区域特征和第二区域特征对监测区域的特征信息进行识别,得到第一识别结果和第二识别结果,然后取第一识别结果和第二识别结果的平均值,得到最终的识别结果。又例如,还可以首先将第一区域特征和第二区域特征进行特征融合,其次可以基于融合后的区域特征对监测区域的特征信息进行识别,得到监测区域的识别结果,但不仅限于此。
例如,当需要对人类的肺部的结节进行监测时,首先可以获取包含人类肺部的CT图像,其次可以基于肺部区域对CT图像进行裁剪得到3D裁剪图像,然后为了能够对3D裁剪图像进行特征提取,可以将3D裁剪图像转换为多张2D图像(即多张图像)其中,多张图像中均包含有肺部的结节,并通过语义分割模型对肺部图像进行上下文语义分割,得到肺部图像的预设区域特征,然后还可以通过语义分割模型对预设区域特征进行上下文解析,进而可以得到肺部图像的第一区域特征,然后可以基于预先构建的依赖关系确定第一区域特征对应的原型,并对第一区域特征以及第一区域特征对应的原型进行处理得到第二区域特征,最后可以分别基于第一区域特征和第二区域特征对肺部图像的特征信息进行识别,得到第一识别结果和第二识别结果,最后可以获取第一识别结果和第二识别结果的平均值,得到最终的识别结果。
在本申请实施例中,在本申请实施例中,采用获取多张图像;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果的方式。容易注意到的是,本申请不仅对图像进行了语义分割,得到区域特征,还结合了原型与区域特征之间的依赖关系,确保最终的区域特征更符合目标部位本身的属性,特征提取准确度更高,进而达到了更准确的对待监测对象的目标部位进行识别的目的,从而实现了提高待监测对象的目标部位的识别准确率的技术效果,进而解决了相关技术中对待监测图像进行识别的识别准确率低的技术问题。
本申请上述实施例中,对图像进行语义分割,得到图像中监测区域的第一区域特征,包括:对图像进行语义分割,得到图像的语义分割结果和全局特征;对语义分割结果和图像进行特征融合,得到融合特征;对全局特征和融合特征进行注意力处理,得到第一区域 特征。
上述的语义分割结果可以是语义掩码M,能够体现图像中的像素是否属于目标区域的像素。当像素属于目标区域的像素时,该像素的语义分割结果可以是1,但不仅限于此,还可以是0。当监测对象为人体肺部时,语义分割结果M中包含有不同体素(例如:结节、血管等),属于集合{0:背景,1:肺部,2:结节,3:血管,4:气管}。
在一种可选的实施例中,首先可以通过语义分割模块对图像进行语义分割,得到图像的语义分割结果和全局特征,其次可以基于目标部位将语义分割结果和图像分割成小块,然后将分割成小块的语义分割结果和图像进行特征融合,也即,可以对包含相同目标部位的,分割成小块后的语义分割结果和图像进行特征融合,得到融合特征,最后可以对全局特征和融合特征进行注意力处理,即可以得到第一区域特征。
本申请上述实施例中,对图像进行语义分割,得到图像的语义分割结果和全局特征,包括:利用U型神经网络模型的编码器模块对图像进行特征提取,得到图像的第一图像特征;从U型神经网络模型的瓶颈层中,提取全局特征;利用U型神经网络模型的编码器模块,对第一图像特征进行解码,得到语义分割结果。
在一种可选的实施例中,首先可以通过U型神经网络模型的编码器模块对图像进行特征提取,得到图像的第一图像特征,其次可以通过U型神经网络模型的解码器模块对第一图像特征进行解码,即可以得到语义分割结果,此外,还可以通过U型神经网络模型的瓶颈层提取图像的全局特征,其中,瓶颈层位于U型神经网络模型的中间层。
本申请上述实施例中,对语义分割结果和图像进行特征融合,得到融合特征,包括:分别对语义分割结果和图像进行切分,得到多个子分割结果和多个子图像;分别对多个子分割结果和多个子图像进行特征提取,得到多个子分割结果的子分割特征,以及多个子图像的子图像特征;将子分割特征和子图像特征进行融合,得到融合特征。
在一种可选的实施例中,在得到图像的语义分割结果后,首先可以基于目标部位对语义分割结果和图像进行切分,得到多个子分割结果和多个子图像,其中,多个子分割结果和对应的多个子图像包含有相同的目标部位,其次可以分别对多个子分割结果和多个子图像进行特征提取,得到多个子分割结果的子分割特征和多个子图像的子图像特征,然后可以将子分割特征和子图像特征进行融合,可以得到融合特征。
本申请上述实施例中,对全局特征和融合特征进行注意力处理,得到第一区域特征,包括:将全局特征和融合特征进行拼接,得到第一拼接特征;利用自注意力模型对第一拼接特征进行自注意力处理,得到第一区域特征。
在一种可选的实施例中,首先可以将全局特征和融合特征进行碎片位置插入(即拼接),即可以得到第一拼接特征token为[q;t1,···,tg]∈R(g+1)D,其中,q为全局特征,t为融合特征,R是维度为(g+1)D的实数集,D表示嵌入维度,g表示融合特征的数量。
在另一种可选的实施例中,首先语义分割结果被切分成小的图像块,和原始图像对应的区域拼接到一起,其次可以通过图像块编码和位置编码,生成一串序列。同时,从卷积 神经网络中抽取出高级语义的特征作为结节的全局特征。
在另一种可选的实施例中,在得到第一拼接特征后,可以通过自注意力模型对第一拼接特征进行自注意力处理,例如,可以通过自注意力模型中的归一化函数(Norm)、自注意力建模(Service Component Architecture,SCA),以及前馈神经网络(Feed Forward Network,FFN)对第一拼接特征进行自注意力处理,即可以得到第一区域特征。
本申请上述实施例中,基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,包括:利用交叉注意力模型对第一区域特征和多个原型进行注意力处理,得到第二区域特征。
在一种可选的实施例中,可以通过交叉注意力模型中的Norm、交叉原型注意力模块(CPA)以及FFN对第一区域特征和多个原型进行注意力处理,即可以得到第二区域特征。
本申请上述实施例中,该方法还包括:获取不同监测区域的全局特征;对不同监测区域的全局特征进行聚类,得到多个特征集合;基于多个特征集合的中心特征,构建多个原型。
在一种可选的实施例中,首先可以获取不同监测区域的全局特征,其次可以对全局特征进行聚类得到多个特征集合{C1,···,CN},其中,C表示聚类后的特征,N表示特征数量。
在另一种可选的实施例中,可以通过最小化目标函数以及多个特征集合的中心特征得到多个原型其中,d为欧几里得函数,p表示全局特征。需要说明的是,第一原型可以表示为PB∈RN/2×D,第二原型可以表示为PM∈RN/2×D
本申请上述实施例中,在基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征之后,该方法还包括:从预先构建的多个原型中,确定与第二区域特征匹配成功的目标原型;对目标原型进行动量更新,得到更新后的区域特征;基于更新后的区域特征对多个原型进行更新。
在一种可选的实施例中,可以通过以下公式对目标原型进行动量更新:
其中,为动量更新后的第一原型,为动量更新后的第二原型,λ为动量因数,一般设置为0.95,但不仅限于此,其中,动量更新可以帮助加速收敛并提高泛化能力。
在另一种可选的实施例中,当得到更新后的目标原型后,可以基于更新后的目标原型得到更新后的区域特征,然后可以基于更新后的区域特征对多个原型进行更新。
本申请上述实施例中,基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果,包括:基于全局特征识别监测区域的特征信息,得到第一子识别结果;基于第一区域特征识别监测区域的特征信息,得到第二子识别结果;基于第二区域特征识别监测区域的特征信息,得到第三子识别结果;对第一子识别结果、第二子识别结果和第三子识别结果进行汇总,得到识别结果。
在一种可选的实施例中,可以通过多层感知器(Multi-Layer Perception,MLP)分别基于全局特征、第一区域特征以及第二区域特征对检测区域的特征信息进行识别,可以分别得到第一识别结果、第二识别结果以及第三识别结果,最后可以将第一识别结果、第二识别结果以及第三识别结果进行汇总,进而可以得到识别结果。例如,可以获取第一识别结果、第二识别结果和第三识别结果的平均值得到识别结果,还可以获取第一识别结果、第二识别结果以及第三识别结果中的较准确的识别结果作为识别结果,但不仅限于此。
图4是根据本申请实施例1的一种可选的图像处理方法的示意图,如图4所示,首先将多张图像输入至U型神经网络模型中,U型神经网络可以对多张图像进行解码和编码后得到语义分割结果,同时U型神经网络模型可以通过瓶颈层输出多张图像的全局特征,其次可以将语义分割结果和多张图像进行切分,得到多个子分割结果和多个子图像,并将多个子分割结果和多个子图像进行特征提取,得到多个子分割特征和子图像特征,然后可以将多个子分割特征和子图像特征进行融合,可以得到融合特征,如图4中的白色小方块,然后可以将全局特征和融合特征进行拼接(即块位置嵌入),即可以得到第一拼接特征,如图4中的白色矩形块和斜线阴影矩形块,然后可以将第一拼接特征输入至自注意力模型中,经过归一化函数、自注意力建模以及前馈神经网络得到第一区域特征,然后可以将第一区域特征输入至交叉注意力模型中,通过归一化函数、交叉原型注意力以及前馈神经网络得到第二区域特征,最后可以通过MLP从全局特征、第一区域特征以及第二区域特征的表征空间映射到类别空间,得到三个MLP的识别结果,最后获取三个MLP的识别结果的均值,即可以得到最终的识别结果。其中,自注意力模型中的查询(Query)、键(Key)和值(Value)相同,来自同一组输入。而交叉注意力模型中的查询与键和值不同。
需要说明的是,图中的平行四边形中有多个特征集合,其中,箭头连接的多边形为原型,在确定原型后,可以确定与第二区域特征匹配成功的目标原型,然后对目标原型进行动量更新,得到更新后的区域特征;基于更新后的区域特征可以对多个原型进行更新。
需要说明的是,本申请所涉及的用户信息(包括但不限于用户设备信息、用户个人信息等)和数据(包括但不限于用于分析的数据、存储的数据、展示的数据等),均为经用户授权或者经过各方充分授权的信息和数据,并且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准,并提供有相应的操作入口,供用户选择授权或者拒绝。
需要说明的是,对于前述的各方法实施例,为了简单描述,故将其都表述为一系列的动作组合,但是本领域技术人员应该知悉,本申请并不受所描述的动作顺序的限制,因为依据本申请,某些步骤可以采用其他顺序或者同时进行。其次,本领域技术人员也应该知悉,说明书中所描述的实施例均属于优选实施例,所涉及的动作和模块并不一定是本申请所必须的。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到根据上述实施例的方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件。基于这样的 理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端设备(可以是手机,计算机,服务器,或者网络设备等)执行本申请各个实施例所述的方法。
实施例2
根据本申请实施例,还提供了一种图像处理方法,需要说明的是,在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
图5是根据本申请实施例2的一种图像处理方法的流程图。如图5所示,该方法可以包括如下步骤:
步骤S502,响应作用于操作界面上的输入指令,在操作界面上显示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;
步骤S504,响应作用于操作界面上的图像处理指令,在操作界面上显示监测区域的识别结果,其中,识别结果是基于第一区域特征和第二区域特征识别监测区域的特征信息得到的,第二区域特征是基于预先构建的多个原型和第一区域特征之间的依赖关系确定的,第一区域特征是对医学图像进行语义分割得到的。
图6是根据本申请实施例2的一种可选的操作界面的示意图,如图6所示,操作界面包括:输入指令输入区域,处理指令输入区域以及显示区域,当需要对图像中的待监测对象的目标部位进行监测时,首先可以在操作界面的输入指令输入区域输入显示指令,则操作界面可以在显示区域中显示多张显示图像,其次可以在处理指令输入区域中输入处理指令,操作界面可以在显示区域显示监测区域的识别结果,其中,识别结果是基于第一区域特征和第二区域特征识别监测区域的特征信息得到的,第二区域特征是基于预先构建的多个原型和第一区域特征之间的依赖关系确定的,第一区域特征是对医学图像进行语义分割得到的。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例3
根据本申请实施例,还提供了一种可以应用于虚拟现实VR设备、增强现实AR设备等虚拟现实场景下的图像处理方法,需要说明的是,在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
图7是根据本申请实施例3的一种图像处理方法的流程图。如图7所示,该方法可以包括如下步骤:
步骤S702,在虚拟现实VR设备或增强现实AR设备的呈现画面上展示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;
步骤S704,对图像进行语义分割,得到图像中监测区域的第一区域特征;
步骤S706,基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;
步骤S708,基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;
步骤S7010,驱动VR设备或AR设备渲染展示识别结果。
在一种可选的实施例中,当需要对待监测对象的目标部位进行监测时,首先可以在虚拟现实VR设备或增强现实AR设备的呈现画面上展示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;其次可以对图像进行语义分割,得到图像中监测区域的第一区域特征;然后可以基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;然后可以基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;最后可以驱动VR设备或AR设备渲染展示识别结果。
可选地,在本实施例中,上述图像处理方法可以应用于由服务器、虚拟现实设备所构成的硬件环境中。在虚拟现实VR设备或增强现实AR设备的呈现画面上展示识别结果,服务器可以为媒体文件运营商对应的服务器,上述网络包括但不限于:广域网、城域网或局域网,上述虚拟现实设备并不限定于:虚拟现实头盔、虚拟现实眼镜、虚拟现实一体机等。
可选地,虚拟现实设备包括:存储器、处理器和传输装置。存储器用于存储应用程序,该应用程序可以用于执行:在虚拟现实VR设备或增强现实AR设备的呈现画面上展示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;驱动VR设备或AR设备渲染展示识别结果。
需要说明的是,该实施例的上述应用在VR设备或AR设备中的图像处理方法可以包括图3所示实施例的方法,以实现驱动VR设备或AR设备展示识别结果的目的。
可选地,该实施例的处理器可以通过传输装置调用上述存储器存储的应用程序以执行上述步骤。传输装置可以通过网络接收服务器发送的媒体文件,也可以用于上述处理器与存储器之间的数据传输。
可选地,在虚拟现实设备中,带有眼球追踪的头戴式显示器,该HMD头显中的屏幕,用于显示展示的视频画面,HMD中的眼球追踪模块,用于获取用户眼球的实时运动轨迹,跟踪系统,用于追踪用户在真实三维空间的位置信息与运动信息,计算处理单元,用于从跟踪系统中获取用户的实时位置与运动信息,并计算出用户头部在虚拟三维空间中的三维坐标,以及用户在虚拟三维空间中的视野朝向等。
在本申请实施例中,虚拟现实设备可以与终端相连接,终端与服务器通过网络进行连接,上述虚拟现实设备并不限定于:虚拟现实头盔、虚拟现实眼镜、虚拟现实一体机等,上述终端并不限定于PC、手机、平板电脑等,服务器可以为媒体文件运营商对应的服务器,上述网络包括但不限于:广域网、城域网或局域网。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例4
根据本申请实施例,还提供了一种图像处理方法,需要说明的是,在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
图8是根据本申请实施例4的一种图像处理方法的流程图。如图8所示,该方法可以包括如下步骤:
步骤S802,通过调用第一接口获取多张图像,其中,第一接口包括第一参数,第一参数的参数值为多张图像,图像的显示内容至少包含待监测对象的目标部位的监测区域;
步骤S804,对图像进行语义分割,得到图像中监测区域的第一区域特征;
步骤S806,基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;
步骤S808,基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;
步骤S8010,通过调用第二接口输出识别结果,其中,第二接口包括第二参数,第二参数的参数值为识别结果。
上述的第一接口可以是向服务器获取多张图像的接口,上述的第二接口可以是向服务器发送识别结果的接口。
在一种可选的实施例中,当需要对待监测对象的目标部位进行监测时,首先可以通过调用第一接口获取多张图像,其中,第一接口包括第一参数,第一参数的参数值为多张图像,图像的显示内容至少包含待监测对象的目标部位的监测区域;其次可以对图像进行语义分割,得到图像中监测区域的第一区域特征;然后可以基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;然后可以基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;最后可以通过调用第二接口输出识别结果,其中,第二接口包括第二参数,第二参数的参数值为识别结果。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例5
根据本申请实施例,还提供了一种肺结节诊断方法,需要说明的是,在附图的流程图 示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
图9是根据本申请实施例5的一种肺结节诊断方法的流程图。如图9所示,该方法可以包括如下步骤:
步骤S902,获取多张医学图像,其中,多张医学图像包含肺结节。
上述的多张医学图像可以是多张2D医学图像,是通过对人体的电子计算机断层摄影(computer tomography,CT)图片的ROI区域(例如可以是肺部区域)进行裁剪并进行转换后得到的多张2D图像,其中,多张医学图像中包含有肺结节。
在一种可选的实施例中,当需要对肺部进行癌症筛查时,首先可以获取包含肺部结节的医学图像,其次可以基于肺部区域对医学图像进行裁剪得到肺部图像,最后可以将肺部图像进行转换得到多张2D医学图像,其中,多张医学图像中包含有肺结节。
步骤S904,对医学图像进行语义分割,得到医学图像中肺结节的第一结节特征;
在一种可选的实施例中,当获取到医学图像后,首先可以通过语义分割模型对肺部图像进行语义分割,即可以得到肺结节的第一结节特征。例如,可以通过语义分割模型对肺部图像进行上下文语义分割,即可以得到肺部图像中肺结节的第一结节特征。又例如,首先可以通过语义分割模型对肺部图像进行上下文语义分割,得到肺部图像中肺结节的第一预设结节特征,其次还可以通过语义分割模型对第一预设结节特征进行上下文解析,进而可以得到肺部图像中肺结节的第一结节特征,但不仅限于此。
例如,在上下文分割阶段,结节上下文信息对良性和恶性诊断具有重要影响。例如,与输血管相关的结节比孤立的结节更可能是恶性的。因此,可以使用U型神经网络(UNet)解析输入图像(即医学图像)的语义掩码m(即上下文语义分割图像),其中,输入图像是对原始CT图像中肺结节ROI区域进行裁剪得到的图像,输入是一个三维的volume,是由多张2D图像(slice)组成。从而允许随后对结节及其周围结构两者进行上下文建模。具体地,m的每个体素属于{0:背景,1:肺,2:结节,3:血管,4:气管}。这种分割过程可以收集对于准确诊断至关重要的综合上下文信息。为了诊断的目的,可以从UNet的瓶颈中提取全局特征作为结节嵌入q,这将在后面的诊断阶段使用。
需要说明的是,在上下文分割阶段,肺结节良恶性鉴别时需要的上下文内容包括肺部正常组织、结节、血管和气管,这些信息不仅反映了结节本身的形状、位置、和大小等信息,同时也反映了结节和周围组织的结构的关系。本申请使用U型卷积神经网络对上述的上下文语义信息进行像素级别的识别,可以终得到每一个结节的上下文语义分割图。
步骤S906,基于预先构建的多个原型和第一结节特征之间的依赖关系,确定肺结节的第二结节特征,其中,不同原型用于表征不同类型的肺结节。
在一种可实施例中,首先可以构建不同原型与第一结节特征之间的依赖关系,其次当得到第一结节特征后,可以基于依赖关系确定第一结节特征对应的原型,然后可以对第一 结节特征以及第一结节特征对应的原型进行处理,得到肺部图像的肺结节的第二结节特征。
例如,在结节内部上下文解析阶段,本申请设计了一种基于注意力机制的上下文解析模块,用来深入剖析结节,整合其上下文信息,提升良恶性鉴别能力。具体地,针对不同结节,其上下文语义分割图像被切分成小的图像块,和原始图像(即输入图像)对应的区域拼接到一起得到多个重叠块,通过一个图像块编码和位置编码,生成一串上下文特征(token)。同时,从卷积神经网络中抽取出高级语义的特征作为结节的全局表征,也称之为结节token。通过设计一种上下文自注意力模块,建模结节token和上下文token之间的长距离的依赖关系,从上下文信息中提取相关的良恶性鉴别依据。自注意力模块输出的结节token作为结节良恶性鉴别的新表征。
需要说明的是,在结节内部上下文解析阶段,可以通过聚合由分割模型产生的上下文信息来增强结节的区别表示。具体地,可以通过重叠块嵌入,将上下文掩码符号化到一组序列中。输入图像也被分割成小块,并被嵌入上下文tokens中以保持原始图像信息。此外,以可学习的方式添加位置编码以保持位置信息。可以将结节嵌入token预先附加到上下文序列,表示为[q;t1,···,tg]∈R(g+1)D。其中,g是上下文tokens的数量,D表示嵌入维度。然后可以同时对这些tokens进行自注意力建模,称为SCA,用于将上下文信息聚合到结节嵌入中。在最后一个SCA块的输出处的结节嵌入token用作更新的结节表示。对结节嵌入及其背景结构之间的依赖性进行明确建模可以导致更区别的表示的进化,从而改善良性和恶性节瘤之间的区别。
又例如,在结节原型召回学习阶段,本申请设计了一种结节诊断知识原型回顾模块,首先结节的原型定义为具有相似特征的结节代表,通过聚类算法将学习过的肺结节在表征空间上进行聚类,得到的类中心作为该类别的原型。其中,原型具有良恶性区分,良性的原型是通过相似特征的良性结节计算得到,而恶性的原型则是来自于恶性结节。为了利用这些原型知识,本申请设计了一种交叉原型注意力模块构建当前结节和其他原型之间的关系,其中query是来自当前结节的表征,key和value分别是来自原型的表征。该交叉注意力模块输出的query token作为最终的良恶性鉴别表征。
需要说明的是,为了保留先前获取的知识,需要更有效的方法,而不是将所有学习的结节存储在存储器中,这将导致存储和计算资源的浪费。为了简化这个过程,可以将相关节结凝聚成原型的形式。对于一组结节(即多张结节图像),可以将其聚集成N个组{C1,···,CN},通过最小化目标函数其中,d是欧几里得距离函数,p表示结节嵌入,并将每个集群的中心作为原型。考虑到良性和恶性结节之间的差异,可以将原型分为良性组和恶性组,由PB∈RN/2×D和PM∈RN/2×D表示。其中,除了解析内部上下文,还可以捕捉结节和外部原型之间的层间依赖性。这使得PARE能够探索除单个结节外的相关鉴定基础。为了实现这一点,本申请设计了交叉原型注意力(CPA)模块,其利用结节嵌入作为查询以及原型作为关键字和值。其允许结节嵌入选择性地参与原型序列的最相关部分。在最后的CPA模块的输出处的查询的状态作为最终结节表示以预测其恶性标记, “良性”(y=0)或“恶性”(y=1)。其中,PARE为本申请提出的一种对肺结节进行诊断的模型,该模型包括:上下文分割、结节内部上下文解析以及结节原型召回学习三部分。
需要说明的是,本申请还可以在线方式更新原型,从而允许原型快速调整结节嵌入的变化。对于数据为(x,y)的结节嵌入q,选出其最近的原型,然后通过以下动量规则更新:
其中,为动量更新后的良性原型,为动量更新后恶性原型,λ为动量因数,一般设置为0.95,但不仅限于此,其中,动量更新可以帮助加速收敛并提高泛化能力。
步骤S908,基于第一结节特征第二结节特征对肺结节进行诊断,得到肺结节的诊断结果,其中,诊断结果用于表征肺结节是良性结节或恶性结节。
在一种可选的实施例中,当得到第一结节特征和第二结节特征后,可以基于第一结节特征和第二结节特征识别对肺结节进行诊断,得到肺结节的诊断结果。例如,可以分别基于第一结节特征和第二结节特征对肺结节进行诊断,得到第一诊断结果和第二诊断结果,然后将第一诊断结果和第二诊断结果进行对比,选取准确度较高的诊断结果作为最终的诊断结果。又例如,可以分别基于第一结节特征和第二结节特征对对肺结节进行诊断,得到第一诊断结果和第二诊断结果,然后取第一诊断结果和第二诊断结果的平均值,得到最终的诊断结果。又例如,还可以首先将第一结节特征和第二结节特征进行特征融合,其次可以基于融合后的结节特征对对肺结节进行诊断,得到肺结节的诊断结果,但不仅限于此。
需要说明的是,本申请设计了一种深度监督训练模式来提高良恶性鉴别能力。深度监督信号分别在卷积神经网络输出的结节全局表征、自上下文注意力模块输出的结节表征、以及交叉原型注意力模块输出的结节表征。通过添加一个多层感知器(Multi-Layer Perception,MLP),从各自的表征空间映射到良性和恶性两大类别空间。在推理场景下,三个MLP得到良恶性类别概率通过平均的方式集成为最终的鉴别概率。
本申请提出一种放射科医生激励方法,模拟放射科医生的诊断过程,由上下文解析和原型回顾模块组成。上下文解析模块首先对结节的上下文结构进行分段,然后聚合上下文信息,以更全面地理解结节。原型回顾模块利用基于原型的学习来将先前学习的情况压缩为用于比较性分析的原型,其在训练期间以动量方式在线更新。基于这两个模块,本申请的方法利用节结的固有特性和从其他节结积累的外部知识两者来实现合理的诊断。为了满足低剂量和非连续筛选的需要,分别从低剂量和非连续CT收集12852和4029个结节的大规模数据集,每个具有病理学或后续确认的标记。在几个数据集上的实验证明,本申请的方法在低剂量和不连续两种场景上实现了先进的筛选性能。
本申请上述实施例中:通过预先构建的多个原型和第一结节之间的依赖关系,能够从结节及其周围器官组织中提取和聚合丰富的上下文信息;通过预先将学习到的结节诊断知 识压缩成原型,并将其用作参考,协助诊断新的结节,能够最终的区域特征更符合目标部位本身的属性,特征提取准确度更高;基于第一结节特征和第二结节特征进行诊断,能够实现面向低剂量和平扫两大筛查场景的肺结节良恶性鉴别,提高了临床应用的通用性。
表1是根据本申请实施例5的一种可选的超参数的消融比较,在表1中,本申请调查了不同配置对PARE性能对验证集的影响,包括变压器层、原型数量、嵌入维数和深度监督。由表1可知,可以通过增加变换器层的数量、增加原型的数量、使token嵌入的通道尺寸加倍、或使用深度分类监督来获得更高的AUC评分。基于0.931的最高AUC评分,在以下实验中经验地设定L=4、N=40、D=256和DS=真。其中,超参数包括:变压器层(L)、原型数量(N)、嵌入维度(D)以及深度监督(DS)。
表1超参数的消融比较
表2是根据本申请实施例5的一种可选的不同模块的有效性,在表2中,本申请调查了不同方法/模块在验证集合上的消融研究并且观察以下结果:(1)纯分割方法表现得比纯分类方法更好,主要是因为它使得能够在像素级上进行更大的监督,(2)联合分割和分类优于任何单一方法,表明两个任务的互补效果,(3)上下文解析和原型比较两者有助于在强基线上改善性能,从而证明两个模块的有效性,并且(4)与仅仅分割结节相比,分割更多上下文结构(例如血管、肺和气管)提供了轻微的改进。其中,MT表示多任务学习。上下文:帧内上下文解析。原型:原型间的回顾。*表示在分割任务中仅使用结节掩模。
表2不同模块的有效性
在两个筛选场景上与其他方法的比较:表3是根据本申请实施例5的一种可选的NNLST和内部测试集上不同方法的比较,包括基于纯分类的方法、基于纯分割的方法和基于多任务的方法。基于结节尺寸分布,在两个测试组中进行分层评估。这些结果指示该基于分段的方法胜过纯粹的分类方法,主要是由于其对上下文结构进行分段的优越能力。此 外,基于多任务的CA-网胜过任何单任务方法。在NLST和内部测试集上,本申请的PARE方法超过了大多数其他方法。此外,通过利用多个深度监督头的集合,对两个数据集将总体AUC进一步改善至0.931。其中,。表示纯分类;表示纯分割;表示多任务学习;*表示深度监督头集合。需要注意的是:本申请在CA-Net中加入了分割任务。
表3 NNLST和内部测试集上不同方法的比较
LUNGx的外部评价:本申请使用LUNGx作为外部检验来评价的PARE的泛化。值得注意的是,这些比较的方法从未在LUNGx上进行过训练。表4是根据本申请实施例5的一种可选的与其他方法在LUNGx上的对比,从表4可以看出,本申请的PARE模型的AUC最高为0.801,比方法DAR提高了2%。本申请还进行了一项读者研究,将PARE与两位分别具有8年和13年肺结节诊断经验的具有丰富经验的放射科医生进行比较。图10是根据本申请实施例5的一种可选的读者研究与人工智能对比的示意图,图3的结果显示,本申请的方法达到了与放射科医生相当的性能。
表4与其他方法在LUNGx上的对比
需要说明的是,本申请所涉及的用户信息(包括但不限于用户设备信息、用户个人信息等)和数据(包括但不限于用于分析的数据、存储的数据、展示的数据等),均为经用户授权或者经过各方充分授权的信息和数据,并且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准,并提供有相应的操作入口,供用户选择授权或者拒绝。
LDCT和NCCT的泛化:本申请的模型实在LDCT和NCCT数据集的混合上训练的,可以在低剂量和常规剂量应用中表现良好。表5示出了在三种训练配置下获得的模型的泛化性能比较,从表5可以看出,单独在LDCT或NCCT数据集上训练的模型不能很好地推广到其他模态,至少有6%的AUC下降。然而,本申请的混合训练方法在LDCT和NCCT 上都表现最好,几乎没有性能下降。
表5不同训练配置对LDCT和NCCT性能的影响对比
此外通过比较由医生和其他由TotalSegmentator生成的语义掩码标注的从LIDC-IDRI数据中得到的分割结果,和由本申请实施例提供的在分割任务中使用的结节掩模标注的分割结果,PARE在LIDC-IDRI的2630个结节上达到77.9%的Dice得分。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例6
根据本申请实施例,还提供了一种癌症计算机辅助诊断方法,该方法包括:
获取多张医学图像,其中,所述医学图像包含待监测对象的目标监测部位;
对所述医学图像进行语义分割,得到所述医学图像中所述监测部位的第一部位特征;
基于预先构建的多个原型和所述第一部位特征之间的依赖关系,确定所述监测部位的第二部位特征,其中,不同原型用于表征不同类型的监测部位;
基于所述第一部位特征和所述第二部位特征对所述监测部位进行诊断,得到所述监测部位的诊断结果,其中,所述诊断结果用于表征所述监测部位是恶性病症或良性病症。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例7
根据本申请实施例,还提供了一种癌症的计算机辅助诊断系统,包括存储器、处理器以及存储在所述存储器上并在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序可用于执行一种癌症计算机辅助方法,所述方法包括:
获取多张医学图像,其中,所述医学图像包含待监测对象的目标监测部位;
对所述医学图像进行语义分割,得到所述医学图像中所述监测部位的第一部位特征;
基于预先构建的多个原型和所述第一部位特征之间的依赖关系,确定所述监测部位的第二部位特征,其中,不同原型用于表征不同类型的监测部位;
基于所述第一部位特征和所述第二部位特征对所述监测部位进行诊断,得到所述监测部位的诊断结果,其中,所述诊断结果用于表征所述监测部位是恶性病症或良性病症。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例8
根据本申请实施例的另一方面,还提供了一种肺癌计算机辅助诊断方法,包括:
获取多张医学图像,其中,所述医学图像包含肺部病变;
对所述医学图像进行语义分割,得到所述医学图像中肺部病变的第一病变特征;
基于预先构建的多个原型和所述第一结节特征之间的依赖关系,确定所述肺部病变的第二病变特征,其中,不同原型用于表征不同类型的肺部病变;
基于所述第一病变特征和所述第二病变特征对所述肺结节进行诊断,得到所述肺部病变的诊断结果,其中,所述诊断结果用于表征所述肺部病变是良性病症或恶性病症。
在本实施例中,肺部病变可以包括肺结节。若诊断结果为肺部病变为恶性病症,则该肺部病变可能是由肺癌造成,以便医生或诊断设备根据诊断结果确定治疗方案。
实施例9
根据本申请实施例,还提供了一种用于实施上述图像处理方法的图像处理装置。图11是根据本申请实施例6的一种图像处理装置的结构示意图,如图11所示,该装置包括:获取模块1102、分割模块1104、第一确定模块1106和第二确定模块1108。
其中,获取模块用于获取多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;分割模块用于对图像进行语义分割,得到图像中监测区域的第一区域特征;第一确定模块用于基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;第二确定模块用于基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果。
此处需要说明的是,上述获取模块、分割模块、第一确定模块和第二确定模块对应于实施例1中的步骤S302至步骤S308,四个模块与对应的步骤所实现的实例和应用场景相同,但不限于上述实施例1所公开的内容。需要说明的是,上述模块或单元可以是存储在存储器中并由一个或多个处理器处理的硬件组件或软件组件,上述模块也可以作为装置的一部分可以运行在实施例1提供的AR/VR设备中。
本申请上述实施例中,分割模块包括:分割单元、融合单元和第一处理单元。
其中,分割单元用于对图像进行语义分割,得到图像的语义分割结果和全局特征;融合单元,用于对语义分割结果和图像进行特征融合,得到融合特征;第一处理单元用于对全局特征和融合特征进行注意力处理,得到第一区域特征。
本申请上述实施例中,分割单元包括:第一提取子单元、第二提取子单元和解码子单元。
其中,第一提取子单元用于利用U型神经网络模型的编码器模块对图像进行特征提取,得到图像的第一图像特征;第二提取子单元用于从U型神经网络模型的瓶颈层中,提取全局特征;解码子单元用于利用U型神经网络模型的编码器模块,对第一图像特征进行解码,得到语义分割结果。
本申请上述实施例中,融合单元包括:切分子单元、第三提取子单元和融合子单元。
其中,切分子单元用于分别对语义分割结果和图像进行切分,得到多个子分割结果和多个子图像;第三提取子单元用于分别对多个子分割结果和多个子图像进行特征提取,得到多个子分割结果的子分割特征,以及多个子图像的子图像特征;融合子单元用于将子分割特征和子图像特征进行融合,得到融合特征。
本申请上述实施例中,第一处理单元包括:拼接子单元和处理子单元。
其中,拼接子单元用于将全局特征和融合特征进行拼接,得到第一拼接特征;处理子单元用于利用自注意力模型对第一拼接特征进行自注意力处理,得到第一区域特征。
本申请上述实施例中,第一确定模块包括:第二处理单元。
其中,第二处理单元用于利用交叉注意力模型对第一区域特征和多个原型进行注意力处理,得到第二区域特征。
本申请上述实施例中,第一确定模块还包括:获取单元、聚类单元和构建单元。
其中,获取单元用于获取不同监测区域的全局特征;聚类单元用于对不同监测区域的全局特征进行聚类,得到多个特征集合;构建单元用于基于多个特征集合的中心特征,构建多个原型。
本申请上述实施例中,在基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征之后,该装置还包括:第三确定模块、第一更新模块和第二更新模块。
其中,第三确定模块用于从预先构建的多个原型中,确定与第二区域特征匹配成功的目标原型;第一更新模块用于对目标原型进行动量更新,得到更新后的区域特征;第二更新模块用于基于更新后的区域特征对多个原型进行更新。
本申请上述实施例中,第二确定模块包括:第三处理单元、第四处理单元、第五处理单元和汇总单元。
其中,第三处理单元用于基于全局特征识别监测区域的特征信息,得到第一子识别结果;第四处理单元用于基于第一区域特征识别监测区域的特征信息,得到第二子识别结果;第五处理单元用于基于第二区域特征识别监测区域的特征信息,得到第三子识别结果;汇总单元用于对第一子识别结果、第二子识别结果和第三子识别结果进行汇总,得到识别结果。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例10
根据本申请实施例,还提供了一种用于实施上述图像处理方法的图像处理装置。
图12是根据本申请实施例7的一种图像处理装置的结构示意图,如图12所示,该装置包括:第一显示模块1202和第二显示模块1204。
其中,第一显示模块用于响应作用于操作界面上的输入指令,在操作界面上显示多张 图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;第二显示模块用于响应作用于操作界面上的图像处理指令,在操作界面上显示监测区域的识别结果,其中,识别结果是基于第一区域特征和第二区域特征识别监测区域的特征信息得到的,第二区域特征是基于预先构建的多个原型和第一区域特征之间的依赖关系确定的,第一区域特征是对医学图像进行语义分割得到的。
此处需要说明的是,上述第一显示模块和第二显示模块对应于实施例2中的步骤S502至步骤S504,两个模块与对应的步骤所实现的实例和应用场景相同,但不限于上述实施例2所公开的内容。需要说明的是,上述模块或单元可以是存储在存储器中并由一个或多个处理器处理的硬件组件或软件组件,上述模块也可以作为装置的一部分可以运行在实施例1提供的AR/VR设备中。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例11
根据本申请实施例,还提供了一种用于实施上述图像处理方法的图像处理装置。图13是根据本申请实施例8的一种图像处理装置的结构示意图,如图13所示,该装置包括:显示模块1302、分割模块1304、第一确定模块1306、第二确定模块1308和驱动模块13010。
其中,显示模块用于在虚拟现实VR设备或增强现实AR设备的呈现画面上展示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;分割模块用于对图像进行语义分割,得到图像中监测区域的第一区域特征;第一确定模块用于基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;第二确定模块用于基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;驱动模块用于驱动VR设备或AR设备渲染展示识别结果。
此处需要说明的是,上述显示模块、分割模块、第一确定模块、第二确定模块和驱动模块对应于实施例3中的步骤S702至步骤S7010,五个模块与对应的步骤所实现的实例和应用场景相同,但不限于上述实施例3所公开的内容。需要说明的是,上述模块或单元可以是存储在存储器中并由一个或多个处理器处理的硬件组件或软件组件,上述模块也可以作为装置的一部分可以运行在实施例1提供的AR/VR设备中。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例12
根据本申请实施例,还提供了一种用于实施上述图像处理方法的图像处理装置。图14是根据本申请实施例9的一种图像处理装置的结构示意图,如图14所示,该装置包括:获取模块1402、分割模块1404、第一确定模块1406、第二确定模块1408和输出模块14010。
其中,获取模块,用于通过调用第一接口获取多张图像,其中,第一接口包括第一参 数,第一参数的参数值为多张图像,图像的显示内容至少包含待监测对象的目标部位的监测区域;分割模块,用于对图像进行语义分割,得到图像中监测区域的第一区域特征;第一确定模块,用于基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;第二确定模块,用于基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;输出模块,用于通过调用第二接口输出识别结果,其中,第二接口包括第二参数,第二参数的参数值为识别结果。
此处需要说明的是,上述获取模块、分割模块、第一确定模块、第二确定模块和输出模块对应于实施例4中的步骤S802至步骤S8010,五个模块与对应的步骤所实现的实例和应用场景相同,但不限于上述实施例4所公开的内容。需要说明的是,上述模块或单元可以是存储在存储器中并由一个或多个处理器处理的硬件组件或软件组件,上述模块也可以作为装置的一部分可以运行在实施例1提供的AR/VR设备中。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例13
根据本申请实施例,还提供了一种用于实施上述肺结节诊断方法的肺结节诊断装置。图15是根据本申请实施例10的一种肺结节诊断装置的结构示意图,如图15所示,该装置包括:获取模块1502、分割模块1504、确定模块1506、诊断模块1508。
其中,获取模块用于获取多张医学图像,其中,多张医学图像包含肺结节;分割模块用于对医学图像进行语义分割,得到医学图像中肺结节的第一结节特征;确定模块用于基于预先构建的多个原型和第一结节特征之间的依赖关系,确定肺结节的第二结节特征,其中,不同原型用于表征不同类型的肺结节;诊断模块,用于基于第一结节特征第二结节特征对肺结节进行诊断,得到肺结节的诊断结果,其中,诊断结果用于表征肺结节是良性结节或恶性结节。
此处需要说明的是,上述获取模块、分割模块、确定模块、诊断模块对应于实施例5中的步骤S902至步骤S908,五个模块与对应的步骤所实现的实例和应用场景相同,但不限于上述实施例5所公开的内容。需要说明的是,上述模块或单元可以是存储在存储器中并由一个或多个处理器处理的硬件组件或软件组件,上述模块也可以作为装置的一部分可以运行在实施例1提供的AR/VR设备中。
需要说明的是,本申请上述实施例中涉及到的优选实施方案与实施例1提供的方案以及应用场景、实施过程相同,但不仅限于实施例1所提供的方案。
实施例14
本申请的实施例可以提供一种AR/VR设备,该AR/VR设备可以是AR/VR设备群中的任意一个AR/VR设备。可选地,在本实施例中,上述AR/VR设备也可以替换为移动终端等终端设备。
可选地,在本实施例中,上述AR/VR设备可以位于计算机网络的多个网络设备中的至少一个网络设备。
在本实施例中,上述AR/VR设备可以执行图像处理方法中以下步骤的程序代码:在虚拟现实VR设备或增强现实AR设备的呈现画面上展示多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果;驱动VR设备或AR设备渲染展示识别结果。
可选地,图16是根据本申请实施例的一种计算机终端的结构框图。如图16所示,该计算机终端A可以包括:一个或多个(图中仅示出一个)处理器1602、存储器1604、存储控制器、以及外设接口,其中,外设接口与射频模块、音频模块和显示器连接。
其中,存储器可用于存储软件程序以及模块,如本申请实施例中的图像处理方法和装置对应的程序指令/模块,处理器通过运行存储在存储器内的软件程序以及模块,从而执行各种功能应用以及数据处理,即实现上述的图像处理方法。存储器可包括高速随机存储器,还可以包括非易失性存储器,如一个或者多个磁性存储装置、闪存、或者其他非易失性固态存储器。在一些实例中,存储器可进一步包括相对于处理器远程设置的存储器,这些远程存储器可以通过网络连接至终端A。上述网络的实例包括但不限于互联网、企业内部网、局域网、移动通信网及其组合。
处理器可以通过传输装置调用存储器存储的信息及应用程序,以执行下述步骤:获取多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果。
可选的,上述处理器还可以执行如下步骤的程序代码:对图像进行语义分割,得到图像的语义分割结果和全局特征;对语义分割结果和图像进行特征融合,得到融合特征;对全局特征和融合特征进行注意力处理,得到第一区域特征。
可选的,上述处理器还可以执行如下步骤的程序代码:利用U型神经网络模型的编码器模块对图像进行特征提取,得到图像的第一图像特征;从U型神经网络模型的瓶颈层中,提取全局特征;利用U型神经网络模型的编码器模块,对第一图像特征进行解码,得到语义分割结果。
可选的,上述处理器还可以执行如下步骤的程序代码:分别对语义分割结果和图像进行切分,得到多个子分割结果和多个子图像;分别对多个子分割结果和多个子图像进行特征提取,得到多个子分割结果的子分割特征,以及多个子图像的子图像特征;将子分割特 征和子图像特征进行融合,得到融合特征。
可选的,上述处理器还可以执行如下步骤的程序代码:将全局特征和融合特征进行拼接,得到第一拼接特征;利用自注意力模型对第一拼接特征进行自注意力处理,得到第一区域特征。
可选的,上述处理器还可以执行如下步骤的程序代码:利用交叉注意力模型对第一区域特征和多个原型进行注意力处理,得到第二区域特征。
可选的,上述处理器还可以执行如下步骤的程序代码:获取不同监测区域的全局特征;对不同监测区域的全局特征进行聚类,得到多个特征集合;基于多个特征集合的中心特征,构建多个原型。
可选的,上述处理器还可以执行如下步骤的程序代码:从预先构建的多个原型中,确定与第二区域特征匹配成功的目标原型;对目标原型进行动量更新,得到更新后的区域特征;基于更新后的区域特征对多个原型进行更新。
可选的,上述处理器还可以执行如下步骤的程序代码:基于全局特征识别监测区域的特征信息,得到第一子识别结果;基于第一区域特征识别监测区域的特征信息,得到第二子识别结果;基于第二区域特征识别监测区域的特征信息,得到第三子识别结果;对第一子识别结果、第二子识别结果和第三子识别结果进行汇总,得到识别结果。
采用本申请实施例,提供了一种获取多张图像;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果的方式。容易注意到的是,本申请不仅对图像进行了语义分割,得到区域特征,还结合了原型与区域特征之间的依赖关系,确保最终的区域特征更符合目标部位本身的属性,特征提取准确度更高,进而达到了更准确的对待监测对象的目标部位进行识别的目的,从而实现了提高待监测对象的目标部位的识别准确率的技术效果,进而解决了相关技术中对待监测图像进行识别的识别准确率低的技术问题。
本领域普通技术人员可以理解,图所示的结构仅为示意,计算机终端也可以是智能手机(如Android手机、iOS手机等)、平板电脑、掌上电脑以及移动互联网设备(Mobile Internet Devices,MID)、PAD等终端设备。图16并不对上述电子装置的结构造成限定。例如,计算机终端A还可包括比图16中所示更多或者更少的组件(如网络接口、显示装置等),或者具有与图16所示不同的配置。
本领域普通技术人员可以理解上述实施例的各种方法中的全部或部分步骤是可以通过程序来指令终端设备相关的硬件来完成,该程序可以存储于一计算机可读存储介质中,存储介质可以包括:闪存盘、只读存储器(Read-Only Memory,ROM)、随机存取器(Random Access Memory,RAM)、磁盘或光盘等。
实施例14
本申请的实施例还提供了一种计算机可读存储介质。可选地,在本实施例中,上述计 算机可读存储介质可以用于保存上述实施例1所提供的图像处理方法所执行的程序代码。
可选地,在本实施例中,上述计算机可读存储介质可以位于AR/VR设备网络中AR/VR设备终端群中的任意一个计算机终端中,或者位于移动终端群中的任意一个移动终端中。
可选地,在本实施例中,计算机可读存储介质被设置为存储用于执行以下步骤的程序代码:获取多张图像,其中,图像的显示内容至少包含待监测对象的目标部位的监测区域;对图像进行语义分割,得到图像中监测区域的第一区域特征;基于预先构建的多个原型和第一区域特征之间的依赖关系,确定监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;基于第一区域特征和第二区域特征识别监测区域的特征信息,确定监测区域的识别结果。
可选地,在本实施例中,计算机可读存储介质还被设置为存储用于执行以下步骤的程序代码:对图像进行语义分割,得到图像的语义分割结果和全局特征;对语义分割结果和图像进行特征融合,得到融合特征;对全局特征和融合特征进行注意力处理,得到第一区域特征。
可选地,在本实施例中,计算机可读存储介质还被设置为存储用于执行以下步骤的程序代码:利用U型神经网络模型的编码器模块对图像进行特征提取,得到图像的第一图像特征;从U型神经网络模型的瓶颈层中,提取全局特征;利用U型神经网络模型的编码器模块,对第一图像特征进行解码,得到语义分割结果。
可选地,在本实施例中,计算机可读存储介质还被设置为存储用于执行以下步骤的程序代码:分别对语义分割结果和图像进行切分,得到多个子分割结果和多个子图像;分别对多个子分割结果和多个子图像进行特征提取,得到多个子分割结果的子分割特征,以及多个子图像的子图像特征;将子分割特征和子图像特征进行融合,得到融合特征。
可选地,在本实施例中,计算机可读存储介质还被设置为存储用于执行以下步骤的程序代码:将全局特征和融合特征进行拼接,得到第一拼接特征;利用自注意力模型对第一拼接特征进行自注意力处理,得到第一区域特征。
可选地,在本实施例中,计算机可读存储介质还被设置为存储用于执行以下步骤的程序代码:利用交叉注意力模型对第一区域特征和多个原型进行注意力处理,得到第二区域特征。
可选地,在本实施例中,计算机可读存储介质还被设置为存储用于执行以下步骤的程序代码:获取不同监测区域的全局特征;对不同监测区域的全局特征进行聚类,得到多个特征集合;基于多个特征集合的中心特征,构建多个原型。
可选地,在本实施例中,计算机可读存储介质还被设置为存储用于执行以下步骤的程序代码:从预先构建的多个原型中,确定与第二区域特征匹配成功的目标原型;对目标原型进行动量更新,得到更新后的区域特征;基于更新后的区域特征对多个原型进行更新。
可选地,在本实施例中,计算机可读存储介质还被设置为存储用于执行以下步骤的程序代码:基于全局特征识别监测区域的特征信息,得到第一子识别结果;基于第一区域特 征识别监测区域的特征信息,得到第二子识别结果;基于第二区域特征识别监测区域的特征信息,得到第三子识别结果;对第一子识别结果、第二子识别结果和第三子识别结果进行汇总,得到识别结果。
实施例15
本申请的实施例还提供了一种计算机程序产品,包括计算机程序,所述计算机程序在计算机中执行时,令计算机执行本申请的实施例提供的方法。
实施例16
本申请的实施例还提供了一种计算机程序,其中,当所述计算机程序在计算机中执行时,令计算机执行本申请的实施例提供的方法。上述本申请实施例序号仅仅为了描述,不代表实施例的优劣。
在本申请的上述实施例中,对各个实施例的描述都各有侧重,某个实施例中没有详述的部分,可以参见其他实施例的相关描述。
在本申请所提供的几个实施例中,应该理解到,所揭露的技术内容,可通过其它的方式实现。其中,以上所描述的装置实施例仅仅是示意性的,例如所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,单元或模块的间接耦合或通信连接,可以是电性或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可为个人计算机、服务器或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、移动硬盘、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述仅是本申请的优选实施方式,应当指出,对于本技术领域的普通技术人员来说,在不脱离本申请原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也应视为本申请的保护范围。

Claims (20)

  1. 一种图像处理方法,其特征在于,包括:
    获取多张图像,其中,所述图像的显示内容至少包含待监测对象的目标部位的监测区域;
    对所述图像进行语义分割,得到所述图像中所述监测区域的第一区域特征;
    基于预先构建的多个原型和所述第一区域特征之间的依赖关系,确定所述监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;
    基于所述第一区域特征和所述第二区域特征识别所述监测区域的特征信息,确定所述监测区域的识别结果。
  2. 根据权利要求1所述的方法,其特征在于,对所述图像进行语义分割,得到所述图像中所述监测区域的第一区域特征,包括:
    对所述图像进行语义分割,得到所述图像的语义分割结果和全局特征;
    对所述语义分割结果和所述图像进行特征融合,得到融合特征;
    对所述全局特征和所述融合特征进行注意力处理,得到所述第一区域特征。
  3. 根据权利要求2所述的方法,其特征在于,对所述图像进行语义分割,得到所述图像的语义分割结果和全局特征,包括:
    利用U型神经网络模型的编码器模块对所述图像进行特征提取,得到所述图像的第一图像特征;
    从所述U型神经网络模型的瓶颈层中,提取所述全局特征;
    利用所述U型神经网络模型的编码器模块,对所述第一图像特征进行解码,得到所述语义分割结果。
  4. 根据权利要求2或3所述的方法,其特征在于,对所述语义分割结果和所述图像进行特征融合,得到融合特征,包括:
    分别对所述语义分割结果和所述图像进行切分,得到多个子分割结果和多个子图像;
    分别对所述多个子分割结果和所述多个子图像进行特征提取,得到所述多个子分割结果的子分割特征,以及所述多个子图像的子图像特征;
    将所述子分割特征和所述子图像特征进行融合,得到所述融合特征。
  5. 根据权利要求2至4任意一项所述的方法,其特征在于,对所述全局特征和所述融合特征进行注意力处理,得到所述第一区域特征,包括:
    将所述全局特征和所述融合特征进行拼接,得到第一拼接特征;
    利用自注意力模型对所述第一拼接特征进行自注意力处理,得到所述第一区域特征。
  6. 根据权利要求1至5任意一项所述的方法,其特征在于,基于预先构建的多个原型和所述第一区域特征之间的依赖关系,确定所述监测区域的第二区域特征,包括:
    利用交叉注意力模型对所述第一区域特征和所述多个原型进行注意力处理,得到所述第二区域特征。
  7. 根据权利要求6所述的方法,其特征在于,所述方法还包括:
    获取所述不同监测区域的全局特征;
    对所述不同监测区域的全局特征进行聚类,得到多个特征集合;
    基于所述多个特征集合的中心特征,构建所述多个原型。
  8. 根据权利要求1至7任意一项所述的方法,其特征在于,在基于预先构建的多个原型和所述第一区域特征之间的依赖关系,确定所述监测区域的第二区域特征之后,所述方法还包括:
    从预先构建的多个原型中,确定与第二区域特征匹配成功的目标原型;
    对所述目标原型进行动量更新,得到更新后的区域特征;
    基于所述更新后的区域特征对所述多个原型进行更新。
  9. 根据权利要求1至8任意一项所述的方法,其特征在于,基于所述第一区域特征和所述第二区域特征识别所述监测区域的特征信息,确定所述监测区域的识别结果,包括:
    基于全局特征识别所述监测区域的特征信息,得到第一子识别结果;
    基于所述第一区域特征识别所述监测区域的特征信息,得到第二子识别结果;
    基于所述第二区域特征识别所述监测区域的特征信息,得到第三子识别结果;
    对所述第一子识别结果、所述第二子识别结果和所述第三子识别结果进行汇总,得到所述识别结果。
  10. 一种图像处理方法,其特征在于,包括:
    响应作用于操作界面上的输入指令,在所述操作界面上显示多张图像,其中,所述图像的显示内容至少包含待监测对象的目标部位的监测区域;
    响应作用于所述操作界面上的图像处理指令,在所述操作界面上显示所述监测区域的识别结果,其中,所述识别结果是基于第一区域特征和第二区域特征识别所述监测区域的特征信息得到的,所述第二区域特征是基于预先构建的多个原型和所述第一区域特征之间的依赖关系确定的,所述第一区域特征是对医学图像进行语义分割得到的。
  11. 一种图像处理方法,其特征在于,包括:
    通过调用第一接口获取多张图像,其中,所述第一接口包括第一参数,所述第一参数的参数值为所述多张图像,所述图像的显示内容至少包含待监测对象的目标部位的监测区域;
    对所述图像进行语义分割,得到所述图像中所述监测区域的第一区域特征;
    基于预先构建的多个原型和所述第一区域特征之间的依赖关系,确定所述监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;
    基于所述第一区域特征和所述第二区域特征识别所述监测区域的特征信息,确定所 述监测区域的识别结果;
    通过调用第二接口输出所述识别结果,其中,所述第二接口包括第二参数,所述第二参数的参数值为所述识别结果。
  12. 一种癌症计算机辅助诊断方法,包括:
    获取多张医学图像,其中,所述医学图像包含待监测对象的目标监测部位;
    对所述医学图像进行语义分割,得到所述医学图像中所述监测部位的第一部位特征;
    基于预先构建的多个原型和所述第一部位特征之间的依赖关系,确定所述监测部位的第二部位特征,其中,不同原型用于表征不同类型的监测部位;
    基于所述第一部位特征和所述第二部位特征对所述监测部位进行诊断,得到所述监测部位的诊断结果,其中,所述诊断结果用于表征所述监测部位是恶性病症或良性病症。
  13. 一种肺癌计算机辅助诊断方法,包括:
    获取多张医学图像,其中,所述医学图像包含肺部病变;
    对所述医学图像进行语义分割,得到所述医学图像中肺部病变的第一病变特征;
    基于预先构建的多个原型和所述第一结节特征之间的依赖关系,确定所述肺部病变的第二病变特征,其中,不同原型用于表征不同类型的肺部病变;
    基于所述第一病变特征和所述第二病变特征对所述肺结节进行诊断,得到所述肺部病变的诊断结果,其中,所述诊断结果用于表征所述肺部病变是良性病症或恶性病症。
  14. 一种图像处理装置,包括:
    获取模块,用于获取多张图像,其中,所述图像的显示内容至少包含待监测对象的目标部位的监测区域;
    分割模块,用于对所述图像进行语义分割,得到所述图像中所述监测区域的第一区域特征;
    第一确定模块,用于基于预先构建的多个原型和所述第一区域特征之间的依赖关系,确定所述监测区域的第二区域特征,其中,不同原型用于表征不同类型的监测区域;
    第二确定模块,基于所述第一区域特征和所述第二区域特征识别所述监测区域的特征信息,确定所述监测区域的识别结果。
  15. 一种图像处理装置,包括:
    第一显示模块,用于响应作用于操作界面上的输入指令,在所述操作界面上显示多张图像,其中,所述图像的显示内容至少包含待监测对象的目标部位的监测区域;
    第二显示模块,用于响应作用于所述操作界面上的图像处理指令,在所述操作界面上显示所述监测区域的识别结果,其中,所述识别结果是基于第一区域特征和第二区域特征识别所述监测区域的特征信息得到的,所述第二区域特征是基于预先构建的多个原 型和所述第一区域特征之间的依赖关系确定的,所述第一区域特征是对医学图像进行语义分割得到的。
  16. 一种癌症的计算机辅助诊断系统,包括存储器、处理器以及存储在所述存储器上并在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序可用于执行一种癌症计算机辅助诊断方法,所述方法包括:
    获取多张医学图像,其中,所述医学图像包含待监测对象的目标监测部位;
    对所述医学图像进行语义分割,得到所述医学图像中所述监测部位的第一部位特征;
    基于预先构建的多个原型和所述第一部位特征之间的依赖关系,确定所述监测部位的第二部位特征,其中,不同原型用于表征不同类型的监测部位;
    基于所述第一部位特征和所述第二部位特征对所述监测部位进行诊断,得到所述监测部位的诊断结果,其中,所述诊断结果用于表征所述监测部位是恶性病症或良性病症。
  17. 一种电子设备,其特征在于,包括:
    存储器,存储有可执行程序;
    处理器,用于运行所述程序,其中,所述程序运行时执行权利要求1至13中任意一项所述的方法。
  18. 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质包括存储的可执行程序,其中,在所述可执行程序运行时控制所述计算机可读存储介质所在设备执行权利要求1至13中任意一项所述的方法。
  19. 一种计算机程序产品,包括计算机可执行指令,该计算机可执行指令被处理器执行时实现权利要求1至13任意一项所述方法的步骤。
  20. 一种计算机程序,该计算机程序被处理器执行时实现权利要求1至13任意一项所述方法的步骤。
PCT/CN2024/103723 2023-07-04 2024-07-04 图像处理方法、电子设备及存储介质 Ceased WO2025007940A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US19/419,601 US20260105717A1 (en) 2023-07-04 2025-12-15 Image processing method, electronic device, and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310814294.0 2023-07-04
CN202310814294.0A CN117095320B (zh) 2023-07-04 2023-07-04 图像处理方法、电子设备及存储介质

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US19/419,601 Continuation US20260105717A1 (en) 2023-07-04 2025-12-15 Image processing method, electronic device, and storage medium

Publications (1)

Publication Number Publication Date
WO2025007940A1 true WO2025007940A1 (zh) 2025-01-09

Family

ID=88772503

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/103723 Ceased WO2025007940A1 (zh) 2023-07-04 2024-07-04 图像处理方法、电子设备及存储介质

Country Status (4)

Country Link
US (1) US20260105717A1 (zh)
CN (1) CN117095320B (zh)
TW (1) TW202503692A (zh)
WO (1) WO2025007940A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119888241A (zh) * 2025-03-31 2025-04-25 厦门理工学院 多模态协同增强与动态对齐的血管图像分割方法及装置
CN120125662A (zh) * 2025-02-21 2025-06-10 中国铁塔股份有限公司四川省分公司 基于视频图像的空间定位方法和装置、设备及介质

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117095320B (zh) * 2023-07-04 2025-12-30 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法、电子设备及存储介质
TWI908591B (zh) * 2025-01-22 2025-12-11 國立中央大學 用於物件追蹤的電子裝置、方法及電腦程式產品

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10430946B1 (en) * 2019-03-14 2019-10-01 Inception Institute of Artificial Intelligence, Ltd. Medical image segmentation and severity grading using neural network architectures with semi-supervised learning techniques
CN114581459A (zh) * 2022-02-08 2022-06-03 浙江大学 一种基于改进性3D U-Net模型的学前儿童肺部影像感兴趣区域分割方法
CN114612902A (zh) * 2022-03-17 2022-06-10 腾讯科技(深圳)有限公司 图像语义分割方法、装置、设备、存储介质及程序产品
WO2022248727A1 (en) * 2021-05-28 2022-12-01 Deepmind Technologies Limited Generating neural network outputs by cross attention of query embeddings over a set of latent embeddings
CN116188392A (zh) * 2022-12-30 2023-05-30 阿里巴巴(中国)有限公司 图像处理方法、计算机可读存储介质以及计算机终端
CN117095320A (zh) * 2023-07-04 2023-11-21 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法、电子设备及存储介质

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP3567548B1 (en) * 2018-05-09 2020-06-24 Siemens Healthcare GmbH Medical image segmentation
US11263488B2 (en) * 2020-04-13 2022-03-01 International Business Machines Corporation System and method for augmenting few-shot object classification with semantic information from multiple sources
CN114119546B (zh) * 2021-11-25 2025-05-30 推想医疗科技股份有限公司 检测mri影像的方法及装置
CN114898406A (zh) * 2022-06-13 2022-08-12 中国计量大学 一种基于对比聚类的无监督行人重识别方法
CN116206159B (zh) * 2023-03-28 2025-12-09 武汉大学 一种图像分类方法、装置、设备及可读存储介质

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10430946B1 (en) * 2019-03-14 2019-10-01 Inception Institute of Artificial Intelligence, Ltd. Medical image segmentation and severity grading using neural network architectures with semi-supervised learning techniques
WO2022248727A1 (en) * 2021-05-28 2022-12-01 Deepmind Technologies Limited Generating neural network outputs by cross attention of query embeddings over a set of latent embeddings
CN114581459A (zh) * 2022-02-08 2022-06-03 浙江大学 一种基于改进性3D U-Net模型的学前儿童肺部影像感兴趣区域分割方法
CN114612902A (zh) * 2022-03-17 2022-06-10 腾讯科技(深圳)有限公司 图像语义分割方法、装置、设备、存储介质及程序产品
CN116188392A (zh) * 2022-12-30 2023-05-30 阿里巴巴(中国)有限公司 图像处理方法、计算机可读存储介质以及计算机终端
CN117095320A (zh) * 2023-07-04 2023-11-21 阿里巴巴达摩院(杭州)科技有限公司 图像处理方法、电子设备及存储介质

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120125662A (zh) * 2025-02-21 2025-06-10 中国铁塔股份有限公司四川省分公司 基于视频图像的空间定位方法和装置、设备及介质
CN119888241A (zh) * 2025-03-31 2025-04-25 厦门理工学院 多模态协同增强与动态对齐的血管图像分割方法及装置

Also Published As

Publication number Publication date
TW202503692A (zh) 2025-01-16
US20260105717A1 (en) 2026-04-16
CN117095320A (zh) 2023-11-21
CN117095320B (zh) 2025-12-30

Similar Documents

Publication Publication Date Title
WO2025007940A1 (zh) 图像处理方法、电子设备及存储介质
CN116188392B (zh) 图像处理方法、计算机可读存储介质以及计算机终端
Zuo et al. R2AU‐Net: attention recurrent residual convolutional neural network for multimodal medical image segmentation
EP3968222B1 (en) Classification task model training method, apparatus and device and storage medium
WO2024146649A1 (zh) 图像处理方法、存储介质及计算机终端
CN1378677A (zh) 生成电子多媒体报告的一种方法和计算机实现程序
TW201224826A (en) Systems and methods for automated extraction of measurement information in medical videos
JP7449366B2 (ja) 機械学習システムおよび方法、統合サーバ、情報処理装置、プログラムならびに推論モデルの作成方法
CN1615489A (zh) 图像报告方法和系统
CN113724184A (zh) 脑出血预后预测方法、装置、电子设备及存储介质
CN116433605A (zh) 基于云智能的医学影像分析移动增强现实系统和方法
CN115100723B (zh) 面色分类方法、装置、计算机可读程序介质及电子设备
CN117058164A (zh) 图像分割方法、电子设备和存储介质
WO2025001689A1 (zh) 图像配准方法、电子设备以及计算机可读存储介质
CN118864861A (zh) 影像自动分割模型的训练方法、影像自动分割方法及系统
CN116206331B (zh) 图像处理方法、计算机可读存储介质以及计算机设备
CN116935388A (zh) 一种皮肤痤疮图像辅助标注方法与系统、分级方法与系统
CN113779440A (zh) 一种游戏攻略收藏方法、装置、设备及存储介质
Lin et al. Adversarial learning with data selection for cross-domain histopathological breast cancer segmentation
WO2025016321A1 (zh) 图像处理方法、电子设备和计算机可读存储介质
Jiang et al. Enhanced medical image segmentation via dynamic and static attention aggregation
CN118176497A (zh) 基于视频信息查询类似病例
AU2021240232A1 (en) Data collection method and apparatus, device and storage medium
Wang et al. Remote intelligent assisted diagnosis system for hepatic echinococcosis
CN116703837B (zh) 一种基于mri图像的肩袖损伤智能识别方法及装置

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24835418

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE