WO2025201017A1 - 模型训练方法、转换方法、电子设备、介质和程序产品 - Google Patents

模型训练方法、转换方法、电子设备、介质和程序产品

Info

Publication number
WO2025201017A1
WO2025201017A1 PCT/CN2025/081603 CN2025081603W WO2025201017A1 WO 2025201017 A1 WO2025201017 A1 WO 2025201017A1 CN 2025081603 W CN2025081603 W CN 2025081603W WO 2025201017 A1 WO2025201017 A1 WO 2025201017A1
Authority
WO
WIPO (PCT)
Prior art keywords
eye
modality
image
video
model
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/081603
Other languages
English (en)
French (fr)
Inventor
何明光
施丹莉
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hong Kong Polytechnic University HKPU
Original Assignee
Hong Kong Polytechnic University HKPU
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hong Kong Polytechnic University HKPU filed Critical Hong Kong Polytechnic University HKPU
Publication of WO2025201017A1 publication Critical patent/WO2025201017A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • G06N3/0455Auto-encoder networks; Encoder-decoder networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0475Generative networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/70Determining position or orientation of objects or cameras
    • G06T7/77Determining position or orientation of objects or cameras using statistical methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/46Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/18Eye characteristics, e.g. of the iris
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/18Eye characteristics, e.g. of the iris
    • G06V40/193Preprocessing; Feature extraction
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30004Biomedical image processing
    • G06T2207/30041Eye; Retina; Ophthalmic

Definitions

  • the present disclosure relates to the field of image processing technology, and in particular to a method for training an eye dynamic sequence generation model, a method for converting a static ophthalmic image of a source modality into a dynamic ophthalmic video of a target modality, a device for training an eye dynamic sequence generation model, a device for converting a static ophthalmic image of a source modality into a dynamic ophthalmic video of a target modality, an electronic device, a computer-readable storage medium, and a computer program product.
  • Ophthalmic diagnostic technologies used in clinical ophthalmology include Fundus Fluorescein Angiography (FFA), Indocyanine Green Angiography (ICGA) and Optical Coherence Tomography (OCT).
  • FFA Fundus Fluorescein Angiography
  • ICGA Indocyanine Green Angiography
  • OCT Optical Coherence Tomography
  • the implementation of FFA requires the intravenous injection of contrast agents into the patient, which may cause side effects for some patients
  • ICGA is an invasive diagnostic technology.
  • OCT is a non-invasive diagnostic technology that can provide detailed information about the internal structure of the retina, it also has the disadvantages of high equipment and operation costs and the need for professional knowledge.
  • the purpose of the present disclosure is to provide a training method for a cross-modal ocular dynamic sequence generation model, a method for converting a static ophthalmic image of a source modality into a dynamic ophthalmic video of a target modality, a training device, a conversion device, an electronic device, a storage medium and a computer program, so as to at least to some extent overcome the problem that the diagnosis of ocular structure in related technologies is costly or invasive to the human body.
  • a training method for an eye dynamic sequence generation model comprising: performing pixel-level alignment processing on a source modality eye image and a corresponding target modality eye structure diagnosis video; selecting a target modality key frame in the target modality eye structure diagnosis video after the alignment processing, wherein the target modality key frame represents the modal characteristics of the target modality eye structure diagnosis video; using the aligned source modality eye image as a first initial domain and the corresponding target modality key frame as a first target domain, and training an image-to-image generation network to obtain a first model; using the target modality key frame as a second initial domain and the aligned target modality eye structure diagnosis video as a second target domain, and training an image-to-video generation network to obtain a second model for frame-by-frame prediction of the eye dynamic sequence, so as to obtain the eye dynamic sequence generation model based on the first model and the second model.
  • a source modality eye image and a corresponding target modality eye structure diagnostic video are subjected to pixel-level alignment processing, including: for the source module eye image and the target modality eye structure diagnostic video from the same eye in the same examination, a feature segmentation algorithm is used to extract eye inspection features; and pixel-level alignment processing is performed based on the eye inspection features.
  • pixel-level alignment processing is performed based on the eye inspection features, including: the eye inspection features include retinal blood vessels, pixel-level key points of the retinal blood vessels are detected using a key point detector, and the pixel-level key points are used as an alignment reference to perform feature matching between the source module eye image and the target modality eye structure diagnosis video; matching feature points are estimated based on a random sampling consistency algorithm to estimate the homography matrix between the source module eye image and the retinal blood vessels in the target modality eye structure diagnosis video; abnormal frames in the eye structure diagnosis video are excluded based on the homography matrix to obtain the aligned source module eye image and the target modality eye structure diagnosis video.
  • the aligned source modality eye image is used as the first initial domain
  • the corresponding target modality key frame is used as the first target domain
  • the image-to-image generation network is trained to obtain a first model, including: the image-to-image generation network includes a U-shaped structure generator, the lesion supervision loss is used as the loss function, the first initial domain and the first target domain are used to train the U-shaped structure generator, so that the U-shaped structure generator converts the input source modality eye image into a conversion frame similar to the target modality key frame, wherein the U-shaped structure generator includes an encoder and a decoder, the encoder is used to encode the source modality eye image into a representation of a latent space, and the decoder is used to decode the representation of the latent space into the target modality key frame; the conversion frame and the corresponding target modality key frame are used as training data sets, and the perceptual loss and feature matching loss are used as loss functions to continue training the generator to obtain the first model.
  • FIG1 is a schematic diagram showing a structure of a system for converting a static ophthalmic image of a source modality into a dynamic ophthalmic video image of a target modality according to an embodiment of the present disclosure
  • FIG2 shows a flow chart of a training method for a cross-modal eye dynamic sequence generation model according to an embodiment of the present disclosure
  • FIG3 shows a flow chart of another method for training a cross-modal eye dynamic sequence generation model according to an embodiment of the present disclosure
  • FIG4 is a schematic diagram showing a training method for a cross-modal eye dynamic sequence generation model according to an embodiment of the present disclosure
  • FIG5 is a schematic diagram showing another training method for a cross-modal eye dynamic sequence generation model according to an embodiment of the present disclosure
  • FIG6 shows a flow chart of a training method for yet another cross-modal eye dynamic sequence generation model according to an embodiment of the present disclosure
  • FIG7 is a schematic diagram showing a method for converting a static ophthalmic image of a source modality into a dynamic ophthalmic video of a target modality according to an embodiment of the present disclosure
  • FIG8 shows a flow chart of a method for converting a static ophthalmic image of a source modality into a dynamic ophthalmic video of a target modality according to an embodiment of the present disclosure
  • FIG9 is a schematic diagram showing another method for converting a static ophthalmic image of a source modality into a dynamic ophthalmic video of a target modality according to an embodiment of the present disclosure
  • FIG11 is a schematic diagram of a device for converting a static ophthalmic image of a source modality into a dynamic ophthalmic video of a target modality according to an embodiment of the present disclosure
  • FIG12 shows a structural block diagram of a computer device according to an embodiment of the present disclosure.
  • FFA Fundus Fluorescein Angiography
  • ICGA Indocyanine Green Angiography
  • OCT Optical Coherence Tomography
  • color fundus photography is non-invasive and rapid. However, since it can only capture static images, it cannot provide clear contrast between pathological lesions and normal structures, and lacks the ability to detect dynamic processes observed in the above diagnostic methods. Therefore, using machine learning and image-to-video translation technology to generate various dynamic ophthalmic videos from color fundus images has important clinical prospects.
  • FIG1 shows a schematic structural diagram of a warning information mass sending decision system according to an embodiment of the present disclosure, which includes a plurality of terminals 120 and a server cluster 140 .
  • the terminal 120 may be installed with an application for providing group warning information decision-making.
  • the system may further include a management device (not shown in FIG1 ), which is connected to the server cluster 140 via a communication network.
  • the communication network is a wired network or a wireless network.
  • the above-mentioned wireless network or wired network uses standard communication technologies and/or protocols.
  • the network is typically the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of virtual private networks.
  • technologies and/or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network.
  • conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can also be used to encrypt all or some links.
  • customized and/or dedicated data communication technologies may be used to replace or supplement the above-mentioned data communication technologies.
  • a method for training a cross-modal eye dynamic sequence generation model includes:
  • Step S202 performing pixel-level alignment processing on the source modality eye image and the corresponding target modality eye structure diagnosis video.
  • the source modality is a 2D modality and the target modality is a 3D modality.
  • the target modality eye structure diagnosis video includes fluorescein fundus angiography or indocyanine green angiography, and the corresponding eye image includes a fundus image.
  • the eye structure diagnostic video of the target modality includes an optical coherence tomography video
  • the corresponding eye image includes an eye planar scan image
  • some representative frames are selected from the aligned eye structure diagnosis video as key frames. These key frames should be able to represent the modal characteristics of the entire video and maintain consistency with the aligned eye image.
  • the key frame may be the first frame of the eye structure diagnosis video, or may be a video in the eye structure diagnosis video that has an angle mapping relationship with the eye image.
  • step S206 the aligned source modality eye image is used as the first initial domain, and the corresponding target modality key frame is used as the first target domain, and the image-to-image generative network is trained to obtain a first model.
  • pixel-level alignment of an eye image and a corresponding target modality eye structure diagnosis video includes:
  • a feature segmentation algorithm is used to extract eye inspection features.
  • the feature segmentation algorithm is a blood vessel segmentation algorithm
  • the eye inspection feature is a retinal blood vessel, that is, the blood vessel segmentation algorithm is used to perform the retinal blood vessel extraction operation.
  • the vascular segmentation algorithm is a computational method used to extract vascular structures from eye images.
  • the vascular segmentation algorithm includes methods based on threshold processing, edge detection, morphological operations, machine learning and deep learning. These algorithms can help doctors diagnose eye diseases and play an important role in ophthalmic image analysis.
  • the ocular inspection features are retinal blood vessels, that is, pixel-level alignment processing is performed based on the retinal blood vessels.
  • the alignment process can be achieved through image registration technology, such as feature point matching or similarity transformation-based methods.
  • image registration technology such as feature point matching or similarity transformation-based methods.
  • the purpose of this step is to align eye images obtained at different examination time points or with different imaging methods for subsequent comparison and analysis.
  • the eye inspection features include retinal blood vessels, and pixel-level alignment processing based on the eye inspection features includes:
  • a keypoint detector is used to detect pixel-level keypoints of retinal blood vessels, and the pixel-level keypoints are used as alignment benchmarks for feature matching between the source module eye image and the target modality eye structure diagnosis video.
  • the keypoint detector is an AKAZE keypoint detector.
  • AKAZE Accelerated-KAZE
  • KAZE 2D Feature Detection and Descriptor Algorithm
  • the AKAZE algorithm uses a nonlinear scale space to detect key points in retinal vascular images by using Gaussian filters at different image resolutions, and then uses a fast feature detection algorithm to detect image areas with high local symmetry, which are generally considered to be key points in retinal vessels.
  • the matching feature points are estimated based on a random sampling consistency algorithm to estimate the homography matrix between the source module eye image and the retinal blood vessels in the target modality eye structure diagnosis video.
  • the Random Sampling Consensus Algorithm is an iterative method for fitting mathematical models and eliminating outliers. It can effectively process data containing noise and outliers.
  • the RANSAC algorithm can estimate these matching points to determine the homography transformation relationship between the two images, namely the homography matrix.
  • the homography matrix describes the geometric transformation relationship between two images with different perspectives and can be used to achieve image alignment and registration.
  • Abnormal frames in the eye structure diagnosis video are excluded based on the homography matrix to obtain aligned source module eye images and target modality eye structure diagnosis video.
  • Abnormal frames are those with poor image quality or that cannot be aligned with the eye image due to factors such as motion blur and noise. By applying the homography matrix, these abnormal frames can be filtered out, resulting in an aligned source module eye image and target modality eye structure diagnosis video.
  • the angiography video frames in the ocular structure diagnosis video of the same eye are first aligned and registered with each other, and then aligned and registered with the eye image.
  • a key point detector is used for feature matching, and RANSAC (random sampling consensus) is used to generate a homography matrix and outlier rejection.
  • RANSAC random sampling consensus
  • validity restrictions can be added and image pairs with poor registration performance can be filtered out. This can be set empirically based on the data set.
  • the aligned source modality eye image is used as the first initial domain, and the corresponding target modality key frame is used as the first target domain.
  • the image-to-image generative network is trained to obtain a first model, including:
  • the image-to-image generation network includes a U-shaped structure generator, which uses the lesion supervision loss as the loss function, and uses the first initial domain and the first target domain to train the U-shaped structure generator so that the U-shaped structure generator converts the input source modality eye image into a conversion frame similar to the target modality key frame, wherein the U-shaped structure generator includes an encoder and a decoder, the encoder is used to encode the source modality eye image into a representation of the latent space, and the decoder is used to decode the representation of the latent space into output data.
  • step S304 the converted frames and the corresponding target modality key frames are used as training data sets, and the perceptual loss and feature matching loss are used as loss functions to continue training the generator.
  • the aligned first initial domain and the first target domain training are input into a pre-designed generative network.
  • the generator is constructed with a series of stacked transposed convolutional layers to incrementally enhance the image resolution.
  • the generator is enriched by integrating jump connections and low-level and high-level features to maintain details and contextual information. This strategy enables the generator to gradually generate complex images similar to real eye structure diagnostic videos.
  • the eye image with a first resolution and the target modality key frame are input into the local enhancement network and upsampled based on the attention mechanism to output a local enhanced image with a second resolution, where the second resolution is N times the first resolution, and N is greater than 1.
  • An attention mechanism is introduced during the upsampling process of G1 and G2 to transfer information from coarse scale to fine scale and enhance task-specific response areas in shallower network layers.
  • an attention mechanism is incorporated into the network architecture to enhance information transfer, ensure that relevant features participate in deeper updates, and improve the overall quality of the generated transformation frames.
  • the source modality eye image is a fundus image
  • the corresponding target modality eye structure diagnostic video is fluorescein fundus angiography.
  • the fundus image 402 is input into the generator 404 in the first model to obtain a conversion frame 406. Referring to Figure 4, it can be seen that there is a high degree of similarity between the conversion frame 406 and the key frame 408 of the real fluorescein fundus angiography corresponding to the fundus image 402.
  • the generative adversarial network is used as an image-to-image generative network to train the first model.
  • the eye image is a plane scanning image
  • the corresponding target modality eye structure diagnosis video is an optical coherence tomography video.
  • the specific processing process includes: inputting the plane scanning image 502 into the generator 504 to obtain the conversion frame 506, inputting the key frame 508 of the real optical coherence tomography video corresponding to the conversion frame 506 and the plane scanning image 502 into the discriminator 510, and outputting the discrimination result to detect the training result of the first model based on the discrimination result.
  • the target modality keyframes are used as the second initial domain, and the aligned target modality eye structure diagnosis video is used as the second target domain.
  • the image-to-video generative network is trained to obtain a second model for frame-by-frame prediction of eye dynamic sequences, including:
  • the loss function used in the training process of the second model is generated based on the reconstruction loss, smoothness constraint, consistency loss and loss in rendering.
  • the target modality key frame is input into the image generation network of the video for prediction processing to obtain the first frame of the generated eye dynamic sequence, including:
  • the motion vector of pixels between two adjacent video frames in the target modality eye structure diagnosis video is calculated as the feature label.
  • the feature labels describing the motion between video frames can be obtained.
  • the motion vectors of these pixels can characterize the motion state between video frames, which is manifested as vascular tissue contrast and dynamic insights, and serve as feature labels for the subsequent input image to video generation network.
  • an eye image is used as the first frame input to the generative network, which outputs the first video frame of the eye dynamic sequence.
  • the eye image and the generated first frame are then coupled as input to generate the next frame, and so on.
  • the generative network processes the feature labels of the ocular structure diagnostic video and generates the first frame of the eye dynamic sequence based on these feature labels.
  • This method can generate an eye dynamic sequence that meets the feature label requirements based on the changes in pixel motion vectors between video frames.
  • it also includes: calculating the difference between the first frame and the last frame in the target modality eye structure diagnosis video based on the temporal consistency restriction mechanism; performing threshold processing on the difference to obtain the corresponding clinical knowledge supervision mask.
  • a clinical knowledge supervision mask is obtained by taking the difference between the first and last frames in the target modality eye structure diagnosis video, while ensuring accuracy without the need for additional manual annotation or model training.
  • a clinical knowledge supervision mask is used to supervise the training of a first model and the training of a second model, including: in the training of the first model, guiding attention to a first region with significant temporal changes in the generated transition frame based on a knowledge-enhanced attention mechanism; in the training of the second model, guiding attention to a second region with significant temporal changes in the generated transition frame based on a knowledge-enhanced attention mechanism to determine a key region based on the first region and/or the second region; performing temporal temporal consistency and change perception operations on the generated eye dynamic sequence based on the clinical knowledge supervision mask, wherein, in the change perception operation, specific supervision is provided for the key region based on the knowledge-aware discriminator loss; and adjusting the pixel misalignment between the source modality eye image and the key region in the target modality eye structure diagnosis video based on the mask-enhanced local normalized cross entropy loss.
  • bidirectional optical flow refers to the method used in computer vision to estimate the direction and speed of optical flow at the pixel level in video sequences.
  • bidirectional optical flow is introduced to perform consistency check to ensure accurate alignment and consistency between generated video frames, thereby ensuring smooth playback of the generated eye dynamic sequence.
  • the first model is used to output key frames
  • the second model is used to convert the static ophthalmic image of the source modality, i.e., the key frame, into a dynamic ophthalmic video of the target modality.
  • the source modality eye image is a plane scanning image
  • the corresponding target modality eye structure diagnostic video is an optical coherence tomography video.
  • the target modality key frame is input into the second model to obtain a simulated optical coherence tomography video 704.
  • a method for converting a static ophthalmic image of a source modality into a dynamic ophthalmic video of a target modality includes:
  • Step S804 Perform cross-modal conversion on the eye image based on the first model to generate a target modality keyframe.
  • Step S806 Input the target modality key frame into the second model to predict the multiple video frames based on the time sequence to obtain the multiple video frames, so as to generate a dynamic ophthalmic video of the target modality based on the multiple video frames.
  • the eye image includes a color fundus image
  • the eye dynamic sequence includes a simulated fluorescein fundus angiography.
  • the following is a process of generating a simulated fluorescein fundus angiography (FFA) based on a color fundus image (CFP), and further describes in detail the scheme of matching the retinal blood vessels in the CFP and the real FFA at the pixel level in the present disclosure, and training the generation network to predict high-resolution FFA video in an autoregressive manner.
  • FFA simulated fluorescein fundus angiography
  • CFP color fundus image
  • the generative model was pre-trained to predict keyframes from FFA videos.
  • the weights were transferred to CFP-FFA video pairs to predict FFA videos from the input CFP.
  • Validation results of this scheme demonstrated realistic generation on xx internal and external test sets subjectively evaluated by three ophthalmologists.
  • DR Diabetic Retinopathy
  • AMD Alge-Related Macular Degeneration
  • multi-disease dataset a challenging multi-class retinal image library with many rare diseases and severe class imbalance
  • adding the generated eye dynamic sequences improved the diagnostic accuracy for DR, AMD, and multiple rare diseases, respectively.
  • the validation results demonstrate that this technology can serve as a novel retinal foundation model and can be immediately implemented to improve automated retinal disease screening processes.
  • Fluorescein angiography is a key method for detecting lesions associated with vascular-retinal barrier breakdown and monitoring DR treatment response. This technology dynamically highlights lesion changes through the injection of dye, which is particularly helpful in highlighting important lesions that are difficult to see clearly on color fundus photographs.
  • FFA Fluorescein angiography
  • This technology dynamically highlights lesion changes through the injection of dye, which is particularly helpful in highlighting important lesions that are difficult to see clearly on color fundus photographs.
  • FFA Fluorescein angiography
  • FFA Fluorescein angiography
  • Preprocessing stage Fundus images CF and angiography videos from the same eye and the same visit are used for matching. Retinal vessels are extracted from the fundus images CF and angiography videos using a vessel segmentation algorithm to achieve pixel-level image matching. First, angiography video frames of the same eye are registered with each other and then with the fundus images CF. Feature matching is performed using the AKAZE keypoint detector, and RANSAC (random sample consensus) is used to generate homography matrices and outlier rejection. In order to exclude incorrectly registered pairs, validity constraints are added and image pairs with poor registration performance are filtered out, which are set empirically based on the dataset.
  • AKAZE keypoint detector AKAZE keypoint detector
  • RANSAC random sample consensus
  • a generative network is trained to generate all video frames.
  • the generative network is constructed with a series of stacked transposed convolutional layers to incrementally enhance image resolution.
  • the generator is enriched by integrating skip connections and low-level and high-level features to maintain details and contextual information.
  • FFA videos are generated in an autoregressive manner, where the input is the first frame, followed by the next frame of the venous phase, and the output will be their next frame.
  • a gradient-guided loss is added to enhance the generation of high-frequency components, including retinal structures and lesions.
  • the input of this process will be the fundus image CF, and the output will be the FFA video from the venous phase to the late phase.
  • Evaluation phase The generated simulated videos will be evaluated against real videos based on the following metrics:
  • SSIM Structural Similarity Measure
  • MS-SSIM Multi-Scale Structural Similarity Measure
  • the eye image is a fundus image
  • the corresponding target modality eye structure diagnosis video is fluorescein fundus angiography.
  • cross-modal eye dynamic sequence generation model training device 1000 according to an embodiment of the present invention with reference to Figure 10.
  • the cross-modal eye dynamic sequence generation model training device 1000 shown in Figure 10 is merely an example and should not limit the functionality and scope of use of the embodiments of the present invention.
  • the training device 1000 for the cross-modal eye dynamic sequence generation model is expressed in the form of a hardware module.
  • the components of the training device 1000 for a cross-modal eye dynamic sequence generation model may include but are not limited to: a processing module 1002, used to perform pixel-level alignment processing on a source modality eye image and a corresponding target modality eye structure diagnosis video; a selection module 1004, used to select a target modality key frame in the target modality eye structure diagnosis video after the alignment processing, the target modality key frame representing the modal features of the target modality eye structure diagnosis video; a first training module 1006, used to train an image-to-image generation network using the aligned source modality eye image as a first initial domain and the corresponding target modality key frame as a first target domain to obtain a first model; a second training module 1008, used to train an image-to-video generation network using the target modality key frame as a second initial domain and the aligned target modality eye structure diagnosis video as a second target domain to obtain a second model for frame-
  • aspects of the present invention may be implemented as systems, methods, or program products. Therefore, various aspects of the present invention may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits,” “modules,” or “systems.”
  • the electronic device 1200 according to this embodiment of the present invention is described below with reference to FIG12.
  • the electronic device is an eye dynamic sequence generator.
  • the electronic device 1200 shown in FIG12 is only an example and should not limit the functions and scope of use of the embodiment of the present invention.
  • the storage unit stores program code, which can be executed by the processing unit 1210, so that the processing unit 1210 performs the steps according to various exemplary embodiments of the present invention described in the "Exemplary Method" section above.
  • the processing unit 1210 can perform the scheme described in steps S202 to S208 as shown in Figure 2.
  • the storage unit 1220 may also include a program/utility 12204 having a set (at least one) of program modules 12205, such program modules 12205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
  • program modules 12205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
  • the bus 1230 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
  • the electronic device 1200 can also communicate with one or more external devices 1270 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 1200, and/or any device that enables the electronic device 1200 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input/output (I/O) interface 1250.
  • the electronic device 1200 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and/or a public network such as the Internet) via a network adapter 1260.
  • networks e.g., a local area network (LAN), a wide area network (WAN), and/or a public network such as the Internet
  • the network adapter 1260 communicates with other modules of the electronic device 1200 via a bus 1230. It should be understood that, although not shown in the figure, other hardware and/or software modules can be used in conjunction with the electronic device 1200, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
  • the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or an electronic device, etc.) to execute the method according to the embodiments of the present disclosure.
  • a non-volatile storage medium which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.
  • a computing device which can be a personal computer, a server, a terminal device, or an electronic device, etc.
  • a computer-readable storage medium is also provided, on which is stored a program product capable of implementing the methods described above.
  • various aspects of the present invention may also be implemented in the form of a program product comprising program code that, when executed on a terminal device, causes the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.
  • a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof.
  • a readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
  • the program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
  • the program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Python, Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages.
  • the program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
  • the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
  • LAN local area network
  • WAN wide area network
  • Internet service provider e.g., via the Internet using an Internet service provider
  • the technical solution according to the embodiment of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or an electronic device, etc.) to execute the method according to the embodiment of the present disclosure.
  • a non-volatile storage medium which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.
  • a computing device which can be a personal computer, a server, a mobile terminal, or an electronic device, etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • Evolutionary Computation (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • General Engineering & Computer Science (AREA)
  • Biophysics (AREA)
  • Computational Linguistics (AREA)
  • Data Mining & Analysis (AREA)
  • Biomedical Technology (AREA)
  • Multimedia (AREA)
  • Molecular Biology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Biology (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Probability & Statistics with Applications (AREA)
  • Ophthalmology & Optometry (AREA)
  • Human Computer Interaction (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Image Analysis (AREA)
  • Eye Examination Apparatus (AREA)

Abstract

本公开提供了一种模型训练方法、转换方法、电子设备、介质和程序产品,涉及图像处理技术领域。其中,模型训练方法包括:将源模态的眼部图像和对应的目标模态的眼部结构诊断视频进行像素级的对齐处理;选取对齐处理后目标模态的眼部结构诊断视频中的关键帧;将对齐处理后的源模态的眼部图像作为第一初始域,将对应的目标模态的关键帧作为第一目标域进行训练,得到第一模型;将目标模态的关键帧作为第二初始域,将对齐处理后的目标模态的眼部结构诊断视频作为第二目标域进行训练,得到第二模型,以得到眼部动态序列生成模型。通过本公开的技术方案,将源模态的静态眼科图像转换为目标模态的动态眼科视频,并具有非侵入性、安全和低成本的特点。

Description

模型训练方法、转换方法、电子设备、介质和程序产品
本公开要求于2024年03月27日提交的申请号为202410360491.4、名称为“模型训练方法、转换方法、电子设备、介质和程序产品”的中国专利申请的优先权,该中国专利申请的全部内容通过引用全部并入本文。
技术领域
本公开涉及图像处理技术领域,尤其涉及一种眼部动态序列生成模型的训练方法、一种将源模态的静态眼科图像转换为目标模态的动态眼科视频的方法、一种眼部动态序列生成模型的训练装置、一种将源模态的静态眼科图像转换为目标模态的动态眼科视频的装置、一种电子设备、一种计算机可读存储介质和一种计算机程序产品。
背景技术
在临床眼科领域使用的眼科诊断技术包括眼底荧光血管造影(Fluorescein Fundus Angiography,FFA)、吲哚青绿血管造影(Indocyanine Green Angiography,ICGA)和光学相干断层扫描(Optical Coherence Tomography,OCT)等,其中,FFA的实施需要向患者静脉注射造影剂,可能会对某些患者产生副作用,而ICGA是一种侵入性诊断技术,虽然OCT是一种非侵入性诊断技术,并能够提供有关视网膜内部结构的详细信息,但是也存在设备和操作成本高昂,需要专业知识才能操作的缺陷。
因此亟需一种应用于眼科诊断、无创且经济高效的造影成像方案。
需要说明的是,在上述背景技术部分公开的信息仅用于加强对本公开的背景的理解,因此可以包括不构成对本领域普通技术人员已知的现有技术的信息。
发明内容
本公开的目的在于提供一种跨模态眼部动态序列生成模型的训练方法、一种将源模态的静态眼科图像转换为目标模态的动态眼科视频的方法、训练装置、转换装置、电子设备、存储介质和计算机程序,至少在一定程度上克服由于相关技术中眼部结构诊断成本高或对人体具有侵入性的问题。
本公开的其他特性和优点将通过下面的详细描述变得显然,或部分地通过本公开的实践而习得。
根据本公开的一个方面,提供一种眼部动态序列生成模型的训练方法,包括:将源模态眼部图像和对应的目标模态眼部结构诊断视频进行像素级的对齐处理;选取对齐处理后的所述目标模态眼部结构诊断视频中的目标模态关键帧,所述目标模态关键帧表征所述目标模态眼部结构诊断视频的模态特征;将对齐处理后的所述源模态眼部图像作为第一初始域,将对应的所述目标模态关键帧作为第一目标域,对图像到图像的生成网络进行训练,得到第一模型;将所述目标模态关键帧作为第二初始域,将对齐处理后的所述目标模态眼部结构诊断视频作为第二目标域,对图像到视频的生成网络进行训练,得到用于对所述眼部动态序列进行逐帧预测的第二模型,以基于所述第一模型和所述第二模型得到所述眼部动态序列生成模型。
在本公开的一个实施例中,将源模态眼部图像和对应的目标模态眼部结构诊断视频进行像素级的对齐处理,包括:针对同一检查中,来自于同一眼睛的所述源模块眼部图像和所述目标模态眼部结构诊断视频,采用特征分割算法进行眼部检验特征的提取操作;基于所述眼部检验特征进行像素级的对齐处理。
在本公开的一个实施例中,基于所述眼部检验特征进行像素级的对齐处理,包括:所述眼部检验特征包括视网膜血管,使用关键点检测器检测所述视网膜血管的像素级关键点,将所述像素级关键点作为对齐基准进行所述源模块眼部图像和所述目标模态眼部结构诊断视频之间的特征匹配;基于随机抽样一致性算法对匹配特征点进行估计,估计所述源模块眼部图像和所述目标模态眼部结构诊断视频中的所述视网膜血管之间的单应性矩阵;基于所述单应性矩阵排除所述眼部结构诊断视频中的异常帧,得到对齐的所述源模块眼部图像和所述目标模态眼部结构诊断视频。
在本公开的一个实施例中,将对齐处理后的所述源模态眼部图像作为第一初始域,将对应的所述目标模态关键帧作为第一目标域,对图像到图像的生成网络进行训练,得到第一模型,包括:所述图像到图像的生成网络包括U型结构的生成器,将病变监督损失作为损失函数,采用所述第一初始域和所述第一目标域训练所述U型结构的生成器,以使所述U型结构的生成器将输入的所述源模态眼部图像转换为与所述目标模态关键帧相像的转换帧,其中,所述U型结构的生成器包括编码器和解码器,所述编码器用于将所述源模态眼部图像编码为潜在空间的表示,所述解码器用于将所述潜在空间的表示解码为所述目标模态关键帧;将所述转换帧和对应的所述目标模态关键帧作为训练数据集,将感知损失和特征匹配损失作为损失函数继续训练所述生成器,得到所述第一模型。
在本公开的一个实施例中,所述生成器包括全局生成器网络和局部增强网络,训练所述生成器,还包括:将具有第一分辨率的所述眼部图像和所述目标模态关键帧输入所述局部增强网络基于注意力机制进行上采样,以输出具有第二分辨率的局部增强图像,所述第二分辨率为所述第一分辨率的N倍,N大于1。
在本公开的一个实施例中,将所述目标模态关键帧作为第二初始域,将对齐处理后的所述目标模态眼部结构诊断视频作为第二目标域,对图像到视频的生成网络进行训练,得到用于对所述眼部动态序列进行逐帧预测的第二模型,包括:将所述目标模态关键帧输入所述图像到视频的生成网络进行预测处理,得到生成的眼部动态序列的第一帧;将所述第一帧和所述目标模态关键帧进行耦合后输入所述图像到视频的生成网络进行预测处理,得到所述生成的眼部动态序列的第二帧;基于时序逐帧预测剩余的视频帧,直至所述图像到视频的生成网络预测出所述生成的眼部动态序列的所有视频帧;迭代训练所述图像到视频的生成网络,以在对应的损失函数确定所述生成的眼部动态序列和作为金标准的所述目标模态眼部结构诊断视频之间的差异小于误差阈值时,将训练完毕的所述图像到视频的生成网络确定为所述第二模型。
在本公开的一个实施例中,将所述目标模态关键帧输入所述图像到视频的生成网络进行预测处理,得到生成的眼部动态序列的第一帧,包括:计算所述目标模态眼部结构诊断视频中相邻两个视频帧之间像素的运动矢量,作为特征标签;将所述关键帧和所述第一帧对应的特征标签输入所述图像到视频的生成网络,以基于所述特征标签对所述关键帧进行扭曲预测,输出所述眼部动态序列的第一帧。
在本公开的一个实施例中,还包括:基于双向光流计算对所述眼部动态序列中的相邻帧进行前后一致性检查,以将所述相邻帧进行对齐。
在本公开的一个实施例中,还包括:基于颞侧一致性限制机制计算所述目标模态眼部结构诊断视频中的第一帧与最后一帧之间的差值;对所述差值进行阈值处理,得到对应的临床知识监督掩膜。
在本公开的一个实施例中,还包括:使用所述临床知识监督掩膜监督所述第一模型的训练和所述第二模型的训练。
在本公开的一个实施例中,使用所述临床知识监督掩膜监督所述第一模型的训练和所述第二模型的训练,包括:在所述第一模型的训练中,基于知识增强注意力机制引导关注生成的转换帧中具有显著时间变化的第一区域;在所述第二模型的训练中,基于所述知识增强注意力机制引导关注生成的转换帧中具有显著时间变化的第二区域,以基于所述第一区域和/或所述第二区域确定关键区域;基于所述临床知识监督掩膜对生成的眼部动态序列进行时序颞侧的一致性和改变感知操作,其中,在所述改变感知操作中,基于知识感知判别器损失为所述关键区域提供特定监督;以及基于掩膜增强的局部归一化交叉熵损失调整所述源模态眼部图像与所述目标模态眼部结构诊断视频中的所述关键区域之间的像素错位。
在本公开的一个实施例中,所述目标模态的眼部结构诊断视频包括荧光素眼底血管造影或吲哚青绿血管造影,对应的所述眼部图像包括眼底图像;所述目标模态的眼部结构诊断视频包括光学相干断层扫描视频,对应的所述眼部图像包括眼部平面扫描图像。
在本公开的一个实施例中,所述生成网络包括生成对抗网络、所述生成对抗网络的衍生网络、扩散模型、所述扩散模型的衍生模型、变分自编码器、所述变分自编码器的衍生结构中的至少一种。
根据本公开的另一方面,提供一种将源模态的静态眼科图像转换为目标模态的动态眼科视频方法,包括:将源模态的静态眼科图像输入眼部动态序列生成模型,所述眼部动态序列生成模型包括第一模型和第二模型,其中,所述第一模型和第二模型基于生成网络生成;基于所述第一模型将所述眼部图像进行跨模态转换,生成目标模态关键帧;将所述目标模态关键帧输入所述第二模型以对多个视频帧基于时序进行预测,得到所述多个视频帧,以基于所述多个视频帧生成所述目标模态的动态眼科视频。
根据本公开的再一方面,提供一种跨模态眼部动态序列生成模型的训练装置,包括:处理模块,用于将源模态眼部图像和对应的目标模态眼部结构诊断视频进行像素级的对齐处理;选取模块,用于选取对齐处理后的所述目标模态眼部结构诊断视频中的目标模态关键帧,所述目标模态关键帧表征所述目标模态眼部结构诊断视频的模态特征;第一训练模块,用于将对齐处理后的所述源模态眼部图像作为第一初始域,将对应的所述目标模态关键帧作为第一目标域,对图像到图像的生成网络进行训练,得到第一模型;第二训练模块,用于将所述目标模态关键帧作为第二初始域,将对齐处理后的所述目标模态眼部结构诊断视频作为第二目标域,对图像到视频的生成网络进行训练,得到用于对所述眼部动态序列进行逐帧预测的第二模型,以基于所述第一模型和所述第二模型得到所述眼部动态序列生成模型。
根据本公开的又一方面,提供一种将源模态的静态眼科图像转换为目标模态的动态眼科视频装置,包括:输入模块,用于将源模态的静态眼科图像输入眼部动态序列生成模型,所述眼部动态序列生成模型包括第一模型和第二模型,其中,所述第一模型和第二模型基于生成网络生成;转换模块,用于基于所述第一模型将所述眼部图像进行跨模态转换,生成目标模态关键帧;预测模块,用于将所述目标模态关键帧输入所述第二模型以对多个视频帧基于时序进行预测,得到所述多个视频帧,以基于所述多个视频帧生成所述目标模态的动态眼科视频。
根据本公开的又一方面,提供一种电子设备,包括:处理器;以及存储器,用于存储处理器的可执行指令;所述处理器配置为经由执行所述可执行指令来执行上述的跨模态眼部动态序列生成模型的训练方法或将源模态的静态眼科图像转换为目标模态的动态眼科视频方法。
根据本公开的又一方面,提供一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现上述的跨模态眼部动态序列生成模型的训练方法或将源模态的静态眼科图像转换为目标模态的动态眼科视频方法。
根据本公开的又一方面,提供一种计算机程序产品,其上存储有计算机程序,所述计算机程序被处理器执行时实现上述的跨模态眼部动态序列生成模型的训练方法或将源模态的静态眼科图像转换为目标模态的动态眼科视频方法。
本公开的实施例所提供的跨模态眼部动态序列生成模型的训练和将源模态的静态眼科图像转换为目标模态的动态眼科视频方案,通过对齐处理,能够将眼部图像与目标模态的眼部结构诊断视频在像素级别上对齐,确保它们具有一致的空间信息,以在模型训练过程中能够准确捕捉到视网膜血管等的病变特征,进一步地,选取能够表征眼部结构诊断视频模态特征的关键帧,保持与眼部图像对齐处理后的一致性,通过生成模型的训练,可以学习到源模态眼部图像到目标模态眼部结构诊断视频之间的映射关系,从而实现逼真的眼部动态序列的生成,眼部动态序列的生成方式不但具有非侵入性、安全和低成本的特点,并且生成的眼部动态序列结果能够保证真实性和准确性,以作为临床诊断的可靠判断依据。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性和解释性的,并不能限制本公开。
附图说明
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本公开的实施例,并与说明书一起用于解释本公开的原理。显而易见地,下面描述中的附图仅仅是本公开的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1示出本公开实施例中一种将源模态的静态眼科图像转换为目标模态的动态眼科视频系统结构的示意图;
图2示出本公开实施例中一种跨模态眼部动态序列生成模型的训练方法流程图;
图3示出本公开实施例中另一种跨模态眼部动态序列生成模型的训练方法流程图;
图4示出本公开实施例中一种跨模态眼部动态序列生成模型的训练方法示意图;
图5示出本公开实施例另一种跨模态眼部动态序列生成模型的训练方法示意图;
图6示出本公开实施例再一种跨模态眼部动态序列生成模型的训练方法流程图;
图7示出本公开实施例中一种将源模态的静态眼科图像转换为目标模态的动态眼科视频方法示意图;
图8示出本公开实施例中一种将源模态的静态眼科图像转换为目标模态的动态眼科视频方法流程图;
图9示出本公开实施例中另一种将源模态的静态眼科图像转换为目标模态的动态眼科视频方法示意图;
图10示出本公开实施例中一种跨模态眼部动态序列生成模型的训练装置示意图;
图11示出本公开实施例中一种将源模态的静态眼科图像转换为目标模态的动态眼科视频装置示意图;
图12示出本公开实施例中一种计算机设备的结构框图。
具体实施方式
现在将参考附图更全面地描述示例实施方式。然而,示例实施方式能够以多种形式实施,且不应被理解为限于在此阐述的范例;相反,提供这些实施方式使得本公开将更加全面和完整,并将示例实施方式的构思全面地传达给本领域的技术人员。所描述的特征、结构或特性可以以任何合适的方式结合在一个或更多实施方式中。
此外,附图仅为本公开的示意性图解,并非一定是按比例绘制。图中相同的附图标记表示相同或类似的部分,因而将省略对它们的重复描述。附图中所示的一些方框图是功能实体,不一定必须与物理或逻辑上独立的实体相对应。可以采用软件形式来实现这些功能实体,或在一个或多个硬件模块或集成电路中实现这些功能实体,或在不同网络和/或处理器装置和/或微控制器装置中实现这些功能实体。
在临床眼科领域,眼底荧光血管造影(Fluorescein Fundus Angiography,FFA)是研究视网膜循环动力学的关键技术,它用于诊断糖尿病视网膜病变、高血压视网膜病变和黄斑变性等疾病,然而,这项技术需要静脉注射造影剂,这可能会产生相关的副作用,可能不适合某些患者。吲哚青绿血管造影(Indocyanine Green Angiography,ICGA)是一种侵入性诊断技术,主要用于脉络膜疾病、黄斑病变、眼内肿瘤和视网膜血管疾病,光学相干断层扫描(Optical Coherence Tomography,OCT)提供了有关视网膜内部结构的详细信息。虽然这些技术能够得到眼睛内部结构的详细信息,但是成本高昂,需要专业知识才能操作,有些技术还会引起人体不适。
相比之下,彩色眼底摄影无创且快速,然而由于其只能捕获静态图像,无法在病理病变和正常结构之间提供清晰的对比,且缺乏上述诊断方法中观察到的动态过程的能力,因此,利用机器学习和图像到视频的翻译技术从彩色眼底图像生成各种动态眼科视频具有重要的临床前景。
本申请提供的方案,通过对齐处理,能够将源模态的眼部图像与目标模态的眼部结构诊断视频在像素级别上对齐,确保它们具有一致的空间信息,以在模型训练过程中能够准确捕捉到视网膜血管等的病变特征,进一步地,选取能够表征眼部结构诊断视频模态特征的关键帧,保持与眼部图像对齐处理后的一致性,通过对生成模型的训练,可以学习到源模态眼部图像到目标模态眼部结构诊断视频之间的映射关系,从而实现逼真的眼部动态序列的生成,眼部动态序列的生成方式不但具有非侵入性、安全和低成本的特点,并且生成的眼部动态序列结果能够保证真实性和准确性,以作为临床诊断的可靠判断依据。
图1示出本公开实施例中一种预警信息群发决策系统的结构示意图,包括多个终端120和服务器集群140。
终端120可以是手机、游戏主机、平板电脑、电子书阅读器、智能眼镜、MP4(Moving Picture Experts Group Audio Layer IV,动态影像专家压缩标准音频层面4)播放器、智能家居设备、AR(Augmented Reality,增强现实)设备、VR(Virtual Reality,虚拟现实)设备等移动终端,或者,终端120也可以是个人计算机(Personal Computer,PC),比如膝上型便携计算机和台式计算机等等。
其中,终端120中可以安装有用于提供的预警信息群发决策的应用程序。
终端120与服务器集群140之间通过通信网络相连。可选的,通信网络是有线网络或无线网络。
服务器集群140是一台服务器,或者由若干台服务器组成,或者是一个虚拟化平台,或者是一个云计算服务中心。服务器集群140用于为提供预警信息群发决策应用程序提供后台服务。可选地,服务器集群140承担主要计算工作,终端120承担次要计算工作;或者,服务器集群140承担次要计算工作,终端120承担主要计算工作;或者,终端120和服务器集群140之间采用分布式计算架构进行协同计算。
在一些可选的实施例中,服务器集群140用于存储将源模态的静态眼科图像转换为目标模态的动态眼科视频的程序等。
可选地,不同的终端120中安装的应用程序的客户端是相同的,或两个终端120上安装的应用程序的客户端是不同控制系统平台的同一类型应用程序的客户端。基于终端平台的不同,该应用程序的客户端的具体形态也可以不同,比如,该应用程序客户端可以是手机客户端、PC客户端或者全球广域网(World Wide Web,Web)客户端等。
本领域技术人员可以知晓,上述终端120的数量可以更多或更少。比如上述终端可以仅为一个,或者上述终端为几十个或几百个,或者更多数量。本申请实施例对终端的数量和设备类型不加以限定。
可选的,该系统还可以包括管理设备(图1未示出),该管理设备与服务器集群140之间通过通信网络相连。可选的,通信网络是有线网络或无线网络。
可选的,上述的无线网络或有线网络使用标准通信技术和/或协议。网络通常为因特网、但也可以是任何网络,包括但不限于局域网(Local Area Network,LAN)、城域网(Metropolitan Area Network,MAN)、广域网(Wide Area Network,WAN)、移动、有线或者无线网络、专用网络或者虚拟专用网络的任何组合)。在一些实施例中,使用包括超文本标记语言(Hyper Text Mark-up Language,HTML)、可扩展标记语言(Extensible MarkupLanguage,XML)等的技术和/或格式来代表通过网络交换的数据。此外还可以使用诸如安全套接字层(Secure Socket Layer,SSL)、传输层安全(Transport Layer Security,TLS)、虚拟专用网络(Virtual Private Network,VPN)、网际协议安全(Internet ProtocolSecurity,IPsec)等常规加密技术来加密所有或者一些链路。在另一些实施例中,还可以使用定制和/或专用数据通信技术取代或者补充上述数据通信技术。
下面,将结合附图及实施例对本示例实施方式中的跨模态眼部动态序列生成模型的训练方法和将源模态的静态眼科图像转换为目标模态的动态眼科视频方法进行更详细的说明。
下面,将结合附图及实施例对本示例实施方式中的眼部动态序列生成模型以及生成方法的各个步骤进行更详细的说明。
如图2所示,根据本公开的一个实施例的跨模态眼部动态序列生成模型的训练方法,包括:
步骤S202,将源模态眼部图像和对应的目标模态眼部结构诊断视频进行像素级的对齐处理。
在一些实施例中,源模态为2D模态,目标模态为3D模态。
在一些实施例中,源模态和目标模态均为2D模态或均为3D模态。
在一些实施例中,目标模态的眼部结构诊断视频包括眼底荧光血管造影、吲哚青绿血管造影或光学相干断层扫描等。
在一些实施例中,目标模态的眼部结构诊断视频包括荧光素眼底血管造影或吲哚青绿血管造影,对应的眼部图像包括眼底图像。
在一些实施例中,目标模态的眼部结构诊断视频包括光学相干断层扫描视频,对应的眼部图像包括眼部平面扫描图像。
示例性地,将成对的眼部图像和相应的荧光素钠血管造影或静态脉络膜荧光血管造影或光学相干断层扫描图像在像素级别上进行对齐,从每个组合中,抽样选取代表不同阶段(早期、中期、晚期等)的帧,以创建平衡的仿真视频,并确保在不同阶段之间保持一致性。
步骤S204,选取对齐处理后的目标模态眼部结构诊断视频中的目标模态关键帧,目标模态关键帧表征目标模态眼部结构诊断视频的模态特征。
其中,从对齐处理后的眼部结构诊断视频中,选取一些代表性的帧作为关键帧。这些关键帧应该能够表征整个视频的模态特征,并且能够保持与眼部图像对齐处理后的一致性。
在一些实施例中,关键帧可以为眼部结构诊断视频的第一帧,也可以为眼部结构诊断视频中与眼部图像具有角度映射关系的视频中。
步骤S206,将对齐处理后的源模态眼部图像作为第一初始域,将对应的目标模态关键帧作为第一目标域,对图像到图像的生成网络进行训练,得到第一模型。
其中,图像到图像的生成网络可以为源模态图像到目标模态图像的生成网络。
将对齐处理后的眼部图像作为第一初始域,将对应的关键帧作为第一目标域,使用图像到图像的生成网络对这两个域进行训练,以学习生成两个模态的图像之间的映射关系,目标是基于眼部图像生成逼真的关键帧,使其与对齐处理后的关键字在视觉上无法区分。
步骤S208,将目标模态关键帧作为第二初始域,将对齐处理后的目标模态眼部结构诊断视频作为第二目标域,对图像到视频的生成网络进行训练,得到用于对眼部动态序列进行逐帧预测的第二模型,以基于第一模型和第二模型得到眼部动态序列生成模型。
其中,将关键帧作为第二初始域,将对齐处理后的眼部结构诊断视频作为第二目标域,使用图像到视频的生成网络对这两个域进行训练,以学习条件概率分布,目标是生成具有高质量和多样性的眼部结构诊断视频。
进一步地,基于第一模型和第二模型,构建出跨模态眼部动态序列生成模型,该模型可以接受源模态眼部图像作为输入,并生成相应的眼部结构诊断视频,从而实现眼部动态序列的生成。
在一些实施例中,生成网络包括生成对抗网络、生成对抗网络的衍生网络、扩散模型、扩散模型的衍生模型、变分自编码器、变分自编码器的衍生结构中的至少一种。
在该实施例中,通过对齐处理,能够将眼部图像与目标模态的眼部结构诊断视频在像素级别上对齐,确保它们具有一致的空间信息,以在模型训练过程中能够准确捕捉到视网膜血管等的病变特征,进一步地,选取能够表征眼部结构诊断视频模态特征的关键帧,保持与眼部图像对齐处理后的一致性,通过对生成网络的训练,可以学习到源模态眼部图像到目标模态眼部结构诊断视频之间的映射关系,从而实现逼真的眼部动态序列的生成,眼部动态序列的生成方式不但具有非侵入性、安全和低成本的特点,并且生成的眼部动态序列结果能够保证真实性和准确性,从而能够在临床诊断中作为可靠判断依据。
在本公开的一个实施例中,将眼部图像和对应的目标模态的眼部结构诊断视频进行像素级的对齐处理,包括:
针对同一检查中,来自于同一眼睛的源模块眼部图像和目标模态眼部结构诊断视频,采用特征分割算法进行眼部检验特征的提取操作。
在一些实施例中,眼部检验特征包括视网膜血管、晶状体、眼睑结构、虹膜特征等。
在一些实施例中,特征分割算法为血管分割算法,眼部检验特征为视网膜血管,即采用血管分割算法进行视网膜血管的提取操作。
其中,血管分割算法是一种用于从眼部图像中提取血管结构的计算方法,血管分割算法包括基于阈值处理、边缘检测、形态学操作、机器学习和深度学习等方法,这些算法可以帮助医生诊断眼部疾病,并在眼科图像分析中起到重要作用。
完成提取操作,基于眼部检验特征进行像素级的对齐处理。
在一些实施例中,眼部检验特征为视网膜血管,即基于视网膜血管进行像素级的对齐处理。
其中,对齐处理过程可以通过图像配准技术来实现,例如基于特征点匹配或基于相似性变换的方法,该步骤的目的是将不同检查时间点或不同成像方式获得的眼部图像进行对齐,以便进行后续的比较和分析。
在该实例中,通过基于视网膜血管进行像素级的对齐处理,能够更加准确地分析和比较来自同一眼睛的多次检查结果,基于这些特征进行模型训练,使生成的眼部动态序列能够正确恢复具有病变特征的血管结构。
在本公开的一个实施例中,眼部检验特征包括视网膜血管,基于眼部检验特征进行像素级的对齐处理,包括:
使用关键点检测器检测视网膜血管的像素级关键点,将像素级关键点作为对齐基准进行源模块眼部图像和目标模态眼部结构诊断视频之间的特征匹配。
在一些实施例中,关键点检测器为AKAZE关键点检测器。
其中,AKAZE(Accelerated-KAZE)是一种加速的KAZE(2D特征检测与描述子算法)算法,AKAZE算法使用非线性尺度空间来检测视网膜血管图像中的关键点,通过在不同的图像分辨率下使用高斯滤波器实现,然后使用快速特征检测算法来检测具有高局部对称性的图像区域,这些区域通常被认为是视网膜血管中的关键点。
基于随机抽样一致性算法对匹配特征点进行估计,估计源模块眼部图像和目标模态眼部结构诊断视频中的视网膜血管之间的单应性矩阵。
其中,随机抽样一致性算法RANSAC是一种用于拟合数学模型并排除异常值的迭代方法,它可以有效地处理包含噪声和异常值的数据,通过在源模块眼部图像和目标模态眼部结构诊断视频的视频帧中提取特征点,并利用这些特征点进行匹配,RANSAC算法可以对这些匹配点进行估计,以确定两个图像之间的单应性变换关系,即单应性矩阵。单应性矩阵描述了两个透视不同的图像之间的几何变换关系,通过该矩阵可以实现图像的对齐和配准。
基于单应性矩阵排除眼部结构诊断视频中的异常帧,得到对齐的源模块眼部图像和目标模态眼部结构诊断视频。
其中,异常帧指的是由于运动模糊、噪声等因素导致的图像质量较差或无法与眼部图像对齐的视频帧。通过应用单应性矩阵,可以筛选掉这些异常帧,从而得到对齐的源模块眼部图像和目标模态眼部结构诊断视频。
在该实施例中,在对齐处理过程中,先将同一眼的眼部结构诊断视频中的血管造影视频帧之间进行对齐注册,然后再与眼部图像进行对齐注册,在一些实施例中,使用关键点检测器进行特征匹配,使用RANSAC(随机抽样一致性)生成单应性矩阵和异常值拒绝,进一步地,为了排除错误注册的对,还可添加有效性限制,并过滤掉注册性能较差的图像对,这可根据数据集进行经验设置。
如图3所示,在本公开的一个实施例中,将对齐处理后的源模态眼部图像作为第一初始域,将对应的目标模态关键帧作为第一目标域,对图像到图像的生成网络进行训练,得到第一模型,包括:
步骤S302,图像到图像的生成网络包括U型结构的生成器,将病变监督损失作为损失函数,采用第一初始域和第一目标域训练U型结构的生成器,以使U型结构的生成器将输入的源模态眼部图像转换为与目标模态关键帧相像的转换帧,其中,U型结构的生成器包括编码器和解码器,编码器用于将源模态眼部图像编码为潜在空间的表示,解码器用于将潜在空间的表示解码为输出数据。
步骤S304,将转换帧和对应的目标模态关键帧作为训练数据集,将感知损失和特征匹配损失作为损失函数继续训练生成器。
本领域的技术人员能够理解的是,U型结构所适应的所有生成网络,包括但不限于生成对抗网络以及扩散模型等,均在本公开的保护范围内。
其中,将对齐的第一初始域和第一目标域训练输入到预先设计的生成网络中,在一些实施例中,生成器构建有一系列堆叠的转置卷积层,以增量增强图像分辨率,通过整合跳跃连接和低级别与高级别特征,丰富了生成器,以维持细节和上下文信息,这种策略使生成器逐渐产生类似真实眼部结构诊断视频的复杂图像。
步骤S306,迭代训练生成器,直至模型将转换帧判定为关键帧的概率大于概率阈值,将训练完成的生成器作为第一模型。
在该实施例中,为了充分利用大量的血管造影视频,通过训练一个生成网络来生成所有视频帧,模型可以在训练过程中使用最小最大博弈将图像转换为不同的领域,进一步地,使用病变监督损失、感知损失和特征匹配损失等作为损失函数,进行生成器的迭代训练,得到第一模型,第一模型使用训练好的生成器生成与关键帧相似的转换帧,以保证后续将源模态的静态眼科图像转换为目标模态的动态眼科视频的可靠性。
在本公开的一个实施例中,生成器包括全局生成器网络和局部增强网络,训练生成器,还包括:
将具有第一分辨率的眼部图像和目标模态关键帧输入局部增强网络基于注意力机制进行上采样,以输出具有第二分辨率的局部增强图像,第二分辨率为第一分辨率的N倍,N大于1。
在一些实施例中,G使用了从粗到细的结构,包括全局生成器网络G1和局部增强网络G2,G1保持了输入和输出样本的分辨率,N=4,即G2输出的帧是输入的4倍,实现了更高的分辨率。
在G1和G2的上采样过程中引入了注意力机制,以将信息从粗尺度传递到细尺度,增强较浅网络层中的特定任务响应区域。
在该实施例中,通过将注意力机制纳入网络架构中,以增强信息传递,确保相关特征参与更深层次的更新,提高生成的转换帧的整体质量。
如图4所示,作为第一模型的一种实施方式,源模态眼部图像为眼底图像,对应的目标模态的眼部结构诊断视频为荧光素眼底血管造影,将眼底图像402输入第一模型中的生成器404,得到转换帧406,参考图4可知,转换帧406和眼底图像402对应的真实的荧光素眼底血管造影的关键帧408之间具有高度的相似性。
另外,在一些实施例中,生成网络为生成对抗网络,生成对抗网络包括生成器,还可以包括判别器,在一些实施方式中,采用具有相似网络结构的三个判别器(D1、D2、D3)来处理不同尺度的图像,确保了高分辨率真实图像和生成的FFA图像之间的区分,而不会过拟合。为了稳定训练过程,引入感知损失和特征匹配损失来评估每个判别器(D1、D2和D3)的每个特征提取层。
如图5所示,将生成对抗网络作为图像到图像的生成网络进行第一模型的训练,眼部图像为平面扫描图像,对应的目标模态的眼部结构诊断视频为光学相干断层扫描视频,具体处理过程包括:将平面扫描图像502输入生成器504,得到转换帧506,将转换帧506和平面扫描图像502对应的真实的光学相干断层扫描视频的关键帧508输入判别器510,输出判别结果,以基于判别结果检测第一模型的训练结果。
如图6所示,在本公开的一个实施例中,将目标模态关键帧作为第二初始域,将对齐处理后的目标模态眼部结构诊断视频作为第二目标域,对图像到视频的生成网络进行训练,得到用于对眼部动态序列进行逐帧预测的第二模型,包括:
步骤S602,将目标模态关键帧输入图像到视频的生成网络进行预测处理,得到眼部动态序列的第一帧。
其中,图像到视频的生成网络是一种深度学习模型,该生成器也包括编码器和解码器,其中,编码器将输入数据转换为潜在空间中的分布参数,解码器使用这些参数从潜在空间中采样,并将其映射回数据空间,利用图像到视频的生成网络生成网络来对眼部动态序列进行预测,并生成第一帧。
在一些实施例中,图像到视频的生成网络可以为条件变分自动编码器(Conditional Variational Autoencoder,CVAE),条件变分自动编码器是一种结合了条件信息的变分自动编码器模型,能够在生成数据时更好地控制输出结果,具有较好的生成和学习能力。
步骤S604,将第一帧和目标模态关键帧进行耦合后输入图像到视频的生成网络进行预测处理,得到生成的眼部动态序列的第二帧。
其中,通过将第一帧与关键帧耦合起来,可以更好地作为输入数据来预测眼部动态序列的第二帧。
步骤S606,基于时序逐帧预测剩余的视频帧,直至图像到视频的生成网络预测出眼部动态序列的所有视频帧。
其中,生成网络通过迭代预测来生成整个眼部动态序列的视频帧。
步骤S608,迭代训练图像到视频的生成网络,以在对应的损失函数确定生成的眼部动态序列和作为金标准的目标模态眼部结构诊断视频之间的差异小于误差阈值时,将训练完毕的图像到视频的生成网络确定为第二模型。
其中,金标准指在医学和临床研究领域被认为是最可信、最权威的标准或方法,用于确定某种疾病、诊断或治疗方法的准确性、有效性或可靠性。
在一些实施例中,通过迭代训练,将图像到视频的生成网络不断进行优化,直到其生成的眼部动态序列视频帧与真实眼部结构诊断视频的差异小于误差阈值。
另外,在第二模型的训练过程中使用的损失函数基于重建损失、平滑度约束、一致性损失和绘制中的损失生成。
在该实施例中,通过训练图像到视频的生成网络,以基于单个静态关键帧生成动态视频,得到第二模型,该过程涉及流量预测和视频帧生成,利用图像到视频的生成网络来实现眼部动态序列的预测和训练,通过迭代训练来不断优化模型,以使其生成的眼部动态序列视频帧与真实眼部结构诊断视频之间的差异最小化。
在本公开的一个实施例中,将目标模态关键帧输入图像到视频的生成网络进行预测处理,得到生成的眼部动态序列的第一帧,包括:
计算目标模态眼部结构诊断视频中相邻两个视频帧之间像素的运动矢量,作为特征标签。
其中,通过计算目标模态眼部结构诊断视频中相邻两个视频帧之间的像素运动矢量,可以得到描述视频帧间运动情况的特征标签,这些像素的运动矢量可以表征出视频帧间的运动状态,表现为血管组织对比度和动态见解,作为后续输入图像到视频的生成网络的特征标签。
将关键帧和第一帧对应的特征标签输入图像到视频的生成网络,以基于特征标签对关键帧进行扭曲预测,输出眼部动态序列的第一帧。
其中,图像到视频的生成网络对关键帧进行处理,即基于特征标签进行扭曲预测,以生成眼部动态序列的第一帧,生成网络能够利用特征标签信息来生成与关键帧匹配、符合特征标签要求的眼部动态序列的第一帧。
在一些实施例中,基于特征标签利用预测的光流来扭曲初始帧,产生初始的未来帧,后处理网络细化了帧,解决了遮挡或部分缺失等问题。
在该实施例中,使用眼部图像作为第一个帧输入生成网络,并输出第一个眼部动态序列的视频帧,然后将眼部图像与生成的第一帧耦合作为输入来生成下一个帧,依此类推,通过利用生成网络来处理眼部结构诊断视频的特征标签,并根据这些特征标签来生成眼部动态序列的第一帧。通过这种方法,可以实现根据视频帧间像素运动矢量的变化情况,来生成符合特征标签要求的眼部动态序列。
在一些实施例中,为了确保生成的眼部动态序列的平滑性,还可以进一步引入多帧输入和平滑处理方式,示例性地,通过滑动窗口的方式将作为金标准的目标模态眼部结构诊断视频的三个连续帧输入到第二模型中,将在滑动窗口中聚合生成的帧执行三帧平均,以为每个生成的帧提供更长的时间上下文,通过该处理方式,有助于使相邻视频帧之间的过渡更平滑,进而保证生成的眼部动态序列的连续性。
在本公开的一个实施例中,也可以基于颞侧一致性限制机制生成平滑的眼部动态序列。
在本公开的一个实施例中,还包括:基于颞侧一致性限制机制计算目标模态眼部结构诊断视频中的第一帧与最后一帧之间的差值;对差值进行阈值处理,得到对应的临床知识监督掩膜。
在该实施例中,通过目标模态眼部结构诊断视频中的第一帧与最后一帧之间的差值得到临床知识监督掩膜,在保证精度的同时,无需进行额外的手动注释或模型训练。
在本公开的一个实施例中,还包括:使用临床知识监督掩膜监督第一模型的训练和第二模型的训练。
在本公开的一个实施例中,在模型训练过程中,还可以应用数据增强,应用过程包括:在训练时随机选择生成的帧或作为金标准的帧作为输入,以增强其在各种场景下的适应性和鲁棒性。
在本公开的一个实施例中,数据增强包括但不限于随机裁剪、缩放和颜色增强等。
在本公开的一个实施例中,使用临床知识监督掩膜监督第一模型的训练和第二模型的训练,包括:在第一模型的训练中,基于知识增强注意力机制引导关注生成的转换帧中具有显著时间变化的第一区域;在第二模型的训练中,基于知识增强注意力机制引导关注生成的转换帧中具有显著时间变化的第二区域,以基于第一区域和/或第二区域确定关键区域;基于临床知识监督掩膜对生成的眼部动态序列进行时序颞侧的一致性和改变感知操作,其中,在改变感知操作中,基于知识感知判别器损失为关键区域提供特定监督;以及基于掩膜增强的局部归一化交叉熵损失调整源模态眼部图像与目标模态眼部结构诊断视频中的关键区域之间的像素错位。
在该实施例中,通过使用临床知识监督掩膜监督第一模型的训练和第二模型的训练,能够改善关键区域的源模态眼部图像与目标模态眼部结构诊断视频之间的像素错位,进而保证训练得到的眼部动态序列生成模型生成眼部动态序列操作的准确性和可靠性。
在一些实施例中,在模型训练过程中,还可以添加梯度引导损失来增强高频组件的生成,包括视网膜结构和病变,这一过程的输入将是眼部图像,输出将是某一阶段的眼部动态序列,可以使用PyTorch开发深度学习算法。
在本公开的一个实施例中,还包括:基于双向光流计算对眼部动态序列中的相邻帧进行前后一致性检查,以将相邻帧进行对齐。
其中,双向光流是指在计算机视觉中用来估计视频序列中像素级别的光流方向和速度。
在该实施例中,通过引入双向光流进行前后一致性检查,确保生成视频帧之间的精确对齐和一致性,以保证生成的眼部动态序列播放的流畅性。
如图7所示,在训练得到第一模型和第二模型后,第一模型用于输出关键帧,第二模型用于将源模态的静态眼科图像,即关键帧,转换为目标模态的动态眼科视频,源模态眼部图像为平面扫描图像,对应的目标模态眼部结构诊断视频为光学相干断层扫描视频,在基于平面扫描图像702得到关键帧后,将目标模态关键帧输入第二模型,得到仿真的光学相干断层扫描视频704。
如图8所示,根据本公开的一个实施例的将源模态的静态眼科图像转换为目标模态的动态眼科视频方法,包括:
步骤S802,将源模态的静态眼科图像输入眼部动态序列生成模型,眼部动态序列生成模型包括第一模型和第二模型,其中,第一模型和第二模型基于生成网络生成。
步骤S804,基于第一模型将眼部图像进行跨模态转换,生成目标模态关键帧。
步骤S806,将目标模态关键帧输入第二模型以对多个视频帧基于时序进行预测,得到多个视频帧,以基于多个视频帧生成目标模态的动态眼科视频。
在该实施例中,通过利用多种机器学习模型,通过第一模型提取眼部图像的特征并生成关键帧,再将目标模态关键帧输入第二模型进行逐帧预测和优化,最终得到完整的眼部动态序列。
在一些实施例中,眼部图像包括彩色眼底图像,眼部动态序列包括仿真的荧光素眼底血管造影,下面基于彩色眼底图像(Color Fundus Photography,CFP)生成仿真的荧光素眼底血管造影FFA的过程,对本公开中的在像素级别上匹配CFP和真实的FFA中的视网膜血管,并对生成网络进行训练,以自回归方式预测高分辨率的FFA视频的方案进行进一步具体描述。
在第一阶段,对生成模型进行预训练,以预测FFA视频的关键帧,在第二阶段,将权重转移到CFP-FFA视频对上,以预测输入CFP的FFA视频中。通过对上述方案的验证结果表明,在由三位眼科医生主观评估的xx个内部和外部测试集上,实现了逼真的生成。此外,在DR(Diabetic Retinopathy,糖尿病视网膜病变)数据集(包括EyePACS、MESSIDOR2、IDRID等)、AMD(Age-Related Macular Degeneration,年龄相关性黄斑变性)数据集以及多疾病数据集(具有许多罕见疾病和严重类别不平衡的挑战性多类视网膜图像库)上进行测试时,添加生成的眼部动态序列可以分别改善DR、AMD和多种罕见疾病的诊断准确性。虽然未来还需要进行更多研究来替代传统侵入性的血管造影检查进行临床应用,但验证结果表明,这项技术可以作为一种新型的视网膜基础模型,并可以立即实施以改进自动视网膜疾病筛查流程。
荧光素血管造影(FFA)是检测血管-视网膜屏障破坏相关病变和监测DR治疗反应的关键方法,该技术通过注射染料动态突出病变变化,特别有助于突出难以在彩色眼底照片上清晰看到的重要病变,然而,FFA是一种侵入性检查,需要静脉注射染料,并可能引起恶心、心脏病发作和过敏性休克等严重副作用。因此不适用于常规社区脉络膜视网膜条件筛查,并且只能在受过专业技术人员密切监测的情况下进行。因此,在DR和AMD患病率增加的地区,开发非侵入性、安全和低成本的FFA替代方案至关重要。
预处理阶段:将来自同一眼和同一次访问的眼底图像CF和血管造影视频用于匹配。通过使用血管分割算法从眼底图像CF和血管造影视频中提取视网膜血管,以实现像素级别的图像匹配。首先,同一眼的血管造影视频帧之间进行注册,然后再与眼底图像CF图像进行注册。使用AKAZE关键点检测器进行特征匹配,使用RANSAC(随机抽样一致性)生成单应性矩阵和异常值拒绝。为了排除错误注册的对,通过添加有效性限制,并过滤掉注册性能较差的图像对,这将根据数据集进行经验设置。
第一模型训练阶段:为了充分利用大量的血管造影视频,训练一个生成网络来生成所有视频帧,在一些实施例中,生成网络为生成器构建有一系列堆叠的转置卷积层,以增量增强图像分辨率。通过整合跳跃连接和低级别与高级别特征,丰富了生成器,以维持细节和上下文信息。
这种策略使生成器逐渐产生类似真实FFA图像的复杂图像。此外,以自回归方式生成FFA视频,其中输入是第一个帧,随后是静脉期的下一个帧,输出将是它们的下一个帧。
第二模型训练阶段:眼底图像CF到血管造影的映射利用第一阶段的权重,将使用眼底图像CF图像作为第一个帧,并输出第一个FFA的视频帧,然后将眼底图像CFP与生成的第一帧耦合作为输入来生成下一个帧,依此类推,将图像调整为768×768。
另外,添加了梯度引导损失来增强高频组件的生成,包括视网膜结构和病变。这一过程的输入将是眼底图像CF,输出将是从静脉期到晚期的FFA视频。
评价阶段:生成的仿真视频将根据以下指标与真实视频进行评估:
平均绝对误差(MAE):MAE计算生成图像与相应真实图像之间的平均绝对像素差异,量化像素值的整体差异,表示生成准确细节的水平。
峰值信噪比(PSNR):PSNR是对重建质量的人类感知的近似度量。它衡量了信号的最大可能功率与干扰其的噪声的功率之间的比率。
结构相似度测量(SSIM):SSIM评估图像之间的结构相似性,值为1表示完全相似,0表示没有相似性。SSIM提供了生成图像与真实图像之间的视觉相似性和一致性的见解。
多尺度结构相似性测量(MS-SSIM):MS-SSIM在融合不同观察条件和图像分辨率的变化方面提供了更大的灵活性。
其中,SSIM、MS-SSIM和PSNR越高,生成图像的质量就越好。
如图9所示,眼部图像为眼底图像,对应的目标模态的眼部结构诊断视频为荧光素眼底血管造影,在基于眼底图像902得到关键帧后,将目标模态关键帧输入第二模型,得到仿真的荧光素眼底血管造影904。
需要注意的是,上述附图仅是根据本发明示例性实施例的方法所包括的处理的示意性说明,而不是限制目的。易于理解,上述附图所示的处理并不表明或限制这些处理的时间顺序。另外,也易于理解,这些处理可以是例如在多个模块中同步或异步执行的。
下面参照图10来描述根据本发明的实施方式的跨模态眼部动态序列生成模型的训练装置1000。图10所示的跨模态眼部动态序列生成模型的训练装置1000仅仅是一个示例,不应对本发明实施例的功能和使用范围带来任何限制。
跨模态眼部动态序列生成模型的训练装置1000以硬件模块的形式表现。跨模态眼部动态序列生成模型的训练装置1000的组件可以包括但不限于:处理模块1002,用于将源模态眼部图像和对应的目标模态眼部结构诊断视频进行像素级的对齐处理;选取模块1004,用于选取对齐处理后的目标模态眼部结构诊断视频中的目标模态关键帧,目标模态关键帧表征目标模态眼部结构诊断视频的模态特征;第一训练模块1006,用于将对齐处理后的源模态眼部图像作为第一初始域,将对应的目标模态关键帧作为第一目标域,对图像到图像的生成网络进行训练,得到第一模型;第二训练模块1008,用于将目标模态关键帧作为第二初始域,将对齐处理后的目标模态眼部结构诊断视频作为第二目标域,对图像到视频的生成网络进行训练,得到用于对眼部动态序列进行逐帧预测的第二模型,以基于第一模型和第二模型得到眼部动态序列生成模型
下面参照图11来描述根据本发明的实施方式的将源模态的静态眼科图像转换为目标模态的动态眼科视频装置1100。图11所示的将源模态的静态眼科图像转换为目标模态的动态眼科视频装置1100仅仅是一个示例,不应对本发明实施例的功能和使用范围带来任何限制。
将源模态的静态眼科图像转换为目标模态的动态眼科视频装置1100以硬件模块的形式表现。将源模态的静态眼科图像转换为目标模态的动态眼科视频装置1100的组件可以包括但不限于:输入模块1102,用于将眼部图像输入眼部动态序列生成模型,眼部动态序列生成模型包括第一模型和第二模型,其中,第一模型和第二模型基于生成网络生成;转换模块1104,用于基于第一模型将眼部图像进行跨模态转换,生成目标模态关键帧;预测模块1106,用于将目标模态关键帧输入第二模型以对多个视频帧基于时序进行预测,得到多个视频帧,以基于多个视频帧生成目标模态的动态眼科视频。
所属技术领域的技术人员能够理解,本发明的各个方面可以实现为系统、方法或程序产品。因此,本发明的各个方面可以具体实现为以下形式,即:完全的硬件实施方式、完全的软件实施方式(包括固件、微代码等),或硬件和软件方面结合的实施方式,这里可以统称为“电路”、“模块”或“系统”。
下面参照图12来描述根据本发明的这种实施方式的电子设备1200。电子设备为眼部动态序列生成端,图12显示的电子设备1200仅仅是一个示例,不应对本发明实施例的功能和使用范围带来任何限制。
如图12所示,电子设备1200以通用计算设备的形式表现。电子设备1200的组件可以包括但不限于:上述至少一个处理单元1210、上述至少一个存储单元1220、连接不同系统组件(包括存储单元1220和处理单元1210)的总线1230。
其中,存储单元存储有程序代码,程序代码可以被处理单元1210执行,使得处理单元1210执行本说明书上述“示例性方法”部分中描述的根据本发明各种示例性实施方式的步骤。例如,处理单元1210可以执行如图2中所示的步骤S202至步骤S208所描述的方案。
存储单元1220可以包括易失性存储单元形式的可读介质,例如随机存取存储单元(RAM)12201和/或高速缓存存储单元12202,还可以进一步包括只读存储单元(ROM)12203。
存储单元1220还可以包括具有一组(至少一个)程序模块12205的程序/实用工具12204,这样的程序模块12205包括但不限于:操作系统、一个或者多个应用程序、其它程序模块以及程序数据,这些示例中的每一个或某种组合中可能包括网络环境的实现。
总线1230可以为表示几类总线结构中的一种或多种,包括存储单元总线或者存储单元控制器、外围总线、图形加速端口、处理单元或者使用多种总线结构中的任意总线结构的局域总线。
电子设备1200也可以与一个或多个外部设备1270(例如键盘、指向设备、蓝牙设备等)通信,还可与一个或者多个使得用户能与该电子设备1200交互的设备通信,和/或与使得该电子设备1200能与一个或多个其它计算设备进行通信的任何设备(例如路由器、调制解调器等等)通信。这种通信可以通过输入/输出(I/O)接口1250进行。并且,电子设备1200还可以通过网络适配器1260与一个或者多个网络(例如局域网(LAN),广域网(WAN)和/或公共网络,例如因特网)通信。如图所示,网络适配器1260通过总线1230与电子设备1200的其它模块通信。应当明白,尽管图中未示出,可以结合电子设备1200使用其它硬件和/或软件模块,包括但不限于:微代码、设备驱动器、冗余处理单元、外部磁盘驱动阵列、RAID系统、磁带驱动器以及数据备份存储系统等。
通过以上的实施方式的描述,本领域的技术人员易于理解,这里描述的示例实施方式可以通过软件实现,也可以通过软件结合必要的硬件的方式来实现。因此,根据本公开实施方式的技术方案可以以软件产品的形式体现出来,该软件产品可以存储在一个非易失性存储介质(可以是CD-ROM,U盘,移动硬盘等)中或网络上,包括若干指令以使得一台计算设备(可以是个人计算机、服务器、终端装置、或者电子设备等)执行根据本公开实施方式的方法。
在本公开的示例性实施例中,还提供了一种计算机可读存储介质,其上存储有能够实现本说明书上述方法的程序产品。在一些可能的实施方式中,本发明的各个方面还可以实现为一种程序产品的形式,其包括程序代码,当程序产品在终端设备上运行时,程序代码用于使终端设备执行本说明书上述“示例性方法”部分中描述的根据本发明各种示例性实施方式的步骤。
根据本发明的实施方式的用于实现上述方法的程序产品,其可以采用便携式紧凑盘只读存储器(CD-ROM)并包括程序代码,并可以在终端设备,例如个人电脑上运行。然而,本发明的程序产品不限于此,在本文件中,可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。
程序产品可以采用一个或多个可读介质的任意组合。可读介质可以是可读信号介质或者可读存储介质。可读存储介质例如可以为但不限于电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。可读存储介质的更具体的例子(非穷举的列表)包括:具有一个或多个导线的电连接、便携式盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。
计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了可读程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。可读信号介质还可以是可读存储介质以外的任何可读介质,该可读介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。
可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于无线、有线、光缆、RF等等,或者上述的任意合适的组合。
可以以一种或多种程序设计语言的任意组合来编写用于执行本发明操作的程序代码,程序设计语言包括面向对象的程序设计语言—诸如Python、Java、C++等,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算设备上执行、部分地在用户设备上执行、作为一个独立的软件包执行、部分在用户计算设备上部分在远程计算设备上执行、或者完全在远程计算设备或服务器上执行。在涉及远程计算设备的情形中,远程计算设备可以通过任意种类的网络,包括局域网(LAN)或广域网(WAN),连接到用户计算设备,或者,可以连接到外部计算设备(例如利用因特网服务提供商来通过因特网连接)。
应当注意,尽管在上文详细描述中提及了用于动作执行的设备的若干模块或者单元,但是这种划分并非强制性的。实际上,根据本公开的实施方式,上文描述的两个或更多模块或者单元的特征和功能可以在一个模块或者单元中具体化。反之,上文描述的一个模块或者单元的特征和功能可以进一步划分为由多个模块或者单元来具体化。
此外,尽管在附图中以特定顺序描述了本公开中方法的各个步骤,但是,这并非要求或者暗示必须按照该特定顺序来执行这些步骤,或是必须执行全部所示的步骤才能实现期望的结果。附加的或备选的,可以省略某些步骤,将多个步骤合并为一个步骤执行,以及/或者将一个步骤分解为多个步骤执行等。通过以上的实施方式的描述,本领域的技术人员易于理解,这里描述的示例实施方式可以通过软件实现,也可以通过软件结合必要的硬件的方式来实现。因此,根据本公开实施方式的技术方案可以以软件产品的形式体现出来,该软件产品可以存储在一个非易失性存储介质(可以是CD-ROM,U盘,移动硬盘等)中或网络上,包括若干指令以使得一台计算设备(可以是个人计算机、服务器、移动终端、或者电子设备等)执行根据本公开实施方式的方法。本领域技术人员在考虑说明书及实践这里公开的发明后,将容易想到本公开的其它实施方案。本申请旨在涵盖本公开的任何变型、用途或者适应性变化,这些变型、用途或者适应性变化遵循本公开的一般性原理并包括本公开未公开的本技术领域中的公知常识或惯用技术手段。说明书和实施例仅被视为示例性的,本公开的真正范围和精神由所附的权利要求指出。

Claims (19)

  1. 一种跨模态眼部动态序列生成模型的训练方法,包括:
    将源模态眼部图像和对应的目标模态眼部结构诊断视频进行像素级的对齐处理;
    选取对齐处理后的所述目标模态眼部结构诊断视频中的目标模态关键帧,所述目标模态关键帧表征所述目标模态眼部结构诊断视频的模态特征;
    将对齐处理后的所述源模态眼部图像作为第一初始域,将对应的所述目标模态关键帧作为第一目标域,对图像到图像的生成网络进行训练,得到第一模型;
    将所述目标模态关键帧作为第二初始域,将对齐处理后的所述目标模态眼部结构诊断视频作为第二目标域,对图像到视频的生成网络进行训练,得到用于对所述眼部动态序列进行逐帧预测的第二模型,以基于所述第一模型和所述第二模型得到所述眼部动态序列生成模型。
  2. 根据权利要求1所述的跨模态眼部动态序列生成模型的训练方法,其中,将源模态眼部图像和对应的目标模态眼部结构诊断视频进行像素级的对齐处理,包括:
    针对同一检查中,来自于同一眼睛的所述源模块眼部图像和所述目标模态眼部结构诊断视频,采用特征分割算法进行眼部检验特征的提取操作;
    基于所述眼部检验特征进行像素级的对齐处理。
  3. 根据权利要求2所述的跨模态眼部动态序列生成模型的训练方法,其中,基于所述眼部检验特征进行像素级的对齐处理,包括:
    所述眼部检验特征包括视网膜血管,使用关键点检测器检测所述视网膜血管的像素级关键点,将所述像素级关键点作为对齐基准进行所述源模块眼部图像和所述目标模态眼部结构诊断视频之间的特征匹配;
    基于随机抽样一致性算法对匹配特征点进行估计,估计所述源模块眼部图像和所述目标模态眼部结构诊断视频中的所述视网膜血管之间的单应性矩阵;
    基于所述单应性矩阵排除所述眼部结构诊断视频中的异常帧,得到对齐的所述源模块眼部图像和所述目标模态眼部结构诊断视频。
  4. 根据权利要求1至3中任一项所述的跨模态眼部动态序列生成模型的训练方法,其中,将对齐处理后的所述源模态眼部图像作为第一初始域,将对应的所述目标模态关键帧作为第一目标域,对图像到图像的生成网络进行训练,得到第一模型,包括:
    所述图像到图像的生成网络包括U型结构的生成器,将病变监督损失作为损失函数,采用所述第一初始域和所述第一目标域训练所述U型结构的生成器,以使所述U型结构的生成器将输入的所述源模态眼部图像转换为与所述目标模态关键帧相像的转换帧,其中,所述U型结构的生成器包括编码器和解码器,所述编码器用于将所述源模态眼部图像编码为潜在空间的表示,所述解码器用于将所述潜在空间的表示解码为所述目标模态关键帧;
    将所述转换帧和对应的所述目标模态关键帧作为训练数据集,将感知损失和特征匹配损失作为损失函数继续训练所述生成器,得到所述第一模型。
  5. 根据权利要求4所述的跨模态眼部动态序列生成模型的训练方法,其中,所述生成器包括全局生成器网络和局部增强网络,训练所述生成器,还包括:
    将具有第一分辨率的所述眼部图像和所述目标模态关键帧输入所述局部增强网络基于注意力机制进行上采样,以输出具有第二分辨率的局部增强图像,所述第二分辨率为所述第一分辨率的N倍,N大于1。
  6. 根据权利要求1至5中任一项所述的跨模态眼部动态序列生成模型的训练方法,其中,将所述目标模态关键帧作为第二初始域,将对齐处理后的所述目标模态眼部结构诊断视频作为第二目标域,对图像到视频的生成网络进行训练,得到用于对所述眼部动态序列进行逐帧预测的第二模型,包括:
    将所述关键帧输入所述图像到视频的生成网络进行预测处理,得到生成的眼部动态序列的第一帧;
    将所述第一帧和所述关键帧进行耦合后输入所述图像到视频的生成网络进行预测处理,得到所述生成的眼部动态序列的第二帧;
    基于时序逐帧预测剩余的视频帧,直至所述图像到视频的生成网络预测出所述生成的眼部动态序列的所有视频帧;
    迭代训练所述图像到视频的生成网络,以在对应的损失函数确定所述生成的眼部动态序列和作为金标准的所述目标模态眼部结构诊断视频之间的差异小于误差阈值时,将训练完毕的所述图像到视频的生成网络确定为所述第二模型。
  7. 根据权利要求6所述的跨模态眼部动态序列生成模型的训练方法,其中,将所述关键帧输入所述图像到视频的生成网络进行预测处理,得到生成的眼部动态序列的第一帧,包括:
    计算所述目标模态眼部结构诊断视频中相邻两个视频帧之间像素的运动矢量,作为特征标签;
    将所述关键帧和所述第一帧对应的特征标签输入所述图像到视频的生成网络,以基于所述特征标签对所述关键帧进行扭曲预测,输出所述眼部动态序列的第一帧。
  8. 根据权利要求6所述的跨模态眼部动态序列生成模型的训练方法,其中,还包括:
    基于双向光流计算对所述眼部动态序列中的相邻帧进行前后一致性检查,以将所述相邻帧进行对齐。
  9. 根据权利要求4或权利要求6所述的跨模态眼部动态序列生成模型的训练方法,其中,还包括:
    基于颞侧一致性限制机制计算所述目标模态眼部结构诊断视频中的第一帧与最后一帧之间的差值;
    对所述差值进行阈值处理,得到对应的临床知识监督掩膜。
  10. 根据权利要求9所述的跨模态眼部动态序列生成模型的训练方法,其中,还包括:
    使用所述临床知识监督掩膜监督所述第一模型的训练和所述第二模型的训练。
  11. 根据权利要求10所述的跨模态眼部动态序列生成模型的训练方法,其中,使用所述临床知识监督掩膜监督所述第一模型的训练和所述第二模型的训练,包括:
    在所述第一模型的训练中,基于知识增强注意力机制引导关注生成的转换帧中具有显著时间变化的第一区域;
    在所述第二模型的训练中,基于所述知识增强注意力机制引导关注生成的转换帧中具有显著时间变化的第二区域,以基于所述第一区域和/或所述第二区域确定关键区域;
    基于所述临床知识监督掩膜对生成的眼部动态序列进行时序颞侧的一致性和改变感知操作,其中,在所述改变感知操作中,基于知识感知判别器损失为所述关键区域提供特定监督;以及
    基于掩膜增强的局部归一化交叉熵损失调整所述源模态眼部图像与所述目标模态眼部结构诊断视频中的所述关键区域之间的像素错位。
  12. 根据权利要求1至8中任一项所述的跨模态眼部动态序列生成模型的训练方法,其中,
    所述目标模态的眼部结构诊断视频包括荧光素眼底血管造影或吲哚青绿血管造影,对应的所述眼部图像包括眼底图像;
    所述目标模态的眼部结构诊断视频包括光学相干断层扫描视频,对应的所述眼部图像包括眼部平面扫描图像。
  13. 根据权利要求1至8中任一项所述的跨模态眼部动态序列生成模型的训练方法,其中,
    所述生成网络包括生成对抗网络、所述生成对抗网络的衍生网络、扩散模型、所述扩散模型的衍生模型、变分自编码器、所述变分自编码器的衍生结构中的至少一种。
  14. 一种将源模态的静态眼科图像转换为目标模态的动态眼科视频的方法,包括:
    将源模态的静态眼科图像输入眼部动态序列生成模型,所述眼部动态序列生成模型包括第一模型和第二模型,其中,所述第一模型和第二模型基于生成网络生成;
    基于所述第一模型将所述眼部图像进行跨模态转换,生成目标模态关键帧;
    将所述目标模态关键帧输入所述第二模型以对多个视频帧基于时序进行预测,得到所述多个视频帧,以基于所述多个视频帧生成所述目标模态的动态眼科视频。
  15. 一种跨模态眼部动态序列生成模型的训练装置,包括:
    处理模块,用于将源模态眼部图像和对应的目标模态眼部结构诊断视频进行像素级的对齐处理;
    选取模块,用于选取对齐处理后的所述目标模态眼部结构诊断视频中的目标模态关键帧,所述目标模态关键帧表征所述目标模态眼部结构诊断视频的模态特征;
    第一训练模块,用于将对齐处理后的所述源模态眼部图像作为第一初始域,将对应的所述目标模态关键帧作为第一目标域,对图像到图像的生成网络进行训练,得到第一模型;
    第二训练模块,用于将所述目标模态关键帧作为第二初始域,将对齐处理后的所述目标模态眼部结构诊断视频作为第二目标域,对图像到视频的生成网络进行训练,得到用于对所述眼部动态序列进行逐帧预测的第二模型,以基于所述第一模型和所述第二模型得到所述眼部动态序列生成模型。
  16. 一种将源模态的静态眼科图像转换为目标模态的动态眼科视频的装置,包括:
    输入模块,用于将源模态的静态眼科图像输入眼部动态序列生成模型,所述眼部动态序列生成模型包括第一模型和第二模型,其中,所述第一模型和第二模型基于生成网络生成;
    转换模块,用于基于所述第一模型将所述眼部图像进行跨模态转换,生成目标模态关键帧;
    预测模块,用于将所述目标模态关键帧输入所述第二模型以对多个视频帧基于时序进行预测,得到所述多个视频帧,以基于所述多个视频帧生成所述目标模态的动态眼科视频。
  17. 一种电子设备,包括:
    处理器;以及
    存储器,用于存储所述处理器的可执行指令;
    其中,所述处理器配置为经由执行所述可执行指令来执行权利要求1~13中任意一项所述的跨模态眼部动态序列生成模型的训练方法或权利要求14所述的将源模态的静态眼科图像转换为目标模态的动态眼科视频的方法。
  18. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现权利要求1~13中任意一项所述的跨模态眼部动态序列生成模型的训练方法或权利要求14所述的将源模态的静态眼科图像转换为目标模态的动态眼科视频的方法。
  19. 一种计算机程序产品,其上存储有计算机程序,所述计算机程序被处理器执行时实现权利要求1~13中任意一项所述的跨模态眼部动态序列生成模型的训练方法或权利要求14所述的将源模态的静态眼科图像转换为目标模态的动态眼科视频的方法。
PCT/CN2025/081603 2024-03-27 2025-03-10 模型训练方法、转换方法、电子设备、介质和程序产品 Pending WO2025201017A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410360491.4A CN120726443A (zh) 2024-03-27 2024-03-27 模型训练方法、转换方法、电子设备、介质和程序产品
CN202410360491.4 2024-03-27

Publications (1)

Publication Number Publication Date
WO2025201017A1 true WO2025201017A1 (zh) 2025-10-02

Family

ID=97164995

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/081603 Pending WO2025201017A1 (zh) 2024-03-27 2025-03-10 模型训练方法、转换方法、电子设备、介质和程序产品

Country Status (2)

Country Link
CN (1) CN120726443A (zh)
WO (1) WO2025201017A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111901595A (zh) * 2020-06-29 2020-11-06 北京大学 一种基于深度神经网络的视频编码方法及装置、介质
KR102303626B1 (ko) * 2021-01-15 2021-09-17 정지수 단일 이미지에 기반하여 비디오 데이터를 생성하기 위한 방법 및 컴퓨팅 장치
CN114708459A (zh) * 2022-04-07 2022-07-05 国网甘肃省电力公司超高压公司 基于vae-gan的视频重构的方法、装置及存储介质
KR102472299B1 (ko) * 2021-11-03 2022-11-30 주식회사 웨이센 정지영상 데이터로부터 동영상 데이터를 생성하기 위한 의료 인공지능 모델의 학습 방법
US20240087179A1 (en) * 2022-09-09 2024-03-14 Nec Laboratories America, Inc. Video generation with latent diffusion probabilistic models

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111901595A (zh) * 2020-06-29 2020-11-06 北京大学 一种基于深度神经网络的视频编码方法及装置、介质
KR102303626B1 (ko) * 2021-01-15 2021-09-17 정지수 단일 이미지에 기반하여 비디오 데이터를 생성하기 위한 방법 및 컴퓨팅 장치
KR102472299B1 (ko) * 2021-11-03 2022-11-30 주식회사 웨이센 정지영상 데이터로부터 동영상 데이터를 생성하기 위한 의료 인공지능 모델의 학습 방법
CN114708459A (zh) * 2022-04-07 2022-07-05 国网甘肃省电力公司超高压公司 基于vae-gan的视频重构的方法、装置及存储介质
US20240087179A1 (en) * 2022-09-09 2024-03-14 Nec Laboratories America, Inc. Video generation with latent diffusion probabilistic models

Also Published As

Publication number Publication date
CN120726443A (zh) 2025-09-30

Similar Documents

Publication Publication Date Title
EP3485411B1 (en) Processing fundus images using machine learning models
CN112966792B (zh) 血管图像分类处理方法、装置、设备及存储介质
KR102709315B1 (ko) 병변 이미지 시각화 장치 및 방법
US20230260652A1 (en) Self-Supervised Machine Learning for Medical Image Analysis
EP3850638A1 (en) Processing fundus camera images using machine learning models trained using other modalities
Su et al. CAVE: Cerebral artery–vein segmentation in digital subtraction angiography
CN114170118B (zh) 基于由粗到精学习的半监督多模态核磁共振影像合成方法
KR20190136577A (ko) 심층 신경망을 이용하여 영상을 분류하는 방법 및 이를 이용한 장치
US20230263493A1 (en) Maskless 2D/3D Artificial Subtraction Angiography
CN112786163A (zh) 一种超声图像处理显示方法、系统及存储介质
CN118172562B (zh) 无造影剂心肌梗死图像分割方法、设备及介质
CN115910366A (zh) 一种基于多模态临床诊疗数据的病情分析系统
Wang et al. MEMO: dataset and methods for robust multimodal retinal image registration with large or small vessel density differences
JP7798900B2 (ja) 造影状態判別装置、造影状態判別方法、及びプログラム
WO2026045653A1 (zh) 图像处理方法、脂肪肝计算机辅助诊断方法、设备、系统、计算机存储介质及计算机程序产品
KR20240110047A (ko) 인공 지능을 이용한 직접 의학적 치료법 예측
CN115965785A (zh) 图像分割方法、装置、设备、程序产品及介质
WO2025201017A1 (zh) 模型训练方法、转换方法、电子设备、介质和程序产品
US11288800B1 (en) Attribution methodologies for neural networks designed for computer-aided diagnostic processes
CN112862752A (zh) 一种图像处理显示方法、系统电子设备及存储介质
KR20210020983A (ko) 심층 신경망을 이용하여 영상을 분류하는 방법 및 이를 이용한 장치
JP2022526126A (ja) 訓練された深層神経網モデルの再現性能を改善する方法及びそれを用いた装置
Sun et al. Deep learning for segmenting ischemic stroke infarction in non-contrast CT scans by utilizing asymmetry
Palaniappan et al. Enhancement of Medical Imaging Technique for Diabetic Retinopathy: Realistic Synthetic Image Generation Using GenAI
Guo et al. A saliency detection-inspired method for optic disc and cup segmentation

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25775597

Country of ref document: EP

Kind code of ref document: A1