WO2020197239A1 - 뉴럴 네트워크를 이용한 결측 영상 데이터 대체 방법 및 그 장치 - Google Patents

뉴럴 네트워크를 이용한 결측 영상 데이터 대체 방법 및 그 장치 Download PDF

Info

Publication number
WO2020197239A1
WO2020197239A1 PCT/KR2020/003995 KR2020003995W WO2020197239A1 WO 2020197239 A1 WO2020197239 A1 WO 2020197239A1 KR 2020003995 W KR2020003995 W KR 2020003995W WO 2020197239 A1 WO2020197239 A1 WO 2020197239A1
Authority
WO
WIPO (PCT)
Prior art keywords
image data
neural network
domains
missing
target domain
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2020/003995
Other languages
English (en)
French (fr)
Inventor
예종철
이동욱
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Korea Advanced Institute of Science and Technology KAIST
Original Assignee
Korea Advanced Institute of Science and Technology KAIST
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from KR1020190125886A external-priority patent/KR102359474B1/ko
Application filed by Korea Advanced Institute of Science and Technology KAIST filed Critical Korea Advanced Institute of Science and Technology KAIST
Publication of WO2020197239A1 publication Critical patent/WO2020197239A1/ko
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/50Image enhancement or restoration using two or more images, e.g. averaging or subtraction
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • G06N20/20Ensemble learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0475Generative networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/094Adversarial learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00Two-dimensional [2D] image generation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/60Image enhancement or restoration using machine learning, e.g. neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/77Retouching; Inpainting; Scratch removal
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10004Still image; Photographic image
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10072Tomographic images
    • G06T2207/10088Magnetic resonance imaging [MRI]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10141Special mode during image acquisition
    • G06T2207/10152Varying illumination
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20212Image combination
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person

Definitions

  • the present invention relates to a technology for replacing missing image data using a neural network, and more specifically, missing image data capable of reconstructing missing image data of a target domain using a neural network using image data of each of multiple domains as input. It relates to an alternative method and apparatus thereof.
  • missing data is sometimes replaced with a substituted value, a process called imputation.
  • the data set can be used as an input to a standard technique designed for a complete data set.
  • GAN Generative Adversarial Network
  • the general GAN framework is composed of two neural networks, generator G and discriminator D. If the discriminator finds a feature to distinguish the fake from the real sample through training, the constructor learns how to remove and synthesize the features the discriminator uses to judge the fake and the real sample. Thus, GANs can produce more realistic samples in which the distinguisher cannot distinguish between real and fake. GANs are showing remarkable achievements in a variety of computer vision tasks, such as image creation and image conversion.
  • conditional GANs control output by adding some information labels as parameters of additional generators.
  • the generator learns how to produce fake samples with specific conditions or characteristics (labels or more detailed tags associated with the image) instead of generating generic samples from unknown noise distributions.
  • a successful application of conditional GAN is for conversion between images such as pix2pix for paired data and CycleGAN for unpaired data.
  • CycleGAN J.-Y. Zhu, T. Park, P. Isola, and AA Efros. Unpaired imageto-image translation using cycle-consistent adversarial networks.arXiv preprint, 2017.) and DiscoGAN (T. Kim, M. Cha, and H. Kim, JK Lee, and J. Kim.Learning to discover cross-domain relations with generative adversarial networks.arXiv preprint arXiv:1703.05192, 2017.) use cycle coherence loss to preserve key properties between input and output images. Try to do it. However, such a framework can only learn the relationship between two different domains at once.
  • StarGAN (Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo.StarGAN: Unified generative adversarial networks for multidomain image-to-image translation.arXiv preprint, 1711, 2017.) and Radial GAN (J. Yoon, J. Jordon, and M. van der Schaar. RadialGAN: Leveraging multiple datasets to improve target-specific predictive models using generative adversarial networks.arXiv preprint arXiv:1802.06403, 2018.) It's a modern framework for handling multiple domains using constructors. For example, in StarGAN, the deep connection between the input image and the mask vector representing the target domain helps to map the input to the reconstructed image in the target domain.
  • the distinguisher should be designed to play another role for domain classification. Specifically, the distinguisher determines not only the authenticity of the sample but also the class of the sample.
  • This GAN-based image transmission technology is closely related to image data replacement because image conversion can be regarded as a process capable of estimating a missing image database by modeling an image manifold structure.
  • image conversion can be regarded as a process capable of estimating a missing image database by modeling an image manifold structure.
  • image imputation and translation there is a fundamental difference between image imputation and translation.
  • CycleGAN and StarGAN are interested in transmitting one image as another image without considering the remaining domain data set as shown in FIGS. 1A and 1B.
  • missing data does not occur frequently, and the goal is to estimate the missing data using another clean data set.
  • Embodiments of the present invention provide a method and apparatus for replacing missing image data capable of improving reconstruction performance by restoring missing image data of a target domain using a neural network using image data of each of multiple domains as input. to provide.
  • a method for replacing missing image data includes: receiving input image data for at least two or more domains among preset multiple domains; And restoring missing image data of a preset target domain by using a neural network receiving the two or more input image data as inputs.
  • the neural network combines the fake image data of the first target domain and the real image data generated by inputting at least two or more real image data of the multi-domains, and reconstructs the combined image data as an input.
  • the image and the real image data can be trained using a multi-cycle coherence loss that should be similar.
  • input image data for the at least two or more domains and information on the target domain may be received together.
  • the neural network is a multiplex including a generative adversarial network (GAN), a convolutional neural network, a convolutional framelet-based neural network, and a pooling layer and an unpooling layer. It may include at least one of resolution neural networks.
  • GAN generative adversarial network
  • convolutional neural network a convolutional neural network
  • convolutional framelet-based neural network a convolutional framelet-based neural network
  • pooling layer and an unpooling layer may include at least one of resolution neural networks.
  • the neural network may include a bypass connection from the pooling layer to the unpooling layer.
  • a method for replacing missing image data includes: receiving input image data for at least two or more domains and information on a target domain among preset multiple domains; And restoring missing image data of the target domain by using a neural network for inputting the two or more input image data and information on the target domain.
  • the neural network combines the fake image data of the first target domain and the real image data generated by inputting at least two or more real image data of the multi-domains, and reconstructs the combined image data as an input.
  • the image and the real image data can be trained using a multi-cycle coherence loss that should be similar.
  • An apparatus for replacing missing image data includes: a receiver configured to receive input image data for at least two or more domains among preset multiple domains; And a replacement unit for restoring missing image data of a preset target domain using a neural network that receives the two or more input image data as inputs.
  • the neural network combines the fake image data of the first target domain and the real image data generated by inputting at least two or more real image data of the multi-domains, and reconstructs the combined image data as an input.
  • the image and the real image data can be trained using a multi-cycle coherence loss that should be similar.
  • the receiving unit may receive input image data for the at least two or more domains and information on the target domain together.
  • the neural network is a multiplex including a generative adversarial network (GAN), a convolutional neural network, a convolutional framelet-based neural network, and a pooling layer and an unpooling layer. It may include at least one of resolution neural networks.
  • GAN generative adversarial network
  • convolutional neural network a convolutional neural network
  • convolutional framelet-based neural network a convolutional framelet-based neural network
  • pooling layer and an unpooling layer may include at least one of resolution neural networks.
  • the neural network may include a bypass connection from the pooling layer to the unpooling layer.
  • a method for replacing missing image data includes: receiving input image data for at least two or more domains among preset multi-domains; And reconstructing missing image data of a preset target domain corresponding to the two or more input image data by using a neural network learned by a predefined multi-cycle coherence loss.
  • the data acquisition method that is actually used for cancer diagnosis in the medical field can be used as it is without modification, and the missing data problem that may occur at this time can be replaced without additional cost and photography. It can significantly save both time and money costs for both sides.
  • a missing image when a missing image occurs in an image set of various contrasts necessary for cancer diagnosis, it may be used for missing data, or may be used to replace missing data in various lighting direction data sets. It may be used to replace missing data in human face data of various facial expressions.
  • the present invention provides data missing from various camera angle data, data missing from data according to the resolution of an image, data missing from data according to the degree of noise of the image, and data missing from data according to the artistic style or type of the image. It can be used universally for missing image data that occurs when various domains exist, such as data missing from data and font type data of letters.
  • FIG. 1 shows an exemplary diagram of an image conversion task according to the prior art and the present invention.
  • FIG. 2 shows an exemplary diagram for explaining a process of training a neural network in the present invention.
  • FIG. 3 is a flowchart illustrating an operation of a method for replacing missing image data according to an embodiment of the present invention.
  • FIG. 5 shows a configuration of an apparatus for replacing missing image data according to an embodiment of the present invention.
  • Embodiments of the present invention make it a gist to restore missing image data of a target domain by using a neural network that uses image data of each of multiple domains as input.
  • the present invention combines the fake image data of the target domain and the input image data generated from the input image data of the multiple domains, and the image restored from the combined image data of the multiple domains and the original input image data must be similar.
  • a learning model may be generated, and missing image data of the target domain may be restored using the neural network of the generated learning model.
  • the neural network includes a generative adversarial network (GAN), a convolutional neural network, a convolutional framelet-based neural network, a pooling layer and an unpooling layer.
  • GAN generative adversarial network
  • the multi-resolution neural network may include various types of neural networks such as U-Net, and may include all types of neural networks that can be used in the present invention.
  • the multi-resolution neural network may include a bypass connection from the pooling layer to the unpooling layer.
  • the present invention describes a Collaborative Generative Adversarial Network (CollaGAN) framework that processes multiple inputs to produce a more realistic and feasible output.
  • the CollaGAN framework of the present invention processes multiple inputs from multiple domains, compared to Star-GAN, which processes a single input and a single output, as shown in FIG. 1C.
  • the image replacement technology of the present invention provides many advantages over existing methods.
  • the basic video manifold can achieve synergy effect from multiple input data sets that share the same manifold structure rather than a single input. Therefore, the estimate of the missing value using CollaGAN is more accurate.
  • CollaGAN still maintains a first-generation architecture similar to StarGAN, which has higher memory efficiency than CycleGAN.
  • C may mean a complementary set. This mapping can be expressed as in Equation 1 below.
  • k ⁇ a,b,c,d ⁇ may mean a target domain index that guides generating an output for an appropriate target domain k.
  • the present invention allows the generator to learn various mappings for multiple target domains by randomly selecting these combinations during training.
  • the multi-cycle coherence loss is a loss in which the original input image data and the reconstructed image from the combined multi-domain image data must be similar, by combining the fake image data of the target domain and the input image data generated from the input image data of the multiple domains. It can mean.
  • the cycle coherence loss of the forward generator x ⁇ k can be expressed as ⁇ Equation 4> and ⁇ Equation 5> below.
  • Discriminator loss The discriminator plays two roles, one to classify the source whether it is real or fake, and the other to classify the domain types of classes a, b, c, and d.
  • the discriminator loss can consist of two parts. As shown in FIG. 2, the discriminator loss can be realized by using a discriminator having two paths, D gan and D clsf , which share the same neural network weights except for the last layer.
  • the present invention replaces the original GAN loss with Least Square GAN (X. Mao, Q. Li, H. Xie, RY Lau, Z. Wang, and SP Smolley. Least. squares generative adversarial networks.In Computer Vision (ICCV), 2017 IEEE International Conference on, pages 2813-2821.IEEE, 2017.)
  • the distinguisher D gan can be optimized by minimizing the loss of ⁇ Equation 6> below, and the constructor can be optimized by minimizing the loss of ⁇ Equation 7> below.
  • k may be defined through Equation 5 above.
  • the domain classification loss consists of two parts, L clsf real and L clsf fake , which may be cross entropy loss for domain classification of real and fake images, respectively.
  • the purpose of training constructor G is to generate an image that is properly classified into the target domain. Therefore, the present invention first needs the best classifier D clsf that is trained only with real data so that it can properly guide the generator. Therefore, the present invention minimizes the loss L clsf real to train the delimiter D clsf , and then trains the generator G while fixing D clsf , in order for the generator to be trained to generate an accurately classified sample, thereby making L clsf fake . Minimize it.
  • D clsf (k;x k ) may mean a probability that the real input x k can be accurately classified as class k.
  • generator G must be trained to generate a fake sample properly classified by D clsf . Therefore, it is necessary to minimize the loss expressed as in Equation 9 below for the generator G.
  • Structural Similarity Index Loss Structural Similarity Index (SSIM) is one of the most advanced indicators for measuring image quality.
  • the l 2 loss which is widely used in the image restoration task, has been reported in the prior art as causing blurring artifacts in the result.
  • SSIM is one of the perceptual metrics and can be differentiated, so it can be backpropagate.
  • the SSIM for the pixel p can be expressed as in Equation 10 below.
  • ⁇ X means the mean X
  • ⁇ 2 X means the variance of X
  • ⁇ XX * means the covariance of X and X*
  • C 1 and C 2 are variables for stabilizing the division
  • L denotes a dynamic range of pixel intensity
  • k 1 and k 2 may be 0.01 and 0.03.
  • P denotes a set of pixel positions
  • may denote the cardinality of P.
  • SSIM loss can be applied as an additional multiple cycle consistency loss as shown in Equation 12 below.
  • the mask vector is a binary matrix with the same dimension as the input image and is easily connected to the input image.
  • the mask vector has an N-class channel dimension that can represent a target domain as a one-hot vector along the channel dimension. This could be a simplified version of the mask vector introduced in the original StarGAN. That is, the mask vector may be target domain information on missing image data to be reconstructed or replaced by using multi-domain input image data input to the neural network.
  • MR contrast synthesis A total of 280 axial brain images can be scanned from 10 subjects by a multi-dynamic multi-echo sequence and an additional T2 FLAIR sequence.
  • the data set includes four MR contrast image types, e.g. T1-FLAIR(T1F), T2-weighted(T2w), T2-FLAIR(T2F), and T2-FLAIR*(T2F*). can do.
  • T1-FLAIR T1F
  • T2-weighted T2w
  • T2-FLAIR T2F
  • MR contrast of T2-FLAIR* The diagram image type may be obtained by an additional scan having a different MR scan parameter of the third contrast image type T2F. Details of the MR acquisition parameters can be found in the supplementary data.
  • CMU Multi-PIE A subset of the Carnegie Mellon University Multi-Pose Illumination and Expression Face Database can be used for the illumination conversion task.
  • the dataset can be selected in five lighting conditions: -90 degrees (right), -45 degrees, 0 degrees (front), 45 degrees and 90 degrees (left) in the front direction of the usual (neutral) facial expressions of 250 participants. .
  • the image can be cropped into a screen with a certain pixel size with the face in the center.
  • RaFD Radboud Faces Database
  • the present invention includes two networks of a generator G and a distinguisher D as shown in FIG. 2.
  • the present invention can redesign the generator and the distinguisher according to the attributes of each task.
  • the constructor is based on the U-net structure and consists of an encoder part and a decoder part, and each part between the encoder and decoder is connected by a contracting path.
  • the constructor follows the Net structure, and instead of a batch normalization layer that performs normalization operations and a ReLU (rectified linear unit) layer that performs nonlinear function operations, instance normalization.
  • Layers D. Ulyanov, A. Vedaldi, and V. Lempitsky. Instance normalization: The missing ingredient for fast stylization.arXiv preprint arXiv:1607.08022, 2016.
  • Ricky-ReLU layer K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification.In Proceedings of the IEEE international conference on computer vision, pages 1026-1034, 2015.) can be used respectively.
  • MR Contrast Conversion There are various MR contrasts, such as T1 weight contrast and T2 weight contrast.
  • the specific MR contrast scan is determined by MRI scan parameters such as repetition time (TR) and echo time (TE).
  • TR repetition time
  • TE echo time
  • the pixel intensity of an MR contrast image is determined by the physical properties of the tissue called MR parameters of the tissue, such as T1, T2, and proton density.
  • the MR parameter has a voxel-wise property. This means that in the case of a convolutional neural network, pixel-by-pixel processing is as important as processing information from neighborhood or large field of view (FOV). So, instead of using a single convolution, the constructor can use two convolution branches with 1 ⁇ 1 and 3 ⁇ 3 filters capable of handling multi-scale feature information. The two convolution branches are connected similarly to the inception network.
  • Lighting Conversion For the lighting conversion task, you can use the original U-Net structure with an instance normalization layer instead of a batch normalization layer.
  • Facial expression conversion Multiple facial images with various expressions are input for the facial expression conversion task. Since the subject's head movements exist between facial expressions, the images are not strictly aligned according to the pixel direction. If the original U-net is used for the task between facial expression images, the performance of the generator is degraded because information of various facial expressions is mixed in the initial stage of the network. For reference, features of facial expressions must be blended in the middle of the constructor, which either computes features in a large FOV or downsamples them to a pulling layer already. Thus, the constructor can be redesigned with 8 encoder branches for every 8 facial expressions and connected after the encoding process in the intermediate stage of the constructor. The structure of the decoder is similar to the decoder part of U-net, except that more convolutional layers are added using a residual block.
  • the distinguisher may generally consist of a series of convolution layers and a Leaky-ReLU layer. As shown in FIG. 2, the distinguisher has two output headers, one of which may be a real or fake classification header and the other may be a classification header for a domain.
  • the distinguisher can use PatchGAN to classify whether a local video patch is real or fake. Dropout is very effective to prevent overfitting of the distinguisher. Exceptionally, the MR contrast transform's distinguisher has a branch for multi-scale-processing.
  • the neural network in the present invention is not limited to the aforementioned neural network, and may include all kinds of networks to which the present invention can be applied.
  • the neural network of the present invention includes a generative adversarial network (GAN), a convolutional neural network, a convolutional framelet-based neural network, a pooling layer and an unpooling layer.
  • GAN generative adversarial network
  • a multi-resolution neural network including a layer may include various types of neural networks such as U-Net, and the multi-resolution neural network may include a bypass connection from a pooling layer to an unpooling layer.
  • the present invention first trains a classifier for a real image with a corresponding label for 10 epochs, and then trains a generator and a distinguisher.
  • the MR contrast conversion task, lighting conversion, and facial expression conversion task can take about 6 hours, 12 hours, and 1 day, respectively, using an NVIDIA GTX 1080 GPU.
  • the YCbCr color code may be used instead of the RGB color code for the lighting conversion task, and the YCbCr coding may consist of Y-luminance and CbCr-color.
  • the Y-luminance channel There are 5 different lighting images, 3 different lighting images almost share CbCr coding, the only difference is the Y-luminance channel. Therefore, only the Y-luminance channel is processed for the illumination conversion task, and the reconstructed image can be applied to the RGB coded image.
  • an RGB channel is used for a facial expression conversion task, and an MR contrast data set may be configured as a single channel image.
  • the method according to the embodiment of the present invention trains the discriminator network and the generator network using multi-cycle coherence loss, and when a learning model of the generator network is generated through this training process, the generated generator network, for example, CollaGAN is used.
  • the generated generator network for example, CollaGAN
  • missing image data can be replaced or restored. That is, in the method according to the embodiment of the present invention, input image data of multiple domains and information about a target domain, for example, a mask vector, are input in a neural network of a learning model generated through a training process using a multi-cycle coherence loss. After receiving, the missing image data for the target domain may be reconstructed using the learning model of the neural network.
  • the method of the present invention will be described with reference to FIG. 3 as follows.
  • FIG. 3 is a flowchart illustrating an operation of a method for replacing missing image data according to an embodiment of the present invention, and may include all the above-described contents.
  • the method for replacing missing image data receives input image data for at least two or more domains among preset multi-domains (S310 and S320).
  • step S320 the neural network for replacing the missing image data of the MR contrast image restores the missing image data for at least one of the remaining two target domains using two input image data of the four domains.
  • input image data for two domains can be received, and a neural network to replace missing image data in the MR contrast image uses the input image data of three of the four domains, and the remaining one target
  • input image data for three domains may be received.
  • step S320 may receive input image data for at least two or more domains through a training process for the corresponding input image even when it is desired to restore missing image data for an illumination image or a facial expression image.
  • the input of the neural network learned in advance through the method may be determined by a business operator or an individual providing the technology of the present invention.
  • step S320 not only input image data for two or more domains, but also information on a target domain to be reconstructed, for example, a mask vector, may be received together.
  • step S320 When input image data for at least two or more domains is received in step S320, the missing image data of the preset target domain is restored using a neural network that receives input image data for two or more domains as input. (S330, S340).
  • the neural network of step S330 receives information on the target domain as an input, and receives missing image data for the target domain based on the input image data for two or more domains and a learning model of the neural network. It can be reconstructed, and the neural network is trained using a multi-cycle coherence loss, as described above, so that a learning model can be generated.
  • the neural network in step S330 may be a generator network trained in FIG. 2, and such a neural network is, as described above, Generative Adversarial Networks (GAN), convolutional neural networks, and convolutional framelets. framelet)-based neural network, a multi-resolution neural network including a pooling layer and an unpooling layer, for example, various types of neural networks such as U-Net, and a multi-resolution neural network May include a bypass connection from the pooling layer to the unpooling layer, and may include all types of neural networks to which the present invention can be applied.
  • GAN Generative Adversarial Networks
  • U-Net unpooling layer
  • a multi-resolution neural network May include a bypass connection from the pooling layer to the unpooling layer, and may include all types of neural networks to which the present invention can be applied.
  • A, B, C, D, and 4 domains are defined as the entire image data set, and when the data D is missing, the data of the A, B, and C domains are converted into a neural network, for example, of a generator network. Use it as an input to restore the D image.
  • the generator network is trained with the goal of being identified as a real image for the discriminator network to discriminate, and the discriminator network is trained in the direction of distinguishing between fake and real images. And, the generator network conducts training in a direction to deceive the corresponding discriminator network.
  • the trained generator network is learned to provide a very realistic and realistic image, so that the missing image data of the desired target domain can be restored by inputting the multi-domain input image data.
  • the method according to an embodiment of the present invention defines the entire image data set for the entire domain, and replaces or restores the image data of the desired target domain by using the image data of the multiple domains as input to the neural network. can do.
  • the method according to the embodiment of the present invention uses a neural network to solve the problem of missing data, and it is possible to receive multiple inputs of an image for the purpose of many-to-one image conversion. Use cycle coherence loss.
  • FIG. 5 illustrates a configuration of an apparatus for replacing missing image data according to an exemplary embodiment of the present invention, and illustrates a conceptual configuration of an apparatus for performing the method of FIGS. 1 to 4.
  • an apparatus 500 includes a receiving unit 510 and a replacement unit 520.
  • the receiving unit 510 receives input image data for at least two or more domains among preset multi-domains.
  • the receiving unit 510 receives the missing image data for at least one of the remaining two target domains by using a neural network for replacing the missing image data of the MR contrast image using two input image data of the four domains.
  • a neural network to replace missing image data in the MR contrast image uses the input image data of three of the four domains to When trained to restore missing image data for three target domains, input image data for three domains may be received.
  • the receiving unit 510 may receive not only input image data for two or more domains, but also information on a target domain to be restored, for example, a mask vector.
  • the replacement unit 520 uses a neural network that receives input image data for two or more domains as inputs to store missing image data of a target domain set in advance. Restore.
  • the replacement unit 520 receives information on the target domain as input, and receives missing image data for the target domain based on input image data for two or more domains and a learning model of the neural network. Can be restored.
  • the neural network can be trained using multi-cycle coherence loss as described above, thereby generating a learning model, and the neural network is Generative Adversarial Networks (GANs), convolutional neural networks, and convoluted networks.
  • GANs Generative Adversarial Networks
  • a convolution framelet-based neural network, a multi-resolution neural network including a pooling layer and an unpooling layer may be included, and the multi-resolution neural network may include a pooling layer to an unpooling layer. It may include a bypass connection.
  • the device of FIG. 5 may include all the contents described with reference to FIGS. 1 to 4, and such matters will be apparent to those skilled in the art.
  • the apparatus described above may be implemented as a hardware component, a software component, and/or a combination of a hardware component and a software component.
  • the devices and components described in the embodiments are, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a Field Programmable Gate Array (FPGA). , A programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions, such as one or more general purpose computers or special purpose computers.
  • the processing device may execute an operating system (OS) and one or more software applications executed on the operating system.
  • OS operating system
  • the processing device may access, store, manipulate, process, and generate data in response to the execution of software.
  • the processing device is a plurality of processing elements and/or multiple types of processing elements. It can be seen that it may include.
  • the processing device may include a plurality of processors or one processor and one controller.
  • other processing configurations are possible, such as a parallel processor.
  • the software may include a computer program, code, instructions, or a combination of one or more of these, configuring the processing unit to behave as desired or processed independently or collectively. You can command the device.
  • Software and/or data may be interpreted by a processing device or to provide instructions or data to a processing device, of any type of machine, component, physical device, virtual equipment, computer storage medium or device. , Or may be permanently or temporarily embodyed in a transmitted signal wave.
  • the software may be distributed over networked computer systems and stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
  • the method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded in a computer-readable medium.
  • the computer-readable medium may include program instructions, data files, data structures, and the like alone or in combination.
  • the program instructions recorded on the medium may be specially designed and configured for the embodiment, or may be known and usable to those skilled in computer software.
  • Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, and magnetic media such as floptical disks.
  • -A hardware device specially configured to store and execute program instructions such as magneto-optical media, and ROM, RAM, flash memory, and the like.
  • Examples of the program instructions include not only machine language codes such as those produced by a compiler, but also high-level language codes that can be executed by a computer using an interpreter or the like.
  • the hardware device described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Artificial Intelligence (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • General Engineering & Computer Science (AREA)
  • Biomedical Technology (AREA)
  • Molecular Biology (AREA)
  • General Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Biophysics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Medical Informatics (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)

Abstract

뉴럴 네트워크를 이용한 결측 영상 데이터 대체 방법 및 그 장치가 개시된다. 본 발명의 일 실시예에 따른 결측 영상 데이터 대체 방법은 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신하는 단계; 및 상기 두 개 이상의 입력 영상 데이터를 입력으로 하는 뉴럴 네트워크를 이용하여 미리 설정된 타겟 도메인의 결측 영상 데이터를 복원하는 단계를 포함하며, 상기 뉴럴 네트워크는 상기 다중 도메인들 중 적어도 두 개 이상의 진짜 영상 데이터를 입력으로 하여 생성된 제1 타겟 도메인의 가짜 영상 데이터와 상기 진짜 영상 데이터를 조합하고, 상기 조합된 영상 데이터를 입력으로 하여 복원된 영상과 상기 진짜 영상 데이터가 유사해야 하는 다중 사이클 일관성 손실을 이용하여 트레이닝될 수 있다.

Description

뉴럴 네트워크를 이용한 결측 영상 데이터 대체 방법 및 그 장치
본 발명은 뉴럴 네트워크를 이용한 결측 영상 데이터 대체 기술에 관한 것으로서, 보다 구체적으로 다중 도메인들 각각의 영상 데이터를 입력으로 사용하는 뉴럴 네트워크를 이용하여 타겟 도메인의 결측 영상 데이터를 복원할 수 있는 결측 영상 데이터 대체 방법 및 그 장치에 관한 것이다.
많은 영상 처리와 컴퓨터 비전 어플리케이션에서, 원하는 출력을 생성하기 위해서는 복수의 입력 영상 셋을 필요로 한다. 예를 들어, 뇌 자기공명영상(MRI)에서는 정확한 암 마진의 진단과 세분화를 위하여 T1, T2, FLAIR(FLuid-Attenuated Inversion Recovery) 대조도(contrast)를 갖는 MR 영상들이 모두 필요하다. 다중 뷰 카메라 영상에서 3D 볼륨을 생성할 때, 대부분의 알고리즘들은 이미 정해진 화각(view angle) 셋을 요구한다. 하지만, 입력 데이터 완전한 셋은 취득 비용과 시간, 데이터 셋의 시스템적 오류 등으로 인해 얻기 어려운 경우가 많다. 예를 들면, Magnetic Resonance Image Compilation 시퀀스를 이용한 합성 MR 대조도(contrast) 생성에서, 합성 T2-FLAIR 대조도(contrast) 영상에 시스템적 오류가 존재하여 오진단으로 이어지는 경우가 많다. 또한 결측 데이터는 상당한 바이어스들을 야기할 수 있어서, 데이터 처리와 분석에 오류를 만들고 통계 효율을 감소시킬 수 있다.
임상 환경에서 종종 실현 가능하지 않은 예상치 못한 상황에서 모든 데이터 셋을 다시 획득하기 보다는, 결측 데이터(missing data)를 대체 값(substituted value)으로 대체하는 경우가 있으며, 이 프로세스를 대체(imputation)라 한다. 모든 결측 값들이 대체되면, 데이터 셋은 완전 데이터 셋을 위해 설계된 표준 기술의 입력으로 사용할 수 있다.
평균 대체(mean imputation), 회귀 대체(regression imputation), 통계적 대체(stochastic imputation) 등과 같이 전체 셋에 대한 모델링 가정에 기초하여 결측 데이터를 대체하는 몇 가지 표준 방법들이 있다. 하지만, 이러한 표준 알고리즘은 영상과 같은 고차원 데이터에 대한 한계가 있으며, 이는 영상 대체가 고차원적인 데이터 매니폴드에 대한 지식을 필요로 하기 때문이다.
영상 간(image-to-image) 변환 문제에도 유사한 기술적 문제가 있으며, 이 문제의 목표는 주어진 영상의 특정 측면을 다른 영상으로 바꾸는 것이다. 초고해상도(super resolution), 노이즈 제거작업(denoising), 블러링 제거작업(deblurring), 스타일 전송(style transfer), 의미론적 세분화(semantic segmentation), 깊이 예측(depth prediction)과 같은 태스크는 한 도메인에서 다른 도메인에 있는 해당 영상으로 영상 매핑하는 것일 수 있다. 여기서, 각 도메인은 해상도, 얼굴 표정, 빛의 각도 등 다른 측면을 가지며, 도메인 간 변환할 영상 데이터 셋의 고유한(intrinsic) 매니폴드 구조에 대해 알아야 한다. 최근에 이러한 태스크는 생성적 적대 네트워크(GAN; Generative Adversarial Network)에 의해 크게 향상되고 있다.
일반적인 GAN 프레임워크는 생성자(generator) G와 구별자(discriminator) D 두 가지 뉴럴 네트워크로 구성된다. 구별자가 트레이닝을 통하여 가짜와 진짜 샘플을 구별하기 위한 특징을 찾는다면, 생성자는 구별자가 가짜와 진짜를 판단하기 위해 사용하는 특징을 제거하고 합성하는 방법을 학습한다. 따라서, GANs는 구별자가 진짜와 가짜를 구별할 수 없는 좀 더 실제적인 샘플을 생성할 수 있다. GANs는 영상 생성, 영상 변환 등과 같은 다양한 컴퓨터 비전 작업에서 놀라운 성과를 보여주고 있다.
기존의 GAN과 달리, 조건부 GAN(Co-GAN)은 일부 정보 라벨을 추가적인 생성자의 파라미터로 더하여 출력을 제어한다. 여기서 생성자는 알려지지 않은 노이즈 분포로부터 일반적인 샘플을 생성하는 대신에 특정 조건 또는 특성(영상과 연관된 라벨 또는 보다 상세한 태그)을 가진 가짜 샘플을 생산하는 방법을 학습한다. 조건부 GAN의 성공적인 어플리케이션은 쌍을 이룬 데이터의 경우 pix2pix, 쌍을 이루지 않은 데이터의 경우 CycleGAN과 같은 영상간 변환을 위한 것이다.
CycleGAN(J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired imageto-image translation using cycle-consistent adversarial networks. arXiv preprint, 2017.)과 DiscoGAN(T. Kim, M. Cha, H. Kim, J. K. Lee, and J. Kim. Learning to discover cross-domain relations with generative adversarial networks. arXiv preprint arXiv:1703.05192, 2017.)은 사이클 일관성 손실을 이용하여 입력과 출력 영상 사이의 주요 속성을 보전하려고 한다. 그러나, 이러한 프레임워크는 한 번에 두 개의 서로 다른 도메인 사이의 관계를 학습할 수 있을 뿐이다. 이러한 접근법은 도 1a에 도시된 바와 같이 각 도메인 쌍이 별도의 도메인 쌍을 필요로 하고 N개의 구분되는 도메인을 처리하기 위해 총 N × (N-1)개의 생성자를 필요로 하기 때문에 다중 도메인을 처리할 때 확장성 한계가 있다. 종래 일 실시예 기술은 다중 도메인 번역 아이디어를 일반화하기 위하여 도 1b에 도시된 바와 같이 단일 생성자로 복수의 도메인 간 번역 매핑을 학습할 수 있는 이른바 StarGAN을 제안하였으며, 최근에 비슷한 다중 도메인 전송 네트워크가 제안된 바도 있다.
StarGAN(Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo. StarGAN: Unified generative adversarial networks for multidomain image-to-image translation. arXiv preprint, 1711, 2017.)와 Radial GAN(J. Yoon, J. Jordon, and M. van der Schaar. RadialGAN: Leveraging multiple datasets to improve target-specific predictive models using generative adversarial networks. arXiv preprint arXiv:1802.06403, 2018.)은 단일 생성자를 사용하여 여러 도메인을 처리하는 최근의 프레임워크이다. 예를 들어, StarGAN에서는 입력 영상과 타겟 도메인(target domain)을 나타내는 마스크 벡터로부터의 깊이 있는 연결은 입력을 타겟 도메인에서 재구성된 영상에 매핑하는데 도움이 된다. 여기서, 구별자는 도메인 분류를 위해 또 다른 역할을 하도록 설계되어야 한다. 구체적으로는 구별자는 샘플의 진위 여부 뿐만 아니라 샘플의 클래스도 판별한다.
이러한 GAN 기반의 영상 전송 기술은 영상 변환이 영상 매니폴드 구조를 모델링하여 결측 영상 데이터베이스를 추정할 수 있는 프로세스로 간주될 수 있으므로, 영상 데이터 대체와 밀접한 관련이 있다. 그러나 영상 대체(imputation)와 번환(translation) 사이에는 근본적인 차이점이 있다. 예를 들어, CycleGAN과 StarGAN은 도 1a와 도 1b에 도시된 바와 같이 남은 도메인 데이터 셋을 고려하지 않고 한 영상을 다른 영상으로 전송하는데 관심이 있다. 그러나 영상 대체 문제에서는 결측 데이터가 자주 발생되지 않으며 다른 클린 데이터 셋을 활용하여 결측 데이터를 추정하는 것을 목표로 한다.
본 발명의 실시예들은, 다중 도메인들 각각의 영상 데이터를 입력으로 사용하는 뉴럴 네트워크를 이용하여 타겟 도메인의 결측 영상 데이터를 복원함으로써, 복원 성능을 향상시킬 수 있는 결측 영상 데이터 대체 방법 및 그 장치를 제공한다.
본 발명의 일 실시예에 따른 결측 영상 데이터 대체 방법은 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신하는 단계; 및 상기 두 개 이상의 입력 영상 데이터를 입력으로 하는 뉴럴 네트워크를 이용하여 미리 설정된 타겟 도메인의 결측 영상 데이터를 복원하는 단계를 포함한다.
상기 뉴럴 네트워크는 상기 다중 도메인들 중 적어도 두 개 이상의 진짜 영상 데이터를 입력으로 하여 생성된 제1 타겟 도메인의 가짜 영상 데이터와 상기 진짜 영상 데이터를 조합하고, 상기 조합된 영상 데이터를 입력으로 하여 복원된 영상과 상기 진짜 영상 데이터가 유사해야 하는 다중 사이클 일관성 손실을 이용하여 트레이닝될 수 있다.
상기 수신하는 단계는 상기 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터와 상기 타겟 도메인에 대한 정보를 함께 수신할 수 있다.
상기 뉴럴 네트워크는 생성적 적대 네트워크(GAN; Generative Adversarial Networks), 컨볼루션 뉴럴 네트워크, 컨볼루션 프레임렛(convolution framelet) 기반의 뉴럴 네트워크 및 풀링(pooling) 레이어와 언풀링(unpooling) 레이어를 포함하는 다중 해상도 뉴럴 네트워크 중 적어도 하나를 포함할 수 있다.
상기 뉴럴 네트워크는 상기 풀링 레이어에서 상기 언풀링 레이어로의 바이패스 연결을 포함할 수 있다.
나아가, 본 발명의 다른 일 실시예에 따른 결측 영상 데이터 대체 방법은 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터와 타겟 도메인에 대한 정보를 수신하는 단계; 및 상기 두 개 이상의 입력 영상 데이터와 상기 타겟 도메인에 대한 정보를 입력으로 하는 뉴럴 네트워크를 이용하여 상기 타겟 도메인의 결측 영상 데이터를 복원하는 단계를 포함한다.
상기 뉴럴 네트워크는 상기 다중 도메인들 중 적어도 두 개 이상의 진짜 영상 데이터를 입력으로 하여 생성된 제1 타겟 도메인의 가짜 영상 데이터와 상기 진짜 영상 데이터를 조합하고, 상기 조합된 영상 데이터를 입력으로 하여 복원된 영상과 상기 진짜 영상 데이터가 유사해야 하는 다중 사이클 일관성 손실을 이용하여 트레이닝될 수 있다.
본 발명의 일 실시예에 따른 결측 영상 데이터 대체 장치는 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신하는 수신부; 및 상기 두 개 이상의 입력 영상 데이터를 입력으로 하는 뉴럴 네트워크를 이용하여 미리 설정된 타겟 도메인의 결측 영상 데이터를 복원하는 대체부를 포함한다.
상기 뉴럴 네트워크는 상기 다중 도메인들 중 적어도 두 개 이상의 진짜 영상 데이터를 입력으로 하여 생성된 제1 타겟 도메인의 가짜 영상 데이터와 상기 진짜 영상 데이터를 조합하고, 상기 조합된 영상 데이터를 입력으로 하여 복원된 영상과 상기 진짜 영상 데이터가 유사해야 하는 다중 사이클 일관성 손실을 이용하여 트레이닝될 수 있다.
상기 수신부는 상기 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터와 상기 타겟 도메인에 대한 정보를 함께 수신할 수 있다.
상기 뉴럴 네트워크는 생성적 적대 네트워크(GAN; Generative Adversarial Networks), 컨볼루션 뉴럴 네트워크, 컨볼루션 프레임렛(convolution framelet) 기반의 뉴럴 네트워크 및 풀링(pooling) 레이어와 언풀링(unpooling) 레이어를 포함하는 다중 해상도 뉴럴 네트워크 중 적어도 하나를 포함할 수 있다.
상기 뉴럴 네트워크는 상기 풀링 레이어에서 상기 언풀링 레이어로의 바이패스 연결을 포함할 수 있다.
본 발명의 또 다른 일 실시예에 따른 결측 영상 데이터 대체 방법은 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신하는 단계; 및 미리 정의된 다중 사이클 일관성 손실에 의해 학습된 뉴럴 네트워크를 이용하여 상기 두 개 이상의 입력 영상 데이터에 대응하는 미리 설정된 타겟 도메인의 결측 영상 데이터를 복원하는 단계를 포함한다.
본 발명의 실시예들에 따르면, 다중 도메인들 각각의 영상 데이터를 입력으로 사용하는 뉴럴 네트워크를 이용하여 타겟 도메인의 결측 영상 데이터를 복원함으로써, 복원 성능을 향상시킬 수 있다.
본 발명의 실시예들에 따르면, 현재 의료계에서 암 진단에 실제로 사용되고 있는 데이터 획득 방법을 수정하지 않고 그대로 사용하고, 이 때 발생 가능한 결측 데이터 문제를 추가적인 비용과 촬영 없이 대체할 수 있기 때문에 환자와 병원 측 모두에 시간적 비용과 금전적 비용을 획기적으로 절약할 수 있다.
본 발명의 실시예들에 따르면, 암 진단에 필요한 다양한 대조도의 영상 셋에서 결측이 발생했을 경우 결측 대체를 위해 사용될 수도 있고, 다양한 조명 방향 데이터 셋에서 결측된 데이터를 대체하기 위해 사용할 수도 있으며, 다양한 표정의 사람 얼굴 데이터에서 결측된 데이터를 대체하기 위해 사용될 수도 있다. 나아가, 본 발명은 이외에도 다양한 카메라 각도 데이터에서 결측된 데이터, 영상의 해상도에 따른 데이터에서 결측된 데이터, 영상의 노이즈 정도에 따른 데이터에서 결측된 데이터, 영상의 예술적 스타일이나 종류에 따른 데이터에서 결측된 데이터 및 글자의 폰트 타입 데이터에서 결측된 데이터 등 다양한 도메인이 존재할 때 발생하는 결측 영상 데이터에 대해 범용적으로 사용할 수 있다.
도 1은 종래 기술과 본 발명에 따른 영상 변환 태스크에 대한 일 예시도를 나타낸 것이다.
도 2는 본 발명에서 뉴럴 네트워크를 트레이닝하는 과정을 설명하기 위한 일 예시도를 나타낸 것이다.
도 3은 본 발명의 일 실시예에 따른 결측 영상 데이터 대체 방법에 대한 동작 흐름도를 나타낸 것이다.
도 4는 MR 대조도 대체 결과에 대한 일 예시도를 나타낸 것이다.
도 5는 본 발명의 일 실시예에 따른 결측 영상 데이터 대체 장치에 대한 구성을 나타낸 것이다.
본 발명의 이점 및 특징, 그리고 그것들을 달성하는 방법은 첨부되는 도면과 함께 상세하게 후술되어 있는 실시예들을 참조하면 명확해질 것이다. 그러나, 본 발명은 이하에서 개시되는 실시예들에 한정되는 것이 아니라 서로 다른 다양한 형태로 구현될 것이며, 단지 본 실시예들은 본 발명의 개시가 완전하도록 하며, 본 발명이 속하는 기술분야에서 통상의 지식을 가진 자에게 발명의 범주를 완전하게 알려주기 위해 제공되는 것이며, 본 발명은 청구항의 범주에 의해 정의될 뿐이다.
본 명세서에서 사용된 용어는 실시예들을 설명하기 위한 것이며, 본 발명을 제한하고자 하는 것은 아니다. 본 명세서에서, 단수형은 문구에서 특별히 언급하지 않는 한 복수형도 포함한다. 명세서에서 사용되는 "포함한다(comprises)" 및/또는 "포함하는(comprising)"은 언급된 구성요소, 단계, 동작 및/또는 소자는 하나 이상 의 다른 구성요소, 단계, 동작 및/또는 소자의 존재 또는 추가를 배제하지 않는다.
다른 정의가 없다면, 본 명세서에서 사용되는 모든 용어(기술 및 과학적 용어를 포함)는 본 발명이 속하는 기술분야에서 통상의 지식을 가진 자에게 공통적으로 이해될 수 있는 의미로 사용될 수 있을 것이다. 또한, 일반적으로 사용되는 사전에 정의되어 있는 용어들은 명백하게 특별히 정의되어 있지 않는 한 이상적으로 또는 과도하게 해석되지 않는다.
이하, 첨부한 도면들을 참조하여, 본 발명의 바람직한 실시예들을 보다 상세하게 설명하고자 한다. 도면 상의 동일한 구성요소에 대해서는 동일한 참조 부호를 사용하고 동일한 구성요소에 대해서 중복된 설명은 생략한다.
본 발명의 실시예들은, 다중 도메인들 각각의 영상 데이터를 입력으로 사용하는 뉴럴 네트워크를 이용하여 타겟 도메인의 결측 영상 데이터를 복원하는 것을 그 요지로 한다.
여기서, 본 발명은 다중 도메인들의 입력 영상 데이터로부터 생성된 타겟 도메인의 가짜 영상 데이터와 입력 영상 데이터를 조합하고, 조합된 다중 도메인의 영상 데이터로부터 복원된 영상과 오리지널 입력 영상 데이터가 유사해야 하는 다중 사이클 일관성 손실을 이용하여 뉴럴 네트워크를 트레이닝함으로써, 학습 모델을 생성하고, 생성된 학습 모델의 뉴럴 네트워크를 이용하여 타겟 도메인의 결측 영상 데이터를 복원할 수 있다.
본 발명에서의 뉴럴 네트워크는 생성적 적대 네트워크(GAN; Generative Adversarial Networks), 컨볼루션 뉴럴 네트워크, 컨볼루션 프레임렛(convolution framelet) 기반의 뉴럴 네트워크, 풀링(pooling) 레이어와 언풀링(unpooling) 레이어를 포함하는 다중 해상도 뉴럴 네트워크 예를 들어, U-Net 등과 같은 다양한 종류의 뉴럴 네트워크를 포함할 수 있으며, 이 뿐만 아니라 본 발명에서 사용할 수 있는 모든 종류의 뉴럴 네트워크를 포함할 수 있다. 이 때, 다중 해상도 뉴럴 네트워크는 풀링 레이어에서 언풀링 레이어로의 바이패스 연결을 포함할 수 있다.
본 발명은 보다 현실적이고 실현 가능한 출력을 생성하기 위해 다중 입력을 처리하는 공동 생성적 적대 네트워크(CollaGAN; Collaborative Generative Adversarial Network) 프레임워크에 대해 설명한다. 본 발명의 CollaGAN 프레임워크는 도 1c에 도시된 바와 같이, 단일 입력과 단일 출력을 처리하는 Star-GAN에 비해, 다중 도메인으로부터의 다중 입력을 처리한다. 본 발명의 영상 대체 기술은 기존 방법에 비해 많은 장점을 제공한다.
첫째, 기본적인 영상 매니폴드는 단일 입력보다는 동일한 매니폴드 구조를 공유하는 다중 입력 데이터 셋에서 시너지 효과를 얻을 수 있다. 따라서, CollaGAN을 이용한 결측값의 추정치는 보다 정확하다.
둘째, CollaGAN은 여전히 CycleGAN에 비해 메모리 효율이 높은 StarGAN과 유사한 1세대 아키텍처를 유지하고 있다.
이러한 본 발명에 대해 상세히 설명하면 다음과 같다.
다중 입력을 이용한 영상 대체(imputation)
설명의 편의를 위하여, a, b, c, d의 4가지 타입(N=4)의 도메인이 있다고 가정한다. 본 발명은 단일 생성자를 이용하여 다중 입력을 처리하기 위하여, 다른 타입의 다중 영상들의 셋 {x a} C={x b, x c, x d}으로부터 공동 매핑(collaborative mapping)을 통하여 생성자를 트레이닝시키고 타겟 도메인 x^ a의 출력 영상을 합성한다. 여기서, C는 상보 셋(complementary set)을 의미할 수 있다. 이 매핑은 아래 <수학식 1>과 같이 나타낼 수 있다.
[수학식 1]
Figure PCTKR2020003995-appb-img-000001
여기서, k∈{a,b,c,d}는 적절한 타겟 도메인인 k에 대한 출력을 생성하도록 가이드하는 타겟 도메인 지수를 의미할 수 있다.
복수 입력과 단일 출력 조합에 대한 조합 수가 N개이므로, 본 발명은 트레이닝 중에 이러한 조합을 무작위로 선택하여 생성자가 복수 타겟 도메인에 대한 다양한 매핑을 학습할 수 있도록 한다.
네트워크 손실
다중 사이클 일관성 손실: 본 발명의 실시예에 따른 방법의 핵심 개념 중 하나는 다중 입력에 대한 사이클 일관성이다. 입력은 복수의 영상이므로, 사이클 손실은 재정의해야 한다. 포워드 생성자 G의 출력을 x^ a라고 가정하면, 도 2에 도시된 바와 같이 생성자의 백워드 흐름(backward flow)에 대한 다른 입력으로서 N-1개의 새로운 조합들을 생성할 수 있다. 예를 들어, N = 4인 경우 아래 <수학식 2>와 같이 다중 입력과 단일 출력의 세 가지 조합이 있어, 생성자의 백워드 흐름을 이용하여 오리지널 도메인의 세가지 영상을 재구성할 수 있다.
[수학식 2]
Figure PCTKR2020003995-appb-img-000002
여기서, 연관된 다중 사이클 일관성 손실은 아래 <수학식 3>과 같이 나타낼 수 있다.
[수학식 3]
Figure PCTKR2020003995-appb-img-000003
여기서, ||·|| 1은 l 1―norm을 의미할 수 있다.
다중 사이클 일관성 손실은 다중 도메인들의 입력 영상 데이터로부터 생성된 타겟 도메인의 가짜 영상 데이터와 입력 영상 데이터를 조합하고, 조합된 다중 도메인의 영상 데이터로부터 복원된 영상과 오리지널 입력 영상 데이터가 유사해야 하는 손실을 의미할 수 있다.
일반적으로, 포워드 생성자(forward generator) x^ k의 사이클 일관성 손실은 아래 <수학식 4> 및 <수학식 5>와 같이 나타낼 수 있다.
[수학식 4]
Figure PCTKR2020003995-appb-img-000004
[수학식 5]
Figure PCTKR2020003995-appb-img-000005
구별자 손실 : 구별자는 두 가지 역할을 수행하는데, 하나는 진짜인지 가짜인지 소스를 분류하는 것이고, 다른 하나는 클래스 a, b, c, d의 도메인 타입을 분류하는 것이다. 따라서, 구별자 손실은 두 부분으로 구성될 수 있다. 도 2에 도시된 바와 같이, 구별자 손실은 마지막 레이어들을 제외하고 동일한 뉴럴 네트워크 웨이트(weights) 를 공유하는 D gan과 D clsf의 두 가지 경로를 가진 구별자를 사용하여 실현할 수 있다.
특히, 적대적 손실은 생성된 영상을 가능한 진짜로 만들기 위해 필요하다. 레귤러 GAN 손실은 학습 프로세스 중에 소멸되는 그래디언트 문제를 야기할 수 있다. 본 발명은 이러한 문제를 극복하고 트레이닝의 견고성(robustness)을 향상시키기 위해 오리지널 GAN 손실 대신 Least Square GAN(X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. P. Smolley. Least squares generative adversarial networks. In Computer Vision (ICCV), 2017 IEEE International Conference on, pages 2813-2821. IEEE, 2017.)의 적대적 손실을 활용할 수 있다. 특히, 구별자 D gan은 아래 <수학식 6>의 손실을 최소화함으로써, 최적화될 수 있고, 생성자는 아래 <수학식 7>의 손실을 최소화함으로써, 최적화될 수 있다.
[수학식 6]
Figure PCTKR2020003995-appb-img-000006
[수학식 7]
Figure PCTKR2020003995-appb-img-000007
여기서, x^ k|k는 상기 수학식 5를 통해 정의될 수 있다.
다음, 도메인 분류 손실은 L clsf real과 L clsf fake의 두 부분으로 구성되며, 그것들은 각각 진짜 영상과 가짜 영상의 도메인 분류를 위한 교차 엔트로피 손실일 수 있다. 생성자 G를 트레이닝하는 목적은 타겟 도메인으로 적절하게 분류된 영상을 생성하는 것이다. 그러므로, 본 발명은 먼저 생성자를 적절하게 가이드할 수 있도록 진짜 데이터로만 트레이닝되는 최고의 구분자(classifier) D clsf가 필요하다. 따라서, 본 발명은 구분자 D clsf를 트레이닝시키기 위해 손실 L clsf real를 최소화하며, 그리고 나서 생성자가 정확하게 분류된 샘플을 생성하도록 트레이닝되기 위하여, D clsf를 고정하면서 생성자 G를 트레이닝함으로써, L clsf fake를 최소화한다.
구체적으로 D clsf를 최적화하기 위하여, 아래 <수학식 8>과 같이 D clsf에 대하여 L clsf real를 최소화해야 한다.
[수학식 8]
Figure PCTKR2020003995-appb-img-000008
여기서, D clsf(k;x k)는 진짜 입력 x k를 클래스 k로 정확하게 분류할 수 있는 확률을 의미할 수 있다.
반면, 생성자 G는 D clsf에 의해 적절하게 분류된 가짜 샘플을 생성하도록 트레이닝되어야 한다. 따라서, 생성자 G에 대하여 아래 <수학식 9>와 같이 나타낸 손실을 최소화해야 한다.
[수학식 9]
Figure PCTKR2020003995-appb-img-000009
구조 유사도 지수 손실 : 구조 유사도 지수(SSIM; Structural Similarity Index)는 영상 품질을 측정하는 최첨단 지표 중 하나이다. 영상 복원 태스크에 널리 사용되는 l 2손실은 결과에서 블러링 아티팩트(blurring artifacts)의 원인이 되는 것으로 종래 기술에서 보고된 바 있다. SSIM은 지각적 측정기준(perceptual metrics) 중 하나이며 차별화가 가능하므로, 역전파(backpropagate)될 수 있다. 픽셀 p에 대한 SSIM은 아래 <수학식 10>과 같이 나타낼 수 있다.
[수학식 10]
Figure PCTKR2020003995-appb-img-000010
여기서, μX는 평균 X를 의미하고, σ 2 X는 X의 분산을 의미하며, σ XX *는 X와 X*의 공분산을 의미하고, C 1C 2는 분할을 안정화시키기 위한 변수들로, C 1 = ( k 1 L) 2C 2 = ( k 2 L) 2를 의미하며, L은 픽셀 강도의 동적 범위를 의미하고, k 1과 k 2는 0.01과 0.03일 수 있다.
SSIM은 0과 1 사이에 정의되므로 SSIM에 대한 손실 함수는 아래 <수학식 11>과 같이 나타낼 수 있다.
[수학식 11]
Figure PCTKR2020003995-appb-img-000011
여기서, P는 픽셀 위치 셋을 의미하고, |P|는 P의 카디널리티(cardinality)를 의미할 수 있다.
SSIM 손실은 아래 <수학식 12>와 같이 추가적인 다중 사이클 일관성 손실(multiple cycle consistency loss)로서 적용될 수 있다.
[수학식 12]
Figure PCTKR2020003995-appb-img-000012
마스크 벡터(Mask Vector)
단일 생성자를 사용하기 위하여, 생성자를 가이드할 마스크 벡터 형태로 타겟 라벨(label)을 추가해야 한다. 마스크 벡터는 입력 영상과 동일한 차원을 가진 이진 매트릭스로, 입력 영상과 쉽게 연결된다. 마스크 벡터는 채널 차원을 따라 원 핫 벡터(one-hot vector)로 타겟 도메인을 나타낼 수 있는 N 클래스의 채널 차원을 가지고 있다. 이는 오리지널 StarGAN에서 도입된 마스크 벡터의 단순화된 버전일 수 있다. 즉, 마스크 벡터는 뉴럴 네트워크로 입력되는 다중 도메인의 입력 영상 데이터를 이용하여 복원 또는 대체할 결측 영상 데이터에 대한 타겟 도메인 정보일 수 있다.
데이터 셋
MR 대조도(contrast) 합성(synthesis): 총 280 축 방향 뇌 영상이 10명의 피실험자로부터 멀티-다이나믹 멀티-에코(multi-dynamic multi-echo) 시퀀스와 추가적인 T2 FLAIR의 시퀀스에 의해 스캔될 수 있다. 데이터 셋에는 4가지의 MR 대조도(contrast) 영상 타입 예를 들어, T1-FLAIR(T1F), T2-weighted(T2w), T2-FLAIR(T2F), 그리고 T2-FLAIR*(T2F*)를 포함할 수 있다. 이 때, T1-FLAIR(T1F), T2-weighted(T2w) 및 T2-FLAIR(T2F)의 3가지의 MR 대조도 영상 타입은 MAGnetic Venocation image Compilation에서 획득될 수 있으며, T2-FLAIR*의 MR 대조도 영상 타입은 세 번째 대조도(contrast) 영상 타입(T2F) 의 다른 MR 스캔 파라미터를 가진 추가 스캔에 의해 획득될 수 있다. MR 획득 파라미터의 세부 사항은 보충 데이터에서 확인할 수 있다.
CMU Multi-PIE : 조명(illumination) 변환 태스크를 위해 카네기 멜론 대학교 Multi-Pose Illumination과 Expression Face Database의 서브셋을 사용할 수 있다. 데이터셋은 250명의 참가자의 평소(중립적) 표정의 정면 방향으로 -90도(오른쪽), -45도, 0도(정면), 45도와 90도(왼쪽)의 다섯 가지 조명 조건으로 선정될 수 있다. 영상은 얼굴이 정중앙에 위치하는 일정 픽셀 크기의 화면으로 잘라낼 수 있다.
RaFD(Radboud Faces Database): RaFD에는 67명의 참가자들로부터 수집된 8개의 다른 얼굴 표정들 예를 들어, 중립, 분노, 경멸, 혐오, 공포, 행복, 슬픔, 그리고 놀라움이 포함될 수 있다.. 또한, 세 가지 다른 시선 방향이 있으며, 따라서 총 1,608개의 영상들이 트레이닝, 유효성 검사 및 테스트 셋에 대한 피실험자에 의해 나누어 질 수 있다.
네트워크 구현
본 발명은 도 2에 도시된 바와 같이 생성자 G와 구별자 D의 2개의 네트워크를 포함한다. 각 태스크에 대해 최고의 성능을 얻기 위해, 본 발명은 각 태스크의 속성에 맞게 생성자와 구별자를 재설계할 수 있다.
생성자는 U-net 구조에 기초하며, 인코더 부분과 디코더 부분으로 구성되고, 인코더와 디코더 사이의 각 파트는 컨트랙팅 경로(contracting path)로 연결된다. 생성자는 Net 구조를 따르며, 정규화(normalization) 연산을 수행하는 배치 노말라이제이션(batch normalization) 레이어와 비선형 함수(nonlinear function) 연산을 수행하는 ReLU(rectified linear unit) 레이어 대신 인스턴스 노말라이제이션(instance normalization) 레이어(D. Ulyanov, A. Vedaldi, and V. Lempitsky. Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022, 2016.)와 리키-ReLU 레이어(K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026-1034, 2015.)가 각각 사용될 수 있다.
MR 대조도 변환: T1 웨이트 대조도(contrast), T2 웨이트 대조도(contrast) 등 다양한 MR 대조도(contrast)가 존재한다. 구체적인 MR 대조도(contrast) 스캔은 반복시간(TR; repetition time), 에코시간(TE; echo time) 등과 같은 MRI 스캔 파라미터에 의해 결정된다. MR 대조도(contrast) 영상의 픽셀 강도는 T1, T2, 양성자 밀도 등과 같이 조직의 MR 파라미터라 불리는 조직의 물리적 특성에 의해 결정된다. MR 파라미터는 복셀 방향(voxel-wise) 속성을 가진다. 이는 컨볼루션 뉴럴 네트워크의 경우, 픽셀단위 처리가 주변(neighborhood) 또는 큰 시야(FOV; Field of View)로부터 정보를 처리하는 것만큼이나 중요하다는 것을 의미한다. 따라서 단일 컨볼루션을 사용하는 대신, 생성자는 다중 스케일 특성 정보를 다룰 수 있는 1 × 1, 3 × 3 필터를 가진 두 개의 컨볼루션 분기(convolution branch)를 이용할 수 있다. 두 컨볼루션 분기는 인셉션 네트워크(inception network)와 유사하게 연결되어 있다
조명 번환: 조명 변환 태스크를 위해, 배치 노말라이제이션(batch normalization) 레이어 대신에 인스턴스 노말라이제이션(instance normalization) 레이어가 있는 오리지널의 U-Net 구조를 이용할 수 있다.
얼굴 표정 번환: 얼굴 표정 번환 태스크를 위해 다양한 표정을 가진 복수의 얼굴 영상이 입력된다. 얼굴 표정들 사이에 피실험자의 머리 움직임이 존재하기 때문에 영상이 픽셀 방향에 따라 엄격하게 정렬되지는 않는다. 얼굴 표정 영상 간 태스크에 오리지널 U-net을 사용하면, 네트워크 초기 단계에서 여러 얼굴 표정의 정보가 뒤섞여 있기 때문에 생성자의 성능이 떨어진다. 참고로 말하면 얼굴 표정의 특징은 대형 FOV에서 특징을 계산하거나 이미 풀링 레이어(pulling layer)로 다운샘플링하는 생성자의 중간 단계에서 혼합해야 한다. 따라서, 생성자는 8개의 얼굴 표정마다 8개의 인코더 분기로 재설계되어 생성자 중간 단계에서 인코딩 프로세스 후에 연결될 수 있다. 디코더의 구조는 잔여 블록(residual block)을 사용하여 더 많은 컨볼루션 레이어(convolutional layer)를 추가하는 것을 제외하고 U-net의 디코더 부분과 유사하다.
구별자는 일반적으로 일련의 컨볼루션 레이어(convolution layer)와 Leaky-ReLU 레이어로 구성될 수 있다. 도 2에 도시된 바와 같이, 구별자는 두 개의 출력 헤더를 가지고 있는데, 하나는 진짜 또는 가짜의 분류 헤더이고 다른 하나는 도메인에 대한 분류 헤더일 수 있다. 구별자는 PatchGAN를 활용하여 로컬 영상 패치가 진짜인지 가짜인지 분류할 수 있다. 드롭아웃(dropout)은 구별자의 오버피팅을 방지하기 위해 매우 효과적이다. 예외적으로, MR 대조도(contrast) 변환의 구별자는 다중 스케일 프로세싱(multi-scale-processing)를 위한 분기를 가지고 있다.
물론, 본 발명에서의 뉴럴 네트워크는 상술한 뉴럴 네트워크로 한정하지 않으며, 본 발명을 적용할 수 있는 모든 종류의 네트워크를 포함할 수 있다. 예를 들어, 본 발명의 뉴럴 네트워크는 생성적 적대 네트워크(GAN; Generative Adversarial Networks), 컨볼루션 뉴럴 네트워크, 컨볼루션 프레임렛(convolution framelet) 기반의 뉴럴 네트워크, 풀링(pooling) 레이어와 언풀링(unpooling) 레이어를 포함하는 다중 해상도 뉴럴 네트워크 예를 들어, U-Net 등과 같은 다양한 종류의 뉴럴 네트워크를 포함할 수 있으며, 다중 해상도 뉴럴 네트워크는 풀링 레이어에서 언풀링 레이어로의 바이패스 연결을 포함할 수 있다.
네트워크 트레이닝(Network training)
모든 모델은 0.00001의 학습 레이트, β1 = 0.9, β2 = 0.999를 가진 Adam 을 사용하여 최적화될 수 있다. 상술한 바와 같이, 구분자의 성능은 진짜 라벨에만 연결되어야 하며, 이는 진짜 데이터를 사용해서만 트레이닝을 받아야 함을 의미한다. 따라서, 본 발명은 먼저 10에포크(epoch) 동안 해당하는 라벨로 진짜 영상에 대한 구분자(classifier)를 트레이닝시키고, 그 후 생성자와 구별자를 트레이닝시킨다. MR 대조도(contrast) 변환 태스크, 조명 변환과 얼굴 표정 변환 태스크는 NVIDIA GTX 1080 GPU를 사용하여 각각 약 6시간, 12시간, 1일이 소요될 수 있다. 조명 변환 태스크에는 RGB 색상 코드 대신 YCbCr 색상 코드가 사용될 수 있으며, YCbCr 코딩은 Y-휘도와 CbCr-색상으로 구성될 수 있다. 5개의 다른 조명 영상들이 있으며, 3개의 다른 조명 영상들은 CbCr 코딩을 거의 공유하고 있으며 유일한 차이점은 Y-휘도 채널이다. 따라서, 조명 변환 태스크를 위해 Y-휘도 채널만 프로세싱되고, 재구성된 영상은 RGB 코딩된 영상에 적용될 수 있다. 본 발명은 얼굴 표정 변환 태스크에 RGB 채널을 사용하고, MR 대조도(contrast) 데이터 셋은 단일 채널 영상으로 구성될 수 있다.
본 발명의 실시예에 따른 방법은 구별자 네트워크와 생성자 네트워크를 다중 사이클 일관성 손실을 이용하여 트레이닝하고, 이러한 트레이닝 과정을 통해 생성자 네트워크의 학습 모델이 생성되면 생성된 생성자 네트워크 예를 들어, CollaGAN을 이용하여 결측 영상 데이터를 대체 또는 복원할 수 있다. 즉, 본 발명의 실시예에 따른 방법은 다중 사이클 일관성 손실을 이용한 트레이닝 과정을 통해 생성된 학습 모델의 뉴럴 네트워크에서 다중 도메인의 입력 영상 데이터와 타겟 도메인에 대한 정보 예를 들어, 마스크 벡터를 입력으로 수신하고, 뉴럴 네트워크의 학습 모델을 이용하여 타겟 도메인에 대한 결측 영상 데이터를 복원할 수 있다. 이러한 본 발명의 방법에 대해 도 3을 참조하여 설명하면 다음과 같다.
도 3은 본 발명의 일 실시예에 따른 결측 영상 데이터 대체 방법에 대한 동작 흐름도를 나타낸 것으로, 상술한 모든 내용을 포함할 수 있다.
도 3을 참조하면, 본 발명의 일 실시예에 따른 결측 영상 데이터 대체 방법은 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신한다(S310, S320).
여기서, 단계 S320은 MR 대조도 영상의 결측 영상 데이터를 대체하기 위한 뉴럴 네트워크가 네 개의 도메인들 중 두 개의 입력 영상 데이터를 이용하여 나머지 두 개의 타겟 도메인들 중 적어도 하나에 대한 결측 영상 데이터를 복원하도록 트레이닝된 경우 두 개의 도메인들에 대한 입력 영상 데이터를 수신할 수 있으며, MR 대조도 영상의 결측 영상 데이터를 대체하기 위한 뉴럴 네트워크가 네 개의 도메인들 중 세 개의 입력 영상 데이터를 이용하여 나머지 한 개의 타겟 도메인에 대한 결측 영상 데이터를 복원하도록 트레이닝된 경우 세 개의 도메인들에 대한 입력 영상 데이터를 수신할 수 있다. 물론, 단계 S320은 조명 영상이나 얼굴 표정 영상에 대한 결측 영상 데이터를 복원하고자 하는 경우에도 해당 입력 영상에 대한 트레이닝 과정을 통해 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신할 수 있으며, 트레이닝 과정을 통해 미리 학습된 뉴럴 네트워크의 입력에 대한 것은 본 발명의 기술을 제공하는 사업자 또는 개인에 의해 결정될 수 있다.
나아가, 단계 S320은 두 개 이상의 도메인들에 대한 입력 영상 데이터 뿐만 아니라 복원하고자 하는 타겟 도메인에 대한 정보 예를 들어, 마스크 벡터를 함께 수신할 수도 있다.
단계 S320에 의해 적어도 두 개 이상 도메인들에 대한 입력 영상 데이터가 수신되면 수신된 두 개 이상 도메인들에 대한 입력 영상 데이터를 입력으로 하는 뉴럴 네트워크를 이용하여 미리 설정된 타겟 도메인의 결측 영상 데이터를 복원한다(S330, S340).
여기서, 단계 S330의 뉴럴 네트워크는 타겟 도메인에 대한 정보를 입력으로 수신하고, 수신된 타겟 도메인에 대한 결측 영상 데이터를 입력된 두 개 이상의 도메인들에 대한 입력 영상 데이터와 뉴럴 네트워크의 학습 모델에 기초하여 복원할 수 있으며, 뉴럴 네트워크는 상술한 바와 같이, 다중 사이클 일관성 손실을 이용하여 트레이닝됨으로써, 학습 모델이 생성될 수 있다.
단계 S330에서의 뉴럴 네트워크는 도 2에서 트레이닝된 생성자 네트워크일 수 있으며, 이러한 뉴럴 네트워크는 상술한 바와 같이, 생성적 적대 네트워크(GAN; Generative Adversarial Networks), 컨볼루션 뉴럴 네트워크, 컨볼루션 프레임렛(convolution framelet) 기반의 뉴럴 네트워크, 풀링(pooling) 레이어와 언풀링(unpooling) 레이어를 포함하는 다중 해상도 뉴럴 네트워크 예를 들어, U-Net 등과 같은 다양한 종류의 뉴럴 네트워크를 포함할 수 있으며, 다중 해상도 뉴럴 네트워크는 풀링 레이어에서 언풀링 레이어로의 바이패스 연결을 포함할 수 있고, 본 발명을 적용할 수 있는 모든 종류의 뉴럴 네트워크를 포함할 수 있다.
예를 들어, 본 발명은 A, B, C, D, 4개의 도메인을 전체 영상 데이터 셋으로 정의하고, D라는 데이터가 결측되었을 때 A, B, C 도메인의 데이터를 뉴럴 네트워크 예컨대, 생성자 네트워크의 입력으로 사용하여 D 영상을 복원한다. 복원된 영상 D(fake image)의 경우 구별자 네트워크가 판별하기에 실제 영상(real image)으로 판별되는 것을 목표로 생성자 네트워크를 학습하며, 구별자 네트워크는 가짜 영상과 진짜 영상을 구별하는 방향으로 트레이닝하고, 생성자 네트워크는 해당 구별자 네트워크를 속이는 방향으로 트레이닝을 진행한다. 최종적으로 트레이닝된 생성자 네트워크는 아주 현실적이고 실제와 같은 영상을 제공하도록 학습됨으로써, 다중 도메인 입력 영상 데이터를 입력으로 하여 원하는 타겟 도메인의 결측 영상 데이터를 복원할 수 있다.
이와 같이, 본 발명의 실시예에 따른 방법은 전체 도메인에 대해 전체 영상 데이터 셋을 정의하고, 존재하는 다중 도메인의 영상 데이터들을 뉴럴 네트워크의의 입력으로 사용하여 원하는 타겟 도메인의 영상 데이터를 대체 또는 복원할 수 있다.
이러한 본 발명의 실시예에 따른 방법은 데이터의 결측 문제를 해결하기 위해서 뉴럴 네트워크를 사용하며, 다대일 영상변환을 목적으로 영상의 입력을 다중으로 받는 것이 가능하고, 이 과정에서 안정적인 트레이닝을 위해 다중 사이클 일관성 손실을 이용한다.
본 발명의 실시예에 따른 방법을 이용하여 결측 영상 데이터를 복원하게 되면 단일 입력 영상을 사용하여 복원하는 다른 알고리즘과 비교하여 훨씬 우수한 성능으로 복원이 가능하다. 예를 들어, 도 4에 도시된 바와 같이, MR 대조도 영상 데이터 셋에서 1장의 입력만 사용하는 CycleGAN과 StarGAN의 성능이 떨어지는 것을 확인할 수 있으며, 본 발명의 실시예에 따른 방법(proposed)은 성능이 우수한 것을 확인할 수 있다.
도 5는 본 발명의 일 실시예에 따른 결측 영상 데이터 대체 장치에 대한 구성을 나타낸 것으로, 도 1 내지 도 4의 방법을 수행하는 장치에 대한 개념적인 구성을 나타낸 것이다.
도 5를 참조하면, 본 발명의 실시예에 따른 장치(500)는 수신부(510) 및 대체부(520)를 포함한다.
수신부(510)는 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신한다.
여기서, 수신부(510)는 MR 대조도 영상의 결측 영상 데이터를 대체하기 위한 뉴럴 네트워크가 네 개의 도메인들 중 두 개의 입력 영상 데이터를 이용하여 나머지 두 개의 타겟 도메인들 중 적어도 하나에 대한 결측 영상 데이터를 복원하도록 트레이닝된 경우 두 개의 도메인들에 대한 입력 영상 데이터를 수신할 수 있으며, MR 대조도 영상의 결측 영상 데이터를 대체하기 위한 뉴럴 네트워크가 네 개의 도메인들 중 세 개의 입력 영상 데이터를 이용하여 나머지 한 개의 타겟 도메인에 대한 결측 영상 데이터를 복원하도록 트레이닝된 경우 세 개의 도메인들에 대한 입력 영상 데이터를 수신할 수 있다.
나아가, 수신부(510)는 두 개 이상의 도메인들에 대한 입력 영상 데이터 뿐만 아니라 복원하고자 하는 타겟 도메인에 대한 정보 예를 들어, 마스크 벡터를 함께 수신할 수도 있다.
대체부(520)는 적어도 두 개 이상 도메인들에 대한 입력 영상 데이터가 수신되면 수신된 두 개 이상 도메인들에 대한 입력 영상 데이터를 입력으로 하는 뉴럴 네트워크를 이용하여 미리 설정된 타겟 도메인의 결측 영상 데이터를 복원한다.
여기서, 대체부(520)는 타겟 도메인에 대한 정보를 입력으로 수신하고, 수신된 타겟 도메인에 대한 결측 영상 데이터를 입력된 두 개 이상의 도메인들에 대한 입력 영상 데이터와 뉴럴 네트워크의 학습 모델에 기초하여 복원할 수 있다.
이 때, 뉴럴 네트워크는 상술한 바와 같이, 다중 사이클 일관성 손실을 이용하여 트레이닝됨으로써, 학습 모델이 생성될 수 있으며, 뉴럴 네트워크는 생성적 적대 네트워크(GAN; Generative Adversarial Networks), 컨볼루션 뉴럴 네트워크, 컨볼루션 프레임렛(convolution framelet) 기반의 뉴럴 네트워크, 풀링(pooling) 레이어와 언풀링(unpooling) 레이어를 포함하는 다중 해상도 뉴럴 네트워크를 포함할 수 있고, 다중 해상도 뉴럴 네트워크는 풀링 레이어에서 언풀링 레이어로의 바이패스 연결을 포함할 수 있다.
비록, 도 5 장치에서 그 설명이 생략되었더라도, 도 5의 장치는 상기 도 1 내지 도 4에서 설명한 내용을 모두 포함할 수 있으며, 이러한 사항은 본 발명의 기술 분야에 종사하는 당업자에게 있어서 자명하다.
이상에서 설명된 장치는 하드웨어 구성요소, 소프트웨어 구성요소, 및/또는 하드웨어 구성요소 및 소프트웨어 구성요소의 조합으로 구현될 수 있다. 예를 들어, 실시예들에서 설명된 장치 및 구성요소는, 예를 들어, 프로세서, 콘트롤러, ALU(arithmetic logic unit), 디지털 신호 프로세서(digital signal processor), 마이크로컴퓨터, FPGA(Field Programmable Gate Array), PLU(programmable logic unit), 마이크로프로세서, 또는 명령(instruction)을 실행하고 응답할 수 있는 다른 어떠한 장치와 같이, 하나 이상의 범용 컴퓨터 또는 특수 목적 컴퓨터를 이용하여 구현될 수 있다. 처리 장치는 운영 체제(OS) 및 상기 운영 체제 상에서 수행되는 하나 이상의 소프트웨어 어플리케이션을 수행할 수 있다. 또한, 처리 장치는 소프트웨어의 실행에 응답하여, 데이터를 접근, 저장, 조작, 처리 및 생성할 수도 있다. 이해의 편의를 위하여, 처리 장치는 하나가 사용되는 것으로 설명된 경우도 있지만, 해당 기술분야에서 통상의 지식을 가진 자는, 처리 장치가 복수 개의 처리 요소(processing element) 및/또는 복수 유형의 처리 요소를 포함할 수 있음을 알 수 있다. 예를 들어, 처리 장치는 복수 개의 프로세서 또는 하나의 프로세서 및 하나의 콘트롤러를 포함할 수 있다. 또한, 병렬 프로세서(parallel processor)와 같은, 다른 처리 구성(processing configuration)도 가능하다.
소프트웨어는 컴퓨터 프로그램(computer program), 코드(code), 명령(instruction), 또는 이들 중 하나 이상의 조합을 포함할 수 있으며, 원하는 대로 동작하도록 처리 장치를 구성하거나 독립적으로 또는 결합적으로(collectively) 처리 장치를 명령할 수 있다. 소프트웨어 및/또는 데이터는, 처리 장치에 의하여 해석되거나 처리 장치에 명령 또는 데이터를 제공하기 위하여, 어떤 유형의 기계, 구성요소(component), 물리적 장치, 가상 장치(virtual equipment), 컴퓨터 저장 매체 또는 장치, 또는 전송되는 신호 파(signal wave)에 영구적으로, 또는 일시적으로 구체화(embody)될 수 있다. 소프트웨어는 네트워크로 연결된 컴퓨터 시스템 상에 분산되어서, 분산된 방법으로 저장되거나 실행될 수도 있다. 소프트웨어 및 데이터는 하나 이상의 컴퓨터 판독 가능 기록 매체에 저장될 수 있다.
실시예에 따른 방법은 다양한 컴퓨터 수단을 통하여 수행될 수 있는 프로그램 명령 형태로 구현되어 컴퓨터 판독 가능 매체에 기록될 수 있다. 상기 컴퓨터 판독 가능 매체는 프로그램 명령, 데이터 파일, 데이터 구조 등을 단독으로 또는 조합하여 포함할 수 있다. 상기 매체에 기록되는 프로그램 명령은 실시예를 위하여 특별히 설계되고 구성된 것들이거나 컴퓨터 소프트웨어 당업자에게 공지되어 사용 가능한 것일 수도 있다. 컴퓨터 판독 가능 기록 매체의 예에는 하드 디스크, 플로피 디스크 및 자기 테이프와 같은 자기 매체(magnetic media), CD-ROM, DVD와 같은 광기록 매체(optical media), 플롭티컬 디스크(floptical disk)와 같은 자기-광 매체(magneto-optical media), 및 롬(ROM), 램(RAM), 플래시 메모리 등과 같은 프로그램 명령을 저장하고 수행하도록 특별히 구성된 하드웨어 장치가 포함된다. 프로그램 명령의 예에는 컴파일러에 의해 만들어지는 것과 같은 기계어 코드뿐만 아니라 인터프리터 등을 사용해서 컴퓨터에 의해서 실행될 수 있는 고급 언어 코드를 포함한다. 상기된 하드웨어 장치는 실시예의 동작을 수행하기 위해 하나 이상의 소프트웨어 모듈로서 작동하도록 구성될 수 있으며, 그 역도 마찬가지이다.
이상과 같이 실시예들이 비록 한정된 실시예와 도면에 의해 설명되었으나, 해당 기술분야에서 통상의 지식을 가진 자라면 상기의 기재로부터 다양한 수정 및 변형이 가능하다. 예를 들어, 설명된 기술들이 설명된 방법과 다른 순서로 수행되거나, 및/또는 설명된 시스템, 구조, 장치, 회로 등의 구성요소들이 설명된 방법과 다른 형태로 결합 또는 조합되거나, 다른 구성요소 또는 균등물에 의하여 대치되거나 치환되더라도 적절한 결과가 달성될 수 있다.
그러므로, 다른 구현들, 다른 실시예들 및 특허청구범위와 균등한 것들도 후술하는 특허청구범위의 범위에 속한다.

Claims (13)

  1. 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신하는 단계; 및
    상기 두 개 이상의 입력 영상 데이터를 입력으로 하는 뉴럴 네트워크를 이용하여 미리 설정된 타겟 도메인의 결측 영상 데이터를 복원하는 단계
    를 포함하는 결측 영상 데이터 대체 방법.
  2. 제1항에 있어서,
    상기 뉴럴 네트워크는
    상기 다중 도메인들 중 적어도 두 개 이상의 진짜 영상 데이터를 입력으로 하여 생성된 제1 타겟 도메인의 가짜 영상 데이터와 상기 진짜 영상 데이터를 조합하고, 상기 조합된 영상 데이터를 입력으로 하여 복원된 영상과 상기 진짜 영상 데이터가 유사해야 하는 다중 사이클 일관성 손실을 이용하여 트레이닝되는 것을 특징으로 하는 결측 영상 데이터 대체 방법.
  3. 제1항에 있어서,
    상기 수신하는 단계는
    상기 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터와 상기 타겟 도메인에 대한 정보를 함께 수신하는 것을 특징으로 하는 결측 영상 데이터 대체 방법.
  4. 제1항에 있어서,
    상기 뉴럴 네트워크는
    생성적 적대 네트워크(GAN; Generative Adversarial Networks), 컨볼루션 뉴럴 네트워크, 컨볼루션 프레임렛(convolution framelet) 기반의 뉴럴 네트워크 및 풀링(pooling) 레이어와 언풀링(unpooling) 레이어를 포함하는 다중 해상도 뉴럴 네트워크 중 적어도 하나를 포함하는 것을 특징으로 하는 결측 영상 데이터 대체 방법.
  5. 제4항에 있어서,
    상기 뉴럴 네트워크는
    상기 풀링 레이어에서 상기 언풀링 레이어로의 바이패스 연결을 포함하는 것을 특징으로 하는 결측 영상 데이터 대체 방법.
  6. 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터와 타겟 도메인에 대한 정보를 수신하는 단계; 및
    상기 두 개 이상의 입력 영상 데이터와 상기 타겟 도메인에 대한 정보를 입력으로 하는 뉴럴 네트워크를 이용하여 상기 타겟 도메인의 결측 영상 데이터를 복원하는 단계
    를 포함하는 결측 영상 데이터 대체 방법.
  7. 제6항에 있어서,
    상기 뉴럴 네트워크는
    상기 다중 도메인들 중 적어도 두 개 이상의 진짜 영상 데이터를 입력으로 하여 생성된 제1 타겟 도메인의 가짜 영상 데이터와 상기 진짜 영상 데이터를 조합하고, 상기 조합된 영상 데이터를 입력으로 하여 복원된 영상과 상기 진짜 영상 데이터가 유사해야 하는 다중 사이클 일관성 손실을 이용하여 트레이닝되는 것을 특징으로 하는 결측 영상 데이터 대체 방법.
  8. 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신하는 수신부; 및
    상기 두 개 이상의 입력 영상 데이터를 입력으로 하는 뉴럴 네트워크를 이용하여 미리 설정된 타겟 도메인의 결측 영상 데이터를 복원하는 대체부
    를 포함하는 결측 영상 데이터 대체 장치.
  9. 제8항에 있어서,
    상기 뉴럴 네트워크는
    상기 다중 도메인들 중 적어도 두 개 이상의 진짜 영상 데이터를 입력으로 하여 생성된 제1 타겟 도메인의 가짜 영상 데이터와 상기 진짜 영상 데이터를 조합하고, 상기 조합된 영상 데이터를 입력으로 하여 복원된 영상과 상기 진짜 영상 데이터가 유사해야 하는 다중 사이클 일관성 손실을 이용하여 트레이닝되는 것을 특징으로 하는 결측 영상 데이터 대체 장치.
  10. 제8항에 있어서,
    상기 수신부는
    상기 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터와 상기 타겟 도메인에 대한 정보를 함께 수신하는 것을 특징으로 하는 결측 영상 데이터 대체 장치.
  11. 제8항에 있어서,
    상기 뉴럴 네트워크는
    생성적 적대 네트워크(GAN; Generative Adversarial Networks), 컨볼루션 뉴럴 네트워크, 컨볼루션 프레임렛(convolution framelet) 기반의 뉴럴 네트워크 및 풀링(pooling) 레이어와 언풀링(unpooling) 레이어를 포함하는 다중 해상도 뉴럴 네트워크 중 적어도 하나를 포함하는 것을 특징으로 하는 결측 영상 데이터 대체 장치.
  12. 제11항에 있어서,
    상기 뉴럴 네트워크는
    상기 풀링 레이어에서 상기 언풀링 레이어로의 바이패스 연결을 포함하는 것을 특징으로 하는 결측 영상 데이터 대체 장치.
  13. 미리 설정된 다중 도메인들 중 적어도 두 개 이상의 도메인들에 대한 입력 영상 데이터를 수신하는 단계; 및
    미리 정의된 다중 사이클 일관성 손실에 의해 학습된 뉴럴 네트워크를 이용하여 상기 두 개 이상의 입력 영상 데이터에 대응하는 미리 설정된 타겟 도메인의 결측 영상 데이터를 복원하는 단계
    를 포함하는 결측 영상 데이터 대체 방법.
PCT/KR2020/003995 2019-03-25 2020-03-24 뉴럴 네트워크를 이용한 결측 영상 데이터 대체 방법 및 그 장치 Ceased WO2020197239A1 (ko)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
KR10-2019-0033826 2019-03-25
KR20190033826 2019-03-25
KR1020190125886A KR102359474B1 (ko) 2019-03-25 2019-10-11 뉴럴 네트워크를 이용한 결측 영상 데이터 대체 방법 및 그 장치
KR10-2019-0125886 2019-10-11

Publications (1)

Publication Number Publication Date
WO2020197239A1 true WO2020197239A1 (ko) 2020-10-01

Family

ID=69953911

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2020/003995 Ceased WO2020197239A1 (ko) 2019-03-25 2020-03-24 뉴럴 네트워크를 이용한 결측 영상 데이터 대체 방법 및 그 장치

Country Status (2)

Country Link
US (1) US11748851B2 (ko)
WO (1) WO2020197239A1 (ko)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112926534A (zh) * 2021-04-02 2021-06-08 北京理工大学重庆创新中心 一种基于变换域信息融合的sar图形船只目标检测方法
CN113177448A (zh) * 2021-04-19 2021-07-27 西安交通大学 数模联合驱动的轴承混合工况无监督域适应诊断方法及系统
CN115409755A (zh) * 2022-11-03 2022-11-29 腾讯科技(深圳)有限公司 贴图处理方法和装置、存储介质及电子设备

Families Citing this family (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111833238B (zh) * 2020-06-01 2023-07-25 北京百度网讯科技有限公司 图像的翻译方法和装置、图像翻译模型的训练方法和装置
DE102020211475A1 (de) * 2020-09-14 2022-03-17 Robert Bosch Gesellschaft mit beschränkter Haftung Kaskadierte Cluster-Generator-Netze zum Erzeugen synthetischer Bilder
US12003719B2 (en) 2020-11-26 2024-06-04 Electronics And Telecommunications Research Institute Method, apparatus and storage medium for image encoding/decoding using segmentation map
CN112767519B (zh) * 2020-12-30 2022-04-19 电子科技大学 结合风格迁移的可控表情生成方法
CN112633234B (zh) * 2020-12-30 2024-08-23 广州华多网络科技有限公司 人脸去眼镜模型训练、应用方法及其装置、设备和介质
KR102593489B1 (ko) * 2021-04-29 2023-10-24 주식회사 딥브레인에이아이 기계 학습을 이용한 데이터 생성 방법 및 이를 수행하기 위한 컴퓨팅 장치
CN113239782B (zh) * 2021-05-11 2023-04-28 广西科学院 一种融合多尺度gan和标签学习的行人重识别系统及方法
CN113689348B (zh) * 2021-08-18 2023-12-26 中国科学院自动化研究所 多任务图像复原方法、系统、电子设备及存储介质
CN116188255A (zh) * 2021-11-25 2023-05-30 北京字跳网络技术有限公司 基于gan网络的超分图像处理方法、装置、设备及介质
US12430725B2 (en) * 2022-05-13 2025-09-30 Adobe Inc. Object class inpainting in digital images utilizing class-specific inpainting neural networks
CN114972082A (zh) * 2022-05-13 2022-08-30 天津大学 一种对高比例负荷缺失数据的恢复与评估方法
CN115983495B (zh) * 2023-02-20 2023-08-11 哈尔滨工业大学(深圳)(哈尔滨工业大学深圳科技创新研究院) 基于RFR-Net的全球中性大气温度密度预测方法及设备
CN116452908A (zh) * 2023-03-03 2023-07-18 中国人民解放军陆军工程大学 大气电场监测站点丢失数据还原方法
CN119904893A (zh) * 2025-04-01 2025-04-29 杭州杰峰软件有限公司 一种鸟类识别方法、装置及电子设备、存储介质

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20140207493A1 (en) * 2011-08-26 2014-07-24 The Regents Of The University Of California Systems and methods for missing data imputation
US9324022B2 (en) * 2014-03-04 2016-04-26 Signal/Sense, Inc. Classifying data with deep learning neural records incrementally refined through expert input
WO2017223560A1 (en) * 2016-06-24 2017-12-28 Rensselaer Polytechnic Institute Tomographic image reconstruction via machine learning
US20170372193A1 (en) * 2016-06-23 2017-12-28 Siemens Healthcare Gmbh Image Correction Using A Deep Generative Machine-Learning Model

Family Cites Families (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9633282B2 (en) * 2015-07-30 2017-04-25 Xerox Corporation Cross-trained convolutional neural networks using multimodal images
US20180247201A1 (en) 2017-02-28 2018-08-30 Nvidia Corporation Systems and methods for image-to-image translation using variational autoencoders
US10970829B2 (en) * 2017-08-24 2021-04-06 Siemens Healthcare Gmbh Synthesizing and segmenting cross-domain medical images
KR102089151B1 (ko) 2017-08-30 2020-03-13 한국과학기술원 확장된 뉴럴 네트워크를 이용한 영상 복원 방법 및 장치
US10825219B2 (en) * 2018-03-22 2020-11-03 Northeastern University Segmentation guided image generation with adversarial networks
US10572770B2 (en) * 2018-06-15 2020-02-25 Intel Corporation Tangent convolution for 3D data
US11556581B2 (en) * 2018-09-04 2023-01-17 Inception Institute of Artificial Intelligence, Ltd. Sketch-based image retrieval techniques using generative domain migration hashing
US11158069B2 (en) * 2018-12-11 2021-10-26 Siemens Healthcare Gmbh Unsupervised deformable registration for multi-modal images
US11580869B2 (en) * 2019-09-23 2023-02-14 Revealit Corporation Computer-implemented interfaces for identifying and revealing selected objects from video
US11551652B1 (en) * 2019-11-27 2023-01-10 Amazon Technologies, Inc. Hands-on artificial intelligence education service
US11610599B2 (en) * 2019-12-06 2023-03-21 Meta Platforms Technologies, Llc Systems and methods for visually guided audio separation

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20140207493A1 (en) * 2011-08-26 2014-07-24 The Regents Of The University Of California Systems and methods for missing data imputation
US9324022B2 (en) * 2014-03-04 2016-04-26 Signal/Sense, Inc. Classifying data with deep learning neural records incrementally refined through expert input
US20170372193A1 (en) * 2016-06-23 2017-12-28 Siemens Healthcare Gmbh Image Correction Using A Deep Generative Machine-Learning Model
WO2017223560A1 (en) * 2016-06-24 2017-12-28 Rensselaer Polytechnic Institute Tomographic image reconstruction via machine learning

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
LEE, DONGWOOK ET AL.: "CollaGAN: Collaborative GAN for Missing Image Data Imputation", COMPUTER VISION AND PATTERN RECOGNITION (CS.CV, vol. 2, 26 February 2019 (2019-02-26), pages 7 - 16, XP033686873 *

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112926534A (zh) * 2021-04-02 2021-06-08 北京理工大学重庆创新中心 一种基于变换域信息融合的sar图形船只目标检测方法
CN112926534B (zh) * 2021-04-02 2023-04-28 北京理工大学重庆创新中心 一种基于变换域信息融合的sar图形船只目标检测方法
CN113177448A (zh) * 2021-04-19 2021-07-27 西安交通大学 数模联合驱动的轴承混合工况无监督域适应诊断方法及系统
CN113177448B (zh) * 2021-04-19 2023-06-13 西安交通大学 数模联合驱动的轴承混合工况无监督域适应诊断方法及系统
CN115409755A (zh) * 2022-11-03 2022-11-29 腾讯科技(深圳)有限公司 贴图处理方法和装置、存储介质及电子设备

Also Published As

Publication number Publication date
US11748851B2 (en) 2023-09-05
US20200311874A1 (en) 2020-10-01

Similar Documents

Publication Publication Date Title
KR102359474B1 (ko) 뉴럴 네트워크를 이용한 결측 영상 데이터 대체 방법 및 그 장치
US11748851B2 (en) Method of replacing missing image data by using neural network and apparatus thereof
WO2021054706A1 (en) Teaching gan (generative adversarial networks) to generate per-pixel annotation
WO2021118270A1 (en) Method and electronic device for deblurring blurred image
WO2020076135A1 (ko) 암 영역에 대한 딥러닝 모델 학습 장치 및 방법
WO2021080145A1 (ko) 이미지 채움 장치 및 방법
WO2020032420A1 (en) Method for training and testing data embedding network to generate marked data by integrating original data with mark data, and training device and testing device using the same
Wang et al. Jpeg compression-aware image forgery localization
Peng et al. Raune-net: a residual and attention-driven underwater image enhancement method
WO2022086147A1 (en) Method for training and testing user learning network to be used for recognizing obfuscated data created by obfuscating original data to protect personal information and user learning device and testing device using the same
WO2021010671A9 (ko) 뉴럴 네트워크 및 비국소적 블록을 이용하여 세그멘테이션을 수행하는 질병 진단 시스템 및 방법
WO2021241994A1 (ko) Rgb-d 카메라의 트래킹을 통한 3d 모델의 생성 방법 및 장치
WO2024162581A1 (ko) 개선된 적대적 어텐션 네트워크 시스템 및 이를 이용한 이미지 생성 방법
WO2025095396A1 (ko) 병리 이미지 분석을 위한 딥러닝 모델의 학습 방법 및 이를 수행하는 컴퓨팅 시스템
WO2024111915A1 (ko) 화질 전환을 이용한 인공지능에 의한 의료영상 변환방법 및 그 장치
WO2021125521A1 (ko) 순차적 특징 데이터 이용한 행동 인식 방법 및 그를 위한 장치
WO2023224350A2 (ko) 3차원 볼륨 영상으로부터 랜드마크를 검출하기 위한 방법 및 장치
WO2021150016A1 (en) Methods and systems for performing tasks on media using attribute specific joint learning
Chen et al. Improved U-Net3+ with spatial–spectral transformer for multispectral image reconstruction
WO2023204342A1 (ko) 딥러닝 기반 이미지 변환 방법 및 그 시스템
WO2020235861A1 (ko) 집중 레이어를 포함하는 생성기를 기반으로 예측 이미지를 생성하는 장치 및 그 제어 방법
WO2023204449A1 (en) Learning method and learning device for training obfuscation network capable of obfuscating original data for privacy to achieve information restriction obfuscation and testing method and testing device using the same
WO2023146141A1 (ko) 인공 지능을 이용하여 의료 영상에서 병변을 판단하는 방법, 이를 수행하는 인공 지능 신경망 시스템 및 이를 컴퓨터에서 실행시키기 위한 프로그램이 기록된 컴퓨터로 읽을 수 있는 기록 매체
Fu et al. PiD: Generalized AI-Generated Images Detection with Pixelwise Decomposition Residuals
Yang et al. An error-activation-guided blind metric for stitched panoramic image quality assessment

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20777191

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20777191

Country of ref document: EP

Kind code of ref document: A1