WO2020175446A1 - 学習方法、学習システム、学習済みモデル、プログラム及び超解像画像生成装置 - Google Patents

学習方法、学習システム、学習済みモデル、プログラム及び超解像画像生成装置 Download PDF

Info

Publication number
WO2020175446A1
WO2020175446A1 PCT/JP2020/007383 JP2020007383W WO2020175446A1 WO 2020175446 A1 WO2020175446 A1 WO 2020175446A1 JP 2020007383 W JP2020007383 W JP 2020007383W WO 2020175446 A1 WO2020175446 A1 WO 2020175446A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
learning
resolution
generator
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2020/007383
Other languages
English (en)
French (fr)
Inventor
工藤 彰
嘉郎 北村
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fujifilm Corp
Original Assignee
Fujifilm Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fujifilm Corp filed Critical Fujifilm Corp
Priority to EP20762514.6A priority Critical patent/EP3932318B1/en
Priority to JP2021502251A priority patent/JP7105363B2/ja
Publication of WO2020175446A1 publication Critical patent/WO2020175446A1/ja
Priority to US17/400,142 priority patent/US12217387B2/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T3/00Geometric image transformations in the plane of the image
    • G06T3/40Scaling of whole images or parts thereof, e.g. expanding or contracting
    • G06T3/4046Scaling of whole images or parts thereof, e.g. expanding or contracting using neural networks
    • AHUMAN NECESSITIES
    • A61MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61BDIAGNOSIS; SURGERY; IDENTIFICATION
    • A61B5/00Measuring for diagnostic purposes; Identification of persons
    • AHUMAN NECESSITIES
    • A61MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61BDIAGNOSIS; SURGERY; IDENTIFICATION
    • A61B6/00Apparatus or devices for radiation diagnosis; Apparatus or devices for radiation diagnosis combined with radiation therapy equipment
    • A61B6/02Arrangements for diagnosis sequentially in different planes; Stereoscopic radiation diagnosis
    • A61B6/03Computed tomography [CT]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • G06N3/0455Auto-encoder networks; Encoder-decoder networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/047Probabilistic or stochastic networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0475Generative networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/088Non-supervised learning, e.g. competitive learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/094Adversarial learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T3/00Geometric image transformations in the plane of the image
    • G06T3/40Scaling of whole images or parts thereof, e.g. expanding or contracting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T3/00Geometric image transformations in the plane of the image
    • G06T3/40Scaling of whole images or parts thereof, e.g. expanding or contracting
    • G06T3/4053Scaling of whole images or parts thereof, e.g. expanding or contracting based on super-resolution, i.e. the output image resolution being higher than the sensor resolution
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/20Image enhancement or restoration using local operators
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00Image enhancement or restoration
    • G06T5/70Denoising; Smoothing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2200/00Indexing scheme for image data processing or generation, in general
    • G06T2200/04Indexing scheme for image data processing or generation, in general involving 3D image data
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10072Tomographic images
    • G06T2207/10081Computed x-ray tomography [CT]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30004Biomedical image processing

Definitions

  • the present invention relates to a learning method, a learning system, a learned model, a program, and a super-resolution image generation device, and more particularly to a machine learning technique and an image processing technique for realizing super-resolution image generation.
  • Patent Document 1 describes a method for realizing a super-resolution image generation by learning a generation model using a hostile generation network (GAN: Generat i ve Adversar i a l Networks).
  • GAN includes a generation network called a generator that creates data and an identification network called a discriminator that identifies data. The discriminator discriminates whether the input data is the correct data from the training data or the data derived from the output of the generator. The goal is to enable the generator to generate data close to the learning data by updating the generator and discriminator alternately during learning and increasing the accuracy of both.
  • Non-Patent Document 2 describes a method of learning a pair of an input image and an output image using GAN.
  • Non-Patent Document 3 describes a study introducing a self-attention mechanism into GAN.
  • the self-attention mechanism is a mechanism that adds global information to the feature map output from the hidden layer of the network.
  • the method described in Non-Patent Document 3 introduces a self-attention mechanism in both the generator and the discriminator network, and the It is possible to generate high-resolution images for input data.
  • Patent Document 1 US Patent Application Publication 2018/0075581
  • Non-Patent Document 1 Ian ⁇ Goodfe L Low, Jean Pouget-Abadie, Mehdi Mirza, Bing
  • Non-Patent Document 2 Phi Uip Iso La, Jun-Yan Zhu, Tinghui Zhou, Alexei A. Efro s Image-to-Image Translation with Conditional Adversarial Networks ”, CVPR2016
  • Non-Patent Document 3 Han Zhang, Ian Goodfe L Low, Dimitris Metaxas, Augustus Od ena “Self-Attention Generative Adversarial Networks”, arXi v: 1805.08 318
  • Non-Patent Document 3 has the following problems.
  • the present invention has been made in view of the above circumstances, and is capable of handling input data of any size without being restricted by the image size at the time of learning and suppressing the amount of calculation at the time of image generation. It is an object of the present invention to provide a learning method and learning system for a generative model, a program, a trained model, and a super-resolution image generation device capable of performing the above.
  • a learning method is a learning method that performs machine learning of a generation model that estimates a second image that includes image information having a higher resolution than the first image from the first image, Adversarial including a generator, which is a generative model, and a discriminator, which is an identification model that identifies whether the given data is the data of the correct answer image for learning or the data derived from the output from the generator.
  • Adversarial including a generator, which is a generative model, and a discriminator, which is an identification model that identifies whether the given data is the data of the correct answer image for learning or the data derived from the output from the generator.
  • the second learning image which is an image that is the correct answer image corresponding to the first learning image, is used as learning data, and the input of the generator is either the first learning image or the second learning image.
  • This is a learning method that includes providing only the first learning image and implementing the self-attention mechanism only in the discriminator network of the generator and discriminator.
  • each network _ of the generator and discriminator _ is a tatami: convolutional _ _ network__ ⁇ 0 2020/175446 4 ⁇ (: 171? 2020 /007383
  • the first image is a three-dimensional tomographic image
  • the second image has a resolution in the slice thickness direction of at least the three-dimensional tomographic image higher than that of the first image. It can be configured to be a resolution.
  • the second learning image is an image acquired by using a computer tomography apparatus, and the first learning image is the second learning image. It can be configured to be an image generated by image processing based on.
  • the image processing for generating the first learning image from the second learning image includes a process of down-sampling the second learning image. be able to.
  • the image processing for generating the first learning image from the second learning image is performed by performing interpolation processing on the image obtained by the downsampling processing and applying It can be configured to include processing to sample.
  • the image processing for generating the first learning image from the second learning image may include a smoothing process using a Gaussian filter.
  • each of the first learning image and the second learning image in a plurality of types of learning data used for machine learning may be configured to have the same size. it can.
  • the second image is a high-frequency component image indicating the information of the high-frequency component
  • the generator is required to increase the resolution of the input image. It can be configured to estimate a high frequency component and output a high frequency component image showing information on the high frequency component.
  • a learning method further comprising: adding a high-frequency component image output from the generator and an image input to the generator, the virtual second obtained by the addition Image of discriminator ⁇ 0 2020/175446 5 ⁇ (: 171? 2020 /007383
  • a program according to another aspect of the present disclosure is a program for causing a computer to execute the learning method according to any one aspect of the present disclosure.
  • a trained model according to another aspect of the present disclosure is a trained model that has been trained by performing the learning method according to any one of the embodiments of the present disclosure, and includes a first image to a first image. Is a generative model that estimates the second image containing high-resolution image information.
  • a super-resolution image generation device includes a generation model that is a learned model that has been learned by performing the learning method according to any one aspect of the present disclosure, and is input. This is a super-resolution image generation device that generates a fourth image including image information with higher resolution than the third image.
  • the third image may have a different image size from the first learning image.
  • a first interpolation processing unit that performs interpolation processing on a third image to generate an interpolation image, and a high frequency generated by the interpolation image and the generation model
  • a first addition unit that adds the component and the interpolating image is input to the generation model, and the generation model can generate a high frequency component necessary for increasing the resolution of the interpolation image. ..
  • a learning system is a learning system that performs machine learning of a generation model that estimates a second image including image information having a higher resolution than the first image from the first image.
  • a generator which is a generation model
  • a discriminator which is an identification model that identifies whether the given data is the data of the correct answer image for learning or the data derived from the output from the generator, and
  • the self-attention mechanism is installed only in the discriminator network and includes the first resolution information that has a lower resolution than the second image.
  • Image and the second learning image that contains the second resolution information having a higher resolution than the first learning image and is the correct answer image corresponding to the first learning image.
  • This is a learning system in which only the first learning image of the first learning image and the second learning image is given to the input of the generator, and learning of the adversarial generation network is performed.
  • a learning data generating unit that generates learning data is further provided, and the learning data generating unit determines a fixed size area from an original original image including the second resolution information.
  • the fixed size area cutout part that cuts out the image, and the downsampling processing part that downsamples the image of the fixed size area cut out by the fixed size area cutout part, are included.
  • An image in the size region may be used as the second learning image, and the first learning image may be generated by performing downsampling on the second learning image.
  • the learning data generating unit further includes: a second interpolation processing unit that performs an interpolation process on the image obtained by the downsampling process; and a Gaussian filter.
  • a smoothing processing unit that performs smoothing may be included.
  • the generator estimates a high-frequency component necessary to increase the resolution of the input image, and outputs a high-frequency component image indicating information of the high-frequency component
  • the second addition unit that adds the high-frequency component image output from the generator and the image input to the generator can be further included.
  • a learning system is a learning system that performs machine learning of a generative model that estimates a second image including image information with higher resolution than the first image from the first image. , which includes at least one processor, which identifies the generator that is the generation model and whether the given data is the data of the correct answer image for training or the data derived from the output of the generator.
  • a discriminator which is an identification model, and a hostile generation network including ⁇ 0 2020/175446 7 ⁇ (: 171? 2020 /007383
  • the self-attention mechanism is installed only in the discriminator network, and the first learning image including the first resolution information, which has a lower resolution than the second image, ,
  • the second learning image that contains the second resolution information with a higher resolution than the first learning image and is the correct answer image corresponding to the first learning image, and capture as learning data
  • the first learning image of the first learning image and the second learning image is given to the input of the generator, and the adversarial generation network is learned.
  • the present invention it is possible to obtain a generation model capable of generating a high-resolution image for input data of arbitrary size. Further, according to the present invention, it is possible to obtain a generation model capable of suppressing the amount of calculation at the time of image generation, and it is possible to realize high-accuracy image generation using a learned model.
  • FIG. 2 is a diagram for explaining a slice interval and a slice thickness of a crocodile image.
  • FIG. 3 is a diagram for explaining a slice interval and a slice thickness of a crocodile image.
  • FIG. 4 is a diagram for explaining a slice interval and a slice thickness of a crocodile image.
  • FIG. 5 is a diagram for explaining a slice interval and a slice thickness of a crocodile image.
  • FIG. 6 is a functional block diagram showing an example of a super-resolution image generation device according to an embodiment of the present invention.
  • FIG. 7 is a block diagram showing a configuration example of a learning system according to an embodiment of the present invention. ⁇ 0 2020/175446 8 ⁇ (: 171? 2020 /007383
  • FIG. 8 is a functional block diagram showing a configuration example of a learning data generation unit.
  • Fig. 9 is a table showing an example of combinations of conditions of the Gaussian filter corresponding to the slice interval and the assumed slice thickness applied when generating the learning data.
  • FIG. 10 is a flowchart showing an example of a procedure of processing for generating learning data.
  • Fig. 11 is a conceptual diagram of the processing in the learning unit to which ⁇ 8 1 ⁇ 1 is applied.
  • Fig. 12 is a conceptual diagram showing an example of an identification network applied to a discriminator.
  • FIG. 13 is a schematic diagram showing an example of a generation network applied to a generator.
  • FIG. 14 is an explanatory diagram of an operation of adding a low resolution image to the output of the generator to generate a virtual high resolution image.
  • Fig. 15 is a diagram for explaining the discriminating operation by the discriminator during learning.
  • Fig. 16 is a flow chart showing an example of the processing procedure in the learning unit.
  • FIG. 17 is an example of an image showing the effect of the embodiment.
  • FIG. 18 is a diagram for explaining another effect of the embodiment.
  • FIG. 19 is a functional block diagram schematically showing a flow of processing by the learning system according to the second embodiment.
  • FIG. 20 is a functional block diagram schematically showing a flow of processing by the learning system according to the third embodiment.
  • FIG. 21 is a functional block diagram schematically showing a flow of processing by the learning system according to the fourth embodiment.
  • FIG. 22 is a block diagram showing an example of a hardware configuration of a computer.
  • the super-resolution image generation device generates virtual high-resolution image data from low-resolution image data.
  • “Generate” includes the concept of the term “estimate”.
  • image data a thick slice (Thick slice) acquired using a CT device is targeted for CT image data acquired using a Computed Tomography (CT) device.
  • CT Computed Tomography
  • Thick slice image data refers to low resolution CT image data with a relatively large slice interval and slice thickness.
  • CT image data with a slice interval and slice thickness exceeding 4 mm correspond to thick slice image data.
  • Thick slice image data may be referred to as “thick slice image”, “six rice data”, or “thick data”.
  • the thin slice image data is high-resolution CT image data with a small slice interval and small slice thickness.
  • CT image data with a slice interval and slice thickness of about 1 mm corresponds to thin slice image data.
  • the image data of thin rice may be referred to as “thin slice image”, “thin slice data”, or “thin data”.
  • a virtual thin slice image generated from a thick slice image is referred to as a virtual thin slice (VTS) image.
  • VTS virtual thin slice
  • RTS real thin slice
  • Figure 1 is an image diagram of each data of the thick slice image and VTS image.
  • the left figure in Fig. 1 is the thick slice image, and the right figure is the VTS image.
  • VTS images can generate higher quality reconstructed images than thick slice images.
  • the z-axis direction is the body axis direction.
  • the CT data may have various slice intervals and slice thickness data depending on the model of the CT apparatus used for imaging and the setting of the output slice condition.
  • Figs. 2 to 5 are diagrams for explaining examples of slice intervals and slice thicknesses of CT images.
  • the slice interval is the distance between the center positions of the thicknesses of a slice and the slices adjacent thereto.
  • the slice interval is synonymous with the distance between slices.
  • the slice thickness is the length in the thickness direction of one slice at the center of the imaging area. Slice thickness is synonymous with slice thickness (S I ice thickness).
  • S D slice interval
  • S T slice thickness
  • the slice thickness direction is the Z-axis direction.
  • FIG. 3 is an explanatory diagram schematically showing a T image M 1; Here, for simplicity, a group of three layered image layers is schematically shown.
  • FIG. 3 is an explanatory diagram schematically showing a T image M3.
  • the slice interval S D is larger than the slice thickness S T, adjacent tomographic images are separated from each other and there is a gap between layers.
  • the CT image m4 shown in Fig. 5 has more information in the Z direction than the other CT images M1 to M3 shown in Figs. 2 to 4.
  • the 0-Cho image 1/14 has a higher resolution in the Z direction than any of the CT images I M 1, M 2, and M 3.
  • the slice interval and slice thickness of the CT image are set under various conditions according to the facility using the CT apparatus, the preference of the doctor, and the like.
  • CT images are for diagnosis ⁇ 0 2020/175446 1 1 ⁇ (: 171? 2020 /007383
  • high-resolution 0x images have a large amount of data and put pressure on the storage capacity of the storage, they may be saved at a lower resolution in order to reduce the capacity. For example, in the case of old 0-day data, the number of slices taken is reduced and saved in the database.
  • the thick slice image has poor quality in a reconstructed image or a volume rendering image viewed from a lateral direction with a plane parallel to the body axis as a cross section, and thus has a problem that it is difficult to use for sufficient observation and analysis. is there.
  • the super-resolution image generation apparatus uses, for example, from Fig. 5 to
  • the slice interval is 1 01 111 and the slice thickness is Performs image generation processing to generate V resolution 3 images with resolution.
  • FIG. 6 is a functional block diagram showing an example of the super-resolution image generation device according to the embodiment of the present invention.
  • the super-resolution image generation device 10 includes an interpolation processing unit 12, a generator 14 that is a learned model of a hierarchical neural network, and an addition unit 16.
  • “Neural network” is a mathematical model of information processing that simulates the mechanism of the cranial nervous system. Processing using a neural network can be realized using a computer.
  • the neural network can be configured as a program module. In this specification, the neural network may be simply referred to as “network”.
  • the interpolation processing unit 12 performs spline interpolation on the input low-resolution thick slice image D (3 ⁇ , and generates an interpolation image D. It is output from the interpolation processing unit 1 2.
  • the interpolated image is a blurred image in one direction and is an example of a low-resolution image.
  • the number of pixels of the interpolated image is set to match the number of pixels of the final virtual thin slice image. It is preferable to set.
  • the interpolated image data output from the interpolation processing unit 12 is input to the generator 14 ⁇ 0 2020/175446 12 12 (:171? 2020 /007383
  • the generator 14 is a generative model learned by machine learning using an adversarial generative network (081 ⁇ !). The learning method for obtaining the generator 14 will be described later.
  • the trained model may be restated as a program module.
  • the generator 14 generates (estimates) high-frequency component information necessary for generating a high-resolution image from the input image, and outputs high-frequency component information.
  • the addition unit 16 adds the map of the high-frequency component information output from the generator 14 and the interpolated image I data itself, which is the input data of the generator 14, to generate a virtual thin slice image V data. ..
  • FIG. 6 shows an example in which the input to the generator 14 is an interpolated image finder and the output of the generator 14 is high frequency component information.
  • [ ⁇ is also possible.
  • the output of the generator 14 is a virtual thin slice image.
  • the high-frequency component information is information that can be combined with the original image to generate a high-resolution image, so the high-frequency component information map is called the “high-frequency component image”.
  • the high-frequency component image is an image containing high-resolution image information, and can be understood as being substantially similar to a “high-resolution image”.
  • the thick slice image " ⁇ " is an example of the “third image” in the present disclosure.
  • the virtual thin slice image V-chome is an example of the “fourth image” in the present disclosure.
  • the high-frequency component output from the generator 14 is an example of “image information having higher resolution than the third image” in the present disclosure.
  • the interpolation processing unit 12 is an example of the “first interpolation processing unit” in the present disclosure.
  • the addition unit 16 is an example of the “first addition unit” in the present disclosure.
  • FIG. 7 is a block diagram showing a configuration example of the learning system 20 according to the embodiment of the present invention.
  • the learning system 20 includes an image storage unit 24, a learning data generation unit 30 and a learning unit 40.
  • the learning system 20 consists of one or more It can be realized by a computer system including a computer. That is, the functions of the image storage unit 24, the learning data generation unit 30, and the learning unit 40 can be realized by a combination of hardware and software of the computer.
  • each of the image storage unit 24, the learning data generation unit 30, and the learning unit 40 is configured as a separate device will be described. However, these functions may be realized by one converter.
  • the processing function may be shared by two or more converters.
  • the image storage unit 24, the learning data generation unit 30, and the learning unit 40 may be connected to each other via a communication line.
  • connection includes not only wired connection but also the concept of wireless connection.
  • the communication line may be an oral area network or a wide area network.
  • the learning data generation and the learning of the generative model can be performed without being physically and temporally bound to each other.
  • the image storage unit 24 includes a mass storage device that stores a CT reconstructed image (CT image) taken by a medical X-ray CT apparatus.
  • CT image CT reconstructed image
  • the image storage unit 24 may be, for example, a storage in a medical image management system represented by PACS (Picture Archiving and Communication Systems).
  • PACS Picture Archiving and Communication Systems
  • the image storage unit 24 stores data of a plurality of thin slice images, which are real high-resolution images taken by using a CT device (not shown).
  • the CT image stored in the image storage unit 24 is a medical image of a human body (subject), and is a three-dimensional tomographic image including a plurality of tomographic images.
  • each slice image is an image that is parallel to the X and Y directions that are orthogonal to each other.
  • the Z direction, which is orthogonal to the X and Y directions, is the body axis direction of the subject and is also called the slice thickness direction.
  • the CT image stored in the image storage unit 24 may be an image of each part of the human body or an image of the whole body.
  • the learning data generation unit 30 generates the learning data necessary for the learning unit 40 to perform learning.
  • Learning data is training data used for machine learning and is synonymous with “learning data” or “training data”.
  • Mechanics of this embodiment ⁇ 0 2020/175446 14 ⁇ (: 171? 2020 /007383
  • the learning data generation unit 30 acquires the original real high resolution image from the image storage unit 24 and performs downsampling processing on the real high resolution image to obtain various low resolution images. (Pseudo thick slice image) is artificially generated.
  • the learning data generation unit 30 uses, for example, 1
  • the orientation conversion is performed on the original thin slice data that has been isotropically sliced, and a fixed size area is randomly extracted.
  • Virtual of Slice data is generated, and data of virtual 80!!! slices with a slice interval of 80!!! is generated.
  • the fixed-size area may be a three-dimensional area in which the number of pixels in the X-axis direction, the XV-axis direction, and the X-axis direction is, for example, "160 X 160 X 160".
  • Learning me by the data generating unit 3 0, the low-resolution image of a fixed size for learning! _ ⁇ and real high-resolution image 1 to 1 of the image pair corresponding thereto is generated.
  • a learning data generation unit is prepared in advance.
  • the learning unit 40 uses an adversarial generation network (08 1 ⁇ 1) 4 as a learning model.
  • the architecture of the learning unit 40 is based on a structure obtained by extending the architecture described in Non-Patent Document 2 from two-dimensional to three-dimensional data.
  • ⁇ Eight-four 41 is composed of a generator network called data generator 4 2 ⁇ that creates data and an identification network called discriminator 4 40 that identifies the input data. That is, the generator 420 is a generative model that generates image data, and the discriminator 440 is ⁇ 0 2020/175446 15 ⁇ (: 171? 2020 /007383
  • the term "generator” is synonymous with the terms “generator,” “generator,” and “generation model.”
  • the term “discriminator” is synonymous with terms such as “identifier,” “discriminator,” and “discriminant model.”
  • the learning unit 40 repeats adversarial learning using the generator 420 and the discriminator 4 4 based on the input learning data, thereby improving the performance of both models. Learn generator 4 2 ⁇ .
  • the discriminator 440 of this example is equipped with a self-attention mechanism.
  • the layer that introduces the self-attention mechanism in the network of the discriminator 440 may be a part of the plurality of convolutional layers or all of them. Details of the configuration and operation of the discriminator 44 including the self-attention mechanism, and an example of the learning method of ⁇ 4 1 will be described later.
  • the learning unit 40 includes an error calculation unit 50 and an optimizer 52.
  • the error calculator 50 evaluates the error between the output of the discriminator 440 and the correct answer using the loss function.
  • the optimizer 52 performs a process of updating the network parameters based on the calculation result of the error calculator 50.
  • Network parameters include the filter coefficient (coupling weight between nodes) and the bias of nodes used in the processing of each layer.
  • the optimizer 52 calculates the update amount of the parameters of the generator 42 and the discriminator 440 from the calculation result of the error calculator 50, and performs the parameter calculation processing of the parameter calculator. According to the calculation result, the parameter updating process is performed to update the parameters of each network of generator 42 and discriminator 440.
  • the optimizer 52 updates the parameters based on an algorithm such as the gradient descent method.
  • FIG. 8 is a functional block diagram showing a configuration example of the learning data generation unit 30. Learning ⁇ 0 2020/175446 16 ⁇ (: 171? 2020 /007383
  • the data generation unit 30 includes a fixed size region cutout unit 31, a downsampling processing unit 32, an upsampling processing unit 34, and a learning data storage unit 38.
  • the fixed size area cutout 31 is the original real high resolution image input.
  • the down-sampling processing unit 32 down-samples the real high-resolution images 1 to 11 in the axial direction to obtain a low-resolution thick slice image. Generates 1.
  • thinning processing may be simply performed so as to reduce the slice in the axial direction at a constant rate. In this example, only down-sampling in the axial direction is performed and down-sampling is not performed in the X-axis direction and the down-axis direction. It is possible.
  • Thick slice image generated by the downsampling processing unit 32 1 is input to the upsampling processing unit 34.
  • the upsampling processing unit 34 is a thick slice image. 1 is axially upsampled to generate a low-resolution thin-slice image, low-resolution image !_ ⁇ 1.
  • the processing of upsampling may be, for example, a combination of spline interpolation and Gaussian filter processing.
  • the upsampling processing unit 34 includes an interpolation processing unit 35 and a Gaussian filter processing unit 36.
  • the interpolation processing unit 35 uses, for example, a thick slice image. Perform spline interpolation on 1.
  • the interpolation processing unit 35 may be the same processing unit as the interpolation processing unit 12 described in FIG.
  • the Gaussian filter processing unit 36 performs smoothing by applying a Gaussian filter to the image output from the interpolation processing unit 35.
  • the interpolation processing unit 35 shown in FIG. 8 is an example of the “second interpolation processing unit” in the present disclosure.
  • the Gaussian fill processing unit 36 is an example of the “smoothing processing unit” in the present disclosure.
  • the low resolution image !_ ⁇ 1 output from the upsampling processing unit 34 is a real high resolution image. It is preferable that the data has the same number of pixels as 1. here ⁇ 0 2020/175446 17 ⁇ (: 171? 2020/007383
  • the low resolution image !- 0 1 and the real high resolution image 1 ⁇ 1 1 are the same size.
  • the low-resolution image !_ ⁇ 1 is of lower quality (that is, lower resolution) than the real high-resolution image [3 ⁇ 4 ! 1].
  • a pair of the low-resolution image !_ ⁇ 1 generated in this way and the real high-resolution images 1 to 11 from which it was generated are linked and stored in the learning data storage unit 38.
  • the original real high-resolution images 0 to 11 are examples of the "original original image” in the present disclosure.
  • the real high-resolution images 1 to 11 are examples of the “second learning image” in the present disclosure.
  • the low-resolution image !_ ⁇ 1 is an example of the “first learning image” in the present disclosure.
  • the image information of the low resolution image !_ ⁇ 1 is an example of “first resolution information” in the present disclosure.
  • the image information of the real high-resolution images 1 to 11 is an example of “second resolution information” in the present disclosure.
  • the learning data generation unit 30 uses one original real high-resolution image ⁇ 1 ⁇ 1
  • Multiple real high-resolution images can be created by changing the clipping position of the fixed size area from 1 By generating a low-resolution image! _ 0 corresponding to each of the real high-resolution image 1 to 1, it is possible to generate a plurality of image bare-.
  • the learning data generation unit 30 changes various slices by changing the combination of the slice interpolation magnification in the upsampling processing unit 34 and the condition of the Gaussian filter applied to the upsampling processing unit 34. It is possible to generate a low resolution image of the condition.
  • the slice interpolation magnification corresponds to the downsampling condition in the downsampling processing unit 32.
  • FIG. 9 is a chart showing an example of combinations of slice intervals and Gaussian filter conditions corresponding to the assumed slice thickness applied when generating training data.
  • the low-resolution image !_ 0 has two slice intervals of 41111 and 8
  • the slice interpolation magnification during learning is two patterns, 4 times and 8 times.
  • the slice thickness is 0 according to the slice interval. The range of ⁇ 0 2020/175446 18 ⁇ (: 171? 2020 /007383
  • FIG. 10 is a flow chart showing an example of a procedure of processing for generating learning data. Each step of the flow chart shown in FIG. 10 is executed by a computer including a processor functioning as a learning data generation unit 30.
  • the computer is And a memory.
  • the computer may include ⁇ 11 (0 “8 ⁇ 11 _1 ⁇ 3 ⁇ “00633 _1 Hit 11 ⁇ 1;).
  • the learning data generation method is as follows: the original image acquisition process (step 31), the fixed size region cutting process (step 32), the down sample process (step 33), and the up-sample process. It includes a sample process (step 34), and a learning data storage process (step 35).
  • step 31 the learning data generation unit 30 acquires original real high-resolution images 0 1 to 1 from the image storage unit 24.
  • an isotropic real high-resolution image ⁇ 1 to 1 with a slice interval of 101 111 and a slice thickness of 10! Is acquired.
  • step 32 the fixed size region cutout unit 3 1 performs a process of cutting out the fixed size region from the input original real high resolution image 0 to 1 to obtain the fixed size region real high resolution image. Generates 1.
  • step 33 the downsampling processing unit 32 determines whether the high resolution image
  • the slice interval is 411111, Thick slice image equivalent to Is generated.
  • step 34 the upsampling processor 34 upsamples the thick slice image 1 obtained by the downsampling to obtain a low quality si ⁇ 0 2020/175 446 19 ⁇ (: 17 2020 /007383
  • the interpolation process and the Gaussian filter process are performed by applying the slice interpolation scale factor corresponding to the slice interval and the conditions of the Gaussian filter.
  • step 35 the learning data generation unit 30 sets the low resolution image !- 0 1 generated in step 34 and the real high resolution images 1 to 1 which is the generation data thereof as an image pair. These data are linked and stored as learning data in the learning data storage unit 38.
  • learning data generation section 30 ends the flowchart of FIG.
  • step 35 in the case of generating a plurality of learning data from the same original real high-resolution image ⁇ 1 to 1 by changing the position of the cutout area, after step 35, return to step 32 and proceed to step 32. Repeat the process from 3 2 to step 3 5.
  • the learning data generation unit 30 repeats the processing from Step 31 to Step 35 on the plurality of original real high-resolution images stored in the image storage unit 24, Can generate many learning data
  • the generator 14 mounted on the super-resolution image generation device 10 is a generation model obtained by performing learning by 088 1 ⁇ 1.
  • the configuration of the learning unit 40 and the learning method will be described in detail below.
  • FIG. 11 is a conceptual diagram of the processing in the learning unit 40 to which 088 1 ⁇ 1 is applied.
  • FIG. 11 shows an example in which a pair of a low resolution image 1-01 and a real high resolution image 1 is input to the learning unit 40 as learning data. ⁇ 0 2020/175446 20 ⁇ (: 171? 2020 /007383
  • the input to the generator 42 is a low resolution image !_ ⁇ 1.
  • the generator 42 generates a virtual high resolution image V 1 ⁇ 1 1 from the input low resolution image !_ 0 1 and outputs it.
  • Virtual high resolution image 1 corresponds to the virtual thin slice image (3 images).
  • the input to the discriminator 440 is the virtual high-resolution image ⁇ ! 1 generated by the generator 420 and the low-resolution image from which this virtual high-resolution image ⁇ ! 1 is generated.
  • _ ⁇ 1 pair, or a pair of real high-resolution image 1 ⁇ 1 1 and low-resolution image !_ 0 1 which are learning data are given.
  • the discriminator 440 is designed so that the input image pair is a real high resolution image. 3 pairs) (whether it is training data) or a fake bear (3 1 ⁇ 6 pairs) containing a virtual high resolution image ⁇ ! 1 derived from the output of the generator 420. , Output the identification result.
  • the error calculator 50 evaluates the error between the output of the discriminator 440 and the correct answer using the loss function.
  • the optimizer 52 performs a process of automatically adjusting network parameters based on the calculation result of the error calculator 50.
  • the network parameters include the weight of the connection between nodes and the bias of the nodes.
  • the optimizer 52 uses the calculation result of the error calculation unit 50 as the parameter calculation process to calculate the update amount of each parameter of the generator 4 2 ⁇ and discriminator 4 40 and the calculation result of the parameter calculation unit. Accordingly, the parameter updating process for updating the network parameters of the generator 42 and discriminator 440 is performed.
  • the optimizer 52 updates the parameters based on an algorithm such as the gradient descent method.
  • the technique described in Non-Patent Document 1 or the like may be adopted for the basic mechanism of learning regarding error evaluation and parameter updating.
  • the generator 420 learns to fool the discriminator 440 to produce a more refined virtual high-resolution image, which the discriminator 440 identifies more accurately. Learn to do.
  • a self-attention mechanism is mounted on the network applied to the discriminator 440 in the present embodiment.
  • the self-attention mechanism is a method that improves the calculation efficiency by considering the global part in the image.
  • Non-Patent Document 3 The content of the self-attention mechanism is described in Non-Patent Document 3. However, in Non-Patent Document 3, the self-attention mechanism is added to both the network of the generator and the discriminator, whereas in the present embodiment, the self-attention mechanism is mounted on the generator 4 2°. However, the method differs from the method described in Non-Patent Document 3 in that the self-attention mechanism is mounted only on the four discriminators.
  • the self-attention mechanism will be briefly described with reference to the contents of Non-Patent Document 3.
  • the self-attention mechanism generates a query command (X) and a key 9 (X) from the convolutional feature map ⁇ IV! , Calculates a value (similarity) indicating which pixel is similar to other pixels. In this way, the map of the similarity calculated for all the pixels of the feature map IV! (X) is called the “attention map”.
  • the attention map plays a role of finding and enhancing a region having similar features in the image.
  • the convolution operation of the convolutional layers that make up the identification network local information is overlapped, but by introducing an attention map, it is possible to consider the global (global) information. ..
  • the attention map is multiplied by the weight II (X) to obtain the self-attention feature map 318 1 ⁇ /1 ( ⁇ ). Then, multiply the self-attention feature map 3 1//1 ( ⁇ ) by the scale parameter ⁇ ⁇ and add it to the convolutional feature map ⁇ IV! (X) that is the original input feature map to the next layer. hand over. That is, the final output V passed to the next layer is given by the following equation.
  • FIG 12 is a conceptual diagram showing an example of an identification network applied to the discriminator 4 4D.
  • the discriminator 4 4D network is a hierarchical neural network classified as a deep dual neural network, and includes multiple convolution layers.
  • the network of the discriminator 4 4D is composed of a tatami: convolutional neural network (CNN).
  • arrows indicated by the symbols C O 1, C 0 2 C 0 5 represent "tatami: convolutional layer".
  • the rectangles shown on the input side and/or output side of each layer represent the set of feature maps.
  • the vertical length of the rectangle represents the size (number of pixels) of the feature map, and the horizontal width of the rectangle represents the number of channels.
  • a self-attention mechanism is introduced from the convolutional layer C 0 2 to each of the subsequent layers.
  • a self-attention feature map is generated for each C N N feature map of the 1 2 8 channels output from the convolutional layer C 0 2.
  • Figure 12 shows the addition of the self-attention feature map for 128 channels corresponding to each CNN feature map for 128 channels. For each channel, each CNN feature map and self-attention feature map are added, and the output is input to the next convolutional layer. The same applies to the convolutional layers C 0 3 and C 0 4.
  • the CNN feature map input to the self-attention mechanism is ⁇ 0 2020/175446 23 ⁇ (: 171? 2020 /007383
  • the whole image is converted into a one-dimensional array and calculated.
  • the number of input channels is ⁇ 3, the total number of pixels is ⁇ It is input to the self-attention mechanism as a vector in which the elements of each 1 ⁇ 1 pixel are arranged in one dimension.
  • the actual crocodile image data is three-dimensional data, and multidimensional data can be calculated in the same one-dimensional array as above.
  • the same processing algorithm can be applied to both 2D image data and 3D image data by calculating them in a 1D array.
  • Figure 13 is a conceptual diagram showing an example of a generation network applied to the generator 420.
  • the network of the generator 420 is also composed of a tatami: convolutional neural network.
  • the generator 420 preferably has an encoder/decoder structure in which an encoder unit and a decoder unit are combined.
  • Figure 13 shows an example of a II-shaped network called the ⁇ -N 61: structure.
  • the "N 6" in the "Lee” notation is a shorthand notation for "Network (1/ ⁇ / ⁇ "10").
  • each of the arrows indicated by the symbols ⁇ 1, ⁇ 2 ⁇ 10 represents the “convolutional layer”.
  • the arrows indicated by the symbols 1, 2, 11, II 3, and II 4 represent the convolutional layers that perform “convolution and upsampling”.
  • the generator 4 2 ⁇ is a high-frequency component image as high resolution information necessary for high resolution from the input low resolution image !_ ⁇ .
  • Fig. 14 by adding the low-resolution image !_ ⁇ , which is the input data to the generator 4 2°, and the high-frequency component images V 1 to 1 0 generated by the generator 4 2 ⁇ , High resolution image Is obtained.
  • the virtual high resolution image The slice interval and slice thickness of 1 are the low resolution images! ⁇ 0 2020/175446 24 ⁇ (: 171? 2020 /007383
  • Virtual high-resolution image equivalent to slice spacing and slice thickness 1 is sharper in the direction than the low resolution image !_ 0 1.
  • the learning section 40 is provided with an adder section 46 that adds the input of the generator 420 and the output of the generator 420, and the output of the adder section 46 is provided to the discriminator 440. It is configured to input.
  • the adder 46 is an example of the “second adder” in the present disclosure.
  • the addition unit 46 is not shown in FIGS. 7 and 11.
  • the virtual high resolution images V 1 to 11 input to the discriminator 440 are examples of the “virtual second image” in the present disclosure.
  • the low-resolution image !_ ⁇ given to the input of the generator 42 is an example of the "first image” in the present disclosure.
  • the high-frequency component image output from the generator 42 is an example of the "second image” in the present disclosure.
  • FIG. 15 is a diagram for explaining an identification operation by the discriminator 440 during learning. Illustration of the addition unit 46 is omitted in FIG.
  • the operating state 70 shown in the left diagram of FIG. 15 shows an example in which a positive sample (positive example) is input to the discriminator 440, and the operating state shown in the right diagram of FIG.
  • the operating state 7 0 1 ⁇ 1 shows an example when a negative sample (negative example) is input to the discriminator 4 40.
  • the discriminator 440 is a high-resolution input image that is a real ⁇ image taken by a ⁇ device not shown, or a virtual image generated by a generator 242. ⁇ 3 images, or learn to answer correctly.
  • the generator 420 generates a virtual (3 images that resembles a real ⁇ 3 images captured by an unillustrated device, and makes the identification of the discriminator 4 4 incorrect. To be learned.
  • the discriminator 4 4 0 and the generator 4 2 ⁇ enhance each other's accuracy, and the generator 4 2 ⁇ is not identified as a fake (virtual high resolution image) by the discriminator 4 4 0. This makes it possible to generate virtual high-resolution images V 1 to 1 that are closer to real ⁇ 3 images.
  • the learned generator 4 2° obtained by such learning is applied as the generator 14 of the super-resolution image generation device 10 described in Fig. 6.
  • FIG. 16 is a flow chart showing an example of a processing procedure in the learning unit 40. Each step of the flow chart shown in Fig. 16 is executed by a computer including a processor functioning as a learning unit 40.
  • step 311 the learning unit 40 acquires the learning data.
  • the learning unit 40 reads the learning data from the learning data generation unit 30 described in FIG.
  • the learning unit 40 can acquire the learning data in units of mini-batch including a plurality of learning data.
  • step 312 the learning unit 40 inputs the low-resolution image of the learning data to the generator 42. ⁇ 0 2020/175446 26 ⁇ (: 171? 2020 /007383
  • step 313 the generator 420 generates a virtual high resolution image from the input low resolution image.
  • the output from the generator 420 may be the high frequency component images V 1 to 10 needed to create a virtual high resolution image.
  • the high frequency component image ⁇ ! (3 and the low resolution image are added to generate the virtual high resolution images V 1 to 1.
  • step 3 14 the learning unit 40 inputs the data to the discriminator 4 40.
  • the input to the discriminator 440 is a pair of training data containing a real high resolution image as a correct image (real pair) or a fake pair containing a virtual high resolution image derived from the general-purpose image 4 2 (5). Either of them is given selectively.
  • step 315 the discriminator 440 identifies the data.
  • step 316 the error calculator 50 calculates the error of the identification result and sends the result to the optimizer 52.
  • step 317 the optimizer 52 calculates the update amount of the network parameter based on the calculated error.
  • step 318 the optimizer 52 performs parameter update processing in accordance with the parameter update amount calculated in step 317.
  • Parameter-Evening update process is carried out in units of mini-batch.
  • learning section 40 determines whether or not to end learning.
  • the learning end condition may be determined based on the value of the error, or may be determined based on the number of parameter update times. As a method based on the error value, for example, the learning end condition may be that the error converges within a specified range. As a method based on the number of updates, for example, the learning end condition may be that the number of updates reaches a specified number.
  • step 3 1 9 If the judgment result in step 3 1 9 is 1 ⁇ 1 0 judgment, the learning unit 40
  • step 3 19 If the result of the judgment in step 3 19 is a judgment of 3 6 3 Finish the mouth chart.
  • the portion of the trained generator 4 2 G thus obtained is applied as the generator 14 of the super-resolution image generation device 10.
  • FIG. 17 is an example of an image showing the effect of this embodiment.
  • Figure 17 shows an example image showing the effect of introducing the self-attention mechanism.
  • An image V H A 1 shown in the center of the upper part of FIG. 17 is an example of a virtual high resolution image generated by using a learned model (generator 14) to which the learning method according to the present embodiment is applied.
  • Image L 1 shown in the upper left of Fig. 17 is an example of the low-resolution image used for input to generator 14.
  • the image G T 1 shown in the upper right of Fig. 17 is the correct image (Ground t ruth) corresponding to the input image L ⁇ 1.
  • the image V H N 2 shown in the lower center of Fig. 17 is an example of a virtual high-resolution image generated using the learned model according to the comparative example.
  • the trained model according to the comparative example was trained using a discriminator having no attention mechanism.
  • the image L I 2 shown in the lower left of Fig. 17 is an example of the low-resolution image used for input to the trained model according to the comparative example.
  • the image G T 2 shown in the lower right of Fig. 17 is the correct image (Ground t ruth) corresponding to the input image L ⁇ 2.
  • the image V H A 1 shown in the upper center of Fig. 17 is very close to the correct image G T 1. Further, it can be seen that the image V H A 1 has reduced local noise as compared to the image V H N 2 of the comparative example in the lower stage.
  • FIG. 18 is a diagram for explaining another effect of the present embodiment.
  • FIG. 18 is an image example showing that the trained model (generator 14) to which the learning method according to the present embodiment is applied can execute the image generation process (super-resolution process) regardless of the input size. It is shown.
  • the second image from the left, ⁇ ! !8 3, in Figure 18 is an example of a virtual high-resolution image generated using the generator 1 4 from the input of image !_ ⁇ 3.
  • the third image from the left !_ ⁇ 4 in Fig. 18 is another example of the image input to the generator 14.
  • This image !_ ⁇ 4 has a smaller image size than the image !_ ⁇ 3 shown on the far left.
  • the images V 1 to 18 shown on the far right of Fig. 18 are examples of virtual high-resolution images generated using the generator 1 4 from the input of image !_ ⁇ 4.
  • the generator 14 can perform the estimation process even on an input of an image having a size different from the image size used at the time of learning.
  • the generator 14 can generate highly accurate images for input data of any image size. That is, according to the present embodiment, the image generation processing can be performed on an image of an arbitrary size without being restricted by the fixed-size image size used as the learning data at the time of learning. According to the present embodiment, it is possible to divide the incoming data into memories of arbitrary size and perform the processing.
  • the discriminator 440 Since the discriminator 440 is used only for learning and does not need to be mounted on the super-resolution image generation device 10, an addition mechanism is added to the discriminator 440. However, there is no problem because learning is performed with a fixed size.
  • the virtual high-resolution image ⁇ derived from the pair of the real high-resolution image 8 ! and the low-resolution image !_ 0 1 or the generator 4 2 ⁇ as the input to the discriminator 4 40.
  • input of low resolution image !_ ⁇ 1 to the discriminator 440 is not mandatory.
  • At least the real high resolution images 1 to 1 or the virtual high resolution images V 1 to 1 may be input to the discriminator 440.
  • a (virtual) high-frequency component image is generated from the generator 420.
  • the virtual high-resolution images V 1 to 1 are obtained by adding and.
  • the second embodiment is a form in which the generator 420 outputs a virtual high resolution image ⁇ !.
  • FIG. 19 is a functional block diagram schematically showing the flow of processing by the learning system 20 according to the second embodiment.
  • parts that are the same as or similar to the configurations shown in FIGS. 7, 8, and 11 to 14 are denoted by the same reference numerals, and detailed description thereof will be omitted.
  • the contents of the learning data generation unit 30 are the same as in Fig. 8.
  • the generator 42 of the learning unit 40 in the second embodiment is a virtual high resolution image from the low resolution image !_ ⁇ 1. To generate.
  • the low-resolution image !_ 0 1 is input to the discriminator 440 in pairs with the real high-resolution image 1 to 11 or the virtual high-resolution image ⁇ ! 1!.
  • the discriminator 440 is a high-resolution image of the input image 1 ⁇ 1 1, and a virtual high-resolution image. Identify which one.
  • the low resolution image !_ ⁇ 1 does not have to be input in the discriminator 4440.
  • the second embodiment it is possible to obtain a product model for generating a virtual high-resolution images V 1 ⁇ 1 1 from the low-resolution image! _ 0 1 (generator 4 2 0).
  • the generator 420 generated by the second embodiment is incorporated in the super-resolution image generation device 10, the adder unit 16 shown in FIG. 6 can be omitted.
  • FIG. 20 is a functional block diagram schematically showing the flow of processing by the learning system 20 according to the third embodiment.
  • the input to the generator 42 is a thick slice image.
  • the output from the generator 420 is a virtual high resolution image ⁇ ! In this case, the thick slice image
  • the pair of real high-resolution images 1 to 1 are the learning data.
  • Thick slice image in Figure 20 Are examples of the “first image” and the “first learning image” in the present disclosure. ⁇ 0 2020/175446 30 ⁇ (: 171? 2020 /007383
  • the generator 420 is a thick slice image.
  • the thick slice image A generation model (generator 420) that generates virtual high-resolution images V 1 to 1 1 can be obtained from.
  • FIG. 21 is a functional block diagram schematically showing the flow of processing by the learning system 20 according to the fourth embodiment.
  • the learning system 20 according to the fourth embodiment shown in FIG. To output the high-frequency component images V 1 to 10 from the generator 4 2°, and input the high-frequency component image to the discriminator 4 40 for identification.
  • Discriminator 4 The learning data generator 30 has a high-frequency component extractor 33 in order to create a high-frequency component image for learning that is used as input to the four inputs.
  • the high-frequency component extraction unit 33 extracts high-frequency components from the real high-resolution images 1 to 1 to generate real high-frequency component images 8 1 to 10. Extraction of high frequency components is performed using a high pass filter. Similar to the real high-resolution image 8, the real high-frequency component images 1 to 10 have a slice interval of 101111 and a slice thickness of 101111.
  • the pair of ⁇ is the learning data.
  • Real high-frequency component image in Fig. 21 ⁇ is an example of the “second learning image” in the present disclosure.
  • the real high-frequency component image [3 ⁇ 4 ! ⁇ generated by the high-frequency component extraction unit 33 is input to the discriminator 440 of the learning unit 40.
  • the generator 4 2 ⁇ of the learning unit 40 is the input thick slice image. From the real high frequency component image
  • the discriminator 44 D includes a pair of a real high frequency component image RH FC and a thick slice image LK, or a virtual high resolution image VH FC and a thick slice image LK derived from the output of the generator 42 G. The pair is entered.
  • the discriminator 44 D identifies whether the input high frequency component image is the real high frequency component image R H F C or the virtual high frequency component image V H F C.
  • the generator 42G is learned to generate a high-frequency component image from the low-resolution image, the slice image L K.
  • a high-resolution image can be obtained by adding the high-frequency component image generated by the generator 42 G and the thick slice image L K input to the generator 42 G.
  • FIG. 22 is a block diagram showing an example of the hardware configuration of the computer used in the learning system 20.
  • the computer 500 may be a personal computer, a workstation, or a server computer.
  • the computer 500 can be used as any one of the super-resolution image generation device 10, the image storage unit 24, the learning data generation unit 30, and the learning unit 40, or as a device having a plurality of these functions.
  • the computer 500 includes a communication unit 5 1 2, a storage unit 5 1 4, an operation unit 5 1 6, a CPU (Central Processing Unit) 5 1 8, a GPU (Graphics Processing Unit) 5 1 9, and a RAM (Random Access Memory). ) 520, a ROM (Read Only Memory) 522, and a display unit 524. Note that G P U (Graphics Processing Unit) 5 19 may be omitted.
  • the communication unit 5 12 is an interface that performs communication processing with an external device by wire or wirelessly and exchanges information with the external device.
  • Storage 5 14 is, for example, a hard disk device, an optical disk, a magneto-optical
  • the storage device includes a disk, a semiconductor memory, or an appropriate combination thereof.
  • the storage 5 14 stores various programs and data necessary for image processing such as learning processing and/or image generation processing.
  • the program stored in the storage 5 14 is read out to the RAM 520, and the CPU 5 18 executes it, so that the computer functions as a means for performing various processes defined by the program.
  • the operation unit 5 16 is an input interface that receives various operation inputs to the computer 500.
  • the operation unit 5 16 may be, for example, a keyboard, a mouse, a touch panel, operation buttons, a voice input device, or an appropriate combination thereof.
  • the CPU 5 18 reads various programs stored in the ROM 522, the storage 5 14 or the like, and executes various processes.
  • RAM520 is used as a work area for CPU 518.
  • the RAM 520 is also used as a storage unit that temporarily stores the read program and various data.
  • the display unit 524 is an output interface that displays various kinds of information.
  • the display unit 524 may be, for example, a liquid crystal display, an organic EL (organic electro-luminescence) display, a projector, or an appropriate combination thereof.
  • a program that causes a computer to implement at least one of the learning data generation function, the learning function, and the image generation function described in each of the above-described embodiments is recorded on an optical disk, a magnetic disk, or a semiconductor. It is possible to record the program on a computer-readable medium, which is a non-transitory information storage medium that is a memory or other tangible material, and to provide the program through this information storage medium.
  • a program signal as a download service by using a telecommunication line such as the Internet.
  • a telecommunication line such as the Internet.
  • a part or all of the processing function of at least one of the learning data generation function, the learning function, and the image generation function described in each of the above embodiments is provided as an application server, and the processing is performed through a telecommunication line. It is also possible to provide services that provide functions.
  • the computer functioning as the learning data generation unit 30 is understood as a learning data generation device.
  • a computer that functions as the learning unit 40 is understood as a learning device.
  • the hardware structure of a processing unit that executes various processes such as the high frequency component extraction unit 33 of 1 is, for example, various processors as described below.
  • the various processors include a CPU that is a general-purpose processor that executes programs and functions as various processing units, a GPU that is a processor specialized for image processing, and a F PGA (Field Programmable Gate Array).
  • a CPU that is a general-purpose processor that executes programs and functions as various processing units
  • a GPU that is a processor specialized for image processing
  • F PGA Field Programmable Gate Array
  • To execute specific processing such as programmable logic device (Programmable Logic Dev ice: PLD), which is a processor whose circuit configuration can be changed after manufacturing the device, and AS IC (App 11 cat i on Specific Integra ⁇ ed Circuit) It includes a dedicated electric circuit, which is a processor having a circuit structure designed specifically for the purpose.
  • One processing unit may be configured by one of these various processors, or may be configured by two or more processors of the same type or different types.
  • one processing unit may be configured by a plurality of FPGAs, a combination of CPU and FPGA, or a combination of CPU and GPU.
  • the plurality of processing units may be configured by one processor. Multiple processing units in one As an example of configuring with a processor, firstly, as represented by a computer such as a client or a server, one processor is configured with a combination of one or more CPUs and software, and this processor has a plurality of processors. There is a form that functions as a processing unit.
  • a processor that realizes the functions of the entire system, including multiple processing units, with a single C (Integrated Circuit) chip, as typified by a system-on-chip (S ⁇ C).
  • S ⁇ C system-on-chip
  • the various processing units are configured by using one or more of the above various processors as a hardware structure.
  • the hardware structure of these various processors is, more specifically, an electrical circuit in which circuit elements such as semiconductor elements are combined.
  • the generation model learning method according to the present disclosure is not limited to CT images and can be applied to various three-dimensional tomographic images.
  • MR images acquired by MR Resonance Imaging (MR) device MR images acquired by PET (Positron Emission Tomography) device
  • PET images acquired by PET (Positron Emission Tomography) device PET images acquired by PET (Positron Emission Tomography) device
  • OCT images acquired by OCT (Optical Coherence Tomography) device 3D ultrasound It may be a three-dimensional ultrasonic image acquired by an imaging device.
  • the generation model learning method according to the present disclosure can be applied to various two-dimensional images as well as three-dimensional tomographic images.
  • it may be an X-ray image.
  • it is not limited to medical images, but can be applied to ordinary camera images.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Computational Linguistics (AREA)
  • General Engineering & Computer Science (AREA)
  • Medical Informatics (AREA)
  • Veterinary Medicine (AREA)
  • Pathology (AREA)
  • Heart & Thoracic Surgery (AREA)
  • Surgery (AREA)
  • Animal Behavior & Ethology (AREA)
  • Public Health (AREA)
  • Probability & Statistics with Applications (AREA)
  • High Energy & Nuclear Physics (AREA)
  • Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
  • Optics & Photonics (AREA)
  • Radiology & Medical Imaging (AREA)
  • Image Processing (AREA)
  • Image Analysis (AREA)
  • Apparatus For Radiation Diagnosis (AREA)
  • Measuring And Recording Apparatus For Diagnosis (AREA)

Abstract

任意サイズの入力データに対応でき、画像生成時の計算量を抑制することが可能な生成モデルの学習方法及び学習システム、プログラム、学習済みモデル、並びに超解像画像生成装置を提供する。本開示の一態様に係る学習方法は、第1画像から第1画像よりも高解像の画像情報を含む第2画像を推定する生成モデルの機械学習を行う学習方法であって、生成モデルであるジェネレータと、与えられたデータが学習用の正解画像のデータであるかジェネレータからの出力に由来するデータであるかを識別する識別モデルであるディスクリミネータと、を含む敵対的生成ネットワークを用い、ジェネレータ及びディスクリミネータのうち、ディスクリミネータのネットワークに限定してセルフアテンション機構を実装して学習を行う。

Description

明 細 書
発明の名称 :
学習方法、 学習システム、 学習済みモデル、 プログラム及び超解像画像生 成装置
技術分野
[0001 ] 本発明は、 学習方法、 学習システム、 学習済みモデル、 プログラム及び超 解像画像生成装置に係り、 特に、 超解像の画像生成を実現する機械学習技術 及び画像処理技術に関する。
背景技術
[0002] 近年、 多層のニューラルネッ トワークを用いて機械学習を行うことにより 、 画像を生成する技術が提案されている。 特許文献 1 には敵対的生成ネッ ト ワーク ( G A N : Generat i ve Adversar i a l Networks) を用いて生成モデルを 学習し、 超解像画像生成を実現する手法が記載されている。 非特許文献 1 に は G A Nに関する研究が記載されている。 G A Nは、 データを作り出すジェ ネレータと呼ばれる生成ネッ トワークと、 データを識別するディスクリミネ —夕と呼ばれる識別ネッ トワークとを含む。 ディスクリミネータは入力され たデータが学習データからの正解のデータであるか、 ジェネレータの出力に 由来するデータであるかを識別する。 学習の際にジェネレータとディスクリ ミネータとを交互に更新し、 両者の精度を高めていくことにより、 最終的に はジェネレータが学習データに近いデータを生成できるようにすることを目 指す。
[0003] 非特許文献 2には、 G A Nを用いて入力画像と出力画像とのペアを学習す る手法が記載されている。 非特許文献 3には、 G A Nにセルフアテンション 機構を導入した研究が記載されている。 セルフアテンション機構は、 ネッ ト ワークの隠れ層から出力される特徴マップに大域的な情報を付加する仕組み である。 非特許文献 3に記載された方法は、 ジェネレータとディスクリミネ —夕のネッ トワークの双方にセルフアテンション機構を導入し、 特定のサイ ズの入カデータに対して高解像度の画像生成を可能としている。
先行技術文献
特許文献
[0004] 特許文献 1 :米国特許出願公開 2018/0075581号
非特許文献
[0005] 非特許文献 1 : Ian 丄 Goodfe L Low, Jean Pouget-Abad i e, Mehdi Mirza, Bing
Xu, David Warde-Far Ley, Sherj i L Ozair, Aaron Courv i L Le, Yoshua Bengi o “Generative Adversarial Nets” , arXiv: 1406.2661
非特許文献 2 : Phi Uip Iso La, Jun-Yan Zhu, Tinghui Zhou, Alexei A. Efro s Image-to-Image Translation with Conditional Adversarial Networks ” ,CVPR2016
非特許文献 3 : Han Zhang, Ian Goodfe L Low, Dimitris Metaxas, Augustus Od ena “Self-Attention Generative Adversarial Networks” , arXi v: 1805.08 318
発明の概要
発明が解決しようとする課題
[0006] しかしながら、 非特許文献 3に記載された方法には以下のような課題があ る。
[0007] [課題 1 ] 非特許文献 3に記載された方法では、 ジェネレータがアテンシ ョン機構を持つため、 学習時と学習後の推定時とでジェネレータに入力させ るデータが同じ入カサイズである必要がある。 つまり、 学習済みのジェネレ —夕に入力できるデータのサイズが固定サイズに制約され、 任意の入カサイ ズに対応できない。
[0008] [課題 2] 非特許文献 3に記載された方法では、 ジェネレータがアテンシ ョン機構を持つため、 画像生成時 (推定時) にジェネレータの計算量が増加 する。 特に、 入力画像サイズが大きくなった際に計算量が指数的に増大する \¥0 2020/175446 3 卩(:171? 2020 /007383
[0009] 本発明はこのような事情に鑑みてなされたもので、 学習時の画像サイズに 制約されることなく、 任意サイズの入カデータに対応でき、 画像生成時の計 算量を抑制することが可能な生成モデルの学習方法及び学習システム、 プロ グラム、 学習済みモデル、 並びに超解像画像生成装置を提供することを目的 とする。
課題を解決するための手段
[0010] 本開示の一態様に係る学習方法は、 第 1画像から第 1画像よりも高解像の 画像情報を含む第 2画像を推定する生成モデルの機械学習を行う学習方法で あって、 生成モデルであるジェネレータと、 与えられたデータが学習用の正 解画像のデータであるかジェネレータからの出力に由来するデータであるか を識別する識別モデルであるディスクリミネータと、 を含む敵対的生成ネッ トワークを用いることと、 第 2画像よりも解像度が低い第 1解像度情報を含 む第 1学習用画像と、 第 1学習用画像よりも解像度が高い第 2解像度情報を 含む第 2学習用画像であって第 1学習用画像に対応する正解画像となる第 2 学習用画像と、 を学習データとして用いることと、 ジェネレータの入力には 、 第 1学習用画像及び第 2学習用画像のうち第 1学習用画像のみを与えるこ とと、 ジェネレータ及びディスクリミネータのうち、 ディスクリミネータの ネッ トワークに限定してセルフアテンション機構を実装することと、 を含む 学習方法である。
[001 1 ] 本態様によれば、 セルフアテンション機構の導入によって、 学習において 画像の大局的な情報が考慮され、 精度の高い学習が行われる。 本態様によれ ば、 ディスクリミネータに限定してセルフアテンション機構を導入したこと により、 かつ、 ジェネレータの計算量を増加することなく、 生成画像の精度 を向上することが可能になる。 また、 ジェネレータはセルフアテンション機 構を備えていないため、 任意サイズの入カデータに対して高精度の画像生成 を行うことができる。
[0012] 本開示の他の態様に係る学習方法において、 ジェネレータ及びディスクリ ミネ _夕のそれぞれのネッ トワ _クは、 畳:み込みニュ _ラルネッ トワ _クで \¥0 2020/175446 4 卩(:171? 2020 /007383
ある構成とすることができる。
[0013] 本開示の更に他の態様に係る学習方法において、 第 1画像は 3次元断層画 像であり、 第 2画像は少なくとも 3次元断層画像のスライス厚方向の解像度 が第 1画像よりも高解像である構成とすることができる。
[0014] 本開示の更に他の態様に係る学習方法において、 第 2学習用画像は、 コン ピュータ断層撮影装置を用いて取得された画像であり、 第 1学習用画像は、 第 2学習用画像を基に画像処理によって生成された画像である構成とするこ とができる。
[0015] 本開示の更に他の態様に係る学習方法において、 第 2学習用画像から第 1 学習用画像を生成する画像処理は、 第 2学習用画像をダウンサンプルする処 理を含む構成とすることができる。
[0016] 本開示の更に他の態様に係る学習方法において、 第 2学習用画像から第 1 学習用画像を生成する画像処理は、 ダウンサンプルの処理によって得られた 画像に補間処理を施してアツプサンプルする処理を含む構成とすることがで きる。
[0017] 本開示の更に他の態様に係る学習方法において、 第 2学習用画像から第 1 学習用画像を生成する画像処理は、 ガウシアンフィルタを用いる平滑化処理 を含む構成とすることができる。
[0018] 本開示の更に他の態様に係る学習方法において、 機械学習に使用する複数 種類の学習データにおける第 1学習用画像及び第 2学習用画像の各々は同一 サイズである構成とすることができる。
[0019] 本開示の更に他の態様に係る学習方法において、 第 2画像は、 高周波成分 の情報を示す高周波成分画像であり、 ジェネレータは、 入力された画像の解 像度を高めるために必要な高周波成分を推定し、 高周波成分の情報を示す高 周波成分画像を出力する構成とすることができる。
[0020] 本開示の更に他の態様に係る学習方法において、 ジェネレータから出力さ れた高周波成分画像と、 ジェネレータに入力された画像とを加算すること、 をさらに含み、 加算によって得られる仮想第 2画像をディスクリミネータの \¥0 2020/175446 5 卩(:171? 2020 /007383
入力に与える構成とすることができる。
[0021 ] 本開示の他の態様に係るプログラムは、 本開示のいずれか一態様に係る学 習方法をコンピュータに実行させるためのプログラムである。
[0022] 本開示の他の態様に係る学習済みモデルは、 本開示のいずれか一態様に係 る学習方法を実施して学習された学習済みモデルであって、 第 1画像から第 1画像よりも高解像の画像情報を含む第 2画像を推定する生成モデルである
[0023] 本開示の他の態様に係る超解像画像生成装置は、 本開示のいずれか一態様 に係る学習方法を実施して学習された学習済みモデルである生成モデルを備 え、 入力される第 3画像から第 3画像よりも高解像の画像情報を含む第 4画 像を生成する超解像画像生成装置である。
[0024] 本態様に係る超解像画像生成装置によれば、 任意サイズの入カデータに対 して高精度の画像生成が可能である。
[0025] 本開示の更に他の態様に係る超解像画像生成装置において、 第 3画像は、 第 1学習用画像と異なる画像サイズである構成とすることができる。
[0026] 本開示の他の態様に係る超解像画像生成装置において、 第 3画像に補間処 理を行い、 補間画像を生成する第 1補間処理部と、 補間画像と生成モデルが 生成する高周波成分とを加算する第 1加算部と、 を含み、 補間画像が生成モ デルに入力され、 生成モデルが補間画像の解像度を高めるために必要な高周 波成分を生成する構成とすることができる。
[0027] 本開示の他の態様に係る学習システムは、 第 1画像から第 1画像よりも高 解像の画像情報を含む第 2画像を推定する生成モデルの機械学習を行う学習 システムであって、 生成モデルであるジェネレータと、 与えられたデータが 学習用の正解画像のデータであるかジェネレータからの出力に由来するデー 夕であるかを識別する識別モデルであるディスクリミネータと、 を含む敵対 的生成ネッ トワークを備え、 ジェネレータ及びディスクリミネータのうち、 ディスクリミネータのネッ トワークに限定してセルフアテンション機構が実 装されており、 第 2画像よりも解像度が低い第 1解像度情報を含む第 1学習 \¥0 2020/175446 6 卩(:171? 2020 /007383
用画像と、 第 1学習用画像よりも解像度が高い第 2解像度情報を含む第 2学 習用画像であって第 1学習用画像に対応する正解画像となる第 2学習用画像 と、 を学習データとして取り込み、 ジェネレータの入力に、 第 1学習用画像 及び第 2学習用画像のうち第 1学習用画像のみが与えられ、 敵対的生成ネッ トワークの学習が行われる学習システムである。
[0028] 本開示の他の態様に係る学習システムにおいて、 学習データを生成する学 習データ生成部をさらに備え、 学習データ生成部は、 第 2解像度情報を含む オリジナルの元画像から固定サイズ領域を切り出す固定サイズ領域切出部と 、 固定サイズ領域切出部によって切り出された固定サイズ領域の画像をダウ ンサンプルするダウンサンプル処理部と、 を含み、 固定サイズ領域切出部に よって切り出された固定サイズ領域の画像を第 2学習用画像とし、 第 2学習 用画像に対してダウンサンプルの処理を行うことによって第 1学習用画像を 生成する構成とすることができる。
[0029] 本開示の更に他の態様に係る学習システムにおいて、 学習データ生成部は 、 さらに、 ダウンサンプルの処理によって得られた画像に補間処理を施す第 2補間処理部と、 ガウシアンフィルタを用いて平滑化を行う平滑化処理部と 、 を含む構成とすることができる。
[0030] 本開示の更に他の態様に係る学習システムにおいて、 ジェネレータは、 入 力された画像の解像度を高めるために必要な高周波成分を推定して高周波成 分の情報を示す高周波成分画像を出力する構成であり、 ジェネレータから出 力された高周波成分画像とジェネレータに入力された画像とを加算する第 2 加算部をさらに備える構成とすることができる。
[0031 ] 本開示の他の態様に係る学習システムは、 第 1画像から第 1画像よりも高 解像の画像情報を含む第 2画像を推定する生成モデルの機械学習を行う学習 システムであって、 少なくとも 1つのプロセッサを含み、 プロセッサは、 生 成モデルであるジェネレータと、 与えられたデータが学習用の正解画像のデ —夕であるかジェネレータからの出力に由来するデータであるかを識別する 識別モデルであるディスクリミネータと、 を含む敵対的生成ネッ トワークを \¥0 2020/175446 7 卩(:171? 2020 /007383
備え、 ジェネレータ及びディスクリミネータのうち、 ディスクリミネータの ネッ トワークに限定してセルフアテンション機構が実装されており、 第 2画 像よりも解像度が低い第 1解像度情報を含む第 1学習用画像と、 第 1学習用 画像よりも解像度が高い第 2解像度情報を含む第 2学習用画像であって第 1 学習用画像に対応する正解画像となる第 2学習用画像と、 を学習データとし て取り込み、 ジェネレータの入力に、 第 1学習用画像及び第 2学習用画像の うち第 1学習用画像のみが与えられ、 敵対的生成ネッ トワークの学習が行わ れる学習システムである。
発明の効果
[0032] 本発明によれば、 任意サイズの入カデータに対して高解像度の画像生成が 可能な生成モデルを得ることができる。 また、 本発明によれば、 画像生成時 の計算量を抑制することが可能な生成モデルを得ることができ、 学習済みモ デルを用いて高精度の画像生成を実現できる。
図面の簡単な説明
Figure imgf000009_0001
门 51 1〇6) 画像のそれぞれのデータのイメージ図である。
[図 2]図 2は、 〇丁画像のスライス間隔及びスライス厚を説明するための図で ある。
[図 3]図 3は、 〇丁画像のスライス間隔及びスライス厚を説明するための図で ある。
[図 4]図 4は、 〇丁画像のスライス間隔及びスライス厚を説明するための図で ある。
[図 5]図 5は、 〇丁画像のスライス間隔及びスライス厚を説明するための図で ある。
[図 6]図 6は、 本発明の実施形態に係る超解像画像生成装置の例を示す機能ブ ロック図である。
[図 7]図 7は、 本発明の実施形態に係る学習システムの構成例を示すブロック 図である。 \¥0 2020/175446 8 卩(:171? 2020 /007383
[図 8]図 8は学習データ生成部の構成例を示す機能ブロック図である。
[図 9]図 9は、 学習データを生成する際に適用されるスライス間隔と想定スラ イス厚に対応したガウシアンフィルタの条件の組み合わせの例を示す図表で ある。
[図 10]図 1 0は、 学習データを生成する処理の手順の例を示すフローチヤー 卜である。
[図 1 1]図 1 1は、 ◦八 1\1を適用した学習部における処理の概念図である。
[図 12]図 1 2は、 ディスクリミネータに適用される識別ネッ トワークの例を 示す概念図である。
[図 13]図 1 3は、 ジェネレータに適用される生成ネッ トワークの例を示す概 念図である。
[図 14]図 1 4は、 ジェネレータの出力に低解像度画像を加えて仮想高解像度 画像を生成する動作の説明図である。
[図 15]図 1 5は、 学習時におけるディスクリミネータによる識別の動作を説 明するための図である。
[図 16]図 1 6は、 学習部における処理の手順の例を示すフローチヤートであ る。
[図 17]図 1 7は、 実施形態の効果を示す画像の例である。
[図 18]図 1 8は、 実施形態の他の効果を説明するための図である。
[図 19]図 1 9は、 第 2実施形態に係る学習システムによる処理の流れを概略 的に示す機能ブロック図である。
[図 20]図 2 0は、 第 3実施形態に係る学習システムによる処理の流れを概略 的に示す機能ブロック図である。
[図 21]図 2 1は、 第 4実施形態に係る学習システムによる処理の流れを概略 的に示す機能ブロック図である。
[図 22]図 2 2は、 コンピユータのハードウェア構成の例を示すブロック図で ある。
発明を実施するための形態 [0034] 以下、 添付図面に従って本発明の好ましい実施の形態について詳説する。
[0035] 《第 1実施形態》
本発明の実施形態に係る超解像画像生成装置は、 低解像度の画像データか ら仮想的な高解像度の画像データを生成する。 「生成する」 とは 「推定する 」 という用語の概念を含む。 ここでは画像データの具体例として、 コンピュ —夕断層撮影 (CT : Computed Tomography) 装置を用いて取得される C T画 像のデータを対象とし、 CT装置を用いて取得されたシックスライス (Thick slice) の画像データから仮想的なシンスライス (Thin slice) の画像デー 夕を生成する超解像画像生成装置を例示する。
[0036] シックスライスの画像データとは、 スライス間隔及びスライス厚が比較的 大きい低解像度の CT画像データをいう。 例えば、 スライス間隔及びスライ ス厚が 4 m mを超える C T画像データはシックスライスの画像データに該当 する。 シックスライスの画像データを 「シックスライス画像」 、 「シックス ライスデータ」 、 又は 「シックデータ」 と表記する場合がある。
[0037] シンスライスの画像データとは、 スライス間隔及びスライス厚が小さい高 解像度の CT画像データである。 例えば、 スライス間隔及びスライス厚が 1 mm程度の C T画像データはシンスライスの画像データに該当する。 シンス ライスの画像データを 「シンスライス画像」 、 「シンスライスデータ」 、 又 は 「シンデータ」 と表記する場合がある。
[0038] 本実施形態において、 シックスライス画像から生成される仮想的なシンス ライス画像を仮想シンスライス (VTS : Virtual Thin Slice) 画像という 。 これに対し、 C T装置を用いた撮影によって取得された本物のシンスライ ス画像をリアルシンスライス (RTS : Real Thin SI ice) 画像と呼ぶ。
[0039] [CT画像データの説明]
図 1は、 シックスライス画像と VTS画像のそれぞれのデータのイメージ 図である。 図 1の左図がシックスライス画像であり、 右図が VTS画像であ る。 VTS画像は、 シックスライス画像に比べて高品質な再構築画像を生成 することが可能である。 図 1 において z軸方向は体軸方向である。 [0040] CTデータは、 撮影に使用した CT装置の機種により、 また、 出カスライ スの条件の設定などにより、 様々なスライス間隔及びスライス厚のデータが 存在し得る。
[0041] 図 2〜図 5は、 CT画像のスライス間隔及びスライス厚の例を説明するた めの図である。 スライス間隔とは、 あるスライスとそれに隣接するスライス とのそれぞれの厚さの中心位置同士間の距離をいう。 スライス間隔はスライ ス間距離と同義である。 スライス厚とは、 撮影領域の中心位置における 1つ のスライスの厚さ方向の長さをいう。 スライス厚は、 スライスシックネス (S I ice thickness) と同義である。 図 2〜図 5において、 スライス間隔を S D と表示し、 スライス厚を S Tと表示する。 なお、 スライスの厚み方向は Z軸 方向である。
[0042] 図 2は、 スライス間隔 S D = 4 m m、 スライス厚 S T = 4 m mの場合の C
T画像丨 M 1 を模式的に示す説明図である。 ここでは簡単のために 3層の断 層画像群を模式的に示している。
[0043] 図 3は、 スライス間隔 S D = 4 m m、 スライス厚 S T = 6 m mの場合の C
T画像 I M2を模式的に示す説明図である。 図 3の場合、 隣り合うスライス 同士でスライス厚の範囲が才ーバーラップしている。
[0044] 図 4は、 スライス間隔 S D = 8 m m、 スライス厚 S T = 4 m mの場合の C
T画像丨 M3を模式的に示す説明図である。 図 4の例の場合、 スライス厚 S Tよりもスライス間隔 S Dの方が大きいため、 隣り合う断層画像同士が離間 し、 層間に隙間がある。
[0045] 図 5は、 スライス間隔 S D = 1 m m、 スライス厚 S T = 1 m mの C T画像
I M 4を模式的に示す説明図である。 図 5に示す CT画像丨 m4は、 図 2か ら図 4に示した他の CT画像丨 M 1〜丨 M3よりも Z方向の情報量が多い。 すなわち、 〇丁画像丨 1\/14は、 CT画像 I M 1、 丨 M2、 及び丨 M3のいず れよりも Z方向の解像度が相対的に高い。
[0046] CT画像のスライス間隔及びスライス厚は、 CT装置を使用する施設、 医 師等の好みなどに応じて様々な条件で設定される。 CT画像は、 診断のため \¥0 2020/175446 1 1 卩(:171? 2020 /007383
には高解像度であることが好ましいが、 スライス間隔を小さくすると、 被検 者に対する被ばく量が増えてしまうという問題点がある。 また、 高解像度の 〇丁画像はデータ量が大きく、 ストレージの記憶容量を圧迫するため、 容量 削減のために低解像度化して保存される場合もある。 例えば、 古い 0丁デー 夕は撮影スライス枚数を削減してデータベースに保存することが行われてい る。
[0047] しかし、 シックスライス画像は、 体軸と平行な面を断面とする側面方向か ら見た再構築画像やボリュームレンダリング画像において品質が悪く、 十分 な観察や解析に利用し難いという課題がある。
[0048] 本実施形態に係る超解像画像生成装置は、 図 3〜図 5に示すような様々な スライス条件 (スライス間隔及びスライス厚) の低解像度の <3丁画像から、 例えば、 図 5に示すようなスライス間隔が 1 01 111、 スライス厚が
Figure imgf000013_0001
解像度の V丁 3画像を生成する画像生成処理を行う。
[0049] [超解像画像生成装置における画像生成アルゴリズムの例]
図 6は、 本発明の実施形態に係る超解像画像生成装置の例を示す機能ブロ ック図である。 超解像画像生成装置 1 〇は、 補間処理部 1 2と、 階層型ニュ —ラルネッ トワークの学習済みモデルであるジェネレータ 1 4と、 加算部 1 6と、 を含む。 「ニューラルネッ トワーク」 とは、 脳神経系の仕組みを模擬 した情報処理の数理モデルである。 ニューラルネッ トワークを用いた処理は 、 コンビュータを用いて実現することができる。 ニューラルネッ トワークは 、 プログラムモジュールとして構成され得る。 本明細書においてニューラル ネッ トワークを単に 「ネッ トワーク」 と表記する場合がある。
[0050] 補間処理部 1 2は、 入力された低解像度のシックスライス画像丁(3 <に対 してスプライン補間を行い、 補間画像丨 丁を生成する。 補間処理部 1 2か ら出力される補間画像丨 丁は、 方向にボケた画像であり、 低解像度画像 の一例である。 なお、 補間画像丨 丁の画素数は、 最終的に生成する仮想シ ンスライス画像 丁の画素数と一致させておくことが好ましい。
[0051 ] 補間処理部 1 2から出力された補間画像丨 丁は、 ジェネレータ 1 4に入 \¥0 2020/175446 12 卩(:171? 2020 /007383
力される。 ジェネレータ 1 4は、 敵対的生成ネッ トヮーク (〇八1\!) を用い た機械学習によって学習された生成モデルである。 ジェネレータ 1 4を得る ための学習方法については後述する。 学習済みモデルは、 プログラムモジュ —ルと言い換えてもよい。
[0052] ジェネレータ 1 4は、 入力された画像から高解像度画像の生成に必要な高 周波成分情報を生成 (推定) し、 高周波成分情報を出力する。
[0053] 加算部 1 6は、 ジェネレータ 1 4から出力された高周波成分情報のマップ とジェネレータ 1 4の入カデータである補間画像 I 丁そのものとを加算し て、 仮想シンスライス画像 V丁を生成する。
[0054] 図 6では、 ジェネレータ 1 4への入力が補間画像丨 丁であり、 ジェネレ —夕 1 4の出力が高周波成分情報である例を示すが、 ジェネレータ 1 4への 入力がシックスライス画像丁〇[<である形態も可能である。 また、 ジェネレ —夕 1 4の出力が仮想シンスライス画像 丁である形態も可能である。 高周 波成分情報は、 元となる画像と加算することによって高解像度画像を生成す ることができる情報であるため、 高周波成分情報のマップを 「高周波成分画 像」 と呼ぶ。 高周波成分画像は、 高解像度の画像情報を含む画像であり、 実 質的に 「高解像度画像」 と同様のものとして理解することができる。
[0055] 図 6においてシックスライス画像丁〇<は本開示における 「第 3画像」 の —例である。 仮想シンスライス画像 V丁は本開示における 「第 4画像」 の一 例である。 ジェネレータ 1 4から出力される高周波成分は本開示における 「 前記第 3画像よりも高解像の画像情報」 の一例である。 補間処理部 1 2は本 開示における 「第 1補間処理部」 の一例である。 加算部 1 6は本開示におけ る 「第 1加算部」 の—例である。
[0056] [学習システムの構成例]
次に、 ジェネレータ 1 4を生成するための学習方法について説明する。
[0057] 図 7は、 本発明の実施形態に係る学習システム 2 0の構成例を示すブロッ ク図である。 学習システム 2 0は、 画像保管部 2 4と、 学習データ生成部 3 0と、 学習部 4 0と、 を含む。 学習システム 2 0は、 1台又は複数台のコン ピュータを含むコンビュータシステムによって実現することができる。 すな わち、 画像保管部 24、 学習データ生成部 30、 及び学習部 40の機能は、 コンビュータのハードウェアとソフトウェアの組み合わせによって実現でき る。 ここでは、 画像保管部 24、 学習データ生成部 30、 及び学習部 40の 各々が別々の装置として構成される例を説明するが、 これらの機能は 1台の コンビュータで実現してもよいし、 2以上の複数台のコンビュータで処理の 機能を分担して実現してもよい。 例えば、 画像保管部 24、 学習データ生成 部 30、 及び学習部 40は、 通信回線を介して互いに接続されていてもよい 。 「接続」 という用語は、 有線接続に限らず、 無線接続の概念も含む。 通信 回線は、 口ーカルエリアネッ トワークであってもよいし、 ワイ ドエリアネッ トワークであってもよい。
[0058] このように構成することで、 学習データの生成と生成モデルの学習とを物 理的にも時間的にも互いに束縛されることなく実施することができる。
[0059] 画像保管部 24は、 医療用 X線 CT装置によって撮影された CT再構成画 像 (CT画像) を保存する大容量ストレージ装置を含む。 画像保管部 24は 、 例えば、 PACS (Picture Archiving and Communication Systems)に代 表される医用画像管理システムにおけるストレージであってよい。 画像保管 部 24には、 不図示の CT装置を用いて撮影されたリアル高解像度画像であ る複数のシンスライス画像のデータが保管されている。
[0060] 画像保管部 24に保管される CT画像は、 人体 (被検体) を撮影した医療 画像であり、 複数の断層画像を含む 3次元断層画像である。 ここでは、 各断 層画像は互いに直交する X方向及び Y方向に平行な画像である。 X方向及び Y方向に直交する Z方向は、 被検体の体軸方向であり、 スライス厚方向とも いう。 画像保管部 24に保管される CT画像は、 人体の部位毎の画像であっ てもよいし、 全身を撮影した画像であってもよい。
[0061] 学習データ生成部 30は、 学習部 40が学習を行うために必要な学習デー 夕を生成する。 学習データとは、 機械学習に用いる訓練用のデータであり、 「学習用データ」 或いは 「訓練データ」 と同義である。 本実施形態の機械学 \¥0 2020/175446 14 卩(:171? 2020 /007383
習においては、 入力用の低解像度画像と、 その低解像度画像に対応する正解 の高解像度画像と、 を紐付けした画像ペアの学習データを多数使用する。 こ のような画像ペアは、 リアル高解像度画像であるシンスライスのデータを元 に、 画像処理によって人工的に生成することが可能である。
[0062] 学習データ生成部 3 0は、 画像保管部 2 4からオリジナルのリアル高解像 度画像を取得し、 リアル高解像度画像にダウンサンプルの処理を実施するこ とにより、 多様な低解像度画像 (擬似的なシックスライス画像) を人工的に 生成する。 学習データ生成部 3 0は、 例えば、 1
Figure imgf000016_0001
に等方化したオリジナ ルのシンスライスのデータに対して、 姿勢変換を行い、 無作為に固定サイズ 領域を切り出した後、 スライス間隔が 4
Figure imgf000016_0003
の仮想的な
Figure imgf000016_0002
スライスのデ —夕、 及びスライス間隔が 8〇!〇!の仮想的な 8〇!〇!スライスのデータを生成 する。 固定サイズ領域は、 X軸方向 X V軸方向 X 軸方向の画素数が、 例え ば 「160 X 160 X 160」 の 3次元領域であってよい。 学習データ生成部 3 0によ って、 学習用の固定サイズの低解像度画像 !_〇とこれに対応するリアル高解 像度画像 1~1の画像ペアが生成される。
[0063] 学習部 4 0による学習の処理を実施するために、 事前に学習データ生成部
3 0を用いてオリジナルのリアル高解像度画像から複数の学習データを生成 しておき、 学習データセッ トとしてストレージに保存しておくことが好まし い。
[0064] 学習データ生成部 3 0によって生成された低解像度画像 !_ 0及びリアル高 解像度画像
Figure imgf000016_0004
学習部 4 0に入力される。
[0065] 学習部 4 0は、 学習モデルとしての敵対的生成ネッ トワーク (〇八1\1) 4
1 を含む。 学習部 4 0のアーキテクチャは、 非特許文献 2に記載のアーキテ クチャを 2次元から 3次元のデータへ拡張した構造をベースとしている。 ◦ 八 4 1は、 データを作り出すジェネレータ 4 2◦と呼ばれる生成ネッ トワ —クと、 入力されたデータを識別するディスクリミネータ 4 4 0と呼ばれる 識別ネッ トワークと、 を含んで構成される。 すなわち、 ジェネレータ 4 2 0 は、 画像データを生成する生成モデルであり、 ディスクリミネータ 4 4 0は \¥0 2020/175446 15 卩(:171? 2020 /007383
データを識別する識別モデルである。 「ジェネレータ」 という用語は 「生成 部」 、 「生成器」 及び 「生成モデル」 などの用語と同義である。 「ディスク リミネータ」 という用語は 「識別部」 、 「識別器」 及び 「識別モデル」 など の用語と同義である。
[0066] 学習部 4 0は、 入力された学習データに基づいて、 ジェネレータ 4 2〇と ディスクリミネータ 4 4とを用いた敵対的な学習を繰り返すことにより、 双 方のモデルの性能を高めながらジェネレータ 4 2◦を学習する。
[0067] 本例のディスクリミネータ 4 4 0には、 セルフアテンション機構が実装さ れている。 ディスクリミネータ 4 4 0のネッ トワークにおいてセルフアテン ション機構を導入する層は、 複数の畳み込み層のうちの一部であってもよい し、 全部であってもよい。 セルフアテンション機構を含むディスクリミネー 夕 4 4口の構成及び動作、 並びに◦ 4 1の学習方法の例について詳細は 後述する。
[0068] 学習部 4 0は、 誤差演算部 5 0と、 オプティマイザ 5 2と、 を含む。 誤差 演算部 5 0は、 損失関数を用いてディスクリミネータ 4 4 0の出力と正解と の誤差を評価する。 オプティマイザ 5 2は、 誤差演算部 5 0の演算結果を基 に、 ネッ トワークのパラメータを更新する処理を行う。 ネッ トワークのパラ メータは、 各層の処理に用いるフィルタのフィルタ係数 (ノード間の結合の 重み) 及びノードのバイアスなどを含む。
[0069] オプティマイザ 5 2は、 誤差演算部 5 0の演算結果からジヱネレータ 4 2 ◦及びディスクリミネータ 4 4 0のそれぞれのネッ トワークのパラメータの 更新量を算出するパラメータ演算処理と、 パラメータ演算部の算出結果に従 い、 ジェネレータ 4 2◦及びディスクリミネータ 4 4 0のそれぞれのネッ ト ワークのパラメータを更新するパラメータ更新処理と、 を行う。 オプティマ イザ 5 2は、 勾配降下法などのアルゴリズムに基づきパラメータの更新を行 う。
[0070] [学習データの生成について]
図 8は学習データ生成部 3 0の構成例を示す機能ブロック図である。 学習 \¥0 2020/175446 16 卩(:171? 2020 /007383
データ生成部 3 0は、 固定サイズ領域切出部 3 1 と、 ダウンサンプル処理部 3 2と、 アップサンプル処理部 3 4と、 学習データ記憶部 3 8と、 を含む。
[0071 ] 固定サイズ領域切出部 3 1は、 入力されたオリジナルのリアル高解像度画
1から無作為に固定サイズ領域を切り出す処理を行う。 固定サイズ 領域切出部 3 1 によって切り出された固定サイズ領域のリアル高解像度画像 ダウンサンプル処理部 3 2に送られる。
[0072] ダウンサンプル処理部 3 2は、 リアル高解像度画像 1~1 1 を 軸方向にダ ウンサンプルして、 低解像度のシックスライス画像
Figure imgf000018_0001
1 を生成する。 ダウ ンサンプルの処理としては、 例えば、 単純に 軸方向のスライスを一定の割 合で削減するように間引き処理を実施すればよい。 なお、 この例では 軸方 向のダウンサンプルのみを行い、 X軸方向及び丫軸方向についてはダウンサ ンプルを行わないものとするが、 X軸方向及び丫軸方向についてもダウンサ ンプルを実施する形態も可能である。
[0073] ダウンサンプル処理部 3 2によって生成されたシックスライス画像
Figure imgf000018_0002
1 は、 アップサンプル処理部 3 4に入力される。
[0074] アップサンプル処理部 3 4は、 シックスライス画像
Figure imgf000018_0003
1 を 軸方向にア ップサンプルして、 低品質のシンスライス画像である低解像度画像 !_〇 1 を 生成する。 アップサンプルの処理は、 例えば、 スプライン補間とガウシアン フィルタ処理との組み合わせであってよい。 アップサンプル処理部 3 4は、 補間処理部 3 5と、 ガウシアンフィルタ処理部 3 6と、 を含む。 補間処理部 3 5は、 例えば、 シックスライス画像
Figure imgf000018_0004
1 に対してスプライン補間を行う 。 補間処理部 3 5は、 図 6で説明した補間処理部 1 2と同様の処理部であっ てよい。 ガウシアンフィルタ処理部 3 6は、 補間処理部 3 5から出力された 画像にガウシアンフィルタを適用して平滑化を行う。 図 8に示す補間処理部 3 5は本開示における 「第 2補間処理部」 の一例である。 ガウシアンフィル 夕処理部 3 6は本開示における 「平滑化処理部」 の一例である。
[0075] アップサンプル処理部 3 4から出力される低解像度画像 !_〇 1は、 リアル 高解像度画像
Figure imgf000018_0005
1 と同じ画素数のデータとすることが好ましい。 ここでは \¥0 2020/175446 17 卩(:171? 2020 /007383
、 低解像度画像 !- 0 1 とリアル高解像度画像 1~1 1は同ーサイズである。 低 解像度画像 !_〇 1は、 リアル高解像度画像 [¾ ! ! 1 と比較して低品質の (つま り、 低解像度の) 画像である。 こうして生成された低解像度画像 !_〇 1 と、 その生成元となったリアル高解像度画像 1~1 1 とのペアを紐付けして学習デ —夕記憶部 3 8に記憶する。
[0076] オリジナルのリアル高解像度画像〇 1~1 1は本開示における 「オリジナル の元画像」 の一例である。 リアル高解像度画像 1~1 1は本開示における 「第 2学習用画像」 の一例である。 低解像度画像 !_〇 1は本開示における 「第 1 学習用画像」 の一例である。 低解像度画像 !_〇 1の画像情報は本開示におけ る 「第 1解像度情報」 の一例である。 リアル高解像度画像 1~1 1の画像情報 は本開示における 「第 2解像度情報」 の一例である。
[0077] 学習データ生成部 3 0は、 1つのオリジナルのリアル高解像度画像〇 1~1
1から固定サイズ領域の切り出し位置を変えて、 複数のリアル高解像度画像
Figure imgf000019_0001
それぞれのリアル高解像度画像 1~1に対応する低解像度 画像 !_ 0を生成することにより、 複数の画像べアを生成することができる。
[0078] また、 学習データ生成部 3 0は、 アップサンプル処理部 3 4におけるスラ イス補間倍率と、 アップサンプル処理部 3 4に適用するガウシアンフィルタ の条件との組み合わせを変えることにより、 多様なスライス条件の低解像度 画像を生成することができる。 なお、 スライス補間倍率は、 ダウンサンプル 処理部 3 2におけるダウンサンプルの条件に対応している。
[0079] 学習の際には多様なスライス条件のデータを与えることが好ましい。 本実 施形態では、 図 9に示すような、 多様なスライス条件に対応する低解像度画 像を用いて学習を行う。 図 9は、 学習データを生成する際に適用されるスラ イス間隔と想定スライス厚に対応したガウシアンフィルタの条件の組み合わ せの例を示す図表である。
[0080] 本例では、 低解像度画像 !_ 0のスライス間隔は、 4〇1 111と8 |11 111の 2通り とする。 つまり、 学習時のスライス補間倍率は、 4倍か 8倍かの 2パターン である。 スライス厚は、 スライス間隔に対応させて 0
Figure imgf000019_0002
の範囲と \¥0 2020/175446 18 卩(:171? 2020 /007383
する。 ガウシアンフィルタの標準偏差 £7を図 9に記載の数値範囲無いえラン ダムに与えることで、 擬似的に多様なスライス厚を想定した低解像度画像が 生成され得る。
[0081 ] オリジナルのリアル高解像度画像を複数種類用いることで多様な学習デー 夕を多数用意することが可能である。
[0082] [学習データを生成する処理の手順の例]
図 1 0は、 学習データを生成する処理の手順の例を示すフローチヤートで ある。 図 1 0に示すフローチヤートの各ステップは、 学習データ生成部 3 0 として機能するプロセッサを含むコンビユータによって実行される。 コンビ ユータは、
Figure imgf000020_0001
及びメモリを備える。 コンピ ユータは、 〇 11 (0「8卩11 _1〇3 卩「00633 _1叩 11门丨1;) を含んでもよい。
[0083] 図 1 0に示すように、 学習データ生成方法は、 オリジナル画像取得工程 ( ステップ 3 1) 、 固定サイズ領域切出工程 (ステップ 3 2) 、 ダウンサンプ ルエ程 (ステップ 3 3) 、 アップサンプルエ程 (ステップ 3 4) 、 及び学習 データ記憶工程 (ステップ 3 5) を含む。
[0084] ステップ 3 1 において、 学習データ生成部 3 0は画像保管部 2 4からオリ ジナルのリアル高解像度画像〇 1~1を取得する。 ここでは、 スライス間隔が 1 01 111、 スライス厚が 1 〇! の等方化されたリアル高解像度画像〇 1~1を取 得する。
[0085] ステップ 3 2において、 固定サイズ領域切出部 3 1は、 入力されたオリジ ナルのリアル高解像度画像〇 1~1から固定サイズ領域を切り出す処理を行い 、 固定サイズ領域のリアル高解像度画像
Figure imgf000020_0002
1 を生成する。
[0086] ステップ 3 3において、 ダウンサンプル処理部 3 2はリアル高解像度画像
Figure imgf000020_0003
、 1 を生成する。 ここ では、 図 9で説明したように、 スライス間隔が 4〇1 111、
Figure imgf000020_0004
に相当す るシックスライス画像
Figure imgf000020_0005
が生成される。
[0087] ステップ 3 4において、 アップサンプル処理部 3 4はダウンサンプルによ って得られたシックスライス画像 1 をアップサンプルして、 低品質のシ \¥0 2020/175446 19 卩(:17 2020 /007383
ンスライス画像に相当する低解像度画像 !_ 0 1 を生成する。 ここでは、 図 9 で説明したように、 スライス間隔に対応したスライス補間倍率とガウシアン フィルタの条件を適用して補間処理とガウシアンフィルタ処理とが行われる
[0088] ステップ 3 5において、 学習データ生成部 3 0はステップ 3 4にて生成さ れた低解像度画像 !- 0 1 とその生成元データであるリアル高解像度画像 1~1 とを画像ペアとして紐付けし、 これらのデータを学習データとして学習デー 夕記憶部 3 8に記憶する。
[0089] ステップ 3 5の後、 学習データ生成部 3 0は、 図 8のフローチャートを終 了する。
[0090] なお、 同じオリジナルのリアル高解像度画像〇 1~1から切出領域の箇所を 変えて複数の学習データを生成する場合には、 ステップ 3 5の後に、 ステッ プ3 2に戻り、 ステップ 3 2からステップ 3 5の処理を繰り返す。
[0091 ] また、 同じ固定サイズ領域のリアル高解像度画像
Figure imgf000021_0001
から異なるスライス 条件又は異なる想定スライス厚の低解像度画像を生成する場合には、 ステッ プ3 5の後に、 ステップ 3 3又はステップ 3 4に戻り、 処理の条件を変更し て、 ステップ 3 3又はステップ 3 4からの処理を繰り返す。
[0092] 学習データ生成部 3 0は、 画像保管部 2 4に保管されている複数のオリジ ナルのリアル高解像度画像に対して、 ステップ 3 1からステップ 3 5の処理 を繰り返し実行することにより、 多数の学習データを生成することができる
[0093] [学習アーキテクチャ]
既述のとおり、 本実施形態に係る超解像画像生成装置 1 0に搭載されるジ エネレータ 1 4は、 〇八 1\1による学習を実施して獲られる生成モデルである 。 以下、 学習部 4 0の構成と学習方法について詳述する。
[0094] 図 1 1は、 〇八 1\1を適用した学習部 4 0における処理の概念図である。 図
1 1 には、 学習用データとして、 低解像度画像 1- 0 1 とリアル高解像度画像 1のペアが学習部 4 0に入力された例が示されている。 \¥0 2020/175446 20 卩(:171? 2020 /007383
[0095] ジェネレータ 4 2◦への入力は低解像度画像 !_〇 1である。 ジェネレータ 4 2◦は、 入力された低解像度画像 !_ 0 1から仮想高解像度画像 V 1~1 1 を生 成して出力する。 仮想高解像度画像
Figure imgf000022_0001
1は、 仮想シンスライス画像 ( 丁 3画像) に相当する。 ディスクリミネータ 4 4 0への入力には、 ジェネレー 夕 4 2〇によって生成された仮想高解像度画像▽! ! 1 と、 この仮想高解像度 画像▽! ! 1の生成元となった低解像度画像 !_〇 1のペア、 又は、 学習データ であるリアル高解像度画像 1~1 1 と低解像度画像 !_ 0 1のペアが与えられる
[0096] ディスクリミネータ 4 4 0は、 入力された画像ペアがリアル高解像度画像
Figure imgf000022_0002
3 丨ペア) であるか (学習データであるか) 、 ジェネレータ 4 2〇の出力に由来する仮想高解像度画像▽! ! 1 を含む偽物べ ア ( 3 1< 6ペア) であるかを識別し、 識別結果を出力する。
[0097] 誤差演算部 5 0は、 損失関数を用いてディスクリミネータ 4 4 0の出力と 正解との誤差を評価する。 オプティマイザ 5 2は、 誤差演算部 5 0の演算結 果を基に、 ネッ トワークのパラメータを自動調整する処理を行う。 ネッ トワ —クのパラメータには、 ノード間の結合の重みとノードのバイアスが含まれ る。 オプティマイザ 5 2は、 誤差演算部 5 0の演算結果からジヱネレータ 4 2◦及びディスクリミネータ 4 4 0のそれぞれのネッ トワークのパラメータ の更新量を算出するパラメータ演算処理と、 パラメータ演算部の算出結果に 従い、 ジェネレータ 4 2◦及びディスクリミネータ 4 4 0のそれぞれのネッ トワークのパラメータを更新するパラメータ更新処理と、 を行う。 オプティ マイザ 5 2は、 勾配降下法などのアルゴリズムに基づきパラメータの更新を 行う。 誤差の評価とパラメータの更新に関する学習の基本的な仕組みの部分 は非特許文献 1等に記載の技術を採用してよい。
[0098] ジェネレータ 4 2 0は、 ディスクリミネータ 4 4 0を欺くように、 より精 織な仮想高解像度画像を生成するように学習し、 ディスクリミネータ 4 4 0 はより正確に真偽を識別するように学習する。
[0099] そして、 最終的には、 ジヱネレータ 4 2 0の部分を超解像画像生成装置 1 \¥0 2020/175446 21 卩(:171? 2020 /007383
0における画像生成モジュールであるジェネレータ 1 4として利用する。
[0100] 本実施形態におけるディスクリミネータ 4 4 0に適用されるネッ トワーク には、 セルフアテンション機構が実装される。 セルフアテンション機構は、 画像内における大局的な部分を考慮することで計算効率を向上させる手法で ある。
[0101 ] [セルフアテンション機構を含むディスクリミネータ 4 4口の説明] セルフアテンション機構の内容は、 非特許文献 3に記載されている。 ただ し、 非特許文献 3では、 ジェネレータとディスクリミネータの両方のネッ ト ワークにそれぞれセルフアテンション機構を追加しているのに対し、 本実施 形態ではジェネレータ 4 2◦にはセルフアテンション機構を実装せず、 ディ スクリミネータ 4 4口に限定してセルフアテンション機構を実装している点 で非特許文献 3に記載の手法と異なる。
[0102] セルフアテンション機構について、 非特許文献 3の内容を参照して簡単に 概説する。 セルフアテンション機構は、 前層の隠れ層から出力された畳み込 み特徴マップ〇 IV! ( X ) からクエリ 干 ( X ) とキー9 ( X ) を生成し、 こ れらを用いて各画素について、 他のどの画素に似ているかを示す値 (類似度 ) を計算する。 こうして特徴マップ〇 IV! ( X ) の全画素に対応して計算さ れた類似度のマップが 「アテンションマップ」 と呼ばれる。
[0103] アテンションマップは、 画像内において特徴が似ている領域を見つけ出し て強調する役割を果たす。 識別ネッ トワークを構成する畳み込み層の畳み込 み演算では、 局所的な情報を重ねていくが、 アテンションマップを導入する ことで大局的 (全域的) な部分の情報を考慮することが可能になる。
[0104] このアテンションマップに重み II ( X ) を掛け合わせて、 セルフアテンシ ョン特徴マップ 3八 1\/1 (〇) を得る。 そして、 セルフアテンション特徴マ ップ 3八 1\/1 (〇) にスケールパラメータ· ^を掛けて、 元の入力特徴マップ である畳み込み特徴マップ〇 IV! ( X ) に足し合わせて次の層へ渡す。 つま り、 次層に渡す最終的な出力 Vは次式で与えられる。
[0105] )/ = 7 〇 + X このようなセルフアテンション機構を含むディスクリミネータ 4 4 Dのネ ッ トワークにおいては、 セルフアテンション機構の f ( X ) 、 g ( X ) 、 及 び h ( X ) のパラメータも学習される。
[0106] [識別ネッ トワークの例]
図 1 2は、 ディスクリミネータ 4 4 Dに適用される識別ネッ トワークの例 を示す概念図である。 ディスクリミネータ 4 4 Dのネッ トワークは、 深層二 ユーラルネッ トワークに分類される階層型ニユーラルネッ トワークであり、 複数の畳:み込み層を含む。 ディスクリミネータ 4 4 Dのネッ トワークは畳:み 込みニユーラルネッ トワーク ( C N N : Convo lut i ona l Neura l Network) に よって構成される。
[0107] 図 1 2において C O 1、 C 0 2 C 0 5の符号で示す白抜き矢印は 「 畳:み込み層」 を表している。 各層の入力側及び/又は出力側に示す矩形は、 特徴マップのセッ トを表している。 矩形の縦方向の長さは、 特徴マップのサ イズ (画素数) を表しており、 矩形の横方向の幅はチャンネル数を表してい る。 なお、 本例のディスクリミネータ 4 4 Dは、 プーリング層が存在せず、 例えば、 4 X 4 X 4のサイズのフィルタの畳み込みをストライ ド = 2で実施 することにより、 特徴マップの画像サイズが小さくなっていく。 例えば、 C N Nの処理を実施する際、 畳み込み後の最小画像サイズは入カデータのサイ ズの 1 / 1 6とすることができる。
[0108] 図 1 2に示す例では、 畳み込み層 C 0 2から後段の各層にセルフアテンシ ョン機構が導入されている。 例えば、 畳み込み層 C 0 2から出力された 1 2 8チヤンネルの C N N特徴マップの各々に対してセルフアテンション特徴マ ップが生成される。 図 1 2には 1 2 8チャンネルの C N N特徴マップの各々 に対応する 1 2 8チャンネル分のセルフアテンション特徴マップが付加され ている様子が示されている。 チャンネル毎にそれぞれの C N N特徴マップと セルフアテンション特徴マップが加算され、 その出力が次の畳み込み層に入 力される。 畳み込み層 C 0 3及び C 0 4についても同様である。
[0109] なお、 セルフアテンション機構に入力させる C N N特徴マップは、 入力と \¥0 2020/175446 23 卩(:171? 2020 /007383
なる画像全体を 1次元の配列に直して計算する。 入カチャンネル数が<3、 総 ピクセル数が
Figure imgf000025_0001
〇 1\1個の各ピクセルの要素を 1次元に配列したべクトルとしてセルフアテンシヨン機構に入力される。
[01 10] 実際の〇丁画像データは 3次元データであり、 多次元のデータは上記と同 様に 1次元の配列にして計算を行うことができる。 2次元の画像データと 3 次元の画像データは、 どちらも 1次元の配列にして計算することで、 同様の 処理アルゴリズムを適用できる。
[01 1 1 ] [生成ネッ トワークの例]
図 1 3は、 ジェネレータ 4 2〇に適用される生成ネッ トワークの例を示す 概念図である。 ジェネレータ 4 2 0のネッ トワークも畳:み込みニューラルネ ッ トワークで構成される。 ジェネレータ 4 2 0は、 エンコーダ部とデコーダ 部とを組み合わせたエンコーダーデコーダ構造を持つ構成が好ましい。 図 1 3では、 □- N 6 1:構造と呼ばれる II字型のネッ トワークの例が示されてい る。 「リー ㊀ 」 の表記における 「N 6 」 は 「ネッ トワーク ( 士1/\/〇「1〇 」 の簡易表記である。
[01 12] 図 1 3において〇 1、 〇2 〇 1 0の符号で示す矢印の各々は 「畳み 込み層」 を表している。 リ 1、 11 2、 II 3及び II 4の符号で示す矢印は 「畳 み込みとアップサンプリング」 を行う畳み込み層を表している。 図 1 2で説 明したディスクリミネータ 4 4 0と同様に、 図 1 3に示すジェネレータ 4 2 ◦は、 プーリング層が存在せず、 フィルタの畳み込みをストライ ド = 2で実 施することにより、 エンコーダ部分において特徴マップの画像サイズが小さ くなっていく。
[01 13] ジェネレータ 4 2◦は、 入力された低解像度画像 !_〇から高解像化に必要 な高解像度情報としての高周波成分画像
Figure imgf000025_0002
図 1 4に示すように、 ジェネレータ 4 2◦への入カデータである低解像度画像 !_ 〇と、 ジェネレータ 4 2◦によって生成された高周波成分画像 V 1~1 〇とを 足し合わせることにより、 仮想高解像度画像
Figure imgf000025_0003
が得られる。 なお、 仮想高 解像度画像
Figure imgf000025_0004
1のスライス間隔及びスライス厚は、 低解像度画像 !_〇 1の \¥0 2020/175446 24 卩(:171? 2020 /007383
スライス間隔及びスライス厚と同等であるが、 仮想高解像度画像
Figure imgf000026_0001
1は、 低解像度画像 !_ 0 1 と比較して 方向によりシャープな画像となる。
[01 14] 学習部 4 0は、 ジェネレータ 4 2〇の入力とジェネレータ 4 2〇の出力と を足し合わせる加算部 4 6を備えており、 加算部 4 6の出力をディスクリミ ネータ 4 4〇に入力させる構成となっている。 加算部 4 6は本開示における 「第 2加算部」 の一例である。 なお、 図 7及び図 1 1では加算部 4 6の図示 が省略される。 ディスクリミネータ 4 4 0に入力する仮想高解像度画像 V 1~1 1は本開示における 「仮想第 2画像」 の一例である。
[01 15] ジェネレータ 4 2◦の入力に与えられる低解像度画像 !_〇は本開示におけ る 「第 1画像」 の一例である。 ジェネレータ 4 2◦から出力される高周波成 分画像は本開示における 「第 2画像」 の一例である。
[01 16] 図 1 3及び図 1 4ではジェネレータ 4 2〇の出力が高周波成分画像 V 1~1 〇である例を説明したが、 ジェネレータ 4 2◦の出力が仮想高解像度画像 V 1~1となる形態も可能である。 この場合、 加算部 4 6は不要となる。 かかる態 様については第 2実施形態として後述する。
[01 17] [学習時におけるディスクリミネータ 4 4口の識別動作]
図 1 5は、 学習時におけるディスクリミネータ 4 4 0による識別の動作を 説明するための図である。 図 1 5において加算部 4 6の図示は省略される。
[01 18] 図 1 5の左図に示す動作状態 7 0 は、 ディスクリミネータ 4 4 0にポジ ティブサンプル (正例) が入力された場合の例を示し、 図 1 5の右図に示す 動作状態 7 0 1\1はディスクリミネータ 4 4 0にネガティプサンプル (負例) が入力された場合の例を示す。
[01 19] 学習データの画像ペアであるリアル高解像度画像 1~1 1 と、 これに対応す る低解像度画像 1- <3 1 とが入力されている場合の例である。 この場合、 ディ スクリミネータ 4 4。が、 入力された高解像度画像をリアル高解像度画像 1~1 1であると識別した場合は、 ディスクリミネータ 4 4 0の出力 (識別結果 ) が正解であり、 仮想高解像度画像▽! ! 1であると識別した場合は不正解で ある。 \¥0 2020/175446 25 卩(:171? 2020 /007383
[0120] —方、 図 1 5の右図に示す動作状態 7 0 1\1の場合は、 ディスクリミネータ 4 4 0に、 ジェネレータ 4 2 0由来の仮想高解像度画像▽! ! 1 と、 その生成 元のデータである低解像度画像 !_ <3 1 と、 の画像ペアが入力されている。 こ の場合、 ディスクリミネータ 4 4 0が、 入力された高解像度画像をリアル高 解像度画像
Figure imgf000027_0001
1であると識別した場合は不正解であり、 仮想高解像度画像 ▽ ! ! 1であると識別した場合は正解である。
[0121 ] ディスクリミネータ 4 4 0は、 入力された高解像度画像が不図示の〇丁装 置によって撮影された本物の〇丁画像であるか、 又はジェネレータ 4 2◦に よって生成された仮想の <3丁画像であるか、 の識別を正解するように学習さ れる。 一方、 ジェネレータ 4 2 0は、 不図示の〇丁装置によって撮影された リアルな <3丁画像に似せた仮想の(3丁画像を生成し、 ディスクリミネータ 4 4口の識別を不正解とするように学習される。
[0122] 学習が進行すると、 ディスクリミネータ 4 4 0とジェネレータ 4 2◦とが 互いに精度を高め合い、 ジェネレータ 4 2◦はディスクリミネータ 4 4 0に 偽物 (仮想高解像度画像) と識別されない、 より本物の <3丁画像に近い仮想 高解像度画像 V 1~1を生成できるようになる。
[0123] このような学習によって獲得された学習済みのジェネレータ 4 2◦が図 6 で説明した超解像画像生成装置 1 0のジェネレータ 1 4として適用される。
[0124] [学習システム 2 0を用いた学習方法]
図 1 6は、 学習部 4 0における処理の手順の例を示すフローチヤートであ る。 図 1 6に示すフローチヤートの各ステップは、 学習部 4 0として機能す るプロセッサを含むコンビュータによって実行される。
[0125] ステップ 3 1 1 において、 学習部 4 0は学習データを取得する。 学習部 4 〇は図 8で説明した学習データ生成部 3 0から学習データを読み込む。 学習 部 4 0は複数の学習データを含むミニバッチの単位で学習データを取得する ことができる。
[0126] ステップ 3 1 2において、 学習部 4 0はジェネレータ 4 2◦に学習データ の低解像度画像を入力する。 \¥0 2020/175446 26 卩(:171? 2020 /007383
[0127] ステップ 3 1 3において、 ジェネレータ 4 2〇は入力された低解像度画像 から仮想高解像度画像を生成する。 ジェネレータ 4 2 0からの出力は仮想高 解像度画像を作るために必要な高周波成分画像 V 1~1 〇であってよい。 この 場合、 図 1 4で説明したとおり、 高周波成分画像▽! ! (3と低解像度画像と が加算されて仮想高解像度画像 V 1~1が生成される。
[0128] ステップ 3 1 4において、 学習部 4 0はディスクリミネータ 4 4 0へのデ —夕入力を行う。 ディスクリミネータ 4 4 0への入力には、 正解画像として のリアル高解像度画像を含む学習データのペア (リアルペア) 、 又は、 ジェ ネレ—夕 4 2(5由来の仮想高解像度画像を含むフェイクペアのいずれかが選 択的に与えられる。
[0129] ステップ 3 1 5において、 ディスクリミネータ 4 4 0はデータの識別を行 う。
[0130] ステップ 3 1 6において、 誤差演算部 5 0は識別結果の誤差を算出し、 そ の結果をオプティマイザ 5 2へ送る。
[0131 ] ステップ 3 1 7において、 オプティマイザ 5 2は、 算出された誤差を基に ネッ トワークのパラメータの更新量を算出する。
[0132] ステップ 3 1 8において、 オプティマイザ 5 2は、 ステップ 3 1 7にて算 出されたパラメータの更新量に従い、 パラメータの更新処理を行う。 パラメ —夕の更新処理はミニバッチの単位で実施される。
[0133] ステップ 3 1 9において、 学習部 4 0は学習を終了するか否かの判別を行 う。 学習終了条件は、 誤差の値に基づいて定められていてもよいし、 パラメ —夕の更新回数に基づいて定められていてもよい。 誤差の値に基づく方法と しては、 例えば、 誤差が規定の範囲内に収束していることを学習終了条件と してよい。 更新回数に基づく方法としては、 例えば、 更新回数が規定回数に 到達したことを学習終了条件としてよい。
[0134] ステップ 3 1 9の判定結果が 1\1〇判定である場合、 学習部 4 0はステップ
3 1 1 に戻り、 学習終了条件を満たすまで、 学習の処理を繰り返す。
[0135] ステップ 3 1 9の判定結果が丫 6 3判定である場合、 学習部は図 1 6のフ 口ーチヤートを終了する。
[0136] こうして得られた学習済みのジェネレータ 4 2 Gの部分を超解像画像生成 装置 1 0のジェネレータ 1 4として適用する。
[0137] [第 1実施形態による効果]
図 1 7は、 本実施形態の効果を示す画像の例である。 図 1 7には、 セルフ アテンション機構の導入の効果を示す画像例が示されている。 図 1 7の上段 中央に示す画像 V H A 1は本実施形態に係る学習方法を適用した学習済みモ デル (ジェネレータ 1 4) を用いて生成された仮想高解像度画像の例である 。 図 1 7の上段左に示す画像 L 丨 1はジェネレータ 1 4への入力に用いた低 解像画像の例である。 図 1 7の上段右に す画像 G T 1は、 入力の画像 L 丨 1 に対応する正解画像 (Ground t ruth) である。
[0138] 図 1 7の下段中央に示す画像 V H N 2は、 比較例に係る学習済みモデルを 用いて生成さ仮想高解像度画像の例である。 比較例に係る学習済みモデルは 、 アテンション機構を持たないディスクリミネータを用いて学習を行ったも のである。 図 1 7の下段左に示す画像 L I 2は比較例に係る学習済みモデル への入力に用いた低解像画像の例である。 図 1 7の下段右に示す画像 G T 2 は、 入力の画像 L 丨 2に対応する正解画像 (Ground t ruth) である。
[0139] 図 1 7の上段中央に示す画像 V H A 1は、 正解の画像 G T 1 に極めて近い 画像となっている。 また、 画像 V H A 1は、 下段の比較例の画像 V H N 2に 比べて、 局所的なノイズが低減されていることがわかる。
[0140] すなわち、 本実施形態によれば、 セルフアテンション機構をディスクリミ ネータ 4 4 Dに導入した効果により、 比較例においては局所的に発生してい たノイズを低減することができる。
[0141 ] 図 1 8は、 本実施形態の他の効果を説明するための図である。 図 1 8には 、 本実施形態に係る学習方法を適用した学習済みモデル (ジェネレータ 1 4 ) が入カサイズによらず画像生成の処理 (超解像処理) を実行可能であるこ とを示す画像例が示されている。
[0142] 図 1 8の最左に示す画像 L I 3は、 本実施形態に係る学習方法を適用した \¥0 2020/175446 28 卩(:171? 2020 /007383
学習済みモデル (ジェネレータ 1 4) に入力された画像の例である。 図 1 8 の左から 2番目の画像▽! !八 3は、 画像 !_ 丨 3の入力からジェネレータ 1 4 を用いて生成された仮想高解像度画像の例である。
[0143] 図 1 8の左から 3番目の画像 !_ 丨 4はジェネレータ 1 4に入力された画像 の他の例である。 この画像 !_ 丨 4は、 最左に示す画像 !_ 丨 3よりも画像サイ ズが小さいものである。 図 1 8の最右に示す画像 V 1~1八 4は、 画像 !_ 丨 4の 入力からジェネレータ 1 4を用いて生成された仮想高解像度画像の例である
[0144] 図 1 8に示すように、 ジェネレータ 1 4は、 学習時に用いた画像サイズと は異なるサイズの画像の入力に対しても推定の処理を実施することができる 。 ジェネレータ 1 4は、 任意の画像サイズの入カデータに対して、 高精度の 画像生成が可能である。 すなわち、 本実施形態によれば、 学習時の学習デー 夕として用いた固定サイズの画像サイズに制約されずに、 任意サイズの画像 に対しても画像生成処理が実施可能である。 本実施形態によれば、 入カデー 夕を任意サイズのメモリに分割して処理を行うことができる。
[0145] なお、 ディスクリミネータ 4 4 0は学習の際に使用するだけであり、 超解 像画像生成装置 1 0に搭載する必要がないため、 ディスクリミネータ 4 4 0 にアテンシヨン機構を追加しても固定サイズで学習を行うため問題はない。
[0146] 《変形例》
上述した第 1実施形態では、 ディスクリミネータ 4 4 0への入力としてリ アル高解像度画像 8 ! !と低解像度画像 !_ 0 1のペア、 又はジェネレータ 4 2 ◦に由来する仮想高解像度画像▽! !と低解像度画像 !_〇 1のペアが与えられ ているが、 ディスクリミネータ 4 4 0に対する低解像度画像 !_〇 1の入力は 必須ではない。 ディスクリミネータ 4 4 0には、 少なくともリアル高解像度 画像 1~1、 又は仮想高解像度画像 V 1~1が入力されればよい。
[0147] 《第 2実施形態》
第 1実施形態では図 1 4のように、 ジェネレータ 4 2 0から (仮想的な) 高周波成分画像
Figure imgf000030_0001
低解像度画像 !_〇と高周波成分画像▽! ! \¥0 2020/175446 29 卩(:171? 2020 /007383
〇とを加算することによって仮想高解像度画像 V 1~1を得ている。 これに対 し、 第 2実施形態は、 ジェネレータ 4 2 0が仮想高解像度画像▽! !を出力す る形態である。
[0148] 図 1 9は、 第 2実施形態に係る学習システム 2 0による処理の流れを概略 的に示す機能ブロック図である。 なお、 図 1 9において図 7、 図 8、 図 1 1 から図 1 4に示す構成と共通又は類似する部分には同一の符号を付し、 その 詳細な説明は省略する。
[0149] 図 1 9において、 学習データ生成部 3 0の内容は図 8と同様である。 第 2 実施形態における学習部 4 0のジェネレータ 4 2◦は、 低解像度画像 !_〇 1 から仮想高解像度画像
Figure imgf000031_0001
を生成する。
[0150] また、 低解像度画像 !_ 0 1は、 リアル高解像度画像 1~1 1又は仮想高解像 度画像▽! ! 1 とペアでディスクリミネータ 4 4 0に入力される。 ディスクリ ミネータ 4 4 0は、 入力された画像がリアル高解像度画像 1~1 1、 及び仮想 高解像度画像
Figure imgf000031_0002
1のいずれであるかを識別する。 なお、 ディスクリミネー 夕 4 4 0には、 低解像度画像 !_〇 1は入力されなくてもよい。
[0151 ] 第 2実施形態によれば、 低解像度画像 !_ 0 1から仮想高解像度画像 V 1~1 1 を生成する生成モデル (ジェネレータ 4 2 0) を得ることができる。 第 2実 施形態によって生成されたジェネレータ 4 2 0を超解像画像生成装置 1 0に 組み込む場合には、 図 6に示した加算部 1 6を省略することができる。
[0152] 《第 3実施形態》
図 2 0は、 第 3実施形態に係る学習システム 2 0による処理の流れを概略 的に示す機能ブロック図である。 図 2 0において、 図 1 9に示す構成と共通 又は類似する部分には同一の符号を付し、 その詳細な説明は省略する。 図 2 0に示す第 3実施形態は、 ジェネレータ 4 2◦への入力がシックスライス画
Figure imgf000031_0003
ジェネレータ 4 2 0からの出力が仮想高解像度画像▽! !であ る。 この場合、 シックスライス画像
Figure imgf000031_0004
とリアル高解像度画像 1~1のペアが 学習データとなる。 図 2 0においてシックスライス画像
Figure imgf000031_0005
は本開示におけ る 「第 1画像」 及び 「第 1学習用画像」 の一例である。 \¥0 2020/175446 30 卩(:171? 2020 /007383
[0153] 第 3実施形態によれば、 ジェネレータ 4 2〇はシックスライス画像
Figure imgf000032_0001
ら仮想高解像度画像
Figure imgf000032_0002
1 を生成するように学習される。 したがって、 第 3 実施形態の学習を行うことにより、 シックスライス画像
Figure imgf000032_0003
から仮想高解像 度画像 V 1~1 1 を生成する生成モデル (ジェネレータ 4 2 0) を得ることがで きる。
[0154] 《第 4実施形態》
図 2 1は、 第 4実施形態に係る学習システム 2 0による処理の流れを概略 的に示す機能ブロック図である。 図 2 1 において、 図 1 9に示す構成と共通 又は類似する部分には同一の符号を付し、 その詳細な説明は省略する。 図 2 1 に示す第 4実施形態に係る学習システム 2 0は、 ジェネレータ 4 2〇にシ ックスライス画像
Figure imgf000032_0004
を入力してジェネレータ 4 2◦から高周波成分画像 V 1~1 〇を出力させ、 ディスクリミネータ 4 4 0に高周波成分画像を入力して 識別を行う。 ディスクリミネータ 4 4口の入力に用いる学習用の高周波成分 画像を作るために、 学習データ生成部 3 0は、 高周波成分抽出部 3 3を備え ている。
[0155] 高周波成分抽出部 3 3は、 リアル高解像度画像 1~1から高周波成分を抽出 し、 リアル高周波成分画像 8 1~1 〇を生成する。 高周波成分の抽出は、 ハイ パスフィルタを用いて行われる。 リアル高周波成分画像 1~1 〇は、 リアル 高解像度画像 8 と同様に、 スライス間隔が 1 01 111、 スライス厚が 1 01 111で ある。
[0156] 第 4実施形態では、 シックスライス画像
Figure imgf000032_0005
とリアル高周波成分画像 1~1
〇のペアが学習データとなる。 図 2 1 においてリアル高周波成分画像
Figure imgf000032_0006
〇は本開示における 「第 2学習用画像」 の一例である。
[0157] 高周波成分抽出部 3 3が生成したリアル高周波成分画像[¾ ! ! 〇は、 学習 部 4 0のディスクリミネータ 4 4〇に入力される。
[0158] 学習部 4 0のジェネレータ 4 2◦は、 入力されたシックスライス画像
Figure imgf000032_0007
から、 リアル高周波成分画像
Figure imgf000032_0008
分画像 V ! ! 〇を生成する。 ここでは、 ジェネレータ 4 2 0は、 スライス間 隔が 1 、 スライス厚が 1 の仮想高周波成分画像 1~1 (3を生成する
[0159] ディスクリミネータ 44 Dには、 リアル高周波成分画像 RH FCとシック スライス画像 L Kとのペア、 又は、 ジェネレータ 42 Gの出力に由来する仮 想高解像度画像 VH FCとシックスライス画像 L Kとのペアが入力される。
[0160] ディスクリミネータ 44 Dは、 入力された高周波成分画像がリアル高周波 成分画像 R H F C、 及び仮想高周波成分画像 V H F Cのいずれであるかを識 別する。
[0161] 第 4実施形態によれば、 ジェネレータ 42 Gは、 低解像度の画像であるシ ックスライス画像 L Kから高周波成分画像を生成するように学習される。 図 1 4で説明したように、 ジェネレータ 42 Gが生成した高周波成分画像とジ ェネレータ 42 Gの入力であるシックスライス画像 L Kとを加算処理するこ とで、 高解像度画像を得ることができる。
[0162] 《コンビユータのハードウェア構成の例》
図 22は、 学習システム 20に用いられるコンビユータのハードウェア構 成の例を示すブロック図である。 コンビユータ 500は、 パーソナルコンビ ユータであってもよいし、 ワークステーシヨンであってもよく、 また、 サー バコンピユータであってもよい。 コンビユータ 500は、 超解像画像生成装 置 1 0、 画像保管部 24、 学習データ生成部 30、 及び学習部 40のいずれ か、 又はこれらの複数の機能を備えた装置として用いることができる。
[0163] コンビユータ 500は、 通信部 5 1 2、 ストレージ 5 1 4、 操作部 5 1 6 、 C P U (Central Processing Unit) 5 1 8、 G P U (Graphics Process i n g Unit) 5 1 9、 RAM (Random Access Memory) 520、 ROM (Read On ly Memory) 522、 及び表示部 524を備える。 なお、 G P U (Graphics P rocessing Unit) 5 1 9は省略されてもよい。
[0164] 通信部 5 1 2は、 有線又は無線により外部装置との通信処理を行い、 外部 装置との間で情報のやり取りを行うインターフェースである。
[0165] ストレージ 5 1 4は、 例えば、 ハードディスク装置、 光ディスク、 光磁気 ディスク、 若しくは半導体メモリ、 又はこれらの適宜の組み合わせを用いて 構成される記憶装置を含んで構成される。 ストレージ 5 1 4には、 学習処理 及び/又は画像生成処理等の画像処理に必要な各種プログラムやデータ等が 記憶される。 ストレージ 5 1 4に記憶されているプログラムが RAM520 に口ードされ、 これを C P U 5 1 8が実行することにより、 コンビュータは 、 プログラムで規定される各種の処理を行う手段として機能する。
[0166] 操作部 5 1 6は、 コンピュータ 500に対する各種の操作入力を受け付け る入カインターフェースである。 操作部 5 1 6は、 例えば、 キーボード、 マ ウス、 タッチパネル、 操作ボタン、 若しくは、 音声入力装置、 又はこれらの 適宜の組み合わせであつてよい。
[0167] C P U 5 1 8は、 ROM 522又はストレージ 5 1 4等に記憶された各種 のプログラムを読み出し、 各種の処理を実行する。 RAM520は、 C P U 5 1 8の作業領域として使用される。 また、 RAM520は、 読み出された プログラム及び各種のデータを一時的に記憶する記憶部として用いられる。
[0168] 表示部 524は、 各種の情報が表示される出カインターフェースである。
表示部 524は、 例えば、 液晶ディスプレイ、 有機 E L (organic electro-l uminescence: 0 E L) ディスプレイ、 若しくは、 プロジェクタ、 又はこれら の適宜の組み合わせであつてよい。
[0169] 《コンピュータを動作させるプログラムについて》
上述の各実施形態で説明した学習データ生成機能、 学習機能、 及び画像生 成機能のうち少なくとも 1つの処理機能の一部又は全部をコンピュータに実 現させるプログラムを、 光ディスク、 磁気ディスク、 若しくは、 半導体メモ リその他の有体物たる非一時的な情報記憶媒体であるコンピュータ可読媒体 に記録し、 この情報記憶媒体を通じてプログラムを提供することが可能であ る。
[0170] またこのような有体物たる非一時的な情報記憶媒体にプログラムを記憶さ せて提供する態様に代えて、 インターネッ トなどの電気通信回線を利用して プログラム信号をダウンロードサービスとして提供することも可能である。 [0171] また、 上述の各実施形態で説明した学習データ生成機能、 学習機能、 及び 画像生成機能のうち少なくとも 1つの処理機能の一部又は全部をアプリケー ションサーバとして提供し、 電気通信回線を通じて処理機能を提供するサー ビスを行うことも可能である。
[0172] 学習データ生成部 30として機能するコンピュータは学習データ生成装置 と理解される。 学習部 40として機能するコンピュータは学習装置と理解さ れる。
[0173] 《各処理部のハードウェア構成について》
図 6の補間処理部 1 2、 ジェネレータ 1 4、 及び加算部 1 6、 図 7の画像 保管部 24、 学習データ生成部 30、 学習部 40、 GAN4 1、 ジヱネレー 夕 42G、 ディスクリミネータ 44 D、 誤差演算部 50、 及びオプティマイ ザ 52、 図 8の固定サイズ領域切出部 3 1、 ダウンサンプル処理部 32、 ア ップサンプル処理部 34、 補間処理部 35、 及びガウシアンフィルタ処理部 36、 並びに図 2 1の高周波成分抽出部 33などの各種の処理を実行する処 理部 (processing unit) のハードウェア的な構造は、 例えば、 次に示すよう な各種のプロセッサ (processor) である。
[0174] 各種のプロセッサには、 プログラムを実行して各種の処理部として機能す る汎用的なプロセッサである C P U、 画像処理に特化したプロセッサである G P U、 F PGA (Field Programmable Gate Array) などの製造後に回路構 成を変更可能なプロセッサであるプログラマブルロジックデバイス (Program mab le Logic Dev ice : P L D) 、 AS I C (App 11 cat i on Spec i f i c Integra† ed Circuit) などの特定の処理を実行させるために専用に設計された回路構 成を有するプロセッサである専用電気回路などが含まれる。
[0175] 1つの処理部は、 これら各種のプロセッサのうちの 1つで構成されていて もよいし、 同種又は異種の 2つ以上のプロセッサで構成されてもよい。 例え ば、 1つの処理部は、 複数の F P G A、 或いは、 C P Uと F P G Aの組み合 わせ、 又は C P Uと G P Uの組み合わせによって構成されてもよい。 また、 複数の処理部を 1つのプロセッサで構成してもよい。 複数の処理部を 1つの プロセッサで構成する例としては、 第一に、 クライアントやサーバなどのコ ンピユータに代表されるように、 1つ以上の C P Uとソフトウェアの組み合 わせで 1つのプロセッサを構成し、 このプロセッサが複数の処理部として機 能する形態がある。 第二に、 システムオンチップ (System On Chip : S〇 C ) などに代表されるように、 複数の処理部を含むシステム全体の機能を 1つ の丨 C (Integrated Circuit) チップで実現するプロセッサを使用する形態 がある。 このように、 各種の処理部は、 ハードウェア的な構造として、 上記 各種のプロセッサを 1つ以上用いて構成される。
[0176] さらに、 これらの各種のプロセッサのハードウェア的な構造は、 より具体 的には、 半導体素子などの回路素子を組み合わせた電気回路 (circuitry) で ある。
[0177] 《その他》
ここでは C T画像の超解像の生成モデルの学習方法を説明したが、 本開示 による生成モデルの学習方法は、 CT画像に限らず、 各種の 3次元断層画像 に適用することができる。 例えば、 MR 丨 (Magnetic Resonance Imaging) 装置により取得される M R画像、 P ET (Positron Emission Tomography) 装置により取得される P E T画像、 OCT (Optical Coherence Tomography ) 装置により取得される OCT画像、 3次元超音波撮影装置により取得され る 3次元超音波画像等であってもよい。
[0178] また、 本開示による生成モデルの学習方法は、 3次元断層画像に限らず、 各種の 2次元画像に適用することができる。 例えば、 X線画像であってもよ い。 また、 医療画像に限定されず、 通常のカメラ画像に適用することができ る。
[0179] 本発明の技術的範囲は、 上記の実施形態に記載の範囲には限定されない。
各実施形態における構成等は、 本発明の趣旨を逸脱しない範囲で、 各実施形 態間で適宜組み合わせることができる。
符号の説明
[0180] 1 0 超解像画像生成装置 \¥02020/175446 35 卩(:171? 2020 /007383
1 2 補間処理部
1 4 ジェネレータ
1 6 加算部
20 学習システム
24 画像保管部
30 学習データ生成部
3 1 固定サイズ領域切出部
32 ダウンサンプル処理部
33 高周波成分抽出部
34 アップサンプル処理部
35 補間処理部
36 ガウシアンフィルタ処理部
38 学習データ記憶部
40 学習部
4 1 敵対的生成ネッ トヮーク (〇八1\!)
420 ジェネレータ
440 ディスクリミネータ
46 加算部
50 誤差演算部
52 オプティマイザ
70 動作状態
709 動作状態
500 コンビュータ
5 1 2 通信部
5 1 4 ストレージ
5 1 6 操作部
5 1 8
Figure imgf000037_0001
5 1 9 0 II \¥02020/175446 36 2020/007383
Figure imgf000038_0001
524 表示部
〇 01 ~〇 05 畳:み込み層
Figure imgf000038_0002
畳:み込み層
071 画像
◦丁 2 画像
I IV! 1 ~丨 IV! 4 07画像
丨 丁 補間画像
!_ 丨 1〜 !_ 丨 4 画像
!_ [<、 1_ [< 1 シックスライス画像
!_〇、 1_〇 1 低解像度画像
Figure imgf000038_0003
リアル高解像度画像
Figure imgf000038_0004
リアル高解像度画像
リアル高周波成分画像
3八 1\/1 セルフアテンション特徴マップ
30 スライス間隔
37 スライス厚
TCK シックスライス画像
▽ !!、
Figure imgf000038_0005
仮想高解像度画像
V 1~1 1 画像
V 1~1 2 画像
V 1~1 3 画像
V 1~1 4 画像
V 1~1 0 高周波成分画像
V丁 仮想シンスライス画像
31〜35 学習データ生成処理のステップ
31 1〜ステップ 31 9 学習処理のステップ

Claims

\¥0 2020/175446 37 卩(:17 2020 /007383 請求の範囲
[請求項 1 ] 第 1画像から前記第 1画像よりも高解像の画像情報を含む第 2画像 を推定する生成モデルの機械学習を行う学習方法であって、
前記生成モデルであるジェネレータと、 与えられたデータが学習用 の正解画像のデータであるか前記ジェネレータからの出力に由来する データであるかを識別する識別モデルであるディスクリミネータと、 を含む敵対的生成ネッ トワークを用いることと、
前記第 2画像よりも解像度が低い第 1解像度情報を含む第 1学習用 画像と、 前記第 1学習用画像よりも解像度が高い第 2解像度情報を含 む第 2学習用画像であって前記第 1学習用画像に対応する前記正解画 像となる前記第 2学習用画像と、 を学習データとして用いることと、 前記ジェネレータの入力には、 前記第 1学習用画像及び前記第 2学 習用画像のうち前記第 1学習用画像のみを与えることと、
前記ジェネレータ及び前記ディスクリミネータのうち、 前記ディス クリミネータのネッ トワークに限定してセルフアテンション機構を実 装することと、
を含む学習方法。
[請求項 2] 前記ジェネレータ及び前記ディスクリミネータのそれぞれのネッ ト ワ _クは、 畳:み込みニユ _ラルネッ トワ _クである、 請求項 1 に記載 の学習方法。
[請求項 3] 前記第 1画像は 3次元断層画像であり、
前記第 2画像は少なくとも前記 3次元断層画像のスライス厚方向の 解像度が前記第 1画像よりも高解像である、 請求項 1又は 2に記載の 学習方法。
[請求項 4] 前記第 2学習用画像は、 コンピュータ断層撮影装置を用いて取得さ れた画像であり、
前記第 1学習用画像は、 前記第 2学習用画像を基に画像処理によつ て生成された画像である、 請求項 1から 3のいずれか一項に記載の学 \¥0 2020/175446 38 卩(:171? 2020 /007383
習方法。
[請求項 5] 前記第 2学習用画像から前記第 1学習用画像を生成する前記画像処 理は、 前記第 2学習用画像をダウンサンプルする処理を含む、 請求項 4に記載の学習方法。
[請求項 6] 前記第 2学習用画像から前記第 1学習用画像を生成する前記画像処 理は、 前記ダウンサンプルの処理によって得られた画像に補間処理を 施してアツプサンプルする処理を含む、 請求項 5に記載の学習方法。
[請求項 7] 前記第 2学習用画像から前記第 1学習用画像を生成する前記画像処 理は、 ガウシアンフィルタを用いる平滑化処理を含む、 請求項 4から 6のいずれか一項に記載の学習方法。
[請求項 8] 前記機械学習に使用する複数種類の前記学習データにおける前記第
1学習用画像及び前記第 2学習用画像の各々は同ーサイズである、 請 求項 1から 7のいずれか一項に記載の学習方法。
[請求項 9] 前記第 2画像は、 高周波成分の情報を示す高周波成分画像であり、 前記ジェネレータは、 入力された画像の解像度を高めるために必要 な高周波成分を推定し、 前記高周波成分の情報を示す高周波成分画像 を出力する、 請求項 1から 8のいずれか一項に記載の学習方法。
[請求項 10] 前記ジェネレータから出力された前記高周波成分画像と、 前記ジェ ネレータに入力された前記画像とを加算すること、 をさらに含み、 前記加算によって得られる仮想第 2画像を前記ディスクリミネータ の入力に与える、 請求項 9に記載の学習方法。
[請求項 1 1 ] 請求項 1から 1 〇のいずれか一項に記載の学習方法をコンピュータ に実行させるためのプログラム。
[請求項 12] 非一時的かつコンピュータ読取可能な記録媒体であって、 前記記録 媒体に格納された指令がコンピュータによって読み取られた場合に請 求項 1 1 に記載のプログラムをコンピュータに実行させる記録媒体。
[請求項 13] 請求項 1から 1 0のいずれか一項に記載の学習方法を実施して学習 された学習済みモデルであって、 前記第 1画像から前記第 1画像より \¥0 2020/175446 39 卩(:171? 2020 /007383
も高解像の画像情報を含む第 2画像を推定する前記生成モデルである 学習済みモデル。
[請求項 14] 請求項 1から 1 〇のいずれか一項に記載の学習方法を実施して学習 された学習済みモデルである前記生成モデルを備え、 入力される第 3 画像から前記第 3画像よりも高解像の画像情報を含む第 4画像を生成 する超解像画像生成装置。
[請求項 15] 前記第 3画像は、 前記第 1学習用画像と異なる画像サイズである、 請求項 1 4に記載の超解像画像生成装置。
[請求項 16] 前記第 3画像に補間処理を行い、 補間画像を生成する第 1補間処理 部と、
前記補間画像と前記生成モデルが生成する高周波成分とを加算する 第 1加算部と、 を含み、
前記補間画像が前記生成モデルに入力され、
前記生成モデルが前記補間画像の解像度を高めるために必要な前記 高周波成分を生成する、 請求項 1 4又は 1 5に記載の超解像画像生成 装置。
[請求項 17] 第 1画像から前記第 1画像よりも高解像の画像情報を含む第 2画像 を推定する生成モデルの機械学習を行う学習システムであって、 前記生成モデルであるジェネレータと、 与えられたデータが学習用 の正解画像のデータであるか前記ジェネレータからの出力に由来する データであるかを識別する識別モデルであるディスクリミネータと、 を含む敵対的生成ネッ トワークを備え、
前記ジェネレータ及び前記ディスクリミネータのうち、 前記ディス クリミネータのネッ トワークに限定してセルフアテンション機構が実 装されており、
前記第 2画像よりも解像度が低い第 1解像度情報を含む第 1学習用 画像と、 前記第 1学習用画像よりも解像度が高い第 2解像度情報を含 む第 2学習用画像であって前記第 1学習用画像に対応する前記正解画 \¥0 2020/175446 40 卩(:171? 2020 /007383
像となる前記第 2学習用画像と、 を学習データとして取り込み、 前記ジェネレータの入力に、 前記第 1学習用画像及び前記第 2学習 用画像のうち前記第 1学習用画像のみが与えられ、 前記敵対的生成ネ ッ トワークの学習が行われる、
学習システム。
[請求項 18] 前記学習データを生成する学習データ生成部をさらに備え、
前記学習データ生成部は、
前記第 2解像度情報を含むオリジナルの元画像から固定サイズ領域 を切り出す固定サイズ領域切出部と、
前記固定サイズ領域切出部によって切り出された前記固定サイズ領 域の画像をダウンサンプルするダウンサンプル処理部と、
を含み、
前記固定サイズ領域切出部によって切り出された前記固定サイズ領 域の画像を前記第 2学習用画像とし、
前記第 2学習用画像に対して前記ダウンサンプルの処理を行うこと によって前記第 1学習用画像を生成する、 請求項 1 7に記載の学習シ ステム。
[請求項 19] 前記学習データ生成部は、 さらに、
前記ダウンサンプルの処理によって得られた画像に補間処理を施す 第 2補間処理部と、
ガウシアンフィルタを用いて平滑化を行う平滑化処理部と、 を含む、 請求項 1 8に記載の学習システム。
[請求項 20] 前記ジェネレータは、 入力された画像の解像度を高めるために必要 な高周波成分を推定して前記高周波成分の情報を示す高周波成分画像 を出力する構成であり、
前記ジェネレータから出力された前記高周波成分画像と前記ジェネ レータに入力された前記画像とを加算する第 2加算部をさらに備える 、 請求項 1 7から 1 9のいずれか一項に記載の学習システム。
PCT/JP2020/007383 2019-02-28 2020-02-25 学習方法、学習システム、学習済みモデル、プログラム及び超解像画像生成装置 Ceased WO2020175446A1 (ja)

Priority Applications (3)

Application Number Priority Date Filing Date Title
EP20762514.6A EP3932318B1 (en) 2019-02-28 2020-02-25 Learning method, learning system, learned model, program, and super-resolution image generation device
JP2021502251A JP7105363B2 (ja) 2019-02-28 2020-02-25 学習方法、学習システム、学習済みモデル、プログラム及び超解像画像生成装置
US17/400,142 US12217387B2 (en) 2019-02-28 2021-08-12 Learning method, learning system, learned model, program, and super resolution image generating device

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2019036374 2019-02-28
JP2019-036374 2019-02-28

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US17/400,142 Continuation US12217387B2 (en) 2019-02-28 2021-08-12 Learning method, learning system, learned model, program, and super resolution image generating device

Publications (1)

Publication Number Publication Date
WO2020175446A1 true WO2020175446A1 (ja) 2020-09-03

Family

ID=72238587

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2020/007383 Ceased WO2020175446A1 (ja) 2019-02-28 2020-02-25 学習方法、学習システム、学習済みモデル、プログラム及び超解像画像生成装置

Country Status (4)

Country Link
US (1) US12217387B2 (ja)
EP (1) EP3932318B1 (ja)
JP (1) JP7105363B2 (ja)
WO (1) WO2020175446A1 (ja)

Cited By (18)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2022091869A1 (ja) * 2020-10-26 2022-05-05 キヤノン株式会社 医用画像処理装置、医用画像処理方法及びプログラム
CN114693897A (zh) * 2021-04-28 2022-07-01 上海联影智能医疗科技有限公司 用于医学图像的无监督层间超分辨率
KR102428326B1 (ko) * 2021-12-21 2022-08-02 서울시립대학교 산학협력단 인공지능 기반의 결함 탐지 방법 및 시스템
JPWO2022163402A1 (ja) * 2021-01-26 2022-08-04
JP2022155027A (ja) * 2021-03-30 2022-10-13 富士フイルム株式会社 画像処理装置、学習装置、放射線画像撮影システム、画像処理方法、学習方法、画像処理プログラム、及び学習プログラム
JP2022155026A (ja) * 2021-03-30 2022-10-13 富士フイルム株式会社 画像処理装置、学習装置、放射線画像撮影システム、画像処理方法、学習方法、画像処理プログラム、及び学習プログラム
JP2022161003A (ja) * 2021-04-07 2022-10-20 キヤノンメディカルシステムズ株式会社 医用画像処理方法、医用画像処理装置、x線ct装置、および医用画像処理プログラム
JP2022161004A (ja) * 2021-04-07 2022-10-20 キヤノンメディカルシステムズ株式会社 医用データ処理方法、モデル生成方法、医用データ処理装置、および医用データ処理プログラム
JP2023533907A (ja) * 2020-10-02 2023-08-07 グーグル エルエルシー 自己注意ベースのニューラルネットワークを使用した画像処理
JP2023553004A (ja) * 2020-12-01 2023-12-20 ビーダブリューエックスティー・アドバンスド・テクノロジーズ・エルエルシー 付加製造のためのディープラーニングベースの画像拡張
CN117456297A (zh) * 2019-03-31 2024-01-26 华为技术有限公司 图像生成方法、神经网络的压缩方法及相关装置、设备
JP2024031119A (ja) * 2022-08-25 2024-03-07 富士フイルム株式会社 画像処理装置、画像処理方法、画像処理プログラム、及び内視鏡システム
JP2024031118A (ja) * 2022-08-25 2024-03-07 富士フイルム株式会社 画像処理装置及びその作動方法並びに内視鏡システム
WO2024057768A1 (ja) 2022-09-14 2024-03-21 富士フイルム株式会社 画像生成装置、学習装置、画像処理装置、画像生成方法、学習方法及び画像処理方法
WO2024253035A1 (ja) * 2023-06-09 2024-12-12 日本電気株式会社 モデル生成装置、変換装置、モデル生成方法、変換方法、および記録媒体
WO2025014322A1 (ko) * 2023-07-13 2025-01-16 주식회사 클라리파이 3차원 의료영상의 단면간 해상도 향상 장치 및 방법
US12499511B2 (en) 2020-10-26 2025-12-16 Canon Kabushiki Kaisha Medical-image processing apparatus, medical-image processing method, and program for the same
JP7855375B2 (ja) 2021-04-07 2026-05-08 キヤノンメディカルシステムズ株式会社 医用画像処理方法、医用画像処理装置、x線ct装置、および医用画像処理プログラム

Families Citing this family (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP3932318B1 (en) * 2019-02-28 2026-01-14 FUJIFILM Corporation Learning method, learning system, learned model, program, and super-resolution image generation device
CN113343705B (zh) * 2021-04-26 2022-07-05 山东师范大学 一种基于文本语义的细节保持图像生成方法及系统
KR102943885B1 (ko) * 2021-12-08 2026-03-26 엘지디스플레이 주식회사 영상 처리 장치와 영상 처리 방법 및 이를 기반으로 한 표시장치
DE102021214741B3 (de) * 2021-12-20 2023-02-23 Siemens Healthcare Gmbh Verfahren zum Generieren von synthetischen Röntgenbildern, Steuereinheit und Computerprogramm
US12363141B2 (en) * 2022-04-19 2025-07-15 Akamai Technologies, Inc. Real-time detection and prevention of online new-account creation fraud and abuse
CN114547017B (zh) * 2022-04-27 2022-08-05 南京信息工程大学 一种基于深度学习的气象大数据融合方法
CN114693831B (zh) * 2022-05-31 2022-09-02 深圳市海清视讯科技有限公司 一种图像处理方法、装置、设备和介质
US12614242B2 (en) * 2022-06-03 2026-04-28 Intel Corporation Methods and apparatus to implement dual-attention vision transformers for interactive image segmentation
KR20240069418A (ko) * 2022-11-11 2024-05-20 삼성전자주식회사 두 이미지 간 대응점을 찾기 위한 장치와 그 동작 방법 및 이를 학습하는 방법
WO2024186659A1 (en) * 2023-03-06 2024-09-12 Intuitive Surgical Operations, Inc. Generation of high resolution medical images using a machine learning model

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20180075581A1 (en) 2016-09-15 2018-03-15 Twitter, Inc. Super resolution using a generative adversarial network
US20190057488A1 (en) * 2017-08-17 2019-02-21 Boe Technology Group Co., Ltd. Image processing method and device

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10262243B2 (en) 2017-05-24 2019-04-16 General Electric Company Neural network point cloud generation system
EP3447721A1 (en) * 2017-08-24 2019-02-27 Agfa Nv A method of generating an enhanced tomographic image of an object
EP3932318B1 (en) * 2019-02-28 2026-01-14 FUJIFILM Corporation Learning method, learning system, learned model, program, and super-resolution image generation device
EP3859599B1 (en) * 2020-02-03 2023-12-27 Robert Bosch GmbH Training a generator neural network using a discriminator with localized distinguishing information

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20180075581A1 (en) 2016-09-15 2018-03-15 Twitter, Inc. Super resolution using a generative adversarial network
US20190057488A1 (en) * 2017-08-17 2019-02-21 Boe Technology Group Co., Ltd. Image processing method and device

Non-Patent Citations (4)

* Cited by examiner, † Cited by third party
Title
HAN ZHANGIAN GOODFELLOWDIMITRIS METAXASAUGUSTUS ODENA: "Self-Attention Generative Adversarial Networks", ARXIV: 1805.08318
IAN J.GOODFELLOW, JEAN POUGET-ABADIEMEHDI MIRZABING XUDAVID WARDE-FARLEYSHERJIL OZAIRAARON COURVILLEYOSHUA BENGIO: "Generative Adversarial Nets", ARXIV: 1406.2661
PATHAK,HARSH, NILESH ET AL.: "Efficient Super Resolution for Large-Scale Images Using Attentional GAN", 2018 IEEE INTERNATIONAL CONFERENCE ON BIG DATA (BIG DATA, 2018, pages 1777 - 1786, XP033508444, DOI: 10.1109/BigData.2018.8622477 *
PHILLIP ISOLAJUN-YAN ZHUTINGHUI ZHOUALEXEI A. EFROS: "Image-to-Image Translation with Conditional Adversarial Networks", CVPR2016

Cited By (32)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117456297A (zh) * 2019-03-31 2024-01-26 华为技术有限公司 图像生成方法、神经网络的压缩方法及相关装置、设备
US11983903B2 (en) 2020-10-02 2024-05-14 Google Llc Processing images using self-attention based neural networks
JP7536893B2 (ja) 2020-10-02 2024-08-20 グーグル エルエルシー 自己注意ベースのニューラルネットワークを使用した画像処理
US12125247B2 (en) 2020-10-02 2024-10-22 Google Llc Processing images using self-attention based neural networks
JP2023533907A (ja) * 2020-10-02 2023-08-07 グーグル エルエルシー 自己注意ベースのニューラルネットワークを使用した画像処理
JP2022070035A (ja) * 2020-10-26 2022-05-12 キヤノン株式会社 医用画像処理装置、医用画像処理方法及びプログラム
US12499511B2 (en) 2020-10-26 2025-12-16 Canon Kabushiki Kaisha Medical-image processing apparatus, medical-image processing method, and program for the same
JP7562369B2 (ja) 2020-10-26 2024-10-07 キヤノン株式会社 医用画像処理装置、医用画像処理方法及びプログラム
WO2022091869A1 (ja) * 2020-10-26 2022-05-05 キヤノン株式会社 医用画像処理装置、医用画像処理方法及びプログラム
JP7832942B2 (ja) 2020-12-01 2026-03-18 ビーダブリューエックスティー・アドバンスド・テクノロジーズ・エルエルシー 付加製造のためのディープラーニングベースの画像拡張
JP2023553004A (ja) * 2020-12-01 2023-12-20 ビーダブリューエックスティー・アドバンスド・テクノロジーズ・エルエルシー 付加製造のためのディープラーニングベースの画像拡張
US12579720B2 (en) 2021-01-26 2026-03-17 Fujifilm Corporation Method of generating trained model, machine learning system, program, and medical image processing apparatus
JP7765415B2 (ja) 2021-01-26 2025-11-06 富士フイルム株式会社 学習済みモデルの生成方法、機械学習システム、プログラムおよび医療画像処理装置
JPWO2022163402A1 (ja) * 2021-01-26 2022-08-04
JP7657635B2 (ja) 2021-03-30 2025-04-07 富士フイルム株式会社 画像処理装置、学習装置、放射線画像撮影システム、画像処理方法、学習方法、画像処理プログラム、及び学習プログラム
JP2022155026A (ja) * 2021-03-30 2022-10-13 富士フイルム株式会社 画像処理装置、学習装置、放射線画像撮影システム、画像処理方法、学習方法、画像処理プログラム、及び学習プログラム
JP2022155027A (ja) * 2021-03-30 2022-10-13 富士フイルム株式会社 画像処理装置、学習装置、放射線画像撮影システム、画像処理方法、学習方法、画像処理プログラム、及び学習プログラム
JP7542478B2 (ja) 2021-03-30 2024-08-30 富士フイルム株式会社 画像処理装置、学習装置、放射線画像撮影システム、画像処理方法、学習方法、画像処理プログラム、及び学習プログラム
JP2022161004A (ja) * 2021-04-07 2022-10-20 キヤノンメディカルシステムズ株式会社 医用データ処理方法、モデル生成方法、医用データ処理装置、および医用データ処理プログラム
US12299841B2 (en) 2021-04-07 2025-05-13 Canon Medical Systems Corporation Medical data processing method, model generation method, medical data processing apparatus, and computer-readable non-transitory storage medium storing medical data processing program
JP7855375B2 (ja) 2021-04-07 2026-05-08 キヤノンメディカルシステムズ株式会社 医用画像処理方法、医用画像処理装置、x線ct装置、および医用画像処理プログラム
US12243127B2 (en) 2021-04-07 2025-03-04 Canon Medical Systems Corporation Medical image processing method, medical image processing apparatus, and computer readable non-volatile storage medium storing medical image processing program
JP2022161003A (ja) * 2021-04-07 2022-10-20 キヤノンメディカルシステムズ株式会社 医用画像処理方法、医用画像処理装置、x線ct装置、および医用画像処理プログラム
CN114693897A (zh) * 2021-04-28 2022-07-01 上海联影智能医疗科技有限公司 用于医学图像的无监督层间超分辨率
KR102428326B1 (ko) * 2021-12-21 2022-08-02 서울시립대학교 산학협력단 인공지능 기반의 결함 탐지 방법 및 시스템
US12205244B2 (en) 2022-08-25 2025-01-21 Fujifilm Corporation Image processing apparatus, image processing method, non-transitory computer readable medium, and endoscope system
JP2024031119A (ja) * 2022-08-25 2024-03-07 富士フイルム株式会社 画像処理装置、画像処理方法、画像処理プログラム、及び内視鏡システム
JP2024031118A (ja) * 2022-08-25 2024-03-07 富士フイルム株式会社 画像処理装置及びその作動方法並びに内視鏡システム
JP7821701B2 (ja) 2022-08-25 2026-02-27 富士フイルム株式会社 画像処理装置及びその作動方法並びに内視鏡システム
WO2024057768A1 (ja) 2022-09-14 2024-03-21 富士フイルム株式会社 画像生成装置、学習装置、画像処理装置、画像生成方法、学習方法及び画像処理方法
WO2024253035A1 (ja) * 2023-06-09 2024-12-12 日本電気株式会社 モデル生成装置、変換装置、モデル生成方法、変換方法、および記録媒体
WO2025014322A1 (ko) * 2023-07-13 2025-01-16 주식회사 클라리파이 3차원 의료영상의 단면간 해상도 향상 장치 및 방법

Also Published As

Publication number Publication date
US20210374911A1 (en) 2021-12-02
JPWO2020175446A1 (ja) 2021-12-23
EP3932318A1 (en) 2022-01-05
EP3932318A4 (en) 2022-04-20
US12217387B2 (en) 2025-02-04
JP7105363B2 (ja) 2022-07-22
EP3932318B1 (en) 2026-01-14

Similar Documents

Publication Publication Date Title
WO2020175446A1 (ja) 学習方法、学習システム、学習済みモデル、プログラム及び超解像画像生成装置
CN109978037B (zh) 图像处理方法、模型训练方法、装置、和存储介质
US9892361B2 (en) Method and system for cross-domain synthesis of medical images using contextual deep network
Kudo et al. Virtual thin slice: 3D conditional GAN-based super-resolution for CT slice interval
CN108701220B (zh) 用于处理多模态图像的系统和方法
CN110036409B (zh) 使用联合深度学习模型进行图像分割的系统和方法
Du et al. Accelerated super-resolution MR image reconstruction via a 3D densely connected deep convolutional neural network
US20240005498A1 (en) Method of generating trained model, machine learning system, program, and medical image processing apparatus
JP7795352B2 (ja) 画像処理方法、画像処理装置およびプログラム
CN113313728B (zh) 一种颅内动脉分割方法及系统
JP2021184594A (ja) ビデオフレームの補間装置及び方法
Bhardwaj et al. A review in wavelet transforms based medical image fusion
US12579720B2 (en) Method of generating trained model, machine learning system, program, and medical image processing apparatus
CN116348911A (zh) 图像分割方法和系统
CN117746034A (zh) 基于频率感知和特征交互的超声图像分割系统及方法
CN114586065A (zh) 用于分割图像的方法和系统
JP2023505676A (ja) 機械学習方法のデータ拡張
JPWO2020175445A1 (ja) 学習方法、学習装置、生成モデル及びプログラム
US20230046302A1 (en) Blood flow field estimation apparatus, learning apparatus, blood flow field estimation method, and program
CN115272250A (zh) 确定病灶位置方法、装置、计算机设备和存储介质
Ghoshal et al. Fast 3D Volumetric Image Reconstruction from 2D MRI Slices by Parallel Processing
WO2023032438A1 (ja) 回帰推定装置および方法、プログラム並びに学習済みモデルの生成方法
US12579716B2 (en) MRI reconstruction based on contrastive learning
JP2025149188A (ja) 画像処理装置、方法およびプログラム、学習装置、方法およびプログラム並びに解析装置
Shen Prior-informed machine learning for biomedical imaging and perception

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20762514

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021502251

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 2020762514

Country of ref document: EP

Effective date: 20210928

WWG Wipo information: grant in national office

Ref document number: 2020762514

Country of ref document: EP