WO2022101962A1 - 学習装置、学習方法およびプログラム - Google Patents
学習装置、学習方法およびプログラム Download PDFInfo
- Publication number
- WO2022101962A1 WO2022101962A1 PCT/JP2020/041850 JP2020041850W WO2022101962A1 WO 2022101962 A1 WO2022101962 A1 WO 2022101962A1 JP 2020041850 W JP2020041850 W JP 2020041850W WO 2022101962 A1 WO2022101962 A1 WO 2022101962A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- label
- feature amount
- data
- unit
- reconstruction
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
- G06N3/0455—Auto-encoder networks; Encoder-decoder networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/084—Backpropagation, e.g. using gradient descent
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/774—Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/98—Detection or correction of errors, e.g. by rescanning the pattern or by human intervention; Evaluation of the quality of the acquired patterns
Definitions
- the present invention relates to a learning device, a learning method, and a program technique.
- an object of the present invention is to provide a technique capable of clearly separating data into arbitrary features.
- One aspect of the present invention is a classification unit that classifies latent variables, which are features obtained from learning data used for training, using label features having label information used for classification, and decodes the latent variables.
- the decoding parameter is optimized so as to minimize the classification error between the label feature amount and the non-label feature amount by using the decoding unit that generates the reconstructed data using the predetermined decoding parameter and the label feature amount. It is a learning device provided with an optimization unit to be optimized.
- the classification unit classifies latent variables, which are features obtained from the learning data used for training, using the label feature amount having the label information used for classification, and the decoding unit describes the above.
- the latent variable is decoded and the reconstruction data is generated using a predetermined decoding parameter, and the optimization unit uses the label feature amount to minimize the classification error between the label feature amount and the non-label feature amount. It is a learning method that optimizes the decoding parameters as described above.
- One aspect of the present invention includes a step in which a computer extracts a feature amount from the target data, a reconstruction step in which the extracted feature amount is reconstructed and the reconstructed data is acquired, and the target data and the reconstructed data. It has a step of outputting the reconstruction error which is the difference from the above as the degree to which the target data has the characteristics common to the predetermined data group, and the reconstruction step is from the data belonging to the predetermined data group.
- the obtained feature amount was separated into a first partial feature amount and a second partial feature amount, and the second partial feature amount was extracted from another data belonging to the predetermined data group. Learning that exchanges with the second partial feature amount, acquires the feature amount after exchange, and optimizes so that the difference between the data obtained by reconstructing the feature amount after exchange and the data belonging to the predetermined data group becomes small.
- the method is a step in which a computer extracts a feature amount from the target data, a reconstruction step in which the extracted feature amount is reconstructed and the reconstructed data is acquired, and
- a computer is made to classify a latent variable, which is a feature amount obtained from training data used for training, using a label feature amount having label information used for classification, and the latent variable is decoded.
- the reconstruction data is generated using a predetermined decoding parameter, and the label feature amount is used to optimize the decoding parameter so as to minimize the classification error between the label feature amount and the non-label feature amount. It is a program.
- FIG. 1 is a diagram showing an example of the configuration of the learning device of the embodiment.
- the learning device 1 includes a sampling unit 11, a classification unit 2, a processing unit 3, and an optimization unit 27.
- the classification unit 2 includes an encoding unit 12, a label feature amount extraction unit 13, and a non-label feature amount extraction unit 14.
- the processing unit 3 includes a label feature amount exchange unit 15, a feature coupling unit 16, a decoding unit 17, a reconstruction error calculation unit 18, a decoding unit 19, a reconstruction error calculation unit 20, a non-label feature quantity exchange unit 21, and a feature coupling unit. It includes 22, a decoding unit 23, an encoding unit 24, a label feature amount extraction unit 25, and a classification error calculation unit 26.
- the learning device 1 separates the input data into a label feature amount and a feature amount other than the label.
- the sampling unit 11 samples the input data ⁇ x 1 , y 1 ⁇ , ..., ⁇ x B , y B ⁇ of the batch size B (B is an integer of 1 or more) from the learning data ⁇ x i , y i ⁇ .
- z i, label is a label feature quantity z i
- label [zi , 1 , ..., z i, C ] composed of C parameters (C is an integer of 1 or more)
- Wo_label is a feature quantity z i
- wo_label [zi , C + 1 , ..., Z i, M ] other than the label composed of MC parameters (M is an integer of 2 or more).
- the encoding unit 12 outputs the feature amount 101 to the label feature amount extracting unit 13, the non-label feature amount extracting unit 14, and the decoding unit 19.
- the latent variable is a feature quantity obtained by encoding when an autoencoder is used.
- the label feature amount extraction unit 13 extracts the label feature amount 102 ⁇ zi , label ⁇ .
- the label feature amount extraction unit 13 outputs the extracted label feature amount 102 to the label feature amount exchange unit 15, the feature coupling unit 22, and the classification error calculation unit 26.
- the non-label feature amount extraction unit 14 extracts the non-label feature amount 103 ⁇ zi , wo_label ⁇ .
- the feature amount extraction unit 14 other than the label outputs the extracted feature amount 103 other than the label to the feature combination unit 16, the feature amount exchange unit 21 other than the label, and the feature connection unit 22.
- the label information attached to the learning data and the label feature amount 102 are input to the label feature amount exchange unit 15.
- the label feature amount exchange unit 15 randomly exchanges (swaps) each parameter of the label feature amounts zi and label with the same label sample in the batch process.
- the exchanged one is referred to as (zi , label ) swap .
- the label feature amount exchange unit 15 outputs the exchanged label feature amount 104 to the feature coupling unit 16.
- the label feature amount exchange unit 15 may be exchanged with another sample having the same label, not limited to batch processing.
- the feature combining unit 16 combines the label feature amount 104 exchanged by the label feature amount exchange unit 15 and the non-label feature amount 103 extracted by the non-label feature amount extraction unit 14, and decodes the combined feature amount. Output to 17.
- the decoding unit 17 decodes the feature amount to obtain the reconstructed data 105 ⁇ ( xi ) (swap_label) ⁇ ⁇ .
- the decoding unit 17 outputs the reconstruction data 105 to the reconstruction error calculation unit 18.
- the reconstruction error calculation unit 18 calculates the reconstruction error 106 ⁇ L rec, swap ⁇ between the input data x i and the reconstruction data ( xi ) ⁇ obtained by decoding by the following equation (1).
- d is an arbitrary function for calculating the distance between two vectors, for example, a mean square error sum, a mean absolute error sum, or the like.
- the reconstruction error calculation unit 18 outputs the calculated reconstruction error 106 to the optimization unit 27.
- the decoding unit 19 decodes the feature amount 101 to obtain the reconstructed data 107 ⁇ (x i ) ⁇ ⁇ .
- the decoding unit 19 outputs the reconstruction data 107 to the reconstruction error calculation unit 20.
- the reconstruction error calculation unit 20 uses the following equation (2) to obtain a reconstruction error 108 ⁇ L rec, org ⁇ between the input data x i and the reconstruction data ( zi ) (swap_label) ⁇ output by the decoding unit 19. calculate.
- the non-label feature amount exchange unit 21 randomly exchanges each parameter of the non-label feature amount zi and wo_label with the sample in the batch process.
- the exchanged one is referred to as (zi , wo_label ) swap .
- the feature amount exchange unit 21 other than the label generates a feature amount ⁇ ( zi ) swap_wo_label ⁇ in which the label feature amount zi , label is combined with the exchanged (zi , wo_label ) swap .
- the feature amount exchange unit 21 other than the label outputs the feature amount 110 other than the exchanged label to the feature coupling unit 22.
- the feature combining unit 22 combines the label feature amount 102 extracted by the label feature amount extraction unit 13 with the non-label feature amount 110 exchanged by the non-label feature amount exchange unit 21.
- the feature coupling unit 22 outputs the combined feature amount to the decoding unit 23.
- the decoding unit 23 decodes the combined feature amount ⁇ ( zi ) swap_wo_label ⁇ to obtain the reconstructed data 111 ⁇ ( xi ) (swap_wo_label) ⁇ ⁇ .
- the decoding unit 23 outputs the reconstructed data 111 to the encoding unit 24.
- the encoding unit 24 re-encodes the reconstructed data 111 ⁇ (x i ) (swap_wo_label) ⁇ ⁇ to obtain the feature amount 112.
- the encoding unit 24 outputs the feature amount 112 to the label feature amount extraction unit 25.
- the label feature amount extraction unit 25 extracts the label feature amount ⁇ (zi , label ) (swap_wo_label) ⁇ ⁇ from the feature amount 112, and outputs the extracted label feature amount 113 to the classification error calculation unit 26.
- Label information, the label feature amount 102 extracted by the label feature amount extraction unit 13, and the label feature amount 113 extracted by the label feature amount extraction unit 25 are input to the classification error calculation unit 26.
- the classification error calculation unit 26 calculates the classification error 109 ⁇ L label, org ⁇ from the label feature amount 102 ⁇ zi , label ⁇ by the following equation (3).
- (z yi, label ) ⁇ is the average of the label features z i, label of the sample whose label information is y i in the batch sample
- K is the number of classification labels. be.
- the classification error calculation unit 26 calculates the classification error 114 ⁇ L label, swap ⁇ from the label feature amount 113 ⁇ (zi , label ) (swap_wo_label) ⁇ ⁇ by the following equation (4).
- the optimization unit 27 calculates the objective function L weighted by each error by the following equation (5).
- ⁇ is a predetermined weighting coefficient.
- the optimization unit 27 updates the parameters of the encoding unit (12, 24) and the decoding unit (17, 19, 23) by, for example, the gradient method.
- the optimization unit 27 determines, for example, whether or not the objective function L has converged, or whether or not the predetermined number of processes has been completed.
- FIG. 1 the configuration and processing shown in FIG. 1 are examples, and are not limited to this. Further, the configuration of FIG. 1 includes a functional unit that is used and a functional unit that is not used, depending on the intended use. Further, the encoding units 17, 19 and 23 may be integrated or separate. The feature coupling portions 18 and 22 may be integrated or separate. The reconstruction error calculation units 18 and 20 may be integrated or separate.
- the learning device 1 is configured by using a processor such as a CPU (Central Processing Unit) and a memory, for example.
- the learning device 1 functions as a sampling unit 11, an encoding unit 2, a classification unit 3, and an optimization unit 27 by executing a program by the processor. All or part of each function of the learning device 1 may be realized by using hardware such as ASIC (Application Specific Integrated Circuit), PLD (Programmable Logic Device), and FPGA (Field Programmable Gate Array).
- the above program may be recorded on a computer-readable recording medium.
- Computer-readable recording media include, for example, flexible disks, magneto-optical disks, ROMs, CD-ROMs, portable media such as semiconductor storage devices (for example, SSD: Solid State Drive), hard disks and semiconductor storage built into computer systems. It is a storage device such as a device.
- the above program may be transmitted over a telecommunication line.
- FIG. 2 is a diagram showing an outline of the processing of the present embodiment.
- the encoder g102 corresponds to the encoding unit 12 in FIG.
- the encoder g102 and the decoder g105 are, for example, autoencoders.
- Input data g101 is input to the encoder g102.
- the learning device 1 performs learning by regarding the bottleneck portion of the autoencoder as a feature.
- the label feature amount extraction unit 13 and the non-label feature amount extraction unit 14 separate the features into two, a label feature amount g103 and a non-label feature amount g104.
- the label feature amount g103 and the non-label feature amount g104 are input to the decoder g105.
- the decoder g105 corresponds to the decoding unit 19 in FIG.
- the optimization unit 27 minimizes the classification error (CE loss; Cross-entropy loss) by using the label feature amount g103.
- the optimization unit 27 uses the label feature amount g103 and the non-label feature amount g104 to minimize the reconstruction error.
- FIG. 3 is a flowchart showing an example of processing procedures at the time of learning and at the time of classification according to the present embodiment.
- the sampling unit 11 samples the input data of batch size B from the learning data (step S11).
- the encoding unit 12 encodes the input data to obtain a feature amount (step S12).
- the label feature amount extraction unit 13 extracts the label feature amount, and the non-label feature amount extraction unit 14 extracts the feature amount other than the label, thereby separating the feature amount into two (step S13).
- the optimization unit 27 minimizes the classification error by using the label feature amount g103 (step S14).
- the optimization unit 27 minimizes the reconstruction error by using the label feature amount g103 and the non-label feature amount g104 (step S15).
- the optimization unit 27 updates the parameters of the encoding unit (12, 24) and the decoding unit (17, 19, 23) by, for example, the gradient method (step S16).
- the optimization unit 27 determines, for example, whether or not the objective function L has converged, or whether or not the predetermined number of processes has been completed (step S16).
- the optimization unit 27 ends the processing when the objective function L converges or when the processing for a predetermined number of times is completed (step S17; YES).
- the optimization unit 27 repeats the processes of steps S11 to S16 when the objective function L has not converged or the predetermined number of processes have not been completed (step S17; NO).
- FIGS. 4 to 6 show an example showing the effect of this embodiment.
- the learning data and the data to be classified are examples of image data.
- the label feature amount is a type of a number (0 to 9), and the feature amount other than the label is a number shape.
- FIG. 4 is a diagram showing an example of a label feature amount and a feature amount other than the label according to the present embodiment.
- the vertical axis is the label feature amount g201 and the non-label feature amount g202.
- the horizontal direction is the original image g203 and the reconstructed image g204 when the features are changed. The image in the frame g205 will be described later.
- FIG. 5 is a diagram showing an example of an original image and a reconstructed image according to the present embodiment. In the horizontal direction, the original images g211 and g213 and the reconstructed images g212 and g214 are shown.
- FIG. 6 is a diagram showing an example of a reconstructed image when a non-label feature amount is exchanged with the original image according to the present embodiment.
- the original images g221 and g223 are reconstructed images g212 and g214 when other than the label feature amount is exchanged.
- the image in the frame g225 will be described later.
- the features are separated into two, a label feature and a feature other than the label. Further, in the learning device 1, the classification error is minimized by using the label feature amount. Further, in the learning device 1, the reconstruction error is minimized by using the label feature amount and the feature amount other than the label.
- the label information can be clearly extracted as an expression on a continuous space.
- FIG. 7 is a diagram showing an outline of the processing of the present embodiment.
- the encoder g107 corresponds to the encoding unit 24 of FIG.
- the encoder g107 is, for example, an autoencoder.
- the reconstructed data g106 is input to the encoder g107.
- the encoder g102 and the encoder g107 may be integrated or separate.
- the feature amount exchange unit 21 other than the label exchanges the feature amount other than the label between batches.
- the decoding unit 23 decodes the feature amount obtained by combining the feature amount other than the label exchanged with the label feature amount.
- the encoding unit 24 re-encodes the decoded features.
- the optimization unit 27 minimizes the classification error by using the label feature amount g103' obtained as a result of re-encoding.
- FIG. 8 is a flowchart showing an example of processing procedures at the time of learning and at the time of classification according to the present embodiment.
- the learning device 1 performs the processes of steps S11 to S13. Subsequently, the feature amount exchange unit 21 other than the label exchanges the feature amount other than the label between batches (step S21).
- the decoding unit 23 decodes the feature amount obtained by combining the feature amount other than the label exchanged with the label feature amount (step S22).
- the encoding unit 24 re-encodes the decoded features (step S23).
- the optimization unit 27 minimizes the classification error by using the re-encoded label feature amount g103'(step S24).
- the learning device 1 performs the processes of steps S16 to S17.
- FIGS. 9 to 11 an example showing the effect of this embodiment is shown in FIGS. 9 to 11.
- the learning data and the data to be classified are examples of image data.
- FIG. 9 is a diagram showing an example of a label feature amount and a non-label feature amount according to the present embodiment.
- FIG. 10 is a diagram showing an example of an original image and a reconstructed image according to the present embodiment.
- FIG. 11 is a diagram showing an example of a reconstructed image when a non-label feature amount is exchanged with the original image according to the present embodiment.
- the numbers do not change to other numbers, that is, the label information is not included in the features other than the label.
- the features are separated into two, a label feature and a feature other than the label. Further, in the learning device 1, features other than labels are exchanged between batches. Further, in the learning device 1, the exchanged data is decoded and the decoded reconstruction data is re-encoded. Further, in the learning device 1, the label feature amount g103' obtained by re-encoding was used to minimize the classification error.
- label information is included in features other than labels, it may result in different label data when reconstructed.
- FIG. 12 is a diagram showing an outline of the processing of the present embodiment.
- the label feature amount exchange unit 15 randomly exchanges the label feature amount between the same labels in the batch.
- the decoding unit 17 decodes the feature amount obtained by combining the exchanged label feature amount and the feature amount other than the label.
- the optimization unit 27 uses the reconstruction data decoded by the decoding unit 17 to minimize the reconstruction error.
- FIG. 13 is a flowchart showing an example of processing procedures at the time of learning and at the time of classification according to the third embodiment.
- the learning device 1 performs the processes of steps S11 to S13.
- the label feature amount exchange unit 15 randomly exchanges the label feature amount g103 between the same labels in the batch (step S31).
- the decoding unit 17 decodes the feature amount obtained by combining the exchanged label feature amount g103 and the non-label feature amount g104 (step S32).
- the optimization unit 27 uses the exchanged and decoded reconstruction data to minimize the reconstruction error (step S33).
- the learning device 1 performs the processes of steps S16 to S17.
- the features are separated into two, a label feature and a feature other than the label. Further, in the learning device 1, the label feature amount is exchanged between the same labels in the batch. Further, in the learning device 1, the exchanged data is decoded, and the decoded reconstruction data is used to minimize the reconstruction error.
- the label feature amount is exchanged with other same label data and reconstructed.
- this reconstruction since only the label information must be included in the exchanged label feature amount, it is possible to prevent the label feature amount from containing information other than the label.
- common features between the exchanged samples can be extracted.
- the training data without label information is divided into two features (first partial feature amount (label feature amount) and second partial feature amount (feature amount other than label)), and the label feature amount is determined.
- the common feature is, for example, in the case of an image group of a dog, the information of a dog is a common feature, and in the case of an image group of handwritten characters of a certain person, the information of how to write the person is a common feature.
- a natural image such as Image, which is a data set
- the concept of a natural image is a common feature.
- the present embodiment can be applied to the learning data to which the label is not attached.
- the learning device 1 extracts, for example, the feature amount from the target data, reconstructs the extracted feature amount, acquires the reconstructed data, and reconstructs the difference between the target data and the reconstructed data.
- the configuration error is output as the degree to which the target data has the characteristics that the predetermined data group has in common.
- the learning device 1 separates the feature amount obtained from the data belonging to a predetermined data group into a first partial feature amount and a second partial feature amount, and the second part is described above.
- the feature amount is exchanged with a second partial feature amount extracted from another data belonging to a predetermined data group, and the feature amount after exchange is acquired. Then, the learning device 1 optimizes so that the difference between the data in which the feature amount is reconstructed after the exchange and the data belonging to the predetermined data group becomes small.
- FIGS. 14 to 16 show an example showing the effect when the processing of the second embodiment and the processing of the present embodiment are performed in addition to the first embodiment.
- the learning data and the data to be classified are examples of image data.
- FIG. 14 is a diagram showing an example of a label feature amount and a non-label feature amount when the processing of the second embodiment and the processing of the present embodiment are performed in addition to the first embodiment.
- FIG. 15 is a diagram showing an example of an original image and a reconstructed image when the processing of the second embodiment and the processing of the present embodiment are performed in addition to the first embodiment.
- FIG. 16 is a diagram showing an example of a reconstructed image when the processing of the second embodiment and the processing of the present embodiment are performed in addition to the first embodiment and the original image and the label feature amount are exchanged. be.
- the label information is not included in the features other than the label. Further, when the processing of the present embodiment is performed in addition to the first embodiment, information other than the label features is not included in the label feature amount. Thereby, according to the second embodiment and the present embodiment, the label information and the non-label information can be clearly separated.
- the data to be separated from the features is not limited to the image data, but may be other data. Further, the image data may be a still image or a moving image.
- each of the above-described embodiments since the data can be separated into arbitrary features, it is possible to generate data having specific features or edit and reconstruct the specific features. Thereby, each of the above-described embodiments can generate or edit data for any feature (data disentanglement).
- the label information and other information can be separated, and the label information can be extracted as a value in a continuous space, so that it can be applied to recognition of unlearned classes and the like.
- each of the above-described embodiments can improve the accuracy of Few-shot learning for recognizing a class of a small number of data.
- the present invention can be applied to data feature separation, data generation, data editing, data class recognition, transfer learning, and the like.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Evolutionary Computation (AREA)
- Computing Systems (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Software Systems (AREA)
- General Health & Medical Sciences (AREA)
- Multimedia (AREA)
- Biophysics (AREA)
- Molecular Biology (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Mathematical Physics (AREA)
- Computational Linguistics (AREA)
- Biomedical Technology (AREA)
- Life Sciences & Earth Sciences (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Databases & Information Systems (AREA)
- Medical Informatics (AREA)
- Quality & Reliability (AREA)
- Image Analysis (AREA)
Abstract
Description
上記事情に鑑み、本発明は、データを任意の特徴に明確に分離することができる技術の提供を目的としている。
図1は、実施形態の学習装置の構成の一例を示す図である。図1のように、学習装置1は、サンプリング部11、分類部2、処理部3、および最適化部27を備える。
本実施形態では、エンコード部12が同一層で特徴を分離する。なお、本実施形態では、バッチ内で交換させない。
図2は、本実施形態の処理の概要を示す図である。エンコーダg102は、図1のエンコード部12に対応する。エンコーダg102とデコーダg105は、例えばオートエンコーダである。エンコーダg102には、入力データg101が入力される。
ラベル特徴量抽出部13とラベル以外特徴量抽出部14は、特徴をラベル特徴量g103とラベル以外特徴量g104との2つに分離する。
ラベル特徴量g103とラベル以外特徴量g104とは、デコーダg105に入力される。デコーダg105は、図1のデコード部19に対応する。
最適化部27は、ラベル特徴量g103とラベル以外特徴量g104とを用いて、再構成誤差を最小化する。
図3は、本実施形態に係る学習時と分類時の処理手順例を示すフローチャートである。
ラベル以外の特徴から、さらに精度よくラベル特徴を除外するテクニックを本実施形態で示す。ラベル以外の特徴にラベル特徴が含まれると、デコードした結果得られる出力値が、違うラベルの出力値になると考えられる。また、同じラベルをもつデータ間であれば、ラベル以外の特徴を交換しても同じクラスの出力値にデコードされる。そこで、本実施形態では、学習装置1が、ラベル以外の特徴をバッチ内で交換させて学習する。
ラベル以外特徴量交換部21は、ラベル以外特徴量をバッチ間で交換する。
デコード部23は、ラベル特徴量と交換されたラベル以外特徴量を結合した特徴量をデコードする。
エンコード部24は、デコードされた特徴量を再エンコードする。
最適化部27は、再エンコードした結果得られたラベル特徴量g103’を用いて、クラス分類誤差を最小化する。
図8は、本実施形態に係る学習時と分類時の処理手順例を示すフローチャートである。
続けて、ラベル以外特徴量交換部21は、ラベル以外特徴量をバッチ間で交換する(ステップS21)。デコード部23は、ラベル特徴量と交換されたラベル以外特徴量を結合した特徴量をデコードする(ステップS22)。エンコード部24は、デコードされた特徴量を再エンコードする(ステップS23)。
続けて、学習装置1は、ステップS16~S17の処理を行う。
図9は、本実施形態に係るラベル特徴量とラベル以外特徴量の一例を示す図である。図10は、本実施形態に係る原画と再構成した画像の一例を示す図である。図11は、本実施形態に係る原画とラベル特徴量以外を交換した時の再構成した画像の一例を示す図である。
ラベル特徴量から、さらにラベル特徴量以外の情報を取り除くテクニックを本実施形態で説明する。同一のラベルが付与されるデータ間であれば、ラベル特徴を交換してもデコードされた結果得られるクラスは同一である。そこで、本実施形態では、学習装置1が、ラベル特徴をバッチ内の同一ラベル間で交換させて学習する。図12は、本実施形態の処理の概要を示す図である。
ラベル特徴量交換部15は、ラベル特徴量をバッチ内の同一ラベル間でランダムに交換する。
デコード部17は、交換されたラベル特徴量とラベル以外特徴量を結合した特徴量をデコードする。
最適化部27は、デコード部17でデコードされた再構成データを用いて、再構成誤差を最小化する。
ラベル特徴量交換部15は、ラベル特徴量g103をバッチ内の同一ラベル間でランダムに交換する(ステップS31)。デコード部17は、交換されたラベル特徴量g103とラベル以外特徴量g104を結合した特徴量をデコードする(ステップS32)。
学習装置1は、ステップS16~S17の処理を行う。
図14は、第1の実施形態に加えて第2の実施形態の処理と本実施形態の処理を行う場合のラベル特徴量とラベル以外特徴量の一例を示す図である。図15は、第1の実施形態に加えて第2の実施形態の処理と本実施形態の処理を行う場合の原画と再構成した画像の一例を示す図である。図16は、第1の実施形態に加えて第2の実施形態の処理と本実施形態の処理を行う場合の原画とラベル特徴量以外を交換した時の再構成した画像の一例を示す図である。
なお、上述した各実施例において、特徴を分離する対象のデータが画像データに限らず、他のデータであってもよい。また、画像データは、静止画であっても動画であってもよい。
Claims (7)
- 学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類する分類部と、
前記潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成するデコード部と、
前記ラベル特徴量を用いて、前記ラベル特徴量と前記ラベル情報との分類誤差を最小化するように前記デコードパラメータを最適化する最適化部と、
を備える学習装置。 - 前記ラベル特徴量は、C(Cは1以上の整数)個のパラメータから構成され、
前記ラベル特徴量の各パラメータを、バッチ処理内の同一ラベルの前記学習データとランダムに交換するラベル特徴量交換部と、
前記交換されたラベル特徴量と、ラベル以外特徴量とを結合する特徴結合部と、
前記潜在変数と、前記結合された特徴量を前記デコード部によってデコードして生成された再構成データ再構成データとの再構成誤差を算出する再構成誤差算出部と、を更に備え、
請求項1に記載の学習装置。 - 前記分類部は、オートエンコーダを備える、
請求項1または請求項2に記載の学習装置。 - 分類部が、学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類し、
デコード部が、前記潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成し、
最適化部が、前記ラベル特徴量を用いて、前記ラベル特徴量と前記ラベル情報との分類誤差を最小化するように前記デコードパラメータを最適化する、
学習方法。 - コンピュータが、
対象データから特徴量を抽出するステップと、
抽出された特徴量を再構成し再構成データを取得する再構成ステップと、
前記対象データと前記再構成データとの差である再構成誤差を、所定のデータ群が共通して有する特徴を前記対象データが有する度合いとして出力するステップと、を有し、
前記再構成ステップは、
前記所定のデータ群に属するデータから得られた特徴量を、第一の部分特徴量と、第二の部分特徴量と、に分離し、
前記第二の部分特徴量を、前記所定のデータ群に属する別のデータから抽出された第二の部分特徴量と交換し、交換後特徴量を取得し、前記交換後特徴量を再構成したデータと、前記所定のデータ群に属するデータとの差が小さくなるよう最適化する、
学習方法。 - コンピュータに、
学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類させ、
前記潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成させ、
前記ラベル特徴量を用いて、前記ラベル特徴量と前記ラベル情報との分類誤差を最小化するように前記デコードパラメータを最適化させる、
プログラム。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2020/041850 WO2022101962A1 (ja) | 2020-11-10 | 2020-11-10 | 学習装置、学習方法およびプログラム |
| JP2022561708A JP7513918B2 (ja) | 2020-11-10 | 2020-11-10 | 学習装置、学習方法およびプログラム |
| US18/035,540 US20230410472A1 (en) | 2020-11-10 | 2020-11-10 | Learning device, learning method and program |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2020/041850 WO2022101962A1 (ja) | 2020-11-10 | 2020-11-10 | 学習装置、学習方法およびプログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022101962A1 true WO2022101962A1 (ja) | 2022-05-19 |
Family
ID=81600893
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2020/041850 Ceased WO2022101962A1 (ja) | 2020-11-10 | 2020-11-10 | 学習装置、学習方法およびプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20230410472A1 (ja) |
| JP (1) | JP7513918B2 (ja) |
| WO (1) | WO2022101962A1 (ja) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2019061512A (ja) * | 2017-09-27 | 2019-04-18 | 株式会社Abeja | データの特徴を利用してデータを処理するシステム |
| JP2019200551A (ja) * | 2018-05-15 | 2019-11-21 | 株式会社日立製作所 | データから潜在因子を発見するニューラルネットワーク |
| JP2020160743A (ja) * | 2019-03-26 | 2020-10-01 | 日本電信電話株式会社 | 評価装置、評価方法、および、評価プログラム |
-
2020
- 2020-11-10 JP JP2022561708A patent/JP7513918B2/ja active Active
- 2020-11-10 WO PCT/JP2020/041850 patent/WO2022101962A1/ja not_active Ceased
- 2020-11-10 US US18/035,540 patent/US20230410472A1/en not_active Abandoned
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2019061512A (ja) * | 2017-09-27 | 2019-04-18 | 株式会社Abeja | データの特徴を利用してデータを処理するシステム |
| JP2019200551A (ja) * | 2018-05-15 | 2019-11-21 | 株式会社日立製作所 | データから潜在因子を発見するニューラルネットワーク |
| JP2020160743A (ja) * | 2019-03-26 | 2020-10-01 | 日本電信電話株式会社 | 評価装置、評価方法、および、評価プログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| JP7513918B2 (ja) | 2024-07-10 |
| JPWO2022101962A1 (ja) | 2022-05-19 |
| US20230410472A1 (en) | 2023-12-21 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Ahmed et al. | Classification and reconstruction of optical quantum states with deep neural networks | |
| CN112734634B (zh) | 换脸方法、装置、电子设备和存储介质 | |
| Gu et al. | Lofgan: Fusing local representations for few-shot image generation | |
| CN114612289B (zh) | 风格化图像生成方法、装置及图像处理设备 | |
| Tirupattur et al. | Thoughtviz: Visualizing human thoughts using generative adversarial network | |
| Perarnau et al. | Invertible conditional gans for image editing | |
| Qi et al. | Avt: Unsupervised learning of transformation equivariant representations by autoencoding variational transformations | |
| CN118053090A (zh) | 使用潜在扩散模型生成视频 | |
| Creswell et al. | Adversarial information factorization | |
| US20200074273A1 (en) | Method for training deep neural network (dnn) using auxiliary regression targets | |
| KR20200052453A (ko) | 딥러닝 모델 학습 장치 및 방법 | |
| KR102332114B1 (ko) | 이미지 처리 방법 및 장치 | |
| CN116612416A (zh) | 一种指代视频目标分割方法、装置、设备及可读存储介质 | |
| JP7436928B2 (ja) | 学習装置、学習方法およびプログラム | |
| Jorge et al. | Empirical Evaluation of Variational Autoencoders for Data Augmentation. | |
| CN120597947A (zh) | 相关于数据生成框架的方法及装置 | |
| Tsoumplekas et al. | A complete survey on contemporary methods, emerging paradigms and hybrid approaches for few-shot learning | |
| Wang et al. | Denoising reuse: Exploiting inter-frame motion consistency for efficient video generation | |
| WO2024039572A1 (en) | Artificial intelligence computing systems for efficiently learning underlying features of data | |
| JP7513918B2 (ja) | 学習装置、学習方法およびプログラム | |
| JP2019082847A (ja) | データ推定装置、データ推定方法及びプログラム | |
| WO2025049487A1 (en) | Systems and methods for enhancing ultrasound images using artificial intelligence | |
| JP7376812B2 (ja) | データ生成方法、データ生成装置及びプログラム | |
| Nguyen et al. | Qc-stylegan-quality controllable image generation and manipulation | |
| Fakhari et al. | An image restoration architecture using abstract features and generative models |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20961490 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2022561708 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18035540 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20961490 Country of ref document: EP Kind code of ref document: A1 |





