WO2022101962A1 - 学習装置、学習方法およびプログラム - Google Patents

学習装置、学習方法およびプログラム Download PDF

Info

Publication number
WO2022101962A1
WO2022101962A1 PCT/JP2020/041850 JP2020041850W WO2022101962A1 WO 2022101962 A1 WO2022101962 A1 WO 2022101962A1 JP 2020041850 W JP2020041850 W JP 2020041850W WO 2022101962 A1 WO2022101962 A1 WO 2022101962A1
Authority
WO
WIPO (PCT)
Prior art keywords
label
feature amount
data
unit
reconstruction
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2020/041850
Other languages
English (en)
French (fr)
Inventor
忍 工藤
隆一 谷田
英明 木全
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to PCT/JP2020/041850 priority Critical patent/WO2022101962A1/ja
Priority to JP2022561708A priority patent/JP7513918B2/ja
Priority to US18/035,540 priority patent/US20230410472A1/en
Publication of WO2022101962A1 publication Critical patent/WO2022101962A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/04—Architecture, e.g. interconnection topology
    • G06N3/045—Combinations of networks
    • G06N3/0455—Auto-encoder networks; Encoder-decoder networks
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/08—Learning methods
    • G06N3/084—Backpropagation, e.g. using gradient descent
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/08—Learning methods
    • G06N3/09—Supervised learning
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/40—Extraction of image or video features
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/774—Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/98—Detection or correction of errors, e.g. by rescanning the pattern or by human intervention; Evaluation of the quality of the acquired patterns

Definitions

  • the present invention relates to a learning device, a learning method, and a program technique.
  • an object of the present invention is to provide a technique capable of clearly separating data into arbitrary features.
  • One aspect of the present invention is a classification unit that classifies latent variables, which are features obtained from learning data used for training, using label features having label information used for classification, and decodes the latent variables.
  • the decoding parameter is optimized so as to minimize the classification error between the label feature amount and the non-label feature amount by using the decoding unit that generates the reconstructed data using the predetermined decoding parameter and the label feature amount. It is a learning device provided with an optimization unit to be optimized.
  • the classification unit classifies latent variables, which are features obtained from the learning data used for training, using the label feature amount having the label information used for classification, and the decoding unit describes the above.
  • the latent variable is decoded and the reconstruction data is generated using a predetermined decoding parameter, and the optimization unit uses the label feature amount to minimize the classification error between the label feature amount and the non-label feature amount. It is a learning method that optimizes the decoding parameters as described above.
  • One aspect of the present invention includes a step in which a computer extracts a feature amount from the target data, a reconstruction step in which the extracted feature amount is reconstructed and the reconstructed data is acquired, and the target data and the reconstructed data. It has a step of outputting the reconstruction error which is the difference from the above as the degree to which the target data has the characteristics common to the predetermined data group, and the reconstruction step is from the data belonging to the predetermined data group.
  • the obtained feature amount was separated into a first partial feature amount and a second partial feature amount, and the second partial feature amount was extracted from another data belonging to the predetermined data group. Learning that exchanges with the second partial feature amount, acquires the feature amount after exchange, and optimizes so that the difference between the data obtained by reconstructing the feature amount after exchange and the data belonging to the predetermined data group becomes small.
  • the method is a step in which a computer extracts a feature amount from the target data, a reconstruction step in which the extracted feature amount is reconstructed and the reconstructed data is acquired, and
  • a computer is made to classify a latent variable, which is a feature amount obtained from training data used for training, using a label feature amount having label information used for classification, and the latent variable is decoded.
  • the reconstruction data is generated using a predetermined decoding parameter, and the label feature amount is used to optimize the decoding parameter so as to minimize the classification error between the label feature amount and the non-label feature amount. It is a program.
  • FIG. 1 is a diagram showing an example of the configuration of the learning device of the embodiment.
  • the learning device 1 includes a sampling unit 11, a classification unit 2, a processing unit 3, and an optimization unit 27.
  • the classification unit 2 includes an encoding unit 12, a label feature amount extraction unit 13, and a non-label feature amount extraction unit 14.
  • the processing unit 3 includes a label feature amount exchange unit 15, a feature coupling unit 16, a decoding unit 17, a reconstruction error calculation unit 18, a decoding unit 19, a reconstruction error calculation unit 20, a non-label feature quantity exchange unit 21, and a feature coupling unit. It includes 22, a decoding unit 23, an encoding unit 24, a label feature amount extraction unit 25, and a classification error calculation unit 26.
  • the learning device 1 separates the input data into a label feature amount and a feature amount other than the label.
  • the sampling unit 11 samples the input data ⁇ x 1 , y 1 ⁇ , ..., ⁇ x B , y B ⁇ of the batch size B (B is an integer of 1 or more) from the learning data ⁇ x i , y i ⁇ .
  • z i, label is a label feature quantity z i
  • label [zi , 1 , ..., z i, C ] composed of C parameters (C is an integer of 1 or more)
  • Wo_label is a feature quantity z i
  • wo_label [zi , C + 1 , ..., Z i, M ] other than the label composed of MC parameters (M is an integer of 2 or more).
  • the encoding unit 12 outputs the feature amount 101 to the label feature amount extracting unit 13, the non-label feature amount extracting unit 14, and the decoding unit 19.
  • the latent variable is a feature quantity obtained by encoding when an autoencoder is used.
  • the label feature amount extraction unit 13 extracts the label feature amount 102 ⁇ zi , label ⁇ .
  • the label feature amount extraction unit 13 outputs the extracted label feature amount 102 to the label feature amount exchange unit 15, the feature coupling unit 22, and the classification error calculation unit 26.
  • the non-label feature amount extraction unit 14 extracts the non-label feature amount 103 ⁇ zi , wo_label ⁇ .
  • the feature amount extraction unit 14 other than the label outputs the extracted feature amount 103 other than the label to the feature combination unit 16, the feature amount exchange unit 21 other than the label, and the feature connection unit 22.
  • the label information attached to the learning data and the label feature amount 102 are input to the label feature amount exchange unit 15.
  • the label feature amount exchange unit 15 randomly exchanges (swaps) each parameter of the label feature amounts zi and label with the same label sample in the batch process.
  • the exchanged one is referred to as (zi , label ) swap .
  • the label feature amount exchange unit 15 outputs the exchanged label feature amount 104 to the feature coupling unit 16.
  • the label feature amount exchange unit 15 may be exchanged with another sample having the same label, not limited to batch processing.
  • the feature combining unit 16 combines the label feature amount 104 exchanged by the label feature amount exchange unit 15 and the non-label feature amount 103 extracted by the non-label feature amount extraction unit 14, and decodes the combined feature amount. Output to 17.
  • the decoding unit 17 decodes the feature amount to obtain the reconstructed data 105 ⁇ ( xi ) (swap_label) ⁇ ⁇ .
  • the decoding unit 17 outputs the reconstruction data 105 to the reconstruction error calculation unit 18.
  • the reconstruction error calculation unit 18 calculates the reconstruction error 106 ⁇ L rec, swap ⁇ between the input data x i and the reconstruction data ( xi ) ⁇ obtained by decoding by the following equation (1).
  • d is an arbitrary function for calculating the distance between two vectors, for example, a mean square error sum, a mean absolute error sum, or the like.
  • the reconstruction error calculation unit 18 outputs the calculated reconstruction error 106 to the optimization unit 27.
  • the decoding unit 19 decodes the feature amount 101 to obtain the reconstructed data 107 ⁇ (x i ) ⁇ ⁇ .
  • the decoding unit 19 outputs the reconstruction data 107 to the reconstruction error calculation unit 20.
  • the reconstruction error calculation unit 20 uses the following equation (2) to obtain a reconstruction error 108 ⁇ L rec, org ⁇ between the input data x i and the reconstruction data ( zi ) (swap_label) ⁇ output by the decoding unit 19. calculate.
  • the non-label feature amount exchange unit 21 randomly exchanges each parameter of the non-label feature amount zi and wo_label with the sample in the batch process.
  • the exchanged one is referred to as (zi , wo_label ) swap .
  • the feature amount exchange unit 21 other than the label generates a feature amount ⁇ ( zi ) swap_wo_label ⁇ in which the label feature amount zi , label is combined with the exchanged (zi , wo_label ) swap .
  • the feature amount exchange unit 21 other than the label outputs the feature amount 110 other than the exchanged label to the feature coupling unit 22.
  • the feature combining unit 22 combines the label feature amount 102 extracted by the label feature amount extraction unit 13 with the non-label feature amount 110 exchanged by the non-label feature amount exchange unit 21.
  • the feature coupling unit 22 outputs the combined feature amount to the decoding unit 23.
  • the decoding unit 23 decodes the combined feature amount ⁇ ( zi ) swap_wo_label ⁇ to obtain the reconstructed data 111 ⁇ ( xi ) (swap_wo_label) ⁇ ⁇ .
  • the decoding unit 23 outputs the reconstructed data 111 to the encoding unit 24.
  • the encoding unit 24 re-encodes the reconstructed data 111 ⁇ (x i ) (swap_wo_label) ⁇ ⁇ to obtain the feature amount 112.
  • the encoding unit 24 outputs the feature amount 112 to the label feature amount extraction unit 25.
  • the label feature amount extraction unit 25 extracts the label feature amount ⁇ (zi , label ) (swap_wo_label) ⁇ ⁇ from the feature amount 112, and outputs the extracted label feature amount 113 to the classification error calculation unit 26.
  • Label information, the label feature amount 102 extracted by the label feature amount extraction unit 13, and the label feature amount 113 extracted by the label feature amount extraction unit 25 are input to the classification error calculation unit 26.
  • the classification error calculation unit 26 calculates the classification error 109 ⁇ L label, org ⁇ from the label feature amount 102 ⁇ zi , label ⁇ by the following equation (3).
  • (z yi, label ) ⁇ is the average of the label features z i, label of the sample whose label information is y i in the batch sample
  • K is the number of classification labels. be.
  • the classification error calculation unit 26 calculates the classification error 114 ⁇ L label, swap ⁇ from the label feature amount 113 ⁇ (zi , label ) (swap_wo_label) ⁇ ⁇ by the following equation (4).
  • the optimization unit 27 calculates the objective function L weighted by each error by the following equation (5).
  • ⁇ is a predetermined weighting coefficient.
  • the optimization unit 27 updates the parameters of the encoding unit (12, 24) and the decoding unit (17, 19, 23) by, for example, the gradient method.
  • the optimization unit 27 determines, for example, whether or not the objective function L has converged, or whether or not the predetermined number of processes has been completed.
  • FIG. 1 the configuration and processing shown in FIG. 1 are examples, and are not limited to this. Further, the configuration of FIG. 1 includes a functional unit that is used and a functional unit that is not used, depending on the intended use. Further, the encoding units 17, 19 and 23 may be integrated or separate. The feature coupling portions 18 and 22 may be integrated or separate. The reconstruction error calculation units 18 and 20 may be integrated or separate.
  • the learning device 1 is configured by using a processor such as a CPU (Central Processing Unit) and a memory, for example.
  • the learning device 1 functions as a sampling unit 11, an encoding unit 2, a classification unit 3, and an optimization unit 27 by executing a program by the processor. All or part of each function of the learning device 1 may be realized by using hardware such as ASIC (Application Specific Integrated Circuit), PLD (Programmable Logic Device), and FPGA (Field Programmable Gate Array).
  • the above program may be recorded on a computer-readable recording medium.
  • Computer-readable recording media include, for example, flexible disks, magneto-optical disks, ROMs, CD-ROMs, portable media such as semiconductor storage devices (for example, SSD: Solid State Drive), hard disks and semiconductor storage built into computer systems. It is a storage device such as a device.
  • the above program may be transmitted over a telecommunication line.
  • FIG. 2 is a diagram showing an outline of the processing of the present embodiment.
  • the encoder g102 corresponds to the encoding unit 12 in FIG.
  • the encoder g102 and the decoder g105 are, for example, autoencoders.
  • Input data g101 is input to the encoder g102.
  • the learning device 1 performs learning by regarding the bottleneck portion of the autoencoder as a feature.
  • the label feature amount extraction unit 13 and the non-label feature amount extraction unit 14 separate the features into two, a label feature amount g103 and a non-label feature amount g104.
  • the label feature amount g103 and the non-label feature amount g104 are input to the decoder g105.
  • the decoder g105 corresponds to the decoding unit 19 in FIG.
  • the optimization unit 27 minimizes the classification error (CE loss; Cross-entropy loss) by using the label feature amount g103.
  • the optimization unit 27 uses the label feature amount g103 and the non-label feature amount g104 to minimize the reconstruction error.
  • FIG. 3 is a flowchart showing an example of processing procedures at the time of learning and at the time of classification according to the present embodiment.
  • the sampling unit 11 samples the input data of batch size B from the learning data (step S11).
  • the encoding unit 12 encodes the input data to obtain a feature amount (step S12).
  • the label feature amount extraction unit 13 extracts the label feature amount, and the non-label feature amount extraction unit 14 extracts the feature amount other than the label, thereby separating the feature amount into two (step S13).
  • the optimization unit 27 minimizes the classification error by using the label feature amount g103 (step S14).
  • the optimization unit 27 minimizes the reconstruction error by using the label feature amount g103 and the non-label feature amount g104 (step S15).
  • the optimization unit 27 updates the parameters of the encoding unit (12, 24) and the decoding unit (17, 19, 23) by, for example, the gradient method (step S16).
  • the optimization unit 27 determines, for example, whether or not the objective function L has converged, or whether or not the predetermined number of processes has been completed (step S16).
  • the optimization unit 27 ends the processing when the objective function L converges or when the processing for a predetermined number of times is completed (step S17; YES).
  • the optimization unit 27 repeats the processes of steps S11 to S16 when the objective function L has not converged or the predetermined number of processes have not been completed (step S17; NO).
  • FIGS. 4 to 6 show an example showing the effect of this embodiment.
  • the learning data and the data to be classified are examples of image data.
  • the label feature amount is a type of a number (0 to 9), and the feature amount other than the label is a number shape.
  • FIG. 4 is a diagram showing an example of a label feature amount and a feature amount other than the label according to the present embodiment.
  • the vertical axis is the label feature amount g201 and the non-label feature amount g202.
  • the horizontal direction is the original image g203 and the reconstructed image g204 when the features are changed. The image in the frame g205 will be described later.
  • FIG. 5 is a diagram showing an example of an original image and a reconstructed image according to the present embodiment. In the horizontal direction, the original images g211 and g213 and the reconstructed images g212 and g214 are shown.
  • FIG. 6 is a diagram showing an example of a reconstructed image when a non-label feature amount is exchanged with the original image according to the present embodiment.
  • the original images g221 and g223 are reconstructed images g212 and g214 when other than the label feature amount is exchanged.
  • the image in the frame g225 will be described later.
  • the features are separated into two, a label feature and a feature other than the label. Further, in the learning device 1, the classification error is minimized by using the label feature amount. Further, in the learning device 1, the reconstruction error is minimized by using the label feature amount and the feature amount other than the label.
  • the label information can be clearly extracted as an expression on a continuous space.
  • FIG. 7 is a diagram showing an outline of the processing of the present embodiment.
  • the encoder g107 corresponds to the encoding unit 24 of FIG.
  • the encoder g107 is, for example, an autoencoder.
  • the reconstructed data g106 is input to the encoder g107.
  • the encoder g102 and the encoder g107 may be integrated or separate.
  • the feature amount exchange unit 21 other than the label exchanges the feature amount other than the label between batches.
  • the decoding unit 23 decodes the feature amount obtained by combining the feature amount other than the label exchanged with the label feature amount.
  • the encoding unit 24 re-encodes the decoded features.
  • the optimization unit 27 minimizes the classification error by using the label feature amount g103' obtained as a result of re-encoding.
  • FIG. 8 is a flowchart showing an example of processing procedures at the time of learning and at the time of classification according to the present embodiment.
  • the learning device 1 performs the processes of steps S11 to S13. Subsequently, the feature amount exchange unit 21 other than the label exchanges the feature amount other than the label between batches (step S21).
  • the decoding unit 23 decodes the feature amount obtained by combining the feature amount other than the label exchanged with the label feature amount (step S22).
  • the encoding unit 24 re-encodes the decoded features (step S23).
  • the optimization unit 27 minimizes the classification error by using the re-encoded label feature amount g103'(step S24).
  • the learning device 1 performs the processes of steps S16 to S17.
  • FIGS. 9 to 11 an example showing the effect of this embodiment is shown in FIGS. 9 to 11.
  • the learning data and the data to be classified are examples of image data.
  • FIG. 9 is a diagram showing an example of a label feature amount and a non-label feature amount according to the present embodiment.
  • FIG. 10 is a diagram showing an example of an original image and a reconstructed image according to the present embodiment.
  • FIG. 11 is a diagram showing an example of a reconstructed image when a non-label feature amount is exchanged with the original image according to the present embodiment.
  • the numbers do not change to other numbers, that is, the label information is not included in the features other than the label.
  • the features are separated into two, a label feature and a feature other than the label. Further, in the learning device 1, features other than labels are exchanged between batches. Further, in the learning device 1, the exchanged data is decoded and the decoded reconstruction data is re-encoded. Further, in the learning device 1, the label feature amount g103' obtained by re-encoding was used to minimize the classification error.
  • label information is included in features other than labels, it may result in different label data when reconstructed.
  • FIG. 12 is a diagram showing an outline of the processing of the present embodiment.
  • the label feature amount exchange unit 15 randomly exchanges the label feature amount between the same labels in the batch.
  • the decoding unit 17 decodes the feature amount obtained by combining the exchanged label feature amount and the feature amount other than the label.
  • the optimization unit 27 uses the reconstruction data decoded by the decoding unit 17 to minimize the reconstruction error.
  • FIG. 13 is a flowchart showing an example of processing procedures at the time of learning and at the time of classification according to the third embodiment.
  • the learning device 1 performs the processes of steps S11 to S13.
  • the label feature amount exchange unit 15 randomly exchanges the label feature amount g103 between the same labels in the batch (step S31).
  • the decoding unit 17 decodes the feature amount obtained by combining the exchanged label feature amount g103 and the non-label feature amount g104 (step S32).
  • the optimization unit 27 uses the exchanged and decoded reconstruction data to minimize the reconstruction error (step S33).
  • the learning device 1 performs the processes of steps S16 to S17.
  • the features are separated into two, a label feature and a feature other than the label. Further, in the learning device 1, the label feature amount is exchanged between the same labels in the batch. Further, in the learning device 1, the exchanged data is decoded, and the decoded reconstruction data is used to minimize the reconstruction error.
  • the label feature amount is exchanged with other same label data and reconstructed.
  • this reconstruction since only the label information must be included in the exchanged label feature amount, it is possible to prevent the label feature amount from containing information other than the label.
  • common features between the exchanged samples can be extracted.
  • the training data without label information is divided into two features (first partial feature amount (label feature amount) and second partial feature amount (feature amount other than label)), and the label feature amount is determined.
  • the common feature is, for example, in the case of an image group of a dog, the information of a dog is a common feature, and in the case of an image group of handwritten characters of a certain person, the information of how to write the person is a common feature.
  • a natural image such as Image, which is a data set
  • the concept of a natural image is a common feature.
  • the present embodiment can be applied to the learning data to which the label is not attached.
  • the learning device 1 extracts, for example, the feature amount from the target data, reconstructs the extracted feature amount, acquires the reconstructed data, and reconstructs the difference between the target data and the reconstructed data.
  • the configuration error is output as the degree to which the target data has the characteristics that the predetermined data group has in common.
  • the learning device 1 separates the feature amount obtained from the data belonging to a predetermined data group into a first partial feature amount and a second partial feature amount, and the second part is described above.
  • the feature amount is exchanged with a second partial feature amount extracted from another data belonging to a predetermined data group, and the feature amount after exchange is acquired. Then, the learning device 1 optimizes so that the difference between the data in which the feature amount is reconstructed after the exchange and the data belonging to the predetermined data group becomes small.
  • FIGS. 14 to 16 show an example showing the effect when the processing of the second embodiment and the processing of the present embodiment are performed in addition to the first embodiment.
  • the learning data and the data to be classified are examples of image data.
  • FIG. 14 is a diagram showing an example of a label feature amount and a non-label feature amount when the processing of the second embodiment and the processing of the present embodiment are performed in addition to the first embodiment.
  • FIG. 15 is a diagram showing an example of an original image and a reconstructed image when the processing of the second embodiment and the processing of the present embodiment are performed in addition to the first embodiment.
  • FIG. 16 is a diagram showing an example of a reconstructed image when the processing of the second embodiment and the processing of the present embodiment are performed in addition to the first embodiment and the original image and the label feature amount are exchanged. be.
  • the label information is not included in the features other than the label. Further, when the processing of the present embodiment is performed in addition to the first embodiment, information other than the label features is not included in the label feature amount. Thereby, according to the second embodiment and the present embodiment, the label information and the non-label information can be clearly separated.
  • the data to be separated from the features is not limited to the image data, but may be other data. Further, the image data may be a still image or a moving image.
  • each of the above-described embodiments since the data can be separated into arbitrary features, it is possible to generate data having specific features or edit and reconstruct the specific features. Thereby, each of the above-described embodiments can generate or edit data for any feature (data disentanglement).
  • the label information and other information can be separated, and the label information can be extracted as a value in a continuous space, so that it can be applied to recognition of unlearned classes and the like.
  • each of the above-described embodiments can improve the accuracy of Few-shot learning for recognizing a class of a small number of data.
  • the present invention can be applied to data feature separation, data generation, data editing, data class recognition, transfer learning, and the like.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Evolutionary Computation (AREA)
  • Computing Systems (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Software Systems (AREA)
  • General Health & Medical Sciences (AREA)
  • Multimedia (AREA)
  • Biophysics (AREA)
  • Molecular Biology (AREA)
  • General Engineering & Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Mathematical Physics (AREA)
  • Computational Linguistics (AREA)
  • Biomedical Technology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Quality & Reliability (AREA)
  • Image Analysis (AREA)

Abstract

本発明の一態様は、学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類する分類部と、潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成するデコード部と、ラベル特徴量を用いて、ラベル特徴量と前記ラベル情報との分類誤差を最小化するようにデコードパラメータを最適化する最適化部と、を備える学習装置である。

Description

学習装置、学習方法およびプログラム
 本発明は、学習装置、学習方法およびプログラムの技術に関する。
 ラベル特徴を抽出するWcとラベル以外特徴を抽出するWuの2つのニューラルネットワークで構成され、ラベル特徴を更にクラス分類用のニューラルネットワークへ入力し、クラス分類タスクを解く学習方法が提案されている。そして、この提案の学習方法では、ラベル特徴の再構成とラベル以外特徴の再構成を1:1で加重和したもので入力xを復元する(例えば非特許文献1参照)。
Thomas Robert, Nicolas Thome, Matthieu Cord、"HybridNet: Classification and Reconstruction Cooperation for Semi-Supervised Learning"、2018、インターネット検索、<URL: https://arxiv.org/abs/1807.11407>
 しかしながら、従来技術は、ラベル特徴のクラス分類を解く際に、ラベル特徴の特徴を更にクラス分類用のNWへ入力しているため、この処理でクラス以外の情報が消失する可能性がある。このため、従来技術では、ラベル特徴がクラス以外の情報を含んでいたとしてもそれを検知できない。このように、従来技術では、学習時に特徴が漏れるため、データを任意の特徴に明確に分離することができない場合があるという問題があった。
 上記事情に鑑み、本発明は、データを任意の特徴に明確に分離することができる技術の提供を目的としている。
 本発明の一態様は、学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類する分類部と、前記潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成するデコード部と、前記ラベル特徴量を用いて、前記ラベル特徴量とラベル以外特徴量との分類誤差を最小化するように前記デコードパラメータを最適化する最適化部と、を備える学習装置である。
 本発明の一態様は、分類部が、学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類し、デコード部が、前記潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成し、最適化部が、前記ラベル特徴量を用いて、前記ラベル特徴量とラベル以外特徴量との分類誤差を最小化するように前記デコードパラメータを最適化する、学習方法である。
 本発明の一態様は、コンピュータが、対象データから特徴量を抽出するステップと、抽出された前記特徴量を再構成し再構成データを取得する再構成ステップと、前記対象データと前記再構成データとの差である再構成誤差を、所定のデータ群が共通して有する特徴を前記対象データが有する度合いとして出力するステップを有し、前記再構成ステップは、前記所定のデータ群に属するデータから得られた特徴量を、第一の部分特徴量と、第二の部分特徴量と、に分離し、前記第二の部分特徴量を、前記所定のデータ群に属する別のデータから抽出された第二の部分特徴量と交換し、交換後特徴量を取得し、前記交換後特徴量を再構成したデータと、前記所定のデータ群に属するデータとの差が小さくなるよう最適化する、学習方法である。
 本発明の一態様は、コンピュータに、学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類させ、前記潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成させ、前記ラベル特徴量を用いて、前記ラベル特徴量とラベル以外特徴量との分類誤差を最小化するように前記デコードパラメータを最適化させる、プログラムである。
 本発明により、データを任意の特徴に明確に分離することができる。
実施形態の学習装置の構成の一例を示す図である。 第1の実施形態の処理の概要を示す図である。 第1の実施形態に係る学習時と分類時の処理手順例を示すフローチャートである。 第1の実施形態に係るラベル特徴量とラベル以外特徴量の一例を示す図である。 第1の実施形態に係る原画と再構成した画像の一例を示す図である。 第1の実施形態に係る原画とラベル特徴量以外を交換した時の再構成した画像の一例を示す図である。 第2の実施形態の処理の概要を示す図である。 第2の実施形態に係る学習時と分類時の処理手順例を示すフローチャートである。 第2の実施形態に係るラベル特徴量とラベル以外特徴量の一例を示す図である。 第2の実施形態に係る原画と再構成した画像の一例を示す図である。 第2の実施形態に係る原画とラベル特徴量以外を交換した時の再構成した画像の一例を示す図である。 第3の実施形態の処理の概要を示す図である。 第3の実施形態に係る学習時と分類時の処理手順例を示すフローチャートである。 第1の実施形態に加えて第2の実施形態の処理と第3の実施形態の処理を行う場合のラベル特徴量とラベル以外特徴量の一例を示す図である。 第1の実施形態に加えて第2の実施形態の処理と第3の実施形態の処理を行う場合の原画と再構成した画像の一例を示す図である。 第1の実施形態に加えて第2の実施形態の処理と第3の実施形態の処理を行う場合の原画とラベル特徴量以外を交換した時の再構成した画像の一例を示す図である。
 本発明の実施形態について、図面を参照して詳細に説明する。
 図1は、実施形態の学習装置の構成の一例を示す図である。図1のように、学習装置1は、サンプリング部11、分類部2、処理部3、および最適化部27を備える。
 分類部2は、エンコード部12、ラベル特徴量抽出部13、およびラベル以外特徴量抽出部14を備える。
 処理部3は、ラベル特徴量交換部15、特徴結合部16、デコード部17、再構成誤差算出部18、デコード部19、再構成誤差算出部20、ラベル以外特徴量交換部21、特徴結合部22、デコード部23、エンコード部24、ラベル特徴量抽出部25、および分類誤差算出部26を備える。
 学習装置1は、入力されたデータをラベル特徴量とラベル以外特徴量とに分離する。なお、以下の説明において、学習データを{xi,yi}(xiは入力データ、yiはラベル(クラス)情報)(i=1,…,N)とする。
 サンプリング部11は、学習データ{xi,yi}からバッチサイズB(Bは1以上の整数)の入力データ{x1,y1},…,{xB,yB}をサンプリングする。
 エンコード部12は、サンプルされた入力データxiをエンコードして、各データについてM個のパラメータから構成される特徴量101{Zi=[zi,label,zi,wo_label]}を得る。ここで、zi,labelはC個(Cは1以上の整数)のパラメータから構成されるラベル特徴量zi,label=[zi,1,…,zi,C]であり、zi,wo_labelはM-C個(Mは2以上の整数)のパラメータから構成されるラベル以外特徴量zi,wo_label=[zi,C+1,…,zi,M]である。エンコード部12は、特徴量101を、ラベル特徴量抽出部13とラベル以外特徴量抽出部14とデコード部19とに出力する。なお、潜在変数とは、オートエンコーダを使用する場合、エンコードして得られる特徴量である。
 ラベル特徴量抽出部13は、ラベル特徴量102{zi,label}を抽出する。ラベル特徴量抽出部13は、抽出したラベル特徴量102を、ラベル特徴量交換部15と特徴結合部22と分類誤差算出部26とに出力する。
 ラベル以外特徴量抽出部14は、ラベル以外特徴量103{zi,wo_label}を抽出する。ラベル以外特徴量抽出部14は、抽出したラベル以外特徴量103を、特徴結合部16とラベル以外特徴量交換部21と特徴結合部22とに出力する。
 ラベル特徴量交換部15には、学習データに付与されているラベル情報と、ラベル特徴量102とが入力される。ラベル特徴量交換部15は、ラベル特徴量zi,labelの各パラメータについてバッチ処理内の同一ラベルサンプルとランダムに交換(スワップ)する。交換したものを(zi,label)swapとする。ラベル特徴量交換部15は、交換したラベル特徴量104を特徴結合部16に出力する。なお、ラベル特徴量交換部15には、バッチ処理内に限らず、同一ラベルの別のサンプルと交換するようにしてもよい。
 特徴結合部16は、ラベル特徴量交換部15によって交換されたラベル特徴量104と、ラベル以外特徴量抽出部14によって抽出されたラベル以外特徴量103とを結合し、結合した特徴量をデコード部17に出力する。
 デコード部17は、特徴量をデコードして再構成データ105{(xi)(swap_label)^}を得る。デコード部17は、再構成データ105を再構成誤差算出部18に出力する。
 再構成誤差算出部18は、入力データxiと、デコードして得られた再構成データ(xi)^との再構成誤差106{Lrec,swap}を次式(1)によって算出する。なお、式(1)においてdは、2つのベクトル間の距離を算出する任意の関数であり、例えば平均二乗誤差和や平均絶対誤差和等である。再構成誤差算出部18は、算出した再構成誤差106を最適化部27に出力する。
Figure JPOXMLDOC01-appb-M000002
 デコード部19は、特徴量101をデコードして再構成データ107{(xi)^}を得る。デコード部19は、再構成データ107を再構成誤差算出部20に出力する。
 再構成誤差算出部20は、入力データxiと、デコード部19が出力する再構成データ(zi)(swap_label)^との再構成誤差108{Lrec,org}を次式(2)によって算出する。
Figure JPOXMLDOC01-appb-M000003
 ラベル以外特徴量交換部21は、ラベル以外特徴量zi,wo_labelの各パラメータについてバッチ処理内のサンプルとランダムに交換する。交換したものを(zi,wo_label)swapとする。ラベル以外特徴量交換部21は、ラベル特徴量zi,labelと交換された(zi,wo_label)swapとを結合した特徴量{(zi)swap_wo_label}を生成する。ラベル以外特徴量交換部21は、交換したラベル以外特徴量110を特徴結合部22に出力する。
 特徴結合部22は、ラベル特徴量抽出部13によって抽出されたラベル特徴量102と、ラベル以外特徴量交換部21によって交換されたラベル以外特徴量110とを結合する。特徴結合部22は、結合した特徴量をデコード部23に出力する。
 デコード部23は、合された特徴量{(zi)swap_wo_label}をデコードして再構成データ111{(xi)(swap_wo_label)^}を得る。デコード部23は、再構成データ111をエンコード部24に出力する。
 エンコード部24は、再構成データ111{(xi)(swap_wo_label)^}を再エンコードして、特徴量112を得る。エンコード部24は、特徴量112をラベル特徴量抽出部25に出力する。
 ラベル特徴量抽出部25は、特徴量112からラベル特徴量{(zi,label)(swap_wo_label)^}を抽出し、抽出したラベル特徴量113を分類誤差算出部26に出力する。
 分類誤差算出部26には、ラベル情報と、ラベル特徴量抽出部13が抽出したラベル特徴量102と、ラベル特徴量抽出部25が抽出したラベル特徴量113とが入力される。分類誤差算出部26は、ラベル特徴量102{zi,label}から、次式(3)によって分類誤差109{Llabel,org}を算出する。式(3)において、(zyi,label) ̄は、バッチサンプルの中でラベル情報がyiであるサンプルのラベル特徴量zi,labelを平均化したものであり、Kは分類ラベル数である。
Figure JPOXMLDOC01-appb-M000004
 また、分類誤差算出部26は、ラベル特徴量113{(zi,label)(swap_wo_label)^}から、次式(4)によって分類誤差114{Llabel,swap}を算出する。
Figure JPOXMLDOC01-appb-M000005
 最適化部27は、各誤差を重み付けした目的関数Lを次式(5)によって算出する。なお、式(5)において、λは所定の重み係数である。
Figure JPOXMLDOC01-appb-M000006
 さらに、最適化部27は、例えば勾配法によりエンコード部(12,24)およびデコード部(17,19,23)のパラメータを更新する。最適化部27は、例えば目的関数Lが収束したか否かを判別、または所定回数の処理が終了したか否かを判別する。
 なお、図1に示した構成や処理は一例であり、これに限らない。また、図1の構成は、用途によって、使用する機能部と使用しない機能部とがある。また、エンコード部17、19、23は、一体であっても別であってもよい。特徴結合部18、22は、一体であっても別であってもよい。再構成誤差算出部18,20は、一体であっても別であってもよい。
 なお、学習装置1は、例えばCPU(Central Processing Unit)等のプロセッサーとメモリーとを用いて構成される。学習装置1は、プロセッサーがプログラムを実行することによって、サンプリング部11、エンコード部2、分類部3および最適化部27として機能する。なお、学習装置1の各機能の全て又は一部は、ASIC(Application Specific Integrated Circuit)やPLD(Programmable Logic Device)やFPGA(Field Programmable Gate Array)等のハードウェアを用いて実現されても良い。上記のプログラムは、コンピュータ読み取り可能な記録媒体に記録されても良い。コンピュータ読み取り可能な記録媒体とは、例えばフレキシブルディスク、光磁気ディスク、ROM、CD-ROM、半導体記憶装置(例えばSSD:Solid State Drive)等の可搬媒体、コンピューターシステムに内蔵されるハードディスクや半導体記憶装置等の記憶装置である。上記のプログラムは、電気通信回線を介して送信されてもよい。
 (第1の実施例)
 本実施形態では、エンコード部12が同一層で特徴を分離する。なお、本実施形態では、バッチ内で交換させない。
 図2は、本実施形態の処理の概要を示す図である。エンコーダg102は、図1のエンコード部12に対応する。エンコーダg102とデコーダg105は、例えばオートエンコーダである。エンコーダg102には、入力データg101が入力される。
 学習装置1は、オートエンコーダのボトルネック部分を特徴とみなして学習を行う。
 ラベル特徴量抽出部13とラベル以外特徴量抽出部14は、特徴をラベル特徴量g103とラベル以外特徴量g104との2つに分離する。
 ラベル特徴量g103とラベル以外特徴量g104とは、デコーダg105に入力される。デコーダg105は、図1のデコード部19に対応する。
 最適化部27は、ラベル特徴量g103を用いて、クラス分類誤差(CE loss;Cross-entropy loss)を最小化する。
 最適化部27は、ラベル特徴量g103とラベル以外特徴量g104とを用いて、再構成誤差を最小化する。
 次に、学習時と分類時の処理手順例を説明する。
 図3は、本実施形態に係る学習時と分類時の処理手順例を示すフローチャートである。
 サンプリング部11は、学習データからバッチサイズBの入力データをサンプルする(ステップS11)。エンコード部12は、入力データをエンコードして特徴量を得る(ステップS12)。
 ラベル特徴量抽出部13がラベル特徴量を抽出し、ラベル以外特徴量抽出部14がラベル以外特徴量を抽出することで、特徴量を2つに分離する(ステップS13)。
 最適化部27は、ラベル特徴量g103を用いて、クラス分類誤差を最小化する(ステップS14)。最適化部27は、ラベル特徴量g103とラベル以外特徴量g104とを用いて、再構成誤差を最小化する(ステップS15)。
 最適化部27は、例えば勾配法によりエンコード部(12,24)およびデコード部(17,19,23)のパラメータを更新する(ステップS16)。最適化部27は、例えば目的関数Lが収束したか否かを判別、または所定回数の処理が終了したか否かを判別する(ステップS16)。最適化部27は、目的関数Lが収束した場合または所定回数の処理が終了した場合(ステップS17;YES)、処理を終了する。最適化部27は、目的関数Lが収束していない場合または所定回数の処理が終了していない場合(ステップS17;NO)、ステップS11~S16の処理を繰り返す。
 次に、本実施形態の効果を示す一例を図4~6に示す。なお、図4~図6では、学習データ、分類すべきデータが画像データの例である。また、ラベル特徴量は数字の種類(0~9)であり、ラベル以外特徴量は数字の形状である。
 図4は、本実施形態に係るラベル特徴量とラベル以外特徴量の一例を示す図である。縦軸は、ラベル特徴量g201と、ラベル以外特徴量g202である。横方向は、原画g203と、特徴をそれぞれ変化させた時の再構成した画像g204である。なお、枠g205内の画像については、後述する。
 図5は、本実施形態に係る原画と再構成した画像の一例を示す図である。横方向は、原画g211、g213と、再構成した画像g212、g214である。
 図6は、本実施形態に係る原画とラベル特徴量以外を交換した時の再構成した画像の一例を示す図である。横方向は、原画g221、g223と、ラベル特徴量以外を交換した時の再構成した画像g212、g214である。なお、枠g225内の画像については、後述する。
 本実施形態では、このように構成された学習装置1では、特徴をラベル特徴とラベル以外の特徴との2つに分離するようにした。また、学習装置1では、ラベル特徴量を用いてクラス分類誤差を最小化するようにした。また、学習装置1では、ラベル特徴量とラベル以外特徴量とを用いて再構成誤差を最小化するようにした。
 これにより、本実施形態によれば、オートエンコーダにより再構成するため、特徴の漏れがない。また、本実施形態によれば、ラベル情報が連続空間上の表現として明確に抽出することができる。
 (第2の実施例)
 ラベル以外の特徴から、さらに精度よくラベル特徴を除外するテクニックを本実施形態で示す。ラベル以外の特徴にラベル特徴が含まれると、デコードした結果得られる出力値が、違うラベルの出力値になると考えられる。また、同じラベルをもつデータ間であれば、ラベル以外の特徴を交換しても同じクラスの出力値にデコードされる。そこで、本実施形態では、学習装置1が、ラベル以外の特徴をバッチ内で交換させて学習する。
 図7は、本実施形態の処理の概要を示す図である。エンコーダg107は、図1のエンコード部24に対応する。エンコーダg107は、例えばオートエンコーダである。エンコーダg107には、再構成されたデータg106が入力される。なお、エンコーダg102とエンコーダg107は、一体であっても別であってもよい。
 第2の実施形態では、第1の実施形態に加えて、以下の処理を行う。
 ラベル以外特徴量交換部21は、ラベル以外特徴量をバッチ間で交換する。
 デコード部23は、ラベル特徴量と交換されたラベル以外特徴量を結合した特徴量をデコードする。
 エンコード部24は、デコードされた特徴量を再エンコードする。
 最適化部27は、再エンコードした結果得られたラベル特徴量g103’を用いて、クラス分類誤差を最小化する。
 次に、学習時と分類時の処理手順例を説明する。
 図8は、本実施形態に係る学習時と分類時の処理手順例を示すフローチャートである。
 学習装置1は、ステップS11~S13の処理を行う。
 続けて、ラベル以外特徴量交換部21は、ラベル以外特徴量をバッチ間で交換する(ステップS21)。デコード部23は、ラベル特徴量と交換されたラベル以外特徴量を結合した特徴量をデコードする(ステップS22)。エンコード部24は、デコードされた特徴量を再エンコードする(ステップS23)。
 続けて、最適化部27は、再エンコードされたラベル特徴量g103’を用いてクラス分類誤差を最小化する(ステップS24)。
 続けて、学習装置1は、ステップS16~S17の処理を行う。
 次に、本実施形態の効果を示す一例を図9~11に示す。なお、図9~図11では、学習データ、分類すべきデータが画像データの例である。
 図9は、本実施形態に係るラベル特徴量とラベル以外特徴量の一例を示す図である。図10は、本実施形態に係る原画と再構成した画像の一例を示す図である。図11は、本実施形態に係る原画とラベル特徴量以外を交換した時の再構成した画像の一例を示す図である。
 図11のようにラベル以外特徴量を交換して再構成しても、他の数字に変化しない、すなわちラベル以外特徴量にラベル情報が入っていない。
 本実施形態では、このように構成された学習装置1では、特徴をラベル特徴とラベル以外の特徴との2つに分離するようにした。また、学習装置1では、ラベル以外特徴量をバッチ間で交換するようにした。また、学習装置1では、交換されたデータをデコードし、デコードされた再構成データを再エンコードするようにした。また、学習装置1では、再エンコードされて得られたラベル特徴量g103’を用いてクラス分類誤差を最小化するようにした。
 ラベル以外の特徴にラベルの情報が入っていると再構成した時に異なるラベルのデータになる場合がある。これに対して、本実施形態によれば、再構成された画像を再エンコードしてクラス分類誤差が小さくなるようにすることでラベル以外の特徴にラベルの情報が含まれなくすることができる。
 (第3の実施例)
 ラベル特徴量から、さらにラベル特徴量以外の情報を取り除くテクニックを本実施形態で説明する。同一のラベルが付与されるデータ間であれば、ラベル特徴を交換してもデコードされた結果得られるクラスは同一である。そこで、本実施形態では、学習装置1が、ラベル特徴をバッチ内の同一ラベル間で交換させて学習する。図12は、本実施形態の処理の概要を示す図である。
 第3の実施形態では、第1の実施形態に加えて、以下の処理を行う。
 ラベル特徴量交換部15は、ラベル特徴量をバッチ内の同一ラベル間でランダムに交換する。
 デコード部17は、交換されたラベル特徴量とラベル以外特徴量を結合した特徴量をデコードする。
 最適化部27は、デコード部17でデコードされた再構成データを用いて、再構成誤差を最小化する。
 次に、第1の実施形態に加えて本実施形態の処理を行う場合の学習時と分類時の第1の処理手順例を説明する。図13は、第3の実施形態に係る学習時と分類時の処理手順例を示すフローチャートである。
 学習装置1は、ステップS11~S13の処理を行う。
 ラベル特徴量交換部15は、ラベル特徴量g103をバッチ内の同一ラベル間でランダムに交換する(ステップS31)。デコード部17は、交換されたラベル特徴量g103とラベル以外特徴量g104を結合した特徴量をデコードする(ステップS32)。
 最適化部27は、交換されデコードされた再構成データを用いて、再構成誤差を最小化する(ステップS33)。
 学習装置1は、ステップS16~S17の処理を行う。
 本実施形態では、このように構成された学習装置1では、特徴をラベル特徴とラベル以外の特徴との2つに分離するようにした。また、学習装置1では、ラベル特徴量をバッチ内の同一ラベル間で交換するようにした。また、学習装置1では、交換されたデータをデコードし、デコードされた再構成データを用いて再構成誤差を最小化するようにした。
 以上のように、本実施形態によれば、ラベル特徴量を他の同一ラベルデータと交換して再構成するようにした。この再構成では、交換したラベル特徴量にラベル情報のみが含まれていなければならないため、ラベル特徴量にラベル以外の情報が含まれなくすることができる。
 なお、本実施形態によれば、交換したサンプル間の共通特徴を抽出できる。本実施形態では、ラベル情報がない学習データを2つの特徴(第一の部分特徴量(ラベル特徴量)と、第二の部分特徴量(ラベル以外特徴量))に分けて、ラベル特徴量をランダムに交換して再構成誤差を算出することで、その学習データの潜在的な共通特徴を求めることができる。なお、共通特徴とは、例えば、犬の画像群であれば、犬という情報が共通特徴であり、ある人の手書き文字の画像群であれば、その人の書き方の情報が共通特徴であり、あるいはデータセットであるImagenetのような自然画像を学習データであれば、自然画像という概念が共通特徴である。これにより、本実施形態は、ラベルが付与されていない学習データにも適用ができる。
 この場合の処理は、学習装置1が、例えば、対象データから特徴量を抽出し、抽出された特徴量を再構成し再構成データを取得し、対象データと再構成データとの差である再構成誤差を、所定のデータ群が共通して有する特徴を前記対象データが有する度合いとして出力する。学習装置1は、再構成の際、所定のデータ群に属するデータから得られた特徴量を、第一の部分特徴量と、第二の部分特徴量と、に分離し、前記第二の部分特徴量を、所定のデータ群に属する別のデータから抽出された第二の部分特徴量と交換し、交換後特徴量を取得する。そして、学習装置1は、交換後特徴量を再構成したデータと、所定のデータ群に属するデータとの差が小さくなるよう最適化する。
 次に、第1の実施形態に加えて第2の実施形態の処理と本実施形態の処理を行う場合の効果を示す一例を図14~16に示す。なお、図14~図16では、学習データ、分類すべきデータが画像データの例である。
 図14は、第1の実施形態に加えて第2の実施形態の処理と本実施形態の処理を行う場合のラベル特徴量とラベル以外特徴量の一例を示す図である。図15は、第1の実施形態に加えて第2の実施形態の処理と本実施形態の処理を行う場合の原画と再構成した画像の一例を示す図である。図16は、第1の実施形態に加えて第2の実施形態の処理と本実施形態の処理を行う場合の原画とラベル特徴量以外を交換した時の再構成した画像の一例を示す図である。
 図14~図16のように、第1の実施形態に加えて第2の実施形態の処理を行う場合は、ラベル以外の特徴にラベルの情報が乗っていない。また、第1の実施形態に加えて本実施形態の処理を行う場合は、ラベル特徴量にラベル特徴以外の情報が乗っていない。これにより、第2実施形態と本実施形態とによれば、ラベル情報とラベル以外情報を明確に分離することができる。
 (変形例)
 なお、上述した各実施例において、特徴を分離する対象のデータが画像データに限らず、他のデータであってもよい。また、画像データは、静止画であっても動画であってもよい。
 また、上述した各実施形態によれば、データを任意の特徴に分離できるため、特定の特徴を持ったデータを生成したり、特定の特徴を編集して再構成したりすることができる。これにより、上述した各実施形態は、任意の特徴についてデータ生成したり編集することができる(データのDisentanglement)。
 また、上述した各実施形態によれば、ラベル情報とそれ以外の情報に分離し、更にラベル情報を連続空間での値として抽出できるため、未学習クラスの認識等へ応用が可能である。これにより、上述した各実施形態は、少数データのクラスを認識するFew-shot学習の精度を向上させることができる。
 通常の転移学習では、例えばImagenetのクラス分類問題で学習する等、クラス分類タスクに特化した特徴を再利用する。しかし、別のタスクで必要な情報が失われてしまう可能性がある。これに対して、上述した各実施形態によれば、データを再現するのに過不足なく特徴を得ているため、様々なタスクへ転移学習しても必要な情報が失われないため、精度を向上させることができる。これにより、上述した各実施形態は、転移学習の精度を向上させることができる。
 以上、この発明の実施形態について図面を参照して詳述してきたが、具体的な構成はこの実施形態に限られるものではなく、この発明の要旨を逸脱しない範囲の設計等も含まれる。
 本発明は、データの特徴の分離、データの生成、データの編集、データのクラスの認識、転移学習等に適用可能である。
1…学習装置、2…分類部、3…処理部、11…サンプリング部、12…エンコード部、13…ラベル特徴量抽出部、14…ラベル以外特徴量抽出部、15…ラベル特徴量交換部、16…特徴結合部、17…デコード部、18…再構成誤差算出部、19…デコード部、20…再構成誤差算出部、21…ラベル以外特徴量交換部、22…特徴結合部、23…デコード部、24…エンコード部、25…ラベル特徴量抽出部、26…分類誤差算出部、27…最適化部

Claims (7)

  1.  学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類する分類部と、
     前記潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成するデコード部と、
     前記ラベル特徴量を用いて、前記ラベル特徴量と前記ラベル情報との分類誤差を最小化するように前記デコードパラメータを最適化する最適化部と、
     を備える学習装置。
  2.  前記ラベル特徴量は、C(Cは1以上の整数)個のパラメータから構成され、
     前記ラベル特徴量の各パラメータを、バッチ処理内の同一ラベルの前記学習データとランダムに交換するラベル特徴量交換部と、
     前記交換されたラベル特徴量と、ラベル以外特徴量とを結合する特徴結合部と、
     前記潜在変数と、前記結合された特徴量を前記デコード部によってデコードして生成された再構成データ再構成データとの再構成誤差を算出する再構成誤差算出部と、を更に備え、
     請求項1に記載の学習装置。
  3.  前記分類部は、オートエンコーダを備える、
     請求項1または請求項2に記載の学習装置。
  4.  前記再構成誤差は、次式においてLrec,swapであり、
    Figure JPOXMLDOC01-appb-M000001
     前記xiは前記潜在変数であり、前記(xi)(swap_wo_label)^は前記再構成データであり、B(Bは1以上の整数)はバッチサイズであり、前記dは2つのベクトル間の距離を算出する任意の関数である、
     請求項2に記載の学習装置。
  5.  分類部が、学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類し、
     デコード部が、前記潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成し、
     最適化部が、前記ラベル特徴量を用いて、前記ラベル特徴量と前記ラベル情報との分類誤差を最小化するように前記デコードパラメータを最適化する、
     学習方法。
  6.  コンピュータが、
     対象データから特徴量を抽出するステップと、
     抽出された特徴量を再構成し再構成データを取得する再構成ステップと、
     前記対象データと前記再構成データとの差である再構成誤差を、所定のデータ群が共通して有する特徴を前記対象データが有する度合いとして出力するステップと、を有し、
     前記再構成ステップは、
     前記所定のデータ群に属するデータから得られた特徴量を、第一の部分特徴量と、第二の部分特徴量と、に分離し、
     前記第二の部分特徴量を、前記所定のデータ群に属する別のデータから抽出された第二の部分特徴量と交換し、交換後特徴量を取得し、前記交換後特徴量を再構成したデータと、前記所定のデータ群に属するデータとの差が小さくなるよう最適化する、
     学習方法。
  7.  コンピュータに、
     学習に用いる学習データから得られた特徴量である潜在変数を、分類に用いられるラベル情報を有するラベル特徴量を用いて分類させ、
     前記潜在変数をデコードして所定のデコードパラメータを用いて再構成データを生成させ、
     前記ラベル特徴量を用いて、前記ラベル特徴量と前記ラベル情報との分類誤差を最小化するように前記デコードパラメータを最適化させる、
     プログラム。
PCT/JP2020/041850 2020-11-10 2020-11-10 学習装置、学習方法およびプログラム Ceased WO2022101962A1 (ja)

Priority Applications (3)

Application Number Priority Date Filing Date Title
PCT/JP2020/041850 WO2022101962A1 (ja) 2020-11-10 2020-11-10 学習装置、学習方法およびプログラム
JP2022561708A JP7513918B2 (ja) 2020-11-10 2020-11-10 学習装置、学習方法およびプログラム
US18/035,540 US20230410472A1 (en) 2020-11-10 2020-11-10 Learning device, learning method and program

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2020/041850 WO2022101962A1 (ja) 2020-11-10 2020-11-10 学習装置、学習方法およびプログラム

Publications (1)

Publication Number Publication Date
WO2022101962A1 true WO2022101962A1 (ja) 2022-05-19

Family

ID=81600893

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2020/041850 Ceased WO2022101962A1 (ja) 2020-11-10 2020-11-10 学習装置、学習方法およびプログラム

Country Status (3)

Country Link
US (1) US20230410472A1 (ja)
JP (1) JP7513918B2 (ja)
WO (1) WO2022101962A1 (ja)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019061512A (ja) * 2017-09-27 2019-04-18 株式会社Abeja データの特徴を利用してデータを処理するシステム
JP2019200551A (ja) * 2018-05-15 2019-11-21 株式会社日立製作所 データから潜在因子を発見するニューラルネットワーク
JP2020160743A (ja) * 2019-03-26 2020-10-01 日本電信電話株式会社 評価装置、評価方法、および、評価プログラム

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019061512A (ja) * 2017-09-27 2019-04-18 株式会社Abeja データの特徴を利用してデータを処理するシステム
JP2019200551A (ja) * 2018-05-15 2019-11-21 株式会社日立製作所 データから潜在因子を発見するニューラルネットワーク
JP2020160743A (ja) * 2019-03-26 2020-10-01 日本電信電話株式会社 評価装置、評価方法、および、評価プログラム

Also Published As

Publication number Publication date
JP7513918B2 (ja) 2024-07-10
JPWO2022101962A1 (ja) 2022-05-19
US20230410472A1 (en) 2023-12-21

Similar Documents

Publication Publication Date Title
Ahmed et al. Classification and reconstruction of optical quantum states with deep neural networks
CN112734634B (zh) 换脸方法、装置、电子设备和存储介质
Gu et al. Lofgan: Fusing local representations for few-shot image generation
CN114612289B (zh) 风格化图像生成方法、装置及图像处理设备
Tirupattur et al. Thoughtviz: Visualizing human thoughts using generative adversarial network
Perarnau et al. Invertible conditional gans for image editing
Qi et al. Avt: Unsupervised learning of transformation equivariant representations by autoencoding variational transformations
CN118053090A (zh) 使用潜在扩散模型生成视频
Creswell et al. Adversarial information factorization
US20200074273A1 (en) Method for training deep neural network (dnn) using auxiliary regression targets
KR20200052453A (ko) 딥러닝 모델 학습 장치 및 방법
KR102332114B1 (ko) 이미지 처리 방법 및 장치
CN116612416A (zh) 一种指代视频目标分割方法、装置、设备及可读存储介质
JP7436928B2 (ja) 学習装置、学習方法およびプログラム
Jorge et al. Empirical Evaluation of Variational Autoencoders for Data Augmentation.
CN120597947A (zh) 相关于数据生成框架的方法及装置
Tsoumplekas et al. A complete survey on contemporary methods, emerging paradigms and hybrid approaches for few-shot learning
Wang et al. Denoising reuse: Exploiting inter-frame motion consistency for efficient video generation
WO2024039572A1 (en) Artificial intelligence computing systems for efficiently learning underlying features of data
JP7513918B2 (ja) 学習装置、学習方法およびプログラム
JP2019082847A (ja) データ推定装置、データ推定方法及びプログラム
WO2025049487A1 (en) Systems and methods for enhancing ultrasound images using artificial intelligence
JP7376812B2 (ja) データ生成方法、データ生成装置及びプログラム
Nguyen et al. Qc-stylegan-quality controllable image generation and manipulation
Fakhari et al. An image restoration architecture using abstract features and generative models

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20961490

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2022561708

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 18035540

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20961490

Country of ref document: EP

Kind code of ref document: A1