WO2022009331A1 - 学習装置、学習方法、およびプログラム - Google Patents

学習装置、学習方法、およびプログラム Download PDF

Info

Publication number
WO2022009331A1
WO2022009331A1 PCT/JP2020/026683 JP2020026683W WO2022009331A1 WO 2022009331 A1 WO2022009331 A1 WO 2022009331A1 JP 2020026683 W JP2020026683 W JP 2020026683W WO 2022009331 A1 WO2022009331 A1 WO 2022009331A1
Authority
WO
WIPO (PCT)
Prior art keywords
motion
encoder
skeleton
time
learning
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2020/026683
Other languages
English (en)
French (fr)
Inventor
芳陸 謝
豪 入江
達史 松林
浩太 日高
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to JP2022534553A priority Critical patent/JP7410441B2/ja
Priority to PCT/JP2020/026683 priority patent/WO2022009331A1/ja
Publication of WO2022009331A1 publication Critical patent/WO2022009331A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/04—Architecture, e.g. interconnection topology
    • G06N3/0475—Generative networks
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/04—Architecture, e.g. interconnection topology
    • G06N3/045—Combinations of networks
    • G06N3/0455—Auto-encoder networks; Encoder-decoder networks
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/08—Learning methods
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/08—Learning methods
    • G06N3/09—Supervised learning
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/08—Learning methods
    • G06N3/094—Adversarial learning
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T13/00—Animation
    • G06T13/20—Three-dimensional [3D] animation
    • G06T13/40—Three-dimensional [3D] animation of characters, e.g. humans, animals or virtual beings
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00—Image analysis
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00—Image analysis
    • G06T7/20—Analysis of motion

Definitions

  • the present invention relates to a learning device, a learning method, and a program.
  • Non-Patent Document 1 discloses a technique for separating dynamic elements (motion) and static elements (skeleton and viewpoint) by using deep learning, and reconstructing the dynamic elements and static elements. ..
  • the conventional technique is considered to reflect individuality such as body shape, but it is not considered to reflect individuality included in movement. There are ways to reflect the motion of a person who imitates the movement of a celebrity, but many are exaggerated expressions of some habits, and the scenes that can be used are limited.
  • the present invention has been made in view of the above, and an object of the present invention is to generate a movement that reflects the unique movement of the target.
  • the learning device of one aspect of the present invention has an input unit for inputting time-series data recording the movement of a moving body, an encoder that extracts the characteristics of the moving body independent of time using the time-series data, and time-varying.
  • a learning unit that learns a model including two or more encoders that extract motion characteristics, a classifier that identifies the moving object from the motion characteristics, and a decoder that outputs time-series data that records new motions of the moving object.
  • the learning unit extracts motion features common to the encoder that extracts the unique motion features of the moving object by hostile learning for two or more encoders that extract the motion features. To learn.
  • the learning method of one aspect of the present invention is a learning method executed by a computer, in which time-series data recording the movement of a moving object is input and the time-series data is used to display the characteristics of the moving object independent of time.
  • learning a model to include and learning the model for two or more encoders that extract the characteristics of the movement, the movement common to the encoder that extracts the characteristics of the unique movement of the moving object by hostile learning. Learn an encoder to extract features.
  • FIG. 1 is a functional block diagram showing an example of the configuration of the learning device of the present embodiment.
  • FIG. 2 is a diagram showing an example of time-series two-dimensional data input by the learning device.
  • FIG. 3 is a diagram showing an example of division of a data set used for learning.
  • FIG. 4 is a diagram showing an example of data classification.
  • FIG. 5 is a diagram showing an example of a data set used for training.
  • FIG. 6 is a diagram showing an example of the configuration of the model of the present embodiment.
  • FIG. 7 is a diagram illustrating an example of learning a discriminative model.
  • FIG. 8 is a diagram illustrating a code input to the decoder.
  • FIG. 9 is a diagram showing an example of learning of a reproduction model that generates a common motion.
  • FIG. 1 is a functional block diagram showing an example of the configuration of the learning device of the present embodiment.
  • FIG. 2 is a diagram showing an example of time-series two-dimensional data input by the learning
  • FIG. 10 is a diagram showing an example of learning of a reproduction model that generates a common motion.
  • FIG. 11 is a diagram showing an example of learning of a reproduction model that generates a unique motion.
  • FIG. 12 is a diagram showing an example of learning of a reproduction model that generates a unique motion.
  • FIG. 13 is a diagram showing an example of learning of a reproduction model that generates a unique motion.
  • FIG. 14 is a diagram showing an example of learning of a reproduction model that generates a unique motion.
  • FIG. 15 is a diagram showing an example of making a target celebrity make a new movement.
  • FIG. 16 is a diagram showing an embodiment for determining whether or not the person in the video is the person himself / herself.
  • FIG. 17 is a diagram showing an example of generating a skeleton having a general-purpose movement without individuality.
  • FIG. 18 is a diagram showing an example of the hardware configuration of the learning device.
  • FIG. 1 is a functional block diagram showing an example of the configuration of the learning device 1 of the present embodiment.
  • the learning device 1 shown in the figure includes a data input unit 11, a data arrangement unit 12, a data block 13, a model 14, and a learning unit 15, and outputs a trained model 16.
  • the data input unit 11 inputs time-series two-dimensional data that records the movement of a person (skeleton), and divides it into three types: training data, validation data, and test data.
  • a human skeleton will be described as an example, but the description is not limited to this.
  • skeleton skeleton
  • data that can detect a time change of a skeleton (skeleton) from an image of a moving object having a skeleton such as an animal, a doll, or a robot.
  • a motion marker may be used, or the input itself can be converted into a human image and handled internally as skeleton information.
  • the data input unit 11 inputs a large number of skeleton blocks of a large number of movements of a large number of people, and divides the large number of input skeleton blocks into three types of training data, validation data, and test data, as shown in FIG. ..
  • four-fifths of the data for M characters were set as training data, and one-fifth was set as validation data, and the data for the remaining N characters was used as test data.
  • the amount of test data can be any number.
  • the data arrangement unit 12 classifies the skeleton block input by the data input unit 11 by a person (skeleton), a movement (Motion), and a visual angle (View-Angle), and stores the skeleton block in the data block 13.
  • FIG. 4 shows an example of the arrangement format classified by the data arrangement unit 12.
  • the person is the uppercase alphabet "A, B, C, D, ..., N”
  • the movement is the number "1, 2, 3, 4, ..., q”
  • the visual angle Is represented by the symbol "!,”, #, &, ..., (", and the individuality of a person's movement is represented by the lowercase alphabet" _a, _b, _c, ..., _n.
  • Movement is the type of movement of the skeleton. For example, movement 1 is walking, movement 2 is dancing, movement 3 is soccer movement, and the like.
  • Visual angle is the limbs that make up the skeleton. Indicates the angle of the observation direction. For example, there are seven angles of 90 degrees to the left, 60 degrees to the left, 30 degrees to the left, front, 30 degrees to the right, 60 degrees to the right, and 90 degrees to the right.
  • the skeleton block "B1! _B" is a skeleton in which Mr. B is moving 1, is a skeleton seen from an angle "!, And includes the individuality of Mr. B's movement. Is shown.
  • the data block 13 stores a learning data set with a label indicating a person for each skeleton block.
  • the model 14 includes a unique motion feature encoder (Unique Motion Encoder) 141, a common motion encoder (Common Motion Encoder) 142, a visual angle encoder (View-Angle Encoder) 143, and a body type encoder (Skeleton Encoder) 14. It has a classifier 145 and a motion reproduction decoder 146. Model 14 is also called a neural network.
  • the unique motion feature encoder 141 and the common motion encoder 142 are encoders that extract time-varying motion features.
  • the intrinsic movement feature encoder 141 learns the movement of the skeleton so as to reduce the identification loss (error) so that the person can be identified.
  • the common motion encoder 142 learns the motion of the skeleton so as to increase the identification loss so that the person cannot be identified.
  • two encoders for extracting the characteristics of time-varying movements are provided, but three or more encoders may be provided.
  • the visual angle encoder 143 and the body shape encoder 144 are encoders that extract time-independent skeleton features.
  • the visual angle encoder 143 extracts the visual angle feature.
  • the body shape encoder 144 extracts the characteristics of the body shape of the skeleton. In the present embodiment, two encoders for extracting the characteristics of the skeleton that do not depend on time are provided, but one may be provided, or three or more encoders may be provided.
  • the classifier 145 identifies a person from the characteristics of the movement of the skeleton.
  • the motion reproduction decoder 146 outputs a skeleton block in which the motion of the source skeleton is reconstructed into the target skeleton from the features extracted by inputting the source skeleton block and the target skeleton block.
  • the learning unit 15 performs hostile learning by the discriminator 145 based on the characteristics of the movements extracted by the encoders 141 and 142, and learns the discriminative model. More specifically, the discrimination result by the motion feature classifier 145, which is a combination of the motion eigen feature extracted by the eigenmotion feature encoder 141 and the common motion extracted by the common motion encoder 142, is learned to reduce the discrimination loss. , The discrimination result by the common motion classifier 145 extracted by the common motion encoder 142 is learned so as to increase the discrimination loss.
  • FIG. 7 shows an example of learning a discriminative model.
  • the skeleton block represented by “A1! _A” was input to the encoders 141 and 142 as training data.
  • This skeleton block is time-series data of the skeleton when Mr. A's movement 1 is viewed from the angle indicated by "!.
  • the label "A” is attached to this skeleton block.
  • This skeleton block is input to the intrinsic motion feature encoder 141 and the common motion encoder 142, and the motion feature that combines the motion intrinsic feature and the common motion output from the encoders 141 and 142 is input to the classifier 145, and the classifier 145 inputs the motion feature. Get the identified output result.
  • the obtained output result is compared with the label "A" given to the skeleton block, and the gradient of the error function is back-propagated as shown by the broken arrow in the figure so as to reduce the discrimination loss, and the discriminative model is trained. ..
  • a negative constant is applied to the gradient transmitted to the common motion encoder 142 in the previous stage through the Gradient Reversal Layer (GRL) 147.
  • the gradient transmitted to the unique motion feature encoder 141 is input as it is.
  • GRL is described in Non-Patent Document 2.
  • the unique features for the T frame are extracted from the unique motion feature encoder 141, and the common motion for the T frame is extracted from the common motion encoder 142.
  • the unique movement reflects the unique feature.
  • the visual angle is extracted from the visual angle encoder 143.
  • the body shape is extracted from the body shape encoder 144.
  • Visual angle and body shape are time-independent features.
  • a code in which the visual angle and the body shape are connected by T frames along the time axis is input to the motion reproduction decoder 146 to the motion features extracted by the encoders 141 and 142.
  • the motion reproduction decoder 146 outputs a skeleton block obtained by viewing the body shape skeleton extracted from the moving body shape encoder 144 extracted by the encoders 141 and 142 from the visual angle extracted from the visual angle encoder 143.
  • the output skeleton block has the same format as the input skeleton block. If only the motion features extracted by the common motion encoder 142 are used, a skeleton block of common motion that does not have the unique characteristics of a person can be obtained.
  • a skeleton block of the unique motion obtained by adding the unique feature of a person to the common motion can be obtained.
  • the learning unit 15 inputs the skeleton block of Mr. A to the common motion encoder 142, inputs the skeleton block of Mr. B to the visual angle encoder 143 and the body shape encoder 144, and the motion of Mr. A (common).
  • the model 14 was trained so as to output the skeleton block of Mr. B's skeleton as seen from the visual angle of Mr. B. Specifically, the skeleton block represented by "A1" _a "was input to the common motion encoder 142, and the skeleton block represented by" B2! _B "was input to the visual angle encoder 143 and the body shape encoder 144.
  • the feature of motion 1 was extracted by the common motion encoder 142, the visual angle for viewing Mr.
  • the learning unit 15 inputs the skeleton block of Mr. A to the common motion encoder 142 and the visual angle encoder 143, inputs the skeleton block of Mr. B to the body encoder 144, and moves the motion of Mr. A (common).
  • the model 14 was trained so as to output the skeleton block of Mr. B's skeleton as seen from the visual angle of Mr. A. Specifically, the skeleton block represented by "A1" _a "was input to the common motion encoder 142 and the visual angle encoder 143, and the skeleton block represented by" B2! _B "was input to the body type encoder 144.
  • the feature of motion 1 was extracted by the common motion encoder 142, the visual angle for viewing Mr.
  • A's skeleton was extracted by the visual angle encoder 143, and the body shape of Mr. B's skeleton was extracted by the body shape encoder 144.
  • the output result of the skeleton block "B1" "" in which the skeleton of Mr. B's body shape in motion 1 is viewed from the visual angle "" " is obtained from the motion reproduction decoder 146.
  • the obtained output result is compared with the skeleton block "B1" _b "which is close to the correct answer, and the error is back-propagated as shown by the broken line arrow in the figure to be trained by the reproduction model.
  • the learning unit 15 inputs the skeleton block of Mr. A to the common motion encoder 142, and inputs the skeleton block of Mr. B to the unique motion feature encoder 141, the visual angle encoder 143, and the body shape encoder 144.
  • the model 14 was learned so as to output the skeleton block of Mr. B's skeleton, which has Mr. B's personality and moves Mr. A, as seen from Mr. B's visual angle.
  • the skeleton block represented by "A1" _a is input to the common motion encoder 142
  • the learning unit 15 inputs the skeleton block of Mr. A to the common motion encoder 142, and inputs the skeleton block having the individuality of Mr. B who has the same movement as Mr. A to the unique motion feature encoder 141. , Input another skeleton block of Mr. B into the visual angle encoder 143 and the body shape encoder 144, and see the skeleton of Mr. B who has the same movement as Mr. A who has the individuality of Mr. B from the visual angle of Mr. B.
  • the model 14 was trained to output blocks. Specifically, the skeleton block indicated by "A1" _a "is input to the common motion encoder 142, the skeleton block indicated by" B1!
  • the skeleton block input to the unique motion feature encoder 141 is different between the example of FIG. 11 and the example of FIG.
  • the motion 1 of the skeleton block input to the common motion encoder 142 and the motion 1 of the skeleton block input to the intrinsic motion feature encoder 141 are the same.
  • the example of FIG. 12 can extract the unique feature of the movement more clearly than the example of FIG.
  • the learning unit 15 inputs the skeleton block of Mr. A to the common motion encoder 142 and the visual angle encoder 143, and inputs the skeleton block of Mr. B to the unique motion feature encoder 141 and the body shape encoder 144.
  • the model 14 was learned so as to output the skeleton block of Mr. B's skeleton, which has Mr. B's personality and moves, as seen from Mr. A's visual angle.
  • the skeleton block indicated by "A1! _A” is input to the common motion encoder 142 and the visual angle encoder 143, and the skeleton block indicated by "B2" _b "is used as the unique motion feature encoder 141 and the body shape encoder 144. I input it.
  • the common motion encoder 142 extracts the feature of motion 1
  • the visual angle encoder 143 extracts the visual angle for viewing Mr. A's skeleton
  • the unique motion feature encoder 141 extracts the unique feature of Mr. B's motion
  • the body shape encoder 144 The body shape of Mr. B's skeleton was extracted.
  • Is obtained from the motion reproduction decoder 146 The obtained output result is compared with the correct skeleton block "B1! _B", and the error is back-propagated as shown by the broken line arrow in the figure to be trained by the reproduction model.
  • the learning unit 15 inputs the skeleton block of Mr. A to the common movement encoder 142 and the visual angle encoder 143, and the skeleton block having the individuality of Mr. B who makes the same movement as Mr. A is a unique movement feature.
  • Input to the encoder 141 input the skeleton block of Mr. B to the body type encoder 144, and see the skeleton block of Mr. B who has the same movement as Mr. A who has the individuality of Mr. B from the visual angle of Mr. A.
  • the model 14 was trained to output. Specifically, the skeleton block indicated by "A1! _A" is input to the common motion encoder 142 and the visual angle encoder 143, and the skeleton block indicated by "B1!
  • _B is input to the unique motion feature encoder 141.
  • the skeleton block represented by B2 "_b” was input to the body encoder 144.
  • the common motion encoder 142 extracts the feature of motion 1
  • the visual angle encoder 143 extracts the visual angle for viewing Mr. A's skeleton
  • the unique motion feature encoder 141 extracts the unique feature of motion that depends on Mr. B's motion 1.
  • the body shape of Mr. B's skeleton was extracted by the body shape encoder 144.
  • the skeleton block input to the unique motion feature encoder 141 is different between the example of FIG. 13 and the example of FIG.
  • the motion 1 of the skeleton block input to the common motion encoder 142 and the motion 1 of the skeleton block input to the intrinsic motion feature encoder 141 are the same.
  • the example of FIG. 14 can extract the unique feature of the movement more clearly than the example of FIG.
  • the trained model 16 is obtained by repeating the learning of the model 14 by the learning unit 15.
  • FIG. 15 is an example of using a reproduction model to make a target celebrity make a new movement.
  • time-series 2D data of the skeleton showing the movement of the celebrity is input to the unique motion feature encoder 141 and the body shape encoder 144, and the time-series 2D data of the skeleton of the source subject who made the movement that the celebrity wants to reproduce is input to the common motion encoder 142. Is input to the visual angle encoder 143.
  • time-series two-dimensional data can be obtained in which the skeleton of the celebrity is made to perform the movement of the subject of the source by adding the personality of the celebrity.
  • the source subject data is input to the body encoder 144 instead of the celebrity data, the source subject can generate data that imitates the movement of the celebrity.
  • FIG. 16 is an example of using an identification model to identify whether a person in a video is genuine and fairly spoofed.
  • the video showing the person to be judged is image-recognized by OpenPose etc. and the motion data is extracted.
  • the extracted motion data is input to the encoders 141 and 142 of the identification model, it can be determined whether or not the person in the video is the person himself / herself.
  • the discriminative model has already learned the movement of the person to be determined.
  • FIG. 17 is an example of using a reproduction model to generate a skeleton that has a general-purpose movement without individuality.
  • the time-series two-dimensional data of the skeleton of the source subject is input to the common motion encoder 142, the visual angle encoder 143, and the body shape encoder 144.
  • the common motion encoder 142 is learned to extract common motions that have been filtered for individuality.
  • time-series two-dimensional data of the skeleton having a common movement excluding the individuality of the movement of the subject of the source can be obtained.
  • time-series two-dimensional data of skeletons that perform general-purpose movements without individuality it can be used to create a movement database. Further, by inputting time-series two-dimensional data of various visual angles, it is possible to generate time-series two-dimensional data of various visual angles.
  • the learning device 1 of the present embodiment uses the data input unit 11 for inputting the time-series data recording the movement of the skeleton and the time-series data to extract the characteristics of the skeleton independent of time. It outputs encoders 143 and 144, two or more encoders 141 and 142 that extract time-varying motion characteristics, a classifier 145 that identifies a skeleton from motion characteristics, and time-series data that records new motions of the skeleton.
  • a learning unit 15 for learning a model 14 including a motion reproduction decoder 146 is provided. The learning unit 15 extracts the motion features common to the skeleton's unique motion features 141 by hostile learning for the encoders 141 and 142 that extract the motion features. Learn 142.
  • two or more encoders 141 and 142 for extracting motion characteristics are prepared, and by hostile learning, the intrinsic motion feature encoder 141 advances learning to strengthen individual identification and extracts unique motions. Since the common motion encoder 142 advances learning that does not identify individuals and extracts common motions, it is possible to generate time-series two-dimensional data that records new motions in which target-specific motions are added to common motions. As a result, for example, in the server space of a digital twin, anyone can easily become the target person.
  • the learning device 1 described above includes, for example, a central processing unit (CPU) 901, a memory 902, a storage 903, a communication device 904, an input device 905, and an output device 906, as shown in FIG.
  • CPU central processing unit
  • a general-purpose computer system can be used.
  • the learning device 1 is realized by the CPU 901 executing a predetermined program loaded on the memory 902.
  • This program can be recorded on a computer-readable recording medium such as a magnetic disk, an optical disk, or a semiconductor memory, or can be distributed via a network.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • Biophysics (AREA)
  • Computational Linguistics (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Molecular Biology (AREA)
  • Computing Systems (AREA)
  • Biomedical Technology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Mathematical Physics (AREA)
  • Software Systems (AREA)
  • Health & Medical Sciences (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Multimedia (AREA)
  • Image Analysis (AREA)
  • Processing Or Creating Images (AREA)

Abstract

本実施形態の学習装置1は、スケルトンの動きを記録した時系列データを入力するデータ入力部11と、時系列データを用いて、時間に依存しないスケルトンの特徴を抽出するエンコーダ143,144、時間変動する動きの特徴を抽出する2つ以上のエンコーダ141,142、動きの特徴からスケルトンを識別する識別器145、およびスケルトンの新たな動きを記録した時系列データを出力する動き再生デコーダ146を含むモデル14を学習する学習部15を備える。学習部15は、動きの特徴を抽出するエンコーダ141,142について、敵対的学習により、スケルトンの固有的な動きの特徴を抽出する固有動き特徴エンコーダ141と共通する動きの特徴を抽出する共通動きエンコーダ142を学習する。

Description

学習装置、学習方法、およびプログラム
 本発明は、学習装置、学習方法、およびプログラムに関する。
 故人や有名人のフェイク映像を生成する方法として、口の部分だけ映像を生成し、生成した映像を本人の映像に部分的に合成する技術や、ターゲットのアバター(コンピュータグラフィックスで作成したキャラクタ)に動きを真似させる技術が知られている。
 非特許文献1には、深層学習を用いて、動的要素(モーション)と静的要素(スケルトンとビューアングル)を分離し、動的要素と静的要素を再構成する技術が開示されている。
Kfir Aberman, et al., "Learning Character-Agnostic Motion for Motion Retargeting in 2D", ACM Trans. Graph., Vol. 38, No. 4, Article 75. Yaroslav Ganin and Victor Lempitsky, "Unsupervised Domain Adaptation by Backpropagation", Proceedings of the 32nd International Conference on Machine Learning
 従来の技術は、体型などの個性を反映することは考慮されているが、動きに含まれる個性を反映することは考慮されていない。有名人の動きを真似する人のモーションを反映する方法が考えられるが、多くは一部の癖を誇張する表現であり、利用できるシーンは限られている。
 本発明は、上記に鑑みてなされたものであり、ターゲットの固有の動きを反映した動きを生成することを目的とする。
 本発明の一態様の学習装置は、動体の動きを記録した時系列データを入力する入力部と、前記時系列データを用いて、時間に依存しない前記動体の特徴を抽出するエンコーダ、時間変動する動きの特徴を抽出する2つ以上のエンコーダ、前記動きの特徴から前記動体を識別する識別器、および前記動体の新たな動きを記録した時系列データを出力するデコーダを含むモデルを学習する学習部を備え、前記学習部は、前記動きの特徴を抽出する2つ以上のエンコーダについて、敵対的学習により、前記動体の固有的な動きの特徴を抽出するエンコーダと共通する動きの特徴を抽出するエンコーダを学習する。
 本発明の一態様の学習方法は、コンピュータが実行する学習方法であって、動体の動きを記録した時系列データを入力し、前記時系列データを用いて、時間に依存しない前記動体の特徴を抽出するエンコーダ、時間変動する動きの特徴を抽出する2つ以上のエンコーダ、前記動きの特徴から前記動体を識別する識別器、および前記動体の新たな動きを記録した時系列データを出力するデコーダを含むモデルを学習し、前記モデルを学習する際、前記動きの特徴を抽出する2つ以上のエンコーダについて、敵対的学習により、前記動体の固有的な動きの特徴を抽出するエンコーダと共通する動きの特徴を抽出するエンコーダを学習する。
 本発明によれば、ターゲットの固有の動きを反映した動きを生成できる。
図1は、本実施形態の学習装置の構成の一例を示す機能ブロック図である。 図2は、学習装置が入力する時系列2次元データの一例を示す図である。 図3は、学習に用いるデータセットの分割の一例を示す図である。 図4は、データの分類の一例を示す図である。 図5は、学習に用いるデータセットの一例を示す図である。 図6は、本実施形態のモデルの構成の一例を示す図である。 図7は、識別モデルの学習の一例を説明する図である。 図8は、デコーダに入力するコードを説明する図である。 図9は、共通動きを生成する再生モデルの学習の一例を示す図である。 図10は、共通動きを生成する再生モデルの学習の一例を示す図である。 図11は、固有動きを生成する再生モデルの学習の一例を示す図である。 図12は、固有動きを生成する再生モデルの学習の一例を示す図である。 図13は、固有動きを生成する再生モデルの学習の一例を示す図である。 図14は、固有動きを生成する再生モデルの学習の一例を示す図である。 図15は、ターゲットとなる有名人に新しい動きをさせる実施例を示す図である。 図16は、映像内の人物が本人であるか否かを判定する実施例を示す図である。 図17は、個性を無くした汎用的な動きをするスケルトンを生成する実施例を示す図である。 図18は、学習装置のハードウェア構成の一例を示す図である。
 以下、本発明の実施の形態について図面を用いて説明する。
 図1は、本実施形態の学習装置1の構成の一例を示す機能ブロック図である。同図に示す学習装置1は、データ入力部11、データ配置部12、データブロック13、モデル14、および学習部15を備えて、学習済みモデル16を出力する。
 データ入力部11は、人(スケルトン)の動きを記録した時系列2次元データを入力し、トレーニングデータ、バリデーションデータ、およびテストデータの3つのタイプに分割する。時系列2次元データは、図2に示すように、15ポイントのスケルトンをTフレーム分(例えばT=64)有するスケルトンブロックである。スケルトンの1つのポイントはx座標値とy座標値を有する。つまり、1つのスケルトンで15×2=30の座標値を有する。なお、以下では、人のスケルトンを例に説明するが、これに限るものではない。例えば、動物、人形、あるいはロボットなど骨格を有する動体を撮影した画像から骨格(スケルトン)の時間変化を検出できるデータを利用できる。人の骨格に限らず、モーションマーカーを利用してもよいし、入力自体を人の映像にして、内部で骨格情報にして扱う事もできる。
 データ入力部11は、多数の人物の多数の動きのスケルトンブロックを入力し、図3に示すように、入力した多数のスケルトンブロックをトレーニングデータ、バリデーションデータ、およびテストデータの3つのタイプに分割する。図3の例では、Mキャラクタ分のデータについて、5分の4をトレーニングデータ、5分の1をバリデーションデータとして設定し、残りのNキャラクタ分のデータをテストデータとした。テストデータの量はいくつでもよい。
 データ配置部12は、データ入力部11の入力したスケルトンブロックを人物(Skeleton)、動き(Motion)、および視覚角度(View-Angle)で分類し、データブロック13に格納する。図4にデータ配置部12が分類する配置フォーマットの一例を示す。図4では、各スケルトンブロックについて、人物を大文字アルファベット「A,B,C,D,・・・,N」、動きを数字「1,2,3,4,・・・,q」、視覚角度を記号「!,”,#,&,・・・,(」、人の動きの個性を小文字アルファベット「_a,_b,_c,・・・,_n」で表した。人物によって体型(スケルトン)が異なる。動きとは、スケルトンの動きのタイプ(種別)である。例えば、動き1は歩く、動き2はダンス、動き3はサッカーの動作などである。視覚角度とは、スケルトンを構成する肢体を観察する方向の角度を示す。例えば、左90度、左60度、左30度、正面、右30度、右60度、右90度の7つの角度である。角度は、肢体が向く方向を示すものでもよい。一例として、スケルトンブロック「B1!_b」は、Bさんが動き1をしているスケルトンであって、角度「!」からみたスケルトンであり、Bさんの動きの個性を含むことを示す。
 データブロック13は、図5に示すように、スケルトンブロックごとに人物を示すラベルを付与した学習用のデータセットを格納する。
 モデル14は、図6に示すように、固有動き特徴エンコーダ(Unique Motion Encoder)141、共通動きエンコーダ(Common Motion Encoder)142、視覚角度エンコーダ(View-Angle Encoder)143、体型エンコーダ(Skeleton Encoder)144、識別器(Classifier)145、および動き再生デコーダ(Motion Reconstruction Decoder)146を有する。モデル14は、ニューラルネットワークとも言う。
 固有動き特徴エンコーダ141と共通動きエンコーダ142は、時間変動する動きの特徴を抽出するエンコーダである。固有動き特徴エンコーダ141は、人物を識別できるように、識別ロス(誤差)を低くするようにスケルトンの動きを学習する。共通動きエンコーダ142は、人物を識別できないように、識別ロスを高くするようにスケルトンの動きを学習する。本実施形態では、時間変動する動きの特徴を抽出するエンコーダを2つ備えたが、3つ以上備えてもよい。
 視覚角度エンコーダ143と体型エンコーダ144は、時間に依存しないスケルトンの特徴を抽出するエンコーダである。視覚角度エンコーダ143は、視覚角度の特徴を抽出する。体型エンコーダ144は、スケルトンの体型の特徴を抽出する。本実施形態では、時間に依存しないスケルトンの特徴を抽出するエンコーダを2つ備えたが、1つでもよいし、3つ以上備えてもよい。
 識別器145は、スケルトンの動きの特徴から人物を識別する。
 動き再生デコーダ146は、ソースとなるスケルトンブロックとターゲットとなるスケルトンブロックを入力して抽出された特徴から、ソースのスケルトンの動きをターゲットのスケルトンへ再構成したスケルトンブロックを出力する。
 エンコーダ141,142と識別器145で動きの特徴から人物を識別する識別モデルを構成し、エンコーダ141~144と動き再生デコーダ146で動きの特徴とスケルトンの特徴からスケルトンの動きを再構成する再生モデルを構成する。
 続いて、学習部15による識別モデルの学習について説明する。
 学習部15は、エンコーダ141,142の抽出した動きの特徴で識別器145による敵対的学習を行い、識別モデルを学習する。より具体的には、固有動き特徴エンコーダ141の抽出した動きの固有特徴と共通動きエンコーダ142の抽出した共通動きを合わせた動き特徴の識別器145による識別結果は識別ロスを低くするように学習し、共通動きエンコーダ142の抽出した共通動きの識別器145による識別結果は識別ロスを高くするように学習する。
 図7に、識別モデルの学習の一例を示す。図7の例では、学習データとして「A1!_a」で示されるスケルトンブロックをエンコーダ141,142に入力した。このスケルトンブロックは、Aさんの動き1を「!」で示される角度から見たスケルトンの時系列データである。このスケルトンブロックにはラベル「A」が付与されている。このスケルトンブロックを固有動き特徴エンコーダ141と共通動きエンコーダ142に入力し、エンコーダ141,142から出力された動きの固有特徴と共通動きを合わせた動き特徴を識別器145に入力し、識別器145で識別した出力結果を得る。得られた出力結果とスケルトンブロックに付与されたラベル「A」を比較し、識別ロスを低くするように、誤差関数の勾配を図中の破線矢印のように逆伝搬させて識別モデルに学習させる。誤差逆伝搬時に、Gradient Reversal Layer(GRL)147を通し、前段の共通動きエンコーダ142へ伝わる勾配に負の定数をかける。固有動き特徴エンコーダ141へ伝わる勾配はそのまま入力する。これにより、固有動き特徴エンコーダ141は人を識別できるように学習させると同時に、共通動きエンコーダ142は人を識別できないように学習させることができる。なお、GRLについては非特許文献2に記載がある。
 続いて、学習部15による再生モデルの学習について説明する。再生モデルの学習の説明に際して、図8を参照し、動き再生デコーダ146に入力するコードの一例について説明する。
 スケルトンブロックが入力されると、固有動き特徴エンコーダ141からはTフレーム分の固有特徴が抽出され、共通動きエンコーダ142からはTフレーム分の共通動きが抽出される。固有特徴と共通動きを合わせると共通動きに固有特徴を反映した固有動きとなる。
 視覚角度エンコーダ143からは視覚角度が抽出される。体型エンコーダ144からは体型が抽出される。視覚角度および体型は、時間に依存しない特徴である。
 図8に示すように、エンコーダ141,142で抽出された動きの特徴に、視覚角度と体型を時間軸に沿ってTフレーム分連結したコードを動き再生デコーダ146に入力する。動き再生デコーダ146からは、エンコーダ141,142で抽出された動きをする体型エンコーダ144から抽出された体型のスケルトンを視覚角度エンコーダ143から抽出された視覚角度から見たスケルトンブロックが出力される。出力されるスケルトンブロックは、入力されたスケルトンブロックと同じ形式である。なお、共通動きエンコーダ142で抽出された動きの特徴のみを用いると、人の固有の特徴を持たない共通動きのスケルトンブロックが得られる。共通動きエンコーダ142で抽出された共通動きに、固有動き特徴エンコーダ141で抽出された固有特徴を加えることで、共通動きに人の固有の特徴を加えた固有動きのスケルトンブロックが得られる。
 図9および図10を参照し、共通動きを生成する再生モデルの学習例について説明する。
 図9の例では、学習部15は、Aさんのスケルトンブロックを共通動きエンコーダ142に入力し、Bさんのスケルトンブロックを視覚角度エンコーダ143と体型エンコーダ144に入力して、Aさんの動き(共通動き)をするBさんのスケルトンをBさんの視覚角度から見たスケルトンブロックを出力するようにモデル14を学習した。具体的には、「A1”_a」で示されるスケルトンブロックを共通動きエンコーダ142に入力し、「B2!_b」で示されるスケルトンブロックを視覚角度エンコーダ143と体型エンコーダ144に入力した。共通動きエンコーダ142で動き1の特徴を抽出し、視覚角度エンコーダ143でBさんのスケルトンを見る視覚角度を抽出し、体型エンコーダ144でBさんのスケルトンの体型を抽出した。エンコーダ142~144で抽出した特徴を再構成して、動き1をするBさんの体型のスケルトンを視覚角度「!」からみたスケルトンブロック「B1!’」の出力結果を動き再生デコーダ146から得る。得られた出力結果と正解に近いスケルトンブロック「B1!_b」を比較し、誤差を図中の破線矢印のように逆伝搬させて再生モデルに学習させる。
 図10の例では、学習部15は、Aさんのスケルトンブロックを共通動きエンコーダ142と視覚角度エンコーダ143に入力し、Bさんのスケルトンブロックを体型エンコーダ144に入力して、Aさんの動き(共通動き)をするBさんのスケルトンをAさんの視覚角度から見たスケルトンブロックを出力するようにモデル14を学習した。具体的には、「A1”_a」で示されるスケルトンブロックを共通動きエンコーダ142と視覚角度エンコーダ143に入力し、「B2!_b」で示されるスケルトンブロックを体型エンコーダ144に入力した。共通動きエンコーダ142で動き1の特徴を抽出し、視覚角度エンコーダ143でAさんのスケルトンを見る視覚角度を抽出し、体型エンコーダ144でBさんのスケルトンの体型を抽出した。エンコーダ142~144で抽出した特徴を再構成して、動き1をするBさんの体型のスケルトンを視覚角度「”」からみたスケルトンブロック「B1”’」の出力結果を動き再生デコーダ146から得る。得られた出力結果と正解に近いスケルトンブロック「B1”_b」を比較し、誤差を図中の破線矢印のように逆伝搬させて再生モデルに学習させる。
 図11~図14を参照し、固有動きを生成する再生モデルの学習例について説明する。
 図11の例では、学習部15は、Aさんのスケルトンブロックを共通動きエンコーダ142に入力し、Bさんのスケルトンブロックを固有動き特徴エンコーダ141、視覚角度エンコーダ143、および体型エンコーダ144に入力して、Bさんの個性を持ったAさんの動きをするBさんのスケルトンをBさんの視覚角度から見たスケルトンブロックを出力するようにモデル14を学習した。具体的には、「A1”_a」で示されるスケルトンブロックを共通動きエンコーダ142に入力し、「B2!_b」で示されるスケルトンブロックを固有動き特徴エンコーダ141、視覚角度エンコーダ143、および体型エンコーダ144に入力した。共通動きエンコーダ142で動き1の特徴を抽出し、固有動き特徴エンコーダ141でBさんの動きの固有特徴を抽出し、視覚角度エンコーダ143でBさんのスケルトンを見る視覚角度を抽出し、体型エンコーダ144でBさんのスケルトンの体型を抽出した。エンコーダ141~144で抽出した特徴を再構成して、Bさんの個性を持った動き1をするBさんの体型のスケルトンを視覚角度「!」からみたスケルトンブロック「B1!_b’」の出力結果を動き再生デコーダ146から得る。得られた出力結果と正解のスケルトンブロック「B1!_b」を比較し、誤差を図中の破線矢印のように逆伝搬させて再生モデルに学習させる。
 図12の例では、学習部15は、Aさんのスケルトンブロックを共通動きエンコーダ142に入力し、Aさんと同じ動きをするBさんの個性を持ったスケルトンブロックを固有動き特徴エンコーダ141に入力し、Bさんの別のスケルトンブロックを視覚角度エンコーダ143と体型エンコーダ144に入力して、Bさんの個性を持ったAさんと同じ動きをするBさんのスケルトンをBさんの視覚角度から見たスケルトンブロックを出力するようにモデル14を学習した。具体的には、「A1”_a」で示されるスケルトンブロックを共通動きエンコーダ142に入力し、「B1!_b」で示されるスケルトンブロックを固有動き特徴エンコーダ141に入力し、「B2!_b」で示されるスケルトンブロックを視覚角度エンコーダ143と体型エンコーダ144に入力した。共通動きエンコーダ142で動き1の特徴を抽出し、固有動き特徴エンコーダ141でBさんの動き1に依存する動きの固有特徴を抽出し、視覚角度エンコーダ143でBさんのスケルトンを見る視覚角度を抽出し、体型エンコーダ144でBさんのスケルトンの体型を抽出した。エンコーダ141~144で抽出した特徴を再構成して、Bさんの個性を持った動き1をするBさんの体型のスケルトンを視覚角度「!」からみたスケルトンブロック「B1!_b’」の出力結果を動き再生デコーダ146から得る。得られた出力結果と正解のスケルトンブロック「B1!_b」を比較し、誤差を図中の破線矢印のように逆伝搬させて再生モデルに学習させる。なお、固有動き特徴エンコーダ141へは、Bさんが動き1をする他のスケルトンブロック(例えば「B1”_b」で示されるスケルトンブロック)を入力しても学習可能である。
 図11の例と図12の例とでは、固有動き特徴エンコーダ141に入力するスケルトンブロックが異なっている。図12の例では、共通動きエンコーダ142に入力するスケルトンブロックの動き1と固有動き特徴エンコーダ141に入力するスケルトンブロックの動き1が同じである。その結果、図12の例は、図11の例よりも、より鮮明に動きの固有特徴を抽出できる。
 図13の例では、学習部15は、Aさんのスケルトンブロックを共通動きエンコーダ142と視覚角度エンコーダ143に入力し、Bさんのスケルトンブロックを固有動き特徴エンコーダ141と体型エンコーダ144に入力して、Bさんの個性を持ったAさんの動きをするBさんのスケルトンをAさんの視覚角度から見たスケルトンブロックを出力するようにモデル14を学習した。具体的には、「A1!_a」で示されるスケルトンブロックを共通動きエンコーダ142と視覚角度エンコーダ143に入力し、「B2”_b」で示されるスケルトンブロックを固有動き特徴エンコーダ141と体型エンコーダ144に入力した。共通動きエンコーダ142で動き1の特徴を抽出し、視覚角度エンコーダ143でAさんのスケルトンを見る視覚角度を抽出し、固有動き特徴エンコーダ141でBさんの動きの固有特徴を抽出し、体型エンコーダ144でBさんのスケルトンの体型を抽出した。エンコーダ141~144で抽出した特徴を再構成して、Bさんの個性を持った動き1をするBさんの体型のスケルトンを視覚角度「!」からみたスケルトンブロック「B1!_b’」の出力結果を動き再生デコーダ146から得る。得られた出力結果と正解のスケルトンブロック「B1!_b」を比較し、誤差を図中の破線矢印のように逆伝搬させて再生モデルに学習させる。
 図14の例では、学習部15は、Aさんのスケルトンブロックを共通動きエンコーダ142と視覚角度エンコーダ143に入力し、Aさんと同じ動きをするBさんの個性を持ったスケルトンブロックを固有動き特徴エンコーダ141に入力し、Bさんのスケルトンブロックを体型エンコーダ144に入力して、Bさんの個性を持ったAさんと同じ動きをするBさんのスケルトンをAさんの視覚角度から見たスケルトンブロックを出力するようにモデル14を学習した。具体的には、「A1!_a」で示されるスケルトンブロックを共通動きエンコーダ142と視覚角度エンコーダ143に入力し、「B1!_b」で示されるスケルトンブロックを固有動き特徴エンコーダ141に入力し、「B2”_b」で示されるスケルトンブロックを体型エンコーダ144に入力した。共通動きエンコーダ142で動き1の特徴を抽出し、視覚角度エンコーダ143でAさんのスケルトンを見る視覚角度を抽出し、固有動き特徴エンコーダ141でBさんの動き1に依存する動きの固有特徴を抽出し、体型エンコーダ144でBさんのスケルトンの体型を抽出した。エンコーダ141~144で抽出した特徴を再構成して、Bさんの個性を持った動き1をするBさんの体型のスケルトンを視覚角度「!」からみたスケルトンブロック「B1!_b’」の出力結果を動き再生デコーダ146から得る。得られた出力結果と正解のスケルトンブロック「B1!_b」を比較し、誤差を図中の破線矢印のように逆伝搬させて再生モデルに学習させる。なお、固有動き特徴エンコーダ141へは、Bさんが動き1をする他のスケルトンブロック(例えば「B1”_b」で示されるスケルトンブロック)を入力しても学習可能である。
 図13の例と図14の例とでは、固有動き特徴エンコーダ141に入力するスケルトンブロックが異なっている。図14の例では、共通動きエンコーダ142に入力するスケルトンブロックの動き1と固有動き特徴エンコーダ141に入力するスケルトンブロックの動き1が同じである。その結果、図14の例は、図13の例よりも、より鮮明に動きの固有特徴を抽出できる。
 学習部15によるモデル14の学習を繰り返して、学習済みモデル16が得られる。
 次に、本実施形態の学習装置1で学習させたモデルを利用した実施例について説明する。
 図15は、再生モデルを利用して、ターゲットとなる有名人に新しい動きをさせる実施例である。
 有名人の動きを示すスケルトンの時系列2次元データを固有動き特徴エンコーダ141と体型エンコーダ144に入力し、有名人に再現させたい動きをしたソースの被写体のスケルトンの時系列2次元データを共通動きエンコーダ142と視覚角度エンコーダ143に入力する。これにより、ソースの被写体の動きに有名人の個性を加えた動きを有名人のスケルトンに行わせた時系列2次元データが得られる。例えば、有名人のアバターに得られたスケルトンの時系列2次元データを適用することで、ソースの被写体の動きに有名人の動きの個性が加えられたリアルなアバターを生成できる。
 なお、有名人のデータの代わりに、ソースの被写体のデータを体型エンコーダ144に入力すれば、ソースの被写体が有名人の動きをまねたデータを生成できる。
 図16は、識別モデルを利用して、映像内の人物が本物であるかなりすましであるかを識別する実施例である。
 判定対象の人物が写ったビデオをOpenPoseなどで画像認識して動きデータを抽出する。抽出した動きデータを識別モデルのエンコーダ141,142に入力すると、ビデオに写った人物が本人であるか否かを判定できる。なお、識別モデルは、判定対象の人物の動きを学習済みである。
 図17は、再生モデルを利用して、個性を無くした汎用的な動きをするスケルトンを生成する実施例である。
 ソースの被写体のスケルトンの時系列2次元データを共通動きエンコーダ142、視覚角度エンコーダ143、および体型エンコーダ144に入力する。共通動きエンコーダ142は、個性をフィルタリングした共通動きを抽出するように学習されている。これにより、ソースの被写体の動きの個性を取り除いた共通動きをするスケルトンの時系列2次元データが得られる。個性を無くした汎用的な動きをするスケルトンの時系列2次元データを集めることで、動きのデータベースの作成に活用できる。また、様々な視覚角度の時系列2次元データを入力することで、様々な視覚角度の時系列2次元データを生成できる。
 以上説明したように、本実施形態の学習装置1は、スケルトンの動きを記録した時系列データを入力するデータ入力部11と、時系列データを用いて、時間に依存しないスケルトンの特徴を抽出するエンコーダ143,144、時間変動する動きの特徴を抽出する2つ以上のエンコーダ141,142、動きの特徴からスケルトンを識別する識別器145、およびスケルトンの新たな動きを記録した時系列データを出力する動き再生デコーダ146を含むモデル14を学習する学習部15を備える。学習部15は、動きの特徴を抽出するエンコーダ141,142について、敵対的学習により、スケルトンの固有的な動きの特徴を抽出する固有動き特徴エンコーダ141と共通する動きの特徴を抽出する共通動きエンコーダ142を学習する。本実施形態では、動きの特徴を抽出する2つ以上のエンコーダ141,142を用意し、敵対的学習により、固有動き特徴エンコーダ141は個人識別を強くする学習を進めて固有の動きを抽出させるとともに、共通動きエンコーダ142は個人識別させない学習を進めて共通する動きを抽出させるので、ターゲット固有の動きを共通する動きに加えた新たな動きを記録した時系列2次元データを生成することができる。その結果、例えば、デジタルツインのサーバ空間内で、誰でも簡単にターゲットとする人物になりきることができる。
 上記説明した学習装置1には、例えば、図18に示すような、中央演算処理装置(CPU)901と、メモリ902と、ストレージ903と、通信装置904と、入力装置905と、出力装置906とを備える汎用的なコンピュータシステムを用いることができる。このコンピュータシステムにおいて、CPU901がメモリ902上にロードされた所定のプログラムを実行することにより、学習装置1が実現される。このプログラムは磁気ディスク、光ディスク、半導体メモリ等のコンピュータ読み取り可能な記録媒体に記録することも、ネットワークを介して配信することもできる。
 1…学習装置
 11…データ入力部
 12…データ配置部
 13…データブロック
 14…モデル
 15…学習部
 16…学習済みモデル
 141…固有動き特徴エンコーダ(Unique Motion Encoder)
 142…共通動きエンコーダ(Common Motion Encoder)
 143…視覚角度エンコーダ視覚角度エンコーダ(View-Angle Encoder)
 144…体型エンコーダ(Skeleton Encoder)
 145…識別器(Classifier)
 146…動き再生デコーダ(Motion Reconstruction Decoder)

Claims (7)

  1.  動体の動きを記録した時系列データを入力する入力部と、
     前記時系列データを用いて、時間に依存しない前記動体の特徴を抽出するエンコーダ、時間変動する動きの特徴を抽出する2つ以上のエンコーダ、前記動きの特徴から前記動体を識別する識別器、および前記動体の新たな動きを記録した時系列データを出力するデコーダを含むモデルを学習する学習部を備え、
     前記学習部は、前記動きの特徴を抽出する2つ以上のエンコーダについて、敵対的学習により、前記動体の固有的な動きの特徴を抽出するエンコーダと共通する動きの特徴を抽出するエンコーダを学習する
     学習装置。
  2.  請求項1に記載の学習装置であって、
     前記学習部は、前記識別器の誤差を低くするように誤差関数の勾配を逆伝搬させて学習し、誤差逆伝搬時に、前記共通する動きの特徴を抽出するエンコーダへは負の定数を乗算した前記勾配を逆伝搬する
     学習装置。
  3.  請求項1または2に記載の学習装置であって、
     前記時間に依存しない前記動体の特徴は、前記動体の体型および前記動体の角度である
     学習装置。
  4.  コンピュータが実行する学習方法であって、
     動体の動きを記録した時系列データを入力し、
     前記時系列データを用いて、時間に依存しない前記動体の特徴を抽出するエンコーダ、時間変動する動きの特徴を抽出する2つ以上のエンコーダ、前記動きの特徴から前記動体を識別する識別器、および前記動体の新たな動きを記録した時系列データを出力するデコーダを含むモデルを学習し、
     前記モデルを学習する際、前記動きの特徴を抽出する2つ以上のエンコーダについて、敵対的学習により、前記動体の固有的な動きの特徴を抽出するエンコーダと共通する動きの特徴を抽出するエンコーダを学習する
     学習方法。
  5.  請求項4に記載の学習方法であって、
     前記モデルを学習する際、前記識別器の誤差を低くするように誤差関数の勾配を逆伝搬させて学習し、誤差逆伝搬時に、前記共通する動きの特徴を抽出するエンコーダへは負の定数を乗算した前記勾配を逆伝搬する
     学習方法。
  6.  請求項4または5に記載の学習方法であって、
     前記時間に依存しない前記動体の特徴は、前記動体の体型および前記動体の角度である
     学習方法。
  7.  請求項1ないし3のいずれかに記載の学習装置の各部としてコンピュータを動作させるプログラム。
PCT/JP2020/026683 2020-07-08 2020-07-08 学習装置、学習方法、およびプログラム Ceased WO2022009331A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2022534553A JP7410441B2 (ja) 2020-07-08 2020-07-08 学習装置、学習方法、およびプログラム
PCT/JP2020/026683 WO2022009331A1 (ja) 2020-07-08 2020-07-08 学習装置、学習方法、およびプログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2020/026683 WO2022009331A1 (ja) 2020-07-08 2020-07-08 学習装置、学習方法、およびプログラム

Publications (1)

Publication Number Publication Date
WO2022009331A1 true WO2022009331A1 (ja) 2022-01-13

Family

ID=79552447

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2020/026683 Ceased WO2022009331A1 (ja) 2020-07-08 2020-07-08 学習装置、学習方法、およびプログラム

Country Status (2)

Country Link
JP (1) JP7410441B2 (ja)
WO (1) WO2022009331A1 (ja)

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
YANG ZHUOQIAN; ZHU WENTAO; WU WAYNE; QIAN CHEN; ZHOU QIANG; ZHOU BOLEI; LOY CHEN CHANGE: "TransMoMo: Invariance-Driven Unsupervised Video Motion Retargeting", 2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), IEEE, 13 June 2020 (2020-06-13), pages 5305 - 5314, XP033804707, DOI: 10.1109/CVPR42600.2020.00535 *

Also Published As

Publication number Publication date
JPWO2022009331A1 (ja) 2022-01-13
JP7410441B2 (ja) 2024-01-10

Similar Documents

Publication Publication Date Title
CN116561533B (zh) 一种教育元宇宙中虚拟化身的情感演化方法及终端
Essa Analysis, interpretation and synthesis of facial expressions
CN113807265B (zh) 一种多样化的人脸图像合成方法及系统
Ludl et al. Enhancing data-driven algorithms for human pose estimation and action recognition through simulation
US7257538B2 (en) Generating animation from visual and audio input
Ribet et al. Survey on style in 3d human body motion: Taxonomy, data, recognition and its applications
CN114419204A (zh) 一种视频生成方法、装置、设备和存储介质
CN115359550A (zh) 基于Transformer的步态情绪识别方法、装置、电子设备及存储介质
Marmpena et al. Generating robotic emotional body language with variational autoencoders
Khodabakhsh et al. A taxonomy of audiovisual fake multimedia content creation technology
Neverova Deep learning for human motion analysis
El-Nouby et al. Keep drawing it: Iterative language-based image generation and editing
Li et al. Multi-timescale motion-decoupled spiking transformer for audio-visual zero-shot learning
Khalid et al. Deepfakes catcher: a novel fused truncated densenet model for deepfakes detection
Kakarla et al. A real time facial emotion recognition using depth sensor and interfacing with Second Life based Virtual 3D avatar
JP7410441B2 (ja) 学習装置、学習方法、およびプログラム
CN115862139B (zh) 动作识别方法、装置以及电子设备
CN117934991A (zh) 一种基于身份保持的多类面部表情图片生成技术
Ishikawa et al. 3D face expression estimation and generation from 2D image based on a physically constraint model
Yohannes et al. Virtual reality in puppet game using depth sensor of gesture recognition and tracking
Kumar Das et al. Audio driven artificial video face synthesis using gan and machine learning approaches
CN119169502B (zh) 一种人形机器人的动作识别与模仿方法及系统
Owens Learning visual models from paired audio-visual examples
Pan Perceiving and Simulating Human-World Interactions for Egocentric Agents
Garg et al. Facial emotion recognition and classification using hybridization method

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20944342

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2022534553

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20944342

Country of ref document: EP

Kind code of ref document: A1