WO2021186668A1 - 情報処理装置、情報処理方法およびプログラム - Google Patents

情報処理装置、情報処理方法およびプログラム Download PDF

Info

Publication number
WO2021186668A1
WO2021186668A1 PCT/JP2020/012277 JP2020012277W WO2021186668A1 WO 2021186668 A1 WO2021186668 A1 WO 2021186668A1 JP 2020012277 W JP2020012277 W JP 2020012277W WO 2021186668 A1 WO2021186668 A1 WO 2021186668A1
Authority
WO
WIPO (PCT)
Prior art keywords
domain
subject
image data
image
time
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2020/012277
Other languages
English (en)
French (fr)
Inventor
善数 大貫
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sony Interactive Entertainment Inc
Original Assignee
Sony Interactive Entertainment Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sony Interactive Entertainment Inc filed Critical Sony Interactive Entertainment Inc
Priority to US17/802,699 priority Critical patent/US20230095977A1/en
Priority to JP2022507958A priority patent/JP7277668B2/ja
Priority to PCT/JP2020/012277 priority patent/WO2021186668A1/ja
Publication of WO2021186668A1 publication Critical patent/WO2021186668A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/20Analysis of motion
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0475Generative networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/088Non-supervised learning, e.g. competitive learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/0895Weakly supervised learning, e.g. semi-supervised or self-supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/094Adversarial learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/774Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/18Eye characteristics, e.g. of the iris
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person
    • G06T2207/30201Face

Definitions

  • the present invention relates to an information processing device, an information processing method and a program.
  • GAN Geneative Adversarial Networks
  • a method called cycle GAN described in Non-Patent Document 1 has been proposed.
  • the conversion source and the conversion destination of the training data are not associated with each other, and each feature is mutually learned by having the generator and the classifier learn an image group having a common feature (domain). You can build a model to transform.
  • the present invention provides an information processing device, an information processing method, and a program capable of improving the analysis accuracy of an image containing two types of subjects having different movement tendencies by applying the above cycle GAN method.
  • the purpose is to do.
  • a generator construction unit for constructing a generator that generates one of the image data of the first domain and the image data of the second domain from one of the image data of the second domain by using the cycle GAN.
  • the first domain is defined by a plurality of channels of image data including at least two time-series images including a first subject and a second subject having different movement tendencies
  • the second domain is a first domain.
  • An information processing apparatus defined by a plurality of channels of image data including at least two time-series images including a subject and not a second subject is provided.
  • the generator constructed using the cycle GAN includes an image generator that generates one of the image data of the first domain and the image data of the second domain to the other.
  • the first domain is defined by a plurality of channels of image data including at least two time-series images including a first subject and a second subject having different movement tendencies
  • the second domain is a first domain.
  • an information processing apparatus defined by a plurality of channels of image data including at least two time-series images including the subject of the above and not including the second subject.
  • the step of constructing a generator that produces one of the image data of the first domain and the image data of the second domain to the other using the cycle GAN and by the generator.
  • the first domain includes a first subject and a second subject having different movement tendencies, including a step of generating one of the image data of the first domain and the image data of the second domain to the other.
  • Defined by multi-channel image data containing at least two time-series images including, and a second domain is a multi-channel containing at least two time-series images containing the first subject and not including the second subject.
  • the generator construction unit is provided to construct a generator that generates one of the image data of the first domain and the image data of the second domain from the other by using the cycle GAN.
  • the first domain is defined by a plurality of channels of image data including at least two time-series images including a first subject and a second subject having different movement tendencies
  • the second domain is a second domain.
  • a program for operating a computer as an information processing device defined by image data of a plurality of channels including at least two time-series images including one subject and not including a second subject is provided.
  • the generator constructed using the cycle GAN includes an image generator that generates one of the image data of the first domain and the image data of the second domain to the other.
  • the first domain is defined by a plurality of channels of image data including at least two time-series images including a first subject and a second subject having different movement tendencies
  • the second domain is a second domain.
  • a program for operating a computer as an information processing device defined by image data of a plurality of channels including at least two time-series images including one subject and not including a second subject is provided.
  • FIG. 1 is a diagram showing a schematic configuration of a learning device according to an embodiment of the present invention.
  • the learning device 100 may be a single device, or may be implemented by a plurality of distributed devices working together via a network.
  • the learning device 100 is implemented by, for example, a computer having a communication interface, a processor, and a memory, and the processor operates according to a program stored in the memory or received via the communication interface, as described below. It is an information processing device in which the functions of each part are realized by software.
  • the learning device 100 includes a generator construction unit 110 that constructs a generator that generates image data of the second domain D2 from the image data of the first domain D1 by using the cycle GAN (Generative Adversarial Networks).
  • the first domain D1 includes a plurality of time-series images including a first subject obj1 having a relatively large movement in the image and a second subject obj2 having a relatively small movement in the image. It is defined by the image data of the channel.
  • the second domain D2 is defined by image data of a plurality of channels including at least two time-series images including the first subject obj1 and not including the second subject obj2.
  • the time-series image defining the first domain D1 and the second domain D2 respectively captures the reflected light of the light emitted toward the eyeball. It is an image that was made.
  • the light source for example, an infrared LED is used.
  • the first subject obj1 is a bright spot where light is reflected by the cornea in the eyeball (hereinafter, also referred to as a corneal bright spot).
  • the second subject obj2 is a bright spot reflected by the spectacles located overlapping the eyeball (hereinafter, also referred to as a spectacle bright spot).
  • the image of the first domain D1 obtained by imaging the user wearing the spectacles includes both the first subject obj1 and the second subject obj2.
  • the image of the second domain D2 obtained by photographing the user who does not wear the glasses includes the first subject obj1 and does not include the second subject obj2.
  • the generator construction unit 110 uses the cycle GAN to convert the image of the first domain D1 including the subjects obj1 and obj2 into the image of the second domain D2 containing only the subject obj1, that is, the first domain.
  • the subject obj2 is removed from the image of D1 to generate an image in which only the subject obj1 is left.
  • FIG. 2 is a diagram for explaining the input / output relationship of the generator, the images of the first domain D1 and the second domain D2 are exemplified as time-series images corresponding to each other.
  • the images of the first domain D1 and the second domain D2 used as the teacher data at the time of learning may be acquired at different times, and the subject may be a different time-series image.
  • the time interval between images in a time series image does not have to be the same between domains or in a time series within the same domain. That is, the images of the first domain D1 and the second domain D2 used as the teacher data at the time of learning both include the eyeball and the bright spot of the reflected light as the subject, but each image is of the same user.
  • the first image contained in both images not an image of the eyeball (or even the same user at different times with and without glasses).
  • the subject obj1 of the above is also different in position in the image.
  • the corneal bright spot which is the first subject obj1
  • the spectacle bright spot which is the second subject obj2
  • the corneal bright spot has a relatively large movement in the image
  • the spectacle bright spot which is the second subject obj2
  • the movement of the corneal bright spot (first subject obj1) due to the movement of the eyeball is observed.
  • the spectacle bright spot (second subject obj2) hardly moves because the positional relationship between the spectacles and the camera is fixed.
  • the difference in characteristics between the corneal bright spot (first subject obj1) and the spectacle bright spot (second subject obj2) is small, and they are located close to each other. ing. Therefore, in the conventional cycle GAN in which the domain is defined by a single image, the image of the first domain D1 including the subjects obj1 and obj2 is converted into the image of the second domain D2 including only the subject obj1 as described above. It is difficult to generate a model to do.
  • the first domain D1 and the second domain D2 are defined by image data of a plurality of channels including at least two time-series images, respectively.
  • a time-series image means a series of images taken at time intervals of the same subject.
  • each domain is defined by image data of a plurality of channels including at least two time-series images in the cycle GAN, the corneal bright spot (first subject obj1) is difficult to distinguish from the features in a single image. It is possible to generate a model capable of distinguishing the bright spot of the spectacles (second subject obj2).
  • the first domain D1 and the second domain D2 are defined by the image data of four channels input to the data input unit 120, respectively.
  • Image data of the first domain D1 the same object, namely the user's eye, the captured image (ch1) at time t n, the captured image (ch2) at time t n + Delta] t n1, time t including an image (ch3) captured in n + ⁇ t n2, the most recent blinking at time point t n and the image (ch4) of (blink).
  • the image data of the second domain D2 is a captured image (ch1) at time t m, a captured image (ch2) in time t m + ⁇ t m1, captured in time t m + ⁇ t m2 including an image (ch3), the most recent blink of time t m and an image (ch4) of (blink).
  • the image data of the second domain D2 used for learning is the data paired with the image data of the first domain D1, that is, artificially glasses from the same image. It is not necessary for the image to have the bright spot (second subject obj2) removed.
  • the image data defining the first domain D1 (including the corneal bright spot and the spectacle bright spot) and the image data defining the second domain D2 (including the corneal bright spot and including the spectacle bright spot).
  • (Not) may be pre-labeled and acquired separately, or may be automatically sorted according to the number of bright spots detected by image recognition. For example, when an image obtained by capturing the reflected light of light emitted from four light sources toward the eyeball is acquired, image data (4ch) in which the maximum number of bright spots detected in the time-series image does not exceed 4.
  • image data (4ch) in which the maximum number of bright spots exceeds 4 may define the second domain D2.
  • the data input to the data input unit 120 may be, for example, data already stored in the memory of the learning device 100, or is provided from another device via a network or via a recording medium. Data may be used. Further, the data input to the data input unit 120 may include image data acquired in real time using a camera. In the above example, each domain is defined by 4 channels of image data including time series images, but domains may be defined by less or more channels of image data.
  • the blink image is added as an image showing a special state that does not include the first subject (obj1) in the time-series image that defines the domain. By including the blinking image in the time series image, for example, the conversion accuracy is improved when the newly input image data of the first domain D1 includes the blinking image.
  • FIG. 3 is a diagram schematically showing a model generated by using the cycle GAN in one embodiment of the present invention.
  • the method already known as cycle GAN is used, except that the data defining the domain is defined by the image data of a plurality of channels including at least two time series images as described above. can.
  • X first domain D1
  • Y second domain D2
  • Glasses have generator for generating an image without glasses from the image (G) and generator for generating a spectacles with image from without glasses image (F), and the respective discriminator D X, generates the D Y, following loss functions It is defined as the equation (1) of.
  • L cyc (G, F) means a cycle consistency loss
  • Lidency (G, F) is an identity mapping consistency loss.
  • the identity mapping consistency loss does not necessarily have to be introduced, but when the data defining the domain is defined by the image data of a plurality of channels as in the present embodiment, the identity mapping consistency loss is lost. Since the mixing action between channels is reduced by introducing the above, the effect of improving the accuracy of the estimation model is high.
  • Generator that is generated by the cycle GAN (G, F) and discriminator (D X, D Y) may be implemented by each example ResNet (residuals Network) or U-Net.
  • ResNet residuals Network
  • U-Net ResNet
  • FIG. 4 is a diagram showing a schematic configuration of an image generator according to an embodiment of the present invention.
  • the image generator 200 may be a single device or may be implemented by a plurality of distributed devices working together via a network.
  • the image generator 200 is implemented by, for example, a computer having a communication interface, a processor, and a memory, and the processor operates according to a program stored in the memory or received via the communication interface, as described below. It is an information processing device in which the functions of various parts are realized by software.
  • the image generation device 200 includes an image generation unit 210 that generates image data of the second domain D2 from the image data of the first domain D1 by a generator constructed by using the cycle GAN.
  • the first domain D1 includes a first subject obj1 having a relatively large movement in the image and a second subject obj2 having a relatively small movement in the image. It is defined by a plurality of channels of image data including at least two time series images.
  • the second domain D2 is defined by image data of a plurality of channels including at least two time-series images including the first subject obj1 and not including the second subject obj2.
  • the image generator 210 is an image of the first domain D1 input to the data input unit 220 using a generator constructed using the parameter 131 output by the learning device 100.
  • the image data of the second domain D2 is generated from the data (specifically, the image data of 4 channels as illustrated in FIG. 2).
  • the image generation device 200 outputs the image data 231 generated from the data output unit 230.
  • the output image data is a blink image of a part of the image data of the plurality of channels generated by the image generation unit 210, for example, the image data of the four channels as illustrated in FIG. Except for, the image with the latest captured time may be output.
  • the image output from the image generator 200 is used, for example, for line-of-sight estimation.
  • the user's line-of-sight direction can be estimated from the relationship between the center position of the black eye and the position of the corneal bright spot included in the image.
  • the spectacle bright spot is confused with the corneal bright spot. This prevents the accuracy of line-of-sight estimation from deteriorating, and enables highly accurate line-of-sight estimation even for users wearing glasses.
  • an embodiment of the present invention is used to separate a foreground subject (first subject; a person, a car, etc.) from a background subject (second subject). It may be used.
  • a game controller, a smartphone, and various moving objects can be used to acquire information on the surrounding environment.
  • the embodiment of the present invention can be used to estimate the self-position from the positions of surrounding objects, detect flying objects, and take evasive action.
  • an embodiment of the present invention can also be used to identify the movement of a subject such as a person.
  • the embodiment of the present invention is used for a robot to identify a movement of a person with whom communication is to be performed, or an in-vehicle device to detect a specific movement of a driver and output an alarm. ..
  • the first subject is a subject having a relatively large movement in the image and the second subject is a subject having a relatively small movement in the image
  • at least two subjects have been described.
  • the above example is not limited to the two types of subjects that can be recognized as having different movement tendencies depending on the time-series images.
  • the embodiment of the present invention can be applied even when the first subject moves irregularly around a predetermined position and the second subject moves regularly in one direction in the entire image.
  • the learning device 100 and the image generation device 200 are described as separate devices, but these devices may be the same device.
  • the image data of the second domain D2 is generated from the image data of the first domain D1 by using the parameters of the generator constructed in advance, and the image data of the first domain D1 and the image data of the first domain D1 additionally acquired are obtained.
  • the parameters may be updated using the image data of the domain D2 of 2.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Biophysics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Biomedical Technology (AREA)
  • Molecular Biology (AREA)
  • Multimedia (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Ophthalmology & Optometry (AREA)
  • Human Computer Interaction (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)

Abstract

サイクルGANを用いて、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する生成器を構築する生成器構築部を備え、第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、第2のドメインは、第1の被写体を含み、第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理装置が提供される。

Description

情報処理装置、情報処理方法およびプログラム
 本発明は、情報処理装置、情報処理方法およびプログラムに関する。
 教師なし学習の一手法として、GAN(Generative Adversarial Networks)が知られている。さらに、GANの応用として、非特許文献1に記載されたサイクルGANと呼ばれる手法が提案されている。サイクルGANでは、学習データの変換元と変換先とが対応付けられておらず、共通した特徴(ドメイン)をもつ画像群を生成器および識別器にそれぞれ学習させることによって、それぞれの特徴を相互に変換するモデルを構築することができる。
Jun-Yan Zhu, Taesung Park, Phillip Isola, Alexei A. Efros, "Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks," ICCV, pages 2223-2232, 2017
 本発明は、上記のサイクルGANの手法を応用して、動き方の傾向が互いに異なる2種類の被写体を含む画像の解析精度を向上させることが可能な情報処理装置、情報処理方法およびプログラムを提供することを目的とする。
 本発明のある観点によれば、サイクルGANを用いて、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する生成器を構築する生成器構築部を備え、第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、第2のドメインは、第1の被写体を含み、第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理装置が提供される。
 本発明の別の観点によれば、サイクルGANを用いて構築された生成器によって、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する画像生成部を備え、第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、第2のドメインは、第1の被写体を含み、第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理装置が提供される。
 本発明のさらに別の観点によれば、サイクルGANを用いて、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する生成器を構築するステップと、生成器によって、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成するステップとを含み、第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、第2のドメインは、第1の被写体を含み、第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理方法が提供される。
 本発明のさらに別の観点によれば、サイクルGANを用いて、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する生成器を構築する生成器構築部を備え、第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、第2のドメインは、第1の被写体を含み、第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理装置としてコンピュータを機能させるためのプログラムが提供される。
 本発明のさらに別の観点によれば、サイクルGANを用いて構築された生成器によって、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する画像生成部を備え、第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、第2のドメインは、第1の被写体を含み、第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理装置としてコンピュータを機能させるためのプログラムが提供される。
本発明の一実施形態に係る学習装置の概略的な構成を示す図である。 図1の例における各ドメインの画像を例示する図である。 本発明の一実施形態においてサイクルGANを用いて生成されるモデルを概略的に示す図である。 本発明の一実施形態に係る画像生成装置の概略的な構成を示す図である。
 以下、添付図面を参照しながら、本発明のいくつかの実施形態について詳細に説明する。なお、本明細書および図面において、実質的に同一の機能構成を有する構成要素については、同一の符号を付することにより重複説明を省略する。
 図1は、本発明の一実施形態に係る学習装置の概略的な構成を示す図である。図示された例において、学習装置100は、単一の装置であってもよいし、分散した複数の装置がネットワークを介して協働することによって実装されてもよい。学習装置100は、例えば通信インターフェース、プロセッサ、およびメモリを有するコンピュータによって実装され、プロセッサがメモリに格納された、または通信インターフェースを介して受信されたプログラムに従って動作することによって、以下で説明するような各部の機能がソフトウェア的に実現される情報処理装置である。
 学習装置100は、サイクルGAN(Generative Adversarial Networks)を用いて、第1のドメインD1の画像データから第2のドメインD2の画像データを生成する生成器を構築する生成器構築部110を含む。第1のドメインD1は、画像内での動きが相対的に大きい第1の被写体obj1、および画像内での動きが相対的に小さい第2の被写体obj2を含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される。また、第2のドメインD2は、第1の被写体obj1を含み、第2の被写体obj2を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される。
 具体的には、図2に示すように、本実施形態において、第1のドメインD1および第2のドメインD2をそれぞれ定義する時系列画像は、眼球に向かって照射された光の反射光を撮像した画像である。光源としては、例えば赤外線LEDが用いられる。第1の被写体obj1は、光が眼球内の角膜で反射した輝点(以下、角膜輝点ともいう)である。第2の被写体obj2は、眼球に重複して位置する眼鏡で反射した輝点(以下、眼鏡輝点ともいう)である。眼鏡を着用したユーザーを撮像することによって得られた第1のドメインD1の画像は、第1の被写体obj1および第2の被写体obj2の両方を含む。一方、眼鏡を着用していないユーザーを撮像することによって得られた第2のドメインD2の画像は、第1の被写体obj1を含み、第2の被写体obj2を含まない。生成器構築部110は、サイクルGANを用いて、被写体obj1,obj2を含む第1のドメインD1の画像を、被写体obj1のみを含む第2のドメインD2の画像に変換すること、すなわち第1のドメインD1の画像から被写体obj2を除去し被写体obj1のみを残した画像を生成する。
 なお、図2は生成器の入出力関係を説明するための図であるため、互いに対応する時系列画像として第1のドメインD1および第2のドメインD2の画像が例示されている。しかしながら、学習時の教師データとして用いられる第1のドメインD1および第2のドメインD2の画像は、それぞれ異なる時刻に取得され、被写体も異なる時系列画像でありうる。時系列画像の画像間の時間間隔も、ドメイン間、または同じドメイン内の時系列で同じでなくてもよい。つまり、学習時の教師データとして用いられる第1のドメインD1および第2のドメインD2の画像には、いずれも眼球と反射光の輝点が被写体として含まれるが、それぞれの画像は同一のユーザーの眼球を撮像したものではなく(あるいは、同一のユーザーであっても眼鏡を着用しているときおよび着用していないときの異なる時点で撮像されたものであり)、両方の画像に含まれる第1の被写体obj1も、画像内での位置は異なっている。
 上述のように、第1の被写体obj1である角膜輝点は画像内での動きが相対的に大きく、第2の被写体obj2である眼鏡輝点は画像内での動きが相対的に小さい。より具体的には、ヘッドマウントデバイスなどを用いてユーザーの頭部に対して固定されたカメラで画像を撮像した場合、角膜輝点(第1の被写体obj1)については眼球の移動に伴う動きが頻繁に発生するのに対して、眼鏡輝点(第2の被写体obj2)については眼鏡とカメラとの位置関係が固定されているためほとんど動きが発生しない。ただし、ある時点で撮像された単一の画像では、角膜輝点(第1の被写体obj1)と眼鏡輝点(第2の被写体obj2)との特徴の差が小さく、また互いに近接して位置している。従って、単一の画像によってドメインを定義する従来のサイクルGANでは、上記のように被写体obj1,obj2を含む第1のドメインD1の画像を、被写体obj1のみを含む第2のドメインD2の画像に変換するモデルを生成することは困難である。
 そこで、本実施形態では、第1のドメインD1および第2のドメインD2を、それぞれ少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義する。本明細書において、時系列画像は、同一の被写体について時間間隔をおいて撮像された一連の画像を意味する。上記のように、角膜輝点(第1の被写体obj1)については動きが頻繁に発生するため、時系列画像の間で位置が変化している可能性が高い。一方、眼鏡輝点(第2の被写体obj2)についてはほとんど動きが発生しないため、時系列画像の間で位置が変化していない可能性が高い。従って、サイクルGANにおいて各ドメインを少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義すれば、単一の画像における特徴では判別することが困難な角膜輝点(第1の被写体obj1)と眼鏡輝点(第2の被写体obj2)とを判別可能なモデルを生成することができる。
 具体的には、図1および図2に示された例において、第1のドメインD1および第2のドメインD2は、データ入力部120に入力される4チャンネルの画像データによってそれぞれ定義される。第1のドメインD1の画像データは、同一の被写体、すなわちユーザーの眼球について、時刻tに撮像された画像(ch1)と、時刻t+Δtn1に撮像された画像(ch2)と、時刻t+Δtn2に撮像された画像(ch3)と、時刻tの直近のまばたき(blink)の画像(ch4)とを含む。同様に、第2のドメインD2の画像データは、時刻tに撮像された画像(ch1)と、時刻t+Δtm1に撮像された画像(ch2)と、時刻t+Δtm2に撮像された画像(ch3)と、時刻tの直近のまばたき(blink)の画像(ch4)とを含む。なお、サイクルGANについて知られているように、学習に用いられる第2のドメインD2の画像データは、第1のドメインD1の画像データと対になったデータ、すなわち同一の画像から人為的に眼鏡輝点(第2の被写体obj2)を取り除いたような画像である必要はない。
 上記の例において、第1のドメインD1を定義する画像データ(角膜輝点と眼鏡輝点とを含む)と第2のドメインD2を定義する画像データ(角膜輝点を含み、眼鏡輝点を含まない)とは、予めラベル付けされて別個に取得されてもよいし、画像の認識によって検出された輝点の数によって自動的に振り分けられてもよい。例えば、4つの光源から眼球に向かって照射された光の反射光を撮像した画像が取得される場合、時系列画像の中で検出された輝点の最大数が4を超えない画像データ(4ch)が第1のドメインD1を定義し、同じく輝点の最大数が4を超える画像データ(4ch)が第2のドメインD2を定義してもよい。
 なお、データ入力部120に入力されるデータは、例えば学習装置100のメモリに既に格納されているデータであってもよいし、他の装置からネットワークを介して、または記録媒体を介して提供されるデータであってもよい。また、データ入力部120に入力されるデータは、カメラを用いてリアルタイムで取得された画像データを含んでもよい。上記の例では各ドメインが時系列画像を含む4チャンネルの画像データによって定義されるが、より少ない、またはより多いチャンネルの画像データによってドメインが定義されてもよい。なお、まばたき(blink)の画像は、ドメインを定義する時系列画像の中に、第1の被写体(obj1)を含まない特殊な状態を示す画像として追加される。まばたきの画像を時系列画像に含めることによって、例えば新たに入力された第1のドメインD1の画像データがまばたきの画像を含む場合の変換の精度が向上する。
 図3は、本発明の一実施形態においてサイクルGANを用いて生成されるモデルを概略的に示す図である。本実施形態では、上記のようにドメインを定義するデータが少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される点を除いて、サイクルGANとして既に知られている手法を用いることができる。具体的には、サイクルGANでは、実際の眼鏡あり画像(X;第1のドメインD1)と実際の眼鏡なし画像(Y;第2のドメインD2)の2つの画像データ群の間の特徴の関係を学習する。眼鏡あり画像から眼鏡なし画像を生成する生成器(G)および眼鏡なし画像から眼鏡あり画像を生成する生成器(F)、ならびにそれぞれの識別器D,Dを生成し、ロス関数を以下の式(1)のように定義する。なお、Lcyc(G,F)はサイクル一貫性(cycle consistency)ロスを意味し、Lidentity(G,F)は恒等写像一貫性(identity mapping consistency)ロスである。恒等写像一貫性ロスについては必ずしも導入しなくてもよいものであるが、本実施形態のようにドメインを定義するデータが複数チャンネルの画像データによって定義される場合は、恒等写像一貫性ロスを導入することによってチャンネル間の混合作用が軽減されるため、推定モデルの精度が向上する効果が高い。
Figure JPOXMLDOC01-appb-M000001
 上記のサイクルGANで生成される生成器(G,F)および識別器(D,D)は、それぞれ例えばResNet(残差ネットワーク)やU-Netによって実装することができる。ロス関数の値が最小化されるように構築された生成器(G)を用いて、新たに入力された実際の眼鏡あり画像(X;第1のドメインD1)から眼鏡なし画像(Y;第2のドメインD2)を生成することができる。学習装置100は、データ出力部130から構築された生成器のパラメータ131を出力する。
 図4は、本発明の一実施形態に係る画像生成装置の概略的な構成を示す図である。図示された例において、画像生成装置200は、単一の装置であってもよいし、分散した複数の装置がネットワークを介して協働することによって実装されてもよい。画像生成装置200は、例えば通信インターフェース、プロセッサ、およびメモリを有するコンピュータによって実装され、プロセッサがメモリに格納された、または通信インターフェースを介して受信されたプログラムに従って動作することによって、以下で説明するような各部の機能がソフトウェア的に実現される情報処理装置である。
 画像生成装置200は、サイクルGANを用いて構築された生成器によって、第1のドメインD1の画像データから第2のドメインD2の画像データを生成する画像生成部210を含む。学習装置について既に説明したように、第1のドメインD1は、画像内での動きが相対的に大きい第1の被写体obj1、および画像内での動きが相対的に小さい第2の被写体obj2を含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される。また、第2のドメインD2は、第1の被写体obj1を含み、第2の被写体obj2を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される。
 図示された例において、画像生成部210は、上記の学習装置100で出力されたパラメータ131を用いて構築される生成器を用いて、データ入力部220に入力された第1のドメインD1の画像データ(具体的には、図2に例示されたような4チャンネルの画像データ)から、第2のドメインD2の画像データを生成する。これによって、角膜輝点(第1の被写体obj1)と眼鏡輝点(第2の被写体obj2)とを含む第1のドメインD1の画像データから、眼鏡輝点(第2の被写体obj2)のみが除去された画像データを得ることができる。画像生成装置200は、データ出力部230から生成された画像データ231を出力する。出力される画像データは、画像生成部210で生成された複数チャンネルの画像データのうちの一部、例えば、図2に例示されたような4チャンネルの画像データのうち、まばたき(blink)の画像を除き、撮像された時刻が最も遅い画像を出力してもよい。
 画像生成装置200から出力された画像は、例えば視線推定に利用される。具体的には、画像に含まれる黒目の中心位置と角膜輝点の位置との関係からユーザーの視線方向を推定することができる。上記のように、本実施形態では角膜輝点と眼鏡輝点とを含む画像データから眼鏡輝点のみを除去した画像データを生成することができるため、眼鏡輝点が角膜輝点と混同されることによる視線推定の精度の低下を防止し、眼鏡をかけたユーザーに対しても精度の高い視線推定が可能になる。
 なお、上記では角膜輝点と眼鏡輝点とを含む画像データから眼鏡輝点のみを除去した画像データを生成するための生成器の構築、および生成器を用いた画像データの生成の例について説明したが、他の実施形態は、動き方が異なる2種類の被写体を含む画像から一方の種類の被写体を除去する様々な例に適用可能である。また、上記の例では画像内での動きが相対的に小さい被写体を除去した画像を生成したが、逆に、そのような被写体を付加した画像を生成してもよい(図3に示した例における生成器(F)を利用する)。
 具体的には、例えば、周辺環境を撮像した画像において、前景の被写体(第1の被写体;人や車など)を背景の被写体(第2の被写体)から分離するために本発明の実施形態が利用されてもよい。この場合、例えばゲームコントローラ、スマートフォン、各種の移動体(自動車、電気自動車、ハイブリッド電気自動車、自動二輪車、自転車、パーソナルモビリティ、飛行機、ドローン、船舶、ロボットなど)で周辺環境の情報を取得したり、周辺のオブジェクトの位置から自己位置を推定したり、飛来するオブジェクトを検出して回避行動をとったりするために本発明の実施形態を利用することができる。あるいは、例えば人などの被写体の動きを特定するためにも本発明の実施形態を利用することができる。具体的には、例えばロボットがコミュニケーションの相手の人などの動きを特定したり、車載装置が運転者の特定の動きを検出して警報を出力したりするために本発明の実施形態を利用する。
 また、上記では第1の被写体が画像内での動きが相対的に大きい被写体であり、第2の被写体が画像内での動きが相対的に小さい被写体である例について説明したが、少なくとも2つの時系列画像によって動き方の傾向が互いに異なることが認識できる2種類の被写体であれば、上記の例には限られない。例えば、第1の被写体が所定の位置の周辺で不規則に動き、第2の被写体が画像全体を一方向に規則的に動くような場合にも、本発明の実施形態が適用できる。
 また、上記の例では学習装置100と画像生成装置200とが別個の装置として説明されたが、これらの装置は同一の装置であってもよい。例えば、予め構築された生成器のパラメータを用いて第1のドメインD1の画像データから第2のドメインD2の画像データを生成するとともに、追加で取得された第1のドメインD1の画像データおよび第2のドメインD2の画像データを用いてパラメータを更新してもよい。
 以上、添付図面を参照しながら本発明のいくつかの実施形態について詳細に説明したが、本発明はかかる例に限定されない。本発明の属する技術の分野における通常の知識を有する者であれば、請求の範囲に記載された技術的思想の範疇内において、各種の変更例または修正例に想到し得ることは明らかであり、これらについても、当然に本発明の技術的範囲に属するものと了解される。
 100…学習装置、110…生成器構築部、120…データ入力部、130…データ出力部、131…パラメータ、200…画像生成装置、210…画像生成部、220…データ入力部、230…データ出力部、231…画像データ、D1…第1のドメイン、D2…第2のドメイン、obj1…第1の被写体、obj2…第2の被写体。
 

Claims (8)

  1.  サイクルGAN(Generative Adversarial Networks)を用いて、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する生成器を構築する生成器構築部を備え、
     前記第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、
     前記第2のドメインは、前記第1の被写体を含み、前記第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理装置。
  2.  サイクルGAN(Generative Adversarial Networks)を用いて構築された生成器によって、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する画像生成部を備え、
     前記第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、
     前記第2のドメインは、前記第1の被写体を含み、前記第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理装置。
  3.  前記第1の被写体は、画像内での動きが相対的に大きい被写体であり、
     前記第2の被写体は、画像内での動きが相対的に小さい被写体である、請求項1または請求項2に記載の情報処理装置。
  4.  前記時系列画像は、眼球に向かって照射された光の反射光を撮像した画像であり、
     前記第1の被写体は、前記光が前記眼球内の角膜で反射した輝点であり、
     前記第2の被写体は、前記光が前記眼球に重複して位置する眼鏡で反射した輝点であり、
     前記生成器は、前記第1のドメインの画像データから前記第2のドメインの画像データを生成する、請求項3に記載の情報処理装置。
  5.  前記時系列画像は、前記第1の被写体を含まない少なくとも1つの画像を含む、請求項1から請求項4のいずれか1項に記載の情報処理装置。
  6.  サイクルGAN(Generative Adversarial Networks)を用いて、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する生成器を構築するステップと、
     前記生成器によって、前記第1のドメインの画像データおよび前記第2のドメインの画像データの一方から他方を生成するステップと
     を含み、
     前記第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、
     前記第2のドメインは、前記第1の被写体を含み、前記第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理方法。
  7.  サイクルGAN(Generative Adversarial Networks)を用いて、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する生成器を構築する生成器構築部を備え、
     前記第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、
     前記第2のドメインは、前記第1の被写体を含み、前記第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理装置としてコンピュータを機能させるためのプログラム。
  8.  サイクルGAN(Generative Adversarial Networks)を用いて構築された生成器によって、第1のドメインの画像データおよび第2のドメインの画像データの一方から他方を生成する画像生成部を備え、
     前記第1のドメインは、動き方の傾向が互いに異なる第1の被写体と第2の被写体とを含む少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義され、
     前記第2のドメインは、前記第1の被写体を含み、前記第2の被写体を含まない少なくとも2つの時系列画像を含む複数チャンネルの画像データによって定義される情報処理装置としてコンピュータを機能させるためのプログラム。
     
PCT/JP2020/012277 2020-03-19 2020-03-19 情報処理装置、情報処理方法およびプログラム Ceased WO2021186668A1 (ja)

Priority Applications (3)

Application Number Priority Date Filing Date Title
US17/802,699 US20230095977A1 (en) 2020-03-19 2020-03-19 Information processing apparatus, information processing method, and program
JP2022507958A JP7277668B2 (ja) 2020-03-19 2020-03-19 情報処理装置、情報処理方法およびプログラム
PCT/JP2020/012277 WO2021186668A1 (ja) 2020-03-19 2020-03-19 情報処理装置、情報処理方法およびプログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2020/012277 WO2021186668A1 (ja) 2020-03-19 2020-03-19 情報処理装置、情報処理方法およびプログラム

Publications (1)

Publication Number Publication Date
WO2021186668A1 true WO2021186668A1 (ja) 2021-09-23

Family

ID=77771961

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2020/012277 Ceased WO2021186668A1 (ja) 2020-03-19 2020-03-19 情報処理装置、情報処理方法およびプログラム

Country Status (3)

Country Link
US (1) US20230095977A1 (ja)
JP (1) JP7277668B2 (ja)
WO (1) WO2021186668A1 (ja)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019168608A (ja) * 2018-03-23 2019-10-03 カシオ計算機株式会社 学習装置、音響生成装置、方法及びプログラム
JP2019204392A (ja) * 2018-05-25 2019-11-28 ギリア株式会社 学習用データ判別装置および学習用データ判別プログラム

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110170060A1 (en) * 2010-01-08 2011-07-14 Gordon Gary B Gaze Tracking Using Polarized Light
SE543240C2 (en) * 2018-12-21 2020-10-27 Tobii Ab Classification of glints using an eye tracking system

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019168608A (ja) * 2018-03-23 2019-10-03 カシオ計算機株式会社 学習装置、音響生成装置、方法及びプログラム
JP2019204392A (ja) * 2018-05-25 2019-11-28 ギリア株式会社 学習用データ判別装置および学習用データ判別プログラム

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
EDAMOTO, YUSUKE ET AL.: "Scene Identification from Corneal Surface Reflection Images Using Generative Adversarial Networks", IPSJ SIG TECHNICAL REPORT (CVIM), vol. 2019 -CV, no. 12, 30 May 2019 (2019-05-30), pages 1 - 8, ISSN: 2188-8701 *
MITSUZUMI, YU ET AL.: "A Generative Self-Ensemble Approach to Simulated+Unsupervised Learning", IEICE TECHNICAL REPORT., vol. 118, no. 513, 1 March 2019 (2019-03-01), pages 137 - 142, ISSN: 2432-6380 *
NAKAO, MEGUMI ET AL.: "Metal artifact reduction using CycleGAN for CT images", IEICE TECHNICAL REPORT., vol. 119, no. 193, September 2019 (2019-09-01), pages 63 - 68, ISSN: 2432-6380 *

Also Published As

Publication number Publication date
JPWO2021186668A1 (ja) 2021-09-23
JP7277668B2 (ja) 2023-05-19
US20230095977A1 (en) 2023-03-30

Similar Documents

Publication Publication Date Title
Wisiecka et al. Comparison of webcam and remote eye tracking
Asperti et al. Deep learning for head pose estimation: A survey
WO2019149061A1 (en) Gesture-and gaze-based visual data acquisition system
DE102020102230A1 (de) Missbrauchsindex für erklärbare künstliche intelligenz in computerumgebungen
US12438849B2 (en) Face anonymization using a generative adversarial network
Perry et al. Minenet: A dilated cnn for semantic segmentation of eye features
Abate et al. The limitations for expression recognition in computer vision introduced by facial masks
CN113850169B (zh) 一种基于图像分割和生成对抗网络的人脸属性迁移方法
JP7149202B2 (ja) 行動分析装置および行動分析方法
Karanchery et al. Emotion recognition using one-shot learning for human-computer interactions
Chakraborty et al. How can a robot calculate the level of visual focus of human’s attention
KR20230159262A (ko) 스케일 분리를 통한 비디오의 빠른 객체 감지 방법
Ngo et al. Identity unbiased deception detection by 2d-to-3d face reconstruction
Abedi et al. Engagement measurement based on facial landmarks and spatial-temporal graph convolutional networks
Liang et al. Real time hand movement trajectory tracking for enhancing dementia screening in ageing deaf signers of British sign language
Palmero et al. Multi-rate sensor fusion for unconstrained near-eye gaze estimation
JP7277668B2 (ja) 情報処理装置、情報処理方法およびプログラム
Soundararajan et al. Study on Eye Gaze Detection Using Deep Transfer Learning Approaches.
Saleh et al. Real-time attention-augmented spatio-temporal networks for video-based driver activity recognition
Qiao et al. A review of attention detection in online learning
Abhaya et al. Eye-move: An eye gaze typing application with OpenCV and Dlib library
Tomas et al. Determining Student's Engagement in Synchronous Online Classes Using Deep Learning (Compute Vision) and Machine Learning
Chew et al. TGN-PL: Learning to Socialize Using Privileged Information and Temporal Graph Networks
Coimbra et al. Review of trends in automatic human activity recognition in vehicle based in synthetic data
Abawi et al. Hri-Free: Cognitive Robotic Simulation for Evaluating Embodied Social Attention Models

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20925514

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2022507958

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20925514

Country of ref document: EP

Kind code of ref document: A1