WO2020148810A1 - 情報処理装置 - Google Patents
情報処理装置 Download PDFInfo
- Publication number
- WO2020148810A1 WO2020148810A1 PCT/JP2019/000964 JP2019000964W WO2020148810A1 WO 2020148810 A1 WO2020148810 A1 WO 2020148810A1 JP 2019000964 W JP2019000964 W JP 2019000964W WO 2020148810 A1 WO2020148810 A1 WO 2020148810A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- images
- information processing
- target
- learning
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T3/00—Geometric image transformations in the plane of the image
- G06T3/40—Scaling of whole images or parts thereof, e.g. expanding or contracting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T19/00—Manipulating three-dimensional [3D] models or images for computer graphics
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/70—Determining position or orientation of objects or cameras
- G06T7/73—Determining position or orientation of objects or cameras using feature-based methods
- G06T7/74—Determining position or orientation of objects or cameras using feature-based methods involving reference images or patches
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N7/00—Television systems
- H04N7/18—Closed-circuit television [CCTV] systems, i.e. systems in which the video signal is not broadcast
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20081—Training; Learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20084—Artificial neural networks [ANN]
Definitions
- the present invention relates to an information processing device, an information processing method, a program, and learned data for generating an image representing an object.
- the present invention has been made in consideration of the above circumstances, and one of its objects is an information processing apparatus, an information processing method, which can increase variations of reproduced images for reproducing an object with relatively little effort. To provide programs and learned data.
- An information processing apparatus includes a reference image acquisition unit that acquires a plurality of reference images obtained by photographing an object to be reproduced, and enlarges or reduces each of the plurality of reference images.
- Machine learning is performed using a converted image acquisition unit that acquires a plurality of acquired converted images and a plurality of learning target images that include the plurality of converted images as teacher data, and a reproduced image that represents the appearance of the target is generated.
- a machine learning unit for generating learned data used for performing the learning.
- An information processing method includes a step of acquiring a plurality of reference images obtained by photographing an object to be reproduced, and a plurality of images obtained by enlarging or reducing each of the plurality of reference images. And a learning process for performing a machine learning using a plurality of learning target images including the plurality of converted images as teacher data and generating a reproduced image representing the appearance of the target. Generating the completed data.
- a program includes a step of acquiring a plurality of reference images obtained by photographing an object to be reproduced, and a plurality of conversions obtained by enlarging or reducing each of the plurality of reference images.
- Image-acquired data and learned data used to generate a reproduced image representing the appearance of the object by performing machine learning using a plurality of learning object images including the plurality of converted images as teacher data.
- This program may be provided by being stored in a computer-readable non-transitory information storage medium.
- the learned data includes a plurality of learning target images including a plurality of converted images obtained by enlarging or reducing each of a plurality of reference images obtained by photographing an object to be reproduced. It is learned data that is used as teacher data and that is obtained as a result of executing machine learning and is used to generate a reproduced image showing the appearance of the object.
- An information processing apparatus enlarges or reduces each of an acquisition unit that acquires attitude data that represents the attitude of an object to be reproduced and a plurality of reference images that are obtained by photographing the object.
- the target object takes the posture represented by the posture data.
- a reproduction image generation unit that generates a reproduction image representing
- FIG. 1 It is a block diagram of a configuration of an information processing apparatus according to an embodiment of the present invention. It is a functional block diagram of an information processing apparatus. It is a figure which shows an example of a reference
- FIG. 1 is a configuration block diagram of an information processing device 10 according to an embodiment of the present invention.
- the information processing apparatus 10 is an information processing apparatus such as a personal computer, and includes a control unit 11, a storage unit 12, and an interface unit 13, as shown in FIG.
- the information processing device 10 is also connected to the operation device 14, the display device 15, and the camera 16.
- the control unit 11 includes at least one processor such as a CPU, and executes programs stored in the storage unit 12 to execute various types of information processing. A specific example of the process executed by the control unit 11 in this embodiment will be described later.
- the storage unit 12 includes at least one memory device such as a RAM, and stores a program executed by the control unit 11 and data processed by the program.
- the interface unit 13 is an interface for data communication with the operation device 14, the display device 15, and the camera 16.
- the information processing device 10 is connected to each of the operation device 14, the display device 15, and the camera 16 via the interface unit 13 in a wired or wireless manner.
- the interface unit 13 includes a multimedia interface such as HDMI (High-Definition Multimedia Interface) in order to transmit the video data supplied by the information processing device 10 to the display device 15.
- a data communication interface such as a USB (Universal Serial Bus) is included in order to receive an operation signal indicating the operation content of the user accepted by the operation device 14 and a captured image captured by the camera 16.
- the operation device 14 receives an operation instruction from a user and transmits an operation signal indicating the content to the information processing apparatus 10.
- the display device 15 displays a video corresponding to the video signal transmitted from the information processing device 10 to allow the user to browse the video.
- the camera 16 outputs a photographed image obtained by photographing the situation in the field of view. In the present embodiment, it is used to acquire posture data representing the posture of a person when generating a reproduced image to be described later.
- the information processing apparatus 10 is used to perform machine learning with an image representing the appearance of an object (object) to be reproduced as an input and use the result to reproduce the appearance of the object.
- the information processing apparatus 10 functionally includes a reference image acquisition unit 51, a converted image acquisition unit 52, a machine learning unit 53, and a reproduction image generation unit 54.
- These functions are realized by the control unit 11 executing the programs stored in the storage unit 12.
- This program may be provided to the information processing device 10 via a communication network such as the Internet, or may be provided by being stored in a computer-readable information storage medium such as an optical disc.
- the reference image acquisition unit 51 acquires data of the reference image Ir used as teacher data in machine learning described later.
- the reference image Ir is assumed to be a photographed image obtained by actually photographing an object to be reproduced.
- the object is a person, and this person is referred to as an object person T.
- the reference image Ir includes substantially the entire subject T. However, if the part to be reproduced is limited to a part of the target person T, it is sufficient to include only the part to be reproduced. Further, it is desirable that the reference image Ir does not include color information of the background other than the target person T as much as possible. Therefore, when the subject T is photographed and the reference image Ir is generated, the subject T is photographed in front of the screen of the specific color and the background information is extracted from the reference image Ir by a technique called a green back or the like. You may erase it. Alternatively, an image obtained by extracting only a portion in which the target person T is captured from a captured image obtained by actually capturing the target person T may be used as the reference image Ir.
- the reference image Ir is an image in which the portion to be reproduced by the subject T is located substantially in the center. Therefore, at the time of shooting, the height of the camera is adjusted so that the position of the eyes of the subject T substantially coincides with the vertical center of the reference image Ir. Further, the orientation of the camera and the position where the target person T stands are adjusted so that the position of the target person T substantially coincides with the lateral center. Alternatively, the photographed image may be clipped later so that the eye position of the target person T is adjusted to be located at the center of the reference image Ir.
- FIG. 3 shows an example of the reference image Ir that satisfies the requirements described above. It should be noted that here, the position of the human eye is assumed to be the center, but it is also possible to adjust such that a site other than this is located at the center of the reference image Ir.
- the reference image acquisition unit 51 acquires data of a plurality of reference images Ir obtained by photographing one target person T in a plurality of different postures. However, it is assumed that the distance from the camera to the target person T at the time of shooting remains substantially constant without any change in any reference image Ir. Therefore, when the subject T is photographed, the subject T may take various postures in the same place, and it is not necessary to change the position of the camera or the position of the subject T. ..
- the converted image acquisition unit 52 acquires the data by enlarging the reference image Ir by a predetermined magnification to generate the converted image It.
- This converted image It is also used for machine learning as teacher data together with the reference image Ir.
- the converted image acquisition unit 52 is supposed to generate a plurality of converted images It from one reference image Ir by enlarging one reference image Ir at different magnifications.
- the value of magnification means the magnification for the length of the reference image Ir in each of the vertical and horizontal directions.
- the converted image It is assumed to be an image having the same size as the original reference image Ir, which is obtained by enlarging a partial area of the reference image Ir while maintaining the aspect ratio.
- Each converted image It is generated by enlarging only a part of the original reference image Ir so that the image has the same size as the original reference image Ir.
- the converted image It(9) with a magnification of 4 is generated by enlarging the area of the size of 1/4 in the vertical and horizontal directions included in the original reference image Ir.
- the reference image Ir is an image of 1024 pixels in each of the vertical and horizontal directions
- the converted image It(9) having the same vertical and horizontal directions of 1024 pixels as Ir is generated.
- the area to be enlarged is an area centered on the central position in the reference image Ir.
- a region displaced from the center according to the position of the target person T is selected. It may be an enlargement target.
- the plurality of converted images It obtained by the image processing as described above are closer to the photographed image taken by the subject T in a state where the subject T is slightly closer to the camera than when actually photographing. That is, it is possible to generate an image corresponding to a captured image when the distance from the camera to the target person T is variously changed, without actually capturing the image.
- these converted images It include the reference image Ir and the converted image It(9) having the maximum magnification.
- the image has a size close to a captured image when the distance is changed at equal intervals, and includes the target person T.
- FIG. 4 shows an example of nine converted images It generated from the reference image Ir shown in FIG. 3 in this way.
- the machine learning unit 53 performs machine learning by using the converted image It acquired by the converted image acquisition unit 52 as teacher data. This machine learning is assumed to be learning about the correspondence between the posture and the appearance of the subject T.
- the image used by the machine learning unit 53 for machine learning is referred to as a learning target image.
- the learning target image includes at least the converted image It.
- the learning target image may include the reference image Ir acquired by the reference image acquisition unit 51.
- the learning target image includes both the reference image Ir and the converted image It.
- the reference image acquisition unit 51 acquires 100 reference images Ir obtained by photographing the same target person T, and the converted image acquisition unit 52 enlarges each of the 100 reference images Ir by different magnifications. It is assumed that nine converted images It are generated for one reference image Ir. In this case, a total of 1000 images of 100 reference images Ir and 900 converted images It are learning target images.
- the specific contents of the machine learning executed by the machine learning unit 53 will be described below.
- the posture specifying process is a process of specifying the positional relationship between the respective parts forming the human body, and can be realized by a known technique.
- the posture data obtained as a result of the posture specifying process may be configured by a set of coordinate values indicating positions in the image such as joints and feature points forming a human body.
- FIG. 5 shows an example of a skeleton model representing posture data obtained as a result of performing the posture specifying process on the reference image Ir of FIG.
- the machine learning unit 53 independently and independently executes the posture specifying process not only for the reference image Ir but also for each of the plurality of converted images It obtained by enlarging the reference image Ir.
- the target person T may be included in a size different from that of the reference image Ir and other converted images It, and only a part of the body may be included. Therefore, the result of the posture specifying process may differ between the reference image Ir obtained by photographing the target person T having the same posture and each converted image It.
- the processing content of the posture identification processing and the processing target may change depending on how much the appearance of the target person T is to be reproduced in detail in the reproduction processing described later. For example, if it is necessary to reproduce the movements of the fingers of the subject T, it is necessary to identify the postures of the fingers included in the learning target image and use them for machine learning. Further, when it is desired to reproduce the facial expression of the target person T, the positions and states of the eyes and mouths included in the face of the target person T in the learning target image may be specified as a part of the posture data of the target person T. Good.
- the machine learning unit 53 uses the result of the posture identifying process to select the teacher data to be actually used for machine learning.
- This teacher data selection processing is not essential, and all acquired learning target images may be used as machine data for machine learning. However, by excluding learning target images that do not contribute to learning or are assumed to have a low degree of contribution from the teacher data by this selection processing, the time required for machine learning and the amount of calculation are reduced, and processing efficiency is reduced. Can be improved.
- the machine learning unit 53 may exclude a learning target image that does not include a specific part of the subject T from the teacher data.
- This specific part is, for example, the head of the subject T.
- the head of the target person T is displaced from the center of the reference image Ir, and as a result, the central region of the reference image Ir is enlarged to obtain the converted image It.
- the head of the subject T may not be included in the converted image It. Even if such a converted image It is used as an input for machine learning, it is difficult to effectively learn the correspondence relationship between the posture and the appearance of the subject T.
- the learning target image in the state in which the important part is not included in this way is excluded from the teacher data.
- the selection process based on whether or not such a specific part is included may be performed only on the converted image It. Further, if the posture specifying process for the learning target image fails in the first place, or if it is determined that the specified posture has low accuracy, the learning target image may be excluded from the teacher data.
- the machine learning unit 53 selects only some of the learning target images as teacher data. However, other learning target images may be excluded from the teacher data.
- the machine learning unit 53 performs classification processing on the posture data (posture identification result) obtained from each learning target image, and determines that the postures are close to each other.
- the learning target images to be grouped. This classification process may be realized by a known clustering technique such as the k-means method.
- the classification process is not performed on the entire learning target image, but is independently performed on each of the reference image Ir and the group of the converted images It having the same magnification. That is, the machine learning unit 53 teaches a part of the reference images Ir whose posture data is determined to be similar to the other reference images Ir based on the result of the classification process performed on the group including only the reference images Ir. Exclude from data. Further, converted images It having the same magnification belong to a group of converted images It(1) enlarged at a magnification of 4 1/9 times, a group of converted images It(2) enlarged at a magnification of 4 2/9 times, and the like.
- a group of nine converted images It is subjected to a classification process independently, and the converted images It having other similar orientations are excluded from the teacher data.
- the reason why the classification process is executed independently for each magnification is that the size of the subject T included in the learning target image differs depending on the magnification, and the part that contributes to the similarity in posture also changes accordingly. ..
- the reference image Ir including the whole body of the subject T if the positions of the hands and feet are close, it is highly likely that the posture is determined to be close.
- the positions of the eyes and mouth in the reference image Ir do not change significantly, and thus the position information of the part in the face is calculated. Has a relatively small contribution to the similarity of postures.
- the converted image It(9) with the maximum magnification as illustrated in FIG. 4, only a limited area centering on the face of the target person T is included in the image, and hands, feet, etc. are included in the image. No longer included in. Therefore, the difference in the positions of the hands and feet does not contribute to the similarity of the postures, and the positions of the eyes and the mouth included in the face are important in determining the similarity of the postures.
- a certain reference image Ir is determined to be similar in posture to another reference image Ir and excluded from the teacher data, while the reference image Ir is enlarged.
- the obtained converted image It may be used as teacher data.
- the multiple selection processes described above may be executed in combination with each other.
- learning target images to be excluded from the teacher data may be selected according to various determination criteria.
- the machine learning unit 53 inputs the selected learning target image and the posture data set identified based on the learning target image. As, the user learns what kind of image the appearance of the target person T having the posture becomes.
- Such machine learning can be realized by a machine learning algorithm using a neural network or the like. By this machine learning, learned data for realizing an estimator that estimates an image representing the appearance of the subject T from the posture data representing the posture of the subject T is generated.
- the reproduction image generation unit 54 generates an image (reproduction image) that reproduces the appearance of the subject T using the learned data obtained by the machine learning executed by the machine learning unit 53. Specifically, the reproduction image generation unit 54 acquires posture data (skeleton model data) representing a posture that the subject T wants to take, and inputs the posture data to the estimator using the learned data generated by the machine learning unit 53. The estimator estimates what the appearance will be when the subject T takes the posture represented by the skeleton model data, and outputs an image of the estimation result.
- posture data skeleton model data
- the reproduction image generation unit 54 acquires a captured image obtained by capturing the user U with the camera 16.
- the photographing condition of the user U by the camera 16 is close to the photographing condition when the target person T is photographed in order to prepare the reference image Ir. That is, the camera 16 may be the same as or different from the camera that has photographed the target person T, but it is desirable that the resolution and the like be the same.
- the height of the camera 16 at the time of shooting should be close to the height of the camera at the time of shooting the target person T.
- the distance between the camera 16 and the user U does not have to be constant, and the user U may move toward or away from the camera 16.
- the distance range does not exceed the distance between the camera and the target person T when the image capturing for obtaining the reference image Ir is performed.
- the maximum magnification when the converted image It is generated is 4 times in each of the vertical and horizontal directions. Therefore, it is not preferable to be too close to the camera 16 to a distance at which the magnification exceeds this magnification.
- FIG. 6 shows a state in which such a user U is photographed.
- the reproduction image generation unit 54 performs posture identification processing on the image captured by the camera 16 including the user U. This posture identifying process is based on the same algorithm as the posture identifying process executed on the learning target image by the machine learning unit 53 to generate the learning target data. As a result, posture data representing the posture of the user U at the time of shooting is obtained.
- the reproduced image generation unit 54 inputs the posture data of the user U into an estimator that uses the learned data generated by the machine learning unit 53. As a result, a reproduced image showing that the subject T has the same posture as the user U is generated.
- the reproduction image generation unit 54 draws this reproduction image and displays it on the screen of the display device 15.
- the target person T It is possible to generate and display a reproduced image showing a moving state.
- FIG. 7 shows an example of a reproduction image of the target person T generated by the reproduction image generation unit 54 according to the posture of the user U shown in FIG.
- the reproduced image generation unit 54 uses the learned data obtained as a result of machine learning using a plurality of reference images Ir obtained by photographing the target person T having various postures. It is possible to estimate the appearance of the subject T when the subject T takes a posture that is not actually taken, and generate a reproduced image. Furthermore, the reproduced image generation unit 54 uses the learned data obtained by performing the machine learning by including not only the reference image Ir but also the converted image It as a learning target, Even if the user U approaches the camera 16 to a position that is close to, the appearance of the target person T can be reproduced with relatively high accuracy.
- the state of the target person T who takes the same action can be displayed as a moving image. it can. According to such a process, a situation in which the target person T who does not exist at that location performs an action that responds to a situation that actually occurs, such as receptionist reception, is generated in real time as a video, and the viewer Can be presented to.
- the reproduction image generated by the reproduction image generation unit 54 may be an image obtained by converting the area other than the area including the subject T into transparent pixels. Such pixel conversion can be realized by converting pixels of the same color as the background color of the reference image Ir into a transparent color in the generated reproduced image. By thus combining the reproduced image with the transparent background with the given background image, it is possible to generate an image as if the subject T were in another place that is not actually being performed. When such composition is performed, the brightness of the background image is adjusted according to the brightness of the target person T, Gaussian noise is added to the boundary part between the background image and the target person T, or the like. You may perform the process which makes it inconspicuous.
- the target person T based on the image obtained by photographing the target person T in the fixed position, the target person T also includes the movement in the front-back direction. It is possible to generate an image that reproduces how various actions are performed.
- the embodiment of the present invention is not limited to the above description.
- the object to be reproduced is a person, but other than this, various objects such as animals whose postures change may be used.
- only one target person T is set as the reproduction target, but a plurality of reference images Ir are acquired for each of the plurality of target persons, and machine learning is independently performed to generate learned data. If so, it is also possible to make a plurality of target persons to be reproduced.
- the reproduction image generation unit 54 first accepts the selection of the target person to be reproduced from the user, reads the learned data corresponding to the selected target person, and uses it for generating the reproduction image.
- one information processing device 10 performs both the machine learning and the reproduction image generation process, but the present invention is not limited to this, and these processes may be realized by different information processing devices.
- the information processing device that has performed the machine learning provides the learned data obtained as a result to an external information processing device by storing it in a computer-readable information storage medium or distributing it via a communication network.
- the information processing apparatus that receives the provision can realize the function of the reproduction image generation unit 54 as described above by using the provided learned data.
- the converted image is an enlarged image obtained by enlarging the reference image, but the present invention is not limited to this, and a reduced image obtained by reducing the reference image at a given magnification is used as the converted image. It may be used for machine learning. By using such a reduced image for machine learning, even if the user U moves away from the camera 16 to a position farther than the distance from the camera to the target person T when the reference image is captured, the target person can be accurately measured. It is possible to generate a reproduction image that reproduces T.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Software Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- General Engineering & Computer Science (AREA)
- Computing Systems (AREA)
- Medical Informatics (AREA)
- Evolutionary Computation (AREA)
- Data Mining & Analysis (AREA)
- Artificial Intelligence (AREA)
- Mathematical Physics (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Computer Graphics (AREA)
- Computer Hardware Design (AREA)
- Image Analysis (AREA)
- Image Processing (AREA)
- Closed-Circuit Television Systems (AREA)
Abstract
Description
(Emax)1/X
で増加するように決定される。すなわち、n番目の変換画像It(n)の倍率E(n)は、
E(n)=(Emax)n/X
で計算される。
E(n)=4n/9
Claims (11)
- 再現対象となる対象物を撮影して得られる複数の基準画像を取得する基準画像取得部と、
前記複数の基準画像のそれぞれを拡大又は縮小して得られる複数の変換画像を取得する変換画像取得部と、
前記複数の変換画像を含む複数の学習対象画像を教師データとして用いて機械学習を実行し、前記対象物の外観を表す再現画像を生成するために用いられる学習済みデータを生成する機械学習部と、
を含むことを特徴とする情報処理装置。 - 請求項1に記載の情報処理装置において、
前記複数の学習対象画像は、前記複数の基準画像を含む
ことを特徴とする情報処理装置。 - 請求項1又は2に記載の情報処理装置において、
前記機械学習部は、前記複数の学習対象画像のそれぞれに対して、当該学習対象画像に含まれる前記対象物の姿勢を特定する処理を実行し、当該処理の結果得られる姿勢データと、当該学習対象画像とを教師データとして用いて前記機械学習を実行し、前記対象物の姿勢データを前記再現画像に変換するために用いられる学習済みデータを生成する
ことを特徴とする情報処理装置。 - 請求項3に記載の情報処理装置において、
前記機械学習部は、前記姿勢を特定する処理を実行した結果、前記対象物の所定の部位を含まないと判定された学習対象画像を前記教師データから除外する
ことを特徴とする情報処理装置。 - 請求項3又は4に記載の情報処理装置において、
前記機械学習部は、前記複数の学習対象画像に対して、それぞれの学習対象画像から得られた姿勢データが互いに類似するか否か判定する類否判定処理を実行し、得られた姿勢データが他の学習対象画像から得られた姿勢データに類似すると判定された一部の学習対象画像を、前記教師データから除外する
ことを特徴とする情報処理装置。 - 請求項5に記載の情報処理装置において、
前記変換画像取得部は、1枚の前記基準画像を互いに異なる倍率で拡大又は縮小して複数の前記変換画像を生成し、
前記機械学習部は、前記複数の学習対象画像を、それぞれ互いに同じ倍率で拡大又は縮小された複数の変換画像を含む複数のグループに分類し、当該複数のグループのそれぞれに対して独立に前記類否判定処理を実行して前記教師データから除外する学習対象画像を選択する
ことを特徴とする情報処理装置。 - 請求項1から6のいずれか一項に記載の情報処理装置において、
前記変換画像取得部は、1枚の前記基準画像を互いに異なる倍率で拡大又は縮小して複数の前記変換画像を生成し、かつ、前記互いに異なる倍率は、等比数列にしたがって順に変化する倍率である
ことを特徴とする情報処理装置。 - 再現対象となる対象物を撮影して得られる複数の基準画像を取得するステップと、
前記複数の基準画像のそれぞれを拡大又は縮小して得られる複数の変換画像を取得するステップと、
前記複数の変換画像を含む複数の学習対象画像を教師データとして用いて機械学習を実行し、前記対象物の外観を表す再現画像を生成するために用いられる学習済みデータを生成するステップと、
を含むことを特徴とする情報処理方法。 - 再現対象となる対象物を撮影して得られる複数の基準画像を取得するステップと、
前記複数の基準画像のそれぞれを拡大又は縮小して得られる複数の変換画像を取得するステップと、
前記複数の変換画像を含む複数の学習対象画像を教師データとして用いて機械学習を実行し、前記対象物の外観を表す再現画像を生成するために用いられる学習済みデータを生成するステップと、
をコンピュータに実行させるためのプログラム。 - 再現対象となる対象物を撮影して得られる複数の基準画像のそれぞれを拡大又は縮小して得られる複数の変換画像を含む複数の学習対象画像を教師データとして用いて機械学習を実行した結果得られる、前記対象物の外観を表す再現画像を生成するために用いられる学習済みデータ。
- 再現対象となる対象物の姿勢を表す姿勢データを取得する取得部と、
対象物を撮影して得られる複数の基準画像のそれぞれを拡大又は縮小して得られる複数の変換画像を含む複数の学習対象画像を教師データとして用いて機械学習を実行した結果得られる学習済みデータを用いて、前記姿勢データが表す姿勢を前記対象物が取っている様子を表す再現画像を生成する再現画像生成部と、
を含むことを特徴とする情報処理装置。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2019/000964 WO2020148810A1 (ja) | 2019-01-15 | 2019-01-15 | 情報処理装置 |
| US17/421,266 US12423770B2 (en) | 2019-01-15 | 2019-01-15 | Information processing apparatus |
| JP2020566361A JP7134260B2 (ja) | 2019-01-15 | 2019-01-15 | 情報処理装置 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2019/000964 WO2020148810A1 (ja) | 2019-01-15 | 2019-01-15 | 情報処理装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020148810A1 true WO2020148810A1 (ja) | 2020-07-23 |
Family
ID=71613568
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/000964 Ceased WO2020148810A1 (ja) | 2019-01-15 | 2019-01-15 | 情報処理装置 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US12423770B2 (ja) |
| JP (1) | JP7134260B2 (ja) |
| WO (1) | WO2020148810A1 (ja) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2022168442A1 (ja) * | 2021-02-05 | 2022-08-11 | ソニーグループ株式会社 | 推論装置、推論方法およびプログラム |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2012088787A (ja) * | 2010-10-15 | 2012-05-10 | Canon Inc | 画像処理装置、画像処理方法 |
| JP2018116599A (ja) * | 2017-01-20 | 2018-07-26 | キヤノン株式会社 | 情報処理装置、情報処理方法およびプログラム |
| JP2018129007A (ja) * | 2017-02-10 | 2018-08-16 | 日本電信電話株式会社 | 学習データ生成装置、学習装置、推定装置、学習データ生成方法及びコンピュータプログラム |
| JP2018136632A (ja) * | 2017-02-20 | 2018-08-30 | オムロン株式会社 | 形状推定装置 |
| JP2018169672A (ja) * | 2017-03-29 | 2018-11-01 | 三菱電機インフォメーションシステムズ株式会社 | 教師画像を生成する方法、コンピュータおよびプログラム |
Family Cites Families (49)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6807286B1 (en) * | 2000-04-13 | 2004-10-19 | Microsoft Corporation | Object recognition using binary image quantization and hough kernels |
| GB2395264A (en) * | 2002-11-29 | 2004-05-19 | Sony Uk Ltd | Face detection in images |
| US8504546B2 (en) * | 2006-11-29 | 2013-08-06 | D&S Consultants, Inc. | Method and system for searching multimedia content |
| JP5553141B2 (ja) * | 2009-11-11 | 2014-07-16 | ソニー株式会社 | 画像処理システム、画像処理装置、画像処理方法、およびプログラム |
| KR101327321B1 (ko) * | 2009-12-11 | 2013-11-11 | 한국전자통신연구원 | 모션 캡쳐 데이터의 포즈 검색 장치 및 방법 |
| JP5881136B2 (ja) * | 2010-09-27 | 2016-03-09 | ソニー株式会社 | 情報処理装置及び方法、並びにプログラム |
| US8903167B2 (en) * | 2011-05-12 | 2014-12-02 | Microsoft Corporation | Synthesizing training samples for object recognition |
| WO2013025220A1 (en) * | 2011-08-18 | 2013-02-21 | Nikon Corporation | Image sharpness classification system |
| US9043247B1 (en) * | 2012-02-25 | 2015-05-26 | Symantec Corporation | Systems and methods for classifying documents for data loss prevention |
| US8737767B2 (en) * | 2012-02-28 | 2014-05-27 | Disney Enterprises, Inc. | Perceptually guided capture and stylization of 3D human figures |
| US9443355B2 (en) * | 2013-06-28 | 2016-09-13 | Microsoft Technology Licensing, Llc | Reprojection OLED display for augmented reality experiences |
| KR102013705B1 (ko) * | 2013-08-16 | 2019-08-23 | 한국전자통신연구원 | 승마시뮬레이터에서 사용자 자세 인식 장치 및 방법 |
| WO2015038144A1 (en) * | 2013-09-13 | 2015-03-19 | Intel Corporation | Processing of images of a subject individual |
| US20150243031A1 (en) * | 2014-02-21 | 2015-08-27 | Metaio Gmbh | Method and device for determining at least one object feature of an object comprised in an image |
| JP6369534B2 (ja) * | 2014-03-05 | 2018-08-08 | コニカミノルタ株式会社 | 画像処理装置、画像処理方法、および、画像処理プログラム |
| GB2550806B (en) * | 2015-04-01 | 2021-01-20 | Landmark Graphics Corp | Model generation for real-time rate of penetration prediction |
| US9767386B2 (en) * | 2015-06-23 | 2017-09-19 | Adobe Systems Incorporated | Training a classifier algorithm used for automatically generating tags to be applied to images |
| US10242293B2 (en) * | 2015-08-04 | 2019-03-26 | The Asan Foundation | Method and program for computing bone age by deep neural network |
| WO2017100929A1 (en) * | 2015-12-15 | 2017-06-22 | Applied Recognition Inc. | Systems and methods for authentication using digital signature with biometrics |
| US10026176B2 (en) * | 2016-03-08 | 2018-07-17 | Amazon Technologies, Inc. | Browsing interface for item counterparts having different scales and lengths |
| US10089717B2 (en) * | 2016-04-05 | 2018-10-02 | Flipboard, Inc. | Image scaling using a convolutional neural network |
| US10964078B2 (en) * | 2016-08-10 | 2021-03-30 | Zeekit Online Shopping Ltd. | System, device, and method of virtual dressing utilizing image processing, machine learning, and computer vision |
| US10402689B1 (en) * | 2017-04-04 | 2019-09-03 | Snap Inc. | Generating an image mask using machine learning |
| US10430966B2 (en) * | 2017-04-05 | 2019-10-01 | Intel Corporation | Estimating multi-person poses using greedy part assignment |
| US10410322B2 (en) * | 2017-04-05 | 2019-09-10 | Here Global B.V. | Deep convolutional image up-sampling |
| KR101986327B1 (ko) * | 2017-08-18 | 2019-06-05 | 강다겸 | 자세 가이드 제공 방법 및 그 장치 |
| WO2019050916A1 (en) * | 2017-09-11 | 2019-03-14 | Track160, Ltd. | TECHNIQUES FOR RENDERING THREE-DIMENSIONAL ANIMATED GRAPHICS FROM A VIDEO |
| CN107680088A (zh) * | 2017-09-30 | 2018-02-09 | 百度在线网络技术(北京)有限公司 | 用于分析医学影像的方法和装置 |
| US12377345B2 (en) * | 2017-11-02 | 2025-08-05 | Michael Callahan | Methods and systems for interactive gaming platform scene generation utilizing captured visual data and artificial intelligence-generated environment |
| JP6831769B2 (ja) * | 2017-11-13 | 2021-02-17 | 株式会社日立製作所 | 画像検索装置、画像検索方法、及び、それに用いる設定画面 |
| KR102438201B1 (ko) * | 2017-12-01 | 2022-08-30 | 삼성전자주식회사 | 사진 촬영과 관련된 추천 정보를 제공하는 방법 및 시스템 |
| KR102466942B1 (ko) * | 2017-12-27 | 2022-11-14 | 한국전자통신연구원 | 얼굴 인식을 위한 얼굴 자세 등록 장치 및 방법 |
| GB2569817B (en) * | 2017-12-29 | 2021-06-23 | Snugs Tech Ltd | Ear insert shape determination |
| US10943096B2 (en) * | 2017-12-31 | 2021-03-09 | Altumview Systems Inc. | High-quality training data preparation for high-performance face recognition systems |
| US10296102B1 (en) * | 2018-01-31 | 2019-05-21 | Piccolo Labs Inc. | Gesture and motion recognition using skeleton tracking |
| US11051571B2 (en) * | 2018-02-27 | 2021-07-06 | Levi Strauss & Co. | Apparel design system including garment features with allocation insights |
| US20190272620A1 (en) * | 2018-03-05 | 2019-09-05 | Jared Siegwarth Seip | System and method for image upscaling |
| US10540757B1 (en) * | 2018-03-12 | 2020-01-21 | Amazon Technologies, Inc. | Method and system for generating combined images utilizing image processing of multiple images |
| TWI679557B (zh) * | 2018-05-14 | 2019-12-11 | 國立臺灣師範大學 | 適應性運動姿態感測系統與方法 |
| US20200094401A1 (en) * | 2018-09-21 | 2020-03-26 | Beijing Jingdong Shangke Information Technology Co., Ltd. | System and method for automatic learning of product manipulation |
| CN109556596A (zh) * | 2018-10-19 | 2019-04-02 | 北京极智嘉科技有限公司 | 基于地面纹理图像的导航方法、装置、设备及存储介质 |
| CN111161274B (zh) * | 2018-11-08 | 2023-07-07 | 上海市第六人民医院 | 腹部图像分割方法、计算机设备 |
| US10489683B1 (en) * | 2018-12-17 | 2019-11-26 | Bodygram, Inc. | Methods and systems for automatic generation of massive training data sets from 3D models for training deep learning networks |
| US10839586B1 (en) * | 2019-06-07 | 2020-11-17 | Snap Inc. | Single image-based real-time body animation |
| CN113438410B (zh) * | 2021-05-19 | 2022-06-17 | 荣耀终端有限公司 | 一种高倍率拍摄方法及电子设备 |
| US20220392099A1 (en) * | 2021-05-28 | 2022-12-08 | Disney Enterprises, Inc. | Stable pose estimation with analysis by synthesis |
| US11875496B2 (en) * | 2021-08-25 | 2024-01-16 | Genpact Luxembourg S.à r.l. II | Dimension estimation using duplicate instance identification in a multiview and multiscale system |
| US12299930B2 (en) * | 2021-11-05 | 2025-05-13 | Image Metrics, Ltd. | Real-time body pose estimation system and method in unconstrained video |
| JP2024029417A (ja) * | 2022-08-22 | 2024-03-06 | 日本電気株式会社 | 画像処理装置、画像処理方法、およびプログラム |
-
2019
- 2019-01-15 US US17/421,266 patent/US12423770B2/en active Active
- 2019-01-15 JP JP2020566361A patent/JP7134260B2/ja active Active
- 2019-01-15 WO PCT/JP2019/000964 patent/WO2020148810A1/ja not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2012088787A (ja) * | 2010-10-15 | 2012-05-10 | Canon Inc | 画像処理装置、画像処理方法 |
| JP2018116599A (ja) * | 2017-01-20 | 2018-07-26 | キヤノン株式会社 | 情報処理装置、情報処理方法およびプログラム |
| JP2018129007A (ja) * | 2017-02-10 | 2018-08-16 | 日本電信電話株式会社 | 学習データ生成装置、学習装置、推定装置、学習データ生成方法及びコンピュータプログラム |
| JP2018136632A (ja) * | 2017-02-20 | 2018-08-30 | オムロン株式会社 | 形状推定装置 |
| JP2018169672A (ja) * | 2017-03-29 | 2018-11-01 | 三菱電機インフォメーションシステムズ株式会社 | 教師画像を生成する方法、コンピュータおよびプログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| JP7134260B2 (ja) | 2022-09-09 |
| US20220101491A1 (en) | 2022-03-31 |
| US12423770B2 (en) | 2025-09-23 |
| JPWO2020148810A1 (ja) | 2021-11-18 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN110650368B (zh) | 视频处理方法、装置和电子设备 | |
| CN110363133B (zh) | 一种视线检测和视频处理的方法、装置、设备和存储介质 | |
| JP4799105B2 (ja) | 情報処理装置及びその制御方法、コンピュータプログラム、記憶媒体 | |
| CN114219878A (zh) | 虚拟角色的动画生成方法及装置、存储介质、终端 | |
| JP7071054B2 (ja) | 情報処理装置、情報処理方法およびプログラム | |
| JP4799104B2 (ja) | 情報処理装置及びその制御方法、コンピュータプログラム、記憶媒体 | |
| CN114333046B (zh) | 舞蹈动作评分方法、装置、设备和存储介质 | |
| US11158104B1 (en) | Systems and methods for building a pseudo-muscle topology of a live actor in computer animation | |
| CN112911393B (zh) | 部位识别方法、装置、终端及存储介质 | |
| CN108198130A (zh) | 图像处理方法、装置、存储介质及电子设备 | |
| JP7790522B2 (ja) | プログラム、追跡方法、および追跡装置 | |
| JP2013200735A (ja) | 画像生成装置、画像生成方法及びプログラム | |
| CN118230394A (zh) | 一种基于人脸3d关键点的面部表情驱动方法及装置 | |
| JP2010250419A (ja) | 目の様態を検出する画像処理装置 | |
| WO2020217425A1 (ja) | 教師データ生成装置 | |
| JP7385416B2 (ja) | 画像処理装置、画像処理システム、画像処理方法及び画像処理プログラム | |
| JP7134260B2 (ja) | 情報処理装置 | |
| CN110321009B (zh) | Ar表情处理方法、装置、设备和存储介质 | |
| US8223220B2 (en) | Image processing device, image processing method, and information storage medium | |
| JP2017102748A (ja) | 瞳画像学習装置、瞳位置検出装置及びそのプログラム | |
| JP2016062184A (ja) | 動画像生成システム、動画像生成装置、動画像生成方法、及びコンピュータプログラム | |
| KR102808311B1 (ko) | 영상 처리 방법 및 장치 | |
| EP4411674A1 (en) | Seated posture detection | |
| CN121000814A (zh) | 一种手机背盖显示装置的控制方法、系统及手机 | |
| JP2019139608A (ja) | 画像生成装置及び画像生成プログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19910189 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2020566361 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19910189 Country of ref document: EP Kind code of ref document: A1 |
|
| WWG | Wipo information: grant in national office |
Ref document number: 17421266 Country of ref document: US |