WO2020171237A1 - 画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラム - Google Patents

画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラム Download PDF

Info

Publication number
WO2020171237A1
WO2020171237A1 PCT/JP2020/007392 JP2020007392W WO2020171237A1 WO 2020171237 A1 WO2020171237 A1 WO 2020171237A1 JP 2020007392 W JP2020007392 W JP 2020007392W WO 2020171237 A1 WO2020171237 A1 WO 2020171237A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
posture
try
clothing
user
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2020/007392
Other languages
English (en)
French (fr)
Inventor
五十嵐 健夫
信行 梅谷
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of Tokyo NUC
Original Assignee
University of Tokyo NUC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of Tokyo NUC filed Critical University of Tokyo NUC
Priority to JP2021502253A priority Critical patent/JP7497059B2/ja
Publication of WO2020171237A1 publication Critical patent/WO2020171237A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00Two-dimensional [2D] image generation
    • G06T11/60Creating or editing images; Combining images with text

Definitions

  • the present invention relates to an image processing device, a learned model, an image collecting device, an image processing method, and an image processing program.
  • the present invention provides an image processing device, a learned model, an image collecting device, an image processing method, and an image processing program that can reduce the calculation load when trying on a virtual wearing product.
  • An image processing apparatus uses, as input data, an image of a reference wear product worn by a wearer in a predetermined posture, and an image of a try-on wear product worn by a wearer in a posture common to the predetermined posture.
  • the image of the user wearing the reference fitting is input to the learned model learned by machine learning using the learning data associated with the output as the output data, and the image of the user trying on the fitting is displayed.
  • An image generation unit for generating is provided.
  • the learned model is learned using the learning data in which the image of the reference wear product and the image of the try-on wear product are associated with each other, and the image of the user wearing the reference wear product is input to this learning model.
  • the image generation unit extracts the image of the reference wear product and the image of the body part of the user from the image of the user wearing the reference wear product, and extracts the image of the extracted reference wear product from the image of the reference wear product.
  • An image of the user who tried on the fitting-on article by combining the image of the fitting-on article output from the learned model and the image of the extracted body part of the user by inputting to the learned model It may be generated.
  • the image generation unit extracts depth information of an image of the reference worn product and depth information of an image of the body part of the user from the image of the user wearing the reference worn product, and the extracted depth.
  • the image of the fitting product and the image of the body part of the user may be combined based on the information.
  • the learning data is a first variation in which the wearer wears the reference wearable image worn by the wearer in a predetermined posture as the input data, and the wearer wears it in a posture common to the predetermined posture.
  • An image of the fitting-on product is associated as the output data
  • an image of the fitting-on product of a second variation which the wearing subject wears in a posture common to the predetermined posture, is output to the input data.
  • the image generation unit accepts selection of a variation of the try-on fitting product, inputs the image of the user wearing the reference wear product to the learned model, and attaches the try-on fitting of the selected variation.
  • An image of the user who tried on the item may be generated.
  • the worn article is clothing
  • the learning data uses the image of the reference clothing worn by the wearer in a predetermined posture as the input data, and wears it in a posture common to the predetermined posture. It may include data in which the image of the try-on garment worn by is associated as the output data.
  • the calculation load in the virtual clothing fitting is reduced. be able to.
  • An image acquisition apparatus is a robot mannequin, a camera, a robot control unit that controls the posture of the robot mannequin, and the robot mannequin to which a wearing article is attached based on the control of the robot control unit.
  • An image capturing control unit that captures an image of the attached article with the camera in a state in which the orientation is controlled to a predetermined orientation and stores the captured image in a storage unit in association with the orientation.
  • the image of the reference wear product and the image of the try-on wear product can be accurately associated with each other, and the reliability of the learning data used for learning the learned model can be increased.
  • An image processing method uses, as input data, an image of a reference wear product worn by a wearer in a predetermined posture, and an image of a try-on wear product worn by a wearer in a posture common to the predetermined posture.
  • the image of the user wearing the reference fitting is input to the learned model learned by machine learning using the learning data associated with the output as the output data, and the image of the user trying on the fitting is displayed.
  • An image generation step of generating is included.
  • An image processing program uses, as input data, an image of a reference wear product worn by a wearer in a predetermined posture in a computer, and try-on wear worn by a wearer in a posture common to the predetermined posture.
  • An image of a user wearing the reference wear product is input to a learned model learned by machine learning using learning data in which an image of a product is associated as output data, and a user trying on the try-on wear product.
  • the process for generating the image is executed.
  • FIG. 1 is a block diagram showing a schematic configuration of a first embodiment of an image processing device.
  • FIG. 6 is a diagram for explaining an example of an image generation process.
  • the flowchart which shows the processing content of an image generation process.
  • FIG. 6 is a diagram for explaining an example of an image generation process.
  • FIG. 6 is a diagram for explaining an example of an image generation process. It is a figure which shows an example of the hardware constitutions of an image processing apparatus.
  • the image collection device 1 includes, for example, a camera 10, a robot mannequin 20, a display 30, and an image processing device 100. These devices and devices are connected to each other by a communication line, a wireless communication network, or the like.
  • the configuration shown in FIG. 1 is merely an example, and a part of the configuration may be omitted, or another configuration may be added.
  • the camera 10 is, for example, a digital camera using a solid-state image sensor such as CCD (Charge Coupled Device) or CMOS (Complementary Metal Oxide Semiconductor).
  • the camera 10 is, for example, fixedly arranged in front of the robot mannequin 20, and captures an image of the robot mannequin 20 on which clothes, which is an example of a wearing article, is attached.
  • the camera 10 may be, for example, a stereo camera or a depth camera.
  • the robot mannequin 20 is, for example, a humanoid robot.
  • the robot mannequin 20 is configured so that it can be driven with, for example, four degrees of freedom for both shoulders, four degrees of freedom for both elbows, and two degrees of freedom for rotation of the waist, and can take various postures while wearing clothes. It is possible.
  • the display 30 is, for example, a liquid crystal display, and displays an image generated by the image processing device 100.
  • the image processing apparatus 100 includes, for example, a control unit 110 and a storage unit 120.
  • the control unit 110 is realized, for example, by a hardware processor such as a CPU (Central Processing Unit) executing a program (software).
  • a hardware processor such as a CPU (Central Processing Unit) executing a program (software).
  • some or all of these components are hardware (circuits) such as LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), GPU (Graphics Processing Unit), etc. Part: including circuitry), or may be realized by cooperation of software and hardware.
  • the program may be stored in advance in a storage device such as a HDD or a flash memory of the image processing apparatus 100, or in a removable storage medium such as a DVD or a CD-ROM, and the storage medium is used as a drive device. It may be installed in the HDD or flash memory of the image processing apparatus 100 by being attached.
  • the control unit 110 includes, for example, a robot control unit 112, a shooting control unit 114, a learning unit 116, and an image generation unit 118.
  • the robot controller 112 controls the operation of the robot mannequin 20.
  • the robot control unit 112 controls the posture of the robot mannequin 20, for example, by controlling the amount of movement of the robot mannequin 20 in each degree of freedom.
  • Each degree of freedom of the robot mannequin 20 is centered on a central axis extending in the vertical direction in addition to the degree of freedom of pose of the robot mannequin 20, such as the degree of freedom of both shoulders, the degree of freedom of both elbows, and the degree of rotation of the waist.
  • the degree of freedom of the rotational position of the robot manikin 20 is included.
  • the degree of freedom of the robot mannequin 20 may include, for example, the angle of the robot mannequin 20 with respect to the shooting position/direction of the camera 10.
  • the above description is merely one example of the posture of the robot manikin 20, and the present invention is not limited to this.
  • the image capturing control unit 114 controls the image capturing operation of the camera 10.
  • the imaging control unit 114 captures an image of the robot mannequin 20 in a state where the posture of the robot mannequin 20 is controlled to a predetermined posture based on the control of the robot control unit 112, for example.
  • the imaging control unit 114 captures images of the robot mannequin 20 while changing the rotational position of the robot mannequin 20 for various poses of the robot mannequin 20 under the control of the robot control unit 112, for example.
  • the imaging control unit 114 stores the captured image in the storage unit 120 in association with the posture of the robot mannequin 20.
  • the image capturing control unit 114 captures, for example, an image of the robot mannequin 20 wearing measurement clothing and an image of the robot mannequin 20 wearing try-on clothing.
  • the image capturing control unit 114 divides the image of the measurement clothing from the image of the robot mannequin 20 on which the measurement clothing is worn, and associates the divided image with the posture of the robot mannequin 20 at the time of capturing the image to obtain image data 122. It is stored in the storage unit 120 as.
  • the imaging control unit 114 divides the image of the try-on garment into regions from the image of the robot mannequin 20 wearing the try-on garment, and associates the region-divided image with the posture of the robot mannequin 20 at the time of capturing the image.
  • the data 122 is stored in the storage unit 120.
  • FIG. 2 is a diagram showing an example of the data content of the image data 122.
  • the data attributes of the image data 122 include, for example, the type of image, image information, and the posture of the robot mannequin 20.
  • the type of image is the type of clothing worn by the robot mannequin 20, and includes, for example, measurement clothing and try-on clothing.
  • the measurement garment is an example of a reference garment referred to when synthesizing the try-on garment.
  • the measurement garment is, for example, a garment for acquiring information about the body of the user, and it is preferable that the measurement garment be easily distinguished from the skin color of the human body and easily associated with various try-on garments.
  • the try-on garment is a garment to be virtually tried on, and includes, for example, a plurality of types of garments having different designs.
  • the image information includes information about the brightness value (R, G, B) of each image position of the image captured by the camera 10.
  • the image information may further include depth information for each image position of the image captured by the camera 10.
  • the posture of the robot manikin 20 is a parameter defined by a combination of motion amounts of the robot manikin 20 in each degree of freedom.
  • the learning unit 116 stores, in the storage unit 120, learning data 124 in which an image of measurement clothing is used as input data and an image of measurement clothing and an image of try-on clothing in which the posture of the robot mannequin 20 is common are associated as output data. To do. That is, the learning unit 116 associates the image of the measurement clothing and the image of the try-on clothing captured in the same pose from the same direction and stores them in the storage unit 120 as the learning data 124.
  • the learning unit 116 refers to the image data 122 stored in the storage unit 120, for example, and stores the learning data 124 in the storage unit 120 in association with the image of the measurement clothing and the image of the try-on clothing.
  • the learning unit 116 learns a learning model by machine learning using the learning data 124 and generates a learned model 126.
  • the learned model 126 is composed of, for example, a neural network which is a type of machine learning model.
  • the learning unit 116 increases the number of data of the learning data 124 by adding processes such as parallel movement, enlargement/reduction, rotation, and noise addition to the image of the measurement clothing and the image of the try-on clothing. Data augmentation may be performed.
  • FIG. 3 is a diagram showing an example of data contents of the learning data 124.
  • the learning data 124 includes a plurality (N in the illustrated example) of image information “A1” to “AN” in which the postures of the robot manikin 20 are different from each other as the image information of the measurement clothing. There is. Further, the learning data 124 is associated with the image information “A1” to “AN” of the measurement clothes and the image information of the try-on clothes in which the posture of the robot mannequin 20 is common.
  • the learning data 124 is for each clothing (“Clothing 1”, “Clothing 2”, etc.) in which the image information “A1” to “AN” of the measurement clothing and the robot mannequin 20 have the same posture. It is associated with image information.
  • FIG. 4 is a diagram showing an example of the correspondence between the image of the measurement clothing and the image of the try-on clothing.
  • the image information of the measurement clothing and the image information of the try-on clothing are associated with each other for each posture of the robot mannequin 20.
  • the learning data 124 corresponds to, for example, the image information “A1” of the measurement clothing and the image information “B11” of the trial clothing when the posture of the robot mannequin 20 is “posture 1”.
  • the image information “A2” of the measurement clothing when the posture of the robot mannequin 20 is “posture 2” and the image information “B12” of the try-on clothing are associated with each other.
  • the image generation unit 118 inputs the image of the user wearing the measurement clothing as the learned model 126 and generates the image of the user who tried on the try-on clothing.
  • the image generation unit 118 extracts, for example, the image of the measurement clothing and the image of the body part of the user from the image of the user wearing the measurement clothing, and inputs the extracted image of the measurement clothing to the learned model 126.
  • the image of the try-on garment output from the learned model 126 and the extracted image of the body part of the user are combined.
  • FIG. 5 is a diagram for explaining an example of image generation processing by the image processing apparatus 100.
  • the image processing apparatus 100 first divides the image of the measurement clothing from the image of the robot mannequin 20 wearing the measurement clothing and divides the area into the robot wearing the try-on clothing.
  • the image of the try-on garment is divided into regions from the image of the mannequin 20.
  • the image processing apparatus 100 associates the image of the measurement clothing and the image of the try-on clothing, which are divided into regions, as learning data 124, and causes the learning model to be learned by machine learning using the learning data 124. As a result, the learned model 126 is generated.
  • the image processing apparatus 100 divides the image of the measurement clothing and the image of the user's body part into regions based on the image of the user wearing the measurement clothing.
  • the image processing apparatus 100 may acquire the depth information of the image of the measurement clothing and the depth information of the image of the body part of the user. Thereby, the area of the image of the measurement clothing and the image of the body part of the user are accurately divided.
  • the image processing apparatus 100 inputs the image of the measurement clothing, which has been divided into regions, to the learned model 126. As a result, the image of the try-on clothing corresponding to the posture of the user is output from the learned model 126.
  • the image processing apparatus 100 puts on the try-on garment by synthesizing the image of the try-on garment output from the learned model 126 and the image of the user's body part divided into regions as described above. Generate a user image.
  • the image processing apparatus 100 synthesizes the image of the user wearing the try-on garment so that, for example, the image of the body part of the user is located in front of the image of the measurement garment.
  • the image processing apparatus 100 may generate the image of the user wearing the try-on clothing, for example, based on the depth information of the image of the measurement clothing and the depth information of the image of the body part of the user.
  • the image processing apparatus 100 may, for example, combine the user's face or neck with the neck of the clothing, combine the user's arm with the sleeve of the clothing, or the like. It is possible to accurately combine the images.
  • FIG. 6 is a flowchart showing an example of the learning process of the learned model 126. Prior to the processing of the flowchart shown in FIG. 6, it is assumed that the image of the measurement clothing for each posture of the robot manikin 20 has been captured. Further, the process of the flowchart shown in FIG. 6 is executed by a predetermined operation as a trigger, for example, when the try-on garment is attached to the robot mannequin 20.
  • the robot control unit 112 controls the robot mannequin 20 wearing the trial clothing to a predetermined posture (step S10).
  • the image capturing control unit 114 captures an image of the robot mannequin 20 using the camera 10 (step S12).
  • the imaging control unit 114 divides the image of the try-on clothing from the image of the robot mannequin 20 into regions (step S14).
  • the imaging control unit 114 stores the image of the measurement clothing and the image of the trial clothing in which the posture of the robot mannequin 20 is common in association with each other in the storage unit 120 as learning data 124 (step S16).
  • the learning unit 116 uses, in addition to the learning data 124 stored in the storage unit 120, the data padded by performing data augmentation on the learning data 124, by using the learning model by machine learning. Is learned (step S18).
  • the imaging control unit 114 determines whether the learning is finished (step S20). Then, when the imaging control unit 114 determines that the learning has not ended, the robot control unit 112 changes the posture of the robot mannequin 20 (step S22), and changes the posture of the trial clothing to the image of the measurement clothing. The processes of steps S12 to S20 are repeated until the image association is completed. On the other hand, when the imaging control unit 114 determines that the learning has ended, the processing of this flowchart ends.
  • FIG. 7 is a flowchart showing an example of image generation processing. Prior to the processing of the flowchart shown in FIG. 7, it is assumed that the try-on garment has been selected by the user in advance. Further, the process of the flowchart shown in FIG. 7 is executed by a predetermined operation as a trigger, for example, when an image of the user wearing the measurement clothing is captured.
  • the image generation unit 118 divides the image of the user wearing the measurement clothing into the image of the body part of the user and the image of the measurement clothing (step S30).
  • the image generation unit 118 inputs the image of the measurement clothing that has been divided into regions in the previous step S30 into the learned model 126 (step S32).
  • the image generation unit 118 synthesizes the image of the try-on garment output from the learned model 126 and the image of the user's body part that has been divided into regions in step S30 (step S34).
  • the image generation unit 118 outputs the combined image to the display 30 (step S36). This completes the processing of this flowchart.
  • the image processing device 100 uses the image of the measurement clothing worn by the robot mannequin 20 in a predetermined posture as input data, and outputs the image of the try-on clothing worn by the robot mannequin 20 in a posture common to the predetermined posture.
  • the learned model 126 is learned by machine learning using the learning data 124 associated as data. Further, the image processing apparatus 100 inputs the image of the user wearing the measurement clothing to the learned model 126 to generate the image of the user wearing the try-on clothing. That is, the image of the measurement clothing worn by the user is converted into the image of the try-on clothing using the learned model 126, and the image of the user wearing the try-on clothing is generated using the converted image.
  • the image processing apparatus 100 extracts the image of the measurement clothing and the image of the body part of the user from the image of the user wearing the measurement clothing, and inputs the extracted image of the measurement clothing to the learned model 126. To do. Further, the image processing apparatus 100 synthesizes the image of the try-on garment output from the learned model 126 and the image of the body part of the user extracted previously to obtain the image of the user trying on the try-on garment. To generate. That is, by aligning and synthesizing the image of the try-on clothing converted using the learned model 126 and the image of the user's body part, it is possible to accurately synthesize the image of the user trying on the try-on clothing. it can.
  • the image processing apparatus 100 extracts the depth information of the image of the measurement clothing and the depth information of the image of the body part of the user from the image of the user wearing the measurement clothing, and based on the extracted depth information, The image of the try-on clothes and the image of the body part of the user are combined. Thereby, the image of the user who tried on the try-on clothes can be more accurately combined.
  • the image processing apparatus 100 captures an image of clothing with the camera 10 in a state where the posture of the robot mannequin 20 on which the clothing is attached is controlled to a predetermined orientation, and the captured clothing image is displayed on the robot mannequin 20. It is stored in the storage unit 120 in association with the posture. Thereby, the image of the measurement clothing and the image of the try-on clothing can be accurately associated, and the reliability of the learning data 124 can be improved.
  • the image processing apparatus 100 learns to include images of the robot mannequin 20 viewed from various rotation positions, such as images of the robot mannequin 20 viewed from the front, in addition to images viewed from the front.
  • the data for use 124 is configured. As a result, even when the user takes various postures, it is possible to accurately synthesize the image of the user who tried on the try-on clothing.
  • the image processing apparatus 100 configures the learning data 124 by associating the image of the measurement clothing and the image of the try-on clothing. Accordingly, it is possible to accurately predict the deformation of the try-on garment due to the change in the posture of the user, and it is possible to more accurately synthesize the image of the user wearing the try-on garment.
  • the learning unit 116 uses the image of the measurement clothing as input data, and the images of the trial clothing of a plurality of variations in which the image of the measurement clothing and the posture of the robot manikin 20 are common as output data.
  • the associated learning data 124 is stored in the storage unit 120.
  • the plurality of variations include, for example, the color and size of the try-on garment.
  • the learning data 124 uses, for example, the image of the measurement clothing as input data, associates the image of the first variation try-on clothing in which the posture of the measurement clothing and the robot mannequin 20 are common as output data, and An image of clothing is used as input data, and an image of a second variation of try-on clothing in which the postures of the measurement clothing and the robot mannequin 20 are common is associated as output data.
  • FIG. 8 is a diagram showing an example of the correspondence between the image of the measurement clothing and the image of the try-on clothing.
  • the image information of the measurement clothing and the image information of the try-on clothing are associated with each other for each posture of the robot mannequin 20.
  • the learning data 124 corresponds to, for example, the image information “A1” of the measurement clothing in which the posture of the robot mannequin 20 is “posture 1” and the image information “B11 ⁇ ” of the S-size trial clothing. It is attached.
  • the image information “A1” of the measurement clothing in which the posture of the robot manikin 20 is “posture 1” and the image information “B11 ⁇ ” of the M-size try-on clothing are associated with each other. Further, the image information “A1” of the measurement clothing in which the posture of the robot mannequin 20 is “posture 1” and the image information “B11 ⁇ ” of the L-size try-on clothing are associated with each other.
  • FIG. 9 is a diagram for explaining an example of image generation processing by the image processing apparatus 100.
  • the image processing apparatus 100 divides the image of the user wearing the measuring clothes into the image of the measuring clothes and the image of the body part of the user. Then, the image processing apparatus 100 inputs the image of the measurement clothing, which has been divided into regions, to the learned model 126.
  • the size of the try-on garment output from the learned model 126 (“M size” in the illustrated example) is selected in advance. As a result, an image of the try-on garment corresponding to the size selected in advance is output from the learned model 126.
  • the image processing apparatus 100 synthesizes the image of the try-on garment output from the learned model 126 and the image of the user's body part, which is divided into regions as described above, so that the user wearing the try-on garment. Generate an image of.
  • the image processing apparatus 100 accepts the selection of the variation of the try-on clothing, inputs the image of the user wearing the measurement clothing to the learned model 126, and tries on the try-on clothing of the previously selected variation. Generate an image of the user who made the request. Accordingly, it is possible to generate an image of the user who has tried on the try-on garment by distinguishing it for each variation of the try-on garment selected by the user.
  • the third embodiment is different from the first embodiment in the method of associating the image of the measurement clothing and the image of the try-on clothing. Therefore, in the following description, a configuration different from that of the first embodiment will be mainly described, and a duplicate description of the same or corresponding configuration as that of the first embodiment will be omitted.
  • the learning unit 116 uses the image of measurement clothing as input data, and associates the image of measurement clothing with the image of try-on clothing in which the body shape and posture of the robot manikin 20 are common as output data.
  • the learning data 124 is stored in the storage unit 120.
  • the body shape of the robot mannequin 20 is defined by parameters such as arm circumference, shoulder width, and abdominal circumference.
  • the learning data 124 includes, for example, an image of the measurement clothing worn by the robot mannequin 20 corresponding to a normal body shape, and an image of the try-on clothing in which the measurement clothing and the robot mannequin 20 have the same body shape and posture.
  • the data 124 is configured. That is, the learning data 124 associates the image of the measurement clothing with the image of the try-on clothing so as to correspond to users of a plurality of body types, and the number of data of the learning data 124 is increased.
  • FIG. 10 is a diagram showing an example of the correspondence between the image of the measurement clothing and the image of the try-on clothing.
  • the image information of the measurement clothing and the image information of the try-on clothing are associated with each other for each body type and posture of the robot mannequin 20.
  • the learning data 124 is, for example, for the robot mannequin 20 corresponding to a normal body type, the image information “A1” of the measurement clothing in which the posture of the robot mannequin 20 is “posture 1” and the trial clothing.
  • the image information “B11” is associated.
  • the image information “A1X” of the measurement clothing in which the posture of the robot mannequin 20 is “posture 1” and the image information “B11X” of the try-on clothing are associated with each other. ..
  • FIG. 11 is a diagram for explaining an example of image generation processing by the image processing apparatus 100.
  • the image processing apparatus 100 extracts the image of the measurement clothing and the image of the body part of the user from the image of the user who has the normal figure wearing the measurement clothing. Divide into areas. Then, the image processing apparatus 100 inputs the image of the measurement clothing, which has been divided into regions, to the learned model 126. As a result, an image of the try-on garment close to a normal body shape is output from the learned model 126. After that, the image processing apparatus 100 synthesizes the image of the try-on garment output from the learned model 126 and the image of the user's body part, which is divided into regions as described above, so that the user wearing the try-on garment. Generate an image of.
  • the image processing apparatus 100 includes an image of the obesity-type user wearing measurement clothing, an image of the measurement clothing, and an image of the body part of the user as an area. To divide. Then, the image processing apparatus 100 inputs the image of the measurement clothing, which has been divided into regions, to the learned model 126. As a result, an image of the try-on garment close to the obesity type is output from the learned model 126. After that, the image processing apparatus 100 synthesizes the image of the try-on garment output from the learned model 126 and the image of the user's body part, which is divided into regions as described above, so that the user wearing the try-on garment. Generate an image of.
  • the image processing apparatus 100 learns the learned model 126 using the learning data 124 in which the image of the measurement clothing and the image of the try-on clothing are associated with each other so as to correspond to the users of a plurality of body types. There is. In addition, the image processing apparatus 100 inputs the image of the user wearing the measurement clothing into the learned model 126 to generate images in which the users of various body types wear the try-on clothing. That is, by configuring the learning data 124 so as to correspond to users of various body types, it is possible to generate a highly realistic image of try-on clothing while suppressing the calculation load.
  • the image processing apparatus 100 controls the robot mannequin 20 to collect an image of the measurement clothing and an image of the try-on clothing as the learning data 124.
  • the image of the measurement clothing and the image of the try-on clothing may be collected as the learning data 124 by changing the posture of the subject wearing the clothing.
  • the learned model 126 is learned using the image of the measurement clothing worn by the user.
  • the learned model 126 may be learned using an image of normal clothing worn by the user.
  • the learning data 124 may be configured by estimating the posture of the user based on the image of the ordinary clothing worn by the user and associating the estimated posture of the user with the image of the ordinary clothing.
  • the image in which the ordinary clothes are attached to the robot mannequin 20 may be used, or the subject is caused to wear the ordinary clothes. Images may be used.
  • FIG. 12 is a diagram illustrating an example of the hardware configuration of the image processing apparatus 100 according to the embodiment.
  • the image processing apparatus 100 includes a communication controller 100-1, a CPU 100-2, a RAM (Randome Access Memory) 100-3 used as a working memory, and a ROM (Read Only Memory) 100 for storing a boot program and the like.
  • a storage device 100-5 such as a flash memory or an HDD (Hard Disk Drive), a drive device 100-6, etc. are connected to each other by an internal bus or a dedicated communication line.
  • the communication controller 100-1 communicates with components other than the image processing apparatus 100.
  • a program 100-5a executed by the CPU 100-2 is stored in the storage device 100-5.
  • This program is expanded in the RAM 100-3 by a DMA (Direct Memory Access) controller (not shown) or the like and executed by the CPU 100-2.
  • a DMA Direct Memory Access
  • the robot control unit 112, the shooting control unit 114, the learning unit 116, and the image generation unit 118 are realized.
  • a storage device that stores a program
  • a hardware processor executes a program stored in the storage device
  • the learning data in which the image of the reference wear product worn by the wearer in a predetermined posture is used as input data and the image of the try-on wear product worn by the wearer in a posture common to the predetermined posture is used as output data is used.
  • An image processing device configured to input an image of a user wearing the reference wearing product to a learned model learned by machine learning and generate an image of a user trying on the fitting wear product. ..
  • Image generation system 10... Camera, 20... Robot mannequin, 30... Display, 100... Image processing device, 110... Control part, 112... Robot control part, 114... Shooting control part, 116... Learning part, 118... Image Generation unit, 120... Storage unit, 122... Image data, 124... Learning data, 126... Learned model.

Landscapes

  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Image Processing (AREA)
  • Processing Or Creating Images (AREA)

Abstract

仮想的な装着品の試着を行う場合の計算負荷を低減することができる画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラムを提供する。 画像処理装置100は、所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データ124を用いた機械学習により学習させた学習済みモデル126に対し、参照装着品を装着したユーザの画像を入力して、試着装着品を試着したユーザの画像を生成する画像生成部118を備える。

Description

画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラム 関連出願の相互参照
 本出願は、2019年2月22日に出願された米国仮出願62/809088に基づくもので、ここにその記載内容を援用する。
 本発明は、画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラムに関する。
 従来、人体の動きに合わせて衣類の画像を生成し、生成した衣類の画像と人体の画像とを重ね合わせることで、仮想的な衣類の試着を行う技術が知られている(例えば、非特許文献1参照)。また、近年では、深層学習を用いて仮想的な衣類の試着を実現する技術も提案されている(例えば、非特許文献2参照)。
Kim,J.,&Forsythe,S.″Adoption of Virtual Try - on technology for online apparel shopping.″Journal of Interactive Marketing, Volume 22, Spring 2008, Page 45-59. Christoph Lassner,Gerard Pons-Moll,Peter V.Gehler.The IEEE International Conference on Computer Vision(ICCV), July 2017,Page 853-862.
 しかしながら、従来の技術では、目的とする衣服毎に設定される多数のパラメータを用いたシミュレーションを実行する際に計算負荷が過大となってしまうという問題があった。
 なお、こうした課題は、衣類を仮想的に試着する場合に限らず、例えば、眼鏡や靴などの他の装着品を仮想的に試着する場合にも概ね共通するものであった。
 そこで、本発明は、仮想的な装着品の試着を行う場合の計算負荷を低減することができる画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラムを提供する。
 本発明の一態様に係る画像処理装置は、所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データを用いた機械学習により学習させた学習済みモデルに対し、前記参照装着品を装着したユーザの画像を入力して、前記試着装着品を試着したユーザの画像を生成する画像生成部を備える。
 この態様によれば、参照装着品の画像と試着装着品の画像とを対応付けた学習用データを用いて学習済みモデルを学習させ、この学習モデルに参照装着品を装着したユーザの画像を入力することで、仮想的な装着品の試着を行う場合の計算負荷を低減することができる。
 上記態様において、前記画像生成部は、前記参照装着品を装着したユーザの画像から前記参照装着品の画像と前記ユーザの身体部分の画像とを抽出し、前記抽出した参照装着品の画像を前記学習済みモデルに入力し、前記学習済みモデルから出力された前記試着装着品の画像と、前記抽出したユーザの身体部分の画像とを合成することによって、前記試着装着品を試着したユーザの画像を生成してもよい。
 この態様によれば、試着装着品を試着したユーザの画像を正確に合成することができる。
 上記態様において、前記画像生成部は、前記参照装着品を装着したユーザの画像から前記参照装着品の画像の深度情報と前記ユーザの身体部分の画像の深度情報とを抽出し、前記抽出した深度情報に基づいて、前記試着装着品の画像と前記ユーザの身体部分の画像とを合成してもよい。
 この態様によれば、画像の深度情報を用いることで、試着装着品を装着したユーザの画像をより正確に合成することができる。
 上記態様において、前記学習用データは、所定の姿勢で装着主体が装着した前記参照装着品の画像を前記入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した第1のバリエーションの前記試着装着品の画像を前記出力データとして関連付け、かつ、前記入力データに対して、前記所定の姿勢と共通する姿勢で装着主体が装着した第2のバリエーションの前記試着装着品の画像を前記出力データとして関連付け、前記画像生成部は、前記試着装着品のバリエーションの選択を受け付け、前記参照装着品を装着したユーザの画像を前記学習済みモデルに入力して、前記選択されたバリエーションの前記試着装着品を試着したユーザの画像を生成してもよい。
 この態様によれば、ユーザにより選択された試着装着品のバリエーションごとに区別して、試着装着品を試着したユーザの画像を生成することができる。
 上記態様において、前記装着品は、衣類であり、前記学習用データは、所定の姿勢で装着主体が装着した前記参照衣類の画像を前記入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した前記試着衣類の画像を前記出力データとして関連付けたデータを含んでもよい。
 この態様によれば、参照衣類の画像と試着衣類の画像とを対応付けた学習用データを用いて学習済みモデルを学習させることで、仮想的な衣類の試着を行う場合の計算負荷を低減することができる。
 本発明の一態様に係る学習済みモデルは、所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データを用いて機械学習により学習させた。
 この態様によれば、上記画像処理装置の発明と同様の効果が得られる。
本発明の一態様に係る画像収集装置は、ロボットマネキンと、カメラと、前記ロボットマネキンの姿勢を制御するロボット制御部と、前記ロボット制御部の制御に基づいて装着品を装着させた前記ロボットマネキンの姿勢を所定の姿勢に制御した状態で、前記カメラで前記装着品の画像を撮影し、撮影された画像を前記姿勢と関連付けて記憶部に記憶させる撮影制御部とを備える。
 この態様によれば、参照装着品の画像と試着装着品の画像との対応付けを的確に行うことができ、学習済みモデルの学習に用いられる学習用データの信頼性を高めることができる。
 本発明の一態様に係る画像処理方法は、所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データを用いた機械学習により学習させた学習済みモデルに対し、前記参照装着品を装着したユーザの画像を入力して、前記試着装着品を試着したユーザの画像を生成する画像生成ステップを含む。
 この態様によれば、上記画像処理装置の発明と同様の効果が得られる。
本発明の一態様に係る画像処理プログラムは、コンピュータに、所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データを用いた機械学習により学習させた学習済みモデルに対し、前記参照装着品を装着したユーザの画像を入力して、前記試着装着品を試着したユーザの画像を生成させる処理を実行させる。
 この態様によれば、上記画像処理プログラムの発明と同様の効果が得られる。
 本発明によれば、仮想的な装着品の試着を行う場合の計算負荷を低減することができる。
画像処理装置の第1の実施の形態の概略構成を示すブロック図。 画像データのデータ内容の一例を示す図。 学習用データのデータ内容の一例を示す図。 計測用衣類の画像と試着用衣類の画像との対応付けの一例を示す図。 画像の生成過程の一例を説明するための図。 学習済みモデルの学習処理の一例を示すフローチャート。 画像生成処理の処理内容を示すフローチャート。 計測用衣類の画像と試着用衣類の画像との対応付けの一例を示す図。 画像の生成過程の一例を説明するための図。 計測用衣類の画像と試着用衣類の画像との対応付けの一例を示す図。 画像の生成過程の一例を説明するための図。 画像処理装置のハードウェア構成の一例を示す図である。
 (第1の実施の形態)
 以下、図面を参照し、本発明の画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および画像処理プログラムの実施形態について説明する。
 図1に示すように、画像収集装置1は、例えば、カメラ10と、ロボットマネキン20と、ディスプレイ30と、画像処理装置100とを備える。これらの装置や機器は、通信線や無線通信網等によって互いに接続される。なお、図1に示す構成はあくまで一例であり、構成の一部が省略されてもよいし、更に別の構成が追加されてもよい。
 カメラ10は、例えば、CCD(Charge Coupled Device)やCMOS(Complementary Metal Oxide Semiconductor)等の固体撮像素子を利用したデジタルカメラである。カメラ10は、例えば、ロボットマネキン20の前方に固定して配置され、装着品の一例である衣類を装着したロボットマネキン20の画像を撮影する。カメラ10は、例えば、ステレオカメラであってもよいし、深度カメラであってもよい。
 ロボットマネキン20は、例えば、ヒト型ロボットである。ロボットマネキン20は、例えば、両肩に4自由度、両肘に4自由度、腰の回転に2自由度で駆動可能に構成されており、衣類を装着した状態で様々な姿勢を取ることが可能である。
 ディスプレイ30は、例えば、液晶ディスプレイであり、画像処理装置100により生成された画像を表示する。
 画像処理装置100は、例えば、制御部110と、記憶部120とを備える。制御部110は、例えば、CPU(Central Processing Unit)などのハードウェアプロセッサがプログラム(ソフトウェア)を実行することにより実現される。また、これらの構成要素のうち一部または全部は、LSI(Large Scale Integration)やASIC(Application Specific Integrated Circuit)、FPGA(Field-Programmable Gate Array)、GPU(Graphics Processing Unit)などのハードウェア(回路部:circuitryを含む)によって実現されてもよいし、ソフトウェアとハードウェアの協働によって実現されてもよい。プログラムは、予め画像処理装置100のHDDやフラッシュメモリなどの記憶装置に格納されていてもよいし、DVDやCD-ROMなどの着脱可能な記憶媒体に格納されており、記憶媒体がドライブ装置に装着されることで画像処理装置100のHDDやフラッシュメモリにインストールされてもよい。
 制御部110は、例えば、ロボット制御部112と、撮影制御部114と、学習部116と、画像生成部118とを備える。
 ロボット制御部112は、ロボットマネキン20の動作を制御する。ロボット制御部112は、例えば、ロボットマネキン20の各自由度における動作量を制御することにより、ロボットマネキン20の姿勢を制御する。ロボットマネキン20の各自由度は、両肩の自由度、両肘の自由度、腰の回転の自由度など、ロボットマネキン20のポーズの自由度に加え、鉛直方向に延びる中心軸を中心としたロボットマネキン20の回転位置の自由度を含む。ロボットマネキン20の自由度は、例えば、カメラ10の撮影位置・方向に対するロボットマネキン20の角度を含んでもよい。なお、上記の説明は、ロボットマネキン20の姿勢の一例にすぎず、これに限られない。
 撮影制御部114は、カメラ10による画像の撮影動作を制御する。撮影制御部114は、例えば、ロボット制御部112の制御に基づいてロボットマネキン20の姿勢を所定の姿勢に制御した状態でロボットマネキン20の画像を撮影する。撮影制御部114は、例えば、ロボット制御部112の制御に基づいて、ロボットマネキン20の種々のポーズについて、ロボットマネキン20の回転位置を変化させつつ、ロボットマネキン20の画像を撮影する。撮影制御部114は、撮影された画像をロボットマネキン20の姿勢と関連付けて記憶部120に記憶させる。撮影制御部114は、例えば、計測用衣類を装着したロボットマネキン20の画像、および、試着用衣類を装着したロボットマネキン20の画像を撮影する。撮影制御部114は、計測用衣類を装着したロボットマネキン20の画像から計測用衣類の画像を領域分割し、領域分割した画像を、画像の撮影時におけるロボットマネキン20の姿勢と関連付けて画像データ122として記憶部120に記憶させる。また、撮影制御部114は、試着用衣類を装着したロボットマネキン20の画像から試着用衣類の画像を領域分割し、領域分割した画像を、画像の撮影時におけるロボットマネキン20の姿勢と関連付けて画像データ122として記憶部120に記憶させる。
 図2は、画像データ122のデータ内容の一例を示す図である。図示の例では、画像データ122のデータ属性は、例えば、画像の種類、画像情報、および、ロボットマネキン20の姿勢を含む。画像の種類は、ロボットマネキン20が装着している衣類の種類であり、例えば、計測用衣類、および、試着用衣類を含む。計測用衣類は、試着用衣類を合成する上で参照される参照衣類の一例である。計測用衣類は、例えば、ユーザの身体に関する情報を取得するための衣類であり、人体の肌の色と区別しやすく、様々な試着用衣類と対応付けがしやすいことが好ましい。試着用衣類は、仮想的な試着の対象となる衣類であり、例えば、デザインが異なる複数の種類の衣類を含む。画像情報は、カメラ10により撮影された画像の画像位置ごとの輝度値(R,G,B)に関する情報を含む。画像情報は、さらに、カメラ10により撮影された画像の画像位置ごとの深度情報を含んでもよい。ロボットマネキン20の姿勢は、ロボットマネキン20の各自由度の動作量の組み合わせにより規定されるパラメータである。
 学習部116は、計測用衣類の画像を入力データとし、計測用衣類の画像とロボットマネキン20の姿勢が共通する試着用衣類の画像を出力データとして関連付けた学習用データ124を記憶部120に格納する。すなわち、学習部116は、同じポーズで同じ方向から撮影した計測用衣類の画像と試着用衣類の画像とを対応付けて学習用データ124として記憶部120に格納する。学習部116は、例えば、記憶部120に格納された画像データ122を参照し、計測用衣類の画像と試着用衣類の画像とを対応付けて学習用データ124を記憶部120に格納する。学習部116は、学習用データ124を用いた機械学習により学習モデルを学習させて学習済みモデル126を生成する。学習済みモデル126は、例えば、機械学習のモデルの一種であるニューラルネットワークにより構成されている。学習部116は、計測用衣類の画像および試着用衣類の画像に対して平行移動、拡大縮小、回転、ノイズの付与などの処理を加えることで、学習用データ124のデータ数を水増しする、いわゆるデータオーギュメンテーションを行ってもよい。
 図3は、学習用データ124のデータ内容の一例を示す図である。図示の例では、学習用データ124は、計測用衣類の画像情報として、ロボットマネキン20の姿勢が互いに異なる複数(図示の例ではN個)の画像情報「A1」~「AN」が含まれている。また、学習用データ124は、計測用衣類のそれぞれの画像情報「A1」~「AN」とロボットマネキン20の姿勢が共通する試着用衣類の画像情報が対応付けられている。具体的には、学習用データ124は、計測用衣類のそれぞれの画像情報「A1」~「AN」とロボットマネキン20の姿勢が共通する衣類ごと(「衣類1」、「衣類2」など)の画像情報とが対応付けられている。
 図4は、計測用衣類の画像と試着用衣類の画像との対応付けの一例を示す図である。図4に示すように、学習用データ124は、ロボットマネキン20の姿勢ごとに区別して、計測用衣類の画像情報と試着用衣類の画像情報とが対応付けられている。図示の例では、学習用データ124は、例えば、ロボットマネキン20の姿勢が「姿勢1」である場合の計測用衣類の画像情報「A1」と試着用衣類の画像情報「B11」とが対応付けられ、ロボットマネキン20の姿勢が「姿勢2」である場合の計測用衣類の画像情報「A2」と試着用衣類の画像情報「B12」とが対応付けられている。
 画像生成部118は、計測用衣類を装着したユーザの画像を学習済みモデル126入力して、試着用衣類を試着したユーザの画像を生成する。画像生成部118は、例えば、計測用衣類を装着したユーザの画像から計測用衣類の画像とユーザの身体部分の画像とを抽出し、抽出した計測用衣類の画像を学習済みモデル126に入力し、学習済みモデル126から出力された試着用衣類の画像と、抽出したユーザの身体部分の画像とを合成する。
 図5は、画像処理装置100による画像の生成処理の一例を説明するための図である。
同図に示すように、画像処理装置100は、学習フェーズとして、まず、計測用衣類を装着したロボットマネキン20の画像から計測用衣類の画像を領域分割し、かつ、試着用衣類を装着したロボットマネキン20の画像から試着用衣類の画像を領域分割する。
 次に、画像処理装置100は、領域分割された計測用衣類の画像と試着用衣類の画像とを学習用データ124として対応付け、この学習用データ124を用いた機械学習により学習モデルを学習させることにより、学習済みモデル126を生成する。
 次に、画像処理装置100は、実行フェーズとして、計測用衣類を装着したユーザの画像から、計測用衣類の画像、および、ユーザの身体部分の画像を領域分割する。この場合、画像処理装置100は、計測用衣類の画像の深度情報と、ユーザの身体部分の画像の深度情報とを取得してもよい。これにより、計測用衣類の画像と、ユーザの身体部分の画像との領域分割が正確に行われる。そして、画像処理装置100は、領域分割した計測用衣類の画像を学習済みモデル126に入力する。これにより、ユーザの姿勢に対応する試着用衣類の画像が学習済みモデル126から出力される。
次に、画像処理装置100は、学習済みモデル126から出力された試着用衣類の画像と、上述のように領域分割したユーザの身体部分の画像とを合成することにより、試着用衣類を装着したユーザの画像を生成する。画像処理装置100は、例えば、ユーザの身体部分の画像を、計測用衣類の画像よりも手前に位置するように、試着用衣類を装着したユーザの画像を合成する。画像処理装置100は、例えば、計測用衣類の画像の深度情報と、ユーザの身体部分の画像の深度情報とに基づいて、試着用衣類を装着したユーザの画像を生成してもよい。この場合、画像処理装置100は、例えば、衣類の首の部分にユーザの顔や首を合成できたり、衣類の袖の部分にユーザの腕を合成できたりするなど、試着用異類を装着したユーザの画像を精度よく合成することができる。
 図6は、学習済みモデル126の学習処理の一例を示すフローチャートである。なお、図6に示すフローチャートの処理に先立ち、ロボットマネキン20の姿勢ごとの計測用衣類の画像の撮影は完了しているものとする。また、図6に示すフローチャートの処理は、例えば、ロボットマネキン20に試着用衣類が装着された場合に、所定の操作をトリガーとして実行される。
 図6に示すように、まず、ロボット制御部112は、試着用衣類を装着したロボットマネキン20を所定の姿勢に制御する(ステップS10)。次に、撮影制御部114は、カメラ10を用いてロボットマネキン20の画像を撮影する(ステップS12)。次に、撮影制御部114は、ロボットマネキン20の画像から試着用衣類の画像を領域分割する(ステップS14)。次に、撮影制御部114は、ロボットマネキン20の姿勢が共通する計測用衣類の画像と試着用衣類の画像とを対応付けて学習用データ124として記憶部120に格納する(ステップS16)。次に、学習部116は、記憶部120に格納された学習用データ124に加え、学習用データ124に対してデータオーギュメンテーションを行うことで水増しされたデータを用いて、機械学習により学習モデルを学習させる(ステップS18)。次に、撮影制御部114は、学習が終了したか否かを判定する(ステップS20)。そして、ロボット制御部112は、撮影制御部114により学習が終了していないと判定された場合には、ロボットマネキン20の姿勢を変更し(ステップS22)、計測用衣類の画像に対する試着用衣類の画像の対応付けが完了するまでの間、ステップS12~ステップS20の処理を繰り返す。一方、撮影制御部114は、学習が終了したと判定した場合には、本フローチャートの処理が終了する。
 図7は、画像生成処理の一例を示すフローチャートである。なお、図7に示すフローチャートの処理に先立ち、試着用衣類はユーザにより事前に選択されているものとする。また、図7に示すフローチャートの処理は、例えば、計測用衣類を装着したユーザの画像が撮影された場合に、所定の操作をトリガーとして実行される。
 図7に示すように、まず、画像生成部118は、計測用衣類を装着したユーザの画像から、ユーザの身体部分の画像と計測用衣類の画像とを領域分割する(ステップS30)。次に、画像生成部118は、先のステップS30において領域分割した計測用衣類の画像を学習済みモデル126に入力する(ステップS32)。次に、画像生成部118は、学習済みモデル126から出力された試着用衣類の画像と、先のステップS30において領域分割したユーザの身体部分の画像とを合成する(ステップS34)。次に、画像生成部118は、合成した画像をディスプレイ30に出力する(ステップS36)。これにより、本フローチャートの処理が終了する。
 以上説明したように、上記第1の実施の形態によれば、以下に示す効果を得ることができる。
 (1)画像処理装置100は、所定の姿勢でロボットマネキン20が装着した計測用衣類の画像を入力データとし、所定の姿勢と共通する姿勢でロボットマネキン20が装着した試着用衣類の画像を出力データとして関連付けた学習用データ124を用いた機械学習により学習済みモデル126を学習する。また、画像処理装置100は、学習済みモデル126に対し、計測用衣類を装着したユーザの画像を入力して、試着用衣類を試着したユーザの画像を生成する。すなわち、ユーザが装着した計測用衣類の画像を学習済みモデル126を用いて試着用衣類の画像に変換し、変換後の画像を用いて試着用衣類を試着したユーザの画像を生成する。これにより、人体や衣類の三次元モデルを用意して物理シミュレーションを実行することが不要となるため、仮想的な衣類の試着を行う場合の計算負荷を低減することができる。また、目的とする衣服毎に多数のパラメータを正確に計測することが不要となるため、仮想的な衣類の試着を行う場合の利便性を向上することができる。
 (2)画像処理装置100は、計測用衣類を装着したユーザの画像から計測用衣類の画像とユーザの身体部分の画像とを抽出し、抽出した計測用衣類の画像を学習済みモデル126に入力する。また、画像処理装置100は、学習済みモデル126から出力された試着用衣類の画像と、先に抽出したユーザの身体部分の画像とを合成することによって、試着用衣類を試着したユーザの画像を生成する。すなわち、学習済みモデル126を用いて変換した試着用衣類の画像とユーザの身体部分の画像とを位置合わせして合成することで、試着用衣類を試着したユーザの画像を正確に合成することができる。
 (3)画像処理装置100は、計測用衣類を装着したユーザの画像から計測用衣類の画像の深度情報とユーザの身体部分の画像の深度情報とを抽出し、抽出した深度情報に基づいて、試着用衣類の画像とユーザの身体部分の画像とを合成する。これにより、試着用衣類を試着したユーザの画像をより正確に合成することができる。
 (4)画像処理装置100は、衣類を装着させたロボットマネキン20の姿勢を所定の姿勢に制御した状態で、カメラ10で衣類の画像を撮影し、撮影された衣類の画像をロボットマネキン20の姿勢と関連付けて記憶部120に記憶させる。これにより、計測用衣類の画像と試着用衣類の画像との対応付けを正確に行うことができ、学習用データ124の信頼性を高めることができる。
 (5)画像処理装置100は、ロボットマネキン20を正面から見た画像に加え、ロボットマネキン20を横方向から見た画像など、ロボットマネキン20を様々な回転位置から見た画像を含むように学習用データ124を構成している。これにより、ユーザが様々な姿勢を取る場合であっても、試着用衣類を試着したユーザの画像を正確に合成することができる。
 (6)画像処理装置100は、計測用衣類の画像と試着用衣類の画像とを対応付けて学習用データ124を構成している。これにより、ユーザの姿勢の変化に伴った試着用衣類の変形を正確に予測することが可能となり、試着用衣類を装着したユーザの画像をより正確に合成することができる。
 (第2の実施の形態)
 次に、画像処理装置の第2の実施の形態について図面を参照して説明する。なお、第2の実施の形態は、計測用衣類の画像と試着用衣類の画像との対応付けの方法が第1の実施の形態と異なる。したがって、以下の説明においては、第1の実施の形態と相違する構成について主に説明し、第1の実施の形態と同一のまたは相当する構成については重複する説明を省略する。
 第2の実施の形態に係る学習部116は、計測用衣類の画像を入力データとし、計測用衣類の画像とロボットマネキン20の姿勢が共通する複数のバリエーションの試着用衣類の画像を出力データとして関連付けた学習用データ124を記憶部120に格納する。複数のバリエーションは、例えば、試着用衣類の色やサイズを含む。学習用データ124は、例えば、計測用衣類の画像を入力データとし、計測用衣類とロボットマネキン20の姿勢が共通する第1のバリエーションの試着用衣類の画像を出力データとして関連付け、かつ、計測用衣類の画像を入力データとし、計測用衣類とロボットマネキン20の姿勢が共通する第2のバリエーションの試着用衣類の画像を出力データとして関連付けている。
 図8は、計測用衣類の画像と試着用衣類の画像との対応付けの一例を示す図である。図8に示すように、学習用データ124は、ロボットマネキン20の姿勢ごとに区別して、計測用衣類の画像情報と試着用衣類の画像情報とが対応付けられている。図示の例では、学習用データ124は、例えば、ロボットマネキン20の姿勢が「姿勢1」である計測用衣類の画像情報「A1」とSサイズの試着用衣類の画像情報「B11α」とが対応付けられている。また、ロボットマネキン20の姿勢が「姿勢1」である計測用衣類の画像情報「A1」とMサイズの試着用衣類の画像情報「B11β」とが対応付けられている。また、ロボットマネキン20の姿勢が「姿勢1」である計測用衣類の画像情報「A1」とLサイズの試着用衣類の画像情報「B11γ」とが対応付けられている。
 図9は、画像処理装置100による画像の生成処理の一例を説明するための図である。
同図に示すように、画像処理装置100は、実行フェーズとして、計測用衣類を装着したユーザの画像から、計測用衣類の画像、および、ユーザの身体部分の画像を領域分割する。そして、画像処理装置100は、領域分割した計測用衣類の画像を学習済みモデル126に入力する。この場合、学習済みモデル126から出力される試着用衣類のサイズ(図示の例では「Mサイズ」)が事前に選択されている。これにより、事前に選択されたサイズに対応する試着用衣類の画像が学習済みモデル126から出力される。その後、画像処理装置100は、学習済みモデル126から出力された試着用衣類の画像と、上述のように領域分割したユーザの身体部分の画像とを合成することにより、試着用衣類を装着したユーザの画像を生成する。
 以上説明したように、上記第2の実施の形態によれば、第1の実施の形態の上記(1)~(6)の効果に加えて、以下に示す効果を得ることができる。
 (7)画像処理装置100は、試着用衣類のバリエーションの選択を受け付け、計測用衣類を装着したユーザの画像を学習済みモデル126に入力して、先に選択されたバリエーションの試着用衣類を試着したユーザの画像を生成する。これにより、ユーザにより選択された試着用衣類のバリエーションごとに区別して、試着用衣類を試着したユーザの画像を生成することができる。
 (第3の実施の形態)
 次に、画像処理装置の第3の実施の形態について図面を参照して説明する。なお、第3の実施の形態は、計測用衣類の画像と試着用衣類の画像との対応付けの方法が第1の実施の形態と異なる。したがって、以下の説明においては、第1の実施の形態と相違する構成について主に説明し、第1の実施の形態と同一のまたは相当する構成については重複する説明を省略する。
 第3の実施の形態に係る学習部116は、計測用衣類の画像を入力データとし、計測用衣類の画像とロボットマネキン20の体型並びに姿勢が共通する試着用衣類の画像を出力データとして関連付けた学習用データ124を記憶部120に格納する。ロボットマネキン20の体型は、例えば、腕囲、肩幅、腹囲などのパラメータにより規定される。学習用データ124は、例えば、通常の体型に相当するロボットマネキン20により装着された計測用衣類の画像と、この計測用衣類とロボットマネキン20の体型ならびに姿勢が共通する試着用衣類の画像とを対応付け、かつ、肥満型に相当するロボットマネキン20により装着された計測用衣類の画像と、この計測用衣類とロボットマネキン20の体型ならびに姿勢が共通する試着用衣類の画像と対応付けて学習用データ124を構成している。すなわち、学習用データ124は、複数の体型のユーザに対応するように計測用衣類の画像と試着用衣類の画像とを対応付けており、学習用データ124のデータ数が水増しされている。
 図10は、計測用衣類の画像と試着用衣類の画像との対応付けの一例を示す図である。図10に示すように、学習用データ124は、ロボットマネキン20の体型ならびに姿勢ごとに区別して、計測用衣類の画像情報と試着用衣類の画像情報とが対応付けられている。図示の例では、学習用データ124は、例えば、通常の体型に相当するロボットマネキン20について、ロボットマネキン20の姿勢が「姿勢1」である計測用衣類の画像情報「A1」と試着用衣類の画像情報「B11」とが対応付けられている。また、肥満型に相当するロボットマネキン20について、ロボットマネキン20の姿勢が「姿勢1」である計測用衣類の画像情報「A1X」と試着用衣類の画像情報「B11X」とが対応付けられている。
 図11は、画像処理装置100による画像の生成処理の一例を説明するための図である。
図11(a)に示すように、画像処理装置100は、実行フェーズとして、計測用衣類を装着した通常の体型のユーザの画像から、計測用衣類の画像、および、ユーザの身体部分の画像を領域分割する。そして、画像処理装置100は、領域分割した計測用衣類の画像を学習済みモデル126に入力する。これにより、通常の体型に近い試着用衣類の画像が学習済みモデル126から出力される。その後、画像処理装置100は、学習済みモデル126から出力された試着用衣類の画像と、上述のように領域分割したユーザの身体部分の画像とを合成することにより、試着用衣類を装着したユーザの画像を生成する。
 図11(b)に示すように、画像処理装置100は、実行フェーズとして、計測用衣類を装着した肥満型のユーザの画像から、計測用衣類の画像、および、ユーザの身体部分の画像を領域分割する。そして、画像処理装置100は、領域分割した計測用衣類の画像を学習済みモデル126に入力する。これにより、肥満型に近い試着用衣類の画像が学習済みモデル126から出力される。その後、画像処理装置100は、学習済みモデル126から出力された試着用衣類の画像と、上述のように領域分割したユーザの身体部分の画像とを合成することにより、試着用衣類を装着したユーザの画像を生成する。
 以上説明したように、上記第3の実施の形態によれば、第1の実施の形態の上記(1)~(6)の効果に加えて、以下に示す効果を得ることができる。
 (8)画像処理装置100は、複数の体型のユーザに対応するように計測用衣類の画像と試着用衣類の画像とを対応付けた学習用データ124を用いて学習済みモデル126を学習している。また、画像処理装置100は、計測用衣類を装着したユーザの画像を学習済みモデル126に入力して、様々な体型のユーザが試着用衣類を装着した画像を生成する。すなわち、様々な体型のユーザに対応するように学習用データ124を構成することで、計算負荷を抑えつつ、リアリティの高い試着用衣類の画像を生成することができる。
 (その他の実施の形態)
 なお、上記各実施の形態は、以下のような形態にて実施することもできる。
 ・上記各実施の形態においては、画像処理装置100がロボットマネキン20を制御することにより、計測用衣類の画像および試着用衣類の画像を学習用データ124として収集した。これに代えて、被験者が衣類を着用した状態で姿勢を変更することにより、計測用衣類の画像および試着用衣類の画像を学習用データ124として収集してもよい。
 ・上記各実施の形態においては、ユーザが装着した計測用衣類の画像を用いて学習済みモデル126を学習するようにした。これに代えて、ユーザが装着した普通の衣類の画像を用いて学習済みモデル126を学習するようにしてもよい。例えば、ユーザが装着した普通の衣類の画像に基づいてユーザの姿勢を推定し、推定したユーザの姿勢を普通の衣類の画像と関連付けて学習用データ124を構成してもよい。なお、普通の衣類の画像に基づいてユーザの姿勢を推定する場合には、ロボットマネキン20に普通の衣類を装着させた画像を用いてもよいし、被検者に普通の衣類を装着させた画像を用いてもよい。
 〔ハードウェア構成〕
 図12は、実施の形態の画像処理装置100のハードウェア構成の一例を示す図である。図示するように、画像処理装置100は、通信コントローラ100-1、CPU100-2、ワーキングメモリとして使用されるRAM(Randome Access Memory)100-3、ブートプログラムなどを格納するROM(Read Only Memory)100-4、フラッシュメモリやHDD(Hard Disk Drive)などの記憶装置100-5、ドライブ装置100-6などが、内部バスあるいは専用通信線によって相互に接続された構成となっている。通信コントローラ100-1は、画像処理装置100以外の構成要素との通信を行う。記憶装置100-5には、CPU100-2が実行するプログラム100-5aが格納されている。このプログラムは、DMA(Direct Memory Access)コントローラ(不図示)などによってRAM100-3に展開されて、CPU100-2によって実行される。これによって、ロボット制御部112、撮影制御部114、学習部116、画像生成部118が実現される。
 上記説明した実施形態は、以下のように表現することができる。
 プログラムを記憶した記憶装置と、
 ハードウェアプロセッサと、を備え、
 前記ハードウェアプロセッサは、前記記憶装置に記憶されたプログラムを実行することにより、
 所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データを用いた機械学習により学習させた学習済みモデルに対し、前記参照装着品を装着したユーザの画像を入力して、前記試着装着品を試着したユーザの画像を生成させるように構成されている、画像処理装置。
 以上説明した実施形態は、本発明の理解を容易にするためのものであり、本発明を限定して解釈するためのものではない。実施形態が備える各要素並びにその配置、材料、条件、形状及びサイズ等は、例示したものに限定されるわけではなく適宜変更することができる。また、異なる実施形態で示した構成同士を部分的に置換し又は組み合わせることが可能である。
 1…画像生成システム、10…カメラ、20…ロボットマネキン、30…ディスプレイ、100…画像処理装置、110…制御部、112…ロボット制御部、114…撮影制御部、116…学習部、118…画像生成部、120…記憶部、122…画像データ、124…学習用データ、126…学習済みモデル。

Claims (9)

  1.  所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データを用いた機械学習により学習させた学習済みモデルに対し、前記参照装着品を装着したユーザの画像を入力して、前記試着装着品を試着したユーザの画像を生成する画像生成部を備える、
    画像処理装置。
  2.  前記画像生成部は、
     前記参照装着品を装着したユーザの画像から前記参照装着品の画像と前記ユーザの身体部分の画像とを抽出し、
     前記抽出した参照装着品の画像を前記学習済みモデルに入力し、前記学習済みモデルから出力された前記試着装着品の画像と、前記抽出したユーザの身体部分の画像とを合成することによって、前記試着装着品を試着したユーザの画像を生成する、
    請求項1に記載の画像処理装置。
  3.  前記画像生成部は、前記参照装着品を装着したユーザの画像から前記参照装着品の画像の深度情報と前記ユーザの身体部分の画像の深度情報とを抽出し、
     前記抽出した深度情報に基づいて、前記試着装着品の画像と前記ユーザの身体部分の画像とを合成する、
    請求項2に記載の画像処理装置。
  4.  前記学習用データは、所定の姿勢で装着主体が装着した前記参照装着品の画像を前記入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した第1のバリエーションの前記試着装着品の画像を前記出力データとして関連付け、かつ、前記入力データに対して、前記所定の姿勢と共通する姿勢で装着主体が装着した第2のバリエーションの前記試着装着品の画像を前記出力データとして関連付け、
     前記画像生成部は、
    前記試着装着品のバリエーションの選択を受け付け、
    前記参照装着品を装着したユーザの画像を前記学習済みモデルに入力して、前記選択されたバリエーションの前記試着装着品を試着したユーザの画像を生成する、
    請求項1から3のいずれか1項に記載の画像処理装置。
  5.  前記装着品は、衣類であり、
     前記学習用データは、所定の姿勢で装着主体が装着した前記参照衣類の画像を前記入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した前記試着衣類の画像を前記出力データとして関連付けたデータを含む、
    請求項1から4のいずれか1項に記載の画像処理装置。
  6. 所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データを用いて機械学習により学習させた、
    学習済みモデル。
  7.  ロボットマネキンと、
     カメラと、
    前記ロボットマネキンの姿勢を制御するロボット制御部と、
     前記ロボット制御部の制御に基づいて装着品を装着させた前記ロボットマネキンの姿勢を所定の姿勢に制御した状態で、前記カメラで前記装着品の画像を撮影し、撮影された画像を前記姿勢と関連付けて記憶部に記憶させる撮影制御部とを備える、
    画像収集装置。
  8. 所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データを用いた機械学習により学習させた学習済みモデルに対し、前記参照装着品を装着したユーザの画像を入力して、前記試着装着品を試着したユーザの画像を生成する画像生成ステップを含む、
    画像処理方法。
  9. コンピュータに、
    所定の姿勢で装着主体が装着した参照装着品の画像を入力データとし、前記所定の姿勢と共通する姿勢で装着主体が装着した試着装着品の画像を出力データとして関連付けた学習用データを用いた機械学習により学習させた学習済みモデルに対し、前記参照装着品を装着したユーザの画像を入力して、前記試着装着品を試着したユーザの画像を生成させる、
    処理を実行させる、
    画像処理プログラム。
PCT/JP2020/007392 2019-02-22 2020-02-25 画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラム Ceased WO2020171237A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2021502253A JP7497059B2 (ja) 2019-02-22 2020-02-25 画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラム

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201962809088P 2019-02-22 2019-02-22
US62/809,088 2019-02-22

Publications (1)

Publication Number Publication Date
WO2020171237A1 true WO2020171237A1 (ja) 2020-08-27

Family

ID=72144680

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2020/007392 Ceased WO2020171237A1 (ja) 2019-02-22 2020-02-25 画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラム

Country Status (2)

Country Link
JP (1) JP7497059B2 (ja)
WO (1) WO2020171237A1 (ja)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200380594A1 (en) * 2018-02-21 2020-12-03 Kabushiki Kaisha Toshiba Virtual try-on system, virtual try-on method, computer program product, and information processing device
WO2022161301A1 (zh) * 2021-01-28 2022-08-04 腾讯科技(深圳)有限公司 图像生成方法、装置、计算机设备及计算机可读存储介质

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2019163218A1 (ja) * 2018-02-21 2019-08-29 株式会社東芝 仮想試着システム、仮想試着方法、仮想試着プログラム、情報処理装置、および学習データ

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2019163218A1 (ja) * 2018-02-21 2019-08-29 株式会社東芝 仮想試着システム、仮想試着方法、仮想試着プログラム、情報処理装置、および学習データ

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
EHARA,JUN ET AL.: "Texture Overlay onto Deformable Surface for Virtual Clothing", IEICE TECHNICAL REPORT, vol. 105, no. 536, 2005, pages 129 - 134, XP058167000, DOI: 10.1145/1152399.1152431 *
YASUDA,TOMOMI ET AL.: "A Study on a Virtual Dressing Simulation System toward both Simpleness and Cloth Deformation", IEICE TECHNICAL REPORT, vol. 109, no. 471, 2010, pages 91 - 96 *

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200380594A1 (en) * 2018-02-21 2020-12-03 Kabushiki Kaisha Toshiba Virtual try-on system, virtual try-on method, computer program product, and information processing device
US12223541B2 (en) * 2018-02-21 2025-02-11 Kabushiki Kaisha Toshiba Virtual try-on system, virtual try-on method, computer program product, and information processing device
WO2022161301A1 (zh) * 2021-01-28 2022-08-04 腾讯科技(深圳)有限公司 图像生成方法、装置、计算机设备及计算机可读存储介质
US12293434B2 (en) 2021-01-28 2025-05-06 Tencent Technology (Shenzhen) Company Limited Image generation method and apparatus

Also Published As

Publication number Publication date
JP7497059B2 (ja) 2024-06-10
JPWO2020171237A1 (ja) 2021-12-16

Similar Documents

Publication Publication Date Title
US11741629B2 (en) Controlling display of model derived from captured image
US12017142B2 (en) System and method for real-time calibration of virtual apparel using stateful neural network inferences and interactive body measurements
CN104217350B (zh) 实现虚拟试戴的方法和装置
TWI488071B (zh) 非接觸式三度空間人體資料擷取系統及方法
EP3479296B1 (en) System of virtual dressing utilizing image processing, machine learning, and computer vision
Mueller et al. Real-time hand tracking under occlusion from an egocentric rgb-d sensor
KR101911133B1 (ko) 깊이 카메라를 이용한 아바타 구성
US20220188897A1 (en) Methods and systems for determining body measurements and providing clothing size recommendations
US8976230B1 (en) User interface and methods to adapt images for approximating torso dimensions to simulate the appearance of various states of dress
US12141916B2 (en) Markerless motion capture of hands with multiple pose estimation engines
Wang et al. Real time eye gaze tracking with kinect
JP6980097B2 (ja) サイズ測定システム
CN110892408A (zh) 用于立体视觉和跟踪的系统、方法和装置
JP6369811B2 (ja) 歩行解析システムおよび歩行解析プログラム
Gültepe et al. Real-time virtual fitting with body measurement and motion smoothing
CN103106604A (zh) 基于体感技术的3d虚拟试衣方法
Xu et al. 3d virtual garment modeling from rgb images
JP6008025B2 (ja) 画像処理装置、画像処理方法、およびプログラム
Jatesiktat et al. Anatomical-marker-driven 3D markerless human motion capture
CN111767817A (zh) 一种服饰搭配方法、装置、电子设备及存储介质
JP7497059B2 (ja) 画像処理装置、学習済みモデル、画像収集装置、画像処理方法、および、画像処理プログラム
Cha et al. Mobile. Egocentric human body motion reconstruction using only eyeglasses-mounted cameras and a few body-worn inertial sensors
Ileperuma et al. An enhanced virtual fitting room using deep neural networks
CN115293958A (zh) 衣服变形方法、虚拟试衣方法及相关装置
Ram et al. A review on virtual reality for 3D virtual trial room

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20760298

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021502253

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20760298

Country of ref document: EP

Kind code of ref document: A1