WO2015192733A1 - 实现虚拟试戴的方法和装置 - Google Patents

实现虚拟试戴的方法和装置 Download PDF

Info

Publication number
WO2015192733A1
WO2015192733A1 PCT/CN2015/081264 CN2015081264W WO2015192733A1 WO 2015192733 A1 WO2015192733 A1 WO 2015192733A1 CN 2015081264 W CN2015081264 W CN 2015081264W WO 2015192733 A1 WO2015192733 A1 WO 2015192733A1
Authority
WO
WIPO (PCT)
Prior art keywords
face
current frame
frame
feature points
image
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2015/081264
Other languages
English (en)
French (fr)
Inventor
张斯聪
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Jingdong Century Trading Co Ltd
Beijing Jingdong Shangke Information Technology Co Ltd
Original Assignee
Beijing Jingdong Century Trading Co Ltd
Beijing Jingdong Shangke Information Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Jingdong Century Trading Co Ltd, Beijing Jingdong Shangke Information Technology Co Ltd filed Critical Beijing Jingdong Century Trading Co Ltd
Priority to US15/319,500 priority Critical patent/US10360731B2/en
Publication of WO2015192733A1 publication Critical patent/WO2015192733A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T19/00Manipulating three-dimensional [3D] models or images for computer graphics
    • G06T19/006Mixed reality
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q30/00Commerce
    • G06Q30/06Buying, selling or leasing transactions
    • G06Q30/0601Electronic shopping [e-shopping]
    • G06Q30/0641Electronic shopping [e-shopping] utilising user interfaces specially adapted for shopping
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00Two-dimensional [2D] image generation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00Two-dimensional [2D] image generation
    • G06T11/60Creating or editing images; Combining images with text
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/42Global feature extraction by analysis of the whole pattern, e.g. using frequency domain transformations or autocorrelation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/161Detection; Localisation; Normalisation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/161Detection; Localisation; Normalisation
    • G06V40/167Detection; Localisation; Normalisation using comparisons between temporally consecutive images
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/168Feature extraction; Face representation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/174Facial expression recognition

Definitions

  • the present invention relates to computer technology, and in particular to a method and apparatus for implementing virtual try-on.
  • This method adopts the effect of wearing a virtual product on a pre-generated human body or a part of a human body to give the user a virtual try-on. This method does not have the actual physical information of the user, and the trial wear effect is not good.
  • the method utilizes a special device such as a depth of field sensor to collect actual physical information of the user, and forms a model of the human body or a human body for the user to try on.
  • a special device such as a depth of field sensor
  • This method obtains the actual physical information of the user, it requires special equipment, and is usually available at a special place provided by the merchant.
  • a typical user only has a conventional image capture device such as a camera that is placed on a mobile phone or on a computer.
  • the present invention provides a method and apparatus for implementing virtual try-on, which enables a user to implement virtual try-on using an ordinary image capturing device such as a camera on a mobile phone or a computer.
  • a method of implementing virtual try-on is provided.
  • the method for realizing virtual try-on of the present invention comprises: performing face detection on the collected initial frame, and in the case of detecting a face, generating an item image at an initial position and then superimposing with the initial frame, and outputting the initial position and The specified position of the face in the initial frame overlaps; the face pose detection of the face in the current frame results in the face pose of the current frame; and the item is generated again according to the current position of the item image and the face pose And displaying an image of the object in the image of the object in accordance with the face pose, and then superimposing the image of the article on the current frame and outputting the image.
  • the step of performing face pose detection on a face in the current frame to obtain a face pose of the current frame comprises: determining a plurality of feature points on the face image in the initial frame; for each feature The point is processed as follows: tracking the feature point to determine the position of the feature point in the current frame, and performing affine transformation on the neighborhood of the feature point in the initial frame according to the face pose of the previous frame to obtain the point a projection area of the neighborhood in the current frame, calculating a color offset between the neighborhood in the initial frame and the projection area in the current frame as a tracking deviation of the feature point, for the determined location A plurality of feature points are selected, and a plurality of feature points having a small tracking deviation are selected; and a plurality of feature points having a small tracking deviation are determined according to the position of the initial frame and the position of the current frame to determine a face pose of the current frame.
  • the step of selecting a plurality of feature points with a small tracking deviation for the plurality of feature points includes: using the maximum value and the minimum value as the tracking deviations of the determined plurality of feature points In the initial center, clustering is performed according to the size of the tracking deviation to obtain two types; the corresponding feature points of the two types with less tracking deviation are selected.
  • the method further includes: projecting a feature point corresponding to a class with a large tracking deviation in the two types according to a face pose of the current frame to a current
  • the frame image plane replaces the position of these feature points at the current frame with the projected position.
  • the method before the step of performing face detection on the collected initial frame, the method further includes: when the reset instruction is received, using the collected current frame as the initial frame; After the clustering obtains two types of steps, the method further includes: if the number of the types of feature points with the small tracking deviation is less than the first preset value of the total number of feature points, or is collected in the current frame If the number of feature points occupies less than the second preset value in the total number of feature points collected in the previous frame, the prompt information is output, and then the reset command is received.
  • the item image is a glasses image, a head ornament image, or a neck jewelry image.
  • an apparatus for implementing virtual trial wear is provided.
  • the device for implementing the virtual try-on of the present invention includes: a face detection module, configured to perform face detection on the collected initial frame; and a first output module, configured to: when the face detection module collects a face, Generating an item image at an initial position and then superimposing it with the initial frame, the initial position overlapping with a specified position of a face in the initial frame; a face pose detection module for performing a person on a face in the current frame The face pose detection obtains a face pose of the current frame; the second output module is configured to regenerate the item image according to the current position of the item image and the face pose, and make the item pose in the item image and the person The face poses the same, and then the item image is superimposed on the current frame and output.
  • the face pose detection module is further configured to: determine, in the initial frame, a plurality of feature points on the face image; perform, for each feature point, a process of: tracking the feature points to determine the feature Pointing at the position of the current frame, performing affine transformation on the neighborhood of the feature point in the initial frame according to the face pose of the previous frame to obtain a projection area of the neighborhood in the current frame, and calculating the initial Selecting a color offset between the neighborhood in the frame and the projection area in the current frame as a tracking deviation of the feature point, and selecting a plurality of tracking deviations for the determined plurality of feature points Feature point; according to the tracking deviation
  • a small plurality of feature points determine the face pose of the current frame at the position of the initial frame and at the position of the current frame.
  • the face pose detection module is further configured to: use the maximum value and the minimum value as the initial center for the determined tracking deviation of the plurality of feature points, and perform clustering according to the size of the tracking deviation to obtain two Class; select the feature points corresponding to a class with a small tracking deviation in the two classes.
  • the method further includes: a modifying module, after the face pose detecting module determines the face pose of the current frame, according to the current feature point of the two types with a large tracking deviation
  • the face pose of the frame is projected to the current frame image plane, and the position of the feature point at the current frame is replaced by the projected position.
  • the method further includes: a reset module and a prompting module, wherein: the reset module is configured to receive a reset instruction, and, in the case that the reset command is received, use the collected current frame as the initial frame; After the face pose detection module performs clustering according to the size of the tracking deviation to obtain two types, the number of the types of feature points whose tracking deviation is small is less than the first preset value. If the number of feature points collected in the current frame occupies less than the second preset value in the total number of feature points collected in the previous frame, the prompt information is output.
  • the item image is a glasses image, a head ornament image, or a neck jewelry image.
  • the user by detecting the face pose of each frame and then adjusting the posture of the glasses according to the face pose, the user can complete the virtual try-on using the ordinary image acquisition device, and the user may turn the head to observe more.
  • the wearing effect of the angle has a relatively high authenticity.
  • FIG. 1 is a schematic diagram of the basic steps of a method for implementing virtual try-on according to an embodiment of the present invention
  • FIG. 2 is a schematic diagram of main steps of face pose detection according to an embodiment of the present invention.
  • FIG. 3 is a schematic diagram of collected feature points according to an embodiment of the present invention.
  • 4A and 4B are schematic diagrams showing texture regions in an initial frame and in a current frame, respectively, according to an embodiment of the present invention
  • FIG. 5 is a schematic diagram of a basic structure of an apparatus for implementing virtual try-on according to an embodiment of the present invention.
  • the virtual try-on technology of the embodiment of the present invention can be applied to a mobile phone with a camera, or to a computer connected or built-in camera, including a tablet computer.
  • Trial of glasses, accessories and other items can be achieved.
  • the trial glasses are taken as an example for illustration.
  • the user selects the glasses to try on, and points the camera at his face, clicks on the screen or a designated button, at which point the camera captures the user's avatar and presents the glasses at the eyes of the user's avatar.
  • the user can click on the glasses in the screen and translate them to further adjust their positional relationship with the eyes.
  • the user can turn the neck up and down or left and right to see the wearing effect of the glasses at various angles.
  • the technique of the present embodiment is applied to keep the posture of the glasses in the glasses image on the screen consistent with the posture of the face, thereby enabling the glasses to track the movement of the face to achieve that the glasses are fixedly worn on the face.
  • FIG. 1 is a schematic diagram of the basic steps of a method for implementing virtual try-on according to an embodiment of the present invention. As shown in FIG. 1, the method mainly includes the following steps S11 to S17.
  • Step S11 Acquire an initial frame. It may be that the acquisition is automatically started when the camera is activated or the acquisition is started according to the user's operation instruction. For example, the user clicks on the touch screen or presses any or a designated button on the keyboard.
  • Step S12 Perform face detection on the initial frame.
  • the existing face detection methods can be used to confirm that the initial frame contains a face and determine the approximate range of the face. This approximate range can be represented by a circumscribed rectangle of the face.
  • Step S13 Generate a glasses image and superimpose it with the initial frame.
  • the image of which glasses is specifically generated is selected by the user. For example, the user clicks on one of a plurality of glasses icons that appear in the screen.
  • the point at which the total length of the face range is 0.3 to 0.35:1 from the upper end of the face range is set in advance as the eye position.
  • the initial position of the glasses image is overlapped with the set eye position. The user can fine tune the glasses presented on the person's face by dragging the glasses image.
  • Step S14 collecting the current frame.
  • Step S15 Perform face pose detection on the face in the current frame.
  • the face pose can be implemented by various existing face pose (or face pose) detection techniques.
  • the face pose can be determined by using the rotation parameter R(r0, r1, r2) together with the translation parameter T(t0, t1, t2).
  • the rotation parameter and the translation parameter respectively represent the rotation angle of one plane on three coordinate planes and the translation length on three coordinate axes with respect to the initial position in the spatial Cartesian coordinate system.
  • the initial position of the face image is the position of the face image in the initial frame, so that for each current frame, it is compared with the initial frame to obtain a face pose in the current frame, that is, The above rotation parameters and translation parameters. That is, the face pose of each frame after the initial frame is a pose formed with respect to the face pose in the initial frame.
  • Step S16 The glasses image is again generated based on the current position of the glasses image and the face pose detected in step S15.
  • Step S17 The glasses image generated in step S16 is superimposed with the current frame and then output.
  • the glasses image output at this time is already located near the eyes of the person's face in the current frame because the processing of step S16 is performed. Up to this step, the glasses image has been superimposed on the current frame. For each frame collected thereafter, the same process is followed, that is, the process returns to step S14.
  • the glasses image is superimposed on the current frame, the user can see the state as shown in FIG.
  • the portrait 30 captured in the black and white single line replaces the portrait captured by the actual camera.
  • the person wears glasses 32.
  • This program can not only achieve eyeglasses try-on, but also try on earrings, necklaces and other accessories. For a try-on necklace, the face must be included in the neck.
  • FIG. 2 is a schematic diagram of the main steps of face pose detection in accordance with an embodiment of the present invention. As shown in FIG. 2, the method mainly includes the following steps S20 to S29.
  • Step S20 determining a plurality of feature points on the face image in the initial frame. Since feature point tracking is to be performed in subsequent steps, the selection of feature points in this step is considered to facilitate tracking. You can select points with rich textures or points with large color gradients. Such points are more easily recognized when the position of the face changes. Can refer to the following documents:
  • FIG. 3 is a schematic diagram of acquired feature points in accordance with an embodiment of the present invention.
  • a plurality of small circles, such as circle 31, in FIG. 3 represent acquired feature points.
  • the deviation of the texture region is determined for each feature point, which is actually the tracking error of the feature point.
  • Step S21 Take one feature point as the current feature point. Feature points can be numbered, each time in numerical order. From step S22 to step S24, processing of one feature point is performed.
  • Step S22 Tracking the current feature point to determine the location of the feature point in the current frame.
  • Various existing feature point tracking methods can be used, such as optical flow tracking, template matching, particle filtering, and feature point detection.
  • the optical flow tracking method can adopt the Lucas & Kanade method.
  • the tracking of feature points is improved. For each feature point, compare the difference between its range of neighborhoods (called texture regions in the description of subsequent steps) and the neighborhood of its corresponding range at the current frame to determine whether the feature points are tracked. Be accurate. That is, the processing method in the next step.
  • Step S23 Perform affine transformation on the neighborhood of the feature point in the initial frame according to the face pose of the previous frame to obtain a projection area of the neighborhood in the current frame. Since the face between two frames inevitably has more or less rotation, it is preferable to perform affine transformation to make the partial regions of the two frames comparable.
  • FIG. 4A and FIG. 4B are respectively schematic diagrams of taking texture regions in an initial frame and in a current frame, according to an embodiment of the present invention. Generally, a rectangular area centered on a feature point is used as a texture area of the feature point. As shown in FIGS.
  • the texture point in the feature point 45 (the white point in the figure) in the initial frame 41 (the portion showing the frame in the figure) is the rectangle 42
  • the texture area in the current frame 43 is the trapezoid 44 . This is because the face is rotated to the left by a certain angle to the current frame. If the texture area is still taken around the feature point 45 in Fig. 4B according to the size of the rectangle 42, an oversized range of pixels is collected even in the frame. In other cases, a background image is acquired. Therefore, it is better to perform an affine transformation to project the texture region of the feature point in the initial frame to the current frame plane. The texture points of the points in different frames are comparable. Thus, the feature point in the texture area of the current frame is actually the above-described projection area.
  • Step S24 Calculate a color shift amount between the texture region of the current feature point in the initial frame and the projection region of the texture region in the current frame.
  • the color offset is the tracking deviation of the feature point.
  • the gray value of each pixel in the texture region of the current feature point in the initial frame is connected into a vector by the row or column of the pixel, and the length of the vector is the total number of pixels of the texture region; Pixels of the above-mentioned projection area are connected in rows or columns, and then equally divided according to the total number, and the gray value of each of the cells obtained by the equal division takes the gray value of the relatively large pixel, and the gray value of the possession Connected to another vector whose length is equal to the total number described above.
  • Calculating the distance between the two vectors yields a value that reflects the tracking deviation of the feature points. Since only the tracking deviation is required, the gray value is shorter than the vector obtained by using the RGB value, which helps to reduce the amount of calculation.
  • the vector distance here can be expressed by Euclidean distance, Mahalanobis distance, cosine distance, related system, and the like. After this step, the process proceeds to step S25.
  • Step S25 It is judged whether all the feature points have been processed. If yes, go to step S26, otherwise go back to step S21.
  • Step S26 The tracking deviations of all the feature points are grouped into two types according to the size. Any self-clustering method can be used, such as K-means self-clustering method. In the calculation, the maximum and minimum values of the tracking deviation of all feature points are taken as the initial center, and the clustering is classified into two types: the tracking deviation is larger and smaller.
  • Step S27 According to the clustering result of step S26, a type of feature point with a small tracking deviation is taken as an effective feature point. Accordingly, other feature points are used as invalid feature points.
  • Step S28 Calculate a coordinate transformation relationship of the effective feature point from the initial frame to the current frame.
  • This coordinate transformation relationship is represented by a matrix P.
  • Various existing algorithms can be used, such as the Levenberg-Marquardt algorithm, reference: Z. Zhang. "A flexible new technique For camera calibration". IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(11): 1330-1334, 2000. Reference may also be made to the algorithm in the following documents:
  • Step S29 The face pose of the current frame is obtained according to the coordinate transformation relationship in step S28 and the face pose in the initial frame. That is, the rotation parameter Rn and the translation parameter Tn of the current frame (nth frame) are calculated according to the above-described matrix P and the above-described rotation parameter R and translation parameter T.
  • the above illustrates a calculation method of the face pose in the current frame.
  • Other face pose detection algorithms can also be used in the implementation to obtain the face pose in the current frame.
  • the invalid feature points described above can be corrected using the face pose in the current frame. That is, the new coordinates of the invalid feature points are calculated according to the rotation parameter Rn and the translation parameter Tn described above and the coordinates of the invalid feature points in the initial frame, and the new coordinates are replaced with their coordinates in the current frame.
  • the coordinates of all feature points in the replaced current frame will be used for data processing of the next frame. This helps to improve the accuracy of the next frame processing. It is also possible to use only the valid feature values in the current frame for the processing of the next frame, but this will reduce the amount of data available.
  • the image of the glasses is superimposed on each frame, so that the user can still see that the glasses are "wearing" on the face while turning the head. If the user's head movement is severe, causing the posture to change too much, especially in the case of insufficient light, it is difficult to accurately track the feature points, and the glasses in the screen will also be separated from the position of the eyes. In this case, the user can be prompted to perform a reset operation. For example, clicking the screen or the specified key again, the camera collects the user's avatar and presents the glasses at the eyes of the user's avatar.
  • the user's operation issues a reset command
  • the mobile phone or computer receives the current frame captured by the camera as the initial frame and processes it as described above.
  • the processing result of the cluster is obtained in step S27, and it can be judged if The ratio of the effective feature points is less than a set value, for example, 60%, or the proportion of the feature points collected by the feature points collected in the frame is less than a set value, for example, 30%, and the prompt information is output.
  • the text "Click the screen to reset" prompts the user to "try on” the glasses again.
  • FIG. 5 is a schematic diagram of a basic structure of an apparatus for implementing virtual try-on according to an embodiment of the present invention.
  • the device can be set as software in a mobile phone or a computer.
  • the device 50 for implementing virtual trialing mainly includes a face detecting module 51, a first output module 52, a face pose detecting module 53, and a second output module 54.
  • the face detection module 51 is configured to perform face detection on the collected initial frame.
  • the first output module 52 is configured to generate an item image at the initial position and then the initial frame in the case that the face detection module 51 collects the face. After being superimposed, the initial position overlaps with the specified position of the face in the initial frame;
  • the face pose detection module 53 is configured to perform face pose detection on the face in the current frame to obtain a face pose of the current frame;
  • the module 54 is configured to regenerate the item image according to the current position and the face posture of the item image, and make the item posture in the item image conform to the face posture, and then superimpose the item image with the current frame and output.
  • the face pose detection module 53 is further configured to: determine a plurality of feature points on the face image in the initial frame; perform processing for each feature point: tracking the feature point to determine the position of the feature point in the current frame, According to the face pose of the previous frame, the neighborhood of the feature point in the initial frame is affine transformed to obtain a projection area of the neighborhood in the current frame, and the projection in the neighborhood and the current frame in the initial frame is calculated. a color offset between the regions as a tracking deviation of the feature point; for the determined plurality of feature points, selecting a plurality of feature points with a small tracking deviation; and a plurality of feature points having a smaller tracking deviation in the initial frame The position and the position of the current frame determine the face pose of the current frame.
  • the face pose detection module 53 is further configured to: use the maximum value and the minimum value as the initial center for the determined tracking deviation of the plurality of feature points, and perform clustering according to the size of the tracking deviation to obtain two types; Tracking the corresponding feature points of a class with a small deviation.
  • the device 50 for implementing the virtual try-on can further include a modification module (not shown) for using the face pose detection module to determine the face pose of the current frame, and the tracking deviation between the two types is large.
  • a class of corresponding feature points are projected to the current frame image plane according to the face pose of the current frame, and the position of the feature points at the current frame is replaced by the projected position.
  • the apparatus 50 for implementing virtual trial-wearing may further include a reset module and a prompting module (not shown), wherein: the reset module is configured to receive the reset instruction, and in the case of receiving the reset instruction, the current frame acquired is taken as an initial a frame; the prompting module is configured to perform clustering according to the size of the tracking deviation in the face pose detecting module to obtain two types, and the number of the type of feature points having a small tracking deviation occupies the total number of the feature points is greater than the first preset value. In the case, or when the number of feature points collected in the current frame occupies less than the second preset value in the total number of feature points, the prompt information is output.
  • the reset module is configured to receive the reset instruction, and in the case of receiving the reset instruction, the current frame acquired is taken as an initial a frame
  • the prompting module is configured to perform clustering according to the size of the tracking deviation in the face pose detecting module to obtain two types, and the number of the type of feature points having a small tracking deviation occupies the total number
  • the user by detecting the face pose of each frame and then adjusting the eye gesture according to the face pose, the user can complete the virtual try-on using the ordinary image capture device, and the user may turn the head to Observing the wearing effect of multiple angles, with higher authenticity.
  • the objects of the invention can also be achieved by running a program or a set of programs on any computing device.
  • the computing device can be a well-known general purpose device.
  • the object of the present invention can also be achieved by merely providing a program product comprising program code for implementing the method or apparatus. That is to say, such a program product also constitutes the present invention.
  • a storage medium storing such a program product also constitutes the present invention. It will be apparent that the storage medium may be any known storage medium or any storage medium developed in the future.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Multimedia (AREA)
  • Business, Economics & Management (AREA)
  • Health & Medical Sciences (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • General Health & Medical Sciences (AREA)
  • Human Computer Interaction (AREA)
  • Finance (AREA)
  • Accounting & Taxation (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Strategic Management (AREA)
  • General Business, Economics & Management (AREA)
  • Development Economics (AREA)
  • Marketing (AREA)
  • Economics (AREA)
  • Computer Graphics (AREA)
  • Computer Hardware Design (AREA)
  • General Engineering & Computer Science (AREA)
  • Software Systems (AREA)
  • Image Analysis (AREA)

Abstract

一种实现虚拟试戴的方法和装置,能够使用户利用普通的图像采集装置例如手机上的或者计算机上的摄像头即可实现虚拟试戴。实现虚拟试载的方法包括:对采集的初始帧进行人脸检测(S12),在检测到人脸的情况下,在初始位置生成物品图像然后与所述初始帧叠加后输出(S13),该初始位置与所述初始帧中的人脸的指定位置重叠;对当前帧中的人脸进行人脸姿势检测得到当前帧的人脸姿势(S15);根据所述物品图像的当前位置和所述人脸姿势再次生成物品图像(S16),并使该物品图像中的物品姿势与所述人脸姿势一致,然后将该物品图像与所述当前帧叠加后输出(S17)。

Description

实现虚拟试戴的方法和装置 技术领域
本发明涉及计算机技术,特别地涉及一种实现虚拟试戴的方法和装置。
背景技术
随着电子商务的发展,网上购物成为越来越多的用户的选择。服饰作为主要的消费品之一,也成为很多用户网购的目标。购买服饰时通常需要试穿试戴,于是虚拟试衣以及虚拟试戴技术应运而生。
目前的虚拟试戴技术主要有两类实现途径:
1、人工合成模型试戴
这种方法采用将虚拟商品佩戴在预先生成的人体或人体局部的模型上,给用户虚拟试戴的效果。这种方式没有用户的实际身体信息,试戴效果不佳。
2、特殊装置采集真实人体信息试戴
该方法利用特殊装置例如景深传感器采集用户实际身体信息,形成人体或人体局部的模型,供用户试戴。这种方式虽然获得了用户实际的身体信息,但是需要特殊装置,目前通常在商家提供的专门场所才具备。一般用户仅具有普通的图像采集装置例如设置在手机上的或计算机上的摄像头。
发明内容
有鉴于此,本发明提供一种实现虚拟试戴的方法和装置,能够使用户利用普通的图像采集装置例如手机上的或者计算机上的摄像头即可实现虚拟试戴。
为实现上述目的,根据本发明的一个方面,提供了一种实现虚拟试戴的方法。
本发明的实现虚拟试戴的方法包括:对采集的初始帧进行人脸检测,在检测到人脸的情况下,在初始位置生成物品图像然后与所述初始帧叠加后输出,该初始位置与所述初始帧中的人脸的指定位置重叠;对当前帧中的人脸进行人脸姿势检测得到当前帧的人脸姿势;根据所述物品图像的当前位置和所述人脸姿势再次生成物品图像,并使该物品图像中的物品姿势与所述人脸姿势一致,然后将该物品图像与所述当前帧叠加后输出。
可选地,所述对当前帧中的人脸进行人脸姿势检测得到当前帧的人脸姿势的步骤包括:在所述初始帧中确定人脸图像上的多个特征点;针对每个特征点进行如下处理:对特征点进行跟踪以确定该特征点在当前帧的位置,根据前一帧的人脸姿势,将所述初始帧中的该特征点的邻域进行仿射变换以得到该邻域在当前帧中的投影区域,计算所述初始帧中的所述邻域与当前帧中的所述投影区域之间的颜色偏移量并作为该特征点的跟踪偏差,对于确定的所述多个特征点,选择跟踪偏差较小的多个特征点;根据所述跟踪偏差较小的多个特征点在所述初始帧的位置以及在当前帧的位置确定当前帧的人脸姿势。
可选地,所述对于所述多个特征点,选择跟踪偏差较小的多个特征点的步骤包括:对于确定的所述多个特征点的跟踪偏差,以其中的最大值和最小值作为初始中心,按跟踪偏差的大小进行聚类得到两类;选择所述两类中跟踪偏差较小的一类所对应的特征点。
可选地,所述确定当前帧的人脸姿势的步骤之后,还包括:将所述两类中跟踪偏差较大的一类所对应的特征点按照所述当前帧的人脸姿势投影到当前帧图像平面,以投影位置代替这些特征点在当前帧的位置。
可选地,所述对采集的初始帧进行人脸检测的步骤之前,还包括:在接收到复位指令的情况下,将采集的当前帧作为所述初始帧;所述按跟踪偏差的大小进行聚类得到两类的步骤之后,还包括:在所述跟踪偏差较小的一类特征点的数目占特征点总数目的比例小于第一预设值的情况下,或者,在当前帧中采集到的特征点的数目占在前一帧采集到的特征点总数目的比例小于第二预设值的情况下,输出提示信息,然后接收复位指令。
可选地,所述物品图像为眼镜图像、头部饰品图像、或者颈部饰品图像。
根据本发明的另一方面,提供了一种实现虚拟试戴的装置。
本发明的实现虚拟试戴的装置包括:人脸检测模块,用于对采集的初始帧进行人脸检测;第一输出模块,用于在所述人脸检测模块采集到人脸的情况下,在初始位置生成物品图像然后与所述初始帧叠加后输出,该初始位置与所述初始帧中的人脸的指定位置重叠;人脸姿势检测模块,用于对当前帧中的人脸进行人脸姿势检测得到当前帧的人脸姿势;第二输出模块,用于根据所述物品图像的当前位置和所述人脸姿势再次生成物品图像,并使该物品图像中的物品姿势与所述人脸姿势一致,然后将该物品图像与所述当前帧叠加后输出。
可选地,所述人脸姿势检测模块还用于:在所述初始帧中确定人脸图像上的多个特征点;针对每个特征点进行如下处理:对特征点进行跟踪以确定该特征点在当前帧的位置,根据前一帧的人脸姿势,将所述初始帧中的该特征点的邻域进行仿射变换以得到该邻域在当前帧中的投影区域,计算所述初始帧中的所述邻域与当前帧中的所述投影区域之间的颜色偏移量并作为该特征点的跟踪偏差,对于确定的所述多个特征点,选择跟踪偏差较小的多个特征点;根据所述跟踪偏差较 小的多个特征点在所述初始帧的位置以及在当前帧的位置确定当前帧的人脸姿势。
可选地,所述人脸姿势检测模块还用于:对于确定的所述多个特征点的跟踪偏差,以其中的最大值和最小值作为初始中心,按跟踪偏差的大小进行聚类得到两类;选择所述两类中跟踪偏差较小的一类所对应的特征点。
可选地,还包括修改模块,用于在所述人脸姿势检测模块确定当前帧的人脸姿势之后,将所述两类中跟踪偏差较大的一类所对应的特征点按照所述当前帧的人脸姿势投影到当前帧图像平面,以投影位置代替这些特征点在当前帧的位置。
可选地,还包括复位模块和提示模块,其中:所述复位模块用于接收复位指令,以及在接收到复位指令的情况下,将采集的当前帧作为所述初始帧;所述提示模块用于在所述人脸姿势检测模块按跟踪偏差的大小进行聚类得到两类之后,在所述跟踪偏差较小的一类特征点的数目占特征点总数目的比例小于第一预设值的情况下,或者,在当前帧中采集到的特征点的数目占在前一帧采集到的特征点总数目的比例小于第二预设值的情况下,输出提示信息。
可选地,所述物品图像为眼镜图像、头部饰品图像、或者颈部饰品图像。
根据本发明的技术方案,通过检测各帧的人脸姿势,再按人脸姿势调整眼镜姿势,能够使用户利用普通的图像采集装置就可完成虚拟试戴,并且用户可能转动头部以观察多个角度的佩戴效果,具有比较高的真实性。
附图说明
附图用于更好地理解本发明,不构成对本发明的不当限定。其中:
图1是根据本发明实施例的实现虚拟试戴的方法的基本步骤的示意图;
图2是根据本发明实施例的人脸姿势检测的主要步骤的示意图;
图3是根据本发明实施例的采集到的特征点的示意图;
图4A和图4B分别是根据本发明实施例的在初始帧中和在当前帧中取纹理区域的示意图;
图5是根据本发明实施例的实现虚拟试戴的装置的基本结构的示意图。
具体实施方式
以下结合附图对本发明的示范性实施例做出说明,其中包括本发明实施例的各种细节以助于理解,应当将它们认为仅仅是示范性的。因此,本领域普通技术人员应当认识到,可以对这里描述的实施例做出各种改变和修改,而不会背离本发明的范围和精神。同样,为了清楚和简明,以下的描述中省略了对公知功能和结构的描述。
本发明实施例的虚拟试戴技术可应用于具有摄像头的手机,或者应用于连接或内置摄像头的计算机,包括平板电脑。可实现眼镜、饰品等物品的试戴。本实施例中,以试戴眼镜为例加以说明。使用时,用户选择试戴的眼镜,并将摄像头对准自己的面部,单击屏幕或指定的键,此时摄像头采集用户头像并将眼镜呈现在用户头像的眼睛处。用户可以点击屏幕中的眼镜并将其平移以进一步调整其与眼睛的位置关系。用户可以上下或左右转动颈部以观看各个角度的眼镜佩戴效果。在该过程中,应用本实施例的技术,使屏幕上的眼镜图像中的眼镜的姿势与人脸的姿势保持一致,从而使眼镜能够跟踪人脸移动,以实现眼镜固定地佩戴在脸部。以下对本发明实施例的技术方案做出说明。
图1是根据本发明实施例的实现虚拟试戴的方法的基本步骤的示意图。如图1所示,该方法主要包括如下的步骤S11至步骤S17。
步骤S11:采集初始帧。可以是在摄像头启动的情况下自动开始采集或者根据用户的操作指令开始采集。例如用户点击触摸屏,或者按下键盘上的任意或指定按钮。
步骤S12:对初始帧进行人脸检测。可采用现有的各种人脸检测方式,确认初始帧中包含人脸并确定人脸的大致范围。该大致范围可用人脸的外接矩形来表示。
步骤S13:生成眼镜图像并与初始帧叠加。具体生成哪个眼镜的图像,由用户进行选择。例如用户点击屏幕中出现的多个眼镜图标中的一个。本实施例中,预先设定从人脸范围的上端起占人脸范围上下总长的0.3~0.35∶1处的分点为眼睛位置。在本步骤中,将眼镜图像与初始帧叠加时,要使眼镜图像的初始位置与设定的眼睛位置重叠。用户可以通过拖动眼镜图像对呈现在人脸上的眼镜进行微调。
步骤S14:采集当前帧。
步骤S15:对当前帧中的人脸进行人脸姿势检测。人脸姿势可采用现有的各种人脸姿势(或称人脸姿态)检测技术来实现。人脸姿势可以用旋转参数R(r0,r1,r2)与平移参数T(t0,t1,t2)共同确定。旋转参数与平移参数分别表示在空间直角坐标系中,相对于初始位置,一个平面在三个坐标平面上的旋转角度以及在三个坐标轴上的平移长度。在本实施例中,人脸图像的初始位置是初始帧中人脸图像的位置,这样,对于每个当前帧,是将其与初始帧进行比较而得出当前帧中的人脸姿势,即上述旋转参数和平移参数。即初始帧之后每一帧的人脸姿势是相对于初始帧中的人脸姿势而言形成的姿势。
步骤S16:根据眼镜图像的当前位置和步骤S15中检测到的人脸姿势再次生成眼镜图像。在本步骤中,需使眼镜图像中的眼镜姿势与人 脸姿势一致。因此要以眼镜图像的当前位置为起始位置,按照人脸姿势的旋转参数和平移参数来确定眼镜图像中的眼镜的旋转末值和平移末值然后据此生成眼镜图像。
步骤S17:将步骤S16中生成的眼镜图像与当前帧叠加然后输出。此时输出的眼镜图像因为经过步骤S16的处理,已经位于当前帧中的人脸上的眼睛附近。至本步骤,当前帧上已经叠加了眼镜图像。对于此后采集的每一帧,同样按上述流程处理,即返回步骤S14。
在当前帧上叠加了眼镜图像的情况下,用户即可看到如图3所示的状态。为示意清晰,图中以黑白单线的人像30代替实际摄像头采集的人像。该人像佩戴有眼镜32。本方案不仅可实现眼镜试戴,还可实现耳环、项链等饰品的试戴。对于试戴项链来说,采集的人脸需包括其颈部。
以下结合图2,对本实施例中采用的人脸姿势检测的方式加以说明。图2是根据本发明实施例的人脸姿势检测的主要步骤的示意图。如图2所示,该方法主要包括如下的步骤S20至步骤S29。
步骤S20:在初始帧中确定人脸图像上的多个特征点。因为在后续的步骤中要进行特征点跟踪,因此本步骤中对于特征点的选择要考虑其便于跟踪。可以选择周围纹理丰富的点或者颜色梯度较大的点,这样的点在人脸位置发生变化时仍比较便于被识别。可参考如下文献:
Jean-Yves Bouguet,“Pyramidal Implementation of the Lucas Kanade Feature Tracker Description of the algorithm”,Technical report,Microprocessor Research Labs,Intel Corporation(1999);
Jianbo Shi Carlo Tomasi,“Good features to track”,Proc.IEEE Comput.Soc.Conf.Comput.Vision and Pattern Recogn.,pages593-600,1994。
采集到的特征点如图3所示。图3是根据本发明实施例的采集到的特征点的示意图。图3中的多个小圆圈例如圆圈31表示采集到的特征点。接下来对每个特征点确定其纹理区域的偏差,该偏差实际上即为该特征点的跟踪误差。
步骤S21:取1个特征点作为当前特征点。可以对特征点进行编号,每次按编号顺序来取。从步骤S22至步骤S24,是对一个特征点的处理。
步骤S22:对当前特征点进行跟踪以确定该特征点在当前帧的位置。可采用现有的各种特征点跟踪方法,例如光流跟踪法、模板匹配法、粒子滤波法、特征点检测法等方式。其中光流跟踪法可采用Lucas&Kanade方法。对于特征点跟踪的各类算法,在应用中都存在一定误差,难以保证所有特征点都能被准确地在新的一帧中被定位,所以本实施例中,对特征点的跟踪作出改进,对于每个特征点,比较其在初始帧时其一定范围的邻域(在后续步骤的描述中称作纹理区域)与其在当前帧时相应范围的邻域的差异来确定该特征点是否被跟踪得准确。即接下来的步骤中的处理方式。
步骤S23:根据前一帧的人脸姿势,将初始帧中的该特征点的邻域进行仿射变换以得到该邻域在当前帧中的投影区域。因为两帧之间的人脸不可避免地存在或多或少的旋转,所以最好是进行仿射变换以使两帧的局部区域具有可比性。参考图4A和图4B,图4A和图4B分别是根据本发明实施例的在初始帧中和在当前帧中取纹理区域的示意图。一般是以特征点为中心的矩形区域作为该特征点的纹理区域。如图4A和图4B所示,特征点45(图中白色点)在初始帧41(图中示出帧的局部)中的纹理区域为矩形42,在当前帧43中的纹理区域为梯形44。这是因为到当前帧时,人脸向左转过了一定角度,如果仍按矩形42的大小在图4B中的特征点45周围取纹理区域,就会采集到过大的范围的像素甚至在另外一些情况下会采集到背景图像。所以最好是作仿射变换,将特征点在初始帧中的纹理区域投影到当前帧平面,使特 征点在不同帧中的纹理区域具有可比性。这样,特征点在当前帧的纹理区域实际上是上述投影区域。
步骤S24:计算当前特征点在初始帧中的纹理区域与该纹理区域在当前帧中的投影区域之间的颜色偏移量。该颜色偏移量即为对该特征点的跟踪偏差。在计算时,将当前特征点在初始帧中的纹理区域中的各个像素点的灰度值按像素的行或列连接成一个向量,该向量长度即为该纹理区域的像素点总数量;另将上述投影区域的像素按行或列连接,再按该总数量进行等分,在等分得到的每一格的灰度值取占比较大的像素的灰度值,所有格的灰度值连接成另一个向量,其长度等于上述的总数量。计算这两个向量的距离得到一个数值,该数值的大小即体现特征点的跟踪偏差。因为仅需得到跟踪偏差,所以采用灰度值比采用RGB值得到的向量更短,有助于减少计算量。这里的向量距离可采用欧氏距离、马氏距离、余弦距离、相关系统等来表示。本步骤之后进入步骤S25。
步骤S25:判断所有特征点是否都已处理。若是,则进入步骤S26,否则返回步骤S21。
步骤S26:对所有特征点的跟踪偏差按大小聚为两类。可采用任意一种自聚类方法来实现,例如K均值自聚类方法。计算时以所有特征点的跟踪偏差的最大值和最小值作为初始中心从而聚类为跟踪偏差较大和较小两类。
步骤S27:根据步骤S26的聚类结果,取跟踪偏差较小的一类特征点作为有效特征点。相应地,其他特征点作为无效特征点。
步骤S28:计算有效特征点从初始帧到当前帧的坐标变换关系。该坐标变换关系由一个矩阵P表示。可采用现有的各种算法,例如Levenberg-Marquardt算法,可参考:Z.Zhang.″A flexible new technique  for camera calibration″.IEEE Transactions on Pattern Analysis and Machine Intelligence,22(11):1330-1334,2000.还可以参考如下文献中的算法:
F.Moreno-Noguer,V.Lepetit and P.Fua″EPnP:Efficient Perspective-n-Point Camera Pose Estimation″
X.S.Gao,X.-R.Hou,J.Tang,H.-F.Chang;″Complete Solution Classification for the Perspective-Three-Point Problem″
步骤S29:根据步骤S28中的坐标变换关系和初始帧中的人脸姿势得出当前帧的人脸姿势。即按上述的矩阵P和上述的旋转参数R、平移参数T计算得出当前帧(第n帧)的旋转参数Rn和平移参数Tn。
以上说明了当前帧中的人脸姿势的一种计算方式。在实现中还可以采用其他的人脸姿势检测算法来得到当前帧中的人脸姿势。可以利用当前帧中的人脸姿势对上述的无效特征点进行修正。即按上述的旋转参数Rn和平移参数Tn以及无效特征点在初始帧中的坐标计算这些无效特征点的新坐标,将该新坐标替换它们在当前帧中的坐标。替换后的当前帧中的所有特征点的坐标将用来进行下一帧的数据处理。这有助于提高下一帧处理的精度。也可仅将当前帧中的有效特征值用来进行下一帧的处理,但这会减少可用的数据量。
按上述方式,每一帧上都会叠加眼镜图像,使用户在转动头部的情况下仍可看到眼镜是“佩戴”在脸上。如果用户头部动作比较剧烈,导致姿势变化过大,特别是在光线不足的情况下这样动作,则难以准确跟踪特征点,屏幕中的眼镜也将脱离眼部的位置。在这种情况下,可以提示用户进行复位操作。例如再次单击屏幕或指定的键,此时摄像头采集用户头像并将眼镜呈现在用户头像的眼睛处。在这种情况下,用户的操作发出了复位指令,手机或计算机接收复位指令后,将摄像头采集的当前帧作为上述的初始帧并按上面的方法进行处理。在处理过程中,于步骤S27得到聚类的处理结果,可以对其进行判断,如果 有效特征点的比例小于一个设定值,例如60%,或者在本帧采集到的特征点占上一帧采集到的特征点的比例小于一个设定值,例如30%,则输出提示信息,例如文本“单击屏幕以复位”,提示用户重新“试戴”眼镜。
图5是根据本发明实施例的实现虚拟试戴的装置的基本结构的示意图。该装置作为软件可设置在手机或计算机中。如图5所示,实现虚拟试戴的装置50主要包括人脸检测模块51、第一输出模块52、人脸姿势检测模块53、以及第二输出模块54。
人脸检测模块51用于对采集的初始帧进行人脸检测;第一输出模块52用于在人脸检测模块51采集到人脸的情况下,在初始位置生成物品图像然后与所述初始帧叠加后输出,该初始位置与初始帧中的人脸的指定位置重叠;人脸姿势检测模块53用于对当前帧中的人脸进行人脸姿势检测得到当前帧的人脸姿势;第二输出模块54用于根据物品图像的当前位置和人脸姿势再次生成物品图像,并使该物品图像中的物品姿势与人脸姿势一致,然后将该物品图像与当前帧叠加后输出。
人脸姿势检测模块53还可用于:在初始帧中确定人脸图像上的多个特征点;针对每个特征点进行如下处理:对特征点进行跟踪以确定该特征点在当前帧的位置,根据前一帧的人脸姿势,将初始帧中的该特征点的邻域进行仿射变换以得到该邻域在当前帧中的投影区域,计算初始帧中的邻域与当前帧中的投影区域之间的颜色偏移量并作为该特征点的跟踪偏差;对于确定的多个特征点,选择跟踪偏差较小的多个特征点;根据跟踪偏差较小的多个特征点在初始帧的位置以及在当前帧的位置确定当前帧的人脸姿势。
人脸姿势检测模块53还可用于:对于确定的多个特征点的跟踪偏差,以其中的最大值和最小值作为初始中心,按跟踪偏差的大小进行聚类得到两类;选择上述两类中跟踪偏差较小的一类所对应的特征点。
实现虚拟试戴的装置50还可包括修改模块(图中未示出),用于在所述人脸姿势检测模块确定当前帧的人脸姿势之后,将所述两类中跟踪偏差较大的一类所对应的特征点按照所述当前帧的人脸姿势投影到当前帧图像平面,以投影位置代替这些特征点在当前帧的位置。
实现虚拟试戴的装置50还可包括复位模块和提示模块(图中未示出),其中:复位模块用于接收复位指令,以及在接收到复位指令的情况下,将采集的当前帧作为初始帧;提示模块用于在人脸姿势检测模块按跟踪偏差的大小进行聚类得到两类之后,在跟踪偏差较小的一类特征点的数目占特征点总数目的比例大于第一预设值的情况下,或者,在当前帧中采集到的特征点的数目占特征点总数目的比例小于第二预设值的情况下,输出提示信息。
根据本发明实施例的技术方案,通过检测各帧的人脸姿势,再按人脸姿势调整眼镜姿势,能够使用户利用普通的图像采集装置就可完成虚拟试戴,并且用户可能转动头部以观察多个角度的佩戴效果,具有比较高的真实性。
以上结合具体实施例描述了本发明的基本原理,但是,需要指出的是,对本领域的普通技术人员而言,能够理解本发明的方法和设备的全部或者任何步骤或者部件,可以在任何计算装置(包括处理器、存储介质等)或者计算装置的网络中,以硬件、固件、软件或者它们的组合加以实现,这是本领域普通技术人员在阅读了本发明的说明的情况下运用他们的基本编程技能就能实现的。
因此,本发明的目的还可以通过在任何计算装置上运行一个程序或者一组程序来实现。所述计算装置可以是公知的通用装置。因此,本发明的目的也可以仅仅通过提供包含实现所述方法或者装置的程序代码的程序产品来实现。也就是说,这样的程序产品也构成本发明, 并且存储有这样的程序产品的存储介质也构成本发明。显然,所述存储介质可以是任何公知的存储介质或者将来开发出的任何存储介质。
还需要指出的是,在本发明的装置和方法中,显然,各部件或各步骤是可以分解和/或重新组合的。这些分解和/或重新组合应视为本发明的等效方案。并且,执行上述系列处理的步骤可以自然地按照说明的顺序按时间顺序执行,但是并不需要一定按照时间顺序执行。某些步骤可以并行或彼此独立地执行。
上述具体实施方式,并不构成对本发明保护范围的限制。本领域技术人员应该明白的是,取决于设计要求和其他因素,可以发生各种各样的修改、组合、子组合和替代。任何在本发明的精神和原则之内所作的修改、等同替换和改进等,均应包含在本发明保护范围之内。

Claims (12)

  1. 一种实现虚拟试戴的方法,其特征在于,包括:
    对采集的初始帧进行人脸检测,在检测到人脸的情况下,在初始位置生成物品图像然后与所述初始帧叠加后输出,该初始位置与所述初始帧中的人脸的指定位置重叠;
    对当前帧中的人脸进行人脸姿势检测得到当前帧的人脸姿势;
    根据所述物品图像的当前位置和所述人脸姿势再次生成物品图像,并使该物品图像中的物品姿势与所述人脸姿势一致,然后将该物品图像与所述当前帧叠加后输出。
  2. 根据权利要求1所述的方法,其特征在于,所述对当前帧中的人脸进行人脸姿势检测得到当前帧的人脸姿势的步骤包括:
    在所述初始帧中确定人脸图像上的多个特征点;
    针对每个特征点进行如下处理:
    对特征点进行跟踪以确定该特征点在当前帧的位置,
    根据前一帧的人脸姿势,将所述初始帧中的该特征点的邻域进行仿射变换以得到该邻域在当前帧中的投影区域,
    计算所述初始帧中的所述邻域与当前帧中的所述投影区域之间的颜色偏移量并作为该特征点的跟踪偏差,
    对于确定的所述多个特征点,选择跟踪偏差较小的多个特征点;
    根据所述跟踪偏差较小的多个特征点在所述初始帧的位置以及在当前帧的位置确定当前帧的人脸姿势。
  3. 根据权利要求2所述的方法,其特征在于,所述对于所述多个特征点,选择跟踪偏差较小的多个特征点的步骤包括:
    对于确定的所述多个特征点的跟踪偏差,以其中的最大值和最小值作为初始中心,按跟踪偏差的大小进行聚类得到两类;
    选择所述两类中跟踪偏差较小的一类所对应的特征点。
  4. 根据权利要求3所述的方法,其特征在于,所述确定当前帧的人脸姿势的步骤之后,还包括:
    将所述两类中跟踪偏差较大的一类所对应的特征点按照所述当前帧的人脸姿势投影到当前帧图像平面,以投影位置代替这些特征点在当前帧的位置。
  5. 根据权利要求3所述的方法,其特征在于,
    所述对采集的初始帧进行人脸检测的步骤之前,还包括:在接收到复位指令的情况下,将采集的当前帧作为所述初始帧;
    所述按跟踪偏差的大小进行聚类得到两类的步骤之后,还包括:
    在所述跟踪偏差较小的一类特征点的数目占特征点总数目的比例小于第一预设值的情况下,或者,在当前帧中采集到的特征点的数目占在前一帧采集到的特征点总数目的比例小于第二预设值的情况下,输出提示信息,然后接收复位指令。
  6. 根据权利要求1至5中任一项所述的方法,其特征在于,所述物品图像为眼镜图像、头部饰品图像、或者颈部饰品图像。
  7. 一种实现虚拟试戴的装置,其特征在于,包括:
    人脸检测模块,用于对采集的初始帧进行人脸检测;
    第一输出模块,用于在所述人脸检测模块采集到人脸的情况下,在初始位置生成物品图像然后与所述初始帧叠加后输出,该初始位置与所述初始帧中的人脸的指定位置重叠;
    人脸姿势检测模块,用于对当前帧中的人脸进行人脸姿势检测得到当前帧的人脸姿势;
    第二输出模块,用于根据所述物品图像的当前位置和所述人脸姿势再次生成物品图像,并使该物品图像中的物品姿势与所述人脸姿势一致,然后将该物品图像与所述当前帧叠加后输出。
  8. 根据权利要求7所述的装置,其特征在于,所述人脸姿势检测模块还用于:
    在所述初始帧中确定人脸图像上的多个特征点;
    针对每个特征点进行如下处理:
    对特征点进行跟踪以确定该特征点在当前帧的位置,
    根据前一帧的人脸姿势,将所述初始帧中的该特征点的邻域进行仿射变换以得到该邻域在当前帧中的投影区域,
    计算所述初始帧中的所述邻域与当前帧中的所述投影区域之间的颜色偏移量并作为该特征点的跟踪偏差,
    对于确定的所述多个特征点,选择跟踪偏差较小的多个特征点;
    根据所述跟踪偏差较小的多个特征点在所述初始帧的位置以及在当前帧的位置确定当前帧的人脸姿势。
  9. 根据权利要求8所述的装置,其特征在于,所述人脸姿势检测模块还用于:
    对于确定的所述多个特征点的跟踪偏差,以其中的最大值和最小值作为初始中心,按跟踪偏差的大小进行聚类得到两类;
    选择所述两类中跟踪偏差较小的一类所对应的特征点。
  10. 根据权利要求9所述的装置,其特征在于,还包括修改模块,用于在所述人脸姿势检测模块确定当前帧的人脸姿势之后,将所述两类中跟踪偏差较大的一类所对应的特征点按照所述当前帧的人脸姿势投影到当前帧图像平面,以投影位置代替这些特征点在当前帧的位置。
  11. 根据权利要求9所述的装置,其特征在于,还包括复位模块和提示模块,其中:
    所述复位模块用于接收复位指令,以及在接收到复位指令的情况下,将采集的当前帧作为所述初始帧;
    所述提示模块用于在所述人脸姿势检测模块按跟踪偏差的大小进行聚类得到两类之后,在所述跟踪偏差较小的一类特征点的数目占特征点总数目的比例小于第一预设值的情况下,或者,在当前帧中采集到的特征点的数目占在前一帧采集到的特征点总数目的比例小于第二预设值的情况下,输出提示信息。
  12. 根据权利要求7至11中任一项所述的装置,其特征在于,所述物品图像为眼镜图像、头部饰品图像、或者颈部饰品图像。
PCT/CN2015/081264 2014-06-17 2015-06-11 实现虚拟试戴的方法和装置 Ceased WO2015192733A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US15/319,500 US10360731B2 (en) 2014-06-17 2015-06-11 Method and device for implementing virtual fitting

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201410270449.X 2014-06-17
CN201410270449.XA CN104217350B (zh) 2014-06-17 2014-06-17 实现虚拟试戴的方法和装置

Publications (1)

Publication Number Publication Date
WO2015192733A1 true WO2015192733A1 (zh) 2015-12-23

Family

ID=52098806

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2015/081264 Ceased WO2015192733A1 (zh) 2014-06-17 2015-06-11 实现虚拟试戴的方法和装置

Country Status (4)

Country Link
US (1) US10360731B2 (zh)
CN (1) CN104217350B (zh)
TW (1) TWI554951B (zh)
WO (1) WO2015192733A1 (zh)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
TWI596379B (zh) * 2016-01-21 2017-08-21 友達光電股份有限公司 顯示模組與應用其之頭戴式顯示裝置
CN110288715A (zh) * 2019-07-04 2019-09-27 厦门美图之家科技有限公司 虚拟项链试戴方法、装置、电子设备及存储介质
US10685457B2 (en) 2018-11-15 2020-06-16 Vision Service Plan Systems and methods for visualizing eyewear on a user

Families Citing this family (23)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104217350B (zh) * 2014-06-17 2017-03-22 北京京东尚科信息技术有限公司 实现虚拟试戴的方法和装置
CN104851004A (zh) * 2015-05-12 2015-08-19 杨淑琪 一种饰品试戴的展示装置和展示方法
KR101697286B1 (ko) * 2015-11-09 2017-01-18 경북대학교 산학협력단 사용자 스타일링을 위한 증강현실 제공 장치 및 방법
CN106203300A (zh) * 2016-06-30 2016-12-07 北京小米移动软件有限公司 内容项显示方法及装置
CN106203364B (zh) * 2016-07-14 2019-05-24 广州帕克西软件开发有限公司 一种3d眼镜互动试戴系统及方法
WO2018096661A1 (ja) 2016-11-25 2018-05-31 日本電気株式会社 画像生成装置、顔照合装置、画像生成方法、およびプログラムを記憶した記憶媒体
CN106846493A (zh) * 2017-01-12 2017-06-13 段元文 3d虚拟试戴方法及装置
LU100348B1 (en) * 2017-07-25 2019-01-28 Iee Sa Method and system for head pose estimation
CN107832741A (zh) * 2017-11-28 2018-03-23 北京小米移动软件有限公司 人脸特征点定位的方法、装置及计算机可读存储介质
JP7290930B2 (ja) * 2018-09-27 2023-06-14 株式会社アイシン 乗員モデリング装置、乗員モデリング方法および乗員モデリングプログラム
CN109492608B (zh) * 2018-11-27 2019-11-05 腾讯科技(深圳)有限公司 图像分割方法、装置、计算机设备及存储介质
CN109615593A (zh) * 2018-11-29 2019-04-12 北京市商汤科技开发有限公司 图像处理方法及装置、电子设备和存储介质
US10825260B2 (en) * 2019-01-04 2020-11-03 Jand, Inc. Virtual try-on systems and methods for spectacles
CN110070481B (zh) * 2019-03-13 2020-11-06 北京达佳互联信息技术有限公司 用于面部的虚拟物品的图像生成方法、装置、终端及存储介质
CN111949112A (zh) 2019-05-14 2020-11-17 Oppo广东移动通信有限公司 对象交互方法及装置、系统、计算机可读介质和电子设备
CN110868554B (zh) * 2019-11-18 2022-03-08 广州方硅信息技术有限公司 直播中实时换脸的方法、装置、设备及存储介质
KR20230002738A (ko) 2020-04-15 2023-01-05 와비 파커 인코포레이티드 참조 프레임을 사용하는 안경용 가상 시착 시스템
CN111510769B (zh) * 2020-05-21 2022-07-26 广州方硅信息技术有限公司 视频图像处理方法、装置及电子设备
CN111627106B (zh) * 2020-05-29 2023-04-28 北京字节跳动网络技术有限公司 人脸模型重构方法、装置、介质和设备
CN111915708B (zh) * 2020-08-27 2024-05-28 网易(杭州)网络有限公司 图像处理的方法及装置、存储介质及电子设备
CN112258280B (zh) * 2020-10-22 2024-05-28 恒信东方文化股份有限公司 一种提取多角度头像生成展示视频的方法及系统
TWI770874B (zh) 2021-03-15 2022-07-11 楷思諾科技服務有限公司 使用點擊及捲動以顯示模擬影像之方法
CN113986015B (zh) * 2021-11-08 2024-04-30 北京字节跳动网络技术有限公司 虚拟道具的处理方法、装置、设备和存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN1866292A (zh) * 2005-05-19 2006-11-22 上海凌锐信息技术有限公司 一种动态眼镜试戴方法
CN102867321A (zh) * 2011-07-05 2013-01-09 艾迪讯科技股份有限公司 眼镜虚拟试戴互动服务系统与方法
CN103400119A (zh) * 2013-07-31 2013-11-20 南京融图创斯信息科技有限公司 基于人脸识别技术的混合显示眼镜交互展示方法
US20130322685A1 (en) * 2012-06-04 2013-12-05 Ebay Inc. System and method for providing an interactive shopping experience via webcam
CN104217350A (zh) * 2014-06-17 2014-12-17 北京京东尚科信息技术有限公司 实现虚拟试戴的方法和装置

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP5191665B2 (ja) * 2006-01-17 2013-05-08 株式会社 資生堂 メイクアップシミュレーションシステム、メイクアップシミュレーション装置、メイクアップシミュレーション方法およびメイクアップシミュレーションプログラム
JP5648299B2 (ja) * 2010-03-16 2015-01-07 株式会社ニコン 眼鏡販売システム、レンズ企業端末、フレーム企業端末、眼鏡販売方法、および眼鏡販売プログラム
US20130088490A1 (en) * 2011-04-04 2013-04-11 Aaron Rasmussen Method for eyewear fitting, recommendation, and customization using collision detection
US9236024B2 (en) * 2011-12-06 2016-01-12 Glasses.Com Inc. Systems and methods for obtaining a pupillary distance measurement using a mobile computing device
CN103310342A (zh) * 2012-03-15 2013-09-18 凹凸电子(武汉)有限公司 电子试衣方法和电子试衣装置
US9311746B2 (en) * 2012-05-23 2016-04-12 Glasses.Com Inc. Systems and methods for generating a 3-D model of a virtual try-on product
US20150382123A1 (en) * 2014-01-16 2015-12-31 Itamar Jobani System and method for producing a personalized earphone
US10564628B2 (en) * 2016-01-06 2020-02-18 Wiivv Wearables Inc. Generating of 3D-printed custom wearables

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN1866292A (zh) * 2005-05-19 2006-11-22 上海凌锐信息技术有限公司 一种动态眼镜试戴方法
CN102867321A (zh) * 2011-07-05 2013-01-09 艾迪讯科技股份有限公司 眼镜虚拟试戴互动服务系统与方法
US20130322685A1 (en) * 2012-06-04 2013-12-05 Ebay Inc. System and method for providing an interactive shopping experience via webcam
CN103400119A (zh) * 2013-07-31 2013-11-20 南京融图创斯信息科技有限公司 基于人脸识别技术的混合显示眼镜交互展示方法
CN104217350A (zh) * 2014-06-17 2014-12-17 北京京东尚科信息技术有限公司 实现虚拟试戴的方法和装置

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
TWI596379B (zh) * 2016-01-21 2017-08-21 友達光電股份有限公司 顯示模組與應用其之頭戴式顯示裝置
US10685457B2 (en) 2018-11-15 2020-06-16 Vision Service Plan Systems and methods for visualizing eyewear on a user
CN110288715A (zh) * 2019-07-04 2019-09-27 厦门美图之家科技有限公司 虚拟项链试戴方法、装置、电子设备及存储介质
CN110288715B (zh) * 2019-07-04 2022-10-28 厦门美图之家科技有限公司 虚拟项链试戴方法、装置、电子设备及存储介质

Also Published As

Publication number Publication date
HK1202690A1 (zh) 2015-10-02
CN104217350B (zh) 2017-03-22
US20170154470A1 (en) 2017-06-01
US10360731B2 (en) 2019-07-23
TW201601066A (zh) 2016-01-01
CN104217350A (zh) 2014-12-17
TWI554951B (zh) 2016-10-21

Similar Documents

Publication Publication Date Title
TWI554951B (zh) 實現虛擬試戴的方法和裝置
US11030237B2 (en) Method and apparatus for identifying input features for later recognition
JP6268303B2 (ja) 2d画像分析装置
US10945514B2 (en) Information processing apparatus, information processing method, and computer-readable storage medium
US8036416B2 (en) Method and apparatus for augmenting a mirror with information related to the mirrored contents and motion
US8976160B2 (en) User interface and authentication for a virtual mirror
Wang et al. Real time eye gaze tracking with kinect
CN108022124B (zh) 互动式服饰试穿方法及其显示系统
US10976829B1 (en) Systems and methods for displaying augmented-reality objects
WO2015020703A1 (en) Devices, systems and methods of virtualizing a mirror
Sun et al. Real-time gaze estimation with online calibration
CN105608238A (zh) 服装试穿方法及装置
Malik et al. Simultaneous hand pose and skeleton bone-lengths estimation from a single depth image
Dubey et al. Unsupervised learning of eye gaze representation from the web
WO2018059258A1 (zh) 采用增强现实技术提供手掌装饰虚拟图像的实现方法及其装置
JP2012003724A (ja) 三次元指先位置検出方法、三次元指先位置検出装置、及びプログラム
Prajapat et al. Jewellery Tryon using AR
Brito et al. Recycling a landmark dataset for real-time facial capture and animation with low cost hmd integrated cameras
HK1202690B (zh) 实现虚拟试戴的方法和装置
Sun et al. A deep learning approach to appearance-based gaze estimation under head pose variations
CN114627313A (zh) 配送任务控制方法、机器人和存储介质
Mishra et al. Fingertips detection with nearest-neighbor pose particles from a single RGB image
Guan et al. BiFingerPose: Bimodal Finger Pose Estimation for Touch Devices
CN117648035B (zh) 一种虚拟手势的控制方法及装置
Egashira et al. Vision-based motion capture of interacting multiple people

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15810445

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: IDP00201608459

Country of ref document: ID

WWE Wipo information: entry into national phase

Ref document number: 15319500

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 12/05/2017)

122 Ep: pct application non-entry in european phase

Ref document number: 15810445

Country of ref document: EP

Kind code of ref document: A1