WO2014117446A1 - 基于单个视频摄像机的实时人脸动画方法 - Google Patents
基于单个视频摄像机的实时人脸动画方法 Download PDFInfo
- Publication number
- WO2014117446A1 WO2014117446A1 PCT/CN2013/075117 CN2013075117W WO2014117446A1 WO 2014117446 A1 WO2014117446 A1 WO 2014117446A1 CN 2013075117 W CN2013075117 W CN 2013075117W WO 2014117446 A1 WO2014117446 A1 WO 2014117446A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- face
- expression
- image
- feature point
- feature points
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T13/00—Animation
- G06T13/20—Three-dimensional [3D] animation
- G06T13/40—Three-dimensional [3D] animation of characters, e.g. humans, animals or virtual beings
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
- G06T7/246—Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
- G06T7/251—Analysis of motion using feature-based methods, e.g. the tracking of corners or segments involving models
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/74—Image or video pattern matching; Proximity measures in feature spaces
- G06V10/75—Organisation of the matching processes, e.g. simultaneous or sequential comparisons of image or video features; Coarse-fine approaches, e.g. multi-scale approaches; using context analysis; Selection of dictionaries
- G06V10/755—Deformable models or variational models, e.g. snakes or active contours
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/168—Feature extraction; Face representation
- G06V40/171—Local features and components; Facial parts ; Occluding parts, e.g. glasses; Geometrical relationships
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/174—Facial expression recognition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30196—Human being; Person
- G06T2207/30201—Face
Definitions
- the invention relates to a real-time capture of human face motion and real-time animation technology, in particular to a real-time face animation method based on a single video camera.
- Facial expression capture is an important part of realistic graphics, and it is widely used in the fields of movies, animation, games, web chat and education.
- a face animation system based on face motion capture is used to estimate the user's expressions and actions and map them to another target model.
- active sensing methods are usually used, including placing some signal transmission points on the person's face (Wi ll iams, L. 1990. Performance driven facial animation.
- SIGGRAPH SIGGRAPH, 234-242; Huang , H. , Chai, J. , Tong, X. , and Wu, H. , T. , 2011. Leveraging motion capture and 3d scanning for high-field facial performance acqui sition.
- projection structure light map Zhang, L. , Snavely, N. , Curless, B. , and Seitz, SM 2004.
- Space time faces high resolution capture for modeling and animation.
- a camera array which allows facial data to be acquired from multiple perspectives and converted into volumetric data for 3D reconstruction, including BEELER, T., BICKEL, B., BEARDSLEY, P. , SUMNER, R. , AND GROSS, M. 2010. High-quality single-shot capture of facial geometry.
- each feature point is not only related to its own local information calculation, but also affected by other feature points.
- Different types of geometric constraints are widely used, including restrictions on feature point offsets when expressions change (CHAI, J. -X. , XIAO, J. , AND HODGINS, J. 2003. Vision-based control of 3d facial animation. In Eurographics/SIGGRAPH Symposium on Computer Animation, 193 - 206. ) , Meeting the needs of physics-based deformable mesh models (ESSA, I . BASU, S. , DARRELL, T. , AND PENTLAND, A. 1996. Modeling, tracking and interactive animation of faces and heads : Using input from video.
- an expression fusion model (Blendshapes) is widely used. This is a subspace representation of a facial gesture that includes a series of basic facial expressions that form a linear space for facial expressions.
- Morphing can be performed on the basic face movements (PIGHIN, F., HECKER, J., LISCHINSKI, D., SZELISKI, R., AND SALES IN, DH 1998. Synthesizing reali stic facial expressions From photographs. In Proceedings of SIGGRAPH, 75 - 84. ) or Linear combinations ( LEWIS, JP , AND ANJYO, K. 2010.
- Multi l inear Models represent a fusion model decomposition with multiple control properties (such as individual, expression, and pronunciation).
- An important feature of the expression fusion model is that the expressions of different individuals correspond to similar basic action coefficients in the fusion model. Using this property, many facial animation applications use the expression fusion model to transfer the user's face motion to the virtual alias by passing the basic action coefficients.
- a real-time face animation method based on a single video camera comprising the following steps:
- Image acquisition and calibration Using a video camera to capture a user's multiple 2D images with different poses and expressions, and using 2D image regression to obtain corresponding 2D face feature points for each image, the automatic detection is obtained. Manual adjustment of inaccurate feature points;
- 3D feature point tracking The user inputs the image in real time using the video camera. For the input image, combined with the 3D face feature points of the previous frame, the regression device is used in step 2 to locate the 3D face feature points in the current frame in real time. ;
- Pose expression parameterization Using the position of the 3D face feature point, combined with the user expression fusion model obtained in step 2, iteratively optimizes to obtain the parametric representation of the face pose and expression;
- Alias driver Maps face poses and expression parameters to virtual substitutes to drive animated characters for face animation.
- 1 is a two-dimensional image and a calibration two-dimensional image feature point map acquired in the image acquisition and calibration step of the present invention
- 3 is a three-dimensional feature point map of an image and a positioning input in real time in the three-dimensional feature point tracking step of the present invention
- FIG. 5 is a diagram of the animation of the animation character in the step of driving the avatar in the embodiment of the present invention.
- the core of the present invention is to obtain a three-dimensional feature point of a face from a two-dimensional image, thereby parameterizing the user's face pose and expression, and mapping it to a virtual substitute.
- the method is mainly divided into the following five steps: image acquisition and calibration, data preprocessing, 3D feature point tracking, gesture expression parameterization, and avatar drive. Specifically, it includes the following steps:
- Image Acquisition and Calibration The user mimics various gestures and expressions and uses the video camera to capture the corresponding image.
- a common two-dimensional image regression device is used for each image to obtain corresponding two-dimensional face feature points, and the user can perform manual adjustment on the inaccurate feature points automatically detected.
- the present invention first collects images of a set of different face poses and expressions of the user. This set of images is divided into two parts: rigid motion and non-rigid motion.
- Rigid action means that the user maintains a natural expression while doing 15 different angles of the face pose.
- Non-rigid motions include 15 different expressions at three ym angles. These expressions are relatively large expressions that vary widely between individuals. These expressions are: open mouth, smile, raise eyebrows, disgust, squeeze left eye, squeeze right eye, anger, pout to the left, pout to the right, grin, grin, grin, flip lips, drums and close eye.
- CVPR Computer Vision and Pattern Recognition
- Data preprocessing Using the image of the 2D feature points, the user expression fusion model is generated and the internal parameters of the camera are calibrated, and the 3D feature points of the image are obtained. A two-dimensional image and a three-dimensional feature point are used to train a regression of a three-dimensional feature point from a two-dimensional image.
- the user expression fusion model contains the user's natural expression shape 2?. And 46 FACS expression shapes ( , ⁇ ,... ⁇ . These expression shapes form the linear space of the user's expression, and the user's arbitrary expression S can be obtained by linear interpolation using the basic expression in the fusion model:
- FaceWarehouse a three-dimensional facial expression model FaceWarehouse (CAO, C., WENG, ⁇ ., ZHOU, S., TONG, ⁇ ., AND ZHOU, ⁇ . 2012.
- Facewarehouse a 3d facial expression database for visual computing. Tech. Rep.
- FaceWarehouse includes individual data in 150 different contexts, each of which includes 46 FACS expression shapes.
- FaceWarehouse uses this data to create a bilinear model with two attributes, individual and expression, which form a three-dimensional nuclear tensor G (11K model vertex ⁇ 50 individual ⁇ 45 table) n), using this nuclear tensor, any expression F of any individual can be obtained by using the tensor contraction calculation ⁇ :
- F is the expression obtained by shrinking calculations.
- u is the first 2D feature point position in the first image
- w is the corresponding vertex number on the 3D mesh shape
- n Q represents the operation of projecting the 3D space point to the image coordinates by means of the camera projection matrix Q
- w ⁇ , and w respectively, are the column vectors of individuals and expression coefficients in the tensor, which are the nuclear tensors of FaceWarehouse.
- the user's expression fusion model can be generated as follows:
- ⁇ is the truncated transformation matrix of FaceWarehouse expression attribute, d, is the expression coefficient vector, its first element is 1 and the remaining elements are 0, which is the core tensor of FaceWarehouse and the individual coefficient of w id. 2.
- the camera projection matrix describes the projection of a three-dimensional point in the camera coordinate system to a two-dimensional image position, which is completely dependent on the parameters inside the camera and can be expressed as a projection matrix Q:
- the parameters ⁇ and ⁇ represent the focal length in pixels in the length and width directions, indicating the offset in the direction of X and ). And means the position of the origin of the image, that is, the intersection of the optical axis and the image plane.
- camera calibration eg ZHANG, Z. 2000. A flexible new technique for camera cal ibration. IEEE Trans. Pattern Anal. Mach. Intel l. 22, 11, 1330 - 1334.
- Some standard calibration targets such as checkerboard).
- the user's expression fusion model is obtained, and at the same time, there is a corresponding posture transformation matrix and expression coefficient on each input image, so that the three-dimensional face shape on the image can be obtained:
- F is the generated three-dimensional face shape
- M is the posture transformation matrix, 2?. It is the natural expression shape of the user, which is the basic expression shape in the user expression fusion model, ", ⁇ is the coefficient of the basic expression.
- the three-dimensional feature points of the image can be constructed.
- the present invention replaces the 15 feature points of the outer contour with the inner 15 feature points (as shown in Fig. 2).
- iS,. l we use iS,. l to represent the three-dimensional feature points corresponding to these images.
- the present invention entails enhancing the acquired images and their corresponding three-dimensional feature points. For each acquired image and its 3D feature points we will have 3D feature points S,. Translation along three axes in the camera coordinate system yields additional m-1 three-dimensional feature points for each S.
- the enhanced three-dimensional feature points correspond to additional images.
- the present invention does not actually generate a corresponding image, but merely records the enhanced three-dimensional feature points and transforms to the original feature point S.
- the transformation matrix M the matrix with, S,.
- w raw data is enhanced into one, we define them as ⁇ / ; , M ⁇ 3 ⁇ 4 ⁇ o
- 3D Feature point space which describes the range of variation of the face feature points of the user in three-dimensional space.
- the present invention assigns it different initialization feature points.
- the present invention simultaneously considers the locality and randomness of the data when selecting the initial feature points of the data used for training. For each set of images/feature points (/, M ⁇ , ), first find the G and most approximate feature points from the n original feature point sets ⁇ S, ° ⁇ , and calculate the similarity of the two feature points. Sex, first align the centers of the two feature points, and then calculate the sum of the squares of the distances between the corresponding feature points. The most similar set of feature points we will find is called, g , i ⁇ g ⁇ G].
- H feature points are randomly selected from the enhanced feature points of each S LG , and are denoted as A , l ⁇ i ⁇ H ⁇ .
- the present invention finds an initialization feature point for each pair of images/feature points ⁇ /, M, S ⁇ .
- Each training data represents it as (/,, M, ⁇ ).
- / is a two-dimensional image, which is a transformation matrix in which the feature points are subjected to translation enhancement, which is /, corresponding to the ⁇ 3 feature points
- S, GA is the initial feature occupancy, .
- N ⁇ m.G.H training data.
- G 5
- H 4.
- N training data ⁇ (/,,M,S,,S, ⁇ .
- the present invention uses the information in the image /, to generate a regression function from the initial feature point to the corresponding feature point S.
- the present invention uses a two-layer cascade of regressions in which there are ⁇ -class weak classifiers in the first layer and level atom classifiers in each of the weak classifiers.
- the present invention produces a set of sequence pair features for atomic classifier construction.
- the current feature point S and the image / are used to calculate the appearance vector: at the current feature point S. Randomly selected P sample points, wherein the position of each sample point p is represented as a feature point position in S plus an offset d P ; then the sample point p is projected onto the image by using n Q (M p) Finally, the color value of the corresponding pixel is taken from the image /.
- the P color values constitute the appearance vector of the training data in the first level cascade regression.
- P 2 sequence number features can be generated by calculating the difference between two different position elements.
- each atomic classifier of the second layer in the P 2 sequence number features generated by the first layer, effective features are searched, and the training data is classified. For each training data (/, M, S, S;), first calculate the difference between the current feature point S and the real feature point S, and then project the differences into a random direction to produce a scalar, this These scalars are treated as random variables, and the features with the greatest correlation with this random variable are found from the P 2 serial number features. This step is repeated F times to produce F different features from which the atom classifier is generated.
- F features are assigned a random threshold that divides all training data into 2f classes. For each training data, we compare the eigenvalues and thresholds calculated by the serial number. To decide which category the training data belongs to.
- the present invention refers to all data sets falling in this class as ⁇ 3 ⁇ 4 and calculates the feature point regression output in this class as:
- the training of the repeller is iteratively executed one time, each time generating a cascade of atomic classifiers, forming a weak classifier, and iteratively optimizing the regression output.
- These cascaded weak classifiers form a strong classifier, the one we need.
- Parameter configuration of the present invention selects ⁇ OO SOO ⁇ z ⁇ O ⁇ z ⁇ SO c
- Three-dimensional feature point tracking For the image input by the user in real time, combined with the three-dimensional face feature point of the previous frame, using the regression device in the data pre-processing step, combined with the three-dimensional feature point S' of the previous frame, the present invention can be real-time Position the 3D feature points of the current frame. First find the feature point closest to S ' in the original feature point set ⁇ S, ° ⁇ , then transform the position where S ' is transformed by a rotation and translation (M" of a rigid body, and record the transformed previous frame. The feature point is S '*.
- the use of a regression to locate feature points is divided into two levels of concatenation.
- the appearance vector V is obtained first based on the current frame image /, the current feature point & , the inverse matrix of the transformation matrix 1 ⁇ , and the offset ⁇ d ⁇ recorded during the training.
- the features are calculated based on the sequence numbers recorded in each atom classifier and compared with the thresholds to determine which class belongs to, and the regression output S 3 ⁇ 4 of the class is obtained. Finally, use this output to update the current feature points: S ⁇ St + SS
- the present invention obtains L output feature points through the regression device for the L initial feature points, and finally, the median operation is performed on these output feature points to obtain the final result.
- this feature point is in the 3D feature point space, you need to use the transformation matrix!
- the inverse matrix of ⁇ transforms it into the original image position.
- the input 2D image and the resulting 3D feature point results are shown in Figure 3.
- Pose expression parameterization Using the three-dimensional position of the feature point, combined with the user expression fusion model obtained in the data preprocessing, iteratively optimizes the parameterized expression of the face pose and expression obtained.
- the three-dimensional feature point positions of the current frame are obtained in the previous step, and the present invention uses them to parameterize the motion of the current frame face.
- the action of the face is mainly divided into two parts: a rigid face pose represented by the transformation matrix M, and a face non-rigid expression represented by the expression fusion coefficient a. These two parameters can be obtained by optimizing the following matching energies:
- ⁇ is the three-dimensional position of the first feature point in ⁇ , which is the corresponding vertex number on the three-dimensional face shape. It is the user's natural expression face shape, which is the other basic expression face shape in the expression fusion model, "is the coefficient of the basic expression, and M is the transformation matrix representing the face pose. And WEISE, T., BOUAZIZ, S., LI, H., AND PAULY, M. 2011. Realt ime performance-based facial animat ion. ACM Trans. Graph. 30, 4 (July) , 77 : 1 - 77 : 10. Similarly, the present invention uses an animated prior to enhance temporal continuity in the tracking process. Given the previous "frame emoticon coefficient vector.., a-" ⁇ , combine it with the coefficient a of the current frame into a single vector (a, A circumstances) The present invention describes the probability distribution of the vector as a Gaussian mixture model:
- the hybrid Gaussian model can be trained using a number of pre-generated expression animation sequences (WEISE, T., BOUAZIZ, S., LI, H., AND PAULY, M. 2011. Realtime performance-based facial animation. ACM Trans. Graph. 30, 4 (July), 77 : 1 - 77 : 10. ).
- the Gaussian mixture model can describe an energy for interframe continuity
- the present invention utilizes a two-step iterative method to optimize the energy ⁇ .
- the expression coefficient a of the previous frame is used as the initial value of the current frame and remains unchanged, and then the covariance matrix of the corresponding point pair distribution is calculated by the singular value decomposition (SVD) to calculate the rigid posture, that is, the transformation matrix M.
- the present invention fixes M, and then uses the gradient descent method to calculate the expression coefficient 3.
- the present invention iteratively performs these two steps until the result converges, usually after two iterations, satisfactory results can be obtained.
- the parametric representation of the gestures and expressions we can get the corresponding user's 3D face shape. The result is shown in Figure 4.
- Steadicam driver Map face poses and expression parameters to virtual substitutes to drive animated characters for face animation.
- the parameterized face pose and expression coefficients are obtained, and the present invention can map it to a virtual substitute.
- the present invention maps the parameterized posture M and the expression coefficient a to the substitute body, that is,
- M is the transformation matrix of the face pose
- 3 ⁇ 4 is the user's natural expression face shape ⁇
- ⁇ is the other basic expression face shape in the expression fusion model
- " is the coefficient of the basic expression
- ) is the face shape of the final substitute.
- Example The inventors implemented an embodiment of the present invention on a desktop computer equipped with an Intel Core i 7 (3.5 GHz) central processing unit and a web camera providing 640 x 480 resolution at 30 fps.
- the results of the figures are obtained in the implementation using the parameter settings mentioned in the specific embodiments. In practice, it takes less than 15 milliseconds to complete a frame of capture, parameterization, and avatar mapping on a normal computer.
- the results show that in our current hardware configuration, the present invention can handle various large-scale posture rotations in real time, various exaggerated expressions, and get an animation effect very close to the user input, and has a good user experience.
- the present invention can obtain satisfactory results under different lighting conditions, such as an office, a direct sunlight, and a dimly lit hotel room.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Multimedia (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Software Systems (AREA)
- Human Computer Interaction (AREA)
- Databases & Information Systems (AREA)
- Evolutionary Computation (AREA)
- Computing Systems (AREA)
- Medical Informatics (AREA)
- Artificial Intelligence (AREA)
- Processing Or Creating Images (AREA)
Abstract
本发明公开了一种基于单个视频摄像机的实时人脸动画方法,该方法利用单个视频摄像机,实时跟踪人脸特征点的三维位置,并以此参数化人脸的姿势和表情,最终可以将这些参数映射到替身上以驱动动画角色的脸部动画。本发明不需要高级的采集设备,只需要用户的普通视频摄像机上就可以达到实时的速度;本发明可以准确处理人脸各种大角度旋转、平移,和人脸的各种夸张表情;本发明还可以在不同光照和背景环境下工作,包括室内和有阳光的室外。
Description
基于单个视频摄像机的实时人脸动画方法
技术领域
本发明涉及人脸动作实时捕获和实时动画技术, 尤其涉及一种基于单个视频 摄像机的实时人脸动画方法。
背景技术 说
本发明相关的研究背景简述如下:
1. 人脸动作捕获
书
人脸表情捕获是真实感图形学的重要组成部分, 其被广泛的应用到电影、 动 画、 游戏、 网络聊天和教育等领域中。 基于人脸动作捕获的人脸动画系统用于估 计用户的表情和动作, 并将它们映射到另外一个目标模型上。 为实现这一目标, 目前已经有很多的相关技术。 其中为了直接与用户进行交互, 通常使用主动传感 方法, 包括在人的脸部放置一些信号发送点 (Wi ll iams, L. 1990. Performance driven facial animation. In Proceedings of SIGGRAPH, 234-242; Huang, H. , Chai, J. , Tong, X. , and Wu, H. , T. , 2011. Leveraging motion capture and 3d scanning for high-field facial performance acqui sition. ACM Trans. Graph. 30, 4, 74 : 1-74 : 10. ) , 或者使用投影结构光图谱 (Zhang, L. , Snavely, N. , Curless, B. , and Seitz, S. M. 2004. Space time faces: high resolution capture for modeling and animation. ACM Trans. Graph. 23, 3, 548-558 ; Wei se, T. , Li, H. , Gool, L. V. , and Pauly, M. 2009. Face/off : Live facial puppetry. In Eurograph i c s/S i ggraph Symposium on Computer Animation. ) , 这些方法可以精确的跟踪人脸表面位置, 并获取高分辨率、 高精 度的人脸估计, 但是这些主动传感方法往往需要昂贵的硬件设备支持, 同时由于 信号发送点或者结构光的干扰, 对用户并不具有友好型, 因此并不能广泛用于普 通用户。
另一种系统是被动系统, 它们并不主动向所在环境里发送信号或在人脸放置 信号发送点, 而只根据接收到的颜色信息等来分析、 捕获人脸动作。 其中有一些 方法只用单个摄像机来捕获人脸动作, 包括 Essa, I. , Basu, S. , Darrell, Τ. , and Pentland, A. 1996. Modeling, tracking and interactive animation of
faces and heads: Using input from video. In Computer Animation, 68-79 ; Pighin, F., Szeliski, R., and Salesin, D. 1999. Resynthesizing facial animation through 3d model-based tracking. In International Conference on Computer Vi sion, 143 - 150 ; In Eurograph i c s/S i ggraph Symposium on Computer Animation, 193-206 ; Vlasic, D. , Brand, M. , Pf ister, H. and Popovic, J. 2005. Face transfer with multi linear models.等工作。 这些方 法的缺点在于得到的结果精度较低, 无法处理人脸大幅度的旋转和夸张表情, 此 外对使用的环境也有一定的要求, 如只能在光照均匀、 没有阴影、 高光干扰的室 内环境下使用。
某些方法中则使用了照相机阵列, 这样可以从多个视角获取人脸数据, 并将 其转化成立体数据用于三维重建, 这些工作包括 BEELER, T. , BICKEL, B. , BEARDSLEY, P. , SUMNER, R. , AND GROSS, M. 2010. High-quality single-shot capture of facial geometry. ACM Trans. Graph. 29, 4, 40 : 1 - 40 : 9. ; BRADLEY, D. , HEIDRICH, I , POPA, T. , AND SHEFFER, A. 2010. High resolution passive facial performance capture. ACM Trans. Graph. 29, 4, 41 : 1 - 41 : 10.; BEELER, T. , HAHN, F. , BRADLEY, D. , BICKEL, B. , BEARDSLEY, P. , GOTSMAN, C. , SUMNER, R. I , AND GROSS, M. 2011. High-quality passive facial performance capture using anchor frames. ACM Trans. Graph. 30, 4 75 : 1 - 75 : 10.等, 这些方法可以比较精确的得到三维人脸表情, 但依然存在设备 昂贵、 对环境要求较高等缺点。
2. 基于视觉的人脸特征点跟踪
人脸表情的捕获往往需要跟踪输入图像中人脸的特征点, 如眼角、 鼻子边缘 或者嘴巴边界等位置。 对于一般的输入视频, 光流法 (Optical Flow) 被普遍使 用。 但是由于输入数据的噪声影响, 对那些不是很明显的人脸特征点 (如脸颊上 的点) , 光流定位并不是那么可靠, 往往会因为帧与帧之间的误差累积造成一种 偏移 (Drift ) 的错误。 此外, 光流法在处理快速运动、 光照变化等方面也存在较 大的误差。
为了更精确的特征点跟踪, 一些工作使用特征点之间的几何约束。 这样, 每 个特征点不仅和其自身局部信息计算有关, 还受到其他特征点的影响。 不同类型 的几何约束被广泛使用, 包括对表情变化时特征点偏移的限制 (CHAI, J. -X. ,
XIAO, J. , AND HODGINS, J. 2003. Vision-based control of 3d facial animation. In Eurographics/SIGGRAPH Symposium on Computer Animation, 193 - 206. ) , 满足基于物理的可变形网格模型需求 (ESSA, I. , BASU, S. , DARRELL, T. , AND PENTLAND, A. 1996. Modeling, tracking and interactive animation of faces and heads : Using input from video. In Computer Animation, 68 - 79. ; DECARLO, D. , AND METAXAS, D. 2000. Optical flow constraints on deformable models with applications to face tracking. Int. Journal of Computer Vi sion 38, 2, 99 - 127. ) , 以及一些从大量样本空 间中建立的人脸模型的对应关系 (PIGHIN, F. , SZELISKI, R. , AND SALES IN, D. 1999. Resynthesizing facial animation through 3d model-based tracking. In International Conference on Computer Vision, 143 - 150.; BLANZ, V., AND VETTER, T. 1999. A morphable model for the synthesi s of 3d faces. In Proceedings of SIGGRAPH, 187 - 194.; VLASIC, D. , BRAND, M. , PFISTER, H. , AND POPOVIC 766 , J. 2005. Face transfer with multi linear models. ACM Trans. Graph. 24, 3 (July) , 426 - 433. ) 。 这些方法都能在一定程度上跟 踪图像、 视频中人脸特征点, 但由于它们得到的都是图像上的二维特征点, 因此 在处理旋转上有一定的局限性。
3. 三维人脸模型
我们的工作中在预处理过程中借助了三维人脸模型, 以从二维图像中获取得 到三维信息。
在现有的图形学和视觉应用中, 各种三维人脸模型被广泛应用。 在人脸动画 应用中, 一种表情融合模型 (Blendshapes ) 被广泛应用。 这是一种表示人脸动作 的子空间表达, 其包括一系列的基本人脸表情, 由此组成了人脸表情的线性空 间。 利用融合模型, 可以对其中的基本人脸动作通过变形 (Morphing) (PIGHIN, F. , HECKER, J. , LISCHINSKI, D. , SZELISKI, R. , AND SALES IN, D. H. 1998. Synthesizing reali stic facial expressions from photographs. In Proceedings of SIGGRAPH, 75 - 84. ) 或者线性组合 ( Linear combinations ) ( LEWIS, J. P. , AND ANJYO, K. 2010. Direct manipulation blendshapes. IEEE CG&A 30, 4, 42 - 50.; SEO, J. , IRVING, G. , LEWIS, J. P. , AND NOH, J
2011. Compression and direct manipulation of complex blendshape models. ACM Trans. Graph. 30, 6. ) 等计算得到各种人脸动画效果。
多线性模型 (Multi l inear Models ) 表示一个拥有多种控制属性 (如个体, 表情, 发音嘴型) 的融合模型分解。 表情融合模型的一个重要的特点在于, 不同 个体的表情对应于融合模型中相似的基本动作系数。 利用这一性质, 很多人脸动 画应用使用了表情融合模型, 通过传递基本动作系数, 将用户的人脸动作转移到 虚拟替身中。
发明内容
本发明的目的在于针对现有技术的不足, 提供一种基于单个视频摄像机的实时 人脸动画方法, 本发明可以在普通桌面电脑上供普通用户使用, 在不同环境下实 时, 准确的捕获用户表情并驱动虚拟替身。 具有易使用, 鲁棒, 快速等特点, 可 以运用于在线游戏、 网络聊天和教育等应用中, 具有很高的实用价值。
本发明的目的是通过以下技术方案来实现的: 一种基于单个视频摄像机的实 时人脸动画方法, 包括以下步骤:
( 1 ) 图像采集和标定: 利用视频摄像机拍摄用户的多幅具有不同姿势和表情的二 维图像, 对每个图像利用二维图像回归器得到对应的二维人脸特征点, 对自动检 测得到的不准确特征点进行手动调整;
(2 ) 数据预处理: 利用标定好二维人脸特征点的图像, 进行用户表情融合模型生 成和摄像机内部参数标定, 并由此得到图像的三维特征点; 利用三维特征点和步 骤 1采集的二维图像, 训练获得从二维图像到三维特征点的回归器;
(3 ) 三维特征点跟踪: 用户使用视频摄像机实时输入图像, 对于输入的图像, 结 合上一帧的三维人脸特征点, 利用步骤 2 中得到回归器, 实时定位当前帧中三维 人脸特征点;
(4) 姿势表情参数化: 利用三维人脸特征点位置, 结合步骤 2中得到的用户表情 融合模型, 迭代优化以得到人脸姿势和表情的参数化表达;
(5 ) 替身驱动: 将人脸姿势和表情参数映射到虚拟替身上, 用以驱动动画角色进 行人脸动画。
本发明的有益效果是, 本发明容易使用, 使用者不需要信号发送点或者投影 结构光谱等昂贵的物理设备, 只使用单个摄像机, 在普通的桌面电脑上, 通过一 次性数据采集和预处理, 就可以供用户完成人脸姿势、 表情的捕获和参数化, 并
将参数化结果映射到虚拟替身上以驱动动画角色的人脸动画, 方便普通用户的使 用。 本发明相比于之前的方法, 可以有效的处理视频中的快速运动, 大幅度头部 姿势旋转和夸张表情, 并且可以处理一定的光照条件变化, 可以在不同的环境下 使用 (包括室内、 有阳光直射的室外环境等) 。 此外, 本发明的方法非常高效, 在具体实施实例中, 普通电脑使用少于 15毫秒即可以完成一帧的特征点跟踪、 姿 势表情参数化和替身映射, 拥有非常好的用户体验。
附图说明
图 1是本发明图像采集和标定步骤中采集的一幅二维图像和标定的二维图像 特征点图;
图 2是本发明数据预处理步骤中生成的三维人脸特征点图;
图 3是本发明三维特征点跟踪步骤中实时输入的图像和定位的三维特征点 图;
图 4是本发明姿势表情参数化步骤中生成的三维人脸形状图;
图 5是本发明替身驱动步骤中将图 4中的参数映射到替身上, 驱动动画角色 人脸动画截图。
具体实施方式
本发明核心是从二维图像中得到人脸的三维特征点, 由此参数化用户的人脸 姿势和表情, 并将其映射到虚拟替身。 该方法主要分为以下五个步骤: 图像采集 和标定、 数据预处理、 三维特征点跟踪、 姿势表情参数化、 替身驱动。 具体来 说, 包括以下步骤:
1.图像采集和标定: 用户模仿做出各种姿势和表情, 并利用视频摄像机拍摄 相应图像。 对每个图像利用通用的二维图像回归器, 得到对应的二维人脸特征 点, 对自动检测得到的不准确的特征点, 允许用户进行手动调整。
本发明首先采集用户的一组不同人脸姿势和表情的图像。 这组图像分为两个 部分: 刚性动作和非刚性动作。 刚性动作指用户保持自然表情, 同时做 15个不同 角度的人脸姿势。 我们用欧拉角( ν, ^^, Γο/Ζ)来表示这些角度: ιν以 30。为间 隔从 -90°到 90°采样, 同时保持 pitch和 roll为 0°; pitch以 15°为间隔从 -30°到 30° 采样并去除 0°, 同时保持 IV和 ro/Z为 0° ; ro/Z以 15°为间隔从 -30°到 30°采样并去 除 0°, 同时保持 ιν和/ ^'t 为 0°。 注意到我们并不要求用户做到的姿势角度与要 求的角度配置完全精确, 只需要一个大概的估计即可。
非刚性动作包括三个 ym角度下的 15个不同表情。 这些表情是一个相对比较 大的表情, 在不同个体之间差异较大。 这些表情是: 张嘴, 微笑, 抬眉毛, 厌 恶, 挤左眼, 挤右眼, 愤怒, 向左歪嘴, 向右歪嘴, 露齿笑, 嘟嘴, 撅嘴, 翻嘴 唇, 鼓嘴和闭眼。
对每个用户, 总共采集了 60张图像。 对每张图像, 我们使用二通用的二维图 像回归器来自动定位 75个特征点位置 (如附图 1所示) , 这些特征点主要分为两 类: 60 个内部特征点 (如眼睛、 眉毛, 鼻子和嘴巴部分的特征) , 和 15个外部 轮廓点。 本发明中使用(CAO, X. , WEI, Y. , WEN, F. , AND SUN, J. 2012. Face alignment by explicit shape regression. In Computer Vision and Pattern Recognition (CVPR) , 2887 - 2894. ) 所描述的回归器来自动定位这些特征点。
自动定位的二维特征点会存在一些偏差, 针对定位不精确的特征点, 用户可 以在屏幕上通过简单的鼠标拖拽操作来修正, 具体来说, 即通过鼠标点击选中特 征点, 然后按住鼠标将其拖到图像上正确的位置。
2. 数据预处理: 利用标定好二维特征点的图像, 进行用户表情融合模型生成 和摄像机内部参数标定, 并由此得到图像的三维特征点。 利用二维图像和三维特 征点, 训练从二维图像得到三维特征点的回归器。
2. 1 用户表情融合模型的生成
用户表情融合模型包含用户的自然表情形状2?。和 46 个 FACS 表情形状 ( ,^,…^ 。 这些表情形状构成了该用户表情的线性空间, 用户任意表情 S可以 用融合模型中的基本表情通过线性插值得到:
其中, 2?。是该用户的自然表情形状, 是用户表情融合模型中的基本表情形 状, 《,·是基本表情的系数, 而 则是插值得到的表情人脸形状。
我们借助一个三维人脸表情模型 FaceWarehouse (CAO, C. , WENG, Υ. , ZHOU, S. , TONG, Υ. , AND ZHOU, Κ. 2012. Facewarehouse: a 3d facial expression database for visual computing. Tech. rep. ) 来构建用户表情融合模型。 FaceWarehouse包括 150个不同背景下的个体数据, 每个个体数据包括 46个 FACS 表情形状。 FaceWarehouse 利用这些数据建立了一个包含两个属性, 即个体和表 情的双线性模型, 组成了一个三维核张量 G ( 11K 模型顶点 χ 50 个体 χ 45 表
n ) , 利用这一的核张量表示, 任意个体的任意表情 F可以利用张量收缩计算 ί得 到:
: w;J χ3 w exp
其中, 和 分别是张量中个体和表情系数的列向量 Cr是
FaceWarehouse的核张量, F则是通过收缩计算得到的表情。
我们使用两步来计算用户表情融合模型。 在第一步中, 对 "图像采集和标 定" 中采集的每一个图像, 我们找到一个变换矩阵 个体系数 和表情系数 w ,, 生成三维人脸形状, 使得其相应的三维特征点在图像±的投影与标定的二 维特征点吻合。 这可以通过优化下面的能量来达到要求:
其中, u )是第 张图像中第 个二维特征点位置, w是三维网格形状上对应 的顶点序号, nQ则表示借助摄像机投影矩阵 Q, 将三维空间点投影到图像坐标的 操作, w^,和 w ,分别是张量中个体和表情系数的列向量, 是 FaceWarehouse 的核张量。 我们可以使用坐标下降法来求解 M,, wL,; 和 w ,, 即每次保持其中 两个变量不变, 优化另外的一个变量, 迭代进行这一步骤直到结果收敛。
在第二步中, 由于我们采集的所有图像描述的是同一个人的不同姿势或不同 表情, 因此我们应该保证所有图像中个体系数, 即 —致, 因此我们固定第一步 中得到每个图片的变换矩阵 M,和表情系数 wL„, 同时在所有图像上计算一致的个 体系数 需要优化如下的能:
u
其中, 是统一的个体系数, w是采集的二维图像数目, 其余变量含义与上 式相同。
这两步的优化过程需要迭代计算直到结果收敛, 在一般情况下, 需要三次迭 代即可得到满意的结果。 一旦得到了一致的个体系数 w^, 就可以生成该用户的表 情融合模型, 如下:
B, = Cr χ2 χ3 (Vexp&i ),0 < ί' < 47 ;
其中, ϋ 是 FaceWarehouse 表情属性的截断变换矩阵, d,是表情系数向 量, 其第 项元素为 1而其余元素是 0, 是 FaceWarehouse的核张量, w id 一的个体系数。
2. 2 照相机内部参数标定
照相机投影矩阵描述的是将摄像机坐标系中的三维点投影到二维图像位置, 它完全依赖于照相机内部的参数, 可以被表达为如下的投影矩阵 Q:
其中参数 Λ和 Λ表示以长和宽方向像素为单位的焦距长度, 表示在 X和) 由 方向的偏移, 《。和 则表示图像原点的位置, 即光轴与图像平面的交点。 有很多 照相机标定的方法 (如 ZHANG, Z. 2000. A flexible new technique for camera cal ibration. IEEE Trans. Pattern Anal. Mach. Intel l. 22, 11, 1330 - 1334. ) 可以用来精确计算投影矩阵, 这些方法通常会借助一些标准标定目 标 (如棋盘格) 。
本发明使用一种简单的方法, 不借助特殊的标定目标, 而直接从用户采集的 数据中直接得到投影矩阵 Q。 本发明假设使用的照相机是一个理想的针孔照相 机, 其 / = = /y = 0, (Wq,Vq)就是图像的中心点, 这通过输入图像的大小可以直 接计算出来。 那么照相机的投影矩阵中只剩下一个未知数, 即/。 本发明假设不 同的 /, 利用假设值进行 "用户表情融合模型的生成" , 并计算最终所有采集图 片中拟合的人脸模型对应特征点投影和标定的特征点之间的误差。 该误差相对于 /的函数关系是一个凸函数, 即有一个最小值, 而在最小值两端都是单调。 这 样, 本发明使用二分法快速找到正确的 f 。
2. 3 训练数据构建
在前几步操作中得到了用户的表情融合模型, 同时在每个输入图像上有相应 的姿势变换矩阵和表情系数, 以此可以得到图像上的三维人脸形状:
46
Β0 +∑αΑ
i=l
其中, F是生成的三维人脸形状, M是姿势变换矩阵, 2?。是该用户的自然表 情形状, 是用户表情融合模型中的基本表情形状, 《,·是基本表情的系数。
通过选取三维人脸形状上对应的三维顶点位置, 就可以构成该图像的三维特 征点。 在实时的视频中, 由于人脸的外轮廓点随时在改变, 为了计算效率, 本发 明将外部轮廓的 15个特征点替换为内部的 15个特征点 (如附图 2所示) 。 我们 用 iS,。l来表示这些图像对应的三维特征点。
为了增强训练数据的通用性, 本发明需要增强采集的图像和它们对应的三维 特征点。 对每一个采集图像和它的三维特征点 我们将三维特征点 S,。沿着 照相机坐标系中的三个轴进行平移得到另外 m-1个三维特征点, 对每个 S,。都得到 集合 ,2≤j'≤ }。 增强的三维特征点对应的是另外的图像。 实际操作中, 本发明 并不真正的生成对应的图像, 而只是记录这些增强的三维特征点变换到原来特征 点 S,。的变换矩阵 M), 该矩阵与 , S,。一起可以提供新图像的完整信息, 隐式的 生成增强的图像。 经过数据增强, w个原始数据增强为 个, 我们定义它们为 {/;,M^¾}o 这些增强后的三维特征点集合 ¾,l≤ ≤w,l≤ j'≤ }称之为三维特征点 空间, 它描述了用户在三维空间中人脸特征点的变化范围。
对每个增强后的每组图像 /特征点数据, 本发明给它指定不同的初始化特征 点。 在选择训练使用的数据初始特征点时, 本发明同时考虑了数据的局部性和随 机性。 对每一组图像 /特征点 (/,,M^, ), 首先从 n个原始特征点集合 {S,°} 中找到 G个与 最近似的特征点, 在计算两个特征点的相似性, 首先将两个特征点的 中心对齐, 然后计算对应特征点之间的距离平方和。 我们将找到的最相似的特征 点集合称为记为 ,g,i≤ g < G]。 然后从每个 SLG的增强特征点中随机选取 H个特征 点, 记为 A,l≤i≤H} 。 我们将这些特征点作为该图像 /特征点 {/,,M ,S^的初始 化特征点集 ^。 这样, 本发明为每一对图像 /特征点 {/,,M ,S^找到了 个初始 化特征点。 每个训练数据将其表示为 (/,,M , Λ)。 其中 /,是二维图像, 是 特征点进行平移增强的变换矩阵, 是 /,对应^三维特征点, 而 S,GA是初始特征 占、、。
经过数据增强和训练集构建, 我们生成了 N = ^m.G.H个训练数据。 在我们 所有例子中, 我们选取 = 9,G = 5,H=4。 为了简化, 我们在之后将这 N个训练数 据称为 {(/,,M ,S,,S, }。
2.4 回归器训练
给定了上述的 N个训练数据 /,,M ,S,,S, }, 本发明利用图像 /,中的信息, 训 练生成一个从初始化特征点 到对应特征点 S,的回归函数。 本发明使用两层的级 连回归, 其中在第一层中有 Γ级弱分类器, 而在每一个弱分类器中又有 级原子 分类器。
第一层的级连回归中, 本发明产生一组用于原子分类器构建的序号对特征。 首先利用当前特征点 S和图像 /,, 计算得到外观向量: 在当前特征点 S 空间范围
随机选取的 P个采样点, 其中每一个采样点 p的位置都表示为 S 中某一个特征点 位置加上一个偏移 dP ; 然后利用 nQ (M p)将采样点 p投影到图像上, 最后从图像 /,取到对应像素点的颜色值。 这 P个颜色值就组成该训练数据在第一层级连回归 中的外观向量 .。 对于每一个外观向量 ., 通过计算两两不同位置元素的差, 可 以产生 P2个序列号特征。
在第二层的每一个原子分类器中, 要在第一层生成的 P2个序列号特征中, 寻 找有效特征, 并以此将训练数据进行分类。 对每一个训练数据 (/,,M ,S,,S;), 首 先计算当前特征点 S和真实特征点 S,之间的差异, 然后将这些差异投影到一个随 机方向上产生一个标量, 本将这些标量看做随机变量, 从 P2序列号特征中找到与 这个随机变量相关性最大的特征。 重复这一步骤 F次以产生 F个不同的特征, 根 据这 F个特征, 生成该原子分类器。
在每一个原子分类器中, F个特征被赋予了一个随机的阈值, 这些阈值可以 将所有的训练数据分成 2f类, 对每一个训练数据, 我们比较序列号计算得的特征 值和阈值, 来决定该训练数据属于哪一类。 在每一个类 b中, 本发明将所有落在 这一类中的数据集合称为为 Ω¾, 并计算这一类中的特征点回归输出为:
其中, |Ω¾|表示这一类中训练数据的个数, S,为训练数据的真实特征点, S 是训练数据的当前特征点, 是松弛系数, 用于防止当落在该类中的训练数据过 少导致过拟合现象。
当我们产生原子分类器后, 我们根据原子分类器更新当前所有的训练数据。 即在原子分类器的每个类 b中, 将其对应的回归输出加到落在这一类中训练数据 的当前特征点中, 即 S = S + S¾。
回归器的训练在会迭代执行 Γ次, 每一次中生成 个级连的原子分类器, 组 成弱分类器, 迭代优化回归输出。 这 Γ个级连的弱分类器组成一个强分类器, 即 我们需要的回归器。 参数配置本发明选择 ^ lO SOO^ z ^O^ z ^^^ SO c
3.三维特征点跟踪: 对于用户实时输入的图像, 结合上一帧的三维人脸特征 点, 利用数据预处理步骤中得到回归器, 结合上一帧的三维特征点 S', 本发明可 以实时定位当前帧的三维特征点。
首先在原始特征点集合 {S,°}中找到与 S '最近似的特征点 , 然后通过一个刚 体的旋转和平移 (M" ) , 将 S '变换到 的位置, 记变换后的上一帧特征点为 S '*。 然后在训练集合的三维特征点空间 ¾,l≤ ≤w,l≤ j'≤ }中找到与 S '*最接近的 L个特征点集合 , 并将每一个&作为初始化特征点输入通过整个回归器。
与回归器训练类似, 使用回归器来定位特征点时分为两层级连结构。 在第一 层回归中, 首先根据当前帧图像 /, 当前特征点& , 变换矩阵1^的逆矩阵, 以及 在训练过程中记录的偏移 {d ^, 得到外观向量 V。 在第二层中, 根据每个原子分 类器中记录的序列号计算特征并与阈值进行比较, 来确定属于哪个类, 并得到该 类的回归输出 S¾。 最后利用这个输出来更新当前特征点: S^ St + SS
本发明对 L个初始特征点都通过回归器得到 L个输出特征点, 最后, 对这些 输出特征点取中值操作, 得到最终的结果。 注意到这个特征点是三维特征点空间 中的, 需要利用变换矩阵!^的逆矩阵将其变换到原来的图像位置中去。 输入的二 维图像和定位得到的三维特征点结果如附图 3所示。
4.姿势表情参数化: 利用特征点的三维位置, 结合数据预处理中得到的用户 表情融合模型, 迭代优化以得到的人脸姿势和表情的参数化表达。
在上一步中得到了当前帧的三维特征点位置, 本发明利用它们对当前帧人脸的动 作进行参数化。 人脸的动作主要分为两个部分: 由变换矩阵 M表示的刚性人脸姿 势, 和由表情融合系数 a表示的人脸非刚性表情。 这两部参数可以通过优化以下 的匹配能量得到:
其中, ^^是^中第 个特征点的三维位置, 是三维人脸形状上对应的顶点 序号, 。是用户的自然表情人脸形状, 是表情融合模型中的其他基本表情人脸 形状, 《是基本表情的系数, M是表示人脸姿势的变换矩阵。 与 WEISE, T. , BOUAZIZ, S. , LI, H. , AND PAULY, M. 2011. Realt ime performance-based facial animat ion. ACM Trans. Graph. 30, 4 (July) , 77 : 1 - 77 : 10. 类似, 本 发明使用一个动画先验来增强跟踪过程中的时域连续性。 给定前《帧的表情系数 向量 ..,a- "}, 将其与当前帧的系数 a结合成一个单独的向量 (a,A„)
, 本发明将该向量的概率分布描述为一个高斯混合模型:
S
p (a, An )二 JISN AJ 5, Covs );
=1
其中 AT是高斯分布符号, 是高斯模型的权重系数, s是高斯模型中的平均 值, bVs则是变量的协方差矩阵。 该混合高斯模型可以利用一些事先产生的表情 动画序列训练得到 (WEISE, T. , BOUAZIZ, S. , LI, H. , AND PAULY, M. 2011. Realtime performance- based facial animation. ACM Trans. Graph. 30, 4 (July) , 77 : 1 - 77 : 10. ) 。 该高斯混合模型可以描述用于帧间连续性的一个能
其中, 我们称之为动画先验能量, Ρ Α„)是上述的高斯混合模型。 本发明将该能量与匹配能量结合起来, 形成了最终的能量描述:
Ef ω prior ^ prior
其中《P„OT是权重系数, 用于权衡跟踪准确性和时域连续性, E,是上述的匹配 是动画先验能量。 本发明利用两步迭代的方法来对能量^进行优化。 在第一步中使用上一帧的表情系数 a作为当前帧的初始值并保持不变, 然后对对 应点对分布的协方差矩阵利用奇异值分解 (SVD ) 计算刚性姿势, 即变换矩阵 M。 然后在第二步中本发明固定 M, 然后利用梯度下降法来计算表情系数3。 本 发明迭代执行这两步, 直到结果收敛, 通常情况下经过两次迭代就可以得到满意 的结果。 得到了人脸姿势和表情的参数化表达后, 我们就可以得到对应的用户三 维人脸形状, 结果如附图 4所示。
5.替身驱动: 将人脸姿势和表情参数映射到虚拟替身上, 用以驱动动画角色 进行人脸动画。
其中 M是人脸姿势的变换矩阵, ¾是该用户的自然表情人脸形状 Α,Α,.., ^ 是表情融合模型中的其他基本表情人脸形状, 《,是基本表情的系数, 而 )则是最 终替身的人脸形状。
这样就完成了对动画角色的驱动, 结果如附图 5所示。
实施例
发明人在一台配备 Intel Core i 7 ( 3. 5GHz ) 中央处理器的台式计算机, 及 一个以 30 fps提供 640 x 480分辨率的网络摄像头上实现了本发明的实施实例。 实施中使用具体实施方式中提及的参数设置, 得到了附图中的结果。 实践中在普 通电脑上只需要少于 15毫秒的时间即可完成一帧的捕获、 参数化和替身映射。
发明人邀请了一些用户来测试本方法的原型系统。 结果表明, 在我们目前的 硬件配置上, 本发明可以实时的处理各种大幅度的姿势旋转, 各种夸张的表情, 得到与用户输入非常接近的动画效果, 具有很好的用户体验。 同时, 在不同的光 照条件下, 如办公室, 阳光直射的室外, 光线较暗的宾馆房间, 本发明都可以得 到满意的结果。
Claims
1. 一种基于单个视频摄像机的实时人脸动画方法, 其特征在于, 包括以下步骤:
( 1 ) 图像采集和标定: 利用视频摄像机拍摄用户的多幅具有不同姿势和表情的二 维图像, 对每个图像利用二维图像回归器得到对应的二维人脸特征点, 对自动检 测得到的不准确特征点进行手动调整;
(2 ) 数据预处理: 利用标定好二维人脸特征点的图像, 进行用户表情融合模型生 成和摄像机内部参数标定, 并由此得到图像的三维特征点; 利用三维特征点和步 骤 1采集的二维图像, 训练获得从二维图像到三维特征点的回归器;
(3 ) 三维特征点跟踪: 用户使用视频摄像机实时输入图像, 对于输入的图像, 结 合上一帧的三维人脸特征点, 利用步骤 2 中得到回归器, 实时定位当前帧中三维 人脸特征点;
(4) 姿势表情参数化: 利用三维人脸特征点位置, 结合步骤 2中得到的用户表情 融合模型, 迭代优化以得到人脸姿势和表情的参数化表达;
(5 ) 替身驱动: 将人脸姿势和表情参数映射到虚拟替身上, 用以驱动动画角色进 行人脸动画。
2. 根据权利要求 1所述的实时人脸动画方法, 其特征在于, 所述步骤 1主要包括 以下子步骤:
( 1. 1 ) 用户模仿做出相应表情和姿势, 包括 15 种自然表情下的不同人头姿势, 和 3个姿势下的 15种不同表情, 共 60组不同的姿势表情数据, 利用视频摄像机 拍摄相应的二维图像;
( 1. 2 ) 利用二维图像回归器对每一个二维图像分别进行自动的二维人脸特征点标 定;
( 1. 3 ) 用户对自动标定的人脸特征点中不满意的部分, 对其进行简单的拖拽操 作, 进行人工修复。
3. 根据权利要求 1所述的实时人脸动画方法, 其特征在于, 所述步骤 2主要包括 以下子步骤:
(2. 1 ) 利用已有的三维人脸表情数据库, 对于每一个标定了二维人脸特征点的二 维图像进行拟合, 使用最小二乘方法计算相应的刚性参数、 个体系数和表情系
数; 之后对所有二维图像进行统一优化, 得到统一的个体系数, 计算得到用户的 表情融合模型;
(2. 2) 对针孔照相机模型进行简化假设, 将其简化到只包括一个未知参数, 使用 二分法来确定最合适的照相机参数;
(2. 3 ) 基于上述步骤得到的用户表情融合模型和照相机参数, 拟合每个图像中人 脸刚性参数和表情系数, 得到三维人脸特征点位置; 其后对二维图像和其对应的 三维特征点进行数据增强操作;
(2. 4) 利用步骤 2. 3中生成的二维图像和三维人脸特征点, 训练获得一个利用二 维图像信息生成三维人脸特征点的回归器。
4. 根据权利要求 1所述的实时人脸动画方法, 其特征在于, 所述步骤 3主要包括 以下子步骤:
(3. 1 ) 运行时, 先使用上一帧的三维特征点, 通过一个刚性变换, 将其转换到原 训练数据中与其最接近的特征点位置, 然后在原训练数据中的三维特征点中找到 一组与转换后特征点最接近的一组特征点作为初始特征点;
( 3. 2 ) 对每个当前特征点, 根据特征点位置, 在当前帧图像上采样得到外观向
(3. 3) 在每个原子分类器中, 根据序列对在步骤 3. 2中外观向量计算对应的特征 值, 并根据特征值定位相应的分类, 并使用分类中对应的输出更新当前特征点位 置; 依次通过所有的原子分类器, 得到了回归器给出的输出结果;
(3. 4) 对每个初始特征点, 用步骤 3. 2和步骤 3. 3得到定位的结果, 然后对这些 结果取中值操作, 得到最终的结果。
5. 根据权利要求 4所述的实时人脸动画方法, 其特征在于, 所述步骤 4主要包括 以下子步骤:
(4. 1 ) 保持表情系数不变, 利用奇异值分解算法计算当前人脸形状的刚性姿势, 使得形状上对应的特征点与权利要求 4 中描述的三维人脸特征点之间的误差最 小;
(4. 2 ) 保持姿势不变, 利用梯度下降算法, 拟合当前表情系数, 使得形状上对应 的特征点与权利要求 4中描述的三维人脸特征点之间的误差最小;
(4. 3) 迭代执行步骤 4. 1和 4. 2直到收敛, 最终得到参数化的人脸姿势系数和表 情系数。
6. 根据权利要求 1所述的实时人脸动画方法, 其特征在于, 所述步骤 5主要包括 以下子步骤:
(5.1) 将参数化的表情系数映射到替身的表情融合模型上, 生成对应的人脸表情 形状;
(5.2) 为生成的人脸表情形状赋予参数化的姿势, 得到与用户输入图像匹配的人 脸动作。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US14/517,758 US9361723B2 (en) | 2013-02-02 | 2014-10-17 | Method for real-time face animation based on single video camera |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201310047850.2A CN103093490B (zh) | 2013-02-02 | 2013-02-02 | 基于单个视频摄像机的实时人脸动画方法 |
| CN201310047850.2 | 2013-02-02 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US14/517,758 Continuation US9361723B2 (en) | 2013-02-02 | 2014-10-17 | Method for real-time face animation based on single video camera |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2014117446A1 true WO2014117446A1 (zh) | 2014-08-07 |
Family
ID=48206021
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2013/075117 Ceased WO2014117446A1 (zh) | 2013-02-02 | 2013-05-03 | 基于单个视频摄像机的实时人脸动画方法 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US9361723B2 (zh) |
| CN (1) | CN103093490B (zh) |
| WO (1) | WO2014117446A1 (zh) |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109702741A (zh) * | 2018-12-26 | 2019-05-03 | 中国科学院电子学研究所 | 基于自监督学习神经网络的机械臂视觉抓取系统及方法 |
| CN110348387A (zh) * | 2019-07-12 | 2019-10-18 | 腾讯科技(深圳)有限公司 | 一种图像数据处理方法、装置以及计算机可读存储介质 |
| CN111626158A (zh) * | 2020-05-14 | 2020-09-04 | 闽江学院 | 一种基于自适应下降回归的面部标记点跟踪原型设计方法 |
| CN112669424A (zh) * | 2020-12-24 | 2021-04-16 | 科大讯飞股份有限公司 | 一种表情动画生成方法、装置、设备及存储介质 |
| CN116309986A (zh) * | 2023-02-08 | 2023-06-23 | 杭州相芯科技有限公司 | 一种实时动漫头像驱动方法及系统 |
Families Citing this family (120)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| TWI439960B (zh) | 2010-04-07 | 2014-06-01 | 蘋果公司 | 虛擬使用者編輯環境 |
| WO2013152453A1 (en) | 2012-04-09 | 2013-10-17 | Intel Corporation | Communication using interactive avatars |
| CN103093490B (zh) | 2013-02-02 | 2015-08-26 | 浙江大学 | 基于单个视频摄像机的实时人脸动画方法 |
| US9378576B2 (en) * | 2013-06-07 | 2016-06-28 | Faceshift Ag | Online modeling for real-time facial animation |
| US9547808B2 (en) * | 2013-07-17 | 2017-01-17 | Emotient, Inc. | Head-pose invariant recognition of facial attributes |
| GB2518589B (en) * | 2013-07-30 | 2019-12-11 | Holition Ltd | Image processing |
| WO2015042867A1 (zh) * | 2013-09-27 | 2015-04-02 | 中国科学院自动化研究所 | 一种基于单摄像头与运动捕捉数据的人脸表情编辑方法 |
| CN106104633A (zh) * | 2014-03-19 | 2016-11-09 | 英特尔公司 | 面部表情和/或交互驱动的化身装置和方法 |
| CN103908064A (zh) * | 2014-04-03 | 2014-07-09 | 安徽海聚信息科技有限责任公司 | 一种矫正坐姿的智能书桌及其矫正方法 |
| CN103908063A (zh) * | 2014-04-03 | 2014-07-09 | 安徽海聚信息科技有限责任公司 | 一种矫正坐姿的智能书桌及其矫正方法 |
| CN103942822B (zh) * | 2014-04-11 | 2017-02-01 | 浙江大学 | 一种基于单视频摄像机的面部特征点跟踪和人脸动画方法 |
| KR102103939B1 (ko) * | 2014-07-25 | 2020-04-24 | 인텔 코포레이션 | 머리 회전을 갖는 아바타 얼굴 표정 애니메이션 |
| CN104217454B (zh) * | 2014-08-21 | 2017-11-03 | 中国科学院计算技术研究所 | 一种视频驱动的人脸动画生成方法 |
| WO2016030305A1 (en) * | 2014-08-29 | 2016-03-03 | Thomson Licensing | Method and device for registering an image to a model |
| CN104268918B (zh) * | 2014-10-09 | 2015-06-10 | 佛山精鹰传媒股份有限公司 | 一种二维动画与三维立体动画融合的处理方法 |
| WO2016070354A1 (en) * | 2014-11-05 | 2016-05-12 | Intel Corporation | Avatar video apparatus and method |
| WO2016090605A1 (en) * | 2014-12-11 | 2016-06-16 | Intel Corporation | Avatar selection mechanism |
| CN107004288B (zh) | 2014-12-23 | 2022-03-01 | 英特尔公司 | 非面部特征的面部动作驱动的动画 |
| US9824502B2 (en) | 2014-12-23 | 2017-11-21 | Intel Corporation | Sketch selection for rendering 3D model avatar |
| WO2016101131A1 (en) * | 2014-12-23 | 2016-06-30 | Intel Corporation | Augmented facial animation |
| CN105809660B (zh) * | 2014-12-29 | 2019-06-25 | 联想(北京)有限公司 | 一种信息处理方法及电子设备 |
| KR102450865B1 (ko) * | 2015-04-07 | 2022-10-06 | 인텔 코포레이션 | 아바타 키보드 |
| CN107924452B (zh) * | 2015-06-26 | 2022-07-19 | 英特尔公司 | 用于图像中的脸部对准的组合形状回归 |
| KR102146398B1 (ko) | 2015-07-14 | 2020-08-20 | 삼성전자주식회사 | 3차원 컨텐츠 생성 장치 및 그 3차원 컨텐츠 생성 방법 |
| US9865072B2 (en) | 2015-07-23 | 2018-01-09 | Disney Enterprises, Inc. | Real-time high-quality facial performance capture |
| CN105184249B (zh) * | 2015-08-28 | 2017-07-18 | 百度在线网络技术(北京)有限公司 | 用于人脸图像处理的方法和装置 |
| CN105139450B (zh) * | 2015-09-11 | 2018-03-13 | 重庆邮电大学 | 一种基于人脸模拟的三维虚拟人物构建方法及系统 |
| CN108140105A (zh) * | 2015-09-29 | 2018-06-08 | 比纳里虚拟现实技术有限公司 | 具有脸部表情检测能力的头戴式显示器 |
| WO2017101094A1 (en) * | 2015-12-18 | 2017-06-22 | Intel Corporation | Avatar animation system |
| US20170186208A1 (en) * | 2015-12-28 | 2017-06-29 | Shing-Tung Yau | 3d surface morphing method based on conformal parameterization |
| CN106997450B (zh) * | 2016-01-25 | 2020-07-17 | 深圳市微舞科技有限公司 | 一种表情迁移中的下巴运动拟合方法和电子设备 |
| CN107292219A (zh) * | 2016-04-01 | 2017-10-24 | 掌赢信息科技(上海)有限公司 | 一种驱动眼睛运动的方法及电子设备 |
| CN106023288B (zh) * | 2016-05-18 | 2019-11-15 | 浙江大学 | 一种基于图像的动态替身构造方法 |
| US9912860B2 (en) | 2016-06-12 | 2018-03-06 | Apple Inc. | User interface for camera effects |
| CN108475424B (zh) * | 2016-07-12 | 2023-08-29 | 微软技术许可有限责任公司 | 用于3d面部跟踪的方法、装置和系统 |
| WO2018033137A1 (zh) * | 2016-08-19 | 2018-02-22 | 北京市商汤科技开发有限公司 | 在视频图像中展示业务对象的方法、装置和电子设备 |
| WO2018053682A1 (en) * | 2016-09-20 | 2018-03-29 | Intel Corporation | Animation simulation of biomechanics |
| CN110109592B (zh) | 2016-09-23 | 2022-09-23 | 苹果公司 | 头像创建和编辑 |
| CN106600667B (zh) * | 2016-12-12 | 2020-04-21 | 南京大学 | 一种基于卷积神经网络的视频驱动人脸动画方法 |
| CN106778628A (zh) * | 2016-12-21 | 2017-05-31 | 张维忠 | 一种基于tof深度相机的面部表情捕捉方法 |
| CN106919899B (zh) * | 2017-01-18 | 2020-07-28 | 北京光年无限科技有限公司 | 基于智能机器人的模仿人脸表情输出的方法和系统 |
| TWI724092B (zh) * | 2017-01-19 | 2021-04-11 | 香港商斑馬智行網絡(香港)有限公司 | 確定融合係數的方法和裝置 |
| CN106952217B (zh) * | 2017-02-23 | 2020-11-17 | 北京光年无限科技有限公司 | 面向智能机器人的面部表情增强方法和装置 |
| US10572720B2 (en) * | 2017-03-01 | 2020-02-25 | Sony Corporation | Virtual reality-based apparatus and method to generate a three dimensional (3D) human face model using image and depth data |
| CN107067429A (zh) * | 2017-03-17 | 2017-08-18 | 徐迪 | 基于深度学习的人脸三维重建和人脸替换的视频编辑系统及方法 |
| CN108876879B (zh) * | 2017-05-12 | 2022-06-14 | 腾讯科技(深圳)有限公司 | 人脸动画实现的方法、装置、计算机设备及存储介质 |
| DK180859B1 (en) | 2017-06-04 | 2022-05-23 | Apple Inc | USER INTERFACE CAMERA EFFECTS |
| US20180357819A1 (en) * | 2017-06-13 | 2018-12-13 | Fotonation Limited | Method for generating a set of annotated images |
| CN108304758B (zh) * | 2017-06-21 | 2020-08-25 | 腾讯科技(深圳)有限公司 | 人脸特征点跟踪方法及装置 |
| US10311624B2 (en) | 2017-06-23 | 2019-06-04 | Disney Enterprises, Inc. | Single shot capture to animated vr avatar |
| CN107316020B (zh) * | 2017-06-26 | 2020-05-08 | 司马大大(北京)智能系统有限公司 | 人脸替换方法、装置及电子设备 |
| CN109215131B (zh) * | 2017-06-30 | 2021-06-01 | Tcl科技集团股份有限公司 | 虚拟人脸的驱动方法及装置 |
| JP7047848B2 (ja) * | 2017-10-20 | 2022-04-05 | 日本電気株式会社 | 顔三次元形状推定装置、顔三次元形状推定方法、及び、顔三次元形状推定プログラム |
| US10586368B2 (en) | 2017-10-26 | 2020-03-10 | Snap Inc. | Joint audio-video facial animation system |
| US10430642B2 (en) | 2017-12-07 | 2019-10-01 | Apple Inc. | Generating animated three-dimensional models from captured images |
| CN109903360A (zh) * | 2017-12-08 | 2019-06-18 | 浙江舜宇智能光学技术有限公司 | 三维人脸动画控制系统及其控制方法 |
| CN108171174A (zh) * | 2017-12-29 | 2018-06-15 | 盎锐(上海)信息科技有限公司 | 基于3d摄像的训练方法及装置 |
| US10949648B1 (en) * | 2018-01-23 | 2021-03-16 | Snap Inc. | Region-based stabilized face tracking |
| CN108648280B (zh) * | 2018-04-25 | 2023-03-31 | 深圳市商汤科技有限公司 | 虚拟角色驱动方法及装置、电子设备和存储介质 |
| US10607065B2 (en) * | 2018-05-03 | 2020-03-31 | Adobe Inc. | Generation of parameterized avatars |
| US12033296B2 (en) | 2018-05-07 | 2024-07-09 | Apple Inc. | Avatar creation user interface |
| US10375313B1 (en) * | 2018-05-07 | 2019-08-06 | Apple Inc. | Creative camera |
| US11722764B2 (en) | 2018-05-07 | 2023-08-08 | Apple Inc. | Creative camera |
| DK179874B1 (en) | 2018-05-07 | 2019-08-13 | Apple Inc. | USER INTERFACE FOR AVATAR CREATION |
| CN108648251B (zh) * | 2018-05-15 | 2022-05-24 | 奥比中光科技集团股份有限公司 | 3d表情制作方法及系统 |
| CN108805056B (zh) * | 2018-05-29 | 2021-10-08 | 电子科技大学 | 一种基于3d人脸模型的摄像监控人脸样本扩充方法 |
| US10825223B2 (en) | 2018-05-31 | 2020-11-03 | Microsoft Technology Licensing, Llc | Mixed reality animation |
| DK201870623A1 (en) | 2018-09-11 | 2020-04-15 | Apple Inc. | User interfaces for simulated depth effects |
| US11128792B2 (en) | 2018-09-28 | 2021-09-21 | Apple Inc. | Capturing and displaying images with multiple focal planes |
| US11321857B2 (en) | 2018-09-28 | 2022-05-03 | Apple Inc. | Displaying and editing images with depth information |
| US10817365B2 (en) | 2018-11-09 | 2020-10-27 | Adobe Inc. | Anomaly detection for incremental application deployments |
| CN109493403A (zh) * | 2018-11-13 | 2019-03-19 | 北京中科嘉宁科技有限公司 | 一种基于运动单元表情映射实现人脸动画的方法 |
| CN109621418B (zh) * | 2018-12-03 | 2022-09-30 | 网易(杭州)网络有限公司 | 一种游戏中虚拟角色的表情调整及制作方法、装置 |
| US10943352B2 (en) * | 2018-12-17 | 2021-03-09 | Palo Alto Research Center Incorporated | Object shape regression using wasserstein distance |
| KR102686099B1 (ko) * | 2019-01-18 | 2024-07-19 | 스냅 아이엔씨 | 모바일 디바이스에서 사실적인 머리 회전들 및 얼굴 애니메이션 합성을 위한 방법들 및 시스템들 |
| US11107261B2 (en) | 2019-01-18 | 2021-08-31 | Apple Inc. | Virtual avatar animation based on facial feature movement |
| KR102639725B1 (ko) * | 2019-02-18 | 2024-02-23 | 삼성전자주식회사 | 애니메이티드 이미지를 제공하기 위한 전자 장치 및 그에 관한 방법 |
| US11770601B2 (en) | 2019-05-06 | 2023-09-26 | Apple Inc. | User interfaces for capturing and managing visual media |
| US11706521B2 (en) | 2019-05-06 | 2023-07-18 | Apple Inc. | User interfaces for capturing and managing visual media |
| US10645294B1 (en) | 2019-05-06 | 2020-05-05 | Apple Inc. | User interfaces for capturing and managing visual media |
| CN110288680A (zh) * | 2019-05-30 | 2019-09-27 | 盎锐(上海)信息科技有限公司 | 影像生成方法及移动终端 |
| US11120569B2 (en) * | 2019-06-24 | 2021-09-14 | Synaptics Incorporated | Head pose estimation |
| US12266042B2 (en) * | 2019-08-16 | 2025-04-01 | Sony Group Corporation | Image processing device and image processing method |
| KR102618732B1 (ko) * | 2019-08-27 | 2023-12-27 | 엘지전자 주식회사 | 얼굴 인식 활용 단말기 및 얼굴 인식 활용 방법 |
| CN110599573B (zh) * | 2019-09-03 | 2023-04-11 | 电子科技大学 | 一种基于单目相机的人脸实时交互动画的实现方法 |
| KR102770795B1 (ko) * | 2019-09-09 | 2025-02-21 | 삼성전자주식회사 | 3d 렌더링 방법 및 장치 |
| CN110677598B (zh) * | 2019-09-18 | 2022-04-12 | 北京市商汤科技开发有限公司 | 视频生成方法、装置、电子设备和计算机存储介质 |
| CN110941332A (zh) * | 2019-11-06 | 2020-03-31 | 北京百度网讯科技有限公司 | 表情驱动方法、装置、电子设备及存储介质 |
| CN111105487B (zh) * | 2019-12-19 | 2020-12-22 | 华中师范大学 | 一种虚拟教师系统中的面部合成方法及装置 |
| CN111243065B (zh) * | 2019-12-26 | 2022-03-11 | 浙江大学 | 一种语音信号驱动的脸部动画生成方法 |
| CN111311712B (zh) * | 2020-02-24 | 2023-06-16 | 北京百度网讯科技有限公司 | 视频帧处理方法和装置 |
| CN113436301B (zh) * | 2020-03-20 | 2024-04-09 | 华为技术有限公司 | 拟人化3d模型生成的方法和装置 |
| DK202070625A1 (en) | 2020-05-11 | 2022-01-04 | Apple Inc | User interfaces related to time |
| US11921998B2 (en) | 2020-05-11 | 2024-03-05 | Apple Inc. | Editing features of an avatar |
| US11054973B1 (en) | 2020-06-01 | 2021-07-06 | Apple Inc. | User interfaces for managing media |
| US11494932B2 (en) * | 2020-06-02 | 2022-11-08 | Naver Corporation | Distillation of part experts for whole-body pose estimation |
| CN111798551B (zh) * | 2020-07-20 | 2024-06-04 | 网易(杭州)网络有限公司 | 虚拟表情生成方法及装置 |
| CN112116699B (zh) * | 2020-08-14 | 2023-05-16 | 浙江工商大学 | 一种基于3d人脸跟踪的实时真人虚拟试发方法 |
| CN111970535B (zh) * | 2020-09-25 | 2021-08-31 | 魔珐(上海)信息科技有限公司 | 虚拟直播方法、装置、系统及存储介质 |
| US11212449B1 (en) | 2020-09-25 | 2021-12-28 | Apple Inc. | User interfaces for media capture and management |
| CN112487993A (zh) * | 2020-12-02 | 2021-03-12 | 重庆邮电大学 | 改进型级联回归人脸特征点定位算法 |
| CN112836089B (zh) * | 2021-01-28 | 2023-08-22 | 浙江大华技术股份有限公司 | 运动轨迹的确认方法及装置、存储介质、电子装置 |
| CN112819947B (zh) * | 2021-02-03 | 2025-02-11 | Oppo广东移动通信有限公司 | 三维人脸的重建方法、装置、电子设备以及存储介质 |
| US11756334B2 (en) * | 2021-02-25 | 2023-09-12 | Qualcomm Incorporated | Facial expression recognition |
| US11587288B2 (en) | 2021-03-15 | 2023-02-21 | Tencent America LLC | Methods and systems for constructing facial position map |
| US11778339B2 (en) | 2021-04-30 | 2023-10-03 | Apple Inc. | User interfaces for altering visual media |
| US11539876B2 (en) | 2021-04-30 | 2022-12-27 | Apple Inc. | User interfaces for altering visual media |
| CN113343761A (zh) * | 2021-05-06 | 2021-09-03 | 武汉理工大学 | 一种基于生成对抗的实时人脸表情迁移方法 |
| US12112024B2 (en) | 2021-06-01 | 2024-10-08 | Apple Inc. | User interfaces for managing media styles |
| US11776190B2 (en) | 2021-06-04 | 2023-10-03 | Apple Inc. | Techniques for managing an avatar on a lock screen |
| CN113453034B (zh) * | 2021-06-29 | 2023-07-25 | 上海商汤智能科技有限公司 | 数据展示方法、装置、电子设备以及计算机可读存储介质 |
| CN113505717B (zh) * | 2021-07-17 | 2022-05-31 | 桂林理工大学 | 一种基于人脸面部特征识别技术的在线通行系统 |
| CN113633983B (zh) * | 2021-08-16 | 2024-03-15 | 上海交通大学 | 虚拟角色表情控制的方法、装置、电子设备及介质 |
| CN114415844B (zh) * | 2021-12-17 | 2026-02-10 | 广西壮族自治区公众信息产业有限公司 | 一种动态键盘自动生成表情的输入方法 |
| FR3132583A1 (fr) * | 2022-02-10 | 2023-08-11 | Orange | Création perfectionnée d’un avatar |
| CN114882197B (zh) * | 2022-05-10 | 2023-05-05 | 贵州多彩宝互联网服务有限公司 | 一种基于图神经网络的高精度三维人脸重建方法 |
| CN115330947A (zh) * | 2022-08-12 | 2022-11-11 | 百果园技术(新加坡)有限公司 | 三维人脸重建方法及其装置、设备、介质、产品 |
| US12287913B2 (en) | 2022-09-06 | 2025-04-29 | Apple Inc. | Devices, methods, and graphical user interfaces for controlling avatars within three-dimensional environments |
| CN115393486B (zh) * | 2022-10-27 | 2023-03-24 | 科大讯飞股份有限公司 | 虚拟形象的生成方法、装置、设备及存储介质 |
| CN116071470B (zh) * | 2022-12-23 | 2026-03-10 | 天翼爱动漫文化传媒有限公司 | 一种用于合影的虚拟人姿态生成方法 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6956569B1 (en) * | 2000-03-30 | 2005-10-18 | Nec Corporation | Method for matching a two dimensional image to one of a plurality of three dimensional candidate models contained in a database |
| CN101303772A (zh) * | 2008-06-20 | 2008-11-12 | 浙江大学 | 一种基于单幅图像的非线性三维人脸建模方法 |
| CN101311966A (zh) * | 2008-06-20 | 2008-11-26 | 浙江大学 | 一种基于运行传播和Isomap分析的三维人脸动画编辑与合成方法 |
| CN102831382A (zh) * | 2011-06-15 | 2012-12-19 | 北京三星通信技术研究有限公司 | 人脸跟踪设备和方法 |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2009067560A1 (en) * | 2007-11-20 | 2009-05-28 | Big Stage Entertainment, Inc. | Systems and methods for generating 3d head models and for using the same |
| KR20110021330A (ko) * | 2009-08-26 | 2011-03-04 | 삼성전자주식회사 | 3차원 아바타 생성 장치 및 방법 |
| CN102103756B (zh) * | 2009-12-18 | 2012-10-03 | 华为技术有限公司 | 支持姿态偏转的人脸数字图像漫画夸张方法、装置及系统 |
| CN101783026B (zh) * | 2010-02-03 | 2011-12-07 | 北京航空航天大学 | 三维人脸肌肉模型的自动构造方法 |
| CN102376100A (zh) * | 2010-08-20 | 2012-03-14 | 北京盛开互动科技有限公司 | 基于单张照片的人脸动画方法 |
| CN101944238B (zh) * | 2010-09-27 | 2011-11-23 | 浙江大学 | 基于拉普拉斯变换的数据驱动人脸表情合成方法 |
| KR101514327B1 (ko) * | 2010-11-04 | 2015-04-22 | 한국전자통신연구원 | 얼굴 아바타 생성 장치 및 방법 |
| WO2012126135A1 (en) * | 2011-03-21 | 2012-09-27 | Intel Corporation | Method of augmented makeover with 3d face modeling and landmark alignment |
| CN103093490B (zh) | 2013-02-02 | 2015-08-26 | 浙江大学 | 基于单个视频摄像机的实时人脸动画方法 |
-
2013
- 2013-02-02 CN CN201310047850.2A patent/CN103093490B/zh active Active
- 2013-05-03 WO PCT/CN2013/075117 patent/WO2014117446A1/zh not_active Ceased
-
2014
- 2014-10-17 US US14/517,758 patent/US9361723B2/en active Active
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6956569B1 (en) * | 2000-03-30 | 2005-10-18 | Nec Corporation | Method for matching a two dimensional image to one of a plurality of three dimensional candidate models contained in a database |
| CN101303772A (zh) * | 2008-06-20 | 2008-11-12 | 浙江大学 | 一种基于单幅图像的非线性三维人脸建模方法 |
| CN101311966A (zh) * | 2008-06-20 | 2008-11-26 | 浙江大学 | 一种基于运行传播和Isomap分析的三维人脸动画编辑与合成方法 |
| CN102831382A (zh) * | 2011-06-15 | 2012-12-19 | 北京三星通信技术研究有限公司 | 人脸跟踪设备和方法 |
Cited By (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109702741A (zh) * | 2018-12-26 | 2019-05-03 | 中国科学院电子学研究所 | 基于自监督学习神经网络的机械臂视觉抓取系统及方法 |
| CN110348387A (zh) * | 2019-07-12 | 2019-10-18 | 腾讯科技(深圳)有限公司 | 一种图像数据处理方法、装置以及计算机可读存储介质 |
| CN110348387B (zh) * | 2019-07-12 | 2023-06-27 | 腾讯科技(深圳)有限公司 | 一种图像数据处理方法、装置以及计算机可读存储介质 |
| CN111626158A (zh) * | 2020-05-14 | 2020-09-04 | 闽江学院 | 一种基于自适应下降回归的面部标记点跟踪原型设计方法 |
| CN111626158B (zh) * | 2020-05-14 | 2023-04-07 | 闽江学院 | 一种基于自适应下降回归的面部标记点跟踪原型设计方法 |
| CN112669424A (zh) * | 2020-12-24 | 2021-04-16 | 科大讯飞股份有限公司 | 一种表情动画生成方法、装置、设备及存储介质 |
| CN112669424B (zh) * | 2020-12-24 | 2024-05-31 | 科大讯飞股份有限公司 | 一种表情动画生成方法、装置、设备及存储介质 |
| CN116309986A (zh) * | 2023-02-08 | 2023-06-23 | 杭州相芯科技有限公司 | 一种实时动漫头像驱动方法及系统 |
Also Published As
| Publication number | Publication date |
|---|---|
| US9361723B2 (en) | 2016-06-07 |
| CN103093490B (zh) | 2015-08-26 |
| US20150035825A1 (en) | 2015-02-05 |
| CN103093490A (zh) | 2013-05-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2014117446A1 (zh) | 基于单个视频摄像机的实时人脸动画方法 | |
| Park et al. | Nerfies: Deformable neural radiance fields | |
| US10679046B1 (en) | Machine learning systems and methods of estimating body shape from images | |
| Cao et al. | 3D shape regression for real-time facial animation | |
| Hassner | Viewing real-world faces in 3D | |
| JP2023521952A (ja) | 3次元人体姿勢推定方法及びその装置、コンピュータデバイス、並びにコンピュータプログラム | |
| CN109359514B (zh) | 一种面向deskVR的手势跟踪识别联合策略方法 | |
| WO2021063271A1 (zh) | 人体模型重建方法、重建系统及存储介质 | |
| Kao et al. | Toward 3d face reconstruction in perspective projection: Estimating 6dof face pose from monocular image | |
| CN111127304A (zh) | 跨域图像转换 | |
| CN118314280A (zh) | 一种基于高斯表达的大尺度三维场景实时重建方法 | |
| CN108292362A (zh) | 用于光标控制的手势识别 | |
| CN104317391A (zh) | 一种基于立体视觉的三维手掌姿态识别交互方法和系统 | |
| Ye et al. | Free-viewpoint video of human actors using multiple handheld kinects | |
| CN113688907A (zh) | 模型训练、视频处理方法,装置,设备以及存储介质 | |
| CN116391208B (zh) | 使用场景流估计的非刚性3d物体建模 | |
| CN111862278A (zh) | 一种动画获得方法、装置、电子设备及存储介质 | |
| CN116452715A (zh) | 动态人手渲染方法、装置及存储介质 | |
| CN120707630A (zh) | 一种单目相机下针对任意物体的姿态识别算法及应用系统 | |
| CN117218246A (zh) | 图像生成模型的训练方法、装置、电子设备及存储介质 | |
| Darujati et al. | Facial motion capture with 3D active appearance models | |
| CN111582120A (zh) | 用于捕捉眼球活动特征的方法、终端设备 | |
| Chen et al. | A particle filtering framework for joint video tracking and pose estimation | |
| Anasosalu et al. | Compact and accurate 3-D face modeling using an RGB-D camera: let's open the door to 3-D video conference | |
| CN114581494A (zh) | 基于神经非刚性注册的人脸光流估计方法及装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 13873306 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 13873306 Country of ref document: EP Kind code of ref document: A1 |






