WO2022110855A1 - 人脸重建方法、装置、计算机设备及存储介质 - Google Patents

人脸重建方法、装置、计算机设备及存储介质 Download PDF

Info

Publication number
WO2022110855A1
WO2022110855A1 PCT/CN2021/108629 CN2021108629W WO2022110855A1 WO 2022110855 A1 WO2022110855 A1 WO 2022110855A1 CN 2021108629 W CN2021108629 W CN 2021108629W WO 2022110855 A1 WO2022110855 A1 WO 2022110855A1
Authority
WO
WIPO (PCT)
Prior art keywords
face
point cloud
cloud data
dense point
sample
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/108629
Other languages
English (en)
French (fr)
Inventor
林纯泽
陈祖凯
王权
徐胜伟
钱晨
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Sensetime Technology Development Co Ltd
Original Assignee
Beijing Sensetime Technology Development Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Sensetime Technology Development Co Ltd filed Critical Beijing Sensetime Technology Development Co Ltd
Priority to KR1020237018454A priority Critical patent/KR20230098313A/ko
Priority to JP2023531694A priority patent/JP7525814B2/ja
Publication of WO2022110855A1 publication Critical patent/WO2022110855A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T17/00Three-dimensional [3D] modelling for computer graphics
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/168Feature extraction; Face representation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10028Range image; Depth image; 3D point clouds
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person
    • G06T2207/30201Face

Definitions

  • the present disclosure relates to the technical field of image processing, and in particular, to a face reconstruction method, apparatus, computer equipment, and storage medium.
  • a three-dimensional model of a virtual face can be established according to a real face or one's own preferences, so as to realize the reconstruction of the face, which has a wide range of applications in the fields of games, animation, and virtual social interaction.
  • the player can generate a 3D virtual face model according to the real face included in the image provided by the player through the face reconstruction system provided by the game program, and use the generated 3D virtual face model to participate in a more immersive participation. game.
  • the similarity between the virtual face three-dimensional model obtained based on the face reconstruction method and the real face is low.
  • the embodiments of the present disclosure provide at least a face reconstruction method, apparatus, computer equipment, and storage medium.
  • an embodiment of the present disclosure provides a face reconstruction method, including: acquiring dense point cloud data of an original face included in a target image; using the dense points of the first reference face corresponding to multiple reference images respectively The cloud data is fitted to the dense point cloud data of the original face, and the fitting coefficients corresponding to multiple sets of the dense point cloud data of the first reference face are obtained; based on multiple sets of second reference faces with preset styles The dense point cloud data and the corresponding fitting coefficients of the dense point cloud data of multiple groups of the first reference face, determine the dense point cloud data of the target face model; the second reference faces of the multiple groups are respectively Generated based on the first reference face in the multiple reference images; based on the dense point cloud data of the target face model, a target face model corresponding to the original face of the target image is generated.
  • the fitting coefficient is used as a medium to establish an association relationship between the dense point cloud data of the original face and the dense point cloud data of a plurality of first reference faces.
  • the features of the original face in the target image (such as shape features, etc.) have higher similarity with the original face, and can make the generated target face model have a preset style.
  • coefficient, determining the dense point cloud data of the target face model including: based on the dense point cloud data of multiple groups of the second reference face, generating the mean data of the dense point cloud data of multiple groups of the second reference face;
  • the target face model is generated based on a plurality of sets of dense point cloud data of the second reference face, the mean data, and a plurality of sets of fitting coefficients corresponding to the dense point cloud data of the first reference face respectively dense point cloud data.
  • the dense point cloud data based on multiple groups of the second reference face, the mean data, and the dense point cloud data based on multiple groups of the first reference face are respectively corresponding.
  • Fitting coefficient, generating the dense point cloud data of the target face model including: based on the dense point cloud data of each group of the second reference face in the multiple groups of dense point cloud data of the second reference face, and the mean value data, determine the difference data of the dense point cloud data of each group of the second reference face; Interpolation processing is performed on the difference data corresponding to the dense point cloud data of the second reference face respectively; based on the result of the interpolation processing and the mean data, the dense point cloud data of the target face model is generated.
  • the average feature of the dense point cloud data of multiple groups of second reference faces can be accurately characterized by the mean data; the difference data of the dense point cloud data of each group of the second reference faces can be accurately characterized The degree of difference between the dense point cloud data of each group of second reference faces and the average feature of the dense point cloud data of multiple groups of second reference faces, so that the more accurate difference degree data is used to adjust the average value data, which is more accurate.
  • the obtaining dense point cloud data of the original face included in the target image includes: obtaining the target image including the original face; The target image is processed to obtain dense point cloud data of the original face in the target image.
  • the facial features of the original face in the target image can be more accurately represented by using the dense point cloud data of the original face.
  • the method obtains the dense point cloud data of the first reference face corresponding to the multiple reference images in the following manner: obtaining multiple reference images including the first reference face; For each of the multiple reference images, use a pre-trained neural network to process each of the reference images to obtain dense point cloud data of the first reference face in each reference image .
  • the face features corresponding to the first reference face in the reference image can be more accurately represented by using the dense point cloud data of the first reference face.
  • using a plurality of reference images including the first reference face can cover as wide a face shape feature as possible.
  • the dense point cloud data of the first reference face corresponding to the multiple reference images is used to fit the dense point cloud data of the original face to obtain multiple sets of the first reference.
  • the fitting coefficients corresponding to the dense point cloud data of the face respectively include: performing least squares processing on the dense point cloud data of the original face and the dense point cloud data of the first reference face to obtain multiple sets of The intermediate coefficients corresponding to the dense point cloud data of the first reference face respectively; based on the intermediate coefficients corresponding to the dense point cloud data of each group of the first reference face, determine the dense point cloud data of each group of the first reference face the corresponding fitting coefficients.
  • the fitting coefficient can be used to accurately characterize the fitting situation when the dense point cloud data of the first reference face is used to fit the dense point cloud data of the original face.
  • the fitting coefficient corresponding to the dense point cloud data of each group of the first reference face is determined based on the intermediate coefficient corresponding to the dense point cloud data of each group of the first reference face
  • the method includes: determining, from the dense point cloud data of each group of the first reference face, the first type of dense point cloud data representing the part of the first reference face corresponding to the target face model; The intermediate coefficients corresponding to the first type of dense point cloud data in the dense point cloud data of the reference face are adjusted to obtain the first fitting coefficient; the second type of dense point cloud data in the first reference face dense point cloud data is adjusted The intermediate coefficient corresponding to the point cloud data is determined as the second fitting coefficient; the second type of dense point cloud data is the dense point cloud data of the first reference face except for the first type of dense point cloud data based on the first fitting coefficient and the second fitting coefficient, obtaining the fitting coefficient of the dense point cloud data of each group of the first reference face.
  • the method obtains the dense point cloud data of the second reference face with the preset style in the following manner: comparing the dense point cloud data of the first reference face in the reference image Adjust to obtain the dense point cloud data of the second reference face with the preset style; or, based on the first reference face in the reference image, generate a second reference face with the preset style.
  • the dense point cloud data obtained by fitting the fitting coefficients and the dense point cloud data of a plurality of corresponding first reference faces can be made
  • the dense point cloud data of the original face corresponding to the target image is similar, that is, the obtained fitting coefficient can more accurately represent the coefficient of the first reference face dense point cloud fitting the original face dense point cloud.
  • training the neural network includes: acquiring a sample image set; the sample image set includes a plurality of first sample images including a first sample face; the plurality of first sample images This image is divided into a plurality of first sample image subsets, and each first sample image subset includes images of the first sample faces with the same expression respectively collected from a plurality of preset collection angles; Dense point cloud data of the first sample face of the first sample image in the sample image set; use the neural network to perform feature learning on the first sample image in the sample image set, and obtain the first sample image in the sample image set.
  • the predicted dense point cloud data of the first sample face of a sample image; the neural network is trained by using the dense point cloud data of the first sample face and the predicted dense point cloud data.
  • the sample image set further includes a plurality of second sample images including second sample faces and backgrounds
  • the training of the neural network further includes: acquiring each of the second samples face key point data of the image; using the face key point data of the second sample image and the second sample image, fitting to generate the dense point cloud data of the second sample face of the second sample image; using the The neural network performs feature learning on the second sample image in the sample image set to obtain the predicted dense point cloud data of the second sample face of the second sample image; using the dense point cloud data of the second sample face
  • the neural network is trained on point cloud data and predicted dense point cloud data.
  • the sample image set further includes a third sample image; the third sample image is obtained by performing data enhancement processing on the first sample image;
  • the training of the neural network further includes:
  • the neural network is trained using the dense point cloud data of the third sample face and the predicted dense point cloud data.
  • the data enhancement processing includes at least one of the following: random occlusion processing, Gaussian noise processing, motion blur processing, and color region channel change processing.
  • neural networks with different advantages can be obtained by adjusting the number of the first sample image, the second sample image, and the third sample image, so as to obtain a better neural network according to actual needs; at the same time, because The third sample image is obtained through data enhancement processing, so when the third sample image is included in the sample image, the neural network obtained by training has a stronger ability to process data.
  • the obtained neural network can have better generalization ability.
  • an embodiment of the present disclosure further provides a face reconstruction device, including:
  • the first acquisition module is used to acquire the dense point cloud data of the original face included in the target image
  • the first processing module is used to fit the dense point cloud data of the original face by using the dense point cloud data of the first reference face corresponding to the multiple reference images respectively, and obtain multiple groups of dense point cloud data of the first reference face.
  • a determination module configured to determine the target face based on the respective fitting coefficients corresponding to the dense point cloud data of multiple groups of second reference faces with preset styles and the respective corresponding fitting coefficients of the dense point cloud data of multiple groups of the first reference faces Dense point cloud data of the model; the multiple sets of second reference faces are respectively generated based on the first reference faces in the multiple reference images;
  • a generating module configured to generate a target face model corresponding to the original face of the target image based on the dense point cloud data of the target face model.
  • the determining module corresponds to the dense point cloud data of multiple groups of second reference faces with preset styles and the dense point cloud data of multiple groups of the first reference faces, respectively.
  • determining the dense point cloud data of the target face model it is used to: generate multiple groups of dense point clouds of the second reference face based on the dense point cloud data of the second reference face in multiple groups.
  • the mean value data of the data based on the corresponding fitting coefficients of the dense point cloud data of the second reference face, the mean value data, and the dense point cloud data of the first reference face of the multiple groups, generating the Describe the dense point cloud data of the target face model.
  • the determining module is based on multiple sets of dense point cloud data of the second reference face, the mean data, and multiple sets of dense point cloud data of the first reference face.
  • the corresponding fitting coefficients, when generating the dense point cloud data of the target face model, are used for: based on the dense point cloud data of each group of the second reference faces in the multiple groups of the second reference faces.
  • Dense point cloud data and the mean data determine the difference data of the dense point cloud data of each group of the second reference face; combination coefficient, and perform interpolation processing on the difference data corresponding to the dense point cloud data of multiple groups of the second reference face respectively; point cloud data.
  • the first obtaining module when acquiring the dense point cloud data of the original face included in the target image, is used to: obtain the target image including the original face;
  • the trained neural network processes the target image to obtain dense point cloud data of the original face in the target image.
  • the apparatus further includes a second processing module, configured to obtain the dense point cloud data of the first reference face corresponding to the multiple reference images in the following manner: obtaining the dense point cloud data including the first reference Multiple reference images of human faces; for each of the multiple reference images, use a pre-trained neural network to process each of the reference images to obtain the first reference image in each reference image A dense point cloud data of a reference face.
  • a second processing module configured to obtain the dense point cloud data of the first reference face corresponding to the multiple reference images in the following manner: obtaining the dense point cloud data including the first reference Multiple reference images of human faces; for each of the multiple reference images, use a pre-trained neural network to process each of the reference images to obtain the first reference image in each reference image A dense point cloud data of a reference face.
  • the first processing module uses the dense point cloud data of the first reference face corresponding to the multiple reference images to fit the dense point cloud data of the original face, and obtains multiple sets of
  • the fitting coefficients corresponding to the dense point cloud data of the first reference face are respectively used, it is used to: perform a least-two method on the dense point cloud data of the original face and the dense point cloud data of the first reference face. Multiply processing to obtain intermediate coefficients corresponding to the dense point cloud data of multiple groups of the first reference faces respectively; based on the intermediate coefficients corresponding to the dense point cloud data of each group of the first reference faces, determine the first reference of each group The fitting coefficients corresponding to the dense point cloud data of the face.
  • the first processing module determines, based on the intermediate coefficients corresponding to the dense point cloud data of each group of first reference faces, the corresponding dense point cloud data of each group of the first reference faces.
  • the fitting coefficient is , it is used to: determine the first type of dense point cloud representing the part of the first reference face corresponding to the target face model from the dense point cloud data of each group of the first reference face data; adjusting the intermediate coefficients corresponding to the first type of dense point cloud data in the dense point cloud data of the first reference face to obtain a first fitting coefficient;
  • the intermediate coefficient corresponding to the second type of dense point cloud data in the cloud data is determined as the second fitting coefficient;
  • the second type of dense point cloud data is the dense point cloud data of the first reference face divided by the Dense point cloud data other than one type of dense point cloud data; based on the first fitting coefficient and the second fitting coefficient, the fitting coefficient of each group of the dense point cloud data of the first reference face is obtained.
  • the apparatus further includes an adjustment module, configured to obtain the dense point cloud data of the second reference face with the preset style in the following manner: The dense point cloud data of the face is adjusted to obtain the dense point cloud data of the second reference face with the preset style; Set a virtual face image of the second reference face of the style; use a pre-trained neural network to generate dense point cloud data of the second reference face in the virtual face image.
  • an adjustment module configured to obtain the dense point cloud data of the second reference face with the preset style in the following manner: The dense point cloud data of the face is adjusted to obtain the dense point cloud data of the second reference face with the preset style; Set a virtual face image of the second reference face of the style; use a pre-trained neural network to generate dense point cloud data of the second reference face in the virtual face image.
  • the apparatus further includes a training module, which, when training the neural network, is used to: obtain a sample image set; the sample image set includes a plurality of first sample faces including a first sample face. sample images; the plurality of first sample images are divided into a plurality of first sample image subsets, and each first sample image subset includes images with the same expression collected from a plurality of preset collection angles respectively
  • the image of the first sample face in the sample image set obtain the dense point cloud data of the first sample face of the first sample image in the sample image set;
  • the image is subjected to feature learning, and the predicted dense point cloud data of the first sample face of the first sample image is obtained; using the dense point cloud data and predicted dense point cloud data of the first sample face, the neural network is analyzed.
  • the network is trained.
  • the sample image set further includes a plurality of second sample images including second sample faces and backgrounds
  • the training module is further configured to: acquire the data of each second sample image. face key point data; using the face key point data of the second sample image and the second sample image, fitting to generate the dense point cloud data of the second sample face of the second sample image; using the neural network
  • the network performs feature learning on the second sample image in the sample image set to obtain the predicted dense point cloud data of the second sample face of the second sample image; using the dense point cloud of the second sample face data and predicted dense point cloud data to train the neural network.
  • the sample image set further includes: a third sample image; the third sample image is obtained by performing data enhancement processing on the first sample image;
  • the training module is further configured to: obtain the dense point cloud data of the third sample face of the third sample image; use the neural network to perform feature learning on the third sample image to obtain the third sample image The predicted dense point cloud data of the third sample face;
  • the neural network is trained using the dense point cloud data of the third sample face and the predicted dense point cloud data.
  • the data enhancement processing includes at least one of the following: random occlusion processing, Gaussian noise processing, motion blur processing, and color region channel change processing.
  • an optional implementation manner of the present disclosure further provides a computer device, including a processor and a memory, where the processor is configured to execute machine-readable instructions stored in the memory, and the machine-readable instructions are processed by the memory When executed by the processor, when the machine-readable instructions are executed by the processor, the above-mentioned first aspect or the steps in any possible implementation manner of the first aspect are performed.
  • an optional implementation manner of the present disclosure further provides a computer-readable storage medium, on which a computer program is run to execute the steps in the first aspect or any possible implementation manner of the first aspect .
  • FIG. 1 shows a flowchart of a face reconstruction method provided by an embodiment of the present disclosure
  • FIG. 2 shows a flowchart of a specific method for training a neural network provided by an embodiment of the present disclosure
  • FIG. 3 shows a schematic diagram of a first sample image provided by an embodiment of the present disclosure, and a third sample image determined by using the first sample image;
  • FIG. 4 shows a flowchart of a specific method for obtaining dense point cloud data of a second sample face of a second sample image provided by an embodiment of the present disclosure
  • FIG. 5 shows a specific example diagram of a neural network structure provided by an embodiment of the present disclosure
  • FIG. 6 shows a flowchart of a method for determining fitting coefficients corresponding to multiple reference images respectively
  • FIG. 7 shows a flowchart of a specific method for determining a fitting coefficient corresponding to dense point cloud data of each group of first reference faces provided by an embodiment of the present disclosure
  • FIG. 8 shows a flowchart of a specific method for determining dense point cloud data of a target face model provided by an embodiment of the present disclosure
  • FIG. 9 shows a flowchart of a specific method for generating dense point cloud data of a target face model provided by an embodiment of the present disclosure
  • FIG. 10 shows a schematic diagram of a face reconstruction apparatus provided by an embodiment of the present disclosure
  • FIG. 11 shows a schematic diagram of a computer device provided by an embodiment of the present disclosure.
  • the dense point cloud of the face corresponding to the face is usually obtained based on the face image, and then the virtual face 3D model is obtained based on the face image.
  • the specific style of face dense point cloud is adjusted multiple times to generate virtual images. Since the face dense point clouds corresponding to different faces are different, even if the style of the virtual face to be reconstructed is the same, the adjustment when reconstructing the face dense point cloud corresponding to different faces according to the determined style will not and the adjustment has great uncertainty, which makes it difficult to control the specific direction of the adjustment during the adjustment process, resulting in a large difference between the generated 3D model of the virtual face and the real face. This leads to the problem of low similarity between the 3D model of the virtual face and the real face.
  • the present disclosure provides a face reconstruction method, device, computer equipment, and storage medium.
  • the dense point cloud data of the original face and the dense point cloud data of multiple first reference faces are established.
  • the relationship between the point cloud data, the relationship can represent the dense point cloud data of the second reference face determined based on the dense point cloud data of the first reference face, and the target face model established based on the original face.
  • the association between dense point cloud data makes the generated target face model have the characteristics of the original face in the target image (such as shape features, etc.), and has a higher similarity with the original face, and can make the generated face model.
  • the target face model has a preset style.
  • this solution only needs to generate the dense point cloud data of the second reference face corresponding to multiple reference images, and use the dense point cloud data of the second reference face with the preset style to
  • the target face model generated by the original face does not need to determine the adjustment scheme for different original faces, but uses the dense point cloud data of the same second reference face and the fitting coefficients of different original faces to determine different
  • the target face model of the original face has higher processing efficiency.
  • the execution subject of the face reconstruction method provided by the embodiment of the present disclosure is generally a computer device with a certain computing capability.
  • the computer equipment includes, for example, a terminal device or a server or other processing device, and the terminal device can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), handheld devices, computing devices, in-vehicle devices, wearable devices, etc.
  • the face reconstruction method may be implemented by the processor calling computer-readable instructions stored in the memory.
  • an embodiment of the present disclosure provides a face reconstruction method, the method includes steps S101 to S104, wherein:
  • S103 Determine the dense point cloud of the target face model based on the respective fitting coefficients corresponding to the dense point cloud data of the multiple groups of the second reference face with the preset style and the dense point cloud data of the multiple groups of the first reference face data; multiple sets of second reference faces are respectively generated based on the first reference faces in the multiple reference images;
  • S104 Generate a target face model corresponding to the original face of the target image based on the dense point cloud data of the target face model.
  • This process uses the fitting coefficient as a medium to establish the correlation between the dense point cloud data of the original face and the dense point cloud data of multiple first reference faces.
  • the association between the dense point cloud of the second reference face determined by the dense point cloud of the first reference face and the dense point cloud data of the target face model established based on the dense point cloud data of the original face, so that the generated
  • the dense point cloud data of the target face model has the characteristics of the original face in the target image (such as shape features, etc.), and has a higher similarity with the original face, and can make the generated target face model have presets. style of.
  • the target image is, for example, a pre-acquired image including a human face, for example, an image including a human face acquired when a certain object is photographed with a photographing device such as a camera.
  • a photographing device such as a camera.
  • any face included in the image can be determined as the original face, and the original face can be used as the object of face reconstruction.
  • the acquisition method of the target image is also different.
  • an image including the face of the game player can be obtained through an image obtaining device installed in the game device, or an image including the face of the game player can be selected from an album in the game device.
  • the image of the face of the game player, and the acquired image including the face of the game player is used as the target image.
  • an image including the user's face may be collected by the camera of the terminal device, or an image including the user's face may be selected from an album of the terminal device, or Receive images including the user's face from other applications installed in the terminal device.
  • a video frame image containing a human face can be obtained from multiple frames of video frame images included in a video stream obtained by a live broadcast device; the video frame image containing a human face can be obtained.
  • the target image may have multiple frames; for example, the multiple-frame target image may be obtained by sampling multiple frames of video frame images in the video stream.
  • the following methods can be used: obtaining the target image including the original face; using a pre-trained neural network to process the target image to obtain the original face in the target image Dense point cloud data of human faces.
  • the pre-trained neural network when using a pre-trained neural network to process the target image to obtain dense point cloud data of the original face, includes at least one of the following: Convolutional Neural Networks (CNN) , Back Propagation Neural Network (Back Propagation, BP), Backbone Neural Network (Backbone).
  • CNN Convolutional Neural Networks
  • BP Back Propagation Neural Network
  • Backbone Backbone Neural Network
  • the backbone network of the neural network may be determined first, which is used as the main structure of the neural network.
  • the backbone network may include at least one of the following: an initial network (Inception) , Residual network variant network (the next dimension to RESNET, ResNeXt), starting network variant network (Xception), squeeze and excitation network (Squeeze-and-Excitation Networks, SENet), lightweight network (MobileNet), And a lightweight network (ShuffleNet).
  • a lightweight network can also be selected as the basic model of the convolutional neural network.
  • the lightweight network On the basis of the lightweight network, other network structures are added to form a convolutional neural network, and the A convolutional neural network is constructed for training.
  • the training speed is also faster; in addition, the trained neural network also has the advantages of small size, data processing The advantage of fast speed is more suitable for deployment in embedded devices.
  • the network structure of the above-mentioned neural network is only an example; the specific construction method and structure of the network structure can be determined according to the actual situation, which will not be repeated here, and the above examples do not limit the embodiments of the present disclosure.
  • an embodiment of the present disclosure provides a specific method for training a neural network, including:
  • S201 Obtain a sample image set; the sample image set includes a plurality of first sample images including a first sample face; the plurality of first sample images are divided into a plurality of first sample image subsets, each of which is the same as the first sample image.
  • This subset of images includes images of the faces of the first sample with the same expression, respectively collected from multiple preset collection angles.
  • the corresponding first sample face is, for example, a predetermined image of at least one individual object used for acquiring a face image to train the neural network. human face.
  • first sample image subsets may be determined for multiple different expressions.
  • the multiple different expressions are, for example, happiness, excitement, loss, sadness, and the like.
  • the expression of "happy” is presented from different angles If the first sample face is photographed, a plurality of first sample images corresponding to the "happy" expression can be obtained as a subset of the first sample images.
  • the backgrounds of the first sample faces in different first sample images may be the same.
  • the multiple first sample images of the first sample face can be determined.
  • an image acquisition device may be used to capture and acquire an image, wherein the image acquisition device includes, for example, at least one of a depth camera and a color camera.
  • the faces of I individual objects may be photographed under E expressions to obtain a plurality of first sample images.
  • the face of a certain individual object A may be determined as the first sample face.
  • P is an integer greater than 1
  • the face of the first sample is photographed from P (P is an integer greater than 1) different angles, and the corresponding P different angles under the expression of "sad” are obtained.
  • P is an integer greater than 1
  • the first sample images can be used to train the neural network to detect the multi-angle images in the images.
  • the second sample image can be randomly photographed for different individual objects, or randomly crawled from a preset network platform containing a plurality of images containing human faces, and the crawled images can be used as the second sample face image.
  • a camera or other image photographing device can be used to photograph multiple second sample faces to obtain the second sample image; A plurality of second sample images obtained by shooting.
  • the second sample image includes, for example, H acquired face images including backgrounds; wherein, the backgrounds included in the second sample images interfere with the recognition of the face, and are used to train the neural network to perform the analysis on the face in the target image.
  • the sample image set further includes a third sample image
  • the third sample image may be obtained, for example, by performing data enhancement processing on the first sample image.
  • the data enhancement processing includes at least one of the following: random occlusion processing, Gaussian noise processing, blurring processing, and color region channel change processing.
  • a third sample image may be obtained by performing occlusion processing on a partial area in the first sample image; wherein, the size of the occluded part may be based on the size of the first sample image and the actual size of the image.
  • the first sample image can be processed by using global color region channel change process or random color region channel change process to obtain a third sample image with global or partial region color region channel change.
  • FIG. 3 is a schematic diagram of a first sample image provided by an embodiment of the present disclosure and a third sample image determined by using the first sample image; as shown in FIG. 3 , 31 represents the first sample image; 32 represents the first sample image.
  • the sample image 31 is a third sample image obtained by performing data enhancement processing of blurring processing; 33 represents a third sample image obtained by performing data enhancement processing on the first sample image 31 with partial occlusion, wherein 34 represents the position of partial occlusion.
  • the dense point cloud data of the sample face corresponding to the third sample image is the same as the dense point cloud data of the sample face corresponding to the first sample image from which the third sample image was generated.
  • the specific method for training the neural network also includes:
  • S202 Acquire dense point cloud data of a first sample face of a first sample image in the sample image set.
  • the following method can be used : Obtain the dense point cloud data of the first sample face of the first sample image, the dense point cloud data of the second sample face of the second sample image, and the third sample face of the third sample image in the sample image set dense point cloud data.
  • the first sample face in each first sample image can be obtained.
  • dense point cloud data In the case of using a color camera to obtain the first sample image, for example, a model such as 3D Morphable Model (3DMM) can be used to obtain the dense point cloud data of the first sample face; when using the depth camera to obtain the first sample image
  • 3DMM 3D Morphable Model
  • the dense point cloud data of the face of the first sample can be obtained based on the depth image obtained by the depth camera.
  • an embodiment of the present disclosure provides a specific method for obtaining dense point cloud data of a second sample face of the second sample image, including:
  • S404 Use the dense point cloud data of the second sample face and the predicted dense point cloud data to train the neural network.
  • the second sample images include, for example, H acquired face images including backgrounds, and each second sample image includes determined face key points.
  • the face key points included in the second sample image are used to determine the dense point cloud data of the sample face corresponding to the second sample image, for example, the key points directly marked in the second sample image and used to characterize the face features are included Points, such as multiple key points that characterize facial features, cheekbones, and brow bones; or, include key points corresponding to human faces determined by a key point detection method.
  • the key point detection method includes at least one of the following: Active Shape Model (ASM), Active Appearance Models (AAM), and Cascaded Pose Regression (CPR).
  • ASM Active Shape Model
  • AAM Active Appearance Models
  • CPR Cascaded Pose Regression
  • the fitting model can be used to generate dense point cloud data of the second sample face of the second sample image.
  • the fitting model includes, for example, a 3D deformation statistical model of a human face.
  • the method for training the neural network further includes:
  • S203 Use a neural network to perform feature learning on the first sample image in the sample image set to obtain predicted dense point cloud data of the first sample face of the first sample image.
  • the following method can be used: using the initial The neural network performs feature learning on the first sample image, the second sample image, and the third sample image in the sample image set to obtain the predicted dense point cloud data of the first sample face of the first sample image, and the second sample image.
  • the predicted dense point cloud data of the second sample face of the sample image, and the predicted dense point cloud data of the third sample face of the third sample image can be used: using the initial The neural network performs feature learning on the first sample image, the second sample image, and the third sample image in the sample image set to obtain the predicted dense point cloud data of the first sample face of the first sample image, and the second sample image.
  • the step of acquiring at least one of the first sample image, the second sample image, and the third sample image may be the same as using the initial neural network for the first sample image and the second sample image.
  • the steps of performing feature learning in at least one of the third sample images are performed synchronously, that is, it is possible to directly obtain the characteristics of at least one of the first sample images, the second sample images, and the third sample images by using the initial neural network.
  • the first sample image, the second sample image, and the third sample image can be synchronously performed. Perform feature learning on at least one of the three sample images; or, perform feature learning on at least one of the first sample image, the second sample image, and the third sample image in sequence according to actual needs to obtain the first sample.
  • This embodiment of the present disclosure does not limit the sequential execution order of the above-mentioned sample processing procedures, which may be specifically set according to actual needs.
  • different numbers of the first sample image and the second sample image can be selected according to a preset ratio, and the selected first sample image and the second sample image can be selected according to a preset ratio.
  • This image and the second sample image are input into the initial neural network; when the sample image set includes the first sample image, the second sample image and the third sample image, different numbers of first sample images can be selected according to a preset ratio sample image, second sample image, and third sample image, and input the selected first sample image, second sample image, and third sample image into the initial neural network.
  • the ratio is selected differently, the emphasis on neural network training is also different.
  • the neural network obtained by training has a strong ability to obtain dense point clouds of faces corresponding to faces of different angles in the image;
  • the trained neural network has stronger anti-interference ability to other background parts other than the face in the image, so as to meet different usage requirements.
  • the initial neural network After inputting the sample image into the initial neural network, the initial neural network can perform feature learning on the sample image, and output the dense point cloud data of the predicted face for each sample image;
  • the dense point cloud data of the sample face corresponding to the sample images determines the loss of the neural network, which is used to measure the accuracy of the neural network in generating the dense point cloud data of the face.
  • the dense point cloud data for predicting the face obtained by using the neural network may include, for example, the coordinate values representing the position of the dense point cloud of the face under a preset coordinate system, or, including the information representing the dense point cloud of the face under the preset coordinate system.
  • the preset coordinate system is, for example, a preset face coordinate system.
  • the dense point cloud data of the predicted face may include coordinate values (x, y, z), or, including the coordinate value x on the x-axis, the coordinate value y on the y-axis, and the coordinate value z on the z-axis, and the dense point cloud data of different predicted faces are included in the x-axis, y
  • the coordinate values on the axis and the z-axis have a preset order in the output dense point cloud data of the predicted face.
  • an embodiment of the present disclosure also provides a specific example diagram of a neural network structure, wherein the neural network structure includes: a backbone network 51 , a first fully connected layer 52 , and three groups of second fully connected layers 53 .
  • the sample image can be input into the backbone network to obtain the characteristic data of the sample image. After connecting the layers, they are respectively input to three groups of second fully connected layers; the three groups of second fully connected layers are used to predict the coordinate values of the dense point cloud of the face in the sample image in the face coordinate system.
  • the first group of second fully-connected layers can output the x-axis coordinates of the dense point cloud in the face coordinate system
  • the second group of second fully-connected layers can output the y-axis coordinates of the dense point cloud in the face coordinate system value
  • the third group of second fully connected layers can output the z-axis coordinate value of the dense point cloud in the face coordinate system.
  • the coordinate values of the dense point cloud in the face coordinate system constitute the dense point cloud data of the predicted face in the sample image.
  • the selection ratio of the first sample images is 40% and the selection ratio of the second sample images is 60%.
  • Multiple sample images on which the neural network is trained For example, in the case of using 100 sample images to train the neural network, 40 first sample images and 60 second sample images are selected as multiple sample images.
  • the convolutional neural network is selected as the initial neural network for training, the light-weight network t is used as the basic model, and the split network (split FC) is used as the output layer to construct the convolutional neural network, and the predicted people corresponding to multiple sample images can be obtained.
  • the dense point cloud data of the face, and the obtained dense point cloud data of the predicted face includes multiple coordinate values of points in the dense point cloud on the x-axis, the y-axis, and the z-axis, respectively.
  • the output dense point cloud data of the predicted face can be, for example, in the form of Output in matrix form, expressed as [x 1 ,x 2 ,...,x R ], [y 1 ,y 2 ,...,y R ], and [z 1 ,z 2 ,...,z R ], which is R
  • the coordinate values (x i , y i , z i ) contained in the dense point cloud data of the predicted face corresponding to any one of the different face positions, i ⁇ [1, R] are split into x-axis,
  • the coordinate values x i , yi , and z i on the y-axis and z-axis are output as the ith element in three different matrices, respectively.
  • the method for training the neural network further includes:
  • S204 Use the dense point cloud data of the face of the first sample and the predicted dense point cloud data to train the neural network.
  • the following method can be adopted: using the dense point cloud data of the first sample face and the predicted dense point cloud data, the dense point cloud data of the second sample face.
  • the point cloud data and predicted dense point cloud data, as well as the dense point cloud data and predicted dense point cloud data of the third sample face, are used to train the neural network, and the trained neural network is obtained after the training is completed.
  • the loss of the neural network can be determined based on the difference between the dense point cloud data of the predicted face and the dense point cloud data of the sample face, and the loss is used to train the neural network, and the training direction is to reduce the loss. direction, so that when the neural network processes the image, the dense point cloud data of the predicted face obtained is close enough to the dense point cloud data of the real face.
  • the target image can be input into the neural network to obtain the dense point cloud data corresponding to the original face in the target image.
  • the reference image may be the faces corresponding to different individual objects, and the faces corresponding to different individual objects are different; exemplarily, a plurality of people with different at least one of gender, age, skin color, degree of fatness and thinness, etc.
  • a face image of each person is obtained, and the obtained face image is used as a second sample image.
  • the dense point cloud data of the second sample face generated based on the second sample image can cover as wide a face shape feature as possible.
  • the following methods may be adopted: acquiring multiple reference images including the first reference faces; A pre-trained neural network is used to obtain the dense point cloud data of the first reference face in each reference image.
  • the method of using the pre-trained neural network to obtain the multiple first reference face dense point clouds is similar to the above-mentioned method of using the pre-trained neural network to obtain the original face dense point cloud, and will not be repeated here.
  • the dense point cloud data of the first reference face can be used to fit the dense point cloud data of the original face to obtain The fitting coefficients of the dense point cloud data of the first reference face corresponding to the multiple reference images are obtained.
  • the fitting coefficient can be used as a medium to establish an association relationship between the dense point cloud data of the original face in the target image and the dense point cloud data of the first reference face corresponding to the multiple reference pictures respectively.
  • an embodiment of the present disclosure provides a method for determining fitting coefficients corresponding to multiple reference images, including the following steps S601 to S602.
  • S601 Perform least squares processing on the dense point cloud data of the original face and the dense point cloud data of the first reference face to obtain intermediate coefficients corresponding to the dense point cloud data of multiple groups of the first reference face respectively.
  • the dense point cloud data of the original face is represented as IN mesh
  • the dense point cloud data of the first reference face is represented as BASE mesh .
  • the dense point cloud data BASE mesh of the first reference face correspondingly contains N groups of Dense point cloud data of human faces, denoted as in, Dense point cloud data representing the face corresponding to the ith reference image.
  • N fitting values can be obtained, which are expressed as ⁇ i (i ⁇ [1,N]).
  • ⁇ i represents the fitting value corresponding to the dense point cloud data of the i-th first reference face.
  • the fitting coefficient can also be regarded as the expression of the dense point cloud data of each first reference face when the dense point cloud data of the first reference face corresponding to the multiple reference images is used to express the dense point cloud data of the original face. coefficient.
  • S602 Determine the fitting corresponding to the dense point cloud data of each group of first reference faces based on the intermediate coefficients corresponding to the dense point cloud data of each group of first reference faces in the multiple groups of dense point cloud data of the first reference face coefficient.
  • an embodiment of the present disclosure further provides a specific method for determining a fitting coefficient corresponding to the dense point cloud data of each group of first reference faces, including the following steps S701 to S704.
  • S701 Determine, from the dense point cloud data of each group of the first reference face, the first type of dense point cloud data representing the part of the first reference face corresponding to the part of the target face model.
  • any set of dense point cloud data for multiple sets of first reference face dense point cloud data because is based on the ith reference image, so contains dense point cloud data representing the face position in the ith reference image.
  • the parts of the human face include, for example, the eyebrows, the nose, the eyes, the mouth, the cheekbones, and the lower jaw.
  • the eyebrows may be further divided into the brow tip, the brow center, and the brow peak.
  • the corresponding parts of the face may be The fitting coefficient is adjusted so that the dense point cloud data obtained by fitting the dense point cloud data of the corresponding first reference face based on the fitting coefficient and the dense point cloud data of the original face corresponding to the target image similar.
  • the partial face parts corresponding to the fitting coefficients that need to be adjusted are the target face model parts, for example, the eyes and the mouth may be included.
  • the specific target face model part can be determined according to specific conditions or experience, and will not be repeated here.
  • a set of dense point clouds of the first reference face data For example, the face parts can be divided into R, for example, the dense point cloud data corresponding to the R face parts can be expressed as
  • the dense point cloud data corresponding to the target face model parts include, for example, and That is, the first type of dense point cloud data.
  • the first fitting coefficient corresponding to the part corresponding to the target face model can be determined, that is, some intermediate coefficients that need to be adjusted to achieve a better fitting effect.
  • ⁇ i is obtained by using IN mesh and It is obtained by performing least squares processing, so ⁇ i also includes multiple intermediate coefficients of partial parts corresponding to multiple target face models respectively.
  • the intermediate coefficients in ⁇ i corresponding to the multiple target face model parts can be expressed as ⁇ i-1 and ⁇ i-2 , for example, That is, the first type of dense point cloud data and the corresponding intermediate coefficients.
  • Numerical adjustment of the intermediate coefficients ⁇ i-1 and ⁇ i-2 can make the first fitting coefficient obtained after the adjustment, after fitting the dense point cloud data of the original face, the fitting result is the same as that of the original face.
  • the dense point cloud data is more similar.
  • the numerical adjustment includes, for example, an increase in the numerical value and/or a decrease in the numerical value.
  • S703 Determine the intermediate coefficient corresponding to the second type of dense point cloud data in the dense point cloud data of the first reference face as the second fitting coefficient; the second type of dense point cloud data is the dense point cloud of the first reference face Dense point cloud data other than the first type of dense point cloud data in the data.
  • the intermediate coefficient corresponding to the second type of dense point cloud data in the dense point cloud data of the first reference face may be determined as the second fitting coefficient.
  • the dense point cloud data other than the first type of dense point cloud data in the dense point cloud data of each group of the first reference face can also be used as the second type of dense point cloud data. Since the fitting coefficients corresponding to the second type of dense point cloud data have little influence on the fitting results, or the fitting results are better during fitting, the fitting coefficients corresponding to the second type of dense point cloud data may not be calculated. Adjustments are made to improve efficiency while ensuring the fitting effect.
  • the first fitting coefficient and the second fitting coefficient are By combining the fitting coefficients, fitting coefficients corresponding to multiple face parts can be determined, that is, fitting coefficients of the dense point cloud data of each group of first reference faces.
  • the preset style can be, for example, a cartoon style, an ancient style or an abstract style, etc., which can be set according to actual needs.
  • the second reference face with the preset style may be a cartoon face.
  • the following method can be used: adjusting the dense point cloud data of the first reference face in the reference image to obtain Dense point cloud data of a second reference face with a preset style; or, based on the first reference face in the reference image, generate a virtual face image including a second reference face with a preset style, and use the pre-
  • the trained neural network generates dense point cloud data of the second reference face in the virtual face image.
  • the dense point cloud of the first reference face is adjusted to obtain the dense point cloud data of the second reference face with a preset style
  • the dense point cloud of the first reference face can be adjusted according to the preset style. All or part of the dense point cloud data in the point cloud data is adjusted so that the face reflected by the dense point cloud data of the second reference face has a preset style.
  • the cartoon style includes, for example, zooming in on the eyes
  • the dense point cloud data of the first reference face is adjusted
  • the corresponding eyes are adjusted accordingly.
  • the dense point cloud data of the upper eyelid part is adjusted upward and/or the position coordinate of the dense point cloud corresponding to the lower eyelid part is moved downward, so that the obtained second reference face is dense
  • the dense point cloud data corresponding to the eyes in the point cloud data shows that the eyes are enlarged.
  • the virtual face image is generated based on the first reference face
  • the dense point cloud data of the second reference face is generated by using the neural network obtained by pre-training, exemplarily, according to the preset style
  • the first reference face is subjected to graphic image processing to generate a virtual face image of the second reference face with a preset style.
  • the graphic image processing may include, for example, picture editing, picture drawing, and picture design.
  • the dense point cloud data of the corresponding second reference face may be determined by using a neural network obtained by pre-training.
  • the method of determining the dense point cloud data of the second reference face by using the pre-trained neural network is the same as the above-mentioned method of using the pre-trained neural network to determine the dense point cloud data of the original face and the dense point cloud data of the first reference face.
  • the method of cloud data is similar, and will not be repeated here.
  • the obtained dense point cloud data of the second reference face can be represented as CART mesh .
  • the dense points of the target face model can be determined. cloud data.
  • an embodiment of the present disclosure further provides a specific method for determining dense point cloud data of a target face model, including the following steps S801 to S802.
  • S801 Based on the multiple sets of dense point cloud data of the second reference face, generate mean data of the multiple sets of dense point cloud data of the second reference face.
  • averaging processing may be performed based on the coordinate values of the corresponding parts in the dense point cloud data CART mesh of the second reference face to generate mean data of the dense point cloud data of the second reference face.
  • the mean data is used to represent the mean features of the dense point cloud data of multiple groups of second reference faces.
  • the dense point clouds of multiple groups of faces in the dense point cloud data of the second reference face may be Cloud data is represented as
  • any group of dense point cloud data of the second reference face Multiple face dense point clouds corresponding to different positions can be determined.
  • W different face dense point clouds can be represented as P 1 , P 2 , ..., P W
  • the corresponding coordinate values are expressed as use each group Calculate the mean value of the coordinate values of the corresponding part positions in the middle, and then the mean value data of the corresponding part positions can be obtained.
  • the first face dense point cloud P 1 at different positions for example, the following formula (1) can be used to obtain the mean data
  • the method of determining the mean data of the dense point clouds of other different parts of the face is similar to the above-mentioned method of calculating the mean data of the dense point cloud of the first face, and will not be repeated here.
  • the mean data of different positions can be obtained That is, the mean data of the dense point cloud data of the second reference face can be expressed as
  • S802 Generate dense point cloud data of the target face model based on the respective fitting coefficients corresponding to the dense point cloud data of the multiple groups of the second reference face, the mean data, and the dense point cloud data of the multiple groups of the first reference face.
  • an embodiment of the present disclosure further provides a specific method for generating dense point cloud data of a target face model, including:
  • S901 Determine the difference data of the dense point cloud data of each group of second reference faces based on the dense point cloud data of each group of second reference faces and the mean data in the dense point cloud data of the multiple groups of second reference faces .
  • the difference values can be made between the dense point cloud data of each group of second reference faces and the coordinate values of the corresponding different positions in the mean data to determine the difference data of the dense point cloud data of each group of second reference faces.
  • ⁇ mesh CART mesh -MEAN mesh .
  • the difference data may represent the degree of difference between the dense point cloud data of each group of second reference faces and the average feature of the dense point cloud data of multiple groups of second reference faces, respectively.
  • the difference data ⁇ mesh since there are multiple dense point cloud data of the second reference face, the difference data ⁇ mesh includes multiple sub-difference data corresponding to the dense point cloud data of the multiple second reference faces respectively.
  • the corresponding difference data ⁇ mesh may include, for example, the same as that contained in the dense point cloud data of the N groups of second reference faces.
  • Corresponding sub-difference data That is, the difference data ⁇ mesh contains N sub-difference data, which can be expressed as
  • S902 Perform interpolation processing on the difference data corresponding to the dense point cloud data of the plurality of groups of the second reference face based on the fitting coefficients corresponding to the dense point cloud data of the plurality of groups of the first reference face respectively.
  • the fitting coefficients may be used as weights corresponding to the dense point cloud data of multiple groups of second reference faces, and weighted summation processing is performed on the difference data corresponding to the dense point cloud data of multiple groups of second reference faces. , to realize the process of interpolation processing.
  • the difference data corresponding to the dense point cloud data of multiple sets of second reference faces are weighted and summed by the fitting coefficient, and the obtained result can be expressed as AIM mesh , which is used to compare the dense point cloud of the original face with the first
  • the association relationship between the dense point clouds of the reference face is transferred between the dense point cloud of the original face and the dense point cloud of the second reference face, so that the obtained dense point cloud of the face has the dense point cloud of the original face.
  • the characteristics of the point cloud also have the style of the dense point cloud of the second reference face; in this case, the difference data corresponding to the dense point cloud data of the multiple groups of the second reference face are weighted and summed.
  • the AIM mesh satisfies the following formula (2):
  • ⁇ i ,i ⁇ [1, N ] represents the weights ⁇ 1 , ⁇ 2 , .
  • Difference data corresponding to point cloud data The respective corresponding weights are used to represent the contribution and/or importance of different difference data to the result obtained by the weighted summation, and may be preset or adjusted, and the specific setting and adjustment methods will not be repeated here.
  • S903 Generate dense point cloud data of the target face model based on the result of the interpolation processing and the mean data.
  • the result of the weighted sum can be directly superimposed on the mean data to generate the dense point cloud data of the target face model, and expressed as OUT mesh , that is, the target can be obtained by using the following formula (3) Dense point cloud data OUT mesh of the face model:
  • the dense point cloud data OUT mesh of the target face model that includes both the face features in the target image and the preset style reflected by the dense point cloud data of the second reference face can be obtained.
  • a target face model corresponding to the target image can be generated by using the dense point cloud data of the target face model.
  • a corresponding target face model can be generated by means of rendering, or, a mask having an associated relationship with the dense point cloud data OUT mesh of the target face model can be determined. skin, and use the dense point cloud data OUT mesh of the target face model to generate the target face model corresponding to the target image.
  • the specific method can be selected according to the actual situation, and will not be repeated here.
  • Embodiments of the present disclosure also provide a description of a specific process for obtaining the target virtual face model Mod Aim corresponding to the original face A in the target image Pic A by using the face reconstruction method provided by the embodiments of the present disclosure.
  • the steps of determining the target virtual face model Mod Aim include the following (1) to (4).
  • (3-1) Determine the first type of dense point cloud data corresponding to the part of the target face model as well as
  • (3-4) Determine the intermediate coefficient corresponding to the second type of dense point cloud except the first type of dense point cloud data in the dense point cloud data of the first reference face as the second fitting coefficient, and use the first The fitting coefficient, and the second fitting coefficient determine the fitting coefficient Alpha.
  • (4-1) Determine the mean data MEAN mesh of the dense point cloud data of the plurality of groups of second reference faces.
  • (4-2) Determine the dense point cloud data of the target face model; based on the dense point cloud data CART mesh , the mean data MEAN mesh of multiple groups of the second reference face, and the dense point cloud data of multiple groups of the first reference face The corresponding fitting coefficients Alpha, respectively, generate the dense point cloud data AIM mesh of the target face model.
  • the writing order of each step does not mean a strict execution order but constitutes any limitation on the implementation process, and the specific execution order of each step should be based on its function and possible Internal logic is determined.
  • the embodiment of the present disclosure also provides a face reconstruction apparatus corresponding to the face reconstruction method.
  • a face reconstruction apparatus corresponding to the face reconstruction method.
  • FIG. 10 is a schematic diagram of a face reconstruction apparatus provided by an embodiment of the present disclosure
  • the apparatus includes: a first acquisition module 10, a first processing module 20, a determination module 30, and a generation module 40; wherein,
  • the first acquisition module 10 is used for acquiring the dense point cloud data of the original face included in the target image
  • the first processing module 20 is used to fit the dense point cloud data of the original face with the dense point cloud data of the first reference face corresponding to the multiple reference images respectively, and obtain multiple groups of the first reference face.
  • the determination module 30 is configured to determine the target person based on the respective fitting coefficients corresponding to the dense point cloud data of multiple groups of second reference faces with preset styles and the respective corresponding fitting coefficients of the dense point cloud data of multiple groups of the first reference faces Dense point cloud data of the face model; the multiple sets of second reference faces are respectively generated based on the first reference faces in the multiple reference images;
  • the generating module 40 is configured to generate a target face model corresponding to the original face of the target image based on the dense point cloud data of the target face model.
  • the determining module 30 is based on the dense point cloud data of multiple groups of second reference faces with preset styles and the dense point cloud data of multiple groups of the first reference faces, respectively.
  • the corresponding fitting coefficient when determining the dense point cloud data of the target face model, is used to: generate multiple groups of dense point cloud data of the second reference face based on the multiple groups of dense point cloud data of the second reference face Mean data of cloud data; based on multiple sets of dense point cloud data of the second reference face, the mean data, and the corresponding fitting coefficients of multiple sets of dense point cloud data of the first reference face, generate Dense point cloud data of the target face model.
  • the determining module 30 is based on the dense point cloud data of multiple groups of the second reference face, the mean data, and the dense point cloud of multiple groups of the first reference face.
  • the fitting coefficients corresponding to the data, when generating the dense point cloud data of the target face model, are used for: each group of the second reference faces based on the dense point cloud data of the multiple groups of the second reference faces.
  • the dense point cloud data and the mean value data determine the difference data of the dense point cloud data of each group of the second reference face; Fitting coefficient, performing interpolation processing on the difference data corresponding to the dense point cloud data of the second reference faces respectively; based on the results of the interpolation processing and the mean data, generating the Dense point cloud data.
  • the first obtaining module 10 when acquiring the dense point cloud data of the original face included in the target image, is used to: obtain the target image including the original face; use The pre-trained neural network processes the target image to obtain dense point cloud data of the original face in the target image.
  • the apparatus further includes a second processing module 50, configured to obtain the dense point cloud data of the first reference face corresponding to the multiple reference images in the following manner: multiple reference images of the reference face; for each of the multiple reference images, use a pre-trained neural network to process each of the reference images to obtain the The first reference face is dense point cloud data.
  • the first processing module 20 uses the dense point cloud data of the first reference face corresponding to the multiple reference images to fit the dense point cloud data of the original face, and obtains multiple When the corresponding fitting coefficients of the dense point cloud data of the first reference face are set, it is used to: perform a minimum calculation on the dense point cloud data of the original face and the dense point cloud data of the first reference face. Square processing to obtain multiple sets of intermediate coefficients corresponding to the dense point cloud data of the first reference face; based on the intermediate coefficients corresponding to the dense point cloud data of each set of first reference faces, determine The fitting coefficient corresponding to the dense point cloud data of the reference face.
  • the first processing module 20 determines the dense point cloud data of each group of the first reference faces based on the intermediate coefficients corresponding to the dense point cloud data of each group of the first reference faces.
  • the corresponding fitting coefficient it is used to: determine, from the dense point cloud data of each group of the first reference face, the first type of dense points representing the first reference face corresponding to the part of the target face model cloud data; adjusting the intermediate coefficients corresponding to the first type of dense point cloud data in the dense point cloud data of the first reference face to obtain a first fitting coefficient;
  • the intermediate coefficient corresponding to the second type of dense point cloud data in the point cloud data is determined as the second fitting coefficient;
  • the second type of dense point cloud data is the dense point cloud data of the first reference face.
  • Dense point cloud data other than the first type of dense point cloud data based on the first fitting coefficient and the second fitting coefficient, obtain the fitting coefficient of the dense point cloud data of each group of the first reference face .
  • the apparatus further includes an adjustment module 60, configured to obtain the dense point cloud data of the second reference face with the preset style in the following manner: Adjusting with reference to the dense point cloud data of the face to obtain the dense point cloud data of the second reference face with the preset style; or, based on the first reference face in the reference image, generating A virtual face image of a second reference face with a preset style; using a pre-trained neural network to generate dense point cloud data of the second reference face in the virtual face image.
  • the apparatus further includes a training module 70, which, when training the neural network, is configured to: obtain a sample image set; A sample image; the plurality of first sample images are divided into a plurality of first sample image subsets, and each first sample image subset includes images of the same type obtained from a plurality of preset collection angles respectively.
  • the image of the first sample face of the expression obtain the dense point cloud data of the first sample face of the first sample image in the sample image set; use the initial neural network to The sample image is subjected to feature learning to obtain the predicted dense point cloud data of the first sample face of the first sample image; using the dense point cloud data and predicted dense point cloud data of the first sample face, all The neural network is trained.
  • the sample image set further includes a plurality of second sample images including second sample faces and backgrounds
  • the training module 70 is further configured to: acquire each of the second sample images. face key point data; use the face key point data of the second sample image and the second sample image to fit the dense point cloud data of the second sample face of the second sample image; use the a neural network, which performs feature learning on the second sample image in the sample image set to obtain the predicted dense point cloud data of the second sample face of the second sample image; using the dense points of the second sample face
  • the neural network is trained on cloud data and predicted dense point cloud data.
  • the sample image set further includes: a third sample image; the third sample image is obtained by performing data enhancement processing on the first sample image;
  • the training module 70 is further configured to: obtain the dense point cloud data of the third sample face of the third sample image; use the initial neural network to perform feature learning on the third sample image to obtain the third sample image The predicted dense point cloud data of the third sample face;
  • the neural network is trained using the dense point cloud data of the third sample face and the predicted dense point cloud data.
  • the data enhancement processing includes at least one of the following: random occlusion processing, Gaussian noise processing, motion blur processing, and color region channel change processing.
  • an embodiment of the present disclosure further provides a computer device, including a processor 11 and a memory 12; the processor 11 is configured to execute machine-readable instructions stored in the memory 12, and the machine-readable instructions are processed When the processor 11 is executed, the processor 11 performs the following steps:
  • the dense point cloud data of the original face included in the target image use the dense point cloud data of the first reference face corresponding to the multiple reference images to fit the dense point cloud data of the original face, and obtain multiple sets of first reference
  • the above-mentioned memory 12 includes a memory 121 and an external memory 122; the memory 121 here is also called an internal memory, which is used to temporarily store the operation data in the processor 11 and the data exchanged with the external memory 122 such as the hard disk.
  • the external memory 122 performs data exchange.
  • Embodiments of the present disclosure further provide a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the face reconstruction method described in the above method embodiments are executed.
  • the storage medium may be a volatile or non-volatile computer-readable storage medium.
  • Embodiments of the present disclosure further provide a computer program product, where the computer program product carries program codes, and the instructions included in the program codes can be used to execute the steps of the face reconstruction method described in the foregoing method embodiments.
  • the computer program product carries program codes
  • the instructions included in the program codes can be used to execute the steps of the face reconstruction method described in the foregoing method embodiments.
  • the above-mentioned computer program product can be specifically implemented by means of hardware, software or a combination thereof.
  • the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (Software Development Kit, SDK), etc. Wait.
  • the units described as separate components may or may not be physically separated, and components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution in this embodiment.
  • each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit.
  • the functions, if implemented in the form of software functional units and sold or used as stand-alone products, may be stored in a processor-executable non-volatile computer-readable storage medium.
  • the computer software products are stored in a storage medium, including Several instructions are used to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure.
  • the aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), magnetic disk or optical disk and other media that can store program codes .

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • General Physics & Mathematics (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • General Health & Medical Sciences (AREA)
  • Software Systems (AREA)
  • Geometry (AREA)
  • Computer Graphics (AREA)
  • Multimedia (AREA)
  • Human Computer Interaction (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Biophysics (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Computing Systems (AREA)
  • Molecular Biology (AREA)
  • Evolutionary Computation (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Biomedical Technology (AREA)
  • Artificial Intelligence (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)
  • Image Generation (AREA)

Abstract

本公开提供了一种人脸重建方法、装置、计算机设备及存储介质,其中,该方法包括获取目标图像中包括的原始人脸的稠密点云数据;利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合原始人脸的稠密点云数据,得到多组第一参考人脸的稠密点云数据分别对应的拟合系数;基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据;多组第二参考人脸是基于多张参考图像中的第一参考人脸生成的;基于目标人脸模型的稠密点云数据,生成与目标图像的原始人脸对应的目标人脸模型。

Description

人脸重建方法、装置、计算机设备及存储介质
相关申请的交叉引用
本专利申请要求于2020年11月25日提交的、申请号为202011337942.0、发明名称为“一种人脸重建方法、装置、计算机设备及存储介质”的中国专利申请的优先权,该申请以引用的方式并入本文中。
技术领域
本公开涉及图像处理技术领域,具体而言,涉及一种人脸重建方法、装置、计算机设备及存储介质。
背景技术
通常,能够根据真实人脸或自身喜好建立虚拟人脸三维模型,以实现人脸的重建,在游戏、动漫、虚拟社交等领域具有广泛应用。例如在游戏中,玩家可以通过游戏程序提供的人脸重建系统依照玩家提供的图像中包括的真实人脸生成虚拟人脸三维模型,并利用所生成的虚拟人脸三维模型更有代入感的参与游戏。
目前,基于人脸重建方法得到的虚拟人脸三维模型与真实人脸之间的相似度较低。
发明内容
本公开实施例至少提供一种人脸重建方法、装置、计算机设备及存储介质。
第一方面,本公开实施例提供了一种人脸重建方法,包括:获取目标图像中包括的原始人脸的稠密点云数据;利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合所述原始人脸的稠密点云数据,得到多组所述第一参考人脸的稠密点云数据分别对应的拟合系数;基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据;所述多组第二参考人脸分别是基于所述多张参考图像中的第一参考人脸生成的;基于所述目标人脸模型的稠密点云数据,生成与所述目标图像的所述原始人脸对应的目标人脸模型。
该实施方式中,利用拟合系数作为媒介,建立了原始人脸的稠密点云数据与多个第一参考人脸的稠密点云数据之间的关联关系,该关联关系,能够表征基于第一参考人脸的稠密点云数据确定的第二参考人脸的稠密点云数据、和基于原始人脸建立的目标人脸模型的稠密点云数据之间的关联,使得生成的目标人脸模型具有目标图像中原始人脸的特征(如形状特征等),与原始人脸之间具有更高的相似度,又能够使生成的目标人脸模型具有预设的风格。
一种可选的实施方式中,所述基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据,包括:基于多组所述第二参考人脸的稠密点云数据,生成多组所述第二参考人脸的稠密点云数据的均值数据;基于多组所述第二参考人脸的稠密点云数据、所述均值数据、以及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,生成所述目标人脸模型的稠密点云数据。
一种可选的实施方式中,所述基于多组所述第二参考人脸的稠密点云数据、所述均值数据、以及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,生成所述目 标人脸模型的稠密点云数据,包括:基于多组所述第二参考人脸的稠密点云数据中每组所述第二参考人脸的稠密点云数据、以及所述均值数据,确定每组所述第二参考人脸的稠密点云数据的差值数据;基于多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,对多组所述第二参考人脸的稠密点云数据分别对应的差值数据进行插值处理;基于所述插值处理的结果以及所述均值数据,生成所述目标人脸模型的稠密点云数据。
该实施方式中,利用均值数据可以准确的表征多组第二参考人脸的稠密点云数据的平均特征;利用每组第二参考人脸的稠密点云数据的差值数据,可以准确的表征每组第二参考人脸的稠密点云数据分别与多组第二参考人脸的稠密点云数据平均特征的差异度,从而,利用较为准确的差异度数据对均值数据做出调整,更为简单的确定目标人脸模型的稠密点云数据。
一种可选的实施方式中,所述获取目标图像中包括的原始人脸的稠密点云数据,包括:获取包括所述原始人脸的所述目标图像;利用预先训练的神经网络对所述目标图像进行处理,得到所述目标图像中所述原始人脸的稠密点云数据。
该实施方式中,利用原始人脸的稠密点云数据可以更准确的表征目标图像中原始人脸的人脸特征。
一种可选的实施方式中,所述方法通过以下方式获取所述多张参考图像分别对应的第一参考人脸的稠密点云数据:获取包括第一参考人脸的多张参考图像;针对多张所述参考图像中的每张所述参考图像,利用预先训练的神经网络对每张所述参考图像进行处理,得到每张参考图像中的所述第一参考人脸的稠密点云数据。
该实施方式中,利用第一参考人脸的稠密点云数据,可以更准确的表征参考图像中第一参考人脸分别对应的人脸特征。同时,利用多张包含第一参考人脸的参考图像,可以覆盖尽量广泛的人脸外形特征。
一种可选的实施方式中,所述利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合所述原始人脸的稠密点云数据,得到多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,包括:对所述原始人脸的稠密点云数据以及所述第一参考人脸的稠密点云数据进行最小二乘处理,得到多组所述第一参考人脸的稠密点云数据分别对应的中间系数;基于每组第一参考人脸的稠密点云数据对应的中间系数,确定每组所述第一参考人脸的稠密点云数据对应的拟合系数。
该实施方式中,利用拟合系数,可以准确的表征在利用多个第一参考人脸的稠密点云数据拟合原始人脸的稠密点云数据时的拟合情况。
一种可选的实施方式中,所述基于每组第一参考人脸的稠密点云数据对应的中间系数,确定每组所述第一参考人脸的稠密点云数据对应的拟合系数,包括:从每组所述第一参考人脸的稠密点云数据中确定表征所述第一参考人脸中与目标人脸模型的部位对应的第一类稠密点云数据;对所述第一参考人脸的稠密点云数据中所述第一类稠密点云数据对应的中间系数进行调整,得到第一拟合系数;将所述第一参考人脸的稠密点云数据中第二类稠密点云数据对应的中间系数,确定为第二拟合系数;所述第二类稠密点云数据为所述第一参考人脸的稠密点云数据中除所述第一类稠密点云数据以外的稠密点云数据;基于所述第一拟合系数和所述第二拟合系数,得到每组所述第一参考人脸的稠密点云数据的拟合系数。
一种可选的实施方式中,所述方法通过以下方式获取所述具有预设风格的第二参考人脸的稠密点云数据:对所述参考图像中第一参考人脸的稠密点云数据进行调整,得到所述具有预设风格的第二参考人脸的稠密点云数据;或者,基于所述参考图像中的第一参考人脸,生成包括具有所述预设风格的第二参考人脸的虚拟人脸图像;利用预先训练 的神经网络生成所述虚拟人脸图像中所述第二参考人脸的稠密点云数据。
该实施方式中,通过对部分人脸部位对应的拟合系数做出调整,可以使得基于拟合系数和多个对应的第一参考人脸的稠密点云数据拟合得到的稠密点云数据与目标图像对应的原始人脸的稠密点云数据相近,也即得到的拟合系数可以更准确的表征第一参考人脸稠密点云拟合原始人脸稠密点云时的系数。
一种可选的实施方式中,训练所述神经网络,包括:获取样本图像集合;所述样本图像集合包括包含第一样本人脸的多张第一样本图像;所述多张第一样本图像划分为多个第一样本图像子集,每个第一样本图像子集中包括从多个预设采集角度分别采集得到的具有同种表情的第一样本人脸的图像;获取所述样本图像集合中的第一样本图像的第一样本人脸的稠密点云数据;利用所述神经网络,对所述样本图像集合中的第一样本图像进行特征学习,得到所述第一样本图像的第一样本人脸的预测稠密点云数据;利用所述第一样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
一种可选的实施方式中,所述样本图像集合还包括包含第二样本人脸和背景的多张第二样本图像,所述训练所述神经网络还包括:获取每张所述第二样本图像的人脸关键点数据;利用第二样本图像的人脸关键点数据以及所述第二样本图像,拟合生成所述第二样本图像的第二样本人脸的稠密点云数据;利用所述神经网络,对所述样本图像集合中的第二样本图像进行特征学习,得到所述第二样本图像的第二样本人脸的预测稠密点云数据;利用所述第二样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
一种可选的实施方式中,所述样本图像集合中还包括第三样本图像;所述第三样本图像为对所述第一样本图像进行数据增强处理得到;
所述训练所述神经网络还包括:
获取所述第三样本图像的第三样本人脸的稠密点云数据;利用所述神经网络对所述第三样本图像进行特征学习,得到所述第三样本图像的第三样本人脸的预测稠密点云数据;
利用所述第三样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
一种可选的实施方式中,所述数据增强处理包括下述至少一种:随机遮挡处理、高斯噪声处理、运动模糊处理、以及颜色区域通道改变处理。
该实施方式中,可以通过对第一样本图像、第二样本图像、以及第三样本图像的数量进行调整,得到不同优势的神经网络,以针对实际需求得到更优的神经网络;同时,由于第三样本图像是通过数据增强处理得到的,因此在样本图像中包括第三样本图像时,训练得到的神经网络对数据处理的能力更强。
同时,由于第二样本图像中包括的人脸可以覆盖尽量广泛的人脸外形特征,因此得到的神经网络可以具有较好的泛化能力。
第二方面,本公开实施例还提供一种人脸重建装置,包括:
第一获取模块,用于获取目标图像中包括的原始人脸的稠密点云数据;
第一处理模块,用于利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合所述原始人脸的稠密点云数据,得到多组所述第一参考人脸的稠密点云数据分别对应的拟合系数;
确定模块,用于基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组所 述第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据;所述多组第二参考人脸分别是基于所述多张参考图像中的第一参考人脸生成的;
生成模块,用于基于所述目标人脸模型的稠密点云数据,生成与所述目标图像的所述原始人脸对应的目标人脸模型。
一种可选的实施方式中,所述确定模块在基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据时,用于:基于多组所述第二参考人脸的稠密点云数据,生成多组所述第二参考人脸的稠密点云数据的均值数据;基于多组所述第二参考人脸的稠密点云数据、所述均值数据、以及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,生成所述目标人脸模型的稠密点云数据。
一种可选的实施方式中,所述确定模块在基于多组所述第二参考人脸的稠密点云数据、所述均值数据、以及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,生成所述目标人脸模型的稠密点云数据时,用于:基于多组所述第二参考人脸的稠密点云数据中每组所述第二参考人脸的稠密点云数据、以及所述均值数据,确定每组所述第二参考人脸的稠密点云数据的差值数据;基于多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,对多组所述第二参考人脸的稠密点云数据分别对应的差值数据进行插值处理;基于所述插值处理的结果以及所述均值数据,生成所述目标人脸模型的稠密点云数据。
一种可选的实施方式中,所述第一获取模块在获取目标图像中包括的原始人脸的稠密点云数据时,用于:获取包括所述原始人脸的所述目标图像;利用预先训练的神经网络对所述目标图像进行处理,得到所述目标图像中所述原始人脸的稠密点云数据。
一种可选的实施方式中,所述装置还包括第二处理模块,用于通过以下方式获取所述多张参考图像分别对应的第一参考人脸的稠密点云数据:获取包括第一参考人脸的多张参考图像;针对多张所述参考图像中的每张所述参考图像,利用预先训练的神经网络对每张所述参考图像进行处理,得到每张参考图像中的所述第一参考人脸的稠密点云数据。
一种可选的实施方式中,所述第一处理模块在利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合所述原始人脸的稠密点云数据,得到多组所述第一参考人脸的稠密点云数据分别对应的拟合系数时,用于:对所述原始人脸的稠密点云数据以及所述第一参考人脸的稠密点云数据进行最小二乘处理,得到多组所述第一参考人脸的稠密点云数据分别对应的中间系数;基于每组第一参考人脸的稠密点云数据对应的中间系数,确定每组所述第一参考人脸的稠密点云数据对应的拟合系数。
一种可选的实施方式中,所述第一处理模块在基于每组第一参考人脸的稠密点云数据对应的中间系数,确定每组所述第一参考人脸的稠密点云数据对应的拟合系数时,用于:从每组所述第一参考人脸的稠密点云数据中确定表征所述第一参考人脸中与目标人脸模型的部位对应的第一类稠密点云数据;对所述第一参考人脸的稠密点云数据中所述第一类稠密点云数据对应的中间系数进行调整,得到第一拟合系数;将所述第一参考人脸的稠密点云数据中第二类稠密点云数据对应的中间系数,确定为第二拟合系数;所述第二类稠密点云数据为所述第一参考人脸的稠密点云数据中除所述第一类稠密点云数据以外的稠密点云数据;基于所述第一拟合系数和所述第二拟合系数,得到每组所述第一参考人脸的稠密点云数据的拟合系数。
一种可选的实施方式中,所述装置还包括调整模块,用于通过以下方式获取所述具有预设风格的第二参考人脸的稠密点云数据:对所述参考图像中第一参考人脸的稠密点 云数据进行调整,得到所述具有预设风格的第二参考人脸的稠密点云数据;或者,基于所述参考图像中的第一参考人脸,生成包括具有所述预设风格的第二参考人脸的虚拟人脸图像;利用预先训练的神经网络生成所述虚拟人脸图像中所述第二参考人脸的稠密点云数据。
一种可选的实施方式中,所述装置还包括训练模块,在训练所述神经网络时,用于:获取样本图像集合;所述样本图像集合包括包含第一样本人脸的多张第一样本图像;所述多张第一样本图像划分为多个第一样本图像子集,每个第一样本图像子集中包括从多个预设采集角度分别采集得到的具有同种表情的第一样本人脸的图像;获取所述样本图像集合中的第一样本图像的第一样本人脸的稠密点云数据;利用神经网络,对所述样本图像集合中的第一样本图像进行特征学习,得到所述第一样本图像的第一样本人脸的预测稠密点云数据;利用所述第一样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
一种可选的实施方式中,所述样本图像集合还包括包含第二样本人脸和背景的多张第二样本图像,所述训练模块还用于:获取每张所述第二样本图像的人脸关键点数据;利用第二样本图像的人脸关键点数据以及所述第二样本图像,拟合生成所述第二样本图像的第二样本人脸的稠密点云数据;利用所述神经网络,对所述样本图像集合中的第二样本图像进行特征学习,得到所述第二样本图像的第二样本人脸的预测稠密点云数据;利用所述第二样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
一种可选的实施方式中,所述样本图像集合中还包括:第三样本图像;所述第三样本图像为对所述第一样本图像进行数据增强处理得到;
所述训练模块还用于:获取所述第三样本图像的第三样本人脸的稠密点云数据;利用所述神经网络对所述第三样本图像进行特征学习,得到所述第三样本图像的第三样本人脸的预测稠密点云数据;
利用所述第三样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
一种可选的实施方式中,所述数据增强处理包括下述至少一种:随机遮挡处理、高斯噪声处理、运动模糊处理、以及颜色区域通道改变处理。
第三方面,本公开可选实现方式还提供一种计算机设备,包括处理器和存储器,所述处理器用于执行所述存储器中存储的机器可读指令,所述机器可读指令被所述处理器执行时,所述机器可读指令被所述处理器执行时执行上述第一方面,或第一方面中任一种可能的实施方式中的步骤。
第四方面,本公开可选实现方式还提供一种计算机可读存储介质,其上存储的计算机程序被运行时执行上述第一方面,或第一方面中任一种可能的实施方式中的步骤。
关于上述人脸重建装置、计算机设备、及计算机可读存储介质的效果描述参见上述人脸重建方法的说明,这里不再赘述。
为使本公开的上述目的、特征和优点能更明显易懂,下文特举较佳实施例,并配合所附附图,作详细说明如下。
附图说明
为了更清楚地说明本公开实施例的技术方案,下面将对实施例中所需要使用的附图作简单地介绍。这些附图示出了符合本公开的实施例,并与说明书一起用于说明本公开的技术方案。应当理解,以下附图仅示出了本公开的某些实施例,因此不应被看作是对 范围的限定,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他相关的附图。
图1示出了本公开实施例所提供的一种人脸重建方法的流程图;
图2示出了本公开实施例所提供的一种训练神经网络具体方法的流程图;
图3示出了本公开实施例所提供的第一样本图像、以及利用第一样本图像确定的第三样本图像的示意图;
图4示出了本公开实施例所提供的一种获取第二样本图像的第二样本人脸的稠密点云数据的具体方法的流程图;
图5示出了本公开实施例所提供的一种神经网络结构的具体示例图;
图6示出了一种确定多张参考图像分别对应的拟合系数的方法的流程图;
图7示出了本公开实施例所提供的一种确定每组第一参考人脸的稠密点云数据对应的拟合系数的具体方法的流程图;
图8示出了本公开实施例所提供的一种确定目标人脸模型的稠密点云数据的具体方法的流程图;
图9示出了本公开实施例所提供的一种生成目标人脸模型的稠密点云数据的具体方法的流程图;
图10示出了本公开实施例所提供的一种人脸重建装置的示意图;
图11示出了本公开实施例所提供的一种计算机设备的示意图。
具体实施方式
为使本公开实施例的目的、技术方案和优点更加清楚,下面将结合本公开实施例中附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本公开一部分实施例,而不是全部的实施例。通常在此处描述和示出的本公开实施例的组件可以以各种不同的配置来布置和设计。因此,以下对本公开的实施例的详细描述并非旨在限制要求保护的本公开的范围,而是仅仅表示本公开的选定实施例。基于本公开的实施例,本领域技术人员在没有做出创造性劳动的前提下所获得的所有其他实施例,都属于本公开保护的范围。
经研究发现,在基于人像图像中包括的人脸进行人脸重建得到虚拟人脸三维模型的时候,通常先基于人脸图像获取人脸对应的人脸稠密点云,再根据虚拟人脸三维模型的具体风格对人脸稠密点云进行多次调整,以生成虚拟图像。由于不同人脸对应的人脸稠密点云不同,即使要重建的虚拟人脸的风格相同,在依据确定的风格对不同人脸对应的人脸稠密点云进行人脸重建时进行的调整也不尽相同,且调整具有较大的不确定性,造成了在调整过程中很难控制调整的具体方向,造成所生成的虚拟人脸三维模型与真实人脸之间可能会存在较大的差异,导致虚拟人脸三维模型与真实人脸之间的相似度较低的问题。
另外,对不同人脸进行人脸重建时,需要依据不同人脸对应的人脸稠密点云、及风格对人脸稠密点云的要求,设置针对不同人脸的调整方案,造成对基于不同调整方案对每个人脸的人脸稠密点云进行调整时,都会消耗较多的时间,导致人脸重建的效率较低。
基于上述研究,本公开提供了一种人脸重建方法、装置、计算机设备及存储介质,利用拟合系数作为媒介,建立了原始人脸的稠密点云数据与多个第一参考人脸的稠密点 云数据之间的关联关系,该关联关系能够表征基于第一参考人脸的稠密点云数据确定的第二参考人脸的稠密点云数据、和基于原始人脸建立的目标人脸模型的稠密点云数据之间的关联,使得生成的目标人脸模型具有目标图像中原始人脸的特征(如形状特征等),与原始人脸之间具有更高的相似度,并且能够使生成的目标人脸模型具有预设的风格。
此外,本方案针对不同的风格,只需要生成多张参考图像分别对应的第二参考人脸的稠密点云数据,并利用具有预设风格的第二参考人脸的稠密点云数据,来为原始人脸生成目标人脸模型,无需对不同原始人脸针对性的确定调整方案,而是利用相同的第二参考人脸的稠密点云数据、以及不同原始人脸的拟合系数,确定不同原始人脸的目标人脸模型,具有更高的处理效率。
针对以上方案所存在的缺陷,均是发明人在经过实践并仔细研究后得出的结果,因此,上述问题的发现过程以及下文中本公开针对上述问题所提出的解决方案,都应该是发明人在本公开过程中对本公开做出的贡献。
应注意到:相似的标号和字母在下面的附图中表示类似项,因此,一旦某一项在一个附图中被定义,则在随后的附图中不需要对其进行进一步定义和解释。
为便于对本实施例进行理解,首先对本公开实施例所公开的一种人脸重建方法进行详细介绍,本公开实施例所提供的人脸重建方法的执行主体一般为具有一定计算能力的计算机设备,该计算机设备例如包括:终端设备或服务器或其它处理设备,终端设备可以为用户设备(User Equipment,UE)、移动设备、用户终端、终端、蜂窝电话、无绳电话、个人数字助理(Personal Digital Assistant,PDA)、手持设备、计算设备、车载设备、可穿戴设备等。在一些可能的实现方式中,该人脸重建方法可以通过处理器调用存储器中存储的计算机可读指令的方式来实现。
下面对本公开实施例提供的人脸重建方法加以说明。
参见图1所示,本公开实施例提供一种人脸重建方法,所述方法包括步骤S101至S104,其中:
S101:获取目标图像中包括的原始人脸的稠密点云数据;
S102:利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合原始人脸的稠密点云数据,得到多组第一参考人脸的稠密点云数据分别对应的拟合系数;
S103:基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据;多组第二参考人脸分别是基于多张参考图像中的第一参考人脸生成的;
S104:基于目标人脸模型的稠密点云数据,生成与目标图像的原始人脸对应的目标人脸模型。
本公开实施例利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合目标图像中原始人脸的稠密点云数据时所确定的拟合系数,以及基于参考图像中第一参考人脸生成的且具有预设风格的第二参考人脸的稠密点云数据,得到目标人脸模型的稠密点云数据,并基于目标人脸模型的稠密点云数据,生成与目标图像对应的目标人脸模型,该过程利用拟合系数作为媒介,建立了原始人脸的稠密点云数据与多个第一参考人脸的稠密点云数据之间的关联关系,该关联关系能够表征基于第一参考人脸的稠密点云确定的第二参考人脸的稠密点云、和基于原始人脸的稠密点云数据建立的目标人脸模型的稠密点云数据之间的关联,使得生成的目标人脸模型的稠密点云数据具有目标图像中原始人脸的特征(如形状特征等),与原始人脸之间具有更高的相似度,并且能够使生成的目标人脸模型具有预设的风格。
下面对上述步骤S101至S104加以详细说明。
针对上述步骤S101:
目标图像例如为预先获取的包括人脸的图像,例如,在利用诸如相机等的拍摄设备对某一对象进行拍摄时获取的包括人脸的图像。此时,例如可以将图像中包括的任一张人脸确定为原始人脸,并将原始人脸作为人脸重建的对象。
在将本公开实施例提供的人脸重建方法应用于不同的场景下时,目标图像的获取方法也有所区别。
例如,在将该人脸重建方法应用于游戏中的情况下,可以通过游戏设备中安装的图像获取设备获取包括了游戏玩家的脸部的图像,或者可以从游戏设备中的相册中选择包括了游戏玩家的脸部的图像、并将获取的包括了游戏玩家的脸部的图像作为目标图像。
又例如,在将人脸重建方法应用于手机等终端设备的情况下,可以由终端设备的摄像头采集包括用户脸部的图像,或者从终端设备的相册中选择包括了用户人脸的图像、或者从终端设备中安装的其他应用程序中接收包括用户脸部的图像。
又例如,在将人脸重建方法应用于直播场景下,可以从直播设备获取的视频流中包括的多帧视频帧图像中获取包含人脸的视频帧图像;并将包含人脸的视频帧图像作为目标图像。此处,目标图像例如可以有多帧;多帧目标图像例如可以是对视频流中的多帧视频帧图像进行采样获得。
在获取目标图像中包括的原始人脸的稠密点云数据时,例如可以采用下述方式:获取包括原始人脸的目标图像;利用预先训练的神经网络对目标图像进行处理,得到目标图像中原始人脸的稠密点云数据。
此处,在利用预先训练的神经网络对目标图像进行处理,得到原始人脸的稠密点云数据时,预先训练的神经网络包括下述至少一种:卷积神经网络(Convolutional Neural Networks,CNN)、反向传播神经网络(Back Propagation,BP)、骨干神经网络(Backbone)。
在确定神经网络结构时,例如可以先确定神经网络的骨干网络(Backbone network),其作为神经网络的主体架构,示例性的,骨干网络例如可以包括下述至少一种:起始网络(Inception)、残差网络变体网络(the next dimension to RESNET,ResNeXt)、起始网络变体网络(Xception)、挤压和激励网络(Squeeze-and-Excitation Networks,SENet)、轻量化网络(MobileNet)、以及轻量级网络(ShuffleNet)。
示例性的,神经网络包括卷积神经网络的情况下,还可选用轻量化网络作为卷积神经网络基础模型,在轻量化网络的基础上,增加其他网络结构,构成卷积神经网络,并对构成的卷积神经网络进行训练。该过程由于采用了轻量化网络作为卷积神经网络的一部分,且轻量化网络体积小、数据处理速度快,因此训练的速度也更快;此外,训练得到的神经网络还具有体积小、数据处理速度快的优势,更适于部署在嵌入式设备中。
此处,上述神经网络的网络结构仅为一种示例;网络结构的具体构建方式和结构可以按照实际情况确定,在此不再赘述,上述示例也不构成对本公开实施例的限定。
参见图2所示,本公开实施例提供了一种训练神经网络的具体方法,包括:
S201:获取样本图像集合;样本图像集合包括包含第一样本人脸的多张第一样本图像;多张第一样本图像划分为多个第一样本图像子集,每个第一样本图像子集中包括从多个预设采集角度分别采集得到的具有同种表情的第一样本人脸的图像。
其中,对于样本图像集合包括的第一样本人脸的多张第一样本图像,对应的第一样本人脸例如为预先确定的用于获取人脸图像以训练神经网络的至少一个个体对象的人 脸。
在对第一样本人脸进行拍摄以确定第一样本图像子集时,例如可以针对多个不同的表情确定多个第一样本图像子集。其中,多个不同的表情例如为快乐、兴奋、失落、难过等。在对第一样本人脸进行拍摄以获取第一样本图像子集的情况下,以针对“快乐”表情对第一样本人脸进行拍摄为例,通过从不同角度对呈现出“快乐”表情的第一样本人脸进行拍摄,可以获取到在“快乐”表情下对应的多张第一样本图像,作为第一样本图像子集。此处,在获取第一样本图像时,不同的第一样本图像中第一样本人脸的背景可以相同。
类似的,可以采用相似的方法,获取第一样本人脸在不同的表情下分别对应的第一样本图像子集,在此不再赘述。
在确定多个不同表情分别对应的第一样本图像子集后,即可以确定第一样本人脸的多张第一样本图像。
示例性的,在获取第一样本图像时,可以利用图像采集设备进行拍摄获取图像,其中,图像采集设备例如包括深度相机和彩色相机中至少一种。
例如,可以对I个个体对象的人脸在E个表情下进行拍摄,以得到多张第一样本图像。例如,可以先确定某一个体对象A的人脸作为第一样本人脸。在第一样本人脸在呈现“难过”的表情下,从P(P为大于1的整数)个不同角度对第一样本人脸进行拍摄,得到在“难过”表情下的P个不同角度对应的多张图像,以作为“难过”的表情对应的第一样本图像子集。然后在第一样本人脸呈现其他表情(例如包括E-1个不同表情)下获取P个角度对应的P张图像,作为其他表情对应的第一样本图像子集。相应地,可以针对个体对象A的E个不同表情确定E个不同的第一样本图像子集,相应地,可以获取到个体对象A的P×E张图像。然后,可对其余I-1个个体对象进行拍摄,以获取其余I-1个个体对象分别对应的P×E张图像,也即共获取M(M=I×E×P)张图像,作为多张第一样本图像。
由于获取的多张第一样本图像呈现不同表情的多个拍摄个体对象对应的人脸,而背景等其他的部分较少,因此利用第一样本图像可以训练神经网络对图像中多角度下的人脸对应的稠密点云数据的获取能力。
第二样本图像可以针对不同的个体对象随机拍摄得到,也可以从预设的网络平台随机爬取多张包含人脸的图像,并将爬取得到的图像作为第二样本人脸图像。
示例性的,在针对不同的个体对象随机拍摄得到第二样本图像的情况下,可以利用相机或者其他图像拍摄设备对多个第二样本人脸进行拍摄获取第二样本图像;或者,直接获取预先拍摄得到的多张第二样本图像。第二样本图像例如包括获取的H张包括背景的人脸图像;其中,第二样本图像中包含的背景等对识别人脸产生干扰的部分,用以训练神经网络在对目标图像中人脸进行识别以获取稠密点云数据的情况下,对除人脸以外其他部分的抗干扰能力。
本公开另一实施例中,样本图像集合中还包括第三样本图像,第三样本图像例如可以通过对第一样本图像进行数据增强处理得到。其中,数据增强处理包括下述至少一种:随机遮挡处理、高斯噪声处理、模糊处理、以及颜色区域通道改变处理。利用数据增强处理的方法,可以将数据的特征集中,以减少神经网络对不相关特征的学习,从而提升神经网络在获取稠密点云数据时的性能。
示例性的,在选取随机遮挡处理的情况下,可以对第一样本图像中的部分区域做遮挡处理得到第三样本图像;其中,遮挡部分的大小可以按照第一样本图像的大小和实际 需要进行确定,在此不做限定;在选取高斯噪声处理的情况下,例如可以选取包括起伏噪声、宇宙噪声、热噪声、及散粒噪声中至少一种加入第一样本图像,以使得到第三样本图像中包含至少一种高斯噪声;在选取模糊处理的情况下,例如可以将图像中的至少部分像素点做运动模糊处理,以得到具有运动模糊效果的第三样本图像;在选取颜色区域通道改变处理的情况下,例如可以利用全局颜色区域通道改变处理,或者随机颜色区域通道改变处理对第一样本图像进行处理,得到全局或部分区域颜色区域通道改变的第三样本图像。
图3为本公开实施例提供的第一样本图像、以及利用第一样本图像确定的第三样本图像的示意图;如图3所示,31表示第一样本图像;32表示对第一样本图像31进行模糊处理的数据增强处理得到的第三样本图像;33表示对第一样本图像31进行局部遮挡的数据增强处理得到的第三样本图像,其中,34表示局部遮挡的位置。
在该种情况下,第三样本图像对应的样本人脸的稠密点云数据与生成第三样本图像的第一样本图像对应的样本人脸的稠密点云数据相同。
承接上述S201,训练神经网络的具体方法还包括:
S202:获取样本图像集合中的第一样本图像的第一样本人脸的稠密点云数据。
具体地,在获取样本图像集合中的第一样本图像的第一样本人脸的稠密点云数据以及第二样本图像的第二样本人脸的稠密点云数据时,例如可以采用下述方法:获取样本图像集合中第一样本图像的第一样本人脸的稠密点云数据、第二样本图像的第二样本人脸的稠密点云数据、以及第三样本图像的第三样本人脸的稠密点云数据。
示例性的,针对样本图像集合中的第一样本图像,可以在对第一样本人脸进行拍摄获取对应的第一样本图像后,获取各张第一样本图像中第一样本人脸的稠密点云数据。在利用彩色相机获取第一样本图像的情况下,例如可以利用人脸3D形变统计模型(3D Morphable Model,3DMM)等模型得到第一样本人脸的稠密点云数据;在利用深度相机获取第一样本图像的情况下,例如可以基于深度相机获取的深度图像,得到第一样本人脸的稠密点云数据。
针对样本图像集合中的第二样本图像,参见图4所示,本公开实施例提供了一种获取第二样本图像的第二样本人脸的稠密点云数据的具体方法,包括:
S401:获取每张第二样本图像的人脸关键点数据;
S402:利用第二样本图像的人脸关键点数据以及第二样本图像,拟合生成第二样本图像的第二样本人脸的稠密点云数据;
S403:利用神经网络,对样本图像集合中的第二样本图像进行特征学习,得到第二样本图像的第二样本人脸的预测稠密点云数据;
S404:利用第二样本人脸的稠密点云数据和预测稠密点云数据,对神经网络进行训练。
此处,第二样本图像例如包括获取的H张包括背景的人脸图像,且每张第二样本图像均包括确定的人脸关键点。其中,第二样本图像中包括的人脸关键点用于确定第二样本图像对应的样本人脸的稠密点云数据,例如包括在第二样本图像中直接标注的用于表征人脸特征的关键点,例如表征五官、颧骨、眉骨的多个关键点;或者,包括利用关键点检测方法确定的对应于人脸的关键点。其中,关键点检测方法包括下述至少一种:主动形状模型(Active Shape Model,ASM)、活动表观模型(Active Appearance Models,AAM)、级联姿势回归(Cascaded Pose Regression,CPR)。
在获得第二样本图像的人脸关键点数据的情况下,可以利用拟合模型生成第二样本图像的第二样本人脸的稠密点云数据。其中,拟合模型例如包括人脸3D形变统计模型。
承接上述S202,训练神经网络的方法还包括:
S203:利用神经网络,对样本图像集合中的第一样本图像进行特征学习,得到第一样本图像的第一样本人脸的预测稠密点云数据。
具体地,在确定第一样本图像的第一样本人脸的预测稠密点云数据以及第二样本图像的第二样本人脸的预测稠密点云数据时,例如可以采用下述方法:利用初始神经网络,对样本图像集合中的第一样本图像、第二样本图像、以及第三样本图像进行特征学习,得到第一样本图像的第一样本人脸的预测稠密点云数据、第二样本图像的第二样本人脸的预测稠密点云数据、以及第三样本图像的第三样本人脸的预测稠密点云数据。
此时,需要注意的是,获取第一样本图像、第二样本图像、以及第三样本图像中至少一种样本图像的步骤可以与利用初始神经网络对第一样本图像、第二样本图像、以及第三样本图像中至少一种进行特征学习的步骤同步执行,也即可以直接得到利用初始神经网络对第一样本图像、第二样本图像、以及第三样本图像中至少一种进行特征学习后分别对应的预测稠密点云数据。
此外,在利用初始神经网络对第一样本图像、第二样本图像、以及第三样本图像中至少一种进行特征学习时,可以同步地对第一样本图像、第二样本图像、以及第三样本图像中至少一种进行特征学习;或者,按照实际的需求对第一样本图像、第二样本图像、以及第三样本图像中至少一种按照顺序进行特征学习,以得到第一样本图像的第一样本人脸的预测稠密点云数据、第二样本图像的第二样本人脸的预测稠密点云数据、以及第三样本图像的第三样本人脸的预测稠密点云数据。
本公开实施例不对上述样本处理过程的先后执行顺序进行限定,具体可以根据实际的需要进行设定。
示例性的,对于样本图像集合中包括的第一样本图像和第二样本图像,可以按照预设的比例选取不同数量的第一样本图像和第二样本图像,并将选取的第一样本图像和第二样本图像输入至初始神经网络中;在样本图像集合包括第一样本图像、第二样本图像和第三样本图像的情况下,可以按照预设的比例选取不同数量的第一样本图像、第二样本图像、和第三样本图像,并将选取的第一样本图像、第二样本图像和第三样本图像输入至初始神经网络中。比例选取不同时,对神经网络训练的侧重也有所差异。当第一样本图像和/或第三样本图像的占比较大时,训练得到的神经网络对图像中不同角度的人脸对应的人脸稠密点云获取能力较强;当第二样本图像的占比较大时,训练得到的神经网络对图像中人脸外的其他背景部分的抗干扰能力更强,从而满足不同的使用需求。
在将样本图像输入至初始神经网络后,初始神经网络能够对样本图像进行特征学习,并输出每张样本图像的预测人脸的稠密点云数据;利用预测人脸的稠密点云数据、和每张样本图像对应的样本人脸的稠密点云数据,确定神经网络的损失,该损失用于衡量神经网络在生成人脸的稠密点云数据时的准确性。
利用神经网络得到的预测人脸的稠密点云数据例如可以包括在预设坐标系下表征人脸稠密点云的位置的坐标值,或者,包括在预设坐标系下表征人脸稠密点云的位置的x轴、y轴、z轴各自对应的坐标值。其中,预设的坐标系例如为预设的人脸坐标系。
示例性的,对于任一预测人脸的稠密点云数据,在包括表征人脸稠密点云的位置坐标值的情况下,此预测人脸的稠密点云数据可以包括坐标值(x,y,z),或者,包括在x轴上的坐标值x、y轴上的坐标值y、及z轴上的坐标值z,且不同的预测人脸的稠密 点云数据包含的在x轴、y轴、及z轴上的坐标值在输出的预测人脸的稠密点云数据中具有预设的排列顺序。
参见图5所示,本公开实施例还提供了一种神经网络结构的具体示例图,其中,神经网络结构包括:骨干网络51、第一全连接层52、以及三组第二全连接层53。
在利用此神经网络对样本图像进行处理,确定样本中人脸的预测人脸的稠密点云数据时,例如可以将样本图像输入至骨干网络,得到样本图像的特征数据,特征数据经过第一全连接层后,分别输入至三组第二全连接层;三组第二全连接层用于预测样本图像中人脸的稠密点云在人脸坐标系中的坐标值。例如,第一组第二全连接层能够输出稠密点云在人脸坐标系中x轴的坐标值,第二组第二全连接层能够输出稠密点云在人脸坐标系中y轴的坐标值,第三组第二全连接层能够输出稠密点云在人脸坐标系中z轴的坐标值。稠密点云在人脸坐标系中的坐标值,构成了样本图像中人脸的预测人脸的稠密点云数据。
在确定多张第一样本图像、及多张第二样本图像的情况下,按照第一样本图像选取比例为40%、第二样本图像选取比例为60%的预设比例选取用于对神经网络进行训练的多张样本图像。例如,在利用100张样本图像训练神经网络的情况下,选取40张第一样本图像、及60张第二样本图像作为多张样本图像。
其中,选取卷积神经网络作为初始神经网络进行训练,并将轻量化网络t作为基础模型、分割网络(split FC)作为输出层构建卷积神经网络,可以得到多张样本图像分别对应的预测人脸的稠密点云数据,且得到的预测人脸的稠密点云数据包括稠密点云中的点分别在x轴、y轴、及z轴上的多个坐标值。
此时,针对任一样本图像对应的预测人脸的稠密点云数据,在人脸稠密点云包含R个不同人脸部位的情况下,输出的预测人脸的稠密点云数据例如可以以矩阵的形式输出,表示为[x 1,x 2,…,x R]、[y 1,y 2,…,y R]、及[z 1,z 2,…,z R],也即R个不同人脸部位中任一部位对应的预测人脸的稠密点云数据包含的坐标值(x i,y i,z i),i∈[1,R]被拆分为在x轴、y轴、及z轴上的坐标值x i、y i、及z i,并分别作为三个不同矩阵中第i个元素输出。
承接上述S203,训练神经网络的方法还包括:
S204:利用第一样本人脸的稠密点云数据和预测稠密点云数据,对神经网络进行训练。
具体地,在对初始神经网络进行训练以得到已训练的神经网络时,可以采用下述方法:利用第一样本人脸的稠密点云数据和预测稠密点云数据、第二样本人脸的稠密点云数据和预测稠密点云数据、以及第三样本人脸的稠密点云数据和预测稠密点云数据,对神经网络进行训练,训练完成后得到已训练的神经网络。
示例性的,可以基于预测人脸的稠密点云数据、以及样本人脸的稠密点云数据的差值确定神经网络的损失,并利用损失对神经网络进行训练,训练的方向为使得损失减小的方向,以使神经网络在对图像进行处理时,得到的预测人脸的稠密点云数据,足够与真实人脸的稠密点云数据接近。
在得到已训练的神经网络的情况下,即可将目标图片输入神经网络,得到目标图片中原始人脸对应的稠密点云数据。
针对上述步骤S102:
参考图像例如可以为不同个体对象分别对应的人脸,不同个体对象对应的人脸不 同;示例性的,可以确定性别、年龄、肤色、胖瘦程度等中至少一项不同的多个人,针对多个人中的每个人,获取每个人的人脸图像,并将获取的人脸图像作为第二样本图像。这样,基于第二样本图像生成的第二样本人脸的稠密点云数据,能够覆盖到尽量广泛的人脸外形特征。
在获取多张参考图像对应的多个第一参考人脸的稠密点云数据时,例如可以采用下述方式:获取包括第一参考人脸的多张参考图像;针对多张参考图像中的每张参考图像,利用预先训练的神经网络获取每张参考图像中第一参考人脸的稠密点云数据。其中,利用预先训练的神经网络获取多个第一参考人脸稠密点云的方式与上述利用预先训练的神经网络获取原始人脸稠密点云的方式相似,在此不再赘述。
在确定原始人脸的稠密点云数据、及第一参考人脸的稠密点云数据的情况下,可以利用第一参考人脸的稠密点云数据拟合原始人脸的稠密点云数据,以获取与多张参考图像分别对应的第一参考人脸的稠密点云数据的拟合系数。其中,拟合系数可以作为媒介,建立目标图像中的原始人脸的稠密点云数据与多张参考图片分别对应的第一参考人脸的稠密点云数据之间的关联关系。
参见图6所示,本公开实施例提供了一种确定多张参考图像分别对应的拟合系数的方法,包括以下步骤S601至S602。
S601:对原始人脸的稠密点云数据以及第一参考人脸的稠密点云数据进行最小二乘处理,得到多组第一参考人脸的稠密点云数据分别对应的中间系数。
示例性的,将原始人脸的稠密点云数据表示为IN mesh,第一参考人脸的稠密点云数据表示为BASE mesh。其中,由于第一参考人脸稠密点云是由多个参考图像确定的,因此在存在N张参考图像的情况下,第一参考人脸的稠密点云数据BASE mesh中对应的包含有N组人脸的稠密点云数据,表示为
Figure PCTCN2021108629-appb-000001
其中,
Figure PCTCN2021108629-appb-000002
表示第i张参考图像对应的人脸的稠密点云数据。
利用IN mesh
Figure PCTCN2021108629-appb-000003
Figure PCTCN2021108629-appb-000004
进行最小二乘处理,可以得到N个拟合值,表示为α i(i∈[1,N])。其中,α i表征第i组第一参考人脸的稠密点云数据对应的拟合值。通过N个拟合值,可以确定拟合系数Alpha,例如可以用系数矩阵表示,也即Alpha=[α 12,…,α N]。
此处,在通过第一参考人脸的稠密点云数据拟合原始人脸的稠密点云数据的过程中,要使得通过拟合系数对第一参考人脸的稠密点云数据进行加权求和后的数据,与原始人脸的稠密点云数据的数据尽可能的接近。
该拟合系数又可视为用多个参考图像对应的第一参考人脸的稠密点云数据表达原始人脸的稠密点云数据时,每个第一参考人脸的稠密点云数据的表达系数。
S602:基于多组第一参考人脸的稠密点云数据中每组第一参考人脸的稠密点云数据对应的中间系数,确定每组第一参考人脸的稠密点云数据对应的拟合系数。
具体地,参见图7所示,本公开实施例还提供了一种确定每组第一参考人脸的稠密点云数据对应的拟合系数的具体方法,包括以下步骤S701至S704。
S701:从每组第一参考人脸的稠密点云数据中确定表征第一参考人脸中与目标人脸模型的部位对应的第一类稠密点云数据。
针对多组第一参考人脸的稠密点云数据中的任一组稠密点云数据
Figure PCTCN2021108629-appb-000005
由于
Figure PCTCN2021108629-appb-000006
是基于第i张参考图像得到的,因此
Figure PCTCN2021108629-appb-000007
中包含表征第i张参考图像中人脸部位对应的稠密点云数据。其中,人脸部位例如包括眉部、鼻部、眼部、嘴部、颧骨部位、下颌部位。针对每一人脸部位,还可以有进一步的详细划分,例如眉部还可以划分为眉尖部位、眉心部位、及眉峰部位。
在一种可能的实施方式中,为了使得拟合系数可以更准确的表征第一参考人脸的稠密点云数据拟合原始人脸的稠密点云数据的情况,可以对部分人脸部位对应的拟合系数做出调整,以使得基于拟合系数和多个对应的第一参考人脸的稠密点云数据拟合得到的稠密点云数据与目标图像对应的原始人脸的稠密点云数据相近。此时,需要调整拟合系数对应的部分人脸部位即为目标人脸模型部位,例如可以包括眼部、嘴部。具体的目标人脸模型部位可以根据具体情况或者经验确定,在此不再赘述。
在确定第一参考人脸的稠密点云数据表征的第一参考人脸中目标人脸模型目标对应的第一类稠密点云数据的情况下,以一组第一参考人脸的稠密点云数据
Figure PCTCN2021108629-appb-000008
为例,人脸部位例如可以划分为R个,此时对应于R个人脸部位的稠密点云数据例如可以表示为
Figure PCTCN2021108629-appb-000009
在R个人脸部位中的目标人脸部位包括眼部、嘴部的情况下,目标人脸模型部位对应的稠密点云数据例如包括
Figure PCTCN2021108629-appb-000010
Figure PCTCN2021108629-appb-000011
也即第一类稠密点云数据。
S702:对第一参考人脸的稠密点云数据中第一类稠密点云数据对应的中间系数进行调整,得到第一拟合系数。
在确定第一类稠密点云数据的情况下,可以确定目标人脸模型对应的部分部位对应的第一拟合系数,也即需要做出调整以达到更好拟合效果的部分中间系数。同时,对于中间系数中的任一拟合值α i,由于α i是利用IN mesh
Figure PCTCN2021108629-appb-000012
进行最小二乘处理得到的,因此α i中也包含有多个分别与多个目标人脸模型对应的部分部位的中间系数。
示例性的,在目标人脸模型部位分别为眼部、嘴部的情况下,α i中对应于多个目标人脸模型部位的中间系数例如可以表示为α i-1和α i-2,也即为第一类稠密点云数据
Figure PCTCN2021108629-appb-000013
Figure PCTCN2021108629-appb-000014
对应的中间系数。对中间系数α i-1和α i-2进行数值调整,可以使得基于调整后得到的第一拟合系数在对原始人脸的稠密点云数据进行拟合后,拟合结果与原始人脸的稠密点云数据更加相似。其中,数值调整例如包括数值增加和/或数值减少。
S703:将第一参考人脸的稠密点云数据中第二类稠密点云数据对应的中间系数确定为第二拟合系数;第二类稠密点云数据为第一参考人脸的稠密点云数据中除第一类稠密点云数据以外的稠密点云数据。
此时,可以将第一参考人脸的稠密点云数据中第二类稠密点云数据对应的中间系数确定为第二拟合系数。并且,还可以将每组第一参考人脸的稠密点云数据中除第一类稠密点云数据外的稠密点云数据作为第二类稠密点云数据。由于第二类稠密点云数据对应的拟合系数对拟合结果的影响较小,或者,在拟合时拟合结果较优,因此可以不对第二类稠密点云数据对应的拟合系数做出调整,以在保证拟合效果的情况下提高效率。
S704:基于第一拟合系数和第二拟合系数,得到每组第一参考人脸的稠密点云数据的拟合系数。
由于第一拟合系数对应目标人脸模型部位,第二拟合系数对应多个人脸部位中除目标人脸模型部位的其他人脸部位,因此将第一拟合系数、以及第二拟合系数进行组合可以确定对应于多个人脸部位的拟合系数,也即每组第一参考人脸的稠密点云数据的拟合系数。
针对上述步骤S103:
预设风格例如可以为卡通风格、古代风格或抽象风格等,具体可以根据实际的需要进行设定。示例性的,针对预设风格为卡通风格的情况,具有预设风格的第二参考人脸可以为卡通人脸。
在利用第一参考人脸生成具有预设风格的第二参考人脸的稠密点云数据时,例如可以采用下述方式:对参考图像中第一参考人脸的稠密点云数据进行调整,得到具有预设风格的第二参考人脸的稠密点云数据;或者,基于参考图像中的第一参考人脸,生成包括具有预设风格的第二参考人脸的虚拟人脸图像,并利用预先训练的神经网络生成虚拟人脸图像中第二参考人脸的稠密点云数据。
在对第一参考人脸稠密点云进行调整得到具有预设风格的第二参考人脸的稠密点云数据的情况下,示例性的,可以依据预设的风格对第一参考人脸的稠密点云数据中的全部稠密点云数据或者部分稠密点云数据做出调整,以使得到的第二参考人脸的稠密点云数据反应出的人脸具有预设的风格。
以预设风格为卡通风格为例,在一种可能的实施方式中,卡通风格例如包括对眼部有放大,则在对第一参考人脸的稠密点云数据进行调整时,对眼部对应的稠密点云数据做出上眼睑部位对应的稠密点云的位置坐标上移和/或下眼睑部位对应的稠密点云的位置坐标下移的调整,以使得到的第二参考人脸的稠密点云数据中眼部对应的稠密点云数据体现出眼部被放大。
在基于第一参考人脸生成虚拟人脸图像,并利用预先训练得到的神经网络生成第二参考人脸的稠密点云数据的情况下,示例性的,可以依据预设风格对参考图像中的第一参考人脸进行图形图像处理,以生成具有预设风格的第二参考人脸的虚拟人脸图像。其中,图形图像处理例如可以包括图片编辑、图片绘图、图片设计。
在获得具有预设风格的第二参考人脸的虚拟人脸图像的情况下,可以利用预先训练得到的神经网络确定对应的第二参考人脸的稠密点云数据。其中,利用预先训练得到的神经网络确定第二参考人脸的稠密点云数据的方法,与上述利用预先训练的神经网络确定原始人脸的稠密点云数据、及第一参考人脸的稠密点云数据的方法相似,在此不再赘述。
此时,例如可以将得到的第二参考人脸的稠密点云数据表示为CART mesh
在确定了具有预设风格的第二参考人脸的稠密点云数据、及多组第一参考人脸的稠密点云数据分别对应的拟合系数后,即可以确定目标人脸模型的稠密点云数据。
参见图8所示,本公开实施例还提供了一种确定目标人脸模型的稠密点云数据的具体方法,包括以下步骤S801至S802。
S801:基于多组第二参考人脸的稠密点云数据,生成多组第二参考人脸的稠密点云数据的均值数据。
在一种可能的实施方式中,可以基于第二参考人脸的稠密点云数据CART mesh中对应部位的坐标值进行求均值处理,以生成第二参考人脸的稠密点云数据的均值数据。其中,均值数据用于表征多组第二参考人脸的稠密点云数据的平均特征。示例性的,在第 二参考人脸的稠密点云数据中包含N组人脸稠密点云的情况下,可以将对于第二参考人脸的稠密点云数据中的多组人脸的稠密点云数据表示为
Figure PCTCN2021108629-appb-000015
Figure PCTCN2021108629-appb-000016
其中,针对任一组第二参考人脸的稠密点云数据
Figure PCTCN2021108629-appb-000017
可以确定对应于不同部位位置的多个人脸稠密点云,在包括W个不同的人脸稠密点云的情况下,例如可以将W个不同的人脸稠密点云表示为P 1、P 2、……、P W,分别对应的坐标值表示为
Figure PCTCN2021108629-appb-000018
利用每一组
Figure PCTCN2021108629-appb-000019
中对应部位位置的坐标值求取均值,可以得到对应部位位置的均值数据。此时,针对不同部位位置的第一个人脸稠密点云P 1,例如可以采用下述公式(1)得到均值数据
Figure PCTCN2021108629-appb-000020
Figure PCTCN2021108629-appb-000021
其他不同部位位置的人脸稠密点云确定均值数据的方式与上述对第一个人脸稠密点云求均值数据的方式相似,在此不再赘述。
此时,即可得到不同部位位置的均值数据
Figure PCTCN2021108629-appb-000022
也即第二参考人脸的稠密点云数据的均值数据,可以表示为
Figure PCTCN2021108629-appb-000023
S802:基于多组第二参考人脸的稠密点云数据、均值数据、以及多组第一参考人脸的稠密点云数据分别对应的拟合系数,生成目标人脸模型的稠密点云数据。
具体地,参见图9所示,本公开实施例还提供了一种生成目标人脸模型的稠密点云数据的具体方法,包括:
S901:基于多组第二参考人脸的稠密点云数据中每组第二参考人脸的稠密点云数据、以及均值数据,确定每组第二参考人脸的稠密点云数据的差值数据。
示例性的,可以将每组第二参考人脸的稠密点云数据与均值数据中对应不同部位位置的坐标值做差值,确定每组第二参考人脸的稠密点云数据的差值数据,表示为Δ mesh,也即Δ mesh=CART mesh-MEAN mesh。其中,差值数据可以表征每组第二参考人脸的稠密点云数据分别与多组第二参考人脸的稠密点云数据平均特征的差异度。
其中,由于存在多个第二参考人脸的稠密点云数据,因此差值数据Δ mesh中包含分别对应于多个第二参考人脸的稠密点云数据的多个子差值数据。示例性的,在存在N组第二参考人脸的稠密点云数据的情况下,对应的差值数据Δ mesh例如可以包括与N组第二参考人脸的稠密点云数据中包含的
Figure PCTCN2021108629-appb-000024
分别对应的子差值数据
Figure PCTCN2021108629-appb-000025
也即差值数据Δmesh中包含N个子差值数据,可以表示为
Figure PCTCN2021108629-appb-000026
S902:基于多组第一参考人脸的稠密点云数据分别对应的拟合系数,对多组第二参考人脸的稠密点云数据分别对应的差值数据进行插值处理。
示例性的,可以将拟合系数作为多组第二参考人脸的稠密点云数据对应的权重,对多组第二参考人脸的稠密点云数据分别对应的差值数据进行加权求和处理,实现插值 处理的过程。
在拟合系数包括Alpha=[α 12,…,α N]的情况下,与多张参考图像分别对应的第二参考人脸的第二参考人脸的稠密点云数据的权值即为α 1、α 2、……、α N。利用拟合系数对多组第二参考人脸的稠密点云数据分别对应的差值数据进行加权求和,得到的结果可以表示为AIM mesh,用于将原始人脸的稠密点云与第一参考人脸的稠密点云之间的关联关系转移至原始人脸的稠密点云与第二参考人脸的稠密点云之间,以使得到的人脸稠密点云即具有原始人脸的稠密点云的特征,又具有第二参考人脸的稠密点云的风格;在该种情况下,对多组第二参考人脸的稠密点云数据分别对应的差值数据进行加权求和得到的结果AIM mesh满足下述公式(2):
Figure PCTCN2021108629-appb-000027
其中,λ i,i∈[1,N]表示与第二参考人脸的稠密点云数据的权值α 1、α 2、……、α N和/或多组第二参考人脸的稠密点云数据分别对应的差值数据
Figure PCTCN2021108629-appb-000028
分别对应的权值,用于表征不同差值数据对加权求和所得到的结果的贡献度和/或重要性,可以预先设定或调整,具体的设定及调整方法在此不再赘述。
S903:基于插值处理的结果以及均值数据,生成目标人脸模型的稠密点云数据。
在一种可能的实施方式中,可以直接将加权和的结果叠加至均值数据以生成目标人脸模型的稠密点云数据,并表示为OUT mesh,也即可以利用下述公式(3)得到目标人脸模型的稠密点云数据OUT mesh
OUT mesh=AIM mesh+MEAN mesh   (3)。
此时,即可以得到既包含目标图像中的人脸特征,又包含第二参考人脸的稠密点云数据反应出的预设风格的目标人脸模型的稠密点云数据OUT mesh
针对上述步骤S104,利用目标人脸模型的稠密点云数据,可以生成与目标图像对应的目标人脸模型。
示例性的,可以基于目标人脸模型的稠密点云数据OUT mesh,利用渲染的方式生成对应的目标人脸模型,或者,确定与目标人脸模型的稠密点云数据OUT mesh具有关联关系的蒙皮,并利用目标人脸模型的稠密点云数据OUT mesh生成与目标图像对应的目标人脸模型。具体的方法可以根据实际情况选取,在此不再赘述。
本公开实施例还提供了一种利用本公开实施例提供的人脸重建方法,获取目标图像Pic A中的原始人脸A对应的目标虚拟人脸模型Mod Aim的具体过程的说明。
确定目标虚拟人脸模型Mod Aim的步骤包括下述(1)至(4)。
(1)准备素材;获取多张参考图像Pic cons
(2)稠密点云数据准备;其中,包括下述(2-1)以及(2-2)。
(2-1)利用神经网络确定目标图像Pic A中的原始人脸A对应的稠密点云数据IN mesh
(2-2)利用神经网络确定多张参考图像Pic cons分别对应的第一参考人脸的稠密点云数据BASE mesh
(3)确定拟合系数Alpha;其中,包括下述(3-1)至(3-4)。
(3-1)确定目标人脸模型的部位对应的第一类稠密点云数据
Figure PCTCN2021108629-appb-000029
以及
Figure PCTCN2021108629-appb-000030
(3-2)利用目标人脸模型对应的稠密点云数据拟合原始人脸的稠密点云数据,确定中间系数;
(3-3)对第一类稠密点云数据对应的中间系数中的α i-1以及α i-2调整,得到第一拟合系数;
(3-4)将第一参考人脸的稠密点云数据中除第一类稠密点云数据外的第二类稠密点云对应的中间系数,确定为第二拟合系数,并利用第一拟合系数、以及第二拟合系数确定拟合系数Alpha。
(4)确定目标人脸模型Mod Aim;其中,包括下述(4-1)至(4-3)。
(4-1)确定多组第二参考人脸的稠密点云数据的均值数据MEAN mesh
(4-2)确定目标人脸模型的稠密点云数据;基于多组第二参考人脸的稠密点云数据CART mesh、均值数据MEAN mesh、以及多组第一参考人脸的稠密点云数据分别对应的拟合系数Alpha,生成目标人脸模型的稠密点云数据AIM mesh
(4-3)利用生成的目标人脸模型的稠密点云数据AIM mesh,确定目标人脸模型Mod Aim
此处,值得注意的是,上述(1)至(4)仅是完成人脸重建方法一个具体示例,不对本公开实施例提供的人脸重建方法造成限定。
本领域技术人员可以理解,在具体实施方式的上述方法中,各步骤的撰写顺序并不意味着严格的执行顺序而对实施过程构成任何限定,各步骤的具体执行顺序应当以其功能和可能的内在逻辑确定。
基于同一发明构思,本公开实施例中还提供了与人脸重建方法对应的人脸重建装置,由于本公开实施例中的装置解决问题的原理与本公开实施例上述人脸重建方法相似,因此装置的实施可以参见方法的实施,重复之处不再赘述。
参照图10所示,为本公开实施例提供的一种人脸重建装置的示意图,所述装置包括:第一获取模块10、第一处理模块20、确定模块30、及生成模块40;其中,
第一获取模块10,用于获取目标图像中包括的原始人脸的稠密点云数据;
第一处理模块20,用于利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合所述原始人脸的稠密点云数据,得到多组所述第一参考人脸的稠密点云数据分别对应的拟合系数;
确定模块30,用于基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据;所述多组第二参考人脸分别是基于所述多张参考图像中的第一参考人脸生成的;
生成模块40,用于基于所述目标人脸模型的稠密点云数据,生成与所述目标图像的所述原始人脸对应的目标人脸模型。
一种可选的实施方式中,所述确定模块30在基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据时,用于:基于多组所述第二参考人脸的稠密点云数据,生成多组所述第二参考人脸的稠密点云数据的均值数据;基于多组所述第二参考人脸的稠密点云数据、所述均值数据、以及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,生成所述目标人脸模型的稠密点云数据。
一种可选的实施方式中,所述确定模块30在基于多组所述第二参考人脸的稠密点云数据、所述均值数据、以及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,生成所述目标人脸模型的稠密点云数据时,用于:基于多组所述第二参考人脸的稠密点云数据中每组所述第二参考人脸的稠密点云数据、以及所述均值数据,确定每组所述第二参考人脸的稠密点云数据的差值数据;基于多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,对多组所述第二参考人脸的稠密点云数据分别对应的差值数据进行插值处理;基于所述插值处理的结果以及所述均值数据,生成所述目标人脸模型的稠密点云数据。
一种可选的实施方式中,所述第一获取模块10在获取目标图像中包括的原始人脸的稠密点云数据时,用于:获取包括所述原始人脸的所述目标图像;利用预先训练的神经网络对所述目标图像进行处理,得到所述目标图像中所述原始人脸的稠密点云数据。
一种可选的实施方式中,所述装置还包括第二处理模块50,用于通过以下方式获取所述多张参考图像分别对应的第一参考人脸的稠密点云数据:获取包括第一参考人脸的多张参考图像;针对多张所述参考图像中的每张所述参考图像,利用预先训练的神经网络对每张所述参考图像进行处理,得到每张参考图像中的所述第一参考人脸的稠密点云数据。
一种可选的实施方式中,所述第一处理模块20在利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合所述原始人脸的稠密点云数据,得到多组所述第一参考人脸的稠密点云数据分别对应的拟合系数时,用于:对所述原始人脸的稠密点云数据以及所述第一参考人脸的稠密点云数据进行最小二乘处理,得到多组所述第一参考人脸的稠密点云数据分别对应的中间系数;基于每组第一参考人脸的稠密点云数据对应的中间系数,确定每组所述第一参考人脸的稠密点云数据对应的拟合系数。
一种可选的实施方式中,所述第一处理模块20在基于每组第一参考人脸的稠密点云数据对应的中间系数,确定每组所述第一参考人脸的稠密点云数据对应的拟合系数时,用于:从每组所述第一参考人脸的稠密点云数据中确定表征所述第一参考人脸中与目标人脸模型的部位对应的第一类稠密点云数据;对所述第一参考人脸的稠密点云数据中所述第一类稠密点云数据对应的中间系数进行调整,得到第一拟合系数;将所述第一参考人脸的稠密点云数据中第二类稠密点云数据对应的中间系数,确定为第二拟合系数;所述第二类稠密点云数据为所述第一参考人脸的稠密点云数据中除所述第一类稠密点云数据以外的稠密点云数据;基于所述第一拟合系数和所述第二拟合系数,得到每组所述第一参考人脸的稠密点云数据的拟合系数。
一种可选的实施方式中,所述装置还包括调整模块60,用于通过以下方式获取所述具有预设风格的第二参考人脸的稠密点云数据:对所述参考图像中第一参考人脸的稠密点云数据进行调整,得到所述具有预设风格的第二参考人脸的稠密点云数据;或者,基于所述参考图像中的第一参考人脸,生成包括具有所述预设风格的第二参考人脸的虚拟人脸图像;利用预先训练的神经网络生成所述虚拟人脸图像中所述第二参考人脸的稠 密点云数据。
一种可选的实施方式中,所述装置还包括训练模块70,在训练所述神经网络时,用于:获取样本图像集合;所述样本图像集合包括包含第一样本人脸的多张第一样本图像;所述多张第一样本图像划分为多个第一样本图像子集,每个第一样本图像子集中包括从多个预设采集角度分别采集得到的具有同种表情的第一样本人脸的图像;获取所述样本图像集合中的第一样本图像的第一样本人脸的稠密点云数据;利用初始神经网络,对所述样本图像集合中的第一样本图像进行特征学习,得到所述第一样本图像的第一样本人脸的预测稠密点云数据;利用所述第一样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
一种可选的实施方式中,所述样本图像集合还包括包含第二样本人脸和背景的多张第二样本图像,所述训练模块70还用于:获取每张所述第二样本图像的人脸关键点数据;利用第二样本图像的人脸关键点数据以及所述第二样本图像,拟合生成所述第二样本图像的第二样本人脸的稠密点云数据;利用所述神经网络,对所述样本图像集合中的第二样本图像进行特征学习,得到所述第二样本图像的第二样本人脸的预测稠密点云数据;利用所述第二样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
一种可选的实施方式中,所述样本图像集合中还包括:第三样本图像;所述第三样本图像为对所述第一样本图像进行数据增强处理得到;
所述训练模块70还用于:获取所述第三样本图像的第三样本人脸的稠密点云数据;利用所述初始神经网络对第三样本图像进行特征学习,得到所述第三样本图像的第三样本人脸的预测稠密点云数据;
利用所述第三样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
一种可选的实施方式中,所述数据增强处理包括下述至少一种:随机遮挡处理、高斯噪声处理、运动模糊处理、以及颜色区域通道改变处理。
关于装置中的各模块的处理流程、以及各模块之间的交互流程的描述可以参照上述方法实施例中的相关说明,这里不再详述。
如图11所示,本公开实施例还提供了一种计算机设备,包括处理器11和存储器12;处理器11用于执行存储器12中存储的机器可读指令,所述机器可读指令被处理器11执行时,处理器11执行下述步骤:
获取目标图像中包括的原始人脸的稠密点云数据;利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合原始人脸的稠密点云数据,得到多组第一参考人脸的稠密点云数据分别对应的拟合系数;基于具有预设风格的第二参考人脸的稠密点云数据、及多组第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据;第二参考人脸是基于参考图像中的第一参考人脸生成的;基于目标人脸模型的稠密点云数据,生成与目标图像的原始人脸对应的目标人脸模型。
上述存储器12包括内存121和外部存储器122;这里的内存121也称内存储器,用于暂时存放处理器11中的运算数据,以及与硬盘等外部存储器122交换的数据,处理器11通过内存121与外部存储器122进行数据交换。
上述指令的具体执行过程可以参考本公开实施例中所述的人脸重建方法的步骤,此处不再赘述。
本公开实施例还提供一种计算机可读存储介质,该计算机可读存储介质上存储有 计算机程序,该计算机程序被处理器运行时执行上述方法实施例中所述的人脸重建方法的步骤。其中,该存储介质可以是易失性或非易失的计算机可读取存储介质。
本公开实施例还提供一种计算机程序产品,该计算机程序产品承载有程序代码,所述程序代码包括的指令可用于执行上述方法实施例中所述的人脸重建方法的步骤,具体可参见上述方法实施例,在此不再赘述。
其中,上述计算机程序产品可以具体通过硬件、软件或其结合的方式实现。在一个可选实施例中,所述计算机程序产品具体体现为计算机存储介质,在另一个可选实施例中,计算机程序产品具体体现为软件产品,例如软件开发包(Software Development Kit,SDK)等等。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的系统和装置的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。在本公开所提供的几个实施例中,应该理解到,所揭露的系统、装置和方法,可以通过其它的方式实现。以上所描述的装置实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,又例如,多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些通信接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本公开各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。
所述功能如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个处理器可执行的非易失的计算机可读取存储介质中。基于这样的理解,本公开的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本公开各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
最后应说明的是:以上所述实施例,仅为本公开的具体实施方式,用以说明本公开的技术方案,而非对其限制,本公开的保护范围并不局限于此,尽管参照前述实施例对本公开进行了详细的说明,本领域的普通技术人员应当理解:任何熟悉本技术领域的技术人员在本公开揭露的技术范围内,其依然可以对前述实施例所记载的技术方案进行修改或可轻易想到变化,或者对其中部分技术特征进行等同替换;而这些修改、变化或者替换,并不使相应技术方案的本质脱离本公开实施例技术方案的精神和范围,都应涵盖在本公开的保护范围之内。因此,本公开的保护范围应所述以权利要求的保护范围为准。

Claims (15)

  1. 一种人脸重建方法,包括:
    获取目标图像中包括的原始人脸的稠密点云数据;
    利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合所述原始人脸的稠密点云数据,得到多组所述第一参考人脸的稠密点云数据分别对应的拟合系数;
    基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据;所述多组第二参考人脸分别是基于所述多张参考图像中的第一参考人脸生成的;
    基于所述目标人脸模型的稠密点云数据,生成与所述目标图像的所述原始人脸对应的目标人脸模型。
  2. 根据权利要求1所述的人脸重建方法,其特征在于,所述基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据,包括:
    基于多组所述第二参考人脸的稠密点云数据,生成多组所述第二参考人脸的稠密点云数据的均值数据;
    基于多组所述第二参考人脸的稠密点云数据、所述均值数据、以及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,生成所述目标人脸模型的稠密点云数据。
  3. 根据权利要求2所述的人脸重建方法,其特征在于,所述基于多组所述第二参考人脸的稠密点云数据、所述均值数据、以及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,生成所述目标人脸模型的稠密点云数据,包括:
    基于多组所述第二参考人脸的稠密点云数据中每组所述第二参考人脸的稠密点云数据、以及所述均值数据,确定每组所述第二参考人脸的稠密点云数据的差值数据;
    基于多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,对多组所述第二参考人脸的稠密点云数据分别对应的差值数据进行插值处理;
    基于所述插值处理的结果以及所述均值数据,生成所述目标人脸模型的稠密点云数据。
  4. 根据权利要求1至3任一所述的人脸重建方法,其特征在于,所述获取目标图像中包括的原始人脸的稠密点云数据,包括:
    获取包括所述原始人脸的所述目标图像;
    利用预先训练的神经网络对所述目标图像进行处理,得到所述目标图像中所述原始人脸的稠密点云数据。
  5. 根据权利要求1至4任一项所述的人脸重建方法,其特征在于,通过以下方式获取所述多张参考图像分别对应的第一参考人脸的稠密点云数据:
    获取包括第一参考人脸的多张参考图像;
    针对多张所述参考图像中的每张所述参考图像,利用预先训练的神经网络对每张所述参考图像进行处理,得到每张所述参考图像中的所述第一参考人脸的稠密点云数据。
  6. 根据权利要求1至5任一项所述的人脸重建方法,其特征在于,所述利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合所述原始人脸的稠密点云数据,得到多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,包括:
    对所述原始人脸的稠密点云数据以及所述第一参考人脸的稠密点云数据进行最小二乘处理,得到多组所述第一参考人脸的稠密点云数据分别对应的中间系数;
    基于每组第一参考人脸的稠密点云数据对应的中间系数,确定每组所述第一参考人脸的稠密点云数据对应的拟合系数。
  7. 根据权利要求6所述的人脸重建方法,其特征在于,所述基于每组第一参考人脸的稠密点云数据对应的中间系数,确定每组所述第一参考人脸的稠密点云数据对应的拟合系数,包括:
    从每组所述第一参考人脸的稠密点云数据中确定表征所述第一参考人脸中与目标人脸模型的部位对应的第一类稠密点云数据;
    对所述第一参考人脸的稠密点云数据中所述第一类稠密点云数据对应的中间系数进行调整,得到第一拟合系数;
    将所述第一参考人脸的稠密点云数据中第二类稠密点云数据对应的中间系数,确定为第二拟合系数;所述第二类稠密点云数据为所述第一参考人脸的稠密点云数据中除所述第一类稠密点云数据以外的稠密点云数据;
    基于所述第一拟合系数和所述第二拟合系数,得到每组所述第一参考人脸的稠密点云数据的拟合系数。
  8. 根据权利要求1至7任一项所述的人脸重建方法,其特征在于,通过以下方式获取所述具有预设风格的第二参考人脸的稠密点云数据:
    对所述参考图像中第一参考人脸的稠密点云数据进行调整,得到所述具有预设风格的第二参考人脸的稠密点云数据;或者,
    基于所述参考图像中的第一参考人脸,生成包括具有所述预设风格的第二参考人脸的虚拟人脸图像;利用预先训练的神经网络生成所述虚拟人脸图像中所述第二参考人脸的稠密点云数据。
  9. 根据权利要求4、5、和8中任一项所述的人脸重建方法,其特征在于,训练所述神经网络,包括:
    获取样本图像集合;所述样本图像集合包括包含第一样本人脸的多张第一样本图像;所述多张第一样本图像划分为多个第一样本图像子集,每个第一样本图像子集中包括从多个预设采集角度分别采集得到的具有同种表情的第一样本人脸的图像;
    获取所述样本图像集合中的第一样本图像的第一样本人脸的稠密点云数据;
    利用所述神经网络,对所述样本图像集合中的第一样本图像进行特征学习,得到所述第一样本图像的第一样本人脸的预测稠密点云数据;
    利用所述第一样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
  10. 根据权利要求9所述的人脸重建方法,其特征在于,所述样本图像集合还包括包含第二样本人脸和背景的多张第二样本图像,所述训练所述神经网络还包括:
    获取每张所述第二样本图像的人脸关键点数据;
    利用所述第二样本图像的人脸关键点数据以及所述第二样本图像,拟合生成所述第二样本图像的第二样本人脸的稠密点云数据;
    利用所述神经网络,对所述样本图像集合中的第二样本图像进行特征学习,得到所述第二样本图像的第二样本人脸的预测稠密点云数据;
    利用所述第二样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
  11. 根据权利要求9或10所述的人脸重建方法,其特征在于,所述样本图像集合中还包括第三样本图像;所述第三样本图像为对所述第一样本图像进行数据增强处理得到;所述训练所述神经网络还包括:
    获取所述第三样本图像的第三样本人脸的稠密点云数据;
    利用所述神经网络对所述第三样本图像进行特征学习,得到所述第三样本图像的第三样本人脸的预测稠密点云数据;
    利用所述第三样本人脸的稠密点云数据和预测稠密点云数据,对所述神经网络进行训练。
  12. 根据权利要求11所述的人脸重建方法,其特征在于,所述数据增强处理包括下述至少一种:随机遮挡处理、高斯噪声处理、运动模糊处理、以及颜色区域通道改变处理。
  13. 一种人脸重建装置,包括:
    第一获取模块,用于获取目标图像中包括的原始人脸的稠密点云数据;
    第一处理模块,用于利用多张参考图像分别对应的第一参考人脸的稠密点云数据拟合所述原始人脸的稠密点云数据,得到多组所述第一参考人脸的稠密点云数据分别对应的拟合系数;
    确定模块,用于基于具有预设风格的多组第二参考人脸的稠密点云数据、及多组所述第一参考人脸的稠密点云数据分别对应的拟合系数,确定目标人脸模型的稠密点云数据;所述多组第二参考人脸分别是基于所述多张参考图像中的第一参考人脸生成的;
    生成模块,用于基于所述目标人脸模型的稠密点云数据,生成与所述目标图像的所述原始人脸对应的目标人脸模型。
  14. 一种计算机设备,包括处理器和存储器,所述存储器存储有所述处理器可执行的机器可读指令,所述处理器用于执行所述存储器中存储的机器可读指令,所述机器可读指令被所述处理器执行时,所述处理器执行如权利要求1至12任一项所述的人脸重建方法的步骤。
  15. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被计算机设备运行时,所述计算机设备执行如权利要求1至12任一项所述的人脸重建方法的步骤。
PCT/CN2021/108629 2020-11-25 2021-07-27 人脸重建方法、装置、计算机设备及存储介质 Ceased WO2022110855A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
KR1020237018454A KR20230098313A (ko) 2020-11-25 2021-07-27 안면 재구성 방법, 장치, 컴퓨터 기기 및 저장 매체
JP2023531694A JP7525814B2 (ja) 2020-11-25 2021-07-27 顔再構成方法、装置、コンピュータ装置及び記憶媒体

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202011337942.0A CN112396692B (zh) 2020-11-25 2020-11-25 一种人脸重建方法、装置、计算机设备及存储介质
CN202011337942.0 2020-11-25

Publications (1)

Publication Number Publication Date
WO2022110855A1 true WO2022110855A1 (zh) 2022-06-02

Family

ID=74607056

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/108629 Ceased WO2022110855A1 (zh) 2020-11-25 2021-07-27 人脸重建方法、装置、计算机设备及存储介质

Country Status (5)

Country Link
JP (1) JP7525814B2 (zh)
KR (1) KR20230098313A (zh)
CN (1) CN112396692B (zh)
TW (1) TW202221645A (zh)
WO (1) WO2022110855A1 (zh)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112396692B (zh) * 2020-11-25 2023-11-28 北京市商汤科技开发有限公司 一种人脸重建方法、装置、计算机设备及存储介质
CN113343773B (zh) * 2021-05-12 2022-11-08 上海大学 基于浅层卷积神经网络的人脸表情识别系统
TWI911817B (zh) * 2024-07-31 2026-01-11 廣達電腦股份有限公司 用以訓練3d關鍵點偵測模型的系統及方法

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20150123967A1 (en) * 2013-11-01 2015-05-07 Microsoft Corporation Generating an avatar from real time image data
CN109087340A (zh) * 2018-06-04 2018-12-25 成都通甲优博科技有限责任公司 一种包含尺度信息的人脸三维重建方法及系统
CN109376698A (zh) * 2018-11-29 2019-02-22 北京市商汤科技开发有限公司 人脸建模方法和装置、电子设备、存储介质、产品
CN109636886A (zh) * 2018-12-19 2019-04-16 网易(杭州)网络有限公司 图像的处理方法、装置、存储介质和电子装置
CN112396692A (zh) * 2020-11-25 2021-02-23 北京市商汤科技开发有限公司 一种人脸重建方法、装置、计算机设备及存储介质
CN112396693A (zh) * 2020-11-25 2021-02-23 上海商汤智能科技有限公司 一种面部信息的处理方法、装置、电子设备及存储介质

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2007257324A (ja) * 2006-03-23 2007-10-04 Space Vision:Kk 顔モデル作成システム
JP6207210B2 (ja) 2013-04-17 2017-10-04 キヤノン株式会社 情報処理装置およびその方法
CN109285112A (zh) * 2018-09-25 2019-01-29 京东方科技集团股份有限公司 基于神经网络的图像处理方法、图像处理装置
CN110706339B (zh) * 2019-09-30 2022-12-06 北京市商汤科技开发有限公司 三维人脸重建方法及装置、电子设备和存储介质
CN110717977B (zh) * 2019-10-23 2023-09-26 网易(杭州)网络有限公司 游戏角色脸部处理的方法、装置、计算机设备及存储介质
CN111402399B (zh) * 2020-03-10 2024-03-05 广州虎牙科技有限公司 人脸驱动和直播方法、装置、电子设备及存储介质
CN111695471B (zh) * 2020-06-02 2023-06-27 北京百度网讯科技有限公司 虚拟形象生成方法、装置、设备以及存储介质
CN111784821B (zh) * 2020-06-30 2023-03-14 北京市商汤科技开发有限公司 三维模型生成方法、装置、计算机设备及存储介质

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20150123967A1 (en) * 2013-11-01 2015-05-07 Microsoft Corporation Generating an avatar from real time image data
CN109087340A (zh) * 2018-06-04 2018-12-25 成都通甲优博科技有限责任公司 一种包含尺度信息的人脸三维重建方法及系统
CN109376698A (zh) * 2018-11-29 2019-02-22 北京市商汤科技开发有限公司 人脸建模方法和装置、电子设备、存储介质、产品
CN109636886A (zh) * 2018-12-19 2019-04-16 网易(杭州)网络有限公司 图像的处理方法、装置、存储介质和电子装置
CN112396692A (zh) * 2020-11-25 2021-02-23 北京市商汤科技开发有限公司 一种人脸重建方法、装置、计算机设备及存储介质
CN112396693A (zh) * 2020-11-25 2021-02-23 上海商汤智能科技有限公司 一种面部信息的处理方法、装置、电子设备及存储介质

Also Published As

Publication number Publication date
JP2023551247A (ja) 2023-12-07
CN112396692B (zh) 2023-11-28
TW202221645A (zh) 2022-06-01
JP7525814B2 (ja) 2024-07-31
KR20230098313A (ko) 2023-07-03
CN112396692A (zh) 2021-02-23

Similar Documents

Publication Publication Date Title
TWI773458B (zh) 重建人臉的方法、裝置、電腦設備及存儲介質
CN114972632B (zh) 基于神经辐射场的图像处理方法及装置
CN110956691B (zh) 一种三维人脸重建方法、装置、设备及存储介质
WO2020192568A1 (zh) 人脸图像生成方法、装置、设备及存储介质
TWI778723B (zh) 重建人臉的方法、裝置、電腦設備及存儲介質
JP7525814B2 (ja) 顔再構成方法、装置、コンピュータ装置及び記憶媒体
WO2022231582A1 (en) Photo relighting and background replacement based on machine learning models
CN114202615B (zh) 人脸表情的重建方法、装置、设备和存储介质
CN114445562A (zh) 三维重建方法及装置、电子设备和存储介质
CN111127309B (zh) 肖像风格迁移模型训练方法、肖像风格迁移方法以及装置
CN111784821A (zh) 三维模型生成方法、装置、计算机设备及存储介质
CN113179421B (zh) 视频封面选择方法、装置、计算机设备和存储介质
CN109859857A (zh) 身份信息的标注方法、装置和计算机可读存储介质
CN113095206A (zh) 虚拟主播生成方法、装置和终端设备
CN115239857B (zh) 图像生成方法以及电子设备
WO2024055379A1 (zh) 基于角色化身模型的视频处理方法、系统及相关设备
CN115496844A (zh) 虚拟会议空间显示方法以及装置
CN115100707A (zh) 模型的训练方法、视频信息生成方法、设备以及存储介质
CN111275610B (zh) 一种人脸变老图像处理方法及系统
CN115861041B (zh) 图像风格迁移方法、装置、计算机设备、存储介质和产品
HK40039017A (zh) 一种人脸重建方法、装置、计算机设备及存储介质
CN111222448B (zh) 图像转换方法及相关产品
Raut¹ et al. Lighting for Mixed Reality Sessions
CN120410835A (zh) 图像处理方法、装置、可读存储介质和程序产品
WO2024112378A1 (en) Stylized animatable representation

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21896362

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2023531694

Country of ref document: JP

ENP Entry into the national phase

Ref document number: 20237018454

Country of ref document: KR

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 21896362

Country of ref document: EP

Kind code of ref document: A1

WWR Wipo information: refused in national office

Ref document number: 1020237018454

Country of ref document: KR