WO2025007933A1 - 人脸活化生成模型训练和人脸活化 - Google Patents
人脸活化生成模型训练和人脸活化 Download PDFInfo
- Publication number
- WO2025007933A1 WO2025007933A1 PCT/CN2024/103675 CN2024103675W WO2025007933A1 WO 2025007933 A1 WO2025007933 A1 WO 2025007933A1 CN 2024103675 W CN2024103675 W CN 2024103675W WO 2025007933 A1 WO2025007933 A1 WO 2025007933A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- driving
- facial
- face
- activation
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/774—Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
- G06N3/0455—Auto-encoder networks; Encoder-decoder networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0475—Generative networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/094—Adversarial learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/776—Validation; Performance evaluation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/161—Detection; Localisation; Normalisation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/168—Feature extraction; Face representation
- G06V40/171—Local features and components; Facial parts ; Occluding parts, e.g. glasses; Geometrical relationships
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
Definitions
- the embodiments of this specification relate to an image processing method, and more particularly to a face activation generation model training method.
- the attack method In the IIFAA biosecurity attack detection, it is necessary to intelligently generate sufficiently realistic facial attack materials to conduct security evaluations on the identity verification function of mobile phones to ensure that the face security evaluation is complete and effective.
- the attack method must first be universal and have a relatively stable attack power in different mobile phone black box tests; secondly, it must have a high attack rate and strong aggressiveness, so as to be able to distinguish the strength of the anti-attack ability of the face recognition algorithm and effectively widen the gap in the anti-attack ability of the identity verification function of different mobile phone manufacturers to be tested; in addition, the attack method needs to be generated intelligently through an algorithm, which is different from physical generation such as artificial masks or ordinary PS methods, so that the potential errors found in the model used by the face recognition algorithm can be traced, which is convenient for improving the model, thereby improving the reliability of the model and the interpretability of the model prediction results.
- One of the purposes of the embodiments of this specification is to provide a face activation generation model training method, which can transfer specified actions to specified face images, retain the identity information of the specified face to a large extent, and effectively test the anti-attack ability of face recognition algorithms.
- the embodiment of this specification proposes a method for training a face activation generation model, the method comprising: obtaining a source image and a driving image; inputting the source image and the driving image into the face activation generation model, respectively obtaining face action information of the source image and face action information of the driving image; performing feature extraction on the source image encoding to obtain source image features; projecting the facial action information of the driving image onto the facial action information of the source image in three-dimensional space to transfer the action of the driving image to the source image; generating a target activation image corresponding to the source image according to the facial action information of the driving image, the facial action information of the source image and the source image features; training the face activation generation model with minimizing the difference between the target activation image and the driving image as the training goal.
- the motion migration from the driving image to the source image is realized by projecting the facial motion information in the three-dimensional space, and the three-dimensional position information is associated with the facial texture information, which greatly improves the problem of motion distortion, so that the facial identity information in the source image can be more accurately and completely retained; by minimizing the difference between the target activation image and the driving image, the realism of the target activation image is greatly improved.
- the target activation image generated by the facial activation generation model trained by this method can more effectively test the anti-attack ability of the face recognition algorithm.
- the facial motion information includes facial three-dimensional posture information and facial key point information.
- projecting the facial motion information of the driving image onto the facial motion information of the source image in three-dimensional space so as to migrate the motion of the driving image to the source image specifically includes: correcting the facial key point information of the source image and the facial key point information of the driving image to the same posture by comparing the three-dimensional facial posture information of the source image and the three-dimensional facial posture information of the driving image to obtain source key points and driving key points in three-dimensional space; projecting the driving key points onto the source key points to migrate the motion of the driving image to the source image.
- the three-dimensional facial posture information includes a head rotation angle.
- the face activation generation model is implemented by a generative adversarial network, which includes a discriminator and a generator; taking minimizing the difference between the target activation image and the driving image as a training goal, training the face activation generation model specifically includes: inputting the facial action information of the driving image, the facial action information of the source image and the source image features into the generator, and outputting the target activation image corresponding to the source image; using the discriminator to determine whether the target activation image is the driving image; calculating the generation loss based on the target activation image and the driving image; calculating the discrimination loss based on the discrimination result of the discriminator; and training the generative adversarial network with minimizing the discrimination loss and the generation loss as training goals, respectively.
- a generative adversarial network which includes a discriminator and a generator; taking minimizing the difference between the target activation image and the driving image as a training goal, training the face activation generation model specifically includes: inputting the facial action information of the driving image, the facial action information of the source image and
- calculating and generating a loss based on the target activation image and the drive image specifically includes: calculating a first pixel-based loss between the target activation image and the drive image; calculating A second loss based on a feature map between the target activation image and the driving image; and weighted fusion of the first loss and the second loss as the generated loss.
- the facial key point information of the source image and the facial key point information of the driving image are corrected to the same posture, and the source key points and driving key points in the three-dimensional space are obtained.
- it includes: by comparing the three-dimensional facial posture information of the source image and the three-dimensional facial posture information of the driving image, the facial key point information of the source image and the facial key point information of the driving image are rotated and/or offset corrected to the same posture, and the source key points and driving key points in the three-dimensional space are obtained.
- the facial key point information of the source image and the facial key point information of the driving image are rotated and/or offset-corrected to the same posture, so as to obtain the source key points and driving key points in the three-dimensional space.
- it includes: by comparing the three-dimensional facial posture information of the source image with the three-dimensional facial posture information of the driving image, the facial key point information of the source image and the facial key point information of the driving image are rotated and/or offset-corrected to the same posture using matrix changes, so as to obtain the source key points and driving key points in the three-dimensional space.
- Another purpose of the embodiments of this specification is to provide a face activation method, which can transfer specified actions to specified face images, generate realistic face activation images, and retain the identity information of the specified face to a large extent, thereby effectively testing the anti-attack ability of face recognition algorithms.
- an embodiment of the present specification provides a face activation method, which includes: obtaining a source image to be activated and a driving video containing a target action, and extracting a plurality of frames of driving images from the driving video; for each frame of driving image obtained, inputting the source image and the driving image into the face activation generation model to generate a target activation image corresponding to the source image, wherein the face activation generation model is trained using the steps described in any of the above methods; and connecting all the obtained target activation images to obtain a target activation video.
- Another purpose of the embodiments of this specification is to provide a face activation generation model training device, which can transfer specified actions to specified face images, retain the identity information of the specified face to a large extent, and can effectively test the anti-attack ability of face recognition algorithms.
- the embodiment of this specification provides a face activation generation model training device, the device comprising: a sample acquisition module for acquiring a source image and a driving image; an information extraction module for inputting the source image and the driving image into the face activation generation model to respectively acquire the face action information of the source image and the face action information of the driving image; feature encoding the source image to obtain the source image feature; a generation module for generating a face activation generation model based on three The facial action information of the driving image is projected onto the facial action information of the source image in a three-dimensional space so that the action of the driving image is transferred to the source image; a target activation image corresponding to the source image is generated according to the facial action information of the driving image, the facial action information of the source image and the source image features; a training module is used to train the face activation generation model with minimizing the difference between the target activation image and the driving image as the training goal.
- the facial motion information includes facial three-dimensional posture information and facial key point information.
- the generation module corrects the facial key point information of the source image and the facial key point information of the driving image to the same posture by comparing the three-dimensional facial posture information of the source image and the three-dimensional facial posture information of the driving image, thereby obtaining the source key points and driving key points in three-dimensional space; and projects the driving key points onto the source key points to migrate the movements of the driving image to the source image.
- the three-dimensional facial posture information includes a head rotation angle.
- the face activation generation model is implemented by a generative adversarial network, which includes a discriminator and a generator;
- the training module specifically includes: inputting the facial action information of the driving image, the facial action information of the source image and the source image features into the generator, and outputting a target activation image corresponding to the source image; using the discriminator to determine whether the target activation image is the driving image; calculating the generation loss based on the target activation image and the driving image; calculating the discrimination loss based on the discrimination result of the discriminator; and training the generative adversarial network with minimization of the discrimination loss and the generation loss as training objectives respectively.
- the training module calculates a first pixel-based loss between the target activation image and the driving image; calculates a second feature map-based loss between the target activation image and the driving image; and weightedly fuses the first loss and the second loss as the generated loss.
- the generation module compares the three-dimensional facial posture information of the source image and the three-dimensional facial posture information of the driving image, rotates and/or offsets the facial key point information of the source image and the facial key point information of the driving image to the same posture, and obtains the source key points and driving key points in three-dimensional space.
- the generation module compares the 3D facial posture information of the source image and the 3D facial posture information of the driving image, and uses matrix changes to rotate or/and offset the facial key point information of the source image and the facial key point information of the driving image to the same posture, thereby obtaining the source key point information in the 3D space. Key points and drive keys.
- Another purpose of the embodiments of this specification is to provide a face activation device that can transfer specified actions to specified face images, generate realistic face activation images, and retain the identity information of the specified face to a large extent, thereby effectively testing the anti-attack ability of face recognition algorithms.
- an embodiment of the present specification provides a face activation device, which includes: a sample acquisition module, which is used to obtain a source image to be activated and a driving video containing a target action, and extract a plurality of frames of driving images from the driving video; an activation generation module, which is used to input the source image and the driving image into the face activation generation model for each frame of the driving image obtained, and generate a target activation image corresponding to the source image, wherein the face activation generation model is trained using the steps described in any of the above methods; and a connection module, which is used to connect all the obtained target activation images to obtain a target activation video.
- Another purpose of an embodiment of the present specification is to provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps described in any of the above methods when executing the program.
- the face activation generation model training and face activation method described in the embodiments of this specification have the following beneficial effects.
- Motion transfer is achieved by projecting facial motion information from the driving image to the source image in three-dimensional space, and the three-dimensional position information is associated with the facial texture information. This can greatly improve the motion distortion problem caused by head rotation, so that the facial identity information in the source image can be retained more accurately and completely; by minimizing the difference between the target activation image and the driving image, the realism of the target activation image is greatly improved, and the generated target activation image can more effectively test the anti-attack ability of the face recognition algorithm.
- Key point alignment correction based on the three-dimensional posture information of the face can project the original 2D key point information into the 3D three-dimensional posture information of the face, and transform the face activation into pseudo-3D activation, so that the face key point information and the face texture features are better integrated to generate a more realistic activation image.
- the face activation generation model training and face activation device described in the embodiments of this specification also have the above-mentioned beneficial effects.
- FIG1 exemplarily shows a flowchart of a face activation generation model training method according to an embodiment of this specification in one implementation manner.
- FIG. 2 exemplarily shows a method for training a face activation generation model according to an embodiment of the present specification in a specific embodiment. Steps performed under the implementation method.
- FIG3 exemplarily shows a flowchart of a face activation method according to an embodiment of this specification in one implementation manner.
- FIG. 4 exemplarily shows a schematic diagram of the structure of a face activation generation model training device according to an embodiment of this specification in one implementation manner.
- FIG. 5 exemplarily shows a schematic structural diagram of a face activation device according to an embodiment of this specification in one implementation manner.
- the process of face activation is manifested as the transfer of the movements or expressions of the face in image B to the face in image A. Since expression changes and expression transfers can be manifested as movement changes and movement transfers of the facial features, the movements in the embodiments of this specification include not only movements such as shaking the head and nodding, but also expressions such as smiling, frowning, and closing eyes.
- FIG1 exemplarily shows a flow chart of a method for training a face activation generation model in one embodiment of the present specification.
- the process includes the following steps.
- the source images and driving images used need to contain significantly different facial movements in order to more intuitively observe the effect of movement transfer.
- the source images and driving images can be extracted from an existing image set containing different facial movements, or from images of different people taken by a camera or other acquisition device, or from several videos, where each video records the facial movement changes of the same person.
- each time a video is obtained as a training sample two frames of images with different facial movements are extracted from the video as source image A and driving image B, respectively.
- the purpose of face activation is to transfer the movement in driving image B to the face in source image A.
- the face activation generation model is trained using the face of the same person as sample data, so that the model can be fully trained and thus more easily converged.
- a general face action detection method can be used to obtain the face action information of the source image and the face action information of the driving image respectively.
- Face action information can show the movement of the head and the position change of the facial features in three-dimensional space, which has stronger fidelity and representation than the action information in two-dimensional space.
- the facial motion information includes facial three-dimensional posture information and facial key point information.
- the facial action detection method includes a facial 3D posture detection method and a facial key point detection method.
- the facial 3D posture information can reflect the posture of the head in the 3D space to determine the orientation of the face; while the facial key point information records the position information by adding key points to the key facial features such as eyes, nose, mouth or other key facial areas, thereby expressing facial actions according to the position changes of these key areas.
- the three-dimensional facial posture information includes the head rotation angle.
- the three-dimensional posture information of the face can be represented by the 3D position information of the face, which is used to describe the rotation angle of the head on three mutually perpendicular coordinate axes in the three-dimensional space.
- the face position can reflect the direction of the head and express the head position information.
- These three mutually perpendicular coordinate axes are constructed based on the directions of nodding, shaking and swinging the head to construct the XYZ axis centered on the head, which can reflect the rotation angle of the head on the three coordinate axes respectively, and are summarized as the 3D position information of the face.
- the source image is input into a general image encoder for feature encoding, and the corresponding source image features are output so that the source image can be further processed such as action migration based on the source image features, and the target activation image is generated in cooperation with the image decoder.
- projecting the facial motion information of the driving image onto the facial motion information of the source image in the three-dimensional space so as to transfer the motion of the driving image to the source image specifically comprises: comparing the facial motion information of the source image with the three-dimensional image of the source image; The facial key point information of the source image and the facial key point information of the driving image are corrected to the same posture to obtain the source key points and the driving key points in the three-dimensional space; the driving key points are projected onto the source key points to transfer the actions of the driving image to the source image.
- the head posture of the source image and the head posture of the driving image can be obtained based on this.
- the key point information of the face in the image corresponds to the head posture.
- the key point correction is performed based on the difference in head posture between the two, so that the key points in the source image and the face behind the key points in the driving image are in the same posture, and the source key points and the driving key points are obtained, so as to further compare the key points with the same position and the key points with different positions.
- the key points with different positions represent the changes in the face action and are the objects that need to be focused on.
- the source key points are adjusted according to the key points with different positions, so as to obtain the face action information that needs to be migrated in the source image, and complete the action migration of the driving image.
- Unifying the head posture of the source image with the head posture of the driving image is conducive to adjusting the face to a posture that is more convenient for action migration, obtaining a better face activation effect, and avoiding the influence of different head postures on the activation results.
- the key points corresponding to the head posture of the source image can be fixed, so that the head posture of the source image is used as the target posture, and the head posture of the driving image is made consistent with the target posture by correcting the key points of the driving image;
- the key points corresponding to the head posture of the driving image can be fixed, so that the head posture of the driving image is used as the target posture, and the head posture of the source image is made consistent with the target posture by correcting the key points of the source image; or the target posture that can better show the face activation effect can be pre-set, and then the head posture of the source image and the head posture of the driving image can be corrected to the target posture.
- the facial key point information of the source image and the facial key point information of the driving image are corrected to the same posture, so as to obtain the source key points and driving key points in the three-dimensional space.
- it includes: by comparing the three-dimensional facial posture information of the source image with the three-dimensional facial posture information of the driving image, the facial key point information of the source image and the facial key point information of the driving image are rotated and/or offset corrected to the same posture, so as to obtain the source key points and driving key points in the three-dimensional space.
- the facial key points can be rotated to the key point positions corresponding to the target posture, or the facial key points can be offset.
- the facial key points can be first rotated to the approximate key point positions corresponding to the target posture, and then the facial key point positions can be further adjusted through offset.
- the three-dimensional posture information is used to rotate and/or offset the facial key point information of the source image and the facial key point information of the driving image to the same posture, so as to obtain the source key point and the driving key point in the three-dimensional space.
- the source key point and the driving key point are obtained by comparing the three-dimensional facial posture information of the source image and the three-dimensional facial posture information of the driving image, and using matrix changes to rotate and/or offset the facial key point information of the source image and the driving image to the same posture, so as to obtain the source key point and the driving key point in the three-dimensional space.
- a target posture for key point correction is selected, and the target key points corresponding to the target posture are obtained through a general key point detection method, and then the affine transformation matrices for adjusting the facial key point information of the source image and the facial key point information of the driving image to the target key points are respectively determined, and the facial key point information of the source image and the facial key point information of the driving image are respectively corrected using the corresponding affine transformation matrices to adjust them to the target posture.
- a general image decoder is used to combine the source image features with the facial motion information obtained by motion migration to generate a target activation image.
- the image decoder can be constructed based on the image decoder in the GAN generation network.
- the face activation generative model is trained with the goal of minimizing the difference between the target activation image and the driving image.
- the difference between the images can be obtained by calculating the difference in pixel-by-pixel values in the images.
- the images to be compared can be further processed into images of other levels, and then the pixel-by-pixel value differences are calculated, such as extracting feature maps of the images to be compared.
- calculating the difference between the target activation image and the driving image through multiple approaches and then using the obtained multiple loss training models is conducive to improving the training efficiency of the face activation generation model and accelerating the improvement of the model performance.
- a face activation generation model is implemented by a generative adversarial network, which includes a discriminator and a generator; minimizing the difference between a target activation image and a driving image is used as a training goal, and training the face activation generation model specifically includes: inputting facial motion information of the driving image, facial motion information of the source image, and source image features into the generator, and outputting a target activation image corresponding to the source image; using the discriminator to determine whether the target activation image is a driving image; calculating the generation loss based on the target activation image and the driving image; calculating the discrimination loss based on the discrimination result of the discriminator; and training the generative adversarial network with minimizing the discrimination loss and the generation loss as training goals, respectively.
- the generative adversarial network consists of a discriminator and a generator.
- the discriminator is first trained to have good discrimination ability; then the generator is trained so that the results generated by the generator can deceive the discriminator.
- a discriminator with good ability has stronger fidelity.
- the discriminator after obtaining the target activation image, it is input into the discriminator to determine whether it is a real sample image or a tampered forged image, and obtain a discrimination result; then the discrimination loss is calculated according to the accuracy of the discrimination result, and the discrimination ability of the discriminator is trained by minimizing the discrimination loss.
- the source image and the driving image in the training samples used are the same person.
- face activation the facial movements in the driving image are transferred to the source image and merged with the face in the source image.
- the obtained target activation image should be consistent with the driving image in an ideal situation. Therefore, the generation loss is determined by calculating the difference between the driving image and the target activation image. By minimizing the generation loss, the generation effect of the generator is optimized, so that the driving image and the source image are merged more naturally.
- the generator is continuously optimized, the generated target activation images become more and more realistic, making it increasingly difficult for the discriminator to determine their authenticity.
- the discriminator can no longer determine the authenticity of the target activation images, the generative adversarial training of the face activation generation model is completed.
- the face activation generation model In the process of iterative training by continuously calculating losses, the face activation generation model also learns and transfers the facial texture features in the source image and the driving image. That is, the face activation process associates and fuses the facial motion information representing the three-dimensional position of the face with the facial texture features, realizes face activation in three-dimensional space, and improves the problem of facial motion distortion after motion transfer.
- a GAN network model is used to construct a generative adversarial network to implement a face activation generation model.
- calculating the generation loss based on the target activation image and the driving image specifically includes: calculating a first loss based on pixels between the target activation image and the driving image; calculating a second loss based on feature maps between the target activation image and the driving image; and weightedly fusing the first loss and the second loss as the generation loss.
- the first loss can be obtained by calculating the pixel-by-pixel mean square error loss between the target activation image and the driving image.
- the second loss is obtained by calculating the face similarity between the target activation image and the driving image.
- the face similarity is calculated by combining Arcface (also known as additive angular margin loss function) with cross entropy. Training the face activation generation model by Arcface can effectively increase the inter-class distance while ensuring the inter-class distance, while achieving the effect of reducing the intra-class distance, so that the face activation generation model can generate more accurate activation results more efficiently.
- Arcface also known as additive angular margin loss function
- FIG. 2 exemplarily shows the steps performed in a specific implementation of the face activation generation model training method described in an embodiment of this specification.
- two frames of images are extracted from a video of the same person as the source image and the driving image, respectively, and the facial motion information of the source image and the driving image are extracted respectively.
- the facial key points of the source image and the driving image are extracted respectively by a general facial key point detection method.
- the faces in the source image and the driving image are the same, only the facial key points in the source image can be extracted and copied to the driving image; then the 3D facial pose of the source image and the 3D facial pose of the driving image are extracted respectively by a general facial posture detection method, and the head posture difference of the faces in the source image and the driving image is compared based on the 3D facial posture, and the facial key points are corrected by using matrix changes, so that the head posture in the source image is consistent with the head posture in the driving image, and the corrected source key points and driving key points are obtained.
- the image features in the source image are extracted, and the source image features, source key points, and driving key points are input into the generator built based on the image decoder in the GAN generation network.
- the action is transferred by projecting the driving key points onto the source key points, and the final target activation image is generated in combination with the source image features.
- the face activation generation model is iteratively trained based on the minimization of the difference between the target activation image and the driving image as the training goal, and a face activation generation model without action distortion in the generated activation image is obtained.
- the face activation generation model training method realizes motion transfer by projecting the face motion information from the driving image to the source image in three-dimensional space, which can well improve the motion distortion problem caused by head rotation; by minimizing the difference between the target activation image and the driving image, the realism of the target activation image is greatly improved, so that the face identity information in the source image can be more accurately and completely retained, and the generated target activation image can more effectively test the anti-attack ability of the face recognition algorithm.
- Key point alignment correction based on the three-dimensional posture information of the face can project the original 2D key point information into the 3D three-dimensional posture information of the face, and the face activation is thus transformed into pseudo-3D activation, so that the face key point information and the face texture features are better integrated to generate a more realistic activation image.
- FIG3 exemplarily shows a flow chart of the face activation method described in the embodiment of the present specification in one implementation manner.
- the face in the source image and the face in the driving video may not be the same person.
- the face in the driving video can be replaced with the target face in the source image using the face activation method.
- the source image and the driving image are input into a face activation generation model to generate a target activation image corresponding to the source image, wherein the face activation generation model is trained using the steps described in any of the above methods.
- the face activation generation model extracts the facial motion information of the source image and the driving image, and corrects the respective facial motion information into postures that are conducive to motion transfer in three-dimensional space, thereby generating a highly realistic target activation image. It can be more effectively used to test the anti-attack ability of face recognition algorithms, solve the problem of motion distortion, and more completely retain the facial identity information in the source image.
- an activation video in which the face has been replaced by the face in the source image can be obtained.
- FIG4 exemplarily shows a structural diagram of the face activation generation model training device described in the embodiment of the present specification in one implementation manner.
- the source image and driving image acquired by the sample acquisition module need to contain obviously different facial movements in order to more intuitively observe the effect of movement transfer.
- the sample acquisition module can extract source images and driving images from an existing image set containing different facial movements, or it can capture images of different people or several videos through a camera or other acquisition device, where each video records the facial movement change process of the same person.
- the sample acquisition module acquires a video as a training sample each time, and extracts two frames of images with different facial movements as the source image A and the driving image B respectively.
- the purpose of face activation is to transfer the movement in the driving image B to the face in the source image A.
- the face activation generation model is trained using the face of the same person as sample data, so that the model can be fully trained and thus more easily converged.
- the information extraction module can adopt a general face action detection method to obtain the face action information and Facial motion information that drives the image.
- Facial motion information can show the movement of the head and the position changes of the facial features in three-dimensional space, which has stronger realism and representation than the motion information in two-dimensional space.
- the facial motion information includes 3D facial posture information and facial key point information.
- the facial motion detection method includes a 3D facial posture detection method and a facial key point detection method.
- the 3D facial posture information can reflect the posture of the head in the 3D space to determine the orientation of the face; while the facial key point information records the position information by adding key points to the key facial features such as eyes, nose, mouth or other key facial areas, thereby expressing facial motion according to the position changes of these key areas.
- the three-dimensional facial posture information includes the head rotation angle.
- the three-dimensional facial posture information can be represented by the 3D facial position information, which is used to describe the rotation angle of the head on three mutually perpendicular coordinate axes in three-dimensional space.
- the facial position can reflect the direction of the head and express the head position information.
- These three mutually perpendicular coordinate axes are based on the directions of nodding, shaking and swinging the head to construct an XYZ axis centered on the head, which can reflect the rotation angle of the head on the three coordinate axes respectively, and are summarized as the 3D facial position information.
- the information extraction module also inputs the source image into a general image encoder for feature encoding, and outputs the corresponding source image features, so as to further process the source image such as action migration based on the source image features, and cooperate with the image decoder to generate the target activation image.
- the generation module corrects the facial key point information of the source image and the facial key point information of the driving image to the same posture by comparing the three-dimensional facial posture information of the source image and the three-dimensional facial posture information of the driving image, thereby obtaining the source key points and driving key points in the three-dimensional space; and projects the driving key points onto the source key points to migrate the movements of the driving image to the source image.
- the head posture of the source image and the head posture of the driving image can be obtained based on this.
- the key point information of the face in the image corresponds to the head posture.
- the generation module corrects the key points based on the difference in head posture between the two, so that the key points in the source image and the face behind the key points in the driving image are in the same posture, and the source key points and the driving key points are obtained, so as to further compare the key points with the same position and the key points with different positions.
- the key points with different positions represent the changes in the face action and are the objects that need to be focused on.
- the generation module projects the driving key points onto the source key points, adjusts the source key points according to the key points with different positions, thereby obtaining the face action information that needs to be migrated in the source image and completing the action migration of the driving image.
- Unifying the head posture of the source image with the head posture of the driving image is conducive to adjusting the face to a posture that is more convenient for action migration, obtaining a better face activation effect, and avoiding the influence of different head postures on the activation results.
- the generation module compares the three-dimensional facial posture information of the source image and the three-dimensional facial posture information of the driving image, rotates and/or offsets the facial key point information of the source image and the facial key point information of the driving image to the same posture, and obtains the source key points and driving key points in three-dimensional space.
- the generation module can rotate the facial key points to the key point positions corresponding to the target posture, or offset the facial key points.
- the facial key points can also be rotated to the approximate key point positions corresponding to the target posture, and then the facial key point positions can be further adjusted through offset.
- the generation module compares the three-dimensional facial posture information of the source image with the three-dimensional facial posture information of the driving image, and uses matrix changes to rotate and/or offset the facial key point information of the source image and the facial key point information of the driving image to the same posture, thereby obtaining the source key points and driving key points in three-dimensional space.
- the generation module selects the target posture for key point correction based on the facial key point information of the source image and the facial key point information of the driving image, and obtains the target key points corresponding to the target posture through a general key point detection method, and then determines the affine transformation matrices that adjust the facial key point information of the source image and the facial key point information of the driving image to the target key points, and uses the corresponding affine transformation matrices to correct the facial key point information of the source image and the facial key point information of the driving image to adjust to the target posture.
- the generation module uses a general image decoder to combine the source image features with the facial action information obtained by action migration to generate the target activation image.
- the image decoder in the generation module can be constructed based on the image decoder in the GAN generation network.
- the difference between the images can be obtained by calculating the difference in pixel-by-pixel values in the images.
- the training module can further process the images to be compared into images of other levels, and then calculate the pixel-by-pixel value difference, such as extracting the feature map of the image to be compared.
- the training module calculates the difference between the target activation image and the driving image through multiple ways, and then uses the obtained multiple loss training models to improve the training efficiency of the face activation generation model and accelerate the improvement of the model performance.
- the face activation generation model is implemented by a generative adversarial network, which includes a discriminator and a generator;
- the training module specifically includes: inputting the facial action information of the driving image, the facial action information of the source image and the source image features into the generator, and outputting the target activation image corresponding to the source image; using the discriminator to determine whether the target activation image is the driving image; calculating the generation loss based on the target activation image and the driving image; calculating the discrimination loss based on the discrimination result of the discriminator; and training the generative adversarial network with minimizing the discrimination loss and the generation loss as the training goals respectively.
- the generative adversarial network consists of a discriminator and a generator.
- the discriminator is first trained to have good discrimination ability; then the generator is trained so that the result generated by the generator can deceive the discriminator with good discrimination ability and have stronger fidelity.
- the training module obtains the target activation image, it inputs it into the discriminator to determine whether it is a real sample image or a tampered forged image to obtain the discrimination result; then the discrimination loss is calculated according to the accuracy of the discrimination result, and the discrimination ability of the discriminator is trained by minimizing the discrimination loss.
- the source image and the driving image in the training samples used are the same person.
- the training module transfers the facial movements in the driving image to the source image through face activation and merges them with the face in the source image.
- the obtained target activation image should be consistent with the driving image in an ideal situation. Therefore, the training module determines the generation loss by calculating the difference between the driving image and the target activation image, and optimizes the generation effect of the generator by minimizing the generation loss, so that the driving image and the source image are merged more naturally.
- the generator is continuously optimized, the generated target activation images become more and more realistic, making it increasingly difficult for the discriminator to determine their authenticity.
- the discriminator can no longer determine the authenticity of the target activation image, the generative adversarial training of the face activation generation model is completed.
- the face activation generation model In the process of iterative training by continuously calculating losses, the face activation generation model also learns and transfers the facial texture features in the source image and the driving image. That is, the face activation process associates and fuses the facial motion information representing the three-dimensional position of the face with the facial texture features, realizes face activation in three-dimensional space, and improves the problem of facial motion distortion after motion transfer.
- a GAN network model is used to construct a generative adversarial network to implement a face activation generation model.
- the training module is used to calculate a first pixel-based loss between a target activation image and a driving image; calculate a second feature map-based loss between a target activation image and a driving image; and weightedly fuse the first loss and the second loss as a generated loss.
- the first loss can be obtained by calculating the pixel-by-pixel mean square error loss between the target activation image and the driving image through a training module.
- the second loss is obtained by calculating the facial similarity between the target activation image and the driving image through the training module.
- the face similarity is calculated by combining Arcface (also known as the additive angular spacing loss function) with the cross entropy. Training the face activation generation model through Arcface can effectively increase the inter-class distance while ensuring the inter-class distance, and at the same time achieve the effect of reducing the intra-class distance, so that the face activation generation model can generate more accurate activation results more efficiently.
- Arcface also known as the additive angular spacing loss function
- FIG5 exemplarily shows a schematic structural diagram of the face activation device described in the embodiment of the present specification in one implementation manner.
- FIG5 it includes: a sample acquisition module 40, which is used to obtain a source image to be activated and a driving video containing a target action, and extract a plurality of frames of driving images from the driving video; an activation generation module 42, which is used to input the source image and the driving image into a face activation generation model for each frame of driving image obtained, and generate a target activation image corresponding to the source image, wherein the face activation generation model is trained using the steps described in any of the above methods; and a connection module 44, which is used to connect all the obtained target activation images to obtain a target activation video.
- a sample acquisition module 40 which is used to obtain a source image to be activated and a driving video containing a target action, and extract a plurality of frames of driving images from the driving video
- an activation generation module 42 which is used to input the source image and the driving image into a face activation generation model for each frame of driving image obtained, and generate a target activation image corresponding to the source image, wherein the
- the image can be directly input by installing a local image acquisition device, or a remote image acquisition device can be retrieved from the cloud for acquisition, or a pre-prepared image set can be retrieved from the cloud directly as a model input.
- the face in the source image and the face in the driving video may not be the same person.
- the face activation device can be used to replace the face in the driving video with the target face in the source image.
- the face activation generation model in the activation generation module extracts the face motion information of the source image and the driving image, and corrects the respective face motion information into a posture that is conducive to motion transfer in three-dimensional space, thereby generating a highly realistic target activation image. It can be used more effectively to test the anti-attack ability of face recognition algorithms, solve the problem of motion distortion, and retain the face identity information in the source image more completely.
- connection module connects the generated corresponding target activation image frames based on the order of the extracted driving image frames, so as to obtain an activation video in which the human face has been replaced by the human face in the source image.
- an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps described in any of the above methods when executing the program.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- General Physics & Mathematics (AREA)
- Evolutionary Computation (AREA)
- Multimedia (AREA)
- Artificial Intelligence (AREA)
- Software Systems (AREA)
- Computing Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Computational Linguistics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Mathematical Physics (AREA)
- Molecular Biology (AREA)
- Data Mining & Analysis (AREA)
- Biophysics (AREA)
- Medical Informatics (AREA)
- Databases & Information Systems (AREA)
- Human Computer Interaction (AREA)
- Biomedical Technology (AREA)
- General Engineering & Computer Science (AREA)
- Social Psychology (AREA)
- Psychiatry (AREA)
- Processing Or Creating Images (AREA)
- Image Analysis (AREA)
Abstract
本说明书实施例公开了一种人脸活化生成模型训练方法,所述方法包括:获取源图像和驱动图像,并输入所述人脸活化生成模型中,得到源图像的人脸动作信息和驱动图像的人脸动作信息;对源图像进行特征编码得到源图像特征;在三维空间内将驱动图像的人脸动作信息投射到源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中;根据驱动图像的人脸动作信息、源图像的人脸动作信息和源图像特征生成目标活化图像;以目标活化图像与驱动图像的差异最小化为训练目标,对所述人脸活化生成模型进行训练。本说明书实施例还公开了人脸活化方法。相应地,本说明书实施例公开了人脸活化生成模型训练装置和人脸活化装置。
Description
本说明书实施例涉及一种图像处理方法,尤其涉及一种人脸活化生成模型训练方法。
手机已经成为如今人人必备的通讯工具,随着科技的发展,使用手机进行身份验证的方式经历了从密码到指纹再到人脸的发展,而手机等设备的安全性也正受到各方面的挑战。
在IIFAA生物安全性攻击检测中,需要智能生成足够逼真的人脸攻击物料,来对手机的身份验证功能进行安全评测,以保证人脸安全评测的完备有效。该攻击手段首先要有通用性,在不同的手机黑盒测试中都具备较稳定的攻击力;其次要有高攻击率,具备强攻击性,从而能够区分人脸识别算法的抗攻击能力强弱,有效拉开不同待测手机厂商在身份验证功能的抗攻击性上的差距;另外该攻击手段需要通过算法智能生成,区别于人造面具等物理生成或普通的ps手段,使从人脸识别算法使用的模型中发现的潜在错误有迹可循,便于对模型进行改进,进而提高模型的可靠性与模型预测结果的可解释性。
相比于FOMM等目前主流的人脸活化算法,仍需要更有效的通过动作迁移获得攻击物料的方法,以减少在转头等动作上发生的动作畸变现象,保留完好的人脸身份信息,从而使在评价人脸识别算法的抗攻击性时获得的结果更加可靠。
鉴于此,希望获得一种新的人脸活化生成方法,能够生成更逼真的人脸活化图像,从而更有效地评测人脸识别算法。
发明内容
本说明书实施例的目的之一在于提供一种人脸活化生成模型训练方法,该方法能够将指定的动作迁移到指定的人脸图像中,较大程度地保留指定人脸的身份信息,有效测试人脸识别算法的抗攻击性。
根据上述目的,本说明书实施例提出了一种人脸活化生成模型训练方法,所述方法包括:获取源图像和驱动图像;将所述源图像和所述驱动图像输入所述人脸活化生成模型中,分别获取源图像的人脸动作信息和驱动图像的人脸动作信息;对源图像进行特征
编码,得到源图像特征;在三维空间内将所述驱动图像的人脸动作信息投射到所述源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中;根据所述驱动图像的人脸动作信息、所述源图像的人脸动作信息和所述源图像特征生成源图像对应的目标活化图像;以所述目标活化图像与所述驱动图像的差异最小化为训练目标,对所述人脸活化生成模型进行训练。
在本说明书实施例中,利用三维空间内的人脸动作信息投射实现了从驱动图像到源图像的动作迁移,将三维的位置信息与人脸纹理信息相关联,很好地改善了动作畸变的问题,使源图像中的人脸身份信息可以较准确且完整地保留;通过使目标活化图像与驱动图像的差异最小化,较大程度地提高了目标活化图像的逼真度。综上,通过该方法训练出的人脸活化生成模型所生成的目标活化图像能够更有效地测试人脸识别算法的抗攻击性。
进一步地,在一些实施方式中,所述人脸动作信息包括人脸三维姿态信息以及人脸关键点信息。
更进一步地,在一些实施方式中,在三维空间内将所述驱动图像的人脸动作信息投射到所述源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中具体包括:通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点;将所述驱动关键点投射到所述源关键点上,以将驱动图像的动作迁移至源图像中。
更进一步地,在一些实施方式中,所述人脸三维姿态信息包括头部转动角度。
进一步地,在一些实施方式中,所述人脸活化生成模型通过生成对抗网络实现,所述生成对抗网络包括判别器和生成器;以所述目标活化图像与所述驱动图像的差异最小化为训练目标,对所述人脸活化生成模型进行训练具体包括:将所述驱动图像的人脸动作信息、所述源图像的人脸动作信息和所述源图像特征输入所述生成器,输出源图像对应的目标活化图像;利用判别器判断所述目标活化图像是否为所述驱动图像;根据所述目标活化图像与所述驱动图像计算生成损失;根据所述判别器的判别结果计算判别损失;分别以所述判别损失和所述生成损失最小化为训练目标,对所述生成对抗网络进行训练。
更进一步地,在一些实施方式中,根据所述目标活化图像与所述驱动图像计算生成损失具体包括:计算所述目标活化图像与所述驱动图像之间基于像素的第一损失;计算
所述目标活化图像与所述驱动图像之间基于特征图的第二损失;将所述第一损失和所述第二损失加权融合,作为所述生成损失。
更进一步地,在一些实施方式中,通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点具体包括:通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
更进一步地,在一些实施方式中,通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点具体包括:通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,利用矩阵变化将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
本说明书实施例的另一目的在于提供一种人脸活化方法,该方法可以将指定的动作迁移到指定的人脸图像中,生成逼真的人脸活化图像,且较大程度地保留指定人脸的身份信息,从而有效测试人脸识别算法的抗攻击性。
根据上述目的,本说明书实施例提供了一种人脸活化方法,所述方法包括:获取待活化的源图像和包含目标动作的驱动视频,从所述驱动视频中抽取若干帧驱动图像;针对获得的每帧驱动图像,将所述源图像和所述驱动图像输入所述人脸活化生成模型中,生成源图像对应的目标活化图像,其中,所述人脸活化生成模型是采用如上任一方法所述的步骤训练得到的;将获得的所有目标活化图像进行连接,得到目标活化视频。
本说明书实施例的又一目的在于提供一种人脸活化生成模型训练装置,该装置能够将指定的动作迁移到指定的人脸图像中,较大程度地保留指定人脸的身份信息,可以有效测试人脸识别算法的抗攻击性。
根据上述目的,本说明书实施例提供了一种人脸活化生成模型训练装置,所述装置包括:样本获取模块,用于获取源图像和驱动图像;信息提取模块,用于将所述源图像和所述驱动图像输入所述人脸活化生成模型中,分别获取源图像的人脸动作信息和驱动图像的人脸动作信息;对源图像进行特征编码,得到源图像特征;生成模块,用于在三
维空间内将所述驱动图像的人脸动作信息投射到所述源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中;根据所述驱动图像的人脸动作信息、所述源图像的人脸动作信息和所述源图像特征生成源图像对应的目标活化图像;训练模块,用于以所述目标活化图像与所述驱动图像的差异最小化为训练目标,对所述人脸活化生成模型进行训练。
进一步地,在一些实施方式中,所述人脸动作信息包括人脸三维姿态信息以及人脸关键点信息。
更进一步地,在一些实施方式中,所述生成模块通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点;将所述驱动关键点投射到所述源关键点上,以将驱动图像的动作迁移至源图像中。
更进一步地,在一些实施方式中,所述人脸三维姿态信息包括头部转动角度。
进一步地,在一些实施方式中,所述人脸活化生成模型通过生成对抗网络实现,所述生成对抗网络包括判别器和生成器;所述训练模块具体包括:将所述驱动图像的人脸动作信息、所述源图像的人脸动作信息和所述源图像特征输入所述生成器,输出源图像对应的目标活化图像;利用判别器判断所述目标活化图像是否为所述驱动图像;根据所述目标活化图像与所述驱动图像计算生成损失;根据所述判别器的判别结果计算判别损失;分别以所述判别损失和所述生成损失最小化为训练目标,对所述生成对抗网络进行训练。
更进一步地,在一些实施方式中,所述训练模块计算所述目标活化图像与所述驱动图像之间基于像素的第一损失;计算所述目标活化图像与所述驱动图像之间基于特征图的第二损失;将所述第一损失和所述第二损失加权融合,作为所述生成损失。
更进一步地,在一些实施方式中,所述生成模块通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
更进一步地,在一些实施方式中,所述生成模块通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,利用矩阵变化将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关
键点和驱动关键点。
本说明书实施例的又一目的在于提供一种人脸活化装置,该装置可以将指定的动作迁移到指定的人脸图像中,生成逼真的人脸活化图像,且较大程度地保留指定人脸的身份信息,从而有效测试人脸识别算法的抗攻击性。
根据上述目的,本说明书实施例提供了一种人脸活化装置,所述装置包括:样本采集模块,用于获取待活化的源图像和包含目标动作的驱动视频,从所述驱动视频中抽取若干帧驱动图像;活化生成模块,用于针对获得的每帧驱动图像,将所述源图像和所述驱动图像输入所述人脸活化生成模型中,生成源图像对应的目标活化图像,其中,所述人脸活化生成模型是采用如上任一方法所述的步骤训练得到的;连接模块,用于将获得的所有目标活化图像进行连接,得到目标活化视频。
本说明书实施例的又一目的在于提供一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,其特征在于,所述处理器执行所述程序时实现如上任一方法所述的步骤。
本说明书实施例所述的人脸活化生成模型训练和人脸活化方法具有以下有益效果。
利用三维空间内从驱动图像到源图像的人脸动作信息投射实现动作迁移,将三维的位置信息与人脸纹理信息相关联,能够很好地改善头部转动导致的动作畸变问题,使源图像中的人脸身份信息可以较准确且完整地保留;通过使目标活化图像与驱动图像的差异最小化,较大程度地提高了目标活化图像的逼真度,所生成的目标活化图像能够更有效地测试人脸识别算法的抗攻击性。
基于人脸三维姿态信息进行关键点对齐矫正能够将原来2D的关键点信息投射到3D的人脸三维姿态信息中,人脸活化转变为伪3D的活化,使得人脸关键点信息与人脸纹理特征更好地融合,生成更逼真的活化图像。
本说明书实施例所述的人脸活化生成模型训练和人脸活化装置同样具有上述有益效果。
图1示例性地显示了本说明书实施例所述的人脸活化生成模型训练方法在一种实施方式下的流程示意图。
图2示例性地示出了本说明书实施例所述的人脸活化生成模型训练方法在一种具体
实施方式下执行的步骤。
图3示例性地显示了本说明书实施例所述的人脸活化方法在一种实施方式下的流程示意图。
图4示例性地显示了本说明书实施例所述的人脸活化生成模型训练装置在一种实施方式下的结构示意图。
图5示例性地显示了本说明书实施例所述的人脸活化装置在一种实施方式下的结构示意图。
首先需要说明的是,在本发明实施例中使用的术语是仅仅出于描述特定实施例的目的,而非旨在限制本发明。在本发明实施例和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。另外,在不冲突的情况下,本说明书中的实施例及实施例中的特征可以相互组合。
需要说明的是,人脸活化的过程表现为将图像B中人脸的动作或者表情迁移至图像A中的人脸上,由于表情变化和表情迁移可以表现为脸部五官的动作变化和五官的动作迁移,因此本说明书实施例中的动作不仅包括摇头、点头等运动,也包括微笑、皱眉、闭眼等表情。
下面将结合说明书附图和具体的实施例来对本发明所述的人脸活化生成模型训练和人脸活化方法以及装置进行进一步地详细说明,但是该详细说明不构成对本发明的限制。
在本说明书的一个实施例中,提出了一种人脸活化生成模型训练方法。图1示例性地显示了本说明书实施例所述的人脸活化生成模型训练方法在一种实施方式下的流程示意图。
如图1所示,包括如下步骤。
100:获取源图像和驱动图像。
使用到的源图像和驱动图像需要包含明显不同的人脸动作,以更直观地观察动作迁移的效果。源图像和驱动图像可以从现有的包含不同人脸动作的图像集中提取,也可以来自通过摄像头等采集设备拍摄的不同人的图像或者若干段视频,其中每段视频都记录了同一个人的人脸动作变化过程。
在一些实施例中,每次获取一段视频作为训练样本,从中抽取人脸动作不同的两帧图像分别作为源图像A和驱动图像B,人脸活化的目的是将驱动图像B中的动作迁移到源图像A中的人脸上。针对每一轮训练,利用同一个人的人脸作为样本数据对人脸活化生成模型进行训练,这样模型可以得到充分的训练,从而更容易收敛。
102:将源图像和驱动图像输入人脸活化生成模型中,分别获取源图像的人脸动作信息和驱动图像的人脸动作信息。
在人脸活化生成模型中,可以采取通用的人脸动作检测方法分别获取源图像的人脸动作信息和驱动图像的人脸动作信息。人脸动作信息能够在三维空间内表现头部的运动情况以及五官的位置变化情况,相对于二维空间的动作信息来说具有更强的逼真度和表征性。
在一些实施例中,人脸动作信息包括人脸三维姿态信息以及人脸关键点信息。
相应地,人脸动作检测方法包括人脸三维姿态检测方法和人脸关键点检测方法。人脸三维姿态信息可以体现三维空间内头部的姿态,以确定人脸的朝向;而人脸关键点信息通过在脸部眼睛、鼻子、嘴巴等关键五官或者其它脸部关键区域添加关键点来记录位置信息,从而根据这些关键区域的位置变化表现人脸动作。
在一些更具体的实施例中,人脸三维姿态信息包括头部转动角度。
人脸三维姿态信息可以用人脸3D位姿信息来表示,用于描述头部在三维空间中互相垂直的三个坐标轴上转动的角度,换句话说,人脸位姿可以体现头部的朝向,表达头部位置信息。这三个互相垂直的坐标轴基于点头、摇头以及摆头动作所沿的方向构建以头部为中心的XYZ轴,能够体现头部分别在三个坐标轴上的转动角度,并汇总为人脸3D位姿信息。
104:对源图像进行特征编码,得到源图像特征。
将源图像输入通用的图像编码器进行特征编码,输出对应的源图像特征,以便基于源图像特征对源图像作动作迁移等进一步处理,配合图像解码器生成目标活化图像。
106:在三维空间内将驱动图像的人脸动作信息投射到源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中。
在一些实施例中,在三维空间内将驱动图像的人脸动作信息投射到源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中具体包括:通过比对源图像的人脸三维
姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点;将驱动关键点投射到源关键点上,以将驱动图像的动作迁移至源图像中。
由于人脸三维姿态信息表现了三维空间内头部的姿态,可以基于此得到源图像的头部姿态与驱动图像的头部姿态。图像中的人脸关键点信息与头部姿态是对应的,为了使驱动图像中的动作与源图像中的人脸更好地融合,基于二者的头部姿态差异进行关键点矫正,使源图像中的关键点与驱动图像中的关键点背后所展现的人脸位于同一姿态下,得到源关键点和驱动关键点,以便进一步比对出位置相同与位置不同的关键点,位置不同的关键点即代表了人脸动作的变化,是需要进行重点迁移的对象。通过将驱动关键点投射到源关键点上,根据位置不同的关键点对源关键点进行调整,从而在源图像中获得需要迁移的人脸动作信息,完成驱动图像的动作迁移。将源图像的头部姿态与驱动图像的头部姿态进行统一,有利于将人脸调整至更方便进行动作迁移的姿态,获得更好的人脸活化效果,同时避免由于头部姿态的不同对活化结果产生影响。
需要说明的是,在将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下时,可以固定源图像的头部姿态对应的关键点,以将源图像的头部姿态作为目标姿态,通过矫正驱动图像的关键点使得驱动图像的头部姿态与目标姿态一致;也可以固定驱动图像的头部姿态对应的关键点,以将驱动图像的头部姿态作为目标姿态,通过矫正源图像的关键点使得源图像的头部姿态与目标姿态一致;也可以预先设置可以更好地展现人脸活化效果的目标姿态,再将源图像的头部姿态与驱动图像的头部姿态矫正至该目标姿态下。
在一些更具体的实施例中,通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点具体包括:通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
在对人脸关键点信息进行矫正时,可以将人脸关键点旋转至目标姿态对应的关键点位置,也可以对人脸关键点进行偏置,当然,还可以先将人脸关键点旋转至目标姿态对应的大致关键点位置,再经过偏置对人脸关键点位置进行进一步调整。
在一些更具体的实施例中,通过比对源图像的人脸三维姿态信息和驱动图像的人脸
三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点具体包括:通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,利用矩阵变化将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
基于源图像的人脸关键点信息和驱动图像的人脸关键点信息,选择进行关键点矫正的目标姿态,并通过通用的关键点检测方法获取该目标姿态对应的目标关键点,进而分别确定将源图像的人脸关键点信息和驱动图像的人脸关键点信息调整至目标关键点的仿射变换矩阵,利用对应的仿射变换矩阵分别对源图像的人脸关键点信息和驱动图像的人脸关键点信息进行矫正以调整至目标姿态。
108:根据驱动图像的人脸动作信息、源图像的人脸动作信息和源图像特征生成源图像对应的目标活化图像。
采用通用的图像解码器结合源图像特征与动作迁移获得的人脸动作信息生成目标活化图像。在一些实施例中,图像解码器可以基于GAN生成网络中的图像解码器构建。
110:以目标活化图像与驱动图像的差异最小化为训练目标,对人脸活化生成模型进行训练。
图像之间的差异可以通过计算图像中逐像素值的差异获得,在一些实施例中还可以将需要比对的图像进一步处理为其他层面的图像,再计算逐像素值差异,例如提取待比对图像的特征图。优选地,通过多种途径计算目标活化图像与驱动图像之间的差异,再利用获得的多种损失训练模型有利于提高人脸活化生成模型的训练效率,加快模型性能的提升。
在一些实施例中,人脸活化生成模型通过生成对抗网络实现,生成对抗网络包括判别器和生成器;以目标活化图像与驱动图像的差异最小化为训练目标,对人脸活化生成模型进行训练具体包括:将驱动图像的人脸动作信息、源图像的人脸动作信息和源图像特征输入生成器,输出源图像对应的目标活化图像;利用判别器判断目标活化图像是否为驱动图像;根据目标活化图像与驱动图像计算生成损失;根据判别器的判别结果计算判别损失;分别以判别损失和生成损失最小化为训练目标,对生成对抗网络进行训练。
生成对抗网络由判别器和生成器组成,在训练生成对抗网络的过程中,首先训练判别器,使其具有良好的判别能力;再训练生成器,使得生成器生成的结果能过骗过判别
能力良好的判别器,拥有更强的逼真度。在上述实施例中,获得目标活化图像后将其输入判别器,以判断是真实的样本图像还是经过篡改的伪造图像,获得判别结果;再根据判别结果的准确性计算判别损失,通过使判别损失最小化训练判别器的判别能力。
优选地,针对人脸活化生成模型的每一轮训练,采用的训练样本中源图像与驱动图像为同一个人。通过人脸活化将驱动图像中的人脸动作迁移至源图像中,与源图像中的人脸融合,所获得的目标活化图像在理想情况下应该与驱动图像一致,因此通过计算驱动图像与目标活化图像之间的差异确定生成损失,通过令生成损失最小化优化生成器的生成效果,使驱动图像与源图像融合得更加自然。
随着生成器的不断优化,生成的目标活化图像越来越逼真,使得判别器越来越难以判断其真伪,当判别器已无法判断目标活化图像的真伪时,对人脸活化生成模型的生成对抗式训练完成。
在不断地通过计算损失进行迭代训练的过程中,人脸活化生成模型还学习并迁移了源图像与驱动图像中的人脸纹理特征,即该人脸活化过程将表示人脸三维位置的人脸动作信息与人脸纹理特征进行关联与融合,实现三维空间下的人脸活化,改善了动作迁移后人脸动作发生畸变的问题。
在一些更具体的实施例中,采用GAN网络模型构建生成对抗网络来实现人脸活化生成模型。
在一些更具体的实施例中,根据目标活化图像与驱动图像计算生成损失具体包括:计算目标活化图像与驱动图像之间基于像素的第一损失;计算目标活化图像与驱动图像之间基于特征图的第二损失;将第一损失和第二损失加权融合,作为生成损失。
优选地,第一损失可以通过计算目标活化图像与驱动图像之间逐像素的均方差损失获得。
优选地,第二损失通过计算目标活化图像与驱动图像之间的人脸相似度获得,可以分别提取出目标活化图像的特征向量与驱动图像的特征向量后,采用Arcface(又称加性角度间隔损失函数)与交叉熵结合的方式计算人脸相似度。通过Arcface对人脸活化生成模型进行训练可以在保证类间距离的情况下有效地增大类间距离,同时达到减小类内距离的效果,使得人脸活化生成模型能够更高效地生成更准确的活化结果。
图2示例性地示出了本说明书实施例所述的人脸活化生成模型训练方法在一种具体实施方式下执行的步骤。
如图2所示的,在一些实施例中,从同一人的一段视频中抽取两帧图像分别作为源图像和驱动图像,分别提取源图像的人脸动作信息和驱动图像的人脸动作信息。其中,通过通用的人脸关键点检测方法分别提取出源图像的人脸关键点以及驱动图像的人脸关键点,由于源图像与驱动图像中的人脸相同,可以仅对源图像中的人脸关键点进行提取,并复制到驱动图像中;再通过通用的人脸位姿检测方法分别提取出源图像的人脸3D位姿以及驱动图像的人脸3D位姿,基于人脸3D位姿比对出源图像与驱动图像中人脸的头部姿态差异,利用矩阵变化进行人脸关键点矫正,使得源图像中的头部姿态与驱动图像中的头部姿态一致,得到矫正后的源关键点与驱动关键点。
接着,对源图像中的图像特征进行提取,将源图像特征与源关键点、驱动关键点一起输入基于GAN生成网络中的图像解码器构建的生成器中,通过将驱动关键点投射到源关键点上进行动作迁移,并结合源图像特征生成最终的目标活化图像。然后根据目标活化图像与驱动图像的差异最小化为训练目标,对人脸活化生成模型进行迭代训练,得到生成的活化图像中无动作畸变的人脸活化生成模型。
本说明书实施例提供的人脸活化生成模型训练方法利用三维空间内从驱动图像到源图像的人脸动作信息投射实现动作迁移,能够很好地改善头部转动导致的动作畸变问题;通过使目标活化图像与驱动图像的差异最小化,较大程度地提高了目标活化图像的逼真度,使源图像中的人脸身份信息可以较准确且完整地保留,所生成的目标活化图像能够更有效地测试人脸识别算法的抗攻击性。基于人脸三维姿态信息进行关键点对齐矫正能够将原来2D的关键点信息投射到3D的人脸三维姿态信息中,人脸活化从而转变为伪3D的活化,使得人脸关键点信息与人脸纹理特征更好地融合,生成更逼真的活化图像。
在本说明书的另一个实施例中,提出了一种人脸活化方法,图3示例性地显示了本说明书实施例所述的人脸活化方法在一种实施方式下的流程示意图。
如图3所示,包括如下步骤。
200:获取待活化的源图像和包含目标动作的驱动视频,从驱动视频中抽取若干帧驱动图像。
在采集样本时,可以直接通过安装本地图像采集设备输入图像,也可以从云端调取异地的图像采集设备进行采集,或者从云端调取预先准备好的图像集直接作为模型输入。
在进行人脸活化时,源图像中的人脸与驱动视频中的人脸可以不是同一人,换句话说,使用该人脸活化方法可以将驱动视频中的人脸替换为源图像中的目标人脸。
202:针对获得的每帧驱动图像,将源图像和驱动图像输入人脸活化生成模型中,生成源图像对应的目标活化图像,其中,人脸活化生成模型是采用如上任一方法所述的步骤训练得到的。
人脸活化生成模型通过提取源图像的人脸动作信息和驱动图像的人脸动作信息,在三维空间下将各自的人脸动作信息矫正为有利于进行动作迁移的姿态,从而生成逼真度极高的目标活化图像,能够更有效地用于对人脸识别算法进行抗攻击性测试,并解决动作畸变问题,较为完整地保留源图像中的人脸身份信息。
204:将获得的所有目标活化图像进行连接,得到目标活化视频。
基于抽取的驱动图像帧的顺序对生成的对应目标活化图像帧进行连接,即可得到人脸已替换为源图像中人脸的活化视频。
在本说明书的一个实施例中,提出了一种人脸活化生成模型训练装置,图4示例性地显示了本说明书实施例所述的人脸活化生成模型训练装置在一种实施方式下的结构示意图。
如图4所示,包括:样本获取模块30,用于获取源图像和驱动图像;信息提取模块32,用于将源图像和驱动图像输入人脸活化生成模型中,分别获取源图像的人脸动作信息和驱动图像的人脸动作信息;对源图像进行特征编码,得到源图像特征;生成模块34,用于在三维空间内将驱动图像的人脸动作信息投射到源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中;根据驱动图像的人脸动作信息、源图像的人脸动作信息和源图像特征生成源图像对应的目标活化图像;训练模块36,用于以目标活化图像与驱动图像的差异最小化为训练目标,对人脸活化生成模型进行训练。
样本获取模块获取到的源图像和驱动图像需要包含明显不同的人脸动作,以更直观地观察动作迁移的效果。样本获取模块可以从现有的包含不同人脸动作的图像集中提取源图像和驱动图像,也可以通过摄像头等采集设备拍摄不同人的图像或者若干段视频,其中每段视频都记录了同一个人的人脸动作变化过程。
在一些实施例中,样本获取模块每次获取一段视频作为训练样本,从中抽取人脸动作不同的两帧图像分别作为源图像A和驱动图像B,人脸活化的目的是将驱动图像B中的动作迁移到源图像A中的人脸上。针对每一轮训练,利用同一个人的人脸作为样本数据对人脸活化生成模型进行训练,这样模型可以得到充分的训练,从而更容易收敛。
信息提取模块可以采取通用的人脸动作检测方法分别获取源图像的人脸动作信息和
驱动图像的人脸动作信息。人脸动作信息能够在三维空间内表现头部的运动情况以及五官的位置变化情况,相对于二维空间的动作信息来说具有更强的逼真度和表征性。
在一些实施例中,人脸动作信息包括人脸三维姿态信息以及人脸关键点信息。相应地,人脸动作检测方法包括人脸三维姿态检测方法和人脸关键点检测方法。人脸三维姿态信息可以体现三维空间内头部的姿态,以确定人脸的朝向;而人脸关键点信息通过在脸部眼睛、鼻子、嘴巴等关键五官或者其它脸部关键区域添加关键点来记录位置信息,从而根据这些关键区域的位置变化表现人脸动作。
在一些更具体的实施例中,人脸三维姿态信息包括头部转动角度。人脸三维姿态信息可以用人脸3D位姿信息来表示,用于描述头部在三维空间中互相垂直的三个坐标轴上转动的角度,换句话说,人脸位姿可以体现头部的朝向,表达头部位置信息。这三个互相垂直的坐标轴基于点头、摇头以及摆头动作所沿的方向构建以头部为中心的XYZ轴,能够体现头部分别在三个坐标轴上的转动角度,并汇总为人脸3D位姿信息。
信息提取模块还将源图像输入通用的图像编码器进行特征编码,输出对应的源图像特征,以便基于源图像特征对源图像作动作迁移等进一步处理,配合图像解码器生成目标活化图像。
在一些更具体的实施例中,生成模块通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点;将驱动关键点投射到源关键点上,以将驱动图像的动作迁移至源图像中。
由于人脸三维姿态信息表现了三维空间内头部的姿态,可以基于此得到源图像的头部姿态与驱动图像的头部姿态。图像中的人脸关键点信息与头部姿态是对应的,为了使驱动图像中的动作与源图像中的人脸更好地融合,生成模块基于二者的头部姿态差异进行关键点矫正,使源图像中的关键点与驱动图像中的关键点背后所展现的人脸位于同一姿态下,得到源关键点和驱动关键点,以便进一步比对出位置相同与位置不同的关键点,位置不同的关键点即代表了人脸动作的变化,是需要进行重点迁移的对象。生成模块通过将驱动关键点投射到源关键点上,根据位置不同的关键点对源关键点进行调整,从而在源图像中获得需要迁移的人脸动作信息,完成驱动图像的动作迁移。将源图像的头部姿态与驱动图像的头部姿态进行统一,有利于将人脸调整至更方便进行动作迁移的姿态,获得更好的人脸活化效果,同时避免由于头部姿态的不同对活化结果产生影响。
在一些更具体的实施例中,生成模块通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
生成模块在对人脸关键点信息进行矫正时,可以将人脸关键点旋转至目标姿态对应的关键点位置,也可以对人脸关键点进行偏置,当然,还可以先将人脸关键点旋转至目标姿态对应的大致关键点位置,再经过偏置对人脸关键点位置进行进一步调整。
在一些更具体的实施例中,生成模块通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,利用矩阵变化将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
生成模块基于源图像的人脸关键点信息和驱动图像的人脸关键点信息,选择进行关键点矫正的目标姿态,并通过通用的关键点检测方法获取该目标姿态对应的目标关键点,进而分别确定将源图像的人脸关键点信息和驱动图像的人脸关键点信息调整至目标关键点的仿射变换矩阵,利用对应的仿射变换矩阵分别对源图像的人脸关键点信息和驱动图像的人脸关键点信息进行矫正以调整至目标姿态。
生成模块采用通用的图像解码器结合源图像特征与动作迁移获得的人脸动作信息生成目标活化图像。在一些实施例中,生成模块中的图像解码器可以基于GAN生成网络中的图像解码器构建。
图像之间的差异可以通过计算图像中逐像素值的差异获得,在一些实施例中训练模块还可以将需要比对的图像进一步处理为其他层面的图像,再计算逐像素值差异,例如提取待比对图像的特征图。优选地,训练模块通过多种途径计算目标活化图像与驱动图像之间的差异,再利用获得的多种损失训练模型有利于提高人脸活化生成模型的训练效率,加快模型性能的提升。
在一些实施例中,人脸活化生成模型通过生成对抗网络实现,生成对抗网络包括判别器和生成器;训练模块具体包括:将驱动图像的人脸动作信息、源图像的人脸动作信息和源图像特征输入生成器,输出源图像对应的目标活化图像;利用判别器判断目标活化图像是否为驱动图像;根据目标活化图像与驱动图像计算生成损失;根据判别器的判别结果计算判别损失;分别以判别损失和生成损失最小化为训练目标,对生成对抗网络进行训练。
生成对抗网络由判别器和生成器组成,在训练生成对抗网络的过程中,首先训练判别器,使其具有良好的判别能力;再训练生成器,使得生成器生成的结果能过骗过判别能力良好的判别器,拥有更强的逼真度。在上述实施例中,训练模块获得目标活化图像后将其输入判别器,以判断是真实的样本图像还是经过篡改的伪造图像,获得判别结果;再根据判别结果的准确性计算判别损失,通过使判别损失最小化训练判别器的判别能力。
优选地,针对人脸活化生成模型的每一轮训练,采用的训练样本中源图像与驱动图像为同一个人。训练模块通过人脸活化将驱动图像中的人脸动作迁移至源图像中,与源图像中的人脸融合,所获得的目标活化图像在理想情况下应该与驱动图像一致,因此训练模块通过计算驱动图像与目标活化图像之间的差异确定生成损失,通过令生成损失最小化优化生成器的生成效果,使驱动图像与源图像融合得更加自然。
随着生成器的不断优化,生成的目标活化图像越来越逼真,使得判别器越来越难以判断其真伪,当判别器已无法判断目标活化图像的真伪时,对人脸活化生成模型的生成对抗式训练完成。
在不断地通过计算损失进行迭代训练的过程中,人脸活化生成模型还学习并迁移了源图像与驱动图像中的人脸纹理特征,即该人脸活化过程将表示人脸三维位置的人脸动作信息与人脸纹理特征进行关联与融合,实现三维空间下的人脸活化,改善了动作迁移后人脸动作发生畸变的问题。
在一些更具体的实施例中,采用GAN网络模型构建生成对抗网络来实现人脸活化生成模型。
在一些更具体的实施例中,训练模块用于计算目标活化图像与驱动图像之间基于像素的第一损失;计算目标活化图像与驱动图像之间基于特征图的第二损失;将第一损失和第二损失加权融合,作为生成损失。
优选地,第一损失可以通过训练模块计算目标活化图像与驱动图像之间逐像素的均方差损失获得。
优选地,第二损失通过训练模块计算目标活化图像与驱动图像之间的人脸相似度获得,可以分别提取出目标活化图像的特征向量与驱动图像的特征向量后,采用Arcface(又称加性角度间隔损失函数)与交叉熵结合的方式计算人脸相似度。通过Arcface对人脸活化生成模型进行训练可以在保证类间距离的情况下有效地增大类间距离,同时达到减小类内距离的效果,使得人脸活化生成模型能够更高效地生成更准确的活化结果。
在本说明书的另一个实施例中,提出了一种人脸活化装置,图5示例性地显示了本说明书实施例所述的人脸活化装置在一种实施方式下的结构示意图。
如图5所示,包括:样本采集模块40,用于获取待活化的源图像和包含目标动作的驱动视频,从驱动视频中抽取若干帧驱动图像;活化生成模块42,用于针对获得的每帧驱动图像,将源图像和驱动图像输入人脸活化生成模型中,生成源图像对应的目标活化图像,其中,人脸活化生成模型是采用如上任一方法所述的步骤训练得到的;连接模块44,用于将获得的所有目标活化图像进行连接,得到目标活化视频。
在样本采集模块采集样本时,可以直接通过安装本地图像采集设备输入图像,也可以从云端调取异地的图像采集设备进行采集,或者从云端调取预先准备好的图像集直接作为模型输入。在样本采集模块中,源图像中的人脸与驱动视频中的人脸可以不是同一人,换句话说,使用该人脸活化装置可以将驱动视频中的人脸替换为源图像中的目标人脸。
活化生成模块中的人脸活化生成模型通过提取源图像的人脸动作信息和驱动图像的人脸动作信息,在三维空间下将各自的人脸动作信息矫正为有利于进行动作迁移的姿态,从而生成逼真度极高的目标活化图像,能够更有效地用于对人脸识别算法进行抗攻击性测试,并解决动作畸变问题,较为完整地保留源图像中的人脸身份信息。
连接模块基于抽取的驱动图像帧的顺序对生成的对应目标活化图像帧进行连接,即可得到人脸已替换为源图像中人脸的活化视频。
在本说明书的一个实施例中,还提出了一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,其特征在于,处理器执行程序时实现如上任一方法所述的步骤。
上述对本说明书特定实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。
需要注意的是,以上列举的仅为本发明的具体实施例,显然本发明不限于以上实施例,随之有着许多的类似变化。本领域的技术人员如果从本发明公开的内容直接导出或联想到的所有变形,均应属于本发明的保护范围。
Claims (19)
- 一种人脸活化生成模型训练方法,所述方法包括:获取源图像和驱动图像;将所述源图像和所述驱动图像输入所述人脸活化生成模型中,分别获取源图像的人脸动作信息和驱动图像的人脸动作信息;对源图像进行特征编码,得到源图像特征;在三维空间内将所述驱动图像的人脸动作信息投射到所述源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中;根据所述驱动图像的人脸动作信息、所述源图像的人脸动作信息和所述源图像特征生成源图像对应的目标活化图像;以所述目标活化图像与所述驱动图像的差异最小化为训练目标,对所述人脸活化生成模型进行训练。
- 如权利要求1所述的人脸活化生成模型训练方法,所述人脸动作信息包括人脸三维姿态信息以及人脸关键点信息。
- 如权利要求2所述的人脸活化生成模型训练方法,在三维空间内将所述驱动图像的人脸动作信息投射到所述源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中具体包括:通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点;将所述驱动关键点投射到所述源关键点上,以将驱动图像的动作迁移至源图像中。
- 如权利要求2所述的人脸活化生成模型训练方法,所述人脸三维姿态信息包括头部转动角度。
- 如权利要求1所述的人脸活化生成模型训练方法,所述人脸活化生成模型通过生成对抗网络实现,所述生成对抗网络包括判别器和生成器;以所述目标活化图像与所述驱动图像的差异最小化为训练目标,对所述人脸活化生成模型进行训练具体包括:将所述驱动图像的人脸动作信息、所述源图像的人脸动作信息和所述源图像特征输入所述生成器,输出源图像对应的目标活化图像;利用判别器判断所述目标活化图像是否为所述驱动图像;根据所述目标活化图像与所述驱动图像计算生成损失;根据所述判别器的判别结果计算判别损失;分别以所述判别损失和所述生成损失最小化为训练目标,对所述生成对抗网络进行训练。
- 如权利要求5所述的人脸活化生成模型训练方法,根据所述目标活化图像与所述驱动图像计算生成损失具体包括:计算所述目标活化图像与所述驱动图像之间基于像素的第一损失;计算所述目标活化图像与所述驱动图像之间基于特征图的第二损失;将所述第一损失和所述第二损失加权融合,作为所述生成损失。
- 如权利要求3所述的人脸活化生成模型训练方法,通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点具体包括:通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
- 如权利要求7所述的人脸活化生成模型训练方法,通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点具体包括:通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,利用矩阵变化将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
- 一种人脸活化方法,所述方法包括:获取待活化的源图像和包含目标动作的驱动视频,从所述驱动视频中抽取若干帧驱动图像;针对获得的每帧驱动图像,将所述源图像和所述驱动图像输入所述人脸活化生成模型中,生成源图像对应的目标活化图像,其中,所述人脸活化生成模型是采用如权利要求1至8任一所述的方法训练得到的;将获得的所有目标活化图像进行连接,得到目标活化视频。
- 一种人脸活化生成模型训练装置,所述装置包括:样本获取模块,用于获取源图像和驱动图像;信息提取模块,用于将所述源图像和所述驱动图像输入所述人脸活化生成模型中, 分别获取源图像的人脸动作信息和驱动图像的人脸动作信息;对源图像进行特征编码,得到源图像特征;生成模块,用于在三维空间内将所述驱动图像的人脸动作信息投射到所述源图像的人脸动作信息上,以使驱动图像的动作迁移至源图像中;根据所述驱动图像的人脸动作信息、所述源图像的人脸动作信息和所述源图像特征生成源图像对应的目标活化图像;训练模块,用于以所述目标活化图像与所述驱动图像的差异最小化为训练目标,对所述人脸活化生成模型进行训练。
- 如权利要求10所述的人脸活化生成模型训练装置,所述人脸动作信息包括人脸三维姿态信息以及人脸关键点信息。
- 如权利要求11所述的人脸活化生成模型训练装置,所述生成模块通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点;将所述驱动关键点投射到所述源关键点上,以将驱动图像的动作迁移至源图像中。
- 如权利要求11所述的人脸活化生成模型训练装置,所述人脸三维姿态信息包括头部转动角度。
- 如权利要求10所述的人脸活化生成模型训练装置,所述人脸活化生成模型通过生成对抗网络实现,所述生成对抗网络包括判别器和生成器;所述训练模块具体包括:将所述驱动图像的人脸动作信息、所述源图像的人脸动作信息和所述源图像特征输入所述生成器,输出源图像对应的目标活化图像;利用判别器判断所述目标活化图像是否为所述驱动图像;根据所述目标活化图像与所述驱动图像计算生成损失;根据所述判别器的判别结果计算判别损失;分别以所述判别损失和所述生成损失最小化为训练目标,对所述生成对抗网络进行训练。
- 如权利要求14所述的人脸活化生成模型训练装置,所述训练模块计算所述目标活化图像与所述驱动图像之间基于像素的第一损失;计算所述目标活化图像与所述驱动图像之间基于特征图的第二损失;将所述第一损失和所述第二损失加权融合,作为所述生成损失。
- 如权利要求12所述的人脸活化生成模型训练装置,所述生成模块通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
- 如权利要求16所述的人脸活化生成模型训练装置,所述生成模块通过比对源图像的人脸三维姿态信息和驱动图像的人脸三维姿态信息,利用矩阵变化将源图像的人脸关键点信息和驱动图像的人脸关键点信息经过旋转或者/以及偏置矫正到同一姿态下,得到三维空间下的源关键点和驱动关键点。
- 一种人脸活化装置,所述装置包括:样本采集模块,用于获取待活化的源图像和包含目标动作的驱动视频,从所述驱动视频中抽取若干帧驱动图像;活化生成模块,用于针对获得的每帧驱动图像,将所述源图像和所述驱动图像输入所述人脸活化生成模型中,生成源图像对应的目标活化图像,其中,所述人脸活化生成模型是采用如权利要求1至8任一所述的方法训练得到的;连接模块,用于将获得的所有目标活化图像进行连接,得到目标活化视频。
- 一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,其特征在于,所述处理器执行所述程序时实现上述权利要求1-9中任一所述方法的步骤。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310814024.XA CN116935156B (zh) | 2023-07-04 | 2023-07-04 | 一种人脸活化生成模型训练和人脸活化的方法及装置 |
| CN202310814024.X | 2023-07-04 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025007933A1 true WO2025007933A1 (zh) | 2025-01-09 |
Family
ID=88378267
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/103675 Ceased WO2025007933A1 (zh) | 2023-07-04 | 2024-07-04 | 人脸活化生成模型训练和人脸活化 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN116935156B (zh) |
| WO (1) | WO2025007933A1 (zh) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116935156B (zh) * | 2023-07-04 | 2026-04-28 | 支付宝(杭州)数字服务技术有限公司 | 一种人脸活化生成模型训练和人脸活化的方法及装置 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112233012A (zh) * | 2020-08-10 | 2021-01-15 | 上海交通大学 | 一种人脸生成系统及方法 |
| CN113807265A (zh) * | 2021-09-18 | 2021-12-17 | 山东财经大学 | 一种多样化的人脸图像合成方法及系统 |
| US20230162376A1 (en) * | 2021-11-23 | 2023-05-25 | VIRNECT inc. | Method and system for estimating motion of real-time image target between successive frames |
| CN116311460A (zh) * | 2023-03-23 | 2023-06-23 | 平安科技(深圳)有限公司 | 图像生成方法、装置、设备及存储介质 |
| CN116935156A (zh) * | 2023-07-04 | 2023-10-24 | 支付宝(杭州)信息技术有限公司 | 一种人脸活化生成模型训练和人脸活化的方法及装置 |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112734890B (zh) * | 2020-12-22 | 2023-11-10 | 上海影谱科技有限公司 | 基于三维重建的人脸替换方法及装置 |
-
2023
- 2023-07-04 CN CN202310814024.XA patent/CN116935156B/zh active Active
-
2024
- 2024-07-04 WO PCT/CN2024/103675 patent/WO2025007933A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112233012A (zh) * | 2020-08-10 | 2021-01-15 | 上海交通大学 | 一种人脸生成系统及方法 |
| CN113807265A (zh) * | 2021-09-18 | 2021-12-17 | 山东财经大学 | 一种多样化的人脸图像合成方法及系统 |
| US20230162376A1 (en) * | 2021-11-23 | 2023-05-25 | VIRNECT inc. | Method and system for estimating motion of real-time image target between successive frames |
| CN116311460A (zh) * | 2023-03-23 | 2023-06-23 | 平安科技(深圳)有限公司 | 图像生成方法、装置、设备及存储介质 |
| CN116935156A (zh) * | 2023-07-04 | 2023-10-24 | 支付宝(杭州)信息技术有限公司 | 一种人脸活化生成模型训练和人脸活化的方法及装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116935156A (zh) | 2023-10-24 |
| CN116935156B (zh) | 2026-04-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Chen et al. | Vgan-based image representation learning for privacy-preserving facial expression recognition | |
| CN109684924B (zh) | 人脸活体检测方法及设备 | |
| Ye et al. | Audio-driven talking face video generation with dynamic convolution kernels | |
| CN110147721B (zh) | 一种三维人脸识别方法、模型训练方法和装置 | |
| CN114187165B (zh) | 图像处理方法和装置 | |
| WO2023011013A1 (zh) | 视频图像的拼缝搜索方法、视频图像的拼接方法和装置 | |
| CN112528902B (zh) | 一种基于3d人脸模型的视频监控动态人脸识别方法及装置 | |
| CN112288627A (zh) | 一种面向识别的低分辨率人脸图像超分辨率方法 | |
| Manikandan et al. | Hand gesture detection and conversion to speech and text | |
| CN113570689B (zh) | 人像卡通化方法、装置、介质和计算设备 | |
| CN110909634A (zh) | 可见光与双红外线相结合的快速活体检测方法 | |
| CN117876609B (zh) | 一种多特征三维人脸重建方法、系统、设备及存储介质 | |
| WO2024198475A1 (zh) | 人脸活体识别方法、装置、电子设备及存储介质 | |
| WO2024055379A1 (zh) | 基于角色化身模型的视频处理方法、系统及相关设备 | |
| CN111275778A (zh) | 人脸简笔画生成方法及装置 | |
| CN115601710A (zh) | 基于自注意力网络架构的考场异常行为监测方法及系统 | |
| CN114399824A (zh) | 一种多角度侧面人脸矫正方法、装置、计算机设备和介质 | |
| WO2025007933A1 (zh) | 人脸活化生成模型训练和人脸活化 | |
| CN112016508B (zh) | 人脸识别方法、装置、系统、计算设备及存储介质 | |
| CN113610058A (zh) | 一种人脸特征迁移的面部姿态增强互动方法 | |
| Ramachandra et al. | VoxAtnNet: A 3D point clouds convolutional neural network for generalizable face presentation attack detection | |
| Park et al. | 3D face reconstruction from stereo video | |
| CN110598595B (zh) | 一种基于人脸关键点和姿态的多属性人脸生成算法 | |
| CN109886084A (zh) | 基于陀螺仪的人脸认证方法、电子设备及存储介质 | |
| CN114373044A (zh) | 生成脸部三维模型的方法、装置、计算设备和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24835411 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |