WO2023116208A1 - 数字对象生成方法、装置、设备及存储介质 - Google Patents
数字对象生成方法、装置、设备及存储介质 Download PDFInfo
- Publication number
- WO2023116208A1 WO2023116208A1 PCT/CN2022/128915 CN2022128915W WO2023116208A1 WO 2023116208 A1 WO2023116208 A1 WO 2023116208A1 CN 2022128915 W CN2022128915 W CN 2022128915W WO 2023116208 A1 WO2023116208 A1 WO 2023116208A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- digital object
- target
- video
- audio
- reference digital
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/57—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for processing of video signals
Definitions
- the present disclosure relates to the technical field of artificial intelligence, and in particular to a digital object generation method, device, equipment and storage medium.
- Digital objects for example, digital humans are widely used in fields such as live broadcast, news broadcast, and voice prompts. Usually, it is necessary to drive the digital object based on the audio to be broadcast to perform actions, expressions, etc. that match the audio, so as to obtain the audio-driven video. In related technologies, it is generally necessary to obtain a large number of audio and video of digital objects (hereinafter referred to as audio and video), and then pre-train each digital object to obtain the voice-driven model of the digital object. After obtaining the voice-driven model of the digital object , different audio can be input into the speech-driven model, that is, the video of the digital object driven by different audio can be output. But this way can only get a video of a fixed digital object driven by different audios.
- audio and video digital objects
- Audio and video use the audio and video of the new digital object to retrain the voice-driven model of the new digital object.
- the entire training process requires a large amount of data and takes a long time, which makes it impossible to quickly generate a video of the new digital object driven by a certain audio.
- the disclosure provides a digital object generation method, device, equipment and storage medium.
- a method for generating a digital object comprising: acquiring a target audio and a target digital object; obtaining a reference A first driving video of a digital object driven by the target audio; wherein, the voice-driven model is obtained based on audio and video training of the reference digital object; The pose of the reference digital object is transferred to the target digital object, and a second driving video of the target digital object driven by the target audio is obtained.
- obtaining the first driving video of the reference digital object driven by the target audio includes: extracting phonemes from the target audio, obtaining A phoneme time stamp; inputting the phoneme time stamp into the speech-driven model to obtain the first driving video.
- the voice-driven model is trained based on the audio and video of the reference digital object, comprising: obtaining the audio of the reference digital object, and the audio of the reference digital object synchronized with the audio of the reference digital object Video; extracting phonemes from the audio of the reference digital object to obtain the phoneme time stamp of the reference digital object; using the phoneme time stamp of the reference digital object as a training sample, and using the video of the reference digital object as a sample label , training to obtain the speech-driven model.
- the pose of the reference digital object in each video frame of the first driving video is transferred to the target digital object to obtain the second driving of the target digital object driven by the target audio
- the video includes: extracting texture features related to the texture of the target digital object from the target image of the target digital object; pose features related to the pose of the reference digital object; reconstruct the target digital object according to the texture features and the pose features, wherein the pose of the reconstructed target digital object is the same as that of the reference digital object in the video frame
- the gestures are consistent; obtaining a reconstructed frame corresponding to the video frame according to the reconstructed target digital object, and constructing the second driving video based on the reconstructed frame.
- extracting a pose feature related to the pose of the reference digital object from the video frame includes: using an encoder in an autoencoder of the reference digital object to extract the pose feature from the video frame The pose feature, wherein the self-encoder of the reference digital object is pre-trained based on a plurality of images containing the reference digital object.
- extracting texture features related to the texture of the target digital object from the target image of the target digital object includes: using a decoder in an autoencoder of the target digital object to extract from the The texture feature is extracted from a target image of the target digital object, wherein the autoencoder of the target digital object is pre-trained based on a plurality of images containing the target digital object.
- reconstructing the target digital object according to the texture features and the pose features includes: using a decoder in an autoencoder of the target digital object to reconstruct according to the texture features and the pose features The target digital object.
- the reconstructed frame corresponding to the previous frame is used as the reconstructed frame corresponding to the current frame.
- obtaining the reconstructed frame corresponding to the video frame according to the reconstructed target digital object includes: generating the reconstructed frame according to the reconstructed target digital object and preset background material .
- the pose of the reference digital object includes one or more of the following: facial movements of the reference digital object, facial expressions of the reference digital object, and body movements of the reference digital object.
- a method for generating a digital object comprising: acquiring a first driving video of a reference digital object driven by a target audio; Gesture feature in the video; the reference digital object in the first driving video is replaced by the target digital object, and the gesture feature is loaded into the target digital object, so that the target digital object is driven by the target audio
- the second driver video below.
- a digital object generation device comprising: an acquisition module, used to acquire target audio and a target digital object; a prediction module, used to The voice-driven model is used to obtain the first driving video of the reference digital object driven by the target audio; wherein, the voice-driven model is obtained based on the audio and video training of the reference digital object; the pose migration module is used to transfer the The pose of the reference digital object in each video frame of the first driving video is migrated to the target digital object to obtain a second driving video of the target digital object driven by the target audio.
- the prediction module is used to obtain the first driving video of the reference digital object driven by the target audio according to the target audio and the pre-trained voice-driven model of the reference digital object, and is specifically used to: Extracting phonemes from the target audio to obtain phoneme time stamps; inputting the phoneme time stamps into the speech-driven model to obtain the first driving video.
- the voice-driven model is trained based on the audio and video of the reference digital object, comprising: obtaining the audio of the reference digital object, and the audio of the reference digital object synchronized with the audio of the reference digital object Video; extracting phonemes from the audio of the reference digital object to obtain the phoneme time stamp of the reference digital object; using the phoneme time stamp of the reference digital object as a training sample, and using the video of the reference digital object as a sample label , training to obtain the speech-driven model.
- the pose transfer module is used to transfer the pose of the reference digital object in each video frame of the first driving video to the target digital object, so as to obtain the position of the target digital object in the target digital object.
- the second driving video is driven by audio, it is specifically used to: extract the texture features related to the texture of the target digital object from the target image of the target digital object; for each video frame of the first driving video , extracting pose features related to the pose of the reference digital object from the video frame; reconstructing the target digital object according to the texture features and the pose features, wherein the pose of the reconstructed target digital object is the same as
- the poses of the reference digital objects in the video frames are consistent; obtaining a reconstructed frame corresponding to the video frame according to the reconstructed target digital object; and constructing the second driving video based on the reconstructed frame.
- the pose migration module when used to extract pose features related to the pose of the reference digital object from the video frame, it is specifically used to: utilize the self-encoder of the reference digital object An encoder extracts the pose feature from the video frame, wherein the autoencoder of the reference digital object is pre-trained based on a plurality of images containing the reference digital object.
- the pose migration module when used to extract texture features related to the texture of the target digital object from the target image of the target digital object, it is specifically used to: use the target digital object
- the decoder in the autoencoder extracts the texture feature from a target image of the target digital object, wherein the autoencoder of the target digital object is pre-trained based on a plurality of images containing the target digital object.
- the pose migration module when used to reconstruct the target digital object according to the texture feature and the pose feature, it is specifically used to: use the decoder in the autoencoder of the target digital object according to the The texture feature and the pose feature are used to reconstruct the target digital object.
- the digital object generation device before reconstructing the target digital object according to the texture feature and the pose feature, is further configured to: detect that the pose feature extracted from the current frame is different from the pose feature from the If the pose features extracted from the previous frame of the current frame are consistent, the reconstructed frame corresponding to the previous frame is used as the reconstructed frame corresponding to the current frame.
- the pose migration module when used to obtain the reconstructed frame corresponding to the video frame according to the reconstructed target digital object, it is specifically configured to: according to the reconstructed target digital object and The set background material generates the reconstructed frame.
- the pose of the reference digital object includes one or more of the following: facial movements of the reference digital object, facial expressions of the reference digital object, and body movements of the reference digital object.
- an apparatus for generating a digital object comprising: an acquisition module, configured to acquire target audio and a target digital object; a prediction module, configured to Referring to the voice-driven model of the digital object, the first driving video of the reference digital object driven by the target audio is obtained; wherein, the voice-driven model is obtained based on the audio and video training of the reference digital object; the gesture migration module uses By transferring the pose of the reference digital object in each video frame of the first driving video to the target digital object, a second driving video of the target digital object driven by the target audio is obtained.
- an electronic device the electronic device includes a processor, a memory, and computer instructions stored in the memory that can be executed by the processor, and the processor executes the computer Instructions, the method mentioned in the first aspect above can be implemented.
- a computer-readable storage medium on which a computer program is stored, and when the computer program is executed, the method mentioned in the above-mentioned first aspect is implemented.
- the driving video of the reference digital object driven by the target audio can be obtained first, and for each video frame in the driving video, the The reference digital object in each video frame is replaced with the target digital object, and the pose of the reference digital object is transferred to the target digital object, so as to obtain the driving video of the target digital object driven by the target audio.
- FIG. 1 is a schematic diagram of training a speech-driven model according to an embodiment of the present disclosure.
- Fig. 2(a) is a flowchart of a method for generating a digital object according to an embodiment of the present disclosure.
- Fig. 2(b) is a schematic diagram of a method for generating a digital object according to an embodiment of the present disclosure.
- Fig. 2(c) is a flowchart of a method for generating a digital object according to an embodiment of the present disclosure.
- Fig. 3 is a schematic diagram of using an autoencoder to generate a second driving video of anchor A driven by target audio according to an embodiment of the present disclosure.
- Fig. 4 is a schematic structural diagram of a digital object generating device according to an embodiment of the present disclosure.
- Fig. 5 is a schematic structural diagram of a digital object generation device according to an embodiment of the present disclosure.
- Fig. 6 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure.
- first, second, third, etc. may be used in the present disclosure to describe various information, the information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of the present disclosure, first information may also be called second information, and similarly, second information may also be called first information. Depending on the context, the word “if” as used herein may be interpreted as “at” or “when” or “in response to a determination.”
- Digital objects are widely used in live broadcast, news broadcast, voice prompt and other fields. Usually, it is necessary to drive the digital object based on the audio to be broadcast to perform actions, expressions, etc. that match the audio, so as to obtain the audio-driven video. In related technologies, it is generally necessary to obtain audio and video of a large number of digital objects first, and then pre-train each digital object to obtain a voice-driven model of the digital object. For example, as shown in Figure 1, for a certain digital object A, a large amount of audio of the digital object A can be obtained, as well as a video of the digital object A that matches the audio (it can also be said that the action expression in the video is the same as the audio).
- synchronous may also be referred to as a video synchronized with the audio in this disclosure), and then use the audio as a training sample and the video as a sample label to train the neural network model to obtain the voice-driven model of the digital object A.
- different audio can be input into the voice-driven model, that is, the video of the digital object A driven by different audio can be output.
- an embodiment of the present disclosure provides a method for generating a digital object.
- the driving video of the reference digital object driven by the target audio can be obtained.
- the driving video For each video frame in , the reference digital object in each video frame can be replaced by the target digital object, and the posture of the reference digital object can be transferred to the target digital object to obtain the driving video of the target digital object driven by the target audio.
- the digital object generation method provided by the embodiments of the present disclosure can be executed by various electronic devices, such as mobile phones, notebook computers, tablets, cloud servers or server clusters, etc., which are not limited by the embodiments of the present disclosure.
- the reference digital object in the embodiments of the present disclosure can be a human figure, for example, it can be a human face, or it can also be a character including the whole human body.
- the reference digital object can also be some animal image, or virtual image, for example, it can be Various designs of cartoon images, etc., are not limited in the embodiments of the present disclosure, as long as a large amount of audio and video data of the reference digital object can be obtained for training the voice-driven model of the reference digital object.
- the target digital object in the embodiments of the present disclosure may also be a character image, an animal image, a virtual image, and the like.
- the target digital object and the reference digital object should preferably be the same type of image, for example, both are faces, or both are the whole body, or both are It's an animal image.
- the video of the reference digital object driven by the target audio is called the first driving video
- the video of the target digital object driven by the target audio is called the second driving video
- the digital object generation method provided by the embodiment of the present disclosure may include the following steps: S202. Obtain the first driving video of the reference digital object driven by the target audio; S204. Extract The gesture feature of the reference digital object in the first driving video; S206. Replace the reference digital object in the first driving video with a target digital object, and load the gesture feature into the target digital object , to obtain a second driving video of the target digital object driven by the target audio.
- the second driving video of the target digital object driven by the target audio can be obtained by means of the first driving video.
- the target audio may be a piece of voice data that the user wants to broadcast or broadcast live
- the target digital object is a digital object that the user wants to broadcast the target audio. For example, if the user wishes to broadcast a certain piece of content through the image of anchor A, the voice data corresponding to the content is the target audio, and the image of anchor A is the target digital object.
- the first driving video of the reference digital object driven by the target audio may be obtained, and the first driving video may be a video of the reference digital object driven by the target audio obtained in advance based on various methods.
- step S204 after the first driving video is acquired, for each video frame in the first driving video, the gesture feature of the reference digital object in the video frame can be extracted, wherein the gesture feature can be the gesture of the reference digital object , including related features such as facial movements, facial expressions, or body movements (postures will be explained in detail in the embodiments below).
- step S206 for each video frame in the first driving video, the reference digital object in each video frame can be replaced by the target digital object, and the gesture feature extracted in the video frame is loaded onto the target digital object, thus, the second driving video of the target digital object driven by the target audio is obtained.
- a driving video of the target digital object driven by the target audio may also be obtained.
- the first driving video of the reference digital object driven by the target audio can be obtained based on a pre-trained speech driving model, and the speech driving model can be obtained based on audio and video training of the reference digital object.
- FIG. 2(b) it is a schematic diagram of a method for generating a digital object according to an embodiment of the present disclosure
- Figure 2(c) a flowchart of a method for generating a digital object according to an embodiment of the present disclosure, The method may include the following steps.
- target audio and target digital objects can be acquired.
- the acquired target digital object can be one or more frames of images containing the object, and can also be some characteristic data (for example, shape features) representing the target digital object, based on these characteristic data, the target digital object can be obtained , the embodiments of the present disclosure are not limited.
- the target audio can be input into the pre-trained speech-driven model.
- a first driving video of the reference digital object driven by the target audio is determined by using the voice driving model.
- the posture of the reference digital object in each video frame of the first driving video is synchronized with the target audio, that is to say, the posture of the reference digital object, such as facial expression, demeanor, or action, matches the sound content in the target audio.
- the voice-driven model can be obtained by using audio and video training of reference digital objects.
- the audio of the reference digital object can be used as a training sample and input into the neural network model (any neural network model that meets the training purpose can be used, such as the generation of confrontation network GAN, which is not limited here), and the Videos are used as sample labels to train a speech-driven model for this reference digital object.
- the neural network model any neural network model that meets the training purpose can be used, such as the generation of confrontation network GAN, which is not limited here
- the Videos are used as sample labels to train a speech-driven model for this reference digital object.
- the pose of the reference digital object in each video frame of the first driving video can be transferred to the target digital object, so that the target digital object can be obtained in the target audio Driven under the second drive video.
- the pose of the reference digital object in the video frame can be transferred to the target digital object, so that the video frame A' can be obtained, and the video frame A' is the A video frame of , the pose of the target digital object in the video frame A' is consistent with the pose of the reference digital object in the video frame A.
- the image containing the target digital object and the video frame of the first driving video can be input into the pre-training
- the neural network model can automatically output the digital object after pose transfer.
- the pose-related information of the reference digital object in the video frame of the first driving video can also be extracted through a model (for example, an autoencoder), and then the pose-related information is applied to the target digital object to obtain the target digital object in this pose. object.
- the feature points in face A can be extracted, the distance between each feature point can be determined, and then the feature points in face B can be extracted.
- the distance between the feature points in A and the size ratio between face A and face B adaptively adjust the distance between the corresponding feature points in face B, so that the pose of face B and the pose of face A maintain unanimous.
- the embodiment of the present disclosure trains a voice-driven model of a reference digital object, and then reuses the voice-driven model for other new images, so as to quickly customize the video of the new image driven by the target audio.
- the posture of the reference digital object may include but not limited to one or more of the following: some facial movements of the reference digital object, such as opening the mouth, closing eyes, etc., may also include some facial movements of the reference digital object Expressions, such as happy, angry, surprised, etc., or may also include some body movements of reference digital objects, such as head movements (head up, head down, etc.), hand movements (various gestures, etc.), body movements (turning , bending over, etc.).
- the audio of the digital object is usually directly input into the model as a training sample, and then the video synchronized with the audio is used as a sample label, based on the model-predicted video
- the difference between the pose of the digital object in the video and the pose of the digital object in the actual video synchronized with the audio is continuously adjusted to model parameters to obtain a trained voice-driven model.
- the timbre may interfere with the final recognition result, resulting in the same sentence, if it is recorded by different people. Recording, the posture corresponding to the output sentence will be different.
- the audio of the reference digital object and the reference number synchronized with the audio can be obtained
- the video of the object, then the phoneme can be extracted from the audio of the reference digital object, and the phoneme time stamp of the reference digital object corresponding to the audio is obtained, and then the phoneme time stamp of the reference digital object is used as a training sample, and the reference digital object
- the videos are used as sample labels to train the speech-driven model.
- the phoneme timestamp indicates the time position of the phone in the audio.
- phonemes can also be first extracted from the target audio to obtain the The phoneme time stamp corresponding to the target audio, and then input the phoneme time stamp into the voice-driven model, and use the voice-driven model to predict the first driving video, thereby eliminating the influence of timbre and improving the accuracy of the voice-driven model prediction results.
- the pose of the reference digital object in each video frame of the first driving video is transferred to the target digital object to obtain a second driving video of the target digital object driven by the target audio
- the target image of several frames of the target digital object can be obtained first, and the texture features related to the texture of the target digital object can be extracted from the target image of the target digital object.
- the texture features are mainly related to the shape and surface of the target digital object. Some features of the target digital object can be obtained based on these features.
- a gesture feature related to the gesture of the reference digital object can be extracted from the video frame, wherein the gesture feature is, for example, related to the movement, expression, expression, etc.
- the current pose of the reference digital object can be determined based on the pose feature.
- the target digital object can be reconstructed according to the texture features of the target digital object and the pose features of the reference digital object, wherein the pose of the reconstructed target digital object is consistent with the pose of the reference digital object in the video frame.
- the reconstructed frame corresponding to the video frame can be obtained according to the reconstructed target digital object. For each video frame of the first driving video, the above method can be used to obtain the reconstructed frame corresponding to each video frame, and then the second driving video of the target digital object driven by the target audio can be obtained.
- the pose feature of the reference digital object or the texture feature of the target digital object can be extracted in various ways, for example, it can be extracted through a pre-trained model, or other ways can also be used, which are not limited in the embodiments of the present disclosure.
- the texture features of the target digital object can be extracted, only a small number of images of the target digital object are needed. Compared with the voice-driven model for training a target digital object, the amount of image data required is greatly reduced, and it can be realized when it is impossible. In the scene where a large number of target digital object audio and video are obtained, the driving video driven by the target audio is customized for the target digital object.
- the step of reconstructing the target digital object is performed , which may not only waste processing resources, but also reduce processing efficiency.
- the extracted data from the current frame can be detected first. Whether the pose feature is consistent with the pose feature extracted from the previous frame of the current frame. If they are consistent, the reconstructed frame corresponding to the previous frame can be directly used as the reconstructed frame corresponding to the current frame without the need to reconstruct the target number object process to improve processing efficiency.
- the background material of the video frame of the first driving video can be directly used as the background material of the target digital object , to get the reconstructed frame.
- the background material can also be pre-set according to actual needs, and then a reconstructed frame can be generated according to the reconstructed target digital object and the pre-set background material.
- the background material in different reconstructed frames can be the same or different, for example, the background material can be dynamically changed based on the pose of the target digital object in the reconstructed frame, or based on the target audio synchronized with the reconstructed frame Select the appropriate background material for the sound content.
- the self-encoder includes two parts: an encoder and a decoder.
- the encoder can extract (such as by compressing the input data) a potential spatial representation from the input, and the decoder can reconstruct the input based on the potential spatial representation. Therefore, multi-frame images of the reference digital object can be used as input to train the autoencoder so that its output and input are as consistent as possible, thereby training the autoencoder of the reference digital object.
- the pose features of the reference digital object can then be extracted from each video frame of the first driving video using the encoder in the self-encoder of the reference digital object.
- the encoder in the self-encoder of the reference digital object extracts the gesture features of the reference digital object, so that the extracted gesture features can be combined with the texture features of the digital object extracted by any decoder to reconstruct the digital object.
- the autoencoder of the digital object when the texture features related to the texture of the target digital object are extracted from the image of the target digital object, multiple frames of images of the target digital object can be obtained, and the target digital object can be trained based on the multiple frames of images of the target digital object.
- the autoencoder of the digital object then utilizes the decoder in the autoencoder of the target digital object to extract texture features from the target image of the target digital object.
- the autoencoder of the target digital object also includes an encoder and a decoder.
- the encoder can extract the potential spatial representation of the input image, and the decoder reconstructs the target digital object based on the spatial representation.
- the decoder can learn information related to the texture of the target digital object, and obtain the texture features of the target digital object.
- the texture features of the target digital object are extracted through the decoder in the self-encoder of the target digital object, so that the extracted texture features can be combined with the posture features of the digital object extracted by any encoder to reconstruct the digital object.
- the decoder in the autoencoder of the target digital object may be used to reconstruct the target digital object according to the pose feature and the texture feature. Since the decoder of the target digital object has learned the texture-related information of the target digital object during the training process, the pose features of the reference digital object can be input into the decoder, and the decoder can use the information it has learned. The texture feature of the target digital object and the pose feature are reconstructed to obtain a target digital object whose pose is consistent with that of the reference digital object.
- the voice-driven model of the image of the anchor B can be pre-trained. After obtaining a large number of audios of anchor B and videos synchronized with the audio, phonemes are extracted from the audios of anchor B, and phoneme time stamps corresponding to the audios are obtained. The phoneme time stamp is used as a training sample, and the video synchronized with the audio is used as a sample label to train a preset neural network model to obtain the voice-driven model.
- multiple frames of images of anchor B’s image can be obtained, and the self-encoder of anchor B’s image can be obtained by using the multi-frame images of anchor B’s image; Multi-frame image training to get the self-encoder of the anchor A image.
- these two kinds of self-encoders include an encoder and a decoder. After the image is input to the self-encoder, the encoder can extract the features of the image of the anchor in the image, and the decoder can reconstruct the image of the anchor based on the extracted features. . Therefore, the image of each anchor image can be used as input, so that the image output by the autoencoder is as consistent as possible with the input image, and the autoencoder is trained.
- the trained anchor B image encoder can be used to extract the posture-related gesture features of the anchor B image, such as facial movements, demeanor, and expressions, and the anchor A image decoder can generate corresponding actions and expressions based on the gesture-related gesture features. , The image of anchor A with a demeanor.
- the video frame can be first input into the encoder of the image of anchor B, and the encoder of image of anchor B can be used to extract the action, Posture-related gesture features such as expression and demeanor, and then input the gesture features into the decoder of the image of anchor A, and the decoder of the image of anchor A can be based on the texture features of the image of anchor A learned in advance and the input posture Features Reconstruct the image of anchor A, the posture of the reconstructed image of anchor A is consistent with the posture of the image of anchor B in the video frame, and the reconstructed frame corresponding to the video frame can be obtained by combining the preset background material.
- the above steps can be performed, so that the second driving video of the image of anchor A driven by the target audio can be obtained.
- audio has a high degree of freedom, and a large number of samples are often required to cover various phoneme combinations in speech. Therefore, if you want to train a speech-driven model for a new image, you often need a large number of training samples, that is, a large number of audio and video for the new image. Data, this method is not suitable for scenes where a large amount of audio and video cannot be obtained for a new image.
- the degree of freedom of human faces is relatively low, and all the movements of all faces can be basically covered by a small number of images. Therefore, training an autoencoder for a new image often requires only a small number of images, such as a few frames or dozens of frames. Compared with training a speech-driven model, not only the training samples that need to be consumed are greatly reduced, but also the efficiency of generating digital object videos of new images can be greatly improved.
- the embodiment of the present disclosure also provides a digital object generating device 40, as shown in FIG.
- a first driving video of the reference digital object driven by the target audio is obtained; wherein the voice driving model is based on the audio and video of the reference digital object Obtained by training; the attitude transfer module 43 is used to transfer the attitude of the reference digital object in each video frame of the first driving video to the target digital object, and obtain the target digital object in the target audio Driven under the second drive video.
- the prediction module 42 is used to obtain the first driving video of the reference digital object driven by the target audio according to the target audio and the pre-trained voice-driven model of the reference digital object, specifically for: Extracting phonemes from the target audio to obtain phoneme time stamps; inputting the phoneme time stamps into the speech-driven model to obtain the first driving video.
- the voice-driven model is trained based on the audio and video of the reference digital object, including: acquiring the audio of the reference digital object, and the video of the reference digital object synchronized with the audio of the reference digital object ; Extract phonemes from the audio of the reference digital object to obtain the phoneme timestamp of the reference digital object; use the phoneme timestamp of the reference digital object as a training sample, and use the video of the reference digital object as a sample label,
- the speech-driven model is obtained through training.
- the pose transfer module 43 is configured to transfer the pose of the reference digital object in each video frame of the first driving video to the target digital object, so as to obtain the position of the target digital object in the
- the second driving video is driven by the target audio, it is specifically used to: extract the texture features related to the texture of the target digital object from the target image of the target digital object; frame, extracting pose features related to the pose of the reference digital object from the video frame; reconstructing the target digital object according to the texture features and the pose features, wherein the pose of the reconstructed target digital object Consistent with the posture of the reference digital object in the video frame; obtaining a reconstructed frame corresponding to the video frame according to the reconstructed target digital object; and constructing the second driving video based on the reconstructed frame .
- the pose migration module 43 when used to extract pose features related to the pose of the reference digital object from the video frame, it is specifically used for: using the self-encoder of the reference digital object
- the encoder extracts the pose features from the video frames, wherein the autoencoder of the reference digital object is pre-trained based on a plurality of images containing the reference digital object.
- the pose transfer module 43 when configured to extract texture features related to the texture of the target digital object from the target image of the target digital object, it is specifically configured to: utilize the target digital object
- the decoder in the autoencoder extracts the texture feature from a target image of the target digital object, wherein the autoencoder of the target digital object is pre-trained based on a plurality of images containing the target digital object.
- the pose migration module 43 when used to reconstruct the target digital object according to the pose feature and the texture feature, it is specifically used to: use the decoder in the autoencoder of the target digital object according to The pose features and the texture features reconstruct a target digital object.
- the digital object generation device before reconstructing the target digital object according to the pose feature and the texture feature, is further configured to: detect that the pose feature extracted from the current frame is different from the If the pose features extracted from the previous frame of the current frame are consistent, the reconstructed frame corresponding to the previous frame is used as the reconstructed frame corresponding to the current frame.
- the pose migration module 43 when used to obtain the reconstructed frame corresponding to the video frame according to the reconstructed target digital object, it is specifically configured to: according to the reconstructed target digital object and The preset background material generates the reconstructed frame.
- the pose of the reference digital object includes one or more of the following: facial movements of the reference digital object, facial expressions of the reference digital object, and body movements of the reference digital object.
- the embodiment of the present disclosure also provides another digital object generation device 50.
- the device 50 includes: an acquisition module 51, configured to acquire the first driving video of the reference digital object driven by the target audio
- the extraction module 52 is used to extract the gesture features of the reference digital object in the first driving video
- the loading module 53 is used to replace the reference digital object in the first driving video with a target digital object
- the gesture feature is loaded to the target digital object to obtain a second driving video of the target digital object driven by the target audio.
- the specific implementation steps of the digital object generating device 50 generating the target digital object driven by the target audio can refer to the descriptions in the above-mentioned method embodiments, and will not be repeated here.
- an embodiment of the present disclosure also provides an electronic device 60. As shown in FIG. instructions, and the processor 61 implements the method described in any one of the above embodiments when executing the computer instructions.
- An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any one of the foregoing embodiments is implemented.
- Computer-readable media includes both volatile and non-volatile, removable and non-removable media, and can be implemented by any method or technology for storage of information.
- Information may be computer readable instructions, data structures, modules of a program, or other data.
- Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Flash memory or other memory technology, Compact Disc Read-Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical storage, Magnetic tape cartridge, tape magnetic disk storage or other magnetic storage device or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
- computer-readable media excludes transitory computer-readable media, such as modulated data signals and carrier waves.
- a typical implementing device is a computer, which may take the form of a personal computer, laptop computer, cellular phone, camera phone, smart phone, personal digital assistant, media player, navigation device, e-mail device, game control device, etc. desktops, tablets, wearables, or any combination of these.
- each embodiment in this specification is described in a progressive manner, the same and similar parts of each embodiment can be referred to each other, and each embodiment focuses on the differences from other embodiments.
- the description is relatively simple, and for relevant parts, please refer to part of the description of the method embodiment.
- the device embodiments described above are only illustrative, and the modules described as separate components may or may not be physically separated, and the functions of each module may be integrated in the same or multiple software and/or hardware implementations. Part or all of the modules can also be selected according to actual needs to achieve the purpose of the solution of this embodiment. It can be understood and implemented by those skilled in the art without creative effort.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Theoretical Computer Science (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Human Computer Interaction (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Multimedia (AREA)
- Acoustics & Sound (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- General Physics & Mathematics (AREA)
- Biophysics (AREA)
- Computing Systems (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Molecular Biology (AREA)
- Evolutionary Computation (AREA)
- Data Mining & Analysis (AREA)
- Biomedical Technology (AREA)
- Artificial Intelligence (AREA)
- Life Sciences & Earth Sciences (AREA)
- Processing Or Creating Images (AREA)
Abstract
本说明书实施例提供数字对象生成方法、装置及设备。根据本说明书中的一实施例,可以获取目标音频和目标数字对象;根据所述目标音频和预先训练的参考数字对象的语音驱动模型,得到所述参考数字对象在所述目标音频驱动下的第一驱动视频;其中,所述语音驱动模型基于所述参考数字对象的音频和视频训练得到;将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到所述目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。通过上述方法,可以快速定制一个新形象在目标音频驱动下的视频。
Description
相关申请的交叉引用
本公开要求于2021年12月24日提交的、申请号为202111599988.4的中国专利申请的优先权,该申请的全文以引用的方式并入本文中。
本公开涉及人工智能技术领域,尤其涉及数字对象生成方法、装置、设备及存储介质。
数字对象(例如,数字人)广泛应用于直播、新闻播报、语音提示等领域。通常需要基于想要播报的音频驱动数字对象做出和该音频匹配的动作、表情等,得到该音频驱动的视频。相关技术中,一般需要先获取大量数字对象的音频和视频(以下简称为音视频),然后针对每个数字对象预先训练得到该数字对象的语音驱动模型,在得到该数字对象的语音驱动模型后,可以将不同的音频输入到该语音驱动模型中,即可以输出不同音频驱动下的该数字对象的视频。但是这种方式只能得到一个固定的数字对象在不同音频驱动下的视频,如果要换成一个新的数字对象(例如,从数字人换成数字动物),则需重新获取大量新数字对象的音视频,利用新数字对象的音视频重新训练新数字对象的语音驱动模型,整个训练过程需要的数据量大,耗时较长,导致无法快速生成新数字对象在某个音频驱动下的视频。
发明内容
本公开提供一种数字对象生成方法、装置、设备及存储介质。
根据本公开实施例的第一方面,提供一种数字对象生成方法,所述方法包括:获取目标音频和目标数字对象;根据所述目标音频和预先训练的参考数字对象的语音驱动模型,得到参考数字对象在所述目标音频驱动下的第一驱动视频;其中,所述语音驱动模型基于所述参考数字对象的音频和视频训练得到;将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到所述目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
在一些实施例中,根据目标音频和预先训练的参考数字对象的语音驱动模型,得到 参考数字对象在所述目标音频驱动下的第一驱动视频,包括:从所述目标音频中提取音素,得到音素时间戳;将所述音素时间戳输入到所述语音驱动模型中,得到所述第一驱动视频。
在一些实施例中,所述语音驱动模型基于所述参考数字对象的音频和视频训练得到,包括:获取所述参考数字对象的音频,以及与所述参考数字对象的音频同步的参考数字对象的视频;从所述参考数字对象的音频中提取音素,得到所述参考数字对象的音素时间戳;以所述参考数字对象的音素时间戳作为训练样本,以所述参考数字对象的视频作为样本标签,训练得到所述语音驱动模型。
在一些实施例中,将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频,包括:从所述目标数字对象的目标图像中提取得到与所述目标数字对象的纹理相关的纹理特征;针对所述第一驱动视频的各视频帧,从所述视频帧中提取与所述参考数字对象的姿态相关的姿态特征;根据所述纹理特征和所述姿态特征重新构建目标数字对象,其中,所述重新构建的目标数字对象的姿态与所述视频帧中的参考数字对象的姿态一致;根据所述重新构建的目标数字对象得到与所述视频帧对应的重构帧以及基于所述重构帧来构造所述第二驱动视频。
在一些实施例中,从所述视频帧中提取与所述参考数字对象的姿态相关的姿态特征,包括:利用所述参考数字对象的自编码器中的编码器从所述视频帧中提取所述姿态特征,其中,所述参考数字对象的自编码器基于包含所述参考数字对象的多个图像预先训练得到。
在一些实施例中,从所述目标数字对象的目标图像中提取得到与所述目标数字对象的纹理相关的纹理特征,包括:利用所述目标数字对象的自编码器中的解码器,从所述目标数字对象的目标图像中提取所述纹理特征,其中,所述目标数字对象的自编码器基于包含所述目标数字对象的多个图像预先训练得到。
在一些实施例中,根据所述纹理特征和所述姿态特征重新构建目标数字对象,包括:利用所述目标数字对象的自编码器中的解码器根据所述纹理特征和所述姿态特征重新构建目标数字对象。
在一些实施例中,在根据所述纹理特征和所述姿态特征重新构建目标数字对象之前,还包括:在检测到从当前帧提取的所述姿态特征与从所述当前帧的前一帧提取的所述姿 态特征一致的情况下,将所述前一帧对应的重构帧作为所述当前帧对应的重构帧。
在一些实施例中,根据所述重新构建的目标数字对象得到与所述视频帧对应的重构帧,包括:根据所述重新构建的目标数字对象和预先设置的背景素材生成所述重构帧。
在一些实施例中,所述参考数字对象的姿态包括以下一种或多种:所述参考数字对象的脸部动作、所述参考数字对象的脸部表情、所述参考数字对象的肢体动作。
根据本公开实施例的第二方面,提供一种数字对象生成方法,所述方法包括:获取参考数字对象在目标音频驱动下的第一驱动视频;提取所述参考数字对象在所述第一驱动视频中的姿态特征;将所述第一驱动视频中的参考数字对象替换为目标数字对象,并将所述姿态特征加载到所述目标数字对象,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
根据本公开实施例的第三方面,提供一种数字对象生成装置,所述装置包括:获取模块,用于获取目标音频和目标数字对象;预测模块,用于根据所述目标音频和预先训练的语音驱动模型,得到参考数字对象在所述目标音频驱动下的第一驱动视频;其中,所述语音驱动模型基于所述参考数字对象的音频和视频训练得到;姿态迁移模块,用于将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到所述目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
在一些实施例中,所述预测模块用于根据目标音频和预先训练的参考数字对象的语音驱动模型,得到参考数字对象在所述目标音频驱动下的第一驱动视频时,具体用于:从所述目标音频中提取音素,得到音素时间戳;将所述音素时间戳输入到所述语音驱动模型中,得到所述第一驱动视频。
在一些实施例中,所述语音驱动模型基于所述参考数字对象的音频和视频训练得到,包括:获取所述参考数字对象的音频,以及与所述参考数字对象的音频同步的参考数字对象的视频;从所述参考数字对象的音频中提取音素,得到所述参考数字对象的音素时间戳;以所述参考数字对象的音素时间戳作为训练样本,以所述参考数字对象的视频作为样本标签,训练得到所述语音驱动模型。
在一些实施例中,所述姿态迁移模块用于将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频时,具体用于:从所述目标数字对象的目标图像中提取得到与所述目标数字对象的纹理相关的纹理特征;针对所述第一驱动视频的各视频帧,从所述视频 帧中提取与所述参考数字对象的姿态相关的姿态特征;根据所述纹理特征和所述姿态特征重新构建目标数字对象,其中,所述重新构建的目标数字对象的姿态与所述视频帧中的参考数字对象的姿态一致;根据所述重新构建的目标数字对象得到与所述视频帧对应的重构帧;以及基于所述重构帧来构造所述第二驱动视频。
在一些实施例中,所述姿态迁移模块用于从所述视频帧中提取与所述参考数字对象的姿态相关的姿态特征时,具体用于:利用所述参考数字对象的自编码器中的编码器从所述视频帧中提取所述姿态特征,其中,所述参考数字对象的自编码器基于包含所述参考数字对象的多个图像预先训练得到。
在一些实施例中,所述姿态迁移模块用于从所述目标数字对象的目标图像中提取得到与所述目标数字对象的纹理相关的纹理特征时,具体用于:利用所述目标数字对象的自编码器中的解码器,从所述目标数字对象的目标图像中提取所述纹理特征,其中,所述目标数字对象的自编码器基于包含所述目标数字对象的多个图像预先训练得到。
在一些实施例中,所述姿态迁移模块用于根据所述纹理特征和所述姿态特征重新构建目标数字对象时,具体用于:利用所述目标数字对象的自编码器中的解码器根据所述纹理特征和所述姿态特征重新构建目标数字对象。
在一些实施例中,在根据所述纹理特征和所述姿态特征重新构建目标数字对象之前,所述数字对象生成装置还用于:在检测到从当前帧提取的所述姿态特征与从所述当前帧的前一帧提取的所述姿态特征一致的情况下,将所述前一帧对应的重构帧作为所述当前帧对应的重构帧。
在一些实施例中,所述姿态迁移模块用于根据所述重新构建的目标数字对象得到与所述视频帧对应的重构帧时,具体用于:根据所述重新构建的目标数字对象和预先设置的背景素材生成所述重构帧。
在一些实施例中,所述参考数字对象的姿态包括以下一种或多种:所述参考数字对象的脸部动作、所述参考数字对象的脸部表情、所述参考数字对象的肢体动作。
根据本公开实施例的第四方面,提供一种数字对象生成装置,所述装置包括:获取模块,用于获取目标音频和目标数字对象;预测模块,用于根据所述目标音频和预先训练的参考数字对象的语音驱动模型,得到参考数字对象在所述目标音频驱动下的第一驱动视频;其中,所述语音驱动模型基于所述参考数字对象的音频和视频训练得到;姿态迁移模块,用于将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到所 述目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
根据本公开实施例的第五方面,提供一种电子设备,所述电子设备包括处理器、存储器、存储在所述存储器可供所述处理器执行的计算机指令,所述处理器执行所述计算机指令时,可实现上述第一方面提及的方法。
根据本公开实施例的第六方面,提供一种计算机可读存储介质,所述存储介质上存储有计算机程序,所述计算机程序被执行时实现上述第一方面提及的方法。
本公开实施例中,当想要生成一段目标音频驱动目标数字对象的驱动视频时,可以先获取参考数字对象在该目标音频驱动下的驱动视频,针对该驱动视频中的各视频帧,可以将各视频帧中的参考数字对象替换成目标数字对象,并将参考数字对象的姿态迁移到目标数字对象上,从而得到目标数字对象在该目标音频驱动下的驱动视频。通过这种方法,无需再专门训练目标数字对象的语音驱动模型,并且姿态迁移仅需少量目标数字对象的图像即可实现,可以适用于无法获取大量目标数字对象的音视频数据的场景,并且耗费时间较短,可以实现快速定制一个新形象在目标音频驱动下的视频。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性和解释性的,而非限制本公开。
此处的附图被并入说明书中并构成本说明书的一部分,这些附图示出了符合本公开的实施例,并与说明书一起用于说明本公开的技术方案。
图1是本公开实施例的一种训练语音驱动模型的示意图。
图2(a)是本公开实施例的一种数字对象生成方法的流程图。
图2(b)是本公开实施例的一种数字对象生成方法的示意图。
图2(c)是本公开实施例的一种数字对象生成方法的流程图。
图3是本公开实施例的一种利用自编码器生成主播A在目标音频驱动下的第二驱动视频的示意图。
图4是本公开实施例的一种数字对象生成装置的结构示意图。
图5是本公开实施例的一种数字对象生成装置的结构示意图。
图6是本公开实施例的一种电子设备的结构示意图。
这里将详细地对示例性实施例进行说明,其示例表示在附图中。下面的描述涉及附图时,除非另有表示,不同附图中的相同数字表示相同或相似的要素。以下示例性实施例中所描述的实施方式并不代表与本公开相一致的所有实施方式。相反,它们仅是与如所附权利要求书中所详述的、本公开的一些方面相一致的装置和方法的例子。
在本公开使用的术语是仅仅出于描述特定实施例的目的,而非旨在限制本公开。在本公开和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。还应当理解,本文中使用的术语“和/或”是指并包含一个或多个相关联的列出项目的任何或所有可能组合。另外,本文中术语“至少一种”表示多种中的任意一种或多种中的至少两种的任意组合。
应当理解,尽管在本公开可能采用术语第一、第二、第三等来描述各种信息,但这些信息不应限于这些术语。这些术语仅用来将同一类型的信息彼此区分开。例如,在不脱离本公开范围的情况下,第一信息也可以被称为第二信息,类似地,第二信息也可以被称为第一信息。取决于语境,如在此所使用的词语“如果”可以被解释成为“在……时”或“当……时”或“响应于确定”。
为了使本技术领域的人员更好的理解本公开实施例中的技术方案,并使本公开实施例的上述目的、特征和优点能够更加明显易懂,下面结合附图对本公开实施例中的技术方案作进一步详细的说明。
数字对象广泛应用于直播、新闻播报、语音提示等领域。通常需要基于想要播报的音频驱动数字对象做出和该音频匹配的动作、表情等,得到该音频驱动的视频。相关技术中,一般需要先获取大量数字对象的音视频,然后针对每个数字对象预先训练得到该数字对象的语音驱动模型。比如,如图1所示,针对某个数字对象A,可以获取大量该数字对象A的音频,以及与该音频匹配的数字对象A的视频(也可以说视频中的动作表情等与该音频是同步的,在本公开中也可以称为与该音频同步的视频),然后利用该音频作为训练样本,该视频作为样本标签,对神经网络模型进行训练,得到该数字对象A的语音驱动模型。在得到该数字对象A的语音驱动模型后,可以将不同的音频输入到该语音驱动模型中,即可以输出不同音频驱动下的该数字对象A的视频。
但是这种方式只能得到一个固定的数字对象在不同音频驱动下的视频,如果要换成一个新数字对象,则需重新获取大量新数字对象的音视频,利用新数字对象的音视频重 新训练新数字对象的语音驱动模型,整个训练过程需要的数据量大,耗时较长,导致无法快速生成新数字对象在某个音频驱动下的视频。比如,针对直播场景,如果用户希望快速为某个模特形象定制一个目标音频驱动下的视频,则需要获取该模特形象的大量音视频数据,重新训练一个该模特形象的语音驱动模型,对于无法获取该模特形象大量音视频数据的场景,则无法生成该模特形象的语音驱动模型,并且整个过程耗时较长,无法实现新形象的快速定制。
基于此,本公开实施例提供一种数字对象生成方法,当想要生成一段目标音频驱动目标数字对象的驱动视频时,可以获取参考数字对象在该目标音频驱动下的驱动视频,针对该驱动视频中的各视频帧,可以将各视频帧中的参考数字对象替换成目标数字对象,并将参考数字对象的姿态迁移到目标数字对象上,得到目标数字对象在该目标音频驱动下的驱动视频。通过这种方法,无需再专门训练目标数字对象的语音驱动模型,并且姿态迁移仅需少量目标数字对象的图像即可实现,可以适用于无法获取大量目标数字对象音视频数据的场景,并且耗费时间较短,可以实现快速定制一个新形象在目标音频驱动下的视频。
本公开实施例提供的数字对象生成方法可以通过各种电子设备执行,比如,可以是手机、笔记本电脑、平板、云端服务器或者服务器集群等等,本公开实施例不作限制。
本公开实施例中的参考数字对象可以是人物形象,比如,可以是人脸,或者也可以是包含整个人体的人物,当然参考数字对象也可以是一些动物形象,或者虚拟形象,比如,可以是各类设计的卡通形象等等,本公开实施例不做限制,只要可以获取该参考数字对象的大量音视频数据,用于训练该参考数字对象的语音驱动模型即可。
同样的,本公开实施例中的目标数字对象也可以是人物形象、动物形象、虚拟形象等。但是,为保证参考数字对象的姿态可以准确的迁移到目标数字对象上,目标数字对象和参考数字对象最好可以是同一类型的形象,比如,都是人脸,或者都是整个人体,或者都是某个动物形象。
为了便于区分,以下将参考数字对象在目标音频驱动下的视频称为第一驱动视频,目标数字对象在目标音频驱动下的视频称为第二驱动视频。
在一些实施例中,如图2(a)所示,本公开实施例提供的数字对象生成方法可以包括以下步骤:S202、获取参考数字对象在目标音频驱动下的第一驱动视频;S204、提取所述参考数字对象在所述第一驱动视频中的姿态特征;S206、将所述第一驱动视频中的 参考数字对象替换为目标数字对象,并将所述姿态特征加载到所述目标数字对象,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
由于已经有参考数字对象在目标音频驱动下的第一驱动视频,因而可以借助第一驱动视频得到目标数字对象在目标音频驱动下的第二驱动视频。其中,目标音频可以是用户想要播报或者直播的一段语音数据,目标数字对象为用户希望用于播报该目标音频的一个数字对象。比如,用户希望通过主播A形象来播报某段内容,该内容对应的语音数据即为目标音频,该主播A形象即为目标数字对象。
具体的,在步骤S202中,可以获取参考数字对象在目标音频驱动下的第一驱动视频,第一驱动视频可以是预先基于各种方式获得的参考数字对象在目标音频驱动下的视频。
在步骤S204中,在获取到第一驱动视频后可以针对第一驱动视频中的各视频帧,提取该视频帧中的参考数字对象的姿态特征,其中,姿态特征可以是与参考数字对象的姿态,包括脸部动作、脸部表情或者肢体动作等相关的特征(姿态将在下文实施例中详细解释)。
在步骤S206中,针对第一驱动视频中的各视频帧,可以将各视频帧中的参考数字对象替换为目标数字对象,并将该视频帧中提取到的姿态特征加载到目标数字对象上,从而得到目标数字对象在目标音频驱动下的第二驱动视频。
通过将第一驱动视频中各视频帧的参考数字对象替换成目标数字对象,然后将提取到的参考数字对象的姿态特征加载到目标数字对象上,可以无需专门训练目标数字对象的语音驱动模型,也可以获得目标数字对象在目标音频驱动下的驱动视频。在一些实施例中,参考数字对象在目标音频驱动下的第一驱动视频可以基于预先训练的语音驱动模型得到,该语音驱动模型可以基于参考数字对象的音视频训练得到。比如,如图2(b)所示,为本公开实施例的一种数字对象生成方法的示意图,如图2(c)所示,本公开实施例的一种数字对象生成方法的流程图,该方法可以包括以下步骤。
S302、获取目标音频和目标数字对象。
首先,可以获取目标音频和目标数字对象。其中,获取的目标数字对象可以是包含该对象的一帧或多帧图像,也可以是表征该目标数字对象的一些特征数据(比如,外形特征),基于这些特征数据即可以得到该目标数字对象,本公开实施例不做限制。
S304、根据所述目标音频和预先训练的参考数字对象的语音驱动模型,得到参考数字对象在所述目标音频驱动下的第一驱动视频;其中,所述语音驱动模型基于所述参考 数字对象的音视频训练得到。
在获取目标音频和目标数字对象后,可以将目标音频输入至预先训练的语音驱动模型中。利用该语音驱动模型确定该参考数字对象在目标音频驱动下的第一驱动视频。第一驱动视频的各视频帧中的参考数字对象的姿态与该目标音频同步,也就是说参考数字对象的姿态,如面部表情、神态、或者动作等和目标音频中的声音内容是匹配的。比如,当声音内容呈现比较欢快的语调,参考数字对象的表情也是微笑的表情等。其中,该语音驱动模型可以利用参考数字对象的音视频训练得到。比如,可以将参考数字对象的音频作为训练样本,输入到神经网络模型(可以使用任何满足训练目的的神经网络模型,如生成对抗网络GAN,在此不做限制)中,将与该音频同步的视频作为样本标签,训练得到该参考数字对象的语音驱动模型。
S306、将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到所述目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
在得到参考数字对象在目标音频驱动下的第一驱动视频后,可以将第一驱动视频的各视频帧中的参考数字对象的姿态迁移到目标数字对象上,从而可以得到目标数字对象在目标音频驱动下的第二驱动视频。比如,针对第一驱动视频中的视频帧A,可以将该视频帧中参考数字对象的姿态迁移到目标数字对象上,从而可以得到视频帧A’,该视频帧A’为第二驱动视频中的一个视频帧,该视频帧A’中目标数字对象的姿态和视频帧A中参考数字对象的姿态一致。
其中,将第一驱动视频的各视频帧中参考数字对象的姿态迁移到目标数字对象上的方式有很多,比如,可以将包含目标数字对象的图像和第一驱动视频的视频帧输入至预先训练的神经网络模型中,神经网络模型可以自动输出姿态迁移后的数字对象。或者,也可以通过模型(例如,自编码器)提取第一驱动视频的视频帧中参考数字对象的姿态相关信息,再将这些姿态相关信息作用到目标数字对象上,得到该姿态下的目标数字对象。
举个例子,参考数字对象为人脸A,目标数字对象为人脸B,可以提取人脸A中的特征点,确定各特征点之间的距离,然后提取人脸B中的特征点,基于人脸A中各特征点之间的距离以及人脸A和人脸B之间的大小比例适应性调整人脸B中相应特征点之间的距离,以便人脸B的姿态和人脸A的姿态保持一致。
本公开实施例通过训练一个参考数字对象的语音驱动模型,然后将该语音驱动模型 复用给其他的新形象,从而可以实现快速定制新形象在目标音频驱动下的视频。
在一些实施例中,参考数字对象的姿态可以包括但不限于以下一种或多种:参考数字对象的一些脸部动作,比如,张口、闭眼等,也可以包括参考数字对象的一些脸部表情,比如,开心、愤怒、惊讶等,或者也可以包括参考数字对象的一些肢体动作,比如,头部动作(仰头、低头等)、手部动作(各类手势等)、身体动作(转身、弯腰等)。
相关技术中,在训练某个数字对象的语音驱动模型时,通常是直接将该数字对象的音频作为训练样本输入到模型中,然后将与该音频同步的视频作为样本标签,基于模型预测的视频中该数字对象的姿态和实际的与该音频同步的视频中该数字对象的姿态的差异不断调整模型参数,得到训练后的语音驱动模型。但是,由于同一句话,如果由不同的人录制,其音色会存在差别,如果直接将音频作为训练样本,音色可能会对最终的识别结果造成干扰,导致同一句话,如果是由不同的人录制,输出的这句话对应的姿态会有差别。
所以,在一些实施例中,为了消除音色对模型识别结果的干扰,在根据参考数字对象的音视频训练语音驱动模型时,可以获取所述参考数字对象的音频,以及和该音频同步的参考数字对象的视频,然后可以从参考数字对象的音频中提取音素,得到与该段音频对应的参考数字对象的音素时间戳,然后以该参考数字对象的音素时间戳作为训练样本,以该参考数字对象的视频作为样本标签,训练得到语音驱动模型。音素时间戳指示了音素在音频中的时间位置。
在一些实施例中,在根据目标音频和预先训练的参考数字对象的语音驱动模型,得到参考数字对象在目标音频驱动下的第一驱动视频时,也可以先从目标音频中提取音素,得到与该目标音频对应的音素时间戳,然后将该音素时间戳输入到语音驱动模型中,利用语音驱动模型预测得到第一驱动视频,从而可以消除音色的影响,提高语音驱动模型预测结果的准确度。
在一些实施例中,在将第一驱动视频的各视频帧中的参考数字对象的姿态迁移到目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频时,可以先获取若干帧目标数字对象的目标图像,从目标数字对象的目标图像中提取得到与该目标数字对象的纹理相关的纹理特征,其中,纹理特征主要是与目标数字对象的外形、表面等相关的一些特征,基于这些特征则可以得到目标数字对象的外表的大体情况。然后可以针对第一驱动视频的每一帧视频帧,从该视频帧中提取与参考数字对象的姿态相关的姿态特征,其中,姿态特征是例如与参考数字对象的动作、神态、表情等有关的特征, 基于姿态特征即可以确定参考数字对象当前所处的姿态。然后可以根据目标数字对象的纹理特征以及参考数字对象的姿态特征重新构建目标数字对象,其中,重新构建的目标数字对象的姿态与视频帧中的参考数字对象的姿态一致。并且可以根据重新构建的目标数字对象得到与该视频帧对应的重构帧。针对第一驱动视频的每一帧视频帧,均可以采用上述方法,得到每一帧视频帧对应的重构帧,进而可以得到目标数字对象在目标音频驱动下的第二驱动视频。
其中,参考数字对象的姿态特征或目标数字对象的纹理特征的提取可以采用多种方式,比如,可以通过预先训练模型提取,或者也可以采用其他方式,本公开实施例不做限制。
由于提取目标数字对象的纹理特征,仅需少量的目标数字对象的图像即可以实现,相比于要训练一个目标数字对象的语音驱动模型,其需要的图像数据量大大减小,可以实现在无法获取大量目标数字对象音视频的场景下,为目标数字对象定制目标音频驱动下的驱动视频。
当然,对于第一驱动视频,由于大多数情况下,可能连续多帧视频帧中参考数字对象的姿态都保持不变,因而,如果针对每一帧视频帧,都执行重新构建目标数字对象的步骤,可能会既浪费处理资源,又降低处理效率。为了提高处理效率,在一些实施例中,在根据第一驱动视频的视频帧中的参考数字对象的姿态特征和目标数字对象的纹理特征重新构建目标数字对象之前,可以先检测从当前帧提取的姿态特征与从当前帧的前一帧提取的姿态特征是否一致,如果一致,则可直接将前一帧对应的重构帧作为该当前帧对应的重构帧,而无需再进行重构目标数字对象的过程,以提高处理效率。
在一些实施例中,在根据重新构建的目标数字对象得到与第一驱动视频的视频帧对应的重构帧时,可以直接将第一驱动视频的视频帧的背景素材作为目标数字对象的背景素材,得到重构帧。当然,也可以根据实际需求预先设置好背景素材,然后根据重新构建的目标数字对象和预先设置的背景素材生成重构帧。其中,不同的重构帧中的背景素材可以相同,也可以不同,比如,可以基于重构帧中目标数字对象的姿态去动态变化背景素材,或者也可以基于与该重构帧同步的目标音频的声音内容选取适配的背景素材。
在一些实施例中,在从所述视频帧中提取与参考数字对象的姿态相关的姿态特征时,可以先获取参考数字对象的多帧图像,利用参考数字对象的多帧图像训练得到参考数字对象的自编码器。其中,自编码器包括编码器和解码器两部分,编码器能从输入提取(如通过压缩输入数据)潜在的空间表征,而解码器可以基于该潜在的空间表征重构输入。 所以,可以利用参考数字对象的多帧图像作为输入,对自编码器进行训练,使得其输出和输入尽可能一致,从而训练得到参考数字对象的自编码器。然后可以利用参考数字对象的自编码器中的编码器从第一驱动视频的各视频帧中提取参考数字对象的姿态特征。通过参考数字对象的自编码器中的编码器提取参考数字对象的姿态特征,使得提取到的姿态特征可以结合任一解码器提取到的数字对象的纹理特征重构数字对象。
在一些实施例中,在从目标数字对象的图像中提取得到与目标数字对象的纹理相关的纹理特征时,可以获取目标数字对象的多帧图像,基于该目标数字对象的多帧图像训练得到目标数字对象的自编码器,然后利用该目标数字对象的自编码器中的解码器从目标数字对象的目标图像中提取纹理特征。其中,目标数字对象的自编码器也包括编码器和解码器两部分,编码器可以提取输入图像潜在的空间表征,解码器基于该空间表征重建目标数字对象。在训练过程中,解码器可以学习目标数字对象的纹理相关的信息,得到目标数字对象的纹理特征。通过目标数字对象的自编码器中的解码器提取目标数字对象的纹理特征,使得提取到的纹理特征可以结合任一编码器提取到的数字对象的姿态特征重构数字对象。
在一些实施例中,在根据姿态特征和纹理特征重新构建目标数字对象时,可以利用目标数字对象的自编码器中的解码器根据姿态特征和纹理特征重新构建目标数字对象。由于目标数字对象的解码器在训练过程中,已经学习到了目标数字对象的纹理相关的信息,因而,可以将参考数字对象的姿态特征输入到该解码器中,解码器即可以利用它学习到的目标数字对象的纹理特征和该姿态特征重新构建得到姿态与参考数字对象的姿态一致的目标数字对象。
为了进一步解释本公开实施例提供的数字对象生成方法,以下结合一个具体的实施例加以解释。
当用户无法获取主播A的大量音视频数据,但又希望可以利用主播A形象快速定制主播A形象在目标音频驱动下的视频时,可以采用以下方式实现。
(1)预先训练参考数字对象如主播B形象的语音驱动模型。
由于可以获取到大量主播B的音频以及和该音频同步的视频,因而,可以预先训练主播B形象的语音驱动模型。在获取到大量主播B的音频以及和该音频同步的视频后,从主播B的音频中提取音素,得到与该音频对应的音素时间戳。将该音素时间戳作为训练样本,将与该音频同步的视频作为样本标签,对预先设置的神经网络模型进行训练, 得到该语音驱动模型。
(2)预先训练主播A形象和主播B形象的自编码器。
如图3所示,可以获取多帧主播B形象的图像,利用主播B形象的多帧图像训练得到主播B形象的自编码器;同时可以获取多帧主播A形象的图像,利用主播A形象的多帧图像训练得到主播A形象的自编码器。其中,这两种自编码器均包括编码器和解码器,将图像输入到自编码器后,在编码器可以提取图像中的主播形象的特征,解码器能够基于提取的特征重构该主播形象。因而,可以利用各主播形象的图像作为输入,使得自编码器输出的图像和输入的图像尽可能一致,对自编码器进行训练。训练得到的主播B形象的编码器可以用于提取主播B形象的脸部动作、神态、表情等姿态相关的姿态特征,主播A形象的解码器可以根据该姿态相关的姿态特征生成对应动作、表情、神态的主播A形象。
(3)将目标音频输入到主播B形象的语音驱动模型中,利用语音驱动模型输出主播B形象在目标音频驱动下的第一驱动视频。
(4)如图3所示,针对第一驱动视频中的各视频帧,可以先将该视频帧输入到主播B形象的编码器中,利用主播B形象的编码器提取视频帧中的动作、表情、神态等姿态相关的姿态特征,然后再将该姿态特征输入到主播A形象的解码器中,主播A形象的解码器则可以基于预先学习到的主播A形象的纹理特征,以及输入的姿态特征重新构建主播A形象,重新构建的主播A形象的姿态和该视频帧中主播B形象的姿态一致,并且可以结合预先设置的背景素材,得到与该视频帧对应的重构帧。针对第一驱动视频中的每一帧视频帧,都可以执行上述步骤,从而可以得到主播A形象在目标音频驱动下的第二驱动视频。
通常音频的自由度较高,往往要大量样本才能覆盖语音中的各种音素组合,因而,如果要训练一个新形象的语音驱动模型,往往需要大量的训练样本,即需要新形象的大量音视频数据,这种方式对于无法获取新形象的大量音视频的场景不适用。而通常人脸的自由度比较低,通过少量的图像就能基本覆盖所有人脸的所有动作,因而,训练一个新形象的自编码器往往仅需少量的图像,比如几帧或几十帧,相比训练一个语音驱动模型,不仅需要耗费的训练样本大大减少,也可以大大提高生成新形象的数字对象视频的效率。
相应的,本公开实施例还提供了一种数字对象生成装置40,如图4所示,所述装置 40包括:获取模块41,用于获取目标音频和目标数字对象;预测模块42,用于根据所述目标音频和预先训练的参考数字对象的语音驱动模型,得到参考数字对象在所述目标音频驱动下的第一驱动视频;其中,所述语音驱动模型基于所述参考数字对象的音视频训练得到;姿态迁移模块43,用于将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到所述目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
在一些实施例中,所述预测模块42用于根据目标音频和预先训练的参考数字对象的语音驱动模型,得到参考数字对象在所述目标音频驱动下的第一驱动视频时,具体用于:从所述目标音频中提取音素,得到音素时间戳;将所述音素时间戳输入到所述语音驱动模型中,得到所述第一驱动视频。
在一些实施例中,所述语音驱动模型基于所述参考数字对象的音视频训练得到,包括:获取所述参考数字对象的音频,以及与所述参考数字对象的音频同步的参考数字对象的视频;从所述参考数字对象的音频中提取音素,得到所述参考数字对象的音素时间戳;以所述参考数字对象的音素时间戳作为训练样本,以所述参考数字对象的视频作为样本标签,训练得到所述语音驱动模型。
在一些实施例中,所述姿态迁移模块43用于将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频时,具体用于:从所述目标数字对象的目标图像中提取得到与所述目标数字对象的纹理相关的纹理特征;针对所述第一驱动视频的各视频帧,从所述视频帧中提取与所述参考数字对象的姿态相关的姿态特征;根据所述纹理特征和所述姿态特征重新构建目标数字对象,其中,所述重新构建的目标数字对象的姿态与所述视频帧中的参考数字对象的姿态一致;根据所述重新构建的目标数字对象得到与所述视频帧对应的重构帧;以及基于所述重构帧来构造所述第二驱动视频。
在一些实施例中,所述姿态迁移模块43用于从所述视频帧中提取与所述参考数字对象的姿态相关的姿态特征时,具体用于:利用所述参考数字对象的自编码器中的编码器从所述视频帧中提取所述姿态特征,其中,所述参考数字对象的自编码器基于包含所述参考数字对象的多个图像预先训练得到。
在一些实施例中,所述姿态迁移模块43用于从所述目标数字对象的目标图像中提取得到与所述目标数字对象的纹理相关的纹理特征时,具体用于:利用所述目标数字对象的自编码器中的解码器从所述目标数字对象的目标图像中提取所述纹理特征,其中,所 述目标数字对象的自编码器基于包含所述目标数字对象的多个图像预先训练得到。
在一些实施例中,所述姿态迁移模块43用于根据所述姿态特征和所述纹理特征重新构建目标数字对象时,具体用于:利用所述目标数字对象的自编码器中的解码器根据所述姿态特征和所述纹理特征重新构建目标数字对象。
在一些实施例中,在根据所述姿态特征和所述纹理特征重新构建目标数字对象之前,所述数字对象生成装置还用于:在检测到从当前帧提取的所述姿态特征与从所述当前帧的前一帧提取的所述姿态特征一致的情况下,将所述前一帧对应的重构帧作为所述当前帧对应的重构帧。
在一些实施例中,所述姿态迁移模块43用于根据所述重新构建的目标数字对象得到与所述视频帧对应的重构帧时,具体用于:根据所述重新构建的目标数字对象和预先设置的背景素材生成所述重构帧。
在一些实施例中,所述参考数字对象的姿态包括以下一种或多种:所述参考数字对象的脸部动作、所述参考数字对象的脸部表情、所述参考数字对象的肢体动作。
此外,本公开实施例还提供了另一种数字对象生成装置50,如图5所示,所述装置50包括:获取模块51,用于获取参考数字对象在目标音频驱动下的第一驱动视频;提取模块52,用于提取所述参考数字对象在所述第一驱动视频中的姿态特征;加载模块53,用于将所述第一驱动视频中的参考数字对象替换为目标数字对象,并将所述姿态特征加载到所述目标数字对象,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
其中,所述数字对象生成装置50在生成目标音频驱动的下的目标数字对象的具体实现步骤可以参考上述各方法实施例中描述,在此不再赘述。
进一步的,本公开实施例还提供一种电子设备60,如图6所示,所述电子设备60包括处理器61、存储器62、存储于所述存储器62可供所述处理器61执行的计算机指令,所述处理器61执行所述计算机指令时实现上述实施例中任一项所述的方法。
本公开实施例还提供一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现前述任一实施例所述的方法。
计算机可读介质包括永久性和非永久性、可移动和非可移动媒体,可以由任何方法或技术来实现信息存储。信息可以是计算机可读指令、数据结构、程序的模块或其他数据。计算机的存储介质的例子包括,但不限于相变内存(PRAM)、静态随机存取存储 器(SRAM)、动态随机存取存储器(DRAM)、其他类型的随机存取存储器(RAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、快闪记忆体或其他内存技术、只读光盘只读存储器(CD-ROM)、数字多功能光盘(DVD)或其他光学存储、磁盒式磁带,磁带磁磁盘存储或其他磁性存储设备或任何其他非传输介质,可用于存储可以被计算设备访问的信息。按照本文中的界定,计算机可读介质不包括暂存电脑可读媒体(transitory media),如调制的数据信号和载波。
通过以上的实施方式的描述可知,本领域的技术人员可以清楚地了解到本说明书实施例可借助软件加必需的通用硬件平台的方式来实现。基于这样的理解,本说明书实施例的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品可以存储在存储介质中,如ROM/RAM、磁碟、光盘等,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本说明书实施例各个实施例或者实施例的某些部分所述的方法。
上述实施例阐明的系统、装置、模块或单元,具体可以由计算机芯片或实体实现,或者由具有某种功能的产品来实现。一种典型的实现设备为计算机,计算机的具体形式可以是个人计算机、膝上型计算机、蜂窝电话、相机电话、智能电话、个人数字助理、媒体播放器、导航设备、电子邮件收发设备、游戏控制台、平板计算机、可穿戴设备或者这些设备中的任意几种设备的组合。
本说明书中的各个实施例均采用递进的方式描述,各个实施例之间相同相似的部分互相参见即可,每个实施例重点说明的都是与其他实施例的不同之处。尤其,对于装置实施例而言,由于其基本相似于方法实施例,所以描述得比较简单,相关之处参见方法实施例的部分说明即可。以上所描述的装置实施例仅仅是示意性的,其中所述作为分离部件说明的模块可以是或者也可以不是物理上分开的,在实施本说明书实施例方案时可以把各模块的功能在同一个或多个软件和/或硬件中实现。也可以根据实际的需要选择其中的部分或者全部模块来实现本实施例方案的目的。本领域普通技术人员在不付出创造性劳动的情况下,即可以理解并实施。
以上所述仅是本说明书实施例的具体实施方式,应当指出,对于本技术领域的普通技术人员来说,在不脱离本说明书实施例原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也应视为本说明书实施例的保护范围。
Claims (15)
- 一种数字对象生成方法,其特征在于,所述方法包括:获取目标音频和目标数字对象;根据所述目标音频和预先训练的参考数字对象的语音驱动模型,得到所述参考数字对象在所述目标音频驱动下的第一驱动视频;其中,所述语音驱动模型基于所述参考数字对象的音频和视频训练得到;将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到所述目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
- 根据权利要求1所述的方法,其特征在于,所述根据所述目标音频和预先训练的参考数字对象的语音驱动模型,得到所述参考数字对象在所述目标音频驱动下的第一驱动视频,包括:从所述目标音频中提取音素,得到音素时间戳;将所述音素时间戳输入到所述语音驱动模型中,得到所述第一驱动视频。
- 根据权利要求1或2所述的方法,其特征在于,所述语音驱动模型基于所述参考数字对象的音频和视频训练得到,包括:获取所述参考数字对象的音频,以及与所述参考数字对象的音频同步的所述参考数字对象的视频;从所述参考数字对象的音频中提取音素,得到所述参考数字对象的音素时间戳;以所述参考数字对象的音素时间戳作为训练样本,以所述参考数字对象的视频作为样本标签,训练得到所述语音驱动模型。
- 根据权利要求1-3任一项所述的方法,其特征在于,所述将所述第一驱动视频的各视频帧中的所述参考数对象的姿态迁移到所述目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频,包括:从所述目标数字对象的目标图像中提取得到与所述目标数字对象的纹理相关的纹理特征;针对所述第一驱动视频的各视频帧,从所述视频帧中提取与所述参考数字对象的姿态相关的姿态特征;根据所述纹理特征和所述姿态特征重新构建所述目标数字对象,其中,所述重新构建的目标数字对象的姿态与所述视频帧中的所述参考数字对象的姿态一致;根据所述重新构建的目标数字对象得到与所述视频帧对应的重构帧;以及基于所述重构帧来构造所述第二驱动视频。
- 根据权利要求4所述的方法,其特征在于,所述从所述视频帧中提取与所述参考数字对象的姿态相关的姿态特征,包括:利用所述参考数字对象的自编码器中的编码器从所述视频帧中提取所述姿态特征,其中,所述参考数字对象的自编码器基于包含所述参考数字对象的多个图像预先训练得到。
- 根据权利要求5所述的方法,其特征在于,所述从所述目标数字对象的目标图像中提取得到与所述目标数字对象的纹理相关的纹理特征,包括:利用所述目标数字对象的自编码器中的解码器,从所述目标数字对象的目标图像中提取所述纹理特征,其中,所述目标数字对象的自编码器基于包含所述目标数字对象的多个图像预先训练得到。
- 根据权利要求6所述的方法,其特征在于,所述根据所述纹理特征和所述姿态特征重新构建所述目标数字对象,包括:利用所述目标数字对象的自编码器中的解码器根据所述纹理特征和所述姿态特征重新构建所述目标数字对象。
- 根据权利要求4-7任一项所述的方法,其特征在于,在所述根据所述纹理特征和所述姿态特征重新构建所述目标数字对象之前,还包括:在检测到从当前帧提取的所述姿态特征与从所述当前帧的前一帧提取的所述姿态特征一致的情况下,将所述前一帧对应的重构帧作为所述当前帧对应的重构帧。
- 根据权利要求4-8任一项所述的方法,其特征在于,所述根据所述重新构建的目标数字对象得到与所述视频帧对应的重构帧,包括:根据所述重新构建的目标数字对象和预先设置的背景素材生成所述重构帧。
- 根据权利要求1-9任一项所述的方法,其特征在于,所述参考数字对象的姿态包括以下一种或多种:所述参考数字对象的脸部动作、所述参考数字对象的脸部表情、所述参考数字对象的肢体动作。
- 一种数字对象生成方法,其特征在于,所述方法包括:获取参考数字对象在目标音频驱动下的第一驱动视频;提取所述参考数字对象在所述第一驱动视频中的姿态特征;将所述第一驱动视频中的所述参考数字对象替换为目标数字对象,并将所述姿态特征加载到所述目标数字对象,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
- 一种数字对象生成装置,其特征在于,所述装置包括:获取模块,用于获取目标音频和目标数字对象;预测模块,用于根据所述目标音频和预先训练的参考数字对象的语音驱动模型,得到所述参考数字对象在所述目标音频驱动下的第一驱动视频;其中,所述语音驱动模型基于所述参考数字对象的音频和视频训练得到;姿态迁移模块,用于将所述第一驱动视频的各视频帧中的所述参考数字对象的姿态迁移到所述目标数字对象上,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
- 一种数字对象生成装置,其特征在于,所述装置包括:获取模块,用于获取参考数字对象在目标音频驱动下的第一驱动视频;提取模块,用于提取所述参考数字对象在所述第一驱动视频中的姿态特征;加载模块,用于将所述第一驱动视频中的所述参考数字对象替换为目标数字对象,并将所述姿态特征加载到所述目标数字对象,得到所述目标数字对象在所述目标音频驱动下的第二驱动视频。
- 一种电子设备,其特征在于,所述电子设备包括处理器、存储器、存储于所述存储器可供所述处理器执行的计算机指令,所述处理器执行所述计算机指令时实现如权利要求1-11任一项所述的方法。
- 一种计算机可读存储介质,其特征在于,所述存储介质上存储有计算机程序,所述计算机程序被执行时实现如权利要求1-11任一项所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202111599988.4A CN114330631A (zh) | 2021-12-24 | 2021-12-24 | 数字人生成方法、装置、设备及存储介质 |
| CN202111599988.4 | 2021-12-24 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023116208A1 true WO2023116208A1 (zh) | 2023-06-29 |
Family
ID=81013894
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2022/128915 Ceased WO2023116208A1 (zh) | 2021-12-24 | 2022-11-01 | 数字对象生成方法、装置、设备及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN114330631A (zh) |
| WO (1) | WO2023116208A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117671093A (zh) * | 2023-11-29 | 2024-03-08 | 上海积图科技有限公司 | 数字人视频制作方法、装置、设备及存储介质 |
| CN120495479A (zh) * | 2025-07-17 | 2025-08-15 | 北京达佳互联信息技术有限公司 | 图像处理方法、装置及存储介质 |
Families Citing this family (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114330631A (zh) * | 2021-12-24 | 2022-04-12 | 上海商汤智能科技有限公司 | 数字人生成方法、装置、设备及存储介质 |
| CN115209180B (zh) * | 2022-06-02 | 2024-06-18 | 阿里巴巴(中国)有限公司 | 视频生成方法以及装置 |
| CN115883753A (zh) * | 2022-11-04 | 2023-03-31 | 网易(杭州)网络有限公司 | 视频的生成方法、装置、计算设备及存储介质 |
| CN115550744B (zh) * | 2022-11-29 | 2023-03-14 | 苏州浪潮智能科技有限公司 | 一种语音生成视频的方法和装置 |
| CN116030825A (zh) * | 2022-12-28 | 2023-04-28 | 中国电信股份有限公司 | 数字人驱动视频生成方法、装置、电子设备及存储介质 |
| CN118474410A (zh) * | 2023-02-08 | 2024-08-09 | 华为技术有限公司 | 视频生成方法、装置和存储介质 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2021196643A1 (zh) * | 2020-03-31 | 2021-10-07 | 北京市商汤科技开发有限公司 | 交互对象的驱动方法、装置、设备以及存储介质 |
| WO2021196646A1 (zh) * | 2020-03-31 | 2021-10-07 | 北京市商汤科技开发有限公司 | 交互对象的驱动方法、装置、设备以及存储介质 |
| US20210326372A1 (en) * | 2020-04-17 | 2021-10-21 | Accenture Global Solutions Limited | Human centered computing based digital persona generation |
| CN113747086A (zh) * | 2021-09-30 | 2021-12-03 | 深圳追一科技有限公司 | 数字人视频生成方法、装置、电子设备及存储介质 |
| CN114330631A (zh) * | 2021-12-24 | 2022-04-12 | 上海商汤智能科技有限公司 | 数字人生成方法、装置、设备及存储介质 |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105118082B (zh) * | 2015-07-30 | 2019-05-28 | 科大讯飞股份有限公司 | 个性化视频生成方法及系统 |
| CN109788312B (zh) * | 2019-01-28 | 2022-10-21 | 北京易捷胜科技有限公司 | 一种视频中人物的替换方法 |
| CN112866586B (zh) * | 2021-01-04 | 2023-03-07 | 北京中科闻歌科技股份有限公司 | 一种视频合成方法、装置、设备及存储介质 |
| CN112801861A (zh) * | 2021-01-29 | 2021-05-14 | 恒安嘉新(北京)科技股份公司 | 一种影视作品的制作方法、装置、设备及存储介质 |
| CN113192161B (zh) * | 2021-04-22 | 2022-10-18 | 清华珠三角研究院 | 一种虚拟人形象视频生成方法、系统、装置及存储介质 |
| CN113486785B (zh) * | 2021-07-01 | 2024-08-13 | 株洲霍普科技文化股份有限公司 | 基于深度学习的视频换脸方法、装置、设备及存储介质 |
| CN113507627B (zh) * | 2021-07-08 | 2022-03-25 | 北京的卢深视科技有限公司 | 视频生成方法、装置、电子设备及存储介质 |
| CN113822969B (zh) * | 2021-09-15 | 2023-06-09 | 宿迁硅基智能科技有限公司 | 训练神经辐射场模型和人脸生成方法、装置及服务器 |
-
2021
- 2021-12-24 CN CN202111599988.4A patent/CN114330631A/zh active Pending
-
2022
- 2022-11-01 WO PCT/CN2022/128915 patent/WO2023116208A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2021196643A1 (zh) * | 2020-03-31 | 2021-10-07 | 北京市商汤科技开发有限公司 | 交互对象的驱动方法、装置、设备以及存储介质 |
| WO2021196646A1 (zh) * | 2020-03-31 | 2021-10-07 | 北京市商汤科技开发有限公司 | 交互对象的驱动方法、装置、设备以及存储介质 |
| US20210326372A1 (en) * | 2020-04-17 | 2021-10-21 | Accenture Global Solutions Limited | Human centered computing based digital persona generation |
| CN113747086A (zh) * | 2021-09-30 | 2021-12-03 | 深圳追一科技有限公司 | 数字人视频生成方法、装置、电子设备及存储介质 |
| CN114330631A (zh) * | 2021-12-24 | 2022-04-12 | 上海商汤智能科技有限公司 | 数字人生成方法、装置、设备及存储介质 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117671093A (zh) * | 2023-11-29 | 2024-03-08 | 上海积图科技有限公司 | 数字人视频制作方法、装置、设备及存储介质 |
| CN120495479A (zh) * | 2025-07-17 | 2025-08-15 | 北京达佳互联信息技术有限公司 | 图像处理方法、装置及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN114330631A (zh) | 2022-04-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN114330631A (zh) | 数字人生成方法、装置、设备及存储介质 | |
| US12184923B2 (en) | Action synchronization for target object | |
| CN113077537B (zh) | 一种视频生成方法、存储介质及设备 | |
| CN114581812B (zh) | 视觉语言识别方法、装置、电子设备及存储介质 | |
| KR102826529B1 (ko) | 비디오 스토리 질의응답을 위한 트랜스포머 모델을 구축하는 방법 및 이를 수행하기 위한 컴퓨팅 장치 | |
| WO2025161585A1 (zh) | 基于数字人的视频生成与交互方法、设备、存储介质与程序产品 | |
| US20250200855A1 (en) | Method for real-time generation of empathy expression of virtual human based on multimodal emotion recognition and artificial intelligence system using the method | |
| US20230326369A1 (en) | Method and apparatus for generating sign language video, computer device, and storage medium | |
| CN110505498A (zh) | 视频的处理、播放方法、装置及计算机可读介质 | |
| KR20230095432A (ko) | 텍스트 서술 기반 캐릭터 애니메이션 합성 시스템 | |
| CN117121099B (zh) | 自适应视觉语音识别 | |
| CN118413722B (zh) | 音频驱动视频生成方法、装置、计算机设备以及存储介质 | |
| US20240320519A1 (en) | Systems and methods for providing a digital human in a virtual environment | |
| CN117155884A (zh) | 基于虚拟形象的语音交互方法、电子设备及存储介质 | |
| CN117850733A (zh) | 基于虚拟形象的语音交互方法及智能终端 | |
| HK40062802A (zh) | 数字人生成方法、装置、设备及存储介质 | |
| CN117476007A (zh) | 数字人视频生成方法及装置、设备、存储介质 | |
| Arthur et al. | Towards a practical lip-to-speech conversion system using deep neural networks and mobile application frontend | |
| CN121397323B (zh) | 基于数字人的长视频生成方法、设备及存储介质 | |
| CN121188613B (zh) | 基于小波分解与注意力筛选机制的课堂教学情感识别方法、装置、设备及介质 | |
| US20250392766A1 (en) | Augmented streaming media | |
| CN113744371B (zh) | 一种生成人脸动画的方法、装置、终端及存储介质 | |
| CN120088823A (zh) | 表情迁移的方法、装置电子设备及计算机可读存储介质 | |
| WO2026021142A1 (zh) | 视频交互方法和装置 | |
| CN121334462A (zh) | 一种视频生成方法、装置、存储介质及电子设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22909525 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 21.11.2024) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22909525 Country of ref document: EP Kind code of ref document: A1 |