WO2025213845A1 - 一种视频生成方法、装置、设备、介质、产品 - Google Patents

一种视频生成方法、装置、设备、介质、产品

Info

Publication number
WO2025213845A1
WO2025213845A1 PCT/CN2024/140100 CN2024140100W WO2025213845A1 WO 2025213845 A1 WO2025213845 A1 WO 2025213845A1 CN 2024140100 W CN2024140100 W CN 2024140100W WO 2025213845 A1 WO2025213845 A1 WO 2025213845A1
Authority
WO
WIPO (PCT)
Prior art keywords
audio
image
frame
video
facial
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/140100
Other languages
English (en)
French (fr)
Inventor
张隆昊
胡天舒
梁爽
葛志鹏
唐铭谦
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Publication of WO2025213845A1 publication Critical patent/WO2025213845A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/233Processing of audio elementary streams
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/234Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/234Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
    • H04N21/23418Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/439Processing of audio elementary streams
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • H04N21/44008Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream

Definitions

  • the embodiments of the present disclosure relate to the field of data processing technology, and in particular to a video generation method, apparatus, device, medium, and product.
  • the present disclosure provides a video generation method, the method comprising:
  • an image corresponding to the i-th audio frame based on the facial key points corresponding to the i-th audio frame in the audio sequence, a lip mask result of a target image corresponding to the i-th audio frame in the reference video, at least two reference images selected from the reference video, and the facial key points of each reference image, where i is a positive integer and i ⁇ the total number of frames in the audio sequence;
  • a video corresponding to the audio sequence is generated based on the image corresponding to each audio frame in the audio sequence.
  • the process of determining the at least two reference image frames includes:
  • the at least two reference images are determined based on the sampled image.
  • the at least two reference image frames are determined based on the sampled image and at least one image with a similar posture corresponding to the i-th frame of audio in the reference video;
  • the time corresponding to each of the similar posture images in the reference video is different from the time corresponding to the i-th frame of audio;
  • the time corresponding to the i-th frame of audio is determined based on the time corresponding to the target image in the reference video.
  • the posture representation data corresponding to the i-th frame of audio is determined based on the posture representation data of the target image.
  • the facial key points corresponding to the i-th audio frame include facial key point determination results of at least one audio frame in the audio sequence
  • the at least one frame of audio includes the i-th frame of audio.
  • the facial key point determination result of the audio is a two-dimensional facial key point obtained by projecting the three-dimensional facial key point corresponding to the audio onto a two-dimensional plane;
  • the three-dimensional facial key points corresponding to the audio are obtained by performing three-dimensional facial key point determination processing on the audio based on part or all of the images in the reference video.
  • facial key points of the reference image are two-dimensional facial key points obtained by projecting three-dimensional facial key points of the reference image onto a two-dimensional plane;
  • the three-dimensional facial key points of the reference image are obtained by performing a three-dimensional facial key point determination process on the reference image.
  • the process of determining the image corresponding to the i-th audio frame includes:
  • the pixel usage description information includes pixel adjustment description information and/or pixel fusion weights
  • An image generation process is performed based on the fused image, the lip mask result of the target image, and the facial key points corresponding to the i-th frame of audio to obtain an image corresponding to the i-th frame of audio.
  • the pixel usage description information corresponding to each reference image is determined by using an information prediction module in a lip-sync rendering model
  • the deformation fusion processing is achieved by using the deformation fusion module in the lip rendering model;
  • the image generation process is implemented by utilizing the face generation module in the lip rendering model.
  • the reference video and the audio sequence are both determined based on a sample video
  • the method After generating the image corresponding to the i-th frame of audio, the method further includes:
  • the face generation module in the lip rendering model is updated; and based on the difference representation data between the target image and the image corresponding to the i-th frame of audio, and the difference representation data between the target image and the fused image, the deformation fusion module and the information prediction module in the lip rendering model are updated.
  • the present disclosure provides a video generation device, including:
  • a data acquisition unit configured to acquire reference video and audio sequences
  • an image generation unit configured to generate an image corresponding to the i-th frame of audio based on facial key points corresponding to the i-th frame of audio in the audio sequence, a lip mask result of a target image corresponding to the i-th frame of audio in the reference video, at least two reference images selected from the reference video, and facial key points of each of the reference images; where i is a positive integer and i ⁇ the total number of frames in the audio sequence;
  • the video generation unit is used to generate a video corresponding to the audio sequence based on the image corresponding to each audio frame in the audio sequence.
  • An embodiment of the present disclosure provides an electronic device, characterized in that the device includes: a processor and a memory;
  • the memory is used to store instructions or computer programs
  • the processor is configured to execute the instructions or computer program in the memory, so that the electronic device executes the video generation method provided by the embodiment of the present disclosure.
  • An embodiment of the present disclosure provides a computer-readable medium, characterized in that instructions or computer programs are stored in the computer-readable medium, and when the instructions or computer programs are executed on a device, the device executes the video generation method provided by the embodiment of the present disclosure.
  • An embodiment of the present disclosure provides a computer program product, characterized in that it includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the video generation method provided by the embodiment of the present disclosure.
  • FIG1 is a flow chart of a video generation method provided by an embodiment of the present disclosure.
  • FIG2 is a schematic diagram of a video generation process provided by an embodiment of the present disclosure.
  • FIG3 is a schematic diagram of an image generation process provided by an embodiment of the present disclosure.
  • FIG4 is a schematic structural diagram of a video generating device provided by an embodiment of the present disclosure.
  • FIG5 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure.
  • these application scenarios may have the following requirements: generating a video adapted to a certain audio sequence based on an existing video.
  • the video generation method provided by the embodiments of the present disclosure includes the following steps S1-S3.
  • Figure 1 is a flow chart of a video generation method provided by the embodiments of the present disclosure.
  • the reference video refers to the video required as a basis for video generation, such as the reference video shown in Figure 2, so that the reference video is used to provide other information in addition to facial expression status, such as lip shape status, such as facial identification information similar to facial contour characteristics and facial features distribution characteristics, facial posture information, etc.
  • the reference video described above can be a video specified by the user through certain means, such as a single-person voice-over video, so that the reference video meets certain user requirements, such as facial identification information, facial posture, etc.
  • certain means such as a single-person voice-over video
  • the embodiments of the present disclosure are not limited to this means.
  • the reference video can be a video manually uploaded by the user, a video selected by the user from a number of candidate videos, or a video downloaded by the user through some other means.
  • the reference video described above can be determined based on a sample video, such that the reference video includes some or all images from the sample video.
  • the sample video refers to the video required for use in the model training process, and the embodiments of this disclosure do not limit the implementation of the sample video.
  • the embodiments of the present disclosure do not limit the method for obtaining the above reference video.
  • An audio sequence refers to the audio required for video generation, such as the audio sequence shown in FIG2 , so that the audio sequence is used to provide information such as facial expression status, such as lip shape status, and the like.
  • Scenario 1 When the video generation method provided by the embodiment of the present disclosure is used to perform a certain generation task, such as a video generation task, the above audio sequence can be specified by the user with the help of certain means; and the embodiment of the present disclosure does not limit the implementation method of the audio sequence.
  • the audio sequence satisfies the following constraints: the semantic information of the sentence described by the audio sequence is consistent with the semantic information of the sentence described in the reference video, but the language of the sentence described by the audio sequence is different from the language of the sentence described by the reference video.
  • the audio sequence at least satisfies the following constraints: the sentence described by the audio sequence is partially identical to the sentence described in the reference video.
  • the audio sequence at least satisfies the following constraints: the sentence described by the audio sequence is completely different from the sentence described in the reference video.
  • Scenario 2 When the video generation method provided by the embodiment of the present disclosure is used to implement the model training process, if the above reference video is determined based on the sample video, the above audio sequence can be determined based on the sample video, so that the audio sequence includes part or all of the audio in the sample video.
  • the embodiment of the present disclosure does not limit the association relationship between the above audio sequence and the above reference video.
  • the two can satisfy the following constraints: for the i-th frame of audio in the audio sequence, there is a target image corresponding to the i-th frame of audio in the reference video.
  • the target image refers to an image that exists in the reference video and has a corresponding relationship with the i-th frame of audio, so that the target image is used to provide other information other than the facial expression state for the facial image generation process of the i-th frame of audio, such as facial features, facial posture and other information; and the embodiment of the present disclosure does not limit the implementation method of the target image.
  • the target image can refer to the i-th frame image in the reference video.
  • i is a positive integer, i ⁇ I, and I represents the total number of frames in the audio sequence.
  • this corresponding relationship may specifically include: the correspondence between the i-th frame of audio in the audio sequence and the i-th frame of image in the reference video.
  • this corresponding relationship may specifically include: the correspondence between the i-th frame of audio in the audio sequence and the i-th frame of image in the reference video after frame addition.
  • the reference video after frame addition is obtained by performing a certain frame addition processing on the original reference video so that the total number of frames of the reference video after frame addition is equal to the total number of frames of the audio sequence; and the embodiments of the present disclosure do not limit the implementation of the frame addition processing.
  • it can be implemented by any existing or future method that can increase the total number of frames of a video, such as performing frame insertion processing on the original video or copying and splicing processing on the original video.
  • this correspondence may specifically include: the correspondence between the i-th audio frame in the audio sequence and the i-th image frame in the reference video after frame reduction.
  • the reference video after frame reduction is obtained by performing a certain frame reduction process on the original reference video, so that the total number of frames of the reference video after frame reduction is equal to the total number of frames of the audio sequence; and the embodiments of the present disclosure do not limit the implementation method of the frame reduction process.
  • it can be implemented using any existing or future method that can reduce the total number of frames of a video, such as sampling the original video or cropping the original video.
  • i is a positive integer
  • i ⁇ I and I represents the total number of frames of the audio sequence.
  • the embodiment of the present disclosure does not limit the method for obtaining the above audio sequence.
  • S2 Generate an image corresponding to the i-th frame of audio based on the facial key points corresponding to the i-th frame of audio in the audio sequence, the lip mask result of the target image corresponding to the i-th frame of audio in the reference video, at least two reference images selected from the reference video, and the facial key points of each reference image; i is a positive integer, i ⁇ the total number of frames in the audio sequence.
  • the facial key points corresponding to the i-th frame of audio are used to represent the facial expression state of the object in the reference video in the i-th frame of audio, such as the lip shape. It should be noted that the embodiments of the present disclosure do not limit the implementation of the object.
  • the object can be implemented as any animal or avatar capable of expressing expressions.
  • the disclosed embodiments do not limit the implementation of the facial key points corresponding to the i-th frame of audio.
  • it can be implemented using any existing or future key points that can be used to represent facial expression states.
  • the facial key points corresponding to the i-th frame of audio can be implemented using two-dimensional facial key points.
  • the two-dimensional facial key points are used to describe the facial expression state of an object in two-dimensional space.
  • the facial key points corresponding to the i-th audio frame described above may include the two-dimensional facial key points of the i-th audio frame.
  • the two-dimensional facial key points of the i-th audio frame are used to describe the two-dimensional facial expression state of the subject in the reference video corresponding to the i-th audio frame.
  • the two-dimensional facial key points of the i-th audio frame may be determined based on the i-th audio frame and the reference video.
  • the present disclosure does not limit the method for determining the two-dimensional facial key points of the i-th frame of audio.
  • it can be implemented using any existing or future method capable of generating two-dimensional facial key points based on a video and a frame of audio, such as a method implemented using a pre-built two-dimensional key point generation model.
  • the two-dimensional key point generation model is used to generate two-dimensional facial key points based on its input data, such as video and audio.
  • the present disclosure does not limit the implementation of the two-dimensional key point generation model.
  • the embodiments of the present disclosure also provide a method for determining the two-dimensional facial key points of the i-th frame of audio described above.
  • the two-dimensional facial key points of the i-th frame of audio can be obtained by projecting the three-dimensional facial key points corresponding to the i-th frame of audio onto a two-dimensional plane.
  • the three-dimensional facial key points corresponding to the i-th frame of audio are used to describe the three-dimensional facial expression state of the object in the reference video under the i-th frame of audio; and the three-dimensional facial key points corresponding to the i-th frame of audio are obtained by performing a three-dimensional facial key point determination process on the i-th frame of audio based on part or all of the images in the reference video.
  • the embodiments of the present disclosure do not limit the implementation method of the three-dimensional facial key point determination process.
  • it can be implemented using any existing or future method that can perform three-dimensional facial key point determination based on a number of images and a frame of audio, such as a method implemented using a pre-built three-dimensional key point generation model.
  • the three-dimensional key point generation model is used to perform three-dimensional facial key point generation processing on the input data of the three-dimensional key point generation model, such as video + audio; and the embodiments of the present disclosure do not limit the implementation method of the three-dimensional key point generation model.
  • the process of determining the three-dimensional facial key points corresponding to the i-th frame of audio may include the following steps 11 and 12.
  • Step 11 Predict the 3D facial parameters corresponding to the i-th frame of audio based on the reference video and the i-th frame of audio.
  • the three-dimensional facial parameters corresponding to the i-th frame of audio are used to describe the facial state of the object in the reference video under the i-th frame of audio, such as facial features, facial expressions, facial posture, etc.
  • the embodiments of the present disclosure do not limit the implementation method of the three-dimensional facial parameters corresponding to the i-th frame of audio.
  • the three-dimensional facial parameters corresponding to the i-th frame of audio may include facial identification (ID) parameters, facial expression parameters, and facial posture parameters.
  • ID facial identification
  • the facial identification parameters are used to describe the identifying facial features of the object in the reference video, such as facial contour, distribution of facial features, etc., so that the three-dimensional facial model constructed based on the facial identification parameters is in a state with ID, no posture, and no expression.
  • the facial expression parameters are used to describe the facial expression state under the i-th frame of audio, such as the lip shape state, etc., so that the three-dimensional facial model constructed based on the facial expression parameters is in a state without ID, no posture, and with expression.
  • the facial posture parameters are used to describe the facial posture under the i-th frame of audio, such as side face, front face, etc., so that the three-dimensional facial model constructed based on the facial posture parameters is in a state without ID, with posture, and without expression. It should be noted that the embodiment of the present disclosure does not limit the correlation relationship between these three parameters. For example, these three parameters are in a decoupled state from each other.
  • the present disclosure does not limit the process for determining the three-dimensional facial parameters corresponding to the i-th frame of audio.
  • the process may employ any existing or future method capable of performing three-dimensional facial parameter determination based on a video and a frame of audio, such as a method implemented using a pre-built facial parameter determination model.
  • the facial parameter determination model refers to a pre-built model capable of determining three-dimensional facial parameters, such as a machine learning model.
  • the present disclosure does not limit the implementation of the facial parameter determination model.
  • the three-dimensional facial parameters corresponding to the i-th frame audio above meet the following constraints: 1
  • the facial identification parameters and facial expression parameters in the three-dimensional facial parameters corresponding to the i-th frame audio are determined based on the reference video and the i-th frame audio, so that the facial features described by the facial identification parameters are consistent with the facial features presented in the reference video, and the facial expression state represented by the facial expression parameters meets the facial expression state requirements of the i-th frame audio;
  • the facial posture parameters in the three-dimensional facial parameters corresponding to the i-th frame audio are determined based on the facial posture parameters in the three-dimensional facial parameters of the target image above, so that the facial posture parameters in the three-dimensional facial parameters corresponding to the i-th frame audio are consistent with the facial posture parameters in the three-dimensional facial parameters of the target image.
  • the three-dimensional facial parameters of the target image are obtained by performing three-dimensional facial parameter determination processing on the target image; and the embodiment of the present disclosure does not limit the process of determining the three-dimensional facial parameters of the target image.
  • it can be implemented using any existing or future method that can perform three-dimensional facial parameter determination processing on an image.
  • Step 12 Determine the 3D facial key points corresponding to the i-th frame of audio based on the 3D facial parameters corresponding to the i-th frame of audio.
  • the embodiments of the present disclosure are not limited to the implementation of the above step 12.
  • it can be implemented by using any existing or future method that can determine three-dimensional facial key points based on three-dimensional facial parameters.
  • facial reconstruction processing can be first performed based on a reference video to obtain three-dimensional facial parameters of each frame image in the reference video; then, based on the three-dimensional facial parameters of some or all images in the reference video and the audio features of the i-th frame of audio, the three-dimensional facial parameters corresponding to the i-th frame of audio are predicted; then, based on the three-dimensional facial parameters corresponding to the i-th frame of audio, the three-dimensional facial key points corresponding to the i-th frame of audio are determined, so that the three-dimensional facial key points can better represent the facial state of the object in the reference video in the i-th frame of video, thereby making the facial key points corresponding to the i-th frame of audio determined based on the three-dimensional facial key points more accurate, thereby facilitating improved generation effects.
  • the embodiment of the present disclosure also provides a possible implementation method of the facial key points corresponding to the i-th frame of audio mentioned above.
  • the facial key points corresponding to the i-th frame of audio can include the facial key point determination results of at least one frame of audio in the audio sequence, and the at least one frame of audio includes the i-th frame of audio.
  • the at least one frame of audio refers to the audio present in the audio sequence that is required for reference when generating the facial image of the i-th frame of audio.
  • the embodiment of the present disclosure does not limit the implementation method of the at least one frame of audio.
  • the at least one frame of audio can include the i-Tth frame of audio, the i-T+1th frame of audio, ..., the i-th frame of audio, the i+1th frame of audio, ..., and the i+Tth frame of audio in the audio sequence, so that the at least one frame of audio can better represent the audio information and context information carried by the i-th frame of audio, thereby making the image generated based on the at least one frame of audio more accurate.
  • the facial key point determination result for the audio refers to the facial key points determined for the audio, so that the facial key point determination result can represent the facial state of the subject in the reference video under the audio.
  • the disclosed embodiments do not limit the implementation of the facial key point determination result for the audio.
  • the facial key point determination result for the audio can be implemented using three-dimensional facial key points and/or two-dimensional facial key points.
  • the facial key point determination result for the audio can be two-dimensional facial key points obtained by projecting the three-dimensional facial key points corresponding to the audio onto a two-dimensional plane.
  • the three-dimensional facial key points corresponding to the audio are obtained by performing three-dimensional facial key point determination processing on the audio based on part or all of the images in the reference video. Furthermore, the implementation of the three-dimensional facial key points corresponding to the audio is similar to the implementation of the three-dimensional facial key points corresponding to the i-th audio frame described above and will not be further described here for the sake of brevity.
  • a facial key point determination result for at least one audio frame in the audio sequence can be first determined based on the three-dimensional facial key points corresponding to the at least one audio frame in the audio sequence. Then, based on the facial key point determination result for the at least one audio frame, the facial key points corresponding to the i-th frame of audio can be determined, so that the facial key points corresponding to the i-th frame of audio include the facial key point determination result for the at least one audio frame.
  • facial key points corresponding to the i-th frame of audio to better represent the facial state of the subject in the reference video in the i-th frame of audio, such as facial features, facial expressions, facial posture, contextual information, etc., thereby improving the image generated based on the facial key points corresponding to the i-th frame of audio.
  • the target image refers to an image existing in the reference video, which is used to provide other information in addition to the facial expression state for the facial image generation processing of the i-th frame of audio, such as the i-th frame image in the reference video; and the lip mask result of the target image is used to provide other information in addition to the lip state in the target image; and the embodiment of the present disclosure does not limit the determination process of the lip mask result.
  • it can be specifically: masking the lower half of the face in the target image to obtain the lip mask result of the target image, so that the lip mask result can represent the upper half of the face in the target image, so that the lip mask result can provide some other information in addition to the lip state in the target image, such as skin color, light and other information, thereby making the image generated based on the lip mask result better.
  • the at least two reference image frames refer to images selected from the reference video, so that the at least two reference image frames can provide some useful pixel information to the facial image generation process of the i-th frame of audio, such as pixel information related to posture, lip shape, etc.
  • the embodiments of the present disclosure do not limit the process of determining the at least two reference image frames mentioned above.
  • the at least two reference image frames may be randomly selected from the reference video, such as the ones randomly selected for the i-th frame of audio.
  • the at least two reference image frames corresponding to different frames of audio in the audio sequence mentioned above refer to the same group of images randomly selected from the reference video.
  • the at least two reference image frames corresponding to different frames of audio in the audio sequence refer to different groups of images randomly selected from the reference video.
  • the embodiment of the present disclosure also provides a method for determining the at least two frames of reference images mentioned above.
  • the determination process of the at least two frames of reference images may include the following steps 21 to 23.
  • Step 21 Sort the images in the reference video according to the lip amplitude representation data of each frame image in the reference video to obtain an image sequence.
  • the j-th frame image in the reference video refers to the image that exists in the reference video and is at the j-th arrangement position
  • j is a positive integer
  • j ⁇ J is a positive integer
  • J represents the total number of frames in the reference video.
  • the lip amplitude representation data of the j-th frame image is used to represent the lip opening amplitude of the object in the j-th frame image
  • the embodiments of the present disclosure do not limit the method for determining the lip amplitude representation data of the j-th frame image.
  • it can be implemented using any existing or future method that can perform lip amplitude determination processing for an image, such as a method implemented by a pre-constructed lip amplitude determination model.
  • the lip amplitude determination model refers to a pre-constructed model that can perform lip amplitude determination processing for an image, such as a machine learning model.
  • the embodiment of the present disclosure also provides a method for determining the lip amplitude representation data of the j-th frame image above.
  • the process of determining the lip amplitude representation data of the j-th frame image can be: based on the facial key points of the j-th frame image, such as the mouth key points in the two-dimensional facial key points and/or the three-dimensional facial key points, determine the lip amplitude representation data of the j-th frame image, so that the lip amplitude representation data can more accurately represent the lip opening amplitude of the object in the j-th frame image.
  • the three-dimensional facial key points of the j-th frame image are used to represent the state of the facial state presented in the j-th frame image in the three-dimensional space; and the three-dimensional facial key points of the j-th frame image are obtained by performing a three-dimensional facial key point determination process on the j-th frame image.
  • the embodiment of the present disclosure does not limit the implementation method of the three-dimensional facial key point determination process.
  • the two-dimensional facial key points of the j-th frame image are used to represent the facial state of the object in the j-th frame image in two-dimensional space.
  • the disclosed embodiments do not limit the method for determining the two-dimensional facial key points of the j-th frame image.
  • the two-dimensional facial key points of the j-th frame image can be obtained by projecting the three-dimensional facial key points of the j-th frame image onto a two-dimensional plane. This helps improve the consistency of key point-related information, thereby improving the generation effect.
  • the image sequence can be determined by sorting some or all of the images in the reference video based on the lip-shape amplitude representation data of each frame in the reference video to obtain an image sequence, so that all images in the image sequence are arranged in ascending or descending order according to the lip-shape amplitude representation data, thereby enabling the image sequence to better describe the lip-shape amplitude distribution presented in the reference video.
  • the image sequence can include all images in the reference video; and all images in the image sequence are arranged in ascending or descending order according to the lip-shape amplitude representation data.
  • the lip amplitude representation data of each frame image in the reference video can be determined first; then these images are arranged in ascending or descending order according to the lip amplitude representation data to obtain an image sequence, so that the image sequence includes part or all of the images in the reference video, and the arrangement position of each image in the image sequence is positively correlated or negatively correlated with the lip amplitude representation data of the corresponding image in the reference video, so that the image sequence can better represent the lip changes presented in the reference video.
  • Step 22 Perform equal-interval sampling on the above image sequence to obtain sampled images.
  • the sampled image refers to the image sampled from the image sequence; and the embodiment of the present disclosure does not limit the sampling interval required for reference during the sampling.
  • the sampling interval can be comprehensively determined based on actual needs, such as generation efficiency requirements or generation accuracy requirements.
  • step 22 above Based on the relevant content of step 22 above, it can be known that for some application scenarios, after obtaining the above image sequence, some images can be sampled at equal intervals from the image sequence as sampling images, so that these sampling images can represent the different lip states of the objects in the reference video, so that these sampling images can provide useful lip shape information for the facial image generation process of the i-th frame audio, which is conducive to improving the generation effect.
  • Step 23 Determine at least two reference images based on the sampled image.
  • step 23 may specifically include determining the sampled image as the at least two reference images above, so that the at least two reference images can provide useful information, such as lip shape information, for the facial image generation process of each frame in the audio sequence.
  • the at least two reference images involved in the facial image generation process of different frames in the audio sequence are both the sampled images, thereby ensuring that the at least two reference images involved in the facial image generation process of different frames in the audio sequence remain the same.
  • the embodiment of the present disclosure also provides a possible implementation method of the above step 23.
  • the step 23 can be specifically as follows: based on the above sampling image and at least one posture-similar image corresponding to the i-th frame of audio in the reference video, determine at least two reference images corresponding to the i-th frame of audio, so that the at least two reference images corresponding to the i-th frame of audio include the sampling image and the at least one posture-similar image, so that the at least two reference images corresponding to the i-th frame of audio can provide as comprehensive and useful information as possible for the facial image generation process of the i-th frame of audio, such as lip shape information + posture information, etc., so that the image corresponding to the i-th frame of audio can be better generated based on the at least two reference images.
  • the at least one posture-similar image refers to some images in the reference video that are most similar to the posture corresponding to the i-th frame of audio.
  • the embodiments of the present disclosure do not limit the implementation method of at least one posture-similar image in the above paragraph.
  • the at least one posture-similar image may satisfy the following constraint: the degree of similarity between the posture representation data of each posture-similar image and the posture representation data corresponding to the i-th frame audio reaches a preset similarity requirement.
  • the embodiments of the present disclosure do not limit the implementation method of the preset similarity requirement.
  • the preset similarity requirement may specifically be: the degree of similarity between the posture representation data of each posture-similar image and the posture representation data corresponding to the i-th frame audio exceeds a preset similarity threshold.
  • the preset similarity requirement may specifically be: the arrangement position of each posture-similar image in the sorting result is higher than the preset arrangement position threshold.
  • the posture representation data of the posture-similar image is used to represent the facial posture of the object in the posture-similar image; and the embodiments of the present disclosure do not limit the determination process of the posture representation data of the posture-similar image.
  • it can be implemented using any existing or future method that can perform posture extraction processing on an image.
  • the determination process of the posture representation data of the posture-similar image can be: determining the posture representation data of the posture-similar image based on the facial posture parameters in the three-dimensional facial parameters of the posture-similar image.
  • the three-dimensional facial parameters of the posture-similar image are used to represent the facial state of the object in the posture-similar image in three-dimensional space; and the implementation method of the three-dimensional facial parameters of the posture-similar image is similar to the implementation method of the three-dimensional facial parameters corresponding to the i-th frame of audio mentioned above. It should be noted that the embodiments of the present disclosure do not limit the implementation method of the step of "determining the posture representation data of the posture similar image based on the facial posture parameters in the three-dimensional facial parameters of the posture similar image".
  • it can be specifically: first use the facial posture parameters in the three-dimensional facial parameters of the posture similar image to determine the three-dimensional facial key points of the posture similar image, so that the three-dimensional facial key points of the posture similar image can represent the facial state without ID, without expression and with posture; then determine the three-dimensional facial key points of the posture similar image as the posture representation data of the posture similar image.
  • the posture representation data corresponding to the i-th frame of audio described above is used to represent the facial posture associated with the i-th frame of audio; and the disclosed embodiments do not limit the method for determining the posture representation data.
  • the posture representation data corresponding to the i-th frame of audio is determined based on the posture representation data of the target image described above.
  • the posture representation data of the target image is used to describe the facial posture presented in the target image; and the implementation method of the posture representation data of the target image is similar to the implementation method of the posture representation data of the posture-similar image described above. For the sake of brevity, this will not be repeated here.
  • the embodiment of the present disclosure also provides a possible implementation method of the above-mentioned at least one posture-similar image.
  • the at least one posture-similar image can meet the following constraints: the degree of similarity between the posture representation data of each posture-similar image and the posture representation data corresponding to the i-th frame audio reaches the preset similarity requirement, and the time corresponding to each posture-similar image in the reference video is different from the time corresponding to the i-th frame audio.
  • the time corresponding to the i-th frame audio is determined based on the time corresponding to the target image in the reference video; and the embodiment of the present disclosure does not limit the method for determining the time corresponding to the i-th frame audio.
  • the at least one posture-similar image corresponding to the i-th frame of audio can meet the following constraints: the degree of similarity between the posture representation data of each posture-similar image and the posture representation data corresponding to the i-th frame of audio reaches the preset similarity requirement, and the time corresponding to each posture-similar image in the reference video is different from the time corresponding to the target image in the reference video, so that the at least one posture-similar image corresponding to the i-th frame of audio can include multiple frames of images in the reference video that are closest to the posture corresponding to the i-th frame of audio, except for the target image, thereby effectively avoiding some interference that may be caused by the target image, and further enabling these posture-similar images to better provide useful posture information for the facial image generation process of the i-th frame of audio, which is conducive to improving the
  • the embodiments of the present disclosure provide two possible implementations of the at least two reference images mentioned above.
  • One is fully fixed, for example, by performing lip-sync processing on the reference video to select some images with different lip-sync states from the reference video, such as the sampled images mentioned above, so that these images with different lip-sync states can be subsequently applied to the facial image generation process of each frame of audio in the audio sequence, so that the at least two reference images referenced by the facial image generation process of different frames of audio in the audio sequence remain consistent.
  • the other is partially fixed and partially dynamic, for example, by performing lip-sync processing on the reference video to select some images with different lip-sync states from the reference video as the fixed part, and by performing posture processing on the reference video to select some images with postures close to the posture corresponding to the i-th frame of audio from the reference video as the dynamic part corresponding to the i-th frame of audio, so that the fixed part + the dynamic part corresponding to the i-th frame of audio can be subsequently applied to the facial image generation process of the i-th frame of audio, which is conducive to better generation effect.
  • the facial key points of the reference image are used to describe the facial state presented in the reference image, and the embodiments of the present disclosure do not limit the implementation method of the facial key points of the reference image.
  • the implementation method of the facial key points of the reference image is similar to the implementation method of the facial key points corresponding to the i-th frame of audio described above.
  • the facial key points corresponding to the i-th frame of audio are implemented using two-dimensional facial key points
  • the facial key points of the reference image can also be implemented using two-dimensional facial key points
  • the facial key points of the reference image and the facial key points corresponding to the i-th frame of audio satisfy the following constraint: the facial key points of the reference image and the facial key points corresponding to the i-th frame of audio correspond in key point sequence number.
  • the embodiments of the present disclosure do not limit the implementation method of the facial key points of the reference image.
  • the facial key points corresponding to the i-th frame of audio are obtained by projecting the three-dimensional facial key points onto a two-dimensional plane
  • the facial key points of the reference image can be two-dimensional facial key points obtained by projecting the three-dimensional facial key points of the reference image onto a two-dimensional plane. This ensures that the method for obtaining the facial key points of the reference image is consistent with the method for obtaining the facial key points corresponding to the i-th frame of audio. This effectively avoids the defects caused by using different two-dimensional key point acquisition mechanisms, thereby improving the generation effect.
  • the three-dimensional facial key points of the reference image are used to describe the state of the facial state presented in the reference image in three-dimensional space; and the three-dimensional facial key points of the reference image are obtained by performing three-dimensional facial key point determination processing on the reference image.
  • the embodiments of the present disclosure do not limit the implementation method of the three-dimensional facial key point determination process.
  • it can adopt any existing or future method that can perform three-dimensional facial key point determination processing on an image, such as a method implemented by using a pre-built machine learning model with three-dimensional facial key point determination processing function.
  • the image corresponding to the i-th frame of audio refers to the image generated for the i-th frame of audio, so that the image corresponding to the i-th frame of audio satisfies the following constraints: the facial expression state presented in the image corresponding to the i-th frame of audio meets the facial expression state requirements of the i-th frame of audio, and other information in the image corresponding to the i-th frame of audio except the facial expression state is consistent with the corresponding information in the above target image, so that the image corresponding to the i-th frame of audio can represent the result of facial expression state adjustment processing of the target image based on the i-th frame of audio.
  • the embodiment of the present disclosure also provides a method for determining the image corresponding to the i-th frame of audio above.
  • the process of determining the image corresponding to the i-th frame of audio may include the following steps 31 to 33.
  • Step 31 For any of the at least two reference images, predict pixel usage description information corresponding to the reference image based on the reference image, the facial key points of the reference image, and the facial key points corresponding to the i-th audio frame; the pixel usage description information includes pixel adjustment description information and/or pixel fusion weights.
  • the pixel usage description information corresponding to the reference image is used to describe how pixels in the reference image are used in the facial image generation process of the i-th frame of audio, so that the pixel usage description information can indicate the impact of the pixels in the reference image on the facial image generation process of the i-th frame of audio.
  • the embodiments of the present disclosure do not limit the implementation method of the pixel usage description information corresponding to the above reference image.
  • the pixel usage description information corresponding to the reference image may include the pixel adjustment description information corresponding to the reference image and/or the pixel fusion weight corresponding to the reference image.
  • this pixel adjustment description information is used to indicate how the pixels in the reference image are adjusted during the facial image generation process for the i-th frame of audio.
  • the disclosed embodiments do not limit the implementation of this pixel adjustment description information; for example, this pixel adjustment description information can be implemented using an offset in image pixel information.
  • this pixel adjustment description information satisfies the following constraints: the size of the pixel adjustment description information is the same as the size of the reference image, and the position coordinates of each pixel point in the pixel adjustment description information represent the offset of the position coordinates of the corresponding pixel point in the reference image.
  • the pixel fusion weight is used to indicate the extent to which each pixel in the reference image affects the facial image generation process of the i-th frame of audio; and the embodiments of the present disclosure do not limit the implementation method of the pixel fusion weight.
  • the pixel fusion weight can meet the following constraints: the size of the pixel fusion weight is the same as the size of the reference image, and the weight value of each pixel point in the pixel fusion weight is used to indicate the influence of the corresponding pixel point in the reference image.
  • the embodiment of the present disclosure does not limit the implementation method of the above step 31.
  • the step 31 can be implemented using the information prediction module in the lip rendering model. It can be seen that under one possible implementation method, the step 31 can be specifically as follows: for any reference image of the at least two reference images above, the information prediction module predicts and outputs the pixel usage description information corresponding to the reference image based on the reference image, the facial key points of the reference image, and the facial key points corresponding to the i-th frame audio.
  • the lip rendering model is used to perform lip rendering processing on the input data of the lip rendering model, such as the lip rendering processing shown in Figure 2 or Figure 3; and the embodiment of the present disclosure does not limit the implementation method of the lip rendering model.
  • the lip rendering model can at least include the information prediction module, such as the information prediction module shown in Figure 3.
  • the information prediction module is used to predict the pixel usage description information corresponding to a reference image based on a reference image, the facial key points of the reference image, and the facial key points corresponding to the i-th frame of audio; and the embodiment of the present disclosure does not limit the implementation method of the information prediction module.
  • the information prediction module may include a convolutional neural network (CNN) and a feature injection module, so that the information prediction module has the function of fusing the reference image, the facial key points of the reference image, and the facial key points corresponding to the i-th frame of audio.
  • the embodiment of the present disclosure does not limit the implementation method of the feature injection module.
  • the feature injection module can be implemented using AdaIN (Adaptive Instance Normalization) or SPADE.
  • step 31 it can be known that after obtaining the above at least two reference image frames, the facial key points of each reference image, and the facial key points corresponding to the i-th frame of audio, these data can be input into the lip rendering model, so that the information prediction module in the lip rendering model can predict and output pixel usage description information corresponding to the n-th reference image based on the n-th reference image, the facial key points of the n-th reference image, and the facial key points corresponding to the i-th frame of audio.
  • the pixel usage description information can indicate how the pixels in the n-th reference image are used in the facial image generation process of the i-th frame of audio, such as position offset + influence strength, where n is a positive integer, n ⁇ N, N is a positive integer, and N represents the number of images in the at least two reference image frames, so that image generation processing can be performed on the i-th frame of audio based on the pixel usage description information corresponding to these reference images.
  • Step 32 Based on the pixel usage description information corresponding to the at least two reference images, perform deformation fusion processing on the at least two reference images to obtain a fused image.
  • the fused image refers to an image obtained by performing deformation fusion processing on the at least two reference images based on the pixel usage description information corresponding to the at least two reference images, so that the fused image is adapted to the i-th frame of audio.
  • step 32 can be: after the information prediction module in the lip rendering model outputs the pixel usage description information corresponding to each reference image, the deformation fusion module in the lip rendering model performs deformation fusion processing on all reference images based on the pixel usage description information corresponding to all reference images, obtains and outputs a fused image, so that the fused image can represent the pixel integration and utilization results of these reference images under the i-th frame audio, so that the fused image is adapted to the i-th frame audio, and further, the fused image can meet the following constraints: the facial expression state presented in the fused image meets the facial expression state requirements of the i-th frame audio, and all or part of the information in the fused image other than the facial expression state is consistent with the corresponding information in the above target image.
  • the deformation fusion module is used to perform pixel integration and utilization processing on these reference images based on the pixel usage description information corresponding to each reference image; and the embodiment of the present disclosure does not limit the implementation of the deformation fusion module.
  • the deformation fusion module can be implemented using any information integration network, such as grid_sample.
  • the information prediction module in the lip rendering model predicts and outputs the pixel usage description information corresponding to the nth reference image based on the nth reference image, the facial key points of the nth reference image, and the facial key points corresponding to the i-th frame of audio, where n is a positive integer, n ⁇ N.
  • the deformation fusion module in the lip rendering model can perform deformation fusion processing on all reference images based on the pixel usage description information corresponding to all reference images to obtain and output a fused image, so that the fused image can represent the facial image generation result for the i-th frame of audio, so that the image corresponding to the i-th frame of audio can be determined based on the facial image generation result.
  • Step 33 Perform image generation processing based on the above fused image, the mouth mask result of the target image, and the facial key points corresponding to the i-th frame of audio to obtain the image corresponding to the i-th frame of audio.
  • the step 33 can be specifically as follows: the face generation module in the lip rendering model performs image generation processing based on the above fused image, the lip mask result of the target image, and the facial key points corresponding to the i-th frame of audio, to obtain and output the image corresponding to the i-th frame of audio, so that the image quality of the image corresponding to the i-th frame of audio is better than the fused image, which is conducive to improving the generation effect.
  • the face generation module is used to optimize the generation process of the fused image based on the lip mask result of the target image and the facial key points corresponding to the i-th frame of audio, so that the image output by the face generation module is better than the fused image, so that the image output by the face generation module is more compatible with the i-th frame of audio; and the embodiment of the present disclosure does not limit the implementation of the face generation module.
  • the face generation module can be implemented using any existing or future image generation network.
  • the lip rendering model can generate and output the image corresponding to the i-th frame of audio based on these data, so that the image corresponding to the i-th frame of audio can represent the facial state of the object in the target image under the i-th frame of audio, thereby realizing expression adjustment processing of the target image based on the i-th frame of audio.
  • S3 Generate a video corresponding to the audio sequence based on the image corresponding to each audio frame in the audio sequence.
  • the embodiments of the present disclosure are not limited to the implementation method of S3 above.
  • the S3 can be specifically: after obtaining the image corresponding to each frame of audio in the audio sequence, the video corresponding to the audio sequence can be generated based on the audio sequence and the image corresponding to each frame of audio in the audio sequence, so that the video corresponding to the audio sequence includes the audio sequence and the image corresponding to each frame of audio in the audio sequence, so that the object in the video corresponding to the audio sequence is consistent with the object in the above reference video, and the facial expression state presented by the video corresponding to the audio sequence is consistent with the facial expression state required by the audio sequence, and then the video corresponding to the audio sequence can express the facial state change of the object in the reference video under the audio sequence, so that the facial expression state adjustment of the object in the video can be driven by the audio sequence.
  • the image corresponding to the i-th frame of audio is first generated based on the facial key points corresponding to the i-th frame of audio in the audio sequence, the lip mask result of the target image corresponding to the i-th frame of audio in the reference video, at least two frames of reference images selected from the reference video, and the facial key points of each reference image, so that the facial expression state presented in the image corresponding to the i-th frame of audio meets the expression requirements of the i-th frame of audio, such as the lip mask requirements, and the i-th frame of audio
  • other information presented in the corresponding image such as facial features and facial posture, is consistent with the corresponding information presented in the target image;
  • i is a positive integer, i ⁇ the total number of frames in the audio sequence; then, based on the images corresponding to each audio frame in the audio sequence
  • the at least two reference image frames can as comprehensively present information that is useful in generating a facial image of the i-th audio frame, such as lip shape and facial posture, appearing in the reference video
  • the image generated based on these reference images for the i-th audio frame can better represent the facial state of the object under the i-th audio frame, thereby enabling the video ultimately generated based on these images to better represent the changes in the facial state of the object under the audio sequence, thereby making the ultimately generated video more adapted to the audio sequence, thereby improving the video generation effect.
  • the embodiments of the present disclosure do not limit the execution subject of the video generation method provided by the embodiments of the present disclosure.
  • the video generation method provided by the embodiments of the present disclosure can be applied to a terminal device or a server.
  • the video generation method provided by the embodiments of the present disclosure can also be implemented with the help of the data interaction process between the terminal device and the server.
  • the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc.
  • the server can be a stand-alone server, a cluster server, or a cloud server.
  • the embodiments of the present disclosure do not limit the application scenarios of the video generation method provided by the embodiments of the present disclosure.
  • the video generation method can be used to complete a certain video generation task, such as a video language switching processing task, a sentence modification processing task in a video, or a sentence replacement processing task in a video; and the data processing logic used when using the video generation method to complete the video generation task is similar to the data processing logic shown in S1-S3 above. For the sake of brevity, it will not be repeated here.
  • the video generation method provided in the embodiment of the present disclosure can be applied to a model training scenario.
  • the embodiment of the present disclosure also provides a model training process, which can specifically include at least the following steps 41 to 43.
  • Step 41 Obtain a reference video and audio sequence, where the reference video and the audio sequence are both determined based on the sample video.
  • step 41 please refer to the relevant content of S1 above.
  • a reference video and an audio sequence are determined from a sample video so that the reference video includes part or all of the images in the sample video, and the audio sequence includes part or all of the audio in the sample video.
  • the target image in the reference video and the i-th frame audio in the audio sequence such as the time corresponding to the target image in the sample video is the same as the time corresponding to the i-th frame audio in the sample video, so that the current round of training process can be completed with the help of the reference video and the audio sequence.
  • Step 42 The lip rendering model generates an image corresponding to the i-th frame of audio based on the facial key points corresponding to the i-th frame of audio in the audio sequence, the lip mask result of the target image corresponding to the i-th frame of audio in the reference video, at least two reference images selected from the reference video, and the facial key points of each reference image; i is a positive integer, i ⁇ the total number of frames in the audio sequence.
  • the lip rendering model is used to perform lip rendering processing on input data of the lip rendering model, such as the lip rendering processing shown in FIG. 2 or FIG. 3 .
  • the lip rendering model may include an information prediction module, a deformation fusion module and a face generation module, so that the working principle of the lip rendering model can be: after the facial key points corresponding to the i-th frame audio in the audio sequence, the lip mask result of the target image corresponding to the i-th frame audio in the reference video, at least two frames of reference images corresponding to the i-th frame audio, and the facial key points of each reference image are input into the lip rendering model, the information prediction module in the lip rendering model first generates a lip mask based on these reference images, the facial key points of these reference images, and the i-th frame audio.
  • the facial key points corresponding to the i-th frame audio are used to predict and output the pixel usage description information corresponding to these reference images; the deformation fusion module in the lip rendering model then performs deformation fusion processing on these reference images based on the pixel usage description information to obtain and output a fused image; then, the face generation module in the lip rendering model generates an image corresponding to the i-th frame audio based on the fused image, the facial key points corresponding to the i-th frame audio, and the lip mask result of the target image, so that the performance of the lip rendering model can be measured based on the difference between the image corresponding to the i-th frame audio and the ground truth (GT) corresponding to the i-th frame audio.
  • GT ground truth
  • the ground truth corresponding to the i-th frame audio is used to guide the facial image generation process of the i-th frame audio; and the embodiment of the present disclosure does not limit the implementation method of the ground truth corresponding to the i-th frame audio.
  • the ground truth corresponding to the i-th frame audio can be implemented using the target image.
  • Step 43 Based on the difference representation data between the target image and the image corresponding to the i-th frame of audio, the face generation module in the lip rendering model is updated, and based on the difference representation data between the target image and the image corresponding to the i-th frame of audio, and the difference representation data between the target image and the fused image, the deformation fusion module and the information prediction module in the lip rendering model are updated, where i is a positive integer, i ⁇ the total number of frames in the audio sequence, and the process returns to continue executing step 41 and subsequent steps until a preset stop condition is reached.
  • the difference representation data between the target image and the image corresponding to the i-th frame of audio is used to represent the difference between the true value corresponding to the i-th frame of audio and the image corresponding to the i-th frame of audio.
  • the embodiments of the present disclosure do not limit the implementation method of the difference representation data.
  • it can be implemented using any existing or future method that can measure the difference between two images.
  • the difference representation data can be based on the similarity between the image features of the target image and the image features of the image corresponding to the i-th frame of audio, such as determined by Euclidean distance, cosine distance, etc.
  • this difference representation data can participate in the update process of all modules in the lip rendering model with the help of gradient feedback, such as the update process of the information prediction module, the update process of the deformation fusion module, and the update process of the face generation module.
  • the difference representation data between the target image and the fused image is used to represent the difference between the true value corresponding to the i-th frame of audio and the fused image.
  • the embodiments of the present disclosure do not limit the implementation of the difference representation data; for example, it can be implemented using any existing or future method that can measure the difference between two images.
  • the difference representation data can be determined based on the similarity between the image features of the target image and the image features of the fused image, such as Euclidean distance, cosine distance, etc.
  • the difference representation data can participate in the update process of other modules in the lip rendering model except the face generation module with the help of gradient backpropagation, such as the update process of the information prediction module and the update process of the deformation fusion module.
  • the preset stop condition refers to the condition that needs to be met when stopping the training of the lip rendering model, and the embodiment of the present disclosure does not limit the implementation method of the preset stop condition.
  • the preset stop condition may specifically include: the model loss of the lip rendering model is lower than a preset loss threshold.
  • the preset stop condition may include: the rate of change of the model loss of the lip rendering model is lower than a preset rate of change threshold.
  • the preset stop condition may include: the number of updates of the lip rendering model reaches a preset number threshold.
  • the model loss of the lip rendering model is used to characterize the performance of the lip rendering model; and the loss of the lip rendering model can be determined based on the difference characterization data between the target image and the image corresponding to the i-th frame of audio, and the difference characterization data between the target image and the fused image. It should be noted that the embodiment of the present disclosure does not limit the calculation method of the loss of the lip rendering model.
  • the information prediction module, deformation fusion module and face generation module in the lip rendering model can be trained simultaneously, and the output results of the face generation module are supervised by GT.
  • the output results of the deformation fusion module are also weakly supervised by GT, such as perceptual loss, which is conducive to improving the model training effect.
  • FIG. 4 is a schematic structural diagram of a video generation device provided in the embodiments of the present disclosure. It should be noted that for the technical details of the video generation device provided in the embodiments of the present disclosure, please refer to the relevant content of the video generation method above.
  • the video generation device 400 provided by the embodiment of the present disclosure includes:
  • the data acquisition unit 401 is used to acquire reference video and audio sequences
  • An image generation unit 402 is configured to generate an image corresponding to the i-th audio frame based on the facial key points corresponding to the i-th audio frame in the audio sequence, a lip mask result of a target image corresponding to the i-th audio frame in the reference video, at least two reference images selected from the reference video, and the facial key points of each reference image, where i is a positive integer and i ⁇ the total number of frames in the audio sequence.
  • the video generating unit 403 is configured to generate a video corresponding to the audio sequence according to the image corresponding to each audio frame in the audio sequence.
  • the process of determining the at least two reference image frames includes: sorting the images in the reference video based on the lip amplitude representation data of each frame image in the reference video to obtain an image sequence; sampling the image sequence at equal intervals to obtain sampled images; and determining the at least two reference image frames based on the sampled images.
  • the at least two frames of reference images are determined based on the sampled image and at least one posture-similar image corresponding to the i-th frame of audio in the reference video; the degree of similarity between the posture representation data of each of the posture-similar images and the posture representation data corresponding to the i-th frame of audio meets a preset similarity requirement.
  • the time corresponding to each of the posture-similar images in the reference video is different from the time corresponding to the i-th frame of audio; the time corresponding to the i-th frame of audio is determined based on the time corresponding to the target image in the reference video.
  • the posture representation data corresponding to the i-th frame of audio is determined based on the posture representation data of the target image.
  • the facial key points corresponding to the i-th frame of audio include facial key point determination results of at least one frame of audio in the audio sequence; and the at least one frame of audio includes the i-th frame of audio.
  • the facial key point determination result of the audio is a two-dimensional facial key point obtained by projecting the three-dimensional facial key point corresponding to the audio onto a two-dimensional plane; the three-dimensional facial key point corresponding to the audio is obtained by performing three-dimensional facial key point determination processing on the audio based on part or all of the images in the reference video.
  • the facial key points of the reference image are two-dimensional facial key points obtained by projecting the three-dimensional facial key points of the reference image onto a two-dimensional plane; the three-dimensional facial key points of the reference image are obtained by performing three-dimensional facial key point determination processing on the reference image.
  • the image generation unit 402 is specifically configured to: for any reference image of the at least two reference images, predict pixel usage description information corresponding to the reference image based on the reference image, the facial key points of the reference image, and the facial key points corresponding to the i-th frame of audio; the pixel usage description information includes pixel adjustment description information and/or pixel fusion weights; perform deformation fusion processing on the at least two reference images based on the pixel usage description information corresponding to the at least two reference images to obtain a fused image; perform image generation processing based on the fused image, the mouth mask result of the target image, and the facial key points corresponding to the i-th frame of audio to obtain an image corresponding to the i-th frame of audio.
  • the pixel usage description information corresponding to each reference image is determined using the information prediction module in the lip rendering model; the deformation fusion processing is implemented using the deformation fusion module in the lip rendering model; and the image generation processing is implemented using the face generation module in the lip rendering model.
  • the reference video and the audio sequence are both determined based on a sample video
  • the video generating device 400 further includes:
  • a model updating unit is configured to update a face generation module in the lip-sync rendering model based on difference representation data between the target image and the image corresponding to the i-th frame of audio, and to update a deformation fusion module and an information prediction module in the lip-sync rendering model based on difference representation data between the target image and the image corresponding to the i-th frame of audio, and difference representation data between the target image and the fused image.
  • the working principle of the video generating device 400 is: after obtaining the reference video and audio sequence, first generate the image corresponding to the i-th frame of audio according to the facial key points corresponding to the i-th frame of audio in the audio sequence, the lip shape mask result of the target image corresponding to the i-th frame of audio in the reference video, at least two frames of reference images selected from the reference video, and the facial key points of each reference image, so that the facial expression state presented in the image corresponding to the i-th frame of audio meets the expression requirements of the i-th frame of audio, such as the lip shape requirements, and makes the i-th frame of audio have the same facial expression as the i-th frame of audio.
  • i is a positive integer, i ⁇ the total number of frames in the audio sequence.
  • a video corresponding to the audio sequence is generated, such that the object presented in the video corresponding to the audio sequence is consistent with the object presented in the reference video, and the video corresponding to the audio sequence can represent the changes in the facial state of the object under the audio sequence, thereby achieving the generation of a video adapted to the audio sequence.
  • the at least two reference images can fully present information present in the reference video that is useful in generating the facial image of the i-th audio frame, such as lip shape and facial posture
  • the image generated based on these reference images for the i-th audio frame can better represent the facial state of the object under the i-th audio frame, thereby enabling the video ultimately generated based on these images to better represent the changes in the facial state of the object under the audio sequence, thereby making the ultimately generated video more adapted to the audio sequence, thereby improving the video generation effect.
  • an embodiment of the present disclosure also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the video generation method provided by the embodiment of the present disclosure.
  • Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers.
  • mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers.
  • PDAs personal digital assistants
  • PADs tablet computers
  • PMPs portable multimedia players
  • in-vehicle terminals e.g., in-vehicle navigation terminals
  • fixed terminals such as digital TVs and desktop computers.
  • the electronic device shown in FIG5 is merely
  • electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503.
  • ROM read-only memory
  • RAM random access memory
  • Various programs and data required for the operation of electronic device 500 are also stored in RAM 503.
  • Processing device 501, ROM 502, and RAM 503 are connected to each other via a bus 504.
  • An input/output (I/O) interface 505 is also connected to bus 504.
  • the following devices may be connected to the I/O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509.
  • the communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data.
  • FIG5 shows the electronic device 500 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
  • an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart.
  • the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502.
  • the processing device 501 When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
  • the embodiments of the present disclosure further provide a computer-readable medium, wherein the computer-readable medium stores instructions or computer programs.
  • the instructions or computer programs are executed on a device, the device executes any implementation of the video generation method provided in the embodiments of the present disclosure.
  • the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two.
  • a computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above.
  • Computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
  • a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component.
  • a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.
  • a computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
  • the program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
  • the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network).
  • HTTP Hyper Text Transfer Protocol
  • Examples of communication networks include a local area network ("LAN”), a wide area network ("WAN”), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
  • the computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
  • the computer-readable medium carries one or more programs.
  • the electronic device can perform the method.
  • Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages.
  • the program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.
  • the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
  • LAN local area network
  • WAN wide area network
  • Internet service provider e.g., AT&T, MCI, Sprint, EarthLink, MSN, GTE, etc.
  • each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function.
  • the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
  • each box in the block diagram and/or flowchart, and the combination of the boxes in the block diagram and/or flowchart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
  • the units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit/module does not, in some cases, limit the unit itself.
  • exemplary types of hardware logic components include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
  • FPGAs field programmable gate arrays
  • ASICs application specific integrated circuits
  • ASSPs application specific standard products
  • SOCs systems on chip
  • CPLDs complex programmable logic devices
  • a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment.
  • a machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium.
  • a machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing.
  • a more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
  • RAM random access memory
  • ROM read-only memory
  • EPROM or flash memory erasable programmable read-only memory
  • CD-ROM portable compact disk read-only memory
  • CD-ROM compact disk read-only memory
  • magnetic storage device or any suitable combination of the foregoing.
  • At least one (item) refers to one or more
  • plural refers to two or more.
  • “And/or” is used to describe the association relationship of associated objects, indicating that three relationships may exist.
  • a and/or B can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural.
  • the character “/” generally indicates that the previous and next associated objects are in an “or” relationship.
  • At least one of the following items” or similar expressions refers to any combination of these items, including any combination of single items or plural items.
  • At least one of a, b or c can represent: a, b, c, "a and b", “a and c", “b and c", or "a and b and c", where a, b, c can be single or multiple.
  • the steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two.
  • the software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Processing Or Creating Images (AREA)
  • Image Analysis (AREA)

Abstract

本公开的实施例公开了一种视频生成方法、装置、设备、介质、产品,该方法包括:在获取到参考视频和音频序列之后,先依据该音频序列中第i帧音频对应的脸部关键点、该第i帧音频在该参考视频中对应的目标图像的口型掩码结果、从该参考视频中选择的至少两帧参考图像、以及各该参考图像的脸部关键点,生成该第i帧音频对应的图像,i为正整数,i≤该音频序列中的总帧数;然后,依据该音频序列中各帧音频对应的图像,生成该音频序列对应的视频。

Description

一种视频生成方法、装置、设备、介质、产品
相关申请的交叉引用
本申请要求申请号为202410418012.X,题为“一种视频生成方法、装置、设备、介质、产品”、申请日为2024年4月8日的中国发明专利申请的优先权,通过引用方式将该申请整体并入本文。
技术领域
本公开的实施例涉及数据处理技术领域,尤其涉及一种视频生成方法、装置、设备、介质、产品。
背景技术
目前一些应用场景较为普遍,如视频的语种切换处理、视频中语句修改处理、或者视频中语句替换处理等场景。
发明内容
本公开实施例提供了一种视频生成方法、装置、设备、介质、产品,有利于提高视频生成效果。
为了实现上述目的,本公开实施例提供的技术方案如下:
本公开实施例提供一种视频生成方法,所述方法包括:
获取参考视频和音频序列;
依据所述音频序列中第i帧音频对应的脸部关键点、所述第i帧音频在所述参考视频中对应的目标图像的口型掩码结果、从所述参考视频中选择的至少两帧参考图像、以及各所述参考图像的脸部关键点,生成所述第i帧音频对应的图像;i为正整数,i≤所述音频序列中的总帧数;
依据所述音频序列中各帧音频对应的图像,生成所述音频序列对应的视频。
在一种可能的实施方式下,所述至少两帧参考图像的确定过程,包括:
依据所述参考视频中各帧图像的口型幅度表征数据,对所述参考视频中图像进行排序,得到图像序列;
对所述图像序列进行等间隔采样,得到采样图像;
依据所述采样图像,确定所述至少两帧参考图像。
在一种可能的实施方式下,所述至少两帧参考图像是依据所述采样图像、以及所述第i帧音频在所述参考视频中对应的至少一个姿态相似图像所确定的;
各所述姿态相似图像的姿态表征数据与所述第i帧音频对应的姿态表征数据之间的相似程度达到预设相似需求。
在一种可能的实施方式下,各所述姿态相似图像在所述参考视频中对应的时间不同于所述第i帧音频对应的时间;
所述第i帧音频对应的时间是依据所述目标图像在所述参考视频中对应的时间所确定的。
在一种可能的实施方式下,所述第i帧音频对应的姿态表征数据是依据所述目标图像的姿态表征数据所确定的。
在一种可能的实施方式下,所述第i帧音频对应的脸部关键点包括所述音频序列中至少一帧音频的脸部关键点确定结果;
所述至少一帧音频包括所述第i帧音频。
在一种可能的实施方式下,对于所述至少一帧音频中任一音频,该音频的脸部关键点确定结果是通过将该音频对应的三维脸部关键点投影至二维平面所得到的二维脸部关键点;
该音频对应的三维脸部关键点是通过依据所述参考视频中部分或者全部图像,对该音频进行三维脸部关键点确定处理所得到的。
在一种可能的实施方式下,对于所述至少两帧参考图像中任一参考图像,该参考图像的脸部关键点是通过将该参考图像的三维脸部关键点投影至二维平面所得到的二维脸部关键点;
该参考图像的三维脸部关键点是通过对该参考图像进行三维脸部关键点确定处理所得到的。
在一种可能的实施方式下,所述第i帧音频对应的图像的确定过程,包括:
对于所述至少两帧参考图像中任一参考图像,依据该参考图像、该参考图像的脸部关键点、以及所述第i帧音频对应的脸部关键点,预测该参考图像对应的像素使用描述信息;所述像素使用描述信息包括像素调整描述信息和/或像素融合权重;
依据所述至少两帧参考图像对应的像素使用描述信息,对所述至少两帧参考图像进行形变融合处理,得到融合图像;
依据所述融合图像、所述目标图像的口型掩码结果、以及所述第i帧音频对应的脸部关键点进行图像生成处理,得到所述第i帧音频对应的图像。
在一种可能的实施方式下,各所述参考图像对应的像素使用描述信息均是利用口型渲染模型中信息预测模块所确定的;
所述形变融合处理是利用所述口型渲染模型中形变融合模块所实现的;
所述图像生成处理是利用所述口型渲染模型中脸部生成模块所实现的。
在一种可能的实施方式下,所述参考视频以及所述音频序列均是依据样本视频所确定的;
所述生成所述第i帧音频对应的图像之后,所述方法还包括:
依据所述目标图像与所述第i帧音频对应的图像之间的差异表征数据,更新所述口型渲染模型中脸部生成模块,并依据所述目标图像与所述第i帧音频对应的图像之间的差异表征数据、以及所述目标图像与所述融合图像之间的差异表征数据,更新所述口型渲染模型中形变融合模块和信息预测模块。
本公开实施例提供了一种视频生成装置,包括:
数据获取单元,用于获取参考视频和音频序列;
图像生成单元,用于依据所述音频序列中第i帧音频对应的脸部关键点、所述第i帧音频在所述参考视频中对应的目标图像的口型掩码结果、从所述参考视频中选择的至少两帧参考图像、以及各所述参考图像的脸部关键点,生成所述第i帧音频对应的图像;i为正整数,i≤所述音频序列中的总帧数;
视频生成单元,用于依据所述音频序列中各帧音频对应的图像,生成所述音频序列对应的视频。
本公开实施例提供了一种电子设备,其特征在于,所述设备包括:处理器和存储器;
所述存储器,用于存储指令或计算机程序;
所述处理器,用于执行所述存储器中的所述指令或计算机程序,以使得所述电子设备执行本公开实施例提供的视频生成方法。
本公开实施例提供了一种计算机可读介质,其特征在于,所述计算机可读介质中存储有指令或计算机程序,当所述指令或计算机程序在设备上运行时,使得所述设备执行本公开实施例提供的视频生成方法。
本公开实施例提供了一种计算机程序产品,其特征在于,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行本公开实施例提供的视频生成方法的程序代码。
附图说明
为了更清楚地说明本公开实施例或相关技术中的技术方案,下面将对实施例或相关技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本公开实施例中记载的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其它的附图。
图1为本公开实施例提供的一种视频生成方法的流程图;
图2为本公开实施例提供的一种视频生成流程的示意图;
图3为本公开实施例提供的一种图像生成流程的示意图;
图4为本公开实施例提供的一种视频生成装置的结构示意图;
图5为本公开实施例提供的一种电子设备的结构示意图。
具体实施方式
为了使本技术领域的人员更好地理解本公开实施例方案,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅是本公开实施例一部分实施例,而不是全部的实施例。基于本公开实施例中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本公开实施例保护的范围。
如上所述,对于一些应用场景,如视频的语种切换处理、视频中语句修改处理、或者视频中语句替换处理等场景来说,这些应用场景可能存在以下需求:依据已有的视频,生成与某个音频序列适配的视频。
为了更好地理解本公开实施例所提供的技术方案,下面先结合一些附图对本公开实施例提供的视频生成方法进行说明。如图1所示,本公开实施例提供的视频生成方法,包括下文S1-S3。其中,该图1为本公开实施例提供的一种视频生成方法的流程图。
S1:获取参考视频和音频序列。
其中,参考视频是指在视频生成时所需依据的视频,如图2所示的参考视频,以使该参考视频用于提供除了脸部表情状态,如口型状态以外的其他信息,如类似于脸部轮廓特点以及五官分布特点的脸部标识性信息、脸部姿态信息等。
另外,本公开实施例不限定上文参考视频的实施方式,为了便于理解,下面结合两个场景进行说明。
场景一,当本公开实施例提供的视频生成方法用于执行某个生成任务,如视频生成任务时,上文参考视频可以是由用户通过一定手段所指定的视频,如单人口播视频等,以使该参考视频符合该用户的一些需求,如脸部标识性信息、脸部姿态等方面的需求。需要说明的是,本公开实施例不限定该手段,比如,该参考视频可以是由用户手动上传的视频,或者由用户从一些候选视频中挑选出的视频,或者由用户借助某些方式下载好的视频。
场景二,当本公开实施例提供的视频生成方法用于实现模型训练过程时,上文参考视频可以是依据样本视频所确定的,以使该参考视频包括该样本视频中部分或者全部图像。其中,该样本视频是指在模型训练过程中所需使用的视频;而且本公开实施例不限定该样本视频的实施方式。
此外,本公开实施例不限定上文参考视频的获取方式。
音频序列是指在视频生成时所需依据的音频,如图2所示的音频序列,以使该音频序列用于提供脸部表情状态,如口型状态等信息。
另外,本公开实施例不限定上文音频序列的实施方式,为了便于理解,下面结合两个场景进行说明。
场景一,当本公开实施例提供的视频生成方法用于执行某个生成任务,如视频生成任务时,上文音频序列可以是由用户借助一定手段所指定的;而且本公开实施例不限定该音频序列的实施方式,比如,在一些应用场景,如视频翻译场景下,该音频序列满足以下约束:该音频序列所描述的语句的语义信息与参考视频中所描述的语句的语义信息保持一致,但是该音频序列所描述的语句的语种不同于该参考视频所描述的语句的语种。又如,在一些应用场景,如视频语句修改场景下,该音频序列至少满足以下约束:该音频序列所描述的语句与参考视频中所描述的语句的之间部分相同。还如,在一些应用场景,如视频语句更换场景下,该音频序列至少满足以下约束:该音频序列所描述的语句完全不同于参考视频中所描述的语句。
场景二,当本公开实施例提供的视频生成方法用于实现模型训练过程时,如果上文参考视频是依据样本视频所确定的,则上文音频序列可以是依据该样本视频所确定的,以使该音频序列包括该样本视频中部分或者全部音频。
此外,本公开实施例不限定上文音频序列与上文参考视频之间的关联关系,比如,两者可以满足以下约束:对于该音频序列中的第i帧音频来说,该参考视频中存在该第i帧音频对应的目标图像。其中,该目标图像是指该参考视频中存在的、与该第i帧音频存在对应关系的图像,以使该目标图像用于为该第i帧音频的脸部图像生成过程提供除了脸部表情状态以外的其他信息,如脸部特点、脸部姿态等信息;而且本公开实施例不限定该目标图像的实施方式,比如,当该音频序列中的总帧数与该参考视频中的总帧数相同时,该目标图像可以是指该参考视频中的第i帧图像。其中,i为正整数,i≤I,I表示该音频序列的总帧数。
需要说明的是,本公开实施例不限定上段中对应关系的实施方式,比如,当音频序列的总帧数等于参考视频的总帧数时,这种对应关系具体可以包括:该音频序列中的第i帧音频与该参考视频中的第i帧图像之间的对应关系。又如,当该音频序列的总帧数大于该参考视频的总帧数时,这种对应关系具体可以包括:该音频序列中的第i帧音频与增帧后的参考视频中的第i帧图像之间的对应关系。其中,该增帧后的参考视频是通过针对原有的参考视频进行某种增帧处理所得到的,以使该增帧后的参考视频的总帧数等于该音频序列的总帧数;而且本公开实施例不限定该增帧处理的实施方式,比如,其可以采用现有的或者未来出现的任意一种能够增加一个视频的总帧数的方法,如针对原视频进行插帧处理或者针对原视频进行复制拼接处理等方法进行实施。还如,当该音频序列的总帧数小于该参考视频的总帧数时,这种对应关系具体可以包括:该音频序列中的第i帧音频与减帧后的参考视频中的第i帧图像之间的对应关系。其中,该减帧后的参考视频是通过针对原有的参考视频进行某种减帧处理所得到的,以使该减帧后的参考视频的总帧数等于该音频序列的总帧数;而且本公开实施例不限定该减帧处理的实施方式,比如,其可以采用现有的或者未来出现的任意一种能够减少一个视频的总帧数的方法,如针对原视频进行采样处理或者针对原视频进行裁剪处理等方法进行实施。其中,i为正整数,i≤I,I表示该音频序列的总帧数。
还有,本公开实施例不限定上文音频序列的获取方式。
S2:依据音频序列中第i帧音频对应的脸部关键点、该第i帧音频在参考视频中对应的目标图像的口型掩码结果、从参考视频中选择的至少两帧参考图像、以及各参考图像的脸部关键点,生成该第i帧音频对应的图像;i为正整数,i≤音频序列中的总帧数。
其中,第i帧音频对应的脸部关键点用于表示参考视频中对象在该第i帧音频下所处的脸部表情状态,如口型状态等。需要说明的是,本公开实施例不限定该对象的实施方式,比如,该对象可以采用任意一种能够呈现表情的动物或者虚拟形象进行实施。
另外,本公开实施例不限定上文第i帧音频对应的脸部关键点的实施方式,比如,其可以采用现有的或者未来出现的任意一种能够用于表征脸部表情状态的关键点进行实施。又如,在一些应用场景下,为了更好地提高生成效果,该第i帧音频对应的脸部关键点可以采用二维脸部关键点进行实施。其中,该二维脸部关键点用于描述一个对象在二维空间中所处的脸部表情状态。
可见,在一种可能的实施方式下,上文第i帧音频对应的脸部关键点可以包括该第i帧音频的二维脸部关键点。其中,该第i帧音频的二维脸部关键点用于描述参考视频中对象在该第i帧音频下所处的二维的脸部表情状态;而且该第i帧音频的二维脸部关键点可以是依据该第i帧音频与该参考视频所确定的。
另外,本公开实施例不限定上文第i帧音频的二维脸部关键点的确定方式,比如,其可以采用现有的或者未来出现的任意一种能够依据一个视频以及一帧音频进行二维脸部关键点生成处理的方法,如借助预先构建的二维关键点生成模型所实现的方法进行实施。其中,该二维关键点生成模型用于针对该二维关键点生成模型的输入数据,如视频+音频进行二维脸部关键点生成处理;而且本公开实施例不限定该二维关键点生成模型的实施方式。
此外,为了更好地提高生成效果,本公开实施例还提供了上文第i帧音频的二维脸部关键点的一种确定方式,在该方式下,该第i帧音频的二维脸部关键点可以是通过将该第i帧音频对应的三维脸部关键点投影至二维平面所得到的。其中,该第i帧音频对应的三维脸部关键点用于描述参考视频中对象在该第i帧音频下所处的三维的脸部表情状态;而且该第i帧音频对应的三维脸部关键点是通过依据该参考视频中部分或者全部图像,对该第i帧音频进行三维脸部关键点确定处理所得到的。需要说明的是,本公开实施例不限定该三维脸部关键点确定处理的实施方式,比如,其可以采用现有的或者未来出现的任意一种能够依据一些图像以及一帧音频进行三维脸部关键点确定处理的方法进行实施,如借助预先构建的三维关键点生成模型所实现的方法进行实施。其中,该三维关键点生成模型用于针对该三维关键点生成模型的输入数据,如视频+音频进行三维脸部关键点生成处理;而且本公开实施例不限定该三维关键点生成模型的实施方式。
又如,在一些应用场景下,为了更好地提高生成效果,上文第i帧音频对应的三维脸部关键点的确定过程可以包括下文步骤11-步骤12。
步骤11:依据参考视频以及第i帧音频,预测该第i帧音频对应的三维脸部参数。
其中,该第i帧音频对应的三维脸部参数用于描述参考视频中对象在该第i帧音频下所处的脸部状态,如脸部特点、脸部表情、脸部姿态等。
另外,本公开实施例不限定第i帧音频对应的三维脸部参数的实施方式,比如,该第i帧音频对应的三维脸部参数可以包括脸部标识(Identity document,ID)参数、脸部表情参数以及脸部姿态参数。其中,该脸部标识参数用于描述参考视频中对象所具有的标识性的脸部特点,如脸部轮廓、五官分布等特点,以使基于该脸部标识参数所构建的三维脸部模型处于有ID、无姿态以及无表情这一状态。该脸部表情参数用于描述在该第i帧音频下的脸部表情状态,如口型状态等,以使基于该脸部表情参数所构建的三维脸部模型处于无ID、无姿态以及有表情这一状态。该脸部姿态参数用于描述在该第i帧音频下的脸部姿态,如侧脸、正脸等姿态,以使基于该脸部姿态参数所构建的三维脸部模型处于无ID、有姿态以及无表情这一状态。需要说明的是,本公开实施例不限定这三种参数之间的关联关系,比如,这三种参数之间处于互相解耦状态。
此外,本公开实施例不限定上文第i帧音频对应的三维脸部参数的确定过程,比如,其可以采用现有的或者未来出现的任意一种能够基于一个视频以及一帧音频进行三维脸部参数确定处理的方法,如借助预先构建的脸部参数确定模型所实现的方法进行实施。其中,该脸部参数确定模型是指预先构建的、具有三维脸部参数确定功能的模型,如机器学习模型;而且本公开实施例不限定该脸部参数确定模型的实施方式。
还有,在一些应用场景,如口型修改场景下,为了更好地提高生成效果,上文第i帧音频对应的三维脸部参数满足以下约束:①该第i帧音频对应的三维脸部参数中脸部标识参数和脸部表情参数均是依据参考视频以及第i帧音频所确定的,以使由该脸部标识参数所描述的脸部特点与该参考视频中所呈现的脸部特点保持一致,并使得由该脸部表情参数所表征的脸部表情状态满足该第i帧音频的脸部表情状态需求;②该第i帧音频对应的三维脸部参数中脸部姿态参数是依据上文目标图像的三维脸部参数中脸部姿态参数所确定的,以使该第i帧音频对应的三维脸部参数中脸部姿态参数与该目标图像的三维脸部参数中脸部姿态参数保持一致。其中,该目标图像的三维脸部参数是通过对该目标图像进行三维脸部参数确定处理所得到的;而且本公开实施例不限定该目标图像的三维脸部参数的确定过程,比如,其可以采用现有的或者未来出现的任意一种能够针对一个图像进行三维脸部参数确定处理的方法进行实施。
步骤12:依据第i帧音频对应的三维脸部参数,确定该第i帧音频对应的三维脸部关键点。
需要说明的是,本公开实施例不限定上文步骤12的实施方式,比如,其可以采用现有的或者未来出现的任意一种能够基于三维脸部参数确定三维脸部关键点的方法进行实施。
基于上文步骤11至步骤12的相关内容可知,在一些应用场景,如图2所示的场景下,对于音频序列中第i帧音频来说,可以先依据参考视频进行脸部重建处理,得到该参考视频中各帧图像的三维脸部参数;再依据该参考视频中部分或者全部图像的三维脸部参数以及该第i帧音频的音频特征,预测该第i帧音频对应的三维脸部参数;然后,依据该第i帧音频对应的三维脸部参数,确定该第i帧音频对应的三维脸部关键点,以使该三维脸部关键点能够更好地表示出该参考视频中对象在该第i帧视频下所处的脸部状态,从而使得基于该三维脸部关键点所确定的该第i帧音频对应的脸部关键点更准确,进而有利于提高生成效果。
实际上,为了更好地提高生成效果,本公开实施例还提供了上文第i帧音频对应的脸部关键点的一种可能的实施方式,在该实施方式下,该第i帧音频对应的脸部关键点可以包括音频序列中至少一帧音频的脸部关键点确定结果,而且该至少一帧音频包括该第i帧音频。其中,该至少一帧音频是指该音频序列中存在的、在生成该第i帧音频的脸部图像时所需参考的音频;而且本公开实施例不限定该至少一帧音频的实施方式,比如,该至少一帧音频可以包括该音频序列中第i-T帧音频、第i-T+1帧音频、……、第i帧音频、第i+1帧音频、……、以及第i+T帧音频,以使该至少一帧音频能够更好地表示出由该第i帧音频所携带的音频信息及其上下文信息,从而使得基于该至少一帧音频所生成的图像更准确。
另外,对于上文至少一帧音频中任一音频来说,该音频的脸部关键点确定结果是指针对该音频所确定的脸部关键点,以使该脸部关键点确定结果能够表示出参考视频中对象在该音频下所处的脸部状态;而且本公开实施例不限定该音频的脸部关键点确定结果的实施方式,比如,该音频的脸部关键点确定结果可以采用三维脸部关键点和/或二维脸部关键点进行实施。可见,在一种可能的实施方式下,该音频的脸部关键点确定结果可以是通过将该音频对应的三维脸部关键点投影至二维平面所得到的二维脸部关键点。其中,该音频对应的三维脸部关键点是通过依据该参考视频中部分或者全部图像,对该音频进行三维脸部关键点确定处理所得到的;而且该音频对应的三维脸部关键点的实施方式类似于上文第i帧音频对应的三维脸部关键点的实施方式,为了简要起见,在此不再赘述。
基于上文第i帧音频对应的脸部关键点的相关内容可知,在一种可能的实施方式下,在获取到音频序列中各帧音频对应的三维脸部关键点之后,可以先依据该音频序列中至少一帧音频对应的三维脸部关键点,确定该至少一帧音频的脸部关键点确定结果;再依据该至少一帧音频的脸部关键点确定结果,确定该第i帧音频对应的脸部关键点,以使该第i帧音频对应的脸部关键点包括该至少一帧音频的脸部关键点确定结果,从而使得该第i帧音频对应的脸部关键点能够更好地表示出参考视频中对象在该第i帧音频下所处的脸部状态,如脸部特点、脸部表情、脸部姿态、上下文信息等,从而使得基于该第i帧音频对应的脸部关键点所生成的图像更好。
另外,对于上文第i帧音频在参考视频中对应的目标图像来说,该目标图像是指该参考视频中存在的、用于为该第i帧音频的脸部图像生成处理提供除了脸部表情状态以外的其他信息的图像,如该参考视频中第i帧图像等;而且该目标图像的口型掩码结果用于提供该目标图像中除了口型状态以外的其他信息;而且本公开实施例不限定该口型掩码结果的确定过程,比如,其具体可以为:对该目标图像中下半脸进行掩码处理,得到该目标图像的口型掩码结果,以使该口型掩码结果能够表示出该目标图像中上半脸,从而使得该口型掩码结果能够提供该目标图像中除了口型状态以外的一些其他信息,如肤色、光线等信息,进而使得基于该口型掩码结果所生成的图像更好。
此外,对于上文至少两帧参考图像来说,该至少两帧参考图像是指从参考视频中选择的图像,以使该至少两帧参考图像能够向第i帧音频的脸部图像生成过程提供一些具有使用价值的像素信息,如姿态、口型等方面所涉及的像素信息。
另外,本公开实施例不限定上文至少两帧参考图像的确定过程,比如,在针对第i帧音频进行图像生成时,该至少两帧参考图像可以是从该参考视频中随机选择的,如针对该第i帧音频所随机选择的。可见,在一些应用场景下,上文音频序列中不同帧音频对应的至少两帧参考图像是指从该参考视频中随机选择的同一组图像。又如,在一些应用场景下,该音频序列中不同帧音频对应的至少两帧参考图像是指从该参考视频中随机选择的不同组图像。
此外,为了更好地提高生成效果,本公开实施例还提供了上文至少两帧参考图像的一种确定方式,在该方式下,该至少两帧参考图像的确定过程可以包括下文步骤21-步骤23。
步骤21:依据参考视频中各帧图像的口型幅度表征数据,对该参考视频中图像进行排序,得到图像序列。
其中,参考视频中第j帧图像是指该参考视频中存在的、处于第j个排列位置上的图像,j为正整数,j≤J,J为正整数,J表示该参考视频中的总帧数。
另外,对于参考视频中第j帧图像来说,该第j帧图像的口型幅度表征数据用于表征该第j帧图像中对象的口型张开幅度,而且本公开实施例不限定该第j帧图像的口型幅度表征数据的确定方式,比如,其可以采用现有的或者未来出现的任意一种能够针对一个图像进行口型幅度确定处理的的方法,如借助预先构建的口型幅度确定模型所实现的方法进行实施。其中,该口型幅度确定模型是指预先构建的、能够针对一个图像进行口型幅度确定处理的模型,如机器学习模型等。
此外,为了更好地提高生成效果,本公开实施例还提供了上文第j帧图像的口型幅度表征数据的一种确定方式,在该方式下,该第j帧图像的口型幅度表征数据的确定过程可以为:依据该第j帧图像的脸部关键点,如二维脸部关键点和/或三维脸部关键点中的嘴部关键点,确定该第j帧图像的口型幅度表征数据,以使该口型幅度表征数据能够更准确地表示出该第j帧图像中对象的口型张开幅度。其中,该第j帧图像的三维脸部关键点用于表示该第j帧图像中所呈现的脸部状态在三维空间中所处状态;而且该第j帧图像的三维脸部关键点是通过对该第j帧图像进行三维脸部关键点确定处理所得到的。需要说明的是,本公开实施例不限定该三维脸部关键点确定处理的实施方式。该第j帧图像的二维脸部关键点用于表示该第j帧图像中对象在二维空间中所处的脸部状态;而且本公开实施例不限定该第j帧图像的二维脸部关键点的确定方法,比如,其可以采用现有的或者未来出现的任意一种能够针对一个图像进行二维脸部关键点确定处理的方法进行实施。又如,该第j帧图像的二维脸部关键点可以是通过将该第j帧图像的三维脸部关键点投影至二维平面所得到的,如此有利于提高关键点相关信息的一致性,从而有利于提高生成效果。
还有,对于上文步骤21中的图像序列来说,该图像序列的确定过程可以为:依据参考视频中各帧图像的口型幅度表征数据,对该参考视频中部分或者全部图像进行排序,得到图像序列,以使该图像序列中所有图像是按照口型幅度表征数据进行升序或者降序排列的,从而使得该图像序列能够更好地描述出该参考视频中所呈现的口型幅度分布情况。可见,在一种可能的实施方式下,该图像序列可以包括该参考视频中所有图像;而且该图像序列中所有图像是按照口型幅度表征数据进行升序或者降序排列的。
基于上文步骤21的相关内容可知,对于一些应用场景来说,在获取到参考视频之后,可以先确定该参考视频中各帧图像的口型幅度表征数据;再将这些图像按照口型幅度表征数据进行升序或者降序排列,得到图像序列,以使该图像序列包括该参考视频中部分或者全部图像,并使得该图像序列中各个图像的排列位置与该参考视频中相应的图像的口型幅度表征数据正相关或者负相关,从而使得该图像序列能够更好地表示出该参考视频中所呈现的口型变化情况。
步骤22:对上文图像序列进行等间隔采样,得到采样图像。
其中,采样图像是指从图像序列中采样所得的图像;而且本公开实施例不限定该采样时所需参考的采样间隔,比如,该采样间隔可以依据实际的需求,如生成效率需求或者生成准确性需求综合确定。
基于上文步骤22的相关内容可知,对于一些应用场景来说,在获取到上文图像序列之后,可以从该图像序列中等间隔采样一些图像,作为采样图像,以使这些采样图像能够表示出参考视频中对象的不同口型状态,从而使得这些采样图像能够为第i帧音频的脸部图像生成过程提供有用的口型信息,如此有利于提高生成效果。
步骤23:依据上文采样图像,确定上文至少两帧参考图像。
需要说明的是,本公开实施例不限定上文步骤23的实施方式,比如,该步骤23具体可以为:将上文采样图像,确定为上文至少两帧参考图像,以使该至少两帧参考图像能够为音频序列中各帧图像的脸部图像生成过程提供一些有用信息,如口型信息。可见,在一种可能的实施方式下,该音频序列中不同帧图像的脸部图像生成过程所涉及的至少两帧参考图像均是这些采样图像,从而使得该音频序列中不同帧图像的脸部图像生成过程所涉及的至少两帧参考图像保持相同。
实际上,为了更好地提高生成效果,本公开实施例还提供了上文步骤23的一种可能的实施方式,在该实施方式下,该步骤23具体可以为:依据上文采样图像、以及第i帧音频在参考视频中对应的至少一个姿态相似图像,确定该第i帧音频对应的至少两帧参考图像,以使该第i帧音频对应的至少两帧参考图像包括该采样图像和该至少一个姿态相似图像,从而使得该第i帧音频对应的至少两帧参考图像能够为该第i帧音频的脸部图像生成过程提供尽可能全面的有用信息,如口型信息+姿态信息等,以便后续能够基于该至少两帧参考图像更好地生成该第i帧音频对应的图像。其中,该至少一个姿态相似图像是指该参考视频中存在的、与该第i帧音频对应的姿态最相近的一些图像。
另外,本公开实施例不限定上段中至少一个姿态相似图像的实施方式,比如,该至少一个姿态相似图像可以满足以下约束:各姿态相似图像的姿态表征数据与第i帧音频对应的姿态表征数据之间的相似程度达到预设相似需求。需要说明的是,本公开实施例不限定该预设相似需求的实施方式,比如,该预设相似需求具体可以为:各姿态相似图像的姿态表征数据与第i帧音频对应的姿态表征数据之间的相似程度均超过预设相似度阈值。又如,在依据参考视频中各帧图像的姿态表征数据与第i帧音频对应的姿态表征数据之间的相似程度,将该参考视频中所有图像按照相似程度进行降序排序,得到排序结果之后,该预设相似需求具体可以为:各姿态相似图像在该排序结果中所处的排列位置均比预设排列位置阈值靠前。
此外,对于上文至少一个姿态相似图像中任一姿态相似图像来说,该姿态相似图像的姿态表征数据用于表征该姿态相似图像中对象所处的脸部姿态;而且本公开实施例不限定该姿态相似图像的姿态表征数据的确定过程,比如,其可以采用现有的或者未来出现的任意一种能够针对一个图像进行姿态提取处理的方法进行实施。又如,为了更好地提高生成效果,该姿态相似图像的姿态表征数据的确定过程可以为:依据该姿态相似图像的三维脸部参数中脸部姿态参数,确定该姿态相似图像的姿态表征数据。其中,该姿态相似图像的三维脸部参数用于表征该姿态相似图像中对象在三维空间中所处的脸部状态;而且该姿态相似图像的三维脸部参数的实施方式类似于上文第i帧音频对应的三维脸部参数的实施方式。需要说明的是,本公开实施例不限定“依据该姿态相似图像的三维脸部参数中脸部姿态参数,确定该姿态相似图像的姿态表征数据”这一步骤的实施方式,比如,其具体可以为:先利用该姿态相似图像的三维脸部参数中脸部姿态参数,确定该姿态相似图像的三维脸部关键点,以使该姿态相似图像的三维脸部关键点能够表示出在无ID、无表情以及有姿态下所处的脸部状态;再将该姿态相似图像的三维脸部关键点,确定为该姿态相似图像的姿态表征数据。
还有,对于上文第i帧音频对应的姿态表征数据来说,该姿态表征数据用于表示该第i帧音频下的脸部姿态;而且本公开实施例不限定该姿态表征数据的确定方式,比如,为了更好地提高生成效果,该第i帧音频对应的姿态表征数据是依据上文目标图像的姿态表征数据所确定的。其中,该目标图像的姿态表征数据用于描述该目标图像中所呈现的脸部姿态;而且该目标图像的姿态表征数据的实施方式类似于上文姿态相似图像的姿态表征数据的实施方式,为了简要起见,在此不再赘述。
再者,为了更好的提高生成效果,本公开实施例还提供了上文至少一个姿态相似图像的一种可能的实施方式,在该实施方式下,该至少一个姿态相似图像可以满足以下约束:各姿态相似图像的姿态表征数据与第i帧音频对应的姿态表征数据之间的相似程度达到预设相似需求,而且各姿态相似图像在参考视频中对应的时间不同于该第i帧音频对应的时间。其中,该第i帧音频对应的时间是依据该目标图像在参考视频中对应的时间所确定的;而且本公开实施例不限定该第i帧音频对应的时间的确定方式,比如,其具体可以为:将该目标图像在参考视频中对应的时间,确定为该第i帧音频对应的时间。
基于上段内容可知,在一种可能的实施方式下,当上文第i帧音频对应的时间与目标图像在参考视频中对应的时间保持一致时,该第i帧音频对应的至少一个姿态相似图像可以满足以下约束:各姿态相似图像的姿态表征数据与第i帧音频对应的姿态表征数据之间的相似程度达到预设相似需求,而且各姿态相似图像在参考视频中对应的时间不同于该目标图像在参考视频中对应的时间,以使该第i帧音频对应的至少一个姿态相似图像可以包括该参考视频中存在的、除了该目标图像以外的、与该第i帧音频对应的姿态最接近的多帧图像,从而能够有效地的避免该目标图像可能造成的一些干扰,进而使得这些姿态相似图像能够更好地为该第i帧音频的脸部图像生成过程提供有用的姿态信息,如此有利于提高生成效果。其中,该第i帧音频对应的姿态是指该目标图像中所呈现的脸部姿态。
基于上文步骤21至步骤23的相关内容可知,本公开实施例提供了上文至少两帧参考图像的两种可能的实施方式,一种为全固定,比如,通过对参考视频进行口型方面的相关处理,以从该参考视频中选择一些具有不同口型状态的图像,如上文采样图像,以便后续能够将这些具有不同口型状态的图像应用到音频序列中各帧音频的脸部图像生成过程,以使该音频序列中不同帧音频的脸部图像生成过程所参考的至少两帧参考图像保持一致。另一种为部分固定部分动态,比如,通过对参考视频进行口型方面的相关处理,以从该参考视频中选择一些具有不同口型状态的图像,作为固定部分,同时通过对该参考视频进行姿态方面的相关处理,以从该参考视频中选择一些姿态接近于第i帧音频对应的姿态的图像,作为该第i帧音频对应的动态部分,以便后续能够将该固定部分+该第i帧音频对应的动态部分应用到该第i帧音频的脸部图像生成过程,如此有利于更好地提高生成效果。
另外,对于上文至少两帧参考图像中任一参考图像来说,该参考图像的脸部关键点用于描述该参考图像中所呈现的脸部状态,而且本公开实施例不限定该参考图像的脸部关键点的实施方式,比如,该参考图像的脸部关键点的实施方式类似于上文第i帧音频对应的脸部关键点的实施方式。可见,在一种可能的实施方式下,如果该第i帧音频对应的脸部关键点采用二维脸部关键点进行实施,则该参考图像的脸部关键点也可以采用二维脸部关键点进行实施,而且该参考图像的脸部关键点与该第i帧音频对应的脸部关键点之间满足以下约束:该参考图像的脸部关键点与该第i帧音频对应的脸部关键点在关键点序号上是对应的。
此外,对于上文至少两帧参考图像中任一参考图像来说,本公开实施例不限定该参考图像的脸部关键点的实施方式,比如,如果上文第i帧音频对应的脸部关键点是通过对三维脸部关键点投影至二维平面所得到的,则该参考图像的脸部关键点可以是通过将该参考图像的三维脸部关键点投影至二维平面所得到的二维脸部关键点,以使该参考图像的脸部关键点的获取方式与该第i帧音频对应的脸部关键点的获取方式保持一致,如此能够有效地避免当采用不同二维关键点获取机制时所造成的缺陷,从而有利于提高生成效果。其中,该参考图像的三维脸部关键点用于描述该参考图像中所呈现的脸部状态在三维空间中所处状态;而且该参考图像的三维脸部关键点是通过对该参考图像进行三维脸部关键点确定处理所得到的。需要说明的是,本公开实施例不限定该三维脸部关键点确定处理的实施方式,比如,其可以采用现有的或者未来出现的任意一种能够针对一个图像进行三维脸部关键点确定处理的方法,如借助预先构建的具有三维脸部关键点确定处理功能的机器学习模型所实现的方法进行实施。
还有,对于音频序列中第i帧音频来说,该第i帧音频对应的图像是指针对该第i帧音频所生成的图像,以使该第i帧音频对应的图像满足以下约束:该第i帧音频对应的图像中所呈现的脸部表情状态满足该第i帧音频的脸部表情状态需求,而且该第i帧音频对应的图像中除了脸部表情状态以外的其他信息与上文目标图像中相应的信息保持一致,从而使得该第i帧音频对应的图像能够表示出依据该第i帧音频,对该目标图像进行脸部表情状态调整处理的结果。
实际上,为了更好地提高生成效果,本公开实施例还提供了上文第i帧音频对应的图像的一种确定方式,在该方式下,该第i帧音频对应的图像的确定过程可以包括下文步骤31-步骤33。
步骤31:对于上文至少两帧参考图像中任一参考图像,依据该参考图像、该参考图像的脸部关键点、以及第i帧音频对应的脸部关键点,预测该参考图像对应的像素使用描述信息;该像素使用描述信息包括像素调整描述信息和/或像素融合权重。
其中,参考图像对应的像素使用描述信息用于描述在第i帧音频的脸部图像生成过程中如何使用该参考图像中的像素,以使该像素使用描述信息能够表示出该参考图像中像素对该第i帧音频的脸部图像生成过程产生何种影响。
另外,本公开实施例不限定上文参考图像对应的像素使用描述信息的实施方式,比如,该参考图像对应的像素使用描述信息可以包括该参考图像对应的像素调整描述信息和/或该参考图像对应的像素融合权重。
此外,对于上文参考图像对应的像素调整描述信息来说,该像素调整描述信息用于表示在第i帧音频的脸部图像生成过程中如何调整该参考图像中的像素;而且本公开实施例不限定该像素调整描述信息的实施方式,比如,该像素调整描述信息可以采用图像像素信息的偏移量进行实施。可见,在一种可能的实施方式下,该像素调整描述信息满足以下约束:该像素调整描述信息的尺寸与该参考图像的尺寸相同,而且该像素调整描述信息中每个像素点的位置坐标用于表示该参考图像中相应的像素点的位置坐标的偏移量。
还有,对于上文参考图像对应的像素融合权重来说,该像素融合权重用于表示该参考图像中各像素以何种程度影响第i帧音频的脸部图像生成过程;而且本公开实施例不限定该像素融合权重的实施方式,比如,该像素融合权重可以满足以下约束:该像素融合权重的尺寸与该参考图像的尺寸相同,而且该像素融合权重中每个像素点的权重值用于表示该参考图像中相应的像素点的影响力度。
再者,本公开实施例不限定上文步骤31的实施方式,比如,为了更好地提高生成效果,该步骤31可以采用口型渲染模型中信息预测模块进行实施。可见,在一种可能的实施方式下,该步骤31具体可以为:对于上文至少两帧参考图像中任一参考图像,由该信息预测模块依据该参考图像、该参考图像的脸部关键点、以及第i帧音频对应的脸部关键点,预测并输出该参考图像对应的像素使用描述信息。其中,该口型渲染模型用于针对该口型渲染模型的输入数据进行口型渲染处理,如图2或者图3所示的口型渲染处理;而且本公开实施例不限定该口型渲染模型的实施方式,比如,该口型渲染模型可以至少包括该信息预测模块,如图3所示的信息预测模块。其中,该信息预测模块用于依据一个参考图像、该参考图像的脸部关键点、以及第i帧音频对应的脸部关键点,预测该参考图像对应的像素使用描述信息;而且本公开实施例不限定该信息预测模块的实施方式,比如,为了更好地提高生成效果,该信息预测模块可以包括卷积神经网络(Convolutional Neural Networks,CNN)以及特征注入模块,以使该信息预测模块具有融合该参考图像、该参考图像的脸部关键点、以及第i帧音频对应的脸部关键点的功能。需要说明的是,本公开实施例不限定该特征注入模块的实施方式,比如,该特征注入模块可以采用AdaIN(Adaptive Instance Normalization)或者SPADE进行实施。
基于上文步骤31的相关内容可知,在获取到上文至少两帧参考图像、各个参考图像的脸部关键点、以及第i帧音频对应的脸部关键点之后,可以将这些数据输入口型渲染模型,以使该口型渲染模型中信息预测模块能够依据第n个参考图像、该第n个参考图像的脸部关键点、以及该第i帧音频对应的脸部关键点,预测并输出该第n个参考图像对应的像素使用描述信息,以使该像素使用描述信息能够表示出在该第i帧音频的脸部图像生成过程中如何使用该第n个参考图像中的像素,如位置偏移量+影响力度等,n为正整数,n≤N,N为正整数,N表示该至少两帧参考图像中的图像个数,以便后续能够基于这些参考图像对应的像素使用描述信息,对该第i帧音频进行图像生成处理。
步骤32:依据上文至少两帧参考图像对应的像素使用描述信息,对该至少两帧参考图像进行形变融合处理,得到融合图像。
其中,融合图像是指依据上文至少两帧参考图像对应的像素使用描述信息,对该至少两帧参考图像进行形变融合处理所得的图像,以使该融合图像与第i帧音频相适配。
另外,本公开实施例不限定上文步骤32的实施方式,比如,为了更好地提高生成效果,该步骤32具体可以为:在口型渲染模型中信息预测模块输出各参考图像对应的像素使用描述信息之后,由该口型渲染模型中形变融合模块依据所有参考图像对应的像素使用描述信息,对所有参考图像进行形变融合处理,得到并输出融合图像,以使该融合图像能够表示出在第i帧音频下针对这些参考图像的像素整合使用结果,从而使得该融合图像与该第i帧音频相适配,进而使得该融合图像能够满足以下约束:该融合图像中所呈现的脸部表情状态满足该第i帧音频的脸部表情状态需求,而且该融合图像中除了脸部表情状态以外的其他部分或者全部信息与上文目标图像中相应的信息保持一致。其中,该形变融合模块用于依据各个参考图像对应的像素使用描述信息,对这些参考图像进行像素整合使用处理;而且本公开实施例不限定该形变融合模块的实施方式,比如,该形变融合模块可以采用任意一种信息整合网络,如grid_sample进行实施。
基于上文步骤32的相关内容可知,对于上文口型渲染模型来说,在由该口型渲染模型中信息预测模块依据第n个参考图像、该第n个参考图像的脸部关键点、以及第i帧音频对应的脸部关键点,预测并输出该第n个参考图像对应的像素使用描述信息,n为正整数,n≤N之后,可以由该口型渲染模型中形变融合模块,如图3所示的形变融合模块依据所有参考图像对应的像素使用描述信息,对所有参考图像进行形变融合处理,得到并输出融合图像,以使该融合图像能够表示出针对该第i帧音频的脸部图像生成结果,以便后续能够基于该脸部图像生成结果,确定该第i帧音频对应的图像。
步骤33:依据上文融合图像、目标图像的口型掩码结果、以及第i帧音频对应的脸部关键点进行图像生成处理,得到该第i帧音频对应的图像。
需要说明的是,本公开实施例不限定上文步骤33的实施方式,比如,为了更好地提高生成效果,该步骤33具体可以为:由口型渲染模型中脸部生成模块依据上文融合图像、目标图像的口型掩码结果、以及第i帧音频对应的脸部关键点进行图像生成处理,得到并输出该第i帧音频对应的图像,以使该第i帧音频对应的图像的图像质量优于该融合图像,如此有利于提高生成效果。其中,该脸部生成模块用于基于该目标图像的口型掩码结果以及该第i帧音频对应的脸部关键点,对该融合图像进行优化生成处理,以使该脸部生成模块输出的图像更优于该融合图像,从而使得该脸部生成模块输出的图像与该第i帧音频更适配;而且本公开实施例不限定该脸部生成模块的实施方式,比如,该脸部生成模块可以采用现有的或者未来出现的任意一种图像生成网络进行实施。
基于上文步骤31至步骤33的相关内容可知,对于一些应用场景来说,在获取到上文至少两帧参考图像、各个参考图像的脸部关键点、第i帧音频对应的目标图像的口型掩码结果、以及该第i帧音频对应的脸部关键点之后,可以将这些数据输入口型渲染模型,以使该口型渲染模型能够依据这些数据生成并输出该第i帧音频对应的图像,如图3所示的第i帧音频对应的图像。其中,因该口型渲染模型具有较好的性能,以使借助该口型渲染模型所生成的第i帧音频对应的图像更好,如此有利于提高生成效果。
基于上文S2的相关内容可知,在一些应用场景下,对于音频序列中第i帧音频来说,在获取到该第i帧音频对应的脸部关键点、该第i帧音频在参考视频中对应的目标图像的口型掩码结果、该第i帧音频对应的至少两帧参考图像以及各参考图像的脸部关键点之后,可以由口型渲染模型依据这些数据,生成并输出该第i帧音频对应的图像,以使该第i帧音频对应的图像能够表示出该目标图像中对象在该第i帧音频下所处的脸部状态,如此能够实现基于该第i帧音频对该目标图像进行表情调整处理。
S3:依据音频序列中各帧音频对应的图像,生成该音频序列对应的视频。
需要说明的是,本公开实施例不限定上文S3的实施方式,比如,为了更好地提高生成效果,该S3具体可以为:在获取到音频序列中各帧音频对应的图像之后,可以依据该音频序列以及该音频序列中各帧音频对应的图像,生成该音频序列对应的视频,以使该音频序列对应的视频包括该音频序列以及该音频序列中各帧音频对应的图像,从而使得该音频序列对应的视频中对象与上文参考视频中对象保持一致,并使得该音频序列对应的视频所呈现的脸部表情状态与该音频序列所需求的脸部表情状态保持一致,进而使得该音频序列对应的视频能够表示出该参考视频中对象在该音频序列下所处的脸部状态变化,如此能够实现基于音频序列驱动一个视频中对象的脸部表情状态调整。
基于上文S1至S3的相关内容可知,对于本公开实施例提供的视频生成方法来说,在获取到参考视频和音频序列之后,先依据该音频序列中第i帧音频对应的脸部关键点、该第i帧音频在该参考视频中对应的目标图像的口型掩码结果、从该参考视频中选择的至少两帧参考图像、以及各该参考图像的脸部关键点,生成该第i帧音频对应的图像,以使该第i帧音频对应的图像中所呈现的脸部表情状态满足该第i帧音频的表情需求,如口型需求,并使得该第i帧音频对应的图像中所呈现的除了脸部表情状态以外的其他信息,如脸部特点、脸部姿态等信息与该目标图像中所呈现的相应的信息保持一致;i为正整数,i≤该音频序列中的总帧数;然后,依据该音频序列中各帧音频对应的图像,生成该音频序列对应的视频,以使该音频序列对应的视频中所呈现的对象与该参考视频中所呈现的对象保持一致,并使得该音频序列对应的视频能够表示出该对象在该音频序列下的脸部状态变化,如此能够实现生成与该音频序列适配的视频。其中,因该至少两帧参考图像能够尽可能全面地呈现出该参考视频中出现的、在该第i帧音频的脸部图像生成过程中具有使用价值的信息,如口型、脸部姿态等方面的信息,以使基于这些参考图像针对该第i帧音频所生成的图像能够更好地表示出该对象在该第i帧音频下所处的脸部状态,从而使得基于这些图像最终生成的视频能够更好地表示出该对象在该音频序列下的脸部状态变化,进而使得最终生成的视频与该音频序列更适配,如此有利于提高视频生成效果。
另外,本公开实施例不限定本公开实施例提供的视频生成方法的执行主体,例如,本公开实施例提供的视频生成方法可以应用于终端设备或者服务器。又如,本公开实施例提供的视频生成方法也可以借助终端设备与服务器之间的数据交互过程进行实现。其中,该终端设备可以为智能手机、计算机、个人数字助理(Personal Digital Assitant,PDA)、平板电脑等。服务器可以为独立服务器、集群服务器或云服务器。
此外,本公开实施例不限定本公开实施例提供的视频生成方法的应用场景,比如,该视频生成方法可以用于完成某个视频生成任务,如视频的语种切换处理任务、视频中语句修改处理任务、或者视频中语句替换处理任务;而且在利用该视频生成方法完成该视频生成任务时所采用的数据处理逻辑类似于上文S1-S3所示的数据处理逻辑,为了简要起见,在此不再赘述。
又如,本公开实施例提供的视频生成方法可以应用于模型训练场景。基于此,本公开实施例还提供了一种模型训练过程,其具体可以至少包括下文步骤41-步骤43。
步骤41:获取参考视频和音频序列,该参考视频以及该音频序列均是依据样本视频所确定的。
需要说明的是,步骤41的相关内容请参见上文S1的相关内容。
可见,在一些应用场景下,对于当前轮来说,从一个样本视频中确定参考视频以及音频序列,以使该参考视频包括该样本视频中部分或者全部图像,并使得该音频序列包括该样本视频中部分或者全部音频,还使得该参考视频中的目标图像与该音频序列中第i帧音频之间存在对应关系,如该目标图像在该样本视频中对应的时间与该第i帧音频在该样本视频中对应的时间相同等对应关系,以便后续能够借助该参考视频和该音频序列,完成当前轮的训练过程。
步骤42:由口型渲染模型依据音频序列中第i帧音频对应的脸部关键点、该第i帧音频在参考视频中对应的目标图像的口型掩码结果、从参考视频中选择的至少两帧参考图像、以及各参考图像的脸部关键点,生成该第i帧音频对应的图像;i为正整数,i≤音频序列中的总帧数。
其中,口型渲染模型用于针对该口型渲染模型的输入数据进行口型渲染处理,如图2或者图3所示的口型渲染处理。
另外,本公开实施例不限定该口型渲染模型的实施方式,比如,为了更好地提高生成效果,该口型渲染模型可以包括信息预测模块、形变融合模块和脸部生成模块,以使该口型渲染模型的工作原理可以为:在将音频序列中第i帧音频对应的脸部关键点、该第i帧音频在参考视频中对应的目标图像的口型掩码结果、该第i帧音频对应的至少两帧参考图像、以及各参考图像的脸部关键点输入该口型渲染模型之后,先由该口型渲染模型中信息预测模块依据这些参考图像、这些参考图像的脸部关键点、以及该第i帧音频对应的脸部关键点,预测并输出这些参考图像对应的像素使用描述信息;再由该口型渲染模型中形变融合模块依据这些像素使用描述信息对这些参考图像进行形变融合处理,得到并输出融合图像;然后,由该口型渲染模型中脸部生成模块依据该融合图像、第i帧音频对应的脸部关键点以及该目标图像的口型掩码结果,生成该第i帧音频对应的图像,以便后续能够基于该第i帧音频对应的图像与该第i帧音频对应的真值(Ground Truth,GT)之间的差异性,衡量该口型渲染模型的性能。其中,该第i帧音频对应的真值用于指导该第i帧音频的脸部图像生成处理;而且本公开实施例不限定该该第i帧音频对应的真值的实施方式,比如,该该第i帧音频对应的真值可以采用该目标图像进行实施。
步骤43:依据目标图像与第i帧音频对应的图像之间的差异表征数据,更新口型渲染模型中脸部生成模块,并依据该目标图像与第i帧音频对应的图像之间的差异表征数据、以及该目标图像与融合图像之间的差异表征数据,更新该口型渲染模型中形变融合模块和信息预测模块,i为正整数,i≤音频序列中的总帧数,并返回继续执行上文步骤41及其后续步骤,直至达到预设停止条件。
其中,目标图像与第i帧音频对应的图像之间的差异表征数据用于表征该第i帧音频对应的真值与该第i帧音频对应的图像之间的差异性;而且本公开实施例不限定该差异表征数据的实施方式,比如,其可以采用现有的或者未来出现的任意一种能够衡量两个图像之间的差异性的方法进行实施。又如,为了更好地提高模型训练效果,该差异表征数据可以是依据该目标图像的图像特征与该第i帧音频对应的图像的图像特征之间的相似度,如欧式距离、余弦距离等所确定的。
另外,对于上文目标图像与第i帧音频对应的图像之间的差异表征数据来说,该差异表征数据可以借助梯度回传方式参与口型渲染模型中所有模块的更新过程,如信息预测模块的更新过程、形变融合模块的更新过程和脸部生成模块的更新过程。
此外,目标图像与融合图像之间的差异表征数据用于表征该第i帧音频对应的真值与融合图像之间的差异性;而且本公开实施例不限定该差异表征数据的实施方式,比如,其可以采用现有的或者未来出现的任意一种能够衡量两个图像之间的差异性的方法进行实施。又如,为了更好地提高模型训练效果,该差异表征数据可以是依据该目标图像的图像特征与该融合图像的图像特征之间的相似度,如欧式距离、余弦距离等所确定的。
还有,对于上文目标图像与融合图像之间的差异表征数据来说,该差异表征数据可以借助梯度回传方式参与口型渲染模型中除了脸部生成模块以外的其他模块的更新过程,如信息预测模块的更新过程和形变融合模块的更新过程。
再者,对于上文预设停止条件来说,该预设停止条件是指停止训练口型渲染模型时所需达到的条件,而且本公开实施例不限定该预设停止条件的实施方式,比如,该预设停止条件具体可以包括:该口型渲染模型的模型损失低于预先设定的损失阈值。又如,该预设停止条件可以包括:该口型渲染模型的模型损失的变化率低于预先设定的变化率阈值。还如,该预设停止条件可以包括:该口型渲染模型的更新次数达到预先设定的次数阈值。其中,该口型渲染模型的模型损失用于表征该口型渲染模型的性能;而且该口型渲染模型的损失可以是依据目标图像与第i帧音频对应的图像之间的差异表征数据、以及该目标图像与融合图像之间的差异表征数据所确定的。需要说明的是,本公开实施例不限定该口型渲染模型的损失的计算方式。
基于上文步骤41至步骤43的相关内容可知,在一些应用场景下,对于口型渲染模型来说,该口型渲染模型中的信息预测模块、形变融合模块和脸部生成模块可以同时训练,而且通过GT来监督该脸部生成模块的输出结果,同时也会借助该GT对该形变融合模块的输出结果进行弱监督,如感知损失等,如此有利于提高模型训练效果。
基于本公开实施例提供的视频生成方法,本公开实施例还提供了一种视频生成装置,下面结合图4进行解释和说明。其中,图4为本公开实施例提供的一种视频生成装置的结构示意图。需要说明的是,本公开实施例提供的视频生成装置的技术详情,请参照上文视频生成方法的相关内容。
如图4所示,本公开实施例提供的视频生成装置400,包括:
数据获取单元401,用于获取参考视频和音频序列;
图像生成单元402,用于依据所述音频序列中第i帧音频对应的脸部关键点、所述第i帧音频在所述参考视频中对应的目标图像的口型掩码结果、从所述参考视频中选择的至少两帧参考图像、以及各所述参考图像的脸部关键点,生成所述第i帧音频对应的图像;i为正整数,i≤所述音频序列中的总帧数;
视频生成单元403,用于依据所述音频序列中各帧音频对应的图像,生成所述音频序列对应的视频。
在一种可能的实施方式下,所述至少两帧参考图像的确定过程,包括:依据所述参考视频中各帧图像的口型幅度表征数据,对所述参考视频中图像进行排序,得到图像序列;对所述图像序列进行等间隔采样,得到采样图像;依据所述采样图像,确定所述至少两帧参考图像。
在一种可能的实施方式下,所述至少两帧参考图像是依据所述采样图像、以及所述第i帧音频在所述参考视频中对应的至少一个姿态相似图像所确定的;各所述姿态相似图像的姿态表征数据与所述第i帧音频对应的姿态表征数据之间的相似程度达到预设相似需求。
在一种可能的实施方式下,各所述姿态相似图像在所述参考视频中对应的时间不同于所述第i帧音频对应的时间;所述第i帧音频对应的时间是依据所述目标图像在所述参考视频中对应的时间所确定的。
在一种可能的实施方式下,所述第i帧音频对应的姿态表征数据是依据所述目标图像的姿态表征数据所确定的。
在一种可能的实施方式下,所述第i帧音频对应的脸部关键点包括所述音频序列中至少一帧音频的脸部关键点确定结果;所述至少一帧音频包括所述第i帧音频。
在一种可能的实施方式下,对于所述至少一帧音频中任一音频,该音频的脸部关键点确定结果是通过将该音频对应的三维脸部关键点投影至二维平面所得到的二维脸部关键点;该音频对应的三维脸部关键点是通过依据所述参考视频中部分或者全部图像,对该音频进行三维脸部关键点确定处理所得到的。
在一种可能的实施方式下,对于所述至少两帧参考图像中任一参考图像,该参考图像的脸部关键点是通过将该参考图像的三维脸部关键点投影至二维平面所得到的二维脸部关键点;该参考图像的三维脸部关键点是通过对该参考图像进行三维脸部关键点确定处理所得到的。
在一种可能的实施方式下,所述图像生成单元402,具体用于:对于所述至少两帧参考图像中任一参考图像,依据该参考图像、该参考图像的脸部关键点、以及所述第i帧音频对应的脸部关键点,预测该参考图像对应的像素使用描述信息;所述像素使用描述信息包括像素调整描述信息和/或像素融合权重;依据所述至少两帧参考图像对应的像素使用描述信息,对所述至少两帧参考图像进行形变融合处理,得到融合图像;依据所述融合图像、所述目标图像的口型掩码结果、以及所述第i帧音频对应的脸部关键点进行图像生成处理,得到所述第i帧音频对应的图像。
在一种可能的实施方式下,各所述参考图像对应的像素使用描述信息均是利用口型渲染模型中信息预测模块所确定的;所述形变融合处理是利用所述口型渲染模型中形变融合模块所实现的;所述图像生成处理是利用所述口型渲染模型中脸部生成模块所实现的。
在一种可能的实施方式下,所述参考视频以及所述音频序列均是依据样本视频所确定的;
所述视频生成装置400,还包括:
模型更新单元,用于依据所述目标图像与所述第i帧音频对应的图像之间的差异表征数据,更新所述口型渲染模型中脸部生成模块,并依据所述目标图像与所述第i帧音频对应的图像之间的差异表征数据、以及所述目标图像与所述融合图像之间的差异表征数据,更新所述口型渲染模型中形变融合模块和信息预测模块。
基于上述视频生成装置400的相关内容可知,本公开实施例提供的视频生成装置400的工作原理为:在获取到参考视频和音频序列之后,先依据该音频序列中第i帧音频对应的脸部关键点、该第i帧音频在该参考视频中对应的目标图像的口型掩码结果、从该参考视频中选择的至少两帧参考图像、以及各该参考图像的脸部关键点,生成该第i帧音频对应的图像,以使该第i帧音频对应的图像中所呈现的脸部表情状态满足该第i帧音频的表情需求,如口型需求,并使得该第i帧音频对应的图像中所呈现的除了脸部表情状态以外的其他信息,如脸部特点、脸部姿态等信息与该目标图像中所呈现的相应的信息保持一致;i为正整数,i≤该音频序列中的总帧数;然后,依据该音频序列中各帧音频对应的图像,生成该音频序列对应的视频,以使该音频序列对应的视频中所呈现的对象与该参考视频中所呈现的对象保持一致,并使得该音频序列对应的视频能够表示出该对象在该音频序列下的脸部状态变化,如此能够实现生成与该音频序列适配的视频。其中,因该至少两帧参考图像能够尽可能全面地呈现出该参考视频中出现的、在该第i帧音频的脸部图像生成过程中具有使用价值的信息,如口型、脸部姿态等方面的信息,以使基于这些参考图像针对该第i帧音频所生成的图像能够更好地表示出该对象在该第i帧音频下所处的脸部状态,从而使得基于这些图像最终生成的视频能够更好地表示出该对象在该音频序列下的脸部状态变化,进而使得最终生成的视频与该音频序列更适配,如此有利于提高视频生成效果。
另外,本公开实施例还提供了一种电子设备,所述设备包括处理器以及存储器:所述存储器,用于存储指令或计算机程序;所述处理器,用于执行所述存储器中的所述指令或计算机程序,以使得所述电子设备执行本公开实施例提供的视频生成方法的任一实施方式。
参见图5,其示出了适于用来实现本公开实施例的电子设备500的结构示意图。本公开实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、PDA(个人数字助理)、PAD(平板电脑)、PMP(便携式多媒体播放器)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图5示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图5所示,电子设备500可以包括处理装置(例如中央处理器、图形处理器等)501,其可以根据存储在只读存储器(ROM)502中的程序或者从存储装置508加载到随机访问存储器(RAM)503中的程序而执行各种适当的动作和处理。在RAM503中,还存储有电子设备500操作所需的各种程序和数据。处理装置501、ROM 502以及RAM 503通过总线504彼此相连。输入/输出(I/O)接口505也连接至总线504。
通常,以下装置可以连接至I/O接口505:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置506;包括例如液晶显示器(LCD)、扬声器、振动器等的输出装置507;包括例如磁带、硬盘等的存储装置508;以及通信装置509。通信装置509可以允许电子设备500与其他设备进行无线或有线通信以交换数据。虽然图5示出了具有各种装置的电子设备500,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置509从网络上被下载和安装,或者从存储装置508被安装,或者从ROM502被安装。在该计算机程序被处理装置501执行时,执行本公开实施例的方法中限定的上述功能。
本公开实施例提供的电子设备与上述实施例提供的方法属于同一发明构思,未在本实施例中详尽描述的技术细节可参见上述实施例,并且本实施例与上述实施例具有相同的有益效果。
本公开实施例还提供了一种计算机可读介质,所述计算机可读介质中存储有指令或计算机程序,当所述指令或计算机程序在设备上运行时,使得所述设备执行本公开实施例提供的视频生成方法的任一实施方式。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
在一些实施方式中,客户端、服务器可以利用诸如HTTP(Hyper Text Transfer Protocol,超文本传输协议)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(“LAN”),广域网(“WAN”),网际网(例如,互联网)以及端对端网络(例如,ad hoc端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备可以执行上述方法。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括但不限于面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(LAN)或广域网(WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元/模块的名称在某种情况下并不构成对该单元本身的限定。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
需要说明的是,本说明书中各个实施例采用递进的方式描述,每个实施例重点说明的都是与其他实施例的不同之处,各个实施例之间相同相似部分互相参见即可。对于实施例公开的系统或装置而言,由于其与实施例公开的方法相对应,所以描述的比较简单,相关之处参见方法部分说明即可。
应当理解,在本公开实施例中,“至少一个(项)”是指一个或者多个,“多个”是指两个或两个以上。“和/或”,用于描述关联对象的关联关系,表示可以存在三种关系,例如,“A和/或B”可以表示:只存在A,只存在B以及同时存在A和B三种情况,其中A,B可以是单数或者复数。字符“/”一般表示前后关联对象是一种“或”的关系。“以下至少一项(个)”或其类似表达,是指这些项中的任意组合,包括单项(个)或复数项(个)的任意组合。例如,a,b或c中的至少一项(个),可以表示:a,b,c,“a和b”,“a和c”,“b和c”,或“a和b和c”,其中a,b,c可以是单个,也可以是多个。
还需要说明的是,在本文中,诸如第一和第二等之类的关系术语仅仅用来将一个实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。而且,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、物品或者设备中还存在另外的相同要素。
结合本文中所公开的实施例描述的方法或算法的步骤可以直接用硬件、处理器执行的软件模块,或者二者的结合来实施。软件模块可以置于随机存储器(RAM)、内存、只读存储器(ROM)、电可编程ROM、电可擦除可编程ROM、寄存器、硬盘、可移动磁盘、CD-ROM、或技术领域内所公知的任意其它形式的存储介质中。
对所公开的实施例的上述说明,使本领域专业技术人员能够实现或使用本公开实施例。对这些实施例的多种修改对本领域的专业技术人员来说将是显而易见的,本文中所定义的一般原理可以在不脱离本公开实施例的精神或范围的情况下,在其它实施例中实现。因此,本公开实施例将不会被限制于本文所示的这些实施例,而是要符合与本文所公开的原理和新颖特点相一致的最宽的范围。

Claims (15)

  1. 一种视频生成方法,其中所述方法包括:
    获取参考视频和音频序列;
    依据所述音频序列中第i帧音频对应的脸部关键点、所述第i帧音频在所述参考视频中对应的目标图像的口型掩码结果、从所述参考视频中选择的至少两帧参考图像、以及各所述参考图像的脸部关键点,生成所述第i帧音频对应的图像;i为正整数,i≤所述音频序列中的总帧数;
    依据所述音频序列中各帧音频对应的图像,生成所述音频序列对应的视频。
  2. 根据权利要求1所述的方法,其中所述至少两帧参考图像的确定过程,包括:
    依据所述参考视频中各帧图像的口型幅度表征数据,对所述参考视频中图像进行排序,得到图像序列;
    对所述图像序列进行等间隔采样,得到采样图像;
    依据所述采样图像,确定所述至少两帧参考图像。
  3. 根据权利要求2所述的方法,其中所述至少两帧参考图像是依据所述采样图像、以及所述第i帧音频在所述参考视频中对应的至少一个姿态相似图像所确定的;
    各所述姿态相似图像的姿态表征数据与所述第i帧音频对应的姿态表征数据之间的相似程度达到预设相似需求。
  4. 根据权利要求3所述的方法,其中各所述姿态相似图像在所述参考视频中对应的时间不同于所述第i帧音频对应的时间;
    所述第i帧音频对应的时间是依据所述目标图像在所述参考视频中对应的时间所确定的。
  5. 根据权利要求3所述的方法,其中所述第i帧音频对应的姿态表征数据是依据所述目标图像的姿态表征数据所确定的。
  6. 根据权利要求1所述的方法,其中所述第i帧音频对应的脸部关键点包括所述音频序列中至少一帧音频的脸部关键点确定结果;
    所述至少一帧音频包括所述第i帧音频。
  7. 根据权利要求6所述的方法,其中对于所述至少一帧音频中任一音频,该音频的脸部关键点确定结果是通过将该音频对应的三维脸部关键点投影至二维平面所得到的二维脸部关键点;
    该音频对应的三维脸部关键点是通过依据所述参考视频中部分或者全部图像,对该音频进行三维脸部关键点确定处理所得到的。
  8. 根据权利要求7所述的方法,其中对于所述至少两帧参考图像中任一参考图像,该参考图像的脸部关键点是通过将该参考图像的三维脸部关键点投影至二维平面所得到的二维脸部关键点;
    该参考图像的三维脸部关键点是通过对该参考图像进行三维脸部关键点确定处理所得到的。
  9. 根据权利要求1所述的方法,其中所述第i帧音频对应的图像的确定过程,包括:
    对于所述至少两帧参考图像中任一参考图像,依据该参考图像、该参考图像的脸部关键点、以及所述第i帧音频对应的脸部关键点,预测该参考图像对应的像素使用描述信息;所述像素使用描述信息包括像素调整描述信息和/或像素融合权重;
    依据所述至少两帧参考图像对应的像素使用描述信息,对所述至少两帧参考图像进行形变融合处理,得到融合图像;
    依据所述融合图像、所述目标图像的口型掩码结果、以及所述第i帧音频对应的脸部关键点进行图像生成处理,得到所述第i帧音频对应的图像。
  10. 根据权利要求9所述的方法,其中各所述参考图像对应的像素使用描述信息均是利用口型渲染模型中信息预测模块所确定的;
    所述形变融合处理是利用所述口型渲染模型中形变融合模块所实现的;
    所述图像生成处理是利用所述口型渲染模型中脸部生成模块所实现的。
  11. 根据权利要求10所述的方法,其中所述参考视频以及所述音频序列均是依据样本视频所确定的;
    所述生成所述第i帧音频对应的图像之后,所述方法还包括:
    依据所述目标图像与所述第i帧音频对应的图像之间的差异表征数据,更新所述口型渲染模型中脸部生成模块,并依据所述目标图像与所述第i帧音频对应的图像之间的差异表征数据、以及所述目标图像与所述融合图像之间的差异表征数据,更新所述口型渲染模型中形变融合模块和信息预测模块。
  12. 一种视频生成装置,其中所述装置包括:
    数据获取单元,用于获取参考视频和音频序列;
    图像生成单元,用于依据所述音频序列中第i帧音频对应的脸部关键点、所述第i帧音频在所述参考视频中对应的目标图像的口型掩码结果、从所述参考视频中选择的至少两帧参考图像、以及各所述参考图像的脸部关键点,生成所述第i帧音频对应的图像;i为正整数,i≤所述音频序列中的总帧数;
    视频生成单元,用于依据所述音频序列中各帧音频对应的图像,生成所述音频序列对应的视频。
  13. 一种电子设备,其中所述设备包括:处理器和存储器;
    所述存储器,用于存储指令或计算机程序;
    所述处理器,用于执行所述存储器中的所述指令或计算机程序,以使得所述电子设备执行权利要求1-11任一项所述的方法。
  14. 一种计算机可读介质,其中所述计算机可读介质中存储有指令或计算机程序,当所述指令或计算机程序在设备上运行时,使得所述设备执行权利要求1-11任一项所述的方法。
  15. 一种计算机程序产品,其中所述程序产品包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行权利要求1-11任一项所述的方法的程序代码。
PCT/CN2024/140100 2024-04-08 2024-12-17 一种视频生成方法、装置、设备、介质、产品 Pending WO2025213845A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410418012.XA CN120786094A (zh) 2024-04-08 2024-04-08 一种视频生成方法、装置、设备、介质、产品
CN202410418012.X 2024-04-08

Publications (1)

Publication Number Publication Date
WO2025213845A1 true WO2025213845A1 (zh) 2025-10-16

Family

ID=97295144

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/140100 Pending WO2025213845A1 (zh) 2024-04-08 2024-12-17 一种视频生成方法、装置、设备、介质、产品

Country Status (2)

Country Link
CN (1) CN120786094A (zh)
WO (1) WO2025213845A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200234690A1 (en) * 2019-01-18 2020-07-23 Snap Inc. Text and audio-based real-time face reenactment
CN112927712A (zh) * 2021-01-25 2021-06-08 网易(杭州)网络有限公司 视频生成方法、装置和电子设备
CN116385604A (zh) * 2023-06-02 2023-07-04 摩尔线程智能科技(北京)有限责任公司 视频生成及模型训练方法、装置、设备、存储介质
CN116844215A (zh) * 2023-07-26 2023-10-03 上海墨百意信息科技有限公司 说话人脸的生成方法、装置、电子设备及存储介质
CN117640994A (zh) * 2023-10-18 2024-03-01 厦门黑镜科技有限公司 一种视频生成方法及相关设备

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200234690A1 (en) * 2019-01-18 2020-07-23 Snap Inc. Text and audio-based real-time face reenactment
CN112927712A (zh) * 2021-01-25 2021-06-08 网易(杭州)网络有限公司 视频生成方法、装置和电子设备
CN116385604A (zh) * 2023-06-02 2023-07-04 摩尔线程智能科技(北京)有限责任公司 视频生成及模型训练方法、装置、设备、存储介质
CN116844215A (zh) * 2023-07-26 2023-10-03 上海墨百意信息科技有限公司 说话人脸的生成方法、装置、电子设备及存储介质
CN117640994A (zh) * 2023-10-18 2024-03-01 厦门黑镜科技有限公司 一种视频生成方法及相关设备

Also Published As

Publication number Publication date
CN120786094A (zh) 2025-10-14

Similar Documents

Publication Publication Date Title
CN109816589B (zh) 用于生成漫画风格转换模型的方法和装置
CN113470626B (zh) 一种语音识别模型的训练方法、装置及设备
CN118097157B (zh) 基于模糊聚类算法的图像分割方法及系统
CN114494709B (zh) 特征提取模型的生成方法、图像特征提取方法和装置
CN110097004B (zh) 面部表情识别方法和装置
CN113222050B (zh) 图像分类方法、装置、可读介质及电子设备
CN112800276A (zh) 视频封面确定方法、装置、介质及设备
WO2023116138A1 (zh) 多任务模型的建模方法、推广内容处理方法及相关装置
CN112861935B (zh) 模型生成方法、对象分类方法、装置、电子设备及介质
CN115937020A (zh) 图像处理方法、装置、设备、介质和程序产品
CN113140012B (zh) 图像处理方法、装置、介质及电子设备
CN110633717A (zh) 一种目标检测模型的训练方法和装置
CN112418233B (zh) 图像处理方法、装置、可读介质及电子设备
CN116524532B (zh) 动作识别方法、装置、存储介质以及电子设备
CN115146657A (zh) 模型训练方法、装置、存储介质、客户端、服务器和系统
CN114201674A (zh) 模型训练方法、推广内容的处理方法及相关装置
CN114495227A (zh) 年龄预测网络生成、年龄预测方法、装置、设备和介质
WO2025200625A1 (zh) 一种视频生成方法、装置、设备、介质、产品
WO2025167333A1 (zh) 一种图像生成方法、装置、设备、介质、产品
CN113051400A (zh) 标注数据确定方法、装置、可读介质及电子设备
WO2025213845A1 (zh) 一种视频生成方法、装置、设备、介质、产品
WO2024169893A1 (zh) 模型构建方法、虚拟形象生成方法、装置、设备、介质
CN114595346A (zh) 内容检测模型的训练方法、内容检测方法及装置
WO2025213838A1 (zh) 一种数据处理方法、装置、设备、介质、产品
CN119557417B (zh) 图文匹配模型的训练方法、装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24934908

Country of ref document: EP

Kind code of ref document: A1