WO2025245877A1 - 一种生成视频内容的方法、装置、设备和存储介质 - Google Patents

一种生成视频内容的方法、装置、设备和存储介质

Info

Publication number
WO2025245877A1
WO2025245877A1 PCT/CN2024/096829 CN2024096829W WO2025245877A1 WO 2025245877 A1 WO2025245877 A1 WO 2025245877A1 CN 2024096829 W CN2024096829 W CN 2024096829W WO 2025245877 A1 WO2025245877 A1 WO 2025245877A1
Authority
WO
WIPO (PCT)
Prior art keywords
images
point cloud
video content
virtual camera
cloud data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/096829
Other languages
English (en)
French (fr)
Inventor
魏国强
侯晨
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Youzhuju Network Technology Co Ltd
Original Assignee
Beijing Youzhuju Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Youzhuju Network Technology Co Ltd filed Critical Beijing Youzhuju Network Technology Co Ltd
Priority to PCT/CN2024/096829 priority Critical patent/WO2025245877A1/zh
Priority to CN202480003522.9A priority patent/CN119631417A/zh
Publication of WO2025245877A1 publication Critical patent/WO2025245877A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T15/00Three-dimensional [3D] image rendering
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/85Assembly of content; Generation of multimedia applications
    • H04N21/854Content authoring

Definitions

  • the exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to a method, apparatus, device, and computer-readable storage medium for generating video content.
  • video generation technology has made significant progress, especially video synthesis techniques based on text prompts or image input. These techniques, through machine learning models (e.g., diffusion models), are able to generate videos with rich dynamic content.
  • machine learning models e.g., diffusion models
  • traditional video generation schemes lack effective control over camera movement during video generation, which limits the expressiveness and diversity of video content.
  • a method for generating video content includes: constructing three-dimensional point cloud data based on a target image and depth information associated with the target image; generating a first set of images corresponding to multiple states of the virtual camera based on the three-dimensional point cloud data according to the motion trajectory of a virtual camera; and generating video content using a target model based on the first set of images and cue items.
  • an apparatus for generating video content includes: a point cloud construction module configured to construct three-dimensional point cloud data based on a target image and depth information associated with the target image; an image generation module configured to generate a first set of images corresponding to multiple states of a virtual camera based on the three-dimensional point cloud data according to the motion trajectory of a virtual camera; and a video generation module configured to generate video content using a target model based on the first set of images and prompts.
  • an electronic device in a third aspect of this disclosure, includes at least one processing unit; and at least one memory, the at least one memory being coupled to at least one...
  • the device has a processing unit and stores instructions for execution by at least one processing unit. When executed by at least one processing unit, the instructions cause the device to perform the method of the first aspect.
  • a computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
  • Figure 1 shows a schematic diagram of an example environment in which some embodiments of the present disclosure can be implemented
  • Figure 2 shows a flowchart of a process for generating video content according to some embodiments of the present disclosure
  • Figure 3 illustrates an example architecture of a video generation system according to some embodiments of the present disclosure
  • Figure 4 shows a schematic structural block diagram of an example apparatus for generating video content according to some embodiments of the present disclosure.
  • Figure 5 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure.
  • the term “comprising” and similar terms should be understood as open-ended inclusion, i.e., “including but not limited to”.
  • the term “based on” should be understood as “at least partially based on”.
  • the term “one embodiment” or “the embodiment” should be understood as “at least one embodiment”.
  • the term “some embodiments” should be understood as “at least some embodiments”.
  • Other explicit and implicit definitions may also be included below.
  • the terms “first”, “second”, etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
  • the embodiments of this disclosure may involve user data, data acquisition, and/or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and/or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
  • any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon.
  • a user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
  • some video generation techniques can utilize machine learning models (e.g., diffusion models) to generate videos with rich dynamic content.
  • machine learning models e.g., diffusion models
  • traditional video generation schemes lack effective control over camera movement during video generation, which limits the expressiveness and diversity of the video content.
  • Embodiments of this disclosure propose a scheme for generating video content.
  • three-dimensional point cloud data can be constructed based on a target image and depth information associated with the target image. Further, based on the motion trajectory of a virtual camera, a first set of images corresponding to multiple states of the virtual camera can be generated from the three-dimensional point cloud data. Accordingly, video content can be generated using a target model based on the first set of images and prompts.
  • embodiments of the present disclosure can generate video content that matches preset camera movements, thereby improving the quality of the generated video content.
  • Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented.
  • the example environment 100 may include an electronic device 110.
  • the electronic device 110 may acquire the input target image 120 and the prompt item 130. Further, the electronic device 110 may utilize the video generation system 140 to generate video content 150 based on the target image 120 and the prompt item 130.
  • a prompt item 130 may, for example, include text content describing the video content 150 to be generated, also known as a text prompt or prompt word, etc.
  • such a video generation system 140 may be deployed locally on electronic device 110, or it may be deployed at a suitable remote device.
  • electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR/AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio/video players, digital cameras/camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof.
  • electronic device 110 may also support any type of user-facing interface (such as "wearable" circuitry).
  • FIG 2 shows a flowchart of an example process 200 for generating video content according to some embodiments of the present disclosure.
  • Process 200 may be implemented, for example, at an electronic device 110 as shown in Figure 1.
  • Process 200 will be described below with reference to Figure 1.
  • electronic device 110 constructs three-dimensional point cloud data based on the target image and the depth information associated with the target image.
  • FIG 3 shows a block diagram of an example video generation system 300 according to some embodiments of the present disclosure.
  • Such a video generation system 300 may correspond to the video generation system 140 shown in Figure 1.
  • the electronic device 110 can acquire the target image 120 and can further determine the depth information 305 associated with the target image 120.
  • depth information 305 may, for example, include a depth map corresponding to the target image 120 to indicate the depth value of each pixel.
  • the target image 120 may include, for example, a depth image, and its depth information may be determined directly based on the depth image.
  • the target image 120 may include, for example, a two-dimensional image, and the electronic device 110 may utilize a depth estimation module to determine the depth information 305.
  • the depth estimation module may utilize any suitable depth estimation model to determine the depth information 305 corresponding to the target image 120.
  • the electronic device 110 can convert multiple pixels in the target image 120 into three-dimensional space based on the target image 120 and the depth information 305 to generate three-dimensional point cloud data 315.
  • the construction process of the 3D point cloud data 315 can be represented as follows:
  • electronic device 110 generates a first set of images corresponding to multiple states of the virtual camera based on the motion trajectory of the virtual camera and 3D point cloud data.
  • the electronic device 110 can determine the motion trajectory 320 of the virtual camera, which can be represented, for example, as [K, Pi ], where K represents the intrinsic parameters of the virtual camera, and Pi represents the extrinsic parameters of the virtual camera in multiple states.
  • K represents the intrinsic parameters of the virtual camera
  • Pi represents the extrinsic parameters of the virtual camera in multiple states.
  • extrinsic parameters may include, for example, the position, rotation, translation, and attitude of the virtual camera in three-dimensional space.
  • the electronic device 110 can project the 3D point cloud data onto the corresponding 2D plane based on the external parameters of the virtual camera in multiple states to generate the corresponding first set of images.
  • This process can be represented, for example, as follows:
  • I ⁇ sub> i ⁇ /sub> represents the image corresponding to the i-th state of the virtual camera
  • represents the projection function
  • the electronic device 110 can acquire a set of images corresponding to different states of the camera, and the dynamic switching between these images can correspond to camera movement effects, such as zooming, panning, and rotation.
  • electronic device 110 generates video content based on the first set of images and prompts using a target model.
  • the electronic device 110 may, for example, directly use the first set of projected images as input to the video generation model 335 to generate a corresponding second set of images as multiple video frames of the video content 150.
  • the first set of images generated by projecting the 3D point cloud data 315 may contain empty pixels.
  • the electronic device 110 can also use an image inpainting model to process one or more images in the first set of images to generate a repaired set of images (also known as the third set of images).
  • the image inpainting model can use appropriate image inpainting techniques to fill in the empty pixels in the images.
  • the electronic device 110 may use such a third set of images as a visual image.
  • the input of the frequency generation model 335 is used to generate a second set of images.
  • the electronic device 110 can further project the repaired third set of images into the three-dimensional space corresponding to the three-dimensional point cloud data 315.
  • the electronic device 110 can update the 3D point cloud data 315 by aligning multiple positions corresponding to the third set of images in 3D space.
  • This alignment process can be represented, for example, as:
  • embodiments of this disclosure can find the optimal depth coefficients so that the point cloud representations of preceding and following images in overlapping regions are as consistent as possible. Therefore, embodiments of this disclosure can correct relative depth errors caused by monocular depth estimation, thereby maintaining the coherence and consistency of objects and scenes in the generated video.
  • the electronic device 110 can generate a fourth set of images corresponding to multiple states of the virtual camera based on the updated 3D point cloud data. This process can be represented as:
  • the fourth set of images generated by the electronic device 110 can be, for example, a set of images 325 as shown in Figure 3. Accordingly, the electronic device 110 can use the video generation model 335 to process the set of images 325 and the prompt item 130 to generate a second set of images as multiple video frames.
  • the video generation model 335 may include, for example, a diffusion model.
  • the diffusion model in the noise addition stage, can generate a set of hidden features corresponding to the fourth set of images by adding noise that satisfies a preset distribution. Further, in the denoising stage, the diffusion model can generate a corresponding second set of images based on the set of hidden features. This process can be represented as:
  • Equation (5) is used to obtain the latent noise representation from the rendered image sequence V0325 through a forward diffusion process.
  • 330. is the variance used in the DDIM scheduler (Denoising Diffusion Implicit Models); ⁇ is random noise sampled from the standard normal distribution to perturb the latent representation and increase the diversity of the generated image; t 0 is a time step in the diffusion process that determines the intensity of the noise.
  • Equation (6) describes the use of noise latent representation
  • the steps to generate video through a backdiffusion process It is the denoised latent representation at time step t-1.
  • a ⁇ sub>t ⁇ /sub> is a scheduling parameter used to control the denoising process.
  • is the random noise sampled in each step. This is the noise prediction part of the diffusion model.
  • ⁇ ⁇ sub>t ⁇ /sub> determines whether the denoising process is deterministic or probabilistic, and is usually set to 1 to encourage diversity in the generated results.
  • t is the time step in the diffusion process.
  • the electronic device 110 may also balance the realism and diversity of the video content, for example, by controlling the time step t ⁇ sub> 0 ⁇ /sub>.
  • the time step t ⁇ sub>0 ⁇ /sub> used to generate the latent representation of noise is a key factor influencing this trade-off.
  • a larger t ⁇ sub> 0 ⁇ /sub> value will result in a video that more closely resembles the original guided camera movement, but may sacrifice the dynamism and plausibility of the video content.
  • a smaller t ⁇ sub> 0 ⁇ /sub> value can produce a more plausible video, but may not perfectly match the desired camera movement.
  • embodiments of the present disclosure can control camera motion during video generation by manipulating noise in the latent space, achieving video generation that can control camera motion without additional training.
  • the electronic device 110 can generate video content 150 by sequentially combining the second set of images, such that multiple video frames of the video content 150 can correspond to multiple states [K, Pi ] of the virtual camera. It should be understood that such correspondence is intended to indicate that the video frames have similar camera movement control, rather than to indicate that the video frame completely corresponds to the external parameters of the virtual camera in the corresponding state.
  • the electronic device 110 may also add audio content, Subtitles and other appropriate elements are used to generate the final video content 150.
  • embodiments of this disclosure can utilize explicit rearrangement of image layout in 3D point cloud space and layout priors of noise latent representation to generate videos with rich dynamic content and high realism, while maintaining the efficiency of the processing and the robustness of the model.
  • the embodiments of this disclosure can support complex hybrid camera motions and can be widely applied to fields such as 3D video generation and virtual reality content creation, thereby improving the flexibility and innovation of video content creation.
  • FIG. 4 shows a schematic structural block diagram of an example apparatus 400 for generating video content according to certain embodiments of this disclosure.
  • Apparatus 400 may be implemented as or included in an electronic device.
  • the various modules/components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
  • the device 400 includes a point cloud construction module 410, configured to construct three-dimensional point cloud data based on a target image and depth information associated with the target image; an image generation module 420, configured to generate a first set of images corresponding to multiple states of the virtual camera based on the three-dimensional point cloud data according to the motion trajectory of the virtual camera; and a video generation module 430, configured to generate video content using the target model based on the first set of images and prompts.
  • a point cloud construction module 410 configured to construct three-dimensional point cloud data based on a target image and depth information associated with the target image
  • an image generation module 420 configured to generate a first set of images corresponding to multiple states of the virtual camera based on the three-dimensional point cloud data according to the motion trajectory of the virtual camera
  • a video generation module 430 configured to generate video content using the target model based on the first set of images and prompts.
  • the image generation module 420 is further configured to: determine multiple sets of external parameters of the virtual camera in multiple states based on the motion trajectory of the virtual camera; and generate a first set of images corresponding to the multiple sets of external parameters by projecting three-dimensional point cloud data.
  • the video generation module 430 is further configured to: process at least one image in the first set of images using an image inpainting model to generate a third set of images; and generate a second set of images using a target model based on the third set of images and the prompts.
  • the video generation module 430 is further configured to: project a third set of images onto a three-dimensional space corresponding to the three-dimensional point cloud data; update the three-dimensional point cloud data by aligning multiple positions corresponding to the third set of images in the three-dimensional space; generate a fourth set of images corresponding to multiple states of the virtual camera based on the updated three-dimensional point cloud data; and project the images onto the target.
  • the model provides a fourth set of images and prompts to generate a second set of images.
  • the target model is a diffusion model, which is configured to: generate a set of hidden features corresponding to the fourth set of images by adding noise that satisfies a preset distribution; and generate a corresponding second set of images based on the set of hidden features.
  • the apparatus 400 further includes a depth estimation module configured to process the target image using a depth estimation model to determine depth information.
  • the video content includes multiple video frames corresponding to multiple states of the virtual camera.
  • FIG. 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used with the electronic device 110 shown in Figure 1.
  • the electronic device 500 is in the form of a general-purpose electronic device.
  • Components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560.
  • the processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.
  • Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media.
  • Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof.
  • Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and/or data and accessible within electronic device 500.
  • Electronic device 500 may further include additional removable/non-removable, volatile/ Non-volatile storage media.
  • disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided.
  • each drive may be connected to a bus (not shown) via one or more data media interfaces.
  • Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
  • Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
  • PCs network personal computers
  • Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc.
  • Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc.
  • Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input/output (I/O) interface (not shown).
  • I/O input/output
  • a computer-readable storage medium that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above.
  • a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
  • These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and/or other device to operate in a particular manner.
  • the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions/actions specified in one or more boxes of a flowchart and/or block diagram.
  • each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function.
  • the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
  • each block in the block diagrams and/or flowcharts, and combinations of blocks in the block diagrams and/or flowcharts may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Security & Cryptography (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computer Graphics (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Processing Or Creating Images (AREA)

Abstract

本公开的实施例涉及一种生成视频内容的方法、装置、设备和存储介质。在此提出的方法包括:基于目标图像和与目标图像相关联的深度信息,构建三维点云数据;根据虚拟相机的运动轨迹,基于三维点云数据生成与虚拟相机的多个状态对应的第一组图像;基于第一组图像和提示项,利用目标模型生成视频内容。以此方式,本公开的实施例能够生成与预设运镜匹配的视频内容,从而提高所生成的视频内容的质量。

Description

一种生成视频内容的方法、装置、设备和存储介质 技术领域
本公开的示例实施例总体涉及计算机领域,特别地涉及一种生成视频内容的方法、装置、设备和计算机可读存储介质。
背景技术
近年来,视频生成技术获得了显著进展,尤其是基于文本提示或图像输入的视频合成技术。这些技术通过机器学习模型(例如,扩散模型)能够生成具有丰富动态内容的视频。然而,传统的视频生成方案在生成视频时缺乏对运镜的有效控制,这限制了视频内容的表现力和多样性。
发明内容
在本公开的第一方面,提供了一种生成视频内容的方法。该方法包括:基于目标图像和与目标图像相关联的深度信息,构建三维点云数据;根据虚拟相机的运动轨迹,基于三维点云数据生成与虚拟相机的多个状态对应的第一组图像;基于第一组图像和提示项,利用目标模型生成视频内容。
在本公开的第二方面,提供了一种用于生成视频内容的装置。该装置包括:点云构建模块,被配置为基于目标图像和与目标图像相关联的深度信息,构建三维点云数据;图像生成模块,被配置为根据虚拟相机的运动轨迹,基于三维点云数据生成与虚拟相机的多个状态对应的第一组图像;视频生成模块,被配置为基于第一组图像和提示项,利用目标模型生成视频内容。
在本公开的第三方面,提供了一种电子设备。该设备包括至少一个处理单元;以及至少一个存储器,至少一个存储器被耦合到至少一 个处理单元并且存储用于由至少一个处理单元执行的指令。指令在由至少一个处理单元执行时使设备执行第一方面的方法。
在本公开的第四方面,提供了一种计算机可读存储介质。该计算机可读存储介质上存储有计算机程序,计算机程序可由处理器执行以实现第一方面的方法。
应当理解,本内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
附图说明
结合附图并参考以下详细说明,本公开各实施例的上述和其他特征、优点及方面将变得更加明显。在附图中,相同或相似的附图标识表示相同或相似的元素,其中:
图1示出了能够实施本公开的一些实施例的示例环境的示意图;
图2示出了根据本公开的一些实施例的生成视频内容的过程的流程图;
图3示出了根据本公开的一些实施例的视频生成系统的示例架构;
图4示出了根据本公开的一些实施例的用于生成视频内容的示例装置的示意性结构框图;以及
图5示出了能够实施本公开的多个实施例的电子设备的框图。
具体实施方式
下面将参照附图更详细地描述本公开的实施例。虽然附图中示出了本公开的某些实施例,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例,相反,提供这些实施例是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实施例仅用于示例性作用,并非用于限制本公开的保护范围。
需要注意的是,本文中所提供的任何节/子节的标题并不是限制性的。本文通篇描述了各种实施例,并且任何类型的实施例都可以包括在任何节/子节下。此外,在任一节/子节中描述的实施例可以以任何方式与同一节/子节和/或不同节/子节中描述的任何其他实施例相结合。
在本公开的实施例的描述中,术语“包括”及其类似用语应当理解为开放性包含,即“包括但不限于”。术语“基于”应当理解为“至少部分地基于”。术语“一个实施例”或“该实施例”应当理解为“至少一个实施例”。术语“一些实施例”应当理解为“至少一些实施例”。下文还可能包括其他明确的和隐含的定义。术语“第一”、“第二”等可以指代不同的或相同的对象。下文还可能包括其他明确的和隐含的定义。
本公开的实施例中可能涉及用户的数据、数据的获取和/或使用等。这些方面均遵循相应的法律法规及相关规定。在本公开的实施例中,所有数据的采集、获取、处理、加工、转发、使用等,都是在用户知晓并且确认的前提下进行的。相应地,在实现本公开的各实施例时,均应根据相关法律法规通过适当的方式,将可能所涉及的数据或信息的类型、使用范围、使用场景等告知用户并获得用户的授权。具体的告知和/或授权方式可以根据实际情况和应用场景而变化,本公开的范围在此方面不受限制。
本说明书及实施例中方案,如涉及个人信息处理,则均会在具备合法性基础(例如征得个人信息主体同意,或者为履行合同所必需等)的前提下进行处理,且仅会在规定或者约定的范围内进行处理。用户拒绝处理基本功能所需必要信息以外的个人信息,不会影响用户使用基本功能。
如上文简要提及的,一些视频生成技术可以利用机器学习模型(例如,扩散模型)生成具有丰富动态内容的视频。然而,传统的视频生成方案在生成视频时缺乏对运镜的有效控制,这限制了视频内容的表现力和多样性。
本公开的实施例提出了一种用于生成视频内容的方案。根据该方案,可以基于目标图像和与目标图像相关联的深度信息,构建三维点云数据。进一步地,可以根据虚拟相机的运动轨迹,基于三维点云数据生成与虚拟相机的多个状态对应的第一组图像。相应地,可以基于第一组图像和提示项,利用目标模型生成视频内容。
以此方式,本公开的实施例能够生成与预设运镜匹配的视频内容,从而提高所生成的视频内容的质量。
以下进一步结合附图来详细描述该方案的各种示例实现。
示例环境
图1示出了本公开的实施例能够在其中实现的示例环境100的示意图。如图1所示,示例环境100可以包括电子设备110。
在一些实施例中,电子设备110可以获取输入的目标图像120和提示项130。进一步地,电子设备110可以利用视频生成系统140以基于目标图像120和提示项130来生成视频内容150。这样的提示项130例如可以包括用于描述待生成的视频内容150的文本内容,也成为文本提示项或提示词等。
在一些实施例中,这样的视频生成系统140例如可以部署在电子设备110本地,或者可以部署在适当的远程设备处。
在一些示例中,电子设备110可以是任意类型的移动终端、固定终端或便携式终端,包括移动手机、台式计算机、膝上型计算机、笔记本计算机、上网本计算机、平板计算机、媒体计算机、多媒体平板、掌上电脑、便携式游戏终端、VR/AR设备、个人通信系统(Personal Communication System,PCS)设备、个人导航设备、个人数字助理(Personal Digital Assistant,PDA)、音频/视频播放器、数码相机/摄像机、定位设备、电视接收器、无线电广播接收器、电子书设备、游戏设备或者前述各项的任意组合,包括这些设备的配件和外设或者其任意组合。在一些实施例中,电子设备110也能够支持任意类型的针对用户的接口(诸如“可佩戴”电路等)。
应当理解,仅出于示例性的目的描述环境100中各个元素的结构和功能,而不暗示对于本公开的范围的任何限制。
以下将继续参考附图描述本公开的一些示例实施例。
示例过程
图2示出了根据本公开的一些实施例的用于生成视频内容的示例过程200的流程图。过程200例如可以被实现在如图1所示的电子设备110处。以下将参考图1来描述过程200。
如图2所示,在框210,电子设备110基于目标图像和与目标图像相关联的深度信息,构建三维点云数据。
图3示出了根据本公开的一些实施例的示例视频生成系统300的框图。这样的视频生成系统300可以对应于图1中所示出的视频生成系统140。
具体地,如图3所示,电子设备110可以获取目标图像120,并可以进一步确定与目标图像120相关联的深度信息305。这样的深度信息305例如可以包括与目标图像120对应的深度图,以指示每个像素的深度值。
在一些实施例中,目标图像120例如可以包括深度图像,并且其深度信息可以基于深度图像直接地确定。在一些实施例中,目标图像120例如可以包括二维图像,电子设备110例如可以利用深度估计模块来确定深度信息305。例如,深度估计模块可以利用任何适当的深度估计模型来确定与目标图像120对应的深度信息305。
进一步地,如图3所示,电子设备110可以基于目标图像120和深度信息305来将目标图像120中的多个像素点转换到三维空间,以生成三维点云数据315。
具体地,三维点云数据315的构建过程可以表示为:
其中,表示三维点云数据315,φ表示从二维图像空间映射到三维空间的映射函数,表示目标图像120,D0表示深度信息 305,K表示目标图像120所对应的虚拟相机的内部参数,P0表示虚拟相机的外部参数。
继续参考图2,在框220,电子设备110根据虚拟相机的运动轨迹,基于三维点云数据生成与虚拟相机的多个状态对应的第一组图像。
如图3所示,电子设备110可以确定虚拟相机的运动轨迹320,其例如可以表示为[K,Pi],其中K为虚拟相机的内部参数,Pi表示虚拟相机在多个状态的外部参数。这样的外部参数例如可以包括虚拟相机在三维空间中的位置、旋转、平移和姿态等参数。
进一步地,电子设备110可以基于虚拟相机在多个状态的外部参数,来将三维点云数据投影至对应的二维平面,以生成对应的第一组图像。该过程例如可以表示为:
其中,Ii表示与虚拟相机的第i个状态对应的图像,ψ表示投影函数。
以此方式,电子设备110可以获取与相机的不同状态对应的一组图像,并且这样图像之间的动态切换可以对应于相机的运镜效果,例如,缩放、平移、旋转等。
继续参考图2,在框230,电子设备110基于第一组图像和提示项,利用目标模型生成视频内容。
在一些实施例中,如图3所示,电子设备110例如可以直接将所投影得到的第一组图像作为视频生成模型335的输入,以用于生成对应的第二组图像,以作为视频内容150的多个视频帧。
考虑到三维点云数据315在某些位置点存在空洞,由三维点云数据315投影所生成的第一组图像可能存在空像素。在一些场景中,为了提高视频内容的生成质量,电子设备110还可以利用图像修补模型处理第一组图像中的一个或多个图像,以生成修复后的一组图像(也成为第三组图像)。例如,图像修补模型可以利用适当的图像修补技术(in-painting)来填充图像中的空像素。
在一些实施例中,电子设备110可以将这样的第三组图像作为视 频生成模型335的输入以生成第二组图像。
此外,由于填充后的空像素在三维点云空间可能存在不对齐的情况下,电子设备110还可以进一步地将修补后的第三组图像投影至与三维点云数据315对应的三维空间。
进一步地,电子设备110可以通过在三维空间中对齐与第三组图像对应的多个位置,更新三维点云数据315。该对齐过程例如可以表示为:
其中,分别表示修改后的第三组图像和对应的深度信息,di表示要优化的深度参数,M表示的重叠区域,||·||表示即计算L1损失。
基于这样的方式,本公开的实施例可以寻找最佳的深度系数,使得前后图像在重叠区域的点云表示尽可能一致。由此,本公开的实施例可以修正由于单目深度估计产生的相对深度误差,从而在生成视频中保持物体和场景的连贯性和一致性。
附加地,电子设备110可以基于经更新的三维点云数据,生成与虚拟相机的多个状态对应的第四组图像。该过程可以表示为:
以图3作为示例,电子设备110所生成的第四组图像例如可以为如图3所示的一组图像325。相应地,电子设备110可以利用视频生成模型335来处理该组图像325和提示项130,以生成作为多个视频帧的第二组图像。
在一些实施例中,视频生成模型335例如可以包括扩散模型。具体地,在加噪阶段,该扩散模型可以通过添加满足预设分布的噪声,生成与第四组图像对应的一组隐藏特征。进一步地,在去噪阶段,扩散模型可以基于一组隐藏特征,生成对应的第二组图像。该过程可以表示为:

具体地,公式(5)用于通过前向扩散过程从渲染图像序列V0325中获取噪声潜在表示330。是在DDIM调度器(Denoising Diffusion Implicit Models,去噪扩散隐式模型)中使用的方差;∈是从标准正态分布中采样的随机噪声,用于扰动潜在表示,增加生成图像的多样性;t0是扩散过程中的一个时间步长,决定了噪声的强度。
公式(6)描述了使用噪声潜在表示通过反向扩散过程生成视频的步骤。是在时间步t-1的去噪后的潜在表示。at是一个调度参数,用于控制去噪过程。∈是在每一步中采样的随机噪声。是扩散模型的噪声预测部分。σt确定去噪过程是确定性的还是概率性的,通常设置为1以鼓励生成结果的多样性。t是扩散过程中的时间步长。
在一些实施例中,电子设备110例如还可以通过控制时间步长t0来平衡视频内容的真实性和多样性。在生成过程汇总,用于生成噪声潜在表示的时间步长t0是影响这一权衡的关键因素。较大的t0值会使生成的视频更贴近于原始指导的相机运动,但可能会牺牲视频内容的动态性和合理性。相反,较小的t0值可以产生更合理的视频,但可能不会完全符合期望的相机运动。
以此方式,本公开的实施例能够通过在潜在空间中操作噪声来控制视频生成过程中的相机运动,实现了无需额外训练即可控制相机运动的视频生成。
进一步地,电子设备110可以通过按序组合第二组图像以生成视频内容150,以使得该视频内容150的多个视频帧能够对应于虚拟相机的多个状态[K,Pi]。应当理解的是,这样的对应性旨在表示视频帧具有相似的运镜控制,而不旨在表示该视频帧完全对应于虚拟相机在对应状态下的外部参数。
在一些实施例中,电子设备110例如还可以通过添加音频内容、 字幕等内容等其他适当的元素来生成最终的视频内容150。
由此,本公开的实施例能够利用3D点云空间中图像布局的显式重排和噪声潜在表示的布局先验,来生成具有丰富动态内容和高度逼真度的视频,同时保持了处理过程的高效性和模型的鲁棒性。
此外,本公开的实施例能够支持复杂的混合相机运动,能够广泛应用于3D视频生成、虚拟现实内容创建等领域,从而提高视频内容创作的灵活性和创新性。
示例装置和设备
本公开的实施例还提供了用于实现上述方法或过程的相应装置。图4示出了根据本公开的某些实施例的用于生成视频内容的示例装置400的示意性结构框图。装置400可以被实现为或者被包括在电子设备中。装置400中的各个模块/组件可以由硬件、软件、固件或者它们的任意组合来实现。
如图4所示,装置400包括点云构建模块410,被配置为基于目标图像和与目标图像相关联的深度信息,构建三维点云数据;图像生成模块420,被配置为根据虚拟相机的运动轨迹,基于三维点云数据生成与虚拟相机的多个状态对应的第一组图像;视频生成模块430,被配置为基于第一组图像和提示项,利用目标模型生成视频内容。
在一些实施例中,图像生成模块420还被配置为:基于虚拟相机的运动轨迹,确定虚拟相机在多个状态下的多组外部参数;以及通过投影三维点云数据,生成与多组外部参数对应的第一组图像。
在一些实施例中,视频生成模块430还被配置为:利用图像修补模型处理第一组图像中的至少一个图像,以生成第三组图像;以及基于第三组图像和提示项,利用目标模型生成第二组图像。
在一些实施例中,视频生成模块430还被配置为:将第三组图像投影至与三维点云数据对应的三维空间;通过在三维空间中对齐与第三组图像对应的多个位置,更新三维点云数据;基于经更新的三维点云数据,生成与虚拟相机的多个状态对应的第四组图像;以及向目标 模型提供第四组图像和提示项,以生成第二组图像。
在一些实施例中,目标模型为扩散模型,扩散模型被配置为:通过添加满足预设分布的噪声,生成与第四组图像对应的一组隐藏特征;以及基于一组隐藏特征,生成对应的第二组图像。
在一些实施例中,装置400还包括深度估计模块,被配置为:利用深度估计模型处理目标图像,以确定深度信息。
在一些实施例中,视频内容包括与虚拟相机的多个状态对应的多个视频帧。
图5示出了其中可以实施本公开的一个或多个实施例的电子设备500的框图。应当理解,图5所示出的电子设备500仅仅是示例性的,而不应当构成对本文所描述的实施例的功能和范围的任何限制。图5所示出的电子设备500可以用于如图1所示的电子设备110。
如图5所示,电子设备500是通用电子设备的形式。电子设备500的组件可以包括但不限于一个或多个处理器或处理单元510、存储器520、存储设备530、一个或多个通信单元540、一个或多个输入设备550以及一个或多个输出设备560。处理单元510可以是实际或虚拟处理器并且能够根据存储器520中存储的程序来执行各种处理。在多处理器系统中,多个处理单元并行执行计算机可执行指令,以提高电子设备500的并行处理能力。
电子设备500通常包括多个计算机存储介质。这样的介质可以是电子设备500可访问的任何可以获取的介质,包括但不限于易失性和非易失性介质、可拆卸和不可拆卸介质。存储器520可以是易失性存储器(例如寄存器、高速缓存、随机访问存储器(RAM))、非易失性存储器(例如,只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、闪存)或它们的某种组合。存储设备530可以是可拆卸或不可拆卸的介质,并且可以包括机器可读介质,诸如闪存驱动、磁盘或者任何其他介质,其能够用于存储信息和/或数据并且可以在电子设备500内被访问。
电子设备500可以进一步包括另外的可拆卸/不可拆卸、易失性/ 非易失性存储介质。尽管未在图5中示出,可以提供用于从可拆卸、非易失性磁盘(例如“软盘”)进行读取或写入的磁盘驱动和用于从可拆卸、非易失性光盘进行读取或写入的光盘驱动。在这些情况中,每个驱动可以由一个或多个数据介质接口被连接至总线(未示出)。存储器520可以包括计算机程序产品525,其具有一个或多个程序模块,这些程序模块被配置为执行本公开的各种实施例的各种方法或动作。
通信单元540实现通过通信介质与其他电子设备进行通信。附加地,电子设备500的组件的功能可以以单个计算集群或多个计算机器来实现,这些计算机器能够通过通信连接进行通信。因此,电子设备500可以使用与一个或多个其他服务器、网络个人计算机(PC)或者另一个网络节点的逻辑连接来在联网环境中进行操作。
输入设备550可以是一个或多个输入设备,例如鼠标、键盘、追踪球等。输出设备560可以是一个或多个输出设备,例如显示器、扬声器、打印机等。电子设备500还可以根据需要通过通信单元540与一个或多个外部设备(未示出)进行通信,外部设备诸如存储设备、显示设备等,与一个或多个使得用户与电子设备500交互的设备进行通信,或者与使得电子设备500与一个或多个其他电子设备通信的任何设备(例如,网卡、调制解调器等)进行通信。这样的通信可以经由输入/输出(I/O)接口(未示出)来执行。
根据本公开的示例性实现方式,提供了一种计算机可读存储介质,其上存储有计算机可执行指令,其中计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,还提供了一种计算机程序产品,计算机程序产品被有形地存储在非瞬态计算机可读介质上并且包括计算机可执行指令,而计算机可执行指令被处理器执行以实现上文描述的方法。
这里参照根据本公开实现的方法、装置、设备和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计 算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理单元,从而生产出一种机器,使得这些指令在通过计算机或其他可编程数据处理装置的处理单元执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
可以把计算机可读程序指令加载到计算机、其他可编程数据处理装置、或其他设备上,使得在计算机、其他可编程数据处理装置或其他设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其他可编程数据处理装置、或其他设备上执行的指令实现流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实现的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或指令的一部分,模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实现,上述说明是示例性的,并非穷尽性的,并且也不限于所公开的各实现。在不偏离所说明的各实现的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改 和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实现的原理、实际应用或对市场中的技术的改进,或者使本技术领域的其他普通技术人员能理解本文公开的各个实现方式。

Claims (10)

  1. 一种生成视频内容的方法,包括:
    基于目标图像和与所述目标图像相关联的深度信息,构建三维点云数据;
    根据虚拟相机的运动轨迹,基于所述三维点云数据生成与所述虚拟相机的多个状态对应的第一组图像;以及
    基于所述第一组图像和提示项,利用目标模型生成视频内容。
  2. 根据权利要求1所述的方法,其中根据虚拟相机的运动轨迹基于所述三维点云数据生成与所述虚拟相机的多个状态对应的第一组图像包括:
    基于所述虚拟相机的所述运动轨迹,确定所述虚拟相机在所述多个状态下的多组外部参数;以及
    通过投影所述三维点云数据,生成与所述多组外部参数对应的所述第一组图像。
  3. 根据权利要求1所述的方法,其中基于所述第一组图像和提示项利用目标模型生成视频内容包括:
    利用图像修补模型处理所述第一组图像中的至少一个图像,以生成第三组图像;以及
    基于所述第三组图像和提示项,利用所述目标模型生成所述视频内容中的第二组图像。
  4. 根据权利要求3所述的方法,其中基于所述第三组图像和提示项利用所述目标模型生成所述视频内容中的第二组图像包括:
    将所述第三组图像投影至与所述三维点云数据对应的三维空间;
    通过在所述三维空间中对齐与所述第三组图像对应的多个位置,更新所述三维点云数据;
    基于经更新的所述三维点云数据,生成与所述虚拟相机的所述多个状态对应的第四组图像;以及
    向所述目标模型提供所述第四组图像和所述提示项,以生成所述 第二组图像。
  5. 根据权利要求4所述的方法,其中所述目标模型为扩散模型,所述扩散模型被配置为:
    通过添加满足预设分布的噪声,生成与所述第四组图像对应的一组隐藏特征;以及
    基于所述一组隐藏特征,生成对应的所述第二组图像。
  6. 根据权利要求1所述的方法,还包括:
    利用深度估计模型处理所述目标图像,以确定所述深度信息。
  7. 根据权利要求1所述的方法,其中所述视频内容包括与所述虚拟相机的所述多个状态对应的多个视频帧。
  8. 一种用于生成视频内容的装置,包括:
    点云构建模块,被配置为基于目标图像和与所述目标图像相关联的深度信息,构建三维点云数据;
    图像生成模块,被配置为根据虚拟相机的运动轨迹,基于所述三维点云数据生成与所述虚拟相机的多个状态对应的第一组图像;以及
    视频生成模块,被配置为基于所述第一组图像和提示项,利用目标模型生成视频内容。
  9. 一种电子设备,包括:
    至少一个处理单元;以及
    至少一个存储器,所述至少一个存储器被耦合到所述至少一个处理单元并且存储用于由所述至少一个处理单元执行的指令,所述指令在由所述至少一个处理单元执行时使所述电子设备执行根据权利要求1至7中任一项所述的方法。
  10. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序可由处理器执行以实现根据权利要求1至7中任一项所述的方法。
PCT/CN2024/096829 2024-05-31 2024-05-31 一种生成视频内容的方法、装置、设备和存储介质 Pending WO2025245877A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
PCT/CN2024/096829 WO2025245877A1 (zh) 2024-05-31 2024-05-31 一种生成视频内容的方法、装置、设备和存储介质
CN202480003522.9A CN119631417A (zh) 2024-05-31 2024-05-31 一种生成视频内容的方法、装置、设备和存储介质

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/096829 WO2025245877A1 (zh) 2024-05-31 2024-05-31 一种生成视频内容的方法、装置、设备和存储介质

Publications (1)

Publication Number Publication Date
WO2025245877A1 true WO2025245877A1 (zh) 2025-12-04

Family

ID=94902662

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/096829 Pending WO2025245877A1 (zh) 2024-05-31 2024-05-31 一种生成视频内容的方法、装置、设备和存储介质

Country Status (2)

Country Link
CN (1) CN119631417A (zh)
WO (1) WO2025245877A1 (zh)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114387326A (zh) * 2022-01-12 2022-04-22 腾讯科技(深圳)有限公司 一种视频的生成方法、装置、设备及存储介质
WO2023195301A1 (ja) * 2022-04-04 2023-10-12 ソニーグループ株式会社 表示制御装置、表示制御方法および表示制御プログラム
CN117014651A (zh) * 2022-04-29 2023-11-07 北京字跳网络技术有限公司 一种视频生成方法及装置
CN117853686A (zh) * 2023-12-26 2024-04-09 浙江大学 一种纯文本引导的任意轨迹三维场景构建及漫游视频生成方法及系统

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114387326A (zh) * 2022-01-12 2022-04-22 腾讯科技(深圳)有限公司 一种视频的生成方法、装置、设备及存储介质
WO2023195301A1 (ja) * 2022-04-04 2023-10-12 ソニーグループ株式会社 表示制御装置、表示制御方法および表示制御プログラム
CN117014651A (zh) * 2022-04-29 2023-11-07 北京字跳网络技术有限公司 一种视频生成方法及装置
CN117853686A (zh) * 2023-12-26 2024-04-09 浙江大学 一种纯文本引导的任意轨迹三维场景构建及漫游视频生成方法及系统

Also Published As

Publication number Publication date
CN119631417A (zh) 2025-03-14

Similar Documents

Publication Publication Date Title
CN114073071B (zh) 视频插帧方法及装置、计算机可读存储介质
WO2019205852A1 (zh) 确定图像捕捉设备的位姿的方法、装置及其存储介质
CN111868786B (zh) 跨设备监控计算机视觉系统
CN115209031B (zh) 视频防抖处理方法、装置、电子设备和存储介质
CN118446909A (zh) 新视角合成方法、装置、设备、介质及计算机程序产品
WO2023029418A1 (zh) 图像超分辨率模型训练方法、装置和计算机可读存储介质
CN115311397B (zh) 用于图像渲染的方法、装置、设备和存储介质
CN116723385A (zh) 处理构图的方法、装置、设备和存储介质
US8872832B2 (en) System and method for mesh stabilization of facial motion capture data
WO2026026818A1 (zh) 生成虚拟资源的方法、装置、设备和存储介质
WO2025261521A1 (zh) 生成媒体内容的方法、装置、设备和存储介质
WO2025256532A1 (zh) 发布内容的方法、装置、设备和存储介质
CN118158340B (zh) 一种运镜控制方法、装置、设备和存储介质
WO2025245877A1 (zh) 一种生成视频内容的方法、装置、设备和存储介质
WO2025060587A1 (zh) 一种视频抖动去除方法、电子设备、系统和存储介质
CN119693582A (zh) 三维重建方法、装置、设备以及存储介质
JP6967150B2 (ja) 学習装置、画像生成装置、学習方法、画像生成方法及びプログラム
CN119515674A (zh) 全景图像的生成方法、装置、电子设备及存储介质
CN117893398A (zh) 用于特效交互的方法、装置、设备和存储介质
WO2023240583A1 (zh) 一种跨媒体对应知识的生成方法和装置
KR20240086004A (ko) 디지털 휴먼 실감 가시화를 위한 컴퓨팅 장치 및 방법
CN111563956A (zh) 一种二维图片的三维显示方法、装置、设备及介质
CN120786142A (zh) 生成视频的方法、装置、设备和存储介质
US20250371819A1 (en) Method, apparatus, device and medium for switching scenarios
CN121616730B (zh) 一种图像处理的方法、装置、介质和产品