WO2024255652A1 - 视频生成方法、装置、设备、介质及程序产品 - Google Patents
视频生成方法、装置、设备、介质及程序产品 Download PDFInfo
- Publication number
- WO2024255652A1 WO2024255652A1 PCT/CN2024/097380 CN2024097380W WO2024255652A1 WO 2024255652 A1 WO2024255652 A1 WO 2024255652A1 CN 2024097380 W CN2024097380 W CN 2024097380W WO 2024255652 A1 WO2024255652 A1 WO 2024255652A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- target video
- target
- data
- features
- video
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/41—Higher-level, semantic clustering, classification or understanding of video scenes, e.g. detection, labelling or Markovian modelling of sport events or news items
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/46—Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/23418—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44008—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/85—Assembly of content; Generation of multimedia applications
Definitions
- the present disclosure relates to the field of computer technology, and in particular to a video generation method, device, equipment, medium and program product.
- Multi-modal-to-Video Generation technology can guide video generation through data in other modalities such as text and voice.
- the existing multi-modal video generation technology has low matching accuracy between different modalities, resulting in low accuracy of video generation and reduced user experience.
- the present disclosure proposes a video generation method, device, equipment, storage medium and program product to solve the technical problem of low accuracy of video generation to a certain extent.
- the present disclosure provides a video generation method, comprising:
- Acquire input data wherein the input data includes at least one of audio data or text data;
- a target video is generated based on the target video features.
- a video generation device comprising:
- An acquisition module used for acquiring input data, wherein the input data includes audio data or text data;
- An extraction module used to extract features from the input data to obtain input features of the input data
- a matching module configured to determine target video features based on the input features
- a generating module is used to generate a target video based on the target video features.
- an electronic device characterized in that it includes one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the program includes instructions for executing the method described in the first aspect or the second aspect.
- a non-volatile computer-readable storage medium containing a computer program is provided.
- the processors execute the method described in the first aspect or the second aspect.
- a computer program product comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method described in the first aspect.
- the video generation method, apparatus, device, medium and program product provided by the present disclosure generate a target video by determining corresponding target video features based on input features of input data.
- FIG. 1 is a schematic diagram of a video generation architecture according to an embodiment of the present disclosure.
- FIG. 2 is a schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.
- FIG3 is a schematic flow chart of a video generation method according to an embodiment of the present disclosure.
- FIG. 4 is a schematic diagram of a video generation method according to an embodiment of the present disclosure.
- FIG. 5 is a schematic diagram of a video generating device according to an embodiment of the present disclosure.
- a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information.
- the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
- the prompt information in response to receiving an active request from the user, may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form.
- the pop-up window may also carry a selection control for the user to choose "agree” or “disagree” to provide personal information to the electronic device.
- FIG1 is a schematic diagram of a video generation architecture of an embodiment of the present disclosure.
- the video generation architecture 100 may include a server 110, a terminal 120, and a network 130 that provides a communication link.
- the server 110 and the terminal 120 may be connected via a wired or wireless network 130.
- the server 110 may be an independent physical server, or a server cluster or distributed system consisting of multiple physical servers, or a server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, networks, etc.
- Cloud servers that provide basic cloud computing services such as services, cloud communications, middleware services, security services, and CDN.
- the terminal 120 may be implemented in hardware or software.
- the terminal 120 when the terminal 120 is implemented in hardware, it may be various electronic devices having a display screen and supporting page display, including but not limited to smart phones, tablet computers, e-book readers, laptop portable computers, and desktop computers, etc.
- the terminal 120 device when the terminal 120 device is implemented in software, it may be installed in the electronic devices listed above; it may be implemented as multiple software or software modules (such as software or software modules used to provide distributed services), or it may be implemented as a single software or software module, which is not specifically limited here.
- the video generation method provided in the embodiment of the present application can be executed by the terminal 120 or by the server 110. It should be understood that the number of terminals, networks and servers in FIG1 is only for illustration and is not intended to limit the number of terminals, networks and servers. Any number of terminals, networks and servers may be provided as required.
- FIG2 shows a schematic diagram of the hardware structure of an exemplary electronic device 200 provided by an embodiment of the present disclosure.
- the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210.
- the processor 202, the memory 204, the network module 206, and the peripheral interface 208 are connected to each other in communication within the electronic device 200 through the bus 210.
- Processor 202 may be a central processing unit (CPU), a video generator, a neural network processor (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits.
- processor 202 may be used to perform functions related to the technology described in the present disclosure.
- processor 202 may also include multiple processors integrated into a single logical component. For example, as shown in FIG. 2, processor 202 may include multiple processors 202a, 202b, and 202c.
- the memory 204 may be configured to store data (e.g., instructions, computer codes, etc.). As shown in FIG. 2 , the data stored in the memory 204 may include program instructions (e.g., program instructions for implementing the video generation method of the disclosed embodiment) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). The processor 202 may also access the program instructions and data stored in the memory 204, and execute the program instructions to operate on the data to be processed.
- the memory 204 may include a volatile storage device or a non-volatile storage device.
- the memory 204 may include a random access memory (RAM), a read-only memory (ROM), an optical disk, a magnetic disk, a hard disk, a solid-state drive (SSD), a flash memory, a memory stick, etc.
- RAM random access memory
- ROM read-only memory
- SSD solid-state drive
- flash memory a memory stick, etc.
- the network module 206 can be configured to provide communication with other external devices to the electronic device 200 via a network.
- the network can be any wired or wireless network capable of transmitting and receiving data.
- the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC) etc.), a cellular network, the Internet or a combination thereof. It is understood that the type of network is not limited to the above specific examples.
- the network module 306 can include any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc., in any combination.
- NICs network interface controllers
- the peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to achieve information input and output.
- the peripheral devices can include input devices such as a keyboard, a mouse, a touch pad, a touch screen, a microphone, and various sensors, and output devices such as a display, a speaker, a vibrator, and an indicator light.
- the bus 210 can be configured to transmit information between various components of the electronic device 200 (e.g., the processor 202, the memory 204, the network module 206, and the peripheral interface 208), such as an internal bus (e.g., a processor-memory bus), an external bus (USB port, PCI-E bus), etc.
- an internal bus e.g., a processor-memory bus
- an external bus USB port, PCI-E bus
- the architecture of the electronic device 200 may also include other components necessary for normal operation.
- the architecture of the electronic device 200 may also only include the components necessary for implementing the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.
- cross-modal matching technology is generally used to solve the matching problem between different modal data, such as video, audio, text, compression coding, etc.
- modal data such as video, audio, text, compression coding, etc.
- the matching accuracy between different modalities in the existing technology is not high, resulting in low accuracy of the video generated based on the user's input, which cannot meet the needs of users. Therefore, how to improve the accuracy of video generation has become a technical problem that needs to be solved urgently.
- the embodiments of the present disclosure provide a video generation method, apparatus, device, storage medium and program product.
- the target video is generated by determining the corresponding target video features based on the input features of the input data. Regardless of the modality of the input data, the matching with the video features can be achieved, and the target video can be accurately generated, thereby improving the accuracy of video generation.
- the user may input a piece of text data, such as a lyric, and hope to generate a video corresponding to the lyric, such as an MV.
- the text data may be subjected to feature extraction to obtain text features.
- a target video matching the text features may be determined.
- generating a corresponding target video based on the target video feature can improve the accuracy of cross-modal matching between the text modality and the video modality, thereby improving the accuracy of the generated target video.
- Fig. 3 shows a schematic flow chart of a video generation method according to an embodiment of the present disclosure.
- the video generation method 300 may further include the following steps.
- step S310 input data is acquired, where the input data includes at least one of audio data or text data.
- the input data may be data directly input or determined by a user through user operation based on an interactive interface. For example, a user may directly input text data, audio data, or video data, or may select text data, audio data, or video data, and then obtain it through a network. If it is video data, the audio data or text data in the video data may be obtained. In some embodiments, the input data may come from the same data source, such as the same audio and video data; or may come from different data sources.
- step S320 feature extraction is performed on the input data to obtain input features of the input data.
- the input data may include at least one input data slice.
- the input data slice is to cut the text data, audio data or video data input by the user into segments.
- the segment may include single semantic data with single semantic characteristics.
- the input data slice can cut long input data into short input data, thereby providing a basis for parallelization acceleration.
- the input data slice cuts the input data into segments with a single semantics, avoiding semantic confusion caused by data segments containing multiple semantics in the same segment of data when performing feature matching.
- performing feature extraction on the input data to obtain input features of the input data includes:
- the input data usually includes multiple semantic data segments. Directly extracting features from multiple semantic data segments may lead to semantic confusion and reduce the quality of the final generated video. For example, for the text data "text1, text2”, text1 and text2 have their own semantics respectively. If feature extraction is performed as one data, it may cause semantic confusion. In order to solve the problem of semantic confusion, a multi-semantic data can be split into multiple single semantic data. For example, the text data "text1, text2" can be split into two single semantic data “text1” and “text2” and feature extraction is performed on each of them to obtain two input features F_text1 and F_text2.
- subsequent feature matching can be performed to avoid semantic confusion caused by semantic intersections between each other, which can effectively improve the quality of semantic matching.
- semantic confusion caused by semantic intersections between each other For example, "The sun shines on the incense burner and produces purple smoke, and the waterfall hangs in front of the river from afar" can be split into “The sun shines on the incense burner and produces purple smoke” and "The waterfall hangs in front of the river from afar".
- the input data may include at least one of text data, audio data or video data, and after segmentation, it may include less or more single semantic data, which is not limited here.
- obtaining a plurality of input features corresponding to the input data based on the single semantic data further includes:
- the multimodal feature extraction process is performed on the monosemantic data to obtain the input feature.
- FIG. 4 shows a schematic diagram of a video generation method according to an embodiment of the present disclosure.
- the feature extraction processing method used is consistent with that when obtaining the video feature library, for example, a feature extraction network for obtaining video features in the video feature library can be used to extract features from the input data.
- Feature extraction extracts a feature vector from a single semantic data segment, thereby effectively reducing the dimension of the data.
- step S330 target video features are determined based on the input features.
- the input feature may be a feature of a textual mode or an audio mode, and feature matching may be performed in a video feature library based on the input feature to determine a matching target video feature, as shown in FIG4 .
- the video feature library may include multiple video features, and the video features may have corresponding semantic characteristics.
- the input feature may be matched with the semantics of the video feature to determine the target video feature in the video feature.
- the video feature library may also include text features corresponding to the video features.
- determining the target video feature based on the input feature includes:
- the video feature with the highest semantic relevance is determined as the target video feature.
- the cosine distance may be calculated based on the input feature and all the video features in the video feature library to obtain the semantic relevance, and the video feature corresponding to the minimum cosine distance may be selected as the matching result.
- method 300 may further include:
- Multimodal feature extraction processing is performed based on multiple video data to obtain multiple video features with semantic characteristics.
- Video data can also be pre-classified based on attributes such as theme and style, which can provide a data basis for generating videos with different themes and styles when generating subsequent target videos.
- Multimodal feature extraction can be performed on each video data. For example, feature extraction can be performed on video data based on a multimodal feature extraction network to obtain video features with semantics. Because it is necessary to retrieve and match with data of other modalities such as text and audio in the future, it is not possible to obtain a single modal feature set like the single modal video feature extractor used in the prior art. The multimodal feature set can be continuously updated in subsequent use, and there is no need to re-create the database every time a feature match is made. It should be understood that the method of extracting features from input data is consistent with the method of extracting features to obtain multimodal features.
- step S340 a target video is generated based on the target video features.
- the target video feature has a corresponding video clip, and the corresponding target video can be generated based on the target video feature.
- generating a target video based on the target video features includes:
- the aligned target video slices are spliced based on the source timestamp to obtain the target video.
- the target video feature corresponds to each input feature, that is, to each input data slice (i.e., single semantic data). After retrieving the corresponding target video slice for each input data slice, these target video slices need to be combined into a complete video.
- the timestamps of the input data and the matched target video slices are not always aligned, so it is necessary to align the timestamps of the two and then splice the target video slices to obtain the final target video.
- the source timestamp of the text data may be set based on a preset playback speed.
- a corresponding timestamp can be set for it.
- the text data can be played based on a preset playback speed v.
- the text data text includes multiple text data segments text1, text2, ...texti, ..., and the corresponding data lengths are L1, L2, ...Li, ..., respectively. Then the start timestamp of the text data segment text1 is 0, and the end timestamp is L1/v.
- the start timestamp of the text data segment text2 is L1/v
- the end timestamp is (L1+L2)/v.
- a timestamp can be set for the text data text. It should be understood that the above timestamps are only examples and are not intended to limit the timestamps.
- the set timestamps may include or exclude the start timestamp and/or the end timestamp, and may also include timestamps at other locations, which are not limited here.
- aligning the target timestamp of the target video slice with the source timestamp of the input data to obtain an aligned target video slice further comprises:
- the target timestamp is aligned with the source timestamp based on the target duration and the source duration to obtain an aligned target video slice.
- aligning the target timestamp with the source timestamp based on the target duration and the source duration to obtain an aligned target video slice includes:
- the target video slice whose target duration is equal to the source duration is directly used as the aligned target video slice;
- the target video slice whose target duration is shorter than the source duration is extended by inserting frames to obtain aligned target video slices.
- each target video slice can be processed one by one in turn to align the timestamps of the target video slice with the input data slice.
- the target video slice T is filled into the time period s1s2 corresponding to the input data slice S as the matching result.
- the target video slice T is cropped so that the duration t2'-t1' of the cropped target video slice T' is equal to the source duration s2-s1, and the cropped target video slice T' is used as the matching result to fill in the time period s1s2 corresponding to the input data slice S.
- all target video slices can be arranged in the order of the timestamps of the input data, and these target video slices can be rendered into the final target video.
- the input data may further include indication information for indicating the attributes of the target video.
- determining the target video features based on the input features includes:
- the attributes may include style, theme, etc.
- the style may include funny style, classical style, etc.
- the theme may include natural theme, animal theme, etc.
- the user may indicate the attributes such as style or theme of the generated target video by inputting data. For example, if the user inputs a video data A with style F1, and the indication information includes that the style of the target video is F2, then the target video generated according to the embodiment of the present disclosure is to change the style of the video data A to F2.
- the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server.
- the method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other.
- one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.
- the present disclosure further provides a video generation device, referring to FIG5 , wherein the video generation device includes:
- An acquisition module used to acquire input data, wherein the input data includes at least one of audio data or text data;
- An extraction module used to extract features from the input data to obtain input features of the input data
- a matching module configured to determine target video features based on the input features
- a generating module is used to generate a target video based on the target video features.
- the above device is described by dividing it into various modules according to its functions.
- the functions of each module can be implemented in the same or multiple software and/or hardware.
- the device of the above embodiment is used to implement the corresponding video generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
- the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions.
- the computer instruction is used to cause the computer to execute the video generation method described in any of the above embodiments.
- the computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology.
- Information can be computer-readable instructions, data structures, modules of programs, or other data.
- Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
- PRAM phase change memory
- SRAM static random access memory
- DRAM dynamic random access memory
- RAM random access memory
- ROM read-only memory
- EEPROM electrically erasable programmable
- the computer instructions stored in the storage medium of the above embodiments are used to enable the computer to execute the video generation method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
- DRAM dynamic RAM
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Computational Linguistics (AREA)
- Software Systems (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
本公开提供一种视频生成方法、装置、设备、存储介质及程序产品。该方法包括:获取输入数据,所述输入数据包括音频数据或文本数据中的至少一种;对所述输入数据进行特征提取,得到所述输入数据的输入特征;基于所述输入特征确定目标视频特征;基于所述目标视频特征生成目标视频。
Description
本申请要求于2023年6月13日提交中国国家知识产权局、申请号为202310701132.6、发明名称为“视频生成方法、装置、设备、介质及程序产品”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本公开涉及计算机技术领域,尤其涉及一种视频生成方法、装置、设备、介质及程序产品。
基于多模态的视频生成技术(Multi-Modal-to-Video Generation)可以通过文字、语音等其他模态的数据来指导视频的生成。然而,现有的多模态视频生成技术中针对不同模态之间的匹配准确度不高,导致视频生成的准确性不高,降低了用户体验。
发明内容
本公开提出一种视频生成方法、装置、设备、存储介质及程序产品,以在一定程度上解决视频生成的准确性不高的技术问题。
本公开第一方面,提供了一种视频生成方法,包括:
获取输入数据,所述输入数据包括音频数据或文本数据中的至少一种;
对所述输入数据进行特征提取,得到所述输入数据的输入特征;
基于所述输入特征确定目标视频特征;
基于所述目标视频特征生成目标视频。
本公开第二方面,提供了一种视频生成装置,包括:
获取模块,用于获取输入数据,所述输入数据包括音频数据或文本数据;
提取模块,用于对所述输入数据进行特征提取,得到所述输入数据的输入特征;
匹配模块,用于基于所述输入特征确定目标视频特征;
生成模块,用于基于所述目标视频特征生成目标视频。
本公开第三方面,提供了一种电子设备,其特征在于,包括一个或者多个处理器、存储器;和一个或多个程序,其中所述一个或多个程序被存储在所述存储器中,并且被所述一个或多个处理器执行,所述程序包括用于执行根据第一方面或第二方面所述的方法的指令。
本公开第四方面,提供了一种包含计算机程序的非易失性计算机可读存储介质,当所述计算机程序被一个或多个处理器执行时,使得所述处理器执行第一方面或第二方面所述的方法。
本公开第五方面,提供了一种计算机程序产品,包括计算机程序指令,当所述计算机程序指令在计算机上运行时,使得计算机执行第一方面所述的方法。
从上面所述可以看出,本公开提供的一种视频生成方法、装置、设备、介质及程序产品,通过基于输入数据的输入特征确定对应的目标视频特征,从而生成目标视频。
为了更清楚地说明本公开或相关技术中的技术方案,下面将对实施例或相关技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本公开的实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1为本公开实施例的视频生成架构的示意图。
图2为本公开实施例的示例性电子设备的硬件结构示意图。
图3为本公开实施例的视频生成方法的示意性流程图。
图4为本公开实施例的视频生成方法的示意性原理图。
图5为本公开实施例的视频生成装置的示意图。
为使本公开的目的、技术方案和优点更加清楚明白,以下结合具体实施例,并参照附图,对本公开进一步详细说明。
需要说明的是,除非另外定义,本公开实施例使用的技术术语或者科学术语应当为本公开所属领域内具有一般技能的人士所理解的通常意义。本公开实施例中使用的“第一”、“第二”以及类似的词语并不表示任何顺序、数量或者重要性,而只是用来区分不同的组成部分。“包括”或者“包含”等类似的词语意指出现该词前面的元件或者物件涵盖出现在该词后面列举的元件或者物件及其等同,而不排除其他元件或者物件。“连接”或者“相连”等类似的词语并非限定于物理的或者机械的连接,而是可以包括电性的连接,不管是直接的还是间接的。“上”、“下”、“左”、“右”等仅用于表示相对位置关系,当被描述对象的绝对位置改变后,则该相对位置关系也可能相应地改变。
可以理解的是,在使用本公开各实施例公开的技术方案之前,均应当依据相关法律法规通过恰当的方式对本公开所涉及个人信息的类型、使用范围、使用场景等告知用户并获得用户的授权。
例如,在响应于接收到用户的主动请求时,向用户发送提示信息,以明确地提示用户,其请求执行的操作将需要获取和使用到用户的个人信息。从而,使得用户可以根据提示信息来自主地选择是否向执行本公开技术方案的操作的电子设备、应用程序、服务器或存储介质等软件或硬件提供个人信息。
作为一种可选的但非限定性的实现方式,响应于接收到用户的主动请求,向用户发送提示信息的方式例如可以是弹窗的方式,弹窗中可以以文字的方式呈现提示信息。此外,弹窗中还可以承载供用户选择“同意”或者“不同意”向电子设备提供个人信息的选择控件。
可以理解的是,上述通知和获取用户授权过程仅是示意性的,不对本公开的实现方式构成限定,其它满足相关法律法规的方式也可应用于本公开的实现方式中。
图1示出了本公开实施例的视频生成架构的示意图。参考图1,该视频生成架构100可以包括服务器110、终端120以及提供通信链路的网络130。服务器110和终端120之间可通过有线或无线的网络130连接。其中,服务器110可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络
服务、云通信、中间件服务、安全服务、CDN等基础云计算服务的云服务器。
终端120可以是硬件或软件实现。例如,终端120为硬件实现时,可以是具有显示屏并且支持页面显示的各种电子设备,包括但不限于智能手机、平板电脑、电子书阅读器、膝上型便携计算机和台式计算机等等。终端120设备为软件实现时,可以安装在上述所列举的电子设备中;其可以实现成多个软件或软件模块(例如用来提供分布式服务的软件或软件模块),也可以实现成单个软件或软件模块,在此不做具体限定。
需要说明的是,本申请实施例所提供的视频生成方法可以由终端120来执行,也可以由服务器110来执行。应了解,图1中的终端、网络和服务器的数目仅为示意,并不旨在对其进行限制。根据实现需要,可以具有任意数目的终端、网络和服务器。
图2示出了本公开实施例所提供的示例性电子设备200的硬件结构示意图。如图2所示,电子设备200可以包括:处理器202、存储器204、网络模块206、外围接口208和总线210。其中,处理器202、存储器204、网络模块206和外围接口208通过总线210实现彼此之间在电子设备200的内部的通信连接。
处理器202可以是中央处理器(Central Processing Unit,CPU)、视频生成器、神经网络处理器(NPU)、微控制器(MCU)、可编程逻辑器件、数字信号处理器(DSP)、应用专用集成电路(Application Specific Integrated Circuit,ASIC)、或者一个或多个集成电路。处理器202可以用于执行与本公开描述的技术相关的功能。在一些实施例中,处理器202还可以包括集成为单一逻辑组件的多个处理器。例如,如图2所示,处理器202可以包括多个处理器202a、202b和202c。
存储器204可以配置为存储数据(例如,指令、计算机代码等)。如图2所示,存储器204存储的数据可以包括程序指令(例如,用于实现本公开实施例的视频生成方法的程序指令)以及要处理的数据(例如,存储器可以存储其他模块的配置文件等)。处理器202也可以访问存储器204存储的程序指令和数据,并且执行程序指令以对要处理的数据进行操作。存储器204可以包括易失性存储装置或非易失性存储装置。在一些实施例中,存储器204可以包括随机访问存储器(RAM)、只读存储器(ROM)、光盘、磁盘、硬盘、固态硬盘(SSD)、闪存、存储棒等。
网络模块206可以配置为经由网络向电子设备200提供与其他外部设备的通信。该网络可以是能够传输和接收数据的任何有线或无线的网络。例如,该网络可以是有线网络、本地无线网络(例如,蓝牙、WiFi、近场通信(NFC)等)、蜂窝网络、因特网、或上述的组合。可以理解的是,网络的类型不限于上述具体示例。在一些实施例中,网络模块306可以包括任意数量的网络接口控制器(NIC)、射频模块、接收发器、调制解调器、路由器、网关、适配器、蜂窝网络芯片等的任意组合。
外围接口208可以配置为将电子设备200与一个或多个外围装置连接,以实现信息输入及输出。例如,外围装置可以包括键盘、鼠标、触摸板、触摸屏、麦克风、各类传感器等输入设备以及显示器、扬声器、振动器、指示灯等输出设备。
总线210可以被配置为在电子设备200的各个组件(例如处理器202、存储器204、网络模块206和外围接口208)之间传输信息,诸如内部总线(例如,处理器-存储器总线)、外部总线(USB端口、PCI-E总线)等。
需要说明的是,尽管上述电子设备200的架构仅示出了处理器202、存储器204、网络模块206、外围接口208和总线210,但是在具体实施过程中,该电子设备200的架构还可以包括实现正常运行所必需的其他组件。此外,本领域的技术人员可以理解的是,上述电子设备200的架构中也可以仅包含实现本公开实施例方案所必需的组件,而不必包含图中所示的全部组件。
在现有的视频生成技术中,一般依赖跨模态匹配技术来解决不同模态数据之间的匹配问题,如视频、音频、文本、压缩编码等。然而,现有技术中针对不同模态之间的匹配准确度不高,导致基于用户的输入所生成的视频准确性不高,不能满足用户的需求。因此,如何提高视频生成的准确性成为了亟需解决的技术问题。
鉴于此,本公开实施例提供了一种视频生成方法、装置、设备、存储介质及程序产品。通过基于输入数据的输入特征确定对应的目标视频特征,从而生成目标视频。无论输入数据是何种模态均可以实现与视频特征的匹配,准确地生成目标视频,提高了视频生成的准确性。
具体地,用户可以输入一段文本数据,例如一段歌词,希望生成与该段歌词对应的视频,例如MV。可以基于本公开实施例的视频生成方法,对文本数据进行特征提取得到文本特征。再基于该文本特征确定与该文本特征匹配的目
标视频特征。然后基于该目标视频特征生成对应的目标视频能够提高文本模态与视频模态的跨模态匹配的准确性,从而提高所生成的目标视频的准确度。
参见图3,图3示出了根据本公开实施例的视频生成方法的示意性流程图。图3中,视频生成方法300可以进一步包括如下步骤。
在步骤S310,获取输入数据,所述输入数据包括音频数据或文本数据中的至少一种。
其中,输入数据可以是用户基于交互界面通过用户操作直接输入或确定的数据。例如,用户可以直接输入文本数据,音频数据或视频数据,也可以选择文本数据,音频数据或视频数据,进而经由网络获取,如果是视频数据,获取视频数据中的音频数据或文本数据。在一些实施例中,输入数据可以来自于同一数据源,例如同一音视频数据;也可以来自于不同的数据源。
在步骤S320,对所述输入数据进行特征提取,得到所述输入数据的输入特征。
在一些实施例中,输入数据可以包括至少一个输入数据切片。其中,输入数据切片,即将用户输入的文本数据、音频数据或视频数据切分成片段。该片段可以包括具有单语义特性的单语义数据。一方面,输入数据切片可以将长输入数据切分为短输入数据,从而为并行化加速提供了基础。另一方面,输入数据切片将输入数据切分为具有单一语义的片段,避免进行特征匹配时同一段数据中包含多个语义的数据段所造成语义混淆。
在一些实施例中,对所述输入数据进行特征提取,得到所述输入数据的输入特征,包括:
对所述输入数据进行切分得到单语义数据;
基于所述单语义数据得到与输入数据对应的的多个所述输入特征。
其中,输入数据通常包括多个语义的数据段,直接对多个语义的数据段进行特征提取可能导致语义混淆,降低最终生成视频的质量。例如,对于文本数据“text1,text2”,text1和text2分别具有各自的语义,如果作为一个数据进行特征提取可能会造成语义的混乱。为了解决语义混乱的问题,可以将一个多语义数据拆分为多个单语义数据,例如将文本数据“text1,text2”拆分为两个单语义数据“text1”和“text2”分别进行特征提取得到2个输入特征F_text1和F_text2。然后基于输入特征F_text1和F_text2来进行后续的特征匹配,则可以避免相互之间的语义交叉导致的语义混乱,能够有效的提升语义匹配的质量。
例如可以将“日照香炉生紫烟,遥看瀑布挂前川”拆分成“日照香炉生紫烟”和“遥看瀑布挂前川”。应了解,上述输入数据的切分仅为举例,并不旨在对输入特征的模态和切分得到的单语义数据的数量进行限制,输入数据可以包括文本数据、音频数据或视频数据中的至少一种,切分后可以包括更少或更多的单语义数据,在此不做限制。
在一些实施例中,基于所述单语义数据得到与输入数据对应的的多个所述输入特征,进一步包括:
对所述单语义数据进行所述多模态特征提取处理,得到所述输入特征。
具体地,参见图4,图4示出了根据本公开实施例的视频生成方法的示意性原理图。图4中,在对输入数据进行特征提取时,所使用的特征提取的处理方式与得到视频特征库时一致,例如可以使得到对视频特征库中视频特征的特征提取网络对输入数据进行特征提取。特征提取将从单语义数据段中提取出一个特征向量,从而有效地降低数据的维度。
在步骤S330,基于所述输入特征确定目标视频特征。
其中,输入特征可以是文本模态或音频模态的特征,可以基于该输入特征在视频特征库中进行特征匹配从中确定相匹配的目标视频特征,如图4所示。其中,视频特征库可以包括多个视频特征,视频特征可以具有对应的语义特性。可以将输入特征与视频特征的语义进行匹配,确定视频特征中的目标视频特征。进一步地,视频特征库还可以包括于视频特征对应的文本特征。
在一些实施例中,基于所述输入特征确定目标视频特征,包括:
计算所述输入特征与视频特征库中的视频特征之间的语义相关度;
将所述语义相关度最高的视频特征确定为所述目标视频特征。
具体地,可以基于输入特征与视频特征库中所有的视频特征计算余弦距离得到语义相关度,并选取最小余弦距离对应的视频特征作为匹配结果。
在一些实施例中,方法300还可以包括:
基于多个视频数据进行多模态特征提取处理,得到多个具有语义特性的所述视频特征。
其中,多个视频特征则形成了视频特征库。还可以基于主题、风格等属性对视频数据进行预分类,可以在后续的目标视频生成时为生成不同主题、风格的视频提供数据基础。针对每个视频数据均可以进行多模态特征提取,例如可以基于多模态特征提取网络对视频数据进行特征提取得到具有语义的视频特
征。因为后续需要与文本、音频等其他模态的数据进行检索匹配,所以不能像现有技术中所使用的单一模态的视频特征提取器得到单一模态的特征集合。多模态特征集合可以在后续的使用中不断更新,不需要在每次特征匹配时重新制作数据库。应了解,对输入数据进行特征提取的方式与得到多模态特征的特征提取方式一致。
在步骤S340,基于所述目标视频特征生成目标视频。
其中,目标视频特征具有对应的视频片段,基于目标视频特征则可以生成对应的目标视频。
在一些实施例中,基于所述目标视频特征生成目标视频,包括:
基于所述目标视频特征确定对应的目标视频切片;
将所述目标视频切片的目标时间戳与所述输入数据的源时间戳对齐,得到对齐后的目标视频切片;
基于所述源时间戳将对齐后的所述目标视频切片进行拼接,得到所述目标视频。
其中,目标视频特征对应于每个输入特征,即对应于每个输入数据切片(即单语义数据)。在为每一个输入数据切片检索到对应的目标视频切片后,需要将这些目标视频切片组合成完整的视频。然而,输入数据与匹配到的目标视频切片的时间戳不总是对齐的,所以需要对齐二者的时间戳后进行目标视频切片的拼接,得到最终的目标视频。
在一些实施例中,当输入数据为文本数据时,可以基于预设播放速度设置所述文本数据的源时间戳。
其中,对于本身具有时间戳的输入数据,例如音频数据,则可以将其自身的时间戳作为源时间戳。对于本身不具有时间戳的输入数据,例如文本数据,则可以为其设置相应的时间戳。具体地,对于文本数据text,可以基于预设播放速度v对文本数据进行播放,例如,文本数据text包括多个文本数据段text1,text2,……texti,……,分别对应的数据长度为L1,L2,……Li,……,则文本数据段text1起始时间戳为0,结束时间戳为L1/v,文本数据段text2的起始时间戳为L1/v,结束时间戳为(L1+L2)/v,依此类推,可以对文本数据text设置时间戳。应了解,上述时间戳仅为示例,并不旨在对时间戳进行限制,设置的时间戳可以包括或不包括起始时间戳和/或结束时间戳,也可以包括其他位置的时间戳,在此不做限制。
在一些实施例中,将所述目标视频切片的目标时间戳与所述输入数据的源时间戳对齐,得到对齐后的目标视频切片,进一步包括:
基于所述目标时间戳得到所述目标视频切片的目标时长,以及基于所述源时间戳得到所述输入数据切片的源时长;
基于所述目标时长和所述源时长进行所述目标时间戳与所述源时间戳的对齐,得到对齐后的目标视频切片。
在一些实施例中,基于所述目标时长和所述源时长进行所述目标时间戳与所述源时间戳的对齐,得到对齐后的目标视频切片,包括:
将所述目标时长等于所述源时长的目标视频切片,直接作为对齐后的目标视频切片;
针对所述目标时长大于所述源时长的目标视频切片进行裁剪,得到对齐后的目标视频切片;
针对所述目标时长小于所述源时长的目标视频切片进行插帧延长,得到对齐后的目标视频切片。
具体地,可以逐个依次处理每个目标视频切片,来使目标视频切片与输入数据切片的时间戳对齐。以输入数据切片S的起始时间戳s1、结束时间戳s2和目标视频切片T的起始时间戳t1、结束时间戳t2为例:
若源时长s2-s1=目标时长t2-t1,则将目标视频切片T作为匹配结果填入输入数据切片S对应的时间段s1s2。
若源时长s2-s1<目标时长t2-t1,则对目标视频切片T进行裁剪,使得裁剪后的目标视频切片T’的时长t2'-t1'=源时长s2-s1,并将裁剪后的目标视频切片T’作为匹配结果填入输入数据切片S对应的时间段s1s2。
若源时长s2-s1>目标时长t2-t1,则对目标视频切片T进行插帧延长,使得插帧延长后的目标视频切片T’的时长t2'-t1'=s2-s1,并将插帧延长后的目标视频切片T’作为匹配结果填入输入数据切片S对应的时间段s1s2。
具体地,在完成所有目标视频切片的匹配和时间戳调整对齐后,就可以将所有的目标视频切片按照输入数据的时间戳顺序排列好,并将这些目标视频切片渲染成最终的目标视频。
在一些实施例中,输入数据还可以包括用于指示目标视频的属性的指示信息。在一些实施例中,基于所述输入特征确定目标视频特征,包括:
基于所述输入特征与视频特征库中具有所述属性的视频特征进行匹配,确
定所述视频特征中的目标视频特征。
其中,属性可以包括风格、主题等,例如风格可以包括搞笑风格、古典风格等;主题可以包括自然主题、动物主题等。用户可以通过输入数据指示生成的目标视频的风格或主题等属性,例如用户输入一段风格为F1的视频数据A,且指示信息包括目标视频的风格为F2,则根据本公开实施例所生成的目标视频为将视频数据A的风格变为F2。
需要说明的是,本公开实施例的方法可以由单个设备执行,例如一台计算机或服务器等。本实施例的方法也可以应用于分布式场景下,由多台设备相互配合来完成。在这种分布式场景的情况下,这多台设备中的一台设备可以只执行本公开实施例的方法中的某一个或多个步骤,这多台设备相互之间会进行交互以完成所述的方法。
需要说明的是,上述对本公开的一些实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于上述实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。
基于同一技术构思,与上述任意实施例方法相对应的,本公开还提供了一种视频生成装置,参见图5,所述视频生成装置包括:
获取模块,用于获取输入数据,所述输入数据包括音频数据或文本数据中的至少一种;
提取模块,用于对所述输入数据进行特征提取,得到所述输入数据的输入特征;
匹配模块,用于基于所述输入特征确定目标视频特征;
生成模块,用于基于所述目标视频特征生成目标视频。
为了描述的方便,描述以上装置时以功能分为各种模块分别描述。当然,在实施本公开时可以把各模块的功能在同一个或多个软件和/或硬件中实现。
上述实施例的装置用于实现前述任一实施例中相应的视频生成方法,并且具有相应的方法实施例的有益效果,在此不再赘述。
基于同一技术构思,与上述任意实施例方法相对应的,本公开还提供了一种非暂态计算机可读存储介质,所述非暂态计算机可读存储介质存储计算机指
令,所述计算机指令用于使所述计算机执行如上任一实施例所述的视频生成方法。
本实施例的计算机可读介质包括永久性和非永久性、可移动和非可移动媒体可以由任何方法或技术来实现信息存储。信息可以是计算机可读指令、数据结构、程序的模块或其他数据。计算机的存储介质的例子包括,但不限于相变内存(PRAM)、静态随机存取存储器(SRAM)、动态随机存取存储器(DRAM)、其他类型的随机存取存储器(RAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、快闪记忆体或其他内存技术、只读光盘只读存储器(CD-ROM)、数字多功能光盘(DVD)或其他光学存储、磁盒式磁带,磁带磁磁盘存储或其他磁性存储设备或任何其他非传输介质,可用于存储可以被计算设备访问的信息。
上述实施例的存储介质存储的计算机指令用于使所述计算机执行如上任一实施例所述的视频生成方法,并且具有相应的方法实施例的有益效果,在此不再赘述。
所属领域的普通技术人员应当理解:以上任何实施例的讨论仅为示例性的,并非旨在暗示本公开的范围(包括权利要求)被限于这些例子;在本公开的思路下,以上实施例或者不同实施例中的技术特征之间也可以进行组合,步骤可以以任意顺序实现,并存在如上所述的本公开实施例的不同方面的许多其它变化,为了简明它们没有在细节中提供。
另外,为简化说明和讨论,并且为了不会使本公开实施例难以理解,在所提供的附图中可以示出或可以不示出与集成电路(IC)芯片和其它部件的公知的电源/接地连接。此外,可以以框图的形式示出装置,以便避免使本公开实施例难以理解,并且这也考虑了以下事实,即关于这些框图装置的实施方式的细节是高度取决于将要实施本公开实施例的平台的(即,这些细节应当完全处于本领域技术人员的理解范围内)。在阐述了具体细节(例如,电路)以描述本公开的示例性实施例的情况下,对本领域技术人员来说显而易见的是,可以在没有这些具体细节的情况下或者这些具体细节有变化的情况下实施本公开实施例。因此,这些描述应被认为是说明性的而不是限制性的。
尽管已经结合了本公开的具体实施例对本公开进行了描述,但是根据前面的描述,这些实施例的很多替换、修改和变型对本领域普通技术人员来说将是显而易见的。例如,其它存储器架构(例如,动态RAM(DRAM))可以使用
所讨论的实施例。
本公开实施例旨在涵盖落入所附权利要求的宽泛范围之内的所有这样的替换、修改和变型。因此,凡在本公开实施例的精神和原则之内,所做的任何省略、修改、等同替换、改进等,均应包含在本公开的保护范围之内。
Claims (12)
- 一种视频生成方法,包括:获取输入数据,所述输入数据包括音频数据或文本数据中的至少一种;对所述输入数据进行特征提取,得到所述输入数据的输入特征;基于所述输入特征确定目标视频特征;基于所述目标视频特征生成目标视频。
- 根据权利要求1的方法,其中,对所述输入数据进行特征提取,得到所述输入数据的输入特征,包括:对所述输入数据进行切分得到单语义数据;基于所述单语义数据得到与输入数据对应的的多个所述输入特征。
- 根据权利要求1的方法,其中,所述基于所述输入特征确定目标视频特征,包括:计算所述输入特征与视频特征库中的视频特征之间的语义相关度;将所述语义相关度最高的视频特征确定为所述目标视频特征。
- 根据权利要求1的方法,其中,基于所述目标视频特征生成目标视频,包括:基于所述目标视频特征确定对应的目标视频切片;将所述目标视频切片的目标时间戳与所述输入数据的源时间戳对齐,得到对齐后的目标视频切片;基于所述源时间戳将对齐后的所述目标视频切片进行拼接,得到所述目标视频。
- 根据权利要求4的方法,其中,所述输入数据包括至少一个单语义数据;将所述目标视频切片的目标时间戳与所述输入数据的源时间戳对齐,得到对齐后的目标视频切片,进一步包括:基于所述目标时间戳得到所述目标视频切片的目标时长,以及基于所述源时间戳得到所述单语义数据的源时长;基于所述目标时长和所述源时长进行所述目标时间戳与所述源时间戳的 对齐,得到对齐后的目标视频切片。
- 根据权利要求5的方法,其中,基于所述目标时长和所述源时长进行所述目标时间戳与所述源时间戳的对齐,得到对齐后的目标视频切片,包括:将所述目标时长等于所述源时长的目标视频切片,直接作为对齐后的目标视频切片;针对所述目标时长大于所述源时长的目标视频切片进行裁剪,得到对齐后的目标视频切片;针对所述目标时长小于所述源时长的目标视频切片进行插帧延长,得到对齐后的目标视频切片。
- 根据权利要求1的方法,其中,所述输入数据还包括用于指示目标视频的属性的指示信息;则基于所述输入特征确定目标视频特征,包括:基于所述输入特征与视频特征库中具有所述属性的视频特征进行匹配,确定所述视频特征中的目标视频特征。
- 根据权利要求1的方法,其中,所述输入数据为文本数据,所述方法还包括:基于预设播放速度设置所述文本数据的源时间戳。
- 一种视频生成装置,包括:获取模块,用于获取输入数据,所述输入数据包括音频数据或文本数据;提取模块,用于对所述输入数据进行特征提取,得到所述输入数据的输入特征;匹配模块,用于基于所述输入特征确定目标视频特征;生成模块,用于基于所述目标视频特征生成目标视频。
- 一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时实现如权利要求1至8任意一项所述的方法。
- 一种非暂态计算机可读存储介质,所述非暂态计算机可读存储介质 存储计算机指令,所述计算机指令用于使计算机执行权利要求1至8任一所述方法。
- 一种计算机程序产品,包括计算机程序指令,当所述计算机程序指令在计算机上运行时,使得计算机执行权利要求1至8任一所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310701132.6 | 2023-06-13 | ||
| CN202310701132.6A CN116634246A (zh) | 2023-06-13 | 2023-06-13 | 视频生成方法、装置、设备、介质及程序产品 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024255652A1 true WO2024255652A1 (zh) | 2024-12-19 |
Family
ID=87613441
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/097380 Ceased WO2024255652A1 (zh) | 2023-06-13 | 2024-06-04 | 视频生成方法、装置、设备、介质及程序产品 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN116634246A (zh) |
| WO (1) | WO2024255652A1 (zh) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116634246A (zh) * | 2023-06-13 | 2023-08-22 | 北京字跳网络技术有限公司 | 视频生成方法、装置、设备、介质及程序产品 |
| CN117789099B (zh) * | 2024-02-26 | 2024-05-28 | 北京搜狐新媒体信息技术有限公司 | 视频特征提取方法及装置、存储介质及电子设备 |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200201904A1 (en) * | 2018-12-21 | 2020-06-25 | AdLaunch International Inc. | Generation of a video file |
| CN114242070A (zh) * | 2021-12-20 | 2022-03-25 | 阿里巴巴(中国)有限公司 | 一种视频生成方法、装置、设备及存储介质 |
| CN115237248A (zh) * | 2022-06-20 | 2022-10-25 | 北京有竹居网络技术有限公司 | 虚拟对象的展示方法、装置、设备、存储介质及程序产品 |
| CN115952317A (zh) * | 2022-07-12 | 2023-04-11 | 北京字跳网络技术有限公司 | 视频处理方法、装置、设备、介质及程序产品 |
| CN115967833A (zh) * | 2021-10-09 | 2023-04-14 | 北京字节跳动网络技术有限公司 | 视频生成方法、装置、设备计存储介质 |
| CN116012753A (zh) * | 2022-12-21 | 2023-04-25 | 平安银行股份有限公司 | 视频处理方法、装置、计算机设备及计算机可读存储介质 |
| CN116634246A (zh) * | 2023-06-13 | 2023-08-22 | 北京字跳网络技术有限公司 | 视频生成方法、装置、设备、介质及程序产品 |
-
2023
- 2023-06-13 CN CN202310701132.6A patent/CN116634246A/zh active Pending
-
2024
- 2024-06-04 WO PCT/CN2024/097380 patent/WO2024255652A1/zh not_active Ceased
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200201904A1 (en) * | 2018-12-21 | 2020-06-25 | AdLaunch International Inc. | Generation of a video file |
| CN115967833A (zh) * | 2021-10-09 | 2023-04-14 | 北京字节跳动网络技术有限公司 | 视频生成方法、装置、设备计存储介质 |
| CN114242070A (zh) * | 2021-12-20 | 2022-03-25 | 阿里巴巴(中国)有限公司 | 一种视频生成方法、装置、设备及存储介质 |
| CN115237248A (zh) * | 2022-06-20 | 2022-10-25 | 北京有竹居网络技术有限公司 | 虚拟对象的展示方法、装置、设备、存储介质及程序产品 |
| CN115952317A (zh) * | 2022-07-12 | 2023-04-11 | 北京字跳网络技术有限公司 | 视频处理方法、装置、设备、介质及程序产品 |
| CN116012753A (zh) * | 2022-12-21 | 2023-04-25 | 平安银行股份有限公司 | 视频处理方法、装置、计算机设备及计算机可读存储介质 |
| CN116634246A (zh) * | 2023-06-13 | 2023-08-22 | 北京字跳网络技术有限公司 | 视频生成方法、装置、设备、介质及程序产品 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116634246A (zh) | 2023-08-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20240420458A1 (en) | Cross-modal data processing method and apparatus, device, medium, and program product | |
| WO2024255652A1 (zh) | 视频生成方法、装置、设备、介质及程序产品 | |
| CN114445047B (zh) | 工作流生成方法、装置、电子设备及存储介质 | |
| US11036468B2 (en) | Human-computer interface for navigating a presentation file | |
| WO2019169723A1 (zh) | 测试用例选择方法、装置、设备以及计算机可读存储介质 | |
| CN115756449B (zh) | 一种页面复用方法、装置、存储介质及电子设备 | |
| CN110727417A (zh) | 一种数据处理方法和装置 | |
| CN116188250A (zh) | 图像处理方法、装置、电子设备及存储介质 | |
| CN118741264A (zh) | 生成视频的方法、装置、电子设备及存储介质 | |
| WO2024230570A1 (zh) | 一种人工智能设备对话的控制方法、装置、设备及介质 | |
| CN115098729A (zh) | 视频处理方法、样本生成方法、模型训练方法及装置 | |
| CN117744651A (zh) | 一种语言大模型融合nlu的槽位信息抽取方法及装置 | |
| CN117009482A (zh) | 对话处理方法、装置、电子设备及存储介质 | |
| CN115617420A (zh) | 应用程序的生成方法、装置、设备以及存储介质 | |
| JP7819117B2 (ja) | 画像特殊効果の設定方法、画像識別方法、装置および電子機器 | |
| CN119692303B (zh) | 页面图表调整方法、装置、电子设备及存储介质 | |
| US20250094139A1 (en) | Method of generating code based on large model, electronic device, and storage medium | |
| CN119106216A (zh) | 信息交互方法、装置、电子设备和存储介质 | |
| WO2025092911A1 (zh) | 视频特征提取方法、视频生成方法、装置、介质及设备 | |
| CN115237248B (zh) | 虚拟对象的展示方法、装置、设备、存储介质及程序产品 | |
| CN118632045A (zh) | 视频生成方法、装置、电子设备以及存储介质 | |
| CN118400563A (zh) | 音频存储方法、装置、电子设备和计算机可读介质 | |
| CN120034693A (zh) | 视频生成方法、装置、电子设备和存储介质 | |
| WO2025139712A1 (zh) | 视频处理方法及相关设备 | |
| CN113157360B (zh) | 用于处理api的方法、装置、设备、介质和产品 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24822603 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |