WO2025201137A1 - 一种视频生成方法、装置、设备及存储介质 - Google Patents

一种视频生成方法、装置、设备及存储介质

Info

Publication number
WO2025201137A1
WO2025201137A1 PCT/CN2025/083442 CN2025083442W WO2025201137A1 WO 2025201137 A1 WO2025201137 A1 WO 2025201137A1 CN 2025083442 W CN2025083442 W CN 2025083442W WO 2025201137 A1 WO2025201137 A1 WO 2025201137A1
Authority
WO
WIPO (PCT)
Prior art keywords
video
description content
multimedia
content
draft
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/083442
Other languages
English (en)
French (fr)
Inventor
林宜静
李娉
成舒茵
钟沛汛
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Publication of WO2025201137A1 publication Critical patent/WO2025201137A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/472End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
    • H04N21/47205End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content for manipulating displayed content, e.g. interacting with MPEG-4 objects, editing locally
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/431Generation of visual interfaces for content selection or interaction; Content or additional data rendering
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/431Generation of visual interfaces for content selection or interaction; Content or additional data rendering
    • H04N21/4312Generation of visual interfaces for content selection or interaction; Content or additional data rendering involving specific graphical features, e.g. screen layout, special fonts or colors, blinking icons, highlights or animations
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • H04N21/44016Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving splicing one content stream with another content stream, e.g. for substituting a video clip
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/472End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/84Generation or processing of descriptive data, e.g. content descriptors

Definitions

  • video generation related functions have become more diverse, for example, using video templates to generate videos.
  • the present disclosure provides a video generation method, the method comprising:
  • the video segments corresponding to the multiple description content segments are spliced together to obtain a video draft.
  • the video copy is split into multiple description content segments; wherein the multiple description content segments have a preset sequence relationship.
  • the text description content in the target text editing box is obtained; wherein the text description content includes a first text content extracted from the target audio and video resource, and/or a second text content input based on the target text editing box.
  • the method before splicing the video segments corresponding to the plurality of description content segments based on the preset sequence relationship to obtain a video draft, the method further includes:
  • a result video corresponding to the video draft is generated based on the video editing information.
  • the present disclosure provides a video generation device, the device comprising:
  • a first acquisition module is configured to acquire multimedia description content, wherein the multimedia description content includes text description content and/or voice description content;
  • a first generating module configured to generate a video draft based on the multimedia description content and the multimedia material input by the user
  • the present disclosure provides a computer program product, which includes a computer program/instructions, and the computer program/instructions implement the above method when executed by a processor.
  • FIG1 is a flow chart of a video generation method provided by an embodiment of the present disclosure.
  • FIG4 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure.
  • FIG5 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure.
  • FIG7 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure.
  • FIG8 is a schematic diagram of another video generation page provided by an embodiment of the present disclosure.
  • FIG12 is a schematic structural diagram of a video generating device provided by an embodiment of the present disclosure.
  • FIG13 is a schematic structural diagram of a video generating device provided by an embodiment of the present disclosure.
  • the disclosed embodiments extract corresponding video segments from multimedia materials input by the user by determining description content segments based on the multimedia description content, and then generate a video draft composed of the extracted video segments based on a preset sequence relationship between the description content segments.
  • the disclosed embodiments support generating a video draft by first determining description content segments and then extracting video segments from multimedia materials input by the user based on the description content segments, enriching the video generation method and thus meeting the diverse video generation function requirements of users.
  • an embodiment of the present disclosure provides a video generation method.
  • FIG1 is a flow chart of a video generation method provided by an embodiment of the present disclosure, the method specifically includes:
  • the multimedia description content includes text description content and/or voice description content.
  • the video generation method provided by the embodiments of the present disclosure may be applied to a client.
  • the client may include a client deployed on a smart phone, a client deployed on a tablet computer, and the like.
  • the multimedia description content may include received text description content and may also include received voice description content.
  • the multimedia description content is not limited to any specific format and may be in voice, text, or other formats.
  • the multimedia description content typically includes a description of the main object in the draft video to be generated. For example, if the draft video to be generated is a marketing video for a specific target, the multimedia description content may include description content for that target. In fact, the specific description content included in the multimedia description content is not limited.
  • the multimedia description content can be obtained based on the target text editing box.
  • the text description content in the target text editing box is obtained.
  • the text description content may include the first text content extracted based on the target audio and video resources.
  • the target audio and video resources can be audio resources, video resources, etc. selected from the user's album, or audio resources, video resources, etc. captured based on the shooting page.
  • the first text content extracted from the target audio and video resources can specifically include: first extracting the voice data extracted from the target audio and video resources, and then performing voice recognition on the extracted voice data to identify the corresponding text content, and displaying the identified text content in the target text editing box.
  • the method for obtaining the target audio resource can include: the user clicks the extraction control set on the page where the target text editing box is located to trigger the extraction of the text content corresponding to the voice data in the target audio and video resource.
  • the text content can be obtained by performing voice recognition on the target audio and video resource.
  • the embodiment of the present disclosure does not impose any restrictions on the specific implementation method of extracting the first text content from the target audio and video resource.
  • FIG. 2 a schematic diagram of a video generation page provided in an embodiment of the present disclosure is shown.
  • An extraction control 201 is displayed on the video generation page for obtaining target audio and video resources.
  • FIG 3 a schematic diagram of a user album page provided in an embodiment of the present disclosure is shown.
  • the user album page is displayed, and a plurality of optional audio and video resources are displayed on the user album page.
  • the user determines the target audio and video resource by triggering a selection operation for any one or more audio and video resources.
  • FIG 4 another schematic diagram of a user album page provided in an embodiment of the present disclosure is shown.
  • the user album page displays audio and video resources 401 in a selected state, i.e., the target audio and video resource.
  • the extraction control 402 displayed on the user album page the first text content is triggered to be extracted from the target audio resource.
  • the extracted first text content can be displayed in the target text edit box, as shown in Figure 5, which is a schematic diagram of another video generation page provided by this public embodiment.
  • the first text content is displayed in the target text edit box 501 displayed on the video generation page.
  • the acquired multimedia description content may include second text content entered by the user in the target text edit box.
  • the second text content manually entered by the user in the target text edit box is used to form the multimedia description content.
  • the second text content may include any text content, including short sentences, long sentences, paragraphs, keywords, and the like.
  • the obtained multimedia description content may include not only the first text content extracted from the target audio and video resources, but also the second text content entered by the user in the target text edit box.
  • the order in which the first and second text contents are displayed in the target text edit box can be determined based on demand.
  • the second text content manually entered by the user in the target text editing box is obtained. Then, based on the user triggering the extraction control, the target audio and video resources are determined, and the first text content is extracted from the target audio and video resources and displayed in the target text editing box.
  • the material upload control 204 in the video generation page As shown in FIG. 2 above, the material upload control 204 in the video generation page.
  • the multiple video segments may be highlight segments in the multimedia material
  • the multimedia material may be compressed, and then the compressed multimedia material may be subjected to image recognition by the server to determine highlight segments with higher picture quality, and then the highlight segments in the compressed multimedia material may be replaced with corresponding segments in the multimedia material input by the user, and the highlight segments may be cropped out.
  • the editing trigger operation for the video draft may include displaying an editing control on the video preview page. Specifically, first, a selected operation for at least one video draft is received. When the user triggers the editing control, a video editing page is displayed, and each video clip in the video draft is displayed on the video editing page. The video editing operation and multiple video clips in the selected video draft are displayed on the video editing page. When video editing information for the video draft is received, the video draft is edited. Specifically, the editing operation can be performed on each video clip in the video draft.
  • an editing page with lightweight editing operations will be displayed first, allowing the user to perform preliminary editing operations on the video draft. Then the user can trigger the editing controls on the video editing page to display a video editing page with multiple tracks to further edit the video draft to meet user needs.
  • the export control can be displayed. Specifically, the export control can be displayed on the video editing page. In an optional implementation, when a trigger operation for the export control is received, a result video corresponding to the video draft is generated based on the video editing information.
  • the present disclosure further provides a video generation device.
  • a video generation device Referring to FIG12 , which is a schematic structural diagram of a video generation device provided in an embodiment of the present disclosure, the device includes:
  • a first acquisition module 1201 is configured to acquire multimedia description content, wherein the multimedia description content includes text description content and/or voice description content;
  • the second acquisition module 1202 is used to acquire multimedia materials input by the user;
  • a first generating module 1203 is configured to generate a video draft based on the multimedia description content and the multimedia material input by the user;
  • the video draft includes multiple video clips, and there is a corresponding relationship between the multiple video clips and multiple description content clips.
  • the multiple description content clips are determined based on the multimedia description content, and the multiple video clips are extracted from the multimedia materials input by the user for the multiple description content clips respectively.
  • the display order of the multiple video clips in the video draft is determined based on the preset sequence relationship between the corresponding description content clips.
  • the first generating module includes:
  • a first determining submodule is configured to determine a plurality of description content segments based on the multimedia description content; wherein the plurality of description content segments have a preset sequence relationship;
  • the first splicing submodule is used to splice the video segments corresponding to the multiple description content segments based on the preset sequence relationship to obtain a video draft.
  • the first determining submodule includes:
  • a first generating submodule configured to generate a video text based on the multimedia content
  • the first acquisition module is specifically configured to:
  • the text description content in the target text editing box is obtained; wherein the text description content includes a first text content extracted from the target audio and video resource, and/or a second text content input based on the target text editing box.
  • the first acquisition module includes:
  • a first receiving submodule is configured to receive text content inputted for at least one target attribute of a target object
  • the first receiving submodule includes:
  • a first display submodule is configured to, in response to a third text content inputted for a first target attribute of a target object, display at least one candidate recommendation content corresponding to a second target attribute of the target object; wherein the at least one candidate recommendation content is determined based on the third text content;
  • a second receiving submodule configured to receive a fourth text content selected from the at least one candidate recommended content, and/or a fifth text content input for the second target attribute
  • the first acquisition submodule is specifically configured to:
  • the device further includes:
  • a second generating module is configured to generate voice segments for corresponding video segments based on the plurality of description content segments; wherein the voice segments and the video segments have the same play time information;
  • the first splicing submodule is specifically used to:
  • the video segments with the voice segments are spliced together to obtain a video draft.
  • the device further includes:
  • the third generation module is used to generate at least one candidate video draft in response to the video generation operation performed on the video preview page, based on the multimedia description content and the multimedia material input by the user, and display the at least one candidate video draft on the video preview page.
  • the device further includes:
  • a second display module configured to display the video draft on a video editing page in response to an editing trigger operation on the video draft
  • a first receiving module is configured to receive video editing information for the video draft
  • the fourth generating module is used to generate a result video corresponding to the video draft based on the video editing information in response to the export operation on the video draft.
  • the embodiments of the present disclosure further provide a computer-readable storage medium, which stores instructions.
  • the terminal device implements the video generation method described in the embodiments of the present disclosure.
  • the embodiment of the present disclosure further provides a video generation device, as shown in FIG13 , which may include:
  • the video generation device may include one or more processors 1301, with one processor being used as an example in FIG13 .
  • processor 1301, memory 1302, input device 1303, and output device 1304 may be connected via a bus or other means, with FIG13 using a bus as an example.
  • the processor 1301 will load the executable files corresponding to the processes of one or more applications into the memory 1302 according to the following instructions, and the processor 1301 will run the applications stored in the memory 1302, thereby realizing the various functions of the above-mentioned video generation device.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Databases & Information Systems (AREA)
  • Human Computer Interaction (AREA)
  • Television Signal Processing For Recording (AREA)

Abstract

本公开提供了一种视频生成方法、装置、设备及存储介质,所述方法包括:获取多媒体描述内容,获取用户输入的多媒体素材,基于多媒体描述内容和用户输入的多媒体素材,生成视频草稿,视频草稿中包括多个视频片段,多个视频片段与多个描述内容片段具有对应关系,多个描述内容片段基于多媒体描述内容确定。即通过基于多媒体描述内容确定的描述内容片段,从用户输入的多媒体素材中提取视频片段,进而基于描述内容片段之间的预设顺序关系生成由提取到的视频片段组成的视频草稿。

Description

一种视频生成方法、装置、设备及存储介质
相关申请的交叉引用
本申请要求于2024年3月28日提交的,申请号为202410370301.7、发明名称为“一种视频生成方法、装置、设备及存储介质”的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
技术领域
本公开涉及数据处理领域,尤其涉及一种视频生成方法、装置、设备及存储介质。
背景技术
随着视频生成技术的不断发展,视频生成相关功能也更加多样化。例如,利用视频模板生成视频等。
发明内容
为了解决上述技术问题,本公开实施例提供了一种视频生成方法、装置、设备及存储介质。
第一方面,本公开提供了一种视频生成方法,所述方法包括:
获取多媒体描述内容;其中,所述多媒体描述内容包括文本描述内容和/或语音描述内容;
以及,获取用户输入的多媒体素材;
基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿;
其中,所述视频草稿中包括多个视频片段,所述多个视频片段与多个描述内容片段之间具有对应关系,所述多个描述内容片段为基于所述多媒体描述内容确定,所述多个视频片段为从所述用户输入的多媒体素材中为所述多个描述内容片段分别提取到的,所述多个视频片段在所述视频草稿中的展示顺序为基于对应的描述内容片段之间的预设顺序关系确定。
一种可选的实施方式中,所述基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿,包括:
基于所述多媒体描述内容确定多个描述内容片段;其中,所述多个描述内容片段之间具有预设顺序关系;
从所述用户输入的多媒体素材中,为所述多个描述内容片段分别裁剪出对应的视频片段;
基于所述预设顺序关系,对所述多个描述内容片段分别对应的视频片段进行拼接处理,得到视频草稿。
所述基于所述多媒体描述内容确定多个描述内容片段,包括:
根据所述多媒体内容生成视频文案;
将所述视频文案拆分成多个描述内容片段;其中,所述多个描述内容片段之间具有预设顺序关系。
一种可选的实施方式中,所述获取多媒体描述内容,包括:
获取目标文本编辑框内的文本描述内容;其中,所述文本描述内容包括从目标音视频资源中提取到的第一文本内容,和/或,基于所述目标文本编辑框输入的第二文本内容。
一种可选的实施方式中,所述获取多媒体描述内容,包括:
接收针对目标对象的至少一个目标属性输入的文本内容;
基于所述至少一个目标属性分别对应的文本内容,获取多媒体描述内容。
一种可选的实施方式中,所述接收针对目标对象的至少一个目标属性输入的文本内容,包括:
响应于针对目标对象的第一目标属性输入的第三文本内容,显示所述目标对象的第二目标属性对应的至少一个候选推荐内容;其中,所述至少一个候选推荐内容为基于所述第三文本内容确定;
接收从所述至少一个候选推荐内容中选定的第四文本内容,和/或针对所述第二目标属性输入的第五文本内容;
相应的,所述基于所述至少一个目标属性分别对应的文本内容,获取多媒体描述内容,包括:
基于所述第三文本内容、所述第四文本内容和所述第五文本内容中的至少一种,获取多媒体描述内容。
一种可选的实施方式中,所述基于所述预设顺序关系,对所述多个描述内容片段分别对应的视频片段进行拼接处理,得到视频草稿之前,还包括:
基于所述多个描述内容片段分别为对应的视频片段生成语音片段;其中,所述语音片段与所述视频片段具有相同的播放时间信息;
相应的,所述基于所述预设顺序关系,对所述多个描述内容片段分别对应的视频片段进行拼接处理,得到视频草稿,包括:
基于所述预设顺序关系,对具有所述语音片段的视频片段进行拼接处理,得到视频草稿。
一种可选的实施方式中,所述基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿之后,还包括:
在视频预览页面上展示所述视频草稿;
响应于作用在所述视频预览页面上的视频生成操作,基于所述多媒体描述内容和所述用户输入的多媒体素材,生成至少一个候选视频草稿,并在所述视频预览页面上展示所述至少一个候选视频草稿。
一种可选的实施方式中,所述基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿之后,还包括:
响应于针对所述视频草稿的编辑触发操作,在视频编辑页面上展示所述视频草稿;
接收针对所述视频草稿的视频编辑信息;
响应于针对所述视频草稿的导出操作,基于所述视频编辑信息生成所述视频草稿对应的结果视频。
第二方面,本公开提供了一种视频生成装置,所述装置包括:
第一获取模块,用于获取多媒体描述内容;其中,所述多媒体描述内容包括文本描述内容和/或语音描述内容;
第二获取模块,用于获取用户输入的多媒体素材;
第一生成模块,用于基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿;
其中,所述视频草稿中包括多个视频片段,所述多个视频片段与多个描述内容片段之间具有对应关系,所述多个描述内容片段为基于所述多媒体描述内容确定,所述多个视频片段为从所述用户输入的多媒体素材中为所述多个描述内容片段分别提取到的,所述多个视频片段在所述视频草稿中的展示顺序为基于对应的描述内容片段之间的预设顺序关系确定。
第三方面,本公开提供了一种计算机可读存储介质,所述计算机可读存储介质中存储有指令,当所述指令在终端设备上运行时,使得所述终端设备实现上述的方法。
第四方面,本公开提供了一种视频生成设备,包括:存储器,处理器,及存储在所述存储器上并可在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序时,实现上述的方法。
第五方面,本公开提供了一种计算机程序产品,所述计算机程序产品包括计算机程序/指令,所述计算机程序/指令被处理器执行时实现上述的方法。
附图说明
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本公开的实施例,并与说明书一起用于解释本公开的原理。
为了更清楚地说明本公开实施例或现有技术中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,对于本领域普通技术人员而言,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本公开实施例提供的一种视频生成方法的流程图;
图2为本公开实施例提供的一种视频生成页面的示意图;
图3为本公开实施例提供的一种用户相册页面的示意图;
图4为本公开实施例提供的另一种视频生成页面的示意图;
图5为本公开实施例提供的另一种视频生成页面的示意图;
图6为本公开实施例提供的另一种视频生成页面的示意图;
图7为本公开实施例提供的另一种视频生成页面的示意图;
图8为本公开实施例提供的另一种视频生成页面的示意图;
图9为本公开实施例提供的一种视频预览页面的示意图;
图10为本公开实施例提供的一种视频编辑页面的示意图;
图11为本公开实施例提供的另一种视频编辑页面的示意图;
图12为本公开实施例提供的一种视频生成装置的结构示意图;
图13为本公开实施例提供的一种视频生成设备的结构示意图。
具体实施方式
为了能够更清楚地理解本公开的上述目的、特征和优点,下面将对本公开的方案进行进一步描述。需要说明的是,在不冲突的情况下,本公开的实施例及实施例中的特征可以相互组合。
在下面的描述中阐述了很多具体细节以便于充分理解本公开,但本公开还可以采用其他不同于在此描述的方式来实施;显然,说明书中的实施例只是本公开的一部分实施例,而不是全部的实施例。
随着视频生成技术的不断发展,视频生成相关功能也更加多样化。例如,利用视频模板生成视频等。
为了满足人们多样化的视频生成功能需求,如何进一步丰富视频生成方式,是目前亟需解决的技术问题。
为此,本公开实施例提供了一种视频生成方法,获取多媒体描述内容,其中,多媒体描述内容包括文本描述内容和/或语音描述内容,以及,获取用户输入的多媒体素材,基于多媒体描述内容和用户输入的多媒体素材,生成视频草稿,其中,视频草稿中包括多个视频片段,多个视频片段与多个描述内容片段之间具有对应关系,多个描述内容片段为基于多媒体描述内容确定,多个视频片段为从用户输入的多媒体素材中为多个描述内容片段分别提取到的,多个视频片段在视频草稿中的展示顺序为基于对应的描述内容片段之间的预设顺序关系确定。
本公开实施例通过基于多媒体描述内容确定的描述内容片段,从用户输入的多媒体素材中提取对应的视频片段,进而基于描述内容片段之间的预设顺序关系生成由提取到的视频片段组成的视频草稿。可见,本公开实施例支持先确定描述内容片段,然后基于描述内容片段从用户输入的多媒体素材提取视频片段的方式生成视频草稿,丰富了视频生成方式,从而满足了用户多样化的视频生成功能需求。
基于此,本公开实施例提供了一种视频生成方法,参考图1,为本公开实施例提供的一种视频生成方法的流程图,该方法具体包括:
S101:获取多媒体描述内容。
其中,所述多媒体描述内容包括文本描述内容和/或语音描述内容。
本公开实施例提供的视频生成方法,可以应用于客户端,例如,客户端可以包括部署于智能手机的客户端、部署于平板电脑的客户端等。
本公开实施例中,多媒体描述内容可以包括接收到的文本描述内容,还可以包括接收到的语音描述内容,也就是说,多媒体描述内容的内容形式不做限制,可以为语音形式、文本形式等。多媒体描述内容通常包括用于描述待生成的视频草稿中的主体描述对象,例如,待生成的视频草稿为针对某个对象的营销视频,相应的,多媒体描述内容可以包括针对该对象的描述内容等。事实上,多媒体描述内容中包括的具体描述内容不受限制。
实际应用中,获取多媒体描述内容的实现方式较多,一种可选的实施方式中,可以基于目标文本编辑框获取多媒体描述内容,具体的,获取目标文本编辑框内的文本描述内容,该文本描述内容可以包括基于目标音视频资源提取到的第一文本内容,该目标音视频资源可以为从用户相册中选定的音频资源、视频资源等,或者是基于拍摄页面拍摄得到的音频资源、视频资源等。
其中,从目标音视频资源中提取到的第一文本内容,具体可以包括:首先提取目标音视频资源中提取到的语音数据,然后将该提取到的语音数据进行语音识别,识别得到对应的文本内容,并将识别到的文本内容显示在目标文本编辑框内。对于获取目标音频资源的方式可以包括:用户点击设置在目标文本编辑框所处页面上的提取控件,以触发提取目标音视频资源中的语音数据对应的文本内容。具体的,可以通过对目标音视频资源进行语音识别的方式得到文本内容。对于从目标音视频资源中提取第一文本内容的具体实施方式,本公开实施例不做任何限制。
如图2所示,为本公开实施例提供的一种视频生成页面的示意图。该视频生成页面上显示有提取控件201,用于获取目标音视频资源。如图3所示,为本公开实施例提供的一种用户相册页面的示意图。当用户触发提取控件201时,显示用户相册页面,该用户相册页面上显示有多个可选的音视频资源,用户通过针对任意一个或多个音视频资源触发选定操作,确定目标音视频资源。如图4所示,为本公开实施例提供的另一种用户相册页面的示意图。该用户相册页面上显示有处于选定状态的音视频资源401,即目标音视频资源,通过触发用户相册页面上显示的提取控件402,触发从目标音频资源中提取第一文本内容。
实际应用中,可以将提取到的第一文本内容显示在目标文本编辑框内,如图5所示,为本公共实施例提供的另一种视频生成页面的示意图。该视频生成页面上显示的目标文本编辑框501内显示有第一文本内容。
另一种可选的实施方式中,获取的多媒体描述内容,可以包括用户在目标文本编辑框内输入的第二文本内容,即用户在目标文本编辑框内手动输入的第二文本内容用于构成多媒体描述内容。第二文本内容可以包括任意的文本内容,具体形式可以包括:短句、长句、段落、关键词等。
如上述图2所示的视频生成页面中,显示有目标文本编辑框202,当用户触发目标文本编辑框时,显示上述图5所示的视频生成页面,用户可以在目标文本编辑框501内手动输入第二文本内容。
又一种可选的实施方式中,获取的多媒体描述内容不仅可以包括从目标音视频资源中提取到的第一文本内容,还可以同时包括用户在目标文本编辑框内输入的第二文本内容。针对第一文本内容和第二文本内容在目标文本编辑框内的显示顺序,可以基于需求确定。
一种应用场景中,首先基于用户触发目标文本编辑框,获取用户在目标文本编辑框内手动输入第二文本内容,然后,基于用户触发提取控件,确定目标音视频资源,从目标音视频资源中提取第一文本内容显示在目标文本编辑框内。
另一种应用场景中,可以首先基于用户触发提取控件,确定目标音视频资源,从目标音视频资源中提取第一文本内容,然后在目标文本编辑框内显示第一文本内容的基础上,获取用户在目标文本编辑框内手动输入第二文本内容。
实际应用中,对于第一文本内容和第二文本内容在目标文本编辑框内的显示位置可以基于当前定位的位置,具体例如,鼠标光标所在位置处。
在上述实施例的基础上,获取目标文本编辑框内的文本描述内容包括从目标音视频提取,或者基于用户手动输入,为此,本公开实施例,还可以在视频生成页面上显示输入或提取文本内容入口,以便于用户了解获取文本描述内容的方式。
如上述图2所示的视频生成页面中显示有输入或提取文本内容入口203。
实际应用中,为提高多媒体描述内容的获取效率,进一步提升用户体验,本公开实施例还支持以智能推荐的方式获取多媒体描述内容。
具体的,首先接收针对目标对象的至少一个目标属性输入的文本内容,基于该至少一个目标属性分别对应的文本内容,获取多媒体描述内容。其中,目标对象可以为生成的视频中的中心主体,目标属性可以为该中心主体的任一属性,具体例如名称、特点、优势、价格、适用人群、优惠活动、视频时长等等。
对于接收针对目标对象的至少一个目标属性输入的文本内容,可以包括当接收到针对目标对象的第一目标属性输入的第三文本内容时,基于第三文本内容确定出至少一个候选推荐内容,该至少一个候选推荐内容对应于上述目标对象的第二目标属性。
也就是说,根据接收到用户手动输入目标对象的第一目标属性的第三文本内容,自动显示针对该目标对象的第二属性的候选推荐内容。
其中,第一目标属性和第二目标属性为目标属性中的任意两个不同的属性。第三文本内容可以是任意文本内容,具体例如成语,关键词等。候选推荐内容可以是任意推荐文本内容,具体例如成语,词语等。
实际应用中,在以智能推荐的方式获取多媒体描述内容之前,还可以显示候选推荐对象内容标识,以向用户提示智能生成候选推荐内容功能。具体的,可以在第二属性对应的输入框内显示智能显示候选推荐对象内容标识。
如图6所示,为本公开实施例提供的一种视频生成页面的示意图。该视频生成页面上显示第一目标属性601,当接收到针对第一目标属性的输入的第三文本内容时,在第二属性602对应的输入框内显示候选推荐对象内容标识603,用于提示当前正在智能生成第二属性的候选推荐内容。在显示至少一个候选推荐内容时,候选推荐对象内容标识隐藏。
如图7所示,为本公开实施例提供的另一种视频生成页面的示意图。该视频生成页面上显示有目标对象的属性,其中包含第一目标属性701,当接收到用户在输入框702内输入第一属性的第三文本内容时,显示第二属性703的至少一个候选推荐内容,其中包含候选推荐内容704。
在基于第三文本内容显示至少一个候选推荐内容的基础上,可以将接收到第四文本内容,和/或第五文本内容作为第二属性的文本内容,然后根据第三文本内容、第四文本内容和第五文本内容中的至少一个,获取多媒体内容。
其中,第四文本内容可以是用户从至少一个候选推荐内容中选定的至少一个候选推荐内容,第五文本内容可以是用户手动输入的第二属性的文本内容。如图8所示,为本公开实施例提供的另一种视频生成页面的示意图。当接收到针对上述图7所示的视频生成页面中候选推荐内容704作为第四文本内容的选定操作,第二目标属性对应的输入框内显示第四文本内容,如果接收到针对第二目标属性对应的第五文本内容,则在第二属性对应的输入框内显示第五文本内容。
实际应用中,为向用户提示智能获取功能,还可以显示智能获取入口,如上述图6所示的视频生成页面中显示有智能获取入口604。
S102:获取用户输入的多媒体素材。
本公开实施例中,多媒体素材为用户上传的一个或多个多媒体素材,多媒体素材可以包括音频媒体素材、图片媒体素材。
对于获取用户输入多媒体素材的方式,可以包括用户触发素材上传控件的方式获取用户输入的多媒体素材。
如上述图2所示视频生成页面中的素材上传控件204。
S103:基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿。
其中,所述视频草稿中包括多个视频片段,所述多个视频片段与多个描述内容片段之间具有对应关系,所述多个描述内容片段为基于所述多媒体描述内容确定,所述多个视频片段为从所述用户输入的多媒体素材中为所述多个描述内容片段分别提取到的,所述多个视频片段在所述视频草稿中的展示顺序为基于对应的描述内容片段之间的预设顺序关系确定。
本公开实施例中,视频草稿可以为未编辑的视频初稿,为生成满足用户需求的视频草稿,可以生成多个视频草稿,以供用户选择,具体例如生成5个视频草稿。
由于视频草稿是由多个视频片段拼接而成的,所以,在生成的多个视频草稿中,每个视频草稿中的视频片段的可以相同可以不同,每个视频草稿所包含的视频片段的数量可能相同可能不相同,每个视频草稿的视频时长可能有所不同。
对于生成视频草稿的触发操作,可以包括针对生成视频控件的触发操作。具体的,可以显示视频生成控件。如上述图2所示视频生成页面上显示的视频生成控件205。
一种可选的实施方式中,基于多媒体描述内容确定多个描述内容片段,其中,多个描述内容片段之间具有预设顺序关系,根据用户输入的多媒体素材,从多媒体素材中为多个描述内容片段分别裁剪出对应的视频片段,基于预设顺序关系,对多个描述内容片段分别对应的视频片段进行拼接处理,得到视频草稿。
其中,描述内容片段可以指描述文本片段,基于多媒体描述内容确定多个描述内容片段,可以包括,将获取到的多媒体描述内容直接确定为描述内容片段,多个描述内容片段之间的预设顺序关系可以为多媒体描述内容之间的预设顺序关系,预设顺序关系具体例如段落关系,上下文关系等。
基于多媒体描述内容确定多个描述内容片段,还可以包括,基于获取到的多媒体内容,智能生成视频文案,然后将视频文案拆分成多个具有预设顺序关系的描述内容片段。具体的,对多媒体描述内容进行内容扩展处理,得到扩展后描述内容,将拓展后描述内容作为视频文案,然后将该视频文案拆分成多个具有预设顺序关系的描述内容片段,以提升视频草稿的质量。对于拓展处理多媒体描述内容,可以利用模型进行拓展,具体实施方式本公开实施例不做任何限制。
对于从用户输入的多媒体素材中裁剪出对应于描述内容片段的视频片段,其中,该多个视频片段可以是多媒体素材中的高光片段,具体的,可以将多媒体素材压缩,然后利用服务器将压缩的多媒体素材进行图像识别,确定画面质量较高的高光片段,然后将压缩的多媒体素材中的高光片段替换成用户输入的多媒体素材中的相应片段,并裁剪出高光片段。
然后,基于多个描述内容片段之间的预设顺序关系,对多个描述内容片段分别对应的视频片段进行拼接处理,得到视频草稿。
实际应用中,在对多个描述内容片段分别对应的视频片段拼接得到视频草稿之前,还需要为多个描述内容片段分别对应的视频片段生成语音片段。具体的,可以将多个描述内容片段利用人声朗读生成多个视频片段对应的语音片段,该语音片段与视频片段具有相同的播放时间信息,即语音片段与视频片段保持同步。
对于语音片段与视频片段具有相同的播放时间信息的方式,可以包括,调整人声朗读生成的语音片段的播放时间,具体的,根据视频片段的播放时间,放慢或者加快人声朗读的速度。
本公开实施例提供的视频生成方法中,获取多媒体描述内容,其中,多媒体描述内容包括文本描述内容和/或语音描述内容,以及,获取用户输入的多媒体素材,基于多媒体描述内容和用户输入的多媒体素材,生成视频草稿,其中,视频草稿中包括多个视频片段,多个视频片段与多个描述内容片段之间具有对应关系,多个描述内容片段为基于多媒体描述内容确定,多个视频片段为从用户输入的多媒体素材中为多个描述内容片段分别提取到的,多个视频片段在视频草稿中的展示顺序为基于对应的描述内容片段之间的预设顺序关系确定。
本公开实施例通过基于多媒体描述内容确定的描述内容片段,从用户输入的多媒体素材中提取对应的视频片段,进而基于描述内容片段之间的预设顺序关系生成由提取到的视频片段组成的视频草稿。可见,本公开实施例支持先确定描述内容片段,然后基于描述内容片段从用户输入的多媒体素材提取视频片段的方式生成视频草稿,丰富了视频生成方式,从而满足了用户多样化的视频生成功能需求。
实际应用中,在生成视频草稿之后,还支持针对视频草稿进行预览,以及在视频预览页面上支持针对多媒体素材描述内容和用户输入的多媒体素材生成更多候选视频草稿,以满足用户需求,进一步提升用户体验。
具体的,在视频预览页面上展示至少一个视频草稿,当接收到作用在视频预览页面上的视频生成操作时,基于获取到的多媒体描述内容以及用户输入的多媒体素材,生成至少一个候选视频草稿,并在该视频预览页面上展示候选视频草稿。
其中,在视频预览页面上展示的至少一个视频草稿,可以根据画面相关性,符合当下热点等排序策略展示在视频预览页面上。对于在视频预览页面上的视频生成操作,可以包括在视频预览页面上设置生成控件,以生成候选视频草稿,并在该视频预览页面上展示候选视频草稿。
对于在视频预设页面上展示候选视频草稿,可以通过在视频预设页面上的预设滑动操作展示候选视频草稿。
如图9所示,为本公开实施例提供的一种视频预览页面的示意图。该视频预览页面上显示有至少一个视频草稿,其中包含视频草稿901,以及生成控件902,当触发生成控件902时,基于获取到的多媒体描述内容以及用户输入的多媒体素材,生成至少一个候选视频草稿,并在该视频预览页面上展示候选视频草稿。
实际应用中,在视频预览页面上展示视频草稿的过程中,还支持针对至少一个视频草稿的编辑操作以及导出操作,以便于用户编辑和发布。具体的,当接收到针对至少一个视频草稿的编辑触发操作时,在视频编辑页面上展示该视频草稿,接收针对该视频草稿的视频编辑信息,当接收到针对视频草稿的导出操作时,基于视频编辑信息生成视频草稿对应的结果视频。
其中,针对视频草稿的编辑触发操作,可以包括在视频预览页面上显示编辑控件,具体的,首先接收到针对至少一个视频草稿中的选定操作,当用户触发编辑控件时,显示视频编辑页面,在视频编辑页面上展示视频草稿中的各个视频片段,视频编辑页面上显示有视频编辑操作,以及选定的视频草稿中的多个视频片段,当接收到针对视频草稿的视频编辑信息时,对视频草稿进行编辑操作。具体的,可以对视频草稿中的每个视频片段分别执行编辑操作。
如上述图9所示的视频预览页面中,当选中视频草稿901,以及触发编辑控件903时,显示视频编辑页面,在视频编辑页面上展示视频草稿中的各个视频片段,以及各种视频编辑操作。
一种可选的实施方式中,当用户选中视频草稿,触发视频预览页面上的编辑控件时,显示视频编辑页面,视频编辑页面上显示有视频草稿的多个视频片段,用户可以选定其中一个视频片段,作为待编辑视频片段,当接收到针对待编辑视频片段的视频编辑信息时,对该待编辑视频片段进行视频编辑操作。该视频编辑操作可以包括文字编辑操作,脚本编辑操作,以及音乐编辑操作等轻量编辑操作,该视频编辑编辑页面具体例如轻量编辑页面。
如图10所示,为本公开实施例提供的一种视频编辑页面的示意图。该视频编辑页面上显示有视频草稿的多个视频片段,以及文字编辑操作,脚本编辑操作,以及音乐编辑操作等轻量编辑操作。
另一种可选的实施方式中,在用户触发视频预览页面上编辑控件时,显示视频编辑页面,该视频编辑页面上显示有视频轨道,音频轨道等,该视频编辑页面具体例如多轨道视频编辑页面。通过该多轨道视频编辑页面,可便于对视频草稿进行更多编辑操作,满足用户需求。
如图11所示,为本公开实施例提供的另一种视频编辑页面的示意图。该视频编辑页面上显示有视频轨道,音频轨道等。
实际应用中,还可以在用户选定视频草稿触发视频预览页面上的编辑控件时,首先显示具有轻量编辑操作的编辑页面,便于用户对视频草稿进行初步编辑操作,然后用户可以触发视频编辑页面上的编辑控件,显示具有多轨道的视频编辑页面,以进一步对视频草稿进行编辑,满足用户需求。
对于针对视频草稿的导出操作,可以显示导出控件,具体的,可以在视频编辑页面上显示导出控件,一种可选的实施方式中,当接收到针对导出控件的触发操作时,基于视频编辑信息生成视频草稿对应的结果视频。
实际应用中,还可以在视频预览页面上显示导出控件。一种可选的实施方式中,可以在视频预览页面上展示至少一个视频草稿的过程中,选中目标视频草稿,当针对该目标视频操作触发导出控件时,生成目标视频草稿对应的结果视频。如上述图9所示的视频预览页面中显示的导出控件904所示。
基于上述方法实施例,本公开还提供了一种视频生成装置,参考图12,为本公开实施例提供的一种视频生成装置的结构示意图,所述装置包括:
第一获取模块1201,用于获取多媒体描述内容;其中,所述多媒体描述内容包括文本描述内容和/或语音描述内容;
第二获取模块1202,用于获取用户输入的多媒体素材;
第一生成模块1203,用于基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿;
其中,所述视频草稿中包括多个视频片段,所述多个视频片段与多个描述内容片段之间具有对应关系,所述多个描述内容片段为基于所述多媒体描述内容确定,所述多个视频片段为从所述用户输入的多媒体素材中为所述多个描述内容片段分别提取到的,所述多个视频片段在所述视频草稿中的展示顺序为基于对应的描述内容片段之间的预设顺序关系确定。
一种可选的实施方式中,所述第一生成模块,包括:
第一确定子模块,用于基于所述多媒体描述内容确定多个描述内容片段;其中,所述多个描述内容片段之间具有预设顺序关系;
第一裁剪子模块,用于从所述用户输入的多媒体素材中,为所述多个描述内容片段分别裁剪出对应的视频片段;
第一拼接子模块,用于基于所述预设顺序关系,对所述多个描述内容片段分别对应的视频片段进行拼接处理,得到视频草稿。
一种可选的实施方式中,所述第一确定子模块,包括:
第一生成子模块,用于根据所述多媒体内容生成视频文案;
第一拆分子模块,用于将所述视频文案拆分成多个描述内容片段;其中,所述多个描述内容片段之间具有预设顺序关系。
一种可选的实施方式中,所述第一获取模块,具体用于:
获取目标文本编辑框内的文本描述内容;其中,所述文本描述内容包括从目标音视频资源中提取到的第一文本内容,和/或,基于所述目标文本编辑框输入的第二文本内容。
一种可选的实施方式中,所述第一获取模块,包括:
第一接收子模块,用于接收针对目标对象的至少一个目标属性输入的文本内容;
第一获取子模块,用于基于所述至少一个目标属性分别对应的文本内容,获取多媒体描述内容。
一种可选的实施方式中,所述第一接收子模块,包括:
第一显示子模块,用于响应于针对目标对象的第一目标属性输入的第三文本内容,显示所述目标对象的第二目标属性对应的至少一个候选推荐内容;其中,所述至少一个候选推荐内容为基于所述第三文本内容确定;
第二接收子模块,用于接收从所述至少一个候选推荐内容中选定的第四文本内容,和/或针对所述第二目标属性输入的第五文本内容;
相应的,所述第一获取子模块,具体用于:
基于所述第三文本内容、所述第四文本内容和所述第五文本内容中的至少一种,获取多媒体描述内容。
一种可选的实施方式中,所述装置还包括:
第二生成模块,用于基于所述多个描述内容片段分别为对应的视频片段生成语音片段;其中,所述语音片段与所述视频片段具有相同的播放时间信息;
相应的,所述第一拼接子模块,具体用于:
基于所述预设顺序关系,对具有所述语音片段的视频片段进行拼接处理,得到视频草稿。
一种可选的实施方式中,所述装置还包括:
第一展示模块,用于在视频预览页面上展示所述视频草稿;
第三生成模块,用于响应于作用在所述视频预览页面上的视频生成操作,基于所述多媒体描述内容和所述用户输入的多媒体素材,生成至少一个候选视频草稿,并在所述视频预览页面上展示所述至少一个候选视频草稿。
一种可选的实施方式中,所述装置还包括:
第二展示模块,用于响应于针对所述视频草稿的编辑触发操作,在视频编辑页面上展示所述视频草稿;
第一接收模块,用于接收针对所述视频草稿的视频编辑信息;
第四生成模块,用于响应于针对所述视频草稿的导出操作,基于所述视频编辑信息生成所述视频草稿对应的结果视频。
本公开实施例提供的视频生成方法中,获取多媒体描述内容,其中,多媒体描述内容包括文本描述内容和/或语音描述内容,以及,获取用户输入的多媒体素材,基于多媒体描述内容和用户输入的多媒体素材,生成视频草稿,其中,视频草稿中包括多个视频片段,多个视频片段与多个描述内容片段之间具有对应关系,多个描述内容片段为基于多媒体描述内容确定,多个视频片段为从用户输入的多媒体素材中为多个描述内容片段分别提取到的,多个视频片段在视频草稿中的展示顺序为基于对应的描述内容片段之间的预设顺序关系确定。
本公开实施例通过基于多媒体描述内容确定的描述内容片段,从用户输入的多媒体素材中提取对应的视频片段,进而基于描述内容片段之间的预设顺序关系生成由提取到的视频片段组成的视频草稿。可见,本公开实施例支持先确定描述内容片段,然后基于描述内容片段从用户输入的多媒体素材提取视频片段的方式生成视频草稿,丰富了视频生成方式,从而满足了用户多样化的视频生成功能需求。
除了上述方法和装置以外,本公开实施例还提供了一种计算机可读存储介质,计算机可读存储介质中存储有指令,当所述指令在终端设备上运行时,使得所述终端设备实现本公开实施例所述的视频生成方法。
本公开实施例还提供了一种计算机程序产品,所述计算机程序产品包括计算机程序/指令,所述计算机程序/指令被处理器执行时实现本公开实施例所述的视频生成方法。
另外,本公开实施例还提供了一种视频生成设备,参见图13所示,可以包括:
处理器1301、存储器1302、输入装置1303和输出装置1304。视频生成设备中的处理器1301的数量可以一个或多个,图13中以一个处理器为例。在本公开的一些实施例中,处理器1301、存储器1302、输入装置1303和输出装置1304可通过总线或其它方式连接,其中,图13中以通过总线连接为例。
存储器1302可用于存储软件程序以及模块,处理器1301通过运行存储在存储器1302的软件程序以及模块,从而执行视频生成设备的各种功能应用以及数据处理。存储器1302可主要包括存储程序区和存储数据区,其中,存储程序区可存储操作系统、至少一个功能所需的应用程序等。此外,存储器1302可以包括高速随机存取存储器,还可以包括非易失性存储器,例如至少一个磁盘存储器件、闪存器件、或其他易失性固态存储器件。输入装置1303可用于接收输入的数字或字符信息,以及产生与视频生成设备的用户设置以及功能控制有关的信号输入。
具体在本实施例中,处理器1301会按照如下的指令,将一个或一个以上的应用程序的进程对应的可执行文件加载到存储器1302中,并由处理器1301来运行存储在存储器1302中的应用程序,从而实现上述视频生成设备的各种功能。
需要说明的是,在本文中,诸如“第一”和“第二”等之类的关系术语仅仅用来将一个实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。而且,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、物品或者设备中还存在另外的相同要素。
以上所述仅是本公开的具体实施方式,使本领域技术人员能够理解或实现本公开。对这些实施例的多种修改对本领域的技术人员来说将是显而易见的,本文中所定义的一般原理可以在不脱离本公开的精神或范围的情况下,在其它实施例中实现。因此,本公开将不会被限制于本文所述的这些实施例,而是要符合与本文所公开的原理和新颖特点相一致的最宽的范围。

Claims (13)

  1. 一种视频生成方法,包括:
    获取多媒体描述内容;其中,所述多媒体描述内容包括文本描述内容和/或语音描述内容;
    以及,获取用户输入的多媒体素材;
    基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿;
    其中,所述视频草稿中包括多个视频片段,所述多个视频片段与多个描述内容片段之间具有对应关系,所述多个描述内容片段为基于所述多媒体描述内容确定,所述多个视频片段为从所述用户输入的多媒体素材中为所述多个描述内容片段分别提取到的,所述多个视频片段在所述视频草稿中的展示顺序为基于对应的描述内容片段之间的预设顺序关系确定。
  2. 根据权利要求1所述的方法,其中,所述基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿,包括:
    基于所述多媒体描述内容确定多个描述内容片段;其中,所述多个描述内容片段之间具有预设顺序关系;
    从所述用户输入的多媒体素材中,为所述多个描述内容片段分别裁剪出对应的视频片段;
    基于所述预设顺序关系,对所述多个描述内容片段分别对应的视频片段进行拼接处理,得到视频草稿。
  3. 根据权利要求2所述的方法,其中,所述基于所述多媒体描述内容确定多个描述内容片段,包括:
    根据所述多媒体内容生成视频文案;
    将所述视频文案拆分成多个描述内容片段;其中,所述多个描述内容片段之间具有预设顺序关系。
  4. 根据权利要求1所述的方法,其中,所述获取多媒体描述内容,包括:
    获取目标文本编辑框内的文本描述内容;其中,所述文本描述内容包括从目标音视频资源中提取到的第一文本内容,和/或,基于所述目标文本编辑框输入的第二文本内容。
  5. 根据权利要求1所述的方法,其中,所述获取多媒体描述内容,包括:
    接收针对目标对象的至少一个目标属性输入的文本内容;
    基于所述至少一个目标属性分别对应的文本内容,获取多媒体描述内容。
  6. 根据权利要求5所述的方法,其中,所述接收针对目标对象的至少一个目标属性输入的文本内容,包括:
    响应于针对目标对象的第一目标属性输入的第三文本内容,显示所述目标对象的第二目标属性对应的至少一个候选推荐内容;其中,所述至少一个候选推荐内容为基于所述第三文本内容确定;
    接收从所述至少一个候选推荐内容中选定的第四文本内容,和/或针对所述第二目标属性输入的第五文本内容;
    相应的,所述基于所述至少一个目标属性分别对应的文本内容,获取多媒体描述内容,包括:
    基于所述第三文本内容、所述第四文本内容和所述第五文本内容中的至少一种,获取多媒体描述内容。
  7. 根据权利要求2或3所述的方法,其中,所述基于所述预设顺序关系,对所述多个描述内容片段分别对应的视频片段进行拼接处理,得到视频草稿之前,还包括:
    基于所述多个描述内容片段分别为对应的视频片段生成语音片段;其中,所述语音片段与所述视频片段具有相同的播放时间信息;
    相应的,所述基于所述预设顺序关系,对所述多个描述内容片段分别对应的视频片段进行拼接处理,得到视频草稿,包括:
    基于所述预设顺序关系,对具有所述语音片段的视频片段进行拼接处理,得到视频草稿。
  8. 根据权利要求1所述的方法,其中,所述基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿之后,还包括:
    在视频预览页面上展示所述视频草稿;
    响应于作用在所述视频预览页面上的视频生成操作,基于所述多媒体描述内容和所述用户输入的多媒体素材,生成至少一个候选视频草稿,并在所述视频预览页面上展示所述至少一个候选视频草稿。
  9. 根据权利要求1所述的方法,其中,所述基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿之后,还包括:
    响应于针对所述视频草稿的编辑触发操作,在视频编辑页面上展示所述视频草稿;
    接收针对所述视频草稿的视频编辑信息;
    响应于针对所述视频草稿的导出操作,基于所述视频编辑信息生成所述视频草稿对应的结果视频。
  10. 一种视频生成装置,包括:
    第一获取模块,用于获取多媒体描述内容;其中,所述多媒体描述内容包括文本描述内容和/或语音描述内容;
    第二获取模块,用于获取用户输入的多媒体素材;
    第一生成模块,用于基于所述多媒体描述内容和所述用户输入的多媒体素材,生成视频草稿;
    其中,所述视频草稿中包括多个视频片段,所述多个视频片段与多个描述内容片段之间具有对应关系,所述多个描述内容片段为基于所述多媒体描述内容确定,所述多个视频片段为从所述用户输入的多媒体素材中为所述多个描述内容片段分别提取到的,所述多个视频片段在所述视频草稿中的展示顺序为基于对应的描述内容片段之间的预设顺序关系确定。
  11. 一种计算机可读存储介质,其中存储有指令,当所述指令在终端设备上运行时,使得所述终端设备实现如权利要求1-9任一项所述的方法。
  12. 一种视频生成设备,包括:存储器,处理器,及存储在所述存储器上并可在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序时,实现如权利要求1-9任一项所述的方法。
  13. 一种计算机程序产品,包括计算机程序/指令,所述计算机程序/指令被处理器执行时实现如权利要求1-9任一项所述的方法。
PCT/CN2025/083442 2024-03-28 2025-03-19 一种视频生成方法、装置、设备及存储介质 Pending WO2025201137A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410370301.7A CN120730128A (zh) 2024-03-28 2024-03-28 一种视频生成方法、装置、设备及存储介质
CN202410370301.7 2024-03-28

Publications (1)

Publication Number Publication Date
WO2025201137A1 true WO2025201137A1 (zh) 2025-10-02

Family

ID=97165341

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/083442 Pending WO2025201137A1 (zh) 2024-03-28 2025-03-19 一种视频生成方法、装置、设备及存储介质

Country Status (2)

Country Link
CN (1) CN120730128A (zh)
WO (1) WO2025201137A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109756751A (zh) * 2017-11-07 2019-05-14 腾讯科技(深圳)有限公司 多媒体数据处理方法及装置、电子设备、存储介质
WO2021259322A1 (zh) * 2020-06-23 2021-12-30 广州筷子信息科技有限公司 一种生成视频的系统和方法
CN114501064A (zh) * 2022-01-29 2022-05-13 北京有竹居网络技术有限公司 一种视频生成方法、装置、设备、介质及产品
CN116320605A (zh) * 2022-11-04 2023-06-23 上海积图科技有限公司 视频数据的生成方法、装置、设备及存储介质
CN117082304A (zh) * 2023-08-16 2023-11-17 北京达佳互联信息技术有限公司 视频生成方法、装置、计算机设备及存储介质

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109756751A (zh) * 2017-11-07 2019-05-14 腾讯科技(深圳)有限公司 多媒体数据处理方法及装置、电子设备、存储介质
WO2021259322A1 (zh) * 2020-06-23 2021-12-30 广州筷子信息科技有限公司 一种生成视频的系统和方法
CN114501064A (zh) * 2022-01-29 2022-05-13 北京有竹居网络技术有限公司 一种视频生成方法、装置、设备、介质及产品
CN116320605A (zh) * 2022-11-04 2023-06-23 上海积图科技有限公司 视频数据的生成方法、装置、设备及存储介质
CN117082304A (zh) * 2023-08-16 2023-11-17 北京达佳互联信息技术有限公司 视频生成方法、装置、计算机设备及存储介质

Also Published As

Publication number Publication date
CN120730128A (zh) 2025-09-30

Similar Documents

Publication Publication Date Title
JP7739470B2 (ja) ビデオ編集方法、装置、機器および記憶媒体
CN101459801B (zh) 信息处理装置和方法
EP4304185A1 (en) Multimedia resource clipping method and apparatus, device and storage medium
WO2018214772A1 (zh) 媒体数据处理方法、装置及存储介质
CN112004137A (zh) 一种智能视频创作方法及装置
US20090083642A1 (en) Method for providing graphic user interface (gui) to display other contents related to content being currently generated, and a multimedia apparatus applying the same
US7889967B2 (en) Information editing and displaying device, information editing and displaying method, information editing and displaying program, recording medium, server, and information processing system
WO2024056023A1 (zh) 一种视频编辑方法、装置、设备及存储介质
WO2025223097A1 (zh) 一种视频处理方法、装置、设备及存储介质
WO2025201137A1 (zh) 一种视频生成方法、装置、设备及存储介质
CN118016110B (zh) 一种媒体数据记录与播放方法
JP2011155329A (ja) 映像コンテンツ編集装置,映像コンテンツ編集方法および映像コンテンツ編集プログラム
WO2024235088A1 (zh) 视频编辑方法、装置、电子设备和存储介质
WO2025031162A1 (zh) 一种封面预览方法、装置、设备及存储介质
JP7179387B1 (ja) ハイライト動画生成システム、ハイライト動画生成方法、およびプログラム
US7610554B2 (en) Template-based multimedia capturing
JP3942471B2 (ja) データ編集方法、データ編集装置、データ記録装置および記録媒体
WO2026037041A1 (zh) 视频生成方法、装置、设备及存储介质
WO2026001828A1 (zh) 一种视频生成方法、装置、设备及存储介质
CN118842959A (zh) 一种视频生成方法、装置、设备及存储介质
CN116723361A (zh) 一种视频创作方法、装置、计算机设备及存储介质
WO2024002057A1 (zh) 音频的播放方法、装置和非易失性计算机可读存储介质
WO2023104079A1 (zh) 一种模板更新方法、装置、设备及存储介质
CN121126086A (zh) 一种视频生成方法、装置、设备及存储介质
CN106294600A (zh) 一种数码照片的多媒体编辑方法、展示方法及系统

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25777997

Country of ref document: EP

Kind code of ref document: A1