WO2024255839A1 - 视频处理方法、装置、设备、介质及程序产品 - Google Patents
视频处理方法、装置、设备、介质及程序产品 Download PDFInfo
- Publication number
- WO2024255839A1 WO2024255839A1 PCT/CN2024/099198 CN2024099198W WO2024255839A1 WO 2024255839 A1 WO2024255839 A1 WO 2024255839A1 CN 2024099198 W CN2024099198 W CN 2024099198W WO 2024255839 A1 WO2024255839 A1 WO 2024255839A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- audio
- beat
- video data
- feature
- optical flow
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/4302—Content synchronisation processes, e.g. decoder synchronisation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/78—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/783—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
- G06F16/7834—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using audio features
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/233—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/23418—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/242—Synchronisation processes, e.g. processing of PCR [Programme Clock References]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
- H04N21/4394—Processing of audio elementary streams involving operations for analysing the audio stream, e.g. detecting features or characteristics in audio streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44008—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream
Definitions
- the present disclosure provides a video processing method, apparatus, device, storage medium and program product to solve the technical problem of inaccurate matching of the visual effects provided to video data with the rhythm of audio data to a certain extent.
- the present disclosure provides a video processing method, comprising:
- a video processing module configured to adjust the playback speed of the first video data based on the audio beat feature of the audio data to obtain second video data; wherein the visual beat feature of the object in the second video data matches the audio beat feature at a corresponding time;
- a synthesis module is used to synthesize the second video data and the audio data into a target video.
- an electronic device characterized in that it includes one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the program includes instructions for executing the method described in the first aspect or the second aspect.
- a non-volatile computer-readable storage medium containing a computer program is provided.
- the processors execute the method described in the first aspect or the second aspect.
- FIG. 6 is an example of an audio beat graph according to an embodiment of the present disclosure.
- FIG. 7 is an example of a motion spectrum diagram according to an embodiment of the present disclosure.
- FIG. 8 is an example of a visual beat graph according to an embodiment of the present disclosure.
- FIG. 11 is a schematic diagram of a video processing device according to an embodiment of the present disclosure.
- the audio beat feature can characterize the rhythm of the audio data.
- the audio beat feature has intensity, wherein the beat feature with an intensity greater than or equal to the intensity threshold can be a heavy beat, and the beat feature with an intensity less than the intensity threshold can be a light beat.
- the visual beat feature can characterize the video rhythm of the video data; the visual beat feature has amplitude, the larger the amplitude, the stronger the video rhythm, and the smaller the amplitude, the weaker the video rhythm.
- the audio beat intensity feature can be subjected to peak detection based on the moving average, so that the audio beat intensity feature corresponding to the obtained peak value can be determined as the audio beat feature, such as the audio beat feature 610 in FIG6 .
- the audio beat point is usually a point where the frequency and/or volume suddenly changes, and the final audio beat feature point can be determined based on the peak detection of the moving average. Specifically, for each moment, a certain number of time windows before the moment (for example, a certain number of time windows before the current moment) and/or a certain number of time windows after the moment are selected.
- the audio beat moving average can be obtained by calculating the average value of the audio beat intensity of a certain amount (for example, several time windows after the current moment).
- obtaining a visual beat feature of an object in the first video data based on the optical flow information of the first video data includes:
- the significant discontinuous changes of moving objects in the video data can be used as local significant peaks of the video and determined as visual beat features.
- optical flow information that can reflect visual features such as the direction of motion amplitude between video frames can be used to perform visual beat analysis. The larger the optical flow information, the larger the amplitude of the visual beat feature, and the stronger the corresponding video rhythm.
- the video data can be processed for visual beat recognition based on the optical flow information to obtain a motion spectrum diagram and a visual beat diagram similar to the audio spectrum diagram and audio beat diagram of the audio data in Figures 5-6, as shown in Figures 7-8, Figure 7 shows an example of an optical flow spectrum diagram according to an embodiment of the present disclosure, and Figure 8 shows an example of a visual beat diagram according to an embodiment of the present disclosure.
- the sum of the pixel motion amplitudes corresponding to the pixel motion directions whose angle differences with the optical flow direction are within a preset range is counted to obtain the optical flow intensity corresponding to the optical flow direction.
- the visual beat intensity feature is subjected to peak detection based on a moving average, and the visual beat intensity feature corresponding to the peak value is determined as the visual beat feature.
- a trained visual beat recognition network can be used to perform visual beat recognition on the optical flow distribution features.
- the visual beat recognition network can be obtained by training the neural network based on the visual beat training data, and the visual beat training data can include an optical flow spectrum training graph used as input layer training data and a corresponding visual beat training graph used as output layer training data.
- the neural network is trained to obtain a trained visual beat recognition network.
- visual beat recognition can be performed on the optical flow distribution features in the optical flow spectrum graph based on the visual beat recognition network, and a visual beat graph including visual beat intensity features about time is obtained, as shown in Figure 8.
- the horizontal coordinate of the visual beat graph can be time (in s), and the vertical coordinate can be visual beat intensity. The higher the visual beat intensity at a certain moment, the higher the possibility that the moment is a visual beat point.
- the visual beat intensity feature can be subjected to peak detection based on the moving average, so that the visual beat intensity feature corresponding to the obtained peak value can be determined as the visual beat feature, as shown in the visual beat intensity feature in FIG8.
- Beat feature 810 Specifically, for each moment, a certain number of visual beat intensities before the moment (e.g., several time windows before the current moment) and/or a certain number of time windows after the moment (e.g., several time windows after the current moment) are selected to calculate the average value, and the visual beat moving average can be obtained.
- the average value corresponding to the visual beat moving average is used as a threshold to determine whether the visual beat intensity at each moment is greater than the threshold corresponding to the moment. If the visual beat intensity is greater than the threshold corresponding to the moment, the moment is a visual beat point, and the corresponding visual beat intensity feature is the visual beat feature.
- adjusting the playback speed between video frames in the first video data to align the visual beat feature with the time corresponding to the audio beat feature includes:
- the playback speed of the video frames in the first video data is adjusted so that the filtered visual beat features are aligned with the audio beat features in sequence.
- the second visual beat feature 910 is at 0.6s in the original video, and the second audio beat feature is at 2.4s in the audio, so the visual beat feature needs to be slowed down by adjusting the playback speed meter so that the time of the visual beat feature 910 reaches 2.4s.
- the method before performing optical flow detection on the first video data, the method further includes:
- the first video data may have a fixed video rhythm or may not have a fixed video rhythm.
- the playback speed of the video frame may be directly changed to match the audio rhythm.
- the first video data may be cropped to have a fixed video rhythm. For example, part of the first video data is removed, or part of the data is repeated multiple times.
- FIG10 shows an example of a video processing method according to an embodiment of the present disclosure. As shown in FIG10 , music data and first video data may be determined based on the user's operation, a rhythm analysis may be performed on the music data to obtain the music rhythm, and a video rhythm analysis may be performed on the first video data.
- the first video data is determined based on the video rhythm analysis.
- the first playback speed curve of the first video data is calculated; when it is determined based on the video rhythm analysis that the video data does not have a fixed rhythm, the first video data can be cropped and then the first playback speed curve of the cropped first video data can be calculated.
- the playback speed between video frames in the first video data can be changed so that the time point at which the motion change feature and the beat feature appear are aligned, and the playback speed between video frames is changed from the first playback speed curve to the second playback speed curve, and the second video data with the second playback speed curve is obtained.
- the music data and the second video data are synthesized to obtain a music card point video with a video rhythm that matches the music rhythm.
- the method before adjusting the playback speed between video frames in the first video data, the method further includes: setting the time and/or intensity of the audio beat feature according to user operation or the visual beat feature.
- step S330 the second video data and the audio data are synthesized into a target video.
- method 300 may further include: using special effects to play video frames in the target video corresponding to the beat feature as a heavy beat, wherein the heavy beat includes the intensity of the beat feature being greater than or equal to a preset intensity.
- the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server.
- the method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other.
- one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.
- a data determination module used for determining first video data and audio data
- a video processing module configured to adjust the playback speed of the first video data based on the audio beat feature of the audio data to obtain second video data; wherein the visual beat feature of the object in the second video data matches the audio beat feature at a corresponding time;
- a synthesis module is used to synthesize the second video data and the audio data into a target video.
- the computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology.
- Information can be computer-readable instructions, data structures, modules of programs, or other data.
- Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
- PRAM phase change memory
- SRAM static random access memory
- DRAM dynamic random access memory
- RAM random access memory
- ROM read-only memory
- EEPROM electrically erasable programmable
- the known power/ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures.
- the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art).
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Library & Information Science (AREA)
- Theoretical Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Image Analysis (AREA)
Abstract
本公开提供一种视频处理方法、装置、设备、存储介质及程序产品。该方法包括:确定第一视频数据和音频数据;基于所述音频数据的音频节拍特征调整所述第一视频数据的播放速度,得到第二视频数据;其中,所述第二视频数据中对象的视觉节拍特征与所述音频节拍特征在对应的时间相匹配;将所述第二视频数据和所述音频数据合成为目标视频。
Description
相关申请的交叉引用
本申请要求于2023年6月15日提交的,申请号为202310715849.6、发明名称为“视频处理方法、装置、设备、介质及程序产品”的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
本公开涉及计算机技术领域,尤其涉及一种视频处理方法、装置、设备、介质及程序产品。
用户希望发布的视频数据的视觉效果和音乐的节奏相匹配。现有的应用程序虽然可以合成视频数据和音频数据,然而却只能从给定的音乐数据库中选择音乐进行合成,无法自由地选择其他音乐。而且,合成之后只有视频画面切换处的视觉效果于音乐的节奏匹配,无法使得整个视频数据的视觉效果与整个音频数据的节奏相匹配,导致最终的音视频数据的呈现效果不佳,不能满足用户的需求。
发明内容
本公开提出一种视频处理方法、装置、设备、存储介质及程序产品,以在一定程度上解决提供给视频数据的视觉效果与音频数据的节奏匹配不准确的技术问题。
本公开第一方面,提供了一种视频处理方法,包括:
确定第一视频数据和音频数据;
基于所述音频数据的音频节拍特征调整所述第一视频数据的播放速度,得到第二视频数据;其中,所述第二视频数据中对象的视觉节拍特征与所述音频节拍特征在对应的时间相匹配;
将所述第二视频数据和所述音频数据合成为目标视频。
本公开第二方面,提供了一种视频处理装置,包括:
数据确定模块,用于确定第一视频数据和音频数据;
视频处理模块,用于基于所述音频数据的音频节拍特征调整所述第一视频数据的播放速度,得到第二视频数据;其中,所述第二视频数据中对象的视觉节拍特征与所述音频节拍特征在对应的时间相匹配;
合成模块,用于将所述第二视频数据和所述音频数据合成为目标视频。
本公开第三方面,提供了一种电子设备,其特征在于,包括一个或者多个处理器、存储器;和一个或多个程序,其中所述一个或多个程序被存储在所述存储器中,并且被所述一个或多个处理器执行,所述程序包括用于执行根据第一方面或第二方面所述的方法的指令。
本公开第四方面,提供了一种包含计算机程序的非易失性计算机可读存储介质,当所述计算机程序被一个或多个处理器执行时,使得所述处理器执行第一方面或第二方面所述的方法。
本公开第五方面,提供了一种计算机程序产品,包括计算机程序指令,当所述计算机程序指令在计算机上运行时,使得计算机执行第一方面所述的方法。
为了更清楚地说明本公开或相关技术中的技术方案,下面将对实施例或相关技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本公开的实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1为本公开实施例的视频处理架构的示意图。
图2为本公开实施例的示例性电子设备的硬件结构示意图。
图3为本公开实施例的视频处理方法的示意性流程图。
图4为本公开实施例的运动变化特征和节拍特征的时间对齐的示意性原理图。
图5为本公开实施例的音频频谱图的示例。
图6为本公开实施例的音频节拍图的示例。
图7为本公开实施例的运动频谱图的示例。
图8为本公开实施例的视觉节拍图的示例。
图9为本公开实施例的节拍特征对齐的示例。
图10为本公开实施例的视频处理方法的示例。
图11为本公开实施例的视频处理装置的示意图。
为使本公开的目的、技术方案和优点更加清楚明白,以下结合具体实施例,并参照附图,对本公开进一步详细说明。
需要说明的是,除非另外定义,本公开实施例使用的技术术语或者科学术语应当为本公开所属领域内具有一般技能的人士所理解的通常意义。本公开实施例中使用的“第一”、“第二”以及类似的词语并不表示任何顺序、数量或者重要性,而只是用来区分不同的组成部分。“包括”或者“包含”等类似的词语意指出现该词前面的元件或者物件涵盖出现在该词后面列举的元件或者物件及其等同,而不排除其他元件或者物件。“连接”或者“相连”等类似的词语并非限定于物理的或者机械的连接,而是可以包括电性的连接,不管是直接的还是间接的。“上”、“下”、“左”、“右”等仅用于表示相对位置关系,当被描述对象的绝对位置改变后,则该相对位置关系也可能相应地改变。
可以理解的是,在使用本公开各实施例公开的技术方案之前,均应当依据相关法律法规通过恰当的方式对本公开所涉及个人信息的类型、使用范围、使用场景等告知用户并获得用户的授权。
例如,在响应于接收到用户的主动请求时,向用户发送提示信息,以明确地提示用户,其请求执行的操作将需要获取和使用到用户的个人信息。从而,使得用户可以根据提示信息来自主地选择是否向执行本公开技术方案的操作的电子设备、应用程序、服务器或存储介质等软件或硬件提供个人信息。
作为一种可选的但非限定性的实现方式,响应于接收到用户的主动请求,向用户发送提示信息的方式例如可以是弹窗的方式,弹窗中可以以文字的方式呈现提示信息。此外,弹窗中还可以承载供用户选择“同意”或者“不同意”向电子设备提供个人信息的选择控件。
可以理解的是,上述通知和获取用户授权过程仅是示意性的,不对本公开
的实现方式构成限定,其它满足相关法律法规的方式也可应用于本公开的实现方式中。
图1示出了本公开实施例的视频处理架构的示意图。参考图1,该视频处理架构100可以包括服务器110、终端120以及提供通信链路的网络130。服务器110和终端120之间可通过有线或无线的网络130连接。其中,服务器110可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、安全服务、CDN等基础云计算服务的云服务器。
终端120可以是硬件或软件实现。例如,终端120为硬件实现时,可以是具有显示屏并且支持页面显示的各种电子设备,包括但不限于智能手机、平板电脑、电子书阅读器、膝上型便携计算机和台式计算机等等。终端120设备为软件实现时,可以安装在上述所列举的电子设备中;其可以实现成多个软件或软件模块(例如用来提供分布式服务的软件或软件模块),也可以实现成单个软件或软件模块,在此不做具体限定。
需要说明的是,本申请实施例所提供的视频处理方法可以由终端120来执行,也可以由服务器110来执行。应了解,图1中的终端、网络和服务器的数目仅为示意,并不旨在对其进行限制。根据实现需要,可以具有任意数目的终端、网络和服务器。
图2示出了本公开实施例所提供的示例性电子设备200的硬件结构示意图。如图2所示,电子设备200可以包括:处理器202、存储器204、网络模块206、外围接口208和总线210。其中,处理器202、存储器204、网络模块206和外围接口208通过总线210实现彼此之间在电子设备200的内部的通信连接。
处理器202可以是中央处理器(Central Processing Unit,CPU)、视频处理器、神经网络处理器(NPU)、微控制器(MCU)、可编程逻辑器件、数字信号处理器(DSP)、应用专用集成电路(Application Specific Integrated Circuit,ASIC)、或者一个或多个集成电路。处理器202可以用于执行与本公开描述的技术相关的功能。在一些实施例中,处理器202还可以包括集成为单一逻辑组件的多个处理器。例如,如图2所示,处理器202可以包括多个处理器202a、202b和202c。
存储器204可以配置为存储数据(例如,指令、计算机代码等)。如图2
所示,存储器204存储的数据可以包括程序指令(例如,用于实现本公开实施例的视频处理方法的程序指令)以及要处理的数据(例如,存储器可以存储其他模块的配置文件等)。处理器202也可以访问存储器204存储的程序指令和数据,并且执行程序指令以对要处理的数据进行操作。存储器204可以包括易失性存储装置或非易失性存储装置。在一些实施例中,存储器204可以包括随机访问存储器(RAM)、只读存储器(ROM)、光盘、磁盘、硬盘、固态硬盘(SSD)、闪存、存储棒等。
网络模块206可以配置为经由网络向电子设备200提供与其他外部设备的通信。该网络可以是能够传输和接收数据的任何有线或无线的网络。例如,该网络可以是有线网络、本地无线网络(例如,蓝牙、WiFi、近场通信(NFC)等)、蜂窝网络、因特网、或上述的组合。可以理解的是,网络的类型不限于上述具体示例。在一些实施例中,网络模块306可以包括任意数量的网络接口控制器(NIC)、射频模块、接收发器、调制解调器、路由器、网关、适配器、蜂窝网络芯片等的任意组合。
外围接口208可以配置为将电子设备200与一个或多个外围装置连接,以实现信息输入及输出。例如,外围装置可以包括键盘、鼠标、触摸板、触摸屏、麦克风、各类传感器等输入设备以及显示器、扬声器、振动器、指示灯等输出设备。
总线210可以被配置为在电子设备200的各个组件(例如处理器202、存储器204、网络模块206和外围接口208)之间传输信息,诸如内部总线(例如,处理器-存储器总线)、外部总线(USB端口、PCI-E总线)等。
需要说明的是,尽管上述电子设备200的架构仅示出了处理器202、存储器204、网络模块206、外围接口208和总线210,但是在具体实施过程中,该电子设备200的架构还可以包括实现正常运行所必需的其他组件。此外,本领域的技术人员可以理解的是,上述电子设备200的架构中也可以仅包含实现本公开实施例方案所必需的组件,而不必包含图中所示的全部组件。
用户希望发布的音视频数据的视觉效果和音乐的节奏相匹配,可以称之为视频与音频的“卡点”。而现有的应用程序所提供的卡点功能只能从给定曲库选择,无法自由添加喜欢的音乐,且只将卡点用在视频转场处,对视频只做裁剪无节奏变换;或者,对视频进行固定间隔的变速,例如500ms区间内先加速再减速,变速后区间长度大致不变,无法使得整个视频数据的视觉效果与整个
音频数据的节奏相匹配,导致最终的音视频数据的呈现效果不佳,不能满足用户的需求。因此,如何使视频的视觉效果与音频的节奏相匹配成为了亟需解决的技术问题。
鉴于此,本公开实施例提供了一种视频处理方法、装置、设备、存储介质及程序产品。通过调整视频数据的视频帧之间的播放速度,使得视频数据中对象的视觉节拍特征与音频数据中的音频节拍特征相匹配。使得最终合成的视频数据的视觉节奏与音频数据的节奏同步,提高了音视频数据的呈现效果,提高用户的使用体验。
具体地,由于视频数据的视觉效果往往通过其所显示的对象呈现,则可以基于图像中对象的视觉节拍特征来表征视频数据的视频节奏,例如视频帧之间运动幅度、方向等。然后根据音频数据的音频节拍特征,调整视频数据中视频帧的播放速度,使得视觉节拍特征出现的时间发生变化,变为与音频节拍特征出现的时间同步,从而完成视频节奏与音频节奏的同步,实现任意视频和音频之间的节奏匹配。
参见图3,图3示出了根据本公开实施例的视频处理方法的示意性流程图。本公开实施例的视频处理方法可以部署于客户端。图3中,视频处理方法300可以进一步包括如下步骤。
在步骤S310,确定第一视频数据和音频数据。
其中,第一视频数据和音频数据可以来自同一音视频数据。在一些实施例中,用户可以选择第一音视频数据A,然后基于第一音视频数据得到第一视频数据A1和音频数据A2。例如,针对音视频数据A,用户认为其中的视频视频视频节奏与音频节奏不匹配,则可以根据本公开实施例的方法,对第一音视频数据A进行处理,得到视频节奏与音频节奏向匹配的音视频数据。
第一视频数据和音频数据也可以来自不同的音视频数据。在一些实施例中,用户可以选择第一音视频数据B,然后基于第一音视频数据得到第一视频数据B1;以及用户可以选择音频数据C。例如,针对音视频数据B,用户希望采用音频数据C来替换音视频数据B中的音频数据,但替换后第一数据时间B1的视频节奏将与音频数据C的音频节奏不匹配,则可以根据本公开实施例的方法,对第一视频数据B1和音频数据C进行处理,使其视频节奏与音频节奏向匹配。
在步骤S320,基于所述音频数据的音频节拍特征调整所述第一视频数据的
播放速度,得到第二视频数据;其中,所述第二视频数据中对象的视觉节拍特征与所述音频节拍特征在对应的时间相匹配。
其中,音频节拍特征可以表征音频数据的节奏。音频节拍特征具有强度,其中,强度大于或等于强度阈值的节拍特征可以是重节拍,强度小于强度阈值的节拍特征可以是轻节拍。视觉节拍特征可以表征视频数据的视频节奏;视觉节拍特征具有幅度,幅度越大表示视频节奏越强,幅度越小表示视频节奏越弱。
在一些实施例中,基于所述音频数据的音频节拍特征调整所述第一视频数据的播放速度,得到第二视频数据,包括:
基于所述音频数据进行音频节拍分析,得到所述音频节拍特征;
基于所述第一视频数据的光流信息,得到所述第一视频数据中对象的视觉节拍特征;
调整所述第一视频数据中视频帧之间的播放速度,以将所述视觉节拍特征与所述音频节拍特征对应的时间对齐,得到所述第二视频数据。
具体地,参见图4,图4示出了根据本公开实施例的运动变化特征和节拍特征的时间对齐的示意性原理图。如图4所示,可以对音频数据进行节拍分析(例如频谱分析等),得到音频数据关于时间的音频节拍特征:时间t0处的音频节拍特征m0、时间t1处的音频节拍特征m1、时间t2处的音频节拍特征m2、……、时间ti处的音频节拍特征mi(i为自然数)、……。可以采用光流检测等技术得到第一视频数据中对象的光流信息,进而得到第一视频数据关于时间的视觉节拍特征:时间t0’处的视觉节拍特征n0、时间t1’处的视觉节拍特征n0、时间t2’处的视觉节拍特征n2、……、时间ti’处的视觉节拍特征ni、……。视觉节拍特征n0、n1、n2、……、ni、……可以分别对应视频帧Fn0、Fn1、Fn2、……、Fni、……,视频帧Fn0、Fn1之间的播放速度为v0,视频帧Fn1、Fn2之间的播放速度为v1,……,视频帧Fni、Fni+1之间的播放速度为vi,……。可见,视觉节拍特征m0-mi与视觉节拍特征n0-ni之间并不同步,则可以改变视频帧之间的播放速度,以改变运动变化特征n0-ni出现的时间点t0’-ti’,使得视觉节拍特征n0-ni与音频节拍特征m0-mi出现的时间点对齐,对齐后的视觉节拍特征n0-ni分别对应于时间点t0-ti。相应地,对齐后视觉节拍特征n0-ni对应的视频帧Fn0-Fni之间的播放速度也变为:视频帧Fn0、Fn1之间的播放速度为v0’,视频帧Fn1、Fn2之间的播放速度为v1’,……,视频帧Fni、Fni+1之间的播放速度为vi’,……,即由播放速度曲线v0-vi变为播放速度曲线
v0’-vi’。
在一些实施例中,基于所述音频数据进行音频节拍分析,得到所述音频数据的音频节拍特征,包括:
对所述音频数据进行频率分析得到所述音频数据关于时间的频谱特征,所述频谱特征包括音频频率和对应的声音强度;
基于所述频谱特征进行音频节拍识别,得到关于时间的音频节拍强度特征;
对所述音频节拍强度特征进行基于移动均线的峰值检测,确定峰值所对应的所述音频节拍强度特征为所述音频节拍特征。
具体地,频率分析可以包括傅里叶变换,例如短时傅里叶变换,将时域的音频数据转换为频域的音频频谱图。音频频谱图可以指音频的频率随时间变化的图像,能够描述某个时刻的各个频率处声音的分布情况。参见图5,图5示出了根据本公开实施例的音频频谱图的示例。图5中,音频频谱图的横坐标可以是时间(单位为s),纵坐标可以是频率(单位为HZ),对于某个时刻t的某个频率(例如0-8000HZ中的某个频率)处,可以采用不同的颜色来表示不同的声音强度(单位为dB)。
可以采用训练好的音频节拍识别网络对频谱特征进行音频节拍识别。其中,音频节拍识别网络可以基于训练数据对神经网络进行训练得到,训练数据可以包括音频频谱训练图和对应的音频节拍训练图,将音频频谱训练图作为输入层数据以及将对应的音频节拍训练图作为输出层数据对神经网络进行训练,可以得到训练好的音频节拍识别网络。具体地,可以基于音频节拍识别网络对音频频谱图中的频谱特征进行音频节拍识别,得到包括关于时间的音频节拍强度特征的音频节拍图,如图6所示,图6示出了根据本公开实施例的音频节拍图的示例。图6中,音频节拍图的横坐标可以是时间(单位为s),纵坐标可以是音频节拍强度,某时刻的音频节拍强度越高表示该时刻为音频节拍点的可能性越高。
接着,可以对音频节拍强度特征进行基于移动均线的峰值检测,从而将得到的峰值所对应的音频节拍强度特征确定为音频节拍特征,如图6中的音频节拍特征610。音频节拍点通常是频率和/或音量突变的点,据此可以基于移动均线的峰值检测确定最终的音频节拍特征点。具体地,针对每个时刻选取该时刻之前一定数量(例如当前时刻之前的若干个时间窗口)和/或该时刻之后一定数
量(例如当前时刻之后的若干个时间窗口)的音频节拍强度计算平均值,即可得到音频节拍移动均线。将该音频节拍移动均线对应的平均值作为阈值,判断每个时刻的音频节拍强度是否大于该时刻对应的阈值,如果音频节拍强度大于该时刻对应的阈值,则该时刻为音频节拍点,对应的音频节拍强度特征为音频节拍特征。
在一些实施例中,基于所述第一视频数据的光流信息,得到所述第一视频数据中对象的视觉节拍特征,包括:
针对所述第一视频数据进行光流检测,得到所述第一视频数据的中像素的光流信息;
基于所述光流信息得到所述第一视频数据关于时间的光流分布特征,所述光流分布特征包括多个光流方向和每个所述光流方向上对应的光流强度;
基于所述运动分布特征进行视觉节拍分析,得到所述第一视频数据中对象的视觉节拍特征。
具体地,可以将视频数据中移动对象的显著不连续变化作为视频局部显著峰值,并将其确定为视觉节拍特征。据此,可以采用能够反映视频帧之间的运动幅值方向等视觉特征的光流信息来进行视觉节拍分析。光流信息越大则视觉节拍特征的幅值越大,对应的视频节奏越强。据此,可以基于光流信息将视频数据进行视觉节拍识别处理得到与图5-图6中的音频数据的音频频谱图和音频节拍图类似的运动频谱图和视觉节拍图,如图7-图8所示,图7示出了根据本公开实施例的光流频谱图的示例,图8示出了根据本公开实施例的视觉节拍图的示例。
在一些实施例中,基于所述光流信息得到所述第一视频数据关于时间的光流分布特征,包括:
基于所述光流信息的角度得到每个所述像素的像素运动方向,以及基于所述光流信息的幅值得到每个所述像素的像素运动幅值;
针对每个所述光流方向,统计与所述光流方向的角度差在预设范围内的所述像素运动方向所对应的所述像素运动幅值之和,得到所述光流方向上对应的光流强度。
其中,如图7所示,光流频谱图的横坐标可以包括时间(单位为s),纵坐标可以包括光流方向,该光流方向可以是预设的。对于某一时刻的某个光流方向处,可以采用不同的颜色来表示不同的光流强度。具体地,对于视频数据中
的每个像素(x,y)在时刻t的光流信息Ft,,其像素运动方向可以表示为光流信息Ft的角度φ,像素运动幅值可以表示为光流信息Ft的幅值|Ft(x,y)|。统计所有像素的像素运动方向及像素运动幅值,将与同一光流方向的角度差在预设范围内的的运动幅值求和,从而得到在各个光流方向上的光流幅值分布。即,图7中的光流频谱图中的某一点D(t,θ)可以表示为:
其中,θ为光流方向,φ为像素运动方向,Nbins表示光流方向的个数,如图7中的纵坐标所示,图7中光流频谱图包括6个预设的光流方向。通过统计在光流方向θ附近(例如与光流方向θ相差预设角度2π/Nbins)的所有光流信息的像素运动幅值之和,可以得到在光流方向θ上的光流强度。
在一些实施例中,基于所述光流分布特征进行视觉节拍分析,得到所述视觉节拍特征,包括:
基于所述光流分布特征进行视觉节拍识别,得到关于时间的视觉节拍强度特征;
对所述视觉节拍强度特征进行基于移动均线的峰值检测,确定峰值所对应的所述视觉节拍强度特征为所述视觉节拍特征。
具体地,与音频节拍识别过程类似,可以采用训练好的视觉节拍识别网络对光流分布特征进行视觉节拍识别。其中,视觉节拍识别网络可以基于视觉节拍训练数据对神经网络进行训练得到,视觉节拍训练数据可以包括用作输入层训练数据的光流频谱训练图和对应的用作输出层训练数据的视觉节拍训练图。基于该视觉节拍训练数据对神经网络进行训练,可以得到训练好的视觉节拍识别网络。具体地,可以基于视觉节拍识别网络对光流频谱图中的光流分布特征进行视觉节拍识别,得到包括关于时间的视觉节拍强度特征的视觉节拍图,如图8所示。图8中,视觉节拍图的横坐标可以是时间(单位为s),纵坐标可以是视觉节拍强度,某时刻的视觉节拍强度越高表示该时刻为视觉节拍点的可能性越高。
接着,可以对视觉节拍强度特征进行基于移动均线的峰值检测,从而将得到的峰值所对应的视觉节拍强度特征确定为视觉节拍特征,如图8中的视觉节
拍特征810。具体地,针对每个时刻选取该时刻之前一定数量(例如当前时刻之前的若干个时间窗口)和/或该时刻之后一定数量(例如当前时刻之后的若干个时间窗口)的视觉节拍强度计算平均值,即可得到视觉节拍移动均线。将该视觉节拍移动均线对应的平均值作为阈值,判断每个时刻的视觉节拍强度是否大于该时刻对应的阈值,如果视觉节拍强度大于该时刻对应的阈值,则该时刻为视觉节拍点,对应的视觉节拍强度特征为视觉节拍特征。
在一些实施例中,调整所述第一视频数据中视频帧之间的播放速度,以将所述视觉节拍特征与所述音频节拍特征对应的时间对齐,包括:
对所述视觉节拍特征和所述音频节拍特征进行过滤,得到过滤后的视觉节拍特征和音频节拍特征;
调整所述第一视频数据中视频帧的播放速度,使得过滤后的所述视觉节拍特征与所述音频节拍特征依次对齐。
其中,可以对视觉节拍特征、音频节拍特征进行过滤,例如剔除节拍间隔时间过短的点,使得两种节拍特征数量一致。例如,参见图9,图9示出了根据本公开实施例的节拍特征对齐的示例。图9中,第i个视觉节拍特征,需要变速到第i个音频节拍特征对应的时刻。因此可以计算出一个视频变速曲线,纵坐标表示原始时刻,横坐标表示需要变速到的目标时刻。例如第2个视觉节拍特征910是在原始视频的0.6s,第二个音频节拍特征是在音频的2.4s,因此需要将该视觉节拍特征通过调整播放速度计慢速,使得视觉节拍特征910的时间到2.4s。
在一些实施例中,针对所述第一视频数据进行光流检测之前,还包括:
重复所述第一视频数据的至少部分数据,以使得重复后的所述第一视频数据的对象具有周期性的光流变化。
具体地,第一视频数据可以具有固定的视频节奏,也可以不具有固定的视频节奏。当第一视频数据具有固定的视频节奏时,可以直接改变视频帧的播放速度来与音频节奏匹配。当第一视频数据不具有固定的视频节奏时,可以对第一视频数据进行裁剪使其变为具有固定的视频节奏。例如,将第一视频数据的部分数据移除,或者对部分数据进行重复多次。参见图10,图10示出了根据本公开实施例的视频处理方法的示例。如图10所示,可以基于用户的操作确定音乐数据和第一视频数据,对音乐数据进行节奏分析得到音乐节奏,以及对第一视频数据进行视频节奏分析。其中,基于视频节奏分析确定第一视频数据
具有固定的节奏时,计算第一视频数据的第一播放速度曲线;基于视频节奏分析确定视频数据不具有固定的节奏时,可以对第一视频数据进行裁剪后计算裁剪后第一视频数据的第一播放速度曲线。可以改变第一视频数据中视频帧之间的播放速度,以使得运动变化特征与节拍特征出现的时间点对齐,则视频帧之间的播放速度由第一播放速度曲线变为第二播放速度曲线,得到具有第二播放速度曲线的第二视频数据。最后将音乐数据和第二视频数据合成,即可得到视频节奏与音乐节奏相匹配的音乐卡点视频。
在一些实施例中,调整所述第一视频数据中视频帧之间的播放速度之前,所述方法还包括:根据用户操作或所述视觉节拍特征设置所述音频节拍特征的时间和/或强度。
具体地,用户可以对节拍特征进行增加、移除、移动等操作来改变音频节拍特征在时间轴上的位置,即改变音频节拍特征出现的时间。也可以针对音频节拍特征的强度进行变更,例如,增大或减小音频节拍特征的强度。这样,用户可以根据自由地编辑最终音视频中视频节奏和音频节奏匹配的时间和强度,即实现对卡点效果的自由设置。
在步骤S330,将所述第二视频数据和所述音频数据合成目标视频。
在一些实施例中,方法300还可以包括:采用特效播放所述目标视频中对应于所述节拍特征为重节拍的视频帧,所述重节拍包括所述节拍特征的强度大于或等于预设强度。
具体地,可以对重节拍对应的视频帧采用慢放、放大、抖动等视觉特效进行播放,配合与之匹配的音频节奏,以进一步增加视觉的冲击力,提高目标视频的呈现效果。
需要说明的是,本公开实施例的方法可以由单个设备执行,例如一台计算机或服务器等。本实施例的方法也可以应用于分布式场景下,由多台设备相互配合来完成。在这种分布式场景的情况下,这多台设备中的一台设备可以只执行本公开实施例的方法中的某一个或多个步骤,这多台设备相互之间会进行交互以完成所述的方法。
需要说明的是,上述对本公开的一些实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于上述实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期
望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。
基于同一技术构思,与上述任意实施例方法相对应的,本公开还提供了一种视频处理装置,参见图11,所述视频处理装置包括:
数据确定模块,用于确定第一视频数据和音频数据;
视频处理模块,用于基于所述音频数据的音频节拍特征调整所述第一视频数据的播放速度,得到第二视频数据;其中,所述第二视频数据中对象的视觉节拍特征与所述音频节拍特征在对应的时间相匹配;
合成模块,用于将所述第二视频数据和所述音频数据合成为目标视频。
为了描述的方便,描述以上装置时以功能分为各种模块分别描述。当然,在实施本公开时可以把各模块的功能在同一个或多个软件和/或硬件中实现。
上述实施例的装置用于实现前述任一实施例中相应的视频处理方法,并且具有相应的方法实施例的有益效果,在此不再赘述。
基于同一技术构思,与上述任意实施例方法相对应的,本公开还提供了一种非暂态计算机可读存储介质,所述非暂态计算机可读存储介质存储计算机指令,所述计算机指令用于使所述计算机执行如上任一实施例所述的视频处理方法。
本实施例的计算机可读介质包括永久性和非永久性、可移动和非可移动媒体可以由任何方法或技术来实现信息存储。信息可以是计算机可读指令、数据结构、程序的模块或其他数据。计算机的存储介质的例子包括,但不限于相变内存(PRAM)、静态随机存取存储器(SRAM)、动态随机存取存储器(DRAM)、其他类型的随机存取存储器(RAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、快闪记忆体或其他内存技术、只读光盘只读存储器(CD-ROM)、数字多功能光盘(DVD)或其他光学存储、磁盒式磁带,磁带磁磁盘存储或其他磁性存储设备或任何其他非传输介质,可用于存储可以被计算设备访问的信息。
上述实施例的存储介质存储的计算机指令用于使所述计算机执行如上任一实施例所述的视频处理方法,并且具有相应的方法实施例的有益效果,在此不再赘述。
所属领域的普通技术人员应当理解:以上任何实施例的讨论仅为示例性的,并非旨在暗示本公开的范围(包括权利要求)被限于这些例子;在本公开
的思路下,以上实施例或者不同实施例中的技术特征之间也可以进行组合,步骤可以以任意顺序实现,并存在如上所述的本公开实施例的不同方面的许多其它变化,为了简明它们没有在细节中提供。
另外,为简化说明和讨论,并且为了不会使本公开实施例难以理解,在所提供的附图中可以示出或可以不示出与集成电路(IC)芯片和其它部件的公知的电源/接地连接。此外,可以以框图的形式示出装置,以便避免使本公开实施例难以理解,并且这也考虑了以下事实,即关于这些框图装置的实施方式的细节是高度取决于将要实施本公开实施例的平台的(即,这些细节应当完全处于本领域技术人员的理解范围内)。在阐述了具体细节(例如,电路)以描述本公开的示例性实施例的情况下,对本领域技术人员来说显而易见的是,可以在没有这些具体细节的情况下或者这些具体细节有变化的情况下实施本公开实施例。因此,这些描述应被认为是说明性的而不是限制性的。
尽管已经结合了本公开的具体实施例对本公开进行了描述,但是根据前面的描述,这些实施例的很多替换、修改和变型对本领域普通技术人员来说将是显而易见的。例如,其它存储器架构(例如,动态RAM(DRAM))可以使用所讨论的实施例。
本公开实施例旨在涵盖落入所附权利要求的宽泛范围之内的所有这样的替换、修改和变型。因此,凡在本公开实施例的精神和原则之内,所做的任何省略、修改、等同替换、改进等,均应包含在本公开的保护范围之内。
Claims (12)
- 一种视频处理方法,包括:确定第一视频数据和音频数据;基于所述音频数据的音频节拍特征调整所述第一视频数据的播放速度,得到第二视频数据;其中,所述第二视频数据中对象的视觉节拍特征与所述音频节拍特征在对应的时间相匹配;将所述第二视频数据和所述音频数据合成为目标视频。
- 根据权利要求1的方法,其中,基于所述音频数据的音频节拍特征调整所述第一视频数据的播放速度,得到第二视频数据,包括:基于所述音频数据进行音频节拍分析,得到所述音频节拍特征;基于所述第一视频数据的光流信息,得到所述第一视频数据中对象的视觉节拍特征;调整所述第一视频数据中视频帧之间的播放速度,以将所述视觉节拍特征与所述音频节拍特征对应的时间对齐,得到所述第二视频数据。
- 根据权利要求2的方法,其中,基于所述音频数据进行音频节拍分析,得到所述音频数据的音频节拍特征,包括:对所述音频数据进行频率分析得到所述音频数据关于时间的频谱特征,所述频谱特征包括音频频率和对应的声音强度;基于所述频谱特征进行音频节拍识别,得到关于时间的音频节拍强度特征;对所述音频节拍强度特征进行基于移动均线的峰值检测,确定峰值所对应的所述音频节拍强度特征为所述音频节拍特征。
- 根据权利要求2的方法,其中,基于所述第一视频数据的光流信息,得到所述第一视频数据中对象的视觉节拍特征,包括:针对所述第一视频数据进行光流检测,得到所述第一视频数据中像素的光流信息;基于所述光流信息得到所述第一视频数据关于时间的光流分布特征,所述光流分布特征包括多个光流方向和每个所述光流方向上对应的光流强度;基于所述光流分布特征进行视觉节拍分析,得到所述第一视频数据中对象 的视觉节拍特征。
- 根据权利要求4的方法,其中,基于所述光流信息得到所述第一视频数据关于时间的光流分布特征,包括:基于所述光流信息的角度得到每个所述像素的像素运动方向,以及基于所述光流信息的幅值得到每个所述像素的像素运动幅值;针对每个所述光流方向,统计与所述光流方向的角度差在预设范围内的所述像素运动方向所对应的所述像素运动幅值之和,得到所述光流方向上对应的光流强度。
- 根据权利要求4的方法,其中,基于所述光流分布特征进行视觉节拍分析,得到所述视觉节拍特征,包括:基于所述光流分布特征进行视觉节拍识别,得到关于时间的视觉节拍强度特征;对所述视觉节拍强度特征进行基于移动均线的峰值检测,确定峰值所对应的所述视觉节拍强度特征为所述视觉节拍特征。
- 根据权利要求2的方法,其中,调整所述第一视频数据中视频帧之间的播放速度,以将所述视觉节拍特征与所述音频节拍特征对应的时间对齐,包括:对所述视觉节拍特征和所述音频节拍特征进行过滤,得到过滤后的视觉节拍特征和音频节拍特征;调整所述第一视频数据中视频帧的播放速度,使得过滤后的所述视觉节拍特征与所述音频节拍特征依次对齐。
- 根据权利要求2的方法,其中,针对所述第一视频数据进行光流检测之前,所述方法还包括:重复所述第一视频数据的至少部分数据,以使得重复后的所述第一视频数据的对象具有周期性的光流变化;或者,调整所述第一视频数据中视频帧之间的播放速度之前,所述方法还包括:根据用户操作或所述视觉节拍特征设置所述音频节拍特征的时间和/或强度。
- 一种视频处理装置,包括:数据确定模块,用于确定第一视频数据和音频数据;视频处理模块,用于基于所述音频数据的音频节拍特征调整所述第一视频数据的播放速度,得到第二视频数据;其中,所述第二视频数据中对象的视觉节拍特征与所述音频节拍特征在对应的时间相匹配;合成模块,用于将所述第二视频数据和所述音频数据合成为目标视频。
- 一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时实现如权利要求1至8任意一项所述的方法。
- 一种非暂态计算机可读存储介质,所述非暂态计算机可读存储介质存储计算机指令,所述计算机指令用于使计算机执行权利要求1至8任一所述方法。
- 一种计算机程序产品,包括计算机程序指令,当所述计算机程序指令在计算机上运行时,使得计算机执行权利要求1至8任一所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310715849.6A CN116761027A (zh) | 2023-06-15 | 2023-06-15 | 视频处理方法、装置、设备、介质及程序产品 |
| CN202310715849.6 | 2023-06-15 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024255839A1 true WO2024255839A1 (zh) | 2024-12-19 |
Family
ID=87960301
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/099198 Ceased WO2024255839A1 (zh) | 2023-06-15 | 2024-06-14 | 视频处理方法、装置、设备、介质及程序产品 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN116761027A (zh) |
| WO (1) | WO2024255839A1 (zh) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115988262B (zh) * | 2022-12-14 | 2025-02-25 | 北京有竹居网络技术有限公司 | 用于视频处理的方法、装置、设备和介质 |
| CN116761027A (zh) * | 2023-06-15 | 2023-09-15 | 北京字跳网络技术有限公司 | 视频处理方法、装置、设备、介质及程序产品 |
| CN118447870B (zh) * | 2023-12-28 | 2025-05-02 | 荣耀终端股份有限公司 | 音频处理方法和电子设备 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112380396A (zh) * | 2020-11-11 | 2021-02-19 | 网易(杭州)网络有限公司 | 视频处理方法及装置、计算机可读存储介质和电子设备 |
| CN113473201A (zh) * | 2021-07-29 | 2021-10-01 | 腾讯音乐娱乐科技(深圳)有限公司 | 一种音视频对齐方法、装置、设备及存储介质 |
| US20220261573A1 (en) * | 2021-02-12 | 2022-08-18 | Adobe Inc. | Re-timing a video sequence to an audio sequence based on motion and audio beat detection |
| CN115988262A (zh) * | 2022-12-14 | 2023-04-18 | 北京有竹居网络技术有限公司 | 用于视频处理的方法、装置、设备和介质 |
| CN116761027A (zh) * | 2023-06-15 | 2023-09-15 | 北京字跳网络技术有限公司 | 视频处理方法、装置、设备、介质及程序产品 |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110797055B (zh) * | 2019-10-29 | 2021-09-03 | 北京达佳互联信息技术有限公司 | 多媒体资源合成方法、装置、电子设备及存储介质 |
| CN112822563A (zh) * | 2019-11-15 | 2021-05-18 | 北京字节跳动网络技术有限公司 | 生成视频的方法、装置、电子设备和计算机可读介质 |
-
2023
- 2023-06-15 CN CN202310715849.6A patent/CN116761027A/zh active Pending
-
2024
- 2024-06-14 WO PCT/CN2024/099198 patent/WO2024255839A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112380396A (zh) * | 2020-11-11 | 2021-02-19 | 网易(杭州)网络有限公司 | 视频处理方法及装置、计算机可读存储介质和电子设备 |
| US20220261573A1 (en) * | 2021-02-12 | 2022-08-18 | Adobe Inc. | Re-timing a video sequence to an audio sequence based on motion and audio beat detection |
| CN113473201A (zh) * | 2021-07-29 | 2021-10-01 | 腾讯音乐娱乐科技(深圳)有限公司 | 一种音视频对齐方法、装置、设备及存储介质 |
| CN115988262A (zh) * | 2022-12-14 | 2023-04-18 | 北京有竹居网络技术有限公司 | 用于视频处理的方法、装置、设备和介质 |
| CN116761027A (zh) * | 2023-06-15 | 2023-09-15 | 北京字跳网络技术有限公司 | 视频处理方法、装置、设备、介质及程序产品 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116761027A (zh) | 2023-09-15 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2024255839A1 (zh) | 视频处理方法、装置、设备、介质及程序产品 | |
| KR102792043B1 (ko) | 비디오 생성 장치 및 방법, 전자 장치, 및 컴퓨터 판독가능 매체 | |
| JP7165818B2 (ja) | ニューラルネットワークのトレーニング方法及び装置並びに画像生成方法及び装置 | |
| JP7387890B2 (ja) | 動画ファイルの生成方法、装置、端末及び記憶媒体 | |
| US11688412B2 (en) | Multi-modal framework for multi-channel target speech separation | |
| US12380570B2 (en) | Image processing method and apparatus, and hardware apparatus | |
| CN113870104A (zh) | 超分辨率图像重建 | |
| CN113344776B (zh) | 图像处理方法、模型训练方法、装置、电子设备及介质 | |
| CN114419300A (zh) | 风格化图像生成方法、装置、电子设备及存储介质 | |
| US20220044693A1 (en) | Internet calling method and apparatus, computer device, and storage medium | |
| KR20220106848A (ko) | 비디오 특수 효과 처리 방법 및 장치 | |
| US20230403413A1 (en) | Method and apparatus for displaying online interaction, electronic device and computer readable medium | |
| CN110312162A (zh) | 精选片段处理方法、装置、电子设备及可读介质 | |
| US20240276037A1 (en) | Video generation method and device | |
| CN113160849A (zh) | 歌声合成方法、装置及电子设备和计算机可读存储介质 | |
| CN113921032A (zh) | 音频处理模型的训练方法及装置、音频处理方法及装置 | |
| CN111798866B (zh) | 音频处理网络的训练及立体声重构方法和装置 | |
| WO2024255652A1 (zh) | 视频生成方法、装置、设备、介质及程序产品 | |
| CN109243479B (zh) | 音频信号处理方法、装置、电子设备及存储介质 | |
| CN113808606B (zh) | 语音信号处理方法和装置 | |
| CN115454287A (zh) | 虚拟数字人交互方法、装置、设备及可读存储介质 | |
| CN115237248B (zh) | 虚拟对象的展示方法、装置、设备、存储介质及程序产品 | |
| CN113707163B (zh) | 语音处理方法及其装置和模型训练方法及其装置 | |
| WO2025036372A1 (zh) | 视频处理方法及相关设备 | |
| CN118098258A (zh) | 音频数据处理方法、装置、设备以及可读介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24822787 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |