WO2025201466A1 - 视频处理方法、装置、电子设备及存储介质 - Google Patents
视频处理方法、装置、电子设备及存储介质Info
- Publication number
- WO2025201466A1 WO2025201466A1 PCT/CN2025/085393 CN2025085393W WO2025201466A1 WO 2025201466 A1 WO2025201466 A1 WO 2025201466A1 CN 2025085393 W CN2025085393 W CN 2025085393W WO 2025201466 A1 WO2025201466 A1 WO 2025201466A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- initial
- segment
- video data
- subtitles
- revised
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/488—Data services, e.g. news ticker
- H04N21/4884—Data services, e.g. news ticker for displaying subtitles
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/233—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/23418—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
- H04N21/4394—Processing of audio elementary streams involving operations for analysing the audio stream, e.g. detecting features or characteristics in audio streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44008—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/485—End-user interface for client configuration
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/485—End-user interface for client configuration
- H04N21/4856—End-user interface for client configuration for language selection, e.g. for the menu or subtitles
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/488—Data services, e.g. news ticker
Definitions
- the embodiments of the present disclosure relate to a video processing method, device, electronic device, and storage medium.
- video editing scenarios users use the various video editing functions provided by video editing applications to record and produce videos, ultimately creating a video work that meets their needs.
- the video copy content is usually determined first. After that, the user will record the video (audio) based on the video copy content to form a video file or draft.
- the embodiments of the present disclosure provide a video processing method, apparatus, electronic device, and storage medium to overcome the problems of low modification efficiency and poor modification effect when modifying voice content in a video.
- an embodiment of the present disclosure provides a video processing method, including:
- an embodiment of the present disclosure provides a video processing method and apparatus, including:
- a generating module configured to perform speech recognition on the initial video data to obtain initial subtitles corresponding to the initial video data
- a correction module configured to generate a corrected subtitle in response to an editing operation on the initial subtitle
- a processing module is configured to generate modified video data corresponding to the initial video data based on the modified subtitles, wherein the speech content in the modified video data matches the subtitle content of the modified subtitles.
- an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
- the memory stores computer-executable instructions
- the processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video processing method described in the first aspect and various possible designs of the first aspect.
- an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored.
- a processor executes the computer execution instructions, the video processing method described in the first aspect and various possible designs of the first aspect is implemented.
- an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the video processing method described in the first aspect and various possible designs of the first aspect.
- FIG1 is a diagram of an application scenario of a video processing method provided by an embodiment of the present disclosure
- FIG2 is a flow chart of a video processing method according to an embodiment of the present disclosure.
- FIG3 is a flow chart of a possible implementation of step S103 in the embodiment shown in FIG2 ;
- FIG4 is a schematic diagram of a process for generating corrected video data according to an embodiment of the present disclosure
- FIG5 is a flowchart of a specific implementation method of step S1032 in the embodiment shown in FIG3 ;
- FIG6 is a schematic diagram of a capture of voice track data provided by an embodiment of the present disclosure.
- FIG7 is a flowchart of a specific implementation of step S1033 in the embodiment shown in FIG3 ;
- FIG8 is a schematic diagram of generating a second modified speech segment according to an embodiment of the present disclosure.
- FIG. 1 is an application scenario diagram of the video processing method provided by the embodiment of the present disclosure.
- the video processing method provided by the embodiment of the present disclosure can be applied to an application (APP, Application) with a video editing function, such as a short video application, a video editing application, etc. More specifically, it can be applied to an application scenario in which the voice content in a video draft is modified.
- the execution subject of this embodiment can be a terminal device that runs the above-mentioned application with a video editing function, or a server that deploys the service end corresponding to the above-mentioned application, or other electronic devices that perform similar functions.
- the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communications, cloud databases, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, among which cloud services can be interactive processing services for terminal devices to call.
- cloud services can be interactive processing services for terminal devices to call.
- the terminal device after responding to a user's first editing operation (which may consist of multiple specific sub-operations, such as recording video, recording voice, cropping a video, inserting special effects, and filters), the terminal device completes recording the video and generates video data (video draft).
- the video data includes multiple editing tracks, such as Track_1, Track_2, and Track_3, as shown in the figure.
- Track_1 is video track data for setting image media
- Track_2 is voice track data for setting audio media
- Track_3 is special effects track data for setting special effects.
- a second editing operation is performed on the terminal device.
- the terminal device responds to this second editing operation and completes the modification of the voice content of the video draft, i.e., the modification of Track_1.
- the usual implementation method is to re-record the audio track data of the video, which leads to the problem of inefficient and time-consuming editing process.
- the re-recorded audio track data needs to be manually aligned, it will further lead to problems such as audio overlap, affecting the audio quality of the video.
- Step S101 performing speech recognition on initial video data to obtain initial subtitles corresponding to the initial video data.
- Step S102 In response to the editing operation on the initial subtitles, a revised subtitle is generated.
- the terminal device runs, for example, a video editing application, it loads the initial video data through the video editing application, wherein the initial video data can be a kind of in-program data, such as a video draft created by the above-mentioned video editing application; thereafter, by parsing the voice track data in the initial video data, a text description corresponding to the voice content of the initial video data, i.e., the initial subtitles, can be automatically generated.
- the initial video data can be a kind of in-program data, such as a video draft created by the above-mentioned video editing application.
- the process of generating the initial subtitles corresponding to the voice content of the initial video data can be triggered based on user operations, for example, after the user clicks the “Generate Subtitles” control, the above process is executed and the initial subtitles are generated, or the initial subtitles are automatically generated after the initial video data is loaded, and are displayed or not displayed based on the specific design.
- the implementation method of parsing and generating corresponding subtitles based on the voice content in the video will not be introduced in detail this time.
- the terminal device in response to the editing operation applied by the user to the initial subtitles, completes the modification of the initial subtitles within the video editing application and generates revised subtitles.
- the initial subtitles include text consisting of a number of characters and punctuation marks, such as "XXXABCXXX" (where X represents a placeholder, and its content can be any value).
- the user applies an editing operation to the text segment "ABC” in the initial subtitles and modifies it to "CBA", then the generated revised subtitles are "XXXCBAXXX".
- the terminal device After the terminal device responds to the above-mentioned editing operation, it first modifies the initial subtitles at the text level according to the content input by the editing operation, that is, conventional text modification, and forms text-based revised subtitles.
- the process may further include steps such as checking the legitimacy of the characters input by the editing operation and detecting the number of input words, which can be set as needed.
- Step S103 Based on the modified subtitles, generating modified video data corresponding to the initial video data, wherein the speech content in the modified video data matches the subtitle content of the modified subtitles.
- the terminal device After obtaining the revised subtitles, that is, completing the editing process for the initial subtitles, the terminal device corrects the speech content in the initial video data based on the text content of the generated revised subtitles, so that the speech content in the initial video data matches the subtitle content of the revised subtitles.
- a pre-trained speech model can be used to achieve "text-to-sound" conversion, replacing the revised speech segments in the initial video data, thereby generating revised video data.
- the revised video data is the result of adjustments to the initial video data. Therefore, similar to the initial video data, it is also a type of in-program data, such as a video draft. Subsequently, the revised video data can be further converted into a corresponding video file as needed.
- a possible implementation step of step 103 includes:
- Step S1031 determining a first playback time period of the initial video data based on a changed text segment in the initial subtitles, wherein the changed text segment includes sentences in the initial subtitles that have content differences from the revised subtitles at corresponding positions.
- Step S1032 generating a first revised language segment based on the revised text segment in the revised subtitles, wherein the revised text segment includes sentences in the revised subtitles that have content differences from the original subtitles at corresponding positions.
- Step S1033 Replace the first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment, or delete the first initial voice segment corresponding to the first playback time period in the initial video data to generate revised video data.
- the terminal device determines the revised text segment in the revised subtitles corresponding to the changed text segment based on the position of the changed text segment, and inputs the revised text segment into the speech model to generate a corresponding speech segment, namely, the first revised speech segment.
- the subtitles of the video correspond to a timestamp sequence, which is used to represent the playback timestamps of each text in the subtitles, so that the subtitles and the corresponding video content (video images and video voice) can be played synchronously through the timestamp sequence.
- the initial subtitles in this embodiment also correspond to a timestamp sequence. Therefore, while determining the changed text segment, the first playback time period corresponding to the changed text segment can be obtained based on the timestamp sequence. Afterwards, based on the first playback time period, the corresponding voice data, that is, the first initial voice segment, is intercepted from the initial video data. Thereafter, the first corrected voice segment obtained in the above steps is used to replace the first initial voice segment to generate corrected video data.
- Figure 4 is a schematic diagram of a process for generating corrected video data provided by an embodiment of the present disclosure.
- the video data is composed of multiple track data, including, for example, voice track data, video track data, special effects track data, etc.
- initial subtitles are generated; then, based on the editing operation of the changed text segment in the initial subtitles (shown as the T1-T2 segment in the figure, with the text content being "I say"), revised subtitles are generated, and the revised subtitles contain a revised text segment corresponding to the changed text segment (shown as the T3-T4 segment in the figure, with the text content being "I say”); then, based on the revised text segment, a corresponding first revised voice segment D2 is generated; then, based on the first playback time period corresponding to the changed text segment, the first initial voice segment D1 is extracted from the voice track data of the initial video data; finally, the first revised voice segment D2 is replaced with the position of the first initial voice segment D1, that is, the first playback time period corresponding to the changed text segment, to obtain the adjusted voice track data L2; and based on the adjusted voice track data L2 and other track data, the revised video data is obtained.
- Step S1032B Generate a first corrected speech segment based on the voiceprint feature and the text content of the corrected text segment.
- the voice track data of the initial video data is input into a pre-trained speech model, and the speech model is used to extract voiceprint features corresponding to the human voice portion of the voice track data, i.e., the voiceprint features corresponding to the initial video data.
- the text content of the revised text segment is obtained and input into the pre-trained speech model.
- the speech model generates speech data corresponding to the text content, i.e., a first revised speech segment.
- the speech model extracts the voiceprint features of the human voice in the voice track data of the initial video data, when played, the first revised speech segment can narrate the text content of the revised text segment with a tone and voice similar to the human voice in the initial video data, thereby achieving imitation of the human voice.
- step S1032A include:
- the length of the revised text segment may be the same as that of the changed text segment, that is, for example, the content of the changed text segment is "ABC”, and the revised text segment generated in response to the editing operation is "CBA".
- the length of the revised text segment may also be different from that of the changed text segment, that is, for example, the content of the changed text segment is "ABC”
- the revised text segment generated in response to the editing operation is "CB” or "CBAE".
- step S1033 include:
- Step S1033A Obtain a duration correction coefficient based on the duration ratio of the first corrected speech segment to the first initial speech segment.
- Step S1033B adjusting the playback duration of the first modified voice segment according to the duration correction coefficient to obtain a second modified voice segment having the same playback duration as the first initial voice segment.
- Step S1033C Replace the first initial voice segment with the second revised voice segment to generate revised video data.
- the embodiment of the present disclosure provides two implementation methods for adjusting the playback duration of the first modified voice segment, including:
- Step S1033B-1 Based on the duration correction coefficient, the first modified voice segment is accelerated or decelerated to obtain a second modified voice segment having the same playback duration as the first initial voice segment.
- Step S1033B-2 based on the duration correction coefficient, trimming the speech blank segment in the first corrected speech segment to obtain a second corrected speech segment having the same playback duration as the first initial speech segment.
- the first corrected voice segment is accelerated or decelerated, that is, the overall duration of the first corrected voice segment is compressed or stretched, so that its playback duration is the same as the playback duration of the first initial voice segment.
- the specific implementation method includes extracting or inserting frames at equal intervals on the first corrected voice segment, or filtering the first corrected voice segment, etc. There is no restriction and it can be set as needed.
- FIG 8 is a schematic diagram of generating a second modified voice segment provided by an embodiment of the present disclosure. As shown in Figure 8, for example, referring to the first initial voice segment D1 and the first modified voice segment D2 shown in the figure, the playback duration of the first modified voice segment D2 is L2, which is longer than the playback duration L1 of the first initial voice segment D1.
- the terminal device detects the human voice component in the first modified voice segment D2 and determines that the time periods [t1, t2] and [t3, t4] therein are voice blank segments that do not contain human voice components (i.e., no voice dialogue). Afterwards, the terminal device performs trimming within the intervals [t1, t2] and [t3, t4] until the trimmed playback duration is equal to the playback duration L1 of the first initial voice segment D1, thereby generating the second modified voice segment D3.
- the playback duration of the first corrected speech segment can be adjusted without changing the speaking speed of the human voice in the first corrected speech segment, thereby making the human voice of the generated second corrected speech segment more similar to the human voice of the first initial speech segment and more authentic.
- the content of the video track data in the initial video data may be further modified to achieve the purpose of making the changes in the image content (e.g., the lip movements of the characters) in the video consistent with the changes in the voice content.
- the image content e.g., the lip movements of the characters
- step 103 includes:
- Step S103 - 1 Determine a first playback time period of the initial video data according to the changed text segment in the initial subtitles.
- Step S103 - 2 generating a first revised language segment based on the revised text segment in the revised subtitles, wherein the revised text segment includes sentences in the revised subtitles that have content differences from the original subtitles at corresponding positions.
- Step S103 - 3 According to the first playback time period of the changed text segment, locate the corresponding first initial voice segment from the initial video data.
- Step S103 - 4 Generate a corrected video data segment based on the corrected text segment in the corrected subtitle.
- Step S103-6 Replace the first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment, and replace the first initial video segment based on the revised video data segment to generate revised video data.
- the terminal device inputs the corresponding image processing model based on the modified text segment, and controls the image processing model to generate modified image frames such as lip shape images and facial expression images that match the text content of the modified text segment based on the learned image features, thereby forming a modified video data segment.
- modified image frames such as lip shape images and facial expression images that match the text content of the modified text segment based on the learned image features
- the terminal device inputs the corresponding image processing model based on the modified text segment, and controls the image processing model to generate modified image frames such as lip shape images and facial expression images that match the text content of the modified text segment based on the learned image features, thereby forming a modified video data segment.
- map data representing the object's expression features during the first playback time period of the initial video data is generated; and based on the map data, the modified video data segment is generated.
- the mapping data is an ordered collection of lip-shaped pictures generated by the image processing model. Playing the mapping data in sequence can show the changes in the character's lip shape.
- the corresponding first initial video segment is located from the video track data of the initial video data.
- the specific implementation method is similar to the method of locating the first initial voice segment, and will not be repeated.
- the first initial voice segment is replaced based on the first revised voice segment, and the first initial video segment is replaced based on the revised video data segment to generate the revised video data.
- this step is equivalent to additionally adding modifications to the video image content, such as modifying the character's lip shape, character expression, etc., so that after the voice content of the video is modified, the video picture content is changed accordingly, thereby improving the video quality and picture authenticity of the final generated revised video data.
- step S103 a step of verifying the modification authority and identity may be further included.
- step S103 the following steps may be further included:
- S100A Receive verification voice data input by the user.
- S100B Compare the verification sound data with the voice track data in the initial video data to obtain a verification result.
- the corrected video data corresponding to the initial video data is generated based on the corrected subtitles; if the verification fails, the process ends.
- the terminal device receives the verification sound data input by the user. Specifically, the terminal device responds to the recording command input by the user to collect the user's voice and obtain the verification sound data. Afterwards, the terminal device extracts the voiceprint features of the verification sound data, and then extracts the voiceprint features of the voice track data corresponding to the initial video data, and compares the two. If the voiceprint features of the two are consistent, it means that the user (author) who recorded the initial video data is the same user as the user who is currently modifying the initial video data.
- the subsequent step of generating the corrected video data corresponding to the initial video data based on the corrected subtitles is allowed to be executed; otherwise, the subsequent step of generating the corrected video data is not allowed to be executed, thereby ensuring the security of the video data and preventing the risk of video content being tampered with.
- the revised subtitles are used to realize the modification of the voice content in the initial video data, thereby realizing the rapid synchronization of the video copy content and the video content, without the need to re-record the video, and providing efficiency and accuracy of video modification.
- Step S201 performing speech recognition on the initial video data to obtain initial subtitles corresponding to the initial video data.
- Step S202 In response to the editing operation on the initial subtitles, a revised subtitle is generated.
- Step S203 Obtain the number of pronunciation objects of the changed text segment, where the changed text segment includes sentences in the initial subtitles that have content differences with the revised subtitles at the corresponding positions.
- the number of pronunciation objects represents the number of objects in the initial video data that output the dialogue content of the changed text segment.
- the terminal device can obtain the number of pronunciation objects corresponding to the changed text segment by parsing the text content of the changed text segment, that is, the text content corresponding to the changed text segment is spoken by several characters.
- the number of pronunciation objects is 1, that is, the text content corresponding to the changed text segment is spoken by one character (such as the voice next to it)
- the corrected video data corresponding to the initial video data can be generated directly based on the corrected subtitles.
- the specific implementation method can refer to the implementation method in the embodiment shown in Figure 2, and will not be described in detail.
- the number of pronunciation objects is greater than 1, the subsequent operation can be directly terminated, and a prompt message can be output without modifying the voice content.
- Step S205 If the number of pronunciation objects is greater than 1, the revised text segment corresponding to the changed text segment in the revised subtitles is split based on the pronunciation objects to obtain sub-revised text segments corresponding to each pronunciation object.
- Step S206 Generate corresponding sub-corrected speech segments according to the sub-corrected text segments corresponding to each pronunciation object.
- the terminal device can divide the modified text segment into several sub-intervals, then extract the voiceprint features corresponding to each sub-interval, merge the sub-intervals with similar voiceprint features, thereby separating the speech segments based on the pronunciation objects. Then, the sub-modified text segments are mapped to the corresponding modified text segments in the time interval based on the voice segments corresponding to each character, thereby obtaining the corresponding sub-modified text segments. Since the number of characters is at least two, the sub-modified text segments corresponding to each pronunciation object can be obtained.
- the speech model is input to obtain the corresponding sub-corrected speech segments.
- the corresponding at least two second initial speech segments are located from the speech track data of the initial video data and segmented.
- each sub-corrected speech segment replaces the corresponding second initial speech segment to generate the corrected video data.
- the process of generating the corrected video data based on the sub-corrected text segment in the above steps is similar to the process of generating the corrected video data based on the corrected text segment in the embodiment shown in FIG2.
- the specific implementation method can refer to the introduction of the corresponding part of the embodiment shown in FIG2, and will not be repeated here.
- the corrected text segment is split based on the pronunciation object, and the corresponding sub-corrected voice segments are generated separately, and then the corrected video data is generated by replacing the second initial voice segment.
- FIG11 is a block diagram of the structure of a video processing device provided by an embodiment of the present disclosure.
- the methods described in the above embodiments can be executed by this video processing device, which can be implemented in software and/or hardware and integrated into an electronic device with certain data processing functions.
- Electronic devices may include, but are not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities, such as desktop computers and supercomputers.
- a generating module 31 is configured to perform speech recognition on the initial video data to obtain initial subtitles corresponding to the initial video data;
- a correction module 32 configured to generate a corrected subtitle in response to an editing operation on the initial subtitle
- the processing module 33 when the processing module 33 generates a first revised language segment based on the revised text segment in the revised subtitles, it is specifically used to: obtain the voiceprint features corresponding to the initial video data based on the voice track data of the initial video data; and generate the first revised voice segment based on the voiceprint features and the text content of the revised text segment.
- the processing module 33 when the processing module 33 obtains the voiceprint features corresponding to the initial video data based on the voice track data of the initial video data, it is specifically used to: intercept the key speech segment data in the voice track data of the initial video data, the key speech segment data including the first initial speech segment, and the second initial speech segment of the first length before the first initial speech segment and/or the third initial speech segment of the second length after the first initial speech segment; process the key speech data segments through a pre-trained speech model to generate voiceprint features.
- the processing module 33 is further used to: generate a revised video data segment based on the revised text segment in the revised subtitles; locate the corresponding first initial video segment from the initial video data based on the first playback time period of the changed text segment; when the processing module 33 replaces the first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment, or deletes the first initial voice segment corresponding to the first playback time period in the initial video data to generate the revised video data, the processing module 33 is specifically used to: replace the first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment, and replace the first initial video segment based on the revised video data segment to generate the revised video data.
- the processing module 33 when the processing module 33 replaces the first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment to generate the revised video data, it is specifically used to: obtain a duration correction coefficient based on the duration ratio of the first revised voice segment and the first initial voice segment; adjust the playback duration of the first revised voice segment based on the duration correction coefficient to obtain a second revised voice segment with the same playback duration as the first initial voice segment; and replace the first initial voice segment with the second revised voice segment to generate revised video data.
- the processing module 33 when the processing module 33 adjusts the playback duration of the first corrected voice segment according to the duration correction coefficient to obtain a second corrected voice segment with the same playback duration as the first initial voice segment, it is specifically used to: accelerate or decelerate the first corrected voice segment based on the duration correction coefficient to obtain a second corrected voice segment with the same playback duration as the first initial voice segment; or, based on the duration correction coefficient, trim the voice blank segment in the first corrected voice segment to obtain a second corrected voice segment with the same playback duration as the first initial voice segment.
- the processing module 33 is further used to: obtain the number of pronunciation objects of the changed text segment, the changed text segment includes sentences in the initial subtitles that have content differences with the revised subtitles at the corresponding position, and the number of pronunciation objects represents the number of objects of the dialogue content of the changed text segment output in the initial video data; when the processing module 33 generates revised video data corresponding to the initial video data based on the revised subtitles, it is specifically used to: if the number of pronunciation objects is 1, then generate revised video data corresponding to the initial video data based on the revised subtitles.
- the processing module 33 is further configured to: receive verification sound data input by a user; compare the verification sound data with the voice track data in the initial video data to obtain a verification result;
- the generation module 31, the correction module 32 and the processing module 33 are connected in sequence.
- the video processing device 3 provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.
- the processor 41 executes the computer-executable instructions stored in the memory 42 to implement the video processing method in the embodiments shown in FIG. 2 to FIG. 10 .
- processor 41 and the memory 42 are connected via a bus 43 .
- An embodiment of the present disclosure provides a computer-readable storage medium, which stores computer-executable instructions.
- the computer-executable instructions are executed by a processor, they are used to implement the video processing method provided in any of the embodiments corresponding to Figures 2 to 10 of the present disclosure.
- electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 902 or programs loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of electronic device 900 are also stored in RAM 903. Processing device 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input/output (I/O) interface 905 is also connected to bus 904.
- a processing device e.g., a central processing unit, a graphics processing unit, etc.
- I/O input/output
- an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart.
- the computer program can be downloaded and installed from the network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902.
- the processing device 901 the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
- the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two.
- a computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above.
- Computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
- a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component.
- the computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
- Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages.
- the program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.
- the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
- LAN local area network
- WAN wide area network
- Internet service provider e.g., AT&T, MCI, Sprint, EarthLink, MSN, GTE, etc.
- each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function.
- the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
- each box in the block diagram and/or flowchart, and the combination of the boxes in the block diagram and/or flowchart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
- exemplary types of hardware logic components include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
- FPGAs field programmable gate arrays
- ASICs application specific integrated circuits
- ASSPs application specific standard products
- SOCs systems on chip
- CPLDs complex programmable logic devices
- a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment.
- a machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium.
- a machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing.
- generating the revised video data corresponding to the initial video data based on the revised subtitles includes: determining a first playback time period of the initial video data based on a changed text segment in the initial subtitles, the changed text segment including sentences in the initial subtitles that have content differences from the revised subtitles at a corresponding position; generating a first revised language segment based on the revised text segment in the revised subtitles, the revised text segment including sentences in the revised subtitles that have content differences from the initial subtitles at a corresponding position; replacing a first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment, or deleting the first initial voice segment corresponding to the first playback time period in the initial video data, to generate the revised video data.
- generating a first revised language segment based on the revised text segment in the revised subtitles includes: obtaining voiceprint features corresponding to the initial video data based on the voice track data of the initial video data; and generating the first revised voice segment based on the voiceprint features and the text content of the revised text segment.
- the method further includes: generating a revised video data segment based on the revised text segment in the revised subtitles; locating the corresponding first initial video segment from the initial video data based on the first playback time period of the changed text segment; replacing the first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment, or deleting the first initial voice segment corresponding to the first playback time period in the initial video data to generate the revised video data, including: replacing the first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment, and replacing the first initial video segment based on the revised video data segment to generate the revised video data.
- generating a corrected video data segment based on the corrected text segment in the corrected subtitles includes: generating, based on the text content of the corrected text segment, map data representing the expression features of an object within the first playback time period of the initial video data; and generating the corrected video data segment based on the map data.
- the replacing of the first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment to generate the revised video data includes: obtaining a duration correction coefficient based on a duration ratio between the first revised voice segment and the first initial voice segment; adjusting the playback duration of the first revised voice segment based on the duration correction coefficient to obtain a second revised voice segment having the same playback duration as the first initial voice segment; and replacing the first initial voice segment with the second revised voice segment to generate the revised video data.
- the playback duration of the first modified voice segment is adjusted according to the duration correction coefficient to obtain a second modified voice segment with the same playback duration as the first initial voice segment, including: accelerating or decelerating the first modified voice segment based on the duration correction coefficient to obtain a second modified voice segment with the same playback duration as the first initial voice segment; or trimming the voice blank segment in the first modified voice segment based on the duration correction coefficient to obtain a second modified voice segment with the same playback duration as the first initial voice segment.
- the further step includes: obtaining the number of pronunciation objects of the changed text segment, the changed text segment including sentences in the initial subtitles that have content differences from the revised subtitles at the corresponding position, and the number of pronunciation objects represents the number of objects in the initial video data that output the dialogue content of the changed text segment; generating revised video data corresponding to the initial video data based on the revised subtitles, including: if the number of pronunciation objects is 1, generating revised video data corresponding to the initial video data based on the revised subtitles; if the number of pronunciation objects is greater than 1, splitting the revised text segment corresponding to the changed text segment in the revised subtitles based on the pronunciation objects to obtain sub-revised text segments corresponding to each of the pronunciation objects; generating corresponding sub-revised voice segments based on the sub-revised text segments corresponding to each of the pronunciation objects; locating corresponding at least two second initial voice segments from the initial video data based on the
- a generating module configured to perform speech recognition on the initial video data to obtain initial subtitles corresponding to the initial video data
- a correction module configured to generate a corrected subtitle in response to an editing operation on the initial subtitle
- a processing module is configured to generate modified video data corresponding to the initial video data based on the modified subtitles, wherein the speech content in the modified video data matches the subtitle content of the modified subtitles.
- the processing module is specifically used to: determine a first playback time period of the initial video data based on a changed text segment in the initial subtitles, the changed text segment including sentences in the initial subtitles that have content differences from the revised subtitles at the corresponding position; generate a first revised language segment based on the revised text segment in the revised subtitles, the revised text segment including sentences in the revised subtitles that have content differences from the initial subtitles at the corresponding position; replace a first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment, or delete the first initial voice segment corresponding to the first playback time period in the initial video data to generate revised video data.
- the processing module when it obtains the voiceprint features corresponding to the initial video data based on the voice track data of the initial video data, it is specifically used to: intercept the key speech segment data in the voice track data of the initial video data, the key speech segment data including the first initial speech segment, and the second initial speech segment of the first length before the first initial speech segment and/or the third initial speech segment of the second length after the first initial speech segment; process the key speech data segments through a pre-trained speech model to generate voiceprint features.
- the processing module when the processing module generates a corrected video data segment based on the corrected text segment in the corrected subtitles, it is specifically used to: generate map data representing the expression features of the object in the first playback time period of the initial video data based on the text content of the corrected text segment; and generate the corrected video data segment based on the map data.
- the processing module when the processing module replaces the first initial voice segment corresponding to the first playback time period in the initial video data with the first revised voice segment to generate the revised video data, it is specifically used to: obtain a duration correction coefficient based on the duration ratio of the first revised voice segment and the first initial voice segment; adjust the playback duration of the first revised voice segment based on the duration correction coefficient to obtain a second revised voice segment with the same playback duration as the first initial voice segment; and replace the first initial voice segment with the second revised voice segment to generate the revised video data.
- the processing module when the processing module adjusts the playback duration of the first corrected voice segment according to the duration correction coefficient to obtain the second corrected voice segment with the same playback duration as the first initial voice segment, it is specifically used to: accelerate or decelerate the first corrected voice segment based on the duration correction coefficient to obtain the second corrected voice segment with the same playback duration as the first initial voice segment; or, based on the duration correction coefficient, trim the voice blank segment in the first corrected voice segment to obtain the second corrected voice segment with the same playback duration as the first initial voice segment.
- the processing module is further used to: obtain the number of pronunciation objects of the changed text segment, the changed text segment includes sentences in the initial subtitles that have content differences with the revised subtitles at the corresponding position, and the number of pronunciation objects represents the number of objects in the initial video data that output the dialogue content of the changed text segment; when the processing module generates revised video data corresponding to the initial video data based on the revised subtitles, the processing module is specifically used to: if the number of pronunciation objects is 1, then generate revised video data corresponding to the initial video data based on the revised subtitles.
- the processing module is further used to: receive verification sound data input by a user; compare the verification sound data with the voice track data in the initial video data to obtain a verification result; when the processing module generates the corrected video data corresponding to the initial video data based on the corrected subtitles, the processing module is specifically used to: based on the verification result, if the verification passes, generate the corrected video data corresponding to the initial video data based on the corrected subtitles.
- an electronic device comprising: at least one processor and a memory;
- the memory stores computer-executable instructions
- the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the video processing method described in the first aspect and various possible designs of the first aspect.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Human Computer Interaction (AREA)
- Television Signal Processing For Recording (AREA)
Abstract
本公开实施例提供一种视频处理方法、装置、电子设备及存储介质,该视频处理方法通过对初始视频数据进行语音识别,得到初始视频数据对应的初始字幕;响应于针对初始字幕的编辑操作,生成修正字幕;基于修正字幕,生成初始视频数据对应的修正视频数据,其中,修正视频数据中的语音内容与修正字幕的字幕内容相匹配。通过响应针对编辑操作的编辑操作,在实现针对初始字幕的修改的同时,利用修正字幕来实现对初始视频数据中的语音内容的修正,从而实现视频文案内容与视频内容的快速同步,无需对视频进行重新录制,提供视频修改的效率和准确性。
Description
本申请要求于2024年3月29日递交的中国专利申请第202410382907.2号的优先权,在此全文引用上述中国专利申请公开的内容以作为本申请的一部分。
本公开实施例涉及一种视频处理方法、装置、电子设备及存储介质。
在视频编辑场景中,用户通过使用视频编辑应用所提供的各类视频编辑功能,进行视频的录制和制作,最终生成满足需求的视频作品。在该过程中,视频文案内容通常是最先确定的,之后,用户会基于视频文案内容来进行视频(音频)的录制,来形成视频文件或草稿。
然而,当用户后期需要对视频文案内容进行修改时,通常只能利用应用内的草稿功能,对需要修改的语音内容进行重新录制,进而完成对视频的修改。
因此,在对视频中的语音内容进行修改时,存在修改效率低、修改效果差的问题。
本公开实施例提供一种视频处理方法、装置、电子设备及存储介质,以克服在对视频中的语音内容进行修改时,存在的修改效率低、修改效果差的问题。
第一方面,本公开实施例提供一种视频处理方法,包括:
对初始视频数据进行语音识别,得到所述初始视频数据对应的初始字幕;响应于针对所述初始字幕的编辑操作,生成修正字幕;基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,其中,所述修正视频数据中的语音内容与所述修正字幕的字幕内容相匹配。
第二方面,本公开实施例提供一种视频处理方法装置,包括:
生成模块,用于对初始视频数据进行语音识别,得到所述初始视频数据对应的初始字幕;
修正模块,用于响应于针对所述初始字幕的编辑操作,生成修正字幕;
处理模块,用于基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,其中,所述修正视频数据中的语音内容与所述修正字幕的字幕内容相匹配。
第三方面,本公开实施例提供一种电子设备,包括:处理器和存储器;
所述存储器存储计算机执行指令;
所述处理器执行所述存储器存储的计算机执行指令,使得所述至少一个处理器执行如上第一方面以及第一方面各种可能的设计所述的视频处理方法。
第四方面,本公开实施例提供一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如上第一方面以及第一方面各种可能的设计所述的视频处理方法。
第五方面,本公开实施例提供一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如上第一方面以及第一方面各种可能的设计所述的视频处理方法。
为了更清楚地说明本公开实施例中的技术方案,下面将对实施例描述中所需要使用的附图作一简单地介绍,显而易见地,下面描述中的附图是本公开的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本公开实施例提供的视频处理方法的一种应用场景图;
图2为本公开实施例提供的视频处理方法的流程示意图一;
图3为图2所示实施例中步骤S103的一种可能的实现方式的流程图;
图4为本公开实施例提供的一种生成修正视频数据的过程示意图;
图5为图3所示实施例中步骤S1032的具体实现方式的流程图;
图6为本公开实施例提供的一种语音轨道数据的截取示意图;
图7为图3所示实施例中步骤S1033的具体实现方式的流程图;
图8为本公开实施例提供的一种生成第二修正语音段的示意图;
图9为图2所示实施例中步骤S103的另一种可能的实现方式的流程图;
图10为本公开实施例提供的视频处理方法的流程示意图二;
图11为本公开实施例提供的视频处理装置的结构框图;
图12为本公开实施例提供的一种电子设备的结构示意图;以及
图13为本公开实施例提供的电子设备的硬件结构示意图。
为使本公开实施例的目的、技术方案和优点更加清楚,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本公开一部分实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
需要说明的是,本公开所涉及的用户信息(包括但不限于用户设备信息、用户个人信息等)和数据(包括但不限于用于分析的数据、存储的数据、展示的数据等),均为经用户授权或者经过各方充分授权的信息和数据,并且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准,并提供有相应的操作入口,供用户选择授权或者拒绝。
下面对本公开实施例的应用场景进行解释:
图1为本公开实施例提供的视频处理方法的一种应用场景图,本公开实施例提供的视频处理方法,可以应用于具有视频编辑功能的应用程序(APP,Application)中,例如短视频类应用程序、视频编辑应用程序等。更具体地,可以应用于对视频草稿中的语音内容进行修改的应用场景中。本实施例的执行主体,可以为运行上述具有视频编辑功能的应用程序的终端设备,也可以为部署上述应用程序所对应的服务端的服务器,或者其他起到类似功能的电子设备。
其中,在一些实施例中,终端设备或服务器可以通过运行各种计算机可执行指令或计算机程序来实现本公开实施例提供的视频处理方法。例如,计算机可执行指令可以是程序级的命令、机器指令或软件指令。计算机程序可以是操作系统中的原生程序或软件模块;可以是本地应用程序,即需要在操作系统中安装才能运行的程序,也可以是嵌入至任意APP中的小程序,即基于浏览器环境运行的程序。综上,上述的计算机可执行指令可以是任意形式的指令,上述计算机程序可以是任意形式的应用程序、模块或插件,具体实现形式可以根据需要配置。进一步地,在一些实施例中,服务器可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云存储、云通信、云数据库、云计算、云函数、网络服务、中间件服务、域名服务、安全服务、内容分发网络(Content DeliveryNetwork,CDN)、以及大数据和人工智能平台等基础云计算服务的云服务器,其中,云服务可以是交互处理服务,供终端设备进行调用。
参考图1中所示,以终端设备为例,终端设备在响应用户的第一次编辑操作(可以由多个具体的子操作构成,例如录制视频操作、录制语音操作、视频剪裁操作、插入特效及滤镜操作等),完成视频的录制后,形成视频数据(视频草稿),视频数据包括多个编辑轨道,例如图中所示的Track_1、Track_2、Track_3等,其中,编辑轨道Track_1为视频轨道数据,用于设置图像媒体;编辑轨道Track_2为语音轨道数据,用于设置音频媒体;编辑轨道Track_3为特效轨道数据,用于设置特效。之后,当用户需要对该视频数据中语音内容做进一步修改,例如修正口误、调整名词内容等,需要向终端设备实施第二次编辑操作,终端设备通过响应该第二次编辑操作,完成对视频草稿的语音内容的修改,也即对编辑轨道Track_1的修改。
例如,针对语音内容的第二次编辑操作,通常的实现方式是对视频的音频轨道数据进行重新录制,这导致编辑过程低效、耗时的问题,同时,由于需要对重新录制的音频轨道数据进行人工对齐,还会进一步导致音频重叠等问题,影响视频的音频质量。
本公开实施例提供一种视频处理方法以解决上述问题。
参考图2,图2为本公开实施例提供的视频处理方法的流程示意图一。本实施例的方法可以应用在终端设备、服务器或其他电子设备中,以终端设备为例,该视频处理方法包括:
步骤S101:对初始视频数据进行语音识别,得到初始视频数据对应的初始字幕。
步骤S102:响应于针对初始字幕的编辑操作,生成修正字幕。
示例性地,参考图1所示的应用场景示意图,终端设备在运行例如视频编辑应用程序后,通过视频编辑应用程序加载初始视频数据,其中,初始视频数据可以是一种程序内数据,例如通过上述视频编辑应用程序创建的视频草稿;之后,通过对初始视频数据中的语音轨道数据进行解析,可以自动生成初始视频数据的语音内容对应的文字描述,即初始字幕。该生成初始视频数据的语音内容对应的初始字幕的过程,可以基于用户操作而触发,例如用户点击“生成字幕”控件后,执行上述过程并生成初始字幕,也可以是在加载初始视频数据后,自动生成初始字幕,并基于具体的设计对其进行显示或不显示。对于基于视频中的语音内容,解析生成对应的字幕的实现方式,此次不做具体介绍。
进一步地,响应于用户施加的、针对初始字幕的编辑操作,终端设备在视频编辑应用内完成对初始字幕的修改,生成修正字幕。例如,初始字幕包括由若干个字符和标点组成的文本,例如表示为“XXXABCXXX”(其中X表示占位,其内容可以为任意值)。用户针对该初始字幕中的文本段“ABC”施加编辑操作,将其修改为“CBA”,则生成的修正字幕为“XXXCBAXXX”。终端设备响应上述编辑操作后,首先根据编辑操作所输入的内容,对初始字幕进行文本层面的修改,即常规的文本修改,并形成基于文本的修正字幕。该过程中还可以进一步包括例如对编辑操作所输入内容的字符合法性检验、输入字数检测等步骤,可以根据需要设置。
步骤S103:基于修正字幕,生成初始视频数据对应的修正视频数据,其中,修正视频数据中的语音内容与修正字幕的字幕内容相匹配。
在得到修正字幕后,也即完成对初始字幕的编辑过程后,终端设备根据所生成的修正字幕的文字内容,对初始视频数据中的语音内容进行修正,使初始视频数据中的语音内容与修正字幕的字幕内容相匹配,具体的,可以通过预训练的语音模型,来实现“文字-声音”的转换,实现对初始视频数据中被修正部分的语音段的替换,从而生成修正视频数据,其中,修正视频数据即对初始视频数据的调整结果,因此,与初始视频数据类似,也是一种程序内数据,例如视频草稿。之后,可以根据需要,进一步地将该修正视频数据生成对应的视频文件。
其中,示例性地,如图3所示,步骤103的一种可能的实现步骤包括:
步骤S1031:根据初始字幕中的变更文本段,确定初始视频数据的第一播放时间段,变更文本段包括初始字幕中,与对应位置的修正字幕存在内容差异的语句。
步骤S1032:根据修正字幕中的修正文本段,生成第一修正语言段,修正文本段包括修正字幕中,与对应位置的初始字幕存在内容差异的语句。
步骤S1033:将初始视频数据中的第一播放时间段对应的第一初始语音段替换为第一修正语音段,或者,删除初始视频数据中的第一播放时间段对应的第一初始语音段,以生成修正视频数据。
示例性地,首先,获取变更文本段,变更文本段包括初始字幕中,与对应位置的修正字幕存在内容差异的语句,其中,变更文本段即初始字幕中由于响应编辑操作而发生改变的部分,其中包括由于响应编辑操作而从初始字幕中删除、添加、修改字符的部分,即变更文本段是初始字幕的一部分。在一种可能的实现方式中,在终端设备接收到编辑操作后,由终端设备根据编辑操作所指示的编辑位置,对初始字幕进行标记,从而得到该变更文本段。之后,终端设备根据变更文本段的位置,确定修正字幕中与变更文本段对应的修正文本段,并将该修正文本段输入语音模型,来生成对应的语音段,即第一修正语音段。
另一方面,视频的字幕对应有时间戳序列,该时间戳序列用于表征字幕中各文字的播放时间戳,从而通过时间戳序列控制字幕和对应的视频内容(视频图像和视频语音)能够同步播放。本实施例中的初始字幕也对应有时间戳序列,因此,在确定变更文本段的同时,可以基于该时间戳序列,得到变更文本段对应的第一播放时间段。之后,基于该第一播放时间段,从初始视频数据中截取出对应的语音数据,即第一初始语音段。再之后,利用上述步骤中得到的第一修正语音段对该第一初始语音段进行替换,生成修正视频数据。
图4为本公开实施例提供的一种生成修正视频数据的过程示意图,下面结合图4对上述过程进行更详细介绍,如图4所示,示例性地,视频数据(初始视频数据、修正视频数据)由多个轨道数据构成,其中例如包括语音轨道数据、视频轨道数据、特效轨道数据等。首先,基于初始视频数据的语音轨道数据L1,生成初始字幕;之后,基于针对初始字幕中的变更文本段(图中示为T1-T2段,文本内容为“我是说”)的编辑操作,生成修正字幕,修正字幕中包含与变更文本段对应的修正文本段(图中示为T3-T4段,文本内容为“我说”),再之后,基于修正文本段,生成对应的第一修正语音段D2;再基于中变更文本段对应的第一播放时间段,从初始视频数据的语音轨道数据中,截取出第一初始语音段D1;最后,将第一修正语音段D2替换到第一初始语音段D1所在位置,也即变更文本段对应的第一播放时间段,得到调整后的语音轨道数据L2,基于该调整后的语音轨道数据L2,以及其他的轨道数据,得到修正视频数据。
其中,示例性地,第一修正语音段可以是通过语音模型生成的“AI语音”,具体的,如图5所示,步骤S1032的具体实现方式包括:
步骤S1032A:根据初始视频数据的语音轨道数据,得到初始视频数据对应的声纹特征。
步骤S1032B:根据声纹特征和修正文本段的文本内容,生成第一修正语音段。
示例性地,一方面,将初始视频数据的语音轨道数据,输入预训练的语音模型,通过该语音模型提取语音轨道数据中人声部分对应的声纹特征,即初始视频数据对应的声纹特征。另一方面,获取修正文本段的文本内容,并将该修正文本段的文本内容输入预训练的语音模型,之后,语音模型基于上述步骤中所提取到的声纹特征,来生成文本内容对应的语音数据,即第一修正语音段,由于语音模型提取了初始视频数据的语音轨道数据中的人声的声纹特征,因此,该第一修正语音段在被播放时,能够以与初始视频数据中的人声相似的语调、声色来讲述上述修正文本段的文本内容,从而实现对人声的模仿。
进一步地,针对上述生成第一修正语音段的实现步骤,在获取初始视频数据的语音轨道数据之后,一种可能的实现方式中,使将初始视频数据全部的语音轨道数据,输入语音模型,使语音模型基于全部的语音轨道数据来进行声纹特征提取。而在另一种可能的实现方式中,可以对初始视频数据的部分语音轨道数据进行分割,并送入语音模型,使该语音模型基于部分语音轨道数据完成声纹特征提取。具体的,步骤S1032A的具体实现步骤包括:
步骤S1032A-1:截取初始视频数据的语音轨道数据中的关键语音段数据,关键语音段数据包括第一初始语音段,以及第一初始语音段之前第一长度的第二初始语音段和/或第一初始语音段之后第二长度的第三初始语音段。
步骤S1032A-2:通过预训练的语音模型对关键语音数据段进行处理,生成声纹特征。
示例性地,图6为本公开实施例提供的一种语音轨道数据的截取示意图,参考图6所示,关键语音段数据包括位于第一初始语音段(图中示为A-B段),以及第一初始语音段之前长度为L1的第二初始语音段(图中示为C-A段),和第一初始语音段之后第二长度L2的第三初始语音段(图中示为B-D段)。之后,从语音轨道数据中截取上述关键语音段数据(即C-D段),并送入语音模型进行处理,生成声纹特征。
本实施例中,通过获取第一初始语音段之前的第一初始语音段和之后的第二初始语音段,作为关键语音段数据进行声纹特征提取,相比使用全量的语音轨道数据进行声纹特征提取的方案,减少了数据量,可以提高生成第一修正语音段的速度,同时,由于关键语音段数据与第一初始语音段的内容相关性高,基于关键语音段数据进行声纹提取,可以提高生成的第一修正语音段与第一初始语音段的声音相似度。当然,在其他可能的实现方式中,也可以仅利用第一初始语音段和第二初始语音段(即C-B段),或者,仅利用第一初始语音段和第三初始语音段(即A-D段),来作为关键语音段数据进行声纹特征提取,可以根据需要设置。
进一步地,基于编辑操作的具体实现方式,修正文本段可以与变更文本段的长度相同,即例如,变更文本段的内容为“ABC”,响应编辑操作而生成的修正文本段为“CBA”。在其他可能的实现方式中,修正文本段也可以与变更文本段的长度不同,即例如,变更文本段的内容为“ABC”,响应编辑操作而生成的修正文本段为“CB”或“CBAE”。因此,相应的,在一种可能的实现方式中,当修正文本段的长度与变更文本段的长度不一致时,基于修正文本段所生成的第一修正语音段的播放时长,则可能与变更文本段对应的第一初始语音段的播放时长不一致,从而在基于第一修正语音段替换第一初始语音段时,出现声音重叠或空白的问题。
为解决上述文本,在本公开提供的一种可能的实施例中,如图7所示,步骤S1033的实现步骤包括:
步骤S1033A:根据第一修正语音段和第一初始语音段的时长比,得到时长修正系数。
步骤S1033B:根据时长修正系数,对第一修正语音段的播放时长进行调整,得到与第一初始语音段的播放时长相同的第二修正语音段。
步骤S1033C:利用第二修正语音段替换第一初始语音段,生成修正视频数据。
示例性地,本实施例中,终端设备在得到第一修正语音段后,首先获取第一修正语音段的播放时长与第一初始语音段的播放时长(即变更文本段的第一播放时间段),之后,根据第一修正语音段和第一初始语音段的时长比,得到时长修正系数,具体的,例如,第一修正语音段的播放时长为3秒、第一初始语音段的播放时长为2秒,则根据时长比,得到对应的时长修正系数为1.5。之后,基于该时长修正系数,对第一修正语音段的播放时长进行调整,得到与第一初始语音段的播放时长相同的第二修正语音段,例如,将第一修正语音段的播放时长同样压缩为2秒,生成第一初始语音段的播放时长保持一致的第二修正语音段,之后利用第二修正语音段对第一初始语音段进行替换,生成修正视频数据。
其中,进一步地,本公开实施例提供了两种对第一修正语音段的播放时长进行调整的实现方式,包括:
步骤S1033B-1:基于时长修正系数,对第一修正语音段进行加速或减速处理,得到与第一初始语音段的播放时长相同的第二修正语音段。
步骤S1033B-2:基于时长修正系数,对第一修正语音段中的语音空白段进行剪裁,得到与第一初始语音段的播放时长相同的第二修正语音段。
其中,基于时长修正系数,对第一修正语音段进行加速或减速处理,即对第一修正语音段的整体时长进行压缩或拉伸,从而使其播放时长与第一初始语音段的播放时长相同,具体实现方式包括对第一修正语音段进行等间隔的抽帧或插帧,或者对第一修正语音段进行滤波等,不做限制,可根据需要设置。
而基于时长修正系数,对第一修正语音段中的语音空白段进行剪裁,是指在第一修正语音段的播放时长大于第一初始语音段的情况下,对第一修正语音段中不描述语音内容的音频段进行剪裁,从而生成第二修正语音段。图8为本公开实施例提供的一种生成第二修正语音段的示意图,如图8所示,示例性地,参考图中所示的第一初始语音段D1和第一修正语音段D2,其中,第一修正语音段D2的播放时长为L2,大于第一初始语音段D1播放时长为L1,此种情形下,终端设备通过对第一修正语音段D2中的人声成分进行检测,确定出其中的[t1,t2]、[t3,t4]时间段为不含人声成分(即没有语音对白)的语音空白段,之后,终端设备在[t1,t2]、[t3,t4]区间内,进行剪裁,直至剪裁后的播放时长与第一初始语音段D1播放时长L1相等,从而生成第二修正语音段D3。
本实施例中,通过基于时长修正系数,对第一修正语音段的语音空白段进行剪裁,可以在不改变第一修正语音段中人声语速的情况下,完成对第一修正语音段的播放时长的调整,从而使生成的第二修正语音段的人声与第一初始语音段的人声更加相似、具体更高的真实性。
进一步地,在图3所示实施例的基础上,还可以进一步修改初始视频数据中视频轨道数据上的内容,从而实现视频中图像内容(例如人物口型)变化与语音内容变化一致的目的,示例性地,如图9所示,步骤103的另一种可能的实现步骤包括:
步骤S103-1:根据初始字幕中的变更文本段,确定初始视频数据的第一播放时间段。
步骤S103-2:根据修正字幕中的修正文本段,生成第一修正语言段,修正文本段包括修正字幕中,与对应位置的初始字幕存在内容差异的语句。
步骤S103-3:根据变更文本段的第一播放时间段,从初始视频数据中定位对应的第一初始语音段。
步骤S103-4:根据修正字幕中的修正文本段,生成修正视频数据段。
步骤S103-5:根据变更文本段的第一播放时间段,从初始视频数据中定位对应的第一初始视频段。
步骤S103-6:将初始视频数据中的第一播放时间段对应的第一初始语音段替换为第一修正语音段、基于修正视频数据段替换第一初始视频段,以生成修正视频数据。
示例性地,其中,步骤S103-1至步骤S103-2的实现方式,与图3所示实施例中步骤S1031至步骤S1032的实现方式相同,此处不再赘述。终端设备在基于编辑操作,获取变更文本段之后,一方面生成第一修正语音段,另一方面,还对视频轨道数据上的内容进行修改,生成修正视频数据段,具体的,终端设备一方面将初始视频数据的视频轨道数据输入图像处理模型,使该图像处理模型学习图像特征,具体的,视频轨道数据中的视频帧中的人物面部特征;另一方面,例如基于修正文本段,输入对应的图像处理模型,控制图像处理模型基于上述学习到的图像特征,生成与修正文本段的文本内容匹配的口型图像、面部表情图像等修正图像帧,从而构成修正视频数据段。例如,根据修正文本段的文本内容,生成表征对象表情特征贴图数据初始视频数据的第一播放时间段内的对象表情特征的贴图数据;根据贴图数据,生成修正视频数据段。其中,贴图数据即由图像处理模型生成的、有序的口型图片集合。次序播放贴图数据,可以表现人物口型的变化。之后,基于变更文本段的第一播放时间段,从初始视频数据的视频轨道数据中,定位对应的第一初始视频段,具体实现方式与定位第一初始语音段的方式类似,不做重复介绍。再之后,基于第一修正语音段替换第一初始语音段、基于修正视频数据段替换第一初始视频段,生成修正视频数据,该步骤相比于图3所示实施例中的步骤S1033,相当于额外增加了对视频图像内容的修改,例如修改人物口型、人物表情等,从而在视频的语音内容发生修改后,使视频画面内容随之改变,提高最终生成的修正视频数据的视频质量和画面真实性。
进一步地,可选地,在步骤S103执行之前,还可以包括对修改权限和身份进行验证的步骤,示例性地,在步骤S103之前,还包括:
S100A:接收用户输入的验证声音数据。
S100B:对比验证声音数据和初始视频数据中的语音轨道数据,得到验证结果。
相应的,步骤S103的具体实现方式包括:
根据验证结果,若验证通过,则基于修正字幕,生成初始视频数据对应的修正视频数据;若验证不通过,则结束。
示例性地,首先,终端设备接收用户输入的验证声音数据,具体的,终端设备响应用户输入的录音指令,来采集用户的声音,得到验证声音数据,之后,终端设备提取该验证声音数据的声纹特征,再提取初始视频数据对应的语音轨道数据的声纹特征,二者进行对比,若二者的声纹特征一致,说明录制初始视频数据的用户(作者),与当前修改该初始视频数据的用户是同一用户,因此,允许执行后续基于修正字幕,生成初始视频数据对应的修正视频数据的步骤;否则,则不允许执行后续生成修正视频数据的步骤,从而保证视频数据的安全性,防止视频内容被篡改的风险。
需要说明的是,上述步骤S100A-S100B,可以在步骤S101之前执行,也可以在步骤S101之后、步骤S102之前执行,还可以在步骤S102之后、步骤S103之前执行,具体不做限制。
本实施例中,通过响应针对编辑操作的编辑操作,在实现针对初始字幕的修改的同时,利用修正字幕来实现对初始视频数据中的语音内容的修正,从而实现视频文案内容与视频内容的快速同步,无需对视频进行重新录制,提供视频修改的效率和准确性。
参考图10,图10为本公开实施例提供的视频处理方法的流程示意图二。本实施例在图2所示实施例的基础上,进一步对步骤S103进行细化,并增加了身份校验的步骤,该视频处理方法包括:
步骤S201:对初始视频数据进行语音识别,得到初始视频数据对应的初始字幕。
步骤S202:响应于针对初始字幕的编辑操作,生成修正字幕。
步骤S203:获取变更文本段的发音对象数量,变更文本段包括初始字幕中,与对应位置的修正字幕存在内容差异的语句,发音对象数量表征初始视频数据中输出变更文本段的对话内容的对象数量。
步骤S204:若发音对象数量为1,则基于修正字幕,生成初始视频数据对应的修正视频数据。
示例性地,在终端设备接收到编辑操作,从而确定初始字幕中待修改的变更文本段之后,终端设备通过对变更文本段的文本内容进行解析,可以得到变更文本段所对应的发音对象数量,即变更文本段所对应的文本内容,是由几个角色说出来的。当发音对象数量为1,即变更文本段对应的文本内容,是由一个角色说出(例如旁边语音),则可以直接基于修正字幕,生成初始视频数据对应的修正视频数据,具体实现方式可以参考图2所示实施例中的实现方式,具体不再赘述。当然,在另一种可能的实现方式中,当发音对象数量大于为1,可以直接结束后续操作,并输出提示信息,不对语音内容进行修改。
步骤S205:若发音对象数量大于1,则基于发音对象,对修正字幕中与变更文本段对应的修正文本段进行拆分,得到各发音对象对应的子修正文本段。
步骤S206:根据各发音对象对应的子修正文本段,生成对应的子修正语音段。
步骤S207:根据变更文本段中各发音对象对应的第一播放时间段,从初始视频数据中定位对应的至少两个第二初始语音段。
步骤S208:基于发音对象,将第二初始语音段替换为对应的子修正语音段,生成修正视频数据。
而在另一种情况下,若发音对象数量大于1,即变更文本段对应的文本内容,是由两个以上的角色说出,此时,由于不同角色之间的声纹特征不同,因此,无法直接基于图2所示实施例中提供的方法,来合成修正字幕对应的第一修正语音段,进而生成修正视频数据。为解决上述问题,在本实施例中,示例性地,当终端设备检测到发音对象数量大于1时,首先,则基于发音对象,对修正字幕中与变更文本段对应的修正文本段进行拆分,得到各发音对象对应的子修正文本段,具体实现上,终端设备可以通过将修正文本段划分为若干个子区间,之后通过提取各子区间对应的声纹特征,对声纹特征类似的子区间进行合并,从而基于发音对象的语音段分离,之后,在基于各角色对应的语音段的时间区间,映射至对应的修正文本段,即可得到对应的子修正文本段。由于角色数量至少2个,因此可以得到各发音对象对应的子修正文本段。
之后,针对上述两个子修正文本段,输入语音模型,得到对应的子修正语音段,再之后,利用子修正文本段的第一播放时间段,也即变更文本段中各发音对象对应的第一播放时间段,从初始视频数据的语音轨道数据中,定位对应的至少两个第二初始语音段,并进行分割,最后将各子修正语音段替换对应的第二初始语音段,以生成修正视频数据。上述步骤中基于子修正文本段生成修正视频数据的过程,与图2所示实施例中基于修正文本段,生成修正视频数据的过程类似,具体实现方式可参考图2所示实施例中对应部分的介绍,此次不再赘述。
本实施例中,在变更文本段对应的发音对象数量大于1的情况下,通过基于发音对象对修正文本段进行拆分,并单独生成对应的子修正语音段,进而通过替换第二初始语音段,生成修正视频数据,解决了多发音对象下无法提取准确声纹特征,进而无法生成匹配的修正语音段的问题,提高修正多角色对话内容的视频的语音内容的效果。
本实施例中,步骤S201-步骤S202的实现方式与本公开图2所示实施例中的步骤S101-步骤S102的实现方式相同,在此不再一一赘述。
对应于上文实施例的视频处理方法,图11为本公开实施例提供的视频处理装置的结构框图。上述实施例所介绍的方法,可以由该视频处理装置来执行,该装置可以由软件和/或硬件的方式实现,该装置可以集成在有一定数据处理功能的电子设备中。其中,电子设备可以包括但不限于具备大数据处理能力的移动终端,以及台式计算机、超级计算机等具备大数据处理能力的固定终端。
为了便于说明,仅示出了与本公开实施例相关的部分。参照图11,视频处理装置3包括:
生成模块31,用于对初始视频数据进行语音识别,得到初始视频数据对应的初始字幕;
修正模块32,用于响应于针对初始字幕的编辑操作,生成修正字幕;
处理模块33,用于基于修正字幕,生成初始视频数据对应的修正视频数据,其中,修正视频数据中的语音内容与修正字幕的字幕内容相匹配。
根据本公开的一个或多个实施例,处理模块33,具体用于:根据初始字幕中的变更文本段,确定初始视频数据的第一播放时间段,变更文本段包括初始字幕中,与对应位置的修正字幕存在内容差异的语句;根据修正字幕中的修正文本段,生成第一修正语言段,修正文本段包括修正字幕中,与对应位置的初始字幕存在内容差异的语句;将初始视频数据中的第一播放时间段对应的第一初始语音段替换为第一修正语音段,或者,删除初始视频数据中的第一播放时间段对应的第一初始语音段,以生成修正视频数据。
根据本公开的一个或多个实施例,处理模块33在根据修正字幕中的修正文本段,生成第一修正语言段时,具体用于:根据初始视频数据的语音轨道数据,得到初始视频数据对应的声纹特征;根据声纹特征和修正文本段的文本内容,生成第一修正语音段。
根据本公开的一个或多个实施例,处理模块33在根据初始视频数据的语音轨道数据,得到初始视频数据对应的声纹特征时,具体用于:截取初始视频数据的语音轨道数据中的关键语音段数据,关键语音段数据包括第一初始语音段,以及第一初始语音段之前第一长度的第二初始语音段和/或第一初始语音段之后第二长度的第三初始语音段;通过预训练的语音模型对关键语音数据段进行处理,生成声纹特征。
根据本公开的一个或多个实施例,处理模块33,还用于:根据修正字幕中的修正文本段,生成修正视频数据段;根据变更文本段的第一播放时间段,从初始视频数据中定位对应的第一初始视频段;处理模块33在将初始视频数据中的第一播放时间段对应的第一初始语音段替换为第一修正语音段,或者,删除初始视频数据中的第一播放时间段对应的第一初始语音段,以生成修正视频数据时,具体用于:将初始视频数据中的第一播放时间段对应的第一初始语音段替换为第一修正语音段、基于修正视频数据段替换第一初始视频段,以生成修正视频数据。
根据本公开的一个或多个实施例,处理模块33在根据修正字幕中的修正文本段,生成修正视频数据段时,具体用于:根据修正文本段的文本内容,生成表征初始视频数据的第一播放时间段内的对象表情特征的贴图数据;根据贴图数据,生成修正视频数据段。
根据本公开的一个或多个实施例,处理模块33在所述将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段,以生成所述修正视频数据时,具体用于:根据第一修正语音段和第一初始语音段的时长比,得到时长修正系数;根据时长修正系数,对第一修正语音段的播放时长进行调整,得到与第一初始语音段的播放时长相同的第二修正语音段;利用第二修正语音段替换第一初始语音段,生成修正视频数据。
根据本公开的一个或多个实施例,处理模块33在根据时长修正系数,对第一修正语音段的播放时长进行调整,得到与第一初始语音段的播放时长相同的第二修正语音段时,具体用于:基于时长修正系数,对第一修正语音段进行加速或减速处理,得到与第一初始语音段的播放时长相同的第二修正语音段;或者,基于时长修正系数,对第一修正语音段中的语音空白段进行剪裁,得到与第一初始语音段的播放时长相同的第二修正语音段。
根据本公开的一个或多个实施例,在响应于针对初始字幕的编辑操作,生成修正字幕之后,处理模块33,还用于:获取变更文本段的发音对象数量,变更文本段包括初始字幕中,与对应位置的修正字幕存在内容差异的语句,所述发音对象数量表征所述初始视频数据中输出所述变更文本段的对话内容的对象数量;处理模块33在基于修正字幕,生成初始视频数据对应的修正视频数据时,具体用于:若发音对象数量为1,则基于修正字幕,生成初始视频数据对应的修正视频数据。
根据本公开的一个或多个实施例,处理模块33在基于修正字幕,生成初始视频数据对应的修正视频数据时,还用于:若发音对象数量大于1,则基于发音对象,对修正字幕中与变更文本段对应的修正文本段进行拆分,得到各发音对象对应的子修正文本段;根据各发音对象对应的子修正文本段,生成对应的子修正语音段;根据变更文本段中各发音对象对应的第一播放时间段,从初始视频数据中定位对应的至少两个第二初始语音段;基于发音对象,将第二初始语音段替换为对应的子修正语音段,生成修正视频数据。
根据本公开的一个或多个实施例,处理模块33,还用于:接收用户输入的验证声音数据;对比验证声音数据和初始视频数据中的语音轨道数据,得到验证结果;
处理模块33在基于修正字幕,生成初始视频数据对应的修正视频数据时,具体用于:根据验证结果,若验证通过,则基于修正字幕,生成初始视频数据对应的修正视频数据。
其中,生成模块31、修正模块32和处理模块33依次连接。本实施例提供的视频处理装置3可以执行上述方法实施例的技术方案,其实现原理和技术效果类似,本实施例此处不再赘述。
图12为本公开实施例提供的一种电子设备的结构示意图,如图12所示,该电子设备4包括:
处理器41,以及与处理器41通信连接的存储器42;
存储器42存储计算机执行指令;
处理器41执行存储器42存储的计算机执行指令,以实现如图2-图10所示实施例中的视频处理方法。
其中,可选地,处理器41和存储器42通过总线43连接。
相关说明可以对应参见图2-图10所对应的实施例中的步骤所对应的相关描述和效果进行理解,此处不做过多赘述。
本公开实施例提供一种计算机可读存储介质,计算机可读存储介质中存储有计算机执行指令,计算机执行指令被处理器执行时用于实现本公开图2-图10所对应的实施例中任一实施例提供的视频处理方法。
本公开实施例提供一种计算机程序产品,包括计算机程序,计算机程序被处理器执行时实现本公开图2-图10所对应的实施例中任一实施例提供的视频处理方法。
为了实现上述实施例,本公开实施例还提供了一种电子设备。
参考图13,其示出了适于用来实现本公开实施例的电子设备900的结构示意图,该电子设备900可以为终端设备或服务器。其中,终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、个人数字助理(Personal Digital Assistant,简称PDA)、平板电脑(Portable Android Device,简称PAD)、便携式多媒体播放器(Portable Media Player,简称PMP)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图13示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图13所示,电子设备900可以包括处理装置(例如中央处理器、图形处理器等)901,其可以根据存储在只读存储器(Read Only Memory,简称ROM)902中的程序或者从存储装置908加载到随机访问存储器(Random Access Memory,简称RAM)903中的程序而执行各种适当的动作和处理。在RAM 903中,还存储有电子设备900操作所需的各种程序和数据。处理装置901、ROM 902以及RAM 903通过总线904彼此相连。输入/输出(I/O)接口905也连接至总线904。
通常,以下装置可以连接至I/O接口905:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置906;包括例如液晶显示器(Liquid Crystal Display,简称LCD)、扬声器、振动器等的输出装置907;包括例如磁带、硬盘等的存储装置908;以及通信装置909。通信装置909可以允许电子设备900与其他设备进行无线或有线通信以交换数据。虽然图13示出了具有各种装置的电子设备900,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置909从网络上被下载和安装,或者从存储装置908被安装,或者从ROM 902被安装。在该计算机程序被处理装置901执行时,执行本公开实施例的方法中限定的上述功能。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备执行上述实施例所示的方法。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(Local Area Network,简称LAN)或广域网(Wide Area Network,简称WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元或模块可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元或模块的名称在某种情况下并不构成对该单元本身的限定。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
第一方面,根据本公开的一个或多个实施例,提供了一种视频处理方法,包括:
对初始视频数据进行语音识别,得到所述初始视频数据对应的初始字幕;响应于针对所述初始字幕的编辑操作,生成修正字幕;基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,其中,所述修正视频数据中的语音内容与所述修正字幕的字幕内容相匹配。
根据本公开的一个或多个实施例,所述基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,包括:根据所述初始字幕中的变更文本段,确定所述初始视频数据的第一播放时间段,所述变更文本段包括所述初始字幕中,与对应位置的修正字幕存在内容差异的语句;根据所述修正字幕中的修正文本段,生成第一修正语言段,所述修正文本段包括所述修正字幕中,与对应位置的初始字幕存在内容差异的语句;将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段,或者,删除所述初始视频数据中的第一播放时间段对应的第一初始语音段,以生成所述修正视频数据。
根据本公开的一个或多个实施例,所述根据所述修正字幕中的修正文本段,生成第一修正语言段,包括:根据所述初始视频数据的语音轨道数据,得到所述初始视频数据对应的声纹特征;根据所述声纹特征和所述修正文本段的文本内容,生成所述第一修正语音段。
根据本公开的一个或多个实施例,所述根据所述初始视频数据的语音轨道数据,得到所述初始视频数据对应的声纹特征,包括:截取所述初始视频数据的语音轨道数据中的关键语音段数据,所述关键语音段数据包括所述第一初始语音段,以及所述第一初始语音段之前第一长度的第二初始语音段和/或所述第一初始语音段之后第二长度的第三初始语音段;通过语音模型对所述关键语音数据段进行处理,生成所述声纹特征。
根据本公开的一个或多个实施例,所述方法还包括:根据所述修正字幕中的修正文本段,生成修正视频数据段;根据所述变更文本段的第一播放时间段,从所述初始视频数据中定位对应的第一初始视频段;将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段,或者,删除所述初始视频数据中的第一播放时间段对应的第一初始语音段,以生成所述修正视频数据,包括:将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段、基于所述修正视频数据段替换所述第一初始视频段,以生成所述修正视频数据。
根据本公开的一个或多个实施例,所述根据所述修正字幕中的修正文本段,生成修正视频数据段,包括:根据所述修正文本段的文本内容,生成表征所述初始视频数据的第一播放时间段内的对象表情特征的贴图数据;根据所述贴图数据,以生成所述修正视频数据段。
根据本公开的一个或多个实施例,所述将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段,以生成所述修正视频数据,包括:根据所述第一修正语音段和所述第一初始语音段的时长比,得到时长修正系数;根据所述时长修正系数,对所述第一修正语音段的播放时长进行调整,得到与所述第一初始语音段的播放时长相同的第二修正语音段;利用所述第二修正语音段替换所述第一初始语音段,以生成所述修正视频数据。
根据本公开的一个或多个实施例,所述根据所述时长修正系数,对所述第一修正语音段的播放时长进行调整,得到与所述第一初始语音段的播放时长相同的第二修正语音段,包括:基于所述时长修正系数,对所述第一修正语音段进行加速或减速处理,得到与所述第一初始语音段的播放时长相同的第二修正语音段;或者,基于所述时长修正系数,对所述第一修正语音段中的语音空白段进行剪裁,得到与所述第一初始语音段的播放时长相同的第二修正语音段。
根据本公开的一个或多个实施例,在所述响应于针对所述初始字幕的编辑操作,生成修正字幕之后,还包括:获取变更文本段的发音对象数量,所述变更文本段包括所述初始字幕中,与对应位置的修正字幕存在内容差异的语句,所述发音对象数量表征所述初始视频数据中输出所述变更文本段的对话内容的对象数量;所述基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,包括:若所述发音对象数量为1,则基于所述修正字幕,生成所述初始视频数据对应的修正视频数据;若所述发音对象数量大于1,则基于发音对象,对所述修正字幕中与所述变更文本段对应的修正文本段进行拆分,得到各所述发音对象对应的子修正文本段;根据各所述发音对象对应的子修正文本段,生成对应的子修正语音段;根据所述变更文本段中各发音对象对应的第一播放时间段,从所述初始视频数据中定位对应的至少两个第二初始语音段;基于发音对象,将所述第二初始语音段替换为对应的子修正语音段,以生成所述修正视频数据。
第二方面,根据本公开的一个或多个实施例,提供了一种视频处理装置,包括:
生成模块,用于对初始视频数据进行语音识别,得到所述初始视频数据对应的初始字幕;
修正模块,用于响应于针对所述初始字幕的编辑操作,生成修正字幕;
处理模块,用于基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,其中,所述修正视频数据中的语音内容与所述修正字幕的字幕内容相匹配。
根据本公开的一个或多个实施例,处理模块,具体用于:根据初始字幕中的变更文本段,确定初始视频数据的第一播放时间段,变更文本段包括初始字幕中,与对应位置的修正字幕存在内容差异的语句;根据修正字幕中的修正文本段,生成第一修正语言段,修正文本段包括修正字幕中,与对应位置的初始字幕存在内容差异的语句;将初始视频数据中的第一播放时间段对应的第一初始语音段替换为第一修正语音段,或者,删除初始视频数据中的第一播放时间段对应的第一初始语音段,以生成修正视频数据。
根据本公开的一个或多个实施例,处理模块在根据修正字幕中的修正文本段,生成第一修正语言段时,具体用于:根据初始视频数据的语音轨道数据,得到初始视频数据对应的声纹特征;根据声纹特征和修正文本段的文本内容,生成第一修正语音段。
根据本公开的一个或多个实施例,处理模块在根据初始视频数据的语音轨道数据,得到初始视频数据对应的声纹特征时,具体用于:截取初始视频数据的语音轨道数据中的关键语音段数据,关键语音段数据包括第一初始语音段,以及第一初始语音段之前第一长度的第二初始语音段和/或第一初始语音段之后第二长度的第三初始语音段;通过预训练的语音模型对关键语音数据段进行处理,生成声纹特征。
根据本公开的一个或多个实施例,处理模块,还用于:根据修正字幕中的修正文本段,生成修正视频数据段;根据变更文本段的第一播放时间段,从初始视频数据中定位对应的第一初始视频段;处理模块在将初始视频数据中的第一播放时间段对应的第一初始语音段替换为第一修正语音段,或者,删除初始视频数据中的第一播放时间段对应的第一初始语音段,以生成修正视频数据时,具体用于:将初始视频数据中的第一播放时间段对应的第一初始语音段替换为第一修正语音段、基于修正视频数据段替换第一初始视频段,以生成修正视频数据。
根据本公开的一个或多个实施例,处理模块在根据修正字幕中的修正文本段,生成修正视频数据段时,具体用于:根据修正文本段的文本内容,生成表征初始视频数据的第一播放时间段内的对象表情特征的贴图数据;根据贴图数据,生成修正视频数据段。
根据本公开的一个或多个实施例,处理模块在所述将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段,以生成所述修正视频数据时,具体用于:根据第一修正语音段和第一初始语音段的时长比,得到时长修正系数;根据时长修正系数,对第一修正语音段的播放时长进行调整,得到与第一初始语音段的播放时长相同的第二修正语音段;利用第二修正语音段替换第一初始语音段,生成修正视频数据。
根据本公开的一个或多个实施例,处理模块在根据时长修正系数,对第一修正语音段的播放时长进行调整,得到与第一初始语音段的播放时长相同的第二修正语音段时,具体用于:基于时长修正系数,对第一修正语音段进行加速或减速处理,得到与第一初始语音段的播放时长相同的第二修正语音段;或者,基于时长修正系数,对第一修正语音段中的语音空白段进行剪裁,得到与第一初始语音段的播放时长相同的第二修正语音段。
根据本公开的一个或多个实施例,在响应于针对初始字幕的编辑操作,生成修正字幕之后,处理模块,还用于:获取变更文本段的发音对象数量,变更文本段包括初始字幕中,与对应位置的修正字幕存在内容差异的语句,所述发音对象数量表征所述初始视频数据中输出所述变更文本段的对话内容的对象数量;处理模块在基于修正字幕,生成初始视频数据对应的修正视频数据时,具体用于:若发音对象数量为1,则基于修正字幕,生成初始视频数据对应的修正视频数据。
根据本公开的一个或多个实施例,处理模块在基于修正字幕,生成初始视频数据对应的修正视频数据时,还用于:若发音对象数量大于1,则基于发音对象,对修正字幕中与变更文本段对应的修正文本段进行拆分,得到各发音对象对应的子修正文本段;根据各发音对象对应的子修正文本段,生成对应的子修正语音段;根据变更文本段中各发音对象对应的第一播放时间段,从初始视频数据中定位对应的至少两个第二初始语音段;基于发音对象,将第二初始语音段替换为对应的子修正语音段,生成修正视频数据。
根据本公开的一个或多个实施例,处理模块,还用于:接收用户输入的验证声音数据;对比验证声音数据和初始视频数据中的语音轨道数据,得到验证结果;处理模块在基于修正字幕,生成初始视频数据对应的修正视频数据时,具体用于:根据验证结果,若验证通过,则基于修正字幕,生成初始视频数据对应的修正视频数据。
第三方面,根据本公开的一个或多个实施例,提供了一种电子设备,包括:至少一个处理器和存储器;
所述存储器存储计算机执行指令;
所述至少一个处理器执行所述存储器存储的计算机执行指令,使得所述至少一个处理器执行如上第一方面以及第一方面各种可能的设计所述的视频处理方法。
第四方面,根据本公开的一个或多个实施例,提供了一种计算机可读存储介质,所述计算机可读存储介质中存储有计算机执行指令,当处理器执行所述计算机执行指令时,实现如上第一方面以及第一方面各种可能的设计所述的视频处理方法。
第五方面,根据本公开的一个或多个实施例,提供了一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如上第一方面以及第一方面各种可能的设计所述的视频处理方法。
以上描述仅为本公开的较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本公开中所涉及的公开范围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述公开构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上述特征与本公开中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。
此外,虽然采用特定次序描绘了各操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序执行来执行。在一定环境下,多任务和并行处理可能是有利的。同样地,虽然在上面论述中包含了若干具体实现细节,但是这些不应当被解释为对本公开的范围的限制。在单独的实施例的上下文中描述的某些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的各种特征也可以单独地或以任何合适的子组合的方式实现在多个实施例中。
尽管已经采用特定于结构特征和/或方法逻辑动作的语言描述了本主题,但是应当理解所附权利要求书中所限定的主题未必局限于上面描述的特定特征或动作。相反,上面所描述的特定特征和动作仅仅是实现权利要求书的示例形式。
Claims (13)
- 一种视频处理方法,包括:对初始视频数据进行语音识别,得到所述初始视频数据对应的初始字幕;响应于针对所述初始字幕的编辑操作,生成修正字幕;基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,其中,所述修正视频数据中的语音内容与所述修正字幕的字幕内容相匹配。
- 根据权利要求1所述的方法,其中,所述基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,包括:根据所述初始字幕中的变更文本段,确定所述初始视频数据的第一播放时间段,所述变更文本段包括所述初始字幕中,与对应位置的修正字幕存在内容差异的语句;根据所述修正字幕中的修正文本段,生成第一修正语言段,所述修正文本段包括所述修正字幕中,与对应位置的初始字幕存在内容差异的语句;将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段,或者,删除所述初始视频数据中的第一播放时间段对应的第一初始语音段,以生成所述修正视频数据。
- 根据权利要求2所述的方法,其中,所述根据所述修正字幕中的修正文本段,生成第一修正语言段,包括:根据所述初始视频数据的语音轨道数据,得到所述初始视频数据对应的声纹特征;根据所述声纹特征和所述修正文本段的文本内容,生成所述第一修正语音段。
- 根据权利要求3所述的方法,其中,所述根据所述初始视频数据的语音轨道数据,得到所述初始视频数据对应的声纹特征,包括:截取所述初始视频数据的语音轨道数据中的关键语音段数据,所述关键语音段数据包括所述第一初始语音段,以及所述第一初始语音段之前第一长度的第二初始语音段和/或所述第一初始语音段之后第二长度的第三初始语音段;通过语音模型对所述关键语音数据段进行处理,生成所述声纹特征。
- 根据权利要求2-4任一项所述的方法,还包括:根据所述修正字幕中的修正文本段,生成修正视频数据段;根据所述变更文本段的第一播放时间段,从所述初始视频数据中定位对应的第一初始视频段;将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段,或者,删除所述初始视频数据中的第一播放时间段对应的第一初始语音段,以生成所述修正视频数据,包括:将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段.基于所述修正视频数据段替换所述第一初始视频段,以生成所述修正视频数据。
- 根据权利要求5所述的方法,其中,所述根据所述修正字幕中的修正文本段,生成修正视频数据段,包括:根据所述修正文本段的文本内容,生成表征所述初始视频数据的第一播放时间段内的对象表情特征的贴图数据;根据所述贴图数据,以生成所述修正视频数据段。
- 根据权利要求2-6任一项所述的方法,其中,所述将所述初始视频数据中的第一播放时间段对应的第一初始语音段替换为所述第一修正语音段,以生成所述修正视频数据,包括:根据所述第一修正语音段和所述第一初始语音段的时长比,得到时长修正系数;根据所述时长修正系数,对所述第一修正语音段的播放时长进行调整,得到与所述第一初始语音段的播放时长相同的第二修正语音段;利用所述第二修正语音段替换所述第一初始语音段,以生成所述修正视频数据。
- 根据权利要求7所述的方法,其中,所述根据所述时长修正系数,对所述第一修正语音段的播放时长进行调整,得到与所述第一初始语音段的播放时长相同的第二修正语音段,包括:基于所述时长修正系数,对所述第一修正语音段进行加速或减速处理,得到与所述第一初始语音段的播放时长相同的第二修正语音段;或者,基于所述时长修正系数,对所述第一修正语音段中的语音空白段进行剪裁,得到与所述第一初始语音段的播放时长相同的第二修正语音段。
- 根据权利要求1-8任一项所述的方法,其中,在所述响应于针对所述初始字幕的编辑操作,生成修正字幕之后,所述方法还包括:获取变更文本段的发音对象数量,所述变更文本段包括所述初始字幕中,与对应位置的修正字幕存在内容差异的语句,所述发音对象数量表征所述初始视频数据中输出所述变更文本段的对话内容的对象数量;所述基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,包括:若所述发音对象数量为1,则基于所述修正字幕,生成所述初始视频数据对应的修正视频数据;若所述发音对象数量大于1,则基于发音对象,对所述修正字幕中与所述变更文本段对应的修正文本段进行拆分,得到各所述发音对象对应的子修正文本段;根据各所述发音对象对应的子修正文本段,生成对应的子修正语音段;根据所述变更文本段中各发音对象对应的第一播放时间段,从所述初始视频数据中定位对应的至少两个第二初始语音段;基于发音对象,将所述第二初始语音段替换为对应的子修正语音段,以生成所述修正视频数据。
- 一种视频处理装置,包括:生成模块,被配置为对初始视频数据进行语音识别,得到所述初始视频数据对应的初始字幕;修正模块,被配置为响应于针对所述初始字幕的编辑操作,生成修正字幕;处理模块,被配置为基于所述修正字幕,生成所述初始视频数据对应的修正视频数据,其中,所述修正视频数据中的语音内容与所述修正字幕的字幕内容相匹配。
- 一种电子设备,包括:处理器和存储器,其中,所述存储器存储计算机执行指令;所述处理器执行所述存储器存储的计算机执行指令,使得所述处理器执行如权利要求1至9任一项所述的视频处理方法。
- 一种计算机可读存储介质,存储有计算机执行指令,其中,当处理器执行所述计算机执行指令时,实现如权利要求1至9任一项所述的视频处理方法。
- 一种计算机程序产品,包括计算机程序,其中,所述计算机程序被处理器执行时实现如权利要求1至9任一项所述的视频处理方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410382907.2 | 2024-03-29 | ||
| CN202410382907.2A CN120730132A (zh) | 2024-03-29 | 2024-03-29 | 视频处理方法、装置、电子设备及存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025201466A1 true WO2025201466A1 (zh) | 2025-10-02 |
Family
ID=97165787
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/085393 Pending WO2025201466A1 (zh) | 2024-03-29 | 2025-03-27 | 视频处理方法、装置、电子设备及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN120730132A (zh) |
| WO (1) | WO2025201466A1 (zh) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20160014438A1 (en) * | 2014-07-14 | 2016-01-14 | Hulu, LLC | Caption and Speech Alignment for a Video Delivery System |
| CN111885416A (zh) * | 2020-07-17 | 2020-11-03 | 北京来也网络科技有限公司 | 一种音视频的修正方法、装置、介质及计算设备 |
| CN111885313A (zh) * | 2020-07-17 | 2020-11-03 | 北京来也网络科技有限公司 | 一种音视频的修正方法、装置、介质及计算设备 |
| CN116074583A (zh) * | 2023-02-09 | 2023-05-05 | 武汉简视科技有限公司 | 根据视频剪辑时间点更正字幕文件时间轴的方法及系统 |
| CN116962755A (zh) * | 2023-03-15 | 2023-10-27 | 腾讯科技(深圳)有限公司 | 字幕生成方法、相关设备及存储介质 |
-
2024
- 2024-03-29 CN CN202410382907.2A patent/CN120730132A/zh active Pending
-
2025
- 2025-03-27 WO PCT/CN2025/085393 patent/WO2025201466A1/zh active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20160014438A1 (en) * | 2014-07-14 | 2016-01-14 | Hulu, LLC | Caption and Speech Alignment for a Video Delivery System |
| CN111885416A (zh) * | 2020-07-17 | 2020-11-03 | 北京来也网络科技有限公司 | 一种音视频的修正方法、装置、介质及计算设备 |
| CN111885313A (zh) * | 2020-07-17 | 2020-11-03 | 北京来也网络科技有限公司 | 一种音视频的修正方法、装置、介质及计算设备 |
| CN116074583A (zh) * | 2023-02-09 | 2023-05-05 | 武汉简视科技有限公司 | 根据视频剪辑时间点更正字幕文件时间轴的方法及系统 |
| CN116962755A (zh) * | 2023-03-15 | 2023-10-27 | 腾讯科技(深圳)有限公司 | 字幕生成方法、相关设备及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN120730132A (zh) | 2025-09-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN111741326B (zh) | 视频合成方法、装置、设备及存储介质 | |
| CN112333179B (zh) | 虚拟视频的直播方法、装置、设备及可读存储介质 | |
| US20160021334A1 (en) | Method, Apparatus and System For Regenerating Voice Intonation In Automatically Dubbed Videos | |
| US20110093263A1 (en) | Automated Video Captioning | |
| WO2023197979A1 (zh) | 一种数据处理方法、装置、计算机设备及存储介质 | |
| CN115955585B (zh) | 视频生成方法、装置、电子设备及存储介质 | |
| US20230326369A1 (en) | Method and apparatus for generating sign language video, computer device, and storage medium | |
| JP2012181358A (ja) | テキスト表示時間決定装置、テキスト表示システム、方法およびプログラム | |
| US20230125543A1 (en) | Generating audio files based on user generated scripts and voice components | |
| WO2023218268A1 (en) | Generation of closed captions based on various visual and non-visual elements in content | |
| KR102555698B1 (ko) | 인공지능을 이용한 자동 자막 동기화 방법 및 장치 | |
| US20240404525A1 (en) | Media System with Closed-Captioning Data and/or Subtitle Data Generation Features | |
| CN113035199A (zh) | 音频处理方法、装置、设备及可读存储介质 | |
| CN114783408A (zh) | 一种音频数据处理方法、装置、计算机设备以及介质 | |
| CN119299770A (zh) | 一种视频字幕提取方法、装置及电子设备 | |
| KR102786445B1 (ko) | 인공지능 알고리즘을 이용하여 화면해설방송을 자동으로 생성하는 방법 및 이를 위한 시스템 | |
| KR102541008B1 (ko) | 화면해설 컨텐츠를 제작하는 방법 및 장치 | |
| KR101618777B1 (ko) | 파일 업로드 후 텍스트를 추출하여 영상 또는 음성간 동기화시키는 서버 및 그 방법 | |
| CN118283288A (zh) | 生成视频的方法、装置、电子设备和存储介质 | |
| WO2025201466A1 (zh) | 视频处理方法、装置、电子设备及存储介质 | |
| CN120353967A (zh) | 有声书视频生成方法、装置、电子设备、介质和程序产品 | |
| CN117668296A (zh) | 视频审核的方法、电子设备和计算机可读存储介质 | |
| WO2023218272A1 (en) | Distributor-side generation of captions based on various visual and non-visual elements in content | |
| CN115171645A (zh) | 一种配音方法、装置、电子设备以及存储介质 | |
| CN115174825B (zh) | 一种配音方法、装置、电子设备以及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25775097 Country of ref document: EP Kind code of ref document: A1 |