WO2020093883A1 - 获取视频片段的方法、装置、服务器和存储介质 - Google Patents

获取视频片段的方法、装置、服务器和存储介质 Download PDF

Info

Publication number
WO2020093883A1
WO2020093883A1 PCT/CN2019/113321 CN2019113321W WO2020093883A1 WO 2020093883 A1 WO2020093883 A1 WO 2020093883A1 CN 2019113321 W CN2019113321 W CN 2019113321W WO 2020093883 A1 WO2020093883 A1 WO 2020093883A1
Authority
WO
WIPO (PCT)
Prior art keywords
time point
video data
target
live video
live
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/113321
Other languages
English (en)
French (fr)
Inventor
耿振健
张洋
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Dajia Internet Information Technology Co Ltd
Original Assignee
Beijing Dajia Internet Information Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Dajia Internet Information Technology Co Ltd filed Critical Beijing Dajia Internet Information Technology Co Ltd
Priority to US17/257,447 priority Critical patent/US11375295B2/en
Publication of WO2020093883A1 publication Critical patent/WO2020093883A1/zh
Anticipated expiration legal-status Critical
Priority to US17/830,222 priority patent/US20220303644A1/en
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/21Server components or server architectures
    • H04N21/218Source of audio or video content, e.g. local disk arrays
    • H04N21/2187Live feed
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/233Processing of audio elementary streams
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/234Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
    • H04N21/23418Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/433Content storage operation, e.g. storage operation in response to a pause request, caching operations
    • H04N21/4334Recording operations
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/439Processing of audio elementary streams
    • H04N21/4394Processing of audio elementary streams involving operations for analysing the audio stream, e.g. detecting features or characteristics in audio streams
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/472End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
    • H04N21/47202End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content for requesting content on demand, e.g. video on demand
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/472End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
    • H04N21/47217End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content for controlling playback functions for recorded or on-demand content, e.g. using progress bars, mode or play-point indicators or bookmarks
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/81Monomedia components thereof
    • H04N21/8106Monomedia components thereof involving special audio data, e.g. different tracks for different languages
    • H04N21/8113Monomedia components thereof involving special audio data, e.g. different tracks for different languages comprising music, e.g. song in MP3 format
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/845Structuring of content, e.g. decomposing content into time segments
    • H04N21/8456Structuring of content, e.g. decomposing content into time segments by decomposing the content in the time domain, e.g. in time segments
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/85Assembly of content; Generation of multimedia applications
    • H04N21/858Linking data to content, e.g. by linking an URL to a video object, by creating a hotspot

Definitions

  • the present application relates to the technical field of audio and video, and in particular to a method, device, server, and storage medium for acquiring video clips.
  • a recording button is provided in the live broadcast interface. After detecting an operation instruction indicating that the recording button is operated, the terminal can use the screen recording function provided by the terminal operating system to start recording the video data displayed on the screen, and the terminal detects again After the operation instruction indicating that the record button is operated, the recording of the video data displayed on the screen is ended. This allows you to record video clips with exciting content.
  • the terminal starts to record the video data displayed on the screen after detecting the operation instruction that the recording button is operated, so that the video data displayed on the screen is recorded from the user to the wonderful content to the terminal There will be a period of time, and the wonderful content of this period cannot be recorded, so it will lead to incomplete video clips of the wonderful content.
  • embodiments of the present application provide a method, device, server, and storage medium for acquiring video clips.
  • a method for obtaining video clips including:
  • the target time point pair obtain a target video segment from the live video data.
  • an apparatus for acquiring video clips including:
  • the obtaining unit is configured to obtain live video data of the live broadcasting room
  • the determining unit is configured to determine a target time point pair of the live video data according to the audio data in the live video data and the audio data of the original performer, wherein the target time point pair includes a start time point and an end Point in time
  • the acquiring unit is further configured to acquire a target video segment from the live video data according to the target time point pair.
  • a server including: a processor and a memory for storing processor-executable instructions; wherein, the processor is configured to perform the following method of acquiring a video clip:
  • the target time point pair obtain a target video segment from the live video data.
  • a non-transitory computer-readable storage medium which when the instructions in the storage medium are executed by the processor of the server, enables the server to perform the following method of acquiring video clips :
  • the target time point pair obtain a target video segment from the live video data.
  • an application program which includes one or more instructions, and the one or more instructions may be executed by a processor of a server to complete the following method for obtaining a video clip:
  • the target time point pair obtain a target video segment from the live video data.
  • Fig. 1 is a flow chart showing a method for acquiring video clips according to an exemplary embodiment
  • Fig. 2 is a schematic diagram showing link information of a video clip according to an exemplary embodiment
  • Fig. 3 is a schematic diagram of a first time period according to an exemplary embodiment
  • Fig. 4 is a structural block diagram of an apparatus for acquiring video clips according to an exemplary embodiment
  • Fig. 5 is a structural block diagram of an apparatus for acquiring video clips according to an exemplary embodiment
  • Fig. 6 is a structural block diagram of a server according to an exemplary embodiment
  • Fig. 7 is a structural block diagram of another server according to an exemplary embodiment.
  • An embodiment of the present application provides a method for obtaining a video clip, and the method may be executed by a server.
  • the server may be a background server of the live broadcast application, and the server may be a CDN (Content Delivery Network) server.
  • the server may be provided with a processor, a memory, a transceiver, etc.
  • the processor may be used to obtain video clips, distribute video clips, and other processes, and the memory may be used to store data required during the video clip acquisition and generated data, such as video
  • the video data of the clips, live video data, etc., the transceiver can be used to receive and send data, the data can be live video data, comment information, video clip link information, etc.
  • the terminal used by the anchor is called a user terminal, and the background server of the live broadcast application is called a server.
  • the above live broadcast application is installed in the user terminal.
  • the user terminal After the anchor controls the live broadcast of the singing live room through the live application installed in the user terminal, the user terminal obtains the live video data of the anchor and sends the live video data to the server. After receiving the live video data sent by the user terminal, the server may obtain the target video segment from the received live video data.
  • the user terminal obtains the live video data of the anchor and sends the live video data to the server.
  • the server stores the received live video data, and after the live broadcast of the live concert room ends, the target video segment can be obtained from the stored live video data.
  • An embodiment of the present application provides a method for obtaining video clips. As shown in FIG. 1, the execution flow of the method may be as follows:
  • Step 101 Obtain live video data in a live concert room.
  • the live singing room refers to the live broadcasting room where music is played after the broadcast starts.
  • the live singing room is a live broadcasting room for singing songs and a live broadcasting room for playing musical instruments.
  • the server may receive the live video data sent by the user terminal and save the received live video data.
  • the server can also determine other accounts other than the anchor account in the account for logging into the live broadcast room, and then send the received live video data to the terminal used by the other accounts.
  • the terminal used by each other account is called the login terminal.
  • each of the above-mentioned login terminals also has the above-mentioned live-streaming application installed. After receiving the live-stream video data, each of the login-terminals can play the received live-streaming video data through the live-streaming application installed in the login terminal on the live interface of the live concert room.
  • the live video data includes audio data with sound information and video data with picture information.
  • the above live video data can also be understood as multimedia data.
  • Step 102 Determine the target time point pair in the live video data according to the audio data in the live video data and the audio data of the original performer.
  • this step can be understood as: determining the target time point pair of the live video data.
  • the audio data of the original performer may be the audio data of the original singing song or the audio data of the original performer using the musical instrument to perform.
  • the target time point pair is one or more time point pairs, and each time point pair includes a set of time points, namely a start time point and an end time point.
  • the target time point pair includes a start time point and an end time point that may be a time point characterizing the start of the live video highlight content and a time point characterizing the end of the live video highlight content.
  • the server may directly obtain the audio data of the original performer according to the audio data in the live video data . If the video data and audio data in the live video data are mixed, after the server obtains the live video data, the video data and audio data in the live video data can be separated to obtain the audio data in the live video data, and then based on the live video Audio data in the data to obtain the audio data of the original performer.
  • the server may obtain the introduction information of the live broadcast room, where the broadcast introduction information includes the content of the live broadcast of the anchor, based on the content of the live broadcast of the live broadcast, Get the audio data of the original performer.
  • the server may perform similar matching on the audio data in the live video data and the audio data of the original performer, and based on the similar matching result, determine the target time point pair in the live video data.
  • a time point may be determined first, and then a target time point pair is determined based on the time point.
  • the processing in step 102 may be as follows:
  • the first time point in the live video data is determined according to the audio data in the live video data and the audio data of the original performer. With the first time point as the center, the target time point pair corresponding to the first time point is determined according to the preset intercept time.
  • the preset interception time can be preset and stored in the server, such as 10 seconds.
  • the server may use the audio data in the live video data and the audio data of the original performer to determine the first type of time point. Then obtain the pre-stored preset interception time, reduce the first type of time point by half of the preset interception time, get the start time point corresponding to the first type of time point, and increase the first type of time point by the preset interception time Half the duration, the end time point corresponding to the first time point is obtained. In this way, the start time point and the end time point can form a target time point pair corresponding to the first type of time point.
  • the first type of time point is the 20th second
  • the preset interception time is 10 seconds
  • the preset half of the interception time is 5 seconds
  • the above-mentioned first time point may be a time point characterizing the live broadcast highlight content
  • the start time point and the end time point of the target time point corresponding to the first type time point may be the time point characterizing the start of the live broadcast highlight content and the characterization The point in time when the live content will end.
  • the first type of time point is determined
  • the audio data in the live video data perform voice recognition on the audio data in the live video data to obtain the lyrics of the song.
  • the audio data of the song sung by the original singer is acquired.
  • the similarity of the audio characteristics of the audio data of the song originally sung and the audio data of the live video data is determined as the similarity of each lyrics.
  • the live time point of the audio data in the live video data corresponding to the position of the lyrics with the highest similarity among the lyrics whose similarity is higher than the first preset threshold is determined as the first time point in the live video data.
  • the first preset threshold may be preset and stored in the server, for example, the first preset threshold may be 90%.
  • the server may use a pre-stored voice recognition algorithm to perform voice recognition on the audio data in the live video data to obtain the lyrics of the song sung by the anchor. Then, the obtained lyrics can be used to query in a preset lyrics database, where the lyrics database includes the lyrics and the audio data of the original singing of the song where the lyrics are located, and the audio data of the original singing of the song where the obtained lyrics are located is determined. Then, for any lyrics, the server can determine the audio data of the song sung by the original singer and the audio data of the song sung by the anchor.
  • the audio data of the song sung by the original singer and the audio data of the song sung by the anchor respectively Perform audio feature extraction to determine the similarity between the audio features of the original song and the audio features of the anchor under the lyrics. Then determine the size relationship between the similarity and the first preset threshold, if the similarity is higher than the first preset threshold, determine the location with the highest similarity and the live video data corresponding to the location with the highest similarity in the lyrics The live time point of the audio data in the medium is determined as the first time point of the live video data. If the similarity is less than or equal to the first preset threshold, the process of determining the first type of time point is not performed. In this way, the above-mentioned processing is performed for each lyrics, and the first time point of the live video data can be determined.
  • the above speech recognition algorithm may be any kind of speech recognition algorithm, such as FED (Fast Endpoint Detection, fast endpoint detection algorithm, etc.).
  • FED Fast Endpoint Detection, fast endpoint detection algorithm, etc.
  • the server can identify the audio data in the live video data, determine the name of the work performed by the anchor, and then find the original based on the name of the work Performer performs the audio data of the musical instrument, aligns the audio data in the live video data with the audio data of the original performer's musical instrument, and performs segmentation processing on the two pieces of audio data after the alignment process according to a preset duration, such as The two pieces of audio data are divided into 5 seconds of audio data, the audio data in the live video data are sequentially numbered a1, a2, a3, ..., ai, ..., an, the audio data of the original performer playing the instrument are sequentially numbered For b1, b2, b3, ..., bi, ..., bn.
  • the server can extract the audio features of a1 and b1 respectively, and calculate the similarity of the extracted audio features for the audio features of a1 and b1. If the similarity is higher than the first preset threshold, then in a1 A location with the highest similarity to b1 is determined, and a live broadcast time corresponding to the location with the highest similarity is obtained, and the time is determined as the first type of time. By analogy, the first time point of audio data such as a2 and a3 after a1 can be determined.
  • the first type of time point may also be obtained in a segmented manner.
  • the above audio features may be pitch audio features, pitch audio features, and so on.
  • the above audio feature extraction algorithm may be an algorithm in the prior art.
  • the algorithm for extracting pitch audio features the specific process of extracting audio features is: pre-emphasis-framing-window-finding short-term average energy- seeking autocorrelation, after this process can be obtained Pitch audio features.
  • the main parameters involved in this process are high frequency boost parameters, frame length, frame shift, and unvoiced voiced threshold.
  • Step 103 Obtain the target video segment from the live video data according to the target time point pair.
  • the target video segment refers to a video segment that contains first audio data in the live video data.
  • the first audio data is: audio data of the live video data and the original performer's audio data whose similarity meets certain conditions.
  • the above target video segment may be a video segment live broadcasted from the start time point to the end time point included in the target time point pair in the live video data.
  • the server after the server determines the target time point pair, it can find the time stamp corresponding to the start time point in the target time point pair according to the time stamp of the live video data, and find the target time point pair end time point Timestamp, intercept the video clip between these two timestamps as the target video clip.
  • the target video segment may also be provided to the audience in the live broadcast room, and the corresponding processing may be as follows:
  • Generate link information for the target video clip Send the link information of the target video clip to the login terminals of each account other than the host account in the live concert room, so that the login terminals of other accounts display the link information on the playback interface of the live concert room, or The live broadcast interface displays the link information.
  • the login terminal of each account other than the above sends the above link information. Since the login terminal of each other account is installed with a live broadcast application, the login terminal of other accounts can display the link information on the playback interface of the live broadcast room through the installed live broadcast application, or on the live end interface of the live broadcast room Link information.
  • the playback interface is an interface for displaying playback links for playing back live video data
  • the live broadcast end interface refers to the interface displayed when the live broadcast ends.
  • the server may randomly obtain a picture from the data of the target video segment as the cover of the target video segment, and add a name to the target video segment. For example, you can add The name of the song sung by the anchor is used as the name of the target video segment, and then link information can be generated based on the above cover, name, and data storage address of the target video segment.
  • the link information can be a URL (Uniform Resource Locator).
  • the server may determine each account other than the anchor account in the account for logging into the live broadcasting room, and send link information of the target video clip to the login terminals of the other accounts.
  • the login terminals of other accounts can display the link information of the target video clip on the playback interface of the live concert through the installed live broadcast application, or can display the link information of the target video clip in the live broadcast end interface .
  • the server obtains the link information of two video clips, one is the link information of "Meow Star" and the other is the link information of "Meow Meow”.
  • the login terminals of the other accounts mentioned above can be accessed through The installed live broadcast application displays the link information of these two video clips at the live broadcast end interface.
  • the link information shown in FIG. 2 is two video playback links.
  • the audience in the live concert wants to share a certain link information, they can select the link information, and then click the corresponding sharing option.
  • the terminal used by the audience will display various areas for sharing after detecting the click instruction of the sharing option Options, such as regional options for sharing in an application or current live application.
  • the viewer can select the corresponding regional option, and then click to determine the option.
  • the terminal used by the viewer detects the click operation to determine the option and displays the edit box.
  • the preset content is displayed in the edit box. For example, look at A B song sung by the anchor, etc.
  • the viewer can directly share the content displayed in the edit box, or re-edit the content displayed in the edit box, and then share to the area corresponding to the selected area option. So far, a sharing process has been completed.
  • a process of screening the first type of time point is also provided, and the corresponding processing may be as follows:
  • the second time point of the live video data is determined. If the target time point in the first type time point belongs to the second type time point, the target time point is retained, and if the target time point in the first type time point does not belong to the second type time point, the target time point is deleted. Based on the reserved first time point as the center, the target time point pair corresponding to the reserved first time point is determined according to the preset intercept time.
  • the interactive information may include one or more of comment information, like information, and gift information.
  • the target time point may be any time point in the first type of time point. That is, each time point in the first type of time point is taken as the target time point to determine whether the target time point belongs to the second type of time point, thereby determining whether to retain the target time point or delete the target time point.
  • the server may store the received comment information, like information and gift information, and then may use one or more of the comment information, like information and gift information To determine the second time point in the live video data.
  • the target time point in the first type of time point belongs to the second type time point, if it belongs to the second type time point, the target time point can be retained, if it does not belong to the second type time point, the target time point can be deleted.
  • the server can then center on the reserved first-type time point, reduce the reserved first-type time point by half of the preset intercept time, and obtain the starting time point corresponding to the reserved first-type time point, and will
  • the reserved first-type time point is increased by half of the preset interception time, and the end time point corresponding to the reserved first-type time point is obtained, and the start time point and the end time point form a target time point pair.
  • the first type of time point can be screened based on the interactive information, so that the intercepted video clip has a higher probability of including exciting content.
  • the above-mentioned second type of time point can be understood as: characterizing the time point of frequent audience interaction during the live broadcast process.
  • a method for determining the target time point pair using interactive information is also provided, and the corresponding processing may be as follows:
  • the second time point of the live video data is determined.
  • the first type of time point and the second type of time point are combined, and the combined time point is deduplicated. Taking the time point after deduplication processing as the center and determining the target time point pair corresponding to the time point after deduplication processing according to the preset intercept time.
  • the server may store the received comment information, like information and gift information, and then may use one or more of the comment information, like information and gift information To determine the second time point of live video data.
  • the server can Taking the time point after deduplication as the center, reduce the time point after deduplication by half of the preset interception time, get the starting time point corresponding to the time point after deduplication, and increase the time point after deduplication by preset
  • the half of the interception time is to get the end time point corresponding to the time point after deduplication.
  • the start time point and the end time point form a target time point pair.
  • Method 1 If the amount of gift resources of the live video data in the first time period exceeds the second preset threshold, the intermediate time point or the end time point of the first time period is determined as the second time point in the live video data .
  • the duration of the first time period can also be preset and stored in the server, such as 2 seconds.
  • the second preset threshold may also be preset and stored in the server.
  • the server may determine the first time period in the live video data according to the time stamp of the live video data.
  • the first time period may be a time period of equal duration, and the time interval between adjacent time periods may be equal.
  • the time interval between adjacent time periods may be determined by the start time point of the adjacent time period or the end time point. Furthermore, there may or may not be an overlapping area between adjacent first time periods.
  • live video data is 30-minute video data
  • the first 0 to 2 seconds are the first first time period t1
  • the first 1 to 3 seconds are the second first time period t2
  • the second ⁇ 4 seconds is the third first time period t3, and so on, multiple first time periods are selected.
  • the server can obtain the resources of each gift carried, for example, "yacht" gift 50 gold coins, multiply the number of each gift and the corresponding resources to obtain The resource amount of each gift, and then add the resource amount of each gift to get the gift resource amount of the first time period.
  • the server can then determine the relationship between the amount of gift resources in the first time period and the second preset threshold, and if the amount of gift resources in the first time period is greater than the second preset threshold, it can determine the intermediate time point of the first time period,
  • the intermediate time point may be determined as the second type time point in the live video data, or the end time point of the first time period may be determined, and the end time point may be determined as the second type time point in the live video data.
  • the amount of gift resources can also be determined based on the image recognition method, and the corresponding processing may be as follows:
  • the server may acquire the image of each first time period from the live video data, and then input the image into a preset gift image recognition algorithm, which may be a pre-trained algorithm to recognize The number of each gift image contained in the image, then obtain the resources of each gift, multiply the number of each gift image by the corresponding resource, to obtain the resource amount of each gift, and then add the resource amount of each gift to get the first time The amount of gift resources for the segment.
  • a preset gift image recognition algorithm which may be a pre-trained algorithm to recognize The number of each gift image contained in the image, then obtain the resources of each gift, multiply the number of each gift image by the corresponding resource, to obtain the resource amount of each gift, and then add the resource amount of each gift to get the first time The amount of gift resources for the segment.
  • the amount of gift resources can be used to determine the exciting content.
  • the above-mentioned gift image may refer to an area in the image that represents a gift.
  • the above gift image recognition algorithm may be a trained neural network algorithm. After inputting an image to the above neural network algorithm, the neural network algorithm can output the name of the gift image contained in the image, that is, the name of the gift, and the number of gift images.
  • the duration of the second time period can also be set in advance and stored in the server, such as 2 seconds.
  • the third preset threshold can also be set in advance and stored in the server.
  • the server may determine the second time period in the live video data according to the time stamp of the live video data.
  • the second time period may be a time period with equal duration, and the time interval between adjacent time periods may be equal. Furthermore, there may or may not be an overlapping area between adjacent second time periods.
  • live video data is 30-minute video data
  • the first 0 to 2 seconds are the second time period
  • the first to third seconds are the second second time period
  • the second to fourth seconds are the third Two time periods, and so on, select multiple second time periods.
  • Determine the start time point and end time point of each second time period and then according to the start time point and end time point, determine the number of comment information received in the time interval between the start time point and end time point, and judge The relationship between the number of received comment information and the size of the third preset threshold, if the number of received comment information is greater than the third preset threshold, the intermediate time point of the second time period may be determined, and the intermediate time point may be determined It is the second time point in the live video data, or the end time point of the second time period can be determined, and the end time point is determined as the second time point in the live video data.
  • the number of comment information can be used to determine the exciting content.
  • Method 3 If the number of likes of live video data in the third time period exceeds the fourth preset threshold, the middle time point or end time point of the third time period is determined as the second type of live video data time point.
  • the duration of the third time period can also be preset and stored in the server, such as 2 seconds.
  • the fourth preset threshold can also be preset and stored in the server. In the live broadcast, to like is to click a preset mark in the live broadcast interface.
  • the server may determine the third time period in the live video data according to the time stamp of the live video data.
  • the third time period may be a time period of equal duration, and the time interval between adjacent time periods may be equal. Furthermore, there may or may not be an overlapping area between adjacent third time periods.
  • live video data is 30-minute video data
  • the first 0 to 2 seconds are the third time period
  • the first 1 to 3 seconds are the second third time period
  • the second to fourth seconds are the third Three time periods, and so on, select multiple third time periods.
  • the amount of gift resources, comment information, and likes information all have certain weights, respectively A, B, and C.
  • the amount of gift resources determined by the server is x
  • comment The number of information is y
  • the number of likes information is z
  • the weighted calculation is performed to obtain the weighted value: A * x + B * y + C * z, and determine the relationship between the weighted value and the preset value. If it is greater than the preset value, the middle time point of the fourth time period is determined as the second time point in the live video data. In this way, the second time point determined by the three kinds of interactive information is more accurate.
  • the fourth time period may be a time period with equal duration, and the time interval between adjacent time periods may be equal. Furthermore, there may or may not be an overlapping area between adjacent fourth time periods.
  • two types of interaction information can be selected in the above-mentioned methods one to three to perform weighted calculation to determine the second type of time point, which is the same as the processing method of the interaction information using the three methods, which will not be repeated here.
  • the duration of the first time period, the second time period, the third time period, and the fourth time period may be the same.
  • the first time period, the second The durations of the time period, the third time period, and the fourth time period are relatively short, for example, may be less than 5 seconds.
  • step 102 In addition, in order to ensure that the determined target video segment has no duplicate content, the following processing may also be performed after step 102 and before step 103:
  • the target time point pair if the first start time point is earlier than the second start time point, and the end time point corresponding to the first start time point is earlier than the end time point corresponding to the second start time point, and the second start The time point is earlier than the end time point corresponding to the first start time point, then in the target time point pair, the end time point corresponding to the first start time point is replaced with the end time point corresponding to the second start time point, and the first The second start time point and the end time point corresponding to the second start time point.
  • the first start time point is different from the second start time point
  • the first start time point is any start time point in the target time point pair except the second start time point
  • the second start time point is the target time Any starting time point in the point pair except the first starting time point.
  • the first start time point and the second start time point are start time points included in different time point pairs in the target time point pair.
  • the server determines the target time point pair, it can be determined whether there is a start time point and an end time point with overlapping time ranges. If there is, there is a first start time point and a second start point The time point is satisfied: the first start time point is earlier than the second start time point, and the end time point corresponding to the first start time point is earlier than the end time point corresponding to the second start time point, and the second start time point is earlier than the first
  • the end time point corresponding to the first start time point in the target time point pair, the end time point corresponding to the first start time point can be replaced with the end time point corresponding to the second end time point, and the second start time point and The end time point corresponding to the second start time point is deleted.
  • the end time points corresponding to the first start time point and the first start time point, the end time points corresponding to the second start time point and the second start time point become the first start time point and the second start time point
  • the end time point of, that is, the end time point corresponding to the first start time point is replaced with the end time point corresponding to the second start time point.
  • the first start time point is 10 minutes and 23 seconds
  • the first start time point corresponds to the end time point is 10 minutes and 33 seconds
  • the second start time point is the 10 minutes and 25 seconds
  • the first start time point is 10 minutes and 23 seconds
  • the end time point corresponding to the first start time point becomes 10 minutes and 35 seconds.
  • step 103 in order to ensure that the determined target video segment does not have repeated content, the following processing may also be performed after step 103:
  • start time point of the first video clip in the target video clip is earlier than the start time point of the second video clip, and the end time point of the first video clip is earlier than the end time point of the second video clip, and the The start time point is earlier than the end time point of the first video clip, then the first video clip and the second video clip are merged.
  • the first video segment is any video segment except the second video segment in the target video segment
  • the second video segment is any video segment except the first video segment in the target video segment
  • the first video segment and the second video segment are different video segments in the target video segment.
  • the server can determine whether any two video segments have overlapping parts. If there is an overlapping part, the first video segment and the second video segment exist, and Two video clips satisfy: the start time point of the first video clip is earlier than the start time point of the second video clip, the end time point of the first video clip is earlier than the end time point of the second video clip, and the The start time point is earlier than the end time point of the first video clip.
  • the server may merge the first video segment and the second video segment. In this way, video clips with duplicate content can be combined into one video clip.
  • the first video clip is a video clip from the 10th minute 30 seconds to the 10th minute 40 seconds
  • the second video clip is the video clip from the 10th minute 35 seconds to the 10th minute 45 seconds
  • the merged video clip is the 10th Video clip from 30 seconds to 10 minutes and 45 seconds.
  • the target video segment in order to increase the probability that the target video segment includes exciting content, the target video segment may be filtered based on the interactive information, and after step 103, the following processing may be performed:
  • the target video segment is retained. Or, if the number of comment information of the target video segment exceeds the sixth preset threshold, the target video segment is retained. Or, if the number of likes of the target video segment exceeds the seventh preset threshold, the target video segment is retained.
  • the fifth preset threshold, the sixth preset threshold, and the seventh preset threshold can all be preset and stored in the server.
  • the server may determine the gift resource amount of the target video segment, wherein the method of determining the gift resource amount of the target video segment and the determination of the gift resource amount of the first time period The method is the same and will not be repeated here. Then it is determined whether the amount of gift resources exceeds the fifth preset threshold. If it exceeds the fifth preset threshold, it is retained. If it does not exceed the fifth preset threshold, it indicates that it may not contain exciting content and can be deleted.
  • the server may determine the number of comment information of the target video segment, where the method of determining the number of comment information of the target video segment is the same as the method of determining the number of comment information of the second time period, here No longer. Then it is determined whether the number of comment information exceeds the sixth preset threshold, if it exceeds the sixth preset threshold, it is retained, and if it does not exceed the sixth preset threshold, it indicates that it may not contain exciting content and can be deleted.
  • the server may determine the number of likes of the target video segment, wherein the method of determining the number of likes of the target video segment and the method of determining the number of likes of the second time period The method is the same and will not be repeated here. Then it is determined whether the number of likes information exceeds the seventh preset threshold, if it exceeds the seventh preset threshold, it is retained, if it does not exceed the seventh preset threshold, it means that it may not contain exciting content, and can be deleted.
  • the obtained video clips can be further filtered through the interactive information, so that the probability that the clipped video clips contain exciting content can be increased.
  • the number of target video clips determined in step 103 may be relatively large.
  • the following filtering process may be performed, and the corresponding processing may be as follows:
  • the determined target video clips are sorted in order of the amount of gift resources from large to small, and the preset number of target video clips obtained at the forefront are determined as the final video clip.
  • the determined target video clips are sorted according to the order of the number of comment information from large to small, and a preset number of target video clips at the forefront are acquired and determined as the final video clip.
  • the determined target video clips are sorted in the order of the number of likes information from large to small, and a preset number of target video clips at the forefront are acquired and determined as the final video clip.
  • the preset number may be a preset number used to indicate the final video segment fed back to the terminal.
  • the server may determine the gift resource amount of the target video segment, wherein the method of determining the gift resource amount of the target video segment and the determination of the gift resource amount of the first time period The method is the same and will not be repeated here.
  • the target video clips are sorted according to the order of the amount of gift resources, and the first preset number of target video clips are selected as the final video clip.
  • the server may determine the number of comment information of the target video segment, where the method of determining the number of comment information of the target video segment is the same as the method of determining the number of comment information of the second period I won't repeat them here.
  • the target video clips are sorted according to the order of the number of comment information from large to small, and the preset number of target video clips at the beginning are selected as the final video clip.
  • the server may determine the number of likes of the target video segment, wherein the method of determining the number of likes of the target video segment and the method of determining the number of likes of the second time period The method is the same and will not be repeated here.
  • the target video clips are sorted in order of the number of likes information from large to small, and the preset number of target video clips at the forefront are selected as the final video clip.
  • determining the amount of gift resources, the number of comment information, and the number of likes of video clips can be understood as: determining the amount of gift resources, number of comment information, and likes of live video time corresponding to the video clip number.
  • the audio data in the live video data and the audio data of the original performer may be used to determine the target time point pair in the live video data, and then according to Get the target video clip at the start time point and end time point in the target time point pair. Since the server directly intercepts the video based on the audio data in the live video data and the audio data of the original performer, the video clip is obtained. There is no need to manually operate the recording button, and there will be no time interval from the start of the highlight content to the start of recording the video data displayed on the screen, so the video clips taken are relatively complete.
  • Fig. 4 is a block diagram of a device for acquiring video clips according to an exemplary embodiment. 4, the device includes an acquisition unit 411 and a determination unit 412.
  • the obtaining unit 411 is configured to obtain live video data of a live concert room
  • the determining unit 412 is configured to determine a target time point pair of the live video data according to the audio data in the live video data and the audio data of the original performer, where the target time point pair includes a start time point and End time
  • the obtaining unit 411 is further configured to obtain a target video segment from the live video data according to the target time point pair.
  • the determining unit 412 is configured to:
  • the audio data in the live video data is the audio data of the song sung by the anchor, and the audio data of the original performer is the audio data of the song sung by the original singer;
  • the determining unit 412 is configured to:
  • For each lyrics determine the similarity between the audio characteristics of the audio data of the song sung by the original singer and the audio characteristics of the audio data in the live video data as the similarity of the lyrics;
  • the live broadcast time point of the audio data in the live video data corresponding to the lyrics position with the highest similarity among the lyrics whose similarity is higher than the first preset threshold is determined as the first time point of the live video data.
  • the determining unit 412 is further configured to:
  • the determining unit 412 is configured to:
  • the target time point in the first type time point belongs to the second type time point, the target time point is retained, if the target time point in the first type time point does not belong to the second type time point Time point, then delete the target time point;
  • the determining unit 412 is further configured to:
  • the middle time point or the end time point of the first time period is determined as the second in the live video data Point in time; and / or,
  • the middle time point or the end time point of the second time period is determined as the first in the live video data Type 2 time points; and / or,
  • the middle time point or the end time point of the third time period is determined as the third in the live video data Second time point.
  • the determining unit 412 is further configured to:
  • the amount of gift resources in the first time period is determined according to the number of each gift image.
  • the determining unit 412 is further configured to:
  • the target time point pair if there is a first start time point earlier than a second start time point, and an end time point corresponding to the first start time point is earlier than an end time corresponding to the second start time point Point, and the second start time point is earlier than the end time point corresponding to the first start time point, then in the target time point pair, the end time point corresponding to the first start time point is replaced by An end time point corresponding to the second start time point, and deleting the end time point corresponding to the second start time point and the second start time point, wherein the first start time point and the second start time point The point is a starting time point included in a different time point pair in the target time point pair.
  • the determining unit 412 is further configured to:
  • the device further includes:
  • the sending unit 413 is configured to send the link information to the login terminals of each account other than the anchor account in the live broadcasting room, so that the login terminals of the other accounts play in the live broadcasting room
  • the interface displays the link information, or displays the link information on the live end interface of the live concert room.
  • the obtaining unit 411 is further configured to:
  • the target video segment is retained; or,
  • the target video segment is retained.
  • the audio data in the live video data and the audio data of the original performer may be used to determine the target time point pair in the live video data, and then according to Get the target video clip at the start time point and end time point in the target time point pair. Since the server directly intercepts the video based on the audio data in the live video data and the audio data of the original performer, the video clip is obtained. There is no need to manually operate the recording button, and there will be no time interval from the start of the highlight content to the start of recording the video data displayed on the screen, so the video clips taken are relatively complete.
  • the server 600 may have a relatively large difference due to different configurations or performances, and may include one or more than one central processing unit (CPU) 601 and one Or more than one memory 602, wherein at least one instruction is stored in the memory 602, and the at least one instruction is loaded and executed by the processor 601 to implement the steps of the foregoing method of acquiring a video clip.
  • CPU central processing unit
  • memory 602 wherein at least one instruction is stored in the memory 602, and the at least one instruction is loaded and executed by the processor 601 to implement the steps of the foregoing method of acquiring a video clip.
  • a server including: a processor and a memory for storing processor executable instructions, wherein the processor is configured to be executed to complete the above method of acquiring a video clip step.
  • Fig. 7 is a block diagram of a server 700 according to an exemplary embodiment.
  • the server 700 includes a processing component 722, which further includes one or more processors, and memory resources represented by the memory 732, for storing instructions executable by the processing component 722, such as application programs.
  • the application programs stored in the memory 732 may include one or more modules each corresponding to a set of instructions.
  • the processing component 722 is configured to execute instructions to perform the steps of the above method of acquiring a video clip.
  • the server 700 may also include a power component 726 configured to perform power management of the server 700, a wired or wireless network interface 750 configured to connect the server 700 to the network, and an input output (I / O) interface 758.
  • the server 700 may operate an operating system based on the storage 732, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or similar operating systems.
  • an apparatus for acquiring video clips including: a processor and a memory for storing processor executable instructions, wherein the processor is configured to execute to complete the foregoing acquiring video Fragment method steps.
  • a non-transitory computer-readable storage medium is also provided.
  • the server can be executed to complete the above method of acquiring video clips step.
  • An embodiment of the present application further provides an application program, which includes one or more instructions, and the one or more instructions may be executed by a processor of the server to complete the steps of the foregoing method for acquiring video clips.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Databases & Information Systems (AREA)
  • Human Computer Interaction (AREA)
  • Information Transfer Between Computers (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
  • Reverberation, Karaoke And Other Acoustics (AREA)

Abstract

本申请是关于一种获取视频片段的方法、装置和存储介质,属于音视频技术领域。所述方法包括:在演唱直播间的直播视频数据中获取视频片段时,可以使用直播视频数据中的音频数据和原表演者的音频数据,来确定直播视频数据的目标时间点对,然后根据目标时间点对中的开始时间点和结束时间点,获取目标视频片段。采用本申请,可以使截取的视频片段比较完整。

Description

获取视频片段的方法、装置、服务器和存储介质
本申请要求于2018年11月9日提交中国专利局、申请号为201811334212.8发明名称为“获取视频片段的方法、装置和存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及音视频技术领域,尤其涉及一种获取视频片段的方法、装置、服务器和存储介质。
背景技术
随着计算机技术和网络技术的发展,直播类应用程序越来越多,人们可以通过登录直播类应用程序,观看感兴趣的直播间中主播的直播节目。人们在观看直播节目的过程中,看到精彩内容时,可以对精彩内容的视频片段进行录制,然后在用户所使用终端中存储录制的视频片段或将录制的视频片段分享给其他好友。
相关技术中,直播界面中设置有录制按钮,终端在检测到表示录制按钮被操作的操作指令后,可以使用终端操作系统提供的录屏功能,开始对屏幕显示的视频数据进行录制,终端再次检测到表示录制按钮被操作的操作指令后,结束录制屏幕显示的视频数据。这样可以录制得到精彩内容的视频片段。
在实现本申请的过程中,发明人发现相关技术至少存在以下问题:
由于人们看到精彩内容后,才开始操作录制按钮,终端检测到录制按钮被操作的操作指令后才开始录制屏幕显示的视频数据,这样从用户看到精彩内容到终端开始录制屏幕显示的视频数据会间隔一段时间,这段时间的精彩内容录制不到,所以会导致精彩内容的视频片段不完整。
发明内容
为克服相关技术中存在的问题,本申请实施例提供了一种获取视频片段的方法、装置、服务器和存储介质。
根据本申请实施例的第一方面,提供了一种获取视频片段的方法,包括:
获取演唱直播间的直播视频数据;
根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,其中,所述目标时间点对包括开始时间点和结束时间点;
根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段。
根据本申请实施例的第二方面,提供了一种获取视频片段的装置,包括:
获取单元,被配置为获取演唱直播间的直播视频数据;
确定单元,被配置为根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,其中,所述目标时间点对包括开始时间点和结束时间点;
所述获取单元,还被配置为根据所述目标时间点对,从所述直播视频数据中,获取目 标视频片段。
根据本申请实施例的第三方面,提供了一种服务器,包括:处理器和用于存储处理器可执行指令的存储器;其中,所述处理器被配置为执行下述获取视频片段的方法:
获取演唱直播间的直播视频数据;
根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,其中,所述目标时间点对包括开始时间点和结束时间点;
根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段。
根据本申请实施例的第四方面,提供了一种非临时性计算机可读存储介质,当所述存储介质中的指令由服务器的处理器执行时,使得服务器能够执行下述获取视频片段的方法:
获取演唱直播间的直播视频数据;
根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,其中,所述目标时间点对包括开始时间点和结束时间点;
根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段。
根据本申请实施例的第五方面,提供了一种应用程序,包括一条或多条指令,该一条或多条指令可以由服务器的处理器执行,以完成下述获取视频片段的方法:
获取演唱直播间的直播视频数据;
根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,其中,所述目标时间点对包括开始时间点和结束时间点;
根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段。
由以上可见,应用本申请实施例提供的方案获取视频片段时,能够获取到比较完整的视频片段。
附图说明
为了更清楚地说明本申请实施例和现有技术的技术方案,下面对实施例和现有技术中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1是根据一示例性实施例示出的一种获取视频片段的方法的流程图;
图2是根据一示例性实施例示出的一种显示视频片段的链接信息的示意图;
图3是根据一示例性实施例示出的一种第一时间段的示意图;
图4是根据一示例性实施例示出的一种获取视频片段的装置的结构框图;
图5是根据一示例性实施例示出的一种获取视频片段的装置的结构框图;
图6是根据一示例性实施例示出的一种服务器的结构框图;
图7是根据一示例性实施例示出的另一种服务器的结构框图。
具体实施方式
为使本申请的目的、技术方案、及优点更加清楚明白,以下参照附图并举实施例,对 本申请进一步详细说明。显然,所描述的实施例仅仅是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
本申请实施例提供了一种获取视频片段的方法,该方法的执行主体可以为服务器。服务器可以是直播应用程序的后台服务器,该服务器可以是CDN(Content Delivery Network,内容分发网络)服务器。该服务器中可以设置有处理器、存储器和收发器等,处理器可以用于获取视频片段、分发视频片段等处理,存储器可以用于存储获取视频片段过程中需要的数据以及产生的数据,如视频片段的视频数据、直播视频数据等,收发器可以用于接收以及发送数据,该数据可以是直播视频数据、评论信息、视频片段的链接信息等。
在介绍本申请实施例提供的获取视频片段的方案前,首先介绍一下本申请实施例的应用场景:
为便于描述将主播所使用的终端称为用户终端,将直播应用程序的后台服务器称为服务器。用户终端中安装有上述直播应用程序。
主播通过用户终端中安装的直播应用程序控制演唱直播间开播后,用户终端获取主播的直播视频数据,向上述服务器发送直播视频数据。服务器接收到用户终端发送的直播视频数据后,可以在接收到的直播视频数据中,获取目标视频片段。
或者,主播通过用户终端中安装的直播应用程序控制演唱直播间开播后,用户终端获取主播的直播视频数据,向服务器发送直播视频数据。服务器接收到直播视频数据后,存储所接收到的直播视频数据,在该演唱直播间直播结束后,可以在存储的直播视频数据中,获取目标视频片段。
本申请实施例中提供了一种获取视频片段的方法,如图1所示,该方法的执行流程可以如下:
步骤101,获取演唱直播间的直播视频数据。
其中,演唱直播间指开播后进行音乐演奏的直播间,例如,演唱直播间为演唱歌曲的直播间、弹奏乐器的直播间等。
本申请的一个实施例中,主播通过用户终端中安装的直播应用程序控制演唱直播间开播后,服务器可以接收用户终端发送的直播视频数据,并保存所接收到的直播视频数据。除此之外,上述服务器还可以确定登录该演唱直播间的账户中除主播账户之外的其它各账户,然后向其它各账户使用的终端发送接收到的直播视频数据。为便于表述,将其它各账户使用的终端称为登录终端。另外,上述各登录终端也安装有上述直播应用程序,各登录终端接收到直播视频数据后,可以通过登录终端中安装的直播应用程序在演唱直播间的直播界面,播放接收到的直播视频数据。
本申请的一个实施例中,上述直播视频数据包括带有声音信息的音频数据和带有画面信息的视频数据。也可以将上述直播视频数据理解为多媒体数据。
步骤102,根据直播视频数据中的音频数据和原表演者的音频数据,确定直播视频数 据中的目标时间点对。
由于直播视频数据是与直播时间相关的,所以,本步骤可以理解为:确定直播视频数据的目标时间点对。
其中,原表演者的音频数据可以是原唱演唱歌曲的音频数据,也可以是原表演者使用乐器进行演奏的音频数据。目标时间点对为一个或多个时间点对,每个时间点对包括一组时间点,即开始时间点和结束时间点。
具体的,上述目标时间点对包括的开始时间点和结束时间点可以为表征直播视频精彩内容开始的时间点和表征直播视频精彩内容结束的时间点。
本申请的一个实施例中,如果直播视频数据中视频数据和音频数据是分开的,服务器在获取到直播视频数据后,可以直接根据直播视频数据中的音频数据,来获取原表演者的音频数据。如果直播视频数据中视频数据和音频数据是混流的,服务器获取到直播视频数据后,可以对直播视频数据中的视频数据和音频数据进行分离,得到直播视频数据中的音频数据,然后根据直播视频数据中的音频数据,来获取原表演者的音频数据。
本申请的一个实施例中,服务器在获取到演唱直播间的直播视频数据后,可以获取直播间的开播介绍信息,在该开播介绍信息中包括主播进行直播的内容,基于主播进行直播的内容,获取原表演者的音频数据。
具体的,服务器可以对直播视频数据中的音频数据和原表演者的音频数据进行相似匹配,基于相似匹配结果,确定直播视频数据中的目标时间点对。
本申请的一个实施例中,可以首先确定出一个时间点,再基于该时间点,确定目标时间点对,相应的,步骤102的处理可以如下:
根据直播视频数据中的音频数据和原表演者的音频数据,确定直播视频数据中的第一类时间点。以第一类时间点为中心,根据预设的截取时长,确定第一类时间点对应的目标时间点对。
其中,预设的截取时长可以预先设定的,并且存储至服务器中,如10秒等。
本申请的一个实施中,服务器可以使用直播视频数据中的音频数据和原表演者的音频数据,确定出第一类时间点。然后获取预先存储的预设的截取时长,将第一类时间点减少预设的截取时长的一半,得到第一类时间点对应的开始时间点,并且将第一类时间点增加预设的截取时长的一半,得到第一类时间点对应的结束时间点。这样该开始时间点和结束时间点,可以组成第一类时间点对应的目标时间点对。
例如,假设,第一类时间点为第20秒,预设的截取时长为10秒,则预设的截取时长的一半为5秒,第一类时间对应的开始时间点为第20-5秒=第15秒,第一类时间点对应的结束时间点为第20+5秒=第25秒,这样第15秒和第25秒组成了第一类时间点对应的目标时间点对。
具体的,上述第一类时间点可以为表征直播精彩内容的时间点,第一类时间点对应的目标时间点对中开始时间点和结束时间点可以为表征直播精彩内容开始的时间点以及表 征直播精彩内容结束的时间点。
本申请的一个实施例中,在直播视频数据中的音频数据为主播演唱的歌曲的音频数据、且原表演者的音频数据为原唱演唱的歌曲的音频数据时,确定第一类时间点的方式可以如下:
对直播视频数据中的音频数据进行语音识别,得到歌曲的歌词。根据所得到的歌词,获取原唱演唱的歌曲的音频数据。对于所获得的每句歌词,将原唱演唱的歌曲的音频数据的音频特征和直播视频数据中音频数据的音频特征进行相似度确定,作为每句歌词的相似度。将相似度高于第一预设阈值的歌词中相似度最高的歌词位置对应的直播视频数据中音频数据的直播时间点,确定为直播视频数据中的第一类时间点。
其中,第一预设阈值可以预先设定的,并且存储至服务器中,如第一预设阈值可以为90%等。
本申请的一个实施例中,服务器可以采用预先存储的语音识别算法,对直播视频数据中的音频数据进行语音识别,得到主播演唱的歌曲的歌词。然后可以使用得到的歌词在预设的歌词数据库中进行查询,其中,上述歌词数据库中包括歌词、歌词所在歌曲的原唱的音频数据,确定上述得到的歌词所在的歌曲的原唱的音频数据。然后服务器对于任一句歌词,可以确定原唱演唱的歌曲的音频数据与主播演唱的歌曲的音频数据,按照音频特征提取算法,分别对原唱演唱的歌曲的音频数据和主播演唱的歌曲的音频数据进行音频特征提取,确定在该句歌词下原唱的音频特征与主播的音频特征的相似度。然后判断该相似度与第一预设阈值的大小关系,如果该相似度高于第一预设阈值,在该句歌词中确定相似度最高的位置以及该相似度最高的位置对应的直播视频数据中音频数据的直播时间点,将该时间点确定为直播视频数据的第一类时间点。如果该相似度小于或等于第一预设阈值,则不进行确定第一类时间点的处理。这样对于每一句歌词都进行上述处理,可以确定出直播视频数据的第一类时间点。
这样对于一句歌词,在直播视频数据中的音频数据的音频特征与原唱的音频数据的音频特征相似度高于第一预设阈值的情况下,进一步选择相似度最高的歌词位置,说明主播在演唱该位置的歌词时演唱的比较好。然后确定上述歌词位置在直播视频数据中对应的音频数据,并将所确定的音频数据的播放时间点作为第一类时间点,说明主播在第一类时间点直播时的音频数据与原唱的音频数据相似度最高,也可以说明主播在第一类时间点演唱的比较好,可以判别为精彩瞬间。
具体的,上述语音识别算法可以是任意一种语音识别算法,如FED(Fast Endpoint Detection,快速端点检测算法等。
另外,在直播视频数据中的音频数据为演奏乐器的音频数据时,服务器可以对直播视频数据中的音频数据进行识别,确定出主播演奏的作品的名称,然后基于该作品的名称,查找出原表演者演奏乐器的音频数据,将直播视频数据中的音频数据和原表演者演奏乐器的音频数据进行对齐处理,按照预设时长,对对齐处理后的两段音频数据分别进行分段处 理,如将两段音频数据都分别分段为5秒的音频数据,直播视频数据中的音频数据依次编号为a1,a2,a3,…,ai,…,an,原表演者演奏乐器的音频数据依次编号为b1,b2,b3,…,bi,…,bn。然后服务器可以分别提取a1的音频特征和b1的音频特征,对a1的音频特征和b1的音频特征,计算所提取音频特征的相似度,如果相似度高于第一预设阈值,则在a1中确定与b1相似度最高的位置,并获得该相似度最高的位置对应的直播时间点,将该时间点确定为第一类时间点。依此类推,即可确定出a1之后a2、a3等音频数据的第一类时间点。
另外,上述直播视频数据中的音频数据为演唱的歌曲的音频数据时,也可以按照分段处理的方式来得到第一类时间点。
上述音频特征可以是基音音频特征、音高音频特征等。上述音频特征提取算法可以是现有技术的一种算法。如现有的音乐评分系统中用于提取基音音频特征的算法,具体的提取音频特征的过程为:预加重-分帧-加窗-求短时平均能量-求自相关,经过此过程可得到基音音频特征,这一过程中涉及的主要参数为高频提升参数、帧长、帧移和清浊音阈值等。
步骤103,根据目标时间点对,从直播视频数据中,获取目标视频片段。
其中,目标视频片段指直播视频数据中包含第一音频数据的视频片段,第一音频数据为:直播视频数据的音频数据中与原表演者的音频数据中相似度满足一定条件的音频数据。
具体的,上述目标视频片段可以为直播视频数据中在目标时间点对包括的开始时间点到结束时间点之间直播的视频片段。
本申请的一个实施例中,服务器确定出目标时间点对后,可以根据直播视频数据的时间戳,找到目标时间点对中开始时间点对应的时间戳,并找到目标时间点对中结束时间点的时间戳,截取这两个时间戳之间的视频片段,作为目标视频片段。
本申请的一个实施例中,在获取到目标视频片段后,还可以将目标视频片段提供给演唱直播间的观众,相应的处理可以如下:
生成目标视频片段的链接信息。向演唱直播间的除主播账户之外的其它各账户的登录终端发送目标视频片段的链接信息,以使其它各账户的登录终端在演唱直播间的回放界面显示链接信息,或者在演唱直播间的直播结束界面显示链接信息。
由于主播在直播过程中主播账户会登录直播间,观看直播的观众的账户也会登录直播间,为此,生成目标视频片段的链接信息后,还可以向登录演唱直播间的账户中除主播账户之外的其它各账户的登录终端发送上述链接信息。由于其它各账户的登录终端均安装有直播应用程序,所以其它各账户的登录终端可以通过所安装的直播应用程序在演唱直播间的回放界面显示链接信息,或者在演唱直播间的直播结束界面显示链接信息。
其中,回放界面是用于显示回放直播视频数据的播放链接的界面,直播结束界面指直播间结束直播时显示的界面。
本申请的一个实施例中,服务器在获取到目标视频片段后,可以从目标视频片段的数据中,随机获取一张图片作为目标视频片段的封面,并为目标视频片段添加名称,如,可 以将主播演唱的歌曲的名称作为目标视频片段的名称,然后可以基于上述封面、名称以及目标视频片段的数据存储地址,生成链接信息,该链接信息可以是URL(Uniform Resource Locator,统一资源定位符)。
服务器可以确定登录演唱直播间的账户中除主播账户之外的其它各账户,向其它各账户的登录终端发送目标视频片段的链接信息。其它各账户的登录终端接收到上述链接信息后,可以通过所安装的直播应用程序在演唱直播间的回放界面显示目标视频片段的链接信息,或者可以在直播结束界面中显示目标视频片段的链接信息。例如,如图2所示,服务器获取到两个视频片段的链接信息,一个是《喵星人》的链接信息,另一个是《喵喵喵》的链接信息,上述其它各账户的登录终端可以通过所安装的直播应用程序在直播结束界面显示这两个视频片段的链接信息。具体的,图2中所示的链接信息为两个视频播放链接。
如果演唱直播间的观众想要对某个链接信息进行分享,可以选择该链接信息,然后点击对应的分享选项,观众使用的终端在检测到分享选项的点击指令后,显示进行分享的各种区域选项,如,在某个应用程序、当前直播应用程序进行分享的区域选项等。观众可以选择相应的区域选项,然后通过点击操作确定选项,观众使用的终端在检测到确定选项的点击操作后,显示编辑框,此时编辑框中显示预设的内容,如,快来看A主播演唱的B歌曲等。观众可以直接按照编辑框中显示的内容进行分享,也可以重新编辑编辑框中显示的内容,然后分享至选择的区域选项对应的区域。至此完成了一次分享过程。
本申请的一个实施例中,还提供了对第一类时间点进行筛选的过程,相应的处理可以如下:
根据直播视频数据的除主播账户之外的其它账户的互动信息,确定直播视频数据的第二类时间点。如果第一类时间点中目标时间点属于第二类时间点,则保留目标时间点,如果第一类时间点中目标时间点不属于第二类时间点,则删除目标时间点。以保留下的第一类时间点为中心,根据预设的截取时长,确定保留下的第一类时间点对应的目标时间点对。
其中,互动信息可以包括评论信息、点赞信息和礼物信息中的一种或多种。
上述目标时间点可以为第一类时间点中任一时间点。也就是,分别将第一类时间点中的每一时间点作为目标时间点,判断该目标时间点是否属于第二类时间点,从而确定保留该目标时间点,还是删除该目标时间点。
本申请的一个实施例中,在演唱直播间开播后,服务器可以存储接收到的评论信息、点赞信息和礼物信息,然后可以使用评论信息、点赞信息和礼物信息中的一种或多种,确定直播视频数据中的第二类时间点。
然后判断第一类时间点中目标时间点是否属于第二类时间点,如果属于第二类时间点,则可以保留目标时间点,如果不属于第二类时间点,则可以删除目标时间点。
然后服务器可以以保留下的第一类时间点为中心,将保留下的第一类时间点减少预设的截取时长的一半,得到保留下的第一类时间点对应的开始时间点,并且将保留下的第一类时间点增加预设的截取时长的一半,得到保留下的第一类时间点对应的结束时间点,将 开始时间点和结束时间点组成目标时间点对。这样可以基于互动信息对第一类时间点进行筛选,使截取出的视频片段包括精彩内容的概率更高。
鉴于上述描述,上述第二类时间点可以理解为:表征直播过程中观众互动频繁的时间点。
本申请的一个实施例中,还提供了使用互动信息,确定目标时间点对的方式,相应的处理可以如下:
根据直播视频数据中除主播账户之外的其它账户的互动信息,确定直播视频数据的第二类时间点。对第一类时间点和第二类时间点进行合并,将合并后的时间点进行去重处理。以去重处理后的时间点为中心,根据预设的截取时长,确定去重处理后的时间点对应的目标时间点对。
本申请的一个实施例中,在演唱直播间开播后,服务器可以存储接收到的评论信息、点赞信息和礼物信息,然后可以使用评论信息、点赞信息和礼物信息中的一种或多种,确定直播视频数据的第二类时间点。
然后将第一类时间点和第二类时间点进行合并,得到合并后的时间点,将合并后的时间点中相同的时间点删除,也就是,对时间点进行去重处理,然后服务器可以以去重后的时间点为中心,将去重后的时间点减少预设的截取时长的一半,得到去重后的时间点对应的开始时间点,并且将去重后的时间点增加预设的截取时长的一半,得到去重后的时间点对应的结束时间点。将开始时间点和结束时间点组成目标时间点对。
本申请的一个实施例中,根据直播视频数据中的互动信息,确定第二类时间点的方式有多种,以下给出几种可行的方式::
方式一,如果直播视频数据在第一时间段的礼物资源量超过第二预设阈值,则将第一时间段的中间时间点或结束时间点,确定为直播视频数据中的第二类时间点。
其中,第一时间段的时长也可以预先设定,并且存储至服务器中,如2秒等。第二预设阈值也可以预先设定,并且存储至服务器中。
本申请的一个实施例中,服务器可以在直播视频数据中,依照直播视频数据的时间戳,确定出第一时间段。其中,第一时间段可以是时长相等的时间段,另外相邻时间段之间的时间间隔可以是相等的。相邻时间段之间的时间间隔可以由相邻时间段的开始时间点确定,也可以由结束时间点确定。再者,相邻第一时间段之间可以存在重叠区域,也可以不存在重叠区域。
例如,如图3所示,直播视频数据是30分钟的视频数据,第0~2秒是第一个第一时间段t1,第1~3秒是第二个第一时间段t2,第2~4秒是第三个第一时间段t3,依此类推,选取多个第一时间段。确定每个第一时间段的开始时间点和结束时间点,依据开始时间点和结束时间点,确定在开始时间点和结束时间点之间的时间间隔内接收到的送礼请求中携带的礼物的名称和数目,统计出该时间间隔内各礼物的数目,然后服务器可以获取携带的各礼物的资源,如,“游艇”礼物50个金币,将各礼物的数目与对应的资源分别相乘,得到 各礼物的资源量,然后将各礼物的资源量相加,得到第一时间段的礼物资源量。然后服务器可以判断第一时间段的礼物资源量与第二预设阈值的大小关系,如果第一时间段的礼物资源量大于第二预设阈值,则可以确定第一时间段的中间时间点,将该中间时间点确定为直播视频数据中的第二类时间点,或者可以确定第一时间段的结束时间点,将该结束时间点确定为直播视频数据中的第二类时间点。
另外,还可以基于图像识别方式来确定礼物资源量,相应的处理可以如下:
对直播视频数据中第一时间段的图像进行礼物图像识别,得到识别出的各礼物图像的数目。根据各礼物图像的数目,确定第一时间段的礼物资源量。
本申请的一个实施例中,服务器可以从直播视频数据中获取每个第一时间段的图像,然后将图像输入到预设的礼物图像识别算法中,该算法可以是预先训练得到的算法,识别图像中包含的各礼物图像的数目,然后获取各礼物的资源,将各礼物图像的数目分别乘以对应资源,得到各礼物的资源量,然后将各礼物的资源量相加,得到第一时间段的礼物资源量。
由于礼物资源量越多反映直播的内容越精彩,所以可以使用礼物资源量确定精彩内容。
上述礼物图像可以是指图像中表示礼物的区域。
具体的,上述礼物图像识别算法可以是经过训练得到的神经网络算法。在向上述神经网络算法输入一张图像后,该神经网络算法可以输出该图像中包含的礼物图像的名称,也就是礼物的名称,以及礼物图像的数目。
方式二,如果直播视频数据在第二时间段的评论信息的数目超过第三预设阈值,则将第二时间段的中间时间点或结束时间点,确定为直播视频数据中的第二类时间点。
其中,第二时间段的时长也可以预先设点,并且存储至服务器中,如2秒等。第三预设阈值也可以预先设点,并且存储至服务器中。
本申请的一个实施例中,服务器可以在直播视频数据中,依照直播视频数据的时间戳,确定出第二时间段。其中,第二时间段可以是时长相等的时间段,另外相邻时间段之间的时间间隔可以是相等的。再者,相邻第二时间段之间可以存在重叠区域,也可以不存在重叠区域。
例如,直播视频数据是30分钟的视频数据,第0~2秒是第一个第二时间段,第1~3秒是第二个第二时间段,第2~4秒是第三个第二时间段,依此类推,选取多个第二时间段。确定每个第二时间段的开始时间点和结束时间点,然后依据开始时间点和结束时间点,确定在开始时间点和结束时间点之间的时间间隔内接收到的评论信息的数目,判断接收到的评论信息的数目与第三预设阈值的大小关系,如果接收到的评论信息的数目大于第三预设阈值,则可以确定第二时间段的中间时间点,将该中间时间点确定为直播视频数据中的第二类时间点,或者可以确定第二时间段的结束时间点,将该结束时间点确定为直播视频数据中的第二类时间点。
由于接收到的评论信息越多反映直播的内容越精彩,所以可以使用评论信息的数目确 定精彩内容。
方式三,如果直播视频数据在第三时间段内的点赞的数目超过第四预设阈值,则将第三时间段的中间时间点或结束时间点,确定为直播视频数据的第二类时间点。
其中,第三时间段的时长也可以预先设定,并且存储至服务器中,如2秒等。第四预设阈值也可以预先设定,并且存储至服务器中。在直播时,点赞是指点击直播界面中某个预设标记。
本申请的一个实施例中,服务器可以在直播视频数据中,依照直播视频数据的时间戳,确定出第三时间段。其中,第三时间段可以是时长相等的时间段,另外相邻时间段之间的时间间隔可以是相等的。再者,相邻第三时间段之间可以存在重叠区域,也可以不存在重叠区域。
例如,直播视频数据是30分钟的视频数据,第0~2秒是第一个第三时间段,第1~3秒是第二个第三时间段,第2~4秒是第三个第三时间段,依此类推,选取多个第三时间段。确定每个第三时间段的开始时间点和结束时间点,依据开始时间点和结束时间点,确定在开始时间点到结束时间点之间的时间间隔内,接收到的点赞请求的数目,也就是接收到的点赞信息的数目。判断接收到的点赞请求的数目与第四预设阈值的大小关系,如果接收到的点赞请求的数目大于第四预设阈值,则可以确定第三时间段的中间时间点,将该中间时间点确定为直播视频数据中的第二类时间点,或者可以确定第三时间段的结束时间点,将该结束时间点确定为直播视频数据中的第二类时间点。
由于接收到的点赞信息越多反映直播的内容越精彩,所以可以使用点赞信息的数目确定精彩内容。
另外,也可以同时使用上述方式一至方式三中的互动信息,确定第二类时间点,相应的处理可以如下:
本申请的一个实施例中,礼物资源量、评论信息和点赞信息均对应有一定权值,分别为A、B和C,对于第四时间段,服务器确定出的礼物资源量为x,评论信息的数目为y,点赞信息的数目为z,然后进行加权计算,得到加权值:A*x+B*y+C*z,判断该加权值与预设数值大小关系,如果该加权值大于预设数值,将第四时间段的中间时间点,确定为直播视频数据中的第二类时间点。这样综合考虑了三种互动信息确定出的第二类时间点更准确。
其中,第四时间段可以是时长相等的时间段,另外相邻时间段之间的时间间隔可以是相等的。再者,相邻第四时间段之间可以存在重叠区域,也可以不存在重叠区域。
另外,还可以在上述方式一至方式三中,选取两种互动信息,进行加权计算,确定第二类时间点,与使用三种方式的互动信息的处理方式一样,此处不再赘述。
需要说明的是,上述第一时间段、第二时间段、第三时间段和第四时间段的时长可以一样,为了使确定出的出现精彩内容的位置准确,一般第一时间段、第二时间段、第三时间段和第四时间段的时长都比较短,例如可以小于5秒。
另外,为了使确定出的目标视频片段没有重复内容,还可以在步骤102之后,步骤103之前进行如下处理:
在目标时间点对中,如果存在第一开始时间点早于第二开始时间点,且第一开始时间点对应的结束时间点早于第二开始时间点对应的结束时间点,且第二开始时间点早于第一开始时间点对应的结束时间点,则在目标时间点对中,将第一开始时间点对应的结束时间点替换为第二开始时间点对应的结束时间点,并删除第二开始时间点和第二开始时间点对应的结束时间点。
其中,第一开始时间点和第二开始时间点不一样,第一开始时间点为目标时间点对中除第二开始时间点之外的任一开始时间点,第二开始时间点为目标时间点对中除第一开始时间点之外的任一开始时间点。
也就是,第一开始时间点、第二开始时间点为目标时间点对中不同时间点对包括的开始时间点。
本申请的一个实施例中,在服务器确定出目标时间点对后,可以确定是否存在时间范围有重叠的开始时间点和结束时间点,如果存在,也即存在第一开始时间点和第二开始时间点满足:第一开始时间点早于第二开始时间点,且第一开始时间点对应的结束时间点早于第二开始时间点对应的结束时间点,且第二开始时间点早于第一开始时间点对应的结束时间点,可以在目标时间点对中,将第一开始时间点对应的结束时间点替换为第二结束时间点对应的结束时间点,并且将第二开始时间点和第二开始时间点对应的结束时间点删除。这样第一开始时间点和第一开始时间点对应的结束时间点、第二开始时间点和第二开始时间点对应的结束时间点,变成了第一开始时间点和第二开始时间点对应的结束时间点,也就是,将第一开始时间点对应的结束时间点替换成了第二开始时间点对应的结束时间点。这样在后续获取视频片段时,本来有重复内容的视频片段会合并变成一个视频片段。
例如,第一开始时间点为第10分钟23秒,第一开始时间点对应的结束时间点为第10分钟33秒,第二开始时间点为第10分钟25秒,第二开始时间点对应的结束时间点为第10分钟35秒,最终变成了第一开始时间点为第10分钟23秒,第一开始时间点对应的结束时间点变为第10分钟35秒。
本申请的一个实施例中,为了使确定出的目标视频片段没有重复内容,还可以在步骤103之后进行如下处理:
如果目标视频片段中第一视频片段的开始时间点早于第二视频片段的开始时间点,且第一视频片段的结束时间点早于第二视频片段的结束时间点,且第二视频片段的开始时间点早于第一视频片段的结束时间点,则将第一视频片段和第二视频片段进行合并。
其中,第一视频片段为目标视频片段中除第二视频片段之外的任一视频片段,第二视频片段为目标视频片段中除第一视频片段之外的任一视频片段。
也就是,第一视频片段和第二视频片段为目标视频片段中的不同视频片段。
本申请的一个实施例中,在服务器确定出目标视频片段后,服务器可以判断任意两个 视频片段是否有重叠的部分,如果有重叠部分,也即存在第一视频片段和第二视频片段,且两个视频片段满足:第一视频片段的开始时间点早于第二视频片段的开始时间点,第一视频片段的结束时间点早于第二视频片段的结束时间点,且第二视频片段的开始时间点早于第一视频片段的结束时间点。服务器可以将第一视频片段和第二视频片段进行合并。这样可以将有重复内容的视频片段合并成一个视频片段。
例如,第一视频片段为第10分钟30秒至第10分钟40秒的视频片段,第二视频片段为第10分钟35秒至第10分钟45秒的视频片段,合并后的视频片段为第10分钟30秒至第10分钟45秒的视频片段。
本申请的一个实施例中,为了使目标视频片段包括精彩内容的概率更大,可以基于互动信息,对目标视频片段进行筛选,在步骤103后可以进行如下处理:
如果目标视频片段的礼物资源量超过第五预设阈值,则保留目标视频片段。或者,如果目标视频片段的评论信息的数目超过第六预设阈值,则保留目标视频片段。或者,如果目标视频片段的点赞信息的数目超过第七预设阈值,则保留目标视频片段。
其中,第五预设阈值、第六预设阈值和第七预设阈值均可以预先设定,并且存储至服务器中。
本申请的一个实施例中,服务器在获取到目标视频片段后,可以确定目标视频片段的礼物资源量,其中,确定目标视频片段的礼物资源量的方法与确定第一时间段的礼物资源量的方法相同,此处不再赘述。然后判断该礼物资源量是否超过第五预设阈值,如果超过第五预设阈值,则进行保留,如果未超过第五预设阈值,则说明有可能不包含精彩内容,可以进行删除。
或者,服务器在获取到目标视频片段后,可以确定目标视频片段的评论信息的数目,其中,确定目标视频片段的评论信息的数目与确定第二时间段的评论信息的数目的方法相同,此处不再赘述。然后判断该评论信息的数目是否超过第六预设阈值,如果超过第六预设阈值,则进行保留,如果未超过第六预设阈值,则说明有可能不包含精彩内容,可以进行删除。
或者,服务器在获取到目标视频片段后,可以确定目标视频片段的点赞信息的数目,其中,确定目标视频片段的点赞信息的数目的方法与确定第二时间段的点赞信息的数目的方法相同,此处不再赘述。然后判断该点赞信息的数目是否超过第七预设阈值,如果超过第七预设阈值,则进行保留,如果未超过第七预设阈值,则说明有可能不包含精彩内容,可以进行删除。
这样通过互动信息可以进一步对获取出的视频片段进行过滤,使截取出的视频片段包含精彩内容的概率可以增大。
本申请的一个实施例中,在步骤103中确定出的目标视频片段的数目有可能比较多,在目标视频片段数目超过预设数目时,可以进行如下过滤处理,相应的处理可以如下:
将确定出的目标视频片段按照礼物资源量从大到小的顺序进行排序,获取最前面的预 设数目个目标视频片段,确定为最终的视频片段。或者,将确定出的目标视频片段按照评论信息的数目从大到小的顺序进行排序,获取最前面的预设数目个目标视频片段,确定为最终的视频片段。或者,将确定出的目标视频片段按照点赞信息的数目从大到小的顺序进行排序,获取最前面的预设数目个目标视频片段,确定为最终的视频片段。
其中,预设数目可以为预设的用于指示最终反馈给终端的视频片段的数目。
本申请的一个实施例中,服务器在获取到目标视频片段后,可以确定目标视频片段的礼物资源量,其中,确定目标视频片段的礼物资源量的方法与确定第一时间段的礼物资源量的方法相同,此处不再赘述。将目标视频片段按照礼物资源量从大到小的顺序进行排序,取最前面的预设数目个目标视频片段,确定为最终的视频片段。
或者,服务器在获取到目标视频片段后,可以确定目标视频片段的评论信息的数目,其中,确定目标视频片段的评论信息的数目的方法与确定第二时间段的评论信息的数目的方法相同,此处不再赘述。将目标视频片段按照评论信息的数目从大到小的顺序进行排序,取最前面的预设数目个目标视频片段,确定为最终的视频片段。
或者,服务器在获取到目标视频片段后,可以确定目标视频片段的点赞信息的数目,其中,确定目标视频片段的点赞信息的数目的方法与确定第二时间段的点赞信息的数目的方法相同,此处不再赘述。将目标视频片段按照点赞信息的数目从大到小的顺序进行排序,取最前面的预设数目个目标视频片段,确定为最终的视频片段。
另外,在此过程中,也可以结合多种互动信息进行加权处理。例如,将点赞信息的数目、评论信息的数目和礼物资源量加权之后,对目标视频片段按照加权值从大到小的顺序进行排序,取最前面的预设数目个目标视频片段,确定为终端的视频片段。
需要说明的是,确定视频片段的礼物资源量、评论信息的数目和点赞信息的数目,可以理解为:确定视频片段对应的直播时间时间的礼物资源量、评论信息的数目和点赞信息的数目。
本申请实施例中,在演唱直播间的直播视频数据中获取视频片段时,可以使用直播视频数据中的音频数据和原表演者的音频数据,确定直播视频数据中的目标时间点对,然后根据目标时间点对中的开始时间点和结束时间点,获取目标视频片段。由于是服务器直接基于直播视频数据中的音频数据和原表演者的音频数据进行视频截取,进而获得视频片段。而不需要人工手动操作录制按钮,进而也就不会存在从精彩内容开始到开始录制屏幕显示的视频数据之间的时间间隔,所以截取出的视频片段相对比较完整。
图4是根据一示例性实施例示出的一种获取视频片段的装置框图。参照图4,该装置包括获取单元411和确定单元412。
获取单元411,被配置为获取演唱直播间的直播视频数据;
确定单元412,被配置为根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,其中,所述目标时间点对包括开始时间点和结束时间点;
所述获取单元411,还被配置为根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段。
可选的,所述确定单元412,被配置为:
根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据中的第一类时间点;
以所述第一类时间点为中心,根据预设的截取时长,确定所述第一类时间点对应的目标时间点对。
可选的,所述直播视频数据中的音频数据为主播演唱的歌曲的音频数据,所述原表演者的音频数据为原唱演唱的歌曲的音频数据;
所述确定单元412,被配置为:
对所述直播视频数据中的音频数据进行语音识别,得到歌曲的歌词;
根据所述歌词,获取原唱演唱的歌曲的音频数据;
对于每句歌词,将所述原唱演唱的歌曲的音频数据的音频特征和所述直播视频数据中音频数据的音频特征进行相似度确定,作为歌词的相似度;
将相似度高于第一预设阈值的歌词中相似度最高的歌词位置对应的所述直播视频数据中音频数据的直播时间点,确定为所述直播视频数据的第一类时间点。
可选的,所述确定单元412,还被配置为:
根据所述直播视频数据的除主播账户之外的其它账户的互动信息,确定所述直播视频数据中的第二类时间点;
所述确定单元412,被配置为:
如果所述第一类时间点中目标时间点属于所述第二类时间点,则保留所述目标时间点,如果所述第一类时间点中所述目标时间点不属于所述第二类时间点,则删除所述目标时间点;
以保留下的第一类时间点为中心,根据预设的截取时长,确定所述保留下的第一类时间点对应的目标时间点对。
可选的,所述确定单元412,还被配置为:
如果所述直播视频数据在第一时间段的礼物资源量超过第二预设阈值,则将所述第一时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点;和/或,
如果所述直播视频数据在第二时间段的评论信息的数目超过第三预设阈值,则将所述第二时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点;和/或,
如果所述直播视频数据在第三时间段的点赞的数目超过第四预设阈值,则将所述第三时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点。
可选的,所述确定单元412,还被配置为:
对所述直播视频数据中的所述第一时间段的图像进行礼物图像识别,得到识别出的各 礼物图像的数目;
根据所述各礼物图像的数目,确定所述第一时间段的礼物资源量。
可选的,所述确定单元412,还被配置为:
在所述目标时间点对中,如果存在第一开始时间点早于第二开始时间点,且所述第一开始时间点对应的结束时间点早于所述第二开始时间点对应的结束时间点,且所述第二开始时间点早于所述第一开始时间点对应的结束时间点,则在所述目标时间点对中,将所述第一开始时间点对应的结束时间点替换为所述第二开始时间点对应的结束时间点,并删除所述第二开始时间点和所述第二开始时间点对应的结束时间点,其中,所述第一开始时间点、第二开始时间点为所述目标时间点对中不同时间点对包括的开始时间点。
可选的,所述确定单元412,还被配置为:
生成目标视频片段的链接信息;
如图5所示,所述装置还包括:
发送单元413,被配置为向所述演唱直播间的除主播账户之外的其它各账户的登录终端发送所述链接信息,以使所述其它各账户的登录终端在所述演唱直播间的回放界面显示所述链接信息,或者在所述演唱直播间的直播结束界面显示所述链接信息。
可选的,所述获取单元411,还被配置为:
如果所述目标视频片段的礼物资源量超过第五预设阈值,则保留所述目标视频片段;或者,
如果所述目标视频片段的评论信息的数目超过第六预设阈值,则保留所述目标视频片段;或者,
如果所述目标视频片段的点赞信息的数目超过第七预设阈值,则保留所述目标视频片段。
本申请实施例中,在演唱直播间的直播视频数据中获取视频片段时,可以使用直播视频数据中的音频数据和原表演者的音频数据,确定直播视频数据中的目标时间点对,然后根据目标时间点对中的开始时间点和结束时间点,获取目标视频片段。由于是服务器直接基于直播视频数据中的音频数据和原表演者的音频数据进行视频截取,进而获得视频片段。而不需要人工手动操作录制按钮,进而也就不会存在从精彩内容开始到开始录制屏幕显示的视频数据之间的时间间隔,所以截取出的视频片段相对比较完整。
关于上述实施例中的装置,其中各个单元执行操作的具体方式已经在有关该方法的实施例中进行了详细描述,此处将不做详细阐述说明。
图6是本申请实施例提供的一种服务器的结构示意图,该服务器600可因配置或性能不同而产生比较大的差异,可以包括一个或一个以上处理器(central processing units,CPU)601和一个或一个以上的存储器602,其中,所述存储器602中存储有至少一条指令,所述至少一条指令由所述处理器601加载并执行以实现上述获取视频片段的方法的步骤。
本申请的一个实施例中,还提供了一种服务器,包括:处理器和用于存储处理器可执 行指令的存储器,其中,所述处理器被配置为执行以完成上述获取视频片段的方法的步骤。
图7是根据一示例性实施例示出的一种服务器700的框图。参照图7,服务器700包括处理组件722,其进一步包括一个或多个处理器,以及由存储器732所代表的存储器资源,用于存储可由处理组件722的执行的指令,例如应用程序。存储器732中存储的应用程序可以包括一个或一个以上的每一个对应于一组指令的模块。此外,处理组件722被配置为执行指令,以执行上述方法获取视频片段的方法的步骤。
服务器700还可以包括一个电源组件726被配置为执行服务器700的电源管理,一个有线或无线网络接口750被配置为将服务器700连接到网络,和一个输入输出(I/O)接口758。服务器700可以操作基于存储在存储器732的操作系统,例如Windows ServerTM,Mac OS XTM,UnixTM,LinuxTM,FreeBSDTM或类似操作系统。
本申请的一个实施例中,还提供了一种获取视频片段的装置,包括:处理器和用于存储处理器可执行指令的存储器,其中,所述处理器被配置为执行以完成上述获取视频片段的方法的步骤。
本申请一个实施例中,还提供了一种非临时性计算机可读存储介质,当所述存储介质中的指令由服务器的处理器执行时,使得服务器能够执行以完成上述获取视频片段的方法的步骤。
本申请实施例中,还提供了一种应用程序,包括一条或多条指令,该一条或多条指令可以由服务器的处理器执行,以完成上述获取视频片段的方法的步骤。
本领域技术人员在考虑说明书及实践这里申请的发明后,将容易想到本申请的其它实施方案。本申请旨在涵盖本申请的任何变型、用途或者适应性变化,这些变型、用途或者适应性变化遵循本申请的一般性原理并包括本申请未申请的本技术领域中的公知常识或惯用技术手段。说明书和实施例仅被视为示例性的,本申请的真正范围和精神由下面的权利要求指出。
应当理解的是,本申请并不局限于上面已经描述并在附图中示出的精确结构,并且可以在不脱离其范围进行各种修改和改变。本申请的范围仅由所附的权利要求来限制。
以上所述仅为本申请的较佳实施例而已,并不用以限制本申请,凡在本申请的精神和原则之内,所做的任何修改、等同替换、改进等,均应包含在本申请保护的范围之内。

Claims (28)

  1. 一种获取视频片段的方法,包括:
    获取演唱直播间的直播视频数据;
    根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,其中,所述目标时间点对包括开始时间点和结束时间点;
    根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段。
  2. 根据权利要求1所述的方法,所述根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,包括:
    根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据中的第一类时间点;
    以所述第一类时间点为中心,根据预设的截取时长,确定所述第一类时间点对应的目标时间点对。
  3. 根据权利要求2所述的方法,所述直播视频数据中的音频数据为主播演唱的歌曲的音频数据,所述原表演者的音频数据为原唱演唱的歌曲的音频数据;
    所述根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据中的第一类时间点,包括:
    对所述直播视频数据中的音频数据进行语音识别,得到歌曲的歌词;
    根据所述歌词,获取原唱演唱的歌曲的音频数据;
    对于每句歌词,将所述原唱演唱的歌曲的音频数据的音频特征和所述直播视频数据中音频数据的音频特征进行相似度确定,作为歌词的相似度;
    将相似度高于第一预设阈值的歌词中相似度最高的歌词位置对应的所述直播视频数据中音频数据的直播时间点,确定为所述直播视频数据的第一类时间点。
  4. 根据权利要求2所述的方法,所述方法还包括:
    根据所述直播视频数据的除主播账户之外的其它账户的互动信息,确定所述直播视频数据中的第二类时间点;
    所述以所述第一类时间点为中心,根据预设的截取时长,确定所述第一类时间点对应的目标时间点对,包括:
    如果所述第一类时间点中目标时间点属于所述第二类时间点,则保留所述目标时间点,如果所述第一类时间点中所述目标时间点不属于所述第二类时间点,则删除所述目标时间点;
    以保留下的第一类时间点为中心,根据预设的截取时长,确定所述保留下的第一类时间点对应的目标时间点对。
  5. 根据权利要求4所述的方法,所述根据所述直播视频数据中除主播账户之外的其它账户的互动信息,确定所述直播视频数据中的第二类时间点,包括:
    如果所述直播视频数据在第一时间段的礼物资源量超过第二预设阈值,则将所述第一 时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点;和/或,
    如果所述直播视频数据在第二时间段的评论信息的数目超过第三预设阈值,则将所述第二时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点;和/或,
    如果所述直播视频数据在第三时间段的点赞的数目超过第四预设阈值,则将所述第三时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点。
  6. 根据权利要求5所述的方法,所述方法还包括:
    对所述直播视频数据中的所述第一时间段的图像进行礼物图像识别,得到识别出的各礼物图像的数目;
    根据所述各礼物图像的数目,确定所述第一时间段的礼物资源量。
  7. 根据权利要求1至6任一所述的方法,所述方法还包括:
    在所述目标时间点对中,如果存在第一开始时间点早于第二开始时间点,且所述第一开始时间点对应的结束时间点早于所述第二开始时间点对应的结束时间点,且所述第二开始时间点早于所述第一开始时间点对应的结束时间点,则在所述目标时间点对中,将所述第一开始时间点对应的结束时间点替换为所述第二开始时间点对应的结束时间点,并删除所述第二开始时间点和所述第二开始时间点对应的结束时间点,其中,所述第一开始时间点、第二开始时间点为所述目标时间点对中不同时间点对包括的开始时间点。
  8. 根据权利要求1至6任一所述的方法,所述方法还包括:
    生成目标视频片段的链接信息;
    向所述演唱直播间的除主播账户之外的其它各账户的登录终端发送所述链接信息,以使所述其它各账户的登录终端在所述演唱直播间的回放界面显示所述链接信息,或者在所述演唱直播间的直播结束界面显示所述链接信息。
  9. 根据权利要求1至6任一所述的方法,所述根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段之后,还包括:
    如果所述目标视频片段的礼物资源量超过第五预设阈值,则保留所述目标视频片段;或者,
    如果所述目标视频片段的评论信息的数目超过第六预设阈值,则保留所述目标视频片段;或者,
    如果所述目标视频片段的点赞信息的数目超过第七预设阈值,则保留所述目标视频片段。
  10. 一种获取视频片段的装置,包括:
    获取单元,被配置为获取演唱直播间的直播视频数据;
    确定单元,被配置为根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,其中,所述目标时间点对包括开始时间点和结束时间点;
    所述获取单元,还被配置为根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段。
  11. 根据权利要求10所述的装置,所述确定单元,被配置为:
    根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据中的第一类时间点;
    以所述第一类时间点为中心,根据预设的截取时长,确定所述第一类时间点对应的目标时间点对。
  12. 根据权利要求11所述的装置,所述直播视频数据中的音频数据为主播演唱的歌曲的音频数据,所述原表演者的音频数据为原唱演唱的歌曲的音频数据;
    所述确定单元,被配置为:
    对所述直播视频数据中的音频数据进行语音识别,得到歌曲的歌词;
    根据所述歌词,获取原唱演唱的歌曲的音频数据;
    对于每句歌词,将所述原唱演唱的歌曲的音频特征和所述直播视频数据中音频数据的音频特征进行相似度确定,作为歌词的相似度;
    将相似度高于第一预设阈值的歌词中相似度最高的歌词位置对应的所述直播视频数据中音频数据的直播时间点,确定为所述直播视频数据的第一类时间点。
  13. 根据权利要求11所述的装置,所述确定单元,还被配置为:
    根据所述直播视频数据的除主播账户之外的其它账户的互动信息,确定所述直播视频数据中的第二类时间点;
    所述确定单元,被配置为:
    如果所述第一类时间点中目标时间点属于所述第二类时间点,则保留所述目标时间点,如果所述第一类时间点中所述目标时间点不属于所述第二类时间点,则删除所述目标时间点;
    以保留下的第一类时间点为中心,根据预设的截取时长,确定所述保留下的第一类时间点对应的目标时间点对。
  14. 根据权利要求13所述的装置,所述确定单元,还被配置为:
    如果所述直播视频数据在第一时间段的礼物资源量超过第二预设阈值,则将所述第一时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点;和/或,
    如果所述直播视频数据在第二时间段的评论信息的数目超过第三预设阈值,则将所述第二时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点;和/或,
    如果所述直播视频数据在第三时间段的点赞的数目超过第四预设阈值,则将所述第三时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点。
  15. 根据权利要求14所述的装置,所述确定单元,还被配置为:
    对所述直播视频数据中的所述第一时间段的图像进行礼物图像识别,得到识别出的各 礼物图像的数目;
    根据所述各礼物图像的数目,确定所述第一时间段的礼物资源量。
  16. 根据权利要求10至15任一所述的装置,所述确定单元,还被配置为:
    在所述目标时间点对中,如果存在第一开始时间点早于第二开始时间点,且所述第一开始时间点对应的结束时间点早于所述第二开始时间点对应的结束时间点,且所述第二开始时间点早于所述第一开始时间点对应的结束时间点,则在所述目标时间点对中,将所述第一开始时间点对应的结束时间点替换为所述第二开始时间点对应的结束时间点,并删除所述第二开始时间点和所述第二开始时间点对应的结束时间点,其中,所述第一开始时间点、第二开始时间点为所述目标时间点对中不同时间点对包括的开始时间点。
  17. 根据权利要求10至15任一所述的装置,所述确定单元,还被配置为:
    生成目标视频片段的链接信息;
    所述装置还包括:
    发送单元,被配置为向所述演唱直播间的除主播账户之外的其它各账户的登录终端发送所述链接信息,以使所述其它各账户的登录终端在所述演唱直播间的回放界面显示所述链接信息,或者在所述演唱直播间的直播结束界面显示所述链接信息。
  18. 根据权利要求10至15任一所述的装置,所述获取单元,还被配置为:
    如果所述目标视频片段的礼物资源量超过第五预设阈值,则保留所述目标视频片段;或者,
    如果所述目标视频片段的评论信息的数目超过第六预设阈值,则保留所述目标视频片段;或者,
    如果所述目标视频片段的点赞信息的数目超过第七预设阈值,则保留所述目标视频片段。
  19. 一种服务器,包括:处理器和用于存储所述处理器可执行指令的存储器;
    其中,所述处理器被配置为执行获取视频片段的方法步骤;
    其中,所述获取视频片段的方法包括:
    获取演唱直播间的直播视频数据;
    根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,其中,所述目标时间点对包括开始时间点和结束时间点;
    根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段。
  20. 根据权利要求19所述的服务器,所述根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据的目标时间点对,包括:
    根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据中的第一类时间点;
    以所述第一类时间点为中心,根据预设的截取时长,确定所述第一类时间点对应的目标时间点对。
  21. 根据权利要求20所述的服务器,所述直播视频数据中的音频数据为主播演唱的歌曲的音频数据,所述原表演者的音频数据为原唱演唱的歌曲的音频数据;
    所述根据所述直播视频数据中的音频数据和原表演者的音频数据,确定所述直播视频数据中的第一类时间点,包括:
    对所述直播视频数据中的音频数据进行语音识别,得到歌曲的歌词;
    根据所述歌词,获取原唱演唱的歌曲的音频数据;
    对于每句歌词,将所述原唱演唱的歌曲的音频数据的音频特征和所述直播视频数据中音频数据的音频特征进行相似度确定,作为歌词的相似度;
    将相似度高于第一预设阈值的歌词中相似度最高的歌词位置对应的所述直播视频数据中音频数据的直播时间点,确定为所述直播视频数据的第一类时间点。
  22. 根据权利要求20所述的服务器,所述方法还包括:
    根据所述直播视频数据的除主播账户之外的其它账户的互动信息,确定所述直播视频数据中的第二类时间点;
    所述以所述第一类时间点为中心,根据预设的截取时长,确定所述第一类时间点对应的目标时间点对,包括:
    如果所述第一类时间点中目标时间点属于所述第二类时间点,则保留所述目标时间点,如果所述第一类时间点中所述目标时间点不属于所述第二类时间点,则删除所述目标时间点;
    以保留下的第一类时间点为中心,根据预设的截取时长,确定所述保留下的第一类时间点对应的目标时间点对。
  23. 根据权利要求22所述的服务器,所述根据所述直播视频数据中除主播账户之外的其它账户的互动信息,确定所述直播视频数据中的第二类时间点,包括:
    如果所述直播视频数据在第一时间段的礼物资源量超过第二预设阈值,则将所述第一时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点;和/或,
    如果所述直播视频数据在第二时间段的评论信息的数目超过第三预设阈值,则将所述第二时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点;和/或,
    如果所述直播视频数据在第三时间段的点赞的数目超过第四预设阈值,则将所述第三时间段的中间时间点或结束时间点,确定为所述直播视频数据中的第二类时间点。
  24. 根据权利要求23所述的服务器,所述方法还包括:
    对所述直播视频数据中的所述第一时间段的图像进行礼物图像识别,得到识别出的各礼物图像的数目;
    根据所述各礼物图像的数目,确定所述第一时间段的礼物资源量。
  25. 根据权利要求19至24任一所述的服务器,所述方法还包括:
    在所述目标时间点对中,如果存在第一开始时间点早于第二开始时间点,且所述第一 开始时间点对应的结束时间点早于所述第二开始时间点对应的结束时间点,且所述第二开始时间点早于所述第一开始时间点对应的结束时间点,则在所述目标时间点对中,将所述第一开始时间点对应的结束时间点替换为所述第二开始时间点对应的结束时间点,并删除所述第二开始时间点和所述第二开始时间点对应的结束时间点,其中,所述第一开始时间点、第二开始时间点为所述目标时间点对中不同时间点对包括的开始时间点。
  26. 根据权利要求19至24任一所述的服务器,所述方法还包括:
    生成目标视频片段的链接信息;
    向所述演唱直播间的除主播账户之外的其它各账户的登录终端发送所述链接信息,以使所述其它各账户的登录终端在所述演唱直播间的回放界面显示所述链接信息,或者在所述演唱直播间的直播结束界面显示所述链接信息。
  27. 根据权利要求19至24任一所述的服务器,所述根据所述目标时间点对,从所述直播视频数据中,获取目标视频片段之后,还包括:
    如果所述目标视频片段的礼物资源量超过第五预设阈值,则保留所述目标视频片段;或者,
    如果所述目标视频片段的评论信息的数目超过第六预设阈值,则保留所述目标视频片段;或者,
    如果所述目标视频片段的点赞信息的数目超过第七预设阈值,则保留所述目标视频片段。
  28. 一种非临时性计算机可读存储介质,当所述存储介质中的指令由服务器的处理器执行时,使得服务器能够执行权利要求1至权利要求9任一所述的方法步骤。
PCT/CN2019/113321 2018-11-09 2019-10-25 获取视频片段的方法、装置、服务器和存储介质 Ceased WO2020093883A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
US17/257,447 US11375295B2 (en) 2018-11-09 2019-10-25 Method and device for obtaining video clip, server, and storage medium
US17/830,222 US20220303644A1 (en) 2018-11-09 2022-06-01 Method and device for obtaining video clip, server, and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201811334212.8 2018-11-09
CN201811334212.8A CN109218746B (zh) 2018-11-09 2018-11-09 获取视频片段的方法、装置和存储介质

Related Child Applications (2)

Application Number Title Priority Date Filing Date
US17/257,447 A-371-Of-International US11375295B2 (en) 2018-11-09 2019-10-25 Method and device for obtaining video clip, server, and storage medium
US17/830,222 Continuation US20220303644A1 (en) 2018-11-09 2022-06-01 Method and device for obtaining video clip, server, and storage medium

Publications (1)

Publication Number Publication Date
WO2020093883A1 true WO2020093883A1 (zh) 2020-05-14

Family

ID=64995816

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/113321 Ceased WO2020093883A1 (zh) 2018-11-09 2019-10-25 获取视频片段的方法、装置、服务器和存储介质

Country Status (3)

Country Link
US (2) US11375295B2 (zh)
CN (1) CN109218746B (zh)
WO (1) WO2020093883A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11375295B2 (en) 2018-11-09 2022-06-28 Beijing Dajia Internet Information Technology Co., Ltd. Method and device for obtaining video clip, server, and storage medium
CN119922349A (zh) * 2024-07-22 2025-05-02 书行科技(北京)有限公司 内容发布和内容浏览的方法、装置、设备、介质及产品

Families Citing this family (19)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109788308B (zh) * 2019-02-01 2022-07-15 腾讯音乐娱乐科技(深圳)有限公司 音视频处理方法、装置、电子设备及存储介质
CN109862387A (zh) * 2019-03-28 2019-06-07 北京达佳互联信息技术有限公司 直播的回看视频生成方法、装置及设备
CN110312162A (zh) * 2019-06-27 2019-10-08 北京字节跳动网络技术有限公司 精选片段处理方法、装置、电子设备及可读介质
CN110267054B (zh) * 2019-06-28 2022-01-18 广州酷狗计算机科技有限公司 一种推荐直播间的方法及装置
CN110519610B (zh) * 2019-08-14 2021-08-06 咪咕文化科技有限公司 直播资源处理方法及系统、服务器和客户端设备
CN110602566B (zh) * 2019-09-06 2021-10-01 Oppo广东移动通信有限公司 匹配方法、终端和可读存储介质
CN110659614A (zh) * 2019-09-25 2020-01-07 Oppo广东移动通信有限公司 视频采样方法、装置、设备和存储介质
CN110719440B (zh) * 2019-09-26 2021-08-06 恒大智慧科技有限公司 一种视频回放方法、装置及存储介质
CN110913240B (zh) * 2019-12-02 2022-02-22 广州酷狗计算机科技有限公司 视频截取方法、装置、服务器以及计算机可读存储介质
CN111050205B (zh) * 2019-12-13 2022-03-25 广州酷狗计算机科技有限公司 视频片段获取方法、装置、设备和存储介质
CN110996167A (zh) * 2019-12-20 2020-04-10 广州酷狗计算机科技有限公司 在视频中添加字幕的方法及装置
CN111212299B (zh) * 2020-01-16 2022-02-11 广州酷狗计算机科技有限公司 直播视频教程的获取方法、装置、服务器及存储介质
CN111221452B (zh) * 2020-02-14 2022-02-25 青岛希望鸟科技有限公司 方案讲解控制方法
CN111405311B (zh) * 2020-04-17 2023-03-21 北京达佳互联信息技术有限公司 直播节目保存方法、装置、电子设备和存储介质
CN113747217B (zh) * 2020-05-29 2024-10-22 聚好看科技股份有限公司 显示设备及提升合唱速度的方法
CN116257159A (zh) * 2021-12-10 2023-06-13 腾讯科技(深圳)有限公司 多媒体内容的分享方法、装置、设备、介质及程序产品
CN116800988A (zh) * 2022-03-14 2023-09-22 北京字跳网络技术有限公司 视频生成方法、装置、设备、存储介质和程序产品
CN115119005A (zh) * 2022-06-17 2022-09-27 广州方硅信息技术有限公司 轮播频道直播间的录播方法、服务器及存储介质
CN115866279B (zh) * 2022-09-20 2024-11-29 北京奇艺世纪科技有限公司 直播视频处理方法、装置、电子设备及可读存储介质

Citations (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20050267750A1 (en) * 2004-05-27 2005-12-01 Anonymous Media, Llc Media usage monitoring and measurement system and method
CN103871426A (zh) * 2012-12-13 2014-06-18 上海八方视界网络科技有限公司 对比用户音频与原唱音频相似度的方法及其系统
CN105939494A (zh) * 2016-05-25 2016-09-14 乐视控股(北京)有限公司 音视频片段提供方法及装置
CN106782600A (zh) * 2016-12-29 2017-05-31 广州酷狗计算机科技有限公司 音频文件的评分方法及装置
CN106804000A (zh) * 2017-02-28 2017-06-06 北京小米移动软件有限公司 直播回放方法及装置
CN107018427A (zh) * 2017-05-10 2017-08-04 广州华多网络科技有限公司 直播分享内容处理方法及装置
CN108540854A (zh) * 2018-03-29 2018-09-14 努比亚技术有限公司 直播视频剪辑方法、终端及计算机可读存储介质
CN108600778A (zh) * 2018-05-07 2018-09-28 广州酷狗计算机科技有限公司 媒体流发送方法及装置
CN109218746A (zh) * 2018-11-09 2019-01-15 北京达佳互联信息技术有限公司 获取视频片段的方法、装置和存储介质

Family Cites Families (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7254455B2 (en) * 2001-04-13 2007-08-07 Sony Creative Software Inc. System for and method of determining the period of recurring events within a recorded signal
US8171030B2 (en) * 2007-06-18 2012-05-01 Zeitera, Llc Method and apparatus for multi-dimensional content search and video identification
GB2452315B (en) * 2007-08-31 2012-06-06 Sony Corp A distribution network and method
GB2526955B (en) * 2011-09-18 2016-06-15 Touchtunes Music Corp Digital jukebox device with karaoke and/or photo booth features, and associated methods
US10282390B2 (en) * 2014-02-24 2019-05-07 Sony Corporation Method and device for reproducing a content item
US9848228B1 (en) * 2014-05-12 2017-12-19 Tunespotter, Inc. System, method, and program product for generating graphical video clip representations associated with video clips correlated to electronic audio files
US9852770B2 (en) * 2015-09-26 2017-12-26 Intel Corporation Technologies for dynamic generation of a media compilation
CN106658033B (zh) * 2016-10-26 2019-12-13 广州华多网络科技有限公司 直播内容查询方法、装置和服务器
CN106791892B (zh) * 2016-11-10 2020-05-12 广州华多网络科技有限公司 一种轮麦直播的方法、装置和系统
CN107507628B (zh) * 2017-08-31 2021-01-15 广州酷狗计算机科技有限公司 唱歌评分方法、装置及终端
CN108683927B (zh) * 2018-05-07 2020-10-09 广州酷狗计算机科技有限公司 主播推荐方法、装置及存储介质

Patent Citations (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20050267750A1 (en) * 2004-05-27 2005-12-01 Anonymous Media, Llc Media usage monitoring and measurement system and method
CN103871426A (zh) * 2012-12-13 2014-06-18 上海八方视界网络科技有限公司 对比用户音频与原唱音频相似度的方法及其系统
CN105939494A (zh) * 2016-05-25 2016-09-14 乐视控股(北京)有限公司 音视频片段提供方法及装置
CN106782600A (zh) * 2016-12-29 2017-05-31 广州酷狗计算机科技有限公司 音频文件的评分方法及装置
CN106804000A (zh) * 2017-02-28 2017-06-06 北京小米移动软件有限公司 直播回放方法及装置
CN107018427A (zh) * 2017-05-10 2017-08-04 广州华多网络科技有限公司 直播分享内容处理方法及装置
CN108540854A (zh) * 2018-03-29 2018-09-14 努比亚技术有限公司 直播视频剪辑方法、终端及计算机可读存储介质
CN108600778A (zh) * 2018-05-07 2018-09-28 广州酷狗计算机科技有限公司 媒体流发送方法及装置
CN109218746A (zh) * 2018-11-09 2019-01-15 北京达佳互联信息技术有限公司 获取视频片段的方法、装置和存储介质

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11375295B2 (en) 2018-11-09 2022-06-28 Beijing Dajia Internet Information Technology Co., Ltd. Method and device for obtaining video clip, server, and storage medium
CN119922349A (zh) * 2024-07-22 2025-05-02 书行科技(北京)有限公司 内容发布和内容浏览的方法、装置、设备、介质及产品

Also Published As

Publication number Publication date
US11375295B2 (en) 2022-06-28
CN109218746B (zh) 2020-07-07
CN109218746A (zh) 2019-01-15
US20210258658A1 (en) 2021-08-19
US20220303644A1 (en) 2022-09-22

Similar Documents

Publication Publication Date Title
WO2020093883A1 (zh) 获取视频片段的方法、装置、服务器和存储介质
US12306871B2 (en) Audio identification during performance
JP6060155B2 (ja) 受信データの比較を実行しその比較に基づいて後続サービスを提供する方法及びシステム
KR101578279B1 (ko) 데이터 스트림 내 콘텐트를 식별하는 방법 및 시스템
CN114520931B (zh) 视频生成方法、装置、电子设备及可读存储介质
EP2940644A1 (en) Method, apparatus, device and system for inserting audio advertisement
CN104239442B (zh) 搜索结果的展现方法和装置
CN111460179A (zh) 多媒体信息展示方法及装置、计算机可读介质及终端设备
CN105224581A (zh) 在播放音乐时呈现图片的方法和装置
CN108920585A (zh) 音乐推荐的方法及装置、计算机可读存储介质
CN111416995A (zh) 一种基于场景识别的内容推送方法、系统及智能终端
CN110536147B (zh) 直播处理的方法、装置及系统
CN104866477B (zh) 一种信息处理方法及电子设备
CN116170613A (zh) 音频流处理方法、计算机设备和计算机程序产品
CN111741333A (zh) 直播数据获取方法、装置、计算机设备及存储介质
KR20150116970A (ko) 콘텐츠 재생 장치를 이용한 사용자 특성 기반의 광고 제공 장치 및 방법
CN115440231B (zh) 说话人识别方法、装置、存储介质、客户端和服务器
CN115866279B (zh) 直播视频处理方法、装置、电子设备及可读存储介质
US10536729B2 (en) Methods, systems, and media for transforming fingerprints to detect unauthorized media content items
CN117014677A (zh) 媒体播放对象的剪辑方法、装置、计算机设备和存储介质
CN116932809A (zh) 一种音乐信息展示方法、装置和计算机可读存储介质
JP6299531B2 (ja) 歌唱動画編集装置、歌唱動画視聴システム
KR20180034718A (ko) 마인드맵을 활용한 뮤직 제공 방법 및 이를 실행하는 서버
CN116932810A (zh) 一种音乐信息展示方法、装置和计算机可读存储介质
CN114969437A (zh) 一种视频处理方法及系统

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19881066

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19881066

Country of ref document: EP

Kind code of ref document: A1