WO2023169259A1 - 音乐热度的预测方法、装置、存储介质及电子设备 - Google Patents
音乐热度的预测方法、装置、存储介质及电子设备 Download PDFInfo
- Publication number
- WO2023169259A1 WO2023169259A1 PCT/CN2023/078757 CN2023078757W WO2023169259A1 WO 2023169259 A1 WO2023169259 A1 WO 2023169259A1 CN 2023078757 W CN2023078757 W CN 2023078757W WO 2023169259 A1 WO2023169259 A1 WO 2023169259A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- music
- music data
- data
- popularity
- video
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/78—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/783—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
- G06F16/7834—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using audio features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/20—Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
- G06F16/28—Databases characterised by their database models, e.g. relational or object models
- G06F16/284—Relational databases
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/60—Information retrieval; Database structures therefor; File system structures therefor of audio data
- G06F16/68—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/683—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/95—Retrieval from the web
- G06F16/953—Querying, e.g. by the use of web search engines
- G06F16/9537—Spatial or temporal dependent retrieval, e.g. spatiotemporal queries
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/442—Monitoring of processes or resources, e.g. detecting the failure of a recording device, monitoring the downstream bandwidth, the number of times a movie has been viewed, the storage space available from the internal hard disk
- H04N21/44204—Monitoring of content usage, e.g. the number of times a movie has been viewed, copied or the amount which has been watched
Definitions
- the present disclosure relates to the field of computer data processing technology, such as methods, devices, storage media and electronic equipment for predicting music popularity.
- the present disclosure provides a music popularity prediction method, device, storage medium and electronic device to achieve music popularity prediction in the video dimension.
- the present disclosure provides a method for predicting music popularity, including:
- Extract the audio data in the current video identify the music data corresponding to the audio data, and construct a corresponding relationship between the video and the music data;
- the predicted popularity of the music data is determined.
- the present disclosure also provides a music data prediction device, including:
- Audio extraction module set to extract audio data in the current video
- a metadata matching module configured to identify music data corresponding to the audio data and construct a corresponding relationship between the video and the music data
- the popularity prediction module is configured to determine the predicted popularity of the music data based on the associated data of the current video in the preset time interval, and/or the associated data of the music data in the preset time interval.
- the present disclosure also provides an electronic device, which includes:
- processors one or more processors
- a storage device configured to store one or more programs
- the one or more processors When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the above-mentioned prediction method of music popularity.
- the present disclosure also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the above-mentioned prediction method of music popularity.
- the present disclosure also provides a computer program product, including a computer program carried on a non-transitory computer-readable medium, where the computer program includes program code for executing the above-mentioned prediction method of music popularity.
- Figure 1 is a schematic flow chart of a method for predicting music popularity provided by an embodiment of the present disclosure
- Figure 2 is a schematic structural diagram of a music data prediction device provided by an embodiment of the present disclosure
- FIG. 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure.
- the term “include” and its variations are open-ended, ie, “including but not limited to.”
- the term “based on” means “based at least in part on.”
- the term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; and the term “some embodiments” means “at least some embodiments”. Relevant definitions of other terms will be given in the description below.
- Figure 1 is a schematic flowchart of a method for predicting music popularity provided by an embodiment of the present disclosure.
- the embodiment of the present disclosure is adapted to predict the popularity of voice metadata based on the usage of music data in a video.
- This method can be provided by an embodiment of the present disclosure.
- the music popularity prediction device can be implemented in the form of software and/or hardware, for example, through an electronic device, and the electronic device can be a computer or a server.
- the method in this embodiment includes:
- S120 Determine the predicted popularity of the music data based on the associated data of the current video in the preset time interval and/or the associated data of the music data in the preset time interval.
- the video in this embodiment can be a video uploaded by a user client.
- the video can be a short video uploaded by a client on a short video platform, or it can also be a short video uploaded by a client on a live data platform.
- the current video is determined based on the upload time of each video.
- the latest uploaded video on the video platform (such as a live broadcast platform or a short video platform) is determined as the current video, that is, the video uploaded in real time on the short video platform is processed in real time. , or, according to the preset time interval, determine the video whose upload time is within the preset time interval as the current video.
- Video is obtained by combining continuous image frames and audio data.
- Audio data is extracted from the video to process the audio data, which avoids the interference of the image frame data on the audio data and reduces the amount of data processed. Perform sound extraction and transcoding processing on the current video to obtain audio data.
- the audio data may include, but is not limited to, sounds such as voices, singing, wind, water, background sounds, and noise.
- it also includes determining whether the audio data includes a music event. If the audio data includes a music event, then continuing to perform the step of identifying music data corresponding to the audio data. If the audio data does not include music event, it indicates that no music data is used in the audio data, and the processing of the audio data is canceled to avoid invalid processing of audio data and improve the effectiveness of audio data processing.
- Determine the duration of the audio data The duration of the audio data is the same as the duration of the video.
- the audio data will be divided into multiple audio segments, and music events will be detected for the multiple audio segments respectively. There will be music events.
- the audio segment is subjected to subsequent processing. If the duration of the audio data is less than or equal to the preset duration, the audio data is used as a whole to detect music events.
- Music events may include background sound events and singing events.
- audio data can be input into a music recognition model to obtain a music event recognition result output by the music recognition model, wherein the music recognition model is trained based on audio samples and event tags corresponding to the audio samples. .
- the music recognition model has the ability to identify music events in audio data. For input audio segments, it can Identifies whether the audio segment contains a music event. Audio samples including music events are regarded as positive samples, and audio samples including laughter, chat, noise and other events are regarded as negative samples. Correspondingly, the event labels corresponding to the positive and negative samples can be positive and negative respectively.
- the initial training model is trained based on the audio samples corresponding to the positive and negative samples and the event labels corresponding to the audio samples to obtain the music recognition model.
- the initial training model may include but is not limited to long short-term memory network model, support vector machine model, etc., which are not limited here.
- music data corresponding to the audio data is determined.
- the start and end timestamps corresponding to the music events may be determined, the sub-data corresponding to the music events may be intercepted from the audio data, and the music data corresponding to the music events may be determined based on the intercepted sub-data. Avoid the interference of non-music event data on the recognition of music data and improve the recognition accuracy of music data.
- the training samples of the music recognition model also include the start and end timestamps corresponding to the music event tags in the audio samples. The music recognition model trained through the above training samples can identify whether the input audio segment includes music events, and the location of the music events. Start and end timestamps.
- a music database may be created in advance for storing music data, or for associated storage of music data identifiers and corresponding music features.
- One or more music data corresponding to the audio data are determined by feature matching the feature information of the audio data or audio sub-data with the feature information of multiple music data in the music database.
- the music features of audio data include one or more of audio fingerprint features and music cover features.
- music data can be matched based on audio fingerprint features and/or music cover features.
- identifying the music data corresponding to the audio data includes: extracting audio fingerprint features of the audio data, matching the music fingerprint features in a preset music fingerprint library, and determining whether the music data is consistent with the audio data. Matching music data, wherein the preset music fingerprint library includes pre-stored music data and corresponding music fingerprint features.
- the fingerprint extraction algorithm is called, and fingerprints are extracted from the audio data based on the fingerprint extraction algorithm to obtain audio fingerprint data.
- the audio fingerprint features correspond to the audio data that determines the audio fingerprint features one-to-one.
- the fingerprint extraction algorithm may be a landmark algorithm, which may transform the audio data into the frequency domain. For example, it may be implemented through Fourier transform to extract the energy peak feature landmark of the frequency domain audio data based on the energy peak feature landmark. Construct audio fingerprint features.
- the audio data is intercepted to obtain Intercept the audio segment and extract the music fingerprint characteristics of the intercepted audio segment.
- the audio data is intercepted to obtain Intercept the audio segment and extract the music fingerprint characteristics of the intercepted audio segment.
- the local audio data corresponding to the intercepted music events is used to determine the music fingerprint features, which avoids the interference of the local audio data of the non-music events on the music fingerprint features, and at the same time, reduces the time required to determine the music fingerprints.
- the characteristic audio data volume is conducive to the rapid extraction of music fingerprint features.
- Partial audio data of the track where the music event is located is extracted from the audio data, and music fingerprint features are extracted based on the partial audio data.
- the audio data includes multiple audio tracks.
- the audio data may include a background collection audio track and a voice collection audio track.
- the audio data in the background collection audio track may be background music or voice collection audio track.
- the audio data in may be the host's dialogue voice data; for example, the audio data in the background collection audio track may be noise, and the audio data in the voice collection audio track may be the host's singing voice.
- Music events can be included in different audio tracks at the same time, or music events can be included in one or more audio tracks independently.
- the music database includes audio fingerprint features of multiple music data.
- the audio fingerprint features corresponding to the music data include fingerprint features corresponding to the overall music data and fingerprint features corresponding to the local music data.
- it can be obtained by combining
- the music data is divided into multiple music sub-data. There may be partial data overlap between the multiple music sub-data. The lengths of different music sub-data are the same or different. There is no limit to this. Determine the fingerprint characteristics corresponding to each music sub-data. to form a music database.
- Audio data may correspond to one or more music data. Match the audio fingerprint features with the fingerprint features of multiple music data in the music database one by one. For the matched fingerprint features, obtain the music data identification corresponding to the matched fingerprint features, and obtain the first occurrence of the audio fingerprint features in the music data.
- identifying the music data corresponding to the audio data includes: extracting music cover features of the audio data, matching the music cover features with multiple music data in a preset music library, and determining the The audio data is music data that satisfies the cover conditions.
- Music data includes a large number of elements, such as rhythm, rotation, and harmony. Therefore, different changes can be introduced when reproducing the music data. For example, changes caused by the performances of users of different genders, changes caused by the performances of different musical instruments, and changes caused by different singing styles of different users. ization, changes caused by different languages, changes caused by improvisation during singing or impromptu interaction with the audience, etc. The above changes cause the music data to be performed again to be different from the music data stored in the music database, which in turn leads to differences in the fingerprint features determined respectively, resulting in matching errors.
- music cover features are extracted to match music data that meets the cover conditions, and the music data is determined from the cover dimension, thereby improving the matching accuracy of the music data.
- Music cover features include but are not limited to feature information such as pitch, beat, lyrics, melody, and melody change trends of audio data.
- the preset music library includes music data identification and corresponding music features. The music cover features of the audio data are matched with the preset music library to determine the music data that meets the cover conditions.
- the music data can be one or more indivual.
- the audio fingerprint features and the music cover features can be extracted respectively, and the music data can be matched based on the audio fingerprint features and the music cover features respectively, and the music data can be matched based on the audio fingerprint features, and
- the music data matching the music cover feature determines the target music data of the audio data, where the target music data may be a set of music data matching the audio fingerprint feature, and a set of music data matching the music cover feature. Matching music data through different feature information improves the matching success rate and accuracy of music data.
- the music data corresponding to the audio data and the video corresponding to the audio data are associated with each other.
- the music data identifier and the video identifier are stored in association.
- the video identifier may be a string used to uniquely mark the video, and the video identifier may be one or more constructs based on the video ID, video publishing timestamp, video title, video publishing user ID, etc., without limitation.
- the corresponding relationship between the music data and the video is stored in a database.
- the database can be a MySQL database, a ByteGraph database, a Hive database, etc., and is not limited to this.
- the corresponding relationship between music data and video can be one-to-one, one-to-many, many-to-one or many-to-many, etc.
- the corresponding relationship between the current video and the music data is determined, and the current video and music data are The corresponding relationship is updated to the above database.
- the video identifier can be used as an index, and the music data identifier can be added to the set of associated identifiers of the current video identifier, or the music data identifier can be used as an index, and the current video identifier can be added to the set of associated identifiers of the music data identifier. middle.
- the correspondence between the video and the music data can represent the usage of the music data in the video.
- the popularity of the music data is predicted based on the correspondence between the video and the music data.
- the popularity of the music data can be predicted based on the associated data of the video and/or the associated data of the music data.
- the preset time interval may be 1 hour, 12 hours, 24 hours, one week, one month, etc.
- different preset time intervals are set to obtain the predicted popularity of music data in different intervals.
- the popularity display formats of the music data are different. Different, for example, the heat may increase in a short period of time and last for a short time. For example, the heat may be low in a short period of time but the heat continues to increase or the heat lasts for a long time.
- the highest predicted popularity of the music data in different preset time intervals may be used as the target predicted popularity of the music data.
- the associated data of the current video includes: within the preset time interval, one or more of the number of plays of the current video and the amount of new video creations based on the current video.
- the videos respectively correspond to the publishing timestamp, that is, the timestamp of the uploading platform.
- the publishing timestamp of the video can be used as the starting time, and the time period for obtaining associated data is determined based on the preset time interval.
- the display interface of the client includes a follow shooting control.
- a new video is created based on the current video as a video shooting template, where the video shooting template includes current video data.
- a follow tag is set in the newly added video.
- the follow tag may include the current video identifier.
- the number of new videos set with the follow tag including the current video identifier within the preset time interval is counted to obtain the new video of the current video. Amount of creation.
- Determining the predicted popularity of the music data based on the associated data of the current video in a preset time interval includes: determining the popularity of the music data based on the number of plays of the current video and/or the amount of new video creations. The first predicted popularity.
- determining the first predicted popularity of the music data based on the number of times the current video is played may be when the number of times the current video is played is greater than a first quantitative threshold, indicating that the music data corresponding to the current video satisfies Heat condition, the music data is high-heat music data.
- the first quantitative threshold includes multiple data values, that is, multiple quantitative ranges.
- the first predicted popularity of the music data may be a popularity level or a popularity value.
- determining the first predicted popularity of the music data based on the amount of new video creations may include indicating the music data corresponding to the current video when the amount of new video creations of the current video is greater than a second quantity threshold. If the popularity condition is met, the music data is highly popular music data.
- the second quantitative threshold includes multiple data values, that is, multiple quantitative ranges.
- the first predicted popularity of the music data is determined based on the number of plays of the current video and the amount of new video creations, for example, when the number of plays of the current video is greater than the first quantity threshold, and the new video creations are If the amount of one or more items is greater than the second quantitative threshold, it is determined that the music data is high-popularity music data.
- determining the first predicted popularity of the music data based on the number of plays of the current video and the amount of new video creations includes: based on a first quantity threshold and the number of plays of the current video, Determine a first prediction parameter of the music data, and, based on a second quantity threshold and the amount of new video creations, determine a second prediction parameter of the music data; based on the first prediction parameter and/or the second Prediction parameters determine the first predicted popularity of the music data.
- the first quantity threshold and the second quantity threshold may be a numerical value. When the number of plays of the current video is greater than the first quantity threshold, the first prediction parameter is determined as the first numerical value.
- the second prediction parameter is determined as a second value; the first quantity threshold and the second quantity threshold can also be multiple data values, that is, multiple data ranges, and different data ranges correspond to different prediction parameters.
- the first The prediction parameters corresponding to the data ranges of the quantity threshold and the second quantity threshold may be the same or different, and the numbers of data values corresponding to the first quantity threshold and the second quantity threshold may be the same or different. Compare the number of plays of the current video with the first quantity threshold to determine the quantity range to obtain the first prediction parameter. Compare the number of new video creations with the second quantity threshold to determine the quantity range to obtain the first prediction parameter. Get the second prediction parameter.
- the first prediction parameter or the second prediction parameter is determined as the target prediction data, or the target prediction data is calculated and determined based on the first prediction parameter and the second prediction parameter, and the target prediction data is used to characterize the predicted popularity of the music data.
- the target prediction data may be obtained by performing a weighted calculation based on the first prediction parameter and the second prediction parameter.
- the weights of the first prediction parameter and the second prediction parameter may be preset, and this is not limited.
- the predicted popularity of the music data is determined according to the associated data of the music data, wherein the associated data of the music data in the preset time interval includes: within the preset time interval, the music data exists with the music data.
- the associated data of the music data in the preset time interval includes: within the preset time interval, the music data exists with the music data.
- determining the predicted popularity of the music data based on the associated data of the music data may be to compare the number of videos that have a corresponding relationship with the music data with a third quantity threshold, and determine the number of videos that have a corresponding relationship with the music data. When the number of videos is greater than the third quantity threshold, the music data is determined to be highly popular music data.
- the third quantity threshold includes a plurality of data values, i.e. Multiple quantity ranges, correspondingly, the predicted popularity of music data includes multiple popularity levels, and the number of popularity levels is not limited here.
- Determining the second predicted popularity of the music data based on the number of videos that have a corresponding relationship with the music data may include matching the number of videos that have a corresponding relationship with the music data with a plurality of quantity ranges in a third quantity threshold, and determining the second predicted popularity of the music data.
- the quantity range of the number of videos in which the data has a corresponding relationship is located, and the predicted popularity corresponding to the determined quantity range is determined as the second predicted popularity of the music data.
- the second predicted popularity of the music data may be a popularity level or a popularity value.
- determining the second predicted popularity of the music data based on the number of playbacks of the video that has a corresponding relationship with the music data may be by comparing the number of playbacks of the video that has a corresponding relationship with the music data with a fourth quantity threshold. Comparison is performed, and when the number of plays of the video corresponding to the music data is greater than the fourth quantitative threshold, the music data is determined to be high-popular music data.
- the fourth quantity threshold includes multiple data values, that is, the fourth quantity threshold corresponds to multiple quantity ranges.
- the predicted popularity of the music data includes multiple popularity levels. The number of popularity levels is not limited here.
- Determining the second predicted popularity of the music data based on the number of plays of the video that has a corresponding relationship with the music data may be a plurality of quantity ranges corresponding to the number of times of play of the video that has a corresponding relationship with the music data and a fourth quantity threshold. A comparison is performed to determine a numerical range in which the number of plays of the video corresponding to the music data is located, and the predicted popularity corresponding to the determined numerical range is determined as the second predicted popularity of the music data.
- the second predicted popularity of the music data may also be determined based on the number of videos that have a corresponding relationship with the music data and the number of plays of the videos that have a corresponding relationship with the music data. For example, when the number of videos that have a corresponding relationship with the music data is greater than a third quantitative threshold, and the number of plays of the videos that have a corresponding relationship with the music data is greater than one or more of the fourth quantitative thresholds. , determining that the music data is highly popular music data.
- a third prediction parameter of the music data is determined, and based on the existence of the corresponding relationship with the music data
- the fourth prediction parameter of the music data is determined based on the number of plays of the corresponding video and the fourth quantity threshold; and the second prediction popularity of the music data is determined based on the third prediction parameter and/or the fourth prediction parameter.
- the third quantity threshold and the fourth quantity threshold may be one value.
- the third prediction parameter is determined as the third value.
- the third prediction parameter is determined as the third value.
- the fourth prediction parameter is determined as a fourth value.
- the third quantity threshold and the fourth quantity threshold can also be multiple data values, that is, multiple data ranges. Different data ranges correspond to different prediction parameters respectively.
- the prediction parameters corresponding to the data ranges of the third quantity threshold and the fourth quantity threshold are They may be the same or different, and the numbers of data values corresponding to the third quantity threshold and the fourth quantity threshold may be the same or different.
- the third prediction parameter or the fourth prediction parameter is determined as the target prediction data, or the target prediction data is calculated and determined based on the third prediction parameter and the fourth prediction parameter, and the target prediction data is used to characterize the predicted popularity of the music data.
- the target prediction data may be obtained by performing a weighted calculation based on the third prediction parameter and the fourth prediction parameter.
- the weights of the third prediction parameter and the fourth prediction parameter may be preset, and this is not limited.
- the above-mentioned first quantity threshold, second quantity threshold, third quantity threshold and fourth quantity threshold may be the same or different, the corresponding one or more data values respectively included may be the same or different, and the multiple data ranges formed They can be the same or different, and there is no limit to this.
- the music data can also be determined based on the associated data of the current video and the associated data of the music data for popularity prediction. For example, it can be based on the first predicted popularity and the second predicted popularity of the music data. Determining target predictive popularity of music data. For example, the first predicted popularity and the second predicted popularity may be weighted, and the first predicted popularity and the second predicted popularity may be popularity values, such as popularity values corresponding to the popularity levels.
- the technical solution provided by this embodiment extracts the audio data in the video to identify the music data corresponding to the audio data, and predicts the popularity of the music data through the associated data of the video and/or the associated data of the music data, thereby achieving the goal of extracting the audio data from the video.
- Dimension predicts the popularity of music data without relying on manpower, with high prediction efficiency and low cost.
- the method further includes: obtaining music tags of the music data, inputting the music tags and prediction time intervals of the music data into the popularity prediction model, and obtaining assistance from the music data. Predict popularity; update the predicted popularity of the music data based on the auxiliary predicted popularity.
- Music tags include but are not limited to song type tags, emotion tags, genre tags, language tags, etc.
- Song type tags may include but are not limited to HIPHOP, electronic, rock, etc.
- emotion tags include but are not limited to happy, sad, etc.
- the prediction time interval can be the name of a holiday, such as Spring Festival, Christmas, Lantern Festival, National Day, etc., or it can also be a date interval such as 10.1-10.7, etc. Convert the music tags and prediction time intervals of the above-mentioned music data into input information of the popularity prediction model, for example, it can be input information in vector form, etc., and perform popularity prediction on the music data through the popularity prediction model.
- the output of the popularity prediction model can be is the popularity value of the music data, which is the auxiliary predicted popularity of the music data.
- the auxiliary predicted popularity of the music data may be a probability value.
- the popularity prediction model is trained through the historical popularity of music data and the popularity time interval.
- the predicted popularity of the music data in the above embodiment is optimized, and the accuracy of the predicted popularity of the music data is improved.
- the auxiliary predicted popularity may be accumulated to the predicted popularity of the music data in the above embodiment to obtain the final predicted popularity; or the auxiliary predicted popularity and the predicted popularity of the music data in the above embodiment may be weighted to obtain the final prediction.
- Popularity it may also be that the highest arbitrary degree among the auxiliary predicted popularity and the predicted popularity of the music data in the above embodiment is determined as the final predicted popularity.
- FIG. 2 is a schematic structural diagram of a music data prediction device provided by an embodiment of the present disclosure. As shown in Figure 2, the device includes:
- the audio extraction module 210 is configured to extract the audio data in the current video; the metadata matching module 220 is configured to identify the music data corresponding to the audio data and construct a corresponding relationship between the video and the music data; the popularity prediction module 230. Set to determine the predicted popularity of the music data based on the associated data of the current video in the preset time interval and/or the associated data of the music data in the preset time interval.
- the metadata matching module 220 is configured as:
- Extract audio fingerprint features of the audio data match the music fingerprint features in a preset music fingerprint database, and determine music data matching the audio data, wherein the preset music fingerprint database includes preset music fingerprint features.
- the preset music fingerprint database includes preset music fingerprint features.
- the metadata matching module 220 is configured as:
- Extract the music cover feature of the audio data match the music cover feature with multiple music data in a preset music library, and determine the music data of the audio data that meets the cover condition.
- the associated data of the current video in the preset time interval includes: within the preset time interval, the number of times the current video is played, the amount of new video creations based on the current video One or more of the above; accordingly, the popularity prediction module 230 is configured to: determine the first predicted popularity of the music data based on the number of times the current video is played and/or the amount of new video creations.
- the popularity prediction module 230 includes:
- the first prediction unit is configured to determine that the music data is high-popular music data if the number of plays of the current video is greater than a first quantitative threshold, and/or the amount of new video creation is greater than a second quantitative threshold;
- the second prediction unit is configured to determine the first prediction parameter of the music data based on the first quantity threshold and the number of times the current video is played, and, based on the second quantity threshold and the amount of new video creations, Determine a second prediction parameter of the music data; determine a first prediction popularity of the music data based on the first prediction parameter and/or the second prediction parameter.
- the associated data of the music data in the preset time interval includes: the number of videos corresponding to the music data in the preset time interval, the number of videos corresponding to the music data, One or more of the play times of the related videos; accordingly, the popularity prediction module 230 is set to:
- the second predicted popularity of the music data is determined based on the number of videos corresponding to the music data and/or the number of plays of the videos corresponding to the music data.
- the popularity prediction module 230 includes:
- the third prediction module is configured such that if the number of videos corresponding to the music data is greater than a third quantity threshold, and/or the number of plays of the videos corresponding to the music data is greater than the fourth number threshold, the music data is determined to be highly popular music data; or, a fourth prediction module is configured to determine the third number of the music data based on the number of videos corresponding to the music data and the third quantity threshold.
- three prediction parameters, and, based on the number of plays of the video corresponding to the music data and a fourth quantity threshold, determining a fourth prediction parameter of the music data; based on the third prediction parameter and/or the third prediction parameter Four prediction parameters are used to determine the second predicted popularity of the music data.
- the device also includes:
- the auxiliary prediction module is configured to obtain the music tag of the music data, input the music tag and prediction time interval of the music data into the popularity prediction model, and obtain the auxiliary predicted popularity of the music data; the popularity update module, It is configured to update the predicted popularity of the music data based on the auxiliary predicted popularity.
- the device provided by the embodiments of the present disclosure can execute the method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
- the multiple units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned divisions, as long as they can achieve the corresponding functions; in addition, the names of the multiple functional units are only for the convenience of distinguishing each other. , are not used to limit the protection scope of the embodiments of the present disclosure.
- Terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (Personal Digital Assistant, PDA), tablet computers (Portable Android Device, PAD), portable multimedia players Mobile terminals such as (Portable Media Player, PMP), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital television (TV), desktop computers, etc.
- PDA Personal Digital Assistant
- PAD Portable Multimedia Players Mobile terminals
- PMP Portable Multimedia Player
- vehicle-mounted terminals such as vehicle-mounted navigation terminals
- fixed terminals such as digital television (TV), desktop computers, etc.
- TV digital television
- the electronic device 400 shown in FIG. 3 is only an example and should not bring any limitations to the functions and usage scope of the embodiments of the present disclosure.
- the electronic device 400 may include a processing device (such as a central processing unit, a graphics processor, etc.) 401, which may be configured according to a program stored in a read-only memory (Read-Only Memory, ROM) 402 or from a storage device. 408 loads the program in the random access memory (Random Access Memory, RAM) 403 to perform various appropriate actions and processes. In RAM403, there are also stored electrical Various programs and data required for sub-device 400 operation.
- the processing device 401, ROM 402 and RAM 403 are connected to each other via a bus 404.
- An input/output (I/O) interface 405 is also connected to bus 404.
- the following devices can be connected to the I/O interface 405: input devices 406 including, for example, a touch screen, touch pad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; including, for example, a Liquid Crystal Display (LCD) , an output device 407 such as a speaker, a vibrator, etc.; a storage device 408 including a magnetic tape, a hard disk, etc.; and a communication device 409.
- the communication device 409 may allow the electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data.
- FIG. 3 illustrates electronic device 400 with various means, implementation or availability of all illustrated means is not required. More or fewer means may alternatively be implemented or provided.
- embodiments of the present disclosure include a computer program product including a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the method illustrated in the flowchart.
- the computer program may be downloaded and installed from the network via communication device 409, or from storage device 408, or from ROM 402.
- the processing device 401 When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
- the electronic device provided by the embodiment of the present disclosure belongs to the same concept as the prediction method of music popularity provided by the above embodiment.
- Technical details that are not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same features as the above embodiment. Effect.
- Embodiments of the present disclosure provide a computer storage medium on which a computer program is stored.
- the program is executed by a processor, the method for predicting music popularity provided by the above embodiments is implemented.
- the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the above two.
- the computer-readable storage medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any combination thereof.
- Examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard drives, RAM, ROM, Erasable Programmable Read-Only Memory (EPROM) or flash memory), optical fiber, portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.
- a computer-readable storage medium may be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device.
- a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code therein. Such propagated data signals may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.
- a computer-readable signal medium may also be a computer Any computer-readable medium other than a machine-readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
- Program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (Radio Frequency, RF), etc., or any suitable combination of the above.
- the client and server can communicate using any currently known or future developed network protocol, such as HyperText Transfer Protocol (HTTP), and can communicate with digital data in any form or medium.
- HTTP HyperText Transfer Protocol
- Communications e.g., communications network
- Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any current network for knowledge or future research and development.
- LANs Local Area Networks
- WANs Wide Area Networks
- the Internet e.g., the Internet
- end-to-end networks e.g., ad hoc end-to-end networks
- the above-mentioned computer-readable medium may be included in the above-mentioned electronic device; it may also exist independently without being assembled into the electronic device.
- the above-mentioned computer-readable medium carries one or more programs.
- the electronic device executes the above-mentioned one or more programs.
- Extract the audio data in the current video identify the music data corresponding to the audio data, and construct a corresponding relationship between the video and the music data; based on the associated data of the current video in the preset time interval, and/or, The associated data of the music data in the preset time interval determines the predicted popularity of the music data.
- Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, including but not limited to object-oriented programming languages—such as Java, Smalltalk, C++, and Includes conventional procedural programming languages—such as "C" or similar programming languages.
- the program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.
- the remote computer may be connected to the user computer through any kind of network, including a LAN or WAN, or may be connected to an external computer (eg, through the Internet using an Internet service provider).
- each block in the flowchart or block diagram may represent a module, segment, or portion of code that contains one or more logic functions that implement the specified executable instructions.
- the functions noted in the block may occur out of the order noted in the figures. For example, two blocks represented one after another may actually execute substantially in parallel. OK, they can sometimes be executed in reverse order, depending on the functionality involved.
- each block of the block diagram and/or flowchart illustration, and combinations of blocks in the block diagram and/or flowchart illustration can be implemented by special purpose hardware-based systems that perform the specified functions or operations. , or can be implemented using a combination of specialized hardware and computer instructions.
- the units involved in the embodiments of the present disclosure can be implemented in software or hardware. Among them, the name of the unit/module does not constitute a limitation on the unit itself.
- exemplary types of hardware logic components include: field programmable gate array (Field Programmable Gate Array, FPGA), application specific integrated circuit (Application Specific Integrated Circuit, ASIC), application specific standard product (Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programming Logic Device (CPLD), etc.
- a machine-readable medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
- the machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- Machine-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any suitable combination of the foregoing. Examples of machine-readable storage media would include an electrical connection based on one or more wires, a portable computer disk, a hard drive, RAM, ROM, EPROM or flash memory, optical fiber, CD-ROM, optical storage device, magnetic storage device, or Any suitable combination of the above.
- Example 1 provides a method for predicting music popularity, which method includes:
- Extract the audio data in the current video identify the music data corresponding to the audio data, and construct a corresponding relationship between the video and the music data;
- the predicted popularity of the music data is determined.
- Example 2 provides a method for predicting music popularity, which also includes:
- the identifying the music data corresponding to the audio data includes:
- the fingerprint database includes pre-stored music data and corresponding music fingerprint features.
- Example 3 provides a method for predicting music popularity, which also includes:
- the identification of music data corresponding to the audio data includes: extracting music cover features of the audio data, matching the music cover features with multiple music data in a preset music library, and determining the characteristics of the audio data. Music data that meets the cover conditions.
- Example 4 provides a method for predicting music popularity, which also includes:
- the associated data of the current video in the preset time interval includes: within the preset time interval, one or more of the number of plays of the current video and the amount of new video creations based on the current video;
- Determining the predicted popularity of the music data based on the associated data of the current video in a preset time interval includes: determining the music data based on the number of times the current video is played and/or the amount of new video creations. The first predicted popularity of the data.
- Example 5 provides a method for predicting music popularity, which also includes:
- Determining the first predicted popularity of the music data based on the number of times the current video is played and/or the amount of new video creations includes: if the number of times the current video is played is greater than a first quantity threshold, and/ Or, if the amount of newly added video creations is greater than the second quantity threshold, it is determined that the music data is high-popular music data;
- Prediction parameters based on the first prediction parameter and/or the second prediction parameter, determine the first prediction popularity of the music data.
- Example 6 provides a method for predicting music popularity, which also includes:
- the associated data of the music data in the preset time interval includes: within the preset time interval, the number of videos corresponding to the music data and the number of plays of the videos corresponding to the music data. one or more items;
- Determining the predicted popularity of the music data based on the associated data of the music data in a preset time interval includes: based on the number of videos corresponding to the music data, and/or the relationship with the music data. The number of times the video has a corresponding relationship with the music data is played, and the number of the music data is determined. 2. Predict the popularity.
- Example 7 provides a method for predicting music popularity, which also includes:
- Determining the second predicted popularity of the music data based on the number of videos corresponding to the music data and/or the number of plays of the videos corresponding to the music data includes: If the number of videos corresponding to the music data is greater than a third quantitative threshold, and/or the number of plays of the videos corresponding to the music data is greater than a fourth quantitative threshold, then it is determined that the music The data is highly popular music data;
- determine a third prediction parameter of the music data based on the number of videos corresponding to the music data and a third quantity threshold, and based on the playback of the videos corresponding to the music data
- the number of times and the fourth quantity threshold are used to determine the fourth prediction parameter of the music data; based on the third prediction parameter and/or the fourth prediction parameter, the second prediction popularity of the music data is determined.
- Example 8 provides a method for predicting music popularity, which also includes:
- the method also includes: obtaining music tags of the music data, inputting the music tags and prediction time intervals of the music data into the popularity prediction model, and obtaining auxiliary predicted popularity of the music data; based on the auxiliary The predicted popularity updates the predicted popularity of the music data.
- Example 9 provides a prediction device for music data, and the device includes:
- Audio extraction module set to extract audio data in the current video
- a metadata matching module configured to identify music data corresponding to the audio data and construct a corresponding relationship between the video and the music data
- the popularity prediction module is configured to determine the predicted popularity of the music data based on the associated data of the current video in the preset time interval, and/or the associated data of the music data in the preset time interval.
Landscapes
- Engineering & Computer Science (AREA)
- Databases & Information Systems (AREA)
- Theoretical Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Library & Information Science (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Evolutionary Computation (AREA)
- Evolutionary Biology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Bioinformatics & Computational Biology (AREA)
- Artificial Intelligence (AREA)
- Life Sciences & Earth Sciences (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
本公开提供了一种音乐热度的预测方法、装置、存储介质及电子设备。该方法包括:提取当前视频中的音频数据,识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系;基于所述当前视频在预设时间区间的关联数据,和/或,所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度。
Description
本申请要求在2022年03月08日提交中国专利局、申请号为202210220172.4的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
本公开涉及计算机数据处理技术领域,例如涉及音乐热度的预测方法、装置、存储介质及电子设备。
随着视频技术的不断发展,尤其是短视频被广大用户接受与使用,视频的热度同样带动视频中音乐的热度。
对视频中音乐热度的统计,一般是在音乐或歌曲具有热度后再被发现,并且音乐热度的统计纯依赖于人力,导致存在对音乐热度预测存在检测效率差,人力消耗大的问题。
发明内容
本公开提供了音乐热度的预测方法、装置、存储介质及电子设备,以实现在视频维度上的音乐热度预测。
第一方面,本公开提供了一种音乐热度的预测方法,包括:
提取当前视频中的音频数据,识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系;
基于所述当前视频在预设时间区间的关联数据,和/或,所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度。
第二方面,本公开还提供了一种音乐数据的预测装置,包括:
音频提取模块,设置为提取当前视频中的音频数据;
元数据匹配模块,设置为识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系;
热度预测模块,设置为基于所述当前视频在预设时间区间的关联数据,和/或,所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度。
第三方面,本公开还提供了一种电子设备,所述电子设备包括:
一个或多个处理器;
存储装置,设置为存储一个或多个程序;
当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现上述的音乐热度的预测方法。
第四方面,本公开还提供了一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由计算机处理器执行时用于执行上述的音乐热度的预测方法。
第五方面,本公开还提供了一种计算机程序产品,包括承载在非暂态计算机可读介质上的计算机程序,所述计算机程序包含用于执行上述的音乐热度的预测方法的程序代码。
图1为本公开实施例所提供的一种音乐热度的预测方法的流程示意图;
图2是本公开实施例所提供的一种音乐数据的预测装置的结构示意图;
图3是本公开实施例所提供的一种电子设备的结构示意图。
下面将参照附图描述本公开的实施例。虽然附图中显示了本公开的一些实施例,然而本公开可以通过多种形式来实现,提供这些实施例是为了理解本公开。本公开的附图及实施例仅用于示例性作用。
本公开的方法实施方式中记载的多个步骤可以按照不同的顺序执行,和/或并行执行。此外,方法实施方式可以包括附加的步骤和/或省略执行示出的步骤。本公开的范围在此方面不受限制。
本文使用的术语“包括”及其变形是开放性包括,即“包括但不限于”。术语“基于”是“至少部分地基于”。术语“一个实施例”表示“至少一个实施例”;术语“另一实施例”表示“至少一个另外的实施例”;术语“一些实施例”表示“至少一些实施例”。其他术语的相关定义将在下文描述中给出。
本公开中提及的“第一”、“第二”等概念仅用于对不同的装置、模块或单元进行区分,并非用于限定这些装置、模块或单元所执行的功能的顺序或者相互依存关系。
本公开中提及的“一个”、“多个”的修饰是示意性而非限制性的,本领域技术人员应当理解,除非在上下文另有指出,否则应该理解为“一个或多个”。
图1为本公开实施例所提供的一种音乐热度的预测方法的流程示意图,本公开实施例适应于根据视频中音乐数据的使用情况预测语音元数据的热度,该方法可以由本公开实施例提供的音乐热度的预测装置来执行,该音乐热度的预测装置可以通过软件和/或硬件的形式实现,例如,通过电子设备来实现,该电子设备可以是计算机或者服务器等。如图1,本实施例的方法包括:
S110、提取当前视频中的音频数据,识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系。
S120、基于所述当前视频在预设时间区间的关联数据,和/或,所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度。
本实施例中的视频可以是用户客户端上传的视频,示例性的,该视频可以是短视频平台上,由客户端上传的短视频,或者,还可以是直播数据平台上,客户端上传的直播视频。根据每个视频的上传时间确定当前视频,示例性的,将视频平台(例如直播平台或短视频平台)上最新上传的视频确定为当前视频,即对短视频平台上实时上传的视频进行实时处理,或者,根据预设时间间隔,对上传时间在预设时间间隔内的视频确定为当前视频。
视频由连续图像帧和音频数据组合得到,通过在视频中提取音频数据,以实现对音频数据的处理,避免了图像帧数据对音频数据的干扰,同时减少处理数据量。对当前视频进行抽声处理和转码处理,得到音频数据。
在一些实施例中,音频数据中可以包括但不限于说话声、唱歌声、风声、水声、背景音以及噪声等声音。相应的,得到音频数据之后,还包括,确定音频数据中是否包括音乐事件,若音频数据中包括音乐事件,则继续执行识别所述音频数据对应的音乐数据的步骤,若音频数据中不包括音乐事件,则表明该音频数据中未使用音乐数据,取消对该音频数据的处理,避免对音频数据的无效处理,提高音频数据处理的有效性。确定音频数据时长,音频数据时长与视频时长相同,若音频数据时长大于预设时长,则将音频数据划分为多个音频段,对多个音频段分别进行音乐事件的检测,将存在音乐事件的音频段进行后续处理,若音频数据时长小于或等于预设时长,则将音频数据作为整体进行音乐事件的检测。
音乐事件可以包括背景音事件和唱歌事件。示例性的,可以是将音频数据输入至音乐识别模型中,得到所述音乐识别模型输出的音乐事件识别结果,其中,所述音乐识别模型基于音频样本与所述音频样本对应的事件标签训练得到。音乐识别模型具有在音频数据中识别音乐事件的能力,对于输入的音频段,可
识别该音频段中是否包括音乐事件。将包括音乐事件的音频样本作为正样本,将包括笑声、聊天、噪音等事件的音频样本作为负样本,相应的,正负样本分别对应的事件标签可以是正和负。基于正负样本对应的音频样本和音频样本对应的事件标签对初始训练模型进行训练,得到音乐识别模型。其中,初始训练模型可以包括但不限于长短期记忆网络模型、支持向量机模型等,在此不做限定。
对于包括音乐事件的音频数据,确定音频数据对应的音乐数据。在一些实施例中,若音频数据大于预设时长,可以是确定音乐事件对应的起止时间戳,在音频数据中截取音乐事件对应的子数据,基于截取的子数据确定音乐事件对应的音乐数据,避免非音乐事件的数据对识别音乐数据的干扰,提高音乐数据的识别准确性。相应的,音乐识别模型的训练样本中,还包括音频样本中音乐事件标签对应的起止时间戳,通过上述训练样本训练得到的音乐识别模型可识别输入音频段中是否包括音乐事件,以及音乐事件的起止时间戳。
对应提取的音频数据,或者音频数据中的至少一个子数据,分别进行特征提取,通过提取的特征数据进行音乐数据的匹配,其中,音乐数据为音乐事件所属的完整音乐数据,例如可以是原唱音乐数据或者具有版权的音乐数据。本实施例中,可以是预先创建一音乐数据库,用于存储音乐数据,或者用于关联存储音乐数据标识和对应的音乐特征。通过将音频数据或音频子数据的特征信息与在音乐数据库中多个音乐数据的特征信息进行特征匹配,以确定音频数据对应的一个或多个音乐数据。
音频数据(包括子数据)的音乐特征包括音频指纹特征和音乐翻唱特征的一项或多项,相应的,可基于音频指纹特征和/或音乐翻唱特征进行音乐数据的匹配。
在一些实施例中,识别所述音频数据对应的音乐数据,包括:提取所述音频数据的音频指纹特征,将所述音乐指纹特征在预设音乐指纹库中进行匹配,确定与所述音频数据相匹配的音乐数据,其中,所述预设音乐指纹库中包括预存储的音乐数据和对应的音乐指纹特征。
本实施例中,调用指纹提取算法,基于该指纹提取算法对音频数据进行指纹提取,得到音频指纹数据,音频指纹特征与确定音频指纹特征的音频数据一一对应,在音频数据被划分为多个子数据的情况下,多个子数据分别对应多个音频指纹特征。在一些实施例中,指纹提取算法可以是landmark算法,可以是将音频数据变换至频域,例如可以是通过傅里叶变换实现,提取频域音频数据的能量峰值特征landmark,基于能量峰值特征landmark构造音频指纹特征。
根据音频数据中音乐事件的起止时间戳,对所述音频数据进行截取,得到
截取音频段,提取所述截取音频段的音乐指纹特征。通过剔除非音乐事件的局部音频数据,仅对截取的音乐事件对应的局部音频数据确定音乐指纹特征,避免了非音乐事件部分的局部音频数据对音乐指纹特征的干扰,同时,减少了确定音乐指纹特征的音频数据量,有利于音乐指纹特征的快速提取。
在所述音频数据中提取所述音乐事件所在音轨的局部音频数据,基于所述局部音频数据提取音乐指纹特征。音频数据包括多个音轨,示例性的,音频数据中可以包括背景采集音轨和语音采集音轨,在任一音频段中,背景采集音轨中的音频数据可以是背景音乐,语音采集音轨中的音频数据可以是主持人的对话语音数据;示例性的,背景采集音轨中的音频数据可以是噪声,语音采集音轨中的音频数据可以是主持人的唱歌语音。不同音轨中可以是同时包括音乐事件,也可以其中一个或多个音轨中独立包括音乐事件,通过提取音乐事件所在音轨的局部音频数据,剔除非音乐事件所在音轨的局部音乐数据,减少了非音乐事件的干扰,有利于提高后续音乐指纹特征的提取的准确性。
音乐数据库中包括多个音乐数据的音频指纹特征,在一些实施例中,音乐数据对应的音频指纹特征包括整体音乐数据对应的指纹特征和局部音乐数据对应的指纹特征,相应的,可以是通过将音乐数据划分为多个音乐子数据,多个音乐子数据之间可存在部分数据的重叠,不同音乐子数据的长度相同或不同,对此不作限定,确定每一音乐子数据对应的指纹特征,以形成音乐数据库。
将音频数据的音频指纹特征在音乐数据库中进行匹配,将匹配度最高的音乐数据,确定为音频数据对应的音乐数据,或者,将满足匹配度阈值的音乐数据确定为音频数据对应的音乐数据,音频数据可对应一个或多个音乐数据。将音频指纹特征与音乐数据库中多个音乐数据的指纹特征进行逐一匹配,对于匹配到的指纹特征,获取匹配到的指纹特征对应的音乐数据标识,分别获取音频指纹特征在音乐数据中出现的第一时间信息和匹配到的指纹特征在音乐数据中出现的第二时间信息,并确定第一时间信息和第二时间信息的时间差,将时间差满足时间差阈值的音乐数据标识确定为音频数据对应的音乐数据;或者,基于时间差对音乐数据标识进行排序,在排序中将包含最多相同时间差的音乐数据确定为音频数据对应的音乐数据。
在一些实施例中,识别所述音频数据对应的音乐数据,包括:提取所述音频数据的音乐翻唱特征,将所述音乐翻唱特征与预设音乐库中的多个音乐数据进行匹配,确定所述音频数据的满足翻唱条件的音乐数据。
音乐数据中包括大量元素,例如节奏、旋转与和声等,因此,音乐数据在重现演绎的情况下,可引入不同的变化。示例性的,不同性别的用户的演绎导致的变化,不同乐器的演绎导致的变化,不同用户通过不同演唱风格导致的变
化,不同的语种导致的变化,演唱过程中即兴创作或者与听众的即兴交互等导致的变化等。上述变化导致再次演绎的音乐数据与音乐数据库中存储的音乐数据不同,进而导致分别确定的指纹特征存在不同,导致匹配误差。
本实施例中,通过提取音乐翻唱特征,以匹配到满足翻唱条件的音乐数据,从翻唱维度确定音乐数据,提高音乐数据的匹配准确度。音乐翻唱特征包括但不限于音频数据的音调、节拍、歌词、旋律以及旋律变化趋势等的特征信息。相应的,预设音乐库中包括音乐数据标识以及对应的音乐特征,将音频数据的音乐翻唱特征与预设音乐库中进行匹配,确定满足翻唱条件的音乐数据,该音乐数据可以是一个或多个。
在上述实施例的基础上,对于任一音频数据,可以是分别提取音频指纹特征和音乐翻唱特征,并基于音频指纹特征和音乐翻唱特征分别匹配音乐数据,基于音频指纹特征匹配的音乐数据,以及音乐翻唱特征匹配的音乐数据确定音频数据的目标音乐数据,其中,目标音乐数据可以是音频指纹特征匹配的音乐数据,以及音乐翻唱特征匹配的音乐数据的集合。通过不同的特征信息进行音乐数据的匹配,提高了音乐数据的匹配成功率和准确性。
将音频数据对应的音乐数据,以及音频数据对应的视频构建对应关联,示例性的,将音乐数据标识和视频标识关联存储。其中视频标识可以是用于唯一标记视频的字符串,视频标识可以是基于视频ID、视频发布时间戳、视频标题、视频发布用户ID等的一项或多项构建,对此不作限定。
将音乐数据和视频的对应关系存储至数据库中,该数据库可以是MySQL数据库、ByteGraph数据库、Hive数据库等,对此不作限定。数据库中,音乐数据和视频的对应关系可以是一对一、一对多、多对一或者多对多等,对于当前视频,确定当前视频与音乐数据的对应关系,并将当前视频与音乐数据的对应关系更新至上述数据库中。示例性的,可以是将视频标识作为索引,将音乐数据标识添加到当前视频标识的关联标识集合中,还可以是将音乐数据标识作为索引,将当前视频标识添加到音乐数据标识的关联标识集合中。
视频与音乐数据的对应关系,可表征视频中音乐数据的使用情况,本实施例中,基于视频与音乐数据的对应关系,对音乐数据进行热度预测。本实施例中,可根据视频的关联数据和/或音乐数据的关联数据对音乐数据进行热度预测。获取预设时间区间内视频的关联数据和/或音乐数据的关联数据,用于对音乐数据进行热度预测,其中,预设时间区间可根据预测需求设置。在一些实施例中,预设时间区间可以是1小时、12小时、24小时、一星期或者一个月等。在一些实施例中,设置不同的预设时间区间,以得到音乐数据在不同区间内的预测热度。由于音乐数据的音乐类型、风格等的不同,导致不同音乐数据的热度展示形式
不同,例如可以是短时间内热度增大,持续时间短,例如还可以是短时间内热度较低,但热度持续增大或热度持续时间长等。通过设置进行热度预测的不同预设时间区间,便于对音乐数据进行全面的热度预测,提供热度预测的准确性。在一些实施例中,可以是将音乐数据,在不同预设时间区间的最高预测热度,作为音乐数据的目标预测热度。
当前视频的关联数据包括:在所述预设时间区间内,当前视频的播放次数、基于所述当前视频的新增视频创作量中的一项或多项。视频分别对应发布时间戳,即上传平台的时间戳,可根据视频的发布时间戳作为起始时间,并基于预设时间区间确定关联数据的获取时间段。
在任一客户端请求当前视频时,将当前视频发送至客户端,并将当前视频的播放次数加一,当前视频的播放次数可以是预设时间区间内的总播放次数。在客户端的显示界面中包括跟随拍摄控件,跟随拍摄控件被触发的情况下,基于当前视频作为视频拍摄模板进行新视频创作,其中,视频拍摄模板中包括当前视频数据。相应的,对新增视频中设置跟随标签,该跟随标签可以包括当前视频标识,统计预设时间区间内,设置有包括当前视频标识的跟随标签的新增视频数量,得到当前视频的新增视频创作量。
基于所述当前视频在预设时间区间的关联数据,确定所述音乐数据的预测热度,包括:基于所述当前视频的播放次数和/或所述新增视频创作量,确定所述音乐数据的第一预测热度。
在一些实施例中,基于所述当前视频的播放次数确定所述音乐数据的第一预测热度,可以是在当前视频的播放次数大于第一数量阈值的情况下,表明当前视频对应的音乐数据满足热度条件,该音乐数据为高热度音乐数据。在一些实施例中,第一数量阈值包括多个数据值,即多个数量范围,相应的,音乐数据的预测热度包括多个热度等级,热度等级的数量此处不作限定。基于所述当前视频的播放次数确定所述音乐数据的第一预测热度,可以是将当前视频的播放次数与第一数量阈值中的多个数量范围进行匹配,确定当前视频的播放次数所在的数量范围,将确定的数量范围对应的预测热度确定为音乐数据的第一预测热度。音乐数据的第一预测热度可以是热度等级,也可以是热度数值。
在一些实施例中,基于新增视频创作量确定所述音乐数据的第一预测热度,可以是在当前视频的新增视频创作量大于第二数量阈值的情况下,表明当前视频对应的音乐数据满足热度条件,该音乐数据为高热度音乐数据。在一些实施例中,第二数量阈值包括多个数据值,即多个数量范围,相应的,音乐数据的预测热度包括多个热度等级,热度等级的数量此处不作限定。基于新增视频创作量确定所述音乐数据的第一预测热度,可以是将当前视频的新增视频创作量
与第二数量阈值的多个数量范围进行匹配,确定当前视频的新增视频创作量所在的数量范围,将确定的数量范围对应的预测热度确定为音乐数据的第一预测热度。
在一些实施例中,基于当前视频的播放次数和新增视频创作量共同确定音乐数据的第一预测热度,例如在满足当前视频的播放次数大于第一数量阈值,以及,所述新增视频创作量大于第二数量阈值的一项或多项的情况下,确定所述音乐数据为高热度音乐数据。
在一些实施例中,基于所述当前视频的播放次数和所述新增视频创作量,确定所述音乐数据的第一预测热度,包括:基于第一数量阈值和所述当前视频的播放次数,确定所述音乐数据的第一预测参数,以及,基于第二数量阈值和所述新增视频创作量,确定所述音乐数据的第二预测参数;基于所述第一预测参数和/或第二预测参数,确定所述音乐数据的第一预测热度。其中,第一数量阈值和第二数量阈值可以是一个数值,在当前视频的播放次数大于第一数量阈值时,将第一预测参数确定为第一数值,在新增视频创作量大于第二数量阈值时,将第二预测参数确定为第二数值;第一数量阈值和第二数量阈值还可以分别为多个数据值,即多个数据范围,不同数据范围分别对应不同的预测参数,第一数量阈值和第二数量阈值的数据范围对应的预测参数可以相同或不同,第一数量阈值和第二数量阈值对应的数据值的数量可以相同或不同。将当前视频的播放次数与第一数量阈值进行比对,确定所在的数量范围,以得到第一预测参数,将新增视频创作量与第二数量阈值进行比对,确定所在的数量范围,以得到第二预测参数。
将第一预测参数或第二预测参数确定为目标预测数据,或者基于第一预测参数和第二预测参数计算确定目标预测数据,该目标预测数据用于表征音乐数据的预测热度。可以是基于第一预测参数和第二预测参数进行加权计算得到目标预测数据,第一预测参数和第二预测参数的权重可以是预先设置,对此不作限定。
在上述实施例的基础上,根据音乐数据的关联数据确定音乐数据的预测热度,其中,音乐数据在预设时间区间的关联数据包括:在所述预设时间区间内,与所述音乐数据存在对应关系的视频数量、与所述音乐数据存在对应关系的视频的播放次数中的一项或多项。
在一些实施例中,根据音乐数据的关联数据确定音乐数据的预测热度,可以是将与所述音乐数据存在对应关系的视频数量与第三数量阈值进行比对,在与音乐数据存在对应关系的视频数量大于第三数量阈值的情况下,将音乐数据确定为高热度音乐数据。在一些实施例中,第三数量阈值包括多个数据值,即
多个数量范围,相应的,音乐数据的预测热度包括多个热度等级,热度等级的数量此处不作限定。基于与音乐数据存在对应关系的视频数量确定所述音乐数据的第二预测热度,可以是将与音乐数据存在对应关系的视频数量与第三数量阈值中的多个数量范围进行匹配,确定与音乐数据存在对应关系的视频数量所在的数量范围,将确定的数量范围对应的预测热度确定为音乐数据的第二预测热度。音乐数据的第二预测热度可以是热度等级,也可以是热度数值。
在一些实施例中,根据与所述音乐数据存在对应关系的视频的播放次数确定音乐数据的第二预测热度,可以是将与所述音乐数据存在对应关系的视频的播放次数与第四数量阈值进行比对,在与所述音乐数据存在对应关系的视频的播放次数大于第四数量阈值时,将音乐数据确定为高热度音乐数据。在一些实施例中,第四数量阈值包括多个数据值,即第四数量阈值对应多个数量范围,相应的,音乐数据的预测热度包括多个热度等级,热度等级的数量此处不作限定。根据与所述音乐数据存在对应关系的视频的播放次数确定音乐数据的第二预测热度,可以是将与所述音乐数据存在对应关系的视频的播放次数与第四数量阈值对应的多个数量范围进行比对,确定与所述音乐数据存在对应关系的视频的播放次数所在的数量范围,将确定的数量范围对应的预测热度确定为音乐数据的第二预测热度。
在上述实施例的基础上,还可以是根据与所述音乐数据存在对应关系的视频数量,和与所述音乐数据存在对应关系的视频的播放次数,确定所述音乐数据的第二预测热度。例如满足与所述音乐数据存在对应关系的视频数量大于第三数量阈值,以及,所述与所述音乐数据存在对应关系的视频的播放次数大于第四数量阈值的一项或多项的情况下,确定所述音乐数据为高热度音乐数据。
在上述实施例的基础上,基于所述与所述音乐数据存在对应关系的视频数量和第三数量阈值,确定所述音乐数据的第三预测参数,以及,基于所述与所述音乐数据存在对应关系的视频的播放次数和第四数量阈值,确定所述音乐数据的第四预测参数;基于所述第三预测参数和/或第四预测参数,确定所述音乐数据的第二预测热度。
第三数量阈值和第四数量阈值可以是一个数值,在与所述音乐数据存在对应关系的视频数量大于第三数量阈值时,将第三预测参数确定为第三数值,在与所述音乐数据存在对应关系的视频的播放次数大于第四数量阈值时,将第四预测参数确定为第四数值。第三数量阈值和第四数量阈值还可以分别为多个数据值,即多个数据范围,不同数据范围分别对应不同的预测参数,第三数量阈值和第四数量阈值的数据范围对应的预测参数可以相同或不同,第三数量阈值和第四数量阈值对应的数据值的数量可以相同或不同。将与所述音乐数据存在
对应关系的视频数量与第三数量阈值进行比对,确定所在的数量范围,以得到第三预测参数,将与所述音乐数据存在对应关系的视频的播放次数与第四数量阈值进行比对,确定所在的数量范围,以得到第四预测参数。
将第三预测参数或第四预测参数确定为目标预测数据,或者基于第三预测参数和第四预测参数计算确定目标预测数据,该目标预测数据用于表征音乐数据的预测热度。可以是基于第三预测参数和第四预测参数进行加权计算得到目标预测数据,第三预测参数和第四预测参数的权重可以是预先设置,对此不作限定。上述的第一数量阈值、第二数量阈值、第三数量阈值和第四数量阈值可以是相同或不同,相应的分别包括的一个或多个数据值可以相同或不同,以及形成的多个数据范围可以相同或不同,对此不做限定。
在上述实施例的基础上,还可以是根据当前视频的关联数据和音乐数据的关联数据共同确定音乐数据进行热度预测,示例性的,可以是基于音乐数据的第一预测热度和第二预测热度确定音乐数据的目标预测热度。例如,可以是将第一预测热度和第二预测热度进行加权处理,此处的第一预测热度和第二预测热度可以是热度数值,例如热度等级对应的热度数值。
本实施例提供的技术方案,通过提取视频中的音频数据,以识别音频数据对应的音乐数据,并通过视频的关联数据和/或音乐数据的关联数据对音乐数据进行热度预测,实现了从视频维度对音乐数据进行热度预测,无需依赖于人力,预测效率高,成本低。
在上述实施例的基础上,该方法还包括:获取所述音乐数据的音乐标签,将所述音乐数据的音乐标签、预测时间区间输入至所述热度预测模型中,得到所述音乐数据的辅助预测热度;基于所述辅助预测热度更新所述音乐数据的预测热度。
音乐标签包括但不限于歌曲类型标签、情感标签、曲风标签、语种标签等,其中,歌曲类型标签可以包括但不限于HIPHOP、电子、摇滚等,情感标签包括但不限于开心、悲伤等。预测时间区间可以是节假日名称,例如春节、圣诞、元宵节、国庆等,还可以是诸如10.1-10.7的日期间隔等。将上述音乐数据的音乐标签、预测时间区间转换为热度预测模型的输入信息,例如可以是转换为向量形式的输入信息等,通过热度预测模型对音乐数据进行热度预测,该热度预测模型的输出可以是音乐数据的热度值,该热度值即为音乐数据的辅助预测热度。在一些实施例中,该音乐数据的辅助预测热度可以是概率值。其中,热度预测模型通过音乐数据的历史热度,以及热度时间区间训练得到。
通过辅助预测热度更新所述音乐数据的预测热度,对上述实施例中音乐数据的预测热度进行优化,提高音乐数据的预测热度的准确性。在一些实施例中,
可以是将辅助预测热度累积到上述实施例中音乐数据的预测热度上,得到最终的预测热度;还可以是将辅助预测热度和上述实施例中音乐数据的预测热度进行加权处理,得到最终的预测热度;还可以是将辅助预测热度和上述实施例中音乐数据的预测热度中的最高任度确定为最终的预测热度。
图2是本公开实施例所提供的一种音乐数据的预测装置的结构示意图。如图2所示,所述装置包括:
音频提取模块210,设置为提取当前视频中的音频数据;元数据匹配模块220,设置为识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系;热度预测模块230,设置为基于所述当前视频在预设时间区间的关联数据,和/或,所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度。
在上述实施例的基础上,元数据匹配模块220设置为:
提取所述音频数据的音频指纹特征,将所述音乐指纹特征在预设音乐指纹库中进行匹配,确定与所述音频数据相匹配的音乐数据,其中,所述预设音乐指纹库中包括预存储的音乐数据和对应的音乐指纹特征。
在上述实施例的基础上,元数据匹配模块220设置为:
提取所述音频数据的音乐翻唱特征,将所述音乐翻唱特征与预设音乐库中的多个音乐数据进行匹配,确定所述音频数据的满足翻唱条件的音乐数据。
在上述实施例的基础上,所述当前视频在预设时间区间的关联数据包括:在所述预设时间区间内,所述当前视频的播放次数、基于所述当前视频的新增视频创作量中的一项或多项;相应的,热度预测模块230设置为:基于所述当前视频的播放次数和/或所述新增视频创作量,确定所述音乐数据的第一预测热度。
在上述实施例的基础上,热度预测模块230包括:
第一预测单元,设置为若所述当前视频的播放次数大于第一数量阈值,和/或,所述新增视频创作量大于第二数量阈值,则确定所述音乐数据为高热度音乐数据;或者,第二预测单元,设置为基于第一数量阈值和所述当前视频的播放次数,确定所述音乐数据的第一预测参数,以及,基于第二数量阈值和所述新增视频创作量,确定所述音乐数据的第二预测参数;基于所述第一预测参数和/或第二预测参数,确定所述音乐数据的第一预测热度。
在上述实施例的基础上,所述音乐数据在预设时间区间的关联数据包括:在所述预设时间区间内,与所述音乐数据存在对应关系的视频数量、与所述音乐数据存在对应关系的视频的播放次数中的一项或多项;相应的,热度预测模块230设置为:
基于所述与所述音乐数据存在对应关系的视频数量,和/或,所述与所述音乐数据存在对应关系的视频的播放次数,确定所述音乐数据的第二预测热度。
在上述实施例的基础上,热度预测模块230包括:
第三预测模块,设置为若所述与所述音乐数据存在对应关系的视频数量大于第三数量阈值,和/或,所述与所述音乐数据存在对应关系的视频的播放次数大于第四数量阈值,则确定所述音乐数据为高热度音乐数据;或者,第四预测模块,设置为基于所述与所述音乐数据存在对应关系的视频数量和第三数量阈值,确定所述音乐数据的第三预测参数,以及,基于所述与所述音乐数据存在对应关系的视频的播放次数和第四数量阈值,确定所述音乐数据的第四预测参数;基于所述第三预测参数和/或第四预测参数,确定所述音乐数据的第二预测热度。
在上述实施例的基础上,该装置还包括:
辅助预测模块,设置为获取所述音乐数据的音乐标签,将所述音乐数据的音乐标签、预测时间区间输入至所述热度预测模型中,得到所述音乐数据的辅助预测热度;热度更新模块,设置为基于所述辅助预测热度更新所述音乐数据的预测热度。
本公开实施例所提供的装置可执行本公开任意实施例所提供的方法,具备执行方法相应的功能模块和有益效果。
上述装置所包括的多个单元和模块只是按照功能逻辑进行划分的,但并不局限于上述的划分,只要能够实现相应的功能即可;另外,多个功能单元的名称也只是为了便于相互区分,并不用于限制本公开实施例的保护范围。
下面参考图3,其示出了适于用来实现本公开实施例的电子设备(例如图3中的终端设备或服务器)400的结构示意图。本公开实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、个人数字助理(Personal Digital Assistant,PDA)、平板电脑(Portable Android Device,PAD)、便携式多媒体播放器(Portable Media Player,PMP)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字电视(Television,TV)、台式计算机等等的固定终端。图3示出的电子设备400仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图3所示,电子设备400可以包括处理装置(例如中央处理器、图形处理器等)401,其可以根据存储在只读存储器(Read-Only Memory,ROM)402中的程序或者从存储装置408加载到随机访问存储器(Random Access Memory,RAM)403中的程序而执行多种适当的动作和处理。在RAM403中,还存储有电
子设备400操作所需的多种程序和数据。处理装置401、ROM 402以及RAM 403通过总线404彼此相连。输入/输出(Input/Output,I/O)接口405也连接至总线404。
通常,以下装置可以连接至I/O接口405:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置406;包括例如液晶显示器(Liquid Crystal Display,LCD)、扬声器、振动器等的输出装置407;包括例如磁带、硬盘等的存储装置408;以及通信装置409。通信装置409可以允许电子设备400与其他设备进行无线或有线通信以交换数据。虽然图3示出了具有多种装置的电子设备400,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置409从网络上被下载和安装,或者从存储装置408被安装,或者从ROM 402被安装。在该计算机程序被处理装置401执行时,执行本公开实施例的方法中限定的上述功能。
本公开实施例提供的电子设备与上述实施例提供的音乐热度的预测方法属于同一构思,未在本实施例中详尽描述的技术细节可参见上述实施例,并且本实施例与上述实施例具有相同的效果。
本公开实施例提供了一种计算机存储介质,其上存储有计算机程序,该程序被处理器执行时实现上述实施例所提供的音乐热度的预测方法。
本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、RAM、ROM、可擦式可编程只读存储器(Erasable Programmable Read-Only Memory,EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(Compact Disc Read-Only Memory,CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算
机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、射频(Radio Frequency,RF)等等,或者上述的任意合适的组合。
在一些实施方式中,客户端、服务器可以利用诸如超文本传输协议(HyperText Transfer Protocol,HTTP)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(Local Area Network,LAN),广域网(Wide Area Network,WAN),网际网(例如,互联网)以及端对端网络(例如,ad hoc端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备:
提取当前视频中的音频数据,识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系;基于所述当前视频在预设时间区间的关联数据,和/或,所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括但不限于面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括LAN或WAN—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开多种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执
行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元/模块的名称在一种情况下并不构成对该单元本身的限定。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(Field Programmable Gate Array,FPGA)、专用集成电路(Application Specific Integrated Circuit,ASIC)、专用标准产品(Application Specific Standard Parts,ASSP)、片上系统(System on Chip,SOC)、复杂可编程逻辑设备(Complex Programming Logic Device,CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、RAM、ROM、EPROM或快闪存储器、光纤、CD-ROM、光学储存设备、磁储存设备、或上述内容的任何合适组合。
根据本公开的一个或多个实施例,【示例一】提供了一种音乐热度的预测方法,该方法包括:
提取当前视频中的音频数据,识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系;
基于所述当前视频在预设时间区间的关联数据,和/或,所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度。
根据本公开的一个或多个实施例,【示例二】提供了一种音乐热度的预测方法,还包括:
所述识别所述音频数据对应的音乐数据,包括:
提取所述音频数据的音频指纹特征,将所述音乐指纹特征在预设音乐指纹库中进行匹配,确定与所述音频数据相匹配的音乐数据,其中,所述预设音乐
指纹库中包括预存储的音乐数据和对应的音乐指纹特征。
根据本公开的一个或多个实施例,【示例三】提供了一种音乐热度的预测方法,还包括:
所述识别所述音频数据对应的音乐数据,包括:提取所述音频数据的音乐翻唱特征,将所述音乐翻唱特征与预设音乐库中的多个音乐数据进行匹配,确定所述音频数据的满足翻唱条件的音乐数据。
根据本公开的一个或多个实施例,【示例四】提供了一种音乐热度的预测方法,还包括:
所述当前视频在预设时间区间的关联数据包括:在所述预设时间区间内,所述当前视频的播放次数、基于所述当前视频的新增视频创作量中的一项或多项;
所述基于所述当前视频在预设时间区间的关联数据,确定所述音乐数据的预测热度,包括:基于所述当前视频的播放次数和/或所述新增视频创作量,确定所述音乐数据的第一预测热度。
根据本公开的一个或多个实施例,【示例五】提供了一种音乐热度的预测方法,还包括:
所述基于所述当前视频的播放次数和/或所述新增视频创作量,确定所述音乐数据的第一预测热度,包括:若所述当前视频的播放次数大于第一数量阈值,和/或,所述新增视频创作量大于第二数量阈值,则确定所述音乐数据为高热度音乐数据;
或者,基于第一数量阈值和所述当前视频的播放次数,确定所述音乐数据的第一预测参数,以及,基于第二数量阈值和所述新增视频创作量,确定所述音乐数据的第二预测参数;基于所述第一预测参数和/或第二预测参数,确定所述音乐数据的第一预测热度。
根据本公开的一个或多个实施例,【示例六】提供了一种音乐热度的预测方法,还包括:
所述音乐数据在预设时间区间的关联数据包括:在所述预设时间区间内,与所述音乐数据存在对应关系的视频数量、与所述音乐数据存在对应关系的视频的播放次数中的一项或多项;
所述基于所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度,包括:基于所述与所述音乐数据存在对应关系的视频数量,和/或,所述与所述音乐数据存在对应关系的视频的播放次数,确定所述音乐数据的第
二预测热度。
根据本公开的一个或多个实施例,【示例七】提供了一种音乐热度的预测方法,还包括:
所述基于所述与所述音乐数据存在对应关系的视频数量,和/或,所述与所述音乐数据存在对应关系的视频的播放次数,确定所述音乐数据的第二预测热度,包括:若所述与所述音乐数据存在对应关系的视频数量大于第三数量阈值,和/或,所述与所述音乐数据存在对应关系的视频的播放次数大于第四数量阈值,则确定所述音乐数据为高热度音乐数据;
或者,基于所述与所述音乐数据存在对应关系的视频数量和第三数量阈值,确定所述音乐数据的第三预测参数,以及,基于所述与所述音乐数据存在对应关系的视频的播放次数和第四数量阈值,确定所述音乐数据的第四预测参数;基于所述第三预测参数和/或第四预测参数,确定所述音乐数据的第二预测热度。
根据本公开的一个或多个实施例,【示例八】提供了一种音乐热度的预测方法,还包括:
所述方法还包括:获取所述音乐数据的音乐标签,将所述音乐数据的音乐标签、预测时间区间输入至所述热度预测模型中,得到所述音乐数据的辅助预测热度;基于所述辅助预测热度更新所述音乐数据的预测热度。
根据本公开的一个或多个实施例,【示例九】提供了音乐数据的预测装置,该装置包括:
音频提取模块,设置为提取当前视频中的音频数据;
元数据匹配模块,设置为识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系;
热度预测模块,设置为基于所述当前视频在预设时间区间的关联数据,和/或,所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度。
此外,虽然采用特定次序描绘了多个操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序执行来执行。在一定环境下,多任务和并行处理可能是有利的。同样地,虽然在上面论述中包含了多个实现细节,但是这些不应当被解释为对本公开的范围的限制。在单独的实施例的上下文中描述的一些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的多种特征也可以单独地或以任何合适的子组合的方式实现在多个实施例中。
Claims (12)
- 一种音乐热度的预测方法,包括:提取当前视频中的音频数据,识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系;基于所述当前视频在预设时间区间的关联数据和所述音乐数据在预设时间区间的关联数据中的至少之一,确定所述音乐数据的预测热度。
- 根据权利要求1所述的方法,其中,所述识别所述音频数据对应的音乐数据,包括:提取所述音频数据的音频指纹特征,将所述音乐指纹特征在预设音乐指纹库中进行匹配,确定与所述音频数据相匹配的音乐数据,其中,所述预设音乐指纹库中包括预存储的音乐数据和对应的音乐指纹特征。
- 根据权利要求1所述的方法,其中,所述识别所述音频数据对应的音乐数据,包括:提取所述音频数据的音乐翻唱特征,将所述音乐翻唱特征与预设音乐库中的多个音乐数据进行匹配,确定所述音频数据的满足翻唱条件的音乐数据。
- 根据权利要求1所述的方法,其中,所述当前视频在预设时间区间的关联数据包括:在所述预设时间区间内,所述当前视频的播放次数和基于所述当前视频的新增视频创作量中的至少一项;所述基于所述当前视频在预设时间区间的关联数据,确定所述音乐数据的预测热度,包括:基于所述当前视频的播放次数和所述新增视频创作量中的至少之一,确定所述音乐数据的第一预测热度。
- 根据权利要求4所述的方法,其中,所述基于所述当前视频的播放次数和所述新增视频创作量中的至少之一,确定所述音乐数据的第一预测热度,包括:在满足以下至少之一的情况下,确定所述音乐数据为高热度音乐数据:所述当前视频的播放次数大于第一数量阈值,所述新增视频创作量大于第二数量阈值;或者,基于第一数量阈值和所述当前视频的播放次数,确定所述音乐数据的第一预测参数,以及,基于第二数量阈值和所述新增视频创作量,确定所述音乐数据的第二预测参数;基于所述第一预测参数和所述第二预测参数中的至少之一, 确定所述音乐数据的第一预测热度。
- 根据权利要求1所述的方法,其中,所述音乐数据在预设时间区间的关联数据包括:在所述预设时间区间内,与所述音乐数据存在对应关系的视频数量和与所述音乐数据存在对应关系的视频的播放次数中的至少一项;所述基于所述音乐数据在预设时间区间的关联数据,确定所述音乐数据的预测热度,包括:基于所述与所述音乐数据存在对应关系的视频数量和/或所述与所述音乐数据存在对应关系的视频的播放次数中的至少之一,确定所述音乐数据的第二预测热度。
- 根据权利要求6所述的方法,其中,所述基于所述与所述音乐数据存在对应关系的视频数量和所述与所述音乐数据存在对应关系的视频的播放次数中的至少之一,确定所述音乐数据的第二预测热度,包括:在满足以下至少之一的情况下,确定所述音乐数据为高热度音乐数据:所述与所述音乐数据存在对应关系的视频数量大于第三数量阈值,所述与所述音乐数据存在对应关系的视频的播放次数大于第四数量阈值则;或者,基于所述与所述音乐数据存在对应关系的视频数量和第三数量阈值,确定所述音乐数据的第三预测参数,以及,基于所述与所述音乐数据存在对应关系的视频的播放次数和第四数量阈值,确定所述音乐数据的第四预测参数;基于所述第三预测参数和第四预测参数中的至少之一,确定所述音乐数据的第二预测热度。
- 根据权利要求1所述的方法,还包括:获取所述音乐数据的音乐标签,将所述音乐数据的音乐标签、预测时间区间输入至所述热度预测模型中,得到所述音乐数据的辅助预测热度;基于所述辅助预测热度更新所述音乐数据的预测热度。
- 一种音乐数据的预测装置,包括:音频提取模块,设置为提取当前视频中的音频数据;元数据匹配模块,设置为识别所述音频数据对应的音乐数据,并构建视频与所述音乐数据之间的对应关系;热度预测模块,设置为基于所述当前视频在预设时间区间的关联数据和所述音乐数据在预设时间区间的关联数据中的至少之一,确定所述音乐数据的预 测热度。
- 一种电子设备,包括:至少一个处理器;存储装置,设置为存储至少一个程序,当所述至少一个程序被所述至少一个处理器执行,使得所述至少一个处理器实现如权利要求1-8中任一所述的音乐热度的预测方法。
- 一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由计算机处理器执行时用于执行如权利要求1-8中任一所述的音乐热度的预测方法。
- 一种计算机程序产品,包括承载在非暂态计算机可读介质上的计算机程序,所述计算机程序包含用于执行如权利要求1-8中任一所述的音乐热度的预测方法的程序代码。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202210220172.4A CN114595361B (zh) | 2022-03-08 | 2022-03-08 | 一种音乐热度的预测方法、装置、存储介质及电子设备 |
| CN202210220172.4 | 2022-03-08 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023169259A1 true WO2023169259A1 (zh) | 2023-09-14 |
Family
ID=81807999
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/078757 Ceased WO2023169259A1 (zh) | 2022-03-08 | 2023-02-28 | 音乐热度的预测方法、装置、存储介质及电子设备 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN114595361B (zh) |
| WO (1) | WO2023169259A1 (zh) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114595361B (zh) * | 2022-03-08 | 2023-09-08 | 北京字跳网络技术有限公司 | 一种音乐热度的预测方法、装置、存储介质及电子设备 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20160116356A (ko) * | 2015-03-10 | 2016-10-10 | 연세대학교 산학협력단 | 신호 성분 분석을 이용한 음악 인기도 예측 시스템 및 방법 |
| CN112911331A (zh) * | 2020-04-15 | 2021-06-04 | 腾讯科技(深圳)有限公司 | 针对短视频的音乐识别方法、装置、设备及存储介质 |
| CN112948623A (zh) * | 2021-02-25 | 2021-06-11 | 杭州网易云音乐科技有限公司 | 音乐热度预测方法、装置、计算设备以及介质 |
| CN113837807A (zh) * | 2021-09-27 | 2021-12-24 | 北京奇艺世纪科技有限公司 | 热度预测方法、装置、电子设备及可读存储介质 |
| CN114595361A (zh) * | 2022-03-08 | 2022-06-07 | 北京字跳网络技术有限公司 | 一种音乐热度的预测方法、装置、存储介质及电子设备 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105430494A (zh) * | 2015-12-02 | 2016-03-23 | 百度在线网络技术(北京)有限公司 | 在播放视频的设备中识别视频中音频的方法和装置 |
| CN111723235B (zh) * | 2019-03-19 | 2023-09-26 | 百度在线网络技术(北京)有限公司 | 音乐内容识别方法、装置及设备 |
| CN113890932A (zh) * | 2020-07-02 | 2022-01-04 | 华为技术有限公司 | 一种音频控制方法、系统及电子设备 |
| CN113035163B (zh) * | 2021-05-11 | 2021-08-10 | 杭州网易云音乐科技有限公司 | 音乐作品的自动生成方法及装置、存储介质、电子设备 |
| CN114020960A (zh) * | 2021-11-15 | 2022-02-08 | 北京达佳互联信息技术有限公司 | 音乐推荐方法、装置、服务器及存储介质 |
-
2022
- 2022-03-08 CN CN202210220172.4A patent/CN114595361B/zh active Active
-
2023
- 2023-02-28 WO PCT/CN2023/078757 patent/WO2023169259A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20160116356A (ko) * | 2015-03-10 | 2016-10-10 | 연세대학교 산학협력단 | 신호 성분 분석을 이용한 음악 인기도 예측 시스템 및 방법 |
| CN112911331A (zh) * | 2020-04-15 | 2021-06-04 | 腾讯科技(深圳)有限公司 | 针对短视频的音乐识别方法、装置、设备及存储介质 |
| CN112948623A (zh) * | 2021-02-25 | 2021-06-11 | 杭州网易云音乐科技有限公司 | 音乐热度预测方法、装置、计算设备以及介质 |
| CN113837807A (zh) * | 2021-09-27 | 2021-12-24 | 北京奇艺世纪科技有限公司 | 热度预测方法、装置、电子设备及可读存储介质 |
| CN114595361A (zh) * | 2022-03-08 | 2022-06-07 | 北京字跳网络技术有限公司 | 一种音乐热度的预测方法、装置、存储介质及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN114595361A (zh) | 2022-06-07 |
| CN114595361B (zh) | 2023-09-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN109543064B (zh) | 歌词显示处理方法、装置、电子设备及计算机存储介质 | |
| WO2020253806A1 (zh) | 展示视频的生成方法、装置、设备及存储介质 | |
| CN111798821B (zh) | 声音转换方法、装置、可读存储介质及电子设备 | |
| US20250218418A1 (en) | Audio detection method and apparatus, storage medium and electronic device | |
| CN110211121B (zh) | 用于推送模型的方法和装置 | |
| CN111414512A (zh) | 一种基于语音搜索的资源推荐方法、装置及电子设备 | |
| US20160196812A1 (en) | Music information retrieval | |
| EP3092734A1 (en) | Method and device for identifying a piece of music in an audio stream | |
| CN106157979A (zh) | 一种获取人声音高数据的方法和装置 | |
| WO2023061229A1 (zh) | 视频生成方法及设备 | |
| CN109410972B (zh) | 生成音效参数的方法、装置及存储介质 | |
| CN115116472A (zh) | 音频识别方法、装置、设备及存储介质 | |
| WO2023000782A1 (zh) | 获取视频热点的方法、装置、可读介质和电子设备 | |
| CN111400542B (zh) | 音频指纹的生成方法、装置、设备及存储介质 | |
| WO2020238777A1 (zh) | 音频片段的匹配方法、装置、计算机可读介质及电子设备 | |
| WO2020015411A1 (zh) | 一种训练改编水平评价模型、评价改编水平的方法及装置 | |
| WO2023169259A1 (zh) | 音乐热度的预测方法、装置、存储介质及电子设备 | |
| US20240404548A1 (en) | Method, apparatus, device and storage medium for video recording | |
| US20250307309A1 (en) | Method and apparatus for generating song list, electronic device, and storage medium | |
| CN111898753B (zh) | 音乐转录模型的训练方法、音乐转录方法以及对应的装置 | |
| CN113204685B (zh) | 资源信息获取方法及装置、可读存储介质、电子设备 | |
| CN106775567B (zh) | 一种音效匹配方法及系统 | |
| US11609948B2 (en) | Music streaming, playlist creation and streaming architecture | |
| CN116072147A (zh) | 音乐检测模型训练方法、装置、电子设备及存储介质 | |
| CN116405616A (zh) | 转场位置确定方法、装置、介质和电子设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23765837 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 29.11.2024) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23765837 Country of ref document: EP Kind code of ref document: A1 |