WO2023169258A1 - 音频检测方法、装置、存储介质及电子设备 - Google Patents
音频检测方法、装置、存储介质及电子设备 Download PDFInfo
- Publication number
- WO2023169258A1 WO2023169258A1 PCT/CN2023/078752 CN2023078752W WO2023169258A1 WO 2023169258 A1 WO2023169258 A1 WO 2023169258A1 CN 2023078752 W CN2023078752 W CN 2023078752W WO 2023169258 A1 WO2023169258 A1 WO 2023169258A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- audio
- music
- events
- event
- metadata information
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/0008—Associated control or indicating means
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/0033—Recording/reproducing or transmission of music for electrophonic musical instruments
- G10H1/0041—Recording/reproducing or transmission of music for electrophonic musical instruments in coded form
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/78—Detection of presence or absence of voice signals
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2210/00—Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
- G10H2210/031—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2210/00—Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
- G10H2210/031—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal
- G10H2210/051—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal for extraction or detection of onsets of musical sounds or notes, i.e. note attack timings
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2240/00—Data organisation or data communication aspects, specifically adapted for electrophonic musical tools or instruments
- G10H2240/075—Musical metadata derived from musical analysis or for use in electrophonic musical instruments
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2240/00—Data organisation or data communication aspects, specifically adapted for electrophonic musical tools or instruments
- G10H2240/121—Musical libraries, i.e. musical databases indexed by musical parameters, wavetables, indexing schemes using musical parameters, musical rule bases or knowledge bases, e.g. for automatic composing methods
- G10H2240/131—Library retrieval, i.e. searching a database or selecting a specific musical piece, segment, pattern, rule or parameter set
- G10H2240/141—Library retrieval matching, i.e. any of the steps of matching an inputted segment or phrase with musical database contents, e.g. query by humming, singing or playing; the steps may include, e.g. musical analysis of the input, musical feature extraction, query formulation, or details of the retrieval process
Definitions
- the present disclosure relates to the field of audio processing technology, such as audio detection methods, devices, storage media and electronic equipment.
- the audio detection method cannot collect music-related statistical data in the audio to be detected (such as music duration, music playback start and end time, etc.).
- the present disclosure provides audio detection methods, devices, storage media and electronic equipment to achieve accurate acquisition of statistical data in audio to be detected.
- an audio detection method including:
- Metadata information matching the music event is determined, and statistics in the detected audio are determined based on the metadata information.
- an audio detection device including:
- a music event identification module configured to obtain audio segments in the detected audio and identify music events in the audio segments
- a statistical data determination module is configured to determine metadata information matching the music event, and determine statistical data in the detected audio based on the metadata information.
- the present disclosure also provides an electronic device, which includes:
- processors one or more processors
- a storage device configured to store one or more programs
- the one or more A processor implements the above audio detection method.
- the present disclosure also provides a storage medium containing computer-executable instructions, which when executed by a computer processor are used to perform the above audio detection method.
- the present disclosure also provides a computer program product, including a computer program carried on a non-transitory computer-readable medium, where the computer program includes program code for executing the above audio detection method.
- Figure 1 is a schematic flow chart of an audio detection method provided by an embodiment of the present disclosure
- Figure 2 is a schematic flow chart of another audio detection method provided by an embodiment of the present disclosure.
- Figure 3 is a schematic flow chart of another audio detection method provided by an embodiment of the present disclosure.
- Figure 4 is a schematic flow chart of another audio detection method provided by an embodiment of the present disclosure.
- Figure 5 is a schematic flow chart of another audio detection method provided by an embodiment of the present disclosure.
- Figure 6 is a schematic structural diagram of an audio detection device provided by an embodiment of the present disclosure.
- FIG. 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure.
- the term “include” and its variations are open-ended, ie, “including but not limited to.”
- the term “based on” means “based at least in part on.”
- the term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; and the term “some embodiments” means “at least some embodiments”. Relevant definitions of other terms will be given in the description below.
- Figure 1 is a schematic flow chart of an audio detection method provided by an embodiment of the present disclosure.
- the embodiment of the present disclosure is adapted to automatically obtain statistical data of music events in audio.
- This method can be performed by an audio detection device provided by an embodiment of the present disclosure.
- the audio detection device can be implemented in the form of software and/or hardware, and implemented through electronic equipment.
- the electronic equipment can be a mobile terminal or a personal computer (Personal Computer, PC) terminal, etc.
- the method in this embodiment includes:
- the electronic device may be any electronic device with audio and video playback functions and/or audio and video processing functions, and may include but is not limited to smart phones, wearable devices, computers, servers, and other devices.
- the above-mentioned electronic device can obtain the detected audio in a variety of ways.
- the detected audio can be collected in real time through an audio collection device, or the detected audio can be retrieved from a preset storage location or other devices.
- the embodiment of the present disclosure does not limit the method of obtaining the detected audio.
- Detected audio refers to audio that requires statistical data detection, which can include but is not limited to audio in live videos, audio in videos, broadcast audio, etc., and is not limited to this.
- obtaining the detected audio may be to extract audio data from a video (such as a real-time live video or an offline video) as the detected audio.
- the detected audio is divided into multiple audio segments, and recognition processing is performed on each audio segment.
- the detected audio is real-time data
- the audio collected in real time is divided into audio segments in sequence, and the obtained audio segments are recognized and processed in real time;
- the audio segments can be divided according to the timing of the audio segments.
- Each audio segment is identified and processed in turn.
- the obtained multiple audio segments may be processed in parallel to improve processing efficiency.
- the audio segment may be audio data with a preset time length, and the audio segment may include one or more of music, environmental sounds, speech, noise and other events.
- the duration of the audio segment may be preset, for example, determined based on the recognition accuracy, and is not limited to this.
- the duration of the audio segment may be 20 seconds.
- Music events may refer to sound events characterized by one or more of elements such as rhythm (such as beat, tempo, and articulation), pitch (such as melody and harmony), dynamics (such as the volume of a sound or note), and may include But it is not limited to events such as background music, a cappella singing, etc.
- At least one sound feature can be extracted through any feature extraction method (such as Mel cepstrum coefficient extraction method, linear prediction coefficient extraction method, etc.), Compare the extracted sound features with music features in a music database, and determine whether the audio segment contains music events based on the comparison results, where the music database may refer to a database containing multiple music features.
- the audio segment can be recognized through a music recognition model, and whether the audio segment contains a music event is determined based on the recognition result.
- the music recognition model can use music, chat sounds, noise, etc.
- Audio data including music are used as positive samples, and samples excluding music, such as chat sounds and noise audio data, are used as negative samples.
- the music recognition model is trained based on the above sample data. When the training end conditions are met, a model with music event recognition function is obtained. This embodiment does not limit the method of identifying music events.
- the metadata information is the description information of the music metadata including the music characteristics in the music event.
- the metadata information may be a tag formed by multiple description information of the music metadata, wherein the description information of the music metadata It may include but is not limited to music spectrum information, music name, music type, singer, composer and other information, which is not limited.
- the metadata information may be in the form of music name-singer/performer.
- Music metadata is characterized by metadata information. Metadata information is unique and can uniquely represent music metadata. Using metadata information as a statistical dimension can improve the reliability of music event statistics in audio, and then use metadata information to Statistics of music events in multiple audio segments can improve the accuracy of statistical data corresponding to music events.
- determining the metadata information matching the music event may be by extracting music features in the music event, matching the music features with music features corresponding to multiple music metadata, and matching the successfully matched music elements.
- the metadata information of the data is determined as metadata information matching the music event.
- music features include but are not limited to feature information such as pitch, beat, lyrics, etc.
- the above feature information is extracted for the music event, and the extracted feature information is matched in the preset metadata database to obtain Metadata information matching the music event, wherein the preset metadata database may contain multiple metadata information and feature information corresponding to the metadata information.
- the music signature may be an audio fingerprint signature.
- the audio fingerprint features of the audio segment are extracted, matched in the fingerprint feature database based on the audio fingerprint features, and metadata information matching the music event is determined, where the fingerprint feature database includes music metadata and corresponding fingerprint features.
- the music metadata can correspond to multiple fingerprint features, and the music metadata is divided into multiple music sub-data, and between the multiple music sub-data There may be some overlap of data, and the fingerprint characteristics corresponding to each music sub-data are determined separately.
- matching the audio fingerprint characteristics of the audio segment with the fingerprint characteristics of the music metadata may be to match the audio fingerprint characteristics of the audio segment with the fingerprint characteristics of multiple music subdata in the music metadata respectively.
- the audio fingerprint feature of the music sub-data is successfully matched with the fingerprint feature of any music sub-data, then it is determined that the music sub-data
- the metadata information corresponding to the music event in the audio segment is determined according to the metadata information of the music metadata to which it belongs.
- Statistical data is the result of counting music events in multiple audio segments, which is music statistics.
- the statistical data may include but is not limited to multiple music playback durations in the detected audio, music playback start time and music playback stop time, multiple receiving users during the music playback process (such as audio listening users or viewing users of the video to which the audio belongs).
- Quantity and other information, the type of statistical data in statistical data can be determined according to business needs, and there is no limit to this.
- the audio duration corresponding to all music events corresponding to each metadata information in the detected audio can be counted based on the metadata information corresponding to each music event; it can also be based on the metadata information corresponding to the music event. and the timestamp of the music event, determine the continuous music events corresponding to the metadata information, and the number of continuous music events corresponding to each metadata information in the detected audio, to obtain the application status of music in the detected audio; it can also be statistics
- the audio interval corresponding to each metadata information in the detected audio, as well as the number of receiving users for each audio interval, are used to evaluate the traffic-draining ability of the music corresponding to each metadata information.
- the music event after obtaining the statistical data, it may also include: obtaining the music metadata corresponding to the music event according to the metadata information corresponding to each music event, and comparing the music metadata in the audio segment according to the music metadata corresponding to the music event.
- the music event is repaired and the repaired music event is obtained to avoid the noise included in the detected audio interfering with the music event and causing the music event to be unclear.
- repairing the music event in the audio segment according to the music metadata corresponding to the music event may intercept the music sub-data corresponding to the music event in the music metadata, and replace the audio data of the music event based on the music sub-data.
- the statistical data may also include: performing operations such as cropping and splicing the audio segments according to the metadata information corresponding to the music events in each audio segment to obtain one or more new audio segments. For example, the Audio data corresponding to the same metadata information in the detected audio is trimmed and spliced.
- the audio detection method achieves preliminary identification of music events in multiple audio segments by acquiring audio segments in the detected audio and identifying music events in the audio segments; and determines elements matching the music events.
- Data information realizes the matching and acquisition of reference data, providing a reference basis for obtaining statistical data; statistics of music events in multiple audio segments are performed based on the metadata information obtained by matching, and statistical data in the detected audio is obtained, realizing the Audio identifies and counts music dimensions, which facilitates subsequent analysis of the detected audio based on statistical data, and enables accurate acquisition of statistical data on music events.
- FIG. 2 is a schematic flow chart of another audio detection method provided by an embodiment of the present disclosure.
- the method of this embodiment can be combined with multiple solutions of the audio detection method provided in the above embodiments.
- identifying music events in the audio segment includes: The audio segment is input into a pre-trained music recognition model to obtain a music event recognition result output by the music recognition model, wherein the music recognition model is trained based on audio samples and event tags corresponding to the audio samples.
- the method in this embodiment includes:
- the music recognition model has the ability to identify music events in audio data. For an input audio segment, it can identify whether the audio segment includes a music event.
- the training process of the music recognition model may include: obtaining audio samples and event tags corresponding to the audio samples, where the audio samples may include a variety of different sound events, such as music, laughter, chat, noise and other events, correspondingly , the event tag corresponding to the audio sample can be an event identifier, such as a music identifier, a laughter identifier, a noise identifier, etc. Audio samples including music events are regarded as positive samples, and audio samples including laughter, chat, noise and other events are regarded as negative samples.
- the event labels corresponding to the positive and negative samples can be positive and negative respectively.
- the initial training model is trained based on the audio samples corresponding to the positive and negative samples and the event labels corresponding to the audio samples to obtain the music recognition model.
- the initial training model may include but is not limited to long short-term memory network model, support vector machine model, etc., which are not limited here.
- the audio segments can be input into the pre-trained music recognition model to classify or identify the sound events in the audio segments.
- the music recognition model can quickly output the music event recognition results.
- the pre-trained music recognition model can be used in online applications of audio detection devices without complex calculations, and can quickly obtain music event recognition results, thus improving the speed of audio detection.
- the music recognition model may also output the start and end timestamps of the recognized music events in the audio segment.
- the training samples of the music recognition model also include the start and end timestamps corresponding to the music event tags in the audio samples. The music recognition model trained through the above training samples can identify whether the input audio segment includes music events, and the location of the music events. Start and end timestamps.
- the method further includes: determining whether the duration of the music event in the audio segment is greater than a first preset duration, and if the duration of the music event in the audio segment is not greater than For the first preset duration, the music event is unmarked.
- the duration of the music event may be determined based on the start and end timestamps of the music event.
- the music event recognition result includes a music event
- the audio segment in the detected audio can It can include events such as playing music or singing, but it may also be caused by interfering sounds.
- the interfering sounds can be short text message alerts or mobile phone ringtones.
- the interfering sounds may also include music. This situation indicates that the audio The music events in the segment are not real music events, and the music events need to be unmarked to avoid misjudgment of music events.
- the duration of the music event in the audio segment is judged. If the duration of the music event in the audio segment is greater than the first preset duration, it indicates that the music event meets the music standard, and the music event mark of the audio segment remains unchanged; if the audio If the duration of the music event in the segment is less than or equal to the first preset duration, it indicates that the music event does not meet the music standards, and the music event mark of the audio segment is cancelled.
- the first preset time period may be set based on historical experience. For example, the first preset time period may be 6 seconds.
- the audio detection method inputs audio segments into a pre-trained music recognition model, classifies or identifies sound events in the audio segments, and obtains music event recognition results. Remove music events that are less than or equal to the first preset duration, reduce the interference of misidentified music events, and reduce the increase in statistical workload caused by short-term music events.
- FIG. 3 is a schematic flow chart of another audio detection method provided by an embodiment of the present disclosure.
- the method of this embodiment can be combined with multiple solutions of the audio detection method provided in the above embodiments.
- determining metadata information matching the music event includes: for an audio segment containing a music event, extracting audio fingerprint features of the audio segment; based on the audio fingerprint The features are matched in a fingerprint feature database to determine metadata information matching the music event, where the fingerprint feature database includes music metadata and corresponding fingerprint features.
- the method in this embodiment includes:
- S330 Perform matching in a fingerprint feature database based on the audio fingerprint features to determine metadata information matching the music event, where the fingerprint feature database includes music metadata and corresponding fingerprint features.
- the audio fingerprint feature refers to the digital feature of the music event, which is the music fingerprint feature and is unique. Audio fingerprint features can be extracted from the audio segment through audio fingerprint technology, which includes but is not limited to the Philips algorithm or the Shazam algorithm.
- the fingerprint feature database refers to a database containing music metadata and fingerprint features, which can pre-store multiple music metadata and fingerprint features corresponding to music metadata.
- the fingerprint features corresponding to the music metadata can be used to match the audio fingerprint features. If the match is successful, the metadata matching the music event will be obtained.
- the fingerprint features may include but are not limited to frequency parameters and time parameters corresponding to the frequency spectrum of the music metadata.
- extracting the audio fingerprint features of the audio segment includes: intercepting the audio segment according to the start and end timestamps of the music events in the audio segment to obtain the intercepted audio segment, and extracting the intercepted audio segment. Audio fingerprint characteristics of the audio segment.
- the identification result of the music event includes the start and end timestamps of the music events, and the start timestamp and end timestamp of the music events in the audio segment are obtained; and the corresponding audio is obtained based on the start timestamps and end timestamps of the music events. Intercept the segment and extract the audio data corresponding to the music event in the audio segment.
- the audio data corresponding to the intercepted music event is determined to determine the audio fingerprint characteristics, avoiding the audio data of the non-music event. It interferes with audio fingerprint features and at the same time reduces the amount of audio data required to determine audio fingerprint features, which is beneficial to the rapid extraction of audio fingerprint features.
- extracting the audio fingerprint characteristics of the audio segment includes: extracting audio data of the track where the music event is located in the audio segment, based on the audio data of the track where the music event is located.
- Data extraction audio fingerprint features include: extracting audio data of the track where the music event is located in the audio segment, based on the audio data of the track where the music event is located.
- Data extraction audio fingerprint features include: extracting audio data of the track where the music event is located in the audio segment, based on the audio data of the track where the music event is located.
- Data extraction audio fingerprint features includes: extracting audio data of the track where the music event is located in the audio segment, based on the audio data of the track where the music event is located.
- Data extraction audio fingerprint features includes: extracting audio data of the track where the music event is located in the audio segment, based on the audio data of the track where the music event is located.
- Data extraction audio fingerprint features includes: extracting audio data of the track where the music event is located in the audio segment, based on the audio data of the track
- Different audio tracks can include music events at the same time, or one or more audio tracks can include music events independently.
- the audio detection method provided by the embodiment of the present disclosure extracts the audio fingerprint features of the audio segment from the audio segment containing the music event, and performs matching in the fingerprint feature library based on the extracted audio fingerprint features to determine the metadata that matches the music event.
- Information, metadata information corresponding to music events is obtained through fingerprint feature database matching, and the processing speed is fast, which can save time in audio detection.
- FIG 4 is a schematic flow chart of another audio detection method provided by an embodiment of the present disclosure.
- the method of this embodiment can be combined with multiple solutions of the audio detection method provided in the above embodiments.
- determining the statistical data in the detected audio based on the metadata information includes: according to the start and end timestamps of the music events in each audio segment, corresponding to the same metadata information Music events are merged to obtain statistical data in the detected audio.
- the method in this embodiment includes:
- the start and end timestamps of music events refer to the start timestamp and end timestamp of music events. If the metadata information of the music events is the same, it means that the above-mentioned multiple music events are part of the same piece of music or song. Music events with the same metadata information can be merged, and the statistical data in the detected audio is determined based on the merged music events. , avoid recognition errors caused by dividing the detected audio data into audio segments, and improve the accuracy of statistical data.
- music events corresponding to the same metadata information are merged to obtain statistical data in the detected audio, including: for adjacent music event, if the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is less than the second preset duration, then the adjacent music events will be merged; if the adjacent music events If the metadata information corresponding to the events is different, or the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is greater than or equal to the second preset duration, then the adjacent music events will not be processed. Merge.
- the adjacent music event may be a music event in an adjacent audio segment or an adjacent music event within an audio segment, which is not limited here.
- adjacent music events if the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is less than the second preset duration, it indicates that the adjacent music events belong to the same song and the two music events The intervals between are normal singing or playback pauses, or recognition errors due to audio segment division, then adjacent music events can be merged to calibrate the recognized music events; if the metadata information corresponding to adjacent music events are different, indicating that the adjacent music events do not belong to the same song, the adjacent music events will not be merged to distinguish statistics between different songs; if the metadata information corresponding to the adjacent music events is the same, and the interval between adjacent music events is longer than Or equal to the second preset duration, indicating that adjacent music events belong to the same song but have a long pause time. For example, if the same song is played twice, adjacent music events will not be merged to avoid long playback intervals. The statistics of the same song are entered into the same statistics.
- the audio detection method provided by the embodiment of the present disclosure merges the music events corresponding to the same metadata information according to the start and end timestamps of the music events in each audio segment, so that the audio segment containing the merged music events can be obtained accurately, so as to accurately obtain the Detect music events in audio to improve statistical accuracy.
- the detected audio is the audio in the live video; the method It also includes: determining the viewing data of the live broadcast interval corresponding to each metadata information in the statistical data.
- the live video can be a live video collected in real time or a historical live video. Extract audio from the live video to obtain the detected audio. By identifying and counting music events on audio extracted from live videos, the usage of music metadata in live videos can be obtained.
- the viewing data of the live broadcast interval refers to the viewing statistics of the live broadcast room within the preset time period, which can include but is not limited to the total number of views, the number of independent visits, the average viewing time and other data.
- Statistical data can be used as viewing data matching conditions. According to the matching conditions, the viewing data of the live broadcast interval is matched in the live broadcast database to achieve accurate acquisition of viewing data.
- the live broadcast database can include but is not limited to real-time statistical video viewing data. Through the statistical data of music metadata in live videos and the video viewing data corresponding to the statistical data, it is used to evaluate the role of music metadata in attracting traffic in live videos, or to predict the development trend of music metadata.
- Figure 5 is a schematic flow chart of another audio detection method provided by an embodiment of the present disclosure. Based on the above embodiment, this embodiment provides an example to illustrate the audio detection method in the above embodiment.
- the method in this embodiment includes:
- the audio in the live stream is segmented to obtain multiple audio stream slices (i.e., the above-mentioned audio segments). Multiple audio stream slices can be processed in parallel;
- Recognizing music events on audio stream slices includes: extracting short-term features and long-term features from each audio stream slice, and reducing the dimensionality of the extracted short-term features and long-term features through a dimensionality reduction algorithm to remove short-term features and Redundant information of long-term features to obtain main features.
- the dimensionality of the reduced features is greatly reduced, and the performance will be improved to a certain extent.
- SVM Support Vector Machine
- short-term features include at least one of the following features: Perceptual Linear Predictive Coefficients (PLP), Linear Predictive Cepstrum Coefficients (LPCC), Linear Frequency Cepstral Coefficients (Linear Frequency Cepstral) Coefficients (LFCC), Pitch, Short-time Energy (STE), Sub-Band Energy Distribution (SBED), Brightness (BR) and Bandwidth (BW).
- Long-term characteristics include at least one of the following characteristics: Spectrum Flux (SF), Long-Term Average Spectrum (LTAS), and LPC entropy (LPC entropy).
- the recognition result is a music event, continue to determine whether the duration of the current music event is greater than the first preset duration; if the recognition result is not a music event, then cancel the marking of the music event. If the duration of the current music event is greater than the first preset duration, continue to extract the audio fingerprint features of the music event; if the duration of the current music event If the duration of the event is not greater than the first preset duration, the music event is unmarked.
- the audio fingerprint features are extracted from music events through the audio fingerprint extraction algorithm, and the audio fingerprint features are matched in the fingerprint feature library to obtain metadata information. If the metadata information of adjacent music events is the same, that is, the adjacent music events are the same song, and the interval between adjacent music events is less than the second preset duration, it indicates that the two belong to the same song and there is just a normal singing or playback pause in between. , then merge adjacent music events; if the metadata information is the same and the interval between adjacent music events is not less than the second preset duration, it indicates that although they belong to the same song, the pause time is long and they are not suitable for merging. processing, adjacent music events will not be merged. If the metadata information is not the same, that is, the adjacent music events are not the same song, the adjacent music events will not be merged.
- the method further includes: obtaining statistical data of the merged music events, such as the playback start time, playback end time and other data of the music events. Statistics can be used for music rights billing.
- FIG. 6 is a schematic structural diagram of an audio detection device provided by an embodiment of the present disclosure. As shown in Figure 6, the device includes:
- the music event identification module 610 is configured to obtain the audio segment in the detected audio and identify the music event in the audio segment; the statistical data determination module 620 is configured to determine the metadata information matching the music event, based on the The metadata information determines statistics in the detected audio.
- the music event identification module 610 may also be configured to:
- the audio segment is input into a pre-trained music recognition model to obtain a music event recognition result output by the music recognition model, wherein the music recognition model is trained based on audio samples and event tags corresponding to the audio samples.
- the device may also be configured to:
- the statistical data determination module 620 may also include:
- the fingerprint feature extraction unit is configured to extract the audio fingerprint features of the audio segment for the audio segment containing the music event; the metadata matching unit is configured to perform matching in the fingerprint feature library based on the audio fingerprint features and determine the Metadata information matching the music event, wherein the fingerprint feature database includes music metadata and corresponding fingerprint features.
- the fingerprint feature extraction unit may also be configured to:
- the audio segment is intercepted to obtain the interception Audio segment, extract the audio fingerprint feature of the intercepted audio segment; or, extract the audio data of the audio track where the music event is located in the audio segment, and extract the audio fingerprint feature based on the audio data of the audio track where the music event is located.
- the statistical data determination module 620 may also include:
- the data merging unit is configured to merge music events corresponding to the same metadata information according to the start and end timestamps of the music events in each audio segment to obtain statistical data in the detected audio.
- the data merging unit may also be configured to:
- the adjacent music events For adjacent music events, if the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is less than the second preset duration, the adjacent music events will be merged; if If the metadata information corresponding to the adjacent music events is different, or the metadata information corresponding to the adjacent music events is the same, and the interval duration of the adjacent music events is greater than or equal to the second preset duration, then the metadata information corresponding to the adjacent music events is not the same. Adjacent music events are merged.
- the detected audio is the audio in the live video; the device may also be configured to: determine the viewing data of the live broadcast interval corresponding to each metadata information in the statistical data .
- the audio detection device provided by the embodiment of the present disclosure can execute the audio detection method provided by any embodiment of the present disclosure, and has corresponding functional modules and effects for executing the audio detection method.
- the multiple units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned divisions, as long as they can achieve the corresponding functions; in addition, the names of the multiple functional units are only for the convenience of distinguishing each other. , are not used to limit the protection scope of the embodiments of the present disclosure.
- Terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (Personal Digital Assistant, PDA), tablet computers (Portable Android Device, PAD), portable multimedia players Mobile terminals such as (Portable Media Player, PMP), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital television (TV), desktop computers, etc.
- PDA Personal Digital Assistant
- PAD Portable Multimedia Players Mobile terminals
- PMP Portable Multimedia Player
- vehicle-mounted terminals such as vehicle-mounted navigation terminals
- fixed terminals such as digital television (TV), desktop computers, etc.
- TV digital television
- the electronic device 400 shown in FIG. 7 is only an example and should not bring any limitations to the functions and usage scope of the embodiments of the present disclosure.
- the electronic device 400 may include a processing device (such as a central processing unit, a graphics processor, etc.) 401, which may process data according to a program stored in a read-only memory (Read-Only Memory, ROM) 402 or from a storage device. 408 loads the program in the random access memory (Random Access Memory, RAM) 403 to perform various appropriate actions and processes. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored.
- Processing device 401, ROM 402 and RAM 403 They are connected to each other via bus 404.
- An input/output (I/O) interface 405 is also connected to bus 404.
- the following devices can be connected to the I/O interface 405: input devices 406 including, for example, a touch screen, touch pad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; including, for example, a Liquid Crystal Display (LCD) , an output device 407 such as a speaker, a vibrator, etc.; a storage device 408 including a magnetic tape, a hard disk, etc.; and a communication device 409.
- the communication device 409 may allow the electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data.
- FIG. 7 illustrates electronic device 400 with various means, implementation or availability of all illustrated means is not required. More or fewer means may alternatively be implemented or provided.
- embodiments of the present disclosure include a computer program product including a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the method illustrated in the flowchart.
- the computer program may be downloaded and installed from the network via communication device 409, or from storage device 408, or from ROM 402.
- the processing device 401 When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
- the electronic device provided by the embodiment of the present disclosure belongs to the same concept as the audio detection method provided by the above embodiment.
- Technical details that are not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same effect as the above embodiment. .
- Embodiments of the present disclosure provide a computer storage medium on which a computer program is stored.
- the program is executed by a processor, the audio detection method provided by the above embodiments is implemented.
- the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the above two.
- the computer-readable storage medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any combination thereof.
- Examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard drives, RAM, ROM, Erasable Programmable Read-Only Memory (EPROM) or flash memory), optical fiber, portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.
- a computer-readable storage medium may be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device.
- a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code therein. Such propagated data signals may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.
- the computer-readable signal medium may also be any computer-readable medium other than computer-readable storage media that can transmit, Propagate or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
- Program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (Radio Frequency, RF), etc., or any suitable combination of the above.
- the client and server can communicate using any currently known or future developed network protocol, such as HyperText Transfer Protocol (HTTP), and can communicate with digital data in any form or medium.
- HTTP HyperText Transfer Protocol
- Communications e.g., communications network
- Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any current network for knowledge or future research and development.
- LANs Local Area Networks
- WANs Wide Area Networks
- the Internet e.g., the Internet
- end-to-end networks e.g., ad hoc end-to-end networks
- the above-mentioned computer-readable medium may be included in the above-mentioned electronic device; it may also exist independently without being assembled into the electronic device.
- the above-mentioned computer-readable medium carries one or more programs.
- the electronic device executes the above-mentioned one or more programs.
- Obtain audio segments in the detected audio identify music events in the audio segments; determine metadata information matching the music events, and determine statistical data in the detected audio based on the metadata information.
- Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, including but not limited to object-oriented programming languages—such as Java, Smalltalk, C++, and Includes conventional procedural programming languages—such as "C" or similar programming languages.
- the program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.
- the remote computer may be connected to the user computer through any kind of network, including a LAN or WAN, or may be connected to an external computer (eg, through the Internet using an Internet service provider).
- each block in the flowchart or block diagram may represent a module, segment, or portion of code that contains one or more logic functions that implement the specified executable instructions.
- the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown one after another may actually execute substantially in parallel, or they may sometimes execute in the reverse order, depending on the functionality involved.
- each block in the block diagram and/or flowchart illustration, and combinations of blocks in the block diagram and/or flowchart illustration may It can be implemented with a dedicated hardware-based system that performs the specified function or operation, or it can be implemented with a combination of dedicated hardware and computer instructions.
- the units involved in the embodiments of the present disclosure can be implemented in software or hardware. Among them, the name of the unit/module does not constitute a limitation on the unit itself.
- exemplary types of hardware logic components include: field programmable gate array (Field Programmable Gate Array, FPGA), application specific integrated circuit (Application Specific Integrated Circuit, ASIC), application specific standard product (Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programming Logic Device (CPLD), etc.
- a machine-readable medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
- the machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- Machine-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any suitable combination of the foregoing. Examples of machine-readable storage media would include an electrical connection based on one or more wires, a portable computer disk, a hard drive, RAM, ROM, EPROM or flash memory, optical fiber, CD-ROM, optical storage device, magnetic storage device, or Any suitable combination of the above.
- Example 1 provides an audio detection method, which includes:
- Metadata information matching the music event is determined, and statistics in the detected audio are determined based on the metadata information.
- Example 2 provides an audio detection method, further including:
- the identifying music events in the audio segment includes:
- the audio segment is input into a pre-trained music recognition model to obtain a music event recognition result output by the music recognition model, wherein the music recognition model is trained based on audio samples and event tags corresponding to the audio samples.
- Example 3 provides an audio detection method, Also includes:
- the method further includes:
- Example 4 provides an audio detection method, further including:
- Determining metadata information matching the music event includes:
- Matching is performed in a fingerprint feature database based on the audio fingerprint features to determine metadata information matching the music event, where the fingerprint feature database includes music metadata and corresponding fingerprint features.
- Example 5 provides an audio detection method, further including:
- the extraction of audio fingerprint features of the audio segment includes:
- Audio data of the audio track where the music event is located is extracted from the audio segment, and audio fingerprint features are extracted based on the audio data of the audio track where the music event is located.
- Example 6 provides an audio detection method, further including:
- Determining statistical data in the detected audio based on the metadata information includes:
- music events corresponding to the same metadata information are merged to obtain statistical data in the detected audio.
- Example 7 provides an audio detection method, further including:
- music events corresponding to the same metadata information are merged to obtain statistical data in the detected audio, including:
- the adjacent music events will not be merged.
- Example 8 provides an audio detection method, further including:
- the detected audio is the audio in the live video
- the method also includes:
- the viewing data of the live broadcast interval corresponding to each metadata information in the statistical data is determined.
- Example 9 provides an audio detection device, which includes:
- a music event identification module configured to obtain audio segments in the detected audio and identify music events in the audio segments
- a statistical data determination module is configured to determine metadata information matching the music event, and determine statistical data in the detected audio based on the metadata information.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Software Systems (AREA)
- Human Computer Interaction (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Health & Medical Sciences (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- General Physics & Mathematics (AREA)
- Mathematical Physics (AREA)
- General Engineering & Computer Science (AREA)
- Computing Systems (AREA)
- Medical Informatics (AREA)
- Evolutionary Computation (AREA)
- Data Mining & Analysis (AREA)
- Artificial Intelligence (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Description
Claims (12)
- 一种音频检测方法,包括:获取被检测音频中的音频段,识别所述音频段中的音乐事件;确定与所述音乐事件相匹配的元数据信息,基于所述元数据信息确定所述被检测音频中的统计数据。
- 根据权利要求1所述的方法,其中,所述识别所述音频段中的音乐事件,包括:将所述音频段输入至预先训练的音乐识别模型中,得到所述音乐识别模型输出的音乐事件识别结果,其中,所述音乐识别模型基于音频样本与所述音频样本对应的事件标签训练得到。
- 根据权利要求1所述的方法,在所述识别所述音频段中的音乐事件之后,还包括:确定所述音频段中的音乐事件的时长是否大于第一预设时长,响应于所述音频段中的音乐事件的时长不大于所述第一预设时长,取消标记所述音乐事件。
- 根据权利要求1所述的方法,其中,所述确定与所述音乐事件相匹配的元数据信息,包括:对于包含所述音乐事件的音频段,提取所述音频段的音频指纹特征;基于所述音频指纹特征在指纹特征库中进行匹配,确定与所述音乐事件相匹配的元数据信息,其中,所述指纹特征库中包括音乐元数据和对应的指纹特征。
- 根据权利要求4所述的方法,其中,所述提取所述音频段的音频指纹特征,包括:根据所述音频段中所述音乐事件的起止时间戳,对所述音频段进行截取,得到截取音频段,提取所述截取音频段的音频指纹特征;或者,在所述音频段中提取所述音乐事件所在音轨的音频数据,基于所述音乐事件所在音轨的音频数据提取音频指纹特征。
- 根据权利要求1所述的方法,其中,所述基于所述元数据信息确定所述被检测音频中的统计数据,包括:根据每个音频段中音乐事件的起止时间戳,对相同元数据信息对应的音乐事件进行合并,得到所述被检测音频中的统计数据。
- 根据权利要求6所述的方法,其中,所述根据每个音频段中音乐事件的起止时间戳,对相同元数据信息对应的音乐事件进行合并,得到所述被检测音 频中的统计数据,包括:对于相邻音乐事件,在所述相邻音乐事件对应的元数据信息相同,且所述相邻音乐事件的间隔时长小于第二预设时长的情况下,将所述相邻音乐事件进行合并;在所述相邻音乐事件对应的元数据信息不同,或者,所述相邻音乐事件对应的元数据信息相同,且所述相邻音乐事件的间隔时长大于或等于第二预设时长的情况下,不对所述相邻音乐事件进行合并。
- 根据权利要求1所述的方法,其中,所述被检测音频为直播视频中的音频;所述方法还包括:确定所述统计数据中每个元数据信息所对应的直播区间的观看数据。
- 一种音频检测装置,包括:音乐事件识别模块,设置为获取被检测音频中的音频段,识别所述音频段中的音乐事件;统计数据确定模块,设置为确定与所述音乐事件相匹配的元数据信息,基于所述元数据信息确定所述被检测音频中的统计数据。
- 一种电子设备,包括:至少一个处理器;存储装置,设置为存储至少一个程序;当所述至少一个程序被所述至少一个处理器执行,使得所述至少一个处理器实现如权利要求1-8中任一所述的音频检测方法。
- 一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由计算机处理器执行时用于执行如权利要求1-8中任一所述的音频检测方法。
- 一种计算机程序产品,包括承载在非暂态计算机可读介质上的计算机程序,所述计算机程序包含用于执行如权利要求1-8中任一所述的音频检测方法的程序代码。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/844,887 US20250218418A1 (en) | 2022-03-08 | 2023-02-28 | Audio detection method and apparatus, storage medium and electronic device |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202210220184.7A CN114596878B (zh) | 2022-03-08 | 2022-03-08 | 一种音频检测方法、装置、存储介质及电子设备 |
| CN202210220184.7 | 2022-03-08 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023169258A1 true WO2023169258A1 (zh) | 2023-09-14 |
Family
ID=81807399
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/078752 Ceased WO2023169258A1 (zh) | 2022-03-08 | 2023-02-28 | 音频检测方法、装置、存储介质及电子设备 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20250218418A1 (zh) |
| CN (1) | CN114596878B (zh) |
| WO (1) | WO2023169258A1 (zh) |
Families Citing this family (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114596878B (zh) * | 2022-03-08 | 2026-01-02 | 北京字跳网络技术有限公司 | 一种音频检测方法、装置、存储介质及电子设备 |
| CN115472148B (zh) * | 2022-08-31 | 2025-06-27 | 海尔优家智能科技(北京)有限公司 | 测试结果的确定方法、装置、存储介质及电子装置 |
| CN117746894A (zh) * | 2022-09-13 | 2024-03-22 | 广州视源电子科技股份有限公司 | 朗读事件识别方法、装置、教学设备和存储介质 |
| CN115866279B (zh) * | 2022-09-20 | 2024-11-29 | 北京奇艺世纪科技有限公司 | 直播视频处理方法、装置、电子设备及可读存储介质 |
| CN116072147A (zh) * | 2023-01-09 | 2023-05-05 | 北京达佳互联信息技术有限公司 | 音乐检测模型训练方法、装置、电子设备及存储介质 |
| EP4728513A1 (en) * | 2023-06-14 | 2026-04-22 | Qualcomm Incorporated | Knowledge-based audio scene graph |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2010078984A (ja) * | 2008-09-26 | 2010-04-08 | Sanyo Electric Co Ltd | 楽曲抽出装置および楽曲記録装置 |
| CN105874732A (zh) * | 2014-01-07 | 2016-08-17 | 高通股份有限公司 | 用于识别音频流中的一首音乐的方法和装置 |
| WO2020176057A1 (en) * | 2019-02-25 | 2020-09-03 | Ahmet Aksoy | Music analysis system and method for public spaces |
| CN113032616A (zh) * | 2021-03-19 | 2021-06-25 | 腾讯音乐娱乐科技(深圳)有限公司 | 音频推荐的方法、装置、计算机设备和存储介质 |
| CN113987258A (zh) * | 2021-11-10 | 2022-01-28 | 北京有竹居网络技术有限公司 | 音频的识别方法、装置、可读介质和电子设备 |
| CN114596878A (zh) * | 2022-03-08 | 2022-06-07 | 北京字跳网络技术有限公司 | 一种音频检测方法、装置、存储介质及电子设备 |
Family Cites Families (15)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7336890B2 (en) * | 2003-02-19 | 2008-02-26 | Microsoft Corporation | Automatic detection and segmentation of music videos in an audio/video stream |
| CN102956230B (zh) * | 2011-08-19 | 2017-03-01 | 杜比实验室特许公司 | 对音频信号进行歌曲检测的方法和设备 |
| GB2507551A (en) * | 2012-11-04 | 2014-05-07 | Julian Andrew John Fells | Copyright protection by comparing identifiers of first and second electronic content |
| CN103440330A (zh) * | 2013-09-03 | 2013-12-11 | 网易(杭州)网络有限公司 | 一种音乐节目信息获取方法和设备 |
| CN104091596B (zh) * | 2014-01-20 | 2016-05-04 | 腾讯科技(深圳)有限公司 | 一种乐曲识别方法、系统和装置 |
| US9620105B2 (en) * | 2014-05-15 | 2017-04-11 | Apple Inc. | Analyzing audio input for efficient speech and music recognition |
| US10089987B2 (en) * | 2015-12-21 | 2018-10-02 | Invensense, Inc. | Music detection and identification |
| CN105657535B (zh) * | 2015-12-29 | 2018-10-30 | 北京搜狗科技发展有限公司 | 一种音频识别方法和装置 |
| US10296638B1 (en) * | 2017-08-31 | 2019-05-21 | Snap Inc. | Generating a probability of music using machine learning technology |
| JP7143327B2 (ja) * | 2017-10-03 | 2022-09-28 | グーグル エルエルシー | コンピューティング装置によって実施される方法、コンピュータシステム、コンピューティングシステム、およびプログラム |
| CN111723235B (zh) * | 2019-03-19 | 2023-09-26 | 百度在线网络技术(北京)有限公司 | 音乐内容识别方法、装置及设备 |
| US10796684B1 (en) * | 2019-04-30 | 2020-10-06 | Dialpad, Inc. | Chroma detection among music, speech, and noise |
| CN113724736A (zh) * | 2021-08-06 | 2021-11-30 | 杭州网易智企科技有限公司 | 一种音频处理方法、装置、介质和电子设备 |
| CN114023289B (zh) * | 2021-11-09 | 2025-10-10 | 北京百度网讯科技有限公司 | 音乐识别方法、音乐特征提取模型的训练方法及装置 |
| GB2612994A (en) * | 2021-11-18 | 2023-05-24 | Audoo Ltd | Media identification system |
-
2022
- 2022-03-08 CN CN202210220184.7A patent/CN114596878B/zh active Active
-
2023
- 2023-02-28 US US18/844,887 patent/US20250218418A1/en active Pending
- 2023-02-28 WO PCT/CN2023/078752 patent/WO2023169258A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2010078984A (ja) * | 2008-09-26 | 2010-04-08 | Sanyo Electric Co Ltd | 楽曲抽出装置および楽曲記録装置 |
| CN105874732A (zh) * | 2014-01-07 | 2016-08-17 | 高通股份有限公司 | 用于识别音频流中的一首音乐的方法和装置 |
| WO2020176057A1 (en) * | 2019-02-25 | 2020-09-03 | Ahmet Aksoy | Music analysis system and method for public spaces |
| CN113032616A (zh) * | 2021-03-19 | 2021-06-25 | 腾讯音乐娱乐科技(深圳)有限公司 | 音频推荐的方法、装置、计算机设备和存储介质 |
| CN113987258A (zh) * | 2021-11-10 | 2022-01-28 | 北京有竹居网络技术有限公司 | 音频的识别方法、装置、可读介质和电子设备 |
| CN114596878A (zh) * | 2022-03-08 | 2022-06-07 | 北京字跳网络技术有限公司 | 一种音频检测方法、装置、存储介质及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20250218418A1 (en) | 2025-07-03 |
| CN114596878B (zh) | 2026-01-02 |
| CN114596878A (zh) | 2022-06-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2023169258A1 (zh) | 音频检测方法、装置、存储介质及电子设备 | |
| CN110503961B (zh) | 音频识别方法、装置、存储介质及电子设备 | |
| EP3945435A1 (en) | Dynamic identification of unknown media | |
| CN113596579B (zh) | 视频生成方法、装置、介质及电子设备 | |
| CN108877779B (zh) | 用于检测语音尾点的方法和装置 | |
| US20160196812A1 (en) | Music information retrieval | |
| US9224385B1 (en) | Unified recognition of speech and music | |
| JP7567028B2 (ja) | ターゲットビデオを生成するための方法、装置、サーバ及び媒体 | |
| EP3092734A1 (en) | Method and device for identifying a piece of music in an audio stream | |
| EP3468205A1 (en) | Temporal fraction with use of content identification | |
| CN111859008B (zh) | 一种推荐音乐的方法及终端 | |
| CN109949798A (zh) | 基于音频的广告检测方法以及装置 | |
| CN107679196A (zh) | 一种多媒体识别方法、电子设备及存储介质 | |
| CN112071287A (zh) | 用于生成歌谱的方法、装置、电子设备和计算机可读介质 | |
| WO2022160603A1 (zh) | 歌曲的推荐方法、装置、电子设备及存储介质 | |
| CN114999454A (zh) | 语音交互设备的性能测试方法、装置、设备及可读介质 | |
| US20240404548A1 (en) | Method, apparatus, device and storage medium for video recording | |
| US11609948B2 (en) | Music streaming, playlist creation and streaming architecture | |
| CN116072147A (zh) | 音乐检测模型训练方法、装置、电子设备及存储介质 | |
| CN111798853A (zh) | 语音识别的方法、装置、设备和计算机可读介质 | |
| WO2026040959A1 (zh) | 音频副本检测方法及装置、电子设备、计算机存储介质以及计算机程序产品 | |
| CN119274119B (zh) | 一种视频内容风险检测方法、装置、介质及设备 | |
| CN114595361B (zh) | 一种音乐热度的预测方法、装置、存储介质及电子设备 | |
| CN114329042A (zh) | 数据处理方法、装置、设备、存储介质及计算机程序产品 | |
| CN115565508B (zh) | 歌曲匹配方法、装置、电子设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23765836 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18844887 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 05.12.2024) |
|
| WWP | Wipo information: published in national office |
Ref document number: 18844887 Country of ref document: US |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23765836 Country of ref document: EP Kind code of ref document: A1 |