WO2020073743A1 - 一种音频检测方法、装置、设备及存储介质 - Google Patents

一种音频检测方法、装置、设备及存储介质 Download PDF

Info

Publication number
WO2020073743A1
WO2020073743A1 PCT/CN2019/102172 CN2019102172W WO2020073743A1 WO 2020073743 A1 WO2020073743 A1 WO 2020073743A1 CN 2019102172 W CN2019102172 W CN 2019102172W WO 2020073743 A1 WO2020073743 A1 WO 2020073743A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
voice
audio
audio file
detection
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/102172
Other languages
English (en)
French (fr)
Inventor
李振
黄震川
邹昱
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Guangzhou Baiguoyuan Information Technology Co Ltd
Original Assignee
Guangzhou Baiguoyuan Information Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Guangzhou Baiguoyuan Information Technology Co Ltd filed Critical Guangzhou Baiguoyuan Information Technology Co Ltd
Priority to US17/282,732 priority Critical patent/US11948595B2/en
Priority to SG11202103561TA priority patent/SG11202103561TA/en
Publication of WO2020073743A1 publication Critical patent/WO2020073743A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/16Speech classification or search using artificial neural networks
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/06Decision making techniques; Pattern matching strategies
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/18Artificial neural networks; Connectionist approaches
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/26Recognition of special voice characteristics, e.g. for use in lie detectors; Recognition of animal voices
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/18Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/27Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
    • G10L25/30Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/78Detection of presence or absence of voice signals

Definitions

  • This application relates to the field of computer network technology, such as an audio detection method, device, equipment, and storage medium.
  • PCs personal computers
  • mobile phones mobile phones
  • tablet computers are becoming more and more popular, bringing great convenience to people's life, learning, and work.
  • users can use the device to communicate with other users through the network.
  • Users can send the voice information they need to send to the network through the device, so that other users can receive and play the voice information through the network to achieve the purpose of voice communication.
  • the voice information involves a wide range of content, which may contain unpleasant voice data, such as harsh, high decibel, inappropriate content, etc. These voices are usually caused by individual users
  • the malicious sending aims to interfere with the normal use of other users, so the software operator classifies this kind of voice as illegal voice data.
  • Embodiments of the present application provide an audio detection method, system, device, and storage medium, which automatically detect voice violations to avoid the detection time lag in the conventional method for manually detecting violation voice data.
  • an embodiment of the present application provides an audio detection method, including: acquiring audio file data; determining attribute detection data corresponding to the audio file data; and performing voice violations on the attribute detection data through a pre-trained fully connected network model Detection, generating the voice behavior detection result corresponding to the audio file data.
  • an embodiment of the present application further provides a device, including: a processor and a memory; at least one instruction is stored in the memory, and the instruction is executed by the processor, so that the device executes as the first Aspect audio detection method.
  • an embodiment of the present application further provides a computer-readable storage medium, when instructions in the storage medium are executed by a processor of a device, so that the device can execute the audio detection method according to the first aspect .
  • FIG. 1 is a schematic flowchart of steps in an audio detection method in an embodiment of the present application.
  • FIG. 2 is a schematic diagram of an audio file data detection process in an embodiment of the present application.
  • FIG. 3 is a schematic structural block diagram of an embodiment of an audio detection device in an embodiment of the present application.
  • FIG. 4 is a schematic structural block diagram of a device in an embodiment of the present application.
  • FIG. 1 a schematic flowchart of steps of an audio detection method in an embodiment of the present application is shown, which may include steps 110 to 130.
  • step 110 audio file data is acquired.
  • the embodiment of the present application may obtain the audio file data that needs to be detected currently, to detect whether the currently acquired audio file data includes the violation voice data corresponding to the voice violation.
  • the audio file data can be used to characterize the data contained in the audio file, such as voice data in a sound file.
  • the sound file can be generated by the device according to the user's speech sound, and can be used to carry the voice data that characterizes the user needs to be sent. Data (ie, compliant voice data), etc., the embodiments of the present application do not limit this.
  • the user's speech sound forms voice data
  • the voice data can be input into the software through the software interface, so that the software can determine the voice data as audio file data, that is, the software can obtain the audio file data, so that The software can automatically perform audio detection on the audio file data; alternatively, the software can transmit the acquired audio file data to the software platform through the network, so that the software platform can obtain the audio file data, and then the audio can be obtained through the software platform File data for audio detection.
  • step 120 the attribute detection data corresponding to the audio file data is determined.
  • an embodiment of the present application may perform data processing on the audio file data to determine attribute detection data corresponding to the audio file data, so that subsequent voice violation detection can be performed on the attribute detection data.
  • the attribute detection data may include at least one of the following: user level data, classification probability data, and voiceprint feature data, etc., which are not limited in the embodiments of the present application.
  • user level data, classification probability data, and voiceprint feature data can be obtained in advance and stored in a predetermined location for reading or comparison or determination from a predetermined location when step 120 is performed.
  • the user level data can be used to determine the target user's user level, and the user level can be predetermined according to the target user's consumption habits and historical login behavior, such as by normalizing the target user's historical recharge records and historical login behavior data To determine the user level of the target user.
  • the historical login behavior data can be used to characterize the user's historical login behavior.
  • the historical login duration data as the historical login behavior data can be used to characterize the user's historical login duration;
  • the target user can refer to the user who sent the audio file data, such as It can be a user who sends voice messages during voice communication. The higher the target user's user level, the less likely it is that the target user can send illegal voice data.
  • the classification probability data can be used to characterize the classification probability corresponding to the voice violation behavior, such as the voice violation probability and the voice compliance probability.
  • the voice violation probability may refer to: the probability that the voice data contained in the audio file data is violation voice data;
  • the voice compliance probability may refer to: the probability that the voice data contained in the audio file data is compliance voice data.
  • the size of the voice violation probability may refer to the following factors: harshness, high decibels, audio content information, and so on. For example, if the audio file data contains malicious, harsh, high-decibel, inappropriate content, and other uncomfortable voice data, the probability of voice violations corresponding to the audio file data will increase, that is, the audio will contain a malicious, harsh ear from the user. , High decibels, inappropriate content and other uncomfortable information will lead to an increased probability of violation of the audio.
  • the voiceprint feature data can be used to characterize the voiceprint feature corresponding to the audio file data.
  • the voiceprint feature may refer to the user's voice texture feature.
  • the sound texture feature may refer to the frequency domain feature of the sound generated by subjecting the user's original audio waveform to Fourier transform and other post-processing in the time domain.
  • the attribute detection data determined in the above step 120 may include at least two of the following: user level data, classification probability data, and voiceprint feature data.
  • the user level is used to characterize the user level
  • the classification probability data is used to characterize the classification probability corresponding to the voice violation
  • the voiceprint feature data is used to characterize the voiceprint feature corresponding to the audio file data.
  • the attribute detection data simultaneously includes user level data, classification probability data, and voiceprint feature data, so that the voice behavior detection result is more accurate.
  • step 130 the pre-trained fully-connected network model is used to perform voice violation behavior detection on the attribute detection data to generate a voice behavior detection result corresponding to the audio file data.
  • the attribute detection data after determining the attribute detection data corresponding to the audio file data, the attribute detection data can be used as the input of the pre-trained fully connected network model, and then the attribute detection data can be input to the pre-trained fully connected network In the model, a voice violation behavior is generated through the fully connected network model to generate a voice behavior test result corresponding to audio file data.
  • the voice behavior detection result is a voice violation detection result, it can be determined that the audio file data contains violation voice data, that is, the violation voice data corresponding to the voice violation is detected.
  • the voice behavior detection result is not the voice violation detection result, it can be determined that the audio file data does not contain the violation voice data, and if the voice behavior detection result is the voice normal behavior detection result, it can be determined that the audio file data does not contain the voice violation The illegal voice data corresponding to the behavior.
  • the embodiment of the present application can perform voice violation detection by determining the attribute detection data corresponding to the audio file data, so that the violation voice data corresponding to the voice violation can be detected in time to ensure the user
  • the normal use of the system avoids the time lag caused by the detection of voice violations based on user reports and manual spot checks, guarantees the user's normal experience, meets user needs, and has a low investment cost.
  • the transmission of the user ’s voice data is restricted or suspended to other users.
  • the user ’s violation is determined based on the user ’s audio file data.
  • the user's client sends corresponding information to enable the client to suspend the user's voice chat function.
  • the audio file data after acquiring audio file data, may be sliced, so that at least two frames of audio time-domain information after slice processing may be obtained; subsequently, at least two frames of audio may be obtained
  • the time domain information is subjected to feature extraction to obtain audio feature data, so as to determine attribute detection data corresponding to audio file data for the audio feature data.
  • the audio feature data can be used to characterize the audio features, such as the amplitude features and voiceprint features of the audio.
  • the classification probability data corresponding to the audio file data can be generated according to the amplitude spectrum characteristic data, so that the classification probability data can be subsequently subjected to voice violations Detection.
  • Amplitude spectrum (Magnitude Spectrum, Mags) feature data can be used to characterize the Mags feature of audio. It should be noted that in the process of implementing the embodiments of the present application, the applicant found that Mags features are effective in detecting sound violations through in-depth analysis and experiments. Therefore, as described later, some solutions of the embodiments of the present application combine Mags Features have been expanded accordingly.
  • the step of determining the attribute detection data corresponding to the audio file data may include: slicing the audio file data to obtain at least two frames of audio time domain information; Perform feature extraction on two frames of audio time-domain information to obtain amplitude spectrum feature data and voiceprint feature data; splice the amplitude spectrum feature data and voiceprint feature data to generate feature vector data; through a pre-trained speech classification model , Perform speech classification processing on the feature vector data to obtain classification probability data as the attribute detection data.
  • the feature extraction may include: amplitude spectrum feature extraction, voiceprint feature extraction, etc., which is not limited in the embodiments of the present application.
  • the audio file data to be sent can be obtained, and the acquired audio file data can be sliced using a preset mobile window to obtain multi-frame audio time domain information. Therefore, Mags feature extraction can be performed on at least two frames of audio time domain information to obtain Mags feature data; and voiceprint feature extraction can be performed on at least two frames of audio time domain information to obtain voiceprint features corresponding to the audio file data Data; Then, you can splice the obtained Mags feature data with the voiceprint feature data to generate one-dimensional feature vector data, and use the feature vector data as the input of the speech classification model, and perform speech classification processing through the speech classification model To obtain the classification probability data corresponding to the audio file data.
  • the speech classification model can extract audio features according to different audio inputs, that is, perform feature extraction based on the input feature vector data to obtain input features; then, according to the distribution of input features, the audio input can be assigned a probability value , That is, assign a classification probability data to the input feature vector data and output it as attribute detection data for voice violation detection.
  • the speech classification model will assign a high probability of violation to the input feature vector data, for example, output 90% as the probability of speech violation of feature vector data.
  • the speech classification model will assign a low probability of violation to the input feature vector data, for example, output 1% as the probability of speech violation of feature vector data, that is, the feature vector data
  • the probability of voice compliance is 99%.
  • performing feature extraction on the at least two frames of audio time domain information to obtain amplitude spectrum characteristic data may include: performing frequency domain transformation on the at least two frames of audio time domain information to obtain audio frequency domain information; Perform amplitude spectrum feature extraction based on the audio frequency domain information to obtain amplitude spectrum feature data.
  • the frequency domain transform may include Fourier transform, such as Fast Fourier Transform (Fast Fourier Transformation, FFT), etc., which is not limited in the embodiments of the present application.
  • each small segment can be called a frame of audio time domain information; subsequently, each frame of audio time can be obtained Perform Fourier transform on the domain information to obtain audio frequency domain information corresponding to each frame of audio time domain information, and extract amplitude spectrum features based on the audio frequency domain information to obtain amplitude spectrum feature data corresponding to the audio file data.
  • Mags of audio frequency domain information take the mean and variance, and then use the obtained mean and variance as amplitude spectrum feature data to generate feature vector data based on the amplitude spectrum feature data, so that the feature vector data can be subjected to speech classification processing to obtain all The classification probability data corresponding to the audio file data.
  • the adjacent two frames of audio time domain information may have overlapping parts, that is, there may be overlap between frames.
  • the frame length of one frame of audio time domain information is 25 milliseconds (ms)
  • the frame shift is 10 At milliseconds, there can be 15 milliseconds of overlap between the two frames. It should be noted that the frame length and frame shift can be set according to the accuracy requirements, which is not limited in this example.
  • the obtained audio feature data can also use the obtained audio feature data as attribute detection data to detect voice violations of the audio feature data.
  • the audio feature data is voiceprint feature data
  • the obtained fixed-length data can be determined as the first fixed-length data to Input the first fixed-length data into a pre-trained neural network model for voiceprint feature extraction to obtain voiceprint feature data, and then use the voiceprint feature data as attribute detection data corresponding to audio file data and input it to the full connection
  • voice violation behavior detection is performed on voiceprint feature data to generate a voice behavior detection result corresponding to the audio file data.
  • the audio file data to be detected may include audio file data to be transmitted, audio file data to be played, etc .; the audio file data to be transmitted may be used to characterize the audio file data to be sent during the voice communication process; the audio to be played The file data can be used to characterize the file data of the voice to be played.
  • users can be rated based on user consumption habits, and voice violation detection can be performed based on the user level obtained after the classification to predict voice violations.
  • the audio detection method provided in this embodiment of the present application may further include: acquiring historical behavior data of the target user; and obtaining user level data as the attribute detection data according to the historical behavior data.
  • the historical behavior data includes at least one of the following: historical login data, user consumption behavior data, violation historical data, user consumption behavior data, and recharge historical data.
  • the user when the user level of a certain user needs to be determined, the user can be determined as the target user, and the target user's Historical behavior data, to determine the user level of the target user according to the acquired historical behavior data, that is, determine the user level data of the target user. Subsequently, the user level data may be stored in a database to obtain the user level data as attribute detection data from the database during audio detection.
  • the embodiments of the present application may also acquire historical behavior data corresponding to the target user of the audio file data corresponding to the acquired audio file data during audio detection, so as to determine in real time the attribute detection data based on the acquired historical behavior data User level data, so that the user level data determined in real time can be used for voice violation detection, and the accuracy of voice violation detection can be improved.
  • the above step of determining attribute detection data corresponding to the audio file data includes: acquiring historical behavior data of the target user for the audio file data; normalizing the historical behavior data to obtain User level data of attribute detection data.
  • the user who sent the audio file data may be determined as the target user, and then for the audio file data, the history of the target user may be obtained according to the user ID of the target user Behavior data, such as user consumption behavior data, historical login data, violation historical data and recharge historical data, etc.
  • the user's consumption behavior data can be used to determine the target user's consumption behavior habits information
  • the violation history data can be used to determine the target user's voice violation history information, such as determining whether the target user has a violation history, or determining the target user's violation history times
  • Recharge history data can be used to determine the target user's recharge history information, such as the target user's recharge times, historical recharge amount, etc .
  • user historical login data can be used to determine the target user's historical login behavior, including: number of logins, login duration , Login address, etc.
  • the number of logins can be used to characterize the number of logins of the target user; the login duration can be used to characterize the historical login duration of the target user, for example, it can include the login duration corresponding to each login of the target user; the login address can be used to determine the address of each user login For example, it can be the Internet Protocol (IP) address and Medium Access Control (MAC) address of the device used by the target user when logging in. This example does not limit this.
  • IP Internet Protocol
  • MAC Medium Access Control
  • the acquired historical behavior data can be normalized, such as the number of logins of the target user, the length of login, whether there is a history of violations, recharge history and other information. Normalization and normalization, so that the user level data of the target user can be determined based on the normalization processing result, and then the user level data can be used as attribute detection data and input into a pre-trained fully connected network model for voice violation detection To generate voice behavior detection results.
  • sound texture extraction may be performed based on, for example, Convolutional Neural Network (CNN) to obtain a voiceprint corresponding to the audio file data Feature data; and can perform sound classification based on the Mags feature, that is, use the Mags feature data corresponding to the audio file data to generate a feature vector number, and input the feature vector number into the speech classification model for speech classification processing to obtain the audio file data corresponding to Probability data of the classification; and for the currently acquired audio file data, user level prediction can be performed based on the consumption behavior habit information corresponding to the user's consumption habits to determine the user level data.
  • CNN Convolutional Neural Network
  • fusion detection can be performed based on user level data, classification probability data, and voiceprint feature data, that is, three attribute detection data of user level data, classification probability data, and voiceprint feature data are input to a pre-trained fully connected network model, The three types of attribute detection data of user level data, classification probability data, and voiceprint feature data are fused through the fully connected network model to detect voice violations, and the detection results output by the fully connected network model are obtained, and then the fully connected network model can be The output detection result is used as the voice behavior detection result, so that it can be determined whether the audio file data contains illegal voice data based on the voice behavior detection result, which realizes the prediction of the voice violation behavior and avoids the detection time lag of the voice violation behavior in the related art happening.
  • An audio detection method provided by an embodiment of the present application may further include: determining that the audio file data contains illegal voice data if the voice is the detection result of the voice violation behavior; prohibiting transmission or playback of the violation voice data.
  • the audio file data that the target user currently needs to send can be acquired to perform voice violation detection to determine whether the audio file data sent by the target user contains violation voice data corresponding to the voice violation.
  • the voice behavior detection results are divided into voice violation behavior detection results and normal voice behavior detection results.
  • the voice behavior detection results output by the fully connected network model are voice normal behavior detection results
  • the output voice behavior detection result is a voice violation detection result
  • Sending request corresponding to the audio file data to prohibit the transmission of the illegal voice data contained in the audio file data, so as to avoid the negative impact of the illegal voice data and ensure the normal use of the user.
  • the audio file data acquired in the embodiment of the present application may also be other audio file data, such as audio file data to be played.
  • the playback of the audio file data may be prohibited based on the voice violation detection result.
  • the software detects If the detection result of the voice behavior corresponding to the audio file data to be played is the detection result of the voice violation, the audio file data can be discarded or ignored, that is, the audio file data is not played, to prohibit the playback of the audio file data. Violation of voice data; after detecting that the voice behavior detection result corresponding to the audio file data to be played is the voice normal behavior detection result, the audio file data can be played based on the voice normal behavior detection result.
  • the embodiment of the present application may block the user's voice input corresponding to the illegal voice data.
  • the user can use the software's voice input interface to perform voice input during the use of the software, so that at least one of the software and the software platform corresponding to the software can obtain the voice data it inputs, thereby
  • the audio file data may be formed based on the acquired voice data, and then audio detection may be performed based on the audio file data to determine whether the voice data is illegal voice data.
  • the software or software platform detects that the voice behavior detection result corresponding to the audio file data is the voice violation detection result, it can be determined that the audio file data contains the violation voice data, that is, the voice data input by the user is determined to be the violation voice data, and then The voice input interface of the software can be turned off for the user, so that the user cannot perform voice input through the voice input interface of the software, so as to block the user's voice input.
  • the software and the software platform can also use other methods to block the user's voice input, for example, the voice input can be blocked by closing the software's voice input function, which is not limited in the embodiments of the present application.
  • the attribute detection data after determining the attribute detection data corresponding to the audio file data, the attribute detection data can be stored in a training set as the attribute detection data to be trained, so that it can be obtained from the training set when training the fully connected network model Go to the to-be-trained attribute detection data for training.
  • the audio detection method provided by the embodiment of the present application may further include: acquiring property detection data to be trained; training the property detection data to be trained to obtain a fully connected network model.
  • the attribute detection data to be trained includes various attribute detection data obtained from the training set, such as user level data, classification probability data, and voiceprint feature data.
  • the user level data, classification probability data, and voiceprint feature data can be used as the training data of the fully connected network model, that is, the user level
  • the data, classification probability data and voiceprint feature data are used as the attribute detection data to be trained, and then the model training can be performed using the classification probability data, user level data and voiceprint feature data according to the preset fully connected network structure to obtain the fully connected network model .
  • the fully-connected network model can be used to perform voice violation detection on the input attribute detection data and output voice behavior detection results.
  • the voice behavior detection result can be used to determine whether there is a voice violation behavior to determine whether the audio file data contains violation voice data corresponding to the voice violation behavior.
  • the embodiments of the present application may also use audio file data as training data to perform training based on the audio file data to obtain a corresponding network model.
  • the network model can be used to determine the attribute detection data corresponding to the audio file data, including: a neural network model and a speech classification model, etc. This embodiment of the present application does not limit this.
  • the neural network model can be used to determine the voiceprint feature data corresponding to the audio file data; the speech classification model can be used to determine the classification probability data corresponding to the audio file data.
  • the above audio detection method may further include: acquiring audio file data to be trained from a preset training set; and using a preset mobile window to slice the audio file data to be trained to obtain Frame time domain information; performing frequency domain transformation on the frame time domain information to obtain frame frequency domain information; averaging the frame frequency domain information to obtain second fixed-length data; based on the second fixed-length data and
  • the label data corresponding to the audio file data is trained according to a preset neural network algorithm to obtain the neural network model.
  • the frequency domain transform may include Fourier transform, fast Fourier transform and so on.
  • the audio file data that needs to be trained can be stored in the training set in advance, and the audio file data stored in the training set can be used as the audio file data to be trained.
  • the audio file data to be trained can be obtained from the training set, and then the audio file data to be trained can be sliced using a preset moving window to obtain time domain information of at least two frames, that is, the frame time domain Information, and then the frequency domain transformation of the frame time domain information, such as FFT transformation of the time domain information of at least two frames to obtain the frame frequency domain information; the frame frequency domain information can be averaged, such as the frame frequency domain information The average value is taken to obtain fixed-length data, and the data can be used as the second fixed-length data.
  • the corresponding tag data can be set for the audio data to be trained, so that the tag data and the second fixed-length data are used for training according to a preset neural network algorithm, such as network training according to a preset CNN algorithm until the network converges .
  • a corresponding neural network model can be constructed based on the network parameters obtained by training, so that voiceprint feature extraction can be performed through the neural network model later.
  • the neural network model may include: network parameters and at least two network layers, such as convolutional layers, fully connected layers, etc., which are not limited in the embodiments of the present application.
  • the label data corresponding to the obtained second fixed-length data and audio file data may be input into the CNN model to train the network parameters of the CNN model until the network converges.
  • the tag data may be used to mark whether the audio data to be trained contains illegal voice data corresponding to the voice violation.
  • training may be performed based on the extracted voiceprint feature data to train a speech classification model.
  • the above audio detection method may further include: slicing the acquired audio file data to be trained using a preset moving window to obtain frame time domain information; performing feature extraction on the frame time domain information to obtain amplitude Spectral feature training data and voiceprint feature training data, wherein the feature extraction includes: amplitude spectrum feature extraction and voiceprint feature extraction; performing average processing on the amplitude spectrum feature training data to obtain third fixed-length data; The amplitude spectrum feature training data and the voiceprint feature training data are stitched together to generate feature vector training data; the third fixed-length data and the feature vector training data are trained to obtain the speech classification model.
  • the audio file data to be trained can be sliced by a preset moving window to obtain frame time domain information, and then amplitude spectrum feature extraction can be performed on the frame time domain information Harmonic voiceprint feature extraction to obtain amplitude spectrum feature training data and voiceprint feature training data.
  • the frame frequency domain information can be obtained by performing FFT transformation on the frame time domain information; then, the amplitude frequency spectrum characteristics can be performed on the frame frequency domain information to obtain the amplitude spectrum characteristic training data corresponding to the audio file data to be trained, and can be based on The frame frequency domain information is used to extract voiceprint features to obtain voiceprint feature training data corresponding to the audio file data to be trained.
  • the amplitude spectrum feature training data and the voiceprint feature training data are spliced to form feature vector training data, and the feature vector training data can be used as the training data of the speech classification model to use the feature vector training data for the speech classification model training .
  • the voiceprint feature training data is a 1-dimensional vector (1, 1024) and the Mags feature training data is a 1-dimensional vector (1, 512)
  • these two vectors can be stitched together to form a 1-dimensional feature vector ( 1, 1536), and can use the one-dimensional feature vector (1, 1536) as feature vector training data, input to the preset fully connected network for training, so that a 2-layer fully connected network model can be trained, and then the The trained 2-layer fully connected network model is used as a speech classification model, so that the speech classification model can be used later for speech classification processing.
  • the embodiments of the present application can perform voiceprint feature extraction through a neural network model to obtain voiceprint feature data corresponding to audio file data, and can perform voice classification processing through a voice classification model to obtain classification probability data corresponding to audio file data.
  • voice classification processing through a voice classification model to obtain classification probability data corresponding to audio file data.
  • the user level data is determined, so that voice violation detection can be performed based on the voiceprint feature data, classification probability data, and user level data, that is, the multiple attributes corresponding to the audio file data are fused
  • the detection data for voice violation detection can effectively avoid the time lag and high cost of the related technology based on manual detection of voice violations, reduce the investment cost of voice violation detection, and improve the accuracy of voice violation detection.
  • FIG. 3 shows a structural block diagram of an embodiment of an audio detection device in an embodiment of the present application.
  • the audio detection device includes an audio file data acquisition module 310, an attribute detection data determination module 320, and voice violation detection Module 330.
  • the audio file data acquisition module 310 is configured to acquire audio file data.
  • the attribute detection data determination module 320 is configured to determine attribute detection data corresponding to the audio file data.
  • the voice violation detection module 330 is configured to perform voice violation detection on the attribute detection data through a pre-trained fully connected network model to generate a voice behavior detection result corresponding to the audio file data.
  • the attribute detection data determination module 320 may include a slice processing sub-module, a feature extraction sub-module, a data splicing sub-module, and a classification processing sub-module.
  • the slice processing sub-module is configured to slice the audio file data to obtain at least two frames of audio time domain information.
  • the feature extraction submodule is configured to perform feature extraction on the at least two frames of audio time domain information to obtain amplitude spectrum feature data and voiceprint feature data corresponding to the audio file data.
  • the data stitching submodule is configured to stitch the amplitude spectrum feature data and the voiceprint feature data to generate feature vector data.
  • the classification processing submodule is configured to perform speech classification processing on the feature vector data through a pre-trained speech classification model to obtain classification probability data as the attribute detection data.
  • the feature extraction sub-module includes a frequency domain transformation unit and an amplitude spectrum feature extraction unit.
  • the frequency domain transformation unit is configured to perform frequency domain transformation on the at least two frames of audio time domain information to obtain audio frequency domain information.
  • the amplitude spectrum feature extraction unit is configured to perform amplitude spectrum feature extraction based on the audio frequency domain information to obtain amplitude spectrum feature data corresponding to the audio file data.
  • the classification processing sub-module includes an average processing unit and a classification processing unit.
  • the mean value processing unit is configured to perform mean value processing on the amplitude spectrum characteristic data to obtain fixed-length data.
  • the classification processing unit is set to perform speech classification processing on the feature vector data through a pre-trained neural network model based on the fixed-length data to obtain classification probability data corresponding to the audio file data.
  • the attribute detection data determination module 320 includes a slice processing sub-module, a frequency domain transformation sub-module, a frequency domain mean processing sub-module, and a voiceprint feature extraction sub-module.
  • the slice processing sub-module is configured to slice the audio file data to obtain at least two frames of audio time domain information.
  • the frequency domain transformation sub-module is configured to perform frequency domain transformation on the at least two frames of audio time domain information to obtain audio frequency domain information.
  • the frequency domain mean value processing sub-module is configured to perform mean value processing on the audio frequency domain information to obtain the first fixed-length data.
  • the voiceprint feature extraction sub-module is configured to extract voiceprint features through a pre-trained neural network model based on the first fixed-length data to obtain voiceprint feature data corresponding to the attribute detection data of the audio file data.
  • the audio detection device may further include a historical behavior data acquisition module and a household level data determination module.
  • the historical behavior data acquisition module is set to acquire historical behavior data of the target user.
  • the user level data determination module is configured to obtain user level data as the attribute detection data according to the historical behavior data.
  • the historical behavior data includes at least one of the following: historical login data, user consumption behavior data, violation historical data and recharge historical data.
  • the attribute detection data determination module 320 may include a behavior data acquisition sub-module and a normalization processing sub-module.
  • the behavior data acquisition submodule is configured to acquire historical behavior data of the target user for the audio file data.
  • the normalization processing sub-module is configured to perform normalization processing on the historical behavior data to determine user level data of the target user.
  • the above historical log-in data includes: log-in times, log-in duration, log-in address, etc.
  • the above attribute detection data may include at least two of the following: user level data, classification probability data, and voiceprint feature data, the user level data is used to characterize the user level, and the classification probability data is used to characterize voice violations Corresponding to the classification probability, the voiceprint feature data is used to characterize the voiceprint feature corresponding to the audio file data.
  • the voice violation detection module 330 includes an input sub-module and an output sub-module.
  • the input submodule is configured to input the attribute detection data to the fully connected network model for detection.
  • the output submodule is configured to use the detection result output by the fully connected network model as the voice behavior detection result.
  • the audio detection device may further include a violation voice data determination module, a transmission prohibition module, a static playback module, and a voice input shielding module.
  • the violation voice data determination module is configured to determine that the audio file data contains violation voice data when the voice is the detection result of the voice violation behavior detection result.
  • the transmission prohibited module is set to prohibit transmission of the illegal voice data.
  • the static playback module is set to prohibit playback of illegal voice data.
  • the voice input blocking module is set to block the user's voice input corresponding to the illegal voice data.
  • the audio detection device may further include a training data acquisition module, a slicing module, a frequency domain transformation module, an average processing module, and a neural network training module.
  • the training data acquisition module is set to acquire the audio file data to be trained from the preset training set.
  • the slicing module is set to use a preset moving window to slice the audio file data to be trained to obtain frame time domain information.
  • the frequency domain transform module is configured to perform frequency domain transform on the frame time domain information to obtain frame frequency domain information.
  • the mean value processing module is configured to perform mean value processing on the frame frequency domain information to obtain second fixed-length data.
  • the neural network training module is set to train based on the second fixed-length data and the label data corresponding to the audio file data according to a preset neural network algorithm to obtain the neural network model.
  • the audio detection device may further include a slicing module, a feature extraction module, an average processing module, a training data splicing module, and a speech classification model training module.
  • the slicing module is set to use a preset moving window to slice the acquired audio file data to be trained to obtain frame time domain information.
  • the feature extraction module is configured to perform feature extraction on the time domain information of the frame to obtain amplitude spectrum feature training data and voiceprint feature training data, where the feature extraction includes: amplitude spectrum feature extraction and voiceprint feature extraction.
  • the mean value processing module is configured to perform mean value processing on the amplitude spectrum feature training data to obtain third fixed-length data.
  • the training data splicing module is configured to splice the amplitude spectrum feature training data and voiceprint feature training data to generate feature vector training data.
  • the speech classification model training module is configured to train the third fixed-length data and the feature vector training data to obtain the speech classification model.
  • the audio detection device may further include: a fully connected network model training module.
  • the fully-connected network model training module is set to obtain the attribute detection data to be trained; training the to-be-attended attribute detection data to obtain a fully-connected network model.
  • the attribute detection data includes at least one of the following: user level data, classification probability data, voiceprint feature data, etc., which are not limited in the embodiments of the present application.
  • the audio detection device provided above can execute the audio detection method provided in any embodiment of the present application.
  • the above audio detection device may be integrated in the device.
  • the device may be composed of at least two physical entities, or a physical entity.
  • the device may be a PC, a computer, a mobile phone, a tablet device, a personal digital assistant, a server, a messaging device, a game console, and so on.
  • An embodiment of the present application further provides a device, including: a processor and a memory. At least one instruction is stored in the memory, and the instruction is executed by the processor, so that the device executes the audio detection method described in the foregoing method embodiment.
  • the device includes a processor 40, a memory 41, a display screen 42 with a touch function, an input device 43, an output device 44, and a communication device 45.
  • the number of processors 40 in the device may be at least one, and one processor 40 is taken as an example in FIG. 4.
  • the number of the memory 41 in the device may be at least one, and one memory 41 is taken as an example in FIG. 4.
  • the processor 40, the memory 41, the display screen 42, the input device 43, the output device 44, and the communication device 45 of the device may be connected by a bus or other means. In FIG. 4, the connection by a bus is used as an example.
  • the memory 41 as a computer-readable storage medium may be configured to store software programs, computer executable programs, and modules, such as program instructions / modules corresponding to the audio detection method described in any embodiment of the present application (for example, in an audio detection device Audio file data acquisition module 310, attribute detection data determination module 320, and voice violation detection module 330, etc.).
  • the memory 41 may mainly include a storage program area and a storage data area, wherein the storage program area may store operation devices and application programs required for at least one function; the storage data area may store data created according to the use of the device and the like.
  • the memory 41 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
  • the memory 41 may include memories remotely provided with respect to the processor 40, and these remote memories may be connected to the device through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
  • the display screen 42 is a display screen 42 with a touch function, which may be a capacitive screen, an electromagnetic screen or an infrared screen.
  • the display screen 42 is configured to display data according to the instructions of the processor 40, and is also configured to receive a touch operation acting on the display screen 42 and send corresponding signals to the processor 40 or other devices.
  • the display screen 42 is an infrared screen, it further includes an infrared touch frame.
  • the infrared touch frame is disposed around the display screen 42.
  • the infrared touch frame is also configured to receive an infrared signal, and the The infrared signal is sent to the processor 40 or other devices.
  • the communication device 45 is configured to establish a communication connection with other devices, which may be at least one of a wired communication device and a wireless communication device.
  • the input device 43 is configured to receive input digital or character information, and generate key signal input related to user settings and function control of the device, and is also configured to be a camera for acquiring images and a sound pickup device for acquiring audio data.
  • the output device 44 may include an audio device such as a speaker. It should be noted that the composition of the input device 43 and the output device 44 can be set according to actual conditions.
  • the processor 40 executes various functional applications and data processing of the device by running software programs, instructions, and modules stored in the memory 41, that is, implementing the audio detection method described above.
  • the processor 40 executes at least one program stored in the memory 41, the following operations are implemented: acquiring audio file data; determining attribute detection data corresponding to the audio file data; through a pre-trained fully connected network model, Perform voice violation detection on the attribute detection data to generate a voice behavior detection result corresponding to the audio file data.
  • Embodiments of the present application also provide a computer-readable storage medium, when instructions in the storage medium are executed by a processor of a device, so that the device can execute the audio detection method described in the foregoing method embodiments.
  • the audio detection method includes: acquiring audio file data; determining attribute detection data corresponding to the audio file data; performing a voice violation detection on the attribute detection data through a pre-trained fully connected network model to generate The voice behavior detection result corresponding to the audio file data is described.
  • the present application can be implemented by software and necessary general-purpose hardware, or of course by hardware.
  • the technical solutions of the present application can essentially be embodied in the form of software products, and the computer software products can be stored in computer-readable storage media, such as computer floppy disks, Read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), flash memory (FLASH), hard disk or optical disc, etc., including multiple instructions to make a computer device (which can be a robot, Personal computers, servers, or network devices, etc.) execute the audio detection method described in any embodiment of the present application.
  • a computer device which can be a robot, Personal computers, servers, or network devices, etc.
  • each unit and module included are only divided according to the function logic, but it is not limited to the above division, as long as the corresponding function can be realized; in addition, each function
  • the specific names of the units are only for the purpose of distinguishing each other, and are not used to limit the protection scope of the present application.
  • each part of the present application may be implemented in hardware, software, firmware, or a combination thereof.
  • multiple steps or methods may be implemented by software or firmware stored in a memory and executed by a suitable instruction execution device.
  • a logic gate circuit for implementing a logic function on a data signal
  • Discrete logic circuits dedicated integrated circuits with appropriate combinational logic gates
  • programmable gate arrays PROM
  • field programmable gate arrays Field Programmable Gate Array, FPGA

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Business, Economics & Management (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Game Theory and Decision Science (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Telephonic Communication Services (AREA)
  • Electrically Operated Instructional Devices (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

一种音频检测方法、装置、设备及存储介质,涉及计算机网络技术领域。该音频检测方法包括:获取音频文件数据(110);确定音频文件数据对应的属性检测数据(120);通过预先训练的全连接网络模型,对属性检测数据进行语音违规行为检测,生成音频文件数据对应的语音行为检测结果(130)。

Description

一种音频检测方法、装置、设备及存储介质
本申请要求在2018年10月10日提交中国专利局、申请号为201811178750.2的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
技术领域
本申请涉及计算机网络技术领域,例如一种音频检测方法、装置、设备及存储介质。
背景技术
随着计算机网络技术的快速发展,诸如个人计算机(Personal Computer,PC)、手机、平板电脑等设备越来越普及,给人们的生活、学习及工作带来了极大的便利。
作为设备的一个应用,用户可以使用设备,通过网络与其他用户进行语音沟通,如可以使用设备中所安装的带有语音聊天功能的软件,通过网络与其他用户进行语音聊天,也可以通过加入特定聊天室或聊天群参与多人的语音聊天和娱乐。用户可以通过设备将其所需要发送的语音信息发送给网络,使得其他用户可以通过网络接收到该语音信息并播放,达到语音沟通的目的。在实际聊天环境中,尤其是多人聊天时,语音信息涉及的内容范围较广,其中可能包含令人不适的语音数据,诸如刺耳、高分贝、内容不恰当等,这些语音通常是由个别用户恶意发出旨在干扰其他用户的正常使用,因而软件运营方将这类语音列为违规语音数据。
为了打击违规语音数据,保障用户的正常使用体验,避免用户流失而影响商业运营,软件运营方做了很多努力和探索,但收效有限。相关技术中,经常采用两种方案,一种是在软件上配置有举报入口,供正常用户举办违规用户,软件平台根据举报的线索作相应处理和惩罚;另一种是在平台侧部署人力,通过人工抽查或监控处理违规语音。对于具有大量活跃用户的软件平台,同一时间内经常同时并存数目极大的聊天室,各种违规语音数据很可能会大量随机出现,由此可知,上述两种方案均难以有效制止同一时间内随机出现的大量违规语音数据,也整体上很难避免违规语音影响用户正常体验的情况,并且偏向事后或者事情发生到一定程度后才介入,因此存在时间滞后,并且投入代价大。
发明内容
本申请实施例提供一种音频检测方法、系统、设备以及存储介质,通过自动检测语音违规行为,以避免传统基于人工检测违规语音数据的方法中所存在的检测时间滞后的情况。
第一方面,本申请实施例提供了一种音频检测方法,包括:获取音频文件数据;确定音频文件数据对应的属性检测数据;通过预先训练的全连接网络模型,对属性检测数据进行语音违规行为检测,生成音频文件数据对应的语音行为检测结果。
第二方面,本申请实施例还提供了一种音频检测装置,包括:音频文件数据获取模块,设置为获取音频文件数据;属性检测数据确定模块,设置为确定音频文件数据对应的属性检测数据;语音违规行为检测模块,设置为通过预先训练的全连接网络模型,对属性检测数据进行语音违规行为检测,生成音频文件数据对应的语音行为检测结果。
第三方面,本申请实施例还提供了一种设备,包括:处理器和存储器;所述存储器中存储有至少一条指令,所述指令由所述处理器执行,使得所述设备执行如第一方面所述的音频检测方法。
第四方面,本申请实施例还提供了一种计算机可读存储介质,所述存储介质中的指令由设备的处理器执行时,使得所述设备能够执行如第一方面所述的音频检测方法。
附图说明
图1是本申请一实施例中的一种音频检测方法的步骤流程示意图;
图2是本申请一实施例中的音频文件数据的检测流程示意图;
图3是本申请一实施例中的一种音频检测装置实施例的结构方框示意图;
图4是本申请一实施例中的一种设备的结构方框示意图。
具体实施方式
参照图1,示出了本申请的一实施例中的一种音频检测方法的步骤流程示意图,可以包括步骤110至步骤130。
在步骤110中,获取音频文件数据。
示例性的,本申请实施例在语音违规行为检测过程中,可以获取当前所需要检测的音频文件数据,以检测当前获取到音频文件数据是否包含语音违规行为对应的违规语音数据。其中,音频文件数据可以用于表征音频文件所包含的数据,如可以是声音文件中的语音数据。需要说明的是,声音文件可以是设备依据用户讲话声音生成的,可以用于承载表征用户所需要发送的语音数据,可以包括:语音违规行为对应的违规语音数据、符合预设语音行为规定的语音数据(即合规语音数据)等,本申请实施例对此不作限制。
在一实施例中,用户的讲话声音形成语音数据,该语音数据可以通过软件接口输入到软件中,使得软件可以将该语音数据确定为音频文件数据,即软件可以获取到音频文件数据,从而使得软件可以对该音频文件数据自动进行音频检测;或者,软件可以通过网络,将获取到的音频文件数据传输给软件平台,使得软件平台可以获取到该音频文件数据,进而可以通过软件平台对该音频文件数据进行音频检测。
在步骤120中,确定所述音频文件数据对应的属性检测数据。
本申请一实施例在获取到音频文件数据后,可以对该音频文件数据进行数据处理,确定音频文件数据对应的属性检测数据,以便后续可以对该属性检测数据进行语音违规行为检测。其中,属性检测数据可以包括以下至少一项:用户等级数据、分类概率数据和声纹特征数据等,本申请实施例对此不作限制。
需要说明的是,用户等级数据、分类概率数据和声纹特征数据等均可以预先获得并存储于预定位置中,以供步骤120执行时从预定位置中读取进行比对或确定。用户等级数据可以用于确定目标用户的用户等级,且用户等级可以依据目标用户的消费习惯、历史登录行为预先确定,如可以通过对目标用户的历史充值记录和历史登录行为数据进行归一化处理,确定出该目标用户的用户等级。其中,历史登录行为数据可以用于表征用户的历史登录行为,如作为历史登录行为数据的历史登录时长数据可以用于表征用户历史登录的时长;目标用户可以是指发送音频文件数据的用户,如可以是在语音沟通过程中,发送语音信息的用户等。目标用户的用户等级越高,可以表征该目标用户发送违规语音数据的可能性越小。
分类概率数据可以用于表征语音违规行为对应的分类概率,如语音违规概率、语音合规概率等。其中,语音违规概率可以是指:音频文件数据中所包含的语音数据是违规语音数据的概率;语音合规概率可以是指:音频文件数据中 所包含的语音数据是合规语音数据的概率。在一实施例中,语音违规概率的大小可以参照以下因素:刺耳、高分贝、音频内容信息等。例如,在音频文件数据含有恶意发出的刺耳、高分贝、内容不恰当等令人不适的语音数据的情况下,会增大该音频文件数据对应的语音违规概率,即音频含有用户恶意发出的刺耳、高分贝、内容不恰当等令人不适的信息,会导致该音频的违规概率增大。
声纹特征数据可以用于表征音频文件数据对应的声纹特征。该声纹特征可以是指用户的声音纹理特征。声音纹理特征可以是指:将用户的原始音频波形经过傅里叶变换及其它后处理生成的声音在时域上的频域特征。
在一实施例中,当获取到音频文件数据后,根据该音频文件数据对应的用户信息,查找或者获取该用户对应的用户等级数据、分类概率数据或者声纹特征数据等,这些数据的至少一种作为该音频文件数据的属性检测数据。因此,在本申请的一实施例中,上述步骤120确定的属性检测数据可以包括以下至少两项:用户等级数据、分类概率数据和声纹特征数据。其中,所述用户等级用于表征用户等级,所述分类概率数据用于表征语音违规行为对应的分类概率,声纹特征数据用于表征音频文件数据对应的声纹特征。示例性的,属性检测数据同时包含用户等级数据、分类概率数据和声纹特征数据,以便语音行为检测结果更加准确。
在步骤130中,通过预先训练的全连接网络模型,对所述属性检测数据进行语音违规行为检测,生成所述音频文件数据对应的语音行为检测结果。
本申请实施例中,在确定出音频文件数据对应的属性检测数据后,可以将该属性检测数据作为预先训练的全连接网络模型的输入,随后可将属性检测数据输入到预先训练的全连接网络模型中,以通过该全连接网络模型进行语音违规行为,生成音频文件数据对应的语音行为测结果。在语音行为检测结果为语音违规行为检测结果的情况下,可以确定音频文件数据包含违规语音数据,即检测出语音违规行为对应的违规语音数据。在语音行为检测结果不是语音违规行为检测结果的情况下,可以确定音频文件数据不包含违规语音数据,在语音行为检测结果为语音正常行为检测结果的情况下,可以确定音频文件数据未含有语音违规行为对应的违规语音数据。
综上,本申请实施例在获取到音频文件数据后,可通过确定音频文件数据对应的属性检测数据,来进行语音违规行为检测,从而能够及时检测出语音违规行为对应的违规语音数据,确保用户的正常使用,避免了相关基于用户举报 和人工抽查导致语音违规行为检测的时间滞后的情况,保障用户的正常使用体验,满足用户需求,投入代价小。
在此基础上,当检测到用户拟发出的语音违规后,限制或暂停该用户的语音数据传输给其他用户,在一实施例中,根据该用户的音频文件数据判断出该用户违规时,给该用户的客户端发送相应信息,使该客户端暂停用户的语音聊天功能。
本申请一实施例在获取到音频文件数据后,可以对该音频文件数据进行切片处理,从而可以得到切片处理后的至少两帧的音频时域信息;随后,可以对得到的至少两帧的音频时域信息进行特征提取,得到音频特征数据,以对该音频特征数据确定出音频文件数据对应的属性检测数据。其中,音频特征数据可以用于表征音频特征,如表征音频的振幅特征、声纹特征等。
例如,在音频特征数据为振幅谱(Magnitude Spectrum,Mags)特征数据的情况下,可以依据该振幅谱特征数据生成音频文件数据对应的分类概率数据,以便后续可以对该分类概率数据进行语音违规行为检测。振幅谱(Magnitude Spectrum,Mags)特征数据可以用于表征音频的Mags特征。需要说明的是,申请人在实现本申请实施例过程中,经过深入分析和实验发现,Mags特征在声音违规检测中,效果良好,因而如后文所述,本申请实施例的一些方案结合Mags特征作了相应扩展。
在本申请的一实施例中,上述确定所述音频文件数据对应的属性检测数据的步骤,可以包括:对所述音频文件数据进行切片处理,得到至少两帧音频时域信息;对所述至少两帧音频时域信息进行特征提取,得到振幅谱特征数据和声纹特征数据;对所述振幅谱特征数据和所述声纹特征数据进行拼接,生成特征向量数据;通过预先训练的语音分类模型,对所述特征向量数据进行语音分类处理,得到作为所述属性检测数据的分类概率数据。其中,特征提取可以包括:振幅谱特征提取、声纹特征提取等,本申请实施例对此不作限制。
例如,在语音通讯过程中,可以获取待发送的音频文件数据,并可利用预设的移动窗口对获取到的音频文件数据进行切片,得到多帧的音频时域信息。从而,可以对至少两帧的音频时域信息进行Mags特征提取,得到Mags特征数据;并且可以对至少两帧的音频时域信息进行声纹特征提取,得到所述音频文件数据对应的声纹特征数据;随后,可以将得到的Mags特征数据与到声纹特征数据进行拼接,生成一维的特征向量数据,并将该特征向量数据作为语音分类 模型的输入,通过该语音分类模型进行语音分类处理,得到所述音频文件数据对应的分类概率数据。在一实施例中,语音分类模型可以根据不同的音频输入提取音频的特征,即根据输入的特征向量数据进行特征提取,得到输入特征;随后可根据输入特征的分布,分配该音频输入一个概率值,即为输入的特征向量数据分配一个分类概率数据,并输出,以作为语音违规行为检测的属性检测数据。在输入特征与预设的违规样本的特征相似的情况下,语音分类模型会分配给输入的特征向量数据一个高违规概率,例如输出90%作为特征向量数据的语音违规概率。在输入特征与预设的正常样本的特征相似的情况下,语音分类模型会分配给输入的特征向量数据一个低违规概率,例如输出1%作为特征向量数据的语音违规概率,即特征向量数据的语音合规概率为99%。
在一实施例中,对所述至少两帧音频时域信息进行特征提取,得到振幅谱特征数据,可以包括:对所述至少两帧音频时域信息进行频域变换,得到音频频域信息;基于所述音频频域信息进行振幅谱特征提取,得到振幅谱特征数据。其中,频域变换可以包括傅里叶变换,如快速傅氏变换(Fast Fourier Transformation,FFT)等,本申请实施例对此不作限制。
作为本申请的一个示例,在利用预设的移动窗口将音频文件数据切成至少两个小段后,可以将每一小段称为一帧音频时域信息;随后,可以对得到的每帧音频时域信息进行傅里叶变换,得到每帧音频时域信息对应的音频频域信息,可以基于该音频频域信息进行振幅谱特征提取,得到所述音频文件数据对应的振幅谱特征数据,如对音频频域信息的Mags取均值和方差,然后将取到的均值和方差作为振幅谱特征数据,以基于振幅谱特征数据生成特征向量数据,从而可以对该特征向量数据进行语音分类处理,得到所述音频文件数据对应的分类概率数据。其中,相邻的两帧音频时域信息可以有交叠部分,即帧与帧之间可以有交叠,如在一帧音频时域信息的帧长为25毫秒(ms),帧移为10毫秒时,两帧之间可以有15毫秒的交叠。需要说明的是,帧长和帧移可以依据精确度要求进行设置,本示例对此不作限制。
当然,也可以将得到的音频特征数据作为属性检测数据,以对该音频特征数据进行语音违规行为检测,如在音频特征数据为声纹特征数据的情况下,可以将该声纹特征数据作为属性检测数据,以对该音频特征数据进行语音违规行为检测,本申请实施例对此不作限制。
在本申请的一实施例中,上述确定所述音频文件数据对应的属性检测数据 的步骤,可以包括:对所述音频文件数据进行切片处理,得到至少两帧音频时域信息;对所述至少两帧音频时域信息进行频域变换,得到音频频域信息;对所述音频频域信息进行均值处理,得到第一定长数据;基于所述第一定长数据,通过预先训练的神经网络模型进行声纹特征提取,得到作为所述属性检测数据的声纹特征数据。例如,在获取到待检测的音频文件数据后,可以利用移动窗口对该音频文件数据进行切片处理,得到至少两帧音频时域信息,随后对得到的至少两帧音频时域信息进行FFT变换,得到音频频域信息,对该音频频域信息进行均值处理,如对该音频频域信息取均值,得到定长的数据,并可将得到的定长的数据确定为第一定长数据,以将该第一定长数据输入到预先训练的神经网络模型中进行声纹特征提取,得到声纹特征数据,然后可将该声纹特征数据作为音频文件数据对应的属性检测数据,输入到全连接网络模型中,以对声纹特征数据进行语音违规行为检测,生成所述音频文件数据对应的语音行为检测结果。其中,待检测的音频文件数据可以包括待传输的音频文件数据、待播放的音频文件数据等;待传输的音频文件数据可以用于表征语音通讯过程中待发送的音频文件数据;待播放的音频文件数据可以用于表征待播放语音的文件数据。
本申请一实施例可以基于用户消费习惯对用户进行分级,以基于分级后得到的用户级别进行语音违规行为检测,以预测出语音违规行为。在一实施例中,本申请实施例提供的音频检测方法还可以包括:获取目标用户的历史行为数据;根据所述历史行为数据得到作为所述属性检测数据的用户等级数据。其中,所述历史行为数据包括以下至少一项:历史登录数据、用户消费行为数据、违规历史数据、用户消费行为数据和充值历史数据等。
示例性的,在需要确定某一个用户的用户等级时,可以将该用户确定为目标用户,并可以根据该目标用户的用户标识,如用户账号、用户名等,从数据库中获取该目标用户的历史行为数据,以根据获取到的历史行为数据确出该目标用户的用户等级,即确定出目标用户的用户等级数据。随后可以将用户等级数据存储到的数据库中,以在音频检测时从该数据库中获取到该用户等级数据作为属性检测数据。
当然,本申请实施例也可在音频检测时,针对获取到的音频文件数据,获取该音频文件数据对应目标用户的历史行为数据,以根据获取到的历史行为数据实时确定出作为属性检测数据的用户等级数据,从而可以采用实时确定出的 用户等级数据进行语音违规行为检测,提高语音违规行为检测的准确性。例如,上述确定所述音频文件数据对应的属性检测数据的步骤,包括:针对所述音频文件数据,获取目标用户的历史行为数据;对所述历史行为数据进行归一化处理,得到作为所述属性检测数据的用户等级数据。
作为本申请的一个示例,在获取到音频文件数据后,可以将发送该音频文件数据的用户确定为目标用户,随后可针对该音频文件数据,依据目标用户的用户标识,获取该目标用户的历史行为数据,如用户消费行为数据,历史登录数据、违规历史数据和充值历史数据等其中至少一项数据。其中,用户消费行为数据可以用于确定目标用户的消费行为习惯信息;违规历史数据可以用于确定目标用户的语音违规历史信息,如确定目标用户是否有违规历史,或者确定目标用户的违规历史次数等;充值历史数据可以用于确定目标用户的充值历史信息,如目标用户的充值次数,历史充值金额等;用户历史登录数据可以用于确定目标用户的历史登录行为,包括:登录次数、登录时长、登录地址等。登录次数可以用于表征目标用户的登录数量;登录时长可以用于表征目标用户的历史登录时长,如可以包括目标用户每一次登录对应的登录时长;登录地址可以用于确定用户每一次登录的地址,如可以是目标用户登录时所使用的设备的互联网协议(Internet Protocol,IP)地址、媒体访问控制(Medium Access Control,MAC)地址等,本示例对此不作限制。
在获取到目标用户的历史行为数据后,可对获取到的历史行为数据进行归一化处理,如对目标用户的目标用户的登陆数量、登陆时长、是否有违规历史、充值历史等信息进行数值化和归一化,从而可以基于归一化处理结果确定出目标用户的用户等级数据,随后可将该用户等级数据作为属性检测数据,输入到预先训练的全连接网络模型中进行语音违规行为检测,生成语音行为检测结果。
本申请一实施例中,通过预先训练的全连接网络模型,对所述属性检测数据进行语音违规行为检测,生成所述音频文件数据对应的语音行为检测结果,包括:将所述属性检测数据输入到所述全连接网络模型进行检测;将所述全连接网络模型输出的检测结果作为所述语音行为检测结果。
作为本申请的一实施例,在获取到音频文件数据后,如图2所示,可以基于诸如卷积神经网络(Convolutional Neural Network,CNN)进行声音纹理提取,得到该音频文件数据对应的声纹特征数据;并可以基于Mags特征进行声音分类,即采用音频文件数据对应的Mags特征数据生成特征向量数,并将该特征向量数 输入到语音分类模型进行语音分类处理,得到所述音频文件数据对应的分类概率数据;以及可以针对当前获取到的音频文件数据,基于用户消费习惯对应的消费行为习惯信息进行用户等级预测,确定出用户等级数据。随后,可以基于用户等级数据、分类概率数据以及声纹特征数据进行融合检测,即将用户等级数据、分类概率数据以及声纹特征数据这三种属性检测数据输入到预先训练好的全连接网络模型,以通过全连接网络模型融合用户等级数据、分类概率数据以及声纹特征数据这三种属性检测数据,来进行语音违规行为检测,得到全连接网络模型输出的检测结果,然后可以将全连接网络模型输出的检测结果作为语音行为检测结果,从而可以基于该语音行为检测结果确定出音频文件数据是否包含违规语音数据,实现了语音违规行为的预测,避免了相关技术中语音违规行为的检测时间滞后的情况。
本申请一实施例提供的音频检测方法还可以包括:在所述语音为检测结果为语音违规行为检测结果的情况下,确定所述音频文件数据包含违规语音数据;禁止传输或播放所述违规语音数据。
示例性的,在语音通讯过程中,可以获取目标用户当前所需要发送的音频文件数据进行语音违规行为检测,以判断该目标用户发送的音频文件数据是否包含有语音违规行为对应的违规语音数据。语音行为检测结果分为语音违规行为检测结果和语音正常行为检测结果,在全连接网络模型输出的语音行为检测结果为语音正常行为检测结果的情况下,可以确定当前获取到的音频文件数据不包含违规语音数据,随后可以基于语音正常行为检测结果发送该音频文件数据,使得与该目标用户进行语音通讯的其他用户可以接收到该音频文件数据并播放,达到语音通讯的目的;在全连接网络模型输出的语音行为检测结果为语音违规行为检测结果的情况下,可以确定当前获取到的音频文件数据包含违规语音数据,随后可以基于语音违规行为检测结果,禁止该音频文件数据的发送,如驳回该音频文件数据对应的发送请求,以禁止传输该音频文件数据中所包含的违规语音数据,从而可以避免违规语音数据所带来的负面影响,确保用户的正常使用。
当然,本申请实施例中获取到的音频文件数据还可以是其他音频文件数据,如可以是待播放的音频文件数据等。在检测到待播放的音频文件数据对应的语音行为检测结果为语音违规行为检测结果的情况下,可以基于语音违规行为检测结果,禁止该音频文件数据的播放,示例性的,在软件在检测到待播放的音 频文件数据对应的语音行为检测结果为语音违规行为检测结果的情况下,可以丢弃或忽略该音频文件数据,即不对该音频文件数据进行播放,以禁止播放该音频文件数据中所包含的违规语音数据;在检测到待播放的音频文件数据对应的语音行为检测结果为语音正常行为检测结果后,可以基于语音正常行为检测结果对音频文件数据进行播放等。
此外,本申请实施例在确定出音频文件数据包含违规语音数据后,可以屏蔽该违规语音数据对应用户的语音输入。在一实施例中,用户可以在使用软件过程中,可以通过该软件的语音输入接口进行语音输入,使得软件和该软件对应的软件平台中至少之一可以获取到其所输入的语音数据,从而可以基于获取到的语音数据形成音频文件数据,随后可基于该音频文件数据进行音频检测,以确定语音数据是否是违规语音数据。在软件或者软件平台检测到音频文件数据对应的语音行为检测结果为语音违规行为检测结果的情况下,可以确定音频文件数据包含违规语音数据,即确定该用户输入的语音数据为违规语音数据,然后可以针对该用户关闭软件的语音输入接口,使得该用户不可以通过该软件的语音输入接口进行语音输入,以屏蔽该用户的语音输入。当然,软件和软件平台中至少之一也可以采用其他方式来屏蔽用户的语音输入,如可以通过关闭软件的语音输入功能来实现语音输入的屏蔽等,本申请实施例对此不作限制。
本申请一实施例在确定出音频文件数据对应的属性检测数据后,可以将该属性检测数据存储到一个训练集中,作为待训练属性检测数据,以便在训练全连接网络模型时可以从训练集中获取到该待训练属性检测数据进行训练。示例性的,本申请实施例提供的音频检测方法还可以包括:获取待训练属性检测数据;对所述待训练属性检测数据进行训练,得到全连接网络模型。其中,所述待训练属性检测数据包括从训练集获取到的各种属性检测数据,如用户等级数据、分类概率数据和声纹特征数据等。
例如,在确定出音频文件数据对应的用户等级数据、分类概率数据和声纹特征数据后,可以将用户等级数据、分类概率数据以及声纹特征数据作为全连接网络模型的训练数据,即将用户等级数据、分类概率数据以及声纹特征数据作为待训练属性检测数据,然后可按照预设的全连接网络结构,采用分类概率数据、用户等级数据以及声纹特征数据进行模型训练,得到全连接网络模型。该全连接网络模型可以用于对输入的属性检测数据进行语音违规行为检测,输出语音行为检测结果。该语音行为检测结果可以用于判断是否存在语音违规行 为,以确定音频文件数据是否包含语音违规行为对应的违规语音数据。
当然,本申请实施例也可以将音频文件数据作为训练数据,以基于该音频文件数据进行训练,得到相应的网络模型。该网络模型可以用于确定音频文件数据对应的属性检测数据,包括:神经网络模型和语音分类模型等,本申请实施例对此不作限制。其中,神经网络模型可以用于确定音频文件数据对应的声纹特征数据;语音分类模型可以用于确定音频文件数据对应的分类概率数据。
在本申请的一实施例中,上述音频检测方法还可以包括:从预设的训练集中,获取待训练音频文件数据;采用预设的移动窗口,对所述待训练音频文件数据进行切片,得到帧时域信息;对所述帧时域信息进行频域变换,得到帧频域信息;对所述帧频域信息进行均值处理,得到第二定长数据;基于所述第二定长数据和所述音频文件数据对应的标签数据,按照预设的神经网络算法进行训练,得到所述神经网络模型。其中,频域变换可以包括傅里叶变换、快速傅里叶变换等。
本申请一实施例可以预先将需要进行训练的音频文件数据存储到训练集中,并可将训练集中存储的音频文件数据作为待训练音频文件数据。在模型训练过程中,可以从该训练集中获取待训练音频文件数据,然后可采用预设的移动窗口对该待训练音频文件数据进行切片,得到至少两帧的时域信息,即得到帧时域信息,随后可对帧时域信息进行频域变换,如对至少两帧的时域信息进行FFT变换,得到帧频域信息;可以对该帧频域信息进行均值处理,如对帧频域信息取均值,得到定长的数据,并可以将该数据作为第二定长数据。
此外,可以为待训练音频数据设置对应的标签数据,从而采用该标签数据和第二定长数据,按照预设的神经网络算法进行训练,如按照预设的CNN算法进行网络训练,直到网络收敛。在网络收敛的情况下,可以基于训练得到的网络参数构建对应的神经网络模型,以便后续可以通过该神经网络模型进行声纹特征提取。其中,神经网络模型可以包括:网络参数和至少两个网络层,如卷积层、全连接层等,本申请实施例对此不作限制。
例如,在神经网络模型的训练过程中,可以将得到的第二定长数据和音频文件数据对应的标签数据输入到CNN模型中,以训练该CNN模型的网络参数,直到网络收敛。其中,标签数据可以用于标记待训练音频数据是否包含语音违规行为对应的违规语音数据。
本申请一实施例在训练过程中,可以基于提取到的声纹特征数据进行训练, 以训练出语音分类模型。示例性的,上述音频检测方法还可以包括:采用预设的移动窗口,对获取到的待训练音频文件数据进行切片,得到帧时域信息;对所述帧时域信息进行特征提取,得到振幅谱特征训练数据和声纹特征训练数据,其中,所述特征提取包括:振幅谱特征提取和声纹特征提取;对所述振幅谱特征训练数据进行均值处理,得到第三定长数据;对所述振幅谱特征训练数据和声纹特征训练数据进行拼接,生成特征向量训练数据;对第三定长数据和所述特征向量训练数据进行训练,得到所述语音分类模型。
本申请一实施例在获取到待训练音频文件数据后,可以预设的移动窗口对该待训练音频文件数据进行切片,得到帧时域信息,然后可对该帧时域信息进行振幅谱特征提取和声纹特征提取,得到振幅谱特征训练数据和声纹特征训练数据。例如,可以通过对帧时域信息进行FFT变换,得到帧频域信息;然后,可对该帧频域信息进行振幅谱特征,得到待训练音频文件数据对应的振幅谱特征训练数据,并可基于该帧频域信息进行声纹特征提取,得到待训练音频文件数据对应的声纹特征训练数据。
随后,对振幅谱特征训练数据和声纹特征训练数据进行拼接,形成特征向量训练数据,并可将该特征向量训练数据作为语音分类模型的训练数据,以采用特征向量训练数据进行语音分类模型训练。例如,在声纹特征训练数据是1维向量(1,1024),Mags特征训练数据是1维向量(1,512)时,可以通过将这两个向量拼接在一起,组成1维特征向量(1,1536),并可以将该1维特征向量(1,1536)作为特征向量训练数据,输入到预设的全连接网络进行训练,从而可以训练出一个2层全连接网络模型,进而可以将训练出的这个2层全连接网络模型作为语音分类模型,以便后续可以采用该语音分类模型进行语音分类处理。
综上,本申请实施例可以通过神经网络模型进行声纹特征提取,得到音频文件数据对应的声纹特征数据,并可通过语音分类模型进行语音分类处理,得到音频文件数据对应的分类概率数据,通过对目标用户的历史行为数据进行归一化处理,确定用户等级数据,从而可以依据声纹特征数据、分类概率数据以及用户等级数据进行语音违规行为检测,即融合音频文件数据对应的多种属性检测数据进行语音违规行为检测,能够有效避免相关技术基于人工检测语音违规行为中所存在的时间滞后、代价大等情况,减少语音违规行为检测的投入代价,提高语音违规行为检测的准确度。
需要说明的是,对于方法实施例,为了简单描述,故将其都表述为一系列的动作组合,但是本领域技术人员应该知悉,本申请实施例并不受所描述的动作顺序的限制,因为依据本申请实施例,某些步骤可以采用其他顺序或者同时进行。
参照图3,示出了本申请一实施例中的一种音频检测装置实施例的结构方框示意图,该音频检测装置包括音频文件数据获取模块310、属性检测数据确定模块320以及语音违规行为检测模块330。
音频文件数据获取模块310,设置为获取音频文件数据。
属性检测数据确定模块320,设置为确定所述音频文件数据对应的属性检测数据。
语音违规行为检测模块330,设置为通过预先训练的全连接网络模型,对所述属性检测数据进行语音违规行为检测,生成所述音频文件数据对应的语音行为检测结果。
在本申请的一实施例中,属性检测数据确定模块320可以包括切片处理子模块、特征提取子模块、数据拼接子模块以及分类处理子模块。
切片处理子模块,设置为对所述音频文件数据进行切片处理,得到至少两帧音频时域信息。
特征提取子模块,设置为对所述至少两帧音频时域信息进行特征提取,得到所述音频文件数据对应的振幅谱特征数据和声纹特征数据。
数据拼接子模块,设置为对所述振幅谱特征数据和所述声纹特征数据进行拼接,生成特征向量数据。
分类处理子模块,设置为通过预先训练的语音分类模型,对所述特征向量数据进行语音分类处理,得到作为所述属性检测数据的分类概率数据。
在一实施例中,所述特征提取子模块包括频域变换单元以及振幅谱特征提取单元。
频域变换单元,设置为对所述至少两帧音频时域信息进行频域变换,得到音频频域信息。
振幅谱特征提取单元,设置为基于所述音频频域信息进行振幅谱特征提取,得到所述音频文件数据对应的振幅谱特征数据。
在一实施例中,分类处理子模块包括均值处理单元以及分类处理单元。
均值处理单元,设置为对所述振幅谱特征数据进行均值处理,得到定长数 据。
分类处理单元,设置为基于所述定长数据,通过预先训练的神经网络模型,对所述特征向量数据进行语音分类处理,得到所述音频文件数据对应的分类概率数据。
在本申请的一实施例中,属性检测数据确定模块320包括切片处理子模块、频域变换子模块、频域均值处理子模块以及声纹特征提取子模块。
切片处理子模块,设置为对所述音频文件数据进行切片处理,得到至少两帧音频时域信息。
频域变换子模块,设置为对所述至少两帧音频时域信息进行频域变换,得到音频频域信息。
频域均值处理子模块,设置为对所述音频频域信息进行均值处理,得到第一定长数据。
声纹特征提取子模块,设置为基于所述第一定长数据,通过预先训练的神经网络模型进行声纹特征提取,得到所述音频文件数据对应作为所述属性检测数据的声纹特征数据。
在本申请的一实施例中,音频检测装置还可以包括历史行为数据获取模块以及户等级数据确定模块。
历史行为数据获取模块,设置为获取目标用户的历史行为数据。
用户等级数据确定模块,设置为根据所述历史行为数据得到作为所述属性检测数据的用户等级数据。其中,所述历史行为数据包括以下至少一项:历史登录数据、用户消费行为数据、违规历史数据和充值历史数据。
本申请一实施例中,属性检测数据确定模块320可以包括行为数据获取子模块以及归一化处理子模块。
行为数据获取子模块,设置为针对所述音频文件数据,获取目标用户的历史行为数据。
归一化处理子模块,设置为对所述历史行为数据进行归一化处理,确定所述目标用户的用户等级数据。
本申请实施例中,上述历史登录数据包括:登录次数、登录时长、登录地址等等。示例性的,上述属性检测数据可以包括以下至少两项:用户等级数据、分类概率数据和声纹特征数据,所述用户等级数据用于表征用户等级,所述分类概率数据用于表征语音违规行为对应的分类概率,声纹特征数据用于表征音 频文件数据对应的声纹特征。
在本申请的一实施例中,语音违规行为检测模块330包括输入子模块以及输出子模块。
输入子模块,设置为将所述属性检测数据输入到所述全连接网络模型进行检测。
输出子模块,设置为将所述全连接网络模型输出的检测结果作为所述语音行为检测结果。
在一实施例中,该音频检测装置还可以包括违规语音数据确定模块、禁止传输模块、静止播放模块以及语音输入屏蔽模块。
违规语音数据确定模块,设置为在所述语音为检测结果为语音违规行为检测结果的情况下,确定所述音频文件数据包含违规语音数据。
禁止传输模块,设置为禁止传输所述违规语音数据。
静止播放模块,设置为禁止播放违规语音数据。
语音输入屏蔽模块,设置为屏蔽所述违规语音数据对应用户的语音输入。
在以实施例中,该音频检测装置还可以包括训练数据获取模块、切片模块、频域变换模块、均值处理模块以及神经网络训练模块。
训练数据获取模块,设置为从预设的训练集中,获取待训练音频文件数据。
切片模块,设置为采用预设的移动窗口,对所述待训练音频文件数据进行切片,得到帧时域信息。
频域变换模块,设置为对所述帧时域信息进行频域变换,得到帧频域信息。
均值处理模块,设置为对所述帧频域信息进行均值处理,得到第二定长数据。
神经网络训练模块,设置为基于所述第二定长数据和所述音频文件数据对应的标签数据,按照预设的神经网络算法进行训练,得到所述神经网络模型。
在本申请的一实施例中,音频检测装置还可以包括切片模块、特征提取模块、均值处理模块、训练数据拼接模块以及语音分类模型训练模块。
切片模块,设置为采用预设的移动窗口,对获取到的待训练音频文件数据进行切片,得到帧时域信息。
特征提取模块,设置为对所述帧时域信息进行特征提取,得到振幅谱特征训练数据和声纹特征训练数据,其中,所述特征提取包括:振幅谱特征提取和声纹特征提取。
均值处理模块,设置为对所述振幅谱特征训练数据进行均值处理,得到第三定长数据。
训练数据拼接模块,设置为对所述振幅谱特征训练数据和声纹特征训练数据进行拼接,生成特征向量训练数据。
语音分类模型训练模块,设置为对第三定长数据和所述特征向量训练数据进行训练,得到所述语音分类模型。
在一实施例中,该音频检测装置还可以包括:全连接网络模型训练模块。该全连接网络模型训练模块,设置为获取待训练属性检测数据;对所述待训练属性检测数据进行训练,得到全连接网络模型。其中,所述属性检测数据包括以下至少一项:用户等级数据、分类概率数据和声纹特征数据等,本申请实施例对此不作限制。
需要说明的是,上述提供的音频检测装置可执行本申请任意实施例所提供的音频检测方法。
在一实施例中,上述音频检测装置可以集成在设备中。该设备可以是至少两个物理实体构成,也可以是一个物理实体构成,如设备可以是PC、电脑、手机、平板设备、个人数字助理、服务器、消息收发设备、游戏控制台等。
本申请实施例还提供一种设备,包括:处理器和存储器。存储器中存储有至少一条指令,且指令由所述处理器执行,使得所述设备执行如上述方法实施例中所述的音频检测方法。
参照图4,示出了本申请一实施例中的一种设备的结构方框示意图。如图4所示,该设备包括:处理器40、存储器41、具有触摸功能的显示屏42、输入装置43、输出装置44以及通信装置45。该设备中处理器40的数量可以是至少一个,图4中以一个处理器40为例。该设备中存储器41的数量可以是至少一个,图4中以一个存储器41为例。该设备的处理器40、存储器41、显示屏42、输入装置43、输出装置44以及通信装置45可以通过总线或者其他方式连接,图4中以通过总线连接为例。
存储器41作为一种计算机可读存储介质,可设置为存储软件程序、计算机可执行程序以及模块,如本申请任意实施例所述的音频检测方法对应的程序指令/模块(例如,音频检测装置中的音频文件数据获取模块310、属性检测数据确定模块320以及语音违规行为检测模块330等)。存储器41可主要包括存储程序区和存储数据区,其中,存储程序区可存储操作装置、至少一个功能所需 的应用程序;存储数据区可存储根据设备的使用所创建的数据等。此外,存储器41可以包括高速随机存取存储器,还可以包括非易失性存储器,例如至少一个磁盘存储器件、闪存器件、或其他非易失性固态存储器件。在一些实例中,存储器41可包括相对于处理器40远程设置的存储器,这些远程存储器可以通过网络连接至设备。上述网络的实例包括但不限于互联网、企业内部网、局域网、移动通信网及其组合。
显示屏42为具有触摸功能的显示屏42,其可以是电容屏、电磁屏或者红外屏。一般而言,显示屏42设置为根据处理器40的指示显示数据,还设置为接收作用于显示屏42的触摸操作,并将相应的信号发送至处理器40或其他装置。在一实施例中,在显示屏42为红外屏的情况下,其还包括红外触摸框,该红外触摸框设置在显示屏42的四周,该红外触摸框还设置为接收红外信号,并将该红外信号发送至处理器40或者其他设备。
通信装置45,设置为与其他设备建立通信连接,其可以是有线通信装置和无线通信装置中至少一种。
输入装置43设置为接收输入的数字或者字符信息,以及产生与设备的用户设置以及功能控制有关的键信号输入,还设置为获取图像的摄像头以及获取音频数据的拾音设备。输出装置44可以包括扬声器等音频设备。需要说明的是,输入装置43和输出装置44的组成可以根据实际情况设定。
处理器40通过运行存储在存储器41中的软件程序、指令以及模块,从而执行设备的各种功能应用以及数据处理,即实现上述的音频检测方法。
在一实施例中,处理器40执行存储器41中存储的至少一个程序时,实现如下操作:获取音频文件数据;确定所述音频文件数据对应的属性检测数据;通过预先训练的全连接网络模型,对所述属性检测数据进行语音违规行为检测,生成所述音频文件数据对应的语音行为检测结果。
本申请实施例还提供一种计算机可读存储介质,所述存储介质中的指令由设备的处理器执行时,使得设备能够执行如上述方法实施例所述的音频检测方法。示例性的,该音频检测方法包括:获取音频文件数据;确定所述音频文件数据对应的属性检测数据;通过预先训练的全连接网络模型,对所述属性检测数据进行语音违规行为检测,生成所述音频文件数据对应的语音行为检测结果。
需要说明的是,对于装置、设备、存储介质实施例而言,由于其与方法实施例基本相似,所以描述的比较简单,相关之处参见方法实施例的部分说明即 可。
通过以上关于实施方式的描述,所属领域的技术人员可以清楚地了解到,本申请可借助软件及必需的通用硬件来实现,当然也可以通过硬件实现,。基于这样的理解,本申请的技术方案本质上或者说对相关技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品可以存储在计算机可读存储介质中,如计算机的软盘、只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、闪存(FLASH)、硬盘或光盘等,包括多个指令用以使得一台计算机设备(可以是机器人,个人计算机,服务器,或者网络设备等)执行本申请任意实施例所述的音频检测方法。
值得注意的是,上述音频检测装置中,所包括的每个单元和模块只是按照功能逻辑进行划分的,但并不局限于上述的划分,只要能够实现相应的功能即可;另外,每个功能单元的具体名称也只是为了便于相互区分,并不用于限制本申请的保护范围。
应当理解,本申请的每个部分可以用硬件、软件、固件或它们的组合来实现。在上述实施方式中,多个步骤或方法可以用存储在存储器中且由合适的指令执行装置执行的软件或固件来实现。例如,如果用硬件来实现,和在另一实施方式中一样,可用本领域公知的下列技术中的任一项或他们的组合来实现:具有用于对数据信号实现逻辑功能的逻辑门电路的离散逻辑电路,具有合适的组合逻辑门电路的专用集成电路,可编程门阵列(Programmable Gate Array,PGA),现场可编程门阵列(Field Programmable Gate Array,FPGA)等。
在本说明书的描述中,参考术语“一个实施例”、“一些实施例”、“一实施例”、“示例”、“具体示例”、或“一些示例”等的描述意指结合该实施例或示例描述的具体特征、结构、材料或者特点包含于本申请的至少一个实施例或示例中。在本说明书中,对上述术语的示意性表述不一定指的是相同的实施例或示例。而且,描述的具体特征、结构、材料或者特点可以在任何的至少一个实施例或示例中以合适的方式结合。

Claims (14)

  1. 一种音频检测方法,包括:
    获取音频文件数据;
    确定所述音频文件数据对应的属性检测数据;
    通过预先训练的全连接网络模型,对所述属性检测数据进行语音违规行为检测,生成所述音频文件数据对应的语音行为检测结果。
  2. 根据权利要求1所述的方法,其中所述确定所述音频文件数据对应的属性检测数据,包括:
    对所述音频文件数据进行切片处理,得到至少两帧音频时域信息;
    对所述至少两帧音频时域信息进行特征提取,得到振幅谱特征数据和声纹特征数据;
    对所述振幅谱特征数据和所述声纹特征数据进行拼接,生成特征向量数据;
    通过预先训练的语音分类模型,对所述特征向量数据进行语音分类处理,得到作为所述属性检测数据的分类概率数据。
  3. 根据权利要求2所述的方法,其中所述对所述至少两帧音频时域信息进行特征提取,得到振幅谱特征数据,包括:
    对所述至少两帧音频时域信息进行频域变换,得到音频频域信息;
    基于所述音频频域信息进行振幅谱特征提取,得到所述振幅谱特征数据。
  4. 根据权利要求1所述的方法,其中所述确定所述音频文件数据对应的属性检测数据,包括:
    对所述音频文件数据进行切片处理,得到至少两帧音频时域信息;
    对所述至少两帧音频时域信息进行频域变换,得到音频频域信息;
    对所述音频频域信息进行均值处理,得到第一定长数据;
    基于所述第一定长数据,通过预先训练的神经网络模型进行声纹特征提取,得到作为所述属性检测数据的声纹特征数据。
  5. 根据权利要求4所述的方法,还包括:
    从预设的训练集中,获取待训练音频文件数据;
    采用预设的移动窗口,对所述待训练音频文件数据进行切片,得到帧时域信息;
    对所述帧时域信息进行频域变换,得到帧频域信息;
    对所述帧频域信息进行均值处理,得到第二定长数据;
    基于所述第二定长数据和所述音频文件数据对应的标签数据,按照预设的 神经网络算法进行训练,得到所述神经网络模型。
  6. 根据权利要求1所述的方法,还包括:
    获取目标用户的历史行为数据,其中,所述历史行为数据包括以下至少一项:历史登录数据、用户消费行为数据、违规历史数据,和充值历史数据;
    根据所述历史行为数据得到作为所述属性检测数据的用户等级数据。
  7. 根据权利要求1所述的方法,其中,所述属性检测数据包括以下至少两项:用户等级数据、分类概率数据,和声纹特征数据,所述用户等级数据用于表征用户等级,所述分类概率数据用于表征语音违规行为对应的分类概率,所述声纹特征数据用于表征音频文件数据对应的声纹特征。
  8. 根据权利要求1至7任一项所述的方法,其中,所述通过预先训练的全连接网络模型,对所述属性检测数据进行语音违规行为检测,生成所述音频文件数据对应的语音行为检测结果,包括:
    将所述属性检测数据输入到所述全连接网络模型进行检测;
    将所述全连接网络模型输出的检测结果作为所述语音行为检测结果。
  9. 根据权利要求8所述的方法,还包括:
    在所述语音为检测结果为语音违规行为检测结果的情况下,确定所述音频文件数据包含违规语音数据;
    禁止传输或播放所述违规语音数据;或者,屏蔽所述违规语音数据对应用户的语音输入。
  10. 根据权利要求2或3所述的方法,还包括:
    采用预设的移动窗口,对获取到的待训练音频文件数据进行切片,得到帧时域信息;
    对所述帧时域信息进行特征提取,得到振幅谱特征训练数据和声纹特征训练数据,其中,所述特征提取包括:振幅谱特征提取和声纹特征提取;
    对所述振幅谱特征训练数据进行均值处理,得到第三定长数据;
    对所述振幅谱特征训练数据和声纹特征训练数据进行拼接,生成特征向量训练数据;
    对所述第三定长数据和所述特征向量训练数据进行训练,得到所述语音分类模型。
  11. 根据权利要求1至7任一项所述的方法,还包括:
    获取待训练属性检测数据;
    对所述待训练属性检测数据进行训练,得到全连接网络模型。
  12. 一种音频检测装置,包括:
    音频文件数据获取模块,设置为获取音频文件数据;
    属性检测数据确定模块,设置为确定所述音频文件数据对应的属性检测数据;
    语音违规行为检测模块,设置为通过预先训练的全连接网络模型,对所述属性检测数据进行语音违规行为检测,生成所述音频文件数据对应的语音行为检测结果。
  13. 一种设备,包括:处理器和存储器;
    所述存储器中存储有至少一条指令,所述指令由所述处理器执行,使得所述设备执行如权利要求1至11任一项所述的音频检测方法。
  14. 一种计算机可读存储介质,所述存储介质中的指令由设备的处理器执行时,使得所述设备能够执行如权利要求1至11任一项所述的音频检测方法。
PCT/CN2019/102172 2018-10-10 2019-08-23 一种音频检测方法、装置、设备及存储介质 Ceased WO2020073743A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
US17/282,732 US11948595B2 (en) 2018-10-10 2019-08-23 Method for detecting audio, device, and storage medium
SG11202103561TA SG11202103561TA (en) 2018-10-10 2019-08-23 Audio detection method and apparatus, and device and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201811178750.2 2018-10-10
CN201811178750.2A CN109065069B (zh) 2018-10-10 2018-10-10 一种音频检测方法、装置、设备及存储介质

Publications (1)

Publication Number Publication Date
WO2020073743A1 true WO2020073743A1 (zh) 2020-04-16

Family

ID=64763727

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/102172 Ceased WO2020073743A1 (zh) 2018-10-10 2019-08-23 一种音频检测方法、装置、设备及存储介质

Country Status (4)

Country Link
US (1) US11948595B2 (zh)
CN (1) CN109065069B (zh)
SG (1) SG11202103561TA (zh)
WO (1) WO2020073743A1 (zh)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111883139A (zh) * 2020-07-24 2020-11-03 北京字节跳动网络技术有限公司 用于筛选目标语音的方法、装置、设备和介质
CN113782036A (zh) * 2021-09-10 2021-12-10 北京声智科技有限公司 音频质量评估方法、装置、电子设备和存储介质
CN114283843A (zh) * 2021-09-27 2022-04-05 腾讯科技(深圳)有限公司 神经网络模型融合监测方法及装置
CN116705031A (zh) * 2023-07-06 2023-09-05 中国电信股份有限公司技术创新中心 违规音频检测方法及装置、电子设备、存储介质

Families Citing this family (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109065069B (zh) * 2018-10-10 2020-09-04 广州市百果园信息技术有限公司 一种音频检测方法、装置、设备及存储介质
CN109949827A (zh) * 2019-03-15 2019-06-28 上海师范大学 一种基于深度学习与强化学习的室内声学行为识别方法
CN112182441A (zh) * 2019-07-02 2021-01-05 中国移动通信集团贵州有限公司 违规数据的检测方法及装置
JP7290507B2 (ja) * 2019-08-06 2023-06-13 本田技研工業株式会社 情報処理装置、情報処理方法、認識モデルならびにプログラム
CN114125506B (zh) * 2020-08-28 2024-03-19 上海哔哩哔哩科技有限公司 语音审核方法及装置
US11948599B2 (en) * 2022-01-06 2024-04-02 Microsoft Technology Licensing, Llc Audio event detection with window-based prediction
CN114863917A (zh) * 2022-04-02 2022-08-05 深圳市大梦龙途文化传播有限公司 游戏语音检测方法、装置、设备及计算机可读存储介质
CN114882881A (zh) * 2022-05-18 2022-08-09 深圳前海微众银行股份有限公司 违规音频的识别方法、装置、设备、存储介质及程序产品
CN115733584B (zh) * 2022-11-29 2025-09-23 航天科工通信技术研究院有限责任公司 一种基于帧数据串行传输的帧数据接收方法
CN116186265B (zh) * 2023-03-01 2025-09-19 上海喜马拉雅科技有限公司 音频的违规审核方法、装置、电子设备及可读存储介质
CN116863921A (zh) * 2023-07-18 2023-10-10 中国工商银行股份有限公司 语音识别模型训练方法、语音识别方法、装置和设备

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101226743A (zh) * 2007-12-05 2008-07-23 浙江大学 基于中性和情感声纹模型转换的说话人识别方法
US20100088088A1 (en) * 2007-01-31 2010-04-08 Gianmario Bollano Customizable method and system for emotional recognition
CN107610707A (zh) * 2016-12-15 2018-01-19 平安科技(深圳)有限公司 一种声纹识别方法及装置
CN107919137A (zh) * 2017-10-25 2018-04-17 平安普惠企业管理有限公司 远程审批方法、装置、设备及可读存储介质
CN108053840A (zh) * 2017-12-29 2018-05-18 广州势必可赢网络科技有限公司 一种基于pca-bp的情绪识别方法及系统
CN108269574A (zh) * 2017-12-29 2018-07-10 安徽科大讯飞医疗信息技术有限公司 语音信号处理方法及装置、存储介质、电子设备
CN109065069A (zh) * 2018-10-10 2018-12-21 广州市百果园信息技术有限公司 一种音频检测方法、装置、设备及存储介质

Family Cites Families (28)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9020114B2 (en) * 2002-04-29 2015-04-28 Securus Technologies, Inc. Systems and methods for detecting a call anomaly using biometric identification
US20060248019A1 (en) * 2005-04-21 2006-11-02 Anthony Rajakumar Method and system to detect fraud using voice data
US8930261B2 (en) * 2005-04-21 2015-01-06 Verint Americas Inc. Method and system for generating a fraud risk score using telephony channel based audio and non-audio data
US9300790B2 (en) * 2005-06-24 2016-03-29 Securus Technologies, Inc. Multi-party conversation analyzer and logger
US8886663B2 (en) * 2008-09-20 2014-11-11 Securus Technologies, Inc. Multi-party conversation analyzer and logger
CN101770774B (zh) * 2009-12-31 2011-12-07 吉林大学 基于嵌入式的开集说话人识别方法及其系统
CN201698746U (zh) * 2010-06-25 2011-01-05 北京安慧音通科技有限责任公司 携行多功能音频检测仪
CN102572839B (zh) * 2010-12-14 2016-03-02 中国移动通信集团四川有限公司 一种控制语音通信的方法和系统
CN102436806A (zh) * 2011-09-29 2012-05-02 复旦大学 一种基于相似度的音频拷贝检测的方法
CN102820033B (zh) * 2012-08-17 2013-12-04 南京大学 一种声纹识别方法
US20140123166A1 (en) * 2012-10-26 2014-05-01 Tektronix, Inc. Loudness log for recovery of gated loudness measurements and associated analyzer
CN103796183B (zh) * 2012-10-26 2017-08-04 中国移动通信集团上海有限公司 一种垃圾短信识别方法及装置
CN104282303B (zh) * 2013-07-09 2019-03-29 威盛电子股份有限公司 利用声纹识别进行语音辨识的方法及其电子装置
CN103731832A (zh) * 2013-12-26 2014-04-16 黄伟 防电话、短信诈骗的系统和方法
CN105827787B (zh) * 2015-01-04 2019-12-17 中国移动通信集团公司 一种号码标记方法及装置
US10142471B2 (en) * 2015-03-02 2018-11-27 Genesys Telecommunications Laboratories, Inc. System and method for call progress detection
CN104616666B (zh) * 2015-03-03 2018-05-25 广东小天才科技有限公司 一种基于语音分析改善对话沟通效果的方法及装置
US10008209B1 (en) 2015-09-25 2018-06-26 Educational Testing Service Computer-implemented systems and methods for speaker recognition using a neural network
WO2017096473A1 (en) * 2015-12-07 2017-06-15 Syngrafii Inc. Systems and methods for an advanced moderated online event
CN107492382B (zh) * 2016-06-13 2020-12-18 阿里巴巴集团控股有限公司 基于神经网络的声纹信息提取方法及装置
CN105869630B (zh) * 2016-06-27 2019-08-02 上海交通大学 基于深度学习的说话人语音欺骗攻击检测方法及系统
CN106791024A (zh) * 2016-11-30 2017-05-31 广东欧珀移动通信有限公司 语音信息播放方法、装置及终端
US20190052471A1 (en) * 2017-08-10 2019-02-14 Microsoft Technology Licensing, Llc Personalized toxicity shield for multiuser virtual environments
US10574597B2 (en) * 2017-09-18 2020-02-25 Microsoft Technology Licensing, Llc Conversational log replay with voice and debugging information
CN107527617A (zh) * 2017-09-30 2017-12-29 上海应用技术大学 基于声音识别的监控方法、装置及系统
GB2571548A (en) * 2018-03-01 2019-09-04 Sony Interactive Entertainment Inc User interaction monitoring
CN108419091A (zh) 2018-03-02 2018-08-17 北京未来媒体科技股份有限公司 一种基于机器学习的视频内容审核方法及装置
CN108428447B (zh) 2018-06-19 2021-02-02 科大讯飞股份有限公司 一种语音意图识别方法及装置

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20100088088A1 (en) * 2007-01-31 2010-04-08 Gianmario Bollano Customizable method and system for emotional recognition
CN101226743A (zh) * 2007-12-05 2008-07-23 浙江大学 基于中性和情感声纹模型转换的说话人识别方法
CN107610707A (zh) * 2016-12-15 2018-01-19 平安科技(深圳)有限公司 一种声纹识别方法及装置
CN107919137A (zh) * 2017-10-25 2018-04-17 平安普惠企业管理有限公司 远程审批方法、装置、设备及可读存储介质
CN108053840A (zh) * 2017-12-29 2018-05-18 广州势必可赢网络科技有限公司 一种基于pca-bp的情绪识别方法及系统
CN108269574A (zh) * 2017-12-29 2018-07-10 安徽科大讯飞医疗信息技术有限公司 语音信号处理方法及装置、存储介质、电子设备
CN109065069A (zh) * 2018-10-10 2018-12-21 广州市百果园信息技术有限公司 一种音频检测方法、装置、设备及存储介质

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111883139A (zh) * 2020-07-24 2020-11-03 北京字节跳动网络技术有限公司 用于筛选目标语音的方法、装置、设备和介质
CN113782036A (zh) * 2021-09-10 2021-12-10 北京声智科技有限公司 音频质量评估方法、装置、电子设备和存储介质
CN113782036B (zh) * 2021-09-10 2024-05-31 北京声智科技有限公司 音频质量评估方法、装置、电子设备和存储介质
CN114283843A (zh) * 2021-09-27 2022-04-05 腾讯科技(深圳)有限公司 神经网络模型融合监测方法及装置
CN116705031A (zh) * 2023-07-06 2023-09-05 中国电信股份有限公司技术创新中心 违规音频检测方法及装置、电子设备、存储介质

Also Published As

Publication number Publication date
CN109065069B (zh) 2020-09-04
CN109065069A (zh) 2018-12-21
US11948595B2 (en) 2024-04-02
US20220005493A1 (en) 2022-01-06
SG11202103561TA (en) 2021-05-28

Similar Documents

Publication Publication Date Title
WO2020073743A1 (zh) 一种音频检测方法、装置、设备及存储介质
JP7083559B2 (ja) コグニティブ仮想検出器
US11631340B2 (en) Adaptive team training evaluation system and method
US20200228521A1 (en) Authenticating a user device via a monitoring device
US10250641B2 (en) Natural language dialog-based security help agent for network administrator
CN107276982B (zh) 一种异常登录检测方法及装置
US20140222995A1 (en) Methods and System for Monitoring Computer Users
US11869511B2 (en) Using speech mannerisms to validate an integrity of a conference participant
US20160218933A1 (en) Impact analyzer for a computer network
WO2022178942A1 (zh) 情绪识别方法、装置、计算机设备和存储介质
WO2018068396A1 (zh) 语音质量评价方法和装置
CN109346061A (zh) 音频检测方法、装置及存储介质
Li et al. A survey on amazon alexa attack surfaces
US11769520B2 (en) Communication issue detection using evaluation of multiple machine learning models
US12413667B2 (en) Detecting synthetic sounds in call audio
Schönherr et al. Exploring accidental triggers of smart speakers
CN114979549A (zh) 在线会议的隐私保护方法、系统、设备及存储介质
Williams et al. Privacy-preserving occupancy estimation
CN114333802B (zh) 语音处理方法、装置、电子设备及计算机可读存储介质
CN110347797A (zh) 文本信息的侦测方法、系统、设备及存储介质
TWI744036B (zh) 聲音辨識模型訓練方法及系統與電腦可讀取媒體
CN112687293A (zh) 一种基于机器学习及数据挖掘的智能坐席训练方法和系统
CN118173083A (zh) 车辆语音交互功能的测试方法、系统、设备及介质
CN105828135B (zh) 音视频播放系统中的播放控制方法、装置及播放设备
CN117407912A (zh) 访问对话时的隐私保护方法、装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19871266

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19871266

Country of ref document: EP

Kind code of ref document: A1