WO2020103550A1 - 音频信号的评分方法、装置、终端设备及计算机存储介质 - Google Patents
音频信号的评分方法、装置、终端设备及计算机存储介质Info
- Publication number
- WO2020103550A1 WO2020103550A1 PCT/CN2019/106502 CN2019106502W WO2020103550A1 WO 2020103550 A1 WO2020103550 A1 WO 2020103550A1 CN 2019106502 W CN2019106502 W CN 2019106502W WO 2020103550 A1 WO2020103550 A1 WO 2020103550A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- audio signal
- signal
- vocal
- original
- feature
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/60—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for measuring the quality of voice signals
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2210/00—Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
- G10H2210/031—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal
- G10H2210/091—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal for performance evaluation, i.e. judging, grading or scoring the musical qualities or faithfulness of a performance, e.g. with respect to pitch, tempo or other timings of a reference performance
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2250/00—Aspects of algorithms or signal processing methods without intrinsic musical character, yet specifically adapted for or used in electrophonic musical processing
- G10H2250/311—Neural networks for electrophonic musical instruments or musical processing, e.g. for musical recognition or control, automatic composition or improvisation
Definitions
- the present application relates to the field of audio scoring, and in particular to audio signal scoring methods, devices, terminal equipment, and computer storage media.
- the scoring process of the audio signal by the singing platform may include: separating the original vocal audio signal from the vocal and rhythm to obtain the original vocal vocal signal, which includes the vocal signal and the accompaniment signal; extracting the original Sing the pitch and rhythm of the vocal signal; extract the pitch and rhythm of the audio signal; compare the pitch of the original vocal signal with the pitch of the audio signal, and compare the rhythm and audio signal of the original vocal signal To compare the rhythm of; based on the above pitch comparison and rhythm comparison, the score of the user's audio signal is obtained.
- the inventor found that in the above scoring process, by separating the vocal signal and the accompaniment signal in the original vocal audio signal, the original vocal signal obtained usually has a large amount of accompaniment residue, then it will lead to the extraction of the original vocal The pitch and rhythm of the vocal signal are inaccurate. When comparing the pitch and rhythm of the original vocal signal and the audio signal, the result of the comparison is inaccurate, resulting in an inaccurate score of the audio signal.
- the present application provides an audio signal scoring method, device, terminal device, and computer storage medium.
- a method for scoring audio signals including:
- the first original vocal vocal signal which is an audio signal containing accompaniment and original vocal vocals
- the score of the target audio signal is obtained.
- an audio signal scoring device including:
- the separating unit is configured to perform audio separation on the original singing audio signal to obtain a first original singing vocal signal, which is an audio signal containing accompaniment and original singing vocals;
- a suppression unit configured to perform noise suppression processing on the first original vocal signal to obtain a second original vocal signal
- the acquiring unit is configured to perform acquiring the score of the target audio signal based on the difference between the second original vocal vocal signal and the target audio signal.
- a scoring terminal device for audio signals including: a processor and a memory for storing processor executable instructions;
- the processor is configured to: when the one or more programs are executed by the one or more processors, so that the one or more processors achieve the scoring of the audio signal according to any one of the above method.
- a non-transitory computer-readable storage medium When instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal can execute any of the audios described above Signal scoring method.
- an application / computer program product including:
- the first original vocal vocal signal which is an audio signal containing accompaniment and original vocal vocals
- the score of the target audio signal is obtained.
- a second original vocal signal with less accompaniment residuals can be obtained, so that the second original vocal vocal with less residual accompaniment can be obtained
- the signal is more accurate, which can reduce the difference between the second original vocal vocal signal and the target audio signal, thereby making the target audio signal score more accurate.
- Fig. 1 is a flow chart showing a method for scoring audio signals according to an exemplary embodiment
- Fig. 2 is a flow chart of a method for scoring audio signals according to an exemplary embodiment
- Fig. 3 is a block diagram of a logical structure of an audio signal scoring device according to an exemplary embodiment
- Fig. 4 is a structural block diagram of a terminal device according to an exemplary embodiment
- Fig. 5 is a structural block diagram of a terminal device according to an exemplary embodiment.
- Fig. 1 is a flowchart of an audio signal scoring method according to an exemplary embodiment. As shown in Fig. 1, the audio signal scoring method may be applied to a terminal device and includes the following steps.
- step S101 audio separation is performed on the original vocal audio signal to obtain a first original vocal vocal signal, and the original vocal audio signal is an audio signal containing accompaniment and original vocal vocals.
- step S102 the first original vocal signal is subjected to noise suppression processing to obtain a second original vocal signal.
- step S103 based on the difference between the second original vocal signal and the target audio signal, a score of the target audio signal is obtained.
- performing audio separation on the original singing audio signal to obtain the first original singing vocal signal including:
- the original vocal audio signal is input into a deep neural network model to obtain the first original vocal vocal signal.
- the deep neural network model is used to separate the accompaniment and the original vocal in the original vocal audio signal.
- obtaining the score of the target audio signal based on the difference between the second original vocal signal and the target audio signal includes:
- a first high-order statistical feature of each first feature is obtained, and each first high-order statistical feature is that each first feature is not less than a second-order Statistics;
- each second high-order statistical feature being a statistic of no less than second order for each second feature;
- the difference between each first higher-order statistical feature and the corresponding second higher-order statistical feature is used to obtain the score of the target audio signal.
- the method before acquiring the first high-order statistical feature of each first feature based on at least one first feature of the second original vocal signal, the method further includes:
- the pitch of the target audio signal is adjusted up by an octave.
- obtaining the score of the target audio signal based on the difference between each first higher-order statistical feature and the corresponding second higher-order statistical feature includes:
- the score of the target audio signal is obtained.
- the method before obtaining the score of the target audio signal based on the difference between the second original vocal signal and the target audio signal, the method further includes:
- the original audio signal is an audio signal containing accompaniment and singer's human voice.
- Fig. 2 is a flowchart of a method for scoring an audio signal according to an exemplary embodiment. As shown in Fig. 2, the method for scoring an audio signal is used in a terminal device and includes the following steps.
- step S201 the terminal device obtains the original singing audio signal based on the song name.
- the song name is the name of the song that the singer will sing.
- the terminal device can provide the original singing audio signal or accompaniment audio signal of the song, record the song, search the song, and score the song through the K song interface.
- the terminal device can access The song database obtains the original vocal audio signal from the song database, wherein the original vocal audio signal is an audio signal containing accompaniment and original vocal vocals, that is, an audio signal including the original singer's voice of the song.
- a large number of original singing audio signals can be stored in the song database.
- the song database can be a song database built in the terminal device or a song database on the network.
- the terminal device can search the song name in the song database to obtain the original singing audio signal corresponding to the song name.
- the user can input the song name on the K song interface to trigger the song search function of the terminal device, so that the terminal The device can access the song database and obtain the original singing audio signal corresponding to the song name from the song database according to the song name.
- the audio signal scoring method proposed in the embodiment of the present application is based on a vocal signal scoring method, and the original singing audio signal acquired in step S201 includes not only the vocal signal but also the accompaniment signal of the song Then, you need to process the original vocal audio signal to obtain the original vocal signal, see steps S202 and S203.
- step S202 the terminal device performs audio separation on the original vocal audio signal to obtain a first original vocal vocal signal, and the original vocal audio signal is an audio signal including accompaniment and original vocal vocals.
- the terminal device can separate the vocal and accompaniment of the original singing audio signal through the deep neural network model.
- the separated vocal signal may contain a large amount of accompaniment residual signals; then, the terminal device can use post-processing of deep learning Technology, which maps the separated vocal signal to a pure vocal signal to eliminate part of the accompaniment residual signal, the pure vocal signal is the original vocal signal, considering that the original vocal signal may also carry a small amount Residual accompaniment, then, in order to further eliminate the residual accompaniment in the original vocal signal, the terminal device also needs to process the original vocal signal, see step S203.
- step S203 the terminal device performs noise suppression processing on the first original vocal signal to obtain a second original vocal signal.
- the terminal device may use a smoothing technique to eliminate the residual accompaniment in the first original vocal signal to obtain the second original vocal signal.
- this step S203 may be a step in which the terminal device performs first-order recursive smoothing filtering on the first original vocal signal.
- the terminal device may use a first-order recursive smoothing filtering algorithm
- the first original vocal signal is filtered to eliminate glitches in the first original vocal signal, thereby obtaining a smoother second original vocal signal, because the obtained second original vocal signal is more Smooth, so it can not only achieve the purpose of eliminating the residual accompaniment in the first original vocal signal, but also ensure the continuity of the second original vocal signal.
- the terminal device may also use other filtering algorithms to filter the first original vocal signal.
- the embodiment of the present application does not specifically limit the other filtering algorithms.
- step S204 the terminal device extracts at least one first feature of the second original vocal signal.
- the at least one first characteristic may include pitch, rhythm, timbre (spectral centroid, signal bandwidth, etc.) and loudness.
- the terminal device may use a fundamental frequency extraction method to extract at least one first feature of the second original vocal signal.
- the fundamental frequency extraction method may be any feature extraction algorithm such as pyin, which is not done in the embodiments of the present application limited.
- the terminal device may use a median filtering method and other related smoothing techniques to solve the problem of half frequency and frequency multiplication when extracting the pitch.
- step S205 the terminal device obtains the first high-order statistical feature of each first feature based on at least one first feature of the second original vocal signal, and each first high-order statistical feature is each One feature is not less than second-order statistics.
- the statistic not lower than the second order may be an average value or a variance.
- This step S205 may be specifically implemented by using the following process: the terminal device may perform regularization on the pitch, rhythm, timbre and loudness characteristics of the second original vocal signal, and calculate each first characteristic based on the regularization result Corresponding first high-order statistical feature.
- this step S205 will be described as follows: the terminal device extracts all sounds from the second original vocal signal High regularity together, regularize all rhythms extracted from the target audio signal, then calculate the average of all pitches, and the average of all rhythms, then the average of all pitches is also the pitch Corresponding to the first higher-order statistical feature, the average value of all rhythms is also the first higher-order statistical feature corresponding to the rhythm.
- step S204 and step S205 are processes for the terminal device to acquire at least one first high-order statistical characteristic of the second original vocal vocal signal
- steps S201 to S205 are for the terminal device to obtain original songs Process of at least one second higher-order statistical feature of audio
- the at least one first higher-order statistical feature obtained by the terminal device through the processes of steps S201 to S205 can be used as a standard feature to compare with at least one first Compare the two higher-order statistical features and score the target audio signal based on the result of the comparison (see steps S211 and S212), thereby avoiding the invitation of professional musicians to mark features such as pitch to provide the target audio signal Standard features, which can reduce the labor cost of marking standard features.
- step S206 the terminal device obtains a singer audio signal, and the singer audio signal is an audio signal containing accompaniment and a singer's human voice.
- the terminal device can obtain the singer audio signal by recording, for example, when the singer clicks "record" on the K-song interface, the terminal device records the sound currently received by the terminal device to obtain the singer Audio signal, the sound received by the terminal device may include accompaniment sound and singer's vocals; when the terminal device starts playing the accompaniment of the song, the terminal device starts recording the sound currently received by the terminal device to obtain the singer audio signal.
- the audio signal scoring method proposed in the embodiment of the present application is based on the vocal signal scoring method, and the singer audio signal obtained in step S206 is an audio signal including accompaniment and singer, then, The audio signal of the singer is processed to obtain the vocal signal of the singer, see step S207.
- step S207 the terminal device obtains a target audio signal based on the singer audio signal.
- the target audio signal is an audio signal of the singer's human voice in the singer's audio signal
- the terminal device may implement this step S207 through the processes shown in steps S207A and S207B as follows:
- step S207A the terminal device inputs the singer audio signal into a deep neural network model to obtain a singer vocal signal, and the deep neural network model is used to separate the accompaniment and singer vocal in the singer audio signal .
- step S207A the terminal device can separate the rhythm and vocals of the singer's audio signal in the same way as in step S202, the terminal device can separate the rhythm and vocals of the original singing audio signal. In this manner, the embodiments of the present application will not repeat them here.
- step S207B the terminal device performs noise suppression processing on the singer's vocal signal to obtain a target audio signal.
- step S207B the manner in which the terminal device performs noise suppression processing on the singer's audio signal may be the same as the manner in which the terminal device performs noise suppression processing on the first original vocal signal in step S203 Therefore, the embodiments of the present application will not be repeated here.
- step S208 the terminal device performs vocal detection on the singer's audio signal to obtain the start time and end time of the target audio signal.
- the terminal device can extract at least one vocal feature from the singer's audio signal, the vocal feature is used to distinguish human voices from the sounds of creatures or objects other than humans, and the vocal characteristics can include energy and fundamental frequency As well as MPEG7 timbre, etc., the terminal device can use the vocal feature to perform vocal detection on the singer's audio signal, for example, when the terminal device extracts the value of each vocal feature from the singer's audio signal, the If the range is set, or if the value of some of the extracted human voice features meets the preset range, the singer audio signal is considered to contain human voice, otherwise, the singer audio signal is considered not to contain human voice.
- the terminal device detects the time when the vocal in the singer's audio signal starts to appear as the start time of the target audio signal, when the terminal device After determining the start time, continue to detect the singer's audio signal, and will continue to detect the time when the vocal disappears in the singer's audio signal, as the end time of the target audio signal, the terminal device can be based on the start Time and end time, to perform the steps of subsequent feature extraction.
- step S209 the terminal device extracts at least one second feature of the target audio signal.
- step S209 the terminal device extracts at least one second feature of the target audio signal
- step S204 the terminal device extracts at least one first feature of the second original vocal signal
- step S210 based on at least one second feature of the target audio signal, the terminal device acquires second high-order statistical features of each second feature, and each second high-order statistical feature is that each second feature is not low Statistics for the second order.
- step S210 the terminal device obtains the second higher-order statistical feature of each second feature
- step S205 the terminal device obtains the first higher-order statistical feature of each first feature
- the terminal device before obtaining the second high-order statistical feature of the pitch of the target audio signal, the terminal device also needs to determine whether these pitches need to be lowered or adjusted up, for example, when the average pitch of the target audio signal When the average pitch of the second original vocal signal is an octave, the pitch of the target audio signal is reduced by an octave; when the average pitch of the target audio signal is lower than the second original vocal vocal When the average pitch of the signal is one octave, increase the pitch of the target audio signal by one octave, when the average pitch of the target audio signal is the same as the average pitch of the second original vocal signal, or , When the difference between the two average pitches does not exceed one octave, the terminal device can directly perform this step S210 to obtain the second high-order statistical characteristic of the pitch of the target audio signal, or directly The average value of the pitch of the signal serves as the second high-order statistical characteristic of the pitch.
- the terminal device taking the terminal device extracting 10 pitches from the target audio signal, for example, the average pitch of the second original vocal signal is B2, when the average of the 10 pitches is B3 Tone adjustment, the pitch of the target audio signal is one octave higher than the average pitch of the second original vocal signal, so the terminal device reduces the 10 pitches by one octave; when the 10 When the average value of the pitches is the B1 key, the pitch of the target audio signal is lower than the average value of the pitch of the second original vocal signal by an octave, so the terminal device raises the 10 pitches One octave; then, the terminal device may acquire the second high-order statistical characteristic of the pitch of the target audio signal based on the 10 pitches after the decrease or increase.
- the pitch When an ordinary singer sings a song with a higher pitch, the pitch may not reach the pitch of the original sing. When singing a song with a lower pitch, the pitch may be higher than the original sing, resulting in ordinary singing. Compared with the original vocal signal, the target audio signal of the singer is too different in pitch, which makes the average singer's score on the pitch feature lower, which may affect the target audio signal of the ordinary singer. Comprehensive rating reduces the interest of the ordinary singer.
- this application reduces the pitch of the target audio signal with a higher pitch by one octave, or increases the pitch of the target audio signal with a lower pitch by an octave, which will make the goal of ordinary singers
- the difference between the pitch of the audio signal and the pitch of the second original vocal signal is not so large, which can improve the score of the ordinary singer in terms of pitch, so that the general score of the ordinary singer will not be too much Low, can improve the K-song confidence of ordinary singers, thereby ensuring the interest of K-songs of ordinary singers.
- steps S209 and S210 are processes for the terminal device to obtain at least one second high-order statistical characteristic of the target audio signal
- the processes of steps S206 to S210 are processes for the terminal device to obtain at least one singer's audio signal A process of second-order high-level statistical features.
- step S211 the terminal device maps the difference between each first higher-order statistical feature and the corresponding second higher-order statistical feature into a scoring interval to obtain a score for each second higher-order statistical feature.
- the scoring interval may be a scoring interval of [0,100], and of course, the scoring interval may also be other scoring intervals.
- the terminal device may subtract each first higher-order statistical feature from the second higher-order statistical feature corresponding to the first higher-order statistical feature to obtain each first higher-order statistical feature and the The difference value of the second higher-order statistical feature, and taking the absolute value of the difference value, and then mapping the absolute value into the scoring interval.
- the variance is used as the first high-order statistical feature and the second high-order statistical feature, and the first high-order statistical feature is 0.5, and the second high-order statistical feature corresponding to the first high-order statistical feature is 0.48.
- 0.02; mapping 0.02 to the scoring interval of [0,100] can make the 0.02 mapping to 98, then, 98 points This is the score for this second higher-order statistic.
- the score of each second higher-order statistical feature can be obtained, and when the difference between the first higher-order statistical feature and its corresponding second higher-order statistical feature is greater, the lower the score, when the first higher-order statistical feature is The smaller the difference between the statistical feature and its corresponding second-order statistical feature, the higher the score.
- the embodiment of the present application does not specifically limit the scoring interval and the method of mapping the difference value of each second higher-order feature to the scoring interval.
- step S212 the terminal device obtains the score of the target audio signal based on the weight of each second higher-order statistical feature and the score of each second higher-order statistical feature.
- the scores of the second high-order statistical features may be fused to obtain the score of the target audio signal.
- the terminal device may set a weight for each second higher-order statistical feature, and the terminal device may set a weight for each second higher-order statistical feature before acquiring the score for each second higher-order statistical feature
- the higher-order statistical features are set with weights, and after obtaining the score of each second higher-order statistical feature, the weight may be set for each second higher-order statistical feature.
- each second higher-order statistical feature The time for setting the weight is not specifically limited.
- This step S212 can be implemented by the following process: the terminal device multiplies the score of each second higher-order statistical feature and the corresponding weight to obtain the product corresponding to each second higher-order statistical feature, and then obtains each The sum of the products is used as the score of the target audio signal.
- the sum value can also be mapped into a scoring interval. The higher the sum value, the higher the score to obtain a score that meets the scoring criteria.
- the scoring interval may be [0,100], so that the scoring can meet the scoring criteria on a 100-point scale.
- the terminal device by performing noise suppression processing on the first original vocal signal separated from the original vocal audio signal, the terminal device can obtain the second original vocal signal with less accompaniment residual, thereby enabling the terminal device Features such as pitch and rhythm extracted from the second original vocal signal are more accurate.
- the accuracy of the comparison result can be ensured.
- the terminal device uses the accurate comparison result to score the target audio signal, the accuracy of the score of the target audio signal can be ensured.
- the terminal device uses the difference between the high-order statistical characteristics of the second original vocal signal and the target audio signal to score the target audio signal, which improves the accuracy and robustness of the score.
- the terminal device determines the start time and end time of the target audio signal through human voice detection, thereby avoiding manually labeling the start time and end time of the target audio signal, thereby reducing labor costs.
- the terminal device reduces or increases the pitch of the target audio signal according to the difference in average pitch between the target audio signal of the singer and the second original vocal signal, so as to reduce the target audio signal and the second The difference in pitch between the vocal signals of the two original singers, which can improve the score of the ordinary singer in terms of pitch, so that the target audio signal of the singer with a lower or higher pitch will not be too low , Can increase the K-song confidence of singers with lower or higher pitch, thereby ensuring the interest of K-song of singers with lower or higher pitch.
- the terminal device adopts a related smoothing technique such as a median filter method to solve the problem of half frequency and frequency multiplication when extracting the pitch.
- Fig. 3 is a block diagram of a logical structure of an audio signal scoring device according to an exemplary embodiment.
- the device includes a separation unit 301, a suppression unit 302 and an acquisition unit 303.
- the separation unit 301 is connected to the suppression unit 302, and is configured to perform audio separation on the original vocal audio signal to obtain a first original vocal vocal signal.
- the original vocal audio signal is audio including accompaniment and original vocal vocals signal;
- the suppression unit 302 is connected to the acquisition unit 303, and is configured to perform noise suppression processing on the first original vocal signal to obtain a second original vocal signal;
- the acquiring unit 303 is configured to perform acquiring the score of the target audio signal based on the difference between the second original vocal vocal signal and the target audio signal.
- the separation unit 301 is further configured to perform input of the original singing audio signal into the deep neural network model to obtain the first original singing vocal signal, and the deep neural network model is used to separate the original singing audio Accompaniment and original singing in the signal.
- the obtaining unit 303 includes:
- the first obtaining subunit is configured to perform at least one first feature based on the second original vocal signal to obtain the first high-order statistical feature of each first feature, each first high-order statistical feature is A statistic whose first characteristic is not lower than second order;
- the second obtaining subunit is configured to perform at least one second feature based on the target audio signal to obtain second high-order statistical features of each second feature, each second high-order statistical feature being each second feature No less than second-order statistics;
- the third acquiring subunit is configured to perform acquiring the score of the target audio signal based on the difference between each first higher-order statistical feature and the corresponding second higher-order statistical feature.
- the audio signal scoring device further includes:
- a reduction unit configured to perform a reduction of the pitch of the target audio signal by an octave when the average pitch of the target audio signal is one octave higher than the average pitch of the second original vocal signal ;
- a height-adjusting unit configured to perform an increase in the pitch of the target audio signal by an octave when the average pitch of the target audio signal is lower than the average pitch of the second original vocal signal by an octave .
- the third obtaining subunit is further configured to perform mapping of the difference between each first higher-order statistical feature and the corresponding second higher-order statistical feature into a scoring interval to obtain each second higher-order statistical feature Scores of statistical features;
- the score of the target audio signal is obtained.
- the audio signal scoring device further includes a detection unit configured to perform vocal detection on the singer audio signal to obtain the start time and end time of the target audio signal, the singer audio signal includes Audio signals of accompaniment and singer's vocals.
- Fig. 4 is a structural block diagram of a terminal device according to an exemplary embodiment.
- the terminal device 400 may be: a smartphone, a tablet computer, an MP3 player (Moving Pictures Experts Group Audio Audio Layer III, motion picture expert compression standard audio level 3), MP4 (Moving Pictures Experts Group Audio Audio Layer IV, motion picture expert compression standard Audio level 4) Player, laptop or desktop computer.
- the terminal device 400 may also be called other names such as user equipment, portable terminal equipment, laptop terminal equipment, and desktop terminal equipment.
- the terminal device 400 includes a processor 401 and a memory 402.
- the processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, and so on.
- the processor 401 may adopt at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), programmable logic array (Programmable Logic Array, PLA) achieve.
- the processor 401 may also include a main processor and a coprocessor.
- the main processor is a processor for processing data in a wake-up state, also called a central processing unit (Central Processing Unit, CPU); the coprocessor is A low-power processor for processing data in the standby state.
- CPU Central Processing Unit
- the processor 401 may be integrated with a graphics processor (Graphics Processing Unit, GPU), and the GPU is used to render and draw the content required to be displayed on the display screen.
- the processor 401 may further include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
- AI artificial intelligence
- the memory 402 may include one or more computer-readable storage media, which may be non-transitory.
- the memory 402 may also include high-speed random access memory, as well as non-volatile memory, such as one or more magnetic disk storage devices, flash memory storage devices.
- the non-transitory computer-readable storage medium in the memory 402 is used to store at least one instruction for execution by the processor 401 to implement the audio signal provided by the method embodiment in the present application Scoring method.
- the terminal device 400 may optionally further include: a peripheral device interface 403 and at least one peripheral device.
- the processor 401, the memory 402 and the peripheral device interface 403 may be connected by a bus or a signal line.
- Each peripheral device may be connected to the peripheral device interface 403 through a bus, a signal line, or a circuit board.
- the peripheral device includes at least one of a radio frequency circuit 404, a touch display screen 405, a camera 406, an audio circuit 407, a positioning component 408, and a power supply 409.
- the peripheral device interface 403 may be used to connect at least one peripheral device related to Input / Output (I / O) to the processor 401 and the memory 402.
- the processor 401, the memory 402, and the peripheral device interface 403 are integrated on the same chip or circuit board; in some other embodiments, any one of the processor 401, the memory 402, and the peripheral device interface 403 or Both can be implemented on a separate chip or circuit board, which is not limited in this embodiment.
- the radio frequency circuit 404 is used to receive and transmit radio frequency (Radio Frequency) signals, also called electromagnetic signals.
- the radio frequency circuit 404 communicates with a communication network and other communication devices through electromagnetic signals.
- the radio frequency circuit 404 converts the electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal.
- the radio frequency circuit 404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on.
- the radio frequency circuit 404 can communicate with other terminal devices through at least one wireless communication protocol.
- the wireless communication protocol includes but is not limited to: metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G and 5G), wireless local area networks and / or wireless fidelity (WiFi) networks.
- the radio frequency circuit 404 may further include a circuit related to short-range wireless communication (Near Field Communication, NFC), which is not limited in this application.
- NFC Near Field Communication
- the display screen 405 is used to display a user interface (User Interface, UI).
- the UI may include graphics, text, icons, video, and any combination thereof.
- the display screen 405 also has the ability to collect touch signals on or above the surface of the display screen 405.
- the touch signal can be input to the processor 401 as a control signal for processing.
- the display screen 405 can also be used to provide virtual buttons and / or virtual keyboards, also called soft buttons and / or soft keyboards.
- the display screen 405 may be one, and the front panel of the terminal device 400 is provided; in other embodiments, the display screen 405 may be at least two, respectively provided on different surfaces of the terminal device 400 or in a folded design.
- the display screen 405 may be a flexible display screen, which is provided on the curved surface or folding surface of the terminal device 400. Even, the display screen 405 can also be set as a non-rectangular irregular figure, that is, a special-shaped screen.
- the display screen 405 may be made of liquid crystal display (Liquid Crystal) (LCD), organic light-emitting diode (Organic Light-Emitting Diode, OLED) and other materials.
- LCD liquid crystal display
- OLED Organic Light-Emitting Diode
- the camera component 406 is used to collect images or videos.
- the camera assembly 406 includes a front camera and a rear camera.
- the front camera is set on the front panel of the terminal device, and the rear camera is set on the back of the terminal device.
- the camera assembly 406 may also include a flash.
- the flash can be a single-color flash or a dual-color flash. Dual color temperature flash refers to the combination of warm light flash and cold light flash, which can be used for light compensation at different color temperatures.
- the audio circuit 407 may include a microphone and a speaker.
- the microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 401 for processing, or input them to the radio frequency circuit 404 to implement voice communication.
- the microphone can also be an array microphone or an omnidirectional acquisition microphone.
- the speaker is used to convert the electrical signal from the processor 401 or the radio frequency circuit 404 into sound waves.
- the speaker can be a traditional thin-film speaker or a piezoelectric ceramic speaker.
- the speaker When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible by humans, but also convert electrical signals into sound waves inaudible to humans for distance measurement and other purposes.
- the audio circuit 407 may also include a headphone jack.
- the positioning component 408 is used to locate the current geographic location of the terminal device 400 to implement navigation or location-based services (Location Based Services, LBS).
- the positioning component 508 may be a positioning component based on the Global Positioning System (GPS) of the United States, the Beidou system of China, the Grenas system of Russia, or the Galileo system of the European Union.
- GPS Global Positioning System
- the power supply 409 is used to supply power to various components in the terminal device 400.
- the power source 409 may be alternating current, direct current, disposable batteries, or rechargeable batteries.
- the rechargeable battery may support wired charging or wireless charging.
- the rechargeable battery can also be used to support fast charging technology.
- the terminal device 400 further includes one or more sensors 410.
- the one or more sensors 410 include, but are not limited to: an acceleration sensor 411, a gyro sensor 412, a pressure sensor 413, a fingerprint sensor 414, an optical sensor 415, and a proximity sensor 416.
- the acceleration sensor 411 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal device 400.
- the acceleration sensor 411 can be used to detect components of gravity acceleration on three coordinate axes.
- the processor 401 may control the touch screen 405 to display the user interface in a landscape view or a portrait view according to the gravity acceleration signal collected by the acceleration sensor 411.
- the acceleration sensor 411 can also be used for game or user movement data collection.
- the gyro sensor 412 can detect the body direction and the rotation angle of the terminal device 400, and the gyro sensor 412 can cooperate with the acceleration sensor 411 to collect a 3D action of the user on the terminal device 400.
- the processor 401 can realize the following functions according to the data collected by the gyro sensor 412: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
- the pressure sensor 413 may be disposed on the side frame of the terminal device 400 and / or the lower layer of the touch display 405.
- the pressure sensor 413 can detect the user's grip signal to the terminal device 400, and the processor 401 can perform left-right hand recognition or shortcut operation according to the grip signal collected by the pressure sensor 413.
- the processor 401 controls the operability control on the UI interface according to the user's pressure operation on the touch display screen 405.
- the operability control includes at least one of a button control, a scroll bar control, an icon control, and a menu control.
- the fingerprint sensor 414 is used to collect the user's fingerprint, and the processor 401 identifies the user's identity according to the fingerprint collected by the fingerprint sensor 414, or the fingerprint sensor 414 identifies the user's identity according to the collected fingerprint. When the user's identity is recognized as a trusted identity, the processor 401 authorizes the user to perform related sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, paying, and changing settings.
- the fingerprint sensor 414 may be provided on the front, back, or side of the terminal device 400. When a physical button or manufacturer logo is provided on the terminal device 400, the fingerprint sensor 414 may be integrated with the physical button or manufacturer logo.
- the optical sensor 415 is used to collect the ambient light intensity.
- the processor 401 can control the display brightness of the touch display 405 according to the ambient light intensity collected by the optical sensor 415. Specifically, when the ambient light intensity is high, the display brightness of the touch display 405 is increased; when the ambient light intensity is low, the display brightness of the touch display 405 is decreased.
- the processor 401 may also dynamically adjust the shooting parameters of the camera assembly 406 according to the ambient light intensity collected by the optical sensor 415.
- the proximity sensor 416 also called a distance sensor, is usually provided on the front panel of the terminal device 400.
- the proximity sensor 416 is used to collect the distance between the user and the front of the terminal device 400.
- the processor 401 controls the touch display 405 to switch from the bright screen state to the breathing state; when the proximity sensor 416 When it is detected that the distance between the user and the front of the terminal device 400 gradually becomes larger, the processor 401 controls the touch display screen 405 to switch from the breath-hold state to the bright-screen state.
- FIG. 5 is a structural block diagram of a terminal device according to an exemplary embodiment.
- the terminal device 500 may have a relatively large difference due to different configurations or performance, and may include one or more processors (Central Processing Units, CPU ) 501 and one or more memories 502, wherein at least one instruction is stored in the memory 502, and the at least one instruction is loaded and executed by the processor 501 to implement the following process:
- processors Central Processing Units, CPU
- the original vocal audio signal is an audio signal containing accompaniment and original vocal vocals; perform noise suppression on the first original vocal vocal
- a second original vocal signal is obtained; based on the difference between the second original vocal signal and the target audio signal, a score of the target audio signal is obtained.
- the processor 501 is specifically configured to: input the original singing audio signal into a deep neural network model to obtain the first original singing vocal signal, and the deep neural network model is used to separate the Accompaniment and original singing in the original singing audio signal.
- the processor 501 is specifically configured to: based on at least one first feature of the second original singing vocal signal, acquire a first high-order statistical feature of each first feature, each first high-order statistical feature
- the statistical feature is a statistic of no less than second order for each first feature; based on at least one second feature of the target audio signal, a second higher order statistical feature for each second feature is acquired, each second higher order
- the statistical feature is a statistic of each second feature not lower than the second order; based on the difference between each first high order statistical feature and the corresponding second high order statistical feature, the score of the target audio signal is obtained.
- the processor 501 before acquiring the second high-order statistical feature of each second feature based on the at least one second feature of the target audio signal, is specifically configured to: When the average pitch of the target audio signal is one octave higher than the average pitch of the second original vocal signal, the pitch of the target audio signal is reduced by one octave; when the average pitch of the target audio signal When the height is lower than the average pitch of the second original vocal signal by an octave, the pitch of the target audio signal is adjusted up by an octave.
- the processor 501 is specifically configured to: map the difference between each first higher-order statistical feature and the corresponding second higher-order statistical feature into a scoring interval to obtain each second higher-order statistical feature The score of the target audio signal based on the weight of each second higher-order statistical feature and the score of each second higher-order statistical feature.
- the processor 501 Before acquiring the score of the target audio signal based on the difference between the second original vocal vocal signal and the target audio signal, the processor 501 is specifically configured to: perform vocal detection on the singer's audio signal To obtain the start time and end time of the target audio signal, and the singer audio signal is an audio signal containing accompaniment and the vocal of the singer.
- the terminal device may also have components such as a wired or wireless network interface, a keyboard, and an input-output interface for input and output.
- the terminal device may also include other components for implementing device functions, which are not repeated here.
- a non-transitory computer-readable storage medium is also provided, for example, a memory including instructions that can be executed by a processor in the terminal to complete the audio signal scoring method in the foregoing embodiments.
- the computer-readable storage medium may be read-only memory (Read-Only Memory, ROM), random-access memory (Random Access Memory, RAM), read-only compact disc (Compact Disc Read-Only Memory, CD-ROM), Magnetic tapes, floppy disks, optical data storage devices, etc.
Landscapes
- Engineering & Computer Science (AREA)
- Human Computer Interaction (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Quality & Reliability (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Auxiliary Devices For Music (AREA)
- Reverberation, Karaoke And Other Acoustics (AREA)
Abstract
提供一种音频信号的评分方法、装置、终端设备及计算机存储介质。该方法包括对原唱音频信号进行音频分离,得到第一原唱人声信号,原唱音频信号为包含伴奏和原唱的人声的音频信号(S101);对第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号(S102);基于第二原唱人声信号与目标音频信号之间的差异,获取目标音频信号的得分(S103)。通过对从原唱信号中分离的第一原唱人声信号进行噪声抑制处理,可以获得伴奏残留较少的第二原唱人声信号,进而使得目标音频信号的得分比较准确。
Description
相关申请的交叉引用
本申请要求在2018年11月19日提交中国专利局、申请号为201811376670.8、发明名称为“音频信号的评分方法、装置、电子设备及计算机存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及音频评分领域,尤其涉及音频信号的评分方法、装置、终端设备及计算机存储介质。
目前,在人们的生活中出现了越来越多的唱歌平台,使得人们可以在这些唱歌平台上唱歌,丰富人们的业余生活,很多唱歌平台上都设有唱歌评分功能,可以通过对用户的音频信号进行采集,并基于采集到的音频信号进行评分,使得用户可以知道自己的唱歌水平,提高了唱歌平台的趣味性,从而使得唱歌平台可以吸引大量的用户。
相关技术中,唱歌平台对音频信号的评分过程可以包括:将原唱音频信号进行人声和节奏的分离,得到原唱人声信号,该原唱音频信号包括人声信号和伴奏信号;提取原唱人声信号的音高和节奏;提取音频信号的音高和节奏;将原唱人声信号的音高和音频信号的音高进行对比,并且,将原唱人声信号的节奏和音频信号的节奏进行对比;基于上述音高的对比以及节奏的对比,得到该用户的音频信号的得分。
发明人发现在上述评分过程中,通过分离原唱音频信号内的人声信号和伴奏信号,得到的原唱人声信号中通常会带有大量的伴奏残留,那么,就会导致提取的原唱人声信号的音高和节奏不准确,进而当对该原唱人声信号与音频信号进行音高的对比以及节奏的对比时,会使得对比结果不准确,从而导致音频信号的得分不准确。
发明内容
为克服相关技术中存在的问题,本申请提供一种音频信号的评分方法、装置、终端设备及计算机存储介质。
根据本申请实施例的第一方面,提供一种音频信号的评分方法,包括:
对原唱音频信号进行音频分离,得到第一原唱人声信号,该原唱音频信号为包含伴奏和原唱的人声的音频信号;
对该第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号;
基于该第二原唱人声信号与目标音频信号之间的差异,获取该目标音频信号的得分。
根据本申请实施例的第二方面,提供一种音频信号的评分装置,包括:
分离单元,被配置为执行对原唱音频信号进行音频分离,得到第一原唱人声信号,该原唱音频信号为包含伴奏和原唱的人声的音频信号;
抑制单元,被配置为执行对该第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号;
获取单元,被配置为执行基于该第二原唱人声信号与目标音频信号之间的差异,获取该目标音频信号的得分。
根据本申请实施例的第三方面,提供一种音频信号的评分终端设备,包括:处理器、用于存储处理器可执行指令的存储器;
其中,所述处理器被配置为:当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现上述任意一项所述的音频信号的评分方法。
根据本申请实施例的第四方面,提供一种非临时性计算机可读存储介质,当该存储介质中的指令由移动终端的处理器执行时,使得移动终端能够执行上述任意一项所述音频信号的评分方法。
根据本申请实施例的第五方面,提供一种应用程序/计算机程序产品,包括:
对原唱音频信号进行音频分离,得到第一原唱人声信号,该原唱音频信号为包含伴奏和原唱的人声的音频信号;
对该第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号;
基于该第二原唱人声信号与目标音频信号之间的差异,获取该目标音频信号的得分。
通过对从原唱信号中分离的第一原唱人声信号进行噪声抑制处理,可以 获得伴奏残留较少的第二原唱人声信号,从而使得该伴奏残留较少的第二原唱人声信号比较准确,可以降低第二原唱人声信号与目标音频信号之间的差异,进而使得目标音频信号的得分比较准确。
图1是根据一示例性实施例示出的一种音频信号的评分方法的流程图;
图2是根据一示例性实施例示出的一种音频信号的评分方法的流程图;
图3是根据一示例性实施例示出的一种音频信号的评分装置的逻辑结构框图;
图4是根据一示例性实施例示出的一种终端设备的结构框图;
图5是根据一示例性实施例示出的一种终端设备的结构框图。
图1是根据一示例性实施例示出的一种音频信号的评分方法的流程图,如图1所示,该音频信号的评分方法可以应用于终端设备中,包括以下步骤。
在步骤S101中,对原唱音频信号进行音频分离,得到第一原唱人声信号,该原唱音频信号为包含伴奏和原唱的人声的音频信号。
在步骤S102中,对该第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号。
在步骤S103中,基于该第二原唱人声信号与目标音频信号之间的差异,获取该目标音频信号的得分。
可选地,对原唱音频信号进行音频分离,得到第一原唱人声信号,包括:
将该原唱音频信号输入至深度神经网络模型中,得到该第一原唱人声信号,该深度神经网络模型用于分离该原唱音频信号中的伴奏和原唱。
可选地,基于该第二原唱人声信号与目标音频信号之间的差异,获取该目标音频信号的得分,包括:
基于该第二原唱人声信号的至少一个第一特征,获取每个第一特征的第一高阶统计特征,每个第一高阶统计特征为每个第一特征不低于二阶的统计量;
基于该目标音频信号的至少一个第二特征,获取每个第二特征的第二高阶统计特征,每个第二高阶统计特征为每个第二特征不低于二阶的统计量; 基于每个第一高阶统计特征与对应的第二高阶统计特征的差异,获取该目标音频信号的得分。
可选地,基于该第二原唱人声信号的至少一个第一特征,获取每个第一特征的第一高阶统计特征之前,还包括:
当该目标音频信号的平均音高高于该第二原唱人声信号的平均音高一个八度时,将该目标音频信号的音高降低一个八度;
当该目标音频信号的平均音高低于该第二原唱人声信号的平均音高一个八度时,将该目标音频信号的音高调高一个八度。
可选地,基于每个第一高阶统计特征与对应的第二高阶统计特征的差异,获取该目标音频信号的得分,包括:
将每个第一高阶统计特征与对应的第二高阶统计特征的差异映射到评分区间内,得到每个第二高阶统计特征的得分;
基于每个第二高阶统计特征的权重以及该每个第二高阶统计特征的得分,获取目标音频信号的得分。
可选地,基于该第二原唱人声信号与目标音频信号之间的差异,获取该目标音频信号的得分之前,还包括:
对演唱者音频信号进行人声检测,得到该目标音频信号的起始时间和结束时间,该原音频信号为包含伴奏和演唱者的人声的音频信号。
图2是根据一示例性实施例示出的一种音频信号的评分方法的流程图,如图2所示,该音频信号的评分方法用于终端设备中,包括以下步骤。
在步骤S201中,终端设备基于歌曲名称,获取原唱音频信号。
该歌曲名称为演唱者将要歌唱的歌曲名称,该终端设备可以通过K歌界面来提供播放歌曲的原唱音频信号或伴奏音频信号、录制歌曲、搜索歌曲以及歌曲评分等功能,该终端设备可以访问歌曲数据库,从歌曲数据库中获取原唱音频信号,其中,原唱音频信号为包含伴奏和原唱的人声的音频信号,也即是,包括歌曲原唱者声音的音频信号。该歌曲数据库中可以存储有大量的原唱音频信号,该歌曲数据库可以是该终端设备自带的歌曲数据库,还可以是网络上的歌曲数据库。
该终端设备可以在歌曲数据库中搜索歌曲名称,从而获取该歌曲名称对应的原唱音频信号,例如,用户可以在该K歌界面上输入歌曲名称,触发该终端设备的搜索歌曲功能,使得该终端设备可以访问歌曲数据库,并根据该 歌曲名称从歌曲数据库中获取与该歌曲名称对应的原唱音频信号。
需要说明的是,本申请实施例提出的音频信号的评分方法,是基于人声信号的评分方法,而本步骤S201获取的原唱音频信号中不仅包括人声信号,还包括该歌曲的伴奏信号,那么,则需要对该原唱音频信号进行处理,以获取原唱的人声信号,参见步骤S202和S203。
在步骤S202中,该终端设备对该原唱音频信号进行音频分离,得到第一原唱人声信号,该原唱音频信号为包含伴奏和原唱的人声的音频信号。
该终端设备可以通过该深度神经网络模型对该原唱音频信号进行人声和伴奏的分离,分离后的人声信号可能含有大量的伴奏残留信号;然后,该终端设备可以利用深度学习的后处理技术,将分离的人声信号映射为纯净的人声信号,来消除一部分伴奏残留信号,该纯净的人声信号即为原唱的人声信号,考虑到该原唱人声信号还可能携带少量的残留伴奏,那么,为了进一步消除该原唱的人声信号中的残留伴奏,该终端设备还需要对该原唱的人声信号进行处理,参见步骤S203。
在步骤S203中,该终端设备对该第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号。
当该第一原唱人声信号中带有毛刺时,该第一原唱人声信号就会不平滑,即可认为,该第一原唱人声信号带有残留伴奏。那么,该终端设备可以采用平滑技术来消除该第一原唱人声信号中的残留伴奏,得到第二原唱人声信号。
平滑技术也称为滤波技术,那么,本步骤S203可以是该终端设备对该第一原唱人声信号进行一阶递归平滑滤波的步骤,具体地,该终端设备可以使用一阶递归平滑滤波算法对该第一原唱人声信号进行滤波,以消除该第一原唱人声信号中的毛刺,从而得到较为平滑的第二原唱人声信号,由于得到的第二原唱人声信号较为平滑,所以不仅可以达到消除掉该第一原唱人声信号中的残留伴奏的目的,而且还能够保证该第二原唱人声信号的连续性。当然,该终端设备还可以采用其他的滤波算法对该第一原唱人声信号进行滤波,本申请实施例对其他滤波算法不做具体限定。
在步骤S204中,该终端设备提取该第二原唱人声信号的至少一个第一特征。
该至少一个第一特征可以包括音高、节奏、音色(谱质心、信号带宽等)以及响度。该终端设备可以使用基频提取方法,提取该第二原唱人声信号的 至少一个第一特征,该基频提取方法可以是pyin等任一种特征提取算法,本申请实施例对此不做限定。
需要说明的,该终端设备在提取该第二原唱人声信号的音高时,可以采用中值滤波法等相关平滑技术来解决提取音高时的半频和倍频的问题。
在步骤S205中,该终端设备基于该第二原唱人声信号的至少一个第一特征,获取每个第一特征的第一高阶统计特征,每个第一高阶统计特征为每个第一特征不低于二阶的统计量。
该不低于二阶的统计量,可以是平均值,还可以是方差。本步骤S205具体可以采用下述过程实现:该终端设备可以对该第二原唱人声信号的音高、节奏、音色以及响度等特征分别进行规整,基于规整的结果计算出每个第一特征对应的第一高阶统计特征。具体地,以第一特征为音高和节奏,第一高阶统计特征为平均值为例,对本步骤S205做如下说明:该终端设备将从该第二原唱人声信号中提取的所有音高规整在一起,将从该目标音频信号中提取的所有节奏规整在一起,然后,计算所有音高的平均值,以及所有节奏的平均值,那么,所有音高的平均值也即是音高对应的第一高阶统计特征,所有节奏的平均值也即是节奏对应的第一高阶统计特征。
值得注意的是,步骤S204和步骤S205的过程为该终端设备获取该第二原唱人声信号的至少一个第一高阶统计特征的过程,步骤S201至S205的过程为该终端设备获取原唱音频的至少一个第二高阶统计特征的过程,那么,可以将该终端设备通过步骤S201至S205的过程获取的至少一个第一高阶统计特征作为标准特征,来与目标音频信号的至少一个第二高阶统计特征作比较,并可以基于该比较的结果对该目标音频信号进行评分(见步骤S211和S212),从而可以避免邀请专业音乐人士来标注音高等特征,来为该目标音频信号提供标准特征,从而可以降低标注标准特征的人工成本。
在步骤S206中,该终端设备获取演唱者音频信号,该演唱者音频信号为包含伴奏和演唱者的人声的音频信号。
该终端设备可以通过录制,来获取该演唱者音频信号,例如,当演唱者在该K歌界面是点击“录制”后,该终端设备录制该终端设备当前接收到的声音,以获取该演唱者音频信号,该终端设备接收的声音可以包括伴奏声音和演唱者的人声;当终端设备开始播放该歌曲的伴奏时,该终端设备开始录制该终端设备当前接收到的声音,以获取该演唱者音频信号。
需要说明的是,本申请实施例提出的音频信号的评分方法,是基于人声信号的评分方法,而本步骤S206获取的演唱者音频信号为包含伴奏和演唱者的音频信号,那么,则需要对该演唱者音频信号进行处理,以获取演唱者的人声信号,参见步骤S207。
在步骤S207中,该终端设备基于该演唱者音频信号,获取目标音频信号。
其中,该目标音频信号为该演唱者音频信号中的演唱者的人声的音频信号,该终端设备可以通过如下步骤S207A和步骤S207B所示的过程来实现本步骤S207:
在步骤S207A中,该终端设备将该演唱者音频信号输入至深度神经网络模型中,得到演唱者人声信号,该深度神经网络模型用于分离该演唱者音频信号中的伴奏和演唱者人声。
在本步骤S207A中,该终端设备对该演唱者音频信号进行节奏和人声分离的方式,与步骤S202中,该终端设备对该原唱音频信号进行节奏和人声分离的方式可以是同一种方式,本申请实施例在此不做赘述。
在步骤S207B中,该终端设备对该演唱者人声信号进行噪声抑制处理,得到目标音频信号。
在本步骤S207B中,该终端设备对该演唱者音频信号进行噪声抑制处理的方式,与步骤S203中,该终端设备对该第一原唱人声信号进行噪声抑制处理的方式可以是同一种方式,本申请实施例在此不做赘述。
在步骤S208中,该终端设备对该演唱者音频信号进行人声检测,得到该目标音频信号的起始时间和结束时间。
该终端设备可以从该演唱者音频信号中提取至少一个人声特征,该人声特征用于区别人类的声音和除人类以外的生物或物体发出的声音,该人声特征可以包括能量、基频以及MPEG7音色等,该终端设备可以利用该人声特征对该演唱者音频信号进行人声检测,例如,当该终端设备从该演唱者音频信号中提取出的每个人声特征的值均符合预设范围,或者提取出的所有人声特征中的部分人声特征的值符合预设范围时,则认为该演唱者音频信号包含人声,否则,认为该演唱者音频信号不包含人声。那么,当该终端设备利用上述方法检测该演唱者音频信号时,将该终端设备检测到该演唱者音频信号中人声开始出现的时间,作为该目标音频信号的起始时间,当该终端设备确定该起始时间后,继续对该演唱者音频信号进行检测,将继续检测到该演唱者 音频信号中人声消失的时间,作为该目标音频信号的结束时间,该终端设备可以根据该起始时间和结束时间,来执行后续特征提取的步骤。
在步骤S209中,该终端设备提取该目标音频信号的至少一个第二特征。
在本步骤S209中,该终端设备提取该目标音频信号的至少一个第二特征的方式,与步骤S204中,该终端设备提取该第二原唱人声信号的至少一个第一特征的方式可以是同一种方式,本申请实施例在此不做赘述。
在步骤S210中,该终端设备基于该目标音频信号的至少一个第二特征,获取每个第二特征的第二高阶统计特征,每个第二高阶统计特征为每个第二特征不低于二阶的统计量。
在本步骤S210中,该终端设备获取每个第二特征的第二高阶统计特征的方式,与步骤S205中,该终端设备获取每个第一特征的第一高阶统计特征的方式可以是同一种方式,本申请实施例在此不做赘述。
需要说明的是,该终端设备在得到该目标音频信号的音高的第二高阶统计特征之前,还需要判断这些音高是否需要降低或者调高,例如,当该目标音频信号的平均音高高于该第二原唱人声信号的平均音高一个八度时,将该目标音频信号的音高降低一个八度;当该目标音频信号的平均音高低于所述第二原唱人声信号的平均音高一个八度时,将该目标音频信号的音高调高一个八度,当该目标音频信号的平均音高和该第二原唱人声信号的平均音高相同时,或者是,这两个平均音高的差别不超过一个八度时,该终端设备可以直接执行本步骤S210以得到该目标音频信号的音高的第二高阶统计特征,或者是,直接将该目标音频信号的音高的平均值作为音高的第二高阶统计特征。
具体地,以该终端设备从该目标音频信号中提取10个音高,该第二原唱人声信号的音高的平均值为B2调为例,当这10个音高的平均值为B3调时,则该目标音频信号的音高高于该第二原唱人声信号的音高的平均值一个八度,所以,该终端设备将该10个音高降低一个八度;当这10个音高的平均值为B1调时,则该目标音频信号的音高低于该第二原唱人声信号的音高的平均值一个八度,所以,该终端设备将该10个音高调高一个八度;然后,该终端设备可以基于降低或者调高后的10个音高,获取该目标音频信号的音高的第二高阶统计特征。
普通演唱者在歌唱音高较高的歌曲时,音高可能达不到原唱的音高,在歌唱音高较低的歌曲时,音高可能高于原唱的音高,从而导致普通演唱者的 目标音频信号与原唱人声信号相比,在音高方面差别过大,使得普通演唱者在音高特征上的评分较低,从而可能会影响对该普通演唱者的目标音频信号的综合评分,降低该普通演唱者的兴趣度。但是,本申请将音高较高的目标音频信号的音高降低一个八度,或者,将音高较低的目标音频信号的音高调高一个八度,那么,将会使得普通演唱者的目标音频信号的音高与第二原唱人声信号的音高上的差别没有那么大,从而可以提高该普通演唱者在音高方面的评分,从而可以使得该普通演唱者的综合得分不会太低,可以提高普通演唱者的K歌信心,进而保证了普通演唱者K歌的兴趣。
值得注意的是,步骤S209和步骤S210的过程为该终端设备获取目标音频信号的至少一个第二高阶统计特征的过程,而步骤S206至S210的过程为该终端设备获取演唱者音频信号的至少一个第二高阶统计特征的过程。
在步骤S211中,该终端设备将每个第一高阶统计特征与对应的第二高阶统计特征的差异映射到评分区间内,得到每个第二高阶统计特征的得分。
该评分区间可以是[0,100]的评分区间,当然,该评分区间还可以是其他分值区间。在一种实施方式中,该终端设备可以将每个第一高阶统计特征与该第一高阶统计特征对应的第二高阶统计特征做减法,得到每个第一高阶统计特征与该第二高阶统计特征的差值,并对该差值取绝对值,然后,将该绝对值映射到评分区间内。
具体地,以方差作为第一高阶统计特征和第二高阶统计特征,并且该第一高阶统计特征为0.5,与该第一高阶统计特征对应的第二高阶统计特征为0.48为例,对本步骤S211做如下描述:首先,0.5-0.48=0.02;其次,|0.02|=0.02;将该0.02映射至[0,100]的评分区间内,可以使得该0.02映射为98,那么,98分即为这个第二高阶统计的得分。
通过上述映射过程,可以得到每个第二高阶统计特征的得分,并且,当第一高阶统计特征与其对应的第二高阶统计特征差异越大,从而得分越低,当第一高阶统计特征与其对应的第二高阶统计特征差异越小,从而得分越高。
本申请实施例对于该评分区间以及将每个第二高阶特征的差值映射到该评分区间的方法不做具体限定。
在步骤S212中,该终端设备基于该每个第二高阶统计特征的权重以及所述每个第二高阶统计特征的得分,获取目标音频信号的得分。
在步骤S212中,可以融合各个第二高阶统计特征的得分来得到目标音频 信号的得分。在一种可能实现方式中,该终端设备可以为该每个第二高阶统计特征设置权重,该终端设备可以在获取该每个第二高阶统计特征的得分之前,为该每个第二高阶统计特征设置权重,还可以在获取该每个第二高阶统计特征的得分之后,为该每个第二高阶统计特征设置权重,本申请实施例对每个第二高阶统计特征设置权重的时间不做具体限定。
本步骤S212可以通过下述过程来实现:该终端设备将该每个第二高阶统计特征的得分和对应的权重进行乘法运算,得到每个第二高阶统计特征对应的乘积,再获取各个乘积的和值,将该和值作为目标音频信号的得分。当然,还可以将该和值映射到一个评分区间内,和值越高,评分越高,以得到一个满足评分标准的得分。例如,该评分区间可以为[0,100],使得评分能够满足百分制评分标准。
本申请实施例通过对从原唱音频信号中分离的第一原唱人声信号进行噪声抑制处理,可以使得该终端设备获得伴奏残留较少的第二原唱人声信号,从而使得该终端设备从该第二原唱人声信号中提取出的音高以及节奏等特征比较准确,当将这些比较准确的特征与演唱者的目标音频信号的特征相比较时,可以保证比较结果的准确性,进而当该终端设备利用准确的比较结果对该目标音频信号进行评分时,可以保证目标音频信号的得分的准确性。并且,该终端设备采用该第二原唱人声信号以及目标音频信号的高阶统计特征之间的差异,对该目标音频信号进行评分,提高了评分的准确性和鲁棒性。并且,该终端设备通过人声检测,来确定目标音频信号的起始时间和结束时间,从而可以避免利用人工来标注目标音频信号的起始时间和结束时间,进而可以降低人工成本。并且,该终端设备根据演唱者的目标音频信号与该第二原唱人声信号在平均音高上的差异,对目标音频信号的音高进行降低或者提高,以降低该目标音频信号与该第二原唱人声信号之间在音高方面的差异,从而可以提高该普通演唱者在音高方面的评分,使得音高较低或者较高的演唱者的目标音频信号的得分不会太低,可以提高音高较低或者较高的演唱者的K歌信心,进而保证了音高较低或者较高的演唱者K歌的兴趣。并且,该终端设备在提取该第二原唱人声信号的音高时,采用了中值滤波法等相关平滑技术来解决提取音高时的半频和倍频的问题。
图3是根据一示例性实施例示出的一种音频信号的评分装置的逻辑结构框图。参照图3,该装置包括分离单元301,抑制单元302和获取单元303。
其中,分离单元301与抑制单元302相连接,被配置为执行对原唱音频信号进行音频分离,得到第一原唱人声信号,该原唱音频信号为包含伴奏和原唱的人声的音频信号;
抑制单元302与获取单元303相连接,被配置为执行对该第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号;
获取单元303,被配置为执行基于该第二原唱人声信号与目标音频信号之间的差异,获取该目标音频信号的得分。
可选地,该分离单元301,还被配置为执行将该原唱音频信号输入至深度神经网络模型中,得到该第一原唱人声信号,该深度神经网络模型用于分离该原唱音频信号中的伴奏和原唱。
可选地,该获取单元303,包括:
第一获取子单元,被配置为执行基于该第二原唱人声信号的至少一个第一特征,获取每个第一特征的第一高阶统计特征,每个第一高阶统计特征为每个第一特征不低于二阶的统计量;
第二获取子单元,被配置为执行基于该目标音频信号的至少一个第二特征,获取每个第二特征的第二高阶统计特征,每个第二高阶统计特征为每个第二特征不低于二阶的统计量;
第三获取子单元,被配置为执行基于每个第一高阶统计特征与对应的第二高阶统计特征的差异,获取该目标音频信号的得分。
可选地,该音频信号的评分装置,还包括:
降低单元,被配置为执行当所述目标音频信号的平均音高高于所述第二原唱人声信号的平均音高一个八度时,将所述目标音频信号的音高降低一个八度;
调高单元,被配置为执行当所述目标音频信号的平均音高低于所述第二原唱人声信号的平均音高一个八度时,将所述目标音频信号的音高调高一个八度。
可选地,该第三获取子单元,还被配置为执行将该每个第一高阶统计特征与对应的第二高阶统计特征的差异映射到评分区间内,得到每个第二高阶统计特征的得分;
基于该每个第二高阶统计特征的权重以及该每个第二高阶统计特征的得分,获取目标音频信号的得分。
可选地,该音频信号的评分装置,还包括检测单元,被配置为执行对演唱者音频信号进行人声检测,得到该目标音频信号的起始时间和结束时间,该演唱者音频信号为包含伴奏和演唱者的人声的音频信号。
关于上述实施例中的装置,其中各个单元执行操作的具体方式已经在有关该方法的实施例中进行了详细描述,此处将不做详细阐述说明。
图4是根据一示例性实施例示出的一种终端设备的结构框图。该终端设备400可以是:智能手机、平板电脑、MP3播放器(Moving Picture Experts Group Audio Layer III,动态影像专家压缩标准音频层面3)、MP4(Moving Picture Experts Group Audio Layer IV,动态影像专家压缩标准音频层面4)播放器、笔记本电脑或台式电脑。终端设备400还可能被称为用户设备、便携式终端设备、膝上型终端设备、台式终端设备等其他名称。
通常,终端设备400包括有:处理器401和存储器402。
处理器401可以包括一个或多个处理核心,比如4核心处理器、8核心处理器等。处理器401可以采用数字信号处理(Digital Signal Processing,DSP)、现场可编程门阵列(Field-Programmable Gate Array,FPGA)、可编程逻辑阵列(Programmable Logic Array,PLA)中的至少一种硬件形式来实现。处理器401也可以包括主处理器和协处理器,主处理器是用于对在唤醒状态下的数据进行处理的处理器,也称中央处理器(Central Processing Unit,CPU);协处理器是用于对在待机状态下的数据进行处理的低功耗处理器。在一些实施例中,处理器401可以在集成有图像处理器(Graphics Processing Unit,GPU),GPU用于负责显示屏所需要显示的内容的渲染和绘制。一些实施例中,处理器401还可以包括人工智能(Artificial Intelligence,AI)处理器,该AI处理器用于处理有关机器学习的计算操作。
存储器402可以包括一个或多个计算机可读存储介质,该计算机可读存储介质可以是非暂态的。存储器402还可包括高速随机存取存储器,以及非易失性存储器,比如一个或多个磁盘存储设备、闪存存储设备。在一些实施例中,存储器402中的非暂态的计算机可读存储介质用于存储至少一个指令,该至少一个指令用于被处理器401所执行以实现本申请中方法实施例提供的音频信号的评分方法。
在一些实施例中,终端设备400还可选包括有:外围设备接口403和至少一个外围设备。处理器401、存储器402和外围设备接口403之间可以通过 总线或信号线相连。各个外围设备可以通过总线、信号线或电路板与外围设备接口403相连。具体地,外围设备包括:射频电路404、触摸显示屏405、摄像头406、音频电路407、定位组件408和电源409中的至少一种。
外围设备接口403可被用于将输入/输出(Input/Output,I/O)相关的至少一个外围设备连接到处理器401和存储器402。在一些实施例中,处理器401、存储器402和外围设备接口403被集成在同一芯片或电路板上;在一些其他实施例中,处理器401、存储器402和外围设备接口403中的任意一个或两个可以在单独的芯片或电路板上实现,本实施例对此不加以限定。
射频电路404用于接收和发射射频(Radio Frequency,RF)信号,也称电磁信号。射频电路404通过电磁信号与通信网络以及其他通信设备进行通信。射频电路404将电信号转换为电磁信号进行发送,或者,将接收到的电磁信号转换为电信号。可选地,射频电路404包括:天线系统、RF收发器、一个或多个放大器、调谐器、振荡器、数字信号处理器、编解码芯片组、用户身份模块卡等等。射频电路404可以通过至少一种无线通信协议来与其它终端设备进行通信。该无线通信协议包括但不限于:城域网、各代移动通信网络(2G、3G、4G及5G)、无线局域网和/或无线保真(Wireless Fidelity,WiFi)网络。在一些实施例中,射频电路404还可以包括近距离无线通信(Near Field Communication,NFC)有关的电路,本申请对此不加以限定。
显示屏405用于显示用户界面(User Interface,UI)。该UI可以包括图形、文本、图标、视频及其它们的任意组合。当显示屏405是触摸显示屏时,显示屏405还具有采集在显示屏405的表面或表面上方的触摸信号的能力。该触摸信号可以作为控制信号输入至处理器401进行处理。此时,显示屏405还可以用于提供虚拟按钮和/或虚拟键盘,也称软按钮和/或软键盘。在一些实施例中,显示屏405可以为一个,设置终端设备400的前面板;在另一些实施例中,显示屏405可以为至少两个,分别设置在终端设备400的不同表面或呈折叠设计;在再一些实施例中,显示屏405可以是柔性显示屏,设置在终端设备400的弯曲表面上或折叠面上。甚至,显示屏405还可以设置成非矩形的不规则图形,也即异形屏。显示屏405可以采用液晶显示屏(Liquid Crystal Display,LCD)、有机发光二极管(Organic Light-Emitting Diode,OLED)等材质制备。
摄像头组件406用于采集图像或视频。可选地,摄像头组件406包括前 置摄像头和后置摄像头。通常,前置摄像头设置在终端设备的前面板,后置摄像头设置在终端设备的背面。在一些实施例中,后置摄像头为至少两个,分别为主摄像头、景深摄像头、广角摄像头、长焦摄像头中的任意一种,以实现主摄像头和景深摄像头融合实现背景虚化功能、主摄像头和广角摄像头融合实现全景拍摄以及虚拟现实(Virtual Reality,VR)拍摄功能或者其它融合拍摄功能。在一些实施例中,摄像头组件406还可以包括闪光灯。闪光灯可以是单色温闪光灯,也可以是双色温闪光灯。双色温闪光灯是指暖光闪光灯和冷光闪光灯的组合,可以用于不同色温下的光线补偿。
音频电路407可以包括麦克风和扬声器。麦克风用于采集用户及环境的声波,并将声波转换为电信号输入至处理器401进行处理,或者输入至射频电路404以实现语音通信。出于立体声采集或降噪的目的,麦克风可以为多个,分别设置在终端设备400的不同部位。麦克风还可以是阵列麦克风或全向采集型麦克风。扬声器则用于将来自处理器401或射频电路404的电信号转换为声波。扬声器可以是传统的薄膜扬声器,也可以是压电陶瓷扬声器。当扬声器是压电陶瓷扬声器时,不仅可以将电信号转换为人类可听见的声波,也可以将电信号转换为人类听不见的声波以进行测距等用途。在一些实施例中,音频电路407还可以包括耳机插孔。
定位组件408用于定位终端设备400的当前地理位置,以实现导航或基于位置的服务(Location Based Service,LBS)。定位组件508可以是基于美国的全球定位系统(Global Positioning System,GPS)、中国的北斗系统、俄罗斯的格雷纳斯系统或欧盟的伽利略系统的定位组件。
电源409用于为终端设备400中的各个组件进行供电。电源409可以是交流电、直流电、一次性电池或可充电电池。当电源409包括可充电电池时,该可充电电池可以支持有线充电或无线充电。该可充电电池还可以用于支持快充技术。
在一些实施例中,终端设备400还包括有一个或多个传感器410。该一个或多个传感器410包括但不限于:加速度传感器411、陀螺仪传感器412、压力传感器413、指纹传感器414、光学传感器415以及接近传感器416。
加速度传感器411可以检测以终端设备400建立的坐标系的三个坐标轴上的加速度大小。比如,加速度传感器411可以用于检测重力加速度在三个坐标轴上的分量。处理器401可以根据加速度传感器411采集的重力加速度 信号,控制触摸显示屏405以横向视图或纵向视图进行用户界面的显示。加速度传感器411还可以用于游戏或者用户的运动数据的采集。
陀螺仪传感器412可以检测终端设备400的机体方向及转动角度,陀螺仪传感器412可以与加速度传感器411协同采集用户对终端设备400的3D动作。处理器401根据陀螺仪传感器412采集的数据,可以实现如下功能:动作感应(比如根据用户的倾斜操作来改变UI)、拍摄时的图像稳定、游戏控制以及惯性导航。
压力传感器413可以设置在终端设备400的侧边框和/或触摸显示屏405的下层。当压力传感器413设置在终端设备400的侧边框时,可以检测用户对终端设备400的握持信号,由处理器401根据压力传感器413采集的握持信号进行左右手识别或快捷操作。当压力传感器413设置在触摸显示屏405的下层时,由处理器401根据用户对触摸显示屏405的压力操作,实现对UI界面上的可操作性控件进行控制。可操作性控件包括按钮控件、滚动条控件、图标控件、菜单控件中的至少一种。
指纹传感器414用于采集用户的指纹,由处理器401根据指纹传感器414采集到的指纹识别用户的身份,或者,由指纹传感器414根据采集到的指纹识别用户的身份。在识别出用户的身份为可信身份时,由处理器401授权该用户执行相关的敏感操作,该敏感操作包括解锁屏幕、查看加密信息、下载软件、支付及更改设置等。指纹传感器414可以被设置终端设备400的正面、背面或侧面。当终端设备400上设置有物理按键或厂商Logo时,指纹传感器414可以与物理按键或厂商Logo集成在一起。
光学传感器415用于采集环境光强度。在一个实施例中,处理器401可以根据光学传感器415采集的环境光强度,控制触摸显示屏405的显示亮度。具体地,当环境光强度较高时,调高触摸显示屏405的显示亮度;当环境光强度较低时,调低触摸显示屏405的显示亮度。在另一个实施例中,处理器401还可以根据光学传感器415采集的环境光强度,动态调整摄像头组件406的拍摄参数。
接近传感器416,也称距离传感器,通常设置在终端设备400的前面板。接近传感器416用于采集用户与终端设备400的正面之间的距离。在一个实施例中,当接近传感器416检测到用户与终端设备400的正面之间的距离逐渐变小时,由处理器401控制触摸显示屏405从亮屏状态切换为息屏状态; 当接近传感器416检测到用户与终端设备400的正面之间的距离逐渐变大时,由处理器401控制触摸显示屏405从息屏状态切换为亮屏状态。
图5是根据一示例性实施例示出的一种终端设备的结构框图,该终端设备500可因配置或性能不同而产生比较大的差异,可以包括一个或一个以上处理器(Central Processing Units,CPU)501和一个或一个以上的存储器502,其中,所述存储器502中存储有至少一条指令,所述至少一条指令由所述处理器501加载并执行以实现下列过程:
对原唱音频信号进行音频分离,得到第一原唱人声信号,所述原唱音频信号为包含伴奏和原唱的人声的音频信号;对所述第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号;基于所述第二原唱人声信号与目标音频信号之间的差异,获取所述目标音频信号的得分。
可选地,所述处理器501具体用于:将所述原唱音频信号输入至深度神经网络模型中,得到所述第一原唱人声信号,所述深度神经网络模型用于分离所述原唱音频信号中的伴奏和原唱。
可选地,所述处理器501具体用于:基于所述第二原唱人声信号的至少一个第一特征,获取每个第一特征的第一高阶统计特征,每个第一高阶统计特征为每个第一特征不低于二阶的统计量;基于所述目标音频信号的至少一个第二特征,获取每个第二特征的第二高阶统计特征,每个第二高阶统计特征为每个第二特征不低于二阶的统计量;基于每个第一高阶统计特征与对应的第二高阶统计特征的差异,获取所述目标音频信号的得分。
在本申请实施例中,在所述基于所述目标音频信号的至少一个第二特征,获取每个第二特征的第二高阶统计特征之前,所述处理器501具体用于:当所述目标音频信号的平均音高高于所述第二原唱人声信号的平均音高一个八度时,将所述目标音频信号的音高降低一个八度;当所述目标音频信号的平均音高低于所述第二原唱人声信号的平均音高一个八度时,将所述目标音频信号的音高调高一个八度。
可选地,所述处理器501具体用于:将所述每个第一高阶统计特征与对应的第二高阶统计特征的差异映射到评分区间内,得到每个第二高阶统计特征的得分;基于所述每个第二高阶统计特征的权重以及所述每个第二高阶统计特征的得分,获取目标音频信号的得分。
在所述基于所述第二原唱人声信号与目标音频信号之间的差异,获取所 述目标音频信号的得分之前,所述处理器501具体用于:对演唱者音频信号进行人声检测,得到所述目标音频信号的起始时间和结束时间,所述演唱者音频信号为包含伴奏和演唱者的人声的音频信号。
当然,该终端设备还可以具有有线或无线网络接口、键盘以及输入输出接口等部件,以便进行输入输出,该终端设备还可以包括其他用于实现设备功能的部件,在此不做赘述。
在示例性实施例中,还提供了一种非临时性计算机可读存储介质,例如包括指令的存储器,上述指令可由终端中的处理器执行以完成上述实施例中的音频信号的评分方法。例如,该计算机可读存储介质可以是只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、只读光盘(Compact Disc Read-Only Memory,CD-ROM)、磁带、软盘和光数据存储设备等。
应当理解的是,本申请并不局限于上面已经描述并在附图中示出的精确结构,并且可以在不脱离其范围进行各种修改和改变。本申请的范围仅由所附的权利要求来限制。
Claims (14)
- 一种音频信号的评分方法,包括:对原唱音频信号进行音频分离,得到第一原唱人声信号,所述原唱音频信号为包含伴奏和原唱的人声的音频信号;对所述第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号;基于所述第二原唱人声信号与目标音频信号之间的差异,获取所述目标音频信号的得分。
- 根据权利要求1所述的音频信号的评分方法,所述对原唱音频信号进行音频分离,得到第一原唱人声信号,包括:将所述原唱音频信号输入至深度神经网络模型中,得到所述第一原唱人声信号,所述深度神经网络模型用于分离所述原唱音频信号中的伴奏和原唱。
- 根据权利要求1所述的音频信号的评分方法,所述基于所述第二原唱人声信号与目标音频信号之间的差异,获取所述目标音频信号的得分,包括:基于所述第二原唱人声信号的至少一个第一特征,获取每个第一特征的第一高阶统计特征,每个第一高阶统计特征为每个第一特征不低于二阶的统计量;基于所述目标音频信号的至少一个第二特征,获取每个第二特征的第二高阶统计特征,每个第二高阶统计特征为每个第二特征不低于二阶的统计量;基于每个第一高阶统计特征与对应的第二高阶统计特征的差异,获取所述目标音频信号的得分。
- 根据权利要求3所述的音频信号的评分方法,所述基于所述目标音频信号的至少一个第二特征,获取每个第二特征的第二高阶统计特征之前,还包括:当所述目标音频信号的平均音高高于所述第二原唱人声信号的平均音高一个八度时,将所述目标音频信号的音高降低一个八度;当所述目标音频信号的平均音高低于所述第二原唱人声信号的平均音高一个八度时,将所述目标音频信号的音高调高一个八度。
- 根据权利要求3所述的音频信号的评分方法,所述基于每个第一高阶统计特征与对应的第二高阶统计特征的差异,获取所述目标音频信号的得分,包括:将所述每个第一高阶统计特征与对应的第二高阶统计特征的差异映射到评分区间内,得到每个第二高阶统计特征的得分;基于所述每个第二高阶统计特征的权重以及所述每个第二高阶统计特征的得分,获取目标音频信号的得分。
- 根据权利要求1所述的音频信号的评分方法,所述基于所述第二原唱人声信号与目标音频信号之间的差异,获取所述目标音频信号的得分之前,还包括:对演唱者音频信号进行人声检测,得到所述目标音频信号的起始时间和结束时间,所述演唱者音频信号为包含伴奏和演唱者的人声的音频信号。
- 一种音频信号的评分装置,包括:分离单元,被配置为执行对原唱音频信号进行音频分离,得到第一原唱人声信号,所述原唱音频信号为包含伴奏和原唱人声的音频信号;抑制单元,被配置为执行采用平滑技术对所述第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号;获取单元,被配置为执行基于所述第二原唱人声信号与目标音频信号之间的差异,获取所述目标音频信号的得分。
- 一种终端设备,包括:处理器用于存储处理器可执行指令的存储器;其中,当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现下列过程:对原唱音频信号进行音频分离,得到第一原唱人声信号,所述原唱音频信号为包含伴奏和原唱的人声的音频信号;对所述第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号;基于所述第二原唱人声信号与目标音频信号之间的差异,获取所述目标音频信号的得分。
- 根据权利要求8所述的终端设备,所述处理器具体用于:将所述原唱音频信号输入至深度神经网络模型中,得到所述第一原唱人声信号,所述深度神经网络模型用于分离所述原唱音频信号中的伴奏和原唱。
- 根据权利要求8所述的终端设备,所述处理器具体用于:基于所述第二原唱人声信号的至少一个第一特征,获取每个第一特征的第一高阶统计特征,每个第一高阶统计特征为每个第一特征不低于二阶的统计量;基于所述目标音频信号的至少一个第二特征,获取每个第二特征的第二高阶统计特征,每个第二高阶统计特征为每个第二特征不低于二阶的统计量;基于每个第一高阶统计特征与对应的第二高阶统计特征的差异,获取所述目标音频信号的得分。
- 根据权利要求10所述的终端设备,所述处理器具体用于:当所述目标音频信号的平均音高高于所述第二原唱人声信号的平均音高一个八度时,将所述目标音频信号的音高降低一个八度;当所述目标音频信号的平均音高低于所述第二原唱人声信号的平均音高一个八度时,将所述目标音频信号的音高调高一个八度。
- 根据权利要求10所述的终端设备,所述处理器具体用于:将所述每个第一高阶统计特征与对应的第二高阶统计特征的差异映射到评分区间内,得到每个第二高阶统计特征的得分;基于所述每个第二高阶统计特征的权重以及所述每个第二高阶统计特征的得分,获取目标音频信号的得分。
- 根据权利要求8所述的终端设备,所述处理器具体用于:对演唱者音频信号进行人声检测,得到所述目标音频信号的起始时间和结束时间,所述演唱者音频信号为包含伴奏和演唱者的人声的音频信号。
- 一种非临时性计算机可读存储介质,当所述存储介质中的指令由移动终端的处理器执行时,使得移动终端能够执行一种音频信号的评分方法,所述方法包括:对原唱音频信号进行音频分离,得到第一原唱人声信号,所述原唱音频信号为包含伴奏和原唱的人声的音频信号;对所述第一原唱人声信号进行噪声抑制处理,得到第二原唱人声信号;基于所述第二原唱人声信号与目标音频信号之间的差异,获取所述目标音频信号的得分。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201811376670.8 | 2018-11-19 | ||
| CN201811376670.8A CN109300485B (zh) | 2018-11-19 | 2018-11-19 | 音频信号的评分方法、装置、电子设备及计算机存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020103550A1 true WO2020103550A1 (zh) | 2020-05-28 |
Family
ID=65144253
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/106502 Ceased WO2020103550A1 (zh) | 2018-11-19 | 2019-09-18 | 音频信号的评分方法、装置、终端设备及计算机存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN109300485B (zh) |
| WO (1) | WO2020103550A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2021254961A1 (en) * | 2020-06-16 | 2021-12-23 | Sony Group Corporation | Audio transposition |
| CN116844569A (zh) * | 2023-07-06 | 2023-10-03 | 北京达佳互联信息技术有限公司 | 音频数据处理方法、装置、电子设备及存储介质 |
Families Citing this family (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109300485B (zh) * | 2018-11-19 | 2022-06-10 | 北京达佳互联信息技术有限公司 | 音频信号的评分方法、装置、电子设备及计算机存储介质 |
| CN110010159B (zh) * | 2019-04-02 | 2021-12-10 | 广州酷狗计算机科技有限公司 | 声音相似度确定方法及装置 |
| CN110189770B (zh) * | 2019-06-18 | 2021-06-25 | 北京达佳互联信息技术有限公司 | 语音数据处理方法、装置、终端、服务器及介质 |
| CN110277106B (zh) * | 2019-06-21 | 2021-10-22 | 北京达佳互联信息技术有限公司 | 音频质量确定方法、装置、设备及存储介质 |
| CN110660383A (zh) * | 2019-09-20 | 2020-01-07 | 华南理工大学 | 一种基于歌词歌声对齐的唱歌评分方法 |
| CN110728968A (zh) * | 2019-10-14 | 2020-01-24 | 腾讯音乐娱乐科技(深圳)有限公司 | 一种音频伴奏信息的评估方法、装置及存储介质 |
| US20230186782A1 (en) * | 2020-06-05 | 2023-06-15 | Sony Group Corporation | Electronic device, method and computer program |
| CN113870897A (zh) * | 2021-10-19 | 2021-12-31 | 广州酷狗计算机科技有限公司 | 音频数据教学测评方法及其装置、设备、介质、产品 |
| CN113823270B (zh) * | 2021-10-28 | 2024-05-03 | 杭州网易云音乐科技有限公司 | 节奏评分的确定方法、介质、装置和计算设备 |
| CN114566191A (zh) * | 2022-02-25 | 2022-05-31 | 腾讯音乐娱乐科技(深圳)有限公司 | 录音的修音方法及相关装置 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006154212A (ja) * | 2004-11-29 | 2006-06-15 | Ntt Advanced Technology Corp | 音声評価方法および評価装置 |
| CN104133851A (zh) * | 2014-07-07 | 2014-11-05 | 小米科技有限责任公司 | 音频相似度的检测方法和检测装置、电子设备 |
| CN104464727A (zh) * | 2014-12-11 | 2015-03-25 | 福州大学 | 一种基于深度信念网络的单通道音乐的歌声分离方法 |
| CN104538011A (zh) * | 2014-10-30 | 2015-04-22 | 华为技术有限公司 | 一种音调调节方法、装置及终端设备 |
| CN109300485A (zh) * | 2018-11-19 | 2019-02-01 | 北京达佳互联信息技术有限公司 | 音频信号的评分方法、装置、电子设备及计算机存储介质 |
Family Cites Families (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP6141737B2 (ja) * | 2013-09-30 | 2017-06-07 | 株式会社第一興商 | ストレッチチューニングを考慮して歌唱採点を行うカラオケ装置 |
| CN103971674B (zh) * | 2014-05-22 | 2017-02-15 | 天格科技(杭州)有限公司 | 一种演唱实时评分方法 |
| CN104183245A (zh) * | 2014-09-04 | 2014-12-03 | 福建星网视易信息系统有限公司 | 一种演唱者音色相似的歌星推荐方法与装置 |
| JP6365483B2 (ja) * | 2015-09-24 | 2018-08-01 | ブラザー工業株式会社 | カラオケ装置,カラオケシステム,及びプログラム |
| CN105590633A (zh) * | 2015-11-16 | 2016-05-18 | 福建省百利亨信息科技有限公司 | 一种用于歌曲评分的曲谱生成方法和设备 |
| CN106024005B (zh) * | 2016-07-01 | 2018-09-25 | 腾讯科技(深圳)有限公司 | 一种音频数据的处理方法及装置 |
| CN107146497A (zh) * | 2016-08-02 | 2017-09-08 | 浙江大学 | 一种钢琴考级评分系统 |
| CN107767850A (zh) * | 2016-08-23 | 2018-03-06 | 冯山泉 | 一种演唱评分方法及系统 |
| CN106782600B (zh) * | 2016-12-29 | 2020-04-24 | 广州酷狗计算机科技有限公司 | 音频文件的评分方法及装置 |
| CN107818796A (zh) * | 2017-11-16 | 2018-03-20 | 重庆师范大学 | 一种音乐考试评定方法及系统 |
| CN108008930B (zh) * | 2017-11-30 | 2020-06-30 | 广州酷狗计算机科技有限公司 | 确定k歌分值的方法和装置 |
-
2018
- 2018-11-19 CN CN201811376670.8A patent/CN109300485B/zh active Active
-
2019
- 2019-09-18 WO PCT/CN2019/106502 patent/WO2020103550A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006154212A (ja) * | 2004-11-29 | 2006-06-15 | Ntt Advanced Technology Corp | 音声評価方法および評価装置 |
| CN104133851A (zh) * | 2014-07-07 | 2014-11-05 | 小米科技有限责任公司 | 音频相似度的检测方法和检测装置、电子设备 |
| CN104538011A (zh) * | 2014-10-30 | 2015-04-22 | 华为技术有限公司 | 一种音调调节方法、装置及终端设备 |
| CN104464727A (zh) * | 2014-12-11 | 2015-03-25 | 福州大学 | 一种基于深度信念网络的单通道音乐的歌声分离方法 |
| CN109300485A (zh) * | 2018-11-19 | 2019-02-01 | 北京达佳互联信息技术有限公司 | 音频信号的评分方法、装置、电子设备及计算机存储介质 |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2021254961A1 (en) * | 2020-06-16 | 2021-12-23 | Sony Group Corporation | Audio transposition |
| US20230215454A1 (en) * | 2020-06-16 | 2023-07-06 | Sony Group Corporation | Audio transposition |
| US12412593B2 (en) * | 2020-06-16 | 2025-09-09 | Sony Group Corporation | Audio transposition |
| CN116844569A (zh) * | 2023-07-06 | 2023-10-03 | 北京达佳互联信息技术有限公司 | 音频数据处理方法、装置、电子设备及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109300485B (zh) | 2022-06-10 |
| CN109300485A (zh) | 2019-02-01 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020103550A1 (zh) | 音频信号的评分方法、装置、终端设备及计算机存储介质 | |
| CN110688082B (zh) | 确定音量的调节比例信息的方法、装置、设备及存储介质 | |
| CN108008930B (zh) | 确定k歌分值的方法和装置 | |
| CN110277106B (zh) | 音频质量确定方法、装置、设备及存储介质 | |
| CN110956971B (zh) | 音频处理方法、装置、终端及存储介质 | |
| CN111048111B (zh) | 检测音频的节奏点的方法、装置、设备及可读存储介质 | |
| CN108831423B (zh) | 提取音频数据中主旋律音轨的方法、装置、终端及存储介质 | |
| CN110867194B (zh) | 音频的评分方法、装置、设备及存储介质 | |
| WO2022111168A1 (zh) | 视频的分类方法和装置 | |
| WO2019128593A1 (zh) | 搜索音频的方法和装置 | |
| CN110837557B (zh) | 摘要生成方法、装置、设备及介质 | |
| CN112086102B (zh) | 扩展音频频带的方法、装置、设备以及存储介质 | |
| CN108922562A (zh) | 演唱评价结果显示方法及装置 | |
| CN112667844B (zh) | 检索音频的方法、装置、设备和存储介质 | |
| CN113362836B (zh) | 训练声码器方法、终端及存储介质 | |
| CN109192223B (zh) | 音频对齐的方法和装置 | |
| CN111081277B (zh) | 音频测评的方法、装置、设备及存储介质 | |
| CN111428079B (zh) | 文本内容处理方法、装置、计算机设备及存储介质 | |
| CN112597331B (zh) | 显示音域匹配信息的方法、装置、设备和存储介质 | |
| CN112992107B (zh) | 训练声学转换模型的方法、终端及存储介质 | |
| CN113257222B (zh) | 合成歌曲音频的方法、终端及存储介质 | |
| CN111125424B (zh) | 提取歌曲核心歌词的方法、装置、设备及存储介质 | |
| CN109003627B (zh) | 确定音频得分的方法、装置、终端及存储介质 | |
| CN111063372B (zh) | 确定音高特征的方法、装置、设备及存储介质 | |
| CN110910862B (zh) | 音频调整方法、装置、服务器及计算机可读存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19886105 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19886105 Country of ref document: EP Kind code of ref document: A1 |