WO2023005409A1 - 设备确定方法及设备确定系统 - Google Patents

设备确定方法及设备确定系统 Download PDF

Info

Publication number
WO2023005409A1
WO2023005409A1 PCT/CN2022/096393 CN2022096393W WO2023005409A1 WO 2023005409 A1 WO2023005409 A1 WO 2023005409A1 CN 2022096393 W CN2022096393 W CN 2022096393W WO 2023005409 A1 WO2023005409 A1 WO 2023005409A1
Authority
WO
WIPO (PCT)
Prior art keywords
audio
target
signal
direct
direct mixing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/096393
Other languages
English (en)
French (fr)
Inventor
郝斌
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Qingdao Haier Technology Co Ltd
Haier Smart Home Co Ltd
Original Assignee
Qingdao Haier Technology Co Ltd
Haier Smart Home Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Qingdao Haier Technology Co Ltd, Haier Smart Home Co Ltd filed Critical Qingdao Haier Technology Co Ltd
Publication of WO2023005409A1 publication Critical patent/WO2023005409A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/21Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being power information
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • G10L2015/223Execution procedure of a spoken command
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02DCLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D30/00Reducing energy consumption in communication networks
    • Y02D30/70Reducing energy consumption in communication networks in wireless communication networks

Definitions

  • the present disclosure relates to the field of smart home, and in particular, to a device determination method and a device determination system.
  • the voice interaction module equipped with each device judges that after the wake-up, each device performs simple processing on the voice signal received by itself: AEC, noise reduction, demixing
  • AEC noise reduction
  • demixing The purpose is to more accurately estimate the energy mean value of the sound source signal in the received signal. Since the sensitivity of the microphone sensor of each device is not necessarily the same, most solutions require additional energy calibration, that is, to select a certain device as a reference.
  • the ratio of the energy of the equipment to be calibrated to the reference equipment is calculated.
  • the energy values calculated by other devices except the reference device must be multiplied by the corresponding ratio.
  • the energy average value is particularly easy to be affected by reverberation factors , that is, the received signal of the current frame includes the target speech signal and the historical speech signal, and the mean value calculated at this time cannot accurately represent the distance information.
  • Embodiments of the present disclosure provide a device determination method and a device determination system to at least solve the technical problem of low accuracy in judging the distance between each device and the target sound source based on the energy average value due to the inconsistency of sensor sensitivity of each device.
  • a method for determining a device including: acquiring the direct mixing ratio of target audio signals received by multiple first audio devices, where the direct mixing ratio is the target audio signal received by each audio device The energy ratio of direct audio and reverberant audio in the signal; determining a target direct mixing ratio from multiple direct mixing ratios; determining a target audio device corresponding to the target direct mixing ratio from multiple first audio devices.
  • a method for determining a device including: receiving a target audio signal, and calculating a direct mixing ratio of the received target audio signal, where the direct mixing ratio is the direct mixing ratio received by the first audio device The energy ratio of direct audio and reverberation audio in the target audio signal; send the direct mixing ratio to the server; receive the judgment instruction sent by the server, and determine whether the first audio device is the target audio device according to the judgment instruction, and the target audio device is to be awakened For the audio device, the judging instruction is generated according to the direct mixing ratios sent by the multiple first audio devices.
  • an audio processing device at least includes a communication module, a processor and an audio collection module, wherein: the audio collection module is configured to receive a target audio signal; the processor, It is set to calculate the direct mixing ratio of the received target audio signal, where the direct mixing ratio is the energy ratio of direct audio and reverberant audio in the target audio signal received by the audio device; and determine the audio processing device according to the judgment instruction received by the communication module Whether it is a target audio device, the target audio device is an audio device to be awakened, wherein the judgment instruction is generated according to the direct mixing ratio sent by multiple audio processing devices; the communication module is set to send the direct mixing ratio to the server, and receive the server sending judgment instruction.
  • a device determination system including a plurality of audio processing devices and a server, wherein: each audio processing device in the plurality of audio processing devices is configured to receive The target audio, and set to calculate the direct mixing ratio of the received target audio, where the direct mixing ratio is the energy ratio of the direct audio received by each audio processing device and the reverberation audio; the server is set to obtain multiple audio Process the direct mixing ratio of the target audio signal received by the device to obtain multiple direct mixing ratios; select the target direct mixing ratio from multiple direct mixing ratios; determine the target audio corresponding to the target direct mixing ratio from multiple audio processing devices equipment.
  • a non-volatile storage medium includes a stored program, and when the program is running, the device where the non-volatile storage medium is located is controlled to perform a device selection method .
  • a processor configured to run a program, wherein the device selection method is executed when the program is running.
  • multiple direct mixing ratios are obtained by obtaining the direct mixing ratios of the target audio signals received by multiple first audio devices, and the direct mixing ratio is the direct mixing ratio of the target audio signals received by each audio device.
  • the energy ratio of audio and reverberation audio select the target direct mixing ratio from a plurality of direct mixing ratios; determine the audio equipment corresponding to the target direct mixing ratio from a plurality of first audio equipment, and determine the direct mixing ratio corresponding to the target
  • the audio device is the target audio device.
  • FIG. 1 is a schematic flowchart of a method for determining equipment according to an embodiment of the present application
  • FIG. 2 is a schematic flowchart of a device positioning method according to an embodiment of the present application
  • FIG. 3 is a schematic structural diagram of a system for determining equipment according to an embodiment of the present application.
  • FIG. 4 is a schematic structural diagram of another device determining system according to an embodiment of the present application.
  • Fig. 5a is a schematic flowchart of a device positioning method according to an embodiment of the present application.
  • Fig. 5b is a schematic structural diagram of an audio processing device according to an embodiment of the present application.
  • Fig. 6 is a schematic structural diagram of a server according to an embodiment of the present application.
  • a method embodiment of a device determination method is provided. It should be noted that the steps shown in the flow charts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and, Although a logical order is shown in the flowcharts, in some cases the steps shown or described may be performed in an order different from that shown or described herein.
  • FIG. 1 is a device determination method according to an embodiment of the present disclosure. As shown in FIG. 1, the method includes the following steps:
  • Step S102 obtaining the direct mixing ratios of the target audio signals received by multiple first audio devices to obtain multiple direct mixing ratios, the direct mixing ratios being the target audio signals received by each of the first audio devices The energy ratio of direct audio and reverberant audio in ;
  • the above-mentioned direct audio is the audio signal directly received by each audio device in the audio signal sent by the target sound source
  • the above-mentioned reverberation audio includes reverberation and multipath reflection, which is the audio signal sent by the target sound source.
  • the audio signal received by the audio device after the audio signal is reflected can be measured by the noise field.
  • the second audio device among the plurality of first audio devices is an audio device with a self-play module, wherein the self-play module is set to play audio
  • the second audio device also needs to The received signal is double-talk detection.
  • the second audio device will detect each frame audio signal in the target audio signal received by the second audio device, and judge whether the target audio signal contains an echo signal of the audio played by the self-play module;
  • the audio signal of the target frame contains an echo signal
  • calculate the proportion of the audio signal containing the echo signal in the audio signal of the target frame process the audio signal of the target frame according to the ratio of the audio signal and the preset threshold; determine the target according to the processed audio signal of the target frame An audio signal; wherein, processing the audio signal of the target frame includes: eliminating an echo signal in the audio signal of the target frame or eliminating the audio signal of the target frame.
  • the first audio device will detect whether each frame of audio signal contains the echo signal of the audio played by the self-playing module, and count the proportion of the audio signal containing the echo signal in each frame of audio signal, and compare the ratio with the preset threshold; If the ratio is not smaller than the preset threshold, the echo signal in the audio signal containing the echo signal is eliminated; and in the case of the ratio smaller than the preset threshold, the audio signal containing the echo signal is eliminated.
  • the above preset threshold may be set by the target user.
  • the above-mentioned first audio device or the second audio device may be a refrigerator, a washing machine, a dishwasher, a robot, a mobile terminal and other household appliances or smart devices integrated with an audio collection module.
  • the above-mentioned target audio may be an audio with a wake-up word instruction issued by a target audio source at a certain time point, wherein the target audio source may be a user.
  • the above-mentioned self-playing module is a module set to play audio.
  • the above-mentioned self-broadcasting module can be a stereo, etc.
  • the above-mentioned audio equipment with a self-broadcasting module can be a smart TV, a smart speaker, a mobile phone, etc.
  • the user because the user is not familiar with the wake-up word instruction, or is interrupted by other things, there may be blank parts in the target audio without audio data. Therefore, multiple first audio devices receive the After the target audio sent by the sound source, it is also necessary to eliminate the blank part in the target audio, and use the target audio after eliminating the blank part as the new target audio, wherein the blank part is a part without audio information in the target audio.
  • determining the direct mixing ratio of the target audio received by each first audio device needs to complete the following steps: determine the corresponding target audio signals received by multiple audio collection modules of each first audio device The target frequency domain signal; determine the degree of linear correlation between the frequency domain signals corresponding to any two audio acquisition modules in the plurality of audio acquisition modules; determine the scattering noise field corresponding to each first audio device; according to the linear The degree of correlation and the scattered noise field determine the direct mixing ratio of the target audio signal received by each first audio device.
  • the above-mentioned audio collection module may be a microphone.
  • the specific way to obtain the direct-to-mix ratio of the target audio signals received by multiple first audio devices is as follows: each first audio device receives Extract the multi-frame audio signal from the original audio signal to obtain the target audio signal; respectively calculate the direct-mixing ratios of the multi-frame audio signals in the target audio signal; calculate the multi-frame audio signals according to the respective direct-mixing ratios of the multi-frame audio signals The average value of the direct mixing ratio, and the average value is used as the direct mixing ratio of the target audio signal received by each first audio receiving device.
  • the method for determining the degree of linear correlation between the frequency domain signals corresponding to any two audio collection modules among the multiple audio collection modules is: determining any two audio collection modules among the multiple audio collection modules The cross power spectral density between the corresponding frequency domain signals; determine the self power spectral density of the frequency domain signal corresponding to each audio acquisition module; determine the audio equipment according to the cross power density spectrum and the self power density spectrum corresponding to each audio acquisition module The degree of linear correlation between the frequency domain signals corresponding to any two audio acquisition modules in .
  • determining the degree of linear correlation of the audio collection modules in each audio device includes: determining the interaction between the first audio collection module and the second audio collection module Power spectral density; determine the self-power spectral density of the first audio collection module, and the self-power spectral density of the second audio collection module; according to the cross-spectrum, the self-spectrum of the first audio collection module, and the self-spectrum of the second audio collection module Determine the degree of linear correlation in frequency between the target audio received by the first audio collection module and the second audio collection module.
  • the above linear correlation degree can be measured by a coherence function.
  • the above cross power spectral density is used for the power spectral density between two frequency domain functions.
  • the real part is the co-spectral density (referred to as "co-spectrum")
  • the imaginary part is the orthogonal spectral density.
  • the above self-power spectral density is used to reflect the correlation function in the time domain to express the internal relationship between the random signal itself and other signals at different times.
  • the degree of linear correlation in frequency between the target audio received by the first audio collection module and the second audio collection module can be measured by a coherence function, wherein the coherence function refers to the degree of linear correlation between components of the two processes at each frequency.
  • the expression of the cross-power spectral density between the first audio acquisition module and the second audio acquisition module is as follows:
  • x and y represent the first audio collection module and the second audio collection module respectively, and ⁇ is a smoothing factor.
  • l represents the time information of the audio data received
  • f is the frequency information of the audio data received
  • X (l, f) represents the frequency domain data of the target audio received by the first audio acquisition module
  • Y * (l, f ) represents the conjugate data of the frequency domain data of the target audio received by the second audio acquisition module.
  • the value of ⁇ may be larger, for example, 0.5.
  • the expression of the self-power spectral density of the target audio frequency that the first audio frequency collection module receives is as follows:
  • x represents the first audio collection module
  • is a smoothing factor.
  • l represents the time information of the audio data received
  • f is the frequency information of the audio data received
  • X (l, f) represents the frequency domain data of the target audio received by the first audio acquisition module
  • X * (l, f ) represents the conjugate data of the frequency domain data of the target audio received by the first audio collection module.
  • y represents the second audio collection module
  • is a smoothing factor.
  • l represents the time information of the audio data received
  • f is the frequency information of the audio data received
  • Y(l, f) represents the frequency domain data of the target audio received by the second audio acquisition module
  • Y * (l, f ) represents the conjugate data of the frequency domain data of the target audio received by the second audio acquisition module.
  • the expression of the coherence function between the first audio acquisition module and the second audio acquisition module can be obtained as follows:
  • the value of the coherence function will not change.
  • the first audio collection module and the second audio collection module are audio collection modules on the same audio equipment, the sensitivity of the sensors of the two audio collection modules can be considered the same, that is, the first audio collection module and the second audio collection module
  • the amplification or reduction factor of the received target audio data must also be the same.
  • the method for determining the scattered noise field corresponding to each audio device is: determining the distance between the first audio collection module and the second audio collection module, and the sound velocity of the target audio data; according to the distance and The velocity of sound, to determine the scattered noise field.
  • the expression of the scattering noise field is as follows:
  • R n (f) sinc(2 ⁇ fd/c)
  • sinc is a sinc function
  • f is the frequency information of the target audio
  • d is the distance between two audio acquisition modules
  • c is the speed of sound.
  • the distances between the respective audio collection modules of all the audio devices described in the present application can be the same value, which can further reduce the direct mixing of different devices to calculate the received target audio than the impact.
  • the distances between the respective audio collection modules of different audio devices may also be different values.
  • the distances between different audio modules are different values, it can be seen from the calculation formula of the above-mentioned scattering noise field that for any two audio acquisition modules, when the distances between the audio acquisition modules are different, the scattering noise Field values are different.
  • the distance value difference between different audio collection modules is small, the influence on the calculation result due to the above-mentioned difference in distance value can be ignored.
  • the expression of the coherent relative diffusion ratio CDR can be further obtained:
  • the coherent relative spread ratio can be regarded as the direct-mixing ratio of the components of the target audio at a certain frequency. It can be understood that at each time point, the frequency of the target audio is a continuous range of values. In order to facilitate calculation, it is necessary to determine multiple frequencies from this continuous range of frequency values, and use the multiple frequencies to replace the original target audio data.
  • fl and fh respectively represent the minimum frequency and maximum frequency after sampling the target audio data of this frame.
  • the target audio is also a piece of continuous data in the time domain.
  • the direct mixing ratio of the target audio received by the target audio device can be obtained as:
  • the value range of VAD(l) is 0 and 1
  • the value range of DTD(l) is 0 and 1, which is used to indicate whether there is an echo in the current frame. When there is an echo, the value is 0, and when there is no echo, the value is 1
  • A indicates the target The total number of audio frames, lb and lt represent the first and last frame of the target audio, respectively.
  • the method for determining the direct-mixing ratio of the target audio received by each audio device is: separately calculate the multiple audio collection modules, The degree of linear correlation in frequency between the target audio received by any two audio acquisition modules, and according to the coherence function and the scattering noise field, determine multiple primary direct-mixing ratios; average the multiple primary direct-mixing ratios to obtain The average direct mixing ratio is used as the direct mixing ratio of the target audio device received by the audio device.
  • Step S104 selecting a target direct mixing ratio from multiple direct mixing ratios
  • the direct mix ratio with the largest direct mix ratio is the target direct mix ratio.
  • step S106 an audio device corresponding to the target direct mixing ratio is determined from the plurality of first audio devices, and the audio device corresponding to the target direct mixing ratio is used as the target audio device.
  • the audio device corresponding to the target direct mixing ratio is the audio device closest to the user (that is, the target audio source).
  • the target audio device after the target audio device is determined, the target audio device will be woken up, and will perform an action corresponding to the indication information after receiving the indication information. For example, when the target audio device is a washing machine, after receiving an instruction instruction to wash clothes, the washing machine will start to wash clothes according to the requirements in the instruction instruction.
  • the above-mentioned multiple first audio devices can communicate with each other.
  • a certain audio device can be selected from the plurality of first audio devices randomly or according to preset rules as a device for collecting and comparing various direct mixing ratios, and select the target direct mixing ratio, and the target direct mixing ratio target audio device.
  • the reverberation field is replaced by the scattering noise field to ensure that the references of each device in the distributed wake-up are consistent, and the obtained coherent relative diffusion ratio CDR can represent the ratio of the sound source signal to the reverberation; the reverberation The field is consistent, so the reference of the direct mixing ratio DRR of each device is consistent, so the distance from the sound source can be directly expressed without energy calibration. At the same time, it avoids directly solving the energy mean value, and also weakens the influence of noise and reverberation factors.
  • a method embodiment of a device positioning method is provided. It should be noted that the steps shown in the flow charts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and, Although a logical order is shown in the flowcharts, in some cases the steps shown or described may be performed in an order different from that shown or described herein.
  • Fig. 2 is a device positioning method according to an embodiment of the present disclosure. As shown in Fig. 2, the method includes the following steps:
  • Step S202 receiving target audio from the target audio source, wherein the target audio includes a wake-up instruction
  • the above-mentioned target audio source may be a user, and the target audio may be a wake-up instruction issued by the user.
  • the audio device when the audio device is an audio device with a self-playing module, after the audio device receives the target audio from the target audio source, before calculating the direct mixing ratio of the received target audio, it also needs Perform the following operations: when the audio device receives the target audio, it performs double-talk detection to obtain the detection result.
  • the double-talk detection is used to detect whether the audio device receives the echo of the audio played by the self-play module at the same time when receiving the target audio; when the detection When the result is that the audio device only receives the target audio, the direct-mix ratio of the audio device is calculated; when the detection result is that the audio device receives the target audio, the echo is received at the same time, and the direct-mix ratio of the audio device is not calculated.
  • the above-mentioned self-playing module is a module set to play audio.
  • the above-mentioned self-broadcasting module may be a stereo, etc.
  • the above-mentioned audio equipment with a self-broadcasting module may be a smart TV, a smart speaker, and the like.
  • the user because the user is not familiar with the wake-up word instruction, or is interrupted by other things, etc., there may be blank parts in the target audio without audio data.
  • the target audio sent by the sound source it is also necessary to eliminate the blank part in the target audio, and use the target audio after eliminating the blank part as the new target audio, wherein the blank part is a part without audio information in the target audio.
  • Step S204 determining the direct mixing ratio of each audio device among the plurality of first audio devices receiving the target audio, wherein the direct mixing ratio is the energy ratio of the direct audio received by the audio device and the reverberant audio;
  • determining the direct mixing ratio of the target audio received by each audio device needs to complete the following steps: determine the frequency linearity between the target audio received by the audio collection module in each audio device Correlation degree; determine the scattering noise field corresponding to each audio device; determine the direct mixing ratio of the target audio received by each audio device according to the coherence function and the scattering noise field.
  • the above-mentioned audio collection module may be a microphone.
  • determining the coherence function of the audio collection module in each audio device includes: determining the first audio collection module and the second audio collection module Cross-power spectral density between modules; determine the self-power spectral density of the first audio collection module, and the self-power spectral density of the second audio collection module; according to the cross-spectrum, the self-spectrum of the first audio collection module, and the second audio
  • the autospectrum of the acquisition module determines the degree of linear correlation in frequency between the target audio received by the first audio acquisition module and the second audio acquisition module.
  • the above cross power spectral density is used for the power spectral density between two frequency domain functions.
  • the real part is the co-spectral density (referred to as "co-spectrum")
  • the imaginary part is the orthogonal spectral density.
  • the above self-power spectral density is used to reflect the correlation function in the time domain to express the internal relationship between the random signal itself and other signals at different times.
  • the degree of linear correlation in frequency between the target audio received by the first audio collection module and the second audio collection module can be measured by a coherence function, wherein the coherence function refers to the degree of linear correlation between components of the two processes at each frequency.
  • the expression of the cross-power spectral density between the first audio acquisition module and the second audio acquisition module is as follows:
  • x and y represent the first audio collection module and the second audio collection module respectively, and ⁇ is a smoothing factor.
  • l represents the time information of the audio data received
  • f is the frequency information of the audio data received
  • X (l, f) represents the frequency domain data of the target audio received by the first audio acquisition module
  • Y * (l, f ) represents the conjugate data of the frequency domain data of the target audio received by the second audio acquisition module.
  • the expression of the self-power spectral density of the target audio frequency that the first audio frequency collection module receives is as follows:
  • x represents the first audio collection module
  • is a smoothing factor.
  • l represents the time information of the audio data received
  • f is the frequency information of the audio data received
  • X (l, f) represents the frequency domain data of the target audio received by the first audio acquisition module
  • X * (l, f ) represents the conjugate data of the frequency domain data of the target audio received by the first audio collection module.
  • y represents the second audio collection module
  • is a smoothing factor.
  • l represents the time information of the audio data received
  • f is the frequency information of the audio data received
  • Y(l, f) represents the frequency domain data of the target audio received by the second audio acquisition module
  • Y * (l, f ) represents the conjugate data of the frequency domain data of the target audio received by the second audio collection module.
  • the expression of the coherence function between the first audio acquisition module and the second audio acquisition module can be obtained as follows:
  • the value of the coherence function will not change.
  • the first audio collection module and the second audio collection module are audio collection modules on the same audio equipment, the sensitivity of the sensors of the two audio collection modules can be considered the same, that is, the first audio collection module and the second audio collection module
  • the amplification or reduction factor of the received target audio data must also be the same.
  • the method for determining the scattered noise field corresponding to each audio device is: determining the distance between the first audio collection module and the second audio collection module, and the sound velocity of the target audio data; according to the distance and The velocity of sound, to determine the scattered noise field.
  • the expression of the scattering noise field is as follows:
  • R n (f) sinc(2 ⁇ fd/c)
  • sinc is a sinc function
  • f is the frequency information of the target audio
  • d is the distance between two audio acquisition modules
  • c is the speed of sound.
  • the expression of the coherent relative diffusion ratio CDR can be further obtained:
  • the coherent relative spread ratio can be regarded as the direct-mixing ratio of the components of the target audio at a certain frequency. It can be understood that at each time point, the frequency of the target audio is a continuous range of values. In order to facilitate calculation, it is necessary to determine multiple frequencies from this continuous range of frequency values, and use the multiple frequencies to replace the original target audio data.
  • fl and fh respectively represent the minimum frequency and maximum frequency after sampling the target audio data of this frame.
  • the target audio is also a piece of continuous data in the time domain.
  • the direct mixing ratio of the target audio received by the target audio device can be obtained as:
  • the value range of VAD(l) is 0 and 1
  • the value range of DTD(l) is 0 and 1, which is used to indicate whether there is an echo in the current frame. When there is an echo, the value is 0, and when there is no echo, the value is 1
  • A indicates the target The total number of audio frames, lb and lt represent the first and last frame of the target audio, respectively.
  • the method for determining the direct-mixing ratio of the target audio received by each audio device is: separately calculate the multiple audio collection modules, The degree of linear correlation in frequency between the target audio received by any two audio acquisition modules, and according to the coherence function and the scattering noise field, determine multiple primary direct-mixing ratios; average the multiple primary direct-mixing ratios to obtain The average direct mixing ratio is used as the direct mixing ratio of the target audio device received by the audio device.
  • Step S206 selecting a target direct mixing ratio from multiple direct mixing ratios
  • the direct mix ratio with the largest direct mix ratio is the target direct mix ratio.
  • Step S208 determining an audio device corresponding to the target direct mixing ratio from the plurality of first audio devices, and using the audio device corresponding to the target direct mixing ratio as the target audio device.
  • the audio device corresponding to the target direct mixing ratio is the closest audio device to the user, that is, the target audio source.
  • the device determination system includes: multiple audio processing devices 30 and a server 32, wherein: each of the multiple audio processing devices 30 processes
  • the device 30 is configured to receive the target audio sent by the target audio source, and is configured to calculate the direct mixing ratio of the received target audio, wherein the direct mixing ratio is the energy ratio of the direct audio and the reverberation audio received by the audio device; the server 32, set to obtain the direct mixing ratio of the target audio signal received by multiple audio processing devices 30 to obtain multiple direct mixing ratios; select the target direct mixing ratio from multiple direct mixing ratios; select the target direct mixing ratio from multiple audio processing devices 30 Determine the audio processing device 30 corresponding to the target direct mixing ratio, and determine the audio processing device 30 corresponding to the target direct mixing ratio as the target audio device.
  • the above audio processing device 30 may be, for example, the first audio device in other embodiments.
  • the server 32 may be installed in a certain audio processing device 30 among the multiple audio processing devices 30 .
  • an audio processing device as shown in FIG. , and the communication module 306, wherein:
  • the audio collection module 302 is configured to receive the target audio signal; the processor 304 is configured to calculate the direct mixing ratio of the received target audio signal, and the direct mixing ratio is the direct audio and the mixing ratio in the target audio signal received by the audio processing device 30 and determine whether the audio processing device 30 is the target audio device according to the judgment instruction received by the communication module 306, and wake up the audio processing device 30 when it is determined that the audio processing device 30 is the target audio device, wherein the judgment The instruction is used to indicate whether the audio processing device 30 is the target audio device; the communication module is configured to send the direct mixing ratio to the server 32, and receive a judgment instruction generated by the server according to the direct mixing ratio.
  • the audio processing device 30 may execute the device determination method shown in FIG. 5a. As shown in Figure 5a, the method includes:
  • Step S502 receiving the target audio signal, and calculating the direct mixing ratio of the received target audio signal, where the direct mixing ratio is the energy ratio of the direct audio and the reverberant audio in the target audio signal received by each audio processing device;
  • Step S504 sending the direct mixing ratio to the server
  • Step S506 receiving a judging instruction sent by the server, and determining whether the audio processing device is a target audio device according to the judging instruction, wherein the target audio device is an audio device to be awakened, and the judging instruction is generated according to the direct mixing ratio sent by multiple audio processing devices.
  • the server 32 of the above-mentioned device determination system may be an additional device, such as a hardware device such as a mobile phone, or may be a cloud server.
  • the server 32 is configured to select a target direct mixing ratio from a plurality of direct mixing ratios, and determine the audio processing device 30 corresponding to the target direct mixing ratio from a plurality of audio processing devices 30, and the audio corresponding to the target direct mixing ratio.
  • the processing device 30 serves as a target audio device, wherein the target audio device enters a wake-up mode triggered by a wake-up instruction.
  • the above-mentioned server 32 in addition to the processor 320, also includes a communication module 326 configured to receive the direct mixing ratio sent by each audio device, and a communication module 326 configured to input control instructions.
  • An input module 322, and a display module 324 configured to display device information of each audio device.
  • a non-volatile storage medium includes a stored program.
  • the device where the non-volatile storage medium is located is controlled to perform the following device determination method: obtain The direct mixing ratio of the target audio signal received by a plurality of first audio devices obtains a plurality of direct mixing ratios, and the direct mixing ratio is the energy ratio of direct audio and reverberant audio in the target audio signal received by each audio device; Selecting a target direct mixing ratio from multiple direct mixing ratios; determining an audio device corresponding to the target direct mixing ratio from multiple first audio devices, and determining the audio device corresponding to the target direct mixing ratio as the target audio device.
  • a processor configured to run a program, and execute the following device determination method when the program is running: obtain the direct mixing ratio of target audio signals received by multiple first audio devices, Get a plurality of direct mixing ratios, the direct mixing ratio is the energy ratio of the direct audio and the reverberation audio in the target audio signal received by each audio device; select the target direct mixing ratio from multiple direct mixing ratios; select the target direct mixing ratio from multiple first An audio device corresponding to the target direct mixing ratio is determined in an audio device, and the audio device corresponding to the target direct mixing ratio is determined as the target audio device.
  • the disclosed technical content can be realized in other ways.
  • the device embodiments described above are only illustrative.
  • the division of the units may be a logical function division.
  • multiple units or components may be combined or may be Integrate into another system, or some features may be ignored, or not implemented.
  • the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of units or modules may be in electrical or other forms.
  • the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
  • each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, each unit may exist separately physically, or two or more units may be integrated into one unit.
  • the above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
  • the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
  • the technical solution of the present disclosure is essentially or part of the contribution to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium , including several instructions to make a computer device (which may be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present disclosure.
  • the aforementioned storage media include: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disc, etc., which can store program codes. .

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Measurement Of Velocity Or Position Using Acoustic Or Ultrasonic Waves (AREA)

Abstract

本公开提供了一种设备确定方法及设备确定系统。其中,该方法包括:获取多个第一音频设备接收到的目标音频信号的直混比,得到多个直混比,直混比为每个音频设备接收到的目标音频信号中的直达音频和混响音频的能量比;从多个直混比中选择目标直混比;从多个第一音频设备中确定与目标直混比对应的音频设备,并确定与目标直混比对应的音频设备为目标音频设备。

Description

设备确定方法及设备确定系统
本公开要求于2021年7月26日提交中国专利局、申请号为202110845692.X、发明名称为“设备确定方法及设备确定系统”的中国专利申请的优先权,其全部内容通过引用结合在本公开中。
技术领域
本公开涉及智能家居领域,具体而言,涉及一种设备确定方法及设备确定系统。
背景技术
语音交互的应用经过多年的发展,历经市场变动,目前单个设备的交互已经不再满足需求,同一场景下多个设备的协同唤醒已经逐渐成为主流应用场景。在该应用场景中,由于各个设备的相对位置无法预知,因此只靠DOA估计无法判断声源的具体位置,即利用声波的相位信息只能判断声源的方位,无法得知各个设备与音源的距离;声波在空间传播,幅值逐渐衰减,因此,能量信息是能够表征距离变化的指标。
在现有技术中,同一场景下,当唤醒词指令发出时,各个设备自配的语音交互模块判断唤醒后,每个设备对自己接收的语音信号进行简单的处理:AEC、降噪、去混响等,目的是更准确估计出接收信号中声源信号的能量均值,由于各个设备的麦克风传感器灵敏度不一定一致,大多数解决方案都需要额外做能量校准,即选取某个设备作为参考,在消声室中,声源到被校准设备和参考设备的距离一致时,计算被校准设备的能量与参考设备的比值。在分布式唤醒模块判别时,除参考设备外其他设备所计算的能量值都要乘以对应的比值。
可以看出,依靠能量均值判断设备距离音源的远近,需要额外的能量校准的工作,不仅工作量提升,而且能量校准的精度必须较高;此外,能量均值,还特别容易被混响因素所影响,即当前帧的接收信号包括目标语音信号和历史语音信号,此时计算的均值不能很准确表征距离信息。
针对上述的问题,目前尚未提出有效的解决方案。
发明内容
本公开实施例提供了一种设备确定方法及设备确定系统,以至少解决由于各个设备的传感器灵敏度不一致造成的依据能量均值判断各设备距离目标音源的远近精确度低的技术问题。
根据本公开实施例的一个方面,提供了一种设备确定方法,包括:获取多个第一音频设备接收到的目标音频信号的直混比,直混比为每个音频设备接收到的目标音频信号中的直达音频和混响音频的能量比;从多个直混比中确定目标直混比;从多个第一音频设备中确定与目标直混比对应的目标音频设备。
根据本公开实施例的另一方面,还提供了一种设备确定方法,包括:接收目标音频信号,并计算接收到的目标音频信号的直混比,直混比为第一音频设备接收到的目标音频信号中的直达音频和混响音频的能量比;发送直混比至服务器;接收服务器发送的判断指令,依据判断指令确定第一音频设备是否为目标音频设备,目标音频设备为待唤醒的音频设备,判断指令根据多个第一音频设备发送的直混比生成。
根据本公开实施例的另一方面,还提供了一种音频处理设备,音频处理设备至少包括通信模块,处理器和音频采集模块,其中:音频采集模块,设置为接收目标音频信号;处理器,设置为计算接收到的目标音频信号的直混比,直混比为音频设备接收到的目标音频信号中的直达音频和混响音频的能量比;以及依据通信模块接收的判断指令确定音频处理设备是否为目标音频设备,目标音频设备为待唤醒的音频设备,其中,判断指令根据多个音频处理设备发送的直混比生成;通信模块,设置为将直混比发送至服务器,并接收服务器发送的判断指令。
根据本公开实施例的另一方面,还提供了一种设备确定系统,包括多个音频处理设备和服务器,其中:多个音频处理设备中的每个音频处理设备,设置为接收由目标音源发出的目标音频,以及设置为计算接收到的目标音频的直混比,其中,直混比为每个音频处理设备接收到的直达音频和混响音频的能量比;服务器,设置为获取多个音频处理设备接收到的目标音频信号的直混比,得到多个直混比;从多个直混比中选择目标直混比;从多个音频处理设备中确定与目标直混比对应的目标音频设备。
根据本公开实施例的另一方面,还提供了一种非易失性存储介质,非易失性存储介质包括存储的程序,在程序运行时控制非易失性存储介质所在设备执行设备选择方法。
根据本公开实施例的另一方面,还提供了一种处理器,处理器设置为运行程序,其中,程序运行时执行设备选择方法。
在本公开实施例中,采用获取多个第一音频设备接收到的目标音频信号的直混比,得到多个直混比,直混比为每个音频设备接收到的目标音频信号中的直达音频和混响音频的能量比;从多个直混比中选择目标直混比;从多个第一音频设备中确定与目标直混比对应的音频设备,并确定与目标直混比对应的音频设备为目标音频设备的方式,通过使用直混比来衡量音频设备距离目标音源的远近,达到了消除音频设备的麦克风 传感器灵敏度对距离判断准确度的影响的目的,从而实现了准确判断应当唤醒的音频设备的技术效果,进而解决了由于各个设备的传感器灵敏度不一致造成的依据能量均值判断各设备距离目标音源的远近精确度低技术问题。
附图说明
此处所说明的附图用来提供对本公开的进一步理解,构成本申请的一部分,本公开的示意性实施例及其说明用于解释本公开,并不构成对本公开的不当限定。在附图中:
图1是根据本申请实施例的一种设备确定方法的流程示意图;
图2是根据本申请实施例的一种设备定位方法的流程示意图;
图3是根据本申请实施例的一种设备确定系统的结构示意图;
图4是根据本申请实施例的另一种设备确定系统的结构示意图;
图5a是根据本申请实施例的一种设备定位方法的流程示意图;
图5b是根据本申请实施例的一种音频处理设备的结构示意图;
图6是根据本申请实施例的一种服务器的结构示意图。
具体实施方式
为了使本技术领域的人员更好地理解本公开方案,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本公开一部分的实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都应当属于本公开保护的范围。
需要说明的是,本公开的说明书和权利要求书及上述附图中的术语“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便这里描述的本公开的实施例能够以除了在这里图示或描述的那些以外的顺序实施。此外,术语“包括”和“具有”以及他们的任何变形,意图在于覆盖不排他的包含,例如,包含了一系列步骤或单元的过程、方法、系统、产品或设备不必限于清楚地列出的那些步骤或单元,而是可包括没有清楚地列出的或对于这些过程、方法、产品或设备固有的其它步骤或单元。
实施例1
根据本公开实施例,提供了一种设备确定方法的方法实施例,需要说明的是,在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
图1是根据本公开实施例的设备确定方法,如图1所示,该方法包括如下步骤:
步骤S102,获取多个第一音频设备接收到的目标音频信号的直混比,得到多个直混比,所述直混比为所述每个第一音频设备接收到的所述目标音频信号中的直达音频和混响音频的能量比;
在本申请的一些实施例中,上述直达音频为由目标音源发出的音频信号中直接被各个音频设备接收到的音频信号,上述混响音频中包括混响和多径反射,为目标音源发出的音频信号经过反射后被音频设备接收的音频信号,可以用噪声场来衡量。
在本申请的一些实施例中,当多个第一音频设备中的第二音频设备为带有自播模块的音频设备时,其中,自播模块设置为播放音频,第二音频设备还需要对接收到的信号进行双讲检测。具体地,第二音频设备会对第二音频设备接收到的目标音频信号中的各帧音频信号进行检测,判断目标音频信号中是否包含自播模块播放音频的回音信号;当目标音频信号中的目标帧音频信号包含回音信号时,计算目标帧音频信号中包含回音信号的音频信号比例;根据音频信号比例和预设阈值,对目标帧音频信号进行处理;根据处理后的目标帧音频信号确定目标音频信号;其中,对目标帧音频信号进行处理包括:消除目标帧音频信号中的回音信号或消除目标帧音频信号。
具体地,第一音频设备会检测每一帧音频信号中是否包含自播模块播放音频的回音信号,并统计每一帧音频信号中包含回音信号的音频信号比例,并比较比例和预设阈值;在比例不小于预设阈值的情况下,消除包含回音信号的音频信号中的回音信号;而在比例小于预设阈值的情况下,消除包含回音信号的音频信号。需要说明的是,上述预设阈值可由目标用户自行设定。
在本申请的一些实施例中,上述第一音频设备或第二音频设备可以是集成了音频采集模块的冰箱,洗衣机,洗碗机或者机器人,移动终端等家电或智能设备。
在本申请的一些实施例中,上述目标音频可以为由目标音源在一个确定的时间点发出的带有唤醒词指令的音频,其中目标音源可以为用户。
需要说明的是,上述自播模块为设置为播放音频的模块。具体地,上述自播模块可以为音响等,上述带有自播模块的音频设备可以是智能电视,智能音箱,手机等。
在本申请的一些实施例中,由于用户对唤醒词指令不熟悉,或被其他事情打断等, 导致目标音频中可能存在没有音频数据的空白部分,因此,多个第一音频设备接收由目标音源发出的目标音频后,还需要消除目标音频中的空白部分,并将消除空白部分后的目标音频作为新的目标音频,其中,空白部分为目标音频中无音频信息的部分。
在本申请的一些实施例中,确定每个第一音频设备接收到的目标音频的直混比需要完成以下步骤:确定每个第一音频设备的多个音频采集模块接收到的目标音频信号对应的目标频域信号;确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的线性相关程度;确定每个第一音频设备对应的散射噪声场;依据所述线性相关程度和所述散射噪声场,确定每个第一音频设备接收到的所述目标音频信号的直混比。
在本申请的一些实施例中,上述音频采集模块可以为麦克风。
在本申请的一些实施例中,依据线性相关程度和散射噪声场,获取多个第一音频设备接收到的目标音频信号的直混比的具体方式为:从每个第一音频设备接收到的原始音频信号中提取多帧音频信号,得到目标音频信号;分别计算目标音频信号中的多帧音频信号各自对应的直混比;根据多帧音频信号各自对应的直混比,计算多帧音频信号的直混比的平均值,并将平均值作为每个第一音频接收设备接收到的目标音频信号的直混比。
在本申请的一些实施例中,确定多个音频采集模块中任意两个音频采集模块对应的频域信号之间的线性相关程度的方法为:确定多个音频采集模块中任意两个音频采集模块对应的频域信号之间的互功率谱密度;确定每个音频采集模块对应的频域信号的自功率谱密度;根据互功率密度谱和各音频采集模块对应的自功率密度谱,确定音频设备中任意两个音频采集模块对应的频域信号之间的线性相关程度。
具体地,当音频设备中音频采集模块的数量为两个时,确定每个音频设备中的音频采集模块的线性相关程度,包括:确定第一音频采集模块和第二音频采集模块之间的互功率谱密度;确定第一音频采集模块的自功率谱密度,以及第二音频采集模块的自功率谱密度;依据互谱,第一音频采集模块的自谱,以及第二音频采集模块的自谱确定第一音频采集模块和第二音频采集模块接收到的目标音频之间的频率上的线性相关程度。
在本申请的一些实施例中,上述线性相关程度可以用相干函数来衡量。
需要说明的是,上述互功率谱密度用于两个频域函数之间的功率谱密度。其实部为共谱密度(简称“共谱”),虚部为正交谱密度,上述自功率谱密度用于反映相关函数在时域内表达随机信号自身与其他信号在不同时刻的内在联系,上述第一音频采集模块和第二音频采集模块接收到的目标音频之间的频率上的线性相关程度可以使用相 干函数来衡量,其中,相干函数指两过程在各频率上分量间的线性相关程度。
具体地,第一音频采集模块和第二音频采集模块之间的互功率谱密度的表达式如下:
P xy(l,f)=αP xy(l,f)+(1-α)X(l,f)Y *(l,f)
其中,x和y分别表示第一音频采集模块和第二音频采集模块,α为平滑因子。l表示接收到的音频数据的时间信息,f为接收到的音频数据的频率信息,X(l,f)表示第一音频采集模块接收到的目标音频的频域数据,Y *(l,f)表示第二音频采集模块接收到的目标音频的频域数据的共轭数据。
在本申请的一些实施例中,为了体现直混比变化的瞬时性,α的取值可以大一些,例如,可以取0.5。
第一音频采集模块接收到的目标音频的自功率谱密度的表达式如下:
P x(l,f)=αP x(l,f)+(1-α)X(l,f)X *(l,f)
其中,x表示第一音频采集模块,α为平滑因子。l表示接收到的音频数据的时间信息,f为接收到的音频数据的频率信息,X(l,f)表示第一音频采集模块接收到的目标音频的频域数据,X *(l,f)表示第一音频采集模块接收到的目标音频的频域数据的共轭数据。
第二音频采集模块接收到的目标音频的自功率谱密度的表达式如下:
P y(l,f)=αP y(l,f)+(1-α)Y(l,f)Y *(l,f)
其中,y表示第二音频采集模块,α为平滑因子。l表示接收到的音频数据的时间信息,f为接收到的音频数据的频率信息,Y(l,f)表示第二音频采集模块接收到的目标音频的频域数据,Y *(l,f)表示第二音频采集模块接收到的目标音频的频域数据的共轭数据。
依据上述互功率谱密度和自功率谱密度,可得第一音频采集模块和第二音频采集模块之间的相干函数的表达式如下:
Figure PCTCN2022096393-appb-000001
从上述表达式中可以看出,当第一音频采集模块和第二音频采集模块对接收到的目标音频放大或缩小相同的倍数时,相干函数的值不会发生改变。而由于第一音频采集模块和第二音频采集模块是同一音频设备上的音频采集模块,因此两个音频采集模块的传感器的灵敏度可认为相同,即第一音频采集模块和第二音频采集模块对接收到的目标音频数据的放大或缩小倍数也一定相同。
在本申请的一些实施例中,确定每个音频设备对应的散射噪声场的方法为:确定第一音频采集模块和第二音频采集模块之间的距离,以及目标音频数据的音速;依据距离和音速,确定散射噪声场。
具体地,散射噪声场表达式如下:
R n(f)=sinc(2πfd/c)
上式中,sinc为sinc函数,f为目标音频的频率信息,d为两个音频采集模块之间的距离,c为声速。
在本申请的一些实施例中,本申请中所述的全部音频设备各自的音频采集模块之间的距离可以为相同的值,这样可以进一步降低不同设备自身对计算接收到的目标音频的直混比的影响。
在本申请的一些实施例中,不同的音频设备各自的音频采集模块之间的距离也可以为不同值。当不同的音频模块之间的距离为不同值时,从上述散射噪声场的计算公式中可以看出,对任意两个音频采集模块而言,当音频采集模块之间的距离不同时,散射噪声场的值不同。但由于不同音频采集模块之间的距离取值差异较小,因此可以忽略由于上述距离值不同对计算结果的影响。
在得到了相干函数和散射噪声场的表达式后,可以进一步得到相干相对扩散比CDR的表达式:
Figure PCTCN2022096393-appb-000002
其中,上式中的Re{}和含义为取实部进行下一步的计算。
在本申请的一些实施例中,相干相对扩散比可以认为是目标音频在某个频率上的分量的直混比。可以理解的,在每个时间点上,目标音频的频率都是一段连续的取值范围,为了便于计算,需要从这一段连续的频率取值范围中确定多个频率,并用所述多个频率来替代原来的目标音频数据。
进一步地,可得当前帧的直混比
Figure PCTCN2022096393-appb-000003
其中,fl和fh分别表示在这一帧的目标音频数据采样后的最小的频率和最大的频率。
可以理解地,目标音频在时域上也是一段连续的数据,为了便于计算,同样需要将目标音频数据在时域上由一段连续的时间变为多个时间点,其中,每个时间点可以用一帧来代指。
综上所述,可得目标音频设备接收到的目标音频的直混比为:
Figure PCTCN2022096393-appb-000004
其中,其中,VAD(l)的取值范围为0和1,用于表示当前帧是否为空白帧,并在判定当前帧为空白帧时取值0,判定当前帧为非空白帧时取值为1;DTD(l)的取值范围为0和1,用于表示当前帧中是否存在回声,当存在回声时,取值为0,当不存在回声时,取值为1;A表示目标音频的总帧数,lb和lt分别表示目标音频的第一帧和最后一帧。
在本申请的一些实施例中,当音频设备中音频采集模块的数量大于两个时,确定每个音频设备接收到的目标音频的直混比的方法为:分别计算多个音频采集模块中,任意两个音频采集模块接收到的目标音频之间的频率上的线性相关程度,并依据相干函数和散射噪声场,确定多个初级直混比;对多个初级直混比取平均值,获取平均直混比,并将平均直混比作为音频设备接收到的目标音频设备的直混比。
步骤S104,从多个直混比中选择目标直混比;
具体地,直混比的比值最大的直混比即为目标直混比。
步骤S106,从多个第一音频设备中确定与目标直混比对应的音频设备,并与目标直混比对应的音频设备作为目标音频设备。
在本申请的一些实施例中,与目标直混比对应的音频设备即为距离用户(也就是目标音源)最近的音频设备。
在本申请的一些实施例中,在确定了目标音频设备后,目标音频设备会被唤醒,并在接收到指示信息后执行与指示信息对应的动作。例如,当目标音频设备为洗衣机时,在接收到指示清洗衣物的指示指令后,洗衣机会按照指示指令中的要求开始清洗衣物。
在本申请的一些实施例中,上述多个第一音频设备之间是可以互相通信的。这样可以随机或按照预设规则的从所述多个第一音频设备中选择某个音频设备作为收集并比较各个直混比的设备,并选择出目标直混比,以及与目标直混比对应的目标音频设备。
通过上述步骤,可以实现准确判断应当唤醒的音频设备的技术效果,进而解决了由于各个设备的传感器灵敏度不一致造成的依据能量均值判断各设备距离目标音源的远近精确度低技术问题。
另外,本申请实施例中用散射噪声场来替代混响场,保证分布式唤醒中各个设备 的参考一致,求得的相干相对扩散比率CDR,可以表征声源信号相对混响的比例;混响场一致,因此各个设备的直混比DRR的参考是一致的,因此无需再做能量校准便可直接表示与声源的距离。同时,避免直接求解能量均值,还减弱了噪声、混响的因素影响。
实施例2
根据本公开实施例,提供了一种设备定位方法的方法实施例,需要说明的是,在附图的流程图示出的步骤可以在诸如一组计算机可执行指令的计算机系统中执行,并且,虽然在流程图中示出了逻辑顺序,但是在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤。
图2是根据本公开实施例的设备定位方法,如图2所示,该方法包括如下步骤:
步骤S202,接收由目标音源发出的目标音频,其中,目标音频中包括唤醒指令;
在本申请的一些实施例中,上述目标音源可以是用户,目标音频可以为用户发出的唤醒指令。
在本申请的一些实施例中,当音频设备为带有自播模块的音频设备时,音频设备接收由目标音源发出的目标音频后,在计算接收到的目标音频的直混比之前,还需要进行以下操作:音频设备在接收目标音频时,进行双讲检测,得到检测结果,其中,双讲检测用于检测音频设备在接收目标音频时是否同时接收了自播模块播放音频的回音;当检测结果为音频设备仅接收了目标音频时,计算音频设备的直混比;当检测结果为音频设备接收了目标音频时,同时接收了回音,不计算音频设备的直混比。
需要说明的是,上述自播模块为设置为播放音频的模块。具体地,上述自播模块可以为音响等,上述带有自播模块的音频设备可以是智能电视,智能音箱等。
在本申请的一些实施例中,由于用户对唤醒词指令不熟悉,或被其他事情打断等,导致目标音频中可能存在没有音频数据的空白部分,因此,多个第一音频设备接收由目标音源发出的目标音频后,还需要消除目标音频中的空白部分,并将消除空白部分后的目标音频作为新的目标音频,其中,空白部分为目标音频中无音频信息的部分。
步骤S204,确定接收到目标音频的多个第一音频设备中各个音频设备的直混比,其中,直混比为音频设备接收到的直达音频和混响音频的能量比;
在本申请的一些实施例中,确定每个音频设备接收到的目标音频的直混比需要完成以下步骤:确定每个音频设备中的音频采集模块接收到的目标音频之间的频率上的线性相关程度;确定每个音频设备对应的散射噪声场;依据相干函数和散射噪声场,确定每个音频设备接收到的目标音频的直混比。
在本申请的一些实施例中,上述音频采集模块可以为麦克风。
在本申请的一些实施例中,当音频设备中音频采集模块的数量为两个时,确定每个音频设备中的音频采集模块的相干函数,包括:确定第一音频采集模块和第二音频采集模块之间的互功率谱密度;确定第一音频采集模块的自功率谱密度,以及第二音频采集模块的自功率谱密度;依据互谱,第一音频采集模块的自谱,以及第二音频采集模块的自谱确定第一音频采集模块和第二音频采集模块接收到的目标音频之间的频率上的线性相关程度。
需要说明的是,上述互功率谱密度用于两个频域函数之间的功率谱密度。其实部为共谱密度(简称“共谱”),虚部为正交谱密度,上述自功率谱密度用于反映相关函数在时域内表达随机信号自身与其他信号在不同时刻的内在联系,上述第一音频采集模块和第二音频采集模块接收到的目标音频之间的频率上的线性相关程度可以使用相干函数来衡量,其中,相干函数指两过程在各频率上分量间的线性相关程度。
具体地,第一音频采集模块和第二音频采集模块之间的互功率谱密度的表达式如下:
P xy(l,f)=αP xy(l,f)+(1-α)X(l,f)Y *(l,f)
其中,x和y分别表示第一音频采集模块和第二音频采集模块,α为平滑因子。l表示接收到的音频数据的时间信息,f为接收到的音频数据的频率信息,X(l,f)表示第一音频采集模块接收到的目标音频的频域数据,Y *(l,f)表示第二音频采集模块接收到的目标音频的频域数据的共轭数据。
第一音频采集模块接收到的目标音频的自功率谱密度的表达式如下:
P x(l,f)=αP x(l,f)+(1-α)X(l,f)X *(l,f)
其中,x表示第一音频采集模块,α为平滑因子。l表示接收到的音频数据的时间信息,f为接收到的音频数据的频率信息,X(l,f)表示第一音频采集模块接收到的目标音频的频域数据,X *(l,f)表示第一音频采集模块接收到的目标音频的频域数据的共轭数据。
第二音频采集模块接收到的目标音频的自功率谱密度的表达式如下:
P y(l,f)=αP y(l,f)+(1-α)Y(l,f)Y *(l,f)
其中,y表示第二音频采集模块,α为平滑因子。l表示接收到的音频数据的时间信息,f为接收到的音频数据的频率信息,Y(l,f)表示第二音频采集模块接收到的目标音频的频域数据,Y *(l,f)表示第二音频采集模块接收到的目标音频的频域数据的共轭数据。
依据上述互功率谱密度和自功率谱密度,可得第一音频采集模块和第二音频采集模块之间的相干函数的表达式如下:
Figure PCTCN2022096393-appb-000005
从上述表达式中可以看出,当第一音频采集模块和第二音频采集模块对接收到的目标音频放大或缩小相同的倍数时,相干函数的值不会发生改变。而由于第一音频采集模块和第二音频采集模块是同一音频设备上的音频采集模块,因此两个音频采集模块的传感器的灵敏度可认为相同,即第一音频采集模块和第二音频采集模块对接收到的目标音频数据的放大或缩小倍数也一定相同。
在本申请的一些实施例中,确定每个音频设备对应的散射噪声场的方法为:确定第一音频采集模块和第二音频采集模块之间的距离,以及目标音频数据的音速;依据距离和音速,确定散射噪声场。
具体地,散射噪声场表达式如下:
R n(f)=sinc(2πfd/c)
上式中,sinc为sinc函数,f为目标音频的频率信息,d为两个音频采集模块之间的距离,c为声速。
在得到了相干函数和散射噪声场的表达式后,可以进一步得到相干相对扩散比CDR的表达式:
Figure PCTCN2022096393-appb-000006
其中,上式中的Re{}和含义为取实部进行下一步的计算。
在本申请的一些实施例中,相干相对扩散比可以认为是目标音频在某个频率上的分量的直混比。可以理解的,在每个时间点上,目标音频的频率都是一段连续的取值范围,为了便于计算,需要从这一段连续的频率取值范围中确定多个频率,并用所述多个频率来替代原来的目标音频数据。
进一步地,可得当前帧的直混比
Figure PCTCN2022096393-appb-000007
其中,fl和fh分别表示在这一帧的目标音频数据采样后的最小的频率和最大的频率。
可以理解地,目标音频在时域上也是一段连续的数据,为了便于计算,同样需要将目标音频数据在时域上由一段连续的时间变为多个时间点,其中,每个时间点可以用一帧来代指。
综上所述,可得目标音频设备接收到的目标音频的直混比为:
Figure PCTCN2022096393-appb-000008
其中,其中,VAD(l)的取值范围为0和1,用于表示当前帧是否为空白帧,并在判定当前帧为空白帧时取值0,判定当前帧为非空白帧时取值为1;DTD(l)的取值范围为0和1,用于表示当前帧中是否存在回声,当存在回声时,取值为0,当不存在回声时,取值为1;A表示目标音频的总帧数,lb和lt分别表示目标音频的第一帧和最后一帧。
在本申请的一些实施例中,当音频设备中音频采集模块的数量大于两个时,确定每个音频设备接收到的目标音频的直混比的方法为:分别计算多个音频采集模块中,任意两个音频采集模块接收到的目标音频之间的频率上的线性相关程度,并依据相干函数和散射噪声场,确定多个初级直混比;对多个初级直混比取平均值,获取平均直混比,并将平均直混比作为音频设备接收到的目标音频设备的直混比。
步骤S206,从多个直混比中选择目标直混比;
具体地,直混比的比值最大的直混比即为目标直混比。
步骤S208,从多个第一音频设备中确定与目标直混比对应的音频设备,并将与目标直混比对应的音频设备作为目标音频设备。
在本申请的一些实施例中,与目标直混比对应的音频设备即为距离用户,也就是目标音源,最近的音频设备。
实施例3
根据本公开实施例,提供了一种设备确定系统,如图3所示,该设备确定系统包括:多个音频处理设备30和服务器32,其中:多个音频处理设备30中的每个音频处理设备30,设置为接收由目标音源发出的目标音频,以及设置为计算接收到的目标音频的直混比,其中,直混比为音频设备接收到的直达音频和混响音频的能量比;服务器32,设置为获取多个音频处理设备30接收到的目标音频信号的直混比,得到多个直混比;从多个直混比中选择目标直混比;从多个音频处理设备30中确定与目标直混比对应的音频处理设备30,并确定与目标直混比对应的音频处理设备30为目标音频设备。
需要说明的是,上述音频处理设备30例如可以是其他实施例中的第一音频设备。
可选地,如图3所示,服务器32可以安装在多个音频处理设备30中的某个音频处理设备30中。
在本申请的一些实施例中,还提供了一种如图5b所示的音频处理设备,其中,该音频处理设备30即为第一音频处理设备,包括多个音频采集模块302,处理器304,以及通信模块306,其中:
音频采集模块302,设置为接收目标音频信号;处理器304,设置为计算接收到的目标音频信号的直混比,直混比为音频处理设备30接收到的目标音频信号中的直达音频和混响音频的能量比;以及依据通信模块306接收的判断指令确定音频处理设备30是否为目标音频设备,并在确定音频处理设备30为目标音频设备的情况下,唤醒音频处理设备30,其中,判断指令用于指示音频处理设备30是否为目标音频设备;通信模块,设置为将直混比发送至服务器32,并接收服务器依据直混比生成的判断指令。
在本申请的一些实施例中,上述音频处理设备30可以执行如图5a所示的设备确定方法。如图5a所示,该方法包括:
步骤S502,接收目标音频信号,并计算接收到的目标音频信号的直混比,直混比为每个音频处理设备接收到的目标音频信号中的直达音频和混响音频的能量比;
步骤S504,发送直混比至服务器;
步骤S506,接收服务器发送的判断指令,依据判断指令确定音频处理设备是否为目标音频设备,其中,目标音频设备为待唤醒的音频设备,判断指令根据多个音频处理设备发送的直混比生成。
在本申请的一些实施例中,如图4所示,上述设备确定系统的服务器32可以为一个额外的设备,如手机等硬件设备,也可以为云端服务器。所述服务器32设置为从多个直混比中选择目标直混比,以及从多个音频处理设备30中确定与目标直混比对应的音频处理设备30,并与目标直混比对应的音频处理设备30作为目标音频设备,其中,目标音频设备在唤醒指令的触发下进入唤醒模式。
在本申请的一些实施例中,上述服务器32如图6所示,除处理器320外,还包括设置为接收各个音频设备发来的直混比的通讯模块326,以及设置为输入控制指令的输入模块322,和设置为展示各个音频设备的设备信息的展示模块324。
根据本公开实施例,还提供了一种非易失性存储介质,非易失性存储介质包括存储的程序,在程序运行时控制非易失性存储介质所在设备执行下述设备确定方法:获取多个第一音频设备接收到的目标音频信号的直混比,得到多个直混比,直混比为每个音频设备接收到的目标音频信号中的直达音频和混响音频的能量比;从多个直混比中选择目标直混比;从多个第一音频设备中确定与目标直混比对应的音频设备,并确定与目标直混比对应的音频设备为目标音频设备。
根据本公开实施例,还提供了一种处理器,处理器设置为运行程序,在程序运行时执行下述设备确定方法:获取多个第一音频设备接收到的目标音频信号的直混比,得到多个直混比,直混比为每个音频设备接收到的目标音频信号中的直达音频和混响音频的能量比;从多个直混比中选择目标直混比;从多个第一音频设备中确定与目标直混比对应的音频设备,并确定与目标直混比对应的音频设备为目标音频设备。
上述本公开实施例序号仅仅为了描述,不代表实施例的优劣。
在本公开的上述实施例中,对各个实施例的描述都各有侧重,某个实施例中没有详述的部分,可以参见其他实施例的相关描述。
在本申请所提供的几个实施例中,应该理解到,所揭露的技术内容,可通过其它的方式实现。其中,以上所描述的装置实施例仅仅是示意性的,例如所述单元的划分,可以为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,单元或模块的间接耦合或通信连接,可以是电性或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本公开各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本公开的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可为个人计算机、服务器或者网络设备等)执行本公开各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、移动硬盘、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述仅是本公开的优选实施方式,应当指出,对于本技术领域的普通技术人员来说,在不脱离本公开原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也应视为本公开的保护范围。

Claims (18)

  1. 一种设备确定方法,包括:
    获取多个第一音频设备接收到的目标音频信号的直混比,所述直混比为所述每个第一音频设备接收到的所述目标音频信号中的直达音频和混响音频的能量比;
    从所述多个直混比中确定目标直混比;
    从所述多个第一音频设备中确定与所述目标直混比对应的目标音频设备。
  2. 根据权利要求1所述的方法,其中,每个第一音频设备包括多个音频采集模块,所述获取多个第一音频设备接收到的目标音频信号的直混比,包括:
    确定每个第一音频设备的多个音频采集模块接收到的目标音频信号对应的目标频域信号;
    确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的线性相关程度;
    确定每个第一音频设备对应的散射噪声场;
    依据所述线性相关程度和所述散射噪声场,确定每个第一音频设备接收到的所述目标音频信号的直混比。
  3. 根据权利要求2所述的方法,其中,所述确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的线性相关程度,包括:
    确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的互功率谱密度;
    确定每个音频采集模块对应的频域信号的自功率谱密度;
    根据所述互功率密度谱和各音频采集模块对应的自功率密度谱,确定所述音频设备中所述任意两个音频采集模块对应的频域信号之间的线性相关程度。
  4. 根据权利要求1-3任一项所述的方法,其中,所述获取多个第一音频设备接收到的目标音频信号的直混比,包括:
    从每个第一音频设备接收到的原始音频信号中提取多帧音频信号,得到目标音频信号;
    分别计算所述目标音频信号中的多帧音频信号各自对应的直混比;
    根据所述多帧音频信号各自对应的直混比,计算所述多帧音频信号的直混比的平均值,并将所述平均值作为每个第一音频接收设备接收到的所述目标音频信 号的直混比。
  5. 根据权利要求1-4任一项所述的方法,其中,当所述多个第一音频设备中的第二音频设备为带有自播模块的音频设备时,其中,所述自播模块设置为播放音频,所述方法还包括:
    对所述第二音频设备接收到的所述目标音频信号中的各帧音频信号进行检测,判断所述目标音频信号中是否包含所述自播模块播放音频的回音信号;
    当所述目标音频信号中的目标帧音频信号包含所述回音信号时,计算所述目标帧音频信号中包含所述回音信号的音频信号比例;
    根据所述音频信号比例和预设阈值,对所述目标帧音频信号进行处理;
    根据处理后的目标帧音频信号确定所述目标音频信号;
    其中,对所述目标帧音频信号进行处理包括:消除所述目标帧音频信号中的回音信号或消除所述目标帧音频信号。
  6. 根据权利要求1-5任一项所述的方法,其中,确定所述目标音频设备之后,所述方法还包括:
    控制所述目标音频设备进入唤醒模式,其中,所述目标音频设备在唤醒模式中用于接收指示信息并执行与所述指示信息对应的动作。
  7. 一种设备确定方法,包括:
    接收目标音频信号,并计算接收到的所述目标音频信号的直混比,所述直混比为第一音频设备接收到的所述目标音频信号中的直达音频和混响音频的能量比;
    发送所述直混比至服务器;
    接收服务器发送的判断指令,依据所述判断指令确定所述第一音频设备是否为目标音频设备,所述目标音频设备为待唤醒的音频设备,所述判断指令根据多个第一音频设备发送的所述直混比生成。
  8. 一种音频处理设备,所述音频处理设备至少包括通信模块,处理器和音频采集模块,其中:
    所述音频采集模块,设置为接收目标音频信号;
    所述处理器,设置为计算接收到的所述目标音频信号的直混比,所述直混比为所述音频处理设设备接收到的所述目标音频信号中的直达音频和混响音频的能量比;以及依据所述通信模块接收的判断指令确定所述音频处理设备是否为目标 音频设备,所述目标音频设备为待唤醒的音频设备,其中,所述判断指令根据多个音频处理设备发送的所述直混比生成;
    所述通信模块,设置为将所述直混比发送至服务器,并接收所述服务器发送的所述判断指令。
  9. 根据权利要求8所述的设备,其中,所述音频采集模块的数量为多个,所述处理器设置为通过如下方式计算接收到的所述目标音频信号的直混比:
    确定所述多个音频采集模块接收到的目标音频信号对应的目标频域信号;
    确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的线性相关程度;
    确定所述音频处理设备对应的散射噪声场;
    依据所述线性相关程度和所述散射噪声场,确定接收到的所述目标音频信号的直混比。
  10. 根据权利要求9所述的设备,其中,所述处理器设置为通过如下方式确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的线性相关程度:
    确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的互功率谱密度;
    确定每个音频采集模块对应的频域信号的自功率谱密度;
    根据所述互功率密度谱和各音频采集模块对应的自功率密度谱,确定所述音频设备中所述任意两个音频采集模块对应的频域信号之间的线性相关程度。
  11. 一种设备确定系统,包括多个音频处理设备和服务器,其中:
    多个音频处理设备中的每个音频处理设备,设置为接收由目标音源发出的目标音频,以及设置为计算接收到的所述目标音频的直混比,其中,所述直混比为每个音频处理设备接收到的直达音频和混响音频的能量比;
    服务器,设置为获取多个音频处理设备接收到的目标音频信号的直混比,得到多个直混比;从所述多个直混比中选择目标直混比;从所述多个音频处理设备中确定与所述目标直混比对应的目标音频设备。
  12. 根据权利要求11所述的系统,其中,每个所述音频设备包括多个音频采集模块,所述处理器设置为通过如下方式获取多个音频处理设备接收到的目标音频信号的直混比:
    确定每个所述音频设备的多个音频采集模块接收到的目标音频信号对应的目标频域信号;
    确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的线性相关程度;
    确定每个所述音频设备对应的散射噪声场;
    依据所述线性相关程度和所述散射噪声场,确定每个所述音频设备接收到的所述目标音频信号的直混比。
  13. 根据权利要求12所述的系统,其中,所述处理器设置为通过如下方式确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的线性相关程度:
    确定所述多个音频采集模块中任意两个音频采集模块对应的频域信号之间的互功率谱密度;
    确定每个音频采集模块对应的频域信号的自功率谱密度;
    根据所述互功率密度谱和各音频采集模块对应的自功率密度谱,确定所述音频设备中所述任意两个音频采集模块对应的频域信号之间的线性相关程度。
  14. 根据权利要求11至13中任一项所述的系统,其中,所述服务器设置为通过如下方式获取多个音频处理设备接收到的目标音频信号的直混比:
    从每个音频设备接收到的原始音频信号中提取多帧音频信号,得到目标音频信号;
    分别计算所述目标音频信号中的多帧音频信号各自对应的直混比;
    根据所述多帧音频信号各自对应的直混比,计算所述多帧音频信号的直混比的平均值,并将所述平均值作为每个音频接收设备接收到的所述目标音频信号的直混比。
  15. 根据权利要求11至14中任一项所述的系统,其中,当所述多个音频设备中的第二音频设备为带有自播模块的音频设备时,其中,所述自播模块设置为播放音频,所述处理器还设置为:
    对所述第二音频设备接收到的所述目标音频信号中的各帧音频信号进行检测,判断所述目标音频信号中是否包含所述自播模块播放音频的回音信号;
    当所述目标音频信号中的目标帧音频信号包含所述回音信号时,计算所述目标帧音频信号中包含所述回音信号的音频信号比例;
    根据所述音频信号比例和预设阈值,对所述目标帧音频信号进行处理;
    根据处理后的目标帧音频信号确定所述目标音频信号;
    其中,对所述目标帧音频信号进行处理包括:消除所述目标帧音频信号中的回音信号或消除所述目标帧音频信号。
  16. 根据权利要求11-15任一项所述的系统,其中,所述服务器还设置为在确定所述目标音频设备之后,控制所述目标音频设备进入唤醒模式,其中,所述目标音频设备在唤醒模式中用于接收指示信息并执行与所述指示信息对应的动作。
  17. 一种非易失性存储介质,所述非易失性存储介质包括存储的程序,其中,在所述程序运行时控制所述非易失性存储介质所在设备执行权利要求1至6中任意一项所述设备确定方法。
  18. 一种处理器,所述处理器设置为运行程序,其中,所述程序运行时执行权利要求1至6中任意一项所述设备确定方法。
PCT/CN2022/096393 2021-07-26 2022-05-31 设备确定方法及设备确定系统 Ceased WO2023005409A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202110845692.XA CN113674761B (zh) 2021-07-26 2021-07-26 设备确定方法及设备确定系统
CN202110845692.X 2021-07-26

Publications (1)

Publication Number Publication Date
WO2023005409A1 true WO2023005409A1 (zh) 2023-02-02

Family

ID=78540169

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/096393 Ceased WO2023005409A1 (zh) 2021-07-26 2022-05-31 设备确定方法及设备确定系统

Country Status (2)

Country Link
CN (1) CN113674761B (zh)
WO (1) WO2023005409A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120412649A (zh) * 2025-07-02 2025-08-01 宁波蛙声科技有限公司 基于音视频联合的发言人实时追踪定位方法及系统

Families Citing this family (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11950252B2 (en) * 2020-07-02 2024-04-02 Qualcomm Incorporated Early termination of uplink communication repetitions with multiple transport blocks
CN113674761B (zh) * 2021-07-26 2023-07-21 青岛海尔科技有限公司 设备确定方法及设备确定系统
CN115166632B (zh) * 2022-06-20 2025-02-11 青岛海尔科技有限公司 声源朝向的确定方法和装置、存储介质及电子装置
CN117292691A (zh) * 2022-06-20 2023-12-26 青岛海尔科技有限公司 一种音频能量分析方法和相关装置
CN116206618B (zh) * 2022-12-29 2024-11-19 海尔优家智能科技(北京)有限公司 设备唤醒方法、存储介质及电子装置
CN118365506B (zh) * 2024-06-18 2024-11-19 北京象帝先计算技术有限公司 Mmu配置方法、图形处理系统、电子组件及设备

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107102296A (zh) * 2017-04-27 2017-08-29 大连理工大学 一种基于分布式麦克风阵列的声源定位系统
CN110288997A (zh) * 2019-07-22 2019-09-27 苏州思必驰信息科技有限公司 用于声学组网的设备唤醒方法及系统
CN110364161A (zh) * 2019-08-22 2019-10-22 北京小米智能科技有限公司 响应语音信号的方法、电子设备、介质及系统
CN112599126A (zh) * 2020-12-03 2021-04-02 海信视像科技股份有限公司 一种智能设备的唤醒方法、智能设备及计算设备
CN112634890A (zh) * 2020-12-17 2021-04-09 北京百度网讯科技有限公司 用于唤醒播放设备的方法、装置、设备以及存储介质
CN113160842A (zh) * 2021-03-06 2021-07-23 西安电子科技大学 一种基于mclp的语音去混响方法及系统
CN113488031A (zh) * 2021-06-30 2021-10-08 青岛海尔科技有限公司 确定电子设备的方法、装置、存储介质及电子装置
CN113674761A (zh) * 2021-07-26 2021-11-19 青岛海尔科技有限公司 设备确定方法及设备确定系统

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111025233B (zh) * 2019-11-13 2023-09-15 阿里巴巴集团控股有限公司 一种声源方向定位方法和装置、语音设备和系统
CN111179909B (zh) * 2019-12-13 2023-01-10 航天信息股份有限公司 一种多麦远场语音唤醒方法及系统
CN111081246B (zh) * 2019-12-24 2022-06-24 北京达佳互联信息技术有限公司 直播机器人唤醒方法、装置、电子设备及存储介质

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107102296A (zh) * 2017-04-27 2017-08-29 大连理工大学 一种基于分布式麦克风阵列的声源定位系统
CN110288997A (zh) * 2019-07-22 2019-09-27 苏州思必驰信息科技有限公司 用于声学组网的设备唤醒方法及系统
CN110364161A (zh) * 2019-08-22 2019-10-22 北京小米智能科技有限公司 响应语音信号的方法、电子设备、介质及系统
CN112599126A (zh) * 2020-12-03 2021-04-02 海信视像科技股份有限公司 一种智能设备的唤醒方法、智能设备及计算设备
CN112634890A (zh) * 2020-12-17 2021-04-09 北京百度网讯科技有限公司 用于唤醒播放设备的方法、装置、设备以及存储介质
CN113160842A (zh) * 2021-03-06 2021-07-23 西安电子科技大学 一种基于mclp的语音去混响方法及系统
CN113488031A (zh) * 2021-06-30 2021-10-08 青岛海尔科技有限公司 确定电子设备的方法、装置、存储介质及电子装置
CN113674761A (zh) * 2021-07-26 2021-11-19 青岛海尔科技有限公司 设备确定方法及设备确定系统

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120412649A (zh) * 2025-07-02 2025-08-01 宁波蛙声科技有限公司 基于音视频联合的发言人实时追踪定位方法及系统

Also Published As

Publication number Publication date
CN113674761B (zh) 2023-07-21
CN113674761A (zh) 2021-11-19

Similar Documents

Publication Publication Date Title
CN113674761B (zh) 设备确定方法及设备确定系统
CN109597022B (zh) 声源方位角运算、定位目标音频的方法、装置和设备
JP5998306B2 (ja) 部屋寸法推定の決定
JP7352740B2 (ja) 風雑音減衰のための方法及び装置
CN114830681B (zh) 用于减少环境噪声补偿系统中的误差的方法
CN109218535A (zh) 智能调节音量的方法、装置、存储介质及终端
US9743211B2 (en) Method and apparatus for determining a position of a microphone
CN112684413B (zh) 声源寻向方法和xr设备
CN113766073A (zh) 会议系统中的啸叫检测
CN106157967A (zh) 脉冲噪声抑制
US20160091604A1 (en) Apparatus, system and method for space status detection based on acoustic signal
CN110534129A (zh) 干声和环境声音的分离
KR101882423B1 (ko) 적어도 제1 쌍의 룸 임펄스 응답에 기초하여, 믹싱 시간 전체를 추정하는 장치 및 방법, 대응하는 컴퓨터 프로그램
US10845479B1 (en) Movement and presence detection systems and methods using sonar
WO2020043037A1 (zh) 语音转录设备、系统、方法、及电子设备
CN107743704A (zh) 使用跨装置麦克风的移动装置环境检测
CN111415678B (zh) 对移动设备或可穿戴设备进行开放或封闭空间环境分类
CN107340864A (zh) 一种基于声波的虚拟输入方法
JP5395399B2 (ja) 携帯端末、拍位置推定方法および拍位置推定プログラム
CN112397082A (zh) 估计回声延迟的方法、装置、电子设备和存储介质
Defrance et al. Finding the onset of a room impulse response: Straightforward?
CN116206618B (zh) 设备唤醒方法、存储介质及电子装置
JP4968397B1 (ja) トイレ装置、ゲーム装置、プログラム及びコンピュータ読み取り可能な記録媒体
CN113055785B (zh) 音量调节方法、系统和装置
CN108540904B (zh) 一种改善音箱音效的方法和装置

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22848019

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 22848019

Country of ref document: EP

Kind code of ref document: A1