WO2020150864A1 - 音频信号检测方法、装置及音视频会议系统 - Google Patents
音频信号检测方法、装置及音视频会议系统 Download PDFInfo
- Publication number
- WO2020150864A1 WO2020150864A1 PCT/CN2019/072543 CN2019072543W WO2020150864A1 WO 2020150864 A1 WO2020150864 A1 WO 2020150864A1 CN 2019072543 W CN2019072543 W CN 2019072543W WO 2020150864 A1 WO2020150864 A1 WO 2020150864A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- howling
- audio
- tone
- mode
- storage medium
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M9/00—Arrangements for interconnection not involving centralised switching
- H04M9/08—Two-way loud-speaking telephone systems with means for conditioning the signal, e.g. for suppressing echoes for one or both directions of traffic
- H04M9/082—Two-way loud-speaking telephone systems with means for conditioning the signal, e.g. for suppressing echoes for one or both directions of traffic using echo cancellers
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M9/00—Arrangements for interconnection not involving centralised switching
- H04M9/08—Two-way loud-speaking telephone systems with means for conditioning the signal, e.g. for suppressing echoes for one or both directions of traffic
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R3/00—Circuits for transducers
- H04R3/02—Circuits for transducers for preventing acoustic reaction, i.e. acoustic oscillatory feedback
Definitions
- the present disclosure relates to audio signal processing, and in particular, to an audio signal detection method, device, and audio and video conference system.
- an audio signal detection method including: acquiring an audio signal; performing a time-frequency domain transformation on the audio signal to obtain a transformed signal; analyzing the transformed signal to obtain Test results.
- the time domain-frequency domain transform includes short-time Fourier transform
- the time window of the short-time Fourier transform ranges from 20-60 mS (milliseconds);
- the time window includes Hanning window and Brackman window
- analyzing the transformed signal includes determining at least one of the following characteristics:
- howling includes a strong tone pattern and a repetitive pattern.
- the strong tone mode is determined by at least features (1), (2), (3).
- the second mode is determined by the feature (3.4).
- determining the second mode includes determining that the time interval of the initial frame satisfies a periodic mode.
- the periodic pattern includes: the initial frame sequence has a periodicity in the time domain, and the periodicity includes: a first period and an integer multiple of the first period.
- an audio signal detection device including a non-transitory storage medium, the non-transitory storage medium including an instruction set, the instruction set can be executed by a processor to achieve the above audio Signal detection method.
- the audio signal detection device further includes: a notch filtering device for filtering howling signals when strong howling is detected.
- the audio signal detection device further includes: a volume adjustment device for reducing the volume of the speaker and microphone when repeated howling is detected.
- the audio signal detection device further includes a reminding device for reminding the user to turn off unnecessary speakers or microphones.
- an audio conference system including at least one speaker and at least one microphone; a non-transitory storage medium, the non-transitory storage medium includes an instruction set, the instruction set can be executed by a processor To achieve the above audio signal detection method.
- an audio and video conference system including at least one speaker and at least one microphone; a non-transitory storage medium, the non-transitory storage medium includes an instruction set, and the instruction set can be used by a processor Execute to realize the above-mentioned audio signal detection method.
- Figure 1 is a schematic diagram of how howling occurs
- Figure 2A shows a howling signal in a strong tone mode based on some embodiments
- Figure 2B shows a repetitive pattern howling signal based on some embodiments
- Fig. 3 is a flowchart of a signal processing algorithm for a strong tone mode according to some embodiments
- Figure 1 shows a schematic diagram of the occurrence of "Howling", that is, there are two or more devices 1 in use in a conference room or a certain range of space, and each device includes a microphone 11 and a speaker 12, audio A feedback process of pick-up-feedback-output-pickup occurs continuously between the microphone 11 of one device and the speaker 12 of the other device, and a howling phenomenon finally occurs.
- the feedback of the microphone 11 and speaker 12 of the same device is generally eliminated by an echo canceller, but it may also cause howling when the echo canceller does not work well.
- the inventor performs a time domain-frequency domain conversion on the audio signal.
- the short-time Fourier transform is used to process the audio signal, so that not only the audio signal can be processed in the time domain.
- the signal can be observed and analyzed, and the audio signal can also be observed and analyzed in the frequency domain.
- the inventor found through research that: for the audio signal with howling, the signal after short-time Fourier transform presents two main patterns: tone pattern and repetitive pattern. pattern);
- the strong tone mode it may include more than one howling frequency, as shown in Figure 2A.
- the horizontal axis is the time axis and the vertical axis is the frequency axis.
- the repetitive mode the appearance of howling sounds periodically repeats in time, as shown in Fig. 2B, which is similar to Fig. 2A.
- the horizontal axis is the time axis and the vertical axis is the frequency axis.
- the following parameters are obtained for the signal after the short-time Fourier transform:
- the first parameter the median of the max-power tone duration (MED_MPTD);
- the second parameter the maximum duration of the maximum intensity tone (max of the max-power tone duration; MAX_MPTD);
- the third parameter The detected maximum intensity tone is dominant in the entire frequency spectrum
- the fourth parameter the detection signal is repeated.
- the first parameter MED_MPED represents the median of the duration of the maximum intensity tone in a period of time (Median).
- Median the signal in Figure 2A will display strong horizontal signal bars as shown in the figure. By comparing these signal bars The statistical analysis can get the first parameter MED_MPED.
- the short-time Fourier transform is performed on the original audio signal, for each frame (for example, but not limited to, the frame length is selected as 20mS), by calculating and recording the The frequency bin with the maximum signal power (signal power) is recorded as the maximum-power frequency bin.
- the duration sequence corresponding to the above frequency point sequence is [1,4,4,4,4,1,2,2,1...] (index "54" lasts for 1 frame, "66" Lasted 4 frames).
- the first parameter MED_MPED is defined as the median maximum intensity tone duration
- the second parameter MAX_MPED is defined as the maximum maximum intensity tone duration
- the above three parameters can be used to determine whether a howling in a strong tone mode has occurred.
- the following logical formula can be used to make the judgment:
- the frequency value corresponding to MED_MPTD can also be used as a reference parameter, because the frequency with the highest energy in human speech is mostly within 600 Hz. Therefore, if a dominant strong tone at 2000 Hz is found Mode, then it is certain that an abnormal sound has occurred.
- the presence of a strong tone mode above 750 Hz can be set as one of the criteria for howling. Those skilled in the art will understand that 750 Hz is an optional value, and any value can be selected. Other suitable frequency values, such as 1000 Hz, are used as the criterion.
- Figure 3 shows a flow chart of signal processing in the strong tone mode based on some embodiments.
- 2.56S is used as a detection (judgment) cycle (128 frames, 20mS per frame), and a strong tone mode higher than 750 Hz is set as one of the criteria.
- the aforementioned 128 frames and 750 Hz are only a selectable parameter, and any other suitable parameters are possible.
- 256 frames can be selected as a detection (judgment) period
- 64 frames are used as a detection (judgment) cycle
- a strong tone mode higher than 1000 Hz can be selected as one of the criteria.
- the initial frame presents a periodic distribution in the frame sequence; in some embodiments, the time sequence number of the recorded appearance frame is [....9,28,47,66,85,104,123 ...], it can be seen that the time interval between adjacent starting frames (the difference in the number of adjacent frames) is [...19,19,19,19,19,19,19...], it can be seen that the adjacent The time difference (frame number difference) between the starting frames is 19 frames.
- the signal strength of the appearing frame may be set to be more than twice the average signal strength as the criterion for the starting frame.
- the initial frame time interval sequence is [...19,19,19, 38,19,19,19...], this is because one of the starting frames was not detected due to noise interference or other reasons, but it is easy to find that 38 is an integer multiple of 19, so it can still be based on the above
- the sequence determines the periodicity of the above-mentioned initial frame time interval sequence.
- FIG. 4 shows a flow chart of detecting howling in a repeated pattern based on some embodiments.
- 2.56S is used as a detection (judgment) cycle (128 frames, 20mS per frame).
- a detection (judgment) cycle (128 frames, 20mS per frame).
- 256 frames can be selected as a detection (judgment) period, or 64 frames can be selected.
- a detection (judgment) cycle As a detection (judgment) cycle.
- time window (20mS) in the above embodiment is only a selectable value, and any time window suitable for short-time Fourier transform can be applied to the technical solution of the present disclosure.
- a time window of 40mS may be used.
- a time window of 60mS may be used.
- the time window can be a common Hanning window, Brackman window, or any other window well known to those skilled in the art.
- the above-mentioned loud-sound mode howling and repetitive mode howling detection can be performed simultaneously in a short time interval based on the above four parameters.
- the detection can be performed in a time interval of 1.28S.
- any suitable time interval can be set, such as 2.56S.
- a notch filter may be set to filter the howling signal after the strong howling is detected to eliminate the strong howling.
- the pickup/pronunciation of the speaker and microphone can be lowered, thereby reducing feedback and eliminating repeated howling.
- a reminder to the user can be set on the PC/PAD/LAPTOP user interface, so that the user turns off unnecessary audio devices, solves the hardware condition of howling, and thus eliminates howling.
- the present disclosure can be provided as methods, devices, or computer program products. Therefore, the present disclosure may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
- a computer-usable storage media including but not limited to disk storage, CD-ROM, optical storage, etc.
- Computer-readable media includes permanent and non-permanent, removable and non-removable media, and information storage can be realized by any method or technology.
- the information can be computer-readable instructions, data structures, program modules, or other data.
- Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disc (DVD) or other optical storage, Magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media can be used to store information that can be accessed by computing devices. According to the definition in this article, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
Landscapes
- Engineering & Computer Science (AREA)
- Signal Processing (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Acoustics & Sound (AREA)
- Otolaryngology (AREA)
- General Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Multimedia (AREA)
- Circuit For Audible Band Transducer (AREA)
Abstract
本披露提供一种音频信号检测方法,包括:获取音频信号;对音频信号进行时域-频域变换,并得到变换信号;对变换信号进行分析,得到检测结果。
Description
本披露涉及音频信号处理,具体地,涉及一种音频信号检测方法、装置及音视频会议系统。
在现有的音频或音视频会议中,参会者在使用声音外放(免提)功能时,由于音频信号之间的串扰(Cross Talk),常常出现令人烦恼的“啸叫(Howling)”的异常音。如图1所示意,当一个会议室或一定范围的空间内包括了至少两套设备,每套设备包括一个扬声器和一个麦克风,并且这两套设备都参与了使用的时候,音频信号会在两套设备的扬声器和麦克风之间形成正反馈,并最终导致“啸叫”。
在现有技术的音频或者音视频会议系统中,缺乏对于啸叫的高精度,低复杂度的可靠检测和提醒方案。因此,即使出现了啸叫现象,一般的用户往往难以意识到这是由于对设备的不当使用(在一定的空间范围内使用了两套设备参与会议)而导致的,并且往往会抱怨会议系统出现了问题。另外,即使没有出现啸叫,仍然不推荐在一个会议室或一定范围的空间内的两套设备同时使用(参与会议),因为这样会让音频信号在两套设备的扬声器和麦克风之间形成回波,影响会议质量。
然而,啸叫的检测却是困难和具有技术挑战的,一方面,检测的准确率必须得到保证,即不能把用户的语音或者普通环境杂音当作啸叫信号,从而影响通话质量;另一方面,啸叫的类型不是单一的,并且啸叫往往会和语音以及杂音信号混合在一起,使得对啸叫的检测更加困难。
基于以上所述,我们需要一种准确、可靠的检测啸叫异常音的方案,并及时提醒参会者关闭不必要的设备,从而结束啸叫,提高会议质量。
发明内容
根据本披露一方面的实施例,提供一种音频信号检测方法,包括:获取音频信号;对所述音频信号进行时域-频域变换,并得到变换信号;对所述变换信号进行分析,得到检测结果。
在一些实施例中,时域-频域变换包括短时傅里叶变换;
在一些实施例中,短时傅里叶变换的时间窗长范围为20-60mS(毫秒);
在一些实施例中,时间窗包括汉宁窗,布拉克曼窗;
在一些实施例中,对所述变换信号进行分析包括:确定以下特征中的至少一个:
(1)最大强度音持续中间值(median of the max-power tone duration);
(2)最大强度音持续最大值(max of the max-power tone duration);
(3)最大强度音在整个频谱中占主导地位;
(4)起始帧的周期重复。
在一些实施例中,所述的检测结果包括啸叫。
在一些实施例中,啸叫包括强音模式和重复模式。
在一些实施例中,强音模式至少由特征(1)、(2)、(3)确定。
在一些实施例中,强音模式还包括判据:强音模式中的最大强度频率点的频率高于750Hz。
在一些实施例中,第二模式由所述特征(3.4)所确定。
在一些实施例中,确定所述第二模式包括:确定所述起始帧的时间间隔满足周期模式。
在一些实施例中,周期模式包括:起始帧序列在时间域上具有周期性,周期性包括:第一周期和第一周期的整数倍。
根据本披露另一方面的实施例,提供一种音频信号检测装置,包括非暂态存储介质,所述非暂态存储介质包括指令集,所述指令集可被处理器执行以实现上述的音频信号检测方法。
在一些实施例中,音频信号检测装置还包括:陷波滤波装置,用于在检测到强音啸叫时过滤啸叫信号。
在一些实施例中,音频信号检测装置还包括:音量调节装置,用于在检测到重复啸叫时调低扬声器和麦克风的音量。
在一些实施例中,音频信号检测装置还包括:提醒装置,用于提醒用户关闭不必要的扬声器或麦克风。
根据本披露另一方面的实施例,提供一种音频会议系统,包括至少一个扬声器和至少一个麦克风;非暂态存储介质,非暂态存储介质包括指令集,所述指令集可被处理器执行以实现上述的音频信号检测方法。
根据本披露另一方面的实施例,提供一种音视频会议系统,包括至少一个扬声器和至少一个麦克风;非暂态存储介质,非暂态存储介质包括指令集,所述指令集可被处理器执行以实现上述的音频信号检测方法。
此处所说明的附图用来提供对本披露的进一步理解,构成本披露的一部分,本披露的示意性实施例及其说明用于解释本披露,并不构成对本披露的不当限定。在附图中:
图1是啸叫发生的原理示意图;
图2A示出了基于一些实施例的强音模式啸叫信号;
图2B示出了基于一些实施例的重复模式啸叫信号;
图3是根据一些实施例的强音模式的信号处理算法流程图;
图4是根据一些实施例的重复模式的信号处理算法流程图;
当结合附图来阅读时,将更好地理解前述概述以及某些实施例的以下详细描述。就图示出一些实施例的功能框的简图而言,功能框未必指示硬件电路之间的分割。因而,例如,可在单件硬件(例如通用信号处理器或一块随机存取存储器、硬盘等)或多件硬件中实施功能框中的一个或多个(例如处理器或存储器)。类似地,程序可为独立的程序,可结合成操作系统中的例程,可为安装好的软件包中的函数等。应当理解,一些实施例不限于图中显示的布置和工具。
如本披露所用,以单数叙述或以词语“一个”或“一种”开头的要素或步骤应理解为不排除所述要素或步骤的复数,除非明确陈述了这种排除。此外,对“一个实施例”的引用不意于被解释为排除也结合了所叙述的特征的另外的实施例的存在。除非明确陈述了相反的情况,否则“包括”、“包含”或“具有”具有特定属性的要素或多个要素的实施例可包括不具有那个属性的另外的这样的要素。
图1示出了“啸叫(Howling)”发生的示意图,即在一个会议室或一定范围的空间内有两个或以上的设备1在使用,每个设备都包括麦克风11和扬声器12,音频信号在一个设备的麦克风11与另一个设备的扬声器12之间持续地发生拾音-反馈-输出-拾音的反馈过程并最终出现啸叫现象。(另,同一设备的麦克风11和扬声器12的反馈一般有回波抵消器予以消除,但回波抵消器工作不好的时候也有可能引起啸叫。)
为了对啸叫信号进行检测,发明人对音频信号进行时域-频域的转换,在一些实施例中,采用短时傅里叶变换对音频信号进行处理,从而不仅可以在时间域上对音频信号进行观察分析,还可以在频率域上观察分析音频信号。
在一些实施例中,发明人通过研究发现:对于发生了啸叫的音频信号,经过短时傅里叶变换后的信号呈现两种主要的模式:强音模式(tone pattern)和重复模式(repetitive pattern);在强音模式下,可能包括不止一个啸叫频率,如图2A所示意,在图2A中,横轴为时间轴,纵轴为频率轴。而在重复模式下,啸叫音的出现在时间上呈现周期性重复的特征,如图2B所示意,和图2A相类似,在图2B中,横轴为时间轴,纵轴为频率轴。
在一些实施例中,为了确定上述的两种啸叫模式,对于短时傅里叶变换后的信号,获取以下参数:
第一参数:最大强度音持续中间值(median of the max-power tone duration;MED_MPTD);
第二参数:最大强度音持续最大值(max of the max-power tone duration;MAX_MPTD);
第三参数:检测最大强度音在整个频谱占主导地位;
第四参数:检测信号重复。
通过对上述四个参数的检测和判断,本披露可以实现对于啸叫的准确可靠的检测。
第一参数MED_MPED代表在一段时间内最大强度音持续时间的中值(Median),当啸叫发生时,如图2A中的信号会显示如图所示的强水平信号条,通过对这些信号条的统计分析可以得到第一参数MED_MPED。
在一些实施例中,具体而言,在对原始音频信号进行短时傅里叶变换后,对于每一帧(例如但不限于,选择帧长为20mS),通过计算并记录每一帧中的具有最大信号强度(signal power)的频率点(frequency bin),记为最大强度频率点(max-power frequency bin)。一旦到达预设的时间范围(例如经过128帧,总时间为20mS*128=2.56S),计算并记录频率点(max-power frequency bin index)的持续时间。在一些实施例中,记录得到的频率点序列为[54,66,66,66,66,45,49,49,43…],其中:54即为最大强度频率 点的索引(index),假如每个频率点的带宽是25Hz,54就代表54*25=1350Hz。在一些实施例中,和上面频率点序列对应的持续时间序列为[1,4,4,4,4,1,2,2,1…](索引“54”持续了1帧,“66”持续了4帧)。
在一些实施例中,第一参数MED_MPED被定义为最大强度音持续时间中值,而第二参数MAX_MPED被定义为最大强度音持续时间最大值。
在一些实施例中,第三参数用于确定频率点是否在整个频谱中占据主导地位,例如,如果待确定的频率点的强度比其它所有频率点强度之和还要大,则可以确定该频率点在整个频谱中占据主导地位。
在一些实施例中,可以利用以上三种参数来确定是否发生了强音模式的啸叫,具体地,可以采用例如以下的逻辑式来进行判断:
(MED_MPTD>6)&&(frequency bin for MAX_MPTD is same as the frequency bin for MED_MPTD)&&(TotalFrames_DominantTone>5)&&(Enough_Signal_Energy)。
需要指出的是,上述逻辑式仅仅是一种可以采用的判断逻辑式,本领域技术人员可以理解的是,可以采用任何其它的合适的逻辑式来进行判断,例如:
(MED_MPTD>=5)&&(Median_Tone_Freq>1500)&&(frequency bin for MAX_MPTD is same as the frequency bin for MED_MPTD)&&(TotalFrames_DominantTone>0)&&(Enough_Signal_Energy)。
在一些实施例中,对应MED_MPTD的频率值也可以作为一个参考的参数,因为人类语音中能量最大的频率大多数处于600赫兹之内,因此,如果发现了一个处于2000Hz的占有主导地位的强音模式,那么可以肯定发生了异常音。在一些实施例中,可以设置在750赫兹以上的强音模式的出现作为发生啸叫的判据之一,本领域技术人员可以理解的是,750赫兹是一个可选的值,还可以选择任何其它合适的频率值,例如1000Hz作为判据。
图3给出了基于一些实施例的强音模式的信号处理流程图。在图3的流 程中,以2.56S为一个检测(判断)周期(128帧,20mS每帧),并且设置高于750Hz的强音模式作为判据之一。本领域技术人员可以理解的是,上述的128帧,750Hz仅仅是一种可以选择的参数,任何合适的其它参数都是可以的,例如,可以选择256帧作为一个检测(判断)周期,或者选择64帧作为一个检测(判断)周期,或者,也可以选择高于1000Hz的强音模式作为判据之一。
上述实施例的技术方案可以准确、可靠地检测强音模式的啸叫,但是,对于重复模式(图2B)的啸叫则无法确定。基于此,本披露采用第四参数来确定是否发生了重复模式的啸叫。
在一些实施例中,在每个频点上,检测当前帧的信号强度(signal power)是否比前两帧的强度大得多。如果当前帧有足够多的频点的信号强度比前两帧大得多,则定义为起始帧(onset frame)。在重复模式啸叫中,起始帧在帧序列中呈现出周期性分布;在一些实施例中,记录得到的出现帧时间序号为[….9,28,47,66,85,104,123…],可以看出:相邻的起始帧之间的时间间隔(相邻帧数差)为[…19,19,19,19,19,19,19…],可以看出,相邻的起始帧之间的时间差(帧数差)为19帧,在帧长为20mS的情况下,所对应的具体的时间差为19*20mS=380mS,这种情况下,起始帧在时间轴上呈现出周期为380mS的序列特征,而在实际中,这种重复模式的啸叫在用户听起来呈现出规律的:“哒..哒..哒..”的声音。
在一些实施例中,可以设置出现帧的信号强度比平均信号强度强两倍以上作为起始帧的判据。
在一些实施例中,上述的起始帧序列之间可能会呈现出轻微的“变形(deviation)”,例如,在一些实施例中,起始帧时间间隔序列为[…19,19,19,38,19,19,19…],这是由于其中一个起始帧由于噪音的干扰或者其它原因而没有被检测到,但是很容易发现的是,38是19的整数倍,所以仍然可以基于上述序列确定上述起始帧时间间隔序列的周期性。
需要指出的是,在重复模式的啸叫检测中,需要对信号进行足够长时间 的检测才能确定出现帧的周期模式,而一旦确定周期性的起始帧模式,则可以判定出现了重复模式的啸叫。
图4示出了基于一些实施例的重复模式啸叫的检测流程图,在图4的流程中,以2.56S为一个检测(判断)周期(128帧,20mS每帧)。本领域技术人员可以理解的是,上述的128帧仅仅是一种可以选择的参数,任何合适的其它参数都是可以的,例如,可以选择256帧作为一个检测(判断)周期,或者选择64帧作为一个检测(判断)周期。
需要指出的是,上述实施例中的时间窗(20mS)仅是一个可以选择的值,任何适于短时傅里叶变换的的时间窗均可以应用于本披露的技术方案,在一些实施例中,可以采用40mS的时间窗,在另一些实施例中,可以采用60mS的时间窗。
具体地,时间窗可以选择常见的汉宁窗,布拉克曼窗,或者任何其它本领域技术人员所熟知的窗。
在一些实施例中,可以基于上述的四种参数在很短的时间间隔内同时进行上述的强音模式啸叫和重复模式啸叫检测,例如,可以在1.28S的时间间隔内进行检测,本领域技术人员可以理解的是:可以设置任何合适的时间间隔(检测间隔),例如2.56S。
在一些实施例中,可以设置陷波滤波器(notch filter),从而当强音啸叫被检测到后过滤啸叫信号,消除强音啸叫。
在一些实施例中,当重复啸叫被检测到后可以调低扬声器和麦克风的拾音/发音大小,从而降低反馈,消除重复啸叫。
在一些实施例中,可以在PC/PAD/LAPTOP用户界面上设置对于用户的提醒,使得用户关闭不需要的音频设备,解决啸叫出现的硬件条件并因而消除啸叫。
本披露的技术方案具有上述的技术优势以及因此而带来的广泛的应用优势。这些应用优势包括但不限于:
(1)极高的检测性能,在实际使用测试中,本披露的技术方案可以实现0误检率和高达90%的检出率;
(2)低复杂度,本披露技术方案所需要的计算能力要求非常低,绝大多数硬件设备都可以实现;从而本披露技术方案可以以非常低的成本应用于现有的音频会议系统或者音视频会议系统中;
(3)多样化的配置,在对于不同模式的啸叫检出后,可以设置多种技术手段或提醒手段来消除啸叫。
要理解的是,以上描述意于为示例性,而不是限制性的。例如,上面描述的实施例(和/或它们的各方面)可与彼此结合起来使用。另外,可在不偏离一些实施例的范围的情况下做出许多修改,以使具体情况或内容适于一些实施例的教导。虽然本文描述的材料的尺寸和类型意于限定一些实施例的参数,但实施例决不是限制性的,而是示例性实施例。在审阅以上描述之后,许多其它实施例对本领域技术人员将是显而易见的。因此,应当参照所附权利要求以及这样的权利要求所涵盖的等效体的全部范围来确定一些实施例的范围。在所附权利要求中,用语“包括”和“在其中”用作相应的用语“包含”和“其中”的易懂语言等效体。此外,在所附权利要求中,用语“第一”、“第二”和“第三”等仅作为标记使用,并且它们不意于对它们的对象施加数字要求。另外,不以手段加功能的格式来书写所附权利要求的限制,除非且直到这样的权利要求限制清楚地使用短语“用于…的手段”,跟随没有另外的结构的功能陈述。
还需要说明的是,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、商品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、商品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、商品或者设备中还存在另外的相同要素。
本领域内的技术人员应明白,本披露的一些实施例可提供为方法、设备、或计算机程序产品。因此,本披露可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本披露可采用在一个或多个其中包含有计算机可用程序代码的计算机可用存储介质(包括但不限于磁盘 存储器、CD-ROM、光学存储器等)上实施的计算机程序产品的形式。
计算机可读介质包括永久性和非永久性、可移动和非可移动媒体可以由任何方法或技术来实现信息存储。信息可以是计算机可读指令、数据结构、程序的模块或其他数据。计算机的存储介质的例子包括,但不限于相变内存(PRAM)、静态随机存取存储器(SRAM)、动态随机存取存储器(DRAM)、其他类型的随机存取存储器(RAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、快闪记忆体或其他内存技术、只读光盘只读存储器(CD-ROM)、数字多功能光盘(DVD)或其他光学存储、磁盒式磁带,磁带磁磁盘存储或其他磁性存储设备或任何其他非传输介质,可用于存储可以被计算设备访问的信息。按照本文中的界定,计算机可读介质不包括暂存电脑可读媒体(transitory media),如调制的数据信号和载波。
本书面描述使用示例来公开一些实施例,包括最佳模式,并且还使本领域任何技术人员能够实践一些实施例,包括制造和使用任何装置或系统,以及实行任何结合的方法。一些实施例的保护范围由权利要求限定,并且可包括本领域技术人员想到的其它示例。如果这样的其它示例具有不异于权利要求的字面语言的结构要素,或者如果它们包括与权利要求的字面语言无实质性差异的等效结构要素,则它们意于处在权利要求的范围之内。
Claims (18)
- 一种音频信号检测方法,包括:获取音频信号;对所述音频信号进行时域-频域变换,并得到变换信号;对所述变换信号进行分析,得到检测结果。
- 根据权利要求1所述的方法,其中,所述的时域-频域变换包括短时傅里叶变换。
- 根据权利要求2所述的方法,其中,所述短时傅里叶变换的时间窗长范围为20-60mS。
- 根据权利要求3所述的方法,其中,所述短时傅里叶变换的时间窗长为20mS。
- 根据权利要求1所述的方法,其中,对所述变换信号进行分析包括:确定以下特征中的至少一个:(1)最大强度音持续中间值(median of the max-power tone duration);(2)最大强度音持续最大值(max of the max-power tone duration);(3)最大强度音在整个频谱中占主导地位;(4)起始帧在帧序列中呈现的周期性特征。
- 根据权利要求5所述的方法,其中,所述的检测结果包括啸叫。
- 根据权利要求6所述的方法,其中,所述的啸叫包括强音模式和重复模式。
- 根据权利要求7所述的方法,其中,所述强音模式至少由所述特征(1)、(2)、(3)确定。
- 根据权利要求8所述的方法,其中,所述强音模式还包括判据:强音模式中的最大强度频率点的频率高于750Hz。
- 根据权利要求7所述的方法,其中,所述第二模式由所述特征(4)所确定。
- 根据权利要求10所述的方法,其中,确定所述第二模式包括:确定所述出现帧的时间间隔满足周期模式。
- 根据权利要求11所述的方法,其中,所述周期模式包括:所述出现帧序列在时间域上具有周期性,所述的周期性包括:第一周期和第一周期的整数倍。
- 一种音频信号检测装置,包括非暂态存储介质,所述非暂态存储介质包括指令集,所述指令集可被处理器执行以实现如权利要求1-13任一所述的方法。
- 根据权利要求13所述的装置,还包括:陷波滤波装置,用于在检测到强音啸叫时过滤啸叫信号。
- 根据权利要求13所述的装置,还包括:音量调节装置,用于在检测到重复啸叫时调低扬声器和麦克风的音量。
- 根据权利要求13所述的装置,还包括:提醒装置,用于提醒用户关闭不必要的扬声器或麦克风。
- 一种音频会议系统,包括:至少一个扬声器和至少一个麦克风;非暂态存储介质,所述非暂态存储介质包括指令集,所述指令集可被处理器执行以实现如权利要求1-13任一所述的方法。
- 一种音视频会议系统,包括:至少一个扬声器和至少一个麦克风;非暂态存储介质,所述非暂态存储介质包括指令集,所述指令集可被处理器执行以实现如权利要求1-13任一所述的方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2019/072543 WO2020150864A1 (zh) | 2019-01-21 | 2019-01-21 | 音频信号检测方法、装置及音视频会议系统 |
| CN201980089645.8A CN113383537B (zh) | 2019-01-21 | 2019-01-21 | 音频信号检测方法、装置及音视频会议系统 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2019/072543 WO2020150864A1 (zh) | 2019-01-21 | 2019-01-21 | 音频信号检测方法、装置及音视频会议系统 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020150864A1 true WO2020150864A1 (zh) | 2020-07-30 |
Family
ID=71735360
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/072543 Ceased WO2020150864A1 (zh) | 2019-01-21 | 2019-01-21 | 音频信号检测方法、装置及音视频会议系统 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN113383537B (zh) |
| WO (1) | WO2020150864A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN118629383A (zh) * | 2024-08-08 | 2024-09-10 | 宁波方太厨具有限公司 | 主动降噪系统及其控制方法、异音检测方法、装置 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050220313A1 (en) * | 2004-03-30 | 2005-10-06 | Yamaha Corporation | Howling frequency component emphasis method and apparatus |
| CN103871418A (zh) * | 2014-03-06 | 2014-06-18 | 北京飞利信电子技术有限公司 | 一种扩声系统啸叫频点的检测方法及装置 |
| CN105812993A (zh) * | 2014-12-29 | 2016-07-27 | 联芯科技有限公司 | 啸叫检测和抑制方法及其装置 |
| CN108184192A (zh) * | 2017-12-27 | 2018-06-19 | 中山大学花都产业科技研究院 | 一种自适应声反馈抑制方法 |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101248341A (zh) * | 2005-08-25 | 2008-08-20 | 罗伯特·博世有限公司 | 评估啸叫噪音干扰性的方法和装置 |
-
2019
- 2019-01-21 CN CN201980089645.8A patent/CN113383537B/zh active Active
- 2019-01-21 WO PCT/CN2019/072543 patent/WO2020150864A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050220313A1 (en) * | 2004-03-30 | 2005-10-06 | Yamaha Corporation | Howling frequency component emphasis method and apparatus |
| CN103871418A (zh) * | 2014-03-06 | 2014-06-18 | 北京飞利信电子技术有限公司 | 一种扩声系统啸叫频点的检测方法及装置 |
| CN105812993A (zh) * | 2014-12-29 | 2016-07-27 | 联芯科技有限公司 | 啸叫检测和抑制方法及其装置 |
| CN108184192A (zh) * | 2017-12-27 | 2018-06-19 | 中山大学花都产业科技研究院 | 一种自适应声反馈抑制方法 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN118629383A (zh) * | 2024-08-08 | 2024-09-10 | 宁波方太厨具有限公司 | 主动降噪系统及其控制方法、异音检测方法、装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN113383537B (zh) | 2023-04-07 |
| CN113383537A (zh) | 2021-09-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP1913708B1 (en) | Determination of audio device quality | |
| US9361898B2 (en) | Three-dimensional sound compression and over-the-air-transmission during a call | |
| Jeub et al. | Model-based dereverberation preserving binaural cues | |
| JP6246792B2 (ja) | ユーザのグループのうちのアクティブに話しているユーザを識別するための装置及び方法 | |
| MX2012011203A (es) | Procesador de audio espacial y metodo para proveer parametros espaciales en base a una señal de ntrada acustica. | |
| US11380312B1 (en) | Residual echo suppression for keyword detection | |
| CN110364175B (zh) | 语音增强方法及系统、通话设备 | |
| CN107124647A (zh) | 一种全景视频录制时自动生成字幕文件的方法及装置 | |
| US9240190B2 (en) | Formant based speech reconstruction from noisy signals | |
| US10192566B1 (en) | Noise reduction in an audio system | |
| CN114678038A (zh) | 音频噪声检测方法、计算机设备和计算机程序产品 | |
| CN101460994A (zh) | 语音区分 | |
| JP4745916B2 (ja) | 雑音抑圧音声品質推定装置、方法およびプログラム | |
| CN113383537B (zh) | 音频信号检测方法、装置及音视频会议系统 | |
| JP7826314B2 (ja) | パーベイシブリスニング向けに編成されたギャップ | |
| Lugasi et al. | Spatial audio signal enhancement by a two-stage source—system estimation with frequency smoothing for improved perception | |
| Habets et al. | Speech dereverberation using backward estimation of the late reverberant spectral variance | |
| Thomas et al. | Automated suppression of howling noise using sinusoidal model based analysis/synthesis | |
| Tsilfidis et al. | Signal-dependent constraints for perceptually motivated suppression of late reverberation | |
| Uhle et al. | A supervised learning approach to ambience extraction from mono recordings for blind upmixing | |
| Jang et al. | Acoustic Feedback Detection for Online Video Conferencing | |
| Lee et al. | Multiple reverberant sound localization based on rigorous zero-crossing-based ITD selection | |
| CN112309419B (zh) | 多路音频的降噪、输出方法及其系统 | |
| Wu et al. | A multi-microphone speech enhancement algorithm tested using acoustic vector sensors | |
| Romoli et al. | A voice activity detection algorithm for multichannel acoustic echo cancellation exploiting fundamental frequency estimation |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19911905 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19911905 Country of ref document: EP Kind code of ref document: A1 |