WO2020034095A1 - 音频信号处理装置及方法 - Google Patents
音频信号处理装置及方法 Download PDFInfo
- Publication number
- WO2020034095A1 WO2020034095A1 PCT/CN2018/100464 CN2018100464W WO2020034095A1 WO 2020034095 A1 WO2020034095 A1 WO 2020034095A1 CN 2018100464 W CN2018100464 W CN 2018100464W WO 2020034095 A1 WO2020034095 A1 WO 2020034095A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- microphone
- microphones
- audio signal
- audio signals
- axes
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R5/00—Stereophonic arrangements
- H04R5/027—Spatial or constructional arrangements of microphones, e.g. in dummy heads
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R1/00—Details of transducers, loudspeakers or microphones
- H04R1/20—Arrangements for obtaining desired frequency or directional characteristics
- H04R1/32—Arrangements for obtaining desired frequency or directional characteristics for obtaining desired directional characteristic only
- H04R1/326—Arrangements for obtaining desired frequency or directional characteristics for obtaining desired directional characteristic only for microphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R1/00—Details of transducers, loudspeakers or microphones
- H04R1/02—Casings; Cabinets ; Supports therefor; Mountings therein
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R1/00—Details of transducers, loudspeakers or microphones
- H04R1/20—Arrangements for obtaining desired frequency or directional characteristics
- H04R1/32—Arrangements for obtaining desired frequency or directional characteristics for obtaining desired directional characteristic only
- H04R1/40—Arrangements for obtaining desired frequency or directional characteristics for obtaining desired directional characteristic only by combining a number of identical transducers
- H04R1/406—Arrangements for obtaining desired frequency or directional characteristics for obtaining desired directional characteristic only by combining a number of identical transducers microphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R27/00—Public address systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R3/00—Circuits for transducers
- H04R3/005—Circuits for transducers for combining the signals of two or more microphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R5/00—Stereophonic arrangements
- H04R5/04—Circuit arrangements, e.g. for selective connection of amplifier inputs/outputs to loudspeakers, for loudspeaker detection, or for adaptation of settings to personal preferences or hearing impairments
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R2201/00—Details of transducers, loudspeakers or microphones covered by H04R1/00 but not provided for in any of its subgroups
- H04R2201/40—Details of arrangements for obtaining desired directional characteristic by combining a number of identical transducers covered by H04R1/40 but not provided for in any of its subgroups
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R2430/00—Signal processing covered by H04R, not provided for in its groups
- H04R2430/20—Processing of the output signals of the acoustic transducers of an array for obtaining a desired directivity characteristic
Definitions
- the present disclosure relates to an audio signal processing device and a corresponding method.
- microphone arrays are widely used in various front-end devices, such as Automatic Speech Recognition (ASR) and audio / video conference systems.
- ASR Automatic Speech Recognition
- picking up the "best quality" sound signal means that the acquired signal has the largest signal-to-noise ratio (SNR) and the smallest reverberation.
- SNR signal-to-noise ratio
- the common "octopus" structure shown in Fig. 1 is generally adopted: that is, three directional microphones 11 at an angle of 120 degrees are set at three "ends". The sound signal is received by one of the microphones through these three ends, and then the received sound signal is processed by the digital signal processing device.
- the direction of the sound signal is not consistent with the end containing the directional microphone, the sound signal will experience relatively severe attenuation during the reception process. Generally, this problem is called “off-axis (off-axis).
- an audio signal processing apparatus including: a plurality of microphones; a plurality of microphones are arranged close to each other, and the plurality of microphones form a symmetrical structure.
- the projections of the axes of the plurality of microphones in the same horizontal plane form an included angle of 120 degrees.
- the axes of multiple microphones are located in the same horizontal plane, and the axes of any two microphones form an included angle of 120 degrees.
- the axes of the plurality of microphones are parallel to each other and the projection points of the plurality of axes in their vertical planes form three vertices of an equilateral triangle.
- the distance between the ends of any two microphones ranges from 0-5 mm.
- the microphone includes a directional microphone.
- the microphone includes at least one of: a cardioid microphone (Cardioid microphone), a subcardioid microphone (Subcardioid microphone), a supercardioid microphone (Supercardioid microphone), and a supercardioid microphone Type microphone (Hypercardioid microphone), dipole type microphone (Dipole microphone).
- a cardioid microphone Cardioid microphone
- Subcardioid microphone subcardioid microphone
- Supercardioid microphone supercardioid microphone
- Type microphone Hypercardioid microphone
- Dipole microphone dipole type microphone
- an audio signal processing method that uses the audio signal processing device in the present disclosure and includes the steps of: linearly combining audio signals obtained by a plurality of microphones; and based on the combined audio signals , Dynamically select the best pickup direction.
- the matrix A for linear combination is set as:
- ⁇ m is the beam angle
- ⁇ n is the air angle
- ⁇ n ⁇ m + 110 * ⁇ / 180. .
- ⁇ n ⁇ m + ⁇ .
- the combined audio signals are continuously processed based on the set sampling time interval to obtain audio signals in multiple virtual directions; the audio signals in multiple virtual directions are compared, and the direction with the highest signal-to-noise ratio is selected as Pick up direction.
- a short-time Fourier transform is used to process the combined audio signal.
- the set sampling interval is 10-20 ms.
- an audio signal is acquired and output based on the selected pickup direction.
- a non-transitory storage medium stores an instruction set, and when the instruction set is executed by a processor, the processor can perform the following process: linear combination Audio signals obtained by multiple microphones; based on the combined audio signals, the optimal pickup direction is dynamically selected.
- FIG. 1 is a schematic diagram of a conference system device in the prior art
- FIG. 1-1 shows a pickup attenuation curve of the conference system device in FIG. 1;
- FIG. 3 is a schematic diagram of setting a plurality of microphones according to some embodiments.
- FIG. 4 is a schematic diagram of a plurality of microphone settings according to some embodiments.
- FIG. 5 is a schematic diagram of a plurality of microphone settings according to some embodiments.
- FIG. 7 is a flowchart of exemplary steps of an algorithm according to some embodiments.
- FIG. 8 is an audio signal spectrum obtained according to some embodiments.
- the functional blocks of some embodiments do not necessarily indicate division between hardware circuits.
- one or more of the functional blocks may be implemented in a single piece of hardware (such as a general-purpose signal processor or a piece of random access memory, a hard disk, etc.) or multiple pieces of hardware.
- the program may be an independent program, which may be combined into a routine in an operating system, or a function in an installed software package, and the like. It should be understood that some embodiments are not limited to the arrangements and tools shown in the figures.
- FIG. 4 shows three superimposed directional microphones 41, 42, and 43, and FIG. 4 shows a "top-down" perspective
- the three directional microphones are 41, 42 and 43 in order from top to bottom.
- the axes of the directional microphones 41, 42 and 43 (lines perpendicular to the center of their pickup plane) are parallel to the plane of FIG. If the directional microphones 41, 42, 43 are projected in the plane of FIG. 4, they also form a triple symmetrical arrangement, and the axes 411, 421, and 431 of the three directional microphones are formed in two in the projection plane of FIG. 4. ⁇ 2/3 included angle (indicated by the dotted line on the right side in Figure 4).
- FIG. 5 shows three directional microphones 51, 52, and 53.
- a triple symmetrical arrangement is formed between these three directional microphones.
- the axes 511, 521, 531 (lines perpendicular to the center of the pickup plane) of the three directional microphones are parallel to each other, and the three projection points of the axes 511, 521, 531 in the plane perpendicular to them constitute an equilateral side Triangle T.
- the distance range D between 51 and 52 shown in the figure
- D 2mm can be selected.
- Directional microphones include but are not limited to: Cardioid microphones, Subcardioid microphones, Supercardioid microphones, Hypercardioid microphones , Dipole microphone (Dipole microphone) to form the microphone setup shown in Figure 3-5. It can be understood that: you can choose the same type of directional microphone, such as a cardioid directional microphone, to form any of the microphone settings in Figure 3-5; you can also choose a combination of different types of directional microphones to form Figure 3-5 Any of the microphone settings.
- the technical solution of the present disclosure will simultaneously pick up and combine audio signals from multiple microphones.
- the distance between a plurality of microphones is set as small as possible; thus, the time difference between the audio signals reaching different microphones can be reduced as much as possible, so that the audio signals of the plurality of microphones are combined "simultaneously" Physically possible first.
- a “virtual microphone” is constituted by “simultaneously” linearly combining signals from three microphones of a physical entity (such as a heart-type directional microphone).
- the coefficients of the linear combination are represented by the vector ⁇ :
- ⁇ m represents the beam angle (that is, the direction of the audio signal that is desired to be obtained), and ⁇ n represents the null angle (that is, the direction of the audio signal that is not desired to be obtained).
- ⁇ m and ⁇ n are selected as:
- ⁇ n ⁇ m + 110 * ⁇ / 180
- FIG. 6 shows the sound pickup effect of the technical solution of the present disclosure in a 60-degree direction under this setting. It can be seen by comparing FIG. 1-1 that the technical solution of the present disclosure has no attenuation at all in the direction of 60 degrees. In addition, not only in the 60-degree direction, but also by dynamically selecting an appropriate ⁇ m , the technical solution of the present disclosure can achieve the technical effect of no attenuation in the 360-degree direction.
- ⁇ m and ⁇ n may be selected as:
- the algorithm and microphone settings of the present disclosure can implement any type of virtual first-order differential microphone, including cardioid microphones, cardioid microphones, and subcardioid directional microphones.
- cardioid microphones including cardioid microphones, cardioid microphones, and subcardioid directional microphones.
- Subcardioid microphone supercardioid microphone
- Supercardioid microphone supercardioid microphone
- supercardioid microphone Heypercardioid microphone
- dipole directional microphone Dipole microphone
- the above-mentioned combination of audio signals is frequency-independent, that is to say: the beamforming mode is the same for any frequency, so the technical solution of the present disclosure does not "amplify" white noise in the low frequency band, thus The technical solution of the present disclosure can also solve the WNG problem.
- the beam selection algorithm further compares in real time and selects the beam direction with the highest signal-to-noise ratio (SNR) from the virtual beams in multiple directions as the audio output source.
- SNR signal-to-noise ratio
- FIG. 7 shows a flowchart of a beam selection algorithm in some embodiments.
- an audio signal frame is transformed into a frequency domain signal by a Short-Time Fourier Transform.
- step 72 determine whether each frequency point (Frequency Frequency Bin) contains an audio signal; if not, proceed directly to step 75 to increase the frequency point; if there is, proceed to step 73, at the current frequency point, select the one with the largest signal. Noise ratio signal, record the corresponding beam index. And in step 74 and step 75, the maximum signal-to-noise ratio number of signals and the frequency interval are sequentially increased.
- step 76 it is determined whether the current total frequency point has been traversed. If not, repeat the above steps 72-75. If yes, then select the signal with the maximum SNR from all virtual beams in step 77, and in step 78 outputs the above-mentioned signal having the maximum SNR as a speech signal.
- FIG. 8 shows an audio signal spectrum obtained by the technical solution of the present disclosure.
- the red spectral line is an audio signal obtained by a virtual microphone of the technical solution of the present disclosure
- the blue spectral line is an audio signal obtained by a traditional physical microphone. It is shown that the SNR of the signals obtained by the technical solution of the present disclosure is better than that of the conventional technology in each spectrum segment.
- the technical solution of the present disclosure can also solve the WNG problem.
- the effective pickup range of audio devices using the settings and algorithms of the present disclosure can be 3x times that of the prior art devices. Therefore, even for a large conference room Using Daisy chain to combine only a few audio devices can achieve effective pickup in the entire area.
- the microphone settings and algorithms of the present disclosure are used in a multi-party conference call, so that when the main speaker speaks, there are noises from other participants in different positions from the main speaker (such as when making a call)
- the problem You can dynamically set and select the direction of ⁇ m in the direction of the main speaker and ⁇ n in the direction of the noise, so that the audio signal can be obtained only from the direction of the main speaker, and the noise emitted by the noise direction is completely different. Will be picked up by the microphone.
- the microphone settings and algorithms of the present disclosure are used in a voice shopping device, especially a voice shopping device in a public place (such as a vending machine), so as to solve the problem that the shopper cannot be accurately identified in a noisy public place. Problems with audio signals.
- the ⁇ m is dynamically set and selected in the direction in which the shopper speaks in real time.
- the technical solution of the present disclosure has a good suppression effect on the background noise so that it can be accurately picked Voice signals from shoppers.
- the microphone settings and algorithms of the present disclosure are adopted in a smart speaker, especially when used in a home environment, when there is noise around and other voice signal sources, similar to the above description, it can accurately pick up from Command the sender's voice signal to avoid noise from noise sources, and also have a good suppression effect on background sound.
- Computer-readable media includes both permanent and non-persistent, removable and non-removable media.
- Information can be stored by any method or technology.
- Information may be computer-readable instructions, data structures, modules of a program, or other data.
- Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), and read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, Magnetic tape cartridges, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media may be used to store information that can be accessed by computing devices.
- computer-readable media does not include temporary computer-readable media, such as modulated data signals and carrier waves.
Landscapes
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Otolaryngology (AREA)
- General Health & Medical Sciences (AREA)
- Circuit For Audible Band Transducer (AREA)
Abstract
本披露提供一种音频信号处理装置,包括:多个麦克风;多个麦克风两两接近设置,并且多个麦克风形成对称性结构。
Description
本披露涉及一种音频信号处理装置及相应的方法。
为了获取高质量的声音信号,麦克风阵列被广泛地应用于各种前端设备,例如自动语音识别(Automatic speech recognition:ASR)以及音/视频会议系统。一般而言,拾取“最佳质量”的声音信号意味着所获取的信号具有最大的信噪比(Signal-to-noise ratio:SNR)以及最小的回响(Reverberation)。
在现有的会议系统的拾音系统中,一般采用如图1所示的常见的“八爪鱼”结构:即设置三个互相成120度夹角的指向型麦克风11在三个“端”部,声音信号通过这三个端部被其中某一麦克风所接收,然后通过数字信号处理装置对接收到的声音信号进行处理。然而,在这种设计中,如果声音信号的方向不是和包含指向型麦克风的端部相一致的话,声音信号在接收过程中会遭遇比较严重的衰减,通常而言这种问题称为“偏离轴线(off-axis)”。举例而言,如果声音信号来自两个端部的角平分线方向(60度方向),例如图1中所示意的A方向,那么所获取的声音信号在此方向会衰减至3dB,如图1-1的衰减曲线所示意。在这种情况下,如果一个发言者处于图1中的A方向的位置,那么他的声音信号在拾音过程中将被极大地衰减,从而可能使得会议另一端(可能处于另一个城市)的人员无法清楚地听到他的话语。另一方面,在会议过程中,常常还会出现除了发言者外的噪音信号。在特殊的情况下,例如位于发言者不同方位的其他参会者所发出的噪音(例如打电话),而如果发言者处于图1的A方位而恰巧噪音来自图1中的B方向(其中一个麦克风的端部方向),那么发言者的声音信号在拾音过程中将被抑制,并且噪声信号将被完整无衰减地拾取,其结果会导致会议另一端的人员根本无法获得有效信息。
在另一种设计方案中,如图2所示意,采用三个全向型麦克风形成一个环形结构,其中全向型麦克风之间的间距大约在2cm左右,这种设计虽然可以部分解决上述的声音信号偏离轴线所导致的衰减问题,但是这种设计会放大低频白噪声,从而产生所谓的white-noise-gain(WNG)问题。
基于以上,需要一种新的音频信号处理装置和方法,以解决上述的这些技术问题。
发明内容
根据本披露的一个方面的实施例,提供一种音频信号处理装置,包括:多个麦克风;多个麦克风两两接近设置,并且所述多个麦克风形成对称性结构。
在一些实施例中,多个麦克风为三个。
在一些实施例中,多个麦克风的轴线在同一水平面内的投影两两形成120度夹角。
在一些实施例中,多个麦克风的轴线位于同一水平面内,所述任意两个麦克风的轴线形成120度夹角。
在一些实施例中,多个麦克风为三个,并且多个麦克风构成叠加模式。
在一些实施例中,多个麦克风的轴线两两平行并且所述多个轴线在其垂直平面内的投影点形成等边三角形的三个顶点。
在一些实施例中,任意两个麦克风的端部之间的距离范围为0-5mm。
在一些实施例中,麦克风包括指向型麦克风。
在一些实施例中,麦克风包括以下的至少一种:心型指向型麦克风(Cardioid microphone)、亚心型指向型麦克风(Subcardioid microphone)、过心型指向型麦克风(Supercardioid microphone)、超心型指向型麦克风(Hypercardioid microphone)、偶极型指向型麦克风(Dipole microphone)。
根据本披露的另一方面,提供一种音频信号处理方法,,其使用本披露中的音频信号处理装置,并且包括步骤:线性组合由多个麦克风所获得的音频信号;基于组合后的音频信号,动态选择最佳拾音方向。
在一些实施例中,用于线性组合的矩阵A设置为:
其中:θ
m为波束角度,θ
n为空角度。
在一些实施例中,当以虚拟Hyper-cardioid microphone模式组合多个麦克风的音频信号时,θ
n=θ
m+110*π/180。。
在一些实施例中,当以虚拟Cardioid microphone模式组合多个麦克风的音频信号时,θ
n=θ
m+π。
在一些实施例中,基于设定的采样时间间隔连续地对组合后的音频信号进行处理,得到多个虚拟方向的音频信号;比较多个虚拟方向的音频信号,选择信噪比最高的方向作为拾音方向。
在一些实施例中,采用短时傅里叶变换对组合后的音频信号进行处理。
在一些实施例中,设定的采样时间间隔为10-20ms。
在一些实施例中,基于选择出的拾音方向获取音频信号并输出。
根据本披露的另一方面,提供一种非暂态存储介质,所述非暂态存储介质存储有指令集,所述指令集由处理器执行时使得所述处理器能够执行如下过程:线性组合由多个麦克风所获得的音频信号;基于组合后的音频信号,动态选择最佳拾音方向。
此处所说明的附图用来提供对本披露的进一步理解,构成本披露的一部分,本披露的示意性实施例及其说明用于解释本披露,并不构成对本披露的不当限定。在附图中:
图1是现有技术中的一种会议系统装置示意;
图1-1示出了图1中的会议系统装置的拾音衰减曲线;
[根据细则91更正 19.11.2018]
图2是现有技术中的一种会议系统装置示意;
图2是现有技术中的一种会议系统装置示意;
图3是根据一些实施例的多个麦克风设置示意;
图4是根据一些实施例的多个麦克风设置示意;
图5是根据一些实施例的多个麦克风设置示意;
图6是根据一些实施例的本披露的拾音曲线;
图7是根据一些实施例的算法的示例性步骤流程图;
图8是根据一些实施例所获得的音频信号图谱
当结合附图来阅读时,将更好地理解前述概述以及某些实施例的以下详细描述。就图示出一些实施例的功能框的简图而言,功能框未必指示硬件电路之间的分割。因而,例如,可在单件硬件(例如通用信号处理器或一块随机存取存储器、硬盘等)或多件硬件中实施功能框中的一个或多个(例如处理器或存储器)。类似地,程序可为独立的程序,可结合成操作系统中的例程,可为安装好的软件包中的函数等。应当理解,一些实施例不限于图中显示的布置和工具。
如本披露所用,以单数叙述或以词语“一个”或“一种”开头的要素或步骤应理解为不排除所述要素或步骤的复数,除非明确陈述了这种排除。此外,对“一个实施例”的引用不意于被解释为排除也结合了所叙述的特征的另外的实施例的存在。除非明确陈述了相反的情况,否则“包括”、“包含”或“具有”具有特定属性的要素或多个要素的实施例可包括不具有那个属性 的另外的这样的要素。
一些实施例提供如图3所示的音频信号处理装置的麦克风设置,图3示出了三个指向型麦克风31、32和33,它们整体上构成三重对称设置,三个指向型麦克风的轴线311、321、331(也即是垂直于其拾音平面中心的线)位于同一平面内,并且两两构成π2/3夹角。并且,指向型麦克风31、32、33的端部之间的距离范围D(如图中所示出的31和32之间)为0-5mm。作为优选,可以选择D=2mm。
另一些实施例提供如图4所示意的音频信号处理装置的麦克风设置,图4示出了三个叠加的指向型麦克风41、42和43,图4示出了“自上而下”的视角,三个指向型麦克风自上而下依次为41、42和43。指向型麦克风41、42和43的轴线(垂直于其拾音平面中心的线)平行于图4的平面。而如果将指向型麦克风41,42,43投影在图4的平面内,它们之间也构成三重对称设置,三个指向型麦克风的轴线411、421和431在图4的投影平面内两两构成π2/3夹角(图4右侧虚线轴线示意)。
另一些实施例提供如图5所示意的音频信号处理装置的麦克风设置,图5示出了示出了三个指向型麦克风51、52和53。这三个指向型麦克风之间形成三重对称设置。三个指向型麦克风的轴线511、521、531(垂直于其拾音平面中心的线)之间互相平行,并且轴线511、521、531在与它们垂直的平面内的三个投影点构成等边三角形T。并且,指向型麦克风51、52、53的端部之间的距离范围D(如图中所示出的51和52之间)为0-5mm。作为优选,可以选择D=2mm。
在上述实施例中,本领域技术人员可以选择合适的指向型麦克风来构成图3-5所示出的麦克风设置。指向型麦克风包括但不限于:心型指向型麦克风(Cardioid microphone)、亚心型指向型麦克风(Subcardioid microphone)、过心型指向型麦克风(Supercardioid microphone)、超心型指向型麦克风(Hypercardioid microphone)、偶极型指向型麦克风(Dipole microphone)来构成图3-5所示的麦克风设置。可以理解的是:可以选择同一种指向型麦克风,例如心型指向型麦克风来构成图3-5中的任意一种麦克风设置;也可以选择不同类型的指向型麦克风组合来构成图3-5中的任意一种麦克风设 置。
当采用了上述图3-5中所示的麦克风设置时,结合下面将要描述的本披露的算法,本披露的技术方案可以实现在任意方向的无损拾音效果,从而可以解决“偏离轴线”以及“WNG”的问题。
不同于传统方案中的由某一麦克风来拾音的设置,本披露的技术方案将同时地(simultaneously)拾取并组合来自多个麦克风的音频信号。在本披露技术方案中,将多个麦克风之间的距离设置地尽可能地小;从而可以尽可能地减少音频信号到达不同的麦克风之间的时间差,使得“同时”组合多个麦克风的音频信号在物理结构上首先成为可能。
在本披露技术中,通过“同时地”线性组合三个来自物理实体的麦克风(例如心脏型指向型麦克风)的信号来构成一个“虚拟麦克风(Virtual Microphone)”。线性组合的系数由矢量μ表示:
μ=inv(A)*b,其中:
b=[0 0 1]
T
θ
m代表波束角度(也即是希望所获得的音频信号的方向),而θ
n代表空角度(也即是不希望获得的音频信号的方向)。
在一些实施例中,如果希望线性组合三个麦克风的信号以构成一个虚拟的超心脏型指向型麦克风,选择θ
m和θ
n的关系为:
θ
n=θ
m+110*π/180
图6示出了在此设置下本披露技术方案在60度方向上的拾音效果。对比图1-1可以看出,本披露的技术方案,在60度方向上的拾音完全没有任何衰减。另外,不仅仅在60度方向,通过动态选择合适的θ
m,本披露的技术方案,可以实现在360度方向上均没有衰减的技术效果。
在另一些实施例中,如果希望线性组合三个麦克风的信号以构成一个虚拟的心脏型指向型麦克风,可以选择θ
m和θ
n的关系为:
θ
n=θ
m+π
通过上述算法以及选择合适的θ
m和θ
n的关系,本披露的算法和麦克风设置可以实现任意类型的虚拟一阶差分麦克风,包括心型指向型麦克风(Cardioid microphone)、次心型指向型麦克风(Subcardioid microphone)、过心型指向型麦克风(Supercardioid microphone)、超心型指向型麦克风(Hypercardioid microphone)、偶极型指向型麦克风(Dipole microphone)等。
另一方面,上述的音频信号的组合是独立于频率的,也即是说:波束形成模式对于任意频率都是相同的,从而本披露的技术方案不会“放大”低频段的白噪声,从而本披露的技术方案还可以解决WNG问题。
一旦虚拟麦克风的波束形成之后,波束选择算法进一步实时地比较并且从多个方向的虚拟波束中选择具有最高信噪比(SNR)的波束方向作为音频输出源。
图7示出了一些实施例中的波束选择算法的流程图,首先,在步骤71,通过短时傅里叶变换(Short-Time Fourier Transform)将音频信号帧变换为频域信号。
在步骤72,判断每个频率点(Frequency Bin)是否包含音频信号;如果没有,则直接进入步骤75,将频率点递增;如果有,进入到步骤73,在当前的频率点,选择具有最大信噪比的信号,记录与之相对应的波束指数。并且在步骤74和步骤75分别依次将最大信噪比信号数和频率间隔递增。
在步骤76,判断是否已经遍历当前的总频率点,如果不是,则重复上述的步骤72-75,如果是,则在步骤77从所有的虚拟波束中选择出具有最大SNR的信号,并在步骤78将上述具有最大SNR的信号作为语音信号输出。
图8示出了本披露技术方案所获得的音频信号图谱,其中红色谱线为本披露技术方案的虚拟麦克风所获得的音频信号,蓝色谱线为传统的物理麦克风所获得的音频信号,可以看出,在各个谱段,本披露的技术方案所获得信 号的SNR均优于传统技术,另一方面,本披露技术方案还可以解决WNG问题。
本披露的技术方案具有上述的技术优势以及因此而带来的广泛的应用优势。这些应用优势包括:
(1)极小的尺寸;目前最小的心脏型指向型麦克风的尺寸可以达到3mm*1.5mm(直径,厚度),在本披露的组合方式下,例如图3-5所示的麦克风组合设置的尺寸整体上可以控制在5mm的范围内,这使得采用本披露的各种装置可以获得体积优势;
(2)极高的信噪比;如上所述,采用本披露设置和算法的音频装置可以获得远远高于现有技术的信噪比;
(3)极大的有效拾音范围和易于组合性,采用本披露设置和算法的音频装置的有效拾音范围可以是现有技术装置的3x倍,因此,即使对于一个面积较大的会议室,采用菊链(Daisy chain)方式组合仅仅几个音频装置即可实现全部面积内的有效拾音。
在一些实施例中,在多方会议电话中采用本披露的麦克风设置和算法,从而可以解决在主发言者发言时在和主发言者不同的方位有别的参与人发出噪音(例如在打电话)的问题。可以实时动态地设置并选择将θ
m对准主发言者的方向,而将θ
n对准发出噪音的方向,从而可以只从主发言者方向获得音频信号,而噪音方向所发出的噪音完全不会被麦克风所拾取。
在一些实施例中,在语音购物装置中采用本披露的麦克风设置和算法,尤其是处于公共场合的语音购物装置(例如贩售机),从而可以解决在嘈杂的公共场合无法准确识别购物者的音频信号的问题。一方面,和上述相类似,实时动态地设置并选择将θ
m对准购物者发言的方向,另一方面,本披露的技术方案具有很好地对背景噪声的抑制作用,从而可以准确地拾取来自购物者的语音信号。
在一些实施例中,在智能音箱中采用本披露的麦克风设置和算法,尤其是在家庭环境中使用时,当周围存在噪声和别的语音信号源,和上述描述相类似,可以准确地拾取来自命令发出者的语音信号而避开来自噪声源的噪音,并且对背景音也有很好的抑制效果。
要理解的是,以上描述意于为示例性,而不是限制性的。例如,上面描述的实施例(和/或它们的各方面)可与彼此结合起来使用。另外,可在不偏离一些实施例的范围的情况下做出许多修改,以使具体情况或内容适于一些实施例的教导。虽然本文描述的材料的尺寸和类型意于限定一些实施例的参数,但实施例决不是限制性的,而是示例性实施例。在审阅以上描述之后,许多其它实施例对本领域技术人员将是显而易见的。因此,应当参照所附权利要求以及这样的权利要求所涵盖的等效体的全部范围来确定一些实施例的范围。在所附权利要求中,用语“包括”和“在其中”用作相应的用语“包含”和“其中”的易懂语言等效体。此外,在所附权利要求中,用语“第一”、“第二”和“第三”等仅作为标记使用,并且它们不意于对它们的对象施加数字要求。另外,不以手段加功能的格式来书写所附权利要求的限制,除非且直到这样的权利要求限制清楚地使用短语“用于…的手段”,跟随没有另外的结构的功能陈述。
还需要说明的是,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、商品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、商品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、商品或者设备中还存在另外的相同要素。
本领域内的技术人员应明白,本披露的一些实施例可提供为方法、设备、或计算机程序产品。因此,本披露可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本披露可采用在一个或多个其中包含有计算机可用程序代码的计算机可用存储介质(包括但不限于磁盘存储器、CD-ROM、光学存储器等)上实施的计算机程序产品的形式。
计算机可读介质包括永久性和非永久性、可移动和非可移动媒体可以由任何方法或技术来实现信息存储。信息可以是计算机可读指令、数据结构、程序的模块或其他数据。计算机的存储介质的例子包括,但不限于相变内存(PRAM)、静态随机存取存储器(SRAM)、动态随机存取存储器(DRAM)、其他类型的随机存取存储器(RAM)、只读存储器(ROM)、电可擦除可编程 只读存储器(EEPROM)、快闪记忆体或其他内存技术、只读光盘只读存储器(CD-ROM)、数字多功能光盘(DVD)或其他光学存储、磁盒式磁带,磁带磁磁盘存储或其他磁性存储设备或任何其他非传输介质,可用于存储可以被计算设备访问的信息。按照本文中的界定,计算机可读介质不包括暂存电脑可读媒体(transitory media),如调制的数据信号和载波。
本书面描述使用示例来公开一些实施例,包括最佳模式,并且还使本领域任何技术人员能够实践一些实施例,包括制造和使用任何装置或系统,以及实行任何结合的方法。一些实施例的保护范围由权利要求限定,并且可包括本领域技术人员想到的其它示例。如果这样的其它示例具有不异于权利要求的字面语言的结构要素,或者如果它们包括与权利要求的字面语言无实质性差异的等效结构要素,则它们意于处在权利要求的范围之内。
Claims (23)
- 一种音频信号处理装置,包括:多个麦克风;所述多个麦克风两两接近设置,并且所述多个麦克风形成对称性结构。
- 根据权利要求1所述的装置,其中,所述多个麦克风为三个。
- 根据权利要求2所述的装置,其中,所述的多个麦克风的轴线在同一水平面内的投影两两形成120度夹角。
- 根据权利要求3所述的装置,其中所述的多个麦克风的轴线位于同一水平面内,所述任意两个麦克风的轴线形成120度夹角。
- 根据权利要求3所述的装置,其中,所述的麦克风构成叠加模式。
- 根据权利要求2所述的装置,所述的多个麦克风的轴线两两平行并且所述多个轴线在其垂直平面内的投影点形成等边三角形的三个顶点。
- 根据权利要求1-6任一所述的装置,其中,所述任意两个麦克风的端部之间的距离范围为0-5mm。
- 根据权利要求7的装置,其中:所述的麦克风包括以下的至少一种:心型指向型麦克风(Cardioid microphone)、次心型指向型麦克风(Subcardioid microphone)、过心型指向型麦克风(Supercardioid microphone)、超心型指向型麦克风(Hypercardioid microphone)、偶极型指向型麦克风(Dipole microphone)。
- 一种音频信号处理方法,使用如权利要求1-8任一所述的装置,所述方法包括:同时线性组合由多个麦克风所获得的音频信号;基于组合后的音频信号,动态选择最佳拾音方向。
- 根据权利要求10所述的方法,其中,当以Hyper-cardioid microphone模式组合多个麦克风的音频信号时,θ n=θ m+110*π/180。
- 根据权利要求10所述的方法,其中,当以Cardioid microphone模式组合多个麦克风的音频信号时,θ n=θ m+π。
- 根据权利要求11或12所述的方法,还包括:基于设定的采样时间间隔连续地对组合后的音频信号进行处理,得到多个方向的音频信号;比较所述多个方向的音频信号,选择信噪比最高的方向作为拾音方向。
- 根据权利要求13所述的方法,其中,采用短时傅里叶变换对组合后的音频信号进行处理。
- 根据权利要求14所述的方法,其中,所述设定的采样时间间隔为10-20ms。
- 根据权利要求13所述的方法,还包括:基于选择出的拾音方向获取音频信号并输出。
- 一种多方会议电话,其特征在于:包括如权利要求1-8任一所述的装置。
- 根据权利要求17所述的多方会议电话,其特征在于:使用权利要求9-16任一所述的方法。
- 一种语音购物装置,其特征在于:包括如权利要求1-8任一所述的装置。
- 根据权利要求19所述的语音购物装置,其特征在于:使用权利要求9-16任一所述的方法。
- 一种智能音箱,其特征在于:包括如权利要求1-8任一所述的装置。
- 根据权利要求21所述的智能音箱,其特征在于:使用权利要求9-16任一所述的方法。
- 一种音频信号处理装置,其包括:处理器和非暂态存储介质,所述非暂态存储介质存储有指令集,所述指令集由处理器执行时使得所述装置能够执行如权利要求9-16任一所述的方法。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2018/100464 WO2020034095A1 (zh) | 2018-08-14 | 2018-08-14 | 音频信号处理装置及方法 |
| CN201880094783.0A CN112292870A (zh) | 2018-08-14 | 2018-08-14 | 音频信号处理装置及方法 |
| US17/143,787 US11778382B2 (en) | 2018-08-14 | 2021-01-07 | Audio signal processing apparatus and method |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2018/100464 WO2020034095A1 (zh) | 2018-08-14 | 2018-08-14 | 音频信号处理装置及方法 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/143,787 Continuation US11778382B2 (en) | 2018-08-14 | 2021-01-07 | Audio signal processing apparatus and method |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020034095A1 true WO2020034095A1 (zh) | 2020-02-20 |
Family
ID=69524631
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/100464 Ceased WO2020034095A1 (zh) | 2018-08-14 | 2018-08-14 | 音频信号处理装置及方法 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US11778382B2 (zh) |
| CN (1) | CN112292870A (zh) |
| WO (1) | WO2020034095A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220028404A1 (en) * | 2019-02-12 | 2022-01-27 | Alibaba Group Holding Limited | Method and system for speech recognition |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102227918A (zh) * | 2008-12-17 | 2011-10-26 | 雅马哈株式会社 | 声音收集装置 |
| CN203608356U (zh) * | 2013-12-02 | 2014-05-21 | 吴东亮 | 一种用于会议室的阵列话筒 |
| CN105764011A (zh) * | 2016-04-08 | 2016-07-13 | 甄钊 | 用于3d沉浸式环绕声音乐与影视拾音的传声器阵列装置 |
| CN106842131A (zh) * | 2017-03-17 | 2017-06-13 | 浙江宇视科技有限公司 | 麦克风阵列声源定位方法及装置 |
Family Cites Families (23)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6584203B2 (en) | 2001-07-18 | 2003-06-24 | Agere Systems Inc. | Second-order adaptive differential microphone array |
| KR100499124B1 (ko) | 2002-03-27 | 2005-07-04 | 삼성전자주식회사 | 직교 원형 마이크 어레이 시스템 및 이를 이용한 음원의3차원 방향을 검출하는 방법 |
| GB0321722D0 (en) | 2003-09-16 | 2003-10-15 | Mitel Networks Corp | A method for optimal microphone array design under uniform acoustic coupling constraints |
| US7515721B2 (en) | 2004-02-09 | 2009-04-07 | Microsoft Corporation | Self-descriptive microphone array |
| US8090117B2 (en) | 2005-03-16 | 2012-01-03 | James Cox | Microphone array and digital signal processing system |
| EP1994788B1 (en) * | 2006-03-10 | 2014-05-07 | MH Acoustics, LLC | Noise-reducing directional microphone array |
| GB0619825D0 (en) * | 2006-10-06 | 2006-11-15 | Craven Peter G | Microphone array |
| US8903106B2 (en) | 2007-07-09 | 2014-12-02 | Mh Acoustics Llc | Augmented elliptical microphone array |
| US9326064B2 (en) | 2011-10-09 | 2016-04-26 | VisiSonics Corporation | Microphone array configuration and method for operating the same |
| EP2592845A1 (en) | 2011-11-11 | 2013-05-15 | Thomson Licensing | Method and Apparatus for processing signals of a spherical microphone array on a rigid sphere used for generating an Ambisonics representation of the sound field |
| EP2848007B1 (en) * | 2012-10-15 | 2021-03-17 | MH Acoustics, LLC | Noise-reducing directional microphone array |
| US9197962B2 (en) | 2013-03-15 | 2015-11-24 | Mh Acoustics Llc | Polyhedral audio system based on at least second-order eigenbeams |
| CN104464739B (zh) * | 2013-09-18 | 2017-08-11 | 华为技术有限公司 | 音频信号处理方法及装置、差分波束形成方法及装置 |
| US9734822B1 (en) * | 2015-06-01 | 2017-08-15 | Amazon Technologies, Inc. | Feedback based beamformed signal selection |
| KR20170035504A (ko) * | 2015-09-23 | 2017-03-31 | 삼성전자주식회사 | 전자 장치 및 전자 장치의 오디오 처리 방법 |
| US9961437B2 (en) | 2015-10-08 | 2018-05-01 | Signal Essence, LLC | Dome shaped microphone array with circularly distributed microphones |
| EP3420735B1 (en) * | 2016-02-25 | 2020-06-10 | Dolby Laboratories Licensing Corporation | Multitalker optimised beamforming system and method |
| WO2017174136A1 (en) * | 2016-04-07 | 2017-10-12 | Sonova Ag | Hearing assistance system |
| US10356514B2 (en) * | 2016-06-15 | 2019-07-16 | Mh Acoustics, Llc | Spatial encoding directional microphone array |
| US10477304B2 (en) * | 2016-06-15 | 2019-11-12 | Mh Acoustics, Llc | Spatial encoding directional microphone array |
| US10827263B2 (en) * | 2016-11-21 | 2020-11-03 | Harman Becker Automotive Systems Gmbh | Adaptive beamforming |
| US10304475B1 (en) * | 2017-08-14 | 2019-05-28 | Amazon Technologies, Inc. | Trigger word based beam selection |
| US9973849B1 (en) * | 2017-09-20 | 2018-05-15 | Amazon Technologies, Inc. | Signal quality beam selection |
-
2018
- 2018-08-14 CN CN201880094783.0A patent/CN112292870A/zh active Pending
- 2018-08-14 WO PCT/CN2018/100464 patent/WO2020034095A1/zh not_active Ceased
-
2021
- 2021-01-07 US US17/143,787 patent/US11778382B2/en active Active
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102227918A (zh) * | 2008-12-17 | 2011-10-26 | 雅马哈株式会社 | 声音收集装置 |
| CN203608356U (zh) * | 2013-12-02 | 2014-05-21 | 吴东亮 | 一种用于会议室的阵列话筒 |
| CN105764011A (zh) * | 2016-04-08 | 2016-07-13 | 甄钊 | 用于3d沉浸式环绕声音乐与影视拾音的传声器阵列装置 |
| CN106842131A (zh) * | 2017-03-17 | 2017-06-13 | 浙江宇视科技有限公司 | 麦克风阵列声源定位方法及装置 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220028404A1 (en) * | 2019-02-12 | 2022-01-27 | Alibaba Group Holding Limited | Method and system for speech recognition |
| US12315527B2 (en) * | 2019-02-12 | 2025-05-27 | Alibaba Group Holding Limited | Method and system for speech recognition |
Also Published As
| Publication number | Publication date |
|---|---|
| CN112292870A (zh) | 2021-01-29 |
| US11778382B2 (en) | 2023-10-03 |
| US20210127208A1 (en) | 2021-04-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN108370470B (zh) | 会议系统以及会议系统中的语音获取方法 | |
| US11496830B2 (en) | Methods and systems for recording mixed audio signal and reproducing directional audio | |
| US8787587B1 (en) | Selection of system parameters based on non-acoustic sensor information | |
| US9161149B2 (en) | Three-dimensional sound compression and over-the-air transmission during a call | |
| KR101547035B1 (ko) | 다중 마이크에 의한 3차원 사운드 포착 및 재생 | |
| CN110337819A (zh) | 来自设备中具有不对称几何形状的多个麦克风的空间元数据的分析 | |
| JP2020500480A5 (zh) | ||
| BR112015014380B1 (pt) | Filtro e método para filtragem espacial informada utilizando múltiplas estimativas da direção de chegada instantânea | |
| CN104010265A (zh) | 音频空间渲染设备及方法 | |
| CN115547354B (zh) | 波束形成方法、装置及设备 | |
| US10332530B2 (en) | Coding of a soundfield representation | |
| CN110035372A (zh) | 扩声系统的输出控制方法、装置、扩声系统及计算机设备 | |
| US20250193611A1 (en) | Ear-worn device with neural network for noise reduction and/or spatial focusing using multiple input audio signals | |
| CN118435278A (zh) | 用于提供空间音频的装置、方法和计算机程序 | |
| US11778382B2 (en) | Audio signal processing apparatus and method | |
| Ba et al. | Enhanced MVDR beamforming for arrays of directional microphones | |
| CN115515038A (zh) | 波束形成方法、装置及设备 | |
| CN110858485A (zh) | 语音增强方法、装置、设备及存储介质 | |
| CN115508777B (zh) | 说话人定位方法、装置及设备 | |
| CN115424633B (zh) | 说话人定位方法、装置及设备 | |
| US12548582B2 (en) | Apparatus, methods and computer programs for audio focusing | |
| US12621389B2 (en) | Conference terminal and echo cancellation method | |
| US20250141998A1 (en) | Conference terminal and echo cancellation method | |
| KR102343811B1 (ko) | 음성 검출 방법 | |
| US20250203310A1 (en) | Spatial Audio Processing |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18930377 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18930377 Country of ref document: EP Kind code of ref document: A1 |


