WO2022052256A1 - 音频信号的处理方法、终端设备及存储介质 - Google Patents
音频信号的处理方法、终端设备及存储介质 Download PDFInfo
- Publication number
- WO2022052256A1 WO2022052256A1 PCT/CN2020/125633 CN2020125633W WO2022052256A1 WO 2022052256 A1 WO2022052256 A1 WO 2022052256A1 CN 2020125633 W CN2020125633 W CN 2020125633W WO 2022052256 A1 WO2022052256 A1 WO 2022052256A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- transfer function
- audio signal
- low
- frequency transfer
- frequency
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R1/00—Details of transducers, loudspeakers or microphones
- H04R1/10—Earpieces; Attachments therefor ; Earphones; Monophonic headphones
- H04R1/1091—Details not provided for in groups H04R1/1008 - H04R1/1083
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R3/00—Circuits for transducers
- H04R3/04—Circuits for transducers for correcting frequency response
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R1/00—Details of transducers, loudspeakers or microphones
- H04R1/20—Arrangements for obtaining desired frequency or directional characteristics
- H04R1/22—Arrangements for obtaining desired frequency or directional characteristics for obtaining desired frequency characteristic only
- H04R1/222—Arrangements for obtaining desired frequency or directional characteristics for obtaining desired frequency characteristic only for microphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R3/00—Circuits for transducers
- H04R3/005—Circuits for transducers for combining the signals of two or more microphones
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04R—LOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
- H04R2460/00—Details of hearing devices, i.e. of ear- or headphones covered by H04R1/10 or H04R5/033 but not provided for in any of their subgroups, or of hearing aids covered by H04R25/00 but not provided for in any of its subgroups
- H04R2460/13—Hearing devices using bone conduction transducers
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02D—CLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
- Y02D30/00—Reducing energy consumption in communication networks
- Y02D30/70—Reducing energy consumption in communication networks in wireless communication networks
Definitions
- the present invention relates to the technical field of audio processing, and in particular, to an audio signal processing method, a terminal device and a computer-readable storage medium.
- Bone conduction is a method of sound conduction, which converts sound into mechanical vibrations of different frequencies, and transmits sound waves through the human skull, bone labyrinth, inner ear lymph, auger, and auditory center. Compared with the classic sound conduction method that generates sound waves through the diaphragm, bone conduction saves many steps of sound wave transmission, and can achieve clear sound restoration in a noisy environment, and sound waves will not affect others due to diffusion in the air. . Due to the above advantages of bone conduction, devices for collecting user voice signals based on bone conduction appear.
- the audio signal collected by the bone conduction vibration pickup device may suffer from serious high-frequency attenuation. As a result, the audio signal collected by the bone conduction audio collection device is incomplete.
- the main purpose of the present invention is to provide an audio signal processing method, a terminal device and a computer-readable storage medium, so as to achieve the purpose of improving the integrity of the voice signal collected by the terminal device.
- the present invention provides a processing method of an audio signal, and the processing method of the audio signal comprises the following steps:
- the bone conduction audio signal is extended in frequency domain based on the low frequency transfer function and the high frequency transfer function, and the extended initial audio is used as output audio.
- the step of obtaining the low-frequency transfer function matched by the bone conduction audio signal and the high-frequency transfer function corresponding to the low-frequency transfer function includes:
- the low-frequency transfer function matching the low-frequency characteristic and the high-frequency transfer function corresponding to the low-frequency transfer function are acquired.
- the audio signal processing method is applied to a terminal device, and the step of acquiring the low-frequency transfer function matching the low-frequency characteristic and the high-frequency transfer function corresponding to the low-frequency transfer function includes:
- the server is configured to obtain, according to the received low-frequency characteristic, the low-frequency transfer function stored in the cloud database that matches the low-frequency characteristic, and the low-frequency transfer function the high-frequency transfer function corresponding to the function, and sending the low-frequency transfer function and the high-frequency transfer function to the terminal device;
- the low-frequency transfer function and the high-frequency transfer function sent by the server are received.
- the audio signal processing method is applied to a terminal device, and the step of acquiring a low-frequency transfer function matched by the bone conduction audio signal and a high-frequency transfer function corresponding to the low-frequency transfer function includes:
- the server is configured to obtain, according to the received bone conduction audio signal, the low frequency stored in the cloud database that matches the initial audio, and the low frequency transfer function the corresponding high-frequency transfer function, and send the acquired low-frequency transfer function and the high-frequency transfer function to the terminal device;
- the low-frequency transfer function and the high-frequency transfer function sent by the server are received.
- the low-frequency transfer function and the high-frequency transfer function are in one-to-one correspondence, and are associated and stored in a database.
- the method further includes:
- the low frequency characteristic of the initial audio, the low frequency transfer function and the high frequency transfer function are stored in association.
- the step of determining the low-frequency transfer function and the high-frequency transfer function according to the bone conduction audio signal and the air conduction audio signal includes:
- the low frequency transfer function is determined from the first low frequency characteristic and the second low frequency characteristic
- the high frequency transfer function is determined from the first high frequency characteristic and the second high frequency characteristic
- the low-frequency transfer function and the high-frequency transfer function are obtained based on a training speech signal, wherein the training speech signal includes the bone conduction audio signal and the air conduction audio signal corresponding to the same speech.
- the present invention also provides a terminal device, the terminal device includes: a memory, a processor, and an audio signal processing program stored in the memory and running on the processor, the The audio signal processing program implements the steps of the audio signal processing method as described above when executed by the processor.
- the present invention also provides a computer-readable storage medium, where an audio signal processing program is stored on the computer-readable storage medium, and the audio signal processing program is executed by a processor to realize the above-mentioned The steps of the audio signal processing method.
- An audio signal processing method, a terminal device, and a computer-readable storage medium proposed by an embodiment of the present invention first acquire a bone conduction audio signal collected by a bone conduction vibration pickup device, and then acquire a low-frequency transfer function matched by the bone conduction audio signal. , and the high-frequency transfer function corresponding to the low-frequency transfer function, wherein the low-frequency transfer function and the high-frequency transfer function are the corresponding transfer functions between the bone conduction audio signal and the air conduction audio signal corresponding to the same voice, and
- the bone conduction audio signal is expanded in the frequency domain based on the low-frequency transfer function and the high-frequency transfer function, and the expanded initial audio is used as the output audio.
- the high-frequency part of the bone conduction audio signal collected by the bone conduction vibration pickup device is expanded, so that the terminal device can obtain a complete audio signal based on the initial audio frequency collected by the bone conduction vibration pickup device. This achieves the effect of improving the integrity of the audio signal collected by the device.
- FIG. 1 is a schematic diagram of a terminal structure of a hardware operating environment involved in an embodiment of the present invention
- FIG. 2 is a schematic flowchart of an embodiment of an audio signal processing method according to the present invention.
- FIG. 3 is a schematic flowchart of another embodiment of an audio signal processing method according to the present invention.
- the audio signal collected by the bone conduction vibration pickup device may suffer from serious high-frequency attenuation. As a result, the audio signal collected by the bone conduction audio collection device is incomplete.
- an embodiment of the present invention proposes an audio signal processing method, and its main solution includes the following steps:
- the bone conduction audio signal is extended in the frequency domain based on the low-frequency transfer function and the high-frequency transfer function, and the extended initial audio is used as output audio.
- the terminal device can obtain a complete audio signal based on the initial audio collected by the bone conduction vibration pickup device. . This achieves the effect of improving the integrity of the audio signal collected by the device.
- FIG. 1 is a schematic diagram of a terminal structure of a hardware operating environment involved in an embodiment of the present invention.
- the terminal in this embodiment of the present invention may be a terminal device such as a bone conduction headset.
- the terminal may include: a processor 1001 , such as a CPU, a network interface 1004 , a user interface 1003 , a memory 1005 , and a communication bus 1002 .
- the communication bus 1002 is used to realize the connection and communication between these components.
- the user interface 1003 may include a bone conduction vibration pickup device, etc., and the optional user interface 1003 may also include a standard wired interface and a wireless interface.
- the network interface 1004 may include a standard wired interface and a wireless interface (eg, a WI-FI interface).
- the memory 1005 may be high-speed RAM memory, or may be non-volatile memory, such as disk memory.
- the memory 1005 may also be a storage device independent of the aforementioned processor 1001 .
- terminal structure shown in FIG. 1 does not constitute a limitation on the terminal, and may include more or less components than the one shown, or combine some components, or arrange different components.
- the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module and an audio signal processing program.
- the network interface 1004 is mainly used to connect to the background server, and perform data communication with the background server;
- the processor 1001 can be used to call the audio signal processing program stored in the memory 1005, and perform the following operations:
- the bone conduction audio signal is extended in the frequency domain based on the low-frequency transfer function and the high-frequency transfer function, and the extended initial audio is used as output audio.
- processor 1001 can call the audio signal processing program stored in the memory 1005, and also perform the following operations:
- the low-frequency transfer function matching the low-frequency characteristic and the high-frequency transfer function corresponding to the low-frequency transfer function are acquired.
- processor 1001 can call the audio signal processing program stored in the memory 1005, and also perform the following operations:
- the server is configured to obtain, according to the received low-frequency characteristic, the low-frequency transfer function stored in the cloud database that matches the low-frequency characteristic, and the low-frequency transfer function the high-frequency transfer function corresponding to the function, and sending the low-frequency transfer function and the high-frequency transfer function to the terminal device;
- the low-frequency transfer function and the high-frequency transfer function sent by the server are received.
- processor 1001 can call the audio signal processing program stored in the memory 1005, and also perform the following operations:
- the server is configured to obtain, according to the received bone conduction audio signal, the low frequency stored in the cloud database that matches the initial audio, and the low frequency transfer function the corresponding high-frequency transfer function, and send the acquired low-frequency transfer function and the high-frequency transfer function to the terminal device;
- the low-frequency transfer function and the high-frequency transfer function sent by the server are received.
- processor 1001 can call the audio signal processing program stored in the memory 1005, and also perform the following operations:
- the low frequency characteristic of the initial audio, the low frequency transfer function and the high frequency transfer function are stored in association.
- processor 1001 can call the audio signal processing program stored in the memory 1005, and also perform the following operations:
- the low frequency transfer function is determined from the first low frequency characteristic and the second low frequency characteristic
- the high frequency transfer function is determined from the first high frequency characteristic and the second high frequency characteristic
- the audio signal processing method includes the following steps:
- Step S10 acquiring the bone conduction audio signal collected by the bone conduction vibration pickup device
- Step S20 obtaining the low-frequency transfer function matched by the bone conduction audio signal, and the high-frequency transfer function corresponding to the low-frequency transfer function, wherein the low-frequency transfer function and the high-frequency transfer function are bone conduction corresponding to the same voice
- Step S30 Perform frequency domain expansion on the bone conduction audio signal based on the low-frequency transfer function and the high-frequency transfer function, and use the expanded initial audio as an output audio.
- Bone conduction is a method of sound conduction, which converts sound into mechanical vibrations of different frequencies, and transmits sound waves through the human skull, bone labyrinth, inner ear lymph, auger, and auditory center. Compared with the classic sound conduction method that generates sound waves through the diaphragm, bone conduction saves many steps of sound wave transmission, and can achieve clear sound restoration in a noisy environment, and sound waves will not affect others due to diffusion in the air. . Due to the above advantages of bone conduction, devices for collecting user voice signals based on bone conduction appear.
- the audio signal collected by the bone conduction vibration pickup device may suffer from serious high-frequency attenuation. As a result, the audio signal collected by the bone conduction audio collection device is incomplete.
- the embodiment of the present invention provides an audio signal processing method.
- the above-mentioned audio signal processing method is applied to a terminal device, and the above-mentioned terminal device is provided with a bone conduction shock pickup device.
- the above-mentioned bone conduction vibration pickup device is used to collect the vibration wave of the object that vibrates due to the sound source voice, and convert the vibration wave into a voice signal.
- the above-mentioned terminal device is set as a bone conduction earphone.
- the bone conduction vibration pickup device disposed on the bone conduction earphone can The vibration waves of the user's limbs are collected, and the vibration waves are converted into audio signals.
- the terminal device When the terminal device is provided with a bone conduction vibration pickup device, audio can be collected through the bone conduction vibration pickup device, and then the initial audio frequency collected by the bone conduction vibration pickup device is acquired.
- a pre-stored low frequency transfer function matching the bone conduction audio signal and a pre-stored high frequency transfer function corresponding to the low frequency transfer function can be acquired.
- the above-mentioned high-frequency transfer function and the above-mentioned low-frequency transfer function are pre-stored data.
- the low-frequency transfer function and the high-frequency transfer function are transfer functions corresponding to the bone conduction audio signal and the air conduction audio signal corresponding to the same sound source.
- the low frequency characteristic of the initial audio can be acquired. Then query the pre-stored low frequency matching the low frequency characteristic in the database
- the collected signal is first converted to convert the time domain signal into a frequency domain signal.
- the terminal device may first acquire the bone conduction audio signal collected by the bone conduction vibration pickup device, wherein the bone conduction audio signal directly acquired by the terminal device is a time domain signal. Therefore, after obtaining the above-mentioned initial audio, the bone conduction audio signal can be converted from a time-domain signal to a frequency-domain signal based on FFT (fast Fourier transform, fast Fourier transform). Further, signal analysis can be performed on the converted bone conduction audio signal, so as to extract the low frequency characteristic of the bone conduction audio signal.
- FFT fast Fourier transform, fast Fourier transform
- the pre-stored low-frequency characteristic and the low-frequency transfer function may be associated and stored in the database.
- the database also stores the high-frequency transfer functions that correspond to the low-frequency transfer functions one-to-one and are stored in association. Therefore, after the low-frequency transfer function is obtained, the high-frequency transfer function corresponding to the low-frequency transfer function can also be obtained. function.
- the low-frequency characteristic may also be sent to a server, wherein the server is configured to, according to the received low-frequency characteristic, Obtain the low-frequency transfer function that matches the low-frequency characteristic and the high-frequency transfer function corresponding to the low-frequency transfer function stored in the cloud database, and send the low-frequency transfer function and the high-frequency transfer function to the terminal device, and then receive the low-frequency transfer function and the high-frequency transfer function sent by the server. Since the above-mentioned transfer functions are stored in the cloud server, the database administrator can more conveniently update the above-mentioned low-frequency transfer functions and high-frequency transfer functions stored in the database, so as to adapt to more application scenarios and users with different sound characteristics.
- the above-mentioned low-frequency transfer function and the high-frequency transfer function are obtained based on a training speech signal, wherein the training speech signal includes the bone conduction audio signal and the air conduction audio signal corresponding to the same speech.
- the voice output from the same sound source can be simultaneously collected through the bone conduction vibration pickup device and the microphone.
- the audio signal corresponding to the speech collected by the bone conduction vibration pickup device is described as a bone conduction audio signal.
- the audio signal corresponding to the audio collected by the microphone or the like is described as an air conduction audio signal (that is, the audio signal corresponding to the sound wave propagating through the air collected by the microphone).
- the bone conduction audio signal and the air conduction audio signal are then converted from time domain signals to frequency domain signals.
- the low frequency transfer function is determined according to the first low frequency characteristic and the second low frequency characteristic
- the high frequency transfer function is determined according to the first high frequency characteristic and the second high frequency characteristic.
- the bone conduction audio signal and the air conduction audio signal can be recorded at the same time, and the bone conduction audio signal and the air conduction audio signal can be converted from the time domain to the frequency domain through FFT, and extracted
- the frequency, amplitude and other characteristics of the frequency domain signal determine the low frequency characteristics ( ⁇ preset frequency value, such as 2KHz) and high frequency characteristics (> preset frequency value) generated by each voice.
- the low-frequency characteristics are generally only changes in harmonic energy, while the high-frequency characteristics are complex changes such as attenuation and frequency loss.
- x 1 (n) represents the bone conduction audio signal
- the low frequency part is represented in the frequency domain
- x 2 (n) represents the bone conduction audio signal
- the high frequency part is represented in the frequency domain
- y 1 (n) represents the air conduction In the audio signal
- the low frequency part is represented in the frequency domain
- y 2 (n) represents the air conduction audio signal
- the high frequency part is represented in the frequency domain.
- the frequency domain transfer function h 1 (n) corresponding to the low frequency part can be calculated according to the following formula:
- h 1 (n) x 1 (n)/y 1 (n);
- the time domain transfer function h 1 (t) corresponding to the low frequency part and the time domain transfer function h 2 (t) corresponding to the high frequency part can be obtained through inverse transformation.
- the time domain transfer function h 1 (t) corresponding to the low frequency part and the time domain transfer function h 2 (t) corresponding to the high frequency part can only be used in pairs and cannot be used alone.
- each group of training speech can generate the time domain transfer function h 1 (t) corresponding to the low frequency part and the time domain transfer function h 2 (t) corresponding to the high frequency part, and store them in the database.
- the The audio signal is expanded in the frequency domain.
- the bone conduction audio signal x(t) can be extended in the frequency domain based on the following formula:
- y(t), h1(t) and h2(t) are the expanded audio signal, the low-frequency transfer function and the high-frequency transfer function, respectively.
- the expanded bone conduction audio signal can be used as the output audio of the terminal device.
- the terminal device is a bone conduction earphone
- the above-mentioned bone conduction earphone can be used as the audio collection device of the mobile terminal, and then the bone conduction earphone can send the expanded bone conduction audio signal as output audio to the terminal device connected to itself.
- the low-frequency transfer function and the high-frequency transfer function of the initial audio matching may be dynamically updated according to the initial audio received in real time. For example, 10-50 updates per second.
- the bone conduction audio signal collected by the bone conduction vibration pickup device is obtained first, and then the low frequency transfer function matched with the bone conduction audio signal and the high frequency transfer function corresponding to the low frequency transfer function are obtained.
- the low-frequency transfer function and the high-frequency transfer function are the corresponding transfer functions between the bone conduction audio signal and the air conduction audio signal corresponding to the same voice, and are based on the low-frequency transfer function and the high-frequency transfer function.
- the frequency domain expansion is performed on the bone conduction audio signal, and the expanded initial audio frequency is used as the output audio frequency, because the low frequency transfer function and the high frequency transfer function can be used to expand the bone conduction audio signal collected by the bone conduction vibration pickup device.
- the high-frequency part enables the terminal device to obtain a complete audio signal based on the initial audio collected by the bone conduction vibration pickup device. This achieves the effect of improving the integrity of the audio signal collected by the device.
- the step S20 includes:
- Step S21 sending the low-frequency characteristic to the server, wherein the server is set to obtain the low-frequency transfer function stored in the cloud database and matching the low-frequency characteristic according to the received low-frequency characteristic, and the the high-frequency transfer function corresponding to the low-frequency transfer function, and sending the low-frequency transfer function and the high-frequency transfer function to the terminal device;
- Step S22 Receive the low-frequency transfer function and the high-frequency transfer function sent by the server.
- the bone conduction audio signal may be sent to a server, wherein the server is configured to acquire, according to the received bone conduction audio signal, a low frequency stored in a cloud database that matches the initial audio , and the high-frequency transfer function corresponding to the low-frequency transfer function, and send the acquired low-frequency transfer function and the high-frequency transfer function to the terminal device, and receive the low-frequency transfer function sent by the server. function and the high frequency transfer function.
- the server compares the low-frequency characteristic of the above-mentioned initial audio with the pre-stored low-frequency characteristic.
- the pre-stored low-frequency characteristic may be the low-frequency characteristic of the bone conduction audio signal in the pre-stored training speech signal.
- the server can obtain the low-frequency transfer function and the high-frequency transfer function determined based on the training voice signal that match the above-mentioned bone conduction audio signal. And send the acquired low-frequency transfer function and high-frequency transfer function to the terminal device.
- an embodiment of the present invention also provides a terminal device, the terminal device includes a memory, a processor, and an audio signal processing program stored on the memory and running on the processor, the audio signal processing program When executed by the processor, the steps of the audio signal processing method described in each of the above embodiments are implemented.
- the terminal device is a bone conduction earphone.
- an embodiment of the present invention further provides a computer-readable storage medium, where an audio signal processing program is stored on the computer-readable storage medium, and when the audio signal processing program is executed by a processor, the above-described various embodiments are implemented. The steps of the audio signal processing method.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Human Computer Interaction (AREA)
- Computational Linguistics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Otolaryngology (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- General Health & Medical Sciences (AREA)
- Quality & Reliability (AREA)
- Details Of Audible-Bandwidth Transducers (AREA)
Abstract
一种音频信号的处理方法,包括以下步骤:获取骨传导拾震器件采集的骨导音频信号(S10);获取骨导音频信号匹配的低频传递函数,以及低频传递函数对应的高频传递函数,其中,低频传递函数和高频传递函数为同一语音对应的骨导音频信号与气导音频信号之间对应的传递函数(S20);基于低频传递函数及高频传递函数对骨导音频信号进行频域扩展,并将扩展后的初始音频作为输出音频(S30)。达成了提升终端设备采集的语音信号的完整性的效果。
Description
本申请要求于2020年9月10日提交中国专利局、申请号为202010953528.6、发明名称为“音频信号的处理方法、终端设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本发明涉及音频处理技术领域,尤其涉及音频信号的处理方法、终端设备及计算机可读存储介质。
骨传导是一种声音传导方式,即将声音转化为不同频率的机械振动,通过人的颅骨、骨迷路、内耳淋巴液、螺旋器、听觉中枢来传递声波。相对于通过振膜产生声波的经典声音传导方式,骨传导省去了许多声波传递的步骤,能在嘈杂的环境中实现清晰的声音还原,而且声波也不会因为在空气中扩散而影响到他人。由于骨传导存在上述优势,因而出现了基于骨传导采集用户声音信号的设备。
但是,在现有的骨传导设备中,源于骨传导拾震器件自身的硬件缺陷,会导致通过骨传导拾震器件采集的音频信号出现高频衰减严重的现象。这样导致骨传导音频采集设备采集的音频信号不完整。
上述内容仅用于辅助理解本发明的技术方案,并不代表承认上述内容是现有技术。
发明内容
本发明的主要目的在于提供一种音频信号的处理方法、终端设备及计算机可读存储介质,旨在达成提升终端设备采集的语音信号的完整性的目的。
为实现上述目的,本发明提供一种音频信号的处理方法,所述音频信号的处理方法包括以下步骤:
获取骨传导拾震器件采集的骨导音频信号;
获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数,其中,所述低频传递函数和所述高频传递函数为同一语音对应的骨导音频信号与气导音频信号之间对应的传递函数;
基于所述低频传递函数及所述高频传递函数对所述骨导音频信号进行频域扩展,并将 扩展后的所述初始音频作为输出音频。
可选地,所述获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数的步骤包括:
获取所述骨导音频信号的低频特性;
获取与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数。
可选地,所述音频信号的处理方法应用于终端设备,所述获取与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数的步骤包括:
将所述低频特性发送至服务器,其中,所述服务器设置为根据接收到的所述低频特性,获取云端数据库中保存的,与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数,并将所述低频传递函数及所述高频传递函数发送至所述终端设备;
接收所述服务器发送的所述低频传递函数及所述高频传递函数。
可选地,所述音频信号的处理方法应用于终端设备,所述获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数的步骤包括:
将所述骨导音频信号发送至服务器,其中,所述服务器设置为根据接收到的所述骨导音频信号,获取云端数据库中保存的与所述初始音频匹配的低频,以及所述低频传递函数对应的所述高频传递函数,并将获取到的所述低频传递函数和所述高频传递函数发送至所述终端设备;
接收所述服务器发送的所述低频传递函数和所述高频传递函数。
可选地,所述低频传递函数与所述高频传递函数一一对应,关联保存于数据库中。
可选地,所述获取骨传导拾震器件采集的骨导音频信号的步骤之后,还包括:
获取麦克风采集到的气导音频信号;
根据所述骨导音频信号及所述气导音频信号确定所述低频传递函数及所述高频传递函数;
关联保存所述初始音频的低频特性、所述低频传递函数及所述高频传递函数。
可选地,所述根据所述骨导音频信号及所述气导音频信号确定所述低频传递函数及所述高频传递函数的步骤包括:
获取所述骨导音频信号的第一低频特性,以及所述气导音频信号的第二低频特性;
获取所述骨导音频信号的第一高频特性,以及所述气导音频信号的第二高频特性;
根据第一低频特性及所述第二低频特性确定所述低频传递函数,以及根据所述第一高频特性和第二高频特性确定所述高频传递函数。
可选地,所述低频传递传递函数及所述高频传递函数基于训练语音信号得到,其中,所述训练语音信号包括同一语音对应的所述骨传导音频信号和所述气导音频信号。
此外,为实现上述目的,本发明还提供一种终端设备,所述终端设备包括:存储器、处理器及存储在所述存储器上并可在所述处理器上运行的音频信号处理程序,所述音频信号处理程序被所述处理器执行时实现如上所述的音频信号的处理方法的步骤。
此外,为实现上述目的,本发明还提供一种计算机可读存储介质,所述计算机可读存储介质上存储有音频信号处理程序,所述音频信号处理程序被处理器执行时实现如上所述的音频信号的处理方法的步骤。
本发明实施例提出的一种音频信号的处理方法、终端设备及计算机可读存储介质,先获取骨传导拾震器件采集的骨导音频信号,然后获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数,其中,所述低频传递函数和所述高频传递函数为同一语音对应的骨导音频信号与气导音频信号之间对应的传递函数,并基于所述低频传递函数及所述高频传递函数对所述骨导音频信号进行频域扩展,并将扩展后的所述初始音频作为输出音频,由于可以通过低频传递函数和高频传递函数,扩展骨传导拾震器件采集的骨导音频信号的高频部分,从而使得终端设备可以基于骨传导拾震器件采集到的初始音频得到完整的音频信号。这样达成了提高设备采集的音频信号的完整性的效果。
图1是本发明实施例方案涉及的硬件运行环境的终端结构示意图;
图2为本发明音频信号的处理方法的一实施例的流程示意图;
图3为本发明音频信号的处理方法的另一实施例的流程示意图。
本发明目的的实现、功能特点及优点将结合实施例,参照附图做进一步说明。
应当理解,此处所描述的具体实施例仅仅用以解释本发明,并不用于限定本发明。
由于在现有的骨传导设备中,源于骨传导拾震器件自身的硬件缺陷,会导致通过骨传导拾震器件采集的音频信号出现高频衰减严重的现象。这样导致骨传导音频采集设备采集的音频信号不完整。
为解决上述缺陷,本发明实施例提出一种音频信号的处理方法,其主要解决方案包括以下步骤:
获取骨传导拾震器件采集的骨导音频信号;
获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数,其中,所述低频传递函数和所述高频传递函数为同一语音对应的骨导音频信号与气导音频信号之间对应的传递函数;
基于所述低频传递函数及所述高频传递函数对所述骨导音频信号进行频域扩展,并将扩展后的所述初始音频作为输出音频。
由于可以通过低频传递函数和高频传递函数,扩展骨传导拾震器件采集的骨导音频信号的高频部分,从而使得终端设备可以基于骨传导拾震器件采集到的初始音频得到完整的音频信号。这样达成了提高设备采集的音频信号的完整性的效果。
如图1所示,图1是本发明实施例方案涉及的硬件运行环境的终端结构示意图。
本发明实施例终端可以是骨传导耳机等终端设备。
如图1所示,该终端可以包括:处理器1001,例如CPU,网络接口1004,用户接口1003,存储器1005,通信总线1002。其中,通信总线1002用于实现这些组件之间的连接通信。用户接口1003可以包括骨传导拾震器件等,可选用户接口1003还可以包括标准的有线接口、无线接口。网络接口1004可选的可以包括标准的有线接口、无线接口(如WI-FI接口)。存储器1005可以是高速RAM存储器,也可以是稳定的存储器(non-volatile memory),例如磁盘存储器。存储器1005可选的还可以是独立于前述处理器1001的存储装置。
本领域技术人员可以理解,图1中示出的终端结构并不构成对终端的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置。
如图1所示,作为一种计算机存储介质的存储器1005中可以包括操作系统、网络通信模块、用户接口模块以及音频信号处理程序。
在图1所示的终端中,网络接口1004主要用于连接后台服务器,与后台服务器进行数据通信;处理器1001可以用于调用存储器1005中存储的音频信号处理程序,并执行以下操作:
获取骨传导拾震器件采集的骨导音频信号;
获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数,其中,所述低频传递函数和所述高频传递函数为同一语音对应的骨导音频信号与气导音频信号之间对应的传递函数;
基于所述低频传递函数及所述高频传递函数对所述骨导音频信号进行频域扩展,并将扩展后的所述初始音频作为输出音频。
进一步地,处理器1001可以调用存储器1005中存储的音频信号处理程序,还执行以下操作:
获取所述骨导音频信号的低频特性;
获取与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数。
进一步地,处理器1001可以调用存储器1005中存储的音频信号处理程序,还执行以下操作:
将所述低频特性发送至服务器,其中,所述服务器设置为根据接收到的所述低频特性,获取云端数据库中保存的,与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数,并将所述低频传递函数及所述高频传递函数发送至所述终端设备;
接收所述服务器发送的所述低频传递函数及所述高频传递函数。
进一步地,处理器1001可以调用存储器1005中存储的音频信号处理程序,还执行以下操作:
将所述骨导音频信号发送至服务器,其中,所述服务器设置为根据接收到的所述骨导音频信号,获取云端数据库中保存的与所述初始音频匹配的低频,以及所述低频传递函数对应的所述高频传递函数,并将获取到的所述低频传递函数和所述高频传递函数发送至所述终端设备;
接收所述服务器发送的所述低频传递函数和所述高频传递函数。
进一步地,处理器1001可以调用存储器1005中存储的音频信号处理程序,还执行以下操作:
获取麦克风采集到的气导音频信号;
根据所述骨导音频信号及所述气导音频信号确定所述低频传递函数及所述高频传递函数;
关联保存所述初始音频的低频特性、所述低频传递函数及所述高频传递函数。
进一步地,处理器1001可以调用存储器1005中存储的音频信号处理程序,还执行以下操作:
获取所述骨导音频信号的第一低频特性,以及所述气导音频信号的第二低频特性;
获取所述骨导音频信号的第一高频特性,以及所述气导音频信号的第二高频特性;
根据第一低频特性及所述第二低频特性确定所述低频传递函数,以及根据所述第一高频特性和第二高频特性确定所述高频传递函数。
参照图2,在本发明音频信号的处理方法的一实施例中,所述音频信号的处理方法包括以下步骤:
步骤S10、获取骨传导拾震器件采集的骨导音频信号;
步骤S20、获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数,其中,所述低频传递函数和所述高频传递函数为同一语音对应的骨导音频信号与气导音频信号之间对应的传递函数;
步骤S30、基于所述低频传递函数及所述高频传递函数对所述骨导音频信号进行频域扩展,并将扩展后的所述初始音频作为输出音频。
骨传导是一种声音传导方式,即将声音转化为不同频率的机械振动,通过人的颅骨、骨迷路、内耳淋巴液、螺旋器、听觉中枢来传递声波。相对于通过振膜产生声波的经典声音传导方式,骨传导省去了许多声波传递的步骤,能在嘈杂的环境中实现清晰的声音还原,而且声波也不会因为在空气中扩散而影响到他人。由于骨传导存在上述优势,因而出现了基于骨传导采集用户声音信号的设备。
但是,在现有的骨传导设备中,源于骨传导拾震器件自身的硬件缺陷,会导致通过骨传导拾震器件采集的音频信号出现高频衰减严重的现象。这样导致骨传导音频采集设备采集的音频信号不完整。
为解决现有的骨传导音频采集设备,只能适用于低频环境,而无法采集音源语音的高频部分的缺陷,本发明实施例提出一种音频信号的处理方法。
在本实施例中,上述音频信号的处理方法应用于终端设备,上述终端设备设置有骨传导拾震器件。其中,上述骨传导拾震器件用于采集因音源语音而产生振动的物体的振动波,并将该振动波转换为语音信号。
示例性地,上述终端设备设置为骨传导耳机,当用户佩戴骨传导耳机时,在用户说话 的过程中,会导致肢体同时产生振动,因此,设置于骨传导耳机上的骨传导拾震器件可以采集用户肢体的振动波,并将该振动波转换为音频信号。
当终端设备上设置有骨传导拾震器件时,可以通过骨传导拾震器件采集音频,然后获取骨传导拾震器件采集的初始音频后。
进一步地,当获取到该骨导音频信号后,可以获取与该骨导音频信号匹配的预存的低频传递函数,以及该低频传递函数对应的预存的高频传递函数。可以理解的是,上述高频传递函数和上述低频传递函数为预先保存的数据。其中,所述低频传递函数和所述高频传递函数为同一音源对应骨传导音频信号与气导音频信号之间对应的传递函数。
具体地,当获取到该骨导音频信号后,可以获取该初始音频的低频特性。进而查询数据库中与该低频特性匹配的预存低频
示例性地,当获取到骨传导拾震器件采集的骨导音频信号时,先对采集到的信号进行转换,以将时域信号转换为频域信号。需要说明的是,终端设备可以先获取通过骨传导拾震器件采集的骨导音频信号,其中,终端设备直接获取到的上述骨导音频信号为时域信号。因此,在获取到上述初始音频后,可以基于FFT(fast Fourier transform,快速傅里叶变换),将骨导音频信号从时域信号转换为频域信号。进一步地,可以对转换后的骨导音频信号进行信号分析,从而提取该骨导音频信号的低频特性。其中,上述低频特性可以包括该骨导音频信号的谐波能量变化,以及幅值和频率分布情况等特性。进而,当获取到该骨导音频信号的低频特性后,可以基于该低频特性,在数据库中查找与该低频特性匹配的低频传递函数,例如,低频各频率点不丢失,且每频率△SPL(灵敏度差值)=±0.1dB,则认为该低频特性与预存的低频特性匹配。其中,数据库中可以将预存的低频特性与低频传递函数关联保存。使得当获取到的骨导音频信号的低频特性,与数据库中保存的预存的低频特性匹配时,或该匹配的低频特性关联的低频传递函数。进一步的,数据库中还保存有与该低频传递函数一一对应,且关联保存的高频传递函数,因此,当获取到该低频传递函数后,还可以获取该低频传递函数对应的高频传递函数。
可选地,作为一种实现方式,在获取到上述骨导音频信号的低频特性后,也可以将所述低频特性发送至服务器,其中,所述服务器设置为根据接收到的所述低频特性,获取云端数据库中保存的,与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数,并将所述低频传递函数及所述高频传递函数发送至所述终端设备,然后接收所述服务器发送的所述低频传递函数及所述高频传递函数。由于上述传递函数保存云端服务器,使得数据库管理者可以更加便捷地更新数据库中保存的上述低频传递 函数和高频传递函数,以适应更多的应用场景和不同声音特征的用户。
需要说明的是,上述低频传递传递函数及所述高频传递函数基于训练语音信号得到,其中,所述训练语音信号包括同一语音对应的所述骨传导音频信号和所述气导音频信号。
示例性地,在训练过程中,可以通过骨传导拾震器件以及麦克风同时采集同一音源输出的语音。在本示例中,将骨传导拾震器件采集到的该语音对应的音频信号描述为骨导音频信号。将麦克风等采集到的该音频对应的音频信号描述为气导音频信号(即麦克风采集的通过空气传播的声波对应的音频信号)。然后将该骨导音频信号及该气导音频信号从时域信号转换为频域信号。并对转换为频域信号后的骨导音频信号和气导音频信号进行信号分析,获取骨导音频信号的第一低频特性和第一高频特性,以及所述气导音频信号对应第二低频特性和第二高频特性。进而根据第一低频特性及所述第二低频特性确定所述低频传递函数,以及根据所述第一高频特性和第二高频特性确定所述高频传递函数。
具体地,在本示例中,在一训练过程中,可以同时录制骨导音频信号及气导音频信号,并将骨导音频信号及气导音频信号,通过FFT从时域转换频域,并提取频域信号中的频率、幅值等特性,确定每份语音产生的低频特性(<预设频率值,例如2KHz)和高频特性(>预设频率值)。其中,低频特性一般只是谐波能量的变化,而高频特性就是复杂的衰减、频率丢失等变化。进一步地,x
1(n)表示骨导音频信号中,低频部分用频域表示,x
2(n)表示骨导音频信号中,高频部分用频域表示;y
1(n)表示气导音频信号中,低频部分用频域表示,y
2(n)表示气导音频信号中,高频部分用频域表示。进一步地,可以根据以下公式计算低频部分对应的频域传递函数h
1(n):
h
1(n)=x
1(n)/y
1(n);
根据以下公式计算高频部分对应的频域传递函数h
2(n):
h
2(n)=x
2(n)/y
2(n)
进一步的,在本示例中,可以再通过逆变换得到低频部分对应的时域传递函数h
1(t)和高频部分对应的时域传递函数h
2(t)。其中,低频部分对应的时域传递函数h
1(t)和高频部分对应的时域传递函数h
2(t)只能配对使用,不可单独使用。并且,每一组训练语音就会均可以生成低频部分对应的时域传递函数h
1(t)和高频部分对应的时域传递函数h
2(t)并将其关联保存至数据库中。
进一步地,当获取到骨传导拾震器件当前时刻采集到的骨导音频信号对应的低频传递函数和高频传递函数后,可以基于所述低频传递函数及所述高频传递函数对所述骨导音频信号进行频域扩展。
具体地,当获取到骨导音频信号x(t)后,可以基于以下公式对骨导音频信号x(t)进行频域扩展:
y(t)=x(t)/[h1(t)*h2(t)]
其中,y(t)、h1(t)和h2(t)分别为扩展后的音频信号、低频传递函数和高频传递函数。当对骨导音频信号进行扩展后,可以将扩展后骨导音频信号作为终端设备的输出音频。
示例性地,当终端设备为骨传导耳机时,上述骨传导耳机可以作为移动终端的音频采集设备,进而骨传导耳机可以将扩展后骨导音频信号作为输出音频发送至与自身连接的终端设备。
需要说明的是,为提高通过本实施例记载的音频信号的处理方法得到的输出音频的质量,可以根据实时接收到的初始音频,动态更新该初始音频匹配的低频传递函数和高频传递函数。例如,每秒更新10-50次。
在本实施例公开的技术方案中,先获取骨传导拾震器件采集的骨导音频信号,然后获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数,其中,所述低频传递函数和所述高频传递函数为同一语音对应的骨导音频信号与气导音频信号之间对应的传递函数,并基于所述低频传递函数及所述高频传递函数对所述骨导音频信号进行频域扩展,并将扩展后的所述初始音频作为输出音频,由于可以通过低频传递函数和高频传递函数,扩展骨传导拾震器件采集的骨导音频信号的高频部分,从而使得终端设备可以基于骨传导拾震器件采集到的初始音频得到完整的音频信号。这样达成了提高设备采集的音频信号的完整性的效果。
参照图3,基于上述实施例,在另一实施例中,所述步骤S20包括:
步骤S21、将所述低频特性发送至服务器,其中,所述服务器设置为根据接收到的所述低频特性,获取云端数据库中保存的,与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数,并将所述低频传递函数及所述高频传递函数发送至所述终端设备;
步骤S22、接收所述服务器发送的所述低频传递函数及所述高频传递函数。
在本实施例中,可以将所述骨导音频信号发送至服务器,其中,所述服务器设置为根据接收到的所述骨导音频信号,获取云端数据库中保存的与所述初始音频匹配的低频,以及所述低频传递函数对应的所述高频传递函数,并将获取到的所述低频传递函数和所述高 频传递函数发送至所述终端设备,接收所述服务器发送的所述低频传递函数和所述高频传递函数。
需要说明的是,服务器接收到上述骨导音频信号后,对比上述初始音频的低频特性与预存的低频特性。其中,上述预存的低频特性可以是预存的训练语音信号中的骨传导音频信号的低频特性。进而使得服务器可以获取与上述骨导音频信号匹配的,基于训练语音信号确定的低频传递函数和高频传递函数。并将获取到的低频传递函数和高频传递函数发送至终端设备。
在本实施例公开的技术方案中,由于确定低频传递函数和高频传递函数的过程可以由服务器完成,这样达成了降低终端设备的运行开销的效果。
此外,本发明实施例还提出一种终端设备,所述终端设备包括存储器、处理器及存储在所述存储器上并可在所述处理器上运行的音频信号处理程序,所述音频信号处理程序被所述处理器执行时实现如上各个实施例所述的音频信号的处理方法的步骤。
可选地,所述终端设备为骨传导耳机。
此外,本发明实施例还提出一种计算机可读存储介质,所述计算机可读存储介质上存储有音频信号处理程序,所述音频信号处理程序被处理器执行时实现如上各个实施例所述的音频信号的处理方法的步骤。
需要说明的是,在本文中,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者系统不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者系统所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者系统中还存在另外的相同要素。
上述本发明实施例序号仅仅为了描述,不代表实施例的优劣。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本发明的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在如上所述的一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端设备(可以是骨传导耳 机等)执行本发明各个实施例所述的方法。
以上仅为本发明的优选实施例,并非因此限制本发明的专利范围,凡是利用本发明说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本发明的专利保护范围内。
Claims (10)
- 一种音频信号的处理方法,其特征在于,所述音频信号的处理方法包括以下步骤:获取骨传导拾震器件采集的骨导音频信号;获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数,其中,所述低频传递函数和所述高频传递函数为同一语音对应的骨导音频信号与气导音频信号之间对应的传递函数;基于所述低频传递函数及所述高频传递函数对所述骨导音频信号进行频域扩展,并将扩展后的所述初始音频作为输出音频。
- 如权利要求1所述的音频信号的处理方法,其特征在于,所述获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数的步骤包括:获取所述骨导音频信号的低频特性;获取与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数。
- 如权利要求2所述的音频信号的处理方法,其特征在于,所述音频信号的处理方法应用于终端设备,所述获取与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数的步骤包括:将所述低频特性发送至服务器,其中,所述服务器设置为根据接收到的所述低频特性,获取云端数据库中保存的,与所述低频特性匹配的所述低频传递函数,以及所述低频传递函数对应的所述高频传递函数,并将所述低频传递函数及所述高频传递函数发送至所述终端设备;接收所述服务器发送的所述低频传递函数及所述高频传递函数。
- 如权利要求1所述的音频信号的处理方法,其特征在于,所述音频信号的处理方法应用于终端设备,所述获取所述骨导音频信号匹配的低频传递函数,以及所述低频传递函数对应的高频传递函数的步骤包括:将所述骨导音频信号发送至服务器,其中,所述服务器设置为根据接收到的所述骨导音频信号,获取云端数据库中保存的与所述初始音频匹配的低频,以及所述低频传递函数对应的所述高频传递函数,并将获取到的所述低频传递函数和所述高频传递函数发送至所 述终端设备;接收所述服务器发送的所述低频传递函数和所述高频传递函数。
- 如权利要求1所述的音频信号的处理方法,其特征在于,所述低频传递函数与所述高频传递函数一一对应,关联保存于数据库中。
- 如权利要求1所述的音频信号的处理方法,其特征在于,所述获取骨传导拾震器件采集的骨导音频信号的步骤之后,还包括:获取麦克风采集到的气导音频信号;根据所述骨导音频信号及所述气导音频信号确定所述低频传递函数及所述高频传递函数;关联保存所述初始音频的低频特性、所述低频传递函数及所述高频传递函数。
- 如权利要求6所述的音频信号的处理方法,其特征在于,所述根据所述骨导音频信号及所述气导音频信号确定所述低频传递函数及所述高频传递函数的步骤包括:获取所述骨导音频信号的第一低频特性,以及所述气导音频信号的第二低频特性;获取所述骨导音频信号的第一高频特性,以及所述气导音频信号的第二高频特性;根据第一低频特性及所述第二低频特性确定所述低频传递函数,以及根据所述第一高频特性和第二高频特性确定所述高频传递函数。
- 如权利要求1所述的音频信号的处理方法,其特征在于,所述低频传递传递函数及所述高频传递函数基于训练语音信号得到,其中,所述训练语音信号包括同一语音对应的所述骨传导音频信号和所述气导音频信号。
- 一种终端设备,其特征在于,所述终端设备包括:存储器、处理器及存储在所述存储器上并可在所述处理器上运行的音频信号处理程序,所述音频信号处理程序被所述处理器执行时实现如权利要求1至8中任一项所述的音频信号的处理方法的步骤。
- 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质上存储有音频信号处理程序,所述音频信号处理程序被处理器执行时实现如权利要求1至8中任一项所 述的音频信号的处理方法的步骤。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/044,921 US12335679B2 (en) | 2020-09-10 | 2020-10-31 | Audio signal processing method, terminal device and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202010953528.6A CN112017677B (zh) | 2020-09-10 | 2020-09-10 | 音频信号的处理方法、终端设备及存储介质 |
| CN202010953528.6 | 2020-09-10 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022052256A1 true WO2022052256A1 (zh) | 2022-03-17 |
Family
ID=73522769
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/125633 Ceased WO2022052256A1 (zh) | 2020-09-10 | 2020-10-31 | 音频信号的处理方法、终端设备及存储介质 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US12335679B2 (zh) |
| CN (1) | CN112017677B (zh) |
| WO (1) | WO2022052256A1 (zh) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112017677B (zh) * | 2020-09-10 | 2024-02-09 | 歌尔科技有限公司 | 音频信号的处理方法、终端设备及存储介质 |
| CN113205824B (zh) * | 2021-04-30 | 2022-11-11 | 紫光展锐(重庆)科技有限公司 | 声音信号处理方法、装置、存储介质、芯片及相关设备 |
| CN113314134B (zh) * | 2021-05-11 | 2022-11-11 | 紫光展锐(重庆)科技有限公司 | 一种骨传导信号补偿方法及装置 |
| EP4241459B1 (en) * | 2021-05-14 | 2026-03-18 | Shenzhen Shokz Co., Ltd. | Systems and methods for audio signal generation |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20040264718A1 (en) * | 2003-06-03 | 2004-12-30 | Mitsubishi Denki Kabushiki Kaisha | Acoustic signal processing unit |
| CN105721973A (zh) * | 2016-01-26 | 2016-06-29 | 王泽玲 | 一种骨传导耳机及其音频处理方法 |
| CN108476369A (zh) * | 2015-09-07 | 2018-08-31 | 3D声音实验室 | 用于开发适用于个体的头部相关传递函数的方法和系统 |
| CN108496285A (zh) * | 2015-08-31 | 2018-09-04 | 诺拉控股有限公司 | 听觉刺激的个性化 |
| CN110996215A (zh) * | 2020-02-26 | 2020-04-10 | 恒玄科技(北京)有限公司 | 确定耳机降噪参数的方法、装置以及计算机可读介质 |
| US20200219525A1 (en) * | 2019-01-04 | 2020-07-09 | Samsung Electronics Co., Ltd. | Processing method of audio signal and electronic device supporting the same |
Family Cites Families (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2012147077A (ja) * | 2011-01-07 | 2012-08-02 | Jvc Kenwood Corp | 骨伝導型音声伝達装置 |
| JP6123503B2 (ja) * | 2013-06-07 | 2017-05-10 | 富士通株式会社 | 音声補正装置、音声補正プログラム、および、音声補正方法 |
| WO2016129717A1 (ko) * | 2015-02-11 | 2016-08-18 | 재단법인 다차원 스마트 아이티 융합시스템 연구단 | 골전도를 이용하는 착용형 장치 |
| CN109640212A (zh) * | 2019-02-20 | 2019-04-16 | 广州明医医疗科技有限公司 | 音质改善方法及骨传导耳机 |
| CN111631728B (zh) * | 2020-05-26 | 2023-01-24 | 广州大学 | 一种骨传导传递函数的测量方法、装置及存储介质 |
| CN112017677B (zh) * | 2020-09-10 | 2024-02-09 | 歌尔科技有限公司 | 音频信号的处理方法、终端设备及存储介质 |
| US12080313B2 (en) * | 2022-06-29 | 2024-09-03 | Analog Devices International Unlimited Company | Audio signal processing method and system for enhancing a bone-conducted audio signal using a machine learning model |
-
2020
- 2020-09-10 CN CN202010953528.6A patent/CN112017677B/zh active Active
- 2020-10-31 US US18/044,921 patent/US12335679B2/en active Active
- 2020-10-31 WO PCT/CN2020/125633 patent/WO2022052256A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20040264718A1 (en) * | 2003-06-03 | 2004-12-30 | Mitsubishi Denki Kabushiki Kaisha | Acoustic signal processing unit |
| CN108496285A (zh) * | 2015-08-31 | 2018-09-04 | 诺拉控股有限公司 | 听觉刺激的个性化 |
| CN108476369A (zh) * | 2015-09-07 | 2018-08-31 | 3D声音实验室 | 用于开发适用于个体的头部相关传递函数的方法和系统 |
| CN105721973A (zh) * | 2016-01-26 | 2016-06-29 | 王泽玲 | 一种骨传导耳机及其音频处理方法 |
| US20200219525A1 (en) * | 2019-01-04 | 2020-07-09 | Samsung Electronics Co., Ltd. | Processing method of audio signal and electronic device supporting the same |
| CN110996215A (zh) * | 2020-02-26 | 2020-04-10 | 恒玄科技(北京)有限公司 | 确定耳机降噪参数的方法、装置以及计算机可读介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN112017677A (zh) | 2020-12-01 |
| US12335679B2 (en) | 2025-06-17 |
| US20230276165A1 (en) | 2023-08-31 |
| CN112017677B (zh) | 2024-02-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2022052256A1 (zh) | 音频信号的处理方法、终端设备及存储介质 | |
| US9892721B2 (en) | Information-processing device, information processing method, and program | |
| US10529358B2 (en) | Method and system for reducing background sounds in a noisy environment | |
| CN111065035B (zh) | 一种骨传导耳机测试方法及测试系统 | |
| US12424196B2 (en) | Signal processing apparatus, method, and system | |
| CN114125625B (zh) | 降噪调整方法、耳机及计算机可读存储介质 | |
| CN112017687A (zh) | 一种骨传导设备的语音处理方法、装置及介质 | |
| WO2015058484A1 (zh) | 降噪耳机及其降噪方法 | |
| WO2022041485A1 (zh) | 音频信号的处理方法、电子设备及存储介质 | |
| WO2022174727A1 (zh) | 啸叫抑制方法、装置、助听器及存储介质 | |
| CN109314814A (zh) | 主动降噪方法及耳机 | |
| WO2024002896A1 (en) | Audio signal processing method and system for enhancing a bone-conducted audio signal using a machine learning model | |
| CN109905808B (zh) | 用于调节智能语音设备的方法和装置 | |
| CN112017639A (zh) | 语音信号的检测方法、终端设备及存储介质 | |
| CN108600893A (zh) | 军事环境音频分类系统、方法及军用降噪耳机 | |
| CN108200492A (zh) | 语音控制优化方法、装置以及集成入耳式麦克风的耳机和穿戴设备 | |
| CN111464930B (zh) | 耳机的啸叫检测方法、检测装置及存储介质 | |
| CN115086851B (zh) | 人耳骨传导传递函数测量方法、装置、终端设备以及介质 | |
| CN207884862U (zh) | 基于人耳仿真结构的音频设备 | |
| CN117278922B (zh) | 自调式听觉补偿装置、方法及电脑程序产品 | |
| CN116193321A (zh) | 声音信号处理方法、装置、设备及存储介质 | |
| US20260095705A1 (en) | Method for tuning anc based on hearing threshold level and apparatus for tuning anc based on hearing threshold level | |
| WO2020220271A1 (zh) | 存储器、声学单元及其音频处理方法、装置、设备和系统 | |
| TWI837748B (zh) | 耳機裝置、其補償方法及電腦程式產品 | |
| CN114786083A (zh) | 降噪方法、装置、耳机设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20953041 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20953041 Country of ref document: EP Kind code of ref document: A1 |
|
| WWG | Wipo information: grant in national office |
Ref document number: 18044921 Country of ref document: US |