WO2013102403A1 - Procédé et dispositif de traitement de signal audio, et terminal - Google Patents
Procédé et dispositif de traitement de signal audio, et terminal Download PDFInfo
- Publication number
- WO2013102403A1 WO2013102403A1 PCT/CN2012/086953 CN2012086953W WO2013102403A1 WO 2013102403 A1 WO2013102403 A1 WO 2013102403A1 CN 2012086953 W CN2012086953 W CN 2012086953W WO 2013102403 A1 WO2013102403 A1 WO 2013102403A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- signal
- audio signal
- video signal
- audio
- currently received
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/18—Vocoders using multiple modes
- G10L19/20—Vocoders using multiple modes using sound class specific coding, hybrid encoders or object based coding
Definitions
- Audio signal processing method device and terminal
- the present invention relates to the field of terminals, and in particular, to an audio signal processing method, apparatus, and terminal. Background technique
- Time domain coding is performed on the waveform of the audio signal.
- time domain coding there are typical coding standards such as International Telecommunication Union (ITU, International Telecommunication Union) G.729, G.723.1 and G.728. These coding standards widely use Code Excited Linear Prediction (CELP).
- CELP Code Excited Linear Prediction
- the technology in principle, is modeled according to the human occurrence mechanism, and uses the inherent characteristics of the human glottis and the channel to remove redundant information in the audio signal, thereby greatly reducing the audio quality while greatly reducing the quality.
- CELP Code Excited Linear Prediction
- the principle of frequency domain coding is to encode the audio signal in the frequency domain using the human ear's acceptance principle of sound. Focus on encoding the frequency bands of human concern, and adopt a roughly quantified or unquantized strategy for bands that are masked by other bands or that are not easily perceived by humans.
- frequency domain coding is that certain redundancy is removed according to the characteristics of the human ear, so the coding effect on various audio signals is almost the same, especially for the encoding quality of music and other signals is higher than the time domain coding.
- the human vocalization mechanism is not considered in the coding, and the vocal redundancy cannot be removed, so the coding effect is much lower than the time domain coding based on CELP technology.
- Time domain coding of technology.
- Low bit rate audio coding based on time domain coding can provide higher quality speech coding quality for videophone applications at a lower bit rate, ensuring clearer and easier to understand voice communication capabilities in videophones.
- videophones are often accompanied by other sounds (non-speech) while the voice communication is being performed. For example, the party should let the other party listen to music or other sounds.
- the low bit rate based on time domain coding is adopted. Audio coding results in poor coding quality and severe distortion of the sound.
- Embodiments of the present invention provide an audio signal processing method, apparatus, and terminal, which are used to solve the problem that a single encoding causes a poor sound transmission quality in a sound transmission process.
- a low code rate audio coding method comprising:
- the audio signal Determining, according to the received video signal, the audio signal as a voice signal or a non-speech signal; if the audio signal is determined to be a voice signal, encoding the audio signal by using low-rate audio coding based on time domain coding;
- the audio signal is encoded using low bit rate audio coding based on frequency domain coding.
- a low bit rate audio encoding device comprising:
- a first receiving module configured to receive an audio signal
- a second receiving module configured to receive a video signal
- a determining module configured to determine, according to the received video signal, the audio signal as a voice signal or a non-speech signal
- a first encoding module configured to: when the determining module determines that the audio signal is a voice signal, encoding the audio signal by using low-rate audio coding based on time domain coding;
- a second encoding module configured to encode the audio signal by using low frequency rate audio coding based on frequency domain coding when the determining module determines that the audio signal is a non-speech signal.
- a terminal comprising the low bit rate audio encoding device.
- the type of the received audio signal is determined by the received video signal, and when the received audio signal is determined to be a voice signal, the time domain coding method is used.
- the speech signal and the non-speech signal are separately encoded, and the sound is transmitted.
- FIG. 1 is a flow chart of steps of an audio signal processing method according to Embodiment 1 of the present invention
- FIG. 2 is a schematic diagram of a code stream according to Embodiment 1 of the present invention
- FIG. 3 is a schematic structural diagram of an audio signal processing apparatus according to Embodiment 2 of the present invention
- FIG. 4 is a schematic structural diagram of a terminal according to Embodiment 3 of the present invention.
- image capture in the videophone is used, and based on the information of the image, whether the audio is irregular audio or voice is discriminated, thereby guiding the audio encoding.
- the audio coding quality is improved when the code rate is unchanged.
- Embodiment 1 is a diagrammatic representation of Embodiment 1:
- the first embodiment of the present invention provides an audio signal processing method, which may be, but is not limited to, applied to the field of videophone audio coding.
- the steps of the method are as shown in FIG. 1, and include:
- Step 101 Receive a signal.
- this step includes: receiving a video signal while receiving an audio signal.
- the video signal may be obtained by taking a camera configured in the videophone for shooting in a set area.
- Step 102 Determine a type of the audio signal.
- the audio signal may be determined to be a voice signal or a non-speech signal according to the received video signal.
- this step it may be determined whether a specified image exists in the currently received video signal (current video frame), that is, whether the specified image is included in the currently captured setting area of the camera, and specifically, may be determined according to the pixel information. Whether a specified image exists in the currently received video signal (current video frame), and if a specified image exists in the video signal, determining a received video signal (previous video frame) having the shortest time from the video signal:
- the received audio signal is a voice signal, otherwise, it determines the currently received The audio signal is a non-speech signal.
- the currently received audio signal may refer to an audio signal received between the time when the type of the audio signal is determined this time and the time when the type of the audio signal is determined next time.
- the time to acquire a frame of video frames is very short, such as 20 milliseconds, the processing speed of the video signal is very fast, and an audio signal is used during the call using the videophone.
- the time is generally longer, so a delay in the beginning of the audio signal can be ignored.
- the type of the audio signal received during the time is set to be a voice signal or a non-voice signal during the time when the video signal is first determined by the video signal.
- the specified image may be, but not limited to, a vocal organ such as a lip or a throat.
- the lip area can be based on the human voice. The area occupied by the area surrounded by the lips and the lower lip changes, determining whether the absolute value of the change in the lip area satisfies a set threshold, such as greater than the first threshold, determining that the current audio signal is a human voice signal, Otherwise, it is determined that the current audio signal is not a human voice signal and belongs to a non-speech signal.
- the upper (lower) lip moves up and down according to the characteristics of the upper and lower lips when the human voice is sounded, and whether the absolute value of the displacement of the upper (or lower) lip movement satisfies the set threshold, such as whether it is greater than the second threshold, and When it is determined that the absolute value of the displacement of the upper (or lower) lip movement satisfies the set threshold, it is determined that the current audio signal is a voice signal sent by a human; otherwise, it is determined that the current audio signal is not a voice signal sent by a human, and belongs to a non-speech signal.
- the currently received audio signal is a non-speech signal. If it is determined that there is a specified image in the currently received video signal, and the specified image does not exist in the received video signal, it is determined that the currently received audio signal is a voice signal.
- the type of the currently received audio signal may be determined only according to the currently received video signal, and specifically, may be determined. Whether the specified image exists in the currently received video signal, and if not, determines that the currently received audio signal is a non-speech signal, otherwise, determines that the currently received audio signal is a voice signal.
- the specified image can be identified from the video frame using existing image recognition methods. For example, when recognizing a lip, there may be a large difference between the color of the lips and the skin of the caller and other organs. In the captured video frame, the red component (R component) and the green component (G component) in the lip image pixel. The difference between the two components is significantly different from that of other blocks, and the difference between the R component and the G component is used as a method of recognizing the lip image from the video frame.
- G(x, y) + R(x, y) R(x, y)
- R , ) represent the value of the R component at the pixel
- G(x, y) represents the value of the G component at the pixel. Represents the difference between the red and green components of the pixel.
- the image can be binarized with components.
- the binarization threshold can be based on the optimal threshold for binarization (multiple skin color, gender, age). By sorting the binarized pixel information and removing the scattered noise points, the estimated area of the lips (the area surrounded by the upper and lower lips) can be obtained, and the image of the lips can be recognized.
- the relative displacement of the current video frame and the image specified in the previous video frame can be determined by:
- the binarized dot matrix corresponding to the region is cropped according to the coordinate point of the region, and the binarized dot matrix corresponding to the lip region is represented by P.
- the area of the dot matrix can be expressed.
- the binarized pixel value is h ⁇ x in the previous video frame, and the binarized pixel value in the current video frame is x, y), which can be calculated by the following formula (2)
- D The difference between a video frame and the lip area of the current video frame, denoted by D:
- the current audio signal is a voice signal sent by a human; otherwise, it is determined that the current audio signal is not a voice signal sent by a human, and belongs to a non-speech signal.
- Step 103 Encode the audio signal.
- the audio signal is encoded by using low-rate audio coding based on time domain coding.
- an existing coding method may be used, for example, according to ITU G.729/728/ 723.1, 3GPP AMR-NB/WB or other coding methods based on CELP technology Row coding, otherwise, when determining that the audio signal is a non-speech signal, the audio signal is encoded by low-rate audio coding based on frequency domain coding.
- an existing coding mode such as usage awareness, may be used. The weight of the force, the coding method of lattice vector quantization in the fast Fourier transform (FFT) domain.
- FFT fast Fourier transform
- Step 104 Quantize the output of the encoded data.
- the encoded data can be quantized, the code stream is organized and output. And the identifier bit can be set in the code stream header, and the code stream obtained by using the time domain coding and the code stream obtained by using the frequency domain coding are distinguished for subsequent decoding operations. Specifically, as shown in FIG.
- the code stream with the identification bit is used, the CELP coding is used for the speech signal (the coding method based on the CELP technology), and the transform domain coding is used for the non-speech signal (the coding method based on the frequency domain coding)
- an identifier bit may be set in the code stream header, and the identifier bit is 0, and the code stream is identified as a CELP code stream (voice code stream), and the identifier bit is 1, and the code stream is a transform domain. Coded stream (non-voice stream).
- Embodiment 2 is a diagrammatic representation of Embodiment 1:
- Embodiment 2 of the present invention provides an audio signal processing apparatus, which may be, but is not limited to, applied to the field of videophone audio coding.
- the structure of the apparatus is as shown in FIG. 3, and includes:
- the module 14 is configured to: when the determining module determines that the audio signal is a voice signal, encode the audio signal by using low-rate audio coding based on time domain coding; and the second encoding module 15 is configured to determine the audio in the determining module.
- the audio signal is a non-speech signal
- the audio signal is encoded using low bit rate audio coding based on frequency domain coding.
- the determining module 13 is specifically configured to determine whether a specified image exists in the currently received video signal, and if a specified image exists in the video signal, determine a received video signal that is the shortest time from the video signal: A specified image exists in the received video signal, and when the absolute value of the relative displacement of the image specified in the received video signal and the image specified in the currently received video signal satisfies a set threshold, determining the currently received
- the audio signal is a voice signal, otherwise, indeed The currently received audio signal is a non-speech signal.
- the determining module 13 is further configured to: when determining that the specified image does not exist in the currently received video signal, determine that the currently received audio signal is a non-speech signal; and, determining that a specified one exists in the currently received video signal And an image, and when the specified image does not exist in the received video signal, determining that the currently received audio signal is a voice signal.
- the determining module 13 is specifically configured to determine whether a specified image exists in the currently received video signal, and if not, determine that the currently received audio signal is a non-speech signal, otherwise, determine that the currently received audio signal is a voice signal. .
- the device also includes:
- the code stream output module 16 is configured to quantize the data obtained after the encoding, and organize the code stream output, where the code stream includes an identifier bit for identifying the encoding mode of the data corresponding to the code stream.
- the identifier bit may be set to 0, and the code stream is identified as a code stream obtained by using time domain coding, and the identifier bit is set to 1 , and the code stream is identified as a code stream obtained by using frequency domain coding.
- the third embodiment of the present invention provides a terminal, and the structure of the terminal may be as shown in FIG. 4, where the device provided by the second embodiment of the present invention may be integrated, and the terminal may further include a video signal acquisition module. 21 and audio signal acquisition module 22:
- the video signal acquisition module 21 is configured to provide a video signal to the second receiving module
- the audio signal acquisition module 22 is configured to provide an audio signal to the first receiving module.
- the terminal may further include an audio signal output module 23 for outputting the encoded audio signal.
- the terminal may further include a video signal output module 24 for outputting a video signal. That is, the terminal may transmit only the encoded audio signal, or may transmit the video signal while transmitting the encoded audio signal.
- the device provided in Embodiment 2 of the present invention can be integrated in a videophone, the device can be independent of the camera of the videophone, and the second receiving module of the device can be collected by using a camera (which can be used as a video signal acquisition module).
- the video signal determines the type of audio signal.
- the camera of the videophone can also be integrated in the device as a second receiving module for collecting video signals to determine the type of audio signal.
- the sound can be determined by using a video signal.
- the type of the frequency signal thereby determining the encoding method of the audio signal, improving the audio encoding quality, and reducing the sound distortion.
- the spirit and scope of the Ming thus, it is intended that the present invention cover the modifications and modifications of the invention
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201210001235.3 | 2012-01-04 | ||
| CN201210001235.3A CN103198834B (zh) | 2012-01-04 | 2012-01-04 | 一种音频信号处理方法、装置及终端 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2013102403A1 true WO2013102403A1 (fr) | 2013-07-11 |
Family
ID=48721308
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2012/086953 Ceased WO2013102403A1 (fr) | 2012-01-04 | 2012-12-19 | Procédé et dispositif de traitement de signal audio, et terminal |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN103198834B (fr) |
| WO (1) | WO2013102403A1 (fr) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108831472A (zh) * | 2018-06-27 | 2018-11-16 | 中山大学肿瘤防治中心 | 一种基于唇语识别的人工智能发声系统及发声方法 |
| CN111081264A (zh) * | 2019-12-06 | 2020-04-28 | 北京明略软件系统有限公司 | 一种语音信号处理方法、装置、设备及存储介质 |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105280188B (zh) * | 2014-06-30 | 2019-06-28 | 美的集团股份有限公司 | 基于终端运行环境的音频信号编码方法和系统 |
| CN105979469B (zh) * | 2016-06-29 | 2020-01-31 | 维沃移动通信有限公司 | 一种录音处理方法及终端 |
| CN115334349B (zh) * | 2022-07-15 | 2024-01-02 | 北京达佳互联信息技术有限公司 | 音频处理方法、装置、电子设备及存储介质 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20030009325A1 (en) * | 1998-01-22 | 2003-01-09 | Raif Kirchherr | Method for signal controlled switching between different audio coding schemes |
| US6754373B1 (en) * | 2000-07-14 | 2004-06-22 | International Business Machines Corporation | System and method for microphone activation using visual speech cues |
| US20040267521A1 (en) * | 2003-06-25 | 2004-12-30 | Ross Cutler | System and method for audio/video speaker detection |
| US20070174051A1 (en) * | 2006-01-24 | 2007-07-26 | Samsung Electronics Co., Ltd. | Adaptive time and/or frequency-based encoding mode determination apparatus and method of determining encoding mode of the apparatus |
| CN101615393A (zh) * | 2008-06-25 | 2009-12-30 | 汤姆森许可贸易公司 | 对语音和/或非语音音频输入信号编码或解码的方法和设备 |
| CN101656070A (zh) * | 2008-08-22 | 2010-02-24 | 展讯通信(上海)有限公司 | 一种语音检测方法 |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7860718B2 (en) * | 2005-12-08 | 2010-12-28 | Electronics And Telecommunications Research Institute | Apparatus and method for speech segment detection and system for speech recognition |
-
2012
- 2012-01-04 CN CN201210001235.3A patent/CN103198834B/zh active Active
- 2012-12-19 WO PCT/CN2012/086953 patent/WO2013102403A1/fr not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20030009325A1 (en) * | 1998-01-22 | 2003-01-09 | Raif Kirchherr | Method for signal controlled switching between different audio coding schemes |
| US6754373B1 (en) * | 2000-07-14 | 2004-06-22 | International Business Machines Corporation | System and method for microphone activation using visual speech cues |
| US20040267521A1 (en) * | 2003-06-25 | 2004-12-30 | Ross Cutler | System and method for audio/video speaker detection |
| US20070174051A1 (en) * | 2006-01-24 | 2007-07-26 | Samsung Electronics Co., Ltd. | Adaptive time and/or frequency-based encoding mode determination apparatus and method of determining encoding mode of the apparatus |
| CN101615393A (zh) * | 2008-06-25 | 2009-12-30 | 汤姆森许可贸易公司 | 对语音和/或非语音音频输入信号编码或解码的方法和设备 |
| CN101656070A (zh) * | 2008-08-22 | 2010-02-24 | 展讯通信(上海)有限公司 | 一种语音检测方法 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108831472A (zh) * | 2018-06-27 | 2018-11-16 | 中山大学肿瘤防治中心 | 一种基于唇语识别的人工智能发声系统及发声方法 |
| CN108831472B (zh) * | 2018-06-27 | 2022-03-11 | 中山大学肿瘤防治中心 | 一种基于唇语识别的人工智能发声系统及发声方法 |
| CN111081264A (zh) * | 2019-12-06 | 2020-04-28 | 北京明略软件系统有限公司 | 一种语音信号处理方法、装置、设备及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN103198834B (zh) | 2016-12-14 |
| CN103198834A (zh) | 2013-07-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US9218820B2 (en) | Audio fingerprint differences for end-to-end quality of experience measurement | |
| US8428959B2 (en) | Audio packet loss concealment by transform interpolation | |
| US20210089863A1 (en) | Method and apparatus for recurrent auto-encoding | |
| CN101165778B (zh) | 音频信号的双变换编码方法和装置 | |
| CN108197572B (zh) | 一种唇语识别方法和移动终端 | |
| US8386266B2 (en) | Full-band scalable audio codec | |
| CN110838894B (zh) | 语音处理方法、装置、计算机可读存储介质和计算机设备 | |
| US8831932B2 (en) | Scalable audio in a multi-point environment | |
| TW201143445A (en) | Controlling video encoding using audio information | |
| CN101165777A (zh) | 快速点阵向量量化 | |
| CN105551512A (zh) | 音频格式转换方法和装置 | |
| WO2013102403A1 (fr) | Procédé et dispositif de traitement de signal audio, et terminal | |
| KR20180040716A (ko) | 음질 향상을 위한 신호 처리방법 및 장치 | |
| WO2021244418A1 (fr) | Procédé de codage audio et appareil de codage audio | |
| WO2023197809A1 (fr) | Procédé de codage et de décodage de signal audio haute fréquence et appareils associés | |
| WO2021244417A1 (fr) | Procédé de codage audio et dispositif de codage audio | |
| CN112767955A (zh) | 音频编码方法及装置、存储介质、电子设备 | |
| CN115831132A (zh) | 音频编解码方法、装置、介质及电子设备 | |
| CN116137151B (zh) | 低码率网络连接中提供高质量音频通信的系统和方法 | |
| CN101547010B (zh) | 编码解码系统、方法及装置 | |
| JP2009522914A (ja) | 通信網を介して加入者端末機に送信されるオーディオ信号の出力品質改善のためのオーディオ信号の処理方法およびこの方法を採用したオーディオ信号処理装置 | |
| WO2014000559A1 (fr) | Procédé de traitement de signaux vocaux ou audio et appareil de codage associé | |
| JP4437011B2 (ja) | 音声符号化装置 | |
| CN101990082B (zh) | 一种实现可视电话的方法及装置 | |
| HK40043832B (zh) | 音频编码方法及装置、存储介质、电子设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 12864111 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 12864111 Country of ref document: EP Kind code of ref document: A1 |