WO2013102403A1 - 一种音频信号处理方法、装置及终端 - Google Patents

一种音频信号处理方法、装置及终端 Download PDF

Info

Publication number
WO2013102403A1
WO2013102403A1 PCT/CN2012/086953 CN2012086953W WO2013102403A1 WO 2013102403 A1 WO2013102403 A1 WO 2013102403A1 CN 2012086953 W CN2012086953 W CN 2012086953W WO 2013102403 A1 WO2013102403 A1 WO 2013102403A1
Authority
WO
WIPO (PCT)
Prior art keywords
signal
audio signal
video signal
audio
currently received
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2012/086953
Other languages
English (en)
French (fr)
Inventor
刘霖
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
China Mobile Communications Group Co Ltd
Original Assignee
China Mobile Communications Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by China Mobile Communications Corp filed Critical China Mobile Communications Corp
Publication of WO2013102403A1 publication Critical patent/WO2013102403A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • G10L19/18Vocoders using multiple modes
    • G10L19/20Vocoders using multiple modes using sound class specific coding, hybrid encoders or object based coding

Definitions

  • Audio signal processing method device and terminal
  • the present invention relates to the field of terminals, and in particular, to an audio signal processing method, apparatus, and terminal. Background technique
  • Time domain coding is performed on the waveform of the audio signal.
  • time domain coding there are typical coding standards such as International Telecommunication Union (ITU, International Telecommunication Union) G.729, G.723.1 and G.728. These coding standards widely use Code Excited Linear Prediction (CELP).
  • CELP Code Excited Linear Prediction
  • the technology in principle, is modeled according to the human occurrence mechanism, and uses the inherent characteristics of the human glottis and the channel to remove redundant information in the audio signal, thereby greatly reducing the audio quality while greatly reducing the quality.
  • CELP Code Excited Linear Prediction
  • the principle of frequency domain coding is to encode the audio signal in the frequency domain using the human ear's acceptance principle of sound. Focus on encoding the frequency bands of human concern, and adopt a roughly quantified or unquantized strategy for bands that are masked by other bands or that are not easily perceived by humans.
  • frequency domain coding is that certain redundancy is removed according to the characteristics of the human ear, so the coding effect on various audio signals is almost the same, especially for the encoding quality of music and other signals is higher than the time domain coding.
  • the human vocalization mechanism is not considered in the coding, and the vocal redundancy cannot be removed, so the coding effect is much lower than the time domain coding based on CELP technology.
  • Time domain coding of technology.
  • Low bit rate audio coding based on time domain coding can provide higher quality speech coding quality for videophone applications at a lower bit rate, ensuring clearer and easier to understand voice communication capabilities in videophones.
  • videophones are often accompanied by other sounds (non-speech) while the voice communication is being performed. For example, the party should let the other party listen to music or other sounds.
  • the low bit rate based on time domain coding is adopted. Audio coding results in poor coding quality and severe distortion of the sound.
  • Embodiments of the present invention provide an audio signal processing method, apparatus, and terminal, which are used to solve the problem that a single encoding causes a poor sound transmission quality in a sound transmission process.
  • a low code rate audio coding method comprising:
  • the audio signal Determining, according to the received video signal, the audio signal as a voice signal or a non-speech signal; if the audio signal is determined to be a voice signal, encoding the audio signal by using low-rate audio coding based on time domain coding;
  • the audio signal is encoded using low bit rate audio coding based on frequency domain coding.
  • a low bit rate audio encoding device comprising:
  • a first receiving module configured to receive an audio signal
  • a second receiving module configured to receive a video signal
  • a determining module configured to determine, according to the received video signal, the audio signal as a voice signal or a non-speech signal
  • a first encoding module configured to: when the determining module determines that the audio signal is a voice signal, encoding the audio signal by using low-rate audio coding based on time domain coding;
  • a second encoding module configured to encode the audio signal by using low frequency rate audio coding based on frequency domain coding when the determining module determines that the audio signal is a non-speech signal.
  • a terminal comprising the low bit rate audio encoding device.
  • the type of the received audio signal is determined by the received video signal, and when the received audio signal is determined to be a voice signal, the time domain coding method is used.
  • the speech signal and the non-speech signal are separately encoded, and the sound is transmitted.
  • FIG. 1 is a flow chart of steps of an audio signal processing method according to Embodiment 1 of the present invention
  • FIG. 2 is a schematic diagram of a code stream according to Embodiment 1 of the present invention
  • FIG. 3 is a schematic structural diagram of an audio signal processing apparatus according to Embodiment 2 of the present invention
  • FIG. 4 is a schematic structural diagram of a terminal according to Embodiment 3 of the present invention.
  • image capture in the videophone is used, and based on the information of the image, whether the audio is irregular audio or voice is discriminated, thereby guiding the audio encoding.
  • the audio coding quality is improved when the code rate is unchanged.
  • Embodiment 1 is a diagrammatic representation of Embodiment 1:
  • the first embodiment of the present invention provides an audio signal processing method, which may be, but is not limited to, applied to the field of videophone audio coding.
  • the steps of the method are as shown in FIG. 1, and include:
  • Step 101 Receive a signal.
  • this step includes: receiving a video signal while receiving an audio signal.
  • the video signal may be obtained by taking a camera configured in the videophone for shooting in a set area.
  • Step 102 Determine a type of the audio signal.
  • the audio signal may be determined to be a voice signal or a non-speech signal according to the received video signal.
  • this step it may be determined whether a specified image exists in the currently received video signal (current video frame), that is, whether the specified image is included in the currently captured setting area of the camera, and specifically, may be determined according to the pixel information. Whether a specified image exists in the currently received video signal (current video frame), and if a specified image exists in the video signal, determining a received video signal (previous video frame) having the shortest time from the video signal:
  • the received audio signal is a voice signal, otherwise, it determines the currently received The audio signal is a non-speech signal.
  • the currently received audio signal may refer to an audio signal received between the time when the type of the audio signal is determined this time and the time when the type of the audio signal is determined next time.
  • the time to acquire a frame of video frames is very short, such as 20 milliseconds, the processing speed of the video signal is very fast, and an audio signal is used during the call using the videophone.
  • the time is generally longer, so a delay in the beginning of the audio signal can be ignored.
  • the type of the audio signal received during the time is set to be a voice signal or a non-voice signal during the time when the video signal is first determined by the video signal.
  • the specified image may be, but not limited to, a vocal organ such as a lip or a throat.
  • the lip area can be based on the human voice. The area occupied by the area surrounded by the lips and the lower lip changes, determining whether the absolute value of the change in the lip area satisfies a set threshold, such as greater than the first threshold, determining that the current audio signal is a human voice signal, Otherwise, it is determined that the current audio signal is not a human voice signal and belongs to a non-speech signal.
  • the upper (lower) lip moves up and down according to the characteristics of the upper and lower lips when the human voice is sounded, and whether the absolute value of the displacement of the upper (or lower) lip movement satisfies the set threshold, such as whether it is greater than the second threshold, and When it is determined that the absolute value of the displacement of the upper (or lower) lip movement satisfies the set threshold, it is determined that the current audio signal is a voice signal sent by a human; otherwise, it is determined that the current audio signal is not a voice signal sent by a human, and belongs to a non-speech signal.
  • the currently received audio signal is a non-speech signal. If it is determined that there is a specified image in the currently received video signal, and the specified image does not exist in the received video signal, it is determined that the currently received audio signal is a voice signal.
  • the type of the currently received audio signal may be determined only according to the currently received video signal, and specifically, may be determined. Whether the specified image exists in the currently received video signal, and if not, determines that the currently received audio signal is a non-speech signal, otherwise, determines that the currently received audio signal is a voice signal.
  • the specified image can be identified from the video frame using existing image recognition methods. For example, when recognizing a lip, there may be a large difference between the color of the lips and the skin of the caller and other organs. In the captured video frame, the red component (R component) and the green component (G component) in the lip image pixel. The difference between the two components is significantly different from that of other blocks, and the difference between the R component and the G component is used as a method of recognizing the lip image from the video frame.
  • G(x, y) + R(x, y) R(x, y)
  • R , ) represent the value of the R component at the pixel
  • G(x, y) represents the value of the G component at the pixel. Represents the difference between the red and green components of the pixel.
  • the image can be binarized with components.
  • the binarization threshold can be based on the optimal threshold for binarization (multiple skin color, gender, age). By sorting the binarized pixel information and removing the scattered noise points, the estimated area of the lips (the area surrounded by the upper and lower lips) can be obtained, and the image of the lips can be recognized.
  • the relative displacement of the current video frame and the image specified in the previous video frame can be determined by:
  • the binarized dot matrix corresponding to the region is cropped according to the coordinate point of the region, and the binarized dot matrix corresponding to the lip region is represented by P.
  • the area of the dot matrix can be expressed.
  • the binarized pixel value is h ⁇ x in the previous video frame, and the binarized pixel value in the current video frame is x, y), which can be calculated by the following formula (2)
  • D The difference between a video frame and the lip area of the current video frame, denoted by D:
  • the current audio signal is a voice signal sent by a human; otherwise, it is determined that the current audio signal is not a voice signal sent by a human, and belongs to a non-speech signal.
  • Step 103 Encode the audio signal.
  • the audio signal is encoded by using low-rate audio coding based on time domain coding.
  • an existing coding method may be used, for example, according to ITU G.729/728/ 723.1, 3GPP AMR-NB/WB or other coding methods based on CELP technology Row coding, otherwise, when determining that the audio signal is a non-speech signal, the audio signal is encoded by low-rate audio coding based on frequency domain coding.
  • an existing coding mode such as usage awareness, may be used. The weight of the force, the coding method of lattice vector quantization in the fast Fourier transform (FFT) domain.
  • FFT fast Fourier transform
  • Step 104 Quantize the output of the encoded data.
  • the encoded data can be quantized, the code stream is organized and output. And the identifier bit can be set in the code stream header, and the code stream obtained by using the time domain coding and the code stream obtained by using the frequency domain coding are distinguished for subsequent decoding operations. Specifically, as shown in FIG.
  • the code stream with the identification bit is used, the CELP coding is used for the speech signal (the coding method based on the CELP technology), and the transform domain coding is used for the non-speech signal (the coding method based on the frequency domain coding)
  • an identifier bit may be set in the code stream header, and the identifier bit is 0, and the code stream is identified as a CELP code stream (voice code stream), and the identifier bit is 1, and the code stream is a transform domain. Coded stream (non-voice stream).
  • Embodiment 2 is a diagrammatic representation of Embodiment 1:
  • Embodiment 2 of the present invention provides an audio signal processing apparatus, which may be, but is not limited to, applied to the field of videophone audio coding.
  • the structure of the apparatus is as shown in FIG. 3, and includes:
  • the module 14 is configured to: when the determining module determines that the audio signal is a voice signal, encode the audio signal by using low-rate audio coding based on time domain coding; and the second encoding module 15 is configured to determine the audio in the determining module.
  • the audio signal is a non-speech signal
  • the audio signal is encoded using low bit rate audio coding based on frequency domain coding.
  • the determining module 13 is specifically configured to determine whether a specified image exists in the currently received video signal, and if a specified image exists in the video signal, determine a received video signal that is the shortest time from the video signal: A specified image exists in the received video signal, and when the absolute value of the relative displacement of the image specified in the received video signal and the image specified in the currently received video signal satisfies a set threshold, determining the currently received
  • the audio signal is a voice signal, otherwise, indeed The currently received audio signal is a non-speech signal.
  • the determining module 13 is further configured to: when determining that the specified image does not exist in the currently received video signal, determine that the currently received audio signal is a non-speech signal; and, determining that a specified one exists in the currently received video signal And an image, and when the specified image does not exist in the received video signal, determining that the currently received audio signal is a voice signal.
  • the determining module 13 is specifically configured to determine whether a specified image exists in the currently received video signal, and if not, determine that the currently received audio signal is a non-speech signal, otherwise, determine that the currently received audio signal is a voice signal. .
  • the device also includes:
  • the code stream output module 16 is configured to quantize the data obtained after the encoding, and organize the code stream output, where the code stream includes an identifier bit for identifying the encoding mode of the data corresponding to the code stream.
  • the identifier bit may be set to 0, and the code stream is identified as a code stream obtained by using time domain coding, and the identifier bit is set to 1 , and the code stream is identified as a code stream obtained by using frequency domain coding.
  • the third embodiment of the present invention provides a terminal, and the structure of the terminal may be as shown in FIG. 4, where the device provided by the second embodiment of the present invention may be integrated, and the terminal may further include a video signal acquisition module. 21 and audio signal acquisition module 22:
  • the video signal acquisition module 21 is configured to provide a video signal to the second receiving module
  • the audio signal acquisition module 22 is configured to provide an audio signal to the first receiving module.
  • the terminal may further include an audio signal output module 23 for outputting the encoded audio signal.
  • the terminal may further include a video signal output module 24 for outputting a video signal. That is, the terminal may transmit only the encoded audio signal, or may transmit the video signal while transmitting the encoded audio signal.
  • the device provided in Embodiment 2 of the present invention can be integrated in a videophone, the device can be independent of the camera of the videophone, and the second receiving module of the device can be collected by using a camera (which can be used as a video signal acquisition module).
  • the video signal determines the type of audio signal.
  • the camera of the videophone can also be integrated in the device as a second receiving module for collecting video signals to determine the type of audio signal.
  • the sound can be determined by using a video signal.
  • the type of the frequency signal thereby determining the encoding method of the audio signal, improving the audio encoding quality, and reducing the sound distortion.
  • the spirit and scope of the Ming thus, it is intended that the present invention cover the modifications and modifications of the invention

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)

Description

一种音频信号处理方法、 装置及终端
本申请要求于 2012 年 1 月 4 日提交中国专利局、 申请号为 201210001235.3、 发明名称为 "一种音频信号处理方法、 装置及终端" 的中国 专利申请的优先权, 其全部内容通过引用结合在本申请中。
技术领域
本发明涉及终端领域, 尤其涉及一种音频信号处理方法、 装置及终端。 背景技术
随着第三代移动通信技术(3G, 3rd-generation )的快速发展, 可视电话逐 步在 3G网络中得到了较多的应用。 在目前的可视电话技术中, 低码率音频编 码技术是可视电话应用中的一个关键技术。
在低码率音频编码领域, 存在 2个主要的技术路线, 一种是时域编码, 一 种是频域编码。
时域编码是针对音频信号的波形, 进行编码。针对时域编码比较典型的有 国际电信联盟 ( ITU, International Telecommunication Union ) G.729、 G.723.1 和 G.728等编码标准,这些编码标准广泛采用了码激励线性预测( CELP, Code Excited Linear Prediction )技术, 从原理上根据人类的发生机理建模, 利用人 类声门、 声道固有的特性, 去除音频信号里面的冗余信息, 从而在保持较高的 音频质量的同时, 大幅度的降低了音频编码所需的比特率。
在这类音频编码方法中, 最致命的缺陷在于该方法主要适用于人类发声 (语音信号), 对于杂乱无章(包括音乐、噪声以及其他声音)的音频信号(非 语音信号), 编码效果较差。
频域编码的原理在于, 利用人耳对于声音的接受原理, 在频域对于音频信 号进行编码。 重点编码人类关注的频段, 而对于被其他频段掩蔽或是人类不易 感知的频段, 采用粗略量化或是不量化的策略。
频域编码的优势在于根据人耳的特性,去除了一定的冗余, 因此对各种音 频信号的编码效果几乎相当, 尤其对于音乐等信号的编码质量要高于时域编 码。但是在语音信号上,其编码时并未考虑人类发声机理,无法去除发声冗余, 因此编码效果要远低于基于 CELP技术的时域编码。
现有的可视电话技术中,由于语音信息相对重要,因此通常采用基于 CELP 技术的时域编码。基于时域编码的低码率音频编码可以在 ^艮低的码率上为可视 电话应用提供较高质量的语音编码质量,确保可视电话中较为清晰、 易懂的语 音通信能力。 但是, 可视电话在进行语音通信的同时, 经常会伴随其他的声音 (非语音), 比如通话方要让对方听音乐或是其他声音的情况, 此时, 采用基 于时域编码的低码率音频编码导致编码质量^艮差, 声音失真严重。
发明内容
本发明实施例提供一种音频信号处理方法、装置及终端, 用于解决声音传 输过程中采用单一编码导致声音传输质量较差的问题。
一种低码率音频编码方法, 所述方法包括:
在接收音频信号的同时, 接收视频信号;
根据接收到的视频信号, 确定所述音频信号为语音信号或非语音信号; 若确定所述音频信号为语音信号,则利用基于时域编码的低码率音频编码 对所述音频信号进行编码;
若确定所述音频信号为非语音信号,则利用基于频域编码的低码率音频编 码对所述音频信号进行编码。
一种低码率音频编码装置, 所述装置包括:
第一接收模块, 用于接收音频信号;
第二接收模块, 用于接收视频信号;
确定模块, 用于根据接收到的视频信号,确定所述音频信号为语音信号或 非语音信号;
第一编码模块, 用于在确定模块确定所述音频信号为语音信号时, 利用基 于时域编码的低码率音频编码对所述音频信号进行编码;
第二编码模块, 用于在确定模块确定所述音频信号为非语音信号时, 利用 基于频域编码的低码率音频编码对所述音频信号进行编码。
一种终端, 所述终端包括上述低码率音频编码装置。
根据本发明实施例提供的方案,在对音频信号进行编码时,通过接收到的 视频信号确定接收到的音频信号的种类,在确定接收到的音频信号为语音信号 时, 利用时域编码的方式对该音频信号进行编码,在确定接收到的音频信号为 非语音信号时, 利用频域编码的方式对该音频信号进行编码,从而对识别出的 语音信号和非语音信号分别进行编码, 并实现声音的传输。
附图说明
图 1为本发明实施例一提供的音频信号处理方法的步骤流程图; 图 2为本发明实施例一提供的码流示意图;
图 3为本发明实施例二提供的音频信号处理装置的结构示意图; 图 4为本发明实施例三提供的终端的结构示意图。
具体实施方式
本发明实施例中, 在可视电话环境下, 利用可视电话中的图像捕捉, 根据 图像的信息, 判别音频是无规律音频还是语音, 从而指导音频编码。 实现在编 码码率不变的情况下, 提高音频编码质量。
下面结合说明书附图和各实施例对本发明方案进行说明。
实施例一:
本发明实施例一提供一种音频信号处理方法,该方法可以但不限于应用于 可视电话音频编码领域, 该方法的步骤如图 1所示, 包括:
步骤 101、 接收信号。
在本步骤中, 不仅需要接收音频信号, 还需要接收音频信号。 因此, 本步 骤包括: 在接收音频信号的同时, 接收视频信号。 所述视频信号可以是可视电 话中配置的摄像头针对设定区域进行拍摄获得的。
步骤 102、 确定音频信号的种类。
在本步骤中, 可以根据接收到的视频信号,确定所述音频信号为语音信号 或非语音信号。
在本步骤中, 可以确定当前接收到的视频信号(当前视频帧)中是否存在 指定的图像, 即确定摄像头当前拍摄的设定区域中是否包含指定的图像, 具体 的, 可以根据像素信息, 确定当前接收到的视频信号(当前视频帧)中是否存 在指定的图像, 若该视频信号中存在指定的图像,确定距离该视频信号时间最 短的一个已接收的视频信号 (上一视频帧):
若该已接收的视频信号中存在指定的图像,在该已接收的视频信号中指定 的图像与当前接收到的视频信号中指定的图像的相对位移的绝对值满足设定 的阈值时, 确定当前接收到的音频信号为语音信号, 否则, 确定当前接收到的 音频信号为非语音信号。
所述当前接收到的音频信号可以是指在本次确定出音频信号种类的时刻 到下次确定出音频信号种类的时刻之间接收到的音频信号。此时, 由于在目前 技术和设备硬件能力下, 采集一帧视频帧的时间非常短, 如 20毫秒, 对视频 信号的处理速度非常快,且在利用可视电话进行通话过程中,一段音频信号的 时间一般较长, 因此可以对音频信号开始的一段延迟忽略不计。 当然, 也可以 在利用可视电话进行的一次通话过程中,在利用视频信号初次确定音频信号种 类的时间内, 设定该时间内接收到的音频信号的种类为语音信号或非语音信 号。
为了利用视频信号确定音频信号的种类,所述指定的图像可以但不限于是 嘴唇、喉咙等发声器官。 并可以在当前视频帧与上一视频帧中指定的图像的相 对位移的绝对值满足设定的阈值时, 具体的, 所述指定的图像为嘴唇时, 可以 根据人类发声时, 嘴唇面积(上嘴唇和下嘴唇围成的区域所占的面积 )会发生 变化的特点, 判断嘴唇面积变化的绝对值是否满足设定的阈值,如大于第一阈 值, 确定当前音频信号是人类发出的语音信号, 否则, 确定当前音频信号不是 人类发出的语音信号,属于非语音信号。当然,也可以根据人类发声时,上(下) 嘴唇会发生上下移动的特点, 判断上(或下 )嘴唇移动的位移的绝对值是否满 足设定的阈值, 如是否大于第二阈值, 并在判断上(或下)嘴唇移动的位移的 绝对值满足设定的阈值时, 确定当前音频信号是人类发出的语音信号, 否则, 确定当前音频信号不是人类发出的语音信号, 属于非语音信号。
进一步的, 若确定当前接收到的视频信号中不存在指定的图像, 可以确定 当前接收到的音频信号为非语音信号。若确定当前接收到的视频信号中存在指 定的图像,且所述已接收的视频信号中不存在指定的图像,确定当前接收到的 音频信号为语音信号。
当然,除了可以结合上一视频帧和当前视频帧来确定当前接收到的音频信 号的种类,也可以仅根据当前接收到的视频信号来确定当前接收到的音频信号 的种类, 具体的, 可以确定当前接收到的视频信号中是否存在指定的图像, 若 不存在, 确定当前接收到的音频信号为非语音信号, 否则, 确定当前接收到的 音频信号为语音信号。 可以采用现有的图像识别方法从视频帧中识别指定的图像。例如,在识别 嘴唇时, 可以根据嘴唇在色彩上与通话者皮肤及其他器官存在较大差异,在采 集到的视频帧中,嘴唇图像像素中的红色分量( R分量)与绿色分量( G分量) 的差异与其他区块有明显的不同的特点, 利用 R分量与 G分量的差异作为从 视频帧中识别嘴唇图像的方法。
具体的, 可以通过如下公式(1 ) 实现嘴唇图像的识别: h(x, y) = ^ ^ ^ ( 1 )
G(x, y) + R(x, y) 其中, R , ) 表示在像素点 上的 R分量值, G(x, y) 表示在像素点 上的 G分量值。 表示像素点( 上的红、 绿分量的差异。
可以利用 分量对图像进行二值化, 二值化的门限值可以根据多人训 练得到 (可以以不同肤色, 不同性别, 不同年龄的人)二值化的最佳门限值。 对二值化后的像素信息进行整理,去除零散的噪声点即可以得到嘴唇的估计区 域(上嘴唇和下嘴唇围成的区域), 实现对嘴唇图像的识别。
且进一步的,可以通过以下方法确定当前视频帧与上一视频帧中指定的图 像的相对位移:
若在当前视频帧搜索到嘴唇区域(嘴唇图像)后, 根据该区域的坐标点, 裁切出该区域对应的二值化点阵,设嘴唇区域对应的二值化点阵用 P表示,该 点阵的面积可以用 表示。 对于点阵 P 中任意一个像素点 在上一视 频帧二值化像素值为 h \x, , 在当前视频帧的二值化像素值为 x, y) , 可以通 过如下公式(2 )计算上一视频帧和当前视频帧中嘴唇区域的差别, 用 D表示:
D =
Figure imgf000007_0001
并可以在确定 D满足设定的阈值时, 确定当前音频信号是人类发出的语 音信号,否则,确定当前音频信号不是人类发出的语音信号,属于非语音信号。
步骤 103、 对音频信号进行编码。
在确定所述音频信号为语音信号时,利用基于时域编码的低码率音频编码 对所述音频信号进行编码, 具体的, 可以采用现有的编码方式, 如根据 ITU G.729/728/723.1 , 3GPP AMR-NB/WB或是其他基于 CELP技术的编码方式进 行编码, 否则, 在确定所述音频信号为非语音信号时, 利用基于频域编码的低 码率音频编码对所述音频信号进行编码, 具体的, 可以采用现有的编码方式, 如使用感知力口权, 在快速傅里叶变换(FFT, Fast Fourier Transform )域进行格 型矢量量化的编码方式。
步骤 104、 对编码后的数据量化输出。
在对音频信号进行编码后, 可以对编码后获得的数据进行量化,组织码流 并输出。且可以在码流头设置标识位, 对采用时域编码获得的码流和对采用频 域编码获得的码流进行区分, 用于后续的解码操作。 具体的, 如图 2所示为带 有标识位的码流,在对语音信号采用 CELP编码(基于 CELP技术的编码方式), 对非语音信号采用变换域编码(基于频域编码的编码方式)时,在编码完成后, 可以在码流头设置一个标识位, 该标识位为 0, 标识该码流是 CELP码流(语 音码流), 该标识位为 1 , 标识该码流是变换域编码码流(非语音码流)。
在解码端, 可以根据标识位, 选择使用变换域解码器还是 CELP解码器, 从而得到正确的解码码流。
与本发明实施例一基于同一发明构思, 提供以下的装置和终端。
实施例二:
本发明实施例二提供一种音频信号处理装置,该装置可以但不限于应用于 可视电话音频编码领域, 该装置的结构如图 3所示, 包括:
第一接收模块 11用于接收音频信号;第二接收模块 12用于接收视频信号; 确定模块 13用于根据接收到的视频信号, 确定所述音频信号为语音信号或非 语音信号;第一编码模块 14用于在确定模块确定所述音频信号为语音信号时, 利用基于时域编码的低码率音频编码对所述音频信号进行编码;第二编码模块 15 用于在确定模块确定所述音频信号为非语音信号时, 利用基于频域编码的 低码率音频编码对所述音频信号进行编码。
所述确定模块 13具体用于确定当前接收到的视频信号中是否存在指定的 图像, 若该视频信号中存在指定的图像,确定距离该视频信号时间最短的一个 已接收的视频信号: 若该已接收的视频信号中存在指定的图像,在该已接收的 视频信号中指定的图像与当前接收到的视频信号中指定的图像的相对位移的 绝对值满足设定的阈值时, 确定当前接收到的音频信号为语音信号, 否则, 确 定当前接收到的音频信号为非语音信号。
所述确定模块 13还用于在确定当前接收到的视频信号中不存在指定的图 像时, 确定当前接收到的音频信号为非语音信号; 以及, 在确定当前接收到的 视频信号中存在指定的图像, 且所述已接收的视频信号中不存在指定的图像 时, 确定当前接收到的音频信号为语音信号。
所述确定模块 13具体用于确定当前接收到的视频信号中是否存在指定的 图像, 若不存在, 确定当前接收到的音频信号为非语音信号, 否则, 确定当前 接收到的音频信号为语音信号。
所述装置还包括:
码流输出模块 16用于对编码后获得的数据进行量化, 并组织码流输出, 所述码流中包括标识位, 用于标识该码流对应的数据的编码方式。 如, 可以将 标识位设置为 0, 标识该码流为采用时域编码获得的码流,将标识位设置为 1 , 标识该码流为采用频域编码获得的码流。
实施例三、
本发明实施例三提供一种终端, 该终端的结构可以如图 4所示, 该终端中 可以集成有本发明实施例二提供的装置,且所述终端中还可以包括进一步包括 视频信号采集模块 21和音频信号采集模块 22:
视频信号采集模块 21用于向所述第二接收模块提供视频信号;
音频信号采集模块 22用于向所述第一接收模块提供音频信号。
所述终端还可以包括音频信号输出模块 23用于输出编码后的音频信号。 当然, 所述终端还可以进一步包括视频信号输出模块 24用于输出视频信号。 即所述终端可以仅传输编码后的音频信号,也可以在传输编码后的音频信号的 同时, 传输视频信号。
具体的, 本发明实施例二提供的装置可以集成在可视电话中, 该装置可以 独立于可视电话的摄像头,且该装置的第二接收模块可以利用摄像头(可以作 为视频信号采集模块)采集的视频信号来确定音频信号的种类。 当然, 可视电 话的摄像头也可以作为第二接收模块集成在该装置中,用于采集视频信号来确 定音频信号的种类。
根据本发明实施例一至实施例三提供的方案,可以通过视频信号来确定音 频信号的种类, 从而确定对音频信号的编码方法, 提高音频编码质量, 减少声 音失真。 明的精神和范围。这样,倘若本发明的这些修改和变型属于本发明权利要求及 其等同技术的范围之内, 则本发明也意图包含这些改动和变型在内。

Claims

权 利 要 求
1、 一种音频信号处理方法, 其特征在于, 所述方法包括:
在接收音频信号的同时, 接收视频信号;
根据接收到的视频信号, 确定所述音频信号为语音信号或非语音信号; 若确定所述音频信号为语音信号,则利用基于时域编码的低码率音频编码 对所述音频信号进行编码;
若确定所述音频信号为非语音信号,则利用基于频域编码的低码率音频编 码对所述音频信号进行编码。
2、 如权利要求 1所述的方法, 其特征在于, 根据接收到的视频信号, 确 定所述音频信号为语音信号或非语音信号, 具体包括:
确定当前接收到的视频信号中是否存在指定的图像,若该视频信号中存在 指定的图像, 则确定距离该视频信号时间最短的一个已接收的视频信号: 若该已接收的视频信号中存在指定的图像,且在该已接收的视频信号中指 定的图像与当前接收到的视频信号中指定的图像的相对位移的绝对值满足设 定的阈值, 则确定当前接收到的音频信号为语音信号, 否则, 确定当前接收到 的音频信号为非语音信号。
3、 如权利要求 2所述的方法, 其特征在于, 所述方法还包括:
若确定当前接收到的视频信号中不存在指定的图像,则确定当前接收到的 音频信号为非语音信号;
若确定当前接收到的视频信号中存在指定的图像,且所述已接收的视频信 号中不存在指定的图像, 则确定当前接收到的音频信号为语音信号。
4、 如权利要求 1所述的方法, 其特征在于, 根据接收到的视频信号, 确 定所述音频信号为语音信号或非语音信号, 具体包括:
确定当前接收到的视频信号中是否存在指定的图像, 若不存在, 则确定当 前接收到的音频信号为非语音信号, 若存在, 则确定当前接收到的音频信号为 语音信号。
5、 如权利要求 1至 4中任一项所述的方法, 其特征在于, 利用基于时域 编码的低码率音频编码对所述音频信号进行编码之后,或利用基于频域编码的 低码率音频编码对音频信号进行编码之后, 所述方法还包括: 对编码后获得的数据进行量化,并组织码流输出,所述码流中包括标识位, 用于标识该码流对应的数据的编码方式。
6、 一种音频信号处理装置, 其特征在于, 所述装置包括:
第一接收模块, 用于接收音频信号;
第二接收模块, 用于接收视频信号;
确定模块, 用于根据接收到的视频信号,确定所述音频信号为语音信号或 非语音信号;
第一编码模块, 用于在确定模块确定所述音频信号为语音信号时, 利用基 于时域编码的低码率音频编码对所述音频信号进行编码;
第二编码模块, 用于在确定模块确定所述音频信号为非语音信号时, 利用 基于频域编码的低码率音频编码对所述音频信号进行编码。
7、 如权利要求 6所述的装置, 其特征在于,
所述确定模块,具体用于确定当前接收到的视频信号中是否存在指定的图 像, 若该视频信号中存在指定的图像,确定距离该视频信号时间最短的一个已 接收的视频信号:
若该已接收的视频信号中存在指定的图像,且在该已接收的视频信号中指 定的图像与当前接收到的视频信号中指定的图像的相对位移的绝对值满足设 定的阈值, 则确定当前接收到的音频信号为语音信号, 否则, 确定当前接收到 的音频信号为非语音信号。
8、 如权利要求 7所述的装置, 其特征在于,
所述确定模块,还用于在确定当前接收到的视频信号中不存在指定的图像 时, 确定当前接收到的音频信号为非语音信号; 以及, 在确定当前接收到的视 频信号中存在指定的图像, 且所述已接收的视频信号中不存在指定的图像时, 确定当前接收到的音频信号为语音信号。
9、 如权利要求 6所述的装置, 其特征在于,
所述确定模块,具体用于确定当前接收到的视频信号中是否存在指定的图 像, 若不存在, 则确定当前接收到的音频信号为非语音信号, 若存在, 则确定 当前接收到的音频信号为语音信号。
10、 如权利要求 6所述的装置, 其特征在于, 所述装置还包括: 码流输出模块, 用于对编码后获得的数据进行量化, 并组织码流输出, 所 述码流中包括标识位, 用于标识该码流对应的数据的编码方式。
11、 一种终端, 其特征在于, 所述终端包括如权利要求 6至 10中任一项 所述的音频信号处理装置。
12、 如权利要求 11所述的终端, 其特征在于, 所述终端还包括视频信号 采集模块和音频信号采集模块:
视频信号采集模块, 用于向所述第二接收模块提供视频信号;
音频信号采集模块, 用于向所述第一接收模块提供音频信号。
13、 如权利要求 11所述的终端, 其特征在于, 所述终端还包括音频信号 输出模块, 用于输出编码后的音频信号。
14、 如权利要求 13所述的终端, 其特征在于, 所述终端还包括视频信号 输出模块, 用于输出视频信号。
PCT/CN2012/086953 2012-01-04 2012-12-19 一种音频信号处理方法、装置及终端 Ceased WO2013102403A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201210001235.3 2012-01-04
CN201210001235.3A CN103198834B (zh) 2012-01-04 2012-01-04 一种音频信号处理方法、装置及终端

Publications (1)

Publication Number Publication Date
WO2013102403A1 true WO2013102403A1 (zh) 2013-07-11

Family

ID=48721308

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2012/086953 Ceased WO2013102403A1 (zh) 2012-01-04 2012-12-19 一种音频信号处理方法、装置及终端

Country Status (2)

Country Link
CN (1) CN103198834B (zh)
WO (1) WO2013102403A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108831472A (zh) * 2018-06-27 2018-11-16 中山大学肿瘤防治中心 一种基于唇语识别的人工智能发声系统及发声方法
CN111081264A (zh) * 2019-12-06 2020-04-28 北京明略软件系统有限公司 一种语音信号处理方法、装置、设备及存储介质

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105280188B (zh) * 2014-06-30 2019-06-28 美的集团股份有限公司 基于终端运行环境的音频信号编码方法和系统
CN105979469B (zh) * 2016-06-29 2020-01-31 维沃移动通信有限公司 一种录音处理方法及终端
CN115334349B (zh) * 2022-07-15 2024-01-02 北京达佳互联信息技术有限公司 音频处理方法、装置、电子设备及存储介质

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20030009325A1 (en) * 1998-01-22 2003-01-09 Raif Kirchherr Method for signal controlled switching between different audio coding schemes
US6754373B1 (en) * 2000-07-14 2004-06-22 International Business Machines Corporation System and method for microphone activation using visual speech cues
US20040267521A1 (en) * 2003-06-25 2004-12-30 Ross Cutler System and method for audio/video speaker detection
US20070174051A1 (en) * 2006-01-24 2007-07-26 Samsung Electronics Co., Ltd. Adaptive time and/or frequency-based encoding mode determination apparatus and method of determining encoding mode of the apparatus
CN101615393A (zh) * 2008-06-25 2009-12-30 汤姆森许可贸易公司 对语音和/或非语音音频输入信号编码或解码的方法和设备
CN101656070A (zh) * 2008-08-22 2010-02-24 展讯通信(上海)有限公司 一种语音检测方法

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7860718B2 (en) * 2005-12-08 2010-12-28 Electronics And Telecommunications Research Institute Apparatus and method for speech segment detection and system for speech recognition

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20030009325A1 (en) * 1998-01-22 2003-01-09 Raif Kirchherr Method for signal controlled switching between different audio coding schemes
US6754373B1 (en) * 2000-07-14 2004-06-22 International Business Machines Corporation System and method for microphone activation using visual speech cues
US20040267521A1 (en) * 2003-06-25 2004-12-30 Ross Cutler System and method for audio/video speaker detection
US20070174051A1 (en) * 2006-01-24 2007-07-26 Samsung Electronics Co., Ltd. Adaptive time and/or frequency-based encoding mode determination apparatus and method of determining encoding mode of the apparatus
CN101615393A (zh) * 2008-06-25 2009-12-30 汤姆森许可贸易公司 对语音和/或非语音音频输入信号编码或解码的方法和设备
CN101656070A (zh) * 2008-08-22 2010-02-24 展讯通信(上海)有限公司 一种语音检测方法

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108831472A (zh) * 2018-06-27 2018-11-16 中山大学肿瘤防治中心 一种基于唇语识别的人工智能发声系统及发声方法
CN108831472B (zh) * 2018-06-27 2022-03-11 中山大学肿瘤防治中心 一种基于唇语识别的人工智能发声系统及发声方法
CN111081264A (zh) * 2019-12-06 2020-04-28 北京明略软件系统有限公司 一种语音信号处理方法、装置、设备及存储介质

Also Published As

Publication number Publication date
CN103198834B (zh) 2016-12-14
CN103198834A (zh) 2013-07-10

Similar Documents

Publication Publication Date Title
US9218820B2 (en) Audio fingerprint differences for end-to-end quality of experience measurement
US8428959B2 (en) Audio packet loss concealment by transform interpolation
US20210089863A1 (en) Method and apparatus for recurrent auto-encoding
CN101165778B (zh) 音频信号的双变换编码方法和装置
CN108197572B (zh) 一种唇语识别方法和移动终端
US8386266B2 (en) Full-band scalable audio codec
CN110838894B (zh) 语音处理方法、装置、计算机可读存储介质和计算机设备
US8831932B2 (en) Scalable audio in a multi-point environment
TW201143445A (en) Controlling video encoding using audio information
CN101165777A (zh) 快速点阵向量量化
CN105551512A (zh) 音频格式转换方法和装置
WO2013102403A1 (zh) 一种音频信号处理方法、装置及终端
KR20180040716A (ko) 음질 향상을 위한 신호 처리방법 및 장치
WO2021244418A1 (zh) 一种音频编码方法和音频编码装置
WO2023197809A1 (zh) 一种高频音频信号的编解码方法和相关装置
WO2021244417A1 (zh) 一种音频编码方法和音频编码装置
CN112767955A (zh) 音频编码方法及装置、存储介质、电子设备
CN115831132A (zh) 音频编解码方法、装置、介质及电子设备
CN116137151B (zh) 低码率网络连接中提供高质量音频通信的系统和方法
CN101547010B (zh) 编码解码系统、方法及装置
JP2009522914A (ja) 通信網を介して加入者端末機に送信されるオーディオ信号の出力品質改善のためのオーディオ信号の処理方法およびこの方法を採用したオーディオ信号処理装置
WO2014000559A1 (zh) 语音频信号处理方法和编码装置
JP4437011B2 (ja) 音声符号化装置
CN101990082B (zh) 一种实现可视电话的方法及装置
HK40043832B (zh) 音频编码方法及装置、存储介质、电子设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 12864111

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 12864111

Country of ref document: EP

Kind code of ref document: A1