EP4550318A1 - Audiosignalcodierungsverfahren und -vorrichtung sowie elektronische vorrichtung und speichermedium - Google Patents

Audiosignalcodierungsverfahren und -vorrichtung sowie elektronische vorrichtung und speichermedium Download PDF

Info

Publication number
EP4550318A1
EP4550318A1 EP22948636.0A EP22948636A EP4550318A1 EP 4550318 A1 EP4550318 A1 EP 4550318A1 EP 22948636 A EP22948636 A EP 22948636A EP 4550318 A1 EP4550318 A1 EP 4550318A1
Authority
EP
European Patent Office
Prior art keywords
audio
encoding
audio signal
mixed
rate
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP22948636.0A
Other languages
English (en)
French (fr)
Other versions
EP4550318A4 (de
Inventor
Shuo Gao
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Xiaomi Mobile Software Co Ltd
Original Assignee
Beijing Xiaomi Mobile Software Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Xiaomi Mobile Software Co Ltd filed Critical Beijing Xiaomi Mobile Software Co Ltd
Publication of EP4550318A1 publication Critical patent/EP4550318A1/de
Publication of EP4550318A4 publication Critical patent/EP4550318A4/de
Pending legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/008Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • G10L19/167Audio streaming, i.e. formatting and decoding of an encoded audio signal representation into a data stream for transmission or storage purposes
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • G10L19/18Vocoders using multiple modes
    • G10L19/24Variable rate codecs, e.g. for generating different qualities using a scalable representation such as hierarchical encoding or layered encoding

Definitions

  • the disclosure relates to a field of communication technologies, in particular to an audio signal encoding method, an audio signal encoding apparatus, an electronic device and a storage medium.
  • the audio signal after acquiring an audio signal, the audio signal is subjected to a uniform encoding process.
  • a number of bits available for each audio channel is different, which causes the number of bits available for each audio channel to exceed or be less than a number of bits necessary for encoding, resulting in the waste of bits or the inability to provide an audio service that matches the encoding rate for remote users, which is an urgent problem to be solved.
  • Embodiments of the disclosure provide an audio signal encoding method, an audio signal encoding apparatus, an electronic device and a storage medium, to encode the audio signal according to a number of audio channels and an encoding rate, which may make full use of bits available during the encoding process, avoid a waste of bits, and provide audio services that match the encoding rate for remote users.
  • an audio signal encoding method includes: obtaining a scene-based audio signal; determining a number of audio channels of the audio signal and an encoding rate; and generating an encoded codestream by encoding the audio signal according to the number of audio channels and the encoding rate.
  • a scene-based audio signal is obtained, and a number of audio channels of the audio signal and an encoding rate are determined, and then an encoded codestream is generated by encoding the audio signal according to the number of audio channels and the encoding rate.
  • the audio signal is thus encoded according to the number of audio channels and the encoding rate, and during the encoding process, the bits available may be fully utilized, avoiding a waste of bits, and providing audio services that match the encoding rate for remote users.
  • generating the encoded codestream by encoding the audio signal according to the number of audio channels and the encoding rate includes: performing a down-mixed processing on the audio signal according to the number of audio channels and the encoding rate to generate a down-mixed parameter and a down-mixed audio channel signal; encoding the down-mixed audio channel signal to generate an encoding parameter; and generating the encoded codestream by performing codestream multiplexing on the down-mixed parameter and the encoding parameter.
  • performing the down-mixed processing on the audio signal according to the number of audio channels and the encoding rate to generate the down-mixed parameter and the down-mixed audio channel signal includes: determining a target control parameter for the audio signal according to the number of audio channels and the encoding rate; determining a down-mixed processing algorithm according to the target control parameter; and performing the down-mixed processing on the audio signal according to the down-mixed processing algorithm to generate the down-mixed parameter and the down-mixed audio channel signal.
  • determining the target control parameter for the audio signal according to the number of audio channels and the encoding rate includes: calculating an initial average rate of each channel according to the number of audio channels and the encoding rate; determining a target average rate according to the initial average rate and a preset average rate threshold; and determining the target control parameter for the audio signal according to the initial average rate and the target average rate.
  • the method before encoding the audio signal, the method further includes: performing a pre-emphasis preprocessing and/or a high-pass filtering preprocessing on the audio signal.
  • an audio signal encoding apparatus includes: a signal obtaining unit, configured to obtain a scene-based audio signal; an information determining unit, configured to determine a number of audio channels of the audio signal and an encoding rate; and an encoding processing unit, configured to generate an encoded codestream by encoding the audio signal according to the number of audio channels and the encoding rate.
  • the encoding processing unit includes: a down-mixed processing module, configured to perform a down-mixed processing on the audio signal according to the number of audio channels and the encoding rate to generate a down-mixed parameter and a down-mixed audio channel signal; a parameter generating module, configured to encode the down-mixed audio channel signal to generate an encoding parameter; and a codestream generating module, configured to generate the encoded codestream by performing codestream multiplexing on the down-mixed parameter and the encoding parameter.
  • the down-mixed processing module includes: a parameter determining sub-module, configured to determine a target control parameter for the audio signal according to the number of audio channels and the encoding rate; an algorithm determining sub-module, configured to determine a down-mixed processing algorithm according to the target control parameter; and a down-mixed processing sub-module, configured to perform the down-mixed processing on the audio signal according to the down-mixed processing algorithm to generate the down-mixed parameter and the down-mixed audio channel signal.
  • the parameter determining sub-module is further configured to: calculate an initial average rate of each channel according to the number of audio channels and the encoding rate; determine a target average rate according to the initial average rate and a preset average rate threshold; and determine the target control parameter for the audio signal according to the initial average rate and the target average rate.
  • the apparatus further includes: a preprocessing unit, configured to perform a pre-emphasis preprocessing and/or a high-pass filtering preprocessing on the audio signal.
  • a preprocessing unit configured to perform a pre-emphasis preprocessing and/or a high-pass filtering preprocessing on the audio signal.
  • an electronic device includes at least one processor, and a memory communicatively connected to the at least one processor.
  • the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the at least one processor to execute the method described in the first aspect.
  • a non-transitory computer-readable storage medium having computer instructions stored thereon is provided.
  • the computer instructions are configured to cause a computer to execute the method described in the first aspect.
  • a computer program product including computer instructions is provided.
  • the computer instructions are executed by a processor, the method described in the first aspect is implemented.
  • first and second in the specification and the claims of this disclosure and the drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
  • the terms “first” and “second” are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with the terms “first” and “second” may explicitly or implicitly include one or more of these features.
  • data so used may be interchanged under appropriate circumstances, so that the embodiments of the disclosure described herein may be implemented in other orders than those illustrated or described herein.
  • the implementations described in the following exemplary embodiments do not represent all implementations consistent with the disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the disclosure as detailed in the appended claims.
  • the term “at least one” in the disclosure may also be described as one or more, and the term “multiple” may be two, three, four, or more, which is not limited in the disclosure.
  • “first”, “second”, and “third”, and “A”, “B”, “C” and “D” are used to distinguish different technical features of the type, the technical features described using the “first”, “second”, and “third”, and “A”, “B”, “C” and “D” do not indicate any order of precedence or magnitude.
  • the correspondences shown in the tables in this disclosure may be configured or may be predefined.
  • the values of information in the tables are merely examples and may be configured to other values, which are not limited by the disclosure.
  • the correspondences illustrated in certain rows in the tables in this disclosure may not be configured.
  • the above tables may be adjusted appropriately, such as splitting, combining, and the like.
  • the names of the parameters shown in the titles of the above tables may be other names that may be understood by the communication device, and the values or representations of the parameters may be other values or representations that may be understood by the communication device.
  • Each of the above tables may also be implemented with other data structures, such as, arrays, queues, containers, stacks, linear tables, pointers, chained lists, trees, graphs, structures, classes, heaps, and Hash tables.
  • the first generation (1G) mobile communication technology is the first generation wireless cellular technology, which belongs to an analog mobile communication network.
  • 1G is upgraded to 2G
  • a mobile phone switches from an analog communication to a digital communication, and a global system for mobile communication (GSM) network standard is adopted.
  • GSM global system for mobile communication
  • a voice encoder adopts an adaptive multi rate (AMR) narrow band speech codec, an enhanced full rate (EFR), a full rate (FR), and a half rate (HR), and a communication provides a single-channel narrowband voice service.
  • AMR adaptive multi rate
  • EFR enhanced full rate
  • FR full rate
  • HR half rate
  • a 3G mobile communication system was proposed by the International Telecommunication Union (ITU) for international mobile communications in 2000, which can adopt Time Division-Synchronous Code Division Multiple Access (TD-SCDMA), Code Division Multiple Access 2000 (CDMA2000), or Wideband Code Division Multiple Access (WCDMA), and the voice encoder of which adopts an adaptive multi-rate wideband (AMR-WB) to provide a single-channel broadband voice service.
  • ITU International Telecommunication Union
  • TD-SCDMA Time Division-Synchronous Code Division Multiple Access
  • CDMA2000 Code Division Multiple Access 2000
  • WCDMA Wideband Code Division Multiple Access
  • AMR-WB adaptive multi-rate wideband
  • a 4G is improved base on the 3G technology.
  • Both data and voice are transmitted in an all-IP manner, for providing a real-time high definition (HD)+Voice service of voice audio.
  • An enhanced voice service (EVS) codec adopted by the 4G can balance a high-quality compression of voice and audio.
  • EVS enhanced voice service
  • the voice and audio communication service provided above have expanded from a narrowband signal to an ultra-wideband service or even a full-band service, but they are all signal-audio channel services.
  • stereo audio has a sense of orientation and distribution for each sound source, and can improve a clarity.
  • an upgrade of a terminal device signal collection device, an improvement of performance of a signal processor, and an upgrade of a terminal playback device three signal formats, namely audio channel-based multi-channel audio signals, object-based audio signals, and scene-based audio signals, can provide three-dimensional audio services.
  • An immersive voice and audio service (IVAS) codec that is being standardized by the 3rd Generation Partnership Project (3GPP) SA4 can support encoding and decoding requirements of the above three signal formats.
  • Terminal devices that can support 3D audio services include a mobile phone, a computer, a Pad, a conference system device, an augmented reality/virtual reality (AR/VR) device, a vehicle, etc.
  • a Firs-Order Ambisonics/High-Order Ambisonics (FOA/HOA) signal is a main scene-based audio signal.
  • the FOA/HOA signal represents audio information collected at a certain position in an audio scene and is an immersive audio format whose audio quality gradually gets better with the increase of order.
  • Different Ambisonics orders represent different numbers of audio signal components. That is, for an N-order Ambisonics signal, the number of Ambisonics coefficients is (N+1)*(N+1).
  • Table 1 the relationship between the Ambionics signal order and the Ambionics coefficient Ambisonics order Ambisonics coefficient/number of audio channels 0 1 1 4 2 9 3 16 4 25 5 36 6 49
  • the embodiment of the disclosure provides an audio signal encoding method and an audio signal encoding apparatus to solve the problems existing in the related art at least to some extent, so as to make full use of the available bits, provide the audio services that match the encoding rate for the remote user, and improve the user experience.
  • FIG. 1 is a flowchart of an audio signal encoding method according to an embodiment of the disclosure.
  • the method includes but is not limited to the following steps.
  • step S1 a scene-based audio signal is obtained.
  • the local user when a local user establishes a voice communication with any remote user, the local user can establish a voice communication with a terminal device of the any remote user through a terminal device of the local user.
  • the terminal device of the local user may obtain sound information of an environment where the local user is located in real time and obtain the scene-based audio signal.
  • the sound information of the environment where the local user is located includes sound information made by the local user and sound information of surrounding things.
  • the sound information of surrounding things may be, for example, sound information of vehicle driving, sound information of birds, sound information of wind, and sound information of other users around the local user, and so on.
  • the terminal device is an entity on a user side for receiving or transmitting signals.
  • the terminal device may be a mobile phone, a computer, a Pad, a watch, an interphone, a conference system device, an augmented reality/virtual reality (AR/VR) device, a vehicle, etc.
  • the terminal device may also be referred to as a user equipment (UE), a mobile station (MS), a mobile terminal (MT), and the like.
  • the terminal device may be a vehicle with communication functions, a smart vehicle, a mobile phone, a wearable device, a Pad, a computer with wireless transceiver functions, a VR terminal device, an AR terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in smart grid, a wireless terminal device in transportation safety, a wireless terminal device in smart city, a wireless terminal device in smart home, etc.
  • the specific technology and specific device form adopted by the terminal device are not limited in embodiments of the disclosure.
  • the terminal device of the local user when acquiring the scene-based audio signal, can acquire the sound information of the environment where the local user is located via a recording apparatus, such as a microphone, arranged in the terminal device or cooperating with the terminal device, and then generate the scene-based audio signal to obtain the scene-based audio signal.
  • a recording apparatus such as a microphone
  • the scene-based audio signal may be an audio signal in a FOA format or an audio signal in a HOA format.
  • step S2 a number of audio channels of the audio signal and an encoding rate are determined.
  • the number of audio channels of the audio signal and the encoding rate are determined.
  • the scene-based audio signal is an audio signal in the FOA format
  • the number of audio channels of the audio signal is 4, which may be represented by W, X, Y and Z, in which W represents a component containing all sounds in all directions in a sound field superimposed with the same gain and phase, X represents a component in a front-back direction in the sound field, Y represents a component in a left-right direction in the sound field, and Z represents a component in an up-down direction in the sound field.
  • the selected encoding rate is 96kbps.
  • an encoded codestream is generated by encoding the audio signal according to the number of audio channels and the encoding rate.
  • the scene-based audio signal is obtained, the number of audio channels of the audio signal and the encoding rate are determined, and the encoded codestream is generated by encoding the audio signal according to the number of audio channels and the encoding rate.
  • the encoding rate of each audio channel may be determined according to the number of audio channels and the encoding rate. For example, an average encoding rate of each audio channel, the maximum encoding rate of each audio channel, or the encoding rate of each audio channel may be determined. The average encoding rate of each audio channel may be determined by dividing the encoding rate by the number of audio channels, the maximum encoding rate of each audio channel is equal to the encoding rate, and the encoding rate of each audio channel is the encoding rate.
  • the number of bits available for each audio channel may be considered at the different encoding rates according to the encoding rate of each audio channel, so that the bits available is able to be fully utilized during the encoding process, to avoid the waste of bits and provide the audio services matching the encoding rate for the remote user.
  • the generated encoded codestream is able to provide clear, stable and understandable audio services when the encoding rate is low, and is able to provide high-definition, stable and immersive audio services when the encoding rate is high. In this way, it can provide the remote user with the audio services matching the encoding rate, thus improving the user experience.
  • the method before encoding the audio signal, the method further includes: performing a pre-emphasis preprocessing and/or a high-pass filtering preprocessing on the audio signal.
  • the pre-emphasis preprocessing may be performed on the audio signa, which may enhance a high-frequency portion of the audio information and increase a high-frequency resolution of the audio information.
  • the high-pass filtering preprocessing may be performed on the audio signal, to filter signal components in the audio signal lower than a certain frequency threshold.
  • a starting frequency in the high-pass filtering processing may be set as required, for example, the starting frequency may be set as 20Hz.
  • an audio signal component of the required encoding frequency band may be obtained.
  • an influence of an ultra-low frequency signal on encoding processing effects may be avoided.
  • the scene-based audio signal is obtained, the number of audio channels of the audio signal and the encoding rate are determined, and the encoded codestream is generated by encoding the audio signal according to the number of audio channels and the encoding rate.
  • the audio signal is encoded according to the number of audio channels and the encoding rate, and the bits available are able to be fully utilized during the encoding process, so that the waste of bits may be avoided, and the audio services that match the encoding rate may be provided for the remote user.
  • FIG. 3 is a flowchart of an audio signal encoding method according to an embodiment of the disclosure.
  • the method includes but is not limited to the following steps.
  • a scene-based audio signal is obtained.
  • step S20 a number of audio channels of the audio signal and an encoding rate are determined.
  • steps S10 and S20 can be referred to the related descriptions in the above embodiments, and the same contents will not be repeated here.
  • a down-mixed processing is performed on the audio signal according to the number of audio channels and the encoding rate, to generate a down-mixed parameter and a down-mixed audio channel signal.
  • the down-mixed audio channel signal is encoded to generate an encoding parameter.
  • the encoded codestream is generated by performing codestream multiplexing on the down-mixed parameter and the encoding parameter.
  • the scene-based audio signal is obtained, the number of audio channels of the audio signal and the encoding rate are determined, and the encoded codestream is generated by encoding the audio signal according to the number of audio channels and the encoding rate.
  • Encoding the audio signal according to the number of audio channels and the encoding rate may include performing the down-mixed processing on the audio signal according to the number of audio channels and the encoding rate to generate the down-mixed parameter and the down-mixed audio channel signal.
  • the down-mixed audio channel signal is then encoded to generate the encoding parameter.
  • the encoded codestream is generated by codestream multiplexing according to the down-mixed parameter and the encoding parameter.
  • the audio signal is subjected to a uniform down-mixed process, and the number of audio channels after down-mixed is less than the initial number of audio channels. All the remaining channels are encoded by a core encoder, and down-mixed parameters generated by the down-mixed processing and output parameters of the core encoder are performed by codestream multiplexing to output the encoded codestream.
  • the uniform down-mixed processing of the audio signal does not consider that the number of bits available for each audio channel is different under different encoding rates, resulting in the number of audio channels after the down-mixed processing does not match the number of audio channels that the core encoder is able to encode. Therefore, when the number of audio channels after the down-mixed processing is much less than the number of input audio channels, better audio services cannot be provided to remote users at a high encoding rate (because the number of bits available for each audio channel exceeds the number of bits necessary for encoding, which may lead to the waste of bits).
  • the remote users cannot be provided with audio services that match the encoding rate at a low encoding rate (because the number of bits available for each audio channel is much less than the number of bits necessary for encoding, which may lead to a poor encoding quality of each audio channel).
  • a scene-based audio signal (an audio signal in the FOA format or an audio signal in the HOA format) is input to an encoder end, the encoder end may determine the number of audio channels of the audio signal and the encoding rate and input the encoding rate, the number of audio channels and the audio signal to a pattern analysis module, or the encoder end may perform a high-pass filtering preprocessing on the audio signal and then input the preprocessed audio signal into the pattern analysis module.
  • the pattern analysis module may output a control parameter according to the selected encoding rate and the number of audio channels, and use the control parameter to guide a down-mixed processing module to select a corresponding down-mixed processing algorithm.
  • the down-mixed processing module outputs a down-mixed parameter and a down-mixed audio channel signal after processing the audio signal.
  • An encoding parameter is output after encoding the down-mixed audio channel signal by the core encoder.
  • the encoding parameter and the down-mixed parameter are input to a codestream multiplexer to output an encoded codestream.
  • a matching down-mixed processing algorithm is adaptively selected according to the number of audio channels of the input audio signal and the number of bits available, so that the number of audio channels after the down-mixed processing matches the number of audio channels that may be encoded by the core encoder at this encoding rate, and a full (optimal) utilization of bits available may be achieved. That is, at a low rate, it may ensure the provision of clear, stable and understandable audio services, and at a high rate, it may ensure the provision of high-definition, stable immersive audio services, which may improve the user experience.
  • the encoded codestream may be sent to a decoder end for decoding, so that the remote terminals may obtain sound information transmitted by the local terminal.
  • step S30 of performing the down-mixed processing on the audio signal according to the number of audio channels and the encoding rate to generate the down-mixed parameter and the down-mixed audio channel signal includes the following steps.
  • a target control parameter for the audio signal is determined according to the number of audio channels and the encoding rate.
  • the target control parameter for the audio signal when performing the down-mixed processing on the audio signal according to the number of audio channels and the encoding rate, the target control parameter for the audio signal may be determined according to the number of audio channels and the encoding rate.
  • the encoding rate of each audio channel may be determined according to the number of audio channels and the encoding rate. For example, an average encoding rate of each audio channel, the maximum encoding rate of each audio channel, or the encoding rate of each audio channel may be determined. The average encoding rate of each audio channel is determined by dividing the encoding rate by the number of audio channels, the maximum encoding rate of each audio channel is equal to the encoding rate, and the encoding rate of each channel is the encoding rate.
  • the target control parameter for the audio signal is determined according to the encoding rate of each audio channel.
  • the target control parameter for the audio signal when determining the target control parameter for the audio signal according to the number of audio channels and the encoding rate, by pre-setting corresponding relationships between the number of audio channels and the encoding rate and the control parameter, in a case of determining the number of audio channels of the audio signal and the encoding rate, the target control parameter for the audio signal may be determined.
  • a target number of audio channels may be determined according to the number of audio channels and the encoding rate, and then the target control parameter for the audio signal may be determined according to the target number of audio channels.
  • the target number of audio channels is determined according to the number of audio channels and the encoding rate.
  • N thresholds of the average encoding rate are preset, where N is a positive integer, and N+1 threshold ranges are determined by the N thresholds. Different threshold ranges are set to correspond to different numbers of audio channels after the down-mixed process.
  • an initial average encoding rate is calculated according to the number of audio channels and the encoding rate, and the target number of audio channels may be determined according to the threshold range to which the initial average rate belongs, and then the target control parameter for the audio signal is determined according to the target number of audio channels.
  • an average rate that is able to be allocated to each audio channel after the down-mixed processing may be obtained, and the target control parameter for the audio signal may be determined according to the target number of audio channels and/or the average rate that is able to be allocated to each audio channel after the down-mixed processing.
  • corresponding relationships between the target number of audio channels and/or the average rate that is able to be allocated to each audio channel after the down-mixed processing and the control parameter may be preset, and the target control parameter for the audio signal may be determined according to the target number of audio channels and/or the average rate that is able to be allocated to each audio channel after the down-mixed processing.
  • a down-mixed processing algorithm is determined according to the target control parameter.
  • the down-mixed processing is performed on the audio signal according to the down-mixed processing algorithm, to generate the down-mixed parameter and the down-mixed audio channel signal.
  • the audio signal may be performed by the down-mixed processing according to the down-mixed processing algorithm to generate the down-mixed parameter and the down-mixed audio channel signal.
  • step S301 of determining the target control parameter for the audio signal according to the number of audio channels and the encoding rate includes the following steps.
  • an initial average rate of each audio channel is calculated according to the number of audio channels and the encoding rate.
  • a target average rate is determined according to the initial average rate and a preset average rate threshold.
  • the target control parameter for the audio signal is determined according to the initial average rate and the target average rate.
  • the initial average rate of each audio channel may be calculated by dividing the encoding rate by the number of audio channels. For example, if the number of audio channels is 4 and the encoding rate is 96kbps, the initial average rate of each audio channel is calculated to be 24kbps according to the number of audio channels and the encoding rate.
  • the target average rate may be determined according to the initial average rate and the preset average rate threshold.
  • the preset average rate threshold may be set according to the scene-based audio signal. For example, a first average rate threshold Thres1 is set to 13.2kbps, and a second average rate threshold Thres2 is set to 32kbps. According to the above two average rate thresholds, ranges corresponding to the average rate is divided into three average rate ranges, as follows,
  • the target average rate is determined according to the initial average rate and the preset average rate threshold. If the average rate threshold range is determined according to the average rate threshold, the corresponding number of output audio channels is set for each average rate threshold range, so that the corresponding target number of output audio channels may be determined according to the average rate threshold range to which the initial average rate belongs.
  • the target average rate may be calculated according to the target number of output audio channels and the encoding rate.
  • the number of output audio channels corresponding to the average rate range 1 is 2, the number of output audio channels corresponding to the average rate range 2 is 3, and the number of output audio channels corresponding to the average rate range 3 is 4.
  • the initial average rate is 24kbps and belongs to the average rate range 2
  • the target number of output audio channels is 3
  • the number of output audio channels after the down-mixed processing matches the number of audio channels that may be encoded by the core encoder at this encoding rate, and the optimal use of available bits may be achieved. That is, at a low rate, it may ensure the provision of clear, stable and understandable audio services, and at a high rate, it may ensure the provision of high-definition, stable immersive audio services, which may improve the user experience.
  • three different types of down-mixed processing algorithms may be selected for scene-based audio signals. After the selected down-mixed processing, an average rate available for each audio channel in the average rate range 1 and the average rate range 2 are increasing after the down-mixed processing.
  • the average rate range 3 chooses not to perform the down-mixed processing because the encoding rate is rich enough, that is, an input signal is directly used as an output signal of the down-mixed processing, which means that the average rate available for each audio channel after the down-mixed processing remains unchanged.
  • Table 2 shows some kinds of scene-based audio signals, initial average rates (average rates that may be allocated to each audio channel initially), preset average rate thresholds, as well as corresponding numbers of output audio channels (numbers of audio channels after the down-mixed processing) and determined target average rates (average rates that may be allocated to each audio channel after the down-mixed processing).
  • the average rate that may be allocated to each audio channel after the down-mixed processing is greater than or equal to an average number of bits available for each audio channel, which may make full use of the available bits, avoid the waste of bits, and provide audio services that match the encoding rate for the remote users.
  • a target average rate is determined according to an initial average rate and a preset average rate threshold.
  • an average rate threshold closest to the initial average rate may be determined as the target average rate, or the initial average rate may be directly determined as the target average rate, or an average rate threshold, among average rate thresholds greater than the initial average rate, closest to the initial average rate may be determined as the target average rate, which is not specifically limited in the embodiment of the disclosure.
  • corresponding relationships between the initial average rate and the target average rate and the control parameter may be preset. For example, corresponding relationships between the initial average rate and the target average rate and the control parameter are set, or corresponding relationships between the control parameter and a difference between the initial average rate and the target average rate are set, or corresponding relationships between the control parameter and an absolute value of a difference between the initial average rate and the target average rate are set, or corresponding relationships between the control parameter and a sum of the initial average rate and the target average rate, etc., which is not specifically limited in the embodiment of the disclosure.
  • a down-mixed processing algorithm is to design a down-mixed conversion matrix according to a target number of output audio channels and a number of audio channels for acquiring scene-based audio signals. For example, if the number of audio channels is N and the target number of output audio channels is M, the conversion matrix is M*N, and both N and M are positive integers, and M is less than or equal to N.
  • M * 1 M * N * N * 1 where [M*1] represents a matrix of M times 1, [M*N] represents a matrix of M times N, and [N* 1] represents a matrix of N times 1.
  • the embodiment of the disclosure provides an exemplary embodiment.
  • the scene-based audio signal obtained is an audio signal in a FOA format
  • the number of audio channels is 4, namely, W, X, Y, Z
  • the selected encoding rate is 96kbps.
  • the target number of output audio channels is 3, where W represents a component containing all sounds in all directions in a sound field superimposed with the same gain and phase, X represents a component in a front-back direction in the sound field, Y represents a component in a left-right direction in the sound field, and Z represents a component in an up-down direction in the sound field.
  • the schematic diagram of the coordinate is shown in FIG. 2 .
  • the target number of audio channels is 3 after the down-mixed processing
  • the component Z in the up-down direction is omitted, and only three channel components, W, X and Y are reserved.
  • FIG. 8 is a structural diagram of an audio signal encoding apparatus provided by an embodiment of the disclosure.
  • the audio signal encoding apparatus 1 includes: a signal obtaining unit 11, an information determining unit 12, and an encoding processing unit 13.
  • the signal obtaining unit 11 is configured to obtain a scene-based audio signal.
  • the information determining unit 12 is configured to determine a number of audio channels of the audio signal and an encoding rate.
  • the encoding processing unit 13 is configured to generate an encoded codestream by encoding the audio signal according to the number of audio channels and the encoding rate.
  • the signal obtaining unit 11 obtains a scene-based audio signal
  • the information determining unit 12 determines a number of audio channels of the audio signal and an encoding rate
  • the encoding processing unit 13 generates a coded codestream by encoding the audio signal according to the number of audio channels and the encoding rate.
  • the audio signal is encoded according to the number of audio channels and the encoding rate, and the bits available may be fully utilized during the encoding process, so that the waste of bits may be avoided, and audio services that match the encoding rate may be provided for remote users.
  • the encoding processing unit 13 includes: a down-mixed processing module 131, a parameter generating module 132, and a codestream generating module 133.
  • the down-mixed processing module 131 is configured to perform a down-mixed processing on the audio signal according to the number of audio channels and the encoding rate to generate a down-mixed parameter and a down-mixed audio channel signal.
  • the parameter generating module 132 is configured to encode the down-mixed audio channel signal to generate an encoding parameter.
  • the codestream generating module 133 is configured to generate the encoded codestream by performing codestream multiplexing on the down-mixed parameter and the encoding parameter.
  • the down-mixed processing module 131 includes: a parameter determining sub-module 1311, an algorithm determining sub-module 1312, and a down-mixed processing sub-module 1313.
  • the parameter determining sub-module 1311 is configured to determine a target control parameter for the audio signal according to the number of audio channels and the encoding rate.
  • the algorithm determining sub-module 1312 is configured to determine a down-mixed processing algorithm according to the target control parameter.
  • the down-mixed processing sub-module 1313 is configured to perform the down-mixed processing on the audio signal according to the down-mixed processing algorithm to generate the down-mixed parameter and the down-mixed audio channel signal.
  • the parameter determining sub-module 1311 is further configured to:
  • the audio signal encoding apparatus 1 further includes: a preprocessing unit 14.
  • the preprocessing unit 14 is configured to perform a pre-emphasis preprocessing and/or a high-pass filtering preprocessing on the audio signal.
  • the audio signal encoding apparatus may execute the audio signal encoding methods as described in some of the above embodiments, and its beneficial effects are the same as those of the above audio signal encoding methods, which are not repeated here.
  • FIG. 12 is a structural diagram of an electronic device 100 for performing an audio signal encoding method illustrated by an exemplary embodiment.
  • the electronic device 100 may be a mobile phone, a computer, a digital broadcasting terminal, a message transceiver device, a game console, a tablet device, a medical device, a fitness device or a personal digital assistant.
  • the electronic device 100 may include one or more of the following components: a processing component 101, a memory 102, a power component 103, a multimedia component 104, an audio component 105, an input/output (I/O) interface 106, a sensor component 107, and a communication component 108.
  • the processing component 101 typically controls overall operations of the electronic device 100, such as the operations associated with display, telephone calls, data communications, camera operations, and recording operations.
  • the processing component 101 may include one or more processors 1011 to perform all or part of the steps in the above described methods.
  • the processing component 101 may include one or more modules which facilitate the interaction between the processing component 101 and other components.
  • the processing component 101 may include a multimedia module to facilitate the interaction between the multimedia component 104 and the processing component 101.
  • the memory 102 is configured to store various types of data to support the operation of the electronic device 100. Examples of such data include instructions for any applications or methods operated on the electronic device 100, contact data, phonebook data, messages, pictures, video, etc.
  • the memory 102 may be implemented using any type of volatile or non-volatile memory devices, or a combination thereof, such as a Static Random-Access Memory (SRAM), an Electrically-Erasable Programmable Read Only Memory (EEPROM), an Erasable Programmable Read Only Memory (EPROM), a Programmable Read Only Memory (PROM), a Read Only Memory (ROM), a magnetic memory, a flash memory, a magnetic or optical disk.
  • SRAM Static Random-Access Memory
  • EEPROM Electrically-Erasable Programmable Read Only Memory
  • EPROM Erasable Programmable Read Only Memory
  • PROM Programmable Read Only Memory
  • ROM Read Only Memory
  • the power component 103 provides power to various components of the electronic device 100.
  • the power component 103 may include a power management system, one or more power sources, and any other components associated with the generation, management, and distribution of power in the electronic device 100.
  • the multimedia component 104 includes a touch screen providing an output interface between the electronic device 100 and the user.
  • the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP).
  • the touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel.
  • the touch sensor may not only sense a boundary of a touch or swipe action, but also sense a period of wakeup time and a pressure associated with the touch or swipe action.
  • the multimedia component 104 includes a front-facing camera and/or a rear-facing camera. When the electronic device 100 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and/or the rear-facing camera can receive external multimedia data.
  • Each front-facing camera and rear-facing camera may be a fixed optical lens system or has focal length and optical zoom capability.
  • the audio component 105 is configured to output and/or input audio signals.
  • the audio component 105 includes a microphone (MIC) configured to receive an external audio signal when the electronic device 100 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode.
  • the received audio signal may be further stored in the memory 102 or transmitted via the communication component 108.
  • the audio component 105 further includes a speaker to output audio signals.
  • the I/O interface 106 provides an interface between the processing component 101 and peripheral interface modules, such as a keyboard, a click wheel, buttons, and the like.
  • the buttons may include, but are not limited to, a home button, a volume button, a starting button, and a locking button.
  • the sensor component 107 includes one or more sensors to provide status assessments of various aspects of the electronic device 100. For instance, the sensor component 107 may detect an open/closed status of the electronic device 100, relative positioning of components, e.g., the display and the keypad, of the electronic device 100, a change in position of the electronic device 100 or a component of the electronic device 100, a presence or absence of user contact with the electronic device 100, an orientation or an acceleration/deceleration of the electronic device 100, and a change in temperature of the electronic device 100.
  • the sensor component 107 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact.
  • the sensor component 107 may also include a light sensor, such as a Complementary Metal Oxide Semiconductor (CMOS) or Charge-Coupled Device (CCD) image sensor, for use in imaging applications.
  • CMOS Complementary Metal Oxide Semiconductor
  • CCD Charge-Coupled Device
  • the sensor component 107 may also include an accelerometer sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
  • the communication component 108 is configured to facilitate communication, wired or wirelessly, between the electronic device 100 and other devices.
  • the electronic device 100 can access a wireless network based on a communication standard, such as Wi-Fi, 2G or 3G, or a combination thereof.
  • the communication component 108 receives a broadcast signal or broadcast associated information from an external broadcast management system via a broadcast channel.
  • the communication component 108 further includes a Near Field Communication (NFC) module to facilitate short-range communication.
  • the NFC module may be implemented based on a Radio Frequency Identification (RFID) technology, an Infrared Data Association (IrDA) technology, an Ultra-Wide Band (UWB) technology, a Blue Tooth (BT) technology, and other technologies.
  • RFID Radio Frequency Identification
  • IrDA Infrared Data Association
  • UWB Ultra-Wide Band
  • BT Blue Tooth
  • the electronic device 100 may be implemented with one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic components, for performing the above described methods.
  • ASICs Application Specific Integrated Circuits
  • DSPs Digital Signal Processors
  • DSPDs Digital Signal Processing Devices
  • PLDs Programmable Logic Devices
  • FPGAs Field Programmable Gate Arrays
  • controllers micro-controllers, microprocessors or other electronic components, for performing the above described methods.
  • the electronic device 100 provided by the embodiment of the disclosure may execute the audio signal encoding method as described in some of the above embodiments, and its beneficial effects are the same as those of the above audio signal encoding method, which will not be repeated here.
  • the disclosure also provides a storage medium.
  • the storage medium may be a ROM, a Random Access Memory (RAM), Compact Disc-ROM (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
  • RAM Random Access Memory
  • CD-ROM Compact Disc-ROM
  • the storage medium may be a ROM, a Random Access Memory (RAM), Compact Disc-ROM (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
  • the disclosure also provides a computer program product.
  • the computer program is executed by the processor of the electronic device, the electronic device is caused to perform the audio signal encoding method as described above.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Mathematical Physics (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)
  • Signal Processing For Digital Recording And Reproducing (AREA)
EP22948636.0A 2022-06-30 Audiosignalcodierungsverfahren und -vorrichtung sowie elektronische vorrichtung und speichermedium Pending EP4550318A4 (de)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2022/103170 WO2024000534A1 (zh) 2022-06-30 2022-06-30 音频信号的编码方法、装置、电子设备和存储介质

Publications (2)

Publication Number Publication Date
EP4550318A1 true EP4550318A1 (de) 2025-05-07
EP4550318A4 EP4550318A4 (de) 2025-05-07

Family

ID=

Also Published As

Publication number Publication date
WO2024000534A1 (zh) 2024-01-04
CN117643073A (zh) 2024-03-01

Similar Documents

Publication Publication Date Title
EP4044578B1 (de) Tonverarbeitungsverfahren und elektronische vorrichtung
EP3139640A2 (de) Verfahren und vorrichtung zum erreichen von objektaudioaufzeichnung und elektronische vorrichtung
CN110865782B (zh) 数据传输方法、装置及设备
US20100053212A1 (en) Portable device having image overlay function and method of overlaying image in portable device
JP6195674B2 (ja) ネットワーク環境に基づく映像画質の調整方法、装置、プログラム、及び記録媒体
CN104782121A (zh) 多区域视频会议编码
WO2009051857A2 (en) System and method for video coding using variable compression and object motion tracking
US9930467B2 (en) Sound recording method and device
CN105392056B (zh) 电视情景模式的确定方法及装置
CN110166797B (zh) 视频转码方法、装置、电子设备及存储介质
WO2023216119A1 (zh) 音频信号编码方法、装置、电子设备和存储介质
EP4550318A1 (de) Audiosignalcodierungsverfahren und -vorrichtung sowie elektronische vorrichtung und speichermedium
CN116418995B (zh) 编解码资源的调度方法及电子设备
WO2024168556A1 (zh) 音频处理方法、装置
CN109348141A (zh) 视频生成、视频播放方法、装置、电子设备及存储介质
KR100879648B1 (ko) 절전형 화상통화 기능을 가지는 휴대용 단말기 및 휴대용단말기의 절전형 화상통화 방법
EP4697327A1 (de) Verfahren und vorrichtung zur verarbeitung von audiocodestromsignalen, elektronische vorrichtung und speichermedium
CN120374404A (zh) 一种图像处理方法、装置、电子设备、存储介质及芯片
US11800041B2 (en) Image processing method and apparatus, electronic device, and storage medium
EP4428807A1 (de) Verfahren und vorrichtung zur bildverarbeitung und medium
HK40069177B (zh) 音频处理的方法及电子设备
CN121010532A (zh) 图像处理方法、装置、电子设备、存储介质及程序产品
KR20060069954A (ko) 이동단말기의 화상통화를 이용한 멀티미디어 데이터 전송방법
CN117597936A (zh) 音频信号格式确定方法、装置
CN119011883A (zh) 直播音频流的处理方法、播放方法、装置及计算机设备

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250115

A4 Supplementary search report drawn up and despatched

Effective date: 20250408

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)
GRAP Despatch of communication of intention to grant a patent

Free format text: ORIGINAL CODE: EPIDOSNIGR1

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: GRANT OF PATENT IS INTENDED

RIC1 Information provided on ipc code assigned before grant

Ipc: G10L 19/008 20130101AFI20251030BHEP

Ipc: H04S 3/00 20060101ALI20251030BHEP

Ipc: G10L 19/16 20130101ALI20251030BHEP

Ipc: G10L 19/24 20130101ALN20251030BHEP

INTG Intention to grant announced

Effective date: 20251110

GRAJ Information related to disapproval of communication of intention to grant by the applicant or resumption of examination proceedings by the epo deleted

Free format text: ORIGINAL CODE: EPIDOSDIGR1

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: EXAMINATION IS IN PROGRESS

17Q First examination report despatched

Effective date: 20260316

INTC Intention to grant announced (deleted)