WO2010000161A1 - 一种基于即时通讯系统的语音通话方法及装置 - Google Patents
一种基于即时通讯系统的语音通话方法及装置 Download PDFInfo
- Publication number
- WO2010000161A1 WO2010000161A1 PCT/CN2009/071931 CN2009071931W WO2010000161A1 WO 2010000161 A1 WO2010000161 A1 WO 2010000161A1 CN 2009071931 W CN2009071931 W CN 2009071931W WO 2010000161 A1 WO2010000161 A1 WO 2010000161A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- voice
- tone
- information
- instant messaging
- messaging client
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/04—Real-time or near real-time messaging, e.g. instant messaging [IM]
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/0018—Speech coding using phonetic or linguistical decoding of the source; Reconstruction using text-to-speech synthesis
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/003—Changing voice quality, e.g. pitch or formants
- G10L21/007—Changing voice quality, e.g. pitch or formants characterised by the process used
- G10L21/013—Adapting to target pitch
- G10L2021/0135—Voice conversion or morphing
Definitions
- the invention belongs to the field of communications, and in particular relates to a voice call method and device based on an instant messaging system. Background of the invention
- the instant messaging system has many other additional functions, such as voice calling, in addition to basic instant messaging capabilities.
- voice calling in addition to basic instant messaging capabilities.
- the use of instant messaging systems for voice calls has become one of the communication tools used by the general public.
- existing voice calls can only use their original voice to make calls, and cannot change the original voice of the caller.
- the identity of the hidden party is hidden, lacking novelty and entertainment, and cannot meet the individual needs of the user.
- Embodiments of the present invention provide a voice-switching voice call method based on an instant messaging system, which aims to solve the problem of a voice-tuned voice call method based on an instant messaging system.
- a voice call method based on an instant messaging system comprising the steps of: establishing a tone-tuned voice call channel between at least two instant messaging clients; performing transposition processing on the input original voice information to obtain a pitch-adjusted voice;
- the pitch voice call channel sends the tone change voice to the instant messaging client B.
- Another object of the present invention is to provide a voice communication device based on an instant messaging system, where the device includes: a request sending unit, configured to establish a tone-tuned voice call channel;
- a voice collection unit configured to collect input original voice information
- a tone adjustment processing unit configured to perform transposition processing on the original voice information collected by the voice collection unit to obtain a tone-modulated voice
- a voice sending unit configured to send the tone-tuned voice obtained by the tone changing processing unit by using the tone-tuned voice call channel established by the request sending unit.
- the voice signal collected in the instant messaging system is first subjected to voice tone modulation processing, and the voice-changing voice call based on the instant messaging system is realized, which brings great entertainment effect to the voice communication based on the instant communication occasion.
- bringing new value-added service growth points to traditional instant messaging services, increasing users' dependence on instant messaging products, thereby enhancing product competitiveness and providing a new business experience for voice call users.
- FIG. 1 is a basic flowchart of a method according to an embodiment of the present invention
- FIG. 3 is another detailed flowchart of a method according to an embodiment of the present invention.
- FIG. 4 is a flowchart of processing, by the instant messaging client B, the tone-switched voice call data sent by the instant messaging client A according to an embodiment of the present invention
- FIG. 5 is a basic structural diagram of an apparatus according to an embodiment of the present invention.
- FIG. 6 is a detailed structural diagram of an apparatus according to an embodiment of the present invention. Mode for carrying out the invention
- a tone change between at least two instant messaging clients can be established.
- the voice call channel for example, can establish a voice-tuned voice call channel between the instant messaging client VIII, the instant messaging client B, and the instant messaging client C.
- a variable tone voice channel is established between A and the instant messaging client B.
- the instant messaging client A sends a voice tone change request to the instant messaging client B, and establishes a tone-adjusted voice call channel with the instant messaging client B, and then performs transposition processing on the collected original voice to obtain a corresponding voice corresponding to the original voice.
- the tone voice is transmitted, and the tone voice is sent to the instant messaging client B through the established tone voice communication channel, thereby realizing the tone voice call between the instant messaging clients in the instant messaging system.
- FIG. 1 is a basic flowchart of an embodiment of the present invention. As shown in FIG. 1 , this embodiment takes an example of establishing a tone-tuned voice communication channel between the instant messaging client A and the instant messaging client B. The process may include the following steps:
- Step S101 Establish a tone-adjusted voice call channel between the instant messaging client A and the instant messaging client B.
- Step S102 Perform transposition processing on the input original speech signal to obtain a pitch-adjusted speech.
- Step S103 the tone-adjusted voice is sent to the instant messaging client B through the voice-adjusted voice call channel.
- the instant messaging client A and the instant messaging client B may be implemented in multiple implementations, such as a terminal in the form of a web, or a terminal in a wireless form, which is not specifically limited in the embodiment of the present invention. .
- the operations of the foregoing steps S102 and S103 can be performed by the instant messaging client A.
- the instant messaging client A directly transposes the input original voice information to obtain a pitch-adjusted voice; Transmitting the tone-changing voice to the instant messaging client B through the server relay or P2P mode; or performing the preset tone-changing processing device such as a server, for example, the preset server receives the instant messaging client
- the original voice information sent by the terminal A after which the original voice information is subjected to a tone-modulation process to obtain a tone-modulated voice; how to transmit the tone-changing voice to the instant messaging client through the voice-switched voice call channel, how to implement the embodiment of the present invention Not specifically limited.
- the following is an example of performing a voice call operation between two clients.
- FIG. 2 is a detailed flowchart of an embodiment of the present invention, which is described in detail as follows:
- the instant messaging client A sends a tone-changing voice call request to the instant messaging client B.
- the instant messaging client B After receiving the voice-tuned voice call request sent by the instant messaging client A, the instant messaging client B responds to the voice-tuned voice call request and returns the response message to the instant messaging client A.
- the instant messaging client A receives the voice-changing voice call response returned by the instant messaging client B, a tone-tuned voice call channel is established between the instant messaging client B and the instant messaging client B.
- the instant messaging client A and the instant messaging client B establish a tone-adjusted voice call channel under the coordination of the instant messaging server.
- instant messaging client A can send a voice-tuned voice call request to instant messaging client B transparently or opaquely. If the instant messaging client A transparently sends a tone-change voice call request to the instant messaging client B, the process does not need to be displayed on the instant messaging client B interface.
- the instant messaging client A performs transposition processing on the original voice collected, and obtains the pitch-adjusted voice corresponding to the original voice.
- a plurality of voice transposition modes are provided, such as changing the pitch of the voice, the gender change sound (the male voice becomes a female voice, the female voice becomes a male voice), the age change sound (the voice of the teenager becomes the voice of the elderly), and the original voice of the user is used.
- the voice of a famous person on the user's voice
- the background sound is added (strictly speaking, adding a background sound to the user's voice does not belong to the voice transposition processing, but belongs to the mixing technique, but the pitch-switched voice call defined by the present invention includes such an application) and the like.
- the speech modulating process can use a linear prediction (LP) analysis synthetic speech model to decompose the digital speech signal into a spectral envelope portion (represented by Linear predictive coding (LPC) coefficients) and excitation. Part (represented by the residual of LPC); Then, the formant frequency and the spectral tilt parameter are extracted on the LPC coefficient, and the speech conversion is realized by vector quantization code calligraphy.
- LP linear prediction
- the frequency envelope conversion can use vector quantization
- the conversion of the prosody mainly the pitch period
- TD-PSOLA time domain pitch synchronous overlap-add
- determining the operation of the voice tone modulation mode to be currently used may include: determining current tone change information; thus, the voice tone adjustment mode currently to be used may be determined according to current tone change information.
- the current tone change information may specifically include user selection information, and/or authorized tone change information.
- the user selection information is a selection made by the user in the provided voice tone changing mode; and the authorized tone changing information is a tone changing information authorized by the tone changing party user in the instant messaging system.
- the voice tone adjustment mode is determined by the tone change mode of the tone change party user in the instant messaging system, and the voice tone adjustment mode may be used as the value added service item.
- the user sends the user's authorized tone changing mode query information to the server through the instant messaging client A, and the server returns the authorized tone changing information according to the identity of the user in the instant messaging system, that is, the user The voice transposition method available to the user.
- the instant messaging client A further inputs the user selection information according to the authorized tone changing information returned by the server, thereby determining the usable voice tone changing mode according to the user selection information and the authorized tone changing information returned by the server.
- the embodiment may also determine the tone change mode according to the user selection information and the authorization tone change information, and use other service selection logics. When the user has only one voice tone adjustment mode that can be used, the tone change mode may be determined only by the authorized tone change information.
- the speech transposition method of speech transposition processing also considers the user personality characteristic information, that is, mainly the segment features in the original speech of the user.
- the transposition mode is determined by the service selection logic according to the user selection information and the user personality characteristic information, or the user selection information, the authorization transposition information, and the user personality characteristic information.
- the service selection logic is defined by the instant messaging service provider to clearly indicate what kind of authorized tone changing information, and what kind of voice-tuning voice service can be enjoyed by the voice communication environment (for example: "Male voice change female voice" is a voice-changing voice service. ), etc., which is used to determine the way the voice is changed.
- the client A After receiving the user selection information, the client A analyzes the original voice signal of the user and obtains the personality feature information.
- the personality feature information cannot meet the requirements of the voice tone modulation process, the user needs to correct the voice tone adjustment request. For example: A user's original voice is thick and hoarse, and the voice tone of his choice is "younger children". The adjustment effect will be poor (it is not easy to identify the other party as "children"), so the system should advise the user to reselect the voice tone mode.
- the voice tone mode confirmation also considers the voice environment information of the other party.
- the tone changing mode is determined by the service selection logic according to the user selection information and the counterpart voice environment information, or the user selection information, the authorized tone changing information, and the counterpart voice environment information.
- Instant messaging client ⁇ to the instant messaging client ⁇ return to tone voice call response, while returning its own voice environment information.
- the voice environment information can be selected by the instant messaging client B user, or analyzed by the instant messaging client B according to the sound signal collected by the microphone.
- the voice tone adjustment mode of the instant messaging client A may be determined by the service selection logic by one or any combination of authorization tone information, user personality feature information, and counterpart voice environment information and user selection information.
- the collected voice information may include signals that are not conducive to processing, transmission, and discrimination, such as echo, noise, and the like, in order to achieve a better voice-tuning voice call effect, the voice heard by the call receiver is improved.
- Quality before performing the speech transposition processing on the digital speech information, denoising the digital speech information, that is, performing one or more combinations of echo cancellation, noise suppression, signal gain adjustment, and the like.
- the instant messaging client A sends the changed tone voice to the instant messaging client B through the established tone voice call channel.
- the instant messaging client A groups and packs the tone-modulated voice before transmitting the tone-tuned voice, obtains the tone-modulated voice data packet, and sends the voice-change voice data packet to the instant messaging client.
- the tone is processed.
- preset coding rules such as G.729, G.729A, G.723.1, etc.
- the obtained tone-modulated speech corresponding to the original voice is compression-encoded.
- the transcoding is obtained by using the channel coding technique.
- the voice bit stream is subjected to redundancy enhancement processing.
- the instant messaging client B sends a voice call request to the instant messaging client A
- the implementation process is the same as above, and details are not described herein. It can be understood that the instant messaging client A and the instant messaging client B can perform a one-way voice call and a two-way tone voice call.
- the above voice calls are based on an instant messaging system on a wired internet or wireless internet.
- FIG. 3 is another detailed flowchart of the embodiment of the present invention.
- the embodiment establishes a voice communication channel between the instant messaging client A and the instant messaging client B, and the instant messaging client A and the instant messaging client B.
- the implementation process of the voice call method is as follows:
- Instant messaging client A sends a voice call request to instant messaging client B.
- the instant messaging client B After receiving the voice call request sent by the instant messaging client A, the instant messaging client B responds to the voice call request and returns the response message to the instant messaging client A. After receiving the voice call returned by the instant messaging client B, the instant messaging client A establishes a voice call channel with the instant messaging client B.
- the voice call channel can be used for voice communication between the instant messaging client A and the instant messaging client B.
- the instant messaging client A sends a tone-changing voice call request to the instant messaging client B. 4. After receiving the voice-tuned voice call request sent by the instant messaging client A, the instant messaging client B responds to the voice-tuned voice call request, and returns the response message to the instant messaging client A. After receiving the voice-changing voice call response returned by the instant messaging client B, the instant messaging client A establishes a tone-adjusted voice call channel with the instant messaging client B.
- the previously established voice call channel can be released.
- instant messaging client A can send a voice-tuned voice call request to instant messaging client B transparently or opaquely. If the instant messaging client A transparently sends a tone-changing voice call request to the instant messaging client B, the process does not need to be displayed on the instant messaging client B interface.
- Instant messaging client A performs transposition processing on the original voice collected to obtain a tone-changing voice corresponding to the original voice.
- the instant messaging client A sends the changed tone voice to the instant messaging client B through the established tone voice call channel.
- a tone-tuned voice call channel is established between the instant messaging client A and the instant messaging client B.
- the instant messaging client A receives the voice-changing voice call response returned by the instant messaging client B, it may not establish the relationship with the instant messaging client B.
- the tone voice sent to the instant messaging client B is sent by using the voice call channel established in step 2 above with the instant messaging client B.
- the operation of establishing a tone-tuned voice call channel in step 4 is omitted.
- whether the bandwidth corresponding to the current voice call channel can be adapted to transmit the pitch-adjusted voice obtained in step 5 is used to determine that the voice call channel needs to be established.
- the processing flow of the call data is the same as that in the normal voice call.
- the processing flow is as shown in FIG. 4, and the details are as follows:
- step S401 call data is received and unpacked
- the packet call data is received through the established tone-tuned voice call channel, the data packet is unpacked according to the same network transmission protocol as the instant messaging client A, and the packet data is assembled to obtain a compressed code stream.
- step S402 the unpacked data is decoded into a voice signal
- the unpacked compressed code stream is decoded by the inverse operation of the instant messaging client A coding operation to obtain an original voice signal that can be recognized by the human ear.
- step S403 a speech signal enhancement process
- the signal enhancement processing may employ a Kalman filter method, a minimum mean square error estimation method for short-term spectral amplitude, or an adaptive filtering method.
- step S404 the enhanced processed speech signal is output.
- the processed speech signal output is enhanced by an output device such as a headphone, a speaker, a sound card, or the like.
- the unpacked data is subjected to inverse redundancy/fault-tolerant processing, and the instant messaging client A is added to the compressed code stream. Redundant signals, modify or discard the erroneous data.
- FIG. 5 is a basic flowchart of an apparatus according to an embodiment of the present invention.
- the device may include a request sending unit 501, a voice collecting unit 502, and a tone modulation process.
- the request sending unit 501 is configured to establish a tone-tuned voice call channel.
- the voice collection unit 502 is configured to collect the input original voice information.
- the pitch adjustment processing unit 503 is configured to perform pitch modulation processing on the original voice information collected by the voice collection unit 502 to obtain a tone-modulated voice.
- the voice transmitting unit 504 transmits the tone-tuned voice obtained by the tone changing processing unit 503 by the tone-tuned voice call channel established by the request transmitting unit 501.
- the voice communication device based on the instant messaging system provided by the embodiment of the present invention is implemented.
- FIG. 6 is a detailed structural diagram of an apparatus according to an embodiment of the present invention. For convenience of description, only parts related to the embodiment of the present invention are shown.
- the device can be used in various instant messaging client devices, such as computers, notebook computers, personal digital assistants (PDAs), smart phones, etc., and can be software units, hardware units or hardware and software running in these devices.
- the combined unit may also be integrated into these devices or run in the application system of these devices as a separate pendant.
- the device may include: a request sending unit 601, a voice collecting unit 602, a tone changing processing unit 603, and a voice. Transmitting unit 604.
- the request sending unit 601 is configured to establish a tone-tuned voice call channel.
- the voice collection unit 602 collects the input original voice information.
- the pitch adjustment processing unit 603 performs a pitch adjustment process on the original voice information collected by the voice collection unit 602 to obtain a tone-modulated voice.
- the voice sending unit 604 sends the tone-tuned voice obtained by the tone-modulating processing unit 603 by the tone-tuned voice call channel established by the requesting sending unit 601.
- the request sending unit 601, the voice collecting unit 602, the tone adjusting processing unit 603, and the voice sending unit 604 may be on the same entity, such as the instant messaging client A, or may not be on the same entity, for example, the request sending unit 601.
- the voice collection unit 602 is on the same entity as the instant messaging client A, and the tone modulation processing unit 603 and the voice sending unit 604 are on a preset tone changing processing device such as a server.
- the specific situation requires specific analysis, which is not specifically limited in the embodiment of the present invention.
- the request sending unit 601 establishes a tone-tuned voice call channel after receiving the tone-change voice call response; wherein the tone-change voice call response is a response to the tone-change voice call request sent by the request sending unit 601.
- the request sending unit 601 is further configured to receive the voice-tuned voice call request information input by the user.
- the voice collection unit 603 is also used to convert the collected voice information into digital voice information.
- the converted digital voice information is recognized and processed by the computer.
- the pitch adjustment processing unit 603 includes: a pitch adjustment information confirmation module 6031, a business logic module 6032, and a voice tone modulation processing module 6033.
- the tone change information confirming module 6031 is configured to determine and output current tone change information, wherein the current tone change information includes user selection information, and/or authorization tone change information.
- the business logic module 6032 is configured to generate service selection logic, wherein the service selection logic is used to perform voice tone modulation processing and output to the voice tone modulation processing module 6033.
- the service selection logic is defined by the instant messaging service provider to clearly indicate what kind of tone-changing information, and what kind of voice-tuning voice service can be enjoyed by the voice communication environment (for example: "Male voice" is a tone-change voice service) Wait.
- the voice tone processing module 6033 is configured to determine the voice tone modulation mode according to the received tone change information outputted from the tone change information receiving module 6031 and the service selection logic output by the service logic module 6032, and change the digital voice information obtained by the voice collection unit 602 according to the voice.
- the method performs a pitch adjustment process to obtain a tone-modulated speech corresponding to the digital voice information and outputs the tone.
- the voice tone processing module 6033 determines the voice tone modulation mode from the service selection logic according to the user selection information included in the tone change information and/or the authorized tone change information.
- the implementation is as described above, and will not be described again.
- the tone modulation processing unit 603 further includes: a user feature acquisition module 6034.
- the user feature obtaining module 6034 is configured to extract personality feature information from the digital voice information obtained by the voice collecting unit 602, and generate and output the personality feature information.
- the voice tone modulation processing module 6033 parses out the user selection information contained in the received current tone change information, and/or the authorized tone change information, and combines the received user personality feature information to determine the voice tone modulation mode by the service selection logic.
- the tone adjustment processing unit 603 further includes: a counterpart environment acquisition module 6035.
- the other party environment obtaining module 6035 is configured to acquire the voice environment information of the counterparty carried by the voice-switching voice call response received by the request sending unit 601.
- the voice-switching voice call response returned by the caller includes the voice environment information
- the request sending unit 601 generates the voice environment information of the partner according to the received voice environment information, and thus, the counterpart environment obtaining module 6035 acquires the request to send.
- the device does not have to include the user feature acquisition module 6034 and the counterpart environment acquisition module 6035.
- the embodiment may also include one or more of the user feature acquisition module 6034 and the counterpart environment acquisition module 6035.
- the transposition processing unit 603 of FIG. 6 includes a user feature acquisition module 6034 and a counterpart environment acquisition module 6035 as an example.
- the voice tone modulation processing module 6033 can use the service selection logic sent by the service logic module 6032, the current tone change information sent by the tone change information confirmation module 6031, and the user special The personalization feature information sent by the acquisition module 6034; or the service selection logic sent by the service logic module 6032, the current transposition information sent by the transposition information confirmation module 6031, and the counterpart speech environment information sent by the counterpart environment acquisition module 6035; or according to the business logic module
- the service selection logic sent by the 6032, the current tone change information sent by the tone change information confirmation module 6031, the personality feature information sent by the user feature acquisition module 6034, and the counterpart voice environment information sent by the counterpart environment acquisition module 6035 determine the voice tone modulation mode.
- the voice call device further includes: a denoising unit 605.
- the denoising unit 605 receives the digital voice information obtained by the voice collecting unit 602, performs denoising processing, and obtains the denoised digital voice information.
- the voice communication device further includes: a coding unit 606, and/or an optimization unit 607.
- Figure 6 shows an example of a voice call device including an encoding unit 606 and an optimization unit 607.
- the coding unit 606 compression-codes the pitch-modulated speech obtained by the pitch-modulation processing unit 603 to obtain a tone-modulated speech bit stream.
- the optimizing unit 607 performs redundancy enhancement processing, and/or grouping and packing processing on the tone-modulated bit stream obtained by the encoding unit 606, and outputs the processed tone-modulated voice data to the voice transmitting unit 604.
- the main purpose of the optimization unit 607 is to avoid signal distortion caused by packet loss, error, etc. during the transmission of the tone-modulated speech, or to facilitate transmission of the tuned speech.
- the optimization unit may perform redundancy enhancement processing on the pitch-adjusted speech obtained by the transposition processing unit 603, and/or grouping and packing processing, and output the processed transposed speech data to the voice transmission. Unit 604.
- the optimization unit 607 can include: The redundancy enhancement processing module 6071 performs redundancy enhancement processing on the tone-modulated bit stream obtained by the coding unit 606 or the tone-modulated speech obtained by the modulation processing unit 603 by using a channel coding technique, and outputs the processed tone-modulated voice bit stream.
- the packet and packet module 6072 groups and receives the received tone voice data to obtain a tone-modulated voice data packet.
- the packetizing and packing module 6072 can receive the pitch-modulated speech and the tuned speech bitstream output by the transcoding processing unit 603, the encoding unit 606, or the redundancy enhancement processing module 6071.
- optimization unit 607 may include only the redundancy enhancement processing module 6071 or the grouping and packaging module 6072.
- the voice communication device further includes:
- the request answering unit 608 receives the voice-tuned voice call request sent by the request sending unit 601, and returns a voice-changing voice call response, generates voice receiving trigger information, and outputs the voice receiving trigger information to the voice receiving unit 609.
- the voice receiving unit 609 after receiving the voice receiving trigger information output by the request response unit 608, if the currently received data packet is a packet processed or packetized, the data packet solution is performed according to the same network transmission protocol as the calling party.
- the packet is assembled and the packet data is assembled to obtain a compressed code stream and output.
- the decoding unit 610 decodes the data obtained by the voice receiving unit 609, that is, the compressed code stream, into a voice signal.
- the speech signal enhancement processing unit 611 decodes the data obtained by the decoding unit 610, that is, decodes the speech signal, obtains the original speech signal, and performs signal enhancement processing to obtain the enhanced speech signal.
- the voice output unit 612 outputs the obtained enhanced voice signal, which may be an earphone, a speaker, a sound card, or the like.
- the voice call device further includes: an inverse redundancy/fault-tolerant processing unit 613.
- the inverse redundancy/fault-tolerant processing unit 613 is configured to remove the redundant signal that the call partner obtained by the voice receiving unit 609 adds to the compressed code stream, and modify or discard the error data therein. In this way, the voice quality can be greatly improved.
- the request response unit 608, the voice receiving unit 609, the decoding unit 610, the voice signal enhancement processing unit 611, the voice output unit 612, the inverse redundancy/fault-tolerant processing unit 613, and the entity in which the voice receiving unit 614 is removed are transmitted with the above request.
- the unit 601, the voice collection unit 602, the tone adjustment processing unit 603, the voice sending unit 604, the denoising unit 605, the encoding unit 606, and the optimization unit 607 have different entities, for example, if the request sending unit 601, the voice collecting unit 602, The tone processing unit 603, the voice sending unit 604, the voice sending unit 604, the denoising unit 605, the encoding unit 606, and the optimizing unit 607, on the instant messaging client A, request the answering unit 608, the voice receiving unit 609, the decoding unit 610,
- the voice signal enhancement processing unit 611, the voice output unit 612, the inverse redundancy/fault-tolerant processing unit 613, and the entity in which the voice removal unit 614 is removed may be the opposite end (receiver) of the instant messaging client A, such as the instant messaging client B.
- the request sending unit 601 and the voice collecting unit 602 are on the same entity, such as the instant messaging client A, and the tone changing processing unit 603 and the voice sending unit 604 are on a preset tone changing processing device such as the server 1, the response unit 608 is requested.
- the voice receiving unit 609, the decoding unit 610, the voice signal enhancement processing unit 611, the voice output unit 612, the inverse redundancy/fault-tolerant processing unit 613, and the entity in which the voice-removing unit 614 is removed may be the opposite end (receiver) of the server 1.
- the instant messaging client is described. The foregoing is only an example and is not intended to limit the application of the embodiments of the present invention.
- the voice signal collected in the instant messaging system is first spoken. Tone tone processing, realizes voice-changing voice call based on instant messaging system, brings great entertainment effect to voice communication based on instant messaging occasion, brings new value-added service growth point to traditional instant messaging service, increases user-to-instantaneous The dependence of communication products, thereby enhancing product competitiveness. And provide a new business experience for voice call users, for example: use variable tone voice calls to achieve the purpose of protecting user identity information.
Landscapes
- Engineering & Computer Science (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Computer Networks & Wireless Communication (AREA)
- Telephonic Communication Services (AREA)
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US12/913,358 US20110044324A1 (en) | 2008-06-30 | 2010-10-27 | Method and Apparatus for Voice Communication Based on Instant Messaging System |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CNA2008100682626A CN101304391A (zh) | 2008-06-30 | 2008-06-30 | 一种基于即时通讯系统的语音通话方法及系统 |
| CN200810068262.6 | 2008-06-30 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US12/913,358 Continuation US20110044324A1 (en) | 2008-06-30 | 2010-10-27 | Method and Apparatus for Voice Communication Based on Instant Messaging System |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2010000161A1 true WO2010000161A1 (zh) | 2010-01-07 |
Family
ID=40114104
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2009/071931 Ceased WO2010000161A1 (zh) | 2008-06-30 | 2009-05-22 | 一种基于即时通讯系统的语音通话方法及装置 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20110044324A1 (zh) |
| CN (1) | CN101304391A (zh) |
| WO (1) | WO2010000161A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109272984A (zh) * | 2018-10-17 | 2019-01-25 | 百度在线网络技术(北京)有限公司 | 用于语音交互的方法和装置 |
Families Citing this family (28)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101304391A (zh) * | 2008-06-30 | 2008-11-12 | 腾讯科技(深圳)有限公司 | 一种基于即时通讯系统的语音通话方法及系统 |
| US9838784B2 (en) | 2009-12-02 | 2017-12-05 | Knowles Electronics, Llc | Directional audio capture |
| US8798290B1 (en) | 2010-04-21 | 2014-08-05 | Audience, Inc. | Systems and methods for adaptive signal equalization |
| CN101888607A (zh) * | 2010-07-15 | 2010-11-17 | 中兴通讯股份有限公司 | 基于widget实现手机聊天的方法及手机 |
| US9021565B2 (en) * | 2011-10-13 | 2015-04-28 | At&T Intellectual Property I, L.P. | Authentication techniques utilizing a computing device |
| CN104144097B (zh) * | 2013-05-07 | 2018-09-07 | 北京音之邦文化科技有限公司 | 语音消息传输系统、发送端、接收端及语音消息传输方法 |
| US9536540B2 (en) | 2013-07-19 | 2017-01-03 | Knowles Electronics, Llc | Speech signal separation and synthesis based on auditory scene analysis and speech modeling |
| CN104376846A (zh) * | 2013-08-16 | 2015-02-25 | 联想(北京)有限公司 | 一种语音调节方法、装置和电子设备 |
| CN104780091B (zh) * | 2014-01-13 | 2019-06-25 | 北京发现角科技有限公司 | 一种具有语音音频处理功能的即时通信方法和系统 |
| CN104980396A (zh) * | 2014-04-03 | 2015-10-14 | 北京千橡网景科技发展有限公司 | 一种用于社交网络的通信方法及系统 |
| CN105208056B (zh) * | 2014-06-18 | 2020-07-07 | 腾讯科技(深圳)有限公司 | 信息交互的方法及终端 |
| CN104200824B (zh) * | 2014-08-25 | 2019-05-03 | 努比亚技术有限公司 | 音频录制方法和装置 |
| WO2016040885A1 (en) | 2014-09-12 | 2016-03-17 | Audience, Inc. | Systems and methods for restoration of speech components |
| US20160093307A1 (en) * | 2014-09-25 | 2016-03-31 | Audience, Inc. | Latency Reduction |
| CN107210824A (zh) | 2015-01-30 | 2017-09-26 | 美商楼氏电子有限公司 | 麦克风的环境切换 |
| CN106506437B (zh) * | 2015-09-07 | 2021-03-16 | 腾讯科技(深圳)有限公司 | 一种音频数据处理方法,及设备 |
| CN105304092A (zh) * | 2015-09-18 | 2016-02-03 | 深圳市海派通讯科技有限公司 | 一种基于智能终端的实时变声方法 |
| US9820042B1 (en) | 2016-05-02 | 2017-11-14 | Knowles Electronics, Llc | Stereo separation and directional suppression with omni-directional microphones |
| CN106161218A (zh) * | 2016-09-28 | 2016-11-23 | 乐视控股(北京)有限公司 | 实时通话中的语音处理方法及装置 |
| CN106406809B (zh) * | 2016-12-21 | 2023-06-23 | 维沃移动通信有限公司 | 一种声音信号处理方法及移动终端 |
| CN107731241B (zh) * | 2017-09-29 | 2021-05-07 | 广州酷狗计算机科技有限公司 | 处理音频信号的方法、装置和存储介质 |
| CN111194545A (zh) * | 2017-10-09 | 2020-05-22 | 深圳传音通讯有限公司 | 一种移动通讯设备通话时改变原始声音的方法和系统 |
| CN108417223A (zh) * | 2017-12-29 | 2018-08-17 | 申子涵 | 在社交网络中发送变调语音的方法 |
| CN109404685B (zh) * | 2018-09-12 | 2022-04-08 | 乐歌人体工学科技股份有限公司 | 增高台 |
| US11943621B2 (en) * | 2018-12-11 | 2024-03-26 | Texas Instruments Incorporated | Secure localization in wireless networks |
| CN111339442A (zh) * | 2020-02-25 | 2020-06-26 | 北京声智科技有限公司 | 线上好友互动方法及装置 |
| KR102548618B1 (ko) * | 2021-01-25 | 2023-06-27 | 박상래 | 음성인식 및 음성합성을 이용한 무선통신장치 |
| US12469509B2 (en) * | 2023-04-04 | 2025-11-11 | Meta Platforms Technologies, Llc | Voice avatars in extended reality environments |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050043951A1 (en) * | 2002-07-09 | 2005-02-24 | Schurter Eugene Terry | Voice instant messaging system |
| CN1719514A (zh) * | 2004-07-06 | 2006-01-11 | 中国科学院自动化研究所 | 基于语音分析与合成的高品质实时变声方法 |
| CN1805478A (zh) * | 2005-01-14 | 2006-07-19 | 华为技术有限公司 | 一种实现通话中变声的系统及方法 |
| CN101304391A (zh) * | 2008-06-30 | 2008-11-12 | 腾讯科技(深圳)有限公司 | 一种基于即时通讯系统的语音通话方法及系统 |
Family Cites Families (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| TW430778B (en) * | 1998-06-15 | 2001-04-21 | Yamaha Corp | Voice converter with extraction and modification of attribute data |
| US7333507B2 (en) * | 2001-08-31 | 2008-02-19 | Philip Bravin | Multi modal communications system |
| JP2005115896A (ja) * | 2003-10-10 | 2005-04-28 | Nec Corp | 通信装置及び通信方法 |
| WO2005116992A1 (en) * | 2004-05-27 | 2005-12-08 | Koninklijke Philips Electronics N.V. | Method of and system for modifying messages |
| WO2006080149A1 (ja) * | 2005-01-25 | 2006-08-03 | Matsushita Electric Industrial Co., Ltd. | 音復元装置および音復元方法 |
| US20060257827A1 (en) * | 2005-05-12 | 2006-11-16 | Blinktwice, Llc | Method and apparatus to individualize content in an augmentative and alternative communication device |
| US20060116142A1 (en) * | 2006-02-07 | 2006-06-01 | Media Lab Europe (In Voluntary Liquidation) | Well Behaved SMS notifications |
| US7983910B2 (en) * | 2006-03-03 | 2011-07-19 | International Business Machines Corporation | Communicating across voice and text channels with emotion preservation |
| CN101046956A (zh) * | 2006-03-28 | 2007-10-03 | 国际商业机器公司 | 交互式音效产生方法及系统 |
| CN101175102B (zh) * | 2006-11-01 | 2013-01-09 | 鸿富锦精密工业(深圳)有限公司 | 具有音频调变功能的通讯装置及其音频调变的方法 |
| JP5275612B2 (ja) * | 2007-07-18 | 2013-08-28 | 国立大学法人 和歌山大学 | 周期信号処理方法、周期信号変換方法および周期信号処理装置ならびに周期信号の分析方法 |
| JP2009122776A (ja) * | 2007-11-12 | 2009-06-04 | Internatl Business Mach Corp <Ibm> | 仮想世界における情報制御方法および装置 |
| US8977779B2 (en) * | 2009-03-31 | 2015-03-10 | Mytalk Llc | Augmentative and alternative communication system with personalized user interface and content |
-
2008
- 2008-06-30 CN CNA2008100682626A patent/CN101304391A/zh active Pending
-
2009
- 2009-05-22 WO PCT/CN2009/071931 patent/WO2010000161A1/zh not_active Ceased
-
2010
- 2010-10-27 US US12/913,358 patent/US20110044324A1/en not_active Abandoned
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050043951A1 (en) * | 2002-07-09 | 2005-02-24 | Schurter Eugene Terry | Voice instant messaging system |
| CN1719514A (zh) * | 2004-07-06 | 2006-01-11 | 中国科学院自动化研究所 | 基于语音分析与合成的高品质实时变声方法 |
| CN1805478A (zh) * | 2005-01-14 | 2006-07-19 | 华为技术有限公司 | 一种实现通话中变声的系统及方法 |
| CN101304391A (zh) * | 2008-06-30 | 2008-11-12 | 腾讯科技(深圳)有限公司 | 一种基于即时通讯系统的语音通话方法及系统 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109272984A (zh) * | 2018-10-17 | 2019-01-25 | 百度在线网络技术(北京)有限公司 | 用于语音交互的方法和装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN101304391A (zh) | 2008-11-12 |
| US20110044324A1 (en) | 2011-02-24 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2010000161A1 (zh) | 一种基于即时通讯系统的语音通话方法及装置 | |
| US11605394B2 (en) | Speech signal cascade processing method, terminal, and computer-readable storage medium | |
| US10542136B2 (en) | Transcribing audio communication sessions | |
| JP3237566B2 (ja) | 通話方法、音声送信装置及び音声受信装置 | |
| US7529675B2 (en) | Conversational networking via transport, coding and control conversational protocols | |
| US6970935B1 (en) | Conversational networking via transport, coding and control conversational protocols | |
| CN1115917C (zh) | 用作因特网电话的增强型无线电电话及实现电话功能的方法 | |
| US20040267527A1 (en) | Voice-to-text reduction for real time IM/chat/SMS | |
| CN113571079B (zh) | 语音增强方法、装置、设备及存储介质 | |
| EP2245826A1 (en) | Method and apparatus for detecting and suppressing echo in packet networks | |
| WO2007070860A2 (en) | Intelligent codec selection to optimize audio transmission in wireless communications | |
| CN114842857B (zh) | 语音处理方法、装置、系统、设备及存储介质 | |
| Chinna Rao et al. | Real-time implementation and testing of VoIP vocoders with asterisk PBX using wireshark packet analyzer | |
| KR100465318B1 (ko) | 광대역 음성신호의 송수신 장치 및 그 송수신 방법 | |
| CN111225102A (zh) | 一种蓝牙音频信号传输方法和装置 | |
| US20110235632A1 (en) | Method And Apparatus For Performing High-Quality Speech Communication Across Voice Over Internet Protocol (VoIP) Communications Networks | |
| US8489216B2 (en) | Sound mixing apparatus and method and multipoint conference server | |
| FR2861247A1 (fr) | Terminal de telephonie a gestion de la qualite de restituton vocale pendant la reception | |
| JP4437011B2 (ja) | 音声符号化装置 | |
| CN111385780A (zh) | 一种蓝牙音频信号传输方法和装置 | |
| JP4120440B2 (ja) | 通信処理装置、および通信処理方法、並びにコンピュータ・プログラム | |
| CN101179600A (zh) | 一种网络协议语音通信方法、系统和服务器 | |
| JP2002101203A (ja) | 音声処理システム、音声処理方法およびその方法を記憶した記憶媒体 | |
| CN105407243B (zh) | 一种Android平台上使用改进仿射投影算法的回声消除VOIP系统 | |
| WO2019071541A1 (zh) | 语音翻译方法、装置和终端设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 09771931 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 7159/CHENP/2010 Country of ref document: IN |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 27/07/11) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 09771931 Country of ref document: EP Kind code of ref document: A1 |