WO2019119406A1 - 一种缩短呼叫建立时间的方法、装置及对讲机 - Google Patents

一种缩短呼叫建立时间的方法、装置及对讲机 Download PDF

Info

Publication number
WO2019119406A1
WO2019119406A1 PCT/CN2017/117957 CN2017117957W WO2019119406A1 WO 2019119406 A1 WO2019119406 A1 WO 2019119406A1 CN 2017117957 W CN2017117957 W CN 2017117957W WO 2019119406 A1 WO2019119406 A1 WO 2019119406A1
Authority
WO
WIPO (PCT)
Prior art keywords
voice
voices
speech
receiving device
inactive
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2017/117957
Other languages
English (en)
French (fr)
Inventor
黄妮
谢汉雄
郭飞
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hytera Communications Corp Ltd
Original Assignee
Hytera Communications Corp Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hytera Communications Corp Ltd filed Critical Hytera Communications Corp Ltd
Priority to PCT/CN2017/117957 priority Critical patent/WO2019119406A1/zh
Publication of WO2019119406A1 publication Critical patent/WO2019119406A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04BTRANSMISSION
    • H04B1/00Details of transmission systems, not covered by a single one of groups H04B3/00 - H04B13/00; Details of transmission systems not characterised by the medium used for transmission
    • H04B1/38Transceivers, i.e. devices in which transmitter and receiver form a structural unit and in which at least one part is used for functions of transmitting and receiving
    • H04B1/3827Portable transceivers

Definitions

  • the present invention relates to the field of communications, and more particularly to a method, apparatus, and walkie-talkie for shortening call setup time.
  • the walkie-talkie is a two-way mobile communication tool that can talk without any network support, no call charges, and is suitable for relatively fixed and frequent calls.
  • the voice input by the user can be collected only after the call setup is completed.
  • the process of call setup includes: after the processor in the walkie-talkie detects that the PTT button is pressed, the PTT button is identified, and after establishing a connection with the called party, the voice prompt is used to notify the user that the speech can be started. That is, when the walkie-talkie is used, after the voice prompt is played, the voice input by the user can be collected and sent to the called party.
  • the user generally starts to speak after pressing the PTT button, and the voice that the user said before the prompt tone is played will not be collected, thereby causing the user to lose the voice input before the prompt tone is played, affecting the conversation between the two parties. Integrity.
  • the present invention provides a method, a device and a walkie-talkie for shortening the call setup time, so as to solve the problem that the user usually starts speaking after pressing the PTT button, and the voice that the user said before the prompt tone is played will not be
  • the acquisition causes the user to lose the voice input before the prompt tone is played, which affects the integrity of the call between the two parties.
  • the present invention adopts the following technical solutions:
  • a method for shortening call setup time, applied to a walkie-talkie including:
  • the collected voice is sent to the receiving device.
  • the method further includes:
  • the sending the collected voice to the receiving device includes:
  • the voice is split into a plurality of active voices and a plurality of inactive voices, including:
  • the voice activity detection algorithm is used to split the voice into a plurality of the active voices and a plurality of the inactive voices.
  • sending the user input voice to the receiving device comprises:
  • the method further includes:
  • the sending the collected voice to the receiving device includes:
  • sending the user input voice to the receiving device comprises:
  • a device for shortening call setup time, applied to a walkie-talkie comprising:
  • a detecting unit configured to detect whether there is a touch point on the PTT button of the walkie-talkie; wherein the touch point is a contact point between a user's finger and the PTT button;
  • a determining unit configured to: when the detecting unit detects that the touch point exists on the PTT button, determine whether a time when the touch point exists is greater than a preset value;
  • a collecting unit configured to: when the determining unit determines that the time when the touch point exists is greater than the preset value, collecting a voice
  • a sending unit configured to send the collected voice to the receiving device after detecting that the PTT button is pressed, and the interphone establishes a connection with the receiving device.
  • the method further comprises:
  • a saving unit configured to save the voice after the collecting unit collects a voice
  • a splitting unit configured to split the voice into multiple active voices and multiple inactive voices
  • a voice discarding unit configured to discard each of the inactive voices by a preset proportion of a portion of the inactive voice to obtain a plurality of segment voices
  • a voice combining unit configured to combine a plurality of the active voices and a plurality of the segment voices according to a voice generation time to obtain a user input voice
  • the sending unit is configured to send the collected voice to the receiving device, specifically:
  • the splitting unit comprises:
  • the splitting unit is configured to split the voice into a plurality of the active voices and a plurality of the inactive voices by using a voice activity detection algorithm.
  • the method further comprises:
  • a voice encoding unit configured to: after the collecting unit collects a voice, perform voice encoding on the voice to obtain a coded processing voice;
  • a voice saving unit configured to save the encoded processing voice
  • a voice splitting unit configured to split the encoded processed voice into multiple active voices and multiple inactive voices based on voice activity detection results in a voice encoding process
  • a discarding unit configured to discard each of the inactive voices by a preset proportion of a portion of the inactive voice to obtain a plurality of segment voices
  • a combining unit configured to combine a plurality of the active voices and the plurality of the segment voices according to a voice generation time to obtain a user input voice
  • the sending unit sends the collected voice to the receiving device, it is specifically used to:
  • the sending unit comprises:
  • a coding unit configured to perform channel coding on the user input voice to obtain channel coded speech
  • a voice sending unit configured to send the channel coded voice to the receiving device.
  • a walkie-talkie including a memory and a processor
  • the memory is used to store a program
  • a processor is operative to run a program, wherein when the processor runs the program, it is used to:
  • the collected voice is sent to the receiving device.
  • the processor is used to collect voice, it is further used for:
  • the processor when configured to send the collected voice to the receiving device, specifically:
  • the processor when the processor is used to split the voice into multiple active voices and multiple inactive voices, specifically:
  • the voice activity detection algorithm is used to split the voice into a plurality of the active voices and a plurality of the inactive voices.
  • the processor when the processor is configured to send the user input voice to the receiving device, specifically:
  • the processor is used to collect voice, it is further used for:
  • the processor when configured to send the collected voice to the receiving device, specifically:
  • the processor when the processor is configured to send the user input voice to the receiving device, specifically:
  • the present invention has the following beneficial effects:
  • the invention provides a method, a device and a walkie-talkie for shortening the call setup time.
  • the PTT button When the PTT button is not pressed in the invention, when the contact time between the user's finger and the PTT button is greater than a preset value, the user is started to collect.
  • the input voice avoids the loss of voice input by the user before the prompt tone is played. It solves the problem that the user usually starts to speak after pressing the PTT button, and the voice that the user said before the prompt tone is played will not be collected, and the voice input by the user before the prompt tone is played is lost, which affects the conversation between the two parties. The issue of integrity.
  • each inactive voice is discarded by a preset proportion of part of the inactive voice, which can reduce the voice playing time of the called party, thereby reducing the time of the voice delay, that is, reducing the transmission delay.
  • FIG. 1 is a flowchart of a method for shortening a call setup time provided by the present invention
  • FIG. 2 is a flowchart of another method for shortening a call setup time provided by the present invention.
  • FIG. 3 is a flowchart of still another method for shortening a call setup time according to the present invention.
  • FIG. 4 is a schematic structural diagram of an apparatus for shortening call setup time according to the present invention.
  • FIG. 5 is a schematic structural diagram of another apparatus for shortening call setup time according to the present invention.
  • FIG. 6 is a schematic structural diagram of another apparatus for shortening call setup time according to the present invention.
  • the embodiment of the invention provides a method for shortening the call setup time, which is applied to the walkie-talkie.
  • the method for shortening the call setup time includes:
  • step S101 Check whether there is a touch point on the PTT button of the walkie-talkie; when it is detected that there is a touch point on the PTT button, step S102 is performed.
  • the touch point is the contact point between the user's finger and the PTT button. Detect whether there is a touch point on the PTT button of the walkie-talkie, that is, whether the user's finger is in contact with the PTT button.
  • the PTT button in this embodiment is a PTT button with a touch sensing function.
  • the mode used by the walkie-talkie is PTT mode, and the PTT mode refers to the half-duplex call.
  • the whole call process needs to press the PTT button to speak.
  • the user will first put the finger on the PTT button. After the action is determined, press the PTT button to start talking. It takes tens of milliseconds to put the finger on the PTT button and press the PTT button.
  • step S102 Determine whether the time when the touch point exists is greater than a preset value. When it is determined that the time when the touch point exists is greater than the preset value, step S103 is performed.
  • determining whether the time of the touch point exists is greater than a preset value, that is, determining whether the contact time between the user's finger and the PTT button is greater than a preset value before the user presses the PTT button, wherein the preset value may be a few seconds. Or ten or more seconds.
  • determining whether the time when the touch point exists is greater than a preset value is to prevent the user from accidentally touching the PTT button but not speaking.
  • the process of collecting voice includes: turning on the microphone of the walkie-talkie and the analog-to-digital converter ADC of the codec CODEC.
  • the microphone can receive the sound signal input by the user and convert the sound signal into an analog signal, and the analog converter converts the analog signal into a digital signal and saves the digital signal.
  • the process of establishing a connection between the walkie-talkie and the receiving device includes:
  • the call request information is generated and sent to the receiving device, and the confirmation connection information sent by the receiving device is received.
  • the receiving device After the call request information is sent to the receiving device, the receiving device sends the confirmation connection information to the walkie-talkie through the air interface in the next-slot time slot of the time slot in which the call request information is received.
  • the air interface is an interface between the base station and the mobile terminal in the mobile communication network.
  • the receiving device may be a base station or other intercom.
  • the base station may be a digital mobile radio DMR, a digital trunk PDT, a terrestrial trunked radio TETRA base station, and the walkie-talkie includes a hand platform or a vehicle platform.
  • the embodiment provides a method for shortening the call setup time.
  • the PTT button is not pressed in this embodiment, when the contact time between the user's finger and the PTT button is greater than a preset value, the voice is collected, and the user is avoided. Loss of voice input before the tone is played. It solves the problem that the user usually starts to speak after pressing the PTT button, and the voice that the user said before the prompt tone is played will not be collected, and the voice input by the user before the prompt tone is played is lost, which affects the conversation between the two parties. The issue of integrity.
  • step S103 the method further includes:
  • a buffer is set in the walkie-talkie to store the collected voice data, wherein the voice data of the buffer is stored in a first-in, first-out manner. That is, the voice data that is preferentially stored is preferentially transmitted to the receiving device.
  • the size of the buffer can be determined according to the usage. It mainly considers the allowable voice delay, the statistical length of the lost information, and the size of the space supported by the memory. In the worst case, DMR1:4 power saving is taken as an example. Carrier, plus PTT button software processing time of 100ms, about to open a space that can store 580ms of voice data, but this space will not be used, will remove some inactive voice and then store, 580ms is the largest one The most insured space.
  • the storage space can store 580 ms of voice data, if the received voice data exceeds 580 ms, the excess voice data will sequentially cover the first stored voice data.
  • the storage unit is generally in units of byte bytes. Therefore, when storing voice data, it is necessary to look at how much space 580ms of voice data needs to occupy, and to sample at 16k sampling rate of 8k, 100ms will occupy 800*2 bytes, then 580ms. Take up 800*2*5.8 bytes.
  • step S205 includes:
  • the voice activity detection algorithm is used to split the voice into multiple active voices and multiple inactive voices.
  • VAD Voice Activity Detection
  • voice endpoint detection and speech boundary detection The purpose is to identify and eliminate long silent periods from the sound signal stream, so as to save the channel resources without degrading the quality of the service, which can help reduce the end-to-end delay perceived by the user.
  • the active voice includes the voice spoken by the user, and the inactive voice is muted, that is, the inactive voice does not include the voice spoken by the user.
  • active voice cannot be discarded, and inactive voice can be discarded.
  • S206 Discard each inactive voice by a preset proportion of part of the inactive voice to obtain a plurality of segment voices;
  • the preset proportion of the part of the inactive voice is discarded, wherein the preset ratio may be 0.5, 0.8, or 1.
  • the preset ratio is 0.5, and then discarding The 0.5 of the inactive voice changes a 10 ms inactive voice to a 5 ms segment voice.
  • the voice generation time refers to the time when the user speaks the voice. Now, an example is shown in which a plurality of active voices and a plurality of clip voices are combined according to voice generation time to obtain a process in which a user inputs voice.
  • the voice is split into three active voices and one inactive voice, and the voice generation time sequence of the three active voices and one inactive voice is active voice 1, inactive voice, active voice 2, and active voice 3.
  • step S104 is changed to step S208:
  • sending the user input voice to the receiving device includes:
  • the user inputs voice into speech coding and channel coding to obtain encoded speech, and transmits the encoded speech to the receiving device.
  • Speech coding that is, vocoder coding, is based on a digital model generated by a speech signal, and analyzes digital speech to propose a set of characteristic parameters.
  • the characteristic parameters mainly refer to the excitation parameters that characterize the glottal vibration and the vocal parameters that characterize the characteristics of the vocal tract. These parameters carry the main information of the speech signal, and they need only a small number of bits to be encoded, and can be re-created by these parameters after decoding.
  • the synthesized speech signal can be encoded at a rate as low as 2.4 kbit/s or less.
  • Channel coding is an error correction and error detection coding of a digital signal to be transmitted in a channel. Specifically, in channel coding, a forward error correction FEC check may be added to the digital signal converted by the analog-to-digital converter.
  • discarding a portion of the inactive voice of the preset proportion of each inactive voice can reduce the voice playing time of the called party, thereby reducing the time of the voice delay, that is, reducing the transmission delay.
  • the method further includes:
  • the overall rate is 3.6Kbps, and the storage space occupied by 580ms of voice storage is 261Bytes.
  • Step S104 is changed to step S309:
  • sending the user input voice to the receiving device includes:
  • the channel encoded speech is transmitted to the receiving device.
  • the encoding processing voice is split into a plurality of active voices and a plurality of inactive voices, and the time when multiple active voices and multiple inactive voices are split in the previous embodiment is obtained. Differently, there are different processing methods when splitting to obtain multiple active voices and multiple inactive voices.
  • another embodiment of the present invention provides a device for shortening a call setup time, which is applied to a walkie-talkie.
  • the apparatus for shortening a call setup time includes:
  • the detecting unit 101 is configured to detect whether there is a touch point on the PTT button of the walkie-talkie; wherein the touch point is a contact point between the user's finger and the PTT button;
  • the determining unit 102 is configured to: when the detecting unit detects that there is a touch point on the PTT button, determine whether the time of the touch point exists is greater than a preset value;
  • the collecting unit 103 is configured to: when the determining unit 102 determines that the time when the touch point exists is greater than a preset value, collecting the voice;
  • the sending unit 104 is configured to send the collected voice to the receiving device after detecting that the PTT button is pressed and the intercom establishes a connection with the receiving device.
  • the sending unit sends the collected voice to the receiving device, specifically:
  • the user inputs voice into speech coding and channel coding to obtain coded speech;
  • the encoded speech is sent to the receiving device.
  • the embodiment provides a device for shortening the call setup time.
  • the PTT button when the PTT button is not pressed, when the contact time between the user's finger and the PTT button is greater than a preset value, the voice is collected, and the user is avoided. Loss of voice input before the tone is played. It solves the problem that the user usually starts to speak after pressing the PTT button, and the voice that the user said before the prompt tone is played will not be collected, and the voice input by the user before the prompt tone is played is lost, which affects the conversation between the two parties. The issue of integrity.
  • the method further includes:
  • the saving unit 105 is configured to save the voice after the collecting unit 103 collects the voice
  • the splitting unit 106 is configured to split the voice into multiple active voices and multiple inactive voices;
  • a voice discarding unit 107 configured to discard each inactive voice by a preset proportion of a part of the inactive voice, to obtain a plurality of segment voices;
  • a voice combining unit 108 configured to combine a plurality of active voices and a plurality of segment voices according to a voice generation time to obtain a user input voice
  • the sending unit 104 when the sending unit 104 is configured to send the collected voice to the receiving device, the sending unit 104 is specifically configured to: send the user input voice to the receiving device.
  • the splitting unit 106 includes:
  • the splitting unit is used to split the voice into multiple active voices and multiple inactive voices using a voice activity detection algorithm.
  • discarding a portion of the inactive voice of the preset proportion for each inactive voice can reduce the voice playing time of the called party, thereby reducing the time of the voice delay.
  • the device for shortening the call setup time further includes:
  • the voice encoding unit 109 is configured to: after the collecting unit 103 collects the voice, perform voice encoding on the voice to obtain an encoded processing voice;
  • the voice saving unit 110 is configured to save the encoded processing voice
  • the voice splitting unit 111 is configured to split the encoded processed voice into multiple active voices and multiple inactive voices based on the voice activity detection result in the voice encoding process;
  • a discarding unit 112 configured to discard each inactive voice by a preset proportion of a portion of the inactive voice, to obtain a plurality of segment voices;
  • the combining unit 113 is configured to combine the plurality of active voices and the plurality of segment voices according to the voice generation time to obtain a user input voice;
  • the sending unit 104 sends the collected voice to the receiving device, it is specifically used to:
  • the sending unit 104 includes:
  • a coding unit configured to perform channel coding on the input voice of the user, to obtain channel coded speech
  • a voice sending unit configured to send channel coded voice to the receiving device.
  • the encoding processing voice is split into a plurality of active voices and a plurality of inactive voices, and the time when multiple active voices and multiple inactive voices are split in the previous embodiment is obtained. Differently, there are different processing methods when splitting to obtain multiple active voices and multiple inactive voices.
  • another embodiment of the present invention provides a walkie-talkie, including a memory and a processor;
  • the memory is used to store a program
  • a processor is operative to run a program, wherein when the processor runs the program, it is used to:
  • the collected voice is sent to the receiving device.
  • the processor is used to collect voice, it is further used to:
  • the processor when configured to send the collected voice to the receiving device, specifically:
  • the voice activity detection algorithm is used to split the voice into a plurality of the active voices and a plurality of the inactive voices.
  • the processor when configured to send the user input voice to the receiving device, specifically:
  • the processor is used to collect voice, it is further used to:
  • the processor when configured to send the collected voice to the receiving device, specifically:
  • the processor when configured to send the user input voice to the receiving device, specifically:
  • the PTT button when the PTT button is not pressed, when the contact time between the user's finger and the PTT button is greater than a preset value, the voice is collected, and the loss of the voice input before the user finishes the prompt tone is avoided. It solves the problem that the user usually starts to speak after pressing the PTT button, and the voice that the user said before the prompt tone is played will not be collected, and the voice input by the user before the prompt tone is played is lost, which affects the conversation between the two parties. The issue of integrity.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Telephone Function (AREA)
  • Telephonic Communication Services (AREA)

Abstract

本申请提供了一种缩短呼叫建立时间的方法及对讲机,能够改善用户在呼叫建立后响完提示音才能传输语音的体验。本发明中PTT按键未被按下时,在用户的手指与所述PTT按键的接触时间大于预设数值时,就开始采集语音,避免了用户在提示音播放完之前输入的语音的丢失。此外,本发明中采集语音后,可以保存语音,或者将所述语音进行语音编码,得到编码处理语音后,保存所述编码处理语音。并将语音或者编码处理语音进行拆分,得到活动语音和非活动语音,将非活动语音按照预设比例进行丢弃,能够减少传输延时,在建立呼叫的过程中,相当于缩短呼叫建立时间至0,用户可以达到即按即说。

Description

一种缩短呼叫建立时间的方法、装置及对讲机 技术领域
本发明涉及通信领域,更具体的说,涉及一种缩短呼叫建立时间的方法、装置及对讲机。
背景技术
对讲机是一种双向移动通信工具,在不需要任何网络支持的情况下,就可以通话,没有话费产生,适用于相对固定且频繁通话的场合。
目前,在使用对讲机时,只有在呼叫建立完成之后,才可以对用户输入的语音进行采集。呼叫建立的过程包括:对讲机中的处理器检测PTT按键被按下后,进行PTT按键的识别处理,此后,在与被呼叫方建立连接后,用语音提示音通知用户可以开始讲话。即在使用对讲机时,在语音提示音播放完毕后,用户输入的语音才可以被采集、并发送到被呼叫方。
但是,用户一般在按下PTT按键后,就开始说话,则用户在提示音播放完之前说的语音就不会被采集,进而导致用户在提示音播放完之前输入的语音丢失,影响双方通话的完整性。
发明内容
有鉴于此,本发明提供一种缩短呼叫建立时间的方法、装置及对讲机,以解决用户一般在按下PTT按键后,就开始说话,则用户在提示音播放完之前说的语音就不会被采集,进而导致用户在提示音播放完之前输入的语音丢失,影响双方通话的完整性的问题。
为解决上述技术问题,本发明采用了如下技术方案:
一种缩短呼叫建立时间的方法,应用于对讲机,包括:
检测所述对讲机的PTT按键上是否存在触摸点;其中,所述触摸点为用户的手指与所述PTT按键的接触点;
当检测出所述PTT按键上存在所述触摸点,判断所述触摸点存在的时 间是否大于预设数值;
当判断出所述触摸点存在的时间大于所述预设数值时,采集语音;
当检测到所述PTT按键被按下、且所述对讲机与接收设备建立连接后,将采集的语音发送到所述接收设备。
优选地,所述采集语音后,还包括:
保存所述语音;
将所述语音拆分为多个活动语音和多个非活动语音;
将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
相应的,所述将采集的语音发送到所述接收设备,具体包括:
将所述用户输入语音发送到所述接收设备。
优选地,将所述语音拆分为多个活动语音和多个非活动语音,包括:
采用语音活动检测算法,将所述语音拆分为多个所述活动语音和多个所述非活动语音。
优选地,将所述用户输入语音发送到所述接收设备,包括:
将所述用户输入语音进行语音编码和信道编码,得到编码语音;
将所述编码语音发送到所述接收设备。
优选地,所述采集语音后,还包括:
将所述语音进行语音编码,得到编码处理语音;
保存所述编码处理语音;
基于语音编码过程中的语音活动检测结果,将所述编码处理语音拆分为多个活动语音和多个非活动语音;
将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
相应的,所述将采集的语音发送到所述接收设备,具体包括:
将所述用户输入语音发送到所述接收设备。
优选地,将所述用户输入语音发送到所述接收设备,包括:
将所述用户输入语音进行信道编码,得到信道编码语音;
将所述信道编码语音发送到所述接收设备。
一种缩短呼叫建立时间的装置,应用于对讲机,包括:
检测单元,用于检测所述对讲机的PTT按键上是否存在触摸点;其中,所述触摸点为用户的手指与所述PTT按键的接触点;
判断单元,用于当所述检测单元检测出所述PTT按键上存在所述触摸点,判断所述触摸点存在的时间是否大于预设数值;
采集单元,用于当所述判断单元判断出所述触摸点存在的时间大于所述预设数值,采集语音;
发送单元,用于当检测到所述PTT按键被按下、且所述对讲机与接收设备建立连接后,将采集的语音发送到所述接收设备。
优选地,还包括:
保存单元,用于所述采集单元采集语音后,保存所述语音;
拆分单元,用于将所述语音拆分为多个活动语音和多个非活动语音;
语音丢弃单元,用于将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
语音组合单元,用于将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
相应的,所述发送单元用于将采集的语音发送到所述接收设备时,具体用于:
将所述用户输入语音发送到所述接收设备。
优选地,所述拆分单元包括:
拆分子单元,用于采用语音活动检测算法,将所述语音拆分为多个所述活动语音和多个所述非活动语音。
优选地,还包括:
语音编码单元,用于所述采集单元采集语音后,将所述语音进行语音编码,得到编码处理语音;
语音保存单元,用于保存所述编码处理语音;
语音拆分单元,用于基于语音编码过程中的语音活动检测结果,将所述编码处理语音拆分为多个活动语音和多个非活动语音;
丢弃单元,用于将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
组合单元,用于将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
相应的,发送单元将采集的语音发送到所述接收设备时,具体用于:
将所述用户输入语音发送到所述接收设备。
优选地,所述发送单元包括:
编码单元,用于将所述用户输入语音进行信道编码,得到信道编码语音;
语音发送单元,用于将所述信道编码语音发送到所述接收设备。
一种对讲机,包括存储器和处理器;
其中,所述存储器用于存储程序;
处理器用于运行程序,其中,当所述处理器运行所述程序时用于:
检测所述对讲机的PTT按键上是否存在触摸点;其中,所述触摸点为用户的手指与所述PTT按键的接触点;
当检测出所述PTT按键上存在所述触摸点,判断所述触摸点存在的时间是否大于预设数值;
当判断出所述触摸点存在的时间大于所述预设数值,采集语音;
当检测到所述PTT按键被按下、且所述对讲机与接收设备建立连接后,将采集的语音发送到所述接收设备。
优选地,所述处理器用于采集语音后,还用于:
保存所述语音;
将所述语音拆分为多个活动语音和多个非活动语音;
将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
将多个所述活动语音和多个所述片段语音按照语音生成时间进行组 合,得到用户输入语音;
相应的,所述处理器用于将采集的语音发送到所述接收设备时,具体用于:
将所述用户输入语音发送到所述接收设备。
优选地,所述处理器用于将所述语音拆分为多个活动语音和多个非活动语音时,具体用于:
采用语音活动检测算法,将所述语音拆分为多个所述活动语音和多个所述非活动语音。
优选地,所述处理器用于将所述用户输入语音发送到所述接收设备时,具体用于:
将所述用户输入语音进行语音编码和信道编码,得到编码语音;
将所述编码语音发送到所述接收设备。
优选地,所述处理器用于采集语音后,还用于:
将所述语音进行语音编码,得到编码处理语音;
保存所述编码处理语音;
基于语音编码过程中的语音活动检测结果,将所述编码处理语音拆分为多个活动语音和多个非活动语音;
将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
相应的,所述处理器用于将采集的语音发送到所述接收设备时,具体用于:
将所述用户输入语音发送到所述接收设备。
优选地,所述处理器用于将所述用户输入语音发送到所述接收设备时,具体用于:
将所述用户输入语音进行信道编码,得到信道编码语音;
将所述信道编码语音发送到所述接收设备。
相较于现有技术,本发明具有以下有益效果:
本发明提供了一种缩短呼叫建立时间的方法、装置及对讲机,本发明中PTT按键未被按下时,在用户的手指与所述PTT按键的接触时间大于预设数值时,就开始采集用户输入的语音,避免了用户在提示音播放完之前输入的语音的丢失。解决了用户一般在按下PTT按键后,就开始说话,则用户在提示音播放完之前说的语音就不会被采集,进而导致用户在提示音播放完之前输入的语音丢失,影响双方通话的完整性的问题。
此外,本发明中将每个非活动语音丢弃预设比例的部分非活动语音,能够减少被呼叫方的语音播放时间,进而能够减少语音延迟的时间,即能够减少传输延时。
附图说明
为了更清楚地说明本发明实施例或现有技术中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本发明的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1为本发明提供的一种缩短呼叫建立时间的方法的方法流程图;
图2为本发明提供的另一种缩短呼叫建立时间的方法的方法流程图;
图3为本发明提供的又一种缩短呼叫建立时间的方法的方法流程图;
图4为本发明提供的一种缩短呼叫建立时间的装置的结构示意图;
图5为本发明提供的另一种缩短呼叫建立时间的装置的结构示意图;
图6为本发明提供的又一种缩短呼叫建立时间的装置的结构示意图。
具体实施方式
下面将结合本发明实施例中的附图,对本发明实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本发明一部分实施例,而不是全部的实施例。基于本发明中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本发明保护的 范围。
本发明实施例提供了一种缩短呼叫建立时间的方法,应用于对讲机,参照图1,缩短呼叫建立时间的方法包括:
S101、检测对讲机的PTT按键上是否存在触摸点;当检测出PTT按键上存在触摸点,执行步骤S102。
其中,触摸点为用户的手指与PTT按键的接触点。检测对讲机的PTT按键上是否存在触摸点,即检测用户的手指是否与PTT按键接触。本实施例中的PTT按键为具有触摸传感功能的PTT按键。
需要说明的是,对讲机使用的模式为PTT模式,PTT模式指半双工通话,整个通话过程均需要按下PTT按键才能说话,使用者一般在使用对讲机时,首先会有把手指放在PTT按键上的动作,确定好按键位置后,按下PTT按键开始说话,从将手指放在PTT按键上,到按下PTT按键,需要几十毫秒。
S102、判断触摸点存在的时间是否大于预设数值;当判断出触摸点存在的时间大于预设数值时,执行步骤S103。
具体的,判断触摸点存在的时间是否大于预设数值,即判断用户在未按下PTT按键之前,用户的手指与PTT按键的接触时间是否大于预设数值,其中,预设数值可以是几秒或者十几秒。
需要说明的是,判断触摸点存在的时间是否大于预设数值是为了防止出现用户不小心触碰到PTT按键但是不需要讲话的这种情况。
S103、采集语音。
具体的,采集语音的过程包括:打开对讲机的麦克风以及编译码器CODEC的模数转换器ADC。麦克风能够接收用户输入的声音信号,并将声音信号转换为模拟信号,模拟转换器将模拟信号转换为数字信号,并将数字信号进行保存。
S104、当检测到PTT按键被按下、且对讲机与接收设备建立连接后,将采集的语音发送到接收设备。
可选的,本发明的另一实施例中,对讲机与接收设备建立连接的过程包括:
生成并发送呼叫请求信息到接收设备,接收接收设备发送的确认连接 信息。
将呼叫请求信息到接收设备后,接收设备会通过空口在接收到呼叫请求信息的时隙的隔壁时隙发送确认连接信息到对讲机。其中,空口为移动通信网络中,基站与移动终端之间的接口。
其中,接收设备可以是基站或者其他对讲机。其中,基站可以是数字移动无线电DMR、数字集群PDT、陆上集群无线电TETRA基站,对讲机包括手台或者车台等。
本实施例提供了一种缩短呼叫建立时间的方法,本实施例中PTT按键未被按下时,在用户的手指与PTT按键的接触时间大于预设数值时,就开始采集语音,避免了用户在提示音播放完之前输入的语音的丢失。解决了用户一般在按下PTT按键后,就开始说话,则用户在提示音播放完之前说的语音就不会被采集,进而导致用户在提示音播放完之前输入的语音丢失,影响双方通话的完整性的问题。
可选的,本发明的另一实施例中,步骤S103后,还包括:
S204、保存语音;
其中,对讲机中开出一个缓冲区,用来存储采集的语音数据,其中,缓冲区的语音数据采用先入先出的方式存储。即优先存储的语音数据,优先被发送到接收设备中。
缓冲区的大小可以根据使用情况来定,主要是考虑容许的语音延迟,丢失信息的统计长度,以及内存支持的空间大小等,以最恶劣的情形DMR1:4省电为例,要发480ms预载波,再加上PTT按键软件处理时间100ms,大约要开一个能够存储580ms的语音数据的空间,但是这个空间并不会全部用上,会去掉一些非活动语音再存储,580ms是开的一个最大的最保险的空间。
需要说明的是,当存储空间能够存储580ms的语音数据时,若接收的语音数据超过580ms时,超出的语音数据会依次覆盖最先存储的语音数据。
举例来说,假设存储了600ms的语音数据,由于存储空间只能存储580ms的语音数据,此时,会将580-600ms之间的数据,覆盖了580ms中0-20ms之间的数据。
其中,存储单元一般以字节byte为单位,因而在存储语音数据时,需要看580ms的语音数据需要占用多大的空间,以8k采样率16byte采样,则100ms会占用800*2个bytes,则580ms占用800*2*5.8个bytes。
S205、将语音拆分为多个活动语音和多个非活动语音;
可选的,本发明的另一实施例中,步骤S205,包括:
采用语音活动检测算法,将语音拆分为多个活动语音和多个非活动语音。
语音活动检测(Voice Activity Detection,VAD)又称语音端点检测,语音边界检测。目的是从声音信号流里识别和消除长时间的静音期,以达到在不降低业务质量的情况下节省话路资源的作用,可以有利于减少用户感觉到的端到端的时延。
其中,活动语音中包含用户说话的语音,非活动语音为静音,即非活动语音中不包含用户说话的语音。
其中,活动语音是不能够丢弃的,非活动语音可以丢弃。
S206、将每个非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
具体的,对于与非活动语音,丢弃预设比例的部分非活动语音,其中,预设比例可以为0.5、0.8或者1,以一段10ms的非活动语音为例,预设比例为0.5,则丢弃该非活动语音的0.5,即将一段10ms的非活动语音更改为一段5ms的片段语音。
需要说明的是,经试验证明,1s的语音里偷走120ms的非活动语音,拼接剩下的语音,组成连续语音后传输,用户很难感知与原有语音的差别。
S207、将多个活动语音和多个片段语音按照语音生成时间进行组合,得到用户输入语音;
其中,语音生成时间是指用户说出该语音的时间。现举例说明将多个活动语音和多个片段语音按照语音生成时间进行组合,得到用户输入语音的过程。
假设将语音拆分为3个活动语音和1个非活动语音,3个活动语音和1个非活动语音的语音生成时间顺序为活动语音1、非活动语音、活动语音2和 活动语音3。
对非活动语音丢弃预设比例的部分非活动语音,即得到1个片段语音,再将活动语音和片段语音组合的过程中,将活动语音1放在第一个,将片段语音放在第二个,将活动语音2和活动语音3进行复制,依次放到片段语音之后。
相应的,步骤S104更改为步骤S208:
将用户输入语音发送到接收设备。
可选的,本发明的另一实施例中,将用户输入语音发送到接收设备,包括:
将用户输入语音进行语音编码和信道编码,得到编码语音,将编码语音发送到接收设备。
语音编码即声码器编码,是以语音信号产生的数字模型为基础,对数字语音进行分析,提出一组特征参数。特征参数主要是指表征声门振动的激励参数和表征声道特性的声道参数,这些参数携带有语音信号的主要信息,编码它们只需较少的比特数,在解码后可以由这些参数重新合成语音信号,其编码速率可低至2.4kbit/s及以下。
信道编码是对要在信道中传送的数字信号进行的纠、检错编码。具体的,在信道编码时,可以将模数转换器转换出的数字信号中添加前向纠错FEC校验。
本实施例中,将每个非活动语音丢弃预设比例的部分非活动语音,能够减少被呼叫方的语音播放时间,进而能够减少语音延迟的时间,即能够减少传输延时。
可选的,本发明的另一实施例中,参照图3,采集语音后,还包括:
S304、将语音进行语音编码,得到编码处理语音;
其中,语音编码的过程请参照上述实施例中的说明,在此不再赘述。
S305、保存编码处理语音;
以2.4Kbps带前向纠错编码FEC的声码器为例,总体速率为3.6Kbps,存储580ms的语音占用的存储空间为261Bytes。
S306、基于语音编码过程中的语音活动检测结果,将编码处理语音拆 分为多个活动语音和多个非活动语音;
其中,在语音编码的过程中,就能够对编码处理语音中包含的语音是活动语音还是非活动语音进行区分。
S307、将每个非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
S308、将多个活动语音和多个片段语音按照语音生成时间进行组合,得到用户输入语音;
步骤S104更改为步骤S309:
将用户输入语音发送到接收设备。
可选的,本发明的另一实施例中,将用户输入语音发送到接收设备,包括:
将用户输入语音进行信道编码,得到信道编码语音;
将信道编码语音发送到接收设备。
本实施例中,先进行语音编码后,在将编码处理语音拆分为多个活动语音和多个非活动语音,同上个实施例中拆分得到多个活动语音和多个非活动语音的时间不同,进而在拆分得到多个活动语音和多个非活动语音时,有不同的处理方法。
可选的,本发明的另一实施例中提供了一种缩短呼叫建立时间的装置,应用于对讲机,参照图4,缩短呼叫建立时间的装置包括:
检测单元101,用于检测对讲机的PTT按键上是否存在触摸点;其中,触摸点为用户的手指与PTT按键的接触点;
判断单元102,用于当检测单元检测出PTT按键上存在触摸点,判断触摸点存在的时间是否大于预设数值;
采集单元103,用于当判断单元102判断出触摸点存在的时间大于预设数值,采集语音;
发送单元104,用于当检测到PTT按键被按下、且对讲机与接收设备建立连接后,将采集的语音发送到接收设备。
可选的,本发明的另一实施例中,发送单元将采集的语音发送到接收设备时,具体用于:
将用户输入语音进行语音编码和信道编码,得到编码语音;
将编码语音发送到接收设备。
本实施例提供了一种缩短呼叫建立时间的装置,本实施例中PTT按键未被按下时,在用户的手指与PTT按键的接触时间大于预设数值时,就开始采集语音,避免了用户在提示音播放完之前输入的语音的丢失。解决了用户一般在按下PTT按键后,就开始说话,则用户在提示音播放完之前说的语音就不会被采集,进而导致用户在提示音播放完之前输入的语音丢失,影响双方通话的完整性的问题。
需要说明的是,本实施例中的各个单元的工作过程,请参照上述实施例中的说明,在此不再赘述。
可选的,本发明的另一实施例中,参照图5,还包括:
保存单元105,用于采集单元103采集语音后,保存语音;
拆分单元106,用于将语音拆分为多个活动语音和多个非活动语音;
语音丢弃单元107,用于将每个非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
语音组合单元108,用于将多个活动语音和多个片段语音按照语音生成时间进行组合,得到用户输入语音;
相应的,发送单元104用于将采集的语音发送到接收设备时,具体用于:将用户输入语音发送到接收设备。
可选的,本发明的另一实施例中,拆分单元106包括:
拆分子单元,用于采用语音活动检测算法,将语音拆分为多个活动语音和多个非活动语音。
本实施例中,将每个非活动语音丢弃预设比例的部分非活动语音,能够减少被呼叫方的语音播放时间,进而能够减少语音延迟的时间。
需要说明的是,本实施例中的各个单元的工作过程,请参照上述实施例中的说明,在此不再赘述。
可选的,本发明的另一实施例中,缩短呼叫建立时间的装置还包括:
语音编码单元109,用于采集单元103采集语音后,将语音进行语音编码,得到编码处理语音;
语音保存单元110,用于保存编码处理语音;
语音拆分单元111,用于基于语音编码过程中的语音活动检测结果,将编码处理语音拆分为多个活动语音和多个非活动语音;
丢弃单元112,用于将每个非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
组合单元113,用于将多个活动语音和多个片段语音按照语音生成时间进行组合,得到用户输入语音;
相应的,发送单元104将采集的语音发送到接收设备时,具体用于:
将用户输入语音发送到接收设备。
可选的,本发明的另一实施例中,发送单元104包括:
编码单元,用于将用户输入语音进行信道编码,得到信道编码语音;
语音发送单元,用于将信道编码语音发送到接收设备。
本实施例中,先进行语音编码后,在将编码处理语音拆分为多个活动语音和多个非活动语音,同上个实施例中拆分得到多个活动语音和多个非活动语音的时间不同,进而在拆分得到多个活动语音和多个非活动语音时,有不同的处理方法。
需要说明的是,本实施例中的各个单元的工作过程,请参照上述实施例中的说明,在此不再赘述。
可选的,本发明的另一实施例中提供了一种对讲机,包括存储器和处理器;
其中,所述存储器用于存储程序;
处理器用于运行程序,其中,当所述处理器运行所述程序时用于:
检测所述对讲机的PTT按键上是否存在触摸点;其中,所述触摸点为用户的手指与所述PTT按键的接触点;
当检测出所述PTT按键上存在所述触摸点,判断所述触摸点存在的时间是否大于预设数值;
当判断出所述触摸点存在的时间大于所述预设数值,采集语音;
当检测到所述PTT按键被按下、且所述对讲机与接收设备建立连接后,将采集的语音发送到所述接收设备。
在上述实施例的基础上,所述处理器用于采集语音后,还用于:
保存所述语音;
将所述语音拆分为多个活动语音和多个非活动语音;
将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
相应的,所述处理器用于将采集的语音发送到所述接收设备时,具体用于:
将所述用户输入语音发送到所述接收设备。
在上述实施例的基础上,所述处理器用于将所述语音拆分为多个活动语音和多个非活动语音时,具体用于:
采用语音活动检测算法,将所述语音拆分为多个所述活动语音和多个所述非活动语音。
在上述实施例的基础上,所述处理器用于将所述用户输入语音发送到所述接收设备时,具体用于:
将所述用户输入语音进行语音编码和信道编码,得到编码语音;
将所述编码语音发送到所述接收设备。
在上述实施例的基础上,所述处理器用于采集语音后,还用于:
将所述语音进行语音编码,得到编码处理语音;
保存所述编码处理语音;
基于语音编码过程中的语音活动检测结果,将所述编码处理语音拆分为多个活动语音和多个非活动语音;
将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
相应的,所述处理器用于将采集的语音发送到所述接收设备时,具体用于:
将所述用户输入语音发送到所述接收设备。
在上述实施例的基础上,所述处理器用于将所述用户输入语音发送到所述接收设备时,具体用于:
将所述用户输入语音进行信道编码,得到信道编码语音;
将所述信道编码语音发送到所述接收设备。
本实施例中PTT按键未被按下时,在用户的手指与PTT按键的接触时间大于预设数值时,就开始采集语音,避免了用户在提示音播放完之前输入的语音的丢失。解决了用户一般在按下PTT按键后,就开始说话,则用户在提示音播放完之前说的语音就不会被采集,进而导致用户在提示音播放完之前输入的语音丢失,影响双方通话的完整性的问题。
对所公开的实施例的上述说明,使本领域专业技术人员能够实现或使用本发明。对这些实施例的多种修改对本领域的专业技术人员来说将是显而易见的,本文中所定义的一般原理可以在不脱离本发明的精神或范围的情况下,在其它实施例中实现。因此,本发明将不会被限制于本文所示的这些实施例,而是要符合与本文所公开的原理和新颖特点相一致的最宽的范围。

Claims (17)

  1. 一种缩短呼叫建立时间的方法,其特征在于,应用于对讲机,包括:检测所述对讲机的PTT按键上是否存在触摸点;其中,所述触摸点为用户的手指与所述PTT按键的接触点;
    当检测出所述PTT按键上存在所述触摸点,判断所述触摸点存在的时间是否大于预设数值;
    当判断出所述触摸点存在的时间大于所述预设数值时,采集语音;
    当检测到所述PTT按键被按下、且所述对讲机与接收设备建立连接后,将采集的语音发送到所述接收设备。
  2. 根据权利要求1所述的方法,其特征在于,所述采集语音后,还包括:
    保存所述语音;
    将所述语音拆分为多个活动语音和多个非活动语音;
    将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
    将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
    相应的,所述将采集的语音发送到所述接收设备,具体包括:
    将所述用户输入语音发送到所述接收设备。
  3. 根据权利要求2所述的方法,其特征在于,将所述语音拆分为多个活动语音和多个非活动语音,包括:
    采用语音活动检测算法,将所述语音拆分为多个所述活动语音和多个所述非活动语音。
  4. 根据权利要求2所述的方法,其特征在于,将所述用户输入语音发送到所述接收设备,包括:
    将所述用户输入语音进行语音编码和信道编码,得到编码语音;
    将所述编码语音发送到所述接收设备。
  5. 根据权利要求1所述的方法,其特征在于,所述采集语音后,还包括:
    将所述语音进行语音编码,得到编码处理语音;
    保存所述编码处理语音;
    基于语音编码过程中的语音活动检测结果,将所述编码处理语音拆分为多个活动语音和多个非活动语音;
    将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
    将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
    相应的,所述将采集的语音发送到所述接收设备,具体包括:
    将所述用户输入语音发送到所述接收设备。
  6. 根据权利要求5所述的方法,其特征在于,将所述用户输入语音发送到所述接收设备,包括:
    将所述用户输入语音进行信道编码,得到信道编码语音;
    将所述信道编码语音发送到所述接收设备。
  7. 一种缩短呼叫建立时间的装置,其特征在于,应用于对讲机,包括:
    检测单元,用于检测所述对讲机的PTT按键上是否存在触摸点;其中,所述触摸点为用户的手指与所述PTT按键的接触点;
    判断单元,用于当所述检测单元检测出所述PTT按键上存在所述触摸点,判断所述触摸点存在的时间是否大于预设数值;
    采集单元,用于当所述判断单元判断出所述触摸点存在的时间大于所述预设数值,采集语音;
    发送单元,用于当检测到所述PTT按键被按下、且所述对讲机与接收设备建立连接后,将采集的语音发送到所述接收设备。
  8. 根据权利要求7所述的装置,其特征在于,还包括:
    保存单元,用于所述采集单元采集语音后,保存所述语音;
    拆分单元,用于将所述语音拆分为多个活动语音和多个非活动语音;
    语音丢弃单元,用于将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
    语音组合单元,用于将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
    相应的,所述发送单元用于将采集的语音发送到所述接收设备时,具体用于:
    将所述用户输入语音发送到所述接收设备。
  9. 根据权利要求8所述的方法装置,其特征在于,所述拆分单元包括:
    拆分子单元,用于采用语音活动检测算法,将所述语音拆分为多个所述活动语音和多个所述非活动语音。
  10. 根据权利要求7所述的方法装置,其特征在于,还包括:
    语音编码单元,用于所述采集单元采集语音后,将所述语音进行语音编码,得到编码处理语音;
    语音保存单元,用于保存所述编码处理语音;
    语音拆分单元,用于基于语音编码过程中的语音活动检测结果,将所述编码处理语音拆分为多个活动语音和多个非活动语音;
    丢弃单元,用于将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
    组合单元,用于将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
    相应的,发送单元将采集的语音发送到所述接收设备时,具体用于:
    将所述用户输入语音发送到所述接收设备。
  11. 根据权利要求10所述的方法装置,其特征在于,所述发送单元包括:
    编码单元,用于将所述用户输入语音进行信道编码,得到信道编码语音;
    语音发送单元,用于将所述信道编码语音发送到所述接收设备。
  12. 一种对讲机,其特征在于,包括存储器和处理器;
    其中,所述存储器用于存储程序;
    处理器用于运行程序,其中,当所述处理器运行所述程序时用于:
    检测所述对讲机的PTT按键上是否存在触摸点;其中,所述触摸点为 用户的手指与所述PTT按键的接触点;
    当检测出所述PTT按键上存在所述触摸点,判断所述触摸点存在的时间是否大于预设数值;
    当判断出所述触摸点存在的时间大于所述预设数值,采集语音;
    当检测到所述PTT按键被按下、且所述对讲机与接收设备建立连接后,将采集的语音发送到所述接收设备。
  13. 根据权利要求12所述的对讲机,其特征在于,所述处理器用于采集语音后,还用于:
    保存所述语音;
    将所述语音拆分为多个活动语音和多个非活动语音;
    将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
    将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
    相应的,所述处理器用于将采集的语音发送到所述接收设备时,具体用于:
    将所述用户输入语音发送到所述接收设备。
  14. 根据权利要求13所述的对讲机,其特征在于,所述处理器用于将所述语音拆分为多个活动语音和多个非活动语音时,具体用于:
    采用语音活动检测算法,将所述语音拆分为多个所述活动语音和多个所述非活动语音。
  15. 根据权利要求13所述的对讲机,其特征在于,所述处理器用于将所述用户输入语音发送到所述接收设备时,具体用于:
    将所述用户输入语音进行语音编码和信道编码,得到编码语音;
    将所述编码语音发送到所述接收设备。
  16. 根据权利要求12所述的对讲机,其特征在于,所述处理器用于采集语音后,还用于:
    将所述语音进行语音编码,得到编码处理语音;
    保存所述编码处理语音;
    基于语音编码过程中的语音活动检测结果,将所述编码处理语音拆分为多个活动语音和多个非活动语音;
    将每个所述非活动语音丢弃预设比例的部分非活动语音,得到多个片段语音;
    将多个所述活动语音和多个所述片段语音按照语音生成时间进行组合,得到用户输入语音;
    相应的,所述处理器用于将采集的语音发送到所述接收设备时,具体用于:
    将所述用户输入语音发送到所述接收设备。
  17. 根据权利要求16所述的对讲机,其特征在于,所述处理器用于将所述用户输入语音发送到所述接收设备时,具体用于:
    将所述用户输入语音进行信道编码,得到信道编码语音;
    将所述信道编码语音发送到所述接收设备。
PCT/CN2017/117957 2017-12-22 2017-12-22 一种缩短呼叫建立时间的方法、装置及对讲机 Ceased WO2019119406A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/CN2017/117957 WO2019119406A1 (zh) 2017-12-22 2017-12-22 一种缩短呼叫建立时间的方法、装置及对讲机

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2017/117957 WO2019119406A1 (zh) 2017-12-22 2017-12-22 一种缩短呼叫建立时间的方法、装置及对讲机

Publications (1)

Publication Number Publication Date
WO2019119406A1 true WO2019119406A1 (zh) 2019-06-27

Family

ID=66992886

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2017/117957 Ceased WO2019119406A1 (zh) 2017-12-22 2017-12-22 一种缩短呼叫建立时间的方法、装置及对讲机

Country Status (1)

Country Link
WO (1) WO2019119406A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7940900B2 (en) * 2006-12-04 2011-05-10 Hewlett-Packard Development Company, L.P. Communication control for device with telephonic functionality
CN104090652A (zh) * 2014-06-13 2014-10-08 北京搜狗科技发展有限公司 一种语音输入方法和装置
CN104796166A (zh) * 2014-01-16 2015-07-22 海能达通信股份有限公司 一种无线配件唤醒方法和无线配件
CN106487408A (zh) * 2015-08-25 2017-03-08 北京信威通信技术股份有限公司 手持通话装置和语音呼叫方法
CN107863981A (zh) * 2017-12-22 2018-03-30 海能达通信股份有限公司 一种缩短呼叫建立时间的方法及对讲机

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7940900B2 (en) * 2006-12-04 2011-05-10 Hewlett-Packard Development Company, L.P. Communication control for device with telephonic functionality
CN104796166A (zh) * 2014-01-16 2015-07-22 海能达通信股份有限公司 一种无线配件唤醒方法和无线配件
CN104090652A (zh) * 2014-06-13 2014-10-08 北京搜狗科技发展有限公司 一种语音输入方法和装置
CN106487408A (zh) * 2015-08-25 2017-03-08 北京信威通信技术股份有限公司 手持通话装置和语音呼叫方法
CN107863981A (zh) * 2017-12-22 2018-03-30 海能达通信股份有限公司 一种缩短呼叫建立时间的方法及对讲机

Similar Documents

Publication Publication Date Title
JP5701916B2 (ja) 電話での会話をテキストに書き起こすための方法及びシステム
US20070225049A1 (en) Voice controlled push to talk system
CN108141498B (zh) 一种翻译方法及终端
CN101496096B (zh) 话音及文本通信系统、方法及设备
CN108197572B (zh) 一种唇语识别方法和移动终端
CN107863981B (zh) 一种缩短呼叫建立时间的方法及对讲机
CN102984666B (zh) 一种通话过程中的通讯录语音信息处理方法及系统
CN100502571C (zh) 通信方法及系统
DK2551847T3 (en) A method of reducing power consumption calls for a mobile terminal and a mobile terminal
CN102149051B (zh) 一种网络对讲的实现方法及系统
TW201246899A (en) Handling a voice communication request
CN106847300B (zh) 一种语音数据处理方法及装置
CN105551491A (zh) 语音识别方法和设备
US20090170504A1 (en) Communication terminal, communication method, and communication program
US11581002B2 (en) Communication method, apparatus, and system for digital enhanced cordless telecommunications (DECT) base station
CN101179635A (zh) 对免提电话进行回声控制的装置、方法和系统
CN110351690B (zh) 一种智能语音系统及其语音处理方法
WO2019119406A1 (zh) 一种缩短呼叫建立时间的方法、装置及对讲机
CN111770231B (zh) 一种dect基站、手柄及通信系统
WO2021150647A1 (en) System and method for data analytics for communications in walkie-talkie network
CN112910508B (zh) 在esco链路上实现立体声通话的方法、装置及服务器
CN105682209A (zh) 一种降低移动终端通话功耗的方法及移动终端
JP2018185758A (ja) 音声対話システムおよび情報処理装置
JP2007274499A (ja) 携帯電話機
CN115623126A (zh) 语音通话方法、系统、装置、计算机设备和存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 17935660

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 17935660

Country of ref document: EP

Kind code of ref document: A1

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 03/11/2020)

122 Ep: pct application non-entry in european phase

Ref document number: 17935660

Country of ref document: EP

Kind code of ref document: A1