WO2024217470A1 - 语音录制方法、装置、电子设备及存储介质 - Google Patents
语音录制方法、装置、电子设备及存储介质 Download PDFInfo
- Publication number
- WO2024217470A1 WO2024217470A1 PCT/CN2024/088407 CN2024088407W WO2024217470A1 WO 2024217470 A1 WO2024217470 A1 WO 2024217470A1 CN 2024088407 W CN2024088407 W CN 2024088407W WO 2024217470 A1 WO2024217470 A1 WO 2024217470A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- voice
- speech
- recording
- user
- electronic device
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
- G10L13/033—Voice editing, e.g. manipulating the voice of the synthesiser
Definitions
- the present application belongs to the field of voice technology, and specifically relates to a voice recording method, device, electronic device and storage medium.
- chat software for voice chat.
- the user can first click the voice button in the chat box to trigger the electronic device to display a "press and speak” control. Then the user can long press the "press and speak” control and start speaking, so that the electronic device starts recording the voice and sends the recorded voice to the contact corresponding to the chat box after the recording is completed.
- the purpose of the embodiments of the present application is to provide a voice recording method, device, electronic device and storage medium, which can ensure the integrity of the voice recorded by the electronic device.
- an embodiment of the present application provides a voice recording method, the method comprising: recording a first voice when a voice recording control exists in a current interface; receiving a first input from a user, the first input being an input to the voice recording control; recording a second voice in response to the first input; and processing the first voice and the second voice to obtain a target voice.
- an embodiment of the present application provides a voice recording device, which includes: a recording module, a receiving module, and a processing module.
- the recording module is used to record a first voice when a voice recording control exists in the current interface.
- the receiving module is used to receive a first input from a user, where the first input is an input to the voice recording control.
- the recording module is also used to record a second voice in response to the first input received by the receiving module.
- the processing module is used to process the first voice and the second voice obtained by the recording module to obtain a target voice.
- an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
- an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
- an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the method described in the first aspect.
- an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.
- the electronic device when there is a voice recording control in the current interface, the electronic device records a first voice, and after receiving a first input to the voice recording control, records a second voice, thereby processing the first voice and the second voice to obtain a target voice.
- the electronic device can first record the first voice when detecting that there is a voice recording control in the current interface, and when actually receiving the user's input to the voice recording control, stop recording the first voice, and start recording the second voice, the electronic device can then process the first voice recorded before the user input and the second voice recorded after the user input to obtain the target voice. In this way, the problem of voice loss that is prone to occur when the electronic device starts recording voice only after the user triggers the input to the recording control is avoided, thereby ensuring the integrity of the voice recorded by the electronic device.
- FIG1 is a flow chart of a method for recording voice according to an embodiment of the present invention.
- FIG. 2(A) is a schematic diagram of an example interface for inputting a voice button provided in an embodiment of the present application
- FIG2(B) is a schematic diagram of an example interface for displaying a “press and speak” control on a conversation interface according to an embodiment of the present application
- FIG3 is a second flow chart of a voice recording method provided in an embodiment of the present application.
- FIG5 is a third flow chart of a voice recording method provided in an embodiment of the present application.
- FIG6 is a fourth flow chart of a voice recording method provided in an embodiment of the present application.
- FIG7 is a fifth flow chart of a voice recording method provided in an embodiment of the present application.
- FIG8 is a sixth flow chart of a voice recording method provided in an embodiment of the present application.
- FIG. 9 is a schematic diagram of an example interface for displaying prompt information on a conversation interface provided by an embodiment of the present application.
- FIG10 is a schematic diagram of a structure of a voice recording device provided in an embodiment of the present application.
- FIG11 is a second structural diagram of a voice recording device provided in an embodiment of the present application.
- FIG12 is a schematic diagram of a hardware structure of an electronic device provided in an embodiment of the present application.
- FIG. 13 is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.
- first, second, etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first”, “second”, etc. are generally of the same type, and do not limit the number of objects.
- the first object can be one
- “and/or” in the specification and claims indicates at least one of the connected objects, and the character “/" generally indicates that the objects connected before and after are in an "or” relationship.
- the voice recording method in the embodiment of the present application can be applied to the scenario of sending voice.
- the electronic device when the electronic device detects that there is a "press and speak” control in the current interface, the electronic device can first record the first voice; when the user makes a first input, for example, when the user clicks the "press and speak” control in the current interface, the electronic device stops recording the first voice and starts recording the second voice; then the electronic device can process the first voice recorded before the user input and the second voice recorded after the user input to obtain the target voice. In this way, the problem of missing voice is avoided, in which the electronic device starts recording voice only after the user triggers the input of the recording control, thereby ensuring the integrity of the voice recorded by the electronic device.
- the execution subject of the voice recording method provided in the embodiment of the present application may be a voice recording device, and the voice recording device may be an electronic device, or a functional module or entity in an electronic device.
- the technical solution provided in the embodiment of the present application is described below using an electronic device as an example.
- the present application embodiment provides a voice recording method
- Figure 1 shows a flow chart of a voice recording method provided by the present application embodiment, which can be applied to electronic devices.
- the voice recording method provided by the present application embodiment may include the following steps 201 to 204.
- Step 201 When a voice recording control exists in the current interface, the electronic device records a first voice.
- the above-mentioned current interface may be an interface in an application currently displayed on the electronic device
- the application may be an application with a voice recording function
- the interface may be a conversation interface in the application with a voice recording function
- the conversation interface may be a conversation interface between the electronic device user and other contacts.
- the above-mentioned application with a voice recording function may include at least one of the following: a communication application, a shopping application, a short video application, etc.
- the specific application can be determined according to actual use requirements, and the embodiment of the present application does not limit it.
- the above-mentioned voice recording control may be a control for starting voice recording in the above-mentioned application with a voice recording function.
- the conversation interface includes a voice control/voice button, and the voice control is used to turn on the voice input function.
- the user can click and input the voice control to trigger the electronic device to display the above-mentioned voice recording control and start recording the first voice.
- the mobile phone displays a conversation interface 10 with contact B.
- the user can click the voice button 11 included in the conversation interface 10 to trigger the mobile phone to display a voice recording control as shown in Figure 2(B) in the conversation interface 10, such as a "press and speak” control 12, so that the mobile phone starts recording the first voice.
- the first voice when there is a voice control in the above-mentioned current interface, can be recorded through a local recording application (for example, a recorder) of the electronic device.
- a local recording application for example, a recorder
- the electronic device may save the first voice in a local recording application or a file management application.
- step 201 may be specifically implemented through the following step 201a.
- Step 201a When a voice recording control exists in the current interface and target behavior information of the user is detected, the electronic device records a first voice.
- the above-mentioned user's target behavior information includes at least one of the following: the user's voice information, the user's line of sight information of looking at the screen, and the user's mouth opening and closing information.
- the above-mentioned user's voice information may include at least one of the following: sound volume information, timbre, pitch, sound frequency, etc.
- the above-mentioned line of sight information of the user looking at the screen may include at least one of the following: line of sight angle, line of sight direction, gaze point position, etc.
- the user's mouth opening and closing information may include at least one of the following: mouth opening and closing state information, mouth opening and closing degree, and mouth shape information of mouth opening and closing.
- the electronic device may detect the user's line of sight information through a front camera to determine whether there is line of sight information of the user looking at the screen.
- the electronic device may detect the user's voice information through a microphone to determine whether the user makes a sound.
- the electronic device may detect the user's mouth opening and closing information through a front camera to determine whether the user has a mouth opening and closing action.
- the electronic device can start recording the first voice.
- the electronic device can obtain the user's facial feature information through a front camera to determine whether the user in front of the electronic device is a user of the electronic device based on the facial feature information, so as to start recording the first voice when it is determined that the user in front of the electronic device is a user of the electronic device and the user's behavior satisfies at least one of the following: the user's eyes are looking at the screen, the user makes a sound, and the user opens and closes his mouth.
- the electronic device when it detects sound information, it can determine whether it is the sound information of the user using the electronic device based on the timbre, so as to record the first voice when it is determined that the sound information is the sound information of the user using the electronic device.
- the user's behavior information is further detected, so that when the user's target behavior information is detected, the first voice is recorded, so that the electronic device can more accurately determine whether the user has a need to record a voice, thereby determining whether to pre-record the first voice. In this way, the flexibility of the electronic device in recording voice is improved.
- Step 202 The electronic device receives a first input from a user.
- the first input is an input to a voice recording control.
- the first input is an input for recording voice to trigger the electronic device to start recording voice.
- the first input includes but is not limited to: a user touches the voice recording control through a touch device such as a finger or a stylus, or a specific gesture input by the user, or a click input.
- a touch device such as a finger or a stylus
- the specific input can be determined according to actual use requirements and is not limited in the embodiment of the present invention.
- the above-mentioned specific gesture can be any one of a single-click gesture, a sliding gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double-press gesture, and a double-click gesture.
- the above-mentioned first input can also be other feasible inputs, for example: when the above-mentioned current interface displays the voice recording control, the user's input on the physical buttons of the electronic device, or the voice command input by the user, to trigger the electronic device to start recording the second voice.
- Step 203 The electronic device records a second voice in response to the first input.
- the electronic device receives the first input during the process of recording the first voice, it stops recording the first voice and starts recording the second voice.
- the electronic device can record the second voice through the application corresponding to the above-mentioned current interface.
- the electronic device may save the second voice in an application corresponding to the above-mentioned current interface, or in a file management application.
- the user may click the voice recording control again during the process of the electronic device recording a second voice to trigger the electronic device to stop recording the second voice.
- the electronic device when the above-mentioned first input is a long press input of the voice recording control by the user, when the user starts the long press input, the electronic device stops recording the first voice and starts recording the second voice, and when the user stops the long press input, the electronic device stops recording the second voice.
- the user can input the "Press and Speak" control 12 in the conversation interface 10, for example: long press input, to trigger the mobile phone to stop recording the first voice and start recording the second voice.
- the user can stop the long press input of the "Press and Speak" control 12 to trigger the mobile phone to stop recording the second voice.
- Step 204 The electronic device processes the first voice and the second voice to obtain a target voice.
- the electronic device may directly concatenate the first voice before the second voice to obtain the target voice.
- the first voice recorded by the electronic device is "Are you free today?"
- the second voice recorded is "Let's go eat”.
- the electronic device can directly splice the first voice before the second voice to obtain the target voice "Are you free today? Let's go eat”.
- the electronic device may perform noise reduction processing on at least one of the first speech and the second speech, and concatenate the first speech and the second speech after the noise reduction processing to obtain the target speech.
- the noise reduction processing here refers to removing noise or noise in the speech.
- the first voice recorded by the electronic device is "The weather is nice today ###", and the second voice recorded is "Let's go out and play", it should be noted that the "#" here represents noise or interference; the electronic device can perform noise reduction processing on the first voice, and splice the first voice and the second voice after noise reduction processing to obtain the target voice "The weather is nice today, let's go out and play”.
- the electronic device can detect whether there is a silent part in the first voice, and In the case of a silent part, the silent part in the first voice is cut to obtain a processed first voice, and the processed first voice and the second voice are spliced to obtain a target voice. It should be noted that the silent part is a part where the user does not make any sound.
- the electronic device when it is detected that the first voice is silent from the 1st second to the 3rd second, the electronic device can cut the silent part from the 1st second to the 3rd second in the first voice to obtain a processed first voice with a duration of 3 seconds, and splice the processed first voice and the second voice to obtain the target voice.
- the electronic device may concatenate the voice within a preset time length in the first voice with the second voice to obtain a target voice, and the preset time length may be a preset time length before the electronic device receives the first input.
- the speech within three seconds before the first input is concatenated with the second speech to obtain the target speech.
- An embodiment of the present application provides a voice recording method, in which, when a voice recording control exists in the current interface, the electronic device records a first voice, and after receiving a first input to the voice recording control, records a second voice, thereby processing the first voice and the second voice to obtain a target voice.
- the electronic device can first record the first voice when detecting that a voice recording control exists in the current interface, and when actually receiving the user's input to the voice recording control, stop recording the first voice, and start recording the second voice, the electronic device can then process the first voice recorded before the user input and the second voice recorded after the user input to obtain the target voice. In this way, the problem of voice loss that is prone to occur when the electronic device starts recording voice only after the user triggers the input to the recording control is avoided, thereby ensuring the integrity of the voice recorded by the electronic device.
- the above step 204 in combination with Figure 1, as shown in Figure 5, can be specifically implemented by the following steps 204a and 204b, or, in combination with Figure 1, as shown in Figure 6, the above step 204 can be specifically implemented by the following steps 204a and 204c.
- Step 204a The electronic device compares the content of the first voice with the second voice.
- the electronic device may convert the contents of the first voice and the second voice into text, and then compare the contents of the first voice and the second voice to obtain the similarity between the first voice and the second voice. If the similarity is greater than a threshold, it is confirmed that the contents of the first voice and the second voice are the same; if the similarity is less than or equal to a threshold, it is confirmed that the contents of the first voice and the second voice are different.
- Step 204b When the contents of the first voice and the second voice are identical, the electronic device deletes the first voice and uses the second voice as the target voice.
- the electronic device can delete the first voice and use the second voice as the target voice.
- the electronic device can also delete the first voice and use the second voice as the target voice.
- the electronic device may also delete the second voice and use the first voice as the target voice.
- Step 204c When the contents of the first voice and the second voice are different, the electronic device concatenates the first voice and the second voice to obtain the target voice.
- the electronic device can concatenate the first voice and the second voice to obtain to the target voice.
- the electronic device after obtaining the first voice and the second voice, can process the first voice and the second voice according to the comparison result by comparing the contents of the first voice and the second voice. In this way, the flexibility of the electronic device in processing the recorded voice is improved.
- the electronic device may display a menu bar for the user to process the spliced voices.
- the above-mentioned menu bar may include at least one of the following: voice changing, voice acceleration, voice deceleration, adding background music, etc.
- the above step 204 in combination with Figure 1, as shown in Figure 7, can be specifically implemented by the following steps 204d and 204e, or, in combination with Figure 1, as shown in Figure 8, the above step 204 can be specifically implemented by the following steps 204d and 204f.
- Step 204d The electronic device performs splicing processing on the first voice and the second voice.
- Step 204e When the concatenated speech is in a semantically coherent state, the electronic device uses the concatenated speech as the target speech.
- the electronic device may use natural language processing (NLP) technology to perform a semantic comparison between the first voice and the second voice to determine whether the characters, words, and sentences in the first voice and the second voice are coherent.
- NLP natural language processing
- the electronic device can respectively obtain semantic feature information of the first voice and the second voice, and then obtain the semantic similarity between the first voice and the second voice based on the semantic feature information. If the similarity is greater than a threshold, it is confirmed that the first voice and the second voice are in a semantically coherent state; if the similarity is less than or equal to a threshold, it is confirmed that the first voice and the second voice are in a semantically incoherent state.
- the electronic device can use the spliced voice as the target voice.
- the electronic device can use the spliced voice as the target voice.
- Step 204f When the concatenated speech is in a semantically incoherent state, the electronic device uses the first speech or the second speech as the target speech.
- the electronic device can use the first voice: "Good morning, you” or the second voice: "Have you eaten?" as the target voice.
- the electronic device may compare the length and quality of the first voice and the second voice.
- the volume, importance of the speech content, etc. are compared to determine the final target speech based on the comparison results.
- the electronic device may use the first voice as the target voice.
- the electronic device may use the first voice as the target voice.
- the electronic device can use the second voice as the target voice.
- the electronic device after obtaining the first voice and the second voice, the electronic device pre-joins the first voice and the second voice, and detects the semantic coherence state of the joined voice, so that the electronic device can perform corresponding processing on the joined voice according to the semantic coherence state of the joined voice. In this way, the flexibility and diversity of the electronic device in processing recorded voices are improved.
- the voice recording method provided in the embodiment of the present application further includes the following steps 301 to 303, or the voice recording method provided in the embodiment of the present application further includes the following steps 301, 302 and 304.
- Step 301 When the first voice and the second voice are in a semantically incoherent state, the electronic device displays a prompt message in the current interface.
- the above prompt information is used to prompt whether to use the spliced speech as the target speech.
- the electronic device may display a window in the above-mentioned current interface, the window including the above-mentioned prompt information, as well as a first control and a second control, the first control being used to determine whether the spliced speech is used as the target speech, and the second control being used to determine whether the first speech or the second speech is used as the target speech.
- the prompt information may be “whether to use the concatenated voice as the sent voice”.
- a window 13 can be displayed in the conversation interface 10, and the window 13 includes a prompt message "Do you want to use the spliced voice as the sent voice", and includes a first control, such as a "Yes” control 14, and includes a second control, such as a "No" control 15.
- Step 302 The electronic device receives a second input of the user to the prompt information.
- the second input may be an input to the first control or the second control.
- the second input includes but is not limited to: a user touches the first control or the second control through a touch device such as a finger or a stylus, or a specific gesture input by the user, or a click input, or other feasible input.
- a touch device such as a finger or a stylus
- the specific input can be determined according to actual use requirements and is not limited in the embodiment of the present invention.
- the above-mentioned specific gesture can be any one of a single-click gesture, a sliding gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double-press gesture, and a double-click gesture.
- Step 303 The electronic device responds to the second input and, if the second input is confirmed to be yes, uses the concatenated speech as the target speech.
- the electronic device can use the spliced voice as the target voice.
- Step 304 In response to the second input, the electronic device uses the first voice or the second voice as the target voice if the second input is confirmed to be negative.
- the electronic device can use the first voice or the second voice as the target voice.
- the user can make a second input to the "yes" control 14 to trigger the mobile phone to use the spliced voice as the target voice, or the user can make a second input to the "no" control 15 to trigger the mobile phone to use the first voice or the second voice as the target voice.
- the electronic device may also delete the first voice and the second voice, and remind the user to re-record the voice.
- the electronic device when the electronic device determines that the spliced speech is in a semantically incoherent state, it displays a prompt message so that the user can flexibly select the speech to be finally sent, that is, the target speech, according to the prompt message. In this way, the diversity and flexibility of the electronic device in sending the recorded speech are improved.
- the electronic device after executing the above step 204d, when the first voice and the second voice are in a semantically incoherent state: the electronic device can execute the above step 204f; or, the electronic device can execute the above steps 301 to 303; or, the electronic device can execute the above steps 301, 302 and 304.
- the voice recording method provided in the embodiment of the present application may further include the following step 401.
- Step 401 When there is a voice recording control in the current interface and the current interface is closed during recording of a first voice, the electronic device deletes the recorded first voice.
- the electronic device may delete the recorded first voice when it is detected that there is no voice recording control in the above-mentioned current interface, for example, when the user triggers the electronic device to display a text input keyboard in the above-mentioned current interface.
- the electronic device may delete the recorded first voice when detecting that the conversation list interface is displayed.
- the electronic device may delete the recorded first voice when receiving an incoming call message, a video call, or a voice call.
- the electronic device detects whether the current interface is closed, so that when the current interface is closed, the electronic device can delete the recorded first voice. In this way, the flexibility of the electronic device in recording voice is improved.
- the voice recording method provided in the embodiment of the present application can be performed by a voice recording device.
- the voice recording device performing the voice recording method is taken as an example to illustrate the voice recording device provided in the embodiment of the present application.
- Fig. 10 shows a possible structural diagram of a voice recording device involved in an embodiment of the present application.
- the voice recording device 70 may include: a recording module 71 , a receiving module 72 and a processing module 73 .
- the recording module 71 is used to record a first voice when a voice recording control is present in the current interface.
- the receiving module 72 is used to receive a first input from a user, where the first input is an input to the voice recording control.
- the recording module 71 is also used to record a second voice in response to the first input received by the receiving module 72.
- the processing module 73 is used to process the recorded voice.
- the first speech and the second speech obtained by the control module 71 are processed to obtain the target speech.
- the embodiment of the present application provides a voice recording device.
- the voice recording device detects that there is a voice recording control in the current interface, it can first record the first voice, and when it actually receives the user's input to the voice recording control, it stops recording the first voice and starts recording the second voice. Then, the voice recording device can process the first voice recorded before the user input and the second voice recorded after the user input to obtain the target voice. In this way, the problem of voice missing that is easy to occur when the voice recording device starts recording voice only after the user triggers the input to the recording control is avoided, thereby ensuring the integrity of the voice recorded by the voice recording device.
- the above-mentioned recording module 71 is specifically used to record the first voice when there is a voice recording control in the current interface and the user's target behavior information is detected.
- the user's target behavior information includes at least one of the following: the user's voice information, the user's line of sight information of looking at the screen, and the user's mouth opening and closing information.
- the processing module 73 is specifically used to compare the content of the first voice and the second voice; when the content of the first voice is the same as that of the second voice, the first voice is deleted and the second voice is used as the target voice; or, when the content of the first voice is different from that of the second voice, the first voice and the second voice are concatenated to obtain the target voice.
- the processing module 73 is specifically used to splice the first speech and the second speech; when the spliced speech is semantically coherent, the spliced speech is used as the target speech; or, when the spliced speech is semantically incoherent, the first speech or the second speech is used as the target speech.
- the voice recording device 70 provided in the embodiment of the present application further includes: a deletion module 74.
- the deletion module 74 is used to delete the recorded first voice when there is a voice recording control in the current interface and the current interface is closed during the recording of the first voice.
- the voice recording device in the embodiment of the present application can be an electronic device or a component in the electronic device, such as an integrated circuit or a chip.
- the electronic device can be a terminal or other devices other than a terminal.
- the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR)/virtual reality (Virtual Reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (Ultra-Mobile Personal Computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc.
- NAS Network Attached Storage
- PC personal computer
- TV television
- teller machine a self-service machine
- the voice recording device in the embodiment of the present application may be a device having an operating system.
- the operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
- the voice recording device provided in the embodiment of the present application can implement each process implemented in the above method embodiment, and will not be described again here to avoid repetition.
- an embodiment of the present application also provides an electronic device 900, including a processor 901 and a memory 902, and the memory 902 stores a program or instruction that can be executed on the processor 901.
- the program or instruction is executed by the processor 901
- the various steps of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
- the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices described above. Sub-device.
- FIG. 13 is a schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.
- the electronic device 100 includes but is not limited to components such as a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110.
- components such as a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110.
- the electronic device 100 may also include a power source (such as a battery) for supplying power to each component, and the power source may be logically connected to the processor 110 through a power management system, so that the power management system can manage charging, discharging, and power consumption management.
- a power source such as a battery
- the electronic device structure shown in FIG13 does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange components differently, which will not be described in detail here.
- the input unit 104 is used to record the first voice when there is a voice recording control in the current interface.
- the user input unit 107 is used to receive a first input from a user, where the first input is an input to a voice recording control.
- the input unit 104 is further configured to record a second voice in response to a first input received by the user input unit 107 .
- the processor 110 is used to process the first speech and the second speech obtained by the input unit 104 to obtain the target speech.
- the embodiment of the present application provides an electronic device.
- the electronic device detects that there is a voice recording control in the current interface, it can first record a first voice, and when it actually receives the user's input to the voice recording control, it stops recording the first voice and starts recording the second voice. Then the electronic device can process the first voice recorded before the user input and the second voice recorded after the user input to obtain the target voice. In this way, the problem of voice loss that is easy to occur when the electronic device starts recording voice only after the user triggers the input to the recording control is avoided, thereby ensuring the integrity of the voice recorded by the electronic device.
- the input unit 104 is specifically used to record a first voice when there is a voice recording control in the current interface and the user's target behavior information is detected, and the user's target behavior information includes at least one of the following: the user's voice information, the user's line of sight information of looking at the screen, and the user's mouth opening and closing information.
- the processor 110 is specifically used to compare the content of the first voice with the second voice; when the content of the first voice is the same as that of the second voice, the first voice is deleted and the second voice is used as the target voice; or, it is specifically used to compare the content of the first voice with the second voice, and when the content of the first voice is different from that of the second voice, the first voice and the second voice are concatenated to obtain the target voice.
- the processor 110 is specifically used to splice the first speech and the second speech; when the spliced speech is in a semantically coherent state, the spliced speech is used as the target speech; or, it is specifically used to splice the first speech and the second speech, and when the spliced speech is in a semantically incoherent state, the first speech or the second speech is used as the target speech.
- the processor 110 is further configured to delete the recorded first voice when a voice recording control exists in the current interface and the current interface is closed during the recording of the first voice.
- the electronic device provided in the embodiment of the present application can implement each process implemented in the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here.
- the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042.
- the graphics processing unit 1041 is used for the video capture mode or the image capture mode.
- the image data of the static picture or video obtained by the image capture device (such as a camera) is processed.
- the display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc.
- the user input unit 107 includes a touch panel 1071 and at least one of other input devices 1072.
- the touch panel 1071 is also called a touch screen.
- the touch panel 1071 may include two parts: a touch detection device and a touch controller.
- Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
- the memory 109 can be used to store software programs and various data.
- the memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc.
- the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both volatile and non-volatile memories.
- the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
- the volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM).
- the memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
- An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored.
- a program or instruction is stored.
- the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
- the processor is the processor in the electronic device described in the above embodiment.
- the readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
- the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
- An embodiment of the present application provides a computer program product, which is stored in a storage medium.
- the program product is executed by at least one processor to implement the various processes of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
- the technical solution of the present application can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM/RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
- a storage medium such as ROM/RAM, a disk, or an optical disk
- a terminal which can be a mobile phone, a computer, a server, or a network device, etc.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- User Interface Of Digital Computer (AREA)
- Telephone Function (AREA)
Abstract
本申请公开了一种语音录制方法、装置、电子设备及存储介质,属于语音技术领域。该方法包括:在当前界面中存在语音录制控件的情况下,录制第一语音;接收用户的第一输入,该第一输入为对语音录制控件的输入;响应于第一输入,录制第二语音;对第一语音和第二语音处理,得到目标语音。
Description
相关申请的交叉引用
本申请主张在2023年04月20日在中国提交的申请号为202310433734.8的中国专利的优先权,其全部内容通过引用包含于此。
本申请属于语音技术领域,具体涉及一种语音录制方法、装置、电子设备及存储介质。
随着电子设备的不断发展,电子设备的功能和应用也越来越丰富,用户会经常使用聊天软件进行语音聊天,例如,在“通讯A”应用的某个聊天框中,用户可以先点击该聊天框中的语音按钮,以触发电子设备显示“按住说话”控件,然后用户可以长按“按住说话”控件并开始说话,从而电子设备开始录制语音,并在录制完成后向该聊天框对应的联系人发送录制的语音。
然而,在上述过程中,经常出现用户已经开始说话,但是没有及时按住语音录制控件(即上述的“按住说话”控件)进行录音的情况,如此,造成电子设备录制的语音缺失。
发明内容
本申请实施例的目的是提供一种语音录制方法、装置、电子设备及存储介质,能够保证电子设备录制语音的完整性。
第一方面,本申请实施例提供了一种语音录制方法,该方法包括:在当前界面中存在语音录制控件的情况下,录制第一语音;接收用户的第一输入,该第一输入为对语音录制控件的输入;响应于第一输入,录制第二语音;对第一语音和第二语音处理,得到目标语音。
第二方面,本申请实施例提供了一种语音录制装置,该语音录制装置包括:录制模块、接收模块和处理模块。录制模块,用于在当前界面中存在语音录制控件的情况下,录制第一语音。接收模块,用于接收用户的第一输入,该第一输入为对语音录制控件的输入。录制模块,还用于响应于接收模块接收的第一输入,录制第二语音。处理模块,用于对录制模块得到的第一语音和第二语音处理,得到目标语音。
第三方面,本申请实施例提供了一种电子设备,该电子设备包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如第一方面所述的方法的步骤。
第四方面,本申请实施例提供了一种可读存储介质,所述可读存储介质上存储程序或指令,所述程序或指令被处理器执行时实现如第一方面所述的方法的步骤。
第五方面,本申请实施例提供了一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如第一方面所述的方法。
第六方面,本申请实施例提供一种计算机程序产品,该程序产品被存储在存储介质中,该程序产品被至少一个处理器执行以实现如第一方面所述的方法。
在本申请实施例中,在当前界面中存在语音录制控件的情况下,电子设备录制第一语音,并在接收到对语音录制控件的第一输入后,录制第二语音,从而对第一语音和第二语音处理,得到目标语音。本方案中,由于电子设备在检测到当前界面中存在语音录制控件的情况下,可以先录制第一语音,并在实际接收到用户对语音录制控件的输入时,停止录制第一语音,并开始录制第二语音,然后电子设备可以将用户输入之前录制的第一语音和用户输入之后录制的第二语音进行处理,以得到目标语音。如此,避免了只有在用户触发对录制控件的输入后电子设备才开始录制语音,容易产生语音缺失的问题,从而保证了电子设备录制语音的完整性。
图1是本申请实施例提供的一种语音录制方法的流程示意图之一;
图2(A)是本申请实施例提供的一种对语音按钮进行输入的界面实例示意图;
图2(B)是本申请实施例提供的一种在会话界面显示“按住说话”控件的界面实例示意图;
图3是本申请实施例提供的一种语音录制方法的流程示意图之二;
图4是本申请实施例提供的一种在会话界面对“按住说话”控件进行输入的界面实例示意图;
图5是本申请实施例提供的一种语音录制方法的流程示意图之三;
图6是本申请实施例提供的一种语音录制方法的流程示意图之四;
图7是本申请实施例提供的一种语音录制方法的流程示意图之五;
图8是本申请实施例提供的一种语音录制方法的流程示意图之六;
图9是本申请实施例提供的一种在会话界面显示提示信息的界面实例示意图;
图10是本申请实施例提供的一种语音录制装置的结构示意图之一;
图11是本申请实施例提供的一种语音录制装置的结构示意图之二;
图12是本申请实施例提供的一种电子设备的硬件结构示意图之一;
图13是本申请实施例提供的一种电子设备的硬件结构示意图之二。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员获得的所有其他实施例,都属于本申请保护的范围。
本申请的说明书和权利要求书中的术语“第一”、“第二”等是用于区别类似的对象,而不用于描述特定的顺序或先后次序。应该理解这样使用的术语在适当情况下可以互换,以便本申请的实施例能够以除了在这里图示或描述的那些以外的顺序实施,且“第一”、“第二”等所区分的对象通常为一类,并不限定对象的个数,例如第一对象可以是一个,
也可以是多个。此外,说明书以及权利要求中“和/或”表示所连接对象的至少其中之一,字符“/”,一般表示前后关联对象是一种“或”的关系。
下面结合附图,通过具体的实施例及其应用场景对本申请实施例提供的语音录制方法进行详细地说明。
本申请实施例中的语音录制方法可以应用于发送语音的场景。
目前,在用户使用“通讯A”应用与联系人B聊天时,若用户想要发送语音,则需要先点击聊天框中的语音按钮,以触发电子设备显示“按住说话”控件,然后用户可以长按该“按住说话”控件并开始说话,从而电子设备开始录制语音,并在录制完成后向联系人B发送录制的语音。然而,在上述过程中,经常出现用户已经开始说话,但是没有及时按住语音录音控件(即上述的“按住说话”控件)进行录音的情况,如此,造成电子设备录制的语音缺失。
在本申请实施例提供的方案中,在电子设备检测到当前界面中存在“按住说话”控件的情况下,电子设备可以先录制第一语音;在用户进行第一输入,例如:用户点击当前界面中的“按住说话”控件时,电子设备停止录制第一语音,并开始录制第二语音;然后电子设备可以将用户输入之前录制的第一语音和用户输入之后录制的第二语音进行处理,以得到目标语音。如此,避免了只有在用户触发对录制控件的输入后电子设备才开始录制语音,容易产生语音缺失的问题,从而保证了电子设备录制语音的完整性。
本申请实施例提供的语音录制方法的执行主体可以为语音录制装置,该语音录制装置可以为电子设备,或电子设备中的功能模块或实体。以下以电子设备为例,对本申请实施例提供的技术方案进行说明。
本申请实施例提供一种语音录制方法,图1示出了本申请实施例提供的一种语音录制方法的流程图,该方法可以应用于电子设备。如图1所示,本申请实施例提供的语音录制方法可以包括下述的步骤201至步骤204。
步骤201、在当前界面中存在语音录制控件的情况下,电子设备录制第一语音。
可选地,本申请实施例中,上述的当前界面可以为电子设备当前显示的某个应用中的界面,该应用可以为具有语音录制功能的应用,该界面可以为具有语音录制功能的应用中的会话界面,该会话界面为电子设备用户与其他联系人的会话界面。
可选地,本申请实施例中,上述的具有语音录制功能的应用可以包括以下至少一项:通讯类应用、购物类应用、短视频类应用等。具体的可以根据实际使用需求确定,本申请实施例不作限制。
可选地,本申请实施例中,上述语音录制控件可以为上述的具有语音录制功能的应用中用于开始录制语音的控件。
可选地,本申请实施例中,电子设备在显示上述应用中的一个会话界面后,该会话界面中包括语音控件/语音按钮,该语音控件用于开启语音输入功能,用户可以对该语音控件进行点击输入,以触发电子设备显示上述的语音录制控件,开始录制第一语音。
举例说明,以电子设备为手机为例进行说明,如图2(A)所示,手机显示与联系人B的会话界面10,用户可以点击该会话界面10中包括的语音按钮11,以触发手机在会话界面10中显示如图2(B)所示的语音录制控件,例如“按住说话”控件12,从而手机开始录制第一语音。
可选地,本申请实施例中,在上述的当前界面中存在语音控件的情况下,可以通过电子设备的本地录音应用(例如:录音机)录制第一语音。
可选地,本申请实施例中,电子设备可以将第一语音保存在本地录音应用中,或者文件管理应用中。
可选地,本申请实施例中,结合图1,如图3所示,上述步骤201具体可以通过下述的步骤201a实现。
步骤201a、在当前界面中存在语音录制控件,且检测到用户的目标行为信息的情况下,电子设备录制第一语音。
本申请实施例中,上述用户的目标行为信息包括以下至少一项:用户的声音信息,用户注视屏幕的视线信息,用户的嘴巴张合信息。
可选地,本申请实施例中,上述用户的声音信息可以包括以下至少一项:声音大小信息、音色、音调、声音频率等。
可选地,本申请实施例中,上述用户注视屏幕的视线信息可以包括以下至少一项:视线角度、视线方向、注视点位置等。
可选地,本申请实施例中,上述用户的嘴巴张合信息可以包括以下至少一项:嘴巴的张合状态信息、嘴巴的张合程度、嘴巴张合的口型信息。
可选地,本申请实施例中,电子设备可以通过前置摄像头来检测用户的视线信息,以确定是否存在用户注视屏幕的视线信息。
可选地,本申请实施例中,电子设备可以通过麦克风来检测用户的声音信息,以确定用户是否发出声音。
可选地,本申请实施例中,电子设备可以通过前置摄像头来检测用户的嘴巴张合信息,以确定用户是否有嘴巴张合动作。
可选地,本申请实施例中,在用户的行为满足以下至少之一:用户的眼睛注视屏幕、用户发出声音、用户有嘴巴张合动作时,电子设备可以开始录制第一语音。
可选地,本申请实施例中,电子设备可以通过前置摄像头获取用户的面部特征信息,以根据该面部特征信息判断电子设备前的用户是否为使用该电子设备的用户,以在确定电子设备前的用户为使用该电子设备的用户,且该用户的行为满足以下至少之一:用户的眼睛注视屏幕、用户发出声音、用户有嘴巴张合动作时,开始录制第一语音。
可选地,本申请实施例中,电子设备在检测到声音信息时,可以根据音色判断是否为使用该电子设备的用户的声音信息,以在确定声音信息为使用该电子设备的用户的声音信息时,录制第一语音。
本申请实施例中,通过在上述的当前界面中存在语音录制控件时,进一步检测用户的行为信息,以在检测到用户的目标行为信息时,录制第一语音,使得电子设备可以更精准地判断出用户是否有录制语音的需求,从而确定是否预先录制第一语音。如此,提高了电子设备录制语音的灵活性。
步骤202、电子设备接收用户的第一输入。
本申请实施例中,上述第一输入为对语音录制控件的输入。
可选地,本申请实施例中,上述第一输入为用于录制语音的输入,以触发电子设备开始录制语音。
可选地,本申请实施例中,上述第一输入包括但不限于:用户通过手指或者手写笔等触控装置对语音录制控件进行触控输入,或者为用户输入的特定手势,或者为点击输入。具体的可以根据实际使用需求确定,本发明实施例不作限定。
可选地,本申请实施例中,上述特定手势可以为单击手势、滑动手势、拖动手势、压力识别手势、长按手势、面积变化手势、双按手势、双击手势中的任意一种。
可选地,本申请实施例中,上述第一输入还可以为其他可行性输入,例如:在上述的当前界面显示语音录制控件的情况下,用户对电子设备的物理按键的输入,或者为用户输入的语音指令,以触发电子设备开始录制第二语音。
步骤203、电子设备响应于第一输入,录制第二语音。
可以理解,电子设备在录制第一语音的过程中,若接收到第一输入,则停止录制第一语音,并开始录制第二语音。
可选地,本申请实施例中,电子设备可以通过上述的当前界面对应的应用录制第二语音。
可选地,本申请实施例中,电子设备可以将第二语音保存在上述的当前界面对应的应用中,或者文件管理应用中。
可选地,本申请实施例中,在上述第一输入为用户对语音录制控件的点击输入的情况下,用户可以在电子设备录制第二语音的过程中,再次点击语音录制控件,以触发电子设备停止录制第二语音。
可选地,本申请实施例中,在上述第一输入为用户对语音录制控件的长按输入的情况下,用户开始长按输入时,电子设备停止录制第一语音,并开始录制第二语音,在用户停止长按输入时,电子设备停止录制第二语音。
举例说明,结合图2(B),如图4所示,用户可以对会话界面10中的“按住说话”控件12进行输入,例如:长按输入,以触发手机停止录制第一语音,并开始录制第二语音,在手机录制第二语音的过程中,用户可以停止对“按住说话”控件12的长按输入,以触发手机停止录制第二语音。
步骤204、电子设备对第一语音和第二语音处理,得到目标语音。
可选地,本申请实施例中,电子设备可以直接将第一语音拼接在第二语音之前,以得到目标语音。
示例性地,电子设备录制的第一语音为“你今天有空吗?”,且录制的第二语音为“我们去吃饭吧”,电子设备可以直接将第一语音拼接在第二语音之前,得到目标语音“你今天有空吗?我们去吃饭吧”。
可选地,本申请实施例中,电子设备可以对第一语音和第二语音中的至少一个进行降噪处理,并对降噪处理后的第一语音和第二语音进行拼接,以得到目标语音。此处的降噪处理是指将语音中的噪音或杂音进行删除。
示例性地,电子设备录制的第一语音为“今天###天气不错”,且录制的第二语音为“我们出去玩吧”,需要说明的是,此处的“#”表示噪音或杂音;电子设备可以对第一语音进行降噪处理,并对降噪处理后的第一语音和第二语音进行拼接,以得到目标语音“今天天气不错,我们出去玩吧”。
可选地,本申请实施例中,电子设备可以检测第一语音中是否有静音部分,并在含有
静音部分的情况下,将第一语音中的静音部分进行裁切,以得到处理后的第一语音,并对处理后的第一语音和第二语音进行拼接处理,得到目标语音。需要说明的是,上述静音部分为用户未发出声音的部分。
示例性地,假设第一语音的时长为6秒,在检测到第一语音的第1秒至第3秒为静音的情况下,电子设备可以将第一语音中第1秒至第3秒的静音部分进行裁切,以得到处理后的时长为3秒的第一语音,并对处理后的第一语音和第二语音进行拼接处理,得到目标语音。
可选地,本申请实施例中,电子设备可以将第一语音中预设时长内的语音与第二语音进行拼接处理,以得到目标语音,该预设时长可以为电子设备接收到第一输入之前的预设时长。
例如:将第一输入之前的三秒内的语音与第二语音进行拼接处理,以得到目标语音。
本申请实施例提供一种语音录制方法,在当前界面中存在语音录制控件的情况下,电子设备录制第一语音,并在接收到对语音录制控件的第一输入后,录制第二语音,从而对第一语音和第二语音处理,得到目标语音。本方案中,由于电子设备在检测到当前界面中存在语音录制控件的情况下,可以先录制第一语音,并在实际接收到用户对语音录制控件的输入时,停止录制第一语音,并开始录制第二语音,然后电子设备可以将用户输入之前录制的第一语音和用户输入之后录制的第二语音进行处理,以得到目标语音。如此,避免了只有在用户触发对录制控件的输入后电子设备才开始录制语音,容易产生语音缺失的问题,从而保证了电子设备录制语音的完整性。
可选地,本申请实施例中,结合图1,如图5所示,上述步骤204具体可以通过下述的步骤204a和步骤204b实现,或者,结合图1,如图6所示,上述步骤204具体可以通过下述的步骤204a和步骤204c实现。
步骤204a、电子设备将第一语音与第二语音进行内容比对。
可选地,本申请实施例中,电子设备可以将第一语音和第二语音的内容均转为文字后,再对比第一语音和第二语音的内容,以获得第一语音和第二语音的相似度。若相似度大于一个阈值,则确认第一语音的内容和第二语音的内容相同;若相似度小于或等于一个阈值,则确认第一语音的内容和第二语音的内容不同。
步骤204b、在第一语音与第二语音的内容相同的情况下,电子设备删除第一语音,并将第二语音作为目标语音。
示例性地,在第一语音的内容为“吃饭了吗”,第二语音的内容同样为“吃饭了吗”的情况下,电子设备可以删除第一语音,并将第二语音作为目标语音。
示例性地,在第一语音的内容为“我们一会儿去商场吧”,第二语音的内容为“一会儿我们去商场吧”的情况下,电子设备同样可以删除第一语音,将第二语音作为目标语音。
可选地,本申请实施例中,在第一语音与第二语音的内容相同时,电子设备也可以删除第二语音,并将第一语音作为目标语音。
步骤204c、在第一语音与第二语音的内容不同的情况下,电子设备对第一语音和第二语音拼接处理,得到目标语音。
示例性地,第一语音的内容为“早上好”,第二语音的内容为“吃饭了吗”,即第一语音与第二语音的内容不同时,电子设备可以将第一语音和第二语音进行拼接处理,以得
到目标语音。
需要说明的是,针对此处对第一语音和第二语音拼接处理的具体说明,可以参见上述实施例的步骤204中对拼接处理的描述,即与上述步骤204中的拼接处理的具体方案相同,此处不再赘述。
本申请实施例中,电子设备在得到第一语音和第二语音之后,通过比对第一语音和第二语音的内容,使得电子设备可以根据比对的结果对第一语音和第二语音进行处理。如此,提高了电子设备处理录制语音的灵活性。
可选地,本申请实施例中,在电子设备对第一语音和第二语音进行拼接处理之后,电子设备可以显示菜单栏,该菜单栏用于用户对拼接后的语音进行处理。
可选地,本申请实施例中,上述菜单栏可以包括以下至少一项:变声、语音加速、语音减速、添加背景音乐等。
可选地,本申请实施例中,结合图1,如图7所示,上述步骤204具体可以通过下述的步骤204d和步骤204e实现,或者,结合图1,如图8所示,上述步骤204具体可以通过下述的步骤204d和步骤204f实现。
步骤204d、电子设备对第一语音和第二语音拼接处理。
需要说明的是,针对此处对第一语音和第二语音拼接处理的具体说明,可以参见上述实施例的步骤204中对拼接处理的描述,即与上述步骤204中的拼接处理的具体方案相同,此处不再赘述。
步骤204e、在拼接后的语音为语义连贯状态的情况下,电子设备将拼接后的语音作为目标语音。
可选地,本申请实施例中,电子设备可以使用自然语言处理(Natural Language Processing,NLP)技术对第一语音和第二语音进行语义比对,以判断第一语音和第二语音中,字、词、以及句子之间是否连贯。
可选地,本申请实施例中,电子设备可以分别获取第一语音和第二语音的语义特征信息,然后根据该语义特征信息,得到第一语音和第二语音的语义相似度,若相似度大于一个阈值,则确认第一语音与第二语音为语义连贯状态;若相似度小于或等于一个阈值,则确认第一语音与第二语音为语义不连贯状态。
示例性地,在第一语音为“你”,第二语音为“起床了吗”,拼接后的语音为“你起床了吗”,即拼接后的语音为语义连贯状态的情况下,电子设备可以将拼接后的语音作为目标语音。
示例性地,在第一语音为“你这会儿忙吗?”,第二语音为“我需要你的帮助”,拼接后的语音为“你这会儿忙吗?我需要你的帮助”,即拼接后的语音为语义连贯状态的情况下,电子设备可以将拼接后的语音作为目标语音。
步骤204f、在拼接后的语音为语义不连贯状态的情况下,电子设备将第一语音或第二语音作为目标语音。
示例性地,在第一语音为“早上好,你”,第二语音为“饭了吗”,拼接后的语音为“早上好,你饭了吗”,即拼接后的语音为语义不连贯状态的情况下,电子设备可以将第一语音:“早上好,你”或第二语音:“饭了吗”作为目标语音。
可选地,本申请实施例中,电子设备可以对第一语音和第二语音的语音长度、语音质
量、语音内容的重要程度等进行对比,以根据对比结果确定最终作为目标语音的语音。
示例性地,在第一语音的语音长度大于第二语音的语音长度时,电子设备可以将第一语音作为目标语音。
示例性地,在第一语音的语音质量高于第二语音的语音质量,例如第一语音的语音清晰度高于第二语音的语音清晰度时,电子设备可以将第一语音作为目标语音。
示例性地,在第一语音的语音内容的重要程度低于第二语音的语音内容的重要程度,例如第一语音的语音内容为“忙完了没?”,第二语音的语音内容为“下午三点会议室开会”时,电子设备可以将第二语音作为目标语音。
本申请实施例中,电子设备在得到第一语音和第二语音之后,通过对第一语音和第二语音进行预拼接处理,并检测拼接后的语音的语义连贯状态,使得电子设备可以根据拼接后的语音的语义连贯状态对拼接后的语音进行对应的处理。如此,提高了电子设备处理录制语音的灵活性和多样性。
可选地,本申请实施例中,本申请实施例提供的语音录制方法还包括下述的步骤301至步骤303,或者,本申请实施例提供的语音录制方法还包括下述的步骤301、步骤302和步骤304。
步骤301、在第一语音和第二语音为语义不连贯状态的情况下,电子设备在当前界面中显示提示信息。
本申请实施例中,上述提示信息用于提示是否将拼接后的语音作为目标语音。
可选地,本申请实施例中,电子设备可以在上述的当前界面中显示一个窗口,该窗口中包括上述的提示信息,以及第一控件和第二控件,该第一控件用于确定将拼接后的语音作为目标语音,第二控件用于确定将第一语音或第二语音作为目标语音。
例如,上述提示信息可以为“是否将拼接后的语音作为发送语音”。
举例说明,如图9所示,手机确定拼接后的语音为语义不连贯状态时,可以在会话界面10中显示窗口13,该窗口13中包括提示信息“是否将拼接后的语音作为发送语音”,并且包括第一控件,例如:“是”控件14,以及包括第二控件,例如:“否”控件15。
步骤302、电子设备接收用户对提示信息的第二输入。
可选地,本申请实施例中,上述第二输入可以为对第一控件或第二控件的输入。
可选地,本申请实施例中,上述第二输入包括但不限于:用户通过手指或者手写笔等触控装置对上述第一控件或第二控件进行触控输入,或者为用户输入的特定手势,或者为点击输入,或者为其他可行性输入。具体的可以根据实际使用需求确定,本发明实施例不作限定。
可选地,本申请实施例中,上述特定手势可以为单击手势、滑动手势、拖动手势、压力识别手势、长按手势、面积变化手势、双按手势、双击手势中的任意一种。
步骤303、电子设备响应于第二输入,在第二输入确认是的情况下,将拼接后的语音作为目标语音。
可以理解,在第二输入为对第一控件的输入时,电子设备可以将拼接后的语音作为目标语音。
步骤304、电子设备响应于第二输入,在第二输入确认否的情况下,将第一语音或第二语音作为目标语音。
可以理解,在第二输入为对第二控件的输入时,电子设备可以将第一语音或第二语音作为目标语音。
举例说明,结合上述图9,用户可以对“是”控件14进行第二输入,以触发手机将拼接后的语音作为目标语音,或者用户可以对“否”控件15进行第二输入,以触发手机将第一语音或第二语音作为目标语音。
需要说明的是,针对电子设备将第一语音或第二语音作为目标语音的具体方案,可以参见上述实施例的上述步骤204f中的描述,此处不再赘述。
可选地,本申请实施例中,在拼接后的语音为语义不连贯状态的情况下,电子设备也可以删除第一语音和第二语音,并提醒用户重新录制语音。
本申请实施例中,电子设备在确定拼接后的语音为语义不连贯状态时,通过显示提示信息,使得用户可以根据该提示信息灵活地选择最终发送的语音,即目标语音。如此,提高了电子设备发送录制语音的多样性和灵活性。
需要说明的是,本申请实施例中,在执行上述步骤204d之后,在第一语音和第二语音为语义不连贯状态的情况下:电子设备可以执行上述步骤204f;或者,电子设备可以执行上述步骤301至步骤303;或者,电子设备可以执行上述步骤301、步骤302和步骤304。
可选地,本申请实施例中,本申请实施例提供的语音录制方法还可以包括下述的步骤401。
步骤401、在当前界面中存在语音录制控件,且录制第一语音的过程中关闭当前界面的情况下,电子设备删除录制的第一语音。
可选地,本申请实施例中,电子设备可以在检测到上述的当前界面中不存在语音录制控件,例如用户触发电子设备在上述的当前界面显示文字输入键盘的情况下,删除录制的第一语音。
可选地,本申请实施例中,电子设备可以在检测到退出上述的当前界面对应的应用时,删除录制的第一语音。
可选地,本申请实施例中,电子设备可以在检测到显示会话列表界面时,删除录制的第一语音。
可选地,本申请实施例中,电子设备可以在接收到来电信息、视频通话、或语音通话时,删除录制的第一语音。
本申请实施例中,电子设备在录制第一语音的过程中,通过检测上述的当前界面是否关闭,使得在上述的当前界面关闭的情况下,电子设备可以删除录制的第一语音。如此,提高了电子设备录制语音的灵活性。
需要说明的是,本申请实施例提供的语音录制方法,执行主体可以为语音录制装置。本申请实施例中以语音录制装置执行语音录制方法为例,说明本申请实施例提供的语音录制装置。
图10示出了本申请实施例中涉及的语音录制装置的一种可能的结构示意图。如图10所示,该语音录制装置70可以包括:录制模块71、接收模块72和处理模块73。
其中,录制模块71,用于在当前界面中存在语音录制控件的情况下,录制第一语音。接收模块72,用于接收用户的第一输入,该第一输入为对语音录制控件的输入。录制模块71,还用于响应于接收模块72接收的第一输入,录制第二语音。处理模块73,用于对录
制模块71得到的第一语音和第二语音处理,得到目标语音。
本申请实施例提供一种语音录制装置,由于语音录制装置在检测到当前界面中存在语音录制控件的情况下,可以先录制第一语音,并在实际接收到用户对语音录制控件的输入时,停止录制第一语音,并开始录制第二语音,然后语音录制装置可以将用户输入之前录制的第一语音和用户输入之后录制的第二语音进行处理,以得到目标语音。如此,避免了只有在用户触发对录制控件的输入后语音录制装置才开始录制语音,容易产生语音缺失的问题,从而保证了语音录制装置录制语音的完整性。
在一种可能的实现方式中,上述录制模块71,具体用于在当前界面中存在语音录制控件,且检测到用户的目标行为信息的情况下,录制第一语音,用户的目标行为信息包括以下至少一项:用户的声音信息,用户注视屏幕的视线信息,用户的嘴巴张合信息。
在一种可能的实现方式中,上述处理模块73,具体用于将第一语音与第二语音进行内容比对;在第一语音与第二语音的内容相同的情况下,删除第一语音,并将第二语音作为目标语音;或者,在第一语音与第二语音的内容不同的情况下,对第一语音和第二语音拼接处理,得到目标语音。
在一种可能的实现方式中,上述处理模块73,具体用于对第一语音和第二语音拼接处理;在拼接后的语音为语义连贯状态的情况下,将拼接后的语音作为目标语音;或者,在拼接后的语音为语义不连贯状态的情况下,将第一语音或第二语音作为目标语音。
在一种可能的实现方式中,结合图10,如图11所示,本申请实施例提供的语音录制装置70还包括:删除模块74。该删除模块74,用于在当前界面中存在语音录制控件,且录制第一语音的过程中关闭当前界面的情况下,删除录制的第一语音。
本申请实施例中的语音录制装置可以是电子设备,也可以是电子设备中的部件,例如集成电路或芯片。该电子设备可以是终端,也可以为除终端之外的其他设备。示例性的,电子设备可以为手机、平板电脑、笔记本电脑、掌上电脑、车载电子设备、移动上网装置(Mobile Internet Device,MID)、增强现实(Augmented Reality,AR)/虚拟现实(Virtual Reality,VR)设备、机器人、可穿戴设备、超级移动个人计算机(Ultra-Mobile Personal Computer,UMPC)、上网本或者个人数字助理(personal digital assistant,PDA)等,还可以为服务器、网络附属存储器(Network Attached Storage,NAS)、个人计算机(personal computer,PC)、电视机(television,TV)、柜员机或者自助机等,本申请实施例不作具体限定。
本申请实施例中的语音录制装置可以为具有操作系统的装置。该操作系统可以为安卓(Android)操作系统,可以为ios操作系统,还可以为其他可能的操作系统,本申请实施例不作具体限定。
本申请实施例提供的语音录制装置能够实现上述方法实施例实现的各个过程,为避免重复,这里不再赘述。
可选地,如图12所示,本申请实施例还提供一种电子设备900,包括处理器901和存储器902,存储器902上存储有可在所述处理器901上运行的程序或指令,该程序或指令被处理器901执行时实现上述方法实施例的各个步骤,且能达到相同的技术效果,为避免重复,这里不再赘述。
需要说明的是,本申请实施例中的电子设备包括上述所述的移动电子设备和非移动电
子设备。
图13为实现本申请实施例的一种电子设备的硬件结构示意图。
该电子设备100包括但不限于:射频单元101、网络模块102、音频输出单元103、输入单元104、传感器105、显示单元106、用户输入单元107、接口单元108、存储器109、以及处理器110等部件。
本领域技术人员可以理解,电子设备100还可以包括给各个部件供电的电源(比如电池),电源可以通过电源管理系统与处理器110逻辑相连,从而通过电源管理系统实现管理充电、放电、以及功耗管理等功能。图13中示出的电子设备结构并不构成对电子设备的限定,电子设备可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置,在此不再赘述。
其中,输入单元104,用于在当前界面中存在语音录制控件的情况下,录制第一语音。
用户输入单元107,用于接收用户的第一输入,该第一输入为对语音录制控件的输入。
输入单元104,还用于响应于用户输入单元107接收的第一输入,录制第二语音。
处理器110,用于对输入单元104得到的第一语音和第二语音处理,得到目标语音。
本申请实施例提供一种电子设备,由于电子设备在检测到当前界面中存在语音录制控件的情况下,可以先录制第一语音,并在实际接收到用户对语音录制控件的输入时,停止录制第一语音,并开始录制第二语音,然后电子设备可以将用户输入之前录制的第一语音和用户输入之后录制的第二语音进行处理,以得到目标语音。如此,避免了只有在用户触发对录制控件的输入后电子设备才开始录制语音,容易产生语音缺失的问题,从而保证了电子设备录制语音的完整性。
可选地,输入单元104,具体用于在当前界面中存在语音录制控件,且检测到用户的目标行为信息的情况下,录制第一语音,用户的目标行为信息包括以下至少一项:用户的声音信息,用户注视屏幕的视线信息,用户的嘴巴张合信息。
可选地,处理器110,具体用于将第一语音与第二语音进行内容比对;在第一语音与第二语音的内容相同的情况下,删除第一语音,并将第二语音作为目标语音;或者,具体用于将第一语音与第二语音进行内容比对,在第一语音与第二语音的内容不同的情况下,对第一语音和第二语音拼接处理,得到目标语音。
可选地,处理器110,具体用于对第一语音和第二语音拼接处理;在拼接后的语音为语义连贯状态的情况下,将拼接后的语音作为目标语音;或者,具体用于对第一语音和第二语音拼接处理,在拼接后的语音为语义不连贯状态的情况下,将第一语音或第二语音作为目标语音。
可选地,处理器110,还用于在当前界面中存在语音录制控件,且录制第一语音的过程中关闭当前界面的情况下,删除录制的第一语音。
本申请实施例提供的电子设备能够实现上述方法实施例实现的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
本实施例中各种实现方式具有的有益效果具体可以参见上述方法实施例中相应实现方式所具有的有益效果,为避免重复,此处不再赘述。
应理解的是,本申请实施例中,输入单元104可以包括图形处理器(Graphics Processing Unit,GPU)1041和麦克风1042,图形处理器1041对在视频捕获模式或图像捕获模式中
由图像捕获装置(如摄像头)获得的静态图片或视频的图像数据进行处理。显示单元106可包括显示面板1061,可以采用液晶显示器、有机发光二极管等形式来配置显示面板1061。用户输入单元107包括触控面板1071以及其他输入设备1072中的至少一种。触控面板1071,也称为触摸屏。触控面板1071可包括触摸检测装置和触摸控制器两个部分。其他输入设备1072可以包括但不限于物理键盘、功能键(比如音量控制按键、开关按键等)、轨迹球、鼠标、操作杆,在此不再赘述。
存储器109可用于存储软件程序以及各种数据。存储器109可主要包括存储程序或指令的第一存储区和存储数据的第二存储区,其中,第一存储区可存储操作系统、至少一个功能所需的应用程序或指令(比如声音播放功能、图像播放功能等)等。此外,存储器109可以包括易失性存储器或非易失性存储器,或者,存储器109可以包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(Read-Only Memory,ROM)、可编程只读存储器(Programmable ROM,PROM)、可擦除可编程只读存储器(Erasable PROM,EPROM)、电可擦除可编程只读存储器(Electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(Random Access Memory,RAM),静态随机存取存储器(Static RAM,SRAM)、动态随机存取存储器(Dynamic RAM,DRAM)、同步动态随机存取存储器(Synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(Double Data Rate SDRAM,DDRSDRAM)、增强型同步动态随机存取存储器(Enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(Synch link DRAM,SLDRAM)和直接内存总线随机存取存储器(Direct Rambus RAM,DRRAM)。本申请实施例中的存储器109包括但不限于这些和任意其它适合类型的存储器。
处理器110可包括一个或多个处理单元;可选的,处理器110集成应用处理器和调制解调处理器,其中,应用处理器主要处理涉及操作系统、用户界面和应用程序等的操作,调制解调处理器主要处理无线通信信号,如基带处理器。可以理解的是,上述调制解调处理器也可以不集成到处理器110中。
本申请实施例还提供一种可读存储介质,所述可读存储介质上存储有程序或指令,该程序或指令被处理器执行时实现上述方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
其中,所述处理器为上述实施例中所述的电子设备中的处理器。所述可读存储介质,包括计算机可读存储介质,如计算机只读存储器ROM、随机存取存储器RAM、磁碟或者光盘等。
本申请实施例另提供了一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现上述方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
应理解,本申请实施例提到的芯片还可以称为系统级芯片、系统芯片、芯片系统或片上系统芯片等。
本申请实施例提供一种计算机程序产品,该程序产品被存储在存储介质中,该程序产品被至少一个处理器执行以实现如上述方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。
需要说明的是,在本文中,术语“包括”、“包含”或者其任何其他变体意在涵盖非
排他性的包含,从而使得包括一系列要素的过程、方法、物品或者装置不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者装置所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者装置中还存在另外的相同要素。此外,需要指出的是,本申请实施方式中的方法和装置的范围不限按示出或讨论的顺序来执行功能,还可包括根据所涉及的功能按基本同时的方式或按相反的顺序来执行功能,例如,可以按不同于所描述的次序来执行所描述的方法,并且还可以添加、省去、或组合各种步骤。另外,参照某些示例所描述的特征可在其他示例中被组合。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以计算机软件产品的形式体现出来,该计算机软件产品存储在一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端(可以是手机,计算机,服务器,或者网络设备等)执行本申请各个实施例所述的方法。
上面结合附图对本申请的实施例进行了描述,但是本申请并不局限于上述的具体实施方式,上述的具体实施方式仅仅是示意性的,而不是限制性的,本领域的普通技术人员在本申请的启示下,在不脱离本申请宗旨和权利要求所保护的范围情况下,还可做出很多形式,均属于本申请的保护之内。
Claims (15)
- 一种语音录制方法,所述方法包括:在当前界面中存在语音录制控件的情况下,录制第一语音;接收用户的第一输入,所述第一输入为对所述语音录制控件的输入;响应于所述第一输入,录制第二语音;对所述第一语音和所述第二语音处理,得到目标语音。
- 根据权利要求1所述的方法,其中,所述在当前界面中存在语音录制控件的情况下,录制第一语音,包括:在所述当前界面中存在所述语音录制控件,且检测到用户的目标行为信息的情况下,录制所述第一语音,所述用户的目标行为信息包括以下至少一项:用户的声音信息,用户注视屏幕的视线信息,用户的嘴巴张合信息。
- 根据权利要求1所述的方法,其中,所述对所述第一语音和所述第二语音处理,得到目标语音,包括:将所述第一语音与所述第二语音进行内容比对;在所述第一语音与所述第二语音的内容相同的情况下,删除所述第一语音,并将所述第二语音作为所述目标语音;或者,在所述第一语音与所述第二语音的内容不同的情况下,对所述第一语音和所述第二语音拼接处理,得到所述目标语音。
- 根据权利要求1或3所述的方法,其中,所述对所述第一语音和所述第二语音处理,得到目标语音,包括:对所述第一语音和所述第二语音拼接处理;在拼接后的语音为语义连贯状态的情况下,将拼接后的语音作为所述目标语音;或者,在拼接后的语音为语义不连贯状态的情况下,将所述第一语音或所述第二语音作为所述目标语音。
- 根据权利要求1所述的方法,其中,所述方法还包括:在所述当前界面中存在所述语音录制控件,且录制所述第一语音的过程中关闭所述当前界面的情况下,删除录制的所述第一语音。
- 一种语音录制装置,所述语音录制装置包括:录制模块、接收模块和处理模块;所述录制模块,用于在当前界面中存在语音录制控件的情况下,录制第一语音;所述接收模块,用于接收用户的第一输入,所述第一输入为对所述语音录制控件的输入;所述录制模块,还用于响应于所述接收模块接收的所述第一输入,录制第二语音;所述处理模块,用于对所述录制模块得到的所述第一语音和所述第二语音处理,得到目标语音。
- 根据权利要求6所述的装置,其中,所述录制模块,具体用于在所述当前界面中存在所述语音录制控件,且检测到用户的目标行为信息的情况下,录制第一语音,所述用户的目标行为信息包括以下至少一项:用户的声音信息,用户注视屏幕的视线信息,用户的嘴巴张合信息。
- 根据权利要求6所述的装置,其中,所述处理模块,具体用于:将所述第一语音与所述第二语音进行内容比对;在所述第一语音与所述第二语音的内容相同的情况下,删除所述第一语音,并将所述第二语音作为所述目标语音;或者,在所述第一语音与所述第二语音的内容不同的情况下,对所述第一语音和所述第二语音拼接处理,得到所述目标语音。
- 根据权利要求6或8所述的装置,其中,所述处理模块,具体用于:对所述第一语音和所述第二语音拼接处理;在拼接后的语音为语义连贯状态的情况下,将拼接后的语音作为所述目标语音;或者,在拼接后的语音为语义不连贯状态的情况下,将所述第一语音或所述第二语音作为所述目标语音。
- 根据权利要求6所述的装置,其中,所述语音录制装置还包括:删除模块;所述删除模块,用于在所述当前界面中存在所述语音录制控件,且录制所述第一语音的过程中关闭所述当前界面的情况下,删除录制的所述第一语音。
- 一种电子设备,包括处理器和存储器,所述存储器存储可在所述处理器上运行的程序或指令,所述程序或指令被所述处理器执行时实现如权利要求1-5中任一项所述的语音录制方法的步骤。
- 一种可读存储介质,所述可读存储介质上存储程序或指令,所述程序或指令被处理器执行时实现如权利要求1-5中任一项所述的语音录制方法的步骤。
- 一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现如权利要求1-5中任一项所述的语音录制方法。
- 一种计算机程序产品,所述程序产品被至少一个处理器执行以实现如权利要求1-5中任一项所述的语音录制方法。
- 一种电子设备,包括所述电子设备被配置成用于执行如权利要求1-5中任一项所述的语音录制方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310433734.8 | 2023-04-20 | ||
| CN202310433734.8A CN116543745A (zh) | 2023-04-20 | 2023-04-20 | 语音录制方法、装置、电子设备及存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024217470A1 true WO2024217470A1 (zh) | 2024-10-24 |
Family
ID=87455285
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/088407 Ceased WO2024217470A1 (zh) | 2023-04-20 | 2024-04-17 | 语音录制方法、装置、电子设备及存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN116543745A (zh) |
| WO (1) | WO2024217470A1 (zh) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116543745A (zh) * | 2023-04-20 | 2023-08-04 | 维沃移动通信有限公司 | 语音录制方法、装置、电子设备及存储介质 |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105760084A (zh) * | 2016-01-25 | 2016-07-13 | 百度在线网络技术(北京)有限公司 | 语音输入的控制方法和装置 |
| CN106383644A (zh) * | 2016-09-19 | 2017-02-08 | 广东小天才科技有限公司 | 一种电子设备的语音识别开启方法及装置 |
| CN106531167A (zh) * | 2016-11-18 | 2017-03-22 | 北京云知声信息技术有限公司 | 一种语音信息的处理方法及装置 |
| CN107886975A (zh) * | 2017-11-07 | 2018-04-06 | 广东欧珀移动通信有限公司 | 音频的处理方法、装置、存储介质及电子设备 |
| KR20180084469A (ko) * | 2017-01-17 | 2018-07-25 | 네이버 주식회사 | 음성 데이터 제공 방법 및 장치 |
| CN108694947A (zh) * | 2018-06-27 | 2018-10-23 | Oppo广东移动通信有限公司 | 语音控制方法、装置、存储介质及电子设备 |
| CN111492426A (zh) * | 2017-12-22 | 2020-08-04 | 瑞典爱立信有限公司 | 注视启动的语音控制 |
| CN116543745A (zh) * | 2023-04-20 | 2023-08-04 | 维沃移动通信有限公司 | 语音录制方法、装置、电子设备及存储介质 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0782759B2 (ja) * | 1988-04-28 | 1995-09-06 | カシオ計算機株式会社 | 音響録音再生装置 |
| CN106653074A (zh) * | 2015-11-02 | 2017-05-10 | 深圳富泰宏精密工业有限公司 | 电子设备及用所述电子设备实时声音回放的方法 |
| CN108156326B (zh) * | 2018-01-02 | 2021-02-02 | 京东方科技集团股份有限公司 | 一种自动启动录音的方法、系统及装置 |
| CN111179908A (zh) * | 2020-01-03 | 2020-05-19 | 苏宁智能终端有限公司 | 智能语音设备的测试方法及系统 |
-
2023
- 2023-04-20 CN CN202310433734.8A patent/CN116543745A/zh active Pending
-
2024
- 2024-04-17 WO PCT/CN2024/088407 patent/WO2024217470A1/zh not_active Ceased
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN105760084A (zh) * | 2016-01-25 | 2016-07-13 | 百度在线网络技术(北京)有限公司 | 语音输入的控制方法和装置 |
| CN106383644A (zh) * | 2016-09-19 | 2017-02-08 | 广东小天才科技有限公司 | 一种电子设备的语音识别开启方法及装置 |
| CN106531167A (zh) * | 2016-11-18 | 2017-03-22 | 北京云知声信息技术有限公司 | 一种语音信息的处理方法及装置 |
| KR20180084469A (ko) * | 2017-01-17 | 2018-07-25 | 네이버 주식회사 | 음성 데이터 제공 방법 및 장치 |
| CN107886975A (zh) * | 2017-11-07 | 2018-04-06 | 广东欧珀移动通信有限公司 | 音频的处理方法、装置、存储介质及电子设备 |
| CN111492426A (zh) * | 2017-12-22 | 2020-08-04 | 瑞典爱立信有限公司 | 注视启动的语音控制 |
| CN108694947A (zh) * | 2018-06-27 | 2018-10-23 | Oppo广东移动通信有限公司 | 语音控制方法、装置、存储介质及电子设备 |
| CN116543745A (zh) * | 2023-04-20 | 2023-08-04 | 维沃移动通信有限公司 | 语音录制方法、装置、电子设备及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116543745A (zh) | 2023-08-04 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11620984B2 (en) | Human-computer interaction method, and electronic device and storage medium thereof | |
| JP7166294B2 (ja) | オーディオ処理方法、装置及び記憶媒体 | |
| CN112540821B (zh) | 信息发送方法和电子设备 | |
| WO2017114020A1 (zh) | 语音输入方法和终端设备 | |
| WO2022142871A1 (zh) | 视频录制方法及装置 | |
| CN112102841B (zh) | 一种音频编辑方法、装置和用于音频编辑的装置 | |
| CN109634501B (zh) | 电子书批注添加方法、电子设备及计算机存储介质 | |
| CN106527883A (zh) | 一种内容分享的方法、装置及终端 | |
| CN108984256A (zh) | 界面显示方法、装置、存储介质及电子设备 | |
| WO2024078514A1 (zh) | 投屏方法、装置、电子设备和存储介质 | |
| CN111756930A (zh) | 通信控制方法、通信控制装置、电子设备和可读存储介质 | |
| CN112667118A (zh) | 显示历史聊天消息的方法、设备以及计算机可读介质 | |
| WO2024001956A1 (zh) | 视频通话方法、装置、第一电子设备及第二电子设备 | |
| WO2025040061A1 (zh) | 截图方法、装置、电子设备及可读存储介质 | |
| WO2018219261A1 (zh) | 文本重组方法、装置、终端设备及计算机可读存储介质 | |
| WO2024217470A1 (zh) | 语音录制方法、装置、电子设备及存储介质 | |
| CN110929484A (zh) | 文本处理方法、装置及存储介质 | |
| WO2024104113A1 (zh) | 截屏方法、截屏装置、电子设备及可读存储介质 | |
| CN112866469A (zh) | 通话内容的记录方法及装置 | |
| CN115473867A (zh) | 消息发送方法、装置、电子设备及存储介质 | |
| CN114979355A (zh) | 麦克风的控制方法、装置及电子设备 | |
| CN108600625A (zh) | 图像获取方法及装置 | |
| CN115080170B (zh) | 信息处理方法、信息处理装置和电子设备 | |
| CN115412634B (zh) | 消息显示方法和装置 | |
| CN117459485A (zh) | 信息处理方法、装置、电子设备及介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24792055 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |