WO2020239013A1 - 一种交互方法和终端设备 - Google Patents
一种交互方法和终端设备 Download PDFInfo
- Publication number
- WO2020239013A1 WO2020239013A1 PCT/CN2020/092888 CN2020092888W WO2020239013A1 WO 2020239013 A1 WO2020239013 A1 WO 2020239013A1 CN 2020092888 W CN2020092888 W CN 2020092888W WO 2020239013 A1 WO2020239013 A1 WO 2020239013A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- terminal device
- input content
- content
- control instruction
- input
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/22—Interactive procedures; Man-machine interfaces
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/26—Speech to text systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M1/00—Substation equipment, e.g. for use by subscribers
- H04M1/72—Mobile telephones; Cordless telephones, i.e. devices for establishing wireless links to base stations without route selection
- H04M1/725—Cordless telephones
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/223—Execution procedure of a spoken command
Definitions
- This application relates to the field of terminals, and in particular to an interaction method and terminal equipment.
- the smart speaker and the mobile phone can communicate through wireless fidelity (WiFi) or the Internet.
- WiFi wireless fidelity
- the user can operate the designated menu on the mobile application (application, APP), Send instructions to the smart speaker to control the smart speaker for music playback and other operations.
- APP mobile application
- the user usually can only perform limited operations that have been pre-made on the mobile phone APP, such as music playback, pause, etc., which limits the functions of the smart speaker.
- the embodiments of the present application provide an interaction method and terminal device, which can enhance the flexibility of the user to remotely operate the intelligent voice interaction device and give full play to the capabilities of the intelligent voice interaction device.
- an embodiment of the present application provides an interaction method, including: a first terminal device receives first input content from a second terminal device, the first input content includes first voice content and/or first text content; The terminal device determines the control instruction corresponding to the first input content; the first terminal device processes the control instruction corresponding to the first input content.
- the first terminal device can receive the first input content from the second terminal device, determine the control instruction corresponding to the first input content, and process the control instruction. In this way, the first terminal device can determine the control instruction based on the first input content (the content that the user remotely enters through the second terminal device), instead of the user operating according to the limited instructions pre-made on the mobile phone APP, which enhances the user
- the flexibility of remote interaction with the first terminal device gives full play to the capabilities of the first terminal device.
- the first terminal device determining the control instruction corresponding to the first input content includes: the first terminal device sends the first input content to the server; the first terminal device receives the control corresponding to the first input content from the server instruction.
- the server may include an automatic speech recognition (ASR) engine and a natural language processing (NLP) engine.
- ASR engine can convert the first speech content into the first text information
- NLP engine can use the first text information The semantics of get control instructions.
- the processing of the control instruction corresponding to the first input content by the first terminal device includes: the first terminal device executes the control instruction corresponding to the first input content; or, the first terminal device sends to the third terminal device The control instruction corresponding to the first input.
- the first terminal device can receive the control instruction issued by the server, and forward it to the third terminal device, thereby realizing control of the third terminal device, giving full play to the capabilities of the first terminal device and improving user experience.
- the method further includes: the first terminal device sends a first response message corresponding to the first input content to the second terminal device.
- the first response message may include the control instruction transcribed by the server and the execution result after the smart speaker executes the control instruction. Therefore, after the user inputs the first input content through the second terminal device, he can learn the control instruction obtained according to the first input content and the execution result of the control instruction, thereby improving the user experience.
- processing the control instruction corresponding to the first input content by the first terminal device includes: the first terminal device receives second input content, and the second input content includes second voice content and/or second text content ; If the first terminal device receives the second input content earlier than the first input content, the first terminal device processes the control instructions corresponding to the second input content, and then processes the control instructions corresponding to the first input content. In other words, when the smart speaker processes different control instructions, it can follow the first-come, first-execute strategy (the first-acquired control instructions are executed first).
- an embodiment of the present application provides an interaction method, including: a second terminal device receives first input content from a user, the first input content includes first voice content and/or first text content; The first terminal device sends the first input content; the second terminal device receives the first response message corresponding to the first input content from the first terminal device; the second terminal device broadcasts a voice or displays the first response message on the display screen.
- the user of the second terminal device after receiving the first input content, sends the first input content to the first terminal device, and then receives the first response message corresponding to the first input content from the first terminal device , And voice broadcast or display the first response message on the display screen.
- the user after the user inputs the first input content through the second terminal device, he can learn the control instruction obtained according to the first input content and the execution result of the control instruction, thereby improving user experience.
- the method further includes: the second terminal device receives a second response message from the first terminal device, the second response message corresponds to the second input content; the second terminal device makes a voice broadcast or displays the first response message on the display screen. 2. Response message.
- the user of the first terminal device can learn the operations of other users on the second terminal device (including the control instruction corresponding to the second input content and the execution result of the control instruction), which improves the user’s understanding of the second terminal device Control the ability to improve the user experience.
- an embodiment of the present application provides a first terminal device, including: a receiving unit, configured to receive first input content from a second terminal device, the first input content includes first voice content and/or first text content
- the determining unit is used to determine the control instruction corresponding to the first input content
- the processing unit is used to process the control instruction corresponding to the first input content.
- the determining unit is configured to: send the first input content to the server through the sending unit; and receive the control instruction corresponding to the first input content from the server through the receiving unit.
- the processing unit is configured to: execute the control instruction corresponding to the first input content; or, send the control instruction corresponding to the first input content to the third terminal device through the sending unit.
- the sending unit is further configured to send a first response message corresponding to the first input content to the second terminal device.
- the processing unit is configured to: receive the second input content through the receiving unit, the second input content includes second voice content and/or second text content; if the first terminal device receives the second input content The time of is earlier than the time of receiving the first input content, and after the control instruction corresponding to the second input content is processed, the control instruction corresponding to the first input content is processed.
- an embodiment of the present application provides a second terminal device, including: a receiving unit, configured to receive first input content from a user, the first input content including first voice content and/or first text content; and a sending unit , Is used to send the first input content to the first terminal device; the receiving unit is also used to receive the first response message corresponding to the first input content from the first terminal device; the processing unit is used to broadcast the voice or display the first input on the display screen A response message.
- the receiving unit is further used for: receiving a second response message from the first terminal device, the second response message corresponding to the second input content; the processing unit is also used for voice broadcasting or displaying the first 2. Response message.
- an embodiment of the present application also provides a device, which may be a first terminal device or a chip.
- the device includes a processor, configured to implement any one of the interaction methods provided in the first aspect.
- the device may also include a memory for storing program instructions and data.
- the memory may be a memory integrated in the device or an off-chip memory provided outside the device.
- the memory is coupled with the processor, and the processor can call and execute program instructions stored in the memory to implement any one of the interaction methods provided in the first aspect.
- the device may also include a communication interface, which is used for the device to communicate with other devices (for example, a second terminal device).
- an embodiment of the present application also provides a device, which may be a second terminal device or a chip.
- the device includes a processor, which is configured to implement any one of the interaction methods provided in the second aspect.
- the device may also include a memory for storing program instructions and data.
- the memory may be a memory integrated in the device or an off-chip memory provided outside the device.
- the memory is coupled with the processor, and the processor can call and execute the program instructions stored in the memory to implement any one of the interaction methods provided in the second aspect.
- the device may also include a communication interface for the device to communicate with other devices (for example, the first terminal device).
- an embodiment of the present application provides a computer-readable storage medium, including instructions, which when run on a computer, cause the computer to execute any one of the interaction methods provided in the first aspect or the second aspect.
- the embodiments of the present application provide a computer program product containing instructions, which when run on a computer, cause the computer to execute any one of the interaction methods provided in the first aspect or the second aspect.
- an embodiment of the present application provides a chip system.
- the chip system includes a processor and may also include a memory, configured to implement any one of the interaction methods provided in the first or second aspect.
- the chip system can be composed of chips, or can include chips and other discrete devices.
- an embodiment of the present application provides an interactive system.
- the system includes the first terminal device in the third aspect and the second terminal device in the fourth aspect.
- Figure 1 is a schematic diagram of a communication architecture between a smart speaker and a mobile phone in the prior art
- FIG. 2 is a schematic diagram of an architecture suitable for an interaction method provided by an embodiment of the present application
- FIG. 3 is a schematic structural diagram of another interaction method applicable to an embodiment of the present application.
- FIG. 4 is a schematic diagram of the internal structure of a second terminal device according to an embodiment of the application.
- FIG. 5 is a schematic diagram of the internal structure of a first terminal device according to an embodiment of the application.
- FIG. 6 is a schematic diagram of signal interaction applicable to the interaction method provided by an embodiment of the present application.
- FIG. 7 is a schematic diagram of a desktop of a mobile phone according to an embodiment of the application.
- FIG. 8 is a schematic diagram of an input interface of a smart speaker APP provided by an embodiment of the application.
- FIG. 9 is a schematic diagram of an input interface of yet another smart speaker APP provided by an embodiment of the application.
- FIG. 10 is a schematic diagram of the internal structure of still another first terminal device according to an embodiment of the application.
- FIG. 11 is a schematic diagram of the internal structure of yet another second terminal device according to an embodiment of the application.
- the embodiments of the present application provide an interaction method and terminal device, which are applied to an interaction system composed of a first terminal device and a second terminal device, and a user can remotely interact with the first terminal device through the second terminal device.
- an interaction system composed of a first terminal device and a second terminal device
- a user can remotely interact with the first terminal device through the second terminal device.
- the first terminal device and the second terminal device can communicate via new radio access technology (New RAT), long term evolution (LTE), Bluetooth (bluetooth, BT), Wifi or other protocols ,
- New RAT new radio access technology
- LTE long term evolution
- Bluetooth bluetooth, BT
- Wifi Wireless Fidelity
- the system may include a first terminal device (for example, a mobile phone 10a), a second terminal device (for example, a smart speaker 10b), and a first terminal device (for example, a smart speaker 10b).
- a network device for example, internet server 11
- a second network device for example, cloud server 12.
- the first terminal device may receive the first input content sent by the second terminal device through the internet server 11, and may send a first response message of the first input content to the second terminal device through the internet server 11, etc.
- the cloud server 12 may be used to parse voice content and/or text content.
- the cloud server 12 may convert voice content into text information through the ASR engine, and convert the text information into control instructions through the NLP engine, so that the second terminal device can respond according to the control instructions.
- the cloud server 12 may be a server corresponding to the smart speaker APP installed on the mobile phone 10a and the smart speaker 10b, or may be a third-party server corresponding to a smart speaker program integrated in another APP, which is not limited in this application.
- the first terminal device and the internet server 11, the second terminal device and the internet server 11, and the second terminal device and the cloud server 12 can communicate through wireless communication.
- the wireless communication method may be, for example, through wireless communication.
- the networked device for example, base station
- the base station may be an evolved node base station (eNB).
- eNB evolved node base station
- the base station can be the next generation node base station (gNB), new radio base station (new radio eNB), macro base station, micro base station, high frequency Base station or transmitting and receiving point (transmission and reception point, TRP), etc.
- the first terminal device may be a user equipment (UE), for example, a mobile phone, a tablet computer, a desktop type, a laptop notebook, or an ultra-mobile personal computer (Ultra-mobile Personal Computer). , UMPC), handheld computers, netbooks, personal digital assistants (personal digital assistant, PDA), vehicle-mounted terminals and other equipment.
- the second terminal device can be various smart home devices or UEs.
- the smart home devices can be, for example, smart speakers, smart TVs, smart refrigerators, smart washing machines, smart rice cookers, smart dishwashers, smart sweeping robots, and so on.
- the interactive system may further include a third terminal device (for example, the cleaning robot 10c), and the third terminal device is connected to the second terminal device.
- the third terminal device may be various smart home devices or UEs and so on.
- the second terminal device in the foregoing communication system architecture may specifically be a mobile phone 100.
- the mobile phone 100 may include a processor 110, an internal memory 120, a camera 130, a display screen 140, a radio frequency module 150, a communication module 160, an antenna 1, an antenna 2, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, etc. .
- the structure illustrated in the embodiment of the present invention does not constitute a limitation on the mobile phone 100. It may include more or fewer components than shown, or combine certain components, or split certain components, or arrange different components.
- the illustrated components can be implemented in hardware, software, or a combination of software and hardware.
- the processor 110 may include one or more processing units.
- the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), and an image signal processor. (image signal processor, ISP), controller, memory, video codec, digital signal processor (digital signal processor, DSP), baseband processor, and/or neural network processor (Neural-network Processing Unit, NPU) Wait.
- AP application processor
- modem processor modem processor
- GPU graphics processing unit
- image signal processor image signal processor
- ISP image signal processor
- controller memory
- video codec digital signal processor
- DSP digital signal processor
- baseband processor baseband processor
- neural network processor Neural-network Processing Unit, NPU
- the controller may be a decision maker who directs the various components of the mobile phone 100 to coordinate work according to instructions. It is the nerve center and command center of the mobile phone 100.
- the controller generates operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
- a memory may also be provided in the processor 110 to store instructions and data.
- the memory in the processor is a cache memory. It can save the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can be directly called from the memory. It avoids repeated access and reduces the waiting time of the processor, thereby improving the efficiency of the system.
- the processor 110 may include an interface.
- the interfaces may include integrated circuit (inter-integrated circuit, I2C) interfaces, integrated circuit built-in audio (inter-integrated circuit sound, I2S) interfaces, pulse code modulation (PCM) interfaces, universal asynchronous transceiver (universal asynchronous transceiver) interfaces, and asynchronous receiver/transmitter, UART) interface, mobile industry processor interface (MIPI), general-purpose input/output (GPIO) interface, subscriber identity module (SIM) interface, And/or Universal Serial Bus (USB) interface, etc.
- I2C integrated circuit
- I2S integrated circuit built-in audio
- PCM pulse code modulation
- MIPI mobile industry processor interface
- GPIO general-purpose input/output
- SIM subscriber identity module
- USB Universal Serial Bus
- the wireless communication function of the mobile phone 100 can be implemented by the antenna 1, the antenna 2, the radio frequency module 150, the communication module 160, a modem, and a baseband processor.
- the antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals.
- Each antenna in the mobile phone 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization.
- the cellular network antenna can be multiplexed into a wireless LAN diversity antenna.
- the antenna can be used in conjunction with a tuning switch.
- the radio frequency module 150 can provide applications on the mobile phone 100 including the second generation (2 th generation, 2G)/third generation (3 th generation, 3G)/fourth generation (4 th generation, 4G)/fifth generation (5 th generation, 5G) and other wireless communication solutions communication processing module. It may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc.
- the radio frequency module receives electromagnetic waves from the antenna 1, filters and amplifies the received electromagnetic waves, and transmits them to the modem for demodulation.
- the radio frequency module 150 can also amplify the signal modulated by the modem, and convert it into electromagnetic waves for radiation by the antenna 1.
- at least part of the functional modules of the radio frequency module 150 may be provided in the processor 110.
- at least part of the functional modules of the radio frequency module 150 and at least part of the modules of the processor 110 may be provided in the same device.
- the modem may include a modulator and a demodulator.
- the modulator is used to modulate the low frequency baseband signal to be sent into a medium and high frequency signal.
- the demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Then the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing.
- the low-frequency baseband signal is processed by the baseband processor and then passed to the application processor.
- the application processor outputs sound signals through audio equipment (not limited to speakers, receivers, etc.), or displays images or videos through the display screen.
- the modem may be a standalone device. In some embodiments, the modem may be independent of the processor and be provided in the same device as the radio frequency module or other functional modules.
- the communication module 160 can provide applications on the mobile phone 100 including wireless local area networks (WLAN) (for example, WiFi), Bluetooth, global navigation satellite system (GNSS), frequency modulation (FM) , Near field communication technology (near field communication, NFC), infrared technology (infrared, IR) and other wireless communication solutions communication processing module.
- WLAN wireless local area networks
- GNSS global navigation satellite system
- FM frequency modulation
- NFC Near field communication technology
- infrared technology infrared, IR
- the communication module 160 may be one or more devices integrating at least one communication processing module.
- the communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor.
- the communication module 160 can also receive the signal to be sent from the processor, perform frequency modulation, amplify it, and convert it into electromagnetic wave radiation via the antenna 2.
- the antenna 1 of the mobile phone 100 is coupled with the radio frequency module 150, and the antenna 2 is coupled with the communication module 160.
- the mobile phone 100 can communicate with the network and other devices through wireless communication technology.
- the wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), broadband Code division multiple access (wideband code division multiple access, WCDMA), time division code division multiple access (time-division code division multiple access, TD-SCDMA), LTE, 5G New Radio (NR), BT, GNSS, WLAN, NFC, FM, and/or IR technology, etc.
- the GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), Beidou navigation satellite system (BDS), quasi-zenith satellite system (quasi -zenith satellite system, QZSS) and/or satellite-based augmentation systems (SBAS).
- GPS global positioning system
- GLONASS global navigation satellite system
- BDS Beidou navigation satellite system
- QZSS quasi-zenith satellite system
- SBAS satellite-based augmentation systems
- the mobile phone 100 implements a display function through a GPU, a display screen 140, and an application processor.
- GPU is a microprocessor for image processing, connected to the display screen and the application processor.
- the GPU is used to perform mathematical and geometric calculations for graphics rendering.
- the processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
- the display screen 140 is used to display images, videos, etc.
- the display screen includes a display panel.
- the display panel can adopt liquid crystal display (LCD), organic light-emitting diode (OLED), active-matrix organic light-emitting diode or active-matrix organic light-emitting diode (active-matrix organic light-emitting diode).
- LCD liquid crystal display
- OLED organic light-emitting diode
- active-matrix organic light-emitting diode active-matrix organic light-emitting diode
- AMOLED Miniled, MicroLed, Micro-oLed, quantum dot light emitting diode (QLED), etc.
- the mobile phone 100 may include 1 or N display screens, and N is a positive integer greater than 1.
- the mobile phone 100 can realize a shooting function through an ISP, a camera 130, a video codec, a GPU, a display screen 140, and an application processor.
- ISP is used to process the data fed back from the camera. For example, when taking a picture, the shutter is opened, the light is transmitted to the photosensitive element of the camera through the lens, the light signal is converted into an electrical signal, and the photosensitive element of the camera transfers the electrical signal to the ISP for processing and is converted into an image visible to the naked eye.
- ISP can also optimize the image noise, brightness, and skin color. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene.
- the ISP may be provided in the camera 130.
- the camera 130 is used to capture still images or videos.
- the object generates an optical image through the lens and projects it to the photosensitive element.
- the photosensitive element may be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor.
- CMOS complementary metal-oxide-semiconductor
- the photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal.
- ISP outputs digital image signals to DSP for processing.
- DSP converts digital image signals into standard RGB, YUV and other formats.
- the mobile phone 100 may include 1 or N cameras, and N is a positive integer greater than 1.
- the internal memory 120 may be used to store computer executable program code, the executable program code including instructions.
- the processor 110 executes various functional applications and data processing of the mobile phone 100 by running instructions stored in the internal memory 120.
- the memory 120 may include a program storage area and a data storage area.
- the storage program area can store an operating system, at least one application program (such as a sound playback function, an image playback function, etc.) required by at least one function.
- the data storage area can store data (such as audio data, phone book, etc.) created during the use of the mobile phone 100.
- the memory 120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, other volatile solid-state storage devices, universal flash storage (UFS), etc. .
- the mobile phone 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor. For example, music playback, recording, etc.
- the audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal.
- the audio module can also be used to encode and decode audio signals.
- the audio module 170 may be disposed in the processor 110, or part of the functional modules of the audio module may be disposed in the processor 110.
- the speaker 170A also called a “speaker” is used to convert audio electrical signals into sound signals.
- the mobile phone 100 can listen to music through a speaker or listen to a hands-free call.
- the receiver 170B also called “earpiece” is used to convert audio electrical signals into sound signals.
- the mobile phone 100 answers a call or voice message, it can receive the voice by bringing the receiver close to the human ear.
- the microphone 170C also called “microphone”, “microphone”, is used to convert sound signals into electrical signals.
- the user can make a sound by approaching the microphone through the human mouth, and input the sound signal into the microphone.
- the mobile phone 100 can be provided with at least one microphone.
- the mobile phone 100 can be equipped with two microphones, which can realize noise reduction in addition to collecting sound signals.
- the mobile phone 100 can also be equipped with three, four or more microphones to collect sound signals, reduce noise, identify the source of sound, and realize the function of directional recording.
- the earphone interface 170D is used to connect wired earphones.
- the earphone interface can be a USB interface, or a 3.5mm open mobile terminal platform (OMTP) standard interface, and a cellular telecommunications industry association of the USA (CTIA) standard interface.
- OMTP open mobile terminal platform
- CTIA cellular telecommunications industry association of the USA
- the first terminal device in the foregoing communication system architecture may be, for example, a smart home device 200.
- the smart home device 200 may include components such as a processor 201, a display screen 202, a storage module 203, a communication module 204, a radio frequency module 205, an antenna 01, an antenna 02, a microphone 206, and a speaker 207.
- a processor 201 a display screen 202
- a storage module 203 a communication module 204
- a radio frequency module 205 a radio frequency module
- antenna 01 an antenna 01
- antenna 02 an antenna 02
- microphone 206 a microphone
- speaker 207 a speaker
- the antenna 01 of the smart home device 200 is coupled with the communication module, and the antenna 02 is coupled with the radio frequency module.
- the smart home device 200 can communicate with the network and other devices through wireless communication technology.
- the wireless communication technology may include LTE, 5G NR, WLAN and so on.
- the smart home device 200 can interact with the internet server 11 and the cloud server 12 (also referred to as the cloud).
- the smart home device 200 described above may have more or fewer components than shown in FIG. 5, two or more components may be combined, or may have different component configurations.
- the various components shown in FIG. 5 may be implemented in hardware, software, or a combination of hardware and software including one or more signal processing or application specific integrated circuits.
- an embodiment of the present application provides an interaction method, taking the first terminal device as a smart speaker and the second terminal device as a mobile phone as an example for description, including:
- the mobile phone receives first input content of a user, where the first input content includes first voice content and/or first text content.
- the user's mobile phone can install the smart speaker APP client, and the user can send the first voice content and/or the first text content to the smart speaker through the APP client on the mobile phone.
- the user can click the icon 702 of the smart speaker APP on the desktop 701 of the mobile phone.
- the mobile phone detects that the user clicks the icon 702 of the smart speaker APP on the desktop 701, it can start the smart speaker APP and display the graphical user interface (GUI) as shown in Figure 8.
- GUI graphical user interface
- the GUI can be called input Interface 801.
- the input interface 801 may include a voice input icon 802 and a voice input prompt message 803 "please input voice", or the voice prompt message may be "please speak the operation you want to perform", etc. (not shown in FIG.
- the prompt message for inputting voice can be displayed after the user clicks on the voice input icon 802; the text input part can include a text box 804 and a prompt message 805 "please enter text" for entering text, or the prompt message can be "please edit out what you want Operations performed” (not shown in Figure 8) and so on.
- the user can operate the smart speaker through the smart speaker applet or official account in other apps of the mobile phone (for example, WeChat APP), or the user can log in to the smart speaker webpage through a browser to operate the smart speaker.
- apps of the mobile phone for example, WeChat APP
- This application does not Make a limit.
- the mobile phone can pick up the voice signal (first voice content) sent by the user through the microphone.
- the user can click on the voice input icon when performing voice input.
- the mobile phone activates the microphone to detect the voice signal (what the user says), and the first voice content may be "please turn on the cleaning robot at home to clean the living room" Or "please clean the living room.”
- the mobile phone can consider that the voice input is complete.
- the mobile phone may prompt the user to complete the voice input through a prompt tone or prompt message.
- the mobile phone can preset the voice length that the user can input, for example, it is stipulated that the voice length input by the user is greater than 1s and less than or equal to 60s.
- the volume of the voice content made by the user is too low (for example, less than 20 dB)
- the user may be prompted to adjust the volume through a prompt tone or prompt message, so that the microphone can better pick up the user's voice content.
- the mobile phone may receive the text content (first text content) input by the user through an input device (such as a touch screen or a keyboard).
- the first text content may be a piece of experience or perception edited by the user himself.
- the first text content may be a short essay or a poem copied by the user from a browser or an application such as WeChat.
- the input method can be displayed around the text box (for example, below or above the text box). For example, Pinyin input method or handwriting input method) or shortcuts (copy, paste, etc.), etc., to facilitate user operations.
- the mobile phone can pick up the voice content (first voice content) uttered by the user through a microphone, and receive the text content (first text content) input by the user through the input device.
- the first text content may be a paragraph of text edited by the user, and the first voice content may be "please read this paragraph of text (that is, the first text content) in an expressive/high/low tone".
- the user can input the first voice content and then the first text content, or the user can input the first text content and then the first voice content, or the user can input the first voice content and the first text content at the same time. Make a limit.
- the first input content may include image information, such as screenshots of webpages, train tickets, and airline ticket booking information.
- image information such as screenshots of webpages, train tickets, and airline ticket booking information.
- the user can input the first voice content or the first text content: "the flight ticket of the scheduled picture" through the smart speaker APP, and input a screenshot of the ticket information.
- the Internet environment can be used to pick up the voice content of the user through the microphone of the mobile phone, and/or the text content input by the user is received through the input device, so that the mobile phone and the smart speaker can realize remote interaction. It enables users to remotely operate smart speakers without barriers, and communicate with other members of the home through smart speakers, giving full play to the capabilities of smart speakers at home.
- the mobile phone sends the first input content to the smart speaker.
- the mobile phone can directly send the voice content input by the user (first voice content) and/or the text content input by the user (first text content) to the smart speaker without corresponding processing of the first input content, such as semantic extraction or operation Operations such as instruction matching.
- the mobile phone can directly send the audio message of the sentence "Please turn on the robot cleaner at home to clean the living room” to the smart speaker without having to follow the " Please turn on the sweeping robot at home to clean the living room.
- the semantics of the sentence match the corresponding operating instructions.
- the smart speaker receives the first input content from the mobile phone.
- the smart speaker determines a control instruction corresponding to the first input content.
- the smart speaker can upload the first input content to a server (for example, a cloud server).
- a server for example, a cloud server.
- the ASR engine of the cloud server can convert the first voice content into the first text information
- the NLP engine of the cloud server can obtain the control instruction according to the semantics of the first text information.
- the control instruction may include the vertical type and slot of the first text information. Exemplarily, if the user says "please book me a ticket to Beijing the next morning" or "I want to book a ticket to Beijing the next morning", the NLP engine can perform intent recognition based on keywords to determine the verticality of the first text message.
- the category is "book air tickets”, and the slots for this vertical category are determined to include: the departure time is “morning after tomorrow" and the destination is "Beijing". Or, if the user says “tell a little joke", the NLP engine can identify the vertical category as “storytelling” based on the intent of the keyword. And determine that the vertical slot includes: story type is "joke”, story length is “short”.
- the cloud server issues a control instruction to the smart speaker, and the smart speaker receives the control instruction issued by the cloud server.
- the NLP engine of the cloud server can obtain a control instruction according to the semantics of the text information, and then the cloud server issues the control instruction to the smart speaker, and the smart speaker receives the control instruction issued by the cloud server.
- the smart speaker processes the control instruction corresponding to the first input content.
- the control instruction corresponding to the first input content may be for the smart speaker, that is, the smart speaker is required to respond according to the control instruction; the control instruction corresponding to the first input content may also be for the third terminal device (for example, a household sweeping robot) , That is, the smart speaker forwards the control instruction to the third terminal device, and the third terminal device responds according to the control instruction.
- the control instruction corresponding to the first input content may also be for the third terminal device (for example, a household sweeping robot) , That is, the smart speaker forwards the control instruction to the third terminal device, and the third terminal device responds according to the control instruction.
- control instruction is for the smart speaker
- the smart speaker executes the control instruction corresponding to the first input content. For example, if the control command is PLAY, the smart speaker will play music.
- the smart speaker can send the control instruction corresponding to the first input content to the third terminal device, so that the third terminal device executes the control instruction corresponding to the first input content, thereby controlling the third terminal device Perform corresponding operations, such as power on, power off, etc.
- the third terminal device may be a terminal device matched with a smart speaker, for example, it may be a sweeping robot matched with a smart speaker, an air conditioner, a refrigerator, a washing machine, a smart curtain, etc.
- the mobile phone when the user needs to remotely operate the home cleaning robot and the mobile phone does not match the home cleaning robot, the mobile phone does not control the cleaning robot instruction and cannot directly control the cleaning robot.
- the mobile phone can control the cleaning robot through the smart speaker.
- the user can input the following voice content through the smart voice APP client of the mobile phone: "Please turn on the cleaning robot at home to clean the living room".
- the smart speaker After the smart speaker receives the voice content, it can upload the voice content to the cloud server.
- the ASR engine of the cloud server can convert the voice content into text information
- the NLP engine of the cloud server can obtain control instructions based on the semantics of the text information.
- the cloud server issues control instructions to the smart speakers, and the smart speakers receive the control instructions issued by the cloud server and forward them to the sweeping robot, thereby achieving control of the sweeping robot, giving full play to the capabilities of the smart speaker and improving users Experience.
- the smart speaker sends a first response message corresponding to the first input content to the mobile phone.
- the control command is for a smart speaker
- the smart speaker executes the control command corresponding to the first input content
- it sends a first response message to the mobile phone.
- the first response message may include the control command transcribed by the cloud service and after the smart speaker executes the control command
- the user can learn the control instruction obtained according to the first input content and the execution result of the control instruction, which improves the user experience. For example, user A can send the voice content of "Play the moonlight in lotus pond" through the APP of the mobile phone.
- the smart speaker receives the voice content, it determines the control instruction corresponding to the voice content through the cloud server, and plays the corresponding control instruction according to the control instruction. Songs, and the playback status (that is, the first response message corresponding to the first input content) is sent to the mobile phone APP.
- control instruction is for the third terminal device
- the smart speaker after the smart speaker receives the execution result corresponding to the first input content from the third terminal device, it can send a first response message to the mobile phone, and the first response message can include the control transcribed by the cloud service
- the instruction and the execution result after the third terminal device executes the control instruction can include the control transcribed by the cloud service
- the mobile phone receives a first response message corresponding to the first input content from the smart speaker.
- the mobile phone can voice broadcast and/or display on the display screen the control instructions and execution results (ie response messages) of the smart speaker transcribed by the cloud server.
- the mobile phone can display the voice content 806 input by the user 805 on the APP, and display the control instruction 808 processed by the smart speaker 807 as “Start playing light music at 18:50 for 30 minutes”, and the execution result 809 as “Start playing light music at 18:50” And the execution result 810 "End playing light music at 19:20”.
- the mobile phone can display the time information 811 of the voice content input by the user 805 on the APP, the time information 812 that the mobile phone receives the control instruction 808 sent by the smart speaker, the time information 813 that the mobile phone receives the execution result 809, and the mobile phone receives Time information 814 of the execution result 810.
- the mobile phone can also voice broadcast the control instruction 808, the execution result 809, and the execution result 810 to the user.
- the mobile phone may not display the control instruction 808, execution result 809, and execution result 810, and only broadcast to the user through voice. For example, when the mobile phone detects that the user is listening to music and the mobile phone screen is in a black screen state, the mobile phone voice broadcasts the response message fed back by the smart speaker without lighting the screen, which can save power consumption.
- interaction method may also include:
- the smart speaker receives second input content, where the second input content includes second voice content and/or second text content.
- the second input content can be the voice content (second voice content) input by other users (different from the user who input the first input content) beside the speaker, or the second input content can be other users through the fourth terminal device (for example, user The second voice content and/or the second text content sent by the mobile phone or smart wearable device, etc.).
- the smart speaker processes the control command corresponding to the second input content and then processes the control command corresponding to the first input content; if the smart speaker receives the second input The content time is later than the time when the first input content is received. After the smart speaker processes the control instruction corresponding to the first input content, it processes the control instruction corresponding to the second input content. In other words, smart speakers follow the first come first execute strategy when processing different control commands.
- Scenario 1 User A sends the voice content of "play music" through the APP of the mobile phone. After the smart speaker receives the voice content, it determines the instruction corresponding to the voice content through the cloud server, plays the music according to the instruction, and can send the playback status To the mobile APP. In the process of playing music, user B uses voice commands to control the pause next to the speaker. The expected result is: the smart speaker pauses the music playback and can send the pause status to the mobile phone APP.
- the smart speaker is playing music, user B controls the pause by voice commands next to the speaker, and then user A sends the voice content of "play music" through the APP of the mobile phone. After the smart speaker receives the voice content, it is determined by the cloud server The expected result of the instruction corresponding to the voice content is: the smart speaker continues to play the music after pausing the music, and can send the playback status to the mobile phone APP.
- User A sends the voice content of "Play Music” through the APP of the mobile phone.
- the smart speaker determines the instruction corresponding to the voice content through the cloud server, and plays music according to the instruction.
- the music playing process For example, after playing 1/3 of a song
- user B controls the music playback by voice beside the speaker (the content requested by user B is exactly the same as the content requested by user A, for example, both require to play “Lotus pond moonlight” ")
- the expected result is: smart speakers continue to play “Lotus Pond Moonlight” without having to start playing "Lotus Pond Moonlight” from the beginning, which can meet user needs and avoid repeating the same instructions.
- Step 609 can be executed before step 603, can also be executed after step 603, or simultaneously with step 603. This embodiment does not do this. Specific restrictions.
- the smart speaker can send a second response message corresponding to the second input content to the mobile phone.
- the mobile phone receives the second response message from the smart speaker, the second response message corresponds to the second input content, the mobile phone voice broadcasts or displays the second response message on the display screen.
- the user of the first terminal device can learn the operations of other users on the second terminal device (including the control instruction corresponding to the second input content and the execution result of the control instruction), which improves the user’s understanding of the second terminal.
- the ability to control the device improves the user experience.
- the second terminal device (for example, a mobile phone) can receive the first input content from the user, and send the first input content to the first terminal device (for example, an intelligent voice interaction device).
- the first terminal device receives the first input content from the second terminal device, determines the control instruction corresponding to the first input content through the server, and processes the control instruction.
- the first terminal device can determine the control instruction based on the content (ie the first input content) remotely input by the user through the second terminal device, instead of the user operating according to the limited instructions pre-made on the mobile phone APP, which enhances The flexibility of remote interaction between the user and the first terminal device gives full play to the capabilities of the first terminal device.
- the methods provided by the embodiments of the present application are introduced from the perspective of the first terminal device, the second terminal device, and the interaction between the first terminal device and the second terminal device.
- the first terminal device and the second terminal device may include a hardware structure and/or a software module, in the form of a hardware structure, a software module, or a hardware structure plus a software module To achieve the above functions. Whether one of the above-mentioned functions is executed in a hardware structure, a software module, or a hardware structure plus a software module depends on the specific application and design constraint conditions of the technical solution.
- FIG. 10 shows a possible structural schematic diagram of the apparatus 10 involved in the foregoing embodiment.
- the apparatus may be a first terminal device, and the first terminal device includes : The receiving unit 1001, the determining unit 1002, and the processing unit 1002.
- the receiving unit 1001 is configured to receive first input content from the second terminal device, the first input content includes the first voice content and/or the first text content;
- the determining unit 1002 is configured to determine the first The control instruction corresponding to the input content;
- the processing unit 1003 is used to process the control instruction corresponding to the first input content.
- the first terminal device may further include a sending unit 1004 (not shown in FIG. 10), configured to send a first response message corresponding to the first input content to the second terminal device.
- the receiving unit 1001 is used to support the first terminal device to perform the process 603 in FIG. 6; the determining unit 1002 is used to support the first terminal device to perform the process 604 in FIG. 6; the processing unit 1003 It is used to support the first terminal device to perform the process 605 in FIG. 6; the sending unit 1004 is used to support the first terminal device to perform the process 606 in FIG. 6.
- FIG. 11 shows a possible structural diagram of the apparatus 11 involved in the above embodiment.
- the apparatus may be a second terminal device, and the second terminal device includes : Receiving unit 1101, sending unit 1102, and processing unit 1103.
- the receiving unit 1101 is configured to receive first input content from a user, the first input content includes first voice content and/or first text content;
- the sending unit 1102 is configured to send to the first terminal device The first input content;
- the receiving unit 1101 is further configured to receive the first response message corresponding to the first input content from the first terminal device;
- the processing unit 1103 is configured to voice broadcast or display the first response message on the display screen.
- the receiving unit 1101 is used to support the second terminal device to perform the processes 601 and 607 in FIG. 6; the sending unit 1102 is used to support the second terminal device to perform the process 602 in FIG. 6; processing The unit 1103 is used to support the second terminal device to execute the process 608 in FIG. 6.
- the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, there may be other division methods.
- the functional modules in the various embodiments of the present application may be integrated into one process. In the device, it can also exist alone physically, or two or more modules can be integrated into one module.
- the above-mentioned integrated modules can be implemented in the form of hardware or software functional modules.
- the receiving unit and the sending unit may be integrated into the transceiver unit.
- the methods provided in the embodiments of the present application may be implemented in whole or in part by software, hardware, firmware, or any combination thereof.
- software When implemented by software, it can be implemented in the form of a computer program product in whole or in part.
- the computer program product includes one or more computer instructions.
- the computer may be a general-purpose computer, a dedicated computer, a computer network, network equipment, user equipment, or other programmable devices.
- the computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center.
- the computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center integrated with one or more available media.
- the usable medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital video disc (DVD)), or a semiconductor medium (for example, a solid state drive (SSD)) )Wait.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Computational Linguistics (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Telephone Function (AREA)
- Telephonic Communication Services (AREA)
Abstract
本申请实施例提供一种交互方法和终端设备,涉及终端领域,能够增强用户远程操作智能语音交互设备的灵活性,充分发挥智能语音交互设备的能力。其方法为:第一终端设备从第二终端设备接收第一输入内容,第一输入内容包括第一语音内容和/或第一文本内容;第一终端设备确定第一输入内容对应的控制指令;第一终端设备处理第一输入内容对应的控制指令。本申请实施例应用于远程交互场景中。
Description
本申请要求在2019年5月31日提交中国国家知识产权局、申请号为201910472665.5的中国专利申请的优先权,发明名称为“一种交互方法和终端设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及终端领域,尤其涉及一种交互方法和终端设备。
当前的智能语音交互设备(例如,智能音箱)是通过语音将用户与设备之间联系起来的。用户可以在智能音箱旁直接通过语音控制智能音箱,让智能音箱播放音乐、讲故事等等。但当用户远离智能音箱时(例如大于5米时),与智能音箱设备便无法取得联系。也就是说,用户在与智能音箱距离较近才能使用智能音箱,用户远离智能音箱时无法使用智能音箱,此时智能音箱处于闲置状态,造成资源浪费。
为解决上述问题,如图1所示,智能音箱与手机可以通过无线保真(wireless fidelity,WiFi)或者互联网(Internet)进行通信,用户可以通过操作手机应用(application,APP)上的指定菜单,发送指令给智能音箱,控制智能音箱进行音乐播放等操作。
但是,上述方案中,用户通常只能进行手机APP上已经预制好的有限的操作,例如音乐播放、暂停等,限制了智能音箱的功能。
发明内容
本申请实施例提供一种交互方法和终端设备,能够增强用户远程操作智能语音交互设备的灵活性,充分发挥智能语音交互设备的能力。
第一方面,本申请实施例提供一种交互方法,包括:第一终端设备从第二终端设备接收第一输入内容,第一输入内容包括第一语音内容和/或第一文本内容;第一终端设备确定第一输入内容对应的控制指令;第一终端设备处理第一输入内容对应的控制指令。
基于本申请实施例提供的方法,第一终端设备可以从第二终端设备接收第一输入内容,确定第一输入内容对应的控制指令,并处理该控制指令。这样一来,可以由第一终端设备基于第一输入内容(用户通过第二终端设备远程输入的内容)确定控制指令,而非用户根据手机APP上预制好的有限的指令进行操作,增强了用户与第一终端设备远程交互的灵活性,充分发挥第一终端设备的能力。
在一种可能的实现方式中,第一终端设备确定第一输入内容对应的控制指令包括:第一终端设备向服务器发送第一输入内容;第一终端设备从服务器接收第一输入内容对应的控制指令。服务器可以包括自动语音识别(automatic speech recognition,ASR)引擎和自然语言处理(natural language processing,NLP)引擎,ASR引擎可以将第一语音内容转化为第一文本信息,NLP引擎可以根据第一文本信息的语义得到控制指令。
在一种可能的实现方式中,第一终端设备处理第一输入内容对应的控制指令包括:第一终端设备执行第一输入内容对应的控制指令;或者,第一终端设备向第三终端设备发送第一输入内容对应的控制指令。第一终端设备可以接收服务器下发的控制指令,并转发给第三终端设备,从而实现对第三终端设备的控制,充分发挥了第一终端设备的能力,提高了用户体验。
在一种可能的实现方式中,该方法还包括:第一终端设备向第二终端设备发送第一输入内容对应的第一响应消息。第一响应消息可以包括服务器转写的控制指令以及智能音箱执行控制指令后的执行结果。以便用户通过第二终端设备输入第一输入内容后,可以获知根据第一输入内容得到的控制指令和针对该控制指令的执行结果,提高用户体验。
在一种可能的实现方式中,第一终端设备处理第一输入内容对应的控制指令包括:第一终端设备接收第二输入内容,第二输入内容包括第二语音内容和/或第二文本内容;若第一终端设备接收第二输入内容的时刻早于接收第一输入内容的时刻,第一终端设备处理第二输入内容对应的控制指令后,处理第一输入内容对应的控制指令。也就是说,智能音箱在处理不同的控制指令时,可以遵循先来先执行(先获取的控制指令先执行)的策略。
第二方面,本申请实施例提供一种交互方法,包括:第二终端设备从用户接收第一输入内容,第一输入内容包括第一语音内容和/或第一文本内容;第二终端设备向第一终端设备发送第一输入内容;第二终端设备从第一终端设备接收第一输入内容对应的第一响应消息;第二终端设备语音播报或在显示屏显示第一响应消息。
基于本申请实施例提供的方法,第二终端设备用户接收第一输入内容后,向第一终端设备发送第一输入内容,而后,从第一终端设备接收第一输入内容对应的第一响应消息,并语音播报或在显示屏显示第一响应消息。这样一来,用户通过第二终端设备输入第一输入内容后,可以获知根据第一输入内容得到的控制指令和针对该控制指令的执行结果,提高用户体验。
在一种可能的实现方式中,方法还包括:第二终端设备从第一终端设备接收第二响应消息,第二响应消息对应第二输入内容;第二终端设备语音播报或在显示屏显示第二响应消息。
这样一来,若其他用户(与输入第一输入内容的用户不同)在音箱旁输入语音内容(第二语音内容),或者其他用户通过第四终端设备(例如手机或智能穿戴设备等)输入第二输入内容,第一终端设备的用户可以获知其他用户对第二终端设备的操作(包括第二输入内容对应的控制指令和针对该控制指令的执行结果),提高了用户对第二终端设备的掌控能力,从而提高了用户体验。
第二方面及其各种可能的实现方式的技术效果可以参见第一方面及其各种可能的实现方式的技术效果,此处不再赘述。
第三方面,本申请实施例提供一种第一终端设备,包括:接收单元,用于从第二终端设备接收第一输入内容,第一输入内容包括第一语音内容和/或第一文本内容;确定单元,用于确定第一输入内容对应的控制指令;处理单元,用于处理第一输入内容对应的控制指令。
在一种可能的实现方式中,确定单元用于:通过发送单元向服务器发送第一输入内容;通过接收单元从服务器接收第一输入内容对应的控制指令。
在一种可能的实现方式中,处理单元用于:执行第一输入内容对应的控制指令;或者,通过发送单元向第三终端设备发送第一输入内容对应的控制指令。
在一种可能的实现方式中,发送单元还用于:向第二终端设备发送第一输入内容对应的第一响应消息。
在一种可能的实现方式中,处理单元用于:通过接收单元接收第二输入内容,第二输入内容包括第二语音内容和/或第二文本内容;若第一终端设备接收第二输入内容的时刻早于接收第一输入内容的时刻,处理第二输入内容对应的控制指令后,处理第一输入内容对应的控制指令。
第四方面,本申请实施例提供一种第二终端设备,包括:接收单元,用于从用户接收第 一输入内容,第一输入内容包括第一语音内容和/或第一文本内容;发送单元,用于向第一终端设备发送第一输入内容;接收单元,还用于从第一终端设备接收第一输入内容对应的第一响应消息;处理单元,用于语音播报或在显示屏显示第一响应消息。
在一种可能的实现方式中,接收单元还用于:从第一终端设备接收第二响应消息,第二响应消息对应第二输入内容;处理单元,还用于语音播报或在显示屏显示第二响应消息。
第五方面,本申请实施例还提供了一种装置,该装置可以是第一终端设备或芯片。该装置包括处理器,用于实现上述第一方面提供的任意一种交互方法。该装置还可以包括存储器,用于存储程序指令和数据,存储器可以是集成在该装置内的存储器,或设置在该装置外的片外存储器。该存储器与该处理器耦合,该处理器可以调用并执行该存储器中存储的程序指令,用于实现上述第一方面提供的任意一种交互方法。该装置还可以包括通信接口,该通信接口用于该装置与其它设备(例如,第二终端设备)进行通信。
第六方面,本申请实施例还提供了一种装置,该装置可以是第二终端设备或芯片。该装置包括处理器,用于实现上述第二方面提供的任意一种交互方法。该装置还可以包括存储器,用于存储程序指令和数据,存储器可以是集成在该装置内的存储器,或设置在该装置外的片外存储器。该存储器与该处理器耦合,该处理器可以调用并执行该存储器中存储的程序指令,用于实现上述第二方面提供的任意一种交互方法。该装置还可以包括通信接口,该通信接口用于该装置与其它设备(例如,第一终端设备)进行通信。
第七方面,本申请实施例提供一种计算机可读存储介质,包括指令,当其在计算机上运行时,使得计算机执行上述第一方面或第二方面提供的任意一种交互方法。
第八方面,本申请实施例提供了一种包含指令的计算机程序产品,当其在计算机上运行时,使得计算机执行上述第一方面或第二方面提供的任意一种交互方法。
第九方面,本申请实施例提供了一种芯片系统,该芯片系统包括处理器,还可以包括存储器,用于实现上述第一方面或第二方面提供的任意一种交互方法。该芯片系统可以由芯片构成,也可以包含芯片和其他分立器件。
第十方面,本申请实施例提供了一种交互系统,所述系统包括第三方面中的第一终端设备和第四方面中的第二终端设备。
图1为现有技术中的一种智能音箱与手机的通信架构示意图;
图2为一种适用于本申请实施例提供的交互方法的架构示意图;
图3为又一种适用于本申请实施例提供的交互方法的架构示意图;
图4为本申请实施例提供的一种第二终端设备的内部结构示意图;
图5为本申请实施例提供的一种第一终端设备的内部结构示意图;
图6为一种适用于本申请实施例提供的交互方法的信号交互示意图;
图7为本申请实施例提供的一种手机的桌面示意图;
图8为本申请实施例提供的一种智能音箱APP的输入界面示意图;
图9为本申请实施例提供的又一种智能音箱APP的输入界面示意图;
图10为本申请实施例提供的又一种第一终端设备的内部结构示意图;
图11为本申请实施例提供的又一种第二终端设备的内部结构示意图。
本申请实施例提供一种交互方法和终端设备,应用于第一终端设备和第二终端设备组成的交互系统中,用户可以通过第二终端设备与第一终端设备远程交互。例如,应用于手机与 (配对的)智能家居产品(例如,智能音箱)组成的交互系统中。第一终端设备和第二终端设备之间可以通过新无线接入(new radio access technical,New RAT)、长期演进(long term evolution,LTE)、蓝牙(bluetooth,BT)、Wifi或其它协议进行通信,本申请不做限定。
如图2所示,为本申请实施例提供的一种交互系统的架构示意图,该系统可以包括第一终端设备(例如,手机10a)、第二终端设备(例如,智能音箱10b)、第一网络设备(例如,internet服务器11)和第二网络设备(例如,云服务器12)。第一终端设备可以通过internet服务器11接收第二终端设备发送的第一输入内容,并可以通过internet服务器11向第二终端设备发送第一输入内容的第一响应消息等。云服务器12可以用于解析语音内容和/或文本内容。例如,云服务器12可以通过ASR引擎将语音内容转换成文本信息,通过NLP引擎将文本信息转换成控制指令,以便第二终端设备根据控制指令做出响应。云服务器12可以是手机10a和智能音箱10b上安装的智能音箱APP对应的服务器,或者可以为集成在其他APP内的智能音箱程序对应的第三方服务器,本申请不做限定。
第一终端设备和internet服务器11之间,第二终端设备和internet服务器11,第二终端设备和云服务器12之间,可以通过无线的通信方式进行通信,无线的通信方式例如可以是通过无线接入网设备(例如,基站)进行通信。在LTE网络中,基站可以为演进型基站(evolved node base station,eNB)。在第五代移动通信技术(5-Generation,5G)网络中,基站可以为下一代基站(next generation node base station,gNB)、新型无线电基站(new radio eNB)、宏基站、微基站、高频基站或发送和接收点(transmission and reception point,TRP)等。
其中,本申请实施例提供的第一终端设备可以是用户设备(user equipment,UE),例如可以为手机、平板电脑、桌面型、膝上型笔记本电脑、超级移动个人计算机(ultra-mobile personal computer,UMPC)、手持计算机、上网本、个人数字助理(personal digital assistant,PDA)、车载终端等设备。第二终端设备可以为各种智能家居设备或UE,智能家居设备例如可以为智能音箱、智能电视、智能冰箱、智能洗衣机、智能电饭煲、智能洗碗机、智能扫地机器人等等。
在一种可能的设计中,如图3所示,交互系统还可以包括第三终端设备(例如扫地机器人10c),第三终端设备与第二终端设备连接。第三终端设备可以为各种智能家居设备或UE等等。
如图4所示,上述通信系统架构中的第二终端设备具体可以为手机100。手机100可以包括处理器110,内部存储器120,摄像头130,显示屏140,射频模块150,通信模块160,天线1,天线2,音频模块170,扬声器170A,受话器170B,麦克风170C,耳机接口170D等。
本发明实施例示意的结构并不构成对手机100的限定。可以包括比图示更多或更少的部件,或者组合某些部件,或者拆分某些部件,或者不同的部件布置。图示的部件可以以硬件,软件或软件和硬件的组合实现。
处理器110可以包括一个或多个处理单元,例如:处理器110可以包括应用处理器(application processor,AP),调制解调处理器,图形处理器(graphics processing unit,GPU),图像信号处理器(image signal processor,ISP),控制器,存储器,视频编解码器,数字信号处理器(digital signal processor,DSP),基带处理器,和/或神经网络处理器(Neural-network Processing Unit,NPU)等。其中,不同的处理单元可以是独立的器件,也可以是集成在同一个处理器中。
控制器可以是指挥手机100的各个部件按照指令协调工作的决策者。是手机100的神经 中枢和指挥中心。控制器根据指令操作码和时序信号,产生操作控制信号,完成取指令和执行指令的控制。
处理器110中还可以设置存储器,用于存储指令和数据。在一些实施例中,处理器中的存储器为高速缓冲存储器。可以保存处理器刚用过或循环使用的指令或数据。如果处理器需要再次使用该指令或数据,可从所述存储器中直接调用。避免了重复存取,减少了处理器的等待时间,因而提高了系统的效率。
在一些实施例中,处理器110可以包括接口。其中接口可以包括集成电路(inter-integrated circuit,I2C)接口,集成电路内置音频(inter-integrated circuit sound,I2S)接口,脉冲编码调制(pulse code modulation,PCM)接口,通用异步收发传输器(universal asynchronous receiver/transmitter,UART)接口,移动产业处理器接口(mobile industry processor interface,MIPI),通用输入输出(general-purpose input/output,GPIO)接口,用户标识模块(subscriber identity module,SIM)接口,和/或通用串行总线(universal serial bus,USB)接口等。
手机100的无线通信功能可以通过天线1,天线2,射频模块150,通信模块160,调制解调器以及基带处理器等实现。
天线1和天线2用于发射和接收电磁波信号。手机100中的每个天线可用于覆盖单个或多个通信频带。不同的天线还可以复用,以提高天线的利用率。例如:可以将蜂窝网天线复用为无线局域网分集天线。在一些实施例中,天线可以和调谐开关结合使用。
射频模块150可以提供应用在手机100上的包括第二代(2
th generation,2G)/第三代(3
th generation,3G)/第四代(4
th generation,4G)/第五代(5
th generation,5G)等无线通信的解决方案的通信处理模块。可以包括至少一个滤波器,开关,功率放大器,低噪声放大器(low noise amplifier,LNA)等。射频模块由天线1接收电磁波,并对接收的电磁波进行滤波,放大等处理,传送至调制解调器进行解调。射频模块150还可以对经调制解调器调制后的信号放大,经天线1转为电磁波辐射出去。在一些实施例中,射频模块150的至少部分功能模块可以被设置于处理器110中。在一些实施例中,射频模块150的至少部分功能模块可以与处理器110的至少部分模块被设置在同一个器件中。
调制解调器可以包括调制器和解调器。调制器用于将待发送的低频基带信号调制成中高频信号。解调器用于将接收的电磁波信号解调为低频基带信号。随后解调器将解调得到的低频基带信号传送至基带处理器处理。低频基带信号经基带处理器处理后,被传递给应用处理器。应用处理器通过音频设备(不限于扬声器,受话器等)输出声音信号,或通过显示屏显示图像或视频。在一些实施例中,调制解调器可以是独立的器件。在一些实施例中,调制解调器可以独立于处理器,与射频模块或其他功能模块设置在同一个器件中。
通信模块160可以提供应用在手机100上的包括无线局域网(wireless local area networks,WLAN)(例如,WiFi)、蓝牙,全球导航卫星系统(global navigation satellite system,GNSS),调频(frequency modulation,FM),近距离无线通信技术(near field communication,NFC),红外技术(infrared,IR)等无线通信的解决方案的通信处理模块。通信模块160可以是集成至少一个通信处理模块的一个或多个器件。通信模块160经由天线2接收电磁波,将电磁波信号调频以及滤波处理,将处理后的信号发送到处理器。通信模块160还可以从处理器接收待发送的信号,对其进行调频,放大,经天线2转为电磁波辐射出去。
在一些实施例中,手机100的天线1和射频模块150耦合,天线2和通信模块160耦合。使得手机100可以通过无线通信技术与网络以及其他设备通信。所述无线通信技术可以包括全球移动通讯系统(global system for mobile communications,GSM),通用分组无线服务(general packet radio service,GPRS),码分多址接入(code division multiple access,CDMA),宽带码分多址(wideband code division multiple access,WCDMA),时分码分多址(time-division code division multiple access,TD-SCDMA),LTE,5G新无线通信(New Radio,NR),BT,GNSS,WLAN,NFC,FM,和/或IR技术等。所述GNSS可以包括全球卫星定位系统(global positioning system,GPS),全球导航卫星系统(global navigation satellite system,GLONASS),北斗卫星导航系统(beidou navigation satellite system,BDS),准天顶卫星系统(quasi-zenith satellite system,QZSS))和/或星基增强系统(satellite based augmentation systems,SBAS)。
手机100通过GPU,显示屏140,以及应用处理器等实现显示功能。GPU为图像处理的微处理器,连接显示屏和应用处理器。GPU用于执行数学和几何计算,用于图形渲染。处理器110可包括一个或多个GPU,其执行程序指令以生成或改变显示信息。
显示屏140用于显示图像,视频等。显示屏包括显示面板。显示面板可以采用液晶显示屏(liquid crystal display,LCD),有机发光二极管(organic light-emitting diode,OLED),有源矩阵有机发光二极体或主动矩阵有机发光二极体(active-matrix organic light emitting diode的,AMOLED),Miniled,MicroLed,Micro-oLed,量子点发光二极管(quantum dot light emitting diodes,QLED)等。在一些实施例中,手机100可以包括1个或N个显示屏,N为大于1的正整数。
手机100可以通过ISP,摄像头130,视频编解码器,GPU,显示屏140以及应用处理器等实现拍摄功能。
ISP用于处理摄像头反馈的数据。例如,拍照时,打开快门,光线通过镜头被传递到摄像头感光元件上,光信号转换为电信号,摄像头感光元件将所述电信号传递给ISP处理,转化为肉眼可见的图像。ISP还可以对图像的噪点,亮度,肤色进行算法优化。ISP还可以对拍摄场景的曝光,色温等参数优化。在一些实施例中,ISP可以设置在摄像头130中。
摄像头130用于捕获静态图像或视频。物体通过镜头生成光学图像投射到感光元件。感光元件可以是电荷耦合器件(charge coupled device,CCD)或互补金属氧化物半导体(complementary metal-oxide-semiconductor,CMOS)光电晶体管。感光元件把光信号转换成电信号,之后将电信号传递给ISP转换成数字图像信号。ISP将数字图像信号输出到DSP加工处理。DSP将数字图像信号转换成标准的RGB,YUV等格式的图像信号。在一些实施例中,手机100可以包括1个或N个摄像头,N为大于1的正整数。
内部存储器120可以用于存储计算机可执行程序代码,所述可执行程序代码包括指令。处理器110通过运行存储在内部存储器120的指令,从而执行手机100的各种功能应用以及数据处理。存储器120可以包括存储程序区和存储数据区。其中,存储程序区可存储操作系统,至少一个功能所需的应用程序(比如声音播放功能,图像播放功能等)等。存储数据区可存储手机100使用过程中所创建的数据(比如音频数据,电话本等)等。此外,存储器120可以包括高速随机存取存储器,还可以包括非易失性存储器,例如至少一个磁盘存储器件,闪存器件,其他易失性固态存储器件,通用闪存存储器(universal flash storage,UFS)等。
手机100可以通过音频模块170,扬声器170A,受话器170B,麦克风170C,耳机接口170D,以及应用处理器等实现音频功能。例如音乐播放,录音等。
音频模块170用于将数字音频信息转换成模拟音频信号输出,也用于将模拟音频输入转换为数字音频信号。音频模块还可以用于对音频信号编码和解码。在一些实施例中,音频模块170可以设置于处理器110中,或将音频模块的部分功能模块设置于处理器110中。
扬声器170A,也称“喇叭”,用于将音频电信号转换为声音信号。手机100可以通过扬声 器收听音乐,或收听免提通话。
受话器170B,也称“听筒”,用于将音频电信号转换成声音信号。当手机100接听电话或语音信息时,可以通过将受话器靠近人耳接听语音。
麦克风170C,也称“话筒”,“传声器”,用于将声音信号转换为电信号。当拨打电话或发送语音信息时,用户可以通过人嘴靠近麦克风发声,将声音信号输入到麦克风。手机100可以设置至少一个麦克风。在一些实施例中,手机100可以设置两个麦克风,除了采集声音信号,还可以实现降噪功能。在一些实施例中,手机100还可以设置三个,四个或更多麦克风,实现采集声音信号,降噪,还可以识别声音来源,实现定向录音功能等。
耳机接口170D用于连接有线耳机。耳机接口可以是USB接口,也可以是3.5mm的开放移动终端平台(open mobile terminal platform,OMTP)标准接口,美国蜂窝电信工业协会(cellular telecommunications industry association of the USA,CTIA)标准接口。
如图5所示,上述通信系统架构中的第一终端设备例如可以为智能家居设备200。智能家居设备200中可以包括处理器201、显示屏202、存储模块203、通信模块204、射频模块205、天线01、天线02、麦克风206以及扬声器207等部件。各部件的功能可以参考上文相关描述,在此不做赘述。
在一些实施例中,智能家居设备200的天线01和通信模块耦合,天线02和射频模块耦合。使得智能家居设备200可以通过无线通信技术与网络以及其他设备通信。所述无线通信技术可以包括LTE,5G NR,WLAN等。从而,智能家居设备200可以与internet服务器11以及云服务器12(也可以称为云端)交互。
可以理解的是,上述智能家居设备200可以具有比图5中所示出的更多的或者更少的部件,可以组合两个或更多的部件,或者可以具有不同的部件配置。图5中所示出的各种部件可以在包括一个或多个信号处理或专用集成电路在内的硬件、软件、或硬件和软件的组合中实现。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行描述。其中,在本申请的描述中,除非另有说明,“至少一个”是指一个或多个,“多个”是指两个或多于两个。另外,为了便于清楚描述本申请实施例的技术方案,在本申请的实施例中,采用了“第一”、“第二”等字样对功能和作用基本相同的相同项或相似项进行区分。本领域技术人员可以理解“第一”、“第二”等字样并不对数量和执行次序进行限定,并且“第一”、“第二”等字样也并不限定一定不同。
为了便于理解,以下结合附图对本申请实施例提供的交互方法进行具体介绍。
如图6所示,本申请实施例提供一种交互方法,以第一终端设备为智能音箱,第二终端设备为手机为例进行说明,包括:
601、手机接收用户的第一输入内容,第一输入内容包括第一语音内容和/或第一文本内容。
用户的手机可以安装智能音箱APP客户端,用户可以通过手机上的APP客户端发送第一语音内容和/或第一文本内容给智能音箱。
举例来说,当用户希望远程操作家中的智能音箱时,如图7所示,用户可以在手机的桌面701上点击智能音箱APP的图标702。当手机检测到用户点击桌面701上的智能音箱APP的图标702的操作后,可以启动智能音箱APP,显示如图8所示的图形用户界面(graphical user interface,GUI),该GUI可以称为输入界面801。该输入界面801可以包括语音输入图标802和输入语音的提示信息803“请输入语音”,或者语音的提示信息可以是“请说出您希望执行 的操作”等(图8中未示出),输入语音的提示信息可以是用户点击语音输入图标802后显示的;文本输入部分可以包括文本框804和输入文本的提示信息805“请输入文字”,或者该提示信息可以是“请编辑出您希望执行的操作”(图8中未示出)等。
可选的,用户可以通过手机的其他APP(例如,微信APP)中的智能音箱小程序或公众号来操作智能音箱,或者用户可以通过浏览器登录智能音箱的网页来操作智能音箱,本申请不做限定。
若第一输入内容包括第一语音内容,手机可以通过麦克风拾取用户发出的语音信号(第一语音内容)。示例性的,用户在进行语音输入时,可以点击语音输入图标,此时手机启动麦克风检测用户发出的语音信号(用户说的话),第一语音内容可以为“请打开家里的扫地机器人清扫客厅”或“请清扫客厅”。用户停止说话后预设时间间隔后,手机可以认为语音输入完毕。可选的,手机可以通过提示音或提示信息提示用户语音输入完毕。可选的,手机可以预设用户可以输入的语音长度,例如规定用户输入的语音长度大于1s小于或等于60s。可选的,若用户发出的语音内容的音量过小(例如,小于20dB),可以通过提示音或提示信息提示用户调整音量大小,以便麦克风可以更好的拾取用户的语音内容。
若第一输入内容包括第一文本内容,手机可以通过输入装置(例如触摸屏或键盘)接收用户输入的文本内容(第一文本内容)。例如,第一文本内容可以是用户自己编辑的一段心得或感悟等。或者,第一文本内容可以是用户从浏览器或微信等应用复制的一篇小短文或一首小诗。示例性的,用户在进行文本输入时,可以点击文本框803,文本框中的光标闪烁提示用户输入文本内容,可选的,文本框周围(例如文本框的下方或上方)可以显示输入法(例如拼音输入法或手写输入法)或快捷方式(复制、粘贴等)等,方便用户的操作。
若第一输入内容包括第一语音内容和第一文本内容,手机可以通过麦克风拾取用户发出的语音内容(第一语音内容),并通过输入装置接收用户输入的文本内容(第一文本内容)。例如,第一文本内容可以是用户编辑的一段文字,第一语音内容可以是“请用声情并茂/高昂/低沉的语气朗读这段文字(即第一文本内容)”。用户可以先输入第一语音内容,再输入第一文本内容,或者用户可以先输入第一文本内容再输入第一语音内容,或者用户可以同时输入第一语音内容和第一文本内容,本申请不做限定。
可选的,第一输入内容可以包括图像信息,例如网页截图、火车票、飞机票的订票信息的截图等。举例来说,用户可以通过智能音箱APP输入第一语音内容或第一文本内容:“预定图片中航次的机票”,并输入机票信息的截图。
这样,当用户不在智能音箱附近的时候,可以利用互联网环境,通过手机的麦克风拾取用户发出的语音内容,和/或,通过输入装置接收用户输入的文本内容,使手机与智能音箱实现远程交互,能够使用户远程无障碍操作智能音箱,并通过智能音箱与家中其他成员进行沟通,充分发挥了家中智能音箱的能力。
602、手机向智能音箱发送第一输入内容。
手机可以直接将用户输入的语音内容(第一语音内容)和/或用户输入的文本内容(第一文本内容)发送给智能音箱,无需对第一输入内容进行相应处理,例如进行语义提取或操作指令匹配等操作。
举例来说,当用户对手机说:“请打开家里的扫地机器人清扫客厅”,手机可以直接将“请打开家里的扫地机器人清扫客厅”这句话的音频信息发送给智能音箱,而无需根据“请打开家里的扫地机器人清扫客厅”这句话的语义匹配相应的操作指令。
603、智能音箱从手机接收第一输入内容。
604、智能音箱确定第一输入内容对应的控制指令。
智能音箱可以将第一输入内容上传服务器(例如,云服务器)。若第一输入内容为第一语音内容,云服务器的ASR引擎可以将第一语音内容转化为第一文本信息,云服务器的NLP引擎可以根据第一文本信息的语义得到控制指令。该控制指令可以包括第一文本信息的垂类和槽位。示例性的,若用户说“请帮我订一张后天早上去北京的机票”或者“我想订后天早上的机票去北京”,NLP引擎可以根据关键字做意图识别确定第一文本信息的垂类为“订机票”,并确定该垂类的槽位包括:出发时间为“后天早上”,目的地为“北京”。或者,若用户说“讲一个小笑话”,NLP引擎可以根据关键字做意图识别确定垂类为“讲故事”。并确定该垂类的槽位包括:故事类型为“笑话”,故事长短为“短”。云服务器向智能音箱下发控制指令,智能音箱接收云服务器下发的控制指令。
若第一输入内容为文本内容,云服务器的NLP引擎可以根据文本信息的语义得到控制指令,而后云服务器向智能音箱下发该控制指令,智能音箱接收云服务器下发的控制指令。
605、智能音箱处理第一输入内容对应的控制指令。
第一输入内容对应的控制指令可以是针对智能音箱的,即需要智能音箱根据控制指令做出响应;第一输入内容对应的控制指令也可以是针对第三终端设备(例如,家用的扫地机器人),即智能音箱转发控制指令给第三终端设备,第三终端设备根据控制指令做出响应。
若控制指令是针对智能音箱的,智能音箱执行第一输入内容对应的控制指令。例如,若控制指令为播放(PLAY),智能音箱播放音乐。
若控制指令是针对第三终端设备的,智能音箱可以向第三终端设备发送第一输入内容对应的控制指令,以便第三终端设备执行第一输入内容对应的控制指令,从而控制第三终端设备进行相应操作,例如开机、关机等等。其中,第三终端设备可以是与智能音箱匹配的终端设备,例如,可以是与智能音箱匹配的扫地机器人,空调、冰箱、洗衣机、智能窗帘等等。
举例来说,当用户需要远程操作家中的扫地机器人而手机没有与家中的扫地机器人匹配时,手机没有控制扫地机器人指令,无法直接控制扫地机器人,此时手机可以通过智能音箱控制扫地机器人。例如,用户可以通过手机的智能语音APP客户端输入如下语音内容:“请打开家里的扫地机器人清扫客厅”。智能音箱接收到这条语音内容后,可以将该语音内容上传云服务器,云服务器的ASR引擎可以将该语音内容转化为文本信息,云服务器的NLP引擎可以根据该文本信息的语义得到控制指令“清扫客厅”,云服务器向智能音箱下发控制指令,智能音箱接收云服务器下发的控制指令,并转发给扫地机器人,从而实现对扫地机器人的控制,充分发挥了智能音箱的能力,提高了用户体验。
606、智能音箱向手机发送第一输入内容对应的第一响应消息。
若控制指令是针对智能音箱的,智能音箱执行第一输入内容对应的控制指令后,向手机发送第一响应消息,第一响应消息可以包括云服务转写的控制指令以及智能音箱执行控制指令后的执行结果,以便用户通过第二终端设备输入第一输入内容后,可以获知根据第一输入内容得到的控制指令和针对该控制指令的执行结果,提高了用户体验。举例来说,用户A可以通过手机的APP端发送“播放荷塘月色”的语音内容,智能音箱接收到该语音内容后,通过云服务器确定该语音内容对应的控制指令,根据该控制指令播放相应歌曲,并将播放状态(即第一输入内容对应的第一响应消息)发送到手机APP端。
若控制指令是针对第三终端设备的,智能音箱从第三终端设备接收第一输入内容对应的执行结果后,可以向手机发送第一响应消息,第一响应消息可以包括云服务转写的控制指令以及第三终端设备执行控制指令后的执行结果。
607、手机从智能音箱接收第一输入内容对应的第一响应消息。
608、手机语音播报或在显示屏显示响应消息。
即手机可以语音播报和/或在显示屏上显示智能音箱通过云服务器转写后的控制指令和执行结果(即响应消息)。
如图9所示,假设用户805语音输入:“吃晚饭的时候播放半小时轻音乐,晚饭18:50开始”。手机可以在APP上显示用户805输入了语音内容806,并显示智能音箱807处理后的控制指令808为“18:50开始播放轻音乐30分钟”,以及执行结果809为“18:50开始播放轻音乐”和执行结果810“19:20结束播放轻音乐”。可选的,手机可以在APP上显示用户805输入语音内容的时间信息811,以及手机接收到智能音箱发送的控制指令808的时间信息812、手机接收到执行结果809的时间信息813和手机接收到执行结果810的时间信息814。进一步的,手机还可以将控制指令808、执行结果809和执行结果810语音播报给用户。或者,手机可以不显示控制指令808、执行结果809和执行结果810,仅通过语音播报给用户。例如,当手机检测到用户正在听音乐且手机屏幕为黑屏状态时,手机语音播报智能音箱反馈的响应消息,无需点亮屏幕,可以节省耗电。
另外,该交互方法还可以包括:
609、智能音箱接收第二输入内容,所述第二输入内容包括第二语音内容和/或第二文本内容。
第二输入内容可以是其他用户(与输入第一输入内容的用户不同)在音箱旁输入的语音内容(第二语音内容),或者第二输入内容可以是其他用户通过第四终端设备(例如用户的手机或智能穿戴设备等)发送的第二语音内容和/或第二文本内容。
若智能音箱接收第二输入内容的时刻早于接收第一输入内容的时刻,智能音箱处理第二输入内容对应的控制指令后,处理第一输入内容对应的控制指令;若智能音箱接收第二输入内容的时刻晚于接收第一输入内容的时刻,智能音箱处理第一输入内容对应的控制指令后,处理第二输入内容对应的控制指令。也就是说,智能音箱在处理不同的控制指令时,遵循先来先执行策略。
以下结合具体场景对智能音箱处理不同的控制指令的先后顺序进行说明:
场景1、用户A通过手机的APP端发送“播放音乐”的语音内容,智能音箱接收到语音内容后,通过云服务器确定该语音内容对应的指令,根据该指令播放音乐,并可以将播放状态发送到手机APP端。在播放音乐的过程中,用户B在音箱旁通过语音指令控制暂停播放,预期结果为:智能音箱暂停音乐播放,并可以将暂停状态发送到手机APP端。
场景2、智能音箱正在进行音乐播放,用户B在音箱旁通过语音指令控制暂停,而后用户A通过手机的APP端发送“播放音乐”的语音内容,智能音箱接收到语音内容后,通过云服务器确定该语音内容对应的指令,预期结果为:智能音箱暂停音乐播放后又继续播放音乐,并可以将播放状态发送到手机APP端。
场景3、用户A通过手机的APP端发送“播放音乐”的语音内容,智能音箱接收到语音内容后,通过云服务器确定该语音内容对应的指令,根据该指令播放音乐,在播放音乐的过程中(例如,播放了一首歌的1/3后),用户B在音箱旁通过语音控制音乐播放(用户B请求播放的内容与用户A请求播放的内容正好一致,例如都要求播放“荷塘月色”),预期结果为:智能音箱继续播放“荷塘月色”,无需重头开始播放“荷塘月色”,能够满足用户需求且避免重复执行相同的指令。
需要说明的是,步骤609和步骤603之间没有必然的执行先后顺序,步骤609可以在 步骤603之前执行,也可以在步骤603之后执行,也可以和步骤603同时执行,本实施例对此不作具体限定。
可以理解的是,智能音箱处理第二输入内容对应的控制指令后,可以向手机发送第二输入内容对应的第二响应消息。手机从智能音箱接收第二响应消息,第二响应消息对应第二输入内容,手机语音播报或在显示屏显示第二响应消息。这样一来,若其他用户(与输入第一输入内容的用户不同)在音箱旁输入语音内容(第二语音内容),或者其他用户通过第四终端设备(例如用户的手机或智能穿戴设备等)输入第二输入内容,第一终端设备的用户可以获知其他用户对第二终端设备的操作(包括第二输入内容对应的控制指令和针对该控制指令的执行结果),提高了用户对第二终端设备的掌控能力,从而提高了用户体验。
基于本申请实施例提供的方法,第二终端设备(例如,手机)可以从用户接收第一输入内容,并向第一终端设备(例如,智能语音交互设备)发送第一输入内容。第一终端设备从第二终端设备接收第一输入内容,通过服务器确定第一输入内容对应的控制指令,并处理该控制指令。这样一来,可以由第一终端设备基于用户通过第二终端设备远程输入的内容(即第一输入内容)确定控制指令,而非用户根据手机APP上预制好的有限的指令进行操作,增强了用户与第一终端设备远程交互的灵活性,充分发挥第一终端设备的能力。
上述本申请提供的实施例中,分别从第一终端设备、第二终端设备以及第一终端设备和第二终端设备之间交互的角度对本申请实施例提供的方法进行了介绍。为了实现上述本申请实施例提供的方法中的各功能,第一终端设备和第二终端设备可以包括硬件结构和/或软件模块,以硬件结构、软件模块、或硬件结构加软件模块的形式来实现上述各功能。上述各功能中的某个功能以硬件结构、软件模块、还是硬件结构加软件模块的方式来执行,取决于技术方案的特定应用和设计约束条件。
在采用对应各个功能划分各个功能模块的情况下,图10示出了上述实施例中所涉及的装置10的一种可能的结构示意图,该装置可以为第一终端设备,该第一终端设备包括:接收单元1001、确定单元1002和处理单元1002。在本申请实施例中,接收单元1001,用于从第二终端设备接收第一输入内容,第一输入内容包括第一语音内容和/或第一文本内容;确定单元1002,用于确定第一输入内容对应的控制指令;处理单元1003,用于处理第一输入内容对应的控制指令。可选的,第一终端设备还可以包括发送单元1004(图10中未示出),用于向第二终端设备发送第一输入内容对应的第一响应消息。
在图6所示的方法实施例中,接收单元1001用于支持第一终端设备执行图6中的过程603;确定单元1002用于支持第一终端设备执行图6中的过程604;处理单元1003用于支持第一终端设备执行图6中的过程605;发送单元1004,用于支持第一终端设备执行图6中的过程606。其中,上述方法实施例涉及的各步骤的所有相关内容均可以援引到对应功能模块的功能描述,在此不再赘述。
在采用对应各个功能划分各个功能模块的情况下,图11示出了上述实施例中所涉及的装置11的一种可能的结构示意图,该装置可以为第二终端设备,该第二终端设备包括:接收单元1101、发送单元1102和处理单元1103。在本申请实施例中,接收单元1101,用于从用户接收第一输入内容,第一输入内容包括第一语音内容和/或第一文本内容;发送单元1102,用于向第一终端设备发送第一输入内容;接收单元1101,还用于从第一终端设备接收第一输入内容对应的第一响应消息;处理单元1103,用于语音播报或在显示屏显示第一响应消息。
在图6所示的方法实施例中,接收单元1101用于支持第二终端设备执行图6中的过程601和607;发送单元1102用于支持第二终端设备执行图6中的过程602;处理单元1103用 于支持第二终端设备执行图6中的过程608。其中,上述方法实施例涉及的各步骤的所有相关内容均可以援引到对应功能模块的功能描述,在此不再赘述。
本申请实施例中对模块的划分是示意性的,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,另外,在本申请各个实施例中的各功能模块可以集成在一个处理器中,也可以是单独物理存在,也可以两个或两个以上模块集成在一个模块中。上述集成的模块既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。示例性地,在本申请实施例中,接收单元和发送单元可以集成至收发单元中。
本申请实施例提供的方法中,可以全部或部分地通过软件、硬件、固件或者其任意组合来实现。当使用软件实现时,可以全部或部分地以计算机程序产品的形式实现。所述计算机程序产品包括一个或多个计算机指令。在计算机上加载和执行所述计算机程序指令时,全部或部分地产生按照本发明实施例所述的流程或功能。所述计算机可以是通用计算机、专用计算机、计算机网络、网络设备、用户设备或者其他可编程装置。所述计算机指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一个计算机可读存储介质传输,例如,所述计算机指令可以从一个网站站点、计算机、服务器或数据中心通过有线(例如同轴电缆、光纤、数字用户线(digital subscriber line,DSL))或无线(例如红外、无线、微波等)方式向另一个网站站点、计算机、服务器或数据中心进行传输。所述计算机可读存储介质可以是计算机可以存取的任何可用介质或者是包含一个或多个可用介质集成的服务器、数据中心等数据存储设备。所述可用介质可以是磁性介质(例如,软盘、硬盘、磁带)、光介质(例如,数字视频光盘(digital video disc,DVD))、或者半导体介质(例如,固态硬盘(solid state drives,SSD))等。
显然,本领域的技术人员可以对本申请实施例进行各种改动和变型而不脱离本申请的精神和范围。这样,倘若本申请实施例的这些修改和变型属于本申请权利要求及其等同技术的范围之内,则本申请也意图包含这些改动和变型在内。
Claims (20)
- 一种交互方法,其特征在于,包括:第一终端设备从第二终端设备接收第一输入内容,所述第一输入内容包括第一语音内容和/或第一文本内容;所述第一终端设备确定所述第一输入内容对应的控制指令;所述第一终端设备处理所述第一输入内容对应的所述控制指令。
- 根据权利要求1所述的交互方法,其特征在于,所述第一终端设备确定所述第一输入内容对应的控制指令包括:所述第一终端设备向服务器发送所述第一输入内容;所述第一终端设备从所述服务器接收所述第一输入内容对应的所述控制指令。
- 根据权利要求1或2所述的交互方法,其特征在于,所述第一终端设备处理所述第一输入内容对应的所述控制指令包括:所述第一终端设备执行所述第一输入内容对应的所述控制指令;或者所述第一终端设备向第三终端设备发送所述第一输入内容对应的所述控制指令。
- 根据权利要求1-3任一项所述的交互方法,其特征在于,所述方法还包括:所述第一终端设备向所述第二终端设备发送所述第一输入内容对应的第一响应消息。
- 根据权利要求1-4任一项所述的交互方法,其特征在于,所述第一终端设备处理所述第一输入内容对应的所述控制指令包括:所述第一终端设备接收第二输入内容,所述第二输入内容包括第二语音内容和/或第二文本内容;若所述第一终端设备接收所述第二输入内容的时刻早于接收所述第一输入内容的时刻,所述第一终端设备处理所述第二输入内容对应的控制指令后,处理所述第一输入内容对应的控制指令。
- 一种交互方法,其特征在于,包括:第二终端设备从用户接收第一输入内容,所述第一输入内容包括第一语音内容和/或第一文本内容;所述第二终端设备向第一终端设备发送所述第一输入内容;所述第二终端设备从所述第一终端设备接收所述第一输入内容对应的第一响应消息;所述第二终端设备语音播报或在显示屏显示所述第一响应消息。
- 根据权利要求6所述的交互方法,其特征在于,所述方法还包括:所述第二终端设备从所述第一终端设备接收第二响应消息,所述第二响应消息对应第二输入内容;所述第二终端设备语音播报或在显示屏显示所述第二响应消息。
- 一种第一终端设备,其特征在于,包括:接收单元,用于从第二终端设备接收第一输入内容,所述第一输入内容包括第一语音内容和/或第一文本内容;确定单元,用于确定所述第一输入内容对应的控制指令;处理单元,用于处理所述第一输入内容对应的所述控制指令。
- 根据权利要求8所述的第一终端设备,其特征在于,所述确定单元用于:通过发送单元向服务器发送所述第一输入内容;通过所述接收单元从所述服务器接收所述第一输入内容对应的所述控制指令。
- 根据权利要求8或9所述的第一终端设备,其特征在于,所述处理单元用于:执行所述第一输入内容对应的所述控制指令;或者通过所述发送单元向第三终端设备发送所述第一输入内容对应的所述控制指令。
- 根据权利要求8-10任一项所述的第一终端设备,其特征在于,所述发送单元还用于:向所述第二终端设备发送所述第一输入内容对应的第一响应消息。
- 根据权利要求8-11任一项所述的第一终端设备,其特征在于,所述处理单元用于:通过所述接收单元接收第二输入内容,所述第二输入内容包括第二语音内容和/或第二文本内容;若所述第一终端设备接收所述第二输入内容的时刻早于接收所述第一输入内容的时刻,处理所述第二输入内容对应的控制指令后,处理所述第一输入内容对应的控制指令。
- 一种第二终端设备,其特征在于,包括:接收单元,用于从用户接收第一输入内容,所述第一输入内容包括第一语音内容和/或第一文本内容;发送单元,用于向第一终端设备发送所述第一输入内容;所述接收单元,还用于从所述第一终端设备接收所述第一输入内容对应的第一响应消息;处理单元,用于语音播报或在显示屏显示所述第一响应消息。
- 根据权利要求13所述的第二终端设备,其特征在于,所述接收单元还用于:从所述第一终端设备接收第二响应消息,所述第二响应消息对应第二输入内容;所述处理单元,还用于语音播报或在显示屏显示所述第二响应消息。
- 一种第一终端设备,其特征在于,所述第一终端设备包括处理器和存储器;所述存储器用于存储计算机执行指令,当所述第一终端设备运行时,所述处理器执行所述存储器存储的所述计算机执行指令,以使所述第一终端设备执行如权利要求1-5中任一项所述的交互方法。
- 一种第二终端设备,其特征在于,所述第二终端设备包括处理器和存储器;所述存储器用于存储计算机执行指令,当所述第二终端设备运行时,所述处理器执行所述存储器存储的所述计算机执行指令,以使所述第二终端设备执行如权利要求6或7所述的交互方法。
- 一种计算机可读存储介质,其特征在于,包括指令,当其在计算机上运行时,使得计算机执行权利要求1-5中任一项所述的交互方法。
- 一种计算机可读存储介质,其特征在于,包括指令,当其在计算机上运行时,使得计算机执行权利要求6或7所述的交互方法。
- 一种芯片系统,其特征在于,包括处理器和存储器,所述处理器执行所述存储器存储的计算机执行指令,以实现如权利要求1-5中任一项所述的交互方法。
- 一种芯片系统,其特征在于,包括处理器和存储器,所述处理器执行所述存储器存储的计算机执行指令,以实现如权利要求6或7所述的交互方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910472665.5 | 2019-05-31 | ||
| CN201910472665.5A CN112017652A (zh) | 2019-05-31 | 2019-05-31 | 一种交互方法和终端设备 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020239013A1 true WO2020239013A1 (zh) | 2020-12-03 |
Family
ID=73506179
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/092888 Ceased WO2020239013A1 (zh) | 2019-05-31 | 2020-05-28 | 一种交互方法和终端设备 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN112017652A (zh) |
| WO (1) | WO2020239013A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117133284A (zh) * | 2023-04-06 | 2023-11-28 | 荣耀终端有限公司 | 一种语音交互方法及电子设备 |
| WO2023231963A1 (zh) * | 2022-06-01 | 2023-12-07 | 华为技术有限公司 | 一种设备控制方法及电子设备 |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115497470A (zh) * | 2021-06-18 | 2022-12-20 | 华为技术有限公司 | 跨设备的对话业务接续方法、系统、电子设备和存储介质 |
| CN113450792A (zh) * | 2021-06-22 | 2021-09-28 | 海信视像科技股份有限公司 | 终端设备的语音控制方法、终端设备及服务器 |
| CN117882130A (zh) * | 2021-06-22 | 2024-04-12 | 海信视像科技股份有限公司 | 一种进行语音控制的终端设备及服务器 |
| CN113393839B (zh) * | 2021-08-16 | 2021-11-12 | 成都极米科技股份有限公司 | 智能终端控制方法、存储介质及智能终端 |
| CN113835350A (zh) * | 2021-09-26 | 2021-12-24 | 北京无线电测量研究所 | 一种智能方舱 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1655206A (zh) * | 2004-02-13 | 2005-08-17 | 邓清辉 | 一种万能家用电器语音声控和电话远程控制装置 |
| CN101272418A (zh) * | 2008-03-25 | 2008-09-24 | 宇龙计算机通信科技(深圳)有限公司 | 一种远程控制通信终端的方法和通信终端 |
| CN102170617A (zh) * | 2011-04-07 | 2011-08-31 | 中兴通讯股份有限公司 | 移动终端及其远程控制方法 |
| CN102427418A (zh) * | 2011-12-09 | 2012-04-25 | 福州海景科技开发有限公司 | 基于语音识别的智能家居的系统 |
| JP2014164241A (ja) * | 2013-02-27 | 2014-09-08 | Nippon Telegraph & Telephone East Corp | 中継システム、中継方法及びプログラム |
| CN105991825A (zh) * | 2015-02-04 | 2016-10-05 | 中兴通讯股份有限公司 | 一种语音控制方法、装置及系统 |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106685772A (zh) * | 2016-12-23 | 2017-05-17 | 北京奇虎科技有限公司 | 一种智能音箱、智能家居系统及其实现方法 |
| CN108235813A (zh) * | 2017-02-28 | 2018-06-29 | 华为技术有限公司 | 一种语音输入的方法和相关设备 |
| CN107835444B (zh) * | 2017-11-16 | 2019-04-23 | 百度在线网络技术(北京)有限公司 | 信息交互方法、装置、音频终端及计算机可读存储介质 |
| CN109143879A (zh) * | 2018-08-10 | 2019-01-04 | 珠海格力电器股份有限公司 | 一种以空调为中心控制家电的方法 |
| CN109408168B (zh) * | 2018-09-25 | 2021-11-19 | 维沃移动通信有限公司 | 一种远程交互方法和终端设备 |
| CN109243444B (zh) * | 2018-09-30 | 2021-06-01 | 百度在线网络技术(北京)有限公司 | 语音交互方法、设备及计算机可读存储介质 |
-
2019
- 2019-05-31 CN CN201910472665.5A patent/CN112017652A/zh active Pending
-
2020
- 2020-05-28 WO PCT/CN2020/092888 patent/WO2020239013A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1655206A (zh) * | 2004-02-13 | 2005-08-17 | 邓清辉 | 一种万能家用电器语音声控和电话远程控制装置 |
| CN101272418A (zh) * | 2008-03-25 | 2008-09-24 | 宇龙计算机通信科技(深圳)有限公司 | 一种远程控制通信终端的方法和通信终端 |
| CN102170617A (zh) * | 2011-04-07 | 2011-08-31 | 中兴通讯股份有限公司 | 移动终端及其远程控制方法 |
| CN102427418A (zh) * | 2011-12-09 | 2012-04-25 | 福州海景科技开发有限公司 | 基于语音识别的智能家居的系统 |
| JP2014164241A (ja) * | 2013-02-27 | 2014-09-08 | Nippon Telegraph & Telephone East Corp | 中継システム、中継方法及びプログラム |
| CN105991825A (zh) * | 2015-02-04 | 2016-10-05 | 中兴通讯股份有限公司 | 一种语音控制方法、装置及系统 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2023231963A1 (zh) * | 2022-06-01 | 2023-12-07 | 华为技术有限公司 | 一种设备控制方法及电子设备 |
| CN117133284A (zh) * | 2023-04-06 | 2023-11-28 | 荣耀终端有限公司 | 一种语音交互方法及电子设备 |
| CN117133284B (zh) * | 2023-04-06 | 2025-01-03 | 荣耀终端有限公司 | 一种语音交互方法及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN112017652A (zh) | 2020-12-01 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020239013A1 (zh) | 一种交互方法和终端设备 | |
| US12019864B2 (en) | Multimedia data playing method and electronic device | |
| US11893359B2 (en) | Speech translation method and terminal when translated speech of two users are obtained at the same time | |
| WO2020192456A1 (zh) | 一种语音交互方法及电子设备 | |
| WO2021185244A1 (zh) | 一种设备交互的方法和电子设备 | |
| CN108695999A (zh) | 数据投屏方法、装置、存储介质及电子设备 | |
| WO2020244492A1 (zh) | 一种投屏显示方法及电子设备 | |
| CN114924682A (zh) | 一种内容接续方法及电子设备 | |
| WO2021204098A1 (zh) | 语音交互方法及电子设备 | |
| US12192885B2 (en) | Method for accessing network by smart home device and related device | |
| CN114996168B (zh) | 一种多设备协同测试方法、测试设备及可读存储介质 | |
| WO2021104114A1 (zh) | 一种提供无线保真WiFi网络接入服务的方法及电子设备 | |
| WO2023071502A1 (zh) | 音量控制方法、装置及电子设备 | |
| CN115842692A (zh) | 一种电子设备的控制方法、介质和电子设备 | |
| US20240103695A1 (en) | Message Reply Method and Apparatus | |
| CN114664306B (zh) | 一种编辑文本的方法、电子设备和系统 | |
| CN116795753A (zh) | 音频数据的传输处理的方法及电子设备 | |
| CN114120987B (zh) | 一种语音唤醒方法、电子设备及芯片系统 | |
| WO2023216922A1 (zh) | 目标设备选择的识别方法、终端设备、系统和存储介质 | |
| CN110737765A (zh) | 多轮对话的对话数据处理方法及相关装置 | |
| CN113950037B (zh) | 一种音频播放方法及终端设备 | |
| CN115412387B (zh) | 一种音频播放方法、系统及电子设备 | |
| WO2022135183A1 (zh) | 设备控制方法、装置和电子设备 | |
| US12619349B2 (en) | Multimedia data playing method and electronic device | |
| CN120434323B (zh) | 通话方法及相关设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20815563 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20815563 Country of ref document: EP Kind code of ref document: A1 |