WO2020048490A1 - 实时语音信息的交互方法、装置、电子设备及存储介质 - Google Patents
实时语音信息的交互方法、装置、电子设备及存储介质 Download PDFInfo
- Publication number
- WO2020048490A1 WO2020048490A1 PCT/CN2019/104421 CN2019104421W WO2020048490A1 WO 2020048490 A1 WO2020048490 A1 WO 2020048490A1 CN 2019104421 W CN2019104421 W CN 2019104421W WO 2020048490 A1 WO2020048490 A1 WO 2020048490A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- voice data
- voice
- electronic device
- data
- audio
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/41—Structure of client; Structure of client peripherals
- H04N21/422—Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS]
- H04N21/42203—Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS] sound input device, e.g. microphone
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/04—Real-time or near real-time messaging, e.g. instant messaging [IM]
- H04L51/046—Interoperability with other network applications or services
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/06—Message adaptation to terminal or network requirements
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/07—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail characterised by the inclusion of specific contents
- H04L51/10—Multimedia information
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/21—Server components or server architectures
- H04N21/218—Source of audio or video content, e.g. local disk arrays
- H04N21/2187—Live feed
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/27—Server based end-user applications
- H04N21/274—Storing end-user multimedia data in response to end-user request, e.g. network recorder
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
- H04N21/4396—Processing of audio elementary streams by muting the audio signal
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/472—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
- H04N21/4722—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content for requesting additional data associated with the content
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/475—End-user interface for inputting end-user data, e.g. personal identification number [PIN], preference data
- H04N21/4758—End-user interface for inputting end-user data, e.g. personal identification number [PIN], preference data for providing answers, e.g. voting
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/478—Supplemental services, e.g. displaying phone caller identification, shopping application
- H04N21/4788—Supplemental services, e.g. displaying phone caller identification, shopping application communicating with other users, e.g. chatting
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/60—Network structure or processes for video distribution between server and client or between remote clients; Control signalling between clients, server and network components; Transmission of management data between server and client, e.g. sending from server to client commands for recording incoming content stream; Communication details between server and client
- H04N21/65—Transmission of management data between client and server
- H04N21/658—Transmission by the client directed to the server
- H04N21/6582—Data stored in the client, e.g. viewing habits, hardware capabilities, credit card number
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/81—Monomedia components thereof
- H04N21/8106—Monomedia components thereof involving special audio data, e.g. different tracks for different languages
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/85—Assembly of content; Generation of multimedia applications
- H04N21/858—Linking data to content, e.g. by linking an URL to a video object, by creating a hotspot
Definitions
- the present application relates to the field of Internet technologies, and in particular, to a method, a device, an electronic device, and a storage medium for real-time voice information interaction.
- the live webcast is a kind of one-to-many communication centered on the audiovisual expression of the anchor.
- the main mode of interactive communication scenes it is necessary to ensure an equal relationship between audiences. In this mode, the audience can only express through words.
- the inventors realized that the audience's level was uneven, and some people's text input speed was slow, or they could not even enter text. As a result, many people could not effectively express their opinions, and the audience experience was poor. Conducive to expanding the reach of online broadcast audience.
- the present application provides a method, device, electronic device and storage medium for real-time voice information interaction.
- An embodiment of the present application provides a real-time voice information interaction method, which is applied to an electronic device.
- the interaction method includes:
- a first electronic device so that the first electronic device displays the voice data in a list form, so that a user of the first electronic device selects and plays the voice data in the list;
- the voice data in the sending queue is displayed locally in a list form, and the sending status of the voice data is displayed.
- the technical solution provided by the embodiments of the present application may include the following beneficial effects: It enables users to upload voice data in a voice manner, and also enables the voice data to fully function as text in a real-time information interaction system, which greatly facilitates text input. Users who are slow or unable to type text improve the user experience.
- Fig. 1 is a flow chart showing a method for interacting with real-time voice information according to an exemplary embodiment
- Fig. 2 is a flow chart showing another method for interacting with real-time voice information according to an exemplary embodiment
- Fig. 3 is a structural block diagram of a device for interacting with real-time voice information according to an exemplary embodiment
- Fig. 4 is a structural block diagram of another real-time voice information interaction device according to an exemplary embodiment
- Fig. 5 is a flow chart showing still another method for interacting with real-time voice information according to an exemplary embodiment
- Fig. 6 is a structural block diagram of another real-time voice information interaction device according to an exemplary embodiment
- Fig. 7 is a flow chart showing still another method for interacting with real-time voice information according to an exemplary embodiment
- Fig. 8 is a flow chart showing still another method for interacting with real-time voice information according to an exemplary embodiment
- Fig. 9 is a structural block diagram of still another real-time voice information interaction device according to an exemplary embodiment.
- Fig. 10 is a structural block diagram of another real-time voice information interaction device according to an exemplary embodiment
- Fig. 11 is a structural block diagram of a server according to an exemplary embodiment
- Fig. 12 is a structural block diagram of an electronic device according to an exemplary embodiment
- Fig. 13 is a structural block diagram of another electronic device according to an exemplary embodiment.
- Fig. 1 is a flow chart showing a method for interacting with real-time voice information according to an exemplary embodiment.
- this specific interaction method is applied to an electronic device. Specifically, the interaction method in this specific embodiment is applied to a viewer of the network live broadcast system.
- the interaction method includes the following steps.
- step S11 the input voice is recorded to obtain at least one voice data.
- the input sound is recorded and converted to obtain at least one voice data, that is, a digital voice signal.
- the corresponding voice signal is obtained from a recording device connected to the audience end and converted to obtain voice data.
- a voice signal sent by the user is obtained, and the voice signal is converted into a piece of audio data every preset duration, that is, the audio data is converted from the voice signal of the preset duration.
- the preset duration can be selected as 20 milliseconds.
- the duration covered by the voice data is the duration of the current recording request. For a specific viewer, it may be the length of time the viewer user presses the recording button.
- the playback volume of the audio and video played by the viewer is reduced to 0, that is, the audio and video are controlled to be muted, so that noise can not interfere with the recording, and a more pure voice can be obtained. data.
- step S12 the voice data is sequentially transmitted to a server connected to the electronic device.
- the specific embodiment can be applied to a network live broadcast system, after the audience terminal obtains the voice data, it is sequentially sent to the server through a long connection with the server, so that the server stores the voice data. After receiving the voice data, the server pushes it to the first electronic device corresponding to the electronic device that recorded the voice data. Since the voice data recorded here is the viewer end of the network live broadcast system, the server corresponding to the viewer end An electronic device is the anchor of the network live broadcast system.
- the server After receiving the voice data, the server sends the voice data to the first electronic device, that is, to the anchor end, and causes the anchor end to display the above-mentioned multiple voice data in a list form.
- the anchor user at the anchor end can select and play the voice data.
- the so-called selection playback means selecting the corresponding voice data from the list for playback.
- step S13 at least one voice data is displayed locally in a list form.
- the recorded voice data is not limited to one. Therefore, for the convenience of users to view, multiple voice data are displayed in a list here.
- a specific display manner may be to display multiple icons only in a list form, and each icon corresponds to one voice data.
- the status of the corresponding voice data is also displayed in the preset position of each voice data. For example, when a voice data is being uploaded to the server, the status of the voice data is displayed as uploading; After the upload is completed, the status of the voice data is displayed as being sent.
- the embodiment of the present application provides a real-time voice information interaction method.
- the interaction method is applied to a real-time information interaction system.
- the interaction method specifically responds to a user's recording request and records the input sound. And convert to obtain at least one voice data, and at least one voice data is stored in a sending queue in the form of a queue; the voice data in the sending queue is sequentially sent to the server of the real-time information interaction system; the sending queue is displayed locally in a list Voice data and display the sending status of voice data.
- the purpose of this operation is to respond to the delete request when the user needs to delete and send a delete request, and delete the delete request.
- the corresponding voice data is deleted. This prevents unsatisfactory voice data from being pushed to other users.
- a deletion control instruction is sent to the server according to the deletion request at this time to control the server to delete the corresponding voice data.
- Fig. 2 is a flow chart showing another method for interacting with real-time voice information according to an exemplary embodiment.
- this specific interaction method is applied to a real-time information interaction system, and the real-time information interaction system may be a network live broadcast system in practical applications. Therefore, specifically, the interaction method in this specific embodiment is applied to the network On the audience side of the live broadcast system, the interaction method includes the following steps.
- step S11 the input voice is recorded to obtain at least one voice data.
- This step is basically the same as that in the previous specific implementation manner, and is not repeated here.
- step S12 the voice data is sequentially transmitted to a server connected to the electronic device.
- This step is basically the same as that in the previous specific implementation manner, and is not repeated here.
- step S13 at least one voice data is displayed locally in a list form.
- the voice data displayed in the form of a list includes not only the voice data recorded locally, but also the server-supplied data that is equal to that used to record voice data locally.
- Voice data recorded by the second electronic device at the location includes not only the voice data recorded by the local audience, but also the voice data recorded by other audiences.
- the equality position here is not complete equality, but actually refers to the status of the basic operation method is equivalent, it has the unequal content of the priority method, for example, for users with higher activity, it has higher priority.
- step S14 the audio and video data pushed by the server is received.
- the audio and video data includes data from the electronic device corresponding to the local voice data.
- the audio data and video data recorded by the first electronic device also come from voice data recorded by a second electronic device having an equal position with the local electronic device.
- the audio and video data comes from the anchor and other viewers of the system, the audio and video data recorded by the anchor user from the anchor, and the voice data from other viewers is uploaded to the server Then, some or all of the voice data that the anchor user chooses to play.
- step S15 the received audio and video data is played locally.
- the audio and video data pushed by the server are played on the viewer.
- the audio and video data includes the audio data and video data recorded by the anchor, and also includes the voice data sent by other viewers.
- step S16 the id of the audio / video data being played is detected.
- the id may match multiple voice data displayed in the list, that is, from the same electronic device.
- step S17 the playback state of the voice data corresponding to the id is displayed.
- the voice data displayed in the list matches the id of the audio and video data that is being played, the voice data is displayed as being in the playing state in the list, so that the user can determine which audio and video data is being played by that voice data. Play at the same time, in order to perform corresponding operations on it, such as playing again or looping.
- step S18 the voice data is controlled to be played again or repeatedly.
- the loop playback instruction is used to control the voice data to be played again. It can also be played repeatedly for an unlimited or limited number of times, so that the user can Know exactly what the corresponding voice data is carrying.
- users can upload voice data in a voice manner, and can also make the voice data fully function as a text in the real-time information interaction system, which greatly facilitates users who have slow or slow text input, thus Improved user experience. It also enables users to get a more advanced experience.
- Fig. 3 is a structural block diagram of a device for interacting with real-time voice information according to an exemplary embodiment.
- the specific interactive device is applied to an electronic device.
- the interactive device in this specific embodiment is applied to the viewer of the network live broadcast system.
- the interactive device includes a voice recording module 10 and a voice sending module 20. And the first display module 30.
- the voice recording module 10 is configured to record the input sound to obtain at least one voice data.
- the input sound is recorded and converted to obtain at least one voice data, that is, a digital voice signal.
- the corresponding voice signal is obtained from a recording device connected to the audience end and converted to obtain voice data.
- the module includes a recording control unit and a data collection unit.
- the recording control unit is configured to obtain a voice signal sent by the user when the user sends a recording request, and convert the voice signal into an audio data every preset time period, that is, the audio data is converted by the voice signal of the preset time period.
- the preset time can be selected as 20 milliseconds in specific practice.
- the data collection unit is configured to collect multiple pieces of audio data generated by each recording and synthesize them into a single voice data file.
- the duration of the voice data coverage is the duration of the current recording request. . For a specific viewer, it may be the length of time the viewer user presses the recording button.
- the module also includes a mute control unit, which is configured to reduce the playback volume of the audio and video played by the viewer when the user sends a recording request to 0, that is, to control the audio and video to remain silent. This can prevent noise from interfering with the recording, thereby obtaining more pure voice data.
- a mute control unit configured to reduce the playback volume of the audio and video played by the viewer when the user sends a recording request to 0, that is, to control the audio and video to remain silent. This can prevent noise from interfering with the recording, thereby obtaining more pure voice data.
- the voice transmitting module 20 is configured to sequentially transmit voice data to a server connected to the electronic device.
- the specific embodiment can be applied to a network live broadcast system, after the audience terminal obtains the voice data, it is sequentially sent to the server through a long connection with the server, so that the server stores the voice data. After receiving the voice data, the server pushes it to the first electronic device corresponding to the electronic device that recorded the voice data. Since the voice data recorded here is the viewer end of the network live broadcast system, the server corresponding to the viewer end An electronic device is the anchor of the network live broadcast system.
- the server After receiving the voice data, the server sends the voice data to the first electronic device, that is, to the anchor end, and causes the anchor end to display the above-mentioned multiple voice data in a list form.
- the anchor user at the anchor end can select and play the voice data.
- the so-called selection playback means selecting the corresponding voice data from the list for playback.
- the first display module 30 is configured to display at least one voice data locally in a list form.
- the recorded voice data is not limited to one. Therefore, for the convenience of users to view, multiple voice data are displayed in a list here.
- a specific display manner may be to display multiple icons only in a list form, and each icon corresponds to one voice data.
- the status of the corresponding voice data is also displayed in the preset position of each voice data. For example, when a voice data is being uploaded to the server, the status of the voice data is displayed as uploading; After the upload is completed, the status of the voice data is displayed as being sent.
- the embodiment of the present application provides a real-time voice information interaction device.
- the interaction device is applied to a real-time information interaction system.
- the interaction device specifically responds to a user's recording request and records the input sound. And convert to obtain at least one piece of voice data, and at least one piece of voice data is stored in a sending queue in the form of a queue; the voice data in the sending queue is sequentially sent to the server of the real-time information interaction system; the sending queue is displayed locally in the form of a list Voice data and display the sending status of voice data.
- a first deletion module (not shown) may be further included.
- the first deletion module is configured to delete the corresponding voice data according to the deletion request of the user.
- the purpose of this operation is to respond to the delete request when the user needs to delete and issue a delete request, and respond to the request.
- Voice data is deleted. This prevents unsatisfactory voice data from being pushed to other users.
- a deletion control instruction is sent to the server according to the deletion request at this time to control the server to delete the corresponding voice data.
- Fig. 4 is a structural block diagram of another real-time voice information interaction device according to an exemplary embodiment.
- the specific interaction method is applied to an electronic device.
- the interaction device in this specific embodiment is applied to the viewer of the network live broadcast system.
- the interaction device is added compared to the previous specific embodiment.
- the first display module is further configured to display at least one voice data locally in a list form.
- the voice data displayed in a list form includes not only the voice data recorded locally, but also the second data that is pushed by the server from an equal position with the electronic device used to record voice data locally.
- Voice data recorded by an electronic device For a webcast system, the list includes not only the voice data recorded by the local audience, but also the voice data recorded by other audiences.
- the audio and video receiving module 40 is configured to receive audio and video data pushed by the server.
- the audio and video data includes data from the electronic device corresponding to the local voice data.
- the audio data and video data recorded by the first electronic device also come from voice data recorded by a second electronic device having an equal position with the local electronic device.
- the audio and video data comes from the anchor and other viewers of the system, the audio and video data recorded by the anchor user from the anchor, and the voice data from other viewers is uploaded to the server Then, some or all of the voice data that the anchor user chooses to play.
- the audio and video playback module 50 is configured to locally play the received audio and video data.
- the audio and video data pushed by the server are played on the viewer.
- the audio and video data includes the audio data and video data recorded by the anchor, and also includes the voice data sent by other viewers.
- the id detection module 60 is configured to detect the id of the audio and video data being played.
- the id may match multiple voice data displayed in the list, that is, from the same electronic device.
- the status display module 70 is configured to display the playback status of the voice data corresponding to the above id.
- the voice data displayed in the list matches the id of the audio and video data that is being played, the voice data is displayed as being in the playing state in the list, so that the user can determine which audio and video data is being played by that voice data Play at the same time, in order to perform corresponding operations on it, such as playing again or looping.
- the loop playback module 80 is configured to control the voice data to be played again or looped.
- the loop playback instruction is used to control the voice data to be played again. It can also be played repeatedly for an unlimited or limited number of times, so that the user can Know exactly what the corresponding voice data is carrying.
- users can upload voice data in a voice manner, and can also make the voice data fully function as a text in the real-time information interaction system, which greatly facilitates users who have slow or slow text input, thus Improved user experience. It also enables users to get a more advanced experience.
- Fig. 5 is a flow chart showing still another method for interacting with real-time voice information according to an exemplary embodiment.
- the interaction method provided in this embodiment is applied to a server of a real-time information interaction system.
- the server is respectively connected to a host and multiple viewers of the network live broadcast system.
- the interaction method specifically includes steps:
- step S21 the voice data sent by the electronic device connected to the server is received.
- the electronic device connected to the server is the viewer. After recording and uploading the voice data at the viewer, the voice data is received in a queue.
- step S22 an id is added to the voice data according to the number of the device transmitting the voice data.
- the device number of the hardware device sending the voice data is detected, and an id is edited according to the detected device number, and the id is added to the corresponding voice data.
- step S23 voice messages are sent to the first electronic device and the second electronic device, respectively.
- the first electronic device corresponds to an electronic device that sends corresponding voice data
- the second electronic device has an equal position with the electronic device that sends corresponding voice data.
- the first electronic device is the anchor
- the second electronic device is the other viewer.
- the voice message sent to the first electronic device also includes the sender information, duration, and id of the voice data, so that the user of the first electronic device, that is, the anchor user of the anchor, can select the voice message to select playback and voice message.
- the voice information sent to the second electronic device is a voice message corresponding to the voice data selected to be played by the anchor user.
- the voice data stored on the display of other electronic devices connected to the server can be made to enable the user to select playback and push the played voice data to other electronic devices.
- Fig. 6 is a structural block diagram of another real-time voice information interaction device according to an exemplary embodiment.
- the interactive device provided in this embodiment is applied to a server of a real-time information interaction system.
- the server is connected to a host and multiple viewers of the network live broadcast system respectively.
- the interaction device specifically includes a data receiving module 110, an id addition module 120, and a message pushing module 130.
- the data receiving module 110 is configured to receive voice data sent by an electronic device connected to the server.
- the electronic device connected to the server is the viewer. After recording and uploading the voice data at the viewer, the voice data is received in a queue.
- the id adding module 120 is configured to add an id to the voice data according to the device number that sends the voice data.
- the device number of the hardware device sending the voice data is detected, and an id is edited according to the detected device number, and the id is added to the corresponding voice data.
- the message push module 130 is configured to send voice messages to the first electronic device and the second electronic device, respectively.
- the first electronic device corresponds to an electronic device that sends corresponding voice data
- the second electronic device has an equal position with the electronic device that sends corresponding voice data.
- the first electronic device is the anchor
- the second electronic device is the other viewer.
- the voice message sent to the first electronic device also includes the sender information, duration, and id of the voice data, so that the user of the first electronic device, that is, the anchor user of the anchor, can select the voice message to select playback and voice message.
- the voice information sent to the second electronic device is a voice message corresponding to the voice data selected to be played by the anchor user.
- the voice data stored on the display of other electronic devices connected to the server can be made to enable the user to select playback and push the played voice data to other electronic devices.
- a second deletion module (not shown) is further included.
- the second deletion module is configured to respond to the deletion request sent by the electronic device sending the voice data, so as to selectively delete the voice data sent by the electronic device, so as to prevent the voice data dissatisfied by the user from being widely transmitted.
- Fig. 7 is a flow chart showing still another method for real-time voice interaction according to an exemplary embodiment.
- the interaction method provided in this embodiment is applied to an electronic device.
- the interaction method is applied to a host that has a long connection between the network live broadcast system and the server.
- the interaction method specifically includes the following steps. :
- step S31 a voice message sent by the server is received.
- the voice message pushed by the server is received through a long connection with the server.
- step S32 at least one voice message is displayed in a list form.
- At least one voice message is displayed in a list form on the display interface for the user to choose to play.
- multiple voice messages are displayed in a list for the anchor user to choose to play voice data corresponding to the corresponding voice message.
- step S33 the voice data corresponding to the voice message is downloaded and played according to the user's selection.
- the user When the user needs to play the corresponding voice message, he can download the voice data corresponding to the voice message by clicking the corresponding voice message, and play the voice data at the same time as the download is completed or downloaded. The corresponding selection playback is completed.
- the anchor user can select and play the uploaded voice data, which increases the anchor's control over the broadcast content and improves the flexibility of the live content.
- this embodiment further includes the following steps:
- step S34 an audio signal for playing voice data is added to the audio stream.
- the audio stream here refers to audio data generated by any audio data played by the local electronic device.
- the audio stream refers to the broadcaster playing locally recorded audio data and the voice data selected for playback, and the voice data comes from the corresponding audience.
- step S35 the audio stream, the id of the voice data, and the video stream are pushed to the server.
- the pushed content also includes the locally recorded video stream and the id of the voice data that is selected to be played.
- step S36 the playback status of the voice data is displayed in the local list.
- At least one voice message is displayed in the local list. While the corresponding voice data is being played, the playback state of the voice message corresponding to the voice data is displayed. For example, a prompt indicating that a voice message is being played is displayed, so that the user can know which voice message corresponds to the voice data being played.
- step S37 the corresponding voice data is played according to the user's selected playback request.
- Fig. 9 is a structural block diagram of another real-time voice interaction device according to an exemplary embodiment.
- the interactive device provided in this embodiment is applied to an electronic device.
- the interactive device is applied to a host that has a long connection between the network live system and the server.
- the interactive device specifically includes a message receiving module 210.
- the message receiving module 210 is configured to receive a voice message sent by a server.
- the voice message pushed by the server is received through a long connection with the server.
- the message display module 220 is configured to display at least one voice message in a list form.
- At least one voice message is displayed in a list form on the display interface for the user to choose to play.
- multiple voice messages are displayed in a list for the anchor user to choose to play voice data corresponding to the corresponding voice message.
- the data download module 230 is configured to download and play the voice data corresponding to the voice message according to the user's selection.
- the user When the user needs to play the corresponding voice message, he can download the voice data corresponding to the voice message by clicking the corresponding voice message, and play the voice data at the same time as the download is completed or downloaded. The corresponding selection playback is completed.
- the anchor user can select and play the uploaded voice data, which increases the anchor's control over the broadcast content and improves the flexibility of the live content.
- the specific embodiment further includes an audio stream processing module 240, an audio stream sending module 250, a second display module 260, and a selected playback module 270.
- the audio stream processing module 240 is configured to add an audio signal for playing voice data to the audio stream.
- the audio stream here refers to audio data generated by any audio data played by the local electronic device.
- the audio stream refers to the broadcaster playing locally recorded audio data and the voice data selected for playback, and the voice data comes from the corresponding audience.
- the audio stream sending module is configured to push the audio stream, the id of the voice data, and the video stream to the server.
- the pushed content also includes the locally recorded video stream and the id of the voice data that is selected to be played.
- the second display module is configured to display the playback status of the voice data in a local list.
- At least one voice message is displayed in the local list. While the corresponding voice data is being played, the playback state of the voice message corresponding to the voice data is displayed. For example, a prompt indicating that a voice message is being played is displayed, so that the user can know which voice message corresponds to the voice data being played.
- the selected playback module is configured to play corresponding voice data according to a user's selected playback request.
- the present application also provides a computer program for performing the operations shown in FIG. 1, FIG. 2, FIG. 5, FIG. 7, or FIG. 8.
- Fig. 11 is a structural block diagram of a server according to an exemplary embodiment.
- the server is provided with at least one processor 1001 and further includes a memory 1002, and the two are connected to each other 1003 through a data bus.
- the memory is configured to store a computer program or instruction
- the processor is configured to acquire and execute the computer program or instruction, so that the electronic device performs an operation as shown in FIG. 5.
- Fig. 12 is a structural block diagram of an electronic device according to an exemplary embodiment.
- the electronic device is provided with at least one processor 1101 and further includes a memory 1102, and the two are connected 1103 through a data bus.
- the memory is configured to store a computer program or instruction
- the processor is configured to acquire and execute the computer program or instruction, so that the electronic device performs the operations shown in FIG. 1, FIG. 2, FIG. 7, or FIG. 8.
- Fig. 13 is a structural block diagram of another electronic device according to an exemplary embodiment.
- the device 1300 may be a mobile phone, a computer, a digital broadcasting terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.
- the device 1300 may include one or more of the following components: a processing component 1302, a memory 1304, a power component 1306, a multimedia component 1308, an audio component 1310, an input / output (I / O) interface 1312, a sensor component 1314, And communication component 1316.
- a processing component 1302 a memory 1304, a power component 1306, a multimedia component 1308, an audio component 1310, an input / output (I / O) interface 1312, a sensor component 1314, And communication component 1316.
- the processing component 1302 generally controls the overall operations of the device 1300, such as operations associated with display, phone calls, data communications, camera operations, and recording operations.
- the processing component 1302 may include one or more processors 1320 to execute instructions to complete all or part of the steps of the method described above.
- the processing component 1302 may include one or more modules to facilitate the interaction between the processing component 1302 and other components.
- the processing component 1302 may include a multimedia module to facilitate the interaction between the multimedia component 1308 and the processing component 1302.
- the memory 1304 is configured to store various types of data to support operation at the device 1300. Examples of such data include instructions for any application or method operating on the device 1300, contact data, phone book data, messages, pictures, videos, and the like.
- the memory 1304 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), Programming read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
- SRAM static random access memory
- EEPROM electrically erasable programmable read-only memory
- EPROM Programming read-only memory
- PROM programmable read-only memory
- ROM read-only memory
- magnetic memory flash memory
- flash memory magnetic disk or optical disk.
- the power supply component 1306 provides power to various components of the device 1300.
- the power component 1306 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 1300.
- the multimedia component 1308 includes a screen that provides an output interface between the device 1300 and a user.
- the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive an input signal from a user.
- the touch panel includes one or more touch sensors to sense touch, swipe, and gestures on the touch panel. The touch sensor can not only sense the boundary of a touch or slide action, but also detect duration and pressure related to the touch or slide operation.
- the multimedia component 1308 includes a front camera and / or a rear camera. When the device 1300 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
- the audio component 1310 is configured to output and / or input audio signals.
- the audio component 1310 includes a microphone (MIC).
- the microphone is configured to receive an external audio signal.
- the received audio signal may be further stored in the memory 1304 or transmitted via the communication component 1316.
- the audio component 1310 further includes a speaker for outputting audio signals.
- the I / O interface 1312 provides an interface between the processing component 1302 and a peripheral interface module.
- the peripheral interface module may be a keyboard, a click wheel, a button, or the like. These buttons may include, but are not limited to: a home button, a volume button, a start button, and a lock button.
- the sensor component 1314 includes one or more sensors for providing status assessment of various aspects of the device 1300.
- the sensor component 1314 can detect the on / off state of the device 1300 and the relative positioning of the components, such as the display and keypad of the device 1300.
- the sensor component 1314 can also detect the change in the position of the device 1300 or a component of the device 1300. The presence or absence of contact with the device 1300, the orientation or acceleration / deceleration of the device 1300, and the temperature change of the device 1300.
- the sensor component 1314 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact.
- the sensor component 1314 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications.
- the sensor component 1314 may further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
- the communication component 1316 is configured to facilitate wired or wireless communication between the device 1300 and other devices.
- the device 1300 may access a wireless network based on a communication standard, such as WiFi, an operator network (such as 2G, 3G, 4G, or 5G), or a combination thereof.
- the communication component 1316 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.
- the communication component 1316 further includes a near field communication (NFC) module to facilitate short-range communication.
- the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
- RFID radio frequency identification
- IrDA infrared data association
- UWB ultra wideband
- Bluetooth Bluetooth
- the device 1300 may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable A gate array (FPGA), controller, microcontroller, microprocessor, or other electronic component is implemented to perform the operations described in FIG. 1, FIG. 2, FIG. 5, FIG. 7, or FIG.
- ASICs application-specific integrated circuits
- DSPs digital signal processors
- DSPDs digital signal processing devices
- PLDs programmable logic devices
- FPGA field programmable A gate array
- controller microcontroller, microprocessor, or other electronic component is implemented to perform the operations described in FIG. 1, FIG. 2, FIG. 5, FIG. 7, or FIG.
- a non-transitory computer-readable storage medium including instructions may be executed by the processor 1320 of the device 1300 to complete the foregoing method.
- the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.
Landscapes
- Engineering & Computer Science (AREA)
- Signal Processing (AREA)
- Multimedia (AREA)
- Databases & Information Systems (AREA)
- Human Computer Interaction (AREA)
- Computer Networks & Wireless Communication (AREA)
- General Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Telephonic Communication Services (AREA)
- Information Transfer Between Computers (AREA)
Abstract
本申请实施例提供了一种实时语音信息的交互方法、装置、电子设备及存储介质,该交互方法和装置具体为响应用户的录音请求,对用户所发声音进行录制并转换,得到至少一份语音数据,至少一份语音数据以队列形式存储于一个发送队列中;将发送队列中的语音数据依次发送到实时信息交互系统的服务器;在本地以列表形式显示发送队列中的语音数据,并显示语音数据的发送状态。通过以上操作,可以使用户以语音方式上传语音数据,还能够使该语音数据在实时信息交互系统中完全起到文字的作用,极大的方便了文字输入慢或者不会文字输入的用户,从而提高了使用体验。
Description
相关申请的交叉引用
本申请要求在2018年09月04日提交中国专利局、申请号为201811027779.0、申请名称为“实时语音信息的交互方法、装置、电子设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及互联网技术领域,尤其涉及一种实时语音信息的交互方法、装置、电子设备及存储介质。
在一些基于互联网的实时信息交互系统中,有一些是一对多的方式进行信息交互的。例如在网络直播系统中,绝大部分情况下一个直播间内只有一个主播,而观众则会有很多,因此,网络直播实现的是一种以主播的影音表达为中心、以一对多进行交流为主要模式的互动交流场景,并需要保证观众之间的平等关系。在这种模式下,观众只能通过文字方式进行表达。
然而,发明人意识到观众的水平良莠不齐,有些人的文字输入速度较慢、甚至不会文字输入,这样一来就使很多人无法有效表达自己的观点,从而使观众的使用体验较差,不利于扩大网络直播的覆盖受众。
发明内容
为克服相关技术中存在的问题,本申请提供一种实时语音信息的交互方法、装置、电子设备及存储介质。
本申请实施例提供一种实时语音信息的交互方法,应用于电子设备,所述交互方法包括:
响应录音请求,对输入的声音进行录制并转换,得到至少一份语音数据,所述至少一份语音数据以队列形式存储于一个发送队列中;
将所述发送队列中的所述语音数据依次发送到与所述电子设备长连接的服务器,以使所述服务器将收到的所述语音数据推送到与录制所述语音数据的电子设备相对应的第一电子设备,以使所述第一电子设备以列表形式显示所述语音数据,使所述第一电子设备的用户对列表中的语音数据选择播放;
在本地以列表形式显示所述发送队列中的所述语音数据,并显示所述语音数据的发送状态。
本申请实施例提供的技术方案可以包括以下有益效果:可以使用户以语音方式上传语音数据,还能够使该语音数据在实时信息交互系统中完全起到文字的作用,极大的方便了文字输入慢或者不会文字输入的用户,从而提高了使用体验。
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本申请的实施例,并与说明书一起用于解释本申请的原理。
图1是根据一示例性实施例示出的一种实时语音信息的交互方法的流程图;
图2是根据一示例性实施例示出的另一种实时语音信息的交互方法的流程图;
图3是根据一示例性实施例示出的一种实时语音信息的交互装置的结构框图;
图4是根据一示例性实施例示出的另一种实时语音信息的交互装置的结构框图;
图5是根据一示例性实施例示出的又一种实时语音信息的交互方法的流程图;
图6是根据一示例性实施例示出的又一种实时语音信息的交互装置的结构框图;
图7是根据一示例性实施例示出的又一种实时语音信息的交互方法的流程图;
图8是根据一示例性实施例示出的又一种实时语音信息的交互方法的流程图;
图9是根据一示例性实施例示出的又一种实时语音信息的交互装置的结构框图;
图10是根据一示例性实施例示出的又一种实时语音信息的交互装置的结构框图;
图11是根据一示例性实施例示出的一种服务器的结构框图;
图12是根据一示例性实施例示出的一种电子设备的结构框图;
图13是根据一示例性实施例示出的另一种电子设备的结构框图。
这里将详细地对示例性实施例进行说明,其示例表示在附图中。以下示例性实施例中所描述的实施方式仅是与如所附权利要求书中所详述的、本申请的一些方面相一致的装置和方法的例子。
图1是根据一示例性实施例示出的一种实时语音信息的交互方法的流程图。
如图1所示,本具体交互方法应用于电子设备中,具体来说本具体实施方式中的交互方法应用于该网络直播系统的观众端,该交互方法包括以下步骤。
在步骤S11中,对输入的声音进行录制,得到至少一份语音数据。
当观众用户通过该观众端发出录音请求时,对输入的声音进行录音并转化,得到至少一份语音数据,即数字化的语音信号。具体来说,是从与该观众端相连接的录制设备中获取相应语音信号并进行转换,从而得到语音数据。具体的实施步骤如下所述:
首先,当用户发出录音请求时,获取用户发出的语音信号,并将语音信号每隔预设时长转换为一份音频数据,即该音频数据由该预设时长的语音信号转换而来,在具体实践时该预设时长可以选择20毫秒。
然后,对于每次的录制所产生的多份音频数据进行汇集,合成为一个独立的语音数据文件,一般说来,该语音数据覆盖的时长为当次录音请求所持续的时长。对于具体的观众端,可以是观众用户对于录音按钮所按压的时长。
另外,在用户发出录音请求时,将该观众端所播放的音视频的播放音量降低,直至降低到0,即控制音视频保持静音,这样可以避免噪音对录音产生干扰,从而得到较为纯净的语音数据。
在步骤S12中,将语音数据依次发送到与电子设备长连接的服务器。
对于由于本具体实施方式可以应用于网络直播系统,因此,在观众端获得语音数据后通过与服务器的长连接依次发送到该服务器,以使服务器存储这些语音数据。服务器在收 到这些语音数据后推送到与录制该语音数据的电子设备相对应的第一电子设备,由于这里录制该语音数据的为网络直播系统的观众端,因此与该观众端相对应的第一电子设备为网络直播系统的主播端。
服务器在接收到语音数据后,将语音数据发送到第一电子设备,即发送到主播端,并使主播端以列表形式显示上述多个语音数据,主播端的主播用户可以对语音数据进行选择播放。所谓选择播放是指从列表中选择相应的语音数据进行播放。
在步骤S13中,在本地以列表形式显示至少一个语音数据。
在实际应用中,所录制的语音数据往往不限于一个,因此,为了方便用户查看,这里以列表形式显示多个语音数据。具体显示方式可以是仅以列表形式显示多个图标,每个图标对应于一个语音数据。除去以列表显示多个语音数据外,还在每个语音数据的预设位置显示相应语音数据的状态,例如,当一个语音数据正在上传服务器时,显示该语音数据的状态为正在上传;如果已经上传完毕,则显示该语音数据的状态为发送完毕。
从上述技术方案可以看出,本申请实施例提供了一种实时语音信息的交互方法,该交互方法应用于实时信息交互系统,该交互方法具体为响应用户的录音请求,对输入的声音进行录制并转换,得到至少一份语音数据,至少一份语音数据以队列形式存储于一个发送队列中;将发送队列中的语音数据依次发送到实时信息交互系统的服务器;在本地以列表形式显示发送队列中的语音数据,并显示语音数据的发送状态。通过以上操作,可以使用户以语音方式上传语音数据,还能够使该语音数据在实时信息交互系统中完全起到文字的作用,极大的方便了文字输入慢或者不会文字输入的用户,从而提高了使用体验。
另外,在本具体实施方式中,还可以包括如下步骤:
根据用户的删除请求删除相应语音数据。
在实际应用中,用户可能在有时候发现所发的语音数据不满意,需要将其删除,因此本操作的目的在于当用户需要删除并发出删除请求时,响应该删除请求,并将该删除请求对应的语音数据予以删除。从而避免不满意的语音数据推送给其他用户。
如果相应语音数据已经上传到服务器,此时根据该删除请求向服务器发送删除控制指令,以控制服务器删除相应的语音数据。
图2是根据一示例性实施例示出的另一种实时语音信息的交互方法的流程图。
如图2所示,本具体交互方法应用于实时信息交互系统中,该实时信息交互系统在实际应用中可以为网络直播系统,因此,具体来说本具体实施方式中的交互方法应用于该网络直播系统的观众端,该交互方法包括以下步骤。
在步骤S11中,对输入的声音进行录制,得到至少一份语音数据。
本步骤与上一具体实施方式中的作用基本相同,这里不再赘述。
在步骤S12中,将语音数据依次发送到与电子设备长连接的服务器。
本步骤与上一具体实施方式中的作用基本相同,这里不再赘述。
在步骤S13中,在本地以列表形式显示至少一个语音数据。
本步骤与上一具体实施方式中的作用基本相同,区别在于,以列表形式显示的语音数据中不仅包括本地录制的语音数据,还包括服务器推送的来自于与用于在本地录制语音数据处于平等位置的第二电子设备所录制的语音数据。对于网络直播系统来说,该列表中不仅包括本地观众端所录制的语音数据,还包括其他观众端所录制的语音数据。
这里的平等位置不是完全的平等,其实是指基本操作方法的地位对等,其具有优先权 方法的不平等内容,比如对于活跃度较高的用户而言,其具有较高的优先权。
在步骤S14中,接收服务器推送的音视频数据。
在本地向服务器上传本地录制的语音数据外,还用于接收服务器下发的音视频数据,还接收服务器推送的音视频数据,音视频数据包括来自于与录制本地的语音数据的电子设备相对应的第一电子设备所录制的音频数据和视频数据,还来自于与本地电子设备具有平等位置的第二电子设备所录制的语音数据。
对于网络直播系统来说,该音视频数据是来自于该系统的主播端和其他观众端,来自主播端的是主播用户所录制的音频数据和视频数据,来自于其他观众端的语音数据是上传到服务器后、由主播用户选择播放的部分或全部语音数据。
在步骤S15中,在本地播放收到的音视频数据。
对于实际应用的网络直播系统来说,即在观众端播放服务器推送的音视频数据,音视频数据包括主播端录制的音频数据和视频数据,还包括其他观众端所发送的语音数据。
在步骤S16中,对正在播放的音视频数据的id进行检测。
具体来说是检测音视频数据中被同时播放的语音数据的id,该id有可能与列表中显示的多个语音数据相匹配,即来自于同一个电子设备。
在步骤S17中,显示与上述id对应的语音数据的播放状态。
即如果列表中显示的语音数据与播放的音视频数据的id相匹配,则在该列表中将该语音数据显示为正在播放状态,从而能够使用户确定那个语音数据正在播放的音视频数据中被同时播放,以便对其进行相应的操作,如再次播放或循环播放。
在步骤S18中,控制语音数据再次播放或循环播放。
在用户需要对于上述id相对应的语音数据再次播放时,可以输入相应的循环播放指令,循环播放指令用于控制该语音数据再次播放,还可以不限次数或者限定次数的循环播放,以便用户能够准确地了解相应语音数据所承载的内容。
通过以上操作,可以使用户以语音方式上传语音数据,还能够使该语音数据在实时信息交互系统中完全起到文字的作用,极大的方便了文字输入慢或者不会文字输入的用户,从而提高了使用体验。还能够使用户得到较为高级的使用体验。
图3是根据一示例性实施例示出的一种实时语音信息的交互装置的结构框图。
如图3所示,本具体交互装置应用于电子设备中,具体来说本具体实施方式中的交互装置应用于该网络直播系统的观众端,该交互装置包括语音录制模块10、语音发送模块20和第一显示模块30。
语音录制模块10被配置为对输入的声音进行录制,得到至少一份语音数据。
当观众用户通过该观众端发出录音请求时,对输入的声音进行录音并转化,得到至少一份语音数据,即数字化的语音信号。具体来说,是从与该观众端相连接的录制设备中获取相应语音信号并进行转换,从而得到语音数据。该模块具体包括录制控制单元和数据汇集单元。
录制控制单元被配置为当用户发出录音请求时,获取用户发出的语音信号,并将语音信号每隔预设时长转换为一份音频数据,即该音频数据由该预设时长的语音信号转换而来,在具体实践时该预设时长可以选择20毫秒。
数据汇集单元则被配置为对于每次的录制所产生的多份音频数据进行汇集,合成为一个独立的语音数据文件,一般说来,该语音数据覆盖的时长为当次录音请求所持续的时长。 对于具体的观众端,可以是观众用户对于录音按钮所按压的时长。
另外,该模块还包括静音控制单元,该静音控制单元被配置为在用户发出录音请求时,将该观众端所播放的音视频的播放音量降低,直至降低到0,即控制音视频保持静音,这样可以避免噪音对录音产生干扰,从而得到较为纯净的语音数据。
语音发送模块20被配置为将语音数据依次发送到与电子设备相连接的服务器。
对于由于本具体实施方式可以应用于网络直播系统,因此,在观众端获得语音数据后通过与服务器的长连接依次发送到该服务器,以使服务器存储这些语音数据。服务器在收到这些语音数据后推送到与录制该语音数据的电子设备相对应的第一电子设备,由于这里录制该语音数据的为网络直播系统的观众端,因此与该观众端相对应的第一电子设备为网络直播系统的主播端。
服务器在接收到语音数据后,将语音数据发送到第一电子设备,即发送到主播端,并使主播端以列表形式显示上述多个语音数据,主播端的主播用户可以对语音数据进行选择播放。所谓选择播放是指从列表中选择相应的语音数据进行播放。
第一显示模块30被配置为在本地以列表形式显示至少一个语音数据。
在实际应用中,所录制的语音数据往往不限于一个,因此,为了方便用户查看,这里以列表形式显示多个语音数据。具体显示方式可以是仅以列表形式显示多个图标,每个图标对应于一个语音数据。除去以列表显示多个语音数据外,还在每个语音数据的预设位置显示相应语音数据的状态,例如,当一个语音数据正在上传服务器时,显示该语音数据的状态为正在上传;如果已经上传完毕,则显示该语音数据的状态为发送完毕。
从上述技术方案可以看出,本申请实施例提供了一种实时语音信息的交互装置,该交互装置应用于实时信息交互系统,该交互装置具体为响应用户的录音请求,对输入的声音进行录制并转换,得到至少一份语音数据,至少一份语音数据以队列形式存储于一个发送队列中;将发送队列中的语音数据依次发送到实时信息交互系统的服务器;在本地以列表形式显示发送队列中的语音数据,并显示语音数据的发送状态。通过以上操作,可以使用户以语音方式上传语音数据,还能够使该语音数据在实时信息交互系统中完全起到文字的作用,极大的方便了文字输入慢或者不会文字输入的用户,从而提高了使用体验。
另外,在本具体实施方式中,还可以包括第一删除模块(未示出)。
第一删除模块被配置为根据用户的删除请求删除相应语音数据。
在实际应用中,用户可能在有时候发现所发的语音数据不满意,需要将其删除,因此本操作的目的在于当用户需要删除并发出删除请求时,响应该删除请求,并将该请求对应的语音数据予以删除。从而避免不满意的语音数据推送给其他用户。
如果相应语音数据已经上传到服务器,此时根据该删除请求向服务器发送删除控制指令,以控制服务器删除相应的语音数据。
图4是根据一示例性实施例示出的另一种实时语音信息的交互装置的结构框图。
如图4所示,本具体交互方法应用于电子设备中,具体来说本具体实施方式中的交互装置应用于该网络直播系统的观众端,该交互装置相比于上一具体实施方式增设了音视频接收模块40、音视频播放模块50、id检测模块60、状态显示模块70和循环播放模块80。
第一显示模块还被配置为在本地以列表形式显示至少一个语音数据。但是具有一定的区别,该区别在于以列表形式显示的语音数据中不仅包括本地录制的语音数据,还包括服务器推送的来自于与用于在本地录制语音数据的电子设备处于平等地位置的第二电子设 备所录制的语音数据。对于网络直播系统来说,该列表中不仅包括本地观众端所录制的语音数据,还包括其他观众端所录制的语音数据。
音视频接收模块40被配置为接收服务器推送的音视频数据。
在本地向服务器上传本地录制的语音数据外,还用于接收服务器下发的音视频数据,还接收服务器推送的音视频数据,音视频数据包括来自于与录制本地的语音数据的电子设备相对应的第一电子设备所录制的音频数据和视频数据,还来自于与本地电子设备具有平等位置的第二电子设备所录制的语音数据。
对于网络直播系统来说,该音视频数据是来自于该系统的主播端和其他观众端,来自主播端的是主播用户所录制的音频数据和视频数据,来自于其他观众端的语音数据是上传到服务器后、由主播用户选择播放的部分或全部语音数据。
音视频播放模块50被配置为在本地播放收到的音视频数据。
对于实际应用的网络直播系统来说,即在观众端播放服务器推送的音视频数据,音视频数据包括主播端录制的音频数据和视频数据,还包括其他观众端所发送的语音数据。
id检测模块60被配置为对正在播放的音视频数据的id进行检测。
具体来说是检测音视频数据中被同时播放的语音数据的id,该id有可能与列表中显示的多个语音数据相匹配,即来自于同一个电子设备。
状态显示模块70被配置为显示与上述id对应的语音数据的播放状态。
即如果列表中显示的语音数据与播放的音视频数据的id相匹配,则在该列表中将该语音数据显示为正在播放状态,从而能够使用户确定那个语音数据正在播放的音视频数据中被同时播放,以便对其进行相应的操作,如再次播放或循环播放。
循环播放模块80被配置为控制语音数据再次播放或循环播放。
在用户需要对于上述id相对应的语音数据再次播放时,可以输入相应的循环播放指令,循环播放指令用于控制该语音数据再次播放,还可以不限次数或者限定次数的循环播放,以便用户能够准确地了解相应语音数据所承载的内容。
通过以上操作,可以使用户以语音方式上传语音数据,还能够使该语音数据在实时信息交互系统中完全起到文字的作用,极大的方便了文字输入慢或者不会文字输入的用户,从而提高了使用体验。还能够使用户得到较为高级的使用体验。
图5是根据一示例性实施例示出的又一种实时语音信息的交互方法的流程图。
如图5所示,本具体实施方式中提供的交互方法应用于实时信息交互系统的服务器,以网络直播系统为例,该服务器分别与网络直播系统的主播端、多个观众端长连接。该交互方法具体包括步骤:
在步骤S21中,接收与服务器长连接的电子设备发送的语音数据。
对于网络直播系统而言,与服务器长连接的电子设备为观众端,在观众端录制语音数据并上传该语音数据后,以队列形式接收该语音数据。
在步骤S22中,根据发送语音数据的设备编号为语音数据附加一个id。
具体而言,在接收到每个语音数据后,对发送语音数据的硬件设备的设备编号进行检测,并根据检测到的设备编号编辑一个id,并将该id附加到相应的语音数据上。
在步骤S23中,向第一电子设备和第二电子设备分别发送语音消息。
这里的第一电子设备与发送相应语音数据的电子设备相对应,第二电子设备则与发送相应语音数据的电子设备具备平等位置。对于网络直播系统而言,发送该语音数据的为观 众端,第一电子设备则为主播端,第二电子设备则为其他观众端。
向第一电子设备发送的语音消息还包括语音数据的发送者信息、时长和id,这样可以使第一电子设备的用户、即主播端的主播用户对语音消息进行选定,以选择播放与语音消息相对应的语音数据。向第二电子设备发送的语音信息是被主播用户选择播放的语音数据所对应的语音消息。
通过上述操作,可以使与服务器连接的其他电子设备显示器所存储的语音数据,以使用户能够选择播放,并将播放的语音数据推送到其他电子设备。
另外,在本具体实施方式中,还包括如下步骤:
响应发送语音数据的电子设备发送的删除请求,以便将该电子设备发送的语音数据进行选择性删除,以避免用户不满意的语音数据被广泛传播。
图6是根据一示例性实施例示出的又一种实时语音信息的交互装置的结构框图。
如图6所示,本具体实施方式中提供的交互装置应用于实时信息交互系统的服务器,以网络直播系统为例,该服务器分别与网络直播系统的主播端、多个观众端长连接。该交互装置具体包括数据接收模块110、id附加模块120和消息推送模块130。
数据接收模块110被配置为接收与服务器长连接的电子设备发送的语音数据。
对于网络直播系统而言,与服务器长连接的电子设备为观众端,在观众端录制语音数据并上传该语音数据后,以队列形式接收该语音数据。
id附加模块120被配置为根据发送语音数据的设备编号为语音数据附加一个id。
具体而言,在接收到每个语音数据后,对发送语音数据的硬件设备的设备编号进行检测,并根据检测到的设备编号编辑一个id,并将该id附加到相应的语音数据上。
消息推送模块130被配置为向第一电子设备和第二电子设备分别发送语音消息。
这里的第一电子设备与发送相应语音数据的电子设备相对应,第二电子设备则与发送相应语音数据的电子设备具备平等位置。对于网络直播系统而言,发送该语音数据的为观众端,第一电子设备则为主播端,第二电子设备则为其他观众端。
向第一电子设备发送的语音消息还包括语音数据的发送者信息、时长和id,这样可以使第一电子设备的用户、即主播端的主播用户对语音消息进行选定,以选择播放与语音消息相对应的语音数据。向第二电子设备发送的语音信息是被主播用户选择播放的语音数据所对应的语音消息。
通过上述操作,可以使与服务器连接的其他电子设备显示器所存储的语音数据,以使用户能够选择播放,并将播放的语音数据推送到其他电子设备。
另外,在本具体实施方式中,还包括第二删除模块(未示出)。
第二删除模块被配置为响应发送语音数据的电子设备发送的删除请求,以便将该电子设备发送的语音数据进行选择性删除,以避免用户不满意的语音数据被广泛传播。
图7是根据一示例性实施例示出的又一种实时语音的交互方法的流程图。
如图7所示,本具体实施方式中提供的交互方法应用于电子设备,对于网络直播系统而言,该交互方法应用于网络直播系统与服务器长连接的主播端,该交互方法具体包括如下步骤:
在步骤S31中,接收服务器发送的语音消息。
具体是通过与服务器的长连接接收服务器推送的语音消息。
在步骤S32中,以列表形式显示至少一条语音消息。
在接收到语音消息后,在显示界面上以列表形式显示至少一条语音消息,以便供用户选择播放。对于实际应用中的网络直播系统来说,通过列表显示多个语音消息以供主播用户选择播放与相应语音消息对应的语音数据。
在步骤S33中,根据用户的选择下载并播放与语音消息对应的语音数据。
在用户需要播放相应语音消息时,既可通过点击相应语音消息的方式下载与该语音消息对应的语音数据,并在下载完成或下载的同时对该语音数据进行播放。即完成相应的选择播放。
通过上述操作,对于网络直播系统而言,可以使主播用户对上传的语音数据进行选择播放,增加了主播对播放内容的控制权,提高了直播内容的灵活性。
另外,如图8所示,本具体实施方式中还包括如下步骤:
在步骤S34中,将播放语音数据的音频信号加入到音频流中。
这里的音频流是指本地电子设备所播放任何音频数据所产生的音频数据。对于网络直播系统而言,该音频流是指主播端播放本地录制的音频数据和选择播放的语音数据,该语音数据来自于相应的观众端。
在步骤S35中,将音频流、语音数据的id和视频流推送到服务器。
在得到上述音频流后将其推送到服务器,推送的内容还包括本地录制的视频流,还包括被选择播放的语音数据的id。
在步骤S36中,在本地列表显示语音数据的播放状态。
本地列表中显示有至少一个语音消息,在播放相应的语音数据的同时,显示与该语音数据相对应的语音消息的播放状态。例如针对某条语音消息显示正在播放的提示,这样能够使用户明确哪条语音消息所对应的语音数据正在播放。
在步骤S37中,根据用户的选定播放请求播放对应的语音数据。
在用户想要重听或仔细收听所播放的语音数据后,可以通过对所提示的语音消息的操作输入选定播放请求,以选定相应语音消息,从而能够对选定的语音消息所对应的语音数据进行反复播放。
图9是根据一示例性实施例示出的又一种实时语音的交互装置的结构框图。
如图9示,本具体实施方式中提供的交互装置应用于电子设备,对于网络直播系统而言,该交互装置应用于网络直播系统与服务器长连接的主播端,该交互装置具体包括消息接收模块210、消息显示模块220和数据下载模块230。
消息接收模块210被配置为接收服务器发送的语音消息。
具体是通过与服务器的长连接接收服务器推送的语音消息。
消息显示模块220被配置为以列表形式显示至少一条语音消息。
在接收到语音消息后,在显示界面上以列表形式显示至少一条语音消息,以便供用户选择播放。对于实际应用中的网络直播系统来说,通过列表显示多个语音消息以供主播用户选择播放与相应语音消息对应的语音数据。
数据下载模块230被配置为根据用户的选择下载并播放与语音消息对应的语音数据。
在用户需要播放相应语音消息时,既可通过点击相应语音消息的方式下载与该语音消息对应的语音数据,并在下载完成或下载的同时对该语音数据进行播放。即完成相应的选择播放。
通过上述操作,对于网络直播系统而言,可以使主播用户对上传的语音数据进行选择 播放,增加了主播对播放内容的控制权,提高了直播内容的灵活性。
另外,如图10所示,本具体实施方式中还包括音频流处理模块240、音频流发送模块250、第二显示模块260和选定播放模块270。
音频流处理模块240被配置为将播放语音数据的音频信号加入到音频流中。
这里的音频流是指本地电子设备所播放任何音频数据所产生的音频数据。对于网络直播系统而言,该音频流是指主播端播放本地录制的音频数据和选择播放的语音数据,该语音数据来自于相应的观众端。
音频流发送模块被配置为将音频流、语音数据的id和视频流推送到服务器。
在得到上述音频流后将其推送到服务器,推送的内容还包括本地录制的视频流,还包括被选择播放的语音数据的id。
第二显示模块被配置为在本地列表显示语音数据的播放状态。
本地列表中显示有至少一个语音消息,在播放相应的语音数据的同时,显示与该语音数据相对应的语音消息的播放状态。例如针对某条语音消息显示正在播放的提示,这样能够使用户明确哪条语音消息所对应的语音数据正在播放。
选定播放模块被配置为根据用户的选定播放请求播放对应的语音数据。
在用户想要重听或仔细收听所播放的语音数据后,可以通过对所提示的语音消息的操作输入选定播放请求,以选定相应语音消息,从而能够对选定的语音消息所对应的语音数据进行反复播放。
本申请还提供一种计算机程序,该计算机程序用于执行如图1、图2、图5、图7或图8所示的操作。
图11是根据一示例性实施例示出的一种服务器的结构框图。
如图11所示,该服务器设置有至少一个处理器1001,还包括存储器1002,两者通过数据总线连接1003。
存储器用于存储计算机程序或指令,处理器用于获取并执行该计算机程序或指令,以使电子设备执行如图5所示的操作。
图12是根据一示例性实施例示出的一种电子设备的结构框图。
如图11所示,该电子设备设置有至少一个处理器1101,还包括存储器1102,两者通过数据总线连接1103。
存储器用于存储计算机程序或指令,处理器用于获取并执行该计算机程序或指令,以使电子设备执行如下图1、图2、图7或图8的操作。
图13是根据一示例性实施例示出的另一种电子设备的结构框图。例如,设备1300可以是移动电话,计算机,数字广播终端,消息收发设备,游戏控制台,平板设备,医疗设备,健身设备,个人数字助理等。
参照图12,设备1300可以包括以下一个或多个组件:处理组件1302,存储器1304,电源组件1306,多媒体组件1308,音频组件1310,输入/输出(I/O)的接口1312,传感器组件1314,以及通信组件1316。
处理组件1302通常控制设备1300的整体操作,诸如与显示,电话呼叫,数据通信,相机操作和记录操作相关联的操作。处理组件1302可以包括一个或多个处理器1320来执行指令,以完成上述的方法的全部或部分步骤。此外,处理组件1302可以包括一个或多个模块,便于处理组件1302和其他组件之间的交互。例如,处理组件1302可以包括多媒 体模块,以方便多媒体组件1308和处理组件1302之间的交互。
存储器1304被配置为存储各种类型的数据以支持在设备1300的操作。这些数据的示例包括用于在设备1300上操作的任何应用程序或方法的指令,联系人数据,电话簿数据,消息,图片,视频等。存储器1304可以由任何类型的易失性或非易失性存储设备或者它们的组合实现,如静态随机存取存储器(SRAM),电可擦除可编程只读存储器(EEPROM),可擦除可编程只读存储器(EPROM),可编程只读存储器(PROM),只读存储器(ROM),磁存储器,快闪存储器,磁盘或光盘。
电源组件1306为设备1300的各种组件提供电力。电源组件1306可以包括电源管理系统,一个或多个电源,及其他与为设备1300生成、管理和分配电力相关联的组件。
多媒体组件1308包括在设备1300和用户之间的提供一个输出接口的屏幕。在一些实施例中,屏幕可以包括液晶显示器(LCD)和触摸面板(TP)。如果屏幕包括触摸面板,屏幕可以被实现为触摸屏,以接收来自用户的输入信号。触摸面板包括一个或多个触摸传感器以感测触摸、滑动和触摸面板上的手势。触摸传感器可以不仅感测触摸或滑动动作的边界,而且还检测与触摸或滑动操作相关的持续时间和压力。在一些实施例中,多媒体组件1308包括一个前置摄像头和/或后置摄像头。当设备1300处于操作模式,如拍摄模式或视频模式时,前置摄像头和/或后置摄像头可以接收外部的多媒体数据。每个前置摄像头和后置摄像头可以是一个固定的光学透镜系统或具有焦距和光学变焦能力。
音频组件1310被配置为输出和/或输入音频信号。例如,音频组件1310包括一个麦克风(MIC),当设备1300处于操作模式,如呼叫模式、记录模式和语音识别模式时,麦克风被配置为接收外部音频信号。所接收的音频信号可以被进一步存储在存储器1304或经由通信组件1316发送。在一些实施例中,音频组件1310还包括一个扬声器,用于输出音频信号。
I/O接口1312为处理组件1302和外围接口模块之间提供接口,上述外围接口模块可以是键盘,点击轮,按钮等。这些按钮可包括但不限于:主页按钮、音量按钮、启动按钮和锁定按钮。
传感器组件1314包括一个或多个传感器,用于为设备1300提供各个方面的状态评估。例如,传感器组件1314可以检测到设备1300的打开/关闭状态,组件的相对定位,例如组件为设备1300的显示器和小键盘,传感器组件1314还可以检测设备1300或设备1300一个组件的位置改变,用户与设备1300接触的存在或不存在,设备1300方位或加速/减速和设备1300的温度变化。传感器组件1314可以包括接近传感器,被配置用来在没有任何的物理接触时检测附近物体的存在。传感器组件1314还可以包括光传感器,如CMOS或CCD图像传感器,用于在成像应用中使用。在一些实施例中,该传感器组件1314还可以包括加速度传感器,陀螺仪传感器,磁传感器,压力传感器或温度传感器。
通信组件1316被配置为便于设备1300和其他设备之间有线或无线方式的通信。设备1300可以接入基于通信标准的无线网络,如WiFi,运营商网络(如2G、3G、4G或5G),或它们的组合。在一个示例性实施例中,通信组件1316经由广播信道接收来自外部广播管理系统的广播信号或广播相关信息。在一个示例性实施例中,通信组件1316还包括近场通信(NFC)模块,以促进短程通信。例如,在NFC模块可基于射频识别(RFID)技术,红外数据协会(IrDA)技术,超宽带(UWB)技术,蓝牙(BT)技术和其他技术来实现。
在示例性实施例中,设备1300可以被一个或多个应用专用集成电路(ASIC)、数字信号处理器(DSP)、数字信号处理设备(DSPD)、可编程逻辑器件(PLD)、现场可编程门阵列(FPGA)、控制器、微控制器、微处理器或其他电子元件实现,用于执行如图1、图2、图5、图7或图8所述的操作。
在示例性实施例中,还提供了一种包括指令的非临时性计算机可读存储介质,例如包括指令的存储器1304,上述指令可由设备1300的处理器1320执行以完成上述方法。例如,非临时性计算机可读存储介质可以是ROM、随机存取存储器(RAM)、CD-ROM、磁带、软盘和光数据存储设备等。
Claims (45)
- 一种实时语音信息的交互方法,应用于电子设备,所述交互方法包括:响应录音请求,对输入的声音进行录制并转换,得到至少一份语音数据,所述至少一份语音数据以队列形式存储于一个发送队列中;将所述发送队列中的所述语音数据依次发送到与所述电子设备长连接的服务器;在本地以列表形式显示所述发送队列中的所述语音数据,并显示所述语音数据的发送状态。
- 如权利要求1所述的交互方法,所述对输入的声音进行录制并转换,得到至少一份语音数据,包括:在对所述声音进行录制时,每隔预设时长产生一份音频数据;将所述录音请求的存续期间内所产生的多份音频数据汇集为所述语音数据。
- 如权利要求2所述的交互方法,所述对输入的声音进行录制并转换,得到至少一份语音数据,还包括:在对所述声音进行录制的同时,控制所述实时信息交互系统在本地播放的视频静音。
- 如权利要求1所述的交互方法,所述语音数据的发送状态包括正在发送状态或发送完毕状态。
- 如权利要求1所述的交互方法,还包括:响应删除请求,将所述删除请求指向的所述语音数据予以删除。
- 如权利要求1~5任一项所述的交互方法,还包括:接收所述服务器推送的音视频数据,所述音视频数据包括与所述服务器长连接的第一电子设备录制的音频数据和视频数据,还包括与所述电子设备处于平等位置的第二电子设备所录制的语音数据;在本地播放所述音频视频数据;在本地以列表形式显示的内容中包括来自于所述第二电子设备所录制的语音数据。
- 如权利要求6所述的交互方法,还包括:对正在播放的所述音视频数据的id进行检测;如果正在播放的所述音视频数据的id与以列表形式显示的所述语音数据相对应,则将与所述id相对应的语音数据的状态显示为正在播放状态。
- 如权利要求7所述的交互方法,还包括:响应循环播放指令,控制与所述id相对应的语音数据再次播放或者循环播放。
- 一种实时语音信息的交互装置,应用于电子设备,所述交互装置包括:语音录制模块,被配置为响应录音请求,对输入的声音进行录制并转换,得到至少一份语音数据,所述至少一份语音数据以队列形式存储于一个发送队列中;语音发送模块,被配置为将所述发送队列中的所述语音数据依次发送到与所述电子设备长连接的服务器;第一显示模块,被配置为在本地以列表形式显示所述发送队列中的所述语音数据,并显示所述语音数据的发送状态。
- 如权利要求9所述的交互装置,所述语音录制模块包括:录制控制单元,被配置为在对所述声音进行录制时,每隔预设时长产生一份音频数据;数据汇集单元,被配置为将所述录音请求的存续期间内所产生的多份音频数据汇集为所述语音数据。
- 如权利要求10所述的交互装置,所述语音录制模块还包括:静音控制单元,被配置为在对所述声音进行录制的同时,控制所述实时信息交互系统在本地播放的视频静音。
- 如权利要求9所述的交互装置,所述语音数据的发送状态包括正在发送状态或发送完毕状态。
- 如权利要求9所述的交互装置,还包括:第一删除模块,被配置为响应删除请求,将所述删除请求指向的所述语音数据予以删除。
- 如权利要求9~13任一项所述的交互装置,还包括:音视频接收模块,被配置为接收所述服务器推送的音视频数据,所述音视频数据包括与所述服务器长连接的第一电子设备录制的音频数据和视频数据,还包括与所述电子设备处于平等位置的第二电子设备所录制的语音数据;音视频播放模块,被配置为在本地播放所述音视频数据;在本地以列表形式显示的内容中包括来自于所述第二电子设备所录制的语音数据。
- 如权利要求14所述的交互装置,还包括:id检测模块,被配置为对正在播放的所述音视频数据的id进行检测;状态显示模块,被配置为如果正在播放的所述音视频数据的id与以列表形式显示的所述语音数据相对应,则将与所述id相对应的语音数据的状态显示为正在播放状态。
- 如权利要求15所述的交互装置,还包括:循环播放模块,被配置为响应循环播放指令,控制与所述id相对应的语音数据再次播放或者循环播放。
- 一种实时语音信息的交互方法,应用于实时信息交互系统的服务器,所述交互方法包括:接收与所述服务器长连接的电子设备以队列形式发送的语音数据;根据发送所述语音数据的设备编号为所述语音数据编制并附加一个与所述设备编号相匹配的id;向与所述服务器长连接、且与所述电子设备相对应的第一电子设备发送语音消息,所述语音消息包括接收到的所述语音数据的发送者信息、时长和所述id;同时,向与所述服务器长连接、且与所述电子设备处于平等位置的第二电子设备发送所述消息。
- 如权利要求17所述的交互方法,还包括:响应所述电子设备发送的删除请求,将与所述删除请求对应的语音信息予以删除。
- 一种实时语音信息的交互装置,应用于实时信息交互系统的服务器,所述交互装置包括:数据接收模块,被配置为接收与所述服务器长连接的电子设备以队列形式发送的语音数据;id附加模块,被配置为根据发送所述语音数据的设备编号为所述语音数据编制并附加一个与所述设备编号相匹配的id;消息推送模块,被配置为向与所述服务器长连接、且与所述电子设备相对应的第一电子设备发送语音消息,所述语音消息包括接收到的所述语音数据的发送者信息、时长和所述id;还被配置为向与所述服务器长连接、且与所述电子设备处于平等位置的第二电子设备发送所述语音消息。
- 如权利要求19所述的交互装置,还包括:第二删除模块,配置为响应所述电子设备发送的删除请求,将与所述删除请求对应的语音信息予以删除。
- 一种实时语音信息的交互方法,应用于电子设备,所述交互方法包括:接收与所述电子设备长连接的服务器发送的语音消息;以列表形式显示从所述服务器接收的至少一条所述语音消息;响应下载请求,从所述服务器下载并播放与被选择的语音消息所对应的所述语音数据。
- 如权利要求21所述的交互方法,还包括:在播放所述语音数据的同时,将播放所述语音数据的音频信号加入到播放本地所采集的音频流中;将所述音频流、所述语音数据的id和本地采集的视频流推送到所述服务器。
- 如权利要求22所述的交互方法,还包括:在播放所述语音数据的同时,在本地列表中显示所述语音消息的播放状态。
- 如权利要求22所述的交互方法,还包括:响应选定播放请求,播放与所述选定播放请求对应的所述语音数据。
- 一种实时语音信息的交互装置,应用于电子设备,所述交互装置包括:消息接收模块,被配置为接收与所述电子设备长连接的服务器发送的语音消息;消息显示模块,被配置为以列表形式显示从所述服务器接收的至少一条所述语音消息;数据下载模块,被配置为响应下载请求,从所述服务器下载并播放与被选择的语音消息所对应的所述语音数据。
- 如权利要求25所述的交互装置,还包括:音频流处理模块,被配置为在播放所述语音数据的同时,将播放所述语音数据的音频信号加入到播放本地所采集的音频流中;音频流发送模块,被配置为将所述音频流、所述语音数据的id和本地采集的视频流推送到所述服务器。
- 如权利要求26所述的交互装置,还包括:第二显示模块,被配置为在播放所述语音数据的同时,在本地列表中显示所述语音消息的播放状态。
- 如权利要求26所述的交互装置,还包括:选定播放模块,被配置为响应选定播放请求,播放与所述选定播放请求对应的所述语音数据。
- 一种电子设备,包括:存储器,用于存储计算机程序,以及执行所述计算机程序产生的候选中间数据以及结果数据;处理器,用于响应录音请求,对输入的声音进行录制并转换,得到至少一份语音数据,所述至少一份语音数据以队列形式存储于一个发送队列中;将所述发送队列中的所述语音 数据依次发送到与所述电子设备长连接的服务器;在本地以列表形式显示所述发送队列中的所述语音数据,并显示所述语音数据的发送状态。
- 如权利要求29所述的电子设备,所述处理器,具体用于在对所述声音进行录制时,每隔预设时长产生一份音频数据;将所述录音请求的存续期间内所产生的多份音频数据汇集为所述语音数据。
- 如权利要求30所述的电子设备,所述处理器,还用于在对所述声音进行录制的同时,控制所述实时信息交互系统在本地播放的视频静音。
- 如权利要求29所述的电子设备,所述语音数据的发送状态包括正在发送状态或发送完毕状态。
- 如权利要求29所述的电子设备,所述处理器,还用于响应删除请求,将所述删除请求指向的所述语音数据予以删除。
- 如权利要求29~33任一项所述的电子设备,所述处理器,还用于接收所述服务器推送的音视频数据,所述音视频数据包括与所述服务器长连接的第一电子设备录制的音频数据和视频数据,还包括与所述电子设备处于平等位置的第二电子设备所录制的语音数据;在本地播放所述音频视频数据;在本地以列表形式显示的内容中包括来自于所述第二电子设备所录制的语音数据。
- 如权利要求34所述的电子设备,所述处理器,还用于对正在播放的所述音视频数据的id进行检测;如果正在播放的所述音视频数据的id与以列表形式显示的所述语音数据相对应,则将与所述id相对应的语音数据的状态显示为正在播放状态。
- 如权利要求35所述的电子设备,所述处理器,还用于响应循环播放指令,控制与所述id相对应的语音数据再次播放或者循环播放。
- 一种服务器,包括:存储器,用于存储计算机程序,以及执行所述计算机程序产生的候选中间数据以及结果数据;处理器,用于接收与所述服务器长连接的电子设备以队列形式发送的语音数据;根据发送所述语音数据的设备编号为所述语音数据编制并附加一个与所述设备编号相匹配的id;向与所述服务器长连接、且与所述电子设备相对应的第一电子设备发送语音消息,所述语音消息包括接收到的所述语音数据的发送者信息、时长和所述id;同时,向与所述服务器长连接、且与所述电子设备处于平等位置的第二电子设备发送所述消息。
- 如权利要求37所述的服务器,所述处理器,还用于响应所述电子设备发送的删除请求,将与所述删除请求对应的语音信息予以删除。
- 一种电子设备,包括:存储器,用于存储计算机程序,以及执行所述计算机程序产生的候选中间数据以及结果数据;处理器,用于接收与所述电子设备长连接的服务器发送的语音消息;以列表形式显示从所述服务器接收的至少一条所述语音消息;响应下载请求,从所述服务器下载并播放与被选择的语音消息所对应的所述语音数据。
- 如权利要求39所述的电子设备,所述处理器,还用于在播放所述语音数据的同时,将播放所述语音数据的音频信号加入到播放本地所采集的音频流中;将所述音频流、所述语音数据的id和本地采集的视频流推送到所述服务器。
- 如权利要求40所述的电子设备,所述处理器,还用于在播放所述语音数据的同时,在本地列表中显示所述语音消息的播放状态。
- 如权利要求40所述的电子设备,所述处理器,还用于响应选定播放请求,播放与所述选定播放请求对应的所述语音数据。
- 一种计算机可读存储介质,其上承载一个或多个计算机指令程序,所述计算机指令程序被一个或多个处理器执行时,所述一个或多个处理器执行权利要求1~8任一项所述的方法。
- 一种计算机可读存储介质,其上承载一个或多个计算机指令程序,所述计算机指令程序被一个或多个处理器执行时,所述一个或多个处理器执行权利要求17或18任一项所述的方法。
- 一种计算机可读存储介质,其上承载一个或多个计算机指令程序,所述计算机指令程序被一个或多个处理器执行时,所述一个或多个处理器执行权利要求21~24任一项所述的方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/257,563 US20210266633A1 (en) | 2018-09-04 | 2019-09-04 | Real-time voice information interactive method and apparatus, electronic device and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201811027779.0 | 2018-09-04 | ||
| CN201811027779.0A CN109039872B (zh) | 2018-09-04 | 2018-09-04 | 实时语音信息的交互方法、装置、电子设备及存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020048490A1 true WO2020048490A1 (zh) | 2020-03-12 |
Family
ID=64623932
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/104421 Ceased WO2020048490A1 (zh) | 2018-09-04 | 2019-09-04 | 实时语音信息的交互方法、装置、电子设备及存储介质 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20210266633A1 (zh) |
| CN (1) | CN109039872B (zh) |
| WO (1) | WO2020048490A1 (zh) |
Families Citing this family (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109039872B (zh) * | 2018-09-04 | 2020-04-17 | 北京达佳互联信息技术有限公司 | 实时语音信息的交互方法、装置、电子设备及存储介质 |
| CN112398890B (zh) * | 2019-08-16 | 2024-11-29 | 北京搜狗科技发展有限公司 | 一种信息推送方法、装置、和用于推送信息的装置 |
| CN111508471B (zh) * | 2019-09-17 | 2021-04-20 | 马上消费金融股份有限公司 | 语音合成方法及其装置、电子设备和存储装置 |
| US11462218B1 (en) | 2020-04-29 | 2022-10-04 | Amazon Technologies, Inc. | Conserving battery while detecting for human voice |
| CN114245195B (zh) * | 2022-01-13 | 2023-11-07 | 百果园技术(新加坡)有限公司 | 直播互动方法、装置、设备、存储介质及程序产品 |
| CN114760274B (zh) * | 2022-06-14 | 2022-09-02 | 北京新唐思创教育科技有限公司 | 在线课堂的语音交互方法、装置、设备及存储介质 |
| CN115586976A (zh) * | 2022-09-09 | 2023-01-10 | 阿里云计算有限公司 | 交互数据处理方法、装置、电子设备及存储介质 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104333770A (zh) * | 2014-11-20 | 2015-02-04 | 广州华多网络科技有限公司 | 一种视频直播的方法以及装置 |
| CN105657482A (zh) * | 2016-03-28 | 2016-06-08 | 广州华多网络科技有限公司 | 一种语音弹幕的实现方法及装置 |
| CN108055577A (zh) * | 2017-12-18 | 2018-05-18 | 北京奇艺世纪科技有限公司 | 一种直播交互方法、系统、装置及电子设备 |
| CN108259989A (zh) * | 2018-01-19 | 2018-07-06 | 广州华多网络科技有限公司 | 视频直播的方法、计算机可读存储介质和终端设备 |
| CN109039872A (zh) * | 2018-09-04 | 2018-12-18 | 北京达佳互联信息技术有限公司 | 实时语音信息的交互方法、装置、电子设备及存储介质 |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20030005462A1 (en) * | 2001-05-22 | 2003-01-02 | Broadus Charles R. | Noise reduction for teleconferencing within an interactive television system |
| US20140298409A1 (en) * | 2006-12-20 | 2014-10-02 | Dst Technologies, Inc. | Secure Processing of Secure Information in a Non-Secure Environment |
| US9226121B2 (en) * | 2013-08-02 | 2015-12-29 | Whatsapp Inc. | Voice communications with real-time status notifications |
| US20150092006A1 (en) * | 2013-10-01 | 2015-04-02 | Filmstrip, Inc. | Image with audio conversation system and method utilizing a wearable mobile device |
| JP6464411B6 (ja) * | 2015-02-25 | 2019-03-13 | Dynabook株式会社 | 電子機器、方法及びプログラム |
| US10085064B2 (en) * | 2016-12-28 | 2018-09-25 | Facebook, Inc. | Aggregation of media effects |
-
2018
- 2018-09-04 CN CN201811027779.0A patent/CN109039872B/zh active Active
-
2019
- 2019-09-04 US US17/257,563 patent/US20210266633A1/en not_active Abandoned
- 2019-09-04 WO PCT/CN2019/104421 patent/WO2020048490A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104333770A (zh) * | 2014-11-20 | 2015-02-04 | 广州华多网络科技有限公司 | 一种视频直播的方法以及装置 |
| CN105657482A (zh) * | 2016-03-28 | 2016-06-08 | 广州华多网络科技有限公司 | 一种语音弹幕的实现方法及装置 |
| CN108055577A (zh) * | 2017-12-18 | 2018-05-18 | 北京奇艺世纪科技有限公司 | 一种直播交互方法、系统、装置及电子设备 |
| CN108259989A (zh) * | 2018-01-19 | 2018-07-06 | 广州华多网络科技有限公司 | 视频直播的方法、计算机可读存储介质和终端设备 |
| CN109039872A (zh) * | 2018-09-04 | 2018-12-18 | 北京达佳互联信息技术有限公司 | 实时语音信息的交互方法、装置、电子设备及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20210266633A1 (en) | 2021-08-26 |
| CN109039872A (zh) | 2018-12-18 |
| CN109039872B (zh) | 2020-04-17 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020048490A1 (zh) | 实时语音信息的交互方法、装置、电子设备及存储介质 | |
| JP6285615B2 (ja) | リモートアシスタンス方法、クライアント、プログラム及び記録媒体 | |
| CN106507207B (zh) | 直播应用中互动的方法及装置 | |
| CN109348239B (zh) | 直播片段处理方法、装置、电子设备及存储介质 | |
| US20220291897A1 (en) | Method and device for playing voice, electronic device, and storage medium | |
| WO2017092247A1 (zh) | 一种播放多媒体数据的方法、装置及系统 | |
| WO2022028234A1 (zh) | 直播间分享方法及装置 | |
| WO2015096648A1 (zh) | 智能电视中的视频分享方法及系统 | |
| WO2017181551A1 (zh) | 视频处理方法及装置 | |
| CN106484856B (zh) | 音频播放方法及装置 | |
| JP6385429B2 (ja) | ストリーム・メディア・データを再生する方法および装置 | |
| CN109245997B (zh) | 语音消息播放方法及装置 | |
| TW200816697A (en) | Method and apparatus for mobile personal video recorder | |
| CN106534953A (zh) | 直播应用中的视频转播方法及控制终端 | |
| CN106911967A (zh) | 直播回放方法及装置 | |
| CN110719530A (zh) | 一种视频播放方法、装置、电子设备及存储介质 | |
| CN113111220A (zh) | 视频处理方法、装置、设备、服务器及存储介质 | |
| JP2017501598A5 (zh) | ||
| CN107333182A (zh) | 多媒体文件的播放方法及装置 | |
| CN107277628A (zh) | 视频预览显示方法及装置 | |
| CN106385614A (zh) | 画面合成方法及装置 | |
| CN108449605B (zh) | 信息同步播放方法、装置、设备、系统及存储介质 | |
| US20220078221A1 (en) | Interactive method and apparatus for multimedia service | |
| CN112764636A (zh) | 视频处理方法、装置、电子设备和计算机可读存储介质 | |
| CN106375846A (zh) | 直播音频的处理方法及装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19857254 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19857254 Country of ref document: EP Kind code of ref document: A1 |