WO2024093442A1 - 用于查看视听内容的方法、装置、设备和存储介质 - Google Patents
用于查看视听内容的方法、装置、设备和存储介质 Download PDFInfo
- Publication number
- WO2024093442A1 WO2024093442A1 PCT/CN2023/113406 CN2023113406W WO2024093442A1 WO 2024093442 A1 WO2024093442 A1 WO 2024093442A1 CN 2023113406 W CN2023113406 W CN 2023113406W WO 2024093442 A1 WO2024093442 A1 WO 2024093442A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- speech
- timeline
- content
- audio
- speaker
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/472—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
- H04N21/47217—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content for controlling playback functions for recorded or on-demand content, e.g. using progress bars, mode or play-point indicators or bookmarks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/78—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/783—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/78—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/783—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
- G06F16/7834—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using audio features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/78—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/783—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
- G06F16/7844—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using original textual content or text extracted from visual content or transcript of audio data
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/22—Interactive procedures; Man-machine interfaces
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L51/00—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
- H04L51/07—User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail characterised by the inclusion of specific contents
- H04L51/10—Multimedia information
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/431—Generation of visual interfaces for content selection or interaction; Content or additional data rendering
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/472—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N7/00—Television systems
- H04N7/14—Systems for two-way working
- H04N7/15—Conference systems
Definitions
- Example embodiments of the present disclosure relate generally to the field of computers, and more particularly to methods, devices, apparatuses, and computer-readable storage media for viewing audiovisual content.
- the Internet has become the main platform for people to obtain and share content.
- people can use the Internet to publish a variety of content, or receive content shared by other users.
- audiovisual content e.g., audio content or video content
- people can use a player to play a speech or a video or audio recording of a meeting shared by other users.
- a method for viewing audio-visual content includes: receiving a selection of a plurality of text segments, the plurality of text segments corresponding to a plurality of parts in target audio-visual content, the plurality of parts at least including a first part and a second part that are not continuous in the target audio-visual content; causing segment audio-visual content to be created based on at least the plurality of parts of the target audio-visual content, wherein the first part and the second part are continuous in the segment audio-visual content; And present a sharing entrance for sharing the audio-visual content clips.
- a device for viewing audio-visual content includes a receiving module configured to receive selections for multiple text segments, the multiple text segments corresponding to multiple parts in the target audio-visual content, the multiple parts at least including a first part and a second part that are discontinuous in the target audio-visual content; a control module configured to create segmented audio-visual content based on at least the multiple parts of the target audio-visual content, wherein the first part and the second part are continuous in the segmented audio-visual content; and a presentation module configured to present a sharing entry for sharing the segmented audio-visual content.
- an electronic device in a third aspect of the present disclosure, includes at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device executes the method of the first aspect.
- a computer-readable storage medium wherein a computer program is stored on the medium, and when the program is executed by a processor, the method of the first aspect is implemented.
- a playback system in a fifth aspect of the present disclosure, includes: a main timeline, which at least indicates the current playback position of the audio-visual content; and at least one speech timeline, which is used to indicate the temporal distribution of the speech content of at least one speaker associated with the audio-visual content.
- FIG1 shows a schematic diagram of a conventional audio-visual content player
- FIGS. 2A to 2C show schematic diagrams of example playback systems according to some embodiments of the present disclosure
- 3A and 3B illustrate example viewing interfaces for audiovisual content according to some embodiments of the present disclosure
- FIGS. 4A and 4B are schematic diagrams showing sharing of audio-visual content segments according to some embodiments of the present disclosure.
- FIG5 illustrates a flow chart of an example process for viewing audiovisual content according to some embodiments of the present disclosure
- FIG6 shows a block diagram of an apparatus for viewing audiovisual content according to some embodiments of the present disclosure.
- FIG. 7 shows a block diagram of a device capable of implementing various embodiments of the present disclosure.
- Figure 1 shows a schematic diagram of a traditional audio-visual content player 100. As shown in Figure 1, in the player 100, people usually need to drag the time axis control to locate the desired playback moment.
- the audiovisual content has a length of more than 1 hour, which makes it difficult for the user to quickly locate the desired playback position through the time axis.
- the embodiments of the present disclosure provide a system for playing audio-visual content (audio content or video content).
- the system may include a main timeline to at least indicate the current playback position of the audio-visual content.
- the system may also include at least one speech timeline, which is used to indicate the distribution of speech content of at least one speaker associated with the audio-visual content in time.
- an embodiment of the present disclosure also provides a solution for viewing audio-visual content.
- a viewing interface for audio-visual content can be provided, wherein the viewing interface includes a playback control for playing the audio-visual content.
- at least one speech timeline can be presented in the playback control, and the at least one speech timeline is used to indicate the distribution of speech content of at least one speaker associated with the audio-visual content in time.
- the embodiments of the present disclosure can provide a speech timeline in the playback system or playback control to provide a time distribution corresponding to the speech content of the speaker associated with the audio-visual content.
- the implementation of the present disclosure can facilitate users to view the part corresponding to a specific speaker, thereby improving the efficiency of users in obtaining desired content.
- embodiments of the present disclosure can utilize a timeline to provide richer information about audio-visual content.
- FIG2A shows a schematic diagram 200A of an example playback system 205 according to some embodiments of the present disclosure.
- the playback system 205 (also referred to as a player 205 or a playback control 205) can be used to play corresponding audio-visual content.
- the playback system 205 can be provided by, for example, an appropriate electronic device, examples of which can include, but are not limited to, a desktop computer, a laptop computer, a smart phone, a tablet computer, a personal digital assistant, or a smart wearable device.
- the audiovisual content may include audio and video files locally stored in the audiovisual system 205, audio and video files stored in the cloud, or audio and video streams.
- it may include a playback stream of recorded audio-visual content (eg, a conference recording), or a live stream of live audio-visual content.
- the playback system 205 may include a main timeline 210.
- the main timeline 210 may indicate the current playback position of the audio-visual content, that is, the playback progress.
- the main timeline 210 may include a playback position indicator 215 to indicate the time point at which the audio-visual content is currently being played.
- the length information of the audiovisual content is fixed, and the total length of the time axis can correspond to the total duration of the time content.
- the play position indicator 215 can be set accordingly according to the corresponding relationship between the position and the play time.
- the play position indicator 215 may, for example, always be set to the far right of the main timeline 210.
- the user may, for example, jump back to the corresponding time point by moving the play position indicator 215.
- the main timeline 210 may also present graphic information corresponding to the audio waveform of the audio-visual content.
- the user can more conveniently understand which parts of the audio-visual content are worth paying attention to, and which parts, such as those with less audio waveforms, can be temporarily ignored.
- a playback system 205 can improve the efficiency of content acquisition for users.
- the audiovisual content may be, for example, recorded content about an online conference. Accordingly, as shown in FIG1 , the main timeline 210 may also present, for example, an interaction identifier 220 corresponding to an interaction behavior in the online conference.
- Such an interaction mark 220 may be set at a corresponding position of the main time axis 210 to indicate that a corresponding interaction behavior has occurred at a corresponding moment.
- different graphics of the interaction mark 220 may correspond to different interaction behaviors.
- the main timeline 210 may include, for example, an interactive identifier 220 for indicating file sharing in an online conference.
- an interactive identifier 220 for indicating file sharing in an online conference.
- the playback system 205 can, for example, guide the user to obtain descriptive information about the file sharing. For example, when the user hovers over the interactive mark 220 with a mouse, the playback system 205 can, for example, indicate the information of the shared file in a floating window, such as the file name, format, size, sharer, etc. In another example, if the user clicks the interactive mark 220, the playback system 205 can, for example, guide the user to obtain the content of the shared file, for example, can guide the user to jump to the online viewing interface of the file.
- the main timeline 210 may include, for example, an interactive identifier 220 for indicating an online chat in an online conference.
- the online chat herein refers to any appropriate chat based on text, emoticons, images, and/or audio, for example, using an instant messaging tool of an online conference.
- the graphic identifier of the interactive identifier 220 may be determined, for example, based on the content of the online chat.
- the graphic identifier of the interactive identifier 220 may be determined, for example, by a graphic identifier (e.g., an avatar) of a user participating in the chat.
- the playback system 205 can guide the user to obtain descriptive information about the online chat. For example, when the user hovers over the interactive mark 220 with a mouse, the playback system 205 can indicate the information of the online chat, such as the participants of the online chat, the chat content, etc., in a floating window. In another example, if the user clicks the interactive mark 220, the playback system 205 can guide the user to obtain the complete content of the previous chat, for example, guide the user to jump to the viewing interface of the chat content in the meeting.
- the main timeline 210 may include, for example, an interactive identifier 220 for indicating comments in an online conference.
- the comments here may include, for example, any appropriate comments based on text, expressions, images, and/or audio.
- a user's like may also be understood as a comment on the corresponding content.
- the graphic identifier of the interactive identifier 220 may be determined, for example, based on the content and/or type of the comment. For example, if it is an expression-based comment, the graphic representation of the interactive identifier 220 may be generated based on the expression.
- the playback system 205 when the user selects the interactive indicator 220, the playback system 205 For example, the user may be guided to obtain descriptive information about the comment. For example, when the user hovers over the interactive mark 220 with a mouse, the playback system 205 may indicate the information of the comment, such as the commenter, comment time, and comment reply, etc., in a floating window. In another example, if the user clicks the interactive mark 220, the playback system 205 may guide the user to jump to the interface for viewing comments to obtain more abundant information about the comment.
- the main timeline 210 may also present speaker information indicating the temporal distribution of speech content of at least one speaker associated with the audiovisual content.
- the main timeline 210 may also be identified as a type of speech timeline.
- the main timeline 210 may assign a corresponding color mark to each speaker. Accordingly, the color distribution on the main timeline 210 may be used to indicate which one or more speakers the corresponding time period corresponds to. It should be understood that other appropriate styles may also be used to use the main timeline 210 to indicate the distribution of the speech content of the speaker in time.
- the playback system 205 may further include a viewing portal 230 for viewing the speech timeline.
- the viewing portal 230 may indicate a graphic identifier (eg, avatar) of one or more speakers associated with the audiovisual content.
- the playback system 205 may present, for example, a speech timeline 240 - 1 and a speech timeline 240 - 2 (individually or collectively referred to as speech timelines 240 ).
- the speech timeline 240 may be used to indicate the distribution of speech content of at least one speaker associated with the audio-visual content over time. For example, if the speaker made a speech at the corresponding moment, the speech timeline 240 may be filled with a first graphic; on the contrary, if the speaker did not make a speech at the corresponding moment, the speech timeline 240 may be filled with a second graphic. Thus, the user can intuitively understand at what moment each speaker made a speech.
- the speech timeline 240 may also be similarly Graphic information corresponding to the audio waveform of the audio-visual content corresponding to the speaker is presented. Based on this method, the user can also intuitively understand at which moments the speaker did not speak and at which moments the speaker spoke frequently. Such information is more helpful for users to quickly obtain the desired content.
- the number of speech timelines 240 may be determined based on the number of speakers participating in the audiovisual content. In some embodiments, the number of such speakers may be determined by the number of terminals participating in the online conference. For example, multiple conference participants may access the online conference through the same terminal (or use the same account), and such multiple participants may be identified as the same speaker, although they may include multiple different speakers.
- the number of such speakers may be determined based on the number of speakers in the audiovisual content. It should be understood that any appropriate speaker recognition technology may be used to determine the corresponding speakers in the audiovisual content, and the present disclosure is not intended to be limited thereto.
- the playback system 205 may present a speech timeline 240 corresponding to all speakers of the audio-visual content.
- the audio-visual content may include two speakers (“speaker 1” and “speaker 2”). Accordingly, the presentation order of the corresponding speech timeline 240-1 and speech timeline 240-2 in the playback system 205 may be determined based on the information of the speakers.
- the presentation order of the speech timeline may be determined based on the text identifier of the speaker, such as a user name or nickname of the speaker, and the presentation order of the speech timeline may be based on the order of the text identifier of the speaker.
- the presentation order of the speech timeline can be determined based on the proportion of the speech content of the speaker. For example, if the speech content proportion of "Speaker 1" reaches “70%”, which is greater than the speech content proportion of "Speaker 2" "30%", then the speech timeline 240-1 can be presented in priority over the speech timeline 240-2.
- the presentation order of the speech timeline can be determined based on the start time of the speech content of the speaker.
- the time is, for example, the first minute after the meeting starts, which is earlier than the start time of the speech content of "Speaker 2", for example, the third minute after the meeting starts. Accordingly, the speech timeline 240-1 can be presented in priority to the speech timeline 240-2.
- the playback system 205 may also present the description information of the corresponding speaker in association with the speech timeline 240.
- the speech timeline 240-1 may have a text identifier (e.g., a user name or nickname) of the corresponding speaker.
- the speech timeline 240-1 may also have a graphic identifier (e.g., an avatar) of the corresponding speaker.
- the playback system 205 may also present the percentage information of the speech content of the corresponding speaker in association with at least one speech timeline.
- the speech timeline 240-1 may include the percentage "XX%" of the speech content of "speaker 1".
- the speech timeline 240 may also present an interaction identifier (not shown in FIG. 2B ) for indicating an interaction behavior associated with a corresponding speaker in the online conference.
- such interactive behaviors refer to corresponding interactive behaviors in which the corresponding speaker participates, such as the file sharing, online chatting or commenting behaviors discussed above.
- the interactive logic of the interactive identifiers presented on the speech timeline 240 may be similar to the interactive identifiers 220 discussed above, which will not be described in detail here.
- the speech timeline 240 - 1 may also be automatically collapsed or expanded in response to the user's selection of the viewing portal 230 .
- the playback system may always provide the speech timeline about all speakers by default, regardless of the selection of the viewing portal 230 .
- the playback system 205 may also provide a search portal 250 for the speech timeline. Using the search portal 250, a user may initiate a viewing request associated with a specific speaker.
- the playback system 205 may present visual elements associated with all speakers associated with the audio-visual content.
- visual elements may include, for example, a text identifier of the speaker (e.g., a user name or nickname). or a graphic identifier (e.g., an avatar).
- the playback system 205 may receive a user's selection of a specific visual element from among the multiple visual elements to determine that the user desires to view the speech timeline of the speaker corresponding to the selected visual element. For example, the user may click on the avatar of "Speaker 1" so that the playback system 205 only presents the speech timeline 240-1 corresponding to "Speaker 1" but not the speech timeline 240-2.
- the user may also provide input indicating the target speaker by viewing the entry 250.
- the user may enter at least part of the nickname or user name of "Speaker 1" to automatically match to "Speaker 1" and cause the playback system 205 to correspondingly present the speech timeline 240-1 corresponding to "Speaker 1" instead of presenting the speech timeline 240-2.
- the search entry 250 may be provided independently of the viewing entry 230.
- the playback system 205 may also provide a search entry 250 for viewing a specific speaker.
- the search portal 250 may be provided, for example, dependent on the viewing portal 230. That is, only when the viewing portal 230 is activated and the speech timelines of all speakers are presented, the search portal 250 is provided accordingly for quickly filtering or locating a specific speech timeline.
- the speech timeline 240 may also support various types of user interactions. For example, as shown in FIG2C , the user may click on a position 260 in the speech timeline 240 - 1 to indicate that the audiovisual content is expected to be played starting from that position.
- the playback system 205 can play the audiovisual content from the time point 270 corresponding to the position 260.
- the playback system 205 can play the audiovisual content continuously from the time point 270. For example, if the time point 270 is the moment "5 minutes and 30 seconds", the audiovisual content will be played continuously from "5 minutes and 30 seconds" until the end.
- the playback system 205 may also play part of the audiovisual content corresponding to "Speaker 1" from time point 270. That is, the playback system 205 may only play part of the audiovisual content of "Speaker 1" corresponding to the speech timeline 240-1, and play it from time point 270, thereby achieving the effect of only listening to a specific speaker.
- the playback system 205 can make the partial audio-visual content corresponding to "Speaker 1" play from the beginning, that is, only play the partial audio-visual content corresponding to "Speaker 1" in the audio-visual content.
- the various features discussed above can be provided independently or in a combination different from that shown in FIGS. 2A to 2C .
- the timeline of the playback system can be a graphical style similar to the timeline of a conventional playback system, and does not necessarily have to be used to indicate an audio waveform.
- the playback system 205 can also be used to play real-time audio-visual content (e.g., live audio and video streams).
- the speech timeline discussed above can be used to indicate the temporal distribution of the historical speech content of at least one speaker associated with the historical portion of the real-time audio-visual content.
- the speech timeline can be presented in a graphical manner: the temporal distribution of the historical speech content of each speaker from the start of the live broadcast to the current moment.
- the embodiments of the present disclosure may also provide a viewing interface for audio-visual content.
- a viewing interface may be, for example, a playback interface for recorded content, or a live interface for real-time content.
- the following uses the "meeting minutes" scenario as an example of viewing audio-visual content, but such a scenario is only exemplary, and the embodiments of the present disclosure may also be applied to other appropriate scenarios.
- FIG. 3A shows an example viewing interface 300 according to some embodiments of the present disclosure.
- the viewing interface 300 may include a playback control 310.
- the playback control 310 may be implemented, for example, using the playback system 205 discussed above.
- the control panel 310 may include, for example, a main timeline 312 and speech timelines 314 - 1 and 314 - 2 (individually or collectively referred to as speech timelines 314 ).
- the viewing interface 300 further includes a text control 320 for presenting text content corresponding to the audio-visual content.
- the text content may be generated based on the audio of the audio-visual content. Taking the audio-visual content as a meeting record as an example, the text content may be generated based on voice recognition of the audio of each speaker in the meeting. Taking the audio-visual content as a real-time live broadcast content as an example, the text content may be generated based on real-time voice recognition of each speaker.
- the user may select the speech timeline 314 - 1 , for example, and accordingly, the text content 322 corresponding to “Speaker 1 ” may be adjusted to be highlighted in the text control 320 relative to other text content 324 of other speakers.
- making the text content 322 highlighted relative to other text content 324 may include, for example, increasing the prominence of the text content 322 displayed in the text control 320.
- the display style e.g., text color, background color, boldness, font size, underline
- the text content 322 may be bolded or highlighted.
- making the text content 322 highlighted relative to the text content 324 may also include, for example, reducing the prominence of the other text content 324 displayed in the text control.
- the display style e.g., text color, background color, boldness, font size, underline
- the text color of the other text content 324 may be gray, thereby forming a contrast with the black text content 322.
- the user may also select a specific position in the speech timeline to trigger the audiovisual content to be played starting from the corresponding moment.
- the text content corresponding to the specific position may also be adjusted to be highlighted in the text control 320.
- the text content presented in the text control 320 always corresponds to the moment at which the audiovisual content is currently playing.
- the text content of the segment corresponding to the moment e.g., the text corresponding to a certain paragraph spoken by the speaker
- the display style of the text content of the segment can also be adjusted to highlight the text content of the segment. For example, one or more words corresponding to the time point can be highlighted to be highlighted.
- embodiments of the present disclosure may also support sharing of audiovisual content segments based on speech timelines.
- a user may select one or more speech timelines (e.g., speech timeline 430 - 1 ) of the multiple speech timelines in the playback control 410 (or the playback system 410 ) for sharing.
- the audio-visual content corresponding to the speech timeline 430-1 can be generated for sharing.
- the entire speech content of "Speaker 1" can be used to generate independent audio-visual content, for example, for sharing with other users or organizations.
- the user can also select one or more time segments in the speech timelines 430-1 and 430-2, for example, time segment 440-1, time segment 440-2, and time segment 440-3. Accordingly, after the user clicks on the sharing entry 420 (i.e., sends a sharing request), multiple discrete audio-visual content segments corresponding to the time segments 440-1, 440-2, and 440-3 can be combined to generate independent segment audio-visual content, for example, for sharing with other users or organizations.
- the embodiments of the present disclosure can support users to more efficiently share audio-visual content by selecting a speech timeline or a time segment, thereby improving the efficiency of audio-visual content sharing and improving the efficiency of information acquisition by the shared party.
- the embodiments of the present disclosure also support users to select non-continuous segments to create, which further improves the flexibility of sharing audio-visual content segments.
- FIG5 shows a flow chart of an example process 500 for viewing audiovisual content according to some embodiments of the present disclosure.
- Process 500 may be implemented at a suitable electronic device. Examples of such electronic devices may include, but are not limited to, desktop computers, laptop computers, smart phones, tablet computers, personal digital assistants, or smart wearable devices.
- the electronic device provides a viewing interface for audio-visual content, where the viewing interface includes a play control for playing the audio-visual content.
- the electronic device presents at least one speech timeline in the playback control, where the at least one speech timeline is used to indicate the temporal distribution of speech content of at least one speaker associated with the audio-visual content.
- the viewing interface further includes a text control for presenting text content corresponding to the audiovisual content, the text content being generated based on the audio of the audiovisual content.
- the method also includes: in response to selection of a first speech timeline in at least one speech timeline, causing first text content corresponding to a first speaker in the text content to be highlighted in the text control relative to second text content of other speakers, wherein the first speech timeline corresponds to the first speaker.
- first text content corresponding to a first speaker in the text content is highlighted in a text control relative to second text content of other speakers: the prominence of the first text content displayed in the text control is increased; and/or the prominence of the second text content displayed in the text control is reduced.
- the method further includes: receiving a selection of a first position in a first speech timeline of at least one speech timeline; and causing text content in the text content corresponding to the first position to be highlighted in the text control.
- presenting at least one speech timeline in the playback control includes: presenting a viewing entry for viewing the speech timeline in the playback control; and presenting at least one speech timeline in the playback control in response to a selection of the viewing entry.
- At least one speech timeline includes multiple speech timelines
- the presentation order of the multiple speech timelines in the playback control is determined based on at least one of the following: text identifiers of multiple speakers corresponding to the multiple speech timelines, The proportion of speeches by multiple speakers, or the start time of speeches by multiple speakers.
- presenting at least one speech timeline in a playback control includes: receiving a viewing request associated with a target speaker; and presenting a target speech timeline corresponding to the target speaker, the target speech timeline being used to indicate the temporal distribution of the target speech content of the target speaker.
- receiving a viewing request associated with a target speaker includes: presenting multiple visual elements associated with multiple speakers associated with audio-visual content; and receiving a viewing request associated with the target speaker based on a preset operation of a target visual element corresponding to the target speaker among the multiple visual elements.
- receiving a view request associated with the target speaker includes receiving a view request associated with the target speaker based on an input indicating the target speaker.
- the method further includes: receiving a selection of a second position in a first speech timeline of at least one speech timeline; and causing a corresponding portion of the audiovisual content to be played from a time point corresponding to the second position.
- the first speech timeline corresponds to the first speaker
- causing at least part of the audio-visual content to be played from a time point corresponding to the second position includes: causing the audio-visual content to be played continuously from the time point; or causing part of the audio-visual content corresponding to the first speaker in the audio-visual content to be played from the time point.
- the method further comprises: presenting description information of the corresponding speaker in association with at least one speech timeline, the description information being generated based on a text identifier and/or a graphic identifier of the speaker.
- the method further includes: presenting the proportion information of the speech content of the corresponding speaker in association with at least one speech timeline.
- the playback controls further include a main timeline for presenting graphical information corresponding to an audio waveform of the audiovisual content.
- the audiovisual content is an audiovisual recording of an online meeting
- the playback control further includes a main timeline, which is used to present a first interaction identifier corresponding to a first interaction behavior in the online meeting.
- At least one speech timeline also presents a timeline for indicating the online meeting A second interaction identifier of a second interaction behavior associated with the corresponding speaker.
- the first interactive behavior and/or the second interactive behavior includes at least one of the following: file sharing, online chatting, and commenting.
- the method further includes: in response to a first selection of the first interaction identifier, presenting first description information for the first interaction behavior; and/or in response to a second selection of the second interaction identifier, presenting second description information for the second interaction behavior.
- the method further includes: receiving a selection of at least one time segment in at least one speech timeline; and based on a first sharing request associated with the at least one time segment, generating a first segment of audiovisual content corresponding to the at least one time segment for sharing.
- the method also includes: receiving a selection of a group of speech timelines in at least one speech timeline, the group of timelines including one or more speech timelines; and based on a second sharing request associated with the group of timelines, causing a second segment of audio-visual content corresponding to the group of speech timelines to be generated for sharing.
- the audiovisual content includes real-time audiovisual content
- at least one speech timeline is used to indicate the temporal distribution of historical speech content of at least one speaker associated with the historical portion of the real-time audiovisual content.
- Fig. 6 shows a schematic structural block diagram of a device 600 for viewing audio-visual content according to some embodiments of the present disclosure.
- the apparatus 600 includes a providing module 610 configured to provide a viewing interface for audio-visual content, where the viewing interface includes a playback control for playing the audio-visual content.
- the apparatus 600 further includes a presentation module 620 configured to present at least one speech timeline in the playback control, wherein the at least one speech timeline is used to indicate the temporal distribution of speech content of at least one speaker associated with the audio-visual content.
- the viewing interface further includes a text control for presenting text content corresponding to the audiovisual content, the text content being generated based on the audio of the audiovisual content. become.
- the presentation module 620 is further configured to: in response to selection of a first speech timeline in at least one speech timeline, cause first text content corresponding to a first speaker in the text content to be highlighted in the text control relative to second text content of other speakers, wherein the first speech timeline corresponds to the first speaker.
- first text content corresponding to a first speaker in the text content is highlighted in a text control relative to second text content of other speakers: the prominence of the first text content displayed in the text control is increased; and/or the prominence of the second text content displayed in the text control is reduced.
- the presentation module 620 is further configured to: receive a selection of a first position in a first speech timeline of at least one speech timeline; and highlight the text content corresponding to the first position in the text content in the text control.
- the presentation module 620 is further configured to: present a viewing entry for viewing the speech timeline in the playback control; and present at least one speech timeline in the playback control in response to a selection of the viewing entry.
- At least one speech timeline includes multiple speech timelines
- the presentation order of the multiple speech timelines in the playback control is determined based on at least one of the following: text identifiers of the multiple speakers corresponding to the multiple speech timelines, the proportion of the speech content of the multiple speakers, or the start time of the speech content of the multiple speakers.
- the presentation module 620 is further configured to: receive a viewing request associated with a target speaker; and present a target speech timeline corresponding to the target speaker, the target speech timeline being used to indicate the temporal distribution of the target speech content of the target speaker.
- the presentation module 620 is further configured to: present multiple visual elements associated with multiple speakers associated with the audio-visual content; and receive a viewing request associated with a target speaker based on a preset operation of a target visual element corresponding to the target speaker among the multiple visual elements.
- presentation module 620 is further configured to: based on the input indicating the target speaker, receive a viewing request associated with the target speaker.
- the presentation module 620 is further configured to: receive a selection of a second position in a first speech timeline of at least one speech timeline; and cause the corresponding portion of the audiovisual content to be played from a time point corresponding to the second position.
- the first speech timeline corresponds to the first speaker
- the presentation module 620 is further configured to: enable the audiovisual content to be played continuously from a point in time; or enable the portion of the audiovisual content corresponding to the first speaker to be played from a point in time.
- the presentation module 620 is further configured to present description information of the corresponding speaker in association with at least one speech timeline, where the description information is generated based on a text identifier and/or a graphic identifier of the speaker.
- the presentation module 620 is further configured to present the proportion information of the speech content of the corresponding speaker in association with at least one speech timeline.
- the playback controls further include a main timeline for presenting graphical information corresponding to an audio waveform of the audiovisual content.
- the audiovisual content is an audiovisual recording of an online meeting
- the playback control further includes a main timeline, which is used to present a first interaction identifier corresponding to a first interaction behavior in the online meeting.
- At least one speech timeline further presents a second interaction identifier for indicating a second interaction behavior associated with the corresponding speaker in the online conference.
- the first interactive behavior and/or the second interactive behavior includes at least one of the following: file sharing, online chatting, and commenting.
- the presentation module 620 is further configured to: present first description information for a first interaction behavior in response to a first selection of a first interaction identifier; and/or present second description information for a second interaction behavior in response to a second selection of a second interaction identifier.
- the presentation module 620 is further configured to: receive a selection of at least one time segment in at least one speech timeline; and based on a first sharing request associated with the at least one time segment, generate a first segment of audio-visual content corresponding to the at least one time segment for sharing.
- the presentation module 620 is further configured to: receive a request for at least one A group of speech timelines is selected from the speech timelines, wherein the group of timelines includes one or more speech timelines; and based on a second sharing request associated with the group of timelines, a second segment of audio-visual content corresponding to the group of speech timelines is generated for sharing.
- the audiovisual content includes real-time audiovisual content
- at least one speech timeline is used to indicate the temporal distribution of historical speech content of at least one speaker associated with the historical portion of the real-time audiovisual content.
- the units included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof.
- one or more units can be implemented using software and/or firmware, such as machine executable instructions stored on a storage medium.
- some or all of the units in the device 600 can be implemented at least in part by one or more hardware logic components.
- exemplary types of hardware logic components include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
- Figure 7 shows a block diagram of a computing device/server 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the computing device/server 700 shown in Figure 7 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein.
- computing device/server 700 is in the form of a general computing device.
- the components of computing device/server 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 760, and one or more output devices 760.
- Processing unit 710 may be an actual or virtual processor and is capable of performing various processes according to a program stored in memory 720. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capabilities of computing device/server 700.
- the computing device/server 700 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device/server 700, including but not limited to volatile and nonvolatile media, removable and non-removable media.
- the memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), Non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof.
- Storage device 730 may be removable or non-removable media and may include machine-readable media such as a flash drive, a disk, or any other media that may be capable of storing information and/or data (e.g., training data for training) and may be accessed within computing device/server 700.
- machine-readable media such as a flash drive, a disk, or any other media that may be capable of storing information and/or data (e.g., training data for training) and may be accessed within computing device/server 700.
- the computing device/server 700 may further include additional removable/non-removable, volatile/non-volatile storage media.
- a disk drive for reading or writing from a removable, non-volatile disk e.g., a “floppy disk”
- an optical drive for reading or writing from a removable, non-volatile optical disk may be provided.
- each drive may be connected to a bus (not shown) by one or more data media interfaces.
- the memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
- the communication unit 740 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device/server 700 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device/server 700 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
- PC network personal computer
- Input device 750 may be one or more input devices, such as a mouse, keyboard, trackball, etc.
- Output device 760 may be one or more output devices, such as a display, speaker, printer, etc.
- Computing device/server 700 may also communicate with one or more external devices (not shown) as needed, such as storage devices, display devices, etc., with one or more devices that enable users to interact with computing device/server 700, or with any device that enables computing device/server 700 to communicate with one or more other computing devices (e.g., network card, modem, etc.) through communication unit 740. Such communication may be performed via an input/output (I/O) interface (not shown).
- I/O input/output
- a computer-readable storage medium on which one or more computer instructions are stored, wherein the one or more computer instructions are executed by a processor to implement the method described above.
- These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions/actions specified in one or more boxes in the flowchart and/or block diagram is generated.
- These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and/or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions/actions specified in one or more boxes in the flowchart and/or block diagram.
- Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions/actions specified in one or more boxes in the flowchart and/or block diagram.
- each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification.
- the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved.
- each square box in the block diagram and/or flow chart, and the combination of the square boxes in the block diagram and/or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Theoretical Computer Science (AREA)
- Library & Information Science (AREA)
- Databases & Information Systems (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- Human Computer Interaction (AREA)
- Computer Networks & Wireless Communication (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Acoustics & Sound (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
Abstract
Description
Claims (32)
- 一种查看视听内容的方法,包括:提供针对视听内容的查看界面,所述查看界面包括用于播放所述视听内容的播放控件;以及在所述播放控件中呈现至少一个发言时间轴,所述至少一个发言时间轴用于指示与所述视听内容相关联的至少一个发言方的发言内容在时间上的分布。
- 根据权利要求1所述的方法,其中所述查看界面还包括文本控件,所述文本控件用于呈现与所述视听内容对应的文本内容,所述文本内容是基于所述视听内容的音频而被生成。
- 根据权利要求2所述的方法,还包括:响应于与所述至少一个发言时间轴中第一发言时间轴的选择,使所述文本内容中与第一发言方对应的第一文本内容在所述文本控件中相对于其它发言方的第二文本内容被强调显示,其中所述第一发言时间轴对应于所述第一发言方。
- 根据权利要求3所述的方法,其中使所述文本内容中与第一发言方对应的第一文本内容在所述文本控件中相对于其它发言方的第二文本内容被强调显示:提高所述第一文本内容在所述文本控件中被显示的突出程度;和/或降低所述第二文本内容在所述文本控件中被显示的突出程度。
- 根据权利要求2所述的方法,还包括:接收针对所述至少一个发言时间轴的第一发言时间轴中第一位置的选择;以及使所述文本内容中与所述第一位置对应的文本内容在所述文本控件中被突出呈现。
- 根据权利要求1所述的方法,其中在所述播放控件中呈现至少一个发言时间轴包括:在所述播放控件中呈现用于查看发言时间轴的查看入口;以及响应于针对所述查看入口的选择,在所述播放控件中呈现所述至少一个发言时间轴。
- 根据权利要求1所述的方法,其中所述至少一个发言时间轴包括多个发言时间轴,并且所述多个发言时间轴在所述播放控件中的呈现顺序是基于以下至少一项而被确定:与所述多个发言时间轴对应的多个发言方的文本标识,所述多个发言方的发言内容的比例,或所述多个发言方的发言内容的起始时间。
- 根据权利要求1所述的方法,其中在所述播放控件中呈现至少一个发言时间轴包括:接收与目标发言方相关联的查看请求;以及呈现与所述目标发言方对应的目标发言时间轴,所述目标发言时间轴用于指示所述目标发言方的目标发言内容在时间上的分布。
- 根据权利要求8所述的方法,其中接收与目标发言方相关联的查看请求包括:呈现与所述视听内容相关联的多个发言方相关联的多个视觉元素;以及基于与所述多个视觉元素中与所述目标发言方对应的目标视觉元素的预设操作,接收与目标发言方相关联的查看请求。
- 根据权利要求8所述的方法,其中接收与目标发言方相关联的查看请求包括:基于指示目标发言方的输入,接收与目标发言方相关联的查看请求。
- 根据权利要求1所述的方法,还包括:接收针对所述至少一个发言时间轴的第一发言时间轴中第二位置的选择;以及使所述视听内容的对应部分从与所述第二位置对应的时间点处被播放。
- 根据权利要求11所述的方法,其中所述第一发言时间轴对应于第一发言方,并且使所述视听内容的至少部分从与所述第二位置对应的时间点处被播放包括:使所述视听内容从所述时间点处被连续播放;或使所述视听内容中与所述第一发言方对应的部分视听内容从所述时间点处被播放。
- 根据权利要求1所述的方法,还包括:与所述至少一个发言时间轴相关联地呈现相应的发言方的描述信息,所述描述信息基于所述发言方的文本标识和/或图形标识而被生成。
- 根据权利要求1所述的方法,还包括:与所述至少一个发言时间轴相关联地呈现相应的发言方的发言内容的占比信息。
- 根据权利要求1所述的方法,其中所述播放控件还包括主时间轴,所述主时间轴用于呈现与所述视听内容的音频波形相对应的图形信息。
- 根据权利要求1所述的方法,其中所述视听内容是针对在线会议的视听记录,并且所述播放控件还包括主时间轴,所述主时间轴用于呈现与所述在线会议中的第一交互行为相对应的第一交互标识。
- 根据权利要求16所述的方法,其中所述至少一个发言时间轴还呈现用于指示所述在线会议中与对应的发言方相关联的第二交互行为的第二交互标识。
- 根据权利要求16或17所述的方法,其中所述第一交互行为和/或所述第二交互行为包括以下至少一项:文件共享、在线聊天、以及评论。
- 根据权利要求16或17所述的方法,还包括:响应于对所述第一交互标识的第一选择,呈现针对所述第一交互行为的第一描述信息;和/或响应于对所述第二交互标识的第二选择,呈现针对所述第二交互 行为的第二描述信息。
- 根据权利要求1所述的方法,还包括:接收针对所述至少一个发言时间轴中的至少一个时间片段的选择;以及基于与所述至少一个时间片段相关联的第一分享请求,使与所述至少一个时间片段对应的第一片段视听内容被生成以用于分享。
- 根据权利要求1所述的方法,还包括:接收针对所述至少一个发言时间轴中的一组发言时间轴的选择,所述一组时间轴包括一个或多个发言时间轴;以及基于与所述一组时间轴相关联的第二分享请求,使与所述一组发言时间轴对应的第二片段视听内容被生成以用于分享。
- 根据权利要求1所述的方法,其中所述视听内容包括实时视听内容,并且所述至少一个发言时间轴用于指示与所述实时视听内容的历史部分相关联的至少一个发言方的历史发言内容在时间上的分布。
- 一种用于查看视听内容的装置,包括:提供模块,被配置为提供针对视听内容的查看界面,所述查看界面包括用于播放所述视听内容的播放控件;以及呈现模块,被配置为在所述播放控件中呈现至少一个发言时间轴,所述至少一个发言时间轴用于指示与所述视听内容相关联的至少一个发言方的发言内容在时间上的分布。
- 一种播放系统,包括:主时间轴,所述主时间轴至少指示视听内容的当前播放位置;以及至少一个发言时间轴,所述至少一个发言时间轴用于指示与所述视听内容相关联的至少一个发言方的发言内容在时间上的分布。
- 根据权利要求24所述的播放系统,其中所述主时间轴还呈现与所述视听内容的音频波形相对应的图形信息。
- 根据权利要求24所述的播放系统,其中所述视听内容是针对 在线会议的视听记录,所述主时间轴还呈现与所述在线会议中的第一交互行为相对应的第一交互标识。
- 根据权利要求26所述的播放系统,其中所述至少一个发言时间轴还呈现用于指示所述在线会议中与对应的发言方相关联的第二交互行为的第二交互标识。
- 根据权利要求26或27所述的播放系统,其中所述第一交互行为和/或所述第二交互行为包括以下至少一项:文件共享、在线聊天、以及评论。
- 根据权利要求26或27所述的播放系统,其中:对所述第一交互标识的第一选择用于触发针对所述第一交互行为的第一描述信息;和/或对所述第二交互标识的第二选择用于触发呈现针对所述第二交互行为的第二描述信息。
- 根据权利要求24所述的播放系统,其中所述至少一个发言时间轴包括多个发言时间轴,并且所述多个发言时间轴在所述播放系统中的呈现顺序是基于以下至少一项而被确定:与所述多个发言时间轴对应的多个发言方的文本标识,所述多个发言方的发言内容的比例,或所述多个发言方的发言内容的起始时间。
- 一种电子设备,包括:至少一个处理单元;以及至少一个存储器,所述至少一个存储器被耦合到所述至少一个处理单元并且存储用于由所述至少一个处理单元执行的指令,所述指令在由所述至少一个处理单元执行时使所述设备执行根据权利要求1至22中任一项所述的方法。
- 一种计算机可读存储介质,其上存储有计算机程序,所述程序被处理器执行时实现根据权利要求1至22中任一项所述的方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2025525050A JP2026513006A (ja) | 2022-10-31 | 2023-08-16 | オーディオビジュアルコンテンツをチェックするための方法、装置、機器及び記憶媒体 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202211352393.3A CN117956233A (zh) | 2022-10-31 | 2022-10-31 | 用于查看视听内容的方法、装置、设备和存储介质 |
| CN202211352393.3 | 2022-10-31 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| WO2024093442A1 true WO2024093442A1 (zh) | 2024-05-10 |
| WO2024093442A9 WO2024093442A9 (zh) | 2025-06-05 |
Family
ID=90798805
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/113406 Ceased WO2024093442A1 (zh) | 2022-10-31 | 2023-08-16 | 用于查看视听内容的方法、装置、设备和存储介质 |
Country Status (3)
| Country | Link |
|---|---|
| JP (1) | JP2026513006A (zh) |
| CN (1) | CN117956233A (zh) |
| WO (1) | WO2024093442A1 (zh) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119922363A (zh) * | 2023-10-30 | 2025-05-02 | 北京字跳网络技术有限公司 | 用于查看媒体内容的方法、装置、设备和存储介质 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112151041A (zh) * | 2019-06-26 | 2020-12-29 | 北京小米移动软件有限公司 | 基于录音机程序的录音方法、装置、设备及存储介质 |
| CN113194349A (zh) * | 2021-04-25 | 2021-07-30 | 腾讯科技(深圳)有限公司 | 视频播放方法、评论方法、装置、设备及存储介质 |
| CN113326387A (zh) * | 2021-05-31 | 2021-08-31 | 引智科技(深圳)有限公司 | 一种会议信息智能检索方法 |
| JP2021184189A (ja) * | 2020-05-22 | 2021-12-02 | i Smart Technologies株式会社 | オンライン会議システム |
| CN114491087A (zh) * | 2022-01-13 | 2022-05-13 | Oppo广东移动通信有限公司 | 文本处理方法、装置、电子设备以及存储介质 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2015001492A1 (en) * | 2013-07-02 | 2015-01-08 | Family Systems, Limited | Systems and methods for improving audio conferencing services |
| CN114387956B (zh) * | 2020-10-21 | 2025-05-02 | 阿里巴巴集团控股有限公司 | 音频信号处理方法、装置及电子设备 |
| CN113329237B (zh) * | 2021-02-02 | 2023-03-21 | 北京意匠文枢科技有限公司 | 一种呈现事件标签信息的方法与设备 |
| CN114936001A (zh) * | 2022-04-14 | 2022-08-23 | 阿里巴巴(中国)有限公司 | 交互方法、装置及电子设备 |
-
2022
- 2022-10-31 CN CN202211352393.3A patent/CN117956233A/zh active Pending
-
2023
- 2023-08-16 WO PCT/CN2023/113406 patent/WO2024093442A1/zh not_active Ceased
- 2023-08-16 JP JP2025525050A patent/JP2026513006A/ja active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112151041A (zh) * | 2019-06-26 | 2020-12-29 | 北京小米移动软件有限公司 | 基于录音机程序的录音方法、装置、设备及存储介质 |
| JP2021184189A (ja) * | 2020-05-22 | 2021-12-02 | i Smart Technologies株式会社 | オンライン会議システム |
| CN113194349A (zh) * | 2021-04-25 | 2021-07-30 | 腾讯科技(深圳)有限公司 | 视频播放方法、评论方法、装置、设备及存储介质 |
| CN113326387A (zh) * | 2021-05-31 | 2021-08-31 | 引智科技(深圳)有限公司 | 一种会议信息智能检索方法 |
| CN114491087A (zh) * | 2022-01-13 | 2022-05-13 | Oppo广东移动通信有限公司 | 文本处理方法、装置、电子设备以及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN117956233A (zh) | 2024-04-30 |
| JP2026513006A (ja) | 2026-04-22 |
| WO2024093442A9 (zh) | 2025-06-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12015683B2 (en) | Method, apparatus and device for issuing and replying to multimedia content | |
| WO2024041549A1 (zh) | 用于会话消息呈现的方法、装置、设备和存储介质 | |
| US20250301025A1 (en) | Method, apparatus, device and storage medium for live streaming interface interaction | |
| US20250157103A1 (en) | Media stream storyboard generation | |
| CN115509412A (zh) | 用于特效交互的方法、装置、设备和存储介质 | |
| US20250298494A1 (en) | Method, apparatus, device and storage medium for sharing audiovisual content | |
| WO2024093442A1 (zh) | 用于查看视听内容的方法、装置、设备和存储介质 | |
| WO2023226853A1 (zh) | 用于作品转发的方法、装置、设备和存储介质 | |
| CN116450859A (zh) | 用于媒体内容推荐的方法、装置、设备和存储介质 | |
| US20250013479A1 (en) | Method, apparatus, device and storage medium for processing information | |
| CN118838673A (zh) | 故事创作方法、装置、设备和存储介质 | |
| WO2024131577A1 (zh) | 用于创建特效的方法、装置、设备和介质 | |
| WO2023226855A1 (zh) | 用于作品转发的方法、装置、设备和存储介质 | |
| CN114205671A (zh) | 基于场景对齐的视频内容剪辑方法及其装置 | |
| WO2024093937A1 (zh) | 用于查看视听内容的方法、装置、设备和存储介质 | |
| CN119172344B (zh) | 交互的方法、装置、设备和存储介质 | |
| JP2026512319A (ja) | メディアコンテンツを閲覧するための方法、装置、機器及び記憶媒体 | |
| WO2024160260A1 (zh) | 视频处理的方法、装置、设备和存储介质 | |
| WO2025123947A1 (zh) | 用于交互的方法、装置、设备和存储介质 | |
| CN120523892A (zh) | 交互方法、装置、设备和存储介质 | |
| WO2022187011A1 (en) | Information search for a conference service | |
| CN120216635A (zh) | 用于交互引导的方法、装置、设备和存储介质 | |
| CN118672475A (zh) | 内容生成方法及相关设备 | |
| CN121858001A (zh) | 界面交互的方法、装置、设备和存储介质 | |
| WO2025195420A1 (zh) | 用于编辑的方法、装置、设备和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23884381 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2025525050 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2025525050 Country of ref document: JP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23884381 Country of ref document: EP Kind code of ref document: A1 |