WO2017101430A1 - 一种语音弹幕生成方法和装置 - Google Patents

一种语音弹幕生成方法和装置 Download PDF

Info

Publication number
WO2017101430A1
WO2017101430A1 PCT/CN2016/089574 CN2016089574W WO2017101430A1 WO 2017101430 A1 WO2017101430 A1 WO 2017101430A1 CN 2016089574 W CN2016089574 W CN 2016089574W WO 2017101430 A1 WO2017101430 A1 WO 2017101430A1
Authority
WO
WIPO (PCT)
Prior art keywords
voice
user comment
comment
user
form user
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/089574
Other languages
English (en)
French (fr)
Inventor
李虎强
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Le Holdings Beijing Co Ltd
LeTV Information Technology Beijing Co Ltd
Original Assignee
Le Holdings Beijing Co Ltd
LeTV Information Technology Beijing Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Le Holdings Beijing Co Ltd, LeTV Information Technology Beijing Co Ltd filed Critical Le Holdings Beijing Co Ltd
Priority to US15/240,152 priority Critical patent/US20170168660A1/en
Publication of WO2017101430A1 publication Critical patent/WO2017101430A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/475End-user interface for inputting end-user data, e.g. personal identification number [PIN], preference data
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/475End-user interface for inputting end-user data, e.g. personal identification number [PIN], preference data
    • H04N21/4756End-user interface for inputting end-user data, e.g. personal identification number [PIN], preference data for rating content, e.g. scoring a recommended movie
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • H04N21/4402Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display
    • H04N21/440236Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display by media transcoding, e.g. video is transformed into a slideshow of still pictures, audio is converted into text

Definitions

  • the present patent application relates to video optimization, and in particular to a voice barrage generation method and a voice barrage generation device.
  • the object of some embodiments of the present invention is to provide a voice barrage generating method and a voice barrage generating device, which can enable a user to listen to the barrage voice by himself or herself, and give the user more More choices and a better viewing experience.
  • an embodiment of the present invention provides a voice barrage generating method, where the generating method includes: collecting a voice form user comment or a text form user comment; and generating the text according to the collected voice form user comment. Forming the user comment or generating the voice form user comment based on the collected text form user comment; establishing a link of the text form user comment and the voice form user comment; and including a link containing the voice form user comment.
  • the textual user comments are displayed to the user in the form of a barrage.
  • An embodiment of the present invention further provides a voice barrage generating device, where the generating device includes: a collecting mode a block for collecting a voice form user comment or a text form user comment; the processing module, configured to: generate the text form user comment according to the voice form user comment collected by the collecting module, or according to the collected The text form user comment generates the voice form user comment; establishing a link of the text form user comment and the voice form user comment; and the text form user comment including the link of the voice form user comment The form of the barrage is displayed to the user.
  • One embodiment of the present invention provides a computer readable storage medium comprising computer executable instructions that, when executed by at least one processor, cause the processor to perform the above method.
  • An embodiment of the present invention provides an electronic device, the electronic device including: a memory for storing instructions, a processor, and a bus system, the processor and the memory being connected by the bus system; Wherein the processor is configured to execute the instructions stored by the memory and to perform the above method.
  • the voice barrage generation method and the voice barrage generating device are used to first collect user comments, which may be in the form of text or voice, and if the user comments in the form of voice are collected, This generates a textual user comment; if a textual user comment is collected, a voice form user comment is generated accordingly. A link between the corresponding voice form user comment and the text form user comment is then established, and finally the text form user comment containing the link of the voice form user comment is displayed to the user in the form of a barrage after the link is established.
  • the voice barrage generation method and the voice barrage generating device provided by the embodiment of the invention can enable the user to listen to the barrage voice by himself, giving the user more choice and a better viewing experience.
  • FIG. 1 is a flowchart of a method for generating a voice barrage according to an embodiment of the present invention
  • FIG. 2 is a schematic structural diagram of a voice barrage generating apparatus according to an embodiment of the present invention.
  • FIG. 1 is a flowchart of a method for generating a voice barrage according to an embodiment of the present invention.
  • a voice barrage generation method includes: collecting a voice form user comment or a text form user comment; generating the text form user comment according to the collected voice form user comment or according to the collection The textual form user comment to generate the voice form user comment; establishing a link of the text form user comment and the voice form user comment; and the text form user to include the link of the voice form user comment Comments are displayed to the user in the form of a barrage.
  • the user comments are collected.
  • the user comments can be in the form of text or voice.
  • the user comments are judged to be voice-based user comments or text-based user comments.
  • the existing text input box can still be used for input.
  • the formal user comment can provide an opportunity for the user to make a voice form comment using any voice input device in the embodiment of the present invention.
  • a voice form user comment is collected, a text form user comment is generated accordingly; if a text form user comment is collected, a voice form user comment is generated accordingly.
  • machine voice can be utilized to voice the user comments to produce a voice form user comment.
  • a link is established between the corresponding voice form user comment and the text form user comment, so that the user can click the text form user comment to call up the corresponding voice form text comment.
  • the textual form user comment containing the link of the voice form user comment is displayed to the user in the form of a barrage after the link is established. When the user sees the text form barrage, clicking the text form barrage can listen to the voice form barrage.
  • the generating method further includes: recording a time period in which the voice form user comment or the text form user comment is collected in the video playing; and when the video is played again, the voice form is included in the time period.
  • the textual user comment of the link of the user comment is displayed to the user in the form of a barrage.
  • the user comment is collected every time the video is played, and the time period of the user comment is recorded every time a user comment is collected, and the time period may be a period of time, for example, 5 seconds.
  • the video is played again, not only the real-time barrage is displayed, but also the user comments that have been recorded for a certain period of time are displayed to the user in the form of a barrage when the video reaches the corresponding time period, so that not only the user comments and Video content is still closely related, and it also allows users to have the feeling of watching videos with multiple people whenever they watch which video.
  • the first threshold can be freely set, which is related to the set length of the time period and the moving speed of the barrage on the screen.
  • the user comments are more and more, and there may be too slow storage space or random selection of text-based user comments with links containing voice-based user comments.
  • each new piece is generated.
  • the textual form of the link for the voice form user comment corresponds to the deletion of a textual user comment that was previously generated as a link containing a voice form user comment.
  • the second threshold may be freely set, related to the storage space, etc., but in some embodiments it should be much larger than the first threshold.
  • FIG. 2 is a schematic structural diagram of a voice barrage generating apparatus according to an embodiment of the present invention.
  • a voice barrage generating device includes: an acquiring module 1 for collecting a voice form user comment or a text form user comment; and a processing module 2, configured to perform the following operations: according to the collecting The voice form user comment collected by the module 1 generates the text form user comment or generates the voice form user comment according to the collected text form user comment; establishing the text form user comment and the voice form user a link to the comment; and the textual user comment containing the link to the voice form user comment is displayed to the user in the form of a barrage.
  • the acquisition module 1 collects user comments and sends the user comments to the processing module 2, and the processing module 2 receives the user comments.
  • User comments can be in text or in speech. If mining The set is a voice form user comment, and the processing module 2 generates a text form user comment accordingly; if the text form user comment is collected, the processing module 2 generates a voice form user comment accordingly.
  • machine voice can be utilized to voice the user comments to produce a voice form user comment.
  • the processing module 2 establishes a link between the corresponding voice form user comment and the text form user comment, so that the user can click the text form user comment to call up the corresponding voice form text comment.
  • the final processing device 2 displays the textual user comments containing the links of the voice form user comments to the user in the form of a barrage after the link is established.
  • the processing module 2 is further configured to: record a time period in which a voice form user comment or a text form user comment is collected during video playback; and when the video is played again, the voice form user comment will be included in the time period.
  • the textual user comments of the link are displayed to the user in the form of a barrage.
  • the processing module 2 is further configured to: if the number of the text form user comments corresponding to the time period of the link containing the voice form user comment is greater than a first threshold, A textual user comment that randomly selects a link containing a voice form user comment from the textual user comment containing the link of the voice form user comment is displayed to the user in the form of a barrage.
  • the textual user comment of the link containing the voice form user comment of the current time period is stopped, thereby randomly displaying the text form user comment of the link containing the voice form user comment for the next time period.
  • the first threshold can be freely set, which is related to the set length of the time period and the moving speed of the barrage on the screen.
  • the processing module 2 is further configured to: each time a new number of user comments of the text form corresponding to the link of the voice form user comment corresponding to the time period is greater than a second threshold A textual user comment containing a link to a voice form user comment corresponds to deleting a textual user comment of a previously generated link containing a voice form user comment.
  • the second threshold may be freely set, related to storage space, etc., but in some embodiments should be much larger than the first threshold.
  • the user can input a text form user comment through the input box, or input a voice form user comment through the voice collection device.
  • the collecting module 1 collects a text form user comment or a voice form user comment at this time, and sends the text form user comment or the voice form user comment to the processing module 2.
  • the processing module 2 records the time period at this time.
  • the processing module 2 When the processing module 2 receives the text form user comment, the processing module 2 generates the text form user comment as a voice form user comment, and establishes a text form user comment and a generated voice form user comment link; the processing module 2 receives When the voice form user comments, the processing module 2 generates a text form user comment by the voice form user comment, and establishes a link between the voice form user comment and the generated text form user comment. The processing module 2 displays the processed textual user comments containing the links of the voice form user comments to the user.
  • the processing module 2 detects the number of the text form user comments corresponding to the link of the voice form user comment corresponding to the time period just recorded, and when the number is greater than the second threshold, deletes one of the earliest generated voices before the deletion User comments in the form of text in the form of a user comment. If the number is not greater than the second threshold, the processing module 2 detects whether the number of the text form user comments including the link of the voice form user comment corresponding to the time period just recorded is greater than a first threshold, and the number is greater than the first threshold At the time, the processing module 2 randomly selects a part of the text form user comment containing the link of the voice form user comment from the text form user comment corresponding to the link of the voice form user comment for the time period just recorded.
  • the form of the barrage is displayed to the user; when the number is not greater than the first threshold, The text module user comments of the processing module 2 corresponding to all the links containing the voice form user comments of the time period are displayed to the user in the form of a barrage.
  • the voice barrage generation method and the voice barrage generating device are used to first collect user comments, which may be in the form of text or voice, and if the user comments in the form of voice are collected, This generates a textual user comment; if a textual user comment is collected, a voice form user comment is generated accordingly. A link is then established between the corresponding voice mode user comment and the text mode user comment, and finally the text form user comment containing the link of the voice form user comment is displayed to the user in the form of a barrage after the link is established.
  • the voice barrage generation method and the voice barrage generating device provided by the embodiment of the invention can enable the user to listen to the barrage voice by himself, giving the user more choice and a better viewing experience.
  • An embodiment of the present invention further provides an electronic device.
  • the electronic device includes a memory, a processor, and a bus system. Wherein the processor and the memory are connected by a bus system for storing instructions for executing the instructions stored by the memory, and for performing the method of the above embodiments, for example, in one embodiment, the processing
  • the method for generating a voice barrage includes: collecting a voice form user comment or a text form user comment; generating the text form user comment according to the collected voice form user comment or according to the collected text form user a comment generating the voice form user comment; establishing a link of the text form user comment with the voice form user comment; and displaying the text form user comment containing the link of the voice form user comment in the form of a barrage To the user.
  • the memory may be a non-transitory computer readable storage medium for storing computer-executable instructions, which when executed by one or more processors, for example, may cause the processor to perform the steps of the above method embodiments, such as , as described in Figure 1, or to cause the processor to execute
  • the functions of the apparatus modules of the above apparatus embodiments, such as the functions of modules 1 through 2 as described in FIG. 2, may also be stored and/or transmitted in any non-transitory computer readable storage medium.
  • an instruction execution system, apparatus, or device such as a computer-based system, a system including a processor, or an instruction execution system, apparatus Or other system in which the device fetches instructions and executes the instructions.
  • Non-transitory computer readable storage medium can be any medium that tangibly contains or stores computer-executable instructions, which can be used by an instruction execution system, device, or system Or use in conjunction with an instruction execution system, device, or device.
  • Non-transitory computer readable storage media may include, but are not limited to, magnetic, optical, and/or semiconductor storage devices. Examples of such storage devices include magnetic disks, CD-ROM based on CD, DVD or Blu-ray technology, and persistent solid-state memories (such as flash memory, solid state drives, etc.).
  • the processor may be a central processing unit ("CPU").
  • the processor 42 can also be other general purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete Hardware components, etc.
  • DSPs digital signal processors
  • ASICs application specific integrated circuits
  • FPGAs off-the-shelf programmable gate arrays
  • the general purpose processor may be a microprocessor or the processor or any conventional processor or the like.
  • each step of the above method or each unit of the device may be completed by an integrated logic circuit of hardware in the processor or an instruction in a form of software.
  • the units of the steps or apparatus of the method disclosed in the embodiments of the present application may be directly implemented as a hardware processor, or may be performed by a combination of hardware and software modules in the processor.
  • the software module can be located in a conventional storage medium such as random access memory, flash memory, read only memory, programmable read only memory or electrically erasable programmable memory, registers, and the like.
  • the storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Human Computer Interaction (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
  • Machine Translation (AREA)

Abstract

本发明实施例涉及视频优化,公开了一种语音弹幕生成方法和一种语音弹幕生成装置,该生成方法包括:采集语音形式用户评论或者文字形式用户评论;根据采集到的所述语音形式用户评论生成所述文字形式用户评论或者根据采集到的所述文字形式用户评论生成所述语音形式用户评论;建立所述文字形式用户评论与所述语音形式用户评论的链接;以及将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。本发明部分实施例可以使用户自选收听到弹幕语音,给用户更多的选择余地和更佳的观看体验。

Description

一种语音弹幕生成方法和装置
本申请要求于2015年12月15日提交中国专利局、申请号为201510932226.X的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本专利申请涉及视频优化,具体地,涉及一种语音弹幕生成方法和一种语音弹幕生成装置。
背景技术
现今视频网站中观看视频大多都有弹幕输入框,观众在弹幕输入框输入文字可以实时显示在视频中,让同时观看该视频的观众可以进行互动,了解其他观看该视频的观众的想法。但是现在弹幕的形式均以文字形式存在,不能满足观众多元化的需求。
发明内容
本发明部分实施例的目的是提供一种语音弹幕生成方法和一种语音弹幕生成装置,该语音弹幕生成方法和语音弹幕生成装置可以使用户自选收听到弹幕语音,给用户更多的选择余地和更佳的观看体验。
为了实现上述目的,本发明实施例实施例提供一种语音弹幕生成方法,该生成方法包括:采集语音形式用户评论或者文字形式用户评论;根据采集到的所述语音形式用户评论生成所述文字形式用户评论或者根据采集到的所述文字形式用户评论生成所述语音形式用户评论;建立所述文字形式用户评论与所述语音形式用户评论的链接;以及将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
本发明实施例还提供一种语音弹幕生成装置,该生成装置包括:采集模 块,用于采集语音形式用户评论或者文字形式用户评论;处理模块,用于执行以下操作:根据所述采集模块采集到的所述语音形式用户评论生成所述文字形式用户评论或者根据采集到的所述文字形式用户评论生成所述语音形式用户评论;建立所述文字形式用户评论与所述语音形式用户评论的链接;以及将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
本发明的一个实施例提供了一种计算机可读存储介质,包括计算机可执行指令,所述计算机可执行指令在由至少一个处理器执行时致使所述处理器执行上述方法。
本发明的一个实施例提供了一种电子设备,所述电子设备包括:存储器,所述存储器用于存储指令;处理器以及总线系统,所述处理器和所述存储器通过所述总线系统连接;其中,所述处理器用于执行所述存储器存储的指令,并且用于执行上述方法。
通过上述技术方案,采用本发明实施例提供的语音弹幕生成方法和语音弹幕生成装置,首先采集用户评论,可以是文字形式也可以是语音形式,如果采集的是语音形式用户评论,则据此生成文字形式用户评论;如果采集的是文字形式用户评论,则据此生成语音形式用户评论。接着将对应的语音形式用户评论与文字形式用户评论之间建立链接,并最后在建立链接之后将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。通过本发明实施例提供的语音弹幕生成方法和语音弹幕生成装置,可以使用户自选收听到弹幕语音,给用户更多的选择余地和更佳的观看体验。
本发明实施例的其它特征和优点将在随后的具体实施方式部分予以详细说明。
附图说明
附图是用来提供对本发明实施例的进一步理解,并且构成说明书的一部分,与下面的具体实施方式一起用于解释本发明,但并不构成对本发明的限制。在附图中:
图1是本发明实施例提供的语音弹幕生成方法的流程图;
图2是本发明实施例提供的语音弹幕生成装置的结构示意图。
附图标记说明
1  采集模块  2  处理模块
具体实施方式
以下结合附图对本发明实施例的具体实施方式进行详细说明。应当理解的是,此处所描述的具体实施方式仅用于说明和解释本发明,并不用于限制本发明。
图1是本发明实施例提供的语音弹幕生成方法的流程图。如图1所示,一种语音弹幕生成方法,该生成方法包括:采集语音形式用户评论或者文字形式用户评论;根据采集到的所述语音形式用户评论生成所述文字形式用户评论或者根据采集到的所述文字形式用户评论生成所述语音形式用户评论;建立所述文字形式用户评论与所述语音形式用户评论的链接;以及将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
首先采集用户评论,用户评论可以是文字形式也可以是语音形式,判断用户评论是语音形式用户评论或是文字形式用户评论,对于文字形式用户评论仍可以利用现有的文字输入框输入,对于语音形式用户评论,在本发明实施例中可以提供给用户利用任意语音输入设备进行语音形式评论的机会。如 果采集的是语音形式用户评论,则据此生成文字形式用户评论;如果采集的是文字形式用户评论,则据此生成语音形式用户评论。例如可以利用机器语音为用户评论进行配音以产生语音形式用户评论。
接着,将对应的语音形式用户评论与文字形式用户评论之间建立链接,可以使得用户点击文字形式用户评论就可以调出对应的语音形式文字评论。最后在建立链接之后将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。用户看到文字形式弹幕,点击该文字形式弹幕就可以收听语音形式弹幕。
另外,上文所述仅是用户评论实时转化为弹幕并显示给用户,但是有些视频由于过时或者因为观看时间的差异,仅仅利用用户评论实时转化为弹幕会导致弹幕很少甚至没有。因此对于该问题,该生成方法还包括:记录在视频播放中采集到语音形式用户评论或者文字形式用户评论的时间段;以及在该视频再次播放时,在所述时间段将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
在每次视频播放时都会采集用户评论,每采集一个用户评论就记录下这个用户评论的时间段,该时间段可以是一段时间,例如5秒等。在视频再次被播放时,不仅显示实时弹幕,而且还会将之前进行过记录时间段的用户评论在视频到达相应的时间段时以弹幕的形式显示给用户,这样不仅可以使用户评论与视频内容仍然紧密相关,还可以使用户无论在何时观看何种视频都可以具有与多人共同观看视频的感觉。
但是,当同一时间段的含有语音形式用户评论的链接的文字形式用户评论越来越多时,可能会出现还没有将该一时间段的含有语音形式用户评论的链接的文字形式用户评论显示完就已经进入下一个时间段。因此,在这种情况下就不能将所有含有语音形式用户评论的链接的文字形式用户评论以弹幕的形式显示给用户。在本发明实施例中,在对应于所述时间段的含有所述 语音形式用户评论的链接的所述文字形式用户评论的数量大于第一阈值的情况下,从该含有语音形式用户评论的链接的文字形式用户评论中随机挑选含有语音形式用户评论的链接的文字形式用户评论以弹幕的形式显示给用户。直到进入下一个时间段,停止随机显示本时间段的含有语音形式用户评论的链接的文字形式用户评论,从而随机显示下一时间段的含有语音形式用户评论的链接的文字形式用户评论。其中,第一阈值可以自由设定,与时间段的设定长度和弹幕在屏幕上的移动速度有关。
考虑到存储空间限制,由于视频存在时间越来越长,用户评论也越来越多,还可能会出现存储空间不足或者随机选取含有语音形式用户评论的链接的文字形式用户评论的速度过慢的情况,因此在本发明实施例中,在对应于所述时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论的数量大于第二阈值的情况下,每生成一条新的含有语音形式用户评论的链接的文字形式用户评论对应删除一条之前最早生成的含有语音形式用户评论的链接的文字形式用户评论。其中第二阈值可以自由设定,与存储空间等有关,但一些实施例中应该远大于第一阈值。
图2是本发明实施例提供的语音弹幕生成装置的结构示意图。如图2所示,一种语音弹幕生成装置,该生成装置包括:采集模块1,用于采集语音形式用户评论或者文字形式用户评论;处理模块2,用于执行以下操作:根据所述采集模块1采集到的所述语音形式用户评论生成所述文字形式用户评论或者根据采集到的所述文字形式用户评论生成所述语音形式用户评论;建立所述文字形式用户评论与所述语音形式用户评论的链接;以及将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
首先采集模块1采集用户评论,并将用户评论发送给处理模块2,处理模块2接收用户评论。用户评论可以是文字形式也可以是语音形式。如果采 集的是语音形式用户评论,处理模块2则据此生成文字形式用户评论;如果采集的是文字形式用户评论,处理模块2则据此生成语音形式用户评论。例如可以利用机器语音为用户评论进行配音以产生语音形式用户评论。
接着,处理模块2将对应的语音形式用户评论与文字形式用户评论之间建立链接,可以使得用户点击文字形式用户评论就可以调出对应的语音形式文字评论。最后处理装置2在建立链接之后将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
所述处理模块2还用于:记录在视频播放中采集到语音形式用户评论或者文字形式用户评论的时间段;以及在该视频再次播放时,在所述时间段将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
当处理模块2中存储的同一时间段的含有语音形式用户评论的链接的文字形式用户评论越来越多时,可能会出现还没有将该一时间段的含有语音形式用户评论的链接的文字形式用户评论显示完就已经进入下一个时间段。因此,在本发明实施例中,处理模块2还用于:在对应于所述时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论的数量大于第一阈值的情况下,从该含有语音形式用户评论的链接的文字形式用户评论中随机挑选含有语音形式用户评论的链接的文字形式用户评论以弹幕的形式显示给用户。直到进入下一个时间段,停止随机显示本时间段的含有语音形式用户评论的链接的文字形式用户评论,从而随机显示下一时间段的含有语音形式用户评论的链接的文字形式用户评论。其中,第一阈值可以自由设定,与时间段的设定长度和弹幕在屏幕上的移动速度有关。
考虑到处理模块2的存储空间限制,由于视频存在时间越来越长,用户评论也越来越多,还可能会出现处理模块2存储空间不足或者随机选取含有语音形式用户评论的链接的文字形式用户评论的速度过慢的情况,因此在本 发明实施例中,处理模块2还用于:在对应于所述时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论的数量大于第二阈值的情况下,每生成一条新的含有语音形式用户评论的链接的文字形式用户评论对应删除一条之前最早生成的含有语音形式用户评论的链接的文字形式用户评论。其中第二阈值可以自由设定,与存储空间等有关,但在一些实施例中应该远大于第一阈值。
下面将通过具体应用详细说明本发明实施例的技术方案:
用户在观看视频时,可以通过输入框输入文字形式用户评论,也可以通过语音采集装置输入语音形式用户评论。采集模块1此时采集文字形式用户评论或者语音形式用户评论,并将文字形式用户评论或者语音形式用户评论发送给处理模块2。处理模块2记录此时的时间段。处理模块2接收到的是文字形式用户评论时,处理模块2将文字形式用户评论生成为语音形式用户评论,并建立文字形式用户评论和生成的语音形式用户评论的链接;在处理模块2接收到的是语音形式用户评论时,处理模块2将语音形式用户评论生成文字形式用户评论,并建立语音形式用户评论与生成的文字形式用户评论的链接。处理模块2将处理过后的含有所述语音形式用户评论的链接的所述文字形式用户评论显示给用户。与此同时,处理模块2检测对应刚刚记录的时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论的数量,在数量大于第二阈值时,删除一条之前最早生成的含有语音形式用户评论的链接的文字形式用户评论。如果数量不大于第二阈值,处理模块2检测检测对应刚刚记录的时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论的数量是否大于第一阈值,在数量大于第一阈值时,处理模块2从对应刚刚记录的时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论中随机挑选一部分含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户;在数量不大于第一阈值时, 处理模块2对应于该时间段的所有含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。用户点击任意一条弹幕(文字形式用户评论),将播放该弹幕(文字形式用户评论)链接的语音形式用户评论。
通过上述技术方案,采用本发明实施例提供的语音弹幕生成方法和语音弹幕生成装置,首先采集用户评论,可以是文字形式也可以是语音形式,如果采集的是语音形式用户评论,则据此生成文字形式用户评论;如果采集的是文字形式用户评论,则据此生成语音形式用户评论。接着将对应的语音模式用户评论与文字模式用户评论之间建立链接,并最后在建立链接之后将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。通过本发明实施例提供的语音弹幕生成方法和语音弹幕生成装置,可以使用户自选收听到弹幕语音,给用户更多的选择余地和更佳的观看体验。
本发明实施例还提供一种电子设备。该电子设备包括:存储器、处理器以及总线系统。其中,处理器和存储器通过总线系统连接,该存储器用于存储指令,该处理器用于执行该存储器存储的指令,并且用于执行上述实施例中的方法,例如,在一个实施例中,该处理器用于执行一种语音弹幕生成方法包括:采集语音形式用户评论或者文字形式用户评论;根据采集到的所述语音形式用户评论生成所述文字形式用户评论或者根据采集到的所述文字形式用户评论生成所述语音形式用户评论;建立所述文字形式用户评论与所述语音形式用户评论的链接;以及将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
其中,存储器可以是非易失性计算机可读存储介质,以用于存储计算机可执行指令,该指令当由一个或多个处理器执行时,例如可以使得处理器执行以上方法实施例的步骤,比如,如图1描述的步骤,或者,使得处理器执 行以上装置实施例的装置模块的功能,比如,如图2描述的模块1至模块2的功能,计算机可执行指令也可以在任何非易失性计算机可读存储介质内存储和/或传输,以便由指令执行系统、装置或设备使用,或者结合指令执行系统、装置或设备使用,其中该指令执行系统、装置或设备诸如基于计算机的系统、包含处理器的系统或可以从指令执行系统、装置或设备获取指令并执行该指令的其他系统。出于本文档的目的,“非易失性计算机可读存储介质”可以是有形地包含或存储计算机可执行指令的任何介质,该计算机可执行指令可以用于由指令执行系统、设备或系统使用或者结合指令执行系统、装置或设备使用。非易失性计算机可读存储介质可以包括但不限于磁的、光的和/或半导体存储装置。这些存储装置的示例包括磁盘、基于CD、DVD或蓝光技术的光盘以及持久性固态存储器(诸如,闪存、固态驱动器等)。
应当理解,在本申请实施例中,该处理器可以是中央处理单元(Central Processing Unit,简称为“CPU”)。该处理器42还可以是其他通用处理器、数字信号处理器(DSP)、专用集成电路(ASIC)、现成可编程门阵列(FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
在实现过程中,上述方法的各步骤或装置的各单元可以通过处理器中的硬件的集成逻辑电路或者软件形式的指令完成。结合本申请实施例所公开的方法的步骤或装置的各单元可以直接体现为硬件处理器执行完成,或者用处理器中的硬件及软件模块组合执行完成。软件模块可以位于随机存储器,闪存、只读存储器,可编程只读存储器或者电可擦写可编程存储器、寄存器等本领域成熟的存储介质中。该存储介质位于存储器,处理器读取存储器中的信息,结合其硬件完成上述方法的步骤。为避免重复,这里不再详细描述。
以上结合附图详细描述了本发明实施例的部分实施方式,但是,本发明并不限于上述实施方式中的具体细节,在本发明的技术构思范围内,可以对 本发明的技术方案进行多种简单变型,这些简单变型均属于本发明的保护范围。
另外需要说明的是,在上述具体实施方式中所描述的各个具体技术特征,在不矛盾的情况下,可以通过任何合适的方式进行组合,为了避免不必要的重复,本发明对各种可能的组合方式不再另行说明。
此外,本发明的各种不同的实施方式之间也可以进行任意组合,只要其不违背本发明的思想,其同样应当视为本发明所公开的内容。

Claims (10)

  1. 一种语音弹幕生成方法,该生成方法包括:
    采集语音形式用户评论或者文字形式用户评论;
    根据采集到的所述语音形式用户评论生成所述文字形式用户评论或者根据采集到的所述文字形式用户评论生成所述语音形式用户评论;
    建立所述文字形式用户评论与所述语音形式用户评论的链接;以及
    将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
  2. 根据权利要求1所述的语音弹幕生成方法,其中,根据采集到的所述文字形式用户评论生成所述语音形式用户评论包括:
    在采集到的用户评论为文字形式时,利用机器语音为所述用户评论进行配音以产生所述语音形式用户评论。
  3. 根据权利要求1或2所述的语音弹幕生成方法,其中,该生成方法还包括:
    记录在视频播放中采集到所述语音形式用户评论或者所述文字形式用户评论的时间段;以及
    在该视频再次播放时,在所述时间段将对应于该时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户,其中该文字形式用户评论是之前在该时间段采集到的或者是根据之前在该时间段采集到的所述语音形式用户评论生成的。
  4. 根据权利要求3所述的语音弹幕生成方法,其中,该生成方法还包括:
    在对应于所述时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论的数量大于第一阈值的情况下,从该含有语音形式用户评论的链接的文字形式用户评论中随机挑选含有语音形式用户评论的链接的文字形式用户评论以弹幕的形式显示给用户。
  5. 根据权利要求3所述的语音弹幕生成方法,其中,该生成方法还包括:
    在对应于所述时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论的数量大于第二阈值的情况下,每生成一条新的含有语音形式用户评论的链接的文字形式用户评论对应删除一条之前最早生成的含有语音形式用户评论的链接的文字形式用户评论。
  6. 一种语音弹幕生成装置,该生成装置包括:
    采集模块(1),用于采集语音形式用户评论或者文字形式用户评论;
    处理模块(2),用于执行以下操作:
    根据所述采集模块(2)采集到的所述语音形式用户评论生成所述文字形式用户评论或者根据采集到的所述文字形式用户评论生成所述语音形式用户评论;
    建立所述文字形式用户评论与所述语音形式用户评论的链接;以及
    将含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户。
  7. 根据权利要求6所述的语音弹幕生成装置,其中,根据采集到的所述文字形式用户评论生成所述语音形式用户评论包括:
    在采集到的用户评论为文字形式时,利用机器语音为所述用户评论进行配音以产生所述语音形式用户评论。
  8. 根据权利要求6或7所述的语音弹幕生成装置,其中,所述处理模块还用于:
    记录在视频播放中采集到语音形式用户评论或者文字形式用户评论的时间段;以及
    在该视频再次播放时,在所述时间段将对应于该时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论以弹幕的形式显示给用户,其 中该文字形式用户评论是之前在该时间段采集到的或者是根据之前在该时间段采集到的所述语音形式用户评论生成的。
  9. 根据权利要求8所述的语音弹幕生成方法,其中,所述处理模块还用于:
    在对应于所述时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论的数量大于第一阈值的情况下,从该含有语音形式用户评论的链接的文字形式用户评论中随机挑选含有语音形式用户评论的链接的文字形式用户评论以弹幕的形式显示给用户。
  10. 根据权利要求8所述的语音弹幕生成方法,其中,所述处理模块还用于:
    在对应于所述时间段的含有所述语音形式用户评论的链接的所述文字形式用户评论的数量大于第二阈值的情况下,每生成一条新的含有语音形式用户评论的链接的文字形式用户评论对应删除一条之前最早生成的含有语音形式用户评论的链接的文字形式用户评论。
PCT/CN2016/089574 2015-12-15 2016-07-10 一种语音弹幕生成方法和装置 Ceased WO2017101430A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US15/240,152 US20170168660A1 (en) 2015-12-15 2016-08-18 Voice bullet screen generation method and electronic device

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201510932226.X 2015-12-15
CN201510932226.XA CN105898603A (zh) 2015-12-15 2015-12-15 一种语音弹幕生成方法和装置

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US15/240,152 Continuation US20170168660A1 (en) 2015-12-15 2016-08-18 Voice bullet screen generation method and electronic device

Publications (1)

Publication Number Publication Date
WO2017101430A1 true WO2017101430A1 (zh) 2017-06-22

Family

ID=57002460

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/089574 Ceased WO2017101430A1 (zh) 2015-12-15 2016-07-10 一种语音弹幕生成方法和装置

Country Status (2)

Country Link
CN (1) CN105898603A (zh)
WO (1) WO2017101430A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109275014A (zh) * 2018-09-13 2019-01-25 武汉斗鱼网络科技有限公司 一种链接弹幕的方法及移动终端
CN116112757A (zh) * 2021-11-11 2023-05-12 上海哔哩哔哩科技有限公司 弹幕处理方法及装置

Families Citing this family (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106341722A (zh) * 2016-09-21 2017-01-18 努比亚技术有限公司 一种视频编辑方法及装置
CN107707985B (zh) * 2017-09-15 2019-12-17 维沃移动通信有限公司 一种弹幕控制方法、移动终端及服务器
CN108174276B (zh) * 2018-01-04 2020-10-20 北京奇艺世纪科技有限公司 一种弹幕显示方法及显示装置
CN111672099B (zh) * 2020-05-28 2023-03-24 腾讯科技(深圳)有限公司 虚拟场景中的信息展示方法、装置、设备及存储介质
CN113535116B (zh) * 2021-08-05 2024-08-27 广州酷狗计算机科技有限公司 音频文件的播放方法、装置、终端及存储介质

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20120253492A1 (en) * 2011-04-04 2012-10-04 Andrews Christopher C Audio commenting system
CN104125483A (zh) * 2014-07-07 2014-10-29 乐视网信息技术(北京)股份有限公司 音频评论信息生成方法和装置,音频评论播放方法和装置
CN104714937A (zh) * 2015-03-30 2015-06-17 北京奇艺世纪科技有限公司 一种评论信息发布方法及装置
US20150199320A1 (en) * 2010-12-29 2015-07-16 Google Inc. Creating, displaying and interacting with comments on computing devices
CN104822093A (zh) * 2015-04-13 2015-08-05 腾讯科技(北京)有限公司 弹幕发布方法和装置
CN104994401A (zh) * 2015-07-03 2015-10-21 王春晖 弹幕处理方法、装置及系统

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104142778B (zh) * 2013-09-25 2017-06-13 腾讯科技(深圳)有限公司 一种文本处理的方法、装置及移动终端
CN104602136A (zh) * 2015-02-28 2015-05-06 科大讯飞股份有限公司 用于外语学习的字幕显示方法及系统

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20150199320A1 (en) * 2010-12-29 2015-07-16 Google Inc. Creating, displaying and interacting with comments on computing devices
US20120253492A1 (en) * 2011-04-04 2012-10-04 Andrews Christopher C Audio commenting system
CN104125483A (zh) * 2014-07-07 2014-10-29 乐视网信息技术(北京)股份有限公司 音频评论信息生成方法和装置,音频评论播放方法和装置
CN104714937A (zh) * 2015-03-30 2015-06-17 北京奇艺世纪科技有限公司 一种评论信息发布方法及装置
CN104822093A (zh) * 2015-04-13 2015-08-05 腾讯科技(北京)有限公司 弹幕发布方法和装置
CN104994401A (zh) * 2015-07-03 2015-10-21 王春晖 弹幕处理方法、装置及系统

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109275014A (zh) * 2018-09-13 2019-01-25 武汉斗鱼网络科技有限公司 一种链接弹幕的方法及移动终端
CN116112757A (zh) * 2021-11-11 2023-05-12 上海哔哩哔哩科技有限公司 弹幕处理方法及装置

Also Published As

Publication number Publication date
CN105898603A (zh) 2016-08-24

Similar Documents

Publication Publication Date Title
WO2017101430A1 (zh) 一种语音弹幕生成方法和装置
EP3311332B1 (en) Automatic recognition of entities in media-captured events
US20200004874A1 (en) Conversational agent dialog flow user interface
WO2017161776A1 (zh) 弹幕推送方法及装置
US10255502B2 (en) Method and a system for generating a contextual summary of multimedia content
US20150156236A1 (en) Synchronize Tape Delay and Social Networking Experience
WO2024146338A1 (zh) 视频生成方法、装置、电子设备及存储介质
US20170168660A1 (en) Voice bullet screen generation method and electronic device
CN103026704A (zh) 信息处理装置、信息处理方法、程序、存储介质以及集成电路
US10425378B2 (en) Comment synchronization in a video stream
CN112114886B (zh) 误唤醒音频的获取方法和装置
WO2023160288A1 (zh) 会议纪要生成方法、装置、电子设备和可读存储介质
CN104902145B (zh) 一种直播流视频的播放方法及装置
CN108882004A (zh) 视频录制方法、装置、设备及存储介质
US20170194032A1 (en) Process for automated video production
Rosa Video enriched retrieval augmented generation using aligned video captions
CN108093311B (zh) 多媒体文件的处理方法、装置、存储介质及电子设备
WO2025092911A1 (zh) 视频特征提取方法、视频生成方法、装置、介质及设备
CN104123112B (zh) 一种图像处理方法及电子设备
CN117556066A (zh) 多媒体内容生成方法和电子设备
CN106954085A (zh) 网络直播方法及装置
CN104837074B (zh) 一种显示时间的设置方法及装置
CN115988240A (zh) 视频生成方法、装置、可读介质及电子设备
CN118573977B (zh) 基于用户点评生成酒店短视频的方法、装置、设备及介质
CN115866351B (zh) 视频弹幕朗读方法、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16874491

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16874491

Country of ref document: EP

Kind code of ref document: A1