WO2019029073A1 - 传屏方法、装置、电子设备及计算机可读存储介质 - Google Patents
传屏方法、装置、电子设备及计算机可读存储介质 Download PDFInfo
- Publication number
- WO2019029073A1 WO2019029073A1 PCT/CN2017/116067 CN2017116067W WO2019029073A1 WO 2019029073 A1 WO2019029073 A1 WO 2019029073A1 CN 2017116067 W CN2017116067 W CN 2017116067W WO 2019029073 A1 WO2019029073 A1 WO 2019029073A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sound information
- information
- speaker
- source device
- screen
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/26—Speech to text systems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/14—Digital output to display device ; Cooperation and interconnection of the display device with other functional units
- G06F3/1454—Digital output to display device ; Cooperation and interconnection of the display device with other functional units involving copying of the display data of a local workstation or window to a remote workstation or window so that an actual copy of the data is displayed simultaneously on two or more displays, e.g. teledisplay
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L2021/02087—Noise filtering the noise being separate speech, e.g. cocktail party
Definitions
- the present invention relates to the field of screen transmission technologies, and in particular, to a screen transmission method, apparatus, electronic device, and computer readable storage medium.
- the screen transmission technology mainly refers to the technology of synchronizing the content displayed on the screen of the mobile phone, the computer, and the like (the desktop data) to the display device such as a projector, a television, a conference tablet, and the like.
- the display device such as a projector, a television, a conference tablet, and the like.
- Mobile phones, computers and other devices have the advantages of convenient operation and strong processing capability.
- the display devices such as conference tablets have the advantages of large screen and good sound effects. Through the screen transmission technology, the advantages of the two can be combined and are used in meetings and other scenarios. Use a lot.
- the present invention provides a screen transmission method, apparatus, electronic device, and computer readable storage medium to overcome the problem of poor speech recognition performance in current conference scenarios.
- a method of screen transmission comprising the following steps:
- the text information is rendered in a screen shot.
- the step of receiving, by the source device, the sound information of the surrounding environment, and combining the sound information of the surrounding environment collected by the source device, and converting the voice information into the text information corresponding to the speaker by feature recognition include:
- the step of receiving, by the source device, the sound information of the surrounding environment, and combining the sound information of the surrounding environment collected by the source device, and converting the voice information into the text information corresponding to the speaker by the feature recognition includes: :
- the sound information corresponding to the speaker is converted into text information.
- the step of analyzing the sound information and the sound information of the surrounding environment collected by the self, and extracting the sound information corresponding to the speaker includes:
- the sound information received from the source device is used as reference information, and the sound information collected by itself is correlated with the reference information to remove environmental noise and/or other speaker's voice information, and the sound information corresponding to the single speaker is extracted.
- the screen shot is a screen obtained by displaying the desktop data sent by the source device, and the step of extracting the sound information corresponding to the single speaker includes:
- voice processing on the voice information corresponding to the speaker, where the voice processing includes gain processing and attenuation processing, so that the volume corresponding to the sound information is within a preset range;
- the processed sound information is associated with the desktop data according to the timestamp.
- the text information includes at least one of the following:
- the step of rendering the text information in the screen shot includes:
- the rendering attribute includes at least one of the following: a font color, a font size, a font thickness, a display orientation, and a personalized mark; the personalized mark includes any one of the following: an underline, and a text highlight color.
- the step of rendering the text information in the screen shot includes:
- the sound information of the surrounding environment sent by the source device that sends the desktop data is matched, and the single speaker corresponding to the voice information is used as the presenter, and the text information of the presenter is highlighted in a form different from other speakers.
- the method further includes:
- the invention also discloses a screen transmission device, comprising:
- the processing module is configured to receive sound information of the surrounding environment collected and transmitted by the source device, and combine the sound information of the surrounding environment collected by the source device to convert the sound information into text information corresponding to the speaker by feature recognition;
- a rendering module configured to render the text information in a screen shot.
- the invention also discloses an electronic device, comprising:
- a memory for storing processor executable instructions
- the processor is configured to perform the screen transmission method according to any of the preceding claims.
- the invention also discloses a computer readable storage medium having stored thereon a computer program, the program being executed by the processor to implement the screen transmission method according to any of the preceding claims.
- the invention receives the sound information of the surrounding environment collected and transmitted by the source device, combines the sound information of the surrounding environment collected by itself, and converts the sound information corresponding to the same speaker into the same text information through the feature recognition, in the screen Rendering the text information in . Since the speaker is usually adjacent to the source device that he/she holds, the speaker of the source device collects the sound information, and the speaker's voice is clearer, which improves the accuracy of distinguishing each speaker from the voice information, and facilitates the same speech. The corresponding voice information of the person converts the text information, which improves the accuracy of the voice recognition.
- FIG. 1 is a flowchart of a screen transmission method according to an exemplary embodiment of the present invention
- FIG. 2a is a diagram showing an example of a conference scenario according to an exemplary embodiment of the present invention.
- FIG. 2b is a detailed diagram of a method for processing sound information according to an exemplary embodiment of the present invention.
- 2c is a detailed illustration of a method of processing sound information, according to an exemplary embodiment of the present invention.
- FIG. 3 is a flowchart of a screen transmission method according to an exemplary embodiment of the present invention.
- FIG. 4 is an effect diagram of rendering text information according to an exemplary embodiment of the present invention.
- FIG. 5 is a logic block diagram of an electronic device according to an exemplary embodiment of the present invention.
- FIG. 6 is a logic block diagram of a screen device according to an exemplary embodiment of the invention.
- first, second, third, etc. may be used to describe various information in the present invention, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other.
- first information may also be referred to as the second information without departing from the scope of the invention.
- second information may also be referred to as the first information.
- word "if” as used herein may be interpreted as "when” or "when” or "in response to determination.”
- Conference tablet and other devices have been widely used in conferences in recent years due to their advantages of large screen, good sound effects, and support for handwriting input.
- the presenter uses the screen transmission technology to synchronize the content displayed on the screen of the mobile phone, computer, and the like that he uses, and the sound played (desktop data) to a display device such as a conference tablet for display.
- participants may use different languages, accents, and speaking speeds, which may result in other participants not fully understanding the meeting information.
- the current speech recognition can convert speech into subtitles, there are many people in the discussion process, and the generated subtitles are also chaotic. Therefore, the application of speech recognition in the conference scene is not effective, which makes the information published or discussed at the conference distorted or omitted, which reduces the efficiency of conference communication.
- the present invention proposes a screen transmission method, as shown in FIG. 1, the method includes:
- S110 Receive sound information of a surrounding environment collected and sent by the source device, and combine the voice information of the surrounding environment collected by the source device to convert the voice information into text information corresponding to the speaker by using feature recognition;
- a device such as a conference tablet (hereinafter referred to as a screen device) needs to be placed at a position convenient for all participants, so the position of the screen device is kept at a certain distance from the participants.
- a small conference diagram is shown.
- Four participants 230 sit down along the round table 240, and the screen device 210 is placed opposite the participant 230 (for example, a wall).
- the conference speaker is configured.
- the source device 220 for example, a computer, a microphone, etc.
- other participants 230 may also configure the source device 220.
- each speaker Since all the participants 230 may speak (referring to the person currently speaking as a speaker), each speaker is far away from the screen device 210, due to the attenuation of the sound wave transmission, the noise of the surrounding environment, and other interferences.
- the quality of the sound collected by the screen device 210 is generally lower than the quality of the sound collected by the source device 220.
- the voice of the speaker collected by the source device 220 is clearer.
- the unit 211 in the screen device 210 in FIG. 2a indicates that the microphone or the like can collect the voice of the speaker. Of course, it can also be an external device for collecting sound (for example, an omnidirectional microphone, etc.) connected to the screen device 210. The invention does not limit this.
- the source device 220 can also collect the voice of the speaker, and send the collected voice information to the screen device 210.
- the unit 212 in the screen device 210 represents the communication device. Of course, it may also be a Bluetooth or wireless network.
- the conference presenter also sends the desktop data to the screen device 210 through the source device 220, and the screen device 210 displays the content of the desktop data of the source device 220 (screening screen).
- the screen device 210 performs comprehensive analysis and processing on the sound information collected by the source device 220 and the sound information collected by the source device 220, thereby accurately identifying the voice information of each speaker, and converting the sound information into text information (may be for each A speaker sets a text message, or records the text information of each speaker in a data, and then renders the text information in a screen, and the rendering effect can be similar to the caption.
- the screen device 210 and the source device 220 simultaneously collect sound information of the surrounding environment, and the screen device 210 analyzes and processes the sound information 0# collected by itself, and recognizes the sound information 0# according to the feature recognition.
- the sound information corresponding to each speaker is converted into the first text information (a first text information may be set for each speaker, or the first text information of each speaker is recorded in a data), and the source device 220
- the sound information 1# collected by itself is transmitted to the screen device 210, and the screen device 210 corrects the first text information according to the sound information 1#, thereby obtaining text information with high accuracy.
- the correction mode may be: converting the sound information 1# into the text information 1#, comparing the first text information with the text information 1# for correction; or the first text information may be reviewed by the sound information 1#;
- the specific manner of correcting the text by sound in the present application is not limited thereto, and other correction methods may be employed.
- the sound information can also be processed in the following manner:
- the sound information corresponding to the speaker is converted into text information.
- the screen device 210 and the source device 220 simultaneously collect the sound information of the surrounding environment, and the source device 220 sends the sound information 1# collected by the source device 220 to the screen device 210, and the screen device 210 collects the information.
- the sound information 0# and the sound information 1# perform an analysis process to extract sound information corresponding to each speaker (the extracted sound information may be for a single speaker or a plurality of speakers), and The voice information corresponding to the speaker is converted into text information. Since the extracted sound information filters out noise or even other speaker's voice information (denoted as pure voice information), the pure voice can be obtained by receiving from the source device 210.
- the sound information is used as reference information to correlate the sound information collected by itself with the reference information.
- the designer can select according to the actual use; thereby removing environmental noise and/or other speakers.
- the sound information so that the sound information (pure voice information) corresponding to a single speaker can be extracted.
- the source device 210 can also perform the process of attenuating and filtering the collected voice information, and only retains the voice information of the speaker using the source device 210, and then sends the voice information to the screen device 210, thereby improving the reference.
- the accuracy of the information can further improve the purity of the sound information corresponding to a single speaker.
- voice processing on the sound information corresponding to the speaker, where the voice processing includes gain processing and/or attenuation processing, so that the volume corresponding to the sound information is within a preset range;
- the processed sound information is associated with the desktop data according to the timestamp.
- the voice of the pure voice information is processed into a preset range, and the peaks and valleys are removed, which can improve the hearing effect.
- the volume information can be converted into text information after the volume adjustment, and there is no interference from the peaks and valleys, and the conversion can be improved.
- the accuracy of the text message is processed into a preset range, and the peaks and valleys are removed, which can improve the hearing effect.
- the text information may include at least one of the following:
- Text information corresponding to the language of the sound information for example, converting Chinese into Chinese, English into English, and Chinese-English mixed into Chinese and English;
- Text information corresponding to the target language for example, if the target language is Chinese, then convert Chinese into Chinese, translate English into Chinese, convert Chinese and English into Chinese, etc.;
- Text information corresponding to the subject language in the voice information for example, if the subject language of the Chinese-English mixed voice information is Chinese, the Chinese-English mixed use is converted into Chinese; if the Chinese-English mixed voice information is English, the subject will be English. The Chinese-English mixed use is converted into English, etc.;
- Text information corresponding to the second language in the voice information for example, if the secondary language of the Chinese-English mixed voice information is Chinese, the Chinese-English mixed use is converted into Chinese; if the Chinese-English mixed voice information is English, the next language is English. The Chinese-English mixed use is converted into English and the like.
- the present invention can generate corresponding text information for each speaker.
- the text information (subtitles) is rendered in the screen shot, different forms of subtitles may be used to distinguish the text information of different speakers:
- the rendering attribute includes at least one of the following: a font color, a font size, a font thickness, a display orientation, and a personalized mark; the personalized mark includes any one of the following: an underline, and a text highlight color.
- the subtitle of one speaker has a background color
- the subtitles of another speaker have no background color.
- the position display may also be such that the position where the subtitles appear is not fixed, or the subtitles of the ordinary spokesperson of the barrage are displayed at both ends, and the subtitles of the presenter are displayed in the middle.
- the speech of the presenter of the meeting is the key point. Therefore, the text information corresponding to the presenter can be highlighted in a form different from other speakers.
- the source device that sends the desktop data is the source device used by the presenter, and can distinguish the voice information from the presenter according to the MAC (Media Access Control) address of the source device.
- the single speaker corresponding to the voice information serves as the presenter, and the text information of the presenter is highlighted in a form different from other speakers.
- subtitles with background color can be considered as subtitles of the presenter, and subtitles without background are ordinary speakers.
- the rendering attribute can also be modified.
- the rendering attribute corresponding to the MAC address is set, and each speaker and corresponding text information are distinguished according to the MAC address of the sent voice information, and then the corresponding rendering attribute is loaded for the text information and rendered to the screen. Subtitles are formed.
- FIG. 4 shows a case where a single presenter screens (the desktop data of only one source device is displayed in the screen device), at present, multiple source device desktop data are received and displayed in one screen device. Obviously, the desktop data of one or more source devices is displayed in the screen device without changing the usage conditions of the screen solution of the present invention. Therefore, the solution of the present invention is also applicable to displaying multiple devices in the screen device. The situation of the source device desktop data.
- the present invention proposes to associate the text information with the desktop data according to the timestamp, and in the subsequent playback of the recorded desktop data, the text information is sequentially displayed at the corresponding time.
- the text information may also appear simultaneously with the aforementioned sound information.
- the method can also be applied to a remote conference, and the local desktop data, voice information, and/or text information is sent to the foreign device, which increases the manner in which the participants attending the conference understand the conference content, and improves the conference communication effect.
- the present invention also provides an embodiment of the screen device.
- Embodiments of the screen device of the present invention can be applied to a conference tablet.
- the device embodiment may be implemented by software, or may be implemented by hardware or a combination of hardware and software.
- the processor of the conference tablet is configured to read the corresponding computer program instructions in the non-volatile memory into the memory.
- FIG. 5 a hardware structure diagram of a conference tablet in which the screen device of the present invention is located, except for the processor, the memory, the network interface, and the non-volatile memory shown in FIG.
- the conference tablet in which the device is located in the embodiment may be further included in other hardware according to the actual function of the screen, and details are not described herein again.
- the screen device 600 includes:
- the processing module 610 is configured to receive sound information of a surrounding environment collected and sent by the source device, and combine the sound information of the surrounding environment collected by the source device to convert the voice information into text information corresponding to the speaker by using feature recognition;
- the rendering module 620 is configured to render the text information in a screen shot.
- an electronic device including:
- a memory for storing processor executable instructions
- the processor is configured to perform the screen transmission method according to any of the preceding claims.
- the present invention further provides a computer readable storage medium having stored thereon a computer program, the program being executed by a processor to implement the screen method according to any of the preceding claims.
- the conference tablet of the invention has the function of screen transmission, and adds functions such as audio to text on the basis of the original screen transmission function, and the function can be realized by using the existing transliteration software, and the transliteration result of the transliteration software is called by the screen transmission function.
- the function of the transliteration software can also be combined in the screen function; of course, other plug-ins capable of realizing the function can be designed according to the actual situation, which is not limited by the present invention.
- the device embodiment since it basically corresponds to the method embodiment, reference may be made to the partial description of the method embodiment.
- the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, ie may be located A place, or it can be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the objectives of the solution of the present invention. Those of ordinary skill in the art can understand and implement without any creative effort.
Landscapes
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Multimedia (AREA)
- Quality & Reliability (AREA)
- Signal Processing (AREA)
- Telephonic Communication Services (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
- Circuits Of Receivers In General (AREA)
Abstract
一种传屏方法、装置(600)、电子设备及计算机可读存储介质,通过接收源端设备(220)采集并发送过来的周围环境的声音信息(1#),结合自身采集的周围环境的声音信息(0#),通过特征识别将与同一发言人对应的声音信息转换成同一文本信息(S110),在投屏画面中渲染文本信息(S120)。由于发言人通常与自己持有的源端设备(220)邻近,因而源端设备(220)采集到声音信息(1#)中该发言人的声音更清晰,提高了从声音信息中区分各发言人的准确性,便于依据同一发言人对应的声音信息转换文本信息,提高了语音识别的准确性。
Description
本发明涉及传屏技术领域,尤其涉及一种传屏方法、装置、电子设备及计算机可读存储介质。
传屏技术主要是指将手机、电脑等设备的屏幕上显示的内容和播放的声音(桌面数据)同步到投影仪、电视机、会议平板等显示设备进行展示的技术。手机、电脑等设备具有操作方便、处理能力强等优势,而会议平板等显示设备具有屏幕大、音效好等优势,通过传屏技术就可以将两者具备的优势结合,在会议等场景下被大量使用。
以会议场景为例,参会人员可能使用不同的语言、口音、语速,导致其他参会人员可能无法完全理解会议信息。目前的语音识别虽然可以将语音转换成字幕,但是,在会议讨论过程中人多口杂,生成的字幕也是混乱的,因此,语音识别在会议场景中的应用效果不佳,使得会议上发布、讨论的信息失真或遗漏,降低了会议沟通的效率。
发明内容
有鉴于此,本发明提供一种传屏方法、装置、电子设备及计算机可读存储介质,以克服目前会议场景中应用语音识别效果不佳的问题。
具体地,本发明是通过如下技术方案实现的:
一种传屏方法,包括以下步骤:
接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息;
在投屏画面中渲染所述文本信息。
一个实施例中,所述接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息的步骤
包括:
对自身采集的周围环境的声音信息进行分析处理,根据特征识别将声音信息转换成与发言人对应的第一文本信息;
接收源端设备采集并发送过来的周围环境的声音信息,根据该声音信息对第一文本信息进行校正。
一个实施例中,所述接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息的步骤包括:
接收源端设备采集并发送过来的周围环境的声音信息,将该声音信息与自身采集的周围环境的声音信息进行分析处理,提取与发言人对应的声音信息;
将与发言人对应的声音信息转换成文本信息。
一个实施例中,所述将该声音信息与自身采集的周围环境的声音信息进行分析处理,提取与发言人对应的声音信息的步骤包括:
以从源端设备接收的声音信息作为参考信息,将自身采集的声音信息与参考信息进行相关性运算,去除环境噪声和/或其他发言人的声音信息,提取与单一发言人对应的声音信息。
一个实施例中,所述投屏画面为对源端设备发送来的桌面数据进行展示所得的画面,所述提取与单一发言人对应的声音信息的步骤之后,还包括:
将发言人对应的声音信息进行语音处理,所述语音处理包括增益处理、衰减处理,使声音信息对应的音量处于预设范围内;
根据时间戳将处理后的声音信息与桌面数据相关联。
一个实施例中,所述文本信息包括以下至少之一:
与声音信息语种对应的文本信息;
与目标语种对应的文本信息;
与声音信息中主语种对应的文本信息;
与声音信息中次语种对应的文本信息。
一个实施例中,在投屏画面中渲染所述文本信息的步骤包括:
匹配与不同发言人对应的渲染属性,依据所述渲染属性在投屏画面中渲染所述文本信息;
其中,所述渲染属性包括以下至少之一:字体颜色、字体大小、字体粗细、显示方位、个性化标记;所述个性化标记包括以下任一:下划线、文字突出显示颜色。
一个实施例中,在投屏画面中渲染所述文本信息的步骤包括:
匹配发送桌面数据的源端设备所发送的周围环境的声音信息,将该声音信息对应的单一发言人作为主讲人,以区别于其他发言人的形式着重显示所述主讲人的文本信息。
一个实施例中,所述通过特征识别将与同一发言人对应的声音信息转换成同一文本信息的步骤之后,还包括:
根据时间戳将文本信息与桌面数据相关联。
本发明还公开了一种传屏装置,包括:
处理模块,用于接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息;
渲染模块,用于在投屏画面中渲染所述文本信息。
本发明还公开了一种电子设备,包括:
处理器;
用于存储处理器可执行指令的存储器;
其中,所述处理器被配置为执行如前任意一项所述的传屏方法。
本发明还公开了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现如前任意一项所述的传屏方法。
本发明通过接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将与同一发言人对应的声音信息转换成同一文本信息,在投屏画面中渲染所述文本信息。由于发言人通常与自己持有的源端设备邻近,因而该源端设备采集到声音信息中该发言人的声音更清晰,提高了从声音信息中区分各发言人的准确性,便于依据同一发言人对应的声音信息转换文本信息,提高了语音识别的准确性。
图1是本发明一示例性实施例示出的一种传屏方法的流程图;
图2a是本发明一示例性实施例示出的会议场景的示例图;
图2b是本发明一示例性实施例示出的对声音信息的处理方法的细化示例图;
图2c是本发明一示例性实施例示出的对声音信息的处理方法的细化示例图;
图3是本发明一示例性实施例示出的一种传屏方法的流程图;
图4是本发明一示例性实施例示出的一种渲染文本信息的效果图;
图5是本发明一示例性实施例示出的一种电子设备的逻辑框图;
图6是本发明一示例性实施例示出的一种传屏装置的逻辑框图。
这里将详细地对示例性实施例进行说明,其示例表示在附图中。下面的描述涉及附图时,除非另有表示,不同附图中的相同数字表示相同或相似的要素。以下示例性实施例中所描述的实施方式并不代表与本发明相一致的所有实施方式。相反,它们仅是与如所附权利要求书中所详述的、本发明的一些方面相一致的装置和方法的例子。
在本发明使用的术语是仅仅出于描述特定实施例的目的,而非旨在限制本发明。在本发明和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。还应当理解,本文中使用的术语“和/或”是指并包含一个或多个相关联的列出项目的任何或所有可能组合。
应当理解,尽管在本发明可能采用术语第一、第二、第三等来描述各种信息,但这些信息不应限于这些术语。这些术语仅用来将同一类型的信息彼此区分开。例如,在不脱离本发明范围的情况下,第一信息也可以被称为第二信息,类似地,第二信息也可以被称为第一信息。取决于语境,如在此所使用的词语“如果”可以被解释成为“在……时”或“当……时”或“响应于确定”。
会议平板等设备因具有大屏幕、音效好、支持手写输入等优势,近年来,在会议场合被广泛使用。通常来说,主讲人通过传屏技术将自己使用的手机、电脑等设备的屏幕上显示的内容和播放的声音(桌面数据)同步到会议平板等显示设备进行展示。然而,参会人员可能使用不同的语言、口音、语速,导致其他参会人员可能无法完全理解会议信息。目前的语音识别虽然可以将语音转换成字幕,但是,在会议讨论过程中人多口杂,生成的字幕也是混乱
的,因此,语音识别在会议场景中的应用效果不佳,使得会议上发布、讨论的信息失真或遗漏,降低了会议沟通的效率。
对此,本发明提出了一种传屏方法,如图1所示,该方法包括:
S110、接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息;
S120、在投屏画面中渲染所述文本信息。
通常来说,会议平板等设备(以下简称传屏设备)需要放置在便于全体与会人员观看的位置,因此传屏设备的位置与参会人员保持一定距离。如图2a所示为一场小型会议示意图,4位参会人员230沿圆桌240依次落坐,传屏设备210放置在参会人员230的对面(例如,墙上),会议的主讲人配置有源端设备220(例如电脑、麦克风等),其他参会人员230也可以配置源端设备220。由于所有参会人员230都可能发言(将当前在说话的人称为发言人),各发言人距离传屏设备210均较远,由于声波传输存在衰减、周围环境的噪声及其它干扰的存在,传屏设备210采集声音的质量通常来说要低于源端设备220采集声音的质量,且发言人有配套的源端设备220时,该源端设备220采集的该发言人的声音会更清晰。
图2a中传屏设备210中的单元211表示麦克风等能采集发言人的声音,当然,也可以是与传屏设备210连接的外置的采集声音的设备(例如,全向麦克风等),本发明对此不作限制。源端设备220也可以采集发言人的声音,并将采集的声音信息发送到传屏设备210,传屏设备210中的单元212表示通信装置,当然,也可能是蓝牙或无线网络等方式发送该声音信息,同时,会议主讲人还通过源端设备220将桌面数据发送给传屏设备210,传屏设备210展示源端设备220桌面数据的内容(投屏画面)。
传屏设备210将自身采集的声音信息与源端设备220采集的声音信息进行综合分析处理,从而能够准确的识别出各发言人的声音信息,并可以将声音信息转化成文本信息(可以针对每一发言人设置一文本信息,或者在一数据中记录各发言人的文本信息),再将该文本信息渲染在投屏画面中,渲染效果可以类似于字幕。
将自身采集的声音信息与源端设备220采集的声音信息进行综合分析处理的方式可以有多种,例如:
对自身采集的周围环境的声音信息进行分析处理,根据特征识别将声音信息转换成与发言人对应的第一文本信息;
接收源端设备采集并发送过来的周围环境的声音信息,根据该声音信息对第一文本信息
进行校正。
如图2b所示,传屏设备210及源端设备220同时采集周围环境的声音信息,传屏设备210将自身采集的声音信息0#进行分析处理,根据特征识别从声音信息0#中识别出与各发言人对应的声音信息并转化成第一文本信息(可以针对每一发言人设置一第一文本信息,或者在一数据中记录各发言人的第一文本信息),源端设备220将自身采集的声音信息1#发送给传屏设备210,传屏设备210根据声音信息1#对第一文本信息进行校正,从而得到准确度高的文本信息。其中,校正方式可以是将声音信息1#转换成文本信息1#,将第一文本信息与文本信息1#进行比较以进行校正;也可以是通过声音信息1#对第一文本信息进行复核;本申请中通过声音校正文本的具体方式不局限于此,还可以采用其它的校正方式。
当然,还可以采用如下方式对声音信息进行处理:
接收源端设备采集并发送过来的周围环境的声音信息,将该声音信息与自身采集的周围环境的声音信息进行分析处理,提取与发言人对应的声音信息;
将与发言人对应的声音信息转换成文本信息。
如图2c所示,传屏设备210及源端设备220同时采集周围环境的声音信息,源端设备220将自身采集的声音信息1#发送给传屏设备210,传屏设备210将自身采集的声音信息0#与声音信息1#进行分析处理,从而提取出与各发言人对应的声音信息(提取的声音信息可以是针对单一发言人的,也可以是包含多个发言人的),将与发言人对应的声音信息转换成文本信息,由于提取的声音信息滤除了噪声甚至其他发言人的声音信息(记为纯净语音信息),纯净语音可以通过以下方式得到:以从源端设备210接收的声音信息作为参考信息,将自身采集的声音信息与参考信息进行相关性运算,相关性运算的方法有多种,设计者可根据实际使用情况选用;从而可以去除环境噪声和/或其他发言人的声音信息,从而就可以提取出与单一发言人对应的声音信息(纯净语音信息)。当然,源端设备210也可以将采集的声音信息进行衰减和滤波等处理,仅保留使用该源端设备210的发言人的声音信息,再将该声音信息发送至传屏设备210,通过提高参考信息的精度,能够进一步提高与单一发言人对应的声音信息的纯度。
根据纯净语音信息转化文本信息的准确性更高,且纯净语音信息还可以做进一步地优化处理。例如,部分发言人的嗓音太小、太大或者嗓音大小波动较大,对这类声音信息进行语音处理将提高听觉效果,特别是录屏后回看(听)和/或将录屏数据发送给异地远程参加会议的与会人员时,处理过程如图3所示:
将发言人对应的声音信息进行语音处理,所述语音处理包括增益处理和/或衰减处理,使声音信息对应的音量处于预设范围内;
根据时间戳将处理后的声音信息与桌面数据相关联。
将纯净语音信息的语音处理到预设范围内,去除尖峰低谷,能够提高听觉效果,当然,也可以进行音量调整后再将声音信息转换成文本信息,没有尖峰低谷的干扰,还可以提高转换成文本信息的准确度。
随着国际化水平越来越高,一场会议中可能会使用多种语言,例如汉语、英语、日语等,从而文本信息可以包括以下至少之一:
与声音信息语种对应的文本信息;例如,将汉语转化成中文、英语转化英文、中英混用的转化成中英文等;
与目标语种对应的文本信息;例如,目标语种是中文,则将汉语转化成中文、英语转化中文、中英混用的转化成中文等;
与声音信息中主语种对应的文本信息;例如,中英混用的声音信息中主语种是中文,则将该中英混用的转化成中文;中英混用的声音信息中主语种是英文,则将该中英混用的转化成英文等;
与声音信息中次语种对应的文本信息;例如,中英混用的声音信息中次语种是中文,则将该中英混用的转化成中文;中英混用的声音信息中次语种是英文,则将该中英混用的转化成英文等。
当然,还可以将方言转成目标语种对应的文本信息,例如,粤语转中文等。
由于会议中可能存在多人同时发言,特别是争论等过程中,可能很难分辨出谁说了什么话,通过前述实施例可知,本发明可以针对每一发言人生成对应的文本信息,因而,在投屏画面中渲染所述文本信息(字幕)时,可以采用不同形式的字幕来区分不同发言人的文本信息:
匹配与不同发言人对应的渲染属性,依据所述渲染属性在投屏画面中渲染所述文本信息;
其中,所述渲染属性包括以下至少之一:字体颜色、字体大小、字体粗细、显示方位、个性化标记;所述个性化标记包括以下任一:下划线、文字突出显示颜色。
如图4所示,一个发言人的字幕带底色,另一发言人的字幕无底色,当然,渲染属性的种类很多,还可以是采用不同的颜色等方式。可以将每一发言人的字幕在屏幕对应的一固定
位置显示,也可以不固定字幕出现的位置,或者类似弹幕普通发言人的字幕在两端显示、主讲人的字幕在中间显示等。
通常来说,会议的主讲人的发言内容是重点,因此,可以将与主讲人对应的文本信息以区别于其他发言人的形式着重显示。可以认为发送桌面数据的源端设备即为主讲人使用的源端设备,可以根据源端设备的MAC(Media Access Control,媒体访问控制)地址等,从而区分出哪些声音信息是主讲人的,将该声音信息对应的单一发言人作为主讲人,以区别于其他发言人的形式着重显示所述主讲人的文本信息。例如,如图4所示,带底色的字幕可以认为是主讲人的字幕,无底色的字幕属于普通发言人的。当然,也可以修改渲染属性,例如,设置与MAC地址对应的渲染属性,根据发送声音信息的MAC地址区分各发言人及对应的文本信息,进而为文本信息加载对应的渲染属性渲染到投屏画面中形成字幕。图4所示虽为由单一主讲人投屏(传屏设备中仅显示一源端设备的桌面数据)的情况,但目前已有在一个传屏设备中接收并展示多个源端设备桌面数据的实现方式,显然,传屏设备中展示一个或多个源端设备的桌面数据,没有改变本发明传屏方案的使用条件,因此,本发明的方案也适用于在传屏设备中展示多个源端设备桌面数据的情况。
开会通常需要进行会议记录,常规的录像或录屏等仅有画面和/或声音,回看录像时枯燥无味,且对于不熟悉会议情况的人,光听声音难以分辨出哪些话是谁说的,为此,本发明提出根据时间戳将文本信息与桌面数据相关联,在后续回放录制的桌面数据时,在对应的时间依次展示文本信息,当然,该文本信息也可以与前述声音信息同时出现,从而,在回看录像时容易辨识各发言人的发言内容;例如,有红、黑、蓝三种颜色的字幕,分别对应甲、乙、丙三个发言人,观看录像时通过将红色字幕与甲的声音对应、黑色字幕与乙的声音对应、蓝色字幕与丙的声音对应,能够轻松的分辨各发言人的发言内容。当然,该方式也可以应用于远程会议中,将本地的桌面数据、声音信息和/或文本信息发送至外地设备中,增加了异地参加会议的人了解会议内容的方式,提高了会议传达效果。
与前述传屏方法的实施例相对应,本发明还提供了传屏装置的实施例。
本发明传屏装置的实施例可以应用在会议平板上。装置实施例可以通过软件实现,也可以通过硬件或者软硬件结合的方式实现。以软件实现为例,作为一个逻辑意义上的装置,是通过其所在会议平板的处理器将非易失性存储器中对应的计算机程序指令读取到内存中运行形成的。从硬件层面而言,如图5所示,为本发明传屏装置所在会议平板的一种硬件结构图,除了图5所示的处理器、内存、网络接口、以及非易失性存储器之外,实施例中装置所在的会议平板通常根据该传屏的实际功能,还可以包括其他硬件,对此不再赘述。
请参考图6,该传屏装置600包括:
处理模块610,用于接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息;
渲染模块620,用于在投屏画面中渲染所述文本信息。
进一步地,本发明还提出了一种电子设备,包括:
处理器;
用于存储处理器可执行指令的存储器;
其中,所述处理器被配置为执行如前任意一项所述的传屏方法。
进一步地,本发明还提出了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现如前任意一项所述的传屏方法。
本发明所述的会议平板具有传屏功能,且在原在传屏功能基础上增加了音频转文字等功能,该功能可以是利用现有的音译软件实现,由传屏功能调用音译软件的音译结果;也可以将音译软件的功能复合在传屏功能中;当然,也可以根据实际情况设计其它能实现该功能的插件,本发明对此不作限定。
上述装置中各个单元的功能和作用的实现过程具体详见上述方法中对应步骤的实现过程,在此不再赘述。
对于装置实施例而言,由于其基本对应于方法实施例,所以相关之处参见方法实施例的部分说明即可。以上所描述的装置实施例仅仅是示意性的,其中所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部模块来实现本发明方案的目的。本领域普通技术人员在不付出创造性劳动的情况下,即可以理解并实施。
以上所述仅为本发明的较佳实施例而已,并不用以限制本发明,凡在本发明的精神和原则之内,所做的任何修改、等同替换、改进等,均应包含在本发明保护的范围之内。
Claims (12)
- 一种传屏方法,其特征在于,包括以下步骤:接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息;在投屏画面中渲染所述文本信息。
- 如权利要求1所述的传屏方法,其特征在于,所述接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息的步骤包括:对自身采集的周围环境的声音信息进行分析处理,根据特征识别将声音信息转换成与发言人对应的第一文本信息;接收源端设备采集并发送过来的周围环境的声音信息,根据该声音信息对第一文本信息进行校正。
- 如权利要求1所述的传屏方法,其特征在于,所述接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息的步骤包括:接收源端设备采集并发送过来的周围环境的声音信息,将该声音信息与自身采集的周围环境的声音信息进行分析处理,提取与发言人对应的声音信息;将与发言人对应的声音信息转换成文本信息。
- 如权利要求3所述的传屏方法,其特征在于,所述将该声音信息与自身采集的周围环境的声音信息进行分析处理,提取与发言人对应的声音信息的步骤包括:以从源端设备接收的声音信息作为参考信息,将自身采集的声音信息与参考信息进行相关性运算,去除环境噪声和/或其他发言人的声音信息,提取与单一发言人对应的声音信息。
- 如权利要求3所述的传屏方法,其特征在于,所述投屏画面为对源端设备发送来的桌面数据进行展示所得的画面,所述提取与单一发言人对应的声音信息的步骤之后,还包括:将发言人对应的声音信息进行语音处理,所述语音处理包括增益处理和/或衰减处理,使声音信息对应的音量处于预设范围内;根据时间戳将处理后的声音信息与桌面数据相关联。
- 如权利要求1所述的传屏方法,其特征在于,所述文本信息包括以下至少之一:与声音信息语种对应的文本信息;与目标语种对应的文本信息;与声音信息中主语种对应的文本信息;与声音信息中次语种对应的文本信息。
- 如权利要求1至6中任一项所述的传屏方法,其特征在于,在投屏画面中渲染所述文本信息的步骤包括:匹配与不同发言人对应的渲染属性,依据所述渲染属性在投屏画面中渲染所述文本信息;其中,所述渲染属性包括以下至少之一:字体颜色、字体大小、字体粗细、显示方位、个性化标记;所述个性化标记包括以下任一:下划线、文字突出显示颜色。
- 如权利要求7所述的传屏方法,其特征在于,在投屏画面中渲染所述文本信息的步骤包括:匹配发送桌面数据的源端设备所发送的周围环境的声音信息,将该声音信息对应的单一发言人作为主讲人,以区别于其他发言人的形式着重显示所述主讲人的文本信息。
- 如权利要求8所述的传屏方法,其特征在于,所述通过特征识别将与同一发言人对应的声音信息转换成同一文本信息的步骤之后,还包括:根据时间戳将文本信息与桌面数据相关联。
- 一种传屏装置,其特征在于,包括:处理模块,用于接收源端设备采集并发送过来的周围环境的声音信息,结合自身采集的周围环境的声音信息,通过特征识别将声音信息转换成与发言人对应的文本信息;渲染模块,用于在投屏画面中渲染所述文本信息。
- 一种电子设备,其特征在于,包括:处理器;用于存储处理器可执行指令的存储器;其中,所述处理器被配置为执行所述权利要求1-9中任意一项所述的传屏方法。
- 一种计算机可读存储介质,其上存储有计算机程序,其特征在于,该程序被处理器执行时实现如权利要求1-9中任意一项所述的传屏方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201710666179.8A CN107527623B (zh) | 2017-08-07 | 2017-08-07 | 传屏方法、装置、电子设备及计算机可读存储介质 |
| CN201710666179.8 | 2017-08-07 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019029073A1 true WO2019029073A1 (zh) | 2019-02-14 |
Family
ID=60680627
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2017/116067 Ceased WO2019029073A1 (zh) | 2017-08-07 | 2017-12-14 | 传屏方法、装置、电子设备及计算机可读存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN107527623B (zh) |
| WO (1) | WO2019029073A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111914115A (zh) * | 2019-05-08 | 2020-11-10 | 阿里巴巴集团控股有限公司 | 一种声音信息的处理方法、装置及电子设备 |
Families Citing this family (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109151642B (zh) * | 2018-09-05 | 2019-12-24 | 北京今链科技有限公司 | 一种智能耳机、智能耳机处理方法、电子设备及存储介质 |
| CN111770319B (zh) * | 2019-10-18 | 2022-04-12 | 北京沃东天骏信息技术有限公司 | 投影方法、装置、系统和存储介质 |
| CN113687803B (zh) * | 2020-05-19 | 2025-01-07 | 华为技术有限公司 | 投屏方法、投屏源端、投屏目的端、投屏系统及存储介质 |
| CN112019786B (zh) * | 2020-08-24 | 2021-05-25 | 上海松鼠课堂人工智能科技有限公司 | 智能教学录屏方法和系统 |
| CN112887781A (zh) * | 2021-01-27 | 2021-06-01 | 维沃移动通信有限公司 | 字幕处理方法及装置 |
| CN112684967A (zh) * | 2021-03-11 | 2021-04-20 | 荣耀终端有限公司 | 一种用于字幕显示的方法及电子设备 |
| CN113746911A (zh) * | 2021-08-26 | 2021-12-03 | 科大讯飞股份有限公司 | 音频处理方法及相关装置、电子设备、存储介质 |
| CN114125358A (zh) * | 2021-11-11 | 2022-03-01 | 北京有竹居网络技术有限公司 | 云会议字幕显示方法、系统、装置、电子设备和存储介质 |
| CN115052126B (zh) * | 2022-08-12 | 2022-10-28 | 深圳市稻兴实业有限公司 | 一种基于人工智能的超高清视频会议分析管理系统 |
| CN115421679B (zh) * | 2022-09-02 | 2026-04-03 | 兴容(上海)信息技术股份有限公司 | 一种多用户投屏方法和系统 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20120245936A1 (en) * | 2011-03-25 | 2012-09-27 | Bryan Treglia | Device to Capture and Temporally Synchronize Aspects of a Conversation and Method and System Thereof |
| CN104240718A (zh) * | 2013-06-12 | 2014-12-24 | 株式会社东芝 | 转录支持设备和方法 |
| CN205647778U (zh) * | 2016-04-01 | 2016-10-12 | 安徽听见科技有限公司 | 一种智能会议系统 |
| CN106057193A (zh) * | 2016-07-13 | 2016-10-26 | 深圳市沃特沃德股份有限公司 | 基于电话会议的会议记录生成方法和装置 |
| CN106911832A (zh) * | 2017-04-28 | 2017-06-30 | 上海与德科技有限公司 | 一种语音记录的方法及装置 |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103456305B (zh) * | 2013-09-16 | 2016-03-09 | 东莞宇龙通信科技有限公司 | 终端和基于多个声音采集单元的语音处理方法 |
| GB2530983A (en) * | 2014-09-30 | 2016-04-13 | Ibm | Content mirroring |
| CN104796584A (zh) * | 2015-04-23 | 2015-07-22 | 南京信息工程大学 | 具有语音识别功能的提词装置 |
| CN106297794A (zh) * | 2015-05-22 | 2017-01-04 | 西安中兴新软件有限责任公司 | 一种语音文字的转换方法及设备 |
| CN106910504A (zh) * | 2015-12-22 | 2017-06-30 | 北京君正集成电路股份有限公司 | 一种基于语音识别的演讲提示方法及装置 |
| CN105704538A (zh) * | 2016-03-17 | 2016-06-22 | 广东小天才科技有限公司 | 一种音视频字幕生成方法及系统 |
| CN105913845A (zh) * | 2016-04-26 | 2016-08-31 | 惠州Tcl移动通信有限公司 | 一种移动终端识别语音生成字幕的方法、系统及移动终端 |
| CN106657865B (zh) * | 2016-12-16 | 2020-08-25 | 联想(北京)有限公司 | 会议纪要的生成方法、装置及视频会议系统 |
-
2017
- 2017-08-07 CN CN201710666179.8A patent/CN107527623B/zh active Active
- 2017-12-14 WO PCT/CN2017/116067 patent/WO2019029073A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20120245936A1 (en) * | 2011-03-25 | 2012-09-27 | Bryan Treglia | Device to Capture and Temporally Synchronize Aspects of a Conversation and Method and System Thereof |
| CN104240718A (zh) * | 2013-06-12 | 2014-12-24 | 株式会社东芝 | 转录支持设备和方法 |
| CN205647778U (zh) * | 2016-04-01 | 2016-10-12 | 安徽听见科技有限公司 | 一种智能会议系统 |
| CN106057193A (zh) * | 2016-07-13 | 2016-10-26 | 深圳市沃特沃德股份有限公司 | 基于电话会议的会议记录生成方法和装置 |
| CN106911832A (zh) * | 2017-04-28 | 2017-06-30 | 上海与德科技有限公司 | 一种语音记录的方法及装置 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111914115A (zh) * | 2019-05-08 | 2020-11-10 | 阿里巴巴集团控股有限公司 | 一种声音信息的处理方法、装置及电子设备 |
| CN111914115B (zh) * | 2019-05-08 | 2024-05-28 | 阿里巴巴集团控股有限公司 | 一种声音信息的处理方法、装置及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN107527623B (zh) | 2021-02-09 |
| CN107527623A (zh) | 2017-12-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2019029073A1 (zh) | 传屏方法、装置、电子设备及计算机可读存储介质 | |
| US12387741B2 (en) | Automated transcript generation from multi-channel audio | |
| CN205647778U (zh) | 一种智能会议系统 | |
| US6771302B1 (en) | Videoconference closed caption system and method | |
| US8120638B2 (en) | Speech to text conversion in a videoconference | |
| US8655654B2 (en) | Generating representations of group interactions | |
| US20120176467A1 (en) | Sharing Participant Information in a Videoconference | |
| US12387738B2 (en) | Distributed teleconferencing using personalized enhancement models | |
| US12229471B2 (en) | Centrally controlling communication at a venue | |
| US12243551B2 (en) | Performing artificial intelligence sign language translation services in a video relay service environment | |
| WO2021057957A1 (zh) | 视频通话方法、装置、计算机设备和存储介质 | |
| US9584761B2 (en) | Videoconference terminal, secondary-stream data accessing method, and computer storage medium | |
| CN110933485A (zh) | 一种视频字幕生成方法、系统、装置和存储介质 | |
| WO2005072336A2 (en) | Method for aiding and enhancing verbal communications | |
| JP2013201505A (ja) | テレビ会議システム及び多地点接続装置並びにコンピュータプログラム | |
| CN120568004A (zh) | 视频会议的处理方法、装置、电子设备、存储介质及产品 | |
| CN111816183B (zh) | 基于音视频录制的语音识别方法、装置、设备及存储介质 | |
| US20200184973A1 (en) | Transcription of communications | |
| TW202009750A (zh) | 即時語音自動同步轉譯字幕直播系統及方法 | |
| TW201410028A (zh) | 影音文字紀錄系統 | |
| TWM574267U (zh) | 即時語音自動同步轉譯字幕直播系統 | |
| KR20180068655A (ko) | 음성 신호에 기초한 문자 생성 장치 및 방법 | |
| CN114245255A (zh) | Tws耳机及其实时传译方法、终端及存储介质 | |
| Patil et al. | MuteTrans: A communication medium for deaf | |
| TWM519272U (zh) | 視訊翻譯系統及客戶裝置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17920770 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 17.06.2020) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17920770 Country of ref document: EP Kind code of ref document: A1 |