WO2020147394A1 - 基于人工智能的电视字幕动态生成方法及相关设备 - Google Patents

基于人工智能的电视字幕动态生成方法及相关设备 Download PDF

Info

Publication number
WO2020147394A1
WO2020147394A1 PCT/CN2019/117013 CN2019117013W WO2020147394A1 WO 2020147394 A1 WO2020147394 A1 WO 2020147394A1 CN 2019117013 W CN2019117013 W CN 2019117013W WO 2020147394 A1 WO2020147394 A1 WO 2020147394A1
Authority
WO
WIPO (PCT)
Prior art keywords
current
language
official language
voice
official
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/117013
Other languages
English (en)
French (fr)
Inventor
朱胜强
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Publication of WO2020147394A1 publication Critical patent/WO2020147394A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L13/00Speech synthesis; Text to speech systems
    • G10L13/02Methods for producing synthetic speech; Speech synthesisers
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/26Speech to text systems
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/431Generation of visual interfaces for content selection or interaction; Content or additional data rendering
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/435Processing of additional data, e.g. decrypting of additional data, reconstructing software from modules extracted from the transport stream
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/439Processing of audio elementary streams
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/441Acquiring end-user identification, e.g. using personal code sent by the remote control or by inserting a card
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/442Monitoring of processes or resources, e.g. detecting the failure of a recording device, monitoring the downstream bandwidth, the number of times a movie has been viewed, the storage space available from the internal hard disk
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/45Management operations performed by the client for facilitating the reception of or the interaction with the content or administrating data related to the end-user or to the client device itself, e.g. learning user preferences for recommending movies, resolving scheduling conflicts
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/488Data services, e.g. news ticker

Definitions

  • This application relates to the field of artificial intelligence, and in particular to an artificial intelligence-based method for dynamically generating TV subtitles and related equipment.
  • the purpose of this application is to address the deficiencies of the prior art, and provide a method and related equipment for dynamically generating TV subtitles based on artificial intelligence, to identify the ethnicity of the current viewer, determine the language of video playback subtitles, and the language of voice playback , Realize automatic switching of the language to be played, help people with disabilities understand the TV content, and improve the efficiency of TV listening.
  • the technical solution of the present application provides an artificial intelligence-based method for dynamically generating TV subtitles and related equipment.
  • This application discloses a method for dynamically generating TV subtitles based on artificial intelligence, which includes the following steps:
  • the application also discloses an artificial intelligence-based TV subtitle dynamic generation device, which includes:
  • Audience image acquisition module configured to acquire a screen image of the current audience, and input the screen image into a face recognition model to obtain the ethnic information of the current audience and the proportion of the number of people corresponding to the ethnicity of the current audience;
  • Subtitle language confirmation module set to obtain current positioning information, and obtain the current official language according to the current positioning information, the ethnic information of the current audience, and the proportion of the number of people corresponding to the ethnicity of the current audience;
  • Subtitle playback module set to obtain voice data, convert the voice data into text through a voice recognition model, translate the text according to the current official language, and subtitle the translated text on the TV screen Play
  • Voice playback module set to switch the current voice to the current official language for voice playback according to the subtitles.
  • the application also discloses a computer device, the computer device includes a memory and a processor, the memory stores computer-readable instructions, and when the computer-readable instructions are executed by one or more of the processors, One or more of the processors perform the following steps:
  • the application also discloses a storage medium that can be read and written by a processor, and the storage medium stores computer instructions.
  • the computer-readable instructions are executed by one or more processors, one or more Each processor performs the following steps:
  • this application determines the language of video playback subtitles and the language of voice playback by identifying the ethnicity of the current viewers, and realizes automatic switching of the playback language, helping people with disabilities understand TV content and improving TV listening efficiency.
  • FIG. 1 is a schematic diagram of a first embodiment of a method for dynamically generating TV subtitles based on artificial intelligence according to an embodiment of the application;
  • FIG. 2 is a schematic diagram of a second embodiment of a method for dynamically generating TV subtitles based on artificial intelligence according to an embodiment of the application;
  • FIG. 3 is a schematic diagram of a third embodiment of a method for dynamically generating TV subtitles based on artificial intelligence according to an embodiment of the application;
  • FIG. 4 is a schematic diagram of a fourth embodiment of a method for dynamically generating TV subtitles based on artificial intelligence according to an embodiment of the application;
  • FIG. 5 is a schematic diagram of a fifth embodiment of a method for dynamically generating TV subtitles based on artificial intelligence according to an embodiment of the application;
  • FIG. 6 is a schematic diagram of a sixth embodiment of a method for dynamically generating TV subtitles based on artificial intelligence according to an embodiment of the application;
  • Figure 7 is a schematic structural diagram of an artificial intelligence-based TV subtitle dynamic generation device according to an embodiment of the application.
  • FIG. 1 The flow of an artificial intelligence-based method for dynamically generating TV subtitles according to the first embodiment of the present application is shown in FIG. 1. This embodiment includes the following steps:
  • Step s101 Obtain a screen image of the current audience, and input the screen image into a face recognition model to obtain the ethnic information of the current audience and the proportion of the number of people corresponding to the ethnicity of the current audience;
  • the acquisition of the current viewer's screen image can be accomplished through a video camera.
  • a video camera is installed on a TV in a port, an airport, and a public place to capture a live screen.
  • the video camera can be used in real time.
  • Obtain the current viewer's screen image As the viewer usually stands in front of the TV when watching TV, the current viewer's screen image can be obtained through the video camera; the current screen image can be captured by the video camera to obtain the current screen image , It is also possible to record a piece of video, and intercept the screen image in the video, and the obtained screen image can be sent to the background system for further analysis and processing.
  • the screen image after acquiring the screen image of the current audience, the screen image can be recognized through a face recognition model, and the recognition is the number of people in the screen, the race corresponding to each person, and The proportion of the number of each race in the total audience.
  • Step s102 obtaining current positioning information, and obtaining the current official language according to the current positioning information, the ethnic information of the current audience, and the proportion of the number of people corresponding to the ethnicity of the current audience;
  • the acquisition of current positioning information can be accomplished through a GPS positioning system.
  • a GPS positioning module is installed in a television, and the country and region where the television is currently located can be obtained through the GPS positioning module.
  • There is a fixed official language but in a country or region, there may be different races, so there may be multiple official languages, and the current audience may also have multiple races, such as yellow and white. People and black people, therefore, the current official language is obtained by statistical analysis of the current positioning information, the current audience’s race, and the proportion of different audience races in the total audience.
  • Step s103 Acquire voice data, convert the voice data into text through a voice recognition model, translate the text according to the current official language, and play the translated text on a TV screen with subtitles;
  • the voice data corresponding to the television programs are usually It is a local language and is not suitable for traveling passengers.
  • the voice data can be converted into text through the voice recognition model, and the text can be converted through the translation software Translate according to the official language in step s102.
  • the translation is completed, the translated text is played on the TV screen. Since the current official language is determined according to the current audience, it is based on the current official language The text to be translated also meets the needs of the current audience, that is, it can be recognized and accepted by the current audience.
  • Step s104 Switch the current voice to the current official language according to the subtitles for voice playback.
  • the current voice can also be switched to the current official language according to the subtitles and the voice is played through the speech synthesis model, that is, when the current official language is determined and After subtitles are played in the current official language, the current voice can be turned off first, and then according to the currently translated subtitles, it is converted into the voice corresponding to the current official language according to the speech synthesis model, so that the current audience can understand it, which is conducive to visual inconvenience. Audience with convenient hearing.
  • the language of video playback subtitles and the language of voice playback are determined, so as to realize automatic switching of the playback language, help people with disabilities understand TV content, and improve TV listening efficiency.
  • Fig. 2 is a schematic flow chart of a method for dynamically generating artificial intelligence-based TV subtitles according to the second embodiment of the application.
  • the step s101 obtaining a screen image of the current viewer, includes:
  • Step s201 preset the shooting period of the current viewer's screen image
  • the shooting period of the picture image can be preset in the TV system, such as once every 30 minutes, or it can be adjusted manually. For example, set the shooting period to be shorter during the peak period of passenger flow. The shooting period is set to be longer, such as shooting once per hour.
  • Step s202 Acquire the screen image of the current audience according to the shooting period of the screen image of the current audience.
  • the television system includes a video camera, and when the shooting period is set in the television system, once the shooting period is reached, the television system will start the video camera to take pictures of the current audience and obtain the current The image of the audience.
  • the current official language can be regularly updated, and the subtitles can be dynamically updated according to the current official language, thereby improving the degree of adaptation of the subtitles to the current audience.
  • Figure 3 is a schematic flow chart of a method for dynamically generating artificial intelligence-based TV subtitles according to a third embodiment of the application.
  • the step s102 is based on the current positioning information, the current viewer’s ethnic information and The proportion of the population corresponding to the race of the current audience obtains the current official language, including:
  • Step s301 preset the correspondence between race and location information, the official language and the population proportion corresponding to the official language, and store the correspondence in a database;
  • the same race may also have different official languages, and the proportion of the population corresponding to different official languages in a certain race is also different.
  • the official The languages include French and English.
  • the French population accounts for 30% and the English population accounts for 70%.
  • ethnic information can be obtained through step s101.
  • the positioning information can be used to determine a region, such as a country or region, and the country or region corresponds to an official language, and different official languages correspond to different population proportions.
  • a region such as a country or region
  • different official languages correspond to different population proportions.
  • the correspondence between ethnicity and location information, the official language and the proportion of the population corresponding to the official language and store the correspondence in the database.
  • the correspondence between white people and Canada is French 30% and English 70%
  • the yellow race corresponds to 90% in Mandarin and 10% in English.
  • Step s302 Query in the database according to the acquired current positioning information and the ethnic information of the current audience to obtain the official language corresponding to the current positioning information and the ethnicity of the current audience and the official language corresponding to the official language The proportion of the population;
  • the current geographic information can be obtained, for example, the current country is China; when the ethnic information is obtained according to step s101, if the current audience is a yellow race, according to the geographic information and people Kind of information is queried in the database in step s301 to obtain the official language corresponding to the current positioning information and the ethnicity of the current audience and the population proportion corresponding to the official language.
  • the corresponding is Mandarin 90% and English 10%
  • the Mandarin and English are the official languages corresponding to the current positioning information and the race of the current audience
  • the 90% and 10% are the proportions of the population corresponding to the official language, that is, speaking Mandarin speakers account for 90% of the population
  • English speakers account for 10% of the population.
  • Step s303 Obtain the current official language according to the proportion of the population corresponding to the official language and the proportion of the population corresponding to the race of the current audience.
  • step s302 since there may be multiple official languages queried in step s302, the proportion of the population corresponding to the official languages will also be different; in addition, the current audience will also have different types and different races. The proportion of the total number of viewers to the total number of viewers will also be different; therefore, statistical analysis can be made according to the proportion of the obtained population corresponding to the official language and the number of races to the total number of viewers.
  • the official language is the subtitle language currently required by most viewers.
  • the current subtitle language is acquired through positioning information and ethnic information, which can improve the current user experience and satisfaction.
  • step s303 obtaining the current official language according to the proportion of the population corresponding to the official language and the proportion of the number of people corresponding to the race of the current audience includes:
  • the current subtitle language is obtained after the weighted summation of the population proportion corresponding to the official language and the proportion of the population corresponding to the race of the current audience, which can improve user experience and satisfaction.
  • Figure 4 is a schematic flow chart of a method for dynamically generating artificial intelligence-based TV subtitles according to the fourth embodiment of the application.
  • voice data is obtained, and the voice data is converted into The text, the text is translated according to the current official language, and the translated text is subtitled on the TV screen, including:
  • Step s401 Acquire voice data and a preset default subtitle language corresponding to the voice data
  • the voice data and the preset default subtitle language corresponding to the voice data can be acquired first.
  • the preset default subtitle language means that the subtitle language can be preset in the television system, such as English, regardless of What language is the voice broadcast in the current TV program, such as Italian, Chinese, or German, will eventually display the preset English in the subtitles, and the settings can be manually modified, such as modifying the settings before the TV program is broadcast.
  • Step s402 Convert the voice data into text through a voice recognition model, and translate the text according to the default subtitle language and the current official language;
  • the voice data can be converted into text through the voice recognition model.
  • the current broadcast is a Chinese program
  • the Chinese program can be passed through
  • the speech recognition model is converted into Chinese, and then translated according to the default subtitle language and the current official language.
  • the translation order can be arbitrary. You can first translate according to the default subtitle language or first according to the current official language. Language translation, after translation, two kinds of subtitles can be obtained.
  • step s403 subtitles are played on the translated text.
  • the translated text can be subtitled.
  • the subtitle needs of a certain number of people can be met.
  • FIG. 5 is a schematic flow chart of a method for dynamically generating artificial intelligence-based TV subtitles according to the fifth embodiment of the application. As shown in the figure, in step s402, according to the default subtitle language and the current official language, Translate the text, including:
  • Step s501 comparing the current official language with the default subtitle language
  • the default subtitle language is set in the TV system in advance, when the current audience is recognized, it is possible that the recognized official language and the preset default subtitle language are the same. For example, in a Chinese airport, a certain A group of Chinese people are watching TV at a moment, and the default subtitle language is Chinese. At this time, the current official language can be compared with the default subtitle language.
  • step s502 if they are consistent, the text is translated according to the current official language, and if they are not consistent, the text is translated sequentially according to the default subtitle language and the current official language.
  • the text can be translated in sequence according to the default subtitle language and the current official language.
  • the order of the translation can be arbitrary, and it can be performed according to the default subtitle language first.
  • For translation it can also be translated according to the current official language, and the translated text can be played with bilingual subtitles. For example, at an airport in China, Chinese subtitles are preset, and a group of British people are watching. After ethnic identification, it is confirmed that the current official language is English, and then dual subtitles in Chinese and English can be played.
  • Figure 6 is a schematic flow chart of a method for dynamically generating artificial intelligence-based TV subtitles according to the sixth embodiment of the application. As shown in the figure, in step s104, the current voice is switched to the current official language according to the subtitles.
  • Voice playback including:
  • Step s601 Obtain the current playback language of the voice data, and compare the current playback language with the current official language;
  • the current playback language of the voice data can be identified through the voice recognition model, such as whether the current playback is Chinese, English, or Japanese, and when the current playback language is obtained, it can be compared with the current official language.
  • step s602 if they are inconsistent, convert the subtitles into the voice of the current official language through a speech synthesis model, and replace the voice of the currently played language with the voice of the current official language for voice playback.
  • the current playing language can be switched to the current official language based on the current subtitles through the speech synthesis model for voice playback; first, the current subtitles can be converted from the speech synthesis model Convert into the voice of the current official language, then stop the voice of the current playing language, and play the voice of the current official language; if the current playing language is consistent with the current official language, then there is no need to perform speech synthesis on the current subtitles No need for voice switching.
  • the visually impaired people can be helped to obtain television information and the user experience can be improved.
  • FIG. 7 The structure of an artificial intelligence-based TV subtitle dynamic generation device according to an embodiment of the present application is shown in FIG. 7, and includes:
  • the playback module 703 is connected to the voice playback module 704;
  • the audience image acquisition module 701 is configured to acquire the screen image of the current audience, and input the screen image into the face recognition model to obtain the ethnic information of the current audience and the relationship with the current audience The proportion of the number of people corresponding to the race;
  • the subtitle language confirmation module 702 is configured to obtain current positioning information, and obtain the current official position according to the current positioning information, the race information of the current viewer, and the proportion of the number of people corresponding to the race of the current viewer Language;
  • the subtitle playback module 703 is configured to obtain voice data, convert the voice data into text through a voice recognition model, translate the text according to the current official language, and display the translated text on the TV screen Perform subtitle playback;
  • the voice playback module 704 is configured to switch the current voice to the current official language for voice playback according to the subtitles.
  • the audience image acquisition module includes:
  • Preset unit set to preset the shooting period of the current audience's screen image
  • the first obtaining unit configured to obtain the screen image of the current audience according to the shooting period of the screen image of the current audience.
  • the subtitle language confirmation module includes:
  • Storage unit set to preset correspondences between race and positioning information, official languages and population proportions corresponding to the official languages, and store the correspondences in a database;
  • Query unit configured to query the database according to the acquired current positioning information and the ethnic information of the current audience, and obtain the official language corresponding to the current positioning information and the ethnicity of the current audience and the official language The corresponding population ratio;
  • the second acquiring unit is configured to acquire the current official language according to the proportion of the population corresponding to the official language and the proportion of the number of people corresponding to the race of the current audience.
  • the subtitle language confirmation module includes:
  • the third acquisition unit configured to query the official language corresponding to the highest confidence probability in the official language, and use the official language corresponding to the highest confidence probability as the current official language.
  • the subtitle playback module includes:
  • Receiving unit configured to obtain voice data and a preset default subtitle language corresponding to the voice data
  • Conversion unit configured to convert the voice data into text through a voice recognition model, and translate the text according to the default subtitle language and the current official language;
  • Play unit set to play the translated text with subtitles.
  • the subtitle playback module includes:
  • the first comparison unit set to compare the current official language with the default subtitle language
  • the first output unit is set to translate the text according to the current official language if they are consistent, and translate the text in turn according to the default subtitle language and the current official language if they are inconsistent.
  • the voice playing module includes:
  • the second comparison unit set to obtain the current playback language of the voice data, and compare the current playback language with the current official language;
  • the second output unit is set to convert the subtitles into the voice of the current official language through a speech synthesis model if they are inconsistent, and replace the voice of the currently played language with the voice of the current official language for voice playback.
  • An embodiment of the present application also discloses a computer device that includes a memory and a processor.
  • the memory stores computer-readable instructions.
  • the computer-readable instructions are executed by one or more of the processors, , Enabling one or more of the processors to execute the steps in the method for dynamically generating TV subtitles in the foregoing embodiments.
  • the embodiment of the present application also discloses a storage medium that can be read and written by a processor, and the memory stores computer-readable instructions.
  • the computer-readable instructions are executed by one or more processors,
  • One or more processors execute the steps in the method for dynamically generating TV subtitles described in the foregoing embodiments.
  • the computer program can be stored in a computer readable storage medium. When executed, it may include the procedures of the above-mentioned method embodiments.
  • the aforementioned storage medium may be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (Read-Only Memory, ROM), or a random access memory (Random Access Memory, RAM), etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Databases & Information Systems (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Studio Circuits (AREA)

Abstract

本申请涉及人工智能领域,本申请公开了一种基于人工智能的电视字幕动态生成方法及相关设备,所述方法包括:获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放。本申请通过对观众人种的识别,确定字幕语言,帮助障碍人士对电视内容的理解,提高电视收听效率。

Description

基于人工智能的电视字幕动态生成方法及相关设备
本申请要求于2019年01月17日提交中国专利局、申请号为201910042720.7、发明名称为“基于人工智能的电视字幕动态生成方法及相关设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及人工智能领域,特别涉及一种基于人工智能的电视字幕动态生成方法及相关设备。
背景技术
目前,电视节目中除了电视剧和电影,综艺节目、新闻和现场直播的节目都不带字幕,在机场、公交车或者火车上环境嘈杂的地方都无法听到节目声音;针对一些听力障碍的人群,家里的电视机如果节目没有字幕也无法更好的欣赏节目。此外,由于节目制作成本限制,很多节目无法提供国语和外语的配音和字幕,中国人无法听懂外国节目,或者外国人无法听到中国语言类节目。而在交通及港口等人流量大的地方,由于人流变换频繁,如果字幕无法动态更新,就无法适应当前的观众的需要。
发明内容
本申请的目的在于针对现有技术的不足,提供一种基于人工智能的电视字幕动态生成方法及相关设备,对当前观看人员的人种的识别,确定视频播放字幕的语言,以及语音播放的语言,实现自动化切换播放的语言,帮助障碍人士对电视内容的理解,提高电视收听效率。
为达到上述目的,本申请的技术方案提供一种基于人工智能的电视字幕动态生成方法及相关设备。
本申请公开了一种基于人工智能的电视字幕动态生成方法,包括以下步骤:
获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;
获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;
获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述 当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;
根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
本申请还公开了一种基于人工智能的电视字幕动态生成装置,所述装置包括:
观众图像获取模块:设置为获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;
字幕语言确认模块:设置为获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;
字幕播放模块:设置为获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;
语音播放模块:设置为根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
本申请还公开了一种计算机设备,所述计算机设备包括存储器和处理器,所述存储器中存储有计算机可读指令,所述计算机可读指令被一个或多个所述处理器执行时,使得一个或多个所述处理器执行以下步骤:
获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;
获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;
获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;
根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
本申请还公开了一种存储介质,所述存储介质可被处理器读写,所述存储介质存储有计算机指令,所述计算机可读指令被一个或多个处理器执行时,使得一个或多个处理器执行以下步骤:
获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;
获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;
获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;
根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
本申请的有益效果是:本申请通过对当前观看人员的人种的识别,确定视频播放字幕的语言,以及语音播放的语言,实现自动化切换播放的语言,帮助障碍人士对电视内容的理解,提高电视收听效率。
附图说明
图1为本申请实施例的一种基于人工智能的电视字幕动态生成方法的第一个实施例示意图;
图2为本申请实施例的一种基于人工智能的电视字幕动态生成方法的第二个实施例示意图;
图3为本申请实施例的一种基于人工智能的电视字幕动态生成方法的第三个实施例示意图;
图4为本申请实施例的一种基于人工智能的电视字幕动态生成方法的第四个实施例示意图;
图5为本申请实施例的一种基于人工智能的电视字幕动态生成方法的第五个实施例示意图;
图6为本申请实施例的一种基于人工智能的电视字幕动态生成方法的第六个实施例示意图;
图7为本申请实施例的一种基于人工智能的电视字幕动态生成装置结构示 意图。
具体实施方式
为了使本申请的目的、技术方案及优点更加清楚明白,以下结合附图及实施例,对本申请进行进一步详细说明。应当理解,此处所描述的具体实施例仅仅用以解释本申请,并不用于限定本申请。
本技术领域技术人员可以理解,除非特意声明,这里使用的单数形式“一”、“一个”、“所述”和“该”也可包括复数形式。应该进一步理解的是,本申请的说明书中使用的措辞“包括”是指存在所述特征、整数、步骤、操作、元件和/或组件,但是并不排除存在或添加一个或多个其他特征、整数、步骤、操作、元件、组件和/或它们的组。
本申请第一个实施例的一种基于人工智能的电视字幕动态生成方法流程如图1所示,本实施例包括以下步骤:
步骤s101,获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;
具体的,所述获取当前观众的画面图像可通过视频摄像头完成,如在港口、机场以及公共场所的电视上安装视频摄像头捕捉现场画面,在公众观看电视的过程中,通过所述视频摄像头可实时获取当前观众的画面图像,由于通常观众在观看电视时都是站在电视机前,因此通过所述视频摄像头就可获取当前观众的画面图像;所述画面图像可以通过视频摄像头抓拍获取当前画面图像,也可通过录制一段视频,并在所述视频中截取画面图像,所述获取的画面图像可以发送到后台系统中进行进一步分析处理。
具体的,当获取到当前观众的画面图像后,可通过人脸识别模型对所述画面图像进行识别,所述识别为识别出所述画面中人物的数量,每个人物对应的人种,以及每个人种的数量在所述观众总数中的占比。
步骤s102,获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;
具体的,所述获取当前定位信息可通过GPS定位系统完成,如在电视机中安装GPS定位模块,通过所述GPS定位模块可获取当前电视机所处的国家和地 区,而国家和地区通常都是有固定的官方语言的,但是在一个国家或者地区中,可能还会有不同的种族,因此官方语言也可能有多种,而当前的观众也可能有多个人种,如黄种人、白种人及黑种人,因此通过对当前的定位信息、当前观众的人种以及不同观众人种在观众总数中的人数占比进行统计分析获得当前的官方语言。
步骤s103,获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;
具体的,位于港口、机场以及公共场所的电视机会播放当地的新闻或者其他电视节目,所述的节目从当地电视台获取或者预先录制在电视播放系统中,因此所述电视节目对应的语音数据通常都是本地语言,不适合往来的旅客观看,当电视机进行电视节目播放时,即获取到语音数据时,可先将所述语音数据通过语音识别模型转换成文字,并通过翻译软件将所述文字根据步骤s102中的官方语言进行翻译,当翻译完成后,将所述翻译好的文字在电视屏幕上进行播放,由于所述当前官方语言是根据当前的观众确定的,因此根据所述当前官方语言进行翻译的文字也是符合当前观众需求的,即能被当前观众认识和接受的。
步骤s104,根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
具体的,当对字幕根据当前官方语言进行翻译及播放后,还可将当前语音根据所述字幕将当前语音切换成当前官方语言并通过语音合成模型进行语音播放,即当确定当前的官方语言并根据当前官方语言进行字幕播放后,可先将当前的语音关闭,然后根据当前翻译的字幕根据语音合成模型转换成当前官方语言对应的语音,这样当前的观众才能听懂,有利于视觉不方便但听觉便利的观众。
本实施例中,通过对当前观看人员的人种的识别,确定视频播放字幕的语言,以及语音播放的语言,实现自动化切换播放的语言,帮助障碍人士对电视内容的理解,提高电视收听效率。
图2为本申请第二个实施例的一种基于人工智能的电视字幕动态生成方法 流程示意图,如图所示,所述步骤s101,获取当前观众的画面图像,包括:
步骤s201,预设当前观众的画面图像的拍摄周期;
具体的,可以在电视系统中预先设置画面图像的拍摄周期,如30分钟拍摄一次,也可进行人工调整,如在客流高峰时期将所述拍摄周期设的短一点,如在客流稀少的时期将所述拍摄周期设的长一点,如设置为1小时拍摄一次。
步骤s202,根据所述当前观众的画面图像的拍摄周期获取当前观众的画面图像。
具体的,所述电视系统包含视频摄像头,当在所述电视系统中设定拍摄周期后,一旦到了所述拍摄周期,所述电视系统就会启动视频摄像头对当前观众的画面进行拍摄,获取当前观众的画面图像。
本实施例中,通过定时获取当前观众的画面图像,可以定时更新当前官方语言,并根据当前官方语言动态更新字幕,提高字幕与当前观众的适配度。
图3为本申请第三个实施例的一种基于人工智能的电视字幕动态生成方法流程示意图,如图所示,所述步骤s102,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言,包括:
步骤s301,预设人种及定位信息与官方语言及与所述官方语言对应的人口比例的对应关系,并将所述对应关系存储在数据库中;
具体的,由于不同的国家或者地区有不同的人种,而同一人种还可能有不同的官方语言,且某一人种中不同的官方语言对应的人口比例也是不同的,如在加拿大国家,官方语言包括法语和英语,而在加拿大的白种人中,法语的人口比例占30%,英语的人口比例占70%。
具体的,人种信息可通过步骤s101获取,所述定位信息可用于确定地域,如国家或者地区,而国家或者地区是和官方语言对应的,且不同的官方语言对应不同的人口比例,因此可预设人种及定位信息与官方语言及与所述官方语言对应的人口比例的对应关系,并将所述对应关系存储在数据库中,例如,白种人和加拿大对应的是法语30%和英语70%,黄种人和中国对应的是普通话90%和英语10%。
步骤s302,根据所述获取的当前定位信息、当前观众的人种信息在所述数 据库中查询,获得与所述当前定位信息及当前观众的人种对应的官方语言及与所述官方语言对应的人口比例;
具体的,当获取到定位信息之后可获取到当前地域信息,如定位到当前国家是中国;当根据步骤s101获取到人种信息后,如当前观众是黄种人,可根据所述地域信息和人种信息在步骤s301的数据库中进行查询,获得与所述当前定位信息及当前观众的人种对应的官方语言及与所述官方语言对应的人口比例,以中国+黄种人为例,对应的是普通话90%和英语10%,所述普通话和英语是与所述当前定位信息及当前观众的人种对应的官方语言,所述90%和10%是与所述官方语言对应的人口比例,即说普通话的人数占人口总数的90%,说英语的人数占人口总数的10%。
步骤s303,根据所述与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比获得当前官方语言。
具体的,由于在步骤s302中查询到的可能有多种官方语言,而所述官方语言对应的人口比例也会有所不同;此外,当前观众的人种也会有不同种类,而不同人种的数量占当前观众总数量的比例也会不同;因此可以根据所述获取的与所述官方语言对应的人口比例及人种的数量占当前观众总数量的比例进行统计分析获取当前字幕所需的官方语言,即为当前大多数观众所需要的字幕语言。
本实施例中,通过定位信息和人种信息获取当前的字幕语言,可以提高当前用户的体验感和满意度。
在一个实施例中,所述步骤s303,根据所述与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比获得当前官方语言,包括:
根据公式:
W=∑ iX iY i
获得每个官方语言的置信概率,其中,W为官方语言的置信概率,X i为第i个人种使用所述官方语言的人口比例,Y i为第i个人种占当前观众总人数的人数占比;
在所述官方语言中,查询最高置信概率对应的官方语言,并将所述最高置信概率对应的官方语言作为当前官方语言。
具体的,当获取到与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比之后,可根据公式:
W=∑ iX iY i
获得每个官方语言的置信概率,其中,W为官方语言的置信概率,X i为第i个人种使用所述官方语言的人口比例,Y i为第i个人种占当前观众总人数的人数占比。以美国新墨西哥州为例,官方语言包括西班牙语和英语,首先通过步骤s101获取人种及所述人种在当前观众总数中的人数占比,通过摄像头拍摄并经过人脸识别后的结果为白种人占40%,其他人种占60%,而在步骤s301的数据库中可查询到白种人说英语的人口比例为70%,说西班牙语的人口比例为30%,其他人种说英语的人口比例为40%,说西班牙语的人口比例为90%,那么本次英语的置信概率为70%*40%+10%*60%=0.34,而本次西班牙语的置信概率为30%*40%+90%*60%=0.66;在所述英语和西班牙语的置信概率中进行查询,并找出最高的置信概率对应的官方语言,由于西班牙语的置信概率大于英语的置信概率,因此可将西班牙语作为当前官方语言。
本实施例中,通过对与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比的加权求和后获取当前字幕语言,可以提高用户体验感和满意度。
图4为本申请第四个实施例的一种基于人工智能的电视字幕动态生成方法流程示意图,如图所示,所述步骤s103,获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放,包括:
步骤s401,获取语音数据,以及与所述语音数据对应的预设的默认字幕语言;
具体的,首先可获取语音数据及与所述语音数据对应的预设的默认字幕语言,所述预设的默认字幕语言是指在电视系统中可以预先设置字幕语言,如事先设置为英语,不管当前电视节目中语音播放的是什么语言,如意大利文、中文或者德文,最终都会在字幕中显示预设的英文,所述的设置可以通过人工修改,如在电视节目播放前进行修改设置。
步骤s402,将所述语音数据通过语音识别模型转换成文字,并根据所述默认字幕语言及所述当前官方语言,对所述文字进行翻译;
具体的,当获取到当面的电视节目后,即获取到语音数据后,可先将所述语音数据通过语音识别模型转换成文字,如当前播放的是中文节目,那么可将所述中文节目通过语音识别模型转换成中文,然后根据所述默认字幕语言及所述当前官方语言进行翻译,所述翻译顺序可以任意,即可先根据所述默认字幕语言进行翻译,也可先根据所述当前官方语言进行翻译,经过翻译后可获得两种翻译字幕。
步骤s403,将翻译后的文字进行字幕播放。
具体的,当将所述语音识别后的文字根据所述默认字幕语言及所述当前官方语言进行翻译后,就可将所述翻译后的文字进行字幕播放了。
本实施例中,通过预设字幕语言,可以满足一定数量人群的字幕需要。
图5为本申请第五个实施例的一种基于人工智能的电视字幕动态生成方法流程示意图,如图所示,所述步骤s402,根据所述默认字幕语言及所述当前官方语言,对所述文字进行翻译,包括:
步骤s501,将所述当前官方语言与所述默认字幕语言进行比较;
具体的,如果预先在电视系统中设置了默认的字幕语言,那么当对当前观众进行识别时,有可能识别过后的官方语言和预设的默认字幕语言是一致的,如在中国的机场,某一时刻有一群中国人在观看电视,而预设的字幕语言是中文,这时可将所述当前的官方语言与默认的字幕语言进行比较。
步骤s502,如果一致,则根据所述当前官方语言将所述文字进行翻译,如果不一致,则根据所述默认字幕语言及所述当前官方语言依次将所述文字进行翻译。
具体的,如果当前的官方语言与默认的字幕语言是一致的,这时可以任意选取当前的官方语言与默认的字幕语言中的一种,例如,选取当前的官方语言进行翻译;而如果当前的官方语言与默认的字幕语言不一致,这时可以根据所述默认字幕语言及所述当前官方语言依次将所述文字进行翻译,所述翻译的顺序可以任意,即可先根据所述默认字幕语言进行翻译,也可先根据所述当前官 方语言进行翻译,并将所述翻译后的文字进行双语字幕播放,如在中国的机场,预设的是中文字幕,而这时有一群英国人在看,经过人种识别后确认当前官方语言是英语,那么可进行中文和英语双字幕播放。
本实施例中,通过当前的官方语言与默认的字幕语言的比较,可以避免重复的字幕播放,也可以进行双语字幕播放,提高字幕播放效率,提高用户体验感。
图6为本申请第六个实施例的一种基于人工智能的电视字幕动态生成方法流程示意图,如图所示,所述步骤s104,根据所述字幕将当前语音切换成所述当前官方语言进行语音播放,包括:
步骤s601,获取语音数据的当前播放语言,并将所述当前播放语言与所述当前官方语言进行比较;
具体的,通过语音识别模型可以识别当前语音数据的播放语言,比如当前播放是的中文、英文还是日文,并当获取到当前播放语言后,可和当前的官方语言进行比较。
步骤s602,如果不一致,将所述字幕通过语音合成模型转换成所述当前官方语言的语音,并将所述当前播放语言的语音替换成所述当前官方语言的语音进行语音播放。
具体的,如果所述当前播放语言与所述当前官方语言不一致,那么可通过语音合成模型根据当前的字幕将当前播放语言切换成当前官方语言进行语音播放;首先可通过语音合成模型将当前的字幕转换成当前官方语言的语音,然后将当前播放语言的语音停掉,并播放当前官方语言的语音;如果所述当前播放语言与所述当前官方语言一致,那么不需要对当前的字幕进行语音合成转换,也不需要进行语音的切换。
本实施例中,通过将所述字幕根据当前官方语言进行语音播放,可以帮助视力障碍人群获取电视信息,提高用户体验感。
本申请实施例的一种基于人工智能的电视字幕动态生成装置结构如图7所示,包括:
观众图像获取模块701、字幕语言确认模块702、字幕播放模块703及语音 播放模块704;其中,观众图像获取模块701与字幕语言确认模块702相连,字幕语言确认模块702与字幕播放模块703相连,字幕播放模块703与语音播放模块704相连;观众图像获取模块701设置为获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;字幕语言确认模块702设置为获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;字幕播放模块703设置为获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;语音播放模块704设置为根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
在一个实施例中,观众图像获取模块,包括:
预置单元:设置为预设当前观众的画面图像的拍摄周期;
第一获取单元:设置为根据所述当前观众的画面图像的拍摄周期获取当前观众的画面图像。
在一个实施例中,字幕语言确认模块,包括:
存储单元:设置为预设人种及定位信息与官方语言及与所述官方语言对应的人口比例的对应关系,并将所述对应关系存储在数据库中;
查询单元:设置为根据所述获取的当前定位信息、当前观众的人种信息在所述数据库中查询,获得与所述当前定位信息及当前观众的人种对应的官方语言及与所述官方语言对应的人口比例;
第二获取单元:设置为根据所述与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比获得当前官方语言。
在一个实施例中,字幕语言确认模块,包括:
计算单元:设置为根据公式:
W=∑ iX iY i
获得每个官方语言的置信概率,其中,W为官方语言的置信概率,X i为第i个人种使用所述官方语言的人口比例,Y i为第i个人种占当前观众总人数的人数占比;
第三获取单元:设置为在所述官方语言中,查询最高置信概率对应的官方语言,并将所述最高置信概率对应的官方语言作为当前官方语言。
在一个实施例中,字幕播放模块,包括:
接收单元:设置为获取语音数据,以及与所述语音数据对应的预设的默认字幕语言;
转换单元:设置为将所述语音数据通过语音识别模型转换成文字,并根据所述默认字幕语言及所述当前官方语言,对所述文字进行翻译;
播放单元:设置为将翻译后的文字进行字幕播放。
在一个实施例中,字幕播放模块,包括:
第一比较单元:设置为将所述当前官方语言与所述默认字幕语言进行比较;
第一输出单元:设置为如果一致,则根据所述当前官方语言将所述文字进行翻译,如果不一致,则根据所述默认字幕语言及所述当前官方语言依次将所述文字进行翻译。
在一个实施例中,语音播放模块,包括:
第二比较单元:设置为获取语音数据的当前播放语言,并将所述当前播放语言与所述当前官方语言进行比较;
第二输出单元:设置为如果不一致,将所述字幕通过语音合成模型转换成所述当前官方语言的语音,并将所述当前播放语言的语音替换成所述当前官方语言的语音进行语音播放。
本申请实施例还公开了一种计算机设备,所述计算机设备包括存储器和处理器,所述存储器中存储有计算机可读指令,所述计算机可读指令被一个或多个所述处理器执行时,使得一个或多个所述处理器执行上述各实施例中所述电视字幕动态生成方法中的步骤。
本申请实施例还公开了一种存储介质,所述存储介质可被处理器读写,所述存储器存储有计算机可读指令,所述计算机可读指令被一个或多个处理器执行时,使得一个或多个处理器执行上述各实施例中所述电视字幕动态生成方法中的步骤。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程, 是可以通过计算机程序来指令相关的硬件来完成,该计算机程序可存储于一计算机可读取存储介质中,该程序在执行时,可包括如上述各方法的实施例的流程。其中,前述的存储介质可为磁碟、光盘、只读存储记忆体(Read-Only Memory,ROM)等非易失性存储介质,或随机存储记忆体(Random Access Memory,RAM)等。
以上所述实施例的各技术特征可以进行任意的组合,为使描述简洁,未对上述实施例中的各个技术特征所有可能的组合都进行描述,然而,只要这些技术特征的组合不存在矛盾,都应当认为是本说明书记载的范围。
以上所述实施例仅表达了本申请的几种实施方式,其描述较为具体和详细,但并不能因此而理解为对本申请专利范围的限制。应当指出的是,对于本领域的普通技术人员来说,在不脱离本申请构思的前提下,还可以做出若干变形和改进,这些都属于本申请的保护范围。因此,本申请专利的保护范围应以所附权利要求为准。

Claims (20)

  1. 一种基于人工智能的电视字幕动态生成方法,包括以下步骤:
    获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;
    获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;
    获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;
    根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
  2. 如权利要求1所述的基于人工智能的电视字幕动态生成方法,其中,所述获取当前观众的画面图像,包括:
    预设当前观众的画面图像的拍摄周期;
    根据所述当前观众的画面图像的拍摄周期获取当前观众的画面图像。
  3. 如权利要求1所述的基于人工智能的电视字幕动态生成方法,其中,所述根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言,包括:
    预设人种及定位信息与官方语言及与所述官方语言对应的人口比例的对应关系,并将所述对应关系存储在数据库中;
    根据所述获取的当前定位信息、当前观众的人种信息在所述数据库中查询,获得与所述当前定位信息及当前观众的人种对应的官方语言及与所述官方语言对应的人口比例;
    根据所述与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比获得当前官方语言。
  4. 如权利要求3所述的基于人工智能的电视字幕动态生成方法,其中,所述根据所述与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比获得当前官方语言,包括:
    根据公式:
    W=∑ iX iY i
    获得每个官方语言的置信概率,其中,W为官方语言的置信概率,X i为第i个人种使用所述官方语言的人口比例,Y i为第i个人种占当前观众总人数的人数占比;
    在所述官方语言中,查询最高置信概率对应的官方语言,并将所述最高置信概率对应的官方语言作为当前官方语言。
  5. 如权利要求1所述的基于人工智能的电视字幕动态生成方法,其中,所述获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放,包括:
    获取语音数据,以及与所述语音数据对应的预设的默认字幕语言;
    将所述语音数据通过语音识别模型转换成文字,并根据所述默认字幕语言及所述当前官方语言,对所述文字进行翻译;
    将翻译后的文字进行字幕播放。
  6. 如权利要求5所述的基于人工智能的电视字幕动态生成方法,其中,所述根据所述默认字幕语言及所述当前官方语言,对所述文字进行翻译,包括:
    将所述当前官方语言与所述默认字幕语言进行比较;
    如果一致,则根据所述当前官方语言将所述文字进行翻译,如果不一致,则根据所述默认字幕语言及所述当前官方语言依次将所述文字进行翻译。
  7. 如权利要求1所述的基于人工智能的电视字幕动态生成方法,其中,所述根据所述字幕将当前语音切换成所述当前官方语言进行语音播放,包括:
    获取语音数据的当前播放语言,并将所述当前播放语言与所述当前官方语言进行比较;
    如果不一致,将所述字幕通过语音合成模型转换成所述当前官方语言的语音,并将所述当前播放语言的语音替换成所述当前官方语言的语音进行语音播放。
  8. 一种基于人工智能的电视字幕动态生成装置,所述装置包括:
    观众图像获取模块:设置为获取当前观众的画面图像,将所述画面图像输 入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;
    字幕语言确认模块:设置为获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;
    字幕播放模块:设置为获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;
    语音播放模块:设置为根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
  9. 根据权利要求8所述的基于人工智能的电视字幕动态生成装置,其中,所述观众图像获取模块,包括:
    预置单元:设置为预设当前观众的画面图像的拍摄周期;
    第一获取单元:设置为根据所述当前观众的画面图像的拍摄周期获取当前观众的画面图像。
  10. 根据权利要求8所述的基于人工智能的电视字幕动态生成装置,其中,所述字幕语言确认模块,包括:
    存储单元:设置为预设人种及定位信息与官方语言及与所述官方语言对应的人口比例的对应关系,并将所述对应关系存储在数据库中;
    查询单元:设置为根据所述获取的当前定位信息、当前观众的人种信息在所述数据库中查询,获得与所述当前定位信息及当前观众的人种对应的官方语言及与所述官方语言对应的人口比例;
    第二获取单元:设置为根据所述与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比获得当前官方语言。
  11. 根据权利要求10所述的基于人工智能的电视字幕动态生成装置,其中,所述字幕语言确认模块,包括:
    计算单元:设置为根据公式:
    W=∑ iX iY i
    获得每个官方语言的置信概率,其中,W为官方语言的置信概率,X i为第i个人种使用所述官方语言的人口比例,Y i为第i个人种占当前观众总人数的人数占比;
    第三获取单元:设置为在所述官方语言中,查询最高置信概率对应的官方语言,并将所述最高置信概率对应的官方语言作为当前官方语言。
  12. 根据权利要求8所述的基于人工智能的电视字幕动态生成装置,其中,所述字幕播放模块,包括:
    接收单元:设置为获取语音数据,以及与所述语音数据对应的预设的默认字幕语言;
    转换单元:设置为将所述语音数据通过语音识别模型转换成文字,并根据所述默认字幕语言及所述当前官方语言,对所述文字进行翻译;
    播放单元:设置为将翻译后的文字进行字幕播放。
  13. 根据权利要求12所述的基于人工智能的电视字幕动态生成装置,其中,所述字幕播放模块,包括:
    第一比较单元:设置为将所述当前官方语言与所述默认字幕语言进行比较;
    第一输出单元:设置为如果一致,则根据所述当前官方语言将所述文字进行翻译,如果不一致,则根据所述默认字幕语言及所述当前官方语言依次将所述文字进行翻译。
  14. 根据权利要求8所述的基于人工智能的电视字幕动态生成装置,其中,所述语音播放模块,包括:
    第二比较单元:设置为获取语音数据的当前播放语言,并将所述当前播放语言与所述当前官方语言进行比较;
    第二输出单元:设置为如果不一致,将所述字幕通过语音合成模型转换成所述当前官方语言的语音,并将所述当前播放语言的语音替换成所述当前官方语言的语音进行语音播放。
  15. 一种计算机设备,所述计算机设备包括存储器和处理器,所述存储器中存储有计算机可读指令,所述计算机可读指令被一个或多个所述处理器执行时,使得一个或多个所述处理器执行以下步骤:
    获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;
    获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;
    获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;
    根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
  16. 根据权利要求15所述的计算机设备,其中,所述根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言时,使得所述处理器执行以下步骤:
    预设人种及定位信息与官方语言及与所述官方语言对应的人口比例的对应关系,并将所述对应关系存储在数据库中;
    根据所述获取的当前定位信息、当前观众的人种信息在所述数据库中查询,获得与所述当前定位信息及当前观众的人种对应的官方语言及与所述官方语言对应的人口比例;
    根据所述与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比获得当前官方语言。
  17. 根据权利要求15所述的计算机设备,其中,所述根据所述字幕将当前语音切换成所述当前官方语言进行语音播放时,使得所述处理器执行以下步骤:
    获取语音数据的当前播放语言,并将所述当前播放语言与所述当前官方语言进行比较;
    如果不一致,将所述字幕通过语音合成模型转换成所述当前官方语言的语音,并将所述当前播放语言的语音替换成所述当前官方语言的语音进行语音播放。
  18. 一种存储介质,所述存储介质可被处理器读写,所述存储介质存储有计算机指令,所述计算机可读指令被一个或多个处理器执行时,使得一个或多个处理器执行以下步骤:
    获取当前观众的画面图像,将所述画面图像输入人脸识别模型,以获取当前观众的人种信息及与所述当前观众的人种对应的人数占比;
    获取当前定位信息,根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言;
    获取语音数据,将所述语音数据通过语音识别模型转换成文字,根据所述当前官方语言将所述文字进行翻译,并在电视屏幕上将所述翻译后的文字进行字幕播放;
    根据所述字幕将当前语音切换成所述当前官方语言进行语音播放。
  19. 根据权利要求18所述的存储介质,其中,所述根据所述当前定位信息、当前观众的人种信息及与所述当前观众的人种对应的人数占比获得当前官方语言时,使得一个或多个所述处理器执行以下步骤:
    预设人种及定位信息与官方语言及与所述官方语言对应的人口比例的对应关系,并将所述对应关系存储在数据库中;
    根据所述获取的当前定位信息、当前观众的人种信息在所述数据库中查询,获得与所述当前定位信息及当前观众的人种对应的官方语言及与所述官方语言对应的人口比例;
    根据所述与所述官方语言对应的人口比例及与所述当前观众的人种对应的人数占比获得当前官方语言。
  20. 根据权利要求18所述的存储介质,其中,所述根据所述字幕将当前语音切换成所述当前官方语言进行语音播放时,使得一个或多个所述处理器执行以下步骤:
    获取语音数据的当前播放语言,并将所述当前播放语言与所述当前官方语言进行比较;
    如果不一致,将所述字幕通过语音合成模型转换成所述当前官方语言的语音,并将所述当前播放语言的语音替换成所述当前官方语言的语音进行语音播放。
PCT/CN2019/117013 2019-01-17 2019-11-11 基于人工智能的电视字幕动态生成方法及相关设备 Ceased WO2020147394A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910042720.7 2019-01-17
CN201910042720.7A CN109905756B (zh) 2019-01-17 2019-01-17 基于人工智能的电视字幕动态生成方法及相关设备

Publications (1)

Publication Number Publication Date
WO2020147394A1 true WO2020147394A1 (zh) 2020-07-23

Family

ID=66943877

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/117013 Ceased WO2020147394A1 (zh) 2019-01-17 2019-11-11 基于人工智能的电视字幕动态生成方法及相关设备

Country Status (2)

Country Link
CN (1) CN109905756B (zh)
WO (1) WO2020147394A1 (zh)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109905756B (zh) * 2019-01-17 2021-11-12 平安科技(深圳)有限公司 基于人工智能的电视字幕动态生成方法及相关设备
CN110519620A (zh) * 2019-08-30 2019-11-29 三星电子(中国)研发中心 在电视机推荐电视节目的方法以及电视机
CN116996748A (zh) * 2023-05-24 2023-11-03 深圳创维-Rgb电子有限公司 电视设置调整方法、电视设置调整设备及可读存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105049950A (zh) * 2014-04-16 2015-11-11 索尼公司 显示信息的方法和系统
US20160021334A1 (en) * 2013-03-11 2016-01-21 Video Dubber Ltd. Method, Apparatus and System For Regenerating Voice Intonation In Automatically Dubbed Videos
CN105843393A (zh) * 2016-03-30 2016-08-10 苏州合欣美电子科技有限公司 一种自适应字幕调整的影音播放器
CN107484034A (zh) * 2017-07-18 2017-12-15 深圳Tcl新技术有限公司 字幕显示方法、终端及计算机可读存储介质
CN109905756A (zh) * 2019-01-17 2019-06-18 平安科技(深圳)有限公司 基于人工智能的电视字幕动态生成方法及相关设备

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20060063575A1 (en) * 2003-03-10 2006-03-23 Cyberscan Technology, Inc. Dynamic theming of a gaming system
US20110191246A1 (en) * 2010-01-29 2011-08-04 Brandstetter Jeffrey D Systems and Methods Enabling Marketing and Distribution of Media Content by Content Creators and Content Providers
CN102547456B (zh) * 2011-12-21 2014-08-20 深圳市九洲电器有限公司 生成多国家多卫星节目文件方法、文件载入方法和机顶盒
US20140337901A1 (en) * 2013-05-07 2014-11-13 Ericsson Television Inc. Network personal video recorder system, method and associated subscriber device
JP5801985B1 (ja) * 2015-01-30 2015-10-28 楽天株式会社 情報処理装置、情報処理方法、プログラム
CN106340294A (zh) * 2016-09-29 2017-01-18 安徽声讯信息技术有限公司 基于同步翻译的新闻直播字幕在线制作系统

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20160021334A1 (en) * 2013-03-11 2016-01-21 Video Dubber Ltd. Method, Apparatus and System For Regenerating Voice Intonation In Automatically Dubbed Videos
CN105049950A (zh) * 2014-04-16 2015-11-11 索尼公司 显示信息的方法和系统
CN105843393A (zh) * 2016-03-30 2016-08-10 苏州合欣美电子科技有限公司 一种自适应字幕调整的影音播放器
CN107484034A (zh) * 2017-07-18 2017-12-15 深圳Tcl新技术有限公司 字幕显示方法、终端及计算机可读存储介质
CN109905756A (zh) * 2019-01-17 2019-06-18 平安科技(深圳)有限公司 基于人工智能的电视字幕动态生成方法及相关设备

Also Published As

Publication number Publication date
CN109905756A (zh) 2019-06-18
CN109905756B (zh) 2021-11-12

Similar Documents

Publication Publication Date Title
US8121849B2 (en) Content filtering for a digital audio signal
CN112601101B (zh) 一种字幕显示方法、装置、电子设备及存储介质
WO2020147394A1 (zh) 基于人工智能的电视字幕动态生成方法及相关设备
US12198700B2 (en) Media system with closed-captioning data and/or subtitle data generation features
JP2012109901A (ja) 資料提示装置
CN110933485A (zh) 一种视频字幕生成方法、系统、装置和存储介质
KR101582574B1 (ko) 실시간 번역을 통한 디지털 방송의 다국어 자막 제공 서비스 장치 및 방법
CN111479124A (zh) 一种实时播放方法和装置
CN207854084U (zh) 一种字幕显示系统
Remael et al. Real-time subtitling in Flanders: Needs and teaching
US12177537B2 (en) Program production apparatus, program production method, and recording medium
Mikul Audio description background paper
Szarkowska Accessibility to the media by hearing impaired audiences in Poland: problems, paradoxes, perspectives.
CN115412702B (zh) 一种会议终端与电视墙一体化设备及系统
US20240348885A1 (en) System and method for question answering
Homma et al. New real-time closed-captioning system for Japanese broadcast news programs
Nicolae Subtitling for the deaf and hard-of-hearing audience in Romania
Costa-Montenegro et al. SubTitleMe, subtitles in cinemas in mobile devices
KR20160044372A (ko) 시각장애인용 화면해설 영상물 시스템
Vera Translating audio description scripts: the way forward? Tentative first stage project results
Ćitić Accessibility of TV programs to persons with disabilities
CN212115549U (zh) 字幕机顶盒
US20150179228A1 (en) Synchronized movie summary
HEILKE SUBTITLING FOR THE DEAF
Heilke Subtitling for the Deaf and Hard of Hearing: the Situation in Germany, the UK and Spain

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19909694

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 18.11.2021)

122 Ep: pct application non-entry in european phase

Ref document number: 19909694

Country of ref document: EP

Kind code of ref document: A1