EP4487319A1 - Circuitry and method for visual speech processing - Google Patents

Circuitry and method for visual speech processing

Info

Publication number
EP4487319A1
EP4487319A1 EP23705432.5A EP23705432A EP4487319A1 EP 4487319 A1 EP4487319 A1 EP 4487319A1 EP 23705432 A EP23705432 A EP 23705432A EP 4487319 A1 EP4487319 A1 EP 4487319A1
Authority
EP
European Patent Office
Prior art keywords
words
expressed
user
expressed words
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23705432.5A
Other languages
German (de)
French (fr)
Inventor
Marc Osswald
Christian Peter BRÄNDLI
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sony Advanced Visual Sensing AG
Sony Semiconductor Solutions Corp
Original Assignee
Sony Advanced Visual Sensing AG
Sony Semiconductor Solutions Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sony Advanced Visual Sensing AG, Sony Semiconductor Solutions Corp filed Critical Sony Advanced Visual Sensing AG
Publication of EP4487319A1 publication Critical patent/EP4487319A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/24Speech recognition using non-acoustical features
    • G10L15/25Speech recognition using non-acoustical features using position of the lips, movement of the lips or face analysis
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/06Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • G10L15/065Adaptation
    • G10L15/07Adaptation to the speaker
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/16Speech classification or search using artificial neural networks
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/02Feature extraction for speech recognition; Selection of recognition unit
    • G10L2015/027Syllables being the recognition units
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • G10L2015/221Announcement of recognition results

Definitions

  • the present disclosure generally pertains to a circuitry and a method, in particular to a circuitry and a method for visual speech processing.
  • the disclosure provides a circuitry for visual speech processing, configured to determine expressed words based on visual data of a user; present the determined expressed words to the user; and adjust the determination of expressed words based on feedback of the user related to the determined expressed words.
  • the disclosure provides a method for visual speech processing that includes determining expressed words based on visual data of a user; presenting the determined expressed words to the user; and adjusting the determination of expressed words based on feedback of the user related to the determined expressed words.
  • Fig. 1 illustrates a block diagram of a device with a circuitry for visual speech processing and receiving user feedback according to an embodiment
  • Fig. 2a illustrates a block diagram of a presentation unit for displaying determined expressed words as text according to an embodiment
  • Fig. 2b illustrates a block diagram of a presentation unit for displaying at least one icon corresponding to determined expressed words according to an embodiment
  • Fig. 2c illustrates a block diagram of a presentation unit for outputting an artificial voice corresponding to determined expressed words according to an embodiment
  • Fig. 3 illustrates a flow diagram of a method for visual speech processing and receiving user feedback according to an embodiment
  • Fig. 4a illustrates a block diagram of a presentation for displaying determined expressed words as text according to an embodiment
  • Fig. 4b illustrates a block diagram of a presentation for displaying at least one icon corresponding to determined expressed words according to an embodiment
  • Fig. 4c illustrates a block diagram of a presentation for outputting an artificial voice corresponding to determined expressed words according to an embodiment
  • Fig. 6a illustrates a block diagram of a lip movement identification unit for identifying a lip movement based on generating a full image of a mouth region of a user according to an embodiment
  • Fig. 6b illustrates a block diagram of a lip movement identification unit for identifying a lip movement based on motion vectors according to an embodiment
  • Fig. 7a illustrates a block diagram of an output unit for displaying determined expressed words as text according to an embodiment
  • Fig. 7c illustrates a block diagram of an output unit for outputting an artificial voice corresponding to determined expressed words according to an embodiment
  • Fig. 7f illustrates a block diagram of an output unit for amplifying a voice according to an embodiment
  • Fig. 8 illustrates a flow diagram of a method for visual speech processing based on event-based visual data according to an embodiment
  • Fig. 9a illustrates a block diagram of a lip movement identification based on generating a full image according to an embodiment
  • Fig. 10a illustrates a block diagram of an output for displaying determined expressed words as text according to an embodiment
  • Fig. 10b illustrates a block diagram of an output for displaying at least one icon corresponding to determined expressed words according to an embodiment
  • Fig. lOd illustrates a block diagram of an output for transmitting determined expressed words to a communication device of another user according to an embodiment
  • Fig. lOe illustrates a block diagram of an output for reducing noise according to an embodiment
  • Fig. lOf illustrates a block diagram of an output for amplifying a voice according to an embodiment
  • Fig. 11 illustrates a block diagram of a general application example according to an embodiment
  • Fig. 12 illustrates a block diagram of an exemplary application for speech enablement for mutes according to an embodiment
  • Fig. 13a illustrates a block diagram of an algorithm for an expressed words determination based on generating an image according to an embodiment
  • Fig. 13b illustrates a block diagram of an algorithm for an expressed words determination based on converting motion to words according to an embodiment
  • Fig. 13c illustrates a block diagram of an algorithm for an expressed words determination based on directly converting events to words according to an embodiment
  • Fig. 14 illustrates an application example of a speech enablement for mutes based on a smartphone according to an embodiment
  • Fig. 15 illustrates an exemplary application of a video call enhancement according to an embodiment
  • Fig. 16a illustrates a block diagram of an algorithm for video call voice enhancement based on fusion of expressed words determined by visual data and expressed words determined by audio data according to an embodiment
  • Fig. 16b illustrates a block diagram of an algorithm for video call voice enhancement based on fusion of visual data and audio data before determining expressed words according to an embodiment
  • Fig. 16c illustrates a block diagram of an algorithm for video call voice enhancement based on removing a noise of an environment according to an embodiment.
  • visual speech enablement can enable speech in a scenario where a speaker cannot be understood because of a loud environment or because the speaker is mute.
  • Some embodiments use a camera that reads the lips of a speaker and can therefore be also used in scenarios where audio-based enhancement methods, such as background noise cancelling or voice amplification, do not work.
  • the circuitry may include any entity capable of processing data, such as a microprocessor, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Tensor Processing Unit (TPU), a Reduced Instruction Set Computer (RISC), a Complex Instruction Set Computer (CISC), an Application-Specific Integrated Circuit (ASIC) and/or a Field-Programmable Gate Array (FPGA).
  • the circuitry may be included on a device of the user whose expressed words are determined, on a device of another user, on a remote data processing apparatus such as a server, and/or may be distributed across several devices and/or data processing apparatus.
  • the visual data may include any data obtained by an optical detector such as a camera.
  • the visual data may be based on a color (e.g. Red-Green-Blue, RGB) image obtained from a color (e.g. RGB) camera such as a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor, on a grayscale image obtained from a grayscale camera such as a CMOS or CCD image sensor, on a thermal image obtained from an infrared camera, on a depth map image obtained from a time-of-flight (ToF) sensor, and/or on an event obtained from an event-based vision sensor.
  • CMOS complementary metal-oxide-semiconductor
  • CCD charge-coupled device
  • the visual data may include a single image frame or a sequence of image frames.
  • the image frame(s) may be acquired with a short exposure such as to obtain a sharp image or with a long exposure such as to obtain an image with a motion blur that indicates a movement.
  • the visual data may include events obtained asynchronously or at a high detection rate from an event-based vision sensor, wherein each event may indicate a change detected at an associated position on the sensor.
  • the expressed words may be words audibly uttered by the user or expressed by the user by moving his lips (and mouth) without generating an audible sound. I.e., the expressed words may be determined regardless of any sound generated by the user.
  • the expressed words may be unconnected or may be included in an expression or in one or more sentences.
  • the expressed words may be words of one or more predetermined languages, e.g., a language in which a user interface of the circuitry is configured or a language of a region in which the user is located, or the language of the words may be determined during or after determining the expressed words.
  • the circuitry may be configured to present the determined expressed words to the user in any suitable form.
  • the determined expressed words may be presented to the user as a text rendered on a display, as one or more images or icons corresponding to the determined expressed words, as an artificial voice audibly output by a loudspeaker, or as tactile or electric stimulus.
  • the feedback of the user may be received via a human-machine interface.
  • the human-machine interface may include a key, a button, a touch pad, a touch screen, a microphone for detecting a voice input of the user or a camera for detecting gesture input of the user.
  • the user may input the feedback in the human-machine interface in any suitable form, and the circuitry may obtain the feedback from the human-machine interface.
  • the circuitry may be configured to extract the feedback of the user from the determined expressed words.
  • the feedback of the user may be related to the determined expressed words by indicating a correction of the expressed words or a confirmation of the expressed words.
  • the circuitry may or may not be configured to request the feedback from the user.
  • the adjusting of the determination of expressed words may serve the purpose of increasing an accuracy of the determination of expressed words.
  • the determination of expressed words may be adjusted such that a probability for determining, based on future visual data that are similar to the visual data, expressed words that are indicated by the feedback of the user is increased.
  • the determination of expressed words includes detecting, in the visual data, a mouth region of the user; and identifying a lip movement in the mouth region of the user.
  • the visual data may include more than the mouth region of the user (e.g., a whole face, a whole body and/or surroundings of the user), and the detection of the mouth region may be based on machine learning such as on an artificial neural network, on a support vector machine, on a wavelet transform and/or on any other suitable technique.
  • the detection of the mouth region may include detecting a region-of-interest (ROI) in the visual data that corresponds to the mouth region.
  • ROI region-of-interest
  • the lip movement may be identified based on a motion blur of the visual data, on a sequence of image frames of the visual data, and/or on events of the visual data.
  • the determination of expressed words includes determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
  • the syllable data may indicate syllables or at least parts of syllables, e.g., that match the identified lip movement.
  • the expressed words may be generated based on the syllables or parts of syllables indicated by the syllable data.
  • the generating of the expressed words is based on a language model.
  • the language model may be specific for a predetermined language in which the expressed words should be generated and may represent language-specific rules including grammatical rules.
  • the language model may be implemented as artificial neural network trained towards a predetermined language.
  • the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data.
  • the generating of the expressed words may include selecting one or more words of the word list that correspond to the syllables or part of syllables indicated by the syllable data.
  • the language model may select the one or more words based on a list of syllables used by a language and/or on syllable probabilities associated with a language.
  • the language model may include the list of syllables and/or the syllable probabilities.
  • the selected word(s) may be selected and/or inflected according to grammatical rules represented by the language model.
  • the word(s) may be selected according to a probability based on a text corpus or on a history of generated words associated with a user.
  • the determination of expressed words includes detecting, in the visual data, a face expression of the user; and generating the expressed words based on the detected face expression.
  • the face expression may include an emotional expression (e.g. angry, happy etc.), the corresponding emotion may be detected, and the expressed words may be generated in accordance with the detected emotion. For example, words from the word list of the language model may be selected, as the determined expressed words, that correspond to the emotion.
  • an emotional expression e.g. angry, happy etc.
  • the expressed words may be generated in accordance with the detected emotion. For example, words from the word list of the language model may be selected, as the determined expressed words, that correspond to the emotion.
  • the face expression may include a movement of a lip, of an eye, of a forehead, of an eyebrow, of a nose and/or of a cheek of the user.
  • the face expression may contribute to a sign language expressed consciously or unconsciously by the user, and determining words expressed by the user based on the face expression may increase a performance of the determination.
  • the determination of expressed words includes obtaining audio data from a microphone; and generating the expressed words based on the audio data.
  • the audio data may represent a sound generated by the user when expressing the words
  • the expressed words may be determined based on both the visual data and the obtained audio data to increase a robustness of the determination of expressed words.
  • the determination of expressed words may be based on the audio data if the audio data and the visual data are correlated (e.g., if a correlation between the audio data and the visual data exceeds a predetermined threshold), and may not be based on the audio data if the audio data and the visual data are not correlated (e.g., if a correlation between the audio data and the visual data does not exceed a predetermined threshold).
  • the determination may be based on the audio data during a training, when the circuitry is trained with visual data of, e.g., a certain user.
  • the determination of expressed words may be based on audio data that correspond to a pronunciation in accordance with the language model, and may also be based on audio data that do not correspond to the language model. For example, even a mute may generate a sound when expressing words by lip movements. Such a sound may aid a determination of words expressed by the mute.
  • Determining expressed words based on audio data in addition to visual data may provide a performance advantage over determining expressed words purely based on visual data.
  • the presenting of the determined expressed words includes displaying the determined expressed words as text.
  • the determined expressed words may be rendered as text and displayed by a display means of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
  • a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
  • the presenting of the determined expressed words includes displaying at least one icon corresponding to the determined expressed words.
  • the at least one icon may represent a meaning of the determined expressed words.
  • the at least one icon may be retrieved from a predetermined database or may be retrieved by a web search upon determining the expressed words.
  • the presenting of the determined expressed words includes generating an artificial voice that utters the determined expressed words; and outputting the generated artificial voice.
  • the artificial voice may be output by loudspeakers of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
  • a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
  • the visual data includes a plurality of events obtained from an eventbased vision sensor.
  • an event included in the visual data may indicate a brightness change detected by a photosensitive element of the event-based vision sensor.
  • the event-based vision sensor may detect events asynchronously or at a high detection rate such as 1,000,000 events per second, without limiting the disclosure to this detection rate.
  • the adjusting of the determination of expressed words includes changing a value of a parameter of the determination of expressed words.
  • the parameter may be adjusted such as to increase a probability that expressed words determined based on future visual data representing a future lip movement that is similar to the lip movement identified in the visual data correspond to the feedback of the user.
  • the determination of expressed words is based on an artificial neural network.
  • the artificial neural network may be trained to detect, in the visual data, a mouth region of the user and/or a face expression of the user.
  • the artificial neural network may be trained to identify a lip movement of the user, to associate a lip movement with one or more syllables and to determine syllable data based on the identified lip movement and/or based on the one or more associated syllables.
  • the artificial neural network may be trained to determine the expressed words by comparing syllables indicated by the syllable data with the syllables, the syllable probabilities, the word list and/or the word probabilities of the language model and/or applying the grammatical rules of the language model.
  • the artificial neural network may be trained based on gradient descent, backpropagation and/or reinforcement learning on a database of visual data that represent lip movements and/or face expressions of users and of associated expressed words. Aspects such as detecting a mouth region, detecting a face expression, identifying a lip movement, associating a lip movement with a syllable and a language model or parts thereof may be trained explicitly in separate training steps or may be trained implicitly during the same training procedure.
  • the adjusting of the parameter may include adjusting, as the parameter, a weight value of the artificial neural network.
  • the adjusting may be based on gradient descent, backpropagation and/or reinforcement learning.
  • the feedback of the user indicates at least one of a label, a loss and a fitness for the artificial neural network.
  • the label may include one or more words that correspond to the visual data.
  • the loss and/or the fitness may include a value that indicates an accuracy of the determined expressed word.
  • the value may indicate an absolute accuracy.
  • the value may indicate a relative accuracy that relates to a previous determination of expressed words.
  • the feedback of the user includes a correction of at least one of the determined expressed words; and the adjusting of the determination of expressed words is based on the correction of the at least one of the determined expressed words.
  • the feedback may indicate one or more words that should be determined as the expressed words based on the visual data
  • the adjusting of the parameter may include adjusting the parameter such as to increase a probability that the one or more words indicated by the feedback are determined as the expressed words.
  • the presenting of the determined expressed words includes recommending an improved lip movement.
  • the improved lip movement may allow a more robust determination of syllable data and/or of expressed words than the lip movement represented by the visual data.
  • the improved lip movement may be recommended to the user on a display means of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses, and may be displayed by an application that is executed on the user terminal.
  • Some embodiments pertain to a method for visual speech processing that includes determining expressed words based on visual data of a user; presenting the determined expressed words to the user; and adjusting the determination of expressed words based on feedback of the user related to the determined expressed words.
  • the method may be configured corresponding to the processing performed by the circuitry described above, and all features described with reference to the processing performed by the circuitry may also be features of the method.
  • Some embodiments pertain to a circuitry for visual speech processing that is configured to obtain event-based visual data of a user from an event-based vision sensor; determine expressed words based on the obtained event-based visual data; and output the determined expressed words.
  • the circuitry may include any entity capable of processing data, such as a microprocessor, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Tensor Processing Unit (TPU), a Reduced Instruction Set Computer (RISC), a Complex Instruction Set Computer (CISC), an Application-Specific Integrated Circuit (ASIC) and/or a Field-Programmable Gate Array (FPGA).
  • the circuitry may be included on a device of the user whose expressed words are determined, on a device of another user, on a remote data processing apparatus such as a server, and/or may be distributed across several devices and/or data processing apparatus.
  • the event-based visual data may represent events obtained from an event-based vision sensor.
  • the event-based vision sensor may capture the events asynchronously or at a high detection rate, wherein each event may indicate a change detected at an associated position on the sensor.
  • an event included in the visual data may indicate a brightness change detected by a photosensitive element of the event-based vision sensor.
  • the event-based vision sensor may detect events asynchronously or at a high detection rate such as 1,000,000 events per second, without limiting the disclosure to this detection rate.
  • the event-based vision sensor may include an event camera, a neuromorphic camera, a silicon retina and/or a dynamic vision sensor.
  • lips of the user may move very fast, and the event-based vision sensor may provide a temporal resolution that is high enough for resolving the lip movement.
  • the event-based vision sensor may provide a more efficient processing of the event-based visual data obtained with a high data rate because redundant data are naturally compressed in the eventbased visual data as compared to data from a camera that generates whole image frames instead of events.
  • the event-based visual data may indicate a lip movement of a user.
  • the event-based vision sensor may capture at least lips of a user that is expressing words, e.g. by speaking.
  • the expressed words may be words audibly uttered by the user or expressed by the user by moving his lips (and mouth) without generating an audible sound. I.e., the expressed words may be determined regardless of any sound generated by the user.
  • the expressed words may be unconnected or may be included in an expression or in one or more sentences.
  • the expressed words may be words of one or more predetermined languages, e.g. a language in which a user interface of the circuitry is configured or a language of a region in which the user is located, or the language of the words may be determined during or after determining the expressed words.
  • the expressed words may be determined based on the event-based visual data by deriving the expressed words at least from a lip movement indicated by the visual data.
  • the determination of expressed words includes identifying, in the eventbased visual data, a lip movement of the user.
  • the lip movement may be captured when the user is expressing words, e.g. by speaking.
  • the lip movement may correspond to a sound generated by the user for expressing words.
  • the identifying of the lip movement includes generating, based on the event-based visual data, at least one full image of a mouth region of the user; and identifying the lip movement based on the generated at least one full image of the mouth region of the user.
  • the mouth region of the user may be identified in the event-based visual data based on a shape of an area with an increased event density, and the at least one full image of the mouth region of the user may be reconstructed from the events represented by the event-based visual data, e.g. based on at least one of temporal smoothing, optimization, gradient estimation and Poisson integration.
  • the at least one full image of the mouth region may include an image in which a motion is represented by motion blur, and the lip movement is identified based on the motion blur.
  • the at least one full image of the mouth region may include a plurality of images of the mouth region, and the lip movement may be detected based on a difference between subsequent images of the plurality of images of the mouth region.
  • the identifying of the lip movement includes identifying, based on the event-based visual data, events of a mouth region of the user; determining motion vectors based on the identified events of the mouth region; and identifying the lip movement based on the determined motion vectors.
  • the events of the mouth region of the user may be events in the event-based visual data that are detected in the mouth region of the user, and the motion vectors may be determined based on correlations between events of the event of the mouth region.
  • the determination of expressed words includes determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
  • the syllable data may indicate syllables or at least parts of syllables, e.g., that match the identified lip movement.
  • the expressed words may be generated based on the syllables or parts of syllables indicated by the syllable data.
  • the determining of the syllable data includes identifying, based on the identified lip movement, at least one viseme; and determining the syllable data based on the at least one viseme.
  • a viseme may indicate a lip shape or transition between lip shapes that is characteristic for one or more specific sounds generated by a user. For example, lips of a user may assume the lip shape or transition of lip shapes indicated by a viseme when the user generates the specific sound or one of the specific sounds. Thus, based on a viseme, a sound generated by the user may be identified or at least narrowed down to a (small) number of sounds. The sound(s) identified by a viseme may be predetermined by human physiology.
  • a syllable may include a predetermined sound or a predetermined sequence of sounds.
  • a viseme or sequence of visemes may indicate a syllable or several alternative syllables indicated by the viseme, or may indicate a sequence of syllables or several alternative sequences of syllables indicated by the sequence of visemes.
  • the generating of the expressed words is based on a language model.
  • the language model may be specific for a predetermined language in which the expressed words should be generated and may represent language-specific rules including grammatical rules.
  • the language model may be implemented as artificial neural network trained towards a predetermined language.
  • the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data.
  • the generating of the expressed words may include selecting one or more words of the word list that correspond to the syllables or part of syllables indicated by the syllable data.
  • the selected word(s) may be selected and/or inflected according to grammatical rules represented by the language model.
  • the word(s) may be selected according to a probability based on a text corpus or on a history of generated words associated with a user.
  • the outputting of the determined expressed words includes displaying the determined expressed words as text.
  • the determined expressed words may be rendered as text and displayed by a display means of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
  • a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
  • the outputting of the determined expressed words includes displaying at least one icon corresponding to the determined expressed words.
  • the at least one icon may represent a meaning of the determined expressed words.
  • the at least one icon may be retrieved from a predetermined database or may be retrieved by a web search upon determining the expressed words.
  • the outputting of the determined expressed words includes generating an artificial voice that utters the determined expressed words; and outputting the generated artificial voice.
  • the artificial voice may be output by loudspeakers of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
  • a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
  • the outputting of the determined expressed words includes transmitting the determined expressed words to communication device of another user.
  • the determined expressed words of the user may be transmitted to a communication device of another user during a phone call or a video call between the user and the other user.
  • the communication device of the other user may output the expressed words in a manner perceivable for the other user.
  • the determination of expressed words is based on an artificial neural network.
  • the artificial neural network may include a deep neural network, a convolution neural network and/or a recurrent neural network, without limiting the disclosure to a specific type of an artificial neural network.
  • the artificial neural network may be trained to identify, in the event-based visual data, a lip movement of the user.
  • the artificial neural network may be trained to generate, based on the event-based visual data, at least one full image of a mouth region of the user and identify the lip movement based on the at least one full image of the mouth region of the user.
  • the artificial neural network may be trained to identify, based on the event-based visual data, events of a mouth region of the user, determine motion vectors based on the identified events of the mouth region of the user and identify the lip movement based on the determined motion vectors.
  • the artificial neural network may be trained to identify at least one viseme based on the lip movement, to associate the lip movement and/or the at least one viseme with one or more syllables and to determine syllable data based on the identified lip movement, based on the at least one viseme and/or based on the one or more associated syllables.
  • the artificial neural network may be trained to include a language model according to a language.
  • the artificial neural network may be trained to include, as the language model, syllables used by the language, syllable probabilities of the language, a word list associated with the language, word probabilities of words from the word list and/or grammatical rules associated with the language.
  • the artificial neural network may be trained to determine the expressed words by comparing syllables indicated by the syllable data with the syllables, the syllable probabilities, the word list and/or the word probabilities of the language model and/or applying the grammatical rules of the language model.
  • the artificial neural network may be trained based on gradient descent, backpropagation and/or reinforcement learning on a database of visual data that represent lip movements users and of associated expressed words. Aspects such as generating a full image of a mouth region, identifying events of a mouth region, determining motion vectors, identifying a lip movement, identifying a viseme, associating a lip movement and/or a viseme with a syllable and a language model or parts thereof may be trained explicitly in separate training steps or may be trained implicitly during the same training procedure.
  • the determination of expressed words is further based on audio data obtained from a microphone.
  • the audio data may represent a sound generated by the user when expressing the words
  • the expressed words may be determined based on both the event-based visual data and the obtained audio data to increase a robustness of the determination of expressed words.
  • the determination of expressed words may be based on the audio data if the audio data and the event-based visual data are correlated (e.g., if a correlation between the audio data and the eventbased visual data exceeds a predetermined threshold), and may not be based on the audio data if the audio data and the event-based visual data are not correlated (e.g., if a correlation between the audio data and the event-based visual data does not exceed a predetermined threshold).
  • the determination may be based on the audio data during a training, when the circuitry is trained with event-based visual data of, e.g., a certain user.
  • the determination of expressed words may be based on audio data that correspond to a pronunciation in accordance with the language model, and may also be based on audio data that do not correspond to the language model.
  • a mute may generate a sound when expressing words by lip movements. Such a sound may aid a determination of words expressed by the mute.
  • Determining expressed words based on audio data in addition to event-based visual data may provide a performance advantage over determining expressed words purely based on event-based visual data.
  • the determining of expressed words based on audio data in addition to event-based visual data may be performed by an artificial neural network trained accordingly.
  • the outputting of the determined expressed words includes reducing, in audio data obtained from a microphone, noise that does not correspond to the determined expressed words; and outputting the audio data in which the noise has been reduced.
  • an artificial voice that pronounces the expressed words may be generated and compared to the audio data, and frequency components of the audio data that are not correlated with the artificial voice may be reduced or cancelled.
  • noise in the audio data may be attenuated or cancelled based on the determined expressed words.
  • Generating the artificial voice recognizing frequency components of the audio data that are correlated or that are not correlated with the artificial voice and/or reducing frequency components of the audio data that are not correlated with the artificial voice may be performed by an artificial neural network trained accordingly, e.g. based on gradient descent, backpropagation and/or reinforcement learning.
  • the outputting of the determined expressed words includes amplifying, in audio data obtained from a microphone, a voice that corresponds to the determined expressed words; and outputting the audio data in which the voice has been amplified.
  • an artificial voice that pronounces the expressed words may be generated and compared to the audio data, and frequency components of the audio data that are correlated with the artificial voice may be amplified.
  • speech in the audio data may be amplified based on the determined expressed words.
  • Generating the artificial voice recognizing frequency components of the audio data that are correlated or that are not correlated with the artificial voice and/or amplifying frequency components of the audio data that are not correlated with the artificial voice may be performed by an artificial neural network trained accordingly, e.g. based on gradient descent, backpropagation and/or reinforcement learning.
  • Some embodiments pertain to a method for visual speech processing that includes obtaining event-based visual data of a user obtained from an event-based vision sensor; determining expressed words based on the obtained event-based visual data; and outputting the determined expressed words.
  • the method may be configured corresponding to the processing performed by the circuitry described above, and all features described with reference to the processing performed by the circuitry may also be features of the method.
  • the methods as described herein are also implemented in some embodiments as a computer program causing a computer and/or a processor to perform the method, when being carried out on the computer and/or processor.
  • a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.
  • Fig. 1 illustrates a block diagram of a device 1 with a circuitry 2 for visual speech processing and receiving user feedback according to an embodiment.
  • the device 1 includes the circuitry 2 for visual speech processing, Central Processing Unit (CPU) 3, a storage unit 4 and an input/output (I/O) unit 5.
  • CPU Central Processing Unit
  • I/O input/output
  • the circuitry 2 includes a determination unit 6, a presentation unit 7 and an adjustment unit 8.
  • the determination unit 6 is configured to determine expressed words expressed by a user based on visual data of the user.
  • the determination unit 6 includes a mouth region detection unit 9, a lip movement identification unit 10, a syllable data determination unit 11, a face expression detection unit 12, an audio data obtaining unit 13 and an expressed words generation unit 14.
  • the determination unit 6 receives visual data of the user via the I/O unit 5.
  • the visual data of the user are acquired with an event-based vision sensor when the user is expressing words and include a plurality of events of a face (including a mouth region) of the user.
  • the determination unit 6 provides the visual data of the user to the mouth region detection unit 9, to the lip movement identification unit 10 and to the face expression detection unit 12.
  • the mouth region detection unit 9 receives the visual data of the user and detects, based on an artificial neural network, a region-of-interest (ROI) in the visual data that corresponds to the mouth region of the user.
  • the mouth region detection unit 9 generates ROI data that indicate the detected ROI and provides the ROI data to the lip movement identification unit 10.
  • ROI region-of-interest
  • the lip movement identification unit 10 receives the visual data of the user and the ROI data.
  • the lip movement identification unit 10 identifies a lip movement in the mouth region of the user, based on motion vectors that are based on correlating events of the visual data within the ROI detected by the mouth region detection unit 9.
  • the lip movement identification unit 10 generates lip movement data that indicate the identified lip movement and provides the lip movement data to the syllable data determination unit 11.
  • the syllable data determination unit 11 receives the lip movement data and determines syllable data based on the identified lip movement indicated by the lip movement data.
  • the syllable data indicate a sequence of syllables that correspond to the identified lip movement.
  • the syllable data also indicate alternative syllables in the sequence of syllables where the syllable data determination unit 11 cannot determine a syllable unambiguously based on the lip movement data.
  • the syllable data determination unit 11 then provides the syllable data to the expressed words generation unit 14.
  • the face expression detection unit 12 receives the visual data of the user and detects, in the visual data, a face expression of the user.
  • the face expression detection unit 12 identifies an emotion represented by the face expression, generates face expression data that indicates the detected face expression of the user and the identified emotion and provides the face expression data to the expressed words generation unit 14.
  • the audio data obtaining unit 13 obtains audio data from a microphone.
  • the audio data are captured by the microphone when the user expresses words.
  • the audio data obtaining unit 13 identifies features in the audio data and generates feature data that indicate the identified features of the audio data.
  • the audio data obtaining unit 13 provides the feature data to the expressed words generation unit 14.
  • the expressed words generation unit 14 receives the syllable data, the face expression data and the feature data.
  • the expressed words generation unit 14 generates expressed words based on an artificial neural network and, for this purpose, provides the syllable data, the face expression data and the feature data to the artificial neural network.
  • the artificial neural network solves ambiguities in the syllable data (i.e., alternative syllables where the syllable data determination unit 11 cannot unambiguously determine a syllable) based on the feature data and on the face expression data.
  • the artificial neural network selects words from a word list of a language model based on the syllable data, on the feature data and on the facial expression data, and inflects the selected words according to grammatical rules of the language model.
  • the expressed words generation unit 14 then outputs the selected (and inflected) words as generated expressed words.
  • the determination unit 6 provides the expressed words generated by the expressed words generation unit 14 as determined expressed words to the presentation unit 7.
  • the presentation unit 7 presents the determined expressed words via the I/O unit 5 to the user and requests feedback from the user regarding the determined expressed words.
  • the presentation unit 7 includes a lip movement recommendation unit 15 that generates a recommendation for an improved lip movement that allows determining the syllable data with less ambiguity than the lip movement identified by the lip movement identification unit 10 based on the visual data.
  • the presentation unit 7 presents the recommendation for the improved lip movement to the user via the I/O unit 5.
  • the adjustment unit 8 receives, via the I/O unit 5, user feedback regarding the determined expressed words, identifies at least one of a label, a loss and a fitness for the artificial neural network of the expressed words generation unit 14 and changes a value of a parameter of the artificial neural network of the expressed words generation unit 14 in accordance with the user feedback.
  • the adjustment unit 8 also identifies in the feedback of the user a correction of at least one of the determined expressed words and changes a value of a parameter of the artificial neural network of the expressed words generation unit 14 in accordance with the user feedback.
  • the adjustment unit 8 adjusts the determination of expressed words based on the user feedback.
  • the CPU 3 controls the device 1 and the circuitry 2.
  • the storage unit 4 includes a memory and a non-volatile storage.
  • the storage unit 4 stores instructions for execution by the CPU 3.
  • the storage unit 4 also stores the word list of the language model and parameters of the artificial neural network of the expressed words generation unit 14.
  • the I/O unit 5 includes a humanmachine interface with presenting means for presenting the determined expressed words to the user and with input means for receiving the feedback from the user.
  • the I/O unit 5 further includes a network interface for communicating with a server in a communication network.
  • Figures 2a, 2b and 2c provide examples of the presentation unit 7 of Fig. 1.
  • Fig. 2a illustrates a block diagram of a presentation unit 7a for displaying the determined expressed words as text according to an embodiment.
  • the presentation unit 7a is an example of the presentation unit 7 of Fig. 1 and includes a text rendering unit 16 that renders the determined expressed words as text.
  • the presentation unit 7a then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 5 to display the rendered text to the user.
  • Fig. 2b illustrates a block diagram of a presentation unit 7b for displaying at least one icon corresponding to the determined expressed words according to an embodiment.
  • the presentation unit 7b is an example of the presentation unit 7 of Fig. 1.
  • the presentation unit 7b retrieves, via the network interface of the I/O unit 5, one or more icons from a server in the communication network. The number and content of the icons retrieved is based on a number and meaning of the determined expressed words and is chosen such that the retrieved icons correspond to the determined expressed words.
  • the presentation unit 7b includes an icon rendering unit 17 that renders the retrieved one or more icons.
  • the presentation unit 7b then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 5 to display the rendered one or more icons to the user.
  • Fig. 2c illustrates a block diagram of a presentation unit 7c for outputting an artificial voice corresponding to the determined expressed words according to an embodiment.
  • the presentation unit 7c is an example of the presentation unit 7 of Fig. 1.
  • the presentation unit 7c includes an artificial voice generation unit 18.
  • the artificial voice generation unit 18 generates an artificial voice that utters the determined expressed words.
  • the presentation unit 7c causes a loudspeaker that is a presenting means of the human-machine interface included in the I/O unit 5 to output the generated artificial voice to the user.
  • the presentation unit 7 of Fig. 1 includes two or all three of the text rendering unit 16 of Fig. 2a, the icon rendering unit 17 of Fig. 2b and the artificial voice generation unit 18 of Fig. 2c.
  • the presentation unit 7 presents, to the user, the determined expressed words as two or all three of a text, one or more icons and an artificial voice.
  • the face expression detection unit 12 is not provided in the determination unit 6 of Fig. 1 and/or the audio data obtaining unit 13 is not provided in the determination unit 6 of Fig. 1 and/or the lip movement recommendation unit 15 is not provided in the presentation unit 7 of Fig. 1.
  • the determination of expressed words does not include detecting, in the visual data, a face expression of the user and generating the expressed words based on the detected face expression and/or the determination of expressed words does not include obtaining audio data from a microphone and generating the expressed words based on the audio data and/or the presenting of the determined expressed words does not include recommending an improved lip movement.
  • the event-based vision sensor that captures the visual data of the user is included in the device 1, and in some embodiments, the event-based vision sensor that captures the visual data of the user is provided separately from the device 1 and the determination unit 6 obtains the visual data from the event-based vision sensor via the network interface included in the I/O unit 5.
  • the event-based vision sensor is only provided as an example. In some embodiments, any suitable camera capable of visual data of a user is used instead of or in addition to the eventbased vision sensor.
  • the human-machine interface with the presenting means and with the input means is not provided in the I/O unit 5 but separately from the device 1. In such a case, the determined expressed words (and a recommendation for an improved lip movement) are transmitted via the network interface of the I/O unit 5 to the presenting means and the feedback of the user is received via the network interface of the I/O unit 5 from the input means.
  • the device 1 that includes the circuitry 2 may be configured as a server and the human-machine interface with the presenting means and with the input means may be provided on a communication device of a user that communicates with the server via a communication network.
  • Fig. 3 illustrates a flow diagram of a method 30 for visual speech processing and receiving user feedback according to an embodiment.
  • the method 30 is an example of a method that is executed by the circuitry 2 of Fig. 1.
  • the method 30 includes, at S31, a determination performed by the determination unit 6 of Fig. 1.
  • the determination at S31 determines expressed words based on visual data of a user that are acquired with an event-based vision sensor when the user is expressing words and that include a plurality of events of a face (including a mouth region) of the user.
  • the determination at S31 includes a mouth region detection at S32, a lip movement identification at S33, a syllable data determination at S34, a face expression detection at S35, obtaining audio data at S36 and an expressed words generation at S37.
  • the mouth region detection at S32 is performed by the mouth region detection unit 9 of Fig. 1 and detects, based on an artificial neural network, a region-of-interest (ROI) in the visual data that corresponds to the mouth region of the user.
  • ROI region-of-interest
  • the lip movement identification at S33 is performed by the lip movement identification unit 10 of Fig. 1 and identifies a lip movement in the mouth region of the user, based on motion vectors that are based on correlating events of the visual data with the ROI detected by the mouth region detection at S32.
  • the syllable data determination at S34 is performed by the syllable data determination unit 11 of Fig. 1 and determines syllable data based on the lip movement identified by the lip movement identification at S33.
  • the syllable data indicate a sequence of syllables that correspond to the identified lip movement as well as alternative syllables in the sequence of syllables where the syllable data determination at S33 cannot determine a syllable unambiguously based on the identified lip movement.
  • the face expression detection at S35 is performed by the face expression detection unit 12 of Fig. 1, detects, in the visual data, a face expression of the user and identifies an emotion represented by the detected face expression.
  • the obtaining of audio data at S36 is performed by the audio data obtaining unit 13 of Fig. 1 and obtains audio data from a microphone.
  • the audio data are captured by the microphone when the user expresses words.
  • the obtaining of audio data further identifies features in the audio data.
  • the expressed words generation at S37 is performed by the expressed words generation unit 24 of Fig. 1 and generates expressed words based on an artificial neural network.
  • the expressed words generation provides the syllable data, the face expression detected at S35 with the emotion identified at S35 and the audio data obtained at S36 with the features identified at S36 to the artificial neural network.
  • the artificial neural network solves ambiguities in the syllable data (i.e., alternative syllables where the syllable data determination at S34 cannot unambiguously determine a syllable) based on the audio data and the features identified therein as well as on the face expression and the emotion identified based thereon.
  • the artificial neural network selects words from a word list of a language model based on the syllable data, on the audio data and the features identified therein as well as on the face expression and the emotion identified based thereon, and inflects the selected words according to grammatical rules of the language model.
  • the words selected (and inflected) by the expressed words generation at S37 are the expressed words determined by the determination at S31.
  • a presentation is performed by the presentation unit 7 of Fig. 1.
  • the presentation presents the determined expressed words via the I/O unit 5 of Fig. 1 to the user and requests feedback from the user regarding the determined expressed words.
  • the presentation includes a lip movement recommendation at S39 that is performed by the lip movement recommendation unit 15 of Fig. 1 and generates a recommendation for an improved lip movement that allows determining the syllable data with less ambiguity than the lip movement identified by the lip movement identification at S33 based on the visual data.
  • the presentation presents the recommendation for the improved lip movement to the user via the VO unit 5 of Fig. 1.
  • an adjustment is performed by the adjustment unit 8 of Fig. 1.
  • the adjustment receives, via the I/O unit 5 of Fig. 1, user feedback regarding the determined expressed words, identifies at least one of a label, a loss and a fitness for the artificial neural network of the expressed words generation of S37, and changes a value of a parameter of the artificial neural network of the expressed words generation of S37 in accordance with the user feedback.
  • the adjustment also identifies in the feedback of the user a correction of at least one of the determined expressed words and changes a value of a parameter of the artificial neural network of the expressed words generation of S37 in accordance with the user feedback.
  • the adjustment adjusts the determination of expressed words based on the user feedback.
  • Figures 4a, 4b and 4c illustrate examples of the presentation at S38 of Fig. 3.
  • Fig. 4a illustrates a block diagram of a presentation at S38a for displaying the determined expressed words as text according to an embodiment.
  • the presentation at S38a is an example of the presentation at S38 of Fig. 3 and is performed by the presentation unit 7a of Fig. 2a.
  • the presentation at S38a includes rendering of text at S41 performed by the text rendering unit 16 of Fig. 2a.
  • the rendering of text at S41 renders the determined expressed words as text.
  • the presentation at S38a then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 5 to display the rendered text to the user.
  • Fig. 4b illustrates a block diagram of a presentation at S38b for displaying at least one icon corresponding to the determined expressed words according to an embodiment.
  • the presentation at S38b is an example of the presentation at S38 of Fig. 3 and is performed by the presentation unit 7b of Fig. 2b.
  • the presentation at S38b retrieves, via the network interface of the I/O unit 5, one or more icons from a server in the communication network. The number and content of the icons retrieved is based on a number and meaning of the determined expressed words and is chosen such that the retrieved icons correspond to the determined expressed words.
  • the presentation at S38b includes rendering of icons at S42 performed by the icon rendering unit 17 of Fig. 2b.
  • the rendering of icons at S42 renders the retrieved one or more icons.
  • the presentation at S38b then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 5 to display the rendered one or more icons to the user.
  • Fig. 4c illustrates a block diagram of a presentation at S38c for outputting an artificial voice corresponding to the determined expressed words according to an embodiment.
  • the presentation at S38c is an example of the presentation S38 of Fig. 3 and is performed by the presentation unit 7c of Fig. 2c.
  • the presentation at S38c includes an artificial voice generation at S43 performed by the artificial voice generation unit 18 of Fig. 2c.
  • the artificial voice generation at S43 generates an artificial voice that utters the determined expressed words.
  • the presentation at S38c then causes a loudspeaker that is a presenting means of the human-machine interface included in the I/O unit 5 to output the generated artificial voice to the user. Note that, in some embodiments, the presentation at S38 of Fig.
  • the presentation at S38 presents, to the user, the determined expressed words as two or all three of a text, one or more icons and an artificial voice.
  • the face expression detection at S35 is not included in the determination at S31 of Fig. 3 and/or the obtaining of audio data at S36 is not included in the determination at S31 of Fig. 3 and/or the lip movement recommendation at S39 is not included in the presentation at S38 of Fig. 3.
  • the determination of expressed words does not include detecting, in the visual data, a face expression of the user and generating the expressed words based on the detected face expression and/or the determination of expressed words does not include obtaining audio data from a microphone and generating the expressed words based on the audio data and/or the presenting of the determined expressed words does not include recommending an improved lip movement.
  • Fig. 5 illustrates a block diagram of a device 100 with a circuitry 102 for visual speech processing based on event-based visual data according to an embodiment.
  • the device 100 includes an event-based vision sensor 101, the circuitry 102, a Central Processing Unit (CPU) 103, a storage unit 104 and an Input/Output (I/O) unit 105.
  • CPU Central Processing Unit
  • I/O Input/Output
  • the event-based vision sensor 101 captures event-based visual data of a user when the user is expressing words by speaking.
  • the event-based visual data indicate events that represent changes in a face (including a mouth region) of the user.
  • the event-based vision sensor 101 provides the event-based visual data to the circuitry 102.
  • the circuitry 102 includes a visual data obtaining unit 106, a determination unit 107 and an output unit 108.
  • the visual data obtaining unit 106 obtains the event-based visual data from the event-based vision sensor 101 and provides the event-based visual data to the determination unit 107.
  • the determination unit 107 determines expressed words based on the obtained event-based visual data.
  • the determination unit 107 includes a lip movement identification unit 109, a syllable data determination unit 110, a face expression detection unit 111, an audio data obtaining unit 112 and an expressed words generation unit 113.
  • the determination unit 107 receives the event-based visual data form the visual data obtaining unit 106 and provides the event-based visual data to the lip movement identification unit 109.
  • the lip movement identification unit 109 receives the event-based visual data and identifies, in the event-based visual data, a lip movement of the user.
  • the lip movement identification unit 109 determines, in the event-based visual data, a region-of-interest (ROI) of a mouth region of the user and, similar to the lip movement identification unit 10 of Fig. 1, identifies the lip movement of the user based on events, of the event-based visual data, from the ROI of the mouth region of the user.
  • the lip movement identification unit 109 generates lip movement data that indicate the identified lip movement of the user and provides the lip movement data to the syllable data determination unit 110.
  • the syllable data determination unit 110 receives the lip movement data and determines syllable data based on the identified lip movement indicated by the lip movement data.
  • the syllable data determination unit 110 is configured similar to the syllable data determinant unit 11 of Fig. 1. Additionally, the syllable data determination unit 110 includes a viseme identification unit 114 that identifies one or more visemes based on the identified lip movement and determines the syllable data based on the at least one viseme.
  • the syllable data determination unit 110 retrieves syllables corresponding to the identified visemes from a viseme database that associates visemes with syllables based on human physiology.
  • the syllable data determination unit 110 generates the syllable data based on the syllables retrieved from the viseme database.
  • the syllable data determination unit 110 provides the syllable data to the expressed words generation unit 113.
  • the face expression detection unit 111 is configured similar to the face expression detection unit 12 of Fig. 1.
  • the face expression detection unit detects, in the event-based visual data, a face expression of the user, identifies an emotion that corresponds to the detected face expression, generates face expression data that indicate the detected face expression and the identified emotion and provides the face expression data to the expressed words generation unit 113.
  • the audio data obtaining unit 112 is configured similar to the audio data obtaining unit 13 of Fig. 1.
  • the audio data obtaining unit 112 obtains, from a microphone, audio data of the user when the user is expressing words by speaking and identifies features in the audio data.
  • the audio data obtaining unit 112 generates feature data that indicate the identified features and provides the feature data to the expressed words generation unit 113.
  • the expressed words generation unit 113 is configured similar to the expressed words generation unit 14 of Fig. 1.
  • the expressed words generation unit 113 receives the syllable data, the face expression data and the feature data and generates expressed words based on the syllable data, the face expression data and the feature data and based on an artificial neural network.
  • the expressed words generation unit 113 selects words from a word list of a language model based on the syllable data, on the face expression data and on the feature data, and inflects the selected words according to grammatical rules of the language model.
  • the expressed words generation unit 14 then outputs the selected (and inflected) words as generated expressed words.
  • the determination unit 107 provides the expressed words generated by the expressed words generation unit 113 as determined expressed words to the output unit 108.
  • the output unit 108 outputs the determined expressed words via the I/O unit 105.
  • the CPU 103, the storage unit 104 and the I/O unit 105 are configured corresponding to the CPU 3, the storage unit 4 and the I/O unit 5, respectively, of Fig. 1.
  • Figures 6a and 6b illustrate examples of the lip movement identification unit 109 of Fig. 5.
  • Fig. 6a illustrates a block diagram of a lip movement identification unit 109a for identifying a lip movement based on generating a full image of the mouth region of the user according to an embodiment.
  • the lip movement identification unit 109a is an example of the lip movement identification unit 109 of Fig. 5.
  • the lip movement identification unit 109a includes a full image generation unit 115 that generates, based on the event-based visual data of the user, a full image of the mouth region of the user, i.e., of the ROI determined by the lip movement identification unit 109a.
  • the lip movement identification unit 109a then identifies the lip movement based on the generated full image of the mouth region of the user.
  • Fig. 6b illustrates a block diagram of a lip movement identification unit 109b for identifying a lip movement based on motion vectors according to an embodiment.
  • the lip movement identification unit 109b is an example of the lip movement identification unit 109 of Fig. 5.
  • the mouth event identification unit 116 includes a mouth event identification unit 116 and a motion vector determination unit 117.
  • the mouth event identification unit 116 identifies, as mouth events, events of the event-based visual data from the mouth region of the user, i.e., from the ROI determined by the lip movement identification unit 109b.
  • the motion vector determination unit 117 determines motion vectors based on correlations between mouth events and, thus, based on the identified events of the mouth region.
  • the lip movement identification unit 109b identifies the lip movement based on the determined motion vectors.
  • Figures 7a to 7f illustrate examples of the output unit 108 of Fig. 5.
  • Fig. 7a illustrates a block diagram of an output unit 108a for displaying the determined expressed words as text according to an embodiment.
  • the output unit 108a is an example of the output unit 108 of Fig. 5 and is configured similar to the presentation unit 16 of Fig. 2a.
  • the output unit 108a includes a text rendering unit 118 that renders the determined expressed words as text.
  • the output unit 108a then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 105 to display the rendered text to the user.
  • Fig. 7b illustrates a block diagram of an output unit 108b for displaying at least one icon corresponding to the determined expressed words according to an embodiment.
  • the output unit 108b is an example of the output unit 108 of Fig. 5 and is configured similar to the presentation unit 7b of Fig. 2b.
  • the output unit 108b retrieves, via the network interface of the I/O unit 105, one or more icons from a server in the communication network. The number and content of the icons retrieved is based on a number and meaning of the determined expressed words and is chosen such that the retrieved icons correspond to the determined expressed words.
  • the output unit 108b includes an icon rendering unit 119 that renders the retrieved one or more icons.
  • the output unit 108b then causes a screen that is a presenting means of the humanmachine interface included in the I/O unit 105 to display the rendered one or more icons to the user.
  • Fig. 7c illustrates a block diagram of an output unit 108c for outputting an artificial voice corresponding to the determined expressed words according to an embodiment.
  • the output unit 108c is an example of the presentation unit 108 of Fig. 5 and is configured similar to the presentation unit 7c of Fig. 2c.
  • the output unit 108c includes an artificial voice generation unit 120.
  • the artificial voice generation unit 120 generates an artificial voice that utters the determined expressed words.
  • the output unit 108c causes a loudspeaker that is a presenting means of the human-machine interface included in the I/O unit 105 to output the generated artificial voice to the user.
  • Fig. 7d illustrates a block diagram of an output unit 108d for transmitting the determined expressed words to a communication device of another user according to an embodiment.
  • the output unit 108d is an example of the output unit 108 of Fig. 5 and includes a transmission preparation unit 121.
  • the transmission preparation unit 121 prepares the determined expressed words for a transmission to the communication device of the other user.
  • the transmission preparation unit 121 encodes the determined expressed words according to a predetermined encoding and inserts the encoded expressed words in a packet of a predetermined protocol.
  • the output unit 108d then transmits the packet via the network interface included in the I/O unit 105 to the communication device of the other user.
  • Fig. 7e illustrates a block diagram of an output unit 108e for reducing noise according to an embodiment.
  • the output unit 108e is an example of the output unit 108 of Fig. 5 and includes a noise reduction unit 122.
  • the noise reduction unit 122 reduces, in the audio data obtained by the audio data obtaining unit 112 from a microphone, noise that does not correspond to the determined expressed words.
  • the output unit 108e then causes a loudspeaker of the humanmachine interface included in the I/O unit 105 to output a sound according to the audio data in which the noise has been reduced.
  • Fig. 7f illustrates a block diagram of an output unit 108f for amplifying a voice according to an embodiment.
  • the output unit 108f is an example of the output unit 108 of Fig. 5 and includes a voice amplification unit 123.
  • the voice amplification unit 123 amplifies, in the audio data obtained by the audio data obtaining unit 112 from a microphone, a voice that corresponds to the determined expressed words.
  • the output unit 108f then causes a loudspeaker of the humanmachine interface included in the I/O unit 105 to output a sound according to the audio data in which the voice has been amplified.
  • the face expression detection unit 111 is not provided in the determination unit 107 of Fig. 5 and/or the audio data obtaining unit 112 is not provided in the determination unit 107 of Fig. 5.
  • the determination of expressed words does not include detecting, in the visual data, a face expression of the user and generating the expressed words based on the detected face expression and/or the determination of expressed words does not include obtaining audio data from a microphone and generating the expressed words based on the audio data.
  • the output unit 108 of Fig. 5 includes any combination of the output units 108a to 108f.
  • the event-based vision sensor 101 that captures the event-based visual data of the user is provided separately from the device 100 and the visual data obtaining unit 106 obtains the event-based visual data from the event-based vision sensor 101 via the network interface included in the I/O unit 105.
  • the human-machine interface with the presenting means is not provided in the I/O unit 105 but separately from the device 100.
  • the determined expressed words are transmitted by the output unit 108 via the network interface of the I/O unit 105 to the presenting means.
  • the device 100 that includes the circuitry 102 may be configured as a server and the human-machine interface with the presenting means may be provided on a communication device of a user that communicates with the server via a communication network.
  • the microphone from which the audio data obtaining unit 112 obtains the audio data is, in some embodiments, provided in the device 100 and is, in some embodiments, provided separately from the device 100.
  • Fig. 8 illustrates a flow diagram of a method 150 for visual speech processing based on eventbased visual data according to an embodiment.
  • the method 150 is an example of a method that is executed by the circuitry 102 of Fig. 5.
  • the method 150 includes, at S 151, an obtaining of event-based visual data that is performed by the visual data obtaining unit 106 of Fig. 5.
  • the obtaining of event-based visual data at SI 51 obtains event-based visual data of a user from the event-based vision sensor 101.
  • a determination is performed by the determination unit 107 of Fig. 5.
  • the determination at SI 52 determines expressed words based on the event-based visual data obtained by the obtaining of event-based visual data at S151.
  • the determination includes a lip movement identification at SI 53, a syllable data determination at SI 54, a face expression detection at SI 55, an obtaining of audio data at SI 56 and an expressed words generation at SI 57.
  • the lip movement identification is performed by the lip movement identification unit 109 of Fig. 5.
  • the lip movement identification identifies, in the event-based visual data, a lip movement of the user and determines, in the event-based visual data, a region-of-interest (ROI) of a mouth region of the user.
  • ROI region-of-interest
  • the syllable data determination is performed by the syllable data determination unit 110 of Fig. 5.
  • the syllable data determination determines syllable data based on the lip movement identified by the lip movement at S153.
  • the syllable data determination includes, at S158, a viseme identification that is performed by the viseme identification unit 114 of Fig. 5.
  • the viseme identification identifies, based on the identified lip movement, at least one viseme, and the syllable data determination determines the syllable data based on the at least one viseme.
  • the face expression detection is performed by the face expression detection unit 111 of Fig. 5.
  • the face expression detection detects, in the event-based visual data, a face expression of the user and identifies an emotion corresponding to the face expression.
  • the obtaining of audio data is performed by the audio data obtaining unit 112 of Fig. 5.
  • the obtaining of audio data obtains audio data from a microphone when the user expressed words by speaking.
  • the obtaining of audio data identifies features in the audio data.
  • the expressed words generation is performed by the expressed words generation unit 113 of Fig. 5.
  • the expressed words generation generates expressed words based on the syllable data determined by the syllable data determination at SI 54, on the face expression detected by the face expression detection at SI 55 and the identified corresponding emotion, and on the audio data obtained by the audio data obtaining at SI 56 and the features identified therein.
  • the expressed words generation unit generates expressed words based on an artificial neural network.
  • the expressed words generation selects, based on the determined syllable data, on the detected face expression and the corresponding emotion and on the obtained audio data and the features identified therein, words from a word list of a language model and inflects the selected words based on grammatical rules of the language model.
  • the determination determines the selected (and inflected) words as determined expressed words of the user.
  • an output is performed by the output unit 108 of Fig. 5.
  • the output outputs the determined expressed words.
  • Figures 9a and 9b illustrate examples of the lip movement identification at S153 of Fig. 8.
  • Fig. 9a illustrates a block diagram of a lip movement identification at S153a based on generating a full image according to an embodiment.
  • the lip movement identification at SI 53 a is an example for the lip movement identification at SI 53 of Fig. 8 and is performed by the lip movement identification unit 109a of Fig. 6a.
  • the lip movement identification at SI 53a includes a full image generation at SI 60 that is performed by the full image generation unit 115 of Fig. 6a.
  • the full image generation generates, based on the event-based visual data and on the ROI of the mouth region of the user determined by the lip movement identification at S 153a, a full image of the mouth region of the user.
  • the lip movement identification then identifies the lip movement of the user based on the generated full image of the mouth region of the user.
  • Fig. 9b illustrates a block diagram of a lip movement identification at SI 53b based on motion vectors according to an embodiment.
  • the lip movement identification at SI 53b is an example of the lip movement identification at SI 53 of Fig. 8 and is performed by the lip movement identification unit 109b of Fig. 6b.
  • the lip movement identification at SI 53b includes a mouth event identification at SI 61 and a motion vector determination at SI 62.
  • the mouth event identification at SI 61 is performed by the mouth event identification unit 116 of Fig. 6b and identifies events of the event-based visual data from the ROI of the mouth region determined by the lip movement identification at SI 53b as mouth events.
  • the motion vector determination at SI 62 is performed by the motion vector determination unit 117 of Fig. 6b and determines motion vectors based on the identified mouth events.
  • the lip movement identification at SI 53b then identifies the lip movement based on the determined motion vectors.
  • Figures 10a to lOf illustrate examples of the output at S159 of
  • Fig. 10a illustrates a block diagram of an output at SI 59a for displaying the determined expressed words as text according to an embodiment.
  • the output at SI 59a is an example of the output at S159 of Fig. 8 and is performed by the output unit 108a of Fig. 7a.
  • the output at S159a includes rendering of text at SI 63 performed by the text rendering unit 118 of Fig. 7a.
  • the rendering of text at SI 63 renders the determined expressed words as text.
  • the output at SI 59a then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 105 to display the rendered text to the user.
  • Fig. 10b illustrates a block diagram of an output at SI 59b for displaying at least one icon corresponding to the determined expressed words according to an embodiment.
  • the output at S159b is an example of the output at S159 of Fig. 8 and is performed by the output unit 108b of Fig. 7b.
  • the output at SI 59b retrieves, via the network interface of the I/O unit 105, one or more icons from a server in the communication network. The number and content of the icons retrieved is based on a number and meaning of the determined expressed words and is chosen such that the retrieved icons correspond to the determined expressed words.
  • the output at SI 59b includes rendering of icons at SI 64 performed by the icon rendering unit 119 of Fig. 7b.
  • the rendering of icons at SI 64 renders the retrieved one or more icons.
  • the output at SI 59b then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 105 to display the rendered one or more icons to the user.
  • Fig. 10c illustrates a block diagram of an output at SI 59c for outputting an artificial voice corresponding to the determined expressed words according to an embodiment.
  • the output at S159c is an example of the output S159 of Fig. 8 and is performed by the output unit 108c of Fig. 7c.
  • the output at S159c includes an artificial voice generation at S165 performed by the artificial voice generation unit 120 of Fig. 7c.
  • the artificial voice generation at SI 65 generates an artificial voice that utters the determined expressed words.
  • the output at SI 59c then causes a loudspeaker that is a presenting means of the human-machine interface included in the I/O unit 105 to output the generated artificial voice to the user.
  • Fig. lOd illustrates a block diagram of an output at S159d for transmitting the determined expressed words to a communication device of another user according to an embodiment.
  • the output at S159d is an example for the output at SI 59 of Fig. 8 and is performed by the output unit 108d of Fig. 7d.
  • the output at S159d includes a transmission preparation at SI 66 that is performed by the transmission preparation unit 121 of Fig. 7d.
  • the transmission preparation at SI 66 prepares the determined expressed words for a transmission to the communication device of the other user.
  • the transmission preparation encodes the determined expressed words according to a predetermined encoding and inserts the encoded expressed words in a packet of a predetermined protocol.
  • the output at S159d then transmits the packet via the network interface included in the I/O unit 105 to the communication device of the other user.
  • Fig. lOe illustrates a block diagram of an output at S159e for reducing noise according to an embodiment.
  • the output at S159e is an example of the output at SI 59 of Fig. 8 and is performed by the output unit 108e of Fig. 7e.
  • the output at S159e includes a noise reduction at S167 that is performed by the noise reduction unit 122 of Fig. 7e.
  • the noise reduction reduces, in the audio data obtained by the obtaining of audio data at S156 of Fig. 8 from a microphone, noise that does not correspond to the determined expressed words.
  • the output at S159e then causes a loudspeaker of the human-machine interface included in the I/O unit 105 to output a sound according to the audio data in which the noise has been reduced.
  • Fig. lOf illustrates a block diagram of an output at S159f for amplifying a voice according to an embodiment.
  • the output at S159f is an example of the output at SI 59 of Fig. 8 and is performed by the output unit 108f of Fig. 7f.
  • the output at S159f includes a voice amplification at S168 that is performed by the voice amplification unit 123 of Fig. 7f.
  • the voice amplification amplifies, in the audio data obtained by the obtaining of audio data at S156 of Fig. 8 from a microphone, a voice that corresponds to the determined expressed words.
  • the output at S159f then causes a loudspeaker of the human-machine interface included in the I/O unit 105 to output a sound according to the audio data in which the voice has been amplified.
  • the output 159 of Fig. 8 includes any combination of the output 159a to 159f.
  • the face expression detection at SI 55 is not included in the determination at S152 of Fig. 8 and/or the obtaining of audio data at S156 is not included in the determination at SI 52 of Fig. 8.
  • the determination of expressed words does not include detecting, in the visual data, a face expression of the user and generating the expressed words based on the detected face expression and/or the determination of expressed words does not include obtaining audio data from a microphone and generating the expressed words based on the audio data.
  • Fig. 11 illustrates a block diagram of a general application example according to an embodiment.
  • the device 200 is an example of the device 1 of Fig. 1 and of the device 100 of Fig. 5.
  • the device 200 includes an event-based vision sensor 201, a circuitry 202 and a loudspeaker 203.
  • the event-based vision sensor 201 captures events of a face (including a mouth region) of a user 204 when the user expresses words by speaking. During speaking, the user 204 moves his lips and, thus, expressed various visemes.
  • the event-based vision sensor 201 records transitions between the visemes in event-based visual data 205 of the user 204 and provides the event-based visual data 205 to the circuitry 202.
  • the circuitry 202 is an example of the circuitry 2 of Fig. 1 and of the circuitry 102 of Fig. 5.
  • the circuitry 202 identifies, based on the event-based vision data 205, lip movements of the user 204 and corresponding visemes, determines expressed words of the user 204 based on the identified lip movements and visemes, and generates audio data 206 of an artificial voice that utters the determined expressed words.
  • the circuitry 202 then provides the audio data 206 to the loudspeaker 203.
  • the loudspeaker 203 outputs a sound that corresponds to the audio data 206 and, thus, to the expressed words of the user 204.
  • Fig. 12 illustrates a block diagram of an exemplary application for speech enablement for mutes according to an embodiment.
  • the circuitry 300 enables a mute user 301 to communicate with a hearing user 302 by expressing words through lip movements.
  • the mute user 301 cannot move his lips as if he was speaking (i.e., according to audible speech) because he did not learn speaking due to a deafness.
  • the circuitry 300 does not require lip movements that correspond to audible speech. Instead, the circuitry 300 can translate any lip movement to an alphabet and the mute user 301 can learn his own language.
  • the circuitry 300 also provides a training app that suggests and teaches an optimal lip language (including lip movements) for the mute user 301.
  • the lip movements of the optimal lip language are optimized in that visemes of the optimal lip language are robust to detect and reduce ambiguities in identifying syllables based on the visemes.
  • the circuitry 300 is an example of the circuitry 2 of Fig. 1, and includes an expressed words determination unit 310, a text rendering unit 311 and an artificial voice generation unit 312.
  • the expressed words determination unit 310 receives, from a user input unit 320 that includes a touchscreen, user feedback data 321, receives, from a visual sensor 322 that includes an eventbased vision sensor, event-based visual data 323, and receives, from a microphone 324, audio data 325.
  • the visual sensor 322 captures, as the event-based visual data 323, lip movements of the user 301 when the user 301 is expressing words by lip movements.
  • the microphone 324 captures, as the audio data 325, sounds generated by the user 301 when the user 301 is expressing words.
  • the expressed words determination unit 310 correlates the audio data 325 with the identified lip movement for increasing a robustness of the determination of expressed words of the user 301.
  • the expressed words determination unit 310 determines, based on the event-based visual data 323 from the visual sensor 322 and on the audio data 325 from the microphone 324, expressed words 313 of the user 301 and provides the determined expressed words 313 to the text rendering unit 311 and to the artificial voice generation unit 312.
  • the text rendering unit 311 receives the determined expressed words 313, renders the determined expressed words 313 as text 314 and causes the rendered text 314 to be displayed, on a screen, as recognized text 315.
  • the artificial voice generation unit 312 receives the determined expressed words 313, generates an artificial voice 316 that utters the determined expressed words 313 and causes the generated artificial voice 316 to be output, on a loudspeaker, as synthesized audio 317.
  • the user 302 reads the recognized text 315 and hears the synthesized audio 317. Thus, the user 302 can understand the words expressed by the mute user 301 by lip movements although the user 302 has not learnt lip reading.
  • the mute user 301 reads the recognized text 315 and controls whether the determined expressed words 313 correspond to the intended words.
  • the user 301 enters a teacher signal 303 as feedback into the user input unit 320 to correct an incorrect determination of expressed words and to confirm a correct determination of expressed words.
  • the teacher signal 303 includes at least one of a label, a loss and a fitness for a machine learning algorithm performed by the expressed words determination unit 310.
  • the user input unit 320 provides the teacher signal 303 as user feedback data 321 to the expressed words determination unit 310.
  • the expressed words determination unit 310 then adjusts a parameter of the machine learning algorithm according to the user feedback data 321.
  • the user 301 can correct incorrectly determined expressed words by the teacher signal 303.
  • the circuitry 300 can learn a “sign language” using the mouth (including lip movements) which allows for real-time communication of the mute (and deaf) user 301 with the hearing user 302 since the listening user 302 does not need to understand sign language.
  • the user feedback provided by the user 301 via the teacher signal 303 improves the performance of the expressed words determination unit 310.
  • the teacher signal 303 improves the determined expressed words 313.
  • the visual sensor 322 includes, instead of or in addition to an event-based vision sensor, a CMOS image sensor or a CCD image sensor.
  • the user input unit 320 includes, instead of or in addition to the touchscreen, any other input means for receiving the teacher signal 303.
  • the determination of expressed words of the user 301 only uses the visual sensor 322 as input and is performed by a processor that includes the circuitry 300; i.e., in some embodiments, the microphone 324 is not provided and the determination of expressed words of the user 301 is not based on the audio data 325.
  • the user 302 only reads the recognized text 315 and does not hear the synthesized audio 317 or, vice versa, the user 302 only hears the synthesized audio 317 and does not read the recognized text 315.
  • the user 302 does neither read the recognized text 315 nor hear the synthesized audio 317, but only the mute user 301 reads the recognized text 315. This may be the case when the mute user 301 is alone and trains the circuitry 300 to correctly determine expressed words based on the lip movements of the user 301 or trains the optimal lip language based on the training app provided by the circuitry 300.
  • the artificial voice generation unit 312 may be omitted, and the circuitry 300 outputs the determined expressed words 313 only as rendered text 314.
  • Figures 13a to 13c illustrate examples of algorithms performed by the expressed words determination unit 310 of Fig. 12 for determining expressed words based on an event-based vision sensor.
  • Fig. 13a illustrates a block diagram of an algorithm for an expressed words determination 330 based on generating an image according to an embodiment.
  • the expressed words determination 330 is performed by the expressed words determination unit 310 of Fig. 12 and includes an events-to-image conversion at S331 and an image-to-words conversion at S332.
  • the events-to-image conversion at S331 receives, from the event-based vision sensor of the visual sensor 322, event-based visual data 333 representing events and generates an image 334 based on events represented by the event-based visual data 333.
  • the generated image 334 is provided to the image-to-words conversion at S332.
  • the image-to-words conversion 332 determines, based on the image 334, expressed words 335.
  • the expressed words 335 are then provided to the text rendering unit 311 and to the artificial voice generation unit 312.
  • Fig. 13b illustrates a block diagram of an algorithm for an expressed words determination 340 based on converting motion to words according to an embodiment.
  • the expressed words determination 340 is performed by the expressed words determination unit 310 of Fig. 12 and includes an events-to-mouth ROI conversion at S341, an events-to-motion conversion at S342 and a motion-to-words conversion at S343.
  • the events-to-mouth ROI conversion at S341 receives, from the event-based vision sensor of the visual sensor 322, event-based visual data 344 representing events, determines a mouth region- of-interest (ROI) in the event-based visual data 344 that corresponds to a mouth region of the user 301, and selects, as mouth events 345, events of the event-based visual data 344 included in the mouth ROI.
  • the mouth events 345 are provided to the event-to-motion conversion at S342.
  • the events-to-motion conversion at S342 determines motion vectors 346 based on the mouth events 345 that indicate lip movements of the user 301.
  • the motion vectors 346 are provided to the motion-to-words conversion at S343.
  • the motion-to-words conversion at S343 determines expressed words 349 based on the motion vectors 346 and on word probabilities 347 provided by a language model 348.
  • the expressed words 349 are then provided to the text rendering unit 311 and to the artificial voice generation unit 312.
  • Fig. 13c illustrates a block diagram of an algorithm for an expressed words determination 350 based on directly converting events to words according to an embodiment.
  • the expressed words determination 350 is performed by the expressed words determination unit 310 of Fig. 12 and includes an events-to-words conversion at S351.
  • the events-to-words conversion at S351 receives, from the event-based vision sensor of the visual sensor 322, event-based visual data 352 representing events and converts the events represented by the event-based visual data 352 to expressed words 353 based on an artificial neural network.
  • the expressed words 353 are then provided to the text rendering unit 311 and to the artificial voice generation unit 312.
  • Fig. 14 illustrates an application example of a speech enablement for mutes based on a smartphone 360 according to an embodiment.
  • the smartphone 360 includes the circuitry 300 of Fig. 12, and further includes a microphone 361, a visual sensor 362 and a loudspeaker 363.
  • the microphone 361 is an example of the microphone 324 of Fig. 12
  • the visual sensor 362 is an example of the visual sensor 322
  • the loudspeaker 363 is an example of the loudspeaker that outputs the synthesized audio 317 of Fig. 12.
  • the mute user 301 is also deaf, so the microphone 361 records the speech 364 of the user 302 (e.g., the sentence: “How are you?”) and translates it to text 365 that is displayed on a screen of the smartphone 360 of the mute user 301.
  • the mute user 301 reads the text 365 (e.g., “How are you?”) and answers with voiceless speech 366 by moving his lips according to a lip language trained by the circuitry 300.
  • the visual sensor 362 of the smartphone 360 captures the voiceless speech 366 of the user 301, and the circuitry 300 translates the voiceless speech 366 of the mute user 301 to audio 367.
  • the loudspeaker 363 outputs the audio 367 (e.g., “I’m fine.”) so that the hearing user 302 can hear and understand the answer.
  • the mute user 301 can have a normal conversation with the hearing user 302.
  • Fig. 15 illustrates an exemplary application of a video call enhancement according to an embodiment.
  • a live caption is created and the voice of a speaker 401 taking a video call in a noisy environment is enhanced.
  • the speaker 401 is in a noisy environment and participates in a video call with a smartphone 402.
  • the smartphone is an example of the device 100 of Fig. 5.
  • the speaker 401 expresses words as a speech (e.g., “It’s loud here!”).
  • the speech of the speaker 401 is captured by the event-based vision sensor 101 of the smartphone 401 and determines the words expressed by the speaker 401 with his speech, thus translating the speech to a live caption, using only the event-based visual data captured by the event-based vision sensor 101.
  • the smartphone 402 records voiceless speech.
  • the smartphone 401 transmits the live caption including the determined expressed words together with video of the speaker 401 via a wireless communication connection 403 based on Wi-Fi to a communication network 404.
  • a communication device 407 of another user receives, via a communication connection 406 to the communication network 404, the video and live caption.
  • the communication device 407 displays the received video and outputs the expressed words of the received live caption.
  • Fig. 16a illustrates a block diagram of an algorithm for video call voice enhancement based on fusion of expressed words determined by visual data and expressed words determined by audio data according to an embodiment.
  • the expressed words determination unit 500a is an example of the circuitry 102 of Fig. 5 and is configured to enhance a speech captured for a video call by generating an artificial voice (computer voice) based on an event-based vision sensor 501 and a microphone 503.
  • the expressed words determination unit 500a is configured to perform an expressed words determination based on visual data at S510, an expressed words determination based on audio data at S511, a fusion at S51 and an artificial voice generation at S513.
  • the expressed words determination based on visual data at S510 receives, from the event-based vision sensor 501, event-based visual data 501 of a user who is speaking (and, thus, expressing words) and determines, based on the event-based visual data 501, expressed words 514 of the user.
  • the expressed words determination based on audio data at S511 receives, from the microphone 503, noisy audio data 504 that indicate a speech of the user who is speaking and environment noise, and determines, based on the noisy audio data 504, expressed words 515 of the user.
  • the expressed words 514 determined at S510 based on the event-based visual data 502 and the expressed words 515 determined at S511 based on the noisy audio data 504 are provided to the fusion at S512.
  • the fusion at S512 fuses the expressed words 514 determined at S510 based on the event-based visual data 502 and the expressed words 515 determined at S511 based on the noisy audio data 504 to determined expressed words 516 and provides the determined expressed words 516 to the artificial voice generation at S513.
  • the artificial voice generation at S513 generates an artificial voice (computer voice) that utters the determined expressed words 516 and provides clear audio data 517 that indicate the artificial voice to a loudspeaker 505.
  • the loudspeaker 505 outputs the clear audio data 517 and, thus, outputs the determined expressed words of the user.
  • the fusion at S512 is performed after determining the expressed words 514 and 515 based on the event-based visual data 502 and on the noisy audio data 504, respectively.
  • the event-based visual data 502 and the noisy audio data 504 are fused before determining expressed words based thereon.
  • the algorithm processes the event-based visual data 502 and the noisy audio data 504 simultaneously based on an artificial neural network. An example of such a case is provided in Fig. 16b.
  • Fig. 16b illustrates a block diagram of an algorithm for video call voice enhancement based on fusion of visual data and audio data before determining expressed words according to an embodiment.
  • the expressed words determination unit 500b is an example of the circuitry 102 of Fig. 5 and is configured to enhance a speech captured for a video call by generating an artificial voice (computer voice) based on an event-based vision sensor 501 and a microphone 503.
  • the expressed words determination unit 500b differs from the expressed words determination unit 500a of Fig. 16a in that it is configured to perform a combined expressed words determination based on visual data and audio data at S520, i.e., expressed words are determined after fusing the event-based visual data 502 and the noisy audio data 504.
  • the expressed words determination unit 500b is also configured to perform the artificial voice generation at S513.
  • the combined expressed words determination based on visual data and audio data at S520 receives the event-based visual data 502 and the noisy audio data 504 fused together, determines, with an artificial neural network, expressed words 521 of the user simultaneously based on the event-based visual data 502 and the noisy audio data 504 and provides the determined expressed words 521 to the artificial voice generation at S513.
  • the artificial voice generation at S513 and the loudspeaker 505 of Fig. 16b correspond to the artificial voice generation at S513 and the loudspeaker 505 of Fig. 16a, respectively.
  • the algorithm performed by the expressed words determination unit 500b thus, performs an audio-video speech recognition.
  • Fig. 16c illustrates a block diagram of an algorithm for video call voice enhancement based on removing a noise of an environment according to an embodiment.
  • the expressed words determination unit 500c is an example of the circuitry 102 of Fig. 5 and is configured to enhance a speech captured for a video call by removing a noise of an environment based on the event-based vision sensor 501 and the microphone 503.
  • the expressed words determination unit 500c is configured to perform the expressed words determination based on visual data at S510, an audio prediction at S530 and a noise filtering at S531.
  • the expressed words determination based on visual data at S510 corresponds to the expressed words determination based on visual data at S510 of Fig. 16a and provides the determined expressed words 514 of the user to the audio prediction at S530.
  • the audio prediction at S530 predicts, based on the determined expressed words 514, predicted audio data 532 of the speech of the user and provides the predicted audio data 532 to the noise filtering at S531.
  • the noise filtering at S531 receives the noisy audio data 504 from the microphone, removes noise from the noisy audio data 504 based on the predicted audio data 532, thus generating clean audio data 533 that correspond to the speech of the user, and provides the clean audio data 533 to the loudspeaker 505.
  • the loudspeaker 505 is configured like the loudspeaker 505 of Fig. 16a.
  • hearing aids In some embodiments, an event-based vison sensor is installed on hearing aids.
  • the hearing aids include a circuitry according to the present disclosure (e.g., the circuitry 102 of Fig. 5) that amplifies a voice of a person that the wearer of the hearing aids is looking at.
  • Such hearing aids improves an amplification of a correct audio source (cocktail party problem) because the wearer of the hearing aids can decide which source to amplify by looking at it.
  • noise protection headphones Further applications of the present technology include noise protection headphones.
  • a circuitry according to the present disclosure e.g., the circuitry 102 of Fig. 5
  • noise protection headphones According to the same concept as for the hearing aids, for example construction workers or factory workers in a loud environment may wear the noise protection headphones which may remove (or at least reduce) a noise from the environment and amplify a voice of another worker.
  • voice-enabled ATMs, kiosks and vending machines in noisy environments include a circuitry according to the present disclosure (e.g., the circuitry 102 of Fig. 5) for obtaining user input of a speaking user based on event-based visual data of the user.
  • determining spoken (or, more generally, expressed) words of a user based on a camera (i.e., on visual data) rather than on a microphone (i.e., on audio data) allows to determine spoken (or expressed) words of the user in a very noisy environment or for a mute user.
  • determining spoken (expressed) words of a user based on event-based visual data obtained from an event-based vision sensor provides at least one of the following advantages:
  • the event-based vision sensor may be capable of capturing fast moving lips that could be a challenge for a normal camera.
  • the event-based vision sensor may compress redundant data (e.g. when lips do not move) and, thereby, make processing more efficient.
  • Event-based visual data captured by the event-based vison sensor may be more private than visual data captured by a normal camera since it is more difficult to reconstruct, based on eventbased visual data form an event-based vision sensor, a clear image that could be used to identify a person.
  • the present technology may provide better video calls everywhere (even in very noisy environments), may provide speech enablement for mutes and/or may provide factory automation.
  • circuitry 2 of Fig. 1 into units 6 to 8 is only made for illustration purposes and that the present disclosure is not limited to any specific division of functions in specific units.
  • circuitry 102 of Fig. 5 into units 106 to 108 is only made for illustration purposes and the present disclosure is not limited to any specific division of functions in specific units.
  • the circuitry 2 and/or the circuitry 102 could be implemented by a respective programmed processor, field programmable gate array (FPGA) and the like.
  • FPGA field programmable gate array
  • the method 30 of Fig. 3 and/or the method 150 of Fig. 8 can also be implemented as a computer program causing a computer and/or a processor, such as circuitry 2 of Fig. 1 and/or circuitry 102 of Fig. 5 discussed above, to perform the method, when being carried out on the computer and/or processor.
  • a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the method described to be performed.
  • a circuitry for visual speech processing configured to: determine expressed words based on visual data of a user; present the determined expressed words to the user; and adjust the determination of expressed words based on feedback of the user related to the determined expressed words.
  • the determination of expressed words includes: obtaining audio data from a microphone; and generating the expressed words based on the audio data.
  • a method for visual speech processing comprising: determining expressed words based on visual data of a user; presenting the determined expressed words to the user; and adjusting the determination of expressed words based on feedback of the user related to the determined expressed words.
  • a computer program comprising program code causing a computer to perform the method according to anyone of (17) to (32), when being carried out on a computer.
  • a non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (17) to (32) to be performed. Note further that the present technology can also be configured as described below.
  • a circuitry for visual speech processing configured to: obtain event-based visual data of a user from an event-based vision sensor; determine expressed words based on the obtained event-based visual data; and output the determined expressed words.
  • the circuitry of (2), wherein the identifying of the lip movement includes: generating, based on the event-based visual data, at least one full image of a mouth region of the user; and identifying the lip movement based on the generated at least one full image of the mouth region of the user.
  • identifying of the lip movement includes: identifying, based on the event-based visual data, events of a mouth region of the user; determining motion vectors based on the identified events of the mouth region; and identifying the lip movement based on the determined motion vectors.
  • determining of the syllable data includes: identifying, based on the identified lip movement, at least one viseme; and determining the syllable data based on the at least one viseme.
  • the circuitry of (14) or (15), wherein the outputting of the determined expressed words includes: amplifying, in the audio data obtained from the microphone, a voice that corresponds to the determined expressed words; and outputting the audio data in which the voice has been amplified.
  • a method for visual speech processing comprising: obtaining event-based visual data of a user from an event-based vision sensor; determining expressed words based on the obtained event-based visual data; and outputting the determined expressed words.
  • the method of ( 18), wherein the identifying of the lip movement includes: generating, based on the event-based visual data, at least one full image of a mouth region of the user; and identifying the lip movement based on the generated at least one full image of the mouth region of the user.
  • identifying of the lip movement includes: identifying, based on the event-based visual data, events of a mouth region of the user; determining motion vectors based on the identified events of the mouth region; and identifying the lip movement based on the determined motion vectors.
  • a computer program comprising program code causing a computer to perform the method according to anyone of (17) to (32), when being carried out on a computer.
  • a non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (17) to (32) to be performed.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

The present disclosure provides a circuitry for visual speech processing, configured to determine expressed words based on visual data of a user; present the determined expressed words to the user; and adjust the determination of expressed words based on feedback of the user related to the determined expressed words.

Description

CIRCUITRY AND METHOD FOR VISUAL SPEECH PROCESSING
TECHNICAL FIELD
The present disclosure generally pertains to a circuitry and a method, in particular to a circuitry and a method for visual speech processing.
TECHNICAL BACKGROUND
It is known to perform visual speech recognition of a person’s speech based on a visual observation of a speaking person. For visual speech recognition, a mouth region of the speaking person is observed and expressed words are derived from lip movements of the speaking person.
Although there exist techniques for visual speech processing, it is generally desirable to provide an improved circuitry and a method for visual speech processing.
SUMMARY
According to a first aspect, the disclosure provides a circuitry for visual speech processing, configured to determine expressed words based on visual data of a user; present the determined expressed words to the user; and adjust the determination of expressed words based on feedback of the user related to the determined expressed words.
According to a second aspect, the disclosure provides a method for visual speech processing that includes determining expressed words based on visual data of a user; presenting the determined expressed words to the user; and adjusting the determination of expressed words based on feedback of the user related to the determined expressed words.
Further aspects are set forth in the dependent claims, the following description and the drawings.
BRIEF DESCRIPTION OF THE DRAWINGS
Embodiments are explained by way of example with respect to the accompanying drawings, in which:
Fig. 1 illustrates a block diagram of a device with a circuitry for visual speech processing and receiving user feedback according to an embodiment;
Fig. 2a illustrates a block diagram of a presentation unit for displaying determined expressed words as text according to an embodiment;
Fig. 2b illustrates a block diagram of a presentation unit for displaying at least one icon corresponding to determined expressed words according to an embodiment; Fig. 2c illustrates a block diagram of a presentation unit for outputting an artificial voice corresponding to determined expressed words according to an embodiment;
Fig. 3 illustrates a flow diagram of a method for visual speech processing and receiving user feedback according to an embodiment;
Fig. 4a illustrates a block diagram of a presentation for displaying determined expressed words as text according to an embodiment;
Fig. 4b illustrates a block diagram of a presentation for displaying at least one icon corresponding to determined expressed words according to an embodiment;
Fig. 4c illustrates a block diagram of a presentation for outputting an artificial voice corresponding to determined expressed words according to an embodiment;
Fig. 5 illustrates a block diagram of a device with a circuitry for visual speech processing based on event-based visual data according to an embodiment;
Fig. 6a illustrates a block diagram of a lip movement identification unit for identifying a lip movement based on generating a full image of a mouth region of a user according to an embodiment;
Fig. 6b illustrates a block diagram of a lip movement identification unit for identifying a lip movement based on motion vectors according to an embodiment;
Fig. 7a illustrates a block diagram of an output unit for displaying determined expressed words as text according to an embodiment;
Fig. 7b illustrates a block diagram of an output unit for displaying at least one icon corresponding to determined expressed words according to an embodiment;
Fig. 7c illustrates a block diagram of an output unit for outputting an artificial voice corresponding to determined expressed words according to an embodiment;
Fig. 7d illustrates a block diagram of an output unit for transmitting determined expressed words to a communication device of another user according to an embodiment;
Fig. 7e illustrates a block diagram of an output unit for reducing noise according to an embodiment;
Fig. 7f illustrates a block diagram of an output unit for amplifying a voice according to an embodiment; Fig. 8 illustrates a flow diagram of a method for visual speech processing based on event-based visual data according to an embodiment;
Fig. 9a illustrates a block diagram of a lip movement identification based on generating a full image according to an embodiment;
Fig. 9b illustrates a block diagram of a lip movement identification based on motion vectors according to an embodiment;
Fig. 10a illustrates a block diagram of an output for displaying determined expressed words as text according to an embodiment;
Fig. 10b illustrates a block diagram of an output for displaying at least one icon corresponding to determined expressed words according to an embodiment;
Fig. 10c illustrates a block diagram of an output for outputting an artificial voice corresponding to determined expressed words according to an embodiment;
Fig. lOd illustrates a block diagram of an output for transmitting determined expressed words to a communication device of another user according to an embodiment;
Fig. lOe illustrates a block diagram of an output for reducing noise according to an embodiment;
Fig. lOf illustrates a block diagram of an output for amplifying a voice according to an embodiment;
Fig. 11 illustrates a block diagram of a general application example according to an embodiment;
Fig. 12 illustrates a block diagram of an exemplary application for speech enablement for mutes according to an embodiment;
Fig. 13a illustrates a block diagram of an algorithm for an expressed words determination based on generating an image according to an embodiment;
Fig. 13b illustrates a block diagram of an algorithm for an expressed words determination based on converting motion to words according to an embodiment;
Fig. 13c illustrates a block diagram of an algorithm for an expressed words determination based on directly converting events to words according to an embodiment;
Fig. 14 illustrates an application example of a speech enablement for mutes based on a smartphone according to an embodiment;
Fig. 15 illustrates an exemplary application of a video call enhancement according to an embodiment; Fig. 16a illustrates a block diagram of an algorithm for video call voice enhancement based on fusion of expressed words determined by visual data and expressed words determined by audio data according to an embodiment;
Fig. 16b illustrates a block diagram of an algorithm for video call voice enhancement based on fusion of visual data and audio data before determining expressed words according to an embodiment; and
Fig. 16c illustrates a block diagram of an algorithm for video call voice enhancement based on removing a noise of an environment according to an embodiment.
DETAILED DESCRIPTION OF EMBODIMENTS
Before a detailed description of the embodiments is given under reference of Fig. 1, general explanations are made.
As mentioned in the outset, it is known to perform visual speech recognition of a person’s speech based on a visual observation of a speaking person. For visual speech recognition, a mouth region of the speaking person is observed and expressed words are derived from lip movements of the speaking person.
In some embodiments of the present disclosure, visual speech enablement can enable speech in a scenario where a speaker cannot be understood because of a loud environment or because the speaker is mute. Some embodiments use a camera that reads the lips of a speaker and can therefore be also used in scenarios where audio-based enhancement methods, such as background noise cancelling or voice amplification, do not work.
Consequently, some embodiments pertain to a circuitry for visual speech processing, configured to: determine expressed words based on visual data of a user; present the determined expressed words to the user; and adjust the determination of expressed words based on feedback of the user related to the determined expressed words.
The circuitry may include any entity capable of processing data, such as a microprocessor, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Tensor Processing Unit (TPU), a Reduced Instruction Set Computer (RISC), a Complex Instruction Set Computer (CISC), an Application-Specific Integrated Circuit (ASIC) and/or a Field-Programmable Gate Array (FPGA). The circuitry may be included on a device of the user whose expressed words are determined, on a device of another user, on a remote data processing apparatus such as a server, and/or may be distributed across several devices and/or data processing apparatus. The visual data may include any data obtained by an optical detector such as a camera. For example, the visual data may be based on a color (e.g. Red-Green-Blue, RGB) image obtained from a color (e.g. RGB) camera such as a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor, on a grayscale image obtained from a grayscale camera such as a CMOS or CCD image sensor, on a thermal image obtained from an infrared camera, on a depth map image obtained from a time-of-flight (ToF) sensor, and/or on an event obtained from an event-based vision sensor.
The visual data may include a single image frame or a sequence of image frames. The image frame(s) may be acquired with a short exposure such as to obtain a sharp image or with a long exposure such as to obtain an image with a motion blur that indicates a movement. The visual data may include events obtained asynchronously or at a high detection rate from an event-based vision sensor, wherein each event may indicate a change detected at an associated position on the sensor.
The expressed words may be words audibly uttered by the user or expressed by the user by moving his lips (and mouth) without generating an audible sound. I.e., the expressed words may be determined regardless of any sound generated by the user.
The expressed words may be unconnected or may be included in an expression or in one or more sentences. The expressed words may be words of one or more predetermined languages, e.g., a language in which a user interface of the circuitry is configured or a language of a region in which the user is located, or the language of the words may be determined during or after determining the expressed words.
The circuitry may be configured to present the determined expressed words to the user in any suitable form. For example, the determined expressed words may be presented to the user as a text rendered on a display, as one or more images or icons corresponding to the determined expressed words, as an artificial voice audibly output by a loudspeaker, or as tactile or electric stimulus.
The feedback of the user may be received via a human-machine interface. The human-machine interface may include a key, a button, a touch pad, a touch screen, a microphone for detecting a voice input of the user or a camera for detecting gesture input of the user. The user may input the feedback in the human-machine interface in any suitable form, and the circuitry may obtain the feedback from the human-machine interface. In some cases, the circuitry may be configured to extract the feedback of the user from the determined expressed words. The feedback of the user may be related to the determined expressed words by indicating a correction of the expressed words or a confirmation of the expressed words. The circuitry may or may not be configured to request the feedback from the user.
The adjusting of the determination of expressed words may serve the purpose of increasing an accuracy of the determination of expressed words. For example, the determination of expressed words may be adjusted such that a probability for determining, based on future visual data that are similar to the visual data, expressed words that are indicated by the feedback of the user is increased.
In some embodiments, the determination of expressed words includes detecting, in the visual data, a mouth region of the user; and identifying a lip movement in the mouth region of the user.
For example, the visual data may include more than the mouth region of the user (e.g., a whole face, a whole body and/or surroundings of the user), and the detection of the mouth region may be based on machine learning such as on an artificial neural network, on a support vector machine, on a wavelet transform and/or on any other suitable technique. The detection of the mouth region may include detecting a region-of-interest (ROI) in the visual data that corresponds to the mouth region.
The lip movement may be identified based on a motion blur of the visual data, on a sequence of image frames of the visual data, and/or on events of the visual data.
In some embodiments, the determination of expressed words includes determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
The syllable data may indicate syllables or at least parts of syllables, e.g., that match the identified lip movement. The expressed words may be generated based on the syllables or parts of syllables indicated by the syllable data.
In some embodiments, the generating of the expressed words is based on a language model.
The language model may be specific for a predetermined language in which the expressed words should be generated and may represent language-specific rules including grammatical rules. For example, the language model may be implemented as artificial neural network trained towards a predetermined language.
In some embodiments, the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data. For example, the generating of the expressed words may include selecting one or more words of the word list that correspond to the syllables or part of syllables indicated by the syllable data. The language model may select the one or more words based on a list of syllables used by a language and/or on syllable probabilities associated with a language. The language model may include the list of syllables and/or the syllable probabilities. The selected word(s) may be selected and/or inflected according to grammatical rules represented by the language model. The word(s) may be selected according to a probability based on a text corpus or on a history of generated words associated with a user.
In some embodiments, the determination of expressed words includes detecting, in the visual data, a face expression of the user; and generating the expressed words based on the detected face expression.
For example, the face expression may include an emotional expression (e.g. angry, happy etc.), the corresponding emotion may be detected, and the expressed words may be generated in accordance with the detected emotion. For example, words from the word list of the language model may be selected, as the determined expressed words, that correspond to the emotion.
For example, the face expression may include a movement of a lip, of an eye, of a forehead, of an eyebrow, of a nose and/or of a cheek of the user. The face expression may contribute to a sign language expressed consciously or unconsciously by the user, and determining words expressed by the user based on the face expression may increase a performance of the determination.
In some embodiments, the determination of expressed words includes obtaining audio data from a microphone; and generating the expressed words based on the audio data.
For example, the audio data may represent a sound generated by the user when expressing the words, and the expressed words may be determined based on both the visual data and the obtained audio data to increase a robustness of the determination of expressed words. The determination of expressed words may be based on the audio data if the audio data and the visual data are correlated (e.g., if a correlation between the audio data and the visual data exceeds a predetermined threshold), and may not be based on the audio data if the audio data and the visual data are not correlated (e.g., if a correlation between the audio data and the visual data does not exceed a predetermined threshold). The determination may be based on the audio data during a training, when the circuitry is trained with visual data of, e.g., a certain user. The determination of expressed words may be based on audio data that correspond to a pronunciation in accordance with the language model, and may also be based on audio data that do not correspond to the language model. For example, even a mute may generate a sound when expressing words by lip movements. Such a sound may aid a determination of words expressed by the mute.
Determining expressed words based on audio data in addition to visual data may provide a performance advantage over determining expressed words purely based on visual data.
In some embodiments, the presenting of the determined expressed words includes displaying the determined expressed words as text.
For example, the determined expressed words may be rendered as text and displayed by a display means of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
In some embodiments, the presenting of the determined expressed words includes displaying at least one icon corresponding to the determined expressed words.
For example, the at least one icon may represent a meaning of the determined expressed words. The at least one icon may be retrieved from a predetermined database or may be retrieved by a web search upon determining the expressed words.
In some embodiments, the presenting of the determined expressed words includes generating an artificial voice that utters the determined expressed words; and outputting the generated artificial voice.
For example, the artificial voice may be output by loudspeakers of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
In some embodiments, the visual data includes a plurality of events obtained from an eventbased vision sensor.
For example, an event included in the visual data may indicate a brightness change detected by a photosensitive element of the event-based vision sensor. The event-based vision sensor may detect events asynchronously or at a high detection rate such as 1,000,000 events per second, without limiting the disclosure to this detection rate.
In some embodiments, the adjusting of the determination of expressed words includes changing a value of a parameter of the determination of expressed words.
For example, the parameter may be adjusted such as to increase a probability that expressed words determined based on future visual data representing a future lip movement that is similar to the lip movement identified in the visual data correspond to the feedback of the user. In some embodiments, the determination of expressed words is based on an artificial neural network.
For example, the artificial neural network may be trained to detect, in the visual data, a mouth region of the user and/or a face expression of the user. The artificial neural network may be trained to identify a lip movement of the user, to associate a lip movement with one or more syllables and to determine syllable data based on the identified lip movement and/or based on the one or more associated syllables.
The artificial neural network may be trained to include a language model according to a language. The artificial neural network may be trained to include, as the language model, syllables used by the language, syllable probabilities of the language, a word list associated with the language, word probabilities of words from the word list and/or grammatical rules associated with the language.
The artificial neural network may be trained to determine the expressed words by comparing syllables indicated by the syllable data with the syllables, the syllable probabilities, the word list and/or the word probabilities of the language model and/or applying the grammatical rules of the language model.
The artificial neural network may be trained based on gradient descent, backpropagation and/or reinforcement learning on a database of visual data that represent lip movements and/or face expressions of users and of associated expressed words. Aspects such as detecting a mouth region, detecting a face expression, identifying a lip movement, associating a lip movement with a syllable and a language model or parts thereof may be trained explicitly in separate training steps or may be trained implicitly during the same training procedure.
The adjusting of the parameter may include adjusting, as the parameter, a weight value of the artificial neural network. The adjusting may be based on gradient descent, backpropagation and/or reinforcement learning.
In some embodiments, the feedback of the user indicates at least one of a label, a loss and a fitness for the artificial neural network.
The label may include one or more words that correspond to the visual data. The loss and/or the fitness may include a value that indicates an accuracy of the determined expressed word. For example, the value may indicate an absolute accuracy. For example, the value may indicate a relative accuracy that relates to a previous determination of expressed words. In some embodiments, the feedback of the user includes a correction of at least one of the determined expressed words; and the adjusting of the determination of expressed words is based on the correction of the at least one of the determined expressed words.
For example, the feedback may indicate one or more words that should be determined as the expressed words based on the visual data, and the adjusting of the parameter may include adjusting the parameter such as to increase a probability that the one or more words indicated by the feedback are determined as the expressed words.
In some embodiments, the presenting of the determined expressed words includes recommending an improved lip movement.
For example, the improved lip movement may allow a more robust determination of syllable data and/or of expressed words than the lip movement represented by the visual data. The improved lip movement may be recommended to the user on a display means of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses, and may be displayed by an application that is executed on the user terminal.
Some embodiments pertain to a method for visual speech processing that includes determining expressed words based on visual data of a user; presenting the determined expressed words to the user; and adjusting the determination of expressed words based on feedback of the user related to the determined expressed words.
The method may be configured corresponding to the processing performed by the circuitry described above, and all features described with reference to the processing performed by the circuitry may also be features of the method.
Some embodiments pertain to a circuitry for visual speech processing that is configured to obtain event-based visual data of a user from an event-based vision sensor; determine expressed words based on the obtained event-based visual data; and output the determined expressed words.
The circuitry may include any entity capable of processing data, such as a microprocessor, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Tensor Processing Unit (TPU), a Reduced Instruction Set Computer (RISC), a Complex Instruction Set Computer (CISC), an Application-Specific Integrated Circuit (ASIC) and/or a Field-Programmable Gate Array (FPGA). The circuitry may be included on a device of the user whose expressed words are determined, on a device of another user, on a remote data processing apparatus such as a server, and/or may be distributed across several devices and/or data processing apparatus. The event-based visual data may represent events obtained from an event-based vision sensor. The event-based vision sensor may capture the events asynchronously or at a high detection rate, wherein each event may indicate a change detected at an associated position on the sensor.
For example, an event included in the visual data may indicate a brightness change detected by a photosensitive element of the event-based vision sensor. The event-based vision sensor may detect events asynchronously or at a high detection rate such as 1,000,000 events per second, without limiting the disclosure to this detection rate.
For example, the event-based vision sensor may include an event camera, a neuromorphic camera, a silicon retina and/or a dynamic vision sensor.
For example, lips of the user may move very fast, and the event-based vision sensor may provide a temporal resolution that is high enough for resolving the lip movement.
The event-based vision sensor may provide a more efficient processing of the event-based visual data obtained with a high data rate because redundant data are naturally compressed in the eventbased visual data as compared to data from a camera that generates whole image frames instead of events.
The event-based visual data may indicate a lip movement of a user. For example, the event-based vision sensor may capture at least lips of a user that is expressing words, e.g. by speaking.
The expressed words may be words audibly uttered by the user or expressed by the user by moving his lips (and mouth) without generating an audible sound. I.e., the expressed words may be determined regardless of any sound generated by the user.
The expressed words may be unconnected or may be included in an expression or in one or more sentences. The expressed words may be words of one or more predetermined languages, e.g. a language in which a user interface of the circuitry is configured or a language of a region in which the user is located, or the language of the words may be determined during or after determining the expressed words.
The expressed words may be determined based on the event-based visual data by deriving the expressed words at least from a lip movement indicated by the visual data.
In some embodiments, the determination of expressed words includes identifying, in the eventbased visual data, a lip movement of the user.
The lip movement may be captured when the user is expressing words, e.g. by speaking. The lip movement may correspond to a sound generated by the user for expressing words. In some embodiments, the identifying of the lip movement includes generating, based on the event-based visual data, at least one full image of a mouth region of the user; and identifying the lip movement based on the generated at least one full image of the mouth region of the user.
For example, the mouth region of the user may be identified in the event-based visual data based on a shape of an area with an increased event density, and the at least one full image of the mouth region of the user may be reconstructed from the events represented by the event-based visual data, e.g. based on at least one of temporal smoothing, optimization, gradient estimation and Poisson integration.
For example, the at least one full image of the mouth region may include an image in which a motion is represented by motion blur, and the lip movement is identified based on the motion blur.
For example, the at least one full image of the mouth region may include a plurality of images of the mouth region, and the lip movement may be detected based on a difference between subsequent images of the plurality of images of the mouth region.
In some embodiments, the identifying of the lip movement includes identifying, based on the event-based visual data, events of a mouth region of the user; determining motion vectors based on the identified events of the mouth region; and identifying the lip movement based on the determined motion vectors.
For example, the events of the mouth region of the user may be events in the event-based visual data that are detected in the mouth region of the user, and the motion vectors may be determined based on correlations between events of the event of the mouth region.
In some embodiments, the determination of expressed words includes determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
The syllable data may indicate syllables or at least parts of syllables, e.g., that match the identified lip movement. The expressed words may be generated based on the syllables or parts of syllables indicated by the syllable data.
In some embodiments, the determining of the syllable data includes identifying, based on the identified lip movement, at least one viseme; and determining the syllable data based on the at least one viseme.
A viseme may indicate a lip shape or transition between lip shapes that is characteristic for one or more specific sounds generated by a user. For example, lips of a user may assume the lip shape or transition of lip shapes indicated by a viseme when the user generates the specific sound or one of the specific sounds. Thus, based on a viseme, a sound generated by the user may be identified or at least narrowed down to a (small) number of sounds. The sound(s) identified by a viseme may be predetermined by human physiology.
A syllable may include a predetermined sound or a predetermined sequence of sounds. Thus, a viseme or sequence of visemes may indicate a syllable or several alternative syllables indicated by the viseme, or may indicate a sequence of syllables or several alternative sequences of syllables indicated by the sequence of visemes.
In some embodiments, the generating of the expressed words is based on a language model.
The language model may be specific for a predetermined language in which the expressed words should be generated and may represent language-specific rules including grammatical rules. For example, the language model may be implemented as artificial neural network trained towards a predetermined language.
In some embodiments, the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data.
For example, the generating of the expressed words may include selecting one or more words of the word list that correspond to the syllables or part of syllables indicated by the syllable data. The selected word(s) may be selected and/or inflected according to grammatical rules represented by the language model. The word(s) may be selected according to a probability based on a text corpus or on a history of generated words associated with a user.
In some embodiments, the outputting of the determined expressed words includes displaying the determined expressed words as text.
For example, the determined expressed words may be rendered as text and displayed by a display means of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
In some embodiments, the outputting of the determined expressed words includes displaying at least one icon corresponding to the determined expressed words.
For example, the at least one icon may represent a meaning of the determined expressed words. The at least one icon may be retrieved from a predetermined database or may be retrieved by a web search upon determining the expressed words. In some embodiments, the outputting of the determined expressed words includes generating an artificial voice that utters the determined expressed words; and outputting the generated artificial voice.
For example, the artificial voice may be output by loudspeakers of a user terminal such as a computer, a notebook, a smartphone, a smartwatch or smart glasses.
In some embodiments, the outputting of the determined expressed words includes transmitting the determined expressed words to communication device of another user.
For example, the determined expressed words of the user may be transmitted to a communication device of another user during a phone call or a video call between the user and the other user. The communication device of the other user may output the expressed words in a manner perceivable for the other user.
In some embodiments, the determination of expressed words is based on an artificial neural network.
For example, the artificial neural network may include a deep neural network, a convolution neural network and/or a recurrent neural network, without limiting the disclosure to a specific type of an artificial neural network.
The artificial neural network may be trained to identify, in the event-based visual data, a lip movement of the user. For example, the artificial neural network may be trained to generate, based on the event-based visual data, at least one full image of a mouth region of the user and identify the lip movement based on the at least one full image of the mouth region of the user. Or, as another example, the artificial neural network may be trained to identify, based on the event-based visual data, events of a mouth region of the user, determine motion vectors based on the identified events of the mouth region of the user and identify the lip movement based on the determined motion vectors.
The artificial neural network may be trained to identify at least one viseme based on the lip movement, to associate the lip movement and/or the at least one viseme with one or more syllables and to determine syllable data based on the identified lip movement, based on the at least one viseme and/or based on the one or more associated syllables.
The artificial neural network may be trained to include a language model according to a language. The artificial neural network may be trained to include, as the language model, syllables used by the language, syllable probabilities of the language, a word list associated with the language, word probabilities of words from the word list and/or grammatical rules associated with the language.
The artificial neural network may be trained to determine the expressed words by comparing syllables indicated by the syllable data with the syllables, the syllable probabilities, the word list and/or the word probabilities of the language model and/or applying the grammatical rules of the language model.
The artificial neural network may be trained based on gradient descent, backpropagation and/or reinforcement learning on a database of visual data that represent lip movements users and of associated expressed words. Aspects such as generating a full image of a mouth region, identifying events of a mouth region, determining motion vectors, identifying a lip movement, identifying a viseme, associating a lip movement and/or a viseme with a syllable and a language model or parts thereof may be trained explicitly in separate training steps or may be trained implicitly during the same training procedure.
In some embodiments, the determination of expressed words is further based on audio data obtained from a microphone.
For example, the audio data may represent a sound generated by the user when expressing the words, and the expressed words may be determined based on both the event-based visual data and the obtained audio data to increase a robustness of the determination of expressed words. The determination of expressed words may be based on the audio data if the audio data and the event-based visual data are correlated (e.g., if a correlation between the audio data and the eventbased visual data exceeds a predetermined threshold), and may not be based on the audio data if the audio data and the event-based visual data are not correlated (e.g., if a correlation between the audio data and the event-based visual data does not exceed a predetermined threshold). The determination may be based on the audio data during a training, when the circuitry is trained with event-based visual data of, e.g., a certain user. The determination of expressed words may be based on audio data that correspond to a pronunciation in accordance with the language model, and may also be based on audio data that do not correspond to the language model.
For example, even a mute may generate a sound when expressing words by lip movements. Such a sound may aid a determination of words expressed by the mute.
Determining expressed words based on audio data in addition to event-based visual data may provide a performance advantage over determining expressed words purely based on event-based visual data. The determining of expressed words based on audio data in addition to event-based visual data may be performed by an artificial neural network trained accordingly.
In some embodiments, the outputting of the determined expressed words includes reducing, in audio data obtained from a microphone, noise that does not correspond to the determined expressed words; and outputting the audio data in which the noise has been reduced.
For example, an artificial voice that pronounces the expressed words may be generated and compared to the audio data, and frequency components of the audio data that are not correlated with the artificial voice may be reduced or cancelled. Thus, noise in the audio data may be attenuated or cancelled based on the determined expressed words.
Generating the artificial voice, recognizing frequency components of the audio data that are correlated or that are not correlated with the artificial voice and/or reducing frequency components of the audio data that are not correlated with the artificial voice may be performed by an artificial neural network trained accordingly, e.g. based on gradient descent, backpropagation and/or reinforcement learning.
In some embodiments, the outputting of the determined expressed words includes amplifying, in audio data obtained from a microphone, a voice that corresponds to the determined expressed words; and outputting the audio data in which the voice has been amplified.
For example, an artificial voice that pronounces the expressed words may be generated and compared to the audio data, and frequency components of the audio data that are correlated with the artificial voice may be amplified. Thus, speech in the audio data may be amplified based on the determined expressed words.
Generating the artificial voice, recognizing frequency components of the audio data that are correlated or that are not correlated with the artificial voice and/or amplifying frequency components of the audio data that are not correlated with the artificial voice may be performed by an artificial neural network trained accordingly, e.g. based on gradient descent, backpropagation and/or reinforcement learning.
Some embodiments pertain to a method for visual speech processing that includes obtaining event-based visual data of a user obtained from an event-based vision sensor; determining expressed words based on the obtained event-based visual data; and outputting the determined expressed words. The method may be configured corresponding to the processing performed by the circuitry described above, and all features described with reference to the processing performed by the circuitry may also be features of the method.
The methods as described herein are also implemented in some embodiments as a computer program causing a computer and/or a processor to perform the method, when being carried out on the computer and/or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.
Returning to Fig. 1, Fig. 1 illustrates a block diagram of a device 1 with a circuitry 2 for visual speech processing and receiving user feedback according to an embodiment. The device 1 includes the circuitry 2 for visual speech processing, Central Processing Unit (CPU) 3, a storage unit 4 and an input/output (I/O) unit 5.
The circuitry 2 includes a determination unit 6, a presentation unit 7 and an adjustment unit 8.
The determination unit 6 is configured to determine expressed words expressed by a user based on visual data of the user. The determination unit 6 includes a mouth region detection unit 9, a lip movement identification unit 10, a syllable data determination unit 11, a face expression detection unit 12, an audio data obtaining unit 13 and an expressed words generation unit 14.
The determination unit 6 receives visual data of the user via the I/O unit 5. The visual data of the user are acquired with an event-based vision sensor when the user is expressing words and include a plurality of events of a face (including a mouth region) of the user.
The determination unit 6 provides the visual data of the user to the mouth region detection unit 9, to the lip movement identification unit 10 and to the face expression detection unit 12.
The mouth region detection unit 9 receives the visual data of the user and detects, based on an artificial neural network, a region-of-interest (ROI) in the visual data that corresponds to the mouth region of the user. The mouth region detection unit 9 generates ROI data that indicate the detected ROI and provides the ROI data to the lip movement identification unit 10.
The lip movement identification unit 10 receives the visual data of the user and the ROI data. The lip movement identification unit 10 identifies a lip movement in the mouth region of the user, based on motion vectors that are based on correlating events of the visual data within the ROI detected by the mouth region detection unit 9. The lip movement identification unit 10 generates lip movement data that indicate the identified lip movement and provides the lip movement data to the syllable data determination unit 11.
The syllable data determination unit 11 receives the lip movement data and determines syllable data based on the identified lip movement indicated by the lip movement data. The syllable data indicate a sequence of syllables that correspond to the identified lip movement. The syllable data also indicate alternative syllables in the sequence of syllables where the syllable data determination unit 11 cannot determine a syllable unambiguously based on the lip movement data. The syllable data determination unit 11 then provides the syllable data to the expressed words generation unit 14.
The face expression detection unit 12 receives the visual data of the user and detects, in the visual data, a face expression of the user. The face expression detection unit 12 identifies an emotion represented by the face expression, generates face expression data that indicates the detected face expression of the user and the identified emotion and provides the face expression data to the expressed words generation unit 14.
The audio data obtaining unit 13 obtains audio data from a microphone. The audio data are captured by the microphone when the user expresses words. The audio data obtaining unit 13 identifies features in the audio data and generates feature data that indicate the identified features of the audio data. The audio data obtaining unit 13 provides the feature data to the expressed words generation unit 14.
The expressed words generation unit 14 receives the syllable data, the face expression data and the feature data. The expressed words generation unit 14 generates expressed words based on an artificial neural network and, for this purpose, provides the syllable data, the face expression data and the feature data to the artificial neural network. The artificial neural network solves ambiguities in the syllable data (i.e., alternative syllables where the syllable data determination unit 11 cannot unambiguously determine a syllable) based on the feature data and on the face expression data. The artificial neural network selects words from a word list of a language model based on the syllable data, on the feature data and on the facial expression data, and inflects the selected words according to grammatical rules of the language model. The expressed words generation unit 14 then outputs the selected (and inflected) words as generated expressed words.
The determination unit 6 provides the expressed words generated by the expressed words generation unit 14 as determined expressed words to the presentation unit 7.
The presentation unit 7 presents the determined expressed words via the I/O unit 5 to the user and requests feedback from the user regarding the determined expressed words. The presentation unit 7 includes a lip movement recommendation unit 15 that generates a recommendation for an improved lip movement that allows determining the syllable data with less ambiguity than the lip movement identified by the lip movement identification unit 10 based on the visual data. The presentation unit 7 presents the recommendation for the improved lip movement to the user via the I/O unit 5.
The adjustment unit 8 receives, via the I/O unit 5, user feedback regarding the determined expressed words, identifies at least one of a label, a loss and a fitness for the artificial neural network of the expressed words generation unit 14 and changes a value of a parameter of the artificial neural network of the expressed words generation unit 14 in accordance with the user feedback. The adjustment unit 8 also identifies in the feedback of the user a correction of at least one of the determined expressed words and changes a value of a parameter of the artificial neural network of the expressed words generation unit 14 in accordance with the user feedback.
Thus, the adjustment unit 8 adjusts the determination of expressed words based on the user feedback.
The CPU 3 controls the device 1 and the circuitry 2. The storage unit 4 includes a memory and a non-volatile storage. The storage unit 4 stores instructions for execution by the CPU 3. The storage unit 4 also stores the word list of the language model and parameters of the artificial neural network of the expressed words generation unit 14. The I/O unit 5 includes a humanmachine interface with presenting means for presenting the determined expressed words to the user and with input means for receiving the feedback from the user. The I/O unit 5 further includes a network interface for communicating with a server in a communication network.
Figures 2a, 2b and 2c provide examples of the presentation unit 7 of Fig. 1.
Fig. 2a illustrates a block diagram of a presentation unit 7a for displaying the determined expressed words as text according to an embodiment. The presentation unit 7a is an example of the presentation unit 7 of Fig. 1 and includes a text rendering unit 16 that renders the determined expressed words as text. The presentation unit 7a then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 5 to display the rendered text to the user.
Fig. 2b illustrates a block diagram of a presentation unit 7b for displaying at least one icon corresponding to the determined expressed words according to an embodiment. The presentation unit 7b is an example of the presentation unit 7 of Fig. 1. The presentation unit 7b retrieves, via the network interface of the I/O unit 5, one or more icons from a server in the communication network. The number and content of the icons retrieved is based on a number and meaning of the determined expressed words and is chosen such that the retrieved icons correspond to the determined expressed words. The presentation unit 7b includes an icon rendering unit 17 that renders the retrieved one or more icons. The presentation unit 7b then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 5 to display the rendered one or more icons to the user.
Fig. 2c illustrates a block diagram of a presentation unit 7c for outputting an artificial voice corresponding to the determined expressed words according to an embodiment. The presentation unit 7c is an example of the presentation unit 7 of Fig. 1. The presentation unit 7c includes an artificial voice generation unit 18. The artificial voice generation unit 18 generates an artificial voice that utters the determined expressed words. The presentation unit 7c causes a loudspeaker that is a presenting means of the human-machine interface included in the I/O unit 5 to output the generated artificial voice to the user.
Note that, in some embodiments, the presentation unit 7 of Fig. 1 includes two or all three of the text rendering unit 16 of Fig. 2a, the icon rendering unit 17 of Fig. 2b and the artificial voice generation unit 18 of Fig. 2c. Thus, in some embodiments, the presentation unit 7 presents, to the user, the determined expressed words as two or all three of a text, one or more icons and an artificial voice.
Note that, in some embodiments, the face expression detection unit 12 is not provided in the determination unit 6 of Fig. 1 and/or the audio data obtaining unit 13 is not provided in the determination unit 6 of Fig. 1 and/or the lip movement recommendation unit 15 is not provided in the presentation unit 7 of Fig. 1. Thus, in some embodiments, the determination of expressed words does not include detecting, in the visual data, a face expression of the user and generating the expressed words based on the detected face expression and/or the determination of expressed words does not include obtaining audio data from a microphone and generating the expressed words based on the audio data and/or the presenting of the determined expressed words does not include recommending an improved lip movement.
Note that, in some embodiments, the event-based vision sensor that captures the visual data of the user is included in the device 1, and in some embodiments, the event-based vision sensor that captures the visual data of the user is provided separately from the device 1 and the determination unit 6 obtains the visual data from the event-based vision sensor via the network interface included in the I/O unit 5.
Note that the event-based vision sensor is only provided as an example. In some embodiments, any suitable camera capable of visual data of a user is used instead of or in addition to the eventbased vision sensor. Note that, in some embodiments, the human-machine interface with the presenting means and with the input means is not provided in the I/O unit 5 but separately from the device 1. In such a case, the determined expressed words (and a recommendation for an improved lip movement) are transmitted via the network interface of the I/O unit 5 to the presenting means and the feedback of the user is received via the network interface of the I/O unit 5 from the input means. For example, the device 1 that includes the circuitry 2 may be configured as a server and the human-machine interface with the presenting means and with the input means may be provided on a communication device of a user that communicates with the server via a communication network.
Fig. 3 illustrates a flow diagram of a method 30 for visual speech processing and receiving user feedback according to an embodiment. The method 30 is an example of a method that is executed by the circuitry 2 of Fig. 1.
The method 30 includes, at S31, a determination performed by the determination unit 6 of Fig. 1. The determination at S31 determines expressed words based on visual data of a user that are acquired with an event-based vision sensor when the user is expressing words and that include a plurality of events of a face (including a mouth region) of the user. The determination at S31 includes a mouth region detection at S32, a lip movement identification at S33, a syllable data determination at S34, a face expression detection at S35, obtaining audio data at S36 and an expressed words generation at S37.
The mouth region detection at S32 is performed by the mouth region detection unit 9 of Fig. 1 and detects, based on an artificial neural network, a region-of-interest (ROI) in the visual data that corresponds to the mouth region of the user.
The lip movement identification at S33 is performed by the lip movement identification unit 10 of Fig. 1 and identifies a lip movement in the mouth region of the user, based on motion vectors that are based on correlating events of the visual data with the ROI detected by the mouth region detection at S32.
The syllable data determination at S34 is performed by the syllable data determination unit 11 of Fig. 1 and determines syllable data based on the lip movement identified by the lip movement identification at S33. The syllable data indicate a sequence of syllables that correspond to the identified lip movement as well as alternative syllables in the sequence of syllables where the syllable data determination at S33 cannot determine a syllable unambiguously based on the identified lip movement. The face expression detection at S35 is performed by the face expression detection unit 12 of Fig. 1, detects, in the visual data, a face expression of the user and identifies an emotion represented by the detected face expression.
The obtaining of audio data at S36 is performed by the audio data obtaining unit 13 of Fig. 1 and obtains audio data from a microphone. The audio data are captured by the microphone when the user expresses words. The obtaining of audio data further identifies features in the audio data.
The expressed words generation at S37 is performed by the expressed words generation unit 24 of Fig. 1 and generates expressed words based on an artificial neural network. For this purpose, the expressed words generation provides the syllable data, the face expression detected at S35 with the emotion identified at S35 and the audio data obtained at S36 with the features identified at S36 to the artificial neural network. The artificial neural network solves ambiguities in the syllable data (i.e., alternative syllables where the syllable data determination at S34 cannot unambiguously determine a syllable) based on the audio data and the features identified therein as well as on the face expression and the emotion identified based thereon. The artificial neural network selects words from a word list of a language model based on the syllable data, on the audio data and the features identified therein as well as on the face expression and the emotion identified based thereon, and inflects the selected words according to grammatical rules of the language model.
The words selected (and inflected) by the expressed words generation at S37 are the expressed words determined by the determination at S31.
At S38, a presentation is performed by the presentation unit 7 of Fig. 1. The presentation presents the determined expressed words via the I/O unit 5 of Fig. 1 to the user and requests feedback from the user regarding the determined expressed words.
The presentation includes a lip movement recommendation at S39 that is performed by the lip movement recommendation unit 15 of Fig. 1 and generates a recommendation for an improved lip movement that allows determining the syllable data with less ambiguity than the lip movement identified by the lip movement identification at S33 based on the visual data. The presentation presents the recommendation for the improved lip movement to the user via the VO unit 5 of Fig. 1.
At S40, an adjustment is performed by the adjustment unit 8 of Fig. 1. The adjustment receives, via the I/O unit 5 of Fig. 1, user feedback regarding the determined expressed words, identifies at least one of a label, a loss and a fitness for the artificial neural network of the expressed words generation of S37, and changes a value of a parameter of the artificial neural network of the expressed words generation of S37 in accordance with the user feedback. The adjustment also identifies in the feedback of the user a correction of at least one of the determined expressed words and changes a value of a parameter of the artificial neural network of the expressed words generation of S37 in accordance with the user feedback. Thus, the adjustment adjusts the determination of expressed words based on the user feedback.
Figures 4a, 4b and 4c illustrate examples of the presentation at S38 of Fig. 3.
Fig. 4a illustrates a block diagram of a presentation at S38a for displaying the determined expressed words as text according to an embodiment. The presentation at S38a is an example of the presentation at S38 of Fig. 3 and is performed by the presentation unit 7a of Fig. 2a. The presentation at S38a includes rendering of text at S41 performed by the text rendering unit 16 of Fig. 2a. The rendering of text at S41 renders the determined expressed words as text. The presentation at S38a then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 5 to display the rendered text to the user.
Fig. 4b illustrates a block diagram of a presentation at S38b for displaying at least one icon corresponding to the determined expressed words according to an embodiment. The presentation at S38b is an example of the presentation at S38 of Fig. 3 and is performed by the presentation unit 7b of Fig. 2b. The presentation at S38b retrieves, via the network interface of the I/O unit 5, one or more icons from a server in the communication network. The number and content of the icons retrieved is based on a number and meaning of the determined expressed words and is chosen such that the retrieved icons correspond to the determined expressed words. The presentation at S38b includes rendering of icons at S42 performed by the icon rendering unit 17 of Fig. 2b. The rendering of icons at S42 renders the retrieved one or more icons. The presentation at S38b then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 5 to display the rendered one or more icons to the user.
Fig. 4c illustrates a block diagram of a presentation at S38c for outputting an artificial voice corresponding to the determined expressed words according to an embodiment. The presentation at S38c is an example of the presentation S38 of Fig. 3 and is performed by the presentation unit 7c of Fig. 2c. The presentation at S38c includes an artificial voice generation at S43 performed by the artificial voice generation unit 18 of Fig. 2c. The artificial voice generation at S43 generates an artificial voice that utters the determined expressed words. The presentation at S38c then causes a loudspeaker that is a presenting means of the human-machine interface included in the I/O unit 5 to output the generated artificial voice to the user. Note that, in some embodiments, the presentation at S38 of Fig. 3 includes two or all three of the rendering of text at S41 of Fig. 4a, the rendering of icons at S42 of Fig. 4b and the artificial voice generation at S43 of Fig. 4c. Thus, in some embodiments, the presentation at S38 presents, to the user, the determined expressed words as two or all three of a text, one or more icons and an artificial voice.
Note that, in some embodiments, the face expression detection at S35 is not included in the determination at S31 of Fig. 3 and/or the obtaining of audio data at S36 is not included in the determination at S31 of Fig. 3 and/or the lip movement recommendation at S39 is not included in the presentation at S38 of Fig. 3. Thus, in some embodiments, the determination of expressed words does not include detecting, in the visual data, a face expression of the user and generating the expressed words based on the detected face expression and/or the determination of expressed words does not include obtaining audio data from a microphone and generating the expressed words based on the audio data and/or the presenting of the determined expressed words does not include recommending an improved lip movement.
Fig. 5 illustrates a block diagram of a device 100 with a circuitry 102 for visual speech processing based on event-based visual data according to an embodiment. The device 100 includes an event-based vision sensor 101, the circuitry 102, a Central Processing Unit (CPU) 103, a storage unit 104 and an Input/Output (I/O) unit 105.
The event-based vision sensor 101 captures event-based visual data of a user when the user is expressing words by speaking. The event-based visual data indicate events that represent changes in a face (including a mouth region) of the user. The event-based vision sensor 101 provides the event-based visual data to the circuitry 102.
The circuitry 102 includes a visual data obtaining unit 106, a determination unit 107 and an output unit 108.
The visual data obtaining unit 106 obtains the event-based visual data from the event-based vision sensor 101 and provides the event-based visual data to the determination unit 107.
The determination unit 107 determines expressed words based on the obtained event-based visual data. The determination unit 107 includes a lip movement identification unit 109, a syllable data determination unit 110, a face expression detection unit 111, an audio data obtaining unit 112 and an expressed words generation unit 113. The determination unit 107 receives the event-based visual data form the visual data obtaining unit 106 and provides the event-based visual data to the lip movement identification unit 109. The lip movement identification unit 109 receives the event-based visual data and identifies, in the event-based visual data, a lip movement of the user. The lip movement identification unit 109 determines, in the event-based visual data, a region-of-interest (ROI) of a mouth region of the user and, similar to the lip movement identification unit 10 of Fig. 1, identifies the lip movement of the user based on events, of the event-based visual data, from the ROI of the mouth region of the user. The lip movement identification unit 109 generates lip movement data that indicate the identified lip movement of the user and provides the lip movement data to the syllable data determination unit 110.
The syllable data determination unit 110 receives the lip movement data and determines syllable data based on the identified lip movement indicated by the lip movement data. The syllable data determination unit 110 is configured similar to the syllable data determinant unit 11 of Fig. 1. Additionally, the syllable data determination unit 110 includes a viseme identification unit 114 that identifies one or more visemes based on the identified lip movement and determines the syllable data based on the at least one viseme. The syllable data determination unit 110 retrieves syllables corresponding to the identified visemes from a viseme database that associates visemes with syllables based on human physiology. The syllable data determination unit 110 generates the syllable data based on the syllables retrieved from the viseme database. The syllable data determination unit 110 provides the syllable data to the expressed words generation unit 113.
The face expression detection unit 111 is configured similar to the face expression detection unit 12 of Fig. 1. The face expression detection unit detects, in the event-based visual data, a face expression of the user, identifies an emotion that corresponds to the detected face expression, generates face expression data that indicate the detected face expression and the identified emotion and provides the face expression data to the expressed words generation unit 113.
The audio data obtaining unit 112 is configured similar to the audio data obtaining unit 13 of Fig. 1. The audio data obtaining unit 112 obtains, from a microphone, audio data of the user when the user is expressing words by speaking and identifies features in the audio data. The audio data obtaining unit 112 generates feature data that indicate the identified features and provides the feature data to the expressed words generation unit 113.
The expressed words generation unit 113 is configured similar to the expressed words generation unit 14 of Fig. 1. The expressed words generation unit 113 receives the syllable data, the face expression data and the feature data and generates expressed words based on the syllable data, the face expression data and the feature data and based on an artificial neural network. The expressed words generation unit 113 selects words from a word list of a language model based on the syllable data, on the face expression data and on the feature data, and inflects the selected words according to grammatical rules of the language model. The expressed words generation unit 14 then outputs the selected (and inflected) words as generated expressed words.
The determination unit 107 provides the expressed words generated by the expressed words generation unit 113 as determined expressed words to the output unit 108.
The output unit 108 outputs the determined expressed words via the I/O unit 105.
The CPU 103, the storage unit 104 and the I/O unit 105 are configured corresponding to the CPU 3, the storage unit 4 and the I/O unit 5, respectively, of Fig. 1.
Figures 6a and 6b illustrate examples of the lip movement identification unit 109 of Fig. 5.
Fig. 6a illustrates a block diagram of a lip movement identification unit 109a for identifying a lip movement based on generating a full image of the mouth region of the user according to an embodiment. The lip movement identification unit 109a is an example of the lip movement identification unit 109 of Fig. 5. The lip movement identification unit 109a includes a full image generation unit 115 that generates, based on the event-based visual data of the user, a full image of the mouth region of the user, i.e., of the ROI determined by the lip movement identification unit 109a. The lip movement identification unit 109a then identifies the lip movement based on the generated full image of the mouth region of the user.
Fig. 6b illustrates a block diagram of a lip movement identification unit 109b for identifying a lip movement based on motion vectors according to an embodiment. The lip movement identification unit 109b is an example of the lip movement identification unit 109 of Fig. 5. The mouth event identification unit 116 includes a mouth event identification unit 116 and a motion vector determination unit 117. The mouth event identification unit 116 identifies, as mouth events, events of the event-based visual data from the mouth region of the user, i.e., from the ROI determined by the lip movement identification unit 109b. The motion vector determination unit 117 determines motion vectors based on correlations between mouth events and, thus, based on the identified events of the mouth region. The lip movement identification unit 109b identifies the lip movement based on the determined motion vectors.
Figures 7a to 7f illustrate examples of the output unit 108 of Fig. 5.
Fig. 7a illustrates a block diagram of an output unit 108a for displaying the determined expressed words as text according to an embodiment. The output unit 108a is an example of the output unit 108 of Fig. 5 and is configured similar to the presentation unit 16 of Fig. 2a. The output unit 108a includes a text rendering unit 118 that renders the determined expressed words as text. The output unit 108a then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 105 to display the rendered text to the user.
Fig. 7b illustrates a block diagram of an output unit 108b for displaying at least one icon corresponding to the determined expressed words according to an embodiment. The output unit 108b is an example of the output unit 108 of Fig. 5 and is configured similar to the presentation unit 7b of Fig. 2b. The output unit 108b retrieves, via the network interface of the I/O unit 105, one or more icons from a server in the communication network. The number and content of the icons retrieved is based on a number and meaning of the determined expressed words and is chosen such that the retrieved icons correspond to the determined expressed words. The output unit 108b includes an icon rendering unit 119 that renders the retrieved one or more icons. The output unit 108b then causes a screen that is a presenting means of the humanmachine interface included in the I/O unit 105 to display the rendered one or more icons to the user.
Fig. 7c illustrates a block diagram of an output unit 108c for outputting an artificial voice corresponding to the determined expressed words according to an embodiment. The output unit 108c is an example of the presentation unit 108 of Fig. 5 and is configured similar to the presentation unit 7c of Fig. 2c. The output unit 108c includes an artificial voice generation unit 120. The artificial voice generation unit 120 generates an artificial voice that utters the determined expressed words. The output unit 108c causes a loudspeaker that is a presenting means of the human-machine interface included in the I/O unit 105 to output the generated artificial voice to the user.
Fig. 7d illustrates a block diagram of an output unit 108d for transmitting the determined expressed words to a communication device of another user according to an embodiment. The output unit 108d is an example of the output unit 108 of Fig. 5 and includes a transmission preparation unit 121. The transmission preparation unit 121 prepares the determined expressed words for a transmission to the communication device of the other user. The transmission preparation unit 121 encodes the determined expressed words according to a predetermined encoding and inserts the encoded expressed words in a packet of a predetermined protocol. The output unit 108d then transmits the packet via the network interface included in the I/O unit 105 to the communication device of the other user.
Fig. 7e illustrates a block diagram of an output unit 108e for reducing noise according to an embodiment. The output unit 108e is an example of the output unit 108 of Fig. 5 and includes a noise reduction unit 122. The noise reduction unit 122 reduces, in the audio data obtained by the audio data obtaining unit 112 from a microphone, noise that does not correspond to the determined expressed words. The output unit 108e then causes a loudspeaker of the humanmachine interface included in the I/O unit 105 to output a sound according to the audio data in which the noise has been reduced.
Fig. 7f illustrates a block diagram of an output unit 108f for amplifying a voice according to an embodiment. The output unit 108f is an example of the output unit 108 of Fig. 5 and includes a voice amplification unit 123. The voice amplification unit 123 amplifies, in the audio data obtained by the audio data obtaining unit 112 from a microphone, a voice that corresponds to the determined expressed words. The output unit 108f then causes a loudspeaker of the humanmachine interface included in the I/O unit 105 to output a sound according to the audio data in which the voice has been amplified.
Note that, in some embodiments, the face expression detection unit 111 is not provided in the determination unit 107 of Fig. 5 and/or the audio data obtaining unit 112 is not provided in the determination unit 107 of Fig. 5. Thus, in some embodiments, the determination of expressed words does not include detecting, in the visual data, a face expression of the user and generating the expressed words based on the detected face expression and/or the determination of expressed words does not include obtaining audio data from a microphone and generating the expressed words based on the audio data.
Note that, in some embodiments, the output unit 108 of Fig. 5 includes any combination of the output units 108a to 108f.
Note that, in some embodiments, the event-based vision sensor 101 that captures the event-based visual data of the user is provided separately from the device 100 and the visual data obtaining unit 106 obtains the event-based visual data from the event-based vision sensor 101 via the network interface included in the I/O unit 105.
Note that, in some embodiments, the human-machine interface with the presenting means is not provided in the I/O unit 105 but separately from the device 100. In such a case, the determined expressed words are transmitted by the output unit 108 via the network interface of the I/O unit 105 to the presenting means. For example, the device 100 that includes the circuitry 102 may be configured as a server and the human-machine interface with the presenting means may be provided on a communication device of a user that communicates with the server via a communication network. Note that the microphone from which the audio data obtaining unit 112 obtains the audio data is, in some embodiments, provided in the device 100 and is, in some embodiments, provided separately from the device 100.
Fig. 8 illustrates a flow diagram of a method 150 for visual speech processing based on eventbased visual data according to an embodiment. The method 150 is an example of a method that is executed by the circuitry 102 of Fig. 5.
The method 150 includes, at S 151, an obtaining of event-based visual data that is performed by the visual data obtaining unit 106 of Fig. 5. The obtaining of event-based visual data at SI 51 obtains event-based visual data of a user from the event-based vision sensor 101.
At SI 52, a determination is performed by the determination unit 107 of Fig. 5. The determination at SI 52 determines expressed words based on the event-based visual data obtained by the obtaining of event-based visual data at S151. The determination includes a lip movement identification at SI 53, a syllable data determination at SI 54, a face expression detection at SI 55, an obtaining of audio data at SI 56 and an expressed words generation at SI 57.
At SI 53, the lip movement identification is performed by the lip movement identification unit 109 of Fig. 5. The lip movement identification identifies, in the event-based visual data, a lip movement of the user and determines, in the event-based visual data, a region-of-interest (ROI) of a mouth region of the user.
At SI 54, the syllable data determination is performed by the syllable data determination unit 110 of Fig. 5. The syllable data determination determines syllable data based on the lip movement identified by the lip movement at S153. The syllable data determination includes, at S158, a viseme identification that is performed by the viseme identification unit 114 of Fig. 5. The viseme identification identifies, based on the identified lip movement, at least one viseme, and the syllable data determination determines the syllable data based on the at least one viseme.
At SI 55, the face expression detection is performed by the face expression detection unit 111 of Fig. 5. The face expression detection detects, in the event-based visual data, a face expression of the user and identifies an emotion corresponding to the face expression.
At S156, the obtaining of audio data is performed by the audio data obtaining unit 112 of Fig. 5. The obtaining of audio data obtains audio data from a microphone when the user expressed words by speaking. The obtaining of audio data identifies features in the audio data.
At SI 57, the expressed words generation is performed by the expressed words generation unit 113 of Fig. 5. The expressed words generation generates expressed words based on the syllable data determined by the syllable data determination at SI 54, on the face expression detected by the face expression detection at SI 55 and the identified corresponding emotion, and on the audio data obtained by the audio data obtaining at SI 56 and the features identified therein. The expressed words generation unit generates expressed words based on an artificial neural network. The expressed words generation selects, based on the determined syllable data, on the detected face expression and the corresponding emotion and on the obtained audio data and the features identified therein, words from a word list of a language model and inflects the selected words based on grammatical rules of the language model.
The determination determines the selected (and inflected) words as determined expressed words of the user.
At SI 59, an output is performed by the output unit 108 of Fig. 5. The output outputs the determined expressed words.
Figures 9a and 9b illustrate examples of the lip movement identification at S153 of Fig. 8.
Fig. 9a illustrates a block diagram of a lip movement identification at S153a based on generating a full image according to an embodiment. The lip movement identification at SI 53 a is an example for the lip movement identification at SI 53 of Fig. 8 and is performed by the lip movement identification unit 109a of Fig. 6a. The lip movement identification at SI 53a includes a full image generation at SI 60 that is performed by the full image generation unit 115 of Fig. 6a. The full image generation generates, based on the event-based visual data and on the ROI of the mouth region of the user determined by the lip movement identification at S 153a, a full image of the mouth region of the user. The lip movement identification then identifies the lip movement of the user based on the generated full image of the mouth region of the user.
Fig. 9b illustrates a block diagram of a lip movement identification at SI 53b based on motion vectors according to an embodiment. The lip movement identification at SI 53b is an example of the lip movement identification at SI 53 of Fig. 8 and is performed by the lip movement identification unit 109b of Fig. 6b. The lip movement identification at SI 53b includes a mouth event identification at SI 61 and a motion vector determination at SI 62. The mouth event identification at SI 61 is performed by the mouth event identification unit 116 of Fig. 6b and identifies events of the event-based visual data from the ROI of the mouth region determined by the lip movement identification at SI 53b as mouth events. The motion vector determination at SI 62 is performed by the motion vector determination unit 117 of Fig. 6b and determines motion vectors based on the identified mouth events. The lip movement identification at SI 53b then identifies the lip movement based on the determined motion vectors. Figures 10a to lOf illustrate examples of the output at S159 of Fig. 8.
Fig. 10a illustrates a block diagram of an output at SI 59a for displaying the determined expressed words as text according to an embodiment. The output at SI 59a is an example of the output at S159 of Fig. 8 and is performed by the output unit 108a of Fig. 7a. The output at S159a includes rendering of text at SI 63 performed by the text rendering unit 118 of Fig. 7a. The rendering of text at SI 63 renders the determined expressed words as text. The output at SI 59a then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 105 to display the rendered text to the user.
Fig. 10b illustrates a block diagram of an output at SI 59b for displaying at least one icon corresponding to the determined expressed words according to an embodiment. The output at S159b is an example of the output at S159 of Fig. 8 and is performed by the output unit 108b of Fig. 7b. The output at SI 59b retrieves, via the network interface of the I/O unit 105, one or more icons from a server in the communication network. The number and content of the icons retrieved is based on a number and meaning of the determined expressed words and is chosen such that the retrieved icons correspond to the determined expressed words. The output at SI 59b includes rendering of icons at SI 64 performed by the icon rendering unit 119 of Fig. 7b. The rendering of icons at SI 64 renders the retrieved one or more icons. The output at SI 59b then causes a screen that is a presenting means of the human-machine interface included in the I/O unit 105 to display the rendered one or more icons to the user.
Fig. 10c illustrates a block diagram of an output at SI 59c for outputting an artificial voice corresponding to the determined expressed words according to an embodiment. The output at S159c is an example of the output S159 of Fig. 8 and is performed by the output unit 108c of Fig. 7c. The output at S159c includes an artificial voice generation at S165 performed by the artificial voice generation unit 120 of Fig. 7c. The artificial voice generation at SI 65 generates an artificial voice that utters the determined expressed words. The output at SI 59c then causes a loudspeaker that is a presenting means of the human-machine interface included in the I/O unit 105 to output the generated artificial voice to the user.
Fig. lOd illustrates a block diagram of an output at S159d for transmitting the determined expressed words to a communication device of another user according to an embodiment. The output at S159d is an example for the output at SI 59 of Fig. 8 and is performed by the output unit 108d of Fig. 7d. The output at S159d includes a transmission preparation at SI 66 that is performed by the transmission preparation unit 121 of Fig. 7d. The transmission preparation at SI 66 prepares the determined expressed words for a transmission to the communication device of the other user. The transmission preparation encodes the determined expressed words according to a predetermined encoding and inserts the encoded expressed words in a packet of a predetermined protocol. The output at S159d then transmits the packet via the network interface included in the I/O unit 105 to the communication device of the other user.
Fig. lOe illustrates a block diagram of an output at S159e for reducing noise according to an embodiment. The output at S159e is an example of the output at SI 59 of Fig. 8 and is performed by the output unit 108e of Fig. 7e. The output at S159e includes a noise reduction at S167 that is performed by the noise reduction unit 122 of Fig. 7e. The noise reduction reduces, in the audio data obtained by the obtaining of audio data at S156 of Fig. 8 from a microphone, noise that does not correspond to the determined expressed words. The output at S159e then causes a loudspeaker of the human-machine interface included in the I/O unit 105 to output a sound according to the audio data in which the noise has been reduced.
Fig. lOf illustrates a block diagram of an output at S159f for amplifying a voice according to an embodiment. The output at S159f is an example of the output at SI 59 of Fig. 8 and is performed by the output unit 108f of Fig. 7f. The output at S159f includes a voice amplification at S168 that is performed by the voice amplification unit 123 of Fig. 7f. The voice amplification amplifies, in the audio data obtained by the obtaining of audio data at S156 of Fig. 8 from a microphone, a voice that corresponds to the determined expressed words. The output at S159f then causes a loudspeaker of the human-machine interface included in the I/O unit 105 to output a sound according to the audio data in which the voice has been amplified.
Note that, in some embodiments, the output 159 of Fig. 8 includes any combination of the output 159a to 159f.
Note that, in some embodiments, the face expression detection at SI 55 is not included in the determination at S152 of Fig. 8 and/or the obtaining of audio data at S156 is not included in the determination at SI 52 of Fig. 8. Thus, in some embodiments, the determination of expressed words does not include detecting, in the visual data, a face expression of the user and generating the expressed words based on the detected face expression and/or the determination of expressed words does not include obtaining audio data from a microphone and generating the expressed words based on the audio data.
Fig. 11 illustrates a block diagram of a general application example according to an embodiment.
The device 200 is an example of the device 1 of Fig. 1 and of the device 100 of Fig. 5. The device 200 includes an event-based vision sensor 201, a circuitry 202 and a loudspeaker 203. The event-based vision sensor 201 captures events of a face (including a mouth region) of a user 204 when the user expresses words by speaking. During speaking, the user 204 moves his lips and, thus, expressed various visemes. The event-based vision sensor 201 records transitions between the visemes in event-based visual data 205 of the user 204 and provides the event-based visual data 205 to the circuitry 202.
The circuitry 202 is an example of the circuitry 2 of Fig. 1 and of the circuitry 102 of Fig. 5. The circuitry 202 identifies, based on the event-based vision data 205, lip movements of the user 204 and corresponding visemes, determines expressed words of the user 204 based on the identified lip movements and visemes, and generates audio data 206 of an artificial voice that utters the determined expressed words. The circuitry 202 then provides the audio data 206 to the loudspeaker 203.
The loudspeaker 203 outputs a sound that corresponds to the audio data 206 and, thus, to the expressed words of the user 204.
Fig. 12 illustrates a block diagram of an exemplary application for speech enablement for mutes according to an embodiment. The circuitry 300 enables a mute user 301 to communicate with a hearing user 302 by expressing words through lip movements.
The mute user 301 cannot move his lips as if he was speaking (i.e., according to audible speech) because he did not learn speaking due to a deafness. However, the circuitry 300 does not require lip movements that correspond to audible speech. Instead, the circuitry 300 can translate any lip movement to an alphabet and the mute user 301 can learn his own language.
The circuitry 300 also provides a training app that suggests and teaches an optimal lip language (including lip movements) for the mute user 301. The lip movements of the optimal lip language are optimized in that visemes of the optimal lip language are robust to detect and reduce ambiguities in identifying syllables based on the visemes.
The circuitry 300 is an example of the circuitry 2 of Fig. 1, and includes an expressed words determination unit 310, a text rendering unit 311 and an artificial voice generation unit 312.
The expressed words determination unit 310 receives, from a user input unit 320 that includes a touchscreen, user feedback data 321, receives, from a visual sensor 322 that includes an eventbased vision sensor, event-based visual data 323, and receives, from a microphone 324, audio data 325.
The visual sensor 322 captures, as the event-based visual data 323, lip movements of the user 301 when the user 301 is expressing words by lip movements. Likewise, the microphone 324 captures, as the audio data 325, sounds generated by the user 301 when the user 301 is expressing words.
Although the mute user 301 cannot speak, he generates, at least for some lip movements, a sound correlated with the lip movements when expressing words by lip movements. The expressed words determination unit 310 correlates the audio data 325 with the identified lip movement for increasing a robustness of the determination of expressed words of the user 301.
The expressed words determination unit 310 determines, based on the event-based visual data 323 from the visual sensor 322 and on the audio data 325 from the microphone 324, expressed words 313 of the user 301 and provides the determined expressed words 313 to the text rendering unit 311 and to the artificial voice generation unit 312.
The text rendering unit 311 receives the determined expressed words 313, renders the determined expressed words 313 as text 314 and causes the rendered text 314 to be displayed, on a screen, as recognized text 315.
The artificial voice generation unit 312 receives the determined expressed words 313, generates an artificial voice 316 that utters the determined expressed words 313 and causes the generated artificial voice 316 to be output, on a loudspeaker, as synthesized audio 317.
The user 302 reads the recognized text 315 and hears the synthesized audio 317. Thus, the user 302 can understand the words expressed by the mute user 301 by lip movements although the user 302 has not learnt lip reading.
The mute user 301 reads the recognized text 315 and controls whether the determined expressed words 313 correspond to the intended words. The user 301 enters a teacher signal 303 as feedback into the user input unit 320 to correct an incorrect determination of expressed words and to confirm a correct determination of expressed words. The teacher signal 303 includes at least one of a label, a loss and a fitness for a machine learning algorithm performed by the expressed words determination unit 310. The user input unit 320 provides the teacher signal 303 as user feedback data 321 to the expressed words determination unit 310.
The expressed words determination unit 310 then adjusts a parameter of the machine learning algorithm according to the user feedback data 321.
Thus, the user 301 can correct incorrectly determined expressed words by the teacher signal 303. The circuitry 300 can learn a “sign language” using the mouth (including lip movements) which allows for real-time communication of the mute (and deaf) user 301 with the hearing user 302 since the listening user 302 does not need to understand sign language. Key to the disclosure is that the user feedback provided by the user 301 via the teacher signal 303 improves the performance of the expressed words determination unit 310. Thus, the teacher signal 303 improves the determined expressed words 313.
Note that, in some embodiments, the visual sensor 322 includes, instead of or in addition to an event-based vision sensor, a CMOS image sensor or a CCD image sensor. Note also that, in some embodiments, the user input unit 320 includes, instead of or in addition to the touchscreen, any other input means for receiving the teacher signal 303.
Note that dashed lines in Fig. 12 indicate features that are optional and/or alternative.
For example, in some embodiments, the determination of expressed words of the user 301 only uses the visual sensor 322 as input and is performed by a processor that includes the circuitry 300; i.e., in some embodiments, the microphone 324 is not provided and the determination of expressed words of the user 301 is not based on the audio data 325.
For example, in some embodiments, the user 302 only reads the recognized text 315 and does not hear the synthesized audio 317 or, vice versa, the user 302 only hears the synthesized audio 317 and does not read the recognized text 315.
In some embodiments, the user 302 does neither read the recognized text 315 nor hear the synthesized audio 317, but only the mute user 301 reads the recognized text 315. This may be the case when the mute user 301 is alone and trains the circuitry 300 to correctly determine expressed words based on the lip movements of the user 301 or trains the optimal lip language based on the training app provided by the circuitry 300.
In some embodiments, the artificial voice generation unit 312 may be omitted, and the circuitry 300 outputs the determined expressed words 313 only as rendered text 314.
Figures 13a to 13c illustrate examples of algorithms performed by the expressed words determination unit 310 of Fig. 12 for determining expressed words based on an event-based vision sensor.
Fig. 13a illustrates a block diagram of an algorithm for an expressed words determination 330 based on generating an image according to an embodiment. The expressed words determination 330 is performed by the expressed words determination unit 310 of Fig. 12 and includes an events-to-image conversion at S331 and an image-to-words conversion at S332.
The events-to-image conversion at S331 receives, from the event-based vision sensor of the visual sensor 322, event-based visual data 333 representing events and generates an image 334 based on events represented by the event-based visual data 333. The generated image 334 is provided to the image-to-words conversion at S332. The image-to-words conversion 332 determines, based on the image 334, expressed words 335. The expressed words 335 are then provided to the text rendering unit 311 and to the artificial voice generation unit 312.
Fig. 13b illustrates a block diagram of an algorithm for an expressed words determination 340 based on converting motion to words according to an embodiment. The expressed words determination 340 is performed by the expressed words determination unit 310 of Fig. 12 and includes an events-to-mouth ROI conversion at S341, an events-to-motion conversion at S342 and a motion-to-words conversion at S343.
The events-to-mouth ROI conversion at S341 receives, from the event-based vision sensor of the visual sensor 322, event-based visual data 344 representing events, determines a mouth region- of-interest (ROI) in the event-based visual data 344 that corresponds to a mouth region of the user 301, and selects, as mouth events 345, events of the event-based visual data 344 included in the mouth ROI. The mouth events 345 are provided to the event-to-motion conversion at S342. The events-to-motion conversion at S342 determines motion vectors 346 based on the mouth events 345 that indicate lip movements of the user 301. The motion vectors 346 are provided to the motion-to-words conversion at S343. The motion-to-words conversion at S343 determines expressed words 349 based on the motion vectors 346 and on word probabilities 347 provided by a language model 348. The expressed words 349 are then provided to the text rendering unit 311 and to the artificial voice generation unit 312.
Fig. 13c illustrates a block diagram of an algorithm for an expressed words determination 350 based on directly converting events to words according to an embodiment. The expressed words determination 350 is performed by the expressed words determination unit 310 of Fig. 12 and includes an events-to-words conversion at S351.
The events-to-words conversion at S351 receives, from the event-based vision sensor of the visual sensor 322, event-based visual data 352 representing events and converts the events represented by the event-based visual data 352 to expressed words 353 based on an artificial neural network. The expressed words 353 are then provided to the text rendering unit 311 and to the artificial voice generation unit 312.
Note that, in some embodiments, the determination of expressed words based on event-based visual data is based on other algorithms instead of or in addition to the algorithms of Fig. 13a to 13c. Fig. 14 illustrates an application example of a speech enablement for mutes based on a smartphone 360 according to an embodiment. The smartphone 360 includes the circuitry 300 of Fig. 12, and further includes a microphone 361, a visual sensor 362 and a loudspeaker 363.
The microphone 361 is an example of the microphone 324 of Fig. 12, the visual sensor 362 is an example of the visual sensor 322 and the loudspeaker 363 is an example of the loudspeaker that outputs the synthesized audio 317 of Fig. 12.
The mute user 301 is also deaf, so the microphone 361 records the speech 364 of the user 302 (e.g., the sentence: “How are you?”) and translates it to text 365 that is displayed on a screen of the smartphone 360 of the mute user 301. The mute user 301 reads the text 365 (e.g., “How are you?”) and answers with voiceless speech 366 by moving his lips according to a lip language trained by the circuitry 300. The visual sensor 362 of the smartphone 360 captures the voiceless speech 366 of the user 301, and the circuitry 300 translates the voiceless speech 366 of the mute user 301 to audio 367. The loudspeaker 363 outputs the audio 367 (e.g., “I’m fine.”) so that the hearing user 302 can hear and understand the answer.
Thus, the mute user 301 can have a normal conversation with the hearing user 302.
Fig. 15 illustrates an exemplary application of a video call enhancement according to an embodiment. In the exemplary application of Fig. 15, a live caption is created and the voice of a speaker 401 taking a video call in a noisy environment is enhanced.
The speaker 401 is in a noisy environment and participates in a video call with a smartphone 402. The smartphone is an example of the device 100 of Fig. 5. The speaker 401 expresses words as a speech (e.g., “It’s loud here!”). The speech of the speaker 401 is captured by the event-based vision sensor 101 of the smartphone 401 and determines the words expressed by the speaker 401 with his speech, thus translating the speech to a live caption, using only the event-based visual data captured by the event-based vision sensor 101. Thus, the smartphone 402 records voiceless speech. The smartphone 401 transmits the live caption including the determined expressed words together with video of the speaker 401 via a wireless communication connection 403 based on Wi-Fi to a communication network 404.
A communication device 407 of another user receives, via a communication connection 406 to the communication network 404, the video and live caption. The communication device 407 displays the received video and outputs the expressed words of the received live caption.
Thus, the speaker 401 can participate in a video call although he is in a noisy environment where his voice cannot be distinguished from surrounding noise. Fig. 16a illustrates a block diagram of an algorithm for video call voice enhancement based on fusion of expressed words determined by visual data and expressed words determined by audio data according to an embodiment.
The expressed words determination unit 500a is an example of the circuitry 102 of Fig. 5 and is configured to enhance a speech captured for a video call by generating an artificial voice (computer voice) based on an event-based vision sensor 501 and a microphone 503. The expressed words determination unit 500a is configured to perform an expressed words determination based on visual data at S510, an expressed words determination based on audio data at S511, a fusion at S51 and an artificial voice generation at S513.
The expressed words determination based on visual data at S510 receives, from the event-based vision sensor 501, event-based visual data 501 of a user who is speaking (and, thus, expressing words) and determines, based on the event-based visual data 501, expressed words 514 of the user.
The expressed words determination based on audio data at S511 receives, from the microphone 503, noisy audio data 504 that indicate a speech of the user who is speaking and environment noise, and determines, based on the noisy audio data 504, expressed words 515 of the user.
The expressed words 514 determined at S510 based on the event-based visual data 502 and the expressed words 515 determined at S511 based on the noisy audio data 504 are provided to the fusion at S512.
The fusion at S512 fuses the expressed words 514 determined at S510 based on the event-based visual data 502 and the expressed words 515 determined at S511 based on the noisy audio data 504 to determined expressed words 516 and provides the determined expressed words 516 to the artificial voice generation at S513.
The artificial voice generation at S513 generates an artificial voice (computer voice) that utters the determined expressed words 516 and provides clear audio data 517 that indicate the artificial voice to a loudspeaker 505.
The loudspeaker 505 outputs the clear audio data 517 and, thus, outputs the determined expressed words of the user.
In the embodiment of Fig. 16a, the fusion at S512 is performed after determining the expressed words 514 and 515 based on the event-based visual data 502 and on the noisy audio data 504, respectively. In some embodiments, instead, the event-based visual data 502 and the noisy audio data 504 are fused before determining expressed words based thereon. In this case, the algorithm processes the event-based visual data 502 and the noisy audio data 504 simultaneously based on an artificial neural network. An example of such a case is provided in Fig. 16b.
Fig. 16b illustrates a block diagram of an algorithm for video call voice enhancement based on fusion of visual data and audio data before determining expressed words according to an embodiment.
The expressed words determination unit 500b is an example of the circuitry 102 of Fig. 5 and is configured to enhance a speech captured for a video call by generating an artificial voice (computer voice) based on an event-based vision sensor 501 and a microphone 503. The expressed words determination unit 500b differs from the expressed words determination unit 500a of Fig. 16a in that it is configured to perform a combined expressed words determination based on visual data and audio data at S520, i.e., expressed words are determined after fusing the event-based visual data 502 and the noisy audio data 504. The expressed words determination unit 500b is also configured to perform the artificial voice generation at S513.
The combined expressed words determination based on visual data and audio data at S520 receives the event-based visual data 502 and the noisy audio data 504 fused together, determines, with an artificial neural network, expressed words 521 of the user simultaneously based on the event-based visual data 502 and the noisy audio data 504 and provides the determined expressed words 521 to the artificial voice generation at S513.
The artificial voice generation at S513 and the loudspeaker 505 of Fig. 16b correspond to the artificial voice generation at S513 and the loudspeaker 505 of Fig. 16a, respectively.
The algorithm performed by the expressed words determination unit 500b, thus, performs an audio-video speech recognition.
Fig. 16c illustrates a block diagram of an algorithm for video call voice enhancement based on removing a noise of an environment according to an embodiment.
The expressed words determination unit 500c is an example of the circuitry 102 of Fig. 5 and is configured to enhance a speech captured for a video call by removing a noise of an environment based on the event-based vision sensor 501 and the microphone 503. The expressed words determination unit 500c is configured to perform the expressed words determination based on visual data at S510, an audio prediction at S530 and a noise filtering at S531. The expressed words determination based on visual data at S510 corresponds to the expressed words determination based on visual data at S510 of Fig. 16a and provides the determined expressed words 514 of the user to the audio prediction at S530.
The audio prediction at S530 predicts, based on the determined expressed words 514, predicted audio data 532 of the speech of the user and provides the predicted audio data 532 to the noise filtering at S531.
The noise filtering at S531 receives the noisy audio data 504 from the microphone, removes noise from the noisy audio data 504 based on the predicted audio data 532, thus generating clean audio data 533 that correspond to the speech of the user, and provides the clean audio data 533 to the loudspeaker 505.
The loudspeaker 505 is configured like the loudspeaker 505 of Fig. 16a.
Further applications of the present technology include hearing aids. In some embodiments, an event-based vison sensor is installed on hearing aids. The hearing aids include a circuitry according to the present disclosure (e.g., the circuitry 102 of Fig. 5) that amplifies a voice of a person that the wearer of the hearing aids is looking at. Such hearing aids improves an amplification of a correct audio source (cocktail party problem) because the wearer of the hearing aids can decide which source to amplify by looking at it.
Further applications of the present technology include noise protection headphones. In some embodiments, a circuitry according to the present disclosure (e.g., the circuitry 102 of Fig. 5) is provided in noise protection headphones. According to the same concept as for the hearing aids, for example construction workers or factory workers in a loud environment may wear the noise protection headphones which may remove (or at least reduce) a noise from the environment and amplify a voice of another worker.
In further embodiments, voice-enabled ATMs, kiosks and vending machines in noisy environments (e.g. train station) include a circuitry according to the present disclosure (e.g., the circuitry 102 of Fig. 5) for obtaining user input of a speaking user based on event-based visual data of the user.
In some embodiments, determining spoken (or, more generally, expressed) words of a user based on a camera (i.e., on visual data) rather than on a microphone (i.e., on audio data) allows to determine spoken (or expressed) words of the user in a very noisy environment or for a mute user. Furthermore, in some embodiments, determining spoken (expressed) words of a user based on event-based visual data obtained from an event-based vision sensor provides at least one of the following advantages:
The event-based vision sensor may be capable of capturing fast moving lips that could be a challenge for a normal camera.
The event-based vision sensor may compress redundant data (e.g. when lips do not move) and, thereby, make processing more efficient.
Event-based visual data captured by the event-based vison sensor may be more private than visual data captured by a normal camera since it is more difficult to reconstruct, based on eventbased visual data form an event-based vision sensor, a clear image that could be used to identify a person.
In some embodiments, the present technology may provide better video calls everywhere (even in very noisy environments), may provide speech enablement for mutes and/or may provide factory automation.
It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding. For example, the ordering of S35 and S36 in the embodiment of Fig. 3 may be exchanged. Also, the ordering of S155 and S156 in the embodiment of Fig. 8 may be exchanged. Other changes of the ordering of method steps may be apparent to the skilled person.
Please note that the division of the circuitry 2 of Fig. 1 into units 6 to 8 is only made for illustration purposes and that the present disclosure is not limited to any specific division of functions in specific units. Likewise, the division of the circuitry 102 of Fig. 5 into units 106 to 108 is only made for illustration purposes and the present disclosure is not limited to any specific division of functions in specific units. For instance, the circuitry 2 and/or the circuitry 102 could be implemented by a respective programmed processor, field programmable gate array (FPGA) and the like.
The method 30 of Fig. 3 and/or the method 150 of Fig. 8 can also be implemented as a computer program causing a computer and/or a processor, such as circuitry 2 of Fig. 1 and/or circuitry 102 of Fig. 5 discussed above, to perform the method, when being carried out on the computer and/or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the method described to be performed.
All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.
In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.
Note that the present technology can also be configured as described below.
(1) A circuitry for visual speech processing, configured to: determine expressed words based on visual data of a user; present the determined expressed words to the user; and adjust the determination of expressed words based on feedback of the user related to the determined expressed words.
(2) The circuitry of (1), wherein the determination of expressed words includes: detecting, in the visual data, a mouth region of the user; and identifying a lip movement in the mouth region of the user.
(3) The circuitry of (2), wherein the determination of expressed words includes: determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
(4) The circuitry of (3), wherein the generating of the expressed words is based on a language model.
(5) The circuitry of (4), wherein the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data.
(6) The circuitry of any one of (1) to (5), wherein the determination of expressed words includes: detecting, in the visual data, a face expression of the user; and generating the expressed words based on the detected face expression. (7) The circuitry of any one of (1) to (6), wherein the determination of expressed words includes: obtaining audio data from a microphone; and generating the expressed words based on the audio data.
(8) The circuitry of any one of (1) to (7), wherein the presenting of the determined expressed words includes displaying the determined expressed words as text.
(9) The circuitry of any one of (1) to (8), wherein the presenting of the determined expressed words includes displaying at least one icon corresponding to the determined expressed words.
(10) The circuitry of any one of ( 1 ) to (9), wherein the presenting of the determined expressed words includes: generating an artificial voice that utters the determined expressed words; and outputting the generated artificial voice.
(11) The circuitry of any one of ( 1 ) to ( 10), wherein the visual data include a plurality of events obtained from an event-based vision sensor.
(12) The circuitry of any one of (1) to (11), wherein the adjusting of the determination of expressed words includes changing a value of a parameter of the determination of expressed words.
(13) The circuitry of any one of ( 1 ) to ( 12), wherein the determination of expressed words is based on an artificial neural network.
(14) The circuitry of (13), wherein the feedback of the user indicates at least one of a label, a loss and a fitness for the artificial neural network.
(15) The circuitry of any one of ( 1 ) to ( 14), wherein the feedback of the user includes a correction of at least one of the determined expressed words; and wherein the adjusting of the determination of expressed words is based on the correction of the at least one of the determined expressed words. (16) The circuitry of any one of ( 1 ) to ( 15), wherein the presenting of the determined expressed words includes recommending an improved lip movement.
(17) A method for visual speech processing, comprising: determining expressed words based on visual data of a user; presenting the determined expressed words to the user; and adjusting the determination of expressed words based on feedback of the user related to the determined expressed words.
(18) The method of ( 17), wherein the determination of expressed words includes: detecting, in the visual data, a mouth region of the user; and identifying a lip movement in the mouth region of the user.
(19) The method of (17), wherein the determination of expressed words includes: determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
(20) The method of ( 19), wherein the generating of the expressed words is based on a language model.
(21) The method of (20), wherein the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data.
(22) The method of any one of (17) to (21), wherein the determination of expressed words includes: detecting, in the visual data, a face expression of the user; and generating the expressed words based on the detected face expression.
(23) The method of any one of (17) to (22), wherein the determination of expressed words includes: obtaining audio data from a microphone; and generating the expressed words based on the audio data.
(24) The method of any one of (17) to (23), wherein the presenting of the determined expressed words includes displaying the determined expressed words as text. (25) The method of any one of (17) to (24), wherein the presenting of the determined expressed words includes displaying at least one icon corresponding to the determined expressed words.
(26) The method of any one of (17) to (25), wherein the presenting of the determined expressed words includes: generating an artificial voice that utters the determined expressed words; and outputting the generated artificial voice.
(27) The method of any one of (17) to (26), wherein the visual data include a plurality of events obtained from an event-based vision sensor.
(28) The method of any one of (17) to (27), wherein the adjusting of the determination of expressed words includes changing a value of a parameter of the determination of expressed words.
(29) The method of any one of (17) to (28), wherein the determination of expressed words is based on an artificial neural network.
(30) The method of (29), wherein the feedback of the user indicates at least one of a label, a loss and a fitness for the artificial neural network.
(31) The method of any one of ( 17) to (30), wherein the feedback of the user includes a correction of at least one of the determined expressed words; and wherein the adjusting of the determination of expressed words is based on the correction of the at least one of the determined expressed words.
(32) The method of any one of (17) to (31), wherein the presenting of the determined expressed words includes recommending an improved lip movement.
(33) A computer program comprising program code causing a computer to perform the method according to anyone of (17) to (32), when being carried out on a computer.
(34) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (17) to (32) to be performed. Note further that the present technology can also be configured as described below.
(1) A circuitry for visual speech processing, configured to: obtain event-based visual data of a user from an event-based vision sensor; determine expressed words based on the obtained event-based visual data; and output the determined expressed words.
(2) The circuitry of (1), wherein the determination of expressed words includes: identifying, in the event-based visual data, a lip movement of the user.
(3) The circuitry of (2), wherein the identifying of the lip movement includes: generating, based on the event-based visual data, at least one full image of a mouth region of the user; and identifying the lip movement based on the generated at least one full image of the mouth region of the user.
(4) The circuitry of (2), wherein the identifying of the lip movement includes: identifying, based on the event-based visual data, events of a mouth region of the user; determining motion vectors based on the identified events of the mouth region; and identifying the lip movement based on the determined motion vectors.
(5) The circuitry of any one of (2) to (4), wherein the determination of expressed words includes: determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
(6) The circuitry of (5), wherein the determining of the syllable data includes: identifying, based on the identified lip movement, at least one viseme; and determining the syllable data based on the at least one viseme.
(7) The circuitry of (5) or (6), wherein the generating of the expressed words is based on a language model.
(8) The circuitry of (7), wherein the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data.
(9) The circuitry of any one of (1) to (8), wherein the outputting of the determined expressed words includes displaying the determined expressed words as text. (10) The circuitry of any one of (1) to (9), wherein the outputting of the determined expressed words includes displaying at least one icon corresponding to the determined expressed words.
(11) The circuitry of any one of ( 1 ) to ( 10), wherein the outputting of the determined expressed words includes: generating an artificial voice that utters the determined expressed words; and outputting the generated artificial voice.
(12) The circuitry of any one of (1) to (11), wherein the outputting of the determined expressed words includes: transmitting the determined expressed words to a communication device of another user.
(13) The circuitry of any one of ( 1 ) to ( 12), wherein the determination of expressed words is based on an artificial neural network.
(14) The circuitry of any one of ( 1 ) to ( 13), wherein the determination of expressed words is further based on audio data obtained from a microphone.
(15) The circuitry of (14), wherein the outputting of the determined expressed words includes: reducing, in the audio data obtained from the microphone, noise that does not correspond to the determined expressed words; and outputting the audio data in which the noise has been reduced.
(16) The circuitry of (14) or (15), wherein the outputting of the determined expressed words includes: amplifying, in the audio data obtained from the microphone, a voice that corresponds to the determined expressed words; and outputting the audio data in which the voice has been amplified.
(17) A method for visual speech processing, comprising: obtaining event-based visual data of a user from an event-based vision sensor; determining expressed words based on the obtained event-based visual data; and outputting the determined expressed words.
(18) The method of (17), wherein the determination of expressed words includes: identifying, in the event-based visual data, a lip movement of the user.
(19) The method of ( 18), wherein the identifying of the lip movement includes: generating, based on the event-based visual data, at least one full image of a mouth region of the user; and identifying the lip movement based on the generated at least one full image of the mouth region of the user.
(20) The method of (18), wherein the identifying of the lip movement includes: identifying, based on the event-based visual data, events of a mouth region of the user; determining motion vectors based on the identified events of the mouth region; and identifying the lip movement based on the determined motion vectors.
(21) The method of any one of (18) to (20), wherein the determination of expressed words includes: determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
(22) The method of (21), wherein the determining of the syllable data includes: identifying, based on the identified lip movement, at least one viseme; and determining the syllable data based on the at least one viseme.
(23) The method of (21 ) or (22), wherein the generating of the expressed words is based on a language model.
(24) The method of (23), wherein the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data.
(25) The method of any one of (17) to (24), wherein the outputting of the determined expressed words includes displaying the determined expressed words as text.
(26) The method of any one of (17) to (25), wherein the outputting of the determined expressed words includes displaying at least one icon corresponding to the determined expressed words.
(27) The method of any one of (17) to (26), wherein the outputting of the determined expressed words includes: generating an artificial voice that utters the determined expressed words; and outputting the generated artificial voice. (28) The method of any one of (17) to (27), wherein the outputting of the determined expressed words includes: transmitting the determined expressed words to a communication device of another user.
(29) The method of any one of (17) to (28), wherein the determination of expressed words is based on an artificial neural network.
(30) The method of any one of (17) to (29), wherein the determination of expressed words is further based on audio data obtained from a microphone.
(31) The method of (30), wherein the outputting of the determined expressed words includes: reducing, in the audio data obtained from the microphone, noise that does not correspond to the determined expressed words; and outputting the audio data in which the noise has been reduced.
(32) The method of (30) or (31), wherein the outputting of the determined expressed words includes: amplifying, in the audio data obtained from the microphone, a voice that corresponds to the determined expressed words; and outputting the audio data in which the voice has been amplified.
(33) A computer program comprising program code causing a computer to perform the method according to anyone of (17) to (32), when being carried out on a computer.
(34) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (17) to (32) to be performed.

Claims

1. A circuitry for visual speech processing, configured to: determine expressed words based on visual data of a user; present the determined expressed words to the user; and adjust the determination of expressed words based on feedback of the user related to the determined expressed words.
2. The circuitry of claim 1, wherein the determination of expressed words includes: detecting, in the visual data, a mouth region of the user; and identifying a lip movement in the mouth region of the user.
3. The circuitry of claim 2, wherein the determination of expressed words includes: determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
4. The circuitry of claim 3, wherein the generating of the expressed words is based on a language model.
5. The circuitry of claim 4, wherein the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data.
6. The circuitry of claim 1, wherein the determination of expressed words includes: detecting, in the visual data, a face expression of the user; and generating the expressed words based on the detected face expression.
7. The circuitry of claim 1, wherein the determination of expressed words includes: obtaining audio data from a microphone; and generating the expressed words based on the audio data.
8. The circuitry of claim 1, wherein the visual data include a plurality of events obtained from an event-based vision sensor.
9. The circuitry of claim 1, wherein the determination of expressed words is based on an artificial neural network.
10. The circuitry of claim 1, wherein the feedback of the user includes a correction of at least one of the determined expressed words; and wherein the adjusting of the determination of expressed words is based on the correction of the at least one of the determined expressed words.
11. A method for visual speech processing, comprising: determining expressed words based on visual data of a user; presenting the determined expressed words to the user; and adjusting the determination of expressed words based on feedback of the user related to the determined expressed words.
12. The method of claim 11, wherein the determination of expressed words includes: detecting, in the visual data, a mouth region of the user; and identifying a lip movement in the mouth region of the user.
13. The method of claim 12, wherein the determination of expressed words includes: determining syllable data based on the identified lip movement; and generating the expressed words based on the syllable data.
14. The method of claim 13, wherein the generating of the expressed words is based on a language model.
15. The method of claim 14, wherein the language model includes a word list; and the generating of the expressed words includes selecting at least one word from the word list based on the determined syllable data.
16. The method of claim 11, wherein the determination of expressed words includes: detecting, in the visual data, a face expression of the user; and generating the expressed words based on the detected face expression.
17. The method of claim 11, wherein the determination of expressed words includes: obtaining audio data from a microphone; and generating the expressed words based on the audio data.
18. The method of claim 11, wherein the visual data include a plurality of events obtained from an event-based vision sensor.
19. The method of claim 11, wherein the determination of expressed words is based on an artificial neural network.
20. The method of claim 11, wherein the feedback of the user includes a correction of at least one of the determined expressed words; and wherein the adjusting of the determination of expressed words is based on the correction of the at least one of the determined expressed words.
EP23705432.5A 2022-03-04 2023-02-21 Circuitry and method for visual speech processing Pending EP4487319A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP22160271 2022-03-04
PCT/EP2023/054267 WO2023165844A1 (en) 2022-03-04 2023-02-21 Circuitry and method for visual speech processing

Publications (1)

Publication Number Publication Date
EP4487319A1 true EP4487319A1 (en) 2025-01-08

Family

ID=80628958

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23705432.5A Pending EP4487319A1 (en) 2022-03-04 2023-02-21 Circuitry and method for visual speech processing

Country Status (2)

Country Link
EP (1) EP4487319A1 (en)
WO (1) WO2023165844A1 (en)

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102826450B1 (en) * 2018-11-15 2025-06-27 삼성전자 주식회사 Apparatus and method for generating personalization lip reading model
US11817080B2 (en) * 2019-09-03 2023-11-14 Google Llc Using corrections, of predicted textual segments of spoken utterances, for training of on-device speech recognition model
DE102020118967A1 (en) * 2020-07-17 2022-01-20 Clinomic GmbH METHOD FOR AUTOMATIC LIP READING USING A FUNCTIONAL COMPONENT AND FOR PROVIDING THE FUNCTIONAL COMPONENT

Also Published As

Publication number Publication date
WO2023165844A1 (en) 2023-09-07

Similar Documents

Publication Publication Date Title
US12032155B2 (en) Method and head-mounted unit for assisting a hearing-impaired user
EP2574220B1 (en) Hand-held communication aid for individuals with auditory, speech and visual impairments
JP6819672B2 (en) Information processing equipment, information processing methods, and programs
CN108702580A (en) Hearing auxiliary with automatic speech transcription
US20170084271A1 (en) Voice processing apparatus and voice processing method
EP2925005A1 (en) Display apparatus and user interaction method thereof
JP7123856B2 (en) Presentation evaluation system, method, trained model and program, information processing device and terminal device
US20230260534A1 (en) Smart glass interface for impaired users or users with disabilities
CN110730360A (en) Video uploading and playing methods and devices, client equipment and storage medium
JP2009178783A (en) Communication robot and control method thereof
US11164341B2 (en) Identifying objects of interest in augmented reality
KR20200044947A (en) Display control device, communication device, display control method and computer program
JP2023117068A (en) Speech recognition device, speech recognition method, speech recognition program, speech recognition system
KR20140093459A (en) Method for automatic speech translation
JP4772315B2 (en) Information conversion apparatus, information conversion method, communication apparatus, and communication method
JP2025100735A (en) Information processing device
EP4487319A1 (en) Circuitry and method for visual speech processing
KR20190067662A (en) Sign language translation system using robot
US20240144955A1 (en) Method for monitoring emotion and behavior during conversation for user in need of protection
WO2017029850A1 (en) Information processing device, information processing method, and program
CN115171284B (en) A method and device for caring for the elderly
CN111507115B (en) Multi-modal language information artificial intelligence translation method, system and equipment
KR102128812B1 (en) Method for evaluating social intelligence of robot and apparatus for the same
JP2018081147A (en) COMMUNICATION DEVICE, SERVER, CONTROL METHOD, AND INFORMATION PROCESSING PROGRAM
KR20170093631A (en) Method of displaying contens adaptively

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20240930

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)