WO2016111644A1 - A method for signal processing of voice of a speaker - Google Patents

A method for signal processing of voice of a speaker Download PDF

Info

Publication number
WO2016111644A1
WO2016111644A1 PCT/SG2015/050519 SG2015050519W WO2016111644A1 WO 2016111644 A1 WO2016111644 A1 WO 2016111644A1 SG 2015050519 W SG2015050519 W SG 2015050519W WO 2016111644 A1 WO2016111644 A1 WO 2016111644A1
Authority
WO
WIPO (PCT)
Prior art keywords
speaker
voice
signals
adjusted
output signal
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/SG2015/050519
Other languages
French (fr)
Inventor
Wong Hoo Sim
Teck Chee Lee
Xiaoting LIU
Efstratios SOFIANOS
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Creative Technology Ltd
Original Assignee
Creative Technology Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Creative Technology Ltd filed Critical Creative Technology Ltd
Priority to SG11201705469XA priority Critical patent/SG11201705469XA/en
Publication of WO2016111644A1 publication Critical patent/WO2016111644A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/003Changing voice quality, e.g. pitch or formants
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/15Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being formant information
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • G10L25/63Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for estimating an emotional state
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/90Pitch determination of speech signals

Definitions

  • the present disclosure generally relates to method for signal processing of voice of a speaker so that emotion of the speaker can be audibly perceived to be adjusted.
  • speech emotion i.e., emotion(s) expressed in speech
  • speech features such as pitch contour and speech rate are typically used on classification.
  • a method for signal processing of voice of a speaker is provided.
  • the voice of the speaker can correspond to audio signals which include one or more voice parameters.
  • the voice parameter(s) can be associated with emotion of the speaker.
  • the method can include receiving the audio signals, processing the audio signals, providing an output and providing an audible feedback.
  • the audio signals can be processed by adjusting one or more of the one or more voice parameters to produce an adjusted speech output signal.
  • an output of the adjusted speech output signal can be provided so that the adjusted speech output signal is audibly perceivable.
  • an audible feedback of the adjusted speech output signal can be provided to the speaker.
  • Speaker emotion can be audibly perceived to be adjusted, in accordance with the adjusted one or more voice parameters, when the adjusted speech output signal is output.
  • the speaker is capable of receiving the audible feedback of the adjusted speech output signal and further adjusting the one or more voice parameters in a manner so as to accentuate speaker emotion.
  • Fig. 1 shows a system which can include an input module, a processing module and an output module, according to an embodiment of the disclosure
  • Fig. 2 shows the system of Fig. 1 in further detail, according to an embodiment of the disclosure
  • Fig. 3 shows exemplary signal waveforms associated with the system of Fig. 1, according to an embodiment of the disclosure.
  • Fig. 4 shows a flow diagram illustrating a method for signal processing in association with the system of Fig. 1, according to an embodiment of the disclosure.
  • speech emotion i.e., emotion(s) expressed in speech
  • predictive data i.e., "future" data
  • speech utterance may be rendered unnatural sounding during real-time processing (e.g. if speech rate is to be changed).
  • An example of a real-time use type situation is where a speaker desires a listener/listener(s) to perceive/discern, via the emotions expressed in the speaker's speech, the speaker's intended emotion (also, otherwise, referable to as "true” emotion) while the speaker is speaking.
  • the speaker's intended emotion also, otherwise, referable to as "true” emotion
  • the tired teacher's speech may not be capable of conveying the intended verbal firmness (i.e., intended emotion or "true” emotion). Therefore, it may be useful for the teacher that students in the classroom are capable of perceiving the intended verbal firmness from the teacher's speech despite the teacher's tiredness.
  • a speaker does not desire the listener/listener(s) to be able to perceive/discern, via the emotions expressed in the speaker's speech, the speaker's intended/"true" emotion while the speaker is speaking.
  • a speaker's speech may inadvertently express negative emotions such as anger or frustration (i.e., "true" emotion).
  • a calm/neutral tone masks the speaker's true emotion (e.g., anger/frustration) and may help facilitate a favorable (or, at least, a more fruitful) outcome for the teleconference session.
  • the system 100 can include an input module 102, a processing module 104 and an output module 106.
  • the input module 102 can be coupled to the processing module 104.
  • the processing module 104 can be coupled to the output module 106.
  • the output module 106 can be coupled to the input module 102.
  • the input module 102 can be configured to generate input signals and communicate the input signals to the processing module 104 for processing.
  • the processing module 104 can be configured to receive and process the input signals to produce processed input signals.
  • the processing module 104 can be further configured to communicate the processed input signals to the output module 106.
  • the output module 106 can be configured to receive and process the processed input signals in a manner so as to produce output signals.
  • the system 100 can correspond to an ecosystem in which the input module 102, the processing module 104 and the output module 106 are three individual/independent devices which can be coupled as described earlier.
  • the system 100 can correspond to a printed circuit board (PCB) on which the input module 102, the processing module 104 and the output module 106 are placed/positioned and coupled as described earlier.
  • PCB printed circuit board
  • the input module 102, the processing module 104 and the output module 106 can be three individual/independent elements carried by the PCB.
  • the system 100 can correspond to a device such as an integrated circuit (IC) chip which includes the input module 102, the processing module 104 and the output module 106.
  • IC integrated circuit
  • the input module 102, the processing module 104 and the output module 106 can form parts, which are coupled as described earlier, of the device (e.g., IC chip).
  • the input module 102 and the processing module 104 can be parts of a device (i.e., a first device), and the output module 106 can be another device (i.e., a second device) coupled thereto.
  • the processing module 104 and the output module 106 can be parts of a device (i.e., a first device), and the input module 102 can be another device (a second device) coupled thereto.
  • the input module 102 and the output module 106 can be parts of a device (i.e., a first device), and the processing module 104 can be another device (i.e., a second device) coupled thereto.
  • the first device can be coupled to the second device.
  • the aforementioned coupling can be based on one or both of wired coupling and wireless coupling.
  • Fig. 2 shows the system 100 of Fig. 1 in further detail, in accordance with an embodiment of the disclosure.
  • the input module 102 can include an input portion 202 and a control portion 204.
  • the input module 102 can, as an option, further include an analyzer portion 206.
  • Each of the input portion 202 and the control portion 204 can be coupled to the processing module 104. Moreover, the input portion 202 can, optionally, be coupled to one or both of the control portion 204 and the analyzer portion 206. Additionally, the analyzer portion 206 can, optionally, be coupled to the processing module 104.
  • the input portion 202 can be configured to receive audio signals.
  • the audio signals can, for example, correspond to the voice of a speaker 208.
  • the audio signals can, for example, correspond to speaker voice (i.e., of the speaker 208).
  • the input portion 202 can, for example, be a microphone which can be configured to capture and receive the voice of a speaker 208 who may, for example, be speaking into the microphone.
  • the input portion 202 can be further configured to process the received audio signals to produce source signals which can be communicated to one or both of the processing module 104 and the analyzer portion 206 for further processing.
  • the control portion 204 can be configured to generate control signals and communicate the control signals to one or both of the processing module 104 and the input portion 202.
  • the control signals can be communicated from the control portion 204 to the processing module 104 to control the manner in which the processing module 104 processes the source signals communicated from the input portion 202 as will be discussed later in further detail in the context of an exemplary scenario.
  • the control portion 204 can be configured to communicate control signals to the input portion 202 to control the manner in which the input portion 202 receives and/or processes the audio signals.
  • control signals can be communicated from the control portion 204 to the input portion 202 to control sensitivity of the input portion 202 (e.g., microphone input sensitivity).
  • control portion 202 can, for example, be a controller device having one or more controls.
  • a control can include button(s), knob(s), slider(s), touch panel(s) and/or sensor(s).
  • examples of a sensor can include an accelerometer, a gyroscope, a magnetometer and a camera based tracker.
  • the control(s) can, for example, be operated by the speaker 208 to generate control signals.
  • the speaker 208 can operate the slider (e.g., by adjusting the slider) to generate control signals.
  • control portion 204 can, for example, be in the form of an "N" Dimensional (ND) controller device where "N" represents an integer and the integer represented by "N” can, for example, correspond to the number of controls available on the control portion 202 or the number of types of controls available on the control portion 202.
  • N represents an integer
  • the control portion 202 can be considered to be a 3 Dimensional controller device.
  • the control portion 202 can be considered to be a 1 Dimensional controller device.
  • the analyzer portion 206 can be configured to receive the source signals from the input portion 202.
  • the analyzer portion 206 can be further configured to analyze the source signals to generate analysis signals.
  • the analyzer portion 206 can yet be further configured to communicate the analysis signals to the processing module 104.
  • the analyzer portion 206 can, for example, correspond to a processor device configured to compute the source signals by way of analyzing to produce analysis signals.
  • the analyzer portion 206 can correspond to a processor device configured to perform the task of data mining (e.g., by way of analyzing) on the received source signals and produce analysis signals based on mined data.
  • the processing module 104 can be configured to process one or both of the received control signals and the received source signals based on the analysis signals.
  • the input signals communicated from the input module 102 can include any one of the source signals, the control signals and the analysis signals, or any combination thereof.
  • the processing module 104 can be configured to receive the source signals, the control signals and/or the analysis signals for further processing to produce processed input signals which can, in turn, be further communicated (i.e., from the processing module 104) to the output module 106 for further processing to produce output signals.
  • the input portion 202 can correspond to a microphone
  • the processing module 104 can correspond to a processor device (e.g., in the form of an integrated circuit chip) and the output module 106 can correspond to one or more speaker drivers capable of sound output.
  • the speaker 208 can, for example, be a natural person generating and providing audio signals to the system 100 by way of speaking into the microphone. In this regard, the voice of the speaker 208 can correspond to the audio signals received by the input portion 202.
  • the voice of the speaker 208 can include one or more voice parameters which can be indicative of the emotion (e.g., anger, joy, sadness, surprise and calmness/neutral) of the speaker 208.
  • the one or more voice parameters can be associated with emotion of the speaker 208.
  • Examples of the aforementioned voice parameter(s) can include pitch contour, speech rate, formant, pitch & variability and energy.
  • the input portion 202 can be configured to process the audio signals to produce source signals.
  • the source signals can, by the same token, include one or more voice parameters.
  • audio signals which can be analog based (i.e., voice of the speaker) can be received by the microphone (i.e., the input portion 202) and processed by the microphone to produce signals in a form (i.e., source signals) suitable for further processing by the processing module 104 and/or the analyzer portion 206.
  • the microphone can process the received audio signals by way of conversion to electric signals. Therefore, the source signals can, for example, correspond to electric signals which are based on conversion from the received audio signals which may be analog type signals.
  • the source signals can be further communicated to the processing module 104 for further processing.
  • the processing module 104 is configured to process the source signals by adjusting one or more of the voice parameters. Adjustment of the voice parameter(s) can, for example, be based on either the control signals or a combination of both of the control signals and the analysis signals.
  • the processing module 104 can be configured to process the source signals by adjusting one or more of the voice parameters based on, for example, either the control signals or a combination of the control signals and the analysis signals so as to produce processed source signals.
  • the processed source signals can correspond to the aforementioned processed input signals in Fig. 1.
  • the processed source signals can be communicated from the processing module 104 to the output module 106 which can process the processed source signals in a manner so as to produce output signals.
  • the output module 106 can correspond to one or more speaker drivers capable of sound output.
  • the output module 106 can be configured to receive and process the processed source signals in a manner so as to produce output signals which can be audibly perceived.
  • the source signals are based on the voice of the speaker 208 and can include one or more voice parameters
  • the source signals can correspond to a speech input signal which can include one or more voice parameters.
  • the speech input signal can be associated with one or more speaker emotions (i.e., of the speaker 208) based on the one or more voice parameters.
  • the speech input signal can be considered to be adjusted in the context of speaker emotion. Therefore, output signals, which can be audibly perceived, communicated from the output module 106 can correspond to an adjusted speech output signal.
  • the adjusted speech output signal when output by the output module 106, can correspond to an audibly perceivable output signal where speaker emotion can be audibly perceived to be adjusted in accordance with the adjusted one or more voice parameters.
  • audibly perceivable speaker emotion associated with the adjusted speech output signal can be considered to be adjusted relative to speaker emotion associated with the input speech signal.
  • the input speech signal is based on voice of the speaker 208.
  • an audible feedback of the adjusted speech output signal can be provided to the speaker 208.
  • the speaker 208 can be capable of receiving an audible feedback of the adjusted speech output signal.
  • the audible feedback can, for example, be based on the speaker 208 hearing the adjusted speech output signal as being output from the output module 106.
  • providing an audible feedback of the adjusted speech output signal to the speaker 208 can be useful in facilitating further adjustment of the one or more voice parameters.
  • the one or more voice parameters can be further adjusted in a manner so as to accentuate, for example, an intended or a desired emotion (i.e., which can be audibly perceived) of the speaker 208.
  • the speaker 208 can be capable of receiving an audible feedback of the adjusted speech output signal and further adjusting the one or more voice parameters in a manner so as to accentuate speaker emotion. Yet more specifically, the speaker 208 can be capable of hearing the adjusted speech output signal as being output by the output module 106 and further adjusting, as desired by the speaker 208, the one or more voice parameters in a manner so as to accentuate speaker emotion.
  • a teacher i.e., the speaker 208 wishes to instill discipline during a lecture session in a lecture theatre/hall
  • the teacher using/operating the input module 102 may speak into the microphone (i.e., the input portion 202) and generate control signals using the control portion 204 (e.g., by sliding the slider).
  • the teacher's voice may be tired sounding.
  • the tired sounding voice (audio signals which can be received by the microphone) may be associated with the emotion of calmness.
  • the tired sounding voice of the teacher can be processed by the microphone to produce the speech input signal (i.e., source signals).
  • the speech input signal can be associated with the emotion of calmness which may not be the intended emotion of the teacher and, further, not ideal for the purpose of instilling discipline.
  • adjustment can be made by way of processing speech input signals based on control signals.
  • the control signals can be generated based on the intent of the teacher to make an adjustment to the speech input signal which can be associated with the emotion of calmness so as produce an adjusted speech output signal which can be audibly perceived to be associated with the emotion of anger.
  • labels can be provided. For example, a label can be provided at one end of the slider indicating "calm voice" and another label can be provided at another end of the slider indicating "extremely angry voice". In this manner, the teacher can slide the slider accordingly to vary the degree/strength of anger (i.e., more towards "calm voice” or more towards “extremely angry voice”) which can be audibly perceived via the adjusted speech output signal.
  • the analyzer portion 206 can be configured to analyze the source signals (i.e., speech input signals) to generate analysis signals.
  • the analyzer portion 206 can be configured to analyze the teacher's voice and extract one or more voice characteristics such as pitch and formant. Based on the voice characteristic(s), it is possible for the analyzer portion 206 to determine useful information associated with the speaker 208 (i.e., the teacher) and generate analysis signals accordingly.
  • useful information associated with the speaker 208 is the gender of the speaker 208. For example, based on the voice characteristic(s) such as the pitch, it can be determined that the teacher is a female.
  • the analyzer portion 206 can be configured to communicate analysis signals indicative of the teacher's gender being a female.
  • the analysis signals can be used for adjusting the control signals as will be discussed later in further detail with reference to Fig. 3. Therefore, the speech input signals can be considered to be processed based on a combination of the control signals and the analysis signals.
  • the control signals can be adjusted by the analysis signals at either the control portion 204 or the processing module 104. Moreover, it is also possible for the control signals to be adjusted by the analysis signals at both the control portion 204 and the processing module 104.
  • the speech input signals can be communicated to the processing module 104 for further processing based on either the control signals or a combination of the control signals and the analysis signals.
  • one or more voice parameters of the speech input signal which can be associated with the emotion of anger (i.e., intended emotion of the speaker 208) can be adjusted based on the control signals or a combination of the control signals and the analysis signals. Therefore, adjusting the voice parameter(s) based on the control signals can be the basis for a form of real-time processing to produce an adjusted speech output signal in which adjustment has been made in the context of the speaker emotion (i.e., associated with anger in this exemplary real-time use type situation).
  • the processing module 104 can process the speech input signals based on a combination of the control signals and analysis signals (communicated from the analyzer portion 206) to produce the adjusted speech output signal
  • the analysis signals can be considered to be another basis for a form of real-time processing to produce an adjusted speech output signal in which adjustment has been made in the context of the speaker motion (i.e., associated with anger in this exemplary real-time use type situation). Processing of the speech input signal based on the control signals or a combination of the control signals and the analysis signals will be discussed later in further detail with reference to Fig. 3.
  • the speech input signals which can be associated with the emotion of calmness can be adjusted by the teacher so that the adjusted speech output signals can be audibly perceived to be associated with the emotion of anger.
  • the teacher can be speaking into the microphone (i.e., the input portion 202) with a tired sounding voice (which can be associated with the emotion of calmness).
  • a tired sounding voice which can be associated with the emotion of calmness.
  • the students hearing the adjusted speech output signal from the one or more speaker drivers may audibly perceive the adjusted speech output signal to be associated with the emotion of anger.
  • Appreciably processing of the speech input signal to produce the adjusted speech output signal is based on real-time based processing.
  • real-time processing of the received audio signals i.e., voice of the teacher in accordance with the above exemplary real-time use type situation
  • the system 100 can be made possible using the system 100 by virtue at least one of the following approaches:
  • the teacher is essentially hearing her (assuming the gender of the teacher is a female as discussed earlier) own voice which has been adjusted to be associated with the emotion of anger from of her original tired sounding voice (associated with the emotion of calmness) as received by the microphone (i.e. the input portion 202).
  • the speech input signal can be associated with an "original/input voice” (e.g., tired sounding voice as received by the microphone) and the adjusted speech output signal can correspond to a "modified voice" (e.g., angry sounding voice as being output by the one or more speaker drivers) relative to the "original/input voice".
  • the teacher upon, hearing her own voice as output via the one or more speaker drivers (i.e., receiving audible feedback), may decide that the emotion of anger may not be clearly audibly perceived enough. The teacher may then decide to make further adjustment(s) to the one or more voice parameters so as to accentuate the emotion of anger.
  • the manner of further adjustment(s) can, for example, be by way of the teacher adjusting her own voice (i.e., speaker voice) when speaking into the microphone to be more angry sounding so as to accentuate the intended or the desired emotion of, for example, anger.
  • her own voice i.e., speaker voice
  • the teacher may realize that she has to make an effort to be less tired sounding when speaking into the microphone.
  • the teacher can make further adjustments via the control portion 204 by sliding the slider even more towards the "extremely angry voice" direction.
  • This can be akin to dynamically adjusting the strength/impact of the effect of, for example, anger by making adjustment(s) (e.g., sliding the slider) using the control portion 204.
  • further adjustment can be by manner of adjusting control signals generated and communicated to the processing module 104.
  • further adjustments can be by manner of both adjusting speaker voice communicated to, and received by, the microphone and adjusting control signals generated, and communicated to, the processing module 104.
  • the system 100 effectively enables a user (e.g., the speaker 208) to change/adjust perceived (i.e., as perceived by one or more listeners such as the aforementioned students) emotion associated with the user's speech.
  • a user e.g., the speaker 208
  • the original tired sounding voice of the speaker 208 e.g., a teacher
  • the original tired sounding voice of the speaker 208 which can be associated with an emotion of calmness can be adjusted/changed so that listeners (e.g., the students) can audibly perceive an emotion of anger instead.
  • system 100 can be considered to be based on real-time processing by virtue of at least one of the following:
  • the speech input signal i.e., source signals based on the audio signals
  • the speech input signal can be processed based on the control signals or a combination of the control signals and the analysis signals. This will be discussed in further detail with reference to Fig. 3 hereinafter.
  • Fig. 3 shows exemplary signal waveforms associated with the system 100, in accordance with an embodiment of the disclosure.
  • Fig. 3 shows an exemplary speech input signal waveform 310 and a corresponding data waveform 320 of a voice parameter.
  • Fig. 3 also shows an adjusted data waveform 330 where the voice parameter has been adjusted and an exemplary speech output signal waveform 340.
  • the voice parameter can, for example, relate to pitch. Therefore, the data waveform 320 can show the pitch data associated with the speech input signal (i.e., as represented by the exemplary speech input signal waveform 310).
  • the pitch data associated with the speech input signal can be indicative of audibly perceivable pitch (e.g., the degree of height or depth of voice tone) of the speaker's 208 voice.
  • the data waveform 320 can be derived from the speech input signal waveform 310 by manner of, for example, a sampling process or an extraction process which can be performed by the processing module 104.
  • the data waveform 320 can include a plurality of data portions (e.g., labels 320a, 320b and 320c). Each of the data portions can include one or more data points (e.g., labels 320d and 320e within data portion 320c)
  • the speaker 208 can provide audio signals to the system 100 by, for example, speaking into the input portion 202.
  • the input portion 202 processes the audio signals and produces the speech input signal (i.e., source signals) as represented by the exemplary speech input signal waveform 310.
  • the speech input signal can be processed by the processing module 104 in a manner so that one or more data points (e.g., labels 320f, 320g and 320h) within one or more data portions (e.g., labels 320a, 320b and 320c) is/are adjusted. Adjustment of the data point(s) can be based on instructions communicated to the processing module 104 in the form of the aforementioned control signals.
  • the speech input signal can be processed based on the control signals.
  • the speech input signal can, as an option, be processed based a combination of the control signals and the analysis signals. This will be discussed in further detail using the earlier discussed example of the analysis signals being indicative of the gender of the speaker 208.
  • the analysis signals can be used to adjust the control signals (at the control portion 204 and/or the processing module 104), so that one or more of the data points can be adjusted, for example, to either a higher degree in or a lesser degree in comparison to the data point(s) being adjusted based on the control signals alone.
  • the speaker 208 is a female teacher (e.g., in the context of the earlier discussed exemplary real-time use type situation)
  • the speaker 208 is, for example, a male teacher
  • the speech input signal can be processed based a combination of the control signals and the analysis signals.
  • the analysis signals can be used to adjust the control signals and the adjusted control signals can, in turn, be used to adjust the speech input signal to produce the adjusted speech output signal.
  • a flow diagram 400 illustrating a method for signal processing in association with the system 100 is shown in accordance with an embodiment of the disclosure.
  • the method for signal processing can be directed at the signal processing of voice of a speaker (e.g., earlier discussed speaker 208).
  • the voice of the speaker can correspond to audio signals which include one or more voice parameters.
  • the one or more voice parameters can be associated with emotion of the speaker.
  • the method can include receiving audio signals 410, processing audio signals to produce an adjusted speech output signal 420, providing an output of an adjusted speech output signal 430 and providing an audible feedback of the adjusted speech output signal 440.
  • the audio signals can be received by the input portion 202 as discussed earlier with reference to the system 100.
  • one or more voice parameters associated with the audio signals can be adjusted to produce an adjusted speech output signal.
  • the received audio signals can be processed by the input portion 202 to produce an input speech signal which can include one or more voice parameters (i.e., corresponding to the one or more voice parameters associated with the audio signals as received by the input portion 202).
  • the input speech signal can be communicated to the processing module 104 for further processing.
  • the processing module 104 can be configured to process the speech input signal (i.e., source signals based on the audio signals) based on the control signals or a combination of the control signals and the analysis signals to produce the adjusted speech output signal.
  • the adjusted speech output signal can, as discussed earlier with reference to the system 100, be communicated to the output module 106 which can be configured to process the speech output signal in a manner so as to produce corresponding output signals which can be audibly perceived.
  • the adjusted speech output signal can, by virtue of the output signals from the output module 106, be considered to be capable of being audibly perceived. More specifically, the adjusted speech output signal can be considered to be audibly perceivable by virtue of the output signals from the output module 106.
  • the audible feedback is provided to the speaker 208 and that the audible feedback corresponds to the speaker 208 hearing the adjusted speech output signal as being output from the output module 106 (e.g., from the one or more speaker drivers).
  • the speaker 208 is essentially hearing his/her own voice which has been adjusted to be associated with the emotion of, for example, anger from his/her, for example, original tired sounding voice which can be associated with the emotion of calmness as received by the input portion 202 (e.g., microphone).
  • an audible feedback of the adjusted speech output signal can be considered to be provided to the speaker 208.
  • speaker emotion is audibly perceivable to be adjusted, in accordance with the adjusted one or more voice parameters, when the adjusted speech output signal is output.
  • the speaker 208 is capable of receiving the audible feedback of the adjusted speech output signal and further adjusting the one or more voice parameters in a manner so as to accentuate speaker emotion.
  • the input module 102 can include the analyzer portion 206
  • the analyzer portion 206 can be included in the processing module 104.
  • the processing module 104 can include the analyzer portion 206 (i.e., the analyzer portion 206 is absent from the input module 102).
  • the analyzer portion 206 can correspond to a sub-processor within the processing module 104 and can operate in analogous manner per earlier discussion with reference to Fig. 2.
  • both the input module 102 and the processing module 104 can include an analyzer portion which can operate in like manner to the analyzer portion 206 per earlier discussion with reference to Fig. 2.
  • the system 100 may also rely on facial visual (e.g., as captured in real-time by the earlier mentioned camera based tracker) of the speaker 208 to aid in real-time processing of the speech input signal so as to produce the adjusted speech output signal.
  • facial visual e.g., as captured in real-time by the earlier mentioned camera based tracker
  • This can be accomplished by detecting facial features and morphing the detected facial features based in data which correlates facial features with emotions (see, for example, "Real-time classification of evoked emotions using facial feature tracking and physiological responses. Bailenson J. N, Pontikakis E. D., Maus I. B. et a I, Int. J. Human-Computer Studies 66 (2008) 303-317.”).

Landscapes

  • Engineering & Computer Science (AREA)
  • Quality & Reliability (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Circuit For Audible Band Transducer (AREA)

Abstract

A method for signal processing of voice of a speaker is provided. The voice of the speaker can correspond to audio signals including one or more voice parameters associable with emotion of the speaker. The method can include receiving the audio signals and processing the audio signals by adjusting one or more of the one or more voice parameters to produce an adjusted speech output signal. The method can further include providing an output of the adjusted speech output signal so that the adjusted speech output signal is audibly perceivable and providing an audible feedback of the adjusted speech output signal to the speaker. The speaker is capable of receiving the audible feedback of the adjusted speech output signal and further adjusting the one or more voice parameters in a manner so as to accentuate speaker emotion.

Description

A METHOD FOR SIGNAL PROCESSING OF VOICE OF A SPEAKER
Field Of Invention
The present disclosure generally relates to method for signal processing of voice of a speaker so that emotion of the speaker can be audibly perceived to be adjusted.
Background
Conventional research relating to speech emotion (i.e., emotion(s) expressed in speech) is based on off-line processing. Additionally, speech features such as pitch contour and speech rate are typically used on classification.
Examples of the aforementioned research are:
[1] An acoustic study of emotions expressed in speech, S. Yildirim, M. Bulut, C. M. Lee et al., Proceedings of the International Conference on Spoken Language Processing (ICSLP Ό4), vol. 1, pp. 2193-2196, 2004; and
[2] Vocal expression of emotion. R. J. Davidson, K. R. Scherer, H. Goldsmith (Eds.). Handbook of the Affective Sciences (pp. 433-456). Scherer, K. R., Johnstone, T. & Klasmeyer, G. (2003). New York and Oxford: Oxford University Press
Therefore, it is notable that conventional research relating to speech emotion fails to provide or suggest a way for real-time processing (i.e., as opposed to off-line processing).
Summary of the Invention
In accordance with an aspect of the disclosure, a method for signal processing of voice of a speaker is provided. The voice of the speaker can correspond to audio signals which include one or more voice parameters. The voice parameter(s) can be associated with emotion of the speaker.
The method can include receiving the audio signals, processing the audio signals, providing an output and providing an audible feedback.
The audio signals can be processed by adjusting one or more of the one or more voice parameters to produce an adjusted speech output signal. In regard to providing an output, an output of the adjusted speech output signal can be provided so that the adjusted speech output signal is audibly perceivable.
In regard to providing an audible feedback, an audible feedback of the adjusted speech output signal can be provided to the speaker.
Speaker emotion can be audibly perceived to be adjusted, in accordance with the adjusted one or more voice parameters, when the adjusted speech output signal is output.
Moreover, the speaker is capable of receiving the audible feedback of the adjusted speech output signal and further adjusting the one or more voice parameters in a manner so as to accentuate speaker emotion.
Brief Description of the Drawings
Embodiments of the disclosure are described hereinafter with reference to the following drawings, in which:
Fig. 1 shows a system which can include an input module, a processing module and an output module, according to an embodiment of the disclosure;
Fig. 2 shows the system of Fig. 1 in further detail, according to an embodiment of the disclosure;
Fig. 3 shows exemplary signal waveforms associated with the system of Fig. 1, according to an embodiment of the disclosure; and
Fig. 4 shows a flow diagram illustrating a method for signal processing in association with the system of Fig. 1, according to an embodiment of the disclosure.
Detailed Description
Representative embodiments of the disclosure, for addressing one or more of the foregoing problems are described.
As mentioned earlier, conventional research relating to speech emotion (i.e., emotion(s) expressed in speech) is based on off-line processing. The present disclosure contemplates that, in the context of speech emotion, real-time processing is avoided owing to the need to rely on predictive data (i.e., "future" data). Moreover, the present disclosure contemplates that there is a possibility that speech utterance may be rendered unnatural sounding during real-time processing (e.g. if speech rate is to be changed).
However, the present disclosure contemplates that real-time processing would be useful for realtime use type situation(s).
An example of a real-time use type situation is where a speaker desires a listener/listener(s) to perceive/discern, via the emotions expressed in the speaker's speech, the speaker's intended emotion (also, otherwise, referable to as "true" emotion) while the speaker is speaking. For example, in a classroom situation where a teacher is tired and is attempting to maintain order in the classroom, it is appreciable that the tired teacher's speech may not be capable of conveying the intended verbal firmness (i.e., intended emotion or "true" emotion). Therefore, it may be useful for the teacher that students in the classroom are capable of perceiving the intended verbal firmness from the teacher's speech despite the teacher's tiredness.
Another example of a useful real-time situation is where a speaker does not desire the listener/listener(s) to be able to perceive/discern, via the emotions expressed in the speaker's speech, the speaker's intended/"true" emotion while the speaker is speaking. For example, during a meeting session or during a teleconference session, it is generally a good idea to maintain a calm/neutral tone while having a discussion with a counterpart/counterpart(s). However, while having a sensitive/stressful discussion, a speaker's speech may inadvertently express negative emotions such as anger or frustration (i.e., "true" emotion). Therefore, it may be useful for the speaker that the speaker's true emotion/emotion(s) is/are not conveyed to his/her counterpart(s) during the teleconference session. Appreciably, a calm/neutral tone masks the speaker's true emotion (e.g., anger/frustration) and may help facilitate a favorable (or, at least, a more fruitful) outcome for the teleconference session.
Other examples are also useful. For example, during a karaoke session, it may be desired for a singer to sing in a manner matching the intended emotion of the song track being sung. In this regard, there is a need to overcome the earlier contemplated problems associated with realtime processing so that application in real-time (i.e., real-time use type situation(s)) can be made possible. This will be discussed in further detail with reference to Fig. 1 to Fig. 4 hereinafter.
Referring to Fig. 1, a system 100 is shown in accordance with an embodiment of the disclosure. The system 100 can include an input module 102, a processing module 104 and an output module 106. The input module 102 can be coupled to the processing module 104. The processing module 104 can be coupled to the output module 106. As an option, the output module 106 can be coupled to the input module 102.
The input module 102 can be configured to generate input signals and communicate the input signals to the processing module 104 for processing.
The processing module 104 can be configured to receive and process the input signals to produce processed input signals. The processing module 104 can be further configured to communicate the processed input signals to the output module 106.
The output module 106 can be configured to receive and process the processed input signals in a manner so as to produce output signals.
In one embodiment, the system 100 can correspond to an ecosystem in which the input module 102, the processing module 104 and the output module 106 are three individual/independent devices which can be coupled as described earlier.
In another embodiment, the system 100 can correspond to a printed circuit board (PCB) on which the input module 102, the processing module 104 and the output module 106 are placed/positioned and coupled as described earlier. Specifically, the input module 102, the processing module 104 and the output module 106 can be three individual/independent elements carried by the PCB.
In yet another embodiment, the system 100 can correspond to a device such as an integrated circuit (IC) chip which includes the input module 102, the processing module 104 and the output module 106. Specifically, the input module 102, the processing module 104 and the output module 106 can form parts, which are coupled as described earlier, of the device (e.g., IC chip). Other examples are also useful. In one example, in the system 100, the input module 102 and the processing module 104 can be parts of a device (i.e., a first device), and the output module 106 can be another device (i.e., a second device) coupled thereto. In another example, in the system 100, the processing module 104 and the output module 106 can be parts of a device (i.e., a first device), and the input module 102 can be another device (a second device) coupled thereto. In yet a further example, the input module 102 and the output module 106 can be parts of a device (i.e., a first device), and the processing module 104 can be another device (i.e., a second device) coupled thereto. Specifically, the first device can be coupled to the second device.
Additionally the aforementioned coupling can be based on one or both of wired coupling and wireless coupling.
Fig. 2 shows the system 100 of Fig. 1 in further detail, in accordance with an embodiment of the disclosure. As shown, the input module 102 can include an input portion 202 and a control portion 204. The input module 102 can, as an option, further include an analyzer portion 206.
Each of the input portion 202 and the control portion 204 can be coupled to the processing module 104. Moreover, the input portion 202 can, optionally, be coupled to one or both of the control portion 204 and the analyzer portion 206. Additionally, the analyzer portion 206 can, optionally, be coupled to the processing module 104.
The input portion 202 can be configured to receive audio signals. The audio signals can, for example, correspond to the voice of a speaker 208. Specifically, the audio signals can, for example, correspond to speaker voice (i.e., of the speaker 208). In this regard, the input portion 202 can, for example, be a microphone which can be configured to capture and receive the voice of a speaker 208 who may, for example, be speaking into the microphone. The input portion 202 can be further configured to process the received audio signals to produce source signals which can be communicated to one or both of the processing module 104 and the analyzer portion 206 for further processing.
The control portion 204 can be configured to generate control signals and communicate the control signals to one or both of the processing module 104 and the input portion 202. In general, the control signals can be communicated from the control portion 204 to the processing module 104 to control the manner in which the processing module 104 processes the source signals communicated from the input portion 202 as will be discussed later in further detail in the context of an exemplary scenario. Additionally, the control portion 204 can be configured to communicate control signals to the input portion 202 to control the manner in which the input portion 202 receives and/or processes the audio signals. For example, control signals can be communicated from the control portion 204 to the input portion 202 to control sensitivity of the input portion 202 (e.g., microphone input sensitivity). In this regard, the control portion 202 can, for example, be a controller device having one or more controls. Examples of a control can include button(s), knob(s), slider(s), touch panel(s) and/or sensor(s). Moreover, examples of a sensor can include an accelerometer, a gyroscope, a magnetometer and a camera based tracker. The control(s) can, for example, be operated by the speaker 208 to generate control signals. For example, the speaker 208 can operate the slider (e.g., by adjusting the slider) to generate control signals. Therefore, the control portion 204 can, for example, be in the form of an "N" Dimensional (ND) controller device where "N" represents an integer and the integer represented by "N" can, for example, correspond to the number of controls available on the control portion 202 or the number of types of controls available on the control portion 202. In one example, where the control portion 202 features controls such as a knob and two sliders, the control portion 202 can be considered to be a 3 Dimensional controller device. In another example, where the control portion 202 features only one type control (e.g., two sliders), the control portion 202 can be considered to be a 1 Dimensional controller device.
The analyzer portion 206 can be configured to receive the source signals from the input portion 202. The analyzer portion 206 can be further configured to analyze the source signals to generate analysis signals. The analyzer portion 206 can yet be further configured to communicate the analysis signals to the processing module 104. In this regard, the analyzer portion 206 can, for example, correspond to a processor device configured to compute the source signals by way of analyzing to produce analysis signals. For example, the analyzer portion 206 can correspond to a processor device configured to perform the task of data mining (e.g., by way of analyzing) on the received source signals and produce analysis signals based on mined data. As will be discussed later in further detail in the context of an exemplary scenario, the processing module 104 can be configured to process one or both of the received control signals and the received source signals based on the analysis signals.
Appreciably, the input signals communicated from the input module 102 can include any one of the source signals, the control signals and the analysis signals, or any combination thereof. The processing module 104 can be configured to receive the source signals, the control signals and/or the analysis signals for further processing to produce processed input signals which can, in turn, be further communicated (i.e., from the processing module 104) to the output module 106 for further processing to produce output signals.
For the sake of clarity, the system 100 will be discussed in further detail based on an exemplary scenario hereinafter.
In one exemplary scenario, the input portion 202 can correspond to a microphone, the processing module 104 can correspond to a processor device (e.g., in the form of an integrated circuit chip) and the output module 106 can correspond to one or more speaker drivers capable of sound output. Moreover, the speaker 208 can, for example, be a natural person generating and providing audio signals to the system 100 by way of speaking into the microphone. In this regard, the voice of the speaker 208 can correspond to the audio signals received by the input portion 202.
The voice of the speaker 208 can include one or more voice parameters which can be indicative of the emotion (e.g., anger, joy, sadness, surprise and calmness/neutral) of the speaker 208. In this regard, the one or more voice parameters can be associated with emotion of the speaker 208.
Examples of the aforementioned voice parameter(s) can include pitch contour, speech rate, formant, pitch & variability and energy.
Appreciably, the input portion 202 can be configured to process the audio signals to produce source signals. In this regard, the source signals can, by the same token, include one or more voice parameters. For example, audio signals, which can be analog based (i.e., voice of the speaker), can be received by the microphone (i.e., the input portion 202) and processed by the microphone to produce signals in a form (i.e., source signals) suitable for further processing by the processing module 104 and/or the analyzer portion 206. More specifically, the microphone can process the received audio signals by way of conversion to electric signals. Therefore, the source signals can, for example, correspond to electric signals which are based on conversion from the received audio signals which may be analog type signals.
As mentioned earlier, the source signals can be further communicated to the processing module 104 for further processing. Preferably, the processing module 104 is configured to process the source signals by adjusting one or more of the voice parameters. Adjustment of the voice parameter(s) can, for example, be based on either the control signals or a combination of both of the control signals and the analysis signals. In this regard, the processing module 104 can be configured to process the source signals by adjusting one or more of the voice parameters based on, for example, either the control signals or a combination of the control signals and the analysis signals so as to produce processed source signals. The processed source signals can correspond to the aforementioned processed input signals in Fig. 1.
The processed source signals can be communicated from the processing module 104 to the output module 106 which can process the processed source signals in a manner so as to produce output signals. Additionally, as mentioned earlier, the output module 106 can correspond to one or more speaker drivers capable of sound output. In this regard, the output module 106 can be configured to receive and process the processed source signals in a manner so as to produce output signals which can be audibly perceived.
Since the source signals are based on the voice of the speaker 208 and can include one or more voice parameters, it is appreciable that the source signals can correspond to a speech input signal which can include one or more voice parameters. The speech input signal can be associated with one or more speaker emotions (i.e., of the speaker 208) based on the one or more voice parameters. By adjusting one or more voice parameters (i.e., associable with speaker emotion) at the processing module 104, the speech input signal can be considered to be adjusted in the context of speaker emotion. Therefore, output signals, which can be audibly perceived, communicated from the output module 106 can correspond to an adjusted speech output signal.
Therefore, the adjusted speech output signal, when output by the output module 106, can correspond to an audibly perceivable output signal where speaker emotion can be audibly perceived to be adjusted in accordance with the adjusted one or more voice parameters. Specifically, audibly perceivable speaker emotion associated with the adjusted speech output signal can be considered to be adjusted relative to speaker emotion associated with the input speech signal. The input speech signal is based on voice of the speaker 208.
Preferably, as signified by arrow 210 in Fig. 2, an audible feedback of the adjusted speech output signal can be provided to the speaker 208. In this regard, the speaker 208 can be capable of receiving an audible feedback of the adjusted speech output signal. The audible feedback can, for example, be based on the speaker 208 hearing the adjusted speech output signal as being output from the output module 106. Appreciably, providing an audible feedback of the adjusted speech output signal to the speaker 208 can be useful in facilitating further adjustment of the one or more voice parameters. The one or more voice parameters can be further adjusted in a manner so as to accentuate, for example, an intended or a desired emotion (i.e., which can be audibly perceived) of the speaker 208. More specifically, the speaker 208 can be capable of receiving an audible feedback of the adjusted speech output signal and further adjusting the one or more voice parameters in a manner so as to accentuate speaker emotion. Yet more specifically, the speaker 208 can be capable of hearing the adjusted speech output signal as being output by the output module 106 and further adjusting, as desired by the speaker 208, the one or more voice parameters in a manner so as to accentuate speaker emotion.
For the sake of further clarity, the above discussed exemplary scenario will be put in context based on an exemplary real-time use type situation hereinafter.
Taking, as an example, a real-time use type situation where a teacher (i.e., the speaker 208) wishes to instill discipline during a lecture session in a lecture theatre/hall, it may be useful for the teacher (who may be tired after a long day) that students in the lecture theatre are able to perceive an emotion (i.e., from the teacher) associated with anger.
The teacher using/operating the input module 102 may speak into the microphone (i.e., the input portion 202) and generate control signals using the control portion 204 (e.g., by sliding the slider). The teacher's voice may be tired sounding. The tired sounding voice (audio signals which can be received by the microphone) may be associated with the emotion of calmness. Hence, the tired sounding voice of the teacher can be processed by the microphone to produce the speech input signal (i.e., source signals). In this regard, the speech input signal can be associated with the emotion of calmness which may not be the intended emotion of the teacher and, further, not ideal for the purpose of instilling discipline.
Appreciably, there may be a need for the teacher to make an adjustment.
Preferably, adjustment can be made by way of processing speech input signals based on control signals. The control signals can be generated based on the intent of the teacher to make an adjustment to the speech input signal which can be associated with the emotion of calmness so as produce an adjusted speech output signal which can be audibly perceived to be associated with the emotion of anger. To aid the teacher in using the control portion 204 to generate control signals as appropriate, labels can be provided. For example, a label can be provided at one end of the slider indicating "calm voice" and another label can be provided at another end of the slider indicating "extremely angry voice". In this manner, the teacher can slide the slider accordingly to vary the degree/strength of anger (i.e., more towards "calm voice" or more towards "extremely angry voice") which can be audibly perceived via the adjusted speech output signal.
Alternatively, adjustment can be made by way of processing speech input signals based on a combination of the control signals and the analysis signals. As mentioned earlier, the analyzer portion 206 can be configured to analyze the source signals (i.e., speech input signals) to generate analysis signals. Specifically, the analyzer portion 206 can be configured to analyze the teacher's voice and extract one or more voice characteristics such as pitch and formant. Based on the voice characteristic(s), it is possible for the analyzer portion 206 to determine useful information associated with the speaker 208 (i.e., the teacher) and generate analysis signals accordingly. An example of useful information associated with the speaker 208 is the gender of the speaker 208. For example, based on the voice characteristic(s) such as the pitch, it can be determined that the teacher is a female. When it is determined that the teacher is a female, the analyzer portion 206 can be configured to communicate analysis signals indicative of the teacher's gender being a female. The analysis signals can be used for adjusting the control signals as will be discussed later in further detail with reference to Fig. 3. Therefore, the speech input signals can be considered to be processed based on a combination of the control signals and the analysis signals. Additionally, the control signals can be adjusted by the analysis signals at either the control portion 204 or the processing module 104. Moreover, it is also possible for the control signals to be adjusted by the analysis signals at both the control portion 204 and the processing module 104.
The speech input signals can be communicated to the processing module 104 for further processing based on either the control signals or a combination of the control signals and the analysis signals. Specifically, one or more voice parameters of the speech input signal which can be associated with the emotion of anger (i.e., intended emotion of the speaker 208) can be adjusted based on the control signals or a combination of the control signals and the analysis signals. Therefore, adjusting the voice parameter(s) based on the control signals can be the basis for a form of real-time processing to produce an adjusted speech output signal in which adjustment has been made in the context of the speaker emotion (i.e., associated with anger in this exemplary real-time use type situation). Moreover, since it is also possible for the processing module 104 to process the speech input signals based on a combination of the control signals and analysis signals (communicated from the analyzer portion 206) to produce the adjusted speech output signal, the analysis signals can be considered to be another basis for a form of real-time processing to produce an adjusted speech output signal in which adjustment has been made in the context of the speaker motion (i.e., associated with anger in this exemplary real-time use type situation). Processing of the speech input signal based on the control signals or a combination of the control signals and the analysis signals will be discussed later in further detail with reference to Fig. 3.
In this regard, it is appreciable that the speech input signals which can be associated with the emotion of calmness (e.g., owing to the tiredness of the teacher) can be adjusted by the teacher so that the adjusted speech output signals can be audibly perceived to be associated with the emotion of anger.
Therefore in this exemplary real-time use type situation, the teacher can be speaking into the microphone (i.e., the input portion 202) with a tired sounding voice (which can be associated with the emotion of calmness). However, instead of audibly perceiving the emotion of calmness, the students hearing the adjusted speech output signal from the one or more speaker drivers may audibly perceive the adjusted speech output signal to be associated with the emotion of anger.
Appreciably processing of the speech input signal to produce the adjusted speech output signal is based on real-time based processing.
Specifically, as discussed thus far, real-time processing of the received audio signals (i.e., voice of the teacher in accordance with the above exemplary real-time use type situation) to produce adjusted speech output signals can be made possible using the system 100 by virtue at least one of the following approaches:
1) Processing the received audio signals based on the control signals.
2) Processing the received audio signals based on the analysis signals in combination with the control signals. Appreciably, apart from the two approaches mentioned above, real-time processing is also possible by providing an audible feedback (i.e., as signified by arrow 210) of the adjusted output signal to the speaker 208 (i.e., the teacher). Audible feedback of the adjusted output signal can be provided to the teacher via the one or more speaker drivers (i.e., the output module 106). Preferably, audible feedback corresponds to the teacher hearing the adjusted speech output signal as being output from the one or more speaker drivers (i.e., the output module 106). Thus, based on the adjusted speech output signal as being output by the one or more speaker drivers, the teacher is essentially hearing her (assuming the gender of the teacher is a female as discussed earlier) own voice which has been adjusted to be associated with the emotion of anger from of her original tired sounding voice (associated with the emotion of calmness) as received by the microphone (i.e. the input portion 202). In this regard, the speech input signal can be associated with an "original/input voice" (e.g., tired sounding voice as received by the microphone) and the adjusted speech output signal can correspond to a "modified voice" (e.g., angry sounding voice as being output by the one or more speaker drivers) relative to the "original/input voice".
The teacher, upon, hearing her own voice as output via the one or more speaker drivers (i.e., receiving audible feedback), may decide that the emotion of anger may not be clearly audibly perceived enough. The teacher may then decide to make further adjustment(s) to the one or more voice parameters so as to accentuate the emotion of anger.
The manner of further adjustment(s) can, for example, be by way of the teacher adjusting her own voice (i.e., speaker voice) when speaking into the microphone to be more angry sounding so as to accentuate the intended or the desired emotion of, for example, anger. Thus the teacher may realize that she has to make an effort to be less tired sounding when speaking into the microphone.
Other ways of further adjustment are also possible.
In one example, the teacher can make further adjustments via the control portion 204 by sliding the slider even more towards the "extremely angry voice" direction. This can be akin to dynamically adjusting the strength/impact of the effect of, for example, anger by making adjustment(s) (e.g., sliding the slider) using the control portion 204. Hence, effectively, further adjustment can be by manner of adjusting control signals generated and communicated to the processing module 104. In another example, further adjustments can be by manner of both adjusting speaker voice communicated to, and received by, the microphone and adjusting control signals generated, and communicated to, the processing module 104.
Therefore, apart from the two earlier approaches mentioned above, it is also possible to aid realtime processing by providing audible feedback of the adjusted speech output signal to the speaker 208 so that the speaker 208 can make further adjustment(s), as desired or appropriate, to one or more voice parameters in a manner so as to accentuate speaker emotion as audibly perceived via the output module 106.
Hence it is appreciable that the system 100 effectively enables a user (e.g., the speaker 208) to change/adjust perceived (i.e., as perceived by one or more listeners such as the aforementioned students) emotion associated with the user's speech. As mentioned earlier as an example, the original tired sounding voice of the speaker 208 (e.g., a teacher) which can be associated with an emotion of calmness can be adjusted/changed so that listeners (e.g., the students) can audibly perceive an emotion of anger instead.
Furthermore, since the system 100 can be considered to be based on real-time processing by virtue of at least one of the following:
1) Processing the received audio signals based on the control signals; and
2) Processing the received audio signals based on the analysis signals in combination with the control signals,
there is no need to rely on predictive data (i.e., "future" data) as adjustment is made possible in realtime. Moreover, by providing audible feedback of the adjusted speech output signal to the speaker 208 so that the speaker 208 can make further adjustment(s), as desired or appropriate, to one or more voice parameters in a manner so as to accentuate speaker emotion as audibly perceived via the output module 106, it is also appreciable that there is no need to rely on predictive data (i.e., "future" data) as real-time compensation by the speaker 208 is made possible. It is also appreciable that based on the discussed approach(es), the possibility of speech utterance being rendered unnatural sounding during real-time processing (e.g. if speech rate is to be changed) can be significantly reduced since real-time adjustment and/or real-time compensation can be possible. Earlier mentioned, the speech input signal (i.e., source signals based on the audio signals) can be processed based on the control signals or a combination of the control signals and the analysis signals. This will be discussed in further detail with reference to Fig. 3 hereinafter.
Fig. 3 shows exemplary signal waveforms associated with the system 100, in accordance with an embodiment of the disclosure.
Specifically, Fig. 3 shows an exemplary speech input signal waveform 310 and a corresponding data waveform 320 of a voice parameter. Fig. 3 also shows an adjusted data waveform 330 where the voice parameter has been adjusted and an exemplary speech output signal waveform 340.
The voice parameter can, for example, relate to pitch. Therefore, the data waveform 320 can show the pitch data associated with the speech input signal (i.e., as represented by the exemplary speech input signal waveform 310). The pitch data associated with the speech input signal can be indicative of audibly perceivable pitch (e.g., the degree of height or depth of voice tone) of the speaker's 208 voice. The data waveform 320 can be derived from the speech input signal waveform 310 by manner of, for example, a sampling process or an extraction process which can be performed by the processing module 104.
As shown, the data waveform 320 can include a plurality of data portions (e.g., labels 320a, 320b and 320c). Each of the data portions can include one or more data points (e.g., labels 320d and 320e within data portion 320c)
In operation, in accordance with an embodiment of the disclosure, the speaker 208 can provide audio signals to the system 100 by, for example, speaking into the input portion 202. The input portion 202 processes the audio signals and produces the speech input signal (i.e., source signals) as represented by the exemplary speech input signal waveform 310. The speech input signal can be processed by the processing module 104 in a manner so that one or more data points (e.g., labels 320f, 320g and 320h) within one or more data portions (e.g., labels 320a, 320b and 320c) is/are adjusted. Adjustment of the data point(s) can be based on instructions communicated to the processing module 104 in the form of the aforementioned control signals. In this manner, the speech input signal can be processed based on the control signals. As mentioned earlier, the speech input signal can, as an option, be processed based a combination of the control signals and the analysis signals. This will be discussed in further detail using the earlier discussed example of the analysis signals being indicative of the gender of the speaker 208.
The analysis signals can be used to adjust the control signals (at the control portion 204 and/or the processing module 104), so that one or more of the data points can be adjusted, for example, to either a higher degree in or a lesser degree in comparison to the data point(s) being adjusted based on the control signals alone.
For example, where the speaker 208 is a female teacher (e.g., in the context of the earlier discussed exemplary real-time use type situation), it may be desired for data point 320f to be adjusted by a smaller degree as indicated by arrow 350. Conversely, where the speaker 208 is, for example, a male teacher, it may be desired for data point 320f to be adjusted by a larger degree as indicated by arrow 360.
In this above manner, the speech input signal can be processed based a combination of the control signals and the analysis signals. Specifically, the analysis signals can be used to adjust the control signals and the adjusted control signals can, in turn, be used to adjust the speech input signal to produce the adjusted speech output signal.
Referring to Fig. 4, a flow diagram 400 illustrating a method for signal processing in association with the system 100 is shown in accordance with an embodiment of the disclosure. The method for signal processing can be directed at the signal processing of voice of a speaker (e.g., earlier discussed speaker 208). The voice of the speaker can correspond to audio signals which include one or more voice parameters. The one or more voice parameters can be associated with emotion of the speaker.
The method can include receiving audio signals 410, processing audio signals to produce an adjusted speech output signal 420, providing an output of an adjusted speech output signal 430 and providing an audible feedback of the adjusted speech output signal 440.
In regard to the step of receiving audio signals 410, the audio signals can be received by the input portion 202 as discussed earlier with reference to the system 100. In regard to the step of processing audio signals to produce an adjusted speech output signal 420, one or more voice parameters associated with the audio signals can be adjusted to produce an adjusted speech output signal. As mentioned earlier with reference to the system 100, the received audio signals can be processed by the input portion 202 to produce an input speech signal which can include one or more voice parameters (i.e., corresponding to the one or more voice parameters associated with the audio signals as received by the input portion 202). The input speech signal can be communicated to the processing module 104 for further processing. Specifically, the processing module 104 can be configured to process the speech input signal (i.e., source signals based on the audio signals) based on the control signals or a combination of the control signals and the analysis signals to produce the adjusted speech output signal.
In regard to the step of providing an output of an adjusted speech output signal 430, the adjusted speech output signal can, as discussed earlier with reference to the system 100, be communicated to the output module 106 which can be configured to process the speech output signal in a manner so as to produce corresponding output signals which can be audibly perceived. In this regard, the adjusted speech output signal can, by virtue of the output signals from the output module 106, be considered to be capable of being audibly perceived. More specifically, the adjusted speech output signal can be considered to be audibly perceivable by virtue of the output signals from the output module 106.
In regard to the step of providing an audible feedback of the adjusted speech output signal 440, it is preferable, in accordance with an embodiment of the disclosure, the audible feedback is provided to the speaker 208 and that the audible feedback corresponds to the speaker 208 hearing the adjusted speech output signal as being output from the output module 106 (e.g., from the one or more speaker drivers). Thus, for example, based on the adjusted speech output signal as being output by the one or more speaker drivers, the speaker 208 is essentially hearing his/her own voice which has been adjusted to be associated with the emotion of, for example, anger from his/her, for example, original tired sounding voice which can be associated with the emotion of calmness as received by the input portion 202 (e.g., microphone). In this regard, an audible feedback of the adjusted speech output signal can be considered to be provided to the speaker 208.
Preferably, speaker emotion is audibly perceivable to be adjusted, in accordance with the adjusted one or more voice parameters, when the adjusted speech output signal is output. Also preferably, the speaker 208 is capable of receiving the audible feedback of the adjusted speech output signal and further adjusting the one or more voice parameters in a manner so as to accentuate speaker emotion.
In the foregoing manner, various embodiments of the disclosure are described for addressing at least one of the foregoing disadvantages. Such embodiments are intended to be encompassed by the following claims, and are not to be limited to specific forms or arrangements of parts so described and it will be apparent to one skilled in the art in view of this disclosure that numerous changes and/or modification can be made, which are also intended to be encompassed by the following claims.
For example, although it is earlier mentioned with reference to Fig. 2 that in accordance with an embodiment of the disclosure, the input module 102 can include the analyzer portion 206, it is appreciable that it is also possible that instead of the input module 102, the analyzer portion 206 can be included in the processing module 104. Specifically, in accordance with an embodiment of the disclosure, the processing module 104 can include the analyzer portion 206 (i.e., the analyzer portion 206 is absent from the input module 102). In such case, the analyzer portion 206 can correspond to a sub-processor within the processing module 104 and can operate in analogous manner per earlier discussion with reference to Fig. 2.
Of course, it is also possible that, in accordance with an embodiment of the disclosure, both the input module 102 and the processing module 104 can include an analyzer portion which can operate in like manner to the analyzer portion 206 per earlier discussion with reference to Fig. 2.
Additionally, the system 100 may also rely on facial visual (e.g., as captured in real-time by the earlier mentioned camera based tracker) of the speaker 208 to aid in real-time processing of the speech input signal so as to produce the adjusted speech output signal. This can be accomplished by detecting facial features and morphing the detected facial features based in data which correlates facial features with emotions (see, for example, "Real-time classification of evoked emotions using facial feature tracking and physiological responses. Bailenson J. N, Pontikakis E. D., Maus I. B. et a I, Int. J. Human-Computer Studies 66 (2008) 303-317.").

Claims

Claims
1. A method for signal processing of voice of a speaker, the voice of the speaker corresponding to audio signals comprising one or more voice parameters associabie with emotion of the speaker, the method comprising:
receiving the audio signals;
processing the audio signals by adjusting one or more of the one or more voice parameters to produce an adjusted speech output signal;
providing an output of the adjusted speech output signal so that the adjusted speech output signal is audibly perceivable; and
providing an audible feedback of the adjusted speech output signal to the speaker, wherein speaker emotion is audibly perceivable to be adjusted, in accordance with the adjusted one or more voice parameters, when the adjusted speech output signal is output, and wherein the speaker is capable of receiving the audible feedback of the adjusted speech output signal and further adjusting the one or more voice parameters in a manner so as to accentuate speaker emotion.
2. The method as in claim 1, wherein further adjustment of the one or more voice parameters to accentuate speaker emotion is based on the speaker adjusting speaker voice.
3. A device configured to receive and process voice of a speaker, the voice of the speaker corresponding to audio signals comprising one or more voice parameters associabie with emotion of the speaker, the device comprising:
an input portion configured to receive the audio signals;
a controller operable to produce control signals;
a processor configured to receive process the audio signals by adjusting, based on the control signals, one or more of the one or more voice parameters to produce an adjusted speech output signal; and
an output portion configured to provide an output of the adjusted speech output signal, wherein the one or more voice parameters are adjusted in a manner so that speaker emotion is audibly perceived to be adjusted when the adjusted speech output signal is output, and wherein the speaker is capable of receiving the audible feedback of the adjusted speech output signal and further adjusting the one or more voice parameters in a manner so as to accentuate speaker emotion.
PCT/SG2015/050519 2015-01-05 2015-12-31 A method for signal processing of voice of a speaker Ceased WO2016111644A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
SG11201705469XA SG11201705469XA (en) 2015-01-05 2015-12-31 A method for signal processing of voice of a speaker

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
SG10201500040U 2015-01-05
SG10201500040U 2015-01-05

Publications (1)

Publication Number Publication Date
WO2016111644A1 true WO2016111644A1 (en) 2016-07-14

Family

ID=55168333

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/SG2015/050519 Ceased WO2016111644A1 (en) 2015-01-05 2015-12-31 A method for signal processing of voice of a speaker

Country Status (2)

Country Link
SG (1) SG11201705469XA (en)
WO (1) WO2016111644A1 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10861483B2 (en) 2018-11-29 2020-12-08 i2x GmbH Processing video and audio data to produce a probability distribution of mismatch-based emotional states of a person

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20080147413A1 (en) * 2006-10-20 2008-06-19 Tal Sobol-Shikler Speech Affect Editing Systems
US20130238337A1 (en) * 2011-07-14 2013-09-12 Panasonic Corporation Voice quality conversion system, voice quality conversion device, voice quality conversion method, vocal tract information generation device, and vocal tract information generation method

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20080147413A1 (en) * 2006-10-20 2008-06-19 Tal Sobol-Shikler Speech Affect Editing Systems
US20130238337A1 (en) * 2011-07-14 2013-09-12 Panasonic Corporation Voice quality conversion system, voice quality conversion device, voice quality conversion method, vocal tract information generation device, and vocal tract information generation method

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
BAILENSON J. N; PONTIKAKIS E. D.; MAUS . B. E: "Real-time classification of evoked emotions using facial feature tracking and physiological responses", INT. J. HUMAN-COMPUTER STUDIES, vol. 66, 2008, pages 303 - 317, XP022762475, DOI: doi:10.1016/j.ijhcs.2007.10.011
S. YILDIRIM; M. BULUT; C. M. LEE ET AL., PROCEEDINGS OF THE INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE PROCESSING (ICSLP '04, vol. 1, 2004, pages 2193 - 2196
SCHERER, K. R.; JOHNSTONE, T.; KLASMEYER, G.: "Handbook of the Affective Sciences", 2003, OXFORD UNIVERSITY PRESS, pages: 433 - 456

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10861483B2 (en) 2018-11-29 2020-12-08 i2x GmbH Processing video and audio data to produce a probability distribution of mismatch-based emotional states of a person

Also Published As

Publication number Publication date
SG11201705469XA (en) 2017-07-28

Similar Documents

Publication Publication Date Title
Fletcher Can haptic stimulation enhance music perception in hearing-impaired listeners?
US8987571B2 (en) Method and apparatus for providing sensory information related to music
Agus et al. Fast recognition of musical sounds based on timbre
Nickerson et al. Teaching speech to the deaf: Can a computer help
JP7036014B2 (en) Speech processing equipment and methods
KR20150104345A (en) Voice synthesys apparatus and method for synthesizing voice
KR20150076128A (en) System and method on education supporting of pronunciation ussing 3 dimensional multimedia
EP4213130B1 (en) Device, system and method for providing a singing teaching and/or vocal training lesson
Huckvale et al. Avatar therapy: an audio-visual dialogue system for treating auditory hallucinations.
WO2016111644A1 (en) A method for signal processing of voice of a speaker
KR102060228B1 (en) A System Providing Vocal Training Service Based On Subtitles
JP2017198922A (en) Karaoke equipment
US20240203435A1 (en) Information processing method, apparatus and computer program
Athanasopoulos et al. King's speech: pronounce a foreign language with style
Athanasopoulos et al. 3D immersive karaoke for the learning of foreign language pronunciation
JP6775218B2 (en) Swallowing information presentation device
Brueggeman et al. Speaker Trait Enhancement for Cochlear Implant Users: A Case Study for Speaker Emotion Perception.
Lim et al. A musical robot that synchronizes with a coplayer using non-verbal cues
Pape et al. Cue-weighting in the perception of intervocalic stop voicing in European Portuguese
TWI852226B (en) Deaf people's interactive performance wearable device and method thereof
US12032807B1 (en) Assistive communication method and apparatus
JP2005099068A (en) Musical instrument and musical sound control method
JP5794507B1 (en) Sound improvement training device and sound improvement training program
KR200287165Y1 (en) language learning appliance for multi-person
Yabu et al. Supporting Communication for Individuals with Speech and Hearing Disorders

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15825643

Country of ref document: EP

Kind code of ref document: A1

DPE1 Request for preliminary examination filed after expiration of 19th month from priority date (pct application filed from 20040101)
WWE Wipo information: entry into national phase

Ref document number: 11201705469X

Country of ref document: SG

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 15825643

Country of ref document: EP

Kind code of ref document: A1