WO2025215744A1 - 装置および方法 - Google Patents
装置および方法Info
- Publication number
- WO2025215744A1 WO2025215744A1 PCT/JP2024/014418 JP2024014418W WO2025215744A1 WO 2025215744 A1 WO2025215744 A1 WO 2025215744A1 JP 2024014418 W JP2024014418 W JP 2024014418W WO 2025215744 A1 WO2025215744 A1 WO 2025215744A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- speaker
- voice
- information
- state
- detection unit
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/003—Changing voice quality, e.g. pitch or formants
- G10L21/007—Changing voice quality, e.g. pitch or formants characterised by the process used
Definitions
- One aspect of the present disclosure relates to an apparatus and a method.
- Patent Document 1 discloses a technology that, if a change in emotion is detected in a customer's voice during a phone conversation after the start of the conversation, changes the representative's voice in real time to accommodate the change in emotion, thereby performing voice conversion tuning.
- the state of one speaker can affect the sound and impression it gives, and the state of the other speaker listening to that sound can also affect how the sound is perceived.
- the other speaker may wish for smooth communication to be achieved, at least for the other speaker, regardless of the state of the other speaker. For this reason, there is a demand for technology that supports smooth communication by applying appropriate processing according to the state of each speaker.
- the present disclosure therefore aims to provide an apparatus and method that enables smooth conversation between speakers via audio terminals.
- the device disclosed herein comprises an acquisition unit that acquires the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal; a first detection unit that detects information related to the state of the first speaker based on the voice of the first speaker; a second detection unit that detects information related to the state of the second speaker based on the biometric information of the second speaker; and an adjustment unit that adjusts the voice based on the information related to the state of the first speaker and the information related to the state of the second speaker.
- smooth conversation between speakers can be realized via audio terminals.
- FIG. 1 is a block diagram showing the configuration of a processing system including an apparatus according to an embodiment of the present disclosure.
- FIG. 2 is a diagram showing an example of the structure of information about emotions, which is an example of information about a state.
- FIG. 3 is a diagram showing information about a reaction, which is an example of information about a state.
- FIG. 4 is a flowchart showing the procedure of an example of a processing method performed by the processing system.
- FIG. 5 is a diagram illustrating an example of a hardware configuration of an apparatus according to an embodiment of the present disclosure.
- FIG. 1 is a block diagram showing the configuration of a processing system including an apparatus according to one embodiment of the present disclosure.
- the processing system 1 shown in FIG. 1 includes a device 10, a first terminal 11, a second terminal 12, a detection device 13, and a user state database 14, which are configured to be able to communicate with each other via a network including a wireless communication network and a fixed communication network.
- the processing system 1 is used in a conversation between a first speaker using the first terminal 11, which is an example of a voice terminal, and a second speaker who converses via the second terminal 12, which is also an example of a voice terminal.
- the first speaker is, for example, a customer of a certain product or service.
- the second speaker is an operator belonging to an organization that is the recipient of an inquiry regarding the product or service.
- the device 10 adjusts at least one of the voice input to the first terminal 11 and the voice input to the second terminal 12 based on the state of the first speaker who is speaking toward the first terminal 11 and the state of the second speaker who is speaking toward the second terminal 12.
- the device 10 of this embodiment relays the voice of the first speaker from the first terminal 11 to the second terminal 12, and relays the voice of the second speaker from the second terminal 12 to the first terminal 11.
- the device 10 adjusts the voice input to the first terminal 11 and the voice input to the second terminal 12.
- the device 10 adjusts the voice based on information obtained from the detection device 13 and the user state database 14. Each component is described in detail below.
- the first terminal 11 and the second terminal 12 are audio terminals used when the first speaker and the second speaker intend to converse.
- the first terminal 11 and the second terminal 12 are devices such as personal computers, smartphones, tablet terminals, feature phones, server devices, game consoles, etc. Note that while FIG. 1 illustrates only one terminal each as the first terminal 11 and the second terminal 12, the processing system 1 may include any number of first terminals 11 and second terminals 12, two or more.
- the first terminal 11 acquires the voice of the first speaker (hereinafter referred to as the first voice).
- the second terminal 12 acquires the voice of the second speaker (hereinafter referred to as the second voice).
- the first terminal 11 and the second terminal 12 each output the acquired voice to the device 10.
- the first terminal 11 acquires the voice of the second speaker adjusted by the device 10 (hereinafter referred to as the second adjusted voice).
- the second terminal 12 acquires the voice of the first speaker adjusted by the device 10 (hereinafter referred to as the first adjusted voice). Note that if the device 10 does not adjust the second voice, the device 10 does not relay the second voice from the second terminal 12, and the second terminal 12 may output the second voice directly to the first terminal 11 via the network.
- Detection device 13 detects the biometric information of the second speaker.
- the biometric information of the second speaker detected by detection device 13 includes at least one of the following information: the voice of the second speaker, the biometric signal of the second speaker, and a video image of the second speaker.
- Detection device 13 includes, for example, a wearable device, a camera, etc. Detection device 13 is not limited to these.
- the wearable device of detection device 13 is worn on the body of the second speaker. For example, when the first speaker and the second speaker are conversing via processing system 1, the wearable device of detection device 13 detects a value related to at least one indicator of the second speaker's heart rate, breathing rate, activity level (amount of movement), and body temperature as the biometric information of the second speaker.
- the camera of detection device 13 captures an image of at least a part of the second speaker's body.
- the camera of detection device 13 may also acquire the voice of the second speaker. For example, when the first speaker and the second speaker are conversing via processing system 1, the camera of detection device 13 detects information regarding the movement of each part of the second speaker's body, such as the amount and direction of movement per unit time, and the second voice as biometric information of the second speaker. Movement of each part includes, for example, eye movement and whole-body movement.
- Detection device 13 outputs the detected biometric information of the second speaker.
- Detection device 13 stores the detected biometric information of the second speaker.
- device 10 may have the configuration and functions of detection device 13.
- the user state database 14 stores a table with indicators and thresholds that can be compared with information about the state of the first speaker (hereinafter referred to as first state information) and information about the state of the second speaker (hereinafter referred to as second state information). Details of the information stored in the user state database 14 will be described later.
- the device 10 is configured to include, as functional components, an acquisition unit 20, a detection unit 30, and an adjustment unit 40.
- the device 10 adjusts the first audio acquired from the first terminal 11.
- the device 10 outputs the first adjusted audio to the second terminal 12.
- the device 10 adjusts the second audio acquired from the second terminal 12.
- the device 10 outputs the second adjusted audio to the first terminal 11.
- the acquisition unit 20 acquires the first voice and the biometric information of the second speaker.
- the acquisition unit 20 has a voice acquisition unit 21 and a state acquisition unit 22.
- the voice acquisition unit 21 acquires the first voice input to the first terminal 11 by the first speaker from the first terminal 11.
- the voice acquisition unit 21 acquires the first voice by accepting the first voice transmitted from the first terminal 11.
- the voice acquisition unit 21 acquires the second voice input to the second terminal 12 by the second speaker from the second terminal 12.
- the voice acquisition unit 21 acquires the second voice by accepting the second voice transmitted from the second terminal 12.
- the status acquisition unit 22 acquires biometric information of the second speaker.
- the status acquisition unit 22 acquires the biometric information of the second speaker from the detection device 13.
- the status acquisition unit 22 may acquire information (second status information) about the state of the second speaker after the first voice or the first adjusted voice reaches the second terminal.
- the status acquisition unit 22 may output a signal to the detection device 13 to acquire the biometric information of the second speaker.
- the detection device 13 may acquire the biometric information of the second speaker using the signal as a trigger.
- the detection unit 30 detects first status information and second status information.
- the first status information and second status information include information regarding the emotions of the first speaker and the second speaker, respectively.
- the detection unit 30 has a first detection unit 31 and a second detection unit 32.
- the first detection unit 31 detects the first status information based on the first voice acquired by the voice acquisition unit 21.
- the first detection unit 31 analyzes an index indicating the voice quality indicated by the first voice.
- an index indicating voice quality the first detection unit 31 analyzes at least one index of the volume, pitch, and clarity indicated by the first voice. In this embodiment, the first detection unit 31 analyzes, for example, the volume, pitch, and clarity indicated by the first voice.
- the first detection unit 31 acquires information about emotions, which is an example of information about the state, from the user state database 14.
- the first detection unit 31 detects the emotion indicated by the first voice as first state information based on the indicators indicated by the first voice and the information about emotions acquired from the user state database 14.
- the first detection unit 31 detects the emotion indicated by the first voice by comparing the indicators of volume, pitch, pitch, and clarity indicated by the first voice with the indicators (rules) corresponding to predetermined emotions.
- FIG. 2 is a diagram showing an example of the structure of information about emotions, which is an example of information about a state.
- the information about emotions obtained from the user state database 14 may be an emotion table showing the correspondence between emotions and voice features (indicators).
- Each emotion such as “joy,” “anger,” or “sadness,” is correlated with an index indicating voice quality (for example, at least one index of "volume,” “pitch,” “pitch,” and “clarity” indicated by the voice).
- the numerical values shown in FIG. 2 indicate the reference value (%) when the maximum value of each index is 100 and the minimum value is 0.
- the first detection unit 31 analyzes the "volume,” “pitch,” “pitch,” and “clarity” indicated by the first voice as 70, 20, 75, and 40, respectively.
- the first detection unit 31 searches the emotion table acquired from the user state database 14 to determine which emotion corresponds to when "volume,” “pitch,” “pitch,” and “clarity” are 70, 20, 75, and 40.
- the first detection unit 31 estimates that the emotion of the first speaker includes anger. Therefore, the first detection unit 31 detects anger from the first voice as information related to the emotion of the first state information. In this way, the first detection unit 31 detects information related to the emotion of the first state information based on the values of each index indicated by the first voice and the emotion table.
- the first detection unit 31 does not have to detect the emotion indicated by the first voice by comparing an index indicating the voice quality of the first voice with each index (rule) corresponding to a predetermined emotion.
- the first detection unit 31 may input each index indicated by the first voice into a machine learning model within the device 10 or external to the device 10, and output information related to the emotion of the first status information.
- the machine learning model may include, for example, at least one of a generative AI model (generative artificial intelligence), a discrimination/determination type AI model, or a model combining these. Below, an example in which a generative AI model is used as the machine learning model will be described.
- the generative AI model outputs, as the first status information, the emotion estimated to be indicated by the first voice based on the index indicating the voice quality of the input first voice.
- the first detection unit 31 may acquire the first status information output by the generative AI model.
- the first detection unit 31 does not need to analyze an index indicating the voice quality of the first voice.
- the first detection unit 31 may input the first voice into a generative AI model to detect an index indicating the voice quality of the first voice.
- the generative AI model may directly output, as first state information, an emotion estimated to be indicated by the first voice based on the first voice input from the first detection unit 31.
- the index indicating voice quality does not need to include any of the indexes of volume, pitch, pitch, and clarity.
- the index indicating voice quality may include, for example, at least one of timbre, resonance, warmth, and degree of nasality.
- the second detection unit 32 detects second status information based on the second voice acquired by the voice acquisition unit 21 and the biometric information of the second speaker acquired by the status acquisition unit 22.
- the second detection unit 32 may detect the emotion indicated by the second voice as the second status information by processing similar to that of the first detection unit 31 described above.
- the second detection unit 32 analyzes an index indicating the voice quality indicated by the second voice.
- an index indicating voice quality the second detection unit 32 analyzes at least one index of the volume, pitch, pitch, and clarity indicated by the second voice.
- the second detection unit 32 acquires information regarding emotion, which is an example of information regarding the status, from the user status database 14.
- the second detection unit 32 detects the emotion indicated by the second voice as the second status information based on each index indicated by the second voice and the information regarding emotion acquired from the user status database 14. In this embodiment, the second detection unit 32 detects the emotion indicated by the second voice by comparing the volume, pitch, pitch, and clarity indicators indicated by the second voice with predetermined indicators (rules) related to voice corresponding to emotions.
- the second detection unit 32 acquires information about reactions, which is an example of information about the state, from the user state database 14.
- the second detection unit 32 detects the reaction indicated by the second voice as second state information based on the indicators indicated by the second speaker's biometric information and the information about the reaction acquired from the user state database 14.
- the second detection unit 32 detects the reaction indicated by the second voice by comparing the indicators of volume, pitch, pitch, and clarity indicated by the second voice with the indicators (rules) corresponding to predetermined reactions.
- FIG. 3 is a diagram showing information about reactions, which is an example of information about a state.
- the information about reactions obtained from the user state database 14 may be a reaction table showing the correspondence between states (reactions) and indicators of biometric information.
- Each reaction that the second speaker may have to a stimulus such as "excitement,” “impatience,” or “shock”
- biometric information may include information other than "heart rate,” “eye movement,” “breathing rate,” and “whole body movement.”
- the second detection unit 32 analyzes the second speaker's biometric information to determine whether the "heart rate,” "eye movement,” “breathing rate,” and “whole body movement” are 95, 20, 25, and 40, respectively.
- the second detection unit 32 searches the reaction table acquired from the user state database 14 to determine which emotion corresponds to the "heart rate,” “eye movement,” “breathing rate,” and “whole body movement” when they are 95, 20, 25, and 40.
- the second detection unit 32 estimates that the second speaker's reaction includes impatience. Therefore, the second detection unit 32 detects an impatient reaction from the second voice as information related to the reaction of the second state information. In this way, the second detection unit 32 detects information related to the reaction of the second state information based on the values of each index indicated by the second speaker's biometric information and the reaction table.
- the second detection unit 32 does not have to detect at least one of the emotions and reactions indicated by the second voice by comparing an index indicating the voice quality of the second voice with each index (rule) corresponding to predetermined emotions and reactions.
- the second detection unit 32 may input the values of each index indicated by the second speaker's biometric information into a generative AI model (an example of a machine learning model) within or external to the device 10, and output at least one of information related to emotions and information related to reactions in the second status information.
- the generative AI model outputs at least one of the emotions and reactions estimated to be indicated by the second voice as second status information based on the input index indicating the voice quality of the second voice.
- the second detection unit 32 may acquire the second status information output by the generative AI model.
- the second detection unit 32 does not need to analyze an index indicating the voice quality of the second voice.
- the second detection unit 32 may input the second voice into a generative AI model to detect an index indicating the voice quality of the second voice.
- the generative AI model may directly output at least one of the emotion and reaction estimated to be indicated by the second voice as second state information.
- the second detection unit 32 estimates the stress value of the second speaker based on the second state information.
- the stress value of the second speaker is an index that quantifies the level of stress of the second speaker.
- the second detection unit 32 estimates the stress value of the second speaker based on, for example, each index of the second voice including at least one of volume, pitch, pitch, and clarity, and at least one of emotional information and reaction information.
- the second detection unit 32 estimates that the stress value is higher the greater the degree of deviation between each index of the second voice including at least one of volume, pitch, pitch, and clarity and the reference value in the emotion table corresponding to the emotion of the second speaker detected by the second detection unit 32.
- the degree of deviation here is, for example, the average value of the difference between each index of the second voice and each reference value in the emotion table corresponding to the emotion of the second speaker, divided by each reference value.
- the degree of deviation here may be, for example, the sum of the differences between each index of the second voice and each reference value in the emotion table corresponding to the emotion of the second speaker.
- the second detection unit 32 estimates that the stress level is higher the greater the degree of deviation between each indicator of the second speaker's biometric information, including, for example, "heart rate,” “eye movement,” “breathing rate,” “whole body movement,” etc., and the reference value in the reaction table corresponding to the second speaker's reaction detected by the second detection unit 32.
- the degree of deviation here is, for example, the average value of the difference between each indicator of the second voice and each reference value in the reaction table corresponding to the second speaker's reaction, divided by each reference value.
- the degree of deviation here may also be, for example, the sum of the differences between each indicator of the second voice and each reference value in the reaction table corresponding to the second speaker's reaction.
- the second detection unit 32 does not have to detect the stress level of the second speaker by comparing the volume, pitch, pitch, and clarity indicators indicated by the second voice with reference values indicated in a predetermined emotion table and reaction table.
- the second detection unit 32 may input the indicators indicated by the second speaker's biometric information and the second state information (at least one of information related to emotions and information related to reactions) into a generative AI model (an example of a machine learning model) within or external to the device 10, and output the stress level of the second speaker.
- the generative AI model estimates the stress level of the second speaker based on the input information and outputs the stress level.
- the second detection unit 32 may also acquire the stress level output by the generative AI model.
- the adjustment unit 40 adjusts the audio based on the first status information and the second status information.
- the audio adjusted by the adjustment unit 40 is the audio of at least one of the first speaker and the second speaker. In this embodiment, the adjustment unit 40 adjusts both the first audio and the second audio.
- the adjustment unit 40 includes a voice conversion unit 41 and a voice output unit 42.
- the voice conversion unit 41 adjusts the first voice based on the first status information and the stress value of the second speaker. For example, if the information regarding the emotions of the first speaker in the first status information includes negative emotions such as "anger” and "sadness," the voice conversion unit 41 adjusts the first voice so that the degree of deviation between each indicator of the first voice, including volume, pitch, pitch, and clarity, and each reference value in the emotion table corresponding to the emotion of the first speaker detected by the first detection unit 31 becomes smaller.
- the voice conversion unit 41 adjusts the first voice so that the degree of deviation between each indicator of the first voice, including volume, pitch, pitch, and clarity, and each indicator of the voice in the emotion table indicated by the emotion of the first speaker detected by the first detection unit 31 becomes smaller.
- the voice conversion unit 41 may adjust the first voice even if the information regarding the emotion of the first speaker in the first status information does not include a negative emotion.
- the voice conversion unit 41 may adjust each indicator of the first voice, including volume, pitch, pitch, and clarity, to a value that does not correspond to any emotion item in the emotion table.
- the voice conversion unit 41 may adjust each indicator of the first voice according to the stress value of the second speaker, regardless of the information regarding the emotion of the first speaker in the first status information.
- the voice conversion unit 41 adjusts the second voice based on the first status information and the second status information. For example, if the information regarding the emotions of the first speaker in the first status information includes negative emotions such as "anger” and "sadness," the voice conversion unit 41 adjusts each indicator of the second voice, including volume, pitch, pitch, and clarity, so as to soothe the emotions of the first speaker detected by the first detection unit 31. At this time, the voice conversion unit 41 may adjust each indicator of the second voice, including volume, pitch, pitch, and clarity, so as to produce a calm second adjusted voice.
- the voice conversion unit 41 may also adjust each indicator of the second voice, including volume, pitch, pitch, and clarity, so as to reduce the degree of deviation between the indicators of the second voice, including volume, pitch, pitch, and clarity, and the reference values in the emotion table corresponding to the emotions of the second speaker detected by the second detection unit 32.
- the voice conversion unit 41 adjusts each indicator of the second voice, including volume, pitch, pitch, and clarity, to reduce the degree of deviation between the indicators and the reference values in the emotion table corresponding to the emotions of the second speaker detected by the second detection unit 32.
- the voice conversion unit 41 adjusts the indices of the second voice, including volume, pitch, pitch, and clarity, so that the degree of deviation between them and the reference values in the reaction table corresponding to the second speaker's reaction detected by the second detection unit 32 becomes smaller.
- the voice conversion unit 41 may adjust the second voice when the information regarding the emotion of the second speaker in the second status information does not include a negative emotion, and when the information regarding the reaction of the second speaker does not include a negative emotion.
- the voice conversion unit 41 may adjust each index of the second voice, including volume, pitch, and clarity, to a value that does not correspond to each emotion item in the emotion table.
- the voice conversion unit 41 may adjust each index of the second voice, including volume, pitch, and clarity, to a value that does not correspond to each reaction item in the reaction table.
- the voice conversion unit 41 may adjust each index of the first voice according to the stress value of the second speaker.
- the voice conversion unit 41 adjusts an index indicating the voice quality indicated by the first voice and the second voice.
- the voice conversion unit 41 adjusts, for example, at least one index of the volume, pitch, pitch, and clarity indicated by the first voice and the second voice.
- the voice conversion unit 41 adjusts at least one index included in the first voice and the second voice according to a preset priority.
- the priority may be determined in advance by, for example, the second speaker.
- the priority may be, for example, in descending order of the degree of deviation of each index from the index of the voice in the emotion table indicated by the detected emotion.
- the voice conversion unit 41 stores the value of the adjusted at least one index.
- the voice conversion unit 41 generates the first adjusted voice by adjusting the voice quality indicators indicated by the first voice. Note that the voice conversion unit 41 may, for example, adjust the first voice to generate the first adjusted voice that sounds like a person other than the first speaker is speaking, as with a voice changer.
- the voice output unit 42 outputs the first adjusted voice adjusted by the voice conversion unit 41 to the second terminal 12.
- the second terminal 12 receives the first adjusted voice.
- the voice conversion unit 41 generates the second adjusted voice by adjusting the voice quality indicators indicated by the second voice.
- the voice conversion unit 41 may, for example, adjust the second voice to generate the second adjusted voice that sounds like a person other than the second speaker is speaking, as with a voice changer.
- the voice output unit 42 outputs the second adjusted voice adjusted by the voice conversion unit 41 to the first terminal 11.
- the first terminal 11 receives the second adjusted voice.
- the voice conversion unit 41 may, for example, input at least one of the first voice and the second voice into a generative AI model (an example of a machine learning model) within or external to the device 10, and output a first adjusted voice and a second adjusted voice.
- the generative AI model may have learned at least one of the above-mentioned voice quality indicators, emotion tables, and reaction tables. For example, based on the input voice, the generative AI model adjusts the voice so that it diverges less from each indicator of one of the emotions shown in the emotion table, and outputs the adjusted voice.
- the adjustment unit 40 may acquire at least one of the first adjusted voice and the second adjusted voice output by the generative AI model.
- FIG. 4 is a flowchart showing the procedure of an example of a processing method by the processing system.
- the processing method shown in FIG. 4 is initiated, for example, when a call start signal is received from the first terminal 11 operated by the first speaker to the second terminal 12 operable by the second speaker, and the device 10 receives information from the second terminal 12 indicating that the signal has been received.
- the processing method may also be initiated at any timing during a conversation between the first and second speakers, when the device 10 receives a signal generated by an input operation by the second speaker to the second terminal 12.
- step S1 the speech acquisition unit 21 of the acquisition unit 20 acquires the speech of the first speaker.
- the speech acquisition unit 21 acquires, from the first terminal 11, the first speech input to the first terminal 11 by the first speaker.
- step S2 the state acquisition unit 22 of the acquisition unit 20 acquires the biometric information of the second speaker.
- the state acquisition unit 22 outputs a signal to the detection device 13 to acquire the biometric information of the second speaker.
- the detection device 13 uses the signal as a trigger to acquire the biometric information of the second speaker.
- the state acquisition unit 22 acquires the biometric information of the second speaker while the first speaker is speaking.
- the first detection unit 31 of the detection unit 30 detects first status information.
- the first detection unit 31 for example, acquires information about emotions, which is an example of status information, from the user status database 14.
- the first detection unit 31 detects the emotion indicated by the first voice as first status information based on the indicators indicated by the first voice and the information about emotions acquired from the user status database 14.
- the second detection unit 32 of the detection unit 30 detects second state information.
- the second detection unit 32 acquires at least one of information about emotions and information about reactions, which are examples of state information, from the user state database 14.
- the second detection unit 32 may detect the emotion indicated by the second voice as first state information based on each indicator indicated by the second voice and the information about emotions acquired from the user state database 14.
- the second detection unit 32 may detect the reaction indicated by the second speaker's biometric information as second state information based on each indicator indicated by the second speaker's biometric information and the information about reactions acquired from the user state database 14.
- the second detection unit 32 estimates the stress value of the second speaker based on the second state information.
- step S5 the voice conversion unit 41 of the adjustment unit 40 selects an indicator of the voice to be adjusted.
- the voice conversion unit 41 selects at least one indicator included in the first voice according to a preset priority.
- step S6 the voice conversion unit 41 of the adjustment unit 40 adjusts the voice of the first speaker (first voice).
- the voice conversion unit 41 adjusts the first voice based on the first state information and the second state information, or the first state information and the stress value.
- step S7 the audio output unit 42 of the adjustment unit 40 outputs the adjusted audio of the first speaker (first adjusted audio).
- the audio output unit 42 outputs the first adjusted audio to the second terminal 12.
- the second speaker listens to the first adjusted audio via the second terminal 12.
- the first adjusted audio has values for the indicators selected in step S5 that have been changed from the first audio.
- the content of the first adjusted audio is the same as the content of the first audio.
- step S8 the voice acquisition unit 21 of the acquisition unit 20 determines whether the voice of the second speaker (second voice) has been acquired from the second terminal 12. For example, if the voice acquisition unit 21 has not acquired the second voice from the second terminal 12 for a predetermined period of time, such as when the call ends, the voice acquisition unit 21 determines that the second voice has not been acquired from the second terminal 12 (step S8: NO), and terminates the flowchart shown in Figure 4, which is the processing method by the processing system 1 and device 10.
- step S8 determines that the second voice has been acquired from the second terminal 12 (step S8: YES) and proceeds to step S9.
- step S9 whether to adjust the voice of the second speaker (second voice). For example, if the emotion of the second speaker detected in step S4 is not a negative emotion, or if the reaction of the second speaker is not a negative reaction, the adjustment unit 40 determines that there is no need to adjust the second voice (step S9: NO), and the voice output unit 42 of the adjustment unit 40 outputs the second voice to the first terminal 11 as is without adjusting it.
- the flowchart shown in Figure 4 which is a processing method by the processing system 1 and device 10, ends.
- step S4 determines that the emotion of the second speaker detected in step S4 if the emotion of the second speaker detected in step S4 is detected as a negative emotion, or if the reaction of the second speaker is detected as a negative reaction, the adjustment unit 40 determines that the second voice needs to be adjusted (step S9: YES) and proceeds to step S10.
- step S9 the voice conversion unit 41 of the adjustment unit 40 adjusts the voice of the second speaker (second voice) in step S10.
- the voice conversion unit 41 adjusts the second voice based on the first state information and the second state information, or the first state information and the stress value.
- step S11 the audio output unit 42 of the adjustment unit 40 outputs the adjusted audio of the second speaker (second adjusted audio).
- the audio output unit 42 outputs the second adjusted audio to the first terminal 11.
- the first speaker listens to the second adjusted audio via the first terminal 11.
- the second adjusted audio has values for the indices selected in step S5 changed from the second audio.
- the content of the second adjusted audio is the same as the content of the second audio.
- the audio conversion unit 41 of the adjustment unit 40 may select audio indices to adjust for the second audio.
- the voice uttered by one speaker and the impression it gives can change depending on the state of that speaker, and the way the other speaker perceives the voice uttered by that speaker can also change depending on the state of the other speaker listening to that voice.
- the other speaker may wish for smooth communication to be achieved at least for the other speaker, regardless of the state of the first speaker. For this reason, there is a demand for technology that supports smooth communication by applying appropriate processing according to the state of each speaker.
- the device 10 of the present disclosure comprises an acquisition unit 20 that acquires the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal (first terminal 11 and second terminal 12); a first detection unit 31 that detects information related to the state of the first speaker based on the voice of the first speaker (first voice); a second detection unit 32 that detects information related to the state of the second speaker based on the biometric information of the second speaker; and an adjustment unit 40 that adjusts the voice based on information related to the state of the first speaker (first state information) and information related to the state of the second speaker (second state information).
- the disclosed method also includes steps of acquiring the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal (steps S1 and S2), detecting status information from the voice of the first speaker (first voice) (step S3), detecting status information from the biometric information of the second speaker (step S4), and adjusting the voice based on the status information of the first speaker (first status information) and the status information of the second speaker (second status information) (steps S5 to S7, S10, and S11).
- the device 10 and method disclosed herein adjust at least one of the first voice and the second voice based on the detected first status information and second status information.
- the device 10 and method can support and realize smooth conversation (communication) between speakers via audio terminals.
- the voice adjusted by the adjustment unit 40 is the voice of at least one of the first and second speakers.
- adjusting the first voice can reduce the psychological burden on the second speaker based on first state information, such as negative emotions from the first speaker.
- adjusting the second voice can adjust the second voice based on second state information including at least one of the emotions and reactions of the second speaker, who has experienced psychological burden based on first state information, such as negative emotions from the first speaker.
- the first speaker can receive the second adjusted voice, or the second speaker can receive the first adjusted voice, thereby realizing smooth conversation between speakers via the voice terminal.
- the second speaker's biometric information includes at least one of the following information: the second speaker's voice, the second speaker's biometric signal, and a video image of the second speaker.
- second status information can be detected from the second speaker's biometric information using various media and various indicators, and at least one of the first voice and the second voice can be appropriately adjusted.
- the acquisition unit 20 acquires information regarding the state of the second speaker after listening to the voice of the first speaker.
- the second detection unit 32 can detect changes in the emotions and reactions of the second speaker who has listened to the first voice. Therefore, the adjustment unit 40 can appropriately adjust at least one of the first voice and the second voice based on the changes in the emotions and reactions of the second speaker.
- the second detection unit 32 estimates the stress level of the second speaker based on information about the second speaker's state, and the adjustment unit 40 adjusts the voice based on the information about the first speaker's state and the stress level.
- the adjustment unit 40 can reflect the first state information and the stress level of the second speaker in its adjustment, thereby making it possible to appropriately adjust at least one of the first voice and the second voice.
- the adjustment unit 40 adjusts at least one indicator of the volume, pitch, pitch, and clarity indicated by the audio.
- the adjustment unit 40 can adjust at least one of the first audio and the second audio with respect to at least one of the above-mentioned indicators.
- the adjustment unit 40 is not limited to adjusting at least one indicator of the volume, pitch, pitch, and clarity, and may also adjust at least one indicator indicating voice quality.
- the adjustment unit 40 adjusts at least one indicator included in the audio according to a preset priority.
- the adjustment unit 40 can select an appropriate indicator according to the priority and adjust at least one of the first audio and the second audio.
- information regarding the state of the first speaker includes information regarding the emotions of the first speaker.
- the emotions of the first speaker can be estimated from the first voice, and at least one of the first voice and the second voice can be adjusted with respect to at least one of the above-mentioned indicators in accordance with the emotions associated with the first voice.
- the device and method disclosed herein have the following configuration:
- an acquisition unit that acquires the voice of a first speaker and biometric information of a second speaker who is conversing with the first speaker via a voice terminal; a first detection unit that detects information about the state based on the voice of the first speaker; a second detection unit that detects information about the condition of the second speaker based on the biometric information of the second speaker; an adjustment unit that adjusts the voice based on information about the state of the first speaker and information about the state of the second speaker;
- biometric information of the second speaker includes at least one of information of the voice of the second speaker, the biometric signal of the second speaker, and a video image of the second speaker.
- [5] The device according to any one of [1] to [4], wherein the adjustment unit estimates a stress value of the second speaker based on information about the second speaker's state, and adjusts the voice based on the information about the first speaker's state and the stress value.
- a method comprising:
- each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (for example, using wires, wirelessly, etc.) and these multiple devices.
- a functional block may also be realized by combining software with the single device or multiple devices.
- Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, election, establishment, comparison, assumption, expectation, regard, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment.
- a functional block (component) that performs transmission functions is called a transmitting unit or transmitter.
- transmitting unit or transmitter As mentioned above, there are no particular limitations on how these functions are implemented.
- device 10 constituting the conversion system according to an embodiment of the present disclosure may function as a computer that performs processing of the control method of the present disclosure.
- Figure 5 is a diagram illustrating an example of the hardware configuration of device 10 according to an embodiment of the present disclosure.
- the above-described device 10 may be physically configured as a computer including a processor 1001, memory 1002, storage 1003, communication device 1004, input device 1005, output device 1006, bus 1007, and the like.
- device 10 may be configured as a computer including at least one processor such as a CPU or GPU, and may also be configured as a computer including multiple processors, or may be configured to include multiple computer devices.
- the first terminal 11, second terminal 12, detection device 13, and the like may also have a similar hardware configuration.
- apparatus can be interpreted as a circuit, device, unit, etc.
- the hardware configuration of apparatus 10 may be configured to include one or more of the apparatuses shown in the figure, or may be configured to exclude some of the apparatuses.
- Each function of device 10 is realized by loading specific software (programs) onto hardware such as processor 1001 and memory 1002, causing processor 1001 to perform calculations, control communications via communication device 1004, and control at least one of reading and writing data from memory 1002 and storage 1003.
- the processor 1001 for example, runs an operating system to control the entire computer.
- the processor 1001 may be configured as a central processing unit (CPU) that includes an interface with peripheral devices, a control unit, an arithmetic unit, registers, etc.
- CPU central processing unit
- the acquisition unit 20, detection unit 30, adjustment unit 40, etc. described above may be realized by the processor 1001.
- the processor 1001 also reads programs (program code), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these.
- the programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments.
- the acquisition unit 20, detection unit 30, and adjustment unit 40 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be used for other functional blocks. While the above-described various processes have been described as being executed by a single processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001.
- the processor 1001 may be implemented on one or more chips.
- the programs may also be transmitted from a network via telecommunications lines.
- Memory 1002 is a computer-readable recording medium and may be composed of, for example, at least one of ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. Memory 1002 may also be called a register, cache, main memory (primary storage device), etc. Memory 1002 can store executable programs (program code), software modules, etc. for implementing a control method according to one embodiment of the present disclosure.
- ROM Read Only Memory
- EPROM Erasable Programmable ROM
- EEPROM Electrical Erasable Programmable ROM
- RAM Random Access Memory
- Memory 1002 may also be called a register, cache, main memory (primary storage device), etc.
- Memory 1002 can store executable programs (program code), software modules, etc. for implementing a control method according to one embodiment of the present disclosure.
- Storage 1003 is a computer-readable recording medium, and may be composed of at least one of an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc.
- Storage 1003 may also be referred to as an auxiliary storage device.
- the above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium that includes at least one of memory 1002 and storage 1003.
- the communication device 1004 is hardware (transmission/reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, or communication module, for example.
- the communication device 1004 may be configured to include high-frequency switches, duplexers, filters, frequency synthesizers, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD).
- FDD frequency division duplex
- TDD time division duplex
- the above-mentioned acquisition unit 20, detection unit 30, and adjustment unit 40 may be realized by the communication device 1004.
- the input device 1005 is an input device (e.g., a keyboard, mouse, microphone, switch, button, sensor, etc.) that accepts input from the outside.
- the output device 1006 is an output device (e.g., a display, speaker, LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
- each device such as the processor 1001 and memory 1002, is connected by a bus 1007 for communicating information.
- the bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
- the device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a field-programmable gate array (FPGA), and some or all of the functional blocks may be realized by this hardware.
- the processor 1001 may be implemented using at least one of these pieces of hardware.
- the notification of information is not limited to the aspects/embodiments described in this disclosure and may be performed using other methods.
- the notification of information may be performed by physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), higher layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, broadcast information (MIB (Master Information Block), SIB (System Information Block))), other signals, or a combination of these.
- RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, etc.
- Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
- the determination may be made based on a value represented by a single bit (0 or 1), a Boolean value (true or false), or a numerical comparison (for example, comparison with a predetermined value).
- notification of specified information is not limited to being done explicitly, but may also be done implicitly (e.g., not notifying the specified information).
- Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
- software, instructions, information, etc. may be transmitted and received via a transmission medium.
- a transmission medium such as coaxial cable, fiber optic cable, twisted pair, or Digital Subscriber Line (DSL)
- wired technology such as coaxial cable, fiber optic cable, twisted pair, or Digital Subscriber Line (DSL)
- wireless technology such as infrared or microwave
- the information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies.
- data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
- a channel and a symbol may be a signal (signaling).
- a signal may be a message.
- a component carrier CC may be called a carrier frequency, a cell, a frequency carrier, etc.
- radio resources may be indicated by an index.
- the names used for the parameters described above are not intended to be limiting in any way. Furthermore, the mathematical formulas using these parameters may differ from those explicitly disclosed in this disclosure.
- the various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
- MS Mobile Station
- UE User Equipment
- a mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.
- determining may encompass a wide variety of actions.
- Determining and “determining” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching a table, database, or other data structure), and ascertaining something that is considered to be a “determination.”
- Determining and “determining” may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and so on.
- judgment and “decision” can include regarding actions such as resolving, selecting, choosing, establishing, and comparing as having been “judgment” or “decision.” In other words, “judgment” and “decision” can include regarding some action as having been “judgment” or “decision.” Furthermore, “judgment (decision)” can be interpreted as “assuming,” “expecting,” “considering,” etc.
- connection refers to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are “connected” or “coupled” to each other.
- the coupling or connection between elements may be physical, logical, or a combination thereof.
- “connected” may be read as "access.”
- two elements may be considered to be “connected” or “coupled” to each other using at least one of one or more wires, cables, and printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.
- the phrase “based on” does not mean “based only on,” unless expressly stated otherwise. In other words, the phrase “based on” means both “based only on” and “based at least on.”
- any reference to an element using a designation such as "first,” “second,” etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
- a and B are different may mean “A and B are different from each other.” Note that this term may also mean “A and B are each different from C.” Terms such as “separate” and “combined” may also be interpreted in the same way as “different.”
- 1... processing system 10... device, 11... first terminal, 12... second terminal, 13... detection device, 14... user status database, 20... acquisition unit, 21... voice acquisition unit, 22... status acquisition unit, 30... detection unit, 31... first detection unit, 32... second detection unit, 40... adjustment unit, 41... voice conversion unit, 42... voice output unit.
Landscapes
- Engineering & Computer Science (AREA)
- Quality & Reliability (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Telephone Function (AREA)
Abstract
本開示によれば、音声端末を介した話者同士の円滑な会話を実現できる。装置10は、第1話者の音声、及び、第1話者と音声端末を介して会話する第2話者の生体情報を取得する取得部20と、第1話者の音声(第1音声)に基づいて状態に関する情報を検出する第1検出部31と、第2話者の生体情報に基づいて状態に関する情報を検出する第2検出部32と、第1話者の状態に関する情報(第1状態情報)、及び、第2話者の状態に関する情報(第2状態情報)に基づいて、音声を調整する調整部40と、を備える。
Description
本開示の一側面は、装置および方法に関する。
特許文献1には、電話対応開始後における会話中の顧客の音声から感情変化が感じ取れた場合、当該顧客の感情変化に対応して、担当者音声をリアルタイムで変更し、音声変換チューニングを行う技術が開示されている。
音声端末を介した話者同士のコミュニケーションにおいては、一方の話者の状態によって当該一方の話者から発される音声及びその印象が変わり得るし、当該音声を聞く他方の話者の状態によって一方の話者から発された音声の捉え方も変わり得る。他方の話者にとっては、一方の話者の状態に左右されず、少なくとも他方の話者にとって円滑なコミュニケーションが実現されることを希望する場合がある。このため、話者同士の状態に応じて適切な処理を施すことで、円滑なコミュニケーションを支援する技術が求められている。
そこで、本開示は、音声端末を介した話者同士の円滑な会話を実現できる装置および方法を提供することを目的とする。
本開示の装置は、第1話者の音声、及び、第1話者と音声端末を介して会話する第2話者の生体情報を取得する取得部と、第1話者の音声に基づいて状態に関する情報を検出する第1検出部と、第2話者の生体情報に基づいて状態に関する情報を検出する第2検出部と、第1話者の状態に関する情報、及び、第2話者の状態に関する情報に基づいて、音声を調整する調整部と、を備える。
本開示の一側面によれば、音声端末を介した話者同士の円滑な会話を実現できる。
添付図面を参照しながら本開示の実施形態を説明する。可能な場合には、同一の部分には同一の符号を付して、重複する説明を省略する。
図1は、本開示の一実施の形態に係る装置を含む処理システムの構成を示すブロック図である。図1に示される処理システム1は、無線通信網及び固定通信網を含むネットワークを介して互いに通信可能に構成された装置10、第1端末11、第2端末12、検出装置13及びユーザ状態データベース14を含む。処理システム1は、音声端末の一例である第1端末11を利用する第1話者と、音声端末の一例である第2端末12を介して会話する第2話者との会話において利用される。第1話者は、例えば、ある商品又はあるサービスの顧客である。第2話者は、当該商品または当該サービスに関する問合せ先の組織に所属するオペレータである。
装置10は、第1端末11に向けて音声を発する第1話者の状態と、第2端末12に向けて音声を発する第2話者の状態とに基づいて、第1端末11に入力された音声及び第2端末12に入力された音声のうち少なくとも一方を調整する。本実施形態の装置10は、第1端末11からの第1話者の音声を第2端末12に中継し、第2端末12からの第2話者の音声を第1端末11に中継する。装置10は、第1端末11に入力された音声及び第2端末12に入力された音声を調整する。装置10は、検出装置13及びユーザ状態データベース14から得られた情報に基づいて、音声を調整する。以下、各構成を詳細に説明する。
第1端末11及び第2端末12は、第1話者と第2話者とが会話しようとする際に使用される音声端末である。第1端末11及び第2端末12は、例えば、パーソナルコンピュータ、スマートフォン、タブレット端末、フィーチャーフォン、サーバ装置、ゲーム機器等の装置である。なお、図1には、第1端末11及び第2端末12としてそれぞれ1台の端末のみが図示されているが、処理システム1は2台以上の任意の台数の第1端末11及び第2端末12を含んでいてもよい。
第1端末11は、第1話者の音声(以下、第1音声と記載)を取得する。第2端末12は、第2話者の音声(以下、第2音声と記載)を取得する。第1端末11及び第2端末12は、取得された各音声を装置10にそれぞれ出力する。第1端末11は、装置10により調整された第2話者の音声(以下、第2調整音声と記載)を取得する。第2端末12は、装置10により調整された第1話者の音声(以下、第1調整音声と記載)を取得する。なお、装置10が第2音声を調整しない場合、装置10は、第2端末12からの第2音声を中継せず、第2端末12は、ネットワークを介して直接第1端末11に第2音声を出力してもよい。
検出装置13は、第2話者の生体情報を検出する。検出装置13が検出する第2話者の生体情報は、第2話者の音声、第2話者の生体信号、及び、第2話者を撮像した動画像の少なくとも1つの情報を含む。検出装置13は、例えば、ウェアラブルデバイス、カメラ等を含む。検出装置13はこれらに限定されない。検出装置13のウェアラブルデバイスは、第2話者の身体に装着される。例えば、処理システム1を介して第1話者と第2話者とが会話しているとき、検出装置13のウェアラブルデバイスは、第2話者の心拍数、呼吸の速さ、活動量(移動量)、体温の少なくとも1つの指標に関する値を第2話者の生体情報として検出する。
検出装置13のカメラは、第2話者の身体の少なくとも一部を撮像する。検出装置13のカメラは、第2話者の音声を取得してもよい。例えば、処理システム1を介して第1話者と第2話者とが会話しているとき、検出装置13のカメラは、第2話者の身体の各部位の単位時間あたりの移動量及び移動方向等の動きに関する情報、及び、第2音声を第2話者の生体情報として検出する。各部位の移動とは、例えば、目の動き及び全身の動きを含む。検出装置13は、検出された第2話者の生体情報を出力する。検出装置13は、検出された第2話者の生体情報を記憶する。なお、装置10は、検出装置13の構成及び機能を有していてもよい。
ユーザ状態データベース14は、第1話者の状態に関する情報(以下、第1状態情報と記載)、及び、第2話者の状態に関する情報(以下、第2状態情報と記載)と比較可能な指標及び閾値を有するテーブルを記憶している。ユーザ状態データベース14が記憶している情報の詳細は後述する。
装置10は、機能的な構成要素として、取得部20、検出部30及び調整部40を含んで構成されている。装置10は、第1端末11から取得された第1音声を調整する。装置10は、第1調整音声を第2端末12に出力する。装置10は、第2端末12から取得された第2音声を調整する。装置10は、第2調整音声を第1端末11に出力する。以下、装置10の各機能部の機能について詳細に説明する。
取得部20は、第1音声、及び、第2話者の生体情報を取得する。取得部20は、音声取得部21と、状態取得部22とを有する。音声取得部21は、第1話者によって第1端末11に入力された第1音声を、第1端末11から取得する。音声取得部21は、第1端末11から送信された第1音声を受け付けることで、第1音声を取得する。音声取得部21は、第2話者によって第2端末12に入力された第2音声を、第2端末12から取得する。音声取得部21は、第2端末12から送信された第2音声を受け付けることで、第2音声を取得する。
状態取得部22は、第2話者の生体情報を取得する。状態取得部22は、検出装置13から第2話者の生体情報を取得する。状態取得部22は、第1音声又は第1調整音声が第2端末に到達した後の第2話者の状態に関する情報(第2状態情報)を取得してもよい。状態取得部22は、音声取得部21によって第1音声又は第2音声が取得されたとき、検出装置13に第2話者の生体情報を取得するように信号を出力してもよい。検出装置13は、当該信号をトリガーとして、第2話者の生体情報を取得してもよい。
検出部30は、第1状態情報及び第2状態情報を検出する。第1状態情報及び第2状態情報は、それぞれ第1話者及び第2話者の感情に関する情報を含む。検出部30は、第1検出部31と、第2検出部32とを有する。第1検出部31は、音声取得部21によって取得された第1音声に基づいて第1状態情報を検出する。第1検出部31は、第1音声が示す声質を示す指標を解析する。第1検出部31は、声質を示す指標の一例として、第1音声が示す音量、音高、ピッチ及び明瞭さのうち少なくとも1つの指標を解析する。本実施形態の第1検出部31は、例えば、第1音声が示す音量、音高、ピッチ及び明瞭さを解析する。
第1検出部31は、ユーザ状態データベース14から状態に関する情報の一例である感情に関する情報を取得する。第1検出部31は、第1音声が示す各指標と、ユーザ状態データベース14から取得した感情に関する情報とに基づいて、第1音声が示す感情を第1状態情報として検出する。本実施形態の第1検出部31は、第1音声が示す音量、音高、ピッチ及び明瞭さの指標と、予め定められた感情に対応する各指標(ルール)とを比較して第1音声が示す感情を検出する。
図2は、状態に関する情報の一例である感情に関する情報の構成の一例を示す図である。図2に示される例では、ユーザ状態データベース14から取得した感情に関する情報は、感情と音声の特徴量(指標)との対応関係を示した感情テーブルであり得る。「喜び」、「怒り」、「哀しみ」等の各感情は、声質を示す指標(例えば、音声が示す「音量」、「音高」、「ピッチ」、及び「明瞭さ」の少なくとも1つの指標)と相関関係がある。図2に示される数値は、各指標の最大値を100、最小値を0とした場合の基準値(%)を示している。
例えば、第1話者が大きな低い声で、まくし立てるように話している場合、第1検出部31は、第1音声が示す「音量」、「音高」、「ピッチ」、及び「明瞭さ」として、それぞれ70、20、75、40と解析する。第1検出部31は、ユーザ状態データベース14から取得した感情テーブルにおいて、「音量」、「音高」、「ピッチ」、及び「明瞭さ」が70、20、75、40である場合に、どの感情に相当するかを検索する。第1検出部31は、ユーザ状態データベース14から取得した感情テーブルに基づいて、第1話者の感情は怒りの感情を含んでいると推定する。よって、第1検出部31は、第1状態情報の感情に関する情報として、第1音声から怒りの感情を検出する。このように、第1検出部31は、第1音声が示す各指標の値と感情テーブルとに基づいて、第1状態情報の感情に関する情報を検出する。
なお、第1検出部31は、第1音声の声質を示す指標と、予め定められた感情に対応する各指標(ルール)とを比較することで第1音声が示す感情を検出しなくてもよい。例えば、第1検出部31は、装置10内または装置10の外部の機械学習モデルに第1音声の示す各指標を入力して、第1状態情報の感情に関する情報を出力させてもよい。機械学習モデルは、例えば、生成AIモデル(Generative Artificial Intelligence)及び識別・判定型のAIモデル、並びに、これらを組み合わせたモデルの少なくとも1つを含んでいてもよい。以下、機械学習モデルとして、生成AIモデルを利用する例を説明する。当該生成AIモデルは、入力された第1音声の声質を示す指標に基づいて、第1音声が示していると推定される感情を、第1状態情報として出力する。第1検出部31は、当該生成AIモデルにより出力された第1状態情報を取得してもよい。
また、第1検出部31は、第1音声の声質を示す指標を解析しなくてもよい。この場合、例えば、第1検出部31は、生成AIモデルに第1音声を入力して第1音声の声質を示す指標を検出させてもよい。生成AIモデルは、第1検出部31から入力された第1音声に基づいて、第1音声が示していると推定される感情を、第1状態情報として直接出力してもよい。なお、声質を示す指標として、音量、音高、ピッチ及び明瞭さのいずれかの指標を含まなくてもよい。声質を示す指標は、例えば、音色、響き、温かみ、及び、鼻に掛かる度合いの少なくとも1つを含んでいてもよい。
第2検出部32は、音声取得部21によって取得された第2音声、及び、状態取得部22によって取得された第2話者の生体情報に基づいて第2状態情報を検出する。第2検出部32は、上述した第1検出部31と同様の処理によって、第2音声が示す感情を第2状態情報として検出してもよい。この場合、第2検出部32は、第2音声が示す声質を示す指標を解析する。第2検出部32は、声質を示す指標の一例として、第2音声が示す音量、音高、ピッチ及び明瞭さのうち少なくとも1つの指標を解析する。第2検出部32は、ユーザ状態データベース14から状態に関する情報の一例である感情に関する情報を取得する。第2検出部32は、第2音声が示す各指標と、ユーザ状態データベース14から取得した感情に関する情報とに基づいて、第2音声が示す感情を第2状態情報として検出する。本実施形態の第2検出部32は、第2音声が示す音量、音高、ピッチ及び明瞭さの指標と、予め定められた感情に対応する音声に関する各指標(ルール)とを比較して第2音声が示す感情を検出する。
第2検出部32は、ユーザ状態データベース14から状態に関する情報の一例である反応に関する情報を取得する。第2検出部32は、第2話者の生体情報が示す各指標と、ユーザ状態データベース14から取得した反応に関する情報とに基づいて、第2音声が示す反応を第2状態情報として検出する。本実施形態の第2検出部32は、第2音声が示す音量、音高、ピッチ及び明瞭さの指標と、予め定められた反応に対応する各指標(ルール)とを比較して第2音声が示す反応を検出する。
図3は、状態に関する情報の一例である反応に関する情報を示す図である。図3に示される例では、ユーザ状態データベース14から取得した反応に関する情報は、状態(反応)と生体情報の指標との対応関係を示した反応テーブルであり得る。「興奮」、「焦り」、「ショック」等の第2話者が刺激(ここでは第1音声)に対して取り得る各反応は、生体情報が示す「心拍数」、「目の動き」、「呼吸の速さ」、及び「全身の動き」の少なくとも1つの指標と相関関係がある。図3に示される「心拍数」、「目の動き」及び「呼吸の速さ」は、各指標の回数であり、「全身の動き」の特徴量は、最大値を100、最小値を0とした場合の基準値(%)を示している。なお、図3に示される「心拍数」、「目の動き」、「呼吸の速さ」、及び「全身の動き」は、生体情報の一例である。生体情報は、「心拍数」、「目の動き」、「呼吸の速さ」、及び「全身の動き」以外の情報を含んでいてもよい。
例えば、第2話者が取り乱しており、頭をかきむしって様々な方向を見ている状態である場合、第2検出部32は、第2話者の生体情報が示す「心拍数」、「目の動き」、「呼吸の速さ」、及び「全身の動き」として、それぞれ95、20、25、40と解析する。第2検出部32は、ユーザ状態データベース14から取得した反応テーブルにおいて、「心拍数」、「目の動き」、「呼吸の速さ」、及び「全身の動き」が95、20、25、40である場合に、どの感情に相当するかを検索する。第2検出部32は、ユーザ状態データベース14から取得した反応テーブルに基づいて、第2話者の反応は焦りを含んでいると推定する。よって、第2検出部32は、第2状態情報の反応に関する情報として、第2音声から焦りの反応を検出する。このように、第2検出部32は、第2話者の生体情報が示す各指標の値と反応テーブルとに基づいて、第2状態情報の反応に関する情報を検出する。
なお、第2検出部32は、第2音声の声質を示す指標と、予め定められた感情及び反応に対応する各指標(ルール)とを比較することで第2音声が示す感情及び反応の少なくとも一方を検出しなくてもよい。例えば、第2検出部32は、装置10内または装置10の外部の生成AIモデル(機械学習モデルの一例)に第2話者の生体情報の示す各指標の値を入力して、第2状態情報における感情に関する情報及び反応に関する情報の少なくとも一方を出力させてもよい。当該生成AIモデルは、入力された第2音声の声質を示す指標に基づいて、第2音声が示していると推定される感情及び反応の少なくとも一方を、第2状態情報として出力する。第2検出部32は、当該生成AIモデルにより出力された第2状態情報を取得してもよい。
また、第2検出部32は、第2音声の声質を示す指標を解析しなくてもよい。この場合、例えば、第2検出部32は、生成AIモデルに第2音声を入力して第2音声の声質を示す指標を検出させてもよい。生成AIモデルは、第2検出部32から入力された第2音声に基づいて、第2音声が示していると推定される感情及び反応の少なくとも一方を、第2状態情報として直接出力してもよい。
第2検出部32は、第2状態情報に基づいて、第2話者のストレス値を推定する。第2話者のストレス値とは、第2話者のストレスの大きさを数値化した指標である。第2検出部32は、例えば、音量、音高、ピッチ及び明瞭さの少なくとも1つを含む第2音声の各指標と、感情に関する情報及び反応に関する情報の少なくとも一方に基づいて、第2話者のストレス値を推定する。第2検出部32は、例えば、音量、音高、ピッチ及び明瞭さの少なくとも1つを含む第2音声の各指標と、第2検出部32によって検出された第2話者の感情に対応する感情テーブルの基準値との乖離度合いが大きいほどストレス値が高いと推定する。ここでの乖離度合いは、例えば、第2音声の各指標と、第2話者の感情に対応する感情テーブルの各基準値との差分を、当該各基準値で除した値の平均値である。ここでの乖離度合いは、例えば、第2音声の各指標と、第2話者の感情に対応する感情テーブルの各基準値との差分の合算値であってもよい。
第2検出部32は、例えば、「心拍数」、「目の動き」、「呼吸の速さ」、「全身の動き」等を含む第2話者の生体情報の各指標と、第2検出部32によって検出された第2話者の反応に対応する反応テーブルの基準値との乖離度合いが大きいほどストレス値が高いと推定する。ここでの乖離度合いは、例えば、第2音声の各指標と、第2話者の反応に対応する反応テーブルの各基準値との差分を、当該各基準値で除した値の平均値である。ここでの乖離度合いは、例えば、第2音声の各指標と、第2話者の反応に対応する反応テーブルの各基準値との差分の合算値であってもよい。
なお、第2検出部32は、第2音声が示す音量、音高、ピッチ及び明瞭さの指標と、予め定められた感情テーブル及び反応テーブルに示される基準値とを比較することで第2話者のストレス値を検出しなくてもよい。例えば、第2検出部32は、装置10内または装置10の外部の生成AIモデル(機械学習モデルの一例)に第2話者の生体情報の示す各指標及び第2状態情報(感情に関する情報及び反応に関する情報の少なくとも一方)を入力して、第2話者のストレス値を出力させてもよい。当該生成AIモデルは、入力された情報に基づいて、第2話者のストレス値を推定し、当該ストレス値出力する。第2検出部32は、当該生成AIモデルにより出力された当該ストレス値を取得してもよい。
調整部40は、第1状態情報、及び、第2状態情報に基づいて、音声を調整する。調整部40が調整する音声は、第1話者及び第2話者の少なくとも一方の音声である。本実施形態の調整部40は、第1音声及び第2音声の双方を調整する。
調整部40は、音声変換部41と、音声出力部42とを有する。音声変換部41は、第1状態情報及び第2話者のストレス値に基づいて第1音声を調整する。例えば、第1状態情報の第1話者の感情に関する情報が「怒り」及び「哀しみ」等のネガティブな感情を含む場合、音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第1音声の各指標と、第1検出部31によって検出された第1話者の感情に対応する感情テーブルの各基準値との乖離度合いが小さくなるように調整する。特に、第2話者のストレス値が予め定められた閾値より高ければ高いほど、又は、第2話者のストレス値が次第に高くなっている場合、音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第1音声の各指標と、第1検出部31によって検出された第1話者の感情が示す感情テーブルの音声の各指標との乖離度合いがより小さくなるように調整する。
なお、音声変換部41は、第1状態情報の第1話者の感情に関する情報がネガティブな感情を含まない場合でも第1音声を調整してもよい。音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第1音声の各指標を、感情テーブルの各感情の項目に該当しない程度の値まで調整してもよい。音声変換部41は、第1状態情報の第1話者の感情に関する情報に関わらず、第2話者のストレス値に応じて、第1音声の各指標を調整してもよい。
音声変換部41は、第1状態情報及び第2状態情報に基づいて第2音声を調整する。例えば、第1状態情報の第1話者の感情に関する情報が「怒り」及び「哀しみ」等のネガティブな感情を含む場合、音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第2音声の各指標を、第1検出部31によって検出された第1話者の感情をなだめる方向に向かうように調整する。このとき、音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第2音声の各指標を落ち着いた第2調整音声となるように調整してもよい。また、音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第2音声の各指標と、第2検出部32によって検出された第2話者の感情に対応する感情テーブルの各基準値との乖離度合いが小さくなるように調整してもよい。
また、例えば、第2状態情報の第2話者の感情に関する情報が「怒り」及び「哀しみ」等のネガティブな感情を含む場合、音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第2音声の各指標と、第2検出部32によって検出された第2話者の感情に対応する感情テーブルの各基準値との乖離度合いが小さくなるように調整する。
また、例えば、第2状態情報の第2話者の反応に関する情報が「焦り」等のネガティブな反応を含む場合、音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第2音声の各指標と、第2検出部32によって検出された第2話者の反応に対応する反応テーブルの各基準値との乖離度合いが小さくなるように調整する。
なお、音声変換部41は、第2状態情報の第2話者の感情に関する情報がネガティブな感情を含まない場合、及び、第2話者の反応に関する情報がネガティブな感情を含まない場合でも第2音声を調整してもよい。音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第2音声の各指標を、感情テーブルの各感情の項目に該当しない程度の値まで調整してもよい。音声変換部41は、音量、音高、ピッチ及び明瞭さを含む第2音声の各指標を、反応テーブルの各反応の項目に該当しない程度の値まで調整してもよい。音声変換部41は、第2話者のストレス値に応じて、第1音声の各指標を調整してもよい。
音声変換部41は、第1音声及び第2音声が示す声質を示す指標を調整する。音声変換部41は、例えば、第1音声及び第2音声が示す音量、音高、ピッチ及び明瞭さのうち少なくとも1つの指標を調整する。音声変換部41は、予め設定された優先度に応じて、第1音声及び第2音声に含まれる少なくとも1つの指標を調整する。優先度は、例えば、第2話者によって予め定められていてもよい。優先度は、例えば、各指標のうち、検出された感情が示す感情テーブルの音声の指標との乖離度合いが大きい順であってもよい。音声変換部41は、調整した当該少なくとも1つの指標の値を記憶する。
音声変換部41は、第1音声が示す声質の指標を調整することで、第1調整音声を生成する。なお、音声変換部41は、例えば、第1音声を調整して、ボイスチェンジャーのように第1話者とは異なる人物が話しているような第1調整音声を生成してもよい。音声出力部42は、音声変換部41によって調整された第1調整音声を第2端末12に向けて出力する。第2端末12は、第1調整音声を受信する。
音声変換部41は、第2音声が示す声質の指標を調整することで、第2調整音声を生成する。なお、音声変換部41は、例えば、第2音声を調整して、ボイスチェンジャーのように第2話者とは異なる人物が話しているような第2調整音声を生成してもよい。音声出力部42は、音声変換部41によって調整された第2調整音声を第1端末11に向けて出力する。第1端末11は、第2調整音声を受信する。
なお、音声変換部41は、例えば、装置10内または装置10の外部の生成AIモデル(機械学習モデルの一例)に第1音声及び第2音声の少なくとも一方を入力して、第1調整音声及び第2調整音声を出力させてもよい。当該生成AIモデルは、上述の声質を示す指標、感情テーブル、反応テーブルの少なくとも何れかを学習していてもよい。当該生成AIモデルは、例えば、入力された音声に基づいて、感情テーブルに示されるいずれかの感情の各指標との乖離度合いが小さくなるように音声を調整し、調整された音声を出力する。調整部40は、当該生成AIモデルにより出力された第1調整音声及び第2調整音声の少なくとも一方を取得してもよい。
上記のように構成された処理システム1及び装置10による処理の手順、すなわち、本実施形態にかかる処理方法の流れについて説明する。図4は、処理システムによる処理方法の一例の手順を示すフローチャートである。図4に示される処理方法は、例えば、第1話者が操作する第1端末11から、第2話者が操作可能な第2端末12に向けて通話開始の信号を受信したタイミングに、第2端末12から当該信号を受信した旨の情報を装置10が受信することで開始される。なお、第1話者と第2話者との会話の任意のタイミングで、第2話者の第2端末12への入力操作によって生成された信号を装置10が受信することで開始されてもよい。
処理方法においては、まず、取得部20の音声取得部21は、ステップS1として、第1話者の音声を取得する。音声取得部21は、第1話者によって第1端末11に入力された第1音声を、第1端末11から取得する。
続いて、取得部20の状態取得部22は、ステップS2として、第2話者の生体情報を取得する。状態取得部22は、音声取得部21によって第1音声が取得されたとき、検出装置13に第2話者の生体情報を取得するように信号を出力する。検出装置13は、当該信号をトリガーとして、第2話者の生体情報を取得する。状態取得部22は、第1話者が話しているときにおける第2話者の生体情報を取得する。
続いて、検出部30の第1検出部31は、ステップS3として、第1状態情報を検出する。第1検出部31は、例えば、ユーザ状態データベース14から状態に関する情報の一例である感情に関する情報を取得する。第1検出部31は、第1音声が示す各指標と、ユーザ状態データベース14から取得した感情に関する情報とに基づいて、第1音声が示す感情を第1状態情報として検出する。
続いて、検出部30の第2検出部32は、ステップS4として、第2状態情報を検出する。第2検出部32は、例えば、ユーザ状態データベース14から状態に関する情報の一例である感情に関する情報及び反応に関する情報の少なくとも一方を取得する。第2検出部32は、第2音声が示す各指標と、ユーザ状態データベース14から取得した感情に関する情報とに基づいて、第2音声が示す感情を第1状態情報として検出してもよい。第2検出部32は、第2話者の生体情報が示す各指標と、ユーザ状態データベース14から取得した反応に関する情報とに基づいて、第2話者の生体情報が示す反応を第2状態情報として検出してもよい。第2検出部32は、第2状態情報に基づいて、第2話者のストレス値を推定する。
続いて、調整部40の音声変換部41は、ステップS5として、調整する音声の指標を選択する。音声変換部41は、予め設定された優先度に応じて、第1音声に含まれる少なくとも1つの指標を選択する。
続いて、調整部40の音声変換部41は、ステップS6として、第1話者の音声(第1音声)を調整する。音声変換部41は、第1状態情報及び第2状態情報、又は、第1状態情報及びストレス値に基づいて第1音声を調整する。
続いて、調整部40の音声出力部42は、ステップS7として、調整された第1話者の音声(第1調整音声)を出力する。音声出力部42は、第2端末12に向けて第1調整音声を出力する。第2話者は、第2端末12を介して第1調整音声を聞く。第1調整音声は、第1音声からステップS5において選択された指標に関する値が変更されている。第1調整音声の内容は、第1音声の内容と同一である。
取得部20の音声取得部21は、ステップS8として、第2端末12から第2話者の音声(第2音声)を取得しているか否かを判定する。例えば、通話が終了する場合等、音声取得部21が第2端末12から第2音声を所定の期間取得していない場合、音声取得部21は、第2端末12から第2音声を取得していないと判定し(ステップS8:NO)、処理システム1及び装置10による処理方法である図4に示されるフローチャートを終了する。
第2話者が応答して、通話が継続する場合等、音声取得部21が第2端末12から第2音声を所定の期間中に取得した場合、音声取得部21は、第2端末12から第2音声を取得していると判定し(ステップS8:YES)、ステップS9に移行する。
音声取得部21によって第2音声を取得していると判定された場合(ステップS8:YES)、調整部40は、ステップS9として、第2話者の音声(第2音声)を調整するか否かを判定する。例えば、調整部40は、ステップS4において検出された第2話者の感情がネガティブな感情ではないと検出された場合、又は、第2話者の反応がネガティブな反応ではないと検出された場合、第2音声を調整する必要がないと判定し(ステップS9:NO)、調整部40の音声出力部42は、第2音声を調整せずに、第2音声のまま第1端末11へ出力する。第2音声の出力後、処理システム1及び装置10による処理方法である図4に示されるフローチャートを終了する。
例えば、調整部40は、ステップS4において検出された第2話者の感情がネガティブな感情であると検出された場合、又は、第2話者の反応がネガティブな反応であると検出された場合、第2音声を調整する必要があると判定し(ステップS9:YES)、ステップS10に移行する。
調整部40が第2音声を調整すると判定した場合(ステップS9:YES)、調整部40の音声変換部41は、ステップS10として、第2話者の音声(第2音声)を調整する。音声変換部41は、第1状態情報及び第2状態情報、又は、第1状態情報及びストレス値に基づいて第2音声を調整する。
続いて、調整部40の音声出力部42は、ステップS11として、調整された第2話者の音声(第2調整音声)を出力する。音声出力部42は、第1端末11に向けて第2調整音声を出力する。第1話者は、第1端末11を介して第2調整音声を聞く。第2調整音声は、第2音声からステップS5において選択された指標に関する値が変更されている。第2調整音声の内容は、第2音声の内容と同一である。なお、ステップS10の前において、調整部40の音声変換部41は、第2音声について調整する音声の指標を選択してもよい。ステップS10が完了した場合、処理システム1及び装置10による処理方法である図4に示されるフローチャートを終了する。
次に、従来の課題の一例を参照しつつ、本開示の装置及び方法の作用効果について説明する。例えば、コールセンター又は窓口など、音声端末を使用する現場では、オペレータは、様々な顧客からのネガティブな感情を受ける場合、又は、顧客によって持ち込まれたトラブルに対応しなければならない場合等がある。このような例では、顧客対応を行うオペレータへの心理的な負担が非常に大きいため、オペレータにおける心理的な負担を軽減する技術が求められている。
また、音声端末を介した話者同士のコミュニケーションにおいては、一方の話者の状態によって当該一方の話者から発される音声及びその印象が変わり得るし、当該音声を聞く他方の話者の状態によって一方の話者から発された音声の捉え方も変わり得る。他方の話者にとっては、一方の話者の状態に左右されず、少なくとも他方の話者にとって円滑なコミュニケーションが実現されることを希望する場合がある。このため、話者同士の状態に応じて適切な処理を施すことで、円滑なコミュニケーションを支援する技術が求められている。
本開示の装置10は、第1話者の音声、及び、第1話者と音声端末(第1端末11及び第2端末12)を介して会話する第2話者の生体情報を取得する取得部20と、第1話者の音声(第1音声)に基づいて状態に関する情報を検出する第1検出部31と、第2話者の生体情報に基づいて状態に関する情報を検出する第2検出部32と、第1話者の状態に関する情報(第1状態情報)、及び、第2話者の状態に関する情報(第2状態情報)に基づいて、音声を調整する調整部40と、を備える。
また、本開示の方法においては、第1話者の音声、及び、第1話者と音声端末を介して会話する第2話者の生体情報を取得するステップ(ステップS1,S2)と、第1話者の音声(第1音声)から状態に関する情報を検出するステップ(ステップS3)と、第2話者の生体情報から状態に関する情報を検出するステップ(ステップS4)と、第1話者の状態に関する情報(第1状態情報)、及び、第2話者の状態に関する情報(第2状態情報)に基づいて、音声を調整するステップ(ステップS5~S7,S10,S11)と、を含む。
本開示の装置10および方法によって、第1音声及び第2音声の少なくとも一方は、検出された第1状態情報及び第2状態情報に基づいて、調整される。話者同士の状態に応じて音声に適切な処理が施されることで、装置10および方法は、音声端末を介した話者同士の円滑な会話(コミュニケーション)を支援及び実現できる。
また、本開示の装置10においては、調整部40が調整する音声は、第1話者及び第2話者の少なくとも一方の音声である。この場合、例えば、第1音声が調整されることで、第1話者からのネガティブな感情等の第1状態情報に基づく第2話者への心理的負担を軽くすることができる。例えば、第2音声が調整されることで、第1話者からのネガティブな感情等の第1状態情報に基づく心理的負担を受けた第2話者において、当該第2話者の感情及び反応の少なくとも一方を含む第2状態情報に基づいて、第2音声が調整される。このため、第1話者は第2調整音声を受信し、又は、第2話者は第1調整音声を受信することができるため、音声端末を介した話者同士の円滑な会話を実現できる。
また、本開示の装置10においては、第2話者の生体情報は、第2話者の音声、第2話者の生体信号、及び、第2話者を撮像した動画像の少なくとも1つの情報を含む。この場合、第2話者の生体情報について様々な媒体を通じて、様々な指標によって第2状態情報を検出することができ、第1音声及び第2音声の少なくとも一方を適切に調整される。
また、本開示の装置10において、取得部20は、第1話者の音声を聞いた後の第2話者の状態に関する情報を取得する。この場合、第2検出部32は、第1音声を聞いた第2話者の感情の変化及び反応の変化を検出することができる。よって、調整部40は、第2話者の感情の変化及び反応の変化に基づいて、第1音声及び第2音声の少なくとも一方を適切に調整できる。
また、本開示の装置10において、第2検出部32は、第2話者の状態に関する情報に基づいて、第2話者のストレス値を推定し、調整部40は、第1話者の状態に関する情報及び当該ストレス値に基づいて音声を調整する。この場合、調整部40は、第1状態情報及び第2話者のストレス値を調整に反映することができるため、第1音声及び第2音声の少なくとも一方を適切に調整できる。
また、本開示の装置10において、調整部40は、音声が示す音量、音高、ピッチ及び明瞭さのうち少なくとも1つの指標を調整する。この場合、調整部40は、第1音声及び第2音声の少なくとも一方を上述の少なくとも1つの指標について調整することができる。調整部40は、音量、音高、ピッチ及び明瞭さの少なくとも1つの指標に限定されず、声質を示す少なくとも1つの指標を調整してもよい。
また、本開示の装置10において、調整部40は、予め設定された優先度に応じて、音声に含まれる少なくとも1つの指標を調整する。この場合、調整部40は、優先度に応じて適切な指標を選択して第1音声及び第2音声の少なくとも一方を調整することができる。
また、本開示の装置10において、第1話者の状態に関する情報とは、第1話者の感情に関する情報を含む。この場合、第1音声から第1話者の感情を推定して、第1音声に伴う感情に応じて、第1音声及び第2音声の少なくとも一方を上述の少なくとも1つの指標について調整することができる。
本開示の装置および方法は、以下の構成を有する。
[1]
第1話者の音声、及び、前記第1話者と音声端末を介して会話する第2話者の生体情報を取得する取得部と、
前記第1話者の音声に基づいて状態に関する情報を検出する第1検出部と、
前記第2話者の生体情報に基づいて状態に関する情報を検出する第2検出部と、
前記第1話者の状態に関する情報、及び、前記第2話者の状態に関する情報に基づいて、音声を調整する調整部と、
を備える、装置。
第1話者の音声、及び、前記第1話者と音声端末を介して会話する第2話者の生体情報を取得する取得部と、
前記第1話者の音声に基づいて状態に関する情報を検出する第1検出部と、
前記第2話者の生体情報に基づいて状態に関する情報を検出する第2検出部と、
前記第1話者の状態に関する情報、及び、前記第2話者の状態に関する情報に基づいて、音声を調整する調整部と、
を備える、装置。
[2]
前記調整部が調整する音声は、前記第1話者及び前記第2話者の少なくとも一方の音声である、上記[1]に記載の装置。
前記調整部が調整する音声は、前記第1話者及び前記第2話者の少なくとも一方の音声である、上記[1]に記載の装置。
[3]
前記第2話者の生体情報は、前記第2話者の音声、前記第2話者の生体信号、及び、前記第2話者を撮像した動画像の少なくとも1つの情報を含む、上記[1]又は[2]に記載の装置。
前記第2話者の生体情報は、前記第2話者の音声、前記第2話者の生体信号、及び、前記第2話者を撮像した動画像の少なくとも1つの情報を含む、上記[1]又は[2]に記載の装置。
[4]
前記取得部は、前記第1話者の音声を聞いた後の前記第2話者の状態に関する情報を取得する、上記[1]~[3]の何れかに記載の装置。
前記取得部は、前記第1話者の音声を聞いた後の前記第2話者の状態に関する情報を取得する、上記[1]~[3]の何れかに記載の装置。
[5]
前記調整部は、前記第2話者の状態に関する情報に基づいて、前記第2話者のストレス値を推定し、前記第1話者の状態に関する情報及び当該ストレス値に基づいて音声を調整する、上記[1]~[4]の何れかに記載の装置。
前記調整部は、前記第2話者の状態に関する情報に基づいて、前記第2話者のストレス値を推定し、前記第1話者の状態に関する情報及び当該ストレス値に基づいて音声を調整する、上記[1]~[4]の何れかに記載の装置。
[6]
前記調整部は、音声が示す音量、音高、ピッチ及び明瞭さのうち少なくとも1つの指標を調整する、上記[1]~[5]の何れかに記載の装置。
前記調整部は、音声が示す音量、音高、ピッチ及び明瞭さのうち少なくとも1つの指標を調整する、上記[1]~[5]の何れかに記載の装置。
[7]
前記調整部は、予め設定された優先度に応じて、音声に含まれる前記少なくとも1つの指標を調整する、上記[6]に記載の装置。
前記調整部は、予め設定された優先度に応じて、音声に含まれる前記少なくとも1つの指標を調整する、上記[6]に記載の装置。
[8]
前記第1話者の状態に関する情報とは、前記第1話者の感情に関する情報を含む、上記[1]~[7]の何れかに記載の装置。
前記第1話者の状態に関する情報とは、前記第1話者の感情に関する情報を含む、上記[1]~[7]の何れかに記載の装置。
[9]
第1話者の音声、及び、前記第1話者と音声端末を介して会話する第2話者の生体情報を取得するステップと、
前記第1話者の音声から状態に関する情報を検出するステップと、
前記第2話者の生体情報から状態に関する情報を検出するステップと、
前記第1話者の状態に関する情報、及び、前記第2話者の状態に関する情報に基づいて、音声を調整するステップと、
を含む、方法。
第1話者の音声、及び、前記第1話者と音声端末を介して会話する第2話者の生体情報を取得するステップと、
前記第1話者の音声から状態に関する情報を検出するステップと、
前記第2話者の生体情報から状態に関する情報を検出するステップと、
前記第1話者の状態に関する情報、及び、前記第2話者の状態に関する情報に基づいて、音声を調整するステップと、
を含む、方法。
上記実施形態の説明に用いたブロック図は、機能単位のブロックを示している。これらの機能ブロック(構成部)は、ハードウェアおよびソフトウェアの少なくとも一方の任意の組み合わせによって実現される。また、各機能ブロックの実現方法は特に限定されない。すなわち、各機能ブロックは、物理的または論理的に結合した1つの装置を用いて実現されてもよいし、物理的または論理的に分離した2つ以上の装置を直接的または間接的に(例えば、有線、無線などを用いて)接続し、これら複数の装置を用いて実現されてもよい。機能ブロックは、上記1つの装置または上記複数の装置にソフトウェアを組み合わせて実現されてもよい。
機能には、判断、決定、判定、計算、算出、処理、導出、調査、探索、確認、受信、送信、出力、アクセス、解決、選択、選定、確立、比較、想定、期待、見做し、報知(broadcasting)、通知(notifying)、通信(communicating)、転送(forwarding)、構成(configuring)、再構成(reconfiguring)、割り当て(allocating、mapping)、割り振り(assigning)などがあるが、これらに限られない。たとえば、送信を機能させる機能ブロック(構成部)は、送信部(transmitting unit)や送信機(transmitter)と呼称される。いずれも、上述したとおり、実現方法は特に限定されない。
例えば、本開示の一実施の形態における変換システムを構成する装置10などは、本開示の制御方法の処理を行うコンピュータとして機能してもよい。図5は、本開示の一実施の形態に係る装置10のハードウェア構成の一例を示す図である。上述の装置10は、物理的には、プロセッサ1001、メモリ1002、ストレージ1003、通信装置1004、入力装置1005、出力装置1006、バス1007などを含むコンピュータ装置として構成されてもよい。なお、装置10は、CPU、GPU等の少なくとも1つのプロセッサを含むコンピュータ装置として構成されていればよく、複数のプロセッサを含むコンピュータ装置として構成されていてもよく、複数のコンピュータ装置を含んで構成されていてもよい。第1端末11、第2端末12、検出装置13等も、同様なハードウェア構成を採りうる。
なお、以下の説明では、「装置」という文言は、回路、デバイス、ユニットなどに読み替えることができる。装置10のハードウェア構成は、図に示した各装置を1つまたは複数含むように構成されてもよいし、一部の装置を含まずに構成されてもよい。
装置10における各機能は、プロセッサ1001、メモリ1002などのハードウェア上に所定のソフトウェア(プログラム)を読み込ませることによって、プロセッサ1001が演算を行い、通信装置1004による通信を制御したり、メモリ1002およびストレージ1003におけるデータの読み出しおよび書き込みの少なくとも一方を制御したりすることによって実現される。
プロセッサ1001は、例えば、オペレーティングシステムを動作させてコンピュータ全体を制御する。プロセッサ1001は、周辺装置とのインターフェース、制御装置、演算装置、レジスタなどを含む中央処理装置(CPU:Central Processing Unit)によって構成されてもよい。例えば、上述の取得部20、検出部30、調整部40等は、プロセッサ1001によって実現されてもよい。
また、プロセッサ1001は、プログラム(プログラムコード)、ソフトウェアモジュール、データなどを、ストレージ1003および通信装置1004の少なくとも一方からメモリ1002に読み出し、これらに従って各種の処理を実行する。プログラムとしては、上述の実施の形態において説明した動作の少なくとも一部をコンピュータに実行させるプログラムが用いられる。例えば、取得部20、検出部30及び調整部40は、メモリ1002に格納され、プロセッサ1001において動作する制御プログラムによって実現されてもよく、他の機能ブロックについても同様に実現されてもよい。上述の各種処理は、1つのプロセッサ1001によって実行される旨を説明してきたが、2以上のプロセッサ1001により同時または逐次に実行されてもよい。プロセッサ1001は、1以上のチップによって実装されてもよい。なお、プログラムは、電気通信回線を介してネットワークから送信されても良い。
メモリ1002は、コンピュータ読み取り可能な記録媒体であり、例えば、ROM(Read Only Memory)、EPROM(Erasable Programmable ROM)、EEPROM(Electrically Erasable Programmable ROM)、RAM(Random Access Memory)などの少なくとも1つによって構成されてもよい。メモリ1002は、レジスタ、キャッシュ、メインメモリ(主記憶装置)などと呼ばれてもよい。メモリ1002は、本開示の一実施の形態に係る制御方法を実施するために実行可能なプログラム(プログラムコード)、ソフトウェアモジュールなどを保存することができる。
ストレージ1003は、コンピュータ読み取り可能な記録媒体であり、例えば、CD-ROM(Compact Disc ROM)などの光ディスク、ハードディスクドライブ、フレキシブルディスク、光磁気ディスク(例えば、コンパクトディスク、デジタル多用途ディスク、Blu-ray(登録商標)ディスク)、スマートカード、フラッシュメモリ(例えば、カード、スティック、キードライブ)、フロッピー(登録商標)ディスク、磁気ストリップなどの少なくとも1つによって構成されてもよい。ストレージ1003は、補助記憶装置と呼ばれてもよい。上述の記憶媒体は、例えば、メモリ1002およびストレージ1003の少なくとも一方を含むデータベース、サーバその他の適切な媒体であってもよい。
通信装置1004は、有線ネットワークおよび無線ネットワークの少なくとも一方を介してコンピュータ間の通信を行うためのハードウェア(送受信デバイス)であり、例えばネットワークデバイス、ネットワークコントローラ、ネットワークカード、通信モジュールなどともいう。通信装置1004は、例えば周波数分割複信(FDD:Frequency Division Duplex)および時分割複信(TDD:Time Division Duplex)の少なくとも一方を実現するために、高周波スイッチ、デュプレクサ、フィルタ、周波数シンセサイザなどを含んで構成されてもよい。例えば、上述の取得部20、検出部30及び調整部40などは、通信装置1004によって実現されてもよい。
入力装置1005は、外部からの入力を受け付ける入力デバイス(例えば、キーボード、マウス、マイクロフォン、スイッチ、ボタン、センサなど)である。出力装置1006は、外部への出力を実施する出力デバイス(例えば、ディスプレイ、スピーカ、LEDランプなど)である。なお、入力装置1005および出力装置1006は、一体となった構成(例えば、タッチパネル)であってもよい。
また、プロセッサ1001、メモリ1002などの各装置は、情報を通信するためのバス1007によって接続される。バス1007は、単一のバスを用いて構成されてもよいし、装置間ごとに異なるバスを用いて構成されてもよい。
また、装置10は、マイクロプロセッサ、デジタル信号プロセッサ(DSP:Digital Signal Processor)、ASIC(Application Specific Integrated Circuit)、PLD(Programmable Logic Device)、FPGA(Field Programmable Gate Array)などのハードウェアを含んで構成されてもよく、当該ハードウェアにより、各機能ブロックの一部または全てが実現されてもよい。例えば、プロセッサ1001は、これらのハードウェアの少なくとも1つを用いて実装されてもよい。
情報の通知は、本開示において説明した態様/実施形態に限られず、他の方法を用いて行われてもよい。例えば、情報の通知は、物理レイヤシグナリング(例えば、DCI(Downlink Control Information)、UCI(Uplink Control Information))、上位レイヤシグナリング(例えば、RRC(Radio Resource Control)シグナリング、MAC(Medium Access Control)シグナリング、報知情報(MIB(Master Information Block)、SIB(System Information Block)))、その他の信号またはこれらの組み合わせによって実施されてもよい。また、RRCシグナリングは、RRCメッセージと呼ばれてもよく、例えば、RRC接続セットアップ(RRC Connection Setup)メッセージ、RRC接続再構成(RRC Connection Reconfiguration)メッセージなどであってもよい。
本開示において説明した各態様/実施形態の処理手順、シーケンス、フローチャートなどは、矛盾の無い限り、順序を入れ替えてもよい。例えば、本開示において説明した方法については、例示的な順序を用いて様々なステップの要素を提示しており、提示した特定の順序に限定されない。
入出力された情報等は特定の場所(例えば、メモリ)に保存されてもよいし、管理テーブルを用いて管理してもよい。入出力される情報等は、上書き、更新、または追記され得る。出力された情報等は削除されてもよい。入力された情報等は他の装置へ送信されてもよい。
判定は、1ビットで表される値(0か1か)によって行われてもよいし、真偽値(Boolean:trueまたはfalse)によって行われてもよいし、数値の比較(例えば、所定の値との比較)によって行われてもよい。
本開示において説明した各態様/実施形態は単独で用いてもよいし、組み合わせて用いてもよいし、実行に伴って切り替えて用いてもよい。また、所定の情報の通知(例えば、「Xであること」の通知)は、明示的に行うものに限られず、暗黙的(例えば、当該所定の情報の通知を行わない)ことによって行われてもよい。
以上、本開示について詳細に説明したが、当業者にとっては、本開示が本開示中に説明した実施形態に限定されるものではないということは明らかである。本開示は、請求の範囲の記載により定まる本開示の趣旨および範囲を逸脱することなく修正および変更態様として実施することができる。したがって、本開示の記載は、例示説明を目的とするものであり、本開示に対して何ら制限的な意味を有するものではない。
ソフトウェアは、ソフトウェア、ファームウェア、ミドルウェア、マイクロコード、ハードウェア記述言語と呼ばれるか、他の名称で呼ばれるかを問わず、命令、命令セット、コード、コードセグメント、プログラムコード、プログラム、サブプログラム、ソフトウェアモジュール、アプリケーション、ソフトウェアアプリケーション、ソフトウェアパッケージ、ルーチン、サブルーチン、オブジェクト、実行可能ファイル、実行スレッド、手順、機能などを意味するよう広く解釈されるべきである。
また、ソフトウェア、命令、情報などは、伝送媒体を介して送受信されてもよい。例えば、ソフトウェアが、有線技術(同軸ケーブル、光ファイバケーブル、ツイストペア、デジタル加入者回線(DSL:Digital Subscriber Line)など)および無線技術(赤外線、マイクロ波など)の少なくとも一方を使用してウェブサイト、サーバ、または他のリモートソースから送信される場合、これらの有線技術および無線技術の少なくとも一方は、伝送媒体の定義内に含まれる。
本開示において説明した情報、信号などは、様々な異なる技術のいずれかを使用して表されてもよい。例えば、上記の説明全体に渡って言及され得るデータ、命令、コマンド、情報、信号、ビット、シンボル、チップなどは、電圧、電流、電磁波、磁界若しくは磁性粒子、光場若しくは光子、またはこれらの任意の組み合わせによって表されてもよい。
なお、本開示において説明した用語および本開示の理解に必要な用語については、同一のまたは類似する意味を有する用語と置き換えてもよい。例えば、チャネルおよびシンボルの少なくとも一方は信号(シグナリング)であってもよい。また、信号はメッセージであってもよい。また、コンポーネントキャリア(CC:Component Carrier)は、キャリア周波数、セル、周波数キャリアなどと呼ばれてもよい。
また、本開示において説明した情報、パラメータなどは、絶対値を用いて表されてもよいし、所定の値からの相対値を用いて表されてもよいし、対応する別の情報を用いて表されてもよい。例えば、無線リソースはインデックスによって指示されるものであってもよい。
上述したパラメータに使用する名称はいかなる点においても限定的な名称ではない。さらに、これらのパラメータを使用する数式等は、本開示で明示的に開示したものと異なる場合もある。様々なチャネル(例えば、PUCCH、PDCCHなど)および情報要素は、あらゆる好適な名称によって識別できるので、これらの様々なチャネルおよび情報要素に割り当てている様々な名称は、いかなる点においても限定的な名称ではない。
本開示においては、「移動局(MS:Mobile Station)」、「ユーザ端末(user terminal)」、「ユーザ装置(UE:User Equipment)」、「端末」などの用語は、互換的に使用され得る。
移動局は、当業者によって、加入者局、モバイルユニット、加入者ユニット、ワイヤレスユニット、リモートユニット、モバイルデバイス、ワイヤレスデバイス、ワイヤレス通信デバイス、リモートデバイス、モバイル加入者局、アクセス端末、モバイル端末、ワイヤレス端末、リモート端末、ハンドセット、ユーザエージェント、モバイルクライアント、クライアント、またはいくつかの他の適切な用語で呼ばれる場合もある。
本開示で使用する「判断(determining)」、「決定(determining)」という用語は、多種多様な動作を包含する場合がある。「判断」、「決定」は、例えば、判定(judging)、計算(calculating)、算出(computing)、処理(processing)、導出(deriving)、調査(investigating)、探索(looking up、search、inquiry)(例えば、テーブル、データベースまたは別のデータ構造での探索)、確認(ascertaining)した事を「判断」「決定」したとみなす事などを含み得る。また、「判断」、「決定」は、受信(receiving)(例えば、情報を受信すること)、送信(transmitting)(例えば、情報を送信すること)、入力(input)、出力(output)、アクセス(accessing)(例えば、メモリ中のデータにアクセスすること)した事を「判断」「決定」したとみなす事などを含み得る。また、「判断」、「決定」は、解決(resolving)、選択(selecting)、選定(choosing)、確立(establishing)、比較(comparing)などした事を「判断」「決定」したとみなす事を含み得る。つまり、「判断」「決定」は、何らかの動作を「判断」「決定」したとみなす事を含み得る。また、「判断(決定)」は、「想定する(assuming)」、「期待する(expecting)」、「みなす(considering)」などで読み替えられてもよい。
「接続された(connected)」、「結合された(coupled)」という用語、またはこれらのあらゆる変形は、2またはそれ以上の要素間の直接的または間接的なあらゆる接続または結合を意味し、互いに「接続」または「結合」された2つの要素間に1またはそれ以上の中間要素が存在することを含むことができる。要素間の結合または接続は、物理的なものであっても、論理的なものであっても、或いはこれらの組み合わせであってもよい。例えば、「接続」は「アクセス」で読み替えられてもよい。本開示で使用する場合、2つの要素は、1またはそれ以上の電線、ケーブルおよびプリント電気接続の少なくとも一つを用いて、並びにいくつかの非限定的かつ非包括的な例として、無線周波数領域、マイクロ波領域および光(可視および不可視の両方)領域の波長を有する電磁エネルギーなどを用いて、互いに「接続」または「結合」されると考えることができる。
本開示において使用する「に基づいて」という記載は、別段に明記されていない限り、「のみに基づいて」を意味しない。言い換えれば、「に基づいて」という記載は、「のみに基づいて」と「に少なくとも基づいて」の両方を意味する。
本開示において使用する「第1の」、「第2の」などの呼称を使用した要素へのいかなる参照も、それらの要素の量または順序を全般的に限定しない。これらの呼称は、2つ以上の要素間を区別する便利な方法として本開示において使用され得る。したがって、第1および第2の要素への参照は、2つの要素のみが採用され得ること、または何らかの形で第1の要素が第2の要素に先行しなければならないことを意味しない。
本開示において、「含む(include)」、「含んでいる(including)」およびそれらの変形が使用されている場合、これらの用語は、用語「備える(comprising)」と同様に、包括的であることが意図される。さらに、本開示において使用されている用語「または(or)」は、排他的論理和ではないことが意図される。
本開示において、例えば、英語でのa, anおよびtheのように、翻訳により冠詞が追加された場合、本開示は、これらの冠詞の後に続く名詞が複数形であることを含んでもよい。
本開示において、「AとBが異なる」という用語は、「AとBが互いに異なる」ことを意味してもよい。なお、当該用語は、「AとBがそれぞれCと異なる」ことを意味してもよい。「離れる」、「結合される」などの用語も、「異なる」と同様に解釈されてもよい。
1…処理システム、10…装置、11…第1端末、12…第2端末、13…検出装置、14…ユーザ状態データベース、20…取得部、21…音声取得部、22…状態取得部、30…検出部、31…第1検出部、32…第2検出部、40…調整部、41…音声変換部、42…音声出力部。
Claims (9)
- 第1話者の音声、及び、前記第1話者と音声端末を介して会話する第2話者の生体情報を取得する取得部と、
前記第1話者の音声に基づいて状態に関する情報を検出する第1検出部と、
前記第2話者の生体情報に基づいて状態に関する情報を検出する第2検出部と、
前記第1話者の状態に関する情報、及び、前記第2話者の状態に関する情報に基づいて、音声を調整する調整部と、
を備える、装置。 - 前記調整部が調整する音声は、前記第1話者及び前記第2話者の少なくとも一方の音声である、請求項1に記載の装置。
- 前記第2話者の生体情報は、前記第2話者の音声、前記第2話者の生体信号、及び、前記第2話者を撮像した動画像の少なくとも1つの情報を含む、請求項1に記載の装置。
- 前記取得部は、前記第1話者の音声を聞いた後の前記第2話者の状態に関する情報を取得する、請求項1に記載の装置。
- 前記第2検出部は、前記第2話者の状態に関する情報に基づいて、前記第2話者のストレス値を推定し、
前記調整部は、前記第1話者の状態に関する情報及び前記ストレス値に基づいて音声を調整する、請求項1に記載の装置。 - 前記調整部は、音声が示す音量、音高、ピッチ及び明瞭さのうち少なくとも1つの指標を調整する、請求項1に記載の装置。
- 前記調整部は、予め設定された優先度に応じて、音声に含まれる前記少なくとも1つの指標を調整する、請求項6に記載の装置。
- 前記第1話者の状態に関する情報とは、前記第1話者の感情に関する情報を含む、請求項1に記載の装置。
- 第1話者の音声、及び、前記第1話者と音声端末を介して会話する第2話者の生体情報を取得するステップと、
前記第1話者の音声から状態に関する情報を検出するステップと、
前記第2話者の生体情報から状態に関する情報を検出するステップと、
前記第1話者の状態に関する情報、及び、前記第2話者の状態に関する情報に基づいて、音声を調整するステップと、
を含む、方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2024/014418 WO2025215744A1 (ja) | 2024-04-09 | 2024-04-09 | 装置および方法 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2024/014418 WO2025215744A1 (ja) | 2024-04-09 | 2024-04-09 | 装置および方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025215744A1 true WO2025215744A1 (ja) | 2025-10-16 |
Family
ID=97350449
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2024/014418 Pending WO2025215744A1 (ja) | 2024-04-09 | 2024-04-09 | 装置および方法 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025215744A1 (ja) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2004252085A (ja) * | 2003-02-19 | 2004-09-09 | Fujitsu Ltd | 音声変換システム及び音声変換プログラム |
| JP2014095753A (ja) * | 2012-11-07 | 2014-05-22 | Hitachi Systems Ltd | 音声自動認識・音声変換システム |
| JP2021107873A (ja) * | 2019-12-27 | 2021-07-29 | パナソニックIpマネジメント株式会社 | 音声特性変更システムおよび音声特性変更方法 |
| JP7164793B1 (ja) * | 2021-11-25 | 2022-11-02 | ソフトバンク株式会社 | 音声処理システム、音声処理装置及び音声処理方法 |
| JP2023105607A (ja) * | 2022-01-19 | 2023-07-31 | 株式会社RevComm | プログラム、情報処理装置及び情報処理方法 |
-
2024
- 2024-04-09 WO PCT/JP2024/014418 patent/WO2025215744A1/ja active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2004252085A (ja) * | 2003-02-19 | 2004-09-09 | Fujitsu Ltd | 音声変換システム及び音声変換プログラム |
| JP2014095753A (ja) * | 2012-11-07 | 2014-05-22 | Hitachi Systems Ltd | 音声自動認識・音声変換システム |
| JP2021107873A (ja) * | 2019-12-27 | 2021-07-29 | パナソニックIpマネジメント株式会社 | 音声特性変更システムおよび音声特性変更方法 |
| JP7164793B1 (ja) * | 2021-11-25 | 2022-11-02 | ソフトバンク株式会社 | 音声処理システム、音声処理装置及び音声処理方法 |
| JP2023105607A (ja) * | 2022-01-19 | 2023-07-31 | 株式会社RevComm | プログラム、情報処理装置及び情報処理方法 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10554800B2 (en) | Audio data routing between multiple wirelessly connected devices | |
| JP2024097525A (ja) | 同期制御装置 | |
| WO2025215744A1 (ja) | 装置および方法 | |
| JP7572890B2 (ja) | 通信制御装置 | |
| JP2024097523A (ja) | 同期制御装置 | |
| JP6934825B2 (ja) | 通信制御システム | |
| JP2020170112A (ja) | 音声認識システム、近距離無線通信デバイス及び情報端末 | |
| JP2025152012A (ja) | 情報処理装置及び情報処理方法 | |
| WO2025210904A1 (ja) | 変換装置および変換方法 | |
| WO2025243399A1 (ja) | 推定装置及び推定方法 | |
| JP7837804B2 (ja) | 遠隔会議制御装置 | |
| WO2025243405A1 (ja) | 装置および方法 | |
| JP2022017868A (ja) | 通信制御装置および通信システム | |
| JP7809817B2 (ja) | 区切り記号挿入装置及び音声認識システム | |
| JP2024154565A (ja) | コミュニケーション支援装置 | |
| JP2026007544A (ja) | 生成装置及び生成方法 | |
| WO2026009320A1 (ja) | 装置および方法 | |
| US20250310709A1 (en) | Techniques for audio compensation for a bluetooth headset | |
| WO2026069461A1 (ja) | 情報処理装置および情報処理方法 | |
| WO2025238753A1 (ja) | 装置、および方法 | |
| JP2025167110A (ja) | 生成システム及び生成方法 | |
| WO2026004005A1 (ja) | 情報処理装置及び情報処理方法 | |
| JP6847006B2 (ja) | 通信制御装置及び端末 | |
| JP2025077703A (ja) | 情報処理装置 | |
| WO2025257982A1 (ja) | プロンプト生成装置及び方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24935124 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2026513822 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2026513822 Country of ref document: JP |