WO2020100647A1 - 情報伝達装置および情報伝達方法 - Google Patents
情報伝達装置および情報伝達方法 Download PDFInfo
- Publication number
- WO2020100647A1 WO2020100647A1 PCT/JP2019/043198 JP2019043198W WO2020100647A1 WO 2020100647 A1 WO2020100647 A1 WO 2020100647A1 JP 2019043198 W JP2019043198 W JP 2019043198W WO 2020100647 A1 WO2020100647 A1 WO 2020100647A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- user
- voice
- unit
- image
- output
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M1/00—Substation equipment, e.g. for use by subscribers
- H04M1/72—Mobile telephones; Cordless telephones, i.e. devices for establishing wireless links to base stations without route selection
- H04M1/724—User interfaces specially adapted for cordless or mobile telephones
- H04M1/72448—User interfaces specially adapted for cordless or mobile telephones with means for adapting the functionality of the device according to specific conditions
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M1/00—Substation equipment, e.g. for use by subscribers
- H04M1/72—Mobile telephones; Cordless telephones, i.e. devices for establishing wireless links to base stations without route selection
- H04M1/724—User interfaces specially adapted for cordless or mobile telephones
- H04M1/72403—User interfaces specially adapted for cordless or mobile telephones with means for local support of applications that increase the functionality
- H04M1/72427—User interfaces specially adapted for cordless or mobile telephones with means for local support of applications that increase the functionality for supporting games or graphical animations
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M1/00—Substation equipment, e.g. for use by subscribers
- H04M1/57—Arrangements for indicating or recording the number of the calling subscriber at the called subscriber's set
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M2250/00—Details of telephonic subscriber devices
- H04M2250/74—Details of telephonic subscriber devices with voice recognition means
Definitions
- the present invention relates to an information transmission device and an information transmission method for transmitting information to a user by voice or display.
- Patent Document 1 Conventionally, there is known a device that transmits information to a user via an agent displayed on the screen (see Patent Document 1, for example).
- the behavior on the display of the agent is changed and the information is presented to the user based on the response information of the user obtained by the interaction between the character type agent displayed on the screen and the user. ..
- An information transmission device allows a voice output unit that outputs a voice, a display unit that displays an image, a voice input unit that inputs a voice uttered by a user, and a voice output by the voice output unit.
- a mode command unit for commanding the first mode or a second mode for prohibiting voice output, and a voice corresponding to the user's utterance input by the voice input unit
- a voice control unit that controls the voice output unit to output, an image control unit that controls the display image on the display unit so that the agent image including the face image of the agent is displayed, and the output of information to the user.
- an output determination unit that determines whether or not it is.
- the image control unit displays the first agent image having the first face image when the voice input unit inputs the voice uttered by the user while the mode instruction unit instructs the second mode.
- the output determination unit determines that information output is necessary without inputting the voice uttered by the user by the voice input unit, it is different from the first face image.
- the display image is controlled so that the second agent image having the second facial image of the facial expression is displayed.
- Another aspect of the present invention is an information transmission method for transmitting information via a display unit, wherein the computer commands a first mode in which voice output is permitted or a second mode in which voice output is prohibited.
- the voice output unit is controlled to output the voice corresponding to the user's utterance, and the display image displayed on the display unit is displayed so that the agent image including the face image of the agent is displayed. Controlling and determining whether to output the information to the user.
- the control of the display image is such that when the voice uttered by the user is input while the second mode is instructed, the first agent image having the first face image is displayed while the second mode is instructed.
- the second agent image having the second face image having a different facial expression from the first face image is displayed. Controlling the displayed image to be displayed.
- the figure which shows an example of the avatar image displayed on the user terminal of FIG. The block diagram which shows the control structure of the user terminal which comprises the information transmission apparatus which concerns on this embodiment.
- the figure which shows an example of the change of the facial expression of the avatar image displayed on the user terminal of FIG. 4 is a flowchart showing an example of processing executed by the controller of FIG.
- FIG. 1 is a diagram showing a schematic configuration of an information transmission system 100 having an information transmission device according to an embodiment of the present invention.
- the information transmission system 100 includes a user terminal 10 carried by a user 1 and a server device 30.
- the user terminal 10 and the server device 30 are connected to a network 3 including a public wireless communication network represented by the Internet network, a mobile phone network, etc., and can communicate with each other via the network 3.
- the network 3 also includes a closed communication network provided for each predetermined management area, such as a wireless LAN and Wi-Fi (registered trademark).
- the user terminal 10 is composed of various mobile terminals such as smartphones, tablet terminals, mobile phones, wearable terminals, etc., which are carried by the user 1 and have display units such as monitors and displays.
- FIG. 1 shows an example in which the user terminal 10 is configured by a smartphone that has a telephone function and is formed into a generally rectangular thin shape.
- the user terminal 10 is a user-specific terminal, and the user terminal 10 is provided with an identification ID (user ID) associated with the user 1 who uses the information transmission system 100.
- the user terminal 10 includes a display 11, a button group 12, a microphone 13, a speaker 14, a camera 15, a communication unit 16, a GPS receiver 17, an actuator 18, and a controller 20.
- the display 11 includes a display device such as a liquid crystal display, an organic EL display, an inorganic EL display.
- the entire display 11 is formed into a substantially rectangular shape, and various images such as characters, symbols and figures are displayed on the display 11.
- the display 11 is composed of, for example, a touch panel, and also functions as an input unit for inputting various kinds of information.
- the arrangement and shape of the display 11 are not limited to those shown in FIG.
- the button group 12 has a plurality of buttons 12a to 12c operated by the user 1, and various commands are input according to the operation of the buttons 12a to 12c.
- One of the buttons (for example, the button 12a) is configured as a manner mode button for instructing and canceling the manner mode, for example.
- the number, shape, and arrangement of the buttons that form the button group 12 are not limited to those described above.
- the microphone 13 constitutes a voice input unit that acquires the utterance of the user 1 as a voice signal.
- the signal acquired by the microphone 13 is input to the controller 20 as audio data via, for example, an A / D converter.
- the speaker 14 constitutes a sound output unit that converts a sound signal transmitted from the controller 20 into sound and outputs the sound.
- the speaker 14 outputs the voice of the other party during the call. Various voice messages, ringtones, music, etc. can be output from the speaker 14.
- the camera 15 is an in-camera that has an image pickup device such as a CCD or a CMOS sensor and takes an image of a person or an object facing the display 11.
- a back camera may be provided on the back surface of the display 11. The arrangement and the number of cameras 15 are not limited to those described above.
- the communication unit 16 wirelessly communicates via the network 3. It is also possible to communicate without going through the network 3.
- the GPS receiver 17 receives positioning signals from a plurality of GPS satellites and measures the absolute position (latitude, longitude, etc.) of the user terminal 10 based on the received positioning signals.
- the actuator 18 is composed of a vibrator or the like that gives vibration to the housing of the user terminal 10.
- the actuator 18 operates when it is necessary to notify the user 1 that a call or mail has arrived in the manner mode.
- the controller 20 is configured to include a microcomputer having an operation unit such as a CPU as an operation circuit, a storage unit such as ROM and RAM, and other peripheral circuits (interface circuit and the like). The functional configuration of the controller 20 will be described later.
- the server device 30 has, as a functional configuration, a communication unit 31, a user information acquisition unit 32, an image signal generation unit 33, and a storage unit 34.
- the communication unit 31 communicates with the user terminal 10 via the network 3.
- the user information acquisition unit 32 acquires user information from the user terminal 10.
- the acquired information is stored in the storage unit 34 in association with the user ID.
- the user information includes preference information such as the user's 1 hobbies and food preferences.
- the preference information is information unique to the user, and includes keywords included in the voice uttered by the user 1 and the voice output to the user 1, input information input from the touch panel (display 11), output information displayed on the display 11 and the like. Based on, the preference information corresponding to each user 1 is determined.
- the image signal generation unit 33 determines an avatar image (see FIG. 2) corresponding to each user 1 and generates an image signal corresponding to the avatar image.
- the personality of the user 1 is digitized based on the user information stored in the storage unit 34. For example, for each of a plurality of personality parameters such as the behavioral power, sociability, cooperation, and compassion of the user 1, their degree (degree of behavioral strength, degree of sociability, etc.) is numerically expressed.
- the avatar image displayed on the display 11 is determined in consideration of the numerical value distribution of the personality parameter, and the image signal corresponding to the avatar image is generated.
- the generated image signal is transmitted to the corresponding user terminal 10 via the communication unit 31.
- FIG. 2 is a diagram showing an example of the avatar image 40 displayed on the display 11.
- the avatar image 40 is an image of an avatar personified as an altercation of the user 1, and is represented in the form of a predetermined character or agent. Particularly, in the present embodiment, the avatar image 40 is represented as the face image of the character. More specifically, the avatar image 40 is displayed in a form that imitates the eyes, eyebrows, and mouth of the character.
- the avatar image 40 is a user-specific image that differs for each user, and is an image that represents the form of the avatar as the alter ego of the user 1, which reflects the character of the user 1 in a good manner. Therefore, the user 1 is likely to have an affinity for the avatar image 40.
- the avatar image 40 of the specific user 1 is changed according to the usage status of the user terminal 10. Note that the form of the avatar image 40 can be set by the user 1 via the user terminal 10 according to his / her preference.
- FIG. 3 is a block diagram showing a control configuration of the user terminal 10 as the information transmission device according to the present embodiment.
- signals from the microphone 13, the camera 15, the communication unit 16, the GPS receiver 17, and the manner mode button 12a are input to the controller 20.
- the CPU of the controller 20 executes a predetermined process based on these input signals and outputs a control signal to the speaker 14, the display 11 and the actuator 18.
- the controller 20 has a signal input unit 21, a surrounding environment estimation unit 22, an output determination unit 23, a voice control unit 24, an image control unit 25, an actuator control unit 26, and a storage unit 27. Have.
- the signal input unit 21 inputs signals from the microphone 13, the camera 15, the GPS receiver 17, and the manner mode button 12a, and receives signals from the server device 30 via the communication unit 16, for example, from the server device 30.
- An image signal corresponding to the transmitted avatar image 40 is input.
- the surrounding environment estimation unit 22 estimates the surrounding environment of the user terminal 10 based on the signal from the GPS receiver 17. This estimation is used to determine whether it is appropriate to output the sound from the speaker 14, that is, whether the user 1 is in the environment in which the user terminal 10 should be in the manner mode.
- the surrounding environment estimation unit 22 can also estimate the surrounding environment based on the signals acquired by the microphone 13 and the camera 15.
- the output determination unit 23 determines whether or not it is necessary to notify the user 1 of information. For example, when the user terminal 10 receives an incoming call or mail via the communication unit 16, or when a push notification is sent from the various servers to the user terminal 10 via the communication unit 16, the user 1 is notified of that fact. It is necessary to determine.
- the voice control unit 24 determines whether or not the manner mode is instructed by operating the manner mode button 12a.
- the manner mode is instructed (during the manner mode)
- the sound output from the speaker 14 is prohibited even when the output determination unit 23 determines that the user 1 needs to be notified of the information. Therefore, for example, no sound is output from the speaker 14 even when an incoming call is received.
- the manner mode is canceled (in the non-manner mode)
- the audio output from the speaker 14 is permitted. Therefore, for example, when there is an incoming call, a voice (ring tone) is output from the speaker 14.
- the voice control unit 24 controls the user terminal 10 based on the surrounding environment estimated by the surrounding environment estimation unit 22 and the user information stored in the storage unit 27 regardless of whether the manner mode button 12a is operated. Determine whether to switch to manner mode.
- the user terminal 10 is automatically set to the manner mode. It should be noted that it is determined whether or not the manner mode should be released according to the surrounding environment and the user information estimated by the surrounding environment estimation unit 22, and if it is determined that the manner mode should be released, the manner mode is automatically released. It can also be configured as follows.
- the voice control unit 24 can also output a voice including an answer or a proposal to the question of the user 1 and a further question to the answer of the user 1 from the speaker 14 in the non-manner mode. That is, the voice can be output so as to have a dialogue with the user 1.
- the voice control unit 24 recognizes the voice signal from the user 1 input via the microphone 13 by referring to the word registered in the dictionary database (not shown) of the storage unit 27. For example, it recognizes a question or an answer to the user terminal 10 uttered by the user 1.
- the response content corresponding to the content of the recognized voice is extracted from the dialogue database (not shown) of the storage unit 27, and the response message is generated.
- the speaker 14 is controlled to output the voice corresponding to the response message.
- the voice signal from the user 1 is transmitted to the server device 30, and the server device 30 generates a response message in consideration of the preference information of the user 1, and the user terminal 10 receives this and outputs it from the speaker 14.
- the voice control unit 24 determines whether or not there is an inquiry from the user 1 to the user terminal 10 based on the signals acquired by the microphone 13 and the camera 15 in the manner mode. More specifically, the voice control unit 24 refers to the dictionary database of the storage unit 27 to identify the word corresponding to the voice signal input from the user 1, and the identified word is registered in the dictionary database in advance. It is determined whether or not any of the words of the plurality of questions is present, and thus the presence or absence of the question is determined. Then, when it is determined that the user 1 makes an inquiry, the speaker 14 outputs the voice of the response to the inquiry even in the manner mode. In this case, the manner mode is temporarily released, and only the response to the inquiry is output from the speaker 14. Note that the manner mode may be released in the same manner as when the manner mode button 12a is operated, instead of temporarily releasing the manner mode.
- the image control unit 25 controls the display image on the display 11. For example, the control signal is output to the display 11 so that the avatar image 40 corresponding to the image signal transmitted from the server device 30 is displayed.
- the avatar image 40 is an image corresponding to the character of the user 1, but the facial expression of the avatar image 40 changes according to the usage status of the user terminal 10.
- FIG. 4 is a diagram showing an example of changes in facial expressions of the avatar image 40 displayed on the display 11.
- avatar images 40 (40A to 40D) having four facial expressions of type A to type D are shown.
- the type A avatar image 40A is, for example, an image when the user terminal 10 is in a standby state, that is, an image in a normal state.
- the avatar image 40A is represented by a standard facial expression in normal times.
- the type B avatar image 40 ⁇ / b> B is an image when the avatar issues a voice imitating a ring tone (for example, a voice such as “two-two”) to the user 1 in the non-manner mode (normal mode) in which the manner mode is canceled.
- the mouth of the avatar image 40B moves in response to the sound being emitted from the speaker 14.
- the avatar image 40B of type B is displayed not only in the non-manner mode, but also when the voice is output from the speaker 14 in response to the inquiry from the user 1 in the manner mode.
- the type C avatar image 40C is an image when notifying the user 1 that there is an incoming call or mail in the manner mode, for example.
- the avatar image 40C has an embarrassing facial expression because there is information to be transmitted to the user 1 but no audio can be output.
- the avatar images 40 displayed on the user terminals 10A, 10B are different from each other. ..
- the type C avatar image 40C is an avatar image displayed on the user terminal 10A of the first user 1A.
- the type D avatar image 40D is an image of the avatar displayed on the user terminal 10B of the second user 1B in the same situation as when the type C is displayed.
- the type D avatar image 40D is represented by an image with an angry expression.
- the type (the image expression) of the avatar images 40C and 40D is different between the type C and the type D because the personalities of the users 1A and 1B are different from each other. That is, the image signal generation unit 33 of the server device 30 generates the image signal in consideration of the personalities of the users 1A and 1B. Therefore, even if the user terminals 10A and 10B are in the same usage state (for example, in the manner mode) and the types of the avatar images 40 are the same (face images of the same character), the displayed avatar images 40 are not displayed.
- the facial expression is different for each of the user terminals 10A and 10B.
- the actuator control unit 26 determines that the output determination unit 23 needs to notify the user 1 of information in the manner mode, the actuator control unit 26 outputs a control signal to the actuator 18 and vibrates the user terminal 10. Thereby, the user 1 can easily notice that there is an incoming call or the like.
- the actuator 18 can always be turned off by changing the setting of the user terminal 10. Further, in the non-manner mode, the actuator 18 can be operated simultaneously with the output of the ring tone.
- FIG. 5 is a flowchart showing an example of processing executed by the CPU of the controller 20 according to a program stored in the storage unit 27 of FIG. 3 in advance.
- the process shown in this flowchart is started, for example, when the power of the user terminal 10 is turned on, and is repeated in a predetermined cycle.
- the server device 30 displays a plurality of avatar images 40 corresponding to the user ID (for example, an avatar image corresponding to the first user 1A).
- the images 40A to 40C) are determined and image signals corresponding to the plurality of avatar images 40 are transmitted to the user terminal 10.
- the user terminal 10 starts the process of FIG. 5 with the image signal stored in the storage unit 27.
- step S1 it is determined whether the manner mode is set manually by operating the manner mode button 12a or automatically according to the surrounding environment.
- the process proceeds to step S2, and it is determined whether or not there is an inquiry from the user 1 to the user terminal 10 based on the signals acquired by the microphone 13 and the camera 15. If the result in step S2 is affirmative, the process proceeds to step S5, and if the result is negative, the process proceeds to step S3.
- step S3 a control signal is output to the display 11 to display the type A avatar image 40A of FIG. 4, for example.
- step S1 determines whether or not there is an inquiry from the user 1 to the user terminal 10 based on the signals acquired by the microphone 13 and the camera 15. If the result in step S4 is affirmative, the process proceeds to step S5, in which a control signal is output to the speaker 14 and a voice of a response to the inquiry is output from the speaker 14.
- step S6 a control signal is output to the display 11 to display, for example, the type B avatar image 40B of FIG.
- step S4 determines whether or not to notify the user 1 based on the signal received via the communication unit 16. Whether or not, that is, whether or not to notify the user 1 is determined.
- step S7 the process advances to step S3 to display the type A avatar image 40A on the display 11, for example.
- the avatar images 40 other than types A to C may be displayed.
- step S7 If the result in step S7 is affirmative, the process proceeds to step S8, a control signal is output to the actuator 18, and the user terminal 10 is vibrated. This allows the user 1 to recognize that there is an incoming call.
- step S9 a control signal is output to the display 11 to display, for example, the type C avatar image 40C of FIG. Thereby, the user 1 can recognize that there is an incoming call or the like by displaying an image.
- the user 1 since the information is transmitted to the user 1 by the change in the facial expression of the avatar image 40, not by the characters or symbols, the user 1 can receive a comfortable service from the user terminal 10 and is highly satisfied with the use of the user terminal 10. You can get a feeling.
- the user terminal 10 as an information transmission device has a speaker 14 for outputting a voice, a display 11 for displaying an image, a microphone 13 for inputting a voice uttered by the user 1, a signal from a manner mode button 12a and the like. Based on the command, a normal mode (non-manner mode) that permits voice output from the speaker 14 or a manner mode that prohibits voice output is issued, and when the non-manner mode is instructed, the user 1 input by the microphone 13 is input.
- a normal mode non-manner mode
- a voice control unit 24 that controls the speaker 14 so as to output a voice corresponding to the utterance, and an image control unit that controls the display image on the display 11 so that the avatar image 40 including the face image of the agent (avatar) is displayed. 25 and an output determination unit 23 that determines whether or not it is necessary to output information to the user 1 (FIG. 3).
- the voice control unit 24 inputs a voice (for example, a voice inquiring to the user terminal 10) uttered by the user 1 when the voice control unit 24 instructs the manner mode, the image control unit 25 in FIG.
- the voice control unit 24 instructs the manner mode
- the voice judging unit 23 does not input the voice uttered by the user 1 and the output judging unit 23 outputs the information. If it is determined that the output is necessary, the display image is controlled so that the avatar image 40C having the facial expression of type C in FIG. 4 is displayed (FIG. 5).
- the mode of the avatar image 40 (when the user 1 makes an inquiry in the manner mode and when there is no inquiry from the user 1 but it is necessary to notify the user 1 that there is an incoming call, etc.) Expression) changes. For this reason, it is possible to appropriately notify the user 1 that there is an incoming call or the like in the manner mode, without outputting voice from the speaker 14. That is, the change in the facial expression is suitable for expressing the emotion of the avatar, and by notifying the user 1 of the change in the emotion of the avatar, the user 1 easily and accurately recognizes that there is an urgent notification, for example. It is possible to Also, the user 1 can interact with the avatar via the user terminal 10, and becomes familiar with the avatar. Then, since the user 1 is notified using the avatar image 40, the comfort of the user 1 is enhanced.
- the user terminal 10 as the information transmission device further includes the storage unit 27 that stores the user information including the preference information of the user 1 and the storage unit 34 of the server device 30 (FIGS. 1 and 3).
- the voice control unit 24 Even when the voice control unit 24 is in the manner mode, when the voice uttered by the user is input through the microphone 13, the voice control unit 24 recognizes the utterance of the user 1 input by the microphone 13 based on the stored user information.
- the speaker 14 is controlled to output the corresponding sound.
- the voice may be surprised, but in the present embodiment, when the user 1 makes a question, the voice is output in the manner mode. User 1 is not surprised by the voice output.
- the avatar images 40B and 40C are images unique to the user 1 based on the stored user information. Therefore, different avatar images (for example, 40C and 40D in FIG. 4) are displayed to the users 1A and 1B having different personalities, and the user 1 is more familiar with the avatars.
- the user terminal 10 as an information transmission device further includes a surrounding environment estimation unit 22 that estimates the surrounding environment (FIG. 3).
- the voice control unit 24 commands the permission or prohibition of voice output based on the stored user information and the surrounding environment estimated by the surrounding environment estimation unit 22.
- the manner mode is automatically set without operating the manner mode button 12a, so that the voice of the utterance is not output via the avatar image 40 in a situation where the voice output is not preferable.
- the user terminal 10 as the information transmission device uses the actuator 18 for generating vibration and the actuator 18 so as to generate vibration when the output determination unit 23 determines that information output is required in the manner mode.
- An actuator control unit 26 for controlling is further provided (FIG. 3). Thereby, the user 1 can easily recognize that there is an incoming call or the like in the manner mode.
- the information transmission method is an information transmission method of transmitting information via the display 11, and the computer (controller 20) prohibits the non-manner mode in which the voice output is permitted or the voice output.
- the speaker 14 is controlled so as to output the voice corresponding to the utterance of the user 1 (step S5), and the avatar including the face image of the agent is controlled. It includes controlling the display image of the display 11 so that the image 40 is displayed (step S3, step S6, step S9), and determining whether or not the output of information to the user 1 is necessary (step S7).
- Controlling the display image is performed when the voice of the user 1 is input and the type B avatar image 40B is displayed while the manner mode is permitted while the manner mode is instructed.
- the display image is controlled so that the type C avatar image 40C is displayed (FIG. 5). In this way, it is possible to accurately transmit information such as an incoming call to the user 1 in the manner mode by changing the display (expression) of the avatar image 40.
- the voice control unit 24 permits the voice output based on the operation of the manner mode button 12a or the surrounding environment estimated by the surrounding environment estimation unit 22 and the user information (first manner mode).
- the mode) or the manner mode (second mode) for prohibiting the audio output is commanded
- the configuration of the mode command section is not limited to this.
- the first mode or the second mode may be instructed only based on the surrounding environment estimated by the surrounding environment estimating unit 22. What is the configuration of the speaker 14 as a voice output unit for outputting a voice, the configuration of the display 11 as a display unit for displaying an image, and the configuration of the microphone 13 as a voice input unit for inputting a voice uttered by the user 1 But it's okay.
- the output determination unit 23 determines that it is necessary to output information to the user 1 when there is an incoming call or email, but the configuration of the output determination unit is not limited to this. For example, even when an emergency alert or the like is received, it may be determined that information output to the user is necessary. It may be determined that it is not necessary to always output information to the user 1 when there is an incoming call or e-mail, but it is necessary to output information according to the other party of the telephone or e-mail. It may be determined that output of information is necessary when a message from the other party is recorded after the incoming call. Even if the controller 20 determines from the contents of the recorded message or mail whether or not the information with high urgency is received, and when the information with high urgency is received, it is determined that the information output is necessary. Good.
- the user information including the preference information of the user 1 is stored in the storage unit 27 of the user terminal 10 and the storage unit 34 of the server device 30, but is stored only in the storage unit 27 of the user terminal 10.
- the configuration of the user information storage unit is not limited to that described above.
- the speaker 14 is controlled to output the voice corresponding to the utterance of the user 1 input from the microphone 13 based on the user information, but the configuration of the voice control unit is not limited to the above.
- the avatar image 40 is displayed as an example of the agent image including the face image of the agent in response to a command from the image control unit 25. More specifically, when the voice uttered by the user 1 (voice asking the user terminal 10) is input in the manner mode, the avatar image 40B of type B as the first agent image having the first face image, When it is determined that the voice output by the user 1 is not input and the information output to the user 1 is necessary in the manner mode, the type C avatar image 40C is displayed as the second agent image having the second face image.
- the type A avatar image 40A is displayed as the third agent image having the third face image. I decided to do it.
- the first face image, the second face image having a different expression from the first face image, and the third face image different from the first face image and the second face image may have any form. Not only the face image but also an image representing the form of the body may be displayed as the first agent image, the second agent image, and the third agent image.
- the third agent image is the same image as the avatar image 40A in the normal mode, it may be an image different from the avatar image 40A.
- the same avatar image 40B is displayed in the case of interacting with the avatar in the manner mode and in the case of interacting with the avatar in the non-manner mode, but different avatar images 40 are displayed. You may do it. Therefore, the configuration of the image control unit is not limited to the above.
- the vibration actuator (vibration actuator) 18 is activated when an incoming call or the like occurs in the manner mode, but the configuration of the actuator control unit 26 is not limited to this.
- the vibration actuator and the actuator controller may be omitted.
- the information transmission system 100 is configured by the user terminal 10 and the server device 30, but the function of the server device 30 may be provided in the user terminal 10 and the server device 30 may be omitted.
- the information transmission system can be configured by the user terminal alone.
- 10 user terminal 11 display, 12a manner mode button, 13 microphone, 14 speaker, 16 communication unit, 18 actuator, 20 controller, 22 surrounding environment estimation unit, 23 output determination unit, 24 voice control unit, 25 image control unit, 26 Actuator control unit, 27 storage unit, 30 server device, 32 user information acquisition unit, 33 image signal generation unit, 34 storage unit, 100 information transmission system
Landscapes
- Engineering & Computer Science (AREA)
- Human Computer Interaction (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Theoretical Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- General Health & Medical Sciences (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- User Interface Of Digital Computer (AREA)
- Telephone Function (AREA)
Abstract
情報伝達装置は、音声出力部(14)による音声出力を許可する第1モードまたは音声出力を禁止する第2モードを指令するモード指令部(24)と、エージェントの顔画像を含むエージェント画像が表示されるように表示部(11)の表示画像を制御する画像制御部(25)と、ユーザへの情報の出力の要否を判定する出力判定部(23)と、を備える。画像制御部(25)は、モード指令部(24)により第2モードが指令されているとき、音声入力部(13)によりユーザの発話による音声が入力されると、第1顔画像を有する第1エージェント画像が表示される一方、モード指令部(24)により第2モードが指令されているとき、音声入力部(13)によりユーザの発話による音声が入力されずに出力判定部(23)により情報の出力が必要と判定されると、第2顔画像を有する第2エージェント画像が表示されるように表示画像を制御する。
Description
本発明は、音声や表示によりユーザに情報を伝達する情報伝達装置および情報伝達方法に関する。
従来、画面上に表示されるエージェントを介してユーザに情報を伝達するようにした装置が知られている(例えば特許文献1参照)。この特許文献1記載の装置では、画面上に表示されるキャラクタ型のエージェントとユーザとの対話により得られるユーザの応答情報から、エージェントの表示上の振る舞いを変更してユーザに対し情報を提示する。
ところで、ユーザにとってエージェントと対話することが好ましくない状況が存在し得るが、上記特許文献1記載の装置では、そのような状況においてユーザに適切に情報を伝達することが困難である。
本発明の一態様である情報伝達装置は、音声を出力する音声出力部と、画像を表示する表示部と、ユーザの発話による音声を入力する音声入力部と、音声出力部による音声出力を許可する第1モードまたは音声出力を禁止する第2モードを指令するモード指令部と、モード指令部により第1モードが指令されているとき、音声入力部により入力されたユーザの発話に対応する音声を出力するように音声出力部を制御する音声制御部と、エージェントの顔画像を含むエージェント画像が表示されるように表示部の表示画像を制御する画像制御部と、ユーザへの情報の出力の要否を判定する出力判定部と、を備える。画像制御部は、モード指令部により第2モードが指令されているとき、音声入力部によりユーザの発話による音声が入力されると、第1顔画像を有する第1エージェント画像が表示される一方、モード指令部により第2モードが指令されているとき、音声入力部によりユーザの発話による音声が入力されずに出力判定部により情報の出力が必要と判定されると、第1顔画像とは異なる表情の第2顔画像を有する第2エージェント画像が表示されるように表示画像を制御する。
本発明の他の態様は、表示部を介して情報を伝達する情報伝達方法であって、コンピュータが、音声出力を許可する第1モードまたは音声出力を禁止する第2モードを指令し、第1モードが指令されているとき、ユーザの発話に対応する音声を出力するように音声出力部を制御し、エージェントの顔画像を含むエージェント画像が表示されるように表示部に表示される表示画像を制御し、ユーザへの情報の出力の要否を判定することを含む。表示画像を制御することは、第2モードが指令されているとき、ユーザの発話による音声が入力されると、第1顔画像を有する第1エージェント画像が表示される一方、第2モードが指令されているとき、音声入力部によりユーザの発話による音声が入力されずに情報の出力が必要と判定されると、第1顔画像とは異なる表情の第2顔画像を有する第2エージェント画像が表示されるように表示画像を制御することを含む。
本発明によれば、ユーザと対話することが好ましくない状況において、エージェント画像の表示を介してユーザに適切に情報を伝達することができる。
以下、図1~図5を参照して本発明の実施形態について説明する。図1は、本発明の実施形態に係る情報伝達装置を有する情報伝達システム100の概略構成を示す図である。図1に示すように、情報伝達システム100は、ユーザ1が携帯するユーザ端末10と、サーバ装置30とを有する。
ユーザ端末10とサーバ装置30とは、インターネット網や携帯電話網等に代表される公衆無線通信網を含むネットワーク3に接続され、ネットワーク3を介して互いに通信可能である。なお、ネットワーク3には、所定の管理地域ごとに設けられた閉鎖的な通信網、例えば無線LAN、Wi-Fi(登録商標)等も含まれる。
ユーザ端末10は、ユーザ1により携帯して使用されるスマートフォンやタブレット端末、携帯電話、さらにはウェアラブル端末等、モニタやディスプレイなどの表示部を有する各種携帯端末により構成される。図1には、電話機能を有し、全体が略矩形の薄型に形成されたスマートフォンによりユーザ端末10を構成した例が示される。ユーザ端末10はユーザ固有の端末であり、ユーザ端末10には、情報伝達システム100を利用するユーザ1に紐づけられた識別ID(ユーザID)が付与される。
ユーザ端末10は、ディスプレイ11と、ボタン群12と、マイク13と、スピーカ14と、カメラ15と、通信ユニット16と、GPS受信機17と、アクチュエータ18と、コントローラ20と、を有する。
ディスプレイ11は、液晶ディスプレイ、有機ELディスプレイ、無機ELディスプレイ等の表示デバイスを備える。ディスプレイ11は全体が略矩形状に構成され、ディスプレイ11に文字、記号、図形等の各種画像が表示される。ディスプレイ11は、例えばタッチパネルにより構成され、各種情報を入力する入力部としても機能する。なお、ディスプレイ11の配置および形状は図1のものに限定されない。
ボタン群12は、ユーザ1により操作される複数のボタン12a~12cを有し、ボタン12a~12cの操作に応じて各種指令が入力される。いずれかのボタン(例えばボタン12a)は、例えばマナーモードを指令および解除するマナーモードボタンとして構成される。なお、ボタン群12を構成するボタンの個数、形状、配置は上述したものに限らない。
マイク13は、ユーザ1の発話を音声信号として取得する音声入力部を構成する。マイク13が取得した信号は、例えばA/D変換器を介し音声データとしてコントローラ20に入力される。
スピーカ14は、コントローラ20から送信される音声信号を音声に変換して出力する音出力部を構成する。スピーカ14は、通話時に通話相手の声を出力する。各種音声メッセージ、着信音、音楽等をスピーカ14から出力することもできる。
カメラ15は、CCDやCMOSセンサ等の撮像素子を有し、ディスプレイ11に面する人や物を撮影するインカメラである。ディスプレイ11の裏面にバックカメラを設けることもできる。なお、カメラ15の配置や個数は上述したものに限定されない。
通信ユニット16は、ネットワーク3を介して無線により通信を行う。ネットワーク3を介さずに通信を行うこともできる。
GPS受信機17は、複数のGPS衛星からの測位信号を受信し、受信した測位信号に基づいてユーザ端末10の絶対位置(緯度、経度など)を測定する。
アクチュエータ18は、ユーザ端末10の筐体に振動を与えるバイブレータなどにより構成される。アクチュエータ18は、マナーモード時に電話やメールの着信があったこと等をユーザ1に通知する必要があるときに、作動する。
コントローラ20は、動作回路としてのCPUなどの演算部と、ROM,RAMなどの記憶部と、その他の周辺回路(インターフェース回路等)とを有するマイクロコンピュータを含んで構成される。コントローラ20の機能的構成については後述する。
サーバ装置30は、機能的構成として、通信部31と、ユーザ情報取得部32と、画像信号生成部33と、記憶部34とを有する。
通信部31は、ネットワーク3を介してユーザ端末10と通信を行う。ユーザ情報取得部32は、ユーザ端末10からユーザ情報を取得する。取得した情報は、ユーザIDに対応付けて記憶部34に記憶される。ユーザ情報にはユーザ1の趣味や食べ物の好み等の嗜好情報が含まれる。嗜好情報はユーザ固有の情報であり、ユーザ1が発した音声やユーザ1に出力する音声等に含まれるキーワードやタッチパネル(ディスプレイ11)から入力される入力情報、ディスプレイ11に表示される出力情報等に基づいて、各ユーザ1に対応する嗜好情報が判断される。
画像信号生成部33は、各ユーザ1に対応するアバタ画像(図2参照)を決定し、アバタ画像に対応する画像信号を生成する。この場合、まず、記憶部34に記憶されたユーザ情報に基づいて、ユーザ1の性格を数値化する。例えば、ユーザ1の行動力、社交性、協調性、思いやりなどの複数の性格パラメータ毎に、それらの程度(行動力の程度、社交性の程度等)をそれぞれ数値化して表す。そして、性格パラメータの数値の分布を考慮して、ディスプレイ11に表示されるアバタ画像を決定し、そのアバタ画像に対応する画像信号を生成する。生成された画像信号は、通信部31を介して、対応するユーザ端末10に送信される。
図2は、ディスプレイ11に表示されるアバタ画像40の一例を示す図である。アバタ画像40は、ユーザ1の分身として擬人化されたアバタの画像であり、所定のキャラクタまたはエージェントの形態で表される。特に、本実施形態では、キャラクタの顔画像としてアバタ画像40が表される。より具体的には、キャラクタの目と眉毛と口とをそれぞれ模した形態によりアバタ画像40が表示される。
アバタ画像40は、ユーザ毎に異なるユーザ固有の画像であり、ユーザ1の性格を良好に反映したユーザ1の分身としてのアバタの形態を表す画像である。このため、ユーザ1はアバタ画像40に親近感を持ちやすい。本実施形態では、特定のユーザ1のアバタ画像40を、ユーザ端末10の使用状況に応じて変更する。なお、アバタ画像40の形態を、ユーザ端末10を介しユーザ1が好みに応じて設定することもできる。
図3は、本実施形態に係る情報伝達装置としてのユーザ端末10の制御構成を示すブロック図である。図3に示すように、コントローラ20には、マイク13と、カメラ15と、通信ユニット16と、GPS受信機17と、マナーモードボタン12aとからの信号が入力される。コントローラ20のCPUは、これらの入力信号に基づいて所定の処理を実行し、スピーカ14とディスプレイ11とアクチュエータ18とに制御信号を出力する。
コントローラ20は、機能的構成として、信号入力部21と、周辺環境推定部22と、出力判定部23と、音声制御部24と、画像制御部25と、アクチュエータ制御部26と、記憶部27とを有する。
信号入力部21は、マイク13とカメラ15とGPS受信機17とマナーモードボタン12aとからの信号を入力するとともに、通信ユニット16を介して受信したサーバ装置30からの信号、例えばサーバ装置30から送信される、アバタ画像40に対応する画像信号を入力する。
周辺環境推定部22は、GPS受信機17からの信号に基づいてユーザ端末10の周辺環境を推定する。この推定は、スピーカ14から音声を出力することが妥当か否か、すなわち、ユーザ1がユーザ端末10をマナーモードにすべき環境にいるか否かを判定するために用いられる。周辺環境推定部22は、マイク13やカメラ15が取得した信号に基づいて周辺環境を推定することもできる。
出力判定部23は、ユーザ1に情報を通知する必要があるか否かを判定する。例えば、通信ユニット16を介してユーザ端末10に電話やメールの着信があったとき、あるいは通信ユニット16を介してユーザ端末10に各種サーバからプッシュ通知があったとき、その旨をユーザ1に通知する必要があると判定する。
音声制御部24は、マナーモードボタン12aの操作によりマナーモードが指令されているか否かを判定する。マナーモードが指令されているときは(マナーモード時)、出力判定部23によりユーザ1に情報を通知する必要があると判定されたときであっても、スピーカ14からの音声出力を禁止する。したがって、例えば電話の着信があってもスピーカ14から音声は出力されない。一方、マナーモードが解除されているときは(非マナーモード時)、スピーカ14からの音声出力を許可する。したがって、例えば電話の着信があると、スピーカ14から音声(着信音)が出力される。
また、音声制御部24は、マナーモードボタン12aの操作の有無に拘らず、周辺環境推定部22により推定された周辺環境と、記憶部27に記憶されたユーザ情報とに基づいてユーザ端末10をマナーモードにすべきか否かを判定する。そして、マナーモードにすべきと判定すると、ユーザ端末10を自動的にマナーモードに設定する。なお、周辺環境推定部22により推定された周辺環境とユーザ情報とに応じてマナーモードを解除すべきか否かを判定し、マナーモードを解除すべきと判定すると、自動的にマナーモードを解除するように構成することもできる。
さらに音声制御部24は、非マナーモード時に、ユーザ1の質問に対する回答や提案、ユーザ1の回答に対するさらなる質問等を含む音声を、スピーカ14から出力させることもできる。すなわち、ユーザ1と対話を行うように音声を出力することができる。この場合、音声制御部24は、記憶部27の辞書データベース(不図示)に登録された単語を参照することで、マイク13を介して入力されたユーザ1からの音声信号を認識する。例えば、ユーザ1の発話によるユーザ端末10に対する質問や回答を認識する。次いで、認識された音声の内容に対応する応答内容を記憶部27の対話データベース(不図示)から抽出し、応答メッセージを生成する。このとき、記憶部27に記憶されたユーザ情報(嗜好情報等)に基づいて情報の絞り込みを行い、絞り込んだ情報を含むような応答メッセージを生成する。そして、応答メッセージに対応する音声を出力するようにスピーカ14を制御する。なお、ユーザ1からの音声信号をサーバ装置30に送信するとともに、ユーザ1の嗜好情報を考慮した応答メッセージをサーバ装置30が生成し、これをユーザ端末10が受信してスピーカ14から出力させるようにしてもよい。
さらにまた、音声制御部24は、マナーモード時にマイク13やカメラ15が取得した信号に基づいて、ユーザ1からユーザ端末10への問いかけの有無を判定する。より詳しくは、音声制御部24は、記憶部27の辞書データベースを参照してユーザ1から入力された音声信号に対応する単語を特定するとともに、特定された単語が、予め辞書データベースに登録された複数の問いかけの単語のいずれかにあたるか否かを判定し、これにより問いかけの有無を判定する。そして、ユーザ1から問いかけがあると判定されると、マナーモード時であっても、問いかけに対する応答の音声をスピーカ14から出力させる。この場合、マナーモードを一時的に解除して、スピーカ14からは問いかけに対する応答のみを出力させる。なお、マナーモードを一時的に解除するのではなく、マナーモードボタン12aが操作されたときと同様にマナーモードを解除してもよい。
画像制御部25は、ディスプレイ11の表示画像を制御する。例えば、サーバ装置30から送信された画像信号に対応するアバタ画像40が表示されるように、ディスプレイ11に制御信号を出力する。アバタ画像40は、ユーザ1の性格に対応する画像であるが、ユーザ端末10の使用状況に応じてアバタ画像40の表情が変化する。
図4は、ディスプレイ11に表示されるアバタ画像40の表情の変化の一例を示す図である。図4では、タイプA~タイプDの4つの表情のアバタ画像40(40A~40D)が示される。タイプAのアバタ画像40Aは、例えばユーザ端末10が待ち受け状態であるときの画像、すなわち通常時の画像である。アバタ画像40Aは、平常時の標準的な表情によって表される。
タイプBのアバタ画像40Bは、例えばマナーモードが解除された非マナーモード(通常モード)において、着信音を模した音声(例えば「ツーツー」などの音声)をアバタがユーザ1に対し発するときの画像であり、スピーカ14から音声が発せられるのに合わせてアバタ画像40Bの口が動く。非マナーモード時だけでなく、マナーモード時にユーザ1からの問いかけに応答してスピーカ14から音声が出力されるときも、タイプBのアバタ画像40Bが表示される。
タイプCのアバタ画像40Cは、例えばマナーモードにおいて、電話やメールの着信があったことをユーザ1に通知するときの画像である。ユーザ1に伝えたい情報があるが音声を出力できないため、アバタ画像40Cは、困ったような表情の画像となる。
複数のユーザ1をそれぞれ第1ユーザ1A、第2ユーザ1Bとし、各ユーザ1A,1Bがそれぞれ固有のユーザ端末10A,10Bを有するとき、ユーザ端末10A,10Bに表示されるアバタ画像40は互いに異なる。例えば、タイプCのアバタ画像40Cは、第1ユーザ1Aのユーザ端末10Aに表示されるアバタの画像である。これに対し、タイプDのアバタ画像40Dは、タイプCが表示されるときと同様の状況における第2ユーザ1Bのユーザ端末10Bに表示されるアバタの画像である。タイプDのアバタ画像40Dは、タイプCのアバタ画像40Cとは異なり、怒ったような表情の画像で表される。
タイプCとタイプDとでアバタ画像40C,40Dの形態(画像の表情)が異なるのは、ユーザ1A,1Bの性格が互いに異なるからである。すなわち、サーバ装置30の画像信号生成部33は、各ユーザ1A,1Bの性格を考慮して画像信号を生成する。このため、ユーザ端末10A,10Bが互いに同一の使用状況(例えばマナーモード時)であり、アバタ画像40の種類が同一(同一キャラクタの顔画像)であったとしても、表示されるアバタ画像40の表情がユーザ端末10A,10B毎に異なる。
これによりマナーモード時に電話やメールの着信があったことを、各ユーザ1A,1Bに良好に通知することができる。つまり、困った表情とした方が伝わりやすいユーザ1Aに対してはアバタ画像40Cを表示させ、怒った表情とした方が伝わりやすいユーザ1Bに対してはアバタ画像40Dを表示させるように構成することで、画像表示を介してユーザ1A,1Bに情報を良好に伝達することができる。
アクチュエータ制御部26は、マナーモード時に出力判定部23がユーザ1に情報を通知する必要があると判定すると、アクチュエータ18に制御信号を出力し、ユーザ端末10を振動させる。これによりユーザ1は電話の着信等があったことを容易に気付くことができる。アクチュエータ18は、ユーザ端末10の設定変更により常にオフ状態とすることもできる。また、非マナーモード時に着信音の出力と同時にアクチュエータ18を作動させることもできる。
図5は、予め図3の記憶部27に記憶されたプログラムに従い、コントローラ20のCPUで実行される処理の一例を示すフローチャートである。このフローチャートに示す処理は、例えばユーザ端末10の電源オンにより開始され、所定周期で繰り返される。ユーザ端末10の電源がオンすると、ユーザ端末10とサーバ装置30との間で通信が開始され、サーバ装置30は、ユーザIDに対応する複数のアバタ画像40(例えば第1ユーザ1Aに対応するアバタ画像40A~40C)を決定するとともに、複数のアバタ画像40に対応する画像信号をユーザ端末10に送信する。ユーザ端末10は、この画像信号を記憶部27に記憶した状態で図5の処理を開始する。
まず、ステップS1で、マナーモードボタン12aの操作により手動で、あるいは周辺環境に応じて自動でマナーモードに設定されたか否かを判定する。ステップS1で否定されるとステップS2に進み、マイク13やカメラ15が取得した信号に基づいて、ユーザ1からユーザ端末10に対し問いかけがあるか否かを判定する。ステップS2で肯定されるとステップS5に進み、否定されるとステップS3に進む。ステップS3では、ディスプレイ11に制御信号を出力し、例えば図4のタイプAのアバタ画像40Aを表示させる。
一方、ステップS1で肯定されるとステップS4に進み、ステップS2と同様、マイク13やカメラ15が取得した信号に基づいて、ユーザ1からユーザ端末10に対し問いかけがあるか否かを判定する。ステップS4で肯定されるとステップS5に進み、スピーカ14に制御信号を出力し、問いかけに対する応答の音声をスピーカ14から出力させる。次いで、ステップS6で、ディスプレイ11に制御信号を出力し、例えば図4のタイプBのアバタ画像40Bを表示させる。
これに対し、ステップS4で否定されるとステップS7に進み、通信ユニット16を介して受信した信号に基づいて、ユーザ端末10に電話の着信あり等の情報を通知(伝達)する必要があるか否か、すなわちユーザ1への通知の要否を判定する。ステップS7で否定されると、ステップS3に進み、例えばタイプAのアバタ画像40Aをディスプレイ11に表示させる。なお、タイプA~タイプC以外のアバタ画像40を表示させるようにしてもよい。
ステップS7で肯定されると、ステップS8に進み、アクチュエータ18に制御信号を出力し、ユーザ端末10を振動させる。これによりユーザ1は、着信があったことを認識することができる。次いで、ステップS9で、ディスプレイ11に制御信号を出力し、例えば図4のタイプCのアバタ画像40Cを表示させる。これによりユーザ1は、画像表示によっても着信等があったことを認識することができる。この場合、文字や記号ではなくアバタ画像40の表情の変化によってユーザ1に情報が伝達されるため、ユーザ1はユーザ端末10から快適なサービスを受けることができ、ユーザ端末10の使用に対する高い満足感が得られる。
本実施形態によれば以下のような作用効果を奏することができる。
(1)情報伝達装置としてのユーザ端末10は、音声を出力するスピーカ14と、画像を表示するディスプレイ11と、ユーザ1の発話による音声を入力するマイク13と、マナーモードボタン12a等からの信号に基づきスピーカ14からの音声出力を許可する通常モード(非マナーモード)または音声出力を禁止するマナーモードを指令するとともに、非マナーモードが指令されているとき、マイク13により入力されたユーザ1の発話に対応する音声を出力するようにスピーカ14を制御する音声制御部24と、エージェント(アバタ)の顔画像を含むアバタ画像40が表示されるようにディスプレイ11の表示画像を制御する画像制御部25と、ユーザ1への情報の出力の要否を判定する出力判定部23と、を備える(図3)。画像制御部25は、音声制御部24によりマナーモードが指令されているとき、マイク13からユーザ1の発話による音声(例えばユーザ端末10への問いかけの音声)が入力されると、例えば図4のタイプBの表情のアバタ画像40Bが表示される一方、音声制御部24によりマナーモードが指令されているとき、マイク13からユーザ1の発話による音声が入力されずに、出力判定部23により情報の出力が必要と判定されると、例えば図4のタイプCの表情のアバタ画像40Cが表示されるように表示画像を制御する(図5)。
(1)情報伝達装置としてのユーザ端末10は、音声を出力するスピーカ14と、画像を表示するディスプレイ11と、ユーザ1の発話による音声を入力するマイク13と、マナーモードボタン12a等からの信号に基づきスピーカ14からの音声出力を許可する通常モード(非マナーモード)または音声出力を禁止するマナーモードを指令するとともに、非マナーモードが指令されているとき、マイク13により入力されたユーザ1の発話に対応する音声を出力するようにスピーカ14を制御する音声制御部24と、エージェント(アバタ)の顔画像を含むアバタ画像40が表示されるようにディスプレイ11の表示画像を制御する画像制御部25と、ユーザ1への情報の出力の要否を判定する出力判定部23と、を備える(図3)。画像制御部25は、音声制御部24によりマナーモードが指令されているとき、マイク13からユーザ1の発話による音声(例えばユーザ端末10への問いかけの音声)が入力されると、例えば図4のタイプBの表情のアバタ画像40Bが表示される一方、音声制御部24によりマナーモードが指令されているとき、マイク13からユーザ1の発話による音声が入力されずに、出力判定部23により情報の出力が必要と判定されると、例えば図4のタイプCの表情のアバタ画像40Cが表示されるように表示画像を制御する(図5)。
これにより、マナーモード時にユーザ1から問いかけがあった場合と、ユーザ1からの問いかけはないが着信等があった旨をユーザ1に通知する必要がある場合とで、アバタ画像40の態様(画像の表情)が変化する。このため、マナーモード時において着信等があった旨を、スピーカ14から音声を出力することなく、ユーザ1に適切に伝達することができる。すなわち、表情の変化はアバタの感情を表すのに適しており、ユーザ1にアバタの感情の変化を伝えることで、ユーザ1は、例えば緊急性を要する通知があった旨を容易かつ正確に認識することが可能である。また、ユーザ1は、ユーザ端末10を介してアバタと対話することが可能であり、アバタに対し親近感を持つようになる。そして、そのアバタ画像40を用いてユーザ1に対し通知するため、ユーザ1の快適性が高まる。
(2)情報伝達装置としてのユーザ端末10は、ユーザ1の嗜好情報を含むユーザ情報を記憶する記憶部27およびサーバ装置30の記憶部34をさらに備える(図1,3)。音声制御部24は、マナーモード時であっても、マイク13を介してユーザの発話による音声が入力されると、記憶されたユーザ情報に基づいて、マイク13により入力されたユーザ1の発話に対応する音声を出力するようにスピーカ14を制御する。マナーモード時に音声が出力されると、ユーザ1はびっくりする可能性があるが、本実施形態では、ユーザ1から問いかけがあったときに、マナーモード時に音声が出力されるように構成するため、音声出力に対しユーザ1が驚くことはない。
(3)アバタ画像40B,40Cは、記憶されたユーザ情報に基づくユーザ1に固有の画像である。したがって、異なる性格のユーザ1A,1Bに対しそれぞれ異なるアバタ画像(例えば図4の40C,40D)が表示されるようになり、ユーザ1のアバタに対する親近感が一層高まる。
(4)情報伝達装置としてのユーザ端末10は、周辺環境を推定する周辺環境推定部22をさらに備える(図3)。音声制御部24は、記憶されたユーザ情報と周辺環境推定部22により推定された周辺環境とに基づいて、音声出力の許可または禁止を指令する。これによりマナーモードボタン12aを操作することなく、自動的にマナーモードに設定されるため、音声出力が好ましくない状況でアバタ画像40を介して発話の音声が出力されることがない。
(5)情報伝達装置としてのユーザ端末10は、振動発生用のアクチュエータ18と、マナーモード時に、出力判定部23により情報の出力が必要と判定されると、振動を発生するようにアクチュエータ18を制御するアクチュエータ制御部26と、をさらに備える(図3)。これによりユーザ1は、マナーモード時に着信等があった旨を容易に認識することができる。
(6)本実施形態に係る情報伝達方法は、ディスプレイ11を介して情報を伝達する情報伝達方法であって、コンピュータ(コントローラ20)が、音声出力を許可する非マナーモードまたは音声出力を禁止するマナーモードを指令し(ステップS1)、非マナーモードが指令されているとき、ユーザ1の発話に対応する音声を出力するようにスピーカ14を制御し(ステップS5)、エージェントの顔画像を含むアバタ画像40が表示されるようにディスプレイ11の表示画像を制御し(ステップS3,ステップS6,ステップS9)、ユーザ1への情報の出力の要否を判定することを含む(ステップS7)。表示画像を制御することは、マナーモードが許可されているときに、ユーザ1の発話による音声が入力されると、タイプBのアバタ画像40Bが表示される一方、マナーモードが指令されているとき、ユーザ1の発話による音声が入力されずに、情報の出力が必要と判定されると、タイプCのアバタ画像40Cが表示されるように表示画像を制御することを含む(図5)。このようにアバタ画像40の表示(表情)の変化を介して、マナーモード時にユーザ1に着信あり等の情報を的確に伝達することができる。
なお、上記実施形態では、マナーモードボタン12aの操作、または周辺環境推定部22により推定された周辺環境とユーザ情報とに基づいて、音声制御部24が音声出力を許可する非マナーモード(第1モード)または音声出力を禁止するマナーモード(第2モード)を指令するようにしたが、モード指令部の構成はこれに限らない。周辺環境推定部22により推定された周辺環境のみに基づいて第1モードまたは第2モードを指令するようにしてもよい。音声を出力する音声出力部としてのスピーカ14の構成、画像を表示する表示部としてのディスプレイ11の構成、およびユーザ1の発話による音声を入力する音声入力部としてのマイク13の構成は、いかなるものでもよい。
上記実施形態では、電話やメールの着信等があったときに出力判定部23がユーザ1への情報の出力が必要と判定するようにしたが、出力判定部の構成はこれに限らない。例えば緊急警報等を受信したときにも、ユーザへの情報の出力が必要と判定してもよい。電話やメールの着信等があったときに常にユーザ1への情報の出力が必要と判定するのではなく、電話やメールの相手に応じて情報の出力が必要と判定してもよい。着信後に相手からのメッセージが録音されたときに、情報の出力が必要と判定してもよい。コントローラ20が、録音されたメッセージやメールの内容から、緊急性が高い情報を受信したか否かを判定し、緊急性が高い情報を受信したときに、情報の出力が必要と判定してもよい。
上記実施形態では、ユーザ1の嗜好情報を含むユーザ情報を、ユーザ端末10の記憶部27とサーバ装置30の記憶部34とに記憶するようにしたが、ユーザ端末10の記憶部27のみに記憶するようにしてもよく、ユーザ情報記憶部の構成は上述したものに限らない。上記実施形態では、ユーザ情報に基づいて、マイク13から入力されたユーザ1の発話に対応する音声を出力するようにスピーカ14を制御したが、音声制御部の構成は上述したものに限らない。
上記実施形態では、画像制御部25からの指令により、エージェントの顔画像を含むエージェント画像の一例としてアバタ画像40を表示するようにした。より具体的には、マナーモード時にユーザ1の発話による音声(ユーザ端末10への問いかけの音声)が入力されると、第1顔画像を有する第1エージェント画像としてタイプBのアバタ画像40Bを、マナーモード時にユーザ1の発話による音声が入力されずにユーザ1への情報の出力が必要と判定されると、第2顔画像を有する第2エージェント画像としてタイプCのアバタ画像40Cを表示するようにした。また、マナーモード時にユーザ1の発話による音声が入力されずにユーザ1への情報の出力が不要と判定されると、第3顔画像を有する第3エージェント画像としてタイプAのアバタ画像40Aを表示するようにした。しかしながら、第1顔画像、第1顔画像とは異なる表情の第2顔画像、および第1顔画像、第2顔画像とは異なる第3顔画像の形態はいかなるものでもよい。顔画像だけでなく身体の形態を表す画像を、第1エージェント画像、第2エージェント画像および第3エージェント画像として表示するようにしてもよい。第3エージェント画像を通常モード時のアバタ画像40Aと同一の画像としたが、アバタ画像40Aと異なる画像としてもよい。上記実施形態では、マナーモード時にアバタと対話する場合と、非マナーモード時にアバタと対話する場合とで、同一のアバタ画像40Bが表示されるようにしたが、互いに異なるアバタ画像40が表示されるようにしてもよい。したがって、画像制御部の構成は上述したものに限らない。
上記実施形態では、マナーモード時に着信等があったときに振動用のアクチュエータ(振動アクチュエータ)18を作動するようにしたが、アクチュエータ制御部26の構成はこれに限らない。振動アクチュエータとアクチュエータ制御部とを省略することもできる。上記実施形態では、ユーザ端末10とサーバ装置30とにより情報伝達システム100を構成したが、サーバ装置30の機能をユーザ端末10に設け、サーバ装置30を省略することもできる。ユーザ端末単独で情報伝達システムを構成することができる。
以上の説明はあくまで一例であり、本発明の特徴を損なわない限り、上述した実施形態および変形例により本発明が限定されるものではない。上記実施形態と変形例の1つまたは複数を任意に組み合わせることも可能であり、変形例同士を組み合わせることも可能である。
10 ユーザ端末、11 ディスプレイ、12a マナーモードボタン、13 マイク、14 スピーカ、16 通信ユニット、18 アクチュエータ、20 コントローラ、22 周辺環境推定部、23 出力判定部、24 音声制御部、25 画像制御部、26 アクチュエータ制御部、27 記憶部、30 サーバ装置、32 ユーザ情報取得部、33 画像信号生成部、34 記憶部、100 情報伝達システム
Claims (6)
- 音声を出力する音声出力部と、
画像を表示する表示部と、
ユーザの発話による音声を入力する音声入力部と、
前記音声出力部による音声出力を許可する第1モードまたは音声出力を禁止する第2モードを指令するモード指令部と、
前記モード指令部により前記第1モードが指令されているとき、前記音声入力部により入力されたユーザの発話に対応する音声を出力するように前記音声出力部を制御する音声制御部と、
エージェントの顔画像を含むエージェント画像が表示されるように前記表示部の表示画像を制御する画像制御部と、
ユーザへの情報の出力の要否を判定する出力判定部と、
を備え、
前記画像制御部は、前記モード指令部により前記第2モードが指令されているとき、前記音声入力部によりユーザの発話による音声が入力されると、第1顔画像を有する第1エージェント画像が表示される一方、前記モード指令部により前記第2モードが指令されているとき、前記音声入力部によりユーザの発話による音声が入力されずに前記出力判定部により情報の出力が必要と判定されると、前記第1顔画像とは異なる表情の第2顔画像を有する第2エージェント画像が表示されるように表示画像を制御することを特徴とする情報伝達装置。 - 請求項1に記載の情報伝達装置において、
ユーザの嗜好情報を含むユーザ情報を記憶するユーザ情報記憶部をさらに備え、
前記音声制御部は、前記モード指令部により前記第2モードが指令されているときであっても、前記音声入力部によりユーザの発話による音声が入力されると、前記ユーザ情報記憶部に記憶されたユーザ情報に基づいて前記音声入力部により入力されたユーザの発話に対応する音声を出力するように前記音声出力部を制御することを特徴とする情報伝達装置。 - 請求項2に記載の情報伝達装置において、
前記第1エージェント画像および前記第2エージェント画像は、前記ユーザ情報記憶部に記憶されたユーザ情報に基づくユーザに固有の画像であることを特徴とする情報伝達装置。 - 請求項2または3に記載の情報伝達装置において、
周辺環境を推定する周辺環境推定部をさらに備え、
前記モード指令部は、前記ユーザ情報記憶部に記憶されたユーザ情報と前記周辺環境推定部により推定された周辺環境とに基づいて、前記第1モードまたは前記第2モードを指令することを特徴とする情報伝達装置。 - 請求項1~4のいずれか1項に記載の情報伝達装置において、
振動アクチュエータと、
前記モード指令部により前記第2モードが指令されているときに、前記出力判定部により情報の出力が必要と判定されると、振動を発生するように前記振動アクチュエータを制御するアクチュエータ制御部と、をさらに備えることを特徴とする情報伝達装置。 - 表示部を介して情報を伝達する情報伝達方法であって、コンピュータが、
音声出力を許可する第1モードまたは音声出力を禁止する第2モードを指令し、
前記第1モードが指令されているとき、ユーザの発話に対応する音声を出力するように音声出力部を制御し、
エージェントの顔画像を含むエージェント画像が表示されるように前記表示部に表示される表示画像を制御し、
ユーザへの情報の出力の要否を判定することを含み、
前記表示画像を制御することは、前記第2モードが指令されているとき、ユーザの発話による音声が入力されると、第1顔画像を有する第1エージェント画像が表示される一方、前記第2モードが指令されているとき、音声入力部によりユーザの発話による音声が入力されずに情報の出力が必要と判定されると、前記第1顔画像とは異なる表情の第2顔画像を有する第2エージェント画像が表示されるように前記表示画像を制御することを含む情報伝達方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/292,636 US11438450B2 (en) | 2018-11-13 | 2019-11-05 | Information transmission apparatus and information transmission method |
| JP2020556058A JP7080991B2 (ja) | 2018-11-13 | 2019-11-05 | 情報伝達装置および情報伝達方法 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2018213237 | 2018-11-13 | ||
| JP2018-213237 | 2018-11-13 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020100647A1 true WO2020100647A1 (ja) | 2020-05-22 |
Family
ID=70731526
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/043198 Ceased WO2020100647A1 (ja) | 2018-11-13 | 2019-11-05 | 情報伝達装置および情報伝達方法 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US11438450B2 (ja) |
| JP (1) | JP7080991B2 (ja) |
| WO (1) | WO2020100647A1 (ja) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH1118145A (ja) * | 1997-06-20 | 1999-01-22 | Sony Corp | 表示方法及び装置並びに通信装置 |
| JP2001217903A (ja) * | 2000-01-31 | 2001-08-10 | Hitachi Ltd | 電話機及び情報端末装置及びその着信通知方法 |
| JP2004020719A (ja) * | 2002-06-13 | 2004-01-22 | Fujikura Ltd | 光ファイバテープ心線 |
| JP2004180261A (ja) * | 2002-10-04 | 2004-06-24 | Nec Saitama Ltd | 携帯電話機及びそれに用いるキャラクタ表示演出方法並びにそのプログラム |
| JP2005196645A (ja) * | 2004-01-09 | 2005-07-21 | Nippon Hoso Kyokai <Nhk> | 情報提示システム、情報提示装置、及び情報提示プログラム |
-
2019
- 2019-11-05 US US17/292,636 patent/US11438450B2/en active Active
- 2019-11-05 WO PCT/JP2019/043198 patent/WO2020100647A1/ja not_active Ceased
- 2019-11-05 JP JP2020556058A patent/JP7080991B2/ja active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH1118145A (ja) * | 1997-06-20 | 1999-01-22 | Sony Corp | 表示方法及び装置並びに通信装置 |
| JP2001217903A (ja) * | 2000-01-31 | 2001-08-10 | Hitachi Ltd | 電話機及び情報端末装置及びその着信通知方法 |
| JP2004020719A (ja) * | 2002-06-13 | 2004-01-22 | Fujikura Ltd | 光ファイバテープ心線 |
| JP2004180261A (ja) * | 2002-10-04 | 2004-06-24 | Nec Saitama Ltd | 携帯電話機及びそれに用いるキャラクタ表示演出方法並びにそのプログラム |
| JP2005196645A (ja) * | 2004-01-09 | 2005-07-21 | Nippon Hoso Kyokai <Nhk> | 情報提示システム、情報提示装置、及び情報提示プログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2020100647A1 (ja) | 2021-10-21 |
| US11438450B2 (en) | 2022-09-06 |
| JP7080991B2 (ja) | 2022-06-06 |
| US20220014618A1 (en) | 2022-01-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8600452B2 (en) | Mobile communications device with distinctive vibration modes | |
| EP2621153B1 (en) | Portable terminal, response message transmitting method and server | |
| US8489062B2 (en) | System and method for sending an emergency message selected from among multiple emergency message types from a wireless communications device | |
| KR101919858B1 (ko) | 모바일 개인 비서를 위한 방법 및 그 전자 장치 | |
| TWI594611B (zh) | 智慧對話方法和使用所述方法的電子裝置 | |
| KR20110036902A (ko) | 전자 장치 이벤트를 위한 통지 프레임워크의 전개 | |
| JP2004531910A5 (ja) | ||
| KR101609585B1 (ko) | 청각 장애인용 이동 통신 단말기 | |
| JP7080991B2 (ja) | 情報伝達装置および情報伝達方法 | |
| US9420111B2 (en) | Communication device, method, and program | |
| CN113726956A (zh) | 一种来电接听控制方法、装置、终端设备及存储介质 | |
| JP7222802B2 (ja) | インターホンシステム | |
| JP4877595B2 (ja) | 携帯端末、予定通知方法、及びプログラム | |
| JP4791470B2 (ja) | 携帯電話機、報知方法、及びプログラム | |
| JP2004135224A (ja) | 携帯電話機 | |
| KR100768666B1 (ko) | 화자에 따라 자연스럽게 동작하는 아바타를 이용한화상통화 방법 및 시스템 | |
| JP6790619B2 (ja) | 発話判定装置、発話判定システム、プログラム及び発話判定方法 | |
| KR200234151Y1 (ko) | 청각 장애인용 전화기 | |
| JP2007049525A (ja) | 通話端末及び通信端末の動作方法 | |
| WO2024070550A1 (ja) | システム、電子機器、システムの制御方法、及びプログラム | |
| KR20040016778A (ko) | 아바타를 이용한 화상통화 방법 및 시스템 | |
| JP2004072292A (ja) | 携帯電話システムおよび携帯電話器 | |
| WO2006033809A1 (en) | Communication device with term analysis capability, associated function triggering and related method | |
| JP5650036B2 (ja) | インターホンシステム | |
| KR20090086648A (ko) | 사용자 입 모양 인식에 의한 아바타 제어 방법 및 시스템 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19884170 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2020556058 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19884170 Country of ref document: EP Kind code of ref document: A1 |