WO2015174172A1 - 制御装置およびメッセージ出力制御システム - Google Patents

制御装置およびメッセージ出力制御システム Download PDF

Info

Publication number
WO2015174172A1
WO2015174172A1 PCT/JP2015/060973 JP2015060973W WO2015174172A1 WO 2015174172 A1 WO2015174172 A1 WO 2015174172A1 JP 2015060973 W JP2015060973 W JP 2015060973W WO 2015174172 A1 WO2015174172 A1 WO 2015174172A1
Authority
WO
WIPO (PCT)
Prior art keywords
message
output
utterance
information
output device
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2015/060973
Other languages
English (en)
French (fr)
Inventor
樹利 杉山
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sharp Corp
Original Assignee
Sharp Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sharp Corp filed Critical Sharp Corp
Priority to CN201580021613.6A priority Critical patent/CN106233378B/zh
Priority to JP2016519160A priority patent/JP6276400B2/ja
Priority to US15/306,819 priority patent/US10127907B2/en
Publication of WO2015174172A1 publication Critical patent/WO2015174172A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J11/00Manipulators not otherwise provided for
    • B25J11/0005Manipulators having means for high-level communication with users, e.g. speech generator, face recognition means
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L13/00Speech synthesis; Text to speech systems
    • G10L13/02Methods for producing synthetic speech; Speech synthesisers
    • G10L13/04Details of speech synthesis systems, e.g. synthesiser structure or memory management
    • G10L13/047Architecture of speech synthesisers
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L13/00Speech synthesis; Text to speech systems
    • G10L13/08Text analysis or generation of parameters for speech synthesis out of text, e.g. grapheme to phoneme translation, prosody generation or stress or intonation determination
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/28Constructional details of speech recognition systems
    • G10L15/30Distributed recognition, e.g. in client-server systems, for mobile phones or network applications
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/22Interactive procedures; Man-machine interfaces
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L2015/088Word spotting
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • G10L2015/221Announcement of recognition results

Definitions

  • the present invention relates to a control device and a message output control system for controlling a message output from an output device to a user.
  • Patent Document 1 discloses a technology in which an interactive robot acquires a user's speech and transmits it to a computer, and the computer determines a response sentence corresponding to the user's speech as an interactive robot response sentence. ing.
  • Japanese Patent Publication Japanese Patent Laid-Open No. 2005-3747 (Publication Date: January 6, 2005)”
  • the present invention has been made in view of the above problems, and an object thereof is to realize a control device and a message output control system capable of preventing a plurality of output devices from outputting the same message to a user. It is in.
  • a control device is a control device that causes a plurality of output devices including a first output device and a second output device to output a message.
  • a control unit that causes the first output device to output a message different from that of the second output device, or causes the first output device not to output a message.
  • FIG. 2 is a diagram showing an overview of the message output control system 100 according to the present embodiment.
  • the message output control system 100 is a system for each of a plurality of robots to present to the user a response message suitable for the content spoken by the user (spoken content).
  • the message output control system 100 includes robots (output devices, first output devices, second output devices) 1 and 2 and a server (control device) 3 as shown in the figure.
  • robots output devices, first output devices, second output devices
  • server control device
  • the number of robots is not limited as long as there are two or more robots. The same applies to the following embodiments.
  • the robot 1 and the robot 2 are devices having a function of interacting with the user by voice, detect the user's utterance, and output (present) a response message to the utterance by voice or text.
  • the server 3 is a device that controls the content of the remarks of the robot 1 and the robot 2.
  • the robot 1 and the robot 2 are information for determining a response message to be answered by the own device in response to the utterance, and information related to the detected utterance
  • the utterance information (output information) is transmitted to the server 3.
  • the server 3 When the server 3 receives the utterance information from the robot 1 or the robot 2, the server 3 determines a response message corresponding to the utterance information and transmits it to the robot that is the transmission source of the utterance information. For example, when the user utters “Tell me tomorrow's weather” to the robot 1, the utterance information about the utterance is transmitted from the robot 1 to the server 3. In this case, the server 3 determines the response message of the robot 1 as the message A (the line “sunny and sometimes cloudy”) according to the utterance content indicated by the utterance information, and tells the robot 1 Send message A. Then, the robot 1 outputs the received message A. That is, the robot 1 says to the user that it is sunny and sometimes cloudy.
  • the server 3 when there are a plurality of pieces of utterance information having the same message among the pieces of utterance information received from a plurality of robots, for example, when a plurality of robots detect the same utterance of the user, the server 3 responds with the same response. In order not to output a message, a response message corresponding to the utterance content and a response message different from that of the other robot is returned to the transmission source robot. For example, when a user's utterance is detected by both robots 1 and 2, utterance information about the same utterance of the user is transmitted from each robot.
  • the robot 2 may also transmit the utterance information to the server 3 almost simultaneously with the robot 1.
  • the server 3 causes the robot 1 to output a message A, and causes the robot 2 to output a message A ′ that is a response message corresponding to the user's utterance content and is different from the message A.
  • the robot 2 speaks differently from the robot 1 such as “the maximum temperature is 24 degrees”.
  • the message output control system 100 when a plurality of robots detect the same utterance, for example, when a user utters toward (or near) a plurality of robots, the plurality of robots. Can return different response messages to the user.
  • FIG. 1 is a block diagram showing a main configuration of the robot 1, the robot 2, and the server 3 according to the present embodiment.
  • the main configuration inside the robot 2 is the same as that of the robot 1. Therefore, illustration and description of the configuration of the robot 2 are omitted.
  • the robot 1 includes a device control unit 10, a microphone 20, a camera 30, an output unit 40, and a device communication unit 50.
  • the microphone 20 converts a user's utterance into voice data and acquires it.
  • the microphone 20 transmits the acquired audio signal to the utterance information creation unit 11 described later.
  • the camera 30 captures a user as a subject.
  • the camera 30 transmits the captured image (user image) to the utterance information creation unit 11 described later.
  • the user specifying unit 13 described later specifies a user from data other than image data, the camera 30 may not be provided.
  • the output unit 40 outputs a message received from the device control unit 10. For example, a speaker or the like may be used as the output unit 40.
  • the device communication unit 50 is for transmitting and receiving data to and from the server 3.
  • the device control unit 10 controls the robot 1 in an integrated manner.
  • the device control unit 10 transmits the utterance information generated by the utterance information creation unit 11 described later to the server 3. Further, the device control unit 10 receives a message output from the robot 1 from the server 3 and causes the output unit 40 to output the message. More specifically, the device control unit 10 includes an utterance information creation unit 11, an utterance content identification unit 12, and a user identification unit 13.
  • the device control unit 10 includes a timer unit (not shown). The timer unit can be realized by, for example, a real time clock of the device control unit 10.
  • the utterance information creation unit 11 creates utterance information. Note that the utterance information creation unit 11 may transmit a response message request to the server 3 together with the utterance information.
  • the utterance information creation unit 11 includes an utterance content identification unit 12 and a user identification unit 13 in order to generate various types of information included in the utterance information.
  • the utterance content identification unit 12 identifies the utterance content of the user and generates information (utterance content information) indicating the utterance content.
  • the utterance content information is not particularly limited as long as the message determination unit 83 of the server 3 described later can identify a response message corresponding to the utterance content indicated by the utterance content information.
  • the utterance content specifying unit 12 specifies and extracts one or more keywords from the voice data of the user's uttered voice as the utterance content information. Note that the utterance content identification unit 12 may use the speech data itself received from the microphone 20 or the speech data extracted from the speech data as speech content information.
  • the user identification unit 13 identifies the user who has spoken.
  • the method for specifying the user is not particularly limited.
  • a captured image may be acquired from the camera 30 and the user who is reflected in the captured image may be specified by image analysis.
  • the user specifying unit 13 may specify a user by analyzing a voiceprint included in the audio data acquired by the microphone 20. Thus, by performing image analysis or voiceprint analysis, it is possible to accurately (accurately) specify the user who spoke.
  • the utterance information creation unit 11 includes the time when the utterance information is created (speech time information), the utterance content information generated by the utterance content identification unit 12, and information (user identification information) indicating the user identified by the user identification unit 13. Create utterance information including. That is, the utterance information creation unit 11 creates utterance information including information indicating “when (time), who (user-specific information), what was spoken (utterance content)”. Then, the utterance information creation unit 11 transmits the created utterance information to the server 3 via the device communication unit 50.
  • the server 3 includes a server communication unit 70, a server control unit 80, a temporary storage unit 82, and a storage unit 90.
  • the server communication unit 70 is for transmitting / receiving data to / from the robot 1 and the robot 2.
  • the temporary storage unit 82 temporarily stores utterance information.
  • the temporary storage unit 82 receives the utterance information and the response message determined according to the utterance information from the message determination unit 83, and stores the utterance information and the response message in association with each other for a predetermined period.
  • the “predetermined period” referred to here is at least a period longer than a period during which utterance information based on the same utterance can reach the server 3 from another robot (that is, a communication and processing time lag between the robots). desirable.
  • the storage unit 90 stores a message table 91 that is read out when a message determination unit 83 described later determines a message (message determination processing).
  • a message table 91 that is read out when a message determination unit 83 described later determines a message (message determination processing).
  • FIG. 3 is a diagram illustrating a data configuration of the message table 91.
  • the message table 91 illustrated in FIG. 3 includes an “utterance content” column, a “message” column, and a “priority” column, and information in the “message” column is associated with information in the “utterance content” column. It has been. Further, information in the “priority” column is associated with information in the “message” column.
  • the “utterance content” column stores information indicating the user's speech content.
  • the information in the “utterance content” column is stored in a format that the message determination unit 83 can collate with the speech content information included in the received speech information.
  • the utterance content information is a keyword extracted from the user's utterance
  • information indicating a keyword (or a combination of a plurality of keywords) indicating the utterance content may be stored as data in the “utterance content” column.
  • the speech content information is speech speech data itself, speech data that can be matched with the speech data may be stored as data in the “speech content” column.
  • the “message” column stores information indicating response messages to be output to the robot 1 and the robot 2. As illustrated, in the message table 91, a plurality of response messages are associated with information in one “utterance content” column.
  • the “priority” column stores information indicating the priority order of each of a plurality of response messages in the “message” column corresponding to one utterance content indicated in the “utterance content” column.
  • the method of assigning the priority is not particularly limited.
  • the information in the “priority” column is not essential. When the information in the “priority” column is not used, one response message may be selected at random from each response message corresponding to the “utterance content” column.
  • the server control unit 80 controls the server 3 in an integrated manner, and includes an utterance information reception unit 81 and a message determination unit (determination unit, control unit, message generation unit) 83.
  • the utterance information receiving unit 81 acquires the utterance information transmitted by the robot 1 and the robot 2 via the server communication unit 70.
  • the utterance information reception unit 81 transmits the acquired utterance information to the message determination unit 83.
  • the message determination unit 83 determines a response message.
  • the message determination unit 83 receives the utterance information from the utterance information reception unit 81, the message determination unit 83 corresponds to the other utterance information stored in the temporary storage unit 82 (utterance information received before the utterance information) and the other utterance information. Read the attached response message.
  • the utterance information received by the utterance information receiving unit 81 that is, the utterance information from which the server 3 determines the response message from now on, is referred to as “target utterance information”.
  • the message determination unit 83 detects speech information corresponding to the same speech as the target speech information from other speech information.
  • the message determination unit 83 includes a time determination unit 84, an utterance content determination unit 85, and a user determination unit 86 in order to perform the above detection.
  • the time determination unit 84 determines whether or not the difference between the time included in the other utterance information and the time included in the target utterance information is within a predetermined range. Note that the time determination unit 84 measures not only the time included in the utterance information but also the time when the server 3 receives the utterance information, and stores this together with the utterance information in order to improve the accuracy of the determination. Also good. Then, it may be determined whether or not the difference between the reception time of the target utterance information and the reception time of other utterance information is within a predetermined range. Thus, for example, even when the time measured by the robot 1 and the time measured by the robot 2 are different, the time can be compared with high accuracy.
  • the utterance content determination unit 85 determines whether the utterance content information of the other utterance information matches the utterance content information of the target utterance information.
  • the user determination unit 86 determines whether or not the user identification information of the target utterance information matches the user identification information of other utterance information.
  • the message determination unit 83 determines that the difference between the target utterance information and the time is within a predetermined range, and the utterance content matches, and Other speech information that matches the user identification information is detected.
  • the other utterance information detected by the message determination unit 83 is referred to as detection target utterance information.
  • the detection target utterance information is estimated to be utterance information created by a plurality of robots detecting utterances of the same user almost simultaneously.
  • the detection target utterance information is estimated to be utterance information created in correspondence with the same utterance as the target utterance information.
  • the message determination unit 83 searches the message table 91 in the storage unit 90 and specifies a response message corresponding to the target utterance information.
  • the message determination unit 83 prevents the response message of the target utterance information from being the same as the response message of the detection target utterance information. Specifically, among the response messages corresponding to the utterance content information of the target utterance information, a response message that is different from the response message of the detection target utterance information and has the highest priority is specified. When there is no detection target utterance information, a response message with the highest priority is specified among the response messages corresponding to the utterance content information of the target utterance information. Then, the message determination unit 83 stores the target utterance information and the identified response message in the temporary storage unit 82 in association with each other.
  • the message determination unit 83 refers to the temporary storage unit 82 and detects other utterance information with the same response message. You may change to Even with this configuration, it is possible to prevent a plurality of robots from outputting the same response message. Even in this configuration, when the time determination unit 84, the utterance content determination unit 85, and the user determination unit 86 perform the determination process using the detected other utterance information and the target utterance information, the two utterance information matches. The determined response message may be changed to another response message.
  • FIG. 4 is a sequence diagram showing data communication between the robot 1 and the robot 2 and the server 3 in time series.
  • the subroutine shown in S300 and S400 of FIG. 4 shows a series of message determination processes (S302 to S310) shown in FIG.
  • the time included in the utterance information of the robot 1 is T1
  • the utterance content information is Q1
  • the user identification information is U1
  • the utterance information including these is described as utterance information (T1, Q1, U1).
  • the time included in the utterance information of the robot 2 is T2
  • the utterance content information is Q2
  • the user identification information is U2
  • the utterance information including these is described as utterance information (T2, Q2, U2).
  • the robot 1 and the robot 2 acquire the utterance with the microphone 20 and transmit it to the utterance information creation unit 11 (S100 and S200).
  • the user who speaks with the camera 30 is photographed, and the photographed image is transmitted to the speech information creating unit 11.
  • the utterance content identification unit 12 of the utterance information creation unit 11 identifies a keyword included in the user's utterance from the voice data.
  • the user specifying unit 13 specifies a user from the captured image.
  • the utterance information creation unit 11 creates utterance information to which the clocked time (the time when the utterance information is created), the utterance content information, and the user specifying information are added (S102 and 202), and transmits the utterance information to the server 3. .
  • the server 3 receives the utterance information (T1, Q1, U1) from the robot 1, it performs a message determination process (S300).
  • the utterance information of the robot 1 is utterance information that is received first by the server 3 as shown in the drawing (received in a state where no other utterance information is stored in the temporary storage unit 82 of the server 3).
  • T1, Q1, U1 is the target utterance information.
  • the message determination unit 83 determines the response message having the highest priority among the response messages corresponding to Q1 as the response message that causes the robot 1 to output. Then, the server 3 transmits the determined message A to the robot 1. When the robot 1 receives the message A, it outputs it from the output unit 40 (S104).
  • the server 3 performs message determination processing on the utterance information (T2, Q2, U2) received from the robot 2 (S400).
  • the message determination unit 83 uses the time determination unit 84, the utterance content determination unit 85, and the user determination unit 86 to transmit the target utterance information (T2, Q2, U2) and other utterance information (T1, Q1, U1). Is determined to match.
  • the server 3 determines a message A ′ different from the response message A corresponding to the utterance information (T1, Q1, U1) and transmits it to the robot 2.
  • the server 3 determines the message A ′ according to Q2 without considering the response message A and transmits it to the robot 2.
  • the robot 2 receives the message A ', it outputs it from the output unit 40 (S204).
  • FIG. 5 is a flowchart showing the flow of the message determination process. Note that the processing order of S302 to S306 in FIG. 5 may be changed.
  • the message determination unit 83 Upon receiving the utterance information (T2, Q2, U2), the message determination unit 83 performs the processing of S302 to 306, thereby the same as the target utterance information (T2, Q2, U2 in this example) among other utterance information Speech information corresponding to the speech is detected.
  • the time determination unit 84 included in the message determination unit 83 includes the time T2 of the target utterance information (T2, Q2, U2) and the time T1 of other utterance information (T1, Q1, U1 in this example). It is determined whether or not the difference is within a predetermined range (S302).
  • the utterance content determination unit 85 determines whether Q2 and Q1 match (S304). When Q2 and Q1 match (YES in S304), the user determination unit 86 determines whether U2 and U1 match (S306). If U2 and U1 match (YES in S306), the message determination unit 83 uses other utterance information (T1, Q1, U1) as utterance information based on the same utterance as the target utterance information (T2, Q2, U2). Detected as (detection target utterance information).
  • the message determination unit 83 determines a message according to Q2. Specifically, the message determination unit 83 specifies the response message corresponding to Q2 by collating the “utterance content” column of the message table 91 with the speech content indicated by Q2. At this time, the message determination unit 83 selects the message with the highest priority among the response messages corresponding to the detection target utterance information (T1, Q1, U1), that is, the messages different from the response message corresponding to Q1. It is determined as a message output by the robot that is the transmission source of the utterance information (S308).
  • the message determination unit 83 detects It is determined that the target utterance information does not exist. If two or more other pieces of utterance information are stored, the processes of S302 to S306 are performed for each piece of other utterance information. Then, the message determination unit 83 refers to the message table 91 to output a response message having the highest priority among the response messages corresponding to the target utterance information, from the robot that is the transmission source of the target utterance information. (S310).
  • the robot may spontaneously output a message to the user instead of presenting a response message in response to the user's utterance.
  • a second embodiment of the present invention will be described with reference to FIG.
  • a message that the robot voluntarily outputs (speaks) is referred to as a “spontaneous message”.
  • FIG. 6 is a diagram showing an outline of the message output control system 200 according to the present embodiment.
  • the message output control system 200 includes a plurality of robots (robot (output device) 1 and robot (output device) 2) and a server (control device) 4.
  • the configuration of the server 4 is the same as that of the server 3 (see FIG. 1), but the server 4 is different from the server 3 in that the robots 1 and 2 output a spontaneous message using a preset event (condition) as a trigger. .
  • the “preset event (condition)” is not particularly limited. In the following description, as an example, an example will be described in which the robot 1 and the robot 2 output a spontaneous message when the approach of the user is detected (the user enters the detection range of the robot).
  • Robot 1 and Robot 2 notify the server 4 when the user enters the detection range of the own device.
  • the server 4 determines a spontaneous message and transmits it to the robot that sent the notification. For example, when the server 4 receives the notification from the robot 1, the server 4 determines the spontaneous message of the robot 1 as the spontaneous message B (the dialogue “Let's play”) and causes the robot 1 to output it.
  • the server 4 determines the spontaneous message of the robot 2 so that the robot 2 does not output the same spontaneous message as the robot 1. For example, the server 4 determines the spontaneous message B ′ (the line “Would you help?”) Different from the spontaneous message B as the spontaneous message of the robot 2 and causes the robot 2 to output it.
  • the message output control system 200 can prevent a plurality of robots from outputting the same spontaneous message to the same user at the same timing.
  • the message presentation is not limited to this example as long as the user can recognize the message.
  • the message may be presented by registering a message for the user on an electronic medium such as an electronic bulletin board.
  • the server 3 or 4 may register a message as a substitute for the robot 1 and the robot 2 (with the names of the robot 1 and the robot 2).
  • the server 3 generates a response message to the user's voice utterance has been shown, but a message corresponding to information written (posted) on the electronic medium by the user may be generated.
  • a message corresponding to information written (posted) on the electronic medium by the user may be generated.
  • the server 3 posts a message of the robot 1 or the robot 2 in response to a comment posted by a user on a short text posting site.
  • the server 3 posts only the message of one of the robots, makes the message of each robot different, etc., and the same message for one comment May be controlled not to be posted.
  • posting of the message of each robot is not limited to the posting of the user, and may be performed when a predetermined condition is satisfied (for example, when a preset date and time is reached).
  • the user specifying unit 13 may specify the user who has spoken to the robot 1 from the position information of the robot 1. For example, when the position information of the robot 1 indicates the position of the user's room where the robot is, the user who has spoken to the robot 1 can be identified as the certain user.
  • the position information may be acquired using GPS (Global Positioning System). When the GPS reception sensitivity such as indoors is lowered, the position information may be corrected using the wireless communication function of the own device. Further, for example, the position of the own apparatus may be specified using an IMES (Indoor ME Messaging System) function or an image acquired by the camera 30.
  • IMES Indoor ME Messaging System
  • the robot 1 and the robot 2 may use the speech data acquired by the microphone 20 as the speech content information without specifying the speech content in the speech content specifying unit 12.
  • a keyword indicating the utterance content may be extracted on the server 3 side.
  • the user may be specified by performing voiceprint authentication of the voice data on the server 3 side, thereby eliminating the need for the robots 1 and 2 to transmit the user specifying information.
  • the server 3 may wait for a predetermined time after receiving the utterance information from any one of the robots, instead of performing the message determination process on a first-come-first-served basis.
  • the server 3 may perform message determination processing on the received utterance information according to a predetermined order or randomly.
  • the server 3 may perform message determination processing in the order in accordance with the priority set for the robot that is the transmission source of each utterance information, or the contents of various information included in the utterance information ( For example, the order in which message determination processing is performed, that is, the order in which messages with high priority are distributed may be determined according to the accuracy of the utterance content.
  • each robot outputs a different message.
  • a configuration may be adopted in which any robot does not output a message (no message is transmitted). This configuration can also prevent a plurality of robots from outputting the same message.
  • control blocks (particularly the utterance information creation unit 11 and the message determination unit 83) of the robot 1, the robot 2, the server 3, and the server 4 are realized by a logic circuit (hardware) formed in an integrated circuit (IC chip) or the like. Alternatively, it may be realized by software using a CPU (Central Processing Unit).
  • a logic circuit hardware
  • IC chip integrated circuit
  • CPU Central Processing Unit
  • the robot 1, the robot 2, and the server 3 include a CPU that executes instructions of a program that is software that realizes each function, and a ROM in which the program and various data are recorded so as to be readable by the computer (or CPU).
  • a computer or CPU
  • the objective of this invention is achieved when a computer (or CPU) reads the said program from the said recording medium and runs it.
  • a “non-temporary tangible medium” such as a tape, a disk, a card, a semiconductor memory, a programmable logic circuit, or the like can be used.
  • the program may be supplied to the computer via an arbitrary transmission medium (such as a communication network or a broadcast wave) that can transmit the program.
  • a transmission medium such as a communication network or a broadcast wave
  • the present invention can also be realized in the form of a data signal embedded in a carrier wave in which the program is embodied by electronic transmission.
  • the control devices (servers 3 and 4) output messages (response messages or spontaneous messages) to a plurality of output devices (robots 1 and 2) including the first output device and the second output device.
  • a determination unit for determining whether a message to be output to the first output device is the same as a message to be output to the second output device or a message output to the second output device (message determination) Unit 83) and control that does not cause the first output device to output a message different from that of the second output device or causes the first output device to output no message when it is determined that the determination unit is the same.
  • Unit messagessage determining unit 83).
  • the control device when the message to be output to the first output device is the same as the message to be output to the second output device or already output, the control device has the same message as the second output device. Can be prevented from being output. Therefore, the control device can prevent a plurality of output devices from outputting the same message.
  • a message generation unit (message determination unit 83) that receives the utterance content information to be generated and generates a message according to each utterance content information, and the determination unit is provided by each of the first output device and the second output device.
  • the message output to the first output device is the same as the message output to the second output device or the message output to the second output device. It may be determined that
  • the control device when the utterance content information received from each of the first output device and the second output device corresponds to the same utterance, the control device outputs so that these output devices do not output the same message.
  • the device can be controlled. Therefore, when a plurality of output devices detect the same utterance and respond (output a message), it is possible to prevent the same message from being output.
  • the determination unit speaks to the first output device and the second output device from each of the first output device and the second output device.
  • User specific information indicating a user is acquired, and when the utterance content indicated by the utterance content information matches and the user indicated by the user specification information matches, the utterance content information corresponds to the same utterance. May be determined.
  • the control device controls not to output the same message as the second output device. be able to. Therefore, the control device can prevent a plurality of output devices from making the same utterance with respect to the same utterance content of the same user.
  • the control device is the control device according to aspect 2 or 3, wherein the determination unit determines the first output device and the second output device for each of the first output device and the second output device.
  • the utterance content information is acquired when the utterance time information indicating the time when the user uttered is acquired, the utterance content indicated by the utterance content information matches, and the time difference indicated by the utterance time information is within a predetermined range. May correspond to the same utterance.
  • the above-mentioned “predetermined range” may be set within a range in which it is estimated that the time when the user speaks is the same in consideration of processing of each output device and a time lag of communication.
  • “the time when the user uttered” needs to substantially match the time when the user actually uttered, and does not need to be completely matched.
  • the time when the output device detects the user's utterance or the time when the control device receives the utterance content information may be set as the “time when the user uttered”.
  • the control device may acquire the utterance time information from the output device, and in the latter case, the reception time may be specified by a time measuring unit provided in the control device.
  • the control device when both the first output device and the second output device output a message in response to an utterance made at the same time (at the same timing), the control device has the same message. Can be controlled not to output. Therefore, the control device can prevent a plurality of output devices from making the same utterance with respect to the same utterance content of the same user.
  • a message output control system (message output control system 100 and message output control system 200) according to aspect 5 of the present invention includes the control device (servers 3 and 4) according to any one of aspects 1 to 4, and A plurality of output devices.
  • control device and the output device may be realized by a computer.
  • the control device or the output is performed by causing the computer to operate as each unit included in the control device or the output device.
  • the control device or the output device control program for realizing the apparatus by a computer and a computer-readable recording medium on which the control program is recorded also fall within the scope of the present invention.
  • the present invention can be suitably used for an output device that presents a message to a user, a control device that determines a message output by the output device, and the like.
  • 1 and 2 robots output device, first output device, second output device), 3 server (control device), 83 message determination unit (determination unit, control unit, message generation unit), 100 and 200 message output control system

Landscapes

  • Engineering & Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Acoustics & Sound (AREA)
  • Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Computational Linguistics (AREA)
  • General Health & Medical Sciences (AREA)
  • Robotics (AREA)
  • Mechanical Engineering (AREA)
  • Manipulator (AREA)
  • Toys (AREA)

Abstract

 複数のロボットがユーザに対し同じメッセージを出力することを防ぐ。ロボット(2)に出力させるメッセージが、ロボット(1)に出力させるメッセージまたは当該ロボット(1)に出力させたメッセージと同じとなるか判定し、同じとなると判定した場合に、ロボット(2)にロボット(1)とは異なるメッセージを出力させるか、またはロボット(2)にはメッセージを出力させないメッセージ決定部(83)を備える。

Description

制御装置およびメッセージ出力制御システム
 本発明は、出力装置がユーザに対して出力するメッセージを制御する制御装置およびメッセージ出力制御システムに関する。
 音声や文字等でメッセージを提示することにより、ユーザとコミュニケーションをとる対話型ロボットが従来から知られている。このような対話型ロボットは、ユーザの問いかけ等に応じて最適な応答メッセージを提示することが望ましい。そのため、ロボットの応答を最適化する種々の技術が開発されている。例えば、特許文献1には、対話型ロボットがユーザの発話音声を取得してコンピュータに送信し、コンピュータがユーザの発話音声に応じた応答文を対話型ロボットの応答文と決定する技術が開示されている。
日本国公開特許公報「特開2005-3747(公開日:2005年1月6日)」
 ユーザと複数のロボットとが対話する場合、上述した従来技術では、全てのロボットが同じメッセージをユーザに提示してしまう可能性がある。これは、各ロボットが取得した音声は同じユーザが発した同じ音声であり、その音声に応じた応答文を決定するコンピュータも同じであるからである。そして、ユーザの発話に対し、複数のロボットが同じ応答をすることは、当然望ましくない。なお、このような問題は、対話型ロボットに限られず、ユーザに対してメッセージを出力する出力装置全般に生じる問題点である。
 本発明は上記問題点を鑑みてなされたものであり、その目的は、複数の出力装置がユーザに対し同じメッセージを出力することを防ぐことが可能な制御装置およびメッセージ出力制御システムを実現することにある。
 上記の課題を解決するために、本発明の一態様に係る制御装置は、第1出力装置と第2出力装置とを含む複数の出力装置にメッセージを出力させる制御装置であって、上記第1出力装置に出力させるメッセージが、上記第2出力装置に出力させるメッセージまたは当該第2出力装置に出力させたメッセージと同じとなるか判定する判定部と、上記判定部が同じとなると判定した場合に、上記第1出力装置に上記第2出力装置とは異なるメッセージを出力させるか、または上記第1出力装置にはメッセージを出力させない制御部と、を備えることを特徴とする。
 本発明の一様態によれば、複数の出力装置がユーザに対し同じメッセージを出力することを防ぐことができるという効果を奏する。
本発明の実施形態1に係るロボットおよびサーバの要部構成を示すブロック図である。 本発明の実施形態1に係るメッセージ提示システムの概要を示す図である。 上記サーバがメッセージの決定に用いるメッセージテーブルのデータ構造を示す図である。 上記ロボットおよび上記サーバの情報の流れを時系列で示したシーケンス図である。 上記サーバの行うメッセージ決定処理の流れを示すフローチャートである。 本発明の実施形態2に係るメッセージ出力制御システムの概要を示す図である。
 〔実施形態1〕
 以下、本発明の第1の実施形態について、図1~5を参照して説明する。まず始めに、本実施形態に係るメッセージ出力制御システム100について図2を用いて説明する。図2は、本実施形態に係るメッセージ出力制御システム100の概要を示す図である。
 ≪システムの概要≫
 メッセージ出力制御システム100は、複数のロボットがそれぞれ、ユーザの発話した内容(発話内容)に適した応答メッセージをユーザに提示するためのシステムである。メッセージ出力制御システム100は図示の通り、ロボット(出力装置、第1出力装置、第2出力装置)1および2と、サーバ(制御装置)3とで構成される。なお、本実施形態ではロボットが2体である例を説明するが、ロボットは2体以上であればその数は問わない。以降の実施形態でも同様である。
 ロボット1およびロボット2は、ユーザと音声で対話する機能を備えた装置であり、ユーザの発話を検出し、該発話に対して音声または文字等で応答メッセージを出力(提示)する。また、サーバ3は、ロボット1およびロボット2の発言内容を制御する装置である。メッセージ出力制御システム100において、ロボット1およびロボット2は、ユーザの発話を検出すると、当該発話に対し自装置が返答すべき応答メッセージを決定するための情報であり、上記検出された発話に関わる情報である発話情報(出力情報)をサーバ3に送信する。
 サーバ3はロボット1またはロボット2から発話情報を受信すると、発話情報に応じた応答メッセージを決定し、発話情報の送信元のロボットに送信する。例えば、ユーザがロボット1に対して、「明日の天気教えて。」と発話した場合、ロボット1からサーバ3に、上記の発話に関する発話情報が送信される。この場合、サーバ3は、発話情報の示す発話内容「明日の天気教えて。」に応じて、ロボット1の応答メッセージをメッセージA(「晴れ時々曇りです」という台詞)と決定し、ロボット1にメッセージAを送信する。そして、ロボット1は受信したメッセージAを出力する。すなわち、ロボット1はユーザに対し「晴れ時々曇りです」と発言する。
 また、サーバ3は、複数のロボットから受信した発話情報の中にメッセージが同じになる発話情報が複数ある場合、例えば、ユーザの同じ発話を複数のロボットが検出した場合、これらのロボットが同じ応答メッセージを出力しないように、発話内容に応じた応答メッセージで、かつ他のロボットと異なる応答メッセージを送信元のロボットに返す。例えば、ユーザの発話がロボット1および2の両方で検出された場合、各ロボットからユーザの同じ発話についての発話情報が送信される。例えば、ロボット2が、上述したロボット1と共に、「明日の天気教えて。」というユーザの発話を検出した場合、ロボット1とほぼ同時に、ロボット2も発話情報をサーバ3に送信する場合がある。このような場合、サーバ3は、ロボット1にはメッセージAを出力させると共に、ロボット2には、ユーザの発話内容に応じた応答メッセージであり、かつメッセージAと異なるメッセージA’を出力させる。例えば、ロボット2は、「最高気温は24度です」のような、ロボット1と異なる発言をする。このように、メッセージ出力制御システム100では、複数のロボットが同じ発話を検出した場合、例えばユーザが複数のロボットに向けて(または複数のロボットの近くで)発話したような場合に、複数のロボットに、ユーザに対して異なる応答メッセージを返させることができる。
 ≪ロボットおよびサーバの要部構成≫
 次に、メッセージ出力制御システム100を構成するロボット1、ロボット2、およびサーバ3の内部構成について図1を用いて説明する。図1は、本実施形態に係るロボット1、ロボット2、およびサーバ3の要部構成を示すブロック図である。なお、本実施形態において、ロボット2の内部の要部構成はロボット1と同様である。そのため、ロボット2の構成については図示および説明を省略する。
 ≪ロボットの要部構成≫
 ロボット1は、装置制御部10と、マイク20と、カメラ30と、出力部40と、装置通信部50とを備えている。マイク20は、ユーザの発話を音声データに変換し取得するものである。マイク20は取得した音声信号を、後述する発話情報作成部11へと送信する。カメラ30は、ユーザを被写体として撮影するものである。カメラ30は撮影画像(ユーザの画像)を、後述する発話情報作成部11へと送信する。なお、後述するユーザ特定部13が画像データ以外のデータからユーザを特定する場合は、カメラ30は設けられなくてもよい。出力部40は、装置制御部10から受信したメッセージを出力するものである。出力部40としては、例えばスピーカ等を用いればよい。装置通信部50は、サーバ3とデータの送受信を行うためのものである。
 装置制御部10は、ロボット1を統括的に制御するものである。装置制御部10は、後述の発話情報作成部11にて生成した発話情報をサーバ3に送信する。また、装置制御部10は、サーバ3からロボット1が出力するメッセージを受信し、当該メッセージを出力部40に出力させる。装置制御部10は、より詳しくは、発話情報作成部11と、発話内容特定部12と、ユーザ特定部13とを含む。また、装置制御部10は、図示しないが計時部を備えている。計時部は、例えば装置制御部10のリアルタイムクロックで実現可能である。
 発話情報作成部11は、発話情報を作成するものである。なお、発話情報作成部11は、発話情報とともに、サーバ3に応答メッセージのリクエストを送信してもよい。発話情報作成部11は、発話情報に含める各種情報を生成するため、発話内容特定部12と、ユーザ特定部13とを含む。
 発話内容特定部12は、ユーザの発話内容を特定し、発話内容を示す情報(発話内容情報)を生成する。発話内容情報は、後述するサーバ3のメッセージ決定部83が、該発話内容情報が示す発話内容に応じた応答メッセージを特定可能であればよく、その形式はとくに限定されない。本実施形態では、発話内容特定部12は、発話内容情報としてユーザの発話音声の音声データから1つ以上のキーワードを特定および抽出する。なお、発話内容特定部12は、マイク20から受信した音声データ自体、または当該音声データからユーザの発話音声を抽出したものを、発話内容情報としてもよい。
 ユーザ特定部13は、発話したユーザを特定する。ユーザの特定方法は、特に限定しないが、例えば、カメラ30から撮影画像を取得し、画像解析により当該撮影画像に写っているユーザを特定すればよい。また、ユーザ特定部13は、マイク20が取得した音声データに含まれる声紋を分析することにより、ユーザを特定してもよい。このように、画像分析または声紋分析を行うことにより、発話したユーザを正確に(精度良く)特定することができる。
 発話情報作成部11は、発話情報作成時の時刻(発話時刻情報)と、発話内容特定部12が生成した発話内容情報と、ユーザ特定部13が特定したユーザを示す情報(ユーザ特定情報)とを含む発話情報を作成する。すなわち、発話情報作成部11は、「いつ(時刻)、誰が(ユーザ特定情報)、何を話した(発話内容)か」を示す情報を含む発話情報を作成する。そして、発話情報作成部11は作成した発話情報を、装置通信部50を介しサーバ3に送信する。
 ≪サーバの要部構成≫
 サーバ3は、サーバ通信部70と、サーバ制御部80と、一時記憶部82と、記憶部90とを備える。サーバ通信部70は、ロボット1およびロボット2とデータの送受信を行うためのものである。
 一時記憶部82は、発話情報を一時的に格納するものである。一時記憶部82は、メッセージ決定部83から発話情報と当該発話情報に応じて決定された応答メッセージとを受信し、上記発話情報と上記応答メッセージとを対応づけて所定の期間記憶しておく。なお、ここでいう「所定の期間」は、少なくとも、他のロボットから同じ発話に基づく発話情報がサーバ3に届き得る期間(すなわち、ロボット間の通信および処理のタイムラグ)以上の期間であることが望ましい。
 記憶部90は、後述するメッセージ決定部83がメッセージを決定する処理(メッセージ決定処理)の際に読み出されるメッセージテーブル91を格納する。以下、メッセージテーブルの構成について図3を用いて説明する。
 ≪メッセージテーブルのデータ構造≫
 図3は、メッセージテーブル91のデータ構成を示す図である。なお、図3においてメッセージテーブル91をテーブル形式のデータ構造にて示したことは一例であって、メッセージテーブル91のデータ構造はテーブル形式に限定されない。図3に例示するメッセージテーブル91は、「発話内容」列と、「メッセージ」列と、「優先度」列とを含み、「発話内容」列の情報に、「メッセージ」列の情報が対応付けられている。さらに、「メッセージ」列の情報に対し、「優先度」列の情報が対応付けられている。
 「発話内容」列は、ユーザの発話内容を示す情報を格納する。「発話内容」列の情報は、メッセージ決定部83が、受信した発話情報に含まれる発話内容情報と照合可能な形式で格納されている。例えば、発話内容情報がユーザの発話から抽出したキーワードである場合、「発話内容」列のデータとして、発話内容を示すキーワード(または複数のキーワードの組み合わせ)を示す情報を格納しておけばよい。また、発話内容情報が、発話の音声データそのものである場合、「発話内容」列のデータとして、上記音声データとマッチング可能な音声データを格納しておけばよい。
 「メッセージ」列は、ロボット1およびロボット2に出力させる応答メッセージを示す情報を格納する。なお、図示の通り、メッセージテーブル91は、1つの「発話内容」列の情報に対し、複数の応答メッセージが対応付けられている。
 「優先度」列は、「発話内容」列に示された1つの発話内容に対応する「メッセージ」列の複数の応答メッセージそれぞれの優先順位を示す情報を格納する。なお、当該優先順位の付し方は特に限定されない。なお、「優先度」列の情報は必須ではない。「優先度」列の情報を用いない場合、「発話内容」列に対応する各応答メッセージからランダムに1つの応答メッセージを選択してもよい。
 サーバ制御部80は、サーバ3を統括的に制御するものであり、発話情報受信部81と、メッセージ決定部(判定部、制御部、メッセージ生成部)83とを含む。発話情報受信部81は、ロボット1およびロボット2が送信した発話情報を、サーバ通信部70を介して取得する。発話情報受信部81は、取得した発話情報をメッセージ決定部83に送信する。
 メッセージ決定部83は応答メッセージを決定する。メッセージ決定部83は、発話情報受信部81から発話情報を受信すると、一時記憶部82に保存された他の発話情報(上記発話情報より前に受信した発話情報)および当該他の発話情報に対応付けられた応答メッセージを読み出す。なお、以降の説明では、発話情報受信部81が受信した発話情報、すなわちサーバ3がこれから応答メッセージを決定する発話情報を「対象発話情報」と称する。次にメッセージ決定部83は、他の発話情報の中に、対象発話情報と同じ発話に対応する発話情報を検出する。メッセージ決定部83は、上記検出を行うため、時刻判定部84と、発話内容判定部85と、ユーザ判定部86とを含む。
 時刻判定部84は、他の発話情報に含まれる時刻と、対象発話情報に含まれる時刻との差が所定範囲内であるか否かを判定する。なお、時刻判定部84は、上記判定の精度を上げるため、発話情報に含まれる時刻だけでなく、サーバ3が当該発話情報を受信した時刻を計時し、これを発話情報と共に記憶しておいてもよい。そして、対象発話情報の受信時刻と、他の発話情報の受信時刻との差が所定範囲内であるか否かを判定してもよい。これにより、例えばロボット1の計時する時刻とロボット2の計時する時刻がずれている場合等でも、精度よく時刻の比較を行うことができる。発話内容判定部85は、他の発話情報の発話内容情報と、対象発話情報の発話内容情報とが一致するか否かを判定する。ユーザ判定部86は、対象発話情報のユーザ特定情報と、他の発話情報のユーザ特定情報とが一致するか否かを判定する。
 メッセージ決定部83は、上記時刻判定部84、発話内容判定部85、およびユーザ判定部86の上記判定結果から、対象発話情報と時刻の差が所定範囲内であり、発話内容が一致し、かつユーザ特定情報が一致する他の発話情報を検出する。以下では、メッセージ決定部83が検出する上記他の発話情報を、検出対象発話情報と呼ぶ。ここで、検出対象発話情報は、複数のロボットが同じユーザの発話をほぼ同時に検出することにより作成された発話情報であると推定される。換言すると、上記検出対象発話情報は、対象発話情報と同一の発話に対応して作成された発話情報であると推定される。そして、メッセージ決定部83は、記憶部90のメッセージテーブル91を検索して、対象発話情報に対応する応答メッセージを特定する。
 応答メッセージの特定において、メッセージ決定部83は、対象発話情報の応答メッセージが、検出対象発話情報の応答メッセージと同じにならないようにする。具体的には、対象発話情報の発話内容情報に対応する応答メッセージのうち、検出対象発話情報の応答メッセージと異なり、かつ優先度が最も高い応答メッセージを特定する。なお、検出対象発話情報が存在しない場合には、対象発話情報の発話内容情報に対応する応答メッセージのうち最も優先度の高い応答メッセージ特定する。そして、メッセージ決定部83は、対象発話情報と特定した応答メッセージとを対応付けて一時記憶部82に記憶させる。
 なお、メッセージ決定部83は、対象発話情報の応答メッセージを決定した後、一時記憶部82を参照して応答メッセージが等しい他の発話情報を検出した場合に、決定した応答メッセージを他の応答メッセージに変更してもよい。この構成であっても、複数のロボットが同じ応答メッセージを出力することを防ぐことができる。なお、この構成においても、検出した他の発話情報と対象発話情報とで、時刻判定部84、発話内容判定部85、およびユーザ判定部86の判定処理を行い、両発話情報が一致した場合に、決定した応答メッセージを他の応答メッセージに変更してもよい。
 ≪装置間のデータ通信の流れ≫
 次に、ロボット1およびロボット2とサーバ3とのデータのやりとりについて、図4を用いて説明する。図4はロボット1およびロボット2と、サーバ3との間のデータ通信を時系列で示したシーケンス図である。なお、図4のS300およびS400に示すサブルーチンは、図5に示す一連のメッセージ決定処理(S302~S310)を示す。なお、以降の説明では、ロボット1の発話情報に含まれる時刻をT1、発話内容情報をQ1、ユーザ特定情報をU1とし、これらを含む発話情報を発話情報(T1、Q1、U1)と記載する。また、ロボット2の発話情報に含まれる時刻をT2、発話内容情報をQ2、ユーザ特定情報をU2とし、これらを含む発話情報を発話情報(T2、Q2、U2)と記載する。
 ユーザが発話すると、ロボット1およびロボット2はマイク20で発話を取得し発話情報作成部11へ送信する(S100およびS200)。また、このときカメラ30にて発話したユーザが撮影され、撮影画像が発話情報作成部11へ送信される。発話情報作成部11の発話内容特定部12は、音声データからユーザの発話に含まれるキーワードを特定する。また、ユーザ特定部13は撮影画像からユーザを特定する。そして、発話情報作成部11は、計時した時刻(発話情報作成時の時刻)と、発話内容情報と、ユーザ特定情報とを付加した発話情報を作成し(S102および202)、サーバ3に送信する。サーバ3はロボット1から発話情報(T1、Q1、U1)を受信すると、メッセージ決定処理を行う(S300)。ここで、ロボット1の発話情報は、図示の通りサーバ3が最初に受信した(他の発話情報がサーバ3の一時記憶部82に記憶されていない状態で受信した)発話情報であり、発話情報(T1、Q1、U1)が対象発話情報である。そのため、検出対象発話情報は存在せず、メッセージ決定部83は、Q1に対応する応答メッセージのうち、最も優先度の高い応答メッセージをロボット1に出力させる応答メッセージとして決定する。そして、サーバ3は、決定したメッセージAをロボット1に送信する。ロボット1はメッセージAを受信すると、これを出力部40から出力する(S104)。
 サーバ3は続いて、ロボット2から受信した発話情報(T2、Q2、U2)について、メッセージ決定処理を行う(S400)。上述の通り、ロボット1の発話情報(T1、Q1、U1)に対するメッセージ決定処理が先に行われているので、サーバ3の一時記憶部82には、他の発話情報(発話情報(T1、Q1、U1))が格納されており、ここでは発話情報(T2、Q2、U2)が対象発話情報となる。したがって、メッセージ決定部83は、時刻判定部84、発話内容判定部85、およびユーザ判定部86により、対象発話情報(T2、Q2、U2)と、他の発話情報(T1、Q1、U1)とが一致するか判定する。そして、一致すると判定した場合、発話情報(T1、Q1、U1)が検出対象発話情報となる。この場合、サーバ3は、発話情報(T1、Q1、U1)に対応する応答メッセージAとは異なるメッセージA’を決定し、ロボット2に送信する。一方、発話情報(T1、Q1、U1)が検出対象発話情報ではない場合、サーバ3は、応答メッセージAについて考慮することなくQ2に応じてメッセージA’を決定しロボット2に送信する。ロボット2はメッセージA’を受信すると、これを出力部40から出力する(S204)。
 ≪メッセージ決定処理≫
 次に、図5を用いて、メッセージ決定部83がメッセージを決定する処理(メッセージ決定処理)の流れの詳細を説明する。なお、ここでは一例として、図4のロボット2の応答メッセージを決定する際の、メッセージ決定処理(S400)について説明する。図5はメッセージ決定処理の流れを示すフローチャートである。なお、図5のS302~306の処理の順序は入れ替わっていてもよい。
 メッセージ決定部83は、発話情報(T2、Q2、U2)を受信すると、S302~306の処理を行うことにより、他の発話情報のうち対象発話情報(この例ではT2、Q2、U2)と同じ発話に対応する発話情報を検出する。具体的には、メッセージ決定部83に含まれる時刻判定部84は、対象発話情報(T2、Q2、U2)の時刻T2と、他の発話情報(この例ではT1、Q1、U1)の時刻T1との差が所定範囲内であるか否かを判定する(S302)。T2とT1との差が所定範囲内である場合(S302でYES)、次に、発話内容判定部85は、Q2とQ1とが一致するか否かを判定する(S304)。Q2とQ1とが一致する場合(S304でYES)、ユーザ判定部86は、U2とU1とが一致するか否かを判定する(S306)。U2とU1とが一致する場合(S306でYES)、メッセージ決定部83は、他の発話情報(T1、Q1、U1)を、対象発話情報(T2、Q2、U2)と同じ発話に基づく発話情報(検出対象発話情報)として検出する。
 そして、メッセージ決定部83は、Q2に応じたメッセージを決定する。具体的には、メッセージ決定部83はメッセージテーブル91の「発話内容」列と、Q2が示す発話内容とを照合することにより、Q2に対応する応答メッセージを特定する。なお、このとき、メッセージ決定部83は、検出対象発話情報(T1、Q1、U1)に対応する応答メッセージ、すなわちQ1に対応する応答メッセージと異なるメッセージのうち、最も優先度が高いメッセージを、対象発話情報の送信元のロボットが出力するメッセージとして決定する(S308)。
 一方、T2とT1との差が所定範囲内でない場合、Q2とQ1とが異なる場合、またはU2とU1とが異なる場合(S302~S306のいずれかでNO)は、メッセージ決定部83は、検出対象発話情報が存在しないと判定する。なお、他の発話情報が2以上記憶されている場合には、他の発話情報のそれぞれについて、S302~306の処理を行う。そして、メッセージ決定部83は、メッセージテーブル91を参照することにより、対象発話情報に対応する応答メッセージのうち、最も優先度の高い応答メッセージを、対象発話情報の送信元のロボットが出力する応答メッセージとして特定する(S310)。
 〔実施形態2〕
 本発明において、ロボットは、ユーザの発話に対応して応答メッセージを提示するのではなく、ユーザに向けて自発的にメッセージを出力してもよい。以下、本発明の第2の実施形態について図6を用いて説明する。以降、ロボットが自発的に出力(発言)するメッセージを「自発メッセージ」と称する。
 図6は、本実施形態に係るメッセージ出力制御システム200の概要を示す図である。メッセージ出力制御システム200はメッセージ出力制御システム100と同様に、複数のロボット(ロボット(出力装置)1およびロボット(出力装置)2)と、サーバ(制御装置)4とを含む。サーバ4の構成は、サーバ3(図1参照)と同様であるが、サーバ4は、予め設定された事象(条件)をトリガとしてロボット1および2に自発メッセージを出力させる点でサーバ3と異なる。なお、上記「予め設定された事象(条件)」については特に限定しない。以降の説明では、一例として、ロボット1およびロボット2が、ユーザの接近を検出する(ユーザが、ロボットの検出範囲に入る)と自発メッセージを出力させる例について説明する。
 ロボット1およびロボット2は、ユーザが自装置の検出範囲に入ると、サーバ4にその旨を通知する。サーバ4は、上記通知を受信すると、自発メッセージを決定して、該通知の送信元のロボットに送信する。例えばサーバ4はロボット1から上記通知を受信すると、ロボット1の自発メッセージを自発メッセージB(「遊ぼうよ」という台詞)と決定し、ロボット1に出力させる。
 ここで、ユーザが、ロボット1に近付いたときに、その近くにロボット2があった場合、ロボット2においても上記ユーザが検出され、その旨がサーバ4に通知される。サーバ4は、このような場合に、ロボット2がロボット1と同じ自発メッセージを出力しないように、ロボット2の自発メッセージを決定する。例えば、サーバ4は、自発メッセージBとは異なる自発メッセージB’(「お手伝いしようか?」という台詞)をロボット2の自発メッセージと決定し、ロボット2に出力させる。このように、メッセージ出力制御システム200では、複数のロボットが同じユーザに対し、同じタイミングで同じ自発メッセージを出力してしまうことを防止できる。
 〔実施形態3〕
 上記各実施形態では、音声発話によってユーザにメッセージを提示する例を示したが、メッセージの提示はユーザが該メッセージを認識できる態様で行われればよく、この例に限られない。例えば、電子掲示板等の電子媒体にユーザ向けのメッセージを登録することによってメッセージを提示してもよい。この場合、サーバ3または4は、ロボット1およびロボット2の代理として(ロボット1およびロボット2の名前で)メッセージを登録してもよい。
 また、上記実施形態では、サーバ3がユーザの音声発話に対する応答メッセージを生成する例を示したが、ユーザが電子媒体に書き込んだ(投稿した)情報に応じたメッセージを生成してもよい。例えば、短文投稿サイトにユーザが投稿したコメントに対し、ロボット1またはロボット2のメッセージをサーバ3が投稿する場合を考える。この場合、サーバ3は、両ロボットのメッセージが同じになると判定したときに、何れかのロボットのメッセージのみを投稿する、各ロボットのメッセージを異ならせる、等により、1つのコメントに対して同じメッセージが投稿されないように制御してもよい。なお、各ロボットのメッセージの投稿は、ユーザの投稿に限られず、所定の条件を満たしたとき(例えば予め設定された日時になったとき)に行ってもよい。
 〔変形例〕
 上記実施形態のメッセージ出力制御システム100および200では、1つのサーバ3または4にてメッセージの決定およびロボットの出力制御等の機能を果たしているが、これらの機能は個別のサーバで実現してもよい。例えば、ロボット1またはロボット2の出力する応答メッセージ(または自発メッセージ)を決定するメッセージ決定サーバと、当該メッセージ決定サーバが決定したメッセージをロボットに出力させる出力制御サーバとの組み合わせであっても、メッセージ出力制御システム100および200と同様の機能を実現できる。
 また、ユーザ特定部13は、ロボット1の位置情報から、該ロボット1に対して発話したユーザを特定してもよい。例えば、ロボット1の位置情報が、該ロボットがあるユーザの居室内の位置を示している場合には、ロボット1に対して発話したユーザを、上記あるユーザと特定することができる。なお、位置情報は、GPS(Global Positioning System)を用いて取得してもよい。また、室内等GPSの受信感度が落ちる場合は、自装置の持つ無線通信機能を用いて位置情報を補正してもよい。また、例えばIMES(Indoor MEssaging System)機能を用いて、またはカメラ30が取得する画像などから自装置の位置を特定してもよい。
 また、メッセージ出力制御システム100において、ロボット1およびロボット2は発話内容特定部12において発話内容を特定せずに、マイク20が取得した音声データ自体を発話内容情報としてもよい。この場合、サーバ3側で発話内容を示すキーワードの抽出を行えばよい。この場合、サーバ3側で音声データの声紋認証を行うことによりユーザを特定してもよく、これによりロボット1および2はユーザ特定情報を送信する必要がなくなる。
 また、サーバ3はロボット1およびロボット2から発話情報を受信した場合、先着順でメッセージ決定処理を行うのではなく、いずれかのロボットから発話情報を受信してから所定時間待機してもよい。そして、サーバ3は待機中に他のロボットから発話情報を受信した場合、予め定めた順番に従って、またはランダムに、上記受信した発話情報についてのメッセージ決定処理をそれぞれ行ってもよい。より具体的には、サーバ3は、それぞれの発話情報の送信元のロボットに設定された優先度に沿った順番でメッセージ決定処理を行ってもよいし、発話情報に含まれる各種情報の内容(例えば、発話内容の特定精度など)に応じて、メッセージ決定処理を行う順番、すなわち優先度の高いメッセージを振り分ける順番を定めてもよい。
 また、上記各実施形態では、各ロボットに異なるメッセージを出力させる例を示したが、何れかのロボットにはメッセージを出力させない(メッセージを送信しない)構成としてもよい。この構成によっても、複数のロボットが同じメッセージを出力しないようにすることができる。
 〔ソフトウェアによる実現例〕
 ロボット1、ロボット2、サーバ3、およびサーバ4の制御ブロック(特に発話情報作成部11およびメッセージ決定部83)は、集積回路(ICチップ)等に形成された論理回路(ハードウェア)によって実現してもよいし、CPU(Central Processing Unit)を用いてソフトウェアによって実現してもよい。
 後者の場合、ロボット1、ロボット2、およびサーバ3は、各機能を実現するソフトウェアであるプログラムの命令を実行するCPU、上記プログラムおよび各種データがコンピュータ(またはCPU)で読み取り可能に記録されたROM(Read Only Memory)または記憶装置(これらを「記録媒体」と称する)、上記プログラムを展開するRAM(Random Access Memory)などを備えている。そして、コンピュータ(またはCPU)が上記プログラムを上記記録媒体から読み取って実行することにより、本発明の目的が達成される。上記記録媒体としては、「一時的でない有形の媒体」、例えば、テープ、ディスク、カード、半導体メモリ、プログラマブルな論理回路などを用いることができる。また、上記プログラムは、該プログラムを伝送可能な任意の伝送媒体(通信ネットワークや放送波等)を介して上記コンピュータに供給されてもよい。なお、本発明は、上記プログラムが電子的な伝送によって具現化された、搬送波に埋め込まれたデータ信号の形態でも実現され得る。
 〔まとめ〕
 本発明の態様1に係る制御装置(サーバ3および4)は、第1出力装置と第2出力装置とを含む複数の出力装置(ロボット1および2)にメッセージ(応答メッセージまたは自発メッセージ)を出力させる制御装置であって、上記第1出力装置に出力させるメッセージが、上記第2出力装置に出力させるメッセージまたは当該第2出力装置に出力させたメッセージと同じとなるか判定する判定部(メッセージ決定部83)と、上記判定部が同じとなると判定した場合に、上記第1出力装置に上記第2出力装置とは異なるメッセージを出力させるか、または上記第1出力装置にはメッセージを出力させない制御部(メッセージ決定部83)と、を備えることを特徴とする。
 上記構成によると制御装置は、第1出力装置に出力させるメッセージが、第2出力装置に出力させる、またはすでに出力させたメッセージと同じとなる場合、第1出力装置が第2出力装置と同じメッセージを出力しないようにすることができる。したがって、制御装置は、複数の出力装置が同じメッセージを出力することを防止できる。
 本発明の態様2に係る制御装置は、上記態様1において、上記第1出力装置および第2出力装置のそれぞれから、上記第1出力装置および第2出力装置に対して発話したユーザの発話内容を示す発話内容情報を受信して、各発話内容情報に応じたメッセージを生成するメッセージ生成部(メッセージ決定部83)を備え、上記判定部は、上記第1出力装置および第2出力装置のそれぞれから受信した発話内容情報が同一の発話に対応している場合に、上記第1出力装置に出力させるメッセージが、上記第2出力装置に出力させるメッセージまたは当該第2出力装置に出力させたメッセージと同じとなると判定してもよい。
 上記構成によると、制御装置は、第1出力装置および第2出力装置のそれぞれから受信した発話内容情報が同一の発話に対応している場合、これらの出力装置が同じメッセージを出力しないように出力装置を制御することができる。したがって、複数の出力装置が同じ発話を検出し応答(メッセージを出力)する場合に、同じメッセージを出力することを防止できる。
 本発明の態様3に係る制御装置は、上記態様2において、上記判定部は、上記第1出力装置および第2出力装置のそれぞれから、上記第1出力装置および第2出力装置に対して発話したユーザを示すユーザ特定情報を取得するとともに、上記発話内容情報が示す発話内容が一致し、かつ上記ユーザ特定情報が示すユーザが一致する場合に、上記発話内容情報が同一の発話に対応していると判定してもよい。
 上記構成によると、制御装置は、第1出力装置が、第2出力装置と同じユーザの同じ発話内容に対してメッセージを出力する場合に、第2出力装置と同じメッセージを出力しないように制御することができる。したがって、制御装置は、複数の出力装置が同じユーザの同じ発話内容に対して同じ発言をすることを防止できる。
 本発明の態様4に係る制御装置は、上記態様2または3において、上記判定部は、上記第1出力装置および第2出力装置のそれぞれについて、上記第1出力装置および第2出力装置に対してユーザが発話した時刻を示す発話時刻情報を取得するとともに、上記発話内容情報が示す発話内容が一致し、かつ上記発話時刻情報が示す時刻の差が所定範囲内である場合に、上記発話内容情報が同一の発話に対応していると判定してもよい。
 ここで、上記「所定範囲」は、各出力装置の処理や通信のタイムラグを考慮した上で、ユーザが発話した時刻が同じであると推定される範囲内に設定すればよい。また、「ユーザが発話した時刻」は、ユーザが実際に発話した時刻と概ね一致していればよく、完全に一致している必要はない。例えば、出力装置がユーザの発話を検出した時刻、または制御装置が発話内容情報を受信した時刻を「ユーザが発話した時刻」としてもよい。前者の場合、制御装置は上記発話時刻情報を出力装置から取得すればよく、後者の場合、制御装置の備える計時部により受信時刻を特定すればよい。
 上記構成により、制御装置は、第1出力装置と第2出力装置との両方が、同じ時刻に(同じタイミングに)なされた発話に対してメッセージを出力する場合に、これらの出力装置が同じメッセージを出力しないように制御することができる。したがって、制御装置は、複数の出力装置が同じユーザの同じ発話内容に対して同じ発言をすることを防止できる。
 本発明の態様5に係るメッセージ出力制御システム(メッセージ出力制御システム100およびメッセージ出力制御システム200)は、上記態様1から4のいずれか1態様に記載の制御装置(サーバ3および4)と、上記複数の出力装置と、を含む。
 本発明の各態様に係る制御装置および出力装置は、コンピュータによって実現してもよく、この場合には、コンピュータを上記制御装置または上記出力装置が備える各部として動作させることにより上記制御装置または上記出力装置をコンピュータにて実現させる上記制御装置または上記出力装置の制御プログラム、およびそれを記録したコンピュータ読み取り可能な記録媒体も、本発明の範疇に入る。
 本発明は上述した各実施形態に限定されるものではなく、請求項に示した範囲で種々の変更が可能であり、異なる実施形態にそれぞれ開示された技術的手段を適宜組み合わせて得られる実施形態についても本発明の技術的範囲に含まれる。さらに、各実施形態にそれぞれ開示された技術的手段を組み合わせることにより、新しい技術的特徴を形成することができる。
 本発明は、ユーザにメッセージを提示する出力装置、当該出力装置の出力するメッセージを決定する制御装置等に好適に利用することができる。
 1および2 ロボット(出力装置、第1出力装置、第2出力装置)、3 サーバ(制御装置)、83 メッセージ決定部(判定部、制御部、メッセージ生成部)、100および200 メッセージ出力制御システム

Claims (5)

  1.  第1出力装置と第2出力装置とを含む複数の出力装置にメッセージを出力させる制御装置であって、
     上記第1出力装置に出力させるメッセージが、上記第2出力装置に出力させるメッセージまたは当該第2出力装置に出力させたメッセージと同じとなるか判定する判定部と、
     上記判定部が同じとなると判定した場合に、上記第1出力装置に上記第2出力装置とは異なるメッセージを出力させるか、または上記第1出力装置にはメッセージを出力させない制御部と、を備えることを特徴とする制御装置。
  2.  上記第1出力装置および第2出力装置のそれぞれから、上記第1出力装置および第2出力装置に対して発話したユーザの発話内容を示す発話内容情報を受信して、各発話内容情報に応じたメッセージを生成するメッセージ生成部を備え、
     上記判定部は、上記第1出力装置および第2出力装置のそれぞれから受信した発話内容情報が同一の発話に対応している場合に、上記第1出力装置に出力させるメッセージが、上記第2出力装置に出力させるメッセージまたは当該第2出力装置に出力させたメッセージと同じとなると判定することを特徴とする請求項1に記載の制御装置。
  3.  上記判定部は、上記第1出力装置および第2出力装置のそれぞれから、上記第1出力装置および第2出力装置に対して発話したユーザを示すユーザ特定情報を取得するとともに、上記発話内容情報が示す発話内容が一致し、かつ上記ユーザ特定情報が示すユーザが一致する場合に、上記発話内容情報が同一の発話に対応していると判定することを特徴とする請求項2に記載の制御装置。
  4.  上記判定部は、上記第1出力装置および第2出力装置のそれぞれについて、上記第1出力装置および第2出力装置に対してユーザが発話した時刻を示す発話時刻情報を取得するとともに、上記発話内容情報が示す発話内容が一致し、かつ上記発話時刻情報が示す時刻の差が所定範囲内である場合に、上記発話内容情報が同一の発話に対応していると判定することを特徴とする請求項2または3に記載の制御装置。
  5.  請求項1から4のいずれか1項に記載の制御装置と、
     上記複数の出力装置と、を含むことを特徴とするメッセージ出力制御システム。
PCT/JP2015/060973 2014-05-13 2015-04-08 制御装置およびメッセージ出力制御システム Ceased WO2015174172A1 (ja)

Priority Applications (3)

Application Number Priority Date Filing Date Title
CN201580021613.6A CN106233378B (zh) 2014-05-13 2015-04-08 控制装置和消息输出控制系统
JP2016519160A JP6276400B2 (ja) 2014-05-13 2015-04-08 制御装置およびメッセージ出力制御システム
US15/306,819 US10127907B2 (en) 2014-05-13 2015-04-08 Control device and message output control system

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2014099730 2014-05-13
JP2014-099730 2014-05-13

Publications (1)

Publication Number Publication Date
WO2015174172A1 true WO2015174172A1 (ja) 2015-11-19

Family

ID=54479716

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2015/060973 Ceased WO2015174172A1 (ja) 2014-05-13 2015-04-08 制御装置およびメッセージ出力制御システム

Country Status (4)

Country Link
US (1) US10127907B2 (ja)
JP (1) JP6276400B2 (ja)
CN (1) CN106233378B (ja)
WO (1) WO2015174172A1 (ja)

Cited By (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2017200076A1 (ja) * 2016-05-20 2017-11-23 日本電信電話株式会社 対話方法、対話システム、対話装置、およびプログラム
WO2017200072A1 (ja) * 2016-05-20 2017-11-23 日本電信電話株式会社 対話方法、対話システム、対話装置、およびプログラム
WO2017200078A1 (ja) * 2016-05-20 2017-11-23 日本電信電話株式会社 対話方法、対話システム、対話装置、およびプログラム
WO2017200077A1 (ja) * 2016-05-20 2017-11-23 日本電信電話株式会社 対話方法、対話システム、対話装置、及びプログラム
JP2018013545A (ja) * 2016-07-19 2018-01-25 トヨタ自動車株式会社 音声対話装置および発話制御方法
JP2018036397A (ja) * 2016-08-30 2018-03-08 シャープ株式会社 応答システムおよび機器
JP2019175432A (ja) * 2018-03-26 2019-10-10 カシオ計算機株式会社 対話制御装置、対話システム、対話制御方法及びプログラム
WO2019216053A1 (ja) * 2018-05-11 2019-11-14 株式会社Nttドコモ 対話装置
WO2020071255A1 (ja) * 2018-10-05 2020-04-09 株式会社Nttドコモ 情報提供装置
JP2021002062A (ja) * 2020-09-17 2021-01-07 シャープ株式会社 応答システム

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP6585733B2 (ja) * 2015-11-17 2019-10-02 株式会社ソニー・インタラクティブエンタテインメント 情報処理装置
JP2019200393A (ja) * 2018-05-18 2019-11-21 シャープ株式会社 判定装置、電子機器、応答システム、判定装置の制御方法、および制御プログラム
CN111191019B (zh) * 2019-12-31 2024-12-20 联想(北京)有限公司 一种信息处理方法、电子设备和信息处理系统
US12380883B2 (en) * 2021-12-02 2025-08-05 Lenovo (Singapore) Pte. Ltd Methods and devices for preventing a sound activated response

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2002287783A (ja) * 2001-03-23 2002-10-04 Ricoh Co Ltd 音声出力タイミング制御装置
JP2009065562A (ja) * 2007-09-07 2009-03-26 Konica Minolta Business Technologies Inc 音出力装置およびこれを含む画像形成装置
JP2009265278A (ja) * 2008-04-23 2009-11-12 Konica Minolta Business Technologies Inc 音声出力管理システムおよび音声出力装置

Family Cites Families (12)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5750200A (en) * 1980-09-09 1982-03-24 Mitsubishi Electric Corp Outdoor loudspeaker system
TWI236610B (en) * 2000-12-06 2005-07-21 Sony Corp Robotic creature device
JP2003204596A (ja) * 2002-01-04 2003-07-18 Matsushita Electric Ind Co Ltd 拡声放送システムおよび拡声放送装置
JP3958253B2 (ja) 2003-06-09 2007-08-15 株式会社シーエーアイメディア共同開発 対話システム
JP4216308B2 (ja) * 2005-11-10 2009-01-28 シャープ株式会社 通話装置および通話プログラム
KR101052963B1 (ko) * 2006-01-27 2011-07-29 교세라 가부시키가이샤 무선 통신 디바이스
CN101187990A (zh) * 2007-12-14 2008-05-28 华南理工大学 一种会话机器人系统
JP4764943B2 (ja) * 2009-12-29 2011-09-07 シャープ株式会社 動作制御装置、動作制御方法、ライセンス提供システム、動作制御プログラム、および記録媒体
JP5614143B2 (ja) * 2010-07-21 2014-10-29 セイコーエプソン株式会社 情報処理システム、印刷装置及び情報処理方法
JP5117599B1 (ja) * 2011-06-30 2013-01-16 株式会社東芝 制御端末およびネットワークシステム
KR101253200B1 (ko) * 2011-08-01 2013-04-10 엘지전자 주식회사 멀티미디어 디바이스와 그 제어 방법
CN103369151B (zh) * 2012-03-27 2016-08-17 联想(北京)有限公司 信息输出方法和装置

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2002287783A (ja) * 2001-03-23 2002-10-04 Ricoh Co Ltd 音声出力タイミング制御装置
JP2009065562A (ja) * 2007-09-07 2009-03-26 Konica Minolta Business Technologies Inc 音出力装置およびこれを含む画像形成装置
JP2009265278A (ja) * 2008-04-23 2009-11-12 Konica Minolta Business Technologies Inc 音声出力管理システムおよび音声出力装置

Cited By (20)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2017200072A1 (ja) * 2016-05-20 2017-11-23 日本電信電話株式会社 対話方法、対話システム、対話装置、およびプログラム
WO2017200078A1 (ja) * 2016-05-20 2017-11-23 日本電信電話株式会社 対話方法、対話システム、対話装置、およびプログラム
WO2017200077A1 (ja) * 2016-05-20 2017-11-23 日本電信電話株式会社 対話方法、対話システム、対話装置、及びプログラム
JPWO2017200077A1 (ja) * 2016-05-20 2018-12-13 日本電信電話株式会社 対話方法、対話システム、対話装置、及びプログラム
JPWO2017200076A1 (ja) * 2016-05-20 2018-12-13 日本電信電話株式会社 対話方法、対話システム、対話装置、およびプログラム
JPWO2017200072A1 (ja) * 2016-05-20 2019-03-14 日本電信電話株式会社 対話方法、対話システム、対話装置、およびプログラム
WO2017200076A1 (ja) * 2016-05-20 2017-11-23 日本電信電話株式会社 対話方法、対話システム、対話装置、およびプログラム
JP2018013545A (ja) * 2016-07-19 2018-01-25 トヨタ自動車株式会社 音声対話装置および発話制御方法
JP2018036397A (ja) * 2016-08-30 2018-03-08 シャープ株式会社 応答システムおよび機器
JP2023055910A (ja) * 2018-03-26 2023-04-18 カシオ計算機株式会社 ロボット、対話システム、情報処理方法及びプログラム
JP2019175432A (ja) * 2018-03-26 2019-10-10 カシオ計算機株式会社 対話制御装置、対話システム、対話制御方法及びプログラム
JP7416295B2 (ja) 2018-03-26 2024-01-17 カシオ計算機株式会社 ロボット、対話システム、情報処理方法及びプログラム
WO2019216053A1 (ja) * 2018-05-11 2019-11-14 株式会社Nttドコモ 対話装置
JP7112487B2 (ja) 2018-05-11 2022-08-03 株式会社Nttドコモ 対話装置
US11430440B2 (en) 2018-05-11 2022-08-30 Ntt Docomo, Inc. Dialog device
JPWO2019216053A1 (ja) * 2018-05-11 2021-01-07 株式会社Nttドコモ 対話装置
JPWO2020071255A1 (ja) * 2018-10-05 2021-09-02 株式会社Nttドコモ 情報提供装置
JP7146933B2 (ja) 2018-10-05 2022-10-04 株式会社Nttドコモ 情報提供装置
WO2020071255A1 (ja) * 2018-10-05 2020-04-09 株式会社Nttドコモ 情報提供装置
JP2021002062A (ja) * 2020-09-17 2021-01-07 シャープ株式会社 応答システム

Also Published As

Publication number Publication date
JPWO2015174172A1 (ja) 2017-04-20
CN106233378A (zh) 2016-12-14
US20170125017A1 (en) 2017-05-04
US10127907B2 (en) 2018-11-13
JP6276400B2 (ja) 2018-02-07
CN106233378B (zh) 2019-10-25

Similar Documents

Publication Publication Date Title
JP6276400B2 (ja) 制御装置およびメッセージ出力制御システム
US12609112B2 (en) Multi-user authentication on a device
JP6630765B2 (ja) 個別化されたホットワード検出モデル
US10991374B2 (en) Request-response procedure based voice control method, voice control device and computer readable storage medium
KR101752119B1 (ko) 다수의 디바이스에서의 핫워드 검출
US11188289B2 (en) Identification of preferred communication devices according to a preference rule dependent on a trigger phrase spoken within a selected time from other command data
US9064495B1 (en) Measurement of user perceived latency in a cloud based speech application
JP6497372B2 (ja) 音声対話装置および音声対話方法
US12367882B1 (en) Speaker disambiguation and transcription from multiple audio feeds
WO2019026617A1 (ja) 情報処理装置、及び情報処理方法
CN111755000B (zh) 语音识别装置、语音识别方法及记录介质
JP2019045831A (ja) 音声処理装置、方法およびプログラム
KR20230075386A (ko) 음성 신호 처리 방법 및 장치
CN106847273B (zh) 语音识别的唤醒词选择方法及装置
JP6775563B2 (ja) 人工知能機器の自動不良検出のための方法およびシステム
EP3833459B1 (en) Systems and devices for controlling network applications
JP2017211610A (ja) 出力制御装置、電子機器、出力制御装置の制御方法、および出力制御装置の制御プログラム
US10847158B2 (en) Multi-modality presentation and execution engine
JP6645779B2 (ja) 対話装置および対話プログラム
US10818298B2 (en) Audio processing
CN113889102B (zh) 指令接收方法、系统、电子设备、云端服务器和存储介质
WO2020110744A1 (ja) 情報処理装置、情報処理方法、およびプログラム
US10505879B2 (en) Communication support device, communication support method, and computer program product
JP2022033824A (ja) 連続発話推定装置、連続発話推定方法、およびプログラム
US20210312908A1 (en) Learning data generation device, learning data generation method and non-transitory computer readable recording medium

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15793219

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2016519160

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 15306819

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 15793219

Country of ref document: EP

Kind code of ref document: A1