WO2025220150A1 - 自動応答システム、自動応答方法およびプログラム - Google Patents

自動応答システム、自動応答方法およびプログラム

Info

Publication number
WO2025220150A1
WO2025220150A1 PCT/JP2024/015272 JP2024015272W WO2025220150A1 WO 2025220150 A1 WO2025220150 A1 WO 2025220150A1 JP 2024015272 W JP2024015272 W JP 2024015272W WO 2025220150 A1 WO2025220150 A1 WO 2025220150A1
Authority
WO
WIPO (PCT)
Prior art keywords
facility
answer
question
information
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/JP2024/015272
Other languages
English (en)
French (fr)
Inventor
悠介 太田
美里 内藤
友樹 渡邊
晋 飯野
和彦 山田
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Mitsubishi Electric Corp
Original Assignee
Mitsubishi Electric Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Mitsubishi Electric Corp filed Critical Mitsubishi Electric Corp
Priority to JP2024559956A priority Critical patent/JP7721018B1/ja
Priority to PCT/JP2024/015272 priority patent/WO2025220150A1/ja
Publication of WO2025220150A1 publication Critical patent/WO2025220150A1/ja
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types

Definitions

  • This disclosure relates to an automatic response system, an automatic response method, and a program that generate answers to questions.
  • Patent Document 1 discloses a route guidance system for use in places like shopping malls and event venues.
  • a display device shows the route, and the robot explains the route to the destination to the user using voice and body movements.
  • Patent Document 1 can guide users, on behalf of employees, on routes to a specified destination. However, users have a wide variety of needs and situations. For this reason, the technology described in Patent Document 1 may not be able to provide information that is appropriate for the user.
  • This disclosure has been made in light of the above, and aims to provide an automatic response system that can provide facility users with more appropriate responses.
  • the automatic response system disclosed herein comprises a question acquisition unit that acquires questions about facility use from facility users who are users of the facility; a situation acquisition unit that acquires additional information indicating the situation of the facility user who asked the question; an answer generation unit that generates an answer to the question using the question, the additional information, and facility data that is data about the facility; and an output unit that outputs the answer.
  • the automated response system disclosed herein has the advantage of being able to provide facility users with more appropriate responses.
  • FIG. 1 is a diagram illustrating a configuration example of an automatic response system according to a first embodiment.
  • FIG. 1 is a diagram showing an example of the configuration of a response generation unit according to a first embodiment;
  • FIG. 1 is a diagram for explaining an automatic response according to the first embodiment.
  • FIG. 1 is a diagram for explaining an automatic response according to the first embodiment.
  • 1 is a flowchart showing an example of an automatic response processing procedure in the automatic response system according to the first embodiment. 5.
  • FIG. 1 is a diagram showing an example of dynamic facility data according to the first embodiment;
  • FIG. 10 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of the first embodiment.
  • FIG. 10 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of the first embodiment.
  • FIG. 10 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of the first embodiment.
  • FIG. 10 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of the first embodiment.
  • FIG. 10 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of the first embodiment.
  • FIG. 10 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of the first embodiment.
  • FIG. 10 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of the first embodiment.
  • FIG. 10 is a diagram showing a specific example of a response
  • FIG. 10 is a diagram showing another specific example of a response that takes the situation into consideration in the automatic response system of the first embodiment.
  • FIG. 1 is a diagram showing an example of the configuration of a computer system that realizes each of the information processing units according to the first embodiment.
  • FIG. 10 is a diagram illustrating a configuration example of an automatic response system according to a second embodiment.
  • FIG. 10 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of the second embodiment.
  • FIG. 10 is a diagram showing a specific example of a response that takes into account the situation in the automatic response system of the second embodiment.
  • FIG. 10 is a diagram illustrating a configuration example of an automatic response system according to a third embodiment.
  • FIG. 1 is a diagram illustrating an example of the configuration of an automatic answering system according to an embodiment.
  • the automatic answering system 10 of the present embodiment accepts questions from facility users regarding facility use and outputs information related to facility guidance as a response to the questions.
  • the facility may be, for example, a train station, a hospital, a commercial facility, a hotel, a building, an event venue, a conference center, a theme park, a railway facility including multiple stations (including the interior of a railway car), a school, a business office, a factory, or the like, but is not limited thereto as long as it is available to multiple users.
  • the facility may be a facility available to an unspecified number of people, such as a train station, a facility available to specific people, such as a school or business office, or a facility available to both specific and unspecified people.
  • the automatic response system 10 includes a data management unit 11, a question acquisition unit 12, a situation acquisition unit 13, an answer generation unit 14, and an output unit 15.
  • the data management unit 11 and the answer generation unit 14 constitute an information processing unit 20, which is an information processing device.
  • the automatic response system 10 is installed, for example, within a facility, but at least some of the components that make up the automatic response system 10 may be installed outside the facility.
  • the data management unit 11 manages facility data, which is data related to facilities.
  • the facility data is data used by the answer generation unit 14 when generating answers, and may include not only data related to the facility itself, but also data about the surrounding area of the facility and data related to the facility outside the facility.
  • facility data includes, but is not limited to, at least one of the following: map information within the facility, facility status information related to the status of the facility, map information around the facility, facility information issued by at least one of the facility manager and facility employees, and weather information for the area including the facility.
  • the map information includes information that enables users to determine routes to potential locations within the facility that they will visit, such as stores and conference rooms.
  • the map information may include information represented by a graph in which links represent corridors within the facility and the potential locations represent nodes.
  • the map information may also include attributes of the corridors, such as whether the corridors include stairs, slopes, elevators, or escalators.
  • the map information within the facility may include image data showing predetermined routes for each potential location that the user will visit.
  • the map information may include image data for each corridor attribute, such as whether the corridors include slopes, elevators, or escalators.
  • Facility status information includes, but is not limited to, congestion information indicating the degree of congestion at the facility, information such as road closures within the facility, information indicating whether equipment within the facility is broken, and road conditions indicating the condition of surrounding roads.
  • Facility-generated information is information that facility managers or employees want to convey to users, such as "It is raining, so the facility is slippery" or "The venue for event X starting at 10:00 today has been changed from X to Y.”
  • Weather information includes, but is not limited to, at least one of the following: weather (sunny, rainy, strong winds, etc.), temperature, and humidity.
  • the facility data may include timetables for trains and other railway vehicles, and if there is a store within the facility, it may include at least one of the following: information about the products sold by the store, information about the food and drink served at the store, information indicating the store's main target demographic, information about the store's recommended products, and store information transmitted by store employees.
  • Store information is information that store employees want to convey to customers, such as "Today, Shop A is having a rainy day sale,” "Shop B is currently running a limited-time sale,” or "Restaurant K has a wide selection of cold drinks available that are perfect for hot days.”
  • Facility data may include static data that does not change over a short period of time, such as map information within the facility or information indicating the store's main target demographic.
  • Facility data may also include dynamic data that may change over a short period of time, such as facility status information, facility transmission information, and weather information.
  • Facility data may also include both static and dynamic data.
  • the data management unit 11 includes a data acquisition unit 111, a pre-processing unit 112, and a data storage unit 113.
  • the data acquisition unit 111 acquires facility data and outputs the acquired facility data to the pre-processing unit 112.
  • the data acquisition unit 111 may acquire static facility data by accepting data input from, for example, a facility manager, operator, or store employee, or by receiving the data from another device (not shown).
  • the data acquisition unit 111 may receive map information about the area around the facility from a device that manages the map information.
  • a terminal device that can be operated by a facility manager, operator, store employee, etc. may accept input of facility data from the facility manager, operator, store employee, etc., and the data acquisition unit 111 may receive the facility data from the terminal device.
  • the data acquisition unit 111 may acquire dynamic data from facility data by, for example, receiving sensor information acquired by a sensor not shown in FIG. 1 , such as a surveillance camera or microphone, and calculating the dynamic data based on the received sensor information.
  • the dynamic data may include data acquired by a sensor installed inside or around the facility.
  • the data acquisition unit 111 may calculate congestion information based on video that is sensor information acquired by a surveillance camera.
  • the data acquisition unit 111 may output sensor information to the preprocessing unit 112, and the preprocessing unit 112 may calculate the dynamic data from the sensor information.
  • another device not shown may calculate the dynamic data based on sensor information, and the data acquisition unit 111 may receive the dynamic data from the other device.
  • the data acquisition unit 111 may acquire facility transmission information and store transmission information as text data by accepting input from the sender, or may accept input of facility transmission information and store transmission information by voice.
  • information broadcast within the facility may be acquired as facility transmission information and store transmission information from a microphone used for in-facility broadcasting at the facility.
  • microphones may be installed near the location where in-facility broadcasting is made, and information broadcast within the facility may be acquired from the microphone as facility transmission information and store transmission information.
  • the preprocessing unit 112 performs preprocessing on the facility data received from the data acquisition unit 111.
  • the preprocessing unit 112 stores the preprocessed facility data in the data storage unit 113 as a database.
  • Preprocessing is, for example, a process of converting the data into a data format accessible by the answer generation unit 14.
  • the preprocessing unit 112 vectorizes the facility data and stores the vectorized facility data in the data storage unit 113.
  • Vectorization includes, for example, a process of converting words in natural language processing into vectors, which are distributed representations according to the meaning of the words.
  • the content of the preprocessing performed by the preprocessing unit 112 is not limited to the example described above, and may be determined depending on the processing method used by the answer generation unit 14. Depending on the processing method used by the answer generation unit 14, the preprocessing unit 112 may not be provided, and the facility data may be stored directly in the data storage unit 113.
  • the data storage unit 113 stores the facility data.
  • the question acquisition unit 12 acquires questions about facility use from facility users (hereinafter also referred to as facility users). More specifically, the question acquisition unit 12 accepts questions about the facility from facility users and outputs the accepted questions to the answer generation unit 14.
  • the question acquisition unit 12 may acquire the questions as voice data, or may acquire the questions as text data by accepting input through operation by the facility user. That is, the question acquisition unit 12 may be equipped with a microphone, or may be equipped with input means such as a touch panel that accepts input through user operation, or may be equipped with both a microphone and input means.
  • the question acquisition unit 12 may also accept image input. In this case, for example, the image input is accepted by capturing an image with a camera. For example, a facility user may present a photo of a product, store, etc.
  • the camera may be a camera used as the situation acquisition unit 13 described below.
  • the question acquisition unit 12 receives a question using an image, it outputs the image data together with the question to the answer generation unit 14.
  • the situation acquisition unit 13 acquires additional information indicating the situation of the facility user who asked the question and outputs the acquired additional information to the answer generation unit 14.
  • the situation acquisition unit 13 is equipped with a sensor such as a camera or an infrared sensor, and acquires the detection results of the sensor as additional information (auxiliary information).
  • the additional information includes video data of the questioner.
  • the situation acquisition unit 13 may acquire the additional information indicating the facility user's situation as text data by accepting input through operation by the facility user, or may acquire the additional information indicating the facility user's situation as audio data.
  • the situation acquisition unit 13 acquires the additional information as text data or audio data, for example, a question inquiring about the situation is presented to the facility user by the output unit 15, and the situation acquisition unit 13 acquires the answer to the question as additional information.
  • the question acquisition unit 12 may have the function of the situation acquisition unit 13.
  • the microphone used by the situation acquisition unit 13 to acquire audio data may be shared with the question acquisition unit 12.
  • the additional information is, for example, situation information indicating the situation of the facility user asking the question.
  • the additional information may be situation information indicating the situation of the facility user asking the question, as well as attribute information indicating the attributes of the facility user asking the question.
  • video data acquired by a camera can be used to understand both the situation and attributes of a facility user.
  • the situation acquisition unit 13 may present the facility user with a question asking about their attributes in addition to their situation via the output unit 15, and the situation acquisition unit 13 may acquire the answer to the question as additional information.
  • the additional information may include situation information and attribute information.
  • the situation information and attribute information may be acquired by different means, such as acquiring video data as situation information and acquiring attribute information as a response to a question asking about their attributes.
  • the method of acquiring the attributes of a facility user is not limited to acquiring them from video data acquired by a camera.
  • the attributes of an individual may be identified using an individual's identification results.
  • the answer generation unit 14 generates an answer regarding facility guidance using the question received from the question acquisition unit 12, the additional information received from the situation acquisition unit 13, and the facility data stored in the data storage unit 113, and outputs the generated answer to the output unit 15. Specifically, the answer generation unit 14 understands the question and generates an answer using the facility data, taking into account the situation indicated by the additional information.
  • the answer generation unit 14 is capable of generating an answer using data in multiple formats, including, for example, audio, text, and images, as input, and is capable of outputting answers in multiple formats. That is, for example, the answer generation unit 14 is capable of, but is not limited to, performing multimodal output in response to multimodal input such as audio data, text data, and video data.
  • performing multimodal output in response to multimodal input will also be referred to as multimodal processing. Details of the answer generation unit 14 will be described later.
  • the output unit 15 presents the answer to the facility user who asked the question by outputting the answer received from the answer generation unit 14.
  • the answer may be output as audio, text data, an image, video, or a combination of two or more of these. Therefore, the output unit 15 includes, for example, at least one of a speaker and a display device such as a display or a touch panel.
  • the output unit 15 outputs the answer according to the format of the answer received from the answer generation unit 14. For example, if the answer received from the answer generation unit 14 is audio data, the output unit 15 outputs the answer as audio. If the answer received from the answer generation unit 14 is at least one of video data and text data, the output unit 15 outputs the answer by display.
  • the output unit 15 includes a touch panel, the touch panel may be used as an input means for the question acquisition unit 12 and as an input means for the situation acquisition unit 13.
  • FIG. 2 is a diagram showing an example of the configuration of the answer generation unit 14 in this embodiment.
  • the answer generation unit 14 includes a situation recognition unit 141 and a generation unit 142.
  • the situation recognition unit 141 recognizes the situation based on the additional information received from the situation acquisition unit 13 and outputs the recognized situation to the generation unit 142.
  • the situation recognition unit 141 recognizes the facility user who asked the question. However, if the facility user is accompanied by someone, such as a family or couple, hereinafter, not only the facility user who actually asked the question but also the accompanying person of the facility user will be treated as the facility user who asked the question.
  • the situation recognized by the situation recognition unit 141 includes, but is not limited to, at least one of the following: what the questioner is wearing (such as the questioner's clothing or accessories worn by the questioner), whether the questioner is using an assistive device such as a wheelchair or cane (whether or not the assistive device is being used), the number of questioners, the questioner's emotions, the size of luggage the questioner is carrying, and the direction of movement of the questioner.
  • the situation recognition unit 141 uses the additional information to recognize the situation and attributes of the questioner, and outputs the recognized situation and attributes to the generation unit 142.
  • the attributes of the questioner may include, for example, at least one of gender, age, generation, presence or absence of a disability, and occupation.
  • the generation unit 142 generates an answer regarding facility guidance using the situation (or situation and attributes) received from the situation recognition unit 141, the question received from the question acquisition unit 12, and facility data stored in the data storage unit 113, and outputs the generated answer to the output unit 15.
  • the answer generation unit 14 generates answers that reflect the questioner's situation. For example, if a questioner using a wheelchair asks how to get to a destination, the answer generation unit 14 can generate an answer that shows a route that does not have stairs and uses ramps or elevators.
  • FIGS. 3 and 4 are diagrams illustrating the automatic response of this embodiment.
  • the automatic response system 10 is equipped with a camera as the situation acquisition unit 13, and a touch panel and speaker as the output unit 15.
  • a keyboard software keyboard
  • the automatic response system 10 is further equipped with a microphone as the question acquisition unit 12.
  • the answer generation unit 14 of the automatic response system 10 recognizes the situation that the questioner 30 is carrying large luggage from video data received from a camera, which is an example of a sensor used as the situation acquisition unit 13, and generates an answer using the recognized situation, the audio data of the question received from the question acquisition unit 12, and the facility data.
  • the situation recognized in the example shown in Figure 3 is not limited to “carrying large luggage,” but may also be “both hands are full” or “difficulty walking.”
  • the answer generation unit 14 since the questioner 30 is carrying large luggage, the answer generation unit 14 generates a route to the station platform, i.e., a route using an elevator, as shown at the bottom of Figure 3, based on the map information within the facility in the facility data, rather than a route using stairs or escalators. Note that there are no particular restrictions on the method of generating the route, and methods such as searching for the shortest route within a range that meets the conditions can be used.
  • the answer generation unit 14 generates a diagram (image) showing the route and text data explaining the diagram as an answer, and outputs the generated answer to the output unit 15, which then outputs the answer.
  • the answer is displayed as text and a diagram showing the route on the touch panel of the output unit 15, but the output method is not limited to this.
  • the answer may also be output as audio by a speaker of the output unit 15, or the answer may be output as audio by a speaker instead of text.
  • the automatic response system 10 is composed of a camera, which is an example of a sensor used as the situation acquisition unit 13, and a main body 40.
  • the main body 40 includes an information processing unit 20, a question acquisition unit 12, and an output unit 15, all of which are not shown in FIG. 4.
  • the sensor used as the situation acquisition unit 13 may be provided separately from the main body 40.
  • the automatic response system 10 may also include multiple sensors as the situation acquisition unit 13; for example, sensors may be provided both on the main body 40 and outside the main body 40.
  • the questioner 30 is using a wheelchair and asks the automatic response system 10, "Where is the exit?"
  • the microphone which is the question acquisition unit 12, receives the question.
  • the answer generation unit 14 generates an answer that includes a slope with no steps as a route to the exit, and the output unit 15 displays the answer.
  • the answer generation unit 14 assumes that the questioner will be leaving the facility, and if it is raining based on the facility data, the answer generation unit 14 may add a sentence that corresponds to the outside weather, such as "It's raining outside, so please be careful on your way home.”
  • the question is about an exit, but by including map information about the area around the facility in the facility data, for example, if a questioner asks, "Please tell me the way to the department store" when the facility is inside a train station, the automatic response system 10 may be able to provide guidance on the way from the station to the department store. In this way, the answer generation unit 14 may respond with information about the area around the facility as information related to facility guidance.
  • the facility data may include information indicating the train lines that stop at the station, the stations on those lines, the nearest stores and tourist attractions to the station, etc., so that a question such as, "Which train line should I take to get to VV?" can be answered by indicating which train line to take and at which station to get off.
  • the automatic response system 10 is realized as a single device, and in the example shown in Figure 4, it is realized by a camera, which is an example of a sensor used as the situation acquisition unit 13, and the main body 40, but the configuration of the device as hardware for the automatic response system 10 is not limited to these examples.
  • the information processing device which is the information processing unit 20 shown in Figure 1
  • the input/output device including the question acquisition unit 12, situation acquisition unit 13, and output unit 15 may be provided in separate locations.
  • the input/output device and the information processing device may each be provided with a transceiver unit for communication, and information may be exchanged via the transceiver unit.
  • Figures 3 and 4 are merely examples, and the situations considered by the automatic response system 10 and the answers presented by the automatic response system 10 are not limited to these examples.
  • the answer generation unit 14 may be, for example, generative AI (Artificial Intelligence).
  • the generative AI may be, for example, multimodal AI such as Gemini (registered trademark) or GPT (Generative Pre-trained Transformer)-4V, image generation AI such as Stable Diffusion, large-scale language model such as GPT-3, or speech generation AI such as Murf. AI.
  • multimodal AI is used for the answer generation unit 14, the answer generation unit 14 has the functions of a situation recognition unit 141 and a generation unit 142.
  • the answer generation unit 14 Before accepting a question from a facility user, such as when the automatic response system 10 is started, an instruction such as "Please consider the situation of the facility user who asked the question and generate video data, text data, and audio data as necessary to provide an answer" is input to the answer generation unit 14.
  • the answer generation unit 14 is also instructed to use facility data as appropriate.
  • the answer generation unit 14 takes into account the situation of the questioner 30, references the facility data, and generates an answer regarding facility guidance using at least one of video data, text data, and audio data.
  • the specific circumstances to be considered may be input to the answer generation unit 14, or data indicating a definition of the circumstances may be included in the facility data.
  • data indicating a definition of the circumstances may be included in the facility data.
  • an instruction such as "The circumstances of the facility user who asked the question include whether the questioner is using an assistive device such as a wheelchair or a cane, the number of people asking, the emotions of the questioner, the clothing of the questioner, the items worn by the questioner, the size of luggage the questioner is carrying, and the direction of movement of the questioner" may be input to the answer generation unit 14.
  • the format of the answer will be determined by the answer generation unit 14.
  • instructions indicating those rules can be input to the answer generation unit 14, or data indicating the rules can be included in the facility data.
  • typical examples, norms, guidelines, etc. of answers corresponding to questions can be stored in the data storage unit 113 as facility data or separately from the facility data, and the answer generation unit 14 can refer to those rules when creating an answer.
  • pre-learning can be performed to learn these rules.
  • facility data is used to generate answers, but part of the facility data may be used as is as part of the answer. For example, if the facility data includes images of products sold by the store, the answer generation unit 14 may include those images in the answer.
  • the answer generation unit 14 may input the situation of the questioner 30 to be considered and the question into the multimodal AI of the answer generation unit 14 in advance, check whether the desired answer is obtained, and if the desired answer is not obtained, instruct the multimodal AI to provide a correct answer corresponding to the situation and question.
  • the correct answer corresponding to the situation and question may also be included in the facility data. For example, as shown in FIG. 3, when a questioner 30 carrying large luggage asks about the route to their destination, if the answer generation unit 14 generates an answer that includes stairs, the answer generation unit 14 may be given an instruction such as "If the questioner has large luggage, use the elevator instead of the stairs," or the instruction may be included in the facility data.
  • the instruction may be broken down into information such as "It is difficult for the questioner to walk if they are carrying large luggage or have both hands full” and information such as "If it is difficult to walk, use the elevator instead of the stairs.”
  • instructions for generating appropriate answers for each situation and attribute may be given to the answer generation unit 14, or these instructions may be included in the facility data.
  • the question acquisition unit 12 may accept the question as text data, and the output unit 15 may output the answer as text data.
  • the question acquisition unit 12 may accept the question as audio, and the question acquisition unit 12 or the answer generation unit 14 may convert the audio data of the question into text data.
  • a large-scale language model may function as the generation unit 142, and a situation recognition model that recognizes the situation from video data may be used as the situation recognition unit 141.
  • the situation recognition model may be, for example, a trained model that has been trained by machine learning to infer the situation from video data.
  • An example of machine learning used to generate the trained model is supervised learning such as a neural network, but it may also be unsupervised learning, reinforcement learning, etc., and is not limited to these.
  • the situation recognition model may also be a model that recognizes attributes.
  • the situation recognition model may be a combination of multiple models, such as a combination of a general emotion recognition model that recognizes emotions from video data and an image recognition model that recognizes the presence or absence of assistive devices, the size of luggage, clothing, etc. from video data.
  • the generation unit 142 generates an answer as text data using, for example, text data indicating the situation (or the situation and attributes) received from the situation recognition unit 141 and the question.
  • the generation unit 142 may further include a voice generation AI that converts the answer generated by the large-scale language model into voice data, or the generation unit 142 may further include an image generation AI that converts the answer generated by the large-scale language model into image data.
  • output may be performed in all supported output formats, or rules for selecting the output format may be determined in advance and the output format may be determined in accordance with the rules.
  • a rule may be established that associates the situation (or situation and attributes) of the asker 30 with the output format, or a rule may be established that outputs in the same format as the format of the question asked by the asker 30.
  • a rule may be established that associates the type of question content (for example, a question about the route to a destination, a question asking about recommended shops, etc.) with the output format.
  • the rules for determining the output format are not limited to these examples.
  • the facility data is vectorized by the preprocessing unit 112 and stored in the data storage unit 113, allowing the answer generation unit 14 to use the facility data as is. Therefore, the facility data can be reflected in the answers of the answer generation unit 14 without the need for additional learning.
  • a general automatic conversation program such as a chatbot may be used for the generation unit 142.
  • a situation recognition model may be used as the situation recognition unit 141, just as when a large-scale language model is used for the generation unit 142, and text data indicating the situation (or situation and attributes) recognized by the situation recognition unit 141 may be input to the generation unit 142.
  • the automatic conversation program may be a rule-based (scenario-based) program in which response rules are defined in advance, or a machine learning program.
  • a rule-based automatic conversation program for example, answers corresponding to the content of questions for each situation (or situation and attributes) are determined in advance and set in the automatic conversation program.
  • the machine learning in a machine learning automatic conversation program may be supervised learning, unsupervised learning, or reinforcement learning.
  • supervised learning is used in the automatic conversation program, a trained model is generated using multiple training datasets containing questions for each situation (or situation and attributes) and corresponding answers, which are correct answer data, and the generation unit 142 generates an answer by inputting the situation (or situation and attributes) and question corresponding to the questioner 30 into the trained model.
  • the answer generated by the automatic conversation program may be converted into voice data by a voice generation AI or into image data by an image generation AI.
  • output may be performed in all supported output formats, or rules for selecting the output format may be defined in advance and the output format may be determined in accordance with the rules.
  • the questioner 30 may not only be an unspecified person visiting the facility, but may also be a specific person, such as an employee, manager, or student if the facility is a school. That is, facility users may include specific people that have been specified in advance. For example, if the facility is intended for specific people, only specific people may be considered as questioners 30. In this case, for example, a facial photograph of the specific person may be included in the facility data, the situation acquisition unit 13 may acquire camera video data as additional information, and the answer generation unit 14 may identify the questioner 30 based on the facial photograph and generate an answer based on the identification result.
  • This video data is an example of personal authentication information for identifying an individual.
  • personal identification information For specific people who may use the facility, their identification information (hereinafter also referred to as personal identification information) may be associated with a facial photograph and stored in the data storage unit 113 as facility data. Furthermore, for each piece of personal identification information, individual information about the person corresponding to that personal identification information may be stored in the data storage unit 113 as facility data. Individual information is information relating to a specific person's individual use of the facility, such as, but not limited to, one or more of the following: schedule, accessible locations, and the scope of information that can be obtained (authority to view information).
  • the facility is a school
  • information indicating the classes each student is taking is included in the individual information of the facility data
  • class location information indicating the time and location (classroom, etc.) of each class is included in the facility data.
  • the answer generation unit 14 can determine the next class the student will take based on the individual information of the student, and the time and location of the class using the class location information.
  • the answer generation unit 14 generates an answer that reflects the information it has determined, such as "The next class is class Y, which will be held in classroom X in building 3." At this time, the answer may also include an image showing the route from the current location to classroom X in building 3. When generating an answer, for example, the situation is taken into consideration, as described above.
  • the facility data may include individual information not only for students but also for school staff.
  • the individual information may include information indicating the schedule of classes taught by staff members, the time and location of the staff members' exam supervision, etc.
  • a response such as, "The next exam supervision will be in classroom Z in building 2, starting at 11:00" is generated.
  • the content of the response described above is merely illustrative, and the content of the response is not limited to the above example.
  • a passcode assigned to each individual may be included in the facility data, and the automatic response system 10 may accept input of the passcode when a question is asked, thereby identifying the individual.
  • the automatic response system 10 may be equipped with a device that performs biometric authentication other than facial recognition, and the device may be used to identify the individual.
  • the personal authentication information may be a passcode, biometric authentication information, etc.
  • attribute-related information is information related to facility use for each attribute, and includes schedules, accessible locations, the range of information that can be obtained (information viewing authority), and contact information.
  • the attributes may include, for example, status (student, professor, assistant professor, teacher, office worker, etc.) and affiliation (faculty, department, selection).
  • the attributes may include job title, affiliation (department, faculty, division), qualifications held, etc.
  • the attributes may include at least one of status, affiliation, and job title.
  • the response generation unit 14 may identify individuals using a method similar to that for identifying individuals described above, understand their attributes based on the individual information, and generate a response according to their attributes.
  • the answer generation unit 14 identifies the department to which the student belongs as an attribute of the student (questioner 30). The answer generation unit 14 also identifies the class corresponding to the identified attribute based on the attribute-related information, and generates an answer indicating the location of the class based on the class location information. For example, if accessible locations are specified for each attribute, when the questioner 30 asks how to get to a destination, the unit 14 obtains a route that travels within the range of accessible locations, and generates the obtained route as the answer.
  • the unit 14 generates an answer for the questioner 30 using information within the range that is permitted to be viewed. For example, if the attribute-related information includes contact information for each attribute, the unit 14 may generate an answer by adding the contact information to a direct answer corresponding to the question, regardless of the content of the question.
  • the contact information may include a notice of a change in class.
  • the questioner 30 may be both an unspecified person, such as a customer of the facility or a visitor to the university, and a specific person, such as an employee or student of the facility.
  • the answer generation unit 14 may change the content of the answer depending on whether the questioner 30 is an unspecified person or a specific person.
  • the attributes of the questioner 30 may include, for example, whether the questioner is a specific person, and rules for generating answers for each attribute, such as accessible locations, the range of information that can be obtained (information viewing authority), and contact information, may be defined in the attribute-related information, as in the above example.
  • the attribute of whether the questioner is a specific person is determined in the same way as when identifying the individual specific person described above, and is achieved, for example, by including facial information or the like for the specific person in the facility data in advance.
  • the attribute-related information may include information that distinguishes between facility data used to generate answers for specific people and facility data used to generate answers for unspecified people.
  • attributes such as affiliation and position may be further included as attributes, as in the above example, and answers may be generated according to these attributes.
  • the answer generation unit 14 will generate, as an answer, information indicating the warehouse where product P is stored based on the facility data if the questioner 30 is an employee, and will generate an answer indicating that the question cannot be answered if the questioner 30 is not an employee. At this time, for example, the answer generation unit 14 will further generate an answer depending on the situation, as described above.
  • the answer generation unit 14 may cause the output unit 15 to output information to the questioner 30 asking for the missing information. That is, for example, if the answer generation unit 14 of the automatic response system 10 cannot recognize the question or if there is insufficient information to generate an answer, the answer generation unit 14 may generate an answer to ask the questioner 30 to repeat the question.
  • the answer generation unit 14 may output at least one of text data and audio data such as "Please repeat the question again" to the output unit 15, and the output unit 15 may output at least one of text and audio.
  • the answer generation unit 14 may output to the output unit 15 at least one of text data and voice data for identifying the platform, such as "Is this the K line platform or the L line platform?", and the output unit 15 may output at least one of text and voice.
  • the answer generation unit 14 may output at least one of text data and voice data for identifying the number of people, such as "How many people are there?", and the output unit 15 may output at least one of text and voice.
  • FIG. 5 is a flowchart showing an example of the automatic response processing procedure in the automatic response system 10 of this embodiment.
  • the automatic response system 10 acquires facility data (step S1).
  • the data acquisition unit 111 acquires the facility data and outputs it to the preprocessing unit 112.
  • the facility data may be static data as described above, dynamic data, or both.
  • the automatic response system 10 stores the facility data (step S2).
  • the preprocessing unit 112 performs preprocessing on the facility data, and stores the preprocessed facility data in the data storage unit 113.
  • Preprocessing is, for example, vectorization as described above, but is not limited to this, and may involve conversion to a data format that can be accessed by the answer generation unit 14, or a data format that can be accessed quickly by the answer generation unit 14. Also, as described above, preprocessing does not necessarily have to be performed.
  • the automatic response system 10 acquires the question content (step S3).
  • the question acquisition unit 12 acquires the question from the questioner 30 as voice or text data, thereby acquiring the question content, and outputs the acquired question to the answer generation unit 14.
  • the automatic response system 10 acquires the situation (step S4).
  • the situation acquisition unit 13 acquires additional information indicating the situation of the questioner 30 and outputs the additional information to the answer generation unit 14.
  • the situation acquisition unit 13 may acquire sensor information acquired by a sensor such as a camera as the additional information, or may acquire the additional information through input by voice or text data from the questioner 30.
  • the automatic response system 10 generates a response (step S5).
  • the response generation unit 14 grasps the situation (the situation of the questioner 30) from the additional information, generates a response to the question based on the situation, and outputs the generated response to the output unit 15.
  • the automatic response system 10 may also grasp attributes (the attributes of the questioner 30) from the additional information, and generate a response to the question based on the situation and attributes.
  • the automatic response system 10 outputs the response (step S6) and ends the automatic response process. More specifically, in step S6, the output unit 15 outputs the response.
  • the output unit 15 outputs the response by performing at least one of audio output and display, depending on the format of the response generated by the response generation unit 14 (whether the response is audio data, text data, or image data).
  • FIG 6 is a flowchart showing an example of the answer generation processing procedure in the answer generation unit 14 shown in step S5 of Figure 5.
  • the answer generation unit 14 recognizes the situation and the question (step S11).
  • the situation recognition unit 141 recognizes the situation based on the accompanying information and outputs it to the generation unit 142.
  • the generation unit 142 recognizes the situation acquired from the situation recognition unit 141 and the question received from the question acquisition unit 12 by converting them into a format required for answer generation processing.
  • the generation unit 142 understands the meaning of the situation and the question.
  • the answer generation unit 14 will understand that "you are being asked about the route to the platform,” and if the recognized situation is “you are carrying large luggage,” it will understand that “it would be better to take a route that avoids stairs as much as possible.”
  • the answer generation unit 14 references the facility data (step S12).
  • the generation unit 142 reads out the facility data required to answer the question from the data storage unit 113. For example, in the example above where the question is "Where is the platform?" and the situation is "I have large luggage,” map information within the station is read out as facility data.
  • the answer generation unit 14 generates an answer (step S13) and ends the answer generation process. More specifically, in step S13, the generation unit 142 uses the referenced facility data to generate an answer to the question that takes the situation into consideration, and outputs the generated answer to the output unit 15. For example, in the example where the question is "Where is the platform?" and the situation is "I have large luggage," the unit uses map information within the station to find a route to the platform that uses an elevator and does not require stairs, and generates a diagram showing the route obtained through the search, along with an explanation of the diagram.
  • the automatic response system 10 of this embodiment can generate an answer that takes into account the situation of the questioner 30. This allows the automatic response system 10 to provide a more appropriate answer to the facility user than when simply answering a specific question from the facility user.
  • the automatic response system 10 may also generate an answer based on attributes in addition to the situation, and in this case, it is possible to provide the facility user with a more appropriate answer that suits the attributes.
  • dynamic data dynamic facility data
  • FIG. 7 is a diagram showing an example of dynamic facility data in this embodiment.
  • the facility is a railway facility (including the inside of a railway vehicle) including a station premises or multiple stations.
  • a questioner 30 asks, "Which car is empty?" In order to generate an answer to this question, it is necessary to know the degree of congestion in each car.
  • the data acquisition unit 111 acquires, as dynamic facility data, video data acquired by cameras 50 installed in each railway vehicle.
  • a timetable indicating when each railway vehicle will arrive at the station is also stored as facility data in the data storage unit 113.
  • the dynamic facility data is updated to the latest data when new data is acquired, for example, but is not limited to this; past dynamic facility data may also be stored in the data storage unit 113.
  • the preprocessing unit 112 may store video data of the interior of the vehicles in the data storage unit 113, or may determine the degree of congestion in each vehicle based on the video data and store the determined result in the data storage unit 113 as facility data.
  • the degree of congestion may be indicated, for example, in two stages: whether or not there is a mixture, or in three stages: crowded, normal, and empty, or by a congestion rate, or in some other way.
  • the answer generation unit 14 recognizes the degree of congestion based on the video data and uses it to generate an answer.
  • the answer generation unit 14 identifies the next arriving train car based on the timetable, and can identify which cars are empty by using the degree of congestion calculated from video data captured inside each car of that train car. This makes it possible to generate an appropriate answer that takes into account the actual situation, such as "Car 2 is empty,” in response to the question, "Which car is empty?"
  • FIG. 7 The example shown in FIG.
  • the automatic response system 10 is installed on a station platform, and the answer generation unit 14 generates an image showing the position of car 2 and the current location as an answer, and the output unit 15 displays the generated image along with the text answer, "Car 2 is empty.”
  • the situation (or the situation and attributes) may also be taken into consideration when generating this answer. For example, in the example shown in Figure 7, if the asker 30 does not have large luggage, a response may be generated to guide them to the emptiest available vehicle, and if the asker 30 has large luggage, a response may be generated to guide them to the nearest available vehicle (even if it is not the emptiest vehicle).
  • the dynamic facility data is not limited to the example shown in Figure 7, but may be information indicating the degree of congestion in restrooms within the facility, information indicating the degree of congestion in a store, or, as described above, facility-originated information, store-originated information, etc.
  • the situation of the questioner 30 may include, for example, the number of questioners 30.
  • the automatic response system 10 will select a quiet restaurant based on the facility data, generate a response indicating that the selected restaurant is recommended, and output the response.
  • the automatic response system 10 will select a restaurant that is suitable for families, such as a family restaurant, generate a response indicating that the selected restaurant is recommended, and output the response.
  • Figures 8 to 14 are diagrams showing specific examples of responses that take the situation into consideration in the automatic response system 10 of this embodiment.
  • a middle-aged or elderly woman wearing elegant clothing and accessories is asking the question, "Do you know of any clothing stores?"
  • the automatic response system 10 recognizes that the situation of the questioner 30 is that she is wearing accessories and elegant clothing, and based on this recognition result, the question, and information in the facility data indicating the products sold by each store and the target audience of each store, it determines that Shop A and Shop B are recommended clothing stores for middle-aged or elderly people.
  • the automatic response system 10 then generates information directing the user to Shop A and Shop B as a response to the question and displays it on the output unit 15. In this way, a response that takes the situation of the questioner 30 into consideration is generated.
  • the automatic response system 10 generates an answer including basis information indicating the basis for the answer based on the situation used to generate the answer.
  • the automatic response system 10 generates basis information of "mother generation” as the basis for deciding to guide the customer to Shop A and Shop B, based on the situation that the customer is a middle-aged or elderly woman wearing elegant clothing and accessories.
  • the automatic response system 10 not only the suggested store names "Shop A” and “Shop B” but also the text "How about a fashionable clothing store popular with the mother generation?" are displayed.
  • touching the "Shop A” and "Shop B" parts on the display screen shown in FIG. 8 may display detailed information about each store. Note that the specific content of the sentence (text) in the answer is not limited to the example shown in FIG. 8.
  • a young woman in a uniform asks, "Do you know of any clothing stores?"
  • the automatic response system 10 recognizes that the questioner 30 is in a situation where she is a young woman in a uniform, and based on this recognition result, the question, and information in the facility data indicating the products sold by each store and the target audience of each store, it determines that Shop C and Shop D are recommended clothing stores for female high school students (high school girls).
  • the automatic response system 10 then generates information directing users to Shop C and Shop D as an answer to the question, and displays this on the output unit 15.
  • the basis information is also displayed. Specifically, the text "high school girl" is generated as the basis information, and an answer including the basis information is displayed.
  • a middle-aged man wearing a classic hat and classic clothing asks, "Do you know of any clothing stores?"
  • the automatic response system 10 recognizes that the situation of the questioner 30 is that of a middle-aged man wearing a classic hat and classic clothing. Based on this recognition result, the question, and the information in the facility data indicating the products sold by each store and the target audience of each store, the automatic response system 10 determines that an e-shop is a recommended clothing store for middle-aged men who like classic fashion. The automatic response system 10 then generates information providing directions to the e-shop as a response to the question and displays it on the output unit 15. In the example shown in FIG. 10, similar to the examples shown in FIGS. 8 and 9, the basis information is also displayed.
  • the text "dandy" is generated as the basis information, and an answer including the basis information is displayed.
  • the gender of the questioner 30 is also taken into consideration, but the gender of the questioner 30 is also an attribute, and these examples can be said to generate an answer that takes into consideration both the situation, such as clothing, and the attribute.
  • FIGS. 11 and 12 show examples in which the facial expression or emotion of the questioner 30 is taken into consideration as the situation in which the questioner 30 finds himself. In both of the examples shown in FIG. 11 and FIG. 12, the questioner 30 is asking where his/her home is.
  • the questioner 30 asks "Where is the platform?" in a calm and normal manner.
  • the automatic answering system 10 recognizes that the questioner 30 is normal, i.e., calm, from the questioner's facial expression, movements, tone of voice, etc., and generates a response while engaging in a dialogue with the questioner 30.
  • the automatic answering system 10 asks the questioner 30, "Where do you want to take the train?”, and the questioner 30 replies, "I want to go to XX," and based on the response from the questioner 30, the automatic answering system 10 generates and outputs the response, "If you are going to XX, please proceed to platform 3."
  • the questioner 30 appears flustered and asks, "Where is the platform?"
  • the automated response system 10 recognizes that the questioner 30 is in a panic and is in a state of panic based on the questioner's facial expression, movements, tone of voice, etc., and outputs a response including multiple types of information at once without engaging in a dialogue with the questioner 30.
  • the automated response system 10 simultaneously displays the information "Platform 3 for direction XX,” “Platform 4 for direction YY,” and “Platform 6 for direction ZZ.”
  • the content of the response and the method of interaction with the questioner 30 may be determined depending on the facial expression or emotions of the questioner 30.
  • FIG. 13 and 14 show examples in which the facial expression or emotion of the questioner 30 and the number of people are taken into consideration as the situation of the questioner 30.
  • the automatic answering system 10 recognizes that the questioner 30 is in a high-energy state (in a good mood and excited state) based on the questioner's facial expression, movements, tone of voice, etc., and also recognizes that the questioner 30 is part of a group, and generates a response that provides directions to lively establishments "Bar A" and "Bar B.”
  • the response also includes evidence information
  • the output unit 15 also displays information such as, "For all of you who are excited, how about a lively bar that offers all-you-can-drink?"
  • the questioner 30 is a man and woman holding hands and asks, "Can you recommend a bar?"
  • the automatic response system 10 recognizes that the questioner 30 is a male-female couple based on the questioner's facial expression, movements, tone of voice, etc., and generates a response recommending the quiet bars "Bar C" and "Bar D.”
  • the response also includes evidence information, just like the examples shown in Figures 8, 9, and 10, and the information "How about a quiet bar for a couple?" is also presented to the output unit 15.
  • the automatic response system 10 may generateSEA-friendly answers that do not distinguish between the two genders, or answers that take both genders into consideration. For example, when providing restroom guidance, an answer that includes both genders may be generated by answering the locations of men's and women's restrooms regardless of the user's gender.
  • rules for generating strig-friendly answers may be stored in the data storage unit 113 as facility data or separately from the facility data, and the answer generation unit 14 may refer to these rules when creating answers.
  • ethical standards to be reflected in answers not limited to strig considerations, may be defined as rules, and these rules may be stored in the data storage unit 113 as facility data or separately from the facility data, and the answer generation unit 14 may refer to these rules when creating answers.
  • a generation AI is used in the answer generation unit 14, these rules may be specified to the generation AI in advance.
  • FIG. 15 is a diagram showing an example configuration of a computer system that realizes each of the information processing units 20 of this embodiment. As shown in Figure 15, this computer system includes a control unit 101, an input unit 102, a memory unit 103, a display unit 104, a communication unit 105, and an output unit 106, which are connected via a system bus 107.
  • the control unit 101 is, for example, a processor such as a CPU (Central Processing Unit), and executes a program describing the processing in the information processing unit 20 of this embodiment.
  • the input unit 102 is composed of, for example, a keyboard, buttons, a mouse, etc., and is used by the user of the computer system to input various information.
  • the memory unit 103 includes various types of memory such as RAM (Random Access Memory) and ROM (Read Only Memory) and storage devices such as a hard disk, and stores programs to be executed by the control unit 101, necessary data obtained during processing, etc.
  • the memory unit 103 is also used as a temporary storage area for programs.
  • the control unit 101 and the memory unit 103 for example, constitute a processing circuit.
  • the processing circuit may be a single circuit or multiple circuits.
  • the display unit 104 is composed of a display, LCD (Liquid Crystal Display), etc., and displays various screens to the user of the computer system. It should be noted that a touch panel in which the input unit 102 and display unit 104 are integrated may also be used.
  • the communication unit 105 is a receiver and transmitter that perform communication processing.
  • the output unit 106 is a speaker or the like. Note that FIG. 15 is just an example, and the configuration of the computer system that realizes each of the information processing units 20 is not limited to the example shown in FIG. 15. For example, the output unit 106 may not be provided.
  • the program is installed into the storage unit 103 from, for example, a CD-ROM or DVD-ROM inserted in a CD (Compact Disc)-ROM drive or DVD (Digital Versatile Disc)-ROM drive (not shown). Then, when the program is executed, the program read from the storage unit 103 is stored in the main memory area of the storage unit 103. In this state, the control unit 101 performs the processing of each of the information processing units 20 of this embodiment in accordance with the program stored in the storage unit 103.
  • the programs describing the processing in each information processing unit 20 are provided on a CD-ROM or DVD-ROM as recording media.
  • this is not limiting.
  • programs provided via a transmission medium such as the Internet via the communications unit 105 it is also possible to use programs provided via a transmission medium such as the Internet via the communications unit 105.
  • the program of this embodiment causes a computer system to execute, for example, the steps of acquiring a question about facility usage from a facility user who is a user of the facility, acquiring additional information indicating the status of the facility user who asked the question, generating an answer to the question using the additional information, the question, and facility data related to the facility, and outputting the answer.
  • the preprocessing unit 112 and answer generation unit 14 shown in FIG. 1 are realized by the control unit 101 shown in FIG. 15 executing a program stored in the storage unit 103.
  • the storage unit 103 is also used to realize the preprocessing unit 112 and answer generation unit 14.
  • the data acquisition unit 111 shown in FIG. 1 is realized by at least one of the communication unit 105 and the input unit 102 shown in FIG. 15. Some of the functions of the data acquisition unit 111 may be realized by the control unit 101 and the storage unit 103.
  • the data storage unit 113 shown in FIG. 1 is part of the storage unit 103 shown in FIG. 15.
  • the information processing unit 20 may be realized by multiple computer systems. For example, the information processing unit 20 may be realized by a cloud computer system.
  • the question acquisition unit 12 of this embodiment is realized by at least one of an input means such as a touch panel or keyboard, and a microphone.
  • the situation acquisition unit 13 of this embodiment is realized by at least one of a sensor such as a camera, an input means such as a touch panel or keyboard, and a microphone.
  • the output unit 15 is realized by at least one of a display such as a touch panel, and a speaker. As described above, if the display that realizes the output unit 15 includes a display and is a touch panel, the touch panel may also function as the question acquisition unit 12. Furthermore, this touch panel may also function as the situation acquisition unit 13.
  • the entire automatic response system 10 may be considered to be the computer system illustrated in FIG. 15.
  • the computer system may include a microphone as the input unit 102, and may further include a sensor as the situation acquisition unit 13.
  • the question acquisition unit 12 is realized by the input unit 102
  • the situation acquisition unit 13 is realized by at least one of the sensor and the input unit 102
  • the output unit 15 is realized by at least one of the display unit 104 and the output unit 106.
  • the automated response system 10 of this embodiment uses the situation of the questioner 30 to generate an answer to the question from the questioner 30. This allows for a more appropriate answer to be provided to the facility user compared to simply answering a specific question from the facility user.
  • the automated response system 10 may also generate an answer based on attributes in addition to the situation, and in this case, it is possible to provide the facility user with a more appropriate answer based on the attributes.
  • by including dynamic facility data as facility data it is possible to provide the facility user with appropriate answers and information based on the current state of the facility.
  • FIG. 16 is a diagram showing an example of the configuration of an automatic response system according to the second embodiment.
  • the automatic response system 10a of this embodiment includes an automatic response device 20a and a terminal device 60.
  • the automatic response device 20a is similar to the information processing unit 20 of the first embodiment, except that a transmitting/receiving unit 16 is added.
  • the terminal device 60 includes a question acquisition unit 12, a situation acquisition unit 13, an output unit 15, and a transmitting/receiving unit 17.
  • Components having the same functions as those of the first embodiment are assigned the same reference numerals as those of the first embodiment, and redundant explanations will be omitted. Below, differences from the first embodiment will be mainly explained.
  • the terminal device 60 is a device that can be operated by the facility user, such as a smartphone, tablet, or personal computer.
  • the terminal device 60 may also be a mobile terminal that can be carried by the facility user.
  • the facility user uses the automatic response system 10a by using the terminal device 60, for example, at home or on the way to the facility. Note that the location where the facility user uses the automatic response system 10a is not limited to this and may be any location, even within the facility.
  • the question acquisition unit 12 of the terminal device 60 acquires questions about the facility from facility users by at least one of voice and text.
  • the questions may include images.
  • the question acquisition unit 12 outputs the acquired questions to the transmission/reception unit 17.
  • the question acquisition unit 12 is, for example, at least one of an input means such as a touch panel or keyboard of a smartphone, tablet, personal computer, etc., and a microphone.
  • an image is included in the question, for example, the facility user specifies image data stored in the terminal device 60, and the input means accepts the specification.
  • the question acquisition unit 12 may convert the voice data into text data and output it to the transmission/reception unit 17, or may output the voice data directly to the transmission/reception unit 17.
  • the situation acquisition unit 13 acquires additional information indicating the situation of the questioner 30, who is a facility user asking a question (the situation in which the questioner finds himself/herself). Similar to embodiment 1, the additional information may also be information indicating the attributes of the facility user.
  • the situation acquisition unit 13 may acquire the additional information in the form of at least one of audio and text, or may acquire image data of the questioner 30 as the additional information.
  • the output unit 15 may output a screen to the questioner 30 asking questions such as the number of people using the facility and whether or not assistive devices will be used, and the answers entered on that screen may be acquired as the additional information.
  • the situation acquisition unit 13 may be realized, for example, by a camera built into the terminal device 60, such as a smartphone, tablet, or personal computer, and may output video data acquired by the camera to the transmission/reception unit 17 as the additional information.
  • This camera may, for example, be an internal camera that captures the questioner from inside the terminal device 60.
  • the output unit 15 outputs the answer by at least one of audio output and display.
  • the output unit 15 is, for example, at least one of the display of the terminal device 60 and the microphone of the terminal device 60, but is not limited to these.
  • the display of the terminal device 60 may be a touch panel. In this case, the touch panel may also function as at least one of the question acquisition unit 12 and the situation acquisition unit 13.
  • the transmission/reception unit 17 communicates with the automatic response device 20a, thereby exchanging information with the automatic response device 20a. For example, the transmission/reception unit 17 transmits the question received from the question acquisition unit 12 and the additional information received from the situation acquisition unit 13 to the automatic response device 20a. The transmission/reception unit 17 also outputs the response received from the automatic response device 20a to the output unit 15.
  • the operations of the question acquisition unit 12, situation acquisition unit 13, output unit 15, and transceiver unit 17 described above may be performed by installing application software that provides facility guidance services on the terminal device 60, or by the terminal device 60 accessing the automatic response device 20a.
  • the automatic response device 20a may function as a web server, and the terminal device 60 may access the web server to realize the above operations.
  • the transmitter/receiver unit 16 of the automatic response device 20a communicates with the terminal device 60 to exchange information with the terminal device 60.
  • the transmitter/receiver unit 16 outputs the question and additional information received from the terminal device 60 to the answer generation unit 14.
  • the transmitter/receiver unit 16 also outputs the answer received from the answer generation unit 14 to the terminal device 60.
  • the answer generation unit 14, like the answer generation unit 14 in embodiment 1, generates an answer using the question, additional information, and facility data, and outputs the generated answer to the transmitter/receiver unit 16.
  • the automatic response device 20a of this embodiment is realized by a computer system, like the information processing unit 20 in embodiment 1. Note that the automatic response device 20a may be installed inside or outside the facility.
  • facility users can ask questions using the terminal device 60, allowing them to use the automatic response system 10a regardless of their location. Therefore, for example, before using a facility, they can obtain information about the facility in advance, either at home or on the way to the facility. This allows facility users to use the facility efficiently based on the information they obtain in advance. Furthermore, by including map information for generating a route to the facility in the facility data, when a facility user asks how to get to the facility, it is possible to present the route from the facility user's current location to the facility as an answer. Furthermore, for example, facility data may also include data on reported lost items.
  • the automatic response system 10a can respond with whether or not a corresponding lost item has been reported.
  • FIGS. 17 and 18 are diagrams showing specific examples of responses that take the situation into consideration in the automatic response system 10a of this embodiment. In both of the examples shown in FIG. 17 and FIG. 18, it is assumed that it is raining around the facility, and that the facility data includes data indicating the weather.
  • the facility is a department store
  • the questioner 30, who is visiting the department store asks a question using the terminal device 60 before arriving at the department store.
  • the questioner 30 is raining, and the questioner 30 is holding an umbrella.
  • the terminal device 60 accepts the question, which is voice data, "Please tell me the way to the department store," and transmits the question and accompanying information, which is video data of the questioner 30, to the automatic response device 20a.
  • the automatic response device 20a determines that even though it is raining, the questioner 30 is holding an umbrella and should be guided along the normal route, and generates the normal route to the department store (for example, the shortest route) as an answer and transmits the answer to the terminal device 60.
  • the terminal device 60 displays the route to the department store as an answer.
  • the questioner 30 is holding an umbrella, but this is not limiting.
  • the automatic response device 20a may generate a response that guides the questioner 30 along a normal route even if it is raining. In other words, the automatic response device 20a may generate a response that guides the questioner 30 along a normal route as long as the questioner 30 is holding an umbrella, regardless of the state of the umbrella.
  • the facility is a conference hall where a briefing session is scheduled.
  • the questioner 30 is dressed formally and does not have an umbrella.
  • the terminal device 60 receives a question in the form of audio data, "Please tell me the way to the briefing session venue," and transmits the question and accompanying information, which is video data of the questioner 30, to the automatic answering device 20a.
  • the automatic answering device 20a determines that it is raining, the questioner 30 is not using an umbrella, and is dressed formally, and therefore should guide the questioner 30 along a covered route that will protect them from the rain.
  • the automatic answering device 20a generates an answer that includes a route that will protect them from the rain and transmits the answer to the terminal device 60.
  • the terminal device 60 displays the route to the conference hall as the answer.
  • the automatic answering device 20a can provide an answer that is appropriate for the questioner 30.
  • facility users ask questions using terminal device 60, and terminal device 60 outputs answers. This allows facility users to use automatic response system 10a even when they are outside the facility.
  • transceiver unit 16 By adding a transceiver unit 16 to the automatic response system 10 shown in embodiment 1 and having the transceiver unit 16 communicate with the terminal device 60, it may be possible to respond to both questions from facility users within the facility as described in embodiment 1 and questions using the terminal device 60 as described in this embodiment.
  • Embodiment 3. 19 is a diagram showing an example of the configuration of an automatic response system according to the third embodiment.
  • the automatic response system 10b of this embodiment is similar to the automatic response system 10 of the first embodiment, except that a rating acquisition unit 18 is added and the data storage unit 113 further stores rating information.
  • Components having the same functions as those of the first embodiment are assigned the same reference numerals as those of the first embodiment, and redundant explanations will be omitted. Below, differences from the first embodiment will be mainly explained.
  • the evaluation acquisition unit 18 acquires an evaluation result from the questioner 30 indicating an evaluation of the answer, and stores the acquired evaluation result in the data storage unit 113 along with the corresponding question and answer.
  • evaluation information is stored in the data storage unit 113
  • this is not limiting, and an evaluation information storage unit that stores evaluation information separate from the data storage unit 113 may be provided.
  • the output unit 15 outputs a question inquiring about the evaluation of the answer from the automatic response system 10b, such as "Did you get the information you wanted?" or "Was the answer appropriate?" Furthermore, if the evaluation from the questioner 30 includes negative words such as "unsatisfactory" or "not good," the evaluation acquisition unit 18 may cause the output unit 15 to output a question further asking the questioner 30 what the problem was, and may accept input from the questioner 30.
  • the evaluation acquisition unit 18 may acquire the evaluation result by voice or as text data.
  • the output unit 15 may present options indicating the evaluation results, and the evaluation acquisition unit 18 may acquire the selection made by the questioner 30 as the evaluation result.
  • the evaluation acquisition unit 18 may evaluate the content of the answer based on the behavior of the questioner 30. For example, after the automatic response system 10b outputs a route to a destination to a questioner 30 who has asked for directions, it may analyze the behavior of the questioner 30 using video data captured of the questioner 30, and if the questioner 30 is still lost, it may evaluate that the answer was difficult to understand.
  • the evaluation information may be used, for example, together with the corresponding questions and answers, for the answer generation unit 14 to re-learn or additionally learn, or may be referenced by the answer generation unit 14 when generating answers.
  • the operator or administrator of the automatic response system 10b may check the evaluation information, use the results to identify areas for improvement, and instruct the answer generation unit 14 on the identified areas for improvement.
  • the operator or administrator of the automatic response system 10b may check the evaluation information, use the results to determine what type of re-learning or additional learning the answer generation unit 14 should perform, and the determined results may be used to perform the re-learning or additional learning.
  • the evaluation acquisition unit 18 is, for example, at least one of an input means via operation and a microphone.
  • the input means, microphone, and other hardware may be shared with the question acquisition unit 12.
  • an evaluation acquisition unit 18 may be added to the automatic answering system 10a of embodiment 2 to reflect the evaluation results, or an evaluation acquisition unit 18 may be added to an automatic answering system that combines embodiments 1 and 2 to reflect the evaluation results.
  • 10, 10a, 10b Automatic response system
  • 11 Data management unit
  • 12 Question acquisition unit
  • 13 Situation acquisition unit
  • 14 Answer generation unit
  • 15 Output unit
  • 16, 17 Transmitting/receiving unit
  • 18 Evaluation acquisition unit
  • 20 Information processing unit
  • 20a Automatic response device
  • 30 Questioner
  • 40 Main unit
  • 50 Camera
  • 60 Terminal device
  • 111 Data acquisition unit
  • 112 Preprocessing unit
  • 113 Data storage unit
  • 142 Generation unit.

Landscapes

  • Engineering & Computer Science (AREA)
  • Databases & Information Systems (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

本開示にかかる自動応答システム(10)は、施設の利用に関する質問を施設の利用者である施設利用者から取得する質問取得部(12)と、質問した施設利用者である質問者の状況を示す付帯情報を取得する状況取得部(13)と、質問と付帯情報と施設に関するデータである施設データとを用いて、質問への回答を生成する回答生成部(14)と、回答を出力する出力部(15)と、を備える。

Description

自動応答システム、自動応答方法およびプログラム
 本開示は、質問への回答を生成する自動応答システム、自動応答方法およびプログラムに関する。
 近年、駅構内、商業施設等において、利用者からの質問への対応、利用者の案内などにより、従業員の時間がとられることが問題となっている。また、利用者が質問しようと思った際に従業員が他の利用者への対応をしていることがあり、利用者が質問したいときに質問できないことも問題となっている。このため、利用者への案内の自動化が検討されている。
 特許文献1には、ショッピングモールやイベント会場などにおいて用いられる道案内システムが開示されている。特許文献1に記載の道案内システムは、利用者が目的地をロボットに伝えると、表示装置が経路を表示するとともに、ロボットが音声および身体動作により利用者に目的地までの経路を説明する。
特開2007-265329号公報
 特許文献1に記載の技術では、従業員の代わりに、利用者に、指定された目的地までの経路を案内することができる。しかしながら、利用者が知りたいことは多様であり、利用者がおかれている状況も多様である。このため、特許文献1に記載の技術では、利用者に適した情報を提供できるとは限らない。
 本開示は、上記に鑑みてなされたものであって、施設の利用者により適切な回答を提供することができる自動応答システムを得ることを目的とする。
 上述した課題を解決し、目的を達成するために、本開示にかかる自動応答システムは、施設の利用に関する質問を施設の利用者である施設利用者から取得する質問取得部と、質問した施設利用者である質問者の状況を示す付帯情報を取得する状況取得部と、質問と付帯情報と施設に関するデータである施設データとを用いて、質問への回答を生成する回答生成部と、回答を出力する出力部と、を備える。
 本開示にかかる自動応答システムは、施設の利用者により適切な回答を提供することができるという効果を奏する。
実施の形態1にかかる自動応答システムの構成例を示す図 実施の形態1の回答生成部の構成例を示す図 実施の形態1の自動応答を説明するための図 実施の形態1の自動応答を説明するための図 実施の形態1の自動応答システムにおける自動応答処理手順の一例を示すフローチャート 図5のステップS5に示した回答生成部における回答生成処理手順の一例を示すフローチャート 実施の形態1の動的な施設データの一例を示す図 実施の形態1の自動応答システムにおける状況を考慮した回答の具体例を示す図 実施の形態1の自動応答システムにおける状況を考慮した回答の具体例を示す図 実施の形態1の自動応答システムにおける状況を考慮した回答の具体例を示す図 実施の形態1の自動応答システムにおける状況を考慮した回答の具体例を示す図 実施の形態1の自動応答システムにおける状況を考慮した回答の具体例を示す図 実施の形態1の自動応答システムにおける状況を考慮した回答の具体例を示す図 実施の形態1の自動応答システムにおける状況を考慮した回答の別の具体例を示す図 実施の形態1の情報処理部のそれぞれを実現するコンピュータシステムの構成例を示す図 実施の形態2にかかる自動応答システムの構成例を示す図 実施の形態2の自動応答システムにおける状況を考慮した回答の具体例を示す図 実施の形態2の自動応答システムにおける状況を考慮した回答の具体例を示す図 実施の形態3にかかる自動応答システムの構成例を示す図
 以下に、実施の形態にかかる自動応答システム、自動応答方法およびプログラムを図面に基づいて詳細に説明する。
実施の形態1.
 図1は、実施の形態にかかる自動応答システムの構成例を示す図である。本実施の形態の自動応答システム10は、施設の利用者からの施設の利用に関する質問を受付け、質問への回答として施設の案内に関する情報を出力する。施設は、例えば、駅構内、病院、商業施設、ホテル、ビル、イベント会場、会議場、テーマパーク、複数の駅を含む鉄道施設(鉄道車両内を含む)、学校、事業所、工場などであるが、複数の利用者が利用可能であればよく、これらに限定されない。施設は、駅構内などのように不特定多数の人が利用可能な施設であってもよいし、学校、事業所などのように特定の人が利用可能な施設であってもよいし、特定の人と不特定の人との両方が利用可能な施設であってもよい。
 図1に示すように自動応答システム10は、データ管理部11、質問取得部12、状況取得部13、回答生成部14および出力部15を備える。データ管理部11および回答生成部14は、情報処理装置である情報処理部20を構成する。自動応答システム10は、例えば、施設内に設置されるが、自動応答システム10を構成する各部のうち少なくとも一部が、施設外に設置されてもよい。
 データ管理部11は、施設に関するデータである施設データを管理する。施設データは、回答生成部14が回答を生成する際に用いられるデータであり、施設自体に関するデータだけでなく、施設の周辺のデータ、施設に関連する施設外のデータを含んでいてもよい。施設データは、例えば、施設内の地図情報(マップ情報)、施設の状態に関する施設状態情報、施設の周辺の地図情報、施設の管理者および施設の従業員のうち少なくとも一方によって発信される施設発信情報、施設を含むエリアの気象情報のうちの少なくとも1つを含むが、これに限定されない。地図情報は、施設内の店舗、会議室などの利用者が訪れる場所の候補までの経路を決定可能な情報を含む。例えば、地図情報は、施設内の通路などをリンクとして上記候補をノードとするグラフにより表された情報を含んでいてもよい。また、通路が階段を含む、スロープを含む、エレベータを含む、エスカレータを含むといった通路の属性を含んでいてもよい。または、施設内の地図情報は、利用者が訪れる場所の候補ごとに、あらかじめ定められた経路を示す画像データを含んでいてもよい。この場合、地図情報は、スロープを含む、エレベータを含む、エスカレータを含むといった通路の属性ごとの画像データを含んでいてもよい。
 施設状態情報は、施設における混雑の度合いを示す混雑情報、施設内の通路の通行止めなどの情報、施設内の設備の故障の有無を示す情報、周辺の道路の状況を示す道路状況などであるが、これらに限定されない。施設発信情報は、例えば、「雨が降っているため、施設内が滑りやすくなっております」、「本日の10時からのXイベントの会場はXからYに変更になっております」などのように、施設の管理者または従業員が利用者に伝えたい情報である。気象情報は、例えば、天候(晴れ、雨、強風など)、気温、湿度のうちの少なくとも1つを含むがこれに限定されない。
 また、施設データは、施設が駅構内の場合には、電車などの鉄道車両の時刻表を含んでいてもよく、施設内に店舗が存在する場合には、店舗が取り扱う商品に関する情報、店舗で提供される飲食物に関する情報、店舗の主なターゲット層を示す情報、店舗のおすすめの商品に関する情報、店舗の従業員によって発信される店舗発信情報などのうちの少なくとも1つを含んでいてもよい。店舗発信情報は、例えば、「本日は、Aショップでは雨の日セールを行っております」、「Bショップではタイムセールを実施中です」、「K飲食店では、暑い日におすすめの冷たいお飲み物をたくさんご用意しております」などのように、店舗の従業員が利用者に伝えたい情報である。
 施設データは、施設内の地図情報、店舗の主なターゲット層を示す情報などのように短期間では変化しない、すなわち一定時間変化しない静的なデータを含んでいてもよい。また、施設データは、施設状態情報、施設発信情報、気象情報などのように短期間で変化する可能性のある、すなわち一定時間の間にも変化する可能性がある動的なデータであってもよい。また、施設データは、静的なデータと動的なデータとの両方を含んでいてもよい。
 データ管理部11は、データ取得部111、前処理部112およびデータ記憶部113を備える。データ取得部111は、施設データを取得し、取得した施設データを前処理部112へ出力する。
 具体的には、データ取得部111は、施設データのうち静的なデータについては、例えば、施設の管理者、運用者、店舗の従業員などから入力されるデータを受付けることで、当該データを取得してもよいし、図示しない他の装置から受信することで当該データを取得してもよい。例えば、データ取得部111は、施設の周辺の地図情報を、当該地図情報を管理する装置から受信してもよい。また、施設の管理者、運用者、店舗の従業員などが操作可能な端末装置が、施設の管理者、運用者、店舗の従業員などから施設データの入力を受付け、データ取得部111が当該端末装置から施設データを受信してもよい。
 また、データ取得部111は、施設データのうち動的なデータについては、例えば、監視カメラ、マイクなどの図1では図示しないセンサから、当該センサによって取得されたセンサ情報を受信し、受信したセンサ情報を基に動的なデータを算出することにより動的なデータを取得してもよい。例えば、動的なデータは、施設内および施設の周辺のうち少なくとも一方に設置されたセンサによって取得されたデータを含んでいてもよい。例えば、データ取得部111は、監視カメラによって取得されたセンサ情報である映像に基づいて混雑情報を算出してもよい。または、データ取得部111は、センサ情報を前処理部112へ出力し、前処理部112がセンサ情報から動的なデータを算出してもよい。または、図示しない他の装置が、センサ情報に基づいて動的なデータを算出し、データ取得部111が当該他の装置から動的なデータを受信してもよい。
 また、例えば、データ取得部111は、動的なデータである施設発信情報、店舗発信情報については、発信者からの入力を受付けることによりテキストデータとして施設発信情報、店舗発信情報を取得してもよいし、音声により施設発信情報、店舗発信情報の入力を受付けてもよい。また、施設において施設内放送を行うためのマイクから、施設内で放送される情報を、施設発信情報、店舗発信情報として取得してもよい。または、施設内放送を行う場所の周辺にマイクを設置し、当該マイクから、施設内で放送される情報を、施設発信情報、店舗発信情報として取得してもよい。
 前処理部112は、データ取得部111から受け取った施設データに前処理を実施する。前処理部112は、例えば、前処理後の施設データをデータベースとしてデータ記憶部113に格納する。前処理は、例えば、回答生成部14がアクセス可能なデータ形式へ変換する処理である。例えば、回答生成部14が、ベクトル化されたデータを読み込む場合には、前処理部112は、施設データをベクトル化し、ベクトル化した施設データをデータ記憶部113に格納する。ベクトル化は、例えば、自然言語処理における単語を、単語が持つ意味に応じた分散表現であるベクトルへ変換する処理を含む。
 なお、前処理部112が行う前処理の内容は、上述した例に限定されず、回答生成部14における処理の方式に応じて決定されればよい。なお、回答生成部14における処理の方式によっては、前処理部112が設けられずに、施設データがそのままデータ記憶部113に記憶されてもよい。データ記憶部113は、施設データを記憶する。
 質問取得部12は、施設の利用に関する質問を施設の利用者(以下、施設利用者とも呼ぶ)から取得する。詳細には、質問取得部12は、施設利用者からの施設に関する質問を受付け、受付けた質問を回答生成部14へ出力する。質問取得部12は、質問を音声データとして取得してもよいし、施設利用者の操作による入力を受付けることでテキストデータとして質問を取得してもよい。すなわち、質問取得部12は、マイクを備えていてもよいし、タッチパネルなどの利用者の操作によって入力を受付ける入力手段を備えていてもよく、マイクと入力手段との両方を備えていてもよい。なお、質問取得部12は、画像の入力を受付けてもよく、この場合、例えば、カメラによって画像を撮影することで画像の入力を受付ける。例えば、施設利用者が、商品、店舗などの写真を提示して、この写真の商品はどこにあるかといった質問、このお店はどこにあるかといった質問を音声またはテキストにより入力してもよい。カメラは、後述する状況取得部13として用いられるカメラであってもよい。質問取得部12は、画像を用いた質問を受付けた場合には、画像データも質問に含めて回答生成部14へ出力する。
 状況取得部13は、質問した施設利用者の状況を示す付帯情報を取得し、取得した付帯情報を回答生成部14へ出力する。例えば、状況取得部13は、カメラ、赤外線センサなどのセンサを備え、当該センサによる検出結果を付帯情報(補助情報)として取得する。例えば、付帯情報は、質問者を撮影した映像データを含む。また、状況取得部13は、施設利用者の操作による入力を受付けることでテキストデータとして施設利用者の状況を示す付帯情報を取得してもよいし、施設利用者の状況を示す付帯情報を音声データとして取得してもよい。状況取得部13が、テキストデータ、音声データとして付帯情報を取得する場合には、例えば、出力部15によって状況を問う質問が施設利用者に提示され、状況取得部13は、当該質問への回答を付帯情報として取得する。テキストデータ、音声データとして付帯情報を取得する場合には、質問取得部12が状況取得部13としての機能を有していてもよい。また、状況取得部13が音声データの取得に用いるマイクは、質問取得部12と共用されてもよい。
 上記のように、付帯情報は、例えば、質問する施設利用者の状況を示す状況情報であるが、これに限らず、付帯情報は、質問する施設利用者の状況を示す状況情報であるとともに、質問する施設利用者の属性を示す属性情報であってもよい。例えば、カメラによって取得された映像データは、施設利用者の状況と属性との両方の把握に用いることができる。また、状況取得部13は、出力部15によって状況に加えて属性を問う質問が施設利用者に提示され、状況取得部13は、当該質問への回答を付帯情報として取得してもよい。これにより、付帯情報が、状況情報と属性情報とを含むようにしてもよい。または、状況情報として映像データが取得され、属性を問う質問への回答として属性情報が取得されるといったように、状況情報と属性情報とが異なる手段で取得されてもよい。なお、施設利用者の属性を取得する方法は、カメラによって取得された映像データから取得する例に限定されず、例えば、後述するように、個人の識別結果を用いて当該個人の属性が把握されてもよい。
 回答生成部14は、質問取得部12から受け取った質問と、状況取得部13から受け取った付帯情報と、データ記憶部113に格納されている施設データとを用いて、施設の案内に関する回答を生成し、生成した回答を出力部15へ出力する。具体的には、回答生成部14は、質問を理解し、施設データを用いて付帯情報が示す状況を考慮した回答を生成する。回答生成部14は、例えば、音声、テキストおよび画像を含む複数の形式のデータを入力として用いて回答を生成することが可能であり、回答を複数の形式で出力することが可能である。すなわち、例えば、回答生成部14は、音声データ、テキストデータ、映像データなどマルチモーダルな入力に対してマルチモーダルな出力を行うことが可能であるが、これに限定されない。以下、マルチモーダルな入力に対してマルチモーダルな出力を行うことをマルチモーダル処理とも呼ぶ。回答生成部14の詳細については後述する。
 出力部15は、回答生成部14から受け取った回答を出力することによって、質問した施設利用者に回答を提示する。回答は、音声として出力されてもよいし、テキストデータとして出力されてもよいし、画像や映像によって出力されてもよいし、これらのうちの2つ以上の組み合わせにより出力されてもよい。したがって、出力部15は、例えば、スピーカと、ディスプレイ、タッチパネルなどの表示装置とのうちの少なくとも一方を備える。なお、出力部15は、回答生成部14がマルチモーダル処理を行うことが可能な場合、回答生成部14から受け取った回答の形式に応じた出力を行う。例えば、回答生成部14から受け取った回答が音声データであれば、出力部15は、音声として回答を出力し、回答生成部14から受け取った回答が映像データおよびテキストデータのうちの少なくとも一方であれば、表示により回答を出力する。なお、出力部15がタッチパネルを備える場合、当該タッチパネルは、質問取得部12における入力手段としても用いられてもよく、状況取得部13における入力手段としても用いられてもよい。
 次に回答生成部14の詳細について説明する。図2は、本実施の形態の回答生成部14の構成例を示す図である。回答生成部14は、例えば、図2に示すように、状況認識部141および生成部142を備える。状況認識部141は、状況取得部13から受け取った付帯情報に基づき、状況を認識し、認識した状況を生成部142へ出力する。なお、状況認識部141が状況を認識する対象は質問した施設利用者であるが、施設利用者が家族連れ、カップルなどのように同行者がいる場合には、以下では、実際に質問した施設利用者だけでなく、当該施設利用者の同行者についても質問した施設利用者として扱う。また、以下、質問した施設利用者を質問者とも呼ぶ。状況認識部141が認識する状況は、例えば、質問者が身に着けているもの(質問者の服装、質問者が身に着けている装飾品など)、質問者が車椅子、杖などの補助具を利用しているか否か(補助具の利用の有無)、質問者の人数、質問者の感情、質問者が持っている荷物の大きさ、質問者の移動方向のうちの少なくとも1つを含むがこれに限定されない。
 また、上述したように付帯情報が、質問者の状況を示すとともに属性を示す情報である場合、状況認識部141は、付帯情報を用いて、質問者の状況および属性を認識し、認識した状況および属性を生成部142へ出力する。質問者の属性は、例えば、性別、年齢、年代、障がいの有無および職業のうちの少なくとも1つを含んでもよい。
 生成部142は、状況認識部141から受け取った状況(または状況および属性)と、質問取得部12から受け取った質問と、データ記憶部113に格納されている施設データとを用いて、施設の案内に関する回答を生成し、生成した回答を出力部15へ出力する。
 回答生成部14は、このように、質問者の状況を反映して回答を生成することにより、例えば、車椅子を利用している質問者が目的地までの行き方を質問した場合に、階段がなく、スロープやエレベータを使用する経路を回答として生成することができる。
 図3および図4は、本実施の形態の自動応答を説明するための図である。図3に示した例では、自動応答システム10が、状況取得部13としてカメラを備え、出力部15としてタッチパネルとスピーカとを備えている。図3に示した例では、タッチパネルの下部に質問を受付けるためのキーボード(ソフトウェアキーボード)が表示されており、このキーボードが表示されている部分が質問取得部12の機能を兼ねている。図3に示した例では、自動応答システム10は、さらに、質問取得部12として、マイクを備えている。図3の上部に示すように、質問者30が大きな荷物を持った状態で、「駅のホームはどこにありますか?」と発話することにより、自動応答システム10へ駅のホームの場所を尋ねている。
 自動応答システム10の回答生成部14は、状況取得部13として用いられるセンサの一例であるカメラから受け取った映像データから質問者30が大きな荷物を持っているという状況を認識し、認識した状況と質問取得部12から受け取った質問の音声データと施設データとを用いて回答を生成する。なお、図3に示した例で認識される状況は、「大きな荷物を持っている」に限らず、「両手がふさがっている」、「歩きにくい」などであってもよい。図3に示した例では、質問者30が大きな荷物を持っていることから、回答生成部14は、駅のホームへの行き方、すなわちホームまでの経路として、階段またはエスカレータを用いる経路ではなく、図3の下部に示すように、エレベータを用いる経路を施設データにおける施設内の地図情報に基づいて生成する。なお、経路の生成方法に特に制約はなく、条件を満たす範囲での最短経路を探索するなどの方法を用いることができる。図3に示した例では、回答生成部14は、回答として、経路を示す図(画像)と、その図を説明するテキストデータとを生成し、生成した回答を出力部15へ出力し、出力部15が回答を出力する。図3に示した例では、出力部15のうちのタッチパネルに、回答としてテキストと経路を示す図とが表示されているが、出力の方法はこれに限定されない。例えば、さらに出力部15のうちのスピーカが音声として回答を出力してもよいし、テキストの代わりにスピーカが音声として回答を出力してもよい。
 図4に示した例では、自動応答システム10は、状況取得部13として用いられるセンサの一例であるカメラと、本体40とで構成される。本体40は、図4では図示しない情報処理部20と、質問取得部12と、出力部15とを備える。このように、状況取得部13として用いられるセンサは、本体40とは別に設けられてもよい。また、自動応答システム10は、状況取得部13として複数のセンサを備えてもよく、例えば、本体40と本体40の外部との両方にそれぞれセンサが設けられてもよい。
 図4に示した例では、質問者30が車椅子を利用しており、「出口はどこですか?」と自動応答システム10に尋ねている。図4に示した例では、質問取得部12であるマイクが質問を受付ける。図4に示した例では、質問者30が車椅子を利用していることから、回答生成部14は、出口までの経路として、段差のないスロープを含む経路を回答として生成し、出力部15が回答を表示している。また、出口の場所を質問された場合などには、回答生成部14は、施設外へ出ると想定し、施設データに基づいて雨が降っている場合には、回答に「外は雨が降っていますので、気を付けてお帰りください。」といったように、外の天候に応じた文章を追加してもよい。
 なお、図4に示した例では、出口を尋ねているが、施設データに施設の周辺の地図情報を含めておくことにより、例えば、質問者が、施設が駅構内である場合に、「デパートまでの経路を教えてください」といった質問をした場合に、自動応答システム10が駅からデパートまでの経路を案内できるようにしてもよい。このように、回答生成部14は、施設の案内に関する情報として施設の周辺に関する情報を回答してもよい。また、施設が駅構内である場合に、当該駅に停車する鉄道車両の路線、当該路線の停車駅、駅の最寄の店舗、観光地などを示す情報を施設データに含めておき、「VVへ行くには何線に乗ればよいですか?」といった質問に対して、何線の鉄道車両に乗ってどこの駅で降りればよいかを回答できるようにしてもよい。
 図3に示した例では、自動応答システム10が1つの装置として実現され、図4に示した例では、状況取得部13として用いられるセンサの一例であるカメラと、本体40とにより実現されるが、自動応答システム10のハードウェアとしての装置の構成はこれらの例に限定されない。例えば、図1に示した情報処理部20である情報処理装置と、質問取得部12、状況取得部13および出力部15を備える入出力装置とが別の場所に設けられてもよい。この場合、入出力装置および情報処理装置は、それぞれ通信を行う送受信部を備え、送受信部により情報のやりとりを行ってもよい。また、図3および図4は例示であり、自動応答システム10が考慮する状況、および自動応答システム10が提示する回答は、これらの例に限定されない。
 図2の説明に戻る。回答生成部14は、例えば、生成AI(Artificial Intelligence)であってもよい。生成AIには、例えば、マルチモーダルAIであるGemini(登録商標)、GPT(Generative Pre-trained Transformer)-4Vなどが用いられてもよいし、画像生成AIであるStable Diffusionなどが用いられてもよいし、大規模言語モデルであるGPT-3などが用いられてもよいし、音声生成AIであるMurf.AIなどが用いられてもよい。回答生成部14にマルチモーダルAIが用いられる場合、回答生成部14が、状況認識部141および生成部142の機能を備える。回答生成部14にマルチモーダルAIが用いられる場合、自動応答システム10の起動時など、施設利用者からの質問を受付ける前に、例えば、「質問した施設利用者の状況を考慮して、必要に応じて映像データ、テキストデータ、音声データを生成して回答してください」という指示を回答生成部14へ入力しておく。また、適宜、施設データも用いることも回答生成部14に指示しておく。これにより、回答生成部14は、質問者30の状況を考慮して、施設データを参照して施設の案内に関する回答を、映像データ、テキストデータおよび音声データのうちの少なくとも1つにより生成する。
 また、状況として具体的に何を考慮するかを、回答生成部14に入力しておくか、または施設データに状況の定義を示すデータを含めておいてもよい。例えば、「質問した施設利用者の状況は、質問者が車椅子、杖などの補助具を利用しているか否か、質問者の人数、質問者の感情、質問者の服装、質問者が身に着けている物、質問者が持っている荷物の大きさ、質問者の移動方向を含みます」という指示を回答生成部14に入力しておいてもよい。
 回答生成部14にマルチモーダルAIが用いられ、上記のような指示をしておくと、回答の形式は、回答生成部14によって決定される。一方、回答の形式の決定を回答生成部14にまかせずに、質問と同じ形式で回答を行うといった指示などのように、特定のルールを定めたい場合には、回答生成部14にそのルールを示す指示を入力しておくか、または施設データにルールを示すデータを含めておいてもよい。例えば、質問に対応する回答の典型例、規範、ガイドラインなどが、施設データとして、または施設データとは別に、データ記憶部113に格納され、回答生成部14が回答を作成する際に当該ルールを参照するようにしてもよい。または、これらのルールを学習するように事前学習が行われてもよい。なお、施設データは、回答の生成に用いられるが、施設データのうちの一部はそのまま回答の一部として用いられてもよい。例えば、施設データとして店舗が扱う商品の画像などが含まれている場合、回答生成部14は、回答に当該画像を含めてもよい。
 なお、例えば、回答生成部14が、図3および図4に例示した回答を生成するために、事前に回答生成部14であるマルチモーダルAIに、考慮すべき質問者30の状況と質問とを入力し、所望の回答が得られるかを確認し、所望の回答が得られない場合には、状況と質問に対応する正しい回答とをマルチモーダルAIに指示しておいてもよい。また、状況と質問に対応する正しい回答とを施設データに含めておいてもよい。例えば、図3に示したように、大きな荷物を持っている質問者30が目的地までの経路を質問した際に、回答生成部14が、階段を含む経路を回答として生成した場合、「質問者が大きな荷物を持っている場合には、階段は使わずにエレベータを使います」といった指示を回答生成部14に与えておいてもよいし、当該指示を施設データに含めておいてもよい。または、指示を分解して、「質問者が大きな荷物を持っていたり、両手がふさがっていると歩きにくい」という情報と「歩きにくいときには、階段は使わずにエレベータを使います」という情報とに分けてもよい。なお、状況に加えて属性を考慮する場合も、同様に、状況および属性ごとに、適切な回答となるような指示を回答生成部14に与えておいてもよいし、当該指示を施設データに含めておいてもよい。
 また、回答生成部14に大規模言語モデルが用いられる場合には、質問取得部12は、テキストデータとして質問を受付け、出力部15はテキストデータとして回答を出力してもよい。また、質問取得部12が質問を音声として受付け、質問取得部12または回答生成部14が、質問の音声データをテキストデータに変換してもよい。例えば、大規模言語モデルが生成部142としての機能を有し、状況認識部141として、映像データから状況を認識する状況認識モデルが用いられてもよい。状況認識モデルは、例えば、機械学習により映像データから状況を推論するように学習された学習済モデルであってもよい。学習済モデルの生成に用いられる機械学習としては、例えば、ニューラルネットワークなどの教師あり学習を例示できるが、教師なし学習、強化学習などであってもよく、これに限定されない。また、状況認識モデルは、さらに属性を認識するモデルであってもよい。
 また、状況認識モデルは、映像データから感情を認識する一般的な感情認識モデルと、映像データから補助具の有無、荷物の大きさ、服装などを認識する画像認識モデルとの組み合わせにより構成されるといったように、複数のモデルの組み合わせであってもよい。生成部142は、例えば、状況認識部141から受け取った状況(または、状況および属性)を示すテキストデータと、質問とを用いて回答をテキストデータとして生成する。または、生成部142としてさらに、音声生成AIが用いられ、大規模言語モデルが生成した回答を音声生成AIが音声データに変換してもよいし、生成部142としてさらに、画像生成AIが用いられ、大規模言語モデルが生成した回答を画像生成AIが画像データに変換してもよい。
 このように、複数のモデルを組み合わせて用いることで、マルチモーダルな出力が実現されてもよい。この場合、例えば、対応可能な全ての出力形式により出力が行われてもよいし、出力の形式を選択するためのルールをあらかじめ定めておき、ルールに従って、出力の形式が決定されてもよい。例えば、質問者30の状況(または、状況および属性)と出力の形式とが対応付けられてルールとして定められてもよいし、質問者30の質問の形式と同一の形式で出力することがルールとして定められてもよい。また、質問の内容の種別(例えば、行き先までの経路に関する質問、おすすめのお店を問う質問など)と出力の形式とが対応付けられてルールが定められていてもよい。出力の形式を決定するためのルールはこれらの例に限定されない。
 なお、回答生成部14に生成AI、大規模言語モデルなどが用いられる場合、施設データが前処理部112によりベクトル化されてデータ記憶部113に蓄積されることにより、回答生成部14が施設データをそのまま利用することができる。このため、追加学習を要せずに、回答生成部14の回答に施設データを反映させることができる。
 また、生成部142に、チャットボットなどの一般的な自動会話プログラムが用いられてもよい。この場合、状況認識部141は、生成部142に大規模言語モデルが用いられる場合と同様に、状況認識部141として状況認識モデルが用いられ、状況認識部141が認識した状況(または、状況および属性)を示すテキストデータが生成部142に入力されてもよい。
 自動会話プログラムは、あらかじめ応答のルールを定めておくルールベース型(シナリオ型)であってもよいし、機械学習型であってもよい。ルールベース型の自動会話プログラムが用いられる場合には、例えば、状況(または、状況および属性)ごとの質問の内容に応じた回答があらかじめ定められ、自動会話プログラムに設定される。機械学習型の自動会話プログラムにおける機械学習は、教師あり学習であってもよいし、教師なし学習であってもよいし、強化学習であってもよい。例えば、自動会話プログラムに教師あり学習が用いられる場合には、状況(または、状況および属性)ごとの質問と対応する正解データである回答とを含む学習用データセットを複数用いて学習済モデルが生成され、生成部142が当該学習済モデルに、質問者30に対応する状況(または、状況および属性)および質問を入力することにより回答を生成する。なお、生成部142に自動会話プログラムが用いられる場合も、自動会話プログラムが生成した回答を、音声生成AIにより音声データに変換してもよいし、画像生成AIにより画像データに変換してもよい。この場合も、対応可能な全ての出力形式により出力が行われてもよいし、出力の形式を選択するためのルールをあらかじめ定めておき、ルールに従って、出力の形式が決定されてもよい。
 なお、質問者30は、施設を訪れる不特定の人だけでなく、施設の従業員、施設の管理者、施設が学校である場合の学生などのように、定められた特定の人であってもよい。すなわち、施設利用者は、あらかじめ定められた特定の人を含んでいてもよい。例えば、施設が、特定の人向けの施設である場合などには、特定の人だけを質問者30として想定してもよい。この場合、例えば、特定の人の顔写真などを施設データに含めておき、状況取得部13が、付帯情報としてカメラの映像データを取得し、回答生成部14が、顔写真に基づいて、質問者30を識別し、識別結果に応じた回答を生成してもよい。この映像データは、個人を識別するための個人認証情報の一例である。例えば、あらかじめ施設を利用する可能性のある特定の人に関して、それぞれの識別情報(以下、個人識別情報とも呼ぶ)と顔写真とを対応付けて施設データとしてデータ記憶部113に格納しておく。また、個人識別情報ごとに、当該個人識別情報に対応する人に関する個別情報を施設データとしてデータ記憶部113に格納しておく。個別情報は、特定の人の個人ごとの施設の利用に関する情報であり、例えば、スケジュール、立ち入り可能な場所、取得可能な情報の範囲(情報の閲覧権限)などのうちの1つ以上であるが、これらに限定されない。
 例えば、施設が学校である場合、各学生が受講している授業を示す情報を施設データの個別情報に含めておき、各授業が行われる時間と場所(教室等)とを示す授業場所情報を施設データに含めておく。これにより、質問者30である学生から「次の授業がある場所はどこ?」といった質問があった場合に、回答生成部14は、当該学生の個別情報に基づき次に受講する授業を把握し、授業場所情報を用いて当該授業の行われる時間と場所とを把握することができる。回答生成部14は、把握した情報を反映して、例えば、「次の授業は3号棟のX番教室で行われるYの授業です」といった回答を生成する。この際に、現在地から3号棟のX番教室までの経路を示す画像が回答に含まれていてもよい。回答の生成の際には、上述したように、例えば、さらに状況が考慮される。
 また、施設が学校である場合に、学生だけでなく学校の職員についても、同様に、個別情報が施設データに含まれていてもよい。例えば、職員の行う授業の予定、職員が試験監督を行う時間と場所とを示す情報などが個別情報として含まれていてもよい。この場合も、学生の例と同様に、例えば、職員からの「私が次に試験監督を務める時間と場所はどこですか?」といった質問に対して、「次に試験監督を務めるのは2号棟のZ番教室で11時からです」といった回答を生成する。上述した回答の内容は例示であり、回答の内容は、上述した例に限定されない。なお、ここでは顔情報を用いて個人を特定する例を説明したが、個人を識別するために、例えば、各個人に定められたパスコードを施設データに含めておき、質問時に自動応答システム10がパスコードの入力を受付けることにより個人が識別されてもよいし、自動応答システム10が顔認証以外の生体認証を行う装置を備え、当該装置を用いて個人が識別されてもよい。すなわち、個人認証情報は、パスコード、生体認証情報などであってもよい。
 また、施設を利用する従業員、学生などの特定の人の属性ごとに、当該属性の質問者30への回答の生成に情報およびルールのうちの少なくとも一方を、属性関連情報として定めておき、属性関連情報を施設データに含めておいてもよい。属性関連情報は、属性ごとの施設の利用に関する情報であり、スケジュール、立ち入り可能な場所、取得可能な情報の範囲(情報の閲覧権限)、連絡事項などを含む。属性は、施設が学校であれば、例えば、身分(学生であるか、教授であるか、助教であるか、教師であるか、事務員であるかなど)、所属(学部、学科、選考)などを含んでいてもよい。また、施設が事業所などである場合には、属性は、役職、所属(学科、学部、部署)、保有する資格などを含んでいてもよい。このように、例えば、属性は、身分、所属および役職のうちの少なくとも1つを含んでいてもよい。例えば、施設データにおける個別情報に個人の属性を含めておくことで、回答生成部14が上述した個人の特定と同様の方法により個人を識別し、個別情報に基づき属性を把握し、属性に応じた回答を生成してもよい。
 例えば、属性関連情報として学科ごとに受講する授業が定められている場合に、上述した例と同様に、学生が「次の授業がある場所はどこ?」と質問すると、回答生成部14が、質問者30である学生の属性として学生の所属する学科を把握する。また、回答生成部14は、属性関連情報に基づき、把握した属性に対応する授業を把握し、授業場所情報に基づき当該授業の場所を示す回答を生成する。また、例えば、属性ごとに立ち入り可能な場所が定められている場合には、質問者30が目的地までの行き方を質問した場合には、立ち入り可能な場所の範囲を移動するように経路を求め、求めた経路を回答として生成する。また、例えば、情報の閲覧権限が属性に含まれている場合には、質問者30の回答を生成する際に閲覧が許可されている範囲の情報を用いて回答を生成する。また、属性ごとの連絡事項が属性関連情報に含まれている場合には、質問の内容に関わらず、質問に対応する直接の回答に連絡事項を付加したものを回答として生成してもよい。例えば、連絡事項として授業の変更のお知らせなどが含まれていてもよい。
 また、施設が、駅構内、商業施設、大学などのように、不特定の人が利用可能な施設である場合、施設の顧客、大学の訪問者などの不特定の人と、施設の従業員、学生などの特定の人との両方を質問者30と想定してもよい。回答生成部14は、不特定の人と、特定の人とで、回答生成部14が回答内容を変えてもよい。この場合、質問者30の属性として、例えば、特定の人であるか否かを含めておき、属性関連情報において上記の例と同様に立ち入り可能な場所、取得可能な情報の範囲(情報の閲覧権限)、連絡事項など属性ごとに回答を生成する際のルールを定義しておいてもよい。特定の人であるか否かの属性の把握は、上述した特定の人の個人を識別する場合と同様であり、例えば、特定の人に関して顔情報などをあらかじめ施設データに含めておくことにより行われる。例えば、属性関連情報は、例えば、特定の人の回答への生成に用いる施設データと、不特定の人への回答の生成に用いる施設データとを区別する情報を含んでいてもよい。また、従業員などの特定の人については、上述した例と同様に、属性としてさらに所属、役職などを含めておき、これらの属性に応じた回答が生成されてもよい。
 例えば、施設が商業施設である場合に、質問者30が「商品Pが収納されている倉庫はどこですか?」と質問した場合に、回答生成部14は、質問者30が従業員である場合には、施設データに基づいて商品Pが格納されている倉庫を示す情報を回答として生成し、質問者30が従業員ではない場合にはその質問には回答できないことを示す回答を生成する。なお、このとき、例えば、回答生成部14は、上述したように、さらに状況に応じて回答を生成する。
 また、回答生成部14は、質問者30からの質問の内容が言語として認識できなかったり、質問者30からの質問が回答を生成するために不十分であったり、質問者30のおかれた状況に関する情報が回答を生成するために不十分であったりする場合には、出力部15に質問者30へ、不足する情報を問う情報を出力させてもよい。すなわち、自動応答システム10の回答生成部14は、例えば、質問を認識できない場合、および回答の生成のために不足する情報がある場合に、質問者30に問い直しを行う回答を生成してもよい。例えば、質問者30からの質問の内容が言語として認識できなかった場合には、回答生成部14は、出力部15に「もう一度質問を繰り返してください」といったテキストデータおよび音声データのうちの少なくとも一方を出力し、出力部15がテキストおよび音声のうちの少なくとも一方を出力してもよい。また、施設が駅構内である場合、回答生成部14は、出力部15に「K線のホームでしょうか、またはL線のホームでしょうか?」といったようにホームを特定するためのテキストデータおよび音声データのうちの少なくとも一方を出力し、出力部15がテキストおよび音声のうちの少なくとも一方を出力してもよい。また、例えば、質問者30の人数が複数であるように認識されるものの質問者30が重なりあっており人数が定かでない場合には、「何名でご利用でしょうか?」といったように人数を特定するためのテキストデータおよび音声データのうちの少なくとも一方を出力し、出力部15がテキストおよび音声のうちの少なくとも一方を出力してもよい。
 次に、本実施の形態の自動応答システム10の動作について説明する。図5は、本実施の形態の自動応答システム10における自動応答処理手順の一例を示すフローチャートである。自動応答システム10は、施設データを取得する(ステップS1)。詳細には、データ取得部111が、施設データを取得し、前処理部112へ出力する。施設データは、上述したように静的なデータであってもよいし、動的なデータであってもよいし、これらの両方であってもよい。
 自動応答システム10は、施設データを記憶する(ステップS2)。詳細には、前処理部112が、施設データに前処理を施し、前処理後の施設データをデータ記憶部113に格納する。前処理は、例えば、上述したようにベクトル化であるが、回答生成部14がアクセスできるデータ形式への変換、または回答生成部14が高速にアクセスすることができるデータ形式への変換であればよく、これに限定されない。また、上述したように、前処理は行われなくてもよい。
 自動応答システム10は、質問内容を取得する(ステップS3)。詳細には、質問取得部12が、音声またはテキストデータとして質問者30からの質問を取得することで、質問内容を取得し、取得した質問を回答生成部14へ出力する。
 自動応答システム10は、状況を取得する(ステップS4)。詳細には、状況取得部13が、質問者30のおかれている状況を示す付帯情報を取得し、付帯情報を回答生成部14へ出力する。例えば、状況取得部13は、カメラをはじめとしたセンサによって取得されたセンサ情報を付帯情報として取得してもよいし、質問者30からの音声またはテキストデータによる入力により付帯情報を取得してもよい。
 自動応答システム10は、回答を生成する(ステップS5)。詳細には、回答生成部14が、付帯情報から状況(質問者30の状況)を把握し、状況に基づいて質問に対する回答を生成し、生成した回答を出力部15へ出力する。なお、上述したように、自動応答システム10は、付帯情報からさらに属性(質問者30の属性)を把握し、状況および属性に基づいて質問に対する回答を生成してもよい。
 自動応答システム10は、回答を出力し(ステップS6)、自動応答処理を終了する。詳細には、ステップS6では、出力部15が、回答を出力する。出力部15は、回答生成部14が生成した回答の形式(音声データであるが、テキストデータであるか、画像データであるか)に応じて、音声による出力と表示とのうちの少なくとも一方を行うことにより、回答を出力する。
 図6は、図5のステップS5に示した回答生成部14における回答生成処理手順の一例を示すフローチャートである。回答生成部14は、状況および質問を認識する(ステップS11)。詳細には、状況認識部141が付帯情報に基づいて状況を認識し、生成部142へ出力する。生成部142は、状況認識部141から取得した状況と質問取得部12から受け取った質問とを回答の生成処理に必要な形式に変換するなどにより認識する。また、例えば、回答生成部14にマルチモーダルAIまたは大規模言語モデルが用いられる場合、生成部142は、状況および質問の意味を理解する。具体的には、質問が「ホームはどこにありますか?」であれば、回答生成部14は、「ホームまでの道のりを聞かれている」と理解し、認識した状況が「大きな荷物を持っている」であれば、「なるべく階段を使わない経路がよい」と理解する。
 回答生成部14は、施設データを参照する(ステップS12)。詳細には、生成部142は、質問の回答に必要な施設データをデータ記憶部113から読み出す。例えば、上述した質問が「ホームはどこにありますか?」であり状況が「大きな荷物を持っている」の例の場合には、施設データとして駅構内の地図情報を読み出す。
 回答生成部14は、回答を生成し(ステップS13)、回答生成処理を終了する。詳細には、ステップS13では、生成部142は、参照した施設データを用いて、状況を考慮した質問への回答を生成し、生成した回答を出力部15へ出力する。例えば、上述した質問が「ホームはどこにありますか?」であり状況が「大きな荷物を持っている」の例の場合には、ホームまでの経路として駅構内の地図情報を用いて階段を用いないエレベータを利用する経路を探索し、探索に得られた経路を示す図と、当該図の説明文とを生成する。
 以上の処理により、本実施の形態の自動応答システム10は、質問者30の状況を考慮した回答を生成することができる。これにより、自動応答システム10は、単に、施設利用者からの特定の質問に回答する場合に比べて、施設利用者により適切な回答を提供することができる。また、自動応答システム10は、状況に加えてさらに属性に基づいて回答を生成してもよく、この場合も、属性に応じたより適切な回答を施設利用者に提供することができる。また、施設データとして動的なデータ(動的な施設データ)も含めておくことで、現在の施設の状態に基づいた適切な回答や情報を施設利用者に提供することができる。
 図7は、本実施の形態の動的な施設データの一例を示す図である。図7に示した例では、施設は、駅構内または複数の駅を含む鉄道施設(鉄道車両内を含む)である。図7に示した例では、質問者30が、「空いているのは何号車ですか?」と質問しており、この質問に対する回答を生成するためには、各車両の混雑の度合いを知る必要がある。図7に示した例では、データ取得部111は、鉄道車両の各車両内に設けられたカメラ50によって取得された映像データを、動的な施設データとして取得する。また、どの編成の鉄道車両がいつ駅に到着するかを示す時刻表も施設データとしてデータ記憶部113に格納されている。なお、動的な施設データは、例えば、新たなデータが取得されると最新のデータに更新されるが、これに限らず、過去の動的な施設データもデータ記憶部113に格納されてもよい。
 前処理部112は、車両内の映像データをデータ記憶部113に格納してもよいし、映像データに基づいて各車両の混雑の度合いを求め、求めた結果を施設データとしてデータ記憶部113に格納してもよい。混在の度合いは、例えば、混在しているか否かの2段階で示されてもよいし、混雑、普通、空いているの3段階で示されてもよいし、混雑率で示されてもよいし、これら以外で示されてもよい。映像データがデータ記憶部113に格納される場合、回答生成部14は、映像データに基づいて混雑の度合いを認識して回答の生成に利用する。
 例えば、図7に示した例では、回答生成部14は、時刻表に基づいて次に到着する鉄道車両を把握し、当該鉄道車両の各車両内を撮影した映像データから算出された混雑の度合いを用いることにより、空いている車両を把握することができる。これにより、「空いているのは何号車ですか?」という質問に対して、例えば、「2号車が空いています。」といったように、実際の状態が考慮された適切な回答を生成することができる。また、図7に示した例では、自動応答システム10が駅のホームに設置される例を示しており、回答生成部14は、2号車の位置と現在位置とを示す画像も回答として生成し、出力部15が、「2号車が空いています」というテキストによる回答とともに生成された画像も表示している。なお、この回答の際には、上述したように、さらに状況(または、状況および属性)が考慮されてもよい。例えば、図7に示した例において、質問者30が大きな荷物を持っていない場合には、空いている車両のうち最も空いている車両を案内するように回答を生成し、質問者30が大きな荷物を持っている場合には、空いている車両のうち最も近い車両(最も空いている車両でなくても)を案内するように回答を生成してもよい。
 また、動的な施設データは、図7に示した例に限らず、施設内のトイレの混雑の度合いを示す情報であってもよいし、店舗の混雑の度合いを示す情報であってもよいし、上述したように施設発信情報、店舗発信情報などであってもよい。
 次に、本実施の形態の自動応答システム10における状況を考慮した回答の具体例について説明する。質問者30のおかれた状況は、上述したように、例えば、質問者30の人数を含んでいてもよい。例えば、質問者30が「おすすめの飲食店を教えてください」と質問した場合に、質問者30が一人の場合には、自動応答システム10は、施設データに基づいて落ち着いた店舗を選択し、選択した店舗がおすすめであることを示す回答を生成し、回答を出力する。一方、質問者30が家族連れである場合には、自動応答システム10は、ファミリレストランなどファミリ向けの店舗を選択し、選択した店舗がおすすめであることを示す回答を生成し、回答を出力する。
 図8~図14は、本実施の形態の自動応答システム10における状況を考慮した回答の具体例を示す図である。図8に示した例では、装飾品を身に着けたエレガントな服装の中高年の女性が、「どこか洋服屋さんは知りませんか?」と質問している。自動応答システム10は、質問者30のおかれた状況として装飾品を身に着けており、エレガントな服装であるということを認識し、この認識結果と、質問と、施設データにおける各店舗の扱う商品を示す情報および各店舗のターゲットを示す情報とに基づいて、中高年におすすめの洋品店は、AショップおよびBショップであることを把握する。そして、自動応答システム10は、質問への回答として、AショップおよびBショップを案内する情報を生成し、出力部15に表示する。これにより、質問者30の状況が考慮された回答が生成される。
 また、図8に示した例では、自動応答システム10は、さらに、回答の生成に用いた状況に基づいて、回答の根拠を示す根拠情報を含む回答を生成している。図8に示した例では、自動応答システム10は、AショップおよびBショップを案内することに決めた根拠として、装飾品を身に着けたエレガントな服装の中高年の女性であるという状況から、「お母さん世代」という根拠情報を生成する。これにより、図8に示した例では、「・Aショップ」および「・Bショップ」という提案する店舗名だけでなく、「お母さん世代に人気なオシャレな洋服店はどうでしょうか?」というテキストが表示される。また、図8に示した表示画面において、「・Aショップ」および「・Bショップ」の部分をタッチすると、各店舗の詳細な情報が表示されるようにしてもよい。なお、回答における文章(テキスト)の具体的な内容は、図8に示した例に限定されない。
 図9に示した例では、制服を着た若い女性が、「どこか洋服屋さんは知りませんか?」と質問している。自動応答システム10は、質問者30のおかれた状況として、制服を着た若い女性であることを認識し、この認識結果と、質問と、施設データにおける各店舗の扱う商品を示す情報および各店舗のターゲットを示す情報とに基づいて、女子高校生(女子高生)におすすめの洋品店は、CショップおよびDショップであることを把握する。そして、自動応答システム10は、質問への回答として、CショップおよびDショップを案内する情報を生成し、出力部15に表示する。図9に示した例では、図8に示した例と同様に、根拠情報も表示されている。具体的には、根拠情報として「女子高生」というテキストが生成され、根拠情報を含む回答が表示される。
 図10に示した例では、クラシックな帽子をかぶりクラシックな服装の中高年の男性が、「どこか洋服屋さんは知りませんか?」と質問している。自動応答システム10は、質問者30のおかれた状況として、クラシックな帽子をかぶりクラシックな服装の中高年の男性であることを認識し、この認識結果と、質問と、施設データにおける各店舗の扱う商品を示す情報および各店舗のターゲットを示す情報とに基づいて、クラシックなファッションを好む中高年の男性におすすめの洋品店は、Eショップであることを把握する。そして、自動応答システム10は、質問への回答として、Eショップを案内する情報を生成し、出力部15に表示する。図10に示した例では、図8および図9に示した例と同様に、根拠情報も表示されている。具体的には、根拠情報として「ダンディな」というテキストが生成され、根拠情報を含む回答が表示される。なお、図8、図9および図10に示した例では、質問者30の性別も考慮されているが、質問者30の性別は属性でもあり、これらの例は、服装などの状況と属性との両方を考慮した回答の生成ともいえる。
 図11および図12は、質問者30のおかれた状況として質問者30の表情または感情が考慮される例を示している。図11および図12に示した例では、いずれも、質問者30は、ホームがどこであるかを尋ねている。
 図11に示した例では、質問者30は、落ち着いた通常の様子で「ホームはどこですか?」と質問している。自動応答システム10は、質問者30の表情、質問者30の動き、声のトーンなどから質問者30は通常であるすなわち冷静であると認識し、質問者30と対話を行いながら回答を生成する。図11に示した例では、自動応答システム10は、「どこ行きの電車に乗りたいですか?」と質問者30に問いかけ、質問者30が「XX方面に行きたいです」と回答し、自動応答システム10は、この質問者30からの回答に基づいて「XX方面は3番ホームにお進みください」という回答を生成して出力している。
 一方、図12に示した例では、質問者30は、慌てた様子で「ホームはどこ?」と質問している。自動応答システム10は、質問者30の表情、質問者30の動き、声のトーンなどから質問者30が焦っており気持ちに余裕のない状況と認識し、質問者30との対話は行わずに、複数種類の情報を含む回答を一度に出力する。図12に示した例では、自動応答システム10は、回答として、「・XX方面 3番ホーム」、「・YY方面 4番ホーム」および「・ZZ方面 6番ホーム」という情報を一度に表示している。図11および図12に例示したように、質問者30の表情または感情に応じて、回答の内容、質問者30とのやりとりの方法などを決めてもよい。
 図13および図14は、質問者30のおかれた状況として、質問者30の表情または感情と人数とが考慮される例を示している。図13に示した例では、質問者30は4人であり、肩を組み、話をしながら楽しそうにしており、「どこか飲み屋を教えてください」と質問している。自動応答システム10は、質問者30の表情、質問者30の動き、声のトーンなどから質問者30がテンションの高い状態(気分が良く盛り上がっている状態)と認識するとともに質問者30が集団であると認識し、にぎやかな店舗である「A屋」および「B飲み屋」を案内する回答を生成する。また、図13に示した例では、図8、図9および図10に示した例と同様に根拠情報も回答に含まれており、「テンション高い皆様には、にぎやかで飲み放題があるお店は、どうでしょうか?」という情報も出力部15に提示される。
 図14に示した例では、質問者30は手を繋いだ男女であり、「どこか飲み屋を教えてください」と質問している。自動応答システム10は、質問者30の表情、質問者30の動き、声のトーンなどから質問者30が男女のカップルであると認識し、落ち着いたバーである「Cバー」および「Dバー」を案内する回答を生成する。また、図14に示した例では、図8、図9および図10に示した例と同様に根拠情報も回答に含まれており、「カップルのお二人には、落ち着いたバーはどうでしょうか?」という情報も出力部15に提示される。
 以上に示した具体例は、例示であり、自動応答システム10における状況を考慮した回答はこれらの例に限定されない。また、図8、図9、図10および図14では、性別を考慮した回答が生成されたが、これに限らず、男性、女性の性別で状況を認識せず、回答にも性別に関する内容を含めないようにしてもよい。すなわち、回答生成部14は、性別に依存しない回答を生成してもよい。これにより、自動応答システム10は、LGBTQ(Lesbian,Gay,Bisexual,Transgender,Queer/Questioning)に配慮した回答を出力することができる。
 また、自動応答システム10は、LGBTQに配慮した回答として双方の性別を区別しない回答、双方の性別を考慮する回答などを生成することとしてもよい。例えば、トイレを案内する際にユーザーの性別によらず男子トイレと女子トイレのある場所をそれぞれ回答するなどして、双方の性別を含む回答をしてもよい。また、LGBTQに配慮した回答を生成するためのルールを、施設データとして、または施設データとは別に、データ記憶部113に格納しておき、回答生成部14が回答を作成する際に当該ルールを参照するようにしてもよい。また、LGBTQへの配慮に限らず、回答に反映させる倫理規範をルールとして定め、当該ルールを、施設データとして、または施設データとは別に、データ記憶部113に格納しておき、回答生成部14が回答を作成する際に当該ルールを参照するようにしてもよい。また、回答生成部14に生成AIが用いられる場合、生成AIにこのルールがあらかじめ指示されてもよい。
 次に、本実施の形態の自動応答システム10のハードウェア構成について説明する。本実施の形態の自動応答システム10における情報処理部20は、コンピュータシステム上で、情報処理部20のそれぞれにおける処理が記述されたコンピュータプログラムであるプログラムが実行されることにより、コンピュータシステムがそれぞれ情報処理部20として機能する。図15は、本実施の形態の情報処理部20のそれぞれを実現するコンピュータシステムの構成例を示す図である。図15に示すように、このコンピュータシステムは、制御部101と入力部102と記憶部103と表示部104と通信部105と出力部106とを備え、これらはシステムバス107を介して接続されている。
 図15において、制御部101は、例えば、CPU(Central Processing Unit)等のプロセッサであり、本実施の形態の情報処理部20における処理が記述されたプログラムを実行する。入力部102は、例えばキーボード、ボタン、マウスなどで構成され、コンピュータシステムの使用者が、各種情報の入力を行うために使用する。記憶部103は、RAM(Random Access Memory),ROM(Read Only Memory)などの各種メモリおよびハードディスクなどのストレージデバイスを含み、上記制御部101が実行すべきプログラム、処理の過程で得られた必要なデータ、などを記憶する。また、記憶部103は、プログラムの一時的な記憶領域としても使用される。制御部101および記憶部103は、例えば、処理回路を構成する。処理回路は、1つの回路であってもよいし複数の回路であってもよい。表示部104は、ディスプレイ、LCD(Liquid Crystal Display:液晶表示パネル)などで構成され、コンピュータシステムの使用者に対して各種画面を表示する。なお、入力部102と表示部104とが一体化されたタッチパネルが用いられてもよい。通信部105は、通信処理を実施する受信機および送信機である。出力部106は、スピーカなどである。なお、図15は、一例であり、情報処理部20のそれぞれを実現するコンピュータシステムの構成は図15に示した例に限定されない。例えば、出力部106が設けられていなくてもよい。
 ここで、本実施の形態のプログラムが実行可能な状態になるまでのコンピュータシステムの動作例について説明する。上述した構成をとるコンピュータシステムには、例えば、図示しないCD(Compact Disc)-ROMドライブまたはDVD(Digital Versatile Disc)-ROMドライブにセットされたCD-ROMまたはDVD-ROMから、プログラムが記憶部103にインストールされる。そして、プログラムの実行時に、記憶部103から読み出されたプログラムが記憶部103の主記憶領域に格納される。この状態で、制御部101は、記憶部103に格納されたプログラムにしたがって、本実施の形態の情報処理部20のそれぞれとしての処理を実行する。
 なお、上記の説明においては、CD-ROMまたはDVD-ROMを記録媒体として、情報処理部20のそれぞれにおける処理を記述したプログラムを提供しているが、これに限らず、コンピュータシステムの構成、提供するプログラムの容量などに応じて、例えば、通信部105を経由してインターネットなどの伝送媒体により提供されたプログラムを用いることとしてもよい。
 本実施の形態のプログラムは、例えば、コンピュータシステムに、施設の利用に関する質問を施設の利用者である施設利用者から取得するステップと、質問した施設利用者である質問者の状況を示す付帯情報を取得するステップと、付帯情報と質問と施設に関するデータである施設データとを用いて、質問への回答を生成するステップと、回答を出力するステップと、を実行させる。
 図1に示した前処理部112および回答生成部14は、図15に示した記憶部103に記憶されたプログラムが図15に示した制御部101により実行されることにより実現される。また、前処理部112および回答生成部14の実現には、記憶部103も用いられる。図1に示したデータ取得部111は、図15に示した通信部105および入力部102のうち少なくとも一方により実現される。データ取得部111の一部の機能は、制御部101および記憶部103により実現されてもよい。図1に示したデータ記憶部113は、図15に示した記憶部103の一部である。情報処理部20は、複数のコンピュータシステムにより実現されてもよい。例えば、情報処理部20は、クラウドコンピュータシステムにより実現されてもよい。
 また、本実施の形態の質問取得部12は、上述したように、例えば、タッチパネル、キーボードなどの入力手段とマイクとのうちの少なくとも一方により実現される。本実施の形態の状況取得部13は、上述したように、例えば、カメラなどのセンサと、タッチパネル、キーボードなどの入力手段とマイクとのうちの少なくとも1つにより実現される。出力部15は、例えば、タッチパネルをはじめとしたディスプレイと、スピーカとのうちの少なくとも1つにより実現される。上述したように、出力部15を実現するティスプレイを含み、当該ディスプレイがタッチパネルである場合には、タッチパネルが、質問取得部12の機能を兼ねていてもよい。さらに、このタッチパネルが、状況取得部13の機能を兼ねていてもよい。なお、質問取得部12および状況取得部13が処理を行う場合には、これらの実現に上述したコンピュータシステムにおける制御部101および記憶部103が用いられてもよい。なお、質問取得部12、状況取得部13、情報処理部20および出力部15が一体化されている場合、自動応答システム10全体を図15に例示したコンピュータシステムとみなしてもよい。この場合、例えば、コンピュータシステムは、入力部102としてマイクを備えるとともに、さらに、状況取得部13としてのセンサを備えてもよい。この場合、例えば、質問取得部12は入力部102により実現され、状況取得部13はセンサおよび入力部102のうちの少なくとも一方により実現され、出力部15は、表示部104および出力部106のうちの少なくとも一方により実現される。
 以上のように、本実施の形態の自動応答システム10は、質問者30の状況を用いて、質問者30からの質問の回答を生成するようにした。単に、施設利用者からの特定の質問に回答する場合に比べて、施設利用者により適切な回答を提供することができる。また、自動応答システム10は、状況に加えてさらに属性に基づいて回答を生成してもよく、この場合も、属性に応じたより適切な回答を施設利用者に提供することができる。また、施設データとして動的な施設データも含めておくことで、現在の施設の状態に基づいた適切な回答や情報を施設利用者に提供することができる。
実施の形態2.
 図16は、実施の形態2にかかる自動応答システムの構成例を示す図である。本実施の形態の自動応答システム10aは、自動応答装置20aと、端末装置60とを備える。自動応答装置20aは、送受信部16が追加される以外は、実施の形態1の情報処理部20と同様である。端末装置60は、質問取得部12、状況取得部13、出力部15および送受信部17を備える。実施の形態1と同一の機能を有する構成要素は、実施の形態1と同一の符号を付して重複する説明を省略する。以下、実施の形態1と異なる点を主に説明する。
 実施の形態1では、主に、施設内に存在する施設利用者が施設に設置されている自動応答システム10を利用する例を説明した。本実施の形態では、施設をこれから利用しようとする利用者も施設利用者として扱う。端末装置60は、施設利用者が操作可能な装置であり、例えば、スマートフォン、タブレット、パーソナルコンピュータなどである。端末装置60は、施設利用者が携帯可能な携帯端末であってもよい。施設利用者は、例えば、自宅、施設へ向かう途中などに端末装置60を用いることにより自動応答システム10aを利用する。なお、施設利用者が自動応答システム10aを利用する場所は、これに限定されず、任意の場所であってよく、施設内であってもよい。
 端末装置60の質問取得部12は、実施の形態1の質問取得部12と同様に、施設利用者からの施設に関する質問を、音声およびテキストのうちの少なくとも一方により取得する。なお、実施の形態1と同様に、質問に画像が含まれていてもよい。質問取得部12は、取得した質問を送受信部17へ出力する。質問取得部12は、例えば、スマートフォン、タブレット、パーソナルコンピュータなどのタッチパネル、キーボードなどの入力手段と、マイクとのうちの少なくとも一方である。質問に画像を含める場合には、例えば、端末装置60に記憶されている画像データを施設利用者が指定し、入力手段が指定を受付ける。なお、質問が音声として入力された場合に、質問取得部12が音声データをテキストデータに変換して、送受信部17へ出力してもよいし、音声データのまま送受信部17へ出力してもよい。
 状況取得部13は、実施の形態1の状況取得部13と同様に、質問する施設利用者である質問者30の状況(質問者のおかれた状況)を示す付帯情報を取得する。実施の形態1と同様に、付帯情報は、さらに施設利用者の属性を示す情報であってもよい。状況取得部13は、音声およびテキストのうちの少なくとも一方により付帯情報を取得してもよいし、質問者30を撮影した撮影データを付帯情報として取得してもよい。また、例えば、質問者30へ、施設を利用する際の、人数、補助具の利用の有無などを問うための画面を出力部15が出力し、その画面において入力された回答を付帯情報として取得してもよい。状況取得部13は、例えば、端末装置60であるスマートフォン、タブレット、パーソナルコンピュータなどに内蔵されたカメラにより実現され、カメラによって取得された映像データを付帯情報として送受信部17へ出力してもよい。このカメラは、例えば、端末装置60の内側から質問者を撮影するインカメラである。
 出力部15は、実施の形態1の出力部15と同様に、音声による出力と表示とのうちの少なくとも一方により回答を出力する。出力部15は、例えば、端末装置60のディスプレイと、端末装置60のマイクとのうちの少なくとも一方であるが、これらに限定されない。端末装置60のディスプレイはタッチパネルであってもよい。この場合、タッチパネルは、質問取得部12および状況取得部13のうちの少なくとも一方の機能を兼ねていてもよい。
 送受信部17は、自動応答装置20aとの間で通信を行うことにより、自動応答装置20aとの間で情報のやりとりを行う。例えば、送受信部17は、質問取得部12から受け取った質問と、状況取得部13から受け取った付帯情報とを自動応答装置20aへ送信する。また、送受信部17は、自動応答装置20aから受信した回答を出力部15へ出力する。
 上記の質問取得部12、状況取得部13、出力部15および送受信部17の動作は、端末装置60に、施設の案内のサービスを提供するアプリケーションソフトウェアがインストールされることにより行われてもよいし、端末装置60が自動応答装置20aにアクセスすることで行われてもよい。例えば、自動応答装置20aがWebサーバとしての機能を有し、端末装置60がWebサーバへアクセスすることで、上記の動作が実現されてもよい。
 自動応答装置20aの送受信部16は、端末装置60との間で通信を行うことにより、端末装置60との間で情報のやりとりを行う。例えば、送受信部16は、端末装置60から受信した質問および付帯情報を回答生成部14へ出力する。また、送受信部16は、回答生成部14から受け取った回答を端末装置60へ出力する。回答生成部14は、実施の形態1の回答生成部14と同様に、質問、付帯情報および施設データを用いて回答を生成し、生成した回答を送受信部16へ出力する。本実施の形態の自動応答装置20aは、実施の形態1の情報処理部20と同様に、コンピュータシステムにより実現される。なお、自動応答装置20aは施設内に設置されていてもよいし、施設外に設置されてもよい。
 本実施の形態では、端末装置60を用いて施設利用者が質問することができるため、施設利用者の位置にかかわらず、自動応答システム10aを利用することができる。このため、例えば、施設を利用する前に、事前に、自宅または施設へ向かう途中で施設に関する情報を得ることができる。このため、施設利用者は、事前に得た情報を基に効率的に施設を利用することができる。また、施設データに、施設への経路を生成するための地図情報を含めておくことで、施設利用者が施設までの行き方を質問した場合に、施設利用者の現在位置から施設への経路を回答として提示することも可能である。また、例えば、施設データとして届け出のあった忘れ物のデータを含めておいてもよい。これにより、施設利用者が、施設を利用した後に施設に忘れ物をした可能性がある場合に、どこに何を忘れたかを示して忘れ物があるかを質問すると、自動応答システム10aは、該当する忘れ物の届け出があるか否かを回答することができる。
 図17および図18は、本実施の形態の自動応答システム10aにおける状況を考慮した回答の具体例を示す図である。図17および図18に示した例では、いずれも施設の周辺では雨が降っており、施設データに天候を示すデータが含まれているとする。
 図17に示した例では、施設がデパートであり、デパートを利用する質問者30がデパートに到着する前に、端末装置60を用いて質問を行う例を示している。図17に示した例では、雨が降っており、質問者30は傘をさしている。端末装置60は、「デパートまでの道を教えてください」という音声データである質問を受付け、質問と質問者30を撮影した映像データである付帯情報とを自動応答装置20aへ送信する。自動応答装置20aは、質問と付帯情報と施設データとを用いて、雨が降っていても質問者30が傘をさしているため通常の経路で案内すると判断し、デパートまでの通常の経路(例えば、最短経路)を回答として生成し、回答を端末装置60へ送信する。端末装置60は、回答としてデパートまでの経路を表示する。なお、図17では、質問者30が傘をさしているが、これに限らず、質問者30が閉じた傘を手にもっている場合も、同様に、自動応答装置20aは、雨が降っていても通常の経路で案内するように回答を生成してもよい。すなわち、自動応答装置20aは、傘の状態にかかわらず、質問者30が傘を持っていれば、通常の経路で案内するように回答を生成するようにしてもよい。
 図18に示した例では、施設は会議場であり、説明会が予定されている。図18に示した例では、質問者30は正装をしており傘を持っていない。端末装置60は、「説明会会場までの道を教えてください」という音声データである質問を受付け、質問と質問者30を撮影した映像データである付帯情報とを自動応答装置20aへ送信する。自動応答装置20aは、質問と付帯情報と施設データとを用いて、雨が降っており質問者30が傘をさしておらず正装であることから、屋根のある雨に濡れない経路で案内すると判断し、雨に濡れない経路を回答として生成し、回答を端末装置60へ送信する。端末装置60は、回答として会議場までの経路を表示する。図17および図18に例示したように、質問者30の状況に応じて、施設までの経路を決定することで、自動応答装置20aは、質問者30に適した回答を提供することができる。
 以上のように、本実施の形態では、施設利用者が端末装置60を用いて質問を行い、端末装置60が回答を出力するようにした。このため、施設利用者は、施設外においても、自動応答システム10aを利用することができる。
 また、実施の形態1に示した自動応答システム10に送受信部16を追加し、送受信部16が端末装置60と通信を行うことで、実施の形態1で述べた施設内の施設利用者からの質問と、本実施の形態で述べた端末装置60を利用した質問との両方に対応できるようにしてもよい。
実施の形態3.
 図19は、実施の形態3にかかる自動応答システムの構成例を示す図である。本実施の形態の自動応答システム10bは、評価取得部18が追加され、データ記憶部113がさらに評価情報を記憶する以外は、実施の形態1の自動応答システム10と同様である。実施の形態1と同一の機能を有する構成要素は、実施の形態1と同一の符号を付して重複する説明を省略する。以下、実施の形態1と異なる点を主に説明する。
 本実施の形態の評価取得部18は、質問者30とのやりとりの終了後、質問者30から回答の評価を示す評価結果を取得し、取得した評価結果を、対応する質問および回答とともにデータ記憶部113に格納する。なお、ここでは、評価情報がデータ記憶部113に格納される例を説明するが、これに限らず、データ記憶部113とは別に評価情報を記憶する評価情報記憶部が設けられてもよい。例えば、出力部15が、回答の出力後に、「お客様の知りたい情報は得られましたか?」、「回答は適切でしたか?」などのように、自動応答システム10bの回答への評価を問う質問を出力する。また、質問者30からの評価が不満足、よくないなど否定的な言葉を含むものであった場合には、評価取得部18は、どのような点が問題であったかをさらに質問者30へ問う質問を出力部15に出力させ、質問者30からの入力を受付けてもよい。評価取得部18は、評価結果を、音声により取得してもよいし、テキストデータとして取得してもよい。また、出力部15が、評価結果を示す選択肢を提示し、評価取得部18が、質問者30の選択結果を評価結果として取得してもよい。
 また、評価取得部18は、質問者30の行動から回答内容を評価してもよい。例えば、自動応答システム10bが、道を聞いた質問者30に対して目的地までの経路を出力後、質問者30を撮影した映像データを用いて質問者30の行動を分析し、質問者30がまだ道に迷っている状況であれば、回答が分かりづらかったと自らの評価を行ってもよい。
 評価情報は、例えば、対応する質問および回答とともに回答生成部14の再学習、追加学習などに用いられてもよいし、回答生成部14が回答を生成する際に参照してもよい。または、自動応答システム10bの運用者、管理者などが、評価情報を確認し、確認した結果を用いて改善点を把握し、回答生成部14に把握した改善点を指示してもよい。または、自動応答システム10bの運用者、管理者などが、評価情報を確認し、確認した結果を用いて、回答生成部14に対して、どのような再学習または追加学習を行うかを決定し、決定した結果を用いて再学習または追加学習が行われてもよい。
 評価取得部18は、質問取得部12と同様に、例えば、操作による入力手段とマイクとのうちの少なくとも一方である。入力手段、マイクなどのハードウェアは、質問取得部12と共用されてもよい。
 以上のように、本実施の形態では、質問者30からの評価を取得し、回答に反映させるため、回答の精度を高めることができる。なお、実施の形態2の自動応答システム10aに評価取得部18を追加し、評価結果を反映するようにしてもよいし、実施の形態1と実施の形態2とを組み合わせた自動応答システムに、評価取得部18を追加し、評価結果を反映するようにしてもよい。
 以上の実施の形態に示した構成は、一例を示すものであり、別の公知の技術と組み合わせることも可能であるし、実施の形態同士を組み合わせることも可能であるし、要旨を逸脱しない範囲で、構成の一部を省略、変更することも可能である。
 10,10a,10b 自動応答システム、11 データ管理部、12 質問取得部、13 状況取得部、14 回答生成部、15 出力部、16,17 送受信部、18 評価取得部、20 情報処理部、20a 自動応答装置、30 質問者、40 本体、50 カメラ、60 端末装置、111 データ取得部、112 前処理部、113 データ記憶部、141 状況認識部、142 生成部。

Claims (22)

  1.  施設の利用に関する質問を施設の利用者である施設利用者から取得する質問取得部と、
     質問した前記施設利用者である質問者の状況を示す付帯情報を取得する状況取得部と、
     前記質問と前記付帯情報と施設に関するデータである施設データとを用いて、前記質問への回答を生成する回答生成部と、
     前記回答を出力する出力部と、
     を備えることを特徴とする自動応答システム。
  2.  前記回答生成部は、音声、テキストおよび画像を含む複数の形式のデータを入力として用いて前記回答を生成することが可能であり、前記回答を複数の形式で出力することが可能であることを特徴とする請求項1に記載の自動応答システム。
  3.  前記状況は、前記質問者の身に着けているもの、前記質問者の補助具の利用の有無、前記質問者の人数、前記質問者の感情、および前記質問者が持っている荷物の大きさ、のうちの少なくとも1つを含むことを特徴とする請求項1または2に記載の自動応答システム。
  4.  前記付帯情報は、前記質問者を撮影した映像データを含むことを特徴とする請求項1から3のいずれか1つに記載の自動応答システム。
  5.  前記施設データは、一定時間変化しない静的なデータと、前記一定時間の間にも変化する可能性がある動的なデータとを含むことを特徴とする請求項1から4のいずれか1つに記載の自動応答システム。
  6.  前記動的なデータは、前記施設内および前記施設の周辺のうち少なくとも一方に設置されたセンサによって取得されたデータを含むことを特徴とする請求項5に記載の自動応答システム。
  7.  前記動的なデータは、前記施設の管理者および前記施設の従業員のうち少なくとも一方によって発信される施設発信情報を含むことを特徴とする請求項5または6に記載の自動応答システム。
  8.  前記動的なデータは、前記施設内の店舗の従業員によって発信される店舗発信情報を含むことを特徴とする請求項5から7のいずれか1つに記載の自動応答システム。
  9.  前記動的なデータは、気象情報を含むことを特徴とする請求項5から8のいずれか1つに記載の自動応答システム。
  10.  前記回答生成部は、
     前記付帯情報に基づき前記状況を認識する状況認識部と、
     前記状況と前記質問と施設データとを用いて、前記質問への前記回答を生成する生成部と、
     を備えることを特徴とする請求項1から9のいずれか1つに記載の自動応答システム。
  11.  前記付帯情報は、さらに前記質問者の属性を示し、
     前記状況認識部は、前記付帯情報に基づき前記状況および前記属性を認識し、
     前記生成部は、前記質問と前記状況と前記属性と施設データとを用いて、前記質問への前記回答を生成することを特徴とする請求項10に記載の自動応答システム。
  12.  前記属性は、性別、年齢、年代、障がいの有無および職業のうちの少なくとも1つを含むことを特徴とする請求項11に記載の自動応答システム。
  13.  前記施設利用者は、あらかじめ定められた特定の人を含み、
     前記施設データは、前記特定の人の個人ごとの前記施設の利用に関する情報である個別情報を含み、
     前記状況取得部は、個人を識別するための個人認証情報を取得し、
     前記回答生成部は、前記個人認証情報を用いて個人を特定し、特定した結果に基づき、対応する前記個別情報を用いて前記回答を生成することを特徴とする請求項11または12に記載の自動応答システム。
  14.  前記属性は、身分、所属および役職のうちの少なくとも1つを含み、
     前記施設データは、前記属性ごとの前記施設の利用に関する情報である属性関連情報と特定の人の個人ごとの前記施設の利用に関する情報である個別情報とを含み、
     前記個別情報は対応する個人の前記属性を含み、
     前記状況取得部は、個人を識別するための個人認証情報を取得し、
     前記回答生成部は、前記個人認証情報を用いて個人を特定し、特定した結果と前記個別情報とに基づき、特定した個人の前記属性を認識し、前記属性に対応する前記属性関連情報を用いて前記回答を生成することを特徴とする請求項12または13に記載の自動応答システム。
  15.  前記回答生成部は、前記回答の生成に用いた前記状況に基づいて、前記回答の根拠を示す根拠情報を含む前記回答を生成することを特徴とする請求項1から14のいずれか1つに記載の自動応答システム。
  16.  前記施設データを、前記回答生成部がアクセス可能な形式のデータに変換する前処理部と、
     前記前処理部による前記変換が行われた後の前記施設データを記憶するデータ記憶部と、
     を備え、
     前記回答生成部は、前記データ記憶部に記憶されている前記施設データを用いて前記回答を生成することを特徴とする請求項1から15のいずれか1つに記載の自動応答システム。
  17.  前記回答生成部は、性別に依存しない前記回答を生成することを特徴とする請求項1から16のいずれか1つに記載の自動応答システム。
  18.  前記回答生成部は、質問を認識できない場合、および回答の生成のために不足する情報がある場合に、前記質問者に問い直しを行う回答を生成することを特徴とする請求項1から17のいずれか1つに記載の自動応答システム。
  19.  前記回答が出力された後に、前記質問者から前記回答の評価結果を取得する評価取得部、
     を備えることを特徴とする請求項1から18のいずれか1つに記載の自動応答システム。
  20.  前記質問取得部、前記状況取得部および前記出力部を備え、前記質問者が操作可能な端末装置と、
     前記回答生成部を備える自動応答装置と、
     を備え、
     前記端末装置は、前記付帯情報および前記質問を前記自動応答装置へ送信し、
     前記自動応答装置は、前記施設データと、前記端末装置から受信した前記付帯情報および前記質問とを用いて前記回答を生成し、生成した前記回答を前記端末装置へ送信し、
     前記端末装置は、前記回答を出力することを特徴とする請求項1から19のいずれか1つに記載の自動応答システム。
  21.  自動応答システムにおける自動応答方法であって、
     施設の利用に関する質問を施設の利用者である施設利用者から取得するステップと、
     質問した前記施設利用者である質問者の状況を示す付帯情報を取得するステップと、
     前記質問と前記付帯情報と施設に関するデータである施設データとを用いて、前記質問への回答を生成するステップと、
     前記回答を出力するステップと、
     を含むことを特徴とする自動応答方法。
  22.  コンピュータシステムに、
     施設の利用に関する質問を施設の利用者である施設利用者から取得するステップと、
     質問した前記施設利用者である質問者の状況を示す付帯情報を取得するステップと、
     前記質問と前記付帯情報と施設に関するデータである施設データとを用いて、前記質問への回答を生成するステップと、
     前記回答を出力するステップと、
     を実行させることを特徴とするプログラム。
PCT/JP2024/015272 2024-04-17 2024-04-17 自動応答システム、自動応答方法およびプログラム Pending WO2025220150A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2024559956A JP7721018B1 (ja) 2024-04-17 2024-04-17 自動応答システム、自動応答方法およびプログラム
PCT/JP2024/015272 WO2025220150A1 (ja) 2024-04-17 2024-04-17 自動応答システム、自動応答方法およびプログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2024/015272 WO2025220150A1 (ja) 2024-04-17 2024-04-17 自動応答システム、自動応答方法およびプログラム

Publications (1)

Publication Number Publication Date
WO2025220150A1 true WO2025220150A1 (ja) 2025-10-23

Family

ID=96656972

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2024/015272 Pending WO2025220150A1 (ja) 2024-04-17 2024-04-17 自動応答システム、自動応答方法およびプログラム

Country Status (2)

Country Link
JP (1) JP7721018B1 (ja)
WO (1) WO2025220150A1 (ja)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2017220181A (ja) * 2016-06-10 2017-12-14 株式会社大林組 案内表示システム、案内表示方法及び案内表示プログラム
JP7411303B1 (ja) * 2023-05-01 2024-01-11 株式会社大正スカイビル 対象物の管理システム

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP7088270B2 (ja) * 2020-11-30 2022-06-21 凸版印刷株式会社 質問応答システム、及び質問応答方法

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2017220181A (ja) * 2016-06-10 2017-12-14 株式会社大林組 案内表示システム、案内表示方法及び案内表示プログラム
JP7411303B1 (ja) * 2023-05-01 2024-01-11 株式会社大正スカイビル 対象物の管理システム

Also Published As

Publication number Publication date
JPWO2025220150A1 (ja) 2025-10-23
JP7721018B1 (ja) 2025-08-08

Similar Documents

Publication Publication Date Title
JP7297326B2 (ja) 情報処理装置、情報処理システム、情報処理方法およびプログラム
Lee et al. The emerging professional practice of remote sighted assistance for people with visual impairments
Sobnath et al. Smart cities to improve mobility and quality of life of the visually impaired
Kanda et al. An affective guide robot in a shopping mall
Cameron et al. Enabling young people with a learning disability to make choices at a time of transition
Bermea et al. Resiliency and adolescent motherhood in the context of residential foster care
JP7207425B2 (ja) 対話装置、対話システムおよび対話プログラム
JP2000029932A (ja) 利用者検知機能を用いた情報案内方法及び利用者検知機能を有する情報案内システム及び情報案内プログラムを格納した記憶媒体
CN109643313A (zh) 信息处理设备、信息处理方法和程序
Tsui et al. Designing speech-based interfaces for telepresence robots for people with disabilities
WO2002057896A2 (en) Interactive virtual assistant
Samim A new paradigm of artificial intelligence to disabilities
Langedijk et al. Persuasive robots in the field
Görland et al. Without it, you will die
JP7721018B1 (ja) 自動応答システム、自動応答方法およびプログラム
US20220297308A1 (en) Control device, control method, and control system
Poli et al. The Contribution of Advanced Technologies to the Tourism Experience of Disabled People: The Greek Case
Macik et al. Smartphoneless context-aware indoor navigation
Kurosu Human-Computer Interaction. User Experience and Behavior: Thematic Area, HCI 2022, Held as Part of the 24th HCI International Conference, HCII 2022, Virtual Event, June 26–July 1, 2022, Proceedings, Part III
Jeffares Robots and virtual agents in frontline public service
Croft et al. Say what you mean, mean what you say-An ethnographic approach to male and female conversations
JP2020032529A (ja) 誘導サービスシステム、誘導サービスプログラム、誘導サービス方法および誘導サービス装置
US20260049835A1 (en) System
Petrova et al. Implementation of audio navigation for smart campus
KR102349665B1 (ko) 사용자 맞춤형 목적지정보 제공 장치 및 방법

Legal Events

Date Code Title Description
ENP Entry into the national phase

Ref document number: 2024559956

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 2024559956

Country of ref document: JP

121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24935955

Country of ref document: EP

Kind code of ref document: A1