WO2025224847A1 - 情報提示装置、情報提示方法、および、情報提示プログラム - Google Patents

情報提示装置、情報提示方法、および、情報提示プログラム

Info

Publication number
WO2025224847A1
WO2025224847A1 PCT/JP2024/015943 JP2024015943W WO2025224847A1 WO 2025224847 A1 WO2025224847 A1 WO 2025224847A1 JP 2024015943 W JP2024015943 W JP 2024015943W WO 2025224847 A1 WO2025224847 A1 WO 2025224847A1
Authority
WO
WIPO (PCT)
Prior art keywords
information
user
words
objects
verbal
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/JP2024/015943
Other languages
English (en)
French (fr)
Inventor
リドウィナ アユ アンダリニ
陽子 石井
徹也 山口
篤 深山
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
NTT Inc USA
Original Assignee
Nippon Telegraph and Telephone Corp
NTT Inc USA
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp, NTT Inc USA filed Critical Nippon Telegraph and Telephone Corp
Priority to PCT/JP2024/015943 priority Critical patent/WO2025224847A1/ja
Publication of WO2025224847A1 publication Critical patent/WO2025224847A1/ja
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/40Processing or translation of natural language
    • G06F40/55Rule-based translation
    • G06F40/56Natural language generation

Definitions

  • the present invention relates to an information presentation device, an information presentation method, and an information presentation program used to generate utterances for a digital human agent.
  • Gemini A Family of Highly Capable Multimodal Models, Gemini Team, Google, [online], [Retrieved April 9, 2024], Internet, ⁇ URL: https://arxiv.org/pdf/2312.11805.pdf> YOLO-World, Tianheng Cheng et al., [online], [Retrieved April 9, 2024], Internet, ⁇ URL: https://www.yoloworld.cc/> SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations, Satwik Kottur et al., [online], [Retrieved April 9, 2024], Internet, ⁇ https://arxiv.org/pdf/2104.08667.pdf> MediaGnosis, [online], [Retrieved April 9, 2024], Internet, ⁇ https://www.rd.ntt/mediagnosis/>
  • the objective of this invention is to solve the above-mentioned problems and provide high-quality dialogue even for users who are not good at self-disclosure.
  • the present invention is characterized by comprising a first information creation unit that acquires non-verbal information about a first user or the first user's surroundings, extracts characteristic objects from the acquired non-verbal information, and creates first information that shows the relationships between the objects in a graph or vector form; a second information creation unit that extracts characteristic features from the profile of a second user who will be speaking to the first user, and creates second information that shows the relationships between the features in a graph or vector form; and a filtering unit that outputs, from among the objects included in the first information, words that represent objects whose distance from the features included in the second information is equal to or less than a predetermined threshold.
  • the present invention makes it possible to provide high-quality dialogue even for users who are not good at self-disclosure.
  • FIG. 1 is a diagram showing an example of a dialogue between agents in a virtual space.
  • FIG. 2 shows an example of a dialogue between agents in a virtual space and a dialogue between a user and an agent in a physical space.
  • FIG. 3 is a diagram illustrating an example of a dialogue system including an information presentation device.
  • FIG. 4 is a diagram showing an outline of the operation of the information presentation device.
  • FIG. 5 is a diagram showing an example of converting non-verbal information into text and graphs.
  • FIG. 6 is a diagram showing an example of profile information converted into text and graphs.
  • FIG. 7 is a diagram showing an example in which a graph of non-verbal information and a graph of agent profile information are arranged on a common sense graph.
  • FIG. 1 is a diagram showing an example of a dialogue between agents in a virtual space.
  • FIG. 2 shows an example of a dialogue between agents in a virtual space and a dialogue between a user and an agent in a physical space.
  • FIG. 8A is a diagram illustrating an example of the configuration of an information processing apparatus according to each embodiment.
  • FIG. 8B is a diagram showing an example of non-verbal information.
  • FIG. 9 is a flowchart illustrating an example of a processing procedure executed by the information presentation device according to the first embodiment.
  • FIG. 10 is a diagram showing an example of converting non-verbal information into text and vectors.
  • FIG. 11 is a diagram showing an example of profile information converted into text and vector.
  • FIG. 12 is a diagram showing an example of calculation of the distance between a word in non-verbal information and a word in the profile information of an agent.
  • FIG. 13 is a flowchart illustrating an example of a processing procedure executed by the information presentation apparatus according to the second embodiment.
  • FIG. 14 is a diagram illustrating an example of a computer that executes an information presentation program.
  • the information presentation device outputs, for example, a topic (e.g., words used in the topic) when agent A of user A converses with agent B of user B in the virtual space shown in Fig. 1.
  • a topic e.g., words used in the topic
  • the information presentation device acquires non-verbal information about Agent B (e.g., information about Agent B or the environment around Agent B).
  • Agent A's profile information information indicating Agent A's personality, thoughts, etc.
  • the information presentation device then outputs words that represent the extracted non-verbal information as words to be used as topics of conversation between Agent A and Agent B. This allows Agent A to have an intimate conversation with Agent B.
  • the agent can say to the interlocutor, "Your hairstyle is cute today! or the like. Also, if the information presentation device extracts "XX-chan's mark” as non-verbal information that is important to the agent, the agent can say to the interlocutor, "You're a fan of XX-chan, aren't you?" or the like.
  • the information output by the information presentation device may be used for dialogue between agents in a virtual space, as described above, or for dialogue between a user and an agent in a physical space.
  • the information presentation device 10 is applied, for example, to the dialogue system shown in Figure 3.
  • the dialogue system comprises the information presentation device 10 and a dialogue device.
  • the dialogue device comprises a dialogue simulation unit that simulates dialogue between agents, a human DTDB (human digital twin database) that stores information indicating the personality and thoughts of users (agents), and a visualization unit that creates video of dialogue between agents.
  • a human DTDB human digital twin database
  • the information presentation device 10 acquires information (profile information) indicating the personality and thoughts of the user (agent) engaging in dialogue from the dialogue device's human DTDB.
  • the information presentation device 10 also acquires non-verbal information about the agent's dialogue partner from the visualization unit or an external device.
  • the information presentation device 10 then extracts non-verbal information important to the agent based on the acquired information.
  • the information presentation device 10 then outputs words highly associated with the text representing the extracted non-verbal information to the dialogue simulation unit as words to be used as topics of dialogue.
  • the conversation device 10 By having the information presentation device 10 output words used in conversation topics in the manner described above, the conversation device can provide high-quality conversations even with conversation partners who are not good at self-disclosure.
  • the information presentation device 10 acquires non-verbal information from agent A's conversation partner, it converts the objects included in the non-verbal information into text and graphs or vectorizes the objects in the non-verbal information.
  • the information presentation device 10 converts an image of the scenery seen by the conversation partner into text using a caption system. Then, the information presentation device 10 applies LLM to the text to graph the objects included in the scenery.
  • the information presentation device 10 also references the common sense DB and graphs or vectorizes the items (text) included in agent A's profile information. For example, the information presentation device 10 graphs the items included in the profile by applying LLM to agent A's profile information, as shown in Figure 6.
  • the information presentation device 10 then references a common sense database (a database that shows the semantic similarity and relevance of words; details will be provided later) and compares the graph or vector of the objects included in the non-verbal information with the graph or vector of the items included in agent A's profile. Then, using the results of this comparison, the information presentation device 10 selects from the objects included in the non-verbal information those whose distance to the items included in agent A's profile (for example, name) is less than a predetermined threshold and outputs them to agent A as important non-verbal information (performing a filtering process).
  • a common sense database a database that shows the semantic similarity and relevance of words; details will be provided later
  • the information presentation device 10 places a graph of objects included in the non-verbal information and a graph of items included in agent A's profile information on a graph (common sense graph) showing the relationships between words in the common sense DB. Then, the information presentation device 10 extracts, from the nodes of the non-verbal information, nodes whose distance from a node of agent A's profile information (for example, the node for the name "Audrey") is less than a predetermined threshold (for example, 2) as non-verbal information important to agent A. For example, the information presentation device 10 extracts "Shibuya Station" and "Billboard" as shown in FIG. 7.
  • the information presentation device 10 may output the above-mentioned "Shibuya Station” and “Billboard” themselves as words to be used in the conversation, or may output words related to "Shibuya Station” and "Billboard.”
  • the information presentation device 10 may output nodes (e.g., "Tokyo” and "design") located in the direction of the group of nodes in Agent A's profile information from among the nodes whose distance from "Shibuya Station” and "Billboard” is less than a predetermined threshold (e.g., 1) on the graph shown in FIG. 7.
  • a predetermined threshold e.g. 1, 1
  • the information presentation device 10 outputs the above-mentioned new topic words as words to be used in the conversation topic.
  • the information presentation device 10 according to the first embodiment performs filtering processing by graphing non-verbal information and agent profile information.
  • the information presentation device 10 includes, for example, an input/output unit 11, a storage unit 12, and a control unit 13.
  • the input/output unit 11 is an interface that handles the input and output of various data.
  • the input/output unit 11 accepts input such as agent profile information and non-verbal information about the agent's interlocutor.
  • agent profile information is information that indicates the agent's profile, the agent's thoughts, etc.
  • Non-verbal information is information that indicates the user and the environment around the user, such as the user's physical characteristics and information about objects around the user.
  • Non-verbal information is input to the information presentation device 10, for example, in the form of a screen capture (image) of the environment around the agent's interlocutor, text listing the objects around the interlocutor, etc.
  • non-verbal information input to the information presentation device 10 may be a screen capture of the conversation partner's line of sight in the virtual space (see reference numeral 801 in Figure 8B), a list of objects around the conversation partner, and information indicating the positions of the objects (see reference numeral 802 in Figure 8B), etc.
  • the memory unit 12 stores data, programs, etc. that are referenced when the control unit 13 executes various processes.
  • the memory unit 12 is realized by semiconductor memory elements such as RAM (Random Access Memory) or flash memory, or by storage devices such as hard disks or optical disks.
  • the memory unit 12 includes a common sense database.
  • the common sense database is realized, for example, by a common sense graph, such as ConceptNet, which represents the relationships between words in a graph.
  • the memory unit 12 also stores non-verbal information received by the input/output unit 11, agent profile information, etc.
  • the control unit 13 is responsible for overall control of the information presentation device 10.
  • the functions of the control unit 13 are realized, for example, by the CPU (Central Processing Unit) executing a program stored in the memory unit 12.
  • the control unit 13 includes, for example, a first information creation unit 131, a second information creation unit 134, and a filtering unit 135.
  • the first information creation unit 131 acquires non-verbal information of the conversation partner, extracts characteristic objects from the acquired non-verbal information, and creates a first graph (a graph of non-verbal information) that graphs the relationships between the objects in the non-verbal information.
  • the first information creation unit 131 includes a text conversion unit 132 and an information creation unit 133.
  • the text conversion unit 132 converts the input non-verbal information into text. For example, if the input non-verbal information is a photograph or screen capture of a physical space or a virtual space, the text conversion unit 132 converts the scenery into text using an automatic caption generation system or the like. Also, for example, if the input non-verbal information is the name of a room in a virtual space, the time, and the names and location information of surrounding objects, the text conversion unit 132 converts the scenery around the conversation partner into text based on this information. The text conversion unit 132 then outputs the converted text.
  • the information creation unit 133 extracts characteristic objects from the non-language information that has been converted into text, and creates a graph of the non-language information that graphs the relationships between the objects in the non-language information.
  • the information creation unit 133 uses a learning model for understanding text (e.g., ChatGPT) to extract topic words from the non-verbal information that has been converted into text, express them as nodes, and create information that graphs the relationships between the nodes (a graph of non-verbal information). For example, the information creation unit 133 creates the graph shown in Figure 5.
  • a learning model for understanding text e.g., ChatGPT
  • the second information creation unit 134 extracts characteristic items from the profile information of the agent, and creates a second graph (a graph of the profile information) that graphically illustrates the relationships between the items in the profile information.
  • the second information creation unit 134 uses a learning model for understanding text to extract characteristic features from the profile information, express them as nodes, and create information (a profile information graph) that graphically depicts the relationships between the nodes. For example, the second information creation unit 134 creates the graph shown in FIG. 6.
  • the filtering unit 135 refers to the common sense graph and extracts, from among the objects included in the non-verbal information graph, words of objects whose distance from items included in the profile information graph is equal to or less than a predetermined threshold, as important words. The filtering unit 135 then extracts and outputs, as words used in conversation topics, nodes whose distance from the important words is within the predetermined threshold and that are located in the direction of the profile information node group.
  • the filtering unit 135 receives the common sense graph, the non-verbal information graph, the profile information graph, and a threshold value as input, and performs the following processing.
  • G cs (V cs ;E cs )
  • G env (V env ;E env )
  • G ag (V ag ;E ag )
  • the filtering unit 135 plots the relationships (edges) between the nodes of the non-verbal information graph and the profile information graph on the common sense graph to create G'cs ( ⁇ Genv, Gag ⁇ ⁇ G'cs ⁇ ).
  • the filtering unit 135 uses the above G'cs to calculate the distance between a node in the graph of non-verbal information and a node in the graph of profile information, and then flags pairs of nodes (words) whose inter-node distance is equal to or less than the above threshold as "important.”
  • the filtering unit 135 flags nodes in the graph of non-verbal information as "important" if their distance from a pre-specified node in the graph of profile information (e.g., a node indicating a name) is less than a threshold.
  • the distance between nodes is If there are no edges between nodes, it is infinite ( ⁇ ). If a node in the graph of non-verbal information matches a node in the graph of profile information, 0 ⁇ [0,...,n, ⁇ ],n ⁇
  • the filtering unit 135 may reflect the type of relationship between the nodes in setting the weight of the edge connecting the nodes. Furthermore, the filtering unit 135 may calculate the number of edges with the shortest distance between the nodes as the distance between the nodes.
  • the filtering unit 135 extracts and outputs, as words used in conversation topics, words of nodes whose distance from a node flagged as "important" (a node of an important word) is less than a predetermined threshold (for example, 1) and which are located in the direction of the group of nodes in the profile information graph.
  • a predetermined threshold for example, 1
  • the first information creation unit 131 acquires non-verbal information of the conversation partner (S1) and converts it into text (S2). Then, the first information creation unit 131 creates a graph of the non-verbal information from the non-verbal information converted into text in S2 (S3).
  • the second information creation unit 134 acquires profile information of the agent (S4). Then, the second information creation unit 134 creates a profile information graph from the profile information acquired in S4 (S5).
  • the filtering unit 135 refers to the common sense graph and extracts, from among the objects included in the non-verbal information graph, objects whose distance to the items included in the profile information graph is less than a predetermined threshold, as important words (S6).
  • the filtering unit 135 extracts and outputs, as words to be used as topics of conversation, the words of nodes located in the direction of the group of nodes in the profile information graph, from among the nodes whose distance from the node of the important word extracted in S6 is less than a predetermined threshold (for example, 1) (S7).
  • a predetermined threshold for example, 1
  • the information presentation device 10 can provide high-quality dialogue even with conversation partners who are not comfortable self-disclosing.
  • the information presentation device 10 according to the second embodiment performs filtering processing by vectorizing non-verbal information and agent profile information.
  • the information presentation device 10 creates a common sense vector that represents the semantic similarity between words in the common sense database as a vector.
  • the common sense vector is created using, for example, Gensim word2vec, doc2vec, etc.
  • the created common sense vector is stored in the storage unit 12. Components that are the same as those in the first embodiment are assigned the same reference numerals and will not be described again.
  • the information creating unit 133 converts the non-language information into text, it refers to the common sense vector and creates information by vectorizing the words included in the non-language information that has been converted into text.
  • the information creation unit 133 extracts topic words from the non-verbal information that has been converted into text, and creates information (non-verbal information vectors) by vectorizing the extracted words using common sense vectors.
  • Figure 10 shows an example of a two-dimensional representation of non-verbal information vectors.
  • V cs ⁇ w cs — 1 , w cs — 2 , . . . , w cs — n ⁇ .
  • the information creation unit 133 checks whether the words included in the text of the non-language information are already included in the common sense vector. Then, the information creation unit 133 adds to Vcs any words included in the text of the non-language information that are not included in the common sense vector. For example, the information creation unit 133 converts the words that need to be added into embedded representations using a technique such as one-hot vectors, and adds them to Vcs .
  • V'cs the common sense vector to which words have been added
  • V'cs the vector of each word in the common sense vector V'cs
  • w'cs the vector of each word in the common sense vector
  • V'cs ⁇ w'cs_1 , w'cs_2 , ..., w'cs_(n+m) ⁇ , where m is the number of added words.
  • the information creation unit 133 stores both Vcs and V' in the storage unit 12.
  • the second information creation unit 134 also refers to the common sense vector and creates information in which words of items included in the agent's profile information are vectorized.
  • the second information creation unit 134 extracts words that represent characteristic items from the agent's profile information, and creates information (profile information vectors) by vectorizing the extracted words using common sense vectors.
  • profile information vectors information vectors
  • Figure 11 shows an example of a two-dimensional representation of profile information vectors.
  • the filtering unit 135 performs filtering processing using the non-language information vector V env , the profile information vector V ag , and a distance threshold.
  • the filtering unit 135 plots the elements ⁇ w env_1 , w env_2 , ..., w env_nenv ⁇ of the non-language information vector V env and the elements ⁇ w ag_1 , w ag_2 , ..., w ag_nag ⁇ of the profile information vector V ag in the same vector space, and measures the distance between the elements of the non-language information vector and the elements of the profile information vector. For example, Euclidean distance is used to measure the distance.
  • the filtering unit 135 then flags elements (words) of the non-verbal information vector whose distance from the profile information vector element is less than a threshold as "important.”
  • the filtering unit 135 calculates the distance between the elements of the non-verbal information vector and the elements of the profile information vector. Then, the filtering unit 135 extracts elements (e.g., "Shibuya", "Billboard") from the elements of the non-verbal information vector whose distance from the elements of the profile information vector is less than a threshold value (e.g., 0.2), and marks them with a flag indicating "important.”
  • elements e.g., "Shibuya", "Billboard
  • the filtering unit 135 extracts and outputs, from among the elements of the non-verbal information vector and the elements of the profile information vector, words (e.g., "Shibuya station” and “Billboard design”) whose distance from the above "Shibuya” and “Billboard” is less than a predetermined threshold (e.g., 0.1) as words to be used as topics of conversation.
  • words e.g., "Shibuya station” and "Billboard design
  • a predetermined threshold e.g., 0.1
  • FIG. 13 An example of a processing procedure executed by the information presentation device 10 will be described using Fig. 13.
  • the processing of S11 and S12 in Fig. 13 is the same as S1 and S2 in Fig. 9, so the description will begin with S13 in Fig. 13.
  • the first information creation unit 131 refers to the common sense vector and creates a vector of non-verbal information from the non-verbal information converted into text in S12 (S13).
  • the second information creation unit 134 acquires the agent's profile information (S14). Then, the second information creation unit 134 references the common sense vector and creates a vector of the profile information acquired in S14 (S15).
  • the filtering unit 135 calculates the distance between the elements of the vector of non-verbal information created in S13 and the elements of the vector of profile information created in S15 (S16).
  • the filtering unit 135 extracts, from among the elements included in the non-verbal information vector, words of elements whose distance to elements of the profile information vector is less than a predetermined threshold, as important words for the agent (S17: Extraction of important words).
  • the filtering unit 135 extracts and outputs, from among the elements of the non-verbal information vector and the elements of the profile information vector, words whose distance from the important words extracted in S17 is less than a predetermined threshold, as words to be used in the conversation topic (S18).
  • the information presentation device 10 can provide high-quality dialogue even with conversation partners who are not comfortable self-disclosing.
  • the information presentation device 10 can extract information highly relevant to the user's personality and preferences from non-verbal information about the user and their surroundings, and use that information to generate dialogue. This makes it possible to have high-quality dialogue tailored to the user, even for users who are not good at self-disclosure, based on the user's non-verbal information.
  • the filtering unit 135 of each embodiment may use a score representing the distance between words (e.g., the shorter the distance, the lower the score) when extracting important words.
  • the filtering unit 135 may extract all words with scores equal to or less than a predetermined threshold, or may extract a predetermined number of words in ascending order of scores.
  • the filtering unit 135 may also use a score representing the distance between words in the same manner as above when extracting words used in conversation topics using the distance from important words.
  • the information presentation device 10 of each embodiment may further include an utterance generation unit (e.g., a dialogue simulation unit in FIG. 3) that generates utterances from the agent to the dialogue partner based on the words output from the filtering unit 135.
  • the utterance generation unit generates and outputs utterances from the agent using the words output from the filtering unit 135 as topics.
  • each unit shown in the figure is conceptual functional units and do not necessarily have to be physically configured as shown.
  • the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
  • all or any part of the processing functions performed by each device can be realized by a CPU and a program executed by the CPU, or can be realized as hardware using wired logic.
  • the information presentation device 10 can be implemented by installing a program (information presentation program) as package software or online software on a desired computer. For example, by executing the program on an information processing device, the information processing device can function as the information presentation device 10.
  • the information processing device referred to here includes mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone Systems), as well as terminals such as PDAs (Personal Digital Assistants).
  • FIG. 14 is a diagram showing an example of a computer that executes an information presentation program.
  • the computer 1000 has, for example, memory 1010 and a CPU 1020.
  • the computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
  • Memory 1010 includes ROM (Read Only Memory) 1011 and RAM (Random Access Memory) 1012.
  • ROM 1011 stores a boot program such as BIOS (Basic Input Output System).
  • Hard disk drive interface 1030 is connected to hard disk drive 1031.
  • Disk drive interface 1040 is connected to disk drive 1041.
  • a removable storage medium such as a magnetic disk or optical disk is inserted into disk drive 1041.
  • Serial port interface 1050 is connected to, for example, a mouse 1110 and keyboard 1120.
  • Video adapter 1060 is connected to, for example, a display 1130.
  • the hard disk drive 1031 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094.
  • the programs that define the processes executed by the information presentation device 10 are implemented as program modules 1093 in which computer-executable code is written.
  • the program modules 1093 are stored, for example, on the hard disk drive 1031.
  • a program module 1093 for executing processes similar to the functional configuration of the information presentation device 10 is stored on the hard disk drive 1031.
  • the hard disk drive 1031 may be replaced by an SSD (Solid State Drive).
  • the data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1031.
  • the CPU 1020 reads the program modules 1093 and program data 1094 stored in memory 1010 or hard disk drive 1031 into RAM 1012 as needed and executes them.
  • the program module 1093 and program data 1094 do not necessarily have to be stored on the hard disk drive 1031; they may instead be stored on a removable storage medium and read by the CPU 1020 via the disk drive 1041 or the like.
  • the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a LAN (Local Area Network) or WAN (Wide Area Network)).
  • the program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computational Linguistics (AREA)
  • General Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

情報提示装置は、第1のユーザまたは第1のユーザの周囲についての非言語情報を取得し、取得した非言語情報から特徴的なオブジェクトを抽出し、オブジェクト同士の関係性をグラフ化またはベクトル化して示した第1の情報を作成する。また、情報提示装置は、第1のユーザの対話相手である第2のユーザのプロフィールから特徴的な事項を抽出し、抽出した事項同士の関係性をグラフ化またはベクトル化して示した第2の情報を作成する。その後、情報提示装置は、第1の情報に含まれるオブジェクトの中から、第2の情報に含まれる事項との距離が所定の閾値以下のオブジェクトを表す単語を抽出し、出力する。

Description

情報提示装置、情報提示方法、および、情報提示プログラム
 本発明は、デジタル・ヒューマン・エージェントの発話の生成に用いられる、情報提示装置、情報提示方法、および、情報提示プログラムに関する。
 従来、ユーザの音声やテキスト入力に基づいて自律的に対話を行うデジタル・ヒューマン・エージェントの開発が進められている。これらの多くは、学習済みの個別のモデルや大規模言語モデル(LLM:Large Language Models)により、ユーザの発話に対する応答を生成する。
Gemini:A Family of Highly Capable Multimodal Models,Gemini Team, Google,[online],[2024年4月9日検索],インターネット,<URL:https://arxiv.org/pdf/2312.11805.pdf> YOLO-World,Tianheng Cheng et al.,[online],[2024年4月9日検索],インターネット,<URL:https://www.yoloworld.cc/> SIMMC 2.0:A Task-oriented Dialog Dataset for Immersive Multimodal Conversations,Satwik Kottur et al.,[online],[2024年4月9日検索],インターネット,<https://arxiv.org/pdf/2104.08667.pdf> MediaGnosis,[online],[2024年4月9日検索],インターネット,<https://www.rd.ntt/mediagnosis/>
 しかし、初対面では人見知りするようなユーザ等、自己開示が苦手なユーザとの対話において、ユーザの言語情報のみからでは当該ユーザの情報が十分に得られない。その結果、ユーザとの対話が定型的な対話にとどまり、対話の質や親密さが高くならないという問題がある。そこで本発明は、前記した問題を解決し、自己開示が苦手なユーザに対しても質の高い対話を提供することを課題とする。
 前記した課題を解決するため、本発明は、第1のユーザまたは前記第1のユーザの周囲についての非言語情報を取得し、取得した前記非言語情報から特徴的なオブジェクトを抽出し、前記オブジェクト同士の関係性をグラフ化またはベクトル化して示した第1の情報を作成する第1の情報作成部と、前記第1のユーザへの発話を行う第2のユーザのプロフィールから特徴的な事項を抽出し、前記事項同士の関係性をグラフ化またはベクトル化して示した第2の情報を作成する第2の情報作成部と、前記第1の情報に含まれるオブジェクトの中から、前記第2の情報に含まれる事項との距離が所定の閾値以下のオブジェクトを表す単語を出力するフィルタリング部とを備えることを特徴とする。
 本発明によれば、自己開示が苦手なユーザに対しても質の高い対話を提供することができる。
図1は、仮想空間におけるエージェント同士の対話の例を示す図である。 図2は、仮想空間におけるエージェント同士の対話と物理空間におけるユーザとエージェントとの対話の例を示す図である。 図3は、情報提示装置を含む対話システムの例を示す図である。 図4は、情報提示装置の動作概要を示す図である。 図5は、非言語情報のテキスト化およびグラフ化の例を示す図である。 図6は、プロフィール情報のテキスト化およびグラフ化の例を示す図である。 図7は、常識グラフ上に、非言語情報のグラフと、エージェントのプロフィール情報のグラフとを配置した例を示す図である。 図8Aは、各実施形態の情報処理装置の構成例を示す図である。 図8Bは、非言語情報の例を示す図である。 図9は、第1の実施形態の情報提示装置が実行する処理手順の例を示すフローチャートである。 図10は、非言語情報のテキスト化およびベクトル化の例を示す図である。 図11は、プロフィール情報のテキスト化およびベクトル化の例を示す図である。 図12は、非言語情報の単語とエージェントのプロフィール情報の単語との距離の算出例を示す図である。 図13は、第2の実施形態の情報提示装置が実行する処理手順の例を示すフローチャートである。 図14は、情報提示プログラムを実行するコンピュータの一例を示す図である。
 以下、図面を参照しながら、本発明を実施するための形態(実施形態)について説明する。本発明は、本実施形態に限定されない。
[概要]
 まず、図1を用いて、本実施形態の情報提示装置の概要を説明する。情報提示装置は、例えば、図1に示す仮想空間においてユーザAのエージェントAがユーザBのエージェントBと対話する際の話題(例えば、話題に用いる単語)を出力する。
 例えば、情報提示装置は、エージェントBの非言語情報(例えば、エージェントBもしくはエージェントBの周囲の環境に関する情報)を取得する。そして、情報提示装置は、エージェントAのプロフィール情報(エージェントAの個性・思考等を示す情報)を参照し、エージェントBの非言語情報の中から、エージェントAにとって重要な非言語情報を抽出する。そして、情報提示装置は、抽出した非言語情報を表す単語を、エージェントAがエージェントBと対話する際の話題に用いる単語として出力する。これにより、エージェントAはエージェントBとの間で親密な対話をすることができる。
 例えば、情報提示装置が、図2に示す例において、エージェントにとって重要な非言語情報として「髪」を抽出した場合、エージェントは対話相手に対し「今日の髪型はかわいいですね!」等の発話を行うことができる。また、情報提示装置が、エージェントにとって重要な非言語情報として「〇〇ちゃんのマーク」を抽出した場合、エージェントは対話相手に対し「〇〇ちゃんのファンですね」等の発話を行うことができる。
 なお、情報提示装置が出力する情報は、上記のように、仮想空間におけるエージェント同士の対話に用いられてもよいし、物理空間におけるユーザとエージェントとの対話に用いられてもよい。
 情報提示装置10は、例えば、図3に示す対話システムに適用される。対話システムは、情報提示装置10と対話装置とを備える。対話装置は、エージェント同士の対話を模擬する対話模擬部、ユーザ(エージェント)の個性・思考を示す情報を蓄積するヒトDTDB(ヒトデジタルツインデータベース)、エージェント同士の対話映像を作成する可視化部を備える。
 上記の対話システムにおいて、情報提示装置10は、対話装置のヒトDTDBから対話を行うユーザ(エージェント)の個性・思考を示す情報(プロフィール情報)を取得する。また、情報提示装置10は、可視化部または外部装置からエージェントの対話相手の非言語情報を取得する。その後、情報提示装置10は、取得した情報に基づき当該エージェントにとって重要な非言語情報を抽出する。そして、情報提示装置10は、抽出した非言語情報を表すテキストとの関連度の高い単語を、対話の話題に用いる単語として対話模擬部に出力する。
 情報提示装置10が上記のようにして対話の話題に用いる単語を出力することで、対話装置は、自己開示が苦手な対話相手に対しても質の高い対話を提供することができる。
 次に、図4を用いて、情報提示装置10の動作概要を説明する。例えば、情報提示装置10は、エージェントAの対話相手の非言語情報を取得すると、当該非言語情報に含まれるオブジェクトをテキスト化し、非言語情報のオブジェクトをグラフ化またはベクトル化する。例えば、情報提示装置10は、図5に示すように、対話相手の見ている風景の画像を、キャプションシステムによりテキスト化する。その後、情報提示装置10は、当該テキストにLLMを適用することにより、風景に含まれるオブジェクトをグラフ化する。
 図4の説明に戻る。また、情報提示装置10は、常識DBを参照し、エージェントAのプロフィール情報に含まれる事項(テキスト)をグラフ化またはベクトル化する。例えば、情報提示装置10は、図6に示すようにエージェントAのプロフィール情報にLLMを適用することにより、プロフィールに含まれる事項をグラフ化する。
 図4の説明に戻る。その後、情報提示装置10は、常識DB(単語同士の意味の近さ、単語同士の関連性を示すDB、詳細は後記)を参照し、非言語情報に含まれるオブジェクトのグラフまたはベクトルと、エージェントAのプロフィールに含まれる事項のグラフまたはベクトルとを比較する。そして、情報提示装置10は、上記の比較結果を用いて、非言語情報に含まれるオブジェクトの中から、エージェントAのプロフィールに含まれる事項(例えば、名前)との距離が所定の閾値以下のオブジェクトを、エージェントAに重要な非言語情報として出力する(フィルタリング処理を行う)。
 例えば、情報提示装置10は、図7に示すように、常識DBの単語同士の関連性を示すグラフ(常識グラフ)上に、非言語情報に含まれるオブジェクトのグラフと、エージェントAのプロフィール情報に含まれる事項のグラフとを配置する。そして、情報提示装置10は、非言語情報のノードの中から、エージェントAのプロフィール情報のノード(例えば、名前「Audrey」のノード)からの距離が所定の閾値(例えば、2)以下のノードを、エージェントAに重要な非言語情報として抽出する。例えば、情報提示装置10は、図7に示す「Shibuya Station」、「Billboard」を抽出する。
 なお、情報提示装置10は、対話の話題に用いる単語として、上記の「Shibuya Station」、「Billboard」そのものを出力してもよいし、「Shibuya Station」、「Billboard」に関連する単語を出力してもよい。
 例えば、情報提示装置10は、図7に示すグラフ上で、「Shibuya Station」、「Billboard」からの距離が所定の閾値以下(例えば、1)のノードのうち、エージェントAのプロフィール情報のノード群の方向に配置されるノード(例えば、「Tokyo」、「design」)を出力してもよい。これにより、情報提示装置10は、エージェントAに重要な非言語情報と関連した、新たな話題の単語を出力できる。その結果、対話装置は、対話の話題を展開しやすくなり、対話の質を向上させることができる。
 なお、以下に説明する各実施形態では、情報提示装置10が、対話の話題に用いる単語として、上記の新たな話題の単語を出力する場合を例に説明する。
[第1の実施形態]
 次に、第1の実施形態の情報提示装置10を説明する。第1の実施形態の情報提示装置10は、非言語情報およびエージェントのプロフィール情報をグラフ化してフィルタリング処理を行う。
[構成例]
 次に、図8Aを用いて、情報提示装置10の構成例を説明する。情報提示装置10は、例えば、入出力部11、記憶部12、および、制御部13を備える。
 入出力部11は、各種データの入出力を司るインタフェースである。入出力部11は、例えば、エージェントのプロフィール情報、エージェントの対話相手の非言語情報等の入力を受け付ける。
 エージェントのプロフィール情報は、前記した通り、エージェントのプロフィール、エージェントの思考等を示す情報である。
 非言語情報は、ユーザおよびユーザの周囲の環境を示す情報であり、例えば、ユーザの身体的特徴、ユーザの周囲に存在する物体の情報等である。非言語情報は、例えば、エージェントの対話相手の周囲の環境のスクリーンキャプチャー(画像)、対話相手の周囲のオブジェクトの一覧を示すテキスト等により、情報提示装置10に入力される。
 例えば、情報提示装置10に入力される非言語情報は、仮想空間における対話相手の視線のスクリーンキャプチャー(図8Bの符号801参照)、対話相手の周辺のオブジェクトのリストとオブジェクトの位置とを示す情報(図8Bの符号802参照)等である。
 図8Aの説明に戻る。記憶部12は、制御部13が各種処理を実行する際に参照されるデータ、プログラム等を記憶する。記憶部12は、RAM(Random Access Memory)、フラッシュメモリ(Flash Memory)等の半導体メモリ素子、または、ハードディスク、光ディスク等の記憶装置によって実現される。
 例えば、記憶部12は、常識DBを備える。常識DBは、例えば、ConceptNet等の単語間の関連性をグラフで表した常識グラフにより実現される。また、記憶部12は、入出力部11で受け付けた非言語情報、エージェントのプロフィール情報等を記憶する。
 制御部13は、情報提示装置10全体の制御を司る。制御部13の機能は、例えば、CPU(Central Processing Unit)が、記憶部12に記憶されるプログラムを実行することにより実現される。
 制御部13は、例えば、第1の情報作成部131と、第2の情報作成部134と、フィルタリング部135とを備える。
[第1の情報作成部]
 第1の情報作成部131は、対話相手の非言語情報を取得し、取得した非言語情報から特徴的なオブジェクトを抽出し、当該非言語情報におけるオブジェクト同士の関係性をグラフ化した第1のグラフ(非言語情報のグラフ)を作成する。第1の情報作成部131は、テキスト化部132と情報作成部133とを備える。
[テキスト化部]
 テキスト化部132は、入力された非言語情報をテキスト化する。例えば、入力された非言語情報が、物理空間・仮想空間の写真、スクリーンキャプチャーである場合、テキスト化部132は、自動キャプション生成システム等を活用して、風景をテキスト化する。また、例えば、入力された非言語情報が、仮想空間の部屋名、時間、周囲のオブジェクト名と位置情報である場合、これらの情報をもとに対話相手の周囲の風景をテキスト化する。そして、テキスト化部132はテキスト化した情報を出力する。
[情報作成部]
 情報作成部133は、テキスト化された非言語情報から特徴的なオブジェクトを抽出し、当該非言語情報におけるオブジェクト同士の関係性をグラフ化した非言語情報のグラフを作成する。
 例えば、情報作成部133は、テキストを理解する学習モデル(例えば、ChatGPT)を用いて、テキスト化された非言語情報からトピックとなる単語を抽出してノードとして表現し、ノード間の関係性をグラフ化した情報(非言語情報のグラフ)を作成する。例えば、情報作成部133は、図5に示すグラフを作成する。
[第2の情報作成部]
 図8Aの説明に戻る。第2の情報作成部134は、エージェントのプロフィール情報から特徴的な事項を抽出し、当該プロフィール情報における事項同士の関係性をグラフ化した第2のグラフ(プロフィール情報のグラフ)を作成する。
 例えば、第2の情報作成部134は、情報作成部133と同様に、テキストを理解する学習モデルを用いて、プロフィール情報から特徴的な事項を抽出してノードとして表現し、ノード間の関係性をグラフ化した情報(プロフィール情報のグラフ)を作成する。例えば、第2の情報作成部134は、図6に示すグラフを作成する。
[フィルタリング部]
 図8Aの説明に戻る。フィルタリング部135は、常識グラフを参照し、非言語情報のグラフに含まれるオブジェクトの中から、プロフィール情報のグラフに含まれる事項との距離が所定の閾値以下のオブジェクトの単語を、重要単語として抽出する。そして、フィルタリング部135は、当該重要単語からの距離が所定の閾値以内、かつ、プロフィール情報のノード群の方向に配置されるノードを、対話の話題に用いる単語として抽出し、出力する。
 例えば、フィルタリング部135は、常識グラフ、非言語情報のグラフ、プロフィール情報のグラフ、閾値を入力とし、以下の処理を実行する。
 なお、
・常識グラフ:Gcs=(Vcs;Ecs)
・非言語情報のグラフ:Genv=(Venv;Eenv)
・プロフィール情報のグラフ:Gag=(Vag;Eag)
とする。
 フィルタリング部135は、非言語情報のグラフおよびプロフィール情報のグラフの各ノードの関係性(エッジ)を常識グラフ上にプロットして、G’csを作る({Genv,Gag}⊆G’cs})。
 その後、フィルタリング部135は、上記のG’csを利用して、非言語情報のグラフのノードとプロフィール情報のグラフのノードとの距離を算出する。そして、フィルタリング部135は、ノード間の距離が、上記の閾値以下のノード(単語)のペアに「重要」を示すフラグを付ける。
 例えば、フィルタリング部135は、予め指定されたプロフィール情報のグラフのノード(例えば、名前を示すノード)からの距離が閾値以下である非言語情報のグラフのノードに「重要」を示すフラグを付ける。
 なお、ノード間の距離は、
・ノード間にエッジが存在しない場合、無限(∞)
・非言語情報のグラフのノードとプロフィール情報のグラフのノードとが一致する場合、0
・[0,…,n,∞],n≦|V(G’cs)|
とする。
 また、ノード間の関係性の種類が複数ある場合、フィルタリング部135は、ノード間の関係性の種類を、ノード間を接続するエッジの重みの設定に反映してもよい。また、フィルタリング部135は、ノード間の最短距離のエッジの数を当該ノード間の距離として算出してもよい。
 フィルタリング部135は、「重要」フラグが付されたノード(重要単語のノード)からの距離が所定の閾値以下(例えば、1)、かつ、プロフィール情報のグラフのノード群の方向に配置されるノードの単語を、対話の話題に用いる単語として抽出し、出力する。
[処理手順の例]
 図9を用いて、情報提示装置10が実行する処理手順の例を説明する。まず、第1の情報作成部131は、対話相手の非言語情報を取得し(S1)、テキスト化する(S2)。そして、第1の情報作成部131は、S2でテキスト化した非言語情報から非言語情報のグラフを作成する(S3)。
 また、第2の情報作成部134は、エージェントのプロフィール情報を取得する(S4)。そして、第2の情報作成部134は、S4で取得したプロフィール情報からプロフィール情報のグラフを作成する(S5)。
 そして、フィルタリング部135は、常識グラフを参照し、非言語情報のグラフに含まれるオブジェクトの中から、プロフィール情報のグラフに含まれる事項との距離が所定の閾値以下のオブジェクトを、重要単語として抽出する(S6)。
 S6の後、フィルタリング部135は、S6で抽出された重要単語のノードからの距離が所定の閾値以下(例えば、1)のノードのうち、プロフィール情報のグラフのノード群の方向に配置されるノードの単語を、対話の話題に用いる単語として抽出し、出力する(S7)。
 情報提示装置10が上記の処理を実行することにより、自己開示が苦手な対話相手に対しても質の高い対話を提供することができる。
[第2の実施形態]
 次に、第2の実施形態の情報提示装置10を説明する。第2の実施形態の情報提示装置10は、非言語情報およびエージェントのプロフィール情報をベクトル化してフィルタリング処理を行う。
 なお、情報提示装置10は、上記のベクトル化のため、例えば、常識DBの単語間の意味の近さをベクトルで表した常識ベクトルを作成する。常識ベクトルは、例えば、Gensim word2vec、doc2vec等により作成される。作成された常識ベクトルは、記憶部12に格納される。第1の実施形態と同じ構成は同じ符号を付して説明を省略する。
[情報作成部]
 情報作成部133は、非言語情報をテキスト化すると、常識ベクトルを参照し、テキスト化された非言語情報に含まれる単語をベクトル化した情報を作成する。
 例えば、情報作成部133は、テキスト化された非言語情報からトピックとなる単語を抽出し、抽出した単語を、常識ベクトルを用いてベクトル化した情報(非言語情報のベクトル)を作成する。なお、図10は、非言語情報のベクトルを2次元で表現した例を示している。
[ベクトル化]
 非言語情報のベクトル化について詳細に説明する。情報作成部133は、非言語情報のテキスト、常識ベクトルの入力を受け付けると、以下の処理を実行する。
 なお、常識ベクトルをVcsとし、常識ベクトルの各単語のベクトルをwcsで表現すると、Vcs={wcs_1,wcs_2,…,wcs_n}となる。
 まず、情報作成部133は、非言語情報のテキストに含まれる単語が、すでに常識ベクトルに入っているか否かを確認する。そして、情報作成部133は、非言語情報のテキストに含まれる単語のうち、常識ベクトルに入っていない単語をVcsに追加する。例えば、情報作成部133は、追加が必要な単語をOne-hotベクトル等の手法で埋め込み表現として変換して、Vcsに追加する。
 単語を追加した常識ベクトルをV’csとし、常識ベクトルV’csの各単語のベクトルをw’csで表現すると、V’cs={w’cs_1,w’cs_2,…,w’cs_(n+m)}となる。なお、mは追加した単語の数である。情報作成部133は、上記のVcsとV’の両方を記憶部12に保存する。
 その後、情報作成部133は、V’csを参照して、非言語情報のベクトルVenv={wenv_1,wenv_2,…,wenv_nenv}を作成する。そして、情報作成部133は、非言語情報のベクトルVenvを出力する。なお、情報作成部133は、非言語情報のベクトルV’envを作成する際、One-hotベクトル等の手法を用いてもよい。
[第2の情報作成部]
 また、第2の情報作成部134は、常識ベクトルを参照し、エージェントのプロフィール情報に含まれる事項の単語をベクトル化した情報を作成する。
 例えば、第2の情報作成部134は、エージェントのプロフィール情報から特徴的な事項の単語を抽出し、抽出した単語を常識ベクトルを用いてベクトル化した情報(プロフィール情報のベクトル)を作成する。なお、図11は、プロフィール情報のベクトルを2次元で表現した例を示している。
 例えば、第2の情報作成部134は、前記した情報作成部133と同様の処理を実行し、V’csを参照して、プロフィール情報のベクトルVag={wag_1,wag_2,…,wag_nag}を作成する。そして、第2の情報作成部134は、プロフィール情報のベクトルVagを出力する。なお、第2の情報作成部134は、プロフィール情報のベクトルVagを作成する際、One-hotベクトル等の手法を用いてもよい。
[フィルタリング部]
 フィルタリング部135は、非言語情報のベクトルVenv、プロフィール情報のベクトルVag、距離の閾値を用いて、フィルタリング処理を行う。
 例えば、フィルタリング部135は、非言語情報のベクトルVenvのエレメント{wenv_1,wenv_2,…,wenv_nenv}とプロフィール情報のベクトルVagのエレメント{wag_1,wag_2,…,wag_nag}とを同じベクトル空間にプロットして、非言語情報のベクトルのエレメントとプロフィール情報のベクトルのエレメントとの距離を測定する。距離の測定には、例えば、ユークリッド距離を用いる。
 そして、フィルタリング部135は、非言語情報のベクトルのエレメントのうち、プロフィール情報のベクトルのエレメントからの距離が閾値以下のエレメント(単語)に「重要」を示すフラグを付ける。
 例えば、フィルタリング部135は、図12に示すように、非言語情報のベクトルのエレメントと、プロフィール情報のベクトルのエレメントとの距離を求める。そして、フィルタリング部135は、非言語情報のベクトルのエレメントの中から、プロフィール情報のベクトルのエレメントとの距離が閾値(例えば、0.2)以下のエレメント(例えば、「Shibuya」、「Billboard」)を抽出し、「重要」を示すフラグを付ける。
 その後、フィルタリング部135は、非言語情報のベクトルのエレメントおよびプロフィール情報のベクトルのエレメントの中から、上記の「Shibuya」、「Billboard」からの距離が所定の閾値(例えば、0.1)以下の単語(例えば、「Shibuya station」、「Billboard design」)を、対話の話題に用いる単語として抽出し、出力する。
[処理手順の例]
 図13を用いて、情報提示装置10が実行する処理手順の例を説明する。図13のS11、S12の処理は、図9のS1、S2と同じなので、図13のS13から説明する。第1の情報作成部131は、常識ベクトルを参照して、S12でテキスト化した非言語情報から非言語情報のベクトルを作成する(S13)。
 また、第2の情報作成部134は、エージェントのプロフィール情報を取得する(S14)。そして、第2の情報作成部134は、常識ベクトルを参照して、S14で取得したプロフィール情報のベクトルを作成する(S15)。
 その後、フィルタリング部135は、S13で作成した非言語情報のベクトルのエレメントと、S15で作成したプロフィール情報のベクトルのエレメントとの距離を算出する(S16)。
 そして、フィルタリング部135は、非言語情報のベクトルに含まれるエレメントの中から、プロフィール情報のベクトルのエレメントとの距離が所定の閾値以下のエレメントの単語を、当該エージェントの重要単語として抽出する(S17:重要単語の抽出)。
 S17の後、フィルタリング部135は、非言語情報のベクトルのエレメントおよびプロフィール情報のベクトルのエレメントの中から、S17で抽出された重要単語からの距離が所定の閾値以下のエレメントの単語を、対話の話題に用いる単語として抽出し、出力する(S18)。
 情報提示装置10が上記の処理を実行することにより、自己開示が苦手な対話相手に対しても質の高い対話を提供することができる。
 つまり、情報提示装置10は、ユーザやユーザの周囲についての非言語情報からユーザの個性や嗜好に関連性の高い情報を抽出して、その情報を対話の生成に活用することができる。これにより、自己開示が苦手なユーザに対してもユーザの非言語情報に基づき当該ユーザに即した質の高い対話が可能になる。
[その他の実施形態]
 なお、各実施形態のフィルタリング部135は、重要単語を抽出する際、単語間の距離を表す点数(例えば、距離が短いほど点数を低くする)を用いてもよい。その場合、フィルタリング部135は、所定の閾値以下の点数の単語をすべて抽出してもよいし、点数が低い順に所定数の単語を抽出してもよい。また、フィルタリング部135は、重要単語からの距離を用いて、対話の話題に用いる単語を抽出する際にも、上記と同様に単語間の距離を表す点数を用いてもよい。
 また、各実施形態の情報提示装置10は、フィルタリング部135から出力された単語に基づき、エージェントから対話相手への発話を生成する発話生成部(例えば、図3に対話模擬部)をさらに備えていてもよい。例えば、発話生成部は、フィルタリング部135から出力された単語を話題として用いたエージェントの発話を生成し、出力する。
[システム構成等]
 また、図示した各部の各構成要素は機能概念的なものであり、必ずしも物理的に図示のように構成されていることを要しない。すなわち、各装置の分散・統合の具体的形態は図示のものに限られず、その全部又は一部を、各種の負荷や使用状況等に応じて、任意の単位で機能的又は物理的に分散・統合して構成することができる。さらに、各装置にて行われる各処理機能は、その全部又は任意の一部が、CPU及び当該CPUにて実行されるプログラムにて実現され、あるいは、ワイヤードロジックによるハードウェアとして実現され得る。
 また、前記した実施形態において説明した処理のうち、自動的に行われるものとして説明した処理の全部又は一部を手動的に行うこともでき、あるいは、手動的に行われるものとして説明した処理の全部又は一部を公知の方法で自動的に行うこともできる。この他、上記文書中や図面中で示した処理手順、制御手順、具体的名称、各種のデータやパラメータを含む情報については、特記する場合を除いて任意に変更することができる。
[プログラム]
 前記した情報提示装置10は、パッケージソフトウェアやオンラインソフトウェアとしてプログラム(情報提示プログラム)を所望のコンピュータにインストールさせることによって実装できる。例えば、上記のプログラムを情報処理装置に実行させることにより、情報処理装置を情報提示装置10として機能させることができる。ここで言う情報処理装置にはスマートフォン、携帯電話機やPHS(Personal Handyphone System)等の移動体通信端末、さらには、PDA(Personal Digital Assistant)等の端末等がその範疇に含まれる。
 図14は、情報提示プログラムを実行するコンピュータの一例を示す図である。コンピュータ1000は、例えば、メモリ1010、CPU1020を有する。また、コンピュータ1000は、ハードディスクドライブインタフェース1030、ディスクドライブインタフェース1040、シリアルポートインタフェース1050、ビデオアダプタ1060、ネットワークインタフェース1070を有する。これらの各部は、バス1080によって接続される。
 メモリ1010は、ROM(Read Only Memory)1011及びRAM(Random Access Memory)1012を含む。ROM1011は、例えば、BIOS(Basic Input Output System)等のブートプログラムを記憶する。ハードディスクドライブインタフェース1030は、ハードディスクドライブ1031に接続される。ディスクドライブインタフェース1040は、ディスクドライブ1041に接続される。例えば磁気ディスクや光ディスク等の着脱可能な記憶媒体が、ディスクドライブ1041に挿入される。シリアルポートインタフェース1050は、例えばマウス1110、キーボード1120に接続される。ビデオアダプタ1060は、例えばディスプレイ1130に接続される。
 ハードディスクドライブ1031は、例えば、OS1091、アプリケーションプログラム1092、プログラムモジュール1093、プログラムデータ1094を記憶する。すなわち、上記の情報提示装置10が実行する各処理を規定するプログラムは、コンピュータにより実行可能なコードが記述されたプログラムモジュール1093として実装される。プログラムモジュール1093は、例えばハードディスクドライブ1031に記憶される。例えば、情報提示装置10における機能構成と同様の処理を実行するためのプログラムモジュール1093が、ハードディスクドライブ1031に記憶される。なお、ハードディスクドライブ1031は、SSD(Solid State Drive)により代替されてもよい。
 また、上述した実施形態の処理で用いられるデータは、プログラムデータ1094として、例えばメモリ1010やハードディスクドライブ1031に記憶される。そして、CPU1020が、メモリ1010やハードディスクドライブ1031に記憶されたプログラムモジュール1093やプログラムデータ1094を必要に応じてRAM1012に読み出して実行する。
 なお、プログラムモジュール1093やプログラムデータ1094は、ハードディスクドライブ1031に記憶される場合に限らず、例えば着脱可能な記憶媒体に記憶され、ディスクドライブ1041等を介してCPU1020によって読み出されてもよい。あるいは、プログラムモジュール1093及びプログラムデータ1094は、ネットワーク(LAN(Local Area Network)、WAN(Wide Area Network)等)を介して接続される他のコンピュータに記憶されてもよい。そして、プログラムモジュール1093及びプログラムデータ1094は、他のコンピュータから、ネットワークインタフェース1070を介してCPU1020によって読み出されてもよい。
 10 情報提示装置
 11 入出力部
 12 記憶部
 13 制御部
 131 第1の情報作成部
 132 テキスト化部
 133 情報作成部
 134 第2の情報作成部
 135 フィルタリング部

Claims (7)

  1.  第1のユーザまたは前記第1のユーザの周囲についての非言語情報を取得し、取得した前記非言語情報から特徴的なオブジェクトを抽出し、前記オブジェクト同士の関係性をグラフ化またはベクトル化して示した第1の情報を作成する第1の情報作成部と、
     前記第1のユーザへの発話を行う第2のユーザのプロフィールから特徴的な事項を抽出し、前記事項同士の関係性をグラフ化またはベクトル化して示した第2の情報を作成する第2の情報作成部と、
     前記第1の情報に含まれるオブジェクトの中から、前記第2の情報に含まれる事項との距離が所定の閾値以下のオブジェクトを抽出し、前記オブジェクトを表す単語を出力するフィルタリング部と
     を備えることを特徴とする情報提示装置。
  2.  前記非言語情報は、
     前記第1のユーザの身体的特徴、または、当該第1のユーザの周囲に存在する物体の情報
     であることを特徴とする請求項1に記載の情報提示装置。
  3.  前記第1の情報作成部は、
     取得した前記非言語情報をテキスト化し、テキスト化した前記非言語情報から特徴的なオブジェクトを表す単語を抽出し、前記単語同士の関係性をグラフ化または前記単語同士の意味の近さをベクトル化することにより前記第1の情報を作成し、
     前記第2の情報作成部は、
     前記第2のユーザのプロフィールから特徴的な事項の単語を抽出し、前記単語同士の関係性をグラフ化または前記単語同士の意味の近さをベクトル化することにより前記第2の情報を作成し、
     前記フィルタリング部は、
     単語同士の意味の近さまたは関連性を示したデータベースを参照し、前記第1の情報に含まれる単語と前記第2の情報に含まれる単語との距離を算出することにより、前記第1の情報に含まれるオブジェクトと、前記第2の情報に含まれる事項との距離を算出する
     ことを特徴とする請求項1に記載の情報提示装置。
  4.  前記フィルタリング部は、
     前記第1の情報に含まれる単語および前記第2の情報に含まれる単語の中から、前記オブジェクトを表す単語からの距離が所定の閾値以下である単語を抽出し、出力する
     ことを請求項3に記載の情報提示装置。
  5.  前記フィルタリング部から出力された単語に基づき、前記第2のユーザから前記第1のユーザへの発話を生成する発話生成部
     をさらに備えることを特徴とする請求項1に記載の情報提示装置。
  6.  情報提示装置により実行される情報提示方法であって、
     第1のユーザまたは前記第1のユーザの周囲についての非言語情報を取得し、取得した前記非言語情報から特徴的なオブジェクトを抽出し、前記オブジェクト同士の関係性をグラフ化またはベクトル化して示した第1の情報を作成する工程と、
     前記第1のユーザへの発話を行う第2のユーザのプロフィールから特徴的な事項を抽出し、前記事項同士の関係性をグラフ化またはベクトル化して示した第2の情報を作成する工程と、
     前記第1の情報に含まれるオブジェクトの中から、前記第2の情報に含まれる事項との距離が所定の閾値以下のオブジェクトを表す単語を抽出し、出力する工程と
     を含むことを特徴とする情報提示方法。
  7.  第1のユーザまたは前記第1のユーザの周囲についての非言語情報を取得し、取得した前記非言語情報から特徴的なオブジェクトを抽出し、前記オブジェクト同士の関係性をグラフ化またはベクトル化して示した第1の情報を作成する工程と、
     前記第1のユーザへの発話を行う第2のユーザのプロフィールから特徴的な事項を抽出し、前記事項同士の関係性をグラフ化またはベクトル化して示した第2の情報を作成する工程と、
     前記第1の情報に含まれるオブジェクトの中から、前記第2の情報に含まれる事項との距離が所定の閾値以下のオブジェクトを表す単語を抽出し、出力する工程と
     をコンピュータに実行させるための情報提示プログラム。
PCT/JP2024/015943 2024-04-23 2024-04-23 情報提示装置、情報提示方法、および、情報提示プログラム Pending WO2025224847A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/JP2024/015943 WO2025224847A1 (ja) 2024-04-23 2024-04-23 情報提示装置、情報提示方法、および、情報提示プログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2024/015943 WO2025224847A1 (ja) 2024-04-23 2024-04-23 情報提示装置、情報提示方法、および、情報提示プログラム

Publications (1)

Publication Number Publication Date
WO2025224847A1 true WO2025224847A1 (ja) 2025-10-30

Family

ID=97489724

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2024/015943 Pending WO2025224847A1 (ja) 2024-04-23 2024-04-23 情報提示装置、情報提示方法、および、情報提示プログラム

Country Status (1)

Country Link
WO (1) WO2025224847A1 (ja)

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9812151B1 (en) * 2016-11-18 2017-11-07 IPsoft Incorporated Generating communicative behaviors for anthropomorphic virtual agents based on user's affect

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9812151B1 (en) * 2016-11-18 2017-11-07 IPsoft Incorporated Generating communicative behaviors for anthropomorphic virtual agents based on user's affect

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
CHIKAMI SHINYA, SHUN HATTORI, CHIAKI KUBOMIRA, HIROYUKI KAMEDA: "Proposal of a Method of Situation-dependent Utterance and Behavioral Functions in Dialogue Robots", ANNUAL CONFERENCE OF JSAI, 9 June 2010 (2010-06-09), XP093365894 *
MATSUI TETSUYA, YAMADA SEIJI: "Designing the Trustworthy Virtual Agent by discriminability", PROCEEDINGS DVD OF THE 32ND ANNUAL CONFERENCE OF THE JAPANESE SOCIETY FOR ARTIFICIAL INTELLIGENCE, THE 2018 ANNUAL CONFERENCE OF THE JAPANESE SOCIETY FOR ARTIFICIAL INTELLIGENCE (32ND), 5 June 2018 (2018-06-05), XP093365897, Retrieved from the Internet <URL:https://confit.atlas.jp/guide/event/jsai2018/subject/4J1-02/detail?lang=en> *

Similar Documents

Publication Publication Date Title
CN110348535B (zh) 一种视觉问答模型训练方法及装置
CN114612290A (zh) 图像编辑模型的训练方法和图像编辑方法
WO2020151689A1 (zh) 对话生成方法、装置、设备及存储介质
US20180226067A1 (en) Modifying a language conversation model
WO2020170912A1 (ja) 生成装置、学習装置、生成方法及びプログラム
JP6980411B2 (ja) 情報処理装置、対話処理方法、及び対話処理プログラム
JP7581502B2 (ja) カスタマイズ可能なチャットボットを実行するための構成可能な会話エンジン
CN117891927A (zh) 基于大语言模型的问答方法、装置、电子设备及存储介质
CN114360488A (zh) 语音合成、语音合成模型训练方法、装置及存储介质
CN112685550A (zh) 智能问答方法、装置、服务器及计算机可读存储介质
CN115762484A (zh) 用于语音识别的多模态数据融合方法、装置、设备及介质
CN118827411A (zh) 一种基于大语言模型构建数字孪生网络的方法及相关装置
CN110245349A (zh) 一种句法依存分析方法、装置及一种电子设备
CN114490967B (zh) 对话模型的训练方法、对话机器人的对话方法、装置和电子设备
US20230103313A1 (en) User assistance system
JPWO2018173943A1 (ja) データ構造化装置、データ構造化方法およびプログラム
JP2020190585A (ja) 自動対話装置、自動対話方法、およびプログラム
JP4824043B2 (ja) 自然言語対話エージェントの知識構造構成方法、知識構造を用いた自動応答の作成方法および自動応答作成装置
JP2022116979A (ja) 文章生成装置、プログラムおよび文章生成方法
JP3950957B2 (ja) 言語処理装置および方法
WO2023119521A1 (ja) 可視化情報生成装置、可視化情報生成方法、及びプログラム
JP7776909B1 (ja) アイデア拡張システム、アイデア拡張方法およびアイデア拡張プログラム
JP7734115B2 (ja) 応対フロー作成支援装置、及び応対フロー作成支援方法
CN114596568B (zh) 一种对扫描图像的智能文字识别方法、装置及存储介质
JP7726275B2 (ja) 対話装置、対話制御方法及び対話プログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24937138

Country of ref document: EP

Kind code of ref document: A1