WO2023181986A1 - 情報処理方法、情報処理プログラム及び情報処理装置 - Google Patents

情報処理方法、情報処理プログラム及び情報処理装置 Download PDF

Info

Publication number
WO2023181986A1
WO2023181986A1 PCT/JP2023/009268 JP2023009268W WO2023181986A1 WO 2023181986 A1 WO2023181986 A1 WO 2023181986A1 JP 2023009268 W JP2023009268 W JP 2023009268W WO 2023181986 A1 WO2023181986 A1 WO 2023181986A1
Authority
WO
WIPO (PCT)
Prior art keywords
level
information processing
breakdown
user
answer
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2023/009268
Other languages
English (en)
French (fr)
Inventor
洋一 松山
真於 佐伯
駿吾 鈴木
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Waseda University
Original Assignee
Waseda University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Waseda University filed Critical Waseda University
Publication of WO2023181986A1 publication Critical patent/WO2023181986A1/ja
Priority to US18/895,299 priority Critical patent/US20250014592A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09BEDUCATIONAL OR DEMONSTRATION APPLIANCES; APPLIANCES FOR TEACHING, OR COMMUNICATING WITH, THE BLIND, DEAF OR MUTE; MODELS; PLANETARIA; GLOBES; MAPS; DIAGRAMS
    • G09B19/00Teaching not covered by other main groups of this subclass
    • G09B19/04Speaking
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L13/00Speech synthesis; Text to speech systems
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L13/00Speech synthesis; Text to speech systems
    • G10L13/08Text analysis or generation of parameters for speech synthesis out of text, e.g. grapheme to phoneme translation, prosody generation or stress or intonation determination
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/10Speech classification or search using distance or distortion measures between unknown speech and reference templates
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination

Definitions

  • the present invention relates to an information processing method, an information processing program, and an information processing device.
  • Patent Document 1 selects a question from pre-prepared questions and outputs it to the interview examinee, analyzes the response to the question using video and audio, and analyzes the content characteristics and communication characteristics of the response. and evaluate the response. Further, the information processing program uses the evaluation of the responses to select questions, and also aggregates the evaluations of all the responses to evaluate the interview.
  • questions are selected according to the evaluation of each response, but since the questions are asked in advance according to the development of the story, the questions are selected based on the successive evaluations of the interviewee. There was a problem in that it was not possible to adaptively select questions at an appropriate level. Further, there is a problem in that no means for verifying the evaluation results is provided, and accurate judgments cannot always be made.
  • An object of the present invention is to provide an information processing method, an information processing program, and an information processing device that ask questions that take the evaluator's level into consideration and improve the accuracy of the evaluator's level determination.
  • one aspect of the present invention provides the following information processing method, information processing program, and information processing apparatus.
  • the determining step determines the level of the interlocutor for each predetermined unit of the answer, and utters a question at a level higher than the determined level.
  • speech control means for controlling the speech of questions at one level among a plurality of predetermined levels to the interlocutor; detection means for detecting a breakdown of the interlocutor in the answer; and determining means for determining the level of the interlocutor based on at least the level at which the breakdown was detected.
  • FIG. 1 is a schematic diagram showing an example of the configuration of an information processing system according to an embodiment.
  • FIG. 2 is a block diagram illustrating a configuration example of an information processing device according to an embodiment.
  • FIG. 3 is a block diagram illustrating a configuration example of a terminal according to an embodiment.
  • FIG. 4 is a schematic diagram showing an example of a screen displayed on the display of the terminal.
  • FIG. 5 is a schematic diagram showing the questions asked by the speech control means of the information processing device and the contents of the user's answers.
  • FIG. 6 is a graph diagram showing the relationship between the ability determination level of the information processing device and the passage of time.
  • FIG. 7 is a flowchart showing an example of the operation of the information processing device.
  • FIG. 1 is a schematic diagram showing an example of the configuration of an information processing system according to an embodiment.
  • this information processing system includes an information processing device 1 that interacts with a user 4 as an interlocutor and determines English conversation ability as a foreign language, and reproduces information generated in the information processing device 1, and 4 and a terminal 2 for receiving a response are connected to each other via a network 3 so as to be able to communicate with each other.
  • the information processing device 1 and the terminal 2 may be configured integrally, and in that case, the network 3 can be omitted.
  • the information processing device 1 is a server-type information processing device that operates in response to requests from the user 4 via the terminal 2, and includes a CPU (Central CPU) having a function for processing information in the main body. It is equipped with electronic components such as a processing unit (Processing Unit) and a flash memory that has a function for storing information.
  • a CPU Central CPU
  • the terminal 2 is a terminal device such as a PC (Personal Computer), a tablet, or a smartphone, and has a CPU, flash memory, and other electronic components such as a speaker, microphone, and camera that have functions for processing information in the main body. Be prepared.
  • the network 3 is a communication network capable of high-speed communication, and is, for example, a wired or wireless communication network such as the Internet, an intranet, or a LAN (Local Area Network).
  • a wired or wireless communication network such as the Internet, an intranet, or a LAN (Local Area Network).
  • the information processing device 1 displays an avatar on the display unit of the terminal 2 via the network 3, causes the avatar to speak a question, and gives the avatar an action.
  • the information processing device 1 collects the answers to the questions of the user 4 through the microphone of the terminal 2, performs speech recognition on the contents, and determines the user's conversational ability in response to the questions. Judgment of ability is performed for each predetermined conversation group, such as for each topic or question, and the level of the judgment result is fed back as appropriate to the selection of the next topic or question.
  • FIG. 2 is a block diagram showing a configuration example of the information processing device 1 according to the embodiment.
  • the information processing device 1 includes a control unit 10 that is composed of a CPU and the like and controls each unit and executes various programs, a storage unit 11 that is composed of a storage medium such as a flash memory and stores information, and a network 3. and a communication section 12 that communicates with the outside via the communication section 12.
  • the control unit 10 executes the ability determination program 110 as an information processing program to be described later, thereby controlling the speech control means 100, the display control means 101, the motion imparting means 102, the voice recognition means 103, the video recognition means 104, and the ability determination means. 105, functions as a breakdown detection means 106, etc.
  • the speech control means 100 controls voice speech on the terminal 2.
  • the utterance control means 100 selects a question according to the level of the user 4 from question information 111 that mainly includes questions of a plurality of levels prepared in advance, and utters the question.
  • the voice utterances include not only questions based on the question information 111, but also greetings to the user 4 in response to the questions, utterances by the user 4, and comments in response to the answers.
  • the display control means 101 displays the avatar on the terminal 2 while adding motion using avatar information 112 that defines the image of the avatar and motion information 113 that defines the motion of the avatar.
  • the motion imparting means 102 applies motions to the avatar displayed on the terminal 2 by the display control means 101 in accordance with the motions of the speech control means 100, the voice recognition means 103, and the video recognition means 104, with reference to the motion information 113.
  • a speech motion is given to the avatar in accordance with the speech.
  • All the motions in the motion information 113 are associated with, for example, a certain utterance, and the motion whose inter-sentence distance is closest to the utterance of the motion to be generated is selected.
  • actions such as listening gestures are given to the avatar according to the actions of the voice recognition means 103, and actions that react to the actions of the user 4 are given to the avatar according to the actions of the video recognition means 104.
  • the motion imparting means 102 may impart motions in accordance with the motions of the ability determining means 105 and the breakdown detecting means 106.
  • the voice recognition means 103 recognizes the voice accompanying the user 4's answer to the question etc. of the speech control means 100 received via the terminal 2, and stores it in the storage unit 11 as answer information 114.
  • the speech recognition means 103 may further perform language understanding processing on the recognized speech, or the language understanding may be performed by the ability determination means 105.
  • methods such as GMM-HMM, DNN-HMM, and End-to-End DNN can be employed, and for language understanding, methods such as keyword extraction, decision trees, and neural networks can be employed.
  • the video recognition means 104 recognizes the video including gestures, glances, gestures, etc. associated with the user 4's answer received via the terminal 2, and stores it in the storage unit 11 as answer information 114.
  • image recognition for example, a means such as a neural network can be adopted.
  • the ability determining means 105 determines the ability of the user 4 based on the content of the answer information 114 and stores it in the storage unit 11 as determination result information 115.
  • a criterion for example, CEFR (Common European Framework of Reference for Languages) is used. Specifically, A1, A2, B1, B2, C1, and C2 are prepared, and the levels increase in this order.
  • the ability can be determined using means such as linear regression, decision trees, and neural networks, for example.
  • the ability determination means 105 feeds back the provisionally determined level to the speech control means 100, and causes the speech control means 100 to select a question according to the level. Let them speak. The feedback operation will be explained in detail later. Further, the ability determining means 105 comprehensively determines the level of the user 4 after obtaining answers to the plurality of questions.
  • the breakdown detection means 106 detects this as a breakdown.
  • breakdown is defined as above based on the case of English conversation ability assessment, but if the content of the assessment is different or the application is different from the assessment, the content of the assessment or the application is different. You may change the breakdown definition accordingly. If we define it in terms of a higher level concept, it is a break that the content of the response of the user 4 to the content sent by the information processing device 1 deviates from the normal response distribution, resulting in an interactive breakdown. Down.
  • the breakdown detection means 106 detects incomprehension from a state where the user is requesting repetition of the utterance, or a state where the user is confused or silent due to thinking about understanding. Regarding the latter, characteristic actions that occur when users do not understand are, for example, averting their gaze, bringing their faces closer, blinking frequently, looking to the side, violently moving their gaze, violently moving their head, becoming silent, etc.
  • the voice recognition means 103 or the image recognition means 104 detects the following eight conditions: a decrease in the volume of the utterance, it is determined that the person is confused or is in a silent state thinking about understanding.
  • the breakdown detection means 106 detects disfluency from a state in which speech production stagnates due to difficulty in recalling linguistic knowledge such as vocabulary, grammar, and pronunciation.
  • linguistic knowledge such as vocabulary, grammar, and pronunciation.
  • silence occurs in the middle of a sentence or clause, it is judged to be a state in which speech production is stagnant. This is because when silence occurs at the beginning or end of a sentence or clause, it is thought to be caused by content recall, such as a lack of background knowledge or a failure to plan an effective discourse structure.
  • the storage unit 11 stores an ability determination program 110 that causes the control unit 10 to operate as each of the above-mentioned means 100-106, question information 111, avatar information 112, operation information 113, answer information 114, determination result information 115, and the like.
  • FIG. 3 is a block diagram showing a configuration example of the terminal 2 according to the embodiment.
  • the terminal 2 includes a control unit 20 that is composed of a CPU, etc., and that controls each unit and executes various programs, a storage unit 21 that is composed of a storage medium such as a flash memory, and stores information, and a network 3.
  • a communication unit 22 that communicates with the outside, a microphone 23 that converts the input voice of the user 4 into an electrical signal, a speaker 24 that converts the signal input from the control unit 20 into voice and outputs it, and the user 4 It is equipped with a camera 25 that takes an image and outputs a video signal, and a display 26 such as an LCD that displays images, videos, characters, etc.
  • the terminal 2 includes an operation unit (not shown) (keyboard, mouse, trackpad, touch panel), etc. that receives operations from the user 4.
  • the terminal 2 may be provided with some or all of the means of the information processing device 1, and the design may be modified as appropriate without departing from the spirit of the invention. May be changed.
  • the user 4 operates the terminal 2 to request a determination of English conversation ability, for example.
  • the terminal 2 communicates with the information processing device 1 via the network 3 and requests the information processing device 1 to determine English conversation ability.
  • the information processing device 1 When the information processing device 1 receives a request for English conversation ability determination from the terminal 2, it issues instructions to the speech control means 100, the display control means 101, and the action imparting means 102, and displays the display of the terminal 2 as shown in FIG. 26 displays the avatar.
  • FIG. 4 is a schematic diagram showing an example of a screen displayed on the display 26 of the terminal 2.
  • the screen 101a has an area 101a 1 that displays an image of the user 4 taken by the camera 25 of the terminal 2, and an area 101a 2 that displays an avatar.
  • the area 101a1 is mainly displayed for reference by the user 4, but may not be displayed.
  • the avatar displayed in the area 101a2 is given an action by the display control means 101 and the action giving means 102, and a sound is output from the speaker 24 by the speech control means 100.
  • the user 4's reactions (voice, facial expressions, actions, etc.) to the screen 101a displayed on the display 26 and the audio output from the speaker 24 are transmitted to the information processing device 1 via the microphone 23 and camera 25 of the terminal 2. are input and recognized by the voice recognition means 103 and video recognition means 104 of the information processing device 1, respectively.
  • the speech control means 100 of the information processing device 1 controls the voice speech of the avatar on the terminal 2.
  • the utterance control means 100 selects a question according to the level of the user 4 from question information 111 that mainly includes questions of a plurality of levels prepared in advance, and utters the question.
  • the display control means 101 of the information processing device 1 displays the avatar on the terminal 2 using avatar information 112 that defines an image of the avatar and motion information 113 that defines the motion of the avatar, and also displays the avatar on the terminal 2.
  • the imparting means 102 imparts motions to the avatar displayed on the terminal 2 by the display control means 101 in accordance with the motions of the speech control means 100, the voice recognition means 103, and the video recognition means 104, with reference to the motion information 113. do.
  • An action such as a listening gesture (listening action) is given to the avatar according to the action of the voice recognition means 103, and an action (reaction) to react to the action of the user 4 is given to the avatar according to the action of the video recognition means 104. .
  • the voice recognition means 103 of the information processing device 1 recognizes the voice accompanying the answer of the user 4 to the question etc. of the speech control means 100 received via the terminal 2, and stores it in the storage unit 11 as answer information 114. .
  • the speech recognition means 103 further processes language understanding for the recognized speech.
  • the image recognition means 104 of the information processing device 1 recognizes the image including gestures, glances, gestures, etc. associated with the answer of the user 4 received via the terminal 2, and stores it in the storage unit 11 as answer information 114. .
  • the ability determining means 105 of the information processing device 1 determines the ability of the user 4 based on the content of the answer information 114 and stores it in the storage unit 11 as determination result information 115.
  • the breakdown detection means 106 of the information processing device 1 detects this as a breakdown. do.
  • FIG. 5 is a schematic diagram showing an example of the contents of questions (S1 to S10) by the speech control means 100 of the information processing device 1 and answers (U1 to U9) by the user 4.
  • Phase 1 shows the questions and answers during the introduction operation
  • Phase 2 shows the questions and answers during the level check operation
  • Phase 3 shows the questions and answers during the push-up operation.
  • FIG. 6 is a graph diagram showing the relationship between the level of ability determination of the information processing device 1 and the passage of time.
  • FIG. 7 is a flowchart showing an example of the operation of the information processing device 1.
  • the information processing device 1 uses the speech control means 100 to perform an introductory operation to have a relatively simple conversation with the user 4, such as greeting or small talk, to relieve tension and to grasp the general level.
  • the display control means 101 and the motion imparting means 102 give the avatar a motion, and the user 4's answer, facial expression, and motion are uttered, for example, by asking a question of level A1 (S100), and the voice recognition means 103 and the video recognition means 104 While recognizing this, the ability of the user 4 is determined by the ability determining means 105 (S101) (phase 1 in FIG. 6).
  • the speech control means 100, the display control means 101, and the action imparting means 102 select a topic such as "What is your favorite season?" (S1) using an avatar. Ask introductory questions.
  • the voice recognition means 103 and the video recognition means 104 recognize the user 4's answer, facial expressions, and actions.
  • the ability determining means 105 determines the ability of the user 4. For example, assume that the CEFR level is determined to be A2.
  • the speech control means 100, the display control means 101, and the action imparting means 102 ask "Are there any activities you like to do in winter?” (S2) Additional questions are asked.
  • the number of times the additional questions are asked may be 0, once, or multiple times as long as the ability determining means 105 can confirm the determination result.
  • the speech control means 100, display control means 101, and action imparting means 102 cause the avatar to say, for example, "Alright. What did you eat for breakfast this?"morning?'' (S4), and asks a topic introduction question according to the A2 level, which is the ability determined in step S101 (S102).
  • the voice recognition means 103 the video recognition means 10 User by 4 Recognize the answer to step 4 as well as facial expressions and actions.
  • the ability of the user 4 is determined by the ability determining means 105 (S103) (phase 2 in FIG. 6). For example, assume that the CEFR level is determined to be A2. Steps S102 and S103 may be performed only once, but in order to confirm that the user 4 can reliably conduct an A2 level conversation, in this embodiment, additional questions are asked multiple times to determine the ability. Note that the ability determination may be performed for each question, or for each predetermined unit such as for each plurality of questions or for each topic.
  • the speech control means 100, the display control means 101, and the action imparting means 102 ask an additional question such as "Do you normally eat breakfast?" (S5).
  • the voice recognition means 103 and the video recognition means 104 recognized the answers, facial expressions, and movements of the user 4. As a result, the answers were made without any problems, so the ability determination means 105 evaluated the ability of the user 4 using the CEFR. It is provisionally determined to be A2 level (S103). Strictly speaking, since the level has not been determined, it is determined that the level is higher than the CEFR A2 level (that is, the lower limit level is A2).
  • the speech control means 100, the display control means 101, and the motion imparting means 102 lower the level than the determination result of step S103. and ask CEFR B1 level topic introduction questions (S104). Specifically, as shown in phase 3 of FIG. 5, the avatar asks a question such as "Have you ever been to a foreign country?" (S7). Note that the level is not limited to being raised by 1, but may be raised by 2 or more, the degree of raising the level may be changed depending on the answer, and furthermore, the level may be lowered depending on the purpose.
  • the breakdown detection means 106 does not detect incomprehension or disfluency in the recognized answer of the user 4 (S105; No), the breakdown detection means 106 would normally ask the question at a higher level (S106), but in this implementation. In order to confirm that the user 4 can reliably conduct a conversation at the B1 level, additional questions are asked multiple times at the B1 level.
  • the speech control means 100, the display control means 101, and the action imparting means 102 say "OK. Which country would you like to visit in the future?” ?” (S8).
  • the breakdown detection means 106 detects disfluency, but in order to confirm whether this is correct, Furthermore, a continuation request is made with content such as "Why is that?" (S8).
  • the breakdown detection means 106 recognizes Since disfluency was detected in the answer of user 4, it is determined that a breakdown has been detected (S105; Yes), and "That's OK. Let's move on.” (S10) make an utterance such as
  • the ability determination means 105 determines that the ability of the user 4 is not CEFR level B1 but level A2 (S107) (as shown in FIG. Phase 3). Further, the ability determining means 105 outputs the determination result to the user 4 or to any administrator other than the user 4 (S108).
  • step S103 the determination result in step S103 is corrected and the level is raised to B2 (S106 ), steps S102 to S106 are repeated, and the ability is finally determined (S107).
  • the ability determining means 105 may feed back the level determined in step S107 to the speech control means 100 as a provisional determination result, and may again select and speak a question according to the level (S102). .
  • the ability determining means 105 may comprehensively determine the level of the user 4 after repeating steps S102 to S107 described above multiple times. In other words, it may be determined that the level is A2 after repeating "(3) Level check operation” and "(4) Pushing up motion" multiple times (second phase 2 and phase 3 in FIG. 6). Further, a cool down phase may be further provided after the ability determination is completed (Cool down after phase 3 in FIG. 6).
  • the ability determining means 105 determines the ability of the user 4, asks a higher level question in order to verify the determined ability, and in response to the question, the breakdown detecting means 106 determines the ability of the user 4. If a breakdown is detected, it is determined that the pre-determined ability is correct rather than the raised level, and if a breakdown is not detected, the level is raised further and the operation is continued to improve user 4's ability. Since the ability is judged, it is possible to ask questions that take into account the level of the user 4 (person being evaluated) whose ability is being judged, and by trying out other levels, the level can be judged accurately and with confidence. be able to.
  • the avatar adds gestures (listening motions, reactions) when speaking and accepting answers, it is possible to ask and answer questions in a more natural manner, which in turn encourages users to self-disclose and collect information. It can be pulled out.
  • the level is determined provisionally based on conversation, topic, etc., and questions are selected according to the level, so you do not have to ask all the questions at random as you would do if the level was not determined provisionally. This improves efficiency in terms of the time required to answer questions.
  • each means 100 to 106 of the control unit 10 are realized by a program, but all or part of each means may be realized by hardware such as an ASIC.
  • the programs used in the above embodiments can also be provided by being stored in a recording medium such as a CD-ROM.
  • the above steps explained in the above embodiments can be replaced, deleted, added, etc. without changing the gist of the present invention.
  • each means 100 to 106 of the control unit 10 of the information processing device 1 do not necessarily need to be realized on the information processing device 1, and may be realized on the terminal 2 without changing the gist of the present invention. It's okay.
  • each function of the terminal 2 does not necessarily need to be realized on the terminal 2, and may be realized on the information processing device 1 without changing the gist of the present invention.
  • an information processing method Provided are an information processing method, an information processing program, and an information processing device that ask questions that take the level of the evaluator into consideration and improve the accuracy of determining the evaluator's level.
  • Information processing device 2 Terminal 3: Network 4: User 10: Control unit 11: Storage unit 12: Communication unit 20: Control unit 21: Storage unit 22: Communication unit 23: Microphone 24: Speaker 25: Camera 26: Display 100: Speech control means 101: Display control means 102: Action imparting means 103: Voice recognition means 104: Image recognition means 105: Ability judgment means 106: Breakdown detection means 110: Ability judgment program 111: Question information 112: Avatar information 113: Operation information 114: Answer information 115: Judgment result information

Landscapes

  • Engineering & Computer Science (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Multimedia (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Computational Linguistics (AREA)
  • Business, Economics & Management (AREA)
  • Entrepreneurship & Innovation (AREA)
  • General Health & Medical Sciences (AREA)
  • Signal Processing (AREA)
  • Educational Administration (AREA)
  • Educational Technology (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Electrically Operated Instructional Devices (AREA)
  • Machine Translation (AREA)

Abstract

【課題】評価者のレベルを考慮した質問をするとともに、評価者のレベル判定の正確性を向上する情報処理方法、情報処理プログラム及び情報処理装置を提供する。 【解決手段】情報処理装置1は、対話者である利用者4に対して予め定めた複数のレベルのうち一のレベルの質問を発話制御する発話制御手段100と、利用者4の質問に対する回答を音声認識する音声認識手段103と、回答において利用者4のブレイクダウンを検出するブレイクダウン検出手段106と、少なくともブレイクダウンを検出したレベルに基づいて利用者4のレベルを判定する能力判定手段105とを有する。

Description

情報処理方法、情報処理プログラム及び情報処理装置
 本発明は、情報処理方法、情報処理プログラム及び情報処理装置に関する。
 従来の技術として、質問に対する面接受験者の応答を評価して面接を採点する情報処理方法が提案されている(例えば、特許文献1参照)。
 特許文献1に開示された情報処理方法は、予め用意された質問から質問を選択して面接受験者に出力し、当該質問に対する応答を映像及び音声で解析し、当該応答の内容特徴と伝達特徴とを抽出し、応答を評価する。また、情報処理プログラムは、当該応答の評価を質問の選択に用いるとともに、すべての応答の評価を集計して面接を評価する。
米国特許第1067188号明細書
 上記情報処理方法によると、各応答の評価に応じて質問を選択するものの、話の展開に応じて事前に決められた質問を行うものであるため、逐次的な面接受験者の評価に基づいて適切なレベルの質問を適応的に選定することができない、という問題があった。また、評価結果を検証する手段が用意されるものではなく、必ずしも正確な判定を行えるとは限らない、という問題があった。
 本発明の目的は、評価者のレベルを考慮した質問をするとともに、評価者のレベル判定の正確性を向上する情報処理方法、情報処理プログラム及び情報処理装置を提供することにある。
 本発明の一態様は、上記目的を達成するため、以下の情報処理方法、情報処理プログラム及び情報処理装置を提供する。
[1]コンピュータに、
 対話者に対して予め定めた複数のレベルのうち一のレベルの質問を発話する発話制御ステップと、
 前記回答において前記対話者のブレイクダウンを検出する検出ステップと、
 少なくとも前記ブレイクダウンを検出したレベルに基づいて前記対話者のレベルを判定する判定ステップとを実行させる情報処理方法。
[2]前記判定ステップは、予め定めた前記回答の単位毎に前記対話者のレベルを判定し、当該レベルより上のレベルの質問を発話する前記[1]に記載の情報処理方法。
[3]ブレイクダウンを検出した場合、前記複数のレベルの質問のうち当該ブレイクダウンを検出したレベル以外のレベルの質問を発話する前記[1]又は[2]に記載の情報処理方法。
[4]前記判定ステップは、前記ブレイクダウンの検出に関わらず予め定めた前記回答の単位毎に前記対話者のレベルを判定し、
 前記発話制御ステップは、前記判定ステップにおいて判定したレベルの質問を発話する前記[1]から[3]のいずれかに記載の情報処理方法。
[5]前記発話制御ステップの発話に連動して動作するアバターを表示制御する表示制御ステップをさらに実行させる前記[1]から[4]のいずれかに記載の情報処理方法。
[6]前記アバターに対して前記回答を傾聴する動作を付与する動作付与ステップをさらに実行させる前記[5]に記載の情報処理方法。
[7]当該アバターに対して前記回答に応答する動作を付与する動作付与ステップをさらに実行させる前記[5]に記載の情報処理方法。
[8]コンピュータを、
 対話者に対して予め定めた複数のレベルのうち一のレベルの質問を発話制御する発話制御手段と、
 前記回答において前記対話者のブレイクダウンを検出する検出手段と、
 少なくとも前記ブレイクダウンを検出したレベルに基づいて前記対話者のレベルを判定する判定手段としてさらに機能させる情報処理プログラム。
[9]対話者に対して予め定めた複数のレベルのうち一のレベルの質問を発話制御する発話制御手段と、
 前記回答において前記対話者のブレイクダウンを検出する検出手段と、
 少なくとも前記ブレイクダウンを検出したレベルに基づいて前記対話者のレベルを判定する判定手段とを有する情報処理装置。
 本願発明によれば、評価者のレベルを考慮した質問をするとともに、評価者のレベル判定の正確性を向上することができる。
図1は、実施の形態に係る情報処理システムの構成の一例を示す概略図である。 図2は、実施の形態に係る情報処理装置の構成例を示すブロック図である。 図3は、実施の形態に係る端末の構成例を示すブロック図である。 図4は、端末のディスプレイに表示される画面例を示す概略図である。 図5は、情報処理装置の発話制御手段による質問と、利用者の回答の内容を示す概略図である。 図6は、情報処理装置の能力判定のレベルと時間経過の関係を示すグラフ図である。 図7は、情報処理装置の動作例を示すフローチャートである。
[実施の形態]
(情報処理システムの構成)
 図1は、実施の形態に係る情報処理システムの構成の一例を示す概略図である。
 この情報処理システムは、一例として、対話者としての利用者4と対話して外国語としての英会話能力を判定する情報処理装置1と、情報処理装置1において生成された情報を再生し、利用者4の応答を受け付けるための端末2とを、ネットワーク3によって互いに通信可能に接続することで構成される。なお、情報処理装置1と端末2とを一体に構成してもよく、その場合はネットワーク3を省略することができる。
 情報処理装置1は、サーバ型の情報処理装置であり、端末2を介した利用者4の要求に応じて動作するものであって、本体内に情報を処理するための機能を有するCPU(Central Processing Unit)や情報を記憶するための機能を有するフラッシュメモリ等の電子部品を備える。
 端末2は、PC(Personal Computer)やタブレット、スマートフォン等の端末装置であって、本体内に情報を処理するための機能を有するCPUやフラッシュメモリ、その他にスピーカー、マイク、カメラ等の電子部品を備える。
 ネットワーク3は、高速通信が可能な通信ネットワークであり、例えば、インターネット、イントラネットやLAN(Local Area Network)等の有線又は無線の通信網である。
 上記構成において、一例として、英会話の能力を判定するため、情報処理装置1は、ネットワーク3を介して端末2の表示部にアバターを表示処理し、質問を発話させるとともにアバターに動作を付与する。情報処理装置1は端末2のマイクを介して利用者4の質問に対する回答を集音し、内容を音声認識して質問に対する会話能力を判定する。能力の判定は、話題毎や質問毎等のように予め定められた会話のまとまり毎に行われ、判定結果のレベルは適宜次の話題や質問の選択にフィードバックされる。また、レベルが判定されると、判定されたレベルを確信のあるものとするため、敢えて判定したレベルより高い質問を行い、利用者4が質問に対して不理解、非流暢性を有する回答をした場合、先に判定したレベルが正しいものと判断し、判定結果を出力する。以降、構成についてさらに詳しく説明する。
(情報処理装置の構成)
 図2は、実施の形態に係る情報処理装置1の構成例を示すブロック図である。
 情報処理装置1は、CPU等から構成され、各部を制御するとともに、各種のプログラムを実行する制御部10と、フラッシュメモリ等の記憶媒体から構成され情報を記憶する記憶部11と、ネットワーク3を介して外部と通信する通信部12とを備える。
 制御部10は、後述する情報処理プログラムとしての能力判定プログラム110を実行することで、発話制御手段100、表示制御手段101、動作付与手段102、音声認識手段103、映像認識手段104、能力判定手段105、ブレイクダウン検出手段106等として機能する。
 発話制御手段100は、端末2における音声発話を制御する。発話制御手段100は、主に予め用意した複数レベルの質問を含む質問情報111から利用者4のレベルに応じて質問を選択して発話する。なお、音声発話は、質問情報111に基づいた質問だけでなく、当該質問に対する利用者4に対する挨拶や利用者4の発話、回答に対する相槌等も含むものとする。
 表示制御手段101は、アバターの画像を定義するアバター情報112及びアバターの動作を定義する動作情報113を用いて端末2にアバターを動作を付与しつつ表示する。
 動作付与手段102は、表示制御手段101によって端末2に表示されるアバターに対して、発話制御手段100、音声認識手段103、映像認識手段104の動作に応じて動作情報113を参照して動作を付与する。例えば、発話制御手段100の動作に応じて発話に合わせてアバターに発話の動作を付与する。動作情報113の全ての動作は、例えば、ある発話文と紐付けられており、生成する動作の発話文と文章間距離が最も近いものが選択される。また、音声認識手段103の動作に応じてアバターに聴くしぐさ等の動作を付与し、映像認識手段104の動作に応じてアバターに利用者4の動作に反応する動作を付与する。なお、動作付与手段102は、能力判定手段105及びブレイクダウン検出手段106の動作に応じて動作を付与するものであってもよい。
 音声認識手段103は、端末2を介して受け付けた発話制御手段100の質問等に対する利用者4の回答に伴う音声を認識し、回答情報114として記憶部11に格納する。音声認識手段103は、認識した音声についてさらに言語理解の処理をするものであってもよいし、言語理解は能力判定手段105により行ってもよい。音声認識としては、例えば、GMM-HMM、DNN-HMM、End-to-End DNN等の手段を採用でき、言語理解としては、キーワード抽出、決定木、ニューラルネットワーク等の手段を採用できる。
 映像認識手段104は、端末2を介して受け付けた利用者4の回答に伴う仕草や目線、ジェスチャー等を含む映像を認識し、回答情報114として記憶部11に格納する。映像認識としては、例えば、ニューラルネットワーク等の手段を採用できる。
 能力判定手段105は、回答情報114の内容に基づいて利用者4の能力を判定して判定結果情報115として記憶部11に格納する。判定基準として、例えば、CEFR(Common European Framework of Reference for Languages)が用いられる。具体的なレベルとしてA1、A2、B1、B2、C1、C2が用意され、当該並順でレベルが上がる。能力判定としては、例えば、線形回帰、決定木、ニューラルネットワーク等の手段を用いて判定することができる。
 また、能力判定手段105は、ブレイクダウン検出手段106がブレイクダウンを検出した場合に、発話制御手段100に暫定的に判定したレベルをフィードバックし、発話制御手段100に当該レベルに応じた質問を選択させて発話させる。当該フィードバック動作については後に詳細に説明する。また、能力判定手段105は、複数の質問の回答が得られた後、利用者4のレベルを総合的に判定する。
 ブレイクダウン検出手段106は、音声認識手段103及び映像認識手段104が認識した利用者4の回答に不理解又は非流暢性、文法正確性の低下が検出された場合、これをブレイクダウンとして検出する。なお、本実施の形態は英会話の能力判定のケースを前提に上記のようにブレイクダウンの定義をしているが、判定内容が異なる場合や判定と異なる用途の場合は、当該判定内容や当該用途に合わせてブレイクダウンの定義を変更してもよい。より上位概念で定義するのであれば、情報処理装置1の発信する内容に対して利用者4の応答の内容が正常時の応答の分布から逸脱し,その結果対話的な破綻が生じることをブレイクダウンとする。
 ブレイクダウン検出手段106は、発話の繰り返しを要求している状態、又は混乱若しくは理解しようと考え込んで黙った状態等から不理解の検出を行う。後者については、ユーザの不理解時に発生する特徴的な動作として、例えば、視線を逸らす、顔を近づける、瞬きが多い、横をむく、視線を激しく動かす、頭部を激しく動かす、無言になる、発話の音量が小さくなる、の8つが音声認識手段103又は映像認識手段104により検出された場合に混乱若しくは理解しようと考え込んで黙った状態であると判断する。
 また、ブレイクダウン検出手段106は、語彙・文法・発音といった言語知識の想起がうまくいかず、発話産出が停滞する状態から非流暢性の検出を行う。特に、沈黙の位置が文や節の途中に生じる場合に発話産出が停滞する状態と判断する。これは文や節の始め又は終わりに沈黙が生じる場合は、背景知識の不足や、効果的な談話構造の計画失敗など、内容想起を起因とすると考えられるためである。
 記憶部11は、制御部10を上述した各手段100-106として動作させる能力判定プログラム110、質問情報111、アバター情報112、動作情報113、回答情報114、判定結果情報115等を記憶する。
(端末2の構成)
 図3は、実施の形態に係る端末2の構成例を示すブロック図である。
 端末2は、CPU等から構成され、各部を制御するとともに、各種のプログラムを実行する制御部20と、フラッシュメモリ等の記憶媒体から構成され情報を記憶する記憶部21と、ネットワーク3を介して外部と通信する通信部22と、入力される利用者4の音声を電気信号に変換するマイク23と、制御部20から入力される信号を音声に変換して出力するスピーカー24と、利用者4を撮像して映像信号を出力するカメラ25と、画像、映像、文字等を表示するLCD等のディスプレイ26を備える。その他、端末2は、利用者4から操作を受け付ける図示しない操作部(キーボード、マウス、トラックパッド、タッチパネル)等を備える。
 なお、情報処理装置1と端末2とをそれぞれ別装置として説明するが、端末2に情報処理装置1の手段の一部又は全部を設けてもよく、発明の趣旨を逸脱しない範囲で適宜設計を変更してよい。
(情報処理装置の動作)
 次に、本実施の形態の作用を、(1)基本動作、(2)導入動作、(3)レベルチェック動作、(4)突き上げ動作に分けて説明する。
(1)基本動作
 まず、利用者4は、例えば、英会話の能力の判定を要求すべく端末2を操作する。端末2は、情報処理装置1とネットワーク3を介して通信し、情報処理装置1に英会話の能力判定を要求する。
 情報処理装置1は、端末2から英会話の能力判定の要求を受け付けると、発話制御手段100、表示制御手段101、動作付与手段102に指示を出して、図4に示すように、端末2のディスプレイ26にアバターを表示処理する。
 図4は、端末2のディスプレイ26に表示される画面例を示す概略図である。
 画面101aは、利用者4を端末2のカメラ25で撮影した映像を表示する領域101aと、アバターを表示する領域101aとを有する。領域101aは、主に利用者4の参照用に表示されるが、表示しないものであってもよい。
 領域101aに表示されるアバターは、表示制御手段101、動作付与手段102により動作が付与されるとともに、発話制御手段100によりスピーカー24から音声が出力される。
 また、ディスプレイ26に表示される画面101a及びスピーカー24から出力される音声に対する利用者4の反応(声や表情や動作等)は、端末2のマイク23及びカメラ25を介して情報処理装置1に入力され、それぞれ情報処理装置1の音声認識手段103及び映像認識手段104によって認識される。
 以降の「(2)導入動作」、「(3)レベルチェック動作」、「(4)突き上げ動作」を説明する前に、情報処理装置1の各手段の基本動作を説明する。
 情報処理装置1の発話制御手段100は、端末2におけるアバターの音声発話を制御する。発話制御手段100は、主に予め用意した複数レベルの質問を含む質問情報111から利用者4のレベルに応じて質問を選択して発話する。
 また、情報処理装置1の表示制御手段101は、アバターの画像を定義するアバター情報112及びアバターの動作を定義する動作情報113を用いて端末2にアバターを表示するとともに、情報処理装置1の動作付与手段102は、表示制御手段101によって端末2に表示されるアバターに対して、発話制御手段100、音声認識手段103、映像認識手段104の動作に応じて動作情報113を参照して動作を付与する。音声認識手段103の動作に応じてアバターに聴くしぐさ等の動作(傾聴動作)を付与し、映像認識手段104の動作に応じてアバターに利用者4の動作に反応する動作(リアクション)を付与する。これらの動作により利用者4の自己開示を促す。
 また、情報処理装置1の音声認識手段103は、端末2を介して受け付けた発話制御手段100の質問等に対する利用者4の回答に伴う音声を認識し、回答情報114として記憶部11に格納する。音声認識手段103は、認識した音声についてさらに言語理解の処理をする。
 また、情報処理装置1の映像認識手段104は、端末2を介して受け付けた利用者4の回答に伴う仕草や目線、ジェスチャー等を含む映像を認識し、回答情報114として記憶部11に格納する。
 また、情報処理装置1の能力判定手段105は、回答情報114の内容に基づいて利用者4の能力を判定して判定結果情報115として記憶部11に格納する。
 また、情報処理装置1のブレイクダウン検出手段106は、音声認識手段103及び映像認識手段104が認識した利用者4の回答に不理解又は非流暢性が検出された場合、これをブレイクダウンとして検出する。
 以降、上記アバターを表示しつつ、利用者4の英会話の能力を判定するための情報処理装置1の具体的な動作を以下に説明する。
(2)導入動作
 図5は、情報処理装置1の発話制御手段100による質問(S1~S10)と、利用者4の回答(U1~U9)の内容例を示す概略図である。フェーズ1は導入動作時の質問及び回答、フェーズ2はレベルチェック動作時の質問及び回答、フェーズ3は突き上げ動作時の質問及び回答の内容を示す。また、図6は、情報処理装置1の能力判定のレベルと時間経過の関係を示すグラフ図である。また、図7は、情報処理装置1の動作例を示すフローチャートである。
 まず、情報処理装置1は、利用者4に対して挨拶やスモールトーク等の比較的簡単な会話を行い、緊張感を解すとともに大まかなレベルを把握する導入動作を行うべく、発話制御手段100、表示制御手段101、動作付与手段102によりアバターに動作をつけつつ、例えば、レベルA1の質問を発話させ(S100)、音声認識手段103、映像認識手段104により利用者4の回答及び表情や動作を認識しつつ、能力判定手段105により利用者4の能力を判定する(S101)(図6のフェーズ1)。
 具体的には発話制御手段100、表示制御手段101、動作付与手段102により、図5のフェーズ1に示すように、アバターにより、例えば、「What is your favorite season?」(S1)といった内容のトピック導入質問を行う。
 これに対して利用者4が、例えば、「My favorite season is winter.」(U1)と回答した場合、音声認識手段103、映像認識手段104により利用者4の回答及び表情や動作を認識する。同時に能力判定手段105により利用者4の能力を判定する。例えば、CEFR A2レベルと判定されたものとする。
 次に、利用者4が確実にA2レベルの会話を行えることを確認するために、発話制御手段100、表示制御手段101、動作付与手段102により「Are there any activities you like to do in winter?」(S2)といった内容の追加質問を行う。
 これに対して利用者4が、例えば、「Uh ... Ski and making snowman.」(U2)と回答した場合、さらに「That sounds like a lot of fun. Could you tell me more about it?」(S3)といった内容の継続依頼を行う。ここで利用者4は「I like skiing with family. I go every year.」(U3)と回答したものとする。
 なお、当該追加質問の回数は、能力判定手段105が判定結果を確認できれば0回でもよいし、1回でも複数回でもよい。
(3)レベルチェック動作
 次に、発話制御手段100、表示制御手段101、動作付与手段102により、図5のフェーズ1に示すように、アバターにより、例えば、「Alright. What did you eat for breakfast this morning?」(S4)といった内容で、ステップS101で判定された能力であるA2レベルに応じてトピック導入質問を行う(S102)。
 これに対して利用者4が、例えば、「I ate uh... Sandwich it is chicken and salad it is very delicious.」(U4)と回答した場合、音声認識手段103、映像認識手段104により利用者4の回答及び表情や動作を認識する。同時に能力判定手段105により利用者4の能力を判定する(S103)(図6のフェーズ2)。例えば、CEFR A2レベルと判定されたものとする。ステップS102及びS103は一度でもよいが、利用者4が確実にA2レベルの会話を行えることを確認するために、本実施の形態では追加質問を複数回行って能力を判定するものとする。なお、1の質問毎に能力判定を行ってもよいし、複数の質問毎、話題毎等の予め定めた単位毎に能力判定を行ってもよい。
 例えば、発話制御手段100、表示制御手段101、動作付与手段102により「Do you usually eat breakfast?」(S5)といった内容の追加質問を行う。
 これに対して利用者4が、例えば、「Uh yes I always eat breakfast.」(U5)と回答した場合、さらに「I see what time do you usually eat breakfast.」(S6)といった内容の追加質問を行う。ここで利用者4は「Uh seven A.M. I wake up and I go to kitchen and I eat breakfast.」(U6)と回答したものとする。
 上記会話において、音声認識手段103、映像認識手段104により利用者4の回答及び表情や動作を認識した結果、回答が問題なく行われているため、能力判定手段105は利用者4の能力をCEFR A2レベルと暫定的に判定する(S103)。厳密に記載すると、レベルが確定したわけではないため、CEFR A2レベル以上(つまり、下限のレベルがA2)と判定する。
(4)突き上げ動作
 次に、上記レベルチェック動作において判定した判定結果が正しいことを検証するため、発話制御手段100、表示制御手段101、動作付与手段102により、ステップS103の判定結果のレベルよりレベルを上げて、CEFR B1レベルのトピック導入質問を行う(S104)。具体的には、図5のフェーズ3に示すように、アバターにより、例えば、「Have you ever been to a foreign country?」(S7)といった内容の質問を行う。なお、レベルを1上げる場合に限らず2以上上げてもよいし、回答次第で上げるレベルの度合いを変更してもよく、さらには目的に応じてレベルを下げるものであってもよい。
 これに対して利用者4が、例えば、「Uh no. I never go to foreign country.」(U7)と回答した場合、音声認識手段103、映像認識手段104により利用者4の回答及び表情や動作を認識する。ブレイクダウン検出手段106は、認識した利用者4の回答に不理解又は非流暢性が検出されないため(S105;No)、本来であればさらにレベルを上げて質問する(S106)が、本実施の形態では利用者4が確実にB1レベルの会話を行えることを確認するために、B1レベルで複数回追加質問を行うものとする。
 次に、利用者4が確実にB1レベルの会話を行えることを確認するために、発話制御手段100、表示制御手段101、動作付与手段102により「Ok. which country would you like to visit in the future?」(S8)といった内容の追加質問を行う。
 これに対して利用者4が、例えば、「I would like visit ... Singapore.」(U8)と回答した場合、ブレイクダウン検出手段106は非流暢性を検出するが、これが正しいか確かめるためにさらに「Why is that?」(S8)といった内容の継続依頼を行う。
 これに対して利用者4が、例えば、「Because I want visit ... I like go to nice ... ah nice ...」(U9)といった回答をした場合、ブレイクダウン検出手段106は、認識した利用者4の回答に非流暢性が検出されたため、ブレイクダウンが検出されたと判断し(S105;Yes)、当該回答に対して「That’s ok. Let’s move on.」(S10)といった発話をする。
 能力判定手段105は、ブレイクダウン検出手段106がブレイクダウンを検出した場合に(S105;Yes)、利用者4の能力がCEFR レベルB1ではなくレベルA2であると判定する(S107)(図6のフェーズ3)。また、能力判定手段105は、判定結果を利用者4に対し、又は利用者4以外の任意の管理者等に対して出力する(S108)。
 一方、複数回のレベルB1での質問に対し、ブレイクダウン検出手段106がブレイクダウンを検出しなかった場合に(S105;No)、ステップS103における判定結果を訂正してレベルをB2に上げ(S106)、ステップS102~S106を繰り返して、最終的に能力を判定する(S107)。
 また、能力判定手段105は、上記ステップS107によって判定されたレベルを発話制御手段100に暫定的な判定結果としてフィードバックし、再度当該レベルに応じた質問を選択させて発話させてもよい(S102)。この場合、能力判定手段105は、上記ステップS102~S107を複数回繰り返してから、利用者4のレベルを総合的に判定してもよい。つまり、「(3)レベルチェック動作」、「(4)突き上げ動作」を複数繰り返してからレベルA2と判断してもよい(図6の2回目のフェーズ2、フェーズ3)。また、能力判定が完了した後にクールダウンのフェーズをさらに設けてもよい(図6のフェーズ3の後、Cool down)。
(実施の形態の効果)
 上記した実施の形態によれば、能力判定手段105により利用者4の能力を判定し、判定した能力を検証するためにレベルを上げた質問を行い、当該質問に対する回答においてブレイクダウン検出手段106がブレイクダウンを検出した場合、上げたレベルではなく予め判定した能力が正しそうであると判定し、ブレイクダウンを検出しなかった場合は、レベルをさらに上げて動作を継続することにより利用者4の能力を判定するようにしたため、能力判定中の利用者4(被評価者)のレベルを考慮した質問をすることができ、他のレベルを試してみることで正確でかつ確信をもってレベルを判定することができる。
 また、能力判定を行う前に挨拶やスモールトーク等の比較的簡単な会話を行う導入動作を採用したため、利用者4をウォームアップして以降の能力判定動作にスムーズに移行することができる。
 また、アバターにより発話の際、回答受付の際にジェスチャー動作(傾聴動作、リアクション)を付与するようにしたため、より自然な質問及び回答動作が可能となり、ひいては利用者の自己開示を促し、情報を引き出すことができる。
 また、会話、話題等を単位として暫定的にレベルを判定し、当該レベルに応じて質問を選択するようにしたため、暫定的にレベルを判定しない場合のように手当たり次第に全ての質問をしなくていいため、質問に要する時間の点で効率性が向上する。
[他の実施の形態]
 なお、本発明は、上記実施の形態に限定されず、本発明の趣旨を逸脱しない範囲で種々な変形が可能である。
 英会話以外の能力判定、例えば、就職面接、精神状態判定、プレゼン能力判定等の能力判定に用いてもよい。さらに能力判定以外の用途、例えば、レストランにおけるメニュー選定、観光案内における観光地選定等のガイド用途に応用することができる。この場合、具体的に能力判定手段105を嗜好判定手段に置き換え、利用者の嗜好を利用者の発言、しぐさ、表情等から判定し、レベルをメニューや観光地に置き換えることで達成される。
 上記実施の形態では制御部10の各手段100~106の機能をプログラムで実現したが、各手段の全て又は一部をASIC等のハードウエアによって実現してもよい。また、上記実施の形態で用いたプログラムをCD-ROM等の記録媒体に記憶して提供することもできる。また、上記実施の形態で説明した上記ステップの入れ替え、削除、追加等は本発明の要旨を変更しない範囲内で可能である。
 また、情報処理装置1の制御部10の各手段100~106の機能は、必ずしも情報処理装置1上で実現する必要はなく、本発明の要旨を変更しない範囲内で、端末2上で実現してもよい。また、同様に端末2の各機能は、必ずしも端末2上で実現する必要はなく、本発明の要旨を変更しない範囲内で、情報処理装置1上で実現してもよい。
産業上利用の可能性
評価者のレベルを考慮した質問をするとともに、評価者のレベル判定の正確性を向上する情報処理方法、情報処理プログラム及び情報処理装置を提供する。
1     :情報処理装置
2     :端末
3     :ネットワーク
4     :利用者
10    :制御部
11    :記憶部
12    :通信部
20    :制御部
21    :記憶部
22    :通信部
23    :マイク
24    :スピーカー
25    :カメラ
26    :ディスプレイ
100   :発話制御手段
101   :表示制御手段
102   :動作付与手段
103   :音声認識手段
104   :映像認識手段
105   :能力判定手段
106   :ブレイクダウン検出手段
110   :能力判定プログラム
111   :質問情報
112   :アバター情報
113   :動作情報
114   :回答情報
115   :判定結果情報

 

Claims (9)

  1.  コンピュータに、
     対話者に対して予め定めた複数のレベルのうち一のレベルの質問を発話する発話制御ステップと、
     前記回答において前記対話者のブレイクダウンを検出する検出ステップと、
     少なくとも前記ブレイクダウンを検出したレベルに基づいて前記対話者のレベルを判定する判定ステップとを実行させる情報処理方法。
  2.  前記判定ステップは、予め定めた前記回答の単位毎に前記対話者のレベルを判定し、当該レベルより上のレベルの質問を発話する請求項1に記載の情報処理方法。
  3.  ブレイクダウンを検出した場合、前記複数のレベルの質問のうち当該ブレイクダウンを検出したレベル以外のレベルの質問を発話する請求項1又は2に記載の情報処理方法。
  4.  前記判定ステップは、前記ブレイクダウンの検出に関わらず予め定めた前記回答の単位毎に前記対話者のレベルを判定し、
     前記発話制御ステップは、前記判定ステップにおいて判定したレベルの質問を発話する請求項1から3のいずれか1項に記載の情報処理方法。
  5.  前記発話制御ステップの発話に連動して動作するアバターを表示制御する表示制御ステップをさらに実行させる請求項1から4のいずれか1項に記載の情報処理方法。
  6.  前記アバターに対して前記回答を傾聴する動作を付与する動作付与ステップをさらに実行させる請求項5に記載の情報処理方法。
  7.  当該アバターに対して前記回答に応答する動作を付与する動作付与ステップをさらに実行させる請求項5に記載の情報処理方法。
  8.  コンピュータを、
     対話者に対して予め定めた複数のレベルのうち一のレベルの質問を発話制御する発話制御手段と、
     前記回答において前記対話者のブレイクダウンを検出する検出手段と、
     少なくとも前記ブレイクダウンを検出したレベルに基づいて前記対話者のレベルを判定する判定手段としてさらに機能させる情報処理プログラム。
  9.  対話者に対して予め定めた複数のレベルのうち一のレベルの質問を発話制御する発話制御手段と、
     前記回答において前記対話者のブレイクダウンを検出する検出手段と、
     少なくとも前記ブレイクダウンを検出したレベルに基づいて前記対話者のレベルを判定する判定手段とを有する情報処理装置。
     

     
PCT/JP2023/009268 2022-03-25 2023-03-10 情報処理方法、情報処理プログラム及び情報処理装置 Ceased WO2023181986A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US18/895,299 US20250014592A1 (en) 2022-03-25 2024-09-24 Information processing method, information processing program, and information processing device

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2022049257A JP7649551B2 (ja) 2022-03-25 2022-03-25 情報処理方法、情報処理プログラム及び情報処理装置
JP2022-049257 2022-03-25

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US18/895,299 Continuation US20250014592A1 (en) 2022-03-25 2024-09-24 Information processing method, information processing program, and information processing device

Publications (1)

Publication Number Publication Date
WO2023181986A1 true WO2023181986A1 (ja) 2023-09-28

Family

ID=88101296

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2023/009268 Ceased WO2023181986A1 (ja) 2022-03-25 2023-03-10 情報処理方法、情報処理プログラム及び情報処理装置

Country Status (3)

Country Link
US (1) US20250014592A1 (ja)
JP (1) JP7649551B2 (ja)
WO (1) WO2023181986A1 (ja)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2007193456A (ja) * 2006-01-17 2007-08-02 Omron Corp 要因推定装置、要因推定プログラム、要因推定プログラムを記録した記録媒体、および要因推定方法
WO2018230345A1 (ja) * 2017-06-15 2018-12-20 株式会社Caiメディア 対話ロボットおよび対話システム、並びに対話プログラム

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2007193456A (ja) * 2006-01-17 2007-08-02 Omron Corp 要因推定装置、要因推定プログラム、要因推定プログラムを記録した記録媒体、および要因推定方法
WO2018230345A1 (ja) * 2017-06-15 2018-12-20 株式会社Caiメディア 対話ロボットおよび対話システム、並びに対話プログラム

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
UEDA, TOMOHIRO ET AL.: "VR English Conversation System that Changes Conversation Progress with Gaze and Gestures", IPSJ SIG TECHNICAL REPORTS, INFORMATION PROCESSING SOCIETY OF JAPAN, vol. 2021-HCI-192, no. 6, 8 March 2021 (2021-03-08), pages 1 - 8, XP009549621 *

Also Published As

Publication number Publication date
US20250014592A1 (en) 2025-01-09
JP7649551B2 (ja) 2025-03-21
JP2023142373A (ja) 2023-10-05

Similar Documents

Publication Publication Date Title
JP4838351B2 (ja) キーワード抽出装置
JP4679254B2 (ja) 対話システム、対話方法、及びコンピュータプログラム
JP6025785B2 (ja) 自然言語理解のための自動音声認識プロキシシステム
US12190859B2 (en) Synthesized speech audio data generated on behalf of human participant in conversation
US9336782B1 (en) Distributed collection and processing of voice bank data
US20020123894A1 (en) Processing speech recognition errors in an embedded speech recognition system
US20080255850A1 (en) Providing Expressive User Interaction With A Multimodal Application
JP5105943B2 (ja) 発話評価装置及び発話評価プログラム
CN115700878A (zh) 通过动态响应打断内容改进会话ai的双工通信
US11314942B1 (en) Accelerating agent performance in a natural language processing system
Lubold et al. Effects of adapting to user pitch on rapport perception, behavior, and state with a social robotic learning companion: N. Lubold et al.
Rebman Jr et al. Speech recognition in the human–computer interface
JP2004037721A (ja) 音声応答システム、音声応答プログラム及びそのための記憶媒体
JP7443826B2 (ja) 対話支援システム、対話支援方法、対話支援プログラム
US12216963B1 (en) Computer system-based pausing and resuming of natural language conversations
KR101004913B1 (ko) 음성인식을 활용한 컴퓨터 주도형 상호대화의 말하기 능력평가 장치 및 그 평가방법
KR20210123545A (ko) 사용자 피드백 기반 대화 서비스 제공 방법 및 장치
WO2017175351A1 (ja) 情報処理装置
da Silva et al. How do illiterate people interact with an intelligent voice assistant?
US20250118298A1 (en) System and method for optimizing a user interaction session within an interactive voice response system
JP2017208003A (ja) 対話方法、対話システム、対話装置、およびプログラム
Ureta et al. At home with Alexa: a tale of two conversational agents
KR100898104B1 (ko) 상호 대화식 학습 시스템 및 방법
JP2010197644A (ja) 音声認識システム
JP2024092451A (ja) 対話支援システム、対話支援方法、およびコンピュータプログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23774585

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23774585

Country of ref document: EP

Kind code of ref document: A1