WO2024262249A1 - 情報処理方法および情報処理システム - Google Patents
情報処理方法および情報処理システム Download PDFInfo
- Publication number
- WO2024262249A1 WO2024262249A1 PCT/JP2024/019314 JP2024019314W WO2024262249A1 WO 2024262249 A1 WO2024262249 A1 WO 2024262249A1 JP 2024019314 W JP2024019314 W JP 2024019314W WO 2024262249 A1 WO2024262249 A1 WO 2024262249A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- information
- performance
- response
- response information
- user
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/0008—Associated control or indicating means
-
- G—PHYSICS
- G09—EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
- G09B—EDUCATIONAL OR DEMONSTRATION APPLIANCES; APPLIANCES FOR TEACHING, OR COMMUNICATING WITH, THE BLIND, DEAF OR MUTE; MODELS; PLANETARIA; GLOBES; MAPS; DIAGRAMS
- G09B15/00—Teaching music
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10G—REPRESENTATION OF MUSIC; RECORDING MUSIC IN NOTATION FORM; ACCESSORIES FOR MUSIC OR MUSICAL INSTRUMENTS NOT OTHERWISE PROVIDED FOR, e.g. SUPPORTS
- G10G1/00—Means for the representation of music
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10G—REPRESENTATION OF MUSIC; RECORDING MUSIC IN NOTATION FORM; ACCESSORIES FOR MUSIC OR MUSICAL INSTRUMENTS NOT OTHERWISE PROVIDED FOR, e.g. SUPPORTS
- G10G3/00—Recording music in notation form, e.g. recording the mechanical operation of a musical instrument
- G10G3/04—Recording music in notation form, e.g. recording the mechanical operation of a musical instrument using electrical means
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L67/00—Network arrangements or protocols for supporting network services or applications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2210/00—Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
- G10H2210/031—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal
- G10H2210/091—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal for performance evaluation, i.e. judging, grading or scoring the musical qualities or faithfulness of a performance, e.g. with respect to pitch, tempo or other timings of a reference performance
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2220/00—Input/output interfacing specifically adapted for electrophonic musical tools or instruments
- G10H2220/155—User input interfaces for electrophonic musical instruments
- G10H2220/265—Key design details; Special characteristics of individual keys of a keyboard; Key-like musical input devices, e.g. finger sensors, pedals, potentiometers, selectors
- G10H2220/311—Key design details; Special characteristics of individual keys of a keyboard; Key-like musical input devices, e.g. finger sensors, pedals, potentiometers, selectors with controlled tactile or haptic feedback effect; output interfaces therefor
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2250/00—Aspects of algorithms or signal processing methods without intrinsic musical character, yet specifically adapted for or used in electrophonic musical processing
- G10H2250/311—Neural networks for electrophonic musical instruments or musical processing, e.g. for musical recognition or control, automatic composition or improvisation
Definitions
- This disclosure relates to technology that notifies users of information related to playing musical instruments.
- Patent Document 1 discloses a configuration in which performance information is generated by analyzing video data generated by filming the user, and an evaluation score is generated by comparing the performance information with reference information.
- an information processing method acquires performance information related to a musical instrument being played by a user, acquires response information in natural language corresponding to the performance information, and executes a notification operation to notify the response information by a guide character displayed on a display device.
- An information processing method acquires performance information relating to a musical instrument being played by a user, and generates a prompt including the performance information, the prompt being used by a trained generative model to generate response information in natural language.
- An information processing system includes an information acquisition unit that acquires performance information related to a user's performance of a musical instrument, a response acquisition unit that acquires response information in natural language corresponding to the performance information, and an operation control unit that causes a guide character displayed on a display device to perform a notification operation to notify the user of the response information.
- a program causes a computer system to function as an information acquisition unit that acquires performance information related to a user's performance of a musical instrument, a response acquisition unit that acquires response information in natural language corresponding to the performance information, and an operation control unit that causes a guide character displayed on a display device to perform a notification operation to notify the user of the response information.
- FIG. 1 is a block diagram illustrating a configuration of an information system according to a first embodiment.
- FIG. 1 is a block diagram illustrating a configuration of an information processing system.
- FIG. 4 is a schematic diagram of a setting screen.
- FIG. 2 is a block diagram illustrating an example of a functional configuration of an information processing system.
- FIG. 1 is a schematic diagram of a basic character string.
- FIG. 2 is a schematic diagram of prompt and response information.
- FIG. FIG. 11 is a schematic diagram of response information.
- FIG. 13 is a schematic diagram of a guide character. 13 is a flowchart of an evaluation informing process.
- FIG. 13 is a schematic diagram of a prompt and response information according to the second embodiment.
- FIG. 11 is a block diagram illustrating a functional configuration of an information processing system according to a second embodiment.
- 13 is a flowchart of an evaluation notification process in the second embodiment.
- FIG. 13 is a schematic diagram of a prompt and response information according to the third embodiment.
- FIG. 13 is a schematic diagram of prompt and response information in a modified example.
- FIG. 13 is a schematic diagram of a prompt in a modified example.
- A: First Embodiment Fig. 1 is a block diagram illustrating the configuration of an information system 100 in the first embodiment.
- the information system 100 is a computer system that guides a user U in playing an electronic musical instrument 20.
- the information system 100 includes an information processing system 10, the electronic musical instrument 20, and a response generation system 30.
- the information processing system 10 is capable of communicating with the response generation system 30 via a communication network 200 such as the Internet.
- the response generation system 30 is a server system that generates response information R corresponding to a prompt P.
- the prompt P is an action instruction expressed in natural language.
- the response information R is text data that expresses a response to the prompt P in natural language.
- the response generation system 30 generates response information R using a trained generative model M.
- the generative model M is a generative probabilistic model that generates response information R in response to a prompt P.
- the generative model M has learned the tendency of response information R to prompts P through prior machine learning (pre-training).
- the generative model M is an interactive large-scale language model (LLM) that is trained specifically for natural language processing tasks such as response generation.
- LLM large-scale language model
- a natural language processing model realized by a transformer model that utilizes a self-attention mechanism is exemplified as the generative model M.
- the electronic musical instrument 20 is an input device that accepts musical pieces played by the user U.
- the electronic musical instrument 20 of the first embodiment is, for example, a keyboard instrument that conforms to the MIDI (Musical Instrument Digital Interface) standard, and has a plurality of keys 21 that correspond to different pitches.
- the electronic musical instrument 20 may also be mounted on the information processing system 10.
- the user U plays a piece of music by operating each key 21 in sequence.
- the electronic musical instrument 20 not only emits musical tones according to the performance by the user U, but also outputs a performance data string D representing the performance to the information processing system 10.
- the performance data string D is, for example, a time series of event data that conforms to the MIDI standard. Specifically, the performance data string D specifies the pitches corresponding to the keys 21 operated by the user U in time series.
- the information processing system 10 transmits a prompt P corresponding to the result of the evaluation of the performance of the electronic musical instrument 20 by the user U to the response generation system 30, and notifies the user U of the response information R received from the response generation system 30.
- the response information R in natural language corresponding to the evaluation of the performance by the user U is notified to the user U.
- FIG. 2 is a block diagram of an information processing system 10.
- the information processing system 10 is realized by an information device such as a smartphone, a tablet terminal, or a personal computer.
- the information processing system 10 comprises a control device 11, a storage device 12, a communication device 13, a sound emitting device 14, a display device 15, and an operation device 16.
- the information processing system 10 may be realized as a single device, or may be realized as multiple devices configured separately from each other.
- the control device 11 is composed of one or more processors that control each element of the information processing system 10.
- the control device 11 is composed of one or more types of processors, such as a CPU (Central Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).
- a CPU Central Processing Unit
- SPU Sound Processing Unit
- DSP Digital Signal Processor
- FPGA Field Programmable Gate Array
- ASIC Application Specific Integrated Circuit
- the storage device 12 is a single or multiple memories that store the programs executed by the control device 11 and various data used by the control device 11.
- the storage device 12 is configured with a known recording medium, such as a magnetic recording medium or a semiconductor recording medium.
- the storage device 12 may be configured with a combination of multiple types of recording media.
- a portable recording medium that is detachable from the information processing system 10, or a recording medium (e.g., cloud storage) that the control device 11 can write to or read from via the communication network 200 may be used as the storage device 12.
- the storage device 12 stores music data C that represents the music piece that the user U is to play.
- the music data C is data that represents the musical score of the music piece.
- the music data C specifies the pitch and duration of each of the multiple notes that make up the music piece.
- the music data C is data that conforms to the MIDI standard.
- the music data C may also specify information such as performance symbols that represent musical expression.
- the communication device 13 communicates with the response generation system 30 under the control of the control device 11. Specifically, the communication device 13 transmits a prompt P to the response generation system 30 and receives response information R transmitted from the response generation system 30.
- the sound emitting device 14 emits sound waves under the control of the control device 11.
- the sound emitting device 14 is, for example, a speaker or headphones.
- an audio signal V representing a voice corresponding to the response information R (hereinafter referred to as "guidance voice") is supplied to the sound emitting device 14.
- the sound emitting device 14 reproduces the guidance voice for the user U.
- a D/A converter that converts the audio signal V from digital to analog and an amplifier that amplifies the audio signal V are omitted from the illustration.
- the sound emitting device 14, which is separate from the information processing system 10, may be connected to the information processing system 10 by wire or wirelessly.
- the display device 15 displays images under the control of the control device 11.
- the display device 15 is configured with a display panel such as a liquid crystal panel or an organic EL (Electroluminescence) panel.
- the operation device 16 is an input device that accepts instructions from the user U.
- the operation device 16 is, for example, an operator operated by the user U, or a touch panel that detects contact by the user U.
- the display device 15 or operation device 16, which is separate from the information processing system 10 may be connected to the information processing system 10 by wire or wirelessly.
- FIG. 3 is a schematic diagram of a setting screen 151 for the user U to specify the response condition X (X1, X2, X3).
- the setting screen 151 is displayed on the display device 15.
- the response condition X includes attributes X1 of a virtual respondent (e.g., a guide character 152 described below) who executes the response represented by the response information R, attributes X2 of the user U, and tone X3 of the response according to the response information R.
- the setting screen 151 is an image in which multiple input fields F (F1, F2, F3) are arranged.
- the input field F1 is an input box in which the user U specifies the attribute X1 of the virtual respondent.
- the user U operates the operation device 16 to input a character string of the respondent's attribute X1 into the input field F1.
- the respondent's occupation piano teacher, etc.
- the user U may select one of multiple options prepared in advance as the attribute X1.
- the input field F2 is an input box in which the user U specifies his/her attribute X2.
- the user U inputs a character string of his/her own attribute X2 into the input field F2 by operating the operation device 16.
- the age group of the user U (elementary school student/junior high school student/high school student/university student/adult, etc.) is input as the attribute X2.
- the user U may select one of multiple options prepared in advance as the attribute X2.
- the input field F3 is an input box in which the user U specifies the tone X3 of the response based on the response information R.
- the user U inputs a character string of the desired tone X3 into the input field F3 by operating the operation device 16.
- the atmosphere (tone) of the remark such as "kindly,” “strictly,” or “frankly,” is specified as the tone X3.
- the user U may specify any one of multiple options prepared in advance as the tone X3.
- FIG. 4 is a block diagram illustrating the functional configuration of the information processing system 10.
- the control device 11 executes a program stored in the storage device 12 to realize multiple functions (information acquisition unit 41, response acquisition unit 42, operation control unit 43).
- the information acquisition unit 41 evaluates the performance of the electronic musical instrument 20 by the user U. Specifically, the information acquisition unit 41 compares the performance data string D supplied from the electronic musical instrument 20 with the music data C stored in the storage device 12, and identifies the differences between the two as performance errors by the user U. In other words, a performance error is a difference between the performance by the user U and the standard performance specified by the music data C.
- the information acquisition unit 41 generates evaluation information Y for each performance error that occurs in the performance by the user U.
- the evaluation information Y is information that represents an evaluation of the performance of the electronic musical instrument 20 by the user U. Specifically, the evaluation information Y includes the importance Y1 of the performance error, the position Y2 of the performance error, the type Y3 of the performance error, and detailed content Y4 of the performance error.
- the importance Y1 of a performance error is the degree of importance of the performance error.
- the importance Y1 is binary information that indicates the level of importance. Note that the importance Y1 may be information that indicates the level of importance in multiple stages.
- the position Y2 of the performance error is the point in the song where the performance error occurred.
- the point in the song where the performance data string D and the song data C differ is identified as the position Y2 of the performance error.
- it is specified by a bar number in the song or the elapsed time from the start of the song. Note that in addition to a specific point in time in the song, a specific section in the song may also be specified as the position Y2 of the performance error.
- the performance error type Y3 refers to a classification that distinguishes the performance error by type. For example, "pitch error,” “rhythm error,” and “hesitation in playing” are specified as the performance error type Y3.
- the aforementioned importance level Y1 is set, for example, according to the performance error type Y3. For example, the importance level Y1 for each performance error type Y3 is stored in advance in the storage device 12.
- the detailed content Y4 of the performance error is the specific content of the performance error made by the user U or the criticism of the performance error.
- examples of the detailed content Y4 of the performance error include performance states such as "playing a B when a C should be played,” “playing earlier than the correct timing,” and “having to play without stopping.”
- the detailed content Y4 of the performance error is expressed, for example, in natural language.
- the response acquisition unit 42 in FIG. 4 acquires response information R in natural language corresponding to evaluation information Y.
- the response acquisition unit 42 in the first embodiment includes an instruction generation unit 421 and an information receiving unit 422.
- the instruction generation unit 421 generates a prompt P for the response generation system 30. Specifically, the instruction generation unit 421 generates a prompt P including a response condition X and evaluation information Y.
- the instruction generation unit 421 uses the basic character string B in FIG. 5 to generate the prompt P.
- the basic character string B is a template that represents a typical character string for the prompt P, and is stored in advance in the storage device 12.
- the basic character string B includes multiple character strings b1 to b6. A blank is set in each of the multiple character strings b1 to b6. Each blank is a portion of the basic character string B into which variable information (response condition X, evaluation information Y) is inserted.
- the instruction generation unit 421 generates the prompt P by inserting the response condition X or evaluation information Y into each blank in the basic character string B.
- FIG. 6 shows an example of a prompt P generated by the instruction generation unit 421 using the basic character string B in FIG. 5.
- the character string b1 of the basic character string B is the portion that represents the condition related to the respondent.
- the instruction generation unit 421 generates an instruction related to the respondent in the prompt P by inserting the attribute X1 specified by the user U on the setting screen 151 into the blank space of the character string b1.
- the prompt P in the first embodiment includes the attribute X1 of a virtual respondent.
- the character string b2 of the basic character string B is a portion that represents the conditions for the user U to whom the response information R is to be notified.
- the instruction generation unit 421 generates an instruction for the user U in the prompt P by inserting the attribute X2 specified by the user U on the setting screen 151 into the blank space of the character string b2.
- the prompt P in the first embodiment includes the attribute X2 of the user U.
- the character string b3 of the basic character string B is a portion expressing a condition regarding the expression of the response based on the response information R.
- the instruction generating unit 421 generates an instruction regarding the expression of the response information R in the prompt P by inserting the tone X3 specified by the user U on the setting screen 151 into the blank space of the character string b3.
- the prompt P in the first embodiment includes the tone X3 related to the response information R.
- the character strings b4 to b6 of the basic character string B are the parts that express the conditions related to the response content according to the response information R.
- the instruction generation unit 421 generates an instruction related to the content of the response information R in the prompt P by inserting the evaluation information Y generated by the information acquisition unit 41 into each blank of the character strings b4 to b6.
- the character string b4 is the portion that notifies of important performance errors.
- the instruction generating unit 421 generates an instruction regarding an important performance error in the prompt P by inserting the position Y2, type Y3, and detailed content Y4 specified by each piece of evaluation information Y with a high importance Y1 into each blank in the character string b4.
- Character strings b5 and b6 are portions that notify of minor performance errors.
- the instruction generation unit 421 generates instructions for minor performance errors in the prompt P by inserting the position Y2, type Y3, and detailed content Y4 specified by each piece of evaluation information Y with a low importance Y1 into each blank space in character strings b5 and b6. Note that the number of performance errors specified in the prompt P is variable.
- the instruction generation unit 421 in FIG. 4 transmits the prompt P generated by the above procedure from the communication device 13 to the response generation system 30.
- the response generation system 30 generates response information R in natural language by processing the prompt P received from the information processing system 10 using the generation model M, and transmits the response information R to the information processing system 10.
- the information receiving unit 422 in FIG. 4 receives the response information R transmitted from the response generation system 30 via the communication device 13. In other words, the information receiving unit 422 obtains the response information R generated by the trained generation model M in response to the prompt P.
- FIG. 6 shows response information R generated from the prompt P described above.
- Response information R represents a natural language response that reflects response condition X and evaluation information Y contained in prompt P.
- response information R represents a natural language response sentence that should be spoken in tone X3 by a respondent with attribute X1 to user U with attribute X2.
- Response information R also represents a natural language response sentence that explains to user U the performance error represented by evaluation information Y.
- response information R distinguishes between important and minor performance errors (importance Y1), and points out the location Y2, type Y3, and detailed content Y4 of the performance error.
- FIG. 6 the attribute X2 of the user U is "adult"
- FIG. 7 illustrates an example of response information R when the attribute X2 is "elementary school student".
- response information R is generated in a polite tone, as in a conversation between adults.
- response information R is generated in a friendly and approachable tone, as in an adult instructing a child.
- the attribute X2 of the user U is included in the prompt P, so a variety of response information R can be generated according to the attribute X2 of the user U.
- the attribute X1 of a virtual respondent is included in the prompt P, so a variety of response information R can be generated according to the respondent's attribute X1.
- tone X3 is "gentle”
- FIG. 8 illustrates an example of response information R when tone X3 is "strict.”
- tone X3 is “gentle” (FIG. 6)
- response information R is generated in a gentle tone so as not to hurt the self-esteem of user U.
- tone X3 is "strict” (FIG. 8)
- response information R is generated in a strict tone.
- tone X3 related to response information R is included in prompt P, so that response information R in a variety of tones can be generated.
- the operation control unit 43 in FIG. 3 executes an operation (hereinafter referred to as the "notification operation") of notifying the user U of the response information R described above.
- the notification operation in the first embodiment is an operation of notifying the user U of the response information R by the guide character 152 (FIG. 9) displayed on the display device 15.
- the operation control unit 43 in the first embodiment includes a voice synthesis unit 431 and a display control unit 432.
- the voice synthesis unit 431 generates a voice signal V of a guidance voice corresponding to the response information R.
- the guidance voice is a voice that reads out the response sentence represented by the response information R.
- the voice synthesis unit 431 generates the voice signal V by executing a voice synthesis process on the response information R.
- Examples of the voice synthesis process include a segment connection type voice synthesis process that connects multiple speech segments, or a statistical model type voice synthesis process that uses a statistical model such as a deep neural network or HMM (Hidden Markov Model).
- the voice synthesis unit 431 supplies the voice signal V to the sound emission device 14. Therefore, the guidance voice represented by the voice signal V is reproduced from the sound emission device 14.
- the display control unit 432 causes the display device 15 to display the guide character 152 of FIG. 9.
- the guide character 152 is an object (agent) placed in the virtual space.
- the guide character 152 is a virtual instructor who instructs the user U on how to play the electronic musical instrument 20 in the virtual space.
- the display control unit 432 causes the guidance character 152 to perform a speaking action in parallel with the reproduction of the guidance voice by the sound emission device 14.
- the speaking action is an action that changes the shape of the mouth of the guidance character 152 in response to the audio signal V (so-called lip-syncing).
- the notification action of the first embodiment includes the speaking action of the guidance character 152 and the reproduction of the guidance voice.
- FIG. 10 is a flowchart of a process (hereinafter referred to as "rating notification process") executed by the control device 11.
- the rating notification process is started in response to an instruction from the user U via the operation device 16.
- the control device 11 When the evaluation notification process is started, the control device 11 (response acquisition unit 42) displays the setting screen 151 of FIG. 3 on the display device 15 (Sa1). The control device 11 (response acquisition unit 42) accepts input of response conditions X (X1, X2, X3) for the setting screen 151 from the user U (Sa2). The response conditions X accepted from the user U are stored in the storage device 12.
- the control device 11 After inputting the response condition X, the user U starts playing the electronic musical instrument 20.
- the control device 11 (information acquisition unit 41) evaluates the performance of the electronic musical instrument 20 by the user U (Sa3). Specifically, the control device 11 generates evaluation information Y for each performance error that occurs during the performance process by the user U.
- the control device 11 (instruction generation unit 421) generates a prompt P including a response condition X and evaluation information Y (Sa4). Specifically, the control device 11 generates the prompt P by inserting the response conditions X (X1, X2, X3) and the evaluation information Y (Y2, Y3, Y4) into each blank of the basic character string B stored in the storage device 12.
- the control device 11 transmits a prompt P to the response generation system 30 via the communication device 13 (Sa5). That is, the control device 11 requests the response generation system 30 to generate response information R.
- the control device 11 (information receiving unit 422) receives the response information R generated and transmitted by the response generation system 30 via the communication device 13 (Sa6).
- the control device 11 (voice synthesis unit 431) generates a voice signal V representing a guidance voice corresponding to the response information R (Sa7), and outputs the voice signal V to the sound emission device 14 (Sa8).
- the control device 11 (display control unit 432) causes the guidance character 152 displayed on the display device 15 to execute a speaking action (Sa9). That is, a notification action is executed that includes the reproduction of the guidance voice (Sa7, Sa8) and the speaking action (Sa9) by the guidance character 152.
- response information R in natural language corresponding to evaluation information Y regarding the performance of the electronic musical instrument 20 is acquired, and a notification action is performed by the guide character 152 to notify the response information R. Therefore, compared to a format in which a numerical value (e.g., an evaluation score) that evaluates the performance of the electronic musical instrument 20 by the user U is displayed, the user U can easily and appropriately understand the evaluation regarding the performance of the electronic musical instrument 20. In addition, a unique customer experience can be provided to the user U in which he or she enjoys the sensation of being instructed on the performance of the electronic musical instrument 20 by the guide character 152.
- a numerical value e.g., an evaluation score
- a prompt P is generated that includes evaluation information Y regarding the performance of the electronic musical instrument 20 by the user U, and response information R in natural language that is generated by the trained generative model M in response to the prompt P is obtained. Therefore, response information R that is statistically valid and linguistically natural in response to the evaluation (evaluation information Y) of the performance of the electronic musical instrument 20 by the user U can be presented to the user U.
- various information regarding performance errors (importance Y1, position Y2, type Y3, and detailed content Y4) is included in the evaluation information Y, so that response information R can be generated that includes a variety of information regarding the performance errors by the user U.
- Second embodiment A second embodiment will be described. Note that, for elements in the following exemplary aspects that have the same functions as those in the first embodiment, the same reference numerals as those in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.
- FIG. 11 is a schematic diagram of the prompt P and response information R in the second embodiment.
- the prompt P in the second embodiment includes identification information G for each performance error and an instruction Z1 to include the identification information G for each performance error in the response information R.
- the identification information G is a code string (tag) for identifying each performance error represented by the evaluation information Y.
- the identification information G includes a code g1 representing the performance error, a code g2 corresponding to the importance Y1, and a number g3 assigned to the performance error.
- the code g2 is set to "i" when the importance Y1 is high, and is set to "n" when the importance Y1 is low.
- the instruction Z1 is a natural language that indicates that the identification information G is to be added to the part of the response information R that points out each performance error.
- FIG. 11 shows an example of response information R generated from the prompt P described above.
- the prompt P including instruction Z1 is processed by the generation model M, and as a result, response information R including identification information G is generated.
- response information R is generated in accordance with instruction Z1.
- the identification information G of each performance error is set immediately after the portion of the response information R that points out the performance error.
- FIG. 12 is a block diagram illustrating the functional configuration of the information processing system 10 in the second embodiment.
- the control device 11 in the second embodiment executes a program stored in the storage device 12, and functions as a judgment processing unit 44 in addition to the same functions as in the first embodiment (information acquisition unit 41, response acquisition unit 42, operation control unit 43).
- the judgment processing unit 44 judges whether the response information R is appropriate.
- the operation of each unit other than the judgment processing unit 44 is the same as in the first embodiment.
- FIG. 13 is a flowchart of the evaluation notification process in the second embodiment.
- the process from the start of the evaluation notification process to the acquisition of response information R (Sb6) is the same as in the first embodiment.
- the control device 11 judgment processing unit 44
- the judgment process Sb in the second embodiment includes a first judgment process Sb1 and a second judgment process Sb2. Note that the order of the first judgment process Sb1 and the second judgment process Sb2 may be reversed.
- the first judgment process Sb1 is a judgment as to whether or not the response information R is appropriate for the prompt P. Specifically, the judgment processing unit 44 judges whether or not all of the performance errors specified in the prompt P are mentioned in the response information R.
- the judgment processing unit 44 compares the prompt P sent to the response generation system 30 with the response information R generated from the prompt P, and judges whether all of the identification information G of the performance error contained in the prompt P is also included in the response information R. If all of the identification information G is included in the response information R, the result of the first judgment process Sb1 is positive. On the other hand, if some of the identification information G included in the prompt P is not included in the response information R, the result of the first judgment process Sb1 is negative.
- the second judgment process Sb2 is a process for judging whether or not the response information R contains a prohibited word.
- a prohibited word is a word that is inappropriate from an educational or social point of view.
- a plurality of prohibited words are stored in advance in the storage device 12. The prohibited words may be set in response to an operation by the user U on the operation device 16.
- the judgment processing unit 44 judges whether or not any of the multiple prohibited words stored in the storage device 12 is included in the response information R. If the response information R does not contain any prohibited words, the result of the second judgment process Sb2 is positive (the response information R is appropriate). On the other hand, if the response information R contains prohibited words, the result of the second judgment process Sb2 is negative (the response information R is inappropriate).
- the control device 11 executes a notification operation including the reproduction of a guidance voice (Sa7, Sa8) and the speech action (Sa9) by the guidance character 152, as in the first embodiment. That is, the control device 11 executes the notification operation when it determines that the response information R is appropriate.
- the identification information G of the response information R is not included in the guidance voice. That is, the identification information G is excluded from the target of voice synthesis by the voice synthesis unit 431.
- the control device 11 instruction generation unit 421) retransmits the prompt P generated in the immediately preceding step Sa4 to the response generation system 30 from the communication device 13 (Sa5).
- the response generation system 30 generates response information R by processing the prompt P received from the information processing system 10 using the generation model M. Even if the prompt P is the same, the response information R generated by the generation model M changes for each generation. In other words, response information R is generated that is separate from the response information R that the control device 11 has determined to be inappropriate.
- the control device 11 receives the response information R generated and transmitted by the response generation system 30 via the communication device 13 (Sa6).
- the transmission of the prompt P (Sb5) and the reception of the response information R (Sb6) are repeated until the response information R is determined to be appropriate in the determination process Sb.
- the notification action is executed when the response information R is determined to be appropriate. Therefore, compared to a form in which the notification action is executed unconditionally, the possibility that inappropriate response information R is notified to the user U can be reduced.
- response information R including identification information G of performance errors is generated by the generative model M. Therefore, it is possible to easily check whether all of the performance errors specified in the prompt P are also included in the response information R by using the identification information G.
- FIG. 14 is a schematic diagram of a prompt P and response information R in the third embodiment.
- the prompt P in the third embodiment includes an instruction Z2 for including action information Q in the response information R.
- the action information Q is information (tag) for identifying an action to be performed in the process of notifying the response information R.
- the instruction Z2 includes a phrase in a natural language representing an action to be performed in the process of notifying the response information R, and action information Q representing the action.
- the phrase representing the action is, for example, a phrase such as "when staring at a student,""when staring at a piano,””whensmiling,” or "when clapping.”
- the action information Q is identification information for identifying each action.
- FIG. 14 illustrates an example of response information R generated from the prompt P described above.
- the prompt P including the instruction Z2 is processed by the generation model M, and as a result, response information R including action information Q is generated.
- response information R in accordance with the instruction Z2 is generated.
- the motion information Q of each motion is set in the portion of the response information R where the motion should be executed.
- motion information Q ⁇ gaze-student> for the motion of gazing at the user U is set at the beginning of the response information R
- the response information R in the third embodiment specifies the motion to be executed in the process of notifying the response information R.
- step Sa9 of the evaluation notification process the control device 11 (display control unit 432) causes the guide character 152 to perform the action specified in the response information R. Specifically, at the point in time when the guidance voice for the portion of the response information R near the action information Q is played, the control device 11 causes the guide character 152 to perform the action represented by the action information Q.
- the third embodiment also achieves the same effects as the first embodiment. Furthermore, in the third embodiment, the guide character 152 performs various actions during the notification operation. Therefore, it is possible to diversify the actions of the guide character 152. Note that in the above explanation, the third embodiment has been explained based on the first embodiment, but the configuration of the second embodiment, which determines whether the response information R is appropriate, is also applicable to the third embodiment.
- the attribute X1 of the respondent, the attribute X2 of the user U, and the tone of the response X3 are exemplified as the response condition X, but the response condition X is not limited to the above examples.
- the language of the response may be specified as the response condition X.
- the response information R is expressed in the language specified as the response condition X.
- the catchphrase X4 of the response in the response information R may be specified as the response condition X.
- FIG. 15 shows a prompt P that specifies the catchphrase X4 "meow” as the response condition X, and response information R generated in response to the prompt P.
- Response information R is generated in which the catchphrase X4 specified as the response condition X is added to the end of each sentence.
- the tone of the response (tone of speech) may be interpreted as including the tone of speech X3 and the catchphrase X4.
- an overall rating X5 resulting from evaluation of the performance of the electronic musical instrument 20 by the user U may be included in the response condition X.
- the response sentence represented by the response information R changes depending on the overall rating X5.
- evaluation information Y is generated for each performance mistake made by user U, but evaluation information Y is not limited to information representing a performance mistake.
- evaluation information Y may represent good points (high evaluation points) of the performance made by user U.
- response information R is generated that praises the performance of user U.
- each of importance Y1, position Y2, type Y3, and detailed content Y4 may be omitted.
- the user U selects the response condition X (X1, X2, X3), but the method of setting the response condition X is not limited to the above-mentioned examples.
- the response condition X for each guide character 152 may be stored in the storage device 12 in advance.
- the user U selects one of the multiple guide characters 152 by operating the operation device 16, for example.
- the response acquisition unit 42 (instruction generation unit 421) generates a prompt P including the response condition X corresponding to the guide character 152 selected by the user U from the multiple response conditions X stored in the storage device 12. According to the above-mentioned embodiment, the burden on the user U in setting the response condition X can be reduced.
- the control device 11 may edit the already-sent prompt P according to a predetermined rule and transmit the edited prompt P.
- the instruction generation unit 421 when the response information R is determined to be inappropriate because it contains a prohibited word, the instruction generation unit 421 generates a prompt P to which the condition that the prohibited word is not contained is added.
- the control device 11 may display a message to the effect that the response information R is inappropriate on the display device 15.
- the method of the judgment process Sb for judging the appropriateness of the response information R is not limited to the above-mentioned example.
- the first judgment process Sb1 using the identification information G may be omitted. That is, in the form in which the judgment processing unit 44 judges the appropriateness of the response information R, the prompt P and the identification information G of the response information R are not required.
- the importance Y1 is set according to the type Y3 of the performance error, but the method of setting the importance Y1 is not limited to the above examples.
- the information acquisition unit 41 may set the importance Y1 according to the history of the user U's past performances. For example, there is a tendency that performance errors that frequently occur in the user U's performances have a high priority for improvement and are of high importance. Taking the above tendency into consideration, the information acquisition unit 41 sets a large value to the importance Y1 of a performance error that occurred frequently in the user U's past performances.
- a prompt P including a response condition X and evaluation information Y is exemplified, but the contents of the prompt P are not limited to the above examples.
- a form in which one of the response condition X and the evaluation information Y is omitted, or a form in which information other than the response condition X and the evaluation information Y is included in the prompt P is also envisioned.
- a process such as obtaining response information R (Sa6) is executed after the user U finishes playing a piece of music.
- the evaluation of the performance (Sa3), the obtaining of response information R (Sa4 to Sa6), and the notification operation (Sa7 to Sa9) may be executed in parallel with the user U playing a piece of music.
- a unit period is, for example, a structural period into which a song is divided according to musical meaning.
- a structural period is, for example, each period such as an intro, verse, bridge, chorus, and outro.
- an evaluation of the performance may be performed in parallel with the performance of a piece of music by the user U, and response information R may be obtained (Sa4 to Sa6) and a notification operation (Sa7 to Sa9) may be performed each time a performance error occurs (i.e., each time evaluation information Y is generated).
- the guidance character 152 notifies the user U of the response information R, but the display of the guidance character 152 may be omitted.
- the display control unit 432 may be omitted from the operation control unit 43.
- a guidance voice represented by the response information R is played back, but the method of notifying the user U of the response information R (notification action) is not limited to the above examples.
- the notification action may be an action of displaying the response sentence represented by the response information R on the display device 15, or an action of printing the response sentence represented by the response information R using a printing device.
- An action of transmitting the response information R to a terminal device owned by the user U is also an example of a notification action.
- the operation control unit 43 may also display the musical score represented by the music data C on the display device 15, and highlight the portion of the musical score that corresponds to the performance error represented by the response information R.
- the operation control unit 43 may contrast the notes played by the user U due to the performance error with the correct notes represented by the music data C on the display device 15.
- the operation control unit 43 may also cause the guide character 152 to perform an operation to indicate the portion of the musical score that corresponds to the performance error.
- the operation control unit 43 may also play back, by the sound output device 14, the performance of the portion of the music represented by the music data C that corresponds to the performance error represented by the response information R. As shown in the above examples, any operation that notifies the user U of the response information R is included in the "notification operation".
- the information acquisition unit 41 evaluates the performance of the electronic musical instrument 20 by the user U, but the information acquisition unit 41 may receive the evaluation information Y generated by an external device via the communication device 13.
- the information acquisition unit 41 receives the evaluation information Y transmitted from the electronic musical instrument 20 via the communication device 13.
- the information acquisition unit 41 is comprehensively expressed as an element that acquires the evaluation information Y regarding the performance of the electronic musical instrument 20 by the user U.
- the "acquisition" of the evaluation information Y includes not only the "generation” of the evaluation information Y (i.e., the evaluation of the performance) but also the "reception" of the evaluation information Y.
- the response information R is generated by a response generation system 30 separate from the information processing system 10.
- the information processing system 10 may generate the response information R.
- the information processing system 10 may be equipped with a generation model M.
- the response acquisition unit 42 generates the response information R by processing the prompt P generated by the instruction generation unit 421 using the generation model M.
- the response acquisition unit 42 is comprehensively expressed as an element that acquires the response information R.
- the "acquisition" of the response information R includes not only the "reception” of the response information R, but also the "generation" of the response information R.
- evaluation information Y that represents an evaluation of the performance by the user U is exemplified, but the information (performance information) acquired by the information acquisition unit 41 is not limited to evaluation information Y.
- text data regarding the status of the performance by the user U is also exemplified as performance information.
- An example of text data as performance information is, for example, the type of instrument played by the user U or the name of the note played by the user U. In other words, an evaluation of the performance (evaluation information Y) is not required to generate performance information.
- video data generated by an imaging device by imaging a performance by the user U, or audio data generated by a sound collection device by collecting musical tones produced by the performance by the user U, are also exemplified as performance information.
- the information acquisition unit 41 may acquire a combination of two or more of the above types of information as performance information.
- performance information is comprehensively expressed as information related to the performance by user U, and evaluation information Y, text data, video data, and audio data are examples of "performance information.”
- a keyboard instrument is given as an example of the electronic musical instrument 20, but the instruments that are the subject of performance evaluation are not limited to keyboard instruments.
- each of the above embodiments can be applied in the same way to any type of musical instrument, such as a string instrument, a wind instrument, or a percussion instrument.
- an electronic musical instrument 20 capable of generating a performance data string D has been exemplified, but the musical instrument that is the subject of the performance evaluation is not limited to the electronic musical instrument 20.
- the information acquisition unit 41 generates evaluation information Y by analyzing an acoustic signal generated by picking up the sound emitted from the natural musical instrument. Any known technology may be adopted for the performance evaluation.
- the subject of evaluation is not limited to playing an instrument.
- the above-mentioned embodiments can be similarly applied to a configuration for evaluating the singing of a user U.
- the information acquisition unit 41 generates evaluation information Y by analyzing an acoustic signal generated by picking up the singing voice. Any known technology can be adopted for singing evaluation.
- the functions of the information processing system 10 exemplified above are realized by the cooperation of one or more processors constituting the control device 11 and the program stored in the storage device 12.
- the program according to the present disclosure can be provided in a form stored in a computer-readable recording medium and installed in a computer.
- the recording medium is, for example, a non-transitory recording medium, and a good example is an optical recording medium (optical disk) such as a CD-ROM, but also includes any known type of recording medium such as a semiconductor recording medium or a magnetic recording medium.
- a non-transitory recording medium includes any recording medium except a transient, propagating signal, and does not exclude volatile recording media.
- the storage medium that stores the program in the distribution device corresponds to the non-transitory recording medium described above.
- An information processing method acquires performance information relating to a user's performance of an instrument, acquires response information in natural language corresponding to the performance information, and executes a notification action of announcing the response information by a guide character displayed on a display device.
- response information in natural language corresponding to performance information relating to an instrument performance is acquired, and a notification action of announcing the response information by a guide character is executed. Therefore, compared to a form in which, for example, a numerical value (e.g., an evaluation score) relating to the user's performance of an instrument is displayed, the user can easily and appropriately understand information relating to the instrument performance. In addition, the user can be given the sensation of being guided (e.g., instructed) on the performance by the guide character.
- Performance information is any information related to a user's performance of an instrument.
- performance information is evaluation information that shows the results of evaluating a user's performance of an instrument.
- Evaluation information is, for example, information that shows a performance mistake made by the user.
- Information that shows a performance mistake is, for example, information such as the importance of the performance mistake, the position of the performance mistake in the music, the type of performance mistake, and the specific content of the performance mistake.
- video data generated by filming a performance by the user, or audio data generated by capturing the sound of the user playing an instrument are also examples of "performance information”.
- "Acquisition" of performance information encompasses both the operation of receiving performance information generated by an external device, and the operation of generating performance information by oneself.
- the "guide character” is a virtual object (agent) displayed on the display device.
- Specific examples of guide characters are virtual living beings such as humans or animals, but inanimate objects such as robots can also be included in the "guide character”.
- Control of the guide character's actions is a control process that causes the guide character to perform an action in response to response information. For example, a process that causes the guide character to perform the action of speaking the response information is exemplified.
- An “announcement action” is an output action for notifying a user of response information.
- Examples of “announcement actions” include a process for displaying the response information on a display device, a process for playing back the sound represented by the response information, a process for moving a virtual guide character in response to the response information, a process for displaying the musical score for a section of a song indicated by the response information, a process for playing back a performance of that section, etc.
- the acquisition of the response information involves generating a prompt including the performance information, and acquiring the response information generated by a trained generative model in response to the prompt.
- a prompt including performance information related to the user's performance of an instrument is generated, and natural language response information generated by a trained generative model in response to the prompt is acquired. Therefore, it is possible to notify the user of response information that is statistically valid and linguistically natural in response to information (e.g., an evaluation) related to the user's performance of an instrument.
- a “prompt” is input information that serves as the basis for a trained generative model to generate response information.
- a “prompt” can also be expressed as an instruction for a trained generative model to generate response information, or an instruction regarding the response information that the generative model should generate.
- a “prompt” includes both a single prompt and a collection of multiple prompts.
- Natural language response information is information that expresses a response to a prompt in natural language. Specifically, a natural language response sentence that corresponds to the performance information is generated as "response information.” For example, a response sentence that instructs or guides the user based on the performance information included in the prompt is an example of response information. "Acquisition" of response information encompasses both the operation of receiving performance information generated by an external device that includes a generative model, and the operation of generating performance information by oneself using the generative model.
- the performance information includes evaluation information that represents an evaluation of the performance.
- the performance information includes at least one of the following: the type of performance error that occurred in the performance, the importance of the performance error, the position of the performance error in the music piece, and the content of the performance error. According to the above aspects, various information related to the performance error is included in the performance information, so response information including a variety of information related to the performance error can be generated.
- the prompt includes an instruction to include in the response information identification information of a performance error that occurred in the performance.
- the response information including the identification information of the performance error is generated by the generative model. Therefore, it is possible to easily check whether all of the performance errors specified in the prompt are also included in the response information using the identification information.
- the prompt includes attributes of the user.
- the user attributes are included in the prompt, a variety of response information can be generated according to the user attributes.
- User attributes include, for example, age, sex, generation, occupation, personality (kind personality, assertive personality, etc.), emotions (angry, sad, etc.), etc. Also, the level of skill in playing an instrument (beginner, intermediate, advanced) is included in "user attributes".
- the prompt includes attributes of a virtual respondent who responds with the response information.
- the attributes of the virtual respondent are included in the prompt, it is possible to generate a variety of response information according to the attributes.
- the "respondent's attributes” include, for example, age, sex, generation, occupation, personality (kind personality, assertive personality, etc.), emotions (angry, sad, etc.), etc.
- the level of proficiency in musical instruction is also included in the "respondent's attribute information.”
- the prompt includes a tone of voice related to the response information.
- the tone of voice related to the response information is included in the prompt, response information with a variety of tones can be generated.
- Tone refers to the tone of the words expressed by the response information.
- the mood (tone) or emotion of the words is included in “tone”.
- examples of “atmosphere” include a gentle tone, a stern tone, a stiff tone, a formal tone, and a frank tone.
- Examples of "emotion” include an angry tone, a sad tone, and a happy tone.
- tone also includes “catchphrases” and "dialects”.
- "Catchphrases” are words that are frequently spoken. For example, adding a particular word (such as “It's meow” or “It's woof") to the end of a sentence is an example of a "catchphrase”.
- "Dialects" are regional differences regarding a particular language.
- the response information specifies an action to be performed in the process of notifying the response information, and in the notifying action, the guidance character is caused to perform the action specified in the response information.
- the guidance character performs various actions in the notifying action. Therefore, it is possible to diversify the actions of the guidance character.
- Actions to be performed in the process of notifying response information are, for example, auxiliary actions of the guide character that are not directly related to the performance information.
- actions to be performed in the process of notifying response information include various actions that a real instructor may perform in the process of instructing a performance, such as gazing at the user, gazing at an instrument in the virtual space, or smiling.
- aspects 1 to 9 the appropriateness of the response information is determined, and if it is determined that the response information is appropriate, the notification action is executed. In the above aspects, if it is determined that the response information is appropriate, the notification action is executed. Therefore, compared to a form in which the notification action is executed unconditionally, it is possible to reduce the possibility that inappropriate response information will be notified to the user.
- Determining the appropriateness of response information includes determining whether the response information is appropriate for the prompt, as well as determining whether the response information is appropriate from an educational or social perspective.
- the former determination is, for example, determining whether all of the performance information included in the prompt is reflected in the response information.
- the latter determination is, for example, determining whether the response information contains words that are educationally or socially inappropriate.
- An information processing method acquires performance information on a user playing an instrument, and generates a prompt including the performance information, where the trained generative model is a prompt for generating response information in natural language.
- a prompt including performance information on a user playing an instrument is generated. Therefore, response information that is statistically valid and linguistically natural can be generated in response to the performance information on the user playing an instrument.
- a unique customer experience can be provided to the user, in which response information that is statistically valid and linguistically natural in response to the performance information is acquired.
- An information processing system includes an information acquisition unit that acquires performance information related to a musical instrument being played by a user, a response acquisition unit that acquires response information in a natural language corresponding to the performance information, and an operation control unit that causes a guide character displayed on a display device to perform a notification operation to notify the user of the response information.
- a program according to one aspect (aspect 13) of the present disclosure causes a computer system to function as an information acquisition unit that acquires performance information related to a user's performance of a musical instrument, a response acquisition unit that acquires response information in natural language corresponding to the performance information, and an operation control unit that causes a guide character displayed on a display device to perform a notification operation to notify the user of the response information.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Multimedia (AREA)
- Acoustics & Sound (AREA)
- Educational Technology (AREA)
- Educational Administration (AREA)
- Business, Economics & Management (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Electrophonic Musical Instruments (AREA)
- Electrically Operated Instructional Devices (AREA)
- Auxiliary Devices For Music (AREA)
Abstract
情報処理システム100は、利用者Uによる電子楽器20の演奏に関する演奏情報を取得する情報取得部41と、演奏情報に応じた自然言語の応答情報Rを取得する応答取得部42と、表示装置15に表示された案内キャラクタにより応答情報Rを報知する報知動作を実行する動作制御部43とを具備する。
Description
本開示は、楽器の演奏に関する情報を利用者に報知する技術に関する。
楽器の演奏に関する情報を利用者に報知する技術が従来から提案されている。例えば特許文献1には、利用者の撮影により生成されたビデオデータの解析により演奏情報を生成し、基準情報と演奏情報との比較により評価スコアを生成する構成が開示されている。
しかし、特許文献1の技術のように評価スコアが表示されるだけでは、利用者は、自身の演奏に関する情報を容易かつ適切に把握できないという問題がある。以上の事情を考慮して、本開示のひとつの態様は、楽器の演奏に関する情報を利用者が容易かつ適切に理解できるようにすることを目的とする。
以上の課題を解決するために、本開示のひとつの態様に係る情報処理方法は、利用者による楽器の演奏に関する演奏情報を取得し、前記演奏情報に応じた自然言語の応答情報を取得し、表示装置に表示された案内キャラクタにより前記応答情報を報知する報知動作を実行する。
本開示の他の態様に係る情報処理方法は、利用者による楽器の演奏に関する演奏情報を取得し、訓練済の生成モデルが自然言語の応答情報を生成するためのプロンプトであって、前記演奏情報を含むプロンプトを生成する。
本開示のひとつの態様に係る情報処理システムは、利用者による楽器の演奏に関する演奏情報を取得する情報取得部と、前記演奏情報に応じた自然言語の応答情報を取得する応答取得部と、表示装置に表示された案内キャラクタに前記応答情報を報知する報知動作を実行させる動作制御部とを具備する。
本開示のひとつの態様に係るプログラムは、利用者による楽器の演奏に関する演奏情報を取得する情報取得部、前記演奏情報に応じた自然言語の応答情報を取得する応答取得部、および、表示装置に表示された案内キャラクタに前記応答情報を報知する報知動作を実行させる動作制御部、としてコンピュータシステムを機能させる。
A:第1実施形態
図1は、第1実施形態における情報システム100の構成を例示するブロック図である。情報システム100は、利用者Uによる電子楽器20の演奏を案内するコンピュータシステムである。情報システム100は、情報処理システム10と電子楽器20と応答生成システム30とを具備する。情報処理システム10は、例えばインターネット等の通信網200を介して応答生成システム30と通信可能である。
図1は、第1実施形態における情報システム100の構成を例示するブロック図である。情報システム100は、利用者Uによる電子楽器20の演奏を案内するコンピュータシステムである。情報システム100は、情報処理システム10と電子楽器20と応答生成システム30とを具備する。情報処理システム10は、例えばインターネット等の通信網200を介して応答生成システム30と通信可能である。
応答生成システム30は、プロンプトPに対応する応答情報Rを生成するサーバシステムである。第1実施形態のプロンプトPは、自然言語で表現された動作指示である。応答情報Rは、プロンプトPに対する応答を自然言語で表現したテキストデータである。
応答生成システム30は、訓練済の生成モデルMにより応答情報Rを生成する。生成モデルMは、プロンプトPに応じた応答情報Rを生成する生成型(generative)の確率モデルである。生成モデルMは、プロンプトPに対する応答情報Rの傾向を、事前の機械学習(pre-trained)により習得済である。具体的には、生成モデルMは、例えば応答生成等の自然言語処理タスクに特化して訓練された対話型の大規模言語モデル(LLM:Large Language Models)である。例えばセルフアテンション機構を利用したトランスフォーマーモデルにより実現された自然言語処理モデルが、生成モデルMとして例示される。
電子楽器20は、利用者Uによる楽曲の演奏を受付ける入力機器である。第1実施形態の電子楽器20は、例えばMIDI(Musical Instrument Digital Interface)規格に準拠した鍵盤楽器であり、相異なる音高に対応する複数の鍵21を具備する。なお、情報処理システム10に電子楽器20が搭載されてもよい。
利用者Uは、各鍵21を順次に操作することで楽曲を演奏する。電子楽器20は、利用者Uによる演奏に応じた楽音を放射するほか、当該演奏を表す演奏データ列Dを情報処理システム10に出力する。演奏データ列Dは、例えばMIDI規格に準拠したイベントデータの時系列である。具体的には、演奏データ列Dは、利用者Uが操作した鍵21に対応する音高を時系列に指定する。
情報処理システム10は、利用者Uによる電子楽器20の演奏を評価した結果に対応するプロンプトPを応答生成システム30に送信し、応答生成システム30から受信する応答情報Rを利用者Uに報知する。すなわち、利用者Uによる演奏の評価に応じた自然言語の応答情報Rが当該利用者Uに報知される。
図2は、情報処理システム10のブロック図である。情報処理システム10は、例えばスマートフォン、タブレット端末またはパーソナルコンピュータ等の情報装置で実現される。情報処理システム10は、制御装置11と記憶装置12と通信装置13と放音装置14と表示装置15と操作装置16とを具備する。なお、情報処理システム10は、単体の装置として実現されるほか、相互に別体で構成された複数の装置でも実現される。
制御装置11は、情報処理システム10の各要素を制御する単数または複数のプロセッサで構成される。例えば、制御装置11は、CPU(Central Processing Unit)、SPU(Sound Processing Unit)、DSP(Digital Signal Processor)、FPGA(Field Programmable Gate Array)、またはASIC(Application Specific Integrated Circuit)等の1種類以上のプロセッサにより構成される。
記憶装置12は、制御装置11が実行するプログラムと、制御装置11が使用する各種のデータとを記憶する単数または複数のメモリである。記憶装置12は、例えば磁気記録媒体または半導体記録媒体等の公知の記録媒体で構成される。記憶装置12は、複数種の記録媒体の組合せで構成されてもよい。また、情報処理システム10に対して着脱される可搬型の記録媒体、または通信網200を介して制御装置11が書込または読出を実行可能な記録媒体(例えばクラウドストレージ)が、記憶装置12として利用されてもよい。
記憶装置12は、利用者Uが演奏する楽曲を表す楽曲データCを記憶する。楽曲データCは、楽曲の楽譜を表すデータである。具体的には、楽曲データCは、楽曲を構成する複数の音符の各々について音高と発音期間とを指定する。例えば、楽曲データCは、MIDI規格に準拠したデータである。なお、音楽的な表情を表す演奏記号等の情報を、楽曲データCが指定してもよい。
通信装置13は、制御装置11による制御のもとで応答生成システム30と通信する。具体的には、通信装置13は、応答生成システム30にプロンプトPを送信し、応答生成システム30から送信された応答情報Rを受信する。
放音装置14は、制御装置11による制御のもとで音波を放射する。放音装置14は、例えばスピーカまたはヘッドホンである。具体的には、応答情報Rに対応する音声(以下「案内音声」という)を表す音声信号Vが放音装置14に供給される。放音装置14は案内音声を利用者Uに対して再生する。なお、音声信号Vをデジタルからアナログに変換するD/A変換器と、音声信号Vを増幅する増幅器とについては、便宜的に図示が省略されている。情報処理システム10とは別体の放音装置14を、情報処理システム10に対して有線または無線により接続してもよい。
表示装置15は、制御装置11による制御のもとで画像を表示する。表示装置15は、例えば液晶パネルまたは有機EL(Electroluminescence)パネル等の表示パネルで構成される。操作装置16は、利用者Uからの指示を受付ける入力機器である。操作装置16は、例えば、利用者Uが操作する操作子、または、利用者Uによる接触を検知するタッチパネルである。なお、情報処理システム10とは別体の表示装置15または操作装置16を、情報処理システム10に対して有線または無線により接続してもよい。
利用者Uは、操作装置16を操作することで、応答情報Rに関する条件(以下「応答条件X」という)を指示可能である。図3は、応答条件X(X1,X2,X3)を利用者Uが指示するための設定画面151の模式図である。表示装置15に設定画面151が表示される。応答条件Xは、応答情報Rが表す応答を実行する仮想的な応答者(例えば後述の案内キャラクタ152)の属性X1と、利用者Uの属性X2と、応答情報Rによる応答の語調X3とを含む。設定画面151は、複数の入力欄F(F1,F2,F3)が配置された画像である。
入力欄F1は、仮想的な応答者の属性X1を利用者Uが指示する入力ボックスである。利用者Uは、操作装置16の操作により応答者の属性X1の文字列を入力欄F1に入力する。例えば応答者の職業(ピアノの先生等)が属性X1として入力される。なお、事前に用意された複数の選択肢の何れかが属性X1として利用者Uにより選択されてもよい。
入力欄F2は、利用者Uの属性X2を利用者Uが指示する入力ボックスである。利用者Uは、操作装置16の操作により自身の属性X2の文字列を入力欄F2に入力する。例えば利用者Uの年代(小学生/中学生/高校生/大学生/大人等)が属性X2として入力される。なお、事前に用意された複数の選択肢の何れかが属性X2として利用者Uにより選択されてもよい。
入力欄F3は、応答情報Rによる応答の語調X3を利用者Uが指定する入力ボックスである。利用者Uは、操作装置16の操作により所望の語調X3の文字列を入力欄F3に入力する。例えば、「優しく」「厳しく」「率直(フランク)に」等、発言の雰囲気(トーン)が語調X3として指示される。なお、事前に用意された複数の選択肢の何れかが語調X3として利用者Uにより指示されてもよい。
図4は、情報処理システム10の機能的な構成を例示するブロック図である。制御装置11は、記憶装置12に記憶されたプログラムを実行することで複数の機能(情報取得部41,応答取得部42,動作制御部43)を実現する。
情報取得部41は、利用者Uによる電子楽器20の演奏を評価する。具体的には、情報取得部41は、電子楽器20から供給される演奏データ列Dと記憶装置12に記憶された楽曲データCとを比較し、両者間の相違点を利用者Uの演奏ミスとして特定する。すなわち、演奏ミスは、利用者Uによる演奏と楽曲データCが指定する標準的な演奏との間の相違である。情報取得部41は、利用者Uによる演奏において発生した演奏ミス毎に評価情報Yを生成する。評価情報Yは、利用者Uによる電子楽器20の演奏に関する評価を表す情報である。具体的には、評価情報Yは、演奏ミスの重要度Y1、演奏ミスの位置Y2、演奏ミスの種別Y3、および演奏ミスの詳細内容Y4を含む。
演奏ミスの重要度Y1は、演奏ミスの重要性の度合である。第1実施形態における重要度Y1は、重要度の高低を表す2値的な情報である。なお、重要度Y1は重要性の度合を多段階で表す情報でもよい。
演奏ミスの位置Y2は、楽曲内において演奏ミスが発生した箇所である。例えば、楽曲のうち演奏データ列Dと楽曲データCとが相違する箇所が、演奏ミスの位置Y2として特定される。例えば、楽曲内の小節番号または楽曲の始点からの経過時間等により指定される。なお、楽曲内の特定の時点のほか、楽曲内の特定の区間が、演奏ミスの位置Y2として指定されてもよい。
演奏ミスの種別Y3は、演奏ミスを種類により区別した分類を意味する。例えば、「音高のミス」「リズムのミス」および「演奏を躊躇している」等が、演奏ミスの種別Y3として指定される。前述の重要度Y1は、例えば演奏ミスの種別Y3に応じて設定される。例えば、演奏ミスの種別Y3毎に重要度Y1が事前に記憶装置12に記憶される。
演奏ミスの詳細内容Y4は、利用者Uによる演奏ミスまたは当該演奏ミスに対する指摘の具体的な内容である。例えば、「『ド』を演奏すべき箇所で『シ』を弾いている」「正しいタイミングよりも早い」「止まらずに弾かないといけない」等の演奏状態が、演奏ミスの詳細内容Y4として例示される。演奏ミスの詳細内容Y4は、例えば自然言語で表現される。
図4の応答取得部42は、評価情報Yに応じた自然言語の応答情報Rを取得する。第1実施形態の応答取得部42は、指示生成部421と情報受信部422とを具備する。指示生成部421は、応答生成システム30に対するプロンプトPを生成する。具体的には、指示生成部421は、応答条件Xと評価情報Yとを含むプロンプトPを生成する。
指示生成部421によるプロンプトPの生成には、図5の基礎文字列Bが利用される。基礎文字列Bは、プロンプトPの定型的な文字列を表すテンプレートであり、記憶装置12に事前に記憶される。基礎文字列Bは、複数の文字列b1~b6を含む。複数の文字列b1~b6の各々には空欄が設定される。各空欄は、基礎文字列Bのうち可変の情報(応答条件X,評価情報Y)が挿入される部分である。指示生成部421は、基礎文字列Bの各空欄に応答条件Xまたは評価情報Yを挿入することでプロンプトPを生成する。図6には、図5の基礎文字列Bを利用して指示生成部421が生成したプロンプトPが例示されている。
基礎文字列Bの文字列b1は、応答者に関する条件を表す部分である。指示生成部421は、設定画面151において利用者Uが指定した属性X1を文字列b1の空欄に挿入することで、プロンプトPのうち応答者に関する指示を生成する。以上の通り、第1実施形態のプロンプトPは、仮想的な応答者の属性X1を含む。
基礎文字列Bの文字列b2は、応答情報Rの報知先となる利用者Uの条件を表す部分である。指示生成部421は、設定画面151において利用者Uが指定した属性X2を文字列b2の空欄に挿入することで、プロンプトPのうち利用者Uに関する指示を生成する。以上の通り、第1実施形態のプロンプトPは、利用者Uの属性X2を含む。
基礎文字列Bの文字列b3は、応答情報Rによる応答の表現に関する条件を表す部分である。指示生成部421は、設定画面151において利用者Uが指定した語調X3を文字列b3の空欄に挿入することで、プロンプトPのうち応答情報Rの表現に関する指示を生成する。以上の通り、第1実施形態のプロンプトPは、応答情報Rに関する語調X3を含む。
基礎文字列Bの文字列b4~b6は、応答情報Rによる応答内容に関する条件を表す部分である。指示生成部421は、情報取得部41が生成した評価情報Yを文字列b4~b6の各々の空欄に挿入することで、プロンプトPのうち応答情報Rの内容に関する指示を生成する。
文字列b4は、重要な演奏ミスを通知する部分である。指示生成部421は、重要度Y1が高い各評価情報Yが指定する位置Y2と種別Y3と詳細内容Y4とを文字列b4の各空欄に挿入することで、プロンプトPのうち重要な演奏ミスに関する指示を生成する。
文字列b5および文字列b6は、軽微な演奏ミスを通知する部分である。指示生成部421は、重要度Y1が低い各評価情報Yが指定する位置Y2と種別Y3と詳細内容Y4とを文字列b5および文字列b6の各空欄に挿入することで、プロンプトPのうち軽微な演奏ミスに関する指示を生成する。なお、プロンプトPにおいて指定される演奏ミスの個数は可変である。
図4の指示生成部421は、以上の手順で生成したプロンプトPを通信装置13から応答生成システム30に送信する。応答生成システム30は、情報処理システム10から受信したプロンプトPを生成モデルMにより処理することで自然言語の応答情報Rを生成し、当該応答情報Rを情報処理システム10に送信する。図4の情報受信部422は、応答生成システム30から送信された応答情報Rを通信装置13により受信する。すなわち、情報受信部422は、訓練済の生成モデルMがプロンプトPに対して生成した応答情報Rを取得する。
図6には、前述のプロンプトPから生成された応答情報Rが図示されている。応答情報Rは、プロンプトPに含まれる応答条件Xおよび評価情報Yが反映された自然言語の応答を表す。具体的には、応答情報Rは、属性X1の応答者が属性X2の利用者Uに対して語調X3により発話すべき自然言語の応答文を表す。また、応答情報Rは、評価情報Yが表す演奏ミスを利用者Uに説明する自然言語の応答文を表す。具体的には、応答情報Rにおいては、重要な演奏ミスと軽微な演奏ミスとを区別して(重要度Y1)、演奏ミスの位置Y2と種別Y3と詳細内容Y4とが指摘される。
図6では利用者Uの属性X2が「大人」であるのに対し、図7には、属性X2が「小学生」である場合の応答情報Rが例示されている。利用者Uの属性X2が「大人」である場合(図6)には、大人同士の会話のように丁寧な口調の応答情報Rが生成される。他方、属性X2が「小学生」である場合(図7)には、大人が子供を指導する場合のように気さくで親しみ易い口調の応答情報Rが生成される。以上の通り、第1実施形態においては利用者Uの属性X2がプロンプトPに含まれるから、利用者Uの属性X2に応じた多様な応答情報Rを生成できる。同様に、第1実施形態においては仮想的な応答者の属性X1がプロンプトPに含まれるから、応答者の属性X1に応じた多様な応答情報Rを生成できる。
また、図6では語調X3が「優しく」であるのに対し、図8には、語調X3が「厳しく」である場合の応答情報Rが例示されている。語調X3が「優しく」である場合(図6)には、利用者Uの自尊心を傷付けないように優しい口調の応答情報Rが生成される。他方、語調X3が「厳しく」である場合(図8)には、厳格な口調の応答情報Rが生成される。以上の通り、第1実施形態においては応答情報Rに関する語調X3がプロンプトPに含まれるから、多様な語調の応答情報Rを生成できる。
図3の動作制御部43は、以上に説明した応答情報Rを利用者Uに報知する動作(以下「報知動作」という)を実行する。第1実施形態の報知動作は、表示装置15に表示された案内キャラクタ152(図9)により応答情報Rを利用者Uに報知する動作である。第1実施形態の動作制御部43は、音声合成部431と表示制御部432とを具備する。
音声合成部431は、応答情報Rに対応する案内音声の音声信号Vを生成する。案内音声は、応答情報Rが表す応答文を読上げる音声である。具体的には、音声合成部431は、応答情報Rに対して音声合成処理を実行することで音声信号Vを生成する。音声合成処理としては、例えば、複数の音声素片を接続する素片接続型の音声合成処理、または、例えば深層ニューラルネットワークまたはHMM(Hidden Markov Model)等の統計モデルを利用した統計モデル型の音声合成処理が例示される。音声合成部431は、音声信号Vを放音装置14に供給する。したがって、音声信号Vが表す案内音声が放音装置14から再生される。
表示制御部432は、図9の案内キャラクタ152を表示装置15に表示させる。案内キャラクタ152は、仮想空間内に配置されたオブジェクト(エージェント)である。具体的には、案内キャラクタ152は、仮想空間内において利用者Uによる電子楽器20の演奏を指導する仮想的な指導者である。
表示制御部432は、放音装置14による案内音声の再生に並行して案内キャラクタ152に発話動作を実行させる。発話動作は、音声信号Vに応じて案内キャラクタ152の口元の形状を変化させる動作(いわゆる口パク)である。第1実施形態の報知動作は、案内キャラクタ152による発話動作と案内音声の再生とを含む。発話動作と案内音声の再生とが並列に実行されることで、案内キャラクタ152が利用者Uを指導しているような感覚を利用者Uに知覚させることが可能である。
図10は、制御装置11が実行する処理(以下「評価報知処理」という)のフローチャートである。例えば操作装置16に対する利用者Uからの指示を契機として評価報知処理が開始される。
評価報知処理が開始されると、制御装置11(応答取得部42)は、図3の設定画面151を表示装置15に表示する(Sa1)。制御装置11(応答取得部42)は、設定画面151に対する応答条件X(X1,X2,X3)の入力を利用者Uから受付ける(Sa2)。利用者Uから受付けた応答条件Xは記憶装置12に記憶される。
応答条件Xを入力した利用者Uは電子楽器20の演奏を開始する。制御装置11(情報取得部41)は、利用者Uによる電子楽器20の演奏を評価する(Sa3)。具体的には、制御装置11は、利用者Uによる演奏の過程で発生した演奏ミス毎に評価情報Yを生成する。
制御装置11(指示生成部421)は、応答条件Xと評価情報Yとを含むプロンプトPを生成する(Sa4)。具体的には、制御装置11は、記憶装置12に記憶された基礎文字列Bの各空欄に応答条件X(X1,X2,X3)と評価情報Y(Y2,Y3,Y4)とを挿入することでプロンプトPを生成する。
制御装置11(指示生成部421)は、通信装置13によりプロンプトPを応答生成システム30に送信する(Sa5)。すなわち、制御装置11は、応答生成システム30に対して応答情報Rの生成を要求する。制御装置11(情報受信部422)は、応答生成システム30が生成および送信した応答情報Rを通信装置13により受信する(Sa6)。
制御装置11(音声合成部431)は、応答情報Rに対応する案内音声を表す音声信号Vを生成し(Sa7)、音声信号Vを放音装置14に出力する(Sa8)。放音装置14に対する音声信号Vの出力に並行して、制御装置11(表示制御部432)は、表示装置15に表示した案内キャラクタ152に発話動作を実行させる(Sa9)。すなわち、案内音声の再生(Sa7,Sa8)と、案内キャラクタ152による発話動作(Sa9)とを含む報知動作が実行される。
以上に説明した通り、第1実施形態においては、電子楽器20の演奏に関する評価情報Yに応じた自然言語の応答情報Rが取得され、案内キャラクタ152により応答情報Rを報知する報知動作が実行される。したがって、例えば利用者Uによる電子楽器20の演奏を評価した数値(例えば評価スコア)が表示される形態と比較して、電子楽器20の演奏に関する評価を利用者Uが容易かつ適切に理解できる。また、案内キャラクタ152により電子楽器20の演奏を指導されている感覚を享受する特有の顧客体験を、利用者Uに提供できる。
また、第1実施形態においては、利用者Uによる電子楽器20の演奏に関する評価情報Yを含むプロンプトPが生成され、訓練済の生成モデルMがプロンプトPに対して生成した自然言語の応答情報Rが取得される。したがって、利用者Uによる電子楽器20の演奏に関する評価(評価情報Y)に対して統計的に妥当で言語的にも自然な応答情報Rを利用者Uに報知できる。第1実施形態においては特に、演奏ミスに関する各種の情報(重要度Y1,位置Y2,種別Y3および詳細内容Y4)が評価情報Yに含まれるから、利用者Uによる演奏ミスに関する多様な情報を含む応答情報Rを生成できる。
B:第2実施形態
第2実施形態を説明する。なお、以下に例示する各態様において機能が第1実施形態と同様である要素については、第1実施形態の説明と同様の符号を流用して各々の詳細な説明を適宜に省略する。
第2実施形態を説明する。なお、以下に例示する各態様において機能が第1実施形態と同様である要素については、第1実施形態の説明と同様の符号を流用して各々の詳細な説明を適宜に省略する。
図11は、第2実施形態におけるプロンプトPおよび応答情報Rの模式図である。第2実施形態のプロンプトPは、第1実施形態と同様の要素に加えて、演奏ミス毎の識別情報Gと、各演奏ミスの識別情報Gを応答情報Rに含ませる指示Z1とを含む。
識別情報Gは、評価情報Yが表す各演奏ミスを識別するための符号列(タグ)である。具体的には、識別情報Gは、演奏ミスを表す符号g1と、重要度Y1に対応する符号g2と、演奏ミスに付与された番号g3とを含む。符号g2は、重要度Y1が高い場合には「i」に設定され、重要度Y1が低い場合には「n」に設定される。他方、指示Z1は、応答情報Rのうち各演奏ミスを指摘する部分に識別情報Gを付加することを表す自然言語である。
図11には、以上に説明したプロンプトPから生成される応答情報Rが例示されている。図11から理解される通り、指示Z1を含むプロンプトPが生成モデルMにより処理される結果、識別情報Gを含む応答情報Rが生成される。すなわち、指示Z1に沿った応答情報Rが生成される。具体的には、応答情報Rのうち各演奏ミスを指摘する部分の直後に、当該演奏ミスの識別情報Gが設定される。
図12は、第2実施形態における情報処理システム10の機能的な構成を例示するブロック図である。第2実施形態の制御装置11は、記憶装置12に記憶されたプログラムを実行することで、第1実施形態と同様の機能(情報取得部41,応答取得部42,動作制御部43)に加えて判定処理部44として機能する。判定処理部44は、応答情報Rの適否を判定する。判定処理部44以外の各部の動作は、第1実施形態と同様である。
図13は、第2実施形態における評価報知処理のフローチャートである。評価報知処理の開始から応答情報Rの取得(Sb6)までの処理は第1実施形態と同様である。応答情報Rを取得すると、制御装置11(判定処理部44)は、応答情報Rの適否を判定する判定処理Sbを実行する。第2実施形態の判定処理Sbは、第1判定処理Sb1と第2判定処理Sb2とを含む。なお、第1判定処理Sb1および第2判定処理Sb2の順序は反転されてもよい。
第1判定処理Sb1は、応答情報RがプロンプトPに対して適切であるか否かの判定である。具体的には、判定処理部44は、プロンプトPにおいて指定された演奏ミスの全部が応答情報Rにおいて言及されているか否かを判定する。
例えば、判定処理部44は、応答生成システム30に送信したプロンプトPと当該プロンプトPから生成された応答情報Rとを比較し、プロンプトPに含まれる演奏ミスの識別情報Gの全部が、応答情報Rにも含まれるか否かを判定する。全部の識別情報Gが応答情報Rに含まれる場合、第1判定処理Sb1の結果は肯定となる。他方、プロンプトPに含まれる識別情報Gの一部が応答情報Rに含まれない場合、第1判定処理Sb1の結果は否定となる。
第2判定処理Sb2は、応答情報Rに禁止語句が含まれているか否かを判定する処理である。禁止語句は、教育的または社会的な観点から不適切な語句である。複数の禁止語句が記憶装置12に事前に記憶される。操作装置16に対する利用者Uからの操作に応じて禁止語句が設定されてもよい。
具体的には、判定処理部44は、記憶装置12に記憶された複数の禁止語句の何れかが応答情報Rに含まれないか否かを判定する。応答情報Rに禁止語句が含まれない場合、第2判定処理Sb2の結果は肯定(応答情報Rは適切)となる。他方、応答情報Rに禁止語句が含まれる場合、第2判定処理Sb2の結果は否定(応答情報Rは不適切)となる。
第1判定処理Sb1および第2判定処理Sb2の双方の結果が肯定である場合、応答情報Rは適切である。したがって、制御装置11(動作制御部43)は、第1実施形態と同様に、案内音声の再生(Sa7,Sa8)と案内キャラクタ152による発話動作(Sa9)とを含む報知動作を実行する。すなわち、制御装置11は、応答情報Rが適切であると判定した場合に報知動作を実行する。なお、音声信号Vの生成(Sa7)において、応答情報Rの識別情報Gは案内音声に含まれない。すなわち、識別情報Gは、音声合成部431による音声合成の対象から除外される。
他方、第1判定処理Sb1または第2判定処理Sb2の結果が否定である場合、応答情報Rは不適切である。したがって、応答情報Rの再生成が実行される。具体的には、制御装置11(指示生成部421)は、直前のステップSa4において生成したプロンプトPを、応答生成システム30に対して通信装置13から再送信する(Sa5)。
応答生成システム30は、情報処理システム10から受信したプロンプトPを生成モデルMにより処理することで応答情報Rを生成する。プロンプトPが共通する場合でも、生成モデルMが生成する応答情報Rは生成毎に変化する。すなわち、制御装置11が不適切と判定した応答情報Rとは別個の応答情報Rが生成される。
制御装置11(情報受信部422)は、応答生成システム30が生成および送信した応答情報Rを通信装置13により受信する(Sa6)。以上の説明から理解される通り、判定処理Sbにおいて応答情報Rが適切と判定されるまで、プロンプトPの送信(Sb5)と応答情報Rの受信(Sb6)とが反復される。
第2実施形態においても第1実施形態と同様の効果が実現される。また、第2実施形態においては、応答情報Rが適切であると判定された場合に報知動作が実行される。したがって、報知動作が無条件に実行される形態と比較して、不適切な応答情報Rが利用者Uに報知される可能性を低減できる。
また、第2実施形態においては演奏ミスの識別情報Gを含む応答情報Rが生成モデルMにより生成される。したがって、プロンプトPにおいて指定された全部の演奏ミスが応答情報Rにも含まれているか否かを、識別情報Gにより簡便に確認できる。
C:第3実施形態
図14は、第3実施形態におけるプロンプトPおよび応答情報Rの模式図である。第3実施形態のプロンプトPは、第1実施形態と同様の要素に加えて、応答情報Rに動作情報Qを含ませる指示Z2を含む。動作情報Qは、応答情報Rの報知の過程において実行されるべき動作を識別するための情報(タグ)である。具体的には、指示Z2は、応答情報Rの報知の過程で実行されるべき動作を表す自然言語の語句と、当該動作を表す動作情報Qとを含む。動作を表す語句は、例えば「生徒を見詰めるとき」「ピアノを見詰めるとき」「微笑むとき」「拍手するとき」等の語句である。動作情報Qは、各動作を識別するための識別情報である。
図14は、第3実施形態におけるプロンプトPおよび応答情報Rの模式図である。第3実施形態のプロンプトPは、第1実施形態と同様の要素に加えて、応答情報Rに動作情報Qを含ませる指示Z2を含む。動作情報Qは、応答情報Rの報知の過程において実行されるべき動作を識別するための情報(タグ)である。具体的には、指示Z2は、応答情報Rの報知の過程で実行されるべき動作を表す自然言語の語句と、当該動作を表す動作情報Qとを含む。動作を表す語句は、例えば「生徒を見詰めるとき」「ピアノを見詰めるとき」「微笑むとき」「拍手するとき」等の語句である。動作情報Qは、各動作を識別するための識別情報である。
図14には、以上に説明したプロンプトPから生成される応答情報Rが例示されている。図14から理解される通り、指示Z2を含むプロンプトPが生成モデルMにより処理される結果、動作情報Qを含む応答情報Rが生成される。すなわち、指示Z2に沿った応答情報Rが生成される。
具体的には、応答情報Rのうち各動作を実行すべき部分に、当該動作の動作情報Qが設定される。例えば、応答情報Rの冒頭には、利用者Uを見詰める動作の動作情報Q=<gaze-student>が設定され、応答情報Rのうち演奏ミスを指摘する部分の直後には、利用者Uを見詰める動作の動作情報Q=<gaze-student>、または、ピアノを見詰める動作の動作情報Q=<gaze-piano>が設定される。また、応答情報Rの最後には、微笑む動作の動作情報Q=<smile>と、拍手する動作の動作情報Q=<applause>とが設定される。以上の説明から理解される通り、第3実施形態の応答情報Rは、応答情報Rの報知の過程で実行されるべき動作を指定する。
評価報知処理の手順は、案内キャラクタ152の制御(Sa9)を除き、第1実施形態と同様である。評価報知処理のステップSa9において、制御装置11(表示制御部432)は、応答情報Rにおいて指定された動作を案内キャラクタ152に実行させる。具体的には、応答情報Rのうち動作情報Qの近傍の部分の案内音声が再生される時点において、制御装置11は、当該動作情報Qが表す動作を案内キャラクタ152に実行させる。
例えば、案内音声の冒頭の再生に並行して、案内キャラクタ152は、利用者Uを見詰める動作(Q=<gaze-student>)を実行する。案内音声のうち演奏ミスを指摘する部分の直後に、案内キャラクタ152は、利用者Uを見詰める動作(Q=<gaze-student>)またはピアノを見詰める動作(Q=<gaze-piano>)を実行する。また、案内音声を最後まで再生すると、案内キャラクタ152は、微笑む動作(Q=<smile>)と拍手する動作(Q=<applause>)とを実行する。なお、応答情報Rの動作情報Qは案内音声に含まれない。すなわち、動作情報Qは、音声合成部431による音声合成の対象から除外される。
第3実施形態においても第1実施形態と同様の効果が実現される。また、第3実施形態においては、報知動作において案内キャラクタ152が種々の動作を実行する。したがって、案内キャラクタ152の動作を多様化することが可能である。なお、以上の説明においては第1実施形態を基礎として第3実施形態を説明したが、応答情報Rの適否を判定する第2実施形態の構成は、第3実施形態にも同様に適用される。
D:変形例
以上に例示した各態様に付加される具体的な変形の態様を以下に例示する。以下の例示から任意に選択された2以上の態様を、相互に矛盾しない範囲で適宜に併合してもよい。
以上に例示した各態様に付加される具体的な変形の態様を以下に例示する。以下の例示から任意に選択された2以上の態様を、相互に矛盾しない範囲で適宜に併合してもよい。
(1)前述の各形態においては、応答条件Xとして、応答者の属性X1と利用者Uの属性X2と応答の語調X3とを例示したが、応答条件Xは以上の例示に限定されない。例えば、応答の言語が応答条件Xとして指定されてもよい。応答情報Rは、応答条件Xとして指定された言語で表現される。
また、応答情報Rにおける応答の口癖X4が応答条件Xとして指定されてもよい。図15には、「ニャー」という口癖X4を応答条件Xとして指定するプロンプトPと、当該プロンプトPに応じて生成される応答情報Rとが図示されている。応答条件Xとして指定された口癖X4が各文の語尾に付加された応答情報Rが生成される。語調X3と口癖X4とを含めて応答の語調(口調)と解釈してもよい。また、図16に例示される通り、利用者Uによる電子楽器20の演奏を評価した結果の総評X5が応答条件Xに含まれてもよい。応答情報Rが表す応答文は、総評X5に応じて変化する。
(2)前述の各形態においては、利用者Uによる演奏ミス毎に評価情報Yを生成したが、評価情報Yは演奏ミスを表す情報に限定されない。例えば、評価情報Yは、利用者Uによる演奏のうち良好な点(高評価ポイント)を表してもよい。高評価ポイントの評価情報Yを含むプロンプトPによれば、利用者Uの演奏を賞賛する応答情報Rが生成される。また、評価情報Yにおいて、重要度Y1、位置Y2、種別Y3および詳細内容Y4の各々は、省略されてもよい。
(3)前述の各形態においては、応答条件X(X1,X2,X3)を利用者Uが選択する形態を例示したが、応答条件Xを設定する方法は以上の例示に限定されない。例えば、案内キャラクタ152毎に応答条件Xが事前に記憶装置12に記憶されてもよい。利用者Uは、例えば操作装置16に対する操作により複数の案内キャラクタ152の何れかを選択する。応答取得部42(指示生成部421)は、記憶装置12に記憶された複数の応答条件Xのうち利用者Uが選択した案内キャラクタ152に対応する応答条件Xを含むプロンプトPを生成する。以上の形態によれば、応答条件Xの設定に関する利用者Uの付加を低減できる。
(4)第2実施形態においては、応答情報Rが不適切である場合に、送信済のプロンプトPが再送信される形態を例示したが、応答情報Rが不適切である場合の動作は、以上の例示に限定されない。例えば、制御装置11(指示生成部421)は、送信済のプロンプトPを所定の規則により編集し、編集後のプロンプトPを送信してもよい。例えば禁止語句を含むことで応答情報Rが不適切と判定された場合、その禁止語句を含まないことが条件として追加されたプロンプトPが、指示生成部421により生成される。また、制御装置11(応答取得部42)は、応答情報Rが不適切である旨のメッセージを表示装置15に表示してもよい。
(5)第2実施形態において、応答情報Rの適否を判定する判定処理Sbの方法は、前述の例示に限定されない。例えば、識別情報Gを利用する第1判定処理Sb1は省略されてもよい。すなわち、判定処理部44が応答情報Rの適否を判定する形態において、プロンプトPおよび応答情報Rの識別情報Gは必須ではない。
(6)前述の各形態においては、重要度Y1が演奏ミスの種別Y3に応じて設定される形態を例示したが、重要度Y1の設定の方法は以上の例示に限定されない。具体的には、情報取得部41は、利用者Uの過去の演奏の履歴に応じて重要度Y1を設定してもよい。例えば、利用者Uの演奏において頻繁に発生する演奏ミスは、改善の優先度が高く重要性が高いという傾向がある。以上の傾向を考慮して、情報取得部41は、利用者Uの過去の演奏において発生回数が多い演奏ミスの重要度Y1を大きい数値に設定する。
(7)前述の各形態においては、応答条件Xと評価情報Yとを含むプロンプトPを例示したが、プロンプトPの内容は以上の例示に限定されない。例えば、応答条件Xおよび評価情報Yの一方が省略された形態、または、応答条件Xおよび評価情報Y以外の情報がプロンプトPに含まれる形態も想定される。
(8)前述の各形態においては、利用者Uによる楽曲の演奏の終了後に応答情報Rの取得(Sa6)等の処理が実行される形態を例示したが、演奏の評価(Sa3)と応答情報Rの取得(Sa4~Sa6)と報知動作(Sa7~Sa9)とは、利用者Uによる楽曲の演奏に並行して実行されてもよい。
例えば、時間軸上で楽曲を区分した複数の単位期間の各々について、以上の処理が実行されてもよい。単位期間は、例えば、音楽的な意味に応じて楽曲を区分した構造期間である。構造期間は、例えばイントロ(intro)、Aメロ(verse)、Bメロ(bridge)、サビ(chorus)およびアウトロ(outro)等の各期間である。
また、利用者Uによる楽曲の演奏に並行して当該演奏の評価を実行し、演奏ミスが発生するたびに(すなわち評価情報Yの生成毎に)、応答情報Rの取得(Sa4~Sa6)と報知動作(Sa7~Sa9)とが実行されてもよい。
(9)前述の各形態においては、案内キャラクタ152が応答情報Rを利用者Uに報知する形態を例示したが、案内キャラクタ152の表示は省略されてもよい。例えば、音声信号Vが表す案内音声の再生のみで応答情報Rを利用者Uに報知する形態も想定される。すなわち、動作制御部43から表示制御部432は省略されてよい。
また、前述の各形態においては、応答情報Rが表す案内音声を再生する形態を例示したが、応答情報Rを利用者Uに報知する方法(報知動作)は以上の例示に限定されない。例えば、報知動作は、応答情報Rが表す応答文を表示装置15に表示する動作、または、応答情報Rが表す応答文を印刷装置により印刷する動作でもよい。利用者Uが所有する端末装置に応答情報Rを送信する動作も、報知動作の一例である。
また、動作制御部43は、楽曲データCが表す楽譜を表示装置15に表示し、楽譜のうち応答情報Rが表す演奏ミスに対応する箇所を強調表示してもよい。動作制御部43は、利用者Uが演奏ミスにより演奏した音符と、楽曲データCが表す正確な音符とを、表示装置15に対比的に表示してもよい。また、動作制御部43は、楽譜のうち演奏ミスに対応する箇所を指示する動作を案内キャラクタ152に実行させてもよい。動作制御部43は、楽曲データCが表す楽曲のうち応答情報Rが表す演奏ミスに対応する箇所の演奏を放音装置14により再生してもよい。以上の例示した通り、応答情報Rを利用者Uに報知する任意の動作が「報知動作」には包含される。
(10)前述の各形態においては、利用者Uによる電子楽器20の演奏を情報取得部41が評価する形態を例示したが、情報取得部41は、外部装置が生成した評価情報Yを通信装置13により受信してもよい。例えば、電子楽器20が評価情報Yを生成する構成において、情報取得部41は、電子楽器20から送信された評価情報Yを通信装置13により受信する。以上の説明から理解される通り、情報取得部41は、利用者Uによる電子楽器20の演奏に関する評価情報Yを取得する要素として包括的に表現される。評価情報Yの「取得」には、評価情報Yの「生成」(すなわち演奏の評価)のほか評価情報Yの「受信」も包含される。
(11)前述の各形態においては、情報処理システム10とは別個の応答生成システム30が応答情報Rを生成する形態を例示したが、情報処理システム10が応答情報Rを生成してもよい。例えば、情報処理システム10に生成モデルMが搭載されてもよい。応答取得部42は、指示生成部421が生成したプロンプトPを生成モデルMにより処理することで応答情報Rを生成する。以上の説明から理解される通り、応答取得部42は、応答情報Rを取得する要素として包括的に表現される。応答情報Rの「取得」には、応答情報Rの「受信」のほか応答情報Rの「生成」も包含される。
(12)前述の各形態においては、利用者Uによる演奏に関する評価を表す評価情報Yを例示したが、情報取得部41が取得する情報(演奏情報)は、評価情報Yに限定されない。例えば、利用者Uによる演奏の状況に関するテキストデータも演奏情報として例示される。演奏情報の一例であるテキストデータは、例えば、利用者Uが演奏する楽器の種類または利用者Uが演奏する音名等である。すなわち、演奏情報の生成にあたり演奏の評価(評価情報Y)は必須ではない。
また、例えば、利用者Uによる演奏の撮像により撮像装置が生成する映像データ、または、利用者Uによる演奏で発音される楽音の収音により収音装置が生成する音響データも、演奏情報として例示される。以上に例示した2種類以上の情報の組合せを、情報取得部41が演奏情報として取得してもよい。
以上の例示から理解される通り、「演奏情報」は、利用者Uによる演奏に関する情報として包括的に表現され、評価情報Yとテキストデータと映像データと音響データとは、「演奏情報」の一例である。
(13)前述の各形態においては、電子楽器20として鍵盤楽器を例示したが、演奏評価の対象となる楽器は鍵盤楽器に限定されない。例えば、弦楽器、管楽器、打楽器等の任意の種類の楽器について、前述の各形態が同様に適用される。
前述の各形態においては演奏データ列Dを生成可能な電子楽器20を例示したが、演奏評価の対象となる楽器は電子楽器20に限定されない。例えば自然楽器の演奏の評価にも前述の各形態が同様に適用される。利用者Uが自然楽器を演奏する形態において、情報取得部41は、自然楽器からの放射音の収音により生成される音響信号を解析することで、評価情報Yを生成する。演奏評価には公知の技術が任意に採用される。
また、評価対象は楽器の演奏に限定されない。例えば、利用者Uによる歌唱を評価する構成にも、前述の各形態は同様に適用される。例えば、情報取得部41は、歌唱音声の収音により生成される音響信号を解析することで、評価情報Yを生成する。歌唱評価には公知の技術が任意に採用される。
(14)以上に例示した情報処理システム10の機能は、前述の通り、制御装置11を構成する単数または複数のプロセッサと、記憶装置12に記憶されたプログラムとの協働により実現される。本開示に係るプログラムは、コンピュータが読取可能な記録媒体に格納された形態で提供されてコンピュータにインストールされ得る。記録媒体は、例えば非一過性(non-transitory)の記録媒体であり、CD-ROM等の光学式記録媒体(光ディスク)が好例であるが、半導体記録媒体または磁気記録媒体等の公知の任意の形式の記録媒体も包含される。なお、非一過性の記録媒体とは、一過性の伝搬信号(transitory, propagating signal)を除く任意の記録媒体を含み、揮発性の記録媒体も除外されない。また、配信装置が通信網を介してプログラムを配信する構成では、当該配信装置においてプログラムを記憶する記憶媒体が、前述の非一過性の記録媒体に相当する。
E:付記
以上に例示した形態から、例えば以下の構成が把握される。
以上に例示した形態から、例えば以下の構成が把握される。
本開示のひとつの態様(態様1)に係る情報処理方法は、利用者による楽器の演奏に関する演奏情報を取得し、前記演奏情報に応じた自然言語の応答情報を取得し、表示装置に表示された案内キャラクタにより前記応答情報を報知する報知動作を実行する。以上の態様においては、楽器の演奏に関する演奏情報に応じた自然言語の応答情報が取得され、案内キャラクタにより応答情報を報知する報知動作が実行される。したがって、例えば利用者による楽器の演奏に関する数値(例えば評価スコア)が表示される形態と比較して、楽器の演奏に関する情報を利用者が容易かつ適切に理解できる。また、案内キャラクタにより演奏を案内(例えば指導)されている感覚を利用者に付与できる。
「演奏情報」は、利用者による楽器の演奏に関する任意の情報である。例えば、利用者による楽器の演奏を評価した結果を表す評価情報が演奏情報として例示される。評価情報は、例えば、利用者による演奏ミスを表す情報である。演奏ミスを表す情報は、例えば、演奏ミスの重要度、楽曲における演奏ミスの位置、演奏ミスの種別、演奏ミスの具体的な内容等の情報である。また、利用者による演奏の撮像により生成された映像データ、または利用者による楽器の演奏音の収音により生成された音響データも、「演奏情報」として例示される。演奏情報の「取得」は、外部装置が生成した演奏情報を受信する動作と、演奏情報を自身が生成する動作との双方を包含する。
「案内キャラクタ」は、表示装置に表示された仮想的なオブジェクト(エージェント)である。案内キャラクタの具体例は、例えば人間または動物等の仮想的な生物であるが、例えばロボット等の無生物的なオブジェクトも「案内キャラクタ」には包含され得る。「案内キャラクタの動作の制御」は、応答情報に応じた動作を案内キャラクタに実行させる制御処理である。例えば、応答情報を発話する動作を案内キャラクタに実行させる処理が例示される。
「報知動作」は、応答情報を利用者に報知するための出力動作である。例えば、応答情報を表示装置に表示する処理、応答情報が表す音声を再生する処理、仮想的な案内キャラクタを応答情報に応じて動作させる処理、楽曲のうち応答情報により指示された区間の楽譜を表示する処理、当該区間の演奏を再生する処理等が、「報知動作」として例示される。
態様1の具体例(態様2)において、前記応答情報の取得においては、前記演奏情報を含むプロンプトを生成し、訓練済の生成モデルが前記プロンプトに対して生成した前記応答情報を取得する。以上の態様においては、利用者による楽器の演奏に関する演奏情報を含むプロンプトが生成され、訓練済の生成モデルがプロンプトに対して生成した自然言語の応答情報が取得される。したがって、利用者による楽器の演奏に関する情報(例えば評価)に対して統計的に妥当で言語的にも自然な応答情報を利用者に報知できる。
「プロンプト」は、訓練済の生成モデルが応答情報を生成するための基礎となる入力情報である。「プロンプト」は、訓練済の生成モデルによる応答情報の生成の指示、または、当該生成モデルが生成すべき応答情報に関する指示とも表現される。「プロンプト」は、単体のプロンプトおよび複数のプロンプトの集合の双方を含む。
「自然言語の応答情報」は、プロンプトに対する応答を自然言語により表現した情報である。具体的には、演奏情報に対応する自然言語の応答文が「応答情報」として生成される。例えば、プロンプトに含まれる演奏情報に沿って利用者を指導または案内する応答文が、応答情報として例示される。応答情報の「取得」は、生成モデルを含む外部装置が生成した演奏情報を受信する動作と、生成モデルを利用して自身が演奏情報を生成する動作との双方を包含する。
態様2の具体例(態様3)において、前記演奏情報は、前記演奏に関する評価を表す評価情報を含む。また、態様3の具体例(態様4)において、前記演奏情報は、前記演奏において発生した演奏ミスの種別、前記演奏ミスの重要度、楽曲における前記演奏ミスの位置、および前記演奏ミスの内容のうち少なくともひとつを含む。以上の態様によれば、演奏ミスに関する各種の情報が演奏情報に含まれるから、演奏ミスに関する多様な情報を含む応答情報を生成できる。
態様2から態様4の何れかの具体例(態様5)において、前記プロンプトは、前記演奏において発生した演奏ミスの識別情報を前記応答情報に含ませる指示を含む。以上の態様によれば、演奏ミスの識別情報を含む応答情報が生成モデルにより生成される。したがって、プロンプトにおいて指定された全部の演奏ミスが応答情報にも含まれているか否かを、識別情報により簡便に確認できる。
態様2から態様5の何れかの具体例(態様6)において、前記プロンプトは、前記利用者の属性を含む。以上の態様においては、利用者の属性がプロンプトに含まれるから、利用者の属性に応じた多様な応答情報を生成できる。
「利用者の属性」は、例えば、年齢、性別、年代、職業、性格(優しい性格、強気な性格等)、感情(怒っている、悲しんでいる等)等である。また、楽器の演奏に対する熟練の度合(初級者,中級者,上級者)も「利用者の属性」に含まれる。
態様2から態様6の何れかの具体例(態様7)において、前記プロンプトは、前記応答情報を応答する仮想的な応答者の属性を含む。以上の態様においては、仮想的な応答者の属性がプロンプトに含まれるから、当該属性に応じた多様な応答情報を生成できる。
「応答者の属性」は、例えば、年齢、性別、年代、職業、性格(優しい性格、強気な性格等)、感情(怒っている、悲しんでいる等)等である。また、演奏の指導に対する熟練の度合(初級者,中級者,上級者)も「応答者の属性情報」に含まれる。
態様2から態様7の何れかの具体例(態様8)において、前記プロンプトは、前記応答情報に関する語調を含む。以上の態様においては、応答情報に関する語調がプロンプトに含まれるから、多様な語調の応答情報を生成できる。
「語調」は、応答情報が表す語句の調子である。例えば、語句の雰囲気(トーン)または感情等が「語調」に包含される。「雰囲気」としては、例えば、優しい口調、厳しい口調、堅苦しい口調、公式(フォーマル)な口調、率直(フランク)な口調等が例示される。また、「感情」としては、例えば、怒った口調、悲しい口調、楽しい口調等が例示される。「語調」は、以上に例示した雰囲気または感情のほか、「口癖」や「方言」等も包含する。「口癖」は、頻繁に発話される言葉である。例えば、特定の言葉(例えば「~だニャー」「~だワン」等)を語尾に付加することが「口癖」として例示される。「方言」は、特定の言語に関する地方毎の相違である。
態様1から態様8の何れかの具体例(態様9)において、前記応答情報は、当該応答情報の報知の過程で実行されるべき動作を指定し、前記報知動作においては、前記応答情報において指定された動作を前記案内キャラクタに実行させる。以上の態様においては、報知動作において案内キャラクタが種々の動作を実行する。したがって、案内キャラクタの動作を多様化することが可能である。
「応答情報の報知の過程で実行されるべき動作」は、例えば演奏情報に直接的には関係しない案内キャラクタの補助的な動作である。例えば、利用者を見詰める動作、仮想空間内の楽器を見詰める動作、または微笑む動作等、現実の指導者が演奏の指導の過程で実行し得る各種の動作が「応答情報の報知の過程で実行されるべき動作」として例示される。
態様1から態様9の何れかの具体例(態様10)において、前記応答情報の適否を判定し、前記応答情報が適切であると判定した場合に、前記報知動作を実行する。以上の態様においては、応答情報が適切であると判定された場合に報知動作が実行される。したがって、報知動作が無条件に実行される形態と比較して、不適切な応答情報が利用者に報知される可能性を低減できる。
「応答情報の適否の判定」は、応答情報がプロンプトに対して適切であるか否かの判定のほか、応答情報が教育的または社会的な観点から適切であるか否かの判定を含む。前者の判定は、例えば、プロンプトに含まれる全部の演奏情報が応答情報に反映されているか否かの判定である。また、後者の判定は、例えば教育的または社会的に適切でない語句が応答情報に含まれているか否かの判定である。
本開示のひとつの態様(態様11)に係る情報処理方法は、利用者による楽器の演奏に関する演奏情報を取得し、訓練済の生成モデルが自然言語の応答情報を生成するためのプロンプトであって、前記演奏情報を含むプロンプトを生成する。以上の態様においては、利用者による楽器の演奏に関する演奏情報を含むプロンプトが生成される。したがって、利用者による楽器の演奏に関する演奏情報に対して統計的に妥当で言語的にも自然な応答情報を生成できる。また、演奏情報に対して統計的に妥当で言語的にも自然な応答情報を取得するという特有の顧客体験を、利用者に提供できる。
本開示のひとつの態様(態様12)に係る情報処理システムは、利用者による楽器の演奏に関する演奏情報を取得する情報取得部と、前記演奏情報に応じた自然言語の応答情報を取得する応答取得部と、表示装置に表示された案内キャラクタに前記応答情報を報知する報知動作を実行させる動作制御部とを具備する。
本開示のひとつの態様(態様13)に係るプログラムは、利用者による楽器の演奏に関する演奏情報を取得する情報取得部、前記演奏情報に応じた自然言語の応答情報を取得する応答取得部、および、表示装置に表示された案内キャラクタに前記応答情報を報知する報知動作を実行させる動作制御部、としてコンピュータシステムを機能させる。
100…情報システム、200…通信網、10…情報処理システム、11…制御装置、12…記憶装置、13…通信装置、14…放音装置、15…表示装置、151…設定画面、152…案内キャラクタ、16…操作装置、20…電子楽器、21…鍵、30…応答生成システム、41…情報取得部、42…応答取得部、421…指示生成部、422…情報受信部、43…動作制御部、431…音声合成部、432…表示制御部、44…判定処理部。
Claims (12)
- 利用者による楽器の演奏に関する演奏情報を取得し、
前記演奏情報に応じた自然言語の応答情報を取得し、
表示装置に表示された案内キャラクタにより前記応答情報を報知する報知動作を実行する
コンピュータシステムにより実現される情報処理方法。 - 前記応答情報の取得においては、
前記演奏情報を含むプロンプトを生成し、
訓練済の生成モデルが前記プロンプトに対して生成した前記応答情報を取得する
請求項1の情報処理方法。 - 前記演奏情報は、前記演奏に関する評価を表す評価情報を含む
請求項2の情報処理方法。 - 前記評価情報は、前記演奏において発生した演奏ミスの種別、前記演奏ミスの重要度、楽曲における前記演奏ミスの位置、および前記演奏ミスの内容のうち少なくともひとつを含む
請求項3の情報処理方法。 - 前記プロンプトは、前記演奏において発生した演奏ミスの識別情報を前記応答情報に含ませる指示を含む
請求項2の情報処理方法。 - 前記プロンプトは、前記利用者の属性を含む
請求項2の情報処理方法。 - 前記プロンプトは、前記応答情報を応答する仮想的な応答者の属性を含む
請求項2の情報処理方法。 - 前記プロンプトは、前記応答情報に関する語調を含む
請求項2の情報処理方法。 - 前記応答情報は、当該応答情報の報知の過程で実行されるべき動作を指定し、
前記報知動作においては、前記応答情報において指定された動作を前記案内キャラクタに実行させる
請求項1の情報処理方法。 - 前記応答情報の適否を判定し、
前記応答情報が適切であると判定した場合に、前記報知動作を実行する
請求項1の情報処理方法。 - 利用者による楽器の演奏に関する演奏情報を取得し、
訓練済の生成モデルが自然言語の応答情報を生成するためのプロンプトであって、前記演奏情報を含むプロンプトを生成する
コンピュータシステムにより実現される情報処理方法。 - 利用者による楽器の演奏に関する演奏情報を取得する情報取得部と、
前記演奏情報に応じた自然言語の応答情報を取得する応答取得部と、
表示装置に表示された案内キャラクタに前記応答情報を報知する報知動作を実行させる動作制御部と
を具備する情報処理システム。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US19/424,098 US20260112343A1 (en) | 2023-06-19 | 2025-12-17 | Information processing method and information processing system |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2023099954A JP2025000223A (ja) | 2023-06-19 | 2023-06-19 | 情報処理方法および情報処理システム |
| JP2023-099954 | 2023-06-19 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US19/424,098 Continuation US20260112343A1 (en) | 2023-06-19 | 2025-12-17 | Information processing method and information processing system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024262249A1 true WO2024262249A1 (ja) | 2024-12-26 |
Family
ID=93935052
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2024/019314 Ceased WO2024262249A1 (ja) | 2023-06-19 | 2024-05-27 | 情報処理方法および情報処理システム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20260112343A1 (ja) |
| JP (1) | JP2025000223A (ja) |
| WO (1) | WO2024262249A1 (ja) |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2007264400A (ja) * | 2006-03-29 | 2007-10-11 | Yamaha Corp | アクセサリ、電子楽器、教習装置およびプログラム |
| JP2012194241A (ja) * | 2011-03-15 | 2012-10-11 | Yamaha Corp | 評価装置 |
| JP2016014781A (ja) * | 2014-07-02 | 2016-01-28 | ヤマハ株式会社 | 歌唱合成装置および歌唱合成プログラム |
| JP2019056871A (ja) * | 2017-09-22 | 2019-04-11 | ヤマハ株式会社 | 再生制御方法および再生制御装置 |
| WO2020071149A1 (ja) * | 2018-10-05 | 2020-04-09 | ソニー株式会社 | 情報処理装置 |
| WO2022070771A1 (ja) * | 2020-09-30 | 2022-04-07 | ヤマハ株式会社 | 情報処理方法、情報処理システムおよびプログラム |
| WO2023105601A1 (ja) * | 2021-12-07 | 2023-06-15 | ヤマハ株式会社 | 情報処理装置、情報処理方法、及びプログラム |
-
2023
- 2023-06-19 JP JP2023099954A patent/JP2025000223A/ja active Pending
-
2024
- 2024-05-27 WO PCT/JP2024/019314 patent/WO2024262249A1/ja not_active Ceased
-
2025
- 2025-12-17 US US19/424,098 patent/US20260112343A1/en active Pending
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2007264400A (ja) * | 2006-03-29 | 2007-10-11 | Yamaha Corp | アクセサリ、電子楽器、教習装置およびプログラム |
| JP2012194241A (ja) * | 2011-03-15 | 2012-10-11 | Yamaha Corp | 評価装置 |
| JP2016014781A (ja) * | 2014-07-02 | 2016-01-28 | ヤマハ株式会社 | 歌唱合成装置および歌唱合成プログラム |
| JP2019056871A (ja) * | 2017-09-22 | 2019-04-11 | ヤマハ株式会社 | 再生制御方法および再生制御装置 |
| WO2020071149A1 (ja) * | 2018-10-05 | 2020-04-09 | ソニー株式会社 | 情報処理装置 |
| WO2022070771A1 (ja) * | 2020-09-30 | 2022-04-07 | ヤマハ株式会社 | 情報処理方法、情報処理システムおよびプログラム |
| WO2023105601A1 (ja) * | 2021-12-07 | 2023-06-15 | ヤマハ株式会社 | 情報処理装置、情報処理方法、及びプログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| US20260112343A1 (en) | 2026-04-23 |
| JP2025000223A (ja) | 2025-01-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US9355634B2 (en) | Voice synthesis device, voice synthesis method, and recording medium having a voice synthesis program stored thereon | |
| CN109949783B (zh) | 歌曲合成方法及系统 | |
| CN111052223B (zh) | 播放控制方法、播放控制装置及记录介质 | |
| CN111418006B (zh) | 声音合成方法、声音合成装置及记录介质 | |
| JP2012215645A (ja) | コンピュータを利用した外国語会話練習システム | |
| CN107851436A (zh) | 语音交互方法和语音交互设备 | |
| EP4213130B1 (en) | Device, system and method for providing a singing teaching and/or vocal training lesson | |
| Martín et al. | Sound synthesis for communicating nonverbal expressive cues | |
| Pardue et al. | Real-time aural and visual feedback for improving violin intonation | |
| JP2018049126A (ja) | 演奏教習装置、演奏教習プログラム、および演奏教習方法 | |
| JP2011028130A (ja) | 音声合成装置 | |
| JP2007256617A (ja) | 楽曲練習装置および楽曲練習システム | |
| Li et al. | Wavbench: Benchmarking reasoning, colloquialism, and paralinguistics for end-to-end spoken dialogue models | |
| JP2025000223A (ja) | 情報処理方法および情報処理システム | |
| KR102634347B1 (ko) | 사용자와의 인터랙션을 이용한 인지 기능 훈련 및 검사 장치 | |
| JP7710657B2 (ja) | 情報処理装置、及びその制御方法 | |
| CN117476180A (zh) | 基于音乐节奏对儿童言语能力进行评估训练的方法及系统 | |
| JP2022181361A (ja) | 学習支援システム | |
| Kaastra | Systematic approaches to the study of cognition in western art music performance | |
| JP2007304489A (ja) | 楽曲練習支援装置、制御方法及びプログラム | |
| JP7587670B1 (ja) | 音情報の学習支援装置、音情報の学習支援方法、音情報の学習支援プログラム及び記録媒体 | |
| JP7823706B2 (ja) | 情報処理システム、電子楽器、情報処理方法、プログラムおよび機械学習システム | |
| US12230244B1 (en) | Graphical user interface for customized storytelling | |
| JP7620741B1 (ja) | 在宅楽器練習用のサーバ端末、再生端末、プログラム、データ処理方法及び楽器練習システム | |
| JP2005316077A (ja) | 情報処理装置およびプログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24825656 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |