WO2021109751A1 - 信息处理装置及非易失性存储介质 - Google Patents

信息处理装置及非易失性存储介质 Download PDF

Info

Publication number
WO2021109751A1
WO2021109751A1 PCT/CN2020/123669 CN2020123669W WO2021109751A1 WO 2021109751 A1 WO2021109751 A1 WO 2021109751A1 CN 2020123669 W CN2020123669 W CN 2020123669W WO 2021109751 A1 WO2021109751 A1 WO 2021109751A1
Authority
WO
WIPO (PCT)
Prior art keywords
score
unit
information processing
trigger word
processing device
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/123669
Other languages
English (en)
French (fr)
Inventor
千葉俊一
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hisense Visual Technology Co Ltd
TVS Regza Corp
Original Assignee
Hisense Visual Technology Co Ltd
TVS Regza Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hisense Visual Technology Co Ltd, TVS Regza Corp filed Critical Hisense Visual Technology Co Ltd
Priority to CN202080005757.3A priority Critical patent/CN113228170B/zh
Publication of WO2021109751A1 publication Critical patent/WO2021109751A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/10Speech classification or search using distance or distortion measures between unknown speech and reference templates
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/28Constructional details of speech recognition systems
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/78Detection of presence or absence of voice signals

Definitions

  • This application relates to an information processing device and a non-volatile storage medium.
  • a user can operate the equipment through voice.
  • This kind of device starts the voice recognition service when it detects the trigger word (Trigger word) sent by the user.
  • Trigger word the trigger word
  • Patent Document 1 JP 2012-008554 A
  • the detection accuracy of the trigger word may be reduced depending on the user's speaking style and the surrounding environment. Since various factors have been taken into account to reduce the detection accuracy, sometimes the user cannot determine the cause of the undetected trigger word.
  • the problem to be solved by this application is to provide an information processing device and a non-volatile storage medium that can assist a user's judgment in an attempt to detect a trigger word.
  • the information processing device of the embodiment of the present application includes: an acquisition unit that acquires the user's voice input to the voice input unit as a sound signal; and a score calculation unit that calculates the score of the sound signal with respect to sound data, where The score becomes a reference for detecting a trigger word from the sound signal, the trigger word is used to start a voice recognition service; and a display control unit that displays the score on a display unit.
  • FIG. 1 is a diagram showing an example of the structure of a voice recognition system according to an embodiment
  • FIG. 2 is a diagram showing an example of the hardware configuration of the television device according to the embodiment
  • FIG. 3 is a diagram showing an example of the functional structure of the television device according to the embodiment.
  • FIG. 4 is a diagram showing an example of a score display screen displayed on the television device according to the embodiment.
  • FIG. 5 is a diagram showing several examples of a method of calculating a score based on the television device according to the embodiment
  • FIG. 6 is a flowchart showing an example of the procedure of trigger word detection processing of the television device according to the embodiment
  • FIG. 7 is a diagram showing an example of a score display screen displayed by the television device according to Modification 1 of the embodiment.
  • FIG. 8 is a diagram showing an example of a functional configuration of a television device according to Modification 2 of the embodiment.
  • FIG. 9 is a diagram showing an example of a score display screen displayed by a television device according to Modification 2 of the embodiment.
  • FIG. 10 is a diagram showing another example of the score display screen displayed by the television device according to Modification 2 of the embodiment.
  • FIG. 11 is a diagram showing an example of a score display screen displayed by a television device according to Modification 3 of the embodiment.
  • 1...Voice recognition system 10, 30... TV device, 11... Input receiving unit, 12... Test function setting unit, 13... Trigger word detection unit, 14... Score calculation unit, 15, 35... Display control unit, 16... Application execution unit, 17...device control unit, 18...communication unit, 19...storage unit, 19a...sound dictionary, 20...sound recognition server, 31...volume judging unit, 40...network.
  • FIG. 1 is a diagram showing an example of the structure of a voice recognition system 1 according to the embodiment.
  • the voice recognition system 1 includes a television device 10 and a voice recognition server 20, and provides voice recognition services to users of the television device 10, for example. With the voice recognition service, the user can perform operations of the television device 10, for example, by voice.
  • the television device 10 and the voice recognition server 20 are connected to each other wirelessly or by wire via a network 40 such as the Internet, for example.
  • the network 40 may also be, for example, a home network based on DLNA (Digital Living Network Alliance, Digital Living Network Alliance) (registered trademark), an intra-home LAN (Local Area Net work, local area network), and the like.
  • the television device 10 as an information processing device can receive a broadcast signal from a broadcasting station and receive various programs, for example.
  • the television device 10 has a voice recognition function, and starts to perform a voice recognition service when a trigger word uttered by a user is detected.
  • the trigger word is a predetermined voice command that becomes a trigger (Trigger) to start the voice recognition service.
  • the voice recognition function of the television device 10 is specifically used to detect the trigger word.
  • the television device 10 uses, for example, the voice recognition function of the voice recognition server 20 to provide the voice recognition service to the user. In this way, the television device 10 functions as a communication device that communicates with the voice recognition server 20.
  • the voice recognition server 20 is configured as, for example, a cloud server or the like placed on the cloud.
  • the voice recognition server 20 may also be configured as a physical structure including a CPU (Central Processing Unit), a ROM (Read Only Memory), and a RAM (Random Access Memory, random access memory).
  • the CPU constituting the cloud server or the computer executes, for example, a program stored in a ROM or the like, thereby realizing functions such as a voice recognition function of the voice recognition server 20.
  • the voice recognition server 20 includes a voice recognition unit 21, a processing unit 22, a communication unit 23, and a storage unit 24.
  • the voice recognition unit 21 analyzes and recognizes a voice signal based on the user's utterance transmitted from the television device 10 via the communication unit 23 and the like. At this time, the voice recognition unit 21 refers to the voice dictionary 24a of the storage unit 24.
  • the processing unit 22 performs various processing based on the recognition result of the sound signal. For example, when the audio signal is an instruction to operate the television device 10, the processing unit 22 transmits the instruction content to the television device 10 via the communication unit 23. For another example, when the audio signal is an instruction to acquire information from the Internet, the processing unit 22 searches for information on the Internet, and transmits the search result to the television device 10 via the communication unit 23. For another example, when the audio signal is a request for dialogue, the processing unit 22 may also transmit the content of the answer to the television device 10 via the communication unit 23.
  • the communication unit 23 performs communication with the television device 10. For example, the communication unit 23 receives a user's voice signal from the television device 10. For another example, the communication unit 23 transmits the processing result based on the processing unit 22 to the television device 10.
  • the storage unit 24 stores various parameters, information, and the like required for the realization of the above-mentioned functions of the voice recognition server 20.
  • the storage unit 24 includes a sound dictionary 24a, and data for analyzing a sound signal from the user is stored in the sound dictionary 24a.
  • the television device 10 also has a voice dictionary for voice recognition.
  • the storage unit 24 of the voice recognition server 20 is configured as a large-capacity storage device, and more detailed and diverse data is stored in the voice dictionary 24a included in the storage unit 24.
  • the voice recognition server 20 with high processing capability take on the main part of the functions related to the voice recognition service, the recognition accuracy and recognition speed of the voice signal from the user can be improved, and a more substantial voice recognition service can be provided.
  • FIG. 2 is a diagram showing an example of the hardware configuration of the television device 10 according to the embodiment.
  • the television device 10 includes an antenna 101, input terminals 102a to 102c, a tuner 103, a demodulator 104, a demultiplexer 105, an A/D (analog/digital) converter 106, a selector 107, Signal processing unit 108, speaker 109, display panel 110, operation unit 111, light receiving unit 112, IP communication unit 113, CPU 114, memory 115, storage 116, microphone 117, and audio I/ F (interface) 118.
  • A/D analog/digital converter
  • the antenna 101 receives broadcasting signals of digital broadcasting, and supplies the received broadcasting signals to the tuner 103 through the input terminal 102a.
  • the tuner 103 selects a broadcast signal of a desired channel from the broadcast signals supplied from the antenna 101, and supplies the selected broadcast signal to the demodulator 104.
  • the demodulator 104 demodulates the broadcast signal supplied from the tuner 103, and supplies the demodulated broadcast signal to the demultiplexer 105.
  • the demultiplexer 105 separates the broadcast signal supplied from the demodulator 104 and generates an image signal and a sound signal, and supplies the generated image signal and sound signal to the selector 107.
  • the selector 107 selects one of a plurality of signals supplied from the demultiplexer 105, the A/D converter 106, and the input terminal 102c, and supplies the selected one signal to the signal processing unit 108.
  • the signal processing unit 108 performs predetermined signal processing on the image signal supplied from the selector 107, and supplies the processed image signal to the display panel 110. In addition, the signal processing unit 108 performs predetermined signal processing on the audio signal supplied from the selector 107, and supplies the processed audio signal to the speaker 109.
  • the speaker 109 outputs voice or various sounds based on the sound signal supplied from the signal processing unit 108. In addition, the speaker 109 changes the volume of the output voice or various sounds based on the control of the CPU 114.
  • the display panel 110 as a display unit displays images such as still images and moving images, other images, text information, and the like based on the image signal supplied from the signal processing unit 108 or the control of the CPU 114.
  • the input terminal 102b receives analog signals such as image signals and audio signals input from the outside.
  • the input terminal 102c receives digital signals such as image signals and audio signals input from the outside.
  • the input terminal 102c may input a digital signal from a recorder (Recorder) equipped with a drive device that drives a storage medium for recording and playback such as BD (Blu-ray Disc) (registered trademark) to perform recording and playback.
  • BD Blu-ray Disc
  • the A/D converter 106 supplies a digital signal to the selector 107, and the digital signal is a signal generated by performing A/D conversion on the analog signal supplied from the input terminal 102b.
  • the operation unit 111 receives the user's operation input.
  • the light receiving unit 112 receives infrared rays from the remote controller 119.
  • the IP communication unit 113 is a communication interface for performing IP (Internet Protocol) communication through the network 40.
  • the CPU 114 as the control unit controls the entire television device 10.
  • the memory 115 is a ROM that stores various computer programs executed by the CPU 114, a RAM that provides a work area for the CPU 114, and the like.
  • a voice recognition program for the television device 10 to detect trigger words, an application program for providing voice recognition services, and the like are stored in the ROM.
  • the storage 116 is a HDD (Hard Disk Drive) or an SSD (Solid State Drive) or the like.
  • the memory 116 stores the signal selected by the selector 107 as recording data, for example.
  • the microphone 117 as a voice input unit acquires the voice of the user's speech and sends it to the audio I/F 118.
  • the audio I/F 118 performs analog/digital conversion on the sound acquired by the microphone 117, and sends it to the CPU 114 as a sound signal. It should be noted that in this case, the digital "sound signal" converted by the audio I/F 118 is sometimes referred to as "sound" for short.
  • FIG. 3 is a diagram showing an example of the functional configuration of the television device 10 according to the embodiment.
  • the aforementioned CPU 114 executes a program stored in, for example, a ROM or the like, thereby realizing the voice recognition function of the television device 10, etc.
  • the program executed in the television device 10 has a module structure including each functional unit described below.
  • the television device 10 as a functional unit for realizing the functions of the television device 10, includes an input receiving unit 11, a test function setting unit 12, a trigger word detection unit 13, a score calculation unit 14, and a display
  • the control unit 15, the application execution unit 16, the device control unit 17, the communication unit 18, and the storage unit 19.
  • the input receiving unit 11 as an acquisition unit receives various inputs from the user. For example, the input receiving unit 11 acquires the user's voice input to the microphone 117 through the audio I/F 118. In addition, for example, the input receiving unit 11 acquires various instructions based on the operation input from the operation unit 111 or the remote controller 119.
  • the test function setting unit 12 sets the test function to become effective. In a state where the test function is enabled, as described later, a score for the sound signal from the user is calculated, and the score is displayed on the display panel 110 of the television device 10.
  • the trigger word detection unit 13 performs acoustic processing such as noise cancellation processing on the obtained user's voice signal. Then, the trigger word detection unit 13 refers to the sound dictionary 19a of the storage unit 19, and detects a trigger word from the sound signal subjected to the sound processing. At this time, the trigger word detection unit 13 calculates the degree of coincidence between the sound data stored in the sound dictionary 19a and used as a reference for trigger word detection and the user's sound signal. Then, when the degree of coincidence between the sound data and the sound signal is equal to or greater than the predetermined value, the trigger word detection unit 13 recognizes that the sound signal includes the trigger word, and determines that the trigger word is detected. When the degree of coincidence between the sound data and the sound signal is less than a predetermined value, the trigger word detection unit 13 recognizes that the acquired sound signal is not a trigger word, and determines that the trigger word is not detected.
  • acoustic processing such as noise cancellation processing
  • the score calculation unit 14 calculates a score for the user's voice signal with respect to the voice data serving as a criterion for trigger word detection. More specifically, the score calculation unit 14 normalizes the degree of coincidence between the calculated sound data and the sound signal and calculates the score. Therefore, when the score is high, the degree of coincidence between the sound data and the sound signal is high, and the pass score becomes a predetermined value or more, which indicates that the trigger word detection unit 13 recognizes that the sound signal represents a trigger word.
  • the display control unit 15 controls various displays on the display panel 110. For example, when the input receiving unit 11 acquires a user's operation input to the remote control 119 or the like, the operation screen corresponding to the operation is displayed on the display panel 110. For another example, when the test function is valid, the display control unit 15 displays the calculated score on the display panel 110. For another example, the display control unit 15 displays a message or an icon in response to the voice on the display panel 110 when the voice recognition service is started by the detection of the trigger word. The message or icon in response to the voice may be, for example, content that urges the user to speak, or may be the display of the recognition result of the user's voice as text data.
  • the application execution unit 16 starts the voice recognition service when the trigger word is detected from the voice signal. More specifically, the application execution unit 16 activates the voice recognition service providing application when the trigger word is detected from the voice signal.
  • the voice recognition service providing application is a user interface for performing information exchange between the voice recognition server 20 and the user. That is, the voice recognition service providing application realizes the communication between the television device 10 and the voice recognition server 20 via the communication unit 18. Furthermore, the voice recognition service providing application transmits the user's voice signal to the voice recognition server 20, and receives a response to the content indicated by the voice signal from the voice recognition server 20.
  • the device control unit 17 controls various parts of the television device 10. For example, after detecting the trigger word, the device control unit 17 controls the speaker 109 to lower the volume. This is to reduce the situation where the user's voice input after the trigger word is disturbed by the voice of the content. For another example, the device control unit 17 controls various parts of the television device 10 based on the commands contained in the user's voice during the process of providing the voice recognition service.
  • the communication unit 18 controls communication with external devices and the like via the network 40.
  • the communication unit 18 controls the communication between the voice recognition server 20 and the television device 10 based on the voice recognition service providing application.
  • the storage unit 19 stores various parameters, information, and the like necessary for the realization of the above-described functions of the television device 10.
  • the storage unit 19 includes a voice dictionary 19a that stores voice data that serves as a reference for detecting trigger words in a voice signal from the user.
  • the sound data contains information on various elements such as phonemes and features included in the trigger word.
  • the trigger word detection unit 13 compares the sound data with the sound signal from the user, and becomes an index for identifying whether the sound signal contains the trigger word.
  • the sound data stored in the sound dictionary 19a may be plural.
  • the plurality of voice data may include various voice data based on gender and age, such as for men, women, and children.
  • FIG. 4 is a diagram showing an example of a score display screen 110a displayed on the television device 10 according to the embodiment.
  • the score display screen 110a is displayed on the display panel 110 when the user validates the test function.
  • the user can, for example, operate the remote control 119 or the like to input an instruction to start the test function.
  • the test function setting unit 12 performs settings to validate the test function.
  • the display control unit 15 displays the score display screen 110a on the display panel 110.
  • the score display screen 110a first displays a message urging the user to speak the trigger word.
  • the trigger word is "nie ie, tie rie bi" (Japanese pronunciation, corresponding to "hey, TV” in Chinese)
  • a message such as "please say nie rie, tie rie bi” is displayed. .
  • the score display screen 110a may also display a message indicating the threshold value of the score, which is a threshold value for detecting the sound made by the user as a trigger word.
  • a message such as "If the score is 50 or more, start voice recognition service.” etc. is displayed.
  • the score display screen 110a may also display the volume setting of the television device 10 at this time. Since the volume emitted by the television device 10 may hinder the detection of the trigger word, by displaying the volume setting, the user's attention can be drawn.
  • the score display screen 110a According to the message on the score display screen 110a, when the user says "nie, tie rie bi", etc., the sound is captured by the microphone 117, converted into a sound signal by the audio I/F 118, and the input receiving unit 11 receives the sound signal . Then, when the trigger word detection unit 13 calculates the degree of agreement between the sound data stored in the sound dictionary 19a stored in the storage unit 19 and the sound signal received by the input receiving unit 11 and subjected to acoustic processing, the score calculation unit 14 passes the agreement The degree is standardized to a numerical value of, for example, 0 to 100, and the score is calculated.
  • the display control unit 15 displays the calculated score on the score display screen 110a in the form of, for example, 0-100 bars.
  • the degree of coincidence between the sound data and the sound signal is insufficient and the score is less than the threshold
  • improving the smoothness of the tongue may be effective
  • speaking slower may be effective
  • speaking louder may be effective.
  • the user can try various speaking methods in order to obtain a higher score while referring to the score displayed on the score display screen 110a.
  • the remote control 119 or the like may lower the volume of the television device 10.
  • the display control unit 15 may display, for example, the maximum value of the previously acquired score on the score display screen 110a.
  • the trigger word detection unit 13 calculates the degree of agreement between the sound data and the sound signal
  • the sound data and the sound signal are decomposed into a plurality of elements of the trigger word, and on this basis, the degree of agreement is obtained for each of these elements.
  • the score calculation unit 14 calculates a score to be displayed on the score display screen 110a based on the plurality of degrees of coincidence. Various methods can be considered for the calculation of the score.
  • FIG. 5 is a diagram showing examples of several score calculation methods of the television device 10 according to the embodiment.
  • the sound data and the sound signal are decomposed into a plurality of phonemes 1 to 5, and the coincidence degree and score are calculated.
  • the sound data and the sound signal may include not only phoneme 1 to phoneme 5, but also information related to other elements such as characteristics and intonation, and the degree of agreement and score can also be calculated for these elements.
  • the trigger word detection unit 13 obtains, for example, the appearance probability X in a sound signal of a plurality of phonemes 1 to 5.
  • These appearance probabilities X are numerical values obtained by comparing the sound signal with the sound data, and correspond to the degree of coincidence between the sound signal and the sound data described above.
  • the appearance probability X is displayed, for example, from 0 to 1.00.
  • the score calculation unit 14 calculates the calculation result Y of the score standardized to these appearance probabilities X. At this time, the score calculation unit 14 normalizes the appearance probability X using the following equations (1) and (2), for example.
  • the following formula (1) is applied when, for example, the degree of coincidence Xn such as the appearance probability X is smaller than the threshold value Tn.
  • the following formula (2) is applied when, for example, the degree of coincidence Xn such as the appearance probability X exceeds the threshold value Tn.
  • a numerical value in the range of 0 to 100 is obtained as a calculation result Yn obtained by normalizing the degree of coincidence Xn. It should be noted that when the degree of coincidence Xn is the same value as the threshold Tn, the calculation result Yn is the same regardless of which of the equations (1) and (2) is used.
  • the sound signal and the sound data include L elements, and for the L number of coincidence degrees Xn, the maximum value An that the degree of coincidence Xn can take and the threshold value Tn that the degree of coincidence Xn should satisfy are respectively set. That is, when the degree of coincidence Xn of a certain element is greater than or equal to the threshold value Tn, it is determined that the sound signal matches the sound data with respect to the element. Then, the matching degree Xn and the threshold value Tn of the elements 1 to L are appropriately substituted into the above-mentioned formula (1) or formula (2), and L calculation results Yn are obtained.
  • the example on the right of Figure 5 (a) and (b) is set as follows: the threshold T for all occurrence probabilities X is 0.90, and the maximum value A that all occurrence probabilities X can take is 1.00, which results in Calculate the result Y.
  • the score calculation unit 14 obtains the score displayed on the score display screen 110a based on these calculation results Y. As mentioned above, there are several methods for this.
  • the score calculation unit 14 adopts the calculation result 30 of the phoneme 5, which is the smallest value among the calculation results Y obtained for the phonemes 1 to 5, as the score displayed on the score display screen 110a.
  • the score calculation unit 14 compares the calculation result Y of phoneme 1 to phoneme 5 that exceeds 50, the calculation result 75 of phoneme 1, and the calculation result 60 of phoneme 3, as shown in FIG. ) As shown in the lower right, the part exceeding 50 is removed as the mantissa and taken as the calculation result 50. On this basis, the average value 44 of the calculation results Y for phoneme 1 to phoneme 5 is adopted as the score displayed on the score display screen 110a.
  • the method of obtaining the score by the score calculation unit 14 is not limited to the examples in (a) and (b) of FIG. 5.
  • the user can directly grasp the difference between the score required in the detection of the trigger word and its own score, as long as it is a score that can be used as an index for obtaining a higher score, any method can be used for calculation.
  • FIG. 6 is a flowchart showing an example of the procedure of the trigger word detection process of the television device 10 according to the embodiment.
  • the input receiving unit 11 receives an instruction to use the test function by the user (step S101). That is, when the user operates the operation unit 111 or the remote controller 119 to instruct to start the test function, the input unit 11 receives the instruction (step S101: Yes), the test function setting unit 12 validates the setting of the test function, and displays the control The section 15 displays the score display screen 110a on the display panel 110 (step S102). If there is no instruction to start the test function by the user (step S101: No), the process of step S102 is not performed and the process proceeds to step S103.
  • the input receiving unit 11 receives a voice signal based on the user's utterance (step S103).
  • the input receiving unit 11 waits until the user speaks (step S103: No).
  • the user speaks toward the microphone 117 of the television device 10 the sound acquired from the microphone 117 is converted into a sound signal through the audio I/F 118.
  • the trigger word detection unit 13 refers to the sound dictionary 19a, and calculates the degree of agreement between the sound data stored in the sound dictionary 19a and the sound signal based on the user's utterance (step S104 ).
  • the score calculation unit 14 confirms whether the setting of the test function is valid (step S105). When the setting of the test function is valid (step S105: Yes), the score calculation unit 14 calculates a score based on the calculated degree of coincidence (step S106). In addition, the display control unit 15 displays the calculated score on the score display screen 110a of the display panel 110 (step S107). When the setting of the test function is not set to be valid (step S105: No), the process of steps S106 to S107 is not performed, and the process proceeds to step S108.
  • the trigger word detection unit 13 determines whether or not the degree of coincidence of all the elements of the sound data and the sound signal is equal to or greater than a threshold value (step S108). When there is an element that the degree of coincidence between the sound data and the sound signal is less than the threshold (step S108: No), the trigger word detection unit 13 determines that the sound signal is not a trigger word and does not perform the trigger word detection processing, and repeats the execution from step S103 Processing.
  • the trigger word detection unit 13 determines that the sound signal includes the trigger word and detects the trigger word (step S109).
  • the application execution unit 17 activates the voice recognition service providing application to start the voice recognition service (step S110).
  • the trigger word detection process of the television device 10 of the embodiment is ended.
  • the user repeatedly tries various measures such as increasing the voice or speaking slowly, so that the television device detects the trigger word.
  • various measures such as increasing the voice or speaking slowly
  • the television device detects the trigger word.
  • the user can only judge which of such trial and error is effective through the start of the provision of the voice recognition service.
  • the score of the sound signal with respect to the sound data is calculated, and the score is displayed on the display panel 110.
  • the user repeats the trial while referring to the change in the score, and can easily recognize the directionality of his own voice as a trigger word that is easily detected.
  • the television device 10 of the embodiment can assist the judgment of the user who is trying to detect the trigger word.
  • the degree of coincidence between the audio data and the audio signal is normalized to calculate the score.
  • the trigger word detection unit 13 calculates the degree of agreement between the sound data and the sound signal.
  • this degree of agreement is calculated based on various elements of various content. Therefore, even if, for example, the calculated degree of agreement is presented to the user as it is, it is difficult for the user to understand the content, and it is difficult for the user to grasp whether his own attempt is close to the detection of the trigger word. Since the television device 10 standardizes this degree of coincidence and presents it to the user, the user can intuitively understand the content and use it as an index for obtaining a higher score.
  • the television device of Modification 1 is different from the above-described embodiment in that the calculated score is displayed for each phoneme.
  • FIG. 7 is a diagram showing an example of a score display screen 110 b displayed on the television device according to Modification 1 of the embodiment.
  • the display control unit included in the television device of Modification 1 displays the score of the sound signal calculated by the score calculation unit for each phoneme included in the sound data on the score display screen 110b.
  • the user can recognize the weakness of his own speech.
  • the phonemes judged to be "ie" and "bi" have low scores.
  • the user can increase the score and detect his own voice as a trigger word.
  • Modification 2 is different from the above-described embodiment in that the television device 30 simultaneously displays the calculated score and the advice to the user.
  • FIG. 8 is a diagram showing an example of a functional configuration of a television device 30 according to Modification 2 of the embodiment.
  • the television device 30 of Modification 2 includes a display control unit 35 and a volume determination unit 31 instead of the configuration of the television device 10 of the above-described embodiment.
  • the volume determination unit 31 determines whether the volume setting of the speaker of the television device 30 exceeds a predetermined value.
  • the display control unit 35 displays the calculated score and displays a message prompting the user to lower the volume setting.
  • FIG. 9 is a diagram showing an example of the score display screen 110c displayed by the television device 30 according to Modification 2 of the embodiment. As shown in Figure 9, the score display screen 110c displays a message such as "The sound of the TV seems to be too loud. Please try to set the volume below 10.” and so on.
  • One of the most accurate and biggest factors that make it difficult to detect trigger words is the sound emitted by the speakers of the television device.
  • the user can notice that the volume of the television device 30 may reduce the detection accuracy, and the trigger word is easy to be detected.
  • the display control unit 35 included in the television device 30 of Modification 2 may also display suggestions for increasing the score to facilitate the detection of trigger words randomly or in a predetermined order.
  • FIG. 10 is a diagram showing another example of the score display screen 110 d displayed by the television device 30 according to Modification 2 of the embodiment. As shown in Figure 10, the score display screen 110d scrolls and displays "please try to speak clearly.”, “please try to speak slowly.”, “please try to speak loudly.” and other general factors that eliminate trigger words that cannot be detected. News.
  • the user may be prompted to try unexpectedly, which helps the user's voice be detected as a trigger word.
  • FIG. 11 is a diagram showing an example of the score display screen 110e displayed by the television device according to Modification 3 of the embodiment.
  • the television device of Modification 3 is set with "nie ie, tie rie bi" (Japanese pronunciation, corresponding to Chinese “Hey, TV”), “mo si mo si, tie rie bi” (Japanese pronunciation, corresponding to Chinese “Hello, TV”), "ha ro,tie rie bi” (Japanese pronunciation, corresponding to Chinese “hello, TV”) and other trigger words.
  • the score calculation unit of the television device of Modification 3 calculates scores for each of these trigger words.
  • the display control unit displays the scores for the plurality of trigger words on the score display screen 110e.
  • the user can display the message on the screen 110e according to the score of the trigger word such as "please say ‘nie, tie rie bi’.”, etc., according to the prompt, for example, speak each trigger word, and refer to the score corresponding to these trigger words.
  • the user gets the highest score among the trigger words "mo si mo si, tie rie bi”. Therefore, the user can easily detect his own voice as the trigger word by selecting and using the trigger word "mo si mo si, tie rie bi" among the multiple trigger words.
  • the voice recognition server 20 as an external device such as the television device 10 provides the main voice recognition service, but the configuration of the embodiment is not limited to this. It is also possible that the television device 10 itself has all functions related to the voice recognition service, and independently provides the voice recognition service.
  • the information processing device having the voice recognition function is the television device 10 or the like, but the configuration of the embodiment is not limited to this.
  • an information processing device or a communication device with a voice recognition function may also be other devices such as a smart speaker (Smart Speaker).
  • the display unit that displays the score of the sound signal for the sound data may be a separate monitor or the like installed in the smart speaker.
  • a program that realizes the various functions described above in the television device 10 or the like is provided as a computer program product in an installable form or an executable form. That is, the above-mentioned program is provided in a state of being included in a computer program product having a non-volatile computer-readable recording medium such as a CD-ROM, a floppy disk (FD), a CD-R, and a DVD.
  • a non-volatile computer-readable recording medium such as a CD-ROM, a floppy disk (FD), a CD-R, and a DVD.
  • the above-mentioned program may be provided or distributed through a network in a state of being stored in a computer connected to a network such as the Internet.
  • the above-mentioned program may also be provided in a state of being installed in a ROM or the like in advance.
  • the CPU of the television device 10 or the like By installing such a program in the television device 10 or the like, the CPU of the television device 10 or the like reads the program from the ROM and executes the above-mentioned respective functional structures on the RAM.
  • the above-mentioned program may be provided as a network application stored in a cloud server or the like, in which case the program can be executed without being installed in the television device 10 or the like.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

一种信息处理装置及非易失性存储介质,其对为了检测触发词而尝试的用户的判断进行辅助。所述信息处理装置具备:获取部,其将输入到声音输入部的用户的声音作为声音信号进行获取;得分计算部,其计算声音信号相对于声音数据的得分,所述得分成为用于从声音信号中检测触发词的基准,所述触发词用于使声音识别服务开始;以及显示控制部,其将得分显示在显示部上。

Description

信息处理装置及非易失性存储介质
本申请要求在2019年12月5日提交日本专利局、申请号为2019-220035、发明名称为“信息处理装置及程序”的日本专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及信息处理装置及非易失性存储介质。
背景技术
在具备声音识别功能的电视装置等设备中,例如用户可以通过声音进行设备的操作。这种设备当检测出用户发出的触发词(Trigger word)时启动声音识别服务。
在先技术文献
专利文献
专利文献1:日本特开2012-008554号公报
发明内容
然而,触发词的检测精度根据用户的说话方式和周围的环境等可能会降低。由于考虑到各种因素导致检测精度的降低,因此有时用户无法判断未检测触发词的原因是什么。
本申请所要解决的课题是提供一种能够对为了检测触发词而尝试的用户的判断进行辅助的信息处理装置和非易失性存储介质。
本申请实施方式的信息处理装置具备:获取部,其将输入到声音输入部的用户的声音作为声音信号进行获取;得分计算部,其计算所述声音信号相对于声音数据的得分,其中,所述得分成为用于从所述声音信号中检测触发词的基准,所述触发词用于使声音识别服务开始;以及显示控制部,其将所 述得分显示在显示部上。
附图说明
图1是表示实施方式涉及的声音识别系统的结构的一例的图;
图2是表示实施方式涉及的电视装置的硬件结构的一例的图;
图3是表示实施方式涉及的电视装置的功能结构的一例的图;
图4是表示实施方式涉及的电视装置显示的得分显示画面的一例的图;
图5是表示基于实施方式涉及的电视装置的得分计算方法的几个例子的图;
图6是表示实施方式涉及的电视装置的触发词检测处理的顺序的一例的流程图;
图7是表示实施方式的变形例1涉及的电视装置显示的得分显示画面的一例的图;
图8是表示实施方式的变形例2的电视装置的功能结构的一例的图;
图9是表示实施方式的变形例2涉及的电视装置显示的得分显示画面的一例的图;
图10是表示实施方式的变形例2涉及的电视装置显示的得分显示画面的其它例的图;
图11是表示实施方式的变形例3涉及的电视装置显示的得分显示画面的一例的图。
附图标记说明
1…声音识别系统,10、30…电视装置,11…输入接收部,12…测试功能设定部,13…触发词检测部,14…得分计算部,15、35…显示控制部,16…应用执行部,17…设备控制部,18…通信部,19…存储部,19a…声音词典,20…声音识别服务器,31…音量判断部,40…网络。
具体实施方式
(声音识别系统的结构)
图1是显示实施方式涉及的声音识别系统1的结构的一例的图。如图1所示,声音识别系统1具备电视装置10和声音识别服务器20,例如向电视装置10的用户提供声音识别服务。通过声音识别服务,用户可以例如通过声音进行电视装置10的操作。
电视装置10与声音识别服务器20例如经由互联网等网络40,通过无线或有线相互连接。网络40还可以是例如基于DLNA(Digital Living Network Alliance,数字生活网络联盟)(注册商标)的家庭网络、家庭内LAN(Local Area Net work,局域网)等。
作为信息处理装置的电视装置10例如可以接收来自广播电台的广播信号而接收各种节目。此外,电视装置10具有声音识别功能,当检测到用户发出的触发词时开始进行声音识别服务。触发词是成为声音识别服务开始的触发(Trigger)的规定的声音指令。电视装置10的声音识别功能专门用于检测该触发词。声音识别服务开始后,电视装置10利用例如声音识别服务器20的声音识别功能,将声音识别服务提供给用户。如此,电视装置10作为与声音识别服务器20进行通信的通信装置发挥作用。
声音识别服务器20构成为例如置于云端上的云服务器等。然而,声音识别服务器20还可以构成为具备CPU(Central Processing Unit,中央处理器)、ROM(Read Only Memory,只读存储器)、以及RAM(Random Access Memory,随机存取存储器)等物理结构的一个以上的计算机。通过构成云服务器或计算机的CPU执行例如存储于ROM等中的程序,从而实现声音识别服务器20的声音识别功能等功能。
作为用于实现声音识别功能等的功能部,声音识别服务器20具备声音识别部21、处理部22、通信部23、以及存储部24。
声音识别部21分析并且识别从电视装置10经由通信部23发送过来的基于用户说话的声音信号等。此时,声音识别部21参照存储部24的声音词典 24a。
处理部22基于声音信号的识别结果进行各种处理。例如,在声音信号为指示电视装置10的操作的情况下,处理部22通过通信部23将指示内容发送到电视装置10。再例如,在声音信号为指示获取来自互联网的信息的情况下,处理部22在互联网上检索信息,经由通信部23将检索结果发送到电视装置10。再例如,声音信号为请求对话的情况下,处理部22还可以经由通信部23将回答的内容发送到电视装置10。
通信部23进行与电视装置10的通信。例如,通信部23从电视装置10接收用户的声音信号。再例如,通信部23将基于处理部22的处理结果发送到电视装置10。
存储部24存储声音识别服务器20的如上所述功能的实现中需要的各种参数和信息等。作为一例,存储部24具备声音词典24a,在该声音词典24a中存储有用于解析来自用户的声音信号的数据。如后所述,电视装置10同样具有用于声音识别的声音词典。但是,声音识别服务器20的存储部24构成为大容量存储装置,存储部24具有的声音词典24a中存储有更详细且多样化的数据。
如此,通过使处理能力高的声音识别服务器20承担声音识别服务相关的功能的主要部分,从而提高来自用户的声音信号的识别精度和识别速度,并且能够提供更充实的内容的声音识别服务。
(电视装置的硬件结构)
图2是表示实施方式涉及的电视装置10的硬件结构的一例的图。
如图2所示,电视装置10具备天线101、输入端子102a~102c、调谐器103、解调器104、解复用器105、A/D(模拟/数字)转换器106、选择器107、信号处理部108、扬声器109、显示面板110、操作部111、受光部112、IP通信部113、CPU 114、内存(memory)115、存储器(storage)116、麦克风117、以及音频(Audio)I/F(接口)118。
天线101接收数字广播的广播信号,将接收的广播信号通过输入端子102a 供给到调谐器103。
调谐器103从由天线101供给的广播信号中对期望的频道的广播信号进行选台,将选台的广播信号供给到解调器104。
解调器104解调从调谐器103供给的广播信号,将解调的广播信号供给到解复用器105。
解复用器105分离从解调器104供给的广播信号并生成图像信号和声音信号,将生成的图像信号和声音信号供给到选择器107。
选择器107从由解复用器105、A/D转换器106、以及输入端子102c供给的多个信号中选择一个,将选择出的一个信号供给到信号处理部108。
信号处理部108对从选择器107供给的图像信号实施规定的信号处理,将处理后的图像信号供给到显示面板110。此外,信号处理部108对从选择器107供给的声音信号实施规定的信号处理,将处理后的声音信号供给到扬声器109。
扬声器109基于从信号处理部108供给的声音信号而输出语音或各种声音。此外,扬声器109基于CPU 114的控制,变更输出的语音或各种声音的音量。
作为显示部的显示面板110基于从信号处理部108供给的图像信号或CPU 114的控制,显示静态图像以及动态图像等图像、其它图像、以及文字信息等。
输入端子102b接收从外部输入的图像信号以及声音信号等模拟信号。此外,输入端子102c接收从外部输入的图像信号以及声音信号等数字信号。例如,输入端子102c可以从搭载了驱动装置的录像机(Recorder)等输入数字信号,该驱动装置驱动BD(Blu-ray Disc)(注册商标)等录像播放用的存储介质而进行录像及播放。
A/D转换器106将数字信号供给到选择器107,该数字信号是通过对从输入端子102b供给的模拟信号实施A/D转换而生成的信号。
操作部111接收用户的操作输入。
受光部112从遥控器119接收红外线。
IP通信部113是用于进行通过网络40的IP(互联网协议)通信的通信接口。
作为控制部的CPU 114控制电视装置10整体。
内存115是存放CPU 114执行的各种计算机程序的ROM、以及为CPU 114提供工作分区(area)的RAM等。例如,ROM中存储有用于电视装置10检测触发词的声音识别程序、以及用于提供声音识别服务的应用程序等。
存储器116是HDD(Hard Disk Drive,硬盘驱动器)或SSD(Solid State Drive,固态硬盘)等。存储器116例如将通过选择器107选择出的信号作为录像数据进行存储。
作为声音输入部的麦克风117获取用户说话的声音,发送到音频I/F 118。
音频I/F 118对麦克风117获取到的声音进行模拟/数字转换,作为声音信号发送到CPU 114。应该说明,如此,以下有时将通过音频I/F 118转换过的数字的“声音信号”简称为“声音”。
(电视装置的功能结构)
接着,使用图3对实施方式的电视装置10的功能结构例进行说明。图3是表示实施方式涉及的电视装置10的功能结构的一例的图。
在电视装置10中,上述的CPU 114执行存储于例如ROM等中的程序,从而实现电视装置10的声音识别功能等。在电视装置10中执行的程序成为包含以下叙述的各个功能部在内的模块儿结构。
如图3所示,电视装置10作为用于实现电视装置10的功能的功能部,具备输入接收部11、测试功能设定部12、触发词检测部13、得分(score)计算部14、显示控制部15、应用执行部16、设备控制部17、通信部18、以及存储部19。
作为获取部的输入接收部11接收来自用户的各种输入。例如,输入接收部11通过音频I/F 118获取输入到麦克风117的用户的声音。此外,例如,输入接收部11获取来自操作部111或遥控器119的基于操作输入的各种指示。
当通过来自操作部111或遥控器119的操作输入来指示测试功能的开始 时,测试功能设定部12设定为测试功能成为有效。在测试功能成为有效的状态下,如后所述,计算针对来自用户的声音信号的得分,该得分显示在电视装置10的显示面板110上。
触发词检测部13对得到的用户的声音信号实施噪音消除处理等音响处理。然后,触发词检测部13参照存储部19的声音词典19a,从实施了音响处理的声音信号中检测触发词。这时,触发词检测部13对存储于声音词典19a的、成为触发词检测的基准的声音数据、与用户的声音信号的一致度进行计算。然后,触发词检测部13在声音数据与声音信号的一致度为规定值以上的情况下,识别为声音信号包含触发词,判断为检测到触发词。触发词检测部13在声音数据与声音信号的一致度小于规定值的情况下,识别为获取到的声音信号不是触发词,判断为未检测到触发词。
得分计算部14在测试功能有效的情况下,计算针对成为触发词检测的基准的声音数据的、用户的声音信号的得分。更具体而言,得分计算部14对计算出的声音数据与声音信号的一致度进行标准化并且计算得分。从而,当得分高时,声音数据与声音信号的一致度高,并且通过得分成为规定值以上,从而表示通过触发词检测部13识别为该声音信号表示触发词。
显示控制部15控制向显示面板110的各种显示。例如,在输入接收部11获取了输入到遥控器119等的用户的操作的情况下,将与该操作对应的操作画面显示在显示面板110上。再例如,显示控制部15在测试功能有效的情况下,将计算出的得分显示在显示面板110上。再例如,显示控制部15在通过触发词的检测而声音识别服务开始时,将对声音进行应答的消息或图标等显示在显示面板110上。对声音进行应答的消息或图标等可以为例如催促用户说话的内容,还可以为将用户的声音的识别结果作为文字数据显示。
应用执行部16当从声音信号检测到触发词时开始进行声音识别服务。更具体而言,应用执行部16当从声音信号检测到触发词时,启动声音识别服务提供应用。声音识别服务提供应用是用于进行声音识别服务器20与用户的信息交换的用户接口。即,声音识别服务提供应用经由通信部18,实现电视装 置10与声音识别服务器20的通信。而且,声音识别服务提供应用将用户的声音信号发送到声音识别服务器20,并且从声音识别服务器20接收对该声音信号所表示的内容的应答。
设备控制部17控制电视装置10的各部分。例如,设备控制部17在检测触发词后,控制扬声器109而降低音量。这是为了减少用户在触发词后说话的声音的输入被内容的声音干扰的情况。再例如,设备控制部17在提供声音识别服务的过程中,基于用户的声音所包含的命令,控制电视装置10的各部分。
通信部18控制经由网络40的与外部设备等的通信。例如,通信部18根据声音识别服务提供应用,控制声音识别服务器20与电视装置10的通信。
存储部19存储电视装置10的如上所述的功能的实现所需的各种参数和信息等。作为一例,存储部19具备声音词典19a,该声音词典19a存储有成为在来自用户的声音信号中检测触发词的基准的声音数据。声音数据具有例如触发词中所包含的关于音素和特征等各种要素的信息,触发词检测部13比较该声音数据与来自用户的声音信号,从而成为用于识别声音信号是否包含触发词的指标。但是,声音词典19a中存储的声音数据可以为复数。例如,多个声音数据中可以包含男性用、女性用、以及儿童用等,基于性别和年龄的各种声音数据。
(电视装置的详细功能)
接着,使用图4和图5对实施方式的电视装置10的功能的详细内容进行说明。图4是显示实施方式涉及的电视装置10显示的得分显示画面110a的一例的图。当用户使测试功能有效时得分显示画面110a显示在显示面板110上。
用户可以例如操作遥控器119等而输入开始测试功能的指示。当输入接收部11接收开始测试功能的指示时,测试功能设定部12进行使测试功能有效的设定。在测试功能被设定为有效时,显示控制部15将得分显示画面110a显示在显示面板110上。
如图4所示,得分显示画面110a上首先显示催促用户说出触发词的消息。例如,在触发词为“nie ie,tie rie bi”(日语的读音,对应于中文的“嘿,电视”)的情况下,显示“请说出‘nie ie,tie rie bi’。”等消息。
此外,得分显示画面110a上还可以显示表示得分的阈值的消息,该得分的阈值是用于将由用户发出的声音作为触发词而检测的阈值。在阈值为例如50的情况下,显示“如果得分为50以上,则开始进行声音识别服务。”等消息。
另外,得分显示画面110a上还可以显示此时的电视装置10的音量设定等。由于电视装置10发出的音量有可能成为触发词检测的阻碍,因此通过显示音量设定,从而能够引起用户的注意。
根据得分显示画面110a的消息,当用户说出“nie ie,tie rie bi”等时,该声音通过麦克风117被获取,通过音频I/F 118转换为声音信号,输入接收部11接收该声音信号。然后,当触发词检测部13计算出存储在存储部19的声音词典19a的声音数据、与输入接收部11接收后实施了音响处理的声音信号的一致度时,得分计算部14通过将该一致度标准化为例如0~100的数值,从而计算得分。显示控制部15将计算出的得分以例如0~100的条(bar)形式显示在得分显示画面110a上。
在声音数据与声音信号的一致度不充分而得分小于阈值的情况下,为了得到更高的得分,例如改善舌头的顺滑度可能有效,说话慢一点可能有效,声音大一点可能有效。用户可以一边参照显示在得分显示画面110a上的得分,一边为了得到更高得分而尝试各种说话方法。也可以操作遥控器119等,降低电视装置10的音量。这时,显示控制部15除了用户的声音的当前得分以外,例如还可以将之前获取到的得分的最大值显示在得分显示画面110a上。
可是,触发词检测部13在计算声音数据与声音信号的一致度时,将声音数据和声音信号分解为触发词所具有的多个要素,在这基础上对这些每个要素求出一致度。得分计算部14根据这些多个一致度而计算用于显示在得分显示画面110a上的得分。得分的计算可以考虑各种方法。
图5是表示实施方式涉及的电视装置10的几个得分计算方法的例子的图。在图5的例子中,为了使说明简单,表示了声音数据和声音信号被分解为多个音素1~音素5而计算一致度和得分的情况。然而,声音数据和声音信号不仅可以包含音素1~音素5,还可以包含特征和语调等其它要素相关的信息,对这些要素也可以计算一致度和得分。
如图5中的(a)、(b)的左图所示,触发词检测部13求出例如多个音素1~音素5的声音信号中的出现概率X。这些出现概率X是通过将声音信号与声音数据比较而得到的数值,相当于上述的声音信号与声音数据的一致度。在图5的(a)、(b)的左图的例子中,出现概率X例如由0至1.00的数值显示。
如图5的(a)、(b)的右图所示,得分计算部14计算对这些出现概率X标准化的得分的计算结果Y。这时,得分计算部14例如用以下的式(1)、式(2)使出现概率X标准化。
以下的式(1)在例如出现概率X等的一致度Xn为小于阈值Tn的情况下适用。
【数学式1】
Figure PCTCN2020123669-appb-000001
以下式(2)在例如出现概率X等的一致度Xn超过阈值Tn的情况下适用。
【数学式2】
Figure PCTCN2020123669-appb-000002
根据上述式(1)、式(2),作为将一致度Xn标准化的计算结果Yn,求出0至100的范围内的数值。应该说明的是,在一致度Xn为与阈值Tn相同值的情况下,不管使用式(1)、式(2)中的哪一个,计算结果Yn都相同。
在此,设定为如下:声音信号和声音数据包含L个要素,关于L个一致度Xn,分别设定有一致度Xn可以取的最大值An以及一致度Xn应该满足的 阈值Tn。即,当某一要素的一致度Xn为阈值Tn以上时,关于该要素,判断为声音信号与声音数据一致。然后,在上述式(1)或式(2)中适当地代入1至L的要素的一致度Xn和阈值Tn,求出L个计算结果Yn。
图5的(a)、(b)的右图的例子是设定为:关于所有的出现概率X的阈值T为0.90,所有的出现概率X可以取的最大值A为1.00,由此得出计算结果Y。得分计算部14基于这些计算结果Y,得出显示在得分显示画面110a上的得分。如上所述,对此存在几种方法。
在图5(a)的例子中,得分计算部14将对音素1~音素5得到的计算结果Y中作为最小值的音素5的计算结果30作为显示在得分显示画面110a上的得分来采用。
在图5(b)的例子中,得分计算部14将对音素1~音素5得到的计算结果Y中超过50的、音素1的计算结果75和音素3的计算结果60,如图5(b)右下所示,将超过50的部分作为尾数去掉而作为计算结果50。在此基础上,将对音素1~音素5的计算结果Y的平均值44作为显示在得分显示画面110a上的得分来采用。
另外,基于得分计算部14的得分的求法并不限定于图5的(a)、(b)的例子。用户可以直接掌握触发词的检测中需要的得分与自身的得分之差,只要是可以作为用于得到更高得分的指标的得分,则可以使用任何方法进行计算。
(电视装置的触发词检测处理)
接着,使用图6对实施方式的电视装置10的触发词检测处理的例子进行说明。图6是表示实施方式涉及的电视装置10的触发词检测处理的顺序的一例的流程图。
如图6所示,输入接收部11接收基于用户的测试功能的使用指示(步骤S101)。即,当用户对操作部111或遥控器119进行操作而指示开始测试功能时,输入部11接收该指示(步骤S101:是),测试功能设定部12使测试功能的设定有效,显示控制部15在显示面板110上显示得分显示画面110a(步骤S102)。在没有基于用户的测试功能的开始指示的情况下(步骤S101:否),不 进行步骤S102的处理而进入步骤S103的处理。
输入接收部11接收基于用户说话的声音信号(步骤S103)。输入接收部11等待直至用户说话(步骤S103:否)。当用户朝向电视装置10的麦克风117说话时,从麦克风117获取到的声音通过音频I/F 118转换至声音信号。当输入接收部11获取该声音信号时(步骤S103:是),触发词检测部13参照声音词典19a,计算存储在声音词典19a中的声音数据与基于用户说话的声音信号的一致度(步骤S104)。
得分计算部14确认测试功能的设定是否有效(步骤S105)。当测试功能的设定有效时(步骤S105:是),得分计算部14基于计算出的一致度而计算得分(步骤S106)。此外,显示控制部15将计算出的得分显示在显示面板110的得分显示画面110a上(步骤S107)。当测试功能的设定未设定为有效时(步骤S105:否),不进行步骤S106~S107的处理而进入步骤S108的处理。
触发词检测部13判断关于声音数据和声音信号的全部要素的一致度是否为阈值以上(步骤S108)。当存在关于声音数据和声音信号的一致度小于阈值的要素时(步骤S108:否),触发词检测部13判断为声音信号不是触发词而不进行触发词的检测处理,反复执行从步骤S103开始的处理。
在关于声音数据和声音信号的全部一致度为阈值以上的情况下(步骤S108:是),触发词检测部13判断为声音信号包含触发词而进行触发词的检测(步骤S109)。应用执行部17启动声音识别服务提供应用来开始进行声音识别服务(步骤S110)。
通过以上步骤,结束实施方式的电视装置10的触发词检测处理。
近年来,已知具备声音识别功能的电视装置等。当检测触发词时,电视装置开始提供声音识别服务。根据用户的说话方式和周围的环境等,有时该触发词的检测精度降低。
在这种情况下,用户反复试验增大声音或缓慢说话等各种措施,以使电视装置检测触发词。但是,用户只能通过声音识别服务的提供开始来判断这种反复试验中的哪一种有效。
根据实施方式的电视装置10,计算声音信号相对于声音数据的得分,将该得分显示在显示面板110上。由此,用户一边参照得分的变化一边重复尝试,从而能够容易地认定自身的声音作为触发词容易被检测的方向性。如此,实施方式的电视装置10能够对为了检测触发词而尝试的用户的判断进行辅助。
根据实施方式的电视装置10,将声音数据与声音信号的一致度标准化而计算得分。为了检测触发词,例如触发词检测部13计算声音数据与声音信号的一致度。然而,这种一致度根据多方面的内容的各种要素来进行计算。因此,即使例如将计算出的一致度原样地提示给用户,用户也难以理解其内容,难以把握自身的尝试是否接近触发词的检测。由于电视装置10将这种一致度进行标准化而提示给用户,因此用户可以直观地理解其内容,并作为用于得到更高的得分的指标。
(变形例1)
接着,使用图7对实施方式的变形例1的电视装置进行说明。变形例1的电视装置与上述的实施方式的不同点在于,将计算出的得分按照每个音素进行显示。
图7是表示实施方式的变形例1涉及的电视装置显示的得分显示画面110b的一例的图。如图7所示,变形例1的电视装置具备的显示控制部将得分计算部按照声音数据中所包含的每个音素来计算出的声音信号的得分显示在得分显示画面110b上。
由此,用户可以识别自身的说话的弱点。例如,在图7所示的例子中,在用户的声音中,判断为“ie”和“bi”的音素的得分低。该用户例如通过关注每个单词的结尾,从而能够提高得分而将自身的声音作为触发词检测出来。
(变形例2)
接着,使用图8~图10对实施方式的变形例2的电视装置30进行说明。变形例2与上述的实施方式的不同点在于,电视装置30同时显示计算出的得分和对用户的建议。
图8是表示实施方式的变形例2的电视装置30的功能结构的一例的图。 如图8所示,变形例2的电视装置30代替上述的实施方式的电视装置10的结构而具备显示控制部35,还具备音量判断部31。
例如在测试功能的设定为有效的情况下,音量判断部31判断电视装置30的扬声器的音量设定是否超过规定值。显示控制部35在音量设定超过了规定值的情况下,显示计算出的得分,并且显示促使用户降低音量设定的消息。
图9是表示实施方式的变形例2涉及的电视装置30显示的得分显示画面110c的一例的图。如图9所示,得分显示画面110c上显示“电视的声音似乎太大。请尝试将音量设定为10以下。”等消息。
触发词难以检测的最准确且最大的因素之一是,电视装置的扬声器发出的声音。通过显示促使降低音量设定的消息,用户可以注意到电视装置30的音量可能降低检测精度,触发词容易被检测。
此外,变形例2的电视装置30具备的显示控制部35还可以将用于提高得分而容易检测触发词的建议随机地或以规定的顺序显示。
图10是表示实施方式的变形例2涉及的电视装置30显示的得分显示画面110d的其它例子的图。如图10所示,得分显示画面110d上例如滚动显示“请尝试清楚地说话。”、“请尝试缓慢地说话。”、“请尝试大声说话。”等消除触发词无法被检测的一般的因素的消息。
由此,例如可以提示用户没有想到的尝试,有助于用户的声音作为触发词被检测。
(变形例3)
接着,使用图11对实施方式的变形例3的电视装置进行说明。变形例3与上述实施方式的区别点是,电视装置对多个触发词显示得分。
图11是表示实施方式的变形例3涉及的电视装置显示的得分显示画面110e的一例的图。如图11所示,变形例3的电视装置中设定有“nie ie,tie rie bi”(日语的读音,对应于中文的“嘿,电视”)、“mo si mo si,tie rie bi”(日语的读音,对应于中文的“你好,电视”)、“ha ro,tie rie bi”(日语的读音,对应于中文的“hello,电视”)等多个触发词。然后,变形例3的 电视装置的得分计算部对这些触发词分别计算得分。显示控制部将关于多个触发词的得分显示在得分显示画面110e上。
用户可以根据促使发出“请说‘nie ie,tie rie bi’。”等规定的触发词的得分显示画面110e上的消息,例如说出各个触发词,参照与这些触发词对应的得分。在图11显示的例子中,在多个触发词中,用户在“mo si mo si,tie rie bi”的触发词中获得最高的得分。因此,该用户通过选择使用多个触发词中的“mo si mo si,tie rie bi”的触发词,可以容易地将自身的声音作为触发词检测。
另外,在上述的实施方式和变形例1~3中,作为电视装置10等外部设备的声音识别服务器20提供主要的声音识别服务,但是实施方式的结构并不限定于此。也可以是,电视装置10等本身具有与声音识别服务的全部相关的功能,并且独立地提供声音识别服务。
此外,虽然在上述的实施方式和变形例1~3中,具备声音识别功能的信息处理装置为电视装置10等,但实施方式的结构并不限定于此。例如,具备声音识别功能的信息处理装置或通信装置还可以为智能音响(Smart speaker)等其它设备。在信息处理装置为智能音响的情况下,显示对于声音数据的声音信号的得分的显示部也可以为安装于智能音响的单独的监视器等。
另外,在电视装置10等中实现上述各种功能的程序作为可安装的形式或可执行的形式的计算机程序产品提供。即,上述程序以包含在具有CD-ROM、软盘(FD)、CD-R、DVD等非易失性计算机可读记录介质的计算机程序产品中的状态被提供。
此外,上述程序可以在存储于连接到互联网等网络上的计算机中的状态下通过网络提供或分配。上述程序还可以以预先安装在ROM等中的状态下提供。
通过将这种程序安装在电视装置10等中,从而电视装置10等的CPU从ROM读取程序,并且在RAM上执行上述各个功能结构。
然而,上述程序可以作为存储在云服务器等中的网络应用被提供,在该情况下,程序无需安装在电视装置10等中就可以被执行。
虽然对本申请的实施方式进行了说明,但该实施方式是作为例子提出的方式,并不限定申请的范围。该新型实施方式可以以其它各种形态实施,在不脱离发明的主旨的范围内可以进行各种省略、替换、变更。这些实施方式其变形包含在申请的范围、主旨中,并且包含在与权利要求书记载的发明等同的范围内。

Claims (15)

  1. 一种信息处理装置,具备:
    获取部,其将输入到声音输入部的用户的声音作为声音信号进行获取;
    得分计算部,其计算所述声音信号相对于声音数据的得分,其中,所述得分成为用于从所述声音信号中检测触发词的基准,所述触发词用于使声音识别服务开始;以及
    显示控制部,其将所述得分显示在显示部上。
  2. 根据权利要求1所述的信息处理装置,其中,
    所述得分计算部将所述声音数据与所述声音信号的一致度进行标准化而计算所述得分。
  3. 根据权利要求2所述的信息处理装置,其中,
    所述信息处理装置具备触发词检测部,所述触发词检测部从所述声音信号检测所述触发词,
    所述触发词检测部将所述声音数据和所述声音信号分解为多个要素,对所述多个要素计算所述一致度,基于所述一致度从所述声音信号检测所述触发词。
  4. 根据权利要求3所述的信息处理装置,其中,
    所述得分计算部分别对所述多个要素的每个要素的所述一致度计算所述得分。
  5. 根据权利要求4所述的信息处理装置,其中,
    所述显示控制部将所述得分中的最小的得分显示在所述显示部上。
  6. 根据权利要求4所述的信息处理装置,其中,
    所述显示控制部将分别对所述一致度计算出的所述得分显示在所述显示部上。
  7. 根据权利要求4所述的信息处理装置,其中,
    所述显示控制部将分别对所述一致度计算出的所述得分的平均值显示在 所述显示部上。
  8. 根据权利要求3至7中任一项所述的信息处理装置,其中,
    所述多个要素为包含在所述触发词中的音素。
  9. 根据权利要求1至8中任一项所述的信息处理装置,其中,
    所述得分计算部对多个所述触发词计算所述得分。
  10. 根据权利要求9所述的信息处理装置,其中,
    所述显示控制部将对多个所述触发词计算出的所述得分显示在所述显示部上。
  11. 根据权利要求1至10中任一项所述的信息处理装置,其中,
    所述显示控制部将用于提高所述得分的建议显示在所述显示部上。
  12. 根据权利要求1至11中任一项所述的信息处理装置,其中,
    所述获取部接收将所述得分显示在所述显示部上的指示的输入。
  13. 根据权利要求1至12中任一项所述的信息处理装置,其中,
    所述信息处理装置具备应用执行部,所述应用执行部在从所述声音信号检测到所述触发词时开始所述声音识别服务。
  14. 根据权利要求1至13中任一项所述的信息处理装置,其中,
    所述声音识别服务由通过网络连接的声音识别服务器提供。
  15. 一种计算机可读的非易失性存储介质,所述存储介质存储有程序,所述程序用于使计算机执行:
    将输入到声音输入部的用户的声音作为声音信号进行获取;
    计算所述声音信号相对于声音数据的得分,其中,所述得分成为用于从所述声音信号中检测触发词的基准,所述触发词用于使声音识别服务开始;以及
    将所述得分显示在显示部上。
PCT/CN2020/123669 2019-12-05 2020-10-26 信息处理装置及非易失性存储介质 Ceased WO2021109751A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
CN202080005757.3A CN113228170B (zh) 2019-12-05 2020-10-26 信息处理装置及非易失性存储介质

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2019-220035 2019-12-05
JP2019220035A JP7248564B2 (ja) 2019-12-05 2019-12-05 情報処理装置及びプログラム

Publications (1)

Publication Number Publication Date
WO2021109751A1 true WO2021109751A1 (zh) 2021-06-10

Family

ID=76220032

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/123669 Ceased WO2021109751A1 (zh) 2019-12-05 2020-10-26 信息处理装置及非易失性存储介质

Country Status (3)

Country Link
JP (1) JP7248564B2 (zh)
CN (1) CN113228170B (zh)
WO (1) WO2021109751A1 (zh)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP7858121B1 (ja) 2025-10-22 2026-05-13 祐樹 竹内 発言記録統合制御システム、方法及びプログラム

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101630448A (zh) * 2008-07-15 2010-01-20 上海启态网络科技有限公司 语言学习客户端及系统
US20130080161A1 (en) * 2011-09-27 2013-03-28 Kabushiki Kaisha Toshiba Speech recognition apparatus and method
CN105556594A (zh) * 2013-12-26 2016-05-04 松下知识产权经营株式会社 声音识别处理装置、声音识别处理方法以及显示装置
CN107924687A (zh) * 2015-09-23 2018-04-17 三星电子株式会社 语音识别设备、用户设备的语音识别方法和非暂时性计算机可读记录介质
CN109739354A (zh) * 2018-12-28 2019-05-10 广州励丰文化科技股份有限公司 一种基于声音的多媒体交互方法及装置

Family Cites Families (23)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH05158493A (ja) * 1991-12-10 1993-06-25 Fujitsu Ltd 音声認識装置
JP2001005480A (ja) * 1999-06-23 2001-01-12 Denso Corp ユーザー発音判定装置及び記録媒体
JP4654513B2 (ja) * 2000-12-25 2011-03-23 ヤマハ株式会社 楽器
JP2006011641A (ja) * 2004-06-23 2006-01-12 Fujitsu Ltd 情報入力方法及びその装置
JP2009124324A (ja) * 2007-11-13 2009-06-04 Sharp Corp 音響機器及び音響機器の制御方法
CN101266593A (zh) * 2008-02-25 2008-09-17 北京理工大学 一种基于网络收集意见的语音及音频质量主观评价方法
CN101547387A (zh) * 2008-03-26 2009-09-30 鸿富锦精密工业(深圳)有限公司 耳机及使用该耳机的音频播放系统
WO2012169679A1 (ko) * 2011-06-10 2012-12-13 엘지전자 주식회사 디스플레이 장치, 디스플레이 장치의 제어 방법 및 디스플레이 장치의 음성인식 시스템
US9536528B2 (en) * 2012-07-03 2017-01-03 Google Inc. Determining hotword suitability
US9275637B1 (en) * 2012-11-06 2016-03-01 Amazon Technologies, Inc. Wake word evaluation
KR20160034973A (ko) * 2013-07-19 2016-03-30 가부시키가이샤 베네세 코포레이션 정보 처리 장치, 정보 처리 방법 및 프로그램
US10789041B2 (en) * 2014-09-12 2020-09-29 Apple Inc. Dynamic thresholds for always listening speech trigger
CN104575504A (zh) * 2014-12-24 2015-04-29 上海师范大学 采用声纹和语音识别进行个性化电视语音唤醒的方法
JP6608254B2 (ja) * 2015-11-25 2019-11-20 オリンパス株式会社 録音機器、アドバイス出力方法およびプログラム
CN105702253A (zh) * 2016-01-07 2016-06-22 北京云知声信息技术有限公司 一种语音唤醒方法及装置
JP2019518985A (ja) * 2016-05-13 2019-07-04 ボーズ・コーポレーションBose Corporation 分散したマイクロホンからの音声の処理
US10957322B2 (en) * 2016-09-09 2021-03-23 Sony Corporation Speech processing apparatus, information processing apparatus, speech processing method, and information processing method
JP6553111B2 (ja) * 2017-03-21 2019-07-31 株式会社東芝 音声認識装置、音声認識方法及び音声認識プログラム
CN109601017B (zh) * 2017-08-02 2024-05-03 松下知识产权经营株式会社 信息处理装置、声音识别系统及信息处理方法
CN107358954A (zh) * 2017-08-29 2017-11-17 成都启英泰伦科技有限公司 一种实时更换唤醒词的设备及方法
KR102485342B1 (ko) * 2017-12-11 2023-01-05 현대자동차주식회사 차량의 환경에 기반한 추천 신뢰도 판단 장치 및 방법
CN108538293B (zh) * 2018-04-27 2021-05-28 海信视像科技股份有限公司 语音唤醒方法、装置及智能设备
CN109036393A (zh) * 2018-06-19 2018-12-18 广东美的厨房电器制造有限公司 家电设备的唤醒词训练方法、装置及家电设备

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101630448A (zh) * 2008-07-15 2010-01-20 上海启态网络科技有限公司 语言学习客户端及系统
US20130080161A1 (en) * 2011-09-27 2013-03-28 Kabushiki Kaisha Toshiba Speech recognition apparatus and method
CN105556594A (zh) * 2013-12-26 2016-05-04 松下知识产权经营株式会社 声音识别处理装置、声音识别处理方法以及显示装置
CN107924687A (zh) * 2015-09-23 2018-04-17 三星电子株式会社 语音识别设备、用户设备的语音识别方法和非暂时性计算机可读记录介质
CN109739354A (zh) * 2018-12-28 2019-05-10 广州励丰文化科技股份有限公司 一种基于声音的多媒体交互方法及装置

Also Published As

Publication number Publication date
JP7248564B2 (ja) 2023-03-29
CN113228170A (zh) 2021-08-06
JP2021089376A (ja) 2021-06-10
CN113228170B (zh) 2023-06-27

Similar Documents

Publication Publication Date Title
US12149781B2 (en) Methods and systems for detecting audio output of associated device
US9477304B2 (en) Information processing apparatus, information processing method, and program
US10991374B2 (en) Request-response procedure based voice control method, voice control device and computer readable storage medium
CN105741836B (zh) 声音识别装置以及声音识别方法
JP4369132B2 (ja) 話者音声のバックグランド学習
CN109643548B (zh) 用于将内容路由到相关联输出设备的系统和方法
KR20170032096A (ko) 전자장치, 전자장치의 구동방법, 음성인식장치, 음성인식장치의 구동 방법 및 컴퓨터 판독가능 기록매체
JPWO2017168936A1 (ja) 情報処理装置、情報処理方法、及びプログラム
US10089980B2 (en) Sound reproduction method, speech dialogue device, and recording medium
JP4729927B2 (ja) 音声検出装置、自動撮像装置、および音声検出方法
CN116583899A (zh) 用户语音简档管理
CN110024027A (zh) 说话人识别
EP2504745B1 (en) Communication interface apparatus and method for multi-user
JP2020095210A (ja) 議事録出力装置および議事録出力装置の制御プログラム
CN111587413A (zh) 信息处理装置、信息处理系统、信息处理方法和程序
JP2014241498A (ja) 番組推薦装置
JP6897678B2 (ja) 情報処理装置及び情報処理方法
JP2018005122A (ja) 検出装置、検出方法及び検出プログラム
CN113228170B (zh) 信息处理装置及非易失性存储介质
JP2001067091A (ja) 音声認識装置
CN107977187B (zh) 一种混响调节方法及电子设备
JP2008275987A (ja) 音声認識装置および会議システム
KR102743866B1 (ko) 전자장치와 그의 제어방법, 및 기록매체
JP2020144209A (ja) 音声処理装置、会議システム、及び音声処理方法
US20250259633A1 (en) Electronic device, server, and system including same

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20897269

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20897269

Country of ref document: EP

Kind code of ref document: A1