WO2024257308A1 - 分類装置、分類方法、記録媒体及び情報表示装置 - Google Patents

分類装置、分類方法、記録媒体及び情報表示装置 Download PDF

Info

Publication number
WO2024257308A1
WO2024257308A1 PCT/JP2023/022279 JP2023022279W WO2024257308A1 WO 2024257308 A1 WO2024257308 A1 WO 2024257308A1 JP 2023022279 W JP2023022279 W JP 2023022279W WO 2024257308 A1 WO2024257308 A1 WO 2024257308A1
Authority
WO
WIPO (PCT)
Prior art keywords
section
sounds
speaker
voice
sound quality
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2023/022279
Other languages
English (en)
French (fr)
Inventor
レイ カク
仁 山本
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Corp
Original Assignee
NEC Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by NEC Corp filed Critical NEC Corp
Priority to PCT/JP2023/022279 priority Critical patent/WO2024257308A1/ja
Priority to JP2025527160A priority patent/JPWO2024257308A1/ja
Publication of WO2024257308A1 publication Critical patent/WO2024257308A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/04Segmentation; Word boundary detection
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/02Preprocessing operations, e.g. segment selection; Pattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal components; Feature selection or extraction

Definitions

  • This disclosure relates to the technical fields of classification devices, classification methods, recording media, and information display devices.
  • Patent Document 1 a device has been proposed that identifies a portion of voice data corresponding to a desired keyword from voice data generated by recording a telephone call without requiring voice recognition processing.
  • the objective of this disclosure is to provide an audio processing device, an audio processing method, a recording medium, and an information display device that aim to improve upon the technology described in prior art documents.
  • One embodiment of the classification device includes a first classification means for classifying a plurality of section sounds generated by dividing an input voice into a plurality of sections into a plurality of first section sounds having a first sound quality and a plurality of second section sounds having a second sound quality different from the first sound quality, and a second classification means for classifying the plurality of second section sounds into one of a plurality of groups based on the plurality of first section sounds or into a new group different from the plurality of groups.
  • a computer classifies a plurality of section sounds generated by dividing an input voice into a plurality of sections into a plurality of first section sounds having a first sound quality and a plurality of second section sounds having a second sound quality different from the first sound quality, and classifies the plurality of second section sounds into one of a plurality of groups based on the plurality of first section sounds or into a new group different from the plurality of groups.
  • a computer program is recorded that causes a computer to execute a classification method that classifies a plurality of section sounds generated by dividing an input voice into a plurality of sections into a plurality of first section sounds having a first sound quality and a plurality of second section sounds having a second sound quality different from the first sound quality, and classifies the plurality of second section sounds into one of a plurality of groups based on the plurality of first section sounds or into a new group different from the plurality of groups.
  • One aspect of the information display device includes a first classification means for classifying a plurality of section sounds generated by dividing a first input voice into a plurality of sections into a plurality of first section sounds having a first sound quality and a plurality of second section sounds having a second sound quality different from the first sound quality, a second classification means for classifying the plurality of second section sounds into one of a plurality of groups based on the plurality of first section sounds or into a new group different from the plurality of groups, a first generation means for generating first speaker information indicating a speaker of each of the plurality of section sounds based on the plurality of groups and the new group, a second generation means for generating character information indicating a character string corresponding to the at least one section sound based on at least one section sound of the plurality of section sounds, and a display means for displaying the speaker indicated by the generated first speaker information and the character string indicated by the generated character information.
  • FIG. 1 is a block diagram showing an example of a configuration of a classification device.
  • FIG. 13 is a block diagram showing another example of the configuration of the classification device.
  • 1 is a flowchart illustrating an operation of a classification device according to the present disclosure.
  • FIG. 1 is a diagram illustrating a concept of a voice detection process.
  • FIG. 13 is a diagram illustrating an example of similarity.
  • FIG. 4 is a diagram showing an example of a display screen.
  • FIG. 13 is a diagram showing an example of a speaker input screen.
  • FIG. 13 is a diagram showing another example of the display screen.
  • the classification device 10 includes a classification unit 11 and a clustering unit 12.
  • the classification unit 11 classifies a plurality of section sounds into a plurality of first section sounds having a first sound quality and a plurality of second section sounds having a second sound quality different from the first sound quality. Note that the first sound quality may be higher than the second sound quality.
  • the multiple section sounds may be generated by dividing the input sound into multiple sections.
  • Existing technology can be applied to the method of dividing the input sound into multiple section sounds.
  • One example is a method of dividing the input sound into multiple section sounds in predetermined time units (e.g., 2-second units).
  • Another example is a method of dividing the input sound into multiple section sounds in predetermined time units with overlap.
  • the input sound may be acquired by a sound collection device such as a microphone.
  • the clustering unit 12 classifies each of the multiple second section sounds into one of multiple groups based on the multiple first section sounds, or into a new group different from the multiple groups.
  • classifying is not limited to dividing according to a predetermined criterion (for example, according to a correct answer determined in advance), but also includes dividing based on the similarity between data (for example, between one section sound and another section sound).
  • classification according to this embodiment may include clustering.
  • the multiple groups based on the multiple first section sounds may be generated (or set) by, for example, the following method.
  • the classification unit 11 may extract features of each of the multiple first section sounds.
  • An example of a section sound feature i.e., a sound feature
  • MFCC Mel-Frequency Cepstrum Coefficients
  • Another example of a section sound feature is an x-vector calculated using a Deep Neural Network (DNN).
  • DNN Deep Neural Network
  • the classification unit 11 may calculate the similarity between one first section sound and another first section sound based on the features of the multiple first section sounds.
  • An example of the similarity between two sounds is the similarity of i-vectors or x-vectors calculated by PLDA (Probabilistic Linear Discriminant Analysis).
  • the similarity may be referred to as an index value indicating the degree of similarity.
  • the extraction of the features of the section sounds and the calculation of the similarity may be performed by the clustering unit 12 instead of the classification unit 11.
  • the clustering unit 12 may perform clustering processing on the multiple first section sounds based on the similarity between the multiple first section sounds.
  • One example of the clustering processing is processing using hierarchical agglomerative clustering (AHC).
  • AHC hierarchical agglomerative clustering
  • the multiple first section sounds are grouped by the clustering processing, and the multiple groups (i.e., multiple groups based on the multiple first section sounds) may be generated (or set).
  • the clustering unit 12 may classify each of the multiple second section sounds into one of multiple groups based on the multiple first section sounds or into a new group different from the multiple groups, for example, by the following method.
  • the clustering unit 12 may extract features of each of the multiple second section sounds.
  • the features of each of the multiple second section sounds may be extracted by the same method as the features of the first section sounds.
  • the clustering unit 12 may calculate a similarity between the one second section sound and one or more first section sounds based on the features of the one second section sound and the features of one or more first section sounds included in one of the multiple groups based on the multiple first section sounds.
  • the clustering unit 12 may further calculate a similarity between the one second section sound and one or more first section sounds based on the features of the one second section sound and the features of one or more first section sounds included in another of the multiple groups based on the multiple first section sounds.
  • the clustering unit 12 may identify the maximum similarity among the similarities calculated as described above.
  • the clustering unit 12 may identify a combination of one second section audio and one first section audio that results in the identified maximum similarity, thereby identifying a group including the one first section audio (i.e., a group based on multiple first section audios). If the identified maximum similarity is equal to or greater than a predetermined threshold, the clustering unit 12 may classify the one second section audio into the identified group. On the other hand, if the identified maximum similarity is less than a predetermined threshold, the clustering unit 12 may classify the one second section audio into a new group different from the multiple groups based on the multiple first section audios.
  • the classification device 10 classifies the multiple section sounds generated by dividing the input voice into multiple sections into multiple first section sounds having a first sound quality and multiple second section sounds having a second sound quality different from the first sound quality, and performs a classification method in which each of the multiple second section sounds is classified into one of multiple groups based on the multiple first section sounds or into a new group different from the multiple groups.
  • the classification device 10 may be realized by a computer reading a computer program recorded on a recording medium.
  • the recording medium can be said to have recorded thereon a computer program that causes the computer to execute a classification method that classifies a plurality of section sounds generated by dividing an input voice into a plurality of sections into a plurality of first section sounds having a first sound quality and a plurality of second section sounds having a second sound quality different from the first sound quality, and classifies each of the plurality of second section sounds into one of a plurality of groups based on the plurality of first section sounds or into a new group different from the plurality of groups.
  • the first feature of the first voice associated with one speaker and the second feature of the second voice associated with the first speaker are similar to each other because the speakers are the same.
  • the first feature of the first voice associated with one speaker and the third feature of the third voice associated with another speaker are not similar to each other because the speakers are different. Therefore, the similarity between the first voice and the second voice based on the first feature and the second feature is higher than the similarity between the first voice and the third voice based on the first feature and the third feature.
  • the similarity between the first voice and the third voice based on the first feature and the third feature is lower than the similarity between the first voice and the second voice based on the first feature and the second feature.
  • the similarity between the first voice and the second voice will be lower than when the sound quality of the first voice and the second voice are high.
  • the similarity between the first voice and the third voice will be higher than when the sound quality of the first voice and the third voice are high.
  • the second feature of the second voice may or may not be similar to the first feature of the first voice.
  • the third feature of the third voice may or may not be similar to the first feature of the first voice. For this reason, when clustering processing is performed on multiple voices with different sound quality, there is a possibility that the multiple voices may not be appropriately separated by speaker.
  • the classification unit 11 classifies the multiple section sounds into multiple first section sounds having a first sound quality and multiple second section sounds having a second sound quality different from the first sound quality. Then, the clustering unit 12 classifies each of the multiple second section sounds into one of multiple groups based on the multiple first section sounds or into a new group different from the multiple groups. In other words, in the classification device 10, after each of the multiple section sounds is classified by sound quality, each of the multiple section sounds is grouped by sound quality. Therefore, the classification device 10 can appropriately separate multiple sounds by speaker.
  • the classification device 20 is used to describe the embodiments of the classification device, classification method, recording medium, and information display device.
  • the classification device 20 includes a calculation device 21, a storage device 22, and a communication device 23.
  • the classification device 20 may further include an input device 24 and an output device 25.
  • the classification device 20 does not have to include at least one of the input device 24 and the output device 25.
  • the calculation device 21, the storage device 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.
  • the computing device 21 may include, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), a TPU (Tensor Processing Unit), and a quantum processor.
  • a CPU Central Processing Unit
  • GPU Graphics Processing Unit
  • FPGA Field Programmable Gate Array
  • TPU Torsor Processing Unit
  • quantum processor a quantum processor
  • the storage device 22 may include, for example, at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, an optical magnetic disk device, an SSD (Solid State Drive), and an optical disk array.
  • the storage device 22 may include a non-transient recording medium.
  • the storage device 22 is capable of storing desired data.
  • the storage device 22 may temporarily store a computer program executed by the arithmetic device 21.
  • the storage device 22 may temporarily store data that is temporarily used by the arithmetic device 21 when the arithmetic device 21 is executing a computer program.
  • the communication device 23 may be capable of communicating with a device external to the classification device 20. Note that the communication device 23 may perform wired communication or wireless communication.
  • the input device 24 is a device capable of accepting information input to the classification device 20 from the outside.
  • the input device 24 may include an operating device (e.g., a keyboard, a mouse, a touch panel, etc.) that can be operated by an operator of the classification device 20.
  • the input device 24 may include a recording medium reading device that can read information recorded on a recording medium that is detachable from the classification device 20, such as a USB (Universal Serial Bus) memory.
  • the communication device 23 may function as an input device.
  • the output device 25 is a device capable of outputting information to the outside of the classification device 20.
  • the output device 25 may output visual information such as characters and images, auditory information such as voice, or tactile information such as vibration, as the above information.
  • the output device 25 may include at least one of a display, a speaker, a printer, and a vibration motor, for example.
  • the output device 25 may be capable of outputting information to a recording medium that is detachable from the classification device 20, such as a USB memory. Note that when the classification device 20 outputs information via the communication device 23, the communication device 23 may function as an output device.
  • the calculation device 21 may have a classification unit 211, a clustering unit 212, a speaker information generation unit 213, a character information generation unit 214, and a display control unit 215 as functional blocks that are logically realized or as processing circuits that are physically realized.
  • At least one of the classification unit 211, clustering unit 212, speaker information generation unit 213, character information generation unit 214, and display control unit 215 may be realized in a form that mixes logical functional blocks and physical processing circuits (i.e., hardware).
  • the classification unit 211, clustering unit 212, speaker information generation unit 213, character information generation unit 214, and display control unit 215 are functional blocks
  • at least some of the classification unit 211, clustering unit 212, speaker information generation unit 213, character information generation unit 214, and display control unit 215 may be realized by the calculation device 21 executing a predetermined computer program.
  • the calculation device 21 may obtain (in other words, read) the above-mentioned specific computer program from the storage device 22.
  • the calculation device 21 may read the above-mentioned specific computer program stored in a computer-readable and non-transient recording medium using a recording medium reading device (not shown) provided in the classification device 20.
  • the calculation device 21 may obtain (in other words, download or read) the above-mentioned specific computer program from a device (not shown) external to the classification device 20 via the communication device 23.
  • the recording medium for recording the above-mentioned specific computer program executed by the calculation device 21 may be at least one of an optical disk, a magnetic medium, a magneto-optical disk, a semiconductor memory, and any other medium capable of storing a program.
  • the "classification unit 211" and “clustering unit 212" are components that correspond to the “classification unit 11" and “clustering unit 12" in the first embodiment described above, respectively.
  • the classification unit 211 detects a section in which speech is present from the input speech (in other words, a section including an acoustic signal corresponding to speech) (step S101).
  • a section in which speech is present is hereinafter referred to as a "speech section" as appropriate.
  • the classification unit 211 may detect the speech section by performing speech detection processing on the input speech. Note that various existing aspects can be applied to the speech detection processing, and therefore detailed explanations thereof will be omitted. Note that when detecting speech related to human speech, the speech detection processing may be referred to as speech detection processing.
  • the input audio may include audio and non-audio sounds (for example, at least one of sounds, music, and animal cries).
  • the input audio may also include silent intervals.
  • audio intervals VS1, VS2, VS3, VS4, and VS5 may be detected from the input audio, for example, as shown in FIG. 4.
  • intervals BS1, BS2, BS3, and BS4 in FIG. 4 are examples of intervals other than audio intervals.
  • intervals other than audio intervals may include, for example, intervals where sounds other than audio exist and silent intervals.
  • the classification unit 211 may divide one input speech into multiple section speeches by extracting detected speech sections (e.g., speech sections VS1, VS2, VS3, VS4, and VS5) from one input speech.
  • one speech section may correspond to one section speech.
  • the classification unit 211 may divide one input speech to generate multiple section speeches, and then perform speech detection processing on each of the multiple section speeches as the input speech.
  • the classification unit 211 may extract a speech section from each of the multiple section speeches.
  • the classification unit 211 may further divide one section speech by extracting a speech section from the one section speech.
  • the classification unit 211 extracts speaker features of the detected speech sections (e.g., speech sections VS1, VS2, VS3, VS4, and VS5) (step S102).
  • the speaker features may be at least one of an i-vector and an x-vector.
  • the classification unit 211 may extract speaker features using a learning model that, when a speech section is input, outputs speaker features of the input speech section.
  • a learning model may be constructed by machine learning using voice data corresponding to the speech section.
  • the classification unit 211 calculates the similarity between each of the multiple voice sections based on the speaker features extracted in the processing of step S102 (step S103).
  • a voice section can be said to be one section of one input voice. Therefore, a voice section may be referred to as a section voice. Therefore, calculating the similarity between each of the multiple voice sections can be said to be calculating the similarity between each of the multiple section voices.
  • the similarity may be the similarity of an i-vector or x-vector calculated by PLDA.
  • the classification unit 211 may calculate the similarity using a learning model that, when two speaker features are input, outputs the similarity between two voice sections that respectively correspond to the two input speaker features. Such a learning model may be constructed by machine learning using learning data corresponding to the speaker features.
  • the speaker features of the speech sections VS1, VS2, VS3, VS4, and VS5 are speaker features SF1, SF2, SF3, SF4, and SF5, respectively.
  • the calculation results of the similarity between each of the speech sections VS1, VS2, VS3, VS4, and VS5 can be represented as the table shown in FIG. 5.
  • the value at the position where the speaker feature SF1 on the vertical axis intersects with the speaker feature SF2 on the horizontal axis (specifically, "0.1") indicates the similarity between the speech section VS1 and the speech section VS2.
  • the more similar the speaker feature of one speech section is to the speaker feature of another speech section the larger the similarity value.
  • the classification unit 211 may calculate distance instead of similarity.
  • the classification unit 211 may calculate distance instead of similarity as an index indicating the degree of similarity between a speaker feature of one voice section and a speaker feature of another voice section. In this case, the more similar the speaker feature of one voice section is to the speaker feature of another voice section, the smaller the value indicating the distance may be.
  • the classification unit 211 classifies the multiple voice sections into a voice section having a first sound quality and a voice section having a second sound quality different from the first sound quality (step S104).
  • the classification unit 211 may classify the multiple voice sections into a voice section having a first sound quality and a voice section having a second sound quality based on the sound quality of each of the multiple voice sections and the reference sound quality. For example, if the sound quality of a voice section is higher than the reference sound quality, the classification unit 211 may classify the voice section into a voice section having a first sound quality.
  • the classification unit 211 may classify the voice section into a voice section having a second sound quality.
  • the reference sound quality may be, for example, an average value of the sound quality of the speech data used to construct a learning model related to speaker features.
  • the first sound quality may be higher than the second sound quality.
  • the voice sections VS1, VS2, VS3, VS4, and VS5 are high-quality voice sections (in other words, voice sections having a first sound quality), and the voice sections VS4 and VS5 are low-quality voice sections (in other words, voice sections having a second sound quality).
  • the speaker related to the voice sections VS1 and VS3 is the first speaker
  • the speaker related to the voice sections VS2, VS4, and VS5 is the second speaker.
  • the classification unit 211 may classify the voice sections VS1, VS2, and VS3 as voice sections having the first sound quality, and classify the voice sections VS4 and VS5 as voice sections having the second sound quality.
  • the clustering unit 212 performs a first clustering process using the multiple voice segments classified as having the first sound quality in the process of step S104 (step S105).
  • One example of the first clustering process is a process using hierarchical agglomerative clustering (AHC).
  • the first clustering process will be specifically described using the similarity shown in FIG. 5.
  • the clustering unit 212 may identify the combination of speech sections that maximizes the similarity among the respective speech sections VS1, VS2, and VS3.
  • the value (i.e., "0.9") at the position where the speaker feature SF1 of speech section VS1 intersects with the speaker feature SF3 of speech section VS3 is the maximum. Therefore, the clustering unit 212 may identify the combination of speech section VS1 and speech section VS3 as the combination of speech sections that maximizes the similarity.
  • the clustering unit 212 may determine whether the similarity between the voice section VS1 and the voice section VS3 is equal to or greater than a first threshold.
  • the first threshold is assumed to be "0.5".
  • the similarity between the voice section VS1 and the voice section VS3 is "0.9", which is greater than the first threshold. Therefore, the clustering unit 212 may classify the voice section VS1 and the voice section VS3 into the same group (which may be referred to as a "cluster").
  • the group including the voice sections VS1 and VS3 is referred to as group C1.
  • the clustering unit 212 may calculate the similarity between the voice sections VS1 and VS3 and the voice section VS2. In this case, the clustering unit 212 may determine the sum of the similarity between the voice sections VS1 and VS2 and the similarity between the voice sections VS2 and VS3 divided by 2 as the similarity between the voice sections VS1 and VS3 and the voice section VS2.
  • the value of the position where the speaker feature SF1 of the speech section VS1 intersects with the speaker feature SF2 of the speech section VS2 is "0.1".
  • the clustering unit 212 may classify the voice section VS2 into a group different from the voice sections VS1 and VS3.
  • the group including the voice section VS2 is referred to as group C2.
  • Groups C1 and C2 may be generated (or set) by performing a first clustering process using the voice sections VS1, VS2, and VS3 as voice sections having a first sound quality. Therefore, it can be said that groups C1 and C2 are multiple groups based on voice sections having a first sound quality.
  • the clustering unit 212 performs a second clustering process (step S106).
  • the clustering unit 212 performs the second clustering process to classify the voice segments having the second sound quality into one of a plurality of groups (e.g., groups C1 and C2) based on the voice segments having the first sound quality, or into a new group different from the plurality of groups.
  • groups C1 and C2 e.g., groups C1 and C2
  • the second clustering process will be specifically described using the similarity shown in FIG. 5.
  • the voice sections VS4 and VS5 correspond to an example of a voice section having the second sound quality. Therefore, by performing the second clustering process, the clustering unit 212 classifies each of the voice sections VS4 and VS5 into either of the above-mentioned groups C1 and C2, or into a new group different from groups C1 and C2.
  • the clustering unit 212 may calculate the similarity between the voice section VS4 and each of the groups C1 and C2, and the similarity between the voice section VS5 and each of the groups C1 and C2.
  • the similarity between the voice section and the group may be calculated based on the similarity between the voice section and one or more voice sections included in the group.
  • the similarity between the voice section VS4 and group C1 may be calculated based on the similarity between the voice section VS4 and each of the voice sections VS1 and VS3 included in group C1.
  • the clustering unit 212 may determine the similarity between the voice section VS4 and group C1 as the similarity between the voice section VS4 and group C1 by dividing the sum of the similarity between the voice section VS1 and the voice section VS4 and the similarity between the voice section VS3 and the voice section VS4 by 2.
  • the value of the position where the speaker feature SF1 of the speech section VS1 and the speaker feature SF4 of the speech section VS4 intersect is "0.5".
  • the similarity between the speech section VS4 and the group C2 is "0.4".
  • the similarity between the speech section VS5 and the group C1 is "0.4".
  • the similarity between the speech section VS5 and the group C2 is "0.5".
  • the clustering unit 212 may identify the combination that has the greatest similarity among the calculated similarities.
  • the combination of the voice section VS5 and group C2 may be identified as the combination that has the greatest similarity.
  • the clustering unit 212 may determine whether the similarity between the voice section VS5 and group C2 is equal to or greater than a second threshold.
  • the second threshold may be the same as the first threshold (i.e., the threshold related to the first clustering process), or may be different.
  • the second threshold is assumed to be "0.5".
  • the similarity between the voice section VS5 and group C2 is "0.5", which is equal to the second threshold. Therefore, the clustering unit 212 may classify the voice section VS5 into group C2. If the similarity between the voice section VS5 and group C2 is smaller than the second threshold, the clustering unit 212 may classify the voice section VS5 into a new group C3 that is different from groups C1 and C2.
  • the clustering unit 212 may recalculate the similarity between the voice section VS4 and each of the groups C1 and C2. Because group C2 includes voice sections VS2 and VS5, the similarity between the voice section VS4 and group C2 is "0.55". The clustering unit 212 may determine whether the similarity between the voice section VS4 and group C2 is equal to or greater than the second threshold. The similarity between the voice section VS4 and group C2 is "0.55", which is greater than the second threshold. Therefore, the clustering unit 212 may classify the voice section VS4 into group C2.
  • the voice sections VS1 and VS3 are classified into group C1, and the voice sections VS2, VS4, and VS5 are classified into group C2.
  • the speaker related to the voice sections VS1 and VS3 is the first speaker
  • the speaker related to the voice sections VS2, VS4, and VS5 is the second speaker. Therefore, it can be seen that the voice sections VS1, VS2, VS3, VS4, and VS5 have been appropriately divided by speaker by the first clustering process and the second clustering process.
  • the speaker information generating unit 213 generates speaker information indicating the speakers of a speech section (e.g., speech sections VS1, VS2, VS3, and VS4) based on the results of the clustering process (specifically, the first clustering process and the second clustering process) by the clustering unit 212.
  • the speaker information generating unit 213 may generate speaker information indicating the speakers in a manner that allows one speaker to be distinguished (or identified) from other speakers.
  • the speaker information generating unit 213 may represent the speakers as, for example, "speaker 1" and "speaker 2", etc.
  • the speaker information generating unit 213 may generate speaker information in which the display manner of a speaker of a voice section (i.e., a voice section having a first sound quality) classified into one group (e.g., one of groups C1 and C2) by the first clustering process is different from that of a speaker of another voice section (i.e., a voice section having a second sound quality) classified into the above-mentioned group by the second clustering process.
  • the speaker information may show characters indicating a speaker who is the speaker of a voice section and characters suggesting the above-mentioned speaker as the speaker of another voice section.
  • the likelihood of clustering can be presented to a user of the classification device 20.
  • the speaker information generating unit 213 does not need to generate speaker information indicating the speaker of the voice section classified into the new group.
  • the speaker information generating unit 213 may represent the speaker of the voice section classified into the new group as, for example, "Unknown”.
  • "Unknown" is an example and is not limited to this.
  • the speaker information generating unit 213 may identify a specific person corresponding to a single speaker by performing at least one of a speaker recognition process and a voice authentication process on one or more voice sections that have been separated as voices relating to a single speaker based on the result of the clustering process performed by the clustering unit 212.
  • the speaker information generating unit 213 may represent the speaker by a specific person's name.
  • the text information generating unit 214 In parallel with the processing of steps S102 to S106 in FIG. 3, the text information generating unit 214 generates text information indicating a character string corresponding to the speech, based on the speech related to the speech section detected in the processing of step S101. Note that various existing methods can be applied to the method of generating text information indicating a character string from speech (so-called transcription), and therefore detailed explanations thereof will be omitted.
  • the display control unit 215 controls the output device 25, which can function as a display device, to display the speaker indicated by the speaker information generated by the speaker information generation unit 213 and the character string indicated by the character information generated by the character information generation unit 214 in association with each other.
  • the output device 25 may display, for example, the image shown in FIG. 6.
  • the display control unit 215 may also control the output device 25 to display an image (e.g., an icon) instead of or in addition to the character indicating the speaker.
  • the classification device 20 may be referred to as an information display device, since it displays the speaker indicated by the speaker information and the character string indicated by the character information.
  • the display control unit 215 may control the output device 25 to display, for example, a speaker input screen as shown in FIG. 7.
  • a user of the classification device 20 may input the name of at least one of the speakers displayed on the speaker input screen via the input device 24.
  • the speaker information generating unit 213 may change the speaker indicated by the speaker information based on the input name.
  • the display control unit 215 may control the output device 25 to display the speaker indicated by the speaker information changed by the speaker information generating unit 213 and the character string indicated by the character information in association with each other.
  • the plurality of voices may not be appropriately divided into groups for each speaker.
  • the voice sections VS1, VS2, VS3, VS4, and VS5 are classified into group C1
  • the voice section VS2 is classified into group C2
  • the voice sections VS4 and VS5 are classified into group C3, which is different from groups C1 and C2. Note that a description of the calculation process will be omitted.
  • the speaker involved in the speech segments VS1 and VS3 is the first speaker
  • the speaker involved in the speech segments VS2, VS4, and VS5 is the second speaker. Therefore, when the first clustering process described above is performed on the speech segments VS1, VS2, VS3, VS4, and VS5, it is found that the speech segments VS1, VS2, VS3, VS4, and VS5 are not divided by speaker.
  • the classification unit 211 classifies multiple voice segments into voice segments having a first sound quality and voice segments having a second sound quality.
  • the clustering unit 212 performs a first clustering process on the multiple voice segments classified into the voice segments having the first sound quality.
  • the clustering unit 212 then classifies the voice segments classified into the voice segments having the second sound quality into one of multiple groups based on the voice segments having the first sound quality, or into a new group different from the multiple groups. In this way, by classifying multiple voice segments by sound quality and then performing a clustering process for each sound quality, multiple voice segments with different sound quality can be appropriately separated for each speaker.
  • the first clustering process by performing the first clustering process on a plurality of voice segments having the first sound quality, which is higher than the second sound quality, it is possible to perform grouping more appropriately than when the first clustering process is performed on a plurality of voice segments having the second sound quality. Since a character string corresponding to the voice is displayed in association with the speaker, this is useful in practical use.
  • the input device 24 may have a microphone.
  • the classification unit 211 may detect a voice section by performing a voice detection process on sounds sequentially acquired by the microphone as the input device 24.
  • the character information generation unit 214 may generate first character information indicating a character string corresponding to the voice, based on the voice related to the detected voice section.
  • the display control unit 215 may control the output device 25 to display the character string indicated by the first character information generated by the character information generation unit 214. As a result, the output device 25 may display, for example, an image shown in FIG. 8. In this case, since the speaker information generation unit 213 has not generated speaker information, information indicating the speaker (for example, a character or an image) is not displayed.
  • the classification unit 211 may extract speaker features for each of the multiple voice sections.
  • the classification unit 211 may calculate the similarity between each of the multiple voice sections based on the extracted speaker features.
  • the classification unit 211 may then classify the multiple voice sections into voice sections having a first sound quality and voice sections having a second sound quality.
  • the clustering unit 212 may perform a first clustering process on the multiple voice sections classified into the voice sections having the first sound quality.
  • the clustering unit 212 may perform a second clustering process to classify the voice sections classified into the voice sections having the second sound quality into one of multiple groups based on the voice sections having the first sound quality, or into a new group different from the multiple groups.
  • the speaker information generating unit 213 may generate first speaker information indicating the speaker of each of the multiple speech segments based on the results of the clustering process (specifically, the first clustering process and the second clustering process).
  • the display control unit 215 may control the output device 25 to display the speaker indicated by the first speaker information generated by the speaker information generating unit 213 and the character string indicated by the first character information generated by the character information generating unit 214.
  • the output device 25 may display, for example, an image as shown in FIG. 6.
  • the classification device 20 may treat the sound acquired by the microphone in a first period as one input sound, and the sound acquired by the microphone in a second period following the first period as another input sound. Alternatively, the classification device 20 may treat the sound acquired by the microphone in the first period and the sound acquired by the microphone in the second period following the first period together as one input sound. In this case, the one input sound is updated by acquiring a new sound by the microphone.
  • This section describes the operation of the classification device 20 when it treats a sound captured by a microphone during a first period and a sound captured by a microphone during a second period following the first period as a single input sound.
  • the classification unit 211 may detect a voice section by performing a voice detection process on one input voice, which is a sound acquired by a microphone during a first period.
  • the classification unit 211 may extract speaker features of each of the multiple voice sections.
  • the classification unit 211 may calculate a similarity between each of the multiple voice sections based on the extracted speaker features.
  • the classification unit 211 may then classify the multiple voice sections into a voice section having a first sound quality and a voice section having a second sound quality.
  • the clustering unit 212 may perform a first clustering process on the multiple voice sections classified into the voice section having the first sound quality.
  • the clustering unit 212 may perform a second clustering process to classify the voice section classified into the voice section having the second sound quality into one of multiple groups based on the voice section having the first sound quality, or into a new group different from the multiple groups.
  • the speaker information generation unit 213 may generate first speaker information indicating a speaker of each of the multiple voice sections based on the result of the clustering process.
  • the character information generating unit 214 may generate first character information indicating a character string corresponding to the voice based on the voice related to the detected voice section.
  • the display control unit 215 may control the output device 25 to display the speaker indicated by the first speaker information generated by the speaker information generating unit 213 and the character string indicated by the first character information generated by the character information generating unit 214.
  • the classification device 20 may update the input audio by combining the sound captured by the microphone during the first period with the sound captured by the microphone during the second period.
  • the classification unit 211 may detect a voice section by performing a voice detection process on the updated input voice. In this case, the classification unit 211 may perform a voice detection process again on the sound acquired by the microphone during the first period included in the updated input voice. Note that the classification unit 211 may not perform a voice detection process on the sound acquired by the microphone during the first period included in the updated input voice.
  • the classification unit 211 may extract speaker features of each of the multiple voice sections. The classification unit 211 may calculate the similarity between each of the multiple voice sections based on the extracted speaker features. Thereafter, the classification unit 211 may classify the multiple voice sections into a voice section having a first sound quality and a voice section having a second sound quality.
  • the clustering unit 212 may perform a first clustering process on the multiple voice sections classified into the voice section having the first sound quality.
  • the clustering unit 212 may perform a second clustering process to classify the voice segments classified as having the second sound quality into one of a plurality of groups based on the voice segments having the first sound quality, or into a new group different from the plurality of groups.
  • the speaker information generating unit 213 may generate second speaker information indicating a speaker of each of the plurality of voice segments based on the result of the clustering process.
  • the character information generating unit 214 may generate second character information indicating a character string corresponding to the voice, based on the voice related to the detected voice segment.
  • the display control unit 215 may control the output device 25 to display the speaker indicated by the second speaker information generated by the speaker information generating unit 213 and the character string indicated by the second character information generated by the character information generating unit 214.
  • the display control unit 215 may control the output device 25 to change the speaker displayed in association with the character string from the speaker indicated by the second speaker information to the speaker indicated by the second speaker information when the second speaker information is generated.
  • the output device 25 may change the speaker displayed in association with the character string, for example, from "Speaker 1" to "Speaker 2.”
  • the classification device 20 of the present embodiment it is possible to display a character string corresponding to an utterance in real time while the utterance is being made. Therefore, for example, it is possible to generate minutes of a meeting in real time.
  • a classification device comprising:
  • the second classification means calculates a second index value indicating a degree of similarity between one of the second section sounds and one of the groups based on one or more first index values indicating a degree of similarity between one of the second section sounds and each of the multiple section sounds, the second index value indicating a degree of similarity between one of the second section sounds and one or more first section sounds of the multiple first section sounds included in one of the multiple groups; the second classification means classifies the one second section speech into a group corresponding to a maximum second index value among a plurality of second index values calculated for each of the plurality of groups, the maximum second index value being equal to or greater than a predetermined threshold.
  • the second classification means calculates a second index value indicating a degree of similarity between one of the second section sounds and one of the groups based on one or more first index values indicating a degree of similarity between one of the second section sounds and each of the multiple section sounds, the second index value indicating a degree of similarity between one of the second section sounds and one or more first section sounds of the multiple first section sounds included in one of the multiple groups;
  • the second classification means classifies the one second section speech into the new group when all of a plurality of second index values calculated for each of the plurality of groups are smaller than a predetermined threshold.
  • Appendix 4 The classification device according to any one of appendix 1 to 3, wherein the classification means classifies the plurality of section sounds into the plurality of first section sounds and the plurality of second section sounds based on a sound quality of each of the plurality of section sounds and a reference sound quality.
  • the computer classifying a plurality of section sounds generated by dividing an input voice into a plurality of sections into a plurality of first section sounds having a first sound quality and a plurality of second section sounds having a second sound quality different from the first sound quality; classifying the second section sounds into any of a plurality of groups based on the first section sounds or into a new group different from the plurality of groups.
  • a first classification means for classifying a plurality of section sounds generated by dividing a first input voice into a plurality of sections into a plurality of first section sounds having a first sound quality and a plurality of second section sounds having a second sound quality different from the first sound quality; a second classification means for classifying the second section sounds into one of a plurality of groups based on the first section sounds or into a new group different from the plurality of groups; a first generating means for generating first speaker information indicating a speaker of each of the plurality of section sounds based on the plurality of groups and the new group; a second generating means for generating character information indicating a character string corresponding to at least one section sound based on at least one section sound of the plurality of section sounds; a display means for displaying a speaker indicated by the generated first speaker information and a character string indicated by the generated character information;
  • An information display device comprising:
  • the display means displays a character string corresponding to a section sound that corresponds to a first section sound classified into a first group among the plurality of groups and the new group, among the at least one section sound, in association with a speaker indicated by the generated first speaker information.
  • the information display device (Appendix 10) The information display device according to claim 8 or 9, wherein the first speaker information indicates a first speaker who is a speaker of a first section voice classified into a second group among the plurality of groups, and indicates characters suggesting the first speaker as a speaker of a second section voice classified into the second group.
  • the first classification means classifies the other plurality of section sounds into a plurality of third section sounds having the first sound quality and a plurality of fourth section sounds having the second sound quality
  • the second classification means classifies the plurality of second section sounds and the plurality of fourth section sounds into any of a plurality of groups based on the plurality of first section sounds and the plurality of third section sounds, or into a new group different from the plurality of groups
  • the first generating means generates second speaker information indicating a speaker of each of the plurality of section sounds and the other plurality of section sounds based on the plurality of groups and the new group;
  • An information display device as described in any one of Appendices 8 to 10, wherein when the speaker of a section audio corresponding to a character string indicated by the generated character information is different between the speaker indicated by the first speaker information

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Computational Linguistics (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

分類装置は、入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、第1音質とは異なる第2音質を有する複数の第2区間音声とに分類する第1分類手段と、複数の第1区間音声に基づく複数の群のいずれか又は該複数の群とは異なる新たな群に、複数の第2区間音声を分類する第2分類手段と、を備える。

Description

分類装置、分類方法、記録媒体及び情報表示装置
 この開示は、分類装置、分類方法、記録媒体及び情報表示装置の技術分野に関する。
 この種の装置として、例えば、通話を録音することで生成された音声データ内から、音声認識処理を要することなく、所望のキーワードに対応する音声データの箇所を特定する装置が提案されている(特許文献1参照)。その他、この開示に関連する技術として、特許文献2乃至4が挙げられる。
特開2011-053563号公報 特開2022-144927号公報 特開2015-040931号公報 特開2013-228472号公報
 この開示は、先行技術文献に記載された技術の改良を目的とする音声処理装置、音声処理方法、記録媒体及び情報表示装置を提供することを課題とする。
 分類装置の一態様は、入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類する第1分類手段と、前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する第2分類手段と、を備える。
 分類方法の一態様は、コンピュータが、入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類し、前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する。
 記録媒体の一態様は、コンピュータに、入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類し、前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する分類方法を実行させるコンピュータプログラムが記録されている。
 情報表示装置の一態様は、第1入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類する第1分類手段と、前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する第2分類手段と、前記複数の群及び前記新たな群に基づいて、前記複数の区間音声各々の話者を示す第1話者情報を生成する第1生成手段と、前記複数の区間音声の少なくとも一つの区間音声に基づいて、前記少なくとも一つの区間音声に相当する文字列を示す文字情報を生成する第2生成手段と、前記生成された第1話者情報により示される話者と、前記生成された文字情報により示される文字列とを表示する表示手段と、を備える。
分類装置の構成の一例を示すブロック図である。 分類装置の構成の他の例を示すブロック図である。 本開示にかかる分類装置の動作を示すフローチャートである。 音声検出処理の概念を示す図である。 類似度の一例を示す図である。 表示画面の一例を示す図である。 話者入力画面の一例を示す図である。 表示画面の他の例を示す図である。
 <第1実施形態>
 分類装置、分類方法及び記録媒体に係る実施形態について図1を参照して説明する。以下では、分類装置10を用いて、分類装置、分類方法及び記録媒体に係る実施形態を説明する。
 図1において、分類装置10は、分類部11及びクラスタリング部12を備える。分類部11は、複数の区間音声を、第1音質を有する複数の第1区間音声と、第1音質とは異なる第2音質を有する複数の第2区間音声とに分類する。尚、第1音質は、第2音質より高音質であってよい。
 複数の区間音声は、入力音声が複数の区間に分割されることによって生成されてよい。入力音声を複数の区間音声に分割する方法には既存の技術を適用可能である。一例としては、入力音声を、所定の時間単位(例えば、2秒単位)で、複数の区間音声に分割する方法が挙げられる。他の例としては、入力音声を、オーバーラップを持たせながら所定の時間単位で、複数の区間音声に分割する方法が挙げられる。尚、入力音声は、例えばマイクロフォン等の集音装置により取得されてよい。
 クラスタリング部12は、複数の第1区間音声に基づく複数の群のいずれか又は該複数の群とは異なる新たな群に、複数の第2区間音声各々を分類する。ここで、「分類する」とは、所定の基準に従って(例えば、事前に決定されている正解に従って)分けることに限らず、データ間(例えば、一の区間音声と他の区間音声との間)の類似度に基づいて分けることも含む概念である。つまり、本実施形態に係る「分類」には、クラスタリングが含まれていてよい。
 複数の第1区間音声に基づく複数の群は、例えば以下の方法により生成(又は、設定)されてよい。分類部11は、複数の第1区間音声各々の特徴を抽出してよい。区間音声の特徴(即ち、音声の特徴)の一例として、音響特徴であるメル周波数ケプストラム係数(MFCC:Mel-Frequency Cepstrum Coefficients)を用いて算出したi-vectorが挙げられる。区間音声の特徴の他の例として、DNN(Deep Neural Network)を用いて算出したx-vectorが挙げられる。
 分類部11は、複数の第1区間音声の特徴に基づいて、一の第1区間音声と他の第1区間音声との間の類似度を算出してよい。2つの音声間の類似度の一例として、PLDA(Probabilistic linear discriminant analysis)により計算されたi-vector又はx-vectorの類似度が挙げられる。尚、類似度は、類似の程度を示す指標値と称されてもよい。尚、区間音声の特徴の抽出及び類似度の計算は、分類部11に代えてクラスタリング部12により行われてよい。
 クラスタリング部12は、複数の第1区間音声に係る類似度に基づいて、複数の第1区間音声に対してクラスタリング処理を行ってよい。クラスタリング処理の一例として、階層型凝集的クラスタリング(Agglomerative Hierarchical Clustering:AHC)を用いた処理が挙げられる。クラスタリング処理により複数の第1区間音声がグループ分けされることで、上記複数の群(即ち、複数の第1区間音声に基づく複数の群)が生成(又は、設定)されてよい。
 クラスタリング部12は、例えば以下の方法により、複数の第1区間音声に基づく複数の群のいずれか又は該複数の群とは異なる新たな群に、複数の第2区間音声各々を分類してよい。クラスタリング部12は、複数の第2区間音声各々の特徴を抽出してよい。複数の第2区間音声各々の特徴は、第1区間音声の特徴と同様の方法により抽出されてよい。
 クラスタリング部12は、一の第2区間音声の特徴と、複数の第1区間音声に基づく複数の群のうち一の群に含まれる一又は複数の第1区間音声の特徴とに基づいて、一の第2区間音声と、一又は複数の第1区間音声との間の類似度を計算してよい。クラスタリング部12は更に、一の第2区間音声の特徴と、複数の第1区間音声に基づく複数の群のうち他の群に含まれる一又は複数の第1区間音声の特徴とに基づいて、一の第2区間音声と、一又は複数の第1区間音声との間の類似度を計算してよい。
 クラスタリング部12は、上記のように計算された類似度のうち、最大の類似度を特定してよい。クラスタリング部12は、該特定された最大の類似度となる、一の第2区間音声と一の第1区間音声との組み合わせを特定することにより、該一の第1区間音声が含まれる群(即ち、複数の第1区間音声に基づく群)を特定してよい。クラスタリング部12は、上記特定された最大の類似度が、所定の閾値以上である場合、上記特定された群に一の第2区間音声を分類してよい。他方で、クラスタリング部12は、上記特定された最大の類似度が、所定の閾値未満である場合、複数の第1区間音声に基づく複数の群とは異なる新たな群に、一の第2区間音声を分類してよい。
 このように、分類装置10では、入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、第1音質とは異なる第2音質を有する複数の第2区間音声とに分類し、複数の第1区間音声に基づく複数の群のいずれか又は該複数の群とは異なる新たな群に、複数の第2区間音声各々を分類する分類方法が行われる。
 分類装置10は、コンピュータが記録媒体に記録されたコンピュータプログラムを読み込むことによって実現されてよい。この場合、記録媒体には、コンピュータに、入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、第1音質とは異なる第2音質を有する複数の第2区間音声とに分類し、複数の第1区間音声に基づく複数の群のいずれか又は該複数の群とは異なる新たな群に、複数の第2区間音声各々を分類する分類方法を実行させるコンピュータプログラムが記録されている、と言える。
 (技術的効果)
 本願発明者の研究によれば次の事が判明している。一の話者に係る第1音声の第1特徴と、一の話者に係る第2音声の第2特徴とは、話者が同一であるがゆえに互いに似ている。一の話者に係る第1音声の第1特徴と、他の話者に係る第3音声の第3特徴とは、話者が異なるがゆえに似ていない。このため、第1特徴及び第2特徴に基づく第1音声と第2音声との間の類似度は、第1特徴及び第3特徴に基づく第1音声と第3音声との間の類似度に比べて高くなる。言い換えれば、第1特徴及び第3特徴に基づく第1音声と第3音声との間の類似度は、第1特徴及び第2特徴に基づく第1音声と第2音声との間の類似度に比べて低くなる。音声が高音質である場合、2つの音声の話者が同一である場合の類似度と、2つの音声の話者が互いに異なる場合の類似度とは明確に異なる。
 例えば、第1音声の音質が高音質であり、第2音声の音質が低音質である場合、第1音声と第2音声との間の類似度は、第1音声及び第2音声の音質が高音質である場合に比べて低くなる。他方で、第1音声の音質が高音質であり、第3音声の音質が低音質である場合、第1音声と第3音声との間の類似度は、第1音声及び第3音声が高音質である場合に比べて高くなる。この結果、第2音声の第2特徴が、第1音声の第1特徴に似ているとも言えるし、似ていないとも言えない状態になり得る。同様に、第3音声の第3特徴が、第1音声の第1特徴に似ているとも言えるし、似ていないとも言える状況になり得る。このため、音質が異なる複数の音声に対してクラスタリング処理を行うと、複数の音声を話者毎に適切に分けられない可能性がある。
 これに対して、分類装置10では、分類部11が、複数の区間音声を、第1音質を有する複数の第1区間音声と、第1音質とは異なる第2音質を有する複数の第2区間音声とに分類する。そして、クラスタリング部12が、複数の第1区間音声に基づく複数の群のいずれか又は該複数の群とは異なる新たな群に、複数の第2区間音声各々を分類する。つまり、分類装置10では、音質によって複数の区間音声各々が分類された後に、音質毎に複数の区間音声各々がグループ分けされる。このため、分類装置10によれば、複数の音声を話者毎に適切に分けることができる。
 <第2実施形態>
 分類装置、分類方法、記録媒体及び情報表示装置に係る実施形態について図2乃至図7を参照して説明する。以下では、分類装置20を用いて、分類装置、分類方法、記録媒体及び情報表示装置に係る実施形態を説明する。
 図2において、分類装置20は、演算装置21、記憶装置22及び通信装置23を備える。分類装置20は更に、入力装置24及び出力装置25を備えていてよい。尚、分類装置20は、入力装置24及び出力装置25の少なくとも一方を備えていなくてもよい。分類装置20において、演算装置21、記憶装置22、通信装置23、入力装置24及び出力装置25は、データバス26を介して接続されていてよい。
 演算装置21は、例えば、CPU(Central Processing Unit)、GPU(Graphics Processing Unit)、FPGA(Field Programmable Gate Array)、TPU(TensorProcessingUnit)、及び、量子プロセッサのうち少なくとも一つを含んでよい。
 記憶装置22は、例えば、RAM(Random Access Memory)、ROM(Read Only Memory)、ハードディスク装置、光磁気ディスク装置、SSD(Solid State Drive)、及び、光ディスクアレイのうち少なくとも一つを含んでよい。つまり、記憶装置22は、一時的でない記録媒体を含んでよい。記憶装置22は、所望のデータを記憶可能である。例えば、記憶装置22は、演算装置21が実行するコンピュータプログラムを一時的に記憶していてよい。記憶装置22は、演算装置21がコンピュータプログラムを実行している場合に演算装置21が一時的に使用するデータを一時的に記憶してよい。
 通信装置23は、分類装置20の外部の装置と通信可能であってもよい。尚、通信装置23は、有線通信を行ってもよいし、無線通信を行ってもよい。
 入力装置24は、外部から分類装置20に対する情報の入力を受け付け可能な装置である。入力装置24は、分類装置20のオペレータが操作可能な操作装置(例えば、キーボード、マウス、タッチパネル等)を含んでよい。入力装置24は、例えばUSB(Universal Serial Bus)メモリ等の、分類装置20に着脱可能な記録媒体に記録されている情報を読み取り可能な記録媒体読取装置を含んでよい。尚、分類装置20に、通信装置23を介して情報が入力される場合(言い換えれば、分類装置20が通信装置23を介して情報を取得する場合)、通信装置23は入力装置として機能してよい。
 出力装置25は、分類装置20の外部に対して情報を出力可能な装置である。出力装置25は、上記情報として、文字や画像等の視覚情報を出力してもよいし、音声等の聴覚情報を出力してもよいし、振動等の触覚情報を出力してもよい。出力装置25は、例えば、ディスプレイ、スピーカ、プリンタ及び振動モータの少なくとも一つを含んでいてよい。出力装置25は、例えばUSBメモリ等の、分類装置20に着脱可能な記録媒体に情報を出力可能であってもよい。尚、分類装置20が通信装置23を介して情報を出力する場合、通信装置23は出力装置として機能してよい。
 演算装置21は、論理的に実現される機能ブロックとして、又は、物理的に実現される処理回路として、分類部211、クラスタリング部212、話者情報生成部213、文字情報生成部214及び表示制御部215を有していてよい。
 尚、分類部211、クラスタリング部212、話者情報生成部213、文字情報生成部214及び表示制御部215の少なくとも一つは、論理的な機能ブロックと、物理的な処理回路(即ち、ハードウェア)とが混在する形式で実現されてよい。分類部211、クラスタリング部212、話者情報生成部213、文字情報生成部214及び表示制御部215の少なくとも一部が機能ブロックである場合、分類部211、クラスタリング部212、話者情報生成部213、文字情報生成部214及び表示制御部215の少なくとも一部は、演算装置21が所定のコンピュータプログラムを実行することにより実現されてよい。
 演算装置21は、上記所定のコンピュータプログラムを、記憶装置22から取得してよい(言い換えれば、読み込んでよい)。演算装置21は、コンピュータで読み取り可能であって且つ一時的でない記録媒体が記憶している上記所定のコンピュータプログラムを、分類装置20が備える図示しない記録媒体読み取り装置を用いて読み込んでもよい。演算装置21は、通信装置23を介して、分類装置20の外部の図示しない装置から上記所定のコンピュータプログラムを取得してもよい(言い換えれば、ダウンロードしてもよい又は読み込んでもよい)。尚、演算装置21が実行する上記所定のコンピュータプログラムを記録する記録媒体としては、光ディスク、磁気媒体、光磁気ディスク、半導体メモリ、及び、その他プログラムを格納可能な任意の媒体の少なくとも一つが用いられてよい。
 尚、「分類部211」及び「クラスタリング212」は、夫々、上述した第1実施形態における「分類部11」及び「クラスタリング部12」に対応する構成要素である。
 分類部211及びクラスタリング部212の動作について図3乃至図5を参照して説明する。尚、分類部211及びクラスタリング部212について、上述した第1実施形態の説明と重複する説明を適宜省略する。
 図3において、分類部211は、入力音声から音声が存在する区間(言い換えれば、音声に対応する音響信号を含む区間)を検出する(ステップS101)。音声が存在する区間を、以降、適宜「音声区間」と称する。ステップS101の処理において、分類部211は、入力音声に対して音声検出処理を行うことによって、音声区間を検出してよい。尚、音声検出処理には、既存の各種態様を適用可能であるので、その詳細についての説明は省略する。尚、人の発話に係る音声を検出する場合、音声検出処理は、発話検出処理と称されてもよい。
 入力音声には、音声と、音声以外の音(例えば、物音、音楽及び動物の鳴声の少なくとも一つ)とが含まれていてよい。また、入力音声には、無音区間が含まれていてよい。ステップS101の処理の結果、例えば図4に示すように、入力音声から音声区間VS1、VS2、VS3、VS4及びVS5が検出されてよい。尚、図4における区間BS1、BS2、BS3及びBS4は、音声区間以外の区間の一例である。尚、音声区間以外の区間には、例えば、音声以外の音が存在する区間及び無音区間が含まれていてよい。
 分類部211は、一の入力音声から、検出された音声区間(例えば、音声区間VS1、VS2、VS3、VS4及びVS5)を抽出することによって、一の入力音声を複数の区間音声に分割してよい。この場合、一の音声区間が一の区間音声に対応してよい。或いは、分類部211は、一の入力音声を分割して複数の区間音声を生成した後に、入力音声としての複数の区間音声各々に対して音声検出処理を行ってもよい。この場合、分類部211は、複数の区間音声各々から音声区間を抽出してもよい。つまり、分類部211は、一の区間音声から音声区間を抽出することによって、一の区間音声を更に分割してもよい。
 ステップS101の処理の後、分類部211は、検出された音声区間(例えば、音声区間VS1、VS2、VS3、VS4及びVS5)の話者特徴量を抽出する(ステップS102)。尚、話者特徴量は、i-vector及びx-vectorの少なくとも一方であってよい。尚、分類部211は、音声区間を入力すると、入力された音声区間の話者特徴量を出力する学習モデルを用いて、話者特徴量を抽出してもよい。このような学習モデルは、音声区間に相当する音声データを用いた機械学習により構築されてよい。
 次に、分類部211は、ステップS102の処理において抽出された話者特徴量に基づいて、複数の音声区間の夫々の間の類似度を計算する(ステップS103)。尚、音声区間は、一の入力音声の一の区間であると言える。このため、音声区間は区間音声と称されてもよい。このため、複数の音声区間の夫々の間の類似度を計算することは、複数の区間音声の夫々の間の類似度を計算することであると言える。尚、類似度は、PLDAにより計算されたi-vector又はx-vectorの類似度であってよい。尚、分類部211は、2つの話者特徴量を入力すると、入力された2つの話者特徴量に夫々対応する2つの音声区間の類似度を出力する学習モデルを用いて、類似度を計算してもよい。このような学習モデルは、話者特徴量に相当する学習データを用いた機械学習により構築されてよい。
 例えば、音声区間VS1、VS2、VS3、VS4及びVS5の話者特徴量を、夫々、話者特徴量SF1、SF2、SF3、SF4及びSF5とする。音声区間VS1、VS2、VS3、VS4及びVS5の夫々の間の類似度の計算結果は、図5に示す表として表すことができる。図5において、例えば、縦軸上の話者特徴量SF1と、横軸上の話者特徴量SF2とが交差する位置の値(具体的には、“0.1”)は、音声区間VS1と音声区間VS2との間の類似度を示している。本実施形態では、一の音声区間の話者特徴量と他の音声区間の話者特徴量とが似ているほど、類似度の値が大きくなる。
 尚、分類部211は、類似度に代えて、距離を計算してもよい。つまり、分類部211は、一の音声区間の話者特徴量と他の音声区間の話者特徴量とが似ている度合いを示す指標として、類似度に代えて距離を計算してもよい。この場合、一の音声区間の話者特徴量と他の音声区間の話者特徴量とが似ているほど、距離を示す値が小さくなってよい。
 ステップS103の処理の後、分類部211は、複数の音声区間を、第1音質を有する音声区間と、第1音質とは異なる第2音質を有する音声区間とに分類する(ステップS104)。ステップS104の処理において、分類部211は、複数の音声区間各々の音質と基準音質とに基づいて、複数の音声区間を、第1音質を有する音声区間と、第2音質を有する音声区間とに分類してよい。例えば、分類部211は、音声区間の音質が基準音質より高い場合、該音声区間を第1音質を有する音声区間に分類してよい。他方で、分類部211は、音声区間の音質が基準音質より低い場合、該音声区間を第2音質を有する音声区間に分類してよい。尚、音声区間の音質と基準音質とが「等しい」場合は、どちらかの場合に含めて扱えばよい。このように構成すれば、複数の音声区間を、第1音質を有する音声区間と第2音質を有する音声区間とに容易に分類することができる。尚、基準音質は、例えば話者特徴量に係る学習モデルを構築するために用いられた音声データの音質の平均値であってよい。尚、第1音質は、第2音質より高音質であってよい。
 例えば、音声区間VS1、VS2、VS3、VS4及びVS5のうち、音声区間VS1、VS2及びVS3が高音質な音声区間(言い換えれば、第1音質を有する音声区間)であり、音声区間VS4及びVS5が低音質な音声区間(言い換えれば、第2音質を有する音声区間)であるものとする。また、音声区間VS1及びVS3に係る話者が第1話者であり、音声区間VS2、VS4及びVS5に係る話者が第2話者であるものとする。この場合、ステップS104の処理において、分類部211は、音声区間VS1、VS2及びVS3を第1音質を有する音声区間に分類するとともに、音声区間VS4及びVS5を第2音質を有する音声区間に分類してよい。
 クラスタリング部212は、ステップS104の処理において第1音質を有する音声区間に分類された複数の音声区間を用いて第1のクラスタリング処理を行う(ステップS105)。第1のクラスタリング処理の一例として、階層型凝集的クラスタリング(AHC)を用いた処理が挙げられる。
 第1のクラスタリング処理について、図5に示す類似度を用いて具体的に説明する。クラスタリング部212は、音声区間VS1、VS2及びVS3の夫々の間の類似度のうち、類似度が最大となる音声区間の組合せを特定してよい。図5に示す例では、音声区間VS1の話者特徴量SF1と音声区間VS3の話者特徴量SF3とが交差する位置の値(即ち、“0.9”)が最大である。このため、クラスタリング部212は、類似度が最大となる音声区間の組合せとして、音声区間VS1と音声区間VS3との組合せを特定してよい。
 次に、クラスタリング部212は、音声区間VS1と音声区間VS3との間の類似度が、第1閾値以上であるか否かを判定してよい。ここで、第1閾値は「0.5」であるものとする。音声区間VS1と音声区間VS3との間の類似度は、「0.9」であり、第1閾値より大きい。このため、クラスタリング部212は、音声区間VS1と音声区間VS3とを同じグループ(“クラスタ”と称されてもよい)に分けてよい。ここで、音声区間VS1及びVS3を含むグループを、グループC1とする。
 次に、クラスタリング部212は、音声区間VS1及びVS3と、音声区間VS2との間の類似度を計算してよい。この場合、クラスタリング部212は、音声区間VS1と音声区間VS2との間の類似度と、音声区間VS2と音声区間VS3との間の類似度との合計値を2で割った値を、音声区間VS1及びVS3と、音声区間VS2との間の類似度としてよい。
 図5において、音声区間VS1の話者特徴量SF1と音声区間VS2の話者特徴量SF2とが交差する位置の値は「0.1」である。音声区間VS2の話者特徴量SF2と音声区間VS3の話者特徴量SF3とが交差する位置の値は「0.1」である。このため、音声区間VS1及びVS3と、音声区間VS2との間の類似度は、“(0.1+0.1)/2=0.1”となる。
 音声区間VS1及びVS3と、音声区間VS2との間の類似度(即ち、“0.1”)は、第1閾値(即ち、“0.5”)より小さい。このため、クラスタリング部212は、音声区間VS2を、音声区間VS1及びVS3とは異なるグループに分けてよい。ここで、音声区間VS2を含むグループを、グループC2とする。
 第1音質を有する音声区間としての音声区間VS1、VS2及びVS3を用いて第1のクラスタリング処理が行われることによって、グループC1及びC2が生成(又は、設定)されてよい。このため、グループC1及びC2は、第1音質を有する音声区間に基づく複数の群である、と言える。
 ステップS105の処理の後、クラスタリング部212は、第2のクラスタリング処理を行う(ステップS106)。ステップS106の処理において、クラスタリング部212は、第2のクラスタリング処理を行うことによって、第2音質を有する音声区間を、第1音質を有する音声区間に基づく複数の群(例えば、グループC1及びC2)のいずれか、又は、該複数の群とは異なる新たな群に分類する。
 第2のクラスタリング処理について、図5に示す類似度を用いて具体的に説明する。上述したように、音声区間VS4及びVS5が第2音質を有する音声区間の一例に相当する。従って、クラスタリング部212は、第2のクラスタリング処理を行うことによって、音声区間VS4及びVS5各々を、上述したグループC1及びC2のいずれか、又は、グループC1及びC2とは異なる新たなグループに分類する。
 第2のクラスタリング処理において、クラスタリング部212は、音声区間VS4とグループC1及びC2各々との間の類似度、並びに、音声区間VS5とグループC1及びC2各々との間の類似度を計算してよい。ここで、音声区間とグループとの間の類似度は、音声区間と、グループに含まれる一又は複数の音声区間との間の類似度に基づいて計算されてよい。
 例えば、音声区間VS4とグループC1との間の類似度は、音声区間VS4と、グループC1に含まれる音声区間VS1及びVS3各々との間の類似度に基づいて計算されてよい。この場合、クラスタリング部212は、音声区間VS1と音声区間VS4との間の類似度と、音声区間VS3と音声区間VS4との間の類似度との合計値を2で割った値を、音声区間VS4とグループC1との間の類似度としてよい。
 図5において、音声区間VS1の話者特徴量SF1と音声区間VS4の話者特徴量SF4とが交差する位置の値は「0.5」である。音声区間VS3の話者特徴量SF3と音声区間VS4の話者特徴量SF4とが交差する位置の値は「0.4」である。このため、音声区間VS4とグループC1との間の類似度は、“(0.5+0.4)/2=0.45”となる。同様の計算により、音声区間VS4とグループC2との間の類似度は「0.4」となる。音声区間VS5とグループC1との間の類似度は「0.4」となる。音声区間VS5とグループC2との間の類似度は「0.5」となる。
 クラスタリング部212は、計算された類似度のうち類似度が最大となる組合せを特定してよい。ここでは、音声区間VS5とグループC2との組合せが、類似度が最大となる組合せとして特定されてよい。次に、クラスタリング部212は、音声区間VS5とグループC2との間の類似度が第2閾値以上であるか否かを判定してよい。第2閾値は、第1閾値(即ち、第1のクラスタリング処理に係る閾値)と同じであってもよいし、異なっていてもよい。ここでは、第2閾値は「0.5」であるものとする。
 音声区間VS5とグループC2との間の類似度は「0.5」であり、第2閾値と等しい。このため、クラスタリング部212は、音声区間VS5をグループC2に分類してよい。仮に、音声区間VS5とグループC2との間の類似度が第2閾値より小さい場合、クラスタリング部212は、音声区間VS5を、グループC1及びC2とは異なる新たなグループC3に分類してよい。
 次に、クラスタリング部212は、音声区間VS4とグループC1及びC2各々との間の類似度を再度計算してよい。グループC2に音声区間VS2及びVS5が含まれるので、音声区間VS4とグループC2との間の類似度は「0.55」となる。クラスタリング部212は、音声区間VS4とグループC2との間の類似度が第2閾値以上であるか否かを判定してよい。音声区間VS4とグループC2との間の類似度は「0.55」であり、第2閾値より大きい。このため、クラスタリング部212は、音声区間VS4をグループC2に分類してよい。
 第1のクラスタリング処理及び第2のクラスタリング処理の結果、音声区間VS1及びVS3がグループC1に分類され、音声区間VS2、VS4及びVS5がグループC2に分類される。上述したように、音声区間VS1及びVS3に係る話者が第1話者であり、音声区間VS2、VS4及びVS5に係る話者が第2話者である。このため、第1のクラスタリング処理及び第2のクラスタリング処理により、音声区間VS1、VS2、VS3、VS4及びVS5が話者毎に適切に分けられたことがわかる。
 話者情報生成部213は、クラスタリング部212によるクラスタリング処理(具体的には、第1のクラスタリング処理及び第2のクラスタリング処理)の結果に基づいて、音声区間(例えば、音声区間VS1、VS2、VS3及びVS4)の話者を示す話者情報を生成する。話者情報生成部213は、一の話者と他の話者とを区別(又は、識別)可能な態様で話者を示す話者情報を生成してよい。話者情報生成部213は、話者を、例えば「話者1」及び「話者2」等と表してよい。
 話者情報生成部213は、第1のクラスタリング処理によって一の群(例えば、グループC1及びC2の一方)に分類された一の音声区間(即ち、第1音質を有する音声区間)の話者と、第2のクラスタリング処理によって上記一の群に分類された他の音声区間(即ち、第2音質を有する音声区間)の話者との表示態様が互いに異なる話者情報を生成してもよい。この場合、話者情報は、一の音声区間の話者である一の話者を示す文字と、他の音声区間の話者として、上記一の話者を示唆する文字とを示してよい。「一の話者を示唆する文字」は、例えば、一の話者を示す文字が「話者1」である場合、「話者1かも」及び「話者1(推測)」の一方であってよい。このように構成すれば、クラスタリングの確からしさを、分類装置20のユーザに提示することができる。尚、第2のクラスタリング処理によって、新たなグループ(例えば、グループC1及びC2とは異なるグループ)に分類された音声区間が存在する場合、話者情報生成部213は、新たなグループに分類された音声区間の話者を示す話者情報を生成しなくてもよい。或いは、第2のクラスタリング処理によって、新たなグループ(例えば、グループC1及びC2とは異なるグループ)に分類された音声区間が存在する場合、話者情報生成部213は、新たなグループに分類された音声区間の話者を、例えば「Unkown」等と表してよい。尚、「Unkown」は一例であり、これに限定されるものではない。
 尚、話者情報生成部213は、クラスタリング部212によるクラスタリング処理の結果に基づいて、一の話者に係る音声として分けられた一又は複数の音声区間に、話者認識処理及び声認証処理の少なくとも一方を施すことにより、一の話者に対応する具体的な人物を特定してもよい。この場合、話者情報生成部213は、話者を具体的な人名で表してよい。
 図3のステップS102からS106の処理と並行して、文字情報生成部214は、ステップS101の処理において検出された音声区間に係る音声に基づいて、音声に相当する文字列を示す文字情報を生成する。尚、音声から文字列を示す文字情報を生成する方法(いわゆる文字起こし)には、既存の各種態様を適用可能であるので、その詳細についての説明は省略する。
 表示制御部215は、話者情報生成部213により生成された話者情報により示される話者と、文字情報生成部214により生成された文字情報により示される文字列とを互いに対応づけて表示するように、表示装置として機能可能な出力装置25を制御する。その結果、出力装置25は、例えば図6に示す画像を表示してよい。尚、表示制御部215は、話者を示す文字に代えて又は加えて、画像(例えば、アイコン)を表示するように、出力装置25を制御してもよい。尚、分類装置20は、話者情報により示される話者と文字情報により示される文字列とを表示するので、情報表示装置と称されてもよい。
 尚、話者情報生成部213が、話者認識処理及び声認証処理を行わない場合、表示制御部215は、例えば図7に示す話者入力画面を表示するように出力装置25を制御してよい。分類装置20のユーザは、入力装置24を介して、話者入力画面に表示される話者のうち少なくとも一人の話者の人名を入力してよい。話者情報生成部213は、入力された人名に基づいて、話者情報により示される話者を変更してよい。表示制御部215は、話者情報生成部213により変更された話者情報により示される話者と、文字情報により示される文字列とを互いに対応づけて表示するように、出力装置25を制御してよい。
 (技術的効果)
 音質が異なる複数の音声に対してクラスタリング処理(上述した第1のクラスタリング処理に相当)を行うと、複数の音声を話者毎に適切に分けられない可能性がある。例えば、音声区間VS1、VS2、VS3、VS4及びVS5に対して、上述した第1のクラスタリング処理を行うと、音声区間VS1及びVS2がグループC1に分類され、音声区間VS2がグループC2に分類され、音声区間VS4及びVS5が、グループC1及びC2とは異なるグループC3に分類される。尚、計算過程の説明は省略する。
 上述したように、音声区間VS1及びVS3に係る話者が第1話者であり、音声区間VS2、VS4及びVS5に係る話者が第2話者である。従って、音声区間VS1、VS2、VS3、VS4及びVS5に対して、上述した第1のクラスタリング処理を行うと、音声区間VS1、VS2、VS3、VS4及びVS5が話者毎に分けられていないことがわかる。
 これに対して、分類装置20では、分類部211により複数の音声区間が、第1音質を有する音声区間と、第2音質を有する音声区間とに分類される。クラスタリング部212により、第1音質を有する音声区間に分類された複数の音声区間に対して第1のクラスタリング処理が行われる。その後、クラスタリング部212により、第2音質を有する音声区間に分類された音声区間が、第1音質を有する音声区間に基づく複数のグループのいずれか、又は、該複数のグループとは異なる新たなグループに分類される。このように、複数の音声区間を音質で分類した上で、音質毎にクラスタリング処理を行うことによって、音質が異なる複数の音声区間を話者毎に適切に分けることができる。
 尚、第2音質より高音質である第1音質を有する複数の音声区間に対して第1のクラスタリング処理が行われることにより、第2音質を有する複数の音声区間に対して第1のクラスタリング処理が行われる場合に比べて、適切にグループ分けをすることができる。音声に相当する文字列が話者に対応づけられて表示されるので、実用上有用である。
 <第3実施形態>
 分類装置、分類方法、記録媒体及び情報表示装置に係る実施形態について図8を参照して説明する。以下では、分類装置20を用いて、分類装置、分類方法、記録媒体及び情報表示装置に係る実施形態を説明する。第3実施形態では、表示制御部215の動作について主に説明する。その他の構成については、上述した第2実施形態と同様であってよい。第3実施形態について、第2実施形態の説明と重複する説明を適宜省略する。
 入力装置24は、マイクロフォンを有していてよい。分類部211は、入力装置24としてのマイクロフォンにより逐次取得される音に対して音声検出処理を行うことによって、音声区間を検出してよい。文字情報生成部214は、検出された音声区間に係る音声に基づいて、音声に相当する文字列を示す第1文字情報を生成してよい。表示制御部215は、文字情報生成部214により生成された第1文字情報により示される文字列を表示するように出力装置25を制御してよい。その結果、出力装置25は、例えば図8に示す画像を表示してよい。この場合、話者情報生成部213により話者情報が生成されていないので、話者を示す情報(例えば、文字又は画像)は表示されない。
 上記処理と並行して、分類部211は、複数の音声区間各々の話者特徴量を抽出してよい。分類部211は、抽出された話者特徴量に基づいて、複数の音声区間の夫々の間の類似度を計算してよい。その後、分類部211は、複数の音声区間を、第1音質を有する音声区間と、第2音質を有する音声区間とに分類してよい。クラスタリング部212は、第1音質を有する音声区間に分類された複数の音声区間に対して第1のクラスタリング処理を行ってよい。クラスタリング部212は、第2のクラスタリング処理を行うことによって、第2音質を有する音声区間に分類された音声区間を、第1音質を有する音声区間に基づく複数のグループのいずれか、又は、該複数のグループとは異なる新たなグループに分類してよい。
 話者情報生成部213は、クラスタリング処理(具体的には、第1のクラスタリング処理及び第2のクラスタリング処理)の結果に基づいて、複数の音声区間各々の話者を示す第1話者情報を生成してよい。表示制御部215は、話者情報生成部213により生成された第1話者情報により示される話者と、文字情報生成部214により生成された第1文字情報により示される文字列を表示するように出力装置25を制御してよい。その結果、出力装置25は、例えば図6に示す画像を表示してよい。
 尚、分類装置20は、第1の期間にマイクロフォンにより取得された音を一の入力音声とし、第1の期間に続く第2の期間にマイクロフォンにより取得された音を他の入力音声として扱ってよい。或いは、分類装置20は、第1の期間にマイクロフォンにより取得された音と、第1の期間に続く第2の期間にマイクロフォンにより取得された音とを合わせて一の入力音声として扱ってよい。この場合、マイクロフォンにより新たな音が取得されることによって、一の入力音声が更新されることとなる。
 分類装置20が、第1の期間にマイクロフォンにより取得された音と、第1の期間に続く第2の期間にマイクロフォンにより取得された音とを合わせて一の入力音声として扱う場合の動作について説明する。
 分類部211は、第1の期間にマイクロフォンにより取得された音である一の入力音声に対して音声検出処理を行うことによって、音声区間を検出してよい。分類部211は、複数の音声区間各々の話者特徴量を抽出してよい。分類部211は、抽出された話者特徴量に基づいて、複数の音声区間の夫々の間の類似度を計算してよい。その後、分類部211は、複数の音声区間を、第1音質を有する音声区間と、第2音質を有する音声区間とに分類してよい。クラスタリング部212は、第1音質を有する音声区間に分類された複数の音声区間に対して第1のクラスタリング処理を行ってよい。クラスタリング部212は、第2のクラスタリング処理を行うことによって、第2音質を有する音声区間に分類された音声区間を、第1音質を有する音声区間に基づく複数のグループのいずれか、又は、該複数のグループとは異なる新たなグループに分類してよい。話者情報生成部213は、クラスタリング処理の結果に基づいて、複数の音声区間各々の話者を示す第1話者情報を生成してよい。文字情報生成部214は、検出された音声区間に係る音声に基づいて、音声に相当する文字列を示す第1文字情報を生成してよい。表示制御部215は、話者情報生成部213により生成された第1話者情報により示される話者と、文字情報生成部214により生成された第1文字情報により示される文字列を表示するように出力装置25を制御してよい。
 第1の期間に続く第2の期間にマイクロフォンにより取得された音が入力された場合、分類装置20は、第1の期間にマイクロフォンにより取得された音と、第2の期間にマイクロフォンにより取得された音とを合わせることによって、一の入力音声を更新してよい。
 分類部211は、更新された一の入力音声に対して音声検出処理を行うことによって、音声区間を検出してよい。この場合、分類部211は、更新された一の入力音声に含まれる第1の期間にマイクロフォンにより取得された音に対して再度音声検出処理を行ってよい。尚、分類部211は、更新された一の入力音声に含まれる第1の期間にマイクロフォンにより取得された音に対して音声検出処理を行わなくてもよい。分類部211は、複数の音声区間各々の話者特徴量を抽出してよい。分類部211は、抽出された話者特徴量に基づいて、複数の音声区間の夫々の間の類似度を計算してよい。その後、分類部211は、複数の音声区間を、第1音質を有する音声区間と、第2音質を有する音声区間とに分類してよい。クラスタリング部212は、第1音質を有する音声区間に分類された複数の音声区間に対して第1のクラスタリング処理を行ってよい。クラスタリング部212は、第2のクラスタリング処理を行うことによって、第2音質を有する音声区間に分類された音声区間を、第1音質を有する音声区間に基づく複数のグループのいずれか、又は、該複数のグループとは異なる新たなグループに分類してよい。話者情報生成部213は、クラスタリング処理の結果に基づいて、複数の音声区間各々の話者を示す第2話者情報を生成してよい。文字情報生成部214は、検出された音声区間に係る音声に基づいて、音声に相当する文字列を示す第2文字情報を生成してよい。表示制御部215は、話者情報生成部213により生成された第2話者情報により示される話者と、文字情報生成部214により生成された第2文字情報により示される文字列を表示するように出力装置25を制御してよい。
 第1文字情報及び第2文字情報の両方に含まれる一の文字列(例えば、第1の期間にマイクロフォンにより取得された音に含まれる一の音声に相当する一の文字列)の話者が、第1話者情報により示される話者と、第2話者情報により示される話者とで異なる場合、表示制御部215は、第2話者情報が生成された際に、上記一の文字列に対応づけて表示される話者を、第2話者情報により示される話者から、第2話者情報により示される話者に変更するように、出力装置25を制御してよい。その結果、出力装置25は、一の文字列に対応づけて表示される話者を、例えば「話者1」から「話者2」に変更してよい。
 (技術的効果)
 本実施形態に係る分類装置20によれば、発話が行われている状況において、発話に相当する文字列をリアルタイムに表示することができる。このため、例えば会議の議事録をリアルタイムに生成することができる。
 <付記>
 以上説明した実施形態に関して、更に以下の付記を開示する。
 (付記1)
 入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類する第1分類手段と、
 前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する第2分類手段と、
 を備える分類装置。
 (付記2)
 前記第2分類手段は、前記複数の区間音声のうち一の区間音声と、前記複数の区間音声各々との類似の程度を夫々示す複数の第1指標値のうち、前記複数の第2区間音声のうち一の第2区間音声と、前記複数の群のうち一の群に含まれる、前記複数の第1区間音声の一又は複数の第1区間音声との類似の程度を示す一又は複数の第1指標値に基づいて、前記一の第2区間音声と前記一の群との類似の程度を示す第2指標値を算出し、
 前記第2分類手段は、前記複数の群各々について算出された複数の第2指標値のうち、最大の第2指標値であって、所定の閾値以上の第2指標値に対応する群に、前記一の第2区間音声を分類する
 付記1に記載の分類装置。
 (付記3)
 前記第2分類手段は、前記複数の区間音声のうち一の区間音声と、前記複数の区間音声各々との類似の程度を夫々示す複数の第1指標値のうち、前記複数の第2区間音声のうち一の第2区間音声と、前記複数の群のうち一の群に含まれる、前記複数の第1区間音声の一又は複数の第1区間音声との類似の程度を示す一又は複数の第1指標値に基づいて、前記一の第2区間音声と前記一の群との類似の程度を示す第2指標値を算出し、
 前記第2分類手段は、前記複数の群各々について算出された複数の第2指標値のいずれもが所定の閾値より小さい場合、前記新たな群に、前記一の第2区間音声を分類する
 付記1又は2に記載の分類装置。
 (付記4)
 前記分類手段は、前記複数の区間音声各々の音質と、基準音質とに基づいて、前記複数の区間音声を、前記複数の第1区間音声と前記複数の第2区間音声とに分類する
 付記1乃至3のいずれか一項に記載の分類装置。
 (付記5)
 前記第2音質は、前記第1音質より低い音質である
 付記1乃至4のいずれか一項に記載の分類装置。
 (付記6)
 コンピュータが、
 入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類し、
 前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する
 分類方法。
 (付記7)
 コンピュータに、
 入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類し、
 前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する
 分類方法を実行させるコンピュータプログラムが記録されている記録媒体。
 (付記8)
 第1入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類する第1分類手段と、
 前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する第2分類手段と、
 前記複数の群及び前記新たな群に基づいて、前記複数の区間音声各々の話者を示す第1話者情報を生成する第1生成手段と、
 前記複数の区間音声の少なくとも一つの区間音声に基づいて、前記少なくとも一つの区間音声に相当する文字列を示す文字情報を生成する第2生成手段と、
 前記生成された第1話者情報により示される話者と、前記生成された文字情報により示される文字列とを表示する表示手段と、
 を備える情報表示装置。
 (付記9)
 前記表示手段は、前記少なくとも一つの区間音声のうち、前記複数の群及び前記新たな群のうち第1群に分類された第1区間音声に該当する区間音声に相当する文字列を、前記生成された第1話者情報により示される話者と対応付けて表示する
 付記8に記載の情報表示装置。
 (付記10)
 前記第1話者情報は、前記複数の群のうち第2群に分類された第1区間音声の話者である第1話者を示すとともに、前記第2群に分類された第2区間音声の話者として、前記第1話者を示唆する文字を示す
 請求項8又は9に記載の情報表示装置。
 (付記11)
 前記第1入力音声が複数の区間に分割されることによって生成された複数の区間音声に、第2入力音声が複数の区間に分割されることによって生成された他の複数の区間音声が加えられた場合、
 前記第1分類手段は、前記他の複数の区間音声を、前記第1音質を有する複数の第3区間音声と、前記第2音質を有する複数の第4区間音声とに分類し、
 前記第2分類手段は、前記複数の第1区間音声及び前記複数の第3区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声及び前記複数の第4区間音声を分類し、
 前記第1生成手段は、前記複数の群及び前記新たな群に基づいて、前記複数の区間音声及び前記他の複数の区間音声各々の話者を示す第2話者情報を生成し、
 前記生成された文字情報により示される一の文字列に相当する一の区間音声の話者が、前記第1話者情報により示される話者と、前記第2話者情報により示される話者とで異なる場合、前記表示手段は、前記第2話者情報が生成された際に、前記一の文字列に対応づけられる話者を、前記第1話者情報により示される話者から、前記第2話者情報により示される話者に変更する
 付記8乃至10のいずれか一項に記載の情報表示装置。
 この開示は、上述した実施形態に限られるものではなく、請求の範囲及び明細書全体から読み取れる発明の要旨或いは思想に反しない範囲で適宜変更可能であり、そのような変更を伴う分類装置、分類方法、記録媒体及び情報表示装置もまたこの開示の技術的範囲に含まれるものである。
 10、20 分類装置
 11、211 分類部
 12、212 クラスタリング部
 213 話者情報生成部
 214 文字情報生成部
 215 表示制御部

Claims (11)

  1.  入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類する第1分類手段と、
     前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する第2分類手段と、
     を備える分類装置。
  2.  前記第2分類手段は、前記複数の区間音声のうち一の区間音声と、前記複数の区間音声各々との類似の程度を夫々示す複数の第1指標値のうち、前記複数の第2区間音声のうち一の第2区間音声と、前記複数の群のうち一の群に含まれる、前記複数の第1区間音声の一又は複数の第1区間音声との類似の程度を示す一又は複数の第1指標値に基づいて、前記一の第2区間音声と前記一の群との類似の程度を示す第2指標値を算出し、
     前記第2分類手段は、前記複数の群各々について算出された複数の第2指標値のうち、最大の第2指標値であって、所定の閾値以上の第2指標値に対応する群に、前記一の第2区間音声を分類する
     請求項1に記載の分類装置。
  3.  前記第2分類手段は、前記複数の区間音声のうち一の区間音声と、前記複数の区間音声各々との類似の程度を夫々示す複数の第1指標値のうち、前記複数の第2区間音声のうち一の第2区間音声と、前記複数の群のうち一の群に含まれる、前記複数の第1区間音声の一又は複数の第1区間音声との類似の程度を示す一又は複数の第1指標値に基づいて、前記一の第2区間音声と前記一の群との類似の程度を示す第2指標値を算出し、
     前記第2分類手段は、前記複数の群各々について算出された複数の第2指標値のいずれもが所定の閾値より小さい場合、前記新たな群に、前記一の第2区間音声を分類する
     請求項1又は2に記載の分類装置。
  4.  前記分類手段は、前記複数の区間音声各々の音質と、基準音質とに基づいて、前記複数の区間音声を、前記複数の第1区間音声と前記複数の第2区間音声とに分類する
     請求項1乃至3のいずれか一項に記載の分類装置。
  5.  前記第2音質は、前記第1音質より低い音質である
     請求項1乃至4のいずれか一項に記載の分類装置。
  6.  コンピュータが、
     入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類し、
     前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する
     分類方法。
  7.  コンピュータに、
     入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類し、
     前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声を分類する
     分類方法を実行させるコンピュータプログラムが記録されている記録媒体。
  8.  第1入力音声が複数の区間に分割されることによって生成された複数の区間音声を、第1音質を有する複数の第1区間音声と、前記第1音質とは異なる第2音質を有する複数の第2区間音声とに分類する第1分類手段と、
     前記複数の第1区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声各々を分類する第2分類手段と、
     前記複数の群及び前記新たな群に基づいて、前記複数の区間音声の話者を示す第1話者情報を生成する第1生成手段と、
     前記複数の区間音声の少なくとも一つの区間音声に基づいて、前記少なくとも一つの区間音声に相当する文字列を示す文字情報を生成する第2生成手段と、
     前記生成された第1話者情報により示される話者と、前記生成された文字情報により示される文字列とを表示する表示手段と、
     を備える情報表示装置。
  9.  前記表示手段は、前記少なくとも一つの区間音声のうち、前記複数の群及び前記新たな群のうち第1群に分類された第1区間音声に該当する区間音声に相当する文字列を、前記生成された第1話者情報により示される話者と対応付けて表示する
     請求項8に記載の情報表示装置。
  10.  前記第1話者情報は、前記複数の群のうち第2群に分類された第1区間音声の話者である第1話者を示すとともに、前記第2群に分類された第2区間音声の話者として、前記第1話者を示唆する文字を示す
     請求項8又は9に記載の情報表示装置。
  11.  前記第1入力音声が複数の区間に分割されることによって生成された複数の区間音声に、第2入力音声が複数の区間に分割されることによって生成された他の複数の区間音声が加えられた場合、
     前記第1分類手段は、前記他の複数の区間音声を、前記第1音質を有する複数の第3区間音声と、前記第2音質を有する複数の第4区間音声とに分類し、
     前記第2分類手段は、前記複数の第1区間音声及び前記複数の第3区間音声に基づく複数の群のいずれか又は前記複数の群とは異なる新たな群に、前記複数の第2区間音声及び前記複数の第4区間音声を分類し、
     前記第1生成手段は、前記複数の群及び前記新たな群に基づいて、前記複数の区間音声及び前記他の複数の区間音声各々の話者を示す第2話者情報を生成し、
     前記生成された文字情報により示される一の文字列に相当する一の区間音声の話者が、前記第1話者情報により示される話者と、前記第2話者情報により示される話者とで異なる場合、前記表示手段は、前記第2話者情報が生成された際に、前記一の文字列に対応づけられる話者を、前記第1話者情報により示される話者から、前記第2話者情報により示される話者に変更する
     請求項8乃至10のいずれか一項に記載の情報表示装置。
PCT/JP2023/022279 2023-06-15 2023-06-15 分類装置、分類方法、記録媒体及び情報表示装置 Ceased WO2024257308A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
PCT/JP2023/022279 WO2024257308A1 (ja) 2023-06-15 2023-06-15 分類装置、分類方法、記録媒体及び情報表示装置
JP2025527160A JPWO2024257308A1 (ja) 2023-06-15 2023-06-15

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2023/022279 WO2024257308A1 (ja) 2023-06-15 2023-06-15 分類装置、分類方法、記録媒体及び情報表示装置

Publications (1)

Publication Number Publication Date
WO2024257308A1 true WO2024257308A1 (ja) 2024-12-19

Family

ID=93851708

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2023/022279 Ceased WO2024257308A1 (ja) 2023-06-15 2023-06-15 分類装置、分類方法、記録媒体及び情報表示装置

Country Status (2)

Country Link
JP (1) JPWO2024257308A1 (ja)
WO (1) WO2024257308A1 (ja)

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008175955A (ja) * 2007-01-17 2008-07-31 Toshiba Corp インデキシング装置、方法及びプログラム
WO2018163279A1 (ja) * 2017-03-07 2018-09-13 日本電気株式会社 音声処理装置、音声処理方法、および音声処理プログラム
JP2019008131A (ja) * 2017-06-23 2019-01-17 日本電信電話株式会社 話者判定装置、話者判定情報生成方法、プログラム
WO2022113218A1 (ja) * 2020-11-25 2022-06-02 日本電信電話株式会社 話者認識方法、話者認識装置および話者認識プログラム
JP2022109867A (ja) * 2021-01-15 2022-07-28 ネイバー コーポレーション 話者識別を結合した話者ダイアライゼーション方法、システム、およびコンピュータプログラム
JP2022133118A (ja) * 2021-03-01 2022-09-13 パナソニックIpマネジメント株式会社 発話分類装置および発話分類方法

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008175955A (ja) * 2007-01-17 2008-07-31 Toshiba Corp インデキシング装置、方法及びプログラム
WO2018163279A1 (ja) * 2017-03-07 2018-09-13 日本電気株式会社 音声処理装置、音声処理方法、および音声処理プログラム
JP2019008131A (ja) * 2017-06-23 2019-01-17 日本電信電話株式会社 話者判定装置、話者判定情報生成方法、プログラム
WO2022113218A1 (ja) * 2020-11-25 2022-06-02 日本電信電話株式会社 話者認識方法、話者認識装置および話者認識プログラム
JP2022109867A (ja) * 2021-01-15 2022-07-28 ネイバー コーポレーション 話者識別を結合した話者ダイアライゼーション方法、システム、およびコンピュータプログラム
JP2022133118A (ja) * 2021-03-01 2022-09-13 パナソニックIpマネジメント株式会社 発話分類装置および発話分類方法

Also Published As

Publication number Publication date
JPWO2024257308A1 (ja) 2024-12-19

Similar Documents

Publication Publication Date Title
JP6556575B2 (ja) 音声処理装置、音声処理方法及び音声処理プログラム
US9934785B1 (en) Identification of taste attributes from an audio signal
CN107492382B (zh) 基于神经网络的声纹信息提取方法及装置
CN102831891B (zh) 一种语音数据处理方法及系统
US10068588B2 (en) Real-time emotion recognition from audio signals
US10224030B1 (en) Dynamic gazetteers for personalized entity recognition
US20220148576A1 (en) Electronic device and control method
CN112912897A (zh) 声音分类系统
JP6246636B2 (ja) パターン識別装置、パターン識別方法およびプログラム
US9251808B2 (en) Apparatus and method for clustering speakers, and a non-transitory computer readable medium thereof
KR20150144031A (ko) 음성 인식을 이용하는 사용자 인터페이스 제공 방법 및 사용자 인터페이스 제공 장치
JPWO2018163279A1 (ja) 音声処理装置、音声処理方法、および音声処理プログラム
CN115512692B (zh) 语音识别方法、装置、设备及存储介质
JP7363107B2 (ja) 発想支援装置、発想支援システム及びプログラム
JP7266390B2 (ja) 行動識別方法、行動識別装置、行動識別プログラム、機械学習方法、機械学習装置及び機械学習プログラム
Alshammri IoT‐Based Voice‐Controlled Smart Homes with Source Separation Based on Deep Learning
JPWO2019244298A1 (ja) 属性識別装置、属性識別方法、およびプログラム
CN120851152A (zh) 用于经由自动多模态图构造的基于知识的音频-文本建模的系统和方法
KR102241436B1 (ko) 임의의 오디오에 사용된 악기를 판단하고 분류하기 위한 학습 방법 및 테스트 방법, 이를 이용한 학습 장치 및 테스트 장치
JP6784255B2 (ja) 音声処理装置、音声処理システム、音声処理方法、およびプログラム
Al Mojaly et al. Detection and classification of voice pathology using feature selection
KR102226427B1 (ko) 호칭 결정 장치, 이를 포함하는 대화 서비스 제공 시스템, 호칭 결정을 위한 단말 장치 및 호칭 결정 방법
WO2024257307A1 (ja) 音声処理装置、音声処理方法、記録媒体及び情報表示装置
KR101737083B1 (ko) 음성 활동 감지 방법 및 장치
Gosztolya et al. Ensemble Bag-of-Audio-Words representation improves paralinguistic classification accuracy

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23941612

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2025527160

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 2025527160

Country of ref document: JP

NENP Non-entry into the national phase

Ref country code: DE