WO2022123742A1 - 話者ダイアライゼーション方法、話者ダイアライゼーション装置および話者ダイアライゼーションプログラム - Google Patents
話者ダイアライゼーション方法、話者ダイアライゼーション装置および話者ダイアライゼーションプログラム Download PDFInfo
- Publication number
- WO2022123742A1 WO2022123742A1 PCT/JP2020/046117 JP2020046117W WO2022123742A1 WO 2022123742 A1 WO2022123742 A1 WO 2022123742A1 JP 2020046117 W JP2020046117 W JP 2020046117W WO 2022123742 A1 WO2022123742 A1 WO 2022123742A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- speaker
- frame
- vector
- label
- learning
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/18—Artificial neural networks; Connectionist approaches
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0272—Voice signal separating
- G10L21/028—Voice signal separating using properties of sound source
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/02—Preprocessing operations, e.g. segment selection; Pattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal components; Feature selection or extraction
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/04—Training, enrolment or model building
Definitions
- the present invention relates to a speaker dialing method, a speaker dialing device, and a speaker dialing program.
- EEND End-to-End Neural Diarization
- the acoustic signal is divided into frames, and a speaker label indicating whether or not a specific speaker exists in the frame is estimated for each frame from the acoustic features extracted from each frame.
- the speaker label for each frame is an S-dimensional vector, and in the frame, 1 when a speaker is speaking and 1 when not speaking. It becomes 0. That is, in EEND, speaker dialization is realized by performing multi-label binary classification of the number of speakers.
- the EEND model used in EEND to estimate the speaker label sequence for each frame is a model based on deep learning composed of layers capable of backpropagation of errors, and the speaker label sequence for each frame is changed at once from the acoustic feature sequence. It can be estimated by through.
- the EEND model includes an RNN (Recurrent Neural Network) layer for time-series modeling. As a result, in EEND, it is possible to estimate the speaker label for each frame by using the acoustic features of not only the frame but also the surrounding frames. Bidirectional LSTM (Long Short-Term Memory) -RNN or Transformer Encoder is used for this RNN layer.
- Non-Patent Document 2 describes RNN Transducer. Further, Non-Patent Document 3 describes acoustic features.
- the present invention has been made in view of the above, and an object of the present invention is to perform online speaker dialiation.
- the speaker dialification method uses a series of acoustic features for each frame of the latest acoustic signal to represent the speaker characteristics of each frame.
- a model for estimating the speaker label of the speaker vector of each frame is learned by using the extraction process for extracting the speaker vector and the speaker label representing the speaker of the speaker vector estimated as the speaker vector. It is characterized by including a learning process generated by.
- FIG. 1 is a diagram for explaining an outline of a speaker dialyrating device.
- FIG. 2 is a schematic diagram illustrating a schematic configuration of a speaker dialyration device.
- FIG. 3 is a diagram for explaining the processing of the speaker dialyration device.
- FIG. 4 is a flowchart showing a speaker dialization processing procedure.
- FIG. 5 is a flowchart showing a speaker dialization processing procedure.
- FIG. 6 is a diagram illustrating a computer that executes a speaker dialyration program.
- FIG. 1 is a diagram for explaining an outline of a speaker dialyrating device.
- the EEND model (online EEND model) of the speaker dialylation device of the present embodiment uses a series of acoustic features for each frame of the latest acoustic signal as an input, and features of the speaker of the latest frame.
- An online EEND model 14a that outputs a speaker vector representing the above is constructed. Specifically, the online EEND model 14a estimates the speaker label of the t-frame using the acoustic characteristics of each frame from the current t-frame to the (t-N) frame continuously traced back (t-N). ..
- This online EEND model 14a has a speaker feature extraction block, a speaker feature update block, and a speaker label estimation block.
- the speaker feature extraction block extracts a speaker vector representing the speaker characteristics of the t-frame using the acoustic features of each frame from the (t-N) th frame to the t-frame.
- the speaker feature extraction block includes, but is not limited to, a Linear (fully connected) layer and an RNN layer, and for example, an input vector is averaged instead of the RNN layer. Layers may be included.
- the speaker feature update block stores the speaker vector in the t-frame and the estimated value of the speaker label estimated by the speaker label estimation block described later with respect to this speaker vector by vector coupling. Further, the speaker feature update block outputs a speaker vector containing information for identifying the speaker as a memorized speaker vector for the input of the vector obtained by vector-coupling the memorized speaker vector and the estimated value of the speaker label. Update the parameters of the model to be used.
- the model includes a Linear (fully connected) layer and an RNN layer.
- the speaker label estimation block outputs the estimated value of the speaker label at the t-frame using the speaker vector and the memory speaker vector.
- the speaker label estimation block includes a Linear (fully connected) layer and a sigmoid layer.
- the speaker dialyration device estimates the speaker label, for example, by determining the estimated value of the output speaker label as a threshold value.
- the speaker dialyration device estimates the speaker label frame by frame using the online EEND model 14a having an autoregressive structure. This allows the speaker dialyration device to estimate the speaker label while updating the memory speaker vector each time a frame is input. Therefore, it is possible to realize online speaker dialification.
- FIG. 2 is a schematic diagram illustrating a schematic configuration of a speaker dialyration device. Further, FIG. 3 is a diagram for explaining the processing of the speaker dialyration device.
- the speaker dialyration device 10 of the present embodiment is realized by a general-purpose computer such as a personal computer, and has an input unit 11, an output unit 12, a communication control unit 13, a storage unit 14, and a control. A unit 15 is provided.
- the input unit 11 is realized by using an input device such as a keyboard or a mouse, and inputs various instruction information such as processing start to the control unit 15 in response to an input operation by the practitioner.
- the output unit 12 is realized by a display device such as a liquid crystal display, a printing device such as a printer, an information communication device, and the like.
- the communication control unit 13 is realized by a NIC (Network Interface Card) or the like, and controls communication via a network between an external device such as a server or a device for acquiring an acoustic signal and the control unit 15.
- NIC Network Interface Card
- the storage unit 14 is realized by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory (Flash Memory), or a storage device such as a hard disk or an optical disk.
- the storage unit 14 may be configured to communicate with the control unit 15 via the communication control unit 13.
- the storage unit 14 stores, for example, the online EEND model 14a used for the speaker dialylation process described later.
- the control unit 15 is realized by using a CPU (Central Processing Unit), an NP (Network Processor), an FPGA (Field Programmable Gate Array), etc., and executes a processing program stored in a memory.
- the control unit 15 serves as an acoustic feature extraction unit 15a, a speaker vector extraction unit 15b, a speaker label generation unit 15c, a learning unit 15d, an estimation unit 15e, and an utterance section estimation unit 15f. Function.
- these functional units may be implemented in different hardware.
- the learning unit 15d may be mounted as a learning device
- the estimation unit 15e may be mounted as an estimation device.
- the control unit 15 may include other functional units.
- the acoustic feature extraction unit 15a extracts the acoustic features for each frame of the acoustic signal including the utterance of the speaker. For example, the acoustic feature extraction unit 15a receives an input of an acoustic signal via the input unit 11 or from a device or the like that acquires an acoustic signal via the communication control unit 13. Further, the acoustic feature extraction unit 15a divides the acoustic signal into frames, extracts the acoustic feature vector by performing discrete Fourier transform or filter bank multiplication on the signal from each frame, and combines the acoustic signals in the frame direction. Output the feature series. In this embodiment, the frame length is 25 ms and the frame shift width is 10 ms.
- the acoustic feature vector is, for example, a 24-dimensional MFCC (Mel Frequency Cepstrum Coefficient), but is not limited to this, and may be, for example, an acoustic feature amount for each other frame such as a mel filter bank output.
- MFCC Mel Frequency Cepstrum Coefficient
- the speaker vector extraction unit 15b extracts a speaker vector representing the speaker characteristics of each frame by using the acoustic feature series for each frame of the latest acoustic signal. Specifically, the speaker vector extraction unit 15b generates a speaker vector by inputting the acoustic feature series acquired from the acoustic feature extraction unit 15a into the speaker feature extraction block shown in FIG.
- the speaker vector extraction unit 15b may be included in the learning unit 15d and the estimation unit 15e, which will be described later.
- FIG. 3 which will be described later, shows an example in which the learning unit 15d and the estimation unit 15e process the speaker vector extraction unit 15b.
- the speaker label generation unit 15c generates a speaker label for each frame using the acoustic feature series. Specifically, as shown in FIG. 3, the speaker label generation unit 15c generates a speaker label for each frame by using the acoustic feature series and the correct answer label of the speaker's utterance section. As a result, a set of the acoustic feature series and the speaker label for each frame is generated as the teacher data used for the processing of the learning unit 15d described later.
- the learning unit 15d generates an online EEND model 14a that estimates the speaker label of the speaker vector of each frame by learning using the speaker vector and the speaker label representing the speaker of the estimated speaker vector. do. Specifically, as shown in FIG. 3, the learning unit 15d learns the online EEND model 14a by using the set of the acoustic feature sequence and the speaker label for each frame as teacher data.
- the online EEND model 14a is composed of a plurality of layers including the RNN layer as shown in FIG.
- a unidirectional LSTM-RNN is applied as the RNN layer.
- N 10
- t—N is a negative value
- the acoustic feature vector is a zero vector.
- the online EEND model 14a outputs the posterior probability of the speaker label for each frame of T ⁇ S dimension.
- the learning unit 15d uses the posterior probability of the speaker label for each frame and the multi-label binary cross entropy with the speaker label for each frame as a loss function, and uses the error back propagation method to obtain the parameters of each layer of the online EEND model 14a. Perform optimization.
- the learning unit 15d uses an online optimization algorithm using a stochastic gradient descent method for optimizing the parameters.
- the learning unit 15d is the t-frame extracted by the speaker vector extraction unit 15b, which is a speaker feature extraction block, using the acoustic features of each of the (tN) th to t-frames of the teacher data.
- the speaker vector of the above and the estimated value of the speaker label estimated by the speaker label estimation block with respect to this speaker vector are connected in a vector and stored.
- the learning unit 15d inputs a vector obtained by vector-coupling the stored speaker vector and the estimated value of the speaker label into the speaker feature update block, and outputs a stored speaker vector including information for identifying the speaker. Update model parameters. Further, the learning unit 15d inputs the speaker vector of the t-frame and the stored speaker vector into the speaker label estimation block, and updates the parameters of the model that outputs the estimated value of the speaker label of the t-frame.
- the learning unit 15d generates the online EEND model 14a by using a plurality of stored combinations of the speaker vector and the speaker label of the estimated speaker vector. This makes it possible to estimate the speaker label while updating the memory speaker vector each time a frame is input.
- the estimation unit 15e estimates the speaker label for each frame of the acoustic signal using the generated online EEND model 14a. Specifically, as shown in FIG. 3, in the estimation unit 15e, each frame from the current t-frame in which the speaker vector extraction unit 15b is continuously traced back (t-N) from the current t-frame of the acoustic feature series. The speaker vector of the t-frame extracted using the acoustic feature of is forward-propagated to the online EEND model 14a.
- the speaker label posterior probability estimate of the speaker label for each frame of the acoustic feature series is output by sequentially propagating from the first frame of the acoustic feature series. do.
- the utterance section estimation unit 15f estimates the utterance section of the speaker in the acoustic signal by using the output speaker label posterior probability. Specifically, the utterance section estimation unit 15f estimates the speaker label using the moving averages of a plurality of frames. That is, the utterance section estimation unit 15f first calculates a moving average of the length 6 of the own frame and the 5 frames immediately before it with respect to the speaker label posterior probability for each frame. This makes it possible to prevent erroneous detection of an unrealistic short utterance section such as an utterance with only one frame.
- the utterance section estimation unit 15f estimates that the frame is the utterance section of the speaker of the dimension when the calculated moving average value is larger than 0.5. Further, the utterance section estimation unit 15f considers a continuous utterance section frame group as one utterance for each speaker, and back-calculates the start time and end time of the utterance section up to a predetermined time from the frame. As a result, it is possible to obtain the utterance start time and the utterance end time up to a predetermined time for each utterance of each speaker.
- FIG. 4 shows a learning processing procedure.
- the flowchart of FIG. 4 is started, for example, at the timing when there is an input instructing the start of the learning process.
- the acoustic feature extraction unit 15a extracts the acoustic features for each frame of the acoustic signal including the speaker's utterance, and outputs the acoustic feature series (step S1).
- the speaker vector extraction unit 15b extracts a speaker vector representing the speaker characteristics of each frame by using the acoustic feature series for each frame of the latest acoustic signal. (Step S2).
- the learning unit 15d has a self-regression structure and estimates the speaker label of the speaker vector of each frame by using the speaker vector and the speaker label representing the speaker of the estimated speaker vector.
- the online EEND model 14a is generated by learning (step S3). As a result, a series of learning processes are completed.
- FIG. 5 shows an estimation processing procedure.
- the flowchart of FIG. 5 is started, for example, at the timing when there is an input instructing the start of the estimation process.
- the acoustic feature extraction unit 15a extracts the acoustic features for each frame of the acoustic signal including the speaker's utterance, and outputs the acoustic feature series (step S1).
- the speaker vector extraction unit 15b extracts a speaker vector representing the speaker characteristics of each frame by using the acoustic feature series for each frame of the latest acoustic signal (step S2).
- the estimation unit 15e estimates the speaker label for each frame of the acoustic signal using the generated online EEND model 14a (step S4). Specifically, the estimation unit 15e outputs the speaker label posterior probability (estimated value of the speaker label) for each frame of the acoustic feature series.
- the utterance section estimation unit 15f estimates the utterance section of the speaker in the acoustic signal using the output speaker label posterior probability (step S5). This completes a series of estimation processes.
- the speaker vector extraction unit 15b expresses the speaker characteristics of each frame by using the acoustic feature series for each frame of the latest acoustic signal. Extract the speaker vector. Further, the learning unit 15d learns the online EEND model 14a that estimates the speaker label of the speaker vector of each frame by using the speaker vector and the speaker label representing the speaker of the estimated speaker vector. Generated by.
- the speaker dialyration device 10 can estimate the speaker label each time a frame is input by the online EEND model 14a having an autoregressive structure. Therefore, it is possible to realize online speaker dialification.
- the learning unit 15d generates an online EEND model 14a by using a plurality of stored combinations of the speaker vector and the speaker label of the estimated speaker vector. This allows the speaker dialyration device 10 to estimate the speaker label while updating the memory speaker vector each time a frame is input. Therefore, online speaker dialiation can be realized with higher accuracy.
- the estimation unit 15e estimates the speaker label for each frame of the acoustic signal using the generated online EEND model 14a. This enables online speaker dialification.
- the utterance section estimation unit 15f estimates the speaker label using the moving average of a plurality of frames. This makes it possible to prevent erroneous detection of an unrealistic short utterance section.
- the speaker dialing device 10 can be implemented by installing a speaker dialing program that executes the above-mentioned speaker dialing process as package software or online software on a desired computer.
- the information processing device can be made to function as the speaker dialyration device 10.
- the information processing device includes a smartphone, a mobile communication terminal such as a mobile phone and a PHS (Personal Handyphone System), and a slate terminal such as a PDA (Personal Digital Assistant).
- the function of the speaker dialyration device 10 may be implemented in the cloud server.
- FIG. 6 is a diagram showing an example of a computer that executes a speaker dialyration program.
- the computer 1000 has, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. Each of these parts is connected by a bus 1080.
- the memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012.
- the ROM 1011 stores, for example, a boot program such as a BIOS (Basic Input Output System).
- BIOS Basic Input Output System
- the hard disk drive interface 1030 is connected to the hard disk drive 1031.
- the disk drive interface 1040 is connected to the disk drive 1041.
- a removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1041.
- a mouse 1051 and a keyboard 1052 are connected to the serial port interface 1050.
- a display 1061 is connected to the video adapter 1060.
- the hard disk drive 1031 stores, for example, the OS 1091, the application program 1092, the program module 1093, and the program data 1094. Each of the information described in the above embodiment is stored in, for example, the hard disk drive 1031 or the memory 1010.
- the speaker dialyration program is stored in the hard disk drive 1031 as, for example, a program module 1093 in which a command executed by the computer 1000 is described.
- the program module 1093 in which each process executed by the speaker dialyration device 10 described in the above embodiment is described is stored in the hard disk drive 1031.
- the data used for information processing by the speaker dialyration program is stored as program data 1094 in, for example, the hard disk drive 1031.
- the CPU 1020 reads the program module 1093 and the program data 1094 stored in the hard disk drive 1031 into the RAM 1012 as needed, and executes each of the above-mentioned procedures.
- the program module 1093 and the program data 1094 related to the speaker dialyration program are not limited to the case where they are stored in the hard disk drive 1031. For example, they are stored in a removable storage medium and are stored in the CPU 1020 via the disk drive 1041 or the like. May be read by.
- the program module 1093 and the program data 1094 related to the speaker dialyration program are stored in another computer connected via a network such as LAN (Local Area Network) or WAN (Wide Area Network), and are stored in the network interface 1070. It may be read out by the CPU 1020 via.
- LAN Local Area Network
- WAN Wide Area Network
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Acoustics & Sound (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Signal Processing (AREA)
- Quality & Reliability (AREA)
- Computational Linguistics (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Circuit For Audible Band Transducer (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
Description
図1は、話者ダイアライゼーション装置の概要を説明するための図である。図1に示すように、本実施形態の話者ダイアライゼーション装置のEENDモデル(オンラインEENDモデル)は、直近の音響信号のフレームごとの音響特徴の系列を入力として、最新のフレームの話者の特徴を表す話者ベクトルを出力するオンラインEENDモデル14aを構築する。具体的には、オンラインEENDモデル14aは、現在のtフレーム目から連続して遡った(t-N)フレーム目までの各フレームの音響特徴を用いて、tフレーム目の話者ラベルを推定する。
図2は、話者ダイアライゼーション装置の概略構成を例示する模式図である。また、図3は、話者ダイアライゼーション装置の処理を説明するための図である。まず、図2に例示するように、本実施形態の話者ダイアライゼーション装置10は、パソコン等の汎用コンピュータで実現され、入力部11、出力部12、通信制御部13、記憶部14、および制御部15を備える。
次に、話者ダイアライゼーション装置10による話者ダイアライゼーション処理について説明する。図4よび図5は、話者ダイアライゼーション処理手順を示すフローチャートである。本実施形態の話者ダイアライゼーション処理は、学習処理と推定処理とを含む。まず、図4は、学習処理手順を示す。図4のフローチャートは、例えば、学習処理の開始を指示する入力があったタイミングで開始される。
上記実施形態に係る話者ダイアライゼーション装置10が実行する処理をコンピュータが実行可能な言語で記述したプログラムを作成することもできる。一実施形態として、話者ダイアライゼーション装置10は、パッケージソフトウェアやオンラインソフトウェアとして上記の話者ダイアライゼーション処理を実行する話者ダイアライゼーションプログラムを所望のコンピュータにインストールさせることによって実装できる。例えば、上記の話者ダイアライゼーションプログラムを情報処理装置に実行させることにより、情報処理装置を話者ダイアライゼーション装置10として機能させることができる。また、その他にも、情報処理装置にはスマートフォン、携帯電話機やPHS(Personal Handyphone System)等の移動体通信端末、さらには、PDA(Personal Digital Assistant)等のスレート端末等がその範疇に含まれる。また、話者ダイアライゼーション装置10の機能を、クラウドサーバに実装してもよい。
11 入力部
12 出力部
13 通信制御部
14 記憶部
14a オンラインEENDモデル
15 制御部
15a 音響特徴抽出部
15b 話者ベクトル抽出部
15c 話者ラベル生成部
15d 学習部
15e 推定部
15f 発話区間推定部
Claims (6)
- 話者ダイアライゼーション装置が実行する話者ダイアライゼーション方法であって、
直近の音響信号のフレームごとの音響特徴の系列を用いて、各フレームの話者特徴を表す話者ベクトルを抽出する抽出工程と、
前記話者ベクトルと推定された該話者ベクトルの話者を表す話者ラベルとを用いて、各フレームの話者ベクトルの話者ラベルを推定するモデルを学習により生成する学習工程と、
を含んだことを特徴とする話者ダイアライゼーション方法。 - 前記学習工程は、前記話者ベクトルと推定された該話者ベクトルの話者ラベルとの記憶された複数の組み合わせを用いて、前記モデルを生成することを特徴とする請求項1に記載の話者ダイアライゼーション方法。
- 生成された前記モデルを用いて、音響信号のフレームごとの話者ラベルを推定する推定工程を、さらに含んだことを特徴とする請求項1に記載の話者ダイアライゼーション方法。
- 前記推定工程は、複数のフレームの移動平均を用いて、前記話者ラベルを推定することを特徴とする請求項3に記載の話者ダイアライゼーション方法。
- 直近の音響信号のフレームごとの音響特徴の系列を用いて、各フレームの話者特徴を表す話者ベクトルを抽出する抽出部と、
前記話者ベクトルと推定された該話者ベクトルの話者を表す話者ラベルとを用いて、各フレームの話者ベクトルの話者ラベルを推定するモデルを学習により生成する学習部と、
を有することを特徴とする話者ダイアライゼーション装置。 - 直近の音響信号のフレームごとの音響特徴の系列を用いて、各フレームの話者特徴を表す話者ベクトルを抽出する抽出ステップと、
前記話者ベクトルと推定された該話者ベクトルの話者を表す話者ラベルとを用いて、各フレームの話者ベクトルの話者ラベルを推定するモデルを学習により生成する学習ステップと、
をコンピュータに実行させるための話者ダイアライゼーションプログラム。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2020/046117 WO2022123742A1 (ja) | 2020-12-10 | 2020-12-10 | 話者ダイアライゼーション方法、話者ダイアライゼーション装置および話者ダイアライゼーションプログラム |
| US18/266,166 US20240038255A1 (en) | 2020-12-10 | 2020-12-10 | Speaker diarization method, speaker diarization device, and speaker diarization program |
| JP2022567984A JP7505582B2 (ja) | 2020-12-10 | 2020-12-10 | 話者ダイアライゼーション方法、話者ダイアライゼーション装置および話者ダイアライゼーションプログラム |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2020/046117 WO2022123742A1 (ja) | 2020-12-10 | 2020-12-10 | 話者ダイアライゼーション方法、話者ダイアライゼーション装置および話者ダイアライゼーションプログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022123742A1 true WO2022123742A1 (ja) | 2022-06-16 |
Family
ID=81973450
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2020/046117 Ceased WO2022123742A1 (ja) | 2020-12-10 | 2020-12-10 | 話者ダイアライゼーション方法、話者ダイアライゼーション装置および話者ダイアライゼーションプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20240038255A1 (ja) |
| JP (1) | JP7505582B2 (ja) |
| WO (1) | WO2022123742A1 (ja) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR102815144B1 (ko) * | 2022-04-29 | 2025-06-04 | 광주과학기술원 | 보조 손실을 이용한 단대단 화자 분리 시스템 및 방법 |
| US12198677B2 (en) * | 2022-05-27 | 2025-01-14 | Tencent America LLC | Techniques for end-to-end speaker diarization with generalized neural speaker clustering |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2019527370A (ja) * | 2017-06-13 | 2019-09-26 | ベイジン ディディ インフィニティ テクノロジー アンド ディベロップメント カンパニー リミティッド | 話者照合の方法、装置、及びシステム |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CA3033675C (en) * | 2016-07-11 | 2022-11-15 | FTR Labs Pty Ltd | Method and system for automatically diarising a sound recording |
| WO2019209569A1 (en) * | 2018-04-23 | 2019-10-31 | Google Llc | Speaker diarization using an end-to-end model |
| US11031017B2 (en) * | 2019-01-08 | 2021-06-08 | Google Llc | Fully supervised speaker diarization |
-
2020
- 2020-12-10 WO PCT/JP2020/046117 patent/WO2022123742A1/ja not_active Ceased
- 2020-12-10 US US18/266,166 patent/US20240038255A1/en not_active Abandoned
- 2020-12-10 JP JP2022567984A patent/JP7505582B2/ja active Active
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2019527370A (ja) * | 2017-06-13 | 2019-09-26 | ベイジン ディディ インフィニティ テクノロジー アンド ディベロップメント カンパニー リミティッド | 話者照合の方法、装置、及びシステム |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2022123742A1 (ja) | 2022-06-16 |
| US20240038255A1 (en) | 2024-02-01 |
| JP7505582B2 (ja) | 2024-06-25 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN108630190B (zh) | 用于生成语音合成模型的方法和装置 | |
| JP6712642B2 (ja) | モデル学習装置、その方法、及びプログラム | |
| CN110689879B (zh) | 端到端语音转写模型的训练方法、系统、装置 | |
| US11264044B2 (en) | Acoustic model training method, speech recognition method, acoustic model training apparatus, speech recognition apparatus, acoustic model training program, and speech recognition program | |
| US11551708B2 (en) | Label generation device, model learning device, emotion recognition apparatus, methods therefor, program, and recording medium | |
| CN107680597B (zh) | 语音识别方法、装置、设备以及计算机可读存储介质 | |
| US11270686B2 (en) | Deep language and acoustic modeling convergence and cross training | |
| JP2018128659A (ja) | 音声対話システム、音声対話方法、および音声対話システムを適合させる方法 | |
| JP2012037619A (ja) | 話者適応化装置、話者適応化方法および話者適応化用プログラム | |
| CN112259089A (zh) | 语音识别方法及装置 | |
| KR20200044388A (ko) | 음성을 인식하는 장치 및 방법, 음성 인식 모델을 트레이닝하는 장치 및 방법 | |
| JP2020042257A (ja) | 音声認識方法及び装置 | |
| US12057105B2 (en) | Speech recognition device, speech recognition method, and program | |
| CN111797220A (zh) | 对话生成方法、装置、计算机设备和存储介质 | |
| JP2020034683A (ja) | 音声認識装置、音声認識プログラムおよび音声認識方法 | |
| JP2021026050A (ja) | 音声認識システム、情報処理装置、音声認識方法、プログラム | |
| CN111557010A (zh) | 学习装置和方法以及程序 | |
| JP7505584B2 (ja) | 話者ダイアライゼーション方法、話者ダイアライゼーション装置および話者ダイアライゼーションプログラム | |
| CN114333790B (zh) | 数据处理方法、装置、设备、存储介质及程序产品 | |
| KR101120765B1 (ko) | 스위칭 상태 스페이스 모델과의 멀티모덜 변동 추정을이용한 스피치 인식 방법 | |
| JP7505582B2 (ja) | 話者ダイアライゼーション方法、話者ダイアライゼーション装置および話者ダイアライゼーションプログラム | |
| JP7700801B2 (ja) | 話者認識方法、話者認識装置および話者認識プログラム | |
| JP7212596B2 (ja) | 学習装置、学習方法および学習プログラム | |
| JP2016122110A (ja) | 音響スコア算出装置、その方法及びプログラム | |
| JP2018128500A (ja) | 形成装置、形成方法および形成プログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20965124 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2022567984 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18266166 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20965124 Country of ref document: EP Kind code of ref document: A1 |