WO2020143263A1 - 一种基于语音样本特征空间轨迹的说话人识别方法 - Google Patents
一种基于语音样本特征空间轨迹的说话人识别方法 Download PDFInfo
- Publication number
- WO2020143263A1 WO2020143263A1 PCT/CN2019/111530 CN2019111530W WO2020143263A1 WO 2020143263 A1 WO2020143263 A1 WO 2020143263A1 CN 2019111530 W CN2019111530 W CN 2019111530W WO 2020143263 A1 WO2020143263 A1 WO 2020143263A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- speech
- speaker
- feature space
- sample
- voice
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/04—Training, enrolment or model building
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/02—Preprocessing operations, e.g. segment selection; Pattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal components; Feature selection or extraction
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/06—Decision making techniques; Pattern matching strategies
- G10L17/08—Use of distortion metrics or a particular distance between probe pattern and reference templates
Definitions
- the invention relates to the field of biometric recognition, and in particular to a speaker recognition method based on the trajectory of the feature space of speech samples.
- audio classification or audio recognition is the core issue of audio perception.
- audio classification manifests as speaker recognition, audio event recognition, audio Event detection, etc.
- Speaker recognition technology is a kind of identity verification technology---biometric recognition technology.
- Biometric recognition technology is a technology that uses biometrics to automatically identify individuals, including fingerprint recognition, iris recognition, gene recognition, and face recognition. Compared with other identity verification technologies, speaker identification is more convenient, natural, and has lower user intrusiveness.
- Speaker recognition uses voice signals for identity recognition, which has the advantages of natural human-computer interaction, easy extraction of voice signals, and remote recognition.
- the existing speaker recognition system includes two stages: training stage and recognition stage.
- the training stage the system uses the collected speaker speech to build a model for the speaker; in the recognition stage, the system matches the input speech with the speaker model to make a decision.
- the speaker recognition system needs to extract features that can reflect the personality of the speaker from the speech signal, and establish an accurate model to distinguish the difference between the speaker and other speakers.
- audio classification technology there are two main types of audio classification technology, one is to generate statistical models, such as the mixed Gaussian model GMM and the hidden Markov model HMM, and the other is based on deep neural network methods, such as DNN, RNN or LSTM. . No matter what kind of technology, a large number of labeled training samples are required.
- the deep neural network method requires higher sample size.
- the GMM or HMM-based method does not give special consideration to the distinguishing information between different audio classes, nor does it consider the sharing of sample data of different classes, such as: MIT Professor Reynold's paper "Speaker Verification Using Adapted Gaussian Mixture Models” (Digital Signal Processing 10 (2000), 19–41.)
- the method mentioned has a high computational complexity; with the support of large samples, the deep neural network method has shown very good performance, such as the paper of Google "End-to-End Text-Dependent Speaker Verification" (2016 IEEE International Conference on Acoustics, Speech and Processing (ICASSP), 2016, Pages: 5115-5119) uses neural networks to extract features and train speech, but neural networks
- the training requires a lot of labeled speech, and the acquisition cost of a large number of samples is very high, and the deep neural network method lacks explanation, which is quite a black box.
- the purpose of the present invention is to provide a speaker recognition method based on the trajectory of the feature space of the speech sample, in which the speech feature space does not depend on the speaker, text and language, so the construction of the speech feature space can be Use any qualified voice data to achieve the sharing of voice data; and the speaker's voice trajectory can be constructed even with a sample, so a large amount of annotated voice data is not required, which overcomes the need to collect a large amount of annotated voice data in the prior art defect.
- Step 2 construct speaker knowledge: use pure speech samples marked with speaker attributes to obtain their distribution information and movement trajectory information in the speech feature space ⁇ ;
- any pure speech samples can be used without any constraints on speakers and language factors.
- K-means or other clustering methods are used to cluster the speech samples in the feature space
- voice features The scale K of the class identifier used in the space determines the granularity of voice feature spatial expression. The larger K is, the finer the voice feature spatial expression.
- the accuracy of spatial expression is related to the size of the data. The richer the data, the more complete the spatial expression. At the same time, the more targeted the data for constructing the speech feature space, the more precise the spatial expression will be for a particular problem.
- step 2) the pure voice samples with speaker attribute annotations are used to mark the voice feature space.
- the Gaussian distribution g k (m k , U k ) is used as the spatial identifier
- the speaker feature space distribution information Obtained as follows:
- the spatial identifier is represented by a multidimensional Gaussian distribution
- m k represents the mean vector of the kth Gaussian distribution
- U k represents the variance matrix of the kth multidimensional Gaussian distribution
- the decision threshold of the ⁇ neighborhood ⁇ t ⁇ g k
- d tk ⁇ of the feature f t of the speech sample refers to the characteristics of the normal distribution, and 2 ⁇ 3 is selected.
- the present invention has the following advantages and beneficial effects:
- a speaker recognition method based on the trajectory of the voice sample feature space provided by the present invention, in which the establishment of the voice feature space is to cluster a large number of voice features without the need for annotated data, to establish data samples of the voice feature space It can be derived from different speakers. There is no exact requirement for the content of the speaker, the age of the speaker, and the language. It overcomes the problem of the need for a large number of labeled voices in the neural network method, and the data collection established in the voice space is easy to implement.
- a speaker recognition method based on the trajectory of the voice sample feature space provided by the present invention based on the location and trajectory information of the speaker's voice feature in the voice feature space, is different from the signal source generation model method, such as Hidden Markov Model (HMM), etc., the positioning is relative, and the generative model is absolute; compared with the deep neural network method, it is interpretable, and each knowledge data has a certain physical semantics, such as the sample features on the space ⁇
- the correlation degree distribution information of P (p 1 , p 2 , ..., p K ) expresses the range of the active space of the sample (the space represented by the identifier subset corresponding to the non-zero elements), and also expresses in the space Distribution.
- a speaker recognition method based on the trajectory of the feature space of the voice samples provided by the present invention is essentially that the voice features are located in the space. For the voice features of different speakers, the location is established on the established voice feature space and the association is used. Degree to represent the speech feature location information of different speakers, and expresses the distinction between different speakers with less calculation, compared to GMM or HMM that requires a generative model to model each speaker The method has a lower computational complexity.
- a speaker recognition method based on the trajectory of the voice sample feature space provided by the present invention wherein the voice feature space identifier subset is a reference system used to locate the speaker's voice feature, which is a relative relationship and is not strictly related to the sample to be recognized The relationship requires that the feature space is shared, and the established voice feature space can be transferred to other speaker data sets for recognition, for example: a speaker's voice feature space of one language can be used as a voice feature for speaker recognition of another language Space to achieve data sharing.
- FIG. 1 is a schematic flowchart of a speaker recognition method in Embodiment 1 of the present invention.
- FIG. 2 is a flowchart of steps for establishing a voice feature space in Embodiment 1 of the present invention.
- FIG. 3 is a flowchart of steps for generating spatial distribution information and trajectory information of speaker voice features in Embodiment 1 of the present invention.
- FIG. 4 is a flowchart of steps for recognizing speech samples to be recognized in Embodiment 1 of the present invention.
- This embodiment provides a speaker recognition method based on the trajectory of the voice sample feature space.
- the schematic flowchart is shown in FIG. 1 and includes the following three steps:
- FIG. 2 it is a flowchart of steps for establishing a voice feature space in this embodiment.
- Aishell contains a total of 400 speakers.
- K is the number of identifiers of the audio feature space, and the number of identifiers K is selected to be 4096, so as to give a higher precision description to the audio feature space;
- FIG. 3 it is a flowchart of steps for generating speaker feature space distribution information in this embodiment.
- 20 wav files are used to label the speech feature space.
- Target speaker speech sample set Y ⁇ (y 1 ,s 1 ),(y 2 ,s 2 ),...
- the spatial distribution of speaker features is calculated as:
- the registered speech of each speaker in the target speaker set is processed to obtain the speech feature distribution information of each speaker.
- FIG. 4 it is a flowchart of steps for recognizing speech samples in this embodiment.
- the positional correlation between the feature f t and the spatial identifier g k (m k , U k ) is:
- This embodiment provides a speaker recognition method based on the trajectory of the feature space of a voice sample, including the following steps:
- Step 1 Use the voice data of the English corpus timit to establish a voice feature space identifier sub-collection
- Step 2 Use the voice data in the aishell corpus to register the target speaker set, as in Example 1;
- Step 3 Recognize the speech samples to be recognized, as in Embodiment 1.
- the obtained recognition effect has a small gap compared with that in Embodiment 1. It can be proved that the speaker speech feature space of another language can be used as the speech feature space of speaker recognition of another language, and data sharing is realized.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Business, Economics & Management (AREA)
- Game Theory and Decision Science (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Image Analysis (AREA)
Abstract
Description
Claims (7)
- 一种基于语音样本特征空间轨迹的说话人识别方法,其中一个语音样本能够视为语音特征空间的一次运动,具有活动空间和空间中的轨迹特性,其特征在于,所述方法包括以下步骤:步骤1)、构建语音特征空间Ω:利用聚类方法将无标注语音样本在特征空间进行聚类,由聚类所得到的子类数据生成该子类数据的某种表达作为语音特征空间的表达Ω={g k,k=1,2,…,K};步骤2)、构建说话人知识:利用有说话人属性标注的纯净语音样本,获得其在语音特征空间Ω上的分布信息以及运动轨迹信息;步骤3)、说话人识别:对于待识别语音样本,首先获得该样本的语音特征空间分布表达以及轨迹,然后利用说话人语音特征空间分布信息计算样本分布与先验分布的差异以及沿轨迹的累计局部分布差异,作为说话人识别的依据并进行判断。
- 根据权利要求1所述的一种基于语音样本特征空间轨迹的说话人识别方法,其特征在于:步骤1)构建语音特征空间Ω的过程中,能够使用任何纯净语音样本,对说话人、语种因素没有任何约束。
- 根据权利要求1所述的一种基于语音样本特征空间轨迹的说话人识别方法,其特征在于:所述语音特征空间表达Ω={g k,k=1,2,…,K}能够是类数据的分布函数、聚类中心矢量或者生成模型这些具有定位能力的标识,称之为特征空间标识子,语音特征空间所使用的类标识子规模K决定语音特征空间表达粒度,K越大,语音特征空间表达越精细。
- 根据权利要求1所述的一种基于语音样本特征空间轨迹的说话人识别方法,其特征在于:步骤2)中,利用有说话人属性标注的纯净语音样本对语音特征空间进行标注,在采用高斯分布g k(m k,U k)作为空间标识子时,说话人特征空间分布信息按以下方式获得:一、计算语音样本每个特征f t与空间标识子g k(m k,U k)的位置关联度,定义为:式中,空间标识子用多维高斯分布来表示,m k表示第k个高斯分布的均值矢 量,U k表示第k个多维高斯分布的方差矩阵;二、计算说话人样本集与空间标识子g k(m k,U k)的位置关联度的期望值:三、计算说话人特征空间分布为:
- 根据权利要求5所述的一种基于语音样本特征空间轨迹的说话人识别方法,其特征在于:所述语音样本特征f t的δ邻域Ψ t={g k|d tk<δ}的判决门限参考正态分布的特性,选取2<δ<3。
- 根据权利要求5所述的一种基于语音样本特征空间轨迹的说话人识别方法,其特征在于,步骤3)中,语音样本f={f 1,f 2,…,f T}的说话人识别过程包括以下步骤:二、确定语音样本f={f 1,f 2,…,f T}在语音特征空间Ω中的运动轨迹Ψ 1Ψ 2…Ψ T,Ψ t={g k|d tk<δ};
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| SG11202103091XA SG11202103091XA (en) | 2019-01-11 | 2019-10-16 | A Speaker Recognition Method Based on Trajectories in Feature Spaces of Voice Samples |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910027145.3 | 2019-01-11 | ||
| CN201910027145.3A CN109545229B (zh) | 2019-01-11 | 2019-01-11 | 一种基于语音样本特征空间轨迹的说话人识别方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020143263A1 true WO2020143263A1 (zh) | 2020-07-16 |
Family
ID=65835222
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/111530 Ceased WO2020143263A1 (zh) | 2019-01-11 | 2019-10-16 | 一种基于语音样本特征空间轨迹的说话人识别方法 |
Country Status (3)
| Country | Link |
|---|---|
| CN (1) | CN109545229B (zh) |
| SG (1) | SG11202103091XA (zh) |
| WO (1) | WO2020143263A1 (zh) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112487978A (zh) * | 2020-11-30 | 2021-03-12 | 清华珠三角研究院 | 一种视频中说话人定位的方法、装置及计算机存储介质 |
| CN113611285A (zh) * | 2021-09-03 | 2021-11-05 | 哈尔滨理工大学 | 基于层叠双向时序池化的语种识别方法 |
| CN117235435A (zh) * | 2023-11-15 | 2023-12-15 | 世优(北京)科技有限公司 | 确定音频信号损失函数的方法及装置 |
| CN117877527A (zh) * | 2024-02-21 | 2024-04-12 | 国能宁夏供热有限公司 | 一种基于通信设备的语音质量分析技术及分析方法 |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109545229B (zh) * | 2019-01-11 | 2023-04-21 | 华南理工大学 | 一种基于语音样本特征空间轨迹的说话人识别方法 |
| US20220392472A1 (en) * | 2019-09-27 | 2022-12-08 | Nec Corporation | Audio signal processing device, audio signal processing method, and storage medium |
| CN111081261B (zh) * | 2019-12-25 | 2023-04-21 | 华南理工大学 | 一种基于lda的文本无关声纹识别方法 |
| CN111128128B (zh) * | 2019-12-26 | 2023-05-23 | 华南理工大学 | 一种基于互补模型评分融合的语音关键词检测方法 |
| CN111933156B (zh) * | 2020-09-25 | 2021-01-19 | 广州佰锐网络科技有限公司 | 基于多重特征识别的高保真音频处理方法及装置 |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5598507A (en) * | 1994-04-12 | 1997-01-28 | Xerox Corporation | Method of speaker clustering for unknown speakers in conversational audio data |
| CN102024455A (zh) * | 2009-09-10 | 2011-04-20 | 索尼株式会社 | 说话人识别系统及其方法 |
| CN102479511A (zh) * | 2010-11-23 | 2012-05-30 | 盛乐信息技术(上海)有限公司 | 一种大规模声纹认证方法及其系统 |
| CN105845141A (zh) * | 2016-03-23 | 2016-08-10 | 广州势必可赢网络科技有限公司 | 基于信道鲁棒的说话人确认模型及说话人确认方法和装置 |
| US20180342250A1 (en) * | 2017-05-24 | 2018-11-29 | AffectLayer, Inc. | Automatic speaker identification in calls |
| CN109065059A (zh) * | 2018-09-26 | 2018-12-21 | 新巴特(安徽)智能科技有限公司 | 用音频特征主成分建立的语音群集来识别说话人的方法 |
| CN109065028A (zh) * | 2018-06-11 | 2018-12-21 | 平安科技(深圳)有限公司 | 说话人聚类方法、装置、计算机设备及存储介质 |
| CN109545229A (zh) * | 2019-01-11 | 2019-03-29 | 华南理工大学 | 一种基于语音样本特征空间轨迹的说话人识别方法 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6067517A (en) * | 1996-02-02 | 2000-05-23 | International Business Machines Corporation | Transcription of speech data with segments from acoustically dissimilar environments |
| CN1302456C (zh) * | 2005-04-01 | 2007-02-28 | 郑方 | 一种声纹识别方法 |
| JP4901657B2 (ja) * | 2007-09-05 | 2012-03-21 | 日本電信電話株式会社 | 音声認識装置、その方法、そのプログラム、その記録媒体 |
-
2019
- 2019-01-11 CN CN201910027145.3A patent/CN109545229B/zh active Active
- 2019-10-16 SG SG11202103091XA patent/SG11202103091XA/en unknown
- 2019-10-16 WO PCT/CN2019/111530 patent/WO2020143263A1/zh not_active Ceased
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5598507A (en) * | 1994-04-12 | 1997-01-28 | Xerox Corporation | Method of speaker clustering for unknown speakers in conversational audio data |
| CN102024455A (zh) * | 2009-09-10 | 2011-04-20 | 索尼株式会社 | 说话人识别系统及其方法 |
| CN102479511A (zh) * | 2010-11-23 | 2012-05-30 | 盛乐信息技术(上海)有限公司 | 一种大规模声纹认证方法及其系统 |
| CN105845141A (zh) * | 2016-03-23 | 2016-08-10 | 广州势必可赢网络科技有限公司 | 基于信道鲁棒的说话人确认模型及说话人确认方法和装置 |
| US20180342250A1 (en) * | 2017-05-24 | 2018-11-29 | AffectLayer, Inc. | Automatic speaker identification in calls |
| CN109065028A (zh) * | 2018-06-11 | 2018-12-21 | 平安科技(深圳)有限公司 | 说话人聚类方法、装置、计算机设备及存储介质 |
| CN109065059A (zh) * | 2018-09-26 | 2018-12-21 | 新巴特(安徽)智能科技有限公司 | 用音频特征主成分建立的语音群集来识别说话人的方法 |
| CN109545229A (zh) * | 2019-01-11 | 2019-03-29 | 华南理工大学 | 一种基于语音样本特征空间轨迹的说话人识别方法 |
Cited By (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112487978A (zh) * | 2020-11-30 | 2021-03-12 | 清华珠三角研究院 | 一种视频中说话人定位的方法、装置及计算机存储介质 |
| CN112487978B (zh) * | 2020-11-30 | 2024-04-16 | 清华珠三角研究院 | 一种视频中说话人定位的方法、装置及计算机存储介质 |
| CN113611285A (zh) * | 2021-09-03 | 2021-11-05 | 哈尔滨理工大学 | 基于层叠双向时序池化的语种识别方法 |
| CN113611285B (zh) * | 2021-09-03 | 2023-11-24 | 哈尔滨理工大学 | 基于层叠双向时序池化的语种识别方法 |
| CN117235435A (zh) * | 2023-11-15 | 2023-12-15 | 世优(北京)科技有限公司 | 确定音频信号损失函数的方法及装置 |
| CN117235435B (zh) * | 2023-11-15 | 2024-02-20 | 世优(北京)科技有限公司 | 确定音频信号损失函数的方法及装置 |
| CN117877527A (zh) * | 2024-02-21 | 2024-04-12 | 国能宁夏供热有限公司 | 一种基于通信设备的语音质量分析技术及分析方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109545229B (zh) | 2023-04-21 |
| CN109545229A (zh) | 2019-03-29 |
| SG11202103091XA (en) | 2021-04-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020143263A1 (zh) | 一种基于语音样本特征空间轨迹的说话人识别方法 | |
| Gao et al. | Transition movement models for large vocabulary continuous sign language recognition | |
| Zhuang et al. | Real-world acoustic event detection | |
| CN103279768B (zh) | 一种基于增量学习人脸分块视觉表征的视频人脸识别方法 | |
| CN111128128B (zh) | 一种基于互补模型评分融合的语音关键词检测方法 | |
| CN101101752A (zh) | 基于视觉特征的单音节语言唇读识别系统 | |
| Tang et al. | Partially supervised speaker clustering | |
| CN113658582B (zh) | 一种音视协同的唇语识别方法及系统 | |
| CN111581348A (zh) | 一种基于知识图谱的查询分析系统 | |
| Elakkiya et al. | Enhanced dynamic programming approach for subunit modelling to handle segmentation and recognition ambiguities in sign language | |
| Fang et al. | A novel approach to automatically extracting basic units from chinese sign language | |
| CN110807370B (zh) | 一种基于多模态的会议发言人身份无感确认方法 | |
| Lang et al. | Study of face detection algorithm for real-time face detection system | |
| Han et al. | Boosted subunits: a framework for recognising sign language from videos | |
| Yao et al. | Real time large vocabulary continuous sign language recognition based on OP/Viterbi algorithm | |
| Liu et al. | Lip event detection using oriented histograms of regional optical flow and low rank affinity pursuit | |
| Van Leeuwen | Speaker linking in large data sets | |
| Lu et al. | Slot transferability for cross-domain slot filling | |
| Feng et al. | Audio-visual human recognition using semi-supervised spectral learning and hidden Markov models | |
| Peng et al. | Fuse after Align: Improving Face-Voice Association Learning via Multimodal Encoder | |
| Das et al. | Unsupervised out-of-distribution dialect detection with Mahalanobis distance | |
| Hibraj et al. | Speaker clustering using dominant sets | |
| Trinh et al. | Audio event classification using SVM with GMM-UBM supervectors | |
| CN117237991B (zh) | 一种多维特征融合的人员识别方法及系统 | |
| CN119964554B (zh) | 基于自监督说话人表征解耦的方言语种识别方法、系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19908700 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19908700 Country of ref document: EP Kind code of ref document: A1 |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 03.11.2021) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19908700 Country of ref document: EP Kind code of ref document: A1 |









