WO2017135487A1 - 글로벌 모델 기반 오디오 객체 분리 방법 및 시스템 - Google Patents

글로벌 모델 기반 오디오 객체 분리 방법 및 시스템 Download PDF

Info

Publication number
WO2017135487A1
WO2017135487A1 PCT/KR2016/001393 KR2016001393W WO2017135487A1 WO 2017135487 A1 WO2017135487 A1 WO 2017135487A1 KR 2016001393 W KR2016001393 W KR 2016001393W WO 2017135487 A1 WO2017135487 A1 WO 2017135487A1
Authority
WO
WIPO (PCT)
Prior art keywords
model
nmf
sound source
audio
matrix
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2016/001393
Other languages
English (en)
French (fr)
Inventor
조충상
김제우
이영한
이혜인
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Korea Electronics Technology Institute
Original Assignee
Korea Electronics Technology Institute
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Korea Electronics Technology Institute filed Critical Korea Electronics Technology Institute
Publication of WO2017135487A1 publication Critical patent/WO2017135487A1/ko
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/008Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • G10L19/18Vocoders using multiple modes
    • G10L19/20Vocoders using multiple modes using sound class specific coding, hybrid encoders or object based coding
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/18Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band

Definitions

  • the present invention relates to audio processing techniques, and more particularly, to a method and system for separating a sound source into a plurality of audio objects.
  • An audio source consists of a number of audio objects, such as vocals, drums, guitars, pianos, and the like. It is possible to separate such audio sources into audio objects.
  • NMF non-negative matrix factorization model
  • an NMF model having the same length as the audio source to be separated is required. Since the length is different for each sound source, the NMF model used differs for each sound source.
  • the audio engineer is designing the NMF model by considering the attributes in addition to the length of the audio source to be separated, which is a very difficult and time-consuming task.
  • the present invention has been made to solve the above problems, and an object of the present invention is to provide a global model-based audio object separation method and system that automatically generates and uses a model used for sound source separation.
  • an audio separation method includes: automatically expanding a first model used for separation of a sound source to generate a second model; And dividing the sound source into a plurality of audio objects by using the generated second model.
  • the step of determining the length of the sound source may further extend, the first model with reference to the length.
  • the first model is a first non-negative matrix factorization model
  • the second model is a second NMF model
  • the generating step repeats the W matrix of the first NMF model. Can be extended.
  • the generating step may be extended by repeatedly arranging all or some units of the W matrix.
  • the generating step may be extended by selecting and arranging a part of the W matrix.
  • a portion of the W matrix may be randomly selected.
  • the generating may include selecting a part of the W matrix based on an analysis result of the sound source.
  • an audio separation system the generation unit for automatically expanding the first model used for separation of the sound source to generate a second model; And a separator configured to separate the sound source into a plurality of audio objects by using the generated second model.
  • FIG. 1 is a view provided in the description of an audio object separation system according to an embodiment of the present invention.
  • 3 is a diagram provided in the description of the extension of the global NMF model
  • FIG. 4 is a view provided to a detailed description of the NMF model expansion engine shown in FIG. 1;
  • 5 to 8 are diagrams provided for explaining a method of extending / converting H into H '.
  • FIG. 9 is a view provided to a detailed description of the index determination method by the audio analysis module.
  • the audio object separation system according to an embodiment of the present invention is a system for generating NMF models of smaller size by expanding NMF models required for separating an input sound source of length T into audio objects.
  • An audio object separation system that performs such a function includes an NMF model extension engine 110 and an NMF model based object separation engine 120 as shown in FIG. 1.
  • the NMF model expansion engine 110 automatically expands the global NMF models 10-1, ... 10-n used for sound source separation according to the sound source length, so that the NMF models 20-1 for sound source separation are used. , ... 20-n).
  • the NMF model based object separation engine 120 separates a sound source into a plurality of audio objects using NMF models 20-1, ... 20-n generated by the NMF model extension engine 110. .
  • Global NMF models 10-1, ... 10-n are small length NMF models provided for audio objects (vocals, drums, guitars, pianos) and are commonly used for all input sources. .
  • the NMF model is based on the W (F by k) and STFT results from a short term term Fourier transform (STFT) operation for a sound source in which several audio objects are mixed (hereinafter referred to as a mixing sound source). Is determined by setting H (k by N).
  • N is the window size used for the STFT operation
  • F is the number of frequency bins in the STFT operation
  • k is the dimension applied to the STFT operation.
  • the parameter that the length T of the mixing sound source affects is "N".
  • the global NMF as shown in FIG.
  • the H (k by N1) matrix of models 10-1, ... 10-n is extended to the H '(k by N2) matrix required for the mixing sound source of length T.
  • W of the global NMF models 10-1,... 10-n is used as it is in generating the NMF models 20-1,. That is, W of the global NMF models 10-1,... 10-n and W of the NMF models 20-1,.
  • FIG. 4 is a diagram provided for a detailed description of the NMF model expansion engine 110 shown in FIG. 1. 4 illustrates a process in which the NMF model extension engine 110 converts the global NMF models 10-1, ... 10-n into NMF models 20-1, ... 20-n. .
  • the NMF model expansion engine 110 determines the length T of the mixing sound source and configures the global NMF models 10-1, 10-n based on the determined length T.
  • H there are many ways to convert H into H 'and it will be explained in detail below.
  • H itself may be used, but the analysis result of the mixing sound source by the audio analysis module 115 may be used.
  • FIG. 5 shows a method of extending / converting H into H '.
  • the expansion / conversion method shown in FIG. 5 is a method of generating H 'by repeatedly arranging H having a column length of N1 but deleting a portion exceeding required N2.
  • FIG. 6 shows another method of extending / converting H into H '.
  • the expansion / conversion method shown in FIG. 6 is a method of generating H 'by repeatedly arranging columns of H having a column length of N1 in units of columns. At this time, the number of repetitions can be set differently for each column to meet the required N2.
  • the expansion / conversion method shown in FIG. 7 is a method of generating H ′ by repeating randomly selecting and arranging one of columns of H having a column length of N1.
  • FIG. 8 shows another method of extending / converting H into H '.
  • the expansion / conversion method shown in FIG. 8 is the same as the method shown in FIG. 7 in that H 'is repeatedly generated by selecting and listing one of the columns of H.
  • the audio analysis module 115 selects the index based on the index determined based on the analysis result of the mixing sound source. This is different from the method presented in FIG.
  • FIG. 9 is a diagram for describing a method of determining an index by the audio analysis module 115.
  • the audio analysis module 115 calculates the absolute value of the STFT calculation result for the mixing sound source, and selects the most similar column of H through similarity analysis while moving the window in the calculated result. Iteratively create indexes.
  • vocals, drums, guitars, pianos referred to as audio objects are also exemplary only.
  • the technical idea of the present invention is also applicable to the case where sound sources are separated into more various audio objects.
  • the audio object separation method and system presented in the embodiments of the present invention can be applied to fields such as audio effects, content production, surveillance systems, etc., as well as fields requiring voice separation or other types of sound source separation.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Electrophonic Musical Instruments (AREA)
  • Reverberation, Karaoke And Other Acoustics (AREA)

Abstract

글로벌 모델 기반 오디오 객체 분리 방법 및 시스템이 제공된다. 본 발명의 실시예에 따른 오디오 분리 방법은, 음원의 분리에 사용되는 제1 모델을 자동으로 확장하여 제2 모델을 생성하고, 생성된 제2 모델을 이용하여 음원을 다수의 오디오 객체로 분리한다. 이에 의해, 작은 길이의 글로벌 NMF 모델을 확장하여 긴 길이의 NMF 모델을 자동 생성하여 오디오 객체 분리에 이용하는 것이 가능해져, 보다 간편하고 짧은 시간으로 모든 음원들에 대해 오디오 객체 분리가 가능해진다.

Description

글로벌 모델 기반 오디오 객체 분리 방법 및 시스템
본 발명은 오디오 처리 기술에 관한 것으로, 더욱 상세하게는 음원을 다수의 오디오 객체들로 분리하는 방법 및 시스템에 관한 것이다.
오디오 음원은 다수의 오디오 객체들, 이를 테면, 보컬, 드럼, 기타, 피아노 등으로 구성된다. 이러한 오디오 음원을 오디오 객체들로 분리하는 것이 가능하다.
현재, 오디오 객체 분리에 있어 가장 많이 사용되는 기법들 중 하나는 NMF 모델(Non-Negative Matrix Factorization Model) 기반의 오디오 객체 분리 기법이다.
NMF 모델 기반으로 오디오 객체를 분리하기 위해서는, 분리하고자 하는 오디오 음원과 동일한 길이의 NMF 모델이 필요한데, 음원 마다 길이가 다르기 때문에, 이용되는 NMF 모델은 음원 마다 다르다.
또한, 분리도를 높이기 위해, 분리하고자 하는 오디오 음원의 길이 외에 속성을 더 고려하여 오디오 엔지니어가 NMF 모델을 직접 설계하고 있는데, 매우 어렵고 장시간이 소요되는 작업이다.
본 발명은 상기와 같은 문제점을 해결하기 위하여 안출된 것으로서, 본 발명의 목적은, 음원 분리에 사용되는 모델을 간편하게 자동으로 생성하여 이용하는 글로벌 모델 기반 오디오 객체 분리 방법 및 시스템을 제공함에 있다.
상기 목적을 달성하기 위한 본 발명의 일 실시예에 따른, 오디오 분리 방법은, 음원의 분리에 사용되는 제1 모델을 자동으로 확장하여 제2 모델을 생성하는 단계; 및 생성된 제2 모델을 이용하여, 상기 음원을 다수의 오디오 객체로 분리하는 단계;를 포함한다.
그리고, 본 발명의 일 실시예에 따른 오디오 분리 방법은, 상기 음원의 길이를 파악하는 단계;를 더 포함하고, 상기 길이를 참조로, 상기 제1 모델을 확장할 수 있다.
또한, 상기 제1 모델은, 제1 NMF 모델(Non-Negative Matrix Factorization Model)이고, 상기 제2 모델은, 제2 NMF 모델이며, 상기 생성 단계는, 상기 제1 NMF 모델의 W 행렬을 반복 나열하여 확장할 수 있다.
그리고, 상기 생성 단계는, 상기 W 행렬의 전부 또는 일부 단위로 반복 나열하여 확장할 수 있다.
또한, 상기 생성 단계는, 상기 W 행렬의 일부를 선택하여 나열함으로써 확장할 수 있다.
그리고, 상기 생성 단계는, 상기 W 행렬의 일부를 랜덤하게 선택할 수 있다.
또한, 상기 생성 단계는, 상기 음원의 분석 결과를 기초로, 상기 W 행렬의 일부를 선택할 수 있다.
한편, 본 발명의 다른 실시예에 따른, 오디오 분리 시스템은, 음원의 분리에 사용되는 제1 모델을 자동으로 확장하여 제2 모델을 생성하는 생성부; 및 생성된 제2 모델을 이용하여, 상기 음원을 다수의 오디오 객체로 분리하는 분리부;를 포함한다.
이상 설명한 바와 같이, 본 발명의 실시예들에 따르면, 작은 길이의 글로벌 NMF 모델을 확장하여 긴 길이의 NMF 모델을 자동 생성하여 오디오 객체 분리에 이용하는 것이 가능해져, 보다 간편하고 짧은 시간으로 모든 음원들에 대해 오디오 객체 분리가 가능해진다.
도 1은 본 발명의 일 실시예에 따른 오디오 객체 분리 시스템의 설명에 제공되는 도면,
도 2는 NMF 모델의 설명에 제공되는 도면,
도 3은 글로벌 NMF 모델의 확장에 대한 설명에 제공되는 도면,
도 4는, 도 1에 도시된 NMF 모델 확장 엔진의 상세 설명에 제공되는 도면,
도 5 내지 도 8은, H를 H'으로 확장/변환하는 방법의 설명에 제공되는 도면들, 그리고,
도 9는 오디오 분석 모듈에 의한 인덱스 결정 방법의 상세 설명에 제공되는 도면이다.
이하에서는 도면을 참조하여 본 발명을 보다 상세하게 설명한다.
도 1은 본 발명의 일 실시예에 따른 오디오 객체 분리 시스템의 설명에 제공되는 도면이다. 본 발명의 실시예에 따른 오디오 객체 분리 시스템은, 입력되는 길이 T의 음원을 오디오 객체들로 분리하기 위해 필요한 NMF 모델들을 그보다 작은 사이즈의 NMF 모델들을 확장하여 생성하는 시스템이다.
이와 같은 기능을 수행하는 본 발명의 실시예에 따른 오디오 객체 분리 시스템은, 도 1에 도시된 바와 같이, NMF 모델 확장 엔진(110) 및 NMF 모델 기반 객체 분리 엔진(120)을 포함한다.
NMF 모델 확장 엔진(110)은 음원 분리에 사용되는 글로벌 NMF 모델들(10-1, ... 10-n)을 음원 길이에 따라 자동으로 확장하여, 음원 분리에 사용할 NMF 모델들(20-1, ... 20-n)을 생성한다.
NMF 모델 기반 객체 분리 엔진(120)은 NMF 모델 확장 엔진(110)에 의해 생성된 NMF 모델들(20-1, ... 20-n)을 이용하여, 음원을 다수의 오디오 객체들로 분리한다.
글로벌 NMF 모델들(10-1, ... 10-n)은 오디오 객체(보컬, 드럼, 기타, 피아노) 별로 마련되어 있는 작은 길이의 NMF 모델들로, 입력되는 모든 음원들에 대해 공통적으로 사용된다.
NMF 모델은, 도 2에 도시된 바와 같이, '여러 오디오 객체들이 믹싱된 음원'(이하, '믹싱 음원'으로 표기)에 대한 STFT(Short Term Fourier Transform) 연산 결과로부터 W(F by k)와 H(k by N)을 설정함으로써 결정된다.
N은 STFT 연산에 사용된 윈도우 사이즈, F는 STFT 연산에서 frequency bin의 개수, k는 STFT 연산에 적용된 차원이다. 믹싱 음원의 길이 T가 영향을 미치는 파라미터는 "N"이다.
이에, 글로벌 NMF 모델들(10-1, ... 10-n)로부터 확장된 NMF 모델들(20-1, ... 20-n)을 생성함에 있어, 도 3에 도시된 바와 같이 글로벌 NMF 모델들(10-1, ... 10-n)의 H(k by N1) 행렬을 길이 T의 믹싱 음원에 대해 요구되는 H'(k by N2) 행렬로 확장한다.
그리고, 글로벌 NMF 모델들(10-1, ... 10-n)의 W는 NMF 모델들(20-1, ... 20-n)을 생성함에 있어 그대로 사용한다. 즉, 글로벌 NMF 모델들(10-1, ... 10-n)의 W와 NMF 모델들(20-1, ... 20-n)의 W는 동일하게 구현한다.
도 4는, 도 1에 도시된 NMF 모델 확장 엔진(110)의 상세 설명에 제공되는 도면이다. 도 4에는 NMF 모델 확장 엔진(110)이 글로벌 NMF 모델들(10-1, ... 10-n)을 NMF 모델들(20-1, ... 20-n)로 변환하는 과정이 나타나 있다.
구체적으로, 도 4에는, NMF 모델 확장 엔진(110)은 믹싱 음원의 길이 T를 파악하고, 파악된 길이 T를 기초로 글로벌 NMF 모델들(10-1, ... 10-n)을 구성하는 작은 사이즈의 H를 긴 사이즈의 H'로 변환하여, NMF 모델들(20-1, ... 20-n)를 생성하는 과정이 나타나 있다.
H를 H'으로 확장하여 변환하는 방법은 매우 다양하며, 이하에서 상세히 설명한다. H를 H'으로 변환함에 있어서는, H 자체만을 이용할 수도 있지만, 오디오 분석 모듈(115)에 의한 믹싱 음원의 분석 결과를 이용할 수도 있다.
도 5는 H를 H'으로 확장/변환하는 방법을 나타내었다. 도 5에 제시된 확장/변환 방법은, 열 길이가 N1인 H를 반복 나열하되, 요구되는 N2를 초과하는 부분은 삭제하여, H'를 생성하는 방법이다.
도 6은 H를 H'으로 확장/변환하는 다른 방법을 나타내었다. 도 6에 제시된 확장/변환 방법은, 열 길이가 N1인 H의 열들을 열 단위로 반복 나열하여, H'을 생성하는 방법이다. 이때, 요구되는 N2를 맞추기 위해 반복 횟수는 열 마다 다르게 설정할 수 있다.
도 7은 H를 H'으로 확장/변환하는 또 다른 방법을 나타내었다. 도 7에 제시된 확장/변환 방법은, 열 길이가 N1인 H의 열들 중 하나를 랜덤하게 선택하여 나열하는 것을 반복하여 H'를 생성하는 방법이다.
도 8은 H를 H'으로 확장/변환하는 또 다른 방법을 나타내었다. 도 8에 제시된 확장/변환 방법은, H의 열들 중 하나를 선택하여 나열하는 것을 반복하여 H'를 생성한다는 점에서 도 7에 제시된 방법과 동일하다.
하지만, 도 8에 제시된 방법에서는 H'에 나열할 H의 열들 중 하나를 랜덤하게 선택하는 것이 아니라, 오디오 분석 모듈(115)에 의해 믹싱 음원의 분석 결과를 기초로 결정된 인덱스에 따라 선택한다는 점에서, 도 7에 제시된 방법과 차이가 있다.
도 9는 오디오 분석 모듈(115)에 의한 인덱스 결정 방법의 상세 설명에 제공되는 도면이다. 도 9에 도시된 바와 같이, 오디오 분석 모듈(115)은 믹싱 음원에 대한 STFT 연산 결과의 절대값을 산출하고, 산출된 결과에서 윈도우를 이동시키면서 유사도 분석을 통해 가장 유사한 H의 열을 선택하는 것을 반복하여 인덱스들을 생성한다.
지금까지, 글로벌 모델 기반 오디오 객체 분리 방법 및 시스템에 대해 바람직한 실시예들을 들어 상세히 설명하였다.
위 실시예들에서는 NMF 모델들을 이용한 오디오 객체 분리를 상정하였는데 예시를 위한 것이다. NMF 모델이 아닌 그로부터 변형된 모델 또는 그와 다른 종류의 모델을 적용하는 경우에도, 본 발명의 기술적 사상이 적용될 수 있음은 물론이다.
또한, 위 실시예들에서, 오디오 객체들로 언급한 보컬, 드럼, 기타, 피아노 역시 예시적인 것에 불과하다. 이 보다 더 다양한 오디오 객체들로 음원을 분리하는 경우에도 본 발명의 기술적 사상이 적용가능하다.
본 발명의 실시예들에서 제시한 오디오 객체 분리 방법 및 시스템은, 오디오 효과, 콘텐츠 제작, 감시 시스템 등과 같은 분야는 물론, 음성 분리나 그 밖의 다른 종류의 음원 분리가 필요한 분야에 적용될 수 있다.
또한, 이상에서는 본 발명의 바람직한 실시예에 대하여 도시하고 설명하였지만, 본 발명은 상술한 특정의 실시예에 한정되지 아니하며, 청구범위에서 청구하는 본 발명의 요지를 벗어남이 없이 당해 발명이 속하는 기술분야에서 통상의 지식을 가진자에 의해 다양한 변형실시가 가능한 것은 물론이고, 이러한 변형실시들은 본 발명의 기술적 사상이나 전망으로부터 개별적으로 이해되어져서는 안될 것이다.

Claims (8)

  1. 음원의 분리에 사용되는 제1 모델을 자동으로 확장하여 제2 모델을 생성하는 단계; 및
    생성된 제2 모델을 이용하여, 상기 음원을 다수의 오디오 객체로 분리하는 단계;를 포함하는 것을 특징으로 하는 오디오 분리 방법.
  2. 청구항 1에 있어서,
    상기 음원의 길이를 파악하는 단계;를 더 포함하고,
    상기 길이를 참조로, 상기 제1 모델을 확장하는 것을 특징으로 하는 오디오 분리 방법.
  3. 청구항 1에 있어서,
    상기 제1 모델은, 제1 NMF 모델(Non-Negative Matrix Factorization Model)이고,
    상기 제2 모델은, 제2 NMF 모델이며,
    상기 생성 단계는,
    상기 제1 NMF 모델의 W 행렬을 반복 나열하여 확장하는 것을 특징으로 하는 오디오 분리 방법.
  4. 청구항 3에 있어서,
    상기 생성 단계는,
    상기 W 행렬의 전부 또는 일부 단위로 반복 나열하여 확장하는 것을 특징으로 하는 오디오 분리 방법.
  5. 청구항 3에 있어서,
    상기 생성 단계는,
    상기 W 행렬의 일부를 선택하여 나열함으로써 확장하는 것을 특징으로 하는 오디오 분리 방법.
  6. 청구항 5에 있어서,
    상기 생성 단계는,
    상기 W 행렬의 일부를 랜덤하게 선택하는 것을 특징으로 하는 오디오 분리 방법.
  7. 청구항 5에 있어서,
    상기 생성 단계는,
    상기 음원의 분석 결과를 기초로, 상기 W 행렬의 일부를 선택하는 것을 특징으로 하는 오디오 분리 방법.
  8. 음원의 분리에 사용되는 제1 모델을 자동으로 확장하여 제2 모델을 생성하는 생성부; 및
    생성된 제2 모델을 이용하여, 상기 음원을 다수의 오디오 객체로 분리하는 분리부;를 포함하는 것을 특징으로 하는 오디오 분리 시스템.
PCT/KR2016/001393 2016-02-05 2016-02-11 글로벌 모델 기반 오디오 객체 분리 방법 및 시스템 Ceased WO2017135487A1 (ko)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
KR1020160014914A KR101864925B1 (ko) 2016-02-05 2016-02-05 글로벌 모델 기반 오디오 객체 분리 방법 및 시스템
KR10-2016-0014914 2016-02-05

Publications (1)

Publication Number Publication Date
WO2017135487A1 true WO2017135487A1 (ko) 2017-08-10

Family

ID=59501027

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2016/001393 Ceased WO2017135487A1 (ko) 2016-02-05 2016-02-11 글로벌 모델 기반 오디오 객체 분리 방법 및 시스템

Country Status (2)

Country Link
KR (1) KR101864925B1 (ko)
WO (1) WO2017135487A1 (ko)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114093369A (zh) * 2021-10-21 2022-02-25 北京捷通华声科技股份有限公司 一种话者分离方法、装置、电子设备与存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20110012946A (ko) * 2009-07-31 2011-02-09 포항공과대학교 산학협력단 소리의 복원 방법, 소리의 복원 방법을 기록한 기록매체 및 소리의 복원 방법을 수행하는 장치
KR20110023688A (ko) * 2009-08-28 2011-03-08 한국전자통신연구원 음악 음원 분리 방법 및 장치
US20150205575A1 (en) * 2014-01-20 2015-07-23 Canon Kabushiki Kaisha Audio signal processing apparatus and method thereof
US20150242180A1 (en) * 2014-02-21 2015-08-27 Adobe Systems Incorporated Non-negative Matrix Factorization Regularized by Recurrent Neural Networks for Audio Processing
KR20150142777A (ko) * 2014-06-11 2015-12-23 전자부품연구원 오디오 소스 분리 방법 및 이를 적용한 오디오 시스템

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR101511553B1 (ko) 2014-02-14 2015-04-13 전자부품연구원 다중 단계 오디오 분리 방법 및 이를 적용한 오디오 시스템

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20110012946A (ko) * 2009-07-31 2011-02-09 포항공과대학교 산학협력단 소리의 복원 방법, 소리의 복원 방법을 기록한 기록매체 및 소리의 복원 방법을 수행하는 장치
KR20110023688A (ko) * 2009-08-28 2011-03-08 한국전자통신연구원 음악 음원 분리 방법 및 장치
US20150205575A1 (en) * 2014-01-20 2015-07-23 Canon Kabushiki Kaisha Audio signal processing apparatus and method thereof
US20150242180A1 (en) * 2014-02-21 2015-08-27 Adobe Systems Incorporated Non-negative Matrix Factorization Regularized by Recurrent Neural Networks for Audio Processing
KR20150142777A (ko) * 2014-06-11 2015-12-23 전자부품연구원 오디오 소스 분리 방법 및 이를 적용한 오디오 시스템

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114093369A (zh) * 2021-10-21 2022-02-25 北京捷通华声科技股份有限公司 一种话者分离方法、装置、电子设备与存储介质

Also Published As

Publication number Publication date
KR20170093474A (ko) 2017-08-16
KR101864925B1 (ko) 2018-06-05

Similar Documents

Publication Publication Date Title
CN110310633B (zh) 多音区语音识别方法、终端设备和存储介质
KR101280253B1 (ko) 음원 분리 방법 및 그 장치
WO2019004671A1 (ko) 인공지능 기반 악성코드 검출 시스템 및 방법
JP7006592B2 (ja) 信号処理装置、信号処理方法および信号処理プログラム
WO2021010613A1 (ko) 다중 디코더를 이용한 심화 신경망 기반의 비-자동회귀 음성 합성 방법 및 시스템
JP7830680B2 (ja) 自己教師あり話者照合のための漸進的対照学習フレームワーク
CN113892136A (zh) 信号提取系统、信号提取学习方法以及信号提取学习程序
WO2018186708A1 (ko) 음원의 하이라이트 구간을 결정하는 방법, 장치 및 컴퓨터 프로그램
CN105761733A (zh) 生成歌词文件的方法和装置
JP2018206261A (ja) 単語分割推定モデル学習装置、単語分割装置、方法、及びプログラム
WO2023234606A1 (ko) 글로벌 스타일 토큰과 예측 모델로 생성한 화자 임베딩 기반의 화자 적응 방법 및 시스템
Lee et al. Combining Multi-Scale Features Using Sample-Level Deep Convolutional Neural Networks for Weakly Supervised Sound Event Detection.
US10679646B2 (en) Signal processing device, signal processing method, and computer-readable recording medium
WO2017135487A1 (ko) 글로벌 모델 기반 오디오 객체 분리 방법 및 시스템
CN116564278B (zh) 歌声检测模型的训练方法、歌声检测方法、设备及介质
JP6747447B2 (ja) 信号検知装置、信号検知方法、および信号検知プログラム
CN110827850B (zh) 音频分离方法、装置、设备及计算机可读存储介质
Tengtrairat et al. Extension of DUET to single-channel mixing model and separability analysis
JP5784075B2 (ja) 信号区間分類装置、信号区間分類方法、およびプログラム
KR102515149B1 (ko) 자연어 이해를 위한 그래프 변환 시스템 및 방법
WO2013168848A1 (ko) 하모닉 주파수 사이의 종속관계를 이용한 암묵 신호 분리 방법 및 이를 위한 디믹싱 시스템
ATE255250T1 (de) Verfahren für die realisation von leistungstests von rechnergeräten über ein telekommunikationsnetzwerk
JP2008278406A (ja) 音源分離装置,音源分離プログラム及び音源分離方法
JP5522393B2 (ja) 音響モデル構築装置、音声認識装置、音響モデル構築方法、およびプログラム
KR101511553B1 (ko) 다중 단계 오디오 분리 방법 및 이를 적용한 오디오 시스템

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16889466

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16889466

Country of ref document: EP

Kind code of ref document: A1