WO2024136409A1 - 화자 분할 방법 및 시스템 - Google Patents

화자 분할 방법 및 시스템 Download PDF

Info

Publication number
WO2024136409A1
WO2024136409A1 PCT/KR2023/021003 KR2023021003W WO2024136409A1 WO 2024136409 A1 WO2024136409 A1 WO 2024136409A1 KR 2023021003 W KR2023021003 W KR 2023021003W WO 2024136409 A1 WO2024136409 A1 WO 2024136409A1
Authority
WO
WIPO (PCT)
Prior art keywords
speaker
voice
feature vector
extracting
feature vectors
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2023/021003
Other languages
English (en)
French (fr)
Inventor
최민석
이봉진
최소연
허희수
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Naver Corp
Original Assignee
Naver Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Naver Corp filed Critical Naver Corp
Priority to JP2025513626A priority Critical patent/JP2025530808A/ja
Publication of WO2024136409A1 publication Critical patent/WO2024136409A1/ko
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/02Preprocessing operations, e.g. segment selection; Pattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal components; Feature selection or extraction
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/02Feature extraction for speech recognition; Selection of recognition unit
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/93Discriminating between voiced and unvoiced parts of speech signals

Definitions

  • the present disclosure relates to a speaker segmentation method and system. Specifically, a speaker segmentation method and system for extracting a central feature vector with the highest reliability among speaker feature vectors and extracting the voice of a specific speaker associated with the extracted central feature vector. It's about.
  • a meeting minutes creation service that automatically generates meeting minutes based on meeting recording files
  • an automatic subtitle creation service that automatically generates subtitles for dramas, movies, videos, etc. are provided.
  • voice data to be recognized when voice data to be recognized includes voices of multiple speakers, speaker segmentation may be performed to distinguish and recognize speech sections of multiple speakers.
  • speaker segmentation may be performed to distinguish and recognize speech sections of multiple speakers.
  • the voice to be recognized is mixed with noise or there is a section where the voices of multiple speakers overlap, there is a problem that the speaker feature vector may be contaminated and an error may occur in the speaker segmentation result.
  • the present disclosure provides a speaker segmentation method, a computer-readable non-transitory recording medium recording commands, and a device (system) to solve the above problems.
  • the present disclosure may be implemented in various ways, including a method, a device (system), or a computer-readable non-transitory recording medium recording instructions.
  • a speaker segmentation method performed by at least one processor includes extracting a first set of speaker feature vectors based on speech sections included in an input signal, the extracted first Based on the set of speaker feature vectors, extracting a first central feature vector with the highest confidence and extracting the first speaker's voice associated with the first central feature vector from the input signal.
  • a computer-readable non-transitory recording medium recording instructions for executing a method according to an embodiment of the present disclosure on a computer is provided.
  • An information processing system includes a memory and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, and the at least one program includes an input Based on the speech section included in the signal, a first set of speaker feature vectors are extracted, and based on the extracted first set of speaker feature vectors, a first central feature vector with the highest reliability is extracted, and the input signal Includes instructions for extracting the voice of the first speaker associated with the first central feature vector from.
  • a central feature vector with high reliability in representing a specific speaker among speaker feature vectors extracted from a speech section even when some of the speaker feature vectors are contaminated, the voice of a specific speaker
  • a representative speaker feature vector that well reflects the features can be extracted.
  • a series of processes of extracting the central feature vector and extracting the voice of another specific speaker associated with the central feature vector are repeated for the residual signal from which the voice of a specific speaker associated with the central feature vector has been removed from the input signal.
  • FIG. 1 is a diagram illustrating an example of a speaker segmentation method according to an embodiment of the present disclosure.
  • Figure 2 is a schematic diagram showing a configuration in which an information processing system according to an embodiment of the present disclosure is connected to enable communication with a plurality of user terminals.
  • Figure 3 is a block diagram showing the internal configuration of a user terminal and an information processing system according to an embodiment of the present disclosure.
  • FIG. 4 is a diagram illustrating an example of extracting a plurality of speaker feature vectors based on speech sections included in an input signal according to an embodiment of the present disclosure.
  • FIG. 5 is a diagram illustrating an example of performing clustering on a speaker feature vector according to an embodiment of the present disclosure.
  • FIG. 6 is a diagram illustrating an example of extracting a central feature vector based on a plurality of speaker feature vectors extracted according to an embodiment of the present disclosure.
  • FIG. 7 is a diagram illustrating an example of extracting a specific speaker's voice from an input signal according to an embodiment of the present disclosure.
  • Figure 8 is a diagram showing an example of a target voice extraction model according to an embodiment of the present disclosure.
  • Figure 9 is a diagram illustrating an example of a method for learning a target voice extraction model according to an embodiment of the present disclosure.
  • Figure 10 is a flowchart showing an example of a speaker segmentation method according to an embodiment of the present disclosure.
  • a modulee' or 'unit' refers to a software or hardware component, and the 'module' or 'unit' performs certain roles.
  • 'module' or 'unit' is not limited to software or hardware.
  • a 'module' or 'unit' may be configured to reside on an addressable storage medium and may be configured to run on one or more processors.
  • a 'module' or 'part' refers to components such as software components, object-oriented software components, class components and task components, processes, functions and properties. , procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables.
  • Components and 'modules' or 'parts' may be combined into smaller components and 'modules' or 'parts' or further components and 'modules' or 'parts'.
  • a 'module' or 'unit' may be implemented with a processor and memory.
  • 'Processor' should be interpreted broadly to include general-purpose processors, central processing units (CPUs), microprocessors, digital signal processors (DSPs), controllers, microcontrollers, state machines, etc.
  • 'processor' may refer to an application-specific integrated circuit (ASIC), programmable logic device (PLD), field programmable gate array (FPGA), etc.
  • ASIC application-specific integrated circuit
  • PLD programmable logic device
  • FPGA field programmable gate array
  • 'Processor' refers to a combination of processing devices, for example, a combination of a DSP and a microprocessor, a combination of a plurality of microprocessors, a combination of one or more microprocessors in combination with a DSP core, or any other such combination of configurations. You may. Additionally, 'memory' should be interpreted broadly to include any electronic component capable of storing electronic information.
  • RAM random access memory
  • ROM read-only memory
  • NVRAM non-volatile random access memory
  • PROM programmable read-only memory
  • EPROM erasable-programmable read-only memory
  • a memory is said to be in electronic communication with a processor if the processor can read information from and/or write information to the memory.
  • the memory integrated into the processor is in electronic communication with the processor.
  • 'system' may include at least one of a server device and a cloud device, but is not limited thereto.
  • a system may consist of one or more server devices.
  • a system may consist of one or more cloud devices.
  • the system may be operated with a server device and a cloud device configured together.
  • 'machine learning model' may include any model used to infer an answer to a given input.
  • the machine learning model may include an artificial neural network model including an input layer (layer), a plurality of hidden layers, and an output layer.
  • each layer may include multiple nodes.
  • a machine learning model may refer to an artificial neural network model
  • an artificial neural network model may refer to a machine learning model.
  • 'display' may refer to any display device associated with a computing device, e.g., any display device capable of displaying any information/data controlled by or provided by the computing device. can refer to.
  • 'each of a plurality of A' or 'each of a plurality of A' may refer to each of all components included in a plurality of A, or may refer to each of some components included in a plurality of A. .
  • 'reliability' may refer to the probability that a specific speaker feature vector is estimated to represent the speech characteristics of a specific speaker. Therefore, in the present disclosure, the speaker feature vector with the highest reliability is a speaker feature vector estimated to best represent the voice characteristics of a specific speaker, or a speaker feature vector estimated to best reflect the voice characteristics of a specific speaker. can refer to.
  • the reliability of speaker feature vectors can be measured by various measures. According to one embodiment, within a vector space where a plurality of speaker feature vectors extracted from a specific input signal are located, the more dense the corresponding speaker feature vectors are in the surrounding area of a specific speaker feature vector, the higher the reliability can be.
  • 'central feature vector' may refer to a speaker feature vector with the highest reliability among a plurality of speaker feature vectors extracted from a specific input signal.
  • the speaker segmentation method analyzes the input signal 110 containing the voices of a plurality of speakers, recognizes each speaker's speech section separately, and extracts the voice of the corresponding speaker.
  • the information processing system may receive an input signal 110.
  • the input signal 110 may include spoken voices and background noise of a plurality of speakers.
  • the information processing system can then detect speech segments from the input signal 110 (120).
  • the voice section may be a section of the input signal 110 that includes the speaker's spoken voice.
  • the information processing system may extract a plurality of speaker feature vectors based on at least a portion of the input signal 110 corresponding to the detected voice section (130). The process of the information processing system detecting a voice section from the input signal 110 and extracting a plurality of speaker feature vectors will be described in detail later with reference to FIG. 4.
  • the information processing system may extract a central feature vector with the highest reliability based on the extracted plurality of speaker feature vectors (140).
  • the specific speaker refers to a predetermined speaker among a plurality of speakers whose speech is included in the input signal 110, but who utters a speech associated with the central feature vector among several speakers whose speech is included. It can refer to any one speaker. Therefore, the central feature vector with the highest reliability among the plurality of speaker feature vectors cannot specify or recognize in advance which speaker the speaker is among the plurality of speakers whose speech is included in the input signal 110, but is arbitrary. It can refer to a speaker feature vector that is estimated to best represent the voice features of a speaker.
  • a speaker identification (SID) process may be additionally performed.
  • the information processing system when speaker feature information extracted from voice samples of some speakers or voice samples of some speakers is obtained in advance, the information processing system additionally performs a speaker identification process to obtain speaker features extracted from the input signal 110. It is possible to estimate which speaker a vector is associated with.
  • the central feature vector may be used to refer to a speaker feature vector that is estimated to best represent the speech characteristics of a predetermined or pre-identified speaker among a plurality of speakers included in the input signal 110.
  • 'specific speaker' will be used to refer to any one speaker among a plurality of speakers whose speech voice is included in the input signal.
  • the information processing system that extracted the central feature vector may extract the voice 152 of a specific speaker associated with the central feature vector from the input signal 110 (150).
  • the information processing system may extract the voice 152 of a specific speaker associated with the central feature vector from the input signal 110 using a target voice extraction model.
  • the process by which the information processing system extracts the voice 152 of a specific speaker associated with the central feature vector from the input signal 110 will be described in detail later with reference to FIGS. 7 and 8.
  • the information processing system may determine whether a voice section remains in the residual signal 154 excluding the voice 152 of a specific speaker extracted from the input signal 110 (160). If it is determined that a voice section remains in the residual signal 154, the voices of other speaker(s) included in the input signal 110 can be extracted by repeatedly performing the above-described series of processes.
  • the information processing system extracts the speaker feature vector 130, centers the remaining voice sections excluding the section containing the extracted voice 152 of a specific speaker, among the voice sections included in the input signal 110.
  • processes such as feature vector extraction (140) and specific speaker voice extraction (150), the voices of other speaker(s) can be extracted.
  • the information processing system may include some of the speaker feature vectors extracted based on the remaining voice sections excluding the section containing the voice 152 of a specific speaker among the plurality of speaker feature vectors extracted based on the input signal 110.
  • the voices of other speaker(s) can be extracted by repeatedly performing processes such as central feature vector extraction (140) and specific speaker voice extraction (150).
  • the information processing system may use the residual signal 154 as the input signal 110 to repeatedly perform the above-described series of processes (120, 130, 140, 150, etc.).
  • the voice 152 of a specific speaker can be separated with high accuracy. There is. Then, a series of processes are repeatedly performed on the residual signal 154 from which the voice 152 of a specific speaker is removed from the input signal 110, so that during this process, sections where the voices of multiple speakers overlap or the voices of multiple speakers overlap. Since the section where voices are closely related can gradually decrease, highly reliable speaker feature vectors can be extracted for the remaining speakers whose speech voices are included in the residual signal 154, allowing high-quality speaker segmentation to be performed.
  • the speaker segmentation method of the present disclosure will be described in more detail.
  • the speaker associated with the voice whose central feature vector is extracted first from the input signal 410 will be referred to as the first speaker.
  • the speaker associated with the voice from which the nth central feature vector is extracted will be referred to as the nth speaker.
  • the speaker segmentation method is described as being performed by an information processing system, but the present disclosure is not limited to this, and at least some steps included in the speaker segmentation method of the present disclosure may be performed by a user terminal. For example, a series of processes (120, 130, 140, 150, etc.) related to the speaker segmentation method of the present disclosure may be performed by the user terminal. However, for convenience of explanation, hereinafter, the speaker segmentation method of the present disclosure will be described assuming that it is performed by an information processing system.
  • Figure 2 is a schematic diagram showing a configuration in which the information processing system 230 according to an embodiment of the present disclosure is connected to communicate with a plurality of user terminals 210_1, 210_2, and 210_3.
  • a plurality of user terminals 210_1, 210_2, and 210_3 are connected to a speaker segmentation service, a voice recognition service, or various applications (e.g., automatic meeting minutes) using a speaker segmentation service and/or voice recognition function through the network 220.
  • It can be connected to an information processing system 230 that can provide a generation application, automatic subtitle playback application, etc.).
  • the plurality of user terminals 210_1, 210_2, and 210_3 may include user terminals to be provided with a speaker segmentation service and/or a voice recognition service.
  • information processing system 230 is one or more server devices capable of storing, providing, and executing computer executable programs (e.g., downloadable applications) and data related to speaker segmentation services and/or speech recognition services, etc. and/or a database, or one or more distributed computing devices and/or distributed databases based on cloud computing services.
  • the user terminals 210_1, 210_2, and 210_3 provide speaker segmentation to the user through a computer executable program (e.g., a downloadable application) related to the speaker segmentation service, speech recognition service, etc., rather than through the network 220. Services and/or voice recognition services may be provided.
  • a computer executable program e.g., a downloadable application
  • the speaker segmentation service and/or voice recognition service provided by the information processing system 230 includes a speaker segmentation application, a voice recognition application, an automatic meeting minutes generation application, and an automatic speaker segmentation application installed on each of the plurality of user terminals 210_1, 210_2, and 210_3. It may be provided to the user through a subtitle creation application, voice editing application, mobile browser application, or web browser.
  • the information processing system 230 provides information corresponding to speaker segmentation requests, voice recognition requests, meeting minutes creation requests, and subtitle creation requests received from the user terminals 210_1, 210_2, and 210_3 through a speaker segmentation application. Or you can perform corresponding processing.
  • a plurality of user terminals 210_1, 210_2, and 210_3 may communicate with the information processing system 230 through the network 220.
  • the network 220 may be configured to enable communication between a plurality of user terminals 210_1, 210_2, and 210_3 and the information processing system 230.
  • the network 220 may be, for example, a wired network such as Ethernet, a wired home network (Power Line Communication), a telephone line communication device, and RS-serial communication, a mobile communication network, a wireless LAN (WLAN), It may consist of wireless networks such as Wi-Fi, Bluetooth, and ZigBee, or a combination thereof.
  • the communication method is not limited, and may include communication methods utilizing communication networks that the network 220 may include (e.g., mobile communication networks, wired Internet, wireless Internet, broadcasting networks, satellite networks, etc.) as well as user terminals (210_1, 210_2, 210_3). ) may also include short-range wireless communication between the network 220 may include (e.g., mobile communication networks, wired Internet, wireless Internet, broadcasting networks, satellite networks, etc.) as well as user terminals (210_1, 210_2, 210_3). ) may also include short-range wireless communication between the network 220 may include (e.g., mobile communication networks, wired Internet, wireless Internet, broadcasting networks, satellite networks, etc.) as well as user terminals (210_1, 210_2, 210_3). ) may also include short-range wireless communication between the network 220 may include (e.g., mobile communication networks, wired Internet, wireless Internet, broadcasting networks, satellite networks, etc.) as well as user terminals (210_1, 210_2,
  • the mobile phone terminal (210_1), tablet terminal (210_2), and PC terminal (210_3) are shown as examples of user terminals, but they are not limited thereto, and the user terminals (210_1, 210_2, 210_3) use wired and/or wireless communication.
  • This is possible and may be any computing device on which a speaker segmentation application, a voice recognition application, an automatic meeting minutes creation application, an automatic subtitle creation application, a voice editing application, a mobile browser application, a web browser, etc. can be installed and executed.
  • user terminals include AI speakers, smartphones, mobile phones, navigation, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, game consoles, It may include wearable devices, IoT (internet of things) devices, VR (virtual reality) devices, AR (augmented reality) devices, set-top boxes, etc.
  • IoT Internet of things
  • VR virtual reality
  • AR augmented reality
  • FIG 220 three user terminals (210_1, 210_2, 210_3) are shown as communicating with the information processing system 230 through the network 220, but this is not limited to this, and a different number of user terminals are connected to the network ( It may be configured to communicate with the information processing system 230 through 220).
  • the information processing system 230 may receive input signals including utterances of a plurality of speakers from a plurality of user terminals 210_1, 210_2, and 210_3. Then, the information processing system 230 extracts a plurality of speaker feature vectors based on the speech section included in the input signal, and selects a central feature vector with the highest reliability based on the extracted plurality of speaker feature vectors. It can be extracted. Then, the information processing system 230 may extract the voice of a specific speaker associated with the central feature vector from the input signal. The information processing system 230 can separate the voices of each speaker included in the input signal by repeating the above-described process several times.
  • the information processing system 230 may transmit the separated voices of each speaker and/or the results of performing additional processing on each speaker's voice to the user terminal 210 as a result of speaker segmentation. For example, the information processing system generates meeting minutes or subtitles by performing Automatic Speech Recognition (ASR) or STT (Speech to Text) based on the voices of each separated speaker, and generates meeting minutes or subtitles on the user terminal ( 210).
  • ASR Automatic Speech Recognition
  • STT Speech to Text
  • FIG. 3 is a block diagram showing the internal configuration of the user terminal 210 and the information processing system 230 according to an embodiment of the present disclosure.
  • the user terminal 210 is any computing device capable of running a speaker segmentation application, a voice recognition application, an automatic meeting minutes creation application, an automatic subtitle creation application, a voice editing application, a mobile browser application, or a web browser, and capable of wired/wireless communication. It may refer to, and may include, for example, the mobile phone terminal 210_1, tablet terminal 210_2, and PC terminal 210_3 of FIG. 2 .
  • the user terminal 210 may include a memory 312, a processor 314, a communication module 316, and an input/output interface 318.
  • information processing system 230 may include memory 332, processor 334, communication module 336, and input/output interface 338. As shown in FIG. 3, the user terminal 210 and the information processing system 230 are configured to communicate information and/or data through the network 220 using respective communication modules 316 and 336. It can be. Additionally, the input/output device 320 may be configured to input information and/or data to the user terminal 210 through the input/output interface 318 or to output information and/or data generated from the user terminal 210.
  • Memories 312 and 332 may include any non-transitory computer-readable recording medium.
  • the memories 312 and 332 are non-perishable mass storage devices such as random access memory (RAM), read only memory (ROM), disk drive, solid state drive (SSD), and flash memory. It may include a (permanent mass storage device).
  • non-perishable mass storage devices such as ROM, SSD, flash memory, disk drive, etc. may be included in the user terminal 210 or the information processing system 230 as a separate persistent storage device that is distinct from memory.
  • the memories 312 and 332 may store an operating system and at least one program code (for example, code installed in the user terminal 210 for a speaker segmentation application, etc.).
  • These software components may be loaded from a computer-readable recording medium separate from the memories 312 and 332.
  • This separate computer-readable recording medium may include a recording medium directly connectable to the user terminal 210 and the information processing system 230, for example, a floppy drive, disk, tape, DVD/CD- It may include computer-readable recording media such as ROM drives and memory cards.
  • software components may be loaded into the memories 312 and 332 through a communication module rather than a computer-readable recording medium. For example, at least one program is loaded into memory 312, 332 based on a computer program installed by files provided over the network 220 by developers or a file distribution system that distributes installation files for applications. It can be.
  • the processors 314 and 334 may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input/output operations. Instructions may be provided to the processors 314 and 334 by memories 312 and 332 or communication modules 316 and 336. For example, processors 314 and 334 may be configured to execute received instructions according to program codes stored in recording devices such as memories 312 and 332.
  • the communication modules 316 and 336 may provide a configuration or function for the user terminal 210 and the information processing system 230 to communicate with each other through the network 220, and may provide a configuration or function for the user terminal 210 and/or information processing.
  • the system 230 may provide a configuration or function for communicating with other user terminals or other systems (for example, a separate cloud system, etc.).
  • a request or data generated by the processor 314 of the user terminal 210 according to a program code stored in a recording device such as the memory 312 e.g., speaker segmentation request, voice recognition request, meeting minutes creation request, Subtitle creation request, etc.
  • a control signal or command provided under the control of the processor 334 of the information processing system 230 is transmitted through the communication module 316 of the user terminal 210 through the communication module 336 and the network 220. It may be received by the user terminal 210.
  • the user terminal 210 may perform additional processing on the voice of each speaker separated from the input signal and/or the voice of each speaker through the communication module 316 from the information processing system 230 (e.g. For example, meeting minutes, subtitles, etc.) can be received as a result of speaker segmentation.
  • the input/output interface 318 may be a means for interfacing with the input/output device 320.
  • input devices may include devices such as cameras, keyboards, microphones, mice, etc., including audio sensors and/or image sensors
  • output devices may include devices such as displays, speakers, haptic feedback devices, etc. You can.
  • the input/output interface 318 may be a means for interfacing with a device that has components or functions for performing input and output, such as a touch screen, integrated into one.
  • the processor 314 of the user terminal 210 uses information and/or data provided by the information processing system 230 or another user terminal when processing instructions of a computer program loaded in the memory 312. A service screen, etc.
  • the input/output device 320 is shown not to be included in the user terminal 210, but the present invention is not limited to this and may be configured as a single device with the user terminal 210. Additionally, the input/output interface 338 of the information processing system 230 is connected to the information processing system 230 or means for interfacing with a device (not shown) for input or output that the information processing system 230 may include. It can be. In FIG.
  • the input/output interfaces 318 and 338 are shown as elements configured separately from the processors 314 and 334, but the present invention is not limited thereto, and the input/output interfaces 318 and 338 may be configured to be included in the processors 314 and 334. there is.
  • the user terminal 210 and information processing system 230 may include more components than those in FIG. 3 . However, there is no need to clearly show most prior art components. According to one embodiment, the user terminal 210 may be implemented to include at least some of the input/output devices 320 described above. Additionally, the user terminal 210 may further include other components such as a transceiver, a global positioning system (GPS) module, a camera, various sensors, and a database. For example, if the user terminal 210 is a smartphone, it may include components generally included in a smartphone, such as an acceleration sensor, a gyro sensor, a camera module, various physical buttons, and a touch screen.
  • GPS global positioning system
  • buttons using a panel, input/output ports, and a vibrator for vibration may be implemented to be further included in the user terminal 210.
  • the processor 314 of the user terminal 210 may be configured to operate an application that provides a speaker segmentation service. At this time, code associated with the corresponding application and/or program may be loaded into the memory 312 of the user terminal 210.
  • the processor 314 inputs or selects input through an input device such as a touch screen, a keyboard, a camera including an audio sensor and/or an image sensor, and a microphone connected to the input/output interface 318.
  • Text, images, videos, voices, and/or motions can be received, and the received text, images, videos, voices, and/or motions can be stored in the memory 312 or stored in the communication module 316 and the network 220. It can be provided to the information processing system 230 through .
  • the processor 314 may receive a user input requesting speaker segmentation and provide it to the information processing system 230 through the communication module 316 and the network 220.
  • the processor 314 receives input indicating the user's selection of an input signal containing speech from multiple speakers, etc. It can be provided to the information processing system 230 through the communication module 316 and the network 220.
  • the processor 314 of the user terminal 210 manages, processes, and/or stores information and/or data received from the input device 320, other user terminals, the information processing system 230, and/or a plurality of external systems. It can be configured to do so. Information and/or data processed by processor 314 may be provided to information processing system 230 via communication module 316 and network 220.
  • the processor 314 of the user terminal 210 may transmit information and/or data to the input/output device 320 through the input/output interface 318 and output the information. For example, the processor 314 may display the received information and/or data on the screen of the user terminal.
  • the processor 334 of the information processing system 230 may be configured to manage, process, and/or store information and/or data received from a plurality of user terminals 210 and/or a plurality of external systems. Information and/or data processed by the processor 334 may be provided to the user terminal 210 through the communication module 336 and the network 220. According to one embodiment, the information processing system 230 extracts a plurality of speaker feature vectors based on voice sections included in the input signal received from the user terminal 210, and based on the extracted plurality of speaker feature vectors Thus, the central feature vector with the highest reliability can be extracted. Then, the information processing system 230 may extract the voice of a specific speaker associated with the central feature vector from the input signal.
  • the information processing system 230 can separate the voices of each speaker included in the input signal by repeating the above-described process several times. Additionally, the information processing system 230 may provide the separated voices of each speaker and/or the results of performing additional processing on each speaker's voice to the user terminal 210 as a result of speaker segmentation. For example, the information processing system can generate meeting minutes or subtitles and transmit them to the user terminal 210 by performing automatic speech recognition (ASR) or STT based on the voices of each separated speaker.
  • ASR automatic speech recognition
  • the processor 334 of the information processing system 230 uses the output device 320, such as a display output capable device (e.g., touch screen, display, etc.), an audio output capable device (e.g., speaker), of the user terminal 210. It may be configured to output processed information and/or data.
  • the processor 334 of the information processing system 230 provides the separated voices of each speaker to the user terminal 210 through the communication module 336 and the network 220 as a result of speaker segmentation, and It may be configured to output the speaker's voice through a sound output capable device of the user terminal 210, etc.
  • the processor 334 of the information processing system 230 may perform automatic speech recognition (ASR) based on the voices of each speaker separated to the user terminal 210 through the communication module 336 and the network 220.
  • ASR automatic speech recognition
  • it may be configured to provide meeting minutes or subtitles generated by performing STT, etc., and output them through a display output capable device of the user terminal 210, etc.
  • FIG. 4 is a diagram illustrating an example of extracting a plurality of speaker feature vectors based on speech sections included in an input signal according to an embodiment of the present disclosure.
  • the information processing system or a processor of the information processing system may receive the input signal 410.
  • the input signal 410 may include the spoken voices (and surrounding noise) of several speakers.
  • the speaker may be a concept that includes not only a human, but also a virtual person or character capable of making synthesized voice utterances, and a software or hardware module capable of generating and outputting sound including voice.
  • the input signal 410 may include audio recorded during a conference with multiple participants.
  • the input signal 410 may include audio extracted from an image such as a drama or movie.
  • the input signal 410 may be various types of sound.
  • the input signal 410 may be a sound in the form of a waveform or a sound in the form of a spectrogram.
  • the information processing system may receive sound in the form of a waveform, perform frequency conversion based on the sound in the received waveform form, and use the sound converted into a spectrogram form as the input signal 410. there is.
  • the information processing system can then detect (420) the speech segment from the input signal (410).
  • the voice section may be a section of the input signal 110 that includes the speaker's spoken voice.
  • the information processing system can detect the voice section included in the input signal 410 by performing voice end point detection (EPD) to detect the start and end points of the utterance.
  • EPD voice end point detection
  • the information processing system may extract a plurality of speaker feature vectors 432 based on the detected voice section (430).
  • the speaker feature vector may include information about the characteristics of the spoken voice included in the voice section. In other words, the speaker feature vector may include information about the speaker who uttered the voice.
  • the information processing system extracts one speaker feature vector for each unit voice section, thereby extracting a plurality of speaker feature vectors 432 based on the detected voice section. For example, assuming an embodiment in which the length of a unit voice section is 0.5 seconds, if the length of the detected first voice section 422 is 1 second, the information processing system creates two voice sections based on the first voice section 422. Speaker feature vectors (in the example shown, E 1 and E 2 ) can be extracted. In addition, when the length of the detected second speech section 424 is 1.5 seconds, the information processing system generates three speaker feature vectors (in the example shown, E 3 , E 4 , and E ) based on the second speech section 424. 5 ) can be extracted.
  • the information processing system can extract a plurality of speaker feature vectors 432 from the detected voice sections using voice section detection and speaker feature vector extraction methods according to unit voice sections of various lengths.
  • the information processing system may extract a plurality of speaker feature vectors 432 based on the speech section using a speaker feature extraction model (Speaker Embedding Extractor).
  • the speaker feature extraction model may be a machine learning model learned to extract speaker feature vectors based on speech sections.
  • the speaker feature extraction model is learned to extract the same/similar speaker feature vectors from speech sections containing the speech voices of the same speaker, and to extract the same/similar speaker feature vectors from speech sections containing the speech voices of different speakers. It may be a machine learning model trained to extract speaker feature vectors from which feature differences can be well distinguished.
  • the information processing system may extract a plurality of speaker feature vectors 432 by processing each voice section (eg, unit voice section) detected from the input signal 410.
  • the information processing system can extract a plurality of speaker feature vectors 432 by calculating the average value of each speech section in the form of a spectrogram.
  • FIG. 5 is a diagram illustrating an example of performing clustering on a speaker feature vector according to an embodiment of the present disclosure.
  • the information processing system may perform clustering on a plurality of speaker feature vectors extracted from speech sections included in the input signal. When clustering is performed on a plurality of speaker feature vectors, similar feature vectors among the plurality of speaker feature vectors may be grouped.
  • an information processing system can perform clustering using any clustering method such as K-means Clustering or Spectral Clustering.
  • the information processing system may estimate the number of clusters from a plurality of extracted speaker feature vectors and group the plurality of speaker feature vectors into the estimated number of clusters.
  • a clustering process is necessarily included to distinguish the voices of each speaker included in the input signal, but according to the speaker segmentation method of the present disclosure, the clustering process can be performed selectively.
  • Figure 5 shows the results of clustering two sets of speaker feature vectors extracted based on two input signals containing speech voices of three speakers.
  • the first clustering result 520 is the result of clustering the first set of speaker feature vectors 510 extracted from the speech section included in the first input signal containing the speech voices of three speakers. Looking at the first clustering result 520, it can be seen that the first set of speaker feature vectors 510 are grouped into three clusters 522, 524, and 526. Speaker feature vectors belonging to the same cluster can be assumed to have been extracted from speech sections containing the same speaker's speech.
  • the second clustering result 540 is the result of clustering the second set of speaker feature vectors 530 extracted from the speech section included in the second input signal containing the speech voices of three speakers.
  • the second input signal may include an overlapping voice section 532 in which the spoken voices of the first speaker and the second speaker overlap.
  • the speaker feature vector 534 extracted based on the overlapping speech section 532 may be contaminated and may not well reflect the features of any one of the three speakers.
  • the speaker feature vector 534 extracted based on the overlapping voice section 532 is a region where the speaker feature vectors extracted from the voice section containing only the first speaker's spoken voice are located and only the second speaker's spoken voice. It may be located between areas where speaker feature vectors extracted from the included speech section are located.
  • the distinction between the area of the speaker feature vector corresponding to the first speaker and the area of the speaker feature vector corresponding to the second speaker becomes unclear due to the speaker feature vector 534 extracted based on the overlapping speech section 532. You can.
  • the second clustering result 540 which is the result of clustering the second set of speaker feature vectors 530 extracted based on the second input signal
  • the second input signal actually contains the speech voices of three speakers.
  • the number of speakers was incorrectly estimated, and it can be seen that the second set of speaker feature vectors 530 are grouped into two clusters 542 and 544.
  • the speaker feature vector may be contaminated by the overlapping speech section 532, which may cause an error in the speaker segmentation result. Additionally, contamination may occur in the speaker feature vector not only when an overlapping voice section 532 exists in the input signal, but also when the input signal includes noise such as ambient noise.
  • a representative speaker feature vector of a specific speaker can be extracted by extracting a central feature vector with the highest reliability based on a plurality of extracted speaker feature vectors. Therefore, even if there is an error in the clustering result, or without performing clustering, speaker segmentation can be performed without error.
  • the process of extracting the central feature vector with the highest reliability based on the extracted plurality of speaker feature vectors will be described later with reference to FIG. 6.
  • FIG. 6 is a diagram illustrating an example of extracting a central feature vector based on a plurality of speaker feature vectors extracted according to an embodiment of the present disclosure.
  • Figure 6 shows an example of a first set of speaker feature vectors 600 extracted based on an input signal containing the spoken voices of three speakers (Speaker A, Speaker B, and Speaker C).
  • the first set of speaker feature vectors 600 includes a first group 610 that includes a plurality of speaker feature vectors extracted based on speech sections containing only speaker A's spoken voice, and a first group 610 that includes only speaker B's spoken voice.
  • the input signal includes a section where the speech voices of two or more speakers overlap
  • the first group 610, the second group 620, and the third group (630) can all be classified into the same cluster. Therefore, when speaker segmentation is performed based only on the clustering result according to the conventional method, errors may occur in the speaker segmentation results.
  • the information processing system may extract a central feature vector with the highest reliability based on the extracted first set of speaker feature vectors 600. According to one embodiment, the information processing system may use the density of speaker feature vectors as a measure of reliability.
  • the information processing system may first determine the area where the speaker feature vectors are most dense within the vector space as the central area. As a specific example, the information processing system selects the region containing the largest number of speaker feature vectors among a plurality of regions of the same size centered on each of the first set of speaker feature vectors 600 extracted within the vector space as the center region. can be decided. In the illustrated example, six speaker feature vectors excluding the first feature vector 650 exist in the first area 652 centered on the first feature vector 650. In addition, in the second area 662 (an area of the same shape and size as the first area 652) centered on the second feature vector 660, five speaker feature vectors excluding the second feature vector 660 are present. exist. Therefore, in this case, the first area 652 may be determined as the center area.
  • the information processing system may determine the central feature vector based on one or more speaker feature vectors included in the central region. For example, the information processing system may determine the average of one or more speaker feature vectors included in the central region as the central feature vector. According to this embodiment, the information processing system stores the seven speaker feature vectors included in the first region 652 determined as the central region.
  • the average can be determined as the central feature vector.
  • the information processing system may extract the speaker feature vector located at the center of the central region or closest to the center as the central feature vector.
  • the information processing system may extract the first feature vector 650 located at the center of the first region 652 determined as the center region as the center feature vector.
  • the first feature vector 650 extracted as the central feature vector is based on the densest region among the first group 610 including a plurality of speaker feature vectors extracted based on a voice section containing only speaker A's speech voice. Extracted, it can be assumed to best represent Speaker A's voice.
  • the information processing system determines the centroids from the cluster containing the largest number of speaker feature vectors as a result of the clustering.
  • Feature vectors can be extracted. For example, as a result of performing clustering on the first set of speaker feature vectors 600, a first cluster including a first group 610, a second group 620, and a third group 630, It may be classified into a second cluster including 4 groups 640.
  • the central region can be determined in the space corresponding to the first cluster containing the largest number of speaker feature vectors, and the central feature vector can be extracted based on one or more speaker feature vectors included in the central region. .
  • a representative speaker feature vector for a specific speaker can be extracted by extracting the central feature vector with the highest reliability, even if there is an error in the clustering result.
  • speaker segmentation can be performed with high accuracy without performing clustering.
  • FIG. 7 is a diagram illustrating an example of extracting a specific speaker's voice 720 from an input signal 710 according to an embodiment of the present disclosure.
  • the information processing system may extract the voice 720 of a specific speaker associated with the central feature vector from the input signal 710.
  • the information processing system may use the target voice extraction model 700 to extract the voice of the first speaker associated with the first central feature vector from the input signal.
  • Specific examples and learning methods of the target voice extraction model 700 used to extract the voice 720 of a specific speaker will be described in detail later with reference to FIGS. 8 and 9.
  • the information processing system may determine whether a voice section remains in the remaining signal excluding the first speaker's voice extracted from the input signal 710. If it is determined that a voice section remains in the remaining signal, at least some of the series of processes described above with reference to FIGS. 4 to 7 may be repeatedly performed.
  • the information processing system extracts a plurality of speaker feature vectors based on the remaining voice sections excluding the voice section containing the extracted voice 720 of a specific speaker among the voice sections included in the input signal 710.
  • a second central feature vector with the highest reliability is extracted based on the plurality of extracted speaker feature vectors, and the second speaker's voice associated with the second central feature vector extracted from the input signal 710 or the residual signal is extracted. It can be extracted.
  • the information processing system may extract a first portion extracted based on the remaining speech section excluding the speech section containing the voice 720 of a specific speaker among the first set of speaker feature vectors extracted based on the input signal 710. Based on the set of speaker feature vectors, a second central feature vector with the highest reliability can be extracted, and the second speaker's voice associated with the second central feature vector extracted from the input signal 710 or the residual signal can be extracted. there is.
  • the information processing system extracts a plurality of speaker feature vectors based on speech sections included in the residual signal, and extracts a second central feature vector with the highest reliability based on the plurality of extracted speaker feature vectors. And, the second speaker's voice associated with the second center feature vector extracted from the input signal 710 or the residual signal can be extracted.
  • the information processing system may transmit the separated voices of each speaker and/or the results of performing additional processing on each speaker's voice to the user terminal as a result of speaker segmentation.
  • the information processing system performs Automatic Speech Recognition (ASR) or STT (Speech to Text) based on the voices of each separated speaker to generate meeting minutes or subtitles and send them to the user terminal. Can be transmitted.
  • ASR Automatic Speech Recognition
  • STT Speech to Text
  • FIG. 8 is a diagram illustrating a specific example of a target voice extraction model 800 according to an embodiment of the present disclosure.
  • the target voice extraction model 800 receives the mixed voice 810 and the target speaker feature vector 820 as input, and outputs a mask 830 associated with the target speaker feature vector 820. It can be.
  • the mixed voice 810 may be a voice that is a mixture of the target speaker's voice and another speaker's voice and/or noise.
  • the mixed voice 810 may be a voice in the form of a waveform or a voice in the form of a spectrogram.
  • the voice in the form of a waveform can be used as the mixed voice 810 by performing frequency conversion on the voice in the form of a waveform and converting it into a spectrogram form.
  • the target speaker feature vector 820 may be a vector containing information about the target speaker. Additionally, the mask 830 associated with the target speaker feature vector 820 may be a mask for removing voices and noise of speakers other than the target speaker from the input signal. Using the output mask 830 and mixed voice 810, the target speaker's voice 840 can be extracted. For example, by multiplying the mixed voice 810 by the mask 830, the target speaker's voice 840 can be extracted.
  • the extracted voice 840 of the target speaker may also be in the form of a spectrogram.
  • the voice of the target speaker in the form of a spectrogram can be converted into a voice in the form of a waveform by performing inverse frequency conversion.
  • the information processing system inputs the input signal as a mixed voice 810 and the extracted central feature vector as a target speaker feature vector 820 to the target speech extraction model 800, and
  • the mask 830 can be estimated, and the voice of a specific speaker can be extracted by multiplying the input signal by the mask 830.
  • the information processing system may use another speaker feature vector extracted by a method similar to or different from the central feature vector extraction method described above as the target speaker feature vector 820.
  • the target voice extraction model 800 shown in FIG. 8 is only an example, and in other embodiments, the target voice extraction model 800 may be implemented differently from the model shown.
  • the target voice extraction model 800 may be implemented to directly output the target speaker's voice 840 instead of outputting the mask 830 associated with the target speaker feature vector 820.
  • FIG. 9 is a diagram illustrating an example of a method for training a target voice extraction model 800 according to an embodiment of the present disclosure.
  • the target voice extraction model 800 receives the mixed voice 810 and the target speaker feature vector 820 as input and learns to output a mask 830 associated with the target speaker feature vector 820. It may be a machine learning model.
  • the target speech extraction model 800 may be an artificial neural network model including at least one of a convolutional neural network (CNN), long short term memory (LSTM), and fully connected (FC) layer.
  • CNN convolutional neural network
  • LSTM long short term memory
  • FC fully connected
  • the information processing system uses a first target speaker reference voice 822, a second target speaker reference voice 812, and another speaker reference voice 814 as training data for training the target voice extraction model 800. ) can be obtained.
  • the first target speaker reference voice 822 and the second target speaker reference voice 812 may be voices that include only the target speaker's speech voice.
  • the first target speaker reference voice 822 and the second target speaker reference voice 812 may be the same or different from each other.
  • the other speaker reference voice 814 may be a voice that includes speech and/or noise from a speaker other than the target speaker.
  • the information processing system may generate a mixed voice 810 by mixing the second target speaker reference voice 812 and another speaker reference voice 814.
  • Each of the voices 810, 812, 814, and 822 described above may be a voice in the form of a wave form or a voice in the form of a spectrogram. According to one embodiment, frequency conversion can be performed on the waveform-type voice and converted into spectrogram-type voice for use.
  • the information processing system may extract the target speaker feature vector 820 based on the first target speaker reference voice 822. Based on the first target speaker reference voice 822, the process of extracting the target speaker feature vector 820 may be performed the same or similar to the method described above with reference to FIG. 4.
  • the information processing system can then input the mixed voice 810 and the target speaker feature vector 820 into the target voice extraction model 800 to estimate the mask 830 associated with the target speaker feature vector 820. . Then, the target speaker's voice 840 can be extracted by multiplying the mixed voice 810 by the estimated mask 830.
  • the information processing system may calculate a prediction loss (Loss) 850 based on a comparison of the extracted target speaker's voice 840 and the second target speaker reference voice 812. Then, the weight corresponding to the connection between the plurality of layers or nodes included in the target speech extraction model 800 can be updated using the calculated prediction loss 850.
  • a prediction loss Loss
  • the target voice extraction model 800 may be implemented differently from the model shown or learned by a different method.
  • FIG. 10 is a flowchart illustrating an example of a speaker segmentation method 1000 according to an embodiment of the present disclosure.
  • the speaker segmentation method 1000 allows a processor (e.g., at least one processor of an information processing system or a user terminal) to generate a first set of speaker feature vectors based on speech sections included in the input signal. It can be initiated by extracting (S1010).
  • the processor may detect speech segments from an input signal and extract a first set of speaker feature vectors based on the detected speech segments.
  • the first set of speaker feature vectors may include a plurality of speaker feature vectors.
  • the processor may extract a plurality of speaker feature vectors based on each of a plurality of unit speech sections included in the input signal.
  • the processor may extract the first central feature vector with the highest reliability based on the extracted first set of speaker feature vectors (S1020).
  • the processor may first determine a first central region of the vector space based on the extracted first set of speaker feature vectors. In one embodiment, the processor may determine the area in the vector space where speaker feature vectors are most densely located as the first center area. For example, the processor may determine as the first center region the region containing the largest number of speaker feature vectors among a plurality of regions of equal size centered on each of the first set of speaker feature vectors extracted within the vector space. You can.
  • the processor may extract the first central feature vector based on one or more speaker feature vectors included in the first central region. For example, the processor may determine the average of one or more speaker feature vectors included in the first central region as the first central feature vector. As another example, the processor may extract a speaker feature vector located at the center of the first center area or closest to the center as the first center feature vector.
  • the processor may perform clustering on the extracted first set of speaker feature vectors before extracting the first central feature vector.
  • the processor may extract the first central feature vector from the cluster that includes the largest number of speaker feature vectors as a result of clustering. For example, the processor determines a first central region within the vector space corresponding to a cluster containing the largest number of speaker feature vectors as a result of clustering, based on one or more speaker feature vectors included in the first central region.
  • the first central feature vector can be extracted.
  • the processor may extract the voice of the first speaker associated with the first central feature vector from the input signal (S1030).
  • the processor may extract the voice of the first speaker associated with the first center feature vector from the input signal using the target voice extraction model.
  • the processor may use the target sound source extraction model to estimate a mask associated with the first center feature vector based on the input signal and the first center feature vector, and based on the input signal and the estimated mask, 1 The speaker's voice can be extracted.
  • the processor may determine whether a voice section remains in the remaining signal excluding the first speaker's voice extracted from the input signal. If it is determined that a voice section remains in the residual signal, the voice of another speaker included in the input signal or the residual signal can be extracted by repeatedly performing the above-described series of processes.
  • the processor may return to step S1010 and repeat a series of processes. Specifically, the processor may extract a second set of speaker feature vectors based on speech sections included in the residual signal. Then, based on the extracted second set of speaker feature vectors, a second central feature vector with the highest reliability is extracted, and the second speaker's voice associated with the second central feature vector is extracted from the input signal or residual signal. You can.
  • the processor may return to step S1020 and repeat a series of processes. Specifically, the processor extracts a second center feature vector with the highest reliability based on a speaker feature vector extracted based on a speech section included in the residual signal among the first set of speaker feature vectors, and extracts the second center feature vector from the input signal or the residual signal. The second speaker's voice associated with the second central feature vector may be extracted from the signal.
  • the processor extracts a second central feature vector with the highest reliability based on the remaining voice sections excluding the voice section containing the first speaker's voice among the voice sections included in the input signal, and extracts the second central feature vector from the input signal or The second speaker's voice associated with the second central feature vector may be extracted from the residual signal.
  • the above-described method may be provided as a computer program stored in a computer-readable recording medium for execution on a computer.
  • Media may be used to continuously store executable programs on a computer, or may be temporarily stored for execution or download.
  • the medium may be a variety of recording or storage means in the form of a single or several pieces of hardware combined. It is not limited to a medium directly connected to a computer system and may be distributed over a network. Examples of media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and There may be something configured to store program instructions, including ROM, RAM, flash memory, etc. Additionally, examples of other media include recording or storage media managed by app stores that distribute applications, sites or servers that supply or distribute various other software, etc.
  • the processing units used to perform the techniques may include one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs). ), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, and other electronic units designed to perform the functions described in this disclosure. , a computer, or a combination thereof.
  • the various illustrative logical blocks, modules, and circuits described in connection with this disclosure may be general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or It may be implemented or performed as any combination of those designed to perform the functions described in.
  • a general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine.
  • a processor may also be implemented as a combination of computing devices, such as a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other configuration.
  • RAM random access memory
  • ROM read-only memory
  • NVRAM non-volatile random access memory
  • PROM on computer-readable media such as programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage devices, etc. It may also be implemented as stored instructions. Instructions may be executable by one or more processors and may cause the processor(s) to perform certain aspects of the functionality described in this disclosure.
  • Computer-readable media includes both computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another.
  • Storage media may be any available media that can be accessed by a computer.
  • such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or the desired program code in the form of instructions or data structures. It can be used to transfer or store data and can include any other media that can be accessed by a computer. Any connection is also properly termed a computer-readable medium.
  • disk and disk include CD, laser disk, optical disk, digital versatile disc (DVD), floppy disk, and Blu-ray disk, where disks are usually magnetic. It reproduces data optically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
  • a software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known.
  • An exemplary storage medium may be coupled to the processor such that the processor may read information from or write information to the storage medium. Alternatively, the storage medium may be integrated into the processor.
  • the processor and storage medium may reside within an ASIC. ASIC may exist within the user terminal. Alternatively, the processor and storage medium may exist as separate components in the user terminal.

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Telephonic Communication Services (AREA)

Abstract

본 개시는 적어도 하나의 프로세서에 의해 수행되는 화자 분할 방법에 관한 것이다. 화자 분할 방법은, 입력 신호에 포함된 음성 구간에 기초하여, 제1 세트의 화자 특징 벡터를 추출하는 단계, 추출된 제1 세트의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제1 중심 특징 벡터를 추출하는 단계 및 입력 신호로부터 제1 중심 특징 벡터와 연관된 제1 화자의 음성을 추출하는 단계를 포함한다.

Description

화자 분할 방법 및 시스템
본 개시는 화자 분할 방법 및 시스템에 관한 것으로, 구체적으로, 화자 특징 벡터 중 가장 높은 신뢰도를 갖는 중심 특징 벡터를 추출하고, 추출된 중심 특징 벡터와 연관된 특정 화자의 음성을 추출하는 화자 분할 방법 및 시스템에 관한 것이다.
IT 기술의 발달 및 음성 인식 기술의 발달로 인해, 다양한 음성 인식 관련 서비스들이 제공되고 있다. 예를 들어, 회의 녹음 파일에 기초하여 자동으로 회의록을 생성해주는 회의록 생성 서비스, 드라마, 영화, 동영상 등의 자막을 자동으로 생성해주는 자동 자막 생성 서비스 등이 제공되고 있다.
이와 같은 음성 인식 기능을 포함하는 서비스에서, 인식 대상 음성 데이터에 여러 화자의 음성이 포함된 경우, 여러 화자의 발화 구간을 구분하여 인식하기 위해 화자 분할이 수행될 수 있다. 다만, 인식 대상 음성에 잡음이 섞이거나 여러 화자의 음성이 중첩된 구간이 있는 경우, 화자 특징 벡터가 오염되어 화자 분할 결과에 오차가 발생할 수 있다는 문제점이 있다.
본 개시는 상기와 같은 문제점을 해결하기 위한 화자 분할 방법, 명령어들을 기록한 컴퓨터 판독가능한 비일시적 기록 매체 및 장치(시스템)를 제공한다.
본 개시는 방법, 장치(시스템) 또는 명령어들을 기록한 컴퓨터 판독가능한 비일시적 기록 매체를 포함한 다양한 방식으로 구현될 수 있다.
본 개시의 일 실시예에 따르면, 적어도 하나의 프로세서에 의해 수행되는, 화자 분할 방법은, 입력 신호에 포함된 음성 구간에 기초하여, 제1 세트의 화자 특징 벡터를 추출하는 단계, 추출된 제1 세트의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제1 중심 특징 벡터를 추출하는 단계 및 입력 신호로부터 제1 중심 특징 벡터와 연관된 제1 화자의 음성을 추출하는 단계를 포함한다.
본 개시의 일 실시예에 따른 방법을 컴퓨터에서 실행하기 위한 명령어들을 기록한 컴퓨터 판독가능한 비일시적 기록 매체가 제공된다.
본 개시의 일 실시예에 따른 정보 처리 시스템은, 메모리 및 메모리와 연결되고, 메모리에 포함된 컴퓨터 판독 가능한 적어도 하나의 프로그램을 실행하도록 구성된 적어도 하나의 프로세서를 포함하고, 적어도 하나의 프로그램은, 입력 신호에 포함된 음성 구간에 기초하여, 제1 세트의 화자 특징 벡터를 추출하고, 추출된 제1 세트의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제1 중심 특징 벡터를 추출하고, 입력 신호로부터 제1 중심 특징 벡터와 연관된 제1 화자의 음성을 추출하기 위한 명령어들을 포함한다.
본 개시의 일부 실시예에 따르면, 음성 구간으로부터 추출된 화자 특징 벡터들 중에서 특정 화자를 나타내는데 있어서 신뢰도가 높은 중심 특징 벡터를 추출함으로써, 화자 특징 벡터들 중 일부가 오염된 경우에도, 특정 화자의 음성 특징을 잘 반영하는 대표성 있는 화자 특징 벡터를 추출할 수 있다.
본 개시의 일부 실시예에 따르면, 입력 신호로부터 중심 특징 벡터와 연관된 특정 화자의 음성이 제거된 잔여 신호에 대해서 중심 특징 벡터의 추출과 이와 연관된 다른 특정 화자의 음성을 추출하는 일련의 과정을 반복하여 실시함으로써, 화자 인식 및 음성 추출 과정에서 여러 화자의 음성이 중첩되는 구간이나 여러 화자의 음성이 밀접한 구간이 줄어들 수 있어, 결과적으로 더 신뢰도를 갖고 하나 이상의 화자를 나타낼 수 있는 화자 특징 벡터를 추출할 수 있다.
본 개시의 효과는 이상에서 언급한 효과로 제한되지 않으며, 언급되지 않은 다른 효과들은 청구범위의 기재로부터 본 개시가 속하는 기술분야에서 통상의 지식을 가진 자(“통상의 기술자”라 함)에게 명확하게 이해될 수 있을 것이다.
본 개시의 실시예들은, 이하 설명하는 첨부 도면들을 참조하여 설명될 것이며, 여기서 유사한 참조 번호는 유사한 요소들을 나타내지만, 이에 한정되지는 않는다.
도 1은 본 개시의 일 실시예에 따른 화자 분할 방법의 예시를 나타내는 도면이다.
도 2는 본 개시의 일 실시예에 따른 정보 처리 시스템이 복수의 사용자 단말과 통신 가능하도록 연결된 구성을 나타내는 개요도이다.
도 3은 본 개시의 일 실시예에 따른 사용자 단말 및 정보 처리 시스템의 내부 구성을 나타내는 블록도이다.
도 4는 본 개시의 일 실시예에 따라 입력 신호에 포함된 음성 구간에 기초하여 복수의 화자 특징 벡터를 추출하는 예시를 나타내는 도면이다.
도 5는 본 개시의 일 실시예에 따라 화자 특징 벡터에 대해 클러스터링을 수행하는 예시를 나타내는 도면이다.
도 6은 본 개시의 일 실시예에 따라 추출된 복수의 화자 특징 벡터에 기초하여 중심 특징 벡터를 추출하는 예시를 나타내는 도면이다.
도 7은 본 개시의 일 실시예에 따라 입력 신호로부터 특정 화자의 음성을 추출하는 예시를 나타내는 도면이다.
도 8은 본 개시의 일 실시예에 따른 목적 음성 추출 모델의 예시를 나타내는 도면이다.
도 9는 본 개시의 일 실시예에 따른 목적 음성 추출 모델을 학습시키는 방법의 예시를 나타내는 도면이다.
도 10은 본 개시의 일 실시예에 따른 화자 분할 방법의 예시를 나타내는 흐름도이다.
이하, 본 개시의 실시를 위한 구체적인 내용을 첨부된 도면을 참조하여 상세히 설명한다. 다만, 이하의 설명에서는 본 개시의 요지를 불필요하게 흐릴 우려가 있는 경우, 널리 알려진 기능이나 구성에 관한 구체적 설명은 생략하기로 한다.
첨부된 도면에서, 동일하거나 대응하는 구성요소에는 동일한 참조부호가 부여되어 있다. 또한, 이하의 실시예들의 설명에 있어서, 동일하거나 대응되는 구성요소를 중복하여 기술하는 것이 생략될 수 있다. 그러나, 구성요소에 관한 기술이 생략되어도, 그러한 구성요소가 어떤 실시예에 포함되지 않는 것으로 의도되지는 않는다.
개시된 실시예의 이점 및 특징, 그리고 그것들을 달성하는 방법은 첨부되는 도면과 함께 후술되어 있는 실시예들을 참조하면 명확해질 것이다. 그러나, 본 개시는 이하에서 개시되는 실시예들에 한정되는 것이 아니라 서로 다른 다양한 형태로 구현될 수 있으며, 단지 본 실시예들은 본 개시가 완전하도록 하고, 본 개시가 통상의 기술자에게 발명의 범주를 완전하게 알려주기 위해 제공되는 것일 뿐이다.
본 명세서에서 사용되는 용어에 대해 간략히 설명하고, 개시된 실시예에 대해 구체적으로 설명하기로 한다. 본 명세서에서 사용되는 용어는 본 개시에서의 기능을 고려하면서 가능한 현재 널리 사용되는 일반적인 용어들을 선택하였으나, 이는 관련 분야에 종사하는 기술자의 의도 또는 판례, 새로운 기술의 출현 등에 따라 달라질 수 있다. 또한, 특정한 경우는 출원인이 임의로 선정한 용어도 있으며, 이 경우 해당되는 발명의 설명 부분에서 상세히 그 의미를 기재할 것이다. 따라서, 본 개시에서 사용되는 용어는 단순한 용어의 명칭이 아닌, 그 용어가 가지는 의미와 본 개시의 전반에 걸친 내용을 토대로 정의되어야 한다.
본 명세서에서의 단수의 표현은 문맥상 명백하게 단수인 것으로 특정하지 않는 한, 복수의 표현을 포함한다. 또한, 복수의 표현은 문맥상 명백하게 복수인 것으로 특정하지 않는 한, 단수의 표현을 포함한다. 명세서 전체에서 어떤 부분이 어떤 구성요소를 포함한다고 할 때, 이는 특별히 반대되는 기재가 없는 한 다른 구성요소를 제외하는 것이 아니라 다른 구성요소를 더 포함할 수 있음을 의미한다.
또한, 명세서에서 사용되는 '모듈' 또는 '부'라는 용어는 소프트웨어 또는 하드웨어 구성요소를 의미하며, '모듈' 또는 '부'는 어떤 역할들을 수행한다. 그렇지만, '모듈' 또는 '부'는 소프트웨어 또는 하드웨어에 한정되는 의미는 아니다. '모듈' 또는 '부'는 어드레싱할 수 있는 저장 매체에 있도록 구성될 수도 있고 하나 또는 그 이상의 프로세서들을 재생시키도록 구성될 수도 있다. 따라서, 일 예로서, '모듈' 또는 '부'는 소프트웨어 구성요소들, 객체지향 소프트웨어 구성요소들, 클래스 구성요소들 및 태스크 구성요소들과 같은 구성요소들과, 프로세스들, 함수들, 속성들, 프로시저들, 서브루틴들, 프로그램 코드의 세그먼트들, 드라이버들, 펌웨어, 마이크로 코드, 회로, 데이터, 데이터베이스, 데이터 구조들, 테이블들, 어레이들 또는 변수들 중 적어도 하나를 포함할 수 있다. 구성요소들과 '모듈' 또는 '부'들은 안에서 제공되는 기능은 더 작은 수의 구성요소들 및 '모듈' 또는 '부'들로 결합되거나 추가적인 구성요소들과 '모듈' 또는 '부'들로 더 분리될 수 있다.
본 개시의 일 실시예에 따르면, '모듈' 또는 '부'는 프로세서 및 메모리로 구현될 수 있다. '프로세서'는 범용 프로세서, 중앙 처리 장치(CPU), 마이크로프로세서, 디지털 신호 프로세서(DSP), 제어기, 마이크로제어기, 상태 머신 등을 포함하도록 넓게 해석되어야 한다. 몇몇 환경에서, '프로세서'는 주문형 반도체(ASIC), 프로그램가능 로직 디바이스(PLD), 필드 프로그램가능 게이트 어레이(FPGA) 등을 지칭할 수도 있다. '프로세서'는, 예를 들어, DSP와 마이크로프로세서의 조합, 복수의 마이크로프로세서들의 조합, DSP 코어와 결합한 하나 이상의 마이크로프로세서들의 조합, 또는 임의의 다른 그러한 구성들의 조합과 같은 처리 디바이스들의 조합을 지칭할 수도 있다. 또한, '메모리'는 전자 정보를 저장 가능한 임의의 전자 컴포넌트를 포함하도록 넓게 해석되어야 한다. '메모리'는 임의 액세스 메모리(RAM), 판독-전용 메모리(ROM), 비-휘발성 임의 액세스 메모리(NVRAM), 프로그램가능 판독-전용 메모리(PROM), 소거-프로그램가능 판독 전용 메모리(EPROM), 전기적으로 소거가능 PROM(EEPROM), 플래쉬 메모리, 자기 또는 광학 데이터 저장장치, 레지스터들 등과 같은 프로세서-판독가능 매체의 다양한 유형들을 지칭할 수도 있다. 프로세서가 메모리로부터 정보를 판독하고/하거나 메모리에 정보를 기록할 수 있다면 메모리는 프로세서와 전자 통신 상태에 있다고 불린다. 프로세서에 집적된 메모리는 프로세서와 전자 통신 상태에 있다.
본 개시에서, '시스템'은 서버 장치와 클라우드 장치 중 적어도 하나의 장치를 포함할 수 있으나, 이에 한정되는 것은 아니다. 예를 들어, 시스템은 하나 이상의 서버 장치로 구성될 수 있다. 다른 예로서, 시스템은 하나 이상의 클라우드 장치로 구성될 수 있다. 또 다른 예로서, 시스템은 서버 장치와 클라우드 장치가 함께 구성되어 동작될 수 있다.
본 개시에서, '기계학습 모델'은 주어진 입력에 대한 해답(answer)을 추론하는데 사용하는 임의의 모델을 포함할 수 있다. 일 실시예에 따르면, 기계학습 모델은 입력 레이어(층), 복수 개의 은닉 레이어 및 출력 레이어를 포함한 인공신경망 모델을 포함할 수 있다. 여기서, 각 레이어는 복수의 노드를 포함할 수 있다. 본 개시에서, 기계학습 모델은 인공신경망 모델을 지칭할 수 있으며, 인공신경망 모델은 기계학습 모델을 지칭할 수 있다.
본 개시에서, '디스플레이'는 컴퓨팅 장치와 연관된 임의의 디스플레이 장치를 지칭할 수 있는데, 예를 들어, 컴퓨팅 장치에 의해 제어되거나 컴퓨팅 장치로부터 제공된 임의의 정보/데이터를 표시할 수 있는 임의의 디스플레이 장치를 지칭할 수 있다.
본 개시에서, '복수의 A의 각각' 또는 '복수의 A 각각'은 복수의 A에 포함된 모든 구성 요소의 각각을 지칭하거나, 복수의 A에 포함된 일부 구성 요소의 각각을 지칭할 수 있다.
본 개시에서, '신뢰도'란, 특정 화자 특징 벡터가 특정 화자의 음성 특징을 대표할 것으로 추정되는 확률을 지칭할 수 있다. 따라서, 본 개시에서, 가장 높은 신뢰도를 갖는 화자 특징 벡터란, 특정 화자의 음성 특징을 가장 잘 대표할 것으로 추정되는 화자 특징 벡터, 또는 특정 화자의 음성 특징을 가장 잘 반영할 것으로 추정되는 화자 특징 벡터를 지칭할 수 있다. 화자 특징 벡터의 신뢰도는 다양한 척도에 의해 측정될 수 있다. 일 실시예에 따르면, 특정 입력 신호로부터 추출된 복수의 화자 특징 벡터가 위치한 벡터 공간 내에서, 특정 화자 특징 벡터의 주위 영역에 해당 화자 특징 벡터들이 많이 밀집되어 있을수록 높은 신뢰도를 가질 수 있다.
본 개시에서, '중심 특징 벡터'는 특정 입력 신호로부터 추출된 복수의 화자 특징 벡터 중 가장 높은 신뢰도를 갖는 화자 특징 벡터를 지칭할 수 있다.
도 1은 본 개시의 일 실시예에 따른 화자 분할 방법의 예시를 나타내는 도면이다. 화자 분할 방법은 복수의 화자의 음성이 포함된 입력 신호(110)를 분석하여 각 화자의 발화 구간을 구분하여 인식하고, 해당 화자의 음성을 추출할 수 있다. 먼저, 정보 처리 시스템은 입력 신호(110)를 수신할 수 있다. 여기서, 입력 신호(110)는 복수의 화자의 발화 음성 및 배경 소음을 포함할 수 있다.
그런 다음, 정보 처리 시스템은 입력 신호(110)로부터 음성 구간을 검출할 수 있다(120). 여기서, 음성 구간은 입력 신호(110) 중 화자의 발화 음성이 포함된 구간일 수 있다. 음성 구간이 검출된 경우, 정보 처리 시스템은 검출된 음성 구간에 대응하는 입력 신호(110)의 적어도 일부에 기초하여 복수의 화자 특징 벡터를 추출할 수 있다(130). 정보 처리 시스템이 입력 신호(110)로부터 음성 구간을 검출하고, 복수의 화자 특징 벡터를 추출하는 과정은 도 4를 참조하여 상세히 후술된다.
정보 처리 시스템은 추출된 복수의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 중심 특징 벡터를 추출할 수 있다(140). 여기서, 특정 화자란 입력 신호(110)에 발화 음성이 포함된 복수의 화자들 중 사전 결정된 한 명의 화자를 지칭하는 것 대신에, 발화가 포함된 여러 명의 화자 중에서 중심 특징 벡터와 연관된 음성을 발화한 임의의 한 화자(any one speaker)를 지칭할 수 있다. 따라서, 복수의 화자 특징 벡터 중 가장 높은 신뢰도를 갖는 중심 특징 벡터는, 입력 신호(110)에 발화 음성이 포함된 복수의 화자들 중 구체적으로 어느 화자인지를 특정하거나 사전에 인지할 수 없으나, 임의의 한 화자의 음성 특징을 가장 잘 대표할 것으로 추정되는 화자 특징 벡터를 지칭할 수 있다.
일 실시예에 따르면, 화자와 연관된 정보를 획득할 수 있는 경우, 화자 식별(Speaker Identification; SID) 과정이 추가적으로 수행될 수도 있다. 예를 들어, 일부 화자의 음성 샘플 또는 일부 화자의 음성 샘플로부터 추출된 화자 특징 정보를 사전에 획득한 경우, 정보 처리 시스템은 화자 식별 과정을 추가적으로 수행함으로써, 입력 신호(110)로부터 추출된 화자 특징 벡터가 어느 화자와 연관되는 지 추정할 수 있다. 이 경우, 중심 특징 벡터는, 입력 신호(110)에 포함된 복수의 화자들 중 사전 결정되거나 사전 식별된 화자의 음성 특징을 가장 잘 대표할 것으로 추정되는 화자 특징 벡터를 지칭하는 의미로 사용될 수도 있다. 다만, 이하의 설명에서는 '특정 화자'를 입력 신호에 발화 음성이 포함된 복수의 화자 중 임의의 한 화자(any one speaker)를 지칭하는 의미로 사용하고자 한다. 정보 처리 시스템이 추출된 복수의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 중심 특징 벡터를 추출하는 과정은 도 6을 참조하여 상세히 후술된다. 상술한 바와 같이, 복수의 화자 특징 벡터에 기초하여 가장 신뢰도가 높은 중심 특징 벡터를 추출함으로써, 화자 특징 벡터들 중 일부가 오염된 경우에도, 특정 화자의 음성 특징을 잘 반영하는 대표성 있는 화자 특징 벡터를 추출할 수 있다.
중심 특징 벡터를 추출한 정보 처리 시스템은, 입력 신호(110)로부터 중심 특징 벡터와 연관된 특정 화자의 음성(152)을 추출할 수 있다(150). 예를 들어, 정보 처리 시스템은 목적 음성 추출 모델을 이용하여 입력 신호(110)로부터 중심 특징 벡터와 연관된 특정 화자의 음성(152)을 추출할 수 있다. 정보 처리 시스템이 입력 신호(110)로부터 중심 특징 벡터와 연관된 특정 화자의 음성(152)을 추출하는 과정은 도 7 및 도 8을 참조하여 상세히 후술된다.
추가적으로, 정보 처리 시스템은 입력 신호(110)로부터 추출된 특정 화자의 음성(152)을 제외한 잔여 신호(154)에 음성 구간이 남아 있는지 여부를 판정할 수 있다(160). 잔여 신호(154)에 음성 구간이 남아 있다고 판정되는 경우, 상술한 일련의 과정을 반복하여 수행함으로써 입력 신호(110)에 포함된 다른 화자(들)의 음성을 추출할 수 있다.
예를 들어, 정보 처리 시스템은 입력 신호(110)에 포함된 음성 구간 중, 추출된 특정 화자의 음성(152)이 포함된 구간을 제외한 잔여 음성 구간에 대해, 화자 특징 벡터 추출(130), 중심 특징 벡터 추출(140), 특정 화자 음성 추출(150) 등의 과정을 반복 수행함으로써, 다른 화자(들)의 음성을 추출할 수 있다.
다른 예로, 정보 처리 시스템은 입력 신호(110)에 기초하여 추출된 복수의 화자 특징 벡터 중 특정 화자의 음성(152)이 포함된 구간을 제외한 잔여 음성 구간에 기초하여 추출된 일부의 화자 특징 벡터에 대해, 중심 특징 벡터 추출(140), 특정 화자 음성 추출(150) 등의 과정을 반복 수행함으로써, 다른 화자(들)의 음성을 추출할 수 있다.
또 다른 예로, 정보 처리 시스템은 잔여 신호(154)를 입력 신호(110)로 하여, 상술한 일련의 과정(120, 130, 140, 150 등)을 반복하여 수행할 수 있다.
상술한 과정을 잔여 신호(154)에 음성 구간이 더 이상 남아 있지 않을 때까지 반복하는 경우, 입력 신호(110)에 발화 음성이 포함된 모든 화자의 음성을 분리하여 인식할 수 있다.
상술한 바와 같이, 먼저, 가장 신뢰도가 높은 화자 특징 벡터를 이용하여, 입력 신호(110)로부터 특정 화자의 음성(152)을 추출함으로써, 특정 화자의 음성(152)을 높은 정확도를 갖고 분리할 수 있다. 그런 다음, 입력 신호(110)로부터 특정 화자의 음성(152)이 제거된 잔여 신호(154)에 대해 일련의 과정을 반복하여 실시함으로써, 이러한 과정 중에 여러 화자의 음성이 중첩되는 구간이나 여러 화자의 음성이 밀접한 구간이 점차적으로 감소할 수 있어, 잔여 신호(154)에 발화 음성이 포함된 나머지 화자에 대해서도 신뢰도가 높은 화자 특징 벡터가 추출될 수 있어, 고품질의 화자 분할을 수행할 수 있다.
이하에서는 도 2 내지 도 10을 참조하여, 본 개시의 화자 분할 방법에 대해 보다 상세히 후술된다. 이하에서는, 설명의 편의를 위해, 입력 신호(410)로부터 가장 먼저 중심 특징 벡터가 추출되는 음성과 연관된 화자를 제1 화자로 지칭하기로 한다. 이와 유사하게, n번째로 중심 특징 벡터가 추출되는 음성과 연관된 화자를 제n 화자로 지칭하기로 한다.
상술한 설명에서, 화자 분할 방법은 정보 처리 시스템에 의해 수행되는 것으로 기재되었으나, 이에 한정되지 않으며, 본 개시의 화자 분할 방법에 포함된 적어도 일부 단계가 사용자 단말에 의해 수행될 수 있다. 예를 들어, 본 개시의 화자 분할 방법과 연관된 일련의 과정(120, 130, 140, 150 등)이 사용자 단말에 의해 수행될 수 있다. 다만, 이하에서는 설명의 편의를 위해, 본 개시의 화자 분할 방법이 정보 처리 시스템에 의해 수행되는 것으로 가정하여 서술하고자 한다.
도 2는 본 개시의 일 실시예에 따른 정보 처리 시스템(230)이 복수의 사용자 단말(210_1, 210_2, 210_3)과 통신 가능하도록 연결된 구성을 나타내는 개요도이다. 도시된 바와 같이, 복수의 사용자 단말(210_1, 210_2, 210_3)은 네트워크(220)를 통해 화자 분할 서비스, 음성 인식 서비스 또는 화자 분할 및/또는 음성 인식 기능을 이용한 다양한 애플리케이션(예를 들어, 자동 회의록 생성 애플리케이션, 자동 자막 재생 애플리케이션 등)을 제공할 수 있는 정보 처리 시스템(230)과 연결될 수 있다. 여기서, 복수의 사용자 단말(210_1, 210_2, 210_3)은 화자 분할 서비스 및/또는 음성 인식 서비스를 제공받을 사용자의 단말을 포함할 수 있다. 일 실시예에서, 정보 처리 시스템(230)은 화자 분할 서비스 및/또는 음성 인식 서비스 등과 관련된 컴퓨터 실행 가능한 프로그램(예를 들어, 다운로드 가능한 애플리케이션) 및 데이터를 저장, 제공 및 실행할 수 있는 하나 이상의 서버 장치 및/또는 데이터베이스, 또는 클라우드 컴퓨팅 서비스 기반의 하나 이상의 분산 컴퓨팅 장치 및/또는 분산 데이터베이스를 포함할 수 있다.
대안적으로, 사용자 단말(210_1, 210_2, 210_3)은 네트워크(220)를 통하지 않고, 화자 분할 서비스, 음성 인식 서비스 등과 관련된 컴퓨터 실행 가능한 프로그램(예를 들어, 다운로드 가능한 애플리케이션)을 통해 사용자에게 화자 분할 서비스 및/또는 음성 인식 서비스를 제공할 수도 있다.
정보 처리 시스템(230)에 의해 제공되는 화자 분할 서비스 및/또는 음성 인식 서비스는, 복수의 사용자 단말(210_1, 210_2, 210_3)의 각각에 설치된 화자 분할 애플리케이션, 음성 인식 애플리케이션, 자동 회의록 생성 애플리케이션, 자동 자막 생성 애플리케이션, 음성 편집 애플리케이션, 모바일 브라우저 애플리케이션 또는 웹 브라우저 등을 통해 사용자에게 제공될 수 있다. 예를 들어, 정보 처리 시스템(230)은 화자 분할 애플리케이션 등을 통해 사용자 단말(210_1, 210_2, 210_3)로부터 수신되는 화자 분할 요청, 음성 인식 요청, 회의록 생성 요청, 자막 생성 요청 등에 대응하는 정보를 제공하거나 대응하는 처리를 수행할 수 있다.
복수의 사용자 단말(210_1, 210_2, 210_3)은 네트워크(220)를 통해 정보 처리 시스템(230)과 통신할 수 있다. 네트워크(220)는 복수의 사용자 단말(210_1, 210_2, 210_3)과 정보 처리 시스템(230) 사이의 통신이 가능하도록 구성될 수 있다. 네트워크(220)는 설치 환경에 따라, 예를 들어, 이더넷(Ethernet), 유선 홈 네트워크(Power Line Communication), 전화선 통신 장치 및 RS-serial 통신 등의 유선 네트워크, 이동통신망, WLAN(Wireless LAN), Wi-Fi, Bluetooth 및 ZigBee 등과 같은 무선 네트워크 또는 그 조합으로 구성될 수 있다. 통신 방식은 제한되지 않으며, 네트워크(220)가 포함할 수 있는 통신망(일례로, 이동통신망, 유선 인터넷, 무선 인터넷, 방송망, 위성망 등)을 활용하는 통신 방식뿐만 아니라 사용자 단말(210_1, 210_2, 210_3) 사이의 근거리 무선 통신 역시 포함될 수 있다.
도 2에서 휴대폰 단말(210_1), 태블릿 단말(210_2) 및 PC 단말 (210_3)이 사용자 단말의 예로서 도시되었으나, 이에 한정되지 않으며, 사용자 단말(210_1, 210_2, 210_3)은 유선 및/또는 무선 통신이 가능하고 화자 분할 애플리케이션, 음성 인식 애플리케이션, 자동 회의록 생성 애플리케이션, 자동 자막 생성 애플리케이션, 음성 편집 애플리케이션, 모바일 브라우저 애플리케이션 또는 웹 브라우저 등이 설치되어 실행될 수 있는 임의의 컴퓨팅 장치일 수 있다. 예를 들어, 사용자 단말은, AI 스피커, 스마트폰, 휴대폰, 내비게이션, 컴퓨터, 노트북, 디지털방송용 단말, PDA(Personal Digital Assistants), PMP(Portable Multimedia Player), 태블릿 PC, 게임 콘솔(game console), 웨어러블 디바이스(wearable device), IoT(internet of things) 디바이스, VR(virtual reality) 디바이스, AR(augmented reality) 디바이스, 셋톱 박스 등을 포함할 수 있다. 또한, 도 2에는 3개의 사용자 단말(210_1, 210_2, 210_3)이 네트워크(220)를 통해 정보 처리 시스템(230)과 통신하는 것으로 도시되어 있으나, 이에 한정되지 않으며, 상이한 수의 사용자 단말이 네트워크(220)를 통해 정보 처리 시스템(230)과 통신하도록 구성될 수도 있다.
일 실시예에 따르면, 정보 처리 시스템(230)은 복수의 사용자 단말(210_1, 210_2, 210_3)로부터 복수의 화자의 발화를 포함하는 입력 신호를 수신할 수 있다. 그리고 나서, 정보 처리 시스템(230)은 입력 신호에 포함된 음성 구간에 기초하여, 복수의 화자 특징 벡터를 추출하고, 추출된 복수의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 중심 특징 벡터를 추출할 수 있다. 그런 다음, 정보 처리 시스템(230)은 입력 신호로부터 중심 특징 벡터와 연관된 특정 화자의 음성을 추출할 수 있다. 정보 처리 시스템(230)은 상술한 과정을 여러 번 반복하여 수행함으로써 입력 신호에 포함된 각 화자의 음성을 분리할 수 있다. 추가적으로, 정보 처리 시스템(230)은 분리된 각 화자의 음성 및/또는 각 화자의 음성에 추가적인 처리를 수행한 결과를 화자 분할의 결과로서 사용자 단말(210)로 전송할 수 있다. 예를 들어, 정보 처리 시스템은 분리된 각 화자의 음성에 기초하여, 자동 음성 인식(Automatic Speech Recognition; ASR) 또는 STT(Speech to Text) 등을 수행함으로써, 회의록 또는 자막 등을 생성하여 사용자 단말(210)로 전송할 수 있다.
도 3은 본 개시의 일 실시예에 따른 사용자 단말(210) 및 정보 처리 시스템(230)의 내부 구성을 나타내는 블록도이다. 사용자 단말(210)은 화자 분할 애플리케이션, 음성 인식 애플리케이션, 자동 회의록 생성 애플리케이션, 자동 자막 생성 애플리케이션, 음성 편집 애플리케이션, 모바일 브라우저 애플리케이션 또는 웹 브라우저 등을 실행 가능하고 유/무선 통신이 가능한 임의의 컴퓨팅 장치를 지칭할 수 있으며, 예를 들어, 도 2의 휴대폰 단말(210_1), 태블릿 단말(210_2), PC 단말(210_3) 등을 포함할 수 있다. 도시된 바와 같이, 사용자 단말(210)은 메모리(312), 프로세서(314), 통신 모듈(316) 및 입출력 인터페이스(318)를 포함할 수 있다. 이와 유사하게, 정보 처리 시스템(230)은 메모리(332), 프로세서(334), 통신 모듈(336) 및 입출력 인터페이스(338)를 포함할 수 있다. 도 3에 도시된 바와 같이, 사용자 단말(210) 및 정보 처리 시스템(230)은 각각의 통신 모듈(316, 336)을 이용하여 네트워크(220)를 통해 정보 및/또는 데이터를 통신할 수 있도록 구성될 수 있다. 또한, 입출력 장치(320)는 입출력 인터페이스(318)를 통해 사용자 단말(210)에 정보 및/또는 데이터를 입력하거나 사용자 단말(210)로부터 생성된 정보 및/또는 데이터를 출력하도록 구성될 수 있다.
메모리(312, 332)는 비-일시적인 임의의 컴퓨터 판독 가능한 기록매체를 포함할 수 있다. 일 실시예에 따르면, 메모리(312, 332)는 RAM(random access memory), ROM(read only memory), 디스크 드라이브, SSD(solid state drive), 플래시 메모리(flash memory) 등과 같은 비소멸성 대용량 저장 장치(permanent mass storage device)를 포함할 수 있다. 다른 예로서, ROM, SSD, 플래시 메모리, 디스크 드라이브 등과 같은 비소멸성 대용량 저장 장치는 메모리와는 구분되는 별도의 영구 저장 장치로서 사용자 단말(210) 또는 정보 처리 시스템(230)에 포함될 수 있다. 또한, 메모리(312, 332)에는 운영체제와 적어도 하나의 프로그램 코드(예를 들어, 사용자 단말(210)에 설치되어 화자 분할 애플리케이션 등을 위한 코드)가 저장될 수 있다.
이러한 소프트웨어 구성요소들은 메모리(312, 332)와는 별도의 컴퓨터에서 판독가능한 기록매체로부터 로딩될 수 있다. 이러한 별도의 컴퓨터에서 판독가능한 기록매체는 이러한 사용자 단말(210) 및 정보 처리 시스템(230)에 직접 연결가능한 기록 매체를 포함할 수 있는데, 예를 들어, 플로피 드라이브, 디스크, 테이프, DVD/CD-ROM 드라이브, 메모리 카드 등의 컴퓨터에서 판독 가능한 기록매체를 포함할 수 있다. 다른 예로서, 소프트웨어 구성요소들은 컴퓨터에서 판독 가능한 기록매체가 아닌 통신 모듈을 통해 메모리(312, 332)에 로딩될 수도 있다. 예를 들어, 적어도 하나의 프로그램은 개발자들 또는 애플리케이션의 설치 파일을 배포하는 파일 배포 시스템이 네트워크(220)를 통해 제공하는 파일들에 의해 설치되는 컴퓨터 프로그램에 기반하여 메모리(312, 332)에 로딩될 수 있다.
프로세서(314, 334)는 기본적인 산술, 로직 및 입출력 연산을 수행함으로써, 컴퓨터 프로그램의 명령을 처리하도록 구성될 수 있다. 명령은 메모리(312, 332) 또는 통신 모듈(316, 336)에 의해 프로세서(314, 334)로 제공될 수 있다. 예를 들어, 프로세서(314, 334)는 메모리(312, 332)와 같은 기록 장치에 저장된 프로그램 코드에 따라 수신되는 명령을 실행하도록 구성될 수 있다.
통신 모듈(316, 336)은 네트워크(220)를 통해 사용자 단말(210)과 정보 처리 시스템(230)이 서로 통신하기 위한 구성 또는 기능을 제공할 수 있으며, 사용자 단말(210) 및/또는 정보 처리 시스템(230)이 다른 사용자 단말 또는 다른 시스템(일례로 별도의 클라우드 시스템 등)과 통신하기 위한 구성 또는 기능을 제공할 수 있다. 일례로, 사용자 단말(210)의 프로세서(314)가 메모리(312) 등과 같은 기록 장치에 저장된 프로그램 코드에 따라 생성한 요청 또는 데이터(예를 들어, 화자 분할 요청, 음성 인식 요청, 회의록 생성 요청, 자막 생성 요청 등)는 통신 모듈(316)의 제어에 따라 네트워크(220)를 통해 정보 처리 시스템(230)으로 전달될 수 있다. 역으로, 정보 처리 시스템(230)의 프로세서(334)의 제어에 따라 제공되는 제어 신호나 명령이 통신 모듈(336)과 네트워크(220)를 거쳐 사용자 단말(210)의 통신 모듈(316)을 통해 사용자 단말(210)에 수신될 수 있다. 예를 들어, 사용자 단말(210)은 정보 처리 시스템(230)으로부터 통신 모듈(316)을 통해 입력 신호로부터 분리된 각 화자의 음성 및/또는 각 화자의 음성에 추가적인 처리를 수행한 결과(예를 들어, 회의록, 자막 등)를 화자 분할의 결과로서 수신할 수 있다.
입출력 인터페이스(318)는 입출력 장치(320)와의 인터페이스를 위한 수단일 수 있다. 일 예로서, 입력 장치는 오디오 센서 및/또는 이미지 센서를 포함한 카메라, 키보드, 마이크로폰, 마우스 등의 장치를, 그리고 출력 장치는 디스플레이, 스피커, 햅틱 피드백 디바이스(haptic feedback device) 등과 같은 장치를 포함할 수 있다. 다른 예로, 입출력 인터페이스(318)는 터치스크린 등과 같이 입력과 출력을 수행하기 위한 구성 또는 기능이 하나로 통합된 장치와의 인터페이스를 위한 수단일 수 있다. 예를 들어, 사용자 단말(210)의 프로세서(314)가 메모리(312)에 로딩된 컴퓨터 프로그램의 명령을 처리함에 있어서 정보 처리 시스템(230)이나 다른 사용자 단말이 제공하는 정보 및/또는 데이터를 이용하여 구성되는 서비스 화면 등이 입출력 인터페이스(318)를 통해 디스플레이에 표시될 수 있다. 도 3에서는 입출력 장치(320)가 사용자 단말(210)에 포함되지 않도록 도시되어 있으나, 이에 한정되지 않으며, 사용자 단말(210)과 하나의 장치로 구성될 수 있다. 또한, 정보 처리 시스템(230)의 입출력 인터페이스(338)는 정보 처리 시스템(230)과 연결되거나 정보 처리 시스템(230)이 포함할 수 있는 입력 또는 출력을 위한 장치(미도시)와의 인터페이스를 위한 수단일 수 있다. 도 3에서는 입출력 인터페이스(318, 338)가 프로세서(314, 334)와 별도로 구성된 요소로서 도시되었으나, 이에 한정되지 않으며, 입출력 인터페이스(318, 338)가 프로세서(314, 334)에 포함되도록 구성될 수 있다.
사용자 단말(210) 및 정보 처리 시스템(230)은 도 3의 구성요소들보다 더 많은 구성요소들을 포함할 수 있다. 그러나, 대부분의 종래기술적 구성요소들을 명확하게 도시할 필요성은 없다. 일 실시예에 따르면, 사용자 단말(210)은 상술된 입출력 장치(320) 중 적어도 일부를 포함하도록 구현될 수 있다. 또한, 사용자 단말(210)은 트랜시버(transceiver), GPS(Global Positioning system) 모듈, 카메라, 각종 센서, 데이터베이스 등과 같은 다른 구성요소들을 더 포함할 수 있다. 예를 들어, 사용자 단말(210)이 스마트폰인 경우, 일반적으로 스마트폰이 포함하고 있는 구성요소를 포함할 수 있으며, 예를 들어, 가속도 센서, 자이로 센서, 카메라 모듈, 각종 물리적인 버튼, 터치패널을 이용한 버튼, 입출력 포트, 진동을 위한 진동기 등의 다양한 구성요소들이 사용자 단말(210)에 더 포함되도록 구현될 수 있다. 일 실시예에 따르면, 사용자 단말(210)의 프로세서(314)는 화자 분할 서비스를 제공하는 애플리케이션 등이 동작하도록 구성될 수 있다. 이 때, 해당 애플리케이션 및/또는 프로그램과 연관된 코드가 사용자 단말(210)의 메모리(312)에 로딩될 수 있다.
화자 분할 애플리케이션 등을 위한 프로그램이 동작되는 동안에, 프로세서(314)는 입출력 인터페이스(318)와 연결된 터치 스크린, 키보드, 오디오 센서 및/또는 이미지 센서를 포함한 카메라, 마이크로폰 등의 입력 장치를 통해 입력되거나 선택된 텍스트, 이미지, 영상, 음성 및/또는 동작 등을 수신할 수 있으며, 수신된 텍스트, 이미지, 영상, 음성 및/또는 동작 등을 메모리(312)에 저장하거나 통신 모듈(316) 및 네트워크(220)를 통해 정보 처리 시스템(230)에 제공할 수 있다. 예를 들어, 프로세서(314)는 화자 분할을 요청하는 사용자 입력을 수신하여, 통신 모듈(316) 및 네트워크(220)를 통해 정보 처리 시스템(230)에 제공할 수 있다. 다른 예로서, 프로세서(314)는 여러 화자의 발화가 포함된 입력 신호 등에 대한 사용자의 선택을 나타내는 입력을 수신하여. 통신 모듈(316) 및 네트워크(220)를 통해 정보 처리 시스템(230)에 제공할 수 있다.
사용자 단말(210)의 프로세서(314)는 입력 장치(320), 다른 사용자 단말, 정보 처리 시스템(230) 및/또는 복수의 외부 시스템으로부터 수신된 정보 및/또는 데이터를 관리, 처리 및/또는 저장하도록 구성될 수 있다. 프로세서(314)에 의해 처리된 정보 및/또는 데이터는 통신 모듈(316) 및 네트워크(220)를 통해 정보 처리 시스템(230)에 제공될 수 있다. 사용자 단말(210)의 프로세서(314)는 입출력 인터페이스(318)를 통해 입출력 장치(320)로 정보 및/또는 데이터를 전송하여, 출력할 수 있다. 예를 들면, 프로세서(314)는 수신한 정보 및/또는 데이터를 사용자 단말의 화면에 디스플레이할 수 있다.
정보 처리 시스템(230)의 프로세서(334)는 복수의 사용자 단말(210) 및/또는 복수의 외부 시스템으로부터 수신된 정보 및/또는 데이터를 관리, 처리 및/또는 저장하도록 구성될 수 있다. 프로세서(334)에 의해 처리된 정보 및/또는 데이터는 통신 모듈(336) 및 네트워크(220)를 통해 사용자 단말(210)에 제공할 수 있다. 일 실시예에 따르면, 정보 처리 시스템(230)은 사용자 단말(210)로부터 수신된 입력 신호에 포함된 음성 구간에 기초하여, 복수의 화자 특징 벡터를 추출하고, 추출된 복수의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 중심 특징 벡터를 추출할 수 있다. 그런 다음, 정보 처리 시스템(230)은 입력 신호로부터 중심 특징 벡터와 연관된 특정 화자의 음성을 추출할 수 있다. 정보 처리 시스템(230)은 상술한 과정을 여러 번 반복하여 수행함으로써 입력 신호에 포함된 각 화자의 음성을 분리할 수 있다. 추가적으로, 정보 처리 시스템(230)은 분리된 각 화자의 음성 및/또는 각 화자의 음성에 추가적인 처리를 수행한 결과를 화자 분할의 결과로서 사용자 단말(210)로 제공할 수 있다. 예를 들어, 정보 처리 시스템은 분리된 각 화자의 음성에 기초하여, 자동 음성 인식(ASR) 또는 STT 등을 수행함으로써, 회의록 또는 자막 등을 생성하여 사용자 단말(210)로 전송할 수 있다.
정보 처리 시스템(230)의 프로세서(334)는 사용자 단말(210)의 디스플레이 출력 가능 장치(예: 터치 스크린, 디스플레이 등), 음성 출력 가능 장치(예: 스피커) 등의 출력 장치(320)를 통해 처리된 정보 및/또는 데이터를 출력하도록 구성될 수 있다. 예를 들어, 정보 처리 시스템(230)의 프로세서(334)는 분리된 각 화자의 음성을 화자 분할의 결과로서 통신 모듈(336) 및 네트워크(220)를 통해 사용자 단말(210)로 제공하고, 각 화자의 음성을 사용자 단말(210)의 사운드 출력 가능 장치 등을 통해 출력하도록 구성될 수 있다. 다른 예로서, 정보 처리 시스템(230)의 프로세서(334)는 통신 모듈(336) 및 네트워크(220)를 통해 사용자 단말(210)로 분리된 각 화자의 음성에 기초하여, 자동 음성 인식(ASR) 또는 STT 등을 수행함으로써 생성된 회의록 또는 자막 등을 제공하고, 사용자 단말(210)의 디스플레이 출력 가능 장치 등을 통해 출력하도록 구성될 수 있다.
도 4는 본 개시의 일 실시예에 따라 입력 신호에 포함된 음성 구간에 기초하여 복수의 화자 특징 벡터를 추출하는 예시를 나타내는 도면이다. 일 실시예에 따르면, 정보 처리 시스템 또는 정보 처리 시스템의 프로세서는 입력 신호(410)를 수신할 수 있다. 여기서, 입력 신호(410)는 여러 명의 화자의 발화 음성(및 주변 소음)을 포함할 수 있다. 또한, 여기서, 화자는 인간 뿐만 아니라, 합성된 음성 발화가 가능한 가상의 인물 또는 캐릭터, 음성을 포함하는 사운드 생성 및 출력이 가능한 소프트웨어 또는 하드웨어 모듈 등을 포함하는 개념일 수 있다. 예를 들어, 입력 신호(410)는 다수가 참가한 회의 도중 녹음된 음성을 포함할 수 있다. 다른 예로, 입력 신호(410)는 드라마 또는 영화 등의 영상으로부터 추출된 음성을 포함할 수 있다.
입력 신호(410)는 다양한 형태의 사운드일 수 있다. 예를 들어, 입력 신호(410)는 웨이브 폼(waveform) 형태의 사운드이거나, 스펙트로그램(spectrogram) 형태의 사운드일 수 있다. 일 실시예에 따르면, 정보 처리 시스템은 웨이브 폼 형태의 사운드를 수신하고, 수신된 웨이브 폼 형태의 사운드를 기초로 주파수 변환을 수행하여 스펙트로그램 형태로 변환된 사운드를 입력 신호(410)로서 사용할 수 있다.
그런 다음, 정보 처리 시스템은 입력 신호(410)로부터 음성 구간을 검출할 수 있다(420). 여기서, 음성 구간은 입력 신호(110) 중 화자의 발화 음성이 포함된 구간일 수 있다. 예를 들어, 정보 처리 시스템은 음성 끝점 검출(End Point Detection; EPD)을 수행하여 발화의 시작점과 끝점을 검출함으로써, 입력 신호(410)에 포함된 음성 구간을 검출할 수 있다.
음성 구간이 검출된 경우, 정보 처리 시스템은 검출된 음성 구간에 기초하여 복수의 화자 특징 벡터(432)를 추출할 수 있다(430). 화자 특징 벡터는 음성 구간에 포함된 발화 음성의 특징에 대한 정보를 포함할 수 있다. 즉, 화자 특징 벡터는 음성을 발화한 화자에 대한 정보를 포함할 수 있다.
일 실시예에 따르면, 정보 처리 시스템은 각 단위 음성 구간 당 하나의 화자 특징 벡터를 추출함으로써, 검출된 음성 구간에 기초하여 복수의 화자 특징 벡터(432)를 추출할 수 있다. 예를 들어, 단위 음성 구간의 길이가 0.5초인 실시예를 가정하면, 검출된 제1 음성 구간(422)의 길이가 1초인 경우, 정보 처리 시스템은 제1 음성 구간(422)에 기초하여 2개의 화자 특징 벡터(도시된 예에서, E1 및 E2)를 추출할 수 있다. 또한, 검출된 제2 음성 구간(424)의 길이가 1.5초인 경우, 정보 처리 시스템은 제2 음성 구간(424)에 기초하여 3개의 화자 특징 벡터(도시된 예에서, E3, E4 및 E5)를 추출할 수 있다.
정보 처리 시스템은 다양한 길이의 단위 음성 구간에 따른 음성 구간 검출 및 화자 특징 벡터 추출 방법을 이용하여 검출된 음성 구간으로부터 복수의 화자 특징 벡터(432)를 추출할 수 있다.
일 실시예에 따르면, 정보 처리 시스템은 화자 특징 추출 모델(Speaker Embedding Extractor)을 이용하여 음성 구간에 기초하여 복수의 화자 특징 벡터(432)를 추출할 수 있다. 여기서, 화자 특징 추출 모델은 음성 구간에 기초하여 화자 특징 벡터를 추출하도록 학습된 기계학습 모델일 수 있다. 예를 들어, 화자 특징 추출 모델은 동일한 화자의 발화 음성을 포함하는 음성 구간들로부터는 동일/유사한 화자 특징 벡터를 추출하도록 학습되고, 상이한 화자의 발화 음성을 포함하는 음성 구간들로부터는 각 화자의 특징 차이가 잘 구분될 수 있는 화자 특징 벡터를 추출하도록 학습된 기계학습 모델일 수 있다.
다른 실시예에 따르면, 정보 처리 시스템은 입력 신호(410)로부터 검출된 각 음성 구간(예를 들어, 단위 음성 구간)을 가공함으로써 복수의 화자 특징 벡터(432)를 추출할 수 있다. 예를 들어, 정보 처리 시스템은 스펙트로그램 형태의 각 음성 구간의 평균 값을 산출함으로써 복수의 화자 특징 벡터(432)를 추출할 수 있다.
도 5는 본 개시의 일 실시예에 따라 화자 특징 벡터에 대해 클러스터링을 수행하는 예시를 나타내는 도면이다. 일 실시예에 따르면, 정보 처리 시스템은 입력 신호에 포함된 음성 구간으로부터 추출된 복수의 화자 특징 벡터에 대해 클러스터링을 수행할 수 있다. 복수의 화자 특징 벡터에 대해 클러스터링이 수행되는 경우, 복수의 화자 특징 벡터 중 유사한 특징 벡터들이 그룹핑될 수 있다. 예를 들어, 정보 처리 시스템은 K-means Clustering, Spectral Clustering 등 임의의 클러스터링 방법을 이용하여 클러스터링을 수행할 수 있다. 일 실시예에 따르면, 정보 처리 시스템은 추출된 복수의 화자 특징 벡터로부터 클러스터의 개수를 추정하고, 복수의 화자 특징 벡터를 추정된 개수의 클러스터로 그룹핑할 수 있다.
종래의 화자 분할 방법에 따르면, 입력 신호에 포함된 각 화자의 음성을 구분하기 위해 클러스터링 과정을 필수적으로 포함하지만, 본 개시의 화자 분할 방법에 따르면, 클러스터링 과정은 선택적으로 수행될 수 있다.
도 5에는 세 명의 화자의 발화 음성이 포함된 두 입력 신호에 기초하여 추출된 두 세트의 화자 특징 벡터에 대해 클러스터링을 수행한 결과가 도시되어 있다. 제1 클러스터링 결과(520)는 세 명의 화자의 발화 음성이 포함된 제1 입력 신호에 포함된 음성 구간으로부터 추출된 제1 세트의 화자 특징 벡터(510)에 대해 클러스터링을 수행한 결과이다. 제1 클러스터링 결과(520)를 살펴보면, 제1 세트의 화자 특징 벡터(510)가 세 개의 클러스터(522, 524, 526)로 그룹핑된 것을 확인할 수 있다. 동일한 클러스터에 속한 화자 특징 벡터들은 서로 동일한 화자의 발화 음성을 포함하는 음성 구간에서 추출된 것으로 추정될 수 있다.
제2 클러스터링 결과(540)는 세 명의 화자의 발화 음성이 포함된 제2 입력 신호에 포함된 음성 구간으로부터 추출된 제2 세트의 화자 특징 벡터(530)에 대해 클러스터링을 수행한 결과이다. 제2 입력 신호는 제1 화자 및 제2 화자의 발화 음성이 중첩된 중첩 음성 구간(532)을 포함할 수 있다. 이 경우, 중첩 음성 구간(532)에 기초하여 추출된 화자 특징 벡터(534)는 오염되어, 세 명의 화자 중 어느 한 화자의 특징을 잘 반영하지 않을 수 있다. 예를 들어, 중첩 음성 구간(532)에 기초하여 추출된 화자 특징 벡터(534)는 제1 화자의 발화 음성만 포함하는 음성 구간으로부터 추출된 화자 특징 벡터들이 위치한 영역과 제2 화자의 발화 음성만 포함하는 음성 구간으로부터 추출된 화자 특징 벡터들이 위치한 영역 사이에 위치할 수 있다. 결국, 중첩 음성 구간(532)에 기초하여 추출된 화자 특징 벡터(534)에 의해, 제1 화자에 대응되는 화자 특징 벡터의 영역과 제2 화자에 대응되는 화자 특징 벡터의 영역의 구분이 불분명해질 수 있다.
제2 입력 신호에 기초하여 추출된 제2 세트의 화자 특징 벡터(530)에 대해 클러스터링을 수행한 결과인 제2 클러스터링 결과(540)를 살펴보면, 제2 입력 신호는 실제로 세 명의 화자의 발화 음성을 포함하고 있음에도 불구하고, 화자의 수가 잘못 추정되어, 제2 세트의 화자 특징 벡터(530)가 두 개의 클러스터(542, 544)로 그룹핑된 것을 확인할 수 있다.
상술한 바와 같이, 입력 신호에 중첩 음성 구간(532)이 존재하는 경우, 중첩 음성 구간(532)에 의해 화자 특징 벡터에 오염이 발생하여, 화자 분할의 결과에 오류가 발생할 수 있다. 또한, 입력 신호에 중첩 음성 구간(532)이 존재하는 경우 뿐만 아니라, 입력 신호에 주변 소음과 같은 잡음이 포함된 경우에도 화자 특징 벡터에 오염이 발생될 수 있다.
본 개시의 일 실시예에 따르면, 추출된 복수의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 중심 특징 벡터를 추출함으로써 특정 화자의 대표성 있는 화자 특징 벡터를 추출할 수 있다. 따라서, 클러스터링 결과에 오류가 있더라도, 혹은 클러스터링을 수행하지 않고도, 오류 없이 화자 분할을 수행할 수 있다. 추출된 복수의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 중심 특징 벡터를 추출하는 과정에 관하여서는 도 6을 참조하여 후술된다.
도 6은 본 개시의 일 실시예에 따라 추출된 복수의 화자 특징 벡터에 기초하여 중심 특징 벡터를 추출하는 예시를 나타내는 도면이다. 도 6에는 세 명의 화자(화자 A, 화자 B 및 화자 C)의 발화 음성을 포함하는 입력 신호에 기초하여 추출된 제1 세트의 화자 특징 벡터(600)의 예시가 도시되어 있다. 제1 세트의 화자 특징 벡터(600)는 화자 A의 발화 음성만 포함된 음성 구간에 기초하여 추출된 복수의 화자 특징 벡터를 포함하는 제1 그룹(610), 화자 B의 발화 음성만 포함된 음성 구간에 기초하여 추출된 복수의 화자 특징 벡터를 포함하는 제2 그룹(620), 화자 A 및 화자 B의 발화가 중첩된 음성 구간에 기초하여 추출된 복수의 화자 특징 벡터를 포함하는 제3 그룹(630), 및 화자 C의 발화 음성만 포함된 음성 구간에 기초하여 추출된 제4 그룹(640)을 포함할 수 있다. 이와 같이, 입력 신호에 두 명 이상의 화자의 발화 음성이 중첩된 구간이 포함된 경우, 화자 특징 벡터들에 대해 클러스터링을 수행하면, 제1 그룹(610), 제2 그룹(620) 및 제3 그룹(630)이 모두 동일한 클러스터로 분류될 수 있다. 따라서, 종래의 방법에 따라, 클러스터링 결과만을 기초로 화자 분할을 수행하는 경우, 화자 분할의 결과에 오류가 발생할 수 있다.
본 개시의 일 실시예에 따르면, 정보 처리 시스템은 추출된 제1 세트의 화자 특징 벡터(600)에 기초하여, 가장 높은 신뢰도를 갖는 중심 특징 벡터를 추출할 수 있다. 일 실시예에 따르면, 정보 처리 시스템은 화자 특징 벡터의 밀집도를 신뢰도의 척도로서 사용할 수 있다.
예를 들어, 정보 처리 시스템은 먼저 벡터 공간 내에서 화자 특징 벡터가 가장 밀집된 영역을 중심 영역으로 결정할 수 있다. 구체적 예로, 정보 처리 시스템은 벡터 공간 내에서 추출된 제1 세트의 화자 특징 벡터(600)의 각각을 중심으로 한 동일한 크기의 복수의 영역 중 가장 많은 수의 화자 특징 벡터를 포함하는 영역을 중심 영역으로 결정할 수 있다. 도시된 예에서, 제1 특징 벡터(650)를 중심으로 한 제1 영역(652)에는 제1 특징 벡터(650)를 제외한 6개의 화자 특징 벡터가 존재한다. 또한, 제2 특징 벡터(660)를 중심으로 한 제2 영역(662)(제1 영역(652)과 동일한 모양 및 크기의 영역)에는 제2 특징 벡터(660)를 제외한 5개의 화자 특징 벡터가 존재한다. 따라서, 이 경우 제1 영역(652)이 중심 영역으로 결정될 수 있다.
그런 다음, 정보 처리 시스템은 중심 영역에 포함된 하나 이상의 화자 특징 벡터에 기초하여, 중심 특징 벡터를 결정할 수 있다. 예를 들어, 정보 처리 시스템은 중심 영역에 포함된 하나 이상의 화자 특징 벡터의 평균을 중심 특징 벡터로 결정할 수 있다. 이 실시예에 따르면, 정보 처리 시스템은 중심 영역으로 결정된 제1 영역(652)에 포함된 7개의 화자 특징 벡터의
평균을 중심 특징 벡터로 결정할 수 있다. 다른 예로, 정보 처리 시스템은 중심 영역의 중심 또는 중심에 가장 근접하여 위치한 화자 특징 벡터를 중심 특징 벡터로 추출할 수 있다. 이 실시예에 따르면, 정보 처리 시스템은 중심 영역으로 결정된 제1 영역(652)의 중심에 위치하는 제1 특징 벡터(650)를 중심 특징 벡터로 추출할 수 있다. 중심 특징 벡터로 추출된 제1 특징 벡터(650)는 화자 A의 발화 음성만 포함된 음성 구간에 기초하여 추출된 복수의 화자 특징 벡터를 포함하는 제1 그룹(610) 중에서도 가장 밀집된 영역에 기초하여 추출되어, 화자 A의 음성을 가장 잘 대표할 것으로 추정될 수 있다.
일 실시예에서, 중심 특징 벡터를 추출하기 전에 제1 세트의 화자 특징 벡터(600)에 대해 클러스터링이 수행된 경우, 정보 처리 시스템은 클러스터링 결과, 가장 많은 수의 화자 특징 벡터를 포함하는 클러스터로부터 중심 특징 벡터를 추출할 수 있다. 예를 들어, 제1 세트의 화자 특징 벡터(600)에 대해 클러스터링을 수행한 결과, 제1 그룹(610), 제2 그룹(620) 및 제3 그룹(630)을 포함하는 제1 클러스터, 제4 그룹(640)을 포함하는 제2 클러스터로 분류될 수 있다. 이 경우, 가장 많은 수의 화자 특징 벡터를 포함하는 제1 클러스터에 대응되는 공간 내에서 중심 영역을 결정하고, 중심 영역에 포함된 하나 이상의 화자 특징 벡터에 기초하여, 중심 특징 벡터를 추출할 수 있다.
상술한 바와 같이, 입력 신호로부터 추출된 복수의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 중심 특징 벡터를 추출함으로써 특정 화자의 대표성 있는 화자 특징 벡터를 추출할 수 있어, 클러스터링 결과에 오류가 있더라도, 혹은 클러스터링을 수행하지 않고도, 높은 정확도를 갖고 화자 분할을 수행할 수 있다.
도 7은 본 개시의 일 실시예에 따라 입력 신호(710)로부터 특정 화자의 음성(720)을 추출하는 예시를 나타내는 도면이다. 정보 처리 시스템은 입력 신호(710)로부터 중심 특징 벡터와 연관된 특정 화자의 음성(720)을 추출할 수 있다. 예를 들어, 정보 처리 시스템은 목적 음성 추출 모델(700)을 이용하여, 입력 신호로부터 제1 중심 특징 벡터와 연관된 제1 화자의 음성을 추출할 수 있다. 특정 화자의 음성(720)을 추출하기 위해 이용되는 목적 음성 추출 모델(700)의 구체적인 예시 및 학습 방법에 관하여서는 도 8 및 도 9를 참조하여 상세히 후술된다.
제1 화자의 음성을 추출한 뒤, 정보 처리 시스템은 입력 신호(710)로부터 추출된 제1 화자의 음성을 제외한 잔여 신호에 음성 구간이 남아 있는지 여부를 판정할 수 있다. 잔여 신호에 음성 구간이 남아 있다고 판정되는 경우, 도 4 내지 도 7을 참조하여 상술한 일련의 과정 중 적어도 일부 과정을 반복하여 수행할 수 있다.
예를 들어, 정보 처리 시스템은 입력 신호(710)에 포함된 음성 구간 중, 추출된 특정 화자의 음성(720)이 포함된 음성 구간을 제외한 잔여 음성 구간에 기초하여, 복수의 화자 특징 벡터를 추출하고, 추출된 복수의 화자 특징 벡터에 기초하여 가장 높은 신뢰도를 갖는 제2 중심 특징 벡터를 추출하고, 입력 신호(710) 또는 잔여 신호로부터 추출된 제2 중심 특징 벡터와 연관된 제2 화자의 음성을 추출할 수 있다.
다른 예로, 정보 처리 시스템은 입력 신호(710)에 기초하여 추출된 제1 세트의 화자 특징 벡터 중 특정 화자의 음성(720)이 포함된 음성 구간을 제외한 잔여 음성 구간에 기초하여 추출된 제1 부분 세트의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제2 중심 특징 벡터를 추출하고, 입력 신호(710) 또는 잔여 신호로부터 추출된 제2 중심 특징 벡터와 연관된 제2 화자의 음성을 추출할 수 있다.
또 다른 예로, 정보 처리 시스템은 잔여 신호에 포함된 음성 구간에 기초하여, 복수의 화자 특징 벡터를 추출하고, 추출된 복수의 화자 특징 벡터에 기초하여 가장 높은 신뢰도를 갖는 제2 중심 특징 벡터를 추출하고, 입력 신호(710) 또는 잔여 신호로부터 추출된 제2 중심 특징 벡터와 연관된 제2 화자의 음성을 추출할 수 있다.
상술한 일련의 과정을 잔여 신호에 음성 구간이 남아 있지 않을 때까지 반복하여 수행하는 경우, 입력 신호(710)에 발화 음성이 포함된 모든 화자의 음성을 분리할 수 있다.
추가적으로, 정보 처리 시스템은 분리된 각 화자의 음성 및/또는 각 화자의 음성에 추가적인 처리를 수행한 결과를 화자 분할의 결과로서 사용자 단말로 전송할 수 있다. 예를 들어, 정보 처리 시스템은 분리된 각 화자의 음성에 기초하여, 자동 음성 인식(Automatic Speech Recognition; ASR) 또는 STT(Speech to Text) 등을 수행함으로써, 회의록 또는 자막 등을 생성하여 사용자 단말로 전송할 수 있다.
도 8은 본 개시의 일 실시예에 따른 목적 음성 추출 모델(800)의 구체적인 예시를 나타내는 도면이다. 일 실시예에 따르면, 목적 음성 추출 모델(800)은 혼합 음성(810) 및 목적 화자 특징 벡터(820)를 입력으로 수신하여, 목적 화자 특징 벡터(820)와 연관된 마스크(830)를 출력하는 모델일 수 있다.
여기서, 혼합 음성(810)은 목적 화자의 음성과 다른 화자의 음성 및/또는 잡음이 혼합된 음성일 수 있다. 예를 들어, 혼합 음성(810)은 웨이브 폼(waveform) 형태의 음성이거나, 스펙트로그램(spectrogram) 형태의 음성일 수 있다. 일 실시예에 따르면, 웨이브 폼 형태의 음성에 대해 주파수 변환을 수행하여 스펙트로그램 형태로 변환함으로써, 스펙트로그램 형태의 음성을 혼합 음성(810)으로서 사용할 수 있다.
목적 화자 특징 벡터(820)는 목적 화자에 대한 정보를 포함하는 벡터일 수 있다. 또한, 목적 화자 특징 벡터(820)와 연관된 마스크(830)는 입력 신호로부터 목적 화자를 제외한 다른 화자의 음성 및 잡음을 제거하기 위한 마스크일 수 있다. 출력된 마스크(830) 및 혼합 음성(810)을 이용하면, 목적 화자의 음성(840)을 추출할 수 있다. 예를 들어, 혼합 음성(810)에 마스크(830)를 곱함으로써, 목적 화자의 음성(840)을 추출할 수 있다.
일 실시예에서, 혼합 음성(810)으로서 스펙트로그램 형태의 음성을 사용한 경우, 추출된 목적 화자의 음성(840) 역시 스펙트로그램 형태일 수 있다. 이 경우, 스펙트로그램 형태의 목적 화자의 음성에 주파수 변환의 역변환을 수행함으로써 웨이브폼 형태의 음성으로 변환할 수 있다.
일 실시예에 따르면, 정보 처리 시스템은 입력 신호를 혼합 음성(810)으로서, 추출된 중심 특징 벡터를 목적 화자 특징 벡터(820)로서 목적 음성 추출 모델(800)에 입력하여, 중심 특징 벡터와 연관된 마스크(830)를 추정할 수 있으며, 입력 신호에 마스크(830)를 곱함으로써 특정 화자의 음성을 추출할 수 있다. 추가적으로 또는 대안적으로, 정보 처리 시스템은 목적 화자 특징 벡터(820)로서 상술한 중심 특징 벡터 추출 방법과 유사하거나 상이한 방법에 의해 추출된 다른 화자 특징 벡터를 이용할 수도 있다.
도 8에 도시된 목적 음성 추출 모델(800)은 일 예시일 뿐이며, 다른 실시예에서 목적 음성 추출 모델(800)은 도시된 모델과 다르게 구현될 수 있다. 예를 들어, 목적 음성 추출 모델(800)은 목적 화자 특징 벡터(820)와 연관된 마스크(830)를 출력하는 대신, 곧바로 목적 화자의 음성(840)을 출력하도록 구현될 수도 있다.
도 9는 본 개시의 일 실시예에 따른 목적 음성 추출 모델(800)을 학습시키는 방법의 예시를 나타내는 도면이다. 일 실시예에 따르면, 목적 음성 추출 모델(800)은 혼합 음성(810) 및 목적 화자 특징 벡터(820)를 입력으로 수신하여, 목적 화자 특징 벡터(820)와 연관된 마스크(830)를 출력하도록 학습된 기계학습 모델일 수 있다. 예를 들어, 목적 음성 추출 모델(800)은 CNN(convolutional neural network), LSTM(long short term memory), FC(fully connected) 레이어 중 적어도 하나를 포함하는 인공신경망 모델일 수 있다.
일 실시예에 따르면, 정보 처리 시스템은 목적 음성 추출 모델(800)을 학습시키기 위한 학습 데이터로서 제1 목적 화자 참조 음성(822), 제2 목적 화자 참조 음성(812), 타 화자 참조 음성(814)을 획득할 수 있다.
제1 목적 화자 참조 음성(822) 및 제2 목적 화자 참조 음성(812)은 목적 화자의 발화 음성만을 포함하는 음성일 수 있다. 제1 목적 화자 참조 음성(822) 및 제2 목적 화자 참조 음성(812)은 서로 동일하거나 상이할 수 있다. 또한, 타 화자 참조 음성(814)은 목적 화자가 아닌 다른 화자의 발화 및/또는 잡음을 포함하는 음성일 수 있다. 정보 처리 시스템은 제2 목적 화자 참조 음성(812)과 타 화자 참조 음성(814)을 혼합함으로써, 혼합 음성(810)을 생성할 수 있다. 상술한 각 음성(810, 812, 814, 822)은 웨이브 폼 형태의 음성이거나, 스펙트로그램 형태의 음성일 수 있다. 일 실시예에 따르면, 웨이브 폼 형태의 음성에 대해 주파수 변환을 수행하여 스펙트로그램 형태의 음성으로 변환하여 사용할 수 있다.
먼저, 정보 처리 시스템은 제1 목적 화자 참조 음성(822)에 기초하여, 목적 화자 특징 벡터(820)를 추출할 수 있다. 제1 목적 화자 참조 음성(822)에 기초하여, 목적 화자 특징 벡터(820)를 추출하는 과정은 도 4를 참조하여 상술한 방법과 동일하거나 유사하게 수행될 수 있다.
그런 다음, 정보 처리 시스템은 혼합 음성(810) 및 목적 화자 특징 벡터(820)를 목적 음성 추출 모델(800)에 입력하여, 목적 화자 특징 벡터(820)와 연관된 마스크(830)를 추정할 수 있다. 그런 다음, 혼합 음성(810)에 추정된 마스크(830)를 곱함으로써 목적 화자의 음성(840)을 추출할 수 있다.
정보 처리 시스템은 추출된 목적 화자의 음성(840)과 제2 목적 화자 참조 음성(812)의 비교에 기초하여, 예측 손실(Loss)(850)을 산출할 수 있다. 그런 다음, 산출된 예측 손실(850)을 이용하여 목적 음성 추출 모델(800)에 포함된 복수의 레이어들 또는 노드들 사이의 연결에 대응하는 가중치를 업데이트할 수 있다.
도 9 및 상술한 학습 방법은 일 예시일 뿐이며, 다른 실시예에서는 목적 음성 추출 모델(800)이 도시된 모델과 다르게 구현되거나 다른 방법에 의해 학습될 수 있다.
도 10은 본 개시의 일 실시예에 따른 화자 분할 방법(1000)의 예시를 나타내는 흐름도이다. 일 실시예에 따르면, 화자 분할 방법(1000)은 프로세서(예를 들어, 정보 처리 시스템 또는 사용자 단말의 적어도 하나의 프로세서)가 입력 신호에 포함된 음성 구간에 기초하여 제1 세트의 화자 특징 벡터를 추출함으로써 개시될 수 있다(S1010). 예를 들어, 프로세서는 입력 신호로부터 음성 구간을 검출하고, 검출된 음성 구간에 기초하여 제1 세트의 화자 특징 벡터를 추출할 수 있다. 여기서, 제1 세트의 화자 특징 벡터는 복수의 화자 특징 벡터를 포함할 수 있다. 일 실시예에 따르면, 프로세서는 입력 신호에 포함된 복수의 단위 음성 구간의 각각에 기초하여, 복수의 화자 특징 벡터를 추출할 수 있다.
그런 다음, 프로세서는 추출된 제1 세트의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제1 중심 특징 벡터를 추출할 수 있다(S1020).
예를 들어, 프로세서는 먼저, 추출된 제1 세트의 화자 특징 벡터에 기초하여, 벡터 공간 중 제1 중심 영역을 결정할 수 있다. 일 실시예에서, 프로세서는 벡터 공간 내에서 화자 특징 벡터가 가장 밀집된 영역을 제1 중심 영역으로 결정할 수 있다. 예를 들어, 프로세서는 벡터 공간 내에서 추출된 제1 세트의 화자 특징 벡터의 각각을 중심으로 한 동일한 크기의 복수의 영역 중 가장 많은 수의 화자 특징 벡터를 포함하는 영역을 제1 중심 영역으로 결정할 수 있다.
제1 중심 영역이 결정된 다음, 프로세서는 제1 중심 영역에 포함된 하나 이상의 화자 특징 벡터에 기초하여, 제1 중심 특징 벡터를 추출할 수 있다. 예를 들어, 프로세서는 제1 중심 영역에 포함된 하나 이상의 화자 특징 벡터의 평균을 제1 중심 특징 벡터로 결정할 수 있다. 다른 예로, 프로세서는 제1 중심 영역의 중심 또는 중심에 가장 근접하여 위치한 화자 특징 벡터를 제1 중심 특징 벡터로서 추출할 수 있다.
일 실시예에 따르면, 프로세서는 제1 중심 특징 벡터를 추출하기 전에, 추출된 제1 세트의 화자 특징 벡터에 대해 클러스터링을 수행할 수도 있다. 클러스터링이 수행된 경우, 프로세서는 클러스터링 결과 가장 많은 수의 화자 특징 벡터를 포함하는 클러스터로부터 제1 중심 특징 벡터를 추출할 수 있다. 예를 들어, 프로세서는 벡터 공간 중 클러스터링 결과 가장 많은 수의 화자 특징 벡터를 포함하는 클러스터에 대응되는 공간 내에서 제1 중심 영역을 결정하고, 제1 중심 영역에 포함된 하나 이상의 화자 특징 벡터에 기초하여, 제1 중심 특징 벡터를 추출할 수 있다.
그런 다음, 프로세서는 입력 신호로부터 제1 중심 특징 벡터와 연관된 제1 화자의 음성을 추출할 수 있다(S1030). 일 실시예에 따르면, 프로세서는 목적 음성 추출 모델을 이용하여, 입력 신호로부터 제1 중심 특징 벡터와 연관된 제1 화자의 음성을 추출할 수 있다. 예를 들어, 프로세서는 목적 음원 추출 모델을 이용하여, 입력 신호 및 제1 중심 특징 벡터를 기초로 제1 중심 특징 벡터와 연관된 마스크를 추정할 수 있으며, 입력 신호 및 추정된 마스크에 기초하여, 제1 화자의 음성을 추출할 수 있다.
추가적으로, 프로세서는 입력 신호로부터 추출된 제1 화자의 음성을 제외한 잔여 신호에 음성 구간이 남아 있는지 여부를 판정할 수 있다. 잔여 신호에 음성 구간이 남아 있다고 판정되는 경우, 상술한 일련의 과정을 반복하여 수행함으로써 입력 신호 또는 잔여 신호에 포함된 다른 화자의 음성을 추출할 수 있다.
예를 들어, 프로세서는 단계 S1010으로 돌아가 일련의 과정을 반복하여 수행할 수 있다. 구체적으로, 프로세서는 잔여 신호에 포함된 음성 구간에 기초하여, 제2 세트의 화자 특징 벡터를 추출할 수 있다. 그런 다음, 추출된 제2 세트의 화자 특징 벡터에 기초하여 가장 높은 신뢰도를 갖는 제2 중심 특징 벡터를 추출하고, 입력 신호 또는 잔여 신호로부터 제2 중심 특징 벡터와 연관된 제2 화자의 음성을 추출할 수 있다.
다른 예로, 프로세서는 단계 S1020로 돌아가 일련의 과정을 반복하여 수행할 수 있다. 구체적으로, 프로세서는 제1 세트의 화자 특징 벡터 중 잔여 신호에 포함된 음성 구간을 기초로 추출된 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제2 중심 특징 벡터를 추출하고, 입력 신호 또는 잔여 신호로부터 제2 중심 특징 벡터와 연관된 제2 화자의 음성을 추출할 수 있다.
또 다른 예로, 프로세서는 입력 신호에 포함된 음성 구간 중 제1 화자의 음성이 포함된 음성 구간을 제외한 잔여 음성 구간에 기초하여, 가장 높은 신뢰도를 갖는 제2 중심 특징 벡터를 추출하고, 입력 신호 또는 잔여 신호로부터 제2 중심 특징 벡터와 연관된 제2 화자의 음성을 추출할 수 있다.
도면에 포함된 흐름도 및 상술한 설명은 일 예시일 뿐이며 본 개시의 범위가 이에 한정되는 것은 아니다. 예를 들어, 다른 실시예에 따르면, 일부 단계가 추가/변경/삭제될 수 있으며, 각 단계의 순서가 변경될 수 있다.
상술한 방법은 컴퓨터에서 실행하기 위해 컴퓨터 판독 가능한 기록매체에 저장된 컴퓨터 프로그램으로 제공될 수 있다. 매체는 컴퓨터로 실행 가능한 프로그램을 계속 저장하거나, 실행 또는 다운로드를 위해 임시 저장하는 것일수도 있다. 또한, 매체는 단일 또는 수개 하드웨어가 결합된 형태의 다양한 기록 수단 또는 저장수단일 수 있는데, 어떤 컴퓨터 시스템에 직접 접속되는 매체에 한정되지 않고, 네트워크 상에 분산 존재하는 것일 수도 있다. 매체의 예시로는, 하드 디스크, 플로피 디스크 및 자기 테이프와 같은 자기 매체, CD-ROM 및 DVD 와 같은 광기록 매체, 플롭티컬 디스크(floptical disk)와 같은 자기-광 매체(magneto optical medium), 및 ROM, RAM, 플래시 메모리 등을 포함하여 프로그램 명령어가 저장되도록 구성된 것이 있을 수 있다. 또한, 다른 매체의 예시로, 애플리케이션을 유통하는 앱 스토어나 기타 다양한 소프트웨어를 공급 내지 유통하는 사이트, 서버 등에서 관리하는 기록매체 내지 저장매체도 들 수 있다.
본 개시의 방법, 동작 또는 기법들은 다양한 수단에 의해 구현될 수도 있다. 예를 들어, 이러한 기법들은 하드웨어, 펌웨어, 소프트웨어, 또는 이들의 조합으로 구현될 수도 있다. 본원의 개시와 연계하여 설명된 다양한 예시적인 논리적 블록들, 모듈들, 회로들, 및 알고리즘 단계들은 전자 하드웨어, 컴퓨터 소프트웨어, 또는 양자의 조합들로 구현될 수도 있음을 통상의 기술자들은 이해할 것이다. 하드웨어 및 소프트웨어의 이러한 상호 대체를 명확하게 설명하기 위해, 다양한 예시적인 구성요소들, 블록들, 모듈들, 회로들, 및 단계들이 그들의 기능적 관점에서 일반적으로 위에서 설명되었다. 그러한 기능이 하드웨어로서 구현되는지 또는 소프트웨어로서 구현되는 지의 여부는, 특정 애플리케이션 및 전체 시스템에 부과되는 설계 요구사항들에 따라 달라진다. 통상의 기술자들은 각각의 특정 애플리케이션을 위해 다양한 방식들로 설명된 기능을 구현할 수도 있으나, 그러한 구현들은 본 개시의 범위로부터 벗어나게 하는 것으로 해석되어서는 안된다.
하드웨어 구현에서, 기법들을 수행하는 데 이용되는 프로세싱 유닛들은, 하나 이상의 ASIC들, DSP들, 디지털 신호 프로세싱 디바이스들(digital signal processing devices; DSPD들), 프로그램가능 논리 디바이스들(programmable logic devices; PLD들), 필드 프로그램가능 게이트 어레이들(field programmable gate arrays; FPGA들), 프로세서들, 제어기들, 마이크로제어기들, 마이크로프로세서들, 전자 디바이스들, 본 개시에 설명된 기능들을 수행하도록 설계된 다른 전자 유닛들, 컴퓨터, 또는 이들의 조합 내에서 구현될 수도 있다.
따라서, 본 개시와 연계하여 설명된 다양한 예시적인 논리 블록들, 모듈들, 및 회로들은 범용 프로세서, DSP, ASIC, FPGA나 다른 프로그램 가능 논리 디바이스, 이산 게이트나 트랜지스터 로직, 이산 하드웨어 컴포넌트들, 또는 본원에 설명된 기능들을 수행하도록 설계된 것들의 임의의 조합으로 구현되거나 수행될 수도 있다. 범용 프로세서는 마이크로프로세서일 수도 있지만, 대안으로, 프로세서는 임의의 종래의 프로세서, 제어기, 마이크로제어기, 또는 상태 머신일 수도 있다. 프로세서는 또한, 컴퓨팅 디바이스들의 조합, 예를 들면, DSP와 마이크로프로세서, 복수의 마이크로프로세서들, DSP 코어와 연계한 하나 이상의 마이크로프로세서들, 또는 임의의 다른 구성의 조합으로서 구현될 수도 있다.
펌웨어 및/또는 소프트웨어 구현에 있어서, 기법들은 랜덤 액세스 메모리(random access memory; RAM), 판독 전용 메모리(read-only memory; ROM), 비휘발성 RAM(non-volatile random access memory; NVRAM), PROM(programmable read-only memory), EPROM(erasable programmable read-only memory), EEPROM(electrically erasable PROM), 플래시 메모리, 컴팩트 디스크(compact disc; CD), 자기 또는 광학 데이터 스토리지 디바이스 등과 같은 컴퓨터 판독가능 매체 상에 저장된 명령들로서 구현될 수도 있다. 명령들은 하나 이상의 프로세서들에 의해 실행 가능할 수도 있고, 프로세서(들)로 하여금 본 개시에 설명된 기능의 특정 양태들을 수행하게 할 수도 있다.
소프트웨어로 구현되는 경우, 상기 기법들은 하나 이상의 명령들 또는 코드로서 컴퓨터 판독 가능한 매체 상에 저장되거나 또는 컴퓨터 판독 가능한 매체를 통해 전송될 수도 있다. 컴퓨터 판독가능 매체들은 한 장소에서 다른 장소로 컴퓨터 프로그램의 전송을 용이하게 하는 임의의 매체를 포함하여 컴퓨터 저장 매체들 및 통신 매체들 양자를 포함한다. 저장 매체들은 컴퓨터에 의해 액세스될 수 있는 임의의 이용 가능한 매체들일 수도 있다. 비제한적인 예로서, 이러한 컴퓨터 판독가능 매체는 RAM, ROM, EEPROM, CD-ROM 또는 다른 광학 디스크 스토리지, 자기 디스크 스토리지 또는 다른 자기 스토리지 디바이스들, 또는 소망의 프로그램 코드를 명령들 또는 데이터 구조들의 형태로 이송 또는 저장하기 위해 사용될 수 있으며 컴퓨터에 의해 액세스될 수 있는 임의의 다른 매체를 포함할 수 있다. 또한, 임의의 접속이 컴퓨터 판독가능 매체로 적절히 칭해진다.
예를 들어, 소프트웨어가 동축 케이블, 광섬유 케이블, 연선, 디지털 가입자 회선 (DSL), 또는 적외선, 무선, 및 마이크로파와 같은 무선 기술들을 사용하여 웹사이트, 서버, 또는 다른 원격 소스로부터 전송되면, 동축 케이블, 광섬유 케이블, 연선, 디지털 가입자 회선, 또는 적외선, 무선, 및 마이크로파와 같은 무선 기술들은 매체의 정의 내에 포함된다. 본원에서 사용된 디스크(disk) 와 디스크(disc)는, CD, 레이저 디스크, 광 디스크, DVD(digital versatile disc), 플로피디스크, 및 블루레이 디스크를 포함하며, 여기서 디스크들(disks)은 보통 자기적으로 데이터를 재생하고, 반면 디스크들(discs) 은 레이저를 이용하여 광학적으로 데이터를 재생한다. 위의 조합들도 컴퓨터 판독가능 매체들의 범위 내에 포함되어야 한다.
소프트웨어 모듈은, RAM 메모리, 플래시 메모리, ROM 메모리, EPROM 메모리, EEPROM 메모리, 레지스터들, 하드 디스크, 이동식 디스크, CD-ROM, 또는 공지된 임의의 다른 형태의 저장 매체 내에 상주할 수도 있다. 예시적인 저장 매체는, 프로세가 저장 매체로부터 정보를 판독하거나 저장 매체에 정보를 기록할 수 있도록, 프로세서에 연결될 수 있다. 대안으로, 저장 매체는 프로세서에 통합될 수도 있다. 프로세서와 저장 매체는 ASIC 내에 존재할 수도 있다. ASIC은 유저 단말 내에 존재할 수도 있다. 대안으로, 프로세서와 저장 매체는 유저 단말에서 개별 구성요소들로서 존재할 수도 있다.
이상 설명된 실시예들이 하나 이상의 독립형 컴퓨터 시스템에서 현재 개시된 주제의 양태들을 활용하는 것으로 기술되었으나, 본 개시는 이에 한정되지 않고, 네트워크나 분산 컴퓨팅 환경과 같은 임의의 컴퓨팅 환경과 연계하여 구현될 수도 있다. 또 나아가, 본 개시에서 주제의 양상들은 복수의 프로세싱 칩들이나 장치들에서 구현될 수도 있고, 스토리지는 복수의 장치들에 걸쳐 유사하게 영향을 받게 될 수도 있다. 이러한 장치들은 PC들, 네트워크 서버들, 및 휴대용 장치들을 포함할 수도 있다.
본 명세서에서는 본 개시가 일부 실시예들과 관련하여 설명되었지만, 본 개시의 발명이 속하는 기술분야의 통상의 기술자가 이해할 수 있는 본 개시의 범위를 벗어나지 않는 범위에서 다양한 변형 및 변경이 이루어질 수 있다. 또한, 그러한 변형 및 변경은 본 명세서에 첨부된 특허청구의 범위 내에 속하는 것으로 생각되어야 한다.

Claims (17)

  1. 적어도 하나의 프로세서에 의해 수행되는, 화자 분할 방법에 있어서,
    입력 신호에 포함된 음성 구간에 기초하여, 제1 세트의 화자 특징 벡터를 추출하는 단계;
    상기 추출된 제1 세트의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제1 중심 특징 벡터를 추출하는 단계; 및
    상기 입력 신호로부터 상기 제1 중심 특징 벡터와 연관된 제1 화자의 음성을 추출하는 단계
    를 포함하는, 화자 분할 방법.
  2. 제1항에 있어서,
    상기 제1 중심 특징 벡터를 추출하는 단계는,
    상기 추출된 제1 세트의 화자 특징 벡터에 기초하여, 벡터 공간 중 제1 중심 영역을 결정하는 단계; 및
    상기 제1 중심 영역에 포함된 하나 이상의 화자 특징 벡터에 기초하여, 상기 제1 중심 특징 벡터를 추출하는 단계
    를 포함하는, 화자 분할 방법.
  3. 제2항에 있어서,
    상기 제1 중심 영역을 결정하는 단계는,
    상기 벡터 공간 내에서 상기 화자 특징 벡터가 가장 밀집된 영역을 상기 제1 중심 영역으로 결정하는 단계
    를 포함하는, 화자 분할 방법.
  4. 제2항에 있어서,
    상기 제1 중심 영역을 결정하는 단계는,
    상기 벡터 공간 내에서 상기 추출된 제1 세트의 화자 특징 벡터의 각각을 중심으로 한 동일한 크기의 복수의 영역 중 가장 많은 수의 화자 특징 벡터를 포함하는 영역을 상기 제1 중심 영역으로 결정하는 단계
    를 포함하는, 화자 분할 방법.
  5. 제2항에 있어서,
    상기 제1 중심 영역에 포함된 하나 이상의 화자 특징 벡터에 기초하여, 제1 중심 특징 벡터를 추출하는 단계는,
    상기 제1 중심 영역에 포함된 하나 이상의 화자 특징 벡터의 평균을 상기 제1 중심 특징 벡터로 결정하는 단계
    를 포함하는, 화자 분할 방법.
  6. 제2항에 있어서,
    상기 제1 중심 영역에 포함된 하나 이상의 화자 특징 벡터에 기초하여, 제1 중심 특징 벡터를 추출하는 단계는,
    상기 제1 중심 영역의 중심 또는 중심에 가장 근접하여 위치한 화자 특징 벡터를 상기 제1 중심 특징 벡터로서 추출하는 단계
    를 포함하는, 화자 분할 방법.
  7. 제1항에 있어서,
    상기 추출된 제1 세트의 화자 특징 벡터에 대해 클러스터링을 수행하는 단계
    를 더 포함하는, 화자 분할 방법.
  8. 제7항에 있어서,
    상기 가장 높은 신뢰도를 갖는 제1 중심 특징 벡터를 추출하는 단계는,
    상기 클러스터링 결과 가장 많은 수의 화자 특징 벡터를 포함하는 클러스터로부터 상기 제1 중심 특징 벡터를 추출하는 단계
    를 포함하는, 화자 분할 방법.
  9. 제7항에 있어서,
    상기 가장 높은 신뢰도를 갖는 제1 중심 특징 벡터를 추출하는 단계는,
    벡터 공간 중 상기 클러스터링 결과 가장 많은 수의 화자 특징 벡터를 포함하는 클러스터에 대응되는 공간 내에서 제1 중심 영역을 결정하는 단계; 및
    상기 제1 중심 영역에 포함된 하나 이상의 화자 특징 벡터에 기초하여, 상기 제1 중심 특징 벡터를 추출하는 단계
    를 포함하는, 화자 분할 방법.
  10. 제1항에 있어서,
    상기 입력 신호로부터 상기 추출된 제1 화자의 음성을 제외한 잔여 신호에 음성 구간이 남아 있는지 여부를 판정하는 단계
    를 더 포함하는, 화자 분할 방법.
  11. 제10항에 있어서,
    상기 잔여 신호에 음성 구간이 남아 있다고 판정되는 경우, 상기 제1 세트의 화자 특징 벡터 중 상기 잔여 신호에 포함된 음성 구간을 기초로 추출된 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제2 중심 특징 벡터를 추출하는 단계; 및
    상기 입력 신호 또는 상기 잔여 신호로부터 상기 제2 중심 특징 벡터와 연관된 제2 화자의 음성을 추출하는 단계
    를 더 포함하는, 화자 분할 방법.
  12. 제10항에 있어서,
    상기 잔여 신호에 포함된 음성 구간에 기초하여, 제2 세트의 화자 특징 벡터를 추출하는 단계;
    상기 추출된 제2 세트의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제2 중심 특징 벡터를 추출하는 단계; 및
    상기 입력 신호 또는 상기 잔여 신호로부터 상기 제2 중심 특징 벡터와 연관된 제2 화자의 음성을 추출하는 단계
    를 더 포함하는, 화자 분할 방법.
  13. 제10항에 있어서,
    상기 잔여 신호에 음성 구간이 남아 있다고 판정되는 경우, 상기 입력 신호에 포함된 음성 구간 중 상기 제1 화자의 음성이 포함된 음성 구간을 제외한 잔여 음성 구간에 기초하여, 가장 높은 신뢰도를 갖는 제2 중심 특징 벡터를 추출하는 단계; 및
    상기 입력 신호 또는 상기 잔여 신호로부터 상기 제2 중심 특징 벡터와 연관된 제2 화자의 음성을 추출하는 단계
    를 더 포함하는, 화자 분할 방법.
  14. 제1항에 있어서,
    입력 신호에 포함된 음성 구간에 기초하여, 제1 세트의 화자 특징 벡터를 추출하는 단계는,
    상기 입력 신호에 포함된 복수의 단위 음성 구간의 각각에 기초하여, 상기 제1 세트의 화자 특징 벡터를 추출하는 단계
    를 포함하는, 화자 분할 방법.
  15. 제1항에 있어서,
    상기 입력 신호로부터 상기 제1 중심 특징 벡터와 연관된 제1 화자의 음성을 추출하는 단계는,
    목적 음성 추출 모델을 이용하여, 상기 입력 신호로부터 상기 제1 중심 특징 벡터와 연관된 상기 제1 화자의 음성을 추출하는 단계
    를 포함하는, 화자 분할 방법.
  16. 제1항에 따른 방법을 컴퓨터에서 실행하기 위한 명령어들을 기록한 컴퓨터 판독 가능한 비일시적 기록 매체.
  17. 정보 처리 시스템으로서,
    메모리; 및
    상기 메모리와 연결되고, 상기 메모리에 포함된 컴퓨터 판독 가능한 적어도 하나의 프로그램을 실행하도록 구성된 적어도 하나의 프로세서
    를 포함하고,
    상기 적어도 하나의 프로그램은,
    입력 신호에 포함된 음성 구간에 기초하여, 제1 세트의 화자 특징 벡터를 추출하고,
    상기 추출된 제1 세트의 화자 특징 벡터에 기초하여, 가장 높은 신뢰도를 갖는 제1 중심 특징 벡터를 추출하고,
    상기 입력 신호로부터 상기 제1 중심 특징 벡터와 연관된 제1 화자의 음성을 추출하기 위한 명령어들을 포함하는, 정보 처리 시스템.
PCT/KR2023/021003 2022-12-19 2023-12-19 화자 분할 방법 및 시스템 Ceased WO2024136409A1 (ko)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2025513626A JP2025530808A (ja) 2022-12-19 2023-12-19 話者分割方法及びシステム

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
KR10-2022-0178297 2022-12-19
KR1020220178297A KR102912207B1 (ko) 2022-12-19 2022-12-19 화자 분할 방법 및 시스템

Publications (1)

Publication Number Publication Date
WO2024136409A1 true WO2024136409A1 (ko) 2024-06-27

Family

ID=91589395

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2023/021003 Ceased WO2024136409A1 (ko) 2022-12-19 2023-12-19 화자 분할 방법 및 시스템

Country Status (3)

Country Link
JP (1) JP2025530808A (ko)
KR (1) KR102912207B1 (ko)
WO (1) WO2024136409A1 (ko)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102868068B1 (ko) * 2024-12-03 2025-10-01 에이치디씨랩스 주식회사 인공신경망을 이용하여 음성데이터로부터 실시간으로 화자를 분리하는 방법 및 시스템

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH09244683A (ja) * 1996-03-11 1997-09-19 Seiko Epson Corp 話者適応化方法および話者適応化装置
US20130006633A1 (en) * 2011-07-01 2013-01-03 Qualcomm Incorporated Learning speech models for mobile device users
JP5486565B2 (ja) * 2011-08-05 2014-05-07 日本電信電話株式会社 話者クラスタリング方法、話者クラスタリング装置、プログラム
JP6350148B2 (ja) * 2014-09-09 2018-07-04 富士通株式会社 話者インデキシング装置、話者インデキシング方法及び話者インデキシング用コンピュータプログラム
KR20220103507A (ko) * 2021-01-15 2022-07-22 네이버 주식회사 화자 식별과 결합된 화자 분리 방법, 시스템, 및 컴퓨터 프로그램

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS58132464A (ja) * 1982-01-27 1983-08-06 Toyoda Autom Loom Works Ltd 研掃装置における加工物の搬入搬出方法及びその装置
JPH10198393A (ja) * 1997-01-08 1998-07-31 Matsushita Electric Ind Co Ltd 会話記録装置
JP2014021957A (ja) * 2012-07-17 2014-02-03 Mitsunobu Watanabe 手前が高くなるキーボード
US10978059B2 (en) 2018-09-25 2021-04-13 Google Llc Speaker diarization using speaker embedding(s) and trained generative model
JP7222828B2 (ja) * 2019-06-24 2023-02-15 株式会社日立製作所 音声認識装置、音声認識方法及び記憶媒体
JP7471139B2 (ja) * 2020-04-30 2024-04-19 株式会社日立製作所 話者ダイアライゼーション装置、及び話者ダイアライゼーション方法
JP7103681B2 (ja) * 2020-12-18 2022-07-20 株式会社ミルプラトー 音声認識プログラム、音声認識方法、音声認識装置および音声認識システム
JP7664549B2 (ja) * 2021-03-01 2025-04-18 パナソニックIpマネジメント株式会社 発話分類装置および発話分類方法

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH09244683A (ja) * 1996-03-11 1997-09-19 Seiko Epson Corp 話者適応化方法および話者適応化装置
US20130006633A1 (en) * 2011-07-01 2013-01-03 Qualcomm Incorporated Learning speech models for mobile device users
JP5486565B2 (ja) * 2011-08-05 2014-05-07 日本電信電話株式会社 話者クラスタリング方法、話者クラスタリング装置、プログラム
JP6350148B2 (ja) * 2014-09-09 2018-07-04 富士通株式会社 話者インデキシング装置、話者インデキシング方法及び話者インデキシング用コンピュータプログラム
KR20220103507A (ko) * 2021-01-15 2022-07-22 네이버 주식회사 화자 식별과 결합된 화자 분리 방법, 시스템, 및 컴퓨터 프로그램

Also Published As

Publication number Publication date
JP2025530808A (ja) 2025-09-17
KR102912207B1 (ko) 2026-01-15
KR20240096049A (ko) 2024-06-26

Similar Documents

Publication Publication Date Title
WO2020139058A1 (en) Cross-device voiceprint recognition
WO2018070780A1 (en) Electronic device and method for controlling the same
WO2019143022A1 (ko) 음성 명령을 이용한 사용자 인증 방법 및 전자 장치
WO2019112342A1 (en) Voice recognition apparatus and operation method thereof cross-reference to related application
WO2020122653A1 (en) Electronic apparatus and controlling method thereof
WO2020204655A1 (en) System and method for context-enriched attentive memory network with global and local encoding for dialogue breakdown detection
WO2017071453A1 (zh) 一种语音识别的方法及装置
WO2021002649A1 (ko) 개별 화자 별 음성 생성 방법 및 컴퓨터 프로그램
WO2021251539A1 (ko) 인공신경망을 이용한 대화형 메시지 구현 방법 및 그 장치
WO2021154018A1 (en) Electronic device and method for controlling the electronic device thereof
WO2023191374A1 (ko) 구조식 이미지를 인식하는 인공 지능 장치 및 그 방법
WO2020054980A1 (ko) 음소기반 화자모델 적응 방법 및 장치
WO2024136409A1 (ko) 화자 분할 방법 및 시스템
WO2025005499A1 (ko) 인공지능 기반 폴리 사운드 제공 장치 및 방법
WO2024005383A1 (en) Online speaker diarization using local and global clustering
WO2021010578A1 (en) Electronic apparatus and method for recognizing speech thereof
WO2021172808A1 (en) System and method for personalization in intelligent multi-modal personal assistants
WO2024058573A1 (ko) 음성 합성 방법 및 시스템
WO2022050459A1 (en) Method, electronic device and system for generating record of telemedicine service
WO2023136505A1 (ko) 회의 안건 체크의 자동화 방법 및 시스템
WO2021107308A1 (ko) 전자 장치 및 이의 제어 방법
WO2024071855A1 (ko) 매크로봇 탐지 서비스 제공 방법 및 시스템
WO2023140720A1 (ko) 인공 지능 대화 서비스 시스템
WO2021256614A1 (ko) 화자가 표지된 텍스트 생성 방법
WO2022215905A1 (ko) 음성 녹음 후의 정보에 기초하여 생성된 음성 기록을 제공하는 방법 및 시스템

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23907692

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2025513626

Country of ref document: JP

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23907692

Country of ref document: EP

Kind code of ref document: A1