WO2025110602A1 - 인공지능을 이용하여 언어학 및 인지과학 기반의 음성기술이 적용된 더빙 생성 시스템 및 방법 - Google Patents

인공지능을 이용하여 언어학 및 인지과학 기반의 음성기술이 적용된 더빙 생성 시스템 및 방법 Download PDF

Info

Publication number
WO2025110602A1
WO2025110602A1 PCT/KR2024/017832 KR2024017832W WO2025110602A1 WO 2025110602 A1 WO2025110602 A1 WO 2025110602A1 KR 2024017832 W KR2024017832 W KR 2024017832W WO 2025110602 A1 WO2025110602 A1 WO 2025110602A1
Authority
WO
WIPO (PCT)
Prior art keywords
dubbing
information
text
artificial intelligence
generation system
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/KR2024/017832
Other languages
English (en)
French (fr)
Inventor
김홍국
김미담
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Aunionai Co Ltd
Gwangju Institute of Science and Technology
Original Assignee
Aunionai Co Ltd
Gwangju Institute of Science and Technology
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Aunionai Co Ltd, Gwangju Institute of Science and Technology filed Critical Aunionai Co Ltd
Publication of WO2025110602A1 publication Critical patent/WO2025110602A1/ko
Anticipated expiration legal-status Critical
Pending legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/26Speech to text systems
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/40Processing or translation of natural language
    • G06F40/58Use of machine translation, e.g. for multi-lingual retrieval, for server-side translation for client devices or for real-time translation
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/04Time compression or expansion
    • G10L21/055Time compression or expansion for synchronising with other signals, e.g. video signals
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/4302Content synchronisation processes, e.g. decoder synchronisation
    • H04N21/4307Synchronising the rendering of multiple content streams or additional data on devices, e.g. synchronisation of audio on a mobile phone with the video output on the TV screen
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/488Data services, e.g. news ticker
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/488Data services, e.g. news ticker
    • H04N21/4884Data services, e.g. news ticker for displaying subtitles

Definitions

  • the present disclosure relates to a dubbing generation system and method using artificial intelligence-based voice technology.
  • dubbing creation service market is expanding.
  • dubbing is created by recording the results of translating the original language, so not only does it take a long time to produce, but the dubbing's ability to convey meaning can vary depending on the producer's level of linguistic and cultural understanding.
  • the purpose of this disclosure is to provide a system and method capable of generating natural dubbing that reflects language-specific characteristics, cultural differences between countries, and the emotions of the speaker through artificial intelligence.
  • a dubbing generation system which includes a server and a terminal, wherein the server includes a communication unit configured to receive image information from the terminal, a processor including an artificial intelligence model that converts an audio signal included in the image information into text and generates dubbing information based on the converted text, and wherein the text is composed of international phonetic symbols.
  • the processor may generate dubbing information for the voice information based on the converted text and the translation result for the voice signal.
  • the processor may convert the translation result into information composed of international phonetic symbols based on the converted text, and generate dubbing information based on the converted information.
  • the processor can remove the audio signal from the image information and combine the dubbing information.
  • a dubbing generation method of a dubbing generation system including a server and a terminal includes a step in which the server receives video information from the terminal; a step in which the server converts a voice signal included in the video information into text; and a step in which the server generates dubbing information based on the converted text, wherein the text may be composed of international phonetic symbols.
  • the dubbing generation system can generate natural dubbing that sounds like it was pronounced by a native speaker.
  • Figure 1 is an overall system diagram of the present disclosure.
  • FIG. 2 is a block diagram of a server included in the dubbing generation system of the present disclosure.
  • FIG. 3 is a block diagram of a terminal included in the dubbing generation system of the present disclosure.
  • FIG. 4 is a block diagram of a processor included in the dubbing generation system of the present disclosure.
  • Figure 5 is a conceptual diagram showing a subtitle and dubbing service.
  • FIG. 6 is a flowchart of a method for generating dubbing using international phonetic symbols according to the present disclosure.
  • FIG. 7 is a conceptual diagram illustrating an embodiment of a dubbing generation system according to the present disclosure that converts a voice signal into text composed of international phonetic symbols.
  • first, second, etc. are used to distinguish one component from another, and the components are not limited by the aforementioned terms.
  • each step is used for convenience of explanation and do not describe the order of each step. Each step may be performed in a different order than specified unless the context clearly indicates a specific order.
  • the 'system according to the present disclosure includes all of various devices that can perform computational processing and provide results to a user.
  • the system according to the present disclosure may include all of a computer, a server device, and a portable terminal, or may be in the form of any one of them.
  • the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser.
  • the above server device is a server that processes information by communicating with an external device, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.
  • the above portable terminal may include, for example, all kinds of handheld-based wireless communication devices such as a PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, a smart phone, and a wearable device such as a watch, a ring, a bracelet, an anklet, a necklace, glasses, contact lenses, or a head-mounted-device (HMD).
  • a PCS Personal Communication System
  • GSM Global System for Mobile communications
  • PDC Personal Digital Cellular
  • PHS Personal Handyphone System
  • PDA Personal Digital Assistant
  • IMT International Mobile Telecommunication
  • CDMA Code Division Multiple Access
  • W-CDMA Wideband Code Division Multiple Access
  • WiBro Wireless Broadband Internet
  • the dubbing generation system according to the present disclosure may include at least one of a server (10) and a terminal (20).
  • the dubbing generation system according to the present disclosure may be implemented solely by the server (10) or the terminal (20), or may be implemented in the form of a system including at least one of the server (10) and the terminal (20).
  • the description of the dubbing generation system described below may be applied to both cases where the dubbing generation system according to the present disclosure is implemented solely by the server (10) or the terminal (20), or may be implemented in the form of a system including at least one of the server (10) and the terminal (20).
  • the server (10) is connected to at least one terminal (20) via a network, transmits information to each of a plurality of terminals, and generates data necessary for dubbing generation learning based on information received from at least one of the terminals (20).
  • the terminal (20) is not limited to the portable terminal described above, and may include a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a processor.
  • FIG. 2 is a block diagram of a server included in the dubbing generation system of the present disclosure.
  • a server (100) may include at least one of a communication unit (110), a storage unit (120), and a processor (130).
  • the communication unit (110) can communicate with at least one of a terminal, an external storage (e.g., a database (140)), an external server, and a cloud server.
  • a terminal e.g., a terminal, a server, and a cloud server.
  • an external storage e.g., a database (140)
  • an external server e.g., a server, and a cloud server.
  • an external server or cloud server may be configured to perform at least a part of the role of the processor (130). That is, data processing or data operations, etc. may be performed on an external server or cloud server, and the present invention does not place any special restrictions on this method.
  • the communication unit (110) can support various communication methods according to the communication standards of the communicating object (e.g., electronic device, external server, device, etc.).
  • the communicating object e.g., electronic device, external server, device, etc.
  • the communication unit (110) may be configured to communicate with a communication target using at least one of the following technologies: WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (BluetoothTM), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus).
  • WLAN Wireless LAN
  • Wi-Fi Wireless-Fidelity
  • Wi-Fi Wireless Fidelity
  • Direct Wireless Fidelity
  • the storage unit (120) may be configured to store various information related to the present invention.
  • the storage unit (120) may be provided in the device itself according to the present invention.
  • at least a part of the storage unit (120) may mean at least one of a database (DB, 140) and a cloud storage (or a cloud server). That is, the storage unit (120) may be sufficient as a space where information required for the device and method according to the present invention is stored, and it may be understood that there is no limitation on the physical space. Accordingly, in the following, the storage unit (120), the database (140), the external storage, and the cloud storage (or the cloud server) will not be separately distinguished, and will all be expressed as the storage unit (120).
  • the processor (130) may be configured to control the overall operation of the device related to the present invention.
  • the processor (130) may process signals, data, information, etc. input or output through the components discussed above, or provide or process appropriate information or functions to the user.
  • the processor (130) includes at least one CPU (Central Processing Unit) and can perform functions according to the present invention.
  • CPU Central Processing Unit
  • At least one component may be added or deleted in accordance with the performance of the components illustrated in FIG. 2.
  • the mutual positions of the components may be changed in accordance with the performance or structure of the device.
  • FIG. 3 is a block diagram of a terminal included in the dubbing generation system of the present disclosure.
  • a terminal (200) may include a communication unit (210), an input unit (220), a display unit (230), a processor (240), etc.
  • the components illustrated in FIG. 3 are not essential for implementing a dubbing generation system according to the present disclosure, and thus, the terminal described in this specification may have more or fewer components than the components listed above.
  • the communication unit (210) may include one or more components that enable communication with an external device, and may include, for example, at least one of a broadcast receiving module, a wired communication module, a wireless communication module, a short-range communication module, and a location information module.
  • the wired communication module may include various wired communication modules such as a Local Area Network (LAN) module, a Wide Area Network (WAN) module, or a Value Added Network (VAN) module, as well as various cable communication modules such as a Universal Serial Bus (USB), a High Definition Multimedia Interface (HDMI), a Digital Visual Interface (DVI), RS-1302 (recommended standard1302), power line communication, or plain old telephone service (POTS).
  • LAN Local Area Network
  • WAN Wide Area Network
  • VAN Value Added Network
  • USB Universal Serial Bus
  • HDMI High Definition Multimedia Interface
  • DVI Digital Visual Interface
  • RS-1302 recommended standard1302
  • POTS plain old telephone service
  • the wireless communication module may include a wireless communication module that supports various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), LTE (Long Term Evolution), 4G, 5G, and 6G, in addition to a WiFi module and a Wireless broadband module.
  • GSM Global System for Mobile Communication
  • CDMA Code Division Multiple Access
  • WCDMA Wideband Code Division Multiple Access
  • UMTS universalal mobile telecommunications system
  • TDMA Time Division Multiple Access
  • LTE Long Term Evolution
  • 4G Long Term Evolution
  • 5G Fifth Generation
  • 6G Wireless broadband module
  • the input unit (220) is for inputting image information (or signal), audio information (or signal), data, or information input from a user, and may include at least one camera, at least one microphone, and at least one user input unit. Voice data or image data collected from the input unit may be analyzed and processed as a user control command.
  • the display unit (230) is for generating output related to visual, auditory or tactile sensations, and may include at least one of a display unit, an audio output unit, a haptic module and a light output unit.
  • the display unit may be formed as a layer structure with a touch sensor or formed as an integral part, thereby implementing a touch screen.
  • This touch screen may function as a user input unit that provides an input interface between the device and a user, and may provide an output interface between the device and a user.
  • the display unit displays (outputs) information processed by this device.
  • the display unit can display execution screen information of an application program (e.g., an application) running on this device, or UI (User Interface) or GUI (Graphical User Interface) information according to such execution screen information.
  • an application program e.g., an application
  • UI User Interface
  • GUI Graphic User Interface
  • the terminal described above may further include an interface unit and a memory.
  • the interface section serves as a passage for various types of external devices connected to the device.
  • the interface section may include at least one of a wired/wireless headset port, an external charger port, a wired/wireless data port, a memory card port, a port for connecting a device equipped with an identification module (SIM), an audio I/O (Input/Output) port, a video I/O (Input/Output) port, and an earphone port.
  • SIM identification module
  • the device may perform appropriate control related to an external device connected to the interface section.
  • the memory can store data supporting various functions of the device, programs for the operation of the processor, can store input/output data (e.g., music files, still images, moving images, etc.), and can store a plurality of application programs (or applications) running on the device, data for the operation of the device, and commands. At least some of these application programs can be downloaded from an external server via wireless communication.
  • Such memory may include at least one type of storage medium among a flash memory type, a hard disk type, an SSD (Solid State Disk type), an SDD (Silicon Disk Drive) type, a multimedia card micro type, a card type memory (for example, an SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk.
  • the memory may be a database separate from the device but connected by a wire or wirelessly.
  • the terminal described above includes a processor (240).
  • the processor may be implemented as a memory storing data for an algorithm for controlling the operation of components within the device or a program reproducing the algorithm, and at least one processor (not shown) performing the above-described operation using the data stored in the memory.
  • the memory and the processor may be implemented as separate chips.
  • the memory and the processor may be implemented as a single chip.
  • the processor may control one or more of the components described above in combination to implement various embodiments according to the present disclosure described in the drawings below on the device.
  • At least one component may be added or deleted in accordance with the performance of the components illustrated in FIGS. 1 to 3.
  • the mutual positions of the components may be changed in accordance with the performance or structure of the device.
  • a processor included in at least one of the server and the terminal may include a plurality of modules for implementing a dubbing generation system to be described later.
  • the processor (300) may include a voice recognition module (310) and an artificial intelligence module (320).
  • the dubbing generation method to be described later is described as being implemented by the operations of the modules, but the performance of the operations of each step described later need not necessarily be performed by the modules.
  • the function related to artificial intelligence is operated through the processor and memory installed in the above-described server and terminal.
  • the processor may be composed of one or more processors.
  • one or more processors may be a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU, a VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU.
  • One or more processors control to process input data according to a predefined operation rule or artificial intelligence model stored in the memory.
  • the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
  • the predefined operation rules or artificial intelligence models are characterized by being created through learning.
  • being created through learning means that the basic artificial intelligence model is learned by using a plurality of learning data by a learning algorithm, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose).
  • Such learning may be performed in the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and/or system.
  • Examples of the learning algorithm include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
  • the artificial intelligence model may be composed of a plurality of neural network layers.
  • Each of the plurality of neural network layers has a plurality of weight values, and performs a neural network operation through an operation between the operation result of the previous layer and the plurality of weights.
  • the plurality of weights of the plurality of neural network layers may be optimized by the learning result of the artificial intelligence model. For example, the plurality of weights may be updated so that a loss value or a cost value obtained from the artificial intelligence model is reduced or minimized during the learning process.
  • the artificial neural network may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or a deep Q-network.
  • DNN deep neural network
  • a processor can implement artificial intelligence.
  • Artificial intelligence refers to a machine learning method based on an artificial neural network that imitates human biological neurons to enable a machine to learn.
  • Artificial intelligence methodologies can be divided into supervised learning in which input data and output data are provided together as training data according to a learning method, so that an answer (output data) to a problem (input data) is determined, unsupervised learning in which only input data is provided without output data, so that an answer (output data) to a problem (input data) is not determined, and reinforcement learning in which a reward is given from an external environment whenever an action is taken in a current state, and learning is performed in a direction to maximize this reward.
  • artificial intelligence methodologies can be categorized by architecture, which is the structure of the learning model.
  • architectures of widely used deep learning technologies can be categorized into convolutional neural networks (CNNs), recurrent neural networks (RNNs), transformers, and generative adversarial networks (GANs).
  • CNNs convolutional neural networks
  • RNNs recurrent neural networks
  • GANs generative adversarial networks
  • the present device and system may include an artificial intelligence model.
  • the artificial intelligence model may be one artificial intelligence model or may be implemented as multiple artificial intelligence models.
  • the artificial intelligence model may be composed of a neural network (or an artificial neural network) and may include a statistical learning algorithm that mimics biological neurons in machine learning and cognitive science.
  • a neural network may refer to a model in which artificial neurons (nodes) that form a network by combining synapses change the strength of the synapses through learning and have problem-solving capabilities.
  • Neurons of a neural network may include a combination of weights or biases.
  • a neural network may include one or more layers composed of one or more neurons or nodes.
  • the device may include an input layer, a hidden layer, and an output layer.
  • a neural network that constitutes the device may infer a desired result (output) from an arbitrary input (input) by changing the weights of neurons through learning.
  • the processor can generate a neural network, train (or learn) a neural network, perform a calculation based on received input data, generate an information signal based on the result of the calculation, or retrain the neural network.
  • the models of the neural network can include various types of models such as CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, R-CNN (Region with Convolution Neural Network), RPN (Region Proposal Network), RNN (Recurrent Neural Network), S-DNN (Stacking-based deep Neural Network), S-SDNN (State-Space Dynamic Neural Network), Deconvolution Network, DBN (Deep Belief Network), RBM (Restrcted Boltzman Machine), Fully Convolutional Network, LSTM (Long Short-Term Memory) Network, Classification Network, etc., but are not limited thereto.
  • the processor can include one or more processors for performing calculations according to the models of the neural network.
  • the neural network can include a deep It may include a neural network
  • Neural networks include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), perceptron, multilayer perceptron, FF (Feed Forward), RBF (Radial Basis Network), DFF (Deep Feed Forward), LSTM (Long Short Term Memory), GRU (Gated Recurrent Unit), AE (Auto Encoder), VAE (Variational Auto) Encoder), DAE (Denoising Auto Encoder), SAE (Sparse Auto Encoder), MC (Markov Chain), HN (Hopfield Network), BM (Boltzmann Machine), RBM (Restricted Boltzmann Machine), DBN (Depp Belief Network), DCN (Deep Convolutional Network), DN (Deconvolutional Network), DCIGN (Deep Convolutional Inverse Graphics Network), Generative Adversarial Network (GAN), Liquid State Machine (LSM), Extreme Learning Machine (ELM), It will be understood by those skilled in the art that any neural network may
  • the processor may be configured to perform a CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, R-CNN (Region with Convolution Neural Network), RPN (Region Proposal Network), RNN (Recurrent Neural Network), S-DNN (Stacking-based deep Neural Network), S-SDNN (State-Space Dynamic Neural Network), Deconvolution Network, DBN (Deep Belief Network), RBM (Restrcted Boltzman Machine), Fully Convolutional Network, LSTM (Long Short-Term Memory) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC/QA for natural language processing, Text Analysis, Dialog System, GPT-3, GPT-4, Visual Analytics for vision processing, Visual Understanding, Video Synthesis, ResNet for data intelligence, Anomaly Detection, Prediction, Time-Series Forecasting,
  • CNN Convolution Neural Network
  • CNN
  • Figure 5 is a conceptual diagram showing a subtitle and dubbing service.
  • a step of receiving voice information is performed.
  • the step of receiving voice information may be a step of receiving video information.
  • the Jamik dubbing generation system according to the present disclosure receives video information and utilizes the voice information included in the video information for dubbing generation.
  • the step of receiving voice information may be a step of receiving voice information itself, not video information.
  • audio information may include all data related to sound included in an image.
  • audio information may include audio signals and background sounds included in image information.
  • the processor receives image information and processes audio information included in the image information.
  • a step is performed to separate the voice signal and background sound information from the voice information.
  • the processor can separate the voice signal and background sound information from the voice information and generate separate files.
  • the voice signal is the speaker's voice included in the video information and is the information that is the target of dubbing generation. Later, the voice signal separated from the voice information is generated as at least one of the dubbing.
  • the separation of the above speech signal and background sound information can be performed through a separate artificial intelligence model trained to separate the speech signal and background sound information from the speech information.
  • the above voice signal can be converted into text through the voice recognition module (310).
  • the text can be made of the language that is the basis for producing the image from which the voice signal is separated.
  • the language that is the basis of video production is called the “source language,” and the language in which dubbing is to be created is called the “target language.”
  • the above translation result is in the form of text in the target language and can be provided as subtitles for the original video.
  • the type of the above target language can be input from the video producer or the user watching the video.
  • a video producer can specify the target language in advance and generate subtitles for at least one target language from the video production stage.
  • the video producer may generate subtitles by specifying a target language desired by the viewer of the video, rather than generating subtitles. Subtitles generated by a specific video viewer may be re-provided to the video viewer who requests subtitles in that target language.
  • the above translation can be performed by a separate artificial intelligence model trained to translate a text in the original language into a target language.
  • the type of artificial intelligence model for translation is not specifically limited.
  • the processor generates the subtitle and synthesizes it into the original image.
  • the processor can set a time section for inserting the subtitle by considering a time section in which a voice signal is generated in the image information.
  • the processor can match a time section in which a voice signal is generated in the image information in units of syllables, words, or sentences.
  • the time information matched in this way is also matched to the converted text after the voice signal is converted into text. Thereafter, the time information can also be matched to the result of translating the converted text.
  • the processor can insert the subtitle into the image information based on the time information.
  • a step is performed to generate a speech signal (hereinafter, dubbing information) in a target language based on at least one of the separated speech signal, the converted text, and the translation result.
  • a speech signal hereinafter, dubbing information
  • the dubbing generation method according to the present disclosure can be implemented in the step of generating the above-described voice signal.
  • an artificial intelligence module (320) receives a voice signal (S210).
  • the information input to the artificial intelligence module (320) may be audio information included in the image information, or an audio signal in which background sound information is separated from the audio information.
  • the above voice signal can be converted into text through the voice recognition module (310).
  • the text can be in the original language.
  • the above converted text composed of the original language is converted into a text composed of the target language through translation.
  • the above translation may be performed by a known artificial intelligence model or by a separately trained artificial intelligence model.
  • the next step is to have the artificial intelligence model convert the voice signal into text composed of the International Phonetic Alphabet (IPA) (S223).
  • IPA International Phonetic Alphabet
  • the above artificial intelligence module (320) may include an artificial intelligence model trained to convert voice information into the International Phonetic Association (IPA).
  • IPA International Phonetic Association
  • the artificial intelligence model may be, but is not limited to, a transformer.
  • the artificial intelligence model can receive voice information such as “Can you lend me ten thousand won?” and convert it into the following international phonetic symbol.
  • the artificial intelligence model can be trained such that 'src' is a speech signal composed of Korean, 'tgt' is an English IPA embedding, and 'output' is an English IPA embedding.
  • the artificial intelligence model generates dubbing information based on the translation result and the text composed of international phonetic symbols (S230).
  • the artificial intelligence model can generate dubbing information for the voice information based on the translated text and the translation result for the voice signal.
  • the AI model converts the translation result into information composed of the international phonetic symbol based on the converted text.
  • the AI model can be trained to receive a text composed of the international phonetic symbol generated based on a speech signal composed of the original language and a translation result for the speech signal composed of the original language, and generate a text composed of the international phonetic symbol and translated into the target language.
  • dubbing information that reflects regional, social, cultural, and linguistic differences between the original language and the target language can be generated.
  • a step of synthesizing the generated voice signal (dubbing information) and the separated background sound information is performed.
  • the processor can use the translation result to generate a subtitle for the image information, and can use the subtitle generation result to synthesize the generated voice signal and the separated background sound information.
  • the processor can search for time information at which the subtitle is output from the image information, and synthesize the generated voice signal and the separated background sound information by utilizing the searched time information.
  • the processor can match the time interval in which the voice signal is generated in the image information in units of syllables, phrases, or sentences.
  • the time information matched in this way is also matched to the converted text after the voice signal is converted into text. Thereafter, the time information can also be matched to the result of translating the converted text.
  • the processor can insert subtitles into the image information based on the time information.
  • the processor can also utilize the above-described time information to generate dubbing.
  • the processor can synthesize the generated voice signal and the separated background sound information by utilizing the result of separating the voice signal and the background sound information from the voice information included in the image information. Specifically, the processor can synthesize the generated voice signal and the separated background sound information by utilizing the time information at which the voice signal was separated from the voice information.
  • the processor When the processor separates the voice signal from the background sound information, it can match the information defining the time interval from which the voice signal is separated to the background sound information.
  • This time interval information can be defined in units of syllables, phrases, or sentences and can be matched to the background sound information. The above time information can be utilized when synthesizing dubbing information to the background sound information.
  • dubbing generation system since background sound separated from voice information is synthesized after dubbing is generated, dubbing can be generated without loss of sound source.
  • dubbing can be generated at a faster speed compared to conventional methods of manually generating dubbing.
  • the disclosed embodiments may be implemented in the form of a recording medium storing instructions executable by a computer.
  • the instructions may be stored in the form of program codes, and when executed by a processor, may generate program modules to perform the operations of the disclosed embodiments.
  • the recording medium may be implemented as a computer-readable recording medium.
  • Computer-readable storage media include all types of storage media that store instructions that can be deciphered by a computer. Examples include ROM (Read Only Memory), RAM (Random Access Memory), magnetic tape, magnetic disk, flash memory, and optical data storage devices.
  • ROM Read Only Memory
  • RAM Random Access Memory
  • magnetic tape magnetic tape
  • magnetic disk magnetic disk
  • flash memory optical data storage devices

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Human Computer Interaction (AREA)
  • Quality & Reliability (AREA)
  • Theoretical Computer Science (AREA)
  • Artificial Intelligence (AREA)
  • General Health & Medical Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)

Abstract

본 개시는 인공지능 기반 음성기술이 적용된 더빙 생성 시스템 및 방법에 관한 것이다. 본 개시에 따른 더빙 생성 시스템은 서버 및 단말기를 포함하고, 상기 서버는 상기 단말기로부터 영상 정보를 수신하도록 이루어지는 통신부, 상기 영상 정보에 포함된 음성 신호를 텍스트로 변환하고, 변환된 텍스트를 기반으로 더빙 정보를 생성하는 인공지능 모델을 포함하는 프로세서를 포함하고, 상기 텍스트는 국제음성기호로 구성되는 것을 특징으로 하는 더빙 생성 시스템을 제공할 수 있다.

Description

인공지능을 이용하여 언어학 및 인지과학 기반의 음성기술이 적용된 더빙 생성 시스템 및 방법
본 개시는 인공지능 기반 음성기술이 적용된 더빙 생성 시스템 및 방법에 관한 것이다.
최근 미디어 콘텐츠 시장과 관련된 비디오 스트리밍 시장이 빠른 속도로 성장하고 있다. 또한, 비디오 스트리밍 시장이 국내가 아닌 글로벌 시장으로 확대됨에 따라, 더빙 생성 기술에 대한 니즈가 증가하고 있다.
구체적으로, 점차 커지는 미디어 콘텐츠 시장에서 1인 창작자는 경쟁력을 높이기 위해 더빙 생성을 통한 고품질의 동영상을 제공하고 있다. 이에 따라, 더빙 생성 서비스 시장이 확대되고 있다.
현재 더빙은 원어를 번역한 결과를 녹음하여 생성되기 때문에, 그 제작 시간이 오래걸릴 뿐만 아니라, 제작자의 언어 및 문화적 이해도에 따라 더빙의 의미 전달력이 상이해질 수 있다.
또한, 더빙 생성의 경우, 원어민의 언어 및 감정을 전달하는데 어려움이 있으며, 이로 인해 영상자체가 부자연스러워질 수 있는 있는 문제가 있다.
이에, 원작이 제작된 국가의 언어 및 문화적 차이를 반영함과 동시에 영상에 삽입하더라도 부자연스러움이 없는 더빙을 생성할 수 있는 기술에 대한 니즈가 존재한다.
본 개시는 인공지능을 통해 언어별 특징, 국가간 문화적 차이, 화자의 감정이반영된 자연스러운 더빙을 생성할 수 있는 시스템 및 방법을 제공하는 것을 그 목적으로 한다.
상술한 목적을 달성하기 위해, 본 개시에 따른 더빙 생성 시스템은 서버 및 단말기를 포함하고, 상기 서버는 상기 단말기로부터 영상 정보를 수신하도록 이루어지는 통신부, 상기 영상 정보에 포함된 음성 신호를 텍스트로 변환하고, 변환된 텍스트를 기반으로 더빙 정보를 생성하는 인공지능 모델을 포함하는 프로세서를 포함하고, 상기 텍스트는 국제음성기호로 구성되는 것을 특징으로 하는 더빙 생성 시스템을 제공할 수 있다.
일 실시 예에 있어서, 상기 프로세서는 상기 변환된 텍스트 및 상기 음성 신호에 대한 번역 결과를 기반으로, 상기 음성 정보에 대한 더빙 정보를 생성할 수 있다.
일 실시 예에 있어서, 상기 프로세서는 상기 변환된 텍스트를 기반으로, 상기 프로세서는 상기 번역 결과를 국제음성기호로 구성된 정보로 변환하고, 변환된 정보를 기반으로 더빙 정보를 생성할 수 있다.
일 실시 예에 있어서, 상기 프로세서는 상기 영상 정보에서 상기 음성 신호를 제거하고, 상기 더빙 정보를 결합할 수 있다.
일 실시 예에 있어서, 상기 인공지능 모델은 트랜스포머(transformer)일 수 있다.
또한, 본 개시에 따른 서버 및 단말기를 포함하는 더빙 생성 시스템의 더빙 생성 방법은, 상기 서버가 상기 단말기로부터 영상 정보를 수신하는 단계; 상기 서버가 상기 영상 정보에 포함된 음성 신호를 텍스트로 변환하는 단계; 및 상기 서버가 상기 변환된 텍스트를 기반으로 더빙 정보를 생성하는 단계를 포함하고, 상기 텍스트는 국제음성기호로 구성될 수 있다.
본 개시에 따르면, 지역적, 사회적, 문화적, 언어적 차이가 반영된 더빙을 생성할 수 있게 된다. 이에 따라, 본 개시에 따른 더빙 생성 시스템은 원어민이 발음하는 듯한 자연스러운 더빙을 생성할 수 있게 된다.
또한, 본 개시에 따르면, 화자가 더빙 생성을 위해 음성을 녹음할 필요없이 빠른 속도로 더빙을 제작하는 것이 가능해진다.
도 1은 본 개시의 전반적 시스템 도면이다.
도 2는 본 개시의 더빙 생성 시스템에 포함된 서버의 블록도이다.
도 3은 본 개시의 더빙 생성 시스템에 포함된 단말기의 블록도이다.
도 4는 본 개시의 더빙 생성 시스템에 포함된 프로세서의 블록도이다.
도 5는 자막 및 더빙 제공 서비스를 나타내는 개념도이다.
도 6은 본 개시에 따른 국제음성기호를 활용하여 더빙을 생성하는 방법의 흐름도이다.
도 7은 본 개시에 따른 더빙 생성 시스템이 음성 신호를 국제음성기호로 구성된 텍스트로 변환하는 일 실시 예를 나타내는 개념도이다.
본 개시 전체에 걸쳐 동일 참조 부호는 동일 구성요소를 지칭한다. 본 개시가 실시예들의 모든 요소들을 설명하는 것은 아니며, 본 개시가 속하는 기술분야에서 일반적인 내용 또는 실시예들 간에 중복되는 내용은 생략한다. 명세서에서 사용되는 '부, 모듈, 부재, 블록'이라는 용어는 소프트웨어 또는 하드웨어로 구현될 수 있으며, 실시예들에 따라 복수의 '부, 모듈, 부재, 블록'이 하나의 구성요소로 구현되거나, 하나의 '부, 모듈, 부재, 블록'이 복수의 구성요소들을 포함하는 것도 가능하다.
명세서 전체에서, 어떤 부분이 다른 부분과 "연결"되어 있다고 할 때, 이는 직접적으로 연결되어 있는 경우뿐 아니라, 간접적으로 연결되어 있는 경우를 포함하고, 간접적인 연결은 무선 통신망을 통해 연결되는 것을 포함한다.
또한 어떤 부분이 어떤 구성요소를 "포함"한다고 할 때, 이는 특별히 반대되는 기재가 없는 한 다른 구성요소를 제외하는 것이 아니라 다른 구성요소를 더 포함할 수 있는 것을 의미한다.
명세서 전체에서, 어떤 부재가 다른 부재 "상에" 위치하고 있다고 할 때, 이는 어떤 부재가 다른 부재에 접해 있는 경우뿐 아니라 두 부재 사이에 또 다른 부재가 존재하는 경우도 포함한다.
제 1, 제 2 등의 용어는 하나의 구성요소를 다른 구성요소로부터 구별하기 위해 사용되는 것으로, 구성요소가 전술된 용어들에 의해 제한되는 것은 아니다.
단수의 표현은 문맥상 명백하게 예외가 있지 않는 한, 복수의 표현을 포함한다.
각 단계들에 있어 식별부호는 설명의 편의를 위하여 사용되는 것으로 식별부호는 각 단계들의 순서를 설명하는 것이 아니며, 각 단계들은 문맥상 명백하게 특정 순서를 기재하지 않는 이상 명기된 순서와 다르게 실시될 수 있다.
이하 첨부된 도면들을 참고하여 본 개시의 작용 원리 및 실시예들에 대해 설명한다.
본 명세서에서 '본 개시에 따른 시스템'은 연산처리를 수행하여 사용자에게 결과를 제공할 수 있는 다양한 장치들이 모두 포함된다. 예를 들어, 본 개시에 따른 시스템은, 컴퓨터, 서버 장치 및 휴대용 단말기를 모두 포함하거나, 또는 어느 하나의 형태가 될 수 있다.
여기에서, 상기 컴퓨터는 예를 들어, 웹 브라우저(WEB Browser)가 탑재된 노트북, 데스크톱(desktop), 랩톱(laptop), 태블릿 PC, 슬레이트 PC 등을 포함할 수 있다.
상기 서버 장치는 외부 장치와 통신을 수행하여 정보를 처리하는 서버로써, 애플리케이션 서버, 컴퓨팅 서버, 데이터베이스 서버, 파일 서버, 게임 서버, 메일 서버, 프록시 서버 및 웹 서버 등을 포함할 수 있다.
상기 휴대용 단말기는 예를 들어, 휴대성과 이동성이 보장되는 무선 통신 장치로서, PCS(Personal Communication System), GSM(Global System for Mobile communications), PDC(Personal Digital Cellular), PHS(Personal Handyphone System), PDA(Personal Digital Assistant), IMT(International Mobile Telecommunication)-2000, CDMA(Code Division Multiple Access)-2000, W-CDMA(W-Code Division Multiple Access), WiBro(Wireless Broadband Internet) 단말, 스마트 폰(Smart Phone) 등과 같은 모든 종류의 핸드헬드(Handheld) 기반의 무선 통신 장치와 시계, 반지, 팔찌, 발찌, 목걸이, 안경, 콘택트 렌즈, 또는 머리 착용형 장치(head-mounted-device(HMD) 등과 같은 웨어러블 장치를 포함할 수 있다.
이하에서는, 본 개시에 따른 더빙 생성 시스템에 대하여 설명한다.
도 1을 참고하면, 본 개시에 따른 더빙 생성 시스템은 서버(10) 및 단말기(20) 중 적어도 하나를 포함할 수 있다. 구체적으로, 본 개시에 따른 더빙 생성 시스템은 서버(10) 또는 단말기(20)에 의해 단독으로 구현되거나, 서버(10) 및 단말기(20) 중 적어도 하나를 포함하는 시스템 형태로 구현될 수 있다. 이하에서 설명하는 더빙 생성 시스템에 관한 설명은 본 개시에 따른 더빙 생성 시스템이 서버(10) 또는 단말기(20)에 의해 단독으로 구현되거나, 서버(10) 및 단말기(20) 중 적어도 하나를 포함하는 시스템 형태로 구현되는 경우에 모두 적용될 수 있다.
서버(10)는 적어도 하나의 단말기(20)와 네트워크로 연결되며, 복수의 단말기 각각에 정보를 전송하고, 단말기(20) 중 적어도 하나로부터 수신된 정보에 기반하여 더빙 생성 학습에 필요한 데이터를 생성한다.
한편, 상기 단말기(20)는 상술한 휴대용 단말기에 한정되지 않고, 프로세서가 탑재된 노트북, 데스크톱(desktop), 랩톱(laptop), 태블릿 PC, 슬레이트 PC 등을 포함할 수 있는 것은 통상의 기술자에게 자명하다.
이하에서는, 본 개시에 따른 더빙 생성 시스템을 구현하기 위한 서버(10) 및 단말기(20) 각각에 대하여 설명한다.
도 2는 본 개시의 더빙 생성 시스템에 포함된 서버의 블록도이다.
본 개시에 따른 서버(100)는 통신부(110), 저장부(120) 및 프로세서(130) 중 적어도 하나를 포함할 수 있다.
통신부(110)는 단말기, 외부 저장소(예를 들어, 데이터베이스(database, 140)), 외부 서버 및 클라우드 서버 중 적어도 하나와 통신을 수행할 수 있다.
한편, 외부 서버 또는 클라우드 서버에서는, 프로세서(130)의 적어도 일부의 역할을 수행하도록 구성될 수 있다. 즉, 데이터 처리 또는 데이터 연산 등의 수행은 외부 서버 또는 클라우드 서버에서 이루어지는 것이 가능하며, 본 발명에서는 이러한 방식에 대한 특별한 제한을 두지 않는다.
한편, 통신부(110)는 통신하는 대상(예를 들어, 전자기기, 외부 서버, 디바이스 등)의 통신 규격에 따라 다양한 통신 방식을 지원할 수 있다.
예를 들어, 통신부(110)는, WLAN(Wireless LAN), Wi-Fi(Wireless-Fidelity), Wi-Fi(Wireless Fidelity) Direct, DLNA(Digital Living Network Alliance), WiBro(Wireless Broadband), WiMAX(World Interoperability for Microwave Access), HSDPA(High Speed Downlink Packet Access), HSUPA(High Speed Uplink Packet Access), LTE(Long Term Evolution), LTE-A(Long Term Evolution-Advanced), 5G(5th Generation Mobile Telecommunication ), 블루투스(Bluetooth™), RFID(Radio Frequency Identification), 적외선 통신(Infrared Data Association; IrDA), UWB(Ultra-Wideband), ZigBee, NFC(Near Field Communication), Wi-Fi Direct, Wireless USB(Wireless Universal Serial Bus) 기술 중 적어도 하나를 이용하여, 통신 대상과 통신하도록 이루어질 수 있다.
다음으로 저장부(120)는, 본 발명과 관련된 다양한 정보를 저장하도록 이루어질 수 있다. 본 발명에서 저장부(120)는 본 발명에 따른 장치 자체에 구비될 수 있다. 이와 다르게, 저장부(120)의 적어도 일부는, 데이터베이스(database: DB, 140) 클라우드 저장소(또는 클라우드 서버) 중 적어도 하나를 의미할 수 있다. 즉, 저장부(120)는 본 발명에 따른 장치 및 방법을 위하여 필요한 정보가 저장되는 공간이면 충분하며, 물리적인 공간에 대한 제약은 없는 것으로 이해될 수 있다. 이에, 이하에서는, 저장부(120), 데이터베이스(140), 외부 저장소, 클라우드 저장소(또는 클라우드 서버)를 별도로 구분하지 않고, 모두 저장부(120)라고 표현하도록 한다.
다음으로, 프로세서(130)는 본 발명과 관련된 장치의 전반적인 동작을 제어하도록 이루어질 수 있다. 프로세서(130)는 위에서 살펴본 구성요소들을 통해 입력 또는 출력되는 신호, 데이터, 정보 등을 처리하거나 사용자에게 적절한 정보 또는 기능을 제공 또는 처리할 수 있다.
프로세서(130)는 적어도 하나의 CPU(Central Processing Unit, 중앙처리장치)를 포함하여, 본 발명에 따른 기능을 수행할 수 있다.
도 2에 도시된 구성 요소들의 성능에 대응하여 적어도 하나의 구성요소가 추가되거나 삭제될 수 있다. 또한, 구성 요소들의 상호 위치는 장치의 성능 또는 구조에 대응하여 변경될 수 있다는 것은 당해 기술 분야에서 통상의 지식을 가진 자에게 용이하게 이해될 것이다.
이하, 본 개시의 더빙 생성 시스템에 포함된 단말기에 대하여 구체적으로 설명한다.
도 3은 본 개시의 더빙 생성 시스템에 포함된 단말기의 블록도이다.
도 3을 참고하면, 본 개시에 따른 단말기(200)는 통신부(210), 입력부(220), 표시부(230) 및 프로세서(240) 등을 포함할 수 있다. 도 3에 도시된 구성요소들은 본 개시에 따른 더빙 생성 시스템을 구현하는데 있어서 필수적인 것은 아니어서, 본 명세서 상에서 설명되는 단말기는 위에서 열거된 구성요소들 보다 많거나, 또는 적은 구성요소들을 가질 수 있다.
상기 구성요소들 중 통신부(210)는 외부 장치와 통신을 가능하게 하는 하나 이상의 구성 요소를 포함할 수 있으며, 예를 들어, 방송 수신 모듈, 유선통신 모듈, 무선통신 모듈, 근거리 통신 모듈, 위치정보 모듈 중 적어도 하나를 포함할 수 있다.
유선 통신 모듈은, 지역 통신(Local Area Network; LAN) 모듈, 광역 통신(Wide Area Network; WAN) 모듈 또는 부가가치 통신(Value Added Network; VAN) 모듈 등 다양한 유선 통신 모듈뿐만 아니라, USB(Universal Serial Bus), HDMI(High Definition Multimedia Interface), DVI(Digital Visual Interface), RS-1302(recommended standard1302), 전력선 통신, 또는 POTS(plain old telephone service) 등 다양한 케이블 통신 모듈을 포함할 수 있다.
무선 통신 모듈은 와이파이(Wifi) 모듈, 와이브로(Wireless broadband) 모듈 외에도, GSM(global System for Mobile Communication), CDMA(Code Division Multiple Access), WCDMA(Wideband Code Division Multiple Access), UMTS(universal mobile telecommunications system), TDMA(Time Division Multiple Access), LTE(Long Term Evolution), 4G, 5G, 6G 등 다양한 무선 통신 방식을 지원하는 무선 통신 모듈을 포함할 수 있다.
입력부(220)는 영상 정보(또는 신호), 오디오 정보(또는 신호), 데이터, 또는 사용자로부터 입력되는 정보의 입력을 위한 것으로서, 적어도 하나의 카메라, 적어도 하나의 마이크로폰 및 사용자 입력부 중 적어도 하나를 포함할 수 있다. 입력부에서 수집한 음성 데이터나 이미지 데이터는 분석되어 사용자의 제어명령으로 처리될 수 있다.
표시부(230)는 시각, 청각 또는 촉각 등과 관련된 출력을 발생시키기 위한 것으로, 디스플레이부, 음향 출력부, 햅팁 모듈 및 광 출력부 중 적어도 하나를 포함할 수 있다. 디스플레이부는 터치 센서와 상호 레이어 구조를 이루거나 일체형으로 형성됨으로써, 터치 스크린을 구현할 수 있다. 이러한 터치 스크린은, 본 장치와 사용자 사이의 입력 인터페이스를 제공하는 사용자 입력부로써 기능함과 동시에, 본 장치와 사용자 간에 출력 인터페이스를 제공할 수 있다.
디스플레이부는 본 장치에서 처리되는 정보를 표시(출력)한다. 예를 들어, 디스플레이부는 본 장치에서 구동되는 응용 프로그램(일 예로, 어플리케이션)의 실행화면 정보, 또는 이러한 실행화면 정보에 따른 UI(User Interface), GUI(Graphic User Interface) 정보를 표시할 수 있다.
상술한 구성요소 외에, 상술한 단말기는 인터페이스부 및 메모리를 더 포함할 수 있다.
인터페이스부는 본 장치에 연결되는 다양한 종류의 외부 기기와의 통로 역할을 수행한다. 이러한 인터페이스부는 유/무선 헤드셋 포트(port), 외부 충전기 포트(port), 유/무선 데이터 포트(port), 메모리 카드(memory card) 포트, 식별 모듈(SIM)이 구비된 장치를 연결하는 포트(port), 오디오 I/O(Input/Output) 포트(port), 비디오 I/O(Input/Output) 포트(port), 이어폰 포트(port) 중 적어도 하나를 포함할 수 있다. 본 장치에서는, 상기 인터페이스부에 연결된 외부 기기와 관련된 적절한 제어를 수행할 수 있다.
메모리는 본 장치의 다양한 기능을 지원하는 데이터와, 프로세서의 동작을 위한 프로그램을 저장할 수 있고, 입/출력되는 데이터들(예를 들어, 음악 파일, 정지영상, 동영상 등)을 저장할 있고, 본 장치에서 구동되는 다수의 응용 프로그램(application program 또는 애플리케이션(application)), 본 장치의 동작을 위한 데이터들, 명령어들을 저장할 수 있다. 이러한 응용 프로그램 중 적어도 일부는, 무선 통신을 통해 외부 서버로부터 다운로드 될 수 있다.
이러한, 메모리는 플래시 메모리 타입(flash memory type), 하드디스크 타입(hard disk type), SSD 타입(Solid State Disk type), SDD 타입(Silicon Disk Drive type), 멀티미디어 카드 마이크로 타입(multimedia card micro type), 카드 타입의 메모리(예를 들어 SD 또는 XD 메모리 등), 램(random access memory; RAM), SRAM(static random access memory), 롬(read-only memory; ROM), EEPROM(electrically erasable programmable read-only memory), PROM(programmable read-only memory), 자기 메모리, 자기 디스크 및 광디스크 중 적어도 하나의 타입의 저장매체를 포함할 수 있다. 또한, 메모리는 본 장치와는 분리되어 있으나, 유선 또는 무선으로 연결된 데이터베이스가 될 수도 있다.
한편, 상술한 단말기는 프로세서(240)를 포함한다. 프로세서는 본 장치 내의 구성요소들의 동작을 제어하기 위한 알고리즘 또는 알고리즘을 재현한 프로그램에 대한 데이터를 저장하는 메모리, 및 메모리에 저장된 데이터를 이용하여 전술한 동작을 수행하는 적어도 하나의 프로세서(미도시)로 구현될 수 있다. 이때, 메모리와 프로세서는 각각 별개의 칩으로 구현될 수 있다. 또는, 메모리와 프로세서는 단일 칩으로 구현될 수도 있다.
한편, 프로세서는 이하의 도면에서 설명되는 본 개시에 따른 다양한 실시 예들을 본 장치 상에서 구현하기 위하여, 위에서 살펴본 구성요소들을 중 어느 하나 또는 복수를 조합하여 제어할 수 있다.
한편, 도 1 내지 3에 도시된 구성 요소들의 성능에 대응하여 적어도 하나의 구성요소가 추가되거나 삭제될 수 있다. 또한, 구성 요소들의 상호 위치는 장치의 성능 또는 구조에 대응하여 변경될 수 있다는 것은 당해 기술 분야에서 통상의 지식을 가진 자에게 용이하게 이해될 것이다.
한편, 도 4와 같이, 서버 및 단말기 중 적어도 하나에 포함된 프로세서는 후술할 더빙 생성 시스템을 구현하기 위한 복수의 모듈을 포함할 수 있다. 구체적으로, 프로세서(300)는 음성 인식 모듈(310) 및 인공지능 모듈(320)을 포함할 수 있다. 후술하는 더빙 생성 방법은 상기 모듈들의 동작에 의해 구현되는 것으로 서술하나, 후술하는 각 단계의 동작의 수행이 반드시 상기 모듈들에 의해 수행될 필요는 없다.
이하에서는, 본 발명에서 서술되는 인공지능에 대하여 구체적으로 설명한다.
본 개시에 따른 인공지능과 관련된 기능은 상술한 서버 및 단말기에 탑재된 프로세서와 메모리를 통해 동작된다. 프로세서는 하나 또는 복수의 프로세서로 구성될 수 있다. 이때, 하나 또는 복수의 프로세서는 CPU, AP, DSP(Digital Signal Processor) 등과 같은 범용 프로세서, GPU, VPU(Vision Processing Unit)와 같은 그래픽 전용 프로세서 또는 NPU와 같은 인공지능 전용 프로세서일 수 있다. 하나 또는 복수의 프로세서는, 메모리에 저장된 기 정의된 동작 규칙 또는 인공지능 모델에 따라, 입력 데이터를 처리하도록 제어한다. 또는, 하나 또는 복수의 프로세서가 인공지능 전용 프로세서인 경우, 인공지능 전용 프로세서는, 특정 인공지능 모델의 처리에 특화된 하드웨어 구조로 설계될 수 있다.
기 정의된 동작 규칙 또는 인공지능 모델은 학습을 통해 만들어진 것을 특징으로 한다. 여기서, 학습을 통해 만들어진다는 것은, 기본 인공지능 모델이 학습 알고리즘에 의하여 다수의 학습 데이터들을 이용하여 학습됨으로써, 원하는 특성(또는, 목적)을 수행하도록 설정된 기 정의된 동작 규칙 또는 인공지능 모델이 만들어짐을 의미한다. 이러한 학습은 본 개시에 따른 인공지능이 수행되는 기기 자체에서 이루어질 수도 있고, 별도의 서버 및/ 또는 시스템을 통해 이루어 질 수도 있다. 학습 알고리즘의 예로는, 지도형 학습(supervised learning), 비지도 형 학습(unsupervised learning), 준지도형 학습(semi-supervised learning) 또는 강화 학습(reinforcement learning)이 있으나, 전술한 예에 한정되지 않는다.
인공지능 모델은, 복수의 신경망 레이어들로 구성될 수 있다. 복수의 신경망 레이어들 각각은 복수의 가중치들 (weight values)을 갖고 있으며, 이전(previous) 레이어의 연산 결과와 복수의 가중치들 간의 연산을 통해 신경 망 연산을 수행한다. 복수의 신경망 레이어들이 갖고 있는 복수의 가중치들은 인공지능 모델의 학습 결과에 의해 최적화될 수 있다. 예를 들어, 학습 과정 동안 인공지능 모델에서 획득한 로스(loss) 값 또는 코스트(cost) 값이 감소 또는 최소화되도록 복수의 가중치들이 갱신될 수 있다. 인공 신경망은 심층 신경망(DNN:Deep Neural Network)를 포함할 수 있으며, 예를 들어, CNN (Convolutional Neural Network), DNN (Deep Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), BRDNN(Bidirectional Recurrent Deep Neural Network) 또는 심층 Q-네트워크 (Deep Q-Networks) 등이 있으나, 전술한 예에 한정되지 않는다.
본 개시의 예시적인 실시예에 따르면, 프로세서는 인공지능을 구현할 수 있다. 인공지능이란 사람의 신경세포(biological neuron)를 모사하여 기계가 학습하도록 하는 인공신경망(Artificial Neural Network) 기반의 기계 학습법을 의미한다. 인공지능의 방법론에는 학습 방식에 따라 훈련데이터로서 입력데이터와 출력데이터가 같이 제공됨으로써 문제(입력데이터)의 해답(출력데이터)이 정해져 있는 지도학습(supervised learning), 및 출력데이터 없이 입력데이터만 제공되어 문제(입력데이터)의 해답(출력데이터)이 정해지지 않는 비지도학습(unsupervised learning), 및 현재의 상태(State)에서 어떤 행동(Action)을 취할 때마다 외부 환경에서 보상(Reward)이 주어지는데, 이러한 보상을 최대화하는 방향으로 학습을 진행하는 강화학습(reinforcement learning)으로 구분될 수 있다. 또한, 인공지능의 방법론은 학습 모델의 구조인 아키텍처에 따라 구분될 수도 있는데, 널리 이용되는 딥러닝 기술의 아키텍처는, 합성곱신경망(CNN; Convolutional Neural Network), 순환신경망(RNN; Recurrent Neural Network), 트랜스포머(Transformer), 생성적 대립 신경망(GAN; generative adversarial networks) 등으로 구분될 수 있다.
본 장치와 시스템은 인공지능 모델을 포함할 수 있다. 인공지능 모델은 하나의 인공지능 모델일 수 있고, 복수의 인공지능 모델로 구현될 수도 있다. 인공지능 모델은 뉴럴 네트워크(또는 인공 신경망)로 구성될 수 있으며, 기계학습과 인지과학에서 생물학의 신경을 모방한 통계학적 학습 알고리즘을 포함할 수 있다. 뉴럴 네트워크는 시냅스의 결합으로 네트워크를 형성한 인공 뉴런(노드)이 학습을 통해 시냅스의 결합 세기를 변화시켜, 문제 해결 능력을 가지는 모델 전반을 의미할 수 있다. 뉴럴 네트워크의 뉴런은 가중치 또는 바이어스의 조합을 포함할 수 있다. 뉴럴 네트워크는 하나 이상의 뉴런 또는 노드로 구성된 하나 이상의 레이어(layer)를 포함할 수 있다. 예시적으로, 장치는 input layer, hidden layer, output layer를 포함할 수 있다. 장치를 구성하는 뉴럴 네트워크는 뉴런의 가중치를 학습을 통해 변화시킴으로써 임의의 입력(input)으로부터 예측하고자 하는 결과(output)를 추론할 수 있다.
프로세서는 뉴럴 네트워크를 생성하거나, 뉴럴 네트워크를 훈련(train, 또는 학습(learn)하거나, 수신되는 입력 데이터를 기초로 연산을 수행하고, 수행 결과를 기초로 정보 신호(information signal)를 생성하거나, 뉴럴 네트워크를 재훈련(retrain)할 수 있다. 뉴럴 네트워크의 모델들은 GoogleNet, AlexNet, VGG Network 등과 같은 CNN(Convolution Neural Network), R-CNN(Region with Convolution Neural Network), RPN(Region Proposal Network), RNN(Recurrent Neural Network), S-DNN(Stacking-based deep Neural Network), S-SDNN(State-Space Dynamic Neural Network), Deconvolution Network, DBN(Deep Belief Network), RBM(Restrcted Boltzman Machine), Fully Convolutional Network, LSTM(Long Short-Term Memory) Network, Classification Network 등 다양한 종류의 모델들을 포함할 수 있으나 이에 제한되지는 않는다. 프로세서는 뉴럴 네트워크의 모델들에 따른 연산을 수행하기 위한 하나 이상의 프로세서를 포함할 수 있다. 예를 들어 뉴럴 네트워크는 심층 뉴럴 네트워크 (Deep Neural Network)를 포함할 수 있다.
뉴럴 네트워크는 CNN(Convolutional Neural Network), RNN(Recurrent Neural Network), 퍼셉트론(perceptron), 다층 퍼셉트론(multilayer perceptron), FF(Feed Forward), RBF(Radial Basis Network), DFF(Deep Feed Forward), LSTM(Long Short Term Memory), GRU(Gated Recurrent Unit), AE(Auto Encoder), VAE(Variational Auto Encoder), DAE(Denoising Auto Encoder), SAE(Sparse Auto Encoder), MC(Markov Chain), HN(Hopfield Network), BM(Boltzmann Machine), RBM(Restricted Boltzmann Machine), DBN(Depp Belief Network), DCN(Deep Convolutional Network), DN(Deconvolutional Network), DCIGN(Deep Convolutional Inverse Graphics Network), GAN(Generative Adversarial Network), LSM(Liquid State Machine), ELM(Extreme Learning Machine), ESN(Echo State Network), DRN(Deep Residual Network), DNC(Differentiable Neural Computer), NTM(Neural Turning Machine), CN(Capsule Network), KN(Kohonen Network) 및 AN(Attention Network)를 포함할 수 있으나 이에 한정되는 것이 아닌 임의의 뉴럴 네트워크를 포함할 수 있음은 통상의 기술자가 이해할 것이다.
본 개시의 예시적인 실시예에 따르면, 프로세서는 GoogleNet, AlexNet, VGG Network 등과 같은 CNN(Convolution Neural Network), R-CNN(Region with Convolution Neural Network), RPN(Region Proposal Network), RNN(Recurrent Neural Network), S-DNN(Stacking-based deep Neural Network), S-SDNN(State-Space Dynamic Neural Network), Deconvolution Network, DBN(Deep Belief Network), RBM(Restrcted Boltzman Machine), Fully Convolutional Network, LSTM(Long Short-Term Memory) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, 자연어 처리를 위한 BERT, SP-BERT, MRC/QA, Text Analysis, Dialog System, GPT-3, GPT-4, 비전 처리를 위한 Visual Analytics, Visual Understanding, Video Synthesis, ResNet 데이터 지능을 위한 Anomaly Detection, Prediction, Time-Series Forecasting, Optimization, Recommendation, Data Creation 등 다양한 인공지능 구조 및 알고리즘을 이용할 수 있으며, 이에 제한되지 않는다. 이하, 첨부된 도면을 참조하여 본 개시의 실시예를 상세하게 설명한다.
이하에서는, 본 개시에 따른 더빙 생성 시스템 및 방법이 제공되는 자막 및 더빙 제공 서비스에 대하여 설명한다.
도 5는 자막 및 더빙 제공 서비스를 나타내는 개념도이다.
도 5를 참조하면, 음성 정보를 입력 받는 단계가 진행된다. 여기서음성 정보를 입력 받는 단계는, 영상 정보를 입력 받는 단계일 수 있다. 본 개시에 따른 자믹 더빙 생성 시스템은 영상 정보를 입력받고, 영상 정보에 포함된 음성 정보를 더빙 생성에 활용한다. 이와 달리, 상기 음성 정보를 입력 받는 단계는 영상 정보가 아니라 음성 정보 자체를 입력 받는 단계일 수 있다.
본 명세서에서 음성 정보는 영상에 포함된 소리와 관련된 모든 데이터를 포함할 수 있다. 일 실시예로, 음성 정보는 영상 정보에 포함된 음성 신호 및 배경음을 포함할 수 있다.
프로세서는 영상 정보를 입력 받아 영상 정보에 포함된 음성 정보를 가공한다.
다음으로, 음성 정보에서 음성 신호 및 배경음 정보를 분리하는 단계가 진행된다.
프로세서는 음성 정보에서 음성 신호와 배경음 정보를 분리하여 별도의 파일을 생성할 수 있다.
음성 신호는 영상 정보에 포함된 화자의 음성이며, 더빙 생성의 대상이되는 정보이다. 추후, 음성 정보에서 분리된 음성 신호는 더빙 중 적어도 하나로 생성된다.
상기 음성 신호와 배경음 정보의 분리는 음성 정보로부터 음성 신호 및 배경음 정보를 분리하도록 훈련된 별도의 인공지능 모델을 통해 수행될 수 있다.
다음으로, 분리된 음성 신호를 텍스트로 변환하는 단계가 진행된다.
상기 음성 신호는 상기 음성 인식 모듈(310)을 통해 텍스트로 변환될 수 있다. 여기서, 상기 텍스트는 상기 음성 신호를 분리한 영상 제작의 기반이 된 언어로 이루어질 수 있다.
본 명세서에서 영상 제작의 기반이 된 언어를 '원어'라하며, 더빙을 생성하고자하는 언어를 "타겟 언어"라 한다.
다음으로, 변환된, 원어로 구성된 텍스트를 타겟 언어로 번역하는 단계가 진행된다.
상기 번역 결과는 타겟 언어로 구성된 텍스트 형태이며, 원본 영상의 자막으로 제공될 수 있다.
상기 타겟 언어의 종류는 영상 제작자 또는 영상을 시청하는 사용자로부터 입력 받을 수 있다.
일 실시 예에 있어서, 영상 제작자는 상기 타겟 언어를 미리 지정하여 영상 제작 단계에서부터 적어도 하나의 타겟 언어에 대한 자막을 생성할 수 있다.
일 실시 예에 있어서, 영상 제작자는 자막을 생성하지 않고, 영상을 시청하는 시청자가 원하는 타겟 언어를 지정하여 자막을 생성할 수 있다. 특정, 영상 시청자에 의해 제작된 자막은 해당 타겟 언어의 자막을 요청하는 영상 시청자에게 다시 제공될 수 있다.
상기 번역 결과는 이후, 더빙 생성에 활용될 수 있다. 이에 대하여는 후술한다.
한편, 상기 번역은 원어로 이루어진 텍스트를 타겟 언어로 번역하도록 훈련된 별도의 인공지능 모델을 통해 수행될 수 있다. 여기서, 번역을 위한 인공지능 모델의 종류는 별도로 한정하지 않는다.
한편, 프로세서는 상기 자막을 생성한 후 원본 영상에 합성한다. 이 때, 프로세서는 영상 정보 내에서 음성 신호가 발생되는 시간 구간을 고려하여, 자막을 삽입할 시간 구간을 설정할 수 있다. 구체적으로, 프로세서는 음절, 어절 또는 문장 단위로 영상 정보내에서 음성 신호가 발생된 시간 구간을 매칭시킬 수 있다. 이러한 방식으로 매칭된 시간 정보는 음성 신호를 텍스트로 변환한 후, 변환된 텍스트에도 매칭된다. 이후, 상기 변환된 텍스트를 번역한 결과에도 상기 시간 정보가 매칭될 수 있다. 프로세서는 상기 시간 정보를 기반으로 자막을 상기 영상 정보에 자막을 삽입할 수 있다.
다음으로, 분리된 음성 신호, 변환된 텍스트 및 번역 결과 중 적어도 하나를 기반으로 타겟 언어로 이루어진 음성 신호(이하, 더빙 정보)를 생성하는 단계가 진행된다.
본 개시에 따른 더빙 생성 방법은 상술한 음성 신호를 생성하는 단계에 구현될 수 있다.
이하, 도 6 및 7을 참조하여, 타겟 언어로 이루어진 더빙 정보를 생성하는 방법에 대하여 구체적으로 설명한다.
도 6을 참조하면, 인공지능 모듈(320)이 음성 신호를 입력 받는 단계가 진행된다(S210).
여기서, 인공지능 모듈(320)이 입력받는 정보는 영상 정보에 포함된 음성 정보이거나, 상기 음성 정보에서 배경음 정보가 분리된 음성 신호일 수 있다.
다음으로, 분리된 음성 신호를 텍스트로 변환하는 단계가 진행된다(S221).
상기 음성 신호는 상기 음성 인식 모듈(310)을 통해 텍스트로 변환될 수 있다. 여기서, 상기 텍스트는 상기 원어로 이루어질 수 있다.
다음으로, 변환된 텍스트를 타겟 언어로 번역하는 단계가 진행된다(S222).
원어로 구성된 상기 변환된 텍스트는 번역을 통해, 타겟 언어로 구성된 텍스트로 변환된다.
상기 번역은 기 공지된 인공 지능 모델을 통해 수행된거나 별도로 훈련된 인공 지능 모델을 통해 수행될 수 있다.
상기 S221 및 S222 단계는 도 5에서 설명한 자막 및 더빙 제공 서비스에서 설명한 번역 결과를 활용함으로써 생략될 수 있다.
상기 텍스트 번역과는 별개로, 다음으로, 인공지능 모델이 음성 신호를 국제음성기호(IPA)로 구성된 텍스트로 변환하는 단계가 진행된다(S223).
상기 인공지능 모듈(320)은 음성 정보를 국제음성기호(International Phonetic Association, IPA)로 변환하도록 훈련된 인공지능 모델을 포함할 수 있다.
일 실시 예에 있어서, 상기 인공지능 모델은 트랜스포머(transformer)일 수 있으나, 이에 한정되지는 않는다.
일 실시 예에 있어서, 도 7을 참조하면, 상기 인공지능 모델은 "만 원만 빌려주시면 안될까요"라는 음성 정보를 입력받아 아래와 같은 국제음성기호로 변환할 수 있다.
Figure PCTKR2024017832-appb-img-000001
구체적으로, 상기 인공지능 모델은 'src'가 한국어로 구성된 음성 신호이고, 'tgt'가 영문 IPA embedding이고, 'output'으로 영문 IPA embedding이 되도록 훈련될 수 있다.
이후, 상기 인공지능 모델은 상기 번역 결과 및 국제음성기호로 구성된 텍스트를 기반으로 더빙 정보를 생성한다(S230).
구체적으로, 인공지능 모델은 상기 변환된 텍스트 및 상기 음성 신호에 대한 번역 결과를 기반으로, 상기 음성 정보에 대한 더빙 정보를 생성할 수 있다.
인공지능 모델은 상기 변환된 텍스트를 기반으로, 상기 번역 결과를 국제음성기호로 구성된 정보로 변환한다. 이를 위해, 인공지능 모델은 원어로 이루어진 음성 신호를 기반으로 생성된, 국제음성기호로 구성된 텍스트와 상기 원어로 이루어진 음성 신호에 대한 번역 결과를 입력 받아 국제음성기호로 구성되며, 타겟 언어로 번역된 텍스트를 생성하도록 훈련될 수 있다. 훈련 과정에서, 국제음성기호를 활용하는 경우, 원어 및 타겟 언어의 지역적, 사회적, 문화적, 언어적 차이가 반영된 더빙 정보를 생성할 수 있게 된다.
또한, 상술한 바와 같이, 더빙 정보 생성을 위한 인공지능 모델 훈련 시 국제음성기호를 활용하는 경우, 자연스럽게 타겟 언어를 발음하는 더빙 정보를 생성할 수 있게 된다.
다시, 도 5를 참조하면, 생성된 음성 신호(더빙 정보)와 분리된 배경음 정보를 합성하는 단계가 진행된다.
프로세서는 상기 번역 결과를 활용하여, 상기 영상 정보에 자막을 생성하고, 상기 자막 생성 결과를 활용하여, 상기 생성된 음성 신호와 상기 분리된 배경음 정보를 합성할 수 있다.
구체적으로, 프로세서는 상기 영상 정보에서 상기 자막이 출력되는 시간 정보를 탐색하고, 탐색된 시간 정보를 활용하여, 상기 생성된 음성 신호와 상기 분리된 배경음 정보를 합성할 수 있다.
프로세서는 음절, 어절 또는 문장 단위로 영상 정보내에서 음성 신호가 발생된 시간 구간을 매칭시킬 수 있다. 이러한 방식으로 매칭된 시간 정보는 음성 신호를 텍스트로 변환한 후, 변환된 텍스트에도 매칭된다. 이후, 상기 변환된 텍스트를 번역한 결과에도 상기 시간 정보가 매칭될 수 있다. 프로세서는 상기 시간 정보를 기반으로 자막을 상기 영상 정보에 자막을 삽입할 수 있다. 또한, 프로세서는 상술한 시간 정보를 더빙 생성에도 활용할 수 있다.
한편, 프로세서는 영상 정보에 포함된 음성 정보에서 상기 음성 신호 및 상기 배경음 정보를 분리한 결과를 활용하여, 상기 생성된 음성 신호와 상기 분리된 배경음 정보를 합성할 수 있다. 구체적으로, 프로세서는 상기 음성 정보에서 상기 음성 신호가 분리된 시간 정보를 활용하여, 상기 생성된 음성 신호와 상기 분리된 배경음 정보를 합성할 수 있다.
프로세서는 음성 신호를 배경음 정보에서 분리할 때, 음성 신호가 분리된 시간 구간을 정의하는 정보를 배경음 정보에 매칭시킬 수 있다. 이러한 시간 구간 정보는 음절, 어절 또는 문장 단위로 정의되어, 배경음 정보에 매칭될 수 있다. 상기 시간 정보는 배경음 정보에 더빙 정보를 합성할 때 활용될 수 있다.
상술한 바와 같이, 본 개시에 따른 더빙 생성 시스템에 따르면, 더빙을 생성한 후 음성 정보에서 분리된 배경음을 합성하기 때문에, 음원 손실 없이 더빙을 생성할 수 있게 된다.
또한, 본 개시에 따르면, 수작업으로 더빙을 생성하는 종래의 방법 대비 빠른 속도로 더빙을 생성할 수 있다.
또한, 본 개시에 따르면, 지역적, 사회적, 문화적, 언어적 차이가 반영된 더빙을 생성할 수 있게 된다.
한편, 개시된 실시예들은 컴퓨터에 의해 실행 가능한 명령어를 저장하는 기록매체의 형태로 구현될 수 있다. 명령어는 프로그램 코드의 형태로 저장될 수 있으며, 프로세서에 의해 실행되었을 때, 프로그램 모듈을 생성하여 개시된 실시예들의 동작을 수행할 수 있다. 기록매체는 컴퓨터로 읽을 수 있는 기록매체로 구현될 수 있다.
컴퓨터가 읽을 수 있는 기록매체로는 컴퓨터에 의하여 해독될 수 있는 명령어가 저장된 모든 종류의 기록 매체를 포함한다. 예를 들어, ROM(Read Only Memory), RAM(Random Access Memory), 자기 테이프, 자기 디스크, 플래쉬 메모리, 광 데이터 저장장치 등이 있을 수 있다.
이상에서와 같이 첨부된 도면을 참조하여 개시된 실시예들을 설명하였다. 본 개시가 속하는 기술분야에서 통상의 지식을 가진 자는 본 개시의 기술적 사상이나 필수적인 특징을 변경하지 않고도, 개시된 실시예들과 다른 형태로 본 개시가 실시될 수 있음을 이해할 것이다. 개시된 실시예들은 예시적인 것이며, 한정적으로 해석되어서는 안 된다.

Claims (10)

  1. 서버 및 단말기를 포함하는 더빙 생성 시스템에 있어서,
    상기 서버는,
    상기 단말기로부터 영상 정보를 수신하도록 이루어지는 통신부;
    상기 영상 정보에 포함된 음성 신호를 텍스트로 변환하고, 변환된 텍스트를 기반으로 더빙 정보를 생성하는 인공지능 모델을 포함하는 프로세서를 포함하고,
    상기 텍스트는 국제음성기호로 구성되는 것을 특징으로 하는 더빙 생성 시스템.
  2. 제1항에 있어서,
    상기 프로세서는,
    상기 변환된 텍스트 및 상기 음성 신호에 대한 번역 결과를 기반으로, 상기 음성 정보에 대한 더빙 정보를 생성하는 것을 특징으로 하는 더빙 생성 시스템.
  3. 제2항에 있어서,
    상기 프로세서는,
    상기 변환된 텍스트를 기반으로, 상기 프로세서는 상기 번역 결과를 국제음성기호로 구성된 정보로 변환하고, 변환된 정보를 기반으로 더빙 정보를 생성하는 것을 특징으로 하는 더빙 생성 시스템.
  4. 제3항에 있어서,
    상기 프로세서는,
    상기 영상 정보에서 상기 음성 신호를 제거하고, 상기 더빙 정보를 결합하는 것을 특징으로 하는 더빙 생성 시스템.
  5. 제4항에 있어서,
    상기 인공지능 모델은 트랜스포머(transformer)인 것을 특징으로 하는 더빙 생성 시스템.
  6. 서버 및 단말기를 포함하는 더빙 생성 시스템의 더빙 생성 방법에 있어서,
    상기 서버가 상기 단말기로부터 영상 정보를 수신하는 단계;
    상기 서버가 상기 영상 정보에 포함된 음성 신호를 텍스트로 변환하는 단계; 및
    상기 서버가 상기 변환된 텍스트를 기반으로 더빙 정보를 생성하는 단계를 포함하고,
    상기 텍스트는 국제음성기호로 구성되는 것을 특징으로 하는 더빙 생성 방법.
  7. 제6항에 있어서,
    상기 더빙 정보를 생성하는 단계는,
    상기 변환된 텍스트 및 상기 음성 신호에 대한 번역 결과를 기반으로, 상기 음성 정보에 대한 더빙 정보를 생성하는 것을 특징으로 하는 더빙 생성 방법.
  8. 제7항에 있어서,
    상기 더빙 정보를 생성하는 단계는,
    상기 변환된 텍스트를 기반으로, 상기 프로세서는 상기 번역 결과를 국제음성기호로 구성된 정보로 변환하고, 변환된 정보를 기반으로 더빙 정보를 생성하는 것을 특징으로 하는 더빙 생성 방법.
  9. 제8항에 있어서,
    상기 영상 정보에서 상기 음성 신호를 제거하고, 상기 더빙 정보를 결합하는 단계를 더 포함하는 것을 특징으로 하는 더빙 생성 방법.
  10. 제9항에 있어서,
    상기 인공지능 모델은 트랜스포머(transformer)인 것을 특징으로 하는 더빙 생성 방법.
PCT/KR2024/017832 2023-11-24 2024-11-12 인공지능을 이용하여 언어학 및 인지과학 기반의 음성기술이 적용된 더빙 생성 시스템 및 방법 Pending WO2025110602A1 (ko)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
KR1020230165413A KR20250077893A (ko) 2023-11-24 2023-11-24 인공지능을 이용하여 언어학 및 인지과학 기반의 음성기술이 적용된 더빙 생성 시스템 및 방법
KR10-2023-0165413 2023-11-24

Publications (1)

Publication Number Publication Date
WO2025110602A1 true WO2025110602A1 (ko) 2025-05-30

Family

ID=95826820

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2024/017832 Pending WO2025110602A1 (ko) 2023-11-24 2024-11-12 인공지능을 이용하여 언어학 및 인지과학 기반의 음성기술이 적용된 더빙 생성 시스템 및 방법

Country Status (2)

Country Link
KR (1) KR20250077893A (ko)
WO (1) WO2025110602A1 (ko)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2826215B2 (ja) * 1990-10-16 1998-11-18 インターナショナル・ビジネス・マシーンズ・コーポレイション 合成音声生成方法及びテキスト音声合成装置
KR101735195B1 (ko) * 2015-08-07 2017-05-12 네이버 주식회사 운율 정보 기반의 자소열 음소열 변환 방법과 시스템 그리고 기록 매체
KR102168529B1 (ko) * 2020-05-29 2020-10-22 주식회사 수퍼톤 인공신경망을 이용한 가창음성 합성 방법 및 장치
KR20220065483A (ko) * 2020-11-13 2022-05-20 주식회사 에스알유니버스 단일음성기호집합을 활용한 인공신경망 기반 다국어 발화 텍스트 음성합성 방법 및 장치
KR102546559B1 (ko) * 2022-03-14 2023-06-26 주식회사 엘젠 영상 콘텐츠 자동 번역 더빙 시스템

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2826215B2 (ja) * 1990-10-16 1998-11-18 インターナショナル・ビジネス・マシーンズ・コーポレイション 合成音声生成方法及びテキスト音声合成装置
KR101735195B1 (ko) * 2015-08-07 2017-05-12 네이버 주식회사 운율 정보 기반의 자소열 음소열 변환 방법과 시스템 그리고 기록 매체
KR102168529B1 (ko) * 2020-05-29 2020-10-22 주식회사 수퍼톤 인공신경망을 이용한 가창음성 합성 방법 및 장치
KR20220065483A (ko) * 2020-11-13 2022-05-20 주식회사 에스알유니버스 단일음성기호집합을 활용한 인공신경망 기반 다국어 발화 텍스트 음성합성 방법 및 장치
KR102546559B1 (ko) * 2022-03-14 2023-06-26 주식회사 엘젠 영상 콘텐츠 자동 번역 더빙 시스템

Also Published As

Publication number Publication date
KR20250077893A (ko) 2025-06-02

Similar Documents

Publication Publication Date Title
US12131586B2 (en) Methods, systems, and machine-readable media for translating sign language content into word content and vice versa
CN114822495B (zh) 声学模型训练方法、装置及语音合成方法
WO2021137657A1 (en) Method and apparatus for personalizing content recommendation model
CN114882862B (zh) 一种语音处理方法及相关设备
CN113948060B (zh) 一种网络训练方法、数据处理方法及相关设备
KR102544249B1 (ko) 발화의 문맥을 공유하여 번역을 수행하는 전자 장치 및 그 동작 방법
WO2023080425A1 (ko) 쿼리문에 관련된 검색 결과를 제공하는 전자 장치 및 방법
WO2022231126A1 (ko) 전자 장치 및 전자 장치의 프로소디 제어를 위한 tts 모델 생성 방법
CN118246537A (zh) 基于大模型的问答方法、装置、设备及存储介质
WO2022092440A1 (ko) 전자 장치 및 그 제어 방법
CN120126459B (zh) 语音识别及模型训练方法、装置、设备及计算机程序产品
WO2025110602A1 (ko) 인공지능을 이용하여 언어학 및 인지과학 기반의 음성기술이 적용된 더빙 생성 시스템 및 방법
WO2025110601A1 (ko) 인공지능 기반 음성기술이 적용된 자막 및 더빙 생성 시스템 및 방법
WO2025079898A1 (ko) 텍스트 태깅된 모션 생성 장치 및 모션 생성 장치의 동작 방법
KR102705393B1 (ko) 문장 내 상태 정보를 이용한 음성 인식 후처리를 수행하는 전자 장치 및 그의 학습 방법
WO2025084457A2 (ko) 생성형 ai 기반의 애플리케이션 자동 제작 서버, 방법 및 프로그램
EP4281848A1 (en) Methods and systems for generating one or more emoticons for one or more users
WO2024029875A1 (ko) 전자 장치, 지능형 서버, 및 화자 적응형 음성 인식 방법
WO2022260485A1 (en) Methods and systems for generating one or more emoticons for one or more users
WO2025143597A1 (ko) 프롬프트 튜닝을 수행하기 위한 전자 장치 및 그 제어 방법
JP7829266B1 (ja) 情報処理システム、情報処理方法及びプログラム
WO2025254252A1 (ko) 스타일 기반의 모션 변환 장치 및 그 방법
Garg et al. Real-time conversion for sign-to-text and text-to-speech communication using machine learning
WO2025121794A1 (ko) 프롬프팅을 위한 전자 장치 및 이의 동작 방법
WO2025143331A1 (ko) 객체 인식을 위한 확률적 비 최대 억제(nms) 방법 및 이를 적용한 장치

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24894495

Country of ref document: EP

Kind code of ref document: A1