WO2021211366A1 - Self-supervised ai-assisted sound effect recommendation for silent video - Google Patents
Self-supervised ai-assisted sound effect recommendation for silent video Download PDFInfo
- Publication number
- WO2021211366A1 WO2021211366A1 PCT/US2021/026550 US2021026550W WO2021211366A1 WO 2021211366 A1 WO2021211366 A1 WO 2021211366A1 US 2021026550 W US2021026550 W US 2021026550W WO 2021211366 A1 WO2021211366 A1 WO 2021211366A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- audio
- positive
- signal
- negative
- visual
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/60—Information retrieval; Database structures therefor; File system structures therefor of audio data
- G06F16/63—Querying
- G06F16/635—Filtering based on additional data, e.g. user or group profiles
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/60—Information retrieval; Database structures therefor; File system structures therefor of audio data
- G06F16/68—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
- G06N3/0442—Recurrent networks, e.g. Hopfield networks characterised by memory or gating, e.g. long short-term memory [LSTM] or gated recurrent units [GRU]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/084—Backpropagation, e.g. using gradient descent
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/0895—Weakly supervised learning, e.g. semi-supervised or self-supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/16—Speech classification or search using artificial neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
Definitions
- the present disclosure relates to sound effect selection for media, specifically aspects of the present disclosure relate to using machine-learning techniques for sound selection in media.
- FIG. 1 A is a simplified diagram of a convolutional neural network for use in a Sound Effect Recommendation Tool according to aspects of the present disclosure.
- FIG. IB is a simplified node diagram of a recurrent neural network for use in a Sound Effect Recommendation Tool according to aspects of the present disclosure.
- FIG. 1C is a simplified node diagram of an unfolded recurrent neural network for use in a Sound Effect Recommendation Tool according to aspects of the present disclosure.
- FIG. ID is a block diagram of a method for training a neural network in development of a Sound Effect Recommendation Tool according to aspects of the present disclosure.
- FIG. 2A is a block diagram depicting a method for training an audio-visual correlation NN using visual input paired with audio containing a noisy mixture of sound sources, for use in the Sound Recommendation tool, according to aspects of the present disclosure.
- FIG. 2B is a block diagram depicting a method that first maps an audio containing a mixture of sound sources into individual sound sources, which are then paired with the visual input for training an audio-visual correlation NN for use in the Sound Recommendation tool, according to aspects of the present disclosure.
- FIG. 3 is a block diagram depicting training of an audio-visual Correlation NN that learns positive and negative correlations simultaneously using triplet inputs containing a visual input, positive correlated audio, and negative uncorrelated audio, for use in the Sound Recommendation tool according to aspects of the present disclosure.
- FIG. 4 is a block diagram showing the training of a NN for learning fine-grained audio-visual correlations based on audio containing a mixture of sound sources, for use in the Sound Recommendation Tool according to aspects of the present disclosure.
- FIG. 5 is a block diagram that depicts a method of using the trained NN in a Sound Effect Recommendation tool for creating a new video with sound, according to aspects of the present disclosure.
- FIG. 6 is a block system diagram depicting a system implementing the training of neural networks and use of the Sound Effect Recommendation Tool according to aspects of the present disclosure.
- DESCRIPTION OF THE SPECIFIC EMBODIMENTS Although the following detailed description contains many specific details for the purposes of illustration, anyone of ordinary skill in the art will appreciate that many variations and alterations to the following details are within the scope of the invention. Accordingly, the exemplary embodiments of the invention described below are set forth without any loss of generality to, and without imposing limitations upon, the claimed invention.
- Neural Networks and machine learning may be applied to sound design to choose appropriate sounds for video sequences that lack sound.
- Three techniques for developing a Sound Effect Recommendation Tool will be discussed herein. First general NN training methods will be discussed. Second, a method will be discussed for training a coarse-grained correlation NN for prediction of sound effects based on a reference video, directly from an audio mixture as well as by mapping the audio mixture to single audio sources using a similarity NN. The third method that will be discussed is for training a fine-grained correlation NN for recommending sound effects based on a reference video. Finally, use of a tool employing the trained Sound Effect Recommendation Networks individually or as a combination will be discussed.
- the Sound Effect Recommendation Tool may include one or more of several different types of neural networks and may have many different layers.
- a classification neural network may consist of one or multiple deep neural networks (DNN), such as convolutional neural networks (CNN) and/or recurrent neural networks (RNN).
- DNN deep neural networks
- CNN convolutional neural networks
- RNN recurrent neural networks
- the Sound Effect Recommendation Tool may be trained using the general training method disclosed herein.
- FIG. 1A depicts an example layout of a convolution neural network according to aspects of the present disclosure.
- the convolution neural network is generated for an input 132 with a size of 4 units in height and 4 units in width giving a total area of 16 units.
- the depicted convolutional neural network has a filter 133 size of 2 units in height and 2 units in width with a stride value of 1 and a channel 136 of size 9.
- the convolutional neural network may have any number of additional neural network node layers 131 and may include such layer types as additional convolutional layers, fully connected layers, pooling layers, max pooling layers, normalization layers, etc. of any size.
- FIG. IB depicts the basic form of an RNN having a layer of nodes 120, each of which is characterized by an activation function S, input U, a recurrent node weight W, and an output V.
- the activation function S is typically a non-linear function known in the art and is not limited to the (hyperbolic tangent (tanh) function.
- the activation function S may be a Sigmoid or ReLU function. As shown in FIG.
- the RNN may be considered as a series of nodes 120 having the same activation function with the value of the activation function S moving through time from SO prior to T to SI after T and S2 after T+l.
- the nodes in a layer of RNN apply the same set of activation functions and weights to a series of inputs.
- the output of each node depends not just on the activation function and weights applied on that node’s input, but also on that node’s previous context.
- the RNN uses historical information by feeding the result from a previous time T to a current time T+l.
- a convolutional RNN may be used, especially when the visual input is a video.
- Another type of RNN that may be used is a Long Short-Term Memory (LSTM) Neural Network which adds a memory block in a RNN node with input gate activation function, output gate activation function and forget gate activation function resulting in a gating memory that allows the network to retain some information for a longer period of time as described by Hochreiter & Schmidhuber “Long Short-term memory” Neural Computation 9(8): 1735-1780 (1997), which is incorporated herein by reference.
- LSTM Long Short-Term Memory
- a neural network begins with initialization of the weights of the NN at 141.
- the initial weights should be distributed randomly.
- an NN with a tanh activation function should have random values distributed between — 1 and 4 where n is the number of inputs to the node. n
- the NN is then provided with a feature vector or input dataset at 142.
- Each of the different feature vectors may be generated by the NN from inputs that have known relationships.
- the NN may be provided with feature vectors that correspond to inputs having known relationships.
- the NN then predicts a distance between the features or inputs at 143. The predicted distance is compared to the known relationship (also known as ground truth) and a loss function measures the total error between the predictions and ground truth over all the training samples at 144.
- the loss function may be a cross entropy loss function, quadratic cost, triplet contrastive function, exponential cost, mean square error etc. Multiple different loss functions may be used depending on the purpose.
- a cross entropy loss function may be used whereas for learning an embedding a triplet contrastive loss function may be employed.
- the NN is then optimized and trained, using known methods of training for neural networks such as backpropagating the result of the loss function and by using optimizers, such as stochastic and adaptive gradient descent etc., as indicated at 145.
- the optimizer tries to choose the model parameters (i.e., weights) that minimize the training loss function (i.e. total error).
- Data is partitioned into training, validation, and test samples.
- the Optimizer minimizes the loss function on the training samples. After each training epoch, the model is evaluated on the validation sample by computing the validation loss and accuracy. If there is no significant change, training can be stopped and the most optimal model resulting from the training may be used to predict the labels or relationships for the test data.
- the neural network may be trained from inputs having known relationships to group related inputs.
- a NN may be trained using the described method to generate a feature vector from inputs having known relationships.
- the automated methods for recommending sound effects for visual scenes is based on learning audio-visual correlations by training on a large number of example videos (such as video games or movie clips), with one or more sound sources mixed together.
- One method to generate training data in order to train a model that learns audio-visual relationships is to generate audio-visual segment pairs from videos with labeled sound sources.
- manually detecting and labeling each sound source to create a large training dataset is not scalable.
- the methods described in this disclosure describe the case when the visual scenes and corresponding sound sources are not explicitly labeled. However, the disclosed methods can be adapted to the case even when such labels are available.
- Audio is first extracted from the video and each second of video frame is paired with the corresponding sound to create pairs of audio-visual training examples from which correlation can be learned.
- Each audio-visual training pair consists of a visual scene having one or more objects and actions, paired with audio comprising one or more sound sources mixed together (henceforth referred to as noisy audio), without any explicit labels or annotations describing the visual elements or sound sources.
- independent methods are disclosed to 1) learn a coarse-grained correlation between the visual input and noisy audio input directly, without separating the noisy audio into its sound sources, 2) learn a coarse-grained correlation by first predicting the dominant single sound sources (henceforth referred to as clean audio) in the noisy audio and using those single sound sources to leam a correlation with the visual input, 3) learn a more fine-grained correlation between local regions of the visual input and regions of the noisy audio input.
- these methods can be used independently or as an ensemble (mixture of models) to recommend sound effects for a visual scene.
- Fig. 2A depicts how a machine learning model is trained to leam the audio-visual correlation given a batch of audio-visual paired samples as training input.
- the visual input 200 may be a still image, video frame, or video segment.
- a noisy audio segment 201 may or may not be extracted from the visual input 200 In some embodiments, the noisy audio segment 201 is 1- second in duration but aspects of the present disclosure are not so limited and audio segments 201 may be greater than 1 -second in other embodiments.
- the raw audio signals 201 may be in any audio or audio/video signal format known in the art for example and without limitation, the audio signals may be file types such as MP3s, WAV, WMA, MP4, OGG, QT, AVI, MKV, etc.
- a corresponding label is applied representing that relationship. For example and without limitation, if the audio input 201 corresponds to an audio recording aligned with the same timeframe as the visual input 200 during the production of a video sequence, or the audio input 201 is a recording of sound made by an object or objects in the visual input 200, then the label 210 has a value 1. If the sound sources included in the audio input 201 is not related to the visual input 200 then a corresponding label representing the lack of a relationship is applied. For example and without limitation, if the audio input 201 and the visual input 200 are from different timeframes of a video sequence, then the label 210 has a value 0.
- the visual input 200 may be optionally transformed 202 (for example, resized, normalized, background subtraction) and applied as an input to a visual neural network 204.
- the visual NN 204 may be for example and without limitation a 2-dimensional or 3 -dimensional CNN having between 8 and 11 convolutional layers, in addition to any number of pooling layers, and may optionally include batch normalization and attention mechanisms.
- the visual NN 204 outputs a visual embedding 206, which is a mathematical representation learned from the visual input 200.
- the noisy audio input 201 is processed by a feature extractor 203 that extracts audio features, such as mel-filter banks or similar 2-dimensional spectral features.
- the audio features may be, optionally, normalized and padded to ensure that the audio features that are input to the audio NN 205 have a fixed dimension.
- the audio NN 205 may be for example and without limitation a CNN having between 8 and 11 convolutional layers with or without batch normalization, in addition to any number of pooling layers, and may, optionally, include batch normalization and attention mechanisms.
- the audio NN 205 outputs an audio embedding 207, which is a mathematical representation learned from the audio input 201.
- One or more subnetwork layers that are part of the NNs 204 and 205 may be chosen suitable to create a representation, or feature vector of the training data.
- the audio and image input subnetworks may produce embeddings in the form of feature vectors having 128 components, though aspects of the present disclosure are not limited to 128 component feature vectors and may encompass other feature vector configurations and embedding configurations.
- the audio embedding 207 and visual embedding 206 are compared by computing a distance value 208 between them. This distance value may be computed by any distance metric such as, but not limited to, Euclidean distance or LI distance. This distance value is a measure of the correlation between the audio-visual input pair. Smaller the distance, higher is the correlation.
- the correlation NN 209 predicts the correlation for an audio-visual input pair as a function of the distance value 208.
- NN 209 may contain one or more linear or non-linear layers.
- the prediction values are compared to the binary labels 210 using a loss function such as cross-entropy loss and the error between the predictions and respective labels is backpropagated through the entire network, including 204, 205, and 209 to improve the predictions.
- the goal of training may be to minimize the cross-entropy loss that measures the error between the predictions and labels, and/or the contrastive loss that minimizes the distance value 208 between correlated audio-visual embeddings while maximizing the distance between uncorrelated embeddings.
- the pairwise contrastive loss function L p airs between an audio-visual pair is given by EQ. 1 : where F (JR) is the output of the visual NN 204 for the reference image and F (A) is output of the audio NN 205 for the Audio signal A.
- the model depicted in Fig. 2A learns an audio embedding 207 and visual embedding 206, in such a way that the distance 208 between correlated audio and visual inputs is small, while the distance between the uncorrelated audio and visual embeddings is large.
- This trained pairwise audio-visual correlation model can be used in the Sound Recommendations tool to generate visual embeddings for any new silent video or image input and audio embeddings for a set of sound samples from which it can recommend the sound effects that are most correlated to the silent visual input by way of having the closest audio-visual embedding distance. The recommended sound effects may then be mixed with the silent visual input to produce a video with sound effects.
- Fig. 2B shows an alternative embodiment to train a machine learning model for learning audio-visual correlation to recommend sounds for visual input.
- the visual input 200 is an image or video frame or video segment and the audio input 201 may be a mixture of one or more audio sources.
- the embodiment in Fig. 2B differs from Fig. 2A in how the training audio-visual pairs are generated.
- the noisy audio input 201 is not directly used for training in this method. Instead, it is first processed by a noisy to clean mapping module 211, which identifies the one or more dominant sound sources that may be included in the audio input
- the noisy to Clean Mapping Module 211 may be trained in different ways. It may be an audio similarity model trained using pairwise similarity or triplet similarity methods. In some embodiments, it may be an audio classifier trained to classify sound sources in an audio mixture. Alternatively, it may be an audio source separation module trained using non negative matrix factorization (NMF), or a neural network trained for audio source separation (for example U-net). Regardless of how it is trained, the purpose of the Noisy to Clean Mapping Module 211 is to identify the top-K dominant reference sound sources that best match or are included in the audio input 201, where K may be any reasonable value such as, but not limited to, a value between 1 and 5.
- NMF non negative matrix factorization
- K sound sources may be considered as positive audio signals with respect to the visual input 200, because they are related to the visual scene.
- Selection module 212 selects K negative reference audio signals that are either complementary or different from the K positive signals.
- the visual input 200 is paired with each of the 2*K predicted clean audio signals to create 2*K audio-visual pairs for training the correlation NN 209 in Fig 2B, as described above for the previous embodiment shown in Fig. 2A.
- One half of the 2*K audio-visual pairs are positive pairs where the audio input is related or similar to the sound produced by one or more objects in the visual scene and each of these positive pairs has a label 210 of value 1.
- the other half of the 2*K audio visual input pairs are negative pairs where the audio input is not related to the visual input 200 and each of these negative pairs has a label 210 of value 0.
- the positive audio signals and negative audio signals may all be part of an audio database containing labeled audio signal files.
- the labeled audio signal files may be organized into a taxonomy where the K clean positive audio signals are part of the same category or sub category as the signals in the audio input 201, whereas the K clean negative audio signals may be part of a different category or sub category than the K positive audio signals.
- the audio-visual correlation is learned by a machine-learning model that takes triplets as inputs and is trained by a triplet contrastive loss function instead of a pairwise loss function.
- the inputs to the correlation NN may be a reference image or video 301, a positive audio signal 302 and a negative audio signal 303.
- the reference image or video 301 may be a still image or part of a reference video sequence as described above in embodiment Fig. 2B.
- the positive audio signal 302 is related the reference image or video 301, for example and without limitation the positive audio may be a recording of sound made by an object or objects in the reference image, the positive audio may be a recording or corresponding audio made during the production of the reference image.
- the negative audio signal 303 is different from the positive audio signal 302 and not related to the reference visual input 301.
- the visual input 301 may be the visual embedding 206 output by a trained correlation NN shown in Fig. 2B
- the positive audio input 302 and negative audio input 303 may be audio negative embeddings 207 output by a trained correlation NN shown in Fig. 2B, for a positive and negative audio signal respectively.
- F(I R ) is the embedding 308 of the neural network in training for the reference visual U R ).
- I A A Y) is the embedding 311 of the neural network in training for the negative audio (A ).
- F(Ap) is the embedding 309 of the neural network in training for the positive audio (A / ⁇ )
- m is a margin that defines the minimum separation between the embeddings for the negative audio and the positive audio.
- L mpiet is optimized during training to maximize the distance between the pairing of the reference visual input 301 and the negative audio 303 and minimize the distance between the reference visual input 301 and the positive audio 302.
- the correlational NN 305 is configured to learn visual and audio embeddings.
- the correlational NN learns embeddings in such a way so as to produce a distance value between the positive audio embedding 309 and reference image or video embedding at 308 that is less than the distance value between the negative audio embedding 311 and reference visual embedding 308.
- the distance may be, without limitation, computed as cosine distance, Euclidean distance, or any other type of pairwise distance function.
- the embedding generated by such a trained correlational NN can be used by a sound recommendation tool to recommend sound effects that can be matched with a visual scene or video segment, as will be discussed below.
- the machine learning models in Fig. 2A, Fig. 2B, and Fig. 3 leam a coarse-grained Audio- Visual correlation by encoding each audio input as well as visual input into a single coarse grained embedding (representation).
- the recommendation performance can be improved by learning a fine-grained correlation that is able to localize the audio sources by correlating the regions within the visual input that may be related to the different sound sources.
- Fig. 4 depicts such a method that learns a fine-grained audio-visual correlation by localizing the audio-visual features. This method may be considered as an extension of the method presented in Fig. 2A.
- the visual input 400 may be a still image, video frame, or video segment.
- the noisy audio input 401 may either be a positive audio segment related to the visual scene 400 in which case the label 410 may have a value of 1, or it may be a negative audio segment that is unrelated to the visual scene 400 with, for example and without limitation, a label 410 of value 0.
- label values of 1 and 0 are discussed explicitly because the described correlation is a binary correlation any labels that can be interpreted to describe a binary relationship may be used.
- the visual input may be optionally preprocessed and transformed by module 402 and the input is used for training the visual NN 404.
- Feature Extraction module 403 extracts 2D audio features, such as filterbank from the audio input 401, which are then used for training the audio NN 405.
- the visual NN 404 and audio NN 405 are multi-layered NN that includes one or more convolutional layers, pooling layers, and optionally recurrent layers and attention layers.
- a visual representation in the form of a 2D or higher dimensional feature map 406 is extracted from the visual NN 404.
- an audio representation in the form of a 2D or higher dimensional feature map 407 is extracted from the audio NN 405.
- These feature maps contain a set of feature vectors that represent higher-level features learned by the NN from different regions of the visual and audio input.
- the visual feature vectors may be optionally consolidated by clustering similar feature vectors together to yield K distinct visual clusters 408, using methods, such as by way of example but not by way of limitation, K-means clustering.
- the audio feature vectors in the audio feature map may be optionally consolidated into K distinct audio clusters 409. The audio feature vectors and visual feature vectors that are (optionally) clustered are then compared and localized by the Multimodal similarity module 411.
- the Multimodal similarity module 411 For each feature vector derived from the visual map, the Multimodal similarity module 411 computes the most correlated feature vector derived from the audio map and the corresponding correlation score, which may be computed by a similarity metric, such as by way of example, but not by way of limitation, cosine similarity.
- the correlation scores between different visual and audio feature vectors are then input to the correlation NN 412, which aggregates the scores to predict the overall correlation score for the audio-visual input pair.
- the prediction value is compared to the label 410 using a loss function such as cross-entropy loss and the error between the predictions and respective labels is backpropagated through the model to improve the prediction.
- the objective of training may be, but not limited, to minimizing the cross-entropy loss that measures the error between the predictions and labels.
- the model in Fig. 4 leams an audio representation and visual representation, in such a way that the representations of correlated audio and visual regions are more similar than that of uncorrelated regions.
- This trained fine grained audio-visual correlation model can then be used in the Sound Recommendations Tool to generate representations for a new silent video or image and a set of sound effect samples and by comparing those audio and visual representations, recommend sound effects that are most correlated to the different visual elements of the silent visual input.
- the video segments have a frame rate of 1 frame per second and as such each frame is used as an input reference image.
- the input image is generated by sampling a video segment with a higher frame down to 1 frame per second and using each frame as an input image.
- an input video segment may have a frame rate of 30 frames per second.
- the input video may be sampled every 15 frames to generate a down sampled 1 frame per second video, then each frame of the down sampled video may be used as input into the NNs.
- the audio database likewise may contain audio segments of 1 second in length, which may be selected from as positive or negative audio signals. Alternatively, the audio signals may be longer than 1 second in length and 1 second of audio may be selected from the longer audio segment.
- the first 1 second of the audio segment may be used or a 1 second sample in the middle of the audio maybe chosen or a 1 second sample at the end of the audio segment may be chosen or a 1 second sample from a random time in the audio segment may be chosen.
- FIG. 5 depicts the use of the Multi-modal Sound Recommendation tool according to aspects of the present disclosure.
- the Multi-modal sound recommendation tool may comprise an audio database 502 and a trained multi-modal correlation neural network 503.
- the input to the Multi-modal correlation NN 503 may be an input image frame or video without sound 501.
- the Multi-modal correlation NN 503 is configured to predict the correlation, quantified by a distance value 504, between the representations of the input image frame or video and each audio segment in an audio database or collection of audio samples. After a correlation value 504 has been generated for each audio segment from the audio database, the correlation values are sorted and filtered by 505 to select the audio segments that are best correlated to the input image/video (indicated by the lowest distance values).
- the sorting and filtering 505 may filter out every audio segment except the top correlated K audio segments, where K may be a reasonable value such as 1, 5, 10 or 20 audio segments. From this sorting and filtering 505 the most correlated audio segments may be selected either automatically or by a user using the correlation values 507. The best matching audio segment may then be recommended to the sound designer for mixing with the input image frame/video. In some alternative embodiments, more than one audio segment is chosen as a best match using their correlation values 507 and these audio segments are all recommended for the silent visual input 506.
- the audio segments in the audio database are subject to a feature extraction and optionally a feature normalization process before they are input to the Multi-modal sound selection NN 503.
- the extracted audio features may be for example and without limitation, filterbank, spectrogram or other similar 2D audio features.
- the input image/video may be subject to some transformations, such as feature normalization, resizing, cropping, before it is input to the Multi-modal sound selection network 503.
- the Multi-modal sound selection NN 503 may be one of the trained models from Fig. 2A, Fig. 2B, Fig 3, or Fig. 4, each configured to output audio-visual representations for the visual input 501 and the corresponding audio inputs, which may be audio segments from the audio database 502. These representations are then used to generate the correlated distance values 504 and select the top-K correlated sounds for the visual input.
- the Multi-modal sound recommendation tool may merge the top most recommended sounds from one or more trained models in Fig. 2A, Fig. 2B, Fig 3, or Fig. 4.
- the audio database 502 may contain a vast number of different audio segments arranged into a taxonomy. Searches of the database using the tool may yield too many correlated sounds, if there are no constraints. Therefore, according to some aspects of the present disclosure the input audio segments from the database 502 may be limited to a category or subcategory in the taxonomy. Alternatively, a visual understanding approach may be applied to limit searches to relevant portions of the database. Neural Networks trained for Object recognition and visual description to identify visual elements and map the visual elements to sound categories/subcategories may be used to limit searches within the audio databases.
- FIG. 6 depicts a multi-modal sound recommendation system for implementing training and the sound selection methods like that shown in Figures throughout the specification for example Figs. 1, 2, 3, 4 and 5.
- the system may include a computing device 600 coupled to a user input device 602.
- the user input device 602 may be a controller, touch screen, microphone, keyboard, mouse, joystick or other device that allows the user to input information including sound data in to the system.
- the user input device may be coupled to a haptic feedback device 621.
- the haptic feedback device 621 may be for example a vibration motor, force feedback system, ultrasonic feedback system, or air pressure feedback system.
- the computing device 600 may include one or more processor units 603, which may be configured according to well-known architectures, such as, e.g., single-core, dual-core, quad- core, multi-core, processor-coprocessor, cell processor, and the like.
- the computing device may also include one or more memory units 604 (e.g., random access memory (RAM), dynamic random access memory (DRAM), read-only memory (ROM), and the like).
- RAM random access memory
- DRAM dynamic random access memory
- ROM read-only memory
- the processor unit 603 may execute one or more programs, portions of which may be stored in the memory 604 and the processor 603 may be operatively coupled to the memory, e.g., by accessing the memory via a data bus 605.
- the programs may include machine learning algorithms 621 configured to adjust the weights and transition values of NNs 610 as discussed above where, the NNs 610 are any of the NNs shown in Figs 2, 3 or 4.
- the Memory 604 may store audio signals 608 that may be the positive, negative or reference audio used in training the NNs 610 with the machine learning algorithms 621. Additionally the reference, positive, and negative audio signals may be stored in the audio database 622. Image frames or videos 609 used in training the NNs 610 may also be stored in the Memory 604.
- the image frames or videos 609 may also be used with the audio database 622 in the operation of the sound recommendation tool as shown in FIG. 5 and described hereinabove.
- the database 622, image frames/video 609, audio signals 608 may be stored as data 618 and machine learning algorithms 621 may be stored as programs 617 in the Mass Store 618 or at a server coupled to the Network 620 accessed through the network interface 614.
- Input audio, image, and/or video may be stored as data 618 in the Mass Store 615.
- the processor unit 603 is further configured to execute one or more programs 617 stored in the mass store 615 or in memory 604, which cause the processor to carry out the one or more of the methods described above.
- the computing device 600 may also include well-known support circuits, such as input/output (I/O) 607, circuits, power supplies (P/S) 611, a clock (CLK) 612, and cache 613, which may communicate with other components of the system, e.g., via the bus 605.
- the computing device may include a network interface 614.
- the processor unit 603 and network interface 614 may be configured to implement a local area network (LAN) or personal area network (PAN), via a suitable network protocol, e.g., Bluetooth, for a PAN.
- the computing device may optionally include a mass storage device 615 such as a disk drive, CD-ROM drive, tape drive, flash memory, or the like, and the mass storage device may store programs and/or data.
- the computing device may also include a user interface 616 to facilitate interaction between the system and a user.
- the user interface may include a monitor, Television screen, speakers, headphones or other devices that communicate information to the user.
- the computing device 600 may include a network interface 614 to facilitate communication via an electronic communications network 620.
- the network interface 614 may be configured to implement wired or wireless communication over local area networks and wide area networks such as the Internet.
- the device 600 may send and receive data and/or requests for files via one or more message packets over the network 620.
- Message packets sent over the network 620 may temporarily be stored in a buffer in memory 604.
- the audio database may be available through the network 620 and stored partially in memory 604 for use.
- the proposed methods provide ways to learn audio-visual correlation (more generally multimodal correlation) in a self-supervised manner without requiring labels or manual annotations.
- the proposed machine learning method leams coarse-grained audio-visual representations based on noisy audio input and uses that to determine coarse-grained multimodal (audio-visual) correlation.
- the proposed machine learning method predicts the clean reference audio sources included in a noisy audio mixture and using the predicted clean audio sources to leam coarse-grained audio-visual representations and determines coarse grained multimodal (audio-visual) correlation.
- the machine learning methods can learn audio-visual representations and determine coarse-grained multimodal (audio-visual) correlations from input triplets consisting of reference image or video, a positive audio signal, and a negative audio signal with respect to the reference visual input.
- the multimodal correlation neural network after being trained can generate a representation (embedding) for a given audio.
- the multimodal correlation neural network after being trained can generate a representation (embedding) for a given image/video.
- the visual representation generated in and audio representation generated in are likely to be close (that is, distance between them is small).
- the visual representation generated and audio representation generated are likely to be dissimilar (that is, distance between them is large).
- a trained correlation NN or Multimodal clustering NN may be used to automatically select and recommend only those sound samples that are most relevant for a visual scene or video.
- the selected sound samples may refer to sounds directly produced by one or more objects in the visual scene and /or may be indirectly associated with one or more objects in the visual scene. While the above is a complete description of the preferred embodiment of the present invention, it is possible to use various alternatives, modifications and equivalents. Therefore, the scope of the present invention should be determined not with reference to the above description but should, instead, be determined with reference to the appended claims, along with their full scope of equivalents. Any feature described herein, whether preferred or not, may be combined with any other feature described herein, whether preferred or not.
- the indefinite article “A” or “An” refers to a quantity of one or more of the item following the article, except where expressly stated otherwise.
- the appended claims are not to be interpreted as including means-plus-function limitations, unless such a limitation is explicitly recited in a given claim using the phrase “means for.”
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- General Physics & Mathematics (AREA)
- Artificial Intelligence (AREA)
- Software Systems (AREA)
- Evolutionary Computation (AREA)
- Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Mathematical Physics (AREA)
- Computing Systems (AREA)
- Biophysics (AREA)
- General Health & Medical Sciences (AREA)
- Biomedical Technology (AREA)
- Life Sciences & Earth Sciences (AREA)
- Molecular Biology (AREA)
- Multimedia (AREA)
- Databases & Information Systems (AREA)
- Library & Information Science (AREA)
- Acoustics & Sound (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Medical Informatics (AREA)
- Human Computer Interaction (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Image Analysis (AREA)
Abstract
Sound effect recommendations for visual input are generated by training machine learning models that learn coarse-grained and fine-grained audio-visual correlations from a reference image, a positive audio signals, and a negative audio signal. A positive audio embedding related to the reference image is generated from the positive audio signal and a negative audio embedding is generated from a negative audio signal. A machine learning algorithm uses the reference image, the positive audio embedding and the negative audio embedding as inputs to train a visual-to-audio correlation neural network to output a smaller distance between the positive audio embedding and the reference image than the negative audio embedding and the reference image.
Description
SELF-SUPERVISED AI-ASSISTED SOUND EFFECT RECOMMENDATION FOR SILENT VIDEO
FIELD OF THE INVENTION
The present disclosure relates to sound effect selection for media, specifically aspects of the present disclosure relate to using machine-learning techniques for sound selection in media.
BACKGROUND OF THE INVENTION
Sound designers for video games and movies often look at objects occurring in video to determine what sounds to apply to the video. Since the inception of sound synchronized movies (colloquially called talkies) sound designers, have been generating corpuses of recorded audio segments. Today, these collections of audio segments are stored in digital audio databases that are searchable by the sound designers.
When a sound designer wants to add a sound effect to a silent video sequence, they have to watch the video sequence and imagine what the sounds occurring within the video might be like. Then the designer must search through the sound database and find sounds that match the context in the visual scene. This makes the sound designing process quite an artistic, iterative process and means that sounds chosen for media sometimes differ radically from reality. In everyday life, most objects create sounds based on their physical properties and not based on an imagined sound design. Thus, sounds can be considered to be almost related to the physical context of their productions. BRIEF DESCRIPTION OF THE DRAWINGS
The teachings of the present disclosure can be readily understood by considering the following detailed description in conjunction with the accompanying drawings, in which:
FIG. 1 A is a simplified diagram of a convolutional neural network for use in a Sound Effect Recommendation Tool according to aspects of the present disclosure.
FIG. IB is a simplified node diagram of a recurrent neural network for use in a Sound Effect Recommendation Tool according to aspects of the present disclosure.
FIG. 1C is a simplified node diagram of an unfolded recurrent neural network for use in a Sound Effect Recommendation Tool according to aspects of the present disclosure. FIG. ID is a block diagram of a method for training a neural network in development of a Sound Effect Recommendation Tool according to aspects of the present disclosure.
FIG. 2A is a block diagram depicting a method for training an audio-visual correlation NN using visual input paired with audio containing a noisy mixture of sound sources, for use in the Sound Recommendation tool, according to aspects of the present disclosure. FIG. 2B is a block diagram depicting a method that first maps an audio containing a mixture of sound sources into individual sound sources, which are then paired with the visual input for training an audio-visual correlation NN for use in the Sound Recommendation tool, according to aspects of the present disclosure.
FIG. 3 is a block diagram depicting training of an audio-visual Correlation NN that learns positive and negative correlations simultaneously using triplet inputs containing a visual input, positive correlated audio, and negative uncorrelated audio, for use in the Sound Recommendation tool according to aspects of the present disclosure.
FIG. 4 is a block diagram showing the training of a NN for learning fine-grained audio-visual correlations based on audio containing a mixture of sound sources, for use in the Sound Recommendation Tool according to aspects of the present disclosure.
FIG. 5 is a block diagram that depicts a method of using the trained NN in a Sound Effect Recommendation tool for creating a new video with sound, according to aspects of the present disclosure.
FIG. 6 is a block system diagram depicting a system implementing the training of neural networks and use of the Sound Effect Recommendation Tool according to aspects of the present disclosure.
DESCRIPTION OF THE SPECIFIC EMBODIMENTS Although the following detailed description contains many specific details for the purposes of illustration, anyone of ordinary skill in the art will appreciate that many variations and alterations to the following details are within the scope of the invention. Accordingly, the exemplary embodiments of the invention described below are set forth without any loss of generality to, and without imposing limitations upon, the claimed invention.
According to aspects of the present disclosure, Neural Networks (NN) and machine learning may be applied to sound design to choose appropriate sounds for video sequences that lack sound. Three techniques for developing a Sound Effect Recommendation Tool will be discussed herein. First general NN training methods will be discussed. Second, a method will be discussed for training a coarse-grained correlation NN for prediction of sound effects based on a reference video, directly from an audio mixture as well as by mapping the audio mixture to single audio sources using a similarity NN. The third method that will be discussed is for training a fine-grained correlation NN for recommending sound effects based on a reference video. Finally, use of a tool employing the trained Sound Effect Recommendation Networks individually or as a combination will be discussed.
GENERAL NN TRAINING
According to aspects of the present disclosure, the Sound Effect Recommendation Tool may include one or more of several different types of neural networks and may have many different layers. By way of example and not by way of limitation a classification neural network may consist of one or multiple deep neural networks (DNN), such as convolutional neural networks (CNN) and/or recurrent neural networks (RNN). The Sound Effect Recommendation Tool may be trained using the general training method disclosed herein.
FIG. 1A depicts an example layout of a convolution neural network according to aspects of the present disclosure. In this depiction, the convolution neural network is generated for an input 132 with a size of 4 units in height and 4 units in width giving a total area of 16 units. The depicted convolutional neural network has a filter 133 size of 2 units in height and 2 units in width with a stride value of 1 and a channel 136 of size 9. For clarity in FIG. 1A only the connections 134 between the first column of channels and their filter windows is depicted. Aspects of the present disclosure, however, are not limited to such implementations. According to aspects of the present disclosure, the convolutional neural
network may have any number of additional neural network node layers 131 and may include such layer types as additional convolutional layers, fully connected layers, pooling layers, max pooling layers, normalization layers, etc. of any size.
For illustrative purposes a RNN is described herein, it should be noted that RNNs differ from a basic NN in the addition of a hidden recurrent layer. FIG. IB depicts the basic form of an RNN having a layer of nodes 120, each of which is characterized by an activation function S, input U, a recurrent node weight W, and an output V. The activation function S is typically a non-linear function known in the art and is not limited to the (hyperbolic tangent (tanh) function. For example, the activation function S may be a Sigmoid or ReLU function. As shown in FIG. 1C, the RNN may be considered as a series of nodes 120 having the same activation function with the value of the activation function S moving through time from SO prior to T to SI after T and S2 after T+l. The nodes in a layer of RNN apply the same set of activation functions and weights to a series of inputs. The output of each node depends not just on the activation function and weights applied on that node’s input, but also on that node’s previous context. Thus, the RNN uses historical information by feeding the result from a previous time T to a current time T+l.
In some embodiments, a convolutional RNN may be used, especially when the visual input is a video. Another type of RNN that may be used is a Long Short-Term Memory (LSTM) Neural Network which adds a memory block in a RNN node with input gate activation function, output gate activation function and forget gate activation function resulting in a gating memory that allows the network to retain some information for a longer period of time as described by Hochreiter & Schmidhuber “Long Short-term memory” Neural Computation 9(8): 1735-1780 (1997), which is incorporated herein by reference.
As seen in FIG. ID Training a neural network (NN) begins with initialization of the weights of the NN at 141. In general, the initial weights should be distributed randomly. For example, an NN with a tanh activation function should have random values distributed between — 1 and 4 where n is the number of inputs to the node. n
After initialization the activation function and optimizer is defined. The NN is then provided with a feature vector or input dataset at 142. Each of the different feature vectors may be
generated by the NN from inputs that have known relationships. Similarly, the NN may be provided with feature vectors that correspond to inputs having known relationships. The NN then predicts a distance between the features or inputs at 143. The predicted distance is compared to the known relationship (also known as ground truth) and a loss function measures the total error between the predictions and ground truth over all the training samples at 144. By way of example and not by way of limitation the loss function may be a cross entropy loss function, quadratic cost, triplet contrastive function, exponential cost, mean square error etc. Multiple different loss functions may be used depending on the purpose. By way of example and not by way of limitation, for training classifiers a cross entropy loss function may be used whereas for learning an embedding a triplet contrastive loss function may be employed. The NN is then optimized and trained, using known methods of training for neural networks such as backpropagating the result of the loss function and by using optimizers, such as stochastic and adaptive gradient descent etc., as indicated at 145. In each training epoch, the optimizer tries to choose the model parameters (i.e., weights) that minimize the training loss function (i.e. total error). Data is partitioned into training, validation, and test samples.
During training, the Optimizer minimizes the loss function on the training samples. After each training epoch, the model is evaluated on the validation sample by computing the validation loss and accuracy. If there is no significant change, training can be stopped and the most optimal model resulting from the training may be used to predict the labels or relationships for the test data.
Thus, the neural network may be trained from inputs having known relationships to group related inputs. Similarly, a NN may be trained using the described method to generate a feature vector from inputs having known relationships.
Self-supervised Audio-Visual Correlation
The automated methods for recommending sound effects for visual scenes is based on learning audio-visual correlations by training on a large number of example videos (such as video games or movie clips), with one or more sound sources mixed together. One method to generate training data in order to train a model that learns audio-visual relationships is to generate audio-visual segment pairs from videos with labeled sound sources. However, manually detecting and labeling each sound source to create a large training dataset is not
scalable. The methods described in this disclosure describe the case when the visual scenes and corresponding sound sources are not explicitly labeled. However, the disclosed methods can be adapted to the case even when such labels are available.
Audio is first extracted from the video and each second of video frame is paired with the corresponding sound to create pairs of audio-visual training examples from which correlation can be learned. Each audio-visual training pair consists of a visual scene having one or more objects and actions, paired with audio comprising one or more sound sources mixed together (henceforth referred to as noisy audio), without any explicit labels or annotations describing the visual elements or sound sources. Given this set of audio-visual training pairs, independent methods are disclosed to 1) learn a coarse-grained correlation between the visual input and noisy audio input directly, without separating the noisy audio into its sound sources, 2) learn a coarse-grained correlation by first predicting the dominant single sound sources (henceforth referred to as clean audio) in the noisy audio and using those single sound sources to leam a correlation with the visual input, 3) learn a more fine-grained correlation between local regions of the visual input and regions of the noisy audio input. After training, these methods can be used independently or as an ensemble (mixture of models) to recommend sound effects for a visual scene. These 3 methods are now described.
Learning Coarse-grained Correlation from Noisy Audio-Visual Pairs
Fig. 2A depicts how a machine learning model is trained to leam the audio-visual correlation given a batch of audio-visual paired samples as training input. The visual input 200 may be a still image, video frame, or video segment. A noisy audio segment 201 may or may not be extracted from the visual input 200 In some embodiments, the noisy audio segment 201 is 1- second in duration but aspects of the present disclosure are not so limited and audio segments 201 may be greater than 1 -second in other embodiments. The raw audio signals 201 according to aspects of the present disclosure may be in any audio or audio/video signal format known in the art for example and without limitation, the audio signals may be file types such as MP3s, WAV, WMA, MP4, OGG, QT, AVI, MKV, etc.
If one or more sound sources included in the audio input 201 is related to the visual input 200 then a corresponding label is applied representing that relationship. For example and without limitation, if the audio input 201 corresponds to an audio recording aligned with the same timeframe as the visual input 200 during the production of a video sequence, or the audio input 201 is a recording of sound made by an object or objects in the visual input 200, then
the label 210 has a value 1. If the sound sources included in the audio input 201 is not related to the visual input 200 then a corresponding label representing the lack of a relationship is applied. For example and without limitation, if the audio input 201 and the visual input 200 are from different timeframes of a video sequence, then the label 210 has a value 0. The visual input 200 may be optionally transformed 202 (for example, resized, normalized, background subtraction) and applied as an input to a visual neural network 204. In some embodiments the visual NN 204 may be for example and without limitation a 2-dimensional or 3 -dimensional CNN having between 8 and 11 convolutional layers, in addition to any number of pooling layers, and may optionally include batch normalization and attention mechanisms. The visual NN 204 outputs a visual embedding 206, which is a mathematical representation learned from the visual input 200.
Similarly, the noisy audio input 201 is processed by a feature extractor 203 that extracts audio features, such as mel-filter banks or similar 2-dimensional spectral features. The audio features may be, optionally, normalized and padded to ensure that the audio features that are input to the audio NN 205 have a fixed dimension. In some embodiments, the audio NN 205 may be for example and without limitation a CNN having between 8 and 11 convolutional layers with or without batch normalization, in addition to any number of pooling layers, and may, optionally, include batch normalization and attention mechanisms. The audio NN 205 outputs an audio embedding 207, which is a mathematical representation learned from the audio input 201.
One or more subnetwork layers that are part of the NNs 204 and 205 may be chosen suitable to create a representation, or feature vector of the training data. In some implementations the audio and image input subnetworks may produce embeddings in the form of feature vectors having 128 components, though aspects of the present disclosure are not limited to 128 component feature vectors and may encompass other feature vector configurations and embedding configurations. The audio embedding 207 and visual embedding 206 are compared by computing a distance value 208 between them. This distance value may be computed by any distance metric such as, but not limited to, Euclidean distance or LI distance. This distance value is a measure of the correlation between the audio-visual input pair. Smaller the distance, higher is the correlation.
The correlation NN 209 predicts the correlation for an audio-visual input pair as a function of the distance value 208. NN 209 may contain one or more linear or non-linear layers. During
each training epoch, the prediction values are compared to the binary labels 210 using a loss function such as cross-entropy loss and the error between the predictions and respective labels is backpropagated through the entire network, including 204, 205, and 209 to improve the predictions. The goal of training may be to minimize the cross-entropy loss that measures the error between the predictions and labels, and/or the contrastive loss that minimizes the distance value 208 between correlated audio-visual embeddings while maximizing the distance between uncorrelated embeddings. The pairwise contrastive loss function L pairs between an audio-visual pair is given by EQ. 1 :
where F (JR) is the output of the visual NN 204 for the reference image and F (A) is output of the audio NN 205 for the Audio signal A.
After many iterations of training including both the negative, uncorrelated audio-visual input pairs and the positive, correlated audio-visual input pairs, the model depicted in Fig. 2A learns an audio embedding 207 and visual embedding 206, in such a way that the distance 208 between correlated audio and visual inputs is small, while the distance between the uncorrelated audio and visual embeddings is large. This trained pairwise audio-visual correlation model can be used in the Sound Recommendations tool to generate visual embeddings for any new silent video or image input and audio embeddings for a set of sound samples from which it can recommend the sound effects that are most correlated to the silent visual input by way of having the closest audio-visual embedding distance. The recommended sound effects may then be mixed with the silent visual input to produce a video with sound effects.
Learning Coarse-grained Audio-Visual Correlation by Predicting Sound Sources
Fig. 2B shows an alternative embodiment to train a machine learning model for learning audio-visual correlation to recommend sounds for visual input. As described in the previous embodiment Fig. 2 A, the visual input 200 is an image or video frame or video segment and the audio input 201 may be a mixture of one or more audio sources. The embodiment in Fig. 2B differs from Fig. 2A in how the training audio-visual pairs are generated. Unlike the previous embodiment, the noisy audio input 201 is not directly used for training in this method. Instead, it is first processed by a noisy to clean mapping module 211, which
identifies the one or more dominant sound sources that may be included in the audio input
201
The Noisy to Clean Mapping Module 211 may be trained in different ways. It may be an audio similarity model trained using pairwise similarity or triplet similarity methods. In some embodiments, it may be an audio classifier trained to classify sound sources in an audio mixture. Alternatively, it may be an audio source separation module trained using non negative matrix factorization (NMF), or a neural network trained for audio source separation (for example U-net). Regardless of how it is trained, the purpose of the Noisy to Clean Mapping Module 211 is to identify the top-K dominant reference sound sources that best match or are included in the audio input 201, where K may be any reasonable value such as, but not limited to, a value between 1 and 5. These K sound sources may be considered as positive audio signals with respect to the visual input 200, because they are related to the visual scene. Given these K positive audio signals, Selection module 212 selects K negative reference audio signals that are either complementary or different from the K positive signals. Thus the result of the Noisy to Clean Mapping Module 211 and selection module 212 together is to predict a total of 2*K “clean” single source reference audio signals 213. These reference audio signals may or may not part of an audio database. The visual input 200 is paired with each of the 2*K predicted clean audio signals to create 2*K audio-visual pairs for training the correlation NN 209 in Fig 2B, as described above for the previous embodiment shown in Fig. 2A. One half of the 2*K audio-visual pairs are positive pairs where the audio input is related or similar to the sound produced by one or more objects in the visual scene and each of these positive pairs has a label 210 of value 1. The other half of the 2*K audio visual input pairs are negative pairs where the audio input is not related to the visual input 200 and each of these negative pairs has a label 210 of value 0. In some embodiments, the positive audio signals and negative audio signals may all be part of an audio database containing labeled audio signal files. The labeled audio signal files may be organized into a taxonomy where the K clean positive audio signals are part of the same category or sub category as the signals in the audio input 201, whereas the K clean negative audio signals may be part of a different category or sub category than the K positive audio signals.
In some embodiments, the audio-visual correlation is learned by a machine-learning model that takes triplets as inputs and is trained by a triplet contrastive loss function instead of a pairwise loss function. As shown in FIG. 3, the inputs to the correlation NN may be a
reference image or video 301, a positive audio signal 302 and a negative audio signal 303. The reference image or video 301 may be a still image or part of a reference video sequence as described above in embodiment Fig. 2B. As described above, the positive audio signal 302 is related the reference image or video 301, for example and without limitation the positive audio may be a recording of sound made by an object or objects in the reference image, the positive audio may be a recording or corresponding audio made during the production of the reference image. As described above, the negative audio signal 303 is different from the positive audio signal 302 and not related to the reference visual input 301. In some embodiments, the visual input 301 may be the visual embedding 206 output by a trained correlation NN shown in Fig. 2B, and the positive audio input 302 and negative audio input 303 may be audio negative embeddings 207 output by a trained correlation NN shown in Fig. 2B, for a positive and negative audio signal respectively.
The visual input 301 may be optionally transformed by operations 304 such as, but not limited to, resizing and normalization, before it is input to the triplet correlation NN 305. Likewise, the positive and negative audio input may be preprocessed to extract audio features 310 that are suitable for training the correlation NN 305. In this embodiment, no additional labels are necessary. The correlation NN 305 is trained through multiple iterations to simultaneously learn a visual embedding and audio embeddings for the positive and negative audio input. The triplet contrastive loss function used to train NN 305 seeks to minimize the distance 306 between the reference visual embedding 308 and the positive audio embedding 309 while simultaneously maximizing the distance 307 between the reference visual embedding 308 and the negative audio embedding 311. The triplet contrastive learning loss function may be expressed as:
Where F(IR) is the embedding 308 of the neural network in training for the reference visual UR). I A A Y) is the embedding 311 of the neural network in training for the negative audio (A ). and F(Ap) is the embedding 309 of the neural network in training for the positive audio (A/·) m is a margin that defines the minimum separation between the embeddings for the negative audio and the positive audio. Lmpiet is optimized during training to maximize the distance between the pairing of the reference visual input 301 and the negative audio 303 and minimize the distance between the reference visual input 301 and the positive audio 302.
After many rounds of training with triplets, including both the negative training set 303 and the positive training set 302, the correlational NN 305 is configured to learn visual and audio embeddings. The correlational NN learns embeddings in such a way so as to produce a distance value between the positive audio embedding 309 and reference image or video embedding at 308 that is less than the distance value between the negative audio embedding 311 and reference visual embedding 308. The distance may be, without limitation, computed as cosine distance, Euclidean distance, or any other type of pairwise distance function. The embedding generated by such a trained correlational NN can be used by a sound recommendation tool to recommend sound effects that can be matched with a visual scene or video segment, as will be discussed below.
Learning Fine-grained Audio-Visual Correlation Through Localization
The machine learning models in Fig. 2A, Fig. 2B, and Fig. 3 leam a coarse-grained Audio- Visual correlation by encoding each audio input as well as visual input into a single coarse grained embedding (representation). When the visual input is a complex scene with multiple objects and the audio input is a mixture of a sound sources, the recommendation performance can be improved by learning a fine-grained correlation that is able to localize the audio sources by correlating the regions within the visual input that may be related to the different sound sources. Fig. 4 depicts such a method that learns a fine-grained audio-visual correlation by localizing the audio-visual features. This method may be considered as an extension of the method presented in Fig. 2A. The visual input 400 may be a still image, video frame, or video segment. As described above for Fig. 2 A, the noisy audio input 401 may either be a positive audio segment related to the visual scene 400 in which case the label 410 may have a value of 1, or it may be a negative audio segment that is unrelated to the visual scene 400 with, for example and without limitation, a label 410 of value 0. Though label values of 1 and 0 are discussed explicitly because the described correlation is a binary correlation any labels that can be interpreted to describe a binary relationship may be used.
The visual input may be optionally preprocessed and transformed by module 402 and the input is used for training the visual NN 404. Similarly, Feature Extraction module 403 extracts 2D audio features, such as filterbank from the audio input 401, which are then used for training the audio NN 405. The visual NN 404 and audio NN 405 are multi-layered NN that includes one or more convolutional layers, pooling layers, and optionally recurrent layers and attention layers. A visual representation in the form of a 2D or higher dimensional feature
map 406 is extracted from the visual NN 404. Similarly, an audio representation in the form of a 2D or higher dimensional feature map 407 is extracted from the audio NN 405. These feature maps contain a set of feature vectors that represent higher-level features learned by the NN from different regions of the visual and audio input.
Some of the feature vectors within the audio and feature maps may be similar. Hence, the visual feature vectors may be optionally consolidated by clustering similar feature vectors together to yield K distinct visual clusters 408, using methods, such as by way of example but not by way of limitation, K-means clustering. Similarly, the audio feature vectors in the audio feature map may be optionally consolidated into K distinct audio clusters 409. The audio feature vectors and visual feature vectors that are (optionally) clustered are then compared and localized by the Multimodal similarity module 411. For each feature vector derived from the visual map, the Multimodal similarity module 411 computes the most correlated feature vector derived from the audio map and the corresponding correlation score, which may be computed by a similarity metric, such as by way of example, but not by way of limitation, cosine similarity. The correlation scores between different visual and audio feature vectors (representing different regions of the input visual scene and audio input) are then input to the correlation NN 412, which aggregates the scores to predict the overall correlation score for the audio-visual input pair. During each training epoch, the prediction value is compared to the label 410 using a loss function such as cross-entropy loss and the error between the predictions and respective labels is backpropagated through the model to improve the prediction. The objective of training may be, but not limited, to minimizing the cross-entropy loss that measures the error between the predictions and labels.
After many iterations of training including both the negative, uncorrelated audio-visual input pairs and the positive, correlated audio-visual input pairs, the model in Fig. 4 leams an audio representation and visual representation, in such a way that the representations of correlated audio and visual regions are more similar than that of uncorrelated regions. This trained fine grained audio-visual correlation model can then be used in the Sound Recommendations Tool to generate representations for a new silent video or image and a set of sound effect samples and by comparing those audio and visual representations, recommend sound effects that are most correlated to the different visual elements of the silent visual input.
In some embodiments, the video segments have a frame rate of 1 frame per second and as such each frame is used as an input reference image. In some alternative embodiments, the
input image is generated by sampling a video segment with a higher frame down to 1 frame per second and using each frame as an input image. For example and without limitation an input video segment may have a frame rate of 30 frames per second. The input video may be sampled every 15 frames to generate a down sampled 1 frame per second video, then each frame of the down sampled video may be used as input into the NNs. The audio database likewise may contain audio segments of 1 second in length, which may be selected from as positive or negative audio signals. Alternatively, the audio signals may be longer than 1 second in length and 1 second of audio may be selected from the longer audio segment. For example and without limitation the first 1 second of the audio segment may be used or a 1 second sample in the middle of the audio maybe chosen or a 1 second sample at the end of the audio segment may be chosen or a 1 second sample from a random time in the audio segment may be chosen.
Multi-Modal Sound Recommendation Tool
FIG. 5 depicts the use of the Multi-modal Sound Recommendation tool according to aspects of the present disclosure. The Multi-modal sound recommendation tool may comprise an audio database 502 and a trained multi-modal correlation neural network 503. The input to the Multi-modal correlation NN 503 may be an input image frame or video without sound 501. The Multi-modal correlation NN 503 is configured to predict the correlation, quantified by a distance value 504, between the representations of the input image frame or video and each audio segment in an audio database or collection of audio samples. After a correlation value 504 has been generated for each audio segment from the audio database, the correlation values are sorted and filtered by 505 to select the audio segments that are best correlated to the input image/video (indicated by the lowest distance values). The sorting and filtering 505 without limitation may filter out every audio segment except the top correlated K audio segments, where K may be a reasonable value such as 1, 5, 10 or 20 audio segments. From this sorting and filtering 505 the most correlated audio segments may be selected either automatically or by a user using the correlation values 507. The best matching audio segment may then be recommended to the sound designer for mixing with the input image frame/video. In some alternative embodiments, more than one audio segment is chosen as a best match using their correlation values 507 and these audio segments are all recommended for the silent visual input 506.
The audio segments in the audio database are subject to a feature extraction and optionally a feature normalization process before they are input to the Multi-modal sound selection NN 503. The extracted audio features may be for example and without limitation, filterbank, spectrogram or other similar 2D audio features. Similarly, the input image/video may be subject to some transformations, such as feature normalization, resizing, cropping, before it is input to the Multi-modal sound selection network 503.
According to some aspects of the present disclosure the Multi-modal sound selection NN 503 may be one of the trained models from Fig. 2A, Fig. 2B, Fig 3, or Fig. 4, each configured to output audio-visual representations for the visual input 501 and the corresponding audio inputs, which may be audio segments from the audio database 502. These representations are then used to generate the correlated distance values 504 and select the top-K correlated sounds for the visual input. According to other alternative aspects of the present disclosure the Multi-modal sound recommendation tool may merge the top most recommended sounds from one or more trained models in Fig. 2A, Fig. 2B, Fig 3, or Fig. 4.
According to some aspects of the present disclosure, the audio database 502 may contain a vast number of different audio segments arranged into a taxonomy. Searches of the database using the tool may yield too many correlated sounds, if there are no constraints. Therefore, according to some aspects of the present disclosure the input audio segments from the database 502 may be limited to a category or subcategory in the taxonomy. Alternatively, a visual understanding approach may be applied to limit searches to relevant portions of the database. Neural Networks trained for Object recognition and visual description to identify visual elements and map the visual elements to sound categories/subcategories may be used to limit searches within the audio databases.
System
FIG. 6 depicts a multi-modal sound recommendation system for implementing training and the sound selection methods like that shown in Figures throughout the specification for example Figs. 1, 2, 3, 4 and 5. The system may include a computing device 600 coupled to a user input device 602. The user input device 602 may be a controller, touch screen, microphone, keyboard, mouse, joystick or other device that allows the user to input information including sound data in to the system. The user input device may be coupled to a
haptic feedback device 621. The haptic feedback device 621 may be for example a vibration motor, force feedback system, ultrasonic feedback system, or air pressure feedback system.
The computing device 600 may include one or more processor units 603, which may be configured according to well-known architectures, such as, e.g., single-core, dual-core, quad- core, multi-core, processor-coprocessor, cell processor, and the like. The computing device may also include one or more memory units 604 (e.g., random access memory (RAM), dynamic random access memory (DRAM), read-only memory (ROM), and the like).
The processor unit 603 may execute one or more programs, portions of which may be stored in the memory 604 and the processor 603 may be operatively coupled to the memory, e.g., by accessing the memory via a data bus 605. The programs may include machine learning algorithms 621 configured to adjust the weights and transition values of NNs 610 as discussed above where, the NNs 610 are any of the NNs shown in Figs 2, 3 or 4. Additionally, the Memory 604 may store audio signals 608 that may be the positive, negative or reference audio used in training the NNs 610 with the machine learning algorithms 621. Additionally the reference, positive, and negative audio signals may be stored in the audio database 622. Image frames or videos 609 used in training the NNs 610 may also be stored in the Memory 604. The image frames or videos 609 may also be used with the audio database 622 in the operation of the sound recommendation tool as shown in FIG. 5 and described hereinabove. The database 622, image frames/video 609, audio signals 608 may be stored as data 618 and machine learning algorithms 621 may be stored as programs 617 in the Mass Store 618 or at a server coupled to the Network 620 accessed through the network interface 614.
Input audio, image, and/or video, may be stored as data 618 in the Mass Store 615. The processor unit 603 is further configured to execute one or more programs 617 stored in the mass store 615 or in memory 604, which cause the processor to carry out the one or more of the methods described above.
The computing device 600 may also include well-known support circuits, such as input/output (I/O) 607, circuits, power supplies (P/S) 611, a clock (CLK) 612, and cache 613, which may communicate with other components of the system, e.g., via the bus 605. The computing device may include a network interface 614. The processor unit 603 and network interface 614 may be configured to implement a local area network (LAN) or personal area
network (PAN), via a suitable network protocol, e.g., Bluetooth, for a PAN. The computing device may optionally include a mass storage device 615 such as a disk drive, CD-ROM drive, tape drive, flash memory, or the like, and the mass storage device may store programs and/or data. The computing device may also include a user interface 616 to facilitate interaction between the system and a user. The user interface may include a monitor, Television screen, speakers, headphones or other devices that communicate information to the user.
The computing device 600 may include a network interface 614 to facilitate communication via an electronic communications network 620. The network interface 614 may be configured to implement wired or wireless communication over local area networks and wide area networks such as the Internet. The device 600 may send and receive data and/or requests for files via one or more message packets over the network 620. Message packets sent over the network 620 may temporarily be stored in a buffer in memory 604. The audio database may be available through the network 620 and stored partially in memory 604 for use.
The proposed methods provide ways to learn audio-visual correlation (more generally multimodal correlation) in a self-supervised manner without requiring labels or manual annotations. The proposed machine learning method leams coarse-grained audio-visual representations based on noisy audio input and uses that to determine coarse-grained multimodal (audio-visual) correlation. The proposed machine learning method predicts the clean reference audio sources included in a noisy audio mixture and using the predicted clean audio sources to leam coarse-grained audio-visual representations and determines coarse grained multimodal (audio-visual) correlation. The machine learning methods can learn audio-visual representations and determine coarse-grained multimodal (audio-visual) correlations from input triplets consisting of reference image or video, a positive audio signal, and a negative audio signal with respect to the reference visual input. The multimodal correlation neural network after being trained can generate a representation (embedding) for a given audio. The multimodal correlation neural network after being trained can generate a representation (embedding) for a given image/video. For a pair of correlated image/video and audio, the visual representation generated in and audio representation generated in are likely to be close (that is, distance between them is small). For a pair of uncorrelated image/video and audio, the visual representation generated and audio representation generated are likely to be dissimilar (that is, distance between them is large). A trained correlation NN or
Multimodal clustering NN may be used to automatically select and recommend only those sound samples that are most relevant for a visual scene or video. The selected sound samples may refer to sounds directly produced by one or more objects in the visual scene and /or may be indirectly associated with one or more objects in the visual scene. While the above is a complete description of the preferred embodiment of the present invention, it is possible to use various alternatives, modifications and equivalents. Therefore, the scope of the present invention should be determined not with reference to the above description but should, instead, be determined with reference to the appended claims, along with their full scope of equivalents. Any feature described herein, whether preferred or not, may be combined with any other feature described herein, whether preferred or not. In the claims that follow, the indefinite article “A” or “An” refers to a quantity of one or more of the item following the article, except where expressly stated otherwise. The appended claims are not to be interpreted as including means-plus-function limitations, unless such a limitation is explicitly recited in a given claim using the phrase “means for.”
Claims
1. A method for training a Sound Effect Recommendation Network, comprising: a)generating a positive audio embedding from a positive audio signal wherein the positive audio signal is related to a reference image; b) generating a negative audio embedding from a negative audio signal; b) using a machine learning algorithm with, the reference image, the positive audio embedding and the negative audio embedding as inputs to train a visual-to-audio correlation neural network to output a smaller distance between the positive audio embedding and the reference image than the negative audio embedding and the reference image.
2. The method of claim 1 further comprising determining the positive audio signal from a noisy positive audio signal using a trained similarity neural network with an audio database and the noisy positive audio signal to select the positive audio signal from the audio database, wherein a reference audio signal is part of the audio database and wherein the reference audio is part of an audio visual sequence having the reference image.
3. The method of claim 2 wherein the positive audio signal is the reference audio signal.
4. The method of claim 2 further comprising determining the negative audio signal using the trained similarity neural network with the audio database wherein the negative signal is not related to the positive audio signal.
5. The method of claim 1 wherein features are extracted from the positive signal, reference signal and negative signal before being used in training with the machine learning algorithm.
6. The method of claim 1 wherein the positive audio signal includes noise.
7. The method of claim 1 wherein the machine learning algorithm used to train the visual-to- audio correlation neural network includes a pairwise loss function.
8. The method of claim 1 wherein the positive signal is part of an audio/video sequence that includes the reference image signal, wherein the positive signal includes noise signals and
wherein the noise signals are other sounds occurring in the audio/video sequence and wherein the negative signal includes noise signals.
9. The method of claim 1 wherein the machine learning algorithm is a self-supervised learning algorithm and wherein the positive, negative and correlated audio are unlabeled or unannotated inputs.
10. A system for training a Sound Effect Recommendation Network, comprising: a Processor; a Memory coupled to the Processor;
Non-transitory instructions embedded in the memory that when executed cause the processor to carry out the method for training a sound effect recommendation network comprising; a) generating a positive audio embedding from a positive audio signal wherein the positive audio signal is related to a reference image; b)generating a negative audio embedding from a negative audio signal;; c) using a machine learning algorithm with, the reference image, the positive audio embedding and the negative audio embedding as inputs to train an image-to-audio correlation neural network to output a smaller distance between the positive audio embedding and the reference image than the negative audio embedding and the reference image.
11. The system of claim 10 lurther comprising determining a positive audio signal from a noisy positive audio signal using a trained similarity neural network with an audio database and the noisy positive audio signal to select the positive audio signal from the audio database, wherein a reference audio signal is part of the audio database and wherein the reference audio is part of an audio visual sequence having the reference image.
12. The system of claim 11 wherein the positive audio signal is the reference audio signal.
13. The system of claim 11 lurther comprising determining the negative audio signal using the trained similarity neural network with the audio database wherein the negative signal is not related to the reference audio signal.
14. The system of claim 10 wherein the audio features are extracted from the positive signal and negative signal before being used in training with the machine learning algorithm.
15. The system of claim 10 wherein the positive audio signal includes noise.
16. The system of claim 10 wherein the machine learning algorithm used to train the visual- to-audio correlation neural network includes a pairwise loss function.
17. The system of claim 10 wherein the positive signal is part of an audio/video sequence that includes the reference image, wherein positive signal includes noise signals and wherein the noise signals are other sounds occurring in the audio/video sequence and wherein the negative signal includes noise signals.
18. The method of claim 10 wherein the machine learning algorithm is a self-supervised learning algorithm and wherein the positive, negative and correlated audio are unlabeled or unannotated inputs.
19. Non-transitory instructions embedded in a computer readable medium that when executed by a computer cause the computer to carry out the method for training a Sound Recommendation Network comprising: a)generating a positive audio embedding from a positive audio signal wherein the positive audio signal is related to a reference image; b)generating a negative audio embedding from a negative audio signal; b) using a machine learning algorithm with, the reference image, the positive audio embedding and a negative audio embedding as inputs to train a visual-to-audio correlation neural network to output a smaller distance between the positive audio embedding and the reference image than the negative audio embedding and the reference image.
20. The non-transitory instructions embedded in a computer readable medium of claim 19 further comprising determining a positive audio signal from a noisy positive audio signal using a trained similarity neural network with an audio database and the noisy positive audio signal to select a positive audio signal from the audio database, wherein a reference audio signal is part of the audio database and wherein the reference audio is part of an audio visual sequence having the reference image.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US16/848,484 | 2020-04-14 | ||
| US16/848,484 US11694084B2 (en) | 2020-04-14 | 2020-04-14 | Self-supervised AI-assisted sound effect recommendation for silent video |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021211366A1 true WO2021211366A1 (en) | 2021-10-21 |
Family
ID=78006074
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2021/026550 Ceased WO2021211366A1 (en) | 2020-04-14 | 2021-04-09 | Self-supervised ai-assisted sound effect recommendation for silent video |
Country Status (2)
| Country | Link |
|---|---|
| US (3) | US11694084B2 (en) |
| WO (1) | WO2021211366A1 (en) |
Families Citing this family (14)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111461235B (en) * | 2020-03-31 | 2021-07-16 | 合肥工业大学 | Audio and video data processing method, system, electronic device and storage medium |
| US11735197B2 (en) * | 2020-07-07 | 2023-08-22 | Google Llc | Machine-learned differentiable digital signal processing |
| US20220391697A1 (en) * | 2021-06-04 | 2022-12-08 | Apple Inc. | Machine-learning based gesture recognition with framework for adding user-customized gestures |
| KR20230047844A (en) * | 2021-10-01 | 2023-04-10 | 삼성전자주식회사 | Method for providing video and electronic device supporting the same |
| US20230177384A1 (en) * | 2021-12-08 | 2023-06-08 | Google Llc | Attention Bottlenecks for Multimodal Fusion |
| US12153879B2 (en) * | 2022-04-19 | 2024-11-26 | International Business Machines Corporation | Syntactic and semantic autocorrect learning |
| CN114840713B (en) * | 2022-05-14 | 2025-10-21 | 云知声智能科技股份有限公司 | Multimodal short video search method, device and storage medium |
| CN117501363A (en) * | 2022-05-30 | 2024-02-02 | 北京小米移动软件有限公司 | A sound effect control method, device and storage medium |
| CN115545117A (en) * | 2022-11-08 | 2022-12-30 | 北京有竹居网络技术有限公司 | Method, device, device and medium for generating negative sample pairs for a contrastive learning model |
| CN115713945B (en) * | 2022-11-10 | 2024-09-27 | 杭州爱华仪器有限公司 | Audio data processing method and prediction method |
| US11983923B1 (en) * | 2022-12-08 | 2024-05-14 | Netflix, Inc. | Systems and methods for active speaker detection |
| KR102624074B1 (en) * | 2023-01-04 | 2024-01-10 | 중앙대학교 산학협력단 | Apparatus and method for video representation learning |
| CN118797126A (en) * | 2023-04-03 | 2024-10-18 | 腾讯科技(深圳)有限公司 | Information recommendation method, device, computer equipment, storage medium and program product |
| CN118567604B (en) * | 2024-07-31 | 2024-11-29 | 歌尔股份有限公司 | Sound effect demonstration control method, device, equipment, platform and storage medium |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140178043A1 (en) * | 2012-12-20 | 2014-06-26 | International Business Machines Corporation | Visual summarization of video for quick understanding |
| US20190289372A1 (en) * | 2018-03-15 | 2019-09-19 | International Business Machines Corporation | Auto-curation and personalization of sports highlights |
| US20200104319A1 (en) * | 2018-09-28 | 2020-04-02 | Sony Interactive Entertainment Inc. | Sound categorization system |
Family Cites Families (53)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2010030978A2 (en) | 2008-09-15 | 2010-03-18 | Aman James A | Session automated recording together with rules based indexing, analysis and expression of content |
| KR20050050419A (en) | 2003-11-25 | 2005-05-31 | 엘지전자 주식회사 | Telephone effect sound output method for mobile station |
| US7818658B2 (en) | 2003-12-09 | 2010-10-19 | Yi-Chih Chen | Multimedia presentation system |
| US20050175315A1 (en) | 2004-02-09 | 2005-08-11 | Glenn Ewing | Electronic entertainment device |
| US9094636B1 (en) | 2005-07-14 | 2015-07-28 | Zaxcom, Inc. | Systems and methods for remotely controlling local audio devices in a virtual wireless multitrack recording system |
| CN101046956A (en) | 2006-03-28 | 2007-10-03 | 国际商业机器公司 | Interactive audio effect generating method and system |
| JP4640463B2 (en) | 2008-07-11 | 2011-03-02 | ソニー株式会社 | Playback apparatus, display method, and display program |
| US8996538B1 (en) | 2009-05-06 | 2015-03-31 | Gracenote, Inc. | Systems, methods, and apparatus for generating an audio-visual presentation using characteristics of audio, visual and symbolic media objects |
| US9384214B2 (en) | 2009-07-31 | 2016-07-05 | Yahoo! Inc. | Image similarity from disparate sources |
| US9111582B2 (en) | 2009-08-03 | 2015-08-18 | Adobe Systems Incorporated | Methods and systems for previewing content with a dynamic tag cloud |
| US9031243B2 (en) | 2009-09-28 | 2015-05-12 | iZotope, Inc. | Automatic labeling and control of audio algorithms by audio recognition |
| US8654250B2 (en) | 2010-03-30 | 2014-02-18 | Sony Corporation | Deriving visual rhythm from video signals |
| US9646209B2 (en) | 2010-08-26 | 2017-05-09 | Blast Motion Inc. | Sensor and media event detection and tagging system |
| CN102480671B (en) | 2010-11-26 | 2014-10-08 | 华为终端有限公司 | Audio processing method and device in video communication |
| US9240215B2 (en) | 2011-09-20 | 2016-01-19 | Apple Inc. | Editing operations facilitated by metadata |
| JP6150320B2 (en) | 2011-12-27 | 2017-06-21 | ソニー株式会社 | Information processing apparatus, information processing method, and program |
| GB2506399A (en) | 2012-09-28 | 2014-04-02 | Frameblast Ltd | Video clip editing system using mobile phone with touch screen |
| TWI498880B (en) | 2012-12-20 | 2015-09-01 | Univ Southern Taiwan Sci & Tec | Automatic Sentiment Classification System with Scale Sound |
| US9378611B2 (en) | 2013-02-11 | 2016-06-28 | Incredible Technologies, Inc. | Automated adjustment of audio effects in electronic game |
| US9338420B2 (en) | 2013-02-15 | 2016-05-10 | Qualcomm Incorporated | Video analysis assisted generation of multi-channel audio data |
| US9373320B1 (en) | 2013-08-21 | 2016-06-21 | Google Inc. | Systems and methods facilitating selective removal of content from a mixed audio recording |
| GB201315142D0 (en) | 2013-08-23 | 2013-10-09 | Ucl Business Plc | Audio-Visual Dialogue System and Method |
| US9736580B2 (en) | 2015-03-19 | 2017-08-15 | Intel Corporation | Acoustic camera based audio visual scene analysis |
| US10388053B1 (en) | 2015-03-27 | 2019-08-20 | Electronic Arts Inc. | System for seamless animation transition |
| GB2581032B (en) | 2015-06-22 | 2020-11-04 | Time Machine Capital Ltd | System and method for onset detection in a digital signal |
| CN105068798A (en) | 2015-07-28 | 2015-11-18 | 珠海金山网络游戏科技有限公司 | System and method for controlling volume and sound effect of game on the basis of Fmod |
| US9858967B1 (en) | 2015-09-09 | 2018-01-02 | A9.Com, Inc. | Section identification in video content |
| US20170092001A1 (en) | 2015-09-25 | 2017-03-30 | Intel Corporation | Augmented reality with off-screen motion sensing |
| US10032081B2 (en) | 2016-02-09 | 2018-07-24 | Oath Inc. | Content-based video representation |
| US20180102143A1 (en) | 2016-10-12 | 2018-04-12 | Lr Acquisition, Llc | Modification of media creation techniques and camera behavior based on sensor-driven events |
| US10459995B2 (en) | 2016-12-22 | 2019-10-29 | Shutterstock, Inc. | Search engine for processing image search queries in multiple languages |
| CN110249387B (en) | 2017-02-06 | 2021-06-08 | 柯达阿拉里斯股份有限公司 | Method for creating audio track accompanying visual image |
| CN108922551B (en) | 2017-05-16 | 2021-02-05 | 博通集成电路(上海)股份有限公司 | Circuit and method for compensating lost frame |
| US10580457B2 (en) | 2017-06-13 | 2020-03-03 | 3Play Media, Inc. | Efficient audio description systems and methods |
| US10423659B2 (en) | 2017-06-30 | 2019-09-24 | Wipro Limited | Method and system for generating a contextual audio related to an image |
| US10522186B2 (en) | 2017-07-28 | 2019-12-31 | Adobe Inc. | Apparatus, systems, and methods for integrating digital media content |
| US11856315B2 (en) | 2017-09-29 | 2023-12-26 | Apple Inc. | Media editing application with anchored timeline for captions and subtitles |
| US11335328B2 (en) * | 2017-10-27 | 2022-05-17 | Google Llc | Unsupervised learning of semantic audio representations |
| US10671854B1 (en) | 2018-04-09 | 2020-06-02 | Amazon Technologies, Inc. | Intelligent content rating determination using multi-tiered machine learning |
| JP6442102B1 (en) | 2018-05-22 | 2018-12-19 | 株式会社フランティック | Information processing system and information processing apparatus |
| US10455297B1 (en) | 2018-08-29 | 2019-10-22 | Amazon Technologies, Inc. | Customized video content summary generation and presentation |
| CN109587554B (en) | 2018-10-29 | 2021-08-03 | 百度在线网络技术(北京)有限公司 | Video data processing method, device and readable storage medium |
| US12334118B2 (en) | 2018-11-19 | 2025-06-17 | Netflix, Inc. | Techniques for identifying synchronization errors in media titles |
| GB2579208B (en) | 2018-11-23 | 2023-01-25 | Sony Interactive Entertainment Inc | Method and system for determining identifiers for tagging video frames with |
| US20200213662A1 (en) | 2018-12-31 | 2020-07-02 | Comcast Cable Communications, Llc | Environmental Data for Media Content |
| US20200242507A1 (en) | 2019-01-25 | 2020-07-30 | International Business Machines Corporation | Learning data-augmentation from unlabeled media |
| US12015637B2 (en) * | 2019-04-08 | 2024-06-18 | Pindrop Security, Inc. | Systems and methods for end-to-end architectures for voice spoofing detection |
| US11030479B2 (en) | 2019-04-30 | 2021-06-08 | Sony Interactive Entertainment Inc. | Mapping visual tags to sound tags using text similarity |
| US10847186B1 (en) | 2019-04-30 | 2020-11-24 | Sony Interactive Entertainment Inc. | Video tagging by correlating visual features to sound tags |
| US11276419B2 (en) | 2019-07-30 | 2022-03-15 | International Business Machines Corporation | Synchronized sound generation from videos |
| US11501102B2 (en) * | 2019-11-21 | 2022-11-15 | Adobe Inc. | Automated sound matching within an audio recording |
| US11615312B2 (en) | 2020-04-14 | 2023-03-28 | Sony Interactive Entertainment Inc. | Self-supervised AI-assisted sound effect generation for silent video using multimodal clustering |
| US11381888B2 (en) * | 2020-04-14 | 2022-07-05 | Sony Interactive Entertainment Inc. | AI-assisted sound effect generation for silent video |
-
2020
- 2020-04-14 US US16/848,484 patent/US11694084B2/en active Active
-
2021
- 2021-04-09 WO PCT/US2021/026550 patent/WO2021211366A1/en not_active Ceased
-
2023
- 2023-07-03 US US18/217,745 patent/US12277501B2/en active Active
-
2025
- 2025-01-08 US US19/013,693 patent/US20250165789A1/en active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140178043A1 (en) * | 2012-12-20 | 2014-06-26 | International Business Machines Corporation | Visual summarization of video for quick understanding |
| US20190289372A1 (en) * | 2018-03-15 | 2019-09-19 | International Business Machines Corporation | Auto-curation and personalization of sports highlights |
| US20200104319A1 (en) * | 2018-09-28 | 2020-04-02 | Sony Interactive Entertainment Inc. | Sound categorization system |
Also Published As
| Publication number | Publication date |
|---|---|
| US11694084B2 (en) | 2023-07-04 |
| US20210319321A1 (en) | 2021-10-14 |
| US12277501B2 (en) | 2025-04-15 |
| US20250165789A1 (en) | 2025-05-22 |
| US20230385646A1 (en) | 2023-11-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11381888B2 (en) | AI-assisted sound effect generation for silent video | |
| US12277501B2 (en) | Training a sound effect recommendation network | |
| US11615312B2 (en) | Self-supervised AI-assisted sound effect generation for silent video using multimodal clustering | |
| Wu et al. | Self-supervised sparse representation for video anomaly detection | |
| US11947593B2 (en) | Sound categorization system | |
| CN119311854B (en) | Data query method and system based on cross-modal similarity text mining | |
| Surís et al. | Cross-modal embeddings for video and audio retrieval | |
| EP3467723B1 (en) | Machine learning based network model construction method and apparatus | |
| US11876986B2 (en) | Hierarchical video encoders | |
| CN112163165A (en) | Information recommendation method, device, equipment and computer readable storage medium | |
| CN114443899A (en) | Video classification method, device, equipment and medium | |
| US20240078785A1 (en) | Method of training image representation model | |
| CN110008365A (en) | An image processing method, apparatus, device and readable storage medium | |
| Pandeya et al. | Music video emotion classification using slow–fast audio–video network and unsupervised feature representation | |
| Gayathri et al. | Classification of Speech Signal Using CNN-LSTM | |
| CN105701516B (en) | An automatic image annotation method based on attribute discrimination | |
| CN113792167B (en) | Cross-media cross-retrieval method based on attention mechanism and modal dependence | |
| Rebecca et al. | Predictive analysis of online television videos using machine learning algorithms | |
| Liang et al. | Enhancing cross-modal voice-face association with heterogeneous hashing network | |
| Zhang | Learning Robust Features for Recognition of Emotions in Images and Videos | |
| CN118673178A (en) | Training method, device, equipment and storage medium for audio and video matching model | |
| CN118235173A (en) | Pre-training of basic computer vision models | |
| Yang et al. | Markov chain models based on genetic algorithms for texture and speech recognition | |
| Wickramasinghe | Visuo-A Deep Learning Video Search Engine | |
| Nilufar | Local feature based pattern classification: from principle to application |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21789024 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 21789024 Country of ref document: EP Kind code of ref document: A1 |