EP2769377A2 - A process for assigning audio data in missing audio parts of a music piece and device for performing the same - Google Patents

A process for assigning audio data in missing audio parts of a music piece and device for performing the same

Info

Publication number
EP2769377A2
EP2769377A2 EP12818580.8A EP12818580A EP2769377A2 EP 2769377 A2 EP2769377 A2 EP 2769377A2 EP 12818580 A EP12818580 A EP 12818580A EP 2769377 A2 EP2769377 A2 EP 2769377A2
Authority
EP
European Patent Office
Prior art keywords
substring
music piece
missing
audio parts
missing audio
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP12818580.8A
Other languages
German (de)
French (fr)
Inventor
Julien ALLALI
Myriam Desainte-Catherine
Pascal FERRARO
Pierre HANNA
Benjamin Martin
Matthias ROBINE
Vinh-Thong TA
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Universite de Bordeaux
Institut Polytechnique de Bordeaux
Original Assignee
Universite de Bordeaux
Institut Polytechnique de Bordeaux
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Universite de Bordeaux, Institut Polytechnique de Bordeaux filed Critical Universite de Bordeaux
Publication of EP2769377A2 publication Critical patent/EP2769377A2/en
Withdrawn legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H1/00Details of electrophonic musical instruments
    • G10H1/0008Associated control or indicating means
    • G10H1/0025Automatic or semi-automatic music composition, e.g. producing random music, applying rules from music theory or modifying a musical piece
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2240/00Data organisation or data communication aspects, specifically adapted for electrophonic musical tools or instruments
    • G10H2240/171Transmission of musical instrument data, control or status information; Transmission, remote access or control of music data for electrophonic musical instruments
    • G10H2240/185Error prevention, detection or correction in files or streams for electrophonic musical instruments
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/005Correction of errors induced by the transmission channel, if related to the coding algorithm

Definitions

  • the invention relates generally to audio data reconstruction and more particularly to a process and device for assigning audio data in missing audio parts of a music piece.
  • Audio data reconstruction has been of major concern for audio signal processing researchers over the last decade and a vast array of solutions has been proposed.
  • Audio signals are often subject to localized audio artefacts and/or distortions, due to recording issues (unexpected noises, clips or clicks) , or to packet losses in network transmissions. This results to corrupted audio excerpts which need to be reconstructed.
  • the most audio data reconstruction systems proposed so far are based on signal features of the music pieces. These systems usually rely on the assumption that the audio signal is stationary during the missing audio part. However, the signal features do not verify this assumption since they may vary very rapidly. Therefore, it is particularly difficult to assign audio data to large missing audio parts of music pieces in order to perform audio data reconstruction. Accordingly, in the known audio data reconstruction systems, audio data reconstruction is performed on missing audio parts which are generally small (for example of a maximum duration of 1 or 2 seconds) in order to maintain a satisfactory reconstructing quality of the missing audio parts.
  • missing data of missing audio parts (substring ( ⁇ ⁇ ) ) of a music piece may be obtained from another part (reference substring ui or U2) of the music piece which has at least one neighbouring substring with musical features similar to the musical features of at least one neighbouring substring of the missing audio parts. If the at least one neighbouring substring of the missing audio parts is similar with the at least one neighbouring substring of the reference substring (ui or u ⁇ ) , the musical sections which are on their proximity are probably also similar as a result of the repetitions of musical features within the music piece.
  • the audio data of the reference substring may substitute the missing audio parts according to the process in a manner that ensures that the reconstructing quality of the missing audio parts remains satisfactory and musical consistency of the music piece is being obtained.
  • the reference substring ( ui or U2) is identified by the application of the similarity measure which compares the musical features of the at least one neighbouring substring of the missing audio parts with the musical features of the at least one neighbouring substring of the reference substring ( ui or U2) and identifies the best candidate reference substring for substituting the substring of missing audio parts.
  • the process comprises the steps of: - locating a first substring (vj) of a predetermined size immediately before the substring ( ⁇ ) of missing audio parts;
  • the reference substring (ui or U2) is identified in the string (u) of musical features representing the music piece.
  • the reference substring (ui or U2) is identified in a database of reference music pieces, which reference music pieces are different from said music piece.
  • the size of the identified reference substring is greater or lesser than the size of the substring ( ⁇ ⁇ ) of missing audio parts.
  • a step of audio data reconstruction is performed in order to ensure smooth audio data transition in the music piece if there is an overlap between the audio data substituting the substring ( ⁇ ⁇ ) of missing audio parts and the rest of the music piece.
  • the musical features comprise tonal features or rhythm features or timbre features or vocal features or lyrics-related features.
  • the music piece is represented by a symbolic music piece and the audio data is represented by symbolic music data.
  • the invention also proposes a device for assigning audio data in missing audio parts of a music piece, which device comprises :
  • an execution unit for executing musical features extraction on the music piece such that the music piece is represented by a string (u) of musical features , each musical feature being represented by a sequence or a set of one or more symbols of a predetermined alphabet;
  • a similarity measure definer unit for defining a similarity measure to measure the similarity between substrings of musical features
  • substring definer unit for defining a substring ( ⁇ ⁇ ) of missing audio parts in the string (u) of musical features
  • a locator unit for locating at least one substring (v) of a predetermined size immediately before, respectively after the substring ( ⁇ ⁇ ) of missing audio parts
  • a similarity measure application unit for applying the similarity measure for identifying a reference substring ( ui or U2) including the at least one substring (v) of a predetermined size at its beginning, respectively at its end and the reference substring ( ui or U2) best matches to the substring ( ⁇ ⁇ ) of missing audio parts;
  • substitution unit for substituting said substring ( ⁇ ⁇ ) of missing audio parts by audio data from the identified reference substring.
  • the locator unit is adapted for locating a first substring (vj) of a predetermined size immediately before the substring ( ⁇ ⁇ ) of missing audio parts and a second substring (v r ) of a predetermined size immediately after the substring ( ⁇ ⁇ ) of missing audio parts.
  • the invention also achieves a computer program with a program code for performing, when the computer program is executed on a computer, a process for assigning audio data in missing audio parts of a music piece, wherein the process comprises the steps of:
  • Figure 1 illustrates a flowchart of a process for assigning audio data in missing audio parts of a music piece according to an embodiment of the invention.
  • Figure 2 illustrates an overview of an algorithm of performing a process for assigning audio data in missing audio parts of a music piece according to an embodiment of the invention .
  • Figure 3 illustrates a device for assigning audio data in missing audio parts of a music piece according to an embodiment of the invention.
  • the process for assigning audio data in missing audio parts of a music piece comprises a first step of executing musical features extraction on a music piece.
  • the musical features comprise tonal features, rhythm features, timbre features, vocal features or lyrics-related features.
  • the musical features comprise tonal features.
  • tonal features extraction on a music piece is known in the prior art.
  • the music piece is first divided into n segments, or audio frames.
  • Constant length frames are chosen in order to optimize the proposed mono-parametric signal representation and to enable our system to be potentially used on diverse musical genres.
  • Harmonic Pitch Class Profile (HPCP) , holding its local tonal context.
  • the dimension value B stands for the precision of the note scale, or tonal resolution, usually set to 12, 24, or 36 bins.
  • Each HPCP feature is normalized by its maximum value.
  • each B-dimensional feature h is defined on [ ⁇ , ⁇ ] 5 .
  • Needleman and Wunsch in their publication "A general method applicable to the search for similarities in the amino acid sequence of two proteins. Journ . of Molecular Biology, v. 48, pp. 443-453, 1970" propose an algorithm that computes a similarity measure between a first string and a second string as a series of elementary operations needed to transform the first string into the second string, and represent the series of transformations by exhibiting an explicit alignment between strings.
  • a variant of the algorithm proposed by Needleman and Wunsch is the so-called local alignment algorithm and is presented in the publication "T. F. Smith and M. S. Waterman. Identification of common molecular subsequences. Journ. of Molecular Biology, v. 147, pp. 195-197, 1981".
  • the above mentioned local alignment algorithm allows identifying and extracting a pair of substrings, one from each of two given strings, which exhibit the highest similarity.
  • the above mentioned local alignment algorithm computes a dynamic programming matrix M such that [z][y] contains the local alignment scores between a substring «[1 ⁇ ⁇ ⁇ /] and a substring v[l - - - y] .
  • ,Y/ l ...
  • part (or) of equation 1 represents the deletion of the symbol u[i]
  • part ( ⁇ ) of equation (1) represents the insertion of the symbol v[j]
  • part ( ⁇ ) of equation (1) represents the substitution of the symbol u[i] by the symbol v[j] .
  • three scores are defined in equation (1) .
  • the particular values assigned to these scores form the scoring scheme of the alignment.
  • the i th symbol of u is denoted by u[i], and u can be written as a concatenation of its symbols w[l]w[2] - ⁇ ⁇ or w[l - - -
  • the j th symbol of v is denoted by v[y]
  • v can be written as a concatenation of its symbols v[l]v[2] ⁇ ⁇ ⁇ [jv
  • the music piece is represented by a string u of HPCP features.
  • Each HPCP feature is represented by a sequence of one or more symbols defined on an alphabet ⁇ .
  • a string of specific symbols (also known as "joker" symbols) in order to define missing audio parts of the music piece.
  • [0,lf U ⁇ (/> ⁇ ⁇
  • ⁇ * the set of all possible substrings whose symbols are defined on ⁇ .
  • a similarity measure is defined between substrings of musical features in order to measure the similarity of the latter.
  • This similarity measure can be derived from the maximum value of the dynamic programming matrix M computed by equation (1) .
  • the symbols u[i] and v[j] of equation (1) are respectively replaced by two HPCP features h 1 and h J .
  • the two HPCP features h 1 and h J have been extracted during the tonal features extraction performed in the music piece (see point 1.1 of the theory of operation) .
  • the score for deleting or inserting a HPCP feature h 1 is denoted by the function C g (h 1 ) according to which :
  • s(h' , h J ) denotes a similarity measure between the HPCP features H and h j .
  • the similarity measure s(h' , h J ) may correspond to the well known Euclidean-based similarity measure for comparing features.
  • Figure 1 illustrates an embodiment of a process for assigning audio data in missing audio parts of a music piece.
  • musical features extraction is performed on the music piece such that the music piece is represented by a string (u) of musical features and each musical feature is represented by a sequence or a set of one or more symbols of a predetermined alphabet.
  • the musical features comprise tonal features, timbre features or rhythm features.
  • the musical features may comprise vocal features or lyrics-related features. It is important to note that some features (for example the lyrics-related features) may be obtained by the user of the system without performing any musical features extraction.
  • the musical features extraction is known to the person skilled in the art and a particular example of musical features extraction is provided in point 1.1 of the theory of operation.
  • the musical features comprise tonal features and after the performance of musical features extraction the music piece is represented by a string u of HPCP features.
  • the musical features extraction is performed by an execution unit 110 (see Figure 3) .
  • a similarity measure is defined in order to measure the similarity between substrings of musical features .
  • This similarity measure can be directly derived from the maximum value of the dynamic programming matrix M of equation (1) .
  • the symbols u[i] and v[j] of equation (1) are respectively replaced by two HPCP features h 1 and h J being extracted during the tonal features extraction .
  • the similarity measure definition is performed by a similarity measure definer unit 120 (see Figure 3) .
  • a substring ( ⁇ ⁇ ) of missing audio parts is defined in said string (u) of musical features.
  • the definition of the substring ( ⁇ ⁇ ) of missing audio parts is succeeded by introducing specific symbols (also known as
  • the substring ( ⁇ ⁇ ) is defined by the user by means of a substring definer unit 130 (see Figure 3) .
  • a step 40 at least one substring (v) of a predetermined size is located immediately before or respectively after said substring ( ⁇ ⁇ ) of missing audio parts .
  • step 40 is achieved by a locator unit 140 which is connected to the execution unit 110 and to the substring unit 130 (see Figure 3) .
  • the similarity measure is applied for identifying a reference substring ( ui or U2) including said at least one substring (v) of a predetermined size at its beginning or respectively at its end, said reference substring ( ui or U2) best matching to the substring ( ⁇ ⁇ ) of missing audio parts .
  • the similarity measure compares the musical features (for example HPCP features) of the at least one neighbouring substring of the missing audio parts with the musical features of the at least one neighbouring substring of the reference substring (ui or u ⁇ ) .
  • the similarity measure is important since it identifies in the string u of the music piece the best candidate reference substring between the reference substring ui and the reference substring u ⁇ , which best candidate reference substring, substitutes the substring of missing audio parts.
  • step 50 is achieved by a similarity measure application unit 150 being connected to the locator unit 140 and to the similarity measure definer unit 120 (see Figure 3) .
  • step 60 the substring ( ⁇ ⁇ ) of missing audio parts is substituted by audio data from the identified reference substring ( ui or U2) .
  • This substitution may be performed by known audio data mixing processes being applied to two audio parts of a music piece.
  • step 60 is achieved by a substitution unit 160 being connected with the similarity measure application unit 150 (see Figure 3) . It is important to note that the above mentioned process allows audio data assignment in large missing audio parts (for example of duration up to 16 seconds) of a music piece and obtains musical consistency to the latter since it is based on musical features.
  • missing data of missing audio parts may be obtained from another part (reference substring ui or U2) of the music piece which has at least one neighbouring substring (located immediately before or respectively after the neighbouring substring) with musical features similar to the musical features of at least one neighbouring substring of the missing audio parts.
  • the at least one neighbouring substring of the missing audio parts is located immediately before or respectively after the missing audio parts.
  • the musical features of the reference substrings may substitute the missing audio parts according to the process in a manner that ensures that the reconstructing quality of the missing audio parts remains satisfactory and musical consistency of the music piece is being obtained.
  • the process may identify a reference substring (ul or u2) which corresponds to that chorus by applying a similarity measure between the musical features which neighbour with the missing audio part and the musical features which neighbour with the reference substring Then, the missing audio parts may be filled by the chorus obtaining a musical consistency in the music piece.
  • the process of the invention is particularly important for the reconstruction of large missing audio parts. However, it is obvious for the person skilled in the art that the process may be applied for reconstructing smaller missing audio parts, for example of a duration of 1 or 2 seconds.
  • the process locates a first substring (vj) of a predetermined size immediately before the substring ( ⁇ ⁇ ) of missing audio parts and a second substring (v r ) of a predetermined size immediately after the substring ( ⁇ ⁇ ) of missing audio parts .
  • This may be performed by the locator unit of Figure 3.
  • the process applies the similarity measure, which can be derived from the maximum value of the dynamic programming matrix M of equation (1), as explained in point 1.2 of the theory of operation, for identifying the reference substring ( ui or U2) including the first substring (vj) at its beginning and the second substring (v r ) at its end and best matching to the substring ( ⁇ ⁇ ) of missing audio parts.
  • the size of the identified reference substring may be greater or lesser than the size of said substring ( ⁇ ⁇ ) of missing audio parts.
  • the size of the identified reference substring is greater than the size of said substring ( ⁇ ⁇ ) of missing audio parts
  • an overlap between the audio data substituting the substring ( ⁇ ⁇ ) of missing audio parts and the rest of the music piece is observed.
  • Such overlap is illustrated by rectangles being depicted by dotted lines in the horizontal section of Figure 2 denoted by (b) .
  • the horizontal section of Figure 2 denoted by (b) illustrates an audio wave-form of the reconstructed music piece, wherein the substring ( ⁇ ⁇ ) of missing audio parts is substituted by audio data from the identified reference substring u ⁇ .
  • a step of audio data reconstruction is performed in order to ensure smooth audio data transition of the music piece in the area where the above mentioned overlap is observed.
  • the audio data reconstruction may be performed by a typical audio data mixing process involving fade-in and fade-out processing between two audio parts and being well known to a person skilled in the art.
  • the reference substring ( ui or U2) is identified in the string ( u ) of musical features representing the music piece or alternatively in a database of reference music pieces, these reference music pieces being different from the music piece.
  • the identification of the reference substring ( ui or U2) in a database of reference music pieces is obvious for the person skilled in the art.
  • the music piece is represented by a symbolic music piece and said audio data is represented by symbolic music data.
  • said audio data is represented by symbolic music data.
  • Figure 2 illustrates an overview of an algorithm of the process according to an embodiment wherein a first substring (vj) of a predetermined size is located immediately before the substring ( ⁇ ⁇ ) of missing audio parts and a second substring (v r ) of a predetermined size is located immediately after said substring ( ⁇ ⁇ ) of missing audio parts .
  • the similarity measure of the process identifies the reference substring ( ui or U2) which best matches to the substring ( ⁇ ⁇ ) of missing audio parts and the substring ( ⁇ ⁇ ) of missing audio parts is substituted by audio data from the identified reference substring ( ui or U2) .
  • the music piece is shown in an audio wave-form having an intermediate part of missing audio data.
  • the intermediate part of missing audio data is located between an initial and a final part of audio data. This intermediate part of missing audio data corresponds to the substring of missing audio parts ⁇ ⁇ .
  • the music piece is represented by a string (u) of musical features and each musical feature is represented by a sequence or a set of one or more symbols of a predetermined alphabet.
  • a musical feature can be an HPCP feature obtained from the performance of musical features extraction.
  • the algorithm computes a triplet , as the result of aligning substring ti with substring vjv v r (align 1 in the horizontal section of Figure 2 denoted by (ii) ) and a triplet (x ⁇ , U2 , v ⁇ ) as a result of aligning substring t ⁇ with substring vjV V r .
  • the triplet is computed according to the maximum value of the dynamic programming matrix M and the alignment of the substrings is performed by applying the known local alignment algorithm (see point 1.2 of the theory of operation) .
  • xi of the triplet denotes the similarity measure (highest similarity) between musical features of substrings Ui, vi and x ⁇ of the triplet ⁇ x ⁇ , u ⁇ , v ⁇ ) denotes the similarity measure (highest similarity) between musical features of substrings u ⁇ , v ⁇ .
  • the algorithm compares xi and x ⁇ and if xi ⁇ X2 then it identifies the reference substring u ⁇ . Otherwise, if xi > X2 it identifies the reference substring ui . In the horizontal section of Figure 2, after the comparison of xi and X2 , the algorithm identifies the reference substring u ⁇ .
  • the algorithm substitutes the section that corresponds to the substring of missing audio parts ( ⁇ ⁇ ) by audio data by the section that corresponds to the identified reference substring u ⁇ .
  • the section that corresponds to the identified reference substring u ⁇ is illustrated in a rectangle.
  • An arrow is used to depict the substitution of the substring of missing audio parts ( ⁇ ⁇ ) by audio data from the identified reference substring u ⁇
  • the algorithm of the process described above achieves to substitute the intermediate part of missing audio data which corresponds to the substring ( ⁇ ⁇ ) of missing audio parts, by audio data of another part of the music piece which corresponds to the identified reference substring U2 (see the horizontal section of Figure 2 denoted by (b) ) .
  • the process for assigning audio data in missing audio parts of a music piece as illustrated in the embodiment of Figure 1 may be realized as a computer program with a program code for performing audio data assignment in missing audio parts when the computer program is executed on a computer.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Auxiliary Devices For Music (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

A process for assigning audio data in missing audio parts of a music piece, comprising the steps of: - execution of musical features extraction on the music piece such that the music piece is represented by a string (u) of musical features, each musical feature being represented by a sequence or a set of one or more symbols of a predetermined alphabet; - definition of a similarity measure for measuring the similarity between substrings of musical features; - definition of a substring (νφ) of missing audio parts in the string (u) of musical features; - location of at least one substring (v) of a predetermined size immediately before, respectively after the substring (νφ) of missing audio parts; - application of the similarity measure for identifying a reference substring (ui or U2) including the at least one substring (v) of a predetermined size at its beginning, respectively at its end, the reference substring (ui or U2) best matching to the substring (νφ) of missing audio parts; - substitution of said substring (νφ) of missing audio parts by audio data from the identified reference substring.

Description

A process for assigning audio data in missing audio parts of a music piece and device for performing the same
Technical field of the invention
The invention relates generally to audio data reconstruction and more particularly to a process and device for assigning audio data in missing audio parts of a music piece.
Background art
Audio data reconstruction has been of major concern for audio signal processing researchers over the last decade and a vast array of solutions has been proposed.
Audio signals are often subject to localized audio artefacts and/or distortions, due to recording issues (unexpected noises, clips or clicks) , or to packet losses in network transmissions. This results to corrupted audio excerpts which need to be reconstructed.
Reconstructing missing audio data in corrupted audio excerpts in order to provide consistent audio signals has thus been challenging for applicative research in order to reconstruct missing audio parts of audio pieces and in particular missing audio parts of music pieces. The problem of missing audio data reconstruction is usually addressed either in the time domain, aiming at recovering missing excerpts in audio pieces, or in the time-frequency domain, aiming at recovering missing frequencies that cause localized distortions of audio pieces. A typical trend for the latter one, often referred to as audio inpainting, is to treat distorted samples as missing and to attempt restore original ones from a local analysis around missing audio parts. Common approaches include linear prediction for sinusoidal models, Bayesian estimators, autoregressive models or non-negative matrix factorization solving. These studies can roughly be summarized as either being based on the distributions of data features around missing audio parts or employing local or global statistical characteristics over audio excerpts.
Particularly in the case of missing audio data in music pieces, the most audio data reconstruction systems proposed so far are based on signal features of the music pieces. These systems usually rely on the assumption that the audio signal is stationary during the missing audio part. However, the signal features do not verify this assumption since they may vary very rapidly. Therefore, it is particularly difficult to assign audio data to large missing audio parts of music pieces in order to perform audio data reconstruction. Accordingly, in the known audio data reconstruction systems, audio data reconstruction is performed on missing audio parts which are generally small (for example of a maximum duration of 1 or 2 seconds) in order to maintain a satisfactory reconstructing quality of the missing audio parts. Thus, there is a need of reconstructing large missing audio parts of a music piece, namely audio gaps with a duration of several seconds (for example up to 16 seconds), by performing a process of audio data assignment which ensures that the reconstructing quality of the missing audio parts remains satisfactory and musical consistency of the music piece is being obtained.
Summary of the invention
It is an object of the present invention to provide a process for assigning audio data to large missing audio parts of a music piece said process ensuring that the reconstructing quality of the missing audio parts remains satisfactory and musical consistency of the music piece is being obtained.
It is another object of the present invention to provide a device for assigning audio data to large missing audio parts of a music piece by applying a data assignment process according to the invention.
These and other objects of the invention are achieved by means of a process for assigning audio data in missing audio parts of a music piece, comprising the steps of:
- executing musical features extraction on the music piece such that the music piece is represented by a string (u) of musical features, each musical feature being represented by a sequence or a set of one or more symbols of a predetermined alphabet; - defining a similarity measure to measure the similarity between substrings of musical features;
- defining a substring (νφ) of missing audio parts in the string (u) of musical features;
- locating at least one substring (v) of a predetermined size immediately before, respectively after the substring (νφ) of missing audio parts;
- applying the similarity measure for identifying a reference substring ( ui or U2) including the at least one substring (v) of a predetermined size at its beginning, respectively at its end, the reference substring ( ui or U2) best matching to the substring (νφ) of missing audio parts;
- substituting said substring (νφ) of missing audio parts by audio data from the identified reference substring. It is important to note that the above mentioned process allows audio data assignment in large missing audio parts (for example of duration up to 16 seconds) of a music piece and obtains musical consistency to the latter since it is based on musical features. The musical features do not present significant variations as it is the case with the signal features in the systems of prior art. Moreover they are temporally organized such that they usually present repetitions (recurrent musical sections like the chorus of a music piece) in the music piece. Particularly, thanks to the above mentioned repetitions of the musical features, missing data of missing audio parts (substring (νφ) ) of a music piece may be obtained from another part (reference substring ui or U2) of the music piece which has at least one neighbouring substring with musical features similar to the musical features of at least one neighbouring substring of the missing audio parts. If the at least one neighbouring substring of the missing audio parts is similar with the at least one neighbouring substring of the reference substring (ui or u) , the musical sections which are on their proximity are probably also similar as a result of the repetitions of musical features within the music piece. Thus, the audio data of the reference substring may substitute the missing audio parts according to the process in a manner that ensures that the reconstructing quality of the missing audio parts remains satisfactory and musical consistency of the music piece is being obtained.
It is important to note that the reference substring ( ui or U2) is identified by the application of the similarity measure which compares the musical features of the at least one neighbouring substring of the missing audio parts with the musical features of the at least one neighbouring substring of the reference substring ( ui or U2) and identifies the best candidate reference substring for substituting the substring of missing audio parts.
In one embodiment, the process comprises the steps of: - locating a first substring (vj) of a predetermined size immediately before the substring (ν ) of missing audio parts;
- locating a second substring (vr) of a predetermined size immediately after the substring (νφ) of missing audio parts;
- applying said similarity measure for identifying said reference substring (ui or U2) including the first substring (vj) at its beginning and the second substring (vr) at its end, the reference substring (ui or U2) best matching to the substring (νφ) of missing audio parts, said steps being executed sequentially or in parallel.
In another embodiment, the reference substring (ui or U2) is identified in the string (u) of musical features representing the music piece.
In a further embodiment, the reference substring (ui or U2) is identified in a database of reference music pieces, which reference music pieces are different from said music piece.
In another embodiment, the size of the identified reference substring is greater or lesser than the size of the substring (νφ) of missing audio parts.
In another embodiment, a step of audio data reconstruction is performed in order to ensure smooth audio data transition in the music piece if there is an overlap between the audio data substituting the substring (νφ) of missing audio parts and the rest of the music piece. In a further embodiment, the musical features comprise tonal features or rhythm features or timbre features or vocal features or lyrics-related features. In another embodiment, the music piece is represented by a symbolic music piece and the audio data is represented by symbolic music data.
The invention also proposes a device for assigning audio data in missing audio parts of a music piece, which device comprises :
- an execution unit for executing musical features extraction on the music piece such that the music piece is represented by a string (u) of musical features , each musical feature being represented by a sequence or a set of one or more symbols of a predetermined alphabet;
- a similarity measure definer unit for defining a similarity measure to measure the similarity between substrings of musical features;
- a substring definer unit for defining a substring (νφ) of missing audio parts in the string (u) of musical features;
- a locator unit for locating at least one substring (v) of a predetermined size immediately before, respectively after the substring (νφ) of missing audio parts; - a similarity measure application unit for applying the similarity measure for identifying a reference substring ( ui or U2) including the at least one substring (v) of a predetermined size at its beginning, respectively at its end and the reference substring ( ui or U2) best matches to the substring (νφ) of missing audio parts;
- a substitution unit for substituting said substring (νφ) of missing audio parts by audio data from the identified reference substring.
In one embodiment, the locator unit is adapted for locating a first substring (vj) of a predetermined size immediately before the substring (νφ) of missing audio parts and a second substring (vr) of a predetermined size immediately after the substring (νφ) of missing audio parts.
The invention also achieves a computer program with a program code for performing, when the computer program is executed on a computer, a process for assigning audio data in missing audio parts of a music piece, wherein the process comprises the steps of:
- executing musical features extraction on the music piece such that the music piece is represented by a string (u) of musical features, each musical feature being represented by a sequence or a set of one or more symbols of a predetermined alphabet ;
- defining a similarity measure to measure the similarity between substrings of musical features; - defining a substring (ν ) of missing audio parts in the string (u) of musical features;
- locating at least one substring (v) of a predetermined size immediately before, respectively after the substring (νφ) of missing audio parts;
- applying the similarity measure for identifying a reference substring (ui or U2) including the at least one substring (v) of a predetermined size at its beginning, respectively at its end, the reference substring (ui or U2) best matching to the substring (νφ) of missing audio parts;
- substituting said substring (νφ) of missing audio parts by audio data from the identified reference substring.
Brief description of the drawings
The above objects and characteristics of the present invention will be more apparent by describing an/several embodiments of the present invention in detail with reference to the accompanying drawings, in which:
Figure 1 illustrates a flowchart of a process for assigning audio data in missing audio parts of a music piece according to an embodiment of the invention.
Figure 2 illustrates an overview of an algorithm of performing a process for assigning audio data in missing audio parts of a music piece according to an embodiment of the invention .
Figure 3 illustrates a device for assigning audio data in missing audio parts of a music piece according to an embodiment of the invention.
Theory of operation 1.1 Tonal features extraction
As it is illustrated in the description of the invention, the process for assigning audio data in missing audio parts of a music piece comprises a first step of executing musical features extraction on a music piece.
In different embodiments, the musical features comprise tonal features, rhythm features, timbre features, vocal features or lyrics-related features.
In the theory of operation disclosed below, it is considered that the musical features comprise tonal features.
Performing tonal features extraction on a music piece is known in the prior art. Particularly, in tonal features extraction, the music piece is first divided into n segments, or audio frames. Constant length frames are chosen in order to optimize the proposed mono-parametric signal representation and to enable our system to be potentially used on diverse musical genres. Each frame is represented by a B-dimensional feature h = {hy,...,hB) that corresponds to a
Harmonic Pitch Class Profile (HPCP) , holding its local tonal context. The dimension value B stands for the precision of the note scale, or tonal resolution, usually set to 12, 24, or 36 bins. Each HPCP feature is normalized by its maximum value. Thus, each B-dimensional feature h is defined on [θ,ΐ]5 . Hence, the music piece can be represented as a sequence u = hlh2 - - - h" of n B-dimensional features. Accordingly, the performance of musical features extraction in the first step of the process of the invention provides a music piece that is represented by a string u of HPCP features.
It is important to note that the sequence of tonal features representing the music piece is considered as a string and allows the application of string matching techniques .
1.2 String matching techniques
Needleman and Wunsch, in their publication "A general method applicable to the search for similarities in the amino acid sequence of two proteins. Journ . of Molecular Biology, v. 48, pp. 443-453, 1970" propose an algorithm that computes a similarity measure between a first string and a second string as a series of elementary operations needed to transform the first string into the second string, and represent the series of transformations by exhibiting an explicit alignment between strings. A variant of the algorithm proposed by Needleman and Wunsch is the so-called local alignment algorithm and is presented in the publication "T. F. Smith and M. S. Waterman. Identification of common molecular subsequences. Journ. of Molecular Biology, v. 147, pp. 195-197, 1981".
The above mentioned local alignment algorithm allows identifying and extracting a pair of substrings, one from each of two given strings, which exhibit the highest similarity.
The above mentioned local alignment algorithm computes a dynamic programming matrix M such that [z][y] contains the local alignment scores between a substring «[1 · · · /] and a substring v[l - - - y] . The dynamic programming matrix M is computed by the following equation: ¾ ( ) wherein u and v represent the two strings to be compared with the initial condition M[i][0] = M[0][j] = 0,Vz = l ...|w|,Y/ = l ...|v|.
Furthermore, part (or) of equation 1 represents the deletion of the symbol u[i] , part (β) of equation (1) represents the insertion of the symbol v[j] , and part ( γ) of equation (1) represents the substitution of the symbol u[i] by the symbol v[j] . Also, in order to evaluate the score of an alignment, three scores are defined in equation (1) . A first score for substituting a symbol u[i] by another symbol v[j] denoted by the function Cm(u[i], v[j] ) , a second score for deleting a symbol u[i] denoted by the function Cg ( u[i] ) and a third score for inserting a symbol v[j] denoted by the function Cg(v[j]). The particular values assigned to these scores form the scoring scheme of the alignment.
Additionally, in equation (1), the ith symbol of u is denoted by u[i], and u can be written as a concatenation of its symbols w[l]w[2] - · · or w[l - - - |w|] where is the length of the string u . Similarly, the jth symbol of v is denoted by v[y] , and v can be written as a concatenation of its symbols v[l]v[2] · · · [jv|] or v[l - - - |v|] where|v| is the length of the string v .
In the process of the invention, after the tonal features extraction performed on the music piece (see point 1.1 of the theory of operation) the music piece is represented by a string u of HPCP features. Each HPCP feature is represented by a sequence of one or more symbols defined on an alphabet ∑ . In the process of the invention we introduce a string of specific symbols (also known as "joker" symbols) in order to define missing audio parts of the music piece. Thus, the alphabet considered in our context is denoted by ∑ = [0,lf U {(/>} · We denote by ∑* the set of all possible substrings whose symbols are defined on ∑ .
In the process of the invention, a similarity measure is defined between substrings of musical features in order to measure the similarity of the latter. This similarity measure can be derived from the maximum value of the dynamic programming matrix M computed by equation (1) . In order to compute the maximum value of the dynamic programming matrix M when HPCP features extraction is performed in the music piece, the symbols u[i] and v[j] of equation (1) are respectively replaced by two HPCP features h1 and hJ . The two HPCP features h1 and hJ have been extracted during the tonal features extraction performed in the music piece (see point 1.1 of the theory of operation) .
It is important to note that the known local alignment algorithm of Smith and Waterman results in a triplet {w, wlr w) according to the maximum value of the dynamic programming matrix M . In the triplet {w, wlf w) , wi and w denote the pair of substrings identified by applying the local alignment algorithm. Moreover, w denotes the similarity measure (highest similarity) between substrings wj and w, which corresponds to the similarity measure applied by the process of the invention.
In an example of the application of the similarity measure of the invention taking into consideration the comparison of two HPCP features h1 and hJ and also the introduction of a string of specific symbols (also known as "joker" symbols) for defining the missing audio parts of a music piece, the score for deleting or inserting a HPCP feature h1 is denoted by the function Cg (h1) according to which :
CgCh =-0.7 if Ιι1≠φ , 0 otherwise,
while the score for substituting a HPCP feature h1 by another feature hJ is denoted by the function Cmih1 , hJ) according to which :
wherein s(h' , hJ) denotes a similarity measure between the HPCP features H and hj . In an example, the similarity measure s(h' , hJ) may correspond to the well known Euclidean-based similarity measure for comparing features.
Description of the embodiments of the invention
Figure 1 illustrates an embodiment of a process for assigning audio data in missing audio parts of a music piece.
In a step 10 of the process, musical features extraction is performed on the music piece such that the music piece is represented by a string (u) of musical features and each musical feature is represented by a sequence or a set of one or more symbols of a predetermined alphabet. different embodiments the musical features comprise tonal features, timbre features or rhythm features. Alternatively, the musical features may comprise vocal features or lyrics-related features. It is important to note that some features (for example the lyrics-related features) may be obtained by the user of the system without performing any musical features extraction.
The musical features extraction is known to the person skilled in the art and a particular example of musical features extraction is provided in point 1.1 of the theory of operation. In that example the musical features comprise tonal features and after the performance of musical features extraction the music piece is represented by a string u of HPCP features. However, it is known to the person skilled in the art how to perform musical features extraction for musical features comprising rhythm features, timbre features or vocal features .
In an embodiment, the musical features extraction is performed by an execution unit 110 (see Figure 3) .
In a step 20, a similarity measure is defined in order to measure the similarity between substrings of musical features . This similarity measure can be directly derived from the maximum value of the dynamic programming matrix M of equation (1) . In order to compute the maximum value of the dynamic programming matrix M , when for example HPCP features extraction is performed in the music piece, the symbols u[i] and v[j] of equation (1) are respectively replaced by two HPCP features h1 and hJ being extracted during the tonal features extraction .
In an embodiment, the similarity measure definition is performed by a similarity measure definer unit 120 (see Figure 3) .
In a step 30, a substring (νφ) of missing audio parts is defined in said string (u) of musical features. The definition of the substring (νφ) of missing audio parts is succeeded by introducing specific symbols (also known as
"joker" symbols) in the substring (νφ) of missing audio parts. It is important to note that such introduction of specific symbols φ is not performed in the known systems for performing audio data reconstruction in missing audio parts of music pieces.
In an embodiment, the substring (νφ) is defined by the user by means of a substring definer unit 130 (see Figure 3) . In a step 40, at least one substring (v) of a predetermined size is located immediately before or respectively after said substring (νφ) of missing audio parts .
In an embodiment, step 40 is achieved by a locator unit 140 which is connected to the execution unit 110 and to the substring unit 130 (see Figure 3) .
In a step 50, the similarity measure is applied for identifying a reference substring ( ui or U2) including said at least one substring (v) of a predetermined size at its beginning or respectively at its end, said reference substring ( ui or U2) best matching to the substring (νφ) of missing audio parts .
It is important to note that the similarity measure compares the musical features (for example HPCP features) of the at least one neighbouring substring of the missing audio parts with the musical features of the at least one neighbouring substring of the reference substring (ui or u) . The similarity measure is important since it identifies in the string u of the music piece the best candidate reference substring between the reference substring ui and the reference substring u, which best candidate reference substring, substitutes the substring of missing audio parts.
In an embodiment, step 50 is achieved by a similarity measure application unit 150 being connected to the locator unit 140 and to the similarity measure definer unit 120 (see Figure 3) .
It is important to note that the steps 40 and 50 may be executed sequentially or in parallel. In a step 60, the substring (νφ) of missing audio parts is substituted by audio data from the identified reference substring ( ui or U2) . This substitution may be performed by known audio data mixing processes being applied to two audio parts of a music piece. In an embodiment, step 60 is achieved by a substitution unit 160 being connected with the similarity measure application unit 150 (see Figure 3) . It is important to note that the above mentioned process allows audio data assignment in large missing audio parts (for example of duration up to 16 seconds) of a music piece and obtains musical consistency to the latter since it is based on musical features. The musical features do not present significant variations as is the case with the signal features in the systems of prior art and they are temporally organized such that they usually present repetitions (recurrent musical sections like the chorus of a music piece) in the music piece. Particularly, thanks to the above mentioned repetitions of the musical features, missing data of missing audio parts (substring (νφ) ) may be obtained from another part (reference substring ui or U2) of the music piece which has at least one neighbouring substring (located immediately before or respectively after the neighbouring substring) with musical features similar to the musical features of at least one neighbouring substring of the missing audio parts. The at least one neighbouring substring of the missing audio parts is located immediately before or respectively after the missing audio parts.
If the at least one neighbouring substring of the missing audio parts is similar with the at least one neighbouring substring of the reference substring (ui or u) , the musical sections which are on their proximity are probably also similar as a result of the repetitions of musical features within the music piece. Thus, the musical features of the reference substrings may substitute the missing audio parts according to the process in a manner that ensures that the reconstructing quality of the missing audio parts remains satisfactory and musical consistency of the music piece is being obtained.
As an example, if the missing audio parts correspond to a chorus of a music piece, the process may identify a reference substring (ul or u2) which corresponds to that chorus by applying a similarity measure between the musical features which neighbour with the missing audio part and the musical features which neighbour with the reference substring Then, the missing audio parts may be filled by the chorus obtaining a musical consistency in the music piece.
The process of the invention is particularly important for the reconstruction of large missing audio parts. However, it is obvious for the person skilled in the art that the process may be applied for reconstructing smaller missing audio parts, for example of a duration of 1 or 2 seconds.
In an embodiment, the process locates a first substring (vj) of a predetermined size immediately before the substring (νφ) of missing audio parts and a second substring (vr) of a predetermined size immediately after the substring (νφ) of missing audio parts . This may be performed by the locator unit of Figure 3. Then, the process applies the similarity measure, which can be derived from the maximum value of the dynamic programming matrix M of equation (1), as explained in point 1.2 of the theory of operation, for identifying the reference substring ( ui or U2) including the first substring (vj) at its beginning and the second substring (vr) at its end and best matching to the substring (νφ) of missing audio parts. The above mentioned steps of location of the first substring (vj) and the second substring (νφ) as well as the step of application of the similarity measure are performed sequentially or in parallel. Then the substring (νφ) of missing audio parts is substituted by audio data from the identified reference substring ( ui or U2) .
It is important to note that the size of the identified reference substring may be greater or lesser than the size of said substring (νφ) of missing audio parts.
Particularly in the case that the size of the identified reference substring is greater than the size of said substring (νφ) of missing audio parts, an overlap between the audio data substituting the substring (νφ) of missing audio parts and the rest of the music piece is observed. Such overlap is illustrated by rectangles being depicted by dotted lines in the horizontal section of Figure 2 denoted by (b) . As it will be explained below, the horizontal section of Figure 2 denoted by (b) illustrates an audio wave-form of the reconstructed music piece, wherein the substring (νφ) of missing audio parts is substituted by audio data from the identified reference substring u. In an embodiment, a step of audio data reconstruction is performed in order to ensure smooth audio data transition of the music piece in the area where the above mentioned overlap is observed. The audio data reconstruction may be performed by a typical audio data mixing process involving fade-in and fade-out processing between two audio parts and being well known to a person skilled in the art.
Also, it is important to note that the reference substring ( ui or U2) is identified in the string ( u ) of musical features representing the music piece or alternatively in a database of reference music pieces, these reference music pieces being different from the music piece. The identification of the reference substring ( ui or U2) in a database of reference music pieces is obvious for the person skilled in the art.
In another embodiment, the music piece is represented by a symbolic music piece and said audio data is represented by symbolic music data. Performing the process of assigning symbolic music data in missing audio parts of the symbolic music piece is obvious for the person skilled in the art. It is important to note that the symbolic music data is obtained by the user of the system without performing musical features extraction .
Figure 2 illustrates an overview of an algorithm of the process according to an embodiment wherein a first substring (vj) of a predetermined size is located immediately before the substring (νφ) of missing audio parts and a second substring (vr) of a predetermined size is located immediately after said substring (νφ) of missing audio parts . As illustrated in Figure 2, the similarity measure of the process, identifies the reference substring ( ui or U2) which best matches to the substring (νφ) of missing audio parts and the substring (νφ) of missing audio parts is substituted by audio data from the identified reference substring ( ui or U2) .
More particularly, in the horizontal section of Figure 2 denoted by (a) , the music piece is shown in an audio wave-form having an intermediate part of missing audio data. The intermediate part of missing audio data is located between an initial and a final part of audio data. This intermediate part of missing audio data corresponds to the substring of missing audio parts νφ.
In a first step of the algorithm (not depicted in Figure 2) , the music piece is represented by a string (u) of musical features and each musical feature is represented by a sequence or a set of one or more symbols of a predetermined alphabet. For example, a musical feature can be an HPCP feature obtained from the performance of musical features extraction.
In a second step, as depicted in the horizontal section of Figure 2 denoted by (i) , a substring of missing audio parts νφ = φ···φ, wherein φ are specific symbols ("joker symbols"), are introduced in the string (u) for defining the substring of missing audio parts. Also, in the horizontal section of Figure 2 denoted by (i) , the above mentioned initial and final part of audio data depicted in the horizontal section of Figure 2 denoted by (a), is represented by substrings tlf t∑ in ∑* , such that u= tivtt2.
In a third step (as shown in the horizontal section of Figure 2 denoted by (ii) ) , the algorithm locates a first substring (vj) of a predetermined size immediately before the substring (νφ) of missing audio parts and a second substring (vr) of a predetermined size immediately after said substring (νφ) of missing audio parts, such that there exist substrings ti 'and t∑f n ∑* for which ti=ti'' vj and t2=vrt2f . Then the algorithm computes a triplet , as the result of aligning substring ti with substring vjv vr (align 1 in the horizontal section of Figure 2 denoted by (ii) ) and a triplet (x∑ , U2 , v∑) as a result of aligning substring t∑ with substring vjV Vr. The triplet is computed according to the maximum value of the dynamic programming matrix M and the alignment of the substrings is performed by applying the known local alignment algorithm (see point 1.2 of the theory of operation) . Moreover, xi of the triplet ( xi ,ui ,vi) denotes the similarity measure (highest similarity) between musical features of substrings Ui, vi and x∑ of the triplet {x∑, u∑, v∑) denotes the similarity measure (highest similarity) between musical features of substrings u∑, v∑. In a third step (see the horizontal section of Figure
2 denoted by (iii) ) , the algorithm compares xi and x∑ and if xi < X2 then it identifies the reference substring u∑. Otherwise, if xi > X2 it identifies the reference substring ui . In the horizontal section of Figure 2, after the comparison of xi and X2 , the algorithm identifies the reference substring u.
In a fourth step, as shown in the horizontal section of Figure 2 denoted by (b) which illustrates an audio wave-form of the reconstructed music piece, the algorithm substitutes the section that corresponds to the substring of missing audio parts (νφ) by audio data by the section that corresponds to the identified reference substring u. The section that corresponds to the identified reference substring u is illustrated in a rectangle. An arrow is used to depict the substitution of the substring of missing audio parts (νφ) by audio data from the identified reference substring u
Thus, the algorithm of the process described above achieves to substitute the intermediate part of missing audio data which corresponds to the substring (νφ) of missing audio parts, by audio data of another part of the music piece which corresponds to the identified reference substring U2 (see the horizontal section of Figure 2 denoted by (b) ) .
The process for assigning audio data in missing audio parts of a music piece as illustrated in the embodiment of Figure 1 may be realized as a computer program with a program code for performing audio data assignment in missing audio parts when the computer program is executed on a computer.
Furthermore, it is important to note that the units illustrated in Figure 3 may be realized by software modules.

Claims

1. A process for assigning audio data in missing audio parts of a music piece, comprising the steps of:
- executing (10) musical features extraction on said music piece such that said music piece is represented by a string (u) of musical features, each musical feature being represented by a sequence or a set of one or more symbols of a predetermined alphabet;
- defining (20) a similarity measure for measuring the similarity between substrings of musical features; - defining (30) a substring (νφ) of missing audio parts in said string (u) of musical features;
- locating (40) at least one substring (v) of a predetermined size immediately before, respectively after said substring (νφ) of missing audio parts;
- applying (50) said similarity measure for identifying a reference substring ( ui or U2) including said at least one substring (v) of a predetermined size at its beginning, respectively at its end, said reference substring ( ui or U2) best matching to the substring (νφ) of missing audio parts, said steps (40) and (50) being executed sequentially or in parallel ; - substituting (60) said substring (νφ) of missing audio parts by audio data from the identified reference substring.
2. The process according to claim 1, comprising the steps of: - locating a first substring (vj) of a predetermined size immediately before said substring (νφ) of missing audio parts ;
- locating a second substring (vr) of a predetermined size immediately after said substring (νφ) of missing audio parts;
- applying said similarity measure for identifying said reference substring ( ui or U2) including the first substring (vj) at its beginning and the second substring (vr) at its end, said reference substring ( ui or U2) best matching to the substring (νφ) of missing audio parts, said steps being executed sequentially or in parallel.
3. The process according to anyone of the claims 1 or 2, wherein said reference substring ( ui or U2) is identified in said string (u) of musical features representing said music piece.
4. The process according to anyone of the claims 1 or 2, wherein said reference substring ( ui or U2) is identified in a database of reference music pieces, said reference music pieces being different from said music piece.
5. The process according to anyone of the preceding claims, wherein the size of said identified reference substring is greater or less than the size of said substring (νφ) of missing audio parts.
6. The process according to claim 5, wherein if there is an overlap between the audio data substituting the substring (νφ) of missing audio parts and the rest of the music piece, a further step of audio data reconstruction is performed in order to ensure smooth audio data transition.
7. The process according to anyone of the claims 1-6, wherein said musical features comprise tonal features.
8. The process according to anyone of the claims 1-6, wherein said musical features comprise rhythm features.
9. The process according to anyone of the claims 1-6, wherein said musical features comprise timbre features.
10. The process according to anyone of the claims 1-6, wherein said musical features comprise vocal or lyrics-related features .
11. The process according to anyone of claims 1-6, wherein said music piece is represented by a symbolic music piece and said audio data is represented by symbolic music data.
12. A device for assigning audio data in missing audio parts of a music piece, comprising:
- an execution unit (110) for executing musical features extraction on said music piece such that the music piece is represented by a string (u) of musical features, each musical feature being represented by a sequence or a set of one or more symbols of a predetermined alphabet;
- a similarity measure definer unit (120) for defining a similarity measure for measuring the similarity between substrings of musical features;
- a substring definer unit (130) for defining a substring (νφ) of missing audio parts in the string (u) of musical features;
- a locator unit (140) for locating at least one substring (v) of a predetermined size immediately before, respectively after said substring (νφ) of missing audio parts; - a similarity measure application unit (150) for applying said similarity measure for identifying a reference substring (ui or U2) including the at least one substring (v) of a predetermined size at its beginning, respectively at its end, said reference substring (ui or U2) best matching to the substring (νφ) of missing audio parts;
- a substitution unit (160) for substituting said substring (νφ) of missing audio parts by audio data from the identified reference substring.
13. The device according to claim 10, wherein said locator unit (140) is adapted for locating a first substring (vj) of a predetermined size immediately before the substring (νφ) of missing audio parts and a second substring (vr) of a predetermined size immediately after the substring (νφ) of missing audio parts.
14. A computer program with a program code for performing, when the computer program is executed on a computer, a process for assigning audio data in missing audio parts of a music piece according to claim 1.
EP12818580.8A 2011-10-17 2012-10-17 A process for assigning audio data in missing audio parts of a music piece and device for performing the same Withdrawn EP2769377A2 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201161547854P 2011-10-17 2011-10-17
PCT/IB2012/055673 WO2013057680A2 (en) 2011-10-17 2012-10-17 A process for assigning audio data in missing audio parts of a music piece and device for performing the same

Publications (1)

Publication Number Publication Date
EP2769377A2 true EP2769377A2 (en) 2014-08-27

Family

ID=47603856

Family Applications (1)

Application Number Title Priority Date Filing Date
EP12818580.8A Withdrawn EP2769377A2 (en) 2011-10-17 2012-10-17 A process for assigning audio data in missing audio parts of a music piece and device for performing the same

Country Status (2)

Country Link
EP (1) EP2769377A2 (en)
WO (1) WO2013057680A2 (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10182093B1 (en) * 2017-09-12 2019-01-15 Yousician Oy Computer implemented method for providing real-time interaction between first player and second player to collaborate for musical performance over network

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20040064308A1 (en) * 2002-09-30 2004-04-01 Intel Corporation Method and apparatus for speech packet loss recovery

Also Published As

Publication number Publication date
WO2013057680A2 (en) 2013-04-25
WO2013057680A3 (en) 2013-07-18

Similar Documents

Publication Publication Date Title
US7017113B2 (en) Method and apparatus for removing redundant information from digital documents
CN110769178B (en) Method, device and equipment for automatically generating goal shooting highlights of football match and computer readable storage medium
CN110335622B (en) Audio single-tone color separation method, device, computer equipment and storage medium
CN118116361A (en) Cross-speaker style transfer for speech synthesis
CN108711422A (en) Audio recognition method, device, computer readable storage medium and computer equipment
CN112927667B (en) Chord identification method, device, equipment and storage medium
CN112331170A (en) Method, device and equipment for analyzing similarity of Buddha music melody and storage medium
Muth et al. Improving DNN-based music source separation using phase features
US8080723B2 (en) Rhythm matching parallel processing apparatus in music synchronization system of motion capture data and computer program thereof
US9323889B2 (en) System and method for processing reference sequence for analyzing genome sequence
WO2016181369A1 (en) Method for determining nucleotide sequence
CN114464260A (en) Assembling method and assembling device for genome at chromosome level
Park et al. Dex-tts: Diffusion-based expressive text-to-speech with style modeling on time variability
Yin et al. Atri: Mitigating multilingual audio text retrieval inconsistencies by reducing data distribution errors
WO2013057680A2 (en) A process for assigning audio data in missing audio parts of a music piece and device for performing the same
US20140121986A1 (en) System and method for aligning genome sequence
CN115394397B (en) Medical report automatic generation method based on cross-mode contrast attention
Han et al. Audio imputation using the non-negative hidden markov model
CN116363402A (en) A detection data waveform alignment method and system based on data feature matching
CN112967734B (en) Multi-voice based music data recognition method, device, equipment and storage medium
US20180060667A1 (en) Comparing video sequences using fingerprints
Zeng et al. End-to-end real-world polyphonic piano audio-to-score transcription with hierarchical decoding
US20150066384A1 (en) System and method for aligning genome sequence
CN117012178B (en) Method and device for generating rhythmic annotation data
Spratley et al. A unified neural architecture for instrumental audio tasks

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

17P Request for examination filed

Effective date: 20140328

AK Designated contracting states

Kind code of ref document: A2

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR

DAX Request for extension of the european patent (deleted)
GRAP Despatch of communication of intention to grant a patent

Free format text: ORIGINAL CODE: EPIDOSNIGR1

INTG Intention to grant announced

Effective date: 20160829

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: GRANT OF PATENT IS INTENDED

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN

18D Application deemed to be withdrawn

Effective date: 20170110