WO2006097299A1 - Method for classifying audio data - Google Patents

Method for classifying audio data Download PDF

Info

Publication number
WO2006097299A1
WO2006097299A1 PCT/EP2006/002398 EP2006002398W WO2006097299A1 WO 2006097299 A1 WO2006097299 A1 WO 2006097299A1 EP 2006002398 W EP2006002398 W EP 2006002398W WO 2006097299 A1 WO2006097299 A1 WO 2006097299A1
Authority
WO
WIPO (PCT)
Prior art keywords
audio data
mood space
mood
comparison
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/EP2006/002398
Other languages
French (fr)
Inventor
Thomas Kemp
Yin Hay Lam
Marta Tolós RIGUEIRO
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sony Deutschland GmbH
Original Assignee
Sony Deutschland GmbH
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sony Deutschland GmbH filed Critical Sony Deutschland GmbH
Priority to US11/908,944 priority Critical patent/US8170702B2/en
Priority to CN200680008774.2A priority patent/CN101142622B/en
Publication of WO2006097299A1 publication Critical patent/WO2006097299A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H1/00Details of electrophonic musical instruments
    • G10H1/0008Associated control or indicating means
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2240/00Data organisation or data communication aspects, specifically adapted for electrophonic musical tools or instruments
    • G10H2240/075Musical metadata derived from musical analysis or for use in electrophonic musical instruments
    • G10H2240/085Mood, i.e. generation, detection or selection of a particular emotional content or atmosphere in a musical piece
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2240/00Data organisation or data communication aspects, specifically adapted for electrophonic musical tools or instruments
    • G10H2240/121Musical libraries, i.e. musical databases indexed by musical parameters, wavetables, indexing schemes using musical parameters, musical rule bases or knowledge bases, e.g. for automatic composing methods
    • G10H2240/155Library update, i.e. making or modifying a musical database using musical parameters as indices

Definitions

  • the present invention relates to a method for classifying audio data.
  • the present invention more particularly relates to a fast music similarity computation method based on e.g. N-dimensional music mood space relationships.
  • the object is achieved according to the present invention by a method for classifying audio data with the features of independent claim 1. Preferred embodiments of the invention method for classifying audio data are within the scope of the dependent subclaims.
  • the object underlying the present invention is also achieved by an apparatus for classifying audio data, by a computer program product, as well as by a computer readable storage medium according to independent claims 18, 19 and 20, respectively.
  • the method for classifying audio data comprises a step (Sl ) of providing audio data in particular as input data, a step (S2) of providing mood space data which define and/or which are descrip- tive or representative for a mood space according to which audio data can be classified, a step (S3) of generating a mood space location within said mood space for said given audio data, a step (S4) of providing at least one compari- son mood space location within said mood space, a step (S5) of comparing said mood space location for said given audio data with said at least one comparison mood space location and thereby generating comparison data, and a step (S6) of providing as a classification result said comparison data in particular as output data which can be used in subsequent classification steps, mainly in detailed comparison steps.
  • said mood space may be or may be modelled by at least one of an Euclidean space model, a Gaussian mixture model, a neural network model, and a decision tree model.
  • said mood space may be or may be modelled by an N-dimensional space or manifold and N may be a given and fixed integer.
  • said comparison data may be alternatively or additionally at least one of being descriptive for, being representative for and comprising at least one of a topology, a metric, a norm, a distance defined in or on said mood space according to a another embodiment of the method for classifying audio data according to the present invention.
  • said comparison data and in particular said topology, metric, norm, and said distance may be obtained based on at least one of said Euclidean space model, said Gaussian mixture model, said neural network model, and said decision tree model according to an advantageous embodiment of the method for classifying audio data according to the present invention.
  • Said comparison data may be derived based on said mood space location within said mood space for said given audio data and they may be based on said comparison mood space location within said mood space according to an additional or alternative embodiment of the method for classifying audio data according to the present invention.
  • Said mood space and / or the model thereof may be defined based on Thayer's music mood model according to an additional or alternative embodiment of the method for classifying audio data according to the present invention.
  • said mood space and/ or the model thereof may be at least two-dimensional and may be defined based on the measured or measurable entities stress SQ describing positive, e.g. happy, and negative, e.g. anxious moods and energy EO describing calm and ener- getic moods as emotional or mood parameters or attributes.
  • said mood space and/or the model thereof are at least three- dimensional and are defined based on the measured or measurable entities for happiness, passion, and excitement.
  • Said step (S4) of providing said at least one comparison mood space location may additionally or alternatively comprise a step of providing at least one additional audio data in particular as additional input data and a step of generating a respective additional mood space location for said additional audio data, and wherein said respective additional mood space location for said additional audio data is used for said at least one comparison mood space location according to an additional or alternative embodiment of the method for classifying audio data according to the present invention.
  • At least two samples of audio data may be compared with respect to each other - one of said samples of audio data being assigned to said derived mood space location and the other one of said of audio data being assigned to said additional mood space location or said comparison mood space location - in particular by comparing said derived mood space location and said additional mood space location or said comparison mood space location.
  • said at least two samples of audio data to be compared with respect to each other may be compared with respect to each other based on said comparison data in a pre-selection process or comparing pre-process and then based on additional features, e. g. based on features more complicated to calculate and/ or based on frequency domain related features, in a more detailed comparing process.
  • said at least two samples of audio data to be compared with respect to each other may be compared with respect to each other in said more detailed comparing process based on said additional features, if said comparison data obtained from said pre-selection process or comparing pre-process are indicative for a sufficient neighbourhood of said at least two samples of audio data.
  • a plurality of more than two samples of audio data may be compared with respect to each other.
  • said given audio data may be compared to a plurality of additional samples of audio data.
  • a comparison list and in particular a play list may be generated which is descriptive for additional samples of audio data of said plurality of additional samples of audio data which are similar to said given audio data.
  • an apparatus for classifying audio data which is adapted and which comprises means for carrying out a method for classifying audio data according to the present invention and the steps thereof.
  • a computer program product is provided comprising computer program means which is adapted to realize the method for classifying audio data according to the present invention and the steps thereof, when it is executed on a computer or a digital signal processing means.
  • a computer readable storage medium which comprises a computer program product according to the present invention.
  • the present invention inter alia relates to a fast music similarity computation method which is in particular based on a N-dimensional music mood space.
  • a N-dimensional music mood space can be used to limit the number of candidates and hence reduce the computation in similarity list generation. For each of the music piece in a huge database, its location in a N- dimensional music mood space is first determined and only music pieces which are close to the music in the mood space are selected and the similarity are computed between the given music and the pre-selected music pieces.
  • a music play list is usually displayed and songs in the play list are usually based on the similarity between the query music and the rest of the music in the data- base.
  • typical commercial music database consists of hundreds of thousands of music.
  • state-of-the-art system usually compute its similarity to all the other music pieces in the database to generate a similarity list.
  • a play list is then generated from the similarity list.
  • the computation required in similarity generation involved about N*N/2 similarity measure computation, where N is the number of songs in the database. For example, if the number of songs in the database is 500,000, then the computation will be 500,000*500,000/2, which is not practical for real applications.
  • a fast music similarity list generation method based on mood space are proposed.
  • the emotion expressed in different music are usually different. Some music are perceived as happy by the listeners, but the other songs might be perceived as sad.
  • listeners generally can distinguish the difference in the degree of emotion expression. For example, one music is happier than the other one, etc.
  • music with different mood usually are considered as dissimilar.
  • the music similarity list generation approach described in this invention proposal exploits such emotion perception as described above.
  • the emotion of music can be described by a N-dimensional mood space.
  • Each dimension describes the extent of a particular emotion attribute.
  • the value of each emotion attribute are first generated.
  • music that are located in the proximity of the given music are first selected.
  • the pre-selection stage instead of computing the similarity of the given music to the rest of the database, only the similarity between the given music and the pre-selected music are computed.
  • Any music emotion /mood model proposed in the literature can be used to construct the N-dimensional mood space. For example, the two-dimensional model proposed by Thayer [ I ].
  • any music can be described by a stress value and an energy value and such values give the coordinates of a given music and hence determine the location of the emotion in the mood space.
  • the coordinates of a music in the mood space is proposed to be generated from any machine learning algorithms such as Neural Network, Decision Tree and Gaussian Mixture Models etc.
  • Gaussian Mixture Models i.e., passion model, happiness model and excite- ment model can be used to model each mood dimension.
  • mood models are trained beforehand. For a given music, each model will generate a score and such score can be used as the coordinates value in the mood space.
  • music pieces that are close to a given music in the mood space are identified by using simple distance measure such as Euclidean distance, Mahalanobis distance or Cosine angles etc.
  • a similarity measure is introduced to compute the similarity between music x and the pre-selected music piece.
  • the similarity measure can be any known similarity measure algorithms, e.g. , each music is modelled by Gaussian Mixture Model. Any model distance criterion (see e.g. [3 ]) can then be used to measure the distance between the two Gaussian Models.
  • the main advantage is the significant reduction in computation to generate music similarity lists for a large database without affecting the similarity ranking performance from the perceptual point of view.
  • Fig. IA is a schematical diagram of a mood space model which can be involved in an embodiment of the inventive method for classifying audio data.
  • Fig. IB is a schematical diagram of a mood space model which can be involved in another embodiment of the inventive method for classifying audio data.
  • Fig. 2 elucidates by means of a schematical diagram a proximity concept which can be involved in the embodiment for the inventive method for classifying audio data as illustrated in Fig. IA.
  • Fig. 3 is a schematical diagram which elucidates basic aspects of the inventive method for analyzing audio data according to a preferred embodiment by means of a flow chart.
  • Fig. IA demonstrates by means of a graphical representation in a schematical manner a model for a mood space M which can be involved for carrying out the method for classifying audio data according to a preferred embodiment of the prevent invention.
  • the mood space M shown in Fig. IA is based, defined and constructed accord- ing to so-called mood space data MSD. Locations or positions within said mood space M and in order to navigate within said mood space M are the entities stress S and energy E. Therefore, the model shown in Fig. IA is a two- dimensional mood space model for said mood space M. In the coordinate system defined by the two axes for stress S and energy E, three locations for three different sets of audio data AD, AD' are indicated. The respective sets of audio data AD, AD' are called x, y, and z, respectively. In the embodiment shown in Fig. IA the first set of audio data AD which is called x serves as given audio data x.
  • the respective location LADx for said first set or sample of audio data x is a function of said measured values S(x), E(x).
  • the location LADx for audio data x is simply the pair of values S(x), E(x), i.e.
  • Fig. IA Under the assumption that a distance function is valid in the Euclidean manner, audio data x and y are close together with respect to each other, whereas audio data z are at a distal position with respect to said first and second audio data x and y, respectively. Additionally certain regions of the complete mood space M can be assigned to certain characteristics moods such as contentment, depression, exuberance, and anxiousness.
  • Fig. IB demonstrates by means of a graphic representation in a schematic way that also more than two dimensions in said mood space M are possible.
  • Fig. IB one has three dimensions with the entities happiness, passion and excitement defining the respective three coordinates within said mood space M.
  • Fig. 2 demonstrates in more detail the notion and the concept of neighbourhood and vicinity for the embodiment already demonstrated in Fig. IA.
  • one has the original audio data x with a respective location or position LADx in said mood space M.
  • a threshold value which might be used in order to realize or define neighbourhoods A(x) for said audio data x within said mood space M.
  • the shown neighbourhood A(x) for said audio data x is a circle with the position LADx for said first audio data x in its centre and having a radius with respect to the distance or matrlc underlying the neighbourhood concept discussed here which is equal to the chosen threshold value.
  • any additional audio data AD within said neighbourhood circle A(x) are assumed to be comparable and similar enough when compared to said first and given audio data x.
  • additional audio data z is too far away with respect to the underlying distance or matric so that z can be classified as being not compa- rable to said given and first audio data x.
  • Such a concept of vicinity or neighbourhood can be used in order to compare a given sample of audio data x with a data base of audio samples, for instance in order to reduce computational burden when comparing audio data samples with respect to each other. In the case shown in Fig.
  • a pre-selection process is carried out based on the con- cept of distance and metric in order to select a much more refined subset from the whole data base containing only a very few samples of audio data which have to be compared with respect to each other or with respect to a given piece of audio data x.
  • Fig. 3 is a schematical block diagram containing a flow chart for the most prominent method steps in order to realize an embodiment of the method for classifying audio data AD according to the present invention.
  • step S2 information is provided with respect to a mood space underlying the inventive method. Therefore in step S2 respective mode space data MSD are provided which define and/ or which are descriptive or representative for said mood space M according to which audio data AD 1 AD' can be classified and compared.
  • a step S3 follows wherein a mood space location LAD for said given audio data AD within said mood space M is generated. Contained is a substep S3a for analyzing said audio data AD, e.g. with respect to a given feature set FS which might be obtained from a respective data base. In the following substep S3b the mood space location LAD for said audio data AD is generated as a function of said audio data AD:
  • LAD : LAD(AD).
  • a comparison mood space location CL is received, for instance also from a data base.
  • Said comparison mood space location CL might be dependent on one or a plurality of additional audio data AD' to which the given audio data AD shall be compared to. Additionally in this case the comparison mood space location CL might also be dependent on the feature set FS underlying the present classification scheme.
  • step S5 the locations LAD for the given sample of audio data AD and the comparison location are compared in order to generate respective comparison data CD.
  • Said comparison data CD might also be realized by indicating a distance between said locations LAD and CL.
  • step S6 the comparison data CD are given as an output O.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

A method for classifying audio data (AD) is provided. For a given piece of audio data (AD) a location or position (LAD) for said given audio data (AD) within a mood space (M) is generated and compared to a comparison mood space location (CL). As a result of the comparison comparison data (CD) are generated and provided as a classification result with respect to said given audio data (AD).

Description

Method for Classifying Audio Data
The present invention relates to a method for classifying audio data. The present invention more particularly relates to a fast music similarity computation method based on e.g. N-dimensional music mood space relationships.
Recently, the classification of audio data and in particular of pieces of music becomes more and more important as many electronic devices and in particular customer devices enable a respective user to store and manage a large plurality of music items and titles. In order to enhance the managing mechanism for such music data basis it is necessary to obtain a comparison between different pieces of audio data or different pieces of music in an easy and fast manner.
Therefore, a variety of mechanisms have been developed in order to extract from an analysis of audio data particular properties and features in order to compare pieces of music by comparing the respective sets or n-tuples of properties and features. However, many of the known features to be evaluated within such a comparison mechanism are difficult to calculate and the computational burden is in some cases not reasonable.
It is an object underlying the present invention to provide a method for classifying audio data which enables a reliable and easy and fast to compute comparison and classification of audio data.
The object is achieved according to the present invention by a method for classifying audio data with the features of independent claim 1. Preferred embodiments of the invention method for classifying audio data are within the scope of the dependent subclaims. The object underlying the present invention is also achieved by an apparatus for classifying audio data, by a computer program product, as well as by a computer readable storage medium according to independent claims 18, 19 and 20, respectively.
The method for classifying audio data according to the present invention comprises a step (Sl ) of providing audio data in particular as input data, a step (S2) of providing mood space data which define and/or which are descrip- tive or representative for a mood space according to which audio data can be classified, a step (S3) of generating a mood space location within said mood space for said given audio data, a step (S4) of providing at least one compari- son mood space location within said mood space, a step (S5) of comparing said mood space location for said given audio data with said at least one comparison mood space location and thereby generating comparison data, and a step (S6) of providing as a classification result said comparison data in particular as output data which can be used in subsequent classification steps, mainly in detailed comparison steps.
It is therefore a key idea of the present invention to obtain from an analysis of given audio data a position or location within a mood space wherein said mood space is pre-defined or given by mood space data. Then the given audio data can be classified or compared by comparing the derived mood space location for said given audio data with said at least one comparison mood space location. The thereby generated comparison data or classification data are provided as a classification result or a comparison result. It is therefore essential to have for a given piece of audio data a position or location, e. g. by means of coordinate n-tuple, which can easily compared with other locations or positions in said mood space, e. g. by simply comparing the respective coordinates of the position or location. Therefore audio data can easily be classified and compared with other audio data.
According to a preferred embodiment of the method for classifying audio data according to the present invention said mood space may be or may be modelled by at least one of an Euclidean space model, a Gaussian mixture model, a neural network model, and a decision tree model.
Additionally or alternatively, according to a further preferred embodiment of the method for classifying audio data according to the present invention said mood space may be or may be modelled by an N-dimensional space or manifold and N may be a given and fixed integer.
Further additionally or alternatively, said comparison data may be alternatively or additionally at least one of being descriptive for, being representative for and comprising at least one of a topology, a metric, a norm, a distance defined in or on said mood space according to a another embodiment of the method for classifying audio data according to the present invention.
Additionally or alternatively, said comparison data and in particular said topology, metric, norm, and said distance may be obtained based on at least one of said Euclidean space model, said Gaussian mixture model, said neural network model, and said decision tree model according to an advantageous embodiment of the method for classifying audio data according to the present invention.
Said comparison data may be derived based on said mood space location within said mood space for said given audio data and they may be based on said comparison mood space location within said mood space according to an additional or alternative embodiment of the method for classifying audio data according to the present invention.
Said mood space and / or the model thereof may be defined based on Thayer's music mood model according to an additional or alternative embodiment of the method for classifying audio data according to the present invention.
According to a further preferred embodiment of the method for classifying audio data according to the present invention said mood space and/ or the model thereof may be at least two-dimensional and may be defined based on the measured or measurable entities stress SQ describing positive, e.g. happy, and negative, e.g. anxious moods and energy EO describing calm and ener- getic moods as emotional or mood parameters or attributes.
Further additionally or alternatively, according to a still further preferred embodiment of the method for classifying audio data according to the present invention said mood space and/or the model thereof are at least three- dimensional and are defined based on the measured or measurable entities for happiness, passion, and excitement.
Said step (S4) of providing said at least one comparison mood space location may additionally or alternatively comprise a step of providing at least one additional audio data in particular as additional input data and a step of generating a respective additional mood space location for said additional audio data, and wherein said respective additional mood space location for said additional audio data is used for said at least one comparison mood space location according to an additional or alternative embodiment of the method for classifying audio data according to the present invention.
At least two samples of audio data may be compared with respect to each other - one of said samples of audio data being assigned to said derived mood space location and the other one of said of audio data being assigned to said additional mood space location or said comparison mood space location - in particular by comparing said derived mood space location and said additional mood space location or said comparison mood space location.
Further additionally or alternatively, according to a still further preferred embodiment of the method for classifying audio data according to the present invention said at least two samples of audio data to be compared with respect to each other may be compared with respect to each other based on said comparison data in a pre-selection process or comparing pre-process and then based on additional features, e. g. based on features more complicated to calculate and/ or based on frequency domain related features, in a more detailed comparing process.
In this case said at least two samples of audio data to be compared with respect to each other may be compared with respect to each other in said more detailed comparing process based on said additional features, if said comparison data obtained from said pre-selection process or comparing pre-process are indicative for a sufficient neighbourhood of said at least two samples of audio data.
Alternatively, a plurality of more than two samples of audio data may be compared with respect to each other.
Alternatively or additionally, said given audio data may be compared to a plurality of additional samples of audio data.
In these cases from said comparison a comparison list and in particular a play list may be generated which is descriptive for additional samples of audio data of said plurality of additional samples of audio data which are similar to said given audio data.
According to a further preferred and advantageous embodiment of the method for classifying audio data according to the present invention music pieces are used as samples of audio data
According to a further aspect of the present invention, an apparatus for classifying audio data is provided which is adapted and which comprises means for carrying out a method for classifying audio data according to the present invention and the steps thereof. According to a further aspect of the present invention a computer program product is provided comprising computer program means which is adapted to realize the method for classifying audio data according to the present invention and the steps thereof, when it is executed on a computer or a digital signal processing means.
Additionally a computer readable storage medium is provided which comprises a computer program product according to the present invention.
These and further aspects of the present invention will be further discussed in the following:
Concept
The present invention inter alia relates to a fast music similarity computation method which is in particular based on a N-dimensional music mood space.
It is proposed that a N-dimensional music mood space can be used to limit the number of candidates and hence reduce the computation in similarity list generation. For each of the music piece in a huge database, its location in a N- dimensional music mood space is first determined and only music pieces which are close to the music in the mood space are selected and the similarity are computed between the given music and the pre-selected music pieces.
Background
Music similarity is a relatively new topic, and at this moment, the interest into it is quite academic. Systems have been developed that compare music pieces with one another using statistics over what is called 'timbre' - a mixture of a variety of low-level features. Various distance measures have been proposed including expensive methods like Monte-Carlo-simulation of samples of a distribution and probability estimation of the artificial samples using the statistics from the other music piece. See e.g. [3] for details.
The state of the art in emotion recognition in music is a rather new topic. While a huge amount of papers have been written about music processing in general, few papers have been published regarding emotion in music. State of the art system used for emotion classification in music classifiers include Gaussian mixtures models, support vector machines, neural networks etc. There are also studies abcmt perception of emotion in music, but the results are still very preliminary. Reference [1] and [2] provides information about the state-of-the art mood detection techniques.
Problem
For applications which involved music retrieval or music suggestion, a music play list is usually displayed and songs in the play list are usually based on the similarity between the query music and the rest of the music in the data- base. Nowadays, typical commercial music database consists of hundreds of thousands of music. For each of the music in the database, state-of-the-art system usually compute its similarity to all the other music pieces in the database to generate a similarity list. Based on the applications, a play list is then generated from the similarity list. The computation required in similarity generation involved about N*N/2 similarity measure computation, where N is the number of songs in the database. For example, if the number of songs in the database is 500,000, then the computation will be 500,000*500,000/2, which is not practical for real applications.
In this proposal, a fast music similarity list generation method based on mood space are proposed. The emotion expressed in different music are usually different. Some music are perceived as happy by the listeners, but the other songs might be perceived as sad. On the other hand, among songs with similar mood or emotion, listeners generally can distinguish the difference in the degree of emotion expression. For example, one music is happier than the other one, etc. In additional, music with different mood usually are considered as dissimilar. The music similarity list generation approach described in this invention proposal exploits such emotion perception as described above.
In this proposal, we first proposed that the emotion of music can be described by a N-dimensional mood space. Each dimension describes the extent of a particular emotion attribute. For each of the music in the database, the value of each emotion attribute are first generated.. According to the coordinates of a particular music in this N-dimensional space, music that are located in the proximity of the given music are first selected. After the pre-selection stage, instead of computing the similarity of the given music to the rest of the database, only the similarity between the given music and the pre-selected music are computed. Any music emotion /mood model proposed in the literature can be used to construct the N-dimensional mood space. For example, the two-dimensional model proposed by Thayer [ I ]. The model adopts the theory that the mood is entrailed from two factors : stress (positive /negative) and energy (calm/ energetic). According to Thayer's mood model, any music can be described by a stress value and an energy value and such values give the coordinates of a given music and hence determine the location of the emotion in the mood space. In Figure Ia, the stress value and energy value of music x is S(x) and E(x) respectively and the mood of x is a function of the emotion attribute, i.e. mood(x) = f(E(x), S(x)), where f can be any function. As mentioned above, two music that are close to each other in the mood space, such as music x and music y, are considered to be similar as they are both considered as "contentment". On the other hand, an "Anxious" music such as z is far away from x in the mood space and anxious music such as z are generally not perceived as similar to a "contentment" music such as x. The similar concept is not limited to Thayer model, it can be extended to any N-dimensional model. For example, in Figure Ib, a three dimensional mood space is depicted. Its coordinates describes the degree of happiness, passion and excitement respectively.
The coordinates of a music in the mood space is proposed to be generated from any machine learning algorithms such as Neural Network, Decision Tree and Gaussian Mixture Models etc. For example, taking Fig. Ib as an example, Gaussian Mixture Models, i.e., passion model, happiness model and excite- ment model can be used to model each mood dimension. Such mood models are trained beforehand. For a given music, each model will generate a score and such score can be used as the coordinates value in the mood space.
After the location of the music in the mood space are determined, music pieces that are close to a given music in the mood space are identified by using simple distance measure such as Euclidean distance, Mahalanobis distance or Cosine angles etc.
For example, in Fig. 2, only music pieces that fall within the proximity area, e.g. circle A, are considered as close to music x in the mood space and music z is considered as too far away and hence dissimilar to music x. According to the distance, the system can either select N music pieces that are close to the given music or a distance threshold can be set and only music distance smaller than the threshold will be selected. To generate a similarity list for music x, a similarity measure is introduced to compute the similarity between music x and the pre-selected music piece. The similarity measure can be any known similarity measure algorithms, e.g. , each music is modelled by Gaussian Mixture Model. Any model distance criterion (see e.g. [3 ]) can then be used to measure the distance between the two Gaussian Models.
Advantages
The main advantage is the significant reduction in computation to generate music similarity lists for a large database without affecting the similarity ranking performance from the perceptual point of view.
The invention will now be explained based on preferred embodiments thereof and by taking reference to the accompanying and schematical figures.
Fig. IA is a schematical diagram of a mood space model which can be involved in an embodiment of the inventive method for classifying audio data.
Fig. IB is a schematical diagram of a mood space model which can be involved in another embodiment of the inventive method for classifying audio data.
Fig. 2 elucidates by means of a schematical diagram a proximity concept which can be involved in the embodiment for the inventive method for classifying audio data as illustrated in Fig. IA.
Fig. 3 is a schematical diagram which elucidates basic aspects of the inventive method for analyzing audio data according to a preferred embodiment by means of a flow chart.
In the following functional and structural similar or equivalent element struc- tures will be denoted with the same reference symbols. Not in each case of their occurrence a detailed description will be repeated.
Fig. IA demonstrates by means of a graphical representation in a schematical manner a model for a mood space M which can be involved for carrying out the method for classifying audio data according to a preferred embodiment of the prevent invention.
The mood space M shown in Fig. IA is based, defined and constructed accord- ing to so-called mood space data MSD. Locations or positions within said mood space M and in order to navigate within said mood space M are the entities stress S and energy E. Therefore, the model shown in Fig. IA is a two- dimensional mood space model for said mood space M. In the coordinate system defined by the two axes for stress S and energy E, three locations for three different sets of audio data AD, AD' are indicated. The respective sets of audio data AD, AD' are called x, y, and z, respectively. In the embodiment shown in Fig. IA the first set of audio data AD which is called x serves as given audio data x. Based on the evaluation of the entities stress S and energy E for said first set of audio data x respective parameter values S(x) and E(x) are generated. Therefore, the respective location LADx for said first set or sample of audio data x is a function of said measured values S(x), E(x). In the simplest case of a representation the location LADx for audio data x is simply the pair of values S(x), E(x), i.e.
LADx := LAD(S(Jc)5E(Jt)) = (S(x),E(x)) .
The same may hold for second and third audio data y and z with measurement values S(y), E(y) and S(z), E(z), respectively. According to the general properties for the locations or positions LADy and LADz in said mood space M the following expressions are given:
LADy := LAD(S (y), E (y)) = (S(y),E(y))
and
LADz := LAD(S(Z), E(z)) = {S(z),E(z)).
As can be seen from the representation of Fig. IA, under the assumption that a distance function is valid in the Euclidean manner, audio data x and y are close together with respect to each other, whereas audio data z are at a distal position with respect to said first and second audio data x and y, respectively. Additionally certain regions of the complete mood space M can be assigned to certain characteristics moods such as contentment, depression, exuberance, and anxiousness.
Fig. IB demonstrates by means of a graphic representation in a schematic way that also more than two dimensions in said mood space M are possible. In the case of Fig. IB one has three dimensions with the entities happiness, passion and excitement defining the respective three coordinates within said mood space M.
Fig. 2 demonstrates in more detail the notion and the concept of neighbourhood and vicinity for the embodiment already demonstrated in Fig. IA. Here one has the original audio data x with a respective location or position LADx in said mood space M. With respect to a given concept of distance or metric one can generate or receive a threshold value which might be used in order to realize or define neighbourhoods A(x) for said audio data x within said mood space M. The shown neighbourhood A(x) for said audio data x is a circle with the position LADx for said first audio data x in its centre and having a radius with respect to the distance or matrlc underlying the neighbourhood concept discussed here which is equal to the chosen threshold value. Any additional audio data AD within said neighbourhood circle A(x) are assumed to be comparable and similar enough when compared to said first and given audio data x. In contrast, additional audio data z is too far away with respect to the underlying distance or matric so that z can be classified as being not compa- rable to said given and first audio data x. Such a concept of vicinity or neighbourhood can be used in order to compare a given sample of audio data x with a data base of audio samples, for instance in order to reduce computational burden when comparing audio data samples with respect to each other. In the case shown in Fig. 2 a pre-selection process is carried out based on the con- cept of distance and metric in order to select a much more refined subset from the whole data base containing only a very few samples of audio data which have to be compared with respect to each other or with respect to a given piece of audio data x.
Fig. 3 is a schematical block diagram containing a flow chart for the most prominent method steps in order to realize an embodiment of the method for classifying audio data AD according to the present invention. After initialization step START a sample of audio data AD is received as an input I in a first method step S l .
Then, in a following step S2 information is provided with respect to a mood space underlying the inventive method. Therefore in step S2 respective mode space data MSD are provided which define and/ or which are descriptive or representative for said mood space M according to which audio data AD1 AD' can be classified and compared.
A step S3 follows wherein a mood space location LAD for said given audio data AD within said mood space M is generated. Contained is a substep S3a for analyzing said audio data AD, e.g. with respect to a given feature set FS which might be obtained from a respective data base. In the following substep S3b the mood space location LAD for said audio data AD is generated as a function of said audio data AD:
LAD := LAD(AD).
In the following step S4 a comparison mood space location CL is received, for instance also from a data base. Said comparison mood space location CL might be dependent on one or a plurality of additional audio data AD' to which the given audio data AD shall be compared to. Additionally in this case the comparison mood space location CL might also be dependent on the feature set FS underlying the present classification scheme.
In the following step S5 the locations LAD for the given sample of audio data AD and the comparison location are compared in order to generate respective comparison data CD. Said comparison data CD might also be realized by indicating a distance between said locations LAD and CL.
In the following step S6 the comparison data CD are given as an output O.
Finally, the process demonstrated in Fig. 3 is terminated either with a process step END- I if a quick and sub-optimal classification is sufficient or with - after a detailed and expensive classification S7 is needed - with an alternative process step END- 2. Cited Literature
[1 ] Dan Liu, Li Lu & Hong- Jiang Zhang, "Automatic mood detection from acoustic music data", Proceedings of the Fourth International Conference on Music Information Retrieval (ISMIR) 2003.
[2] Tao Li & Mitsunori Ogihara, "Detecting emotion in music", Proceedings of the Fourth International Conference on Music Information Retrieval (ISMIR) 2003.
[3] J.J. Aucouturier & F. Pachet, "Finding songs that sound the same", in Proc. Of the IEEE Benelux Workshop on model based processing and coding of audio, Nov 2002.
Reference Symbols
A, A(X) neighbourhood, vicinity, neighbourhood or vicinity w.r.t. mood space location for audio data x
AD audio data, audio data sample
AD1 audio data, audio data sample, additional audio data
CD comparison data
CL comparison mood space location
E, E() energy
FS feature set
I input, input data
LAD, LADx, LADy, mood space location for received audio data AD, x, y,
LADz z, respectively
LAD' additional mood space location for received additional audio data AD'
M mood space
MSD mood space data
O output, output data
S, SQ stress x audio data, audio data sample y audio data, audio data sample
Z audio data, audio data sample

Claims

Claims
1. Method for classifying audio data (AD), comprising: a step (S l ) of providing audio data (AD) in particular as input data (I), a step (S2) of providing mood space data (MSD) which define and/ or which are descriptive or representative for a mood space (M) according to which audio data (AD, AD') can be classified, a step (S3) of generating a mood space location (LAD) within said mood space (M) for said given audio data (AD), a step (S4) of providing at least one comparison mood space location (CL) within said mood space (M) , a step (S5) of comparing said mood space location (LAD) for said given audio data (AD) with said at least one comparison mood space location (CL) and thereby generating comparison data (CD), and a step (S 6) of providing as a classification result said comparison data (CD) in particular as output data (O) .
2. Method according to claim 1 , wherein said mood space (M) is or is modelled by at least one of a Gaussian mixture model, a neural network model, and a decision tree model.
3. Method according to any one of the preceding claims, wherein said mood space (M) is or is modelled by an N-dimensional space or manifold and wherein N is a given and fixed integer.
4. Method according to any one of the preceding claims, wherein said comparison data (CD) are at least one of being descriptive for, being representative for and comprising at least one of a topology, a metric, a norm, a distance defined in or on said mood space (M).
5. Method according to any one of the preceding claims, wherein said comparison data (CD) and in particular said topology, metric, norm, and said distance are obtained based on at least one of said Euclidean space model, said Gaussian mixture model, said neural network model, and said decision tree model.
6. Method according to any one of the preceding claims, wherein said comparison data (CD) are derived based on said mood space location (LAD) within said mood space (M) for said given audio data (AD) and on said comparison mood space location (CL) within said mood space (M).
7. Method according to any one of the preceding claims, wherein said mood space (M) and/ or the model thereof are defined based on Thayer's mood model.
8. Method according to any one of the preceding claims, wherein said mood space (M) and/ or the model thereof are two-dimensional and are defined based on the measured or measurable entities stress (SQ) describing happy and anxious moods and energy (E()) describing calm and energetic moods as emotional or mood parameters or attributes.
9. Method according to any one of the preceding claims, wherein said mood space (M) and/ or the model thereof are three-dimensional and are defined based on the measured or measurable entities for happiness, passion, and excitement.
10. Method according to any one of the preceding claims, wherein said step (S4) of providing said at least one comparison mood space location (CL) comprises: a step of providing at least one additional audio data (AD, AD') in particular as additional input data (I) and - a step of generating a respective additional mood space location (LAD') for said additional audio data (AD'), and wherein said respective additional mood space location (LAD') for said additional audio data (AD') is used for said at least one comparison mood space location (CL).
11. Method according to claim 10, wherein at least two samples of audio data (AD, AD') are compared with respect to each other - one (AD) of said samples of audio data (AD, AD') being assigned to said derived mood space location (LAD) and the other one (AD') of said of audio data (AD, AD') being assigned to said additional mood space location (LAD') or said comparison mood space location (CL) - in particular by comparing said derived mood space location (LAD) and said additional mood space location (LAD') or said comparison mood space location (CL).
12. Method according to claim 11 , wherein said at least two samples of audio data (AD, AD') to be compared with respect to each other are compared with respect to each other based on said comparison data (CD) in a pre-selection process or comparing pre-process and then based on additional features, e. g. based on features more complicated to calculate and/ or based on frequency domain related features, in a more detailed comparing process.
13. Method according to claim 12, wherein said at least two samples of audio data (AD, AD1) to be compared with respect to each other are compared with respect to each other in said more detailed comparing process based on said additional features, if said comparison data (CD) obtained from said pre-selection process or comparing pre- process are indicative for a sufficient neighbourhood of said at least two samples of audio data (AD, AD').
14. Method according to any one of the preceding claims, wherein a plurality of more than two samples of audio data (AD, AD') are compared with respect to each other.
15. Method according to any one of the preceding claims, wherein said given audio data (AD) are compared to a plurality of additional samples of audio data (AD1).
16. Method according to any one of the preceding claims 14 or 15, wherein from said comparison a comparison list and in particular a play list is generated which is descriptive for additional samples of audio data (AD') of said plurality of additional samples of audio data (AD') which are similar to said given audio data (AD).
17. Method according to any one of the preceding claims, wherein music pieces are used as samples of audio data (AD, AD').
18. Apparatus for classifying audio data, which is adapted and which comprises means for carrying out a method for classifying audio data according to any one of the claims 1 to 17 and the steps thereof.
19. Computer program product, comprising computer program means which is adapted to realize a method for classifying audio data according to any one of the claims 1 to 17 and the steps thereof, when it is executed on a computer or a digital signal processing means.
20. Computer readable storage medium, comprising a computer program product according to claim 19
PCT/EP2006/002398 2005-03-18 2006-03-15 Method for classifying audio data Ceased WO2006097299A1 (en)

Priority Applications (2)

Application Number Priority Date Filing Date Title
US11/908,944 US8170702B2 (en) 2005-03-18 2006-03-15 Method for classifying audio data
CN200680008774.2A CN101142622B (en) 2005-03-18 2006-03-15 Method for classifying audio data

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP05005994A EP1703491B1 (en) 2005-03-18 2005-03-18 Method for classifying audio data
EP05005994.8 2005-03-18

Publications (1)

Publication Number Publication Date
WO2006097299A1 true WO2006097299A1 (en) 2006-09-21

Family

ID=34934366

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2006/002398 Ceased WO2006097299A1 (en) 2005-03-18 2006-03-15 Method for classifying audio data

Country Status (5)

Country Link
US (1) US8170702B2 (en)
EP (1) EP1703491B1 (en)
JP (1) JP2006276854A (en)
CN (1) CN101142622B (en)
WO (1) WO2006097299A1 (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7601315B2 (en) 2006-12-28 2009-10-13 Cansolv Technologies Inc. Process for the recovery of carbon dioxide from a gas stream
US20110029112A1 (en) * 2008-01-23 2011-02-03 Sony Corporation Method for deriving animation parameters and animation display device
US10888816B2 (en) 2016-11-01 2021-01-12 Shell Oil Company Process for producing a purified gas stream

Families Citing this family (46)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1531456B1 (en) 2003-11-12 2008-03-12 Sony Deutschland GmbH Apparatus and method for automatic dissection of segmented audio signals
US7842876B2 (en) * 2007-01-05 2010-11-30 Harman International Industries, Incorporated Multimedia object grouping, selection, and playback system
EP1975866A1 (en) 2007-03-31 2008-10-01 Sony Deutschland Gmbh Method and system for recommending content items
US20080300702A1 (en) * 2007-05-29 2008-12-04 Universitat Pompeu Fabra Music similarity systems and methods using descriptors
US8583615B2 (en) * 2007-08-31 2013-11-12 Yahoo! Inc. System and method for generating a playlist from a mood gradient
EP2101501A1 (en) * 2008-03-10 2009-09-16 Sony Corporation Method for recommendation of audio
US8996538B1 (en) 2009-05-06 2015-03-31 Gracenote, Inc. Systems, methods, and apparatus for generating an audio-visual presentation using characteristics of audio, visual and symbolic media objects
US8071869B2 (en) * 2009-05-06 2011-12-06 Gracenote, Inc. Apparatus and method for determining a prominent tempo of an audio work
US8805854B2 (en) * 2009-06-23 2014-08-12 Gracenote, Inc. Methods and apparatus for determining a mood profile associated with media data
WO2011145249A1 (en) * 2010-05-17 2011-11-24 パナソニック株式会社 Audio classification device, method, program and integrated circuit
US20120023403A1 (en) * 2010-07-21 2012-01-26 Tilman Herberger System and method for dynamic generation of individualized playlists according to user selection of musical features
KR101069090B1 (en) * 2011-03-03 2011-09-30 송석명 Prefabricated Rice Garland
CN102693724A (en) * 2011-03-22 2012-09-26 张燕 Noise classification method of Gaussian Mixture Model based on neural network
GB201109731D0 (en) 2011-06-10 2011-07-27 System Ltd X Method and system for analysing audio tracks
US9263060B2 (en) 2012-08-21 2016-02-16 Marian Mason Publishing Company, Llc Artificial neural network based system for classification of the emotional content of digital music
CN103258532B (en) * 2012-11-28 2015-10-28 河海大学常州校区 A kind of Chinese speech sensibility recognition methods based on fuzzy support vector machine
EP2759949A1 (en) * 2013-01-28 2014-07-30 Tata Consultancy Services Limited Media system for generating playlist of multimedia files
US9639871B2 (en) 2013-03-14 2017-05-02 Apperture Investments, Llc Methods and apparatuses for assigning moods to content and searching for moods to select content
US10061476B2 (en) 2013-03-14 2018-08-28 Aperture Investments, Llc Systems and methods for identifying, searching, organizing, selecting and distributing content based on mood
US10623480B2 (en) 2013-03-14 2020-04-14 Aperture Investments, Llc Music categorization using rhythm, texture and pitch
US10225328B2 (en) 2013-03-14 2019-03-05 Aperture Investments, Llc Music selection and organization using audio fingerprints
US9875304B2 (en) 2013-03-14 2018-01-23 Aperture Investments, Llc Music selection and organization using audio fingerprints
US10242097B2 (en) 2013-03-14 2019-03-26 Aperture Investments, Llc Music selection and organization using rhythm, texture and pitch
US11271993B2 (en) 2013-03-14 2022-03-08 Aperture Investments, Llc Streaming music categorization using rhythm, texture and pitch
CN103440863B (en) * 2013-08-28 2016-01-06 华南理工大学 A kind of speech-emotion recognition method based on stream shape
TWI603213B (en) * 2014-01-23 2017-10-21 國立交通大學 Method for selecting music based on face recognition, music selecting system and electronic apparatus
US20220147562A1 (en) 2014-03-27 2022-05-12 Aperture Investments, Llc Music streaming, playlist creation and streaming architecture
CN104700829B (en) * 2015-03-30 2018-05-01 中南民族大学 Animal sounds Emotion identification system and method
US10854180B2 (en) 2015-09-29 2020-12-01 Amper Music, Inc. Method of and system for controlling the qualities of musical energy embodied in and expressed by digital music to be automatically composed and generated by an automated music composition and generation engine
US9721551B2 (en) 2015-09-29 2017-08-01 Amper Music, Inc. Machines, systems, processes for automated music composition and generation employing linguistic and/or graphical icon based musical experience descriptions
US10261964B2 (en) * 2016-01-04 2019-04-16 Gracenote, Inc. Generating and distributing playlists with music and stories having related moods
CN107293308B (en) * 2016-04-01 2019-06-07 腾讯科技(深圳)有限公司 A kind of audio-frequency processing method and device
CN106331741B (en) * 2016-08-31 2019-03-08 徐州视达坦诚文化发展有限公司 A kind of compression method of television broadcast media audio, video data
CN106231357B (en) * 2016-08-31 2017-05-10 浙江华治数聚科技股份有限公司 Method for predicting fragment time of television broadcast media audio/video data
US10750229B2 (en) 2017-10-20 2020-08-18 International Business Machines Corporation Synchronized multi-media streams including mood data
GB201718894D0 (en) 2017-11-15 2017-12-27 X-System Ltd Russel space
US11020560B2 (en) 2017-11-28 2021-06-01 International Business Machines Corporation System and method to alleviate pain
US10426410B2 (en) 2017-11-28 2019-10-01 International Business Machines Corporation System and method to train system to alleviate pain
JP7223848B2 (en) * 2018-11-15 2023-02-16 ソニー・インタラクティブエンタテインメント エルエルシー Dynamic music generation in gaming
US11341945B2 (en) * 2019-08-15 2022-05-24 Samsung Electronics Co., Ltd. Techniques for learning effective musical features for generative and retrieval-based applications
US11037538B2 (en) 2019-10-15 2021-06-15 Shutterstock, Inc. Method of and system for automated musical arrangement and musical instrument performance style transformation supported within an automated music performance system
US11024275B2 (en) 2019-10-15 2021-06-01 Shutterstock, Inc. Method of digitally performing a music composition using virtual musical instruments having performance logic executing within a virtual musical instrument (VMI) library management system
US10964299B1 (en) 2019-10-15 2021-03-30 Shutterstock, Inc. Method of and system for automatically generating digital performances of music compositions using notes selected from virtual musical instruments based on the music-theoretic states of the music compositions
US11615772B2 (en) * 2020-01-31 2023-03-28 Obeebo Labs Ltd. Systems, devices, and methods for musical catalog amplification services
US12536227B2 (en) * 2021-10-20 2026-01-27 Sony Group Corporation Information processing apparatus, information processing method, and program
US12272341B2 (en) * 2021-11-08 2025-04-08 Lemon Inc. Controllable music generation

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20030045953A1 (en) * 2001-08-21 2003-03-06 Microsoft Corporation System and methods for providing automatic classification of media entities according to sonic properties

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1300831B1 (en) 2001-10-05 2005-12-07 Sony Deutschland GmbH Method for detecting emotions involving subspace specialists
US7158931B2 (en) * 2002-01-28 2007-01-02 Phonak Ag Method for identifying a momentary acoustic scene, use of the method and hearing device
EP1531456B1 (en) 2003-11-12 2008-03-12 Sony Deutschland GmbH Apparatus and method for automatic dissection of segmented audio signals

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20030045953A1 (en) * 2001-08-21 2003-03-06 Microsoft Corporation System and methods for providing automatic classification of media entities according to sonic properties

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
KODAMA Y ET AL: "A music recommendation system", CONSUMER ELECTRONICS, 2005. ICCE. 2005 DIGEST OF TECHNICAL PAPERS. INTERNATIONAL CONFERENCE ON LAS VEGAS, NV, USA JAN. 8-12, 2005, PISCATAWAY, NJ, USA,IEEE, 8 January 2005 (2005-01-08), pages 219 - 220, XP010796610, ISBN: 0-7803-8838-0 *
TOLOS M ET AL: "Mood-based navigation through large collections of musical data", CONSUMER COMMUNICATIONS AND NETWORKING CONFERENCE, 2005. CCNC. 2005 SECOND IEEE LAS VEGAS, NV, USA 3-6 JAN. 2005, PISCATAWAY, NJ, USA,IEEE, 3 January 2005 (2005-01-03), pages 71 - 75, XP010787613, ISBN: 0-7803-8784-8 *
YAZHONG FENG ET AL: "Music information retrieval by detecting mood via computational media aesthetics", WEB INTELLIGENCE, 2003. WI 2003. PROCEEDINGS. IEEE/WIC INTERNATIONAL CONFERENCE ON OCT. 13-17, 2003, PISCATAWAY, NJ, USA,IEEE, 13 October 2003 (2003-10-13), pages 235 - 241, XP010662937, ISBN: 0-7695-1932-6 *

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7601315B2 (en) 2006-12-28 2009-10-13 Cansolv Technologies Inc. Process for the recovery of carbon dioxide from a gas stream
US20110029112A1 (en) * 2008-01-23 2011-02-03 Sony Corporation Method for deriving animation parameters and animation display device
US8606384B2 (en) * 2008-01-23 2013-12-10 Sony Corporation Method for deriving animation parameters and animation display device
US10888816B2 (en) 2016-11-01 2021-01-12 Shell Oil Company Process for producing a purified gas stream

Also Published As

Publication number Publication date
EP1703491B1 (en) 2012-02-22
EP1703491A1 (en) 2006-09-20
US8170702B2 (en) 2012-05-01
JP2006276854A (en) 2006-10-12
CN101142622A (en) 2008-03-12
US20090069914A1 (en) 2009-03-12
CN101142622B (en) 2011-10-26

Similar Documents

Publication Publication Date Title
EP1703491B1 (en) Method for classifying audio data
Casey et al. Content-based music information retrieval: Current directions and future challenges
EP1615204B1 (en) Method for classifying music
Yang et al. Ranking-based emotion recognition for music organization and retrieval
JP5115966B2 (en) Music retrieval system and method and program thereof
US7805389B2 (en) Information processing apparatus and method, program and recording medium
EP1895505A1 (en) Method and device for musical mood detection
CN114783456B (en) Song main melody extraction method, song processing method, computer equipment and product
TWI396105B (en) Digital data processing method for personalized information retrieval and computer readable storage medium and information retrieval system thereof
US9576050B1 (en) Generating a playlist based on input acoustic information
Purnama Music genre recommendations based on spectrogram analysis using convolutional neural network algorithm with RESNET-50 and VGG-16 architecture
Dwivedi et al. Generative adversarial networks based framework for music genre classification
US20090106176A1 (en) Information processing apparatus, information processing method, and program
Zeng et al. Content filtering methods for music recommendation: A review
CN116417012A (en) Audio recognition method, computer device and computer program product
Kim et al. A music recommendation system based on personal preference analysis
KR101520572B1 (en) Method and apparatus for multiple meaning classification related music
Pavitha et al. Analysis of clustering algorithms for music recommendation
JPWO2006137271A1 (en) Music search device, music search method, and music search program
JP3934556B2 (en) Method and apparatus for extracting signal identifier, method and apparatus for creating database from signal identifier, and method and apparatus for referring to search time domain signal
Fan et al. Music similarity model based on CRP fusion and Multi-Kernel Integration
Rajadnya et al. Raga classification based on pitch co-occurrence based features
Singh et al. Computational approaches for Indian classical music: A comprehensive review
Asha et al. Assessment and Evaluation of Music Genre Classification by employing various AI Techniques
Shen et al. QUC-tree: Integrating query context information for efficient music retrieval

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application
DPE1 Request for preliminary examination filed after expiration of 19th month from priority date (pct application filed from 20040101)
WWE Wipo information: entry into national phase

Ref document number: 200680008774.2

Country of ref document: CN

NENP Non-entry into the national phase

Ref country code: DE

NENP Non-entry into the national phase

Ref country code: RU

WWW Wipo information: withdrawn in national office

Country of ref document: RU

122 Ep: pct application non-entry in european phase

Ref document number: 06707576

Country of ref document: EP

Kind code of ref document: A1

WWW Wipo information: withdrawn in national office

Ref document number: 6707576

Country of ref document: EP

WWE Wipo information: entry into national phase

Ref document number: 11908944

Country of ref document: US