WO2020098086A1 - 一种音乐自动生成方法、装置及计算机可读存储介质 - Google Patents

一种音乐自动生成方法、装置及计算机可读存储介质 Download PDF

Info

Publication number
WO2020098086A1
WO2020098086A1 PCT/CN2018/123593 CN2018123593W WO2020098086A1 WO 2020098086 A1 WO2020098086 A1 WO 2020098086A1 CN 2018123593 W CN2018123593 W CN 2018123593W WO 2020098086 A1 WO2020098086 A1 WO 2020098086A1
Authority
WO
WIPO (PCT)
Prior art keywords
playing time
audio
music
digital audio
time threshold
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2018/123593
Other languages
English (en)
French (fr)
Inventor
刘奡智
王义文
王健宗
肖京
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Publication of WO2020098086A1 publication Critical patent/WO2020098086A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H1/00Details of electrophonic musical instruments
    • G10H1/0008Associated control or indicating means
    • G10H1/0025Automatic or semi-automatic music composition, e.g. producing random music, applying rules from music theory or modifying a musical piece
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H7/00Instruments in which the tones are synthesised from a data store, e.g. computer organs
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2210/00Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
    • G10H2210/005Musical accompaniment, i.e. complete instrumental rhythm synthesis added to a performed melody, e.g. as output by drum machines
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2210/00Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
    • G10H2210/101Music Composition or musical creation; Tools or processes therefor
    • G10H2210/111Automatic composing, i.e. using predefined musical rules
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2250/00Aspects of algorithms or signal processing methods without intrinsic musical character, yet specifically adapted for or used in electrophonic musical processing
    • G10H2250/131Mathematical functions for musical analysis, processing, synthesis or composition
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2250/00Aspects of algorithms or signal processing methods without intrinsic musical character, yet specifically adapted for or used in electrophonic musical processing
    • G10H2250/311Neural networks for electrophonic musical instruments or musical processing, e.g. for musical recognition or control, automatic composition or improvisation

Definitions

  • the present application relates to the field of intelligent decision-making technology, and in particular to a method and device for automatically generating music and a computer-readable storage medium.
  • Sound is a wave of sound generated by the vibration of an object, propagating through a medium (air or solid, liquid) and can be perceived by human or animal hearing organs.
  • Music belongs to a special sound mode. When playing an instrument, the vibration of the instrument will cause rhythmic vibration of the medium (air molecules), causing the surrounding air to change densely and densely, forming longitudinal waves between densely and densely. This produces what is called Music (this phenomenon will continue until the vibration disappears).
  • the present application provides a method and device for automatically generating music and a computer-readable storage medium, and its main purpose is to improve the accuracy of automatically generated music.
  • the present application also provides a method for automatically generating music.
  • the method includes:
  • the digital audio is stored as training data of a non-time series prediction model.
  • the present application also provides an apparatus for automatically generating music, which includes a memory and a processor, and the memory stores a program that can run on the processor, and the program is processed by the program The following steps are implemented when the device is executed:
  • the digital audio is stored as training data of a non-time series prediction model.
  • the present application also provides a computer-readable storage medium, the computer-readable storage medium stores an automatic music generation program, the program can be executed by one or more processors to achieve the above The steps of the method.
  • the method, device and computer-readable storage medium for automatically generating music proposed by the present application predict the music melody playing time and predict the music melody by different prediction models, thereby improving the robustness and adaptability of the entire model.
  • FIG. 1 is a schematic flowchart of a method for automatically generating music according to an embodiment of the present application
  • FIG. 2 is a schematic structural diagram of an echo state network model provided by an embodiment of this application.
  • FIG. 3 is a schematic diagram of a training process of a DCGAN network model provided by an embodiment of this application;
  • FIG. 4 is a schematic diagram of an internal structure of an automatic music generation device provided by an embodiment of the present application.
  • FIG. 5 is a schematic diagram of modules of a program in an apparatus for automatically generating music provided by an embodiment of the present application.
  • This application provides a method for automatically generating music.
  • FIG. 1 it is a schematic flowchart of a method for automatically generating music according to an embodiment of the present application.
  • This method can use various interactive equipments with a sound card (Digital to Analog Converter, DAC) and a digital-to-analog converter in Chinese, such as mobile phones, tablets, computers, etc., as performance devices to implement the method of this embodiment.
  • DAC Digital to Analog Converter
  • the various types of interactive devices mentioned above can be implemented by software and / or hardware.
  • a method for automatically generating music includes:
  • Step S10 Collect audio signals of music melody and convert the audio signals into digital audio storage.
  • step S10 further includes:
  • S101 Use an audio amplifier to collect the sampling frequency and sampling digits of the audio signal
  • the task of collecting audio signals is to discretize continuous sound waveforms, that is, collecting music analog signals.
  • a continuous signal with limited bandwidth can be replaced by a sequence of discrete sampling points, and this substitution will not lose any information.
  • Fourier theory also pointed out: All complex periodic waveforms are composed of a series of sine waves arranged in harmonics. Complex waveforms can be synthesized by the summation and summation of multiple sine waves. Therefore, the audio signal is discretely sampled according to the system, the audio signal is defined at each exact time point, and the audio signal to be collected can be collected.
  • sampling frequency Sample Rate
  • sampling digit the sampling digit is the amplitude dynamic response data range of each sampling point
  • storage (sampling frequency * sampling digits) / 8 (number of bytes)
  • the audio amplifier uses a sampling frequency of 22.05kHz and a sampling number of 8 bits.
  • the sampling frequency must be at least twice the highest frequency of the signal. The higher the sampling frequency, the smaller the sound distortion and the greater the amount of audio data. Therefore, in general, the upper frequency limit of human ear hearing is about 2OkHz.
  • the sampling frequency In order to ensure that the sound is not distorted, the sampling frequency should be about 4OkHz, but there is no concert to reach the frequency of 20kHz, because high frequency will affect the listener ’s hearing experience Due to the resonance effect of music, the sampling frequency used in the audio amplifier is 22.05kHz.
  • 8-bit, 12-bit and 16-bit are often used for sampling digits.
  • 8-bit quantization level means that each sampling point can represent 256 (2 8 ) different quantization values
  • 16-bit quantization level can represent 65536 different quantization values
  • the higher the sampling and quantization digits the better the sound quality and the larger the data volume. But combined with the performance of the interactive device processor, the audio amplifier uses 8-bit sampling bits for the processing.
  • step S102 further includes:
  • the audio signal is passed through a low-pass filter, and the audio signal higher than the half sampling frequency is band-limited to improve aliasing interference.
  • the aliasing interference phenomenon that is, an input signal higher than the half sampling frequency will produce a lower frequency aliasing signal, in which the half sampling evaluation rate is half of the sampling frequency.
  • the sampling frequency of the audio amplifier is 22.05kHz.
  • an interfering alias signal will be generated.
  • the following data cleaning method is adopted for the aliasing interference signal: after the audio amplifier collects the audio signal, a low-pass filter is added. The collected audio signal is band-limited by a low-pass filter (anti-aliasing filter), which provides sufficient attenuation at the half-sampling frequency to ensure that the sampled signal does not contain spectrum exceeding the half-sampling frequency content.
  • step S102 further includes:
  • the noise emitted by the jitter generator is collected and added to the audio signal to improve quantization error interference.
  • the amplitude value is rounded to the nearest quantization scale value. This operation will cause a quantization error.
  • an error will occur between the real analog value and the selected quantization scale value. , Which is the quantization error.
  • This quantization error results in the inability to perfectly encode a continuous analog function when digitally storing audio signals.
  • Data cleaning method based on quantization error interference when the audio amplifier collects audio signals, a small amount of noise generated by the jitter generator is collected at the same time. Because jitter itself is a small amplitude noise that is uncorrelated with the audio signal, it is added to the audio signal of the interactive device before the audio signal is sampled.
  • the audio signal After adding the dithering signal, the audio signal will be shifted for each quantization level. For the previous waveforms that are adjacent in time, because each cycle is different now, there will be no periodic quantization mode, because the quantization error is closely related to the signal cycle, so the various effects of the final quantization error , Also randomized enough to be removed.
  • the digital signal After adding the low-pass filter and jitter generator to solve the data cleaning problem, the digital signal is finally converted into digital audio and stored in the interactive device by the digital converter, and the audio data collection process ends.
  • Step S20 timing the playing time of the digital audio to determine the relationship between the playing time and the preset playing time threshold
  • Step S30 when it is determined that the playing time of the digital audio is greater than the preset playing time threshold, a time series prediction model is started, and the preset playing time threshold is obtained according to the digital audio training before the preset playing time threshold Later music accompaniment.
  • step S3 also includes:
  • the digitized audio is stored as training data of a non-time series prediction model. Doing so can better provide sufficient training data for non-time series models for subsequent non-time series model training and prediction.
  • Step S40 When it is determined that the complete playing time of the digital audio is less than the preset playing time threshold, store the digital audio as training data of a non-time series prediction model.
  • the next step is to make predictions based on the stored digital audio.
  • the preset playing time threshold is set to 30 seconds, when the player plays continuously
  • the time series model is started to predict the musical accompaniment after 30 seconds.
  • the audio signal is stored as digital audio For training and prediction of non-time series prediction models.
  • the music prediction model of this embodiment uses a time series prediction model and a non-time series prediction model.
  • the specific model prediction methods are as follows:
  • step S30 the time series prediction model is commonly known as online prediction.
  • the model will recursively modify the output connection weight w through the 30 seconds of performance data, and then predict the output regularly So as to achieve the purpose of assisting the player to play.
  • time series prediction model The entire time series prediction model is divided into model training and model prediction. details as follows:
  • Time series prediction is to obtain the true value of a system-related variable within a period of time, and then use the echo state network algorithm to predict the future value of one or some variables of this system.
  • the variables predicted by this model are the sampling frequency and digits of music.
  • the echo state network is a simplified recursive neural network model, which can effectively avoid the disadvantage of slow convergence speed of the recursive neural network learning algorithm. It has the characteristics of high computational complexity and is particularly suitable for use in interactive devices. The main reason for using it for time series forecasting.
  • the echo state network is composed of three parts. As shown in FIG. 2, FIG. 2 is a schematic structural diagram of an echo state network model provided by an embodiment of the present application.
  • the large circle 001 in the middle shows the reserve pool x t , and w t is the estimated value of the reserve pool weight at time t.
  • the left part 002 represents the input neurons of real data, that is, the sampling frequency and the number of bits of music, collectively called the measured value
  • the right part 003 represents the output neuron y t predicted by the model.
  • the reserve pool is composed of a large number of neurons (the number is usually several hundred), and the neurons inside the reserve pool use sparse connections (sparse connection means that the neurons are only partially connected, as shown in the above figure), and between the neurons
  • the connection weights are randomly generated and remain fixed after the connection weights are generated, that is, the connection weights of the reserve pool do not require training.
  • the external data enters the reserve pool through input neurons and is predicted, and finally output y t is output by the output neurons.
  • Kalman filtering is an optimization method for numerical estimation. It can be used in any dynamic system with uncertain information. It can make an educated prediction of the next direction of the system. Therefore, using Kalman filtering to train the echo state network can Efficiently improve the accuracy of the time series prediction model. Combined with the Kalman filter method equation formula, at time t + 1:
  • ⁇ t and ⁇ t are the process noise and measurement noise of Kalman filter at time t , respectively, and their covariance matrices are q t and r t respectively .
  • Model prediction stage time the playing time to determine whether the playing time exceeds the preset playing time threshold
  • the device when the user starts playing using the interactive device, the device starts two steps at the same time. First, time the playing time; second, store the digital audio.
  • the purpose of digital audio storage is to store enough training data for non-time series prediction model training.
  • the preset threshold for playing time is 30 seconds. Once the playing time exceeds the threshold of 30 seconds, the time series prediction model based on the trained echo state network begins to work, outputting musical accompaniment to assist the player to play;
  • the time series prediction model does not work, but the playing data will be converted to digital audio through an interactive device and stored in memory as training data for non-time series prediction model training.
  • the reason for setting the playing time threshold is to ensure that there is enough audio storage to improve the prediction accuracy.
  • the non-time series prediction model corresponds to the time series prediction model.
  • the audio signal will be converted into digital audio and stored in the interactive device. Based on the stored digital audio every time, the interactive device will train and predict it.
  • This method based on offline training and prediction is called non-temporal prediction model.
  • a deep convolutional generation adversarial network method (Deep Convolutional Generative Adversarial Nerworks, DCGAN) is used to predict non-time series.
  • the main steps include:
  • Step S401 is mainly to extract the digital audio previously stored in the interactive device.
  • Step S402 trains the generative confrontation network based on the extracted data.
  • the reason for using this network is because the player ’s energy is limited, and then the amount of digital audio data stored in the interactive device is not large.
  • deep convolution is used to generate an adversarial network Automatically generate data while training music melody to achieve a double effect.
  • the DCGAN network model includes a generating network G and a discriminating network D.
  • the objective function of DCGAN is based on the minimum and maximum values of the generating network G and the discriminating network D.
  • Figure 3 is a schematic diagram 3 of the training process of the DCGAN network model.
  • the generating network G When generating a generator against the network training, first use the generating network G to randomly randomize the audio noise Z (audio noise is stored in the DCGAN in advance Digitized random audio data, which is not regular music melody data) Generate realistic digital audio samples, and at the same time discriminate network D to train a discriminator to discriminate real digital audio X (real digital audio refers to the Melody of digital audio) and the generated digital audio samples.
  • the whole process allows the generator and the discriminator to train at the same time until the loss function values of the generating network G and the discriminating network D reach a preset threshold, which proves that the model training is successful at this time and has the ability to predict the music melody.
  • the digital audio data generated by the generation network of the model has a high degree of similarity to the real sample. Even if the network is discriminated, the difference between the digital audio data generated by the generation network and the real data cannot be distinguished.
  • the loss function of generating network G is:
  • the loss function for discriminating network D is:
  • x represents the input parameter, that is, the digitized audio extracted in step (1)
  • y refers to the digitized audio value predicted by the DCGAN generating network G and the discriminating network D.
  • both the generation network and the discriminant network of DCGAN are convolutional neural networks.
  • the present application also provides an automatic music generation device.
  • FIG. 4 it is a schematic diagram of the internal structure of an apparatus for automatically generating music provided by an embodiment of the present application.
  • the automatic music generation device 1 may be a PC (Personal Computer) or a terminal device such as a smart phone, tablet computer, or portable computer.
  • the automatic music generation device 1 includes at least a memory 11, a processor 12, a communication bus 13, and a network interface 14.
  • the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, and the like.
  • the memory 11 may be an internal storage unit of the automatic music generation device 1 in some embodiments, such as the hard disk of the automatic music generation device 1.
  • the memory 11 may also be an external storage device of the automatic music generation device 1, for example, a plug-in hard disk equipped on the automatic music generation device 1, a smart memory card (Smart Media (SMC), secure digital (Secure) Digital, SD) card, flash card (Flash Card), etc.
  • the memory 11 may also include both the internal storage unit of the music automatic generating apparatus 1 and an external storage device.
  • the memory 11 can be used not only to store application software and various types of data installed in the music automatic generation device 1, such as codes of the music automatic generation program 01, but also to temporarily store data that has been or will be output.
  • the processor 12 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip for running the program code or processing stored in the memory 11 Data, for example, the automatic music generation program 01 is executed.
  • CPU central processing unit
  • controller microcontroller
  • microprocessor or other data processing chip for running the program code or processing stored in the memory 11 Data, for example, the automatic music generation program 01 is executed.
  • the communication bus 13 is used to realize connection and communication between these components.
  • the network interface 14 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the device 1 and other electronic devices.
  • the music automatic generation device 1 may further include a user interface
  • the user interface may include a display (Display), an input unit such as a keyboard (Keyboard), and the optional user interface may also include a standard wired interface and a wireless interface.
  • the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode, organic light emitting diode) touch device, or the like.
  • the display may also be appropriately referred to as a display screen or a display unit, for displaying information processed in the automatic music generation device 1 and for displaying a visual user interface.
  • FIG. 4 only shows the music automatic generation device 1 having the components 11-14 and the music automatic generation program 01.
  • FIG. 1 does not constitute a limitation on the music automatic generation device 1 , May include fewer or more components than shown, or combine certain components, or a different component arrangement.
  • the automatic music generation program 01 is stored in the memory 11; the processor 12 implements the following steps when executing the automatic music generation program 01 stored in the memory 11:
  • Step S10 Collect audio signals of music melody and convert the audio signals into digital audio storage
  • step S10 further includes:
  • S101 Use an audio amplifier to collect the sampling frequency and sampling digits of the audio signal
  • the task of collecting audio signals is to discretize continuous sound waveforms, that is, collecting music analog signals.
  • a continuous signal with limited bandwidth can be replaced by a sequence of discrete sampling points, and this substitution will not lose any information.
  • Fourier theory also pointed out: All complex periodic waveforms are composed of a series of sine waves arranged in harmonics. Complex waveforms can be synthesized by the summation and summation of multiple sine waves. Therefore, the audio signal is discretely sampled according to the system, the audio signal is defined at each exact time point, and the audio signal to be collected can be collected.
  • sampling frequency Sample Rate
  • sampling digit the sampling digit is the amplitude dynamic response data range of each sampling point
  • storage (sampling frequency * sampling digits) / 8 (number of bytes)
  • the audio amplifier uses a sampling frequency of 22.05kHz and a sampling number of 8 bits.
  • the sampling frequency must be at least twice the highest frequency of the signal. The higher the sampling frequency, the smaller the sound distortion and the greater the amount of audio data. Therefore, in general, the upper frequency limit of human ear hearing is about 2OkHz.
  • the sampling frequency In order to ensure that the sound is not distorted, the sampling frequency should be about 4OkHz, but there is no concert to reach the frequency of 20kHz, because high frequency will affect the listener ’s hearing experience Due to the resonance effect of music, the sampling frequency used in the audio amplifier is 22.05kHz.
  • 8-bit, 12-bit and 16-bit are often used for sampling digits.
  • 8-bit quantization level means that each sampling point can represent 256 (2 8 ) different quantization values
  • 16-bit quantization level can represent 65536 different quantization values .
  • the higher the number of sampling and quantization bits the better the sound quality and the larger the data volume.
  • the audio amplifier uses 8-bit sampling bits for processing.
  • step S102 further includes:
  • the audio signal is passed through a low-pass filter, and the audio signal higher than the half sampling frequency is band-limited to improve aliasing interference.
  • the aliasing interference phenomenon that is, an input signal higher than the half sampling frequency will produce a lower frequency aliasing signal, in which the half sampling evaluation rate is half of the sampling frequency.
  • the sampling frequency of the audio amplifier is 22.05kHz.
  • an interfering alias signal will be generated.
  • the following data cleaning method is adopted for the aliasing interference signal: after the audio amplifier collects the audio signal, a low-pass filter is added. The collected audio signal is band-limited by a low-pass filter (anti-aliasing filter), which provides sufficient attenuation at the half-sampling frequency to ensure that the sampled signal does not contain spectrum exceeding the half-sampling frequency content.
  • step S102 further includes:
  • the noise emitted by the jitter generator is collected and added to the audio signal to improve quantization error interference.
  • the amplitude value is rounded to the nearest quantization scale value. This operation will cause a quantization error.
  • an error When quantizing the amplitude of the audio signal, an error will occur between the real analog value and the selected quantization scale value. , Which is the quantization error.
  • This quantization error results in the inability to perfectly encode a continuous analog function when digitally storing audio signals.
  • Data cleaning method based on quantization error interference when an audio amplifier is collecting audio signals, a small amount of noise generated by the jitter generator is also collected. Because the jitter itself is a small amplitude noise that is not related to the audio signal, it is added to the audio signal of the interactive device before the audio signal is sampled.
  • the audio signal After adding the dithering signal, the audio signal will be shifted for each quantization level. For the previous waveforms that are adjacent in time, because each cycle is different now, there will be no periodic quantization mode, because the quantization error is closely related to the signal cycle, so the various effects of the final quantization error , Also randomized enough to be removed.
  • the digital signal After adding the low-pass filter and jitter generator to solve the data cleaning problem, the digital signal is finally converted into digital audio and stored in the interactive device by the digital converter, and the audio data collection process ends.
  • Step S20 timing the playing time of the digital audio to determine the relationship between the playing time and the preset playing time threshold
  • Step S30 when it is determined that the playing time of the digital audio is greater than the preset playing time threshold, a time series prediction model is started, and the preset playing time threshold is obtained according to the digital audio training before the preset playing time threshold Later music accompaniment;
  • step S30 also includes:
  • the digitized audio is stored as training data of a non-time series prediction model. Doing so can better provide sufficient training data for non-time series models for subsequent non-time series model training and prediction.
  • Step S40 When it is determined that the complete playing time of the digital audio is less than the preset playing time threshold, store the digital audio as training data of a non-time series prediction model.
  • the next step is to make predictions based on the stored digital audio, and set the preset playing time threshold to 30 seconds, when the player plays continuously.
  • the preset playing time threshold exceeds 30 seconds
  • the time series model is started to predict the music accompaniment after 30 seconds.
  • the audio signal is stored as digital audio to For non-time series prediction model training and prediction.
  • the music prediction model of this embodiment uses a time series prediction model and a non-time series prediction model.
  • the specific model prediction methods are as follows:
  • step S30 the time series prediction model is commonly known as online prediction.
  • the model will recursively modify the output connection weight w through the 30 seconds of performance data, and then predict the output regularly. So as to achieve the purpose of assisting the player to play.
  • time series prediction model The entire time series prediction model is divided into model training and model prediction. details as follows:
  • Time series prediction is to obtain the true value of a system-related variable within a period of time, and then use the echo state network algorithm to predict the future value of one or some variables of this system.
  • the variables predicted by this model are the sampling frequency and digits of music.
  • the echo state network is a simplified recursive neural network model, which can effectively avoid the disadvantage of slow convergence speed of the recursive neural network learning algorithm. It has the characteristics of high computational complexity and is particularly suitable for use in interactive devices. This is the case in this embodiment. The main reason for using it for time series forecasting.
  • the echo state network is composed of three parts. As shown in FIG. 2, FIG. 2 is a schematic structural diagram of an echo state network model provided by an embodiment of the present application.
  • the large circle 001 in the middle shows the reserve pool x t , and w t is the estimated value of the reserve pool weight at time t.
  • the left part 002 represents the input neurons of real data, that is, the sampling frequency and the number of bits of music, collectively called the measured value
  • the right part 003 represents the output neuron y t predicted by the model.
  • the reserve pool is composed of a large number of neurons (the number is usually several hundred), and the neurons inside the reserve pool use sparse connections (sparse connection means that the neurons are only partially connected, as shown in the above figure), and between the neurons
  • the connection weights are randomly generated and remain fixed after the connection weights are generated, that is, the connection weights of the reserve pool do not require training.
  • the external data enters the reserve pool through input neurons and is predicted, and finally output y t is output by the output neurons.
  • Kalman filtering is an optimization method for numerical estimation. It can be used in any dynamic system with uncertain information. It can make an educated prediction of the next direction of the system. Therefore, using Kalman filtering to train the echo state network can Efficiently improve the accuracy of the time series prediction model. Combined with the Kalman filter method equation formula, at time t + 1:
  • ⁇ t and ⁇ t are the process noise and measurement noise of Kalman filter at time t , respectively, and their covariance matrices are q t and r t respectively .
  • Model prediction stage time the playing time to determine whether the playing time exceeds the preset playing time threshold
  • the device when the user starts playing using the interactive device, the device starts two steps at the same time. First, time the playing time; second, store the digital audio.
  • the purpose of digital audio storage is to store enough training data for non-time series prediction model training.
  • the preset threshold for playing time is 30 seconds. Once the playing time exceeds the threshold of 30 seconds, the time series prediction model based on the trained echo state network begins to work, outputting musical accompaniment to assist the player to play;
  • the time series prediction model does not work, but the playing data will be converted to digital audio through an interactive device and stored in memory as training data for non-time series prediction model training.
  • the reason for setting the playing time threshold is to ensure that there is enough audio storage to improve the prediction accuracy.
  • the non-time series prediction model corresponds to the time series prediction model.
  • the audio signal will be converted into digital audio and stored in the interactive device. Based on the stored digital audio every time, the interactive device will train and predict it.
  • This method based on offline training and prediction is called non-temporal prediction model.
  • a deep convolutional generation adversarial network method (Deep Convolutional Generative Adversarial Nerworks, DCGAN) is used to predict non-time series.
  • the main steps include:
  • Step S401 is mainly to extract the digital audio previously stored in the interactive device.
  • Step S402 trains the generative confrontation network based on the extracted data.
  • the reason for using this network is because the player ’s energy is limited, and then the amount of digital audio data stored in the interactive device is not large.
  • deep convolution is used to generate an adversarial network Automatically generate data while training music melody to achieve a double effect.
  • the DCGAN network model includes a generating network G and a discriminating network D.
  • the objective function of DCGAN is based on the minimum and maximum values of the generating network G and the discriminating network D.
  • Figure 3 is a schematic diagram 3 of the training process of the DCGAN network model.
  • the digital audio data generated by the generation network of the model has a high degree of similarity to the real sample. Even if the network is discriminated, the difference between the digital audio data generated by the generation network and the real data cannot be distinguished.
  • the loss function of the generation network G is :
  • the loss function for discriminating network D is:
  • x represents the input parameter, that is, the digitized audio extracted in step (1)
  • y refers to the digitized audio value predicted by the DCGAN generating network G and the discriminating network D.
  • both the generation network and the discriminant network of DCGAN are convolutional neural networks.
  • the music automatic generation program may also be divided into one or more modules, and the one or more modules are stored in the memory 11 and are controlled by one or more processors (this embodiment is The processor 12) is executed to complete the application.
  • the module referred to in the application refers to a series of computer program instruction segments capable of performing specific functions, and is used to describe the execution process of the music automatic generation program in the music automatic generation device.
  • FIG. 5 it is a schematic diagram of a program module of a music automatic generation program in an embodiment of an automatic music generation device of the present application.
  • the time counting module 20, the time series prediction model 30, and the non-time series prediction model 40 exemplarily:
  • the audio signal collection module 10 is used to collect audio signals of music melody and convert the audio signals into digital audio storage;
  • the playing time timing module 20 is used to time the playing time of the digital audio and determine the relationship between the playing time and the preset playing time threshold;
  • Time series prediction model 30 used to judge that the playing time of the digital audio is greater than the preset playing time threshold, start the time series prediction model, and obtain a preset according to the digital audio training before the preset playing time threshold Music accompaniment after playing time threshold;
  • the non-time series prediction model 40 is used to judge that the complete playing time of the digital audio is less than the preset playing time threshold, and store the digital audio as training data of the non-time series prediction model.
  • the above-mentioned audio signal acquisition module 10, playing time timing module 20, time series prediction model 30, and non-time series prediction model 40 and other program modules are implemented when the functions or operation steps are substantially the same as the above embodiments, and are not repeated here Repeat.
  • the embodiments of the present application also provide a computer-readable storage medium on which a music automatic generation program is stored, and the music automatic generation program may be executed by one or more processors to implement the following operating:
  • the digital audio is stored as training data of a non-time series prediction model.
  • the methods in the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware, but in many cases the former is better Implementation.
  • the technical solution of the present application can be embodied in the form of a software product in essence or part that contributes to the existing technology, and the computer software product is stored in a storage medium (such as ROM / RAM as described above) , Magnetic disks, optical disks), including several instructions to enable a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to perform the method described in each embodiment of the present application.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • General Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Auxiliary Devices For Music (AREA)
  • Electrophonic Musical Instruments (AREA)

Abstract

提供了一种音乐自动生成方法,该方法包括:采集音乐旋律的音频信号,将音频信号转化为数字化音频存储(S10);对数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系(S20);当判断数字化音频的弹奏时间大于预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏(S30);当判断数字化音频的完整弹奏时间小于预设弹奏时间阈值时,将数字化音频存储为非时间序列预测模型的训练数据(S40)。还提供一种音乐自动生成装置以及计算机可读存储介质。由此通过预判断音乐的弹奏时间,分不同的预测模型预测音乐旋律,提高了模型的鲁棒性和自适应性。

Description

一种音乐自动生成方法、装置及计算机可读存储介质
本申请要求于2018年11月12日提交中国专利局,申请号为201811341758.6、发明名称为“一种音乐自动生成方法、装置及计算机可读存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及智能决策技术领域,尤其涉及一种音乐自动生成方法、装置及计算机可读存储介质。
背景技术
声音是由物体振动产生的声波,通过介质(空气或固体、液体)传播并能被人或动物听觉器官所感知的波动现象。音乐属于一种特殊的声音模式,当演奏乐器时,乐器的振动会引起介质(空气分子)有节奏的振动,使周围的空气产生疏密变化,形成疏密相间的纵波,这就产生了所谓的音乐(这种现象会一直延续到振动消失为止)。
科学的音乐旋律预测迄今为止己有很多的方法,从预测性质上分为定量和定性两种。定性分析一般来说就是运用归纳、演绎、分析、综合及抽象与概括等方法进行分析;而定量分析通常包含两个方面的内容:因果关系研究、统计分析。但不管是利用哪种方法进行预测,都属于传统的简单模型预测,其音乐旋律的精准度不高。为了提高预测精准度,通常需要将多种传统预测方法进行比较取最好的方法或将多种预测方法结合起来进行预测,常用的统计分析模型主要有;指数平滑法,趋势外推法,移动平均法等。但当音乐旋律数据以时间序列的形式存在时,这些数据有时是线性关系,有时却是非线性关系,此时即使多种传统的预测方法结合起来,其精度也有待提高。
发明内容
本申请提供一种音乐自动生成方法、装置及计算机可读存储介质,其主要目的在于提高自动生成的音乐的精准度。
为实现上述目的,本申请还提供一种音乐自动生成方法,该方法包括:
采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储;
对所述数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系;
判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏;
判断所述数字化音频的完整弹奏时间小于所述预设弹奏时间阈值时,将所述数字化音频存储为非时间序列预测模型的训练数据。
此外,为实现上述目的,本申请还提供一种音乐自动生成装置,该装置包括存储器和处理器,所述存储器中存储有可在所述处理器上运行的程序,所述程序被所述处理器执行时实现如下步骤:
采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储;
对所述数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系;
判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏;
判断所述数字化音频的完整弹奏时间小于所述预设弹奏时间阈值时,将所述数字化音频存储为非时间序列预测模型的训练数据。
此外,为实现上述目的,本申请还提供一种计算机可读存储介质,所述计算机可读存储介质上存储有音乐自动生成程序,所述程序可被一个或者多个处理器执行,以实现如上所述的方法的步骤。
本申请提出的音乐自动生成方法、装置及计算机可读存储介质,通过预判断音乐旋律的弹奏时间,分不同的预测模型进行预测音乐旋律,提高了整个模型的鲁棒性和自适应性。
附图说明
图1为本申请一实施例提供的音乐自动生成方法的流程示意图;
图2为本申请一实施例提供的回声状态网络模型结构示意图;
图3为本申请一实施例提供的DCGAN网络模型训练流程示意图;
图4为本申请一实施例提供的音乐自动生成装置的内部结构示意图;
图5为本申请一实施例提供的音乐自动生成装置中程序的模块示意图。
本申请目的的实现、功能特点及优点将结合实施例,参照附图做进一步说明。
具体实施方式
应当理解,此处所描述的具体实施例仅仅用以解释本申请,并不用于限定本申请。
本申请提供一种音乐自动生成方法。参照图1所示,为本申请一实施例提供的音乐自动生成方法的流程示意图。本方法可以采用具有声卡(Digital to Analog Converter,DAC),中文称数模转换器的各类交互式装备,如手机、平板、电脑等,作为演奏的装置,实现本实施例的方法。上述各类交互式设备可以由软件和/或硬件实现。
在本实施例中,一种音乐自动生成方法包括:
步骤S10,采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储。
进一步的,步骤S10还包括:
S101:利用音频放大器采集所述音频信号的采样频率和采样数位;
因为音乐属于声音,是通过波形传播,采集音频信号的任务是将连续的声音波形离散化,即采集音乐模拟信号。根据奈奎斯特在1924年指出的采样定理:一个带宽受限的连续信号可以用一个离散的采样点序列替代,这种替代不会丢失任何信息。且傅里叶理论也指出:所有复杂的周期波形都由一系列按谐波排列的正弦波组成,复杂波形可以有多个正弦波的累加求和而合成出来。所以根据系统对音频信号进行离散采样,在各个确切的时间点上定义音频信号,可以采集到所要收集的音频信号。
演奏者通过交互式设备进行弹奏时,采集音频信号,在整个采集过程中,主要采集音频信号的采样频率(Sample Rate,频率是对音乐波形每秒钟所采样的次数)和采样数位,也可叫做采样精度(Quantizing,也称量化级,采样数位是每个采样点的振幅动态响应数据范围)两个方面,因为这二者决定了 数字化音频的质量,即决定了后期深度学习预测音乐模型的鲁棒性。在本实施例中,利用音频放大器对音频信号的采样频率和采样精度进行采集,且结合交互式设备处理器性能和存储能力(存储量=(采样频率*采样数位)/8(字节数)),在不影响本方案深度模型训练的前提下,音频放大器采用22.05kHz的采样频率和8位的采样位数。因为根据奈奎斯特采样定理:采样频率必须至少为信号最高频率的两倍,采样频率越高,声音失真越小、音频数据量也越大。所以综合实际来说,人耳听觉的频率上限在2OkHz左右,为了保证声音不失真,采样频率应在4OkHz左右,但是没有音乐会达到20kHz的频率,因为高频会影响听众的听觉感受,达不到音乐引起共鸣的效果,所以音频放大器中使用的采样频率为22.05kHz。采样数位经常采用的有8位、12位和16位,例如8位量化级表示每个采样点可以表示256(2 8)个不同量化值,而16位量化级则可表示65536个不同量化值,采样量化位数越高音质越好,数据量也越大。但结合交互式设备处理器性能,音频放大器的处理环节采用8位的采样位数。
S102:对所述音频信号进行数据清洗。
即使用最复杂的技术,一个交互式设备的音频系统所能重现出的声音,也仅仅是真实声音的近似声音。而数据清洗是通过各种技术,缩小音频系统所存储的音乐和真实音乐的差距。上述通过音频放大器采集到音频信号,会产生很多干扰,因此需要对采集的音频数据进行清洗;在音频数据的采集阶段加入清洗步骤,减小了音频数据的噪声干扰。
进一步的,步骤S102还包括:
将所述音频信号通过低通滤波器,对高于半采样频率的音频信号进行限带处理,以改善混叠干扰。
混叠干扰现象,即一个高于半采样频率的输入信号将产生一个频率较低的混叠信号,其中,半采样评率为采样频率的一半。例如,音频放大器的采样频率是22.05kHz,当音频信号的频率高于半采样频率11.025kHz时,就会产生一个干扰的混叠信号。针对混叠干扰信号采取如下数据清洗方法:在音频放大器采集完音频信号后,加入一个低通滤波器。将采集的音频信号通过一个低通滤波器(抗混叠滤波器)进行限带处理,这就在半采样频率处提供了足够的衰减,从而确保被采样信号中不包含超过半采样频率的频谱内容。
进一步的,步骤S102还包括:
在采集所述音频信号的同时,采集抖动发生器发出的噪声,并将所述噪声加入到所述音频信号中,以改善量化误差干扰。
在采样时刻,幅度值被舍入到最近的量化分度值上,这一操作将导致量化误差,在对音频信号的幅度进行量化时,真实模拟值与所选的量化分度值会产生误差,也就是量化误差。这种量化误差导致数字化存储音频信号时,不能对一个连续的模拟函数进行完美编码。根据量化误差干扰采取的数据清洗方法:在音频放大器采集音频信号时,同时采集抖动发生器发生的少量噪声。因为抖动本身是一个与音频信号不相关的幅度很小的噪声,它在音频信号采样之前被加入到交互式设备的音频信号中。在加入抖动信号以后,音频信号就会对各个量化级进行平移。对于之前时间上相邻的各个波形来说,因为现在每个周期都是不同的,所以就不会产生周期性的量化模式,因为量化误差是与信号周期息息相关,所以最终量化误差的各种影响,也被随机到足以使其得到去除的程度。
通过加入低通滤波器和抖动发生器解决数据清洗问题后,最后由数字转换器将音频信号转换为数字化音频存储到交互式设备中,音频数据的采集环节结束。
步骤S20,对所述数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系;
步骤S30,当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏。
进一步的,步骤S3还包括:
将所述数字化音频存储作为非时间序列预测模型的训练数据。这样做可以更好为非时间序列模型提供足够的训练数据,以供后续非时间序列模型训练和预测。
步骤S40,当判断所述数字化音频的完整弹奏时间小于所述预设弹奏时间阈值时,将所述数字化音频存储为非时间序列预测模型的训练数据。
交互式设备将用户弹奏的音乐旋律成功存储为数字化音频后,下一步就是根据所存储的数字化音频进行预测,例如将预设弹奏时间阈值设为30秒, 当演奏者不间断的弹奏时间超过预设弹奏时间阈值30秒时,启动时间序列模型预测30秒后的音乐伴奏,当演奏者完整弹奏时间不足预设弹奏时间阈值30秒时,将音频信号存储为数字化音频,以供非时间序列预测模型训练和预测。
本实施例音乐预测模型采用的是时间序列预测模型和非时间序列预测模型,具体模型预测方法分别如下:
在步骤S30中,时间序列预测模型俗称在线预测,当演奏者达到30秒的演奏时间时,模型会通过该30秒的演奏数据,递归的修改输出连接权值w,然后有规律的预测输出,从而达到辅助演奏者演奏的目的。
整个时间序列预测模型分为模型训练和模型预测。具体如下:
训练时间序列训练模型阶段:时间序列预测是先在一段时间内获得一个系统相关变量的真实值,然后用回声状态网络算法对这个系统的某个或某些变量未来的取值进行预测。本模型预测的变量是音乐的采样频率和采样数位。回声状态网络是一种简化的递归神经网络模型,可以有效避免递归神经网络学习算法收敛速度慢的缺点,具有计算复杂性高的特性,特别适合应用于交互式设备中,这是本实施例中采用其进行时间序列预测的主要原因。回声状态网络是由三个部分构成,如图2所示,图2为本申请一实施例提供的回声状态网络模型结构示意图。
结合音乐旋律在某t时刻来说,
中间部分的大圆圈001表示储备池x t,w t是t时刻储备池权值的估计值。
左边部分002表示真实数据的输入神经元,即音乐的采样频率和位数,统称为测量值
Figure PCTCN2018123593-appb-000001
右边部分003表示模型预测的输出神经元y t
储备池是由大量的神经元组成(数量通常为几百个),储备池内部的神经元采用稀疏连接(稀疏连接是指神经元之间只是部分连接,如上图所示),神经元之间的连接权值是随机生成的,并且连接权值生成后就保持固定不变,也就是储备池的连接权值不需要训练。外部数据通过输入神经元进入到储备池后预测,最终由输出神经元输出y t
对于回声状态网络的时间序列预测模型的训练,本实施例使用卡尔曼滤波法。卡尔曼滤波作为一种数值估计的优化方法,应用在任何含有不确定信息的动态系统中,对系统的下一步的走向都能做出有根据的预测,所以使用 卡尔曼滤波训练回声状态网络可以高效的提升时间序列预测模型的准确率。结合卡尔曼滤波法的方程公式,在t+1时刻时有:
w t+1=w tt
Figure PCTCN2018123593-appb-000002
其中α t、β t分别为卡尔曼滤波在t时刻的过程噪声和测量噪声,其协方差矩阵分别为q t、r t。而对于t时刻的时间序列模型,由以下步骤可得:
p t=p t-1+q t-1
Figure PCTCN2018123593-appb-000003
Figure PCTCN2018123593-appb-000004
其中p t是协方差矩阵,k t是卡尔曼滤波器的增益。同理可得t-1、t-2等时刻的状态量。由以上所述,可以更新储备池内的权重,达到训练时间序列预测模型的目的。
模型预测阶段:对弹奏时间计时,判断弹奏时间是否超过预设弹奏时间阈值;
进一步的,本实施例中用户在使用交互式设备开始弹奏时,设备同时启动两个步骤,一、对弹奏时间计时;二、将数字化音频存储。数字化音频存储的目的是为了存储足够训练数据供非时间序列预测模型训练使用。
设定的预设弹奏时间阈值为30秒。一旦弹奏时间超过阈值30秒后,基于已训练好的回声状态网络的时间序列预测模型开始工作,输出音乐伴奏,辅助演奏者弹奏;
当弹奏的完整时间不足30秒时,时间序列预测模型不工作,但弹奏数据会通过交互式设备转为数字化音频存储到内存中,作为训练数据供非时间序列预测模型训练。设定弹奏时间阈值的原因是为了保证具有足够的音频存储量,以提高预测准确率。
在步骤S40中,与时间序列预测模型对应的是非时间序列预测模型。当弹奏者弹奏出音乐旋律时,音频信号会被转化成数字化音频被存储到交互式设备中,基于每次存储的数字化音频,交互式设备都会对其进行训练和预测。这种基于离线训练和预测的方法称为非时间预测模型。本实施例采用深度卷积生成对抗网络法(Deep Convolutional Generative Adversarial Nerworks, DCGAN)对非时间序列进行预测。主要步骤包括:
S401:提取存储的数字化音频;
S402:训练深度卷积生成对抗网络;
S403:根据用户需求播放预测的音乐伴奏。
其中步骤S401主要是将之前在交互式设备存储的数字化音频提取出来。步骤S402根据所提取的数据进行生成式对抗网络的训练。使用该网络的原因是因为弹奏者的精力有限,继而在交互式设备中所存储的数字化音频的数据量并不多,针对这种样本数据量不够多的问题,使用深度卷积生成对抗网络自动生成数据的同时也训练音乐旋律,达到两重的效果。在本实施例中,DCGAN网络模型包含一个生成网络G和一个判别网络D,DCGAN的目标函数是基于生成网络G与判别网络D的最小值和最大值的问题。如图3所示,图3为DCGAN网络模型训练流程示意图3,当生成对抗网络训练一个生成器时,首先利用生成网络G,从随机的数字化音频噪声Z(音频噪声是提前在DCGAN中存储的数字化随机音频数据,其并不是有规律的音乐旋律数据)中生成逼真的数字化音频样本,同时判别网络D训练一个鉴别器来鉴别真实数字化音频X(真实数字化音频是指在步骤一所存储的具有旋律的数字化音频)和生成的数字化音频样本之间的差距。整个过程让生成器和鉴别器同时训练,直到生成网络G与判别网络D的损失函数值都达到预设的某阈值时,证明此时模型训练成功,具有预测音乐旋律的能力。此时模型的生成网络生成的数字化音频数据与真实的样本具有很高的相似度,即使判别网络也无法区分生成网络生成的数字化音频数据和真实数据的差异,
其中,生成网络G的损失函数为:
(1-y)lg(1-D(G(Z)))
判别网络D的损失函数为:
-((1-y)lg(1-D(G(Z)))+ylgD(x))
其中,x表示输入参数,即步骤(1)提取的数字化音频,y指DCGAN的生成网络G与判别网络D所预测的数字化音频值。特别要强调是,DCGAN的生成网络和判别网络都是卷积神经网络。基于以上所述,训练成功的非时间序列预测模型可自动生成音乐伴奏,供演奏者使用学习。
本申请还提供一种音乐自动生成装置。参照图4所示,为本申请一实施例提供的音乐自动生成装置的内部结构示意图。
在本实施例中,音乐自动生成装置1可以是PC(Personal Computer,个人电脑),也可以是智能手机、平板电脑、便携计算机等终端设备。该音乐自动生成装置1至少包括存储器11、处理器12,通信总线13,以及网络接口14。
其中,存储器11至少包括一种类型的可读存储介质,所述可读存储介质包括闪存、硬盘、多媒体卡、卡型存储器(例如,SD或DX存储器等)、磁性存储器、磁盘、光盘等。存储器11在一些实施例中可以是音乐自动生成装置1的内部存储单元,例如该音乐自动生成装置1的硬盘。存储器11在另一些实施例中也可以是音乐自动生成装置1的外部存储设备,例如音乐自动生成装置1上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。进一步地,存储器11还可以既包括音乐自动生成装置1的内部存储单元也包括外部存储设备。存储器11不仅可以用于存储安装于音乐自动生成装置1的应用软件及各类数据,例如音乐自动生成程序01的代码等,还可以用于暂时地存储已经输出或者将要输出的数据。
处理器12在一些实施例中可以是一中央处理器(Central Processing Unit,CPU)、控制器、微控制器、微处理器或其他数据处理芯片,用于运行存储器11中存储的程序代码或处理数据,例如执行音乐自动生成程序01等。
通信总线13用于实现这些组件之间的连接通信。
网络接口14可选的可以包括标准的有线接口、无线接口(如WI-FI接口),通常用于在该装置1与其他电子设备之间建立通信连接。
可选地,该音乐自动生成装置1还可以包括用户接口,用户接口可以包括显示器(Display)、输入单元比如键盘(Keyboard),可选的用户接口还可以包括标准的有线接口、无线接口。可选地,在一些实施例中,显示器可以是LED显示器、液晶显示器、触控式液晶显示器以及OLED(Organic Light-Emitting Diode,有机发光二极管)触摸器等。其中,显示器也可以适当的称为显示屏或显示单元,用于显示在音乐自动生成装置1中处理的信息以及用于显示可视化的用户界面。
图4仅示出了具有组件11-14以及音乐自动生成程序01的音乐自动生成装置1,本领域技术人员可以理解的是,图1示出的结构并不构成对音乐自动生成装置1的限定,可以包括比图示更少或者更多的部件,或者组合某些部件,或者不同的部件布置。
在图4所示的音乐自动生成装置1实施例中,存储器11中存储有音乐自动生成程序01;处理器12执行存储器11中存储的音乐自动生成程序01时实现如下步骤:
步骤S10,采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储;
进一步的,步骤S10还包括:
S101:利用音频放大器采集所述音频信号的采样频率和采样数位;
因为音乐属于声音,是通过波形传播,采集音频信号的任务是将连续的声音波形离散化,即采集音乐模拟信号。根据奈奎斯特在1924年指出的采样定理:一个带宽受限的连续信号可以用一个离散的采样点序列替代,这种替代不会丢失任何信息。且傅里叶理论也指出:所有复杂的周期波形都由一系列按谐波排列的正弦波组成,复杂波形可以有多个正弦波的累加求和而合成出来。所以根据系统对音频信号进行离散采样,在各个确切的时间点上定义音频信号,可以采集到所要收集的音频信号。
演奏者通过交互式设备进行弹奏时,采集音频信号,在整个采集过程中,主要采集音频信号的采样频率(Sample Rate,频率是对音乐波形每秒钟所采样的次数)和采样数位,也可叫做采样精度(Quantizing,也称量化级,采样数位是每个采样点的振幅动态响应数据范围)两个方面,因为这二者决定了数字化音频的质量,即决定了后期深度学习预测音乐模型的鲁棒性。在本实施例中,利用音频放大器对音频信号的采样频率和采样精度进行采集,且结合交互式设备处理器性能和存储能力(存储量=(采样频率*采样数位)/8(字节数)),在不影响本方案深度模型训练的前提下,音频放大器采用22.05kHz的采样频率和8位的采样位数。因为根据奈奎斯特采样定理:采样频率必须至少为信号最高频率的两倍,采样频率越高,声音失真越小、音频数据量也越大。所以综合实际来说,人耳听觉的频率上限在2OkHz左右,为了保证声音不失真,采样频率应在4OkHz左右,但是没有音乐会达到20kHz的频率,因 为高频会影响听众的听觉感受,达不到音乐引起共鸣的效果,所以音频放大器中使用的采样频率为22.05kHz。采样数位经常采用的有8位、12位和16位,例如8位量化级表示每个采样点可以表示256(2 8)个不同量化值,而16位量化级则可表示65536个不同量化值,采样量化位数越高音质越好,数据量也越大。但结合交互式设备处理器性能,音频放大器的处理环节采用8位的采样位数。
S102:对所述音频信号进行数据清洗。
即使用最复杂的技术,一个交互式设备的音频系统所能重现出的声音,也仅仅是真实声音的近似声音。而数据清洗是通过各种技术,缩小音频系统所存储的音乐和真实音乐的差距。上述通过音频放大器采集到音频信号,会产生很多干扰,因此需要对采集的音频数据进行清洗;在音频数据的采集阶段加入清洗步骤,减小了音频数据的噪声干扰。
进一步的,步骤S102还包括:
将所述音频信号通过低通滤波器,对高于半采样频率的音频信号进行限带处理,以改善混叠干扰。
混叠干扰现象,即一个高于半采样频率的输入信号将产生一个频率较低的混叠信号,其中,半采样评率为采样频率的一半。例如,音频放大器的采样频率是22.05kHz,当音频信号的频率高于半采样频率11.025kHz时,就会产生一个干扰的混叠信号。针对混叠干扰信号采取如下数据清洗方法:在音频放大器采集完音频信号后,加入一个低通滤波器。将采集的音频信号通过一个低通滤波器(抗混叠滤波器)进行限带处理,这就在半采样频率处提供了足够的衰减,从而确保被采样信号中不包含超过半采样频率的频谱内容。
进一步的,步骤S102还包括:
在采集所述音频信号的同时,采集抖动发生器发出的噪声,并将所述噪声加入到所述音频信号中,以改善量化误差干扰。
在采样时刻,幅度值被舍入到最近的量化分度值上,这一操作将导致量化误差,在对音频信号的幅度进行量化时,真实模拟值与所选的量化分度值会产生误差,也就是量化误差。这种量化误差导致数字化存储音频信号时,不能对一个连续的模拟函数进行完美编码。根据量化误差干扰采取的数据清洗方法:在音频放大器采集音频信号时,同时采集抖动发生器发生的少量噪 声。因为抖动本身是一个与音频信号不相关的幅度很小的噪声,它在音频信号采样之前被加入到交互式设备的音频信号中。在加入抖动信号以后,音频信号就会对各个量化级进行平移。对于之前时间上相邻的各个波形来说,因为现在每个周期都是不同的,所以就不会产生周期性的量化模式,因为量化误差是与信号周期息息相关,所以最终量化误差的各种影响,也被随机到足以使其得到去除的程度。
通过加入低通滤波器和抖动发生器解决数据清洗问题后,最后由数字转换器将音频信号转换为数字化音频存储到交互式设备中,音频数据的采集环节结束。
步骤S20,对所述数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系;
步骤S30,当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏;
进一步的,步骤S30还包括:
将所述数字化音频存储作为非时间序列预测模型的训练数据。这样做可以更好为非时间序列模型提供足够的训练数据,以供后续非时间序列模型训练和预测。
步骤S40,当判断所述数字化音频的完整弹奏时间小于所述预设弹奏时间阈值时,将所述数字化音频存储为非时间序列预测模型的训练数据。
交互式设备将用户弹奏的音乐旋律成功存储为数字化音频后,下一步就是根据所存储的数字化音频进行预测,将预设弹奏时间阈值设为30秒,当演奏者不间断的弹奏时间超过预设弹奏时间阈值30秒时,启动时间序列模型预测30秒后的音乐伴奏,当演奏者完整弹奏时间不足预设弹奏时间阈值30秒时,将音频信号存储为数字化音频,以供非时间序列预测模型训练和预测。
本实施例音乐预测模型采用的是时间序列预测模型和非时间序列预测模型,具体模型预测方法分别如下:
在步骤S30中,时间序列预测模型俗称在线预测,当演奏者达到30秒的演奏时间时,模型会通过该30秒的演奏数据,递归的修改输出连接权值w,然后有规律的预测输出,从而达到辅助演奏者演奏的目的。
整个时间序列预测模型分为模型训练和模型预测。具体如下:
训练时间序列训练模型阶段:时间序列预测是先在一段时间内获得一个系统相关变量的真实值,然后用回声状态网络算法对这个系统的某个或某些变量未来的取值进行预测。本模型预测的变量是音乐的采样频率和采样数位。回声状态网络是一种简化的递归神经网络模型,可以有效避免递归神经网络学习算法收敛速度慢的缺点,具有计算复杂性高的特性,特别适合应用于交互式设备中,这是本实施例中采用其进行时间序列预测的主要原因。回声状态网络是由三个部分构成,如图2所示,图2为本申请一实施例提供的回声状态网络模型结构示意图。
结合音乐旋律在某t时刻来说,
中间部分的大圆圈001表示储备池x t,w t是t时刻储备池权值的估计值。
左边部分002表示真实数据的输入神经元,即音乐的采样频率和位数,统称为测量值
Figure PCTCN2018123593-appb-000005
右边部分003表示模型预测的输出神经元y t
储备池是由大量的神经元组成(数量通常为几百个),储备池内部的神经元采用稀疏连接(稀疏连接是指神经元之间只是部分连接,如上图所示),神经元之间的连接权值是随机生成的,并且连接权值生成后就保持固定不变,也就是储备池的连接权值不需要训练。外部数据通过输入神经元进入到储备池后预测,最终由输出神经元输出y t
对于回声状态网络的时间序列预测模型的训练,本实施例使用卡尔曼滤波法。卡尔曼滤波作为一种数值估计的优化方法,应用在任何含有不确定信息的动态系统中,对系统的下一步的走向都能做出有根据的预测,所以使用卡尔曼滤波训练回声状态网络可以高效的提升时间序列预测模型的准确率。结合卡尔曼滤波法的方程公式,在t+1时刻时有:
w t+1=w tt
Figure PCTCN2018123593-appb-000006
其中α t、β t分别为卡尔曼滤波在t时刻的过程噪声和测量噪声,其协方差矩阵分别为q t、r t。而对于t时刻的时间序列模型,由以下步骤可得:
p t=p t-1+q t-1
Figure PCTCN2018123593-appb-000007
Figure PCTCN2018123593-appb-000008
Figure PCTCN2018123593-appb-000009
其中p t是协方差矩阵,k t是卡尔曼滤波器的增益。同理可得t-1、t-2等时刻的状态量。由以上所述,可以更新储备池内的权重,达到训练时间序列预测模型的目的。
模型预测阶段:对弹奏时间计时,判断弹奏时间是否超过预设弹奏时间阈值;
进一步的,本实施例中用户在使用交互式设备开始弹奏时,设备同时启动两个步骤,一、对弹奏时间计时;二、将数字化音频存储。数字化音频存储的目的是为了存储足够训练数据供非时间序列预测模型训练使用。
设定的预设弹奏时间阈值为30秒。一旦弹奏时间超过阈值30秒后,基于已训练好的回声状态网络的时间序列预测模型开始工作,输出音乐伴奏,辅助演奏者弹奏;
当弹奏的完整时间不足30秒时,时间序列预测模型不工作,但弹奏数据会通过交互式设备转为数字化音频存储到内存中,作为训练数据供非时间序列预测模型训练。设定弹奏时间阈值的原因是为了保证具有足够的音频存储量,以提高预测准确率。
在步骤S40中,与时间序列预测模型对应的是非时间序列预测模型。当弹奏者弹奏出音乐旋律时,音频信号会被转化成数字化音频被存储到交互式设备中,基于每次存储的数字化音频,交互式设备都会对其进行训练和预测。这种基于离线训练和预测的方法称为非时间预测模型。本实施例采用深度卷积生成对抗网络法(Deep Convolutional Generative Adversarial Nerworks,DCGAN)对非时间序列进行预测。主要步骤包括:
S401:提取存储的数字化音频;
S402:训练深度卷积生成对抗网络;
S403:根据用户需求播放预测的音乐伴奏。
其中步骤S401主要是将之前在交互式设备存储的数字化音频提取出来。步骤S402根据所提取的数据进行生成式对抗网络的训练。使用该网络的原因是因为弹奏者的精力有限,继而在交互式设备中所存储的数字化音频的数据量并不多,针对这种样本数据量不够多的问题,使用深度卷积生成对抗网络 自动生成数据的同时也训练音乐旋律,达到两重的效果。在本实施例中,DCGAN网络模型包含一个生成网络G和一个判别网络D,DCGAN的目标函数是基于生成网络G与判别网络D的最小值和最大值的问题。如图3所示,图3为DCGAN网络模型训练流程示意图3,当生成对抗网络训练一个生成器时,首先利用生成网络G,从随机的数字化音频噪声Z(音频噪声是提前在DCGAN中存储的数字化随机音频数据,其并不是有规律的音乐旋律数据)中生成逼真的数字化音频样本,同时判别网络D训练一个鉴别器来鉴别真实数字化音频X(真实数字化音频是指在步骤一所存储的具有旋律的数字化音频)和生成的数字化音频样本之间的差距。整个过程让生成器和鉴别器同时训练,直到生成网络G与判别网络D的损失函数值都达到预设的某阈值时,证明此时模型训练成功,具有预测音乐旋律的能力。此时模型的生成网络生成的数字化音频数据与真实的样本具有很高的相似度,即使判别网络也无法区分生成网络生成的数字化音频数据和真实数据的差异,其中,生成网络G的损失函数为:
(1-y)lg(1-D(G(Z)))
判别网络D的损失函数为:
-((1-y)lg(1-D(G(Z)))+ylgD(x))
其中,x表示输入参数,即步骤(1)提取的数字化音频,y指DCGAN的生成网络G与判别网络D所预测的数字化音频值。特别要强调是,DCGAN的生成网络和判别网络都是卷积神经网络。基于以上所述,训练成功的非时间序列预测模型可自动生成音乐伴奏,供演奏者使用学习。
可选地,在其他实施例中,音乐自动生成程序还可以被分割为一个或者多个模块,一个或者多个模块被存储于存储器11中,并由一个或多个处理器(本实施例为处理器12)所执行以完成本申请,本申请所称的模块是指能够完成特定功能的一系列计算机程序指令段,用于描述音乐自动生成程序在音乐自动生成装置中的执行过程。
例如,参照图5所示,为本申请音乐自动生成装置一实施例中的音乐自动生成程序的程序模块示意图,该实施例中,音乐自动生成程序可以被分割为音频信号采集模块10、弹奏时间计时模块20,时间序列预测模型30,及非 时间序列预测模型40,示例性地:
音频信号采集模块10,用于采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储;
弹奏时间计时模块20,用于对所述数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系;
时间序列预测模型30,用于判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏;
非时间序列预测模型40,用于判断所述数字化音频的完整弹奏时间小于所述预设弹奏时间阈值时,将所述数字化音频存储为非时间序列预测模型的训练数据。
上述音频信号采集模块10、弹奏时间计时模块20、时间序列预测模型30及非时间序列预测模型40等程序模块被执行时所实现的功能或操作步骤与上述实施例大体相同,在此不再赘述。
此外,本申请实施例还提出一种计算机可读存储介质,所述计算机可读存储介质上存储有音乐自动生成程序,所述音乐自动生成程序可被一个或多个处理器执行,以实现如下操作:
采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储;
对所述数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系;
判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏;
判断所述数字化音频的完整弹奏时间小于所述预设弹奏时间阈值时,将所述数字化音频存储为非时间序列预测模型的训练数据。
本申请计算机可读存储介质具体实施方式与上述音乐自动生成装置和方法各实施例基本相同,在此不作累述。
需要说明的是,上述本申请实施例序号仅仅为了描述,不代表实施例的 优劣。并且本文中的术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、装置、物品或者方法不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、装置、物品或者方法所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、装置、物品或者方法中还存在另外的相同要素。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在如上所述的一个存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端设备(可以是手机,计算机,服务器,或者网络设备等)执行本申请各个实施例所述的方法。
以上仅为本申请的优选实施例,并非因此限制本申请的专利范围,凡是利用本申请说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本申请的专利保护范围内。

Claims (20)

  1. 一种音乐自动生成方法,其特征在于,所述方法包括:
    采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储;
    对所述数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系;
    当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏;
    当判断所述数字化音频的完整弹奏时间小于所述预设弹奏时间阈值时,将所述数字化音频存储为非时间序列预测模型的训练数据。
  2. 根据权利要求1所述的音乐自动生成方法,其特征在于,所述采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储的步骤,包括如下步骤:
    利用音频放大器采集所述音频信号的采样频率和采样数位;
    对所述音频信号进行数据清洗。
  3. 根据权利要求2所述的音乐自动生成方法,其特征在于,所述对所述音频信号进行数据清洗的步骤,包括如下步骤:
    将所述音频信号通过低通滤波器,对高于半采样频率的音频信号进行限带处理,以改善混叠干扰。
  4. 根据权利要求2所述的音乐自动生成方法,其特征在于,所述对所述音频信号进行数据清洗的步骤,包括如下步骤:
    在采集所述音频信号的同时,采集抖动发生器发出的噪声,并将所述噪声加入到所述音频信号中,以改善量化误差干扰。
  5. 根据权利要求1所述的音乐自动生成方法,其特征在于,所述当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏的步骤,还包括如下步骤:
    将所述数字化音频存储为非时间序列预测模型的训练数据。
  6. 根据权利要求2所述的音乐自动生成方法,其特征在于,所述当判断 所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏的步骤,还包括如下步骤:
    将所述数字化音频存储为非时间序列预测模型的训练数据。
  7. 根据权利要求3或4所述的音乐自动生成方法,其特征在于,所述当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏的步骤,还包括如下步骤:
    将所述数字化音频存储为非时间序列预测模型的训练数据。
  8. 一种音乐自动生成装置,其特征在于,所述装置包括存储器和处理器,所述存储器上存储有可在所述处理器上运行的音乐自动生成程序,所述音乐自动生成程序被所述处理器执行时实现如下步骤:
    采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储;
    对所述数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系;
    当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏;
    当判断所述数字化音频的完整弹奏时间小于所述预设弹奏时间阈值时,将所述数字化音频存储为非时间序列预测模型的训练数据。
  9. 根据权利要求8所述的音乐自动生成装置,其特征在于,所述采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储的步骤,包括如下步骤:
    利用音频放大器采集所述音频信号的采样频率和采样数位;
    对所述音频信号进行数据清洗。
  10. 根据权利要求9所述的音乐自动生成装置,其特征在于,所述对所述音频信号进行数据清洗的步骤,包括如下步骤:
    将所述音频信号通过低通滤波器,对高于半采样频率的音频信号进行限带处理,以改善混叠干扰。
  11. 根据权利要求9所述的音乐自动生成装置,其特征在于,所述对所 述音频信号进行数据清洗的,还包括如下步骤:
    在采集所述音频信号的同时,采集抖动发生器发出的噪声,并将所述噪声加入到所述音频信号中,以改善量化误差干扰。
  12. 根据权利要求8所述的音乐自动生成装置,其特征在于,所述当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏的步骤,还包括如下步骤:
    将所述数字化音频存储为非时间序列预测模型的训练数据。
  13. 根据权利要求9所述的音乐自动生成装置,其特征在于,所述当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏的步骤,还包括如下步骤:
    将所述数字化音频存储为非时间序列预测模型的训练数据。
  14. 根据权利要求10或11所述的音乐自动生成装置,其特征在于,所述当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏的步骤,还包括如下步骤:
    将所述数字化音频存储为非时间序列预测模型的训练数据。
  15. 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质上存储有音乐自动生成程序,所述程序可被一个或者多个处理器执行,以实现如下步骤:
    采集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储;
    对所述数字化音频的弹奏时间进行计时,判断弹奏时间与预设弹奏时间阈值的关系;
    当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏;
    当判断所述数字化音频的完整弹奏时间小于所述预设弹奏时间阈值时,将所述数字化音频存储为非时间序列预测模型的训练数据。
  16. 根据权利要求15所述的计算机可读存储介质,其特征在于,所述采 集音乐旋律的音频信号,将所述音频信号转化为数字化音频存储的步骤,包括如下步骤:
    利用音频放大器采集所述音频信号的采样频率和采样数位;
    对所述音频信号进行数据清洗。
  17. 根据权利要求16所述的计算机可读存储介质,其特征在于,所述对所述音频信号进行数据清洗的步骤,包括如下步骤:
    将所述音频信号通过低通滤波器,对高于半采样频率的音频信号进行限带处理,以改善混叠干扰。
  18. 根据权利要求16所述的计算机可读存储介质,其特征在于,所述对所述音频信号进行数据清洗的,还包括如下步骤:
    在采集所述音频信号的同时,采集抖动发生器发出的噪声,并将所述噪声加入到所述音频信号中,以改善量化误差干扰。
  19. 根据权利要求15或16所述的计算机可读存储介质,其特征在于,所述当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏的步骤,还包括如下步骤:
    将所述数字化音频存储为非时间序列预测模型的训练数据。
  20. 根据权利要求17或18所述的计算机可读存储介质,其特征在于,所述当判断所述数字化音频的弹奏时间大于所述预设弹奏时间阈值时,启动时间序列预测模型,根据对预设弹奏时间阈值以前的数字化音频训练得到预设弹奏时间阈值以后的音乐伴奏的步骤,还包括如下步骤:
    将所述数字化音频存储为非时间序列预测模型的训练数据。
PCT/CN2018/123593 2018-11-12 2018-12-25 一种音乐自动生成方法、装置及计算机可读存储介质 Ceased WO2020098086A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201811341758.6A CN109637509B (zh) 2018-11-12 2018-11-12 一种音乐自动生成方法、装置及计算机可读存储介质
CN201811341758.6 2018-11-12

Publications (1)

Publication Number Publication Date
WO2020098086A1 true WO2020098086A1 (zh) 2020-05-22

Family

ID=66067828

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2018/123593 Ceased WO2020098086A1 (zh) 2018-11-12 2018-12-25 一种音乐自动生成方法、装置及计算机可读存储介质

Country Status (2)

Country Link
CN (1) CN109637509B (zh)
WO (1) WO2020098086A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116670751A (zh) * 2020-11-25 2023-08-29 雅马哈株式会社 音响处理方法、音响处理系统、电子乐器及程序

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110222226B (zh) * 2019-04-17 2024-03-12 平安科技(深圳)有限公司 基于神经网络的以词生成节奏的方法、装置及存储介质
CN111753519B (zh) * 2020-06-29 2024-05-28 鼎富智能科技有限公司 一种模型训练和识别方法、装置、电子设备及存储介质
CN112669798B (zh) * 2020-12-15 2021-08-03 深圳芒果未来教育科技有限公司 一种对音乐信号主动跟随的伴奏方法及相关设备

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107293289A (zh) * 2017-06-13 2017-10-24 南京医科大学 一种基于深度卷积生成对抗网络的语音生成方法
CN107644630A (zh) * 2017-09-28 2018-01-30 清华大学 基于神经网络的旋律生成方法及装置
CN107871492A (zh) * 2016-12-26 2018-04-03 珠海市杰理科技股份有限公司 音乐合成方法和系统
US10068557B1 (en) * 2017-08-23 2018-09-04 Google Llc Generating music with deep neural networks

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6175072B1 (en) * 1998-08-05 2001-01-16 Yamaha Corporation Automatic music composing apparatus and method
JP3531507B2 (ja) * 1998-11-25 2004-05-31 ヤマハ株式会社 楽曲生成装置および楽曲生成プログラムを記録したコンピュータで読み取り可能な記録媒体
EP1265221A1 (en) * 2001-06-08 2002-12-11 Sony France S.A. Automatic music improvisation method and device
CN108281127A (zh) * 2017-12-29 2018-07-13 王楠珊 一种音乐练习辅助系统、方法、装置及存储装置

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN107871492A (zh) * 2016-12-26 2018-04-03 珠海市杰理科技股份有限公司 音乐合成方法和系统
CN107293289A (zh) * 2017-06-13 2017-10-24 南京医科大学 一种基于深度卷积生成对抗网络的语音生成方法
US10068557B1 (en) * 2017-08-23 2018-09-04 Google Llc Generating music with deep neural networks
CN107644630A (zh) * 2017-09-28 2018-01-30 清华大学 基于神经网络的旋律生成方法及装置

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
LIU, XIA ET AL.: "Generation of the Music by Computer Simulation", JOURNAL OF SOOCHOW UNIVERSITY (ENGINEERING SCIENCE EDITION), vol. 24, no. 2, 30 April 2004 (2004-04-30), pages 6 - 9, ISSN: 1000-1999 *
WANG, CHENG ET AL.: "Recurrent Neural Network Method for Automatic Generation of Music", JOURNAL OF CHINESE COMPUTER SYSTEMS, vol. 38, no. 10, 31 October 2017 (2017-10-31), pages 2412 - 2414, ISSN: 1000-1220 *

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116670751A (zh) * 2020-11-25 2023-08-29 雅马哈株式会社 音响处理方法、音响处理系统、电子乐器及程序

Also Published As

Publication number Publication date
CN109637509A (zh) 2019-04-16
CN109637509B (zh) 2023-10-03

Similar Documents

Publication Publication Date Title
CN112309365B (zh) 语音合成模型的训练方法、装置、存储介质以及电子设备
WO2020098086A1 (zh) 一种音乐自动生成方法、装置及计算机可读存储介质
CN111444967B (zh) 生成对抗网络的训练方法、生成方法、装置、设备及介质
CN113571080B (zh) 语音增强方法、装置、设备及存储介质
CN109326270B (zh) 音频文件的生成方法、终端设备及介质
CN103189913A (zh) 用于分解多信道音频信号的方法、设备和机器可读存储媒体
CN111680187A (zh) 乐谱跟随路径的确定方法、装置、电子设备及存储介质
WO2021213135A1 (zh) 音频处理方法、装置、电子设备和存储介质
CN114299918B (zh) 声学模型训练与语音合成方法、装置和系统及存储介质
TWI740315B (zh) 聲音分離方法、電子設備和電腦可讀儲存媒體
CN109410972B (zh) 生成音效参数的方法、装置及存储介质
CN112309409A (zh) 音频修正方法及相关装置
US9330649B2 (en) Selecting audio samples of varying velocity level
EP4618071A1 (en) Human voice note recognition model training method, human voice note recognition method, and device
CN110969141A (zh) 一种基于音频文件识别的曲谱生成方法、装置及终端设备
CN115762546B (zh) 音频数据处理方法、装置、设备以及介质
CN106484765A (zh) 利用音频数字冲击以创建数字媒体演示
CN109308903A (zh) 语音模仿方法、终端设备及计算机可读存储介质
CN112420002A (zh) 乐曲生成方法、装置、电子设备及计算机可读存储介质
US20230185846A1 (en) Music streaming, playlist creation and streaming architecture
WO2024082389A1 (zh) 音乐分轨匹配振动的触觉反馈方法、系统及相关设备
CN113436621B (zh) 一种基于gpu语音识别的方法、装置、电子设备及存储介质
CN112767957B (zh) 获得预测模型的方法、语音波形的预测方法及相关装置
CN111179691A (zh) 一种音符时长显示方法、装置、电子设备及存储介质
CN115240696B (zh) 一种语音识别方法及可读存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 18940188

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 18940188

Country of ref document: EP

Kind code of ref document: A1