WO2012011475A1 - 声色変化反映歌声合成システム及び声色変化反映歌声合成方法 - Google Patents
声色変化反映歌声合成システム及び声色変化反映歌声合成方法 Download PDFInfo
- Publication number
- WO2012011475A1 WO2012011475A1 PCT/JP2011/066383 JP2011066383W WO2012011475A1 WO 2012011475 A1 WO2012011475 A1 WO 2012011475A1 JP 2011066383 W JP2011066383 W JP 2011066383W WO 2012011475 A1 WO2012011475 A1 WO 2012011475A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- voice
- singing voice
- singing
- input
- spectral
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
- G10L13/033—Voice editing, e.g. manipulating the voice of the synthesiser
- G10L13/0335—Pitch control
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
- G10L13/033—Voice editing, e.g. manipulating the voice of the synthesiser
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/02—Means for controlling the tone frequencies, e.g. attack or decay; Means for producing special musical effects, e.g. vibratos or glissandos
- G10H1/06—Circuits for establishing the harmonic content of tones, or other arrangements for changing the tone colour
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/08—Text analysis or generation of parameters for speech synthesis out of text, e.g. grapheme to phoneme translation, prosody generation or stress or intonation determination
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2250/00—Aspects of algorithms or signal processing methods without intrinsic musical character, yet specifically adapted for or used in electrophonic musical processing
- G10H2250/315—Sound category-dependent sound synthesis processes [Gensound] for musical use; Sound category-specific synthesis-controlling parameters or control means therefor
- G10H2250/455—Gensound singing voices, i.e. generation of human voices for musical applications, vocal singing sounds or intelligible words at a desired pitch or with desired vocal effects, e.g. by phoneme synthesis
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/06—Elementary speech units used in speech synthesisers; Concatenation rules
Definitions
- the present invention relates to a voice color change reflecting singing voice synthesizing system and a voice color change reflecting singing voice synthesizing method capable of generating a synthesized singing voice imitating the pitch, volume and tone color change of an input singing voice.
- a singing voice synthesis system that can artificially generate human-like singing voices can easily synthesize with various singing voices and can control the expression of singing with high reproducibility, so it is important to expand the possibilities in the production of songs with singing It has become a tool. Since 2007, the number of users enjoying music production using commercially available singing voice synthesis software has increased rapidly, and the singing voice synthesis system has been taken up by various media due to high social interest in expanding its use.
- Non-Patent Document 1 a technique for adjusting numerical parameters in a user's manual operation (mouse)
- Non-Patent Document 2 a technique for morphing voice quality from singing voices of the same lyrics by two singers
- Non-patent Document 3 an emotion morphing technique applied to a plurality of songs of the same singer who sang with different emotions.
- speech synthesis there have been technologies related to voice quality conversion between different speakers (Non-Patent Documents 4 and 5) and research on emotional speech synthesis (Non-Patent Documents 6 and 7).
- Non-Patent Documents 8 to 15 studies that use voice quality conversion accompanying emotional changes.
- speech morphing research that generates an average voice from a plurality of voices (Non-Patent Document 14) and research that estimates a ratio from a plurality of voices and morphs them into voices close to user voices (Non-Patent Document 15). is there.
- Hideki Kenmochi, Hayato Ohshita "Singing Voice Synthesis System VOCALOID-Current Status and Issues", Information Processing Society of Japan 2008-MUS-74-9, Vol. 2008, No. 12, pp. 51-58 (2008) .
- Masamasa Morise “E.morish, an interface that mixes singing voices”, http://www.crestmuse.jp/cmstraight/personal/e.morish/. Toda, T., Black, A. and Tokuda, K .: ⁇ Voice ⁇ conversion based on maximum likelihood estimation of spectral parameter trajectory '', IEEE Trans. On Audio, Speechand Language Processing, Vol. 15, No. 8222, pp. -2235 (2007).
- Yamato Otani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano “The best voice conversion method based on mixed normal distribution model using STRAIGHT mixed excitation source”, IEICE Transactions, Vol. J91-D, No. 4, pp.
- Patent Document 1 and Non-Patent Documents 16 and 17 are techniques for estimating the singing voice synthesis parameters of existing singing voice synthesis software by imitating the pitch and volume from user singing (FIG. 1). Through repeated parameter estimation, the estimation accuracy is improved, and it is now possible to synthesize automatically without readjustment even when the singing voice synthesis system or its sound source (singer's voice) is switched. All you have to do is to give the text of the lyrics using your own singing voice model, and the work of assigning each note can be done almost automatically. You can view the results of this conventional singing voice synthesis at http://staff.aist.go.jp/t.nakano/VocaListener/index-j.html.
- the term “voice quality” refers not only to the acoustic characteristics and auditory differences that can identify an individual, but also to voice differences (whispering, whispering, etc.) caused by different utterance styles, bright voices, and dark voices. It is used in a variety of ways, such as the difference in impression (expression word) on hearing. Therefore, in the present specification, when expressing the change in voice quality during singing, the word “voice color change” is used in distinction from the word voice quality. If the voice color change during user singing can be imitated and reflected in the singing voice synthesis result in accordance with the lyrics and melody, it is thought that it will lead to the realization of more attractive singing voice synthesis.
- Non-Patent Document 1 a technique that allows a user to explicitly handle such a voice color change has been a singing voice synthesis system “Vocaloid” (trademark) shown in Non-Patent Document 1.
- a singing voice synthesis system “Vocaloid” (trademark) shown in Non-Patent Document 1.
- An object of the present invention is to provide a voice color change reflecting singing voice synthesizing system and a voice color change reflecting singing voice synthesizing method capable of reflecting not only changes in pitch and volume of a user song but also voice color changes in a singing voice synthetic song. .
- the voice color reflecting singing voice synthesizing system of the present invention includes a pitch and volume change reflecting singing voice synthesizing system, a synthesized singing voice acoustic signal storage unit, a spectrum envelope estimating unit, a voice color space estimating unit, a trajectory displacement deforming unit, and a first spectrum deforming unit.
- a curve estimation unit, a second spectrum modification curve estimation unit, a spectrum modification curved surface generation unit, and a synthesized acoustic signal generation unit are provided.
- the singing voice synthesizing system reflecting the change in sound and volume is composed of an acoustic signal storage unit of the input singing voice, a singing voice sound source database, and a singing voice in order to synthesize a plurality of various singing voices having the same lyrics as the input singing voice and imitating the pitch and volume.
- a synthesis parameter data estimation unit, a singing voice synthesis parameter data storage unit, a lyrics data storage unit, and a singing voice synthesis unit are provided.
- the sound and volume change reflecting singing voice synthesis system for example, the systems disclosed in Patent Document 1 and Non-Patent Documents 16 and 17 can be used.
- the input singing voice acoustic signal storage unit stores an acoustic signal of the user's input singing voice.
- the singing voice source database stores K singing voice source data of different singing voices (K is an integer of 1 or more) and J singing voice source data of the same singing voice and J types (J is an integer of 2 or more). .
- J singing voice sound source data of the same singing voice and J types (J is an integer of 2 or more) can be easily obtained by using an existing singing voice synthesizing system capable of changing the voice color.
- the singing voice synthesis parameter data estimation unit estimates singing voice synthesis parameter data in which the acoustic signal of the input singing voice is expressed by a plurality of types of parameters including at least a pitch parameter and a volume parameter.
- the singing voice synthesis parameter data storage unit stores singing voice synthesis parameter data.
- the lyrics data storage unit stores lyrics data corresponding to the acoustic signal of the input singing voice.
- the singing voice synthesizing unit outputs an acoustic signal of the synthesized singing voice based on one type of singing voice source data, singing voice synthesis parameter data, and lyrics data selected from the singing voice source database.
- the pitch parameter only needs to indicate a change in pitch.
- the volume parameter only needs to indicate a change in volume.
- the volume parameter is a MIDI standard expression or the dynamics (DYN) of a commercially available singing voice synthesis system.
- the synthesized singing voice signal storage unit is different from the voice signal of the same singing voice synchronized in time with the acoustic signals of K synthesized singing voices of different singing voices synchronized in time generated by the singing voice synthesizing system reflecting pitch and volume change.
- the sound signals of J synthesized singing voices are stored.
- the spectrum envelope estimation unit frequency-analyzes the acoustic signal of the input singing voice and K + J synthesized singing voice signals, and removes the influence of the pitch (F 0 ) on the frequency analysis results of these acoustic signals.
- Estimate the spectral envelope of S K + J + 1).
- the inventor has found that the difference in voice color can be defined as the difference in the shape of the spectrum envelope of the frequency analysis result of the acoustic signal.
- the difference in spectrum envelope shape includes a difference in phonemes and individuality. Therefore, it can be said that the time change of the shape of the spectrum envelope of the frequency analysis result of the acoustic signal in which such components are suppressed is a voice color change. Therefore, the present invention employs a voice space estimation unit and a trajectory displacement deformation unit in order to suppress components of phoneme differences and individuality differences.
- the voice space estimation unit suppresses components other than the component contributing to the voice color change from the time series of the S spectrum envelopes by processing based on the subspace method, and reflects the voice color of the input singing voice and the J types of voice colors (M dimensions).
- M is an estimated voice space.
- This voice color space is a virtual space in which components other than the voice color change are suppressed.
- S acoustic signals correspond to (position) one point on the voice color space at each time, and the time change of the S acoustic signals can be expressed as a trajectory that changes with time in the voice color space.
- the trajectory displacement deformation unit is based on the subspace method, except for components contributing to the tone color change, from the J spectrum envelopes for the acoustic signals of the J synthesized singing voices with different voice colors in the same singing voice.
- the positional relationship between the J kinds of voice colors at each time obtained by the processing is estimated with an M-dimensional vector, and the time trajectory of the positional relationship between the voice colors estimated with the M-dimensional vector is estimated as a voice color change tube.
- the voice color change tube is a polyhedron (polytope) in which J positions of the synthesized voices of J synthesized singing voices having different voice colors in the same singing voice are obtained on the voice color space and include the J positions. The time trajectory of the polyhedron is assumed.
- the trajectory displacement deformation unit obtains the position of the voice color of the input singing voice at each time obtained by suppressing the components other than the component contributing to the voice color change from the spectral envelope of the acoustic signal of the input singing voice by the process based on the subspace method.
- the time trajectory of the voice color position estimated by the vector and estimated by the M-dimensional vector is estimated as the voice color trajectory.
- the trajectory displacement deforming unit displaces (shifts) or deforms at least one of the voice color trajectory of the input singing voice and the voice color changing tube so that all or most of the voice color trajectory of the input singing voice exists in the voice color changing tube.
- the voice color space is an M-dimensional space in this way, it is assumed that J M-dimensional vectors exist in the M-dimensional space at each time t as the synthesis target voice color. It is assumed that the inner side surrounded by J points on the M-dimensional space is a deformable region of the same input singing voice to be synthesized. That is, the polyhedron (M-dimensional polytope) that changes from moment to moment is a region where the tone color can be changed. Therefore, the voice color trajectory of the input singing voice that exists in another place in the voice color space is shifted and scaled so as to enter the voice color change tube as much as possible (enlarge at least one of the voice color trajectory and the voice color change tube without changing the time axis).
- the target position of the synthesis in the voice space at each time is determined. And based on this synthetic
- the first spectral deformation curve estimation unit uses one singing voice source data in the J singing voice source data as the reference singing voice source data and corresponds to the reference singing voice source data.
- the spectrum envelope of the synthesized singing voice acoustic signal is defined as the reference spectral envelope, and the deformation ratio of the J synthesized voice signals of the J synthesized singing voice signals to the reference spectral envelope is determined at each time to obtain J kinds of voice colors.
- J spectrum deformation curves for synthesis corresponding to are estimated.
- the spectrum deformation curve for synthesis shows the change in the deformation ratio obtained at each time.
- the second spectrum deformation curve estimation unit overlaps a point in the voice trajectory of the input singing voice determined by the trajectory displacement deformation unit with a certain voice color in the voice color change tube at a certain time
- the spectral deformation curve at each time corresponding to the voice trajectory of the input singing voice is estimated so as to satisfy the constraint that the spectral envelope of the singing voice acoustic signal matches the spectral envelope of the synthesized singing voice.
- This spectrum deformation curve is for imitating the voice color of the input singing voice on the voice color space.
- the spectrum deformation curved surface generation unit generates a spectrum deformation curved surface by combining the spectrum deformation curves estimated by the second spectrum deformation curve estimation unit at each time.
- the synthesized acoustic signal generation unit generates a modified spectrum envelope by deforming the reference spectrum envelope based on the spectrum deformed curved surface at each time, and generates the modified spectrum envelope and the fundamental frequency (F 0 ) included in the reference singing voice source data. Based on this, a synthesized singing voice signal reflecting the change in voice color of the input singing voice is generated.
- singing voice synthesis imitating the voice color change of the input singing voice can be realized.
- the specific spectrum envelope estimation unit normalizes the volume of S acoustic signals including an input singing voice acoustic signal, J synthesized singing voice acoustic signals, and K synthesized singing voice acoustic signals.
- the spectrum envelope estimation unit performs frequency analysis on the normalized S acoustic signals, and estimates a plurality of pitches and non-periodic components from the frequency analysis result.
- the spectrum envelope estimation unit compares the estimated pitch with the voicedness threshold to determine whether the voice is voiced or unvoiced, and the voiced section has envelopes of a plurality of frequency spectra based on the fundamental frequency F 0 of the acoustic signal.
- the L 1 dimension is estimated (L 1 is a power of 2 + 1), and an unvoiced section estimates envelopes of a plurality of frequency spectra in the L 1 dimension based on a predetermined low frequency.
- the spectrum envelope estimation unit estimates S spectrum envelopes based on the envelopes of the plurality of frequency spectra in the section that is voiced and the envelopes of the plurality of frequency spectra in the section that is unvoiced. If the spectrum envelope estimation unit is configured in this way, it is possible to estimate a spectrum envelope from which the influence of F 0 is removed in a voiced interval. In the unvoiced section, a spectral envelope that appropriately represents the frequency transfer characteristic can be estimated. As a result, it is possible to obtain an advantage that a singing voice can be synthesized with high synthesis quality by using an aperiodic component at the time of synthesis.
- the specific voice space estimation unit obtains S discrete cosine transform coefficients by performing a discrete cosine transform on the S spectral envelopes, and uses a DC component in the discrete cosine transform coefficients for the S spectral envelopes.
- Discrete cosine transform coefficient vectors up to low-order L 2 dimensions excluding a certain 0th dimension (where L 2 ⁇ L 1 and L 2 is a positive integer) are acquired as analysis targets.
- the voice space estimation unit then performs S L 2 in each of T frames in which the S acoustic signals are voiced at the same time (T is the maximum number of seconds of the time length of the acoustic signal ⁇ sampling period).
- a principal component analysis is performed on the dimensional discrete cosine transform coefficient vector, and a principal component coefficient and a cumulative contribution rate are obtained from each principal component analysis.
- the number of seconds of the time length of the acoustic signal is obtained by measuring the length of the acoustic signal to be analyzed by time.
- the voice space estimation unit converts S discrete cosine transform coefficients into S L 2 dimensional principal component scores using the principal component coefficients in T frames, and S L 2 dimensional principal components.
- a principal component score of a higher dimension than the low-order N dimension (N is an integer of 1 or more and L 2 or less determined by R) with a cumulative contribution ratio R% (number of 0 ⁇ R ⁇ 100) is 0.
- the voice space estimation unit inversely transforms the S N-dimensional principal component scores into S new L 2- dimensional discrete cosine transform coefficients using the corresponding principal component coefficients, and T ⁇ S pieces. Principal component analysis is performed on the new vector of L 2 -dimensional discrete cosine transform coefficients to obtain principal component coefficients and cumulative contribution rates. Then, the L 2 dimensional discrete cosine transform coefficient is converted into a principal component score using the acquired principal component coefficient, and the space represented by the M (1 ⁇ M ⁇ L 2 ) dimension of the principal component score is defined as a voice space. Determine. If the voice space is defined in this way using the discrete cosine transform, the power can be concentrated in a low frequency range and can be handled with real numbers compared to the case where the Fourier transform is used.
- the trajectory displacement deformation unit is in the range of 0 to 1 in each dimension with respect to T ⁇ J M-dimensional principal component score vectors for the J synthesized singing voice acoustic signals constituting the voice change tube. Shift / scaling so that the value becomes.
- the trajectory displacement deformation unit shifts the T-dimensional M-dimensional principal component score vector for the input singing voice acoustic signal constituting the voice singing voice trajectory of the input singing voice so as to have a value in the range of 0 to 1 in each dimension.
- By performing the scaling all or most of the timbre trajectory of the input singing voice is present in the timbre change tube.
- By performing the shift / scaling so that the value is in the range of 0 to 1 in each dimension, it is possible to make all or most of the timbre trajectory of the input singing voice exist in the timbre change tube by calculation.
- the specific second spectral deformation curve estimation unit has a function of performing threshold processing by setting an upper limit and a lower limit on the spectral deformation curve at each time corresponding to the timbre trajectory of the input singing voice.
- threshold processing is performed by setting an upper limit and a lower limit on the spectrum deformation curve, unnatural deformation of the voice timbre of the input singing voice can be reduced when the timbre trajectory of the input singing voice is far away from the voice color changing tube.
- the specific spectral deformation curved surface generation unit has a function of performing two-dimensional smoothing on the spectral deformation curved surface.
- two-dimensional smoothing it is possible to suppress an abrupt change in the spectrum envelope, so that the unnaturalness of the synthesized singing voice can be reduced.
- the positional relationship between the plurality of types of voice colors at each time obtained by the processing is estimated with an M-dimensional vector, and the time trajectory of the positional relationship between the voice colors estimated with the M-dimensional vector is estimated as a voice color change tube.
- the position of the timbre of the input singing voice at each time is estimated with an M-dimensional vector obtained by suppressing the components other than the component contributing to the timbre change from the spectral envelope of the acoustic signal of the input singing voice by processing based on the subspace method.
- the time trajectory of the voice color position estimated by the M-dimensional vector is estimated as the voice color trajectory of the input singing voice, and the voice color trajectory of the input singing voice so that all or most of the voice color trajectory of the input singing voice exists in the voice color changing tube.
- At least one of the voice color change tubes is displaced or deformed (trajectory displacement deformation step).
- one singing voice sound source data among the J singing voice sound source data is set as reference singing voice sound source data
- the spectrum envelope of the synthesized singing voice sound signal corresponding to the reference singing voice sound source data is set as a reference spectral envelope
- J synthesis is performed.
- a deformation ratio of the J spectrum envelopes of the singing voice signal to the reference spectrum envelope is obtained at each time to estimate J synthesis spectrum deformation curves corresponding to the J voices (first spectrum deformation curves). Estimation step).
- the spectrum envelope of the acoustic signal of the input singing voice at a certain time is: A spectral deformation curve at each time corresponding to the voice trajectory of the input singing voice is estimated so as to satisfy the constraint that it matches the spectral envelope of the synthesized singing voice of the overlapping voice colors (second spectral deformation curve estimating step).
- a spectrum deformation curved surface is generated by combining the spectrum deformation curves estimated by the second spectrum deformation curve estimation unit (spectrum deformation curved surface generation step).
- the reference spectrum envelope is deformed based on the spectrum deformed curved surface to generate a modified spectrum envelope
- the voice color of the input singing voice is based on the modified spectrum envelope and the fundamental frequency (F 0 ) included in the reference singing voice source data.
- a synthesized singing voice signal reflecting the change is generated (synthetic acoustic signal generation step).
- the above steps are performed by a computer.
- (A) And (B) is a figure used in order to demonstrate that the difference in a voice color can be defined as the difference in the shape of a spectrum envelope. It is a block diagram which shows the structure of an example of a structure of the pitch and volume change reflection singing voice synthesis system used by embodiment of this invention. It is a block diagram which shows the main components of one Embodiment of the voice color change reflection singing voice synthesis system of this invention. It is a flowchart which shows the main algorithm in the case of implement
- FIG. 7 (C) to (E) are diagrams used to explain the operation process of the embodiment. It is each enlarged view of the waveform of the acoustic signal i of FIG.7 (C) thru
- FIG. (G) to (J) are diagrams used to explain the operation process of the embodiment.
- (A) to (E) are enlarged views of the waveforms of the frames shown in FIGS. 7, 10 and 12.
- FIG. It is a flowchart which shows an example of the algorithm in the case of implement
- the voice quality parameters in the existing singing voice synthesis system are automatically estimated according to the user singing as in the techniques described in Patent Document 1 and Non-Patent Documents 16 and 17.
- a way to do this is conceivable.
- this method has low practicality and versatility even though it is feasible. This is because, unlike the pitch and volume, the parameters relating to the voice quality and voice color change are different depending on the singing voice synthesis system, so that it is sufficiently conceivable that the acoustic characteristics that change depending on the parameters differ from system to system.
- some parameters that can be operated are different.
- the difference in voice color corresponds to the difference between the synthetic singing obtained by the above-mentioned applied product “Hatsune Miku” and the synthetic singing obtained by the applied product “Hatsune Miku Append”, which is defined as the difference in the shape of the spectrum envelope. it can.
- the difference in spectrum envelope shape includes a difference in phoneme and a difference in personality, as shown in FIGS. Therefore, it can be said that a time change in which such components are suppressed is a voice color change. If a spectrum envelope time series reflecting such a voice color change can be newly generated, singing voice synthesis imitating the voice color change of the user song can be realized.
- FIG. 2 is a block diagram showing an example of the configuration of the pitch and volume change reflecting singing voice synthesis system 100 used in the present embodiment.
- FIG. 3 is a block diagram showing the main components of one embodiment of the voice color change reflecting singing voice synthesizing system of the present invention.
- FIG. 4 is a flowchart showing a main algorithm of a program when the voice color change reflecting singing voice synthesizing system and the voice color change reflecting singing voice synthesizing method of the present invention are realized using a computer.
- the singing voice synthesis parameter data is repeatedly updated while comparing the synthesized singing (synthesized singing voice acoustic signal) with the input singing (input singing voice acoustic signal).
- the acoustic signal of the singing voice given by the user is referred to as the acoustic signal of the input singing voice
- the acoustic signal of the synthesized singing synthesized by the singing voice synthesizing unit is referred to as the synthesized singing voice acoustic signal.
- the user gives an input singing voice sound signal and its lyrics data to the system as input (step ST1 in FIG. 4).
- K singing voice source data of different singing voices K is an integer of 1 or more
- J singing voice source data of the same singing voice and J types J is an integer of 2 or more
- the acoustic signal of the input singing voice is stored in the acoustic signal storage unit 1 of the input singing voice.
- the acoustic signal of the input singing voice is an acoustic signal of a user's singing voice input from a microphone or the like, an acoustic signal of a ready-made singing voice, or an acoustic signal output by any other singing voice synthesis system.
- the lyrics are in Japanese
- the lyric data is usually character string data of a kanji-kana mixed sentence.
- the lyrics are in English, the data is alphabet string data.
- the lyric data is input to the lyric alignment unit 3 described later.
- the input singing voice acoustic signal analyzing unit 5 analyzes the acoustic signal of the input singing voice.
- the lyric alignment unit 3 converts the input lyric data into lyric data in which syllable boundaries are designated so as to synchronize with the sound signal of the input singing voice, and stores the conversion result in the lyric data storage unit 15.
- the lyrics alignment unit 3 allows the user to manually correct errors when converting kanji-kana mixed sentences into kana character strings.
- the lyrics alignment unit 3 allows a user to manually correct when there is a large error that spans phrases in the assignment of lyrics.
- the singing voice source database 103 stores K singing voice source data of different singing voices (K is an integer of 1 or more) and J singing voice source data of the same singing voice and J types (J is an integer of 2 or more). To do.
- K singing voice sound source data of different singing voices for example, male singing voice, female singing voice, child singing voice, etc.
- K singing voice sound source data of different singing voices are the existing singing voice synthesis system. 1 or the like can be used.
- J singing voice source data of the same singing voice and J types can be used to change the existing voice color like the “Singing Voice Synthesis System VOCALOID” shown in Non-Patent Document 1.
- the singing voice synthesis system 2 can be used.
- singing voice source data of six kinds of voices of DARK, LIGHT, SOFT, SOLID, SWEET and VIVID can be created as J kinds of timbres.
- the singing voice synthesizing unit 101 stores singing voice synthesizing parameter data storage unit 105 that stores singing voice synthesizing parameter data in which the acoustic signal of the input singing voice and the synthesized singing voice are expressed by a plurality of types of parameters including at least a pitch parameter and a volume parameter. Is the input. Then, the singing voice synthesizing unit 101 synthesizes the synthesized singing voice acoustic signal based on the one type of singing voice source data selected from the singing voice source database, the singing voice synthesis parameter data, and the lyrics data, and the synthesized singing voice acoustic signal storage unit. It outputs to 107.
- the synthesized singing voice signal storage unit 107 is the same singing voice whose time is synchronized with the acoustic signals of K synthesized singing voices of different singing voices synchronized in time, which are generated by the singing voice synthesizing system 100 reflecting pitch and volume change. Are stored as J synthesized singing voice signals.
- the operation so far is executed as step ST2 in FIG. 4, and the obtained K + J acoustic signals reflect changes in pitch and volume as shown in FIG. 5B.
- the system for estimating singing voice synthesis parameter data is roughly divided into an input singing voice acoustic signal analysis unit 5, an analysis data storage unit 7, a pitch parameter estimation unit 9, a volume parameter estimation unit 11, and a singing voice synthesis parameter data. And a creation unit 13.
- the input singing voice acoustic signal analysis unit 5 analyzes the pitch, volume, voiced section and vibrato section of the acoustic signal of the input singing voice as feature quantities, and stores the analysis result in the analysis data storage section 7. Note that when a tone deviation amount estimation unit 17, a pitch correction unit 19, a pitch transpose unit, a vibrato adjustment unit, and a smoothing processing unit described later are not provided, it is not necessary to analyze a vibrato section as a feature amount.
- the input singing voice acoustic signal analysis unit 5 may have any configuration as long as it can analyze (extract) the feature amount of the acoustic signal of the input singing voice.
- the input singing voice acoustic signal analysis unit 5 of the present embodiment has the following four functions.
- the first function is a function that estimates the fundamental frequency F 0 from the acoustic signal of the input singing voice at a predetermined cycle and stores it in the analysis data storage unit 7 as feature data of the pitch of the acoustic signal of the input singing voice. is there. Incidentally method of estimating the fundamental frequency F 0 is arbitrary.
- a method of estimating the fundamental frequency F 0 from an unaccompanied song may be used, or a method of estimating the fundamental frequency F 0 from a song with accompaniment may be used.
- the second function estimates the likelihood of voiced sound from the acoustic signal of the input singing voice, and observes and analyzes a section having a higher likelihood of voiced sound than the threshold as a voiced section of the acoustic signal of the input singing voice with reference to a predetermined threshold. This is a function of storing in the data storage unit.
- the third function is a function of observing the volume feature quantity of the acoustic signal of the input singing voice and storing it in the analysis data storage unit as volume feature quantity data.
- the fourth function is a function of observing a section where vibrato exists from pitch feature value data and storing it in the analysis data storage unit as a vibrato section. Any known detection method may be employed as the vibrato detection method.
- the pitch parameter estimation unit 9 is based on the feature value of the pitch of the acoustic signal of the input singing voice read from the analysis data storage unit 7 and the lyrics data in which the syllable boundary stored in the lyrics data storage unit 15 is designated. Assuming that the volume parameter is constant, the pitch parameter that can approximate the pitch feature amount of the synthesized singing voice signal to the pitch feature amount of the singing voice acoustic signal is estimated. Accordingly, the pitch parameter estimation unit 9 synthesizes the temporary singing voice synthesis parameter data created by the singing voice synthesis parameter data creation unit 13 based on the estimated pitch parameter by the singing voice synthesis unit 101, and the sound of the temporarily synthesized singing voice. Get a signal.
- the temporary singing voice synthesis parameter data created by the singing voice synthesis parameter data creation unit 13 is stored in the singing voice synthesis parameter data storage unit 105. Accordingly, the singing voice synthesizing unit 101 outputs an acoustic signal of the synthesized singing voice synthesized by the singing voice synthesizing unit 101 based on the temporary singing voice synthesis parameter data and the lyrics data in accordance with a normal synthesis operation. Then, the pitch parameter estimation unit 9 repeats the estimation of the pitch parameter until the pitch feature quantity of the temporarily synthesized singing voice signal approaches the pitch feature quantity of the input singing voice signal. Note that the pitch parameter estimation method is described in detail in Patent Document 1 and thus omitted.
- the pitch parameter estimation unit 9 has a function of analyzing the pitch feature amount of the temporarily synthesized singing voice signal output from the singing voice synthesis unit 101. Yes.
- the pitch parameter estimation unit 9 repeats the estimation of the pitch parameter a predetermined number of times (specifically, four times). Note that the pitch parameter estimation is repeated until the pitch feature value of the temporarily synthesized singing voice signal converges to the pitch feature value of the input singing voice signal instead of the predetermined number of times.
- the high parameter estimation unit 9 may be configured.
- the pitch of the temporarily synthesized singing voice signal is increased. Automatically approaches the feature value of the pitch of the acoustic signal of the input singing voice, so that the quality and accuracy of the synthesis of the singing voice synthesizing unit 101 become high.
- the volume parameter estimation unit 11 converts the volume feature of the sound signal of the input singing voice relative to the volume feature of the synthesized singing voice signal, and calculates the input singing voice.
- the volume parameter that can approximate the volume feature amount of the synthesized sound signal of the singing voice is estimated to the feature value of the volume of the sound signal that is converted into a relative value.
- the singing voice synthesis parameter data creation unit sings the temporary singing voice synthesis parameter data created based on the pitch parameter estimated by the pitch parameter estimation unit 9 and the volume parameter newly estimated by the volume parameter estimation unit 11. It is stored in the synthesis parameter data storage unit 105.
- the singing voice synthesizing unit 101 synthesizes the temporary singing voice synthesis parameter data and outputs an acoustic signal of the temporarily synthesized singing voice.
- the volume parameter estimation unit 11 repeats the estimation of the volume parameter a predetermined number of times until the volume feature quantity of the temporarily synthesized singing voice signal approaches the volume characteristic quantity converted to the relative value of the input singing voice signal. Similar to the pitch parameter estimation unit 9, the volume parameter estimation unit 11 is also characterized by the volume characteristic of the temporarily synthesized singing voice acoustic signal output from the singing voice synthesis unit 101, as with the input singing voice acoustic signal analysis unit 5. Built-in analysis function.
- the volume parameter estimation unit 11 of the present embodiment repeats the estimation of the volume parameter for a predetermined number of times (specifically, 4 times).
- the volume parameter estimation unit 11 is configured to repeat the estimation of the volume parameter until the volume characteristic quantity of the temporarily synthesized singing voice acoustic signal converges to the volume characteristic quantity converted to the relative value of the input singing voice acoustic signal.
- the estimation accuracy of the volume parameter can be made higher when the estimation is repeated for the volume parameter.
- the singing voice synthesis parameter data creation unit 13 creates singing voice synthesis parameter data based on the pitch parameter that has been estimated and the volume parameter that has been estimated, and stores the singing voice synthesis parameter data in the singing voice synthesis parameter data storage unit 105. .
- the pitch parameter estimated by the pitch parameter estimation unit 9 only needs to be able to indicate a change in pitch.
- the pitch parameter is a parameter element indicating a reference pitch level of a signal of a plurality of partial sections of an acoustic signal of an input singing voice corresponding to each of a plurality of syllables of lyrics data, and A parameter element indicating a temporal relative change in pitch with respect to a reference pitch level, and a parameter element indicating a change width in the pitch direction of a signal in a partial section.
- the data is directly stored in the lyric data storage unit 15.
- lyrics data for which no syllable boundary is specified is input to the singing voice synthesis parameter data creation unit 13
- the lyrics alignment unit 3 is based on the lyrics data for which the syllable boundary is not specified and the acoustic signal of the input singing voice. Then, lyric data in which syllable boundaries are specified is created.
- the tone deviation amount estimation unit 17, the pitch correction unit 19, the pitch transpose unit 21, the vibrato adjustment unit 23, and the smoothing are performed.
- a processing unit 25 is provided.
- the expression of the singing input is expanded by editing the acoustic signal itself of the input singing voice using these. Specifically, the following two types of changing functions can be realized. These change functions may be used depending on the situation, and it is possible to select not to use them.
- Vibrato extent can be changed to your own expression by intuitive operation to make vibrato stronger or weaker.
- the tone deviation amount estimation unit 17 estimates the tone deviation amount from the feature value data of the pitch in the continuous voiced section of the acoustic signal of the input singing voice stored in the analysis data storage unit 7.
- the pitch correction unit 19 corrects the pitch feature value data so as to exclude the tone shift amount estimated by the tone shift amount estimation unit 17 from the pitch feature value data. By estimating the amount of tone deviation and excluding that amount, an acoustic signal of an input singing voice with a low degree of tone deviation can be obtained.
- the pitch transpose unit 21 is used when pitch transposition is performed by adding / subtracting an arbitrary value to / from pitch feature value data.
- the voice range can be easily changed or transposed with respect to the acoustic signal of the input singing voice.
- the vibrato adjusting unit 23 arbitrarily adjusts the vibrato depth in the vibrato section.
- the smoothing processing unit 25 arbitrarily smoothes the pitch feature value data and the volume feature value data outside the vibrato section.
- the smoothing process here is a process equivalent to “adjusting the vibrato depth arbitrarily” outside the vibrato section, and the fluctuations in pitch and volume are increased or decreased outside the vibrato section. It has an effect to do. Since these functions are described in detail in Patent Document 1, they are omitted.
- the spectrum envelope estimation unit 109 performs the input singing voice signal i shown in FIG. 5A and K singing voice signals k 1 to k K of different singing voices (K is an integer of 1 or more) and the same singing voice.
- the frequency analysis is performed on the J synthesized singing voice acoustic signals j 1 to j J of J types (J is an integer of 2 or more), and the pitches (F 0 ) of the frequency analysis results of these acoustic signals are respectively obtained.
- a signal based on an input singing voice signal i, K synthesized singing voice signals k 1 to k K , and J synthesized singing voice signals j 1 to j J. are given the same symbols i, k 1 to k K , and j 1 to j J for convenience.
- the difference in voice color can be defined as the difference in the shape of the spectrum envelope of the frequency analysis result of the acoustic signal.
- the difference in spectrum envelope shape includes a difference in phonemes and individuality. Therefore, it can be said that a time change in which such components are suppressed is a voice color change.
- a spectral envelope is targeted as an acoustic characteristic that well represents a change in voice color.
- the pitch F 0
- the spectral envelope for the frequency analysis results of the acoustic signal of the input singing voice and the acoustic signals of the K + J synthesized singing voices, it is described in the following document.
- the voice analysis and synthesis system STRAIGHT technology is used.
- the spectrum envelope estimation unit 109 executes each step ST of the flowchart of the algorithm for estimating the spectrum envelope using a computer shown in FIG.
- K + J acoustic signals k 1 to k K and j 1 to j J are synthesized using “VocaListener” described in Patent Document 1 and Non-Patent Documents 16 and 17.
- the spectrum envelope of the singer of all acoustic signals at a certain frame time is considered to have only variations corresponding to differences in personality (voice quality) and voice color. This is because “VocaListener” is used to imitate the pitch, volume, and phoneme.
- the spectrum envelope estimation unit 109 first includes an acoustic signal i of the input singing voice, K synthesized singing voice acoustic signals k 1 to k K and J synthesized singing voice acoustic signals j 1 to j J.
- the volume of each acoustic signal (i, k 1 to k K and j 1 to j J ) is normalized (step ST31).
- the spectrum envelope estimation unit 109 performs frequency analysis on the normalized S acoustic signals, and estimates a plurality of pitches (F 0 ) and aperiodic components for each frequency band from the frequency analysis result (step ST32).
- the estimation method of the pitch and the non-periodic component is not particularly limited.
- Fujimura ⁇ Aperiodicity extraction and control using mixed mode excitation and group delay manipulation for a high quality speech analysis, modification and synthesis system STRAIGHT '', MAVEBA 2001, Sept. Methods such as those described in 13-15, firentze Italy, 2001. can be used.
- the spectrum envelope estimation unit 109 compares the estimated pitch with a voicedness threshold value to determine whether it is voiced or unvoiced [step ST33 and FIG. 7 (C)]. This determination is made because it is necessary to separate analysis / synthesis processing in the estimation of the spectral envelope for the voiced and unvoiced intervals.
- the voiced section estimates the envelope of a plurality of frequency spectra in the L 1 dimension (L 1 is a power of 2 + 1 ) based on the fundamental frequency F 0 of each acoustic signal (frequency as a reference for analysis) is unvoiced estimates an envelope of a plurality of frequency spectra by L 1-dimensional based on a predetermined low frequency (frequency to be a reference of the analysis).
- the reference frequency is F 0 in a voiced interval, and is a frequency lower than F 0 sufficient to estimate a spectrum envelope in an unvoiced interval.
- the spectrum envelope estimation unit 109 estimates S spectrum envelopes based on the envelopes of the plurality of frequency spectra in the section that is voiced, the envelopes of the plurality of frequency spectra in the section that is unvoiced, and the aperiodic component [ Step ST34 in FIG. 6: FIG. 7D].
- the estimation of the spectral envelope and the estimation of the non-periodic component are not limited to this embodiment, and any method with high accuracy can be used to increase the accuracy of synthesis.
- L 1 dimensional (frequency resolution) adopts 2049-D, to calculate the steps ST32 ⁇ ST34 in time units (1 ms) your bets i.e. each frame of the processing of FIG.
- the voice space estimation unit 111 and the trajectory displacement deformation unit 113 are employed in order to suppress the components of phonological differences and individual differences.
- the voice color space of M dimensions (M is an integer of 1 or more) reflecting the voice colors of J and J types of voice colors is estimated.
- the voice color space is a virtual space in which components other than the voice color change are suppressed.
- S acoustic signals correspond to one point on the timbre space at each time, and a time change of one point on the timbre space can be expressed as a trajectory changing in time on the timbre space.
- a partial space is configured for each frame.
- different frames are formed in each frame, and all frames cannot be handled in a unified manner. Therefore, by storing only the low-order N dimensions in the partial space for each frame and returning to the original space, components other than those contributing to voice quality / voice color change are suppressed.
- all frames of all synthesized singing voices are connected in series and a principal component analysis is performed at once, and the low-order M-dimensional space is treated as a voice space.
- the voice space estimation unit 111 used in the present embodiment is configured to execute the steps in the flowchart of FIG. 9 showing an algorithm when the voice space estimation unit 111 is realized using a computer.
- the timbre space estimation unit 111 performs discrete cosine transform for each frame Fd on the S spectrum envelopes to obtain S spectrum envelopes [FIG. 7D].
- a discrete cosine transform coefficient (shown as DCT coefficient in FIG. 9) is obtained [FIG. 7E].
- 8A to 8G are enlarged views of the waveforms of the S acoustic signals (i, k 1 to k K and j 1 to j J ) in FIGS. 7C to 7E.
- Is shown. 13A and 13B are enlarged diagrams showing examples of waveforms in the frames Fd and Fe in FIGS. 7D and 7E for easy understanding. is there.
- the signs are changed for distinction, the frames Fd and Fe are frames at the same time.
- step ST42 discrete cosine transform coefficient vectors up to low-order L 2 dimensions (where L 2 ⁇ L 1 and L 2 is a positive integer) are acquired as analysis targets. Steps ST41 and ST42 are executed in each frame of all acoustic signals (step ST4A).
- the voice space estimation unit 111 then includes T frames (T is the maximum and the time length of the acoustic signal) in which the S acoustic signals (i, k 1 to k K , j 1 to j J ) are voiced at the same time.
- T is the maximum and the time length of the acoustic signal
- the principal component analysis is performed on the S L 2 -dimensional discrete cosine transform coefficient vectors to obtain the principal component coefficients and the cumulative contribution rate [step ST43].
- S L 2 -dimensional discrete cosine transform coefficients are converted into S L 2 -dimensional principal component scores for each frame using the principal component coefficients [step ST44: FIG. 10 (F)].
- S principal component scores having a high-dimensional principal component score of 0 are inversely transformed into S new L-dimensional discrete cosine transform coefficient vectors [step ST46: FIG. 10 (G). And FIG. 12 (G)]. Steps ST43 to ST46 (step 4B) are performed in all the T frames described above.
- FIG. 11A shows an enlarged view of the S waveforms in FIG. 10E
- FIG. 11B shows an enlarged view of the S waveforms in FIG. 11C shows an enlarged view of the S waveforms in FIG. 10G
- FIG. 11D shows an enlarged view of the S waveforms in FIG. 12H described later.
- 13C and 13D are enlarged diagrams showing examples of waveforms in the frames Ff and Fg in FIGS. 10F and 10G for easy understanding. It is. Although the signs are changed for distinction, the frames Fd, Fe, Ff, and Fg are frames at the same time.
- FIG. 13E illustrates an enlarged waveform example of the frame Fh in FIG. 12H for easy understanding. Although the signs are changed for distinction, the frames Fd, Fe, Ff, Fg, and Fh are frames at the same time.
- a space expressed by the upper M (1 ⁇ M ⁇ L 2 ) dimensions of the principal component score is defined as a voice space [step ST49: FIG. 12 (I)].
- the voice space is defined in this way using the discrete cosine transform, the spectral envelope can be reproduced while reducing the number of dimensions (L 1 dimension ⁇ L 2 dimension). Note that Fourier transform can be used instead of discrete cosine transform.
- the voice change tube VT obtains J positions of the synthesized voices of J synthesized singing voices having different voice colors in the same singing voice in the voice color space.
- a polyhedron P (polytope) that includes a single position is considered, and a time locus of the polyhedron P is assumed.
- FIG. 12 (I) schematically shows the voice color changing tube VT and the polyhedron P, which are actually three-dimensional.
- transformation part 113 obtained the position of the voice color of the input singing voice in each time obtained by suppressing other than the component which contributes to a voice color change from the spectrum envelope of the acoustic signal i of the input singing voice by the process based on the subspace method.
- the time trajectory of the voice color position estimated with the M-dimensional vector and with the M-dimensional vector is estimated as the voice color trajectory IT of the input singing voice (FIG. 12 (I)).
- the trajectory displacement deforming unit 113 displaces or deforms at least one of the voice color trajectory IT of the input singing voice and the voice color changing tube VT so that all or most of the voice color trajectory IT of the input singing voice exists in the voice color changing tube VT ( FIG. 12 (J)).
- the voice space is an M-dimensional space
- the voice color to be synthesized exists on the M-dimensional space as J M-dimensional vectors at each time t. It is assumed that the inside surrounded by these J points is a deformable region of the same input singing voice to be synthesized. That is, the polyhedron P (M-dimensional polytope) that changes from moment to moment is a region where the tone color can be changed.
- the timbre locus IT of the input singing voice that is also present in another place in the timbre space is shifted and scaled so as to enter the timbre change tube VT as much as possible (at least one of the timbre locus IT and the timbre change tube VT is changed with respect to the time axis).
- the target position of synthesis in the voice space at each time is determined by enlarging or reducing without changing the position and displacing the position. And based on this synthetic
- FIG. 14 shows the details of step ST5 in FIG. 4, and is a flowchart showing an example of a program algorithm when the trajectory displacement deforming unit 113 is realized by a computer.
- the T ⁇ J M-dimensional principal component score vectors for the J synthesized singing voice acoustic signals constituting the voice color change tube VT in step ST51 are in the range of 0 to 1 in each dimension. Shift / scaling to a value.
- step ST52 the T-dimensional M principal component score vector for the input singing voice acoustic signal IT constituting the singing voice trajectory IT is shifted and scaled so as to have a value in the range of 0 to 1 in each dimension.
- step ST52 may be executed before step ST51 is executed.
- FIG. 15 shows details of step ST6 of FIG. 4, and the first spectral deformation curve estimation unit 115, the second spectral deformation curve estimation unit 117, the spectral deformation curved surface generation unit 119, and the synthetic acoustic signal generation of FIG. 6 shows a flowchart of an algorithm of a program used when the unit 121 is realized by a computer.
- FIG. 16 is a diagram used for explaining a process of forming a spectral deformation curve.
- the first spectral deformation curve estimation unit 115 instead of using the spectral envelope as it is, the first spectral deformation curve estimation unit 115 first estimates J spectral deformation curves for synthesis.
- the first spectral deformation curve estimation unit 115 determines one standard voice among J voice colors to be synthesized in the voice space. Specifically, one singing voice sound source data in the J singing voice sound source data is determined as reference singing voice sound source data (step ST61). Then, all frames in which all acoustic signals are voiced, that is, T frames in which the S acoustic signals described above are voiced at the same time (T is the maximum number of seconds of the time length of the acoustic signal ⁇ sampling period). In each, steps ST62 to ST65 are performed.
- a spectrum envelope is associated with each of the J M-dimensional vectors corresponding to the J voice singing sound source data to be synthesized in the voice space.
- a spectrum envelope of the synthesized singing voice signal corresponding to the reference singing voice source data is set as a reference spectrum envelope RS.
- MIKU Append trademark
- an application product of Krypton Future Media Co., Ltd. DARK, LIGHT, SOFT, SOLID "Hatsune Miku” synthesized by the synthesis system of 6 types of singing voice sound source data of 6 voices, SWEET, and VIVID, and the applied product of Krypton Future Media Co., Ltd.
- “Hatsune Miku” (trademark) J singing voice source data is composed of the singing voice source data.
- the spectrum envelope of the acoustic signal corresponding to the singing voice sound source data of “Hatsune Miku” is used as the reference spectrum envelope RS.
- FIG. 16 also shows the spectrum envelopes of SOFT, SWEET, and VIVID.
- the first spectral deformation curve estimation unit 115 obtains the deformation ratios of the J spectral envelopes of the J synthesized singing voice signals with respect to the reference spectral envelope RS at each time, and calculates changes in these deformation ratios as J. Estimated as J synthesis spectrum deformation curves corresponding to the types of voices (step ST63).
- the synthesis spectral deformation curve indicates the change in the deformation ratio obtained at each time. Therefore, as shown in the lowermost area of FIG. 16, the synthesis spectrum deformation curve of the reference spectrum envelope RS corresponding to the singing voice sound source data of “Hatsune Miku” is a straight line.
- step ST64 a spectrum deformation curve in the M-dimensional vector of the input singing voice in the voice space is calculated from the J-dimensional M-color vectors to be synthesized in the voice space and the corresponding spectrum deformation curves for synthesis.
- the second spectrum deformation curve estimation unit 117 determines that one point in the voice color trajectory IT of the input singing voice determined by the trajectory displacement deformation unit 113 and a certain voice color in the voice color change tube VT.
- a spectral deformation curve IS [FIG. 17] at each time is estimated.
- This spectral deformation curve IS is for imitating the voice color of the input singing voice on the voice color space.
- step ST65 threshold processing is performed by setting an upper limit and a lower limit on the spectrum deformation curve IS of the input singing voice at each time (FIG. 17).
- the spectral deformation curve IS exceeding the upper limit and the lower limit is cut.
- the upper and lower limits are determined based on the maximum value and the minimum value of the synthesis spectrum deformation curve for the J voice colors to be synthesized.
- FIG. 17 shows a process until a synthesized acoustic signal is generated using the spectral deformation curve IS.
- the spectrum deformation curved surface generation unit 119 estimates the spectrum deformation curved surface by combining the spectrum deformation curves IS at all times (all frames) (step ST66).
- the synthesized acoustic signal generation unit 121 performs two-dimensional smoothing on the spectrally deformed curved surface (step ST67), and uses the spectrally deformed curved surface that has been two-dimensionally smoothed as a reference for the spectral envelope of the voiced acoustic signal ( In FIG. 17, the spectrum envelope of Hatsune Miku) is deformed (step ST68).
- singing voice synthesis is performed using the deformed spectrum envelope and the fundamental frequency F 0 of the acoustic signal as a reference to generate a synthetic acoustic signal for synthetic singing that imitates the timbre change of the input singing voice (step ST69).
- the synthesized sound signal may be reproduced by the signal reproduction unit 123 or stored in an appropriate storage medium.
- the spectral envelope is not used as it is, but the deformation ratio is obtained based on a standard voice (for example, a Hatsune Miku that is not a Hatsune Miku append).
- the deformation ratio is estimated for each frame. This ratio is the aforementioned spectral deformation curve.
- the input singing voice overlaps each voice color point in the voice color space, it is estimated so as to satisfy the constraint that the spectrum deformation curve of the input singing voice at that time is the same as the spectrum deformation curve of the overlapping voice color.
- a paper entitled “Modelling with implicit surfaces that interpolate” written by Turk, G. and O'Brien, J. F. [ACM Transactions on Graphics, Vol. 21, No. 4, pp. 855-873 (2002 )] Is applied to the application of Variational Interpolation using Radial Basis Function.
- the input singing voice on the voice space is u (t) and each voice color is zj (t)
- the spectrum to imitate the voice color of the input singing voice on the voice space by solving the following constrained equation Get the deformation curve.
- Z ri (f, t) takes a logarithm as in equation (1), and allows the ratio to be linearly converted on the logarithmic axis and the estimation result to take a negative value.
- w k (f, t) is the mixing ratio
- P (•) is an M-variable first-order polynomial (coefficient) with z j (t) or u (t) as variables as the vector x as shown in Equation (5).
- Pm 0, ..., M).
- ⁇ (•)
- 2 log (•), ⁇ (•)
- 3, etc. may be used. Equation (4) corresponds to the above-mentioned constraint, and can be written by the following matrix when the voice space is M 3 dimensions.
- ⁇ jk represents ⁇ (z j (t) ⁇ z k (t)), and (f, t) and (t) are omitted.
- Equation (2) a spectrally deformed curved surface is generated by Equation (2).
- an upper limit and a lower limit are set for each frame to reduce the influence when a user song exists outside the voice color change tube.
- the smoothing process on the time-frequency plane reduces the change that is too steep and maintains the continuity of the spectrum.
- the spectral envelope of the singing voice signal used as a reference is transformed using this spectrally deformed curved surface, and synthesized with "STRAIGHT" to synthesize the synthesized voice signal for a synthetic song that mimics the timbre change of the input singing voice.
- a function for changing the scale of voice tone change by changing the scale of voice tone change It is possible to synthesize a singing voice with a large scale, or conversely, by reducing the voice tone change by reducing the scale.
- singing voice synthesis was performed in which the singer reflected the timbre change from the same plurality of sound sources, such as Hatsune Miku and Hatsune Miku Append.
- the voice color change tube by configuring the voice color change tube with different singers, there is a possibility that the voice quality can be dynamically changed to synthesize a singing voice.
- parameter estimation of the existing singing voice synthesizing system is not performed.
- the voice color change tube is composed of a plurality of voices with different GEN parameters, for example, there is a possibility that it can be applied to parameter estimation.
- the present invention it is possible to estimate a voice color change from an input singing voice that has not been realized so far and imitate it to synthesize a singing voice. That is, according to the present invention, there is an advantage that the user can easily synthesize a richly expressive singing voice, and can express a song from various viewpoints of pitch, volume and voice color.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Electrophonic Musical Instruments (AREA)
- Reverberation, Karaoke And Other Acoustics (AREA)
Abstract
Description
・ 調子はずれ(off Pitch) の補正:音高がずれた音を修正する。
・ ビブラート深さ(vibrato extent) の調整:ビブラートを強く・弱くという直感的操作で、自分好みの表現へ変更できる。
研究2: 井上 徹,西田昌史,藤本雅清,有木康雄:「部分空間と混合分布モデルを用いた声質変換」,電子情報通信学会技術研究報告SP,Vol. 101, No. 86, pp. 1-6 (2001).
上記2つの研究では、話者毎に部分空間を構成することで、音韻性(低次部分空間: 変動が大きな成分)と話者性(高次部分空間: 変動が小さな成分)を分離している。本実施の形態では、フレーム毎に部分空間を構成する。しかしそのままでは、各フレームで異なる空間が構成されることになり、全フレームを統一的に扱えない。そこで、フレーム毎の部分空間における低次N次元のみを保存して、元の空間に戻すことで、声質・声色変化に寄与する成分以外を抑制する。続いて、すべての合成された歌声の全フレームを直列につないで一度に主成分分析を行い、その低次M次元の空間を声色空間として扱う。このような処理によって、異なる歌唱者の全てのフレームが同じ空間上で扱えるだけでなく、歌詞の文脈などの音韻変化に伴う声色変化に関係する成分を、低次元で効率的に表現できる。なお表現力の高い空間を得るために、声色空間を構成する際に用いる歌唱者は多い方が望ましい。すなわちK個の音響信号は多いほうが好ましい。さらに、このような処理による余計な成分の抑制は、入力歌声との対応付けにおいても重要と考えられる。
スケールを大きくして抑揚ある歌声を合成したり、逆にスケールを小さく声色変化を抑えたりして合成できる。
声色変化の中心を変えることで、それぞれの声色を中心とした声色変化に変換できる。
3 歌詞アラインメント部
5 入力歌声音響信号分析部
7 分析データ記憶部
9 音高パラメータ推定部
11 音量パラメータ推定部
13 歌声合成パラメータデータ作成部
15 歌詞データ記憶部
17 調子はずれ量推定部
19 音高補正部
21 音高トランスポーズ部
23 ビブラート調整部
25 スムージング処理部
101 歌声合成部
103 歌声音源データベース
105 歌声合成パラメータデータ記憶部
107 合成歌声音響信号記憶部
109 スペクトル包絡推定部
111 声色空間推定部
113 軌跡変位変形部
115 第1のスペクトル変形曲線推定部
117 第2のスペクトル変形曲線推定部
119 スペクトル変形曲面生成部
121 合成音響信号生成部
123 信号再生部
Claims (10)
- 入力歌声の音響信号を記憶する入力歌声の音響信号記憶部と、異なる歌声のK個(Kは1以上の整数)の歌声音源データと、同一歌声で且つJ種類(Jは2以上の整数)の声色のJ個の歌声音源データが蓄積された歌声音源データベースと、入力歌声の音響信号を少なくとも音高パラメータ及び音量パラメータを含む複数種類のパラメータで表現した歌声合成パラメータデータを推定する歌声合成パラメータデータ推定部と、前記歌声合成パラメータデータを記憶する歌声合成パラメータデータ記憶部と、入力歌声の音響信号に対応した歌詞データを記憶する歌詞データ記憶部と、前記歌声音源データベースから選択した1種類の前記歌声音源データと前記歌声合成パラメータデータと前記歌詞データとに基づいて、合成された歌声の音響信号を出力する歌声合成部とを備えた音高及び音量変化反映歌声合成システムと、
前記音高及び音量変化反映歌声合成システムで生成された、時刻が同期した異なる歌声のK個の合成された歌声の音響信号と時刻が同期した同一歌声で声色が異なるJ個の合成された歌声の音響信号とを記憶する合成歌声音響信号記憶部と、
前記入力歌声の音響信号及び前記K+J個の合成された歌声の音響信号を周波数分析し、これら前記音響信号の周波数分析結果についてそれぞれ音高(F0)の影響を除去したS個(S=K+J+1)のスペクトル包絡を推定するスペクトル包絡推定部と、
前記S個のスペクトル包絡の時間系列から声色変化に寄与する成分以外を部分空間法に基づいた処理により抑制して、前記入力歌声の声色及び前記J種類の声色を反映したM次元(Mは1以上の整数)の声色空間を推定する声色空間推定部と、
前記声色空間内に、前記同一歌声で声色が異なるJ個の合成された歌声の音響信号についての前記J個のスペクトル包絡から、声色変化に寄与する成分以外を部分空間法に基づいた処理により抑制して得た、各時刻における前記J種類の声色の位置関係をM次元のベクトルで推定し且つ前記M次元のベクトルで推定した前記声色の位置関係の時間軌跡を声色変化チューブとして推定し、前記入力歌声の音響信号の前記スペクトル包絡から声色変化に寄与する成分以外を部分空間法に基づいた処理により抑制して得た、各時刻における前記入力歌声の声色の位置をM次元ベクトルで推定し且つ前記M次元ベクトルで推定した前記声色の位置の時間軌跡を入力歌声の声色軌跡として推定し、前記入力歌声の声色軌跡の全部または大部分が前記声色変化チューブ内に存在するように前記入力歌声の声色軌跡及び前記声色変化チューブの少なくとも一方を変位または変形する軌跡変位変形部と、
前記J個の歌声音源データ中の一つの歌声音源データを基準歌声音源データとして、該基準歌声音源データに対応する前記合成された歌声の音響信号の前記スペクトル包絡を基準スペクトル包絡とし、前記J個の合成された歌声の音響信号のJ個のスペクトル包絡の前記基準スペクトル包絡に対する変形比率を各時刻で求めて前記J種類の声色に対応したJ個の合成用スペクトル変形曲線を推定する第1のスペクトル変形曲線推定部と、
前記軌跡変位変形部で定めた前記入力歌声の声色軌跡中の1点と前記声色変化チューブ内のある声色とが、ある時刻で重なったときに、前記ある時刻における前記入力歌声の音響信号のスペクトル包絡が、重なった前記声色の前記合成された歌声のスペクトル包絡と一致するという制約を満たすように、前記入力歌声の声色軌跡に対応する各時刻のスペクトル変形曲線を推定する第2のスペクトル変形曲線推定部と、
各時刻において、前記第2のスペクトル変形曲線推定部が推定した前記スペクトル変形曲線を合わせてスペクトル変形曲面を生成するスペクトル変形曲面生成部と、
各時刻において、前記スペクトル変形曲面に基づいて前記基準スペクトル包絡を変形して変形スペクトル包絡を生成し、該変形スペクトル包絡と前記基準歌声音源データに含まれる基本周波数(F0)に基づいて前記入力歌声の声色の変化を反映した合成された歌声の音響信号を生成する合成音響信号生成部とを備えてなる声色変化反映歌声合成システム。 - 前記スペクトル包絡推定部は、前記入力歌声の音響信号、前記K個の合成された歌声の音響信号及び前記J個の合成された歌声の音響信号からなるS個の音響信号の音量を正規化し、
正規化した前記S個の音響信号を周波数分析して、周波数分析結果から複数の音高及び非周期成分を推定し、
推定した音高を有声らしさの閾値と比較して有声または無声の判定を行い、有声である区間は前記音響信号の基本周波数F0に基づいて前記複数の周波数スペクトルの包絡をL1次元(L1は2の累乗+1の整数)で推定し、無声である区間はあらかじめ定めた低い周波数に基づいて前記複数の周波数スペクトルの包絡をL1次元で推定し、
前記有声である区間の前記複数の周波数スペクトルの包絡と、前記無声である区間の前記複数の周波数スペクトルの包絡とに基づいて前記S個のスペクトル包絡を推定するように構成されている請求項1に記載の声色変化反映歌声合成システム。 - 前記声色空間推定部は、前記S個のスペクトル包絡に対して離散コサイン変換を行ってS個の離散コサイン変換係数を求め、前記S個のスペクトル包絡に対して離散コサイン変換係数における直流成分である第0次元を除いた低次のL2次元(但しL2<L1でL2は正の整数)までの離散コサイン変換係数ベクトルを分析対象として取得し、
前記S個の音響信号が同時刻で有声となるT個のフレーム(Tは最大で、音響信号の時間長の秒数×サンプリング周期)のそれぞれにおいて、前記S個のL2次元の離散コサイン変換係数ベクトルについて主成分分析を行って、それぞれ主成分係数と累積寄与率を取得し、
前記T個のフレームにおいて、前記主成分係数を用いて前記S個の離散コサイン変換係数をS個のL2次元の主成分スコアに変換し、
前記S個のL2次元の主成分スコアに対して、累積寄与率R%(0<R<100の数)となる低次のN次元(NはRによって定まる1以上L2以下の整数)よりも高次元の主成分スコアを0としてS個のN次元の主成分スコアを取得し、
前記S個のN次元の主成分スコアをそれぞれ、対応する前記主成分係数を用いて、S個の新たなL2次元の離散コサイン変換係数に逆変換し、
T×S個の新たなL2次元の離散コサイン変換係数のベクトルに対して主成分分析を行って主成分係数及び累積寄与率を取得し、取得した主成分係数を用いてL2次元の離散コサイン変換係数を主成分スコアに変換し、主成分スコアの上位M(1≦M≦L2)次元までで表現される空間を前記声色空間と定める請求項1に記載の声色変化反映歌声合成システム。 - 前記軌跡変位変形部は、前記声色変化チューブを構成する前記J個の合成された歌声の音響信号についてのT×J個のM次元主成分スコアベクトルに対して、各次元で0~1の範囲の値になるようにシフト・スケーリングを行い、且つ前記入力歌声の声色軌跡を構成する前記入力歌声の音響信号についてのT個のM次元主成分スコアベクトルに対して各次元で0~1の範囲の値になるようにシフト・スケーリングを行うことにより、前記入力歌声の声色軌跡の全部または大部分が前記声色変化チューブ内に存在させることを特徴とする請求項1,2または3に記載の声色変化反映歌声合成システム。
- 前記第2のスペクトル変形曲線推定部は、前記入力歌声の声色軌跡に対応する各時刻の前記スペクトル変形曲線に上限・下限を定めて閾値処理を行う機能を有している請求項1に記載の声色変化反映歌声合成システム。
- 前記スペクトル変形曲面生成部は、前記スペクトル変形曲面に対して二次元平滑化を行うことを特徴とする請求項1に記載の声色変化反映歌声合成システム。
- 入力歌声の音響信号を記憶する入力歌声の音響信号記憶部と、異なる歌声のK個(Kは1以上の整数)の歌声音源データと、同一歌声で且つJ種類(Jは2以上の整数)の声色のJ個の歌声音源データが蓄積された歌声音源データベースと、入力歌声の音響信号を少なくとも音高パラメータ及び音量パラメータを含む複数種類のパラメータで表現した歌声合成パラメータデータを推定する歌声合成パラメータデータ推定部と、前記歌声合成パラメータデータを記憶する歌声合成パラメータデータ記憶部と、入力歌声の音響信号に対応した歌詞データを記憶する歌詞データ記憶部と、前記歌声音源データベースから選択した1種類の前記歌声音源データと前記歌声合成パラメータデータと前記歌詞データとに基づいて、合成された歌声の音響信号を出力する歌声合成部とを備えた音高及び音量変化反映歌声合成システムを用いて、時刻が同期した異なる歌声のK個の合成された歌声の音響信号と時刻が同期した同一歌声で声色が異なるJ個の合成された歌声の音響信号とを生成する合成歌声音響信号生成ステップと、
前記入力歌声の音響信号及び前記K+J個の合成された歌声の音響信号を周波数分析し、これら前記音響信号の周波数分析結果についてそれぞれ音高(F0)の影響を除去したS個(S=K+J+1)のスペクトル包絡を推定するスペクトル包絡推定ステップと、
前記S個のスペクトル包絡の時間系列から声色変化に寄与する成分以外を部分空間法に基づいた処理により抑制して、前記入力歌声の声色及び前記J種類の声色を反映したM次元(Mは1以上の整数)の声色空間を推定する声色空間推定ステップと、
前記声色空間内に、前記同一歌声で声色が異なるJ個の合成された歌声の音響信号についての前記J個のスペクトル包絡から、声色変化に寄与する成分以外を部分空間法に基づいた処理により抑制して得た、各時刻における前記J種類の声色の位置関係をM次元のベクトルで推定し且つ前記M次元のベクトルで推定した前記声色の位置関係の時間軌跡を声色変化チューブとして推定し、前記入力歌声の音響信号の前記スペクトル包絡から声色変化に寄与する成分以外を部分空間法に基づいた処理により抑制して得た、各時刻における前記入力歌声の声色の位置をM次元ベクトルで推定し且つ前記M次元ベクトルで推定した前記声色の位置の時間軌跡を入力歌声の声色軌跡として推定し、前記入力歌声の声色軌跡の全部または大部分が前記声色変化チューブ内に存在するように前記入力歌声の声色軌跡及び前記声色変化チューブの少なくとも一方を変位または変形する軌跡変位変形ステップと、
前記K個の歌声音源データ中の一つの歌声音源データを基準歌声音源データとして、該基準歌声音源データに対応する前記合成された歌声の音響信号の前記スペクトル包絡を基準スペクトル包絡とし、前記K個の合成された歌声の音響信号のJ個のスペクトル包絡の前記基準スペクトル包絡に対する変形比率を各時刻で求めて前記J種類の声色に対応したJ個の合成用スペクトル変形曲線を推定する第1のスペクトル変形曲線推定ステップと、
前記軌跡変位変形ステップで定めた前記入力歌声の声色軌跡中の1点と前記声色変化チューブ内のある声色とが、ある時刻で重なったときに、前記ある時刻における前記入力歌声の音響信号のスペクトル包絡が、重なった前記声色の前記合成された歌声のスペクトル包絡と一致するという制約を満たすように、前記入力歌声の声色軌跡に対応する各時刻のスペクトル変形曲線を推定する第2のスペクトル変形曲線推定ステップと、
各時刻において、前記第2のスペクトル変形曲線推定部が推定した前記スペクトル変形曲線を合わせてスペクトル変形曲面を生成するスペクトル変形曲面生成ステップと、
各時刻において、前記スペクトル変形曲面に基づいて前記基準スペクトル包絡を変形して変形スペクトル包絡を生成し、該変形スペクトル包絡と前記基準歌声音源データに含まれる基本周波数(F0)に基づいて前記入力歌声の声色の変化を反映した合成された歌声の音響信号を生成する合成音響信号生成ステップとをコンピュータが実施することを特徴とする声色変化反映歌声合成方法。 - 前記スペクトル包絡推定ステップでは、前記入力歌声の音響信号、前記J個の合成された歌声の音響信号及び前記K個の合成された歌声の音響信号からなるS個の音響信号の音量を正規化し、
正規化した前記S個の音響信号を周波数分析して、周波数分析結果から複数の周波数スペクトル毎の音高及び非周期成分を推定し、
推定した音高を有声らしさの閾値と比較して有声または無声の判定を行い、有声である区間は前記音響信号のF0に基づいて前記複数の周波数スペクトルの包絡をL1次元(L1は2の累乗+1の整数)で推定し、無声である区間はあらかじめ定めた低い周波数に基づいて前記複数の周波数スペクトルの包絡をL1次元で推定し、
前記有声である区間の前記複数の周波数スペクトルの包絡と、前記無声である区間の前記複数の周波数スペクトルの包絡とに基づいて前記S個のスペクトル包絡を推定する請求項7に記載の声色変化反映歌声合成方法。 - 前記声色空間推定ステップでは、前記S個のスペクトル包絡に対して離散コサイン変換を行ってS個の離散コサイン変換係数を求め、前記S個のスペクトル包絡に対して離散コサイン変換係数における直流成分である第0次元を除いた低次のL2次元(但しL2<L1でL2は正の整数)までの離散コサイン変換係数ベクトルを分析対象として取得し、
前記S個の音響信号が同時刻で有声となるT個のフレーム(Tは最大で音響信号の時間長の秒数×サンプリング周期)のそれぞれにおいて、前記S個のL2次元の離散コサイン変換係数ベクトルについて主成分分析を行って、それぞれ主成分係数と累積寄与率を取得し、
前記T個のフレームにおいて、前記主成分係数を用いて前記S個の離散コサイン変換係数をS個のL2次元の主成分スコアに変換し、
前記S個のL2次元の主成分スコアに対して、累積寄与率R%(0<R<100の数)となる低次のN次元(NはRによって定まる1以上L2以下の整数)よりも高次元の主成分スコアを0としてS個のN次元の主成分スコアを取得し、
前記S個のN次元の主成分スコアをそれぞれ、対応する前記主成分係数を用いて、S個の新たなL2次元の離散コサイン変換係数に逆変換し、
T×S個の新たなL2次元の離散コサイン変換係数のベクトルに対して主成分分析を行って主成分係数及び累積寄与率を取得し、取得した主成分係数を用いてL2次元の離散コサイン変換係数を主成分スコアに変換し、主成分スコアの上記M(1≦M≦L2)次元までで表現される空間を前記声色空間と定める請求項7に記載の声色変化反映歌声合成方法。 - 前記軌跡変位変形ステップでは、前記声色変化チューブを構成する前記K個の合成された歌声の音響信号についてのT×K個のM次元主成分スコアベクトルに対して、各次元で0~1の範囲の値になるようにシフト・スケーリングを行い、且つ前記入力歌声の声色軌跡を構成する前記入力歌声の音響信号についてのT個のM次元主成分スコアベクトルに対して各次元で0~1の範囲の値になるようにシフト・スケーリングを行うことにより、前記入力歌声の声色軌跡の全部または大部分が前記声色変化チューブ内に存在させる請求項7、8または9に記載の声色変化反映歌声合成方法。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US13/810,758 US9009052B2 (en) | 2010-07-20 | 2011-07-19 | System and method for singing synthesis capable of reflecting voice timbre changes |
| GB1302870.9A GB2500471B (en) | 2010-07-20 | 2011-07-19 | System and method for singing synthesis capable of reflecting voice timbre changes |
| JP2012525402A JP5510852B2 (ja) | 2010-07-20 | 2011-07-19 | 声色変化反映歌声合成システム及び声色変化反映歌声合成方法 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2010-163402 | 2010-07-20 | ||
| JP2010163402 | 2010-07-20 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2012011475A1 true WO2012011475A1 (ja) | 2012-01-26 |
Family
ID=45496895
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2011/066383 Ceased WO2012011475A1 (ja) | 2010-07-20 | 2011-07-19 | 声色変化反映歌声合成システム及び声色変化反映歌声合成方法 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US9009052B2 (ja) |
| JP (1) | JP5510852B2 (ja) |
| GB (1) | GB2500471B (ja) |
| WO (1) | WO2012011475A1 (ja) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103295574A (zh) * | 2012-03-02 | 2013-09-11 | 盛乐信息技术(上海)有限公司 | 唱歌语音转换设备及其方法 |
| WO2014088036A1 (ja) * | 2012-12-04 | 2014-06-12 | 独立行政法人産業技術総合研究所 | 歌声合成システム及び歌声合成方法 |
| JP2017045073A (ja) * | 2016-12-05 | 2017-03-02 | ヤマハ株式会社 | 音声合成方法および音声合成装置 |
| WO2018084305A1 (ja) * | 2016-11-07 | 2018-05-11 | ヤマハ株式会社 | 音声合成方法 |
Families Citing this family (26)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10860946B2 (en) * | 2011-08-10 | 2020-12-08 | Konlanbi | Dynamic data structures for data-driven modeling |
| WO2013077843A1 (en) * | 2011-11-21 | 2013-05-30 | Empire Technology Development Llc | Audio interface |
| JP5846043B2 (ja) * | 2012-05-18 | 2016-01-20 | ヤマハ株式会社 | 音声処理装置 |
| US9159310B2 (en) | 2012-10-19 | 2015-10-13 | The Tc Group A/S | Musical modification effects |
| JP5949607B2 (ja) * | 2013-03-15 | 2016-07-13 | ヤマハ株式会社 | 音声合成装置 |
| CN103489443B (zh) * | 2013-09-17 | 2016-06-15 | 湖南大学 | 一种声音模仿方法及装置 |
| US9123315B1 (en) * | 2014-06-30 | 2015-09-01 | William R Bachand | Systems and methods for transcoding music notation |
| JP6728754B2 (ja) * | 2015-03-20 | 2020-07-22 | ヤマハ株式会社 | 発音装置、発音方法および発音プログラム |
| EP3392884A1 (en) * | 2017-04-21 | 2018-10-24 | audEERING GmbH | A method for automatic affective state inference and an automated affective state inference system |
| US10614826B2 (en) | 2017-05-24 | 2020-04-07 | Modulate, Inc. | System and method for voice-to-voice conversion |
| JP7000782B2 (ja) * | 2017-09-29 | 2022-01-19 | ヤマハ株式会社 | 歌唱音声の編集支援方法、および歌唱音声の編集支援装置 |
| GB201719734D0 (en) * | 2017-10-30 | 2018-01-10 | Cirrus Logic Int Semiconductor Ltd | Speaker identification |
| CN108109610B (zh) * | 2017-11-06 | 2021-06-18 | 芋头科技(杭州)有限公司 | 一种模拟发声方法及模拟发声系统 |
| US10186247B1 (en) * | 2018-03-13 | 2019-01-22 | The Nielsen Company (Us), Llc | Methods and apparatus to extract a pitch-independent timbre attribute from a media signal |
| CN108877753B (zh) * | 2018-06-15 | 2020-01-21 | 百度在线网络技术(北京)有限公司 | 音乐合成方法及系统、终端以及计算机可读存储介质 |
| WO2019245916A1 (en) * | 2018-06-19 | 2019-12-26 | Georgetown University | Method and system for parametric speech synthesis |
| JP6747489B2 (ja) * | 2018-11-06 | 2020-08-26 | ヤマハ株式会社 | 情報処理方法、情報処理システムおよびプログラム |
| JP6737320B2 (ja) | 2018-11-06 | 2020-08-05 | ヤマハ株式会社 | 音響処理方法、音響処理システムおよびプログラム |
| WO2021030759A1 (en) | 2019-08-14 | 2021-02-18 | Modulate, Inc. | Generation and detection of watermark for real-time voice conversion |
| US12059533B1 (en) | 2020-05-20 | 2024-08-13 | Pineal Labs Inc. | Digital music therapeutic system with automated dosage |
| JP2023546989A (ja) | 2020-10-08 | 2023-11-08 | モジュレイト インク. | コンテンツモデレーションのためのマルチステージ適応型システム |
| CN112331234A (zh) * | 2020-10-27 | 2021-02-05 | 北京百度网讯科技有限公司 | 歌曲多媒体的合成方法、装置、电子设备及存储介质 |
| US11495200B2 (en) * | 2021-01-14 | 2022-11-08 | Agora Lab, Inc. | Real-time speech to singing conversion |
| CN117561570A (zh) * | 2021-06-29 | 2024-02-13 | 索尼集团公司 | 信息处理装置、信息处理方法和程序 |
| CN116110366A (zh) * | 2021-11-11 | 2023-05-12 | 北京字跳网络技术有限公司 | 音色选择方法、装置、电子设备、可读存储介质及程序产品 |
| US12341619B2 (en) | 2022-06-01 | 2025-06-24 | Modulate, Inc. | User interface for content moderation of voice chat |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0527771A (ja) * | 1991-07-23 | 1993-02-05 | Yamaha Corp | 電子楽器 |
| JP2002268658A (ja) * | 2001-03-09 | 2002-09-20 | Yamaha Corp | 音声分析及び合成装置、方法、プログラム |
| JP2003223178A (ja) * | 2002-01-30 | 2003-08-08 | Nippon Telegr & Teleph Corp <Ntt> | 電子歌唱カード生成方法、受信方法、装置及びプログラム |
| JP2004038071A (ja) * | 2002-07-08 | 2004-02-05 | Yamaha Corp | 歌唱合成装置、歌唱合成方法及び歌唱合成用プログラム |
| JP2004287099A (ja) * | 2003-03-20 | 2004-10-14 | Sony Corp | 歌声合成方法、歌声合成装置、プログラム及び記録媒体並びにロボット装置 |
| JP2005234337A (ja) * | 2004-02-20 | 2005-09-02 | Yamaha Corp | 音声合成装置、音声合成方法、及び音声合成プログラム |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6046395A (en) * | 1995-01-18 | 2000-04-04 | Ivl Technologies Ltd. | Method and apparatus for changing the timbre and/or pitch of audio signals |
| US6336092B1 (en) * | 1997-04-28 | 2002-01-01 | Ivl Technologies Ltd | Targeted vocal transformation |
| US6304846B1 (en) * | 1997-10-22 | 2001-10-16 | Texas Instruments Incorporated | Singing voice synthesis |
| JP2000105595A (ja) * | 1998-09-30 | 2000-04-11 | Victor Co Of Japan Ltd | 歌唱装置及び記録媒体 |
| JP3365354B2 (ja) * | 1999-06-30 | 2003-01-08 | ヤマハ株式会社 | 音声信号または楽音信号の処理装置 |
| JP3858842B2 (ja) * | 2003-03-20 | 2006-12-20 | ソニー株式会社 | 歌声合成方法及び装置 |
| JP3864918B2 (ja) * | 2003-03-20 | 2007-01-10 | ソニー株式会社 | 歌声合成方法及び装置 |
| US8244546B2 (en) | 2008-05-28 | 2012-08-14 | National Institute Of Advanced Industrial Science And Technology | Singing synthesis parameter data estimation system |
-
2011
- 2011-07-19 JP JP2012525402A patent/JP5510852B2/ja active Active
- 2011-07-19 GB GB1302870.9A patent/GB2500471B/en active Active
- 2011-07-19 US US13/810,758 patent/US9009052B2/en active Active
- 2011-07-19 WO PCT/JP2011/066383 patent/WO2012011475A1/ja not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0527771A (ja) * | 1991-07-23 | 1993-02-05 | Yamaha Corp | 電子楽器 |
| JP2002268658A (ja) * | 2001-03-09 | 2002-09-20 | Yamaha Corp | 音声分析及び合成装置、方法、プログラム |
| JP2003223178A (ja) * | 2002-01-30 | 2003-08-08 | Nippon Telegr & Teleph Corp <Ntt> | 電子歌唱カード生成方法、受信方法、装置及びプログラム |
| JP2004038071A (ja) * | 2002-07-08 | 2004-02-05 | Yamaha Corp | 歌唱合成装置、歌唱合成方法及び歌唱合成用プログラム |
| JP2004287099A (ja) * | 2003-03-20 | 2004-10-14 | Sony Corp | 歌声合成方法、歌声合成装置、プログラム及び記録媒体並びにロボット装置 |
| JP2005234337A (ja) * | 2004-02-20 | 2005-09-02 | Yamaha Corp | 音声合成装置、音声合成方法、及び音声合成プログラム |
Cited By (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103295574A (zh) * | 2012-03-02 | 2013-09-11 | 盛乐信息技术(上海)有限公司 | 唱歌语音转换设备及其方法 |
| CN103295574B (zh) * | 2012-03-02 | 2018-09-18 | 上海果壳电子有限公司 | 唱歌语音转换设备及其方法 |
| WO2014088036A1 (ja) * | 2012-12-04 | 2014-06-12 | 独立行政法人産業技術総合研究所 | 歌声合成システム及び歌声合成方法 |
| US9595256B2 (en) | 2012-12-04 | 2017-03-14 | National Institute Of Advanced Industrial Science And Technology | System and method for singing synthesis |
| WO2018084305A1 (ja) * | 2016-11-07 | 2018-05-11 | ヤマハ株式会社 | 音声合成方法 |
| JPWO2018084305A1 (ja) * | 2016-11-07 | 2019-09-26 | ヤマハ株式会社 | 音声合成方法、音声合成装置およびプログラム |
| JP2017045073A (ja) * | 2016-12-05 | 2017-03-02 | ヤマハ株式会社 | 音声合成方法および音声合成装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP5510852B2 (ja) | 2014-06-04 |
| GB2500471A (en) | 2013-09-25 |
| US9009052B2 (en) | 2015-04-14 |
| US20130151256A1 (en) | 2013-06-13 |
| GB201302870D0 (en) | 2013-04-03 |
| GB2500471B (en) | 2018-06-13 |
| JPWO2012011475A1 (ja) | 2013-09-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP5510852B2 (ja) | 声色変化反映歌声合成システム及び声色変化反映歌声合成方法 | |
| JP4241736B2 (ja) | 音声処理装置及びその方法 | |
| CN101308652B (zh) | 一种个性化歌唱语音的合成方法 | |
| CN101578659B (zh) | 音质转换装置及音质转换方法 | |
| JP5471858B2 (ja) | 歌唱合成用データベース生成装置、およびピッチカーブ生成装置 | |
| Umbert et al. | Expression control in singing voice synthesis: Features, approaches, evaluation, and challenges | |
| JP2017107228A (ja) | 歌声合成装置および歌声合成方法 | |
| CN113609255A (zh) | 一种面部动画的生成方法、系统及存储介质 | |
| JP2010009034A (ja) | 歌声合成パラメータデータ推定システム | |
| CN114694632A (zh) | 语音处理装置 | |
| JP2010049196A (ja) | 声質変換装置及び方法、音声合成装置及び方法 | |
| CN101369423A (zh) | 语音合成方法和装置 | |
| JP2010014913A (ja) | 声質変換音声生成装置および声質変換音声生成システム | |
| WO2018084305A1 (ja) | 音声合成方法 | |
| JP2002244689A (ja) | 平均声の合成方法及び平均声からの任意話者音声の合成方法 | |
| JP2002358090A (ja) | 音声合成方法、音声合成装置及び記録媒体 | |
| Hai et al. | Diff-pitcher: Diffusion-based singing voice pitch correction | |
| Bonada et al. | Hybrid neural-parametric f0 model for singing synthesis | |
| JP2006227589A (ja) | 音声合成装置および音声合成方法 | |
| JP3513071B2 (ja) | 音声合成方法及び音声合成装置 | |
| Lee et al. | A comparative study of spectral transformation techniques for singing voice synthesis. | |
| JP2003058908A (ja) | 顔画像制御方法および装置、コンピュータプログラム、および記録媒体 | |
| JP4430174B2 (ja) | 音声変換装置及び音声変換方法 | |
| Kobayashi et al. | Regression approaches to perceptual age control in singing voice conversion | |
| JP5393546B2 (ja) | 韻律作成装置及び韻律作成方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 11809645 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2012525402 Country of ref document: JP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 1302870 Country of ref document: GB Kind code of ref document: A Free format text: PCT FILING DATE = 20110719 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 1302870.9 Country of ref document: GB |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 13810758 Country of ref document: US |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 11809645 Country of ref document: EP Kind code of ref document: A1 |

