WO2007015489A1 - 音声検索装置及び音声検索方法 - Google Patents

音声検索装置及び音声検索方法 Download PDF

Info

Publication number
WO2007015489A1
WO2007015489A1 PCT/JP2006/315228 JP2006315228W WO2007015489A1 WO 2007015489 A1 WO2007015489 A1 WO 2007015489A1 JP 2006315228 W JP2006315228 W JP 2006315228W WO 2007015489 A1 WO2007015489 A1 WO 2007015489A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
voice
pitch
search
feature
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2006/315228
Other languages
English (en)
French (fr)
Inventor
Yasushi Sato
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Kyushu Institute of Technology NUC
Original Assignee
Kyushu Institute of Technology NUC
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Kyushu Institute of Technology NUC filed Critical Kyushu Institute of Technology NUC
Priority to JP2007529275A priority Critical patent/JP4961565B2/ja
Publication of WO2007015489A1 publication Critical patent/WO2007015489A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/90Pitch determination of speech signals
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/003Changing voice quality, e.g. pitch or formants
    • G10L21/007Changing voice quality, e.g. pitch or formants characterised by the process used
    • G10L21/013Adapting to target pitch

Definitions

  • the present invention relates to a voice search device for searching a portion that matches a predetermined voice from stored search target voice data.
  • search keyword voice hereinafter referred to as "query voice”
  • search keyword voice hereinafter referred to as "query voice”
  • Patent Document 1 As a voice search device for searching for a portion that matches a query voice from search target voice data, the one described in Patent Document 1 is known.
  • FIG. 12 is a diagram showing the configuration of the voice search device described in Patent Document 1.
  • this voice search device when a voice signal is input to the voice signal input unit 102 of the search data generation unit 100, the voice signal is stored in the recording unit 201 as search target voice data.
  • the video search index generated by the video search index generation unit 104 is added.
  • a video signal is input to the video signal input unit 101 in synchronization with the audio signal, and stored in the recording unit 201 as accumulated video data.
  • the query voice is input from the keyword input unit 203 of the search processing unit 200, verified with the search target voice data in the keyword pattern matching unit 205, and the voice signal that most closely matches is output from the voice signal output unit 207. Is done.
  • these processes will be outlined.
  • the audio feature pattern extraction unit 103 divides the input audio into 10 msec analysis frames. Then, fast Fourier transform is performed on each analysis frame to generate acoustic characteristic data in the generated frequency band. Furthermore, this acoustic characteristic data is converted into N-dimensional vector data (hereinafter referred to as “feature pattern”) that also includes acoustic feature value forces.
  • feature pattern N-dimensional vector data
  • acoustic feature amount a short-time spectrum in the generation frequency band of the input speech or its logarithmic value, logarithmic energy within a certain time of the input speech, or the like is used.
  • the video search index generation unit 104 extracts the first standard audio pattern from the audio feature pattern storage unit 105.
  • the standard voice pattern is a subword unit (#V, #CV, #CjV, CV, CjV, VC, QC, VQ, W, V #: Is a consonant, V is a vowel, j is a stuttering, Q is a prompting sound, and # is silence.
  • the video search index generation unit 104 determines the similarity between the first standard audio pattern and the audio feature pattern of the input audio for one audio segment to be processed, using the DP matching method or HMM (Hidden It is calculated by speech recognition processing such as Markov Model). Then, the section showing the highest similarity to the first standard speech pattern is detected as a “subword section”. Hereinafter, the similarity between subword sections is referred to as “score”.
  • the video search index generation unit 104 outputs a set of the phoneme symbol of the subword section, the utterance section (start time and end time), and the score as a “video search index”.
  • a subword section is detected, and a video search index related to the detected subword section is output.
  • video search index generation unit 104 moves the audio section to be processed to the next adjacent audio section, A similar process is executed. Then, when the video search index is created over the entire section of the input audio, the process is terminated.
  • FIG. 13 is a diagram showing a part of the lattice structure of the video search index stored in the recording unit 201.
  • the end of each audio segment of the input audio divided in 10 msec units is defined as the end of each video search index generated for that audio segment.
  • the video search indexes in the same audio section are arranged in the order of generation.
  • Such a lattice structure of the video search index is called a “phoneme similarity table”. “Lattice” refers to a list of multiple phoneme and word candidates and their possibilities for various continuous speech segments (see Non-Patent Document 1, p. 198). ).
  • the process of searching for a video scene using query audio is performed as follows. First, the keyword input unit 203 is input with the search keyword “taly”. The keyword conversion unit 204 converts the query speech into a time series of subwords. Next, the keyword pattern matching unit 205 picks up only the subwords constituting the query speech from the phoneme similarity table. Then, the sub-words on the plurality of latencies that have been picked up are connected without gaps in the order of the sub-word sequence obtained by converting the search keyword.
  • the keyword conversion unit 204 when “empty” is input to the keyword input unit 203 as a query voice, the keyword conversion unit 204 generates subword sequences “SO”, “OR”, and “RA”.
  • the keyword pattern matching unit 205 picks up the subwords “ S0 ”, “ 0R ”, and “RA” for the phoneme similarity table power, and connects them without any gaps.
  • take out the sub-word “RA” from the lattice ska at a certain time take out the latitudinal force corresponding to the start time of the subword “RA” and the sub-word “OR” before it, and also apply the latitud force corresponding to the start time of the subword “OR”.
  • Sub Sakai Take out “SO”. Then, “SO”, “OR”, and “RA” are concatenated with the end of the last subword “RA” as a reference.
  • the control unit 202 calculates the time code of the corresponding video signal from the start time of the first subword of the restoration keyword with the highest score. Then, control is performed to reproduce the corresponding portion of the stored video data “search target audio data stored in the storage unit 201.
  • Patent Document 1 JP 2000-236494 A (Patent No. 3252282)
  • Patent Document 2 JP-A-2005-91709
  • Non-Patent Document 1 Sadaaki Furui, “Acoustic and Speech Engineering”, Modern Science, pp. 194-210 Disclosure of Invention
  • the speech recognition is performed based on the similarity between the query speech and the standard speech pattern using the standard speech pattern stored in the speech feature pattern storage unit 105.
  • the processing time for similarity calculation increases or the scale of the arithmetic circuit increases.
  • a query voice is entered without being registered as a standard voice pattern, it cannot be recognized normally, and the voice search function may not work properly.
  • an object of the present invention is to provide a voice search device that does not require a standard voice pattern and that is not affected by individual differences in voice and has high search accuracy.
  • the first configuration of the voice search device is a partial voice data that matches or resembles query voice data (q Uery voice-data) from search target voice data (retrieval voice -data).
  • Voice search device that searches for (partial voice-data), and voice equalization search target voice that equalizes the pitch period of voiced sound of the search target voice data. From the data (pitch-equalized retrieval voice-data), the distance measure for the pitch equalized query voice data that equalizes the pitch period of the voiced voice of the query voice data in the voice feature space (distance measure) ( Or a partial voice search means for searching for partial voice data whose similarity measure is equal to or lower than a predetermined threshold (or higher than a predetermined threshold).
  • a short-time spectrum or logarithmic value thereof in a frequency band where speech is generated, logarithmic energy within a certain time, or the like can be used.
  • a short-time spectrum for example, a time series of feature data in each band obtained using a band filter group of about 10 to 30 channels, a spectrum directly calculated using a short-time FFT It is used as a force feature such as a cepstrum obtained by cepstrum transformation, a correlation data string calculated by a correlation function, an LPC coefficient string obtained on the basis of LPC analysis, a PARCOR coefficient, and an LSP frequency.
  • distance scale various distance scales can be used according to the feature amount. For example, when using a short time extrapolation as a feature value, a simple Euclidean distance, a weighted distance considering auditory sensitivity, statistical analysis such as discriminant analysis, principal component analysis, etc. Euclidean distance, Maharabinos distance, Itakura's Saito distance, COSH scale, WLR scale (weighted likelihood ratio), PWLR scale (power weighted likelihood ratio), LPC cepstrum Euclidean distance, LPC weighted cepstrum Inter-Euclidean distance can be used.
  • the distance measure d (x, y) of the feature quantity (generally vector quantity) x, y does not necessarily satisfy the triangular inequality like the mathematical distance. However, it is desirable to have symmetries and positive values defined by the following equations, and there must be an algorithm that efficiently calculates d (x, y).
  • the "similarity scale” refers to a scale indicating how similar two feature quantities are.
  • the similarity that can be defined by the following equation can be used.
  • X and y represent feature quantities.
  • a second configuration of the speech search apparatus is the first configuration, wherein the pitch equalized query voice data is obtained by equalizing a pitch period of voiced sounds of the query single voice data.
  • Feature data for generating pitch period equalization means to be generated and data obtained by converting the pitch equalization query voice data into time-series data of feature amounts (hereinafter referred to as “query feature-data”).
  • the pitch period equalizing means equalizes the pitch period of the voiced sound of the query voice data.
  • the feature data generation means calculates the feature amount of the query voice data with the equal pitch period, and generates query feature data.
  • the partial speech search means extracts the distance measure (or similarity measure) between the partial speech data of the pitch equalization search target speech data and the query feature data by threshold judgment. Thereby, it is possible to search the search target voice data for voice data that matches or is similar to the query voice data.
  • the partial voice search means uses the pitch equalization search target voice data as time-series data of feature values.
  • Partial voice selection means for sequentially selecting partial data for the same phoneme length as the query voice data (hereinafter referred to as “selected feature data”) from the search target characteristic data converted into the data, while moving the selection position;
  • a feature amount scale calculating means for calculating a distance measure (or similarity measure) between each of the selected feature data and the terry feature data, and the distance measure (or similarity measure) is equal to or less than a predetermined threshold (or a predetermined value). In the case of a threshold value or more), a matching position determining means for outputting a position in the search target voice data corresponding to the selected feature data is provided.
  • the pitch equalization search target voice data in the feature amount space (or similarity measure) from the search target voice data is equal to or less than a predetermined threshold (or higher than the predetermined threshold). Can be extracted.
  • the procedure for “moving the selected position” by the partial voice selecting means is not particularly limited.
  • the partial voice data start position is sequentially moved toward the beginning force end of the search target voice data, or conversely, the partial voice data end position is sequentially moved toward the end force start of the search target voice data. It is possible to adopt such a method.
  • a fourth configuration of the voice search device is characterized in that, in the third configuration, voice search means for storing the search target feature data is provided.
  • search target voice data as search target feature data in advance in the voice storage means, it becomes possible to quickly search for partial voice data similar to the query voice data.
  • the fifth configuration of the speech search device is the pitch equalization search target in the third or fourth configuration, by equalizing the pitch period of the voiced sound of the search target speech data.
  • Second pitch period equalization means for generating voice data
  • second feature data generation for generating the search target feature data by converting the pitch equalization search target voice data into time-series data of feature quantities Means.
  • the pitch cycle is equalized by the second pitch cycle equalization means.
  • the feature amount of the search target speech data with the equal pitch period can be obtained.
  • a sixth configuration of the speech search device is the configuration of the second or fifth configuration, wherein the pitch period equalizing means (or the second pitch period equalizing means) Pitch detection means for detecting the pitch frequency of the data (or the search target voice data), residual calculation means for calculating the difference between the pitch frequency and a predetermined reference frequency, and so that the difference is minimized.
  • a frequency shifter for equalizing a pitch frequency of the terry voice data (or the search target voice data) is provided.
  • the pitch period equalizing means (or the second pitch period equalizing means) can equalize the pitch frequency of the query voice data (or the search target voice data).
  • the search target feature data and the terry feature data are each the pitch equalization It is a time series of subband data obtained by orthogonal transformation of search target voice data and pitch equalization query voice data.
  • An eighth configuration of the speech search apparatus is the configuration in which, in the second or fifth configuration, the query feature data is averaged for each phoneme section and converted into time-series data of an average value.
  • 1 segment division means, and second segment division means for averaging the search target feature data for each phoneme segment and converting it into time-series data of average values, wherein the feature amount scale computation means is:
  • a distance measure (or similarity measure) between the time series data of the average values generated by the first and second section dividing means is calculated.
  • a ninth configuration of the speech search apparatus is the phoneme labeling for the tally speech data (or the search target speech data) in any one of the first to eighth configurations.
  • Phoneme labeling processing means for generating a query phoneme sequence (or search target phoneme sequence) by performing the above, and a distance measure (or similar) between the search target phoneme sequence corresponding to the selected feature data and the query phoneme sequence
  • Phoneme string scale calculating means for determining the scale), a distance measure (or similarity measure) of the feature quantity output by the feature quantity scale calculating means, and a distance measure of the phoneme string output by the phoneme string scale calculating means ( Or a total scale calculating means for calculating a linear sum with the same) (hereinafter referred to as “total distance scale (or total similarity scale)”), and the coincidence position determining means includes the total distance scale (or total scale).
  • Similarity scale is predetermined If the value is less than or equal to the threshold value (or greater than or equal to the predetermined
  • the speech search method is a speech search method for searching partial speech data that matches or is similar to query speech data from search target speech data.
  • the distance from the pitch equalization search voice data that equalizes the pitch period of the voiced sound of the tally voice data in the feature space of the voice from the pitch equalization search target voice data that equalizes the pitch period of the voice. It has a partial voice search step of searching for partial voice data whose scale (or similarity scale) is equal to or less than a predetermined threshold (or higher than a predetermined threshold).
  • a second configuration of the speech search method is the first configuration, wherein the pitch equalized query voice data is obtained by equalizing a pitch period of voiced sounds of the query one voice data.
  • the feature amount force is a predetermined threshold value It is characterized by searching for the following (or more than a predetermined threshold).
  • a third configuration of the speech search method according to the present invention is characterized in that, in the first or second configuration, the partial speech search step is! /, And the pitch equalization search target speech data is characterized.
  • Select partial data hereinafter referred to as “selected feature data” for the same phoneme length as the above-mentioned Taely voice data from the search target feature data converted into a quantity of time-series data while moving the selection position.
  • a matching position determination step of outputting a position in the search target speech data corresponding to the selected feature data is provided.
  • a fourth configuration of the speech search method according to the present invention is characterized in that, in the third configuration, a speech storage step of storing the search target feature data is provided.
  • the pitch equalization search target is obtained by equalizing the pitch period of the voiced sound of the search target speech data in the third or fourth configuration.
  • a data generation step is performed.
  • the query in the second or fifth configuration, in the pitch period equalization step (or second pitch period equalization step), the query A pitch detection step for detecting the pitch frequency of the audio data (or the search target audio data), a residual calculation step for calculating a difference between the pitch frequency and a predetermined reference frequency, and the difference is minimized. And a frequency shift step for equalizing a pitch frequency of the terry voice data (or the search target voice data).
  • a seventh configuration of the speech search method according to the present invention is the configuration according to any one of the first to sixth configurations, wherein each of the search target feature data and the terry feature data is the pitch equalization. It is a time series of subband data obtained by orthogonal transformation of search target voice data and pitch equalization query voice data.
  • An eighth configuration of the speech search method according to the present invention is the above-described second or fifth configuration, wherein the query feature data is averaged for each phoneme section and converted to time-series data of average values. And a second interval dividing step of averaging the search target feature data for each phoneme interval and converting the averaged time-series data into the feature value scale calculating step. Is characterized by calculating a distance measure (or similarity measure) between the time series data of the average values generated in the first and second interval division steps.
  • a ninth configuration of the speech search method is the phoneme labeling of the tally speech data (or the search target speech data) in any one of the first to eighth configurations.
  • Phoneme labeling that generates a query phoneme sequence (or search target phoneme sequence)
  • a feature string scale calculating step for determining a distance measure (or similarity measure) between the search target phoneme string corresponding to the selected feature data and the terry phoneme string, and a feature amount scale calculating step.
  • a linear sum of the distance measure (or similarity measure) of the quantity and the distance measure (or similarity measure) of the phoneme sequence output in the phoneme sequence scale calculation step (hereinafter referred to as the “total distance measure (or overall similarity measure)”)
  • a comprehensive scale calculation step for calculating the matching position determination step wherein the total distance scale (or the total similarity scale) is less than or equal to a predetermined threshold (or greater than or equal to a predetermined threshold).
  • the feature is that the position in the search target speech data corresponding to the selected feature data is output.
  • a program according to the present invention is characterized by causing a computer to function as the voice search device having any one of the first to eighth configurations by being read into a computer and executed.
  • the audio data from which the gender difference and the individual difference in the audio band are removed is used.
  • the accuracy of the voice search can be improved with little influence on gender differences and individual differences in the voice band.
  • the feature amounts of the search target speech data and the tally speech data in which the pitch period is equalized for each phoneme section are averaged, and the speech search is performed by matching the time series of the average value of the feature amounts. This reduces the effects of noise and fluctuations and eliminates the effects of voice expansion and contraction. As a result, the accuracy of voice search can be improved.
  • FIG. 1 is a diagram showing an overall configuration of a voice search device 1 according to Embodiment 1 of the present invention.
  • FIG. 2 is a block diagram showing a configuration of speech coder 2 in FIG. 1.
  • FIG. 3 is a block diagram showing the configuration of pitch period equalizing means 10 in FIG.
  • FIG. 4 is a diagram for explaining the outline of signal processing in pitch detecting means 21 and pitch averaging means 22.
  • FIG. 5 is a diagram showing formant characteristics of voiced sound “A”.
  • FIG. 6 is a diagram showing the autocorrelation, cepstrum waveform and frequency characteristics of unvoiced sound “su”.
  • FIG. 7 is a diagram illustrating an internal configuration of a frequency shifter 23.
  • FIG. 8 is a diagram illustrating another example of the internal configuration of the frequency shifter 23.
  • FIG. 9 is a block diagram showing a configuration of speech decoder 5 in FIG. 1.
  • FIG. 10 is a block diagram showing a configuration of partial speech search means 6 of FIG.
  • FIG. 11 is an explanatory diagram of the number of quantization bits.
  • FIG. 12 is a diagram illustrating a configuration of a voice search device described in Patent Document 1.
  • FIG. 13 is a diagram showing a part of a lattice structure of a video search index stored in a recording unit 201.
  • FIG. 14 is a diagram showing the structure of a lattice connected to calculate a score for each restoration keyword.
  • Input pitch detection means Pitch averaging means Frequency shifter Output pitch detection means Residual calculation means
  • VCO Pitch equalization waveform decoder Inverse quantizer Synthesizer Pitch information decoder Pitch frequency detection means Differencer
  • FIG. 1 is a diagram illustrating the overall configuration of the voice search device according to the first embodiment of the present invention.
  • the speech search apparatus 1 of the first embodiment includes a speech coder 2, speech storage means 3, data reading means 4, speech decoder 5, and partial speech search means 6.
  • the search target voice data and the query voice data are input to the voice encoder 2 as input voice data.
  • the voice encoder 2 equalizes the pitch period of voiced sound with respect to the input voice data and converts it into time-series data (feature data) of feature quantities.
  • feature data time-series data
  • the pitch period information of the input voice data is separated from the feature data, encoded, and output as encoded pitch data.
  • the feature data is output as a subband waveform.
  • the speech encoder 2 encodes the feature data and outputs it as encoded feature data.
  • the speech coder 2 performs a phoneme labeling process on the input speech data, and outputs the phoneme label of each phoneme and the phoneme label data that is the information power of the time interval.
  • the speech storage means 3 stores search target speech data that has been decomposed and encoded by the speech coder 2 into encoded feature data, encoded pitch data, and phoneme label data.
  • the encoded feature data and the encoded pitch data stored in the speech storage means 3 are encoded search target feature data.
  • the data reading unit 4 reads partial data of the search target speech data (encoded feature data, encoded pitch data, and phoneme label data) in the speech storage unit 3 in accordance with the data selection signal.
  • the speech decoder 5 decodes the encoded feature data and the code key pitch data read by the data reading unit 4 and outputs them as feature data or output speech data.
  • the partial voice search means 6 searches partial data that matches or is similar to the query voice data from the encoded search target voice data stored in the voice storage means 3.
  • FIG. 2 is a block diagram showing the configuration of speech encoder 2 in FIG.
  • Speech encoder 2 includes pitch period equalizing means 10, feature data generating means 11, output switching means 12a and 12b, quantizing means 13, pitch equalizing waveform encoder 14, differential bit calculator 15, pitch information code A device 16 and a phoneme labeling means 17 are provided.
  • the pitch period equalizing means 10 equalizes the pitch period of the voiced sound of the input sound data X (t).
  • the Input voice data with equal pitch period (hereinafter referred to as “pitch equalized voice data”) X (t) is output from the output terminal Out_l.
  • the feature data generation means 11 converts the pitch equalized audio data X (t) output from the output terminal Out_l into time-series data of feature amounts.
  • the feature value is out
  • a short time frequency spectrum is used.
  • the feature data generation means 11 includes a resampler 18 and an analyzer (a modified discrete cosine transformer).
  • MDCT Modified Discrete Cosine Transformer
  • the resampler 18 re-standardizes each pitch section of the pitch equalized speech data X (t) output from the output terminal Out_l of the pitch period equalizing means 100 so as to have the same number of samples.
  • the analyzer 19 changes the completely equalized audio data X (t) with a fixed number of pitch sections.
  • feature data X (f). That is, in the present embodiment, the feature data is given as a time series of vector quantities (formula (4)) that also becomes a short-time frequency spectrum.
  • the output switching means 12a sends the output destination of the feature data X (f) generated by the analyzer 19 to the partial voice searching means 6 or the voice storage means 3 in accordance with the switching signal input from the partial voice searching means 6. Switch. Specifically, when search target voice data is input as input voice data, the output destination of the feature data X (f) is switched to the voice storage means 3. When query voice data is input as input voice data, the output destination of the feature data X (f) is switched to the partial voice search means 6.
  • the quantizer 13 quantizes the feature data X (f) according to a predetermined quantization curve.
  • the pitch equalization waveform encoder 14 encodes the feature data X (f) output from the quantizer 13 and outputs it as encoded feature data.
  • an entropy code method such as a Huffman code method or an arithmetic code method is used.
  • the difference bit calculator 15 subtracts the target bit number from the code amount of the sign key characteristic data output from the pitch equalization waveform encoder 14 (hereinafter referred to as “difference bit number”). Is output.
  • the quantizer 13 translates the quantization curve according to the difference bit number and adjusts the code amount of the encoded feature data to be within the target bit number range.
  • the pitch information encoder 16 encodes the residual periodic signal ⁇ and the reference periodic signal AV output from the pitch period equalizing means 10 and outputs the encoded signal as code pitch data. This mark tch pitch
  • Encoding uses entropy code methods such as Huffman code method and arithmetic code method
  • Phoneme labeling means 17 divides input speech data into phoneme sections and performs phoneme labeling on each phoneme section. And it outputs as phoneme label data which also has the information power of a phoneme label and a time section.
  • the output switching unit 12b switches the output destination of the phoneme label data generated by the phoneme labeling processing unit 17 to the partial speech search unit 6 or the speech storage unit 3. Specifically, when search target speech data is input as input speech data, the output destination of phoneme label data is switched to speech storage means 3. When query speech data is input as input speech data, the output destination of phoneme label data is the partial speech search means 6 Can be switched to.
  • FIG. 3 is a block diagram showing a configuration of pitch period equalizing means 10 in FIG.
  • the pitch period equalizing means 10 includes an input pitch detecting means 21, a pitch averaging means 22, a frequency shifter 23, an output pitch detecting means 24, a residual calculating means 25, and a PID controller 26.
  • the input pitch detection means 21 is included in the audio signal from the input audio data X (
  • the input pitch detection means 21 includes a pitch detection means 27, a band pass filter (hereinafter referred to as “BPF”) 28, and a frequency counter 29.
  • BPF band pass filter
  • the h detection means 27 first performs a short-time Fourier transform on this waveform to derive a spectrum waveform X (f) as shown in FIG. 4 (b).
  • a speech waveform includes many frequency components in addition to the pitch, and the spectrum waveform obtained here has many additional components in addition to the fundamental frequency and the harmonic components of the pitch. Has a frequency component. Therefore, the fundamental frequency f of the pitch of this spectrum waveform X (f)
  • the pitch detection means 27 performs Fourier transform again on this spectral waveform X (f).
  • the reciprocal number of the pitch interval of harmonics ⁇ f included in the spectrum waveform X (f) F
  • Pitch detection means 27 detects this peak position F
  • the pitch detection means 27 has the input sound data X (t) from the spectrum waveform X (f).
  • FIG. 5 is a diagram showing the formant characteristics of voiced sound “A”
  • FIG. 6 is a diagram showing the autocorrelation, cepstrum waveform, and frequency characteristics of unvoiced sound “S”.
  • the voiced sound has a formant waveform whose X (f) is large on the low frequency side and small on the high frequency side. Characteristics.
  • unvoiced sounds as shown in Fig. 6, exhibit frequency characteristics that generally increase toward the high frequency side. Therefore, by detecting the overall slope of the spectrum waveform X (f), it is determined whether the input speech data X (t) is voiced or unvoiced.
  • the fundamental frequency f of the pitch output by the means 27 becomes a meaningless value.
  • BPF28 a narrow bandpass filter whose passband can be set from the outside is used. BPF28 passes the fundamental frequency f of the pitch detected by the pitch detector 27
  • BPF28 then filters the input audio data X (t) and outputs a nearly sinusoidal waveform with the fundamental frequency f of the pitch in 0
  • the basic period T of the detected pitch is the output signal of the input pitch detection means 21 (hereinafter referred to as “basic frequency”).
  • the pitch averaging means 22 receives the pitch basic period signal T output from the pitch detection means 27.
  • LPF normal low pass filter
  • the frequency shifter 23 brings the pitch frequency of the input audio data X (t) closer to the reference frequency f.
  • the pitch period of the audio signal is equalized by shifting in the in 0 direction.
  • the output pitch detection means 24 includes the output voice data (hereinafter referred to as "pitch equalized voice data") X (t) output from the frequency shifter 23 and included in the pitch equalized voice data X (t).
  • This output pitch detection means 24 is also basically input.
  • the output pitch detection means 24 includes a BPF 31 and a frequency counter 32.
  • BPF31 a narrow band BPF whose pass band can be set from the outside is used.
  • BPF31 Is the fundamental frequency f of the pitch detected by the pitch detector 27
  • BPF31 then filters the pitch-equalized audio data X (t) out.
  • This period T ′ is output as an output signal of the output pitch detection means 24.
  • the residual calculation means 25 performs pitch flattening from the basic period T 'output by the output pitch detection means 24.
  • the period ⁇ is input to the frequency shifter 23 via the PID controller 26.
  • the PID controller 26 includes an amplifier 34 and a resistor 20 connected in series, and a capacitor 36 connected in parallel to the amplifier 34. This PID controller 26 is for preventing oscillation of a feedback loop comprising the frequency shifter 23, the output pitch detection means 24, and the residual calculation means 25.
  • the PID controller 26 displays an analog circuit, but it may be a digital circuit.
  • FIG. 7 is a diagram showing the internal configuration of the frequency shifter 23.
  • the frequency shifter 23 includes a transmitter 41, a modulator 42, a BPF 43, a voltage controlled oscillator (hereinafter referred to as “VCO”) 44, and a demodulator 45.
  • VCO voltage controlled oscillator
  • the transmitter 41 modulates the input audio data X (t) at a constant frequency for frequency modulation.
  • modulation carrier frequency the frequency of the modulated carrier signal C generated by the transmitter 41 (hereinafter referred to as “modulation carrier frequency”) is normally about 20 kHz.
  • the modulator 42 converts the modulated carrier signal C output from the transmitter 41 into input audio data x (t).
  • BPF 43 is a BPF having a modulation carrier frequency as a lower cutoff frequency and having a passband having a bandwidth larger than the bandwidth of the input audio data.
  • the modulated signal output by the BPF43 force is a signal obtained by cutting out only the upper sideband (see FIG. 7 (iii)).
  • VC044 outputs a signal having the same frequency as that of modulated carrier signal C output from transmitter 41, with a residual period ⁇ ⁇ input from residual calculation means 25 via PID controller 26 (
  • the demodulator 45 demodulates the modulated signal of only the upper side band output from the BPF 43 using the demodulated carrier signal output from the VC044 to restore the audio signal (see FIG. 7 (iv)). At this time, the demodulated carrier signal is modulated with a residual periodic signal. Therefore, when the modulated signal is decoded, the deviation of the reference frequency f force of the pitch frequency of the input audio data X (t) is eliminated.
  • FIG. 8 is a diagram illustrating another example of the internal configuration of the frequency shifter 23.
  • the transmitter 41 and VC044 of FIG. 7 are replaced. Even with this configuration, the pitch period of the input audio data X (t) is equalized to the reference period T, as in FIG.
  • FIG. 9 is a block diagram showing a configuration of speech decoder 5 in FIG.
  • the audio decoder 5 is a device that decodes the audio signal encoded by the audio encoder 2.
  • the speech decoder 5 includes a pitch equalization waveform decoder 51, an inverse quantizer 52, a synthesizer 53, a pitch information decoder 54, a pitch frequency detection means 55, a difference unit 56, an adder 57, a frequency shifter 58, and an output.
  • Switching means 59 is provided.
  • the speech decoder 5 receives the encoded feature data and the encoded pitch data.
  • the sign key characteristic data is sign key characteristic data output from the pitch equalization waveform encoder 14 of FIG.
  • the code key pitch data is the code key pitch data output from the pitch information encoder 16 of FIG.
  • Pitch equalization waveform decoder 51 decodes the encoded feature data, and restores the feature data of each subband after quantization (hereinafter, “quantized feature data” t ⁇ ⁇ ).
  • quantized feature data t ⁇ ⁇
  • the synthesizer 53 performs inverse modified discrete cosine transform (hereinafter referred to as “IMDCT”) on the feature data X (f), and time-series data (hereinafter referred to as “equalized audio signal”). ) Generates x (t).
  • the pitch frequency detecting means 55 detects the pitch frequency of the equalized audio signal X (t) and outputs it as the equalized pitch frequency signal V.
  • the pitch information decoder 54 restores the reference frequency signal AV and the residual frequency signal ⁇ by decoding the encoded pitch data.
  • Differentiator 56 is the reference frequency
  • the adder 57 generates a residual frequency signal ⁇ and a reference frequency change pitch pitch.
  • the signal ⁇ is added and output as a modified residual frequency signal ⁇ ′′.
  • the frequency shifter 58 has the same configuration as the frequency shifter 23 shown in FIG. 7 or FIG.
  • the equalized audio signal X (t) is input to the input terminal In, and the modified residual frequency signal AV "is input to the VC044.
  • the VC044 is the modulated carrier signal output from the transmitter 41.
  • a signal with the same carrier frequency as signal C is applied to the modified residual frequency signal AV input from the adder 57.
  • a signal obtained by frequency modulation by pitch “(hereinafter referred to as“ demodulated carrier signal ”) is output.
  • the frequency of the demodulated carrier signal is the carrier frequency plus the residual frequency.
  • the output switching means 59 determines the output destination of the characteristic data X (f) generated by the inverse quantizer 52 in accordance with the switching signal input from the partial speech search means 6, and the synthesizer 53 or the partial speech search means Switch to 6. Specifically, when performing the partial speech search operation, the output destination of the feature data X (f) is switched to the partial speech search means 6. On the other hand, when the search target speech data is output to the outside, the output destination of the feature data X (f) is switched to the synthesizer 53.
  • FIG. 10 is a block diagram showing the configuration of the partial speech search means 6 of FIG. Partial voice detection
  • the search means 6 includes an operation switching means 61, a partial speech selection means 62, an interval dividing means 63, 64, a feature value scale calculation means 65, a phoneme string scale calculation means 66, an overall scale calculation means 67, and a coincidence position determination means 68. It is equipped with.
  • the operation switching means 61 outputs a switching signal for switching the operation of the voice search device 1 to the search target voice data input / output operation for the voice storage means 3 or the partial voice search operation by the partial voice search means 6. .
  • the partial speech selection means 62 is used to select partial speech data from the search target feature data (more precisely, encoded search target feature data) stored in the speech storage means 3. Output data selection signal. This data selection signal is input to the data reading means 4. The data reading means 4 selects and reads the search target feature data stored in the voice storage means 3 according to the data selection signal.
  • the section dividing unit 63 uses the feature data (subband waveform) of the query speech input from the analyzer 19 of the speech coder 2 as the phoneme label data of the query single speech input from the phoneme labeling processing unit 17. Divide into phoneme segments according to time segment information. Then, the feature data is averaged for each phoneme section, and is output to the feature quantity scale calculation means 65 as time-series data of average values.
  • the section dividing unit 64 uses the feature data (subband waveform) of the search target speech input from the inverse quantizer 52 of the speech decoder 5 as the phoneme label of the search target speech input from the data reading unit 4. Divide into phoneme intervals according to the information of the time interval of the data. And
  • the feature data is averaged and output to the feature quantity scale calculation means 65 as time-series data of average values.
  • the feature quantity scale calculation means 65 calculates a distance scale D (X 1, X 2) between the feature data input from the section division means 63 and 64.
  • the distance scale is expressed as a linear sum of correlation coefficients of the subband waveforms constituting the feature data.
  • the feature data of the query speech is X (f)
  • the feature data of the search target speech is X (f)
  • t represents the jth phoneme interval.
  • X (t) is a special characteristic in the jth phoneme interval.
  • the time average value of the collected data X (t), X (t) is the feature data in the jth phoneme interval
  • X (t) is a time average value.
  • Example l the distance measure D (X, X) between feature data is defined by equation (10).
  • W is a weighting factor.
  • the weighting factor W is set as appropriate.
  • the phoneme string scale calculation means 66 receives the phoneme label data of the tally speech from the phoneme labeling processing means 17 of the speech coder 2 and the phoneme label data of the search target speech from the data reading means 4. Is done.
  • the phoneme string scale calculation means 66 calculates a distance scale D of these phoneme label data using a predetermined interphoneme distance scale table. here,
  • the interphoneme distance scale table is a table of distance scales between two phonemes for all two phoneme combinations.
  • the total scale calculation means 67 is a distance scale D (X, X) between the feature data calculated by the feature quantity scale calculation means 65 and the distance between the phoneme label data calculated by the phoneme string scale calculation means 66.
  • the total distance measure D is calculated by taking the linear sum of the separation measures D. That is, overall The distance measure D is expressed by equation (11).
  • the coincidence position determination means 68 determines whether or not the distance measure D is equal to or less than a predetermined threshold value D.
  • the operation switching means 61 of the partial voice search means 6 outputs a level (for example, H level) indicating the input / output operation of the search target voice data as a switching signal.
  • the output switching means 12a of the speech coder 2 outputs the feature data X (f) generated by the analyzer 19 to the quantizer 13.
  • the output switching unit 12b of the speech coder 2 outputs the phoneme label data generated by the phoneme labeling processing unit 17 to the speech storage unit 3.
  • the output switching means 59 of the speech decoder 5 outputs the feature data X (f) generated by the inverse quantizer 52 to the synthesizer 53.
  • input speech data X (t) is input to speech coder 2 as search target speech data in
  • Input pitch detection means 21 of the pitch period equalization means 10 the input voice data X (t)
  • the pitch frequency is detected from the input audio data X (t), and the fundamental frequency signal V is output to the in pitch pitch averaging means 22.
  • the pitch averaging means 22 converts the basic frequency signal V into an average pitch (in this case, LPF is used so that a weighted average is used), and this is used as a reference frequency signal AV
  • the frequency shifter 23 shifts the frequency of the input audio data X (t), the pitch, etc. in
  • Audio data X (t) is output to the output terminal Out_l.
  • the frequency signal A V is 0 (reset state), and the frequency shifter 23 receives the input audio data X
  • pitch in (t) is directly output to the output terminal Out_l as pitch equalized audio data x (t).
  • the output pitch detection means 24 detects the pitch frequency f ′ of the output audio data output from the frequency shifter 23.
  • the detected pitch frequency f ′ is the pitch frequency signal V ′.
  • the residual calculation means 25 generates the residual frequency signal ⁇ by subtracting the reference frequency signal AV from the pitch frequency signal V and pitch pitch. This residual frequency signal ⁇ is output to the output pitch output terminal Out_2 and also input to the frequency shifter 23 via the PID controller 26.
  • the frequency shifter 23 sets a frequency shift amount in proportion to the residual frequency signal ⁇ pitc input via the PID controller 26. In this case, if the residual frequency signal ⁇ is a positive h pitch value, the pitch is shifted so as to decrease the frequency by an amount proportional to the residual frequency signal ⁇ V.
  • the amount is set. If the residual frequency signal ⁇ force S is a negative value, the shift amount is set to increase the frequency by an amount proportional to the residual frequency signal ⁇ pitch pitch.
  • the pitch period equalizing means 10 the information in the input voice data X (t)
  • the information of (a) to (d) is the noise flag signal V and the pitch period is noise, respectively.
  • Pitch equalized voice data X (t) equalized to the reference period lZf s (the inverse of the weighted average of the past pitch frequencies of the input voice data), the reference frequency signal AV, and the residual frequency signal out pitcn
  • the noise flag signal V is output from the output terminal Out_4.
  • the equalized audio data ⁇ (t) is output from the output terminal Out_l, the reference frequency signal AV is output from the out pitch output terminal Out_3, and the residual frequency signal ⁇ is output from the output terminal Out_2 pitch It is.
  • Pitch equalized speech data x (t) depends on gender differences, individual differences, phonemes, emotions, and conversation content.
  • the voiced out can be obtained by comparing the pitch equalized voice data X (t).
  • pitch ⁇ is the same pitch due to the nature of audio.
  • the voiced sound component of the input speech data X (t) can be encoded with high efficiency.
  • the resampler 18 calculates the resampling period by dividing the reference frequency signal AV by a fixed re-pitch sampling number n in each pitch interval. Then, the pitch equalized audio data X (t) is resampled by the resampling period, and the equal number of samples out
  • the number of samples in the interval is a constant value.
  • the analyzer 19 converts the equal sample number speech data X (t) into a sub-eq of a fixed number of pitch intervals.
  • the frequency spectrum signal X (f) is generated by performing the modified discrete cosine transform for each subframe.
  • the length of one subframe is an integral multiple of one pitch period.
  • the length of the subframe is 1 pitch period (sample number n). Therefore, n frequency vector signals ⁇ X (f), X (f),..., X (f) ⁇ are output.
  • the frequency f is the reference frequency 1st harmonic
  • frequency f is second harmonic of reference frequency
  • frequency f is nth harmonic of reference frequency
  • the frequency spectrum signal of speech waveform data is obtained by performing subband coding by dividing the subframes into subframes that are integral multiples of one pitch period and orthogonally transforming each subframe. It is summarized in the spectrum of harmonics of the reference frequency. And, due to the nature of speech, the waveforms of successive pitch sections within the same phoneme are similar, so the spectrum of the harmonic component of the reference frequency is similar between adjacent subframes. Enhanced.
  • the quantizer 13 quantizes the frequency spectrum signal X (f).
  • the quantizer 13 refers to the noise flag signal V, and in the case of the noise flag signal V force SO (voiced sound) and 1 (
  • the quantization curve is such that the number of quantization bits decreases as the frequency increases. This corresponds to the fact that the frequency characteristic of voiced sound has a characteristic that decreases as the frequency becomes larger in the low frequency range as shown in FIG.
  • the quantization curve is such that the number of quantization bits increases as the frequency increases. This corresponds to the fact that the frequency characteristics of unvoiced sound have characteristics that increase as the frequency range increases, as shown in FIG.
  • an optimal quantization curve is selected corresponding to voiced sound or unvoiced sound.
  • the data format of quantization by the quantizer 13 is expressed by the real part (FL) below the decimal point and the exponent part (EXP) representing the power of 2, as shown in Figs. 11 (a) and 11 (b).
  • EXP exponent part
  • the exponent (EXP) shall be adjusted so that the first bit of the real part (FL) is always 1.
  • n bits are left from the top of the real part (FL), and the remaining bits are set to 0 (see FIG. 11 (d)).
  • the pitch equalization waveform encoder 14 encodes the quantized frequency spectrum signal X (f) output from the quantizer 13 by an entropy encoding method, and outputs encoded feature data. To help. Further, the pitch equalization waveform encoder 14 outputs the code amount (number of bits) of the encoded feature data to the differential bit calculator 15. The difference bit calculator 15 subtracts a predetermined number of target bits from the code amount of the encoded feature data and outputs the number of difference bits. The quantizer 13 moves the quantization curve for voiced sound up and down in parallel translation according to the number of differential bits.
  • the quantization curve for ⁇ f, f, f, f, f ⁇ was ⁇ 6, 5, 4, 3, 2, 1 ⁇
  • the quantizer 13 translates the quantization curve downward by 2 in parallel. As a result, the quantization curve is ⁇ 4, 3, 2, 1, 0, 0 ⁇ . Also, assuming that ⁇ 2 is input as the difference bit number, the quantizer 13 translates the quantization curve upward by 2 in the upward direction. As a result, the quantization curve is ⁇ 8, 7, 6, 5, 4, 3 ⁇ . [0152] By changing the quantization curve of the voiced sound up and down in this way, the code amount of the encoded feature data in each subframe is adjusted to about the target number of bits.
  • the pitch information encoder 16 performs the reference frequency signal AV and the remaining frequency.
  • the phoneme labeling processing means 17 divides the input speech data X (t) into phoneme sections, and m
  • the phoneme labeling processing means 17 outputs, as phoneme label data, phoneme labels obtained by phoneme labeling and information on phoneme intervals representing time intervals for each phoneme label.
  • the encoded feature data, the encoded pitch data, and the phoneme label data generated as described above are output to the speech storage means 3 and stored.
  • the data reading unit 4 reads the encoded feature data and the encoded pitch data from the audio storage unit 3, these data are input to the audio decoder 5.
  • the pitch equalization waveform decoder 51 of the speech decoder 5 decodes the encoded feature data and quantizes each frequency spectrum signal of each subband (hereinafter referred to as "quantized frequency spectrum signal"). ).
  • the synthesizer 53 converts the frequency spectrum signal X (f) into an inverse modified discrete cosine transform (Inverse
  • the pitch frequency detection means 55 detects the pitch frequency of the equalized audio signal X (t) and detects the equalized pitch frequency signal V.
  • the pitch information decoder 54 restores the reference frequency signal AV and the residual frequency signal ⁇ by decoding the sign key pitch data.
  • Differentiator 56 is the reference frequency
  • Signal AV force Equalization Pitch Frequency signal V is subtracted from reference frequency change signal pitch eq Output as AAV.
  • the adder 57 generates a residual frequency signal ⁇ and a reference frequency change pitch pitch.
  • the signal ⁇ is added and output as a modified residual frequency signal ⁇ ′′.
  • the frequency shifter 58 has the same configuration as the frequency shifter 23 shown in FIG. 7 or FIG.
  • the equalized audio signal X (t) is input to the input terminal In, and the modified residual frequency signal AV "is input to the VC044.
  • the VC044 is the modulated carrier signal output from the transmitter 41.
  • a signal obtained by frequency-modulating a signal having the same carrier frequency as that of signal C by the modified residual frequency signal AV "input from the adder 57 (hereinafter referred to as" demodulated carrier signal ") is output.
  • the frequency of the demodulated carrier signal is the carrier frequency plus the residual frequency.
  • the operation switching means 61 of the partial voice search means 6 outputs a level (for example, L level) representing the partial voice search operation as a switching signal.
  • the output switching means 12a of the speech coder 2 outputs the feature data X (f) generated by the analyzer 19 to the partial speech search means 6.
  • the output switching means 12b of the speech coder 2 outputs the phoneme label data generated by the phoneme labeling processing means 17 to the partial speech search means 6.
  • the output switching means 59 of the speech decoder 5 outputs the feature data X (f) generated by the inverse quantizer 52 to the partial speech search means 6.
  • the query speech data is input to speech coder 2 as input speech data X (t).
  • the pitch period equalizing means 1 as described above, the pitch period of the voiced sound of the input sound data X (t) is equalized and output from the output terminal Out_l as the pitch equalized sound data X (t). .
  • the feature data generating means 19 converts the pitch equalized speech data X (t) into feature data X (f) that is a time series cover of a short out time spectrum.
  • the feature data X (f) is output to the partial speech search means 6 via the output switching means 12a.
  • the phoneme labeling processing means 17 converts the input speech data X (t) into sound. Divide into phoneme segments and perform phoneme labeling for each phoneme segment. Then, the phoneme label and the information on the phoneme section are output as phoneme label data.
  • the partial speech selection means 62 of the partial speech search means 6 sequentially converts the encoded feature data, encoded pitch data, and phoneme label data stored in the speech storage means 3 from the head of the data.
  • a data selection signal for reading is output.
  • the length of the partial data to be read out is the same phoneme length as that of the query speech data.
  • the data reading means 4 reads partial data from the voice storage means 3 in accordance with the data selection signal.
  • the phoneme label data read by the data reading means 4 is input to the partial speech search means 6.
  • the encoded feature data and the partial data of the encoded pitch data read by the data reading means 4 are input to the speech decoder 5.
  • the pitch equalization waveform decoder 51 decodes the encoded feature data
  • the inverse quantizer 52 performs inverse quantization, thereby generating feature data and partial speech. Output to search means 6.
  • the partial data of the search target feature data input from the speech decoder 5 to the partial speech search means 6 is referred to as “selected feature data”.
  • the segment dividing means 63 Data is averaged for each phoneme section and converted to average time-series data.
  • the category feature data may be divided into time sections, and an average value may be taken for each time section.
  • the time series data of the average value is input to the feature amount scale calculation means 65.
  • the section dividing unit 64 averages the selected feature data for each phoneme section and calculates the average value. Convert to time series data. The time-series data of the average value is input to the feature amount scale calculation means 65.
  • the feature quantity scale calculation means 65 calculates a distance scale D (X 1, X 2) between the time series data of the average values input from the section dividing means 63 and the section dividing means 64 according to the equation (10).
  • the phoneme string scale calculating means 66 is adapted to perform query speech input from the speech coder 2.
  • the distance measure D between the phoneme label data and the phoneme label data of the search target speech input from the data reading means is calculated using the interphoneme distance measure table.
  • the total scale calculation means 67 is a distance scale D (X, X) between the feature data calculated by the feature quantity scale calculation means 65 and the distance between the phoneme label data calculated by the phoneme string scale calculation means 66.
  • the coincidence position determination means 68 determines whether or not the distance scale D is equal to or smaller than a predetermined threshold D th
  • the operation switching means 61 outputs a level (for example, L level) representing the partial voice search operation as a switching signal.
  • the partial data of the searched search target data is output as output audio data.
  • the present embodiment can also be applied to search for information in a multimedia database in which audio information and video are recorded as a whole.
  • the present invention can be used in an audio database, a multimedia database including audio information, and the like.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

 標準音声パターンを必要とせず、音声の個人差にも影響されず検索精度の高い音声検索装置を提供する。  検索対象音声データの有声音のピッチ周期を等化したピッチ等化検索対象音声データの中から、音声の特徴量空間において、クエリー音声データの有声音のピッチ周期を等化したピッチ等化クエリー音声データに対する距離尺度が所定の閾値以下である部分音声データを検索する部分音声検索手段を備えた構成とする。ピッチ周期を等化することによって、音声帯域の男女差や個人差にほとんど影響されず、高い精度で音声検索を行うことが可能となる。

Description

明 細 書
音声検索装置及び音声検索方法
技術分野
[0001] 本発明は、蓄積された検索対象音声データの中から、所定の音声に合致する部分 を検索するための音声検索装置に関する。
背景技術
[0002] 近年、多くの蓄積映像 ·音声データの中から、視聴者が最も知りたい情報の部分だ けを取り出すマルチメディア 'データベースの要請が強まりつつある。代表的な例とし ては、蓄積された多くの-ユース番組の中から、視聴者が最も知りたい-ユースのみ を取り出す-ユース 'オンデマンド(News On Demand : NOD) 'システムなどがある。
[0003] 力かるマルチメディア 'データベースを構築するためには、テレビ-ユースなどの蓄 積された映像 ·音声データの中から、検索キーワードの音声 (以下「クエリー音声」と いう。 )に合致する部分を検索する音声検索技術が必要とされる。
[0004] 検索対象音声データの中からクエリー音声に合致する部分を検索する音声検索装 置としては、特許文献 1に記載のものが公知である。
[0005] 図 12は、特許文献 1に記載の音声検索装置の構成を表す図である。この音声検索 装置では、検索データ生成部 100の音声信号入力部 102に音声信号が入力される と、当該音声信号は、検索対象音声データとして記録部 201に記憶される。この際、 映像検索インデックス生成部 104が生成する映像検索インデックスが付加される。ま た、音声信号に同期して映像信号入力部 101には映像信号が入力され、記録部 20 1に蓄積映像データとして記憶される。一方、クエリー音声は、検索処理部 200のキ 一ワード入力部 203から入力され、キーワードパターン照合部 205において検索対 象音声データと照合され、もっとも一致する音声信号が音声信号出力部 207から出 力される。以下、これらの処理を概説する。
[0006] まず、音声信号入力部 102に音声信号が入力されると、音声特徴パターン抽出部 103は、入力音声を 10msecの分析フレームに分割する。そして、各分析フレームに ついて、高速フーリエ変換を行い、発生周波数帯域の音響特性データを生成する。 さらに、この音響特性データを、音響特徴量力も構成される N次元のベクトルデータ( 以下「特徴パターン」という。 )に変換する。ここで、音響特徴量としては、入力音声の 発生周波数帯域における短時間スペクトル又はその対数値、入力音声の一定時間 内における対数エネルギー等が用いられる。
[0007] 次に、映像検索インデックス生成部 104は、音声特徴パターン収納部 105から第 1 番目の標準音声パターンを取り出す。
[0008] ここで、音声特徴パターン収納部 105には、 500個の標準音声パターンが予め記 憶されている。標準音声パターンとは、予め複数の話者から収集した発音を分析して 、サブワード単位(#V, #CV, #CjV, CV, CjV, VC, QC, VQ, W, V#:但し、 Cは子 音、 Vは母音、 jは拗音、 Qは促音、 #は無音。)で抽出した音声特徴パターンを統計 処理して標準化したものである。
[0009] 映像検索インデックス生成部 104は、処理対象となる 1つの音声区間に対して、第 1番目の標準音声パターンと入力音声の音声特徴パターンとの類似度を、 DP照合 法や HMM (Hidden Markov Model)等の音声認識処理により計算される。そして、第 1番目の標準音声パターンに対して最も高い類似度を示す区間を「サブワード区間」 として検出する。以下、サブワード区間の類似度を「スコア」という。映像検索インデッ タス生成部 104は、サブワード区間の音素記号、発声区間 (始端時刻、終端時刻)、 及びスコアの組を「映像検索インデックス」として出力する。
[0010] 同様に、第 2番目以降の標準音声パターンについてもサブワード区間を検出し、検 出サブワード区間に関する映像検索インデックスを出力する。
[0011] 当該音声区間において、すべての標準音声パターンに関して映像検索インデック スが生成されたならば、映像検索インデックス生成部 104は、処理対象となる音声区 間を隣接する次の音声区間に移し、同様の処理を実行する。そして、入力音声の全 区間に亘つて映像検索インデックスを作成したところで、処理を終了する。
[0012] 入力音声の音声データと映像検索インデックスは、検索対象音声データとして記録 部 201に記憶される。図 13は記録部 201に記憶された映像検索インデックスのラティ ス構造の一部を示す図である。図 13では、 10msec単位で分割した入力音声の各音 声区間の終端を、その音声区間に対して生成した各映像検索インデックスの終端と し、同一音声区間における映像検索インデックスを生成された順番に配置している。 このような映像検索インデックスのラテイス構造を「音素類似度表」と呼ぶ。尚、「ラティ ス」とは、連続する種々の音声区間に対して、複数の音素や単語の候補とその可能 性を表の形で表したものをいう(非特許文献 1, p. 198参照)。
[0013] クエリー音声を用いて映像シーンを検索する処理は次のように行われる。まず、キ 一ワード入力部 203に検索キーワードであるタエリー音声が入力される。キーワード 変換部 204は、クエリー音声をサブワードの時系列に変換する。次に、キーワードパ ターン照合部 205は、音素類似度表の中から、クエリー音声を構成するサブワードだ けをピックアップする。そして、ピックアップされた複数のラテイス上のサブワードを、検 索キーワードを変換したサブワードの系列順に隙間なく接続する。
[0014] 例えば、クエリー音声としてキーワード入力部 203に「空 (そら)」が入力された場合 、キーワード変換部 204は、サブワードの系列「SO」, 「OR」, 「RA」を生成する。キー ワードパターン照合部 205は、音素類似度表力もサブワード「S0」, 「0R」, 「RA」をピ ックアップして、これを隙間なく接続する。この場合、ある時刻のラテイスカゝらサブヮー ド「RA」を取り出し、サブワード「RA」の始端時刻にあたるラテイス力もその前のサブヮ ード「OR」を取り出し、さらにサブワード「OR」の始端時刻に当たるラテイス力もサブヮ 一 「SO」を取り出す。そして、最後のサブワード「RA」の終端を基準にして「SO」「OR 」「RA」を連結する。
[0015] このようにサブワード(上記例では、「SO」「OR」「RA」)を連結することによって復元 されたキーワードについて、その復元キーワードのスコアの総和を計算する。
[0016] 以下同様に、サブワード「RA」の終端時刻をずらした復元キーワードをすベての時 刻について順次作成し、各復元キーワードについてそのスコアを計算する(図 14参 照)。
[0017] 制御部 202は、スコアが上位となる復元キーワードの先頭サブワードの始端時刻か ら対応する映像信号のタイムコードを算出する。そして、記憶部 201に蓄積された蓄 積映像データ'検索対象音声データの該当部分を再生する制御を行う。
特許文献 1:特開 2000— 236494号公報 (特許第 3252282号公報)
特許文献 2:特開 2005— 91709号公報 非特許文献 1 :古井貞熙, 「音響 ·音声工学」,近代科学社, pp. 194- 210 発明の開示
発明が解決しょうとする課題
[0018] 上記従来の音声検索装置では、音声認識を行うにあたり、音声特徴パターン収納 部 105に格納された標準音声パターンを使用し、クエリー音声と標準音声パターンと の類似度によって音声認識を行う。この場合、認識精度を上げるためには標準音声 ノターンを多く用意する必要がある。しかし、標準音声パターンの数が増えると、類似 度演算の処理時間が増大し又は演算回路の規模が大きくなる。また、標準音声バタ ーンとして登録されて 、な 、クエリー音声が入力された場合には、正常に認識するこ とができないため、音声検索機能が正常に働力ない場合も考えられる。
[0019] また、通常、同じ音素に対する音声であっても男女間で周波数帯域が異なり、また 同性でも個人間で周波数帯域が異なる。従って、標準音声パターンとクエリー音声と の類似度に、これらの差異による影響が現れるため、認識精度に限界がある。
[0020] そこで、本発明の目的は、標準音声パターンを必要とせず、音声の個人差にも影 響されず検索精度の高い音声検索装置を提供することにある。
課題を解決するための手段
[0021] 本発明に係る音声検索装置の第 1の構成は、検索対象音声データ (retrieval voice -data)の中から、クエリー音声データ(qUery voice-data)に一致又は類似する部分音 声テータ (partial voice-data)を検索する音声検索装置 (voice retrieval deviceリでぁ つて、前記検索対象音声データの有声音(voiced sound)のピッチ周期(pitch period) を等化したピッチ等化検索対象音声データ(pitch-equalized retrieval voice-data)の 中から、音声の特徴量空間において、前記クエリー音声データの有声音のピッチ周 期を等化したピッチ等化クエリー音声データに対する距離尺度 (distance measure) ( 又は類似尺度 (likelihood measure) )が所定の閾値以下 (又は所定の閾値以上)であ る部分音声データを検索する部分音声検索手段を備えていることを特徴とする。
[0022] このように、検索対象音声データ及びタエリー音声データのピッチ周期を等化する ことによって、音声帯域の男女差や個人差が除去される。従って、ピッチ周期が等化 された検索対象音声信号及びタエリー音声信号の特徴量空間における距離尺度や 類似尺度は、音声帯域の男女差や個人差にほとんど影響されず、その音声が表す 音素列に依存して定まる。故に、この距離尺度や類似尺度をマッチングの指標として 用いることによって、高い精度で音声検索を行うことが可能となる。
[0023] ここで、「特徴量」とは、音声の発生周波数帯域における短時間スペクトル又はその 対数値、一定時間内での対数エネルギーなどを用いることができる。特徴量として短 時間スペクトルを用いる場合は、例えば、 10〜 30チャンネル程度の帯域フィルタ群 を用いて得られる各帯域の特徴データの時系列、短時間 FFTを用いて直接的に計 算されるスペクトル、ケプストラム変換により得られるケプストラム、相関関数により計 算される相関データ列、 LPC分析を基礎として得られる LPC係数列、 PARCOR係 数、 LSP周波数など力 特徴量として使用される。
[0024] 「距離尺度」とは、特徴量に応じて種々の距離尺度を用いることができる。例えば、 特徴量として短時間スぺ外ルを使用する場合、単純なユークリッド距離、聴覚の感 度を考慮した重み付けを行った距離、判別分析,主成分分析などの統計的分析を行 つて低次元に射影した空間におけるユークリッド距離、マハラビノス距離、板倉'齋藤 距離、 COSH尺度、 WLR尺度 (重みつき尤度比)、 PWLR尺度 (パワー重みつき尤度比) 、 LPCケプストラム間ユークリッド距離、 LPC重みつきケプストラム間ユークリッド距離な どを用いることができる。
[0025] 尚、特徴量 (一般にベクトル量 ) x, yの距離尺度 d(x, y)は、必ずしも数学的な意味 での距離のように三角不等式を満たす必要はない。しかしながら、次式で定義される 対称性と正値性を持つことが望ましぐまた、 d (x、 y)を効率よく計算するアルゴリズム が存在する必要がある。
[0026] [数 1]
(a) 対称性 : d(x, y) = <i(y, x) (1)
(b) 正値性 : ^(x, y) > 0 (x ^ y)
Figure imgf000007_0001
[0027] 「類似尺度」とは、二つの特徴量がどれだけ類似して 、るのかを示す尺度を 、う。例 えば、次式によって定義できる類似度等を用いることができる。ここで、 X, yは特徴量 を表す。
[0028] [数 2] [0029] 本発明に係る音声検索装置の第 2の構成は、前記第 1の構成において、前記クエリ 一音声データの有声音のピッチ周期を等化することにより前記ピッチ等化クエリー音 声データを生成するピッチ周期等化手段と、前記ピッチ等化クエリー音声データを特 徴量の時系列データに変換したデータ(以下「クエリー特徴データ(query feature-da ta)」という。)を生成する特徴データ生成手段と、を備え、前記部分音声検索手段は 、前記ピッチ等化検索対象音声データに含まれる部分音声データのうち、その特徴 量が、前記タエリー特徴データとの間の距離尺度 (又は類似尺度)が所定の閾値以 下 (又は所定の閾値以上)であるものを検索することを特徴とする。
[0030] この構成により、クエリー音声データが入力されると、ピッチ周期等化手段が当該ク エリー音声データの有声音のピッチ周期を等化する。そして、特徴データ生成手段 は、ピッチ周期が等化されたクエリー音声データの特徴量を演算し、クエリー特徴デ ータを生成する。これにより、部分音声検索手段は、ピッチ等化検索対象音声データ の部分音声データとクエリー特徴データとの間の距離尺度 (又は類似尺度)を閾値判 定により抽出する。これにより、クエリー音声データに一致又は類似する音声データ を、検索対象音声データの中から検索することが可能となる。
[0031] 本発明に係る音声検索装置の第 3の構成は、前記第 1又は 2の構成において、前 記部分音声検索手段は、前記ピッチ等化検索対象音声データを特徴量の時系列デ ータに変換した検索対象特徴データの中から、前記クエリー音声データと同じ音素 長分の部分データ (以下「選択特徴データ」という。)を、選択位置を移動させながら 順次選択する部分音声選択手段と、前記各選択特徴データと前記タエリー特徴デー タとの間の距離尺度 (又は類似尺度)を演算する特徴量尺度演算手段と、前記距離 尺度 (又は類似尺度)が所定の閾値以下 (又は所定の閾値以上)の場合、前記選択 特徴データに対応する検索対象音声データ内の位置を出力する一致位置判定手段 と、を備えていることを特徴とする。
[0032] この構成により、検索対象音声データの中から、特徴量空間におけるピッチ等化検 索対象音声データとの (又は類似尺度)が所定の閾値以下 (又は所定の閾値以上) の部分音声データを抽出することが可能となる。
[0033] 部分音声選択手段が「選択位置を移動」させる手順は、特に限定するものではな 、 。例えば、部分音声データの開始位置を検索対象音声データの先頭力 末尾に向 かって逐次移動させる方法や、逆に、部分音声データの終端位置を検索対象音声 データの末尾力 先頭に向力つて逐次移動させる方法などを採ることができる。
[0034] 本発明に係る音声検索装置の第 4の構成は、前記第 3の構成において、前記検索 対象特徴データを記憶する音声記憶手段を備えていることを特徴とする。
[0035] 検索対象音声データを、検索対象特徴データとして、音声記憶手段に予め記憶さ せておくことにより、クエリー音声データに類似する部分音声データを素早く検索する ことが可能となる。
[0036] 本発明に係る音声検索装置の第 5の構成は、前記第 3又は 4の構成において、前 記検索対象音声データの有声音のピッチ周期を等化することにより前記ピッチ等化 検索対象音声データを生成する第 2のピッチ周期等化手段と、前記ピッチ等化検索 対象音声データを特徴量の時系列データに変換することにより、前記検索対象特徴 データを生成する第 2の特徴データ生成手段と、を備えて!/、ることを特徴とする。
[0037] この構成により、音声データベース内の検索対象音声データが有声音のピッチ周 期が等化されていない場合であっても、第 2のピッチ周期等化手段によりピッチ周期 を等化して第 2の特徴データ生成手段により特徴量を算出することによって、ピッチ 周期が等化された検索対象音声データの特徴量を得ることができる。
[0038] 本発明に係る音声検索装置の第 6の構成は、前記第 2又は 5の構成において、前 記ピッチ周期等化手段 (又は第 2のピッチ周期等化手段)は、前記タエリー音声デー タ (又は前記検索対象音声データ)のピッチ周波数の検出を行うピッチ検出手段、前 記ピッチ周波数と所定の基準周波数との差分を演算する残差演算手段、及び、前記 差分が最小となるように、前記タエリー音声データ (又は前記検索対象音声データ) のピッチ周波数を等化する周波数シフタを具備することを特徴とする。
[0039] この構成により、ピッチ周期等化手段 (又は第 2のピッチ周期等化手段)は、クエリー 音声データ (又は前記検索対象音声データ)のピッチ周波数を等化することができる [0040] 本発明に係る音声検索装置の第 7の構成は、前記第 1乃至 6の何れか一の構成に おいて、前記検索対象特徴データ及び前記タエリー特徴データは、それぞれ、前記 ピッチ等化検索対象音声データ及び前記ピッチ等化クエリー音声データを直交変換 して得られるサブバンド ·データの時系列であることを特徴とする。
[0041] このように特徴量としてサブバンドを使用することにより、簡単なフィルタバンクや FF T, DFT等を使用して検索対象特徴データ及び前記タエリー特徴データを高速に求 めることが可能となる。
[0042] 本発明に係る音声検索装置の第 8の構成は、前記第 2又は 5の構成において、前 記クエリー特徴データを、音素区間ごとに平均化し、平均値の時系列データに変換 する第 1の区間分割手段と、前記検索対象特徴データを、音素区間ごとに平均化し 、平均値の時系列データに変換する第 2の区間分割手段と、を備え、前記特徴量尺 度演算手段は、前記第 1及び第 2の区間分割手段が生成する平均値の時系列デー タの間の距離尺度 (又は類似尺度)を演算することを特徴とする。
[0043] このように、音素区間で特徴量を平均化し、その平均値を用いてマッチング判定を 行うことにより、ノイズや揺らぎの影響が低減され、検索精度が向上する。また、各特 徴量は、音素区間ごとに時間的に離散化される。この際に、音声の伸縮の影響が除 去される。従って、マッチング判定は単純な比較計算のみとなり、 DPマッチングのよう に計算量の多い方法を用いる必要がなぐ装置構成の単純化、演算時間の高速ィ匕 が図られる。
[0044] 本発明に係る音声検索装置の第 9の構成は、前記第 1乃至 8の何れか一の構成に おいて、前記タエリー音声データ (又は前記検索対象音声データ)に対して音素ラベ リングを行うことによりクエリー音素列 (又は検索対象音素列)を生成する音素ラベリン グ処理手段と、前記前記選択特徴データに対応する前記検索対象音素列と前記ク エリー音素列との距離尺度 (又は類似尺度)を決定する音素列尺度演算手段と、前 記特徴量尺度演算手段が出力する特徴量の距離尺度 (又は類似尺度)と、前記音 素列尺度演算手段が出力する音素列の距離尺度 (又は類似尺度)との線形和 (以下 「総合距離尺度 (又は総合類似尺度)」という。)を算出する総合尺度演算手段と、を 備え、前記一致位置判定手段は、前記総合距離尺度 (又は総合類似尺度)が所定 の閾値以下 (又は所定の閾値以上)の場合、前記選択特徴データに対応する検索 対象音声データ内の位置を出力することを特徴とする。
[0045] このように、特徴量尺度に加えて音素列尺度をマッチング判定に考慮することにより 、検索精度を高めることができる。
[0046] 本発明に係る音声検索方法は、検索対象音声データの中から、クエリー音声デー タに一致又は類似する部分音声データを検索する音声検索方法であって、前記検 索対象音声データの有声音のピッチ周期を等化したピッチ等化検索対象音声デー タの中から、音声の特徴量空間において、前記タエリー音声データの有声音のピッ チ周期を等化したピッチ等化クエリー音声データに対する距離尺度 (又は類似尺度) が所定の閾値以下 (又は所定の閾値以上)である部分音声データを検索する部分音 声検索ステップを有することを特徴とする。
[0047] 本発明に係る音声検索方法の第 2の構成は、前記第 1の構成において、前記クエリ 一音声データの有声音のピッチ周期を等化することにより前記ピッチ等化クエリー音 声データを生成するピッチ周期等化ステップと、前記ピッチ等化クエリー音声データ を特徴量の時系列データに変換したデータ(以下「クエリー特徴データ」 t 、う。)を生 成する特徴データ生成ステップと、を備え、前記部分音声検索ステップにおいては、 前記ピッチ等化検索対象音声データに含まれる部分音声データのうち、その特徴量 力 前記タエリー特徴データとの間の距離尺度 (又は類似尺度)が所定の閾値以下( 又は所定の閾値以上)であるものを検索することを特徴とする。
[0048] 本発明に係る音声検索方法の第 3の構成は、前記第 1又は 2の構成において、前 記部分音声検索ステップにお!/、ては、前記ピッチ等化検索対象音声データを特徴 量の時系列データに変換した検索対象特徴データの中から、前記タエリー音声デー タと同じ音素長分の部分データ (以下「選択特徴データ」という。)を、選択位置を移 動させながら順次選択する部分音声選択ステップと、前記各選択特徴データと前記 クエリー特徴データとの間の距離尺度 (又は類似尺度)を演算する特徴量尺度演算 ステップと、前記距離尺度 (又は類似尺度)が所定の閾値以下 (又は所定の閾値以 上)の場合、前記選択特徴データに対応する検索対象音声データ内の位置を出力 する一致位置判定ステップと、を有することを特徴とする。 [0049] 本発明に係る音声検索方法の第 4の構成は、前記第 3の構成において、前記検索 対象特徴データを記憶する音声記憶ステップを備えていることを特徴とする。
[0050] 本発明に係る音声検索方法の第 5の構成は、前記第 3又は 4の構成において、前 記検索対象音声データの有声音のピッチ周期を等化することにより前記ピッチ等化 検索対象音声データを生成する第 2のピッチ周期等化ステップと、前記ピッチ等化検 索対象音声データを特徴量の時系列データに変換することにより、前記検索対象特 徴データを生成する第 2の特徴データ生成ステップとを有することを特徴とする。
[0051] 本発明に係る音声検索方法の第 6の構成は、前記第 2又は 5の構成において、前 記ピッチ周期等化ステップ (又は第 2のピッチ周期等化ステップ)においては、前記ク エリー音声データ (又は前記検索対象音声データ)のピッチ周波数の検出を行うピッ チ検出ステップと、前記ピッチ周波数と所定の基準周波数との差分を演算する残差 演算ステップと、前記差分が最小となるように、前記タエリー音声データ (又は前記検 索対象音声データ)のピッチ周波数を等化する周波数シフトステップとを具備するこ とを特徴とする。
[0052] 本発明に係る音声検索方法の第 7の構成は、前記第 1乃至 6の何れか一の構成に おいて、前記検索対象特徴データ及び前記タエリー特徴データは、それぞれ、前記 ピッチ等化検索対象音声データ及び前記ピッチ等化クエリー音声データを直交変換 して得られるサブバンド ·データの時系列であることを特徴とする。
[0053] 本発明に係る音声検索方法の第 8の構成は、前記第 2又は 5の構成において、前 記クエリー特徴データを、音素区間ごとに平均化し、平均値の時系列データに変換 する第 1の区間分割ステップと、前記検索対象特徴データを、音素区間ごとに平均化 し、平均値の時系列データに変換する第 2の区間分割ステップと、を有し、前記特徴 量尺度演算ステップにおいては、前記第 1及び第 2の区間分割ステップにおいて生 成される平均値の時系列データの間の距離尺度 (又は類似尺度)を演算することを 特徴とする。
[0054] 本発明に係る音声検索方法の第 9の構成は、前記第 1乃至 8の何れか一の構成に おいて、前記タエリー音声データ (又は前記検索対象音声データ)に対して音素ラベ リングを行うことによりクエリー音素列 (又は検索対象音素列)を生成する音素ラベリン グステップと、前記選択特徴データに対応する前記検索対象音素列と前記タエリー 音素列との距離尺度 (又は類似尺度)を決定する音素列尺度演算ステップと、前記 特徴量尺度演算ステップにおいて出力される特徴量の距離尺度 (又は類似尺度)と 、前記音素列尺度演算ステップにおいて出力される音素列の距離尺度 (又は類似尺 度)との線形和 (以下「総合距離尺度 (又は総合類似尺度)」と!、う。 )を算出する総合 尺度演算ステップと、を備え、前記一致位置判定ステップにおいては、前記総合距 離尺度 (又は総合類似尺度)が所定の閾値以下 (又は所定の閾値以上)の場合、前 記選択特徴データに対応する検索対象音声データ内の位置を出力することを特徴と する。
[0055] 本発明に係るプログラムは、コンピュータに読み込んで実行することにより、コンビュ ータを前記第 1乃至 8の何れか一の構成の音声検索装置として機能させることを特徴 とする。
発明の効果
[0056] 以上のように、本発明によれば、検索対象音声データ及びタエリー音声データのピ ツチ周期を等化することにより、音声帯域の男女差や個人差が除去した音声データ を用いて、特徴量のマッチングにより音声検索を行うことで、音声帯域の男女差や個 人差にほとんど影響されず、音声検索の精度を向上させることができる。
[0057] また、音素区間ごとにピッチ周期を等化した検索対象音声データ及びタエリー音声 データの特徴量を平均化し、その特徴量の平均値の時間列のマッチング検査によつ て音声検索を行うことで、ノイズや揺らぎの影響が低減されるとともに、音声の伸縮に よる影響が除去される。その結果、音声検索の精度を向上させることができる。
図面の簡単な説明
[0058] [図 1]本発明の実施例 1に係る音声検索装置 1の全体構成を表す図である。
[図 2]図 1の音声符号化器 2の構成を表すブロック図である。
[図 3]図 2のピッチ周期等化手段 10の構成を表すブロック図である。
[図 4]ピッチ検出手段 21及びピッチ平均手段 22における信号処理の概略を説明す る図である。
[図 5]有声音「あ」のフォルマント特性を示す図である。 [図 6]無声音「す」の自己相関及びケプストラム波形並びに周波数特性を示す図であ る。
[図 7]周波数シフタ 23の内部構成を表す図である。
[図 8]周波数シフタ 23の内部構成の他の例を表す図である。
[図 9]図 1の音声復号器 5の構成を表すブロック図である。
[図 10]図 1の部分音声検索手段 6の構成を表すブロック図である。
[図 11]量子化ビット数についての説明図である。
[図 12]特許文献 1に記載の音声検索装置の構成を表す図である。
[図 13]記録部 201に記憶された映像検索インデックスのラテイス構造の一部を示す 図である。
[図 14]各復元キーワードについてそのスコアを計算するために接続されたラテイスの 構造を表す図である。
符号の説明
1 音声検索装置
2 音声符号化器
3 音声記憶手段
4 データ読出手段
5 音声復号器
6 部分音声検索手段
10 ピッチ周期等化手段
11 特徴データ生成手段
12a, 12b 出力切替手段
13 量子化器
14 ピッチ等化波形符号化器
15 差分ビット演算器
16 ピッチ情報符号化器
17 音素ラベリング処理手段
18 リサンプラ アナライザ
抵抗
入力ピッチ検出手段 ピッチ平均手段 周波数シフタ 出力ピッチ検出手段 残差演算手段
PIDコントローラ ピッチ検出手段
BPF
周波数カウンタ
BPF
周波数カウンタ アンプ
コンデンサ
発信器
変調器
BPF
VCO ピッチ等化波形復号器 逆量子化器 シンセサイザ ピッチ情報復号器 ピッチ周波数検出手段 差分器
加算器
周波数シフタ 59 出力切替手段
61 動作切替手段
62 部分音声選択手段
63, 64 区間分割手段
65 特徴量尺度演算手段
66 音素列尺度演算手段
67 総合尺度演算手段
68 一致位置判定手段
発明を実施するための最良の形態
[0060] 以下、本発明を実施するための最良の形態について、図面を参照しながら説明す る。
実施例 1
[0061] 図 1は、本発明の実施例 1に係る音声検索装置の全体構成を表す図である。実施 例 1の音声検索装置 1は、音声符号化器 2、音声記憶手段 3、データ読出手段 4、音 声復号器 5、及び部分音声検索手段 6を備えている。
[0062] 検索対象音声データゃクエリー音声データは、入力音声データとして音声符号ィ匕 器 2に入力される。音声符号化器 2は、入力音声データに対して有声音のピッチ周期 を等化するとともに、特徴量の時系列データ (特徴データ)に変換する。この際、入力 音声データのピッチ周期の情報は特徴データとは分離され、符号化されて符号化ピ ツチデータとして出力される。一方、特徴データは、サブバンド波形として出力される 。またさらに、音声符号化器 2は、特徴データを符号化し、符号化特徴データとして 出力する。また、音声符号化器 2は、入力音声データに対して音素ラベリング処理を 行い、各音素の音素ラベル及び時間区間の情報力 なる音素ラベルデータとして出 力する。
[0063] 音声記憶手段 3は、音声符号化器 2により符号化特徴データ,符号化ピッチデータ ,及び音素ラベルデータに分解され符号化された検索対象音声データを記憶する。 この音声記憶手段 3に記憶された符号化特徴データ及び符号化ピッチデータが、符 号化された検索対象特徴データである。 [0064] データ読出手段 4は、データ選択信号に従って、音声記憶手段 3内の符号化され た検索対象音声データ (符号化特徴データ,符号化ピッチデータ,及び音素ラベル データ)の部分データを読み出す。
[0065] 音声復号器 5は、データ読出手段 4により読み出された符号化特徴データ及び符 号ィ匕ピッチデータを復号し、特徴データ又は出力音声データとして出力する。
[0066] 部分音声検索手段 6は、音声記憶手段 3に蓄積されている符号化された検索対象 音声データから、クエリー音声データに一致又は類似する部分データを検索する。
[0067] 図 2は、図 1の音声符号化器 2の構成を表すブロック図である。音声符号化器 2は、 ピッチ周期等化手段 10、特徴データ生成手段 11、出力切替手段 12a, 12b、量子 化手段 13、ピッチ等化波形符号化器 14、差分ビット演算器 15、ピッチ情報符号ィ匕 器 16、及び音素ラベリング手段 17を備えている。
[0068] ピッチ周期等化手段 10は、入力音声データ X (t)の有声音のピッチ周期を等化す
in
る。ピッチ周期が等化された入力音声データ (以下「ピッチ等化音声データ」という。 ) X (t)は、出力端子 Out_lから出力される。
out
[0069] 特徴データ生成手段 11は、出力端子 Out_lから出力されるピッチ等化音声データ X (t)を特徴量の時系列データに変換する。本実施例にお!ヽては、特徴量として、 out
短時間周波数スペクトルが用いられる。
[0070] 特徴データ生成手段 11は、リサンプラ 18及びアナライザ (変形離散コサイン変換器
(Modified Discrete Cosine Transformer : MDCT) ) 19から構成されている。
[0071] リサンプラ 18は、ピッチ周期等化手段 100の出力端子 Out_lから出力されるピッチ 等化音声データ X (t)の各ピッチ区間について、同一の標本ィ匕数となるように再標
out
本化を行い、完全等化音声データ X (t)として出力する。
eq
[0072] アナライザ 19は、完全等化音声データ X (t)について、一定のピッチ区間数で変
eq
形離散コサイン変換を行い、短時間周波数スペクトル (以下「特徴データ」という。)X ( f)を生成する。すなわち、本実施例においては、特徴データは、短時間周波数スぺ クトルカもなるベクトル量の時系列(式 (4) )として与えられる。
[0073] [数 3]
X(f) = ( l (i) , ¾ (i) , - - - , Xf.„, (i)) (4) [0074] ここで、 tは時刻、 X (t) (i= l, 2, · ··, n)は時刻 tにおける周波数 fのサブバンドの
fi i
短時間スペクトル値を表す。
[0075] 出力切替手段 12aは、部分音声検索手段 6から入力される切替信号に従って、ァ ナライザ 19が生成する特徴データ X(f)の出力先を、部分音声検索手段 6又は音声 記憶手段 3に切り替える。具体的には、入力音声データとして、検索対象音声データ が入力される場合には、特徴データ X(f)の出力先は音声記憶手段 3に切り替えられ る。入力音声データとして、クエリー音声データが入力される場合には、特徴データ X (f)の出力先は部分音声検索手段 6に切り替えられる。
[0076] 量子化器 13は、特徴データ X(f)を所定の量子化曲線に従って量子化する。ピッチ 等化波形符号化器 14は、量子化器 13が出力する特徴データ X(f)を符号ィ匕し、符 号化特徴データとして出力する。この符号化には、ハフマン符号ィ匕法や算術符号ィ匕 法等のエントロピ符号ィ匕法が使用される。
[0077] 差分ビット演算器 15は、ピッチ等化波形符号化器 14が出力する符号ィ匕特徴デー タの符号量から目的ビット数を減算し差分 (以下「差分ビット数」と 、う。 )を出力する。 量子化器 13は、この差分ビット数によって量子化曲線を平行移動させ、符号化特徴 データの符号量が目的ビット数の範囲内となるように調整する。
[0078] ピッチ情報符号化器 16は、ピッチ周期等化手段 10が出力する残差周期信号 Δν 及び基準周期信号 AV を符号化し、符号ィ匕ピッチデータとして出力する。この符 tch pitch
号化には、ハフマン符号ィ匕法や算術符号ィ匕法等のエントロピ符号ィ匕法が使用される
[0079] 音素ラベリング手段 17は、入力音声データを音素区間に区分するとともに、各音素 区間に対して音素ラベリングを行う。そして、音素ラベル及び時間区間の情報力もな る音素ラベルデータとして出力する。
[0080] 出力切替手段 12bは、音素ラベリング処理手段 17が生成する音素ラベルデータの 出力先を、部分音声検索手段 6又は音声記憶手段 3に切り替える。具体的には、入 力音声データとして、検索対象音声データが入力される場合には、音素ラベルデー タの出力先は音声記憶手段 3に切り替えられる。入力音声データとして、クエリー音 声データが入力される場合には、音素ラベルデータの出力先は部分音声検索手段 6 に切り替えられる。
[0081] 図 3は、図 2のピッチ周期等化手段 10の構成を表すブロック図である。ピッチ周期 等化手段 10は、入力ピッチ検出手段 21、ピッチ平均手段 22、周波数シフタ 23、出 力ピッチ検出手段 24、残差演算手段 25、及び PIDコントローラ 26を備えている。
[0082] 入力ピッチ検出手段 21は、入力音声データ X ( から、当該音声信号に含まれる
in
ピッチの基本周波数を検出する。ピッチの基本周波数を検出する方法は、現在まで に種々の方法が考案されている力 本実施例ではその代表的なものを示す。この入 力ピッチ検出手段 21は、ピッチ検出手段 27、バンドパスフィルタ(Band Pass Filter: 以下「BPF」という。) 28、及び周波数カウンタ 29を備えている。
[0083] ピッチ検出手段 27は、入力音声データ X (t)から、ピッチの基本周期 T = 1/f を
in 0 0 検出する。例えば、入力音声データ X (t)が図 4 (a)のような波形であったとする。ピッ
in
チ検出手段 27は、まずこの波形に対して短時間フーリエ変換を行い、図 4 (b)のよう なスペクトル波形 X (f)を導出する。
[0084] 通常、音声波形は、ピッチ以外にも多くの周波数成分を含み、ここで得られるスぺク トル波形は、ピッチの基本周波数及びピッチの高調波成分以外にも、付加的に多く の周波数成分を有する。したがって、このスペクトル波形 X(f)カゝらピッチの基本周波 数 f
0を抽出するのは一般に困難である。そこで、ピッチ検出手段 27は、このスぺタト ル波形 X(f)し対し再度フーリエ変換を行う。これにより、スペクトル波形 X(f)に含ま れるピッチの高調波の間隔 Δ f の逆数 F =
0 0 1Z Δ f の点に鋭 、ピークを持つスぺタト
0
ル波形が得られる(図 4 (c)参照)。ピッチ検出手段 27は、このピークの位置 Fを検出
0 すること〖こよって、ピッチの基本周波数 f = Δ ί
0 0 /2=F
0 Z2を検出する。
[0085] また、ピッチ検出手段 27は、スペクトル波形 X(f)から、入力音声データ X (t)が有
in
声音か無声音かを判別する。有声音の場合には、ノイズフラグ信号 V として 0を出
noise
力する。無声音の場合にはノイズフラグ信号 V として
noise 1を出力する。なお、有声音と 無声音の判別は、スペクトル波形 X(f)の傾き検出によって行われる。図 5は有声音「 あ」のフォルマント特性を示す図であり、図 6は無声音「す」の自己相関及びケプストラ ム波形並びに周波数特性を示す図である。有声音は、図 5のように、スペクトル波形 X(f)は、全体的に低周波側が大きく高周波側に向力つて小さくなるようなフォルマン ト特性を示す。それに対して、無声音は、図 6のように、全体的に高周波側に向かつ て大きくなるような周波数特性を示す。したがって、スペクトル波形 X(f)の全体的な 傾きを検出することによって、入力音声データ X (t)が有声音か無声音かを判別する
in
ことができる。
[0086] 尚、入力音声データ X (t)が無声音の場合、ピッチが存在しないので、ピッチ検出
in
手段 27が出力するピッチの基本周波数 f は無意味な値となる。
0
[0087] BPF28は、通過帯域を外部から設定可能な狭帯域のバンドパスフィルタが使用さ れる。 BPF28は、ピッチ検出手段 27により検出されるピッチの基本周波数 f を通過
0 帯域の中心周波数として設定する(図 4 (d)参照)。そして、 BPF28は、入力音声デ ータ X (t)をフィルタリングし、ピッチの基本周波数 f のほぼ正弦波状の波形を出力 in 0
する(図 4 (e)参照)。
[0088] 周波数カウンタ 29は、 BPF28が出力するほぼ正弦波状の波形のゼロクロス点の時 間間隔をカウントすることにより、ピッチの基本周期 T = l/f
0 0を出力する。この検出 されたピッチの基本周期 Tが入力ピッチ検出手段 21の出力信号 (以下「基本周波数
0
信号」 )として出力される(図 4 (f)参照)。
[0089] ピッチ平均手段 22は、ピッチ検出手段 27が出力するピッチの基本周期信号 Tを
0 平均化するものであり、通常のローパスフィルタ(Low Pass Filter:以下「LPF」という。 )が使用される。ピッチ平均手段 22により、基本周期信号 V が平滑化され、音素内
pitch
では時間的にほぼ一定の信号となる。この平滑化された基本周期が基準周期 T (基 準周波数 f = 1ZT )として使用される(図 4 (g)参照)。
[0090] 周波数シフタ 23は、入力音声データ X (t)のピッチ周波数を基準周波数 f に近づ
in 0 ける方向にシフトさせることにより、音声信号のピッチ周期を等化する。
[0091] 出力ピッチ検出手段 24は、周波数シフタ 23より出力される出力音声データ(以下「 ピッチ等化音声データ」という。) X (t)から、当該ピッチ等化音声データ X (t)に含
out out まれるピッチの基本周期 T 'を検出する。この出力ピッチ検出手段 24も、基本的に入
0
力ピッチ検出手段 21と同様の構成とすることができる。本実施例の場合、出力ピッチ 検出手段 24は、 BPF31及び周波数カウンタ 32を備えている。
[0092] BPF31は、通過帯域を外部から設定可能な狭帯域の BPFが使用される。 BPF31 は、ピッチ検出手段 27により検出されるピッチの基本周波数 f を通過帯域の中心周
0
波数として設定する。そして、 BPF31は、ピッチ等化音声データ X (t)をフィルタリン out
グし、ピッチの基本周波数 f 'のほぼ正弦波状の波形を出力する。周波数カウンタ 32
0
は、 BPF31が出力するほぼ正弦波状の波形のゼロクロス点の時間間隔をカウントす ることにより、ピッチの基本周期 T, = l/f ,を出力する。この検出されたピッチの基
0 0
本周期 T 'が出力ピッチ検出手段 24の出力信号として出力される。
0
[0093] 残差演算手段 25は、出力ピッチ検出手段 24が出力する基本周期 T 'からピッチ平
0
均手段 22が出力する基準周期 Tを引いた残差周期 ΔΤ を出力する。この残差周 s pitch
期 ΔΤ は、 PIDコントローラ 26を介して周波数シフタ 23に入力される。周波数シフ pitch
タ 23は、残差周波数 1Ζ ΔΤ に比例して、入力音声データのピッチ周波数を基準 pitch
周波数 f
0に近づける方向にシフトさせる。
[0094] 尚、 PIDコントローラ 26は、直列接続されたアンプ 34及び抵抗 20、並びに、アンプ 34に対して並列接続されたコンデンサ 36から構成されている。この PIDコントローラ 2 6は、周波数シフタ 23、出力ピッチ検出手段 24、及び残差演算手段 25からなるフィ ードバックループの発振を防止するためのものである。
[0095] 尚、図 3では、 PIDコントローラ 26は、アナログ回路表示しているが、デジタル回路 で構成してもよ ヽ。
[0096] 図 7は周波数シフタ 23の内部構成を表す図である。周波数シフタ 23は、発信器 41 、変調器 42、 BPF43、電圧制御発信器 (Voltage Controlled Oscillator:以下「VCO 」という。)44、及び復調器 45を備えている。
[0097] 発信器 41は、入力音声データ X (t)の周波数変調を行うための一定周波数の変調 in
キャリア信号 Cを出力する。通常、音声信号の帯域は 8kHz程度である(図 7 (i)参照 )。したがって、発信器 41が発生する変調キャリア信号 Cの周波数 (以下「変調キヤリ ァ周波数」という。)としては、通常は 20kHz程度のものが使用される。
[0098] 変調器 42は、発信器 41が出力する変調キャリア信号 Cを入力音声データ x (t)で
1 in 周波数変調し、被変調信号を生成する。この被変調信号は、変調キャリア周波数を 中心として、その両側に音声信号の帯域と同じバンド幅の側波帯 (上側波帯及び下 側波帯)を有する信号である(図 7 (ii)参照)。 [0099] BPF43は、変調キャリア周波数を下限遮断周波数とし、入力音声データの帯域幅 よりも大きいバンド幅の通過域を有する BPFである。これにより、 BPF43力 出力さ れる被変調信号は、上側波帯のみが切り出された信号となる(図 7 (iii)参照)。
[0100] VC044は、発信器 41が出力する変調キャリア信号 Cと同じ周波数の信号を、 PI Dコントローラ 26を介して残差演算手段 25から入力される残差周期 ΔΤ の信号(
pitcn
以下「残差周期信号」という。 ) AV
pitchにより周波数を変調して得られる信号 (以下「復 調キャリア信号」という。)を出力する。
[0101] 復調器 45は、 BPF43が出力する上側波帯のみの被変調信号を、 VC044が出力 する復調キャリア信号により復調し、音声信号を復元する(図 7 (iv)参照)。このとき、 復調キャリア信号は、残差周期信号で変調されている。そのため、被変調信号を復 調する際に、入力音声データ X (t)のピッチ周波数の基準周波数 f力 のずれが消
in s
去される。すなわち、入力音声データ X (t)のピッチ周期は、基準周期 Tに等化され
in s
る。
[0102] 図 8は、周波数シフタ 23の内部構成の他の例を表す図である。図 8においては、図 7の発信器 41と VC044とを入れ替えた構成とされている。この構成によっても、図 7 の場合と同様に、入力音声データ X (t)のピッチ周期は、基準周期 Tに等化すること
in s
ができる。
[0103] 図 9は、図 1の音声復号器 5の構成を表すブロック図である。音声復号器 5は、音声 符号化器 2により符号化された音声信号を復号する装置である。音声復号器 5は、ピ ツチ等化波形復号器 51、逆量子化器 52、シンセサイザ 53、ピッチ情報復号器 54、 ピッチ周波数検出手段 55、差分器 56、加算器 57、周波数シフタ 58、及び出力切替 手段 59を備えている。
[0104] 音声復号器 5には、符号化特徴データ及び符号化ピッチデータが入力される。符 号ィ匕特徴データは、図 2のピッチ等化波形符号化器 14から出力される符号ィ匕特徴 データである。符号ィ匕ピッチデータは、図 2のピッチ情報符号化器 16から出力される 符号ィ匕ピッチデータである。
[0105] ピッチ等化波形復号器 51は、符号化特徴データを復号し、量子化後の各サブバン ドの特徴データ(以下「量子化特徴データ」 t ヽぅ。)を復元する。逆量子化器 52は、 この量子化特徴データを逆量子化し、 n個のサブバンドの特徴データ X(f) = {X(^) , X(f ) , · ··, X(f ) }を復元する。
2 n
[0106] シンセサイザ 53は、特徴データ X(f)を逆変形離散コサイン変換(Inverse Modified Discrete Cosine Transform:以下「IMDCT」という。)し、 1ピッチ区間の時系列デー タ (以下「等化音声信号」という。 ) x (t)を生成する。ピッチ周波数検出手段 55は、こ の等化音声信号 X (t)のピッチ周波数を検出し等化ピッチ周波数信号 V として出力
eq eq
する。
[0107] 一方、ピッチ情報復号器 54は、符号化ピッチデータを復号することにより、基準周 波数信号 AV 及び残差周波数信号 Δν を復元する。差分器 56は、基準周波数
pitch pitcn
信号 AV 力 等化ピッチ周波数信号 V を差し引いた差分を基準周波数変化信号 pitch eq
AAV として出力する。加算器 57は、残差周波数信号 Δν と基準周波数変化 pitch pitch
信号 ΔΑν とを加算してこれを修正残差周波数信号 Δν "として出力する。
pitch pitch
[0108] 周波数シフタ 58は、図 7又は図 8に示した周波数シフタ 23と同様の構成を有する。
この場合、入力端子 Inには等化音声信号 X (t)が入力され、 VC044には修正残差 周波数信号 AV "が入力される。 VC044は発信器 41が出力する変調キャリア信
pitch
号 Cと同じキャリア周波数の信号を、加算器 57から入力される修正残差周波数信号 AV
pitch "により周波数変調して得られる信号 (以下「復調キャリア信号」という。)を出 力するが、この場合、復調キャリア信号の周波数は、キャリア周波数に残差周波数を 加えた周波数となる。
[0109] これにより、周波数シフタ 58において等化音声信号 X (t)の各ピッチ区間のピッチ 周期に揺らぎ成分が加えられ、音声信号 X (t)
res が復元される。
[0110] 出力切替手段 59は、部分音声検索手段 6から入力される切替信号に従って、逆量 子化器 52が生成する特徴データ X(f)の出力先を、シンセサイザ 53又は部分音声検 索手段 6に切り替える。具体的には、部分音声検索動作を行う場合には、特徴デー タ X(f)の出力先は部分音声検索手段 6に切り替えられる。一方、検索対象音声デー タを外部に出力する場合には、特徴データ X(f)の出力先はシンセサイザ 53に切り 替えられる。
[0111] 図 10は、図 1の部分音声検索手段 6の構成を表すブロック図である。部分音声検 索手段 6は、動作切替手段 61、部分音声選択手段 62、区間分割手段 63, 64、特徴 量尺度演算手段 65、音素列尺度演算手段 66、総合尺度演算手段 67、及び一致位 置判定手段 68を備えて 、る。
[0112] 動作切替手段 61は、音声検索装置 1の動作を、音声記憶手段 3に対する検索対 象音声データの入出力動作、又は部分音声検索手段 6による部分音声検索動作に 切り替える切替信号を出力する。
[0113] 部分音声選択手段 62は、音声記憶手段 3に記憶されている検索対象特徴データ( 正確には、符号化された検索対象特徴データ)の中から、部分音声データを選択す るためのデータ選択信号を出力する。このデータ選択信号は、データ読出手段 4に 入力される。データ読出手段 4は、データ選択信号に従って、音声記憶手段 3に記憶 されて!/ヽる検索対象特徴データを選択し読み出す。
[0114] 区間分割手段 63は、音声符号化器 2のアナライザ 19から入力されるクエリー音声 の特徴データ (サブバンド波形)を、音素ラベリング処理手段 17から入力されるクエリ 一音声の音素ラベルデータの時間区間の情報に従って、音素区間ごとに分割する。 そして、それぞれの音素区間ごとに、特徴データを平均化し、平均値の時系列デー タとして特徴量尺度演算手段 65に出力する。
[0115] 区間分割手段 64は、音声復号器 5の逆量子化器 52から入力される検索対象音声 の特徴データ (サブバンド波形)を、データ読出手段 4から入力される検索対象音声 の音素ラベルデータの時間区間の情報に従って、音素区間ごとに分割する。そして
、それぞれの音素区間ごとに、特徴データを平均化し、平均値の時系列データとして 特徴量尺度演算手段 65に出力する。
[0116] 特徴量尺度演算手段 65は、区間分割手段 63, 64から入力される特徴データの間 の距離尺度 D (X , X )を演算する。ここで、距離尺度は、特徴データを構成する各 サブバンド波形の相関係数の線形和として表される。
すなわち、クエリー音声の特徴データを X (f)、検索対象音声の特徴データを X (f) とし、それぞれ式 (5) (6)で表す。
[0117] [数 4] Xfi(f) = (Xqifl(t),Xqih{t),..-,Xil n(t)) (5) x。(f) =
Figure imgf000025_0001
(6)
[0118] 特徴データ X (f) , X (f)の各サブバンド要素の相関係数は式(7)により表される。
ここで、 tは j番目の音素区間を表す。また、 X (t )は、 j番目の音素区間における特
fi ]
徴データ X (t)の時間平均値、 X (t )は、 j番目の音素区間における特徴データ
fi o, fi ]
X (t)を時間平均値である。
o, fi
[0119] [数 5]
R{Xq ,XoJi) = Νσ (Ϊ = 1,2, · - · ,η) J
Figure imgf000025_0002
[0120] 本実施例 lにおいては、特徴データの間の距離尺度 D (X , X )を式(10)により定
q o
義する。
[0121] 園
Figure imgf000025_0003
ここで、 Wは重み係数である。重み係数 Wは、適宜設定される。
[0122] 音素列尺度演算手段 66は、音声符号化器 2の音素ラベリング処理手段 17からタエ リー音声の音素ラベルデータが入力されるとともに、データ読出手段 4から検索対象 音声の音素ラベルデータが入力される。音素列尺度演算手段 66は、これらの音素ラ ベルデータの距離尺度 Dを所定の音素間距離尺度表を用いて演算する。ここで、
2
音素間距離尺度表とは、すべての 2つの音素の組み合わせに対して 2つの音素間の 距離尺度をテーブルとして表したものである。
[0123] 総合尺度演算手段 67は、特徴量尺度演算手段 65が算出する特徴データの間の 距離尺度 D (X , X )と音素列尺度演算手段 66が算出する音素ラベルデータの距
1 q ο
離尺度 Dの線形和をとることによって、総合距離尺度 Dを演算する。すなわち、総合 距離尺度 Dは、式(11)により表される。
[0124] [数 7]
D = W.D^X^ X,,) + W2D2 (11) ここで、 W , Wは重み係数であり、適宜決められる。
1 2
[0125] 一致位置判定手段 68は、距離尺度 Dが所定の閾値 D 以下であるか否かを判定し
th
、D≤D の場合には、当該部分データを選択するデータ選択信号を出力する。
th
[0126] 以上のように構成された本実施例の音声検索装置 1につ 、て、以下その動作を説 明する。
[0127] 〔1〕検索対象音声データの蓄積動作
まず、検索対象音声データを音声記憶手段 3に蓄積する際の動作について説明す る。この場合、部分音声検索手段 6の動作切替手段 61は、切替信号として検索対象 音声データの入出力動作を表すレベル (例えば Hレベル)を出力する。これにより、 音声符号化器 2の出力切替手段 12aは、アナライザ 19が生成する特徴データ X (f) を量子化器 13に出力する。音声符号化器 2の出力切替手段 12bは、音素ラベリング 処理手段 17が生成する音素ラベルデータを音声記憶手段 3に出力する。また、音声 復号器 5の出力切替手段 59は、逆量子化器 52が生成する特徴データ X (f)をシンセ サイザ 53に出力する。
[0128] まず、検索対象音声データとして入力音声データ X (t)が音声符号化器 2へ入力さ in
れると、ピッチ周期等化手段 10の入力ピッチ検出手段 21は、入力音声データ X (t)
in が有声音か無声音かを判別してノイズフラグ信号 V を出力端子 OUT_4へ出力する noise
とともに、入力音声データ X (t)からピッチ周波数を検出し、基本周波数信号 V を in pitch ピッチ平均手段 22に出力する。ピッチ平均手段 22は、基本周波数信号 V を平均 pitch 化し (この場合、 LPFを使用するので加重平均となる。)、これを基準周波数信号 AV
P
として出力する。この基準周波数信号 AV は、出力端子 OUT_3から出力されると itch pitch
ともに、残差演算手段 25に入力される。
[0129] 一方、周波数シフタ 23は、入力音声データ X (t)の周波数をシフトさせ、ピッチ等 in
化音声データ X (t)として出力端子 Out_lへ出力する。初期状態においては、残差 out
周波数信号 A V は 0 (リセット状態)であり、周波数シフタ 23は、入力音声データ X
pitch in (t)がそのままピッチ等化音声データ x (t)として出力端子 Out_lへ出力される。
out
[0130] 次に、出力ピッチ検出手段 24は、周波数シフタ 23が出力する出力音声データのピ ツチ周波数 f 'を検出する。検出されたピッチ周波数 f 'は、ピッチ周波数信号 V '
0 0 pitch として残差演算手段 25に入力される。
[0131] 残差演算手段 25は、ピッチ周波数信号 V ,から基準周波数信号 AV を差し引 pitch pitch くことにより、残差周波数信号 Δν を生成する。この残差周波数信号 Δν は、出 pit en pitch 力端子 Out_2へ出力されるとともに、 PIDコントローラ 26を介して周波数シフタ 23へ 入力される。
[0132] 周波数シフタ 23は、 PIDコントローラ 26を介して入力される残差周波数信号 Δν pitc に比例して、周波数のシフト量を設定する。この場合、残差周波数信号 Δν が正 h pitch 値であれば、残差周波数信号 Δ V に比例した量だけ周波数を下げるようにシフト pitch
量が設定される。残差周波数信号 Δν 力 S負値であれば、残差周波数信号 Δν pitch pitch に比例した量だけ周波数を上げるようにシフト量が設定される。
[0133] このようなフィードバック制御により、入力音声データ X (t)のピッチ周期は、常に基 in
準周期 lZf sに維持され、ピッチ等化音声データ X (t)の
out ピッチ周期は等化される。
[0134] このように、ピッチ周期等化手段 10において、入力音声データ X (t)に含まれる情 in
報は、
(a)有声音か無声音かを示す情報;
(b) 1ピッチ区間の音声波形を表す情報;
(c)基準ピッチ周波数の情報;
(d)各ピッチ区間のピッチ周波数の基準ピッチ周波数力 の偏倚量を表す残差周波 数情報;
に分離される。(a)〜 (d)の情報は、それぞれ、ノイズフラグ信号 V 、ピッチ周期が noise
基準周期 lZf s (入力音声データの過去のピッチ周波数の加重平均の逆数)に等化 されたピッチ等化音声データ X (t)、基準周波数信号 AV 、及び残差周波数信号 out pitcn
Δν として出力される。ノイズフラグ信号 V は出力端子 Out_4から出力され、ピッ pitch noise
チ等化音声データ χ (t)は出力端子 Out_lから出力され、基準周波数信号 AV は out pitch 出力端子 Out_3から出力され、残差周波数信号 Δν は出力端子 Out_2から出力さ pitch れる。
[0135] ピッチ等化音声データ x (t)は、男女差、個人差、音素、感情及び会話内容によ out
つて変化するピッチ周波数のジッタ成分や変化成分が除去された音声信号であり、 抑揚のない平坦的'機械的な音声信号である。したがって、同じ有声音のピッチ等化 音声データ X (t)は、男女差、個人差、音素、感情又は会話内容に無関係にほぼ out
同じ波形が得られるため、ピッチ等化音声データ X (t)を比較することによって有声 out
音についてのマッチングを精度よく行うことが可能となる。
[0136] また、有声音のピッチ等化音声データ X (t)はピッチ周期が基準周期 lZf に等化 out S されているので、一定数のピッチ区間でサブバンド符号ィ匕を行うことにより、ピッチ等 化音声データ X (t)の周波数スペクトル X (f)は、基準周波数の高調波成分のサ out out
ブバンド成分に集約される。音声はピッチ間の波形相関が大きいので、各サブバンド 成分のスペクトル強度の時間変化は緩やかである。したがって、各サブバンド成分を 符号化し、その他の雑音成分を省略することにより、高効率の符号化が可能となる。 また、基準周波数信号 AV 、及び残差周波数信号
pitch Δν は、音声の性質上、同 pitch
一音素内で狭レンジでしか変動しないため、高効率の符号ィ匕が可能である。したが つて、全体として入力音声データ X (t)の有声音成分を高効率で符号ィ匕することが可 能となる。
[0137] 次に、リサンプラ 18は、各ピッチ区間において、基準周波数信号 AV を一定のリ pitch サンプリング数 nで除算することによりリサンプリング周期を計算する。そして、ピッチ 等化音声データ X (t)をそのリサンプリング周期によりリサンプリングし、等標本数音 out
声データ X (t)として出力する。これにより、ピッチ等化音声データ X (t)の 1ピッチ eq out
区間の標本ィ匕数が一定の値とされる。
[0138] 次に、アナライザ 19は、等標本数音声データ X (t)を、一定のピッチ区間数のサブ eq
フレームに区分する。そして、サブフレーム毎に変形離散コサイン変換を行うことによ つて周波数スペクトル信号 X(f)を生成する。
[0139] ここで、 1つのサブフレームの長さは、 1ピッチ周期の整数倍とされる。本実施例で は、サブフレームの長さは 1ピッチ周期(標本ィ匕数 n)とする。従って、 n個の周波数ス ベクトル信号 {X(f ) , X(f ) , · ··, X(f ) }が出力される。周波数 f は基準周波数の第 1高調波、周波数 f は基準周波数の第 2高調波、周波数 f は基準周波数の第 n高調
2 n
波である。
[0140] このように、 1ピッチ周期の整数倍のサブフレームに分割して各サブフレームを直交 変換することによりサブバンド符号ィ匕を行うことで、音声波形データの周波数スぺタト ル信号は基準周波数の高調波のスペクトルに集約される。そして、音声の性質上、 同一の音素内における連続するピッチ区間の波形は類似する、従って、隣接するサ ブフレーム間で基準周波数の高調波成分のスペクトルは類似する、従って、符号ィ匕 効率は高められる。
[0141] 次に、量子化器 13は、周波数スペクトル信号 X(f)を量子化する。ここで、量子化器 13はノイズフラグ信号 V を参照し、ノイズフラグ信号 V 力 SO (有声音)の場合と 1 (
noise noise
無声音)の場合とで量子化曲線を切り換える。
[0142] ノイズフラグ信号 V が 0 (有声音)の場合、量子化曲線は、図 2 (a)に示したように
noise
、周波数が高くなるに従って量子化ビット数が減少するような量子化曲線とされる。こ れは、有声音の周波数特性は、図 5に示したように低周波数域で大きく高周波域に なるに従って減少する特性を有することに対応させたものである。
[0143] 一方、ノイズフラグ信号 V が 1 (無声音)の場合、量子化曲線は、図 2 (b)に示した
noise
ように、周波数が高くなるに従って量子化ビット数が増加するような量子化曲線とされ る。これは、無声音の周波数特性は、図 6に示したように高周波域になるに従って増 加する特性を有することに対応させたものである。
[0144] この量子化曲線の切り換えにより、有声音か無声音かに対応して最適な量子化曲 線が選択される。
[0145] 尚、補足として、量子化ビット数について説明する。量子化器 13による量子化のデ ータフォーマットは図 11 (a) (b)に示したように、小数点以下の実数部(FL)及び 2の 冪乗を表す指数部 (EXP)によって表現される。但し、 0以外の数を表す場合におい て、実数部 (FL)の先頭の 1ビットは必ず 1であるように指数部 (EXP)が調整されるも のとする。
[0146] 例えば、実数部 (FL)が 4ビット、指数部 (EXP)が 2ビットの場合にぉ 、て、 4ビット で量子化する場合、及び 2ビットで量子化する場合は、次のようになる(図 11 (c) , (d )参照)。
[0147] (1)4ビットで量子化する場合
(例 1) X(f)=8=[1000] (但し、 [ ] は 2進数表記を表す。)は、
2 2
FL=[1000], EXP=[100]
2 2
(例 2) X(f)=7=[0100] は、
2
FL=[1110], EXP=[011]
2 2
(例 3) X(f)=3=[1000] は、
2
FL=[1100], EXP=[010]
2 2
[0148] (2) 2ビットで量子化する場合
(例 1) X(f)=8=[1000] は、
2
FL=[1000], EXP=[100]
2 2
(例 2) X(f)=7=[0100] は、
2
FL=[1100] , EXP=[011]
2 2
(例 3) X(f)=3=[1000] は、
2
FL=[1100], EXP=[010]
2 2
[0149] すなわち、 nビットで量子化する場合は、実数部 (FL)の先頭カゝら nビットを残し、残 りのビットは 0とするものとする(図 11 (d)参照)。
[0150] 次に、ピッチ等化波形符号化器 14は、量子化器 13が出力する量子化された周波 数スペクトル信号 X(f)をエントロピ符号化法により符号化し、符号化特徴データを出 力する。また、ピッチ等化波形符号化器 14は、符号化特徴データの符号量 (ビット数 )を差分ビット演算器 15に出力する。差分ビット演算器 15は、符号化特徴データの符 号量から所定の目的ビット数を減算し、差分ビット数を出力する。量子化器 13は、差 分ビット数に応じて、有声音に対する量子化曲線を平行移動的に上下させる。
[0151] 例えば、 {f , f , f , f , f , f }に対する量子化曲線が {6, 5, 4, 3, 2, 1}であった
1 2 3 4 5 6
とし、差分ビット数として 2が入力されたとすると、量子化器 13は、量子化曲線を下方 に 2だけ平行移動する。その結果、量子化曲線は {4, 3, 2, 1, 0, 0}となる。また、差 分ビット数として— 2が入力されたとすると、量子化器 13は、量子化曲線を上方に 2だ け平行移動する。その結果、量子化曲線は {8, 7, 6, 5, 4, 3}となる。 [0152] このように有声音の量子化曲線を上下に変化させることによって、各サブフレーム の符号化特徴データの符号量が目的ビット数程度に調整される。
[0153] 一方、これに並行して、ピッチ情報符号化器 16は、基準周波数信号 AV 及び残
pitch 差周波数信号 Δν を符号化する。
pitch
[0154] 一方、音素ラベリング処理手段 17は、入力音声データ X (t)を音素区間に区分し、 m
各音素区間に対して音素ラベリングを行う。音素区間の分割方法や音素ラベリングの 方法に関しては、音声認識の分野において多くの技術が公知であり、ここではそれら 公知の方法を用いることができる。音素ラベリング処理手段 17は、音素ラベリングに より得られた音素ラベルと各音素ラベルに対する時間区間を表す音素区間の情報を 、音素ラベルデータとして出力する。
[0155] 以上のようにして生成された、符号化特徴データ,符号化ピッチデータ,及び音素 ラベルデータは、音声記憶手段 3に出力され、保存される。
[0156] 〔2〕音声復号器の動作
データ読出手段 4が、音声記憶手段 3から符号化特徴データ及び符号化ピッチデ ータを読み出すと、これらのデータは音声復号器 5に入力される。
[0157] 音声復号器 5のピッチ等化波形復号器 51は、符号化特徴データを復号し、量子化 後の各サブバンドの周波数スペクトル信号 (以下「量子化周波数スペクトル信号」と!ヽ う。)を復元する。逆量子化器 52は、この量子化周波数スペクトル信号を逆量子化し 、 n個のサブバンドの周波数スペクトル信号 X (f) = {X (f ) , X (f ) , · · ·, X (f ) }を復
1 2 n 元する。
[0158] シンセサイザ 53は、周波数スペクトル信号 X(f)を逆変形離散コサイン変換 (Inverse
Modified Discrete Cosine Transform :以下「IMDCT」という。)し、 1ピッチ区間の時 系列データ (以下「等化音声信号」という。 ) x (t)を生成する。ピッチ周波数検出手 段 55は、この等化音声信号 X (t)のピッチ周波数を検出し等化ピッチ周波数信号 V
eq e として出力する。
[0159] 一方、ピッチ情報復号器 54は、符号ィ匕ピッチデータを復号することにより、基準周 波数信号 AV 及び残差周波数信号 Δν を復元する。差分器 56は、基準周波数
pitch pitcn
信号 AV 力 等化ピッチ周波数信号 V を差し引いた差分を基準周波数変化信号 pitch eq AAV として出力する。加算器 57は、残差周波数信号 Δν と基準周波数変化 pitch pitch
信号 ΔΑν とを加算してこれを修正残差周波数信号 Δν "として出力する。
pitch pitch
[0160] 周波数シフタ 58は、図 7又は図 8に示した周波数シフタ 23と同様の構成を有する。
この場合、入力端子 Inには等化音声信号 X (t)が入力され、 VC044には修正残差 周波数信号 AV "が入力される。 VC044は発信器 41が出力する変調キャリア信
pitch
号 Cと同じキャリア周波数の信号を、加算器 57から入力される修正残差周波数信号 AV "により周波数変調して得られる信号 (以下「復調キャリア信号」という。)を出 力するが、この場合、復調キャリア信号の周波数は、キャリア周波数に残差周波数を 加えた周波数となる。
[0161] これにより、周波数シフタ 58において等化音声信号 X (t)の各ピッチ区間のピッチ 周期に揺らぎ成分が加えられ、音声信号 X (t)
res が復元される。
[0162] 〔3〕クエリー音声データによる部分音声データの検索動作
次に、クエリー音声データによる部分音声データの検索動作について説明する。こ の場合、部分音声検索手段 6の動作切替手段 61は、切替信号として部分音声検索 動作を表すレベル (例えば Lレベル)を出力する。これにより、音声符号化器 2の出力 切替手段 12aは、アナライザ 19が生成する特徴データ X(f)を部分音声検索手段 6 に出力する。音声符号化器 2の出力切替手段 12bは、音素ラベリング処理手段 17が 生成する音素ラベルデータを部分音声検索手段 6に出力する。また、音声復号器 5 の出力切替手段 59は、逆量子化器 52が生成する特徴データ X(f)を部分音声検索 手段 6に出力する。
[0163] まず、クエリー音声データは、入力音声データ X (t)として音声符号化器 2に入力さ れる。
[0164] ピッチ周期等化手段 1では、上述のように、入力音声データ X (t)の有声音のピッ チ周期を等化し、ピッチ等化音声データ X (t)として出力端子 Out_lから出力する。
out
また、特徴データ生成手段 19は、上述のように、ピッチ等化音声データ X (t)を短 out 時間スペクトルの時系列カゝらなる特徴データ X(f)に変換する。特徴データ X(f)は、 出力切替手段 12aを介して部分音声検索手段 6へ出力される。
[0165] 一方、音素ラベリング処理手段 17では、上述のように、入力音声データ X (t)を音 素区間に区分し、各音素区間に対して音素ラベリングを行う。そして、音素ラベルと 音素区間の情報を、音素ラベルデータとして出力する。
[0166] 次に、部分音声検索手段 6の部分音声選択手段 62は、音声記憶手段 3に記憶さ れた符号化特徴データ,符号化ピッチデータ,及び音素ラベルデータを、データの 先頭から順に順次読み出すためのデータ選択信号を出力する。このとき、読み出す 部分データの長さは、クエリー音声データと同じ音素長の長さとされる。データ読出 手段 4は、データ選択信号に従って、音声記憶手段 3から部分データを読み出す。
[0167] データ読出手段 4により読み出された音素ラベルデータは、部分音声検索手段 6に 入力される。
[0168] 一方、データ読出手段 4により読み出された符号化特徴データ及び符号化ピッチ データの部分データは、音声復号器 5に入力される。音声復号器 5では、上述のよう に、ピッチ等化波形復号器 51で符号化特徴データを復号し、逆量子化器 52で逆量 子化を行うことにより、特徴データを生成し、部分音声検索手段 6に出力する。
[0169] 以下、音声復号器 5から部分音声検索手段 6に入力される検索対象特徴データの 部分データを「選択特徴データ」と呼ぶ。
[0170] 部分音声検索手段 6においては、音声符号化器 2からタエリー音声の特徴データ( 以下「クエリー特徴データ」という。)及び音素ラベルデータが入力されると、区間分割 手段 63は、クエリー特徴データを音素区間ごとに平均化し、平均値の時系列データ に変換する。この場合、音素ラベルデータに含まれる音素区間の情報に基づき、タエ リー特徴データを時間区間に区分し、各時間区間で平均値をとればよい。この平均 値の時系列データは、特徴量尺度演算手段 65に入力される。
[0171] また、音声復号器 5及びデータ読出手段 4から選択特徴データ及び音素ラベルデ ータが入力されると、区間分割手段 64は、選択特徴データを音素区間ごとに平均化 し、平均値の時系列データに変換する。この平均値の時系列データは、特徴量尺度 演算手段 65に入力される。
[0172] 特徴量尺度演算手段 65は、区間分割手段 63及び区間分割手段 64から入力され る平均値の時系列データの間の距離尺度 D (X , X )を式(10)に従って算出する。
[0173] 一方、音素列尺度演算手段 66は、音声符号化器 2から入力されるクエリー音声の 音素ラベルデータとデータ読出手段から入力される検索対象音声の音素ラベルデー タとの間の距離尺度 Dを音素間距離尺度表を用いて演算する。
2
[0174] 総合尺度演算手段 67は、特徴量尺度演算手段 65が算出する特徴データの間の 距離尺度 D (X , X )と音素列尺度演算手段 66が算出する音素ラベルデータの距
1 q ο
離尺度 Dの線形和をとることによって、総合距離尺度 Dを式(11)により演算する。
2
[0175] 一致位置判定手段 68は、距離尺度 Dが所定の閾値 D 以下であるか否かを判定し th
、 D≤D の場合には、当該部分データを選択するデータ選択信号を出力する。そし th
て、動作切替手段 61は、切替信号として部分音声検索動作を表すレベル (例えば L レベル)を出力する。
[0176] これにより、検索された検索対象データの部分データが、出力音声データとして出 力される。
[0177] 尚、本実施例は、音声情報と映像とがー体として記録されたマルチメディア 'データ ベースにおける情報の検索においても適用することができる。
産業上の利用可能性
[0178] 本発明は、音声データベースや音声情報を含むマルチメディア 'データベース等に おいて利用可能である。

Claims

請求の範囲
[1] 検索対象音声データの中から、クエリー音声データに一致又は類似する部分音声デ ータを検索する音声検索装置であって、
前記検索対象音声データの有声音のピッチ周期を等化したピッチ等化検索対象 音声データの中から、音声の特徴量空間において、前記タエリー音声データの有声 音のピッチ周期を等化したピッチ等化クエリー音声データに対する距離尺度 (又は類 似尺度)が所定の閾値以下 (又は所定の閾値以上)である部分音声データを検索す る部分音声検索手段
を備えて 、ることを特徴とする音声検索装置。
[2] 前記タエリー音声データの有声音のピッチ周期を等化することにより前記ピッチ等化 クエリー音声データを生成するピッチ周期等化手段と、
前記ピッチ等化クエリー音声データを特徴量の時系列データに変換したデータ(以 下「クエリー特徴データ」と 、う。 )を生成する特徴データ生成手段と、
を備え、
前記部分音声検索手段は、前記ピッチ等化検索対象音声データに含まれる部分 音声データのうち、その特徴量が、前記クエリー特徴データとの間の距離尺度 (又は 類似尺度)が所定の閾値以下 (又は所定の閾値以上)であるものを検索すること を特徴とする請求項 1記載の音声検索装置。
[3] 前記部分音声検索手段は、
前記ピッチ等化検索対象音声データを特徴量の時系列データに変換した検索対 象特徴データの中から、前記クエリー音声データと同じ音素長分の部分データ(以下 「選択特徴データ」という。)を、選択位置を移動させながら順次選択する部分音声選 択手段と、
前記各選択特徴データと前記タエリー特徴データとの間の距離尺度 (又は類似尺 度)を演算する特徴量尺度演算手段と、
前記距離尺度 (又は類似尺度)が所定の閾値以下 (又は所定の閾値以上)の場合 、前記選択特徴データに対応する検索対象音声データ内の位置を出力する一致位 置判定手段と、 を備えていることを特徴とする請求項 1又は 2記載の音声検索装置。
[4] 前記検索対象特徴データを記憶する音声記憶手段
を備えて!/、ることを特徴とする請求項 3記載の音声検索装置。
[5] 前記検索対象音声データの有声音のピッチ周期を等化することにより前記ピッチ等 化検索対象音声データを生成する第 2のピッチ周期等化手段と、
前記ピッチ等化検索対象音声データを特徴量の時系列データに変換することによ り、前記検索対象特徴データを生成する第 2の特徴データ生成手段と、
を備えていることを特徴とする請求項 3又は 4記載の音声検索装置。
[6] 前記ピッチ周期等化手段 (又は第 2のピッチ周期等化手段)は、
前記タエリー音声データ (又は前記検索対象音声データ)のピッチ周波数の検出を 行うピッチ検出手段、
前記ピッチ周波数と所定の基準周波数との差分を演算する残差演算手段、 及び、前記差分が最小となるように、前記タエリー音声データ (又は前記検索対象 音声データ)のピッチ周波数を等化する周波数シフタ
を具備することを特徴とする請求項 2又は 5記載の音声検索装置。
[7] 前記検索対象特徴データ及び前記タエリー特徴データは、それぞれ、前記ピッチ等 化検索対象音声データ及び前記ピッチ等化クエリー音声データを直交変換して得ら れるサブバンド'データの時系列であることを特徴とする請求項 1乃至 6の何れか一 記載の音声検索装置。
[8] 前記タエリー特徴データを、音素区間ごとに平均化し、平均値の時系列データに変 換する第 1の区間分割手段と、
前記検索対象特徴データを、音素区間ごとに平均化し、平均値の時系列データに 変換する第 2の区間分割手段と、
を備え、
前記特徴量尺度演算手段は、前記第 1及び第 2の区間分割手段が生成する平均 値の時系列データの間の距離尺度 (又は類似尺度)を演算すること
を特徴とする請求項 2又は 5記載の音声検索装置。
[9] 前記タエリー音声データ (又は前記検索対象音声データ)に対して音素ラベリングを 行うことによりクエリー音素列 (又は検索対象音素列)を生成する音素ラベリング処理 手段と、
前記前記選択特徴データに対応する前記検索対象音素列と前記タエリー音素列と の距離尺度 (又は類似尺度)を決定する音素列尺度演算手段と、
前記特徴量尺度演算手段が出力する特徴量の距離尺度 (又は類似尺度)と、前記 音素列尺度演算手段が出力する音素列の距離尺度 (又は類似尺度)との線形和 (以 下「総合距離尺度 (又は総合類似尺度)」 t 、う。 )を算出する総合尺度演算手段と、 を備え、
前記一致位置判定手段は、前記総合距離尺度 (又は総合類似尺度)が所定の閾 値以下 (又は所定の閾値以上)の場合、前記選択特徴データに対応する検索対象 音声データ内の位置を出力すること
を特徴とする請求項 1乃至 8の何れか一記載の音声検索装置。
[10] 検索対象音声データの中から、クエリー音声データに一致又は類似する部分音声デ ータを検索する音声検索方法であって、
前記検索対象音声データの有声音のピッチ周期を等化したピッチ等化検索対象 音声データの中から、音声の特徴量空間において、前記タエリー音声データの有声 音のピッチ周期を等化したピッチ等化クエリー音声データに対する距離尺度 (又は類 似尺度)が所定の閾値以下 (又は所定の閾値以上)である部分音声データを検索す る部分音声検索ステップ
を有することを特徴とする音声検索方法。
[11] 前記タエリー音声データの有声音のピッチ周期を等化することにより前記ピッチ等化 クエリー音声データを生成するピッチ周期等化ステップと、
前記ピッチ等化クエリー音声データを特徴量の時系列データに変換したデータ(以 下「クエリー特徴データ」と 、う。)を生成する特徴データ生成ステップと、
を備え、
前記部分音声検索ステップにおいては、前記ピッチ等化検索対象音声データに含 まれる部分音声データのうち、その特徴量が、前記クエリー特徴データとの間の距離 尺度 (又は類似尺度)が所定の閾値以下 (又は所定の閾値以上)であるものを検索す ること
を特徴とする請求項 9記載の音声検索方法。
[12] 前記部分音声検索ステップにおいては、
前記ピッチ等化検索対象音声データを特徴量の時系列データに変換した検索対 象特徴データの中から、前記クエリー音声データと同じ音素長分の部分データ(以下 「選択特徴データ」という。)を、選択位置を移動させながら順次選択する部分音声選 択ステップと、
前記各選択特徴データと前記タエリー特徴データとの間の距離尺度 (又は類似尺 度)を演算する特徴量尺度演算ステップと、
前記距離尺度 (又は類似尺度)が所定の閾値以下 (又は所定の閾値以上)の場合 、前記選択特徴データに対応する検索対象音声データ内の位置を出力する一致位 置判定ステップと、
を有することを特徴とする請求項 9又は 10記載の音声検索方法。
[13] 前記検索対象特徴データを記憶する音声記憶ステップ
を備えて 、ることを特徴とする請求項 11記載の音声検索方法。
[14] 前記検索対象音声データの有声音のピッチ周期を等化することにより前記ピッチ等 化検索対象音声データを生成する第 2のピッチ周期等化ステップと、
前記ピッチ等化検索対象音声データを特徴量の時系列データに変換することにより 、前記検索対象特徴データを生成する第 2の特徴データ生成ステップと、 を有することを特徴とする請求項 11又は 12記載の音声検索方法。
[15] 前記ピッチ周期等化ステップ (又は第 2のピッチ周期等化ステップ)にお 、ては、 前記タエリー音声データ (又は前記検索対象音声データ)のピッチ周波数の検出を 行うピッチ検出ステップと、
前記ピッチ周波数と所定の基準周波数との差分を演算する残差演算ステップと、 前記差分が最小となるように、前記タエリー音声データ (又は前記検索対象音声デ ータ)のピッチ周波数を等化する周波数シフトステップと
を具備することを特徴とする請求項 11又は 14記載の音声検索方法。
[16] 前記検索対象特徴データ及び前記タエリー特徴データは、それぞれ、前記ピッチ等 化検索対象音声データ及び前記ピッチ等化クエリー音声データを直交変換して得ら れるサブバンド'データの時系列であることを特徴とする請求項 10乃至 15の何れか 一記載の音声検索方法。
[17] 前記タエリー特徴データを、音素区間ごとに平均化し、平均値の時系列データに変 換する第 1の区間分割ステップと、
前記検索対象特徴データを、音素区間ごとに平均化し、平均値の時系列データに 変換する第 2の区間分割ステップと、
を有し、
前記特徴量尺度演算ステップにおいては、前記第 1及び第 2の区間分割ステップ において生成される平均値の時系列データの間の距離尺度 (又は類似尺度)を演算 すること
を特徴とする請求項 11又は 14記載の音声検索方法。
[18] 前記タエリー音声データ (又は前記検索対象音声データ)に対して音素ラベリングを 行うことによりクエリー音素列 (又は検索対象音素列)を生成する音素ラベリングステツ プと、
前記選択特徴データに対応する前記検索対象音素列と前記タエリー音素列との距 離尺度 (又は類似尺度)を決定する音素列尺度演算ステップと、
前記特徴量尺度演算ステップにおいて出力される特徴量の距離尺度 (又は類似尺 度)と、前記音素列尺度演算ステップにおいて出力される音素列の距離尺度 (又は 類似尺度)との線形和 (以下「総合距離尺度 (又は総合類似尺度)」と!ヽぅ。 )を算出す る総合尺度演算ステップと、
を備え、
前記一致位置判定ステップにお!/ヽては、前記総合距離尺度 (又は総合類似尺度) が所定の閾値以下 (又は所定の閾値以上)の場合、前記選択特徴データに対応する 検索対象音声データ内の位置を出力すること
を特徴とする請求項 10乃至 17の何れか一記載の音声検索方法。
[19] コンピュータに読み込んで実行することにより、コンピュータを請求項 1乃至 8の何れ か一の音声検索装置として機能させることを特徴とするプログラム。
PCT/JP2006/315228 2005-08-01 2006-08-01 音声検索装置及び音声検索方法 Ceased WO2007015489A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2007529275A JP4961565B2 (ja) 2005-08-01 2006-08-01 音声検索装置及び音声検索方法

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2005223155 2005-08-01
JP2005-223155 2005-08-01

Publications (1)

Publication Number Publication Date
WO2007015489A1 true WO2007015489A1 (ja) 2007-02-08

Family

ID=37708770

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2006/315228 Ceased WO2007015489A1 (ja) 2005-08-01 2006-08-01 音声検索装置及び音声検索方法

Country Status (2)

Country Link
JP (1) JP4961565B2 (ja)
WO (1) WO2007015489A1 (ja)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2012242542A (ja) * 2011-05-18 2012-12-10 Nippon Hoso Kyokai <Nhk> 音声比較装置及び音声比較プログラム
JP2019074580A (ja) * 2017-10-13 2019-05-16 Kddi株式会社 音声認識方法、装置およびプログラム
CN111145728A (zh) * 2019-12-05 2020-05-12 厦门快商通科技股份有限公司 语音识别模型训练方法、系统、移动终端及存储介质
US11069373B2 (en) 2017-09-25 2021-07-20 Fujitsu Limited Speech processing method, speech processing apparatus, and non-transitory computer-readable storage medium for storing speech processing computer program
CN114974271A (zh) * 2021-12-29 2022-08-30 昆明理工大学 一种基于声道滤波和声门激励的语音重构方法
CN119564344A (zh) * 2024-10-09 2025-03-07 中国科学院香港创新研究院人工智能与机器人创新中心 基于多模态大语言模型的手术导航方法以及装置

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5999500A (ja) * 1982-11-29 1984-06-08 日本電信電話株式会社 音声認識方法
JP2834471B2 (ja) * 1989-04-17 1998-12-09 日本電信電話株式会社 発音評価法
JP3252282B2 (ja) * 1998-12-17 2002-02-04 松下電器産業株式会社 シーンを検索する方法及びその装置

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5999500A (ja) * 1982-11-29 1984-06-08 日本電信電話株式会社 音声認識方法
JP2834471B2 (ja) * 1989-04-17 1998-12-09 日本電信電話株式会社 発音評価法
JP3252282B2 (ja) * 1998-12-17 2002-02-04 松下電器産業株式会社 シーンを検索する方法及びその装置

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
SINGER H. ET AL.: "Pitch dependent phone modelling for HMM-based speech recognition", THE JOURNAL OF THE ACOUSTICAL SOCIETY OF JAPAN (E), vol. 15, no. 2, March 1994 (1994-03-01), pages 77 - 86, XP003007849 *

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2012242542A (ja) * 2011-05-18 2012-12-10 Nippon Hoso Kyokai <Nhk> 音声比較装置及び音声比較プログラム
US11069373B2 (en) 2017-09-25 2021-07-20 Fujitsu Limited Speech processing method, speech processing apparatus, and non-transitory computer-readable storage medium for storing speech processing computer program
JP2019074580A (ja) * 2017-10-13 2019-05-16 Kddi株式会社 音声認識方法、装置およびプログラム
CN111145728A (zh) * 2019-12-05 2020-05-12 厦门快商通科技股份有限公司 语音识别模型训练方法、系统、移动终端及存储介质
CN111145728B (zh) * 2019-12-05 2022-10-28 厦门快商通科技股份有限公司 语音识别模型训练方法、系统、移动终端及存储介质
CN114974271A (zh) * 2021-12-29 2022-08-30 昆明理工大学 一种基于声道滤波和声门激励的语音重构方法
CN119564344A (zh) * 2024-10-09 2025-03-07 中国科学院香港创新研究院人工智能与机器人创新中心 基于多模态大语言模型的手术导航方法以及装置

Also Published As

Publication number Publication date
JPWO2007015489A1 (ja) 2009-02-19
JP4961565B2 (ja) 2012-06-27

Similar Documents

Publication Publication Date Title
EP2659482B1 (en) Ranking representative segments in media data
JP4005154B2 (ja) 音声復号化方法及び装置
KR100958144B1 (ko) 오디오 압축
McLoughlin Line spectral pairs
CN104252862B (zh) 处理音频信号的方法和装置
US20150262587A1 (en) Pitch Synchronous Speech Coding Based on Timbre Vectors
US6678655B2 (en) Method and system for low bit rate speech coding with speech recognition features and pitch providing reconstruction of the spectral envelope
TW201246183A (en) Extraction and matching of characteristic fingerprints from audio signals
CN110472097A (zh) 乐曲自动分类方法、装置、计算机设备和存储介质
CN1920947B (zh) 用于低比特率音频编码的语音/音乐检测器
WO2006114964A1 (ja) ピッチ周期等化装置及びピッチ周期等化方法、並びに音声符号化装置、音声復号装置及び音声符号化方法
US8942977B2 (en) System and method for speech recognition using pitch-synchronous spectral parameters
Vuppala et al. Improved consonant–vowel recognition for low bit‐rate coded speech
Yu et al. Sparse cepstral codes and power scale for instrument identification
US20040199382A1 (en) Method and apparatus for formant tracking using a residual model
Kim et al. Note-level singing melody transcription for time-aligned musical score generation
JP4961565B2 (ja) 音声検索装置及び音声検索方法
Khemiri et al. Automatic detection of known advertisements in radio broadcast with data-driven ALISP transcriptions
Tarján et al. A bilingual study on the prediction of morph-based improvement
KR20070069631A (ko) 음성 신호에서 음소를 분절하는 방법 및 그 시스템
CN106548784B (zh) 一种语音数据的评价方法及系统
KR100766170B1 (ko) 다중 레벨 양자화를 이용한 음악 요약 장치 및 방법
Kos et al. Online speech/music segmentation based on the variance mean of filter bank energy
WO2004088634A1 (ja) 音声信号圧縮装置、音声信号圧縮方法及びプログラム
JP4213416B2 (ja) ワードスポッティング音声認識装置、ワードスポッティング音声認識方法、ワードスポッティング音声認識用プログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 2007529275

Country of ref document: JP

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 06782105

Country of ref document: EP

Kind code of ref document: A1