WO2017056885A1 - 楽曲処理方法および楽曲処理装置 - Google Patents

楽曲処理方法および楽曲処理装置 Download PDF

Info

Publication number
WO2017056885A1
WO2017056885A1 PCT/JP2016/076266 JP2016076266W WO2017056885A1 WO 2017056885 A1 WO2017056885 A1 WO 2017056885A1 JP 2016076266 W JP2016076266 W JP 2016076266W WO 2017056885 A1 WO2017056885 A1 WO 2017056885A1
Authority
WO
WIPO (PCT)
Prior art keywords
music
target
reproduction
sound
search
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2016/076266
Other languages
English (en)
French (fr)
Inventor
秀樹 高野
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Yamaha Corp
Original Assignee
Yamaha Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Yamaha Corp filed Critical Yamaha Corp
Publication of WO2017056885A1 publication Critical patent/WO2017056885A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H1/00Details of electrophonic musical instruments
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • G10L25/54Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for retrieval
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/90Pitch determination of speech signals

Definitions

  • the present invention relates to a technique for searching and playing back music.
  • inging voice A technique for searching for a song sung by a user from a plurality of candidates by analyzing a voice uttered by the user by singing the song (hereinafter referred to as “singing voice”) has been proposed.
  • the basic frequency sequentially detected from the user's singing voice is converted into an intermediate format and then the similarity is analyzed with each of a plurality of candidate songs, thereby A technique for searching for music corresponding to the above is disclosed.
  • an object of the present invention is to reduce a burden on a user who desires to search for music and to generate sound in parallel with reproduction of the music (for example, singing or playing the music). To do.
  • the computer system corresponds to the target sound by music search processing using pitch information corresponding to the pitch of the target sound whose pitch changes with time.
  • the target music is searched in parallel with the progress of the target sound, and the target music is played back in parallel with the progress of the target sound from the portion corresponding to the progress of the target sound in the target music.
  • the music processing device uses the music search processing using pitch information corresponding to the pitch of the target sound whose pitch changes over time to target music corresponding to the target sound.
  • a music search unit that searches in parallel with the progress of the sound, and a playback control unit that causes the playback device to play back the target music in parallel with the progress of the target sound from the portion of the target music that corresponds to the progress of the target sound To do.
  • FIG. 1 is a configuration diagram of a music processing system according to a first embodiment of the present invention. It is a flowchart of an acoustic analysis process. It is explanatory drawing of an acoustic analysis process. It is a flowchart of a music search process. It is explanatory drawing of note corresponding information. It is a flowchart of a reproduction
  • FIG. 1 is a configuration diagram of a music processing system 100 according to the first embodiment of the present invention.
  • the music processing system 100 includes a terminal device 12 and a music processing device 14.
  • the terminal device 12 and the music processing device 14 communicate with each other via a communication network 16 including a mobile communication network or the Internet.
  • the terminal device 12 is, for example, a portable communication terminal such as a mobile phone or a smartphone, or a portable or stationary communication terminal such as a personal computer.
  • the music processing device 14 is a server device that searches for music sung by the user of the terminal device 12 (hereinafter referred to as “target music”) and causes the terminal device 12 to reproduce the target music.
  • target music searches for music sung by the user of the terminal device 12 (hereinafter referred to as “target music”) and causes the terminal device 12 to reproduce the target music.
  • the music processing apparatus 14 searches for a target music in response to a search request (query) transmitted from the terminal device 12, and transmits a reproduction request including the music file F of the target music to transmit the terminal apparatus. 12 to play the target music.
  • the music file F is data representing the accompaniment sound of the target music, for example.
  • a plurality of terminal devices 12 can communicate with the music processing device 14, but the following description focuses on one terminal device 12 for convenience.
  • the terminal device 12 is realized by a computer system including a control device 20, a storage device 22, a communication device 24, a display device 26, an operation device 28, a sound collection device 32, and a reproduction device 34.
  • the control device 20 is composed of, for example, a CPU (Central Processing ⁇ Unit) and comprehensively controls each element of the terminal device 12.
  • the storage device 22 is realized by a known recording medium such as a magnetic recording medium and a semiconductor recording medium, or a combination of a plurality of types of recording media, and a program executed by the control device 20 and various data used by the control device 20. And remember.
  • the communication device 24 communicates with the music processing device 14 via the communication network 16.
  • the communication device 24 according to the first embodiment transmits a target music search request to the music processing device 14 and receives the music file F of the target music transmitted from the music processing device 14.
  • the display device 26 displays an image instructed from the control device 20. For example, the search result (music name of the target music) by the music processing device 14 is displayed on the display device 26.
  • the operation device 28 is an input device that receives an instruction from a user.
  • the operation device 28 is, for example, a plurality of operators that detect an operation by the user, or a touch panel that detects a user's contact with the display surface of the display device 26.
  • the sound collection device 32 collects ambient sound and generates an acoustic signal XA.
  • the user of the terminal device 12 sings a desired music piece toward the sound collection device 32.
  • the sound collection device 32 generates an acoustic signal XA representing the sound (hereinafter referred to as “singing sound”) in which the user of the terminal device 12 sang music.
  • the A / D converter for converting the acoustic signal XA from analog to digital is not shown for convenience.
  • the control device 20 of the first embodiment functions as an acoustic analysis unit 42 and a reproduction processing unit 44 by executing a program stored in the storage device 22.
  • the acoustic analysis unit 42 generates the pitch information Z corresponding to the pitch of each note included in the singing voice by analyzing the acoustic signal XA supplied from the sound collection device 32.
  • the reproduction processing unit 44 generates an audio signal XB of the target music from the music file F received by the communication device 24 from the music processing device 14 and supplies it to the reproduction device 34.
  • the reproduction device 34 (for example, a speaker or headphones) reproduces the music indicated by the acoustic signal XB supplied from the reproduction processing unit 44. Specifically, the accompaniment sound of the target music (for example, the performance sound of a musical instrument) is reproduced by the reproduction device 34.
  • the D / A converter for converting the acoustic signal XB from digital to analog is not shown for convenience.
  • FIG. 2 is a flowchart of a process (hereinafter referred to as “acoustic analysis process”) in which the acoustic analysis unit 42 according to the first embodiment generates pitch information Z.
  • the acoustic analysis process is started in response to an instruction from the user to the terminal device 12.
  • the acoustic analysis unit 42 sequentially detects the pitch (fundamental frequency) p of the acoustic signal XA (SA1), and stabilizes the acoustic signal XA on the time axis as illustrated in FIG.
  • a section x1 is divided into a variable section x2 (SA2).
  • the stable section x1 is a section where the pitch p is stable in time
  • the fluctuation section x2 is a section where the pitch p fluctuates unstable.
  • a series of sections (for example, a series of sections in which the degree of dispersion of the plurality of pitches p is less than a predetermined value) is a stable section where time-dependent changes in the feature values related to the plurality of pitches p satisfy a specific condition. defined as x1.
  • An arbitrary stable section x1 of the acoustic signal XA corresponds to one musical note of the music.
  • the acoustic analysis unit 42 calculates a representative value (for example, an average value or a median value) of a plurality of pitches p in the stable section x1 as a pitch P of the stable section x1 for each stable section x1 (that is, for each musical note). (SA3).
  • the acoustic analysis unit 42 sequentially generates pitch information Z corresponding to the pitch P for each note of the acoustic signal XA (SA4). As illustrated in FIG. 3, the acoustic analysis unit 42 according to the first embodiment performs a difference value (that is, a pitch difference) of the pitch P between one stable section x1 and the immediately preceding stable section x1 of the acoustic signal XA. ) Pitch information Z corresponding to ⁇ is generated for each stable section x1.
  • a difference value that is, a pitch difference
  • different codes for example, letters “A” to “Z”
  • a code assigned to a range including the pitch difference ⁇ is generated as pitch information Z.
  • the pitch information Z is information expressing the pitch P of the singing voice as a relative value (amount of fluctuation from the previous pitch P), and the absolute value of the pitch P and the stable interval.
  • the time and time length of x1 that is, the singing rhythm by the user) are not reflected in the pitch information Z.
  • the acoustic analysis unit 42 generates a search request every time the pitch information Z is generated (that is, every stable section x1) and transmits the search request from the communication device 24 to the music processing device 14.
  • the search request includes a time series (Z 1, Z 2,%) Of a plurality of pitch information Z generated up to the present time and the time axis of notes corresponding to each pitch information Z.
  • time information TZ (TZ1, TZ2,7) For designating the position of.
  • the time information TZ of any one note specifies, for example, the time of the sounding point of the note (for example, the time of the start point of the stable section x1).
  • a search request including the pitch information Z and the time information TZ for each note is sequentially transmitted from the terminal device 12 to the music processing device 14 in parallel with the progress of the singing voice.
  • the acoustic signal XA may include a section where the pitch p is not properly detected (hereinafter referred to as “non-detection section”). For example, a silent section in which the pitch p cannot be accurately detected due to a lack of volume or the like and a consonant section in which a consonant having no harmonic structure is pronounced are assumed as non-detection sections.
  • the pitch information Z according to the pitch difference between the zone a and the zone immediately before it, the zone b and the zone just before it
  • a search request including pitch information Z corresponding to a pitch difference (that is, zero) from the section a where the pitch is detected is generated.
  • a symbol indicating a non-detection section can be generated as the pitch information Z for a silent section or a consonant section. It is also possible to determine the time information TZ by including the consonant interval corresponding to one consonant in the vowel interval immediately after the consonant as one note.
  • the music processing apparatus 14 is realized by a computer system including a control device 50, a storage device 52, and a communication device 54.
  • the control device 50 is constituted by a CPU, for example, and comprehensively controls each element of the music processing device 14.
  • the communication device 54 communicates with the terminal device 12 via the communication network 16.
  • the communication device 54 of the first embodiment receives the search request transmitted from the terminal device 12 and transmits the music file F of the target music searched in response to the search request to the requesting terminal device 12.
  • the storage device 52 is realized by a known recording medium such as a magnetic recording medium and a semiconductor recording medium, or a combination of a plurality of types of recording media, and a program executed by the control device 50 and various data used by the control device 50. And remember.
  • the storage device 52 of the first embodiment stores search information V and a music file F for each of N pieces of music (hereinafter referred to as “candidate music”) that are candidates for the target music.
  • N is a natural number of 2 or more.
  • the search information V is information used for searching for the target music, and expresses a time series of pitches for each note in the candidate music.
  • the search information V of each candidate song is composed of a plurality of pitch information Y time series (Y1, Y2,7) Corresponding to different notes of the candidate song.
  • the pitch information Y of any one note is a code corresponding to the pitch difference from the immediately preceding note (for example, “A”, similarly to the pitch information Z generated by the acoustic analysis unit 42 of the terminal device 12. " ⁇ " (“Z" characters).
  • the music file F of each candidate music specifies the performance contents of the candidate music.
  • the music file F of the first embodiment includes instruction information E (E1, E2,%) For instructing the pronunciation or mute of each note constituting the accompaniment sound of the candidate music, and each instruction information E.
  • This is a MIDI (Musical Instrument Digital Interface) format file in which time information TE (TE1, TE2,%) That designates processing time points (for example, time intervals of instructions that follow each other) is arranged in time series.
  • TE Time information
  • TE2 Time information
  • the MIDI format is a suitable example of the format of the music file F, and the format of the music file F is arbitrary.
  • the music file F includes a standard reproduction speed (hereinafter referred to as “standard reproduction speed”) ⁇ 0 of the candidate music.
  • the control device 50 functions as a music search unit 62 and a reproduction control unit 64 by executing a program stored in the storage device 52.
  • the music search unit 62 searches for a target music corresponding to the singing voice that is pronounced by the user of the terminal device 12.
  • the reproduction control unit 64 causes the reproduction device 34 of the terminal device 12 to reproduce the target music searched by the music search unit 62.
  • the playback control unit 64 transmits a playback request including the music file F of the target music from the communication device 54 to the terminal device 12 that has transmitted the search request.
  • FIG. 4 is a flowchart of a process in which the music search unit 62 of the first embodiment searches for a target music (hereinafter referred to as “music search process”).
  • the music search process is started when the communication device 54 receives the search request.
  • the music search unit 62 selects one candidate music (hereinafter referred to as “selected candidate music”) from the N candidate music (SB1). Then, the music search unit 62 calculates the similarity index (hereinafter referred to as “similarity index”) M between the singing voice and the selection candidate music through the first process SB2 that compares the singing voice with the selection candidate music.
  • the first process SB2 compares the time series (code string) of the pitch information Y specified by the search information V of the selection candidate music with the time series of the pitch information Z of the singing voice specified by the search request. It is processing to do.
  • the similarity index M of the first embodiment is set according to the edit distance ⁇ between the time series of the pitch information Y of the selection candidate music and the time series of the pitch information Z of the singing voice.
  • the edit distance ⁇ is the number of edits when one of the time series of the pitch information Y and the time series of the pitch information Z is converted into the other by editing (deleting, inserting, replacing) the pitch information Y or the pitch information Z. Means the minimum value of. Therefore, the similarity index M becomes smaller as the singing voice and the selection candidate music are more similar.
  • the music search unit 62 includes a time series of pitch information Y and a time series of pitch information Z of the singing voice for each of a plurality of sections obtained by dividing the selection candidate music on the time axis.
  • the edit distance ⁇ is calculated, and the minimum value of the edit distance ⁇ (that is, the edit distance ⁇ of the section most similar to the singing voice in the selection candidate music) is determined as the similarity index M of the selection candidate music.
  • dynamic programming Dynamic programming
  • the first process SB2 of the first embodiment is a temporal correspondence between each pitch information Y of the selection candidate song and each pitch information Z of the singing voice (that is, the time of each note in the selection candidate song and the singing voice).
  • a route search process for searching for a route indicating a specific response).
  • the similarity index M is also calculated. That is, the note correspondence information C representing the temporal correspondence of each note between the selection candidate song and the singing voice is generated together with the calculation of the similarity index M.
  • the music search unit 62 uses the dynamic programming method to determine the optimum path for each of the pitch information Y of the selection candidate music and the pitch information Z of the singing voice.
  • the note correspondence information C is searched together with the calculation of the edit distance ⁇ by the above, and the notes of the selection candidate music and the notes of the singing voice are associated with each pair of the pitch information Y and the pitch information Z on the optimum route. Is generated.
  • a known route search process such as a Viterbi algorithm may be employed.
  • the music search unit 62 determines whether or not the generation of the similarity index M and the note correspondence information C (that is, the first process SB2) has been executed for the N candidate music (SB3). If the determination result is negative (SB3: NO), the music search unit 62 selects an unprocessed candidate song from among the N candidate songs as a new selection candidate song (SB1), and then selects the selected candidate song. The first process SB2 is executed. On the other hand, when the first process SB2 is executed for N candidate songs (SB3: YES), the song search unit 62 selects the target song from the N candidate songs according to the similarity index M of each candidate song. 2 Process SC is executed.
  • the music search unit 62 identifies one candidate song having a similarity index M having a minimum value (Mmin) among the N candidate songs as a provisional target song (hereinafter referred to as “provisional song”). (SC1), it is determined whether the similarity index Mmin of the provisional music is below the threshold value Mth (SC2). If the similarity index Mmin is lower than the threshold value Mth (SC2: YES), it can be evaluated that the provisional music is sufficiently similar to the singing voice, so the music search unit 62 determines the provisional music as the target music (SC3).
  • the music search process is terminated without determining the target music.
  • the target music similar to the singing voice is searched, and the note correspondence information C indicating the temporal correspondence of each note between the singing voice and the target music is generated.
  • the playback control unit 64 in FIG. 1 transmits a playback request including the music file F of the target music searched by the music searching unit 62 among the N music files F stored in the storage device 52 to the terminal device 12. Then, the reproduction device 34 of the terminal device 12 is caused to reproduce the target music piece.
  • FIG. 6 is a flowchart of processing (hereinafter referred to as “playback control processing”) for the playback control unit 64 to cause the playback device 34 to play back the target song.
  • the reproduction control process is started when the music search unit 62 searches for the target music.
  • the playback control unit 64 acquires the music file F of the target music searched by the music search unit 62 from the storage device 52 (SD1), and generates playback information D that specifies the playback conditions of the target music. (SD2-SD4).
  • the reproduction information D is a reproduction condition for causing the reproduction of the target music by the music file F to follow the singing voice (that is, reproducing the target music at a speed corresponding to the singing voice so that the reproduction of the target music is synchronized with the singing voice). And is variably set according to the characteristics of the singing voice.
  • the reproduction information D of the first embodiment includes a reproduction speed ⁇ B, a reproduction key ⁇ B, and time correspondence information ⁇ .
  • the music search unit 62 sequentially performs the setting of the reproduction speed ⁇ B (SD2), the setting of the reproduction key ⁇ B (SD3), and the setting of the time correspondence information ⁇ (SD4). Then, the playback control unit 64 transmits a playback request including the music file F of the target music and the playback information D from the communication device 54 to the terminal device 12 (SD5). Specific contents and setting method of the reproduction information D will be exemplified below.
  • the reproduction speed ⁇ B is the reproduction speed (tempo) of the target music piece by the reproduction apparatus 34, and is expressed by, for example, the number of beats per unit time (Beats Per Minute).
  • FIG. 7 is an explanatory diagram of processing (SD2) in which the playback control unit 64 sets the playback speed ⁇ B.
  • the reproduction control unit 64 of the first embodiment sets the reproduction speed ⁇ B of the target music according to the sounding speed ⁇ A so as to approximate or match the speed of singing voice progression (hereinafter referred to as “sounding speed”) ⁇ A.
  • the sounding point of each note of the singing voice (for example, the starting point of the stable section x1) and the sounding point of each note specified by the music file F of the target music (in the section corresponding to each note of the singing voice) are shown. It is shown under a common time axis.
  • the pronunciation point of each note (pitch information Z) of the singing voice is specified by the time information TZ (TZ1, TZ2,%) In the search request, and the pronunciation point of each note (instruction information E) of the target song is a music file. It is specified by time information TE (TE1, TE2,...) In F.
  • the temporal correspondence between each note of the singing voice and each note of the target music is specified by the note correspondence information C generated by the first process SB2. As illustrated in FIG.
  • the note pronunciation information C specifies the correspondence of each note.
  • the position of the point on the time axis may be different.
  • the reproduction control unit 64 of the first embodiment refers to the temporal correspondence for each note specified by the note correspondence information C between the singing voice and the target music, and the target music corresponding to the pronunciation speed ⁇ A of the singing voice. Set the playback speed ⁇ B.
  • the playback control unit 64 first sets a time interval (IOI: Inter Onset Interval) QA (IOI) for each note up to the last note specified in the search request (that is, the note immediately before the singer sang). QA1, QA2,...) Are calculated, and the total value LA of the time interval QA to the last note is calculated (SD21).
  • the reproduction control unit 64 calculates a time interval QB (QB1, QB2,%) With respect to the immediately preceding note for each note corresponding to the already sung section in the singing voice of the target music, and sets the time for a plurality of notes.
  • the total value LB of the interval QB is calculated (SD22).
  • the playback control unit 64 multiplies the standard playback speed ⁇ 0 of the target music specified in the music file F by the ratio (LB / LA) of the total value LB to the total value LA, thereby reproducing the target music playback speed ⁇ B.
  • the reproduction key ⁇ B set by the reproduction control unit 64 in step SD3 is a key (key) for reproducing the target music by the reproduction device 34.
  • the search request transmitted by the terminal device 12 of the first embodiment includes the pitch P (absolute value) of the first note of the singing voice. Therefore, it is possible to specify the pitch for each note of the singing voice by sequentially adding the pitch difference specified by each pitch information Z to the pitch P of the first note.
  • the pitch information Z and the pitch P may be included in the search request for each note and transmitted from the terminal device 12 to the music processing device 14.
  • the reproduction control unit 64 sets the reproduction key ⁇ B of the target music according to the pitch difference between the notes indicated by the note correspondence information C between the singing voice and the target music indicated by the music file F.
  • a numerical value obtained by averaging the pitch differences between the pitches of the notes of the singing voice and the pitches of the notes of the target music indicated by the instruction information E of the music file F over a plurality of notes has a predetermined threshold value.
  • the reproduction control unit 64 sets the reproduction adjustment ⁇ B so as to be lower (for example, zero).
  • the playback control unit 64 generates time correspondence information ⁇ in step SD4 of FIG.
  • the time correspondence information ⁇ is the time point indicated by the time information TZ for any one note of the singing voice (hereinafter referred to as “reference note”) and the note associated with the reference note by the note correspondence information C of the target music.
  • This is information that correlates the time point indicated by the time information TE with each other. That is, the time correspondence information ⁇ is information indicating a correspondence between one time in the singing voice (time indicated by the time information TZ) and one time in the target music (time indicated by the time information TE).
  • the time point indicated by the time information TE is expressed, for example, by the number of MIDI clocks from the start point of the target music.
  • the reproduction information D including the reproduction speed ⁇ B, the reproduction key ⁇ B, and the time correspondence information ⁇ exemplified above is transmitted from the music processing device 14 to the terminal device 12 as a reproduction request together with the music file F of the target music.
  • the reproduction processing unit 44 When the communication device 24 of the terminal device 12 receives the reproduction request, the reproduction processing unit 44 generates an acoustic signal XB specified by the music file F according to the reproduction information D and supplies it to the reproduction device 34. Therefore, the target music piece is reproduced by the reproduction device 34 under the reproduction speed ⁇ B and the reproduction key ⁇ B specified by the reproduction information D. That is, the reproduction speed of the music file F of the target music is adjusted to the reproduction speed ⁇ B included in the reproduction information D, and the reproduction key of the music file F is adjusted to the reproduction key ⁇ B included in the reproduction information D.
  • the singing by the user is in progress from the time when the search request is transmitted from the terminal device 12 to the present time. Therefore, if the reproduction of the target music is started immediately after the section where the user sang at the time of transmission of the search request among the target music, the reproduction time of the target music is delayed with respect to the singing by the user. . Therefore, the reproduction processing unit 44 according to the first embodiment starts the reproduction of the target music from the time corresponding to the progress of the singing voice (that is, the time when the user of the terminal device 12 actually sings).
  • the playback start time (hereinafter referred to as “playback start point”) in the target music is controlled according to the time correspondence information ⁇ . Specifically, as illustrated in FIG.
  • the reference note in the target music is recorded over the elapsed time ⁇ from the time point when the reference note specified by the time correspondence information ⁇ is sounded (the time indicated by the time information TZ) to the present time.
  • a time point t2 at which the playback point is advanced in time from the time t1 of the corresponding note at the playback speed ⁇ B is selected as the playback start point of the target music piece. Therefore, the section after the point of time when the user sings is reproduced among the target music pieces at the reproduction speed ⁇ B and the reproduction key ⁇ B equivalent to the singing voice by the user.
  • the target music is reproduced so as to follow the singing voice (that is, the target music is synchronized with the singing voice by reproducing the target music at a speed corresponding to the singing voice).
  • the reproduction control unit 64 of the music processing device 14 in the first embodiment reproduces the target music from the reproduction start point t2 corresponding to the progress of the singing voice in parallel with the progress of the singing voice. 34 functions as an element to be reproduced.
  • FIG. 9 is an explanatory diagram of the operation of the music processing system 100.
  • a search request including a time series of a plurality of pitch information Z and time information TZ of each pitch information Z for each detection of a stable section x1 of the acoustic signal XA (that is, for each note).
  • the music search unit 62 searches for one target music corresponding to the singing voice by the music search process of FIG. 4 (S2).
  • the music search unit 62 transmits a notification of the result of the music search process (hereinafter referred to as “result notification”) to the terminal device 12 (S3).
  • the result notification includes the song name of the target song.
  • the control device 20 of the terminal device 12 receives the result notification, displays the name of the target music specified in the result notification on the display device 26 and presents the search result to the user. For example, it is possible to display on the display device 26 a message such as “song A will be played after about 3 seconds”, for example, for notifying the playback of the target song along with the name of the target song.
  • the user confirms the display on the display device 26 and determines whether or not the target song is a song being sung by himself (correctness of the search result).
  • the target music is not a music being sung (when the search result is incorrect)
  • the user instructs the music change by appropriately operating the operation device 28.
  • the music change instruction is transmitted from the terminal device 12 to the music processing device 14.
  • the target song is a song being sung (when the search result is appropriate)
  • the user does not operate the operation device 28. Therefore, a music change instruction is not transmitted to the music processing device 14.
  • the music search unit 62 of the music processing device 14 determines whether or not the user has instructed to change the music before a predetermined time (hereinafter referred to as “standby time”) elapses from the result notification (S3) (S4).
  • the standby time is set to a time length of about several seconds, for example.
  • the music search unit 62 discards the result of the previous music search process and uses the search request (S1) transmitted from the terminal device 12 thereafter. Then, the music search process is re-executed (S2). Since the pitch information Z of a new note is added to the search request every time it is transmitted, the target music searched in the music search process changes for each music search process.
  • the music search unit 62 of the first embodiment searches for the target music by the music search process in parallel with the progress of the singing voice.
  • the target music searched in the previous music search process is determined as the search result.
  • the reproduction control unit 64 causes the reproduction device 34 of the terminal device 12 to reproduce the target music searched in the immediately preceding music search process (S5).
  • the playback control unit 64 generates a playback request including the music file F of the target music and the playback information D that specifies the playback conditions (playback speed ⁇ B, playback key ⁇ B, time correspondence information ⁇ ) of the target music.
  • the data is transmitted from the communication device 54 to the terminal device 12 via the communication network 16.
  • the reproduction processing unit 44 of the terminal device 12 receives an acoustic signal from the music file F so that the target music is reproduced at the reproduction speed ⁇ B and the reproduction key ⁇ B from the time corresponding to the time correspondence information ⁇ and the reproduction speed ⁇ B of the target music.
  • XB is generated and supplied to the playback device 34. That is, the target music is reproduced so as to follow the singing voice in parallel with the progress of the singing voice.
  • the target music corresponding to the singing voice is searched in parallel with the progress of the singing voice, and singing from the reproduction start point t2 corresponding to the progress of the singing voice among the target music.
  • the target music is played in parallel with the progress of the audio. That is, the search for the target music and the reproduction of the target music are sequentially performed in parallel with the progress of a series of singing voices in which the user sang the desired music. Specifically, when the user continuously sings a desired music piece, reproduction of the target music piece (accompaniment sound) is started in parallel with the singing. Therefore, it is possible to reduce the burden on the user who desires the search for the target music and the singing of the target music in parallel with the reproduction of the target music.
  • the music search process is re-executed when the user gives an instruction to change the music before a predetermined waiting time elapses from the result notification of the music search process. Therefore, it is possible to correct an erroneous search by the music search process and reproduce an appropriate target music.
  • the standby time elapses without the instruction from the user to change the music
  • the playback of the target music searched in the previous music search process is started, so it is appropriate without disturbing the singing by the user
  • a target music can be reproduced.
  • the target music is played at a playback speed ⁇ B corresponding to the singing voice pronunciation speed ⁇ A, the user can continue singing in parallel with the playback of the target music without any discomfort after executing the music search process.
  • the result of the first process SB2 in the music search process (temporal correspondence of each note between the singing voice and the target music) is used for setting the reproduction speed ⁇ B of the target music. Therefore, it is possible to reduce the processing load of the reproduction control unit 64 as compared with a configuration in which the reproduction speed ⁇ B of the target music is set by a process separate from the music search process.
  • the target music is reproduced with the reproduction key ⁇ B corresponding to the pitch difference between the notes corresponding to each other between the singing voice and the target music, after the music search process, There is an advantage that the user can continue singing in parallel with the reproduction without a sense of incongruity.
  • the result of the first process SB2 in the music search process (temporal correspondence of each note between the singing voice and the target music) is used for setting the reproduction tone ⁇ B of the target music. Therefore, it is possible to reduce the processing load of the reproduction control unit 64 as compared with the configuration in which the reproduction tone ⁇ B of the target music is set by a process separate from the music search process.
  • the reproduction start point t2 is specified in the sound processing device 14, and the reproduction information D including the reproduction start point t2 instead of the time correspondence information ⁇
  • a configuration for transmitting to the terminal device 12 (hereinafter referred to as “proportional”) is also assumed.
  • the transmission delay in the communication network 16 changes every moment. Therefore, in comparison, due to the magnitude of the transmission delay, for example, there is a possibility that the reproduction of the target music is started from the time when the user of the terminal device 12 deviates from the place where the user actually sings.
  • the time correspondence information ⁇ for designating the temporal correspondence between the singing voice and the target music and the reproduction speed ⁇ B of the target music are displayed together with the music file F on the terminal device 12. Sent to. Therefore, it is possible to specify the playback start point t2 corresponding to the progress of the singing voice among the target music regardless of the transmission delay in the communication network 16 (the target music of the music file F can be played back in synchronization with the singing voice). ).
  • a proportional configuration can be included in the scope of the present invention.
  • Second Embodiment A second embodiment of the present invention will be described.
  • the detailed description of each is abbreviate
  • the editing distance ⁇ used for calculating the similarity index M tends to be larger as the singing voice is longer (the total number of pitch information Z is larger). Therefore, it can be evaluated that the candidate music in which the edit distance ⁇ is maintained at a numerical value with a small similarity index M even though the time length T of the singing voice is long is sufficiently likely to correspond to the music sung by the user. .
  • the candidate song whose similarity index M calculated in the above procedure is the minimum value Mmin lower than the threshold value Mth is determined as the target song, as in the first embodiment.
  • the same effect as in the first embodiment is realized.
  • the numerical value obtained by dividing the edit distance ⁇ by the time length T of the singing voice is calculated as the similarity index M, it is possible to search for a target song that is likely to be a song sung by the user.
  • the similarity index M between the singing voice and the candidate song is not limited to the edit distance ⁇ exemplified in each of the above embodiments.
  • a known distance measure such as the Euclidean distance between the time series of the pitch information Y of the candidate music and the time series of the pitch information Z of the singing voice or the Earth Mover's distance can be used as the similarity index M. is there. It is also possible to calculate the correlation between the singing voice and the candidate song as the similarity index M. In the configuration in which the correlation between the singing voice and the candidate music is calculated as the similarity index M, the similarity index M becomes a larger numerical value as the singing voice and the candidate music are more similar.
  • a provisional music with the similarity index M having the maximum value Mmax is selected (SC1), and when the maximum value Mmax exceeds the threshold value Mth (SC2: YES), the provisional music is the target. Confirmed as a song.
  • the similarity index M is comprehensively expressed as an index of the degree of similarity (distance or correlation) between the singing voice and the candidate music, and the calculation method of the similarity index M and the singing voice It does not matter whether the numerical value of the similarity index M increases or decreases with respect to the degree of similarity with the candidate music.
  • one target music selected according to the similarity index M is presented to the user of the terminal device 12 as a result notification (S3).
  • the distance measure such as ⁇ is the similarity index M
  • a plurality of candidate songs selected in the ascending order of the similarity index M can be transmitted to the terminal device 12 as a result notification and presented to the user.
  • the playback control unit 64 transmits a playback request including the music file F of the target song to the terminal device 12 (S5). .
  • a list of a plurality of candidate songs searched in the song search process is displayed on the display device 26 of the terminal device 12. Specifically, a list in which the names of a plurality of candidate songs are arranged in the order of the similarity index M is displayed. Depending on the similarity index M, the display mode (for example, display color or size) of each music piece can be made different.
  • the user can select the music intended by the user from the list by operating the operation device 28 as the target music.
  • the display device 26 highlights the target music selected by the user. For example, the target music selected by the user is moved to the top of the list and displayed in a display mode different from other music.
  • the search is completed when the user selects a desired song as the target song.
  • the generation and transmission of the search request is terminated when the user selects the target music, and the music search process is not executed thereafter.
  • the sound utilized for search of a target music is not limited to a singing voice.
  • the music search unit 62 can search for a target music corresponding to a performance sound produced by playing a musical instrument by the user. That is, the same processing as in the above-described embodiments is performed on the acoustic signal XA representing the performance sound collected by the sound collection device 32.
  • any sound (target sound) that the user pronounces (sings or plays) a song is used to search for the target song.
  • the time series of the pitch information Z generated from the acoustic signal XA by the acoustic analysis unit 42 of the terminal device 12 is transmitted to the music processing device 14. It is also possible to mount it on. Specifically, the acoustic signal XA generated by the sound collection device 32 is transmitted from the terminal device 12 to the music processing device 14 via the communication network 16, and the acoustic analysis unit 42 of the music processing device 14 analyzes the acoustic signal XA. A time series of pitch information Z is generated.
  • the acoustic analysis unit 42 is mounted on the terminal device 12 from the viewpoint of reducing the communication load of the communication network 16. Therefore, a configuration in which the pitch information Z is transmitted to the music processing device 14 (that is, a configuration in which the acoustic signal XA does not need to be transmitted to the music processing device 14) is advantageous.
  • the one-to-one correspondence of each note between the candidate music and the singing voice is analyzed by the first process SB2 (generation of note correspondence information C). It is also possible to analyze the correspondence between sequences (hereinafter referred to as “note sequences”) between candidate music and singing voice. For example, by executing the first process SB2 with a predetermined number of notes arranged in time series as a unit (note string), the correspondence of each note string between the candidate music and the singing voice (note correspondence information C) is analyzed. Is done.
  • the “temporal correspondence of each note” analyzed in the first process SB2 for the candidate music and the singing voice (target sound) is between the individual notes exemplified in the above-described embodiments.
  • the correspondence between musical note strings composed of a plurality of musical notes is also included.
  • the server device that communicates with the terminal device 12 is exemplified as the music processing device 14, but the music processing device 14 is realized by a single device (for example, an information terminal such as a mobile phone and a smartphone). Is also possible.
  • the music processing device 14 illustrated in FIG. 10 includes a control device 50, a storage device 52, a sound collection device 32, and a reproduction device 34.
  • the storage device 52 stores the music file F and the search information V for each of the N candidate music pieces in the same manner as the above-described embodiments.
  • the control device 50 generates the pitch information Z corresponding to the acoustic signal XA generated by the sound pickup device 32 by the acoustic analysis unit 42 in FIG.
  • a music search unit 62 that searches for the target music by processing and a playback control unit 64 that causes the playback device 34 to play back the target music searched by the music search unit 62 are realized.
  • the reproduction control unit 64 generates the music file F and the reproduction information D of the target music in the same manner as in the first embodiment, and reproduces the acoustic signal XB corresponding to the music file F and the reproduction information D. 34.
  • the reproduction control unit 64 is comprehensively expressed as an element that causes the reproduction device 34 to reproduce the target music piece.
  • the playback device 34 is installed separately from the music processing device 14 as illustrated in the first embodiment or whether the playback device 34 is installed in the music processing device 14 as illustrated in FIG. is there.
  • the storage device 52 (for example, cloud storage) is installed separately from the music processing device 14, and the music processing device 14 refers to the search information V or the music file F in the storage device 52 via the communication network 16, for example. Configurations can also be employed.
  • the music processing apparatus 14 illustrated in the above-described embodiments is realized by the cooperation of the control device 50 and the program as illustrated in the above-described embodiments.
  • the program according to a preferred aspect of the present invention corresponds to the target sound by the music search process (FIG. 4) using the pitch information Z corresponding to the pitch of each note of the target sound such as singing voice and performance sound.
  • the music search unit 62 that searches for the target music in parallel with the progress of the target sound, and the target music in the playback device 34 in parallel with the progress of the target sound from the time corresponding to the progress of the target sound of the target music.
  • the computer is caused to function as the reproduction control unit 64 for reproduction.
  • the programs exemplified above can be provided in a form stored in a computer-readable recording medium and installed in the computer.
  • the recording medium is, for example, a non-transitory recording medium, and an optical recording medium (optical disk) such as a CD-ROM is a good example, but any known recording medium such as a semiconductor recording medium or a magnetic recording medium This type of recording medium can be included.
  • non-transitory recording medium includes all computer-readable recording media except for transient propagation signals (transitory, “propagating” signal), and does not exclude volatile recording media. . It is also possible to distribute the program to a computer in the form of distribution via a communication network.
  • the present invention is also specified as an operation method (music processing method) of the music processing device 14 according to each of the above-described embodiments.
  • the target sound is obtained by the music search process (FIG. 4) using the pitch information Z corresponding to the pitch of each note of the target sound such as singing voice and performance sound.
  • the target music corresponding to is searched in parallel with the progress of the target sound, and the target music is played back by the playback device 34 in parallel with the progress of the target sound from the time corresponding to the progress of the target sound in the target music ( S5).
  • the computer system uses the music search processing using the pitch information corresponding to the pitch of the target sound whose pitch changes with time, and the target sound.
  • the target music corresponding to is searched in parallel with the progress of the target sound, and the target music is played back by the playback device in parallel with the progress of the target sound from the portion of the target music corresponding to the progress of the target sound.
  • the target music corresponding to the target sound is searched in parallel with the progress of the target sound, and the target music is parallel to the progress of the target sound from the portion corresponding to the progress of the target sound in the target music. Played. That is, the search for the target music and the reproduction of the target music are sequentially performed in parallel with the progression of the target sound of the music. Therefore, it is possible to reduce the burden on the user who desires the search for the target music and the pronunciation of the target sound in parallel with the reproduction of the target music (for example, singing or playing the target music).
  • aspect 2 in the search for the target music, the music search process is re-executed when the user gives an instruction to change the music before a predetermined time elapses from the notification of the result of the music search process.
  • the reproduction apparatus In the reproduction of the target music, when a predetermined time has passed without an instruction to change the music, the reproduction apparatus is made to reproduce the target music searched in the immediately preceding music search process.
  • the music search process is re-executed when the user gives an instruction to change the music before a predetermined time elapses from the result notification of the music search process. Therefore, it is possible to correct an erroneous search by the music search process and reproduce an appropriate target music.
  • the music search processing is performed by comparing the time series of pitch information between the target sound and the candidate music for each of the plurality of candidate music.
  • a first process for generating temporal correspondence between a plurality of notes in a sound and a plurality of notes in a candidate song, and a similarity index that is an index of similarity between the target sound and the candidate song, and a similarity of each candidate song A second process of selecting a target music from a plurality of candidate music according to the index, and in the reproduction of the target music, the time of each note analyzed by the first process between the target sound and the target music With reference to the correspondence, the reproduction speed of the target music corresponding to the sounding speed of the target sound is set, and the reproduction apparatus is caused to reproduce the target music at the reproduction speed.
  • the target music is played at a playback speed corresponding to the sound generation speed of the target sound
  • the user can continue to sound the target sound in parallel with the playback of the target music after the music search process is executed.
  • the result of the first process that is, the temporal correspondence of each note between the target sound and the target music
  • the music search process for searching for the target music is used for setting the playback speed of the target music. . Therefore, it is possible to reduce the processing load of the reproduction control unit as compared with a configuration in which the reproduction speed of the target music is set by a process separate from the music search process.
  • the result of the first process (that is, the temporal correspondence of each note between the target sound and the target music) in the music search process for searching for the target music is used for setting the reproduction tone of the target music. . Therefore, it is possible to reduce the processing load of the playback control unit as compared with a configuration in which the playback tone of the target music is set by a process separate from the music search process.
  • ⁇ Aspect 5> In a preferred example (aspect 5) of aspect 3 or aspect 4, in reproducing the target music, a music file representing the target music, time correspondence information for designating temporal correspondence between the target sound and the target music, and a target The reproduction speed of the music is transmitted to the terminal device equipped with the reproduction device via the communication network.
  • the time correspondence information specifying the temporal correspondence between the target sound and the target music and the reproduction speed of the target music are transmitted to the terminal device together with the music file of the target music. Therefore, regardless of the transmission delay in the communication network, it is possible to specify the portion of the target music corresponding to the progress of the target sound (the target music of the music file can be reproduced so as to synchronize with the progress of the target sound). There is.
  • the music processing apparatus which concerns on the suitable aspect (aspect 6) of this invention is the target corresponding to an object sound by the music search process using the pitch information according to the pitch of the object sound to which a pitch changes temporally.
  • a music search unit that searches for music in parallel with the progress of the target sound, and playback control that causes the playback device to play back the target music in parallel with the progress of the target sound from the portion of the target music that corresponds to the progress of the target sound Part.
  • the target music corresponding to the target sound is searched in parallel with the progress of the target sound, and the target music is parallel to the progress of the target sound from the portion of the target music corresponding to the progress of the target sound. Played.
  • the search for the target music and the reproduction of the target music are sequentially performed in parallel with the progression of the target sound of the music. Therefore, it is possible to reduce the burden on the user who desires the search for the target music and the pronunciation of the target sound in parallel with the reproduction of the target music (for example, singing or playing the target music).

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

コンピュータシステムが、対象音の音高に応じた音高情報を利用した楽曲検索処理により、前記対象音に対応する目標楽曲を当該対象音の進行に並行して検索し、前記目標楽曲のうち当該対象音の進行に対応した部分から前記対象音の進行に並行して前記目標楽曲を再生装置に再生させる。

Description

楽曲処理方法および楽曲処理装置
 本発明は、楽曲を検索および再生する技術に関する。
 楽曲の歌唱により利用者が発音する音声(以下「歌唱音声」という)を解析することで、当該利用者が歌唱した楽曲を複数の候補から検索する技術が従来から提案されている。例えば特許文献1には、利用者の歌唱音声から順次に検出される基本周波数を中間的なフォーマットに変換してから複数の候補楽曲の各々との間で類似性を解析することで、歌唱音声に対応する楽曲を検索する技術が開示されている。
特開2002-132278号公報
 しかし、特許文献1の技術では、歌唱音声に対応する楽曲が検索されるに過ぎない。したがって、楽曲の再生に並行した歌唱を所望する利用者は、第1に、楽曲の一部を歌唱して当該楽曲を検索したうえで、第2に、当該楽曲を最初から再生して歌唱する必要がある。なお、以上の説明では楽曲の歌唱を便宜的に例示したが、例えば楽曲を楽器で演奏する場面でも同様の問題が発生し得る。以上の事情を考慮して、本発明は、楽曲の検索と当該楽曲の再生に並行した音響の発音(例えば当該楽曲の歌唱または演奏)とを所望する利用者の負担を軽減することを目的とする。
 本発明の好適な態様に係る楽曲処理方法は、コンピュータシステムが、音高が時間的に変化する対象音の音高に応じた音高情報を利用した楽曲検索処理により、前記対象音に対応する目標楽曲を当該対象音の進行に並行して検索し、前記目標楽曲のうち当該対象音の進行に対応した部分から前記対象音の進行に並行して前記目標楽曲を再生装置に再生させる。
 本発明の好適な態様に係る楽曲処理装置は、音高が時間的に変化する対象音の音高に応じた音高情報を利用した楽曲検索処理により、対象音に対応する目標楽曲を当該対象音の進行に並行して検索する楽曲検索部と、目標楽曲のうち当該対象音の進行に対応した部分から対象音の進行に並行して目標楽曲を再生装置に再生させる再生制御部とを具備する。
本発明の第1実施形態に係る楽曲処理システムの構成図である。 音響解析処理のフローチャートである。 音響解析処理の説明図である。 楽曲検索処理のフローチャートである。 音符対応情報の説明図である。 再生制御処理のフローチャートである。 再生速度の設定の説明図である。 再生開始点の設定の説明図である。 楽曲処理システムの動作の説明図である。 変形例に係る楽曲処理装置の構成図である。
<第1実施形態>
 図1は、本発明の第1実施形態に係る楽曲処理システム100の構成図である。図1に例示される通り、第1実施形態の楽曲処理システム100は、端末装置12と楽曲処理装置14とを具備する。端末装置12と楽曲処理装置14とは、移動体通信網またはインターネット等を含む通信網16を介して相互に通信する。端末装置12は、例えば携帯電話機もしくはスマートフォン等の可搬型の通信端末、またはパーソナルコンピュータ等の可搬型もしくは据置型の通信端末である。楽曲処理装置14は、端末装置12の利用者が歌唱した楽曲(以下「目標楽曲」という)を検索するとともに目標楽曲を端末装置12に再生させるサーバ装置である。第1実施形態の楽曲処理装置14は、端末装置12から送信される検索要求(クエリ)を契機として目標楽曲を検索し、当該目標楽曲の楽曲ファイルFを含む再生要求を送信することで端末装置12に目標楽曲を再生させる。楽曲ファイルFは、例えば目標楽曲の伴奏音を表すデータである。なお、実際には複数の端末装置12が楽曲処理装置14と通信し得るが、以下の説明では便宜的に1個の端末装置12に着目する。
 図1に例示される通り、端末装置12は、制御装置20と記憶装置22と通信装置24と表示装置26と操作装置28と収音装置32と再生装置34とを具備するコンピュータシステムで実現される。制御装置20は、例えばCPU(Central Processing Unit)で構成され、端末装置12の各要素を統括的に制御する。記憶装置22は、例えば磁気記録媒体および半導体記録媒体等の公知の記録媒体、または、複数種の記録媒体の組合せで実現され、制御装置20が実行するプログラムと制御装置20が使用する各種のデータとを記憶する。通信装置24は、通信網16を介して楽曲処理装置14と通信する。第1実施形態の通信装置24は、目標楽曲の検索要求を楽曲処理装置14に送信し、楽曲処理装置14から送信された目標楽曲の楽曲ファイルFを受信する。
 表示装置26は、制御装置20から指示された画像を表示する。例えば楽曲処理装置14による検索結果(目標楽曲の楽曲名)が表示装置26に表示される。操作装置28は、利用者からの指示を受付ける入力機器である。操作装置28は、例えば利用者による操作を検知する複数の操作子、または、表示装置26の表示面に対する利用者の接触を検知するタッチパネルである。
 収音装置32は、周囲の音響を収音して音響信号XAを生成する。端末装置12の利用者は、収音装置32に向かって所望の楽曲を歌唱する。収音装置32は、端末装置12の利用者が楽曲を歌唱した音声(以下「歌唱音声」という)を表す音響信号XAを生成する。なお、音響信号XAをアナログからデジタルに変換するA/D変換器の図示は便宜的に省略した。
 図1に例示される通り、第1実施形態の制御装置20は、記憶装置22に記憶されたプログラムを実行することで音響解析部42および再生処理部44として機能する。音響解析部42は、収音装置32から供給される音響信号XAを解析することで、歌唱音声に含まれる音符毎の音高に応じた音高情報Zを生成する。再生処理部44は、通信装置24が楽曲処理装置14から受信した楽曲ファイルFから目標楽曲の音響信号XBを生成して再生装置34に供給する。
 再生装置34(例えばスピーカまたはヘッドホン)は、再生処理部44から供給される音響信号XBが示す楽曲を再生する。具体的には、目標楽曲の伴奏音(例えば楽器の演奏音)が再生装置34により再生される。なお、音響信号XBをデジタルからアナログに変換するD/A変換器の図示は便宜的に省略した。
 図2は、第1実施形態の音響解析部42が音高情報Zを生成する処理(以下「音響解析処理」という)のフローチャートである。端末装置12に対する利用者からの指示を契機として音響解析処理が開始される。音響解析処理を開始すると、音響解析部42は、音響信号XAの音高(基本周波数)pを順次に検出し(SA1)、図3に例示される通り、音響信号XAを時間軸上で安定区間x1と変動区間x2とに区分する(SA2)。安定区間x1は、音高pが時間的に安定に推移する区間であり、変動区間x2は、音高pが不安定に変動する区間である。具体的には、複数の音高pに関する特徴量の時間変化が特定の条件を満たす一連の区間(例えば複数の音高pの分散等の散布度が所定値を下回る一連の区間)が安定区間x1として画定される。音響信号XAの任意の1個の安定区間x1は、楽曲の1個の音符に対応する。音響解析部42は、安定区間x1内の複数の音高pの代表値(例えば平均値または中央値)を当該安定区間x1の音高Pとして安定区間x1毎(すなわち楽曲の音符毎)に算定する(SA3)。
 音響解析部42は、音響信号XAの音符毎の音高Pに応じた音高情報Zを順次に生成する(SA4)。図3に例示される通り、第1実施形態の音響解析部42は、音響信号XAの1個の安定区間x1と直前の安定区間x1との間における音高Pの差分値(すなわち音高差)δに応じた音高情報Zを安定区間x1毎に生成する。具体的には、音高差δの数値範囲を所定の単位(例えば100cent単位)で区画した複数の範囲の各々に相異なる符号(例えば“A”~“Z”の文字)が付与され、複数の範囲のうち音高差δを内包する範囲に付与された符号が音高情報Zとして生成される。以上の説明から理解される通り、音高情報Zは歌唱音声の音高Pを相対値(直前の音高Pからの変動量)で表現する情報であり、音高Pの絶対値と安定区間x1の時刻および時間長(すなわち利用者による歌唱のリズム)とは音高情報Zに反映されない。
 音響解析部42は、音高情報Zの生成毎(すなわち安定区間x1毎)に検索要求を生成して通信装置24から楽曲処理装置14に送信する。検索要求は、図1に例示される通り、現時点までに生成された複数の音高情報Zの時系列(Z1,Z2,……)と、各音高情報Zに対応する音符の時間軸上の位置を指定する時間情報TZ(TZ1,TZ2,……)とを包含する。任意の1個の音符の時間情報TZは、例えば当該音符の発音点の時刻(例えば安定区間x1の始点の時刻)を指定する。以上の説明から理解される通り、音高情報Zと時間情報TZとを音符毎に含む検索要求が、歌唱音声の進行に並行して端末装置12から楽曲処理装置14に順次に送信される。
 なお、音響信号XAには、音高pが適正に検出されない区間(以下「非検出区間」という)が含まれ得る。例えば、音量の不足等の理由により音高pを正確に検出できない無音区間と、調波構造を持たない子音が発音されている子音区間とが、非検出区間として想定される。非検出区間の直前の区間aと直後の区間bとで音高pが同一である場合、区間aとその直前の区間との音高差に応じた音高情報Zと、区間bとその直前に音高が検出された区間aとの音高差(すなわちゼロ)に応じた音高情報Zとを含む検索要求が生成される。また、無音区間または子音区間について、非検出区間であることを示す記号を音高情報Zとして生成することも可能である。また、1個の子音に対応する子音区間を、当該子音の直後の母音の区間に含めて1個の音符として、時間情報TZを決定することも可能である。
 図1に例示される通り、楽曲処理装置14は、制御装置50と記憶装置52と通信装置54とを具備するコンピュータシステムで実現される。制御装置50は、例えばCPUで構成され、楽曲処理装置14の各要素を統括的に制御する。通信装置54は、通信網16を介して端末装置12と通信する。第1実施形態の通信装置54は、端末装置12から送信された検索要求を受信するとともに、当該検索要求に応じて検索された目標楽曲の楽曲ファイルFを要求元の端末装置12に送信する。記憶装置52は、例えば磁気記録媒体および半導体記録媒体等の公知の記録媒体、または、複数種の記録媒体の組合せで実現され、制御装置50が実行するプログラムと制御装置50が使用する各種のデータとを記憶する。
 図1に例示される通り、第1実施形態の記憶装置52は、目標楽曲の候補となるN個の楽曲(以下「候補楽曲」という)の各々について検索情報Vと楽曲ファイルFとを記憶する(Nは2以上の自然数)。検索情報Vは、目標楽曲の検索に利用される情報であり、候補楽曲内の音符毎の音高の時系列を表現する。図1に例示される通り、各候補楽曲の検索情報Vは、当該候補楽曲の相異なる音符に対応する複数の音高情報Yの時系列(Y1,Y2,……)で構成される。任意の1個の音符の音高情報Yは、端末装置12の音響解析部42が生成する前述の音高情報Zと同様に、直前の音符との音高差に応じた符号(例えば“A”~“Z”の文字)を指定する。
 他方、各候補楽曲の楽曲ファイルFは、当該候補楽曲の演奏内容を指定する。具体的には、第1実施形態の楽曲ファイルFは、候補楽曲の伴奏音を構成する各音符の発音または消音を指示する指示情報E(E1,E2,……)と、各指示情報Eの処理時点(例えば相前後する指示の時間間隔)を指定する時間情報TE(TE1,TE2,……)とが時系列に配列されたMIDI(Musical Instrument Digital Interface)形式のファイルである。ただし、MIDI形式は楽曲ファイルFの形式の好適な例示であり、楽曲ファイルFの形式は任意である。また、楽曲ファイルFは、候補楽曲の標準的な再生速度(以下「標準再生速度」という)ν0を包含する。
 図1に例示される通り、第1実施形態の制御装置50は、記憶装置52に記憶されたプログラムを実行することで楽曲検索部62および再生制御部64として機能する。楽曲検索部62は、端末装置12の利用者が発音した歌唱音声に対応する目標楽曲を検索する。再生制御部64は、楽曲検索部62が検索した目標楽曲を端末装置12の再生装置34に再生させる。具体的には、再生制御部64は、目標楽曲の楽曲ファイルFを含む再生要求を検索要求の送信元の端末装置12に通信装置54から送信する。
 図4は、第1実施形態の楽曲検索部62が目標楽曲を検索する処理(以下「楽曲検索処理」という)のフローチャートである。通信装置54による検索要求の受信を契機として楽曲検索処理が開始される。
 楽曲検索処理を開始すると、楽曲検索部62は、N個の候補楽曲から1個の候補楽曲(以下「選択候補楽曲」という)を選択する(SB1)。そして、楽曲検索部62は、歌唱音声と選択候補楽曲とを相互に対比する第1処理SB2により歌唱音声と選択候補楽曲との類否の指標(以下「類似指標」という)Mを算定する。第1処理SB2は、選択候補楽曲の検索情報Vが指定する音高情報Yの時系列(符号列)と、検索要求で指定される歌唱音声の音高情報Zの時系列とを相互に対比する処理である。
 第1実施形態の類似指標Mは、選択候補楽曲の音高情報Yの時系列と歌唱音声の音高情報Zの時系列との間の編集距離εに応じて設定される。編集距離εは、音高情報Yまたは音高情報Zの編集(削除,挿入,置換)により音高情報Yの時系列および音高情報Zの時系列の一方を他方に変換するときの編集回数の最小値を意味する。したがって、歌唱音声と選択候補楽曲とが類似するほど類似指標Mは小さい数値となる。具体的には、楽曲検索部62は、選択候補楽曲を時間軸上で区画した複数の区間の各々について当該区間内の音高情報Yの時系列と歌唱音声の音高情報Zの時系列との編集距離εを算定し、編集距離εの最小値(すなわち選択候補楽曲のうち歌唱音声に最も類似する区間の編集距離ε)を当該選択候補楽曲の類似指標Mとして確定する。編集距離εの算定には、例えば特開2013-242409号公報にも開示される通り、例えば動的計画法(Dynamic Programming)が好適に利用される。
 第1実施形態の第1処理SB2は、選択候補楽曲の各音高情報Yと歌唱音声の各音高情報Zとの時間的な対応(すなわち、選択候補楽曲と歌唱音声とにおける各音符の時間的な対応)を示す経路を探索する処理(すなわち経路探索処理)である。経路探索処理のなかで類似指標Mも算定される。すなわち、選択候補楽曲と歌唱音声との間の各音符の時間的な対応を表す音符対応情報Cが類似指標Mの算定とともに生成される。具体的には、図5に例示される通り、楽曲検索部62は、選択候補楽曲の各音高情報Yと歌唱音声の各音高情報Zとを相互に対応させる最適経路を動的計画法による編集距離εの算定とともに探索し、当該最適経路上の音高情報Yと音高情報Zとの組毎に選択候補楽曲の各音符と歌唱音声の各音符とを対応させた音符対応情報Cを生成する。なお、音符対応情報Cの解析には、ビタビアルゴリズム等の公知の経路探索処理も採用され得る。
 楽曲検索部62は、N個の候補楽曲について類似指標Mおよび音符対応情報Cの生成(すなわち第1処理SB2)を実行したか否かを判定する(SB3)。判定結果が否定である場合(SB3:NO)、楽曲検索部62は、N個の候補楽曲のうち未処理の候補楽曲を新たな選択候補楽曲として選択(SB1)したうえで当該選択候補楽曲について第1処理SB2を実行する。他方、N個の候補楽曲について第1処理SB2を実行した場合(SB3:YES)、楽曲検索部62は、各候補楽曲の類似指標Mに応じてN個の候補楽曲から目標楽曲を選択する第2処理SCを実行する。
 具体的には、楽曲検索部62は、N個の候補楽曲のうち類似指標Mが最小値(Mmin)である1個の候補楽曲を暫定的な目標楽曲(以下「暫定楽曲」という)として特定し(SC1)、暫定楽曲の類似指標Mminが閾値Mthを下回るか否かを判定する(SC2)。類似指標Mminが閾値Mthを下回る場合(SC2:YES)には、暫定楽曲が歌唱音声に充分に類似すると評価できるから、楽曲検索部62は、当該暫定楽曲を目標楽曲として確定する(SC3)。他方、類似指標Mminが閾値Mthを上回る場合(SC2:NO)には、暫定楽曲と歌唱音声とは必ずしも類似しないと評価できるから、目標楽曲を確定せずに楽曲検索処理を終了する。以上に例示した楽曲検索処理により、歌唱音声に類似する目標楽曲が検索されるとともに、歌唱音声と目標楽曲との間の各音符の時間的な対応を示す音符対応情報Cが生成される。
 図1の再生制御部64は、記憶装置52に記憶されたN個の楽曲ファイルFのうち楽曲検索部62が検索した目標楽曲の楽曲ファイルFを含む再生要求を端末装置12に送信することで、端末装置12の再生装置34に目標楽曲を再生させる。図6は、再生制御部64が目標楽曲を再生装置34に再生させるための処理(以下「再生制御処理」という)のフローチャートである。楽曲検索部62による目標楽曲の検索を契機として再生制御処理が開始される。
 再生制御処理を開始すると、再生制御部64は、楽曲検索部62が検索した目標楽曲の楽曲ファイルFを記憶装置52から取得し(SD1)、目標楽曲の再生条件を指定する再生情報Dを生成する(SD2~SD4)。再生情報Dは、楽曲ファイルFによる目標楽曲の再生を歌唱音声に追従させる(すなわち目標楽曲の再生が歌唱音声に同期するように目標楽曲を歌唱音声に応じた速度で再生する)ための再生条件を指定する情報であり、歌唱音声の特性に応じて可変に設定される。具体的には、第1実施形態の再生情報Dは、再生速度νBと再生調κBと時間対応情報τとを包含する。楽曲検索部62は、再生速度νBの設定(SD2)と再生調κBの設定(SD3)と時間対応情報τの設定(SD4)とを順次に実行する。そして、再生制御部64は、目標楽曲の楽曲ファイルFと再生情報Dとを含む再生要求を通信装置54から端末装置12に送信する(SD5)。再生情報Dの具体的な内容と設定方法について以下に例示する。
 再生速度νBは、再生装置34による目標楽曲の再生の速度(テンポ)であり、例えば単位時間毎の拍数(Beats Per Minute)で表現される。図7は、再生制御部64が再生速度νBを設定する処理(SD2)の説明図である。第1実施形態の再生制御部64は、歌唱音声の進行の速度(以下「発音速度」という)νAに近似または合致するように、発音速度νAに応じて目標楽曲の再生速度νBを設定する。
 図7には、歌唱音声の各音符の発音点(例えば安定区間x1の始点)と目標楽曲の楽曲ファイルFが指定する各音符(歌唱音声の各音符に対応する区間内)の発音点とが共通の時間軸のもとで図示されている。歌唱音声の各音符(音高情報Z)の発音点は検索要求内の時間情報TZ(TZ1,TZ2,……)で指定され、目標楽曲の各音符(指示情報E)の発音点は楽曲ファイルF内の時間情報TE(TE1,TE2,……)で指定される。歌唱音声の各音符と目標楽曲の各音符との時間的な対応は、前述の第1処理SB2により生成された音符対応情報Cで指定される。図7に例示される通り、楽曲ファイルFで指定される標準再生速度ν0で目標楽曲を再生した場合、歌唱音声と目標楽曲との間では、音符対応情報Cが対応を指定する各音符の発音点の時間軸上の位置が相違し得る。第1実施形態の再生制御部64は、歌唱音声と目標楽曲との間で音符対応情報Cが指定する音符毎の時間的な対応を参照して、歌唱音声の発音速度νAに対応する目標楽曲の再生速度νBを設定する。
 再生制御部64は、まず、検索要求で指定された最後の音符(すなわち、歌唱者が歌唱した直前の音符)までの各音符について直前の音符との時間間隔(IOI:Inter Onset Interval)QA(QA1,QA2,……)を算定し、当該最後の音符までの時間間隔QAの合計値LAを算定する(SD21)。また、再生制御部64は、目標楽曲のうち歌唱音声における歌唱済の区間に対応する各音符について直前の音符との時間間隔QB(QB1,QB2,……)を算定し、複数の音符について時間間隔QBの合計値LBを算定する(SD22)。そして、再生制御部64は、楽曲ファイルFで指定される目標楽曲の標準再生速度ν0に、合計値LAに対する合計値LBの比率(LB/LA)を乗算することで、目標楽曲の再生速度νB(νB=ν0×(LB/LA))を算定する(SD23)。すなわち、目標楽曲のうち歌唱音声の各音符に対応する区間が、時間軸上で比率(LB/LA)に応じて伸縮される。したがって、歌唱音声の発音速度νAが目標楽曲の標準再生速度ν0と比較して速い(LA<LB)ほど、目標楽曲の再生速度νBは増加する。なお、図7に例示される通り、歌唱音声の各音符の時間間隔QAn(n=1,2,……)と目標楽曲の各音符の時間間隔QBnとで相互に対応するもの同士の差分Δn(Δn=QAn-QBn)の合計値(Δ1+Δ2+……)が減少する(理想的にはゼロになる)ように、再生制御部64が目標楽曲の再生速度νBを調整する、と表現することも可能である。
 他方、再生制御部64がステップSD3で設定する再生調κBは、再生装置34による目標楽曲の再生の調(キー)である。第1実施形態の端末装置12が送信する検索要求は、歌唱音声の最初の音符の音高P(絶対値)を包含する。したがって、各音高情報Zで指定される音高差を最初の音符の音高Pに順次に加算することで歌唱音声の音符毎の音高を特定することが可能である。なお、音符毎に音高情報Zと音高Pとを検索要求に含めて端末装置12から楽曲処理装置14に送信することも可能である。再生制御部64は、歌唱音声と楽曲ファイルFが示す目標楽曲との間で音符対応情報Cが対応を示す音符間の音高差に応じて目標楽曲の再生調κBを設定する。具体的には、歌唱音声の各音符の音高と楽曲ファイルFの各指示情報Eが示す目標楽曲の各音符の音高との音高差を複数の音符にわたり平均した数値が所定の閾値を下回る(例えばゼロとなる)ように、再生制御部64は再生調κBを設定する。
 また、再生制御部64は、図6のステップSD4において時間対応情報τを生成する。時間対応情報τは、歌唱音声の任意の1個の音符(以下「基準音符」という)について時間情報TZが示す時点と、目標楽曲のうち音符対応情報Cにより基準音符に対応づけられた音符の時間情報TEが示す時点とを相互に対応させる情報である。すなわち、時間対応情報τは、歌唱音声における1個の時刻(時間情報TZが示す時点)と目標楽曲における1個の時刻(時間情報TEが示す時点)との対応を示す情報である。時間情報TEが示す時点は、例えば目標楽曲の始点からのMIDIクロックの個数で表現される。
 以上に例示した再生速度νBと再生調κBと時間対応情報τとを含む再生情報Dが目標楽曲の楽曲ファイルFとともに再生要求として楽曲処理装置14から端末装置12に送信される。端末装置12の通信装置24が再生要求を受信すると、再生処理部44は、楽曲ファイルFで指定される音響の音響信号XBを再生情報Dに応じて生成して再生装置34に供給する。したがって、再生情報Dが指定する再生速度νBおよび再生調κBのもとで目標楽曲が再生装置34により再生される。すなわち、目標楽曲の楽曲ファイルFの再生の速度が、再生情報Dに含まれる再生速度νBに調整され、楽曲ファイルFの再生の調が、再生情報Dに含まれる再生調κBに調整される。
 端末装置12から検索要求が送信されてから現時点までの期間では利用者による歌唱が進行している。したがって、目標楽曲のうち検索要求の送信の時点で利用者が歌唱した区間の直後から目標楽曲の再生を開始したのでは、目標楽曲の再生時点が利用者による歌唱に対して遅延した状態となる。そこで、第1実施形態の再生処理部44は、歌唱音声の進行に対応した時点(すなわち、端末装置12の利用者が実際に歌唱している時点)から目標楽曲の再生が開始されるように、目標楽曲内の再生の開始の時点(以下「再生開始点」という)を時間対応情報τに応じて制御する。具体的には、図8に例示される通り、時間対応情報τが指定する基準音符の発音の時点(時間情報TZが示す時刻)から現在までの経過時間σにわたり、目標楽曲内で基準音符に対応する音符の時点t1から再生速度νBで再生点を時間的に進行させた時点t2が、目標楽曲の再生開始点として選定される。したがって、利用者による歌唱音声と同等の再生速度νBおよび再生調κBのもとで、目標楽曲のうち利用者が歌唱した時点以降の区間が再生される。すなわち、歌唱音声に追従する(すなわち、歌唱音声に応じた速度で目標楽曲を再生することで目標楽曲が歌唱音声に同期する)ように目標楽曲が再生される。以上の説明から理解される通り、第1実施形態における楽曲処理装置14の再生制御部64は、歌唱音声の進行に対応した再生開始点t2から目標楽曲を歌唱音声の進行に並行して再生装置34に再生させる要素として機能する。
 図9は、楽曲処理システム100の動作の説明図である。図9に例示される通り、音響信号XAの安定区間x1の検出毎(すなわち音符毎)に、複数の音高情報Zの時系列と各音高情報Zの時間情報TZとを包含する検索要求が、端末装置12から楽曲処理装置14に送信される(S1)。通信装置54が検索要求を受信すると、楽曲検索部62は、歌唱音声に対応する1個の目標楽曲を図4の楽曲検索処理により検索する(S2)。楽曲検索部62は、楽曲検索処理の結果の通知(以下「結果通知」という)を端末装置12に送信する(S3)。結果通知は目標楽曲の楽曲名を包含する。端末装置12の制御装置20が結果通知を受信すると、制御装置20は、結果通知で指定された目標楽曲の楽曲名を表示装置26に表示させて検索結果を利用者に提示する。なお、例えば目標楽曲の再生を予告する「約3秒後に楽曲Aを再生します」等のメッセージを目標楽曲の楽曲名とともに表示装置26に表示させることも可能である。
 利用者は、表示装置26の表示を確認し、目標楽曲が自身の歌唱中の楽曲であるか否か(検索結果の正誤)を判定する。そして、目標楽曲が歌唱中の楽曲でない場合(検索結果が誤りである場合)、利用者は、操作装置28を適宜に操作することで楽曲変更を指示する。図9に破線で図示される通り、楽曲変更の指示は、端末装置12から楽曲処理装置14に送信される。目標楽曲が歌唱中の楽曲である場合(検索結果が適正である場合)には、利用者は操作装置28を操作しない。したがって、楽曲処理装置14に楽曲変更の指示は送信されない。
 楽曲処理装置14の楽曲検索部62は、結果通知(S3)から所定時間(以下「待機時間」という)が経過するまでに利用者から楽曲変更が指示されたか否かを判定する(S4)。待機時間は、例えば数秒程度の時間長に設定される。待機時間内に楽曲変更が指示された場合(S4:YES)、楽曲検索部62は、直前の楽曲検索処理の結果を破棄し、端末装置12から以降に送信される検索要求(S1)を利用して楽曲検索処理を再実行する(S2)。検索要求には送信毎に新たな音符の音高情報Zが追加されるから、楽曲検索処理で検索される目標楽曲は楽曲検索処理毎に変化する。以上の説明から理解される通り、第1実施形態の楽曲検索部62は、歌唱音声の進行に並行して楽曲検索処理により目標楽曲を検索する。
 他方、楽曲変更を利用者から指示されずに待機時間が経過した場合には(S4:NO)、直前の楽曲検索処理で検索された目標楽曲が検索結果として確定する。再生制御部64は、直前の楽曲検索処理で検索された目標楽曲を端末装置12の再生装置34に再生させる(S5)。具体的には、再生制御部64は、目標楽曲の楽曲ファイルFと当該目標楽曲の再生条件(再生速度νB,再生調κB,時間対応情報τ)を指定する再生情報Dとを含む再生要求を、通信装置54から通信網16を介して端末装置12に送信する。端末装置12の再生処理部44は、目標楽曲のうち時間対応情報τと再生速度νBとに応じた時点から再生速度νBおよび再生調κBで目標楽曲が再生されるように楽曲ファイルFから音響信号XBを生成して再生装置34に供給する。すなわち、歌唱音声の進行に並行して歌唱音声に追従するように目標楽曲が再生される。
 以上に説明した通り、第1実施形態では、歌唱音声に対応する目標楽曲が歌唱音声の進行に並行して検索され、かつ、目標楽曲のうち歌唱音声の進行に対応した再生開始点t2から歌唱音声の進行に並行して目標楽曲が再生される。すなわち、利用者が所望の楽曲を歌唱した一連の歌唱音声の進行に並行して目標楽曲の検索と当該目標楽曲の再生とが順次に実行される。具体的には、利用者が所望の楽曲を継続的に歌唱していると、途中の時点から歌唱に並行して目標楽曲(伴奏音)の再生が開始される。したがって、目標楽曲の検索と目標楽曲の再生に並行した当該目標楽曲の歌唱とを所望する利用者の負担を軽減することが可能である。
 第1実施形態では、楽曲検索処理の結果通知から所定の待機時間が経過するまでに利用者が楽曲変更を指示した場合に楽曲検索処理が再実行される。したがって、楽曲検索処理による誤検索を修正して適正な目標楽曲を再生することが可能である。また、楽曲変更が利用者から指示されずに待機時間が経過した場合には、直前の楽曲検索処理で検索された目標楽曲の再生が開始されるから、利用者による歌唱を阻害せずに適正な目標楽曲を再生できるという利点がある。
 第1実施形態では、歌唱音声の発音速度νAに対応する再生速度νBで目標楽曲が再生されるから、楽曲検索処理の実行後に、目標楽曲の再生に並行した歌唱を利用者が違和感なく継続できるという利点がある。また、楽曲検索処理のうち第1処理SB2の結果(歌唱音声と目標楽曲との間の各音符の時間的な対応)が、目標楽曲の再生速度νBの設定に流用される。したがって、目標楽曲の再生速度νBを楽曲検索処理とは別個の処理で設定する構成と比較して、再生制御部64の処理負荷を軽減することが可能である。
 第1実施形態では、歌唱音声と目標楽曲との間で相互に対応する音符間の音高差に応じた再生調κBで目標楽曲が再生されるから、楽曲検索処理の実行後に、目標楽曲の再生に並行した歌唱を利用者が違和感なく継続できるという利点がある。また、楽曲検索処理のうち第1処理SB2の結果(歌唱音声と目標楽曲との間の各音符の時間的な対応)が、目標楽曲の再生調κBの設定に流用される。したがって、目標楽曲の再生調κBを楽曲検索処理とは別個の処理で設定する構成と比較して、再生制御部64の処理負荷を軽減することが可能である。
 ところで、目標楽曲の再生時点を端末装置12に指示する構成としては、例えば、音響処理装置14において再生開始点t2を特定し、時間対応情報τの代わりに再生開始点t2を含む再生情報Dを端末装置12に送信する構成(以下「対比例」という)も想定される。しかし、通信網16における伝送遅延は刻々と変動する。したがって、対比例では、伝送遅延の大小に起因して、例えば端末装置12の利用者が実際に歌唱している箇所に対して乖離した時点から目標楽曲の再生が開始される可能性がある。対比例とは対照的に、第1実施形態では、歌唱音声と目標楽曲との間の時間的な対応を指定する時間対応情報τと目標楽曲の再生速度νBとが楽曲ファイルFとともに端末装置12に送信される。したがって、通信網16での伝送遅延の大小に関わらず、目標楽曲のうち歌唱音声の進行に応じた再生開始点t2を特定できる(歌唱音声に同期するように楽曲ファイルFの目標楽曲を再生できる)という利点がある。ただし、対比例の構成も本発明の範囲には包含され得る。
<第2実施形態>
 本発明の第2実施形態を説明する。なお、以下に説明する各例示において作用または機能が第1実施形態と同様である要素については、第1実施形態の説明で使用した符号を流用して各々の詳細な説明を適宜に省略する。
 類似指標Mの算定に利用される編集距離εは、歌唱音声が長い(音高情報Zの総数が多い)ほど大きい数値になるという傾向がある。したがって、歌唱音声の時間長Tが長いにも関わらず編集距離εが類似指標Mが小さい数値に維持される候補楽曲は、利用者が歌唱した楽曲に該当する可能性が充分に高いと評価できる。以上の事情を考慮して、第2実施形態の楽曲検索部62は、第1処理SB2において、編集距離εを時間長Tで除算した数値を類似指標Mとして算定する(M=ε/T)。なお、歌唱音声の時間長Tは、歌唱音声の音符の総数とも換言され得る。N個の候補楽曲のうち、以上の手順で算定された類似指標Mが、閾値Mthを下回る最小値Mminとなる候補楽曲が目標楽曲として確定される点は、第1実施形態と同様である。
 第2実施形態においても第1実施形態と同様の効果が実現される。また、第2実施形態では、編集距離εを歌唱音声の時間長Tで除算した数値が類似指標Mとして算定されるから、利用者が歌唱した楽曲である可能性が高い目標楽曲を検索できるという利点がある。
<変形例>
 以上の各形態は多様に変形され得る。具体的な変形の態様を以下に例示する。以下の例示から任意に選択された2以上の態様は適宜に併合され得る。
(1)歌唱音声と候補楽曲との間の類似指標Mは、前述の各形態で例示した編集距離εに限定されない。例えば、候補楽曲の音高情報Yの時系列と歌唱音声の音高情報Zの時系列との間のユークリッド距離またはEarth Mover's 距離等の公知の距離尺度を類似指標Mとして利用することも可能である。また、歌唱音声と候補楽曲との相関を類似指標Mとして算定することも可能である。歌唱音声と候補楽曲との相関を類似指標Mとして算定する構成では、歌唱音声と候補楽曲とが類似するほど類似指標Mは大きい数値となる。したがって、楽曲検索処理の第2処理SCでは、類似指標Mが最大値Mmaxとなる暫定楽曲が選択され(SC1)、最大値Mmaxが閾値Mthを上回る場合(SC2:YES)に当該暫定楽曲が目標楽曲として確定される。以上の説明から理解される通り、類似指標Mは、歌唱音声と候補楽曲との類似の度合(距離または相関)の指標として包括的に表現され、類似指標Mの算定方法、および、歌唱音声と候補楽曲との類似の度合に対して類似指標Mの数値が増加するか減少するかは不問である。
(2)前述の各形態では、類似指標Mに応じて選定された1個の目標楽曲を結果通知(S3)として端末装置12の利用者に提示したが、類似指標Mの順番(例えば編集距離ε等の距離尺度を類似指標Mとした構成では類似指標Mの昇順)で選択された複数の候補楽曲を結果通知として端末装置12に送信して利用者に提示することも可能である。操作装置28に対する操作で利用者が複数の候補楽曲の何れかを目標楽曲として選択すると、再生制御部64は、当該目標楽曲の楽曲ファイルFを含む再生要求を端末装置12に送信する(S5)。
 例えば、楽曲検索処理で検索された複数の候補楽曲のリストが端末装置12の表示装置26に表示される。具体的には、複数の候補楽曲の楽曲名を類似指標Mの順番で配列したリストが表示される。類似指標Mに応じて各楽曲の表示態様(例えば表示の色またはサイズ)を相違させることも可能である。利用者は、操作装置28を操作することで、自身が意図した楽曲を目標楽曲としてリストから選択可能である。表示装置26は、利用者が選択した目標楽曲を強調表示する。例えば、利用者が選択した目標楽曲がリストの最上位に移行され、他の楽曲とは異なる表示態様で表示される。以上の例示のように複数の候補楽曲を利用者に提示する構成では、利用者が所望の楽曲を目標楽曲として選択した時点で検索が完了する。すなわち、利用者による目標楽曲の選択を契機として検索要求の生成および送信が終了し、以降は楽曲検索処理は実行されない。
(3)前述の各形態では、利用者が発音した歌唱音声に対応する目標楽曲を検索したが、目標楽曲の検索に利用される音響は歌唱音声に限定されない。例えば、利用者による楽器演奏で発音された演奏音に対応する目標楽曲を楽曲検索部62が検索することも可能である。すなわち、収音装置32で収音した演奏音を表す音響信号XAに対して前述の各形態と同様の処理が実行される。以上の説明から理解される通り、利用者が楽曲を発音(歌唱または演奏)した任意の音響(対象音)が目標楽曲の検索に利用される。
(4)前述の各形態では、端末装置12の音響解析部42が音響信号XAから生成した音高情報Zの時系列を楽曲処理装置14に送信したが、音響解析部42を楽曲処理装置14に搭載することも可能である。具体的には、収音装置32が生成した音響信号XAが端末装置12から通信網16を介して楽曲処理装置14に送信され、楽曲処理装置14の音響解析部42が音響信号XAの解析で音高情報Zの時系列を生成する。ただし、以上の構成では音響信号XAを端末装置12から楽曲処理装置14に送信する必要があるから、通信網16の通信負荷を軽減する観点からは、音響解析部42を端末装置12に搭載して音高情報Zを楽曲処理装置14に送信する構成(すなわち、音響信号XAを楽曲処理装置14に送信する必要がない構成)が有利である。
(5)前述の各形態では、候補楽曲と歌唱音声との間で各音符の1対1の対応を第1処理SB2(音符対応情報Cの生成)にて解析したが、複数の音符の時系列(以下「音符列」という)同士の対応を候補楽曲と歌唱音声との間で解析することも可能である。例えば、時系列に配列した所定個の音符を単位(音符列)として第1処理SB2を実行することで、候補楽曲と歌唱音声との間における各音符列の対応(音符対応情報C)が解析される。以上の説明から理解される通り、候補楽曲と歌唱音声(対象音)とについて第1処理SB2で解析される「各音符の時間的な対応」は、前述の各形態で例示した個々の音符同士の1対1の対応のほか、複数の音符で構成される音符列同士の対応も包含する。
(6)前述の各形態では、端末装置12と通信するサーバ装置を楽曲処理装置14として例示したが、楽曲処理装置14を単体の装置(例えば携帯電話機およびスマートフォン等の情報端末)で実現することも可能である。例えば、図10に例示された楽曲処理装置14は、制御装置50と記憶装置52と収音装置32と再生装置34とを具備する。記憶装置52は、前述の各形態と同様に、楽曲ファイルFと検索情報VとをN個の候補楽曲の各々について記憶する。制御装置50は、収音装置32が生成する音響信号XAに応じた音高情報Zを図2の音響解析処理で生成する音響解析部42と、音高情報Zを利用した図4の楽曲検索処理により目標楽曲を検索する楽曲検索部62と、楽曲検索部62が検索した目標楽曲を再生装置34に再生させる再生制御部64とを実現する。具体的には、再生制御部64は、目標楽曲の楽曲ファイルFと再生情報Dとを第1実施形態と同様に生成し、楽曲ファイルFと再生情報Dとに応じた音響信号XBを再生装置34に供給する。以上の説明から理解される通り、再生制御部64は、目標楽曲を再生装置34に再生させる要素として包括的に表現される。第1実施形態の例示のように再生装置34が楽曲処理装置14とは別体として設置されるか、図10の例示のように再生装置34が楽曲処理装置14に設置されるかは不問である。また、なお、記憶装置52(例えばクラウドストレージ)を楽曲処理装置14とは別個に設置し、楽曲処理装置14が例えば通信網16を介して記憶装置52の検索情報Vまたは楽曲ファイルFを参照する構成も採用され得る。
(7)前述の各形態で例示した楽曲処理装置14は、前述の各形態の例示の通り、制御装置50とプログラムとの協働で実現される。本発明の好適な態様に係るプログラムは、歌唱音声および演奏音等の対象音の音符毎の音高に応じた音高情報Zを利用した楽曲検索処理(図4)により、対象音に対応する目標楽曲を当該対象音の進行に並行して検索する楽曲検索部62、および、目標楽曲のうち当該対象音の進行に対応した時点から対象音の進行に並行して目標楽曲を再生装置34に再生させる再生制御部64としてコンピュータを機能させる。以上に例示したプログラムは、コンピュータが読取可能な記録媒体に格納された形態で提供されてコンピュータにインストールされ得る。記録媒体は、例えば非一過性(non-transitory)の記録媒体であり、CD-ROM等の光学式記録媒体(光ディスク)が好例であるが、半導体記録媒体および磁気記録媒体等の公知の任意の形式の記録媒体を包含し得る。なお、「非一過性の記録媒体」とは、一過性の伝搬信号(transitory, propagating signal)を除く全てのコンピュータ読み取り可能な記録媒体を含み、揮発性の記録媒体を除外するものではない。また、通信網を介した配信の形態でプログラムをコンピュータに配信することも可能である。
(8)本発明は、前述の各形態に係る楽曲処理装置14の動作方法(楽曲処理方法)としても特定される。例えば、本発明の一態様に係る楽曲処理方法は、歌唱音声および演奏音等の対象音の音符毎の音高に応じた音高情報Zを利用した楽曲検索処理(図4)により、対象音に対応する目標楽曲を当該対象音の進行に並行して検索し、目標楽曲のうち当該対象音の進行に対応した時点から対象音の進行に並行して目標楽曲を再生装置34に再生させる(S5)。
(9)以上に例示した具体的な形態から把握される本発明の好適な態様を以下に例示する。
<態様1>
 本発明の好適な態様(態様1)に係る楽曲処理方法は、コンピュータシステムが、音高が時間的に変化する対象音の音高に応じた音高情報を利用した楽曲検索処理により、対象音に対応する目標楽曲を当該対象音の進行に並行して検索し、目標楽曲のうち当該対象音の進行に対応した部分から対象音の進行に並行して目標楽曲を再生装置に再生させる。態様1では、対象音に対応する目標楽曲が対象音の進行に並行して検索され、かつ、目標楽曲のうち対象音の進行に対応した部分から当該目標楽曲が当該対象音の進行に並行して再生される。すなわち、楽曲の一連の対象音の進行に並行して目標楽曲の検索と当該目標楽曲の再生とが順次に実行される。したがって、目標楽曲の検索と当該目標楽曲の再生に並行した対象音の発音(例えば目標楽曲の歌唱または演奏)とを所望する利用者の負担を軽減することが可能である。
<態様2>
 態様1の好適例(態様2)では、目標楽曲の検索において、楽曲検索処理の結果通知から所定時間が経過するまでに利用者から楽曲変更が指示された場合に楽曲検索処理を再実行し、目標楽曲の再生において、楽曲変更が指示されずに所定時間が経過した場合に、直前の楽曲検索処理で検索された目標楽曲を再生装置に再生させる。態様2では、楽曲検索処理の結果通知から所定時間が経過するまでに利用者が楽曲変更を指示した場合に楽曲検索処理が再実行される。したがって、楽曲検索処理による誤検索を修正して適正な目標楽曲を再生することが可能である。
<態様3>
 態様1または態様2の好適例(態様3)において、楽曲検索処理は、複数の候補楽曲の各々について、対象音と当該候補楽曲との間で音高情報の時系列を対比することで、対象音における複数の音符と候補楽曲における複数の音符との時間的な対応と、対象音と当該候補楽曲との類否の指標である類似指標とを生成する第1処理と、各候補楽曲の類似指標に応じて複数の候補楽曲から目標楽曲を選択する第2処理とを含み、目標楽曲の再生においては、対象音と目標楽曲との間で第1処理により解析された音符毎の時間的な対応を参照して、対象音の発音速度に対応する当該目標楽曲の再生速度を設定し、目標楽曲を当該再生速度で再生装置に再生させる。態様3では、対象音の発音速度に対応する再生速度で目標楽曲が再生されるから、楽曲検索処理の実行後に、目標楽曲の再生に並行した対象音の発音を利用者が違和感なく継続できるという利点がある。また、目標楽曲を検索する楽曲検索処理のうち第1処理の結果(すなわち、対象音と目標楽曲との間の各音符の時間的な対応)が、目標楽曲の再生速度の設定に流用される。したがって、目標楽曲の再生速度を楽曲検索処理とは別個の処理で設定する構成と比較して、再生制御部の処理負荷を軽減することが可能である。
<態様4>
 態様3の好適例(態様4)では、目標楽曲の再生において、対象音と目標楽曲との間で第1処理により対応が特定された音符間の音高差に応じて当該目標楽曲の再生調を設定し、目標楽曲を当該再生調で再生装置に再生させる。態様4では、対象音と目標楽曲との間で相互に対応する音符間の音高差に応じた再生調で目標楽曲が再生されるから、楽曲検索処理の実行後に、目標楽曲の再生に並行した対象音の発音を利用者が違和感なく継続できるという利点がある。また、目標楽曲を検索する楽曲検索処理のうち第1処理の結果(すなわち、対象音と目標楽曲との間の各音符の時間的な対応)が、目標楽曲の再生調の設定に流用される。したがって、目標楽曲の再生調を楽曲検索処理とは別個の処理で設定する構成と比較して、再生制御部の処理負荷を軽減することが可能である。
<態様5>
 態様3または態様4の好適例(態様5)では、目標楽曲の再生において、目標楽曲を表す楽曲ファイルと、対象音と目標楽曲との間の時間的な対応を指定する時間対応情報と、目標楽曲の再生速度とを、再生装置を具備する端末装置に通信網を介して送信する。以上の態様では、対象音と目標楽曲との間の時間的な対応を指定する時間対応情報と目標楽曲の再生速度とが目標楽曲の楽曲ファイルとともに端末装置に送信される。したがって、通信網での伝送遅延の大小に関わらず、目標楽曲のうち対象音の進行に応じた部分を特定できる(対象音の進行に同期するように楽曲ファイルの目標楽曲を再生できる)という利点がある。
<態様6>
 本発明の好適な態様(態様6)に係る楽曲処理装置は、音高が時間的に変化する対象音の音高に応じた音高情報を利用した楽曲検索処理により、対象音に対応する目標楽曲を当該対象音の進行に並行して検索する楽曲検索部と、目標楽曲のうち当該対象音の進行に対応した部分から対象音の進行に並行して目標楽曲を再生装置に再生させる再生制御部とを具備する。態様6では、対象音に対応する目標楽曲が対象音の進行に並行して検索され、かつ、目標楽曲のうち対象音の進行に対応した部分から当該目標楽曲が当該対象音の進行に並行して再生される。すなわち、楽曲の一連の対象音の進行に並行して目標楽曲の検索と当該目標楽曲の再生とが順次に実行される。したがって、目標楽曲の検索と当該目標楽曲の再生に並行した対象音の発音(例えば目標楽曲の歌唱または演奏)とを所望する利用者の負担を軽減することが可能である。
100……楽曲処理システム、12……端末装置、14……楽曲処理装置、16……通信網、20,50……制御装置、22,52……記憶装置、24,54……通信装置、26……表示装置、28……操作装置、32……収音装置、34……再生装置、42……音響解析部、44……再生処理部、50……制御装置、52……記憶装置、54……通信装置、62……楽曲検索部、64……再生制御部。

Claims (6)

  1.  コンピュータシステムが、
     音高が時間的に変化する対象音の音高に応じた音高情報を利用した楽曲検索処理により、前記対象音に対応する目標楽曲を当該対象音の進行に並行して検索し、
     前記目標楽曲のうち当該対象音の進行に対応した部分から前記対象音の進行に並行して前記目標楽曲を再生装置に再生させる
     楽曲処理方法。
  2.  前記目標楽曲の検索においては、前記楽曲検索処理の結果通知から所定時間が経過するまでに利用者から楽曲変更が指示された場合に前記楽曲検索処理を再実行し、
     前記目標楽曲の再生においては、前記楽曲変更が指示されずに前記所定時間が経過した場合に、直前の楽曲検索処理で検索された目標楽曲を前記再生装置に再生させる
     請求項1の楽曲処理方法。
  3.  前記楽曲検索処理は、
     複数の候補楽曲の各々について、前記対象音と当該候補楽曲との間で音高情報の時系列を対比することで、前記対象音における複数の音符と前記候補楽曲における複数の音符との時間的な対応と、前記対象音と当該候補楽曲との類否の指標である類似指標とを生成する第1処理と、
     前記各候補楽曲の類似指標に応じて前記複数の候補楽曲から前記目標楽曲を選択する第2処理とを含み、
     前記目標楽曲の再生においては、前記対象音と前記目標楽曲との間で前記第1処理により解析された音符毎の時間的な対応を参照して、前記対象音の発音速度に対応する当該目標楽曲の再生速度を設定し、前記目標楽曲を当該再生速度で前記再生装置に再生させる
     請求項1または請求項2の楽曲処理方法。
  4.  前記目標楽曲の再生においては、前記対象音と前記目標楽曲との間で前記第1処理により対応が特定された音符間の音高差に応じて当該目標楽曲の再生調を設定し、前記目標楽曲を当該再生調で前記再生装置に再生させる
     請求項3の楽曲処理方法。
  5.  前記目標楽曲の再生においては、前記目標楽曲を表す楽曲ファイルと、前記対象音と前記目標楽曲との間の時間的な対応を指定する時間対応情報と、前記目標楽曲の再生速度とを、前記再生装置を具備する端末装置に通信網を介して送信する
     請求項3または請求項4の楽曲処理方法。
  6.  音高が時間的に変化する対象音の音高に応じた音高情報を利用した楽曲検索処理により、前記対象音に対応する目標楽曲を当該対象音の進行に並行して検索する楽曲検索部と、
     前記目標楽曲のうち当該対象音の進行に対応した部分から前記対象音の進行に並行して前記目標楽曲を再生装置に再生させる再生制御部と
     を具備する楽曲処理装置。
PCT/JP2016/076266 2015-09-30 2016-09-07 楽曲処理方法および楽曲処理装置 Ceased WO2017056885A1 (ja)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2015-194638 2015-09-30
JP2015194638 2015-09-30

Publications (1)

Publication Number Publication Date
WO2017056885A1 true WO2017056885A1 (ja) 2017-04-06

Family

ID=58423331

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2016/076266 Ceased WO2017056885A1 (ja) 2015-09-30 2016-09-07 楽曲処理方法および楽曲処理装置

Country Status (1)

Country Link
WO (1) WO2017056885A1 (ja)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07140977A (ja) * 1993-11-19 1995-06-02 Roland Corp 自動演奏装置
JP2002244672A (ja) * 2001-02-15 2002-08-30 Daiichikosho Co Ltd カラオケ装置におけるマルチプル予約演奏制御システム
JP2013117688A (ja) * 2011-12-05 2013-06-13 Sony Corp 音響処理装置、音響処理方法、プログラム、記録媒体、サーバ装置、音響再生装置および音響処理システム

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07140977A (ja) * 1993-11-19 1995-06-02 Roland Corp 自動演奏装置
JP2002244672A (ja) * 2001-02-15 2002-08-30 Daiichikosho Co Ltd カラオケ装置におけるマルチプル予約演奏制御システム
JP2013117688A (ja) * 2011-12-05 2013-06-13 Sony Corp 音響処理装置、音響処理方法、プログラム、記録媒体、サーバ装置、音響再生装置および音響処理システム

Similar Documents

Publication Publication Date Title
JP6467887B2 (ja) 情報提供装置および情報提供方法
JP6452229B2 (ja) カラオケ効果音設定システム
JP6784022B2 (ja) 音声合成方法、音声合成制御方法、音声合成装置、音声合成制御装置およびプログラム
JPWO2017056982A1 (ja) 楽曲検索方法および楽曲検索装置
JP6252420B2 (ja) 音声合成装置、及び音声合成システム
JP7355165B2 (ja) 楽曲再生システム、楽曲再生システムの制御方法およびプログラム
JP5311069B2 (ja) 歌唱評価装置及び歌唱評価プログラム
JP6501344B2 (ja) 聴取者評価を考慮したカラオケ採点システム
JP6365483B2 (ja) カラオケ装置,カラオケシステム,及びプログラム
WO2014142200A1 (ja) 音声処理装置
JP6809177B2 (ja) 情報処理システムおよび情報処理方法
JP6415341B2 (ja) ハーモニー歌唱のためのピッチシフト機能を備えたカラオケシステム
WO2017056885A1 (ja) 楽曲処理方法および楽曲処理装置
JP6365561B2 (ja) カラオケシステム、カラオケ装置、及びプログラム
JP5983670B2 (ja) プログラム、情報処理装置、及びデータ生成方法
JP2020122948A (ja) カラオケ装置
JP4048249B2 (ja) カラオケ装置
JP6260565B2 (ja) 音声合成装置、及びプログラム
JP6406182B2 (ja) カラオケ装置、及びカラオケシステム
JP7679870B2 (ja) 信号処理システム、信号処理方法およびプログラム
JP6144593B2 (ja) 歌唱採点システム
JP5704201B2 (ja) カラオケ装置及びカラオケ楽曲処理プログラム
JP6057079B2 (ja) カラオケ装置及びカラオケ用プログラム
JP2009244790A (ja) 歌唱指導機能を備えるカラオケシステム
JP2006276560A (ja) 音楽再生装置および音楽再生方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16851059

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 16851059

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: JP