WO2016017622A1 - リファレンス表示装置、リファレンス表示方法およびプログラム - Google Patents

リファレンス表示装置、リファレンス表示方法およびプログラム Download PDF

Info

Publication number
WO2016017622A1
WO2016017622A1 PCT/JP2015/071342 JP2015071342W WO2016017622A1 WO 2016017622 A1 WO2016017622 A1 WO 2016017622A1 JP 2015071342 W JP2015071342 W JP 2015071342W WO 2016017622 A1 WO2016017622 A1 WO 2016017622A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
sound
timing
guide image
guide
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2015/071342
Other languages
English (en)
French (fr)
Inventor
紀行 畑
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Yamaha Corp
Original Assignee
Yamaha Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Yamaha Corp filed Critical Yamaha Corp
Priority to US15/500,011 priority Critical patent/US10332496B2/en
Publication of WO2016017622A1 publication Critical patent/WO2016017622A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H1/00Details of electrophonic musical instruments
    • G10H1/36Accompaniment arrangements
    • G10H1/361Recording/reproducing of accompaniment for use with an external source, e.g. karaoke systems
    • G10H1/368Recording/reproducing of accompaniment for use with an external source, e.g. karaoke systems displaying animated or moving pictures synchronized with the music or audio part
    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09BEDUCATIONAL OR DEMONSTRATION APPLIANCES; APPLIANCES FOR TEACHING, OR COMMUNICATING WITH, THE BLIND, DEAF OR MUTE; MODELS; PLANETARIA; GLOBES; MAPS; DIAGRAMS
    • G09B15/00Teaching music
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10GREPRESENTATION OF MUSIC; RECORDING MUSIC IN NOTATION FORM; ACCESSORIES FOR MUSIC OR MUSICAL INSTRUMENTS NOT OTHERWISE PROVIDED FOR, e.g. SUPPORTS
    • G10G1/00Means for the representation of music
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2210/00Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
    • G10H2210/005Musical accompaniment, i.e. complete instrumental rhythm synthesis added to a performed melody, e.g. as output by drum machines
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2210/00Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
    • G10H2210/031Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal
    • G10H2210/091Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal for performance evaluation, i.e. judging, grading or scoring the musical qualities or faithfulness of a performance, e.g. with respect to pitch, tempo or other timings of a reference performance
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2220/00Input/output interfacing specifically adapted for electrophonic musical tools or instruments
    • G10H2220/005Non-interactive screen display of musical or status data
    • G10H2220/011Lyrics displays, e.g. for karaoke applications
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2220/00Input/output interfacing specifically adapted for electrophonic musical tools or instruments
    • G10H2220/005Non-interactive screen display of musical or status data
    • G10H2220/015Musical staff, tablature or score displays, e.g. for score reading during a performance

Definitions

  • the present invention relates to a display device, and more particularly to a device for displaying a model (reference).
  • a karaoke apparatus displays lyrics and a model pitch on a display unit (see, for example, Patent Document 1).
  • the pitch is displayed as a so-called piano roll.
  • a piano roll is a linear image on the screen where the vertical axis represents the scale (the piano keyboard is vertical) and the horizontal axis corresponds to the time. Is displayed. Thereby, the singer can grasp
  • each sound is not pronounced discretely, but each sound is smoothly connected or breathed.
  • a conventional piano roll it is difficult to intuitively understand the connection of each sound and the timing of breathing even though the sound generation start timing and stop timing of each sound can be grasped.
  • an object of the present invention is to provide a reference display device, a reference display method, and a program that can intuitively grasp the connection of each sound and the timing of breathing.
  • the reference display device of the present invention includes a display unit, and an image generation unit that generates a guide image indicating a sound generation timing, a pitch, and a sound generation length based on the reference data, and displays the guide image on the display unit.
  • the image generation means displays a guide image in which the sounds in the reference data are connected.
  • the user can visually grasp how to connect each sound.
  • the reference data includes information indicating breathing timing
  • the image generation means further displays a guide image in which sounds before and after the breathing timing are discrete based on the information indicating the breathing timing.
  • the reference display device includes information indicating breathing timing in the reference data, it is possible to clearly distinguish and display a silent section and a breathing section, which is accurate for the user. The position of breathing can be grasped.
  • the image generation means can display a guide image in which a phoneme related to the prompting sound and a phoneme next to the phoneme related to the prompting sound are discrete.
  • the image generation means displays an image indicating that a phoneme related to the sound is present together with a guide image of the phoneme related to the sound. In this case, the user can grasp whether the sound is generated by connecting sounds more easily or whether the sound is generated by stopping the sound.
  • the image generation means displays an image for prompting breathing together with the guide image based on information indicating the breathing timing. Thereby, the user can grasp
  • the image generation means displays an image indicating a sound generation timing of each sound so as to be superimposed on the guide image. For example, when different lyrics are continuously generated with the same pitch, if the guide images are connected, it is difficult for the user to grasp at which timing the next lyrics are to be generated. However, if, for example, a circular image is superimposed and displayed on the guide image at the sound generation timing of each sound, the user can easily understand that the sound is generated at the timing of the circular image.
  • the reference data includes information indicating the timing of the singing technique
  • the image generation unit displays an image prompting the singing technique based on the information indicating the timing of the singing technique.
  • the vibrato section can be more intuitively understood by changing the guide image to a different image (for example, a wavy line) and displaying it.
  • the reference data includes information indicating the volume of each sound
  • the image generation unit changes the guide image to an image corresponding to the volume based on the information indicating the volume of each sound. It is preferable to display. For example, a section with high volume is changed to a thick line, and a section with low volume is changed to a thin line. Alternatively, the section with high volume is changed to a dark line, and the section with low volume is changed to a light line.
  • the image generation means displays an image corresponding to the user (a photograph taken by the user, a character image, etc.) at a position corresponding to the current pronunciation timing, and the image corresponding to the user moves along the guide image.
  • the guide image may be scrolled. In this case, the user can feel as if the character is moving in accordance with his / her pronunciation, and can enjoy singing, language learning, or playing.
  • the character image may be an objective viewpoint (two-dimensional display) or a subjective viewpoint (three-dimensional display). Moreover, when displaying from a subjective viewpoint, for example, when performing a duet singing, it is possible to display a character corresponding to itself and a character corresponding to another singer in parallel. You can feel the atmosphere of singing with other singers.
  • the reference display device or the reference display method of the present invention can intuitively grasp the connection of each sound and the timing of breathing.
  • FIG. 1 is a diagram showing a configuration of a karaoke system provided with a reference display device of the present invention.
  • the karaoke system includes a center (server) 1 connected via a network 2 such as the Internet and a plurality of karaoke stores 3.
  • Each karaoke store 3 is provided with a relay device 5 such as a router connected to the network 2 and a plurality of karaoke devices 7 connected to the network 2 via the relay device 5.
  • the repeater 5 is installed in a management room of a karaoke store.
  • a plurality of karaoke apparatuses 7 are installed in each private room (karaoke box).
  • Each karaoke device 7 is provided with a remote controller 9.
  • the karaoke device 7 can communicate with other karaoke devices 7 via the relay 5 and the network 2.
  • a karaoke system can communicate between karaoke apparatuses 7 installed in different places, and can perform a duet among a plurality of singers.
  • FIG. 2 is a block diagram showing the configuration of the karaoke apparatus.
  • the karaoke device 7 corresponds to a reference display device of the present invention.
  • the karaoke device 7 includes a CPU 11, a RAM 12, an HDD 13, a network interface (I / F) 14, an LCD (touch panel) 15, a microphone 16, an A / D converter 17, a sound source 18, a mixer (effector) 19, and a sound system (SS) 20.
  • the CPU 11 that controls the operation of the entire apparatus includes a RAM 12, an HDD 13, a network interface (I / F) 14, an LCD (touch panel) 15, an A / D converter 17, a sound source 18, a mixer (effector) 19, and a decoder 22 such as an MPEG.
  • the display processing unit 23, the operation unit 25, and the transmission / reception unit 26 are connected.
  • the HDD 13 stores an operation program for the CPU 11.
  • the RAM 12 as a work memory, an area to be read for executing the operation program of the CPU 11, an area to read music data to play karaoke music, an area to read reference data such as a guide melody, a reservation list, a scoring result, etc.
  • An area for temporarily storing the data is set.
  • the HDD 13 stores music data for playing karaoke music. Further, the HDD 13 also stores video data for displaying a background video on the monitor 24. Video data stores both moving images and still images. Music data and video data are periodically distributed from the center 1 and updated.
  • the CPU 11 is a control unit that comprehensively controls the karaoke apparatus and functionally incorporates a sequencer to perform karaoke performance. Further, the CPU 11 performs an audio signal generation process, a video signal generation process, a scoring process, and a piano roll display process. Thereby, CPU11 functions as an image generation means in the present invention.
  • the touch panel 15 and the operation unit 25 are provided on the front surface of the karaoke apparatus.
  • the CPU 11 displays an image corresponding to the operation information on the touch panel 15 based on the operation information input from the touch panel 15 to realize a GUI.
  • the remote controller 9 also realizes the same GUI.
  • the CPU 11 performs various operations based on operation information input from the remote controller 9 via the touch panel 15, the operation unit 25, or the transmission / reception unit 26.
  • the CPU 11 functionally includes a sequencer.
  • the CPU 11 reads music data corresponding to the music number of the reserved music registered in the reserved list in the RAM 12 from the HDD 13, and performs a karaoke performance with the sequencer.
  • the music data includes a header in which a music number is written, a musical sound track in which performance MIDI data is written, a guide melody track in which MIDI data for guide melody is written, and lyrics It consists of a lyrics track in which MIDI data is written, a chorus track in which back chorus playback timing and audio data to be played are written, a breath position track indicating the timing of breathing, a technique position track indicating the timing of the singing technique, and the like.
  • the guide melody track, breath position track, and technique position track correspond to the reference data of the present invention.
  • Reference data is model data for a singer to use as a reference for singing, and includes information indicating the sounding timing, pitch, and sounding length of each sound.
  • the format of the music data is not limited to this example.
  • the format of the reference data is not limited to the MIDI format as described above.
  • the reference data indicating the breath position may be text data indicating the timing of the breath position (elapsed time from the beginning of the music).
  • voice data for example, recorded singing sound
  • the pitch is extracted from the voice data to extract the pitch
  • the sound is generated from the timing and length at which the pitch is extracted. It is also possible to extract timing and pronunciation length.
  • It is also possible to detect the silent section by detecting the volume (power), and when there is a silent section between each sound, the timing at which the silent section is extracted can be extracted as the timing of the breath position It is.
  • the timing (technical position) at which the singing technique is performed is extracted by determining that “vibrato” is being performed for the institution. It is also possible.
  • the musical sound track stores information indicating the type, timing, pitch (key), strength, length, localization (pan), sound effect (effect), etc. of the musical instrument that generates the musical sound.
  • information such as the sounding start timing of each sound corresponding to the model song and the length of the sounding are recorded.
  • the sequencer controls the sound source 18 based on the data of the musical tone track, and generates the musical tone of the karaoke song.
  • the sequencer plays back chorus audio data (compressed audio data such as MP3 attached to music data) at the timing specified by the chorus track. Further, the sequencer synthesizes the character pattern of the lyrics in synchronism with the progress of the song based on the lyrics track, converts the character pattern into a video signal, and inputs it to the display processing unit 23.
  • chorus audio data compressed audio data such as MP3 attached to music data
  • the sound source 18 forms a musical sound signal (digital audio signal) according to data (note event data) input from the CPU 11 by processing of the sequencer.
  • the formed tone signal is input to the mixer 19.
  • the mixer 19 generates an acoustic effect such as an echo on the musical sound signal generated by the sound source 18, the chorus sound, and the singing voice signal of the singer input from the microphone (singing voice input means) 16 via the A / D converter 17. And mixing these signals.
  • a singing voice signal is transmitted from another karaoke device.
  • the singing voice signal received from the other karaoke apparatus is also input to the mixer 19 and mixed with the singing voice signal input from the microphone 16 of the own apparatus.
  • the sound system 20 includes a D / A converter and a power amplifier, converts an input digital signal into an analog signal, amplifies it, and emits the sound from a speaker (musical sound generating means) 21.
  • the effect that the mixer 19 gives to each audio signal and the balance of mixing are controlled by the CPU 11.
  • the CPU 11 reads the video data stored in the HDD 13 and reproduces the background video and the like in synchronism with the generation of musical sounds and the generation of the lyrics telop by the sequencer.
  • the video data of the moving image is encoded in the MPEG format.
  • the CPU 11 can also download picture data representing a singer or video data such as a character from the center 1 and input it to the display processing unit 23.
  • the photograph representing the singer can be taken on the spot with a camera (not shown) provided on the karaoke device or the remote control 9, or with a camera provided on a mobile terminal owned by the user. is there.
  • the CPU 11 inputs the read video data of the background video to the decoder 22.
  • the decoder 22 converts the input data such as MPEG into a video signal and inputs it to the display processing unit 23.
  • the display processing unit 23 receives a piano roll video signal based on the guide melody track, together with the character pattern of the lyrics telop.
  • FIG. 4A-4C are diagrams showing an example of a piano roll.
  • the piano roll has a scale (the piano keyboard is in a vertical position) on the vertical axis and a time on the screen corresponding to the time on the horizontal axis.
  • a corresponding linear image is displayed as a guide image.
  • the singer can grasp
  • the guide images of each sound are connected and displayed, and the guide images are displayed discretely at the breathing timing.
  • the CPU 11 first generates a guide image based on information on the sound generation start timing and the sound generation length of each sound included in the guide melody track. Then, the CPU 11 smoothly connects the guide images for each sound.
  • the inclination of the connection part of each sound is displayed as an image having the same inclination uniformly, for example, by connecting with a time length corresponding to a sixteenth note.
  • the reference data may include information that specifies an inclination according to a change in the pitch of the connected portion of each sound.
  • the CPU 11 disperses the guide image at the breathing timing indicated by the breath position track. For example, in the example of FIG. 4A, since there is a breathing timing between the pronunciation of “Akai” and the beginning of “Hanaga”, a guide image of “Akai” and “Hanaga” Are separated from the guide image.
  • FIG. 3 shows an example in which one piece of music data includes a breath position track and a technique position track.
  • the existing music data remains as it is, and the breath position track and the technique position track are different data. You may prepare. In this case, it is not necessary to prepare new music data including the breath position track and the technique position track.
  • song identification information such as a song number is described in the data of the breath position track and the technique position track.
  • the CPU 11 reads the corresponding breath position track and technique position track, and performs a sequence operation.
  • the CPU 11 displays, for example, a circular image superimposed on the guide image at the sound generation timing of each sound in the guide melody track. As a result, the user can grasp that the sound is generated at the timing when the circle image is displayed.
  • FIG. 4C shows a mode in which the guide image related to the phoneme of the prompt sound and the guide image after the phoneme of the prompt sound are displayed separately.
  • the prompt sound is represented by “t” or “t” in Japanese kana notation, and is silent between the following sound.
  • the CPU 11 extracts the prompt sound from the lyrics track, and disperses the guide image at the timing of emitting the extracted prompt sound.
  • the guide image of “ka” and the guide image of “ta” after that are displayed separately.
  • an image indicating that a phoneme related to the prompt sound is present in this example, a square image described with “tsu” is displayed.
  • the singer can easily grasp whether the sound is connected and pronounced or whether the sound is interrupted due to the presence of a prompt sound.
  • FIG. 5A is an example in which the volume is represented by a guide image.
  • the guide melody track includes information indicating the volume of each sound.
  • the CPU 11 changes the line thickness of the guide image based on information indicating the volume of each sound included in the guide melody track. For example, in the example of FIG. 5A, among “Akai”, the sound of “A” has the highest volume, so the section representing “A” is changed to a thick line. Since the volume of the “ka” sound is low, the section representing “ka” is changed to a thin line.
  • the information indicating the volume included in the guide melody track has three levels of “large”, “standard”, and “small”, and the thickness of the line is changed to three levels.
  • the thickness of the line may be changed in multiple stages.
  • the thickness of the line is changed at the intermediate position of the connected portion.
  • the thickness of the line may be changed at the beginning of each sound or at the end of each sound.
  • the volume may be changed so that the thickness of the line gradually changes from the end of each sound to the beginning of the next sound.
  • FIG. 5B is a mode in which the volume is further expressed by the color of the line.
  • the CPU 11 changes the line color of the guide image based on information indicating the volume of each sound included in the guide melody track. For example, in the example of FIG. 5B, since the volume of the sound “ka” is low, the section representing “ka” is changed to a light line. In this example, when the information indicating the volume included in the guide melody track is “low”, the section of the “low” sound is changed to a light color. The section may be changed to a dark color, or only the color may be changed without changing the thickness of the line.
  • FIG. 6 shows an example in which the singing technique is displayed on the piano roll.
  • the guide image of the vibrato section is changed to a wavy line and displayed.
  • the CPU 11 reads out information indicating the vibrato timing included in the technique position track, and changes the section from the timing to the end of the sound guide image corresponding to the timing to a wavy line. This makes it easier for the singer to grasp the vibrato timing and vibrato length more intuitively.
  • the head of the sound corresponding to the “time” information in this example, the “ka” sound. Delay position. This makes it easier for the singer to intuitively understand the “tame” that intentionally delays the singing of the specific sound (“ka” sound in this example).
  • the guide image is raised from a pitch lower than the reference pitch to the original pitch.
  • the sound of “a” and the sound of “ha” at the beginning are “shrunk”, and the song is sung while being lifted from a pitch lower than the reference pitch. Therefore, the beginning of the section “A” is a guide image that raises the pitch lower than the reference pitch to the original pitch.
  • the guide image is temporarily raised at the position corresponding to the timing of “Kobushi” as shown by the sound “N” in FIG. 6B.
  • “Kobushi” is a singing technique for changing the voice color of a specific sound so as to beat in the middle of pronunciation.
  • the guide image is changed from the position corresponding to the “fall” timing to a lower pitch as shown by the “ga” sound in FIG. 6B. It is also possible to make it a guide image corresponding to the “fall” singing technique by adopting the mode of making it happen.
  • FIG. 6C is an example of displaying an image that prompts a singing technique.
  • CPU11 reads the information which shows the timing of the various singing techniques contained in the technique position track, and displays the image which prompts the singing technique in the position corresponding to the said timing.
  • “A” at the beginning is a mess, and is a part where the song is sung while being lifted from a pitch lower than the reference pitch. Therefore, an image pronounced of a rise in pitch such as “no” is displayed at the beginning of the section “a”.
  • a wavy image such as “m” is separately displayed. Thereby, the singer can grasp
  • the CPU 11 displays an image for prompting breathing together with a guide image based on information indicating breathing timing indicated by the breath position track. For example, in the example of FIG. 6C, an image such as “V” is displayed between the “Akai” section and the “Hanaga” section. Thereby, the singer can grasp
  • a karaoke performance is performed, and a piano roll is displayed as the performance progresses.
  • a singer can connect each sound smoothly while watching the guide image and sing or breathe. Easy to do.
  • the scoring process is performed by comparing the singing voice of the singer with the guide melody track. Scoring is performed by comparing the pitch of the singing voice and the guide melody for each note of the guide melody track. That is, a high score is given when the pitch of the singing voice matches the pitch of the guide melody track for a predetermined time or longer (within the allowable range). The timing of pitch change is also taken into account. Further, points are added based on the presence or absence of singing techniques such as timing of pitch change, vibrato, inflection, and squealing (moving gently from a low pitch).
  • whether or not the singer has performed breath breathing at the breath breathing timing included in the breath position track is also regarded as a point to be scored. Whether or not breathing has been performed is determined when no sound is input from the microphone 16 within a predetermined time including the breathing timing (the input level is less than a predetermined threshold) or when a breathing sound is input from the microphone 16. When the voice is input from the microphone 16 (the input level is equal to or higher than the predetermined threshold value), it is determined that breathing is not performed. Note that whether or not a breathing sound has been collected is determined by comparing with the waveform of the breathing sound by pattern matching or the like, for example.
  • the scoring process may be performed at each karaoke apparatus, but may be performed at the center 1 (or another server). Moreover, when performing duet with another karaoke apparatus via a network, you may perform a scoring process with one karaoke apparatus which processes typically.
  • FIG. 7 is a flowchart showing the operation of the karaoke system.
  • the singer makes a music request (s11).
  • the CPU 11 displays an image prompting whether or not to perform a duet with a singer of another karaoke apparatus connected via the network on the monitor 24, Accept duet singing.
  • a duet song For example, when a singer inputs the name of a specific user using the touch panel 15, the operation unit 25, or the remote control 9, the user related to the name is searched at the center 1 and set as a duet partner.
  • the CPU 11 of the karaoke device reads the requested music data (s12) and creates a piano roll (s13). That is, the CPU 11 generates a guide image based on the information on the sound generation start timing and the sound generation length of each sound included in the guide melody track.
  • the CPU 11 reads out the lyrics track (s14), and associates the lyrics image with each guide image (s15). Further, the CPU 11 reads out breath timing information from the breath position track and reads out the sound generation timing related to the phoneme of the prompt sound (s16). Then, the CPU 11 smoothly connects the guide images for each sound (s17). At this time, the CPU 11 displays the guide images in a discrete manner at the breathing timing indicated by the breath position track and the sounding timing related to the phoneme of the prompt sound.
  • the CPU 11 reads the technique position track (s18) and displays the singing technique on the piano roll (s19). Moreover, CPU11 changes a guide image according to a song technique. For example, the vibrato section is displayed with the guide image changed to a wavy line.
  • the CPU 11 reads information indicating the volume of each sound included in the guide melody track (s20), and changes to a guide image corresponding to the volume of each sound (s21). For example, as shown in FIG. 5A, the thickness of the line is changed according to the volume, or the color of the line is changed according to the volume as shown in FIG. 5B.
  • the aspect which performs a karaoke performance and a piano roll display using the karaoke apparatus 7 was shown, information processing apparatuses (a microphone, a speaker, and a personal computer, a smart phone, a game machine, etc. which a user owns), for example
  • the display device of the present invention can also be realized by using a display device having a structure of a display portion. Note that the music data and the reference data need not be stored in the display device, and may be downloaded from the server each time and used.
  • the CPU 11 may display the character at the current singing position as shown in FIG.
  • the guide image corresponds to the ground image
  • the ground image is interrupted at the breath connection timing.
  • the guide image and the background are scrolled so that the character image 101 moves along the guide image (ground image). Since the ground is interrupted at the breath connection timing, the character image 101 falls from the ground when the breath connection (the state in which no sound is input from the microphone 16) is not detected.
  • the result of the singing score is displayed on the screen. Therefore, the singer can enjoy karaoke like a game.
  • the guide image may be displayed in an objective viewpoint (two-dimensional display, two-dimensional viewpoint) as shown in FIG. 8, but for example, a subjective viewpoint (three-dimensional display, three-dimensional display) as shown in FIG. (Viewpoint) may be displayed.
  • the subjective viewpoint is a display mode that imitates the user's own visual field, and is a kind of three-dimensional viewpoint.
  • a display mode in which the depth direction corresponds to the time axis and the plane direction corresponds to the pitch is shown.
  • the display mode is such that the depth direction corresponds to time and the vertical direction corresponds to a musical scale.
  • FIG. 9 the display mode is such that the depth direction corresponds to time and the vertical direction corresponds to a musical scale.
  • an image (character image or the like) corresponding to the user himself / herself is displayed, and the character image or the like is displayed from behind, and the display mode also corresponds to a three-dimensional viewpoint.
  • the scale may correspond to the left-right direction.
  • the guide image and the background are scrolled so that the character image 101A moves in the depth direction along the guide image.
  • the result of the singing score is displayed on the screen. Therefore, the singer can enjoy karaoke like a game.
  • the change of the pitch of the model of a wind instrument performance is displayed as a guide image, and it is the same also as an aspect which a guide image interrupts at a breathing timing.
  • the effect is obtained.
  • the same effect can be obtained by displaying a guide image indicating the pronunciation timing and length of the model and interrupting the guide image at the breath timing and the prompt sound.
  • an example of displaying lyrics is described, but display of lyrics is not essential in the present invention.
  • the reference display device of the present invention includes a monitor 24 that is a display unit and a CPU 11 that functions as an image generation unit that performs a guide image display process.
  • the CPU 11 stores the information in the HDD 13.
  • a guide image may be generated based on the music data (which is an example of the reference data of the present invention) and the sounds of the guide image may be connected.
  • Other hardware configurations are not essential elements in the present invention.
  • the reference data does not need to be stored in the HDD 13 and may be downloaded and used from the outside (for example, a server) each time.
  • the decoder 22, the display processing unit 23, and the RAM 12 may be incorporated in the CPU 11 as a part of the function of the CPU 11.
  • the guide image is not necessarily displayed as a piano roll (the vertical axis corresponds to the piano keyboard and a solid line is displayed along the horizontal axis direction).
  • any display mode may be used as long as it generates a guide image indicating a sound generation timing, a pitch, and a sound generation length and connects the sounds.
  • the guide image referred to in the present invention is not limited to the elongated line shown in FIGS. 4A to 6, and an image having a certain width in the left-right or vertical direction as shown in the example of FIG. 9. Those extending in one direction (the depth direction in the example of FIG. 9) are also included.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Acoustics & Sound (AREA)
  • Business, Economics & Management (AREA)
  • Educational Administration (AREA)
  • Educational Technology (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Reverberation, Karaoke And Other Acoustics (AREA)
  • Auxiliary Devices For Music (AREA)

Abstract

 各音のつながりおよび息継ぎのタイミングを直感的に把握することができる表示装置およびプログラムを提供する。CPU(11)は、ガイドメロディトラックに含まれている各音の発音開始タイミングおよび発音の長さの情報に基づいて、ガイド画像を生成する。そして、CPU(11)は、各音を滑らかに連結する。その後、CPU(11)は、ブレス位置トラックが示す息継ぎタイミングで音を離散させる。

Description

リファレンス表示装置、リファレンス表示方法およびプログラム
 本発明は、表示装置に関し、特にお手本(リファレンス)の表示を行う装置に関する。
 従来、カラオケ装置は、歌詞やお手本の音程を表示部に表示することが行われている(例えば特許文献1を参照)。音程は、いわゆるピアノロールとして表示される。ピアノロールとは、縦軸が音階(ピアノの鍵盤が縦になった状態)、横軸が時間に対応した画面上に、各音の発音開始タイミングと発音の長さに応じた線状の画像を表示するものである。これにより、歌唱者は、歌唱するタイミングと音程を視覚的に把握することができる。
日本国特開2004-205817号公報
 歌唱、会話、または吹奏楽器の演奏等では、各音を離散的に発音するのではなく、各音を滑らかにつなげたり、息継ぎを行ったりする。従来のピアノロールでは、各音の発音開始タイミングと停止タイミングとを把握することができても、各音のつながりおよび息継ぎのタイミングを直感的に把握することが困難であった。
 そこで、本発明は、各音のつながりおよび息継ぎのタイミングを直感的に把握することができるリファレンス表示装置、リファレンス表示方法およびプログラムを提供することを限定されないひとつの目的とする。
 本発明のリファレンス表示装置は、表示部と、リファレンスデータに基づいて、発音タイミングと音程と発音長を示すガイド画像を生成し、前記表示部に表示する画像生成手段と、を備えている。
 そして、前記画像生成手段は、前記リファレンスデータ中の各音を連結させたガイド画像を表示することを特徴とする。
 これにより、ユーザ(歌唱者、話者、または演奏者等)は、各音のつなげ方を視覚的に把握することができる。
 また、リファレンスデータには、息継ぎタイミングを示す情報が含まれ、前記画像生成手段は、さらに前記息継ぎタイミングを示す情報に基づいて、該息継ぎタイミングの前後の音を離散させたガイド画像を表示する。
 従来のピアノロールでは、各音の発音開始タイミングと停止タイミングとを把握することができても、各音のつながりおよび息継ぎのタイミングを直感的に把握することが困難であった。そこで、息継ぎタイミングでガイド画像を途切れさせることで、ユーザは、各音のつながりおよび息継ぎのタイミングを直感的に把握することができる。また、本発明のリファレンス表示装置は、リファレンスデータに息継ぎタイミングを示す情報が含まれているため、単なる無音区間と息継ぎの区間とを明確に区別して表示させることができ、ユーザに対して正確な息継ぎの位置を把握させることができる。
 また、前記画像生成手段は、促音に係る音素と、該促音に係る音素の次の音素を離散させたガイド画像を表示することも可能である。
 促音は、日本語のかな表記では「っ」「ッ」で表されるものであり、後に続く音との間が無音となるものである。したがって、促音に係る音素とその次の音素を離散させることで、ユーザは、音をつなげて発音するのか、一旦止めて発音するのか、直感的に把握することができる。
 また、前記画像生成手段は、前記促音に係る音素のガイド画像とともに、促音に係る音素が存在する旨を示す画像を表示することが望ましい。この場合、ユーザは、より容易に音をつなげて発音するのか、一旦止めて発音するのか、把握することができる。
 また、前記画像生成手段は、前記息継ぎタイミングを示す情報に基づいて、息継ぎを促す画像を、前記ガイド画像とともに表示させることが望ましい。これにより、ユーザは、より容易に音をつなげて発音するのか、息継ぎを行うのか、把握することができる。
 また、前記画像生成手段は、各音の発音タイミングを示す画像を、前記ガイド画像に重畳して表示することが好ましい。例えば同じ音程で異なる歌詞を続けて発音する場合、ガイド画像が連結されていると、ユーザは、どのタイミングで次の歌詞を発音するのか把握し難い。しかし、各音の発音タイミングにおいて例えばガイド画像の上に円画像を重畳して表示すれば、ユーザは、当該円画像のタイミングで発音を行う旨を把握し易くなる。
 また、前記リファレンスデータには、歌唱技法のタイミングを示す情報が含まれ、前記画像生成手段は、前記歌唱技法のタイミングを示す情報に基づいて、歌唱技法を促す画像を表示することが好ましい。これにより、ユーザは、歌唱技法を行うタイミングを容易に把握することができる。
 また、ビブラートの区間は、ガイド画像を異なる画像(例えば波線)に変更して表示することで、より直感的にビブラート区間を把握し易くすることができる。
 また、リファレンスデータには、各音の音量を示す情報が含まれ、前記画像生成手段は、前記各音の音量を示す情報に基づいて、前記ガイド画像を前記音量に応じた画像に変更して表示することが好ましい。例えば、音量の大きい区間は太い線、音量の小さい区間は細い線に変更する。あるいは、音量の大きい区間は濃い色の線、音量の小さい区間は薄い色の線に変更する。
 なお、画像生成手段は、現在の発音タイミングに応じた位置にユーザに対応する画像(ユーザを撮影した写真、キャラクタ画像等)を表示し、該ユーザに対応する画像が前記ガイド画像に沿って移動するように、前記ガイド画像をスクロールさせる態様としてもよい。この場合、ユーザは、自身の発音に応じてキャラクタを移動させているように感じることができ、歌唱、語学学習、または演奏等を楽しんで行うことができる。
 また、キャラクタ画像は、客観視点(2次元表示)であってもよいが、主観視点(3次元表示)であってもよい。また、主観視点で表示する場合には、例えばデュエット歌唱を行う場合に、自身に相当するキャラクタと他の歌唱者に相当するキャラクタとを並行して表示することも可能であり、ユーザは、他の歌唱者と一緒に歌唱を行っている雰囲気をより感じ取ることができる。
 本発明のリファレンス表示装置またはリファレンス表示方法は、各音のつながりおよび息継ぎのタイミングを直感的に把握することができる。
カラオケシステムの構成を示したブロック図である。 カラオケ装置の構成を示したブロック図である。 リファレンスデータを含む各種データの構造を示す図である。 リファレンスの表示例を示す図である。 リファレンスの表示例を示す図である。 リファレンスの表示例を示す図である。 リファレンスの表示例を示す図である。 リファレンスの表示例を示す図である。 リファレンスの表示例を示す図である。 リファレンスの表示例を示す図である。 リファレンスの表示例を示す図である。 カラオケ装置の動作を示すフローチャートである。 応用例に係るリファレンスの表示態様である。 応用例に係るリファレンスの表示態様である。 リファレンス表示装置の最小構成を示したブロック図である。
 図1は、本発明のリファレンス表示装置を備えたカラオケシステムの構成を示す図である。カラオケシステムは、インターネット等のネットワーク2を介して接続されるセンタ(サーバ)1と、複数のカラオケ店舗3と、からなる。
 各カラオケ店舗3には、ネットワーク2に接続されるルータ等の中継機5と、中継機5を介してネットワーク2に接続される複数のカラオケ装置7が設けられている。中継機5は、カラオケ店舗の管理室内等に設置されている。複数台のカラオケ装置7は、それぞれ個室(カラオケボックス)に1台ずつ設置されている。また、各カラオケ装置7には、それぞれリモコン9が設置されている。
 カラオケ装置7は、中継機5およびネットワーク2を介して他のカラオケ装置7と通信可能になっている。カラオケシステムは、異なる場所に設置されているカラオケ装置7同士で通信を行い、複数の歌唱者間でデュエットを行うことができる。
 図2は、カラオケ装置の構成を示すブロック図である。カラオケ装置7は、本発明のリファレンス表示装置に相当する。カラオケ装置7は、CPU11、RAM12、HDD13、ネットワークインタフェース(I/F)14、LCD(タッチパネル)15、マイク16、A/Dコンバータ17、音源18、ミキサ(エフェクタ)19、サウンドシステム(SS)20、スピーカ21、MPEG等のデコーダ22、表示処理部23、モニタ24、操作部25、および送受信部26を備えている。
 装置全体の動作を制御するCPU11には、RAM12、HDD13、ネットワークインタフェース(I/F)14、LCD(タッチパネル)15、A/Dコンバータ17、音源18、ミキサ(エフェクタ)19、MPEG等のデコーダ22、表示処理部23、操作部25、および送受信部26が接続されている。
 HDD13は、CPU11の動作用プログラムが記憶されている。ワークメモリであるRAM12には、CPU11の動作用プログラムを実行するために読み出すエリア、カラオケ曲を演奏するために楽曲データを読み出すエリア、ガイドメロディ等のリファレンスデータを読み出すエリア、予約リストや採点結果等のデータを一時記憶するエリア、等が設定される。
 また、HDD13は、カラオケ曲を演奏するための楽曲データを記憶している。さらに、HDD13は、モニタ24に背景映像を表示するための映像データも記憶している。映像データは動画、静止画の両方を記憶している。楽曲データや映像データは、定期的にセンタ1から配信され、更新される。
 CPU11は、カラオケ装置を統括的に制御する制御部であり、機能的にシーケンサを内蔵し、カラオケ演奏を行う。また、CPU11は、音声信号生成処理、映像信号生成処理、採点処理、およびピアノロール表示処理を行う。これにより、CPU11は、本発明における画像生成手段として機能する。
 タッチパネル15および操作部25は、カラオケ装置の前面に設けられている。CPU11は、タッチパネル15から入力される操作情報に基づいて、操作情報に応じた画像をタッチパネル15上に表示し、GUIを実現する。また、リモコン9も同様のGUIを実現するものである。CPU11は、タッチパネル15、操作部25、または送受信部26を介してリモコン9から入力される操作情報に基づいて、各種の動作を行う。
 次に、カラオケ演奏を行うための構成について説明する。上述したように、CPU11は、機能的にシーケンサを内蔵している。CPU11は、RAM12の予約リストに登録された予約曲の曲番号に対応する楽曲データをHDD13から読み出し、シーケンサでカラオケ演奏を行う。
 楽曲データは、例えば図3に示すように、曲番号等が書き込まれているヘッダ、演奏用MIDIデータが書き込まれている楽音トラック、ガイドメロディ用MIDIデータが書き込まれているガイドメロディトラック、歌詞用MIDIデータが書き込まれている歌詞トラック、バックコーラス再生タイミングおよび再生すべき音声データが書き込まれているコーラストラック、息継ぎのタイミングを示すブレス位置トラック、歌唱技法のタイミングを示す技法位置トラック、等からなっている。ガイドメロディトラック、ブレス位置トラック、および技法位置トラックは、本発明のリファレンスデータに対応する。リファレンスデータとは、歌唱者が歌唱の参考にするためのお手本データであり、各音を発する発音タイミング、音程、および発音長を示す情報が含まれている。なお、楽曲データの形式としては、この例に限るものではない。また、リファレンスデータの形式も、上述のようなMIDI形式に限るものではない。例えばブレス位置を示すリファレンスデータとしては、ブレス位置のタイミング(楽曲先頭からの時間経過)を示したテキストデータ等であってもよい。また、リファレンスデータが音声データ(例えば歌唱音を録音したもの)である場合には、当該音声データからピッチを抽出して音程を抽出するとともに、該音程が抽出されるタイミングおよび長さから、発音タイミングおよび発音長を抽出することも可能である。また、音量(パワー)を検出することで無音区間を検出し、各音の間に無音区間が存在する場合には、当該無音区間が抽出されたタイミングをブレス位置のタイミングとして抽出することも可能である。また、所定期間内においてピッチが規則的に変動している場合には、当該機関について「ビブラート」が行われていると判定することで、歌唱技法が行われたタイミング(技法位置)を抽出することも可能である。
 楽音トラックは、楽音を発生させる楽器の種類、タイミング、音程(キー)、強さ、長さ、定位(パン)、音響効果(エフェクト)等を示す情報が記録されている。ガイドメロディトラックは、お手本の歌唱に対応する各音の発音開始タイミング、発音の長さ等の情報が記録されている。
 シーケンサは、楽音トラックのデータに基づいて音源18を制御し、カラオケ曲の楽音を発生する。
 また、シーケンサは、コーラストラックの指定するタイミングでバックコーラスの音声データ(楽曲データに付随しているMP3等の圧縮音声データ)を再生する。また、シーケンサは、歌詞トラックに基づいて曲の進行に同期して歌詞の文字パターンを合成し、この文字パターンを映像信号に変換して表示処理部23に入力する。
 音源18は、シーケンサの処理によってCPU11から入力されたデータ(ノートイベントデータ)に応じて楽音信号(デジタル音声信号)を形成する。形成した楽音信号はミキサ19に入力される。
 ミキサ19は、音源18が発生した楽音信号、コーラス音、およびマイク(歌唱音声入力手段)16からA/Dコンバータ17を介して入力された歌唱者の歌唱音声信号に対してエコー等の音響効果を付与するとともに、これらの信号をミキシングする。
 また、異なる場所に設置されているカラオケ装置7同士で通信を行い、デュエットを行う場合には、他のカラオケ装置から歌唱音声信号が送信される。ミキサ19には、当該他のカラオケ装置から受信した歌唱音声信号も入力され、自装置のマイク16から入力された歌唱音声信号とミキシングされる。
 ミキシングされた各デジタル音声信号はサウンドシステム20に入力される。サウンドシステム20は、D/Aコンバータおよびパワーアンプを内蔵しており、入力されたデジタル信号をアナログ信号に変換して増幅し、スピーカ(楽音発生手段)21から放音する。ミキサ19が各音声信号に付与する効果およびミキシングのバランスは、CPU11によって制御される。
 CPU11は、上記シーケンサによる楽音の発生、歌詞テロップの生成と同期して、HDD13に記憶されている映像データを読み出して背景映像等を再生する。動画の映像データは、MPEG形式にエンコードされている。
 また、CPU11は、歌唱者を表す写真、またはキャラクタ等の映像データをセンタ1からダウンロードして表示処理部23に入力することもできる。歌唱者を表す写真は、その場でカラオケ装置またはリモコン9に設けられたカメラ(不図示)で撮影したり、ユーザが所有する携帯端末等に設けられたカメラで撮影したりすることも可能である。
 CPU11は、読み出した背景映像の映像データをデコーダ22に入力する。デコーダ22は、入力されたMPEG等のデータを映像信号に変換して表示処理部23に入力する。表示処理部23には、背景映像の映像信号以外に上記歌詞テロップの文字パターンとともに、ガイドメロディトラックに基づくピアノロールの映像信号も入力される。
 図4A-4Cは、ピアノロールの一例を示す図である。ピアノロールは、図4Aに示すように、縦軸が音階(ピアノの鍵盤が縦になった状態)、横軸が時間に対応した画面上に、各音の発音開始タイミングと発音の長さに応じた線状の画像をガイド画像として表示するものである。これにより、歌唱者は、各音を歌唱するタイミングと音程を視覚的に把握することができる。ここで、本実施形態のピアノロールでは、各音のガイド画像が連結されて表示されるとともに、息継ぎタイミングでガイド画像が離散して表示される。
 CPU11は、まずガイドメロディトラックに含まれている各音の発音開始タイミングおよび発音の長さの情報に基づいて、ガイド画像を生成する。そして、CPU11は、各音のガイド画像を滑らかに連結する。各音の連結部分の傾きは、例えば16分音符に対応する時間長で連結させる等、一律に同じ傾きの画像として表示させる態様とする。ただし、実際には各曲に個別の歌い方が存在し、音程の変化の態様は一律ではない。したがって、各音の連結部分毎に異なる傾きで表示されることが好ましい。この場合、リファレンスデータとして、各音の連結部分の音程の変化に応じた傾きを指定する情報が含まれていてもよい。
 その後、CPU11は、ブレス位置トラックが示す息継ぎタイミングでガイド画像を離散させる。例えば、図4Aの例では、「あかい」の発音の後と、「はなが」の冒頭の発音タイミングとの間に息継ぎタイミングが存在するため、「あかい」のガイド画像と、「はなが」のガイド画像とを離散させる。
 これにより、歌唱者は、各音のつなげ方および息継ぎのタイミングを視覚的に把握することができる。例えば図4Aの例では、歌唱者は、「あかい」の各音の音程を1音ずつ滑らかに変化させながら歌唱を行い、息継ぎを行った後に「はなが」の各音の音程を1音ずつ滑らかに変化させながら歌唱を行う旨を、視覚的に容易に把握することができる。また、CPU11は、ブレス位置トラックが示す息継ぎタイミングでガイド画像を離散させるため(リファレンスデータに息継ぎタイミングを示す情報が含まれているため)、単なる無音区間と息継ぎの区間と明確に区別して表示することができ、ユーザに対して正確な息継ぎの位置を把握させることができる。
 なお、図3では、1つの楽曲データにブレス位置トラックおよび技法位置トラックが含まれている例を示したが、既存の楽曲データはそのままで、ブレス位置トラックおよび技法位置トラックは、別のデータとして用意してもよい。この場合、ブレス位置トラックおよび技法位置トラックが含まれた新たな楽曲データを用意する必要はない。ただし、ブレス位置トラックおよび技法位置トラックのデータには、それぞれ曲番号等の曲識別情報を記載しておく。CPU11は、楽曲データを読み出すときに、対応するブレス位置トラックおよび技法位置トラックを読み出し、シーケンス動作を行う。
 なお、同じ音程で続けて異なる歌詞を発音する場合、ガイド画像が連結されていると、ユーザは、どのタイミングで次の歌詞を発音するのか把握し難い可能性がある。そこで、CPU11は、図4Bに示すように、ガイドメロディトラックにおける各音の発音タイミングにおいて、例えばガイド画像の上に円画像を重畳して表示する。これにより、ユーザは、当該円画像が表示されているタイミングで発音を行う旨を把握することができる。
 次に、図4Cは、促音の音素に係るガイド画像と、該促音の音素の後のガイド画像と、を離散させて表示する態様である。促音は、日本語のかな表記では「っ」「ッ」で表されるものであり、後に続く音との間が無音となるものである。CPU11は、歌詞トラックの中から促音を抽出し、抽出した促音を発するタイミングにおいて、ガイド画像を離散させる。図4Cの例では、「つかったら」の「か」の後に促音が存在するため、「か」のガイド画像と、その後の「た」のガイド画像とを離散させて表示する。また、図4Cの例では、促音に係る音素が存在する旨を示す画像(この例では「っ」と記載された四角画像)を表示する態様としている。
 これにより、歌唱者は、音をつなげて発音するのか、促音が存在して音が途切れるのか、等を容易に把握することができる。
 次に、図5Aは、ガイド画像で音量を表す場合の例である。この場合、ガイドメロディトラックには、各音の音量を示す情報が含まれている。CPU11は、ガイドメロディトラックに含まれている各音の音量を示す情報に基づいて、ガイド画像の線の太さを変更する。例えば、図5Aの例では、「あかい」のうち、「あ」の音が最も音量が大きいため、「あ」を表す区間は太い線に変更する。「か」の音は音量が小さいため、「か」を表す区間は細い線に変更する。この例では、ガイドメロディトラックに含まれている音量を示す情報が、「大」、「標準」、および「小」の3段階であり、線の太さを3段階に変更する態様としているが、さらに多段階に線の太さを変更する態様としてもよい。
 なお、図5Aの例では、連結部分の中間位置において線の太さを変更する態様としているが、各音の冒頭または各音の末尾の位置において線の太さを変更する態様としてもよい。また、音量を各音の末尾から次の音の冒頭まで徐々に線の太さが変化する態様としてもよい。
 次に、図5Bは、さらに線の色で音量を表す態様である。CPU11は、ガイドメロディトラックに含まれている各音の音量を示す情報に基づいて、ガイド画像の線の色を変更する。例えば、図5Bの例では、「か」の音は音量が小さいため、「か」を表す区間は薄い色の線に変更する。この例では、ガイドメロディトラックに含まれている音量を示す情報が「小」の場合に、当該「小」の音の区間を薄い色に変更する態様としているが、音量「大」の音の区間を濃い色に変更してもよいし、線の太さを変更せずに色だけを変更する態様としてもよい。
 次に、図6は、歌唱技法をピアノロール上に表示する例を示したものである。図6Aの例では、ビブラートの区間のガイド画像を波線に変更して表示するものである。この場合、CPU11は、技法位置トラックに含まれているビブラートのタイミングを示す情報を読み出し、当該タイミングに対応する音のガイド画像のうち、当該タイミングから末尾までの区間を波線に変更する。これにより、歌唱者は、より直感的にビブラートを行うタイミングおよびビブラートの長さを把握し易くなる。
 また、図6Bに示すように、技法位置トラックに「タメ」の情報が含まれている場合には、当該「タメ」の情報に対応する音(この例では「か」の音)の先頭の位置を遅らせる。これにより、歌唱者は、当該特定の音(この例では「か」の音)の歌いだしを故意に遅らせる「タメ」を直感的に把握し易くなる。
 また、技法位置トラックに「しゃくり」の情報が含まれている場合には、リファレンスの音程よりも低い音程から元の音程に上昇させるガイド画像とする。図6Bの例では、冒頭の「あ」の音および「は」の音が「しゃくり」であり、リファレンスの音程よりも低い音程から持ち上げつつ歌唱を行う箇所である。したがって「あ」の区間の冒頭は、リファレンスの音程よりも低い音程から元の音程に上昇させるガイド画像とする。これにより、歌唱者は、ガイド画像から「しゃくり」の歌唱技法を直感的に把握できるようになっている。
 また、技法位置トラックに「コブシ」の情報が含まれている場合には、図6Bの「な」の音に示すように、「コブシ」のタイミングに対応する位置でガイド画像を一時的に上昇させる。これにより、特定の音の声色を発音の途中でうなるように変化させる歌唱技法である「コブシ」に対応するガイド画像とすることも可能である。また、技法位置トラックに「フォール」の情報が含まれている場合には、図6Bの「が」の音に示すように、「フォール」のタイミングに対応する位置からガイド画像を低い音程に変化させる態様とすることで、「フォール」の歌唱技法に対応するガイド画像とすることも可能である。
 図6Cは、歌唱技法を促す画像を表示する例である。CPU11は、技法位置トラックに含まれている各種歌唱技法のタイミングを示す情報を読み出し、当該タイミングに対応する位置に歌唱技法を促す画像を表示する。例えば、冒頭の「あ」についてはしゃくりであり、リファレンスの音程よりも低い音程から持ち上げつつ歌唱を行う箇所である。したがって「あ」の区間の冒頭に「ノ」のような音程が上がることを連想させる画像を表示する。また、ビブラートの区間については波線状のガイド画像に加えて、別途「m」のような波線状の画像を表示する。これにより、歌唱者は、どのタイミングでどのような歌唱技法を行うか、容易に把握することができる。
 さらに、CPU11は、ブレス位置トラックが示す息継ぎタイミングを示す情報に基づいて、息継ぎを促す画像を、ガイド画像とともに表示する。例えば、図6Cの例では、「あかい」の区間と「はなが」の区間の間に「V」のような画像を表示する。これにより、歌唱者は、音をつなげて発音するのか、息継ぎを行うのか、より容易に把握することができる。
 以上の様にして、カラオケ演奏が行われ、演奏の進行にしたがってピアノロールが表示される。このように各音が滑らかに連結されたガイド画像を表示することで、従来のピアノロールに比べて、歌唱者は、ガイド画像を見ながら各音を滑らかにつなげて歌唱したり息継ぎを行ったりすることが容易になる。
 次に、採点処理について説明する。採点処理は、歌唱者の歌唱音声をガイドメロディトラックと比較することによって行われる。採点は、ガイドメロディトラックのノート毎に、歌唱音声とガイドメロディの音程(ピッチ)を比較することによって行われる。すなわち、歌唱音声の音程が、所定時間以上、ガイドメロディトラックの音程に合っていた(許容範囲に入っていた)場合には、高い得点を付与する。また、音程変化のタイミングも得点に考慮される。さらに、音程変化のタイミング、ビブラート、抑揚、しゃくり(低い音程からなだらかに移行すること)等の歌唱技法の有無に基づいて加点も行われる。
 さらに、本実施形態の採点処理では、ブレス位置トラックに含まれる息継ぎタイミングにおいて歌唱者が息継ぎを行ったか否かも加点対象とする。息継ぎを行ったか否かは、当該息継ぎタイミングを含む所定時間内においてマイク16から音声が入力されていない(入力レベルが所定閾値未満である)またはマイク16から息継ぎ音が入力された場合に、息継ぎを行ったと判定し、マイク16から音声が入力された(入力レベルが所定閾値以上である)場合に、息継ぎが行われていないと判定する。なお、息継ぎ音が収音されたか否かは、例えばパターンマッチング等で息継ぎ音の波形と対比することで判断する。
 また、本実施形態の採点処理では、技法位置トラックに含まれる各技法のタイミングにおいて、同じ技法を検出した場合に、より高い得点を付与することが好ましい。
 なお、採点処理は、各カラオケ装置において行ってもよいが、センタ1(または他のサーバ)で行ってもよい。また、ネットワークを介して他のカラオケ装置とデュエットを行っている場合には、代表的に処理を行うカラオケ装置1台で採点処理を行ってもよい。
 次に、カラオケシステムの動作についてフローチャートを参照して説明する。図7は、カラオケシステムの動作を示すフローチャートである。
 まず、歌唱者は、楽曲のリクエストを行う(s11)。このとき、デュエット曲が選択された場合に、CPU11は、モニタ24において、ネットワークを介して接続された他のカラオケ装置の歌唱者とデュエットを行うか否かを促す画像を表示し、ネットワーク経由のデュエット歌唱を受け付ける。例えば、歌唱者がタッチパネル15、操作部25、またはリモコン9を用いて特定のユーザの氏名を入力すると、センタ1で当該氏名に係るユーザが検索され、デュエット相手として設定される。
 次に、カラオケ装置のCPU11は、リクエストされた楽曲データを読み出し(s12)、ピアノロールを作成する(s13)。すなわち、CPU11は、ガイドメロディトラックに含まれている各音の発音開始タイミングおよび発音の長さの情報に基づいて、ガイド画像を生成する。
 その後、CPU11は、歌詞トラックを読み出して(s14)、各ガイド画像に歌詞の画像を対応付ける(s15)。また、CPU11は、ブレス位置トラックから息継ぎタイミングの情報を読み出すとともに、促音の音素に係る発音タイミングを読み出す(s16)。そして、CPU11は、各音のガイド画像を滑らかに連結する(s17)。このとき、CPU11は、ブレス位置トラックが示す息継ぎタイミング、および促音の音素に係る発音タイミングにおいて、ガイド画像を離散させて表示する。
 また、CPU11は、技法位置トラックを読み出して(s18)、歌唱技法をピアノロール上に表示する(s19)。また、CPU11は、歌唱技法に応じてガイド画像を変更する。例えば、ビブラートの区間は、ガイド画像を波線に変更して表示する。
 また、CPU11は、ガイドメロディトラックに含まれている各音の音量を示す情報を読み出し(s20)、各音の音量に応じたガイド画像に変更する(s21)。例えば、図5Aに示したように、音量に応じて線の太さを変更したり、図5Bに示したように、音量に応じて線の色を変更したりする。
 なお、本実施形態では、カラオケ装置7を用いてカラオケ演奏およびピアノロールの表示を行う態様を示したが、例えばユーザの所有するPCやスマートフォン、ゲーム機等の情報処理装置(マイク、スピーカ、および表示部の構成を備えたもの)を用いることでも、本発明の表示装置を実現することが可能である。なお、楽曲データおよびリファレンスデータは、表示装置に記憶されている必要はなく、サーバから都度ダウンロードして利用するようにしてもよい。
 なお、CPU11は、図8に示すように、現在の歌唱位置にキャラクタを表示する態様としてもよい。この例では、ガイド画像が地面の画像に対応し、息継ぎタイミングにおいて地面の画像が途切れるようになっている。そして、キャラクタ画像101がガイド画像(地面の画像)に沿って移動するように、ガイド画像および背景をスクロールさせる。息継ぎタイミングにおいては、地面が途切れるため、息継ぎ(マイク16から音声が入力されていない状態)を検出しなかった場合に、キャラクタ画像101が地面から落ちるようになっている。また、この例では、歌唱採点の結果が画面上に表示される。したがって、歌唱者は、ゲームのように楽しんでカラオケを行うことができる。
 また、ガイド画像は、図8のような客観視点(2次元表示、2次元視点)で表示される態様であってもよいが、例えば図9に示すような主観視点(3次元表示、3次元視点)で表示される態様であってもよい。主観視点とは、ユーザ自身の視野を模した表示態様であり、3次元視点の一種である。ここでは、奥行き方向を時間軸に対応させ、平面方向を音程に対応させた表示態様を示す。例えば、図9に示すように、奥行き方向が時間に対応し、上下方向が音階に対応した表示態様である。なお、図9の例では、ユーザ自身に相当する画像(キャラクタ画像等)を表示し、当該キャラクタ画像等を背後から映すように表示する態様であり、当該表示態様も3次元視点に相当する。なお、音階は左右方向に対応していてもよい。この場合、キャラクタ画像101Aがガイド画像に沿って奥行き方向に移動するように、ガイド画像および背景をスクロールさせる。この例でも、歌唱採点の結果が画面上に表示される。したがって、歌唱者は、ゲームのように楽しんでカラオケを行うことができる。
 また、図9に示すように、主観視点で表示する場合には、デュエット歌唱を行う場合に、自身に相当するキャラクタ画像101Aと他の歌唱者に相当するキャラクタ101B(およびキャラクタ画像101C)とを並行して表示することも可能である。これにより、ユーザは、他の歌唱者と一緒に歌唱を行っている雰囲気をより感じ取ることができる。
 また、本実施形態では、カラオケにおけるガイドメロディをガイド画像として表示する例を示したが、例えば吹奏楽器演奏のお手本の音程変化をガイド画像として表示し、息継ぎタイミングでガイド画像が途切れる態様としても同様の効果が得られる。また、例えば語学学習において、お手本の発音タイミングおよび発音長を示すガイド画像を表示し、息継ぎタイミングおよび促音でガイド画像が途切れる態様としても、同様の効果が得られる。本実施形態では、歌詞を表示する例を説明しているが、歌詞の表示は本発明において必須ではない。
 なお、図10に示すように、本発明のリファレンス表示装置は、表示部であるモニタ24と、ガイド画像表示処理を行う画像生成手段として機能するCPU11と、を備え、当該CPU11が、HDD13に記憶されている楽曲データ(本発明のリファレンスデータの一例である。)に基づいてガイド画像を生成して、ガイド画像の各音を連結させる態様とすればよい。他のハードウェア構成は、本発明において必須の要素ではない。
 また、上述したように、リファレンスデータは、HDD13に記憶されている必要はなく、外部(例えばサーバ)から都度ダウンロードして利用するようにしてもよい。また、デコーダ22、表示処理部23、およびRAM12も、CPU11の機能の一部として当該CPU11が内蔵していてもよい。
 なお、ガイド画像は、ピアノロール(縦軸がピアノの鍵盤に対応し、横軸方向に沿って実線が表示されたもの)として表示することは必須ではない。例えば、図8および図9に示したように、発音タイミング、音程、および発音長を示すガイド画像を生成して各音を連結させる態様であれば、どの様な表示態様であってもよい。なお、本発明で言うガイド画像とは、図4Aから図6で示した細長い線に限るものではなく、図9の例に示したように、左右または上下方向にある程度の幅を有した画像が一方向(図9の例では奥行き方向)に延びるものも含む。
 本出願は、2014年7月28日に出願された日本国特願2014-152479に基づき、その優先権を主張するものであり、ここに参照として取り込まれる。
1…センタ
2…ネットワーク
3…カラオケ店舗
5…中継機
7…カラオケ装置
9…リモコン
11…CPU
12…RAM
13…HDD
15…タッチパネル
16…マイク
17…A/Dコンバータ
18…音源
19…ミキサ
20…サウンドシステム
22…デコーダ
23…表示処理部
24…モニタ
25…操作部
26…送受信部

Claims (12)

  1.  表示部と、
     リファレンスデータに基づいて、発音タイミングと音程と発音長を示すガイド画像を生成し、前記表示部に表示する画像生成手段と、
     を備えたリファレンス表示装置であって、
     前記リファレンスデータには、息継ぎタイミングを示す情報が含まれ、
     前記画像生成手段は、前記リファレンスデータ中の各音を連結させ、かつ前記息継ぎタイミングを示す情報に基づいて、該息継ぎタイミングの前後の音を離散させて表現されたガイド画像を表示することを特徴とするリファレンス表示装置。
  2.  前記画像生成手段は、促音に係る音素と、該促音に係る音素の次の音素とが離散させて表現されたガイド画像を表示する請求項1に記載のリファレンス表示装置。
  3.  前記画像生成手段は、前記息継ぎタイミングを示す情報に基づいて、息継ぎを促す画像を、前記ガイド画像とともに表示させる請求項1または2に記載のリファレンス表示装置。
  4.  前記画像生成手段は、各音の発音タイミングを示す画像を、前記ガイド画像に重畳して表示する請求項1から3のいずれかに記載のリファレンス表示装置。
  5.  前記リファレンスデータには、各音の音量を示す情報が含まれ、
     前記画像生成手段は、前記各音の音量を示す情報に基づいて、前記ガイド画像を前記音量に応じた画像に変更して表示する請求項1から4のいずれかに記載のリファレンス表示装置。
  6.  前記画像生成手段は、前記各音の音量を示す情報に基づいて、前記ガイド画像を前記音量に応じてその大きさ、色、濃淡の少なくともいずれかを変更して表示する請求項5に記載のリファレンス表示装置。
  7.  前記ガイド画像は、前記表示部上の平面方向の一方向を時間軸に対応させ、平面方向の他方向を音程に対応させた2次元視点で表示される請求項1から6のいずれかに記載のリファレンス表示装置。
  8.  前記ガイド画像は、細長い線状画像で表現される請求項1から7のいずれかに記載のリファレンス表示装置。
  9.  前記ガイド画像は、前記表示部上の奥行き方向を時間軸に対応させ、平面方向を音程に対応させた3次元視点で表示される請求項1から6のいずれかに記載のリファレンス表示装置。
  10.  前記ガイド画像は、奥行き方向に延びる細長い線状画像、または平面方向に幅を有し奥行き方向に延びる板状画像で表現される請求項9に記載のリファレンス表示装置。
  11.  表示部を備えた情報処理装置におけるリファレンス表示方法であって、
     息継ぎタイミングを示す情報を含むリファレンスデータに基づいて、発音タイミングと音程と発音長を示すガイド画像であって、前記リファレンスデータ中の各音を連結させ、かつ前記息継ぎタイミングを示す情報に基づいて、該息継ぎタイミングの前後の音を離散させて表現されたガイド画像を生成し、前記表示部に表示するリファレンス表示方法。
  12.  表示部を備えた情報処理装置に、
     リファレンスデータに基づいて、発音タイミングと音程と発音長を示すガイド画像を生成し、前記表示部に表示する画像生成ステップを実行させるプログラムであって、
     前記リファレンスデータには、息継ぎタイミングを示す情報が含まれ、
     前記画像生成ステップは、前記リファレンスデータ中の各音を連結させ、かつ前記息継ぎタイミングを示す情報に基づいて、該息継ぎタイミングの前後の音を離散させて表現されたガイド画像を表示することを特徴とするプログラム。
PCT/JP2015/071342 2014-07-28 2015-07-28 リファレンス表示装置、リファレンス表示方法およびプログラム Ceased WO2016017622A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US15/500,011 US10332496B2 (en) 2014-07-28 2015-07-28 Reference display device, reference display method, and program

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2014-152479 2014-07-28
JP2014152479A JP6070652B2 (ja) 2014-07-28 2014-07-28 リファレンス表示装置およびプログラム

Publications (1)

Publication Number Publication Date
WO2016017622A1 true WO2016017622A1 (ja) 2016-02-04

Family

ID=55217522

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2015/071342 Ceased WO2016017622A1 (ja) 2014-07-28 2015-07-28 リファレンス表示装置、リファレンス表示方法およびプログラム

Country Status (4)

Country Link
US (1) US10332496B2 (ja)
JP (1) JP6070652B2 (ja)
TW (1) TWI595476B (ja)
WO (1) WO2016017622A1 (ja)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2017156553A (ja) * 2016-03-02 2017-09-07 ブラザー工業株式会社 カラオケ装置、および、カラオケ制御プログラム

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP6070652B2 (ja) * 2014-07-28 2017-02-01 ヤマハ株式会社 リファレンス表示装置およびプログラム
WO2019240042A1 (ja) * 2018-06-15 2019-12-19 ヤマハ株式会社 表示制御方法、表示制御装置およびプログラム
CN110010162A (zh) * 2019-02-28 2019-07-12 华为技术有限公司 一种歌曲录制方法、修音方法及电子设备
JP7755512B2 (ja) * 2022-02-22 2025-10-16 株式会社第一興商 カラオケ装置

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08190391A (ja) * 1994-11-08 1996-07-23 Mikihiko Yao カラオケ御指南役色変わりテロップ
JP2001318683A (ja) * 2000-05-12 2001-11-16 Victor Co Of Japan Ltd 歌唱情報表示装置及び方法
JP2004212547A (ja) * 2002-12-27 2004-07-29 Yamaha Corp カラオケ装置
JP2007264060A (ja) * 2006-03-27 2007-10-11 Casio Comput Co Ltd カラオケ装置およびカラオケ情報処理のプログラム
JP2007322933A (ja) * 2006-06-02 2007-12-13 Yamaha Corp 指導装置、指導用データ製作装置及びプログラム

Family Cites Families (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH03296789A (ja) * 1990-04-17 1991-12-27 Brother Ind Ltd 楽音発生装置
US5563358A (en) * 1991-12-06 1996-10-08 Zimmerman; Thomas G. Music training apparatus
JP3201202B2 (ja) * 1995-01-12 2001-08-20 ヤマハ株式会社 楽音信号合成装置
US6546229B1 (en) * 2000-11-22 2003-04-08 Roger Love Method of singing instruction
US8672852B2 (en) * 2002-12-13 2014-03-18 Intercure Ltd. Apparatus and method for beneficial modification of biorhythmic activity
JP4182750B2 (ja) 2002-12-25 2008-11-19 ヤマハ株式会社 カラオケ装置
US7271329B2 (en) * 2004-05-28 2007-09-18 Electronic Learning Products, Inc. Computer-aided learning system employing a pitch tracking line
JP2007225916A (ja) * 2006-02-23 2007-09-06 Yamaha Corp オーサリング装置、オーサリング方法およびプログラム
JP5176311B2 (ja) * 2006-12-07 2013-04-03 ソニー株式会社 画像表示システム、表示装置、表示方法
JP6070652B2 (ja) * 2014-07-28 2017-02-01 ヤマハ株式会社 リファレンス表示装置およびプログラム

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08190391A (ja) * 1994-11-08 1996-07-23 Mikihiko Yao カラオケ御指南役色変わりテロップ
JP2001318683A (ja) * 2000-05-12 2001-11-16 Victor Co Of Japan Ltd 歌唱情報表示装置及び方法
JP2004212547A (ja) * 2002-12-27 2004-07-29 Yamaha Corp カラオケ装置
JP2007264060A (ja) * 2006-03-27 2007-10-11 Casio Comput Co Ltd カラオケ装置およびカラオケ情報処理のプログラム
JP2007322933A (ja) * 2006-06-02 2007-12-13 Yamaha Corp 指導装置、指導用データ製作装置及びプログラム

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2017156553A (ja) * 2016-03-02 2017-09-07 ブラザー工業株式会社 カラオケ装置、および、カラオケ制御プログラム

Also Published As

Publication number Publication date
TW201610981A (zh) 2016-03-16
US20170352340A1 (en) 2017-12-07
TWI595476B (zh) 2017-08-11
JP6070652B2 (ja) 2017-02-01
JP2016031394A (ja) 2016-03-07
US10332496B2 (en) 2019-06-25

Similar Documents

Publication Publication Date Title
TWI492216B (zh) 顯示控制裝置、方法及電腦可讀取之記憶媒體
JP6776788B2 (ja) 演奏制御方法、演奏制御装置およびプログラム
WO2019058942A1 (ja) 再生制御方法、再生制御装置およびプログラム
TW201108202A (en) System, method, and apparatus for singing voice synthesis
JP6070652B2 (ja) リファレンス表示装置およびプログラム
JP5151245B2 (ja) データ再生装置、データ再生方法およびプログラム
JP5459331B2 (ja) 投稿再生装置及びプログラム
WO2016017623A1 (ja) リファレンス表示装置、リファレンス表示方法およびプログラム
JP2007310204A (ja) 楽曲練習支援装置、制御方法及びプログラム
JP5387642B2 (ja) 歌詞テロップ表示装置及びプログラム
JP5486941B2 (ja) 聴衆に唱和をうながす気分を楽しむカラオケ装置
JP6369225B2 (ja) カラオケ装置及びカラオケプログラム
JP2006251697A (ja) カラオケ装置
JP5537246B2 (ja) 歌唱位置表示システム
JP6838357B2 (ja) 音響解析方法および音響解析装置
JP6236807B2 (ja) 歌唱音声評価装置および歌唱音声評価システム
JP2014178457A (ja) 歌唱システムおよび歌唱装置
JP2007233078A (ja) 評価装置、制御方法及びプログラム
JP6144593B2 (ja) 歌唱採点システム
JP4033146B2 (ja) カラオケ装置
JP2007322933A (ja) 指導装置、指導用データ製作装置及びプログラム
JP6589400B2 (ja) ネットワークカラオケシステム、カラオケ装置、サーバ、ネットワークカラオケシステムの制御方法、カラオケ装置の制御方法、及びサーバの制御方法
JP2026042560A (ja) カラオケシステム
JP6323236B2 (ja) カラオケ装置及びカラオケ用プログラム
JP6485955B2 (ja) 歌唱音声の放音遅延に対応したカラオケシステム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15827807

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

WWE Wipo information: entry into national phase

Ref document number: 15500011

Country of ref document: US

122 Ep: pct application non-entry in european phase

Ref document number: 15827807

Country of ref document: EP

Kind code of ref document: A1