WO2014142201A1 - 分離用データ処理装置およびプログラム - Google Patents

分離用データ処理装置およびプログラム Download PDF

Info

Publication number
WO2014142201A1
WO2014142201A1 PCT/JP2014/056572 JP2014056572W WO2014142201A1 WO 2014142201 A1 WO2014142201 A1 WO 2014142201A1 JP 2014056572 W JP2014056572 W JP 2014056572W WO 2014142201 A1 WO2014142201 A1 WO 2014142201A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
separation
update
acoustic signal
recorded
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2014/056572
Other languages
English (en)
French (fr)
Inventor
木村 繁樹
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Yamaha Corp
Original Assignee
Yamaha Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Yamaha Corp filed Critical Yamaha Corp
Priority to CN201480014346.5A priority Critical patent/CN105122360A/zh
Priority to KR1020157024425A priority patent/KR20150119013A/ko
Publication of WO2014142201A1 publication Critical patent/WO2014142201A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0272Voice signal separating

Definitions

  • the present invention relates to a technique for separating (emphasizing or suppressing) a specific component (hereinafter referred to as “specific component”) of an acoustic signal.
  • Patent Document 1 discloses a technique for calculating a position where a sound image of an acoustic signal is localized for each frequency component and suppressing the frequency component where the sound image is localized in a specific range from the acoustic signal.
  • an object of the present invention is to easily generate separation data capable of separating (suppressing or enhancing) a specific component of an acoustic signal with high accuracy.
  • a separation data processing device of the present invention includes a storage unit (for example, a storage device 144) that stores separation data applied to separation processing that emphasizes or suppresses a specific component of an acoustic signal.
  • An update data acquisition unit for acquiring update data reflecting an input by a user who has received a reproduction sound of an acoustic signal (for example, an acoustic signal SB) after separation processing to which the separation data is applied; and update data acquisition
  • An update processing unit that updates the separation data in the storage unit using the update data acquired by the unit.
  • the separation data is updated using the update data reflecting the input by the listener of the reproduced sound of the acoustic signal (for example, the acoustic signal SB) after the separation processing to which the separation data is applied. Therefore, there is an advantage that highly reliable separation data capable of separating a specific component of an acoustic signal with high accuracy can be easily generated.
  • the update data acquisition unit records the recorded sound that is recorded in parallel with the reproduction of the acoustic signal in which the specific component is suppressed in the separation process using the separation data (that is, in parallel with the reproduced sound of the acoustic signal after the separation process).
  • the update data including the recorded data according to the singing sound sung by the user may be acquired.
  • the separation data is updated using the update data including the recorded data corresponding to the recorded sound recorded in parallel with the reproduction of the acoustic signal in which the specific component is suppressed.
  • the update data acquisition unit emphasizes the specific component by the recorded signal recorded in parallel with the reproduction of the acoustic signal in which the specific component is suppressed by the separation process using the separation data and the separation process by applying the separation data.
  • Update data including evaluation data corresponding to the comparison result with the acoustic signal and recorded data corresponding to the recorded signal may be acquired.
  • the evaluation data corresponding to the comparison result between the recorded signal recorded in parallel with the reproduction of the acoustic signal in which the specific component is suppressed and the acoustic signal in which the specific component is emphasized is updated as the update data.
  • the update processing unit performs separation according to the update data so that the influence (statistical weighting) of the recorded data on the update of the separation data increases as the recorded signal and the acoustic signal that emphasizes a specific component are more similar. Data may be updated.
  • highly reliable separation data can be efficiently generated by reducing the influence of the recorded signal whose acoustic characteristics deviate from the acoustic signal in which the specific component is emphasized.
  • the update data acquisition unit may include adjustment data in accordance with an adjustment instruction from the listener of the reproduced sound of the acoustic signal after separation processing to which the separation data is applied.
  • the adjustment data according to the adjustment instruction by the listener of the reproduced sound of the acoustic signal after the separation process is applied as update data to the update of the separation data. Therefore, there is an advantage that it is possible to generate highly reliable separation data reflecting the result of the user listening to the reproduced sound of the acoustic signal after the separation processing.
  • the update data acquisition unit may acquire update data from one or more terminal devices via a communication network.
  • the update data acquired from one or more terminal devices via the communication network is used for the update of the separation data, the specific component of the acoustic signal can be separated with high accuracy.
  • the above-described effect that the separation data can be easily generated is particularly remarkable.
  • separation data applied to separation processing for emphasizing or suppressing a specific component of an acoustic signal is stored, and in parallel with reproduction of the acoustic signal in which the specific component is suppressed by the separation processing using the separation data.
  • a separation data processing method for acquiring update data including recorded data corresponding to the recorded sound recorded and updating the separation data using the acquired update data.
  • the data processing apparatus for separation according to each of the above aspects is realized by hardware (electronic circuit) such as DSP (Digital Signal Processor) dedicated to processing of acoustic signals, as well as general-purpose such as CPU (Central Processing Unit) This is also realized by cooperation between the arithmetic processing unit and the program.
  • the program according to the present invention uses a computer having a storage unit that stores separation data to be applied to separation processing for emphasizing or suppressing a specific component of an acoustic signal.
  • the computer is caused to execute an update processing unit that updates the business data.
  • the program according to each aspect described above can be provided in a form stored in a computer-readable recording medium and installed in the computer.
  • the recording medium is, for example, a non-transitory recording medium, and an optical recording medium (optical disk) such as a CD-ROM is a good example, but a known arbitrary one such as a semiconductor recording medium or a magnetic recording medium This type of recording medium can be included.
  • the program of the present invention can be provided in the form of distribution via a communication network and installed in a computer.
  • FIG. 1 is a block diagram of a sound processing system according to a first embodiment of the present invention. It is explanatory drawing of a separation process. It is explanatory drawing of schematic operation
  • FIG. 1 is a block diagram of a sound processing system 100 according to the first embodiment of the present invention.
  • the sound processing system 100 is a communication system including a plurality of terminal devices 12 and a separation data processing device 14A.
  • Each terminal device 12 is a communication terminal such as a mobile phone or a smartphone, and communicates with the separation data processing device 14A via the communication network 16 (for example, a mobile communication network or the Internet).
  • the communication network 16 for example, a mobile communication network or the Internet.
  • each terminal device 12 includes a control device 121, a storage device 122, a communication device 123, a display device 124, an input device 125, a sound collection device 126, and a sound emission. And a computer system including the device 127.
  • the control device 121 is an arithmetic processing device that executes various types of control processing and arithmetic processing by executing a program stored in the storage device 122.
  • the communication device 123 communicates with the separation data processing device 14A via the communication network 16.
  • the communication between the terminal device 12 and the communication network 16 is typically wireless communication. However, for example, when a stationary information processing device is used as the terminal device 12, the terminal device 12 and the communication network 16 Can also be wired.
  • Display device 124 (for example, a liquid crystal display panel) displays an image instructed from control device 121.
  • the input device 125 is a device that receives an instruction from the user to the terminal device 12 and includes, for example, a plurality of operators operated by the user.
  • a touch panel configured integrally with the display device 124 may be employed as the input device 125.
  • the storage device 122 (for example, a semiconductor recording medium) stores a program executed by the control device 121 and various data used by the control device 121.
  • the storage device 122 of the first embodiment stores a plurality of acoustic signals SA corresponding to different music pieces.
  • the acoustic signal SA is a stereo signal of two left and right channels indicating a time waveform of a mixed sound of a plurality of sounds. Specifically, a mixed sound of a singing sound singing the melody of the music and a performance sound of a plurality of musical instruments constituting the accompaniment of the music is expressed by an acoustic signal SA.
  • Each acoustic signal SA stored in the storage device 122 is associated with attribute data MA.
  • the attribute data MA is information related to music. Specifically, attribute information such as a song name and a singer name is specified by the attribute data MA.
  • each terminal device 12 of the first embodiment a separation process for separating (suppressing or enhancing) a specific component of the acoustic signal SA is executed. Specifically, a stereo-type acoustic signal SB is generated that indicates the sound in which the singing sound of the acoustic signal SA is suppressed (ideally removed) and the accompaniment sound is emphasized (ideally extracted).
  • FIG. 2 is an explanatory diagram of separation processing in the first embodiment.
  • FIG. 2 shows the distribution (scatter diagram) of each frequency component of the acoustic signal in the localization-frequency plane.
  • the localization-frequency plane is a coordinate in which a localization axis XL indicating a position (hereinafter referred to as “sound image position”) ⁇ where a sound image of each frequency component of an acoustic signal is localized and a frequency axis XF indicating the frequency of each frequency component are set. It is a plane.
  • the separation processing of the first embodiment suppresses or emphasizes components in a specific range (hereinafter referred to as “separation target range”) R defined in the localization-frequency plane of FIG. 2 in the sound indicated by the acoustic signal SA. It is processing. As shown in FIG.
  • the separation target range R includes a specific range on the frequency axis XF (hereinafter referred to as “target frequency range”) RF and a specific range on the localization axis XL (hereinafter referred to as “target localization range”) RL. It is prescribed by.
  • FIG. 2 illustrates a case where components within the separation target range R are suppressed.
  • the sound collection device 126 generates a recording signal V indicating a time waveform of the user's singing sound. That is, the recorded signal V is an acoustic signal indicating the recorded sound recorded in parallel with the reproduction of the accompaniment sound of the music.
  • the illustration of a D / A converter that converts the acoustic signal SB into an analog signal and an A / D converter that generates a digital recording signal V from the analog signal are omitted for convenience.
  • the separation data processing device 14A in FIG. 1 is a server device (typically a web server) that manages separation data Q applied to separation processing in each terminal device 12, and includes a control device 142 and a storage device 144. And a communication system 146.
  • the separation data processing device 14A can be realized by a plurality of devices configured separately from each other (for example, a plurality of server devices communicating with each other via the communication network 16).
  • the control device 142 executes various control processes and arithmetic processes by executing programs stored in the storage device 144.
  • the communication device 146 communicates with each terminal device 12 via the communication network 16.
  • the storage device 144 stores a program executed by the control device 142 and various data used by the control device 142.
  • a recording medium such as a semiconductor recording medium or a magnetic recording medium or a combination of a plurality of types of recording media may be employed as the storage device 144.
  • the storage device 144 is installed in an external device (for example, an external server device) separate from the separation data processing device 14A, and the separation data processing device 14A writes information to the storage device 144 via the communication network 16.
  • an external device for example, an external server device
  • a configuration for executing reading may be employed.
  • the storage device 144 of the first embodiment stores a plurality of separation data Q corresponding to different music pieces.
  • the separation data Q is setting data applied to the separation processing by each terminal device 12, and is used to specify a separation target by the separation processing in the acoustic signal SA, for example.
  • the separation data Q of the first embodiment designates the separation target range R in FIG. That is, the target frequency range RF (upper limit value and lower limit value on the frequency axis XF) and the target localization range RL (upper limit value and lower limit value on the localization axis XL) are specified by the separation data Q.
  • the separation data Q is individually generated for each music so that the singing sound of the music in the acoustic signal SA is included in the separation target. Since the frequency band and sound image position ⁇ of the singing sound can be different for each music piece, the target frequency range RF and the target localization range RL specified by the separation data Q are different for each music piece.
  • the separation data Q designates one separation target range R is exemplified, but the separation data Q designates a plurality of separation target ranges R (that is, different in the acoustic signal SA).
  • a configuration in which a plurality of components are separated may also be employed. As shown in FIG. 1, each separation data Q is associated with attribute data MB.
  • the attribute data MB is information related to the music (for example, attribute information such as the name of the music and the name of the singer), like the attribute data MA.
  • the attribute data MB can also specify information related to the separation target of the acoustic signal SA (for example, part names such as vocals and guitars).
  • FIG. 3 is an explanatory diagram of a schematic operation of the sound processing system 100.
  • the user of the terminal device 12 instructs the reproduction of a desired song (hereinafter referred to as “target song”) among a plurality of songs whose acoustic signal SA is stored in the storage device 122 by appropriately operating the input device 125. To do.
  • target song a desired song
  • the terminal device 12 requests the separation data processing device 14A for the separation data Q of the target song (S1), and the separation data processing device 14A responds to the request from the terminal device 12.
  • the separation data Q of the target song is transmitted to the requesting terminal device 12 (S2).
  • the terminal device 12 executes the separation process to which the separation data Q provided from the separation data processing device 14A is applied to the acoustic signal SA of the target song (S3). Then, the terminal device 12 generates update data U according to the result of the separation process and transmits it to the separation data processing apparatus 14A (S4).
  • the update data U is generally generated according to an input operation (for example, singing in parallel with the reproduced sound) performed by the user who has received the reproduced sound after the separation process in relation to the reproduced sound.
  • the separation data processing device 14A updates the separation data Q of the target song using the update data U transmitted from the terminal device 12 (S5). The above processing is executed every time each terminal apparatus 12 is instructed to reproduce the target song.
  • the update data U is transmitted from each of the plurality of terminal devices 12 to the separation data processing device 14A. Therefore, the separation data Q of each song stored in the storage device 144 of the separation data processing device 14A is cumulatively updated every time the song is played in each terminal device 12, and the separation data Q is applied.
  • the accuracy with which the singing sound is separated by the process (hereinafter referred to as “reliability of the separation data Q”) improves with time. Details of the configuration and operation outlined above will be described below.
  • FIG. 4 is a functional configuration diagram of each terminal device 12 and the separation data processing device 14A in the first embodiment.
  • the control device 121 of the terminal device 12 executes the program (music reproduction program) stored in the storage device 122 to reproduce the acoustic signal SA and generate the update data U.
  • a plurality of functions are realized.
  • a configuration in which each function of the control device 121 is distributed over a plurality of integrated circuits, or a configuration in which a dedicated electronic circuit (for example, a DSP) realizes a part of the function of the control device 121 may be employed.
  • the separation data acquisition unit 22 acquires the separation data Q of the target song from the separation data processing device 14A. Specifically, the separation data acquisition unit 22 transmits a request including the attribute data MA of the target song selected by the user of the terminal device 12 from the communication device 123 to the separation data processing device 14A (FIG. 3). In step S1), in response to a request from the terminal device 12, the separation data Q transmitted from the separation data processing device 14A is acquired from the communication device 123 (step S2 in FIG. 3).
  • the separation processing unit 24 performs a separation process to which the separation data Q acquired by the separation data acquisition unit 22 is applied to the acoustic signal SA of the target music (step S3 in FIG. 3).
  • the separation processing unit 24 includes a suppression processing unit 242 and an enhancement processing unit 244.
  • the suppression processing unit 242 generates an acoustic signal SB that suppresses the singing sound of the acoustic signal SA of the target song and emphasizes the accompaniment sound.
  • the enhancement processing unit 244 generates an acoustic signal SC in which the singing sound of the acoustic signal SA is enhanced (ideally extracted) and the accompaniment sound is suppressed (ideally removed). That is, the separation process of the first embodiment includes a suppression process that suppresses the singing sound and an enhancement process that emphasizes the singing sound. Specific examples of the separation process will be described in detail below.
  • the separation processing unit 24 calculates the sound image position ⁇ for each frequency component of the acoustic signal SA of the target song.
  • a known technique can be arbitrarily adopted. For example, as disclosed in Patent Document 1, calculation using the intensity ratio of each channel of the acoustic signal SA is suitable.
  • the suppression processing unit 242 is a frequency component within the target frequency range RF specified by the separation data Q among the plurality of frequency components of the acoustic signal SA and is within the target localization range RL specified by the separation data Q.
  • the acoustic signal SB is generated by suppressing the intensity of the frequency component including the sound image position ⁇ .
  • the enhancement processing unit 244 out of the plurality of frequency components of the acoustic signal SA is outside the target localization range RL specified by the frequency component outside the target frequency range RF specified by the separation data Q or the separation data Q.
  • the acoustic signal SC is generated by suppressing the intensity of the frequency component including the sound image position ⁇ . It is also possible to generate the other of the acoustic signal SB and the acoustic signal SC by subtracting one of the acoustic signal SB and the acoustic signal SC from the acoustic signal SA.
  • the reproduction control unit 26 mixes the acoustic signal SB generated by the separation processing unit 24 (suppression processing unit 242) and the recording signal V supplied from the sound collection device 126 and applies various acoustic effects (e.g., echoes). After performing acoustic processing such as amplification of signal intensity, the sound emitting device 127 reproduces the sound wave. The user sings the melody (singing part) of the target song in parallel with the reproduction of the accompaniment sound of the target song indicated by the separated acoustic signal SB.
  • a mixed sound of the accompaniment sound (acoustic signal SB) in which the singing sound is suppressed from the acoustic signal SA of the music and the user's singing sound (recorded signal V) is reproduced from the sound emitting device 127.
  • the analysis processing unit 30 generates update data U for the target song.
  • the update data U of the first embodiment is generated according to the recorded signal V generated by the sound pickup device 126 in parallel with the reproduced sound of the acoustic signal SB (that is, the singing sound synchronized with the accompaniment sound of the music).
  • the analysis processing unit 30 of the first embodiment includes an evaluation processing unit 32 and a conversion processing unit 34.
  • the evaluation processing unit 32 compares the recorded signal V generated by the sound collection device 126 with the acoustic signal SC in which the separation processing unit 24 (enhancement processing unit 244) emphasizes the singing sound, and evaluates the evaluation data DS according to the comparison result. Generate. That is, the evaluation processing unit 32 compares the singing sound included in the acoustic signal SA with the user's singing sound recorded by the sound collecting device 126, and responds to the similarity (typically correlation) between the two. The evaluation data DS is calculated. Since the singing sound included in the acoustic signal SA corresponds to an exemplary or standard singing sound of the music, the evaluation processing unit 32 performs the skill of the user's singing (approximate or different from the exemplary or standard singing).
  • a known singing evaluation technique for example, a technique for comparing feature quantities such as pitch and volume among a plurality of acoustic signals
  • the evaluation (score) indicated by the evaluation data DS generated by the evaluation processing unit 32 is displayed on the display device 124.
  • the conversion processing unit 34 generates recording data DV corresponding to the recording signal V generated by the sound collection device 126. Specifically, the conversion processing unit 34 generates recorded data DV in the MIDI (Musical Instrument Digital Interface) format that represents the recorded signal V. That is, the recorded data DV includes event data that specifies the pitch (note number) and intensity (velocity) of each note of the singing sound indicated by the recorded signal V, and the processing time ( For example, it is time-series data in which a plurality of sets of time data specifying the processing interval of each event data in sequence are arranged. A known conversion technique (Audio-MIDI conversion) is arbitrarily employed to generate the recorded data DV.
  • MIDI Musical Instrument Digital Interface
  • the update data U of the target song including the evaluation data DS generated by the evaluation processing unit 32 and the recorded data DV generated by the conversion processing unit 34 is transmitted from the communication device 123 of the terminal device 12 to the communication network. 16 is transmitted to the separation data processing device 14A via step 16 (step S4 in FIG. 3). That is, the recorded data DV corresponding to the singing sound of the user of the terminal device 12 and the evaluation data DS indicating the evaluation result (skill) of the singing sound are transmitted as the update data U.
  • the control device 142 of the separation data processing device 14A executes the program (separation data update program) stored in the storage device 144, so that the separation data Q for each terminal device 12 is stored.
  • a plurality of functions for updating the separation data Q using the provision and update data U are realized.
  • a configuration in which each function of the control device 142 is distributed over a plurality of integrated circuits, or a configuration in which a dedicated electronic circuit (for example, DSP) realizes a part of the function of the control device 142 may be employed.
  • the separation data providing unit 42 provides the terminal device 12 with the separation data Q for the target song. Specifically, the separation data providing unit 42 sets the attribute data of the target song corresponding to (typically matches) the attribute data MA specified by the request (step S1 in FIG. 3) transmitted from the terminal device 12. MB is searched from the storage device 144, and the separation data Q of the target music corresponding to the attribute data MB is transmitted from the communication device 146 to the requesting terminal device 12 (step S2 in FIG. 3).
  • the update data acquisition unit 44 acquires the update data U of the target song from each of the plurality of terminal devices 12 via the communication network 16 and the communication device 146. Specifically, the update data acquisition unit 44 acquires from the terminal device 12 update data U including the target song evaluation data DS and recorded data DV generated by the analysis processing unit 30 of the terminal device 12.
  • the update processing unit 46 uses the target song separation data Q stored in the storage device 144 using the target song update data U (recorded data DV, evaluation data DS) acquired by the update data acquisition unit 44. (Step S5 in FIG. 3). For example, the update processing unit 46 specifies the pitch range of the singing sound specified by the recording data DV, and the target frequency range RF specified by the separation data Q of the target song approaches the pitch range of the recording data DV. The data Q for separating the target music is updated. In addition, for example, a configuration in which the separation data Q is updated so that each harmonic component having a pitch specified by the recorded data DV as a fundamental frequency is suppressed, or a non-harmonic sound such as a percussion instrument sound using the recorded data DV.
  • a configuration that excludes (recovers) from the suppression target may also be employed.
  • a technique for improving separation accuracy using time series data such as MIDI data is also disclosed in, for example, Japanese Patent Application Laid-Open No. 2012-108453.
  • the recorded data DV used for updating the separation data Q can be stored in the storage device 144 for each music piece.
  • the recorded data DV accumulated for each music is used for various processes such as music search. For example, a configuration in which a musical piece in which recorded data DV similar to the melody specified by the user by the operation of the input device 125 is stored is searched as a target song, or separation data of the target song searched using the recorded data DV
  • a configuration in which the separation data providing unit 42 transmits Q to the terminal device 12 (S2) is preferable.
  • the skill of singing is different for each user of each terminal device 12.
  • the reliability (quality) of the separation data Q is improved, but the recording data DV of a user who is poor at singing is separated.
  • the reliability of the separation data Q may be lowered.
  • the update processing unit 46 of the first embodiment applies the evaluation data DS indicating the skill of the user's singing to the update of the separation data Q by the recorded data DV.
  • the update processing unit 46 records data for updating the separation data Q as the evaluation indicated by the evaluation data DS is higher (the recorded sound indicated by the recorded signal V is more similar to the singing sound indicated by the acoustic signal SC).
  • the separation data Q of the target song is updated according to the recorded data DV (update data U) so that the influence (statistical weighting) of DV increases.
  • the evaluation indicated by the evaluation data DS is below a threshold value (for example, when the user's singing is excessively poor or when the user is not singing at the time of reproducing the acoustic signal SB)
  • the recorded data DV of the update data U is separated. Data Q is not reflected.
  • the evaluation data DS of the first embodiment is the degree to which the recorded data DV (recorded signal V) indicating the user's singing sound contributes to the improvement of the reliability of the separation data Q (singing). It corresponds to an index indicating the sound validity) and is used as a weight value of the recorded data DV when the separation data Q is updated.
  • the above processing is executed each time the user of each terminal device 12 instructs the reproduction of the target song. That is, the separation data processing device 14A updates the data U for updating according to the input operation (singing the target song) of the user who has listened to the reproduced sound of the acoustic signal SB after the separation processing to which the existing separation data Q is applied.
  • the separation data processing device 14A updates the data U for updating according to the input operation (singing the target song) of the user who has listened to the reproduced sound of the acoustic signal SB after the separation processing to which the existing separation data Q is applied.
  • the separation data Q of each music in the storage device 144 is updated using the update data U acquired from each terminal device 12.
  • the separation data Q to which the update data U is applied is sequentially updated.
  • the separation data Q of each song in the storage device 144 is updated using the update data U acquired from each of the plurality of terminal devices 12. That is, the singing of a plurality of users of each terminal device 12 is reflected in the separation data Q. Therefore, there is an advantage that the reliable separation data Q that can separate the specific component (singing sound) of the acoustic signal SA with high accuracy can be easily generated.
  • the influence of the recorded data DV on the update of the separation data Q increases. To do. Therefore, it is possible to efficiently improve the reliability of the separation data Q for each musical piece as compared with the configuration in which the recorded data DV is reflected in the separation data Q regardless of the skill of the user of each terminal device 12. There are also advantages.
  • Second Embodiment A second embodiment of the present invention will be described.
  • elements having the same functions and functions as those of the first embodiment are referred to in the description of the first embodiment, and detailed descriptions thereof are appropriately omitted.
  • FIG. 5 is a functional configuration diagram of each terminal device 12 and separation data processing device 14A in the second embodiment.
  • the control device 121 of the terminal device 12 of the second embodiment includes the same elements (separation data acquisition unit 22, separation processing unit 24, reproduction control unit 26, analysis processing unit as in the first embodiment. 30) In addition to functioning as the display control unit 28.
  • the display control unit 28 causes the display device 124 to display a separation processing image representing the separation processing to which the separation data Q is applied.
  • the separation processing image is an image that presents the distribution (scatter diagram) of each frequency component of the acoustic signal SB (or acoustic signal SA) on the localization-frequency plane and the separation target range R to the user. It is.
  • the user can adjust the separation target range R of the separation processed image by appropriately operating the input device 125 while confirming the separation processing image displayed on the display device 124.
  • the separation target range R By adjusting the separation target range R, the volume of the singing sound in the reproduced sound of the acoustic signal SB after the separation process is increased or decreased.
  • the user operates (adjusts) the input device 125 while listening to the reproduced sound of the acoustic signal SB radiated from the sound emitting device 127, so that the singing sound of the target song is reduced with high accuracy by the separation process.
  • the target frequency range RF and the target localization range RL of the separation target range R are adjusted so that the volume of the singing sound in the reproduced sound is sufficiently reduced.
  • the separation target range R is set so that the singing sound of the target song is suppressed with high accuracy, the acoustic signal SC after the enhancement process is suppressed from mixing components other than the singing sound.
  • the evaluation of the recording signal V tends to increase (the evaluation of the singing sound decreases as the separation target range R is inappropriate). That is, the similarity evaluation of the acoustic signal SC after the enhancement process and the recorded signal V of the user's singing sound (presentation of the singing evaluation result to the user) is performed so that the singing sound of the target song is highly accurately suppressed. This acts as an incentive for causing the user to properly perform the adjustment operation for adjusting the separation target range R.
  • the analysis processing unit 30 of the second embodiment includes an adjustment management unit 36 in addition to the same elements (evaluation processing unit 32, conversion processing unit 34) as in the first embodiment.
  • the adjustment management unit 36 generates adjustment data DC corresponding to the user's adjustment operation for the separation target range R.
  • the adjustment management unit 36 generates adjustment data DC that specifies the separation target range R (target frequency range RF, target localization range RL) after the user's adjustment operation.
  • the generation time of the adjustment data DC by the adjustment management unit 36 is arbitrary. For example, a configuration in which the adjustment data DC is generated when the reproduction of the target song is completed, or a configuration in which the adjustment data DC is generated at a point in the middle of the target song (for example, at the time of an interlude) can be adopted.
  • the update data for the target song including the adjustment data DC generated by the adjustment management unit 36.
  • U is transmitted from the terminal device 12 to the separation data processing device 14A (step S4 in FIG. 3).
  • the update data acquisition unit 44 of the separation data processing device 14A acquires update data U from each of the plurality of terminal devices 12, and the update processing unit 46 acquires each update data acquired by the update data acquisition unit 44.
  • U is used to update the separation data Q of the target song.
  • the update processing unit 46 of the second embodiment updates the separation data Q for the target song in accordance with the recorded data DV and the evaluation data DS of the update data U, and adjusts the terminal device 12 as well as the first embodiment.
  • the adjustment data DC generated by the management unit 36 is used to update the separation data Q for the target song.
  • the update processing unit 46 separates the separation target range R (target frequency range RF, target localization range RL) specified by the separation data Q of the target song so as to approach the separation target range R specified by the adjustment data DC. ).
  • the separation target range R specified by the separation data Q of each music piece is used by each terminal device 12 by repeatedly updating the separation data Q using the adjustment data DC transmitted from the plurality of terminal devices 12. Is adjusted to an average range of a plurality of separation target ranges R specified in the past for the music.
  • the same effect as in the first embodiment is realized.
  • the adjustment data DC corresponding to the adjustment operation by the listener of the reproduced sound after separation processing to which the separation data Q is applied is applied to the update of the separation data Q, the acoustic signal SA
  • the update data U including the evaluation data DS, the recorded data DV, and the adjustment data DC is illustrated.
  • the update data U includes only the adjustment data DC (evaluation data DS and recording).
  • a configuration in which the data DV is omitted) may be employed. That is, the evaluation processing unit 32 or the conversion processing unit 34 can be omitted.
  • ⁇ Third Embodiment> an acoustic processing system 100 in which the separation data processing device 14A receives the update data U from each terminal device 12 via the communication network 16 and updates the separation data Q is illustrated. did.
  • the separation data processing device 14B of the third embodiment executes the generation of the update data U and the update of the separation data Q by itself.
  • FIG. 6 is a block diagram of the separation data processing device 14B of the third embodiment.
  • the separation data processing device 14B is realized by a computer system including a control device 181, a storage device 182, a display device 184, an input device 185, a sound collection device 186, and a sound emission device 187.
  • an information processing device such as a mobile phone, a smartphone, or a personal computer is used as the separation data processing device 14B.
  • the transmission / reception of the separation data Q or the update data U via the communication network 16 is not required in the third embodiment, it does not matter whether the separation data processing device 14B has a communication function.
  • Storage device 182 stores acoustic signal SA and separation data Q for each of a plurality of music pieces.
  • the display device 184 displays the separation processed image in the same manner as the display device 124 of the second embodiment, and the input device 185 gives an instruction from the user to select a target song, as in the input device 125 of the first embodiment. Accept.
  • the sound collecting device 186 generates the recording signal V in the same manner as the sound collecting device 126 in the first embodiment, and the sound emitting device 187 is the acoustic signal SB after the separation processing, like the sound emitting device 127 in the first embodiment. Play.
  • the control device 181 executes the program (music reproduction program, separation data update program) stored in the storage device 182 to reproduce the acoustic signal SA and generate the update data U and update the separation data Q.
  • a plurality of functions (separation processing unit 54, reproduction control unit 56, display control unit 58, update data acquisition unit 60, update processing unit 70) are realized.
  • a configuration in which the functions of the control device 181 are distributed over a plurality of integrated circuits, or a configuration in which a dedicated electronic circuit (for example, a DSP) realizes a part of the functions of the control device 181 may be employed.
  • the separation processing unit 54 performs separation processing (suppression processing) using the target song separation data Q specified by the user.
  • Emphasis processing is performed on the acoustic signal SA of the target song to enhance (ideally extract) the acoustic signal SB and the singing sound that suppress (ideally remove) the singing sound of the acoustic signal SA.
  • the generated acoustic signal SC is generated.
  • the playback control unit 56 mixes the acoustic signal SB generated by the separation processing unit 54 (suppression processing unit 542) and the recording signal V supplied from the sound collection device 186. Then, the sound emitting device 187 is made to reproduce. Similar to the first embodiment, the user sings the melody (singing part) of the music while listening to the accompaniment sound of the target music reproduced by the sound emitting device 187.
  • the update data acquisition unit 60 acquires the update data U for the target song.
  • the update data acquisition unit 60 includes an evaluation processing unit 62, a conversion processing unit 64, and an adjustment management unit 66, like the analysis processing unit 30 of the second embodiment, and the evaluation data DS.
  • the update data U including the recorded data DV and the adjustment data DC is generated.
  • the evaluation processing unit 62 generates the evaluation data DS corresponding to the comparison result between the acoustic signal SC and the recorded signal V in the same manner as the evaluation processing unit 32 of the first embodiment, and the conversion processing unit 64 is the same as that of the first embodiment.
  • the recording data DV corresponding to the recording signal V is generated.
  • the adjustment management unit 66 also generates adjustment data DC that specifies the separation target range R instructed by the user, similarly to the adjustment management unit 36 of the second embodiment.
  • the update processing unit 70 updates the separation data Q of the target song using the update data U acquired (generated) by the update data acquisition unit 60, as in the update processing unit 46 of the first embodiment.
  • the method for updating the separation data Q is the same as in the first embodiment. Every time the user gives an instruction to play a song, the generation of the update data U and the update of the separation data Q are executed. That is, as in the first embodiment, the separation data Q for each music piece is repeatedly updated a plurality of times. Therefore, also in the third embodiment, the same effect as the first embodiment and the second embodiment is realized.
  • the update data acquisition unit 44 of the first embodiment and the update data acquisition unit 60 of the third embodiment reproduce the acoustic signal SB after the separation processing to which the separation data Q is applied. It is included as an element for acquiring the update data U reflecting the input by the user who has listened to the sound (singing to the sound collection device or instruction to the input device). That is, the element (update data acquisition unit) that acquires the update data U acquires the update data U from the external device (terminal device 12) as in the first embodiment, or as in the third embodiment. It does not matter whether the update data U is generated.
  • the separation data Q is prepared for each of a plurality of music pieces, but it is not necessary to prepare the separation data Q for each music piece individually.
  • a configuration in which separation data Q is prepared for each attribute of music for example, singer, genre, recording format
  • Common separation data Q is applied to separation processing of a plurality of music pieces having common attributes. According to the above configuration, since the common separation data Q is applied to a plurality of music pieces, the capacity required for storing the separation data Q is larger than the configuration of holding the separation data Q for each music piece. There is an advantage that it is reduced.
  • the separation target range R does not change over the entire section of the music
  • a configuration in which the separation target range R is changed with time in the music may be employed. That is, the separation data Q specifies a time series (temporal transition) of the separation target range R.
  • the update data U is repeatedly generated in the music and applied to the update of each separation target range R specified by the separation data Q.
  • the specific contents of the separation process are not limited to the examples of the above embodiments, and any known acoustic processing technique (for example, sound source separation technique) that separates (suppresses or emphasizes) a specific component from the acoustic signal SA is arbitrarily used. Adopted.
  • the format of the separation data Q or the update data U is appropriately changed according to the content of the separation processing.
  • time-series data in MIDI format for example, recorded data DV
  • a configuration that separates from the acoustic signal SA as a component may also be employed.
  • the recording data DV corresponding to the recording signal V (or the average data of the plurality of recording data DV transmitted from each terminal device 12) is stored as the separation data Q.
  • the acoustic signal SA of each song is stored in the storage device 122 of the terminal device 12, but the acoustic signal SA of each song is stored in the storage device 144 of the separation data processing device 14A. Can also be stored.
  • the separation data processing device 14A separation data providing unit 42
  • the separation processing unit 24 that applies the separation data Q to the acoustic signal SA of the target song stored in the storage device 144 is subjected to separation data processing. It can also be installed in the device 14A.
  • the acoustic signals (SB, SC) after separation processing by the separation processing unit 24 of the separation data processing device 14A are transmitted to the terminal device 12 via the communication network 16.
  • the acoustic signal SB is reproduced from the sound emitting device 127 by the reproduction control unit 26, and the acoustic signal SC is applied to the evaluation of the recorded signal V by the evaluation processing unit 32.
  • the separation processing unit 24 installed in the separation data processing device 14A can process the acoustic signal SA received from the terminal device 12 via the communication network 16.
  • the suppression processing unit 242 and the enhancement processing unit 244 included in the separation processing unit 24 are installed in the separation data processing device 14A.
  • a configuration in which the enhancement processing unit 244 and the evaluation processing unit 32 are installed in the separation data processing device 14A and the suppression processing unit 242 is installed in the terminal device 12 is employed.
  • the recorded signal V of the singing sound is transmitted from the terminal device 12 to the separation data processing device 14A and received from the terminal device 12 and the acoustic signal SC generated by the enhancement processing unit 244 of the separation data processing device 14A.
  • the song evaluation unit 32 of the separation data processing device 14A generates the evaluation data DS.
  • the acoustic signal SA in the storage device 144 is converted into the separation data.
  • the acoustic signal SB is generated by the suppression processing unit 242 of the terminal device 12 after being applied to the generation of the acoustic signal SC by the enhancement processing unit 244 of the processing device 14A and transmitted from the separation data processing device 14A to the terminal device 12.
  • the acoustic signal SA is stored in the storage device 122 of the terminal device 12 under the configuration in which the enhancement processing unit 244 is installed in the separation data processing device 14A
  • the acoustic signal SA in the storage device 122 is
  • the acoustic signal SB is applied to the generation of the acoustic signal SB by the suppression processing unit 242 of the terminal device 12 and is transmitted from the terminal device 12 to the separation data processing device 14A, and then the acoustic by the enhancement processing unit 244 of the separation data processing device 14A.
  • the suppression processing unit 242 can also be installed in the separation data processing device 14A.
  • the user's adjustment operation on the input device 125 can be reflected in the separation data Q applied to the separation process by the separation processing unit 24 as needed. It is.
  • the specific component that is suppressed from the acoustic signal SA by the separation process is changed in real time according to the adjustment operation by the user. Therefore, the user can execute the adjustment operation while actually listening to the reproduced sound of the acoustic signal SB and confirming the effect of the adjustment operation (whether or not the desired specific component is appropriately suppressed). is there.
  • FIG. 8 is a block diagram of the analysis processing unit 30 according to a modification of the first embodiment.
  • the conversion processing unit 38 in FIG. 8 generates conversion data DA corresponding to the acoustic signal SC generated by the separation processing unit 24 (enhancement processing unit 244).
  • the converted data DA is, for example, time-series data in the MIDI format that represents the acoustic signal SC. 8 selects either the acoustic signal SC generated by the separation processing unit 24 or the converted data DA generated by the conversion processing unit 38 as reference data DREF.
  • the conversion processing unit 34 in FIG. 8 generates conversion data DB corresponding to the recording signal V generated by the sound collection device 126.
  • the conversion data DB is, for example, time-series data in the MIDI format that represents the recorded signal V, similarly to the recorded data DV in the first and second embodiments.
  • the selection unit 84 selects either the recording signal V generated by the sound collection device 126 or the conversion data DB generated by the conversion processing unit 34 as the recording data DV.
  • the evaluation processing unit 32 generates the evaluation data DS by comparing the reference data DREF selected by the selection unit 82 with the recorded data DV selected by the selection unit 84.
  • the operation (selection target) of the selection unit 82 and the selection unit 84 is controlled in accordance with an instruction from the user to the input device 125, for example.
  • the user can select the first evaluation process and the second evaluation process.
  • the selection unit 82 selects the acoustic signal SC as the reference data DREF
  • the selection unit 84 selects the recording signal V as the recording data DV. Therefore, the evaluation processing unit 32 generates the evaluation data DS corresponding to the comparison result between the acoustic signal SC and the recorded signal V (first evaluation process).
  • the selection unit 82 selects the conversion data DA as the reference data DREF, and the selection unit 84 selects the conversion data DB as the recorded data DV. Therefore, the evaluation processing unit 32 generates evaluation data DS corresponding to the comparison result between the conversion data DA and the conversion data DB (second evaluation process). In any of the first evaluation process and the second evaluation process, the evaluation data DS corresponding to the comparison result between the acoustic signal SC (acoustic signal SC or converted data DA) and the recorded signal V (recorded signal V or converted data DB) is obtained. Generated.
  • the configuration in which the update data acquisition unit 44 of the separation data processing device 14A receives the update data U generated by the analysis processing unit 30 of the terminal device 12 is exemplified.
  • the update data acquisition unit 44 of the separation data processing device 14 ⁇ / b> A can generate the update data U with the same configuration and operation as the analysis processing unit 30.
  • the update data acquisition unit 44 receives the recording signal V generated by the sound collection device 126 of the terminal device 12 and the acoustic signal SC in which the singing sound is emphasized from the terminal device 12, and receives the recording signal V and The evaluation data DS corresponding to the comparison result with the acoustic signal SC is generated (evaluation processing unit 32) and the recording data DV corresponding to the recording signal V is generated (conversion processing unit 34). Further, the update data acquisition unit 44 acquires an instruction (adjustment operation) from the user for the input device 125 of the terminal device 12 from the terminal device 12, and generates adjustment data DC corresponding to the adjustment operation (adjustment management unit) 36). The update processing unit 46 updates the separation data Q of the target song according to the update data U (evaluation data DS, recorded data DV, adjustment data DC) generated by the update data acquisition unit 44.
  • the update data acquisition unit 44 of the first embodiment and the second embodiment is performed by a user who has listened to the reproduced sound of the acoustic signal SB after the separation process to which the separation data Q is applied. It is included as an element for acquiring the update data U reflecting the input (singing to the sound collection device 126 and instruction to the input device 125), and the update data U from the external device (terminal device 12) as in the first embodiment. It does not matter whether or not it generates the update data U as illustrated in the modification.
  • the separation of the singing sound of the acoustic signal SA is exemplified, but the component separated from the acoustic signal SA is not limited to the singing sound.
  • the configuration for separating the performance sound of a specific instrument from the acoustic signal SA is preferably used for practice of the performance of the instrument. It is also possible to separate a plurality of components (for example, performance sounds of a plurality of musical instruments) of the acoustic signal SA.
  • the separation processing for the acoustic signal SA of the music is illustrated, but a musical viewpoint is not essential in the present invention.
  • a specific component for example, a voice of a specific speaker among a plurality of speaker's conversational sounds
  • SA for language learning in which a conversational sound or the like of a specific language is recorded.
  • DESCRIPTION OF SYMBOLS 100 ... Acoustic processing system, 12 ... Terminal device, 14 ... Separation data processing device, 121, 142, 181 ... Control device, 122, 144, 182 ... Storage device, 123, 146 ... Communication device, 124,184 ... display device, 125,185 ... input device, 126,186 ... sound collecting device, 127,187 ... sound emitting device, 16 ... communication network, 22 ... separation data acquisition unit, 24 , 54... Separation processing unit, 26, 56... Reproduction control unit, 30... Analysis processing unit, 32, 62... Evaluation processing unit, 34, 64 ... Conversion processing unit, 36, 66. . 42... Separation data providing unit 44, 60... Updating data acquisition unit 46, 70.

Landscapes

  • Engineering & Computer Science (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Computational Linguistics (AREA)
  • Quality & Reliability (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Reverberation, Karaoke And Other Acoustics (AREA)

Abstract

 記憶装置144は、音響信号SAの特定成分を強調または抑圧する分離処理に適用される分離用データQを記憶する。更新用データ取得部44は、分離用データQを適用した分離処理後の音響信号SBの再生音の受聴者による入力が反映された更新用データUを、複数の端末装置12の各々から通信網16を介して取得する。更新処理部46は、更新用データ取得部44が取得した各更新用データUを利用して記憶装置144の分離用データQを更新する。

Description

分離用データ処理装置およびプログラム
 本発明は、音響信号の特定の成分(以下「特定成分」という)を分離(強調または抑圧)する技術に関する。
 複数の音響の混合音を示す音響信号の特定成分を抑圧する技術が従来から提案されている。例えば特許文献1には、音響信号の音像が定位する位置を周波数成分毎に算定し、音像が特定の範囲に定位する周波数成分を音響信号から抑圧する技術が開示されている。
日本国特開2012-163861号公報
 しかし、特許文献1の技術において、音響信号の所望の成分が高精度に抑圧されるように周波数および音像位置を適切に指定することは実際には容易ではない。また、抑圧対象となる歌唱音等の音響成分の周波数や音像位置は例えば楽曲毎に相違し得るから、例えば相異なる楽曲に対応する各音響信号について特定の成分を高精度に抑圧するには、周波数や音像位置を楽曲毎に個別に設定する必要があり現実には困難である。以上の事情を考慮して、本発明は、音響信号の特定成分を高精度に分離(抑圧または強調)可能な分離用データを簡便に生成することを目的とする。
 以上の課題を解決するために、本発明の分離用データ処理装置は、音響信号の特定成分を強調または抑圧する分離処理に適用される分離用データを記憶する記憶部(例えば記憶装置144)と、分離用データを適用した分離処理後の音響信号(例えば音響信号SB)の再生音を受聴した利用者による入力が反映された更新用データを取得する更新用データ取得部と、更新用データ取得部が取得した更新用データを利用して記憶部の分離用データを更新する更新処理部とを具備する。
 以上の態様では、分離用データを適用した分離処理後の音響信号(例えば音響信号SB)の再生音の受聴者による入力が反映された更新用データを利用して分離用データが更新される。したがって、音響信号の特定成分を高精度に分離可能な信頼性の高い分離用データを簡便に生成できるという利点がある。
 更新用データ取得部は、分離用データを適用した分離処理で特定成分が抑圧された音響信号の再生に並行して収録された収録音(すなわち、分離処理後の音響信号の再生音に並行して利用者が歌唱した歌唱音)に応じた収録データを含む更新用データを取得してもよい。
 以上の態様では、特定成分を抑圧した音響信号の再生に並行して収録された収録音に応じた収録データを含む更新用データを利用して分離用データが更新される。
更新用データ取得部は、分離用データを適用した分離処理で特定成分を抑圧した音響信号の再生に並行して収録された収録信号と、分離用データを適用した分離処理で特定成分を強調した音響信号との比較結果に応じた評価データと、収録信号に応じた収録データとを含む更新用データを取得してもよい。
 以上の態様では、特定成分を抑圧した音響信号の再生に並行して収録された収録信号と特定成分を強調した音響信号との比較結果に応じた評価データが更新用データとして分離用データの更新に適用される。したがって、各端末装置の利用者の発音(典型的には歌唱)の巧拙を反映して信頼性の高い分離用データを生成できるという利点がある。
更新処理部は、収録信号と特定成分を強調した音響信号とが類似するほど分離用データの更新に対する収録データの影響(統計的な重み付け)が増加するように、更新用データに応じて分離用データを更新してもよい。
以上の態様によれば、特定成分を強調した音響信号とは音響特性が乖離した収録信号の影響を低減して信頼性の高い分離用データを効率的に生成できるという利点がある。
 更新用データ取得部は、分離用データを適用した分離処理後の音響信号の再生音の受聴者による調整指示に応じた調整データを含んでもよい。
以上の態様では、分離処理後の音響信号の再生音の受聴者による調整指示に応じた調整データが更新用データとして分離用データの更新に適用される。したがって、分離処理後の音響信号の再生音を利用者が聴取した結果を反映した信頼性の高い分離用データを生成できるという利点がある。
 更新用データ取得部は、1以上の端末装置から通信網を介して更新用データを取得してもよい。
 以上の態様では、1以上の端末装置から通信網を介して取得された更新用データが分離用データの更新に利用されるから、音響信号の特定成分を高精度に分離可能な信頼性の高い分離用データを簡便に生成できるという前述の効果は格別に顕著である。
 本発明において、音響信号の特定成分を強調または抑圧する分離処理に適用される分離用データを記憶し、前記分離用データを適用した分離処理で前記特定成分が抑圧された音響信号の再生に並行して収録された収録音に応じた収録データを含む更新用データを取得し、前記取得した更新用データを利用して前記した分離用データを更新する分離用データ処理方法が提供される。
 以上の各態様に係る分離用データ処理装置は、音響信号の処理に専用されるDSP(Digital Signal Processor)などのハードウェア(電子回路)によって実現されるほか、CPU(Central Processing Unit)等の汎用の演算処理装置とプログラムとの協働によっても実現される。具体的には、本発明に係るプログラムは、音響信号の特定成分を強調または抑圧する分離処理に適用される分離用データを記憶する記憶部を具備するコンピュータを、分離用データを適用した分離処理後の音響信号の再生音の受聴者による入力が反映された更新用データを取得する更新用データ取得部、および、更新用データ取得部が取得した各更新用データを利用して記憶部の分離用データを更新する更新処理部をコンピュータに実行させる。
以上の各態様に係るプログラムは、コンピュータが読取可能な記録媒体に格納された形態で提供されてコンピュータにインストールされ得る。記録媒体は、例えば非一過性(non-transitory)の記録媒体であり、CD-ROM等の光学式記録媒体(光ディスク)が好例であるが、半導体記録媒体や磁気記録媒体等の公知の任意の形式の記録媒体を包含し得る。また、例えば、本発明のプログラムは、通信網を介した配信の形態で提供されてコンピュータにインストールされ得る。
本発明の第1実施形態に係る音響処理システムのブロック図である。 分離処理の説明図である。 音響処理システムの概略的な動作の説明図である。 各端末装置および分離用データ処理装置の機能的な構成図である。 第2実施形態における各端末装置および分離用データ処理装置の機能的な構成図である。 本発明の第3実施形態に係る分離用データ処理装置の機能的な構成図である。 変形例に係る各端末装置および分離用データ処理装置の機能的な構成図である。 変形例に係る分離用データ処理装置の機能的な構成図である。
<第1実施形態>
 図1は、本発明の第1実施形態に係る音響処理システム100のブロック図である。図1に示すように、音響処理システム100は、複数の端末装置12と分離用データ処理装置14Aとを具備する通信システムである。各端末装置12は、例えば携帯電話機やスマートフォン等の通信端末であり、通信網16(例えば移動通信網やインターネット)を介して分離用データ処理装置14Aと相互に通信する。
 1個の端末装置12について図1に代表的に図示した通り、各端末装置12は、制御装置121と記憶装置122と通信装置123と表示装置124と入力装置125と収音装置126と放音装置127とを具備するコンピュータシステムで実現される。制御装置121は、記憶装置122に記憶されたプログラムを実行することで各種の制御処理および演算処理を実行する演算処理装置である。通信装置123は、通信網16を介して分離用データ処理装置14Aと通信する。なお、端末装置12と通信網16との間の通信は典型的には無線通信であるが、例えば据置型の情報処理装置を端末装置12として利用する場合には端末装置12と通信網16とが有線通信することも可能である。
 表示装置124(例えば液晶表示パネル)は、制御装置121から指示された画像を表示する。入力装置125は、端末装置12に対する利用者からの指示を受付ける機器であり、例えば利用者が操作する複数の操作子を含んで構成される。なお、表示装置124と一体に構成されたタッチパネルを入力装置125として採用することも可能である。
 記憶装置122(例えば半導体記録媒体)は、制御装置121が実行するプログラムや制御装置121が使用する各種のデータを記憶する。第1実施形態の記憶装置122は、相異なる楽曲に対応する複数の音響信号SAを記憶する。音響信号SAは、複数の音響の混合音の時間波形を示す左右2チャネルのステレオ信号である。具体的には、楽曲の旋律を歌唱した歌唱音と楽曲の伴奏を構成する複数の楽器の演奏音との混合音が音響信号SAで表現される。記憶装置122に記憶された各音響信号SAには属性データMAが対応付けられる。属性データMAは、楽曲に関連する情報である。具体的には、楽曲名や歌手名等の属性情報が属性データMAで指定される。
 第1実施形態の各端末装置12では、音響信号SAの特定成分を分離(抑圧または強調)する分離処理が実行される。具体的には、音響信号SAの歌唱音を抑圧(理想的には除去)するとともに伴奏音を強調(理想的には抽出)した音響を示すステレオ形式の音響信号SBが生成される。図2は、第1実施形態における分離処理の説明図である。定位-周波数平面における音響信号の各周波数成分の分布(散布図)が図2には図示されている。定位-周波数平面は、音響信号の各周波数成分の音像が定位する位置(以下「音像位置」という)θを示す定位軸XLと、各周波数成分の周波数を示す周波数軸XFとが設定された座標平面である。例えば定位軸XL上の原点(θ=0)が受聴者の正面方向に相当する。第1実施形態の分離処理は、音響信号SAが示す音響のうち図2の定位-周波数平面内に画定された特定の範囲(以下「分離対象範囲」という)R内の成分を抑圧または強調する処理である。図2に示すように、分離対象範囲Rは、周波数軸XF上の特定の範囲(以下「対象周波数範囲」という)RFと定位軸XL上の特定の範囲(以下「対象定位範囲」という)RLとで規定される。図2では分離対象範囲R内の成分を抑圧した場合が例示されている。
 図1の放音装置127は、制御装置121が分離処理で生成した音響信号SBに応じた音響(すなわち楽曲の歌唱音を抑圧した音響)を再生する。したがって、利用者は、楽曲の伴奏音を聴取しながら楽曲の旋律(歌唱パート)を歌唱すること(すなわちカラオケ)が可能である。収音装置126は、利用者の歌唱音の時間波形を示す収録信号Vを生成する。すなわち、収録信号Vは、楽曲の伴奏音の再生に並行して収録された収録音を示す音響信号である。なお、音響信号SBをアナログ信号に変換するD/A変換器やアナログ信号からデジタルの収録信号Vを生成するA/D変換器の図示は便宜的に省略した。
 図1の分離用データ処理装置14Aは、各端末装置12での分離処理に適用される分離用データQを管理するサーバ装置(典型的にはウェブサーバ)であり、制御装置142と記憶装置144と通信装置146とを具備するコンピュータシステムで実現される。なお、相互に別体に構成された複数の装置(例えば通信網16を介して相互に通信する複数のサーバ装置)で分離用データ処理装置14Aを実現することも可能である。制御装置142は、記憶装置144に記憶されたプログラムを実行することで各種の制御処理および演算処理を実行する。通信装置146は、通信網16を介して各端末装置12と通信する。
 記憶装置144は、制御装置142が実行するプログラムや制御装置142が使用する各種のデータを記憶する。例えば半導体記録媒体や磁気記録媒体等の記録媒体または複数種の記録媒体の組合せが記憶装置144として採用され得る。なお、分離用データ処理装置14Aとは別体の外部装置(例えば外部サーバ装置)に記憶装置144を設置し、分離用データ処理装置14Aが通信網16を介して記憶装置144に対する情報の書込や読出を実行する構成も採用され得る。
 第1実施形態の記憶装置144は、相異なる楽曲に対応する複数の分離用データQを記憶する。分離用データQは、各端末装置12による分離処理に適用される設定データであり、例えば音響信号SAのうち分離処理による分離対象を指定するために利用される。具体的には、第1実施形態の分離用データQは図2の分離対象範囲Rを指定する。すなわち、対象周波数範囲RF(周波数軸XF上の上限値および下限値)と対象定位範囲RL(定位軸XL上の上限値および下限値)とが分離用データQで指定される。以下の説明では、分離用データQが指定する分離対象範囲Rが楽曲の全区間にわたり変化しない構成(すなわち分離用データQが1種類の分離対象範囲Rのみを指定する構成)を便宜的に例示する。なお、分離用データQが対象周波数範囲RFのみを指定する構成(対象定位範囲RLの指定を省略した構成)も採用され得る。
 音響信号SAのうち楽曲の歌唱音が分離対象に包含されるように分離用データQは楽曲毎に個別に生成される。歌唱音の周波数帯域や音像位置θは楽曲毎に相違し得るから、分離用データQが指定する対象周波数範囲RFおよび対象定位範囲RLは楽曲毎に相違する。なお、以下の説明では分離用データQが1個の分離対象範囲Rを指定する場合を例示するが、分離用データQが複数の分離対象範囲Rを指定する(すなわち音響信号SAのうち相異なる複数の成分を分離する)構成も採用され得る。図1に示すように、各分離用データQには属性データMBが対応付けられる。属性データMBは、属性データMAと同様に、楽曲に関連する情報(例えば楽曲名や歌手名等の属性情報)である。なお、音響信号SAの分離対象に関連する情報(例えばボーカル,ギター等のパート名)を属性データMBが指定することも可能である。
 図3は、音響処理システム100の概略的な動作の説明図である。端末装置12の利用者は、入力装置125を適宜に操作することで、音響信号SAが記憶装置122に記憶された複数の楽曲のうち所望の楽曲(以下「対象曲」という)の再生を指示する。対象曲の再生が指示されると、端末装置12は対象曲の分離用データQを分離用データ処理装置14Aに要求し(S1)、分離用データ処理装置14Aは、端末装置12からの要求に応じて対象曲の分離用データQを要求元の端末装置12に送信する(S2)。端末装置12は、分離用データ処理装置14Aから提供された分離用データQを適用した分離処理を対象曲の音響信号SAに対して実行する(S3)。そして、端末装置12は、分離処理の結果に応じた更新用データUを生成して分離用データ処理装置14Aに送信する(S4)。更新用データUは、概略的には、分離処理後の再生音を受聴した利用者が当該再生音に関連して実行する入力動作(例えば再生音に並行した歌唱)に応じて生成される。分離用データ処理装置14Aは、端末装置12から送信された更新用データUを利用して対象曲の分離用データQを更新する(S5)。対象曲の再生が各端末装置12に指示されるたびに以上の処理が実行される。すなわち、複数の端末装置12の各々から更新用データUが分離用データ処理装置14Aに送信される。したがって、分離用データ処理装置14Aの記憶装置144に記憶された各楽曲の分離用データQは、各端末装置12における当該楽曲の再生毎に累積的に更新され、分離用データQを適用した分離処理で歌唱音が分離される精度(以下「分離用データQの信頼性」という)が経時的に向上する。以上に概説した構成および動作の詳細を以下に説明する。
 図4は、第1実施形態における各端末装置12および分離用データ処理装置14Aの機能的な構成図である。図4に示すように,端末装置12の制御装置121は、記憶装置122に記憶されたプログラム(楽曲再生プログラム)を実行することで、音響信号SAの再生および更新用データUの生成のための複数の機能(分離用データ取得部22,分離処理部24,再生制御部26,解析処理部30)を実現する。なお、制御装置121の各機能を複数の集積回路に分散した構成や、制御装置121の機能の一部を専用の電子回路(例えばDSP)が実現する構成も採用され得る。
 分離用データ取得部22は、対象曲の分離用データQを分離用データ処理装置14Aから取得する。具体的には、分離用データ取得部22は、端末装置12の利用者が選択した対象曲の属性データMAを包含する要求を通信装置123から分離用データ処理装置14Aに送信し(図3のステップS1)、端末装置12からの要求に応じて分離用データ処理装置14Aから送信された分離用データQを通信装置123から取得する(図3のステップS2)。
 図4の分離処理部24は、分離用データ取得部22が取得した分離用データQを適用した分離処理を対象曲の音響信号SAに対して実行する(図3のステップS3)。第1実施形態の分離処理部24は、抑圧処理部242と強調処理部244とを含んで構成される。抑圧処理部242は、対象曲の音響信号SAの歌唱音を抑圧するとともに伴奏音を強調した音響信号SBを生成する。他方、強調処理部244は、音響信号SAの歌唱音を強調(理想的には抽出)するとともに伴奏音を抑圧(理想的には除去)した音響信号SCを生成する。すなわち、第1実施形態の分離処理は、歌唱音を抑圧する抑圧処理と歌唱音を強調する強調処理とを包含する。分離処理の具体例を以下に詳述する。
 分離処理部24は、対象曲の音響信号SAの各周波数成分について音像位置θを算定する。音像位置θの算定には公知の技術が任意に採用され得るが、例えば特許文献1に開示されるように音響信号SAの各チャネルの強度比を利用した演算が好適である。抑圧処理部242は、音響信号SAの複数の周波数成分のうち、分離用データQが指定する対象周波数範囲RF内の周波数成分であり、かつ、分離用データQが指定する対象定位範囲RL内に音像位置θが包含される周波数成分の強度を抑圧することで音響信号SBを生成する。他方、強調処理部244は、音響信号SAの複数の周波数成分のうち、分離用データQが指定する対象周波数範囲RF外の周波数成分、または、分離用データQが指定する対象定位範囲RL外に音像位置θが包含される周波数成分の強度を抑圧することで音響信号SCを生成する。なお、音響信号SBおよび音響信号SCの一方を音響信号SAから減算することで音響信号SBおよび音響信号SCの他方を生成することも可能である。
 再生制御部26は、分離処理部24(抑圧処理部242)が生成した音響信号SBと収音装置126から供給される収録信号Vとを混合するとともに各種の音響効果(例えばエコー)の付与や信号強度の増幅等の音響処理を実行したうえで放音装置127から音波として再生させる。利用者は、分離処理後の音響信号SBが示す対象曲の伴奏音の再生に並行して対象曲の旋律(歌唱パート)を歌唱する。すなわち、楽曲の音響信号SAから歌唱音を抑圧した伴奏音(音響信号SB)と利用者の歌唱音(収録信号V)との混合音が放音装置127から再生される。
 解析処理部30は、対象曲の更新用データUを生成する。第1実施形態の更新用データUは、音響信号SBの再生音に並行して収音装置126が生成した収録信号V(すなわち楽曲の伴奏音に同期した歌唱音)に応じて生成される。図4に例示される通り、第1実施形態の解析処理部30は、評価処理部32と変換処理部34とを含んで構成される。
 評価処理部32は、収音装置126が生成する収録信号Vと分離処理部24(強調処理部244)が歌唱音を強調した音響信号SCとを比較するとともに比較結果に応じた評価データDSを生成する。すなわち、評価処理部32は、音響信号SAに包含される歌唱音と収音装置126が収録した利用者の歌唱音とを比較し、両者間の類似度(典型的には相関)に応じた評価データDSを算定する。音響信号SAに包含される歌唱音は楽曲の模範的または標準的な歌唱音に相当するから、評価処理部32は、利用者の歌唱の巧拙(模範的または標準的な歌唱との近似または相違の度合)を評価する要素とも換言され得る。評価処理部32による評価データDSの生成には公知の歌唱評価技術(例えば複数の音響信号の相互間で音高や音量等の特徴量を比較する技術)が任意に採用され得る。評価処理部32が生成した評価データDSが示す評価(得点)は表示装置124に表示される。
 変換処理部34は、収音装置126が生成した収録信号Vに応じた収録データDVを生成する。具体的には、変換処理部34は、収録信号Vを表現するMIDI(Musical Instrument Digital Interface)形式の収録データDVを生成する。すなわち、収録データDVは、収録信号Vが示す歌唱音の各音符の音高(ノートナンバ)および強度(ベロシティ)を指定して発音または消音を指示するイベントデータと、各イベントデータの処理時点(例えば相前後する各イベントデータの処理間隔)を指定する時間データとの複数組を配列した時系列データである。収録データDVの生成には、公知の変換技術(Audio-MIDI変換)が任意に採用される。
 図4に示すように、評価処理部32が生成した評価データDSと変換処理部34が生成した収録データDVとを包含する対象曲の更新用データUが端末装置12の通信装置123から通信網16を介して分離用データ処理装置14Aに送信される(図3のステップS4)。すなわち、端末装置12の利用者の歌唱音に応じた収録データDVと当該歌唱音の評価結果(巧拙)を示す評価データDSとが更新用データUとして送信される。
 図4に示すように、分離用データ処理装置14Aの制御装置142は、記憶装置144に記憶されたプログラム(分離用データ更新プログラム)を実行することで、各端末装置12に対する分離用データQの提供および更新用データUを利用した分離用データQの更新のための複数の機能(分離用データ提供部42,更新用データ取得部44、更新処理部46)を実現する。なお、制御装置142の各機能を複数の集積回路に分散した構成や、制御装置142の機能の一部を専用の電子回路(例えばDSP)が実現する構成も採用され得る。
 分離用データ提供部42は、対象曲の分離用データQを端末装置12に提供する。具体的には、分離用データ提供部42は、端末装置12から送信された要求(図3のステップS1)で指定された属性データMAに対応(典型的には合致)する対象曲の属性データMBを記憶装置144から検索し、当該属性データMBに対応する対象曲の分離用データQを通信装置146から要求元の端末装置12に送信する(図3のステップS2)。
 更新用データ取得部44は、対象曲の更新用データUを複数の端末装置12の各々から通信網16および通信装置146を介して取得する。具体的には、更新用データ取得部44は、端末装置12の解析処理部30が生成した対象曲の評価データDSと収録データDVとを包含する更新用データUを端末装置12から取得する。
 更新処理部46は、記憶装置144に記憶された対象曲の分離用データQを、更新用データ取得部44が取得した対象曲の更新用データU(収録データDV,評価データDS)を利用して更新する(図3のステップS5)。例えば、更新処理部46は、収録データDVが指定する歌唱音の音高範囲を特定し、対象曲の分離用データQで指定される対象周波数範囲RFが収録データDVの音高範囲に近付くように対象曲の分離用データQを更新する。また、例えば収録データDVで指定される音高を基本周波数とした各倍音成分が抑圧されるように分離用データQを更新する構成や、収録データDVを利用して打楽器音等の非調波音を抑圧対象から除外(回復)する構成も採用され得る。MIDIデータ等の時系列データを利用して分離精度を改善する技術は、例えば日本国特開2012-108453号公報にも開示されている。なお、分離用データQの更新に利用した収録データDVを楽曲毎に記憶装置144に蓄積することも可能である。楽曲毎に蓄積された収録データDVは、楽曲検索等の各種の処理に利用される。例えば、入力装置125の操作で利用者が指定した旋律に類似する収録データDVが蓄積された楽曲を対象曲として検索する構成や、収録データDVを利用して検索された対象曲の分離用データQを分離用データ提供部42が端末装置12に送信(S2)する構成が好適である。
 ところで、各端末装置12の利用者毎に歌唱の巧拙は相違する。歌唱が上手な利用者の収録データDVを分離用データQに反映させた場合には分離用データQの信頼性(品質)は向上するが、歌唱が下手な利用者の収録データDVを分離用データQに反映させた場合には分離用データQの信頼性が却って低下する可能性がある。以上の事情を考慮して、第1実施形態の更新処理部46は、利用者の歌唱の巧拙を示す評価データDSを収録データDVによる分離用データQの更新に適用する。具体的には、更新処理部46は、評価データDSが示す評価が高い(収録信号Vが示す収録音と音響信号SCが示す歌唱音とが類似する)ほど分離用データQの更新に対する収録データDVの影響(統計的な重み付け)が増加するように、収録データDV(更新用データU)に応じて対象曲の分離用データQを更新する。評価データDSが示す評価が閾値を下回る場合(例えば利用者の歌唱が過度に下手な場合や音響信号SBの再生時に利用者が歌唱していない場合)、更新用データUの収録データDVは分離用データQに反映されない。以上の説明から理解される通り、第1実施形態の評価データDSは、利用者の歌唱音を示す収録データDV(収録信号V)が分離用データQの信頼性の向上に寄与する度合(歌唱音の妥当性)を示す指標に相当し、分離用データQの更新時における収録データDVの重み値として利用される。
 図3を参照して説明した通り、各端末装置12の利用者が対象曲の再生を指示するたびに以上の処理が実行される。すなわち、分離用データ処理装置14Aは、既存の分離用データQを適用した分離処理後の音響信号SBの再生音を受聴した利用者の入力動作(対象曲の歌唱)に応じた更新用データUを、複数の端末装置12の各々から複数回にわたり反復的に取得し、各端末装置12から取得した更新用データUを利用して記憶装置144内の各楽曲の分離用データQを更新する。端末装置12から更新用データUを取得するたびに当該更新用データUを適用した分離用データQの更新が逐次的に実行される。
 以上の説明から理解される通り、第1実施形態では、複数の端末装置12の各々から取得した更新用データUを利用して記憶装置144内の各楽曲の分離用データQが更新される。すなわち、各端末装置12の複数の利用者の歌唱が分離用データQに反映される。したがって、音響信号SAの特定成分(歌唱音)を高精度に分離可能な信頼性の高い分離用データQを簡便に生成できるという利点がある。
 第1実施形態では、評価データDSが示す評価が高い(収録信号Vが示す収録音と音響信号SCが示す歌唱音とが類似する)ほど分離用データQの更新に対する収録データDVの影響が増加する。したがって、各端末装置12の利用者の歌唱の巧拙に関わらず収録データDVを分離用データQに反映させる構成と比較して、各楽曲の分離用データQの信頼性を効率的に改善できるという利点もある。
<第2実施形態>
 本発明の第2実施形態を説明する。以下に例示する各形態において作用や機能が第1実施形態と同等である要素については、第1実施形態の説明で参照した符号を流用して各々の詳細な説明を適宜に省略する。
 図5は、第2実施形態における各端末装置12および分離用データ処理装置14Aの機能的な構成図である。図5に示すように、第2実施形態の端末装置12の制御装置121は、第1実施形態と同様の要素(分離用データ取得部22,分離処理部24,再生制御部26,解析処理部30)に加えて表示制御部28として機能する。表示制御部28は、分離用データQを適用した分離処理を表象する分離処理画像を表示装置124に表示させる。分離処理画像は、図2の例示と同様に、定位-周波数平面における音響信号SB(または音響信号SA)の各周波数成分の分布(散布図)と分離対象範囲Rとを利用者に提示する画像である。
 利用者は、表示装置124に表示された分離処理画像を確認しながら入力装置125を適宜に操作することで分離処理画像の分離対象範囲Rを調整することが可能である。分離対象範囲Rを調整することで、分離処理後の音響信号SBの再生音における歌唱音の音量が増減する。利用者は、放音装置127から放射される音響信号SBの再生音を聴取しながら入力装置125を操作(調整操作)することで、対象曲の歌唱音が分離処理で高精度に低減される(すなわち再生音における歌唱音の音量が充分に低減される)ように、分離対象範囲Rの対象周波数範囲RFおよび対象定位範囲RLを調整する。
 対象曲の歌唱音が高精度に抑圧されるように分離対象範囲Rが設定されるほど、強調処理後の音響信号SCでは歌唱音以外の成分の混在が抑制されるから、評価処理部32による収録信号Vの評価が上昇し易い(分離対象範囲Rが不適切であるほど歌唱音の評価が低下する)という傾向がある。すなわち、強調処理後の音響信号SCと利用者の歌唱音の収録信号Vとの類否評価(利用者に対する歌唱評価結果の提示)は、対象曲の歌唱音が高精度に抑圧されるように分離対象範囲Rを調整する調整操作を利用者に適正に実行させるための誘因として作用する。
 図5に示す通り、第2実施形態の解析処理部30は、第1実施形態と同様の要素(評価処理部32,変換処理部34)に加えて調整管理部36を含んで構成される。調整管理部36は、分離対象範囲Rに対する利用者の調整操作に応じた調整データDCを生成する。具体的には、調整管理部36は、利用者の調整操作後の分離対象範囲R(対象周波数範囲RF,対象定位範囲RL)を指定する調整データDCを生成する。なお、調整管理部36による調整データDCの生成時期は任意である。例えば、対象曲の再生が完了した時点で調整データDCを生成する構成や、対象曲の途中の時点(例えば間奏の時点)で調整データDCを生成する構成が採用され得る。
 第2実施形態では、評価処理部32が生成した評価データDSと変換処理部34が生成した収録データDVとに加えて調整管理部36が生成した調整データDCを包含する対象曲の更新用データUが端末装置12から分離用データ処理装置14Aに送信される(図3のステップS4)。分離用データ処理装置14Aの更新用データ取得部44は、複数の端末装置12の各々から更新用データUを取得し、更新処理部46は、更新用データ取得部44が取得した各更新用データUを利用して対象曲の分離用データQを更新する。
 第2実施形態の更新処理部46は、更新用データUの収録データDVおよび評価データDSに応じて第1実施形態と同様に対象曲の分離用データQを更新するほか、端末装置12の調整管理部36が生成した調整データDCを利用して対象曲の分離用データQを更新する。具体的には、更新処理部46は、調整データDCが指定する分離対象範囲Rに近付くように、対象曲の分離用データQが指定する分離対象範囲R(対象周波数範囲RF,対象定位範囲RL)を更新する。複数の端末装置12から送信される調整データDCを利用した分離用データQの更新が反復されることで、各楽曲の分離用データQが指定する分離対象範囲Rは、各端末装置12の利用者が当該楽曲について過去に指定した複数の分離対象範囲Rの平均的な範囲に調整される。
 第2実施形態においても第1実施形態と同様の効果が実現される。また、第2実施形態では、分離用データQを適用した分離処理後の再生音の受聴者による調整操作に応じた調整データDCが当該分離用データQの更新に適用されるから、音響信号SAの特定成分(歌唱音)を高精度に分離可能な信頼性の高い分離用データQを簡便に生成できるという前述の効果は格別に顕著である。なお、以上の例示では、評価データDSと収録データDVと調整データDCとを包含する更新用データUを例示したが、更新用データUが調整データDCのみを包含する構成(評価データDSや収録データDVを省略した構成)も採用され得る。すなわち、評価処理部32または変換処理部34を省略することも可能である。
<第3実施形態>
 第1実施形態および第2実施形態では、分離用データ処理装置14Aが各端末装置12から通信網16を介して更新用データUを受信して分離用データQを更新する音響処理システム100を例示した。第3実施形態の分離用データ処理装置14Bは、更新用データUの生成と分離用データQの更新とを装置単体で実行する。
 図6は、第3実施形態の分離用データ処理装置14Bのブロック図である。図6に示すように、分離用データ処理装置14Bは、制御装置181と記憶装置182と表示装置184と入力装置185と収音装置186と放音装置187とを具備するコンピュータシステムで実現される。例えば携帯電話機やスマートフォンやパーソナルコンピュータ等の情報処理装置が分離用データ処理装置14Bとして利用される。なお、通信網16を介した分離用データQまたは更新用データUの授受は第3実施形態では不要であるから、分離用データ処理装置14Bの通信機能の有無は不問である。
 記憶装置182は、音響信号SAと分離用データQとを複数の楽曲の各々について記憶する。表示装置184は、第2実施形態の表示装置124と同様に分離処理画像を表示し、入力装置185は、第1実施形態の入力装置125と同様に対象曲の選択等の指示を利用者から受付ける。収音装置186は、第1実施形態の収音装置126と同様に収録信号Vを生成し、放音装置187は、第1実施形態の放音装置127と同様に分離処理後の音響信号SBを再生する。
 制御装置181は、記憶装置182に記憶されたプログラム(楽曲再生プログラム,分離用データ更新プログラム)を実行することで、音響信号SAの再生および更新用データUの生成と分離用データQの更新とのための複数の機能(分離処理部54,再生制御部56,表示制御部58,更新用データ取得部60,更新処理部70)を実現する。なお、制御装置181の各機能を複数の集積回路に分散した構成や、制御装置181の機能の一部を専用の電子回路(例えばDSP)が実現する構成も採用され得る。
 分離処理部54(抑圧処理部542,強調処理部544)は、第1実施形態の分離処理部24と同様に、利用者が指定した対象曲の分離用データQを適用した分離処理(抑圧処理,強調処理)を対象曲の音響信号SAに対して実行することで、音響信号SAの歌唱音を抑圧(理想的には除去)した音響信号SBと歌唱音を強調(理想的には抽出)した音響信号SCとを生成する。再生制御部56は、第1実施形態の再生制御部26と同様に、分離処理部54(抑圧処理部542)が生成した音響信号SBと収音装置186から供給される収録信号Vとを混合して放音装置187に再生させる。第1実施形態と同様に、利用者は、放音装置187が再生する対象曲の伴奏音を聴取しながら楽曲の旋律(歌唱パート)を歌唱する。
 更新用データ取得部60は、対象曲の更新用データUを取得する。具体的には、更新用データ取得部60は、第2実施形態の解析処理部30と同様に、評価処理部62と変換処理部64と調整管理部66とを含んで構成され、評価データDSと収録データDVと調整データDCと包含する更新用データUを生成する。評価処理部62は、第1実施形態の評価処理部32と同様に音響信号SCと収録信号Vとの比較結果に応じた評価データDSを生成し、変換処理部64は、第1実施形態の変換処理部34と同様に収録信号Vに応じた収録データDVを生成する。また、調整管理部66は、第2実施形態の調整管理部36と同様に、利用者が指示した分離対象範囲Rを指定する調整データDCを生成する。
 更新処理部70は、第1実施形態の更新処理部46と同様に、更新用データ取得部60が取得(生成)した更新用データUを利用して対象曲の分離用データQを更新する。分離用データQの更新方法は第1実施形態と同様である。利用者が楽曲の再生を指示するたびに更新用データUの生成と分離用データQの更新とが実行される。すなわち、第1実施形態と同様に、各楽曲の分離用データQは複数回にわたり反復的に更新される。したがって、第3実施形態においても第1実施形態や第2実施形態と同様の効果が実現される。
 以上の説明から理解される通り、第1実施形態の更新用データ取得部44および第3実施形態の更新用データ取得部60は、分離用データQを適用した分離処理後の音響信号SBの再生音を受聴した利用者による入力(収音装置に対する歌唱や入力装置に対する指示)が反映された更新用データUを取得する要素として包括される。すなわち、更新用データUを取得する要素(更新用データ取得部)が、第1実施形態のように外部装置(端末装置12)から更新用データUを取得するか第3実施形態のように自身が更新用データUを生成するかは不問である。
<変形例>
 前述の各形態は多様に変形され得る。具体的な変形の態様を以下に例示する。以下の例示から任意に選択された2以上の態様は適宜に併合され得る。
(1)前述の各形態では、複数の楽曲の各々について分離用データQを用意したが、楽曲毎に個別に分離用データQを用意する必要はない。例えば、楽曲の属性(例えば歌唱者,ジャンル,録音形式)毎に分離用データQを用意した構成も好適である。属性が相互に共通する複数の楽曲の分離処理には共通の分離用データQが適用される。以上の構成によれば、複数の楽曲について共通の分離用データQが適用されるから、楽曲毎に分離用データQを保持する構成と比較して、分離用データQの記憶に必要な容量が削減されるという利点がある。
(2)前述の各形態では、分離対象範囲Rが楽曲の全区間にわたり変化しない構成を例示したが、分離対象範囲Rを楽曲内で経時的に変化させる構成も採用され得る。すなわち、分離用データQは分離対象範囲Rの時系列(時間的な遷移)を指定する。以上の構成では、更新用データUが楽曲内で反復的に生成され、分離用データQが指定する各分離対象範囲Rの更新に適用される。
(3)分離処理の具体的な内容は以上の各形態の例示に限定されず、音響信号SAから特定成分を分離(抑圧または強調)する公知の音響処理技術(例えば音源分離技術)が任意に採用される。分離用データQまたは更新用データUの形式は分離処理の内容に応じて適宜に変更される。例えば楽曲の各音符の音高を時系列に指定するMIDI形式の時系列データ(例えば収録データDV)を分離用データQとして利用し、分離用データQが指定する各音高の音響成分を特定成分として音響信号SAから分離する構成も採用され得る。以上の構成では、収録信号Vに応じた収録データDV(または各端末装置12から送信された複数の収録データDVの平均データ)が分離用データQとして記憶される。
(4)所定の条件が成立した場合に分離用データQの更新を終了することも可能である。例えば、所定の回数にわたり分離用データQが更新された場合に分離用データQの更新を終了する(分離用データQを確定する)構成や、更新用データUを適用した更新時の分離用データQの変動が充分に小さい場合(分離用データQが収束した場合)に分離用データQの更新を終了する構成が採用される。
(5)第1実施形態および第2実施形態では、各楽曲の音響信号SAを端末装置12の記憶装置122に記憶したが、各楽曲の音響信号SAを分離用データ処理装置14Aの記憶装置144に格納することも可能である。対象曲の再生が端末装置12に指示されると、分離用データ処理装置14A(分離用データ提供部42)は、対象曲の分離用データQと音響信号SAとを記憶装置144から取得して要求元の端末装置12に送信する。以上の構成によれば、音響信号SAを端末装置12毎に保持する必要がないという利点がある。
 分離用データ処理装置14Aの記憶装置144に音響信号SAを格納した構成では、記憶装置144に記憶された対象曲の音響信号SAに分離用データQを適用する分離処理部24を分離用データ処理装置14Aに設置することも可能である。以上の構成では、分離用データ処理装置14Aの分離処理部24による分離処理後の音響信号(SB,SC)が通信網16を介して端末装置12に送信される。音響信号SBは再生制御部26により放音装置127から再生され、音響信号SCは評価処理部32による収録信号Vの評価に適用される。以上の構成によれば、分離処理部24を各端末装置12に搭載する必要がないという利点がある。なお、分離用データ処理装置14Aに設置された分離処理部24が、端末装置12から通信網16を介して受信した音響信号SAを処理することも可能である。
 また、分離処理部24に包含される抑圧処理部242および強調処理部244の一方のみを分離用データ処理装置14Aに設置することも可能である。例えば、強調処理部244と評価処理部32とを分離用データ処理装置14Aに設置するとともに抑圧処理部242を端末装置12に設置した構成が採用される。以上の構成では、歌唱音の収録信号Vが端末装置12から分離用データ処理装置14Aに送信され、分離用データ処理装置14Aの強調処理部244が生成した音響信号SCと端末装置12から受信した収録信号Vとを対比することで分離用データ処理装置14Aの歌唱評価部32が評価データDSを生成する。なお、強調処理部244を分離用データ処理装置14Aに設置した構成のもとで、記憶装置144に音響信号SAが格納される場合には、記憶装置144内の音響信号SAが、分離用データ処理装置14Aの強調処理部244による音響信号SCの生成に適用され、かつ、分離用データ処理装置14Aから端末装置12に送信されたうえで端末装置12の抑圧処理部242による音響信号SBの生成に適用される。他方、強調処理部244を分離用データ処理装置14Aに設置した構成のもとで、端末装置12の記憶装置122に音響信号SAが格納される場合には、記憶装置122内の音響信号SAが、端末装置12の抑圧処理部242による音響信号SBの生成に適用され、かつ、端末装置12から分離用データ処理装置14Aに送信されたうえで分離用データ処理装置14Aの強調処理部244による音響信号SCの生成に適用される。なお、以上の説明では強調処理部244を分離用データ処理装置14Aに設置した構成を例示したが、抑圧処理部242を分離用データ処理装置14Aに設置することも可能である。
(6)図7に例示される通り、第2実施形態において、入力装置125に対する利用者の調整操作を、分離処理部24が分離処理に適用する分離用データQに随時に反映させることも可能である。図7の構成では、分離処理(抑圧処理)で音響信号SAから抑圧される特定成分が、利用者による調整操作に応じて実時間的に変更される。したがって、利用者は、音響信号SBの再生音を実際に聴取して調整操作の効果(所望の特定成分が適切に抑圧されているか否か)を確認しながら調整操作を実行することが可能である。なお、音響信号SBの再生音を聴取した利用者が分離対象範囲Rの確定のための所定の操作を入力装置125に付与した場合に分離用データ処理装置14Aに対する更新用データUの送信が実行される構成(分離対象範囲Rの確定まで更新用データUを送信しない構成)も好適である。
(7)図8は、第1実施形態の変形例に係る解析処理部30のブロック図である。図8の変換処理部38は、分離処理部24(強調処理部244)が生成した音響信号SCに応じた変換データDAを生成する。変換データDAは、例えば音響信号SCを表現するMIDI形式の時系列データである。図8の選択部82は、分離処理部24が生成した音響信号SCと変換処理部38が生成した変換データDAとの何れかを参照データDREFとして選択する。
 他方、図8の変換処理部34は、収音装置126が生成した収録信号Vに応じた変換データDBを生成する。変換データDBは、第1実施形態や第2実施形態の収録データDVと同様に、例えば収録信号Vを表現するMIDI形式の時系列データである。選択部84は、収音装置126が生成した収録信号Vと変換処理部34が生成した変換データDBとの何れかを収録データDVとして選択する。評価処理部32は、選択部82が選択した参照データDREFと選択部84が選択した収録データDVとを比較することで評価データDSを生成する。
 選択部82および選択部84の動作(選択対象)は、例えば入力装置125に対する利用者からの指示に応じて制御される。利用者は、第1評価処理および第2評価処理を選択することが可能である。第1評価処理が選択された場合、選択部82は音響信号SCを参照データDREFとして選択するとともに選択部84は収録信号Vを収録データDVとして選択する。したがって、評価処理部32は、音響信号SCと収録信号Vとの比較結果に応じた評価データDSを生成する(第1評価処理)。他方、第2評価処理が選択された場合、選択部82は変換データDAを参照データDREFとして選択するとともに選択部84は変換データDBを収録データDVとして選択する。したがって、評価処理部32は、変換データDAと変換データDBとの比較結果に応じた評価データDSを生成する(第2評価処理)。第1評価処理および第2評価処理の何れにおいても、音響信号SC(音響信号SCまたは変換データDA)と収録信号V(収録信号Vまたは変換データDB)との比較結果に応じた評価データDSが生成される。
 なお、図8においては第1実施形態の構成を基礎とした変形例を説明したが、第2実施形態や第3実施形態においても同様の構成(すなわち第1評価処理と第2評価処理とが選択的に実行される構成)が採用され得る。
(8)第1実施形態および第2実施形態では、端末装置12の解析処理部30が生成した更新用データUを分離用データ処理装置14Aの更新用データ取得部44が受信する構成を例示したが、分離用データ処理装置14Aの更新用データ取得部44が解析処理部30と同様の構成および動作により更新用データUを生成することも可能である。具体的には、更新用データ取得部44は、端末装置12の収音装置126が生成した収録信号Vと歌唱音が強調された音響信号SCとを端末装置12から受信し、収録信号Vと音響信号SCとの比較結果に応じた評価データDSを生成する(評価処理部32)とともに収録信号Vに応じた収録データDVを生成する(変換処理部34)。また、更新用データ取得部44は、端末装置12の入力装置125に対する利用者からの指示(調整操作)を端末装置12から取得し、調整操作に応じた調整データDCを生成する(調整管理部36)。更新処理部46は、更新用データ取得部44が生成した更新用データU(評価データDS,収録データDV,調整データDC)に応じて対象曲の分離用データQを更新する。
 以上の説明から理解される通り、第1実施形態および第2実施形態の更新用データ取得部44は、分離用データQを適用した分離処理後の音響信号SBの再生音を受聴した利用者による入力(収音装置126に対する歌唱や入力装置125に対する指示)が反映された更新用データUを取得する要素として包括され、第1実施形態のように外部装置(端末装置12)から更新用データUを取得するか変形例の例示のように自身が更新用データUを生成するかは不問である。
(9)前述の各形態では、音響信号SAの歌唱音の分離を例示したが、音響信号SAから分離される成分は歌唱音に限定されない。例えば音響信号SAのうち特定の楽器の演奏音を分離する構成は、当該楽器の演奏の練習に好適に利用される。また、音響信号SAの複数の成分(例えば複数の楽器の演奏音)を分離対象とすることも可能である。なお、以上の説明では楽曲の音響信号SAに対する分離処理を例示したが、本発明において音楽的な観点は必須ではない。例えば、特定の言語の会話音等が収録された語学学習用の音響信号SAから特定の成分(例えば複数の話者の会話音のうち特定話者の発声音)を分離することも可能である。
 本出願は、2013年3月15日出願の日本特許出願(特願2013-054217)に基づくものであり、その内容はここに参照として取り込まれる。
本発明によれば、音響信号の特定成分を高精度に分離(抑圧または強調)可能な分離用データを簡便に生成することが可能である。
100……音響処理システム、12……端末装置、14……分離用データ処理装置、121,142,181……制御装置、122,144,182……記憶装置、123,146……通信装置、124,184……表示装置、125,185……入力装置、126,186……収音装置、127,187……放音装置、16……通信網、22……分離用データ取得部、24,54……分離処理部、26,56……再生制御部、30……解析処理部、32,62……評価処理部、34,64……変換処理部、36,66……調整管理部。42……分離用データ提供部、44,60……更新用データ取得部、46,70……更新処理部。 

Claims (7)

  1.  音響信号の特定成分を強調または抑圧する分離処理に適用される分離用データを記憶する記憶部と、
     前記分離用データを適用した分離処理で前記特定成分が抑圧された音響信号の再生に並行して収録された収録音に応じた収録データを含む更新用データを取得する更新用データ取得部と、
     前記更新用データ取得部が取得した更新用データを利用して前記記憶部の分離用データを更新する更新処理部と
     を具備する分離用データ処理装置。
  2.  前記更新用データ取得部は、前記分離用データを適用した分離処理で特定成分が抑圧された音響信号の再生に並行して収録された収録信号と、前記分離用データを適用した分離処理で前記特定成分を強調した音響信号との比較結果に応じた評価データと、前記収録信号に応じた前記収録データとを含む前記更新用データを取得する
     請求項1の分離用データ処理装置。
  3.  前記更新処理部は、前記収録信号と前記特定成分を強調した音響信号とが類似するほど前記分離用データの更新に対する前記収録データの影響が増加するように、前記更新用データに応じて前記分離用データを更新する
     請求項2の分離用データ処理装置。
  4.  前記更新用データ取得部は、前記分離用データを適用した分離処理後の音響信号の再生音の受聴者による調整指示に応じた調整データを含む
     請求項1から請求項3の何れかの分離用データ処理装置。
  5.  前記更新用データ取得部は、端末装置から通信網を介して前記更新用データを取得する
     請求項1から請求項4の何れかの分離用データ処理装置。
  6.  音響信号の特定成分を強調または抑圧する分離処理に適用される分離用データを記憶する記憶部を具備するコンピュータを、
     前記分離用データを適用した分離処理で前記特定成分が抑圧された音響信号の再生に並行して収録された収録音に応じた収録データを含む更新用データを取得する更新用データ取得部、および、
     前記更新用データ取得部が取得した更新用データを利用して前記記憶部の分離用データを更新する更新処理部
     として機能させるプログラム。
  7.  音響信号の特定成分を強調または抑圧する分離処理に適用される分離用データを記憶し、
     前記分離用データを適用した分離処理で前記特定成分が抑圧された音響信号の再生に並行して収録された収録音に応じた収録データを含む更新用データを取得し、
     前記取得した更新用データを利用して前記した分離用データを更新する
     分離用データ処理方法。
PCT/JP2014/056572 2013-03-15 2014-03-12 分離用データ処理装置およびプログラム Ceased WO2014142201A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
CN201480014346.5A CN105122360A (zh) 2013-03-15 2014-03-12 分离用数据处理装置以及程序
KR1020157024425A KR20150119013A (ko) 2013-03-15 2014-03-12 분리용 데이터 처리 장치 및 프로그램

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2013054217A JP2014178641A (ja) 2013-03-15 2013-03-15 分離用データ処理装置およびプログラム
JP2013-054217 2013-03-15

Publications (1)

Publication Number Publication Date
WO2014142201A1 true WO2014142201A1 (ja) 2014-09-18

Family

ID=51536852

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2014/056572 Ceased WO2014142201A1 (ja) 2013-03-15 2014-03-12 分離用データ処理装置およびプログラム

Country Status (4)

Country Link
JP (1) JP2014178641A (ja)
KR (1) KR20150119013A (ja)
CN (1) CN105122360A (ja)
WO (1) WO2014142201A1 (ja)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2020120754A1 (en) * 2018-12-14 2020-06-18 Sony Corporation Audio processing device, audio processing method and computer program thereof
TWI902280B (zh) * 2024-05-31 2025-10-21 瑞昱半導體股份有限公司 卡拉ok裝置及其歌聲評分系統

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP7443823B2 (ja) * 2020-02-28 2024-03-06 ヤマハ株式会社 音響処理方法

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2007102103A (ja) * 2005-10-07 2007-04-19 Yamaha Corp オーディオデータ再生装置および携帯端末装置
JP2011209593A (ja) * 2010-03-30 2011-10-20 Brother Industries Ltd 歌声分離装置、及びプログラム
JP2012109924A (ja) * 2010-10-28 2012-06-07 Yamaha Corp 音響処理装置

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102264014A (zh) * 2010-05-26 2011-11-30 陈柚仁 音源分离式无线耳机模块及音源分离方法

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2007102103A (ja) * 2005-10-07 2007-04-19 Yamaha Corp オーディオデータ再生装置および携帯端末装置
JP2011209593A (ja) * 2010-03-30 2011-10-20 Brother Industries Ltd 歌声分離装置、及びプログラム
JP2012109924A (ja) * 2010-10-28 2012-06-07 Yamaha Corp 音響処理装置

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2020120754A1 (en) * 2018-12-14 2020-06-18 Sony Corporation Audio processing device, audio processing method and computer program thereof
TWI902280B (zh) * 2024-05-31 2025-10-21 瑞昱半導體股份有限公司 卡拉ok裝置及其歌聲評分系統

Also Published As

Publication number Publication date
KR20150119013A (ko) 2015-10-23
CN105122360A (zh) 2015-12-02
JP2014178641A (ja) 2014-09-25

Similar Documents

Publication Publication Date Title
JP6201460B2 (ja) ミキシング管理装置
JP5333517B2 (ja) データ処理装置およびプログラム
JP7739348B2 (ja) 音高に依存しない音色属性をメディア信号から抽出する方法、コンピュータ可読記憶媒体及び装置
JP2016070999A (ja) カラオケ効果音設定システム
US8327269B2 (en) Positioning a virtual sound capturing device in a three dimensional interface
JPWO2018155480A1 (ja) 情報処理方法および情報処理装置
JP6102076B2 (ja) 評価装置
JP6288197B2 (ja) 評価装置及びプログラム
d'Escriván Music technology
WO2014142201A1 (ja) 分離用データ処理装置およびプログラム
CN112420006B (zh) 运行模拟乐器组件的方法及装置、存储介质、计算机设备
JP2016102982A (ja) カラオケシステム、プログラム、カラオケ音声再生方法及び音声入力処理装置
KR20150118974A (ko) 음성 처리 장치
JP5678935B2 (ja) 楽器演奏評価装置、楽器演奏評価システム
JP6657866B2 (ja) 音響効果付与装置及び音響効果付与プログラム
JP6316099B2 (ja) カラオケ装置
WO2017135350A1 (ja) 記録媒体、音響処理装置および音響処理方法
JP5510207B2 (ja) 楽音編集装置及びプログラム
EP4597488A1 (en) Audio processing method and apparatus, and electronic device
JP5742472B2 (ja) データ検索装置およびプログラム
JP2019101071A (ja) 歌唱評価装置、歌唱評価プログラム及びカラオケ装置
JP2019045755A (ja) 歌唱評価装置、歌唱評価プログラム、歌唱評価方法及びカラオケ装置
WO2023062865A1 (ja) 情報処理装置および方法、並びにプログラム
CN120636375A (zh) 音频音效处理方法、装置、计算机设备及存储介质
CN119400136A (zh) 一种实时伴奏调音方法、设备、存储介质及程序产品

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 14765431

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 20157024425

Country of ref document: KR

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 14765431

Country of ref document: EP

Kind code of ref document: A1