WO2020174776A1 - 音声信号処理装置 - Google Patents
音声信号処理装置 Download PDFInfo
- Publication number
- WO2020174776A1 WO2020174776A1 PCT/JP2019/045255 JP2019045255W WO2020174776A1 WO 2020174776 A1 WO2020174776 A1 WO 2020174776A1 JP 2019045255 W JP2019045255 W JP 2019045255W WO 2020174776 A1 WO2020174776 A1 WO 2020174776A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- audio signal
- posture
- amount
- variation
- processing device
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/46—Volume control
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
- G06F3/165—Management of the audio stream, e.g. setting of volume, audio stream path
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/02—Means for controlling the tone frequencies, e.g. attack or decay; Means for producing special musical effects, e.g. vibratos or glissandos
- G10H1/06—Circuits for establishing the harmonic content of tones, or other arrangements for changing the tone colour
- G10H1/12—Circuits for establishing the harmonic content of tones, or other arrangements for changing the tone colour by filtering complex waveforms
- G10H1/125—Circuits for establishing the harmonic content of tones, or other arrangements for changing the tone colour by filtering complex waveforms using a digital filter
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/02—Means for controlling the tone frequencies, e.g. attack or decay; Means for producing special musical effects, e.g. vibratos or glissandos
- G10H1/06—Circuits for establishing the harmonic content of tones, or other arrangements for changing the tone colour
- G10H1/14—Circuits for establishing the harmonic content of tones, or other arrangements for changing the tone colour during execution
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/302—Electronic adaptation of stereophonic sound system to listener position or orientation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S7/00—Indicating arrangements; Control arrangements, e.g. balance control
- H04S7/30—Control circuits for electronic adaptation of the sound field
- H04S7/307—Frequency adjustment, e.g. tone control
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2220/00—Input/output interfacing specifically adapted for electrophonic musical tools or instruments
- G10H2220/155—User input interfaces for electrophonic musical instruments
- G10H2220/201—User input interfaces for electrophonic musical instruments for movement interpretation, i.e. capturing and recognizing a gesture or a specific kind of movement, e.g. to control a musical instrument
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2220/00—Input/output interfacing specifically adapted for electrophonic musical tools or instruments
- G10H2220/155—User input interfaces for electrophonic musical instruments
- G10H2220/391—Angle sensing for musical purposes, using data from a gyroscope, gyrometer or other angular velocity or angular movement sensing device
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2220/00—Input/output interfacing specifically adapted for electrophonic musical tools or instruments
- G10H2220/155—User input interfaces for electrophonic musical instruments
- G10H2220/395—Acceleration sensing or accelerometer use, e.g. 3D movement computation by integration of accelerometer data, angle sensing with respect to the vertical, i.e. gravity sensing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04S—STEREOPHONIC SYSTEMS
- H04S2400/00—Details of stereophonic systems covered by H04S but not provided for in its groups
- H04S2400/13—Aspects of volume control, not necessarily automatic, in stereophonic sound systems
Definitions
- the present technology relates to an audio signal processing device. More specifically, the present invention relates to an audio signal processing device that adjusts an audio signal based on a posture variation amount.
- Patent Document 1 Japanese Unexamined Patent Publication No. 2 0 1 7-1 9 9 9 9 5
- Patent Document 2 JP 2 0 1 2 _ 2 3 0 1 3 5
- the present technology is created in view of such a situation, and an object thereof is to adjust a voice signal based on an amount of posture variation according to a situation of surrounding voice. ⁇ 2020/174776 2 ⁇ (: 170?2019/045255
- the present technology has been made to solve the above-mentioned problems, and the first aspect thereof is to acquire an audio signal and set a target value for audio adjustment based on the audio signal.
- a voice signal analysis unit for setting a posture variation analysis unit that obtains posture information and generates a posture variation amount based on the posture information, and a voice signal for the target value according to the posture variation amount.
- An audio signal processing device comprising an audio signal adjusting unit for adjusting. This has the effect of adjusting the voice signal toward the target value according to the amount of posture variation.
- the audio signal adjusting unit adjusts the audio signal so as to reduce the volume of the audio signal when the amount of posture variation becomes larger than a first threshold value, and the target value is adjusted.
- the sound signal may be adjusted so that the volume of the sound signal is returned when the amount of posture variation becomes smaller than the second threshold after the above-mentioned condition. This brings about the effect of controlling the increase and decrease of the sound volume according to the posture variation amount.
- the audio signal adjusting unit adjusts the audio signal so as to narrow a bandwidth of a frequency of the audio signal when the posture variation becomes larger than a first threshold value.
- the voice signal may be adjusted so as to restore the bandwidth of the frequency of the voice signal when the posture variation amount becomes smaller than the second threshold value after reaching the target value. This has the effect of controlling the bandwidth of the frequency according to the amount of posture variation.
- the audio signal adjusting unit adjusts the audio signal so as to reduce a gain of a frequency of the audio signal when the posture variation amount becomes larger than a first threshold value.
- the voice signal may be adjusted so as to restore the gain of the frequency of the voice signal when the posture variation amount becomes smaller than the second threshold value after reaching the target value. This brings about the effect of controlling the frequency gain according to the amount of posture variation.
- the audio signal adjustment unit is configured such that the posture variation amount is larger than a first threshold value, and the posture variation amount is smaller than a second threshold value.
- the audio signal may be adjusted when one of the two states, which is in the open state, continues for a predetermined period. This has the effect of suppressing unnatural behavior by providing hysteretic characteristics in the adjustment of the audio signal.
- the audio signal adjusting unit may adjust the audio signal stepwise with a certain amount of step size. This brings about the effect of suppressing the discomfort in the reproduced sound.
- the first aspect may further include a sensor that detects acceleration or angular velocity and generates the posture information. This brings about the effect of adjusting the audio signal based on the posture information detected in the device.
- a recording/reproducing unit that synchronously records and reproduces the audio signal and the posture information is further provided, and the audio signal analysis unit is based on the reproduced audio signal.
- the attitude variation analysis unit generates an attitude variation amount based on the reproduced posture information
- the audio signal adjustment unit generates the attitude variation amount according to the reproduced posture variation amount.
- the reproduced audio signal may be adjusted. This brings about the effect of adjusting the audio signal at the time of reproduction, based on the audio signal and the posture information that have been recorded once.
- an image signal correction unit that acquires an image signal synchronized with the audio signal and corrects the blur of the image signal according to the posture variation amount may be further provided. .. This brings about the effect of correcting the blurring of the image signal as well as adjusting the audio signal.
- the audio signal adjusting unit adjusts the audio signal when there is a correlation between the attitude variation indicated by the attitude variation and the audio signal variation. You can This brings about an effect of suppressing the adjustment of the audio signal when the audio does not change even if the posture changes.
- FIG. 1 is a diagram showing a configuration example of an audio signal processing device 100 according to a first embodiment of the present technology. 20/174776 4 ⁇ (: 170?2019/045255
- FIG. 2 is a diagram showing a waveform example of each unit of the audio signal processing device 100 according to the first embodiment of the present technology.
- FIG. 3 is a diagram showing an example of a state transition of each section in the first embodiment of the present technology.
- FIG. 4 is a diagram showing an example of a processing procedure of control in a section 1 of the first embodiment of the present technology.
- FIG. 5 is a diagram showing an example of a processing procedure of control in a section 2 of the first embodiment of the present technology.
- FIG. 6 is a diagram showing an example of a processing procedure of control in a section 3 of the first embodiment of the present technology.
- FIG. 7 is a diagram showing an example of a control processing procedure in a section 4 of the first embodiment of the present technology.
- FIG. 8 is a diagram showing an example of a processing procedure of control in a section 5 of the first embodiment of the present technology.
- FIG. 9 is a diagram showing an example of a processing procedure of control in a section 6 of the first embodiment of the present technology.
- FIG. 10 is a diagram showing an example of a processing procedure of posture variation analysis processing in the first embodiment of the present technology.
- FIG. 11 is a diagram showing an example of a processing procedure of an audio signal analysis processing according to the first embodiment of the present technology.
- FIG. 12 is a diagram showing an example of a processing procedure of an audio signal adjustment processing according to the first embodiment of the present technology.
- Fig. 13 is a diagram showing an example of adjusting the frequency spectrum of an audio signal in the second embodiment of the present technology.
- FIG. 14 A diagram showing an example of a processing procedure of an audio signal analysis processing according to the second embodiment of the present technology.
- FIG. 15 A diagram showing an example of a processing procedure of an audio signal adjustment processing in the second embodiment of the present technology. ⁇ 2020/174776 5 ⁇ (: 170?2019/045255
- FIG. 16 is a diagram showing a configuration example of an audio signal processing device 100 according to a third embodiment of the present technology.
- FIG. 17 is a diagram showing a configuration example of an audio signal processing device 100 according to a fourth embodiment of the present technology.
- FIG. 18 is a diagram showing a configuration example of an audio signal processing device 100 according to a fifth embodiment of the present technology.
- FIG. 19 A diagram showing an example of a processing procedure of audio signal analysis processing according to a fifth embodiment of the present technology.
- FIG. 20 A diagram showing an example of a processing procedure of control in a section 1 of a fifth embodiment of the present technology.
- FIG. 21 A diagram showing an example of a processing procedure of control in a section 2 of the fifth embodiment of the present technology.
- FIG. 22 A diagram showing an example of a processing procedure of control in a section 3 of the fifth embodiment of the present technology.
- FIG. 23 is a diagram showing an example of a processing procedure of control in section 4 of the fifth embodiment of the present technology.
- FIG. 24 is a diagram showing an example of a processing procedure of control in section 5 of the fifth embodiment of the present technology.
- FIG. 25 is a diagram showing an example of a processing procedure of control in section 6 of the fifth embodiment of the present technology.
- Second embodiment (example of adjusting frequency spectrum of audio signal according to posture variation)
- FIG. 1 is a diagram showing a configuration example of an audio signal processing device 100 according to the first embodiment of the present technology.
- the audio signal processing device 100 includes an audio signal input unit 110, a sensor signal input unit 120, a posture variation analysis unit 1300, an audio signal analysis unit 1440, A control unit 150, an audio signal adjustment unit 160, and an audio signal output unit 170 are provided.
- the audio signal input section 110 receives an audio signal input to the audio signal processing device 100.
- the audio signal input unit 110 may be a device such as a microphone, for example.
- the audio signal input unit 110 converts the input audio signal from an analog signal to a digital signal (8/O conversion), and converts the digital signal data into an audio signal analysis unit 140 and an audio signal adjustment unit. Supply to 160.
- the sensor signal input unit 120 receives the posture information such as acceleration or angular velocity applied to the audio signal processing device 100.
- the sensor signal input section 120 may be, for example, an acceleration sensor such as a gyro sensor or an angular velocity sensor. In this case, the sensor signal input section 120 is an example of the sensor described in the claims.
- the sensor signal input unit 120 converts the input attitude information into /O and supplies the digital signal data to the attitude variation analysis unit 130.
- the posture variation analysis unit 130 analyzes the posture variation of the audio signal processing device 100 based on the posture information supplied from the sensor signal input unit 120.
- the attitude variation analysis unit 130 generates an attitude variation amount as a result of the analysis, and supplies the attitude variation amount to the control unit 150.
- the audio signal analysis unit 140 is an audio signal supplied from the audio signal input unit 110. ⁇ 2020/174776 7 ⁇ (: 170?2019/045255
- the voice signal analysis unit 140 measures the voice signal level and supplies the maximum signal level as a target value to the control unit 150.
- the control unit 150 controls the audio signal adjustment in the audio signal adjusting unit 160. This control unit 150 determines the gain as the adjustment parameter of the audio signal according to the posture variation amount supplied from the posture variation analysis unit 1300 and the target value supplied from the audio signal analysis unit 140. Supply to the adjusting unit 160.
- the audio signal adjusting unit 1600 multiplies the audio signal supplied from the audio signal input unit 1100 by the gain supplied from the control unit 150, and outputs the audio corresponding to the gain. The signal is adjusted. The audio signal adjusting unit 160 supplies the adjusted audio signal to the audio signal output unit 170.
- the audio signal output unit 170 converts the adjusted audio signal supplied from the audio signal adjustment unit 160 from a digital signal to an analog signal (mouth/eight conversion), and outputs the audio of the analog signal. It outputs a signal.
- FIG. 2 is a diagram showing a waveform example of each part of the audio signal processing device 100 according to the first embodiment of the present technology.
- Reference numeral 3 in the same figure shows the time variation of the posture variation amount output from the posture variation analysis unit 130.
- the box in the figure shows the time change of the audio signal output from the audio signal input unit 110.
- ⁇ indicates the time change of the gain output from the control unit 150, that is, the gain that adjusts the audio signal level according to the audio signal and the amount of posture variation.
- ⁇ indicates the time change of the audio signal after being adjusted by the audio signal adjusting unit 160. This shows that the level of the audio signal is adjusted appropriately.
- the target period is classified into six sections from section 1 to section 6. These sections make state transitions according to the magnitude of posture change and duration.
- Section 1 is a state in which the amount of posture variation is smaller than the threshold 1, and the gain ⁇ 1 ⁇ 2020/174776 8 ⁇ (: 170?2019/045255
- the state in which the amount of posture variation is larger than the threshold value 1 continues for time 1 or more.
- the maximum audio signal level !_ 2 is acquired.
- the target gain 02 is set as described later.
- Section 3 is a state in which the level of the output audio signal drops as the gain is continuously decreased from ⁇ 1 to 0 2 with a predetermined step size 3 1.
- the gain is fixed to 0 2, and the level of the output audio signal is maintained at a lower level.
- the level of the output audio signal rises as the gain is continuously increased from 0 2 to ⁇ 1 with a predetermined step size of 32.
- the same threshold value is used for the posture variation amount, but different threshold values may be used for each judgment.
- FIG. 3 is a diagram showing an example of state transition of each section in the first embodiment of the present technology.
- section 3 while observing the amount of posture variation, if the amount of posture variation is large, it gradually decreases to the target gain, and when the target gain is reached, it transits to section 4. On the other hand, when the amount of posture change becomes small, it transits to section 5.
- section 4 while observing the amount of posture variation, if the amount of posture variation is large, the gain is maintained, and if the amount of posture variation is small, it transits to section 5.
- section 6 while observing the amount of posture variation, if the amount of posture variation is small, it gradually increases to the target gain, and when the target gain is reached, it transits to section 1. On the other hand, when the amount of posture change becomes large, it transits to section 2.
- FIG. 4 is a diagram showing an example of a processing procedure of control in the section 1 of the first embodiment of the present technology.
- the audio signal analysis unit 140 analyzes the audio signal supplied from the audio signal input unit 1 10 and detects a feature amount (step 3911). Then, for example, the maximum level !_ 1 of the audio signal is held as the feature amount (step 3 9 1 2).
- the posture variation analysis unit 130 analyzes the posture of the audio signal processing device 100 based on the posture information supplied from the sensor signal input unit 120, and detects the posture variation amount ( Step 3 9 1 3).
- the control unit 150 compares the posture variation amount with the threshold value 1 (step 3 9 1 4)
- step 3911:NO If the amount of posture change is smaller than the threshold value 1 (step 3911:NO), it is assumed that there is no posture change, and steps 3911 and after are repeated. If the amount of posture variation is greater than the threshold 1, (Step 3 9 1 4: 3) Then, it is assumed that there is a posture change and transits to section 2 (steps 3 9 16).
- FIG. 5 is a diagram showing an example of a processing procedure of control in the section 2 of the first embodiment of the present technology.
- the voice signal analysis unit 140 analyzes the voice signal supplied from the voice signal input unit 110 to detect the feature amount (step 3922). Then, for example, the maximum level !_ 2 of the audio signal is held as the feature amount (step 3922).
- the posture variation analysis unit 130 analyzes the posture of the audio signal processing device 100 based on the posture information supplied from the sensor signal input unit 120, and detects the posture variation amount ( Step 3 9 2 3).
- the control unit 150 compares the posture variation amount with the threshold value 1 (step 3924). ⁇ 2020/174776 10 ⁇ (: 170?2019/045255
- step 3924: N 0 If the amount of posture change is smaller than the threshold value 1 (step 3924: N 0), it is determined that there is no posture change. On the other hand, if the amount of posture change is greater than the threshold value 1 (steps 3 9 2 4 :gray 63), it is assumed that there is a posture change.
- step 3 9 2 1 When there is no fluctuation amount, the processes from step 3 9 2 1 are repeated (step 3 929: N 0) and when the condition without fluctuation amount is continuously detected 1 ⁇ 11 times (step 3 9 2 9: ⁇ 6 3), transition to section 1 (step 3 9 3 0).
- step 3 9 2 6 In a state where there is a posture change, the processing from step 3 9 2 1 is repeated (step 3 9 2 6 :! ⁇ 10), and when a state with a posture change is detected IV! 1 time, a voice signal is detected.
- the target gain ⁇ 2 corresponds to the maximum level !_ 1 of the voice signal due to the gain ⁇ 1 in the steady state in the section 1, for example, corresponding to the maximum level !_ 2 of the voice signal in the section 2. Therefore, it can be calculated by the following formula.
- the waiting time 1 may be elapsed. It is also possible to minimize section 2 by setting the number of detections 1 ⁇ /1 1 or waiting time 1 to 0.
- FIG. 6 is a diagram showing an example of a processing procedure of control in the section 3 of the first embodiment of the present technology.
- the posture variation analysis unit 130 analyzes the posture of the audio signal processing device 100 based on the posture information supplied from the sensor signal input unit 120, and detects the posture variation amount ( Step 3 9 3 1).
- the control unit 150 compares the posture variation amount with the threshold value 1 (step 393 2).
- step 393 2 :N 0 If the amount of posture change is smaller than the threshold value 1 (step 393 2 :N 0), it is assumed that there is no posture change. On the other hand, if the amount of posture change is greater than the threshold value 1 (step 3932: V63), it is assumed that there is a posture change.
- step 393 1 When there is no fluctuation amount, the processing from step 393 1 is repeated (step 3939 8 :! ⁇ ! ⁇ ), and if the condition without fluctuation amount is detected continuously ! ⁇ 1 once. ⁇ 2020/174776 1 1 ⁇ (: 170?2019/045255
- Step 3 9 3 8 ⁇ 6 3
- transition to section 5 Step 3 9 3 9
- the audio signal adjusting unit 160 adjusts the audio signal by multiplying the audio signal by a gain (step 393 5).
- a gain As a specific example, every time it is determined that the amount of fluctuation is large (step 3 9 3 2 : ⁇ 6 3 ), the default step size 31 minutes is decreased from the current gain, and this step size 3 1 is changed to the initial value. ⁇ It is possible to skip section 3 by setting 1-target value 0 2”.
- Step 3939 it is determined whether or not the gain has decreased to the target value 2 (step 3939). Until target value ⁇ 2 is reached (Step 3 9 3 6 :N 0) Steps 3 9 3 1 and after are repeated, and when target value ⁇ 2 is reached (Step 3 9 3 6 :V 6 3) Transition to section 4 Yes (Step 3 9 3 7) 0
- FIG. 7 is a diagram showing an example of a processing procedure of control in section 4 of the first embodiment of the present technology.
- the attitude variation analysis unit 130 analyzes the attitude of the audio signal processing device 100 based on the attitude information supplied from the sensor signal input unit 120, and detects the attitude fluctuation amount ( Steps 3 9 4 1).
- the control unit 150 compares the amount of posture variation with the threshold 1 (step 3942).
- step 3942: N 0 If the amount of posture change is smaller than the threshold value 1 (step 3942: N 0), it is assumed that there is no posture change. On the other hand, if the amount of posture change is larger than the threshold value 1 (step 3942: step 63), it is assumed that there is a posture change, and the processes from step 394 1 are repeated.
- step 394 1 In the state where there is no fluctuation amount, the processing from step 394 1 is repeated (step 3945: N 0), and when the state without fluctuation amount is continuously detected 1 ⁇ 11 times (step 3 9 4 5: ⁇ 6 3), transition to section 5 (step 3 9 4 6).
- FIG. 8 is a diagram showing an example of a processing procedure of control in section 5 according to the first embodiment of the present technology.
- the posture variation analysis unit 1300 is provided with the posture information supplied from the sensor signal input unit 1120. ⁇ 2020/174776 12 ⁇ (: 170?2019/045255
- the posture of the voice signal processor 100 is analyzed and the posture fluctuation amount is detected (step 3951).
- the control unit 150 compares the amount of posture variation with the threshold value 1 (step 3952).
- step 395 2 :N 0 If the amount of posture change is smaller than the threshold 1 (step 395 2 :N 0), it is assumed that there is no posture change. On the other hand, if the amount of posture change is larger than the threshold value 1 (step 3952: step 63), it is assumed that there is a posture change.
- step 395 5 3: N 0 the processing from step 395 1 is repeated (step 395 5 3: N 0) and when the state without fluctuation amount is continuously detected IV! 2 times (step 3 9 5 3 :Take 6 3), and transit to section 6 (step 3 9 5 4).
- Step 395 7 N 0
- Step 395 7: ⁇ 63 3 If a state with posture change is detected N 2 times (Step 395 7: ⁇ 63 3), it transits to section 4 (Step 395 8).
- FIG. 9 is a diagram showing an example of a control processing procedure in the section 6 of the first embodiment of the present technology.
- the posture variation analysis unit 130 analyzes the posture of the audio signal processing device 100 based on the posture information supplied from the sensor signal input unit 120, and detects the posture variation amount ( Step 3 9 6 1).
- the control unit 150 compares the posture variation amount with the threshold value 1 (step 3966).
- step 396 2 :N 0 If the amount of posture change is smaller than the threshold value 1 (step 396 2 :N 0), it is assumed that there is no posture change. On the other hand, when the amount of posture change is larger than the threshold value 1 (step 3966: V63), it is assumed that there is a posture change.
- step 3966 When there is a fluctuation amount, the processing from step 3966 is repeated (step 3966: 1 ⁇ 10), and when the fluctuation amount is detected twice in succession (step 3 9 6 8: ⁇ 6 3), transition to section 2 (step 3 9 6 9).
- the audio signal adjusting unit 160 ⁇ 2020/174776 13 ⁇ (: 170?2019/045255
- step 3 9 6 3 adjusts the audio signal by multiplying the gain.
- step 3966: N 0 every time it is determined that the amount of fluctuation is small (step 3966: N 0), the default step size 3 2 minutes is increased from the current gain, and this step size 3 2 is changed to the initial value. ⁇ It is possible to skip section 6 by setting 1-target value 0 2”.
- Step 3966 it is judged whether or not the gain has increased to the target value ⁇ 1 (step 3966). Until the target value ⁇ 1 is reached (Step 3 9 64: N 0) Steps 3 9 6 1 and after are repeated, and when the target value ⁇ 1 is reached (Step 3 9 6 4 :Head 6 3) Transition to section 1 Yes (Step 3965) Yes
- FIG. 10 is a diagram illustrating an example of a processing procedure of the posture variation analysis processing in the first embodiment of the present technology.
- the attitude variation analysis unit 1300 obtains attitude information such as acceleration or angular velocity from the sensor signal input unit 1120 (step 3818). Then, the posture variation analysis unit 130 calculates the difference between the newly acquired posture information and the previously stored data (step 3812), and calculates the amount of posture variation from the difference value (step 3 8 1 3). Then, the posture variation analysis unit 1330 updates the newly calculated posture variation by holding it as new data (steps 318 and 14).
- FIG. 11 is a diagram illustrating an example of a processing procedure of the audio signal analysis processing according to the first embodiment of the present technology.
- the audio signal input unit 1 1 0 converts the input audio signal into 8/ ⁇ (step 3 82 1).
- the audio signal analysis unit 140 measures the signal level of the audio data supplied from the audio signal input unit 110, detects the maximum value in the measurement section, and supplies it to the control unit 150 (step 3 8 2 2).
- the audio signal analysis unit 140 temporarily stores the maximum value of the signal level (step 382 3).
- the maximum value of the stored signal level is updated every time the speech analysis processing is executed. ⁇ 2020/174776 14 ⁇ (: 170?2019/045255
- FIG. 12 is a diagram showing an example of a processing procedure of the audio signal adjustment processing in the first embodiment of the present technology.
- the attitude fluctuation analysis unit 130 analyzes the attitude of the audio signal processing device 100 based on the attitude information supplied from the sensor signal input unit 120, and detects the attitude fluctuation amount ( Step 3 8 3 1).
- the control unit 150 compares the amount of posture variation with the threshold value 1 (step 3823).
- step 382 3 :N 0 If the amount of posture change is smaller than the threshold value 1 (step 382 3 :N 0), it is determined that there is no posture change. On the other hand, if the amount of posture change is larger than the threshold value 1 (step 383 2 :V63), it is assumed that there is a posture change.
- step 3 8 3 3 If the gain does not reach the target value ⁇ 2 in the state where there is a posture change (step 3 8 3 3 : ⁇ 6 3), the gain is reduced by the predetermined width 3 1 minute (step 3 8 3 Four) . On the other hand, in the state where there is no fluctuation amount, if the gain has not reached the initial value ⁇ 1 (step 3 8 3 5: ⁇ 6 3), the gain is increased by the predetermined width 3 2 minutes (step 3 8 3 6 ).
- the audio signal adjustment unit 160 outputs the audio signal supplied from the audio signal input unit 110 by multiplying the audio signal supplied from the control unit 150 with the updated gain. Adjust the signal (step 3 8 3 7).
- the audio signal is adjusted by the audio signal adjusting unit 160 according to the variation amount.
- the level of the audio signal is adjusted according to the posture variation amount.
- the frequency scan of the audio signal is ⁇ 2020/174776 1 5 (: 170?2019/045255
- FIG. 13 is a diagram showing an example of adjusting the frequency spectrum of an audio signal according to the second embodiment of the present technology.
- the horizontal axis represents frequency and the vertical axis represents gain.
- the audio signal has a constant bandwidth and signal strength distribution as shown by the solid line.
- the noise spectrum may be distributed at frequencies other than the normal bandwidth, or noise spectrum may appear as a strong signal even within the bandwidth.
- the speech signal analysis unit 140 performs a spectrum analysis of the speech signal, and the speech signal adjustment unit 160 is provided with a filter to obtain the result of the spectrum analysis.
- the filter characteristic is switched according to.
- the filter characteristic as indicated by the downward arrow is applied stepwise in the state where there is posture variation, and the filter characteristic is switched to low gain or narrow band.
- the filter characteristics as shown by the upward arrow are applied stepwise in the absence of posture change, the filter characteristics are switched to high gain and wide band, and finally the signal not to be filtered. Output.
- the overall configuration of the second embodiment is similar to that of the first embodiment described above, and detailed description thereof will be omitted.
- FIG. 14 is a diagram illustrating an example of a processing procedure of audio signal analysis processing according to the second embodiment of the present technology.
- the audio signal input unit 1 1 0 converts the input audio signal into 8/O (step 3 84 1).
- the audio signal analysis unit 140 analyzes the frequency spectrum of the audio data supplied from the audio signal input unit 110, and detects the maximum value of the bandwidth and signal strength of the frequency spectrum in the analysis section. And supplies it to the controller 150 (step 3842). ⁇ 2020/174776 16 ⁇ (:170?2019/045255
- the voice signal analysis unit 140 temporarily stores the bandwidth of the frequency spectrum and the signal strength (step 384 3).
- the bandwidth of the stored frequency spectrum and the maximum value of the signal strength are updated each time the speech analysis processing is executed.
- FIG. 15 is a diagram showing an example of a processing procedure of audio signal adjustment processing in the second embodiment of the present technology.
- the attitude variation analysis unit 130 analyzes the attitude of the audio signal processing device 100 based on the attitude information supplied from the sensor signal input unit 120, and detects the amount of attitude fluctuation ( Step 3 8 5 1).
- the control unit 150 compares the amount of posture variation with a threshold value (step 3852). If the posture change amount is smaller than the threshold value (step 3852:N0), it is assumed that there is no posture change. On the other hand, if the posture variation amount is larger than the threshold value (step 3852:V63), it is assumed that there is a posture variation.
- step 3 8 5 3 If the gain does not reach the target value in the state where there is a posture change (step 3 8 5 3 : ⁇ 6 3), the gain is reduced by the predetermined width (step 3 8 5 4). If the current bandwidth does not reach the target value (step 3 8 6 Narrow the band by the specified width (steps 3 8 6 4).
- step 3 8 5 5: ⁇ 6 3 if the gain has not reached the initial value (step 3 8 5 5: ⁇ 6 3), increase the gain by the predetermined width (step 3 8 5 6) .. If the current band has not reached the initial value (step 386 5: ⁇ 6 3), the band is widened by the predetermined width (step 3 8 6 6).
- the audio signal adjusting unit 1600 makes the audio signal supplied from the audio signal inputting unit 1 1 0 by filtering the characteristics of the audio signal supplied from the controlling unit 1 50. Make adjustments (steps 3 8 6 7).
- the audio signal adjusting unit 1600 uses the filter characteristic according to the amount of attitude change generated by the attitude changing analysis unit 1300, and The signal is adjusted.
- the input audio signal is adjusted in real time.
- the voice signal and the sensor signal are recorded, the posture variation is analyzed from the sensor signal when the voice signal is reproduced, and the voice signal is adjusted based on this.
- FIG. 16 is a diagram showing a configuration example of an audio signal processing device 100 according to the third embodiment of the present technology.
- the audio signal processing device 100 includes an audio signal recording/reproducing unit 1 15 and a sensor signal recording/reproducing unit 1 2 5 in addition to the above-described first embodiment. It is different in that it is provided, and is otherwise similar to the first embodiment described above.
- the audio signal recording/reproducing unit 115 is for recording/reproducing audio signals.
- the sensor signal recording/reproducing unit 1 2 5 is for recording/reproducing sensor signals.
- the audio signal recording/reproducing unit 115 and the sensor signal recording/reproducing unit 125 are examples of the recording/reproducing unit described in the claims.
- the audio signal recording/playback unit 1 15 records the audio data from the audio signal input unit 1 10 when recording data.
- the sensor signal recording/reproducing unit 125 records the posture information such as acceleration or angular velocity from the sensor signal input unit 120 at the time of data recording.
- the audio signal recording/reproducing unit 1 15 and the sensor signal recording/reproducing unit 1 25 perform recording in synchronization with each other.
- the audio signal recording/reproducing unit 1 15 and the sensor signal recording/reproducing unit 1 2 5 synchronously reproduce the audio signal and the sensor signal, respectively, at the time of data reproduction, and respectively reproduce the audio signal analyzing unit 1 4 0 and the attitude. Supplied to the fluctuation analysis unit 1300. Subsequent processing is the same as in the above-described first embodiment.
- the audio signal recording/reproducing unit 1 As described above, according to the third embodiment of the present technology, the audio signal recording/reproducing unit 1
- the audio signal can be adjusted during reproduction.
- the sound signal is adjusted based on the posture variation amount, but in the fourth embodiment, the image signal is further corrected based on the posture variation amount.
- FIG. 17 is a diagram showing a configuration example of an audio signal processing device 100 according to the fourth embodiment of the present technology.
- the audio signal processing device 100 is similar to the first embodiment described above except that the image signal input unit 1 81, the image signal correction unit 1 8 2 and the image signal The difference is that a signal output unit 183 is provided, and the other points are the same as in the above-described first embodiment.
- the image signal input unit 181 receives an image signal supplied from an image sensor (not shown) and supplies it as an image frame to the image signal correction unit 182.
- the image signal correction unit 182 performs the keystone correction process by performing the keystone correction processing on the image frame from the image signal input unit 181 based on the amount of posture change from the posture change analysis unit 1300. It is a thing.
- the image signal output unit 183 outputs the image signal that has been subjected to the shake correction by the image signal correction unit 182. As a result, similarly to the adjustment of the audio signal, the image signal is subjected to the shake correction based on the posture variation amount.
- the image signal is corrected in the image signal correction unit 182 based on the amount of posture variation. be able to. That is, at the same time as the image shake correction of the image, it is possible to suppress the noise caused by the shake due to the reproduced sound.
- the noise of the audio signal due to the posture variation is suppressed by adjusting the audio signal according to the amount of posture variation.
- the gain of the audio signal may be reduced even when noise is not generated due to the posture change. Therefore, in the fifth embodiment, the sound signal is not adjusted if the sound does not change even if the posture changes.
- the voice signal is adjusted only when the posture variation and the voice variation have a high correlation.
- FIG. 18 is a diagram showing a configuration example of an audio signal processing device 100 according to the fifth embodiment of the present technology.
- the audio signal analysis unit 140 is provided with the audio signal from the audio signal input unit 110 and the posture variation analysis unit 1300.
- the correlation value with the posture variation amount of is calculated. In the calculation of the correlation value, the correlation value is high when the amount of posture change is large and larger than the voice level in section 1, and the correlation value is low when the amount of posture change is large and the voice level is small.
- the control unit 150 controls the audio signal adjustment unit 160 so that the audio signal adjustment process is performed based on the correlation value from the audio signal analysis unit 1440 only when the correlation value is high. To do.
- FIG. 19 is a diagram illustrating an example of a processing procedure of audio signal analysis processing according to the fifth embodiment of the present technology.
- the audio signal analysis unit 1440 receives an audio signal from the audio signal input unit 110. Then, the correlation value between the posture variation analysis unit 1300 and the posture variation amount is calculated (step 3824). The calculated correlation value is supplied to the control unit 150.
- FIG. 20 is a diagram showing an example of a processing procedure of control in the section 1 of the fifth embodiment of the present technology.
- the control unit 150 analyzes the voice signal. Based on the correlation value from the section 140, it is judged whether or not the posture variation correlates with the voice variation (step 3915). If it is determined that they do not correlate (step 391 5: N 0), it is assumed that there is no posture change, and ⁇ 2020/174776 20 ⁇ (: 170?2019/045255
- step 391 5 ⁇ 6 3
- the posture change is regarded as significant and transition is made to section 2 (step 3 9 16).
- FIG. 21 is a diagram showing an example of a processing procedure of control in the section 2 of the fifth embodiment of the present technology.
- the control unit 150 causes the voice signal analysis to be performed when there is a posture change. Based on the correlation value from the section 140, it is judged whether or not the posture variation is correlated with the voice variation (step 3925). If it is determined that they do not correlate (step 392 5: N 0 ), it is assumed that the posture has not changed, and the processes after step 329 29 are performed. If it is determined that there is a correlation (step 392 5: ⁇ 6 3 ), it is determined that the posture change is significant, and the processes after step 392 6 are performed.
- FIG. 22 is a diagram showing an example of a processing procedure of control in the section 3 of the fifth embodiment of the present technology.
- the voice signal analysis unit 140 outputs a voice when there is a posture change. Perform signal analysis processing (step 393 3). Then, the control unit 150 determines, based on the correlation value from the voice signal analysis unit 140, whether or not the posture variation correlates with the voice variation (step 3939). If it is determined that there is no correlation (step 393 4 :N 0 ), it is assumed that there is no posture change, and the processing from step 393 1 onward is repeated. If it is determined that there is a correlation (step 3 9 3 4: 3) Then, the posture change is regarded as significant, and the voice signal adjusting unit 160 is controlled to perform the voice signal adjusting process (step 3935).
- Fig. 23 is a diagram showing an example of a processing procedure of control in a section 4 of the fifth embodiment of the present technology.
- the sound signal analysis unit 140 outputs the sound when there is a posture change. ⁇ 2020/174776 21 ⁇ (: 170?2019/045255
- step 3944 Performs voice signal analysis processing (step 3944). Then, the control unit 150 determines, based on the correlation value from the voice signal analysis unit 140, whether or not the posture variation correlates with the voice variation (step 3944). If it is determined that they do not correlate (step 3944:N0), it is assumed that the posture has not changed, and the processing from step 3945 is performed. If it is determined that there is a correlation (step 394 4: ⁇ 6 3 ), it is determined that the posture change is significant, and the processes from step 3 94 1 are performed.
- FIG. 24 is a diagram showing an example of a processing procedure of control in section 5 of the fifth embodiment of the present technology.
- the voice signal analysis unit 140 outputs a voice when there is a posture change. Perform signal analysis processing (step 3995). Then, the control unit 150 determines whether or not the posture variation is correlated with the voice variation based on the correlation value from the voice signal analysis unit 140 (step 3995). If it is determined that they do not correlate (step 395 5: N 0), it is assumed that there is no posture change, and the processing from step 395 3 is repeated. If it is determined that there is a correlation (step 395 6: step 63), it is determined that the posture change is significant, and the processing from step 395 7 is performed.
- FIG. 25 is a diagram showing an example of a processing procedure of control in section 6 of the fifth embodiment of the present technology.
- the voice signal analysis unit 140 outputs a voice when there is a posture change.
- Perform signal analysis processing step 3966.
- the control unit 150 determines, based on the correlation value from the voice signal analysis unit 140, whether or not the posture variation correlates with the voice variation (step 3966). If it is determined that they do not correlate (step 396 7: N 0 ), it is assumed that the posture has not changed, and the processing from step 396 1 onward is repeated. If it is judged that there is a correlation (step 3 9 6 7: ⁇ 6 3), the posture change is regarded as significant and step 3 ⁇ 2020/174776 22 ⁇ (: 170?2019/045255
- the audio signal by adjusting the audio signal only when the correlation between the posture variation and the voice variation is high, the voice is reproduced even if the posture variation occurs. If it does not fluctuate, the audio signal may not be adjusted.
- the processing procedure described in the above embodiments may be regarded as a method having a series of these procedures, and a program for causing a computer to execute the series of procedures or a program for storing the program is stored. It may be regarded as a recording medium.
- a recording medium for example, a CD (Compact Disc), an MD (MiniDisc), a DVD (Digital Versa ⁇ i le Disc), a memory card, a Blu-ray (Blu-ray (registered trademark) Disc), or the like can be used. it can.
- the present technology may have the following configurations.
- an audio signal analysis unit that acquires an audio signal and sets a target value for audio adjustment based on the audio signal
- a posture change analysis unit that acquires posture information and generates a posture variation amount based on the posture information
- An audio signal adjustment unit that adjusts the audio signal toward the target value according to the posture variation amount
- An audio signal processing device comprising: 20/174776 23 ⁇ (: 170?2019/045255
- the voice signal adjusting unit adjusts the voice signal so that the volume of the voice signal is reduced when the posture variation amount becomes larger than a first threshold value, and the posture variation amount is adjusted after reaching the target value. Is smaller than the second threshold, the audio signal is adjusted so that the volume of the audio signal is returned.
- the audio signal processing device according to (1) above.
- the audio signal adjusting unit adjusts the audio signal so that the bandwidth of the frequency of the audio signal is narrowed when the posture variation becomes larger than a first threshold value, and after the target value is reached, When the posture variation amount becomes smaller than a second threshold value, the audio signal is adjusted so as to restore the frequency bandwidth of the audio signal.
- the audio signal processing device according to (1) or (2) above.
- the audio signal adjusting unit adjusts the audio signal so as to reduce the gain of the frequency of the audio signal when the posture variation amount becomes larger than a first threshold value, and after the target value is reached, the audio signal adjusting unit adjusts the audio signal.
- the voice signal is adjusted so as to restore the frequency gain of the voice signal
- the audio signal processing device according to any one of (1) to (3) above.
- the audio signal adjustment unit is configured such that one of the states in which the posture variation amount is larger than a first threshold value and the posture variation amount is smaller than a second threshold value continues for a predetermined period. Adjust the audio signal in case
- the audio signal processing device according to any one of (1) to (4) above.
- the audio signal adjusting unit adjusts the audio signal stepwise with a fixed step size.
- the audio signal processing device according to any one of (1) to (5) above.
- the audio signal processing device according to any one of (1) to (6), further including a sensor that detects acceleration or angular velocity and generates the posture information.
- a recording/reproducing unit that synchronously records and reproduces the audio signal and the posture information
- the audio signal analysis unit sets the target value based on the reproduced audio signal: ⁇ 2020/174776 24 ⁇ (: 170?2019/045255
- the posture variation analysis unit generates a posture variation amount based on the reproduced posture information
- the audio signal adjusting unit adjusts the reproduced audio signal according to the reproduced posture variation amount.
- the audio signal processing device according to any one of (1) to (7) above.
- An image signal correction unit is further provided that acquires an image signal synchronized with the audio signal and corrects the blur of the image signal according to the posture variation amount.
- the audio signal processing device according to any one of (1) to (8) above.
- the audio signal adjustment unit adjusts the audio signal when there is a correlation between the attitude change indicated by the attitude change amount and the change of the audio signal.
- the audio signal processing device according to any one of (1) to (9) above.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Signal Processing (AREA)
- Audiology, Speech & Language Pathology (AREA)
- General Health & Medical Sciences (AREA)
- Human Computer Interaction (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Circuit For Audible Band Transducer (AREA)
- Stereophonic System (AREA)
Abstract
周囲の音声の状況に応じて、姿勢変動量に基づいた音声信号の調整を行う。 音声信号処理装置は、音声信号解析部と、姿勢変動解析部と、音声信号調整部とを備える。音声信号解析部は、音声信号を取得して、その音声信号に基づいて音声調整のための目標値を設定する。姿勢変動解析部は、姿勢情報を取得して、その姿勢情報に基づいて姿勢変動量を生成する。音声信号調整部は、姿勢変動量に応じて、目標値に向けて、音声信号を調整する。
Description
\¥0 2020/174776 1 卩(:17 2019/045255 明 細 書
発明の名称 : 音声信号処理装置
技術分野
[0001 ] 本技術は、 音声信号処理装置に関する。 詳しくは、 姿勢変動量に基づいて 音声信号を調整する音声信号処理装置に関する。
背景技術
[0002] 撮像装置においては、 ジャイロセンサなどを用いて動き検出を行い、 その 検出された動きに応じてフレームメモリから画像を切り出すことによって手 振れ補正を行う技術が知られている (例えば、 特許文献 1参照。 ) 。 また、 楽器をシミュレートするゲームシステムなどにおいては、 取得された姿勢変 化量に応じて音声を制御する技術が提案されている (例えば、 特許文献 2参 照。 ) 〇
先行技術文献
特許文献
[0003] 特許文献 1 :特開 2 0 1 7 - 1 9 9 9 9 5号公報
特許文献 2 :特開 2 0 1 2 _ 2 3 0 1 3 5号公報
発明の概要
発明が解決しようとする課題
[0004] 上述の姿勢変化量に応じて音声を制御する従来技術では、 楽器のシミュレ —卜を想定しており、 楽器を操作するための情報として姿勢変化量に着目し ている。 これを撮像装置に適用して、 画像の手振れ補正に倣って音声を補正 しようとした場合、 周囲の音声の状況とは無関係に姿勢変化量だけを利用し て音声を変化させてしまうと、 周囲の音声の状況と不整合が生じて、 視聴者 に違和感を与えるおそれがある。
[0005] 本技術はこのような状況に鑑みて生み出されたものであり、 周囲の音声の 状況に応じて姿勢変動量に基づいた音声信号の調整を行うことを目的とする
〇 2020/174776 2 卩(:170?2019/045255
課題を解決するための手段
[0006] 本技術は、 上述の問題点を解消するためになされたものであり、 その第 1 の側面は、 音声信号を取得して上記音声信号に基づいて音声調整のための目 標値を設定する音声信号解析部と、 姿勢情報を取得して上記姿勢情報に基づ いて姿勢変動量を生成する姿勢変動解析部と、 上記姿勢変動量に応じて上記 目標値に向けて上記音声信号を調整する音声信号調整部とを具備する音声信 号処理装置である。 これにより、 姿勢変動量に応じて目標値に向けて音声信 号を調整するという作用をもたらす。
[0007] また、 この第 1の側面において、 上記音声信号調整部は、 上記姿勢変動量 が第 1の閾値より大きくなると上記音声信号の音量を下げるように上記音声 信号を調整し、 上記目標値に達した後に上記姿勢変動量が第 2の閾値より小 さくなると上記音声信号の音量を戻すように上記音声信号を調整するように してもよい。 これにより、 姿勢変動量に応じて音量の増減を制御するという 作用をもたらす。
[0008] また、 この第 1の側面において、 上記音声信号調整部は、 上記姿勢変動量 が第 1の閾値より大きくなると上記音声信号の周波数の帯域幅を狭くするよ うに上記音声信号を調整し、 上記目標値に達した後に上記姿勢変動量が第 2 の閾値より小さくなると上記音声信号の周波数の帯域幅を戻すように上記音 声信号を調整するようにしてもよい。 これにより、 姿勢変動量に応じて周波 数の帯域幅を制御するという作用をもたらす。
[0009] また、 この第 1の側面において、 上記音声信号調整部は、 上記姿勢変動量 が第 1の閾値より大きくなると上記音声信号の周波数のゲインを下げるよう に上記音声信号を調整し、 上記目標値に達した後に上記姿勢変動量が第 2の 閾値より小さくなると上記音声信号の周波数のゲインを戻すように上記音声 信号を調整するようにしてもよい。 これにより、 姿勢変動量に応じて周波数 のゲインを制御するという作用をもたらす。
[0010] また、 この第 1の側面において、 上記音声信号調整部は、 上記姿勢変動量 が第 1の閾値より大きい状態、 および、 上記姿勢変動量が第 2の閾値より小
〇 2020/174776 3 卩(:170?2019/045255
さい状態の、 何れか一方の状態が所定期間継続した場合に上記音声信号を調 整するようにしてもよい。 これにより、 音声信号の調整においてヒステリシ ス特性を持たせることにより、 不自然な挙動を抑制するという作用をもたら す。
[001 1 ] また、 この第 1の側面において、 上記音声信号調整部は、 一定量のステッ プサイズで段階的に上記音声信号を調整するようにしてもよい。 これにより 、 再生音声における違和感を抑制するという作用をもたらす。
[0012] また、 この第 1の側面において、 加速度または角速度を検出して上記姿勢 情報を生成するセンサをさらに具備してもよい。 これにより、 装置内で検知 された姿勢情報に基づいて音声信号を調整するという作用をもたらす。
[0013] また、 この第 1の側面において、 上記音声信号および上記姿勢情報を同期 して記録および再生する記録再生部をさらに具備し、 上記音声信号解析部は 、 上記再生された音声信号に基づいて上記目標値を設定し、 上記姿勢変動解 析部は、 上記再生された姿勢情報に基づいて姿勢変動量を生成し、 上記音声 信号調整部は、 上記再生された姿勢変動量に応じて上記再生された音声信号 を調整するようにしてもよい。 これにより、 一旦記録された音声信号および 姿勢情報に基づいて、 再生時に音声信号を調整するという作用をもたらす。
[0014] また、 この第 1の側面において、 上記音声信号に同期する画像信号を取得 して上記姿勢変動量に応じて上記画像信号のブレを補正する画像信号補正部 をさらに具備してもよい。 これにより、 音声信号の調整とともに、 画像信号 のブレを補正するという作用をもたらす。
[0015] また、 この第 1の側面において、 上記音声信号調整部は、 上記姿勢変動量 が示す姿勢変動と上記音声信号の変動との間に相関が有る場合に上記音声信 号を調整するようにしてもよい。 これにより、 姿勢変動があっても音声が変 動していない場合には音声信号の調整を抑制するという作用をもたらす。 図面の簡単な説明
[0016] [図 1 ]本技術の第 1の実施の形態における音声信号処理装置 1 〇〇の構成例を 示す図である。
20/174776 4 卩(:170?2019/045255
[図 2]本技術の第 1の実施の形態における音声信号処理装置 1 0 0の各部の波 形例を示す図である。
[図 3]本技術の第 1の実施の形態における各区間の状態遷移の例を示す図であ る。
[図 4]本技術の第 1の実施の形態の区間 1 における制御の処理手順例を示す図 である。
[図 5]本技術の第 1の実施の形態の区間 2における制御の処理手順例を示す図 である。
[図 6]本技術の第 1の実施の形態の区間 3における制御の処理手順例を示す図 である。
[図 7]本技術の第 1の実施の形態の区間 4における制御の処理手順例を示す図 である。
[図 8]本技術の第 1の実施の形態の区間 5における制御の処理手順例を示す図 である。
[図 9]本技術の第 1の実施の形態の区間 6における制御の処理手順例を示す図 である。
[図 10]本技術の第 1の実施の形態における姿勢変動解析処理の処理手順例を 示す図である。
[図 1 1]本技術の第 1の実施の形態における音声信号解析処理の処理手順例を 示す図である。
[図 12]本技術の第 1の実施の形態における音声信号調整処理の処理手順例を 示す図である。
[図 13]本技術の第 2の実施の形態における音声信号の周波数スぺクトルの調 整例を示す図である。
[図 14]本技術の第 2の実施の形態における音声信号解析処理の処理手順例を 示す図である。
[図 15]本技術の第 2の実施の形態における音声信号調整処理の処理手順例を 示す図である。
〇 2020/174776 5 卩(:170?2019/045255
[図 16]本技術の第 3の実施の形態における音声信号処理装置 1 0 0の構成例 を示す図である。
[図 17]本技術の第 4の実施の形態における音声信号処理装置 1 0 0の構成例 を示す図である。
[図 18]本技術の第 5の実施の形態における音声信号処理装置 1 0 0の構成例 を示す図である。
[図 19]本技術の第 5の実施の形態における音声信号解析処理の処理手順例を 示す図である。
[図 20]本技術の第 5の実施の形態の区間 1 における制御の処理手順例を示す 図である。
[図 21]本技術の第 5の実施の形態の区間 2における制御の処理手順例を示す 図である。
[図 22]本技術の第 5の実施の形態の区間 3における制御の処理手順例を示す 図である。
[図 23]本技術の第 5の実施の形態の区間 4における制御の処理手順例を示す 図である。
[図 24]本技術の第 5の実施の形態の区間 5における制御の処理手順例を示す 図である。
[図 25]本技術の第 5の実施の形態の区間 6における制御の処理手順例を示す 図である。
発明を実施するための形態
[0017] 以下、 本技術を実施するための形態 (以下、 実施の形態と称する) につい て説明する。 説明は以下の順序により行う。
1 . 第 1の実施の形態 (姿勢変動量に応じて音声信号のレベルを調整する 例)
2 . 第 2の実施の形態 (姿勢変動量に応じて音声信号の周波数スペクトル を調整する例)
3 . 第 3の実施の形態 (記録された音声信号を再生時に調整する例)
〇 2020/174776 6 卩(:170?2019/045255
4 . 第 4の実施の形態 (画像の手振れ補正を行う例)
5 . 第 5の実施の形態 (姿勢変動と相関が有る場合に音声信号を調整する 例)
[0018] < 1 . 第 1の実施の形態 >
[音声信号処理装置の構成]
図 1は、 本技術の第 1の実施の形態における音声信号処理装置 1 0 0の構 成例を示す図である。
[0019] この音声信号処理装置 1 0 0は、 音声信号入力部 1 1 0と、 センサ信号入 力部 1 2 0と、 姿勢変動解析部 1 3 0と、 音声信号解析部 1 4 0と、 制御部 1 5 0と、 音声信号調整部 1 6 0と、 音声信号出力部 1 7 0とを備える。
[0020] 音声信号入力部 1 1 0は、 音声信号処理装置 1 〇〇に入力される音声信号 を受け付けるものである。 この音声信号入力部 1 1 0は、 例えば、 マイクロ ホンなどの機器であってもよい。 この音声信号入力部 1 1 0は、 入力された 音声信号をアナログ信号からデジタル信号に変換 (八/〇変換) して、 その デジタル信号のデータを音声信号解析部 1 4 0および音声信号調整部 1 6 0 に供給する。
[0021 ] センサ信号入力部 1 2 0は、 音声信号処理装置 1 0 0に加えられる加速度 または角速度などの姿勢情報を受け付けるものである。 このセンサ信号入力 部 1 2 0は、 例えば、 ジャイロセンサなどの加速度センサまたは角速度セン サであってもよい。 この場合、 センサ信号入力部 1 2 0は、 特許請求の範囲 に記載のセンサの一例である。 このセンサ信号入力部 1 2 0は、 入力された 姿勢情報を /〇変換して、 そのデジタル信号のデータを姿勢変動解析部 1 3 0に供給する。
[0022] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から供給された姿勢情 報に基づいて音声信号処理装置 1 0 0の姿勢の変動を解析するものである。 この姿勢変動解析部 1 3 0は、 解析の結果として姿勢変動量を生成して、 そ の姿勢変動量を制御部 1 5 0に供給する。
[0023] 音声信号解析部 1 4 0は、 音声信号入力部 1 1 0から供給された音声信号
〇 2020/174776 7 卩(:170?2019/045255
を解析して、 その音声信号に基づいて音声調整のための目標値を設定するも のである。 この音声信号解析部 1 4 0は、 音声信号レベルを計測し、 その最 大信号レベルを目標値として制御部 1 5 0に供給する。
[0024] 制御部 1 5 0は、 音声信号調整部 1 6 0における音声信号調整を制御する ものである。 この制御部 1 5 0は、 姿勢変動解析部 1 3 0から供給された姿 勢変動量および音声信号解析部 1 4 0から供給された目標値に従って、 音声 信号の調整パラメータとしてのゲインを音声信号調整部 1 6 0に供給する。
[0025] 音声信号調整部 1 6 0は、 音声信号入力部 1 1 0から供給された音声信号 に対して、 制御部 1 5 0から供給されたゲインを掛けることによって、 その ゲインに応じた音声信号の調整を行うものである。 この音声信号調整部 1 6 〇は、 調整後の音声信号を音声信号出力部 1 7 0に供給する。
[0026] 音声信号出力部 1 7 0は、 音声信号調整部 1 6 0から供給された調整後の 音声信号をデジタル信号からアナログ信号に変換 (口/八変換) して、 その アナログ信号の音声信号を出力するものである。
[0027] [信号波形と区間]
図 2は、 本技術の第 1の実施の形態における音声信号処理装置 1 0 0の各 部の波形例を示す図である。
[0028] 同図における 3は、 姿勢変動解析部 1 3 0から出力される姿勢変動量の時 間変化を示している。 同図における匕は、 音声信号入力部 1 1 0から出力さ れる音声信号の時間変化を示している。 同図における〇は、 制御部 1 5 0か ら出力されるゲイン、 すなわち、 音声信号と姿勢変動量に応じて音声信号レ ベルを調整するゲインの時間変化を示している。 同図における は、 音声信 号調整部 1 6 0によって調整された後の音声信号の時間変化を示している。 これにより、 音声信号のレベルが適切に調整されることがわかる。
[0029] この例では、 対象となる期間を区間 1乃至区間 6の 6つの区間に分類して いる。 これらの区間は、 姿勢変動量の大きさや継続時間に応じて状態遷移す る。
[0030] 区間 1は、 姿勢変動量が閾値 1 より小さい状態であり、 ゲイン◦ 1 に対
〇 2020/174776 8 卩(:170?2019/045255
する音声信号のレベル !_ 1 を取得する。 なお、 2週目の区間 1 においては、 新たに音声信号のレベル !_ 3が取得されるものとしている。
[0031 ] 区間 2は、 姿勢変動量が閾値 1 より大きい状態が、 時間丁 1以上継続し ている状態である。 この区間 2において、 音声信号の最大レベル !_ 2を取得 する。 これにより、 後述するように目標ゲイン〇2が設定される。
[0032] 区間 3は、 ゲインを◦ 1から 0 2まで既定のステップサイズ 3 1で下げ続 けることにより、 出力される音声信号のレベルが降下していく状態である。
[0033] 区間 4は、 ゲインを 0 2に固定した状態であり、 出力される音声信号のレ ベルも下がつた状態が維持される。
[0034] 区間 5は、 姿勢変動量が閾値 1 より小さい状態が、 時間丁 2以上継続し ている状態である。
[0035] 区間 6は、 ゲインを 0 2から◦ 1 まで既定のステップサイズ 3 2で上げ続 けることにより、 出力される音声信号のレベルが上昇していく状態である。
[0036] なお、 この例では、 姿勢変動量の閾値を同じ 1 としたが、 それぞれの判 断において異なる閾値を用いるようにしてもよい。
[0037] [状態遷移]
図 3は、 本技術の第 1の実施の形態における各区間の状態遷移の例を示す 図である。
[0038] 区間 1では、 姿勢変動量が閾値より大きくなると区間 2に遷移する。
[0039] 区間 2では、 姿勢変動量が閾値より大きい状態が一定期間継続すると区間
3に遷移する。 一方、 一定期間の経過前に姿勢変動量が小さくなると区間 1 に民る。
[0040] 区間 3では、 姿勢変動量を観測しながら、 姿勢変動量が大きければ目標の ゲインまで徐々に下げていき、 目標のゲインに到達すると区間 4に遷移する 。 一方、 姿勢変動量が小さくなると区間 5に遷移する。
[0041 ] 区間 4では、 姿勢変動量を観測しながら、 姿勢変動量が大きければゲイン を維持し続け、 姿勢変動量が小さくなると区間 5に遷移する。
[0042] 区間 5では、 姿勢変動量が閾値より小さい状態が一定期間継続すると区間
〇 2020/174776 9 卩(:170?2019/045255
6に遷移する。 一方、 一定期間の経過前に姿勢変動量が大きくなると区間 4 に民る。
[0043] 区間 6では、 姿勢変動量を観測しながら、 姿勢変動量が小さければ目標の ゲインまで徐々に上げていき、 目標のゲインに到達すると区間 1 に遷移する 。 一方、 姿勢変動量が大きくなると区間 2に遷移する。
[0044] [動作]
図 4は、 本技術の第 1の実施の形態の区間 1 における制御の処理手順例を 示す図である。
[0045] 音声信号解析部 1 4 0は、 音声信号入力部 1 1 0から供給された音声信号 を解析して、 特徴量を検出する (ステップ 3 9 1 1) 。 そして、 特徴量とし て、 例えば音声信号の最大レベル !_ 1 を保持する (ステップ 3 9 1 2) 。
[0046] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から供給された姿勢情 報に基づいて音声信号処理装置 1 0 0の姿勢を解析し、 姿勢変動量を検出す る (ステップ3 9 1 3) 。
[0047] 制御部 1 5 0は、 姿勢変動量を閾値 1 と比較する (ステップ 3 9 1 4)
。 姿勢変動量が閾値 1 より小さい場合には (ステップ 3 9 1 4 : N 0) 、 姿勢変動が無いものとしてステップ 3 9 1 1以降を繰り返す。 姿勢変動量が 閾値 1 より大きい場合には (ステップ 3 9 1 4 :
3) , 姿勢変動が有 るものとして区間 2に遷移する (ステップ 3 9 1 6) 。
[0048] 図 5は、 本技術の第 1の実施の形態の区間 2における制御の処理手順例を 示す図である。
[0049] 音声信号解析部 1 4 0は、 音声信号入力部 1 1 0から供給された音声信号 を解析して、 特徴量を検出する (ステップ 3 9 2 1) 。 そして、 特徴量とし て、 例えば音声信号の最大レベル !_ 2を保持する (ステップ 3 9 2 2) 。
[0050] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から供給された姿勢情 報に基づいて音声信号処理装置 1 0 0の姿勢を解析し、 姿勢変動量を検出す る (ステップ3 9 2 3) 。
[0051 ] 制御部 1 5 0は、 姿勢変動量を閾値 1 と比較する (ステップ 3 9 2 4)
〇 2020/174776 10 卩(:170?2019/045255
。 姿勢変動量が閾値 1 より小さい場合には (ステップ 3 9 2 4 : N 0) 、 姿勢変動が無いものとする。 一方、 姿勢変動量が閾値 1 より大きい場合に は (ステップ 3 9 2 4 : 丫 6 3) 、 姿勢変動が有るものとする。
[0052] 変動量が無い状態においては、 ステップ 3 9 2 1以降の処理を繰り返し ( ステップ 3 9 2 9 : N 0) 、 変動量が無い状態を連続して 1\1 1回検出すると (ステップ 3 9 2 9 : 丫 6 3) 、 区間 1 に遷移する (ステップ 3 9 3 0) 。
[0053] 姿勢変動が有る状態においては、 ステップ3 9 2 1以降の処理を繰り返し (ステップ 3 9 2 6 : !\1〇) 、 姿勢変動が有る状態を IV! 1回検出すると、 音 声信号の最大レベル 1- 2に応じた目標ゲイン◦ 2を設定して (ステップ 3 9 2 7) 、 区間 3に遷移する (ステップ 3 9 2 8) 。
[0054] ここで、 目標ゲイン〇 2は、 例えば、 区間 1 における定常時のゲイン〇 1 による音声信号の最大レベル !_ 1 に対応して、 区間 2の音声信号の最大レべ ル !_ 2から、 次式により算出することができる。
[0055] なお、 区間 2から区間 3に遷移する条件としては、 IV! 1回の検出とする代 わりに、 待ち時間丁 1の経過としてもよい。 また、 検出回数 1\/1 1 または待ち 時間丁 1 を 0とすることにより、 区間 2を最短化することも可能である。
[0056] 図 6は、 本技術の第 1の実施の形態の区間 3における制御の処理手順例を 示す図である。
[0057] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から供給された姿勢情 報に基づいて音声信号処理装置 1 0 0の姿勢を解析し、 姿勢変動量を検出す る (ステップ3 9 3 1) 。
[0058] 制御部 1 5 0は、 姿勢変動量を閾値 1 と比較する (ステップ 3 9 3 2)
。 姿勢変動量が閾値 1 より小さい場合には (ステップ 3 9 3 2 : N 0) 、 姿勢変動が無いものとする。 一方、 姿勢変動量が閾値 1 より大きい場合に は (ステップ 3 9 3 2 : V 6 3) , 姿勢変動が有るものとする。
[0059] 変動量が無い状態においては、 ステップ 3 9 3 1以降の処理を繰り返し ( ステップ 3 9 3 8 : !\!〇) 、 変動量が無い状態を連続して !\1 1回検出すると
〇 2020/174776 1 1 卩(:170?2019/045255
(ステップ 3 9 3 8 : 丫6 3) 、 区間 5に遷移する (ステップ 3 9 3 9) 。
[0060] 姿勢変動が有る状態においては、 音声信号調整部 1 6 0が、 音声信号に対 してゲインを掛けることによって音声信号の調整を行う (ステップ 3 9 3 5 ) 。 具体例としては、 変動量が大きいと判定されるたびに (ステップ 3 9 3 2 : 丫6 3) 、 現在のゲインから既定のステップサイズ 3 1分を下げ、 この ステップサイズ 3 1 を 「初期値◦ 1 -目標値 0 2」 とすることにより、 区間 3をスキップすることも可能である。
[0061 ] 音声信号調整処理の完了判定として、 例えばゲインが目標値◦ 2まで下が ったか否かを判定する (ステップ 3 9 3 6) 。 目標値◦ 2に到達するまでは (ステップ 3 9 3 6 : N 0) ステップ 3 9 3 1以降を繰り返し、 目標値◦ 2 に到達すると (ステップ 3 9 3 6 : V 6 3) 区間 4に遷移する (ステップ 3 9 3 7) 0
[0062] 図 7は、 本技術の第 1の実施の形態の区間 4における制御の処理手順例を 示す図である。
[0063] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から供給された姿勢情 報に基づいて音声信号処理装置 1 0 0の姿勢を解析し、 姿勢変動量を検出す る (ステップ3 9 4 1) 。
[0064] 制御部 1 5 0は、 姿勢変動量を閾値 1 と比較する (ステップ 3 9 4 2)
。 姿勢変動量が閾値 1 より小さい場合には (ステップ 3 9 4 2 : N 0) 、 姿勢変動が無いものとする。 一方、 姿勢変動量が閾値 1 より大きい場合に は (ステップ 3 9 4 2 : 丫 6 3) 、 姿勢変動が有るものとして、 ステップ 3 9 4 1以降の処理を繰り返す。
[0065] 変動量が無い状態においては、 ステップ 3 9 4 1以降の処理を繰り返し ( ステップ 3 9 4 5 : N 0) 、 変動量が無い状態を連続して 1\1 1回検出すると (ステップ 3 9 4 5 : 丫6 3) 、 区間 5に遷移する (ステップ 3 9 4 6) 。
[0066] 図 8は、 本技術の第 1の実施の形態の区間 5における制御の処理手順例を 示す図である。
[0067] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から供給された姿勢情
〇 2020/174776 12 卩(:170?2019/045255
報に基づいて音声信号処理装置 1 0 0の姿勢を解析し、 姿勢変動量を検出す る (ステップ3 9 5 1) 。
[0068] 制御部 1 5 0は、 姿勢変動量を閾値 1 と比較する (ステップ 3 9 5 2)
。 姿勢変動量が閾値 1 より小さい場合には (ステップ 3 9 5 2 : N 0) 、 姿勢変動が無いものとする。 一方、 姿勢変動量が閾値 1 より大きい場合に は (ステップ 3 9 5 2 : 丫 6 3) 、 姿勢変動が有るものとする。
[0069] 変動量が無い状態においては、 ステップ 3 9 5 1以降の処理を繰り返し ( ステップ 3 9 5 3 : N 0) 、 変動量が無い状態を連続して IV! 2回検出すると (ステップ 3 9 5 3 : 丫 6 3) 、 区間 6に遷移する (ステップ 3 9 5 4) 。
[0070] 姿勢変動が有る状態においては、 ステップ 3 9 5 1以降の処理を繰り返し
(ステップ 3 9 5 7 : N 0) 、 姿勢変動が有る状態を N 2回検出すると (ス テップ3 9 5 7 : 丫 6 3) 、 区間 4に遷移する (ステップ 3 9 5 8) 。
[0071 ] なお、 区間 5から区間 6に遷移する条件としては、 IV! 2回の検出とする代 わりに、 待ち時間丁 2の経過としてもよい。 また、 検出回数 IV! 2または待ち 時間丁 2を 0とすることにより、 区間 5を最短化することも可能である。
[0072] 図 9は、 本技術の第 1の実施の形態の区間 6における制御の処理手順例を 示す図である。
[0073] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から供給された姿勢情 報に基づいて音声信号処理装置 1 0 0の姿勢を解析し、 姿勢変動量を検出す る (ステップ3 9 6 1) 。
[0074] 制御部 1 5 0は、 姿勢変動量を閾値 1 と比較する (ステップ 3 9 6 2)
。 姿勢変動量が閾値 1 より小さい場合には (ステップ 3 9 6 2 : N 0) 、 姿勢変動が無いものとする。 一方、 姿勢変動量が閾値 1 より大きい場合に は (ステップ 3 9 6 2 : V 6 3) , 姿勢変動が有るものとする。
[0075] 変動量が有る状態においては、 ステップ 3 9 6 1以降の処理を繰り返し ( ステップ 3 9 6 8 : 1\1〇) 、 変動量が有る状態を連続して 2回検出すると (ステップ 3 9 6 8 : 丫 6 3) 、 区間 2に遷移する (ステップ 3 9 6 9) 。
[0076] 姿勢変動が無い状態においては、 音声信号調整部 1 6 0が、 音声信号に対
〇 2020/174776 13 卩(:170?2019/045255
してゲインを掛けることによって音声信号の調整を行う (ステップ 3 9 6 3 ) 。 具体例としては、 変動量が小さいと判定されるたびに (ステップ 3 9 6 2 : N 0) 、 現在のゲインから既定のステップサイズ 3 2分を上げ、 このス テップサイズ 3 2を 「初期値◦ 1 -目標値 0 2」 とすることにより、 区間 6 をスキップすることも可能である。
[0077] 音声信号調整処理の完了判定として、 例えばゲインが目標値◦ 1 まで上が ったか否かを判定する (ステップ 3 9 6 4) 。 目標値◦ 1 に到達するまでは (ステップ 3 9 6 4 : N 0) ステップ 3 9 6 1以降を繰り返し、 目標値◦ 1 に到達すると (ステップ 3 9 6 4 : 丫 6 3) 区間 1 に遷移する (ステップ 3 9 6 5) 〇
[0078] 図 1 0は、 本技術の第 1の実施の形態における姿勢変動解析処理の処理手 順例を示す図である。
[0079] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から加速度または角速 度などの姿勢情報を取得する (ステップ 3 8 1 1) 。 そして、 姿勢変動解析 部 1 3 0は、 新たに取得した姿勢情報と、 前回保持したデータとの差分を計 算し (ステップ 3 8 1 2) 、 その差分値から姿勢変動量を計算する (ステッ プ3 8 1 3) 。 そして、 姿勢変動解析部 1 3 0は、 新たに計算した姿勢変動 量を新たなデータとして保持することにより更新する (ステップ 3 8 1 4)
[0080] 図 1 1は、 本技術の第 1の実施の形態における音声信号解析処理の処理手 順例を示す図である。
[0081 ] 音声信号入力部 1 1 0は、 入力された音声信号を八/〇変換する (ステッ プ3 8 2 1) 。 音声信号解析部 1 4 0は、 音声信号入力部 1 1 0から供給さ れた音声データの信号レベルを計測し、 計測区間における最大値を検出して 、 制御部 1 5 0に供給する (ステップ 3 8 2 2) 。
[0082] また、 音声信号解析部 1 4 0は、 信号レベルの最大値を一時的に記憶する (ステップ 3 8 2 3) 。 ここで、 記憶される信号レベルの最大値は、 音声解 析処理を実行するたびに更新される。
〇 2020/174776 14 卩(:170?2019/045255
[0083] 図 1 2は、 本技術の第 1の実施の形態における音声信号調整処理の処理手 順例を示す図である。
[0084] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から供給された姿勢情 報に基づいて音声信号処理装置 1 0 0の姿勢を解析し、 姿勢変動量を検出す る (ステップ3 8 3 1) 。
[0085] 制御部 1 5 0は、 姿勢変動量を閾値 1 と比較する (ステップ 3 8 3 2)
。 姿勢変動量が閾値 1 より小さい場合には (ステップ 3 8 3 2 : N 0) 、 姿勢変動が無いものとする。 一方、 姿勢変動量が閾値 1 より大きい場合に は (ステップ 3 8 3 2 : V 6 3) , 姿勢変動が有るものとする。
[0086] 姿勢変動が有る状態においては、 ゲインが目標値◦ 2に達していなければ (ステップ 3 8 3 3 : 丫 6 3) 、 既定幅 3 1分だけゲインを小さくする (ス テップ3 8 3 4) 。 一方、 変動量が無い状態においては、 ゲインが初期値◦ 1 に達していなければ (ステップ 3 8 3 5 : 丫6 3) 、 既定幅 3 2分だけゲ インを大きくする (ステップ 3 8 3 6) 。
[0087] 音声信号調整部 1 6 0は、 音声信号入力部 1 1 0から供給された音声信号 に対して、 制御部 1 5 0から供給された更新されたゲインを掛けることによ って音声信号の調整を行う (ステップ 3 8 3 7) 。
[0088] このように、 本技術の第 1の実施の形態では、 音声信号解析部 1 4 0によ って設定された目標値に向けて、 姿勢変動解析部 1 3 0によって生成された 姿勢変動量に応じて、 音声信号調整部 1 6 0によって音声信号が調整される 。 これにより、 周囲の音声の状況に合わせて、 姿勢変動量に応じた音声信号 の調整を行うことができる。 例えば、 人体や機体に装着した状態で取得した 音声信号に含まれる、 姿勢変動や振動による雑音を抑制することができる。 また、 姿勢変動に対して段階的に雑音抑制を効かせることにより、 再生音声 に違和感が出ないようにすることができる。
[0089] < 2 . 第 2の実施の形態 >
上述の第 1の実施の形態では、 姿勢変動量に応じて音声信号のレベルを調 整していた。 これに対し、 この第 2の実施の形態では、 音声信号の周波数ス
〇 2020/174776 1 5 卩(:170?2019/045255
ぺクトルを調整する。
[0090] [周波数スぺクトル]
図 1 3は、 本技術の第 2の実施の形態における音声信号の周波数スペクト ルの調整例を示す図である。
[0091 ] ここでは、 音声信号について横軸に周波数、 縦軸にゲインを示している。
平常時であれば、 音声信号は実線に示したように一定の帯域幅および信号強 度の分布を有する。 これに対し、 雑音が発生した場合、 平常時の帯域幅以外 の周波数において雑音スペクトルが分布し、 または、 帯域幅内であっても強 い強度の信号として雑音スぺクトルが現れることがある。
[0092] そのため、 この第 2の実施の形態では、 音声信号解析部 1 4 0において音 声信号のスペクトル解析を行うとともに、 音声信号調整部 1 6 0にフィルタ を設けてスぺクトル解析の結果に応じてフィルタ特性を切り替える。
[0093] 具体的には、 区間 3においては、 姿勢変動が有る状態において下向き矢印 で示すようなフィルタ特性を段階的に適用して、 フィルタ特性を低ゲインか つ狭帯域に切り替える。 一方、 区間 6においては、 姿勢変動が無い状態にお いて上向き矢印で示すようなフィルタ特性を段階的に適用して、 フィルタ特 性を高ゲインかつ広帯域に切り替え、 最終的にフィルタをかけない信号を出 力する。
[0094] なお、 この第 2の実施の形態の全体構成については、 上述の第 1の実施の 形態と同様であるため、 詳細な説明は省略する。
[0095] [動作]
図 1 4は、 本技術の第 2の実施の形態における音声信号解析処理の処理手 順例を示す図である。
[0096] 音声信号入力部 1 1 0は、 入力された音声信号を八/〇変換する (ステツ プ3 8 4 1) 。 音声信号解析部 1 4 0は、 音声信号入力部 1 1 0から供給さ れた音声データの周波数スぺクトルを解析し、 解析区間における周波数スぺ クトルの帯域幅および信号強度の最大値を検出して、 制御部 1 5 0に供給す る (ステツプ3 8 4 2) 。
〇 2020/174776 16 卩(:170?2019/045255
[0097] また、 音声信号解析部 1 4 0は、 周波数スぺクトルの帯域幅および信号強 度を一時的に記憶する (ステップ 3 8 4 3) 。 ここで、 記憶される周波数ス ぺクトルの帯域幅および信号強度の最大値は、 音声解析処理を実行するたび に更新される。
[0098] 図 1 5は、 本技術の第 2の実施の形態における音声信号調整処理の処理手 順例を示す図である。
[0099] 姿勢変動解析部 1 3 0は、 センサ信号入力部 1 2 0から供給された姿勢情 報に基づいて音声信号処理装置 1 0 0の姿勢を解析し、 姿勢変動量を検出す る (ステップ3 8 5 1) 。
[0100] 制御部 1 5 0は、 姿勢変動量を閾値と比較する (ステップ 3 8 5 2) 。 姿 勢変動量が閾値より小さい場合には (ステップ 3 8 5 2 : N 0) 、 姿勢変動 が無いものとする。 一方、 姿勢変動量が閾値より大きい場合には (ステップ 3 8 5 2 : V 6 3) , 姿勢変動が有るものとする。
[0101 ] 姿勢変動が有る状態においては、 ゲインが目標値に達していなければ (ス テップ3 8 5 3 : 丫6 3) 、 既定幅分だけゲインを小さくする (ステップ 3 8 5 4) 。 また、 現在の帯域が目標値に達していなければ (ステップ 3 8 6
既定幅分だけ帯域を狭くする (ステップ 3 8 6 4) 。
[0102] 一方、 変動量が無い状態においては、 ゲインが初期値に達していなければ (ステップ 3 8 5 5 : 丫6 3) 、 既定幅分だけゲインを大きくする (ステッ プ3 8 5 6) 。 また、 現在の帯域が初期値に達していなければ (ステップ 3 8 6 5 : 丫6 3) 、 既定幅分だけ帯域を広くする (ステップ 3 8 6 6) 。
[0103] 音声信号調整部 1 6 0は、 音声信号入力部 1 1 0から供給された音声信号 に対して、 制御部 1 5 0から供給された特性のフィルタを掛けることによっ て音声信号の調整を行う (ステップ 3 8 6 7) 。
[0104] このように、 本技術の第 2の実施の形態によれば、 姿勢変動解析部 1 3 0 によって生成された姿勢変動量に応じたフィルタ特性により、 音声信号調整 部 1 6 0によって音声信号が調整される。
[0105] < 3 . 第 3の実施の形態 >
〇 2020/174776 17 卩(:170?2019/045255
上述の第 1の実施の形態では、 入力された音声信号をリアルタイムに調整 していた。 これに対し、 この第 3の実施の形態では、 音声信号およびセンサ 信号を記録しておいて、 音声信号の再生時にセンサ信号から姿勢変動を解析 し、 これに基づいて音声信号を調整する。
[0106] [音声信号処理装置の構成]
図 1 6は、 本技術の第 3の実施の形態における音声信号処理装置 1 0 0の 構成例を示す図である。
[0107] この第 3の実施の形態における音声信号処理装置 1 0 0は、 上述の第 1の 実施の形態に加えて、 音声信号記録再生部 1 1 5およびセンサ信号記録再生 部 1 2 5を備える点において異なり、 それ以外の点については上述の第 1の 実施の形態と同様である。
[0108] 音声信号記録再生部 1 1 5は、 音声信号の記録再生を行うものである。 セ ンサ信号記録再生部 1 2 5は、 センサ信号の記録再生を行うものである。 な お、 音声信号記録再生部 1 1 5およびセンサ信号記録再生部 1 2 5は、 特許 請求の範囲に記載の記録再生部の一例である。
[0109] 音声信号記録再生部 1 1 5は、 データ記録時に、 音声信号入力部 1 1 0か らの音声データを記録する。 センサ信号記録再生部 1 2 5は、 データ記録時 に、 センサ信号入力部 1 2 0からの加速度または角速度などの姿勢情報を記 録する。 音声信号記録再生部 1 1 5およびセンサ信号記録再生部 1 2 5は、 互いに同期して記録を行う。
[01 10] 音声信号記録再生部 1 1 5およびセンサ信号記録再生部 1 2 5は、 データ 再生時に、 それぞれ音声信号およびセンサ信号を同期して再生し、 それぞれ 音声信号解析部 1 4 0および姿勢変動解析部 1 3 0に供給する。 以降の処理 は上述の第 1の実施の形態と同様である。
[01 1 1 ] このように、 本技術の第 3の実施の形態によれば、 音声信号記録再生部 1
1 5およびセンサ信号記録再生部 1 2 5に記録された音声信号およびセンサ 信号に基づいて、 再生時に音声信号を調整することができる。
[01 12] < 4 . 第 4の実施の形態 >
〇 2020/174776 18 卩(:170?2019/045255
上述の第 1の実施の形態では音声信号について姿勢変動量に基づいた調整 を行っていたが、 この第 4の実施の形態ではさらに画像信号についても姿勢 変動量に基づいた補正を行う。
[01 13] [音声信号処理装置の構成]
図 1 7は、 本技術の第 4の実施の形態における音声信号処理装置 1 0 0の 構成例を示す図である。
[01 14] この第 4の実施の形態における音声信号処理装置 1 0 0は、 上述の第 1の 実施の形態に加えて、 画像信号入力部 1 8 1、 画像信号補正部 1 8 2および 画像信号出力部 1 8 3を備える点において異なり、 それ以外の点については 上述の第 1の実施の形態と同様である。
[01 15] 画像信号入力部 1 8 1は、 (図示しない) 画像センサから供給される画像 信号を受け取り、 画像フレームとして画像信号補正部 1 8 2に供給するもの である。 画像信号補正部 1 8 2は、 姿勢変動解析部 1 3 0からの姿勢変動量 に基づいて、 画像信号入力部 1 8 1からの画像フレームに台形補正処理を行 うことにより、 ブレ補正を行うものである。 画像信号出力部 1 8 3は、 画像 信号補正部 1 8 2においてブレ補正された画像信号を出力するものである。 これにより、 音声信号の調整と同様に、 画像信号についても姿勢変動量に基 づいたブレ補正が行われる。
[01 16] このように、 本技術の第 4の実施の形態によれば、 音声信号の調整に加え て、 画像信号についても画像信号補正部 1 8 2において姿勢変動量に基づい て補正を行うことができる。 すなわち、 画像の手振れ補正と同時に、 再生す る音声から手振れ時の雑音を抑制することができる。
[01 17] < 5 . 第 5の実施の形態 >
上述の第 1の実施の形態では、 姿勢変動量に応じて音声信号の調整を行う ことにより、 姿勢変動に起因する音声信号の雑音を抑制していた。 この場合 、 姿勢変動によって雑音を発生していないときであっても音声信号のゲイン を低下させるおそれが生じ得る。 そのため、 この第 5の実施の形態では、 姿 勢変動があっても音声が変動していない場合には音声信号の調整を行わない
\¥0 2020/174776 19 卩(:17 2019/045255
ように、 姿勢変動と音声変動との相関性が高い場合にのみ音声信号を調整す る。
[01 18] [音声信号処理装置の構成]
図 1 8は、 本技術の第 5の実施の形態における音声信号処理装置 1 0 0の 構成例を示す図である。
[01 19] この第 5の実施の形態における音声信号処理装置 1 0 0では、 音声信号解 析部 1 4 0が音声信号入力部 1 1 0からの音声信号と姿勢変動解析部 1 3 0 からの姿勢変動量との相関値を計算する。 相関値の計算においては、 姿勢変 動量が大きくかつ区間 1の音声レベルより大きい場合に相関値が高くなり、 姿勢変動量が大きくても音声レベルが小さい場合には相関値は低くなる。
[0120] 制御部 1 5 0は、 音声信号解析部 1 4 0からの相関値に基づいて、 相関値 が高い場合にのみ音声信号の調整処理を行うように音声信号調整部 1 6 0を 制御する。
[0121 ] [動作]
図 1 9は、 本技術の第 5の実施の形態における音声信号解析処理の処理手 順例を示す図である。
[0122] この第 5の実施の形態における音声信号解析処理では、 上述の第 1の実施 の形態における処理に加えて、 音声信号解析部 1 4 0が音声信号入力部 1 1 0からの音声信号と姿勢変動解析部 1 3 0からの姿勢変動量との相関値を計 算する (ステップ 3 8 2 4) 。 この計算された相関値は、 制御部 1 5 0に供 給される。
[0123] 図 2 0は、 本技術の第 5の実施の形態の区間 1 における制御の処理手順例 を示す図である。
[0124] この第 5の実施の形態の区間 1 における制御では、 上述の第 1の実施の形 態における処理に加えて、 姿勢変動が有った際に、 制御部 1 5 0が音声信号 解析部 1 4 0からの相関値に基づいて、 姿勢変動が音声変動と相関している か否かを判断する (ステップ 3 9 1 5) 。 相関していないと判断した場合に は (ステップ 3 9 1 5 : N 0) 、 その姿勢変動はなかったものとしてステッ
〇 2020/174776 20 卩(:170?2019/045255
プ3 9 1 1以降の処理を繰り返す。 相関していると判断した場合には (ステ ップ 3 9 1 5 : 丫6 3) 、 その姿勢変動が有意なものとして区間 2に遷移す る (ステップ3 9 1 6) 。
[0125] 図 2 1は、 本技術の第 5の実施の形態の区間 2における制御の処理手順例 を示す図である。
[0126] この第 5の実施の形態の区間 2における制御では、 上述の第 1の実施の形 態における処理に加えて、 姿勢変動が有った際に、 制御部 1 5 0が音声信号 解析部 1 4 0からの相関値に基づいて、 姿勢変動が音声変動と相関している か否かを判断する (ステップ 3 9 2 5) 。 相関していないと判断した場合に は (ステップ 3 9 2 5 : N 0) 、 その姿勢変動はなかったものとしてステッ プ3 9 2 9以降の処理を行う。 相関していると判断した場合には (ステップ 3 9 2 5 : 丫 6 3) 、 その姿勢変動が有意なものとしてステップ 3 9 2 6以 降の処理を行う。
[0127] 図 2 2は、 本技術の第 5の実施の形態の区間 3における制御の処理手順例 を示す図である。
[0128] この第 5の実施の形態の区間 3における制御では、 上述の第 1の実施の形 態における処理に加えて、 姿勢変動が有った際に音声信号解析部 1 4 0が音 声信号解析処理を行う (ステップ 3 9 3 3) 。 そして、 制御部 1 5 0が音声 信号解析部 1 4 0からの相関値に基づいて、 姿勢変動が音声変動と相関して いるか否かを判断する (ステップ 3 9 3 4) 。 相関していないと判断した場 合には (ステップ 3 9 3 4 : N 0) 、 その姿勢変動はなかったものとしてス テップ3 9 3 1以降の処理を繰り返す。 相関していると判断した場合には ( ステップ 3 9 3 4 :
3) , その姿勢変動が有意なものとして音声信号調 整部 1 6 0に音声信号調整処理を行うよう制御する (ステップ 3 9 3 5) 。
[0129] 図 2 3は、 本技術の第 5の実施の形態の区間 4における制御の処理手順例 を示す図である。
[0130] この第 5の実施の形態の区間 4における制御では、 上述の第 1の実施の形 態における処理に加えて、 姿勢変動が有った際に音声信号解析部 1 4 0が音
〇 2020/174776 21 卩(:170?2019/045255
声信号解析処理を行う (ステップ 3 9 4 3) 。 そして、 制御部 1 5 0が音声 信号解析部 1 4 0からの相関値に基づいて、 姿勢変動が音声変動と相関して いるか否かを判断する (ステップ 3 9 4 4) 。 相関していないと判断した場 合には (ステップ 3 9 4 4 : N 0) 、 その姿勢変動はなかったものとしてス テップ3 9 4 5以降の処理を行う。 相関していると判断した場合には (ステ ップ 3 9 4 4 : 丫 6 3) 、 その姿勢変動が有意なものとしてステップ 3 9 4 1以降の処理を行う。
[0131 ] 図 2 4は、 本技術の第 5の実施の形態の区間 5における制御の処理手順例 を示す図である。
[0132] この第 5の実施の形態の区間 5における制御では、 上述の第 1の実施の形 態における処理に加えて、 姿勢変動が有った際に音声信号解析部 1 4 0が音 声信号解析処理を行う (ステップ 3 9 5 5) 。 そして、 制御部 1 5 0が音声 信号解析部 1 4 0からの相関値に基づいて、 姿勢変動が音声変動と相関して いるか否かを判断する (ステップ 3 9 5 6) 。 相関していないと判断した場 合には (ステップ 3 9 5 6 : N 0) 、 その姿勢変動はなかったものとしてス テップ3 9 5 3以降の処理を繰り返す。 相関していると判断した場合には ( ステップ 3 9 5 6 : 丫6 3) 、 その姿勢変動が有意なものとしてステップ 3 9 5 7以降の処理を行う。
[0133] 図 2 5は、 本技術の第 5の実施の形態の区間 6における制御の処理手順例 を示す図である。
[0134] この第 5の実施の形態の区間 6における制御では、 上述の第 1の実施の形 態における処理に加えて、 姿勢変動が有った際に音声信号解析部 1 4 0が音 声信号解析処理を行う (ステップ 3 9 6 6) 。 そして、 制御部 1 5 0が音声 信号解析部 1 4 0からの相関値に基づいて、 姿勢変動が音声変動と相関して いるか否かを判断する (ステップ 3 9 6 7) 。 相関していないと判断した場 合には (ステップ 3 9 6 7 : N 0) 、 その姿勢変動はなかったものとしてス テップ3 9 6 1以降の処理を繰り返す。 相関していると判断した場合には ( ステップ 3 9 6 7 : 丫6 3) 、 その姿勢変動が有意なものとしてステップ 3
〇 2020/174776 22 卩(:170?2019/045255
968以降の処理を行う。
[0135] このように、 本技術の第 5の実施の形態によれば、 姿勢変動と音声変動と の相関性が高い場合にのみ音声信号を調整することにより、 姿勢変動があっ ても音声が変動していない場合には音声信号の調整を行わないようにするこ とができる。
[0136] なお、 上述の実施の形態は本技術を具現化するための一例を示したもので あり、 実施の形態における事項と、 特許請求の範囲における発明特定事項と はそれぞれ対応関係を有する。 同様に、 特許請求の範囲における発明特定事 項と、 これと同一名称を付した本技術の実施の形態における事項とはそれぞ れ対応関係を有する。 ただし、 本技術は実施の形態に限定されるものではな く、 その要旨を逸脱しない範囲において実施の形態に種々の変形を施すこと により具現化することができる。
[0137] また、 上述の実施の形態において説明した処理手順は、 これら一連の手順 を有する方法として捉えてもよく、 また、 これら一連の手順をコンピュータ に実行させるためのプログラム乃至そのプログラムを記憶する記録媒体とし て捉えてもよい。 この記録媒体として、 例えば、 CD (Compact Disc) 、 M D (MiniDisc) 、 DVD (Digital Versa† i le Disc) , メモリカード、 ブル —レイディスク (Blu-ray (登録商標) Disc) 等を用いることができる。
[0138] なお、 本明細書に記載された効果はあくまで例示であって、 限定されるも のではなく、 また、 他の効果があってもよい。
[0139] なお、 本技術は以下のような構成もとることができる。
(1 ) 音声信号を取得して前記音声信号に基づいて音声調整のための目標値 を設定する音声信号解析部と、
姿勢情報を取得して前記姿勢情報に基づいて姿勢変動量を生成する姿勢変 動解析部と、
前記姿勢変動量に応じて前記目標値に向けて前記音声信号を調整する音声 信号調整部と
を具備する音声信号処理装置。
20/174776 23 卩(:170?2019/045255
(2) 前記音声信号調整部は、 前記姿勢変動量が第 1の閾値より大きくなる と前記音声信号の音量を下げるように前記音声信号を調整し、 前記目標値に 達した後に前記姿勢変動量が第 2の閾値より小さくなると前記音声信号の音 量を戻すように前記音声信号を調整する
前記 ( 1) に記載の音声信号処理装置。
(3) 前記音声信号調整部は、 前記姿勢変動量が第 1の閾値より大きくなる と前記音声信号の周波数の帯域幅を狭くするように前記音声信号を調整し、 前記目標値に達した後に前記姿勢変動量が第 2の閾値より小さくなると前記 音声信号の周波数の帯域幅を戻すように前記音声信号を調整する
前記 (1) または (2) に記載の音声信号処理装置。
(4) 前記音声信号調整部は、 前記姿勢変動量が第 1の閾値より大きくなる と前記音声信号の周波数のゲインを下げるように前記音声信号を調整し、 前 記目標値に達した後に前記姿勢変動量が第 2の閾値より小さくなると前記音 声信号の周波数のゲインを戻すように前記音声信号を調整する
前記 (1) から (3) のいずれかに記載の音声信号処理装置。
(5) 前記音声信号調整部は、 前記姿勢変動量が第 1の閾値より大きい状態 、 および、 前記姿勢変動量が第 2の閾値より小さい状態の、 何れか一方の状 態が所定期間継続した場合に前記音声信号を調整する
前記 (1) から (4) のいずれかに記載の音声信号処理装置。
(6) 前記音声信号調整部は、 _定量のステップサイズで段階的に前記音声 信号を調整する
前記 (1) から (5) のいずれかに記載の音声信号処理装置。
(7) 加速度または角速度を検出して前記姿勢情報を生成するセンサをさら に具備する前記 (1) から (6) のいずれかに記載の音声信号処理装置。
(8) 前記音声信号および前記姿勢情報を同期して記録および再生する記録 再生部をさらに具備し、
前記音声信号解析部は、 前記再生された音声信号に基づいて前記目標値を 設¾:し、
〇 2020/174776 24 卩(:170?2019/045255
前記姿勢変動解析部は、 前記再生された姿勢情報に基づいて姿勢変動量を 生成し、
前記音声信号調整部は、 前記再生された姿勢変動量に応じて前記再生され た音声信号を調整する
前記 (1) から (7) のいずれかに記載の音声信号処理装置。
(9) 前記音声信号に同期する画像信号を取得して前記姿勢変動量に応じて 前記画像信号のブレを補正する画像信号補正部をさらに具備する
前記 (1) から (8) のいずれかに記載の音声信号処理装置。
(1 0) 前記音声信号調整部は、 前記姿勢変動量が示す姿勢変動と前記音声 信号の変動との間に相関が有る場合に前記音声信号を調整する
前記 (1) から (9) のいずれかに記載の音声信号処理装置。
符号の説明
[0140] 1 00 音声信号処理装置
1 1 0 音声信号入力部
1 1 5 音声信号記録再生部
1 20 センサ信号入力部
1 25 センサ信号記録再生部
1 30 姿勢変動解析部
1 40 音声信号解析部
1 50 制御部
1 60 音声信号調整部
1 70 音声信号出力部
1 81 画像信号入力部
1 82 画像信号補正部
1 83 画像信号出力部
Claims
[請求項 1 ] 音声信号を取得して前記音声信号に基づいて音声調整のための目標 値を設定する音声信号解析部と、
姿勢情報を取得して前記姿勢情報に基づいて姿勢変動量を生成する 姿勢変動解析部と、
前記姿勢変動量に応じて前記目標値に向けて前記音声信号を調整す る音声信号調整部と
を具備する音声信号処理装置。
[請求項 2] 前記音声信号調整部は、 前記姿勢変動量が第 1の閾値より大きくな ると前記音声信号の音量を下げるように前記音声信号を調整し、 前記 目標値に達した後に前記姿勢変動量が第 2の閾値より小さくなると前 記音声信号の音量を戻すように前記音声信号を調整する
請求項 1記載の音声信号処理装置。
[請求項 3] 前記音声信号調整部は、 前記姿勢変動量が第 1の閾値より大きくな ると前記音声信号の周波数の帯域幅を狭くするように前記音声信号を 調整し、 前記目標値に達した後に前記姿勢変動量が第 2の閾値より小 さくなると前記音声信号の周波数の帯域幅を戻すように前記音声信号 を調整する
請求項 1記載の音声信号処理装置。
[請求項 4] 前記音声信号調整部は、 前記姿勢変動量が第 1の閾値より大きくな ると前記音声信号の周波数のゲインを下げるように前記音声信号を調 整し、 前記目標値に達した後に前記姿勢変動量が第 2の閾値より小さ くなると前記音声信号の周波数のゲインを戻すように前記音声信号を 調整する
請求項 1記載の音声信号処理装置。
[請求項 5] 前記音声信号調整部は、 前記姿勢変動量が第 1の閾値より大きい状 態、 および、 前記姿勢変動量が第 2の閾値より小さい状態の、 何れか _方の状態が所定期間継続した場合に前記音声信号を調整する
〇 2020/174776 26 卩(:170?2019/045255
請求項 1記載の音声信号処理装置。
[請求項 6] 前記音声信号調整部は、 _定量のステツプサイズで段階的に前記音 声信号を調整する
請求項 1記載の音声信号処理装置。
[請求項 7] 加速度または角速度を検出して前記姿勢情報を生成するセンサをさ らに具備する請求項 1記載の音声信号処理装置。
[請求項 8] 前記音声信号および前記姿勢情報を同期して記録および再生する記 録再生部をさらに具備し、
前記音声信号解析部は、 前記再生された音声信号に基づいて前記目 標値を設定し、
前記姿勢変動解析部は、 前記再生された姿勢情報に基づいて姿勢変 動量を生成し、
前記音声信号調整部は、 前記再生された姿勢変動量に応じて前記再 生された音声信号を調整する
請求項 1記載の音声信号処理装置。
[請求項 9] 前記音声信号に同期する画像信号を取得して前記姿勢変動量に応じ て前記画像信号のブレを補正する画像信号補正部をさらに具備する 請求項 1記載の音声信号処理装置。
[請求項 10] 前記音声信号調整部は、 前記姿勢変動量が示す姿勢変動と前記音声 信号の変動との間に相関が有る場合に前記音声信号を調整する 請求項 1記載の音声信号処理装置。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/431,690 US12008288B2 (en) | 2019-02-25 | 2019-11-19 | Audio signal processing device based on orientation |
| CN201980092372.2A CN113491135A (zh) | 2019-02-25 | 2019-11-19 | 音频信号处理装置 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2019031535A JP2020137044A (ja) | 2019-02-25 | 2019-02-25 | 音声信号処理装置 |
| JP2019-031535 | 2019-02-25 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020174776A1 true WO2020174776A1 (ja) | 2020-09-03 |
Family
ID=72238584
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/045255 Ceased WO2020174776A1 (ja) | 2019-02-25 | 2019-11-19 | 音声信号処理装置 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US12008288B2 (ja) |
| JP (1) | JP2020137044A (ja) |
| CN (1) | CN113491135A (ja) |
| WO (1) | WO2020174776A1 (ja) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2020137044A (ja) * | 2019-02-25 | 2020-08-31 | ソニーセミコンダクタソリューションズ株式会社 | 音声信号処理装置 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2017092818A (ja) * | 2015-11-13 | 2017-05-25 | 日本放送協会 | ラウドネス調節装置およびプログラム |
| JP2018148254A (ja) * | 2017-03-01 | 2018-09-20 | 大和ハウス工業株式会社 | インターフェースユニット |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP5812663B2 (ja) | 2011-04-22 | 2015-11-17 | 任天堂株式会社 | 音楽演奏用プログラム、音楽演奏装置、音楽演奏システムおよび音楽演奏方法 |
| WO2014167384A1 (en) * | 2013-04-10 | 2014-10-16 | Nokia Corporation | Audio recording and playback apparatus |
| JP6759680B2 (ja) | 2016-04-26 | 2020-09-23 | ソニー株式会社 | 画像処理装置、撮像装置、画像処理方法、および、プログラム |
| GB2554447A (en) * | 2016-09-28 | 2018-04-04 | Nokia Technologies Oy | Gain control in spatial audio systems |
| CA3230221A1 (en) * | 2017-10-12 | 2019-04-18 | Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E.V. | Optimizing audio delivery for virtual reality applications |
| JP2020137044A (ja) * | 2019-02-25 | 2020-08-31 | ソニーセミコンダクタソリューションズ株式会社 | 音声信号処理装置 |
| US11340861B2 (en) * | 2020-06-09 | 2022-05-24 | Facebook Technologies, Llc | Systems, devices, and methods of manipulating audio data based on microphone orientation |
| US11558707B2 (en) * | 2020-06-29 | 2023-01-17 | Qualcomm Incorporated | Sound field adjustment |
-
2019
- 2019-02-25 JP JP2019031535A patent/JP2020137044A/ja active Pending
- 2019-11-19 CN CN201980092372.2A patent/CN113491135A/zh not_active Withdrawn
- 2019-11-19 WO PCT/JP2019/045255 patent/WO2020174776A1/ja not_active Ceased
- 2019-11-19 US US17/431,690 patent/US12008288B2/en active Active
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2017092818A (ja) * | 2015-11-13 | 2017-05-25 | 日本放送協会 | ラウドネス調節装置およびプログラム |
| JP2018148254A (ja) * | 2017-03-01 | 2018-09-20 | 大和ハウス工業株式会社 | インターフェースユニット |
Also Published As
| Publication number | Publication date |
|---|---|
| US12008288B2 (en) | 2024-06-11 |
| CN113491135A (zh) | 2021-10-08 |
| JP2020137044A (ja) | 2020-08-31 |
| US20220137919A1 (en) | 2022-05-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7188082B2 (ja) | 音響処理装置および方法、並びにプログラム | |
| JP5321263B2 (ja) | 信号処理装置、信号処理方法 | |
| JP2011028061A (ja) | 音声記録装置及び方法、ならびに撮影装置 | |
| US8600078B2 (en) | Audio signal amplitude adjusting device and method | |
| JP2004214707A (ja) | 音響装置および音響特性の変更方法および音響補正用プログラム | |
| US20150271439A1 (en) | Signal processing device, imaging device, and program | |
| WO2020174776A1 (ja) | 音声信号処理装置 | |
| JP2922397B2 (ja) | 車両用音響装置 | |
| JP2008245123A (ja) | 音場補正装置、音場補正方法及び制御プログラム | |
| JP6887315B2 (ja) | 音声処理装置およびその制御方法、プログラム並びに記憶媒体 | |
| JP5762797B2 (ja) | 信号処理装置及び信号処理方法 | |
| JP5202342B2 (ja) | 信号処理装置、方法及びプログラム | |
| JP4269892B2 (ja) | オーディオデータ処理方法及び装置 | |
| JP6213701B1 (ja) | 音響信号処理装置 | |
| JP5030250B2 (ja) | 電子機器及びその制御方法 | |
| US10425731B2 (en) | Audio processing apparatus, audio processing method, and program | |
| JP2009010824A (ja) | 音響装置およびスピーカの駆動方法 | |
| WO2006093256A1 (ja) | 音声再生装置及び方法、並びに、コンピュータプログラム | |
| US10313824B2 (en) | Audio processing device for processing audio, audio processing method, and program | |
| JP4328601B2 (ja) | 音声処理装置、編集装置、制御プログラム及び記録媒体 | |
| JP2011124959A (ja) | 音声信号処理装置 | |
| JP2019161334A (ja) | 音声処理装置 | |
| JP2019070688A (ja) | 撮像装置 | |
| JP2004200934A (ja) | スピーカ装置及びスピーカ装置の制御方法 | |
| JP4285507B2 (ja) | オートゲインコントロール回路 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19916979 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19916979 Country of ref document: EP Kind code of ref document: A1 |