WO2010140254A1 - 映像音声出力装置及び音声定位方法 - Google Patents

映像音声出力装置及び音声定位方法 Download PDF

Info

Publication number
WO2010140254A1
WO2010140254A1 PCT/JP2009/060362 JP2009060362W WO2010140254A1 WO 2010140254 A1 WO2010140254 A1 WO 2010140254A1 JP 2009060362 W JP2009060362 W JP 2009060362W WO 2010140254 A1 WO2010140254 A1 WO 2010140254A1
Authority
WO
WIPO (PCT)
Prior art keywords
attribute
voice
face
feature
video
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2009/060362
Other languages
English (en)
French (fr)
Inventor
和実 菅谷
洋人 河内
禎司 鈴木
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Pioneer Corp
Original Assignee
Pioneer Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Pioneer Corp filed Critical Pioneer Corp
Priority to PCT/JP2009/060362 priority Critical patent/WO2010140254A1/ja
Publication of WO2010140254A1 publication Critical patent/WO2010140254A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/06Decision making techniques; Pattern matching strategies
    • G10L17/10Multimodal systems, i.e. based on the integration of multiple recognition engines or fusion of expert systems
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/441Acquiring end-user identification, e.g. using personal code sent by the remote control or by inserting a card
    • H04N21/4415Acquiring end-user identification, e.g. using personal code sent by the remote control or by inserting a card using biometric characteristics of the user, e.g. by voice recognition or fingerprint scanning
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N5/00Details of television systems
    • H04N5/44Receiver circuitry for the reception of television signals according to analogue transmission standards
    • H04N5/60Receiver circuitry for the reception of television signals according to analogue transmission standards for the sound signals
    • H04N5/602Receiver circuitry for the reception of television signals according to analogue transmission standards for the sound signals for digital sound signals
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2400/00Details of stereophonic systems covered by H04S but not provided for in its groups
    • H04S2400/11Positioning of individual sound objects, e.g. moving airplane, within a sound field
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field

Definitions

  • the present invention relates to an audio localization technology for a video / audio output device that outputs content data including video and audio, and more particularly, to an audio localization technology for performing audio localization according to a speaker position.
  • a human voice When receiving program content such as TV broadcast, displaying video on a display and outputting sound from a speaker, in monaural sound, a human voice can be heard from the position of the speaker. In stereo / surround sound, in many cases, a human voice is localized at the center of the screen so that the human voice can be heard from the center of the screen.
  • Patent Document 1 the position of a speaker is detected, and the volume of sound output from a plurality of speakers is controlled according to the detected position.
  • the present invention has been made in view of the above circumstances, and an example of the problem is to provide an audio localization technique that takes into account the case where the speaker's face is not on the screen.
  • face feature extraction is performed to analyze an input video, detect one or a plurality of face positions, and extract face features of each detected face position.
  • face attribute determination means for determining the attribute of each face feature extracted by the face feature extraction means as a first attribute
  • voice feature extraction means for analyzing the input voice and extracting a voice feature
  • Voice attribute determination means for determining the attribute of the voice feature extracted by the voice feature extraction means as a first attribute, first attributes of each face feature determined by the face attribute determination means, and determination by the voice attribute determination means
  • the degree of fitness is highest when it is determined that the speaker's face is in the screen from the fitness level determination means for determining the fitness level of the first attribute of the voice feature and the determination result of the fitness level determination means.
  • Voice localization hand that localizes voice to speaker's face position When a video-audio output apparatus comprising a.
  • an input image is analyzed, the input image is analyzed, one or a plurality of face positions are detected, and a face feature is extracted for each detected face position.
  • Each of the facial features extracted in the step and the facial feature extraction step is compared with the facial feature information stored in the facial feature database storing facial feature information related to the facial features classified by attribute.
  • a face attribute determining step for determining an attribute of the face feature a voice feature extracting step for analyzing the input voice and extracting a voice feature, a voice feature extracted in the voice feature extracting step, and a voice by attribute
  • a voice attribute determination step for comparing the voice feature information stored in the voice feature database storing the voice feature information about the feature to determine the attribute of the extracted voice feature, and the determination in the face attribute determination step
  • the face of the speaker is displayed on the screen from the attribute of each facial feature, the fitness determination step for determining the fitness of the voice feature attribute determined in the voice attribute determination step, and the determination result of the fitness determination step. If it is determined that there is a voice localization method, the voice localization method includes a voice localization step of localizing the voice to the position of the speaker's face having the highest fitness.
  • FIG. 1 is a schematic configuration diagram of a video / audio output device according to a first embodiment of the present invention. It is an example of the image which the video / audio output device which concerns on the 1st Embodiment of this invention displays. It is an example of the face attribute determination result and the voice attribute determination result by the video / audio output device according to the embodiment of the present invention. It is an example of the face attribute determination result and the voice attribute determination result by the video / audio output device according to the embodiment of the present invention. It is a flowchart which shows the flow of the video / audio output process of the video / audio output device which concerns on the 1st Embodiment of this invention.
  • FIG. 1 is a schematic configuration diagram of a video / audio output device 1 according to an embodiment of the present invention.
  • the video / audio output device 1 is a device that outputs audio with audio localization in accordance with the speaker position. In this embodiment, it is determined whether or not the speaker's face is in the screen. If the speaker's face is in the screen, the voice is output in accordance with the position of the speaker in the screen. The panorama has been changed.
  • “speaker” refers to the person speaking in the video data (on the screen)
  • “speaker position” refers to the position of the speaker on the screen. The position near the speaker's face. “Output the voice with the voice localization in accordance with the speaker position” means outputting the voice so that the voice can be heard from the position of the speaker. For example, as shown in FIG. Is present on the left side of the screen, the volume of the speaker voice output from the speaker SP1 provided on the left side of the screen d10 is increased, and the volume of the speaker voice output from the speaker SP2 provided on the right side of the screen d10 is increased.
  • the sound is output so that the sound can be heard from the position of the speaker on the left side of the screen at a reduced volume.
  • the volume ratio of the speaker voices of the speakers SP1 and SP2 may be set to D: C. .
  • the video / audio output device 1 may be any device as long as it has a function of reproducing content data including video and audio input from the outside and outputting the content data to the outside.
  • a television (TV), a DVD player and recorder, a BD player and recorder, a personal computer (PC), and the like are assumed.
  • the video / audio output apparatus 1 includes a face area detection unit 101, a face feature detection unit 102, an attribute-specific face feature database (hereinafter referred to as attribute-specific face feature DB) 103, a face attribute determination unit 104, and voice feature detection.
  • Unit 105 attribute-specific voice feature database (hereinafter referred to as attribute-specific voice feature DB) 106, voice attribute determination unit 107, conformity determination unit 108, audio localization processing unit 109, video display unit 110, and audio output unit 111.
  • attribute-specific face feature database hereinafter referred to as attribute-specific voice feature DB
  • the face area detection unit 101 detects a human face area from the input video data.
  • the detection of the face area is performed using a known technique. For example, there are a technique for detecting a face area by searching for a skin color area, a technique for detecting a face area by a template matching method for detecting a face area while comparing templates (face pattern images) on images.
  • the face area detection unit 101 outputs the detected face area position information (face area position information) to the face feature detection unit 102.
  • face area position information face area position information
  • each face area position information is output to the face feature detection unit 102, and when no face area is detected, the face position information is output to the face feature detection unit 102.
  • the face feature detection unit 102 detects each human face feature from the input video data based on each face region position information output from the face region detection unit 101. Facial features are detected using a known technique. In this embodiment, a method for detecting the feature amount of a major part of the face is adopted, and the feature amount indicating the positional relationship between the face contour, both eyebrows, both eyes, nose, mouth and the like is used as the facial feature. Detect as. Then, the face feature detection unit 102 outputs the detected face feature to the face attribute determination unit 104.
  • the attribute-specific face feature DB 103 is a database that stores data related to attribute-specific face features.
  • attribute means, for example, sex and age.
  • the data is classified into six attributes of women aged between 49 and 49, and women aged 50 and over, and data relating to facial features is stored for each classified attribute.
  • classification is made into six attributes, but the classification of attributes is not limited to this, and may be further subdivided.
  • the video / audio output apparatus 1 includes the attribute-specific face feature DB 103, but the video / audio output apparatus 1 may not include the attribute-specific face feature DB 103.
  • the video / audio output apparatus 1 may access the attribute-specific face feature DB 103 via a communication network and refer to data stored in the attribute-specific face feature DB 103.
  • the face attribute determination unit 104 compares each face feature output from the face feature detection unit 102 with data related to the face feature by attribute stored in the attribute-specific face feature DB 103, and determines the attribute of each face. It is supposed to be.
  • the determined face attribute result (referred to as a face attribute determination result) is data having a probability belonging to each attribute, and in the present embodiment, for example, data as shown in FIG. Then, the face attribute determination unit 104 outputs the face attribute determination result to the suitability determination unit 108.
  • a method for determining attributes (age, gender) from face features in the face feature detection unit 102, the attribute-specific face feature DB 103, and the face attribute determination unit 104 is performed using a known technique. For example, according to the method described in IEICE, IEICE, PRMU 2001-138 (2001) “Proposal of Gender / Age Estimation Method Using Distance from Average Face”, the average by detected face and attribute A matching attribute may be determined based on the feature point distance with the face.
  • the voice feature detecting unit 105 detects a human voice feature from the input voice data. The detection of the voice feature is performed using a known technique. In this embodiment, the voice frequency is analyzed, the formant frequency and pitch frequency are specified, and the specified formant frequency and pitch frequency are detected as voice characteristics. Then, the voice feature detection unit 105 outputs the detected voice feature to the voice attribute determination unit 107.
  • the attribute-specific voice feature DB 106 is a database that stores data related to voice features by attribute.
  • attribute means, for example, sex and age.
  • attribute means, for example, males under 20 years old, males 20 years old and younger than 49 years old, males 50 years old and older, females under 20 years old, 20
  • the data is classified into six attributes of a woman aged between 49 and 49 and a woman aged 50 and over, and data relating to voice characteristics is stored for each classified attribute.
  • the above six attributes are classified, but the attribute classification is not limited to this.
  • the video / audio output device 1 includes the attribute-specific voice feature DB 103, but the video / audio output device 1 may not include the attribute-specific voice feature DB 106.
  • the video / audio output device 1 may access the attribute-specific voice feature DB 106 via a communication network and refer to data stored in the attribute-specific voice feature DB 106.
  • the voice attribute determination unit 107 determines the voice attribute by comparing the voice feature output from the voice feature detection unit 105 with the data regarding the voice feature by attribute stored in the attribute-specific voice feature DB 106. ing.
  • the determined voice attribute result (referred to as a voice attribute determination result) is data having a probability belonging to each attribute, and in the present embodiment, for example, data as shown in FIG. Then, the voice attribute determination unit 107 outputs the voice attribute determination result to the suitability determination unit 108.
  • a method for determining attributes (age, sex) from voice features in the voice feature detection unit 105, the attribute-specific voice feature DB 106, and the voice attribute determination unit 107 is performed using a known technique. For example, according to the method described in Journal of the Acoustical Society of Japan Vol. 24, No. 6 (1968) “Changes in Pitch Frequency and Formant Frequency of Japanese 5 Vowels by Age and Gender” You may make it discriminate
  • the suitability determination unit 108 determines the degree of matching between the face and the voice. It has become. More specifically, an adaptation determination process is performed to determine whether or not the speaker's face is on the screen, and if the speaker's face is on the screen, which face is the speaker's face.
  • the match determination unit 108 outputs speaker information, which is a result of the match determination process, to the sound localization processing unit 109.
  • FIG. 3 shows each face attribute determination result and face attribute determination result when three face features are detected from the screen.
  • FIG. 4 shows each face attribute detection result when two face features are detected.
  • the face attribute determination result and the face attribute determination result are shown.
  • the matching degree between the face and the voice is calculated by multiplying values belonging to the same attribute for all the attributes and adding the respective multiplied values.
  • the threshold value T1 for determining whether or not the speaker's face is in the screen is 0.5 as an example, and when the threshold T1 is greater than this value, it is determined that the speaker's face is in the screen. Since the large value is A1 and A1 ⁇ T1, it is determined that there is a speaker in the screen, and the speaker is determined to be a person having face 1. In this case, position information indicating the position of the face 1 is output to the speech localization processing unit 109 as speaker information.
  • the value having the highest fitness is B1, but since B1 ⁇ T1, it is determined that there is no speaker in the screen.
  • information indicating that there is no speaker on the screen is output to the speech localization processing unit 109 as speaker information.
  • the voice localization processing unit 109 performs a localization change process of the input voice data based on the speaker position information output from the matching determination unit 108. That is, the volume is adjusted so that the sound is localized at the speaker position on the screen. For example, as shown in FIG. 2, when the speaker A exists on the left side of the screen, the volume of the sound output from the speaker SP1 provided on the left side of the screen d10 is increased and provided on the right side of the screen d10. The sound volume output from the speaker SP2 is reduced.
  • the sound localization processing unit 109 outputs the sound subjected to the localization change process to the sound output unit 111.
  • the video display unit 110 outputs and displays the input video data on a display or the like. Note that the video data to be output and displayed is delayed as necessary in order to synchronize with the audio data to be output.
  • the audio output unit 111 outputs audio data that has been subjected to the localization change process to a speaker.
  • FIG. 5 is a flowchart showing the flow of the video / audio output process of the video / audio output device 1.
  • the video / audio output device 1 analyzes the input video data and performs face attribute determination processing for determining the face attribute (step S10). Specifically, first, the face area detection unit 101 detects a human face area from the input video data, and then the face feature detection unit 102 detects the human face from the input video data based on the detected face area. The feature is detected, and finally, the face attribute determination unit 104 determines the detected face attribute by comparing the detected face feature with the data of the attribute-specific face feature DB 103.
  • the video / audio output device 1 analyzes the input audio data and performs a voice attribute determination process for determining a voice attribute (step S20). Specifically, first, a voice feature of a person is detected from the voice data input by the voice feature detector 105, and then detected by comparing the voice feature detected by the voice attribute determination unit 107 with the data of the voice feature DB 106 by attribute. Determine voice attributes.
  • the suitability determination unit 108 of the video / audio output device 1 performs a suitability determination process for determining the suitability of the face and the voice based on the face attribute determination result and the voice attribute determination result (step S30).
  • step S40 If it is determined that the speaker's face is within the screen as a result of the fitness determination process (step S40: YES), the sound localization process is performed to change the sound to the position of the speaker's face determined to be within the screen. Is performed (step S50). For example, the output value of the speaker voice of the speaker close to the speaker position is raised, and the output value of the speaker voice of the speaker far from the speaker position is lowered.
  • step S40 If it is determined that the speaker's face is not on the screen as a result of the fitness determination process (step S40: NO), the voice localization process is not performed.
  • the video display unit 110 of the video / audio output device 1 outputs video data, and the audio output unit 111 performs the audio localization change when the speaker's face is on the screen.
  • the audio data that has not been subjected to the audio localization change is output (step S60).
  • the face of the speaker is determined. It can be determined whether or not is in the screen. As a result of the determination, if the speaker's face is on the screen, the sound is panned to the specified speaker position, whereas if the speaker's face is not on the screen, the voice panning is changed. Since this is not performed, the viewer does not feel uncomfortable and more natural and realistic viewing is possible.
  • FIG. 6 is a schematic configuration diagram of the video / audio output device 2 according to the second embodiment of the present invention.
  • the audio / video output apparatus 2 is an apparatus that outputs audio with audio localization in accordance with the speaker position, and the audio / video output apparatus 1 has a function of selecting attribute-specific facial feature data based on content program information, and A function for selecting voice characteristics data by attribute based on content program information is added.
  • the same portions are denoted by the same reference numerals and description thereof is omitted.
  • the attribute-specific face feature DB 201 includes data related to attribute-specific face features divided by race.
  • data relating to Western facial feature quantities and data relating to Eastern facial feature quantities are provided.
  • the attribute-specific face feature selection unit 202 is configured to select data relating to the optimal face feature value for each attribute from the program information of the content. For example, if the video content to be reproduced is an American movie, data relating to Western facial features is selected from the program information.
  • the face attribute determination unit 104 determines the face attribute of the detected face feature based on the data regarding the attribute-specific face feature selected by the attribute-specific face feature selection unit 202.
  • the attribute-specific voice feature DB 203 is provided with data related to the voice features by attribute divided by race.
  • data relating to a voice feature value of a Western person and data relating to a voice feature value of an Oriental person are provided.
  • the attribute-specific voice feature selection unit 204 selects data related to the optimum voice feature by attribute from the program information of the content. For example, when the video content to be reproduced is an American movie, data relating to Western voice characteristics is selected from the program information.
  • the voice attribute determination unit 107 determines the voice attribute of the detected voice feature based on the data regarding the voice feature by attribute selected by the attribute-specific voice feature selection unit 204.
  • the video / audio output device 2 of the present embodiment since the data regarding the attribute-specific facial features and the data regarding the attribute-specific voice characteristics are provided for each race, the matching of the face and the voice The accuracy of the degree determination can be further improved. As a result, the speaker position can be more accurately specified, so that the viewer can view with a more natural and realistic feeling.
  • FIG. 7 is a schematic configuration diagram of the video / audio output device 3 according to the third embodiment of the present invention.
  • the audio / video output device 3 is a device that outputs audio with audio localization in accordance with the speaker position, and a function of performing audio localization processing in consideration of viewing environment information is added to the audio / video output device 1.
  • the voice localization processing unit 301 performs a localization change process of the input voice data based on the speaker information output from the matching determination unit 108.
  • the sound localization processing unit 301 performs the localization changing process in consideration of the viewing environment information during the localization changing process.
  • the viewing environment information is, for example, the size of a display screen such as a display or the position of a speaker.
  • the volume of the upper and lower speakers is adjusted not only according to the volume of the left and right speakers as shown in FIG. Is.
  • the viewing environment information may be set by the user instructing the video / audio output device 3 to operate, or the video / audio output device 3 automatically determines the screen size, the position of the speaker, and the like. If it can be determined, it may be automatically set by the video / audio output device 3.
  • the video / audio output device 3 since the sound can be localized and changed to a more accurate position using the viewing environment information, the viewer is more natural and realistic. Viewing is possible.
  • FIG. 8 is a schematic configuration diagram of a video / audio output device 4 according to the fourth embodiment of the present invention.
  • the video / audio output device 4 is a device that outputs audio with audio localization in accordance with the speaker position.
  • the video / audio output device 1 stores data on attribute-specific facial features stored in the attribute-specific facial feature DB 103 and attributes. A function of updating data related to voice features by attribute stored in the different face feature DB 106 is added.
  • the attribute-specific face feature update unit 401 determines that the face feature at that time that is a target of determination when the degree of matching determined by the matching determination unit 108 is a high value that is greater than or equal to a threshold T2 (T2> T1). Is reflected in the data relating to the face feature by attribute in the face feature DB 103 by attribute. As a result, when the face of the same person is input again as video data, the degree of matching between the face and the voice is further increased, and the accuracy of the sound localization can be improved.
  • the attribute-specific voice feature update unit 402 when the degree of fitness determined by the suitability determination unit 108 is a high value equal to or higher than a certain threshold value, determines the voice feature at that time as the attribute of the attribute of the attribute-specific voice feature DB 106. It is to be reflected in the data related to different voice characteristics. Thereby, when the voice of the same person is input again as voice data, the degree of matching between the face and voice is further increased, and the accuracy of voice localization can be improved.
  • the video / audio output device 4 based on the determination of the degree of matching between face and voice, the data related to attribute-specific face features stored in the attribute-specific face feature DB 103 and the attribute-specific face Since the data related to the voice feature for each attribute stored in the feature DB 106 is updated at any time, the data related to the facial feature and the voice feature for each attribute can be made more accurate by reproducing the video content.
  • data related to attributes and facial features and voice features is updated as needed, but in addition to this, data related to individual facial features and voice features is created as needed. It is also possible to update the created data regarding individual facial features and voice features as needed. This makes it possible to identify individuals such as actors and talents who often appear in video content, for example, so that the position of the speaker can be grasped more accurately, and the sound localization can be changed accurately. Can do.
  • Video / audio output device 101 Face region detection unit 102 Face feature detection unit 103, 201 Attribute-specific face feature DB 104 face attribute determination unit 105 voice feature detection unit 106,203 attribute-specific voice feature DB DESCRIPTION OF SYMBOLS 107 Voice attribute determination part 108 Conformity determination part 109,301 Audio localization process part 110 Image

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Human Computer Interaction (AREA)
  • General Health & Medical Sciences (AREA)
  • Theoretical Computer Science (AREA)
  • Biomedical Technology (AREA)
  • Business, Economics & Management (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Game Theory and Decision Science (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Image Analysis (AREA)

Abstract

 映像音声出力装置1は、映像を解析して1又は複数の顔位置を検出する顔領域検出部101と、検出した顔位置それぞれの特徴を抽出する顔特徴検出部102と、抽出したそれぞれの顔特徴と、属性別の顔特徴に関する顔特徴情報を記憶している属性別顔特徴DB103とを比較して、抽出した顔特徴の属性を判定する顔属性判定部104と、音声を解析して、声特徴を抽出する声特徴検出部105と、抽出した声特徴と、属性別の声特徴に関する声特徴情報を記憶している属性別声特徴DB106を比較して、抽出した声特徴の属性を判定する声属性判定部107と、判定した顔特徴の属性と判定した声特徴の属性の適合度を判定する適合度判定部108と、適合度の判定から、話者が画面内にいると判断した場合には、画面内にいると判断した話者の顔位置に音声を定位させる音声定位処理部109と、を備えている。

Description

映像音声出力装置及び音声定位方法
 本発明は、映像及び音声を含むコンテンツデータを出力する映像音声出力装置の音声定位技術に関し、特に、話者位置に応じた音声定位を行う音声定位技術に関する。
 テレビ放送などの番組コンテンツを受信して、ディスプレイに映像を表示するとともにスピーカから音声を出力する場合、モノラル音声においてはスピーカの位置から人の声が聞こえるようになっている。また、ステレオ/サラウンド音声においては、多くの場合、画面中央に人の声を定位させて、画面中央から人の声が聞こえるようになっている。
 しかしながら、一般に、ディスプレイ上の話者位置に人の声が定位していると臨場感が増すことが知られているため、従来においては、映像解析により話者位置を特定し、話者位置に音声を定位させる音声定位技術が開示されている。
 例えば、特許文献1では、話者の位置を検出し、検出した位置に応じて、複数のスピーカから出力する音声の音量を制御している。
特開平11-313272号公報
 しかしながら、上述した従来技術においては、シーンの内容を考慮せずに、話者位置と判定した位置に音声を定位させているため、シーンによっては、臨場感を高めるどころか、却ってストレスを感じてしまう場合がある。例えば、特許文献1では、顔を検出することによって人の位置を特定し、複数の顔を検出した場合には、口の動きが最も大きい顔を話者と判断しているため、話者が画面内にいない場合や話者の口が映像に映っていない場合には、話者の位置が正しく検出できず、誤った位置に音声を定位させてしまうおそれがある。したがって、このような場合には、臨場感を高めるどころか、却って違和感が生じるという問題がある。
 本発明は上記の事情を鑑みてなされたものであり、その課題の一例としては、話者の顔が画面内にない場合も考慮した音声定位技術を提供することにある。
 上記の課題を達成するため、本発明の第1の態様は、入力された映像を解析して、1又は複数の顔位置を検出し、検出した顔位置それぞれの顔特徴を抽出する顔特徴抽出手段と、前記顔特徴抽出手段で抽出したそれぞれの顔特徴の属性を第1属性として判定する顔属性判定手段と、入力された音声を解析して、声特徴を抽出する声特徴抽出手段と、前記声特徴抽出手段で抽出した声特徴の属性を第1属性として判定する声属性判定手段と、前記顔属性判定手段が判定したそれぞれの顔特徴の第1属性と、前記声属性判定手段が判定した声特徴の第1属性の適合度を判定する適合度判定手段と、前記適合度判定手段の判定結果から、話者の顔が画面内にあると判断した場合には、適合度の最も高い話者の顔位置に音声を定位させる音声定位手段と、を備える映像音声出力装置である。
 本発明の第2態様は、入力された映像を解析して、入力された映像を解析して、1又は複数の顔位置を検出し、検出した顔位置それぞれの顔特徴を抽出する顔特徴抽出ステップと、前記顔特徴抽出ステップで抽出したそれぞれの顔特徴と、属性別の顔特徴に関する顔特徴情報を記憶している顔特徴データベースに記憶された顔特徴情報とを比較して、抽出したそれぞれの顔特徴の属性を判定する顔属性判定ステップと、入力された音声を解析して、声特徴を抽出する声特徴抽出ステップと、前記声特徴抽出ステップで抽出した声特徴と、属性別の声特徴に関する声特徴情報を記憶している声特徴データベースに記憶された声特徴情報とを比較して、抽出した声特徴の属性を判定する声属性判定ステップと、前記顔属性判定ステップで判定したそれぞれの顔特徴の属性と、前記声属性判定ステップで判定した声特徴の属性の適合度を判定する適合度判定ステップと、前記適合度判定ステップの判定結果から、話者の顔が画面内にあると判断した場合には、適合度の最も高い話者の顔位置に音声を定位させる音声定位ステップと、を有する音声定位方法である。
本発明の第1の実施の形態に係る映像音声出力装置の概略構成図である。 本発明の第1の実施の形態に係る映像音声出力装置が表示する画像の一例である。 本発明の実施の形態に係る映像音声出力装置による顔属性判定結果及び声属性判定結果の一例である。 本発明の実施の形態に係る映像音声出力装置による顔属性判定結果及び声属性判定結果の一例である。 本発明の第1の実施の形態に係る映像音声出力装置の映像音声出力処理の流れを示すフローチャートである。 本発明の第2の実施の形態に係る映像音声出力装置の概略構成図である。 本発明の第3の実施の形態に係る映像音声出力装置の概略構成図である。 本発明の第4の実施の形態に係る映像音声出力装置の概略構成図である。
 以下、本発明の実施の形態を図面を用いて説明する。
 図1は、本発明の実施の形態に係る映像音声出力装置1の概略構成図である。映像音声出力装置1は、話者位置に合わせた音声定位で音声を出力する装置である。本実施の形態においては、話者の顔が画面内にあるか否かの判断を行い、話者の顔が画面内にある場合には、画面内にいる話者の位置に合わせて音声を定位変更している。
 なお、以下において、「話者」とは、映像データ(画面上)において発話している者をいい、「話者位置」とは、話者の画面上の位置をいうが、より正確には話者の顔付近の位置をいう。「話者位置に合わせた音声定位で音声を出力する」とは、話者の位置から音声が聞こえてくるように音声を出力することをいい、例えば、図2に示すように、話者Aが画面上左側に存在する場合には、画面d10の左側に設けたスピーカSP1から出力される話者音声の音量を大きくし、画面d10の右側に設けたスピーカSP2から出力される話者音声の音量を小さくして、画面左側にいる話者の位置から音声が聞こえてくるように音声を出力することをいう。例えば、話者Aが画面上、水平方向の比がC:Dとなる位置に存在する場合には、スピーカSP1とスピーカSP2の話者音声の音量の比をD:Cに設定してもよい。
 ここで、映像音声出力装置1は、外部から入力された映像及び音声を含むコンテンツデータを再生して外部に出力する機能を有する装置であれば何であってもよく、例えば、具体的には、テレビジョン(TV)、DVDプレーヤ及びレコーダ、BDプレーヤ及びレコーダ、パーソナルコンピュータ(PC)などが想定される。
 映像音声出力装置1は、詳しくは、顔領域検出部101、顔特徴検出部102、属性別顔特徴データベース(以下、属性別顔特徴DBと表記する)103、顔属性判定部104、声特徴検出部105、属性別声特徴データベース(以下、属性別声特徴DBと表記する)106、声属性判定部107、適合判定部108、音声定位処理部109、映像表示部110、及び音声出力部111を備えている。
 顔領域検出部101は、入力した映像データから人の顔領域を検出するようになっている。顔領域の検出は、公知の技術を用いて行われる。例えば、肌色領域を探し出すことにより顔領域を検出する技術やテンプレート(顔パターン画像)を画像上で比較しながら顔領域を検出するテンプレートマッチング法により顔領域を検出する技術などである。そして、顔領域検出部101は、検出した顔領域の位置情報(顔領域位置情報)を顔特徴検出部102に出力するようになっている。なお、複数の顔領域が検出された場合には、それぞれの顔領域位置情報が顔特徴検出部102に出力され、顔領域が検出されない場合には、顔位置情報は顔特徴検出部102に出力されない。
 顔特徴検出部102は、顔領域検出部101から出力されたそれぞれの顔領域位置情報に基づいて、入力した映像データから人の顔特徴をそれぞれ検出するようになっている。顔特徴の検出は、公知の技術を用いて行われる。本実施の形態では、顔の主要な部位に対してその特徴量を検出する方法を採用しており、顔輪郭、両眉、両眼、鼻、口などの位置関係を示す特徴量を顔特徴として検出する。そして、顔特徴検出部102は、検出した顔特徴を顔属性判定部104に出力するようになっている。
 属性別顔特徴DB103は、属性別の顔特徴に関するデータを記憶しているデータベースである。属性別とは、例えば、性別及び年齢別のことを意味し、本実施の形態では、一例として、20歳未満男性、20歳以上49歳以下男性、50歳以上男性、20歳未満女性、20歳以上49歳以下女性、及び50歳以上女性の6つの属性に分類され、分類された属性ごとに顔特徴に関するデータを記憶している。なお、本実施の形態では、6つの属性に分類しているが、属性の分類はこれに限定されず、さらに細分化してもよい。
 また、本実施の形態では、映像音声出力装置1が属性別顔特徴DB103を備える構成としているが、映像音声出力装置1が属性別顔特徴DB103を備えなくてもよい。例えば、映像音声出力装置1が、通信ネットワークを介して属性別顔特徴DB103にアクセスし、属性別顔特徴DB103に記憶されたデータを参照する構成としてもよい。
 顔属性判定部104は、顔特徴検出部102から出力されたそれぞれの顔特徴と、属性別顔特徴DB103に記憶された属性別の顔特徴に関するデータを比較して、それぞれの顔の属性を判定するようになっている。判定された顔属性の結果(顔属性判定結果という)は、各属性に属する確率を備えたデータであり、本実施の形態では、例えば、図3に示すようなデータとなる。そして、顔属性判定部104は、顔属性判定結果を適合判定部108に出力するようになっている。
 なお、顔特徴検出部102、属性別顔特徴DB103及び顔属性判定部104において、顔特徴から属性(年齢、性別)を判定する方法は、公知の技術を用いて行われる。例えば、電子情報通信学会、信学技法、PRMU2001-138(2001年)「平均顔との距離を用いた性別・年齢推定手法の提案」に記載された方法により、検知した顔と属性別の平均顔との特徴点距離に基づいて、合致する属性を判別するようにしてもよい。
 声特徴検出部105は、入力した音声データから人の声特徴を検出するようになっている。声特徴の検出は、公知の技術を用いて行われる。本実施の形態では、声の周波数を分析し、ホルマント周波数やピッチ周波数を特定し、特定したホルマント周波数やピッチ周波数を声特徴として検出する。そして、声特徴検出部105は、検出した声特徴を声属性判定部107に出力するようになっている。
 属性別声特徴DB106は、属性別の声特徴に関するデータを記憶しているデータベースである。属性別とは、例えば、性別及び年齢別のことを意味し、本実施の形態では、一例として、20歳未満男性、20歳以上49歳以下男性、50歳以上男性、20歳未満女性、20歳以上49歳以下女性、50歳以上女性の6つの属性に分類され、分類された属性ごとに声特徴に関するデータを記憶している。なお、本実施の形態では、上述した6つの属性に分類しているが、属性の分類はこれに限定されない。
 また、本実施の形態では、映像音声出力装置1が属性別声特徴DB103を備える構成としているが、映像音声出力装置1が属性別声特徴DB106を備えなくてもよい。例えば、映像音声出力装置1が、通信ネットワークを介して属性別声特徴DB106にアクセスし、属性別声特徴DB106に記憶されたデータを参照する構成としてもよい。
 声属性判定部107は、声特徴検出部105から出力された声特徴と、属性別声特徴DB106に記憶された属性別の声特徴に関するデータを比較して、声の属性を判定するようになっている。判定された声属性の結果(声属性判定結果という)は、各属性に属する確率を備えたデータであり、本実施の形態では、例えば、図3に示すようなデータとなる。そして、声属性判定部107は、声属性判定結果を適合判定部108に出力するようになっている。
 なお、声特徴検出部105、属性別声特徴DB106及び声属性判定部107において、声特徴から属性(年齢、性別)を判定する方法は、公知の技術を用いて行われる。例えば、日本音響学会誌第24巻第6号(1968年)「年齢、性別による日本語5母音のピッチ周波数とホルマント周波数の変化」に記載された方法により、検出した声特徴に合致する属性を判別するようにしてもよい。
 適合判定部108は、顔属性判定部104から出力された顔属性判定結果と、声属性判定部107から出力された声属性判定結果とに基づいて、顔と声の適合度の判定を行うようになっている。詳しくは、話者の顔が画面内にあるか否か、話者の顔が画面内にある場合には、どの顔が話者の顔であるかを判定する適合判定処理を行う。
そして、適合判定部108は、適合判定処理の結果である話者情報を音声定位処理部109に出力するようになっている。
 図3及び図4を用いて、適合判定処理を具体的に説明する。図3は、画面から3人の顔特徴が検出された場合のそれぞれの顔属性判定結果と顔属性判定結果を示しており、図4は、2人の顔特徴が検出された場合のそれぞれの顔属性判定結果と顔属性判定結果を示している。本実施の形態においては、顔と声の適合度は、同一属性に属する値同士の乗算を、すべての属性について行い、それぞれの乗算値を加算することにより算出される。
 例えば、図3に示す場合には、顔1と声の適合度A1は、
 A1=0.7×0.9+0.1×0.05+0.1×0.05+0.05×0+0.05×0+0×0=0.64
となり、顔2と声の適合度A2は、
 A2=0.05×0.9+0.1×0.05+0.1×0.05+0.6×0+0.05×0+0.1×0=0.055
となり、顔3と声の適合度A3は、
 A3=0×0.9+0.05×0.05+0.05×0.05+0.1×0+0.1×0+0.7×0=0.005
となる。ここで、話者の顔が画面内にあるか否かの閾値T1を一例として0.5とし、これ以上の値のときは話者の顔が画面内にあると判断すると、適合度が最も大きい値はA1であり、かつ、A1≧T1であるので、画面内に話者はいると判断し、話者は顔1を有する者であると判断する。なお、この場合には、顔1の位置を示す位置情報が話者情報として音声定位処理部109に出力される。
 また、図4に示す場合には、顔1と声の適合度B1は、
 B1=0.05×0.9+0.1×0.05+0.1×0.05+0.6×0+0.05×0+0.1×0=0.055
となり、顔2と声の適合度B2は、
 B2=0×0.9+0.05×0.05+0.05×0.05+0.1×0+0.1×0+0.7×0=0.005
となる。この場合には、適合度が最も大きい値はB1であるが、B1<T1であるので、画面内に話者はいないと判断する。なお、この場合には、画面内に話者がいないことを示す情報が話者情報として音声定位処理部109に出力される。
 音声定位処理部109は、適合判定部108から出力された話者位置情報に基づいて、入力した音声データの定位変更処理を行うようになっている。すなわち、画面上の話者位置に音声を定位させるように音量の調整を行っている。例えば、図2に示すように、話者Aが画面上左側に存在する場合には、画面d10の左側に設けたスピーカSP1から出力される音声の音量を大きくし、画面d10の右側に設けたスピーカSP2から出力される音声の音量を小さくする。
 また、音声定位処理部109は、定位変更処理をした音声を音声出力部111に出力するようになっている。
 映像表示部110は、入力した映像データをディスプレイ等に出力表示するようになっている。なお、出力表示される映像データは、出力される音声データと同期させるため、必要に応じて遅延させている。
 音声出力部111は、定位変更処理された音声データをスピーカに出力するようになっている。
 次に、図5を参照して、本実施の形態の映像音声出力装置1の映像音声出力処理について説明する。図5は、映像音声出力装置1の映像音声出力処理の流れを示すフローチャートである。
 まず、映像音声出力装置1は、入力された映像データを解析して、顔属性を判定する顔属性判定処理を行う(ステップS10)。詳しくは、まず、顔領域検出部101が入力された映像データから人の顔領域を検出し、次いで、顔特徴検出部102が、検出した顔領域に基づいて、入力した映像データから人の顔特徴を検出し、最後に、顔属性判定部104が、検出した顔特徴と属性別顔特徴DB103のデータを比較して検出した顔の属性を判定する。
 次に、映像音声出力装置1は、入力された音声データを解析して、声属性を判定する声属性判定処理を行う(ステップS20)。詳しくは、まず、声特徴検出部105が入力した音声データから人の声特徴を検出し、次いで、声属性判定部107が検出した声特徴と属性別声特徴DB106のデータを比較して検出した声の属性を判定する。
 次に、映像音声出力装置1の適合判定部108は、顔属性判定結果と声属性判定結果とに基づいて、顔と声の適合度の判定を行う適合度判定処理を行う(ステップS30)。
 適合度判定処理の結果、話者の顔が画面内にあると判定した場合には(ステップS40:YES)、画面内にいると判定した話者の顔位置に音声を定位変更する音声定位処理を行う(ステップS50)。例えば、話者位置に近いスピーカの話者音声の出力値を上げて、話者位置に遠いスピーカの話者音声の出力値を下げる。
 適合度判定処理の結果、話者の顔が画面内にないと判定した場合には(ステップS40:NO)、音声定位処理を行わない。
 次に、映像音声出力装置1の映像表示部110は、映像データを出力し、また、音声出力部111は、話者の顔が画面内にある場合には、音声定位変更を行われた音声データを出力し、話者の顔が画面内にない場合には、音声定位変更を行われていない音声データを出力する(ステップS60)。
 以上説明したように、本実施の形態に係る映像音声出力装置1によれば 映像解析に基づく顔属性判定結果と音声解析に基づく声属性判定結果の適合度を判定することにより、話者の顔が画面内にあるか否かの判断を行うことができる。そして、判断の結果、話者の顔が画面内にある場合には、特定した話者位置に音声を定位変更する一方、話者の顔が画面内にない場合には、音声の定位変更を行わないので、視聴者は違和感を生じることがなく、より自然で臨場感のある視聴が可能となる。
<第2の実施の形態>
 図6は、本発明の第2の実施の形態に係る映像音声出力装置2の概略構成図である。映像音声出力装置2は、話者位置に合わせた音声定位で音声を出力する装置であり、映像音声出力装置1に、コンテンツのプログラム情報に基づいて属性別顔特徴のデータを選択する機能、及びコンテンツのプログラム情報に基づいて属性別声特徴のデータを選択する機能を追加している。なお、以下においては、第1の実施の形態と異なる構成、機能及び処理のみ説明し、その他の構成、機能及び処理に関しては同一部位には同一符号を付して説明を省略する。
 属性別顔特徴DB201は、属性別の顔特徴に関するデータを人種ごとに細分化して備えている。本実施の形態では、西洋人の顔特徴量に関するデータ、及び東洋人の顔特徴量に関するデータを備えている。
 属性別顔特徴選択部202は、コンテンツのプログラム情報から、最適な属性別の顔特徴量に関するデータを選択するようになっている。例えば、プログラム情報から、再生する映像コンテンツがアメリカ映画である場合には、西洋人の顔特徴に関するデータを選択するようになっている。
 なお、顔属性判定部104は、属性別顔特徴選択部202により選択された属性別の顔特徴に関するデータに基づいて、検出した顔特徴の顔属性を判定するようになっている。
 属性別声特徴DB203は、属性別の声特徴に関するデータを人種ごとに細分化して備えている。本実施の形態では、西洋人の声特徴量に関するデータ、及び東洋人の声特徴量に関するデータを備えている。
 属性別声特徴選択部204は、コンテンツのプログラム情報から、最適な属性別の声特徴に関するデータを選択するようになっている。例えば、プログラム情報から、再生する映像コンテンツがアメリカ映画である場合には、西洋人の声特徴に関するデータを選択するようになっている。
 なお、声属性判定部107は、属性別声特徴選択部204により選択された属性別の声特徴に関するデータに基づいて、検出した声特徴の声属性を判定するようになっている。
 以上、本実施の形態の映像音声出力装置2によれば、属性別の顔特徴に関するデータ、及び属性別の声特徴に関するデータを人種ごとに細分化して備えているので、顔と声の適合度判定の正確性をさらに向上させることができる。この結果、話者位置をさらに正確に特定することができるので、視聴者は、より自然で臨場感のある視聴が可能となる。
<第3の実施の形態>
 図7は、本発明の第3の実施の形態に係る映像音声出力装置3の概略構成図である。映像音声出力装置3は、話者位置に合わせた音声定位で音声を出力する装置であり、映像音声出力装置1に、視聴環境情報を加味して音声定位処理を行う機能を追加している。
 すなわち、音声定位処理部301は、適合判定部108から出力された話者情報に基づいて、入力した音声データの定位変更処理を行うようになっている。
 また、音声定位処理部301は、定位変更処理の際、視聴環境情報を加味して定位変更処理を行っている。ここで、視聴環境情報とは、例えば、ディスプレイ等の表示画面の大きさやスピーカの位置などである。例えば、画面の上下左右方向それぞれにスピーカを備える場合には、図2に示すように左右スピーカの音量の調整だけでなく、話者の上下方向の位置に応じて上下スピーカの音量の調整を行うものである。
 なお、視聴環境情報は、ユーザが映像音声出力装置3に操作指示することで設定されるようにしてもよいし、また、映像音声出力装置3が画面の大きさやスピーカの位置などを自動的に判別できる場合には、映像音声出力装置3により自動的に設定されるようにしてもよい。
 以上、本実施の形態に係る映像音声出力装置3によれば、視聴環境情報を用いて、さらに正確な位置に音声を定位変更することができるので、視聴者は、より自然で臨場感のある視聴が可能となる。
<第4の実施の形態>
 図8は、本発明の第4の実施の形態に係る映像音声出力装置4の概略構成図である。映像音声出力装置4は、話者位置に合わせた音声定位で音声を出力する装置であり、映像音声出力装置1に、属性別顔特徴DB103に記憶された属性別の顔特徴に関するデータ、及び属性別顔特徴DB106に記憶された属性別の声特徴に関するデータを更新する機能を追加している。
 すなわち、属性別顔特徴更新部401は、適合判定部108により判定された適合度がある閾値T2(T2>T1)以上の高い値の場合には、判定の対象となったそのときの顔特徴を属性別顔特徴DB103の属性別の顔特徴に関するデータに反映するようになっている。これにより、再度同一人物の顔が映像データとして入力された場合には、顔と声の適合度がさらに高くなり、音声定位の正確性を高めることができる。
 また、属性別声特徴更新部402は、適合判定部108により判定された適合度がある閾値以上の高い値の場合には、判定の対象となったそのときの声特徴を属性別声特徴DB106の属性別の声特徴に関するデータに反映するようになっている。これにより、再度同一人物の声が音声データとして入力された場合には、顔と声の適合度がさらに高くなり、音声定位の正確性を高めることができる。
 以上、本実施の形態に係る映像音声出力装置4によれば、顔と声の適合度の判定に基づいて、属性別顔特徴DB103に記憶された属性別の顔特徴に関するデータ、及び属性別顔特徴DB106に記憶された属性別の声特徴に関するデータを随時更新していくので、映像コンテンツを再生していくことにより属性別の顔特徴や声特徴に関するデータをさらに正確にしていくことができる。
 なお、本実施の形態においては、属性別の顔特徴や声特徴に関するデータを随時更新していくようにしたが、これに加えて、個人別の顔特徴や声特徴に関するデータを随時作成していき、作成した個人別の顔特徴や声特徴に関するデータを随時更新していくようにしてもよい。これにより、例えば、映像コンテンツによく出演する俳優やタレントなど個人を特定することが可能となるので、話者位置をさらに正確に把握することができ、以て正確に音声の定位変更を行うことができる。
 以上、本発明の実施の形態について説明してきたが、本発明は、上述した実施の形態に限られるものではなく、本発明の要旨を逸脱しない範囲において、本発明の実施の形態に対して種々の変形や変更を施すことができ、そのような変形や変更を伴うものもまた、本発明の技術的範囲に含まれるものである。
  1,2,3、4 映像音声出力装置
  101 顔領域検出部
  102 顔特徴検出部
  103,201 属性別顔特徴DB
  104 顔属性判定部
  105 声特徴検出部
  106,203 属性別声特徴DB
  107 声属性判定部
  108 適合判定部
  109,301 音声定位処理部
  110 映像表示部
  111 音声出力部
  202 属性別顔特徴選択部
  204 属性別声特徴選択部
  401 属性別顔特徴更新部
  402 属性別声特徴更新部

Claims (7)

  1.  入力された映像を解析して、1又は複数の顔位置を検出し、検出した顔位置それぞれの顔特徴を抽出する顔特徴抽出手段と、
     前記顔特徴抽出手段で抽出したそれぞれの顔特徴の属性を第1属性として判定する顔属性判定手段と、
     入力された音声を解析して、声特徴を抽出する声特徴抽出手段と、
     前記声特徴抽出手段で抽出した声特徴の属性を第1属性として判定する声属性判定手段と、
     前記顔属性判定手段が判定したそれぞれの顔特徴の第1属性と、前記声属性判定手段が判定した声特徴の第1属性の適合度を判定する適合度判定手段と、
     前記適合度判定手段の判定結果から、話者の顔が画面内にあると判断した場合には、適合度の最も高い話者の顔位置に音声を定位させる音声定位手段と、
    を備えることを特徴とする映像音声出力装置。
  2.  属性別の顔特徴に関する顔特徴情報を記憶している顔特徴データベースと、
     属性別の声特徴に関する声特徴情報を記憶している声特徴データベースと、をさらに備え、
     前記顔属性判定手段は、
     前記顔特徴抽出手段で抽出したそれぞれの顔特徴と、前記顔特徴データベースに記憶された顔特徴情報とを比較して、前記顔特徴の前記第1属性を判定し、
     前記声属性判定手段は、
     前記声特徴抽出手段で抽出した声特徴と、前記声特徴データベースに記憶された声特徴情報とを比較して、前記声特徴の前記第1属性を判定することを特徴とする請求項1記載の映像音声出力装置。
  3.  前記顔特徴データベースは、さらに、所定の第2属性ごとに、前記顔特徴情報を記憶し、
     前記顔特徴データベースは、さらに、所定の第2属性ごとに、前記声特徴情報を記憶し、
     入力されたコンテンツプログラム情報に基づいて、前記顔特徴データベースの中から、再生するコンテンツの種類に合った、前記第2属性の顔特徴情報を選択する顔特徴情報選択手段と、
     入力されたコンテンツプログラム情報に基づいて、前記前記声特徴データベースの中から、再生するコンテンツの種類に合った、前記第2属性の声特徴情報を選択する声特徴情報選択手段と、
    をさらに備え、
     前記顔属性判定手段は、
     前記顔特徴抽出手段で抽出したそれぞれの顔特徴と、前記顔特徴情報選択手段により選択された顔特徴情報とを比較して、抽出したそれぞれの顔特徴の属性を判定し、
     前記声属性判定手段は、
     前記声特徴抽出手段で抽出した声特徴と、前記顔特徴情報選択手段により選択された声特徴情報とを比較して、抽出した声特徴の属性を判定することを特徴とする請求項2記載の映像音声出力装置。
  4.  前記第2属性は、人種であることを特徴とする請求項3記載の映像音声出力装置。
  5.  画面の大きさやスピーカの設置位置に関する視聴環境情報を備え、
     前記音声定位手段は、前記視聴環境情報を加味して、音声を定位させることを特徴とする請求項2乃至4のいずれか1項に記載の映像音声出力装置。
  6.  前記適合度判定手段の判定結果に基づいて、前記顔特徴データベースに記憶された顔特徴情報を更新する顔特徴情報更新手段と、
     前記適合度判定手段の判定結果に基づいて、前記声特徴データベースに記憶された声特徴情報を更新する声特徴情報更新手段と、
    をさらに備えることを特徴とする請求項2乃至5のいずれか1項に記載の映像音声出力装置。
  7.  入力された映像を解析して、1又は複数の顔位置を検出し、検出した顔位置それぞれの顔特徴を抽出する顔特徴抽出ステップと、
     前記顔特徴抽出ステップで抽出したそれぞれの顔特徴と、属性別の顔特徴に関する顔特徴情報を記憶している顔特徴データベースに記憶された顔特徴情報とを比較して、抽出したそれぞれの顔特徴の属性を判定する顔属性判定ステップと、
     入力された音声を解析して、声特徴を抽出する声特徴抽出ステップと、
     前記声特徴抽出ステップで抽出した声特徴と、属性別の声特徴に関する声特徴情報を記憶している声特徴データベースに記憶された声特徴情報とを比較して、抽出した声特徴の属性を判定する声属性判定ステップと、
     前記顔属性判定ステップで判定したそれぞれの顔特徴の属性と、前記声属性判定ステップで判定した声特徴の属性の適合度を判定する適合度判定ステップと、
     前記適合度判定ステップの判定結果から、話者の顔が画面内にあると判断した場合には、適合度の最も高い話者の顔位置に音声を定位させる音声定位ステップと、
    を有することを特徴とする音声定位方法。
PCT/JP2009/060362 2009-06-05 2009-06-05 映像音声出力装置及び音声定位方法 Ceased WO2010140254A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/JP2009/060362 WO2010140254A1 (ja) 2009-06-05 2009-06-05 映像音声出力装置及び音声定位方法

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2009/060362 WO2010140254A1 (ja) 2009-06-05 2009-06-05 映像音声出力装置及び音声定位方法

Publications (1)

Publication Number Publication Date
WO2010140254A1 true WO2010140254A1 (ja) 2010-12-09

Family

ID=43297400

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2009/060362 Ceased WO2010140254A1 (ja) 2009-06-05 2009-06-05 映像音声出力装置及び音声定位方法

Country Status (1)

Country Link
WO (1) WO2010140254A1 (ja)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2014127019A1 (en) * 2013-02-15 2014-08-21 Qualcomm Incorporated Video analysis assisted generation of multi-channel audio data
JP2019152737A (ja) * 2018-03-02 2019-09-12 株式会社日立製作所 話者推定方法および話者推定装置
EP3706442A1 (en) * 2019-03-08 2020-09-09 LG Electronics Inc. Method and apparatus for sound object following
CN112929739A (zh) * 2021-01-27 2021-06-08 维沃移动通信有限公司 发声控制方法、装置、电子设备和存储介质
CN115131405A (zh) * 2022-07-07 2022-09-30 沈阳航空航天大学 一种基于多模态信息的发言人跟踪方法及系统

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH11313272A (ja) * 1998-04-27 1999-11-09 Sharp Corp 映像音声出力装置
JP2000295700A (ja) * 1999-04-02 2000-10-20 Nippon Telegr & Teleph Corp <Ntt> 画像情報を用いた音源定位方法及び装置及び該方法を実現するプログラムを記録した記憶媒体
JP2004056286A (ja) * 2002-07-17 2004-02-19 Fuji Photo Film Co Ltd 画像表示方法
JP2007201818A (ja) * 2006-01-26 2007-08-09 Sony Corp オーディオ信号処理装置、オーディオ信号処理方法及びオーディオ信号処理プログラム

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH11313272A (ja) * 1998-04-27 1999-11-09 Sharp Corp 映像音声出力装置
JP2000295700A (ja) * 1999-04-02 2000-10-20 Nippon Telegr & Teleph Corp <Ntt> 画像情報を用いた音源定位方法及び装置及び該方法を実現するプログラムを記録した記憶媒体
JP2004056286A (ja) * 2002-07-17 2004-02-19 Fuji Photo Film Co Ltd 画像表示方法
JP2007201818A (ja) * 2006-01-26 2007-08-09 Sony Corp オーディオ信号処理装置、オーディオ信号処理方法及びオーディオ信号処理プログラム

Cited By (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2014127019A1 (en) * 2013-02-15 2014-08-21 Qualcomm Incorporated Video analysis assisted generation of multi-channel audio data
CN104995681A (zh) * 2013-02-15 2015-10-21 高通股份有限公司 多声道音频数据的视频分析辅助产生
US9338420B2 (en) 2013-02-15 2016-05-10 Qualcomm Incorporated Video analysis assisted generation of multi-channel audio data
JP2016513410A (ja) * 2013-02-15 2016-05-12 クゥアルコム・インコーポレイテッドQualcomm Incorporated マルチチャネルオーディオデータのビデオ解析支援生成
CN104995681B (zh) * 2013-02-15 2017-10-31 高通股份有限公司 多声道音频数据的视频分析辅助产生
JP2019152737A (ja) * 2018-03-02 2019-09-12 株式会社日立製作所 話者推定方法および話者推定装置
EP3706442A1 (en) * 2019-03-08 2020-09-09 LG Electronics Inc. Method and apparatus for sound object following
KR20200107757A (ko) * 2019-03-08 2020-09-16 엘지전자 주식회사 음향 객체 추종을 위한 방법 및 이를 위한 장치
KR20200107758A (ko) * 2019-03-08 2020-09-16 엘지전자 주식회사 음향 객체 추종을 위한 방법 및 이를 위한 장치
US11277702B2 (en) 2019-03-08 2022-03-15 Lg Electronics Inc. Method and apparatus for sound object following
KR102737006B1 (ko) * 2019-03-08 2024-12-02 엘지전자 주식회사 음향 객체 추종을 위한 방법 및 이를 위한 장치
KR102758939B1 (ko) 2019-03-08 2025-01-23 엘지전자 주식회사 음향 객체 추종을 위한 방법 및 이를 위한 장치
CN112929739A (zh) * 2021-01-27 2021-06-08 维沃移动通信有限公司 发声控制方法、装置、电子设备和存储介质
CN115131405A (zh) * 2022-07-07 2022-09-30 沈阳航空航天大学 一种基于多模态信息的发言人跟踪方法及系统

Similar Documents

Publication Publication Date Title
US8326623B2 (en) Electronic apparatus and display process method
US20210249012A1 (en) Systems and methods for operating an output device
US9251805B2 (en) Method for processing speech of particular speaker, electronic system for the same, and program for electronic system
KR101378493B1 (ko) 영상 데이터에 동기화된 텍스트 데이터 설정 방법 및 장치
KR101958664B1 (ko) 멀티미디어 콘텐츠 재생 시스템에서 다양한 오디오 환경을 제공하기 위한 장치 및 방법
US20050038661A1 (en) Closed caption control apparatus and method therefor
US11122341B1 (en) Contextual event summary annotations for video streams
JP4736511B2 (ja) 情報提供方法および情報提供装置
KR20150093425A (ko) 콘텐츠 추천 방법 및 장치
US11211074B2 (en) Presentation of audio and visual content at live events based on user accessibility
JP2011250100A (ja) 画像処理装置および方法、並びにプログラム
JPWO2011132403A1 (ja) 補聴器フィッティング装置
CN108055592A (zh) 字幕显示方法、装置、移动终端及存储介质
US12427413B2 (en) Automatic in-game subtitles and closed captions
US20140064517A1 (en) Multimedia processing system and audio signal processing method
WO2010140254A1 (ja) 映像音声出力装置及び音声定位方法
CN112601120A (zh) 字幕显示方法及装置
JP2010124391A (ja) 情報処理装置、機能設定方法及び機能設定プログラム
Song et al. How different kinds of sound in videos can influence gaze
WO2010131318A1 (ja) 映像音声出力装置及び音声定位方法
JP7456492B2 (ja) 音声処理装置、音声処理システム、音声処理方法及びプログラム
KR102898795B1 (ko) 실시간 동시 번역 정보 제공이 가능한 커뮤니케이션 기반 사용자 상호 매칭 플랫폼 서비스 제공 방법, 장치 및 시스템
JP2026022883A (ja) 情報処理装置、制御方法、およびプログラム
WO2026002730A1 (en) Data processing apparatus, system and method
JP4264028B2 (ja) 要約番組生成装置、及び要約番組生成プログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 09845537

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 09845537

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: JP