WO2020162220A1 - クレジット区間特定装置、クレジット区間特定方法及びプログラム - Google Patents

クレジット区間特定装置、クレジット区間特定方法及びプログラム Download PDF

Info

Publication number
WO2020162220A1
WO2020162220A1 PCT/JP2020/002458 JP2020002458W WO2020162220A1 WO 2020162220 A1 WO2020162220 A1 WO 2020162220A1 JP 2020002458 W JP2020002458 W JP 2020002458W WO 2020162220 A1 WO2020162220 A1 WO 2020162220A1
Authority
WO
WIPO (PCT)
Prior art keywords
credit
audio signal
partial
section
signal
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2020/002458
Other languages
English (en)
French (fr)
Inventor
康智 大石
川西 隆仁
柏野 邦夫
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to US17/428,612 priority Critical patent/US12494222B2/en
Publication of WO2020162220A1 publication Critical patent/WO2020162220A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • G10L25/57Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for processing of video signals
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q30/00Commerce
    • G06Q30/02Marketing; Price estimation or determination; Fundraising
    • G06Q30/0241Advertisements
    • G06Q30/0242Determining effectiveness of advertisements
    • G06Q30/0246Traffic
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/60Information retrieval; Database structures therefor; File system structures therefor of audio data
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/60Information retrieval; Database structures therefor; File system structures therefor of audio data
    • G06F16/65Clustering; Classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/44Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
    • G06V10/443Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components by matching or filtering
    • G06V10/449Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters
    • G06V10/451Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters with interaction between the filter responses, e.g. cortical complex cells
    • G06V10/454Integrating the filters into a hierarchical structure, e.g. convolutional neural networks [CNN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/46Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/02Feature extraction for speech recognition; Selection of recognition unit
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N17/00Diagnosis, testing or measuring for television systems or their details
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/442Monitoring of processes or resources, e.g. detecting the failure of a recording device, monitoring the downstream bandwidth, the number of times a movie has been viewed, the storage space available from the internal hard disk
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition

Definitions

  • the present invention relates to a credit section identifying device, a credit section identifying method, and a program.
  • Such a survey is conducted by visually recognizing the display of the provided credits on TV broadcasts and the like and transcribing the company name from the provided credits.
  • the provided credit refers to the display of the logo of the sponsor of the broadcast program and narration (for example, "This program will be sent as XXX and the sponsor who is watching").
  • the section in which the provided credit is displayed is only about 1% of the broadcast time. Therefore, in the above-mentioned survey, much time is spent for viewing and listening work such as television broadcasting for identifying the section of the provided credit.
  • the present invention has been made in view of the above points, and it is an object of the present invention to efficiently identify a credit section.
  • the credit section identifying device includes a plurality of first partial voices, each of which is a part of the first voice signal from the first voice signal and has a mutual shift in the time direction.
  • An extraction unit that extracts a signal and whether or not each of the first partial audio signals includes a credit is determined by associating each second partial audio signal extracted from the second audio signal with the presence or absence of credit.
  • a specifying unit that specifies a credit section in the first audio signal.
  • FIG. 6 is a flowchart for explaining an example of a processing procedure of learning processing according to the first embodiment. It is a figure which shows the example of extraction of the audio
  • FIG. 9 is a flowchart for explaining an example of a processing procedure of learning processing according to the second embodiment. It is a figure which shows the example of extraction of the pair of the audio segment and still image of the positive example in 2nd Embodiment. It is a figure which shows the example of a model of the discriminator in 2nd Embodiment. It is a flow chart for explaining an example of a processing procedure of detection processing of a provided credit in a 2nd embodiment. It is a figure which shows the example of extraction of the audio segment from the audio signal for a detection in 2nd Embodiment. It is a figure which shows an example of the evaluation result of each embodiment.
  • FIG. 1 is a diagram illustrating a hardware configuration example of a provided credit section identification device 10 according to the first embodiment.
  • the provided credit section identification device 10 of FIG. 1 is a computer including a drive device 100, an auxiliary storage device 102, a memory device 103, a CPU 104, an interface device 105, and the like, which are connected to each other by a bus B.
  • the program that realizes the processing in the provided credit section identification device 10 is provided by the recording medium 101 such as a CD-ROM.
  • the recording medium 101 storing the program is set in the drive device 100, the program is installed in the auxiliary storage device 102 from the recording medium 101 via the drive device 100.
  • the program does not necessarily have to be installed from the recording medium 101, and may be downloaded from another computer via a network.
  • the auxiliary storage device 102 stores the installed program and also stores necessary files and data.
  • the memory device 103 reads the program from the auxiliary storage device 102 and stores it when an instruction to activate the program is given.
  • the CPU 104 executes the function of the provided credit section identifying device 10 according to the program stored in the memory device 103.
  • the interface device 105 is used as an interface for connecting to a network.
  • FIG. 2 is a diagram showing a functional configuration example of the provided credit section identification device 10 according to the first embodiment.
  • the provided credit section identification device 10 includes a learning data generation unit 11, a learning unit 12, a detection data generation unit 13, a provided credit section estimation unit 14, a time information output unit 15, and the like.
  • Each of these units is realized by a process that causes the CPU 104 to execute one or more programs installed in the provided credit section identification device 10.
  • the provided credit section identifying device 10 also uses the correct answer storage unit 121, the related term storage unit 122, the parameter storage unit 123, and the like.
  • Each of these storage units can be realized using, for example, an auxiliary storage device 102 or a storage device that can be connected to the provided credit section identifying device 10 via a network.
  • the correct answer storage unit 121 provides credits for audio signals (hereinafter, referred to as “learning audio signals”) of learning TV broadcasts (hereinafter referred to as “learning TV broadcasts”) broadcast during a certain period.
  • the time data (start time, end time) indicating the section (hereinafter referred to as “provided credit section”) is stored.
  • the provided credit section may be confirmed in advance by the user's eyes or the like.
  • the related word/phrase storage unit 122 stores related words/phrases that are included in the announcement when the provided credit is displayed (announcement that flows when the provided credit is displayed) and related to the provided credit display. Examples of related words and phrases include words such as "view”, “sponsor", “offer”, and “send (send/send)". Further, a word or the like indicating a company name may be set as the related word or phrase. Note that the related term is preset by the user, for example.
  • the parameter storage unit 123 stores the parameters of the discriminator that identifies the presence or absence of the provided credit in the audio signal.
  • the discriminator is a model that learns the association between a plurality of voice signals (“voice segments” described below) extracted from the learning voice signal and the presence or absence of provided credit.
  • FIG. 3 is a flowchart for explaining an example of a processing procedure of learning processing according to the first embodiment.
  • step S101 the learning data generation unit 11 extracts a positive example voice segment (a portion (partial voice signal) estimated to include a provided credit in the learning voice signal) from the learning voice signal.
  • the learning data generation unit 11 specifies the provided credit section in the learning voice signal based on the time data stored in the correct answer storage unit 121. There may be a plurality of provided credit sections.
  • the learning data generation unit 11 executes voice recognition for each of the identified provision credit sections of the learning voice signal, and generates a voice recognition result (text data) for each provision credit section.
  • FIG. 4 is a diagram showing an example of extraction of audio segments of the positive example in the first embodiment.
  • the part of "Learning sent by sponsor of viewing” in the learning audio signal corresponds to the provided credit section, and among them, "Looking", “Sponsor", “Providing”, “Sending” An example that is a related phrase is shown. Therefore, the audio signals for 3 seconds before and after centering on these related terms are extracted as the audio segment of the positive example.
  • the learning data generation unit 11 extracts a negative example voice segment from a random part other than the provided credit section in the learning voice signal (S102).
  • the length of the negative example voice segment is the same as the length of the positive example voice segment (6 seconds). Further, it is desirable that the number of negative example audio segments is the same as the number of positive example audio segments.
  • the learning unit 12 uses the positive example voice segment extracted in step S101 and the negative example voice segment extracted in step S102 to learn the discriminator for the provided credit section (S103). ..
  • the learning unit 12 frequency-analyzes each voice segment of the positive example or the negative example (for example, a window length of 25 ms, a window shift length of 10 ms), and performs 40 mel filter bank processing to obtain 600 ⁇ . Take 40 mel spectrograms.
  • the learning unit 12 uses, for each voice segment, the mel spectrogram acquired for the voice segment as an input feature amount, and whether or not the voice segment has a provided credit (whether or not the voice segment includes the provided credit).
  • the classifier for example, a convolutional neural network may be used, or another classifier such as SVM (support vector machine) may be used.
  • FIG. 5 is a diagram showing a model example of the discriminator in the first embodiment.
  • FIG. 5 shows an example using a convolutional neural network.
  • the learning unit 12 stores the learned parameters of the discriminator in the parameter storage unit 123 (S104).
  • FIG. 6 is a flowchart for explaining an example of the processing procedure of the provided credit detection processing according to the first embodiment.
  • the processing procedure of FIG. 6 is premised on that the processing procedure of FIG. 3 has been executed.
  • step S201 the detection data generation unit 13 outputs a window from an audio signal (hereinafter, referred to as a “detection audio signal”) of a TV broadcast for a provided credit detection (hereinafter, referred to as a “detection TV broadcast”).
  • FIG. 7 is a diagram showing an example of extracting a voice segment from the detection voice signal according to the first embodiment.
  • FIG. 7 shows an example in which a 6-second audio signal having a 1-second offset is extracted as an audio segment. Note that, in FIG. 7, an example of extracting the audio segment up to the middle of the audio signal for detection is shown for convenience, but the audio segment is extracted for all the audio signals for detection.
  • the provided credit section estimation unit 14 frequency-analyzes each voice segment extracted in step S201 (for example, window length 25 ms, window shift length 10 ms), and performs 40 mel filter bank processing, thereby 600
  • the x40 mel spectrogram is acquired as the feature amount of each audio segment (S202).
  • the provided credit section estimation unit 14 restores (generates) the discriminator learned by the processing procedure of FIG. 3 using the parameters stored in the parameter storage unit 123 (S203).
  • the provided credit section estimation unit 14 inputs the feature amount acquired in step S202 to the discriminator for each voice segment extracted in step S201, and determines whether or not there is a provided credit in each voice segment (each voice segment). It is determined whether or not the segment includes the provided credit) (S204). For example, the provided credit section estimation unit 14 determines that there is a provided credit for a voice segment whose output value of the discriminator is equal to or more than a predetermined threshold value, and provides the provided credit for a voice segment whose output value is smaller than the threshold value. None Determined as “0”. The provided credit section estimation unit 14 generates a binary time-series signal indicating the presence or absence of provided credit in time series by arranging the determination results in time-series order of the voice segments.
  • the provided credit section estimation unit 14 detects, in the binary time-series signal, a section in which the audio segment for which the provided credit display is present continues for a predetermined time or longer as the provided credit display section in which the provided credit is displayed ( Identification) (S205). Specifically, the provided credit section estimation unit 14 applies a median filter to the binary time series signal for the purpose of removing noise. In the time-series signal after the median filter processing, the provided credit section estimation unit 14 has a section in which a voice segment determined to have the provided credit display continues for a predetermined time or more (a signal “1” is a predetermined time or more (for example, a voice segment.
  • the provided credit section estimation unit 14 detects (specifies) the section from 5 minutes 00 seconds to 5 minutes 10 seconds as the provided credit display section.
  • the time information output unit 15 outputs the time information (start time and end time) of the detected provided credit display section (S206).
  • a TV broadcast audio signal is taken as an example, but for example, the first embodiment may be specified for specifying the section of the provided credit in the radio broadcast audio signal. Further, the first embodiment may be applied not only to the provided credits such as a specific commercial (CM) but also to the section of other credits. In this case, the words included in the specific CM may be stored in the related word storage unit 122 as the related words.
  • CM specific commercial
  • the second embodiment will be described.
  • the points different from the first embodiment will be described.
  • the points that are not particularly mentioned in the second embodiment may be the same as in the first embodiment.
  • FIG. 8 is a diagram showing a functional configuration example of the provided credit section identification device 10 according to the second embodiment. 8, the same parts as or the corresponding parts to those in FIG. 2 are designated by the same reference numerals, and the description thereof will be appropriately omitted.
  • a video signal of the TV broadcast for learning that is, a video signal corresponding to (synchronized with) the audio signal for learning.
  • a video signal for learning a video signal corresponding to (synchronized with) the audio signal for learning.
  • an audio signal audio signal for learning.
  • the time data (start time, end time) of the provided credit section is stored.
  • the parameter storage unit 123 stores the parameters of the discriminator that identifies the presence or absence of the provided credit for the pair of the video signal and the audio signal.
  • FIG. 9 is a flow chart for explaining an example of a processing procedure of learning processing in the second embodiment.
  • step S101a the learning data generation unit 11 extracts a positive example speech segment (a portion including a provided credit in the learning speech signal) from the learning speech signal, and at the same time, the still data corresponding to the time of the related phrase.
  • An image is extracted from the learning video signal. Therefore, a pair of a positive audio segment and a still image is extracted.
  • the method of extracting the audio segment of the positive example may be the same as that of the first embodiment.
  • a frame (still image) at the time of the related phrase in the positive example audio segment may be extracted from the learning video signal. Note that a plurality of frames (still images) may be extracted for one audio segment.
  • FIG. 10 is a diagram showing an example of extracting a pair of a positive audio segment and a still image according to the second embodiment.
  • the learning voice signal in FIG. 10 is the same as the learning voice signal in FIG. Therefore, in FIG. 10, the same audio segment as in FIG. 4 is extracted. However, in FIG. 10, a still image at the time when a related phrase appears in each audio segment is extracted from the learning video signal. Note that in FIG. 10, the positional relationship between each audio segment and the still image is irrelevant to the timing of the still image with respect to the audio segment.
  • the learning data generation unit 11 extracts a negative example audio segment from a portion other than the provided credit section in the learning audio signal, and a negative example of a still image corresponding to the center time of the audio segment in the learning video signal. Is extracted as a still image (S102a). Therefore, a pair of the negative audio segment and the still image is extracted.
  • the method of extracting the voice segment in the negative example may be the same as that in the first embodiment.
  • the learning unit 12 uses the pair of the positive example audio segment and still image extracted in step S101a and the negative example audio segment and still image pair extracted in step S102a to provide the provided credit.
  • the discriminator (the association between each pair and the presence or absence of the provided credit) is learned (S103a).
  • the learning unit 12 frequency-analyzes each voice segment of the positive example or the negative example (for example, a window length of 25 ms, a window shift length of 10 ms), and performs 40 mel filter bank processing to obtain 600 ⁇ . Take 40 mel spectrograms.
  • the learning unit 12 uses, for each voice segment, a pair of the mel spectrogram acquired for the voice segment and a still image corresponding to the voice segment as an input feature amount, and whether or not there is a provided credit for the pair (the relevant A discriminator that discriminates (detects) two classes is learned by determining whether or not the pair includes the provided credit.
  • the discriminator for example, a convolutional neural network may be used, or another discriminator such as SVM may be used.
  • FIG. 11 is a diagram showing a model example of the discriminator according to the second embodiment.
  • FIG. 11 shows an example using a convolutional neural network.
  • the learning unit 12 stores the learned parameter of the discriminator in the parameter storage unit 123 (S104a).
  • FIG. 12 is a flowchart for explaining an example of the processing procedure of the provided credit detection processing according to the second embodiment. 12, the same steps as those in FIG. 6 are designated by the same step numbers, and the description thereof will be omitted as appropriate.
  • the processing procedure of FIG. 12 is premised on that the processing procedure of FIG. 9 has been executed.
  • step S201a the detection data generation unit 13 extracts a voice segment from the detection voice signal with a window length of 2N seconds and a window shift length of 1 second, and at the same time, outputs a still image at the central time (third second) of each voice segment.
  • the video signal of the TV broadcast for detection that is, the video signal corresponding to (synchronized with) the audio signal for detection).
  • FIG. 13 is a diagram showing an example of extracting a voice segment and a still image from the detection voice signal according to the second embodiment.
  • an example is shown in which audio signals for 6 seconds having a shift of 1 second are extracted as audio segments, and a still image at the central time of each audio segment is extracted from the video signal for detection.
  • the feature amount (600 ⁇ 40 mel spectrogram) of each audio segment is acquired (S202).
  • the provided credit section estimation unit 14 restores (generates) the discriminator learned by the processing procedure of FIG. 9 using the parameters stored in the parameter storage unit 123 (S203a).
  • the provided credit section estimation unit 14 determines, for each pair of the audio segment and the still image extracted in step S201a, the pair of the feature amount and the still image acquired from the audio segment in step S202 as the identifier. Then, the presence or absence of the provided credit in each pair is determined (S204a). Note that the method of determining the presence/absence of the provided credit may be the same as that in the first embodiment. As a result, a binary time-series signal that time-sequentially shows the presence or absence of the provided credit is generated.
  • the subsequent steps (S205, S205) may be the same as those in the first embodiment.
  • FIG. 14 is a diagram showing an example of the evaluation result of each embodiment.
  • FIG. 14 shows the evaluation result (recall rate) when learning about one week's broadcasting of five terrestrial stations and identifying the section of the provided credit for broadcasting of five terrestrial stations in another week.
  • audio is a case where only an audio signal is used, that is, corresponds to the first embodiment
  • image+audio is a case where an audio signal and a video signal are used, That is, it corresponds to the second embodiment.
  • the provided credit section identifying device 10 is an example of the credit section identifying device.
  • the detection data generation unit 13 is an example of an extraction unit.
  • the provided credit section estimation unit 14 is an example of a specification unit.
  • the detection audio signal is an example of the first audio signal.
  • the audio segment extracted from the detection audio signal is an example of the first partial audio signal.
  • the learning voice signal is an example of the second voice signal.
  • the audio segment extracted from the learning audio signal is an example of the second partial audio signal.
  • the detection video signal is an example of the first video signal.
  • the still image extracted from the detection video signal is an example of the first still image.
  • the learning video signal is an example of the second video signal.
  • the still image extracted from the learning video signal is an example of the second still image.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Business, Economics & Management (AREA)
  • Health & Medical Sciences (AREA)
  • Signal Processing (AREA)
  • Strategic Management (AREA)
  • Finance (AREA)
  • Development Economics (AREA)
  • Accounting & Taxation (AREA)
  • Human Computer Interaction (AREA)
  • Computational Linguistics (AREA)
  • Acoustics & Sound (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Databases & Information Systems (AREA)
  • Game Theory and Decision Science (AREA)
  • General Business, Economics & Management (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Entrepreneurship & Innovation (AREA)
  • Economics (AREA)
  • Marketing (AREA)
  • Biomedical Technology (AREA)
  • General Health & Medical Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Molecular Biology (AREA)
  • Artificial Intelligence (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Evolutionary Computation (AREA)
  • Biodiversity & Conservation Biology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
  • Testing, Inspecting, Measuring Of Stereoscopic Televisions And Televisions (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
  • Image Analysis (AREA)

Abstract

クレジット区間特定装置は、第1の音声信号から、それぞれが前記第1の音声信号の一部であり、相互に時間方向にずれを有する複数の第1の部分音声信号を抽出する抽出部と、前記各第1の部分音声信号にクレジットが含まれるか否かを、第2の音声信号から抽出される各第2の部分音声信号とクレジットの有無との関連付けに基づいて判定することで、前記第1の音声信号におけるクレジットの区間を特定する特定部と、を有することで、クレジットの区間の特定を効率化する。

Description

クレジット区間特定装置、クレジット区間特定方法及びプログラム
 本発明は、クレジット区間特定装置、クレジット区間特定方法及びプログラムに関する。
 従来、テレビ放送等について、いずれの企業がいずれの番組のスポンサーであるかを調査することに経済的な価値が認められている。
 このような調査は、テレビ放送等における提供クレジットの表示を目視で見つけ出し、当該提供クレジットから企業名を書き起こすことで行われている。なお、提供クレジットとは、放送番組のスポンサーのロゴの表示やナレーション(例えば、「この番組は、XXXとご覧のスポンサーの提供でお送りします」等)をいう。
[online]、インターネット<URL:http://www.jppanet.or.jp/documents/video.html>
 しかしながら、提供クレジットが表示される区間は、放送時間の約1%程度に過ぎない。したがって、上記のような調査においては提供クレジットの区間を特定するためのテレビ放送等の視聴作業に多くの時間が費やされてしまう。
 なお、上記では、説明の便宜上、提供クレジットを例として記載したが、例えば、特定のコマーシャル等、提供クレジットだけでなく、他のクレジットの区間を特定したい場合にも同様の課題が生じる。
 本発明は、上記の点に鑑みてなされたものであって、クレジットの区間の特定を効率化することを目的とする。
 そこで上記課題を解決するため、クレジット区間特定装置は、第1の音声信号から、それぞれが前記第1の音声信号の一部であり、相互に時間方向にずれを有する複数の第1の部分音声信号を抽出する抽出部と、前記各第1の部分音声信号にクレジットが含まれるか否かを、第2の音声信号から抽出される各第2の部分音声信号とクレジットの有無との関連付けに基づいて判定することで、前記第1の音声信号におけるクレジットの区間を特定する特定部と、を有する。
 クレジットの区間の特定を効率化することができる。
第1の実施の形態における提供クレジット区間特定装置10のハードウェア構成例を示す図である。 第1の実施の形態における提供クレジット区間特定装置10の機能構成例を示す図である。 第1の実施の形態における学習処理の処理手順の一例を説明するためのフローチャートである。 第1の実施の形態における正例の音声セグメントの抽出例を示す図である。 第1の実施の形態における識別器のモデル例を示す図である。 第1の実施の形態における提供クレジットの検出処理の処理手順の一例を説明するためのフローチャートである。 第1の実施の形態における検出用音声信号からの音声セグメントの抽出例を示す図である。 第2の実施の形態における提供クレジット区間特定装置10の機能構成例を示す図である。 第2の実施の形態における学習処理の処理手順の一例を説明するためのフローチャートである。 第2の実施の形態における正例の音声セグメント及び静止画のペアの抽出例を示す図である。 第2の実施の形態における識別器のモデル例を示す図である。 第2の実施の形態における提供クレジットの検出処理の処理手順の一例を説明するためのフローチャートである。 第2の実施の形態における検出用音声信号からの音声セグメントの抽出例を示す図である。 各実施形態の評価結果の一例を示す図である。
 以下、図面に基づいて本発明の実施の形態を説明する。図1は、第1の実施の形態における提供クレジット区間特定装置10のハードウェア構成例を示す図である。図1の提供クレジット区間特定装置10は、それぞれバスBで相互に接続されているドライブ装置100、補助記憶装置102、メモリ装置103、CPU104、及びインタフェース装置105等を有するコンピュータである。
 提供クレジット区間特定装置10での処理を実現するプログラムは、CD-ROM等の記録媒体101によって提供される。プログラムを記憶した記録媒体101がドライブ装置100にセットされると、プログラムが記録媒体101からドライブ装置100を介して補助記憶装置102にインストールされる。但し、プログラムのインストールは必ずしも記録媒体101より行う必要はなく、ネットワークを介して他のコンピュータよりダウンロードするようにしてもよい。補助記憶装置102は、インストールされたプログラムを格納すると共に、必要なファイルやデータ等を格納する。
 メモリ装置103は、プログラムの起動指示があった場合に、補助記憶装置102からプログラムを読み出して格納する。CPU104は、メモリ装置103に格納されたプログラムに従って提供クレジット区間特定装置10に係る機能を実行する。インタフェース装置105は、ネットワークに接続するためのインタフェースとして用いられる。
 図2は、第1の実施の形態における提供クレジット区間特定装置10の機能構成例を示す図である。図2において、提供クレジット区間特定装置10は、学習データ生成部11、学習部12、検出用データ生成部13、提供クレジット区間推定部14及び時刻情報出力部15等を有する。これら各部は、提供クレジット区間特定装置10にインストールされた1以上のプログラムが、CPU104に実行させる処理により実現される。提供クレジット区間特定装置10は、また、正解記憶部121、関連語句記憶部122及びパラメータ記憶部123等を利用する。これら各記憶部は、例えば、補助記憶装置102、又は提供クレジット区間特定装置10にネットワークを介して接続可能な記憶装置等を用いて実現可能である。
 正解記憶部121には、或る期間に放送された学習用のTV放送(以下、「学習用TV放送」という。)の音声信号(以下、「学習用音声信号」という。)について、提供クレジットの区間(以下、「提供クレジット区間」という。)を示す時刻データ(開始時刻、終了時刻)が記憶されている。なお、提供クレジット区間は、例えば、予めユーザによる目視等によって確認されてもよい。
 関連語句記憶部122には、提供クレジットの表示時のアナウンス(提供クレジット表示の際に流れるアナウンス)に含まれ、提供クレジット表示に関連する関連語句が記憶されている。関連語句の一例として、「ご覧の」、「スポンサー」、「提供」、「お送り(お送りします/お送りしました)」等の語句が挙げられる。また、企業名を示す語句等が関連語句とされてもよい。なお、関連語句は、例えば、予めユーザにより設定される。
 パラメータ記憶部123には、音声信号における提供クレジットの有無を識別する識別器のパラメータが記憶される。識別器は、学習用音声信号から抽出される複数の音声信号(後述の「音声セグメント」)と提供クレジット有無との関連付けを学習したモデルである。
 以下、提供クレジット区間特定装置10が実行する処理手順について説明する。図3は、第1の実施の形態における学習処理の処理手順の一例を説明するためのフローチャートである。
 ステップS101において、学習データ生成部11は、学習用音声信号から、正例の音声セグメント(学習用音声信号において提供クレジットを含むと推定される部分(部分音声信号))を抽出する。
 具体的には、学習データ生成部11は、正解記憶部121に記憶されている時刻データに基づいて、学習用音声信号における提供クレジット区間を特定する。なお、提供クレジット区間は、複数有ってもよい。学習データ生成部11は、学習用音声信号のうち、特定した各提供クレジット区間を対象として音声認識を実行し、提供クレジット区間ごとに音声認識結果(テキストデータ)を生成する。学習データ生成部11は、各テキストデータについて、関連語句記憶部122に記憶されているいずれかの関連語句を含む部分を特定し、学習用音声信号において当該部分に対応する音声信号を正例の音声セグメントとして抽出する。例えば、関連語句を中心とした前後N秒間の部分が正例の音声セグメントとして抽出される。本実施の形態では、N=3とする。但し、Nは、他の値であってもよい。
 図4は、第1の実施の形態における正例の音声セグメントの抽出例を示す図である。図4では、学習用音声信号のうち、「ご覧のスポンサーの提供でお送りしました」の部分が提供クレジット区間に対応し、このうち「ご覧」、「スポンサー」、「提供」、「送り」が関連語句である例が示されている。したがって、これらの関連語句を中心とした前後3秒間の音声信号が正例の音声セグメントとして抽出されている。
 続いて、学習データ生成部11は、学習用音声信号における提供クレジット区間以外のランダムな部分から、負例の音声セグメントを抽出する(S102)。負例の音声セグメントの長さは正例の音声セグメントの長さ(6秒間)と同じである。また、負例の音声セグメントの個数は、正例の音声セグメントの個数と同数であるのが望ましい。
 続いて、学習部12は、ステップS101において抽出された正例の音声セグメントと、ステップS102において抽出された負例の音声セグメントとを用いて、提供クレジット区間に関する識別器の学習を行う(S103)。
 具体的には、学習部12は、正例又は負例の各音声セグメントを周波数分析し(例えば、窓長25ms、窓シフト長10ms)、40個のメルフィルタバンク処理を施すことで、600×40のメルスペクトログラムを取得する。学習部12は、音声セグメントごとに、当該音声セグメントに関して取得されたメルスペクトログラムを入力特徴量として、当該音声セグメントに提供クレジットが有るか無いか(当該音声セグメントに提供クレジットが含まれるか否か)を2クラス識別(検出)する識別器を学習する。すなわち、正例の音声セグメントについては、提供クレジットが有ることが学習され、負例の音声セグメントについては、提供クレジットが無いことが学習される。識別器としては、例えば、畳み込みニューラルネットワークが利用されてもよいし、SVM(support vector machine)などの他の識別器が利用されてもよい。
 図5は、第1の実施の形態における識別器のモデル例を示す図である。図5には、畳み込みニューラルネットワークを利用した例が示されている。
 続いて、学習部12は、学習された識別器のパラメータをパラメータ記憶部123に記憶する(S104)。
 図6は、第1の実施の形態における提供クレジットの検出処理の処理手順の一例を説明するためのフローチャートである。図6の処理手順は、図3の処理手順が実行済みであることが前提となる。
 ステップS201において、検出用データ生成部13は、提供クレジットの検出用のTV放送(以下、「検出用TV放送」という。)の音声信号(以下、「検出用音声信号」という。)から、窓長2N秒、窓シフト長1秒で音声セグメントを抽出する。本実施の形態においてN=3であるため、1秒ずつずれた(相互に時間方向にずれを有する)6秒間の複数の音声セグメントが抽出される。
 図7は、第1の実施の形態における検出用音声信号からの音声セグメントの抽出例を示す図である。図7では、1秒ずつずれを有する6秒間の音声信号が音声セグメントとして抽出される例が示されている。なお、図7では、便宜上、検出用音声信号の途中までの音声セグメントの抽出例が示されているが、検出用音声信号の全部について、音声セグメントの抽出が行われる。
 続いて、提供クレジット区間推定部14は、ステップS201において抽出された各音声セグメントを周波数分析し(例えば、窓長25ms、窓シフト長10ms)、40個のメルフィルタバンク処理を施すことで、600×40のメルスペクトログラムを各音声セグメントの特徴量として取得する(S202)。
 続いて、提供クレジット区間推定部14は、パラメータ記憶部123に記憶されているパラメータを用いて、図3の処理手順によって学習された識別器を復元(生成)する(S203)。
 続いて、提供クレジット区間推定部14は、ステップS201において抽出された音声セグメントごとに、ステップS202において取得された特徴量を当該識別器に入力して、各音声セグメントにおける提供クレジットの有無(各音声セグメントに提供クレジットが含まれるか否か)を判定する(S204)。例えば、提供クレジット区間推定部14は、識別器の出力値が所定の閾値以上である音声セグメントについては提供クレジット有り「1」と判定し、当該出力値が閾値よりも小さい音声セグメントについては提供クレジット無し「0」と判定する。提供クレジット区間推定部14は、判定結果を音声セグメントの時系列順に配列することで、提供クレジットの有無を時系列的に示すバイナリ時系列信号を生成する。
 続いて、提供クレジット区間推定部14は、当該バイナリ時系列信号において、提供クレジット表示ありと判定された音声セグメントが所定時間以上連続する区間を、提供クレジットが表示された提供クレジット表示区間として検出(特定)する(S205)。具体的には、提供クレジット区間推定部14は、ノイズ除去を目的として、バイナリ時系列信号に対して中央値フィルタを適用する。提供クレジット区間推定部14は、中央値フィルタ処理後の時系列信号において、提供クレジット表示有りと判定された音声セグメントが所定時間以上連続する区間(信号「1」が所定時間以上(例えば、音声セグメントの長さ(6秒)×M以上(M≧2))連続して並ぶ区間)を、提供クレジット表示区間として検出(特定)する。本実施の形態のように、音声セグメントが1秒間隔で(すなわち、1秒のずれを有するように)作成された場合、例えば、300番目から310番目に信号「1」が連続して並んでいれば、提供クレジット区間推定部14は、5分00秒から5分10秒の区間を提供クレジット表示区間として検出(特定)する。
 続いて、時刻情報出力部15は、検出され提供クレジット表示区間の時刻情報(開始時刻及び終了時刻)を出力する(S206)。
 なお、上記では、TV放送の音声信号を例として説明したが、例えば、ラジオ放送の音声信号における提供クレジットの区間の特定について第1の実施の形態が特定されてもよい。また、特定のコマーシャル(CM)等、提供クレジットだけでなく、他のクレジットの区間の特定について第1の実施の形態が適用されてもよい。この場合、特定のCMに含まれている語句が、関連語句として関連語句記憶部122に記憶されればよい。
 上述したように、第1の実施の形態によれば、クレジットの区間の特定を効率化することができる。
 次に、第2の実施の形態について説明する。第2の実施の形態では第1の実施の形態と異なる点について説明する。第2の実施の形態において特に言及されない点については、第1の実施の形態と同様でもよい。
 図8は、第2の実施の形態における提供クレジット区間特定装置10の機能構成例を示す図である。図8において、図2と同一部分又は対応する部分には同一符号を付し、その説明は適宜省略する。
 正解記憶部121には、学習用TV放送の映像信号(すなわち、学習用音声信号に対応する(同期した)映像信号。以下、「学習用映像信号」という)及び音声信号(学習用音声信号)に対して、提供クレジット区間の時刻データ(開始時刻、終了時刻)が記憶されている。
 パラメータ記憶部123には、映像信号及び音声信号のペアについて、提供クレジットの有無を識別する識別器のパラメータが記憶される。
 図9は、第2の実施の形態における学習処理の処理手順の一例を説明するためのフローチャートである。
 ステップS101aにおいて、学習データ生成部11は、正例の音声セグメント(学習用音声信号において提供クレジットを含む部分)を学習用音声信号から抽出すると共に、当該音声セグメントにおいて関連語句の時刻に対応する静止画を学習用映像信号から抽出する。したがって、正例の音声セグメントと静止画のペアが抽出される。正例の音声セグメントの抽出方法は第1の実施の形態と同様でよい。正例の静止画としては、学習用映像信号において、正例の音声セグメントにおける関連語句の時刻のフレーム(静止画)が抽出されればよい。なお、1つの音声セグメントに対して複数のフレーム(静止画)が抽出されてもよい。
 図10は、第2の実施の形態における正例の音声セグメント及び静止画のペアの抽出例を示す図である。図10における学習用音声信号は、図4における学習用音声信号と同じである。したがって、図10では、図4と同じ音声セグメントが抽出されている。但し、図10では、各音声セグメントにおいて関連語句の出現する時刻における静止画が学習用映像信号から抽出されている。なお、図10において、各音声セグメントと静止画との位置関係は、当該音声セグメントに対する当該静止画のタイミングとは無関係である。
 続いて、学習データ生成部11は、学習用音声信号における提供クレジット区間以外の部分から負例の音声セグメントを抽出し、学習用映像信号において当該音声セグメントの中心時刻に対応する静止画を負例の静止画として抽出する(S102a)。したがって、負例の音声セグメントと静止画とのペアが抽出される。なお、負例の音声セグメントの抽出方法は、第1の実施の形態と同様でよい。
 続いて、学習部12は、ステップS101aにおいて抽出された正例の音声セグメント及び静止画のペアと、ステップS102aにおいて抽出された負例の音声セグメント及び静止画のペアとを用いて、提供クレジットに関する識別器(これら各ペアと提供クレジットの有無との関連付け)の学習を行う(S103a)。
 具体的には、学習部12は、正例又は負例の各音声セグメントを周波数分析し(例えば、窓長25ms、窓シフト長10ms)、40個のメルフィルタバンク処理を施すことで、600×40のメルスペクトログラムを取得する。学習部12は、音声セグメントごとに、当該音声セグメントに関して取得されたメルスペクトログラムと、当該音声セグメントに対応する静止画とのペアを入力特徴量として、当該ペアに提供クレジットが有るか無いか(当該ペアに提供クレジットが含まれているか否か)を2クラス識別(検出)する識別器を学習する。識別器としては、例えば、畳み込みニューラルネットワークが利用されてもよいし、SVMなどの他の識別器が利用されてもよい。
 図11は、第2の実施の形態における識別器のモデル例を示す図である。図11には、畳み込みニューラルネットワークを利用した例が示されている。
 続いて、学習部12は、学習された識別器のパラメータをパラメータ記憶部123に記憶する(S104a)。
 図12は、第2の実施の形態における提供クレジットの検出処理の処理手順の一例を説明するためのフローチャートである。図12中、図6と同一ステップには同一ステップ番号を付し、その説明は適宜省略する。図12の処理手順は、図9の処理手順が実行済みであることが前提となる。
 ステップS201aにおいて、検出用データ生成部13は、窓長2N秒、窓シフト長1秒で音声セグメントを検出用音声信号から抽出すると共に、各音声セグメントの中心時刻(3秒目)の静止画を、検出用TV放送の映像信号(すなわち、検出用音声信号に対応する(同期した)映像信号)から抽出する。
 図13は、第2の実施の形態における検出用音声信号からの音声セグメント及び静止画の抽出例を示す図である。図13では、1秒ずつずれを有する6秒間の音声信号が音声セグメントとして抽出され、各音声セグメントの中心時刻における静止画が検出用映像信号から抽出される例が示されている。
 続いて、第1の実施の形態と同様に、各音声セグメントの特徴量(600×40のメルスペクトログラム)が取得される(S202)。
 続いて、提供クレジット区間推定部14は、パラメータ記憶部123に記憶されているパラメータを用いて、図9の処理手順によって学習された識別器を復元(生成)する(S203a)。
 続いて、提供クレジット区間推定部14は、ステップS201aにおいて抽出された音声セグメント及び静止画のペアごとに、当該音声セグメントからステップS202において取得された特徴量と当該静止画とのペアを当該識別器に入力して、各ペアにおける提供クレジットの有無を判定する(S204a)。なお、提供クレジットの有無の判定方法は、第1の実施の形態と同様でよい。その結果、提供クレジットの有無を時系列的に示すバイナリ時系列信号が生成される。
 以降(S205、S205)は、第1の実施の形態と同様でよい。
 図14は、各実施形態の評価結果の一例を示す図である。図14には、地上波5局の1週間分の放送について学習し、別の1週間における地上波5局の放送について提供クレジットの区間を特定した際の評価結果(再現率)が示されている。ここで、再現率とは、正解の区間(提供クレジットが実際に表示された区間)に対して、提供クレジット区間特定装置10が、提供クレジットの区間であると判定した区間の割合をいう。例えば、放送の開始から11秒目から20秒目10秒間が正解の区間である場合に、12秒目から20秒目の9秒間が提供クレジットの区間として特定された場合には、再現率は9÷10=0.9となる。
 また、図14の横軸において「音声」は、音声信号のみを利用した場合、すなわち、第1の実施の形態に対応し、「画像+音声」は、音声信号と映像信号を利用した場合、すなわち、第2の実施の形態に対応する。
 図14によれば、「音声」及び「画像+音声」のいずれについても高い再現率が得られている。また、「音声」の場合よりも「画像+音声」の方が、高い再現率が得られていることが分かる。このことから、第2の実施の形態によれば、第1の実施の形態よりも高精度に提供クレジットの区間を特定できることが分かる。
 なお、上記各実施の形態は、インターネット等において配信される動画におけるクレジットの区間の特定に適用されてもよい。
 なお、上記各実施の形態において、提供クレジット区間特定装置10は、クレジット区間特定装置の一例である。検出用データ生成部13は、抽出部の一例である。提供クレジット区間推定部14は、特定部の一例である。検出用音声信号は、第1の音声信号の一例である。検出用音声信号から抽出される音声セグメントは、第1の部分音声信号の一例である。学習用音声信号は、第2の音声信号の一例である。学習用音声信号から抽出される音声セグメントは、第2の部分音声信号の一例である。検出用映像信号は、第1の映像信号の一例である。検出用映像信号から抽出される静止画は、第1の静止画の一例である。学習用映像信号は、第2の映像信号の一例である。学習用映像信号から抽出される静止画は、第2の静止画の一例である。
 以上、本発明の実施の形態について詳述したが、本発明は斯かる特定の実施形態に限定されるものではなく、請求の範囲に記載された本発明の要旨の範囲内において、種々の変形・変更が可能である。
10     提供クレジット区間特定装置
11     学習データ生成部
12     学習部
13     検出用データ生成部
14     提供クレジット区間推定部
15     時刻情報出力部
100    ドライブ装置
101    記録媒体
102    補助記憶装置
103    メモリ装置
104    CPU
105    インタフェース装置
121    正解記憶部
122    関連語句記憶部
123    パラメータ記憶部
B      バス

Claims (6)

  1.  第1の音声信号から、それぞれが前記第1の音声信号の一部であり、相互に時間方向にずれを有する複数の第1の部分音声信号を抽出する抽出部と、
     前記各第1の部分音声信号にクレジットが含まれるか否かを、第2の音声信号から抽出される各第2の部分音声信号とクレジットの有無との関連付けに基づいて判定することで、前記第1の音声信号におけるクレジットの区間を特定する特定部と、
    を有することを特徴とするクレジット区間特定装置。
  2.  前記第2の部分音声信号は、予め設定された語句を含む音声信号であり、
     前記第2の部分音声信号が前記語句を含むか否かは、当該第2の部分音声信号を対象とした音声認識に基づき判定される、
    ことを特徴とする請求項1記載のクレジット区間特定装置。
  3.  前記特定部は、前記各第2の部分音声信号とクレジットの有無とを学習した識別器を用いて、前記各第1の部分音声信号にクレジットが含まれるか否かを判定する、
    ことを特徴とする請求項1又は2記載のクレジット区間特定装置。
  4.  前記抽出部は、前記第1の音声信号に対応する第1の映像信号から、前記各第1の部分音声信号に対応する複数の第1の静止画を抽出し、
     前記特定部は、前記第1の部分音声信号及び前記第1の静止画の各ペアにクレジットが含まれるか否かを、前記各第2の部分音声信号と、前記第2の音声信号に対応する第2の映像信号から抽出される、前記各第2の部分音声信号に対応する第2の静止画とクレジットの有無との関連付けに基づいて判定することで、前記第1の音声信号及び前記第1の映像信号におけるクレジットの区間を特定する、
    ことを特徴とする請求項1乃至3いずれか一項記載のクレジット区間特定装置。
  5.  第1の音声信号から、それぞれが前記第1の音声信号の一部であり、相互に時間方向にずれを有する複数の第1の部分音声信号を抽出する抽出手順と、
     前記各第1の部分音声信号にクレジットが含まれるか否かを、第2の音声信号から抽出される各第2の部分音声信号とクレジットの有無との関連付けに基づいて判定することで、前記第1の音声信号におけるクレジットの区間を特定する特定手順と、
    をコンピュータが実行することを特徴とするクレジット区間特定方法。
  6.  請求項1乃至4いずれか一項記載のクレジット区間特定装置としてコンピュータを機能させることを特徴とするプログラム。
PCT/JP2020/002458 2019-02-07 2020-01-24 クレジット区間特定装置、クレジット区間特定方法及びプログラム Ceased WO2020162220A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US17/428,612 US12494222B2 (en) 2019-02-07 2020-01-24 Sponsorship credit period identification apparatus, sponsorship credit period identification method and program

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2019-020322 2019-02-07
JP2019020322A JP7196656B2 (ja) 2019-02-07 2019-02-07 クレジット区間特定装置、クレジット区間特定方法及びプログラム

Publications (1)

Publication Number Publication Date
WO2020162220A1 true WO2020162220A1 (ja) 2020-08-13

Family

ID=71947176

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2020/002458 Ceased WO2020162220A1 (ja) 2019-02-07 2020-01-24 クレジット区間特定装置、クレジット区間特定方法及びプログラム

Country Status (3)

Country Link
US (1) US12494222B2 (ja)
JP (1) JP7196656B2 (ja)
WO (1) WO2020162220A1 (ja)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008050718A1 (fr) * 2006-10-26 2008-05-02 Nec Corporation Dispositif d'extraction d'informations de droit, procédé d'extraction d'informations de droit et programme
JP2008108166A (ja) * 2006-10-27 2008-05-08 Matsushita Electric Ind Co Ltd 楽曲選択装置、楽曲選択方法

Family Cites Families (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2003012744A1 (en) * 2001-08-02 2003-02-13 Intellocity Usa, Inc. Post production visual alterations
US20040062520A1 (en) * 2002-09-27 2004-04-01 Koninklijke Philips Electronics N.V. Enhanced commercial detection through fusion of video and audio signatures
US8930561B2 (en) * 2003-09-15 2015-01-06 Sony Computer Entertainment America Llc Addition of supplemental multimedia content and interactive capability at the client
JP5581309B2 (ja) * 2008-03-24 2014-08-27 スー カン,ミン 放送サービスシステムの情報処理方法、その情報処理方法を実施する放送サービスシステム及びその情報処理方法に関する記録媒体
US8805689B2 (en) * 2008-04-11 2014-08-12 The Nielsen Company (Us), Llc Methods and apparatus to generate and use content-aware watermarks
US20160073148A1 (en) * 2014-09-09 2016-03-10 Verance Corporation Media customization based on environmental sensing
US9973813B2 (en) * 2015-01-23 2018-05-15 DISH Technologies L.L.C. Commercial-free audiovisual content
US9990350B2 (en) * 2015-11-02 2018-06-05 Microsoft Technology Licensing, Llc Videos associated with cells in spreadsheets
US20180176645A1 (en) * 2016-12-15 2018-06-21 Arris Enterprises Llc Method for providing feedback for television advertisements

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008050718A1 (fr) * 2006-10-26 2008-05-02 Nec Corporation Dispositif d'extraction d'informations de droit, procédé d'extraction d'informations de droit et programme
JP2008108166A (ja) * 2006-10-27 2008-05-08 Matsushita Electric Ind Co Ltd 楽曲選択装置、楽曲選択方法

Also Published As

Publication number Publication date
US20220115031A1 (en) 2022-04-14
JP2020129165A (ja) 2020-08-27
JP7196656B2 (ja) 2022-12-27
US12494222B2 (en) 2025-12-09

Similar Documents

Publication Publication Date Title
US6332122B1 (en) Transcription system for multiple speakers, using and establishing identification
CN109473123B (zh) 语音活动检测方法及装置
CN112437337B (zh) 一种直播实时字幕的实现方法、系统及设备
US20170169827A1 (en) Multimodal speech recognition for real-time video audio-based display indicia application
CN114598933B (zh) 一种视频内容处理方法、系统、终端及存储介质
JP7022782B2 (ja) コンテンツ特徴に基づいたトリガ機能を有するコンピューティングシステム
CN112423081B (zh) 一种视频数据处理方法、装置、设备及可读存储介质
CN113035199B (zh) 音频处理方法、装置、设备及可读存储介质
JP7691055B2 (ja) データ処理方法、装置、電子機器および記憶媒体
US20100042412A1 (en) Skipping radio/television program segments
US20230216598A1 (en) Detection device
CN110072140A (zh) 一种视频信息提示方法、装置、设备及存储介质
JP6966705B2 (ja) Cm区間検出装置、cm区間検出方法、及びプログラム
Tapu et al. Dynamic subtitles: A multimodal video accessibility enhancement dedicated to deaf and hearing impaired users
WO2020162220A1 (ja) クレジット区間特定装置、クレジット区間特定方法及びプログラム
US11645845B2 (en) Device and method for detecting display of provided credit, and program
CN115659211B (zh) 用户画像生成方法、装置、电子设备和计算机可读介质
US11727446B2 (en) Device and method for detecting display of provided credit, and program
CN113206996B (zh) 一种业务录制数据的质检方法及装置
US12417630B1 (en) Connecting computing devices presenting information relating to the same or similar topics
WO2019235406A1 (ja) Cm情報生成装置、cm情報生成方法、及びプログラム
JP2005150943A (ja) 動画像話題分割点決定装置
CN118629407A (zh) 语音识别方法、装置、电子设备和存储介质
CN121217880A (zh) 一种字幕显示方法及相关装置
WO2026014265A1 (ja) 指導データ生成装置、指導データ生成方法、及びプログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20753248

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20753248

Country of ref document: EP

Kind code of ref document: A1

WWG Wipo information: grant in national office

Ref document number: 17428612

Country of ref document: US