WO2018042959A1 - 動画データ解析装置及び動画データ解析方法 - Google Patents
動画データ解析装置及び動画データ解析方法 Download PDFInfo
- Publication number
- WO2018042959A1 WO2018042959A1 PCT/JP2017/027139 JP2017027139W WO2018042959A1 WO 2018042959 A1 WO2018042959 A1 WO 2018042959A1 JP 2017027139 W JP2017027139 W JP 2017027139W WO 2018042959 A1 WO2018042959 A1 WO 2018042959A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image data
- moving image
- data
- scene
- unit
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N5/00—Details of television systems
- H04N5/76—Television signal recording
- H04N5/91—Television signal processing therefor
- H04N5/93—Regeneration of the television signal or of selected parts thereof
Definitions
- the present invention relates to a moving image data analysis apparatus and a moving image data analysis method, and more particularly to a moving image data analysis apparatus and a moving image data analysis method for analyzing a moving image.
- moving image data analysis apparatus that analyzes the contents of moving image data (moving image data, moving image data).
- moving image data moving image data
- Patent Document 1 uncompressed or compressed moving image data is classified into various classes at low cost and with high accuracy by using features as a moving image and audio features attached to the moving image as necessary.
- An apparatus for classifying moving image data to be classified is disclosed.
- Patent Document 1 is a technique for classifying into several classes based on image feature values such as color arrangement descriptors in moving image data and audio features. For this reason, it was not possible to estimate a scene that is a specific scene or the like having the same content in the moving image data.
- This invention is made in view of such a situation, and makes it a subject to provide the moving image data analyzer which eliminates the above-mentioned problem.
- the moving image data analysis apparatus of the present invention is a moving image data analysis device that analyzes moving image data, and includes a still image acquisition unit that acquires still image data corresponding to time series from the moving image data, and the still image acquisition unit.
- a category recognition unit that recognizes an object included in the acquired still image data, recognizes a category of the recognized object, and creates recognition result time-series data; and the recognition recognized by the category recognition unit And a scene estimation unit for estimating a scene using the result time-series data.
- the moving image data analysis method of the present invention is a moving image data analysis method executed by a moving image data analysis device, and the moving image data analysis device acquires still image data corresponding to time series from the moving image data, and is acquired. Recognize an object included in the still image data, recognize a category of the recognized object, create recognition result time-series data, and use the recognition result time-series data to estimate a scene in a specific section It is characterized by performing.
- still image data is acquired from moving image data, an object included in the still image data is recognized, a category of the recognized object is recognized, and recognition result time-series data is created,
- an image data analysis apparatus capable of estimating a scene from a moving image can be provided.
- FIG. 1 is a system configuration diagram of a moving image data analysis apparatus according to an embodiment of the present invention. It is a block diagram which shows the control structure of the moving image data analyzer shown in FIG. It is a flowchart of the moving image data analysis process which concerns on embodiment of this invention. It is a conceptual diagram of the category recognition process shown in FIG. It is a conceptual diagram of the scene aggregation process and title provision process shown in FIG. It is a conceptual diagram of the scene composition process shown in FIG. It is a conceptual diagram of the scene composition process shown in FIG.
- the moving image data analyzing apparatus 1 is an information processing apparatus such as a server that analyzes moving image data 200, a PC (Personal Computer), a NAS (Network Attached Storage), a video recorder, a video distribution server, a portable terminal, a smartphone, and an image forming apparatus. is there.
- the moving image data analysis apparatus 1 includes a control unit 10, an I / O unit 15, a storage unit 19, and the like. Each unit is connected to the control unit 10 and controlled in operation by the control unit 10.
- the control unit 10 includes a general purpose processor (GPP), a central processing unit (CPU), a micro processing unit (MPU), a digital signal processor (DSP), a graphics processing unit (GPU), an application specific processor (ASIC). Information processing means such as a processor for a specific application).
- the control unit 10 reads out a control program stored in the ROM or HDD of the storage unit 19, develops the control program in the RAM, and executes it to operate as each unit of a functional block described later.
- the control unit 10 controls the entire apparatus in accordance with predetermined instruction information input from an external terminal (not shown), a connected keyboard, mouse, or other input unit.
- the I / O unit 15 is a so-called chip set or the like, and is a circuit or the like that includes various interfaces and transmits and receives data.
- the I / O unit 15 may include, for example, a LAN (Local Area Network) board, a wireless transceiver, a USB (Universal Serial Bus) interface, and the like.
- the I / O unit 15 can also transmit and receive data to and from devices such as a network camera and NAS (Network Attached Storage).
- the storage unit 19 is a non-temporary recording medium such as a ROM (Read Only Memory), a RAM (Random Access Memory), a semiconductor memory such as a flash memory, or an HDD (Hard Disk Drive).
- the ROM and HDD of the storage unit 19 store a control program for controlling the operation of the moving image data analysis apparatus 1 in addition to various data including moving image data 200 (FIG. 2) described later.
- the storage unit 19 also stores moving image data analysis setting data (not shown) in which user instructions, scheduling, and the like regarding how to analyze the moving image data 200 are specified.
- the storage unit 19 may be managed in an area divided for each device, for each partition, for each user, and the like.
- control unit 10 and the I / O unit 15 may be integrally formed as a so-called SOC (System On Chip), GPU built-in CPU, or the like.
- the control unit 10 and the I / O unit 15 may incorporate a non-temporary storage medium such as a RAM, a ROM, or a flash memory.
- the control unit 10 of the moving image data analysis apparatus 1 includes a still image acquisition unit 100, a category recognition unit 110, a speech recognition addition unit 120, a scene estimation unit 130, and a scene synthesis unit 140.
- the storage unit 19 stores moving image data 200, recognition result time series data 210, scene estimation data 220, title data 230, and composite scene data 240.
- the still image acquisition unit 100 acquires still image data 300 corresponding to time series from the moving image data 200. Specifically, the still image acquisition unit 100 extracts, for example, image data included in the moving image data 200 in units of frames in time-series order, and as still image data 300 such as bitmap data. Store in the storage unit 19. At this time, when the moving image data 200 is stored in a format such as a combination of a reference image and a difference image, the still image acquisition unit 100 draws a still image of the frame from the reference image and the difference image. Then, the still image data 300 may be created. Further, when the moving image data 200 includes object data, the still image acquisition unit 100 may include this in the still image data 300.
- the category recognition unit 110 recognizes an object included in the still image data 300 acquired by the still image acquisition unit 100, recognizes a category that is the type of the recognized object, and the recognition results are displayed in time series.
- the recorded recognition result time series data 210 is created.
- the category recognition unit 110 recognizes the category of the object using, for example, a convolutional neural network.
- the recognition result of this category may be expressed by a character string such as an object name.
- the types of objects for which the category recognition unit 110 recognizes categories may include things, people, sex, emotions, postures, places, colors, and orientations.
- the voice recognition adding unit 120 performs voice recognition on the voice data 310 included in the moving image data 200, and adds the recognition results to the recognition result time series data 210 in time series.
- This recognition result can be expressed by a character string, for example.
- the voice recognition adding unit 120 may add a character string related to the same category as that recognized by the category recognizing unit 110. Note that the voice recognition adding unit 120 may add the recognition result to the corresponding category recognized by the category recognizing unit 110.
- the scene estimation unit 130 estimates a scene using the recognition result time series data 210 recognized by the category recognition unit 110.
- the scene estimation unit 130 estimates a scene as a specific scene having a common content from a set of words such as names of typical objects that can be included in various scenes using, for example, topic modeling. This word set may be appropriately selected by the user during modeling, or may be created by self-organized learning such as a hop field type neural network.
- the scene estimation unit 130 sets the estimation result in the scene estimation data 220 that is time-series data for scene estimation.
- the scene estimation unit 130 can also aggregate the scene estimation data 220. At this time, the scene estimation unit 130 may set estimation results hierarchically in the scene estimation data 220. In addition, when all scenes are finally collected, the scene estimation unit 130 estimates that this is the title of the moving image data 200 and sets the title data 230.
- the scene synthesis unit 140 extracts and synthesizes a part of the video data 200 corresponding to the same type of scene estimated by the scene estimation unit 130 from the plurality of video data 200.
- the moving image data 200 is moving image data of various formats. Examples of the moving image data 200 include Motion JPEG, MPEG1, 2, 4, 5, VC-1, H.264, and the like. H.264, H.C. Various formats such as H.265 (HEVC) and VP9 can be supported.
- the file of the moving image data 200 may correspond to various containers such as AVI, PS (Program Stream), TS (Transport Stream), MOV, MP4, MKV, ASF, MXF, 3GPP, and FLV. Further, the moving image data 200 may be uploaded by a user via a USB or a network and stored in the storage unit 19, for example.
- the moving image data 200 may be acquired from another server on the cloud, a network camera, a NAS, or the like and stored in the storage unit 19.
- the moving image data 200 may be filed by receiving a television broadcast with a tuner.
- This television broadcasting may be analog broadcasting or digital broadcasting.
- the moving image data 200 may include still image data 300 and audio data 310.
- the still image data 300 is image data of a still image included in the moving image data 200, for example, in units of frames.
- the still image data 300 is data having a GOP (Group Of Picture) structure or the like, which is compressed based on an image change in units of frames such as MPEG, for example, a reference image (I picture or the like) Are drawn by the still image acquisition unit 100 from the combination of the difference image (P picture, B picture, etc.).
- Still image data 300 may be still image data such as JPEG inserted into moving image data 200 at specific intervals.
- the audio data 310 is audio data of various quantization frequencies and various compression methods.
- the audio data 310 may include a plurality of audio streams such as stereo and various surrounds.
- the audio data 310 may be, for example, WAV, MP2, MP3, AC-3, AAC, Ogg, or the like.
- the moving image data 200 may include data indicating a timeline (time series) such as the number of frames and a time code.
- the moving image data 200 may include subtitles, other control data, meta (hidden) data, and the like.
- the moving image data 200 may include recognition result time-series data 210, scene estimation data 220, and title data 230, which will be described later.
- the recognition result time series data 210 is data in which categories of objects recognized by the category recognition unit 110 are recorded in time series.
- the recognition result time-series data 210 may be, for example, data including a set of character strings such as category names recognized in units of frames, a start position where each character string appears, an end position, and the like. That is, the recognition result time series data 210 includes a recognition result corresponding to a time series synchronized with the timeline of the moving image data 200. That is, the recognition result time-series data 210 lists, for example, characters such as recognized category names corresponding to the timeline.
- the recognition result time series data 210 includes data such as character strings recognized from the audio data 310 of the moving image data 200 by the voice recognition adding unit 120 so as to be listed corresponding to the timeline. May be.
- the recognition result time-series data 210 includes, for example, a set of IDs and numbers indicated by a category dictionary set for each type of own device and each type of moving image data 200, and a start position and an end where the IDs and numbers appear. It may be data including a position or the like.
- This dictionary may be created, for example, by reading image data of a plurality of moving images using a self-organizing map or an artificial neural network.
- the scene estimation data 220 is data indicating a scene estimated by the scene estimation unit 130.
- the scene estimation data 220 may include scene data including a scene name, a start position, an end position, and the like, for example.
- the scene estimation data 220 includes scene data including a scene name, a start position, an end position, and the like for each layer when aggregated by, for example, hierarchical clustering. Also good. That is, the scene estimation data 220 may be composed of hierarchical scene data.
- the scene name may be character data, ID, number, or the like. This character data is, for example, a character string or the like indicating the location or situation of the background, such as the name of the person who appears, gender, various actions, situations, or emotions.
- This scene name can be used as an identifier when searching for a necessary part from the moving image data 200. Further, an ID such as a chapter number may be given as the scene name. Further, the start position and the end position may be specified in units of frames, time units of the moving image data 200, or the like.
- Title data 230 is data to which the scene estimation data 220 is aggregated and given by the scene estimation unit 130. Specifically, the title data 230 is title data of the estimated moving image data 200. This title may indicate the term, genre, type, etc. selected as the title of the moving image data 200. Further, for example, when the scene estimation data 220 is aggregated by hierarchical clustering or the like, the title data 230 may be character data corresponding to a scene name aggregated as a root, an ID, a number, or the like.
- the combined scene data 240 is image data of a moving image obtained by extracting the moving image data 200 corresponding to the same type of scene based on the scene estimation data 220 estimated by the scene estimating unit 130.
- the same type of scene is not necessarily the same as the scene name in the scene estimation data 220 but may be a related scene name.
- control unit 10 of the moving image data analysis apparatus 1 executes the control program stored in the storage unit 19, so that the still image acquisition unit 100, the category recognition unit 110, the speech recognition addition unit 120, and the scene estimation unit 130. , And the scene composition unit 140.
- Each unit of the moving image data analysis apparatus 1 described above becomes a hardware resource for executing the image forming method of the present invention.
- the moving image data analysis processing of the present embodiment starts analyzing the moving image data 200 stored in the storage unit 19 when an instruction from the user is acquired or when scheduling is performed by CRON or the like. At this time, analysis of the moving image data 200, appropriate scene estimation, and extraction of scene data are performed. In this scene estimation, scenes of the moving image data 200 are estimated uniformly and accurately, and the titles are estimated by aggregation.
- the control unit 10 mainly executes a program stored in the storage unit 19 using hardware resources in cooperation with each unit.
- the details of the moving image data analysis process will be described step by step with reference to the flowchart of FIG.
- Step S101 First, the still image acquisition unit 100 performs a still image acquisition process.
- the still image acquisition unit 100 acquires still image data 300 corresponding to time series from the moving image data 200.
- the moving image data 200 is data that is compressed based on the image change in units of frames
- the still image acquisition unit 100 creates the still image data 300 by drawing from the combination of the reference image and the difference image. .
- the category recognition unit 110 performs category recognition processing.
- the category recognition unit 110 recognizes an object included in each still image data 300 acquired by the still image acquisition unit 100.
- the category recognizing unit 110 recognizes this object based on the feature amount of various images such as region recognition based on the contour line in the image, recognition from the motion data indicated in the previous and next frames in the moving image data 200, and the like. Is possible. Further, the object region may be recognized by a convolutional neural network described below.
- the category recognition unit 110 recognizes the object category for each still image data 300. Specifically, the category recognition unit 110 performs category recognition of each object included in an image using, for example, a convolutional neural network for each still image data 300. As an example of recognizing this category, things, people, sex, emotions, postures, places, colors, orientations, and the like can be used. For this reason, this convolutional neural network is preferably learned before the category recognition process, and can be further optimized based on the recognition result.
- the category recognition unit 110 recognizes the objects 401 and 402 that are people and the object 403 that is a ball from the still image data 300a acquired from the moving image data 200a. At this time, the category recognition unit 110 may recognize a category indicating that the person's facial expression of the object 401 is laughing and the person of the object 402 has no expression.
- the category recognition unit 110 creates recognition result time-series data 210 in which the recognition results obtained by performing category recognition for the entire moving image are collected in time series order, and stores them in the storage unit 19.
- the recognition result time series data 210 corresponds to a time series synchronized with the timeline of the moving image data 200.
- the still image data 300 may be directly input to the convolutional neural network to perform object recognition and category recognition simultaneously.
- Step S103 the voice recognition adding unit 120 determines whether to recognize voice as well.
- the voice recognition addition unit 120 is designated to recognize voice in the setting of the video data analysis, and when the voice data 310 is included in the video data 200, the voice recognition addition unit 120 determines Yes to perform voice recognition addition. . In other cases, the speech recognition adding unit 120 determines No. In the case of Yes, the speech recognition adding unit 120 advances the processing to step S104. In No, the speech recognition addition part 120 advances a process to step S105.
- Step S104 When performing voice recognition addition, the voice recognition addition unit 120 performs voice recognition addition processing.
- the voice recognition adding unit 120 also performs voice recognition on the voice data 310 related to the entire moving image data 200 and includes it in the recognition result time-series data 210.
- Step S105 the scene estimation unit 130 performs a scene estimation process.
- the scene estimation unit 130 uses the recognition result time-series data 210 to perform scene estimation in a specific section.
- the scene estimation unit 130 uses topic modeling or the like to perform scene estimation for the entire moving image data 200 in units of a very short specific period such as several tens to several hundred frames.
- the scene estimation unit 130 estimates the scene based on a set of words such as names of typical objects that can be included in the scene.
- the scene estimation unit 130 sets the estimation result of each scene as data corresponding to the first layer of the scene estimation data 220.
- the scene data set in the scene estimation data 220 can be used as an identifier.
- Step S106 the scene estimation unit 130 performs scene aggregation processing.
- the scene estimation unit 130 uses the scene estimation result and the recognition result time-series data 210 to perform aggregation by clustering or the like. That is, the scene estimation unit 130 performs scene estimation in a larger unit.
- the scene estimation unit 130 collects similar scene data of the first layer of the scene estimation data 220a estimated from the recognition result time-series data 210a based on the similarity and the like. .
- the scene estimation unit 130 may calculate the similarity by weighting the types of character strings and the like of the recognition result time-series data 210.
- the scene estimation unit 130 sets scene data including a scene name, a start position, an end position, and the like for each layer. Accordingly, it is possible to uniformly set appropriate identifiers at various granularities for a user who uses the moving image data 200.
- the scene estimation unit 130 finally aggregates the scenes into one.
- FIG. 5 shows a concept for clustering scene estimation data 220a.
- the data of the scene estimation data 220a may have a structure in which scene data is set for each layer.
- Step S107 the scene estimation unit 130 determines whether all scenes have been collected.
- the scene estimation unit 130 determines Yes when the scenes are collected into one. In other cases, the scene estimation unit 130 determines No.
- the scene estimation part 130 advances a process to step S108. In No, the scene estimation part 130 returns a process to step S106, and repeats aggregation of a scene.
- Step S108 When all the scenes are collected, the scene estimation unit 130 performs a title assignment process.
- the scene estimation unit 130 sets the scene name when the scenes are combined into one in the title data 230.
- the scene estimation unit 130 finally estimates the scene name of the scene data aggregated into one as the title of the moving image and sets it in the title data 230a.
- Step S109 the scene synthesis unit 140 determines whether to synthesize a scene.
- the scene synthesis unit 140 determines Yes when scene synthesis is specified in the analysis settings of the moving image data 200. In other cases, the scene synthesis unit 140 determines No.
- combination part 140 advances a process to step S110. In No, the scene synthetic
- Step S110 When scene synthesis is designated, the scene synthesis unit 140 performs scene synthesis processing.
- the scene synthesis unit 140 extracts and synthesizes the video data 200 corresponding to the same type of scene estimated by the scene estimation unit 130 from the plurality of video data 200.
- the scene synthesis unit 140 searches for scene names (identifiers) of scene data of each layer of the scene estimation data 220 based on, for example, a search keyword included in the analysis setting of the moving image data 200, and finds a suitable one. get.
- the scene synthesizing unit 140 may search for a keyword or the like from the upper layer or the lower layer of each scene estimation data 220.
- the scene composition unit 140 sets all the searched scene data that have the same scene name even if they are aggregated, depending on the setting, to the lowest layer (granularity) or upper layer (granularity) scene data. Is possible to get.
- Each scene estimation data 220a, b, c may include the same scene name.
- FIG. 6A shows an example in which “cooking” is searched as a search keyword for the scene estimation data 220a, b, and c in accordance with a user instruction.
- portions where scene data exists are schematically shown by lines. Also, only the location of the scene name “Cooking” is shown.
- the combined scene data 240 includes moving image data 200d and scene estimation data 220d obtained by combining the extracted scene portion data with respect to the moving image data 200a, b, c and the scene estimation data 220a, b, c. Is also included.
- the moving image data analysis process according to the embodiment of the present invention is completed.
- the moving image data analysis apparatus 1 is a moving image data analysis apparatus that analyzes the moving image data 200, and obtains still image data 300 corresponding to time series from the moving image data 200.
- a still image acquisition unit 100 that recognizes an object included in the still image data 300 acquired by the still image acquisition unit 100, recognizes a category of the recognized object, and creates recognition result time-series data 210.
- the image processing apparatus includes a recognition unit 110 and a scene estimation unit 130 that estimates a scene using the recognition result time-series data 210 recognized by the category recognition unit 110.
- a scene can be estimated from the moving image data 200. Then, by estimating a scene from the moving image data 200, this scene can be used as an identifier. That is, it becomes possible to uniformly and accurately assign an identifier with a fine granularity to the moving image data 200. This makes it possible to easily search for necessary scenes from the moving image data 200 and effectively use them.
- some moving image data can be categorized by a user and an identifier can be added.
- an identifier can be added.
- the moving image data analysis apparatus 1 analyzes the contents of the moving image data 200 and estimates an appropriate scene, so that this scene can be used as an identifier. . Thereby, it can be expected that the user can easily obtain an appropriate moving image.
- the title of the video set by the poster and the content of the video may be different. If the video is completely different, the user can judge immediately after browsing and stop the playback, but if the title and video content are slightly out of sync, for example, It was necessary to confirm whether it was the target content by operating the seek bar. As described above, the user is stressed by these operations because the expected content cannot be obtained regardless of whether the title and the content of the moving image can be immediately determined. Moreover, such a situation is not a situation that occurs specifically only on the video posting site, but may also occur for general video data that is not properly assigned a title.
- the scene estimation unit 130 aggregates the scene estimation data 220 that is time-series data of scene estimation, and estimates the title of the moving image data 200. It is characterized by doing. By configuring in this way and estimating titles aggregated from all scene estimates, appropriate titles can be automatically assigned.
- the moving image data analysis apparatus 1 indicates that the types of objects that the category recognition unit 110 recognizes categories are things, people, gender, emotions, postures, places, colors, and orientations. Features. With this configuration, it is possible to assign identifiers with fine granularity such as chapters, scene types, characters, and places in movies to the movie data 200, and the movie data 200 can be used more effectively. Become.
- the moving image data analysis apparatus 1 includes a voice recognition adding unit 120 that performs voice recognition on the audio data 310 included in the moving image data 200 and adds the recognition result time-series data 210 to the recognition result time-series data 210 in time-series order.
- the recognition result time-series data 210 includes a recognition result as characters for a time series synchronized with the timeline of the moving image data 200.
- the moving image data analysis apparatus 1 extracts a scene combining unit that extracts and combines moving image data 200 corresponding to the same type of scene estimated by the scene estimating unit 130 from a plurality of moving image data 200. 140 is further provided.
- a scene combining unit that extracts and combines moving image data 200 corresponding to the same type of scene estimated by the scene estimating unit 130 from a plurality of moving image data 200.
- the synthesized scene data 240 obtained by extracting the scene that the user wants to search.
- the moving image data analysis apparatus 1 of the present invention can be applied to various moving image data 200 that requires scene estimation and title estimation. For example, from a plurality of moving image data 200 of a network camera connected to an image forming apparatus or the like, a scene such as a user's use or trouble is estimated, and the use or trouble type is given as a title. Alternatively, the identifier may be extracted and combined. Moreover, you may estimate about the scene corresponding to categories, such as a user's gender distinction, age classification, and facial expression.
- the function of the image forming apparatus used by the user may be included as metadata of the moving image data 200.
- the video data 200 such as a life log that the user has taken with a wearable device, the scene where the person was met, what was spoken, and what action was taken was estimated, and the title Can also be stored. Thereby, it becomes easy to search and extract necessary items from the enormous life log moving image data 200 for use.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
動画データを容易に活用可能にする画像データ解析装置を提供する。 動画データ解析装置(1)は、動画データ(200)の解析を行う。静止画像取得部(100)は、動画データ(200)から時系列に対応した静止画像データ(300)を取得する。カテゴリー認識部(110)は、静止画像取得部(100)により取得された静止画像データ(300)に含まれるオブジェクトを認識し、認識されたオブジェクトのカテゴリーの認識を行い、認識結果時系列データ(210)を作成する。シーン推定部(130)は、カテゴリー認識部(110)により認識された認識結果時系列データ(210)を用いて、特定区間のシーンの推定を行う。
Description
本発明は、動画データ解析装置及び動画データ解析方法に係り、特に動画像を解析するための動画データ解析装置及び動画データ解析方法に関する。
従来から、動画データ(動画の画像データ、動画像データ)の内容を解析する動画データ解析装置が存在する。
特許文献1を参照すると、非圧縮または圧縮された動画像データを、動画像としての特徴や、必要に応じて動画像に付随するオーディオの特徴を用いて、低コストかつ高精度で様々なクラスへ分類する動画像データの分類装置が開示されている。
特許文献1を参照すると、非圧縮または圧縮された動画像データを、動画像としての特徴や、必要に応じて動画像に付随するオーディオの特徴を用いて、低コストかつ高精度で様々なクラスへ分類する動画像データの分類装置が開示されている。
しかしながら、特許文献1の技術は、動画データ内の色配置記述子等の画像特徴値とオーディオの特徴とにより数種類程度のクラスに分類する技術であった。
このため、動画データ内で、内容が共通した特定の場面等であるシーンを推定することはできなかった。
このため、動画データ内で、内容が共通した特定の場面等であるシーンを推定することはできなかった。
本発明は、このような状況に鑑みてなされたものであって、上述の問題点を解消する動画データ解析装置を提供することを課題とする。
本発明の動画データ解析装置は、動画データの解析を行う動画データ解析装置であって、前記動画データから時系列に対応した静止画像データを取得する静止画像取得部と、前記静止画像取得部により取得された前記静止画像データに含まれるオブジェクトを認識し、認識された前記オブジェクトのカテゴリーの認識を行い、認識結果時系列データを作成するカテゴリー認識部と、前記カテゴリー認識部により認識された前記認識結果時系列データを用いて、シーンの推定を行うシーン推定部とを備えることを特徴とする。
本発明の動画データ解析方法は、動画データ解析装置により実行される動画データ解析方法であって、前記動画データ解析装置は、動画データから時系列に対応した静止画像データを取得し、取得された前記静止画像データに含まれるオブジェクトを認識し、認識された前記オブジェクトのカテゴリーの認識を行い、認識結果時系列データを作成し、前記認識結果時系列データを用いて、特定区間のシーンの推定を行うことを特徴とする。
本発明の動画データ解析方法は、動画データ解析装置により実行される動画データ解析方法であって、前記動画データ解析装置は、動画データから時系列に対応した静止画像データを取得し、取得された前記静止画像データに含まれるオブジェクトを認識し、認識された前記オブジェクトのカテゴリーの認識を行い、認識結果時系列データを作成し、前記認識結果時系列データを用いて、特定区間のシーンの推定を行うことを特徴とする。
本発明によれば、動画像データから静止画像データを取得して、この静止画像データに含まれるオブジェクトを認識し、認識されたオブジェクトのカテゴリーの認識を行って認識結果時系列データを作成し、この認識結果時系列データからシーンの推定を行うことで、動画像からシーンを推定可能な画像データ解析装置を提供することができる。
<実施の形態>
〔動画データ解析装置1の全体のシステム構成〕
まず、図1を参照して、動画データ解析装置1の全体のシステム構成について説明する。
動画データ解析装置1は、動画データ200の解析を行うサーバー、PC(Personal Computer)、NAS(Network Attached Storage)、ビデオレコーダー、映像配信サーバー、携帯端末、スマートフォン、画像形成装置等の情報処理装置である。
〔動画データ解析装置1の全体のシステム構成〕
まず、図1を参照して、動画データ解析装置1の全体のシステム構成について説明する。
動画データ解析装置1は、動画データ200の解析を行うサーバー、PC(Personal Computer)、NAS(Network Attached Storage)、ビデオレコーダー、映像配信サーバー、携帯端末、スマートフォン、画像形成装置等の情報処理装置である。
また、動画データ解析装置1は、制御部10、I/O部15、及び記憶部19等を含んでいる。各部は、制御部10に接続され、制御部10によって動作制御される。
制御部10は、GPP(General Purpose Processor)、CPU(Central Processing Unit、中央処理装置)、MPU(Micro Processing Unit)、DSP(Digital Signal Processor)、GPU(Graphics Processing Unit)、ASIC(Application Specific Processor、特定用途向けプロセッサー)等の情報処理手段である。
制御部10は、記憶部19のROMやHDDに記憶されている制御プログラムを読み出して、この制御プログラムをRAMに展開させて実行することで、後述する機能ブロックの各手段として動作させられる。また、制御部10は、図示しない外部の端末、接続されたキーボードやマウス等の入力部から入力された所定の指示情報に応じて、装置全体の制御を行う。
制御部10は、記憶部19のROMやHDDに記憶されている制御プログラムを読み出して、この制御プログラムをRAMに展開させて実行することで、後述する機能ブロックの各手段として動作させられる。また、制御部10は、図示しない外部の端末、接続されたキーボードやマウス等の入力部から入力された所定の指示情報に応じて、装置全体の制御を行う。
I/O部15は、いわゆるチップセット等であり、各種インターフェイスを備え、データを送受信する回路等である。I/O部15は、例えば、LAN(Local Area Network)ボード、無線送受信機、USB(Universal Serial Bus)インターフェイス等を含んでいてもよい。
また、I/O部15は、ネットワークカメラやNAS(Network Attached Storage)等の機器との間で、データを送受信することも可能である。
また、I/O部15は、ネットワークカメラやNAS(Network Attached Storage)等の機器との間で、データを送受信することも可能である。
記憶部19は、ROM(Read Only Memory)、RAM(Random Access Memory)、フラッシュメモリー等の半導体メモリーやHDD(Hard Disk Drive)等の一時的でない記録媒体である。
記憶部19のROMやHDDには、後述する動画データ200(図2)を含む各種データに加え、動画データ解析装置1の動作制御を行うための制御プログラムを格納している。記憶部19は、動画データ200をどのように解析するかについてのユーザーの指示、スケジューリング等が指定された、動画データ解析の設定のデータ(図示せず)も格納している。
なお、記憶部19は、デバイス毎、パーティション毎、ユーザー毎等に分割された領域で管理されていてもよい。
記憶部19のROMやHDDには、後述する動画データ200(図2)を含む各種データに加え、動画データ解析装置1の動作制御を行うための制御プログラムを格納している。記憶部19は、動画データ200をどのように解析するかについてのユーザーの指示、スケジューリング等が指定された、動画データ解析の設定のデータ(図示せず)も格納している。
なお、記憶部19は、デバイス毎、パーティション毎、ユーザー毎等に分割された領域で管理されていてもよい。
なお、動画データ解析装置1において、制御部10及びI/O部15は、いわゆるSOC(System On Chip)、GPU内蔵CPU等のように、一体的に形成されていてもよい。
また、制御部10及びI/O部15は、RAMやROMやフラッシュメモリー等の一時的でない記憶媒体を内蔵していてもよい。
また、制御部10及びI/O部15は、RAMやROMやフラッシュメモリー等の一時的でない記憶媒体を内蔵していてもよい。
〔動画データ解析装置1の制御構成〕
ここで、図2を参照し、動画データ解析装置1の制御構成について説明する。
動画データ解析装置1の制御部10は、静止画像取得部100、カテゴリー認識部110、音声認識付加部120、シーン推定部130、及びシーン合成部140を備えている。
記憶部19は、動画データ200、認識結果時系列データ210、シーン推定データ220、タイトルデータ230、及び合成シーンデータ240を格納している。
ここで、図2を参照し、動画データ解析装置1の制御構成について説明する。
動画データ解析装置1の制御部10は、静止画像取得部100、カテゴリー認識部110、音声認識付加部120、シーン推定部130、及びシーン合成部140を備えている。
記憶部19は、動画データ200、認識結果時系列データ210、シーン推定データ220、タイトルデータ230、及び合成シーンデータ240を格納している。
静止画像取得部100は、動画データ200から時系列に対応した静止画像データ300を取得する。
具体的には、静止画像取得部100は、例えば、動画データ200に含まれる画像データを、時系列順に、フレーム(frame)単位で抜き出して、ビットマップ(bitmap)データ等の静止画像データ300として記憶部19に格納する。この際、静止画像取得部100は、動画データ200が基準画像と差分画像の組み合わせのような形式で格納されていた場合には、この基準画像と差分画像とから、当該フレームの静止画像を描画して、静止画像データ300を作成してもよい。また、静止画像取得部100は、動画データ200にオブジェクトのデータが含まれていた場合には、これを静止画像データ300に含めてもよい。
具体的には、静止画像取得部100は、例えば、動画データ200に含まれる画像データを、時系列順に、フレーム(frame)単位で抜き出して、ビットマップ(bitmap)データ等の静止画像データ300として記憶部19に格納する。この際、静止画像取得部100は、動画データ200が基準画像と差分画像の組み合わせのような形式で格納されていた場合には、この基準画像と差分画像とから、当該フレームの静止画像を描画して、静止画像データ300を作成してもよい。また、静止画像取得部100は、動画データ200にオブジェクトのデータが含まれていた場合には、これを静止画像データ300に含めてもよい。
カテゴリー認識部110は、静止画像取得部100により取得された静止画像データ300に含まれるオブジェクトを認識し、認識されたオブジェクトの種類等であるカテゴリーの認識を行い、この認識結果を時系列順で記録した認識結果時系列データ210を作成する。このため、カテゴリー認識部110は、例えば、畳み込みニューラルネットワーク等を用いて、オブジェクトのカテゴリーを認識する。このカテゴリーの認識結果は、例えば、物体の名称等の文字列で表現されてもよい。
また、カテゴリー認識部110がカテゴリーを認識するオブジェクトの種類は、物、人、性別、感情、姿勢、場所、色、及び向きを含んでいてもよい。
また、カテゴリー認識部110がカテゴリーを認識するオブジェクトの種類は、物、人、性別、感情、姿勢、場所、色、及び向きを含んでいてもよい。
音声認識付加部120は、動画データ200に含まれる音声データ310について音声認識を行い、認識結果時系列データ210に、認識結果を時系列順に付加する。この認識結果は、例えば、文字列で表現することが可能である。また、音声認識付加部120は、カテゴリー認識部110で認識されるのと同程度のカテゴリーに関する文字列を付加してもよい。
なお、音声認識付加部120は、カテゴリー認識部110により認識された対応するカテゴリーに、認識結果を付加してもよい。
なお、音声認識付加部120は、カテゴリー認識部110により認識された対応するカテゴリーに、認識結果を付加してもよい。
シーン推定部130は、カテゴリー認識部110により認識された認識結果時系列データ210を用いて、シーンの推定を行う。
シーン推定部130は、例えば、トピックモデリング等を用いて、各種のシーンに含まれ得る典型的な物体の名称等の単語集合から、内容が共通した特定の場面としてのシーンの推定を行う。この単語集合は、モデリングの際にユーザーにより適切に選択されたり、ホップフィールド型のニューラルネットのような自己組織化学習により作成されたりしてもよい。シーン推定部130は、推定結果を、シーンの推定の時系列データであるシーン推定データ220に設定する。
また、シーン推定部130は、シーン推定データ220を集約することも可能である。この際、シーン推定部130は、シーン推定データ220に、階層的に、推定結果を設定してもよい。また、シーン推定部130は、最終的に全てのシーンが集約された際に、これが動画データ200のタイトルであると推定し、タイトルデータ230に設定する。
シーン推定部130は、例えば、トピックモデリング等を用いて、各種のシーンに含まれ得る典型的な物体の名称等の単語集合から、内容が共通した特定の場面としてのシーンの推定を行う。この単語集合は、モデリングの際にユーザーにより適切に選択されたり、ホップフィールド型のニューラルネットのような自己組織化学習により作成されたりしてもよい。シーン推定部130は、推定結果を、シーンの推定の時系列データであるシーン推定データ220に設定する。
また、シーン推定部130は、シーン推定データ220を集約することも可能である。この際、シーン推定部130は、シーン推定データ220に、階層的に、推定結果を設定してもよい。また、シーン推定部130は、最終的に全てのシーンが集約された際に、これが動画データ200のタイトルであると推定し、タイトルデータ230に設定する。
シーン合成部140は、複数の動画データ200から、シーン推定部130により推定された同一種類のシーンに対応する動画データ200の箇所を抜き出して合成する。
動画データ200は、各種フォーマットの動画の画像データである。動画データ200として、例えば、MotionJPEG、MPEG1、2、4、5、VC-1、H.264、H.265(HEVC)、VP9等の各種フォーマットに対応可能である。また、動画データ200のファイルは、AVI、PS(Program Stream)、TS(Transport Stream)、MOV、MP4、MKV、ASF、MXF、3GPP、FLV等の各種コンテナに対応していてもよい。
また、動画データ200は、例えば、ユーザーによりUSBやネットワーク経由等でアップロードされて、記憶部19に格納されてもよい。また、動画データ200は、クラウド上の他のサーバー、ネットワークカメラ、NAS等から取得されて、記憶部19に格納されてもよい。また、動画データ200は、テレビジョン放送をチューナーで受信して、ファイル化されたものであってもよい。このテレビジョン放送は、アナログ放送でも、デジタル放送でもよい。
また、動画データ200は、例えば、ユーザーによりUSBやネットワーク経由等でアップロードされて、記憶部19に格納されてもよい。また、動画データ200は、クラウド上の他のサーバー、ネットワークカメラ、NAS等から取得されて、記憶部19に格納されてもよい。また、動画データ200は、テレビジョン放送をチューナーで受信して、ファイル化されたものであってもよい。このテレビジョン放送は、アナログ放送でも、デジタル放送でもよい。
また、動画データ200は、静止画像データ300及び音声データ310を含んでいてもよい。
このうち、静止画像データ300は、動画データ200に含まれる、例えば、フレーム単位の静止画の画像データである。また、静止画像データ300は、例えば、MPEG等のようにフレーム単位の画像変化を基に圧縮された、GOP(Group Of Picture)構造等のあるデータの場合には、基準画像(Iピクチャ等)と差分画像(PピクチャやBピクチャ等)の組み合わせから静止画像取得部100により描画されて作成される。また、静止画像データ300は、動画データ200に特定期間毎に挿入されたJPEG等の静止画のデータであってもよい。
また、音声データ310は、各種量子化周波数、各種圧縮方式の音声のデータである。音声データ310は、ステレオ、各種サラウンド等の複数の音声ストリームを含んでいてもよい。音声データ310は、例えば、WAV、MP2、MP3、AC-3、AAC、Ogg等であってもよい。
このうち、静止画像データ300は、動画データ200に含まれる、例えば、フレーム単位の静止画の画像データである。また、静止画像データ300は、例えば、MPEG等のようにフレーム単位の画像変化を基に圧縮された、GOP(Group Of Picture)構造等のあるデータの場合には、基準画像(Iピクチャ等)と差分画像(PピクチャやBピクチャ等)の組み合わせから静止画像取得部100により描画されて作成される。また、静止画像データ300は、動画データ200に特定期間毎に挿入されたJPEG等の静止画のデータであってもよい。
また、音声データ310は、各種量子化周波数、各種圧縮方式の音声のデータである。音声データ310は、ステレオ、各種サラウンド等の複数の音声ストリームを含んでいてもよい。音声データ310は、例えば、WAV、MP2、MP3、AC-3、AAC、Ogg等であってもよい。
なお、動画データ200は、フレーム数やタイムコード等のタイムライン(時系列)を示すデータを含んでいてもよい。
また、動画データ200は、字幕やその他の制御データやメタ(隠蔽)データ等を含んでいてもよい。
また、動画データ200は、後述する認識結果時系列データ210、シーン推定データ220、及びタイトルデータ230を含んでいてもよい。
また、動画データ200は、字幕やその他の制御データやメタ(隠蔽)データ等を含んでいてもよい。
また、動画データ200は、後述する認識結果時系列データ210、シーン推定データ220、及びタイトルデータ230を含んでいてもよい。
認識結果時系列データ210は、カテゴリー認識部110により認識されたオブジェクトのカテゴリーを時系列順に記録したデータである。認識結果時系列データ210は、例えば、フレーム単位で認識されたカテゴリーの名称等の文字列の集合と、各文字列が出現する開始位置、終了位置等を含むデータ等であってもよい。すなわち、認識結果時系列データ210は、動画データ200のタイムラインと同期する時系列に対応した認識結果を含んでいる。すなわち、認識結果時系列データ210は、例えば、認識されたカテゴリーの名称等の文字を、タイムラインに対応して列挙している。
また、認識結果時系列データ210は、音声認識付加部120により動画データ200の音声データ310から認識された文字列等のデータについても、同様に、タイムラインに対応して列挙するように含んでいてもよい。
なお、認識結果時系列データ210は、例えば、自装置の種類や動画データ200の種類別に設定されたカテゴリーの辞書で示されるIDや番号の集合と、このIDや番号の出現する開始位置、終了位置等を含むデータであってもよい。この辞書は、例えば、自己組織化マップや人工ニューラルネットにより複数の動画の画像データを読み込ませて作成されたものであってもよい。
また、認識結果時系列データ210は、音声認識付加部120により動画データ200の音声データ310から認識された文字列等のデータについても、同様に、タイムラインに対応して列挙するように含んでいてもよい。
なお、認識結果時系列データ210は、例えば、自装置の種類や動画データ200の種類別に設定されたカテゴリーの辞書で示されるIDや番号の集合と、このIDや番号の出現する開始位置、終了位置等を含むデータであってもよい。この辞書は、例えば、自己組織化マップや人工ニューラルネットにより複数の動画の画像データを読み込ませて作成されたものであってもよい。
シーン推定データ220は、シーン推定部130により推定されたシーンを示すデータである。
シーン推定データ220は、例えば、シーン名、開始位置、終了位置等を含むシーンデータを含んでいてもよい。
また、シーン推定データ220は、後述するように、例えば、階層的クラスタリング(clustering)等で集約された場合には、各層について、シーン名、開始位置、終了位置等を含むシーンデータを備えていてもよい。すなわち、シーン推定データ220は、階層的なシーンデータで構成されてもよい。
また、シーン名は、文字データ、IDや番号等であってもよい。この文字データは、例えば、登場するヒトの名前や男女別や各種行動や状況や感情等、背景の場所や状況等を示す文字列等である。このシーン名は、動画データ200から、必要な箇所を探す際の識別子として用いることが可能である。また、シーン名として、チャプター番号のようなIDが付与されていてもよい。
また、開始位置及び終了位置は、フレーム単位や動画データ200の時間単位等で指定されていてもよい。
シーン推定データ220は、例えば、シーン名、開始位置、終了位置等を含むシーンデータを含んでいてもよい。
また、シーン推定データ220は、後述するように、例えば、階層的クラスタリング(clustering)等で集約された場合には、各層について、シーン名、開始位置、終了位置等を含むシーンデータを備えていてもよい。すなわち、シーン推定データ220は、階層的なシーンデータで構成されてもよい。
また、シーン名は、文字データ、IDや番号等であってもよい。この文字データは、例えば、登場するヒトの名前や男女別や各種行動や状況や感情等、背景の場所や状況等を示す文字列等である。このシーン名は、動画データ200から、必要な箇所を探す際の識別子として用いることが可能である。また、シーン名として、チャプター番号のようなIDが付与されていてもよい。
また、開始位置及び終了位置は、フレーム単位や動画データ200の時間単位等で指定されていてもよい。
タイトルデータ230は、シーン推定部130によりシーン推定データ220が集約されて付与されたデータである。具体的には、タイトルデータ230は、推定された動画データ200のタイトルのデータである。このタイトルは、動画データ200の題名として選択された用語、ジャンル、種類等を示すものであってもよい。また、タイトルデータ230は、例えば、シーン推定データ220が階層的クラスタリング等で集約された場合は、ルート(根)として集約されたシーン名にあたる文字データ、IDや番号等であってもよい。
合成シーンデータ240は、シーン推定部130により推定されたシーン推定データ220に基づいて、同一種類のシーンに対応する動画データ200が抜き出されて合成された動画の画像データである。この同一種類のシーンとしては、必ずしもシーン推定データ220内のシーン名と同一ではなく、関連あるシーン名であってもよい。
ここで、動画データ解析装置1の制御部10は、記憶部19に記憶された制御プログラムを実行することで、静止画像取得部100、カテゴリー認識部110、音声認識付加部120、シーン推定部130、及びシーン合成部140として機能させられる。
また、上述の動画データ解析装置1の各部は、本発明の画像形成方法を実行するハードウェア資源となる。
また、上述の動画データ解析装置1の各部は、本発明の画像形成方法を実行するハードウェア資源となる。
〔動画データ解析装置1による動画データ解析処理〕
次に、図3~図6Bを参照して、本発明の実施の形態に係る動画データ解析装置1による動画データ解析処理の説明を行う。
本実施形態の動画データ解析処理は、ユーザーによる指示を取得した場合、又は、CRON等によりスケジューリングされた場合等に、記憶部19に格納された動画データ200の解析を開始する。この際、動画データ200の解析と適切なシーン推定、及び、シーンのデータの抜き出しを行う。このシーン推定においては、動画データ200について、一律かつ正確にシーンを推定し、集約してタイトルを推定する。このように、動画データ200の内容を解析することで、動画データ200の活用を意図するユーザーが適切なシーンのような識別子やタイトルが付加された動画データ200を得ることが可能となる。
本実施形態の動画データ解析処理は、主に制御部10が、記憶部19に記憶されたプログラムを、各部と協働し、ハードウェア資源を用いて実行する。
以下で、図3のフローチャートを参照して、動画データ解析処理の詳細をステップ毎に説明する。
次に、図3~図6Bを参照して、本発明の実施の形態に係る動画データ解析装置1による動画データ解析処理の説明を行う。
本実施形態の動画データ解析処理は、ユーザーによる指示を取得した場合、又は、CRON等によりスケジューリングされた場合等に、記憶部19に格納された動画データ200の解析を開始する。この際、動画データ200の解析と適切なシーン推定、及び、シーンのデータの抜き出しを行う。このシーン推定においては、動画データ200について、一律かつ正確にシーンを推定し、集約してタイトルを推定する。このように、動画データ200の内容を解析することで、動画データ200の活用を意図するユーザーが適切なシーンのような識別子やタイトルが付加された動画データ200を得ることが可能となる。
本実施形態の動画データ解析処理は、主に制御部10が、記憶部19に記憶されたプログラムを、各部と協働し、ハードウェア資源を用いて実行する。
以下で、図3のフローチャートを参照して、動画データ解析処理の詳細をステップ毎に説明する。
(ステップS101)
まず、静止画像取得部100が、静止画像取得処理を行う。
静止画像取得部100は、動画データ200から、時系列に対応した静止画像データ300を取得する。
静止画像取得部100は、動画データ200がフレーム単位の画像変化を基に圧縮されたようなデータであった場合は、基準画像と差分画像との組み合わせから描画して静止画像データ300を作成する。
まず、静止画像取得部100が、静止画像取得処理を行う。
静止画像取得部100は、動画データ200から、時系列に対応した静止画像データ300を取得する。
静止画像取得部100は、動画データ200がフレーム単位の画像変化を基に圧縮されたようなデータであった場合は、基準画像と差分画像との組み合わせから描画して静止画像データ300を作成する。
(ステップS102)
次に、カテゴリー認識部110が、カテゴリー認識処理を行う。
カテゴリー認識部110は、静止画像取得部100により取得された各静止画像データ300に含まれるオブジェクトを認識する。カテゴリー認識部110は、画像内の輪郭線による領域認識や、動画データ200中の前後フレームに示された動きのデータからの認識等、各種画像の特徴量に基づいてこのオブジェクトの認識を行うことが可能である。また、下記で説明する畳み込みニューラルネットにより、オブジェクトの領域を認識してもよい。
次に、カテゴリー認識部110が、カテゴリー認識処理を行う。
カテゴリー認識部110は、静止画像取得部100により取得された各静止画像データ300に含まれるオブジェクトを認識する。カテゴリー認識部110は、画像内の輪郭線による領域認識や、動画データ200中の前後フレームに示された動きのデータからの認識等、各種画像の特徴量に基づいてこのオブジェクトの認識を行うことが可能である。また、下記で説明する畳み込みニューラルネットにより、オブジェクトの領域を認識してもよい。
また、カテゴリー認識部110は、各静止画像データ300について、オブジェクトのカテゴリーの認識を行う。具体的には、カテゴリー認識部110は、各静止画像データ300について、例えば、畳み込みニューラルネットワーク等を用いて、画像に含まれるオブジェクトのカテゴリー認識を行う。このカテゴリーを認識する例としては、物、人、性別、感情、姿勢、場所、色、向き等を用いることが可能である。このため、この畳み込みニューラルネットワークは、カテゴリー認識処理の前に学習させておくことが好適であり、又、認識結果に基づいて更に最適化することも可能である。
図4の例により説明すると、カテゴリー認識部110は、動画データ200aから取得した静止画像データ300aから、人であるオブジェクト401、402と、ボールであるオブジェクト403とを認識する。この際、カテゴリー認識部110は、オブジェクト401の人の表情が笑っていて、オブジェクト402の人が無表情である旨のカテゴリーを認識してもよい。
また、カテゴリー認識部110は、動画全体についてカテゴリー認識を行った認識結果について、時系列順にまとめた認識結果時系列データ210を作成し、記憶部19に格納する。この認識結果時系列データ210は、動画データ200のタイムラインと同期する時系列に対応している。
なお、畳み込みニューラルネットに、直接、静止画像データ300を入力してオブジェクト認識とカテゴリー認識とを同時に行ってもよい。
なお、畳み込みニューラルネットに、直接、静止画像データ300を入力してオブジェクト認識とカテゴリー認識とを同時に行ってもよい。
(ステップS103)
次に、音声認識付加部120が、音声も認識するか否かを判定する。音声認識付加部120は、動画データ解析の設定において、音声も認識するよう指定されており、動画データ200に音声データ310が含まれていた場合には、音声認識付加を行うとしてYesと判定する。音声認識付加部120は、それ以外の場合には、Noと判定する。
Yesの場合、音声認識付加部120は、処理をステップS104に進める。
Noの場合、音声認識付加部120は、処理をステップS105に進める。
次に、音声認識付加部120が、音声も認識するか否かを判定する。音声認識付加部120は、動画データ解析の設定において、音声も認識するよう指定されており、動画データ200に音声データ310が含まれていた場合には、音声認識付加を行うとしてYesと判定する。音声認識付加部120は、それ以外の場合には、Noと判定する。
Yesの場合、音声認識付加部120は、処理をステップS104に進める。
Noの場合、音声認識付加部120は、処理をステップS105に進める。
(ステップS104)
音声認識付加を行う場合、音声認識付加部120が、音声認識付加処理を行う。
音声認識付加部120は、動画データ200全体に係る音声データ310についても音声認識を行い、認識結果時系列データ210に含める。
音声認識付加を行う場合、音声認識付加部120が、音声認識付加処理を行う。
音声認識付加部120は、動画データ200全体に係る音声データ310についても音声認識を行い、認識結果時系列データ210に含める。
(ステップS105)
ここで、シーン推定部130が、シーン推定処理を行う。
シーン推定部130は、認識結果時系列データ210を用いて、特定区間におけるシーン推定を行う。
シーン推定部130は、トピックモデリング等を用いて、例えば、動画データ200全体について、数十フレーム~数百フレーム等のごく短い特定期間の単位でシーン推定を行う。この際、シーン推定部130は、シーンに含まれ得る典型的な物体の名称等の単語集合に基づいて、当該シーンの推定を行う。
ここでは、シーン推定部130は、各シーンの推定結果について、シーン推定データ220の第一層にあたるデータに設定する。このシーン推定データ220に設定されたシーンデータは、識別子として使用可能である。
ここで、シーン推定部130が、シーン推定処理を行う。
シーン推定部130は、認識結果時系列データ210を用いて、特定区間におけるシーン推定を行う。
シーン推定部130は、トピックモデリング等を用いて、例えば、動画データ200全体について、数十フレーム~数百フレーム等のごく短い特定期間の単位でシーン推定を行う。この際、シーン推定部130は、シーンに含まれ得る典型的な物体の名称等の単語集合に基づいて、当該シーンの推定を行う。
ここでは、シーン推定部130は、各シーンの推定結果について、シーン推定データ220の第一層にあたるデータに設定する。このシーン推定データ220に設定されたシーンデータは、識別子として使用可能である。
(ステップS106)
次に、シーン推定部130が、シーン集約処理を行う。
シーン推定部130は、シーン推定の結果及び認識結果時系列データ210を用いてクラスタリング等で集約する。すなわち、シーン推定部130は、より大きな単位でのシーン推定を行う。
次に、シーン推定部130が、シーン集約処理を行う。
シーン推定部130は、シーン推定の結果及び認識結果時系列データ210を用いてクラスタリング等で集約する。すなわち、シーン推定部130は、より大きな単位でのシーン推定を行う。
図5の例により説明すると、シーン推定部130は、認識結果時系列データ210aから推定されたシーン推定データ220aの第一層のシーンデータと類似するものを、類似度等を基にまとめていく。この際に、シーン推定部130は、認識結果時系列データ210の文字列等の種類についても重み付けをして類似度を算出してもよい。この際、シーン推定部130は、各層についても、シーン名、開始位置、終了位置等を含むシーンデータを設定してゆく。これによって、動画データ200を活用するユーザーに対して、様々な粒度において、一律に適正な識別子を設定することが可能となる。
また、シーン推定部130は、最終的には、シーンを一つに集約する。
なお、図5においては、シーン推定データ220aをクラスタリングする際の概念について示している。シーン推定データ220aのデータ内では、各層について、シーンデータが設定されるような構造であってもよい。
また、シーン推定部130は、最終的には、シーンを一つに集約する。
なお、図5においては、シーン推定データ220aをクラスタリングする際の概念について示している。シーン推定データ220aのデータ内では、各層について、シーンデータが設定されるような構造であってもよい。
(ステップS107)
次に、シーン推定部130が、全てのシーンを集約したか否かを判定する。シーン推定部130は、シーンが一つに集約された場合に、Yesと判定する。シーン推定部130は、それ以外の場合には、Noと判定する。
Yesの場合、シーン推定部130は、処理をステップS108に進める。
Noの場合、シーン推定部130は、処理をステップS106に戻して、シーンの集約を繰り返す。
次に、シーン推定部130が、全てのシーンを集約したか否かを判定する。シーン推定部130は、シーンが一つに集約された場合に、Yesと判定する。シーン推定部130は、それ以外の場合には、Noと判定する。
Yesの場合、シーン推定部130は、処理をステップS108に進める。
Noの場合、シーン推定部130は、処理をステップS106に戻して、シーンの集約を繰り返す。
(ステップS108)
全てのシーンを集約した場合、シーン推定部130が、タイトル付与処理を行う。
シーン推定部130は、一つに集約された際のシーン名を、タイトルデータ230に設定する。
図5の例によれば、シーン推定部130は、最終的に、一つに集約されたシーンデータのシーン名を動画のタイトルと推定して、タイトルデータ230aに設定する。
全てのシーンを集約した場合、シーン推定部130が、タイトル付与処理を行う。
シーン推定部130は、一つに集約された際のシーン名を、タイトルデータ230に設定する。
図5の例によれば、シーン推定部130は、最終的に、一つに集約されたシーンデータのシーン名を動画のタイトルと推定して、タイトルデータ230aに設定する。
(ステップS109)
次に、シーン合成部140が、シーンを合成するか否かを判定する。シーン合成部140は、動画データ200の解析の設定において、シーンの合成が指定されていた場合に、Yesと判定する。シーン合成部140は、それ以外の場合には、Noと判定する。
Yesの場合、シーン合成部140は、処理をステップS110に進める。
Noの場合、シーン合成部140は、動画データ解析処理を終了する。
次に、シーン合成部140が、シーンを合成するか否かを判定する。シーン合成部140は、動画データ200の解析の設定において、シーンの合成が指定されていた場合に、Yesと判定する。シーン合成部140は、それ以外の場合には、Noと判定する。
Yesの場合、シーン合成部140は、処理をステップS110に進める。
Noの場合、シーン合成部140は、動画データ解析処理を終了する。
(ステップS110)
シーンの合成が指定されていた場合、シーン合成部140が、シーン合成処理を行う。
シーン合成部140は、複数の動画データ200から、シーン推定部130により推定された同一種類のシーンに対応する動画データ200を抜き出して合成する。
シーン合成部140は、例えば、動画データ200の解析の設定に含まれる検索用のキーワード等に基づき、シーン推定データ220の各層のシーンデータのシーン名(識別子)を検索して、適合するものを取得する。この際、シーン合成部140は、キーワード等を、各シーン推定データ220の層の上層から検索しても、下層から検索してもよい。また、シーン合成部140は、検索された全てのシーンデータのうち、集約されても同一のシーン名等であったものについては、設定により、最も下層(粒度)又は上層(粒度)のシーンデータを取得することが可能である。
シーンの合成が指定されていた場合、シーン合成部140が、シーン合成処理を行う。
シーン合成部140は、複数の動画データ200から、シーン推定部130により推定された同一種類のシーンに対応する動画データ200を抜き出して合成する。
シーン合成部140は、例えば、動画データ200の解析の設定に含まれる検索用のキーワード等に基づき、シーン推定データ220の各層のシーンデータのシーン名(識別子)を検索して、適合するものを取得する。この際、シーン合成部140は、キーワード等を、各シーン推定データ220の層の上層から検索しても、下層から検索してもよい。また、シーン合成部140は、検索された全てのシーンデータのうち、集約されても同一のシーン名等であったものについては、設定により、最も下層(粒度)又は上層(粒度)のシーンデータを取得することが可能である。
図6A及び図6Bの例により説明すると、動画データ200aが料理番組、動画データ200bが報道番組、動画データ200cが旅番組のように別のタイトル(ジャンル)が付与されていた場合であっても、それぞれのシーン推定データ220a、b、cには、同様のシーン名が含まれることがある。
図6Aでは、ユーザーの指示により、シーン推定データ220a、b、cについて、「料理」を検索用のキーワードとして検索された例を示している。ここでは、シーン推定データ220a、b、cのそれぞれについて、シーンデータが存在する箇所を線により模式的に示している。また、「料理」のシーン名の箇所のみ示している。
図6Bでは、動画データ200a、b、cのそれぞれについて、検索されたシーンの箇所を抜き出して、合成シーンデータ240dに合成した例を示している。この合成シーンデータ240には、動画データ200a、b、cとシーン推定データ220a、b、cとについて、それぞれの抜き出されたシーンの箇所のデータが合成された動画データ200d及びシーン推定データ220dも含まれる。
以上により、本発明の実施の形態に係る動画データ解析処理を終了する。
図6Aでは、ユーザーの指示により、シーン推定データ220a、b、cについて、「料理」を検索用のキーワードとして検索された例を示している。ここでは、シーン推定データ220a、b、cのそれぞれについて、シーンデータが存在する箇所を線により模式的に示している。また、「料理」のシーン名の箇所のみ示している。
図6Bでは、動画データ200a、b、cのそれぞれについて、検索されたシーンの箇所を抜き出して、合成シーンデータ240dに合成した例を示している。この合成シーンデータ240には、動画データ200a、b、cとシーン推定データ220a、b、cとについて、それぞれの抜き出されたシーンの箇所のデータが合成された動画データ200d及びシーン推定データ220dも含まれる。
以上により、本発明の実施の形態に係る動画データ解析処理を終了する。
以上のように構成することで、以下のような効果を得ることができる。
従来、特許文献1の技術では、動画データからシーンを推定することはできなかった。
これに対して、本発明の実施の形態に係る動画データ解析装置1は、動画データ200の解析を行う動画データ解析装置であって、動画データ200から時系列に対応した静止画像データ300を取得する静止画像取得部100と、静止画像取得部100により取得された静止画像データ300に含まれるオブジェクトを認識し、認識されたオブジェクトのカテゴリーの認識を行い、認識結果時系列データ210を作成するカテゴリー認識部110と、カテゴリー認識部110により認識された認識結果時系列データ210を用いて、シーンの推定を行うシーン推定部130とを備えることを特徴とする。
このように構成することで、動画データ200からシーンを推定可能となる。そして、動画データ200からシーンを推定することで、このシーンを識別子として使用可能となる。すなわち、動画データ200について、細かい粒度での識別子を一律かつ正確に付与することが可能となる。これにより、動画データ200から、必要なシーンを容易に検索し、有効に活用することが可能になる。
従来、特許文献1の技術では、動画データからシーンを推定することはできなかった。
これに対して、本発明の実施の形態に係る動画データ解析装置1は、動画データ200の解析を行う動画データ解析装置であって、動画データ200から時系列に対応した静止画像データ300を取得する静止画像取得部100と、静止画像取得部100により取得された静止画像データ300に含まれるオブジェクトを認識し、認識されたオブジェクトのカテゴリーの認識を行い、認識結果時系列データ210を作成するカテゴリー認識部110と、カテゴリー認識部110により認識された認識結果時系列データ210を用いて、シーンの推定を行うシーン推定部130とを備えることを特徴とする。
このように構成することで、動画データ200からシーンを推定可能となる。そして、動画データ200からシーンを推定することで、このシーンを識別子として使用可能となる。すなわち、動画データ200について、細かい粒度での識別子を一律かつ正確に付与することが可能となる。これにより、動画データ200から、必要なシーンを容易に検索し、有効に活用することが可能になる。
また、動画データには、ユーザーがカテゴリー付けして識別子を付加することが可能なものが存在する。その動画データを活用するに際には、その識別子を頼りに内容を確認して活用することが可能である。
ここで、撮影日時のように自動的に付与することが可能な識別子と異なり、ユーザーが付加するカテゴリーのような識別子は、その妥当性が確保されている保障がなかった。
これに対して、本発明の実施の形態に係る動画データ解析装置1は、動画データ200の内容を解析し、適切なシーンを推定することで、このシーンを識別子として活用することが可能となる。これにより、ユーザーが適切な動画を容易に得られるようになることが期待できる。
ここで、撮影日時のように自動的に付与することが可能な識別子と異なり、ユーザーが付加するカテゴリーのような識別子は、その妥当性が確保されている保障がなかった。
これに対して、本発明の実施の形態に係る動画データ解析装置1は、動画データ200の内容を解析し、適切なシーンを推定することで、このシーンを識別子として活用することが可能となる。これにより、ユーザーが適切な動画を容易に得られるようになることが期待できる。
また、従来、YouTube(登録商標)に代表されるような動画投稿サイトにおいて、投稿者によって設定された動画のタイトルと、動画の内容が異なることがあった。完全に異なる動画であれば、ユーザーが閲覧した後に即座に判断し、再生を中断することができるものの、タイトルと動画の内容が若干ずれている場合などは、その動画を最後まで再生するか、シークバーを操作することで、目的の内容であるか確認する必要があった。
このように、タイトルと動画の内容が異なることを即座に判断できるか否かに関わらず期待した内容を得られないため、ユーザーは、これらの操作によりストレスがかかっていた。また、このような状況は、動画投稿サイトだけで特別に発生する状況ではなく、タイトルが適切に付与されていない一般的な動画データについても発生することがあった。
これに対して、本発明の実施の形態に係る動画データ解析装置1は、シーン推定部130は、シーンの推定の時系列データであるシーン推定データ220を集約し、動画データ200のタイトルを推定することを特徴とする。
このように構成し、全てのシーン推定から集約されたタイトルを推定することで、適切なタイトルを自動的に付与可能となる。
このように、タイトルと動画の内容が異なることを即座に判断できるか否かに関わらず期待した内容を得られないため、ユーザーは、これらの操作によりストレスがかかっていた。また、このような状況は、動画投稿サイトだけで特別に発生する状況ではなく、タイトルが適切に付与されていない一般的な動画データについても発生することがあった。
これに対して、本発明の実施の形態に係る動画データ解析装置1は、シーン推定部130は、シーンの推定の時系列データであるシーン推定データ220を集約し、動画データ200のタイトルを推定することを特徴とする。
このように構成し、全てのシーン推定から集約されたタイトルを推定することで、適切なタイトルを自動的に付与可能となる。
また、本発明の実施の形態に係る動画データ解析装置1は、カテゴリー認識部110がカテゴリーを認識するオブジェクトの種類は、物、人、性別、感情、姿勢、場所、色、向きであることを特徴とする。
このように構成することで、動画データ200に、映画におけるチャプター、シーンの種類、登場人物、場所のような細かな粒度での識別子を付与することができ、より動画データ200を有効活用可能となる。
このように構成することで、動画データ200に、映画におけるチャプター、シーンの種類、登場人物、場所のような細かな粒度での識別子を付与することができ、より動画データ200を有効活用可能となる。
また、本発明の実施の形態に係る動画データ解析装置1は、動画データ200に含まれる音声データ310について音声認識を行い、認識結果時系列データ210に時系列順に付加する音声認識付加部120を更に備え、認識結果時系列データ210は、動画データ200のタイムラインと同期する時系列について、認識結果を文字として含むことを特徴とする。
このように構成することで、動画データ200に含まれる静止画像データ300だけではなく、音声データ310も基にシーンを推定可能となるため、より適切なシーンの推定が可能となる。
このように構成することで、動画データ200に含まれる静止画像データ300だけではなく、音声データ310も基にシーンを推定可能となるため、より適切なシーンの推定が可能となる。
また、本発明の実施の形態に係る動画データ解析装置1は、複数の動画データ200から、シーン推定部130により推定された同一種類のシーンに対応する動画データ200を抜き出して合成するシーン合成部140を更に備えることを特徴とする。
このように構成することで、ユーザーが探したいシーンを抜き出した合成シーンデータ240を容易に得て活用することができる。
たとえば、上述の例では、「料理」の動画を抜き出してアーカイブして、必要なときに動画を閲覧して料理法を参照するといった活用が可能となる。
このように構成することで、ユーザーが探したいシーンを抜き出した合成シーンデータ240を容易に得て活用することができる。
たとえば、上述の例では、「料理」の動画を抜き出してアーカイブして、必要なときに動画を閲覧して料理法を参照するといった活用が可能となる。
〔他の実施の形態〕
なお、本発明の実施の形態においては、ユーザーに投稿された動画データ200について解析する例について記載していた。
しかしながら、本発明の動画データ解析装置1は、シーン推定やタイトル推定が必要な各種動画データ200について適用することが可能である。
たとえば、画像形成装置等に接続されたネットワークカメラの複数の動画データ200から、ユーザーが使用したり、トラブルがあったりした等のシーンを推定し、その使用やトラブルの種類をタイトルとして付与し、又は識別子にして抜き出して合成してもよい。また、使用するユーザーの男女別、年齢別、表情別等のカテゴリーに対応したシーンについても推定してもよい。この場合、ユーザーが使用した画像形成装置等の機能を、動画データ200のメタデータとして含ませておいてもよい。このように構成することで、画像形成装置等の使用状況やトラブルについての分析が容易となる。
また、ユーザーがウェアラブル(wearable)機器で撮像した、いわゆるライフログ(life log)等の動画データ200について、誰と出会い、何を話し、どのような行動を取ったといたシーンを推定して、タイトルを付与して格納しておくことも可能である。これにより、膨大なライフログの動画データ200から、必要なものを検索、抜き出して活用することが容易となる。
なお、本発明の実施の形態においては、ユーザーに投稿された動画データ200について解析する例について記載していた。
しかしながら、本発明の動画データ解析装置1は、シーン推定やタイトル推定が必要な各種動画データ200について適用することが可能である。
たとえば、画像形成装置等に接続されたネットワークカメラの複数の動画データ200から、ユーザーが使用したり、トラブルがあったりした等のシーンを推定し、その使用やトラブルの種類をタイトルとして付与し、又は識別子にして抜き出して合成してもよい。また、使用するユーザーの男女別、年齢別、表情別等のカテゴリーに対応したシーンについても推定してもよい。この場合、ユーザーが使用した画像形成装置等の機能を、動画データ200のメタデータとして含ませておいてもよい。このように構成することで、画像形成装置等の使用状況やトラブルについての分析が容易となる。
また、ユーザーがウェアラブル(wearable)機器で撮像した、いわゆるライフログ(life log)等の動画データ200について、誰と出会い、何を話し、どのような行動を取ったといたシーンを推定して、タイトルを付与して格納しておくことも可能である。これにより、膨大なライフログの動画データ200から、必要なものを検索、抜き出して活用することが容易となる。
また、上述の実施の形態においては、動画データ200に含まれる静止画像データ300と音声データ310とからシーンを推定する例について説明したものの、他のデータを用いることも可能である。
たとえば、動画のメタデータに含まれる位置情報、撮像機器の情報、撮影日付の情報等を、シーンの推定に用いてもよい。
これにより、シーンの推定の妥当性を向上させることが可能となる。
たとえば、動画のメタデータに含まれる位置情報、撮像機器の情報、撮影日付の情報等を、シーンの推定に用いてもよい。
これにより、シーンの推定の妥当性を向上させることが可能となる。
また、上記実施の形態の構成及び動作は例であって、本発明の趣旨を逸脱しない範囲で適宜変更して実行することができることは言うまでもない。
Claims (6)
- 動画データの解析を行う動画データ解析装置であって、
前記動画データから時系列に対応した静止画像データを取得する静止画像取得部と、
前記静止画像取得部により取得された前記静止画像データに含まれるオブジェクトを認識し、認識された前記オブジェクトのカテゴリーの認識を行い、認識結果時系列データを作成するカテゴリー認識部と、
前記カテゴリー認識部により認識された前記認識結果時系列データを用いて、シーンの推定を行うシーン推定部とを備える
ことを特徴とする動画データ解析装置。 - 前記シーン推定部は、
前記シーンの推定の時系列データであるシーン推定データを集約し、前記動画データのタイトルを推定する
ことを特徴とする請求項1に記載の動画データ解析装置。 - 前記カテゴリー認識部がカテゴリーを認識する前記オブジェクトの種類は、物、人、性別、感情、姿勢、場所、色、向きである
ことを特徴とする請求項1に記載の動画データ解析装置。 - 前記動画データに含まれる音声データについて音声認識を行い、前記認識結果時系列データに時系列順に付加する音声認識付加部を更に備え、
前記認識結果時系列データは、前記動画データのタイムラインと同期する時系列について、認識結果を文字として含む
ことを特徴とする請求項1に記載の動画データ解析装置。 - 複数の前記動画データから、前記シーン推定部により推定された同一種類のシーンに対応する前記動画データを抜き出して合成するシーン合成部を更に備える
ことを特徴とする請求項1に記載の動画データ解析装置。 - 動画データ解析装置により実行される動画データ解析方法であって、前記動画データ解析装置は、
動画データから時系列に対応した静止画像データを取得し、
取得された前記静止画像データに含まれるオブジェクトを認識し、認識された前記オブジェクトのカテゴリーの認識を行い、認識結果時系列データを作成し、
前記認識結果時系列データを用いて、特定区間のシーンの推定を行う
ことを特徴とする動画データ解析方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2016166855 | 2016-08-29 | ||
| JP2016-166855 | 2016-08-29 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018042959A1 true WO2018042959A1 (ja) | 2018-03-08 |
Family
ID=61300631
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2017/027139 Ceased WO2018042959A1 (ja) | 2016-08-29 | 2017-07-27 | 動画データ解析装置及び動画データ解析方法 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2018042959A1 (ja) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2000293685A (ja) * | 1999-04-06 | 2000-10-20 | Toyota Motor Corp | シーン認識装置 |
| JP2002010178A (ja) * | 2000-06-19 | 2002-01-11 | Sony Corp | 画像管理システム及び画像管理方法、並びに、記憶媒体 |
| JP2007295218A (ja) * | 2006-04-25 | 2007-11-08 | Nippon Hoso Kyokai <Nhk> | ノンリニア編集装置およびそのプログラム |
| JP2013098790A (ja) * | 2011-11-01 | 2013-05-20 | Canon Inc | 映像編集装置およびその制御方法 |
| JP2013164837A (ja) * | 2011-03-24 | 2013-08-22 | Toyota Infotechnology Center Co Ltd | シーン判定方法およびシーン判定システム |
-
2017
- 2017-07-27 WO PCT/JP2017/027139 patent/WO2018042959A1/ja not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2000293685A (ja) * | 1999-04-06 | 2000-10-20 | Toyota Motor Corp | シーン認識装置 |
| JP2002010178A (ja) * | 2000-06-19 | 2002-01-11 | Sony Corp | 画像管理システム及び画像管理方法、並びに、記憶媒体 |
| JP2007295218A (ja) * | 2006-04-25 | 2007-11-08 | Nippon Hoso Kyokai <Nhk> | ノンリニア編集装置およびそのプログラム |
| JP2013164837A (ja) * | 2011-03-24 | 2013-08-22 | Toyota Infotechnology Center Co Ltd | シーン判定方法およびシーン判定システム |
| JP2013098790A (ja) * | 2011-11-01 | 2013-05-20 | Canon Inc | 映像編集装置およびその制御方法 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN113569088B (zh) | 一种音乐推荐方法、装置以及可读存储介质 | |
| US8750681B2 (en) | Electronic apparatus, content recommendation method, and program therefor | |
| KR102290419B1 (ko) | 디지털 컨텐츠의 시각적 내용 분석을 통해 포토 스토리를 생성하는 방법 및 장치 | |
| US8938393B2 (en) | Extended videolens media engine for audio recognition | |
| US8959071B2 (en) | Videolens media system for feature selection | |
| US9189137B2 (en) | Method and system for browsing, searching and sharing of personal video by a non-parametric approach | |
| US8948515B2 (en) | Method and system for classifying one or more images | |
| US8804999B2 (en) | Video recommendation system and method thereof | |
| US8583647B2 (en) | Data processing device for automatically classifying a plurality of images into predetermined categories | |
| US20210117471A1 (en) | Method and system for automatically generating a video from an online product representation | |
| KR102199446B1 (ko) | 영상 컨텐츠 검색을 지원하는 영상 서비스 장치 및 영상 컨텐츠 검색 지원 방법 | |
| CN110119711A (zh) | 一种获取视频数据人物片段的方法、装置及电子设备 | |
| CN107077595A (zh) | 选择和呈现代表性帧以用于视频预览 | |
| CN103200463A (zh) | 一种视频摘要生成方法和装置 | |
| CN103207917B (zh) | 标注多媒体内容的方法、生成推荐内容的方法及系统 | |
| CN103534755B (zh) | 声音处理装置、声音处理方法、程序及集成电路 | |
| CN103514248B (zh) | 视频记录设备、信息处理系统、信息处理方法和记录介质 | |
| JP2016035607A (ja) | ダイジェストを生成するための装置、方法、及びプログラム | |
| US20100272411A1 (en) | Playback apparatus and playback method | |
| KR102472194B1 (ko) | Ai 기술을 활용한 개인 미디어 컨텐츠 장면분석 시스템 및 그 구동방법 | |
| KR101640317B1 (ko) | 오디오 및 비디오 데이터를 포함하는 영상의 저장 및 검색 장치와 저장 및 검색 방법 | |
| CN117319765A (zh) | 视频处理方法、装置、计算设备及计算机存储介质 | |
| KR20220095591A (ko) | 개인미디어 크리에이터를 위한 클라우드 기반 스튜디오 플랫폼 제공 시스템 | |
| CN118590714A (zh) | 视觉媒体数据处理方法、程序产品、存储介质及电子设备 | |
| JP2025130819A (ja) | 映像処理装置及び映像処理プログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17845968 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| NENP | Non-entry into the national phase |
Ref country code: JP |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17845968 Country of ref document: EP Kind code of ref document: A1 |