WO2014208581A1 - シナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラム - Google Patents

シナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラム Download PDF

Info

Publication number
WO2014208581A1
WO2014208581A1 PCT/JP2014/066800 JP2014066800W WO2014208581A1 WO 2014208581 A1 WO2014208581 A1 WO 2014208581A1 JP 2014066800 W JP2014066800 W JP 2014066800W WO 2014208581 A1 WO2014208581 A1 WO 2014208581A1
Authority
WO
WIPO (PCT)
Prior art keywords
scenario
video
music
scenario generation
generation system
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2014/066800
Other languages
English (en)
French (fr)
Inventor
広海 石先
服部 元
滝嶋 康弘
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
KDDI Corp
Original Assignee
KDDI Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by KDDI Corp filed Critical KDDI Corp
Priority to US14/901,348 priority Critical patent/US10104356B2/en
Publication of WO2014208581A1 publication Critical patent/WO2014208581A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/4302Content synchronisation processes, e.g. decoder synchronisation
    • H04N21/4307Synchronising the rendering of multiple content streams or additional data on devices, e.g. synchronisation of audio on a mobile phone with the video output on the TV screen
    • H04N21/43072Synchronising the rendering of multiple content streams or additional data on devices, e.g. synchronisation of audio on a mobile phone with the video output on the TV screen of multiple content streams on the same device
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N9/00Details of colour television systems
    • H04N9/79Processing of colour television signals in connection with recording
    • H04N9/80Transformation of the television signal for recording, e.g. modulation, frequency changing; Inverse transformation for playback
    • H04N9/802Transformation of the television signal for recording, e.g. modulation, frequency changing; Inverse transformation for playback involving processing of the sound signal
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H1/00Details of electrophonic musical instruments
    • G10H1/36Accompaniment arrangements
    • G10H1/361Recording/reproducing of accompaniment for use with an external source, e.g. karaoke systems
    • G10H1/368Recording/reproducing of accompaniment for use with an external source, e.g. karaoke systems displaying animated or moving pictures synchronized with the music or audio part
    • GPHYSICS
    • G11INFORMATION STORAGE
    • G11BINFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
    • G11B27/00Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
    • G11B27/02Editing, e.g. varying the order of information signals recorded on, or reproduced from, record carriers
    • G11B27/031Electronic editing of digitised analogue information signals, e.g. audio or video signals
    • GPHYSICS
    • G11INFORMATION STORAGE
    • G11BINFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
    • G11B27/00Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
    • G11B27/10Indexing; Addressing; Timing or synchronising; Measuring tape travel
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/41Structure of client; Structure of client peripherals
    • H04N21/422Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS]
    • H04N21/42203Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS] sound input device, e.g. microphone
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/439Processing of audio elementary streams
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/45Management operations performed by the client for facilitating the reception of or the interaction with the content or administrating data related to the end-user or to the client device itself, e.g. learning user preferences for recommending movies, resolving scheduling conflicts
    • H04N21/458Scheduling content for creating a personalised stream, e.g. by combining a locally stored advertisement with an incoming stream; Updating operations, e.g. for OS modules ; time-related management operations
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/475End-user interface for inputting end-user data, e.g. personal identification number [PIN], preference data
    • H04N21/4756End-user interface for inputting end-user data, e.g. personal identification number [PIN], preference data for rating content, e.g. scoring a recommended movie
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/47End-user applications
    • H04N21/482End-user interface for programme selection
    • H04N21/4828End-user interface for programme selection for searching programme descriptors
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/84Generation or processing of descriptive data, e.g. content descriptors
    • H04N21/8405Generation or processing of descriptive data, e.g. content descriptors represented by keywords
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/845Structuring of content, e.g. decomposing content into time segments
    • H04N21/8456Structuring of content, e.g. decomposing content into time segments by decomposing the content in the time domain, e.g. in time segments
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/85Assembly of content; Generation of multimedia applications
    • H04N21/854Content authoring
    • H04N21/8541Content authoring involving branching, e.g. to different story endings
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/85Assembly of content; Generation of multimedia applications
    • H04N21/854Content authoring
    • H04N21/8547Content authoring involving timestamps for synchronizing content
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/85Assembly of content; Generation of multimedia applications
    • H04N21/858Linking data to content, e.g. by linking an URL to a video object, by creating a hotspot
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N9/00Details of colour television systems
    • H04N9/79Processing of colour television signals in connection with recording
    • H04N9/87Regeneration of colour television signals
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N9/00Details of colour television systems
    • H04N9/79Processing of colour television signals in connection with recording
    • H04N9/87Regeneration of colour television signals
    • H04N9/89Time-base error compensation
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2220/00Input/output interfacing specifically adapted for electrophonic musical tools or instruments
    • G10H2220/005Non-interactive screen display of musical or status data
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2240/00Data organisation or data communication aspects, specifically adapted for electrophonic musical tools or instruments
    • G10H2240/075Musical metadata derived from musical analysis or for use in electrophonic musical instruments
    • G10H2240/081Genre classification, i.e. descriptive metadata for classification or selection of musical pieces according to style
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10HELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
    • G10H2240/00Data organisation or data communication aspects, specifically adapted for electrophonic musical tools or instruments
    • G10H2240/121Musical libraries, i.e. musical databases indexed by musical parameters, wavetables, indexing schemes using musical parameters, musical rule bases or knowledge bases, e.g. for automatic composing methods
    • G10H2240/131Library retrieval, i.e. searching a database or selecting a specific musical piece, segment, pattern, rule or parameter set
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N5/00Details of television systems
    • H04N5/76Television signal recording

Definitions

  • the present invention relates to a scenario generation system, a scenario generation method, and a scenario generation program for generating a scenario used for video playback synchronized with music playback.
  • background images of music and lyrics are played on the display screen.
  • the background video is seen not only by the singer but also by the person present at the venue, and is required to be a video suitable for music.
  • Patent Literature 1 connects the motion data to the motion data selected based on the music when generating the video content that matches the music change of the music. A video is being created.
  • Patent Document 2 displays a wide variety of karaoke background images by generating a karaoke background image using a plurality of images extracted from various images uploaded from mobile phones or the like when using karaoke.
  • the user can enjoy karaoke.
  • the device described in Patent Document 3 generates an image according to the meaning of the lyrics that occur as the musical performance progresses.
  • the above-described device generates a video of changing characters according to emotions, such as inflection, while outputting a selected desired musical piece.
  • Patent Document 4 enables intermittent playback of a long story video along with karaoke performance, and plays back video data intermittently in parallel in a time-sharing manner.
  • the present invention has been made in view of such circumstances, a scenario generation system and scenario generation capable of creating a scenario in a scene having a time-series order and reproducing a synchronized video with a natural impression corresponding to music reproduction. It is an object to provide a method and a scenario generation program.
  • a scenario generation system of the present invention is a scenario generation system used for video reproduction synchronized with music reproduction, a situation estimation unit for estimating a situation represented by the music, a time series
  • the video specifying unit that specifies at least one video that matches the estimated situation among the videos that are configured by scenes having the following order, and the scene that configures the specified video are associated with each music segment
  • a scenario generation unit for generating a scenario.
  • the scenario generation unit generates a scenario in which the identified video scene is associated with each piece of music while maintaining a time-series order. It is said. This makes it possible to generate a synchronized video scenario in which videos are matched for each music segment while maintaining the time-series order of the original video scene.
  • the scenario generation system of the present invention is characterized in that the situation estimation unit estimates the situation represented by the song by analyzing the lyrics of the song. In this way, by analyzing the lyrics of the music, it is possible to accurately match the video that matches the music.
  • the scenario generation system of the present invention includes an adjustment video specifying unit that extracts a keyword from the lyrics of the music for each of the music pieces and specifies an adjustment video using the extracted keyword. And a scenario adjusting unit that adjusts the scenario while maintaining the time-series order of the scene specified by the scenario using the adjustment video.
  • an adjustment video specifying unit that extracts a keyword from the lyrics of the music for each of the music pieces and specifies an adjustment video using the extracted keyword.
  • a scenario adjusting unit that adjusts the scenario while maintaining the time-series order of the scene specified by the scenario using the adjustment video.
  • the scenario generation system of the present invention is characterized in that the situation estimation unit estimates the situation represented by the music piece by estimating the contents of each category corresponding to the situation represented by the music piece. Thereby, for example, it is possible to estimate the situation represented by the music with 5W1H as each category and perform appropriate matching with the video.
  • the video adjustment unit compares the scene specified in the scenario and the video for adjustment based on a preset criterion for each segment of the music. , One of them is newly adopted as a scene specified in the scenario. Thereby, it is possible to create a scenario by adopting a preferable video suitable for music.
  • the scenario generation system of the present invention collects evaluations for the synchronized video from viewers of the synchronized videos based on the generated scenario, and stores the collected evaluations in the evaluation information DB. And an evaluation information DB for storing the collected evaluations, and the video adjustment unit determines a combination of scenes specified in the scenario based on the evaluation information stored in the evaluation information DB. It is characterized by correction. Thereby, a scenario can be generated reflecting the user's evaluation.
  • the scenario generation system of the present invention is characterized by including a playback unit that plays back the synchronized video based on the generated scenario.
  • a playback unit that plays back the synchronized video based on the generated scenario.
  • the scenario generation method of the present invention is a scenario generation method for generating a scenario used for video reproduction synchronized with music reproduction, and includes a step of estimating a situation represented by the music and a time-series order.
  • a step of identifying at least one video suitable for the estimated situation among videos composed of scenes, and a step of generating a scenario in which scenes constituting the identified video are associated with each music segment It is characterized by including.
  • a scenario can be created with a scene having a time-series order, and a synchronized video having a natural impression corresponding to music reproduction can be reproduced based on the scenario.
  • the scenario generation program of the present invention is a scenario generation program that generates a scenario used for video reproduction synchronized with music reproduction, and has a process for estimating a situation represented by the music and a time-series order.
  • a scenario can be created with a scene having a time-series order, and a synchronized video having a natural impression corresponding to music reproduction can be reproduced based on the scenario.
  • a scenario can be created with a scene having a time-series order, and a synchronized video with a natural impression corresponding to music reproduction can be reproduced based on the scenario.
  • the scenario generation system of the present invention generates a scenario used for video playback synchronized with music playback. That is, the contents of the song lyrics are selected by selecting the main axis video based on the music status information estimated from the music lyrics (for example, 5W1H information) and combining the main axis video and the adjustment video specified by the keyword extracted from the lyrics line.
  • a scenario that constitutes a background video corresponding to the is generated.
  • a scenario is a video playlist for combining and reproducing a plurality of videos according to the contents when generating videos in synchronization with lyrics.
  • FIG. 1 is a block diagram showing a configuration of the scenario generation system 100.
  • the video DB 110 stores video material. Meta information is added to the video in advance. The meta information is given, for example, for the entire video material or for each scene, and information included in the video, such as person details, seasons, and time zones, is stored in association with text information.
  • a scene means a plurality of short-time videos constituting a video, and is not necessarily limited to one scene or one cut. The association between the video material and the meta information will be described later.
  • the song DB 120 stores song lyrics, sound source files, and meta information.
  • the meta information of music includes, for example, genre, title, artist, and lyrics display time for music time.
  • the genre and the like can be estimated in advance. For example, it is possible to use overall impression word estimation based on lyrics and genre determination based on acoustic data.
  • the overall impression word can be estimated using the number of images searched by the image search engine and the number of unique image contributors.
  • genre determination based on acoustic data can be performed by classifying objects represented by dimension vectors by SVM (Support Vector Vector Machine) that classifies unknown data into each impression category based on learning data (Non-Patent Document). 1 and 2).
  • SVM Small Vector Vector Machine
  • the situation estimation unit 130 analyzes the lyrics of the music and estimates the situation represented by the music. In this way, by analyzing the lyrics of the music, it is possible to estimate the scene or situation represented by the music lyrics. As a result, it is possible to accurately identify a video that matches the music.
  • the situation estimation unit 130 estimates the situation represented by the song by estimating the contents of each category corresponding to the situation represented by the song. For such estimation, the same meta information as the meta information given to the video DB 110 in advance or a category corresponding to 5W1H can be used.
  • the video specifying unit 140 specifies at least one main-axis video that matches the estimated situation among videos configured by scenes having a time-series order. Thereby, a synchronized video scenario can be created in a scene having a time-series order, and a natural-synchronized synchronized video corresponding to music playback can be played back based on the scenario.
  • a plurality of videos may be identified and used as the main axis video. For example, if an OR search is performed from keywords selected from 5W1H, such as male, spring, you, island country, etc., and about 100 videos can be identified, the degree of association between the meta information about the 100 videos and the keywords can be calculated. Since it can be ranked, the higher rank can be selected as the main axis video.
  • the scenario generation unit 150 generates a scenario in which scenes constituting the main axis video are associated with each music segment. That is, the scenes identified as the main-axis videos are combined for each music segment. In that case, it is preferable to generate the scenario while maintaining the time-series order of the scenes constituting the main axis video. This makes it possible to generate a synchronized video scenario in which scenes are matched for each music segment while maintaining the time-series order of a plurality of scenes included in the original main-spindle video.
  • the combination of a series of synchronized images is performed by creating a scenario format.
  • the scenario format will be described later.
  • the scene is played until the time of the music segment elapses from the beginning of the scene. If there is not enough video in the associated scene, it can be supplemented with video for adjustment.
  • the adjustment video specifying unit 160 extracts a keyword from the lyrics of the music for each music segment, and specifies the video for adjustment using the extracted keyword.
  • the video for adjustment may be a set of scenes or a single scene. Basically, since the keyword of lyrics is used, the original scenario is used without adjustment for the music segment (prelude and interlude) where the lyrics do not exist.
  • the scenario adjustment unit 170 uses the adjustment video to adjust the scenario while maintaining the time-series order of the video specified by the scenario. That is, the degree of association between the adjustment video and the scene in the main axis video is evaluated, and the main axis video is corrected according to the evaluation. As a result, the original scenario can be complemented and a variety of scenarios more suitable for music can be created.
  • the scenario adjustment unit 170 compares the scene specified by the scenario and the adjustment video based on preset criteria for each music segment, and adopts either one as a new scene specified by the scenario. It is preferable to do. Thereby, it is possible to create a scenario by adopting a preferable video suitable for music.
  • the scenario adjustment unit 170 corrects the combination of scenes specified by the scenario based on the information stored in the evaluation information DB 195. Thereby, the scenario of a synchronous image
  • the playback unit 180 plays back the synchronized video based on the scenario in synchronization with the playback of the generated music. As a result, it is possible to reproduce a synchronized video suitable for music created by automatically combining videos. For example, it can be used for background images of karaoke.
  • the evaluation information collection unit 190 collects the evaluation of the synchronized video from the viewer of the synchronized video based on the generated scenario, and stores the collected evaluation in the evaluation information DB 195.
  • the evaluation information can be obtained as a binary value of GOOD or BAD, or an evaluation value in five stages. It is also possible to input an evaluation value for the entire video and for each scene. The input evaluation information is stored in the evaluation information DB 195 in association with the scene, the entire video, or the scenario.
  • the evaluation information DB 195 stores evaluation information for the synchronized video reproduced by the generated scenario.
  • the evaluation information stored in the evaluation information DB 195 is extracted and used.
  • FIG. 2 is a flowchart showing the operation of the scenario generation system. Hereinafter, description will be given along steps S1 to S7 shown in FIG.
  • (S1) Search of main axis video The video that becomes the main axis of the background video is searched from the video DB 110.
  • a video can be searched using a keyword included in the music status information.
  • filtering it is possible to search only for videos / scenes including a person, only videos / scenes that do not include a person, and the like.
  • S2 Music division division Based on the lyrics information, the music is divided and an ID is assigned to each music division.
  • ID For example, a unique ID such as Pp: Prelude, P1-Pn: Paragraph, Pi: Interlude, Pe: Postlude is assigned to the music segment.
  • P1-1 may be the ID of the first three lines
  • P1-2 may be the ID of the remaining two lines. The combination is not limited to this.
  • (S4) Video Search Using Lyrics KW Video is searched using KW (keyword) extracted from the lyrics line included in the music segment P.
  • KW keyword
  • a search result can be narrowed down by using a unique image with a large number of contributors (see Non-Patent Document 1) or filtering meta information of videos.
  • one video can be searched for a category by randomly selecting a narrowing result.
  • (S7) Scenario generation Eventually, a video to be played back by music segment is recorded.
  • the created scenario is recorded as a scenario format, and scenario information can be generated in the format.
  • the above operation can be performed by executing a program.
  • FIG. 3 is a diagram showing a processing image of scenario generation and adjustment.
  • the scenario generation system associates each video of the main-axis video (video consisting of scenes of Sp to Se) with the music divisions.
  • the video is searched with the keyword obtained from the lyrics line, and videos K1 to K4 are obtained.
  • the obtained videos K1 to K4 are compared with S1, S2, S3, and S4 to evaluate the degree of association.
  • the relevance is S1 ⁇ K1, S2> K2, S3> K3, and S4 ⁇ K4
  • the video specified in the scenario is replaced with a video with high relevance (S1 ⁇ K1, S4 ⁇ K4).
  • FIG. 4 is a diagram showing the relationship between meta information for the entire video and meta information for each scene.
  • the meta information associated with the video material stored in the video DB 110 includes information for the entire material and information for each scene.
  • meta information for the whole material can be managed with an ID of 001
  • meta information for each scene can be managed with 001, 1, 2, etc.
  • FIG. 5 is a diagram showing a table for associating video materials with meta information.
  • the video identified by ID001 cannot be changed, the subject is “Cherry Blossom”, “Blue Sky”, the place is “Cherry Blossom Tree”, the season is “Spring”, and the time is “Morning”.
  • “Noon” is associated with meta-information that no person is shown.
  • FIG. 6 is a flowchart illustrating a processing example of music status information estimation.
  • the classifier is first learned in advance (step T1).
  • morphological analysis is performed on the lyrics of the specific paragraph (step T2).
  • WHO estimation, WHAT estimation, WHEN estimation, WHERE estimation, WHY estimation, and HOW estimation are performed based on the morphologically analyzed words (steps T3 to T8).
  • the important words of the lyrics in the specific paragraph are extracted (step T9).
  • the processing is completed for all the paragraphs, the processing is completed.
  • the above operation can be performed by executing a program.
  • FIG. 7 is a diagram schematically showing impression estimation processing received from lyrics. As shown in FIG. 7, the entire impression classifier group is learned in advance by the learning lyrics DB. Then, feature words are extracted from the lyrics and applied to the overall impression classifier. The impression received from the lyrics is estimated in consideration of the relevance of the words of the obtained impression words.
  • FIG. 8 is a table showing examples of categories that are set in advance and words representing the state of each category. For example, when the result of the “male” classifier is positive and the result of the “company employee” classifier is positive for the WHO item, “male” and “company employee” are used as the WHO information of the situation information. .
  • the discriminator can be configured using SVM (see Non-Patent Document 2).
  • the music status information estimation process can be divided and applied to the entire lyrics, each paragraph, each lyric line, and the like. It is not always necessary to use all of the estimation results for matching. On the other hand, each combination can also be used.
  • FIG. 9 is a schematic diagram illustrating an example of video specifying processing. As shown in FIG. 9, when searching for videos with 5M1H as the main keywords “male”, “spring”, “island”, “country”, “you”, “enka”, “hometown”, videos 1 to 3 are displayed. Searched as a candidate. Of these, the keyword and the meta information of the video most closely match the video 1, and the video 1 has a lot of related meta information, and there is no conflicting meta information.
  • FIG. 10 is a table showing an example scenario format.
  • the scenario generation unit 150 can generate scenario information in a format as shown in FIG. 10, for example.
  • a table 210 in FIG. 10 is a table in which a song name and meta information are associated with a song ID.
  • the sections Pp to Pe indicate music sections.
  • the content of the scenario is specified by the start time, the prelude or the like, the lyrics, the video ID of the main axis video, and the video ID of the adjustment video for each of the sections Pp to Pe.
  • the scenario generation system described above is preferably configured as a client-server type system in which a configuration other than the playback unit 180 is a server and only the playback unit 180 is a client, in consideration of processing capability and processing speed.
  • a form in which streaming is used by using only a database as a server and another structure as a client may be used, or a form in which all the structures are used as clients may be used.
  • the separation of the server and the client is not limited.
  • the video basically indicates a moving image, but may be a still image. In particular, it is easy to use as a video to be played in a time of about one line of lyrics even if it is a still image, since it does not cause a sense of incompatibility.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computer Security & Cryptography (AREA)
  • Human Computer Interaction (AREA)
  • Databases & Information Systems (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Reverberation, Karaoke And Other Acoustics (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
  • Television Signal Processing For Recording (AREA)

Abstract

 時系列の順序を有するシーンでシナリオを作成でき、楽曲再生に対応した自然な印象の同期映像を再生できるシナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラムを提供する。楽曲再生と同期した映像再生に用いられるシナリオ生成システム100であって、楽曲が表わす状況を推定する状況推定部130と、時系列の順序を有するシーンにより構成される映像のうち、推定された状況に合った映像を少なくとも一つ特定する映像特定部140と、特定された映像を構成するシーンを楽曲の区分ごとに関連付けたシナリオを生成するシナリオ生成部150と、を備える。これにより、時系列の順序を有するシーンでシナリオを作成でき、そのシナリオに基づいて楽曲再生に対応した自然な印象の同期映像を再生できる。

Description

シナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラム
 本発明は、楽曲再生と同期した映像再生に用いられるシナリオを生成するシナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラムに関する。
 現在利用されているカラオケシステムでは、表示画面に楽曲、歌詞の背景映像が再生される。背景映像は、歌唱者のみならず、その場に同席する者が目にするものであり、楽曲に適した映像であることが求められている。
 例えば、特許文献1記載の装置は、楽曲の音楽の変化に合わせた映像コンテンツを生成する際に、楽曲に基づいて選択された動きデータを対象にして、動きデータ間を連結することで、背景映像を作成している。
 特許文献2記載のシステムは、カラオケ利用時に、携帯等からアップロードされた様々な画像の中から抽出した複数の画像を用いてカラオケ背景映像を生成することで、バラエティに富んだカラオケ背景映像を表示しながらユーザがカラオケを楽しめるようにしている。
 特許文献3記載の装置は、演奏曲の進行に伴って生じる歌詞の意味に応じて、映像を生成する。上記の装置は、選択された所望の演奏曲を出力させるとともに、抑揚等、感情に合わせて文字を変化させる映像を作成している。
 特許文献4記載のシステムは、カラオケ演奏とともに長編ストーリー映像を断続再生できるようにし、映像データを時分割的に並行して断続再生している。
特開2010-044484号公報 特開2011-075701号公報 特開2004-354423号公報 特開2003-108155号公報
舟澤慎太郎、石先広海、帆足啓一郎、滝嶋康弘、甲藤二郎、「歌詞特徴を考慮したWeb画像と楽曲同期再生システムの提案」,FIT2009  第8回情報科学技術フォーラム  講演論文集,社団法人情報処理学会、社団法人電子情報通信学会,2009年 8月20日,第2分冊,pp.333~334 舟澤慎太郎, 石先広海, 帆足啓一郎, 滝嶋康弘, 甲藤二郎、「歌詞の印象に基づく楽曲検索のための楽曲自動分類に関する検討」, 第71 回情報処理学会全国大会, 5R-2 (2009). Thorsten Joachims, "SVMLIGHT",[online], 2008年8月14日, Cornell University,[2013年6月18日検索], インターネット〈URL:http://svmlight.joachims.org/〉 "Mecab" ,[online],[2013年6月18日検索], インターネット〈URL: http://mecab.googlecode.com/svn/trunk/mecab/doc/index.html〉
 しかしながら、上記のような装置では、データベースに格納された素材映像を組み合わせて多様な背景映像を作成したりすることは困難である。すなわち、時系列を維持した順序で楽曲の各区分に適した映像で同期映像を再生できず、楽曲再生に対応した自然な印象の同期映像を再生することができない。その結果、楽曲の歌詞の内容に応じて背景映像を変化させて背景映像にストーリー性を確保するのは難しくなる。
 本発明は、このような事情に鑑みてなされたものであり、時系列の順序を有するシーンでシナリオを作成でき、楽曲再生に対応した自然な印象の同期映像を再生できるシナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラムを提供することを目的とする。
 (1)上記の目的を達成するため、本発明のシナリオ生成システムは、楽曲再生と同期した映像再生に用いられるシナリオ生成システムであって、楽曲が表わす状況を推定する状況推定部と、時系列の順序を有するシーンにより構成される映像のうち、前記推定された状況に合った映像を少なくとも一つ特定する映像特定部と、前記特定された映像を構成するシーンを前記楽曲の区分ごとに関連付けたシナリオを生成するシナリオ生成部と、を備えることを特徴としている。これにより、時系列の順序を有するシーンでシナリオを作成でき、そのシナリオに基づいて楽曲再生に対応した自然な印象の同期映像を再生できる。
 (2)また、本発明のシナリオ生成システムは、前記シナリオ生成部が、前記特定された映像のシーンを時系列の順序を保持しつつ前記楽曲の区分ごとに関連付けたシナリオを生成することを特徴としている。これにより、もとの映像のシーンが有する時系列の順序を維持したまま、さらに楽曲の区分ごとに映像をマッチさせた同期映像のシナリオを生成できる。
 (3)また、本発明のシナリオ生成システムは、前記状況推定部が、楽曲の歌詞を解析して前記楽曲が表わす状況を推定することを特徴としている。このように、楽曲の歌詞を解析することで的確に楽曲に合った映像を対応させることができる。
 (4)また、本発明のシナリオ生成システムは、前記楽曲の区分ごとに前記楽曲の歌詞からキーワードを抽出し、前記抽出されたキーワードを用いて調整用の映像を特定する調整用映像特定部と、前記調整用の映像を用いて、前記シナリオで特定されるシーンの時系列の順序を保持しつつ、前記シナリオを調整するシナリオ調整部と、を更に備えることを特徴としている。これにより、調整用の映像を用いてもとのシナリオを補完し、多様でより適切な同期映像のシナリオを作成することができる。
 (5)また、本発明のシナリオ生成システムは、前記状況推定部が、楽曲が表わす状況に相当する各カテゴリの内容を推定することで前記楽曲が表わす状況を推定することを特徴としている。これにより、例えば5W1Hを各カテゴリとして楽曲が表わす状況を推定し、映像との適切なマッチングを行なうことができる。
 (6)また、本発明のシナリオ生成システムは、前記映像調整部が、前記楽曲の区分ごとに、前記シナリオで特定されるシーンと前記調整用の映像とをあらかじめ設定した基準に基づいて比較し、どちらか一方を、新たに前記シナリオで特定されるシーンとして採用することを特徴としている。これにより、より楽曲に適した好ましい映像を採用してシナリオを作成することができる。
 (7)また、本発明のシナリオ生成システムは、前記生成されたシナリオによる同期映像の視聴者から前記同期映像に対する評価を収集し、前記収集された評価を前記評価情報DBに格納させる評価情報収集部と、前記収集された評価を格納する評価情報DBと、を更に備え、前記映像調整部は、前記評価情報DBに格納された評価情報に基づいて、前記シナリオで特定されるシーンの組合せを修正することを特徴としている。これにより、ユーザの評価を反映してシナリオを生成できる。
 (8)また、本発明のシナリオ生成システムは、前記生成されたシナリオに基づいて同期映像を再生する再生部を備えることを特徴としている。これにより、自動で映像を組み合わせて作成された楽曲に適した同期映像を再生できる。例えば、カラオケの背景映像に利用できる。
 (9)また、本発明のシナリオ生成方法は、楽曲再生と同期した映像再生に用いられるシナリオを生成するシナリオ生成方法であって、楽曲が表わす状況を推定するステップと、時系列の順序を有するシーンにより構成される映像のうち、前記推定された状況に合った映像を少なくとも一つ特定するステップと、前記特定された映像を構成するシーンを前記楽曲の区分ごとに関連付けたシナリオを生成するステップと、を含むことを特徴としている。これにより、時系列の順序を有するシーンでシナリオを作成でき、そのシナリオに基づいて楽曲再生に対応した自然な印象の同期映像を再生できる。
 (10)また、本発明のシナリオ生成プログラムは、楽曲再生と同期した映像再生に用いられるシナリオを生成するシナリオ生成プログラムであって、楽曲が表わす状況を推定する処理と、時系列の順序を有するシーンにより構成される映像のうち、前記推定された状況に合った映像を少なくとも一つ特定する処理と、前記特定された映像を構成するシーンを前記楽曲の区分ごとに関連付けたシナリオを生成する処理と、をコンピュータに実行させることを特徴としている。これにより、時系列の順序を有するシーンでシナリオを作成でき、そのシナリオに基づいて楽曲再生に対応した自然な印象の同期映像を再生できる。
 本発明によれば、時系列の順序を有するシーンでシナリオを作成でき、そのシナリオに基づいて楽曲再生に対応した自然な印象の同期映像を再生できる。
本発明のシナリオ生成システムの構成を示すブロック図である。 本発明のシナリオ生成システムの動作を示すフローチャートである。 シナリオの生成および調整の処理イメージを示す図である。 映像全体に対するメタ情報とシーン毎のメタ情報の関係を示す図である。 映像素材とメタ情報とを関連付けるテーブルを示す図である。 楽曲状況情報推定の処理例を示すフローチャートである。 歌詞から受ける印象の推定処理を模式的に示す図である。 事前に設定するカテゴリと、カテゴリ毎の状態を表わす単語の例を示すテーブルである。 映像特定の処理の一例を示す模式図である。 シナリオフォーマット例を示すテーブルである。
 次に、本発明の実施の形態について、図面を参照しながら説明する。説明の理解を容易にするため、各図面において同一の構成要素に対しては同一の参照番号を付し、重複する説明は省略する。
 (シナリオ生成システムの構成)
 本発明のシナリオ生成システムは、楽曲再生と同期した映像再生に用いられるシナリオを生成する。すなわち、楽曲歌詞から推定した楽曲状況情報(例えば5W1Hの情報)に基づいて主軸映像を選定し、主軸映像と、歌詞行から抽出したキーワードによって特定した調整用の映像を組み合わせることで楽曲歌詞の内容に応じた背景映像を構成するシナリオを生成する。シナリオとは、歌詞と同期して映像を生成する際に、内容に応じた複数の映像を組み合わせて再生するための、映像プレイリストである。
 (各部の構成)
 図1は、シナリオ生成システム100の構成を示すブロック図である。映像DB110は、映像素材を格納する。あらかじめ映像には、メタ情報が付与される。メタ情報は、例えば、映像素材全体、もしくはシーンごとに付与されており、人物詳細、季節、時間帯、等映像に含まれる情報がテキスト情報として関連付けられて保存されている。シーンは、映像を構成する複数の短時間の映像を意味し、必ずしも1つの場面や1カットに限定されない。なお、映像素材とメタ情報との関連付けについては後述する。
 楽曲DB120は、楽曲の歌詞、音源ファイル、メタ情報を格納する。楽曲のメタ情報には、例えば、ジャンル、タイトル、アーティスト、楽曲時刻に対する歌詞表示時間が含まれる。ジャンル等は、事前に推定することも可能である。例えば、歌詞による全体印象語推定と音響データによるジャンル判定を利用することができる。また、全体印象語は、画像検索エンジンにより検索された画像数やユニークな画像投稿者数を用いて推定できる。そして、音響データによるジャンル判定は、学習データをもとに未知のデータを各印象カテゴリに分類するSVM(Support Vector Machine)により次元ベクトルで表現されたオブジェクトを分類することで実行できる(非特許文献1、2参照)。一度生成したシナリオ等も、楽曲に関連付けて格納できる。通信機能により格納される楽曲の追加、情報の更新等を行なってもよい。
 状況推定部130は、楽曲の歌詞を解析して楽曲が表わす状況を推定する。このように、楽曲の歌詞をテキスト解析することで楽曲歌詞が表現する情景や状況等を推定できる。その結果、的確に楽曲に合った映像を特定できる。
 状況推定部130は、楽曲が表わす状況に相当する各カテゴリの内容を推定することで楽曲が表わす状況を推定することが好ましい。このような推定には、事前に映像DB110に付与するメタ情報と同一のメタ情報や、5W1Hに該当するカテゴリを利用できる。
 例えば、歌詞の全ての段落に対して、重要語を抽出し、5W1Hの各項目に該当するカテゴリを事前に設定する。そして、設定されたカテゴリ内の状態単語(男、女、10代等)毎にSVM等の識別器を用意し、歌詞入力結果が正となった情報を楽曲状況情報として利用することができる。なお、5W1Hの全部を用いずに、2つあるいは3つのみを重視することで楽曲状況を推定してもよい。
 映像特定部140は、時系列の順序を有するシーンにより構成される映像のうち、推定された状況に合った主軸映像を少なくとも一つ特定する。これにより、時系列の順序を有するシーンで同期映像のシナリオを作成でき、そのシナリオにより楽曲再生に対応した自然な印象の同期映像を再生することができる。なお、複数の映像を特定して、主軸映像としてもよい。例えば、男性、春、君、島国、等5W1Hから選ばれたキーワードからOR検索し、100件程度の映像を特定できたとして、その100件についているメタ情報とキーワードとの関連度を算出すれば順位付けできるため、その上位を主軸映像として選択できる。
 シナリオ生成部150は、主軸映像を構成するシーンを楽曲の区分ごとに関連付けたシナリオを生成する。すなわち、各楽曲区分に対して主軸映像として特定された各シーンを組み合わせる。その際には、主軸映像を構成するシーンの時系列の順序を保持しつつシナリオを生成することが好ましい。これにより、もとの主軸映像が有する複数のシーンの時系列の順序を維持したまま、楽曲の区分ごとにシーンをマッチさせた同期映像のシナリオを生成できる。
 同期映像の一連の映像の組み合わせは、シナリオフォーマットの作成により行なわれる。シナリオフォーマットについては後述する。また、各シーンを各楽曲区分に関連付ける際に楽曲区分よりシーンの再生時間が多い場合があるが、その場合にはシーンの最初から楽曲区分の時間が経過するまで再生することになる。関連付けられたシーンでは映像が足らない場合には調整用の映像で補完できる。
 調整用映像特定部160は、楽曲の区分ごとに楽曲の歌詞からキーワードを抽出し、抽出されたキーワードを用いて調整用の映像を特定する。調整用の映像は、シーンの集合であってもよいし、単体のシーンであってもよい。基本的に歌詞のキーワードを用いるため、歌詞が存在しない楽曲区分(前奏、間奏)は、調整せずもとのシナリオを用いることになる。
 シナリオ調整部170は、調整用の映像を用いて、シナリオで特定される映像の時系列の順序を保持しつつ、シナリオを調整する。すなわち、調整用の映像と主軸映像内のシーンとの関連度を評価して、評価に応じて主軸映像を修正する。これにより、もとのシナリオを補完し、多様でより楽曲に合ったシナリオを作成することができる。
 シナリオ調整部170は、楽曲の区分ごとに、シナリオで特定されるシーンと調整用の映像とをあらかじめ設定した基準に基づいて比較し、どちらか一方を、新たにシナリオで特定されるシーンとして採用することが好ましい。これにより、より楽曲に適した好ましい映像を採用してシナリオを作成することができる。
 また、シナリオ調整部170は、評価情報DB195に格納された情報に基づいて、シナリオで特定されるシーンの組合せを修正することが好ましい。これにより、ユーザの評価を反映して同期映像のシナリオを生成できる。
 再生部180は、生成された楽曲の再生に同期させてシナリオに基づき同期映像を再生する。これにより、自動で映像を組み合わせて作成された楽曲に適した同期映像を再生できる。例えば、カラオケの背景映像に利用できる。
 評価情報収集部190は、生成されたシナリオに基づく同期映像の視聴者から同期映像に対する評価を収集し、収集された評価を評価情報DB195に格納する。例えば、評価情報はGOODまたはBADの2値、または5段階で評価値を取得することができる。また、映像全体、シーンごとに評価値を入力することもできる。入力された評価情報は、シーン、映像全体またはシナリオと関連付けられて、評価情報DB195に格納される。
 評価情報DB195は、生成されたシナリオにより再生された同期映像に対する評価情報を格納する。評価に応じて同期映像のシナリオを調整する際には、評価情報DB195に格納された評価情報が抽出され、用いられる。
 (シナリオ生成システムの動作)
 次に、上記のように構成されたシナリオ生成システムの動作を説明する。図2は、シナリオ生成システムの動作を示すフローチャートである。以下、図2に示すステップS1~S7に沿って説明する。
 (S1)主軸映像の検索
 背景映像の主軸となる映像を映像DB110より検索する。例えば、楽曲状況情報に含まれるキーワードを利用して、映像を検索することができる。歌詞の全体印象語と歌詞の行から得られる印象語から楽曲の状況にあった1映像を検索してもよい(非特許文献1参照)。また、フィルタリングにより、人物を含む映像・シーンのみ、人物を含まない映像・シーンのみ等として検索することもできる。
 (S2)楽曲区分分割
 歌詞情報に基づいて、楽曲を分割し、各楽曲区分にIDを付与する。例えば、Pp:前奏、P1-Pn:段落、Pi:間奏、Pe:後奏のような固有IDを楽曲区分に付与する。これ以外に、段落を複数に分割することもできる。例えば、P1に5行含まれていた場合、P1-1を最初の3行文のIDとし、P1-2を残りの2行分のIDとしてもよい。なお、組合せはこの限りではない。
 (S3)シーン組合せ候補算出
 主軸映像のシーン情報において、時系列の順序を保持しつつ、全ての組合せをリストアップする。例えば、Pp:S1,P1:S3,P2:S4,…,Pe:S39を主軸映像のシーンの組合せとしてリストアップする。
 (S4)歌詞KWによる映像検索
 楽曲区分Pに含まれる歌詞行から抽出したKW(キーワード)を利用して、映像を検索する。複数の映像が検索された場合には、ユニークな投稿者数の多い画像を用いたり(非特許文献1参照)、映像のメタ情報をフィルタリングすることで、検索結果を絞りこむことができる。また絞り込み結果に対してランダムに選択する等で、区分に対して一つの映像を検索することができる。
 (S5)関連度比較
 主軸映像内で組み合わされたシーンと、歌詞KWにより検索された映像とを、各分割区分を単位として、比較基準に基づいて比較し、比較基準に適合、もしくは数値が高いものを選択する。比較基準は例えば、以下の(a)~(e)等を利用できる。
(a)利用回数が少ない物
(b)関連度(非特許文献1参照)が低いもの
(c)登録日時が新しいもの
(d)メタ情報が多く付与されているもの
(e)評価情報DB195の評価数もしくは評価値が高いもの
 (S6)評価情報による組合せ修正
 評価情報DB195に格納された評価値が低いシーンまたはシーンの組み合わせが含まれている場合には、検索されたシーンを利用せずに、他のシーンと置き換えることができる。例えば、関連度比較処理で次点となったシーンや、検索結果が次点となったシーンを順次利用する。
 (S7)シナリオ生成
 最終的に、楽曲区分で再生する映像を記録する。作成されたシナリオは、シナリオフォーマットとして記録され、フォーマットでシナリオ情報を生成できる。なお、上記のような動作は、プログラムを実行することで可能になる。
 (全体の処理イメージ)
 図3は、シナリオの生成および調整の処理イメージを示す図である。図3に示すように、まず、シナリオ生成システムは、楽曲の分割区分に対して、主軸映像(Sp~Seのシーンからなる映像)の各映像を関連付けている。このようにして得られたシナリオに対して、歌詞行から得られたキーワードで映像を検索し、映像K1~K4を得る。そして、得られた映像K1~K4について、それぞれS1、S2、S3、S4と対比し関連度を評価する。関連度を比較して関連度がS1<K1、S2>K2、S3>K3、S4<K4であった場合には、シナリオで特定される映像を関連度の高い映像に入れ替える(S1⇔K1、S4⇔K4)。
 (映像素材とメタ情報との関連付け)
 図4は、映像全体に対するメタ情報とシーン毎のメタ情報の関係を示す図である。図4に示すように、映像DB110に格納される映像素材に関連付けられるメタ情報には、素材全体向けのものと、各シーン向けのものがある。例えば、素材全体向けのメタ情報を001のIDで管理し、各シーン向けのメタ情報を001-1,2等で管理できる。
 図5は、映像素材とメタ情報とを関連付けるテーブルを示す図である。例えば、ID001で特定される映像に対しては、変更不可、被写体が「桜」、「青空」であり、場所が「桜並木」であり、季節が「春」であり、時間が「朝」、「昼」であり、人物は映っていないというメタ情報が関連付けられている。
 (楽曲状況情報推定の処理例)
 図6は、楽曲状況情報推定の処理例を示すフローチャートである。図6に示すように、まず、予め識別器の学習を行なう(ステップT1)。次に、特定段落の歌詞について形態素解析を行なう(ステップT2)。そして、形態素解析された単語をもとに、WHO推定、WHAT推定、WHEN推定、WHERE推定、WHY推定、HOW推定を行なう(ステップT3~T8)。推定結果に基づいて特定段落の歌詞の重要語を抽出する(ステップT9)。そして、全ての段落について処理が終了したかを確認し、全ての段落について処理が終わってない場合にはステップT2に戻る。全ての段落について処理が終了した場合には処理を完了する。なお、上記のような動作は、プログラムを実行することで可能になる。
 図7は、歌詞から受ける印象の推定処理を模式的に示す図である。図7に示すように、予め学習用歌詞DBにより全体印象識別器群を学習させる。そして、歌詞から特徴語を抽出し、全体印象識別器にかける。得られた印象語の単語の関連性を考慮し、歌詞から受ける印象を推定する。
 例えば、「あのこ」、「あにき」、「おやじ」は男性目線の楽曲に頻出であることから、WHOについて「男性」と判断することができる。また、WHENについて、「春」、「冬」、「雪解け」という単語が歌詞にあれば、雪の残る時期と判定できる。WHYについて「おふくろ」「もう5年」のような表現は、「望郷」をテーマにしていると推定できる。
 図8は、事前に設定するカテゴリと、カテゴリ毎の状態を表わす単語の例を示すテーブルである。例えば、WHOの項目に対して、「男」識別器の結果が正、「サラリーマン」識別器の結果が正であった場合、状況情報のWHO情報として、「男」「サラリーマン」が利用される。識別器は、SVMを用いて構成可能である(非特許文献2参照)。楽曲状況情報の推定処理は、歌詞全体、段落毎、歌詞行毎等、分割して適用することができる。推定結果は、必ずしもすべてをマッチングに利用する必要はない。一方、各々の組合せを利用することもできる。
 (映像特定)
 図9は、映像特定の処理の一例を示す模式図である。図9に示すように、5W1Hによるメインキーワードとして「男性」、「春」、「島」、「国」、「君」、「演歌」、「ふるさと」で映像を検索すると、映像1~3が候補として検索される。このうち、最もキーワードと映像のメタ情報とがマッチするのは映像1であり、映像1では関連メタ情報が多く、矛盾するメタ情報が無い。一方、映像2では、関連メタ情報として「北国」、「男」、「冬」があるものの、矛盾するメタ情報として「漁師」、「海」があり、映像3では、「冬」の関連メタ情報に対して「男女」が矛盾するメタ情報となっている。
 (シナリオフォーマット)
 図10は、シナリオフォーマット例を示すテーブルである。シナリオ生成部150は、例えば図10のようなフォーマットでシナリオ情報を生成できる。図10のテーブル210は、楽曲IDに対して曲名とメタ情報とを関連付けたテーブルである。図10のテーブル220において、区分Pp~Peは、楽曲の区分を示している。このシナリオフォーマットは、それぞれの区分Pp~Peに対して開始時刻、前奏等または歌詞、主軸映像の映像ID、調整用映像の映像IDでシナリオの内容を特定している。
 (総括)
 なお、上記のシナリオ生成システムは、処理能力や処理速度を考慮すると、再生部180以外の構成をサーバとし、再生部180のみをクライアントとしたクライアント・サーバ型のシステムで構成することが好ましい。その他、データベースのみをサーバとし、その他の構成をクライアントとしてストリーミングを活用する形態であってもよいし、すべての構成をクライアントとする形態であってもよいし、サーバ、クライアントの切り分けは限定されない。また、上記の説明において、映像は基本的に動画を指すが、静止画であってもよい。特に、歌詞の1行程度の時間に再生する映像としては、静止画であっても違和感を生じさせないので使いやすい。
100 シナリオ生成システム
110 映像DB
120 楽曲DB
130 状況推定部
140 映像特定部
150 シナリオ生成部
160 調整用映像特定部
170 シナリオ調整部
180 再生部
190 評価情報収集部
195 評価情報DB
210、220 テーブル
 
 

Claims (10)

  1.  楽曲再生と同期した映像再生に用いられるシナリオ生成システムであって、
     楽曲が表わす状況を推定する状況推定部と、
     時系列の順序を有するシーンにより構成される映像のうち、前記推定された状況に合った映像を少なくとも一つ特定する映像特定部と、
     前記特定された映像を構成するシーンを前記楽曲の区分ごとに関連付けたシナリオを生成するシナリオ生成部と、を備えることを特徴とするシナリオ生成システム。
  2.  前記シナリオ生成部は、前記特定された映像のシーンを時系列の順序を保持しつつ前記楽曲の区分ごとに関連付けたシナリオを生成することを特徴とする請求項1記載のシナリオ生成システム。
  3.  前記状況推定部は、楽曲の歌詞を解析して前記楽曲が表わす状況を推定することを特徴とする請求項1または請求項2記載のシナリオ生成システム。
  4.  前記楽曲の区分ごとに前記楽曲の歌詞からキーワードを抽出し、前記抽出されたキーワードを用いて調整用の映像を特定する調整用映像特定部と、
     前記調整用の映像を用いて、前記シナリオで特定されるシーンの時系列の順序を保持しつつ、前記シナリオを調整するシナリオ調整部と、を更に備えることを特徴とする請求項1から請求項3のいずれかに記載のシナリオ生成システム。
  5.  前記状況推定部は、楽曲が表わす状況に相当する各カテゴリの内容を推定することで前記楽曲が表わす状況を推定することを特徴とする請求項1から請求項4のいずれかに記載のシナリオ生成システム。
  6.  前記映像調整部は、前記楽曲の区分ごとに、前記シナリオで特定されるシーンと前記調整用の映像とをあらかじめ設定した基準に基づいて比較し、どちらか一方を、新たに前記シナリオで特定されるシーンとして採用することを特徴とする請求項1から請求項5のいずれかに記載のシナリオ生成システム。
  7.  前記生成されたシナリオによる同期映像の視聴者から前記同期映像に対する評価を収集し、前記収集された評価を前記評価情報DBに格納させる評価情報収集部と、
     前記収集された評価を格納する評価情報DBと、を更に備え、
     前記映像調整部は、前記評価情報DBに格納された評価情報に基づいて、前記シナリオで特定されるシーンの組合せを修正することを特徴とする請求項1から請求項6のいずれかに記載のシナリオ生成システム。
  8.  前記生成されたシナリオに基づいて同期映像を再生する再生部を備えることを特徴とする請求項1から請求項7のいずれかに記載のシナリオ生成システム。
  9.  楽曲再生と同期した映像再生に用いられるシナリオを生成するシナリオ生成方法であって、
     楽曲が表わす状況を推定するステップと、
     時系列の順序を有するシーンにより構成される映像のうち、前記推定された状況に合った映像を少なくとも一つ特定するステップと、
     前記特定された映像を構成するシーンを前記楽曲の区分ごとに関連付けたシナリオを生成するステップと、を含むことを特徴とするシナリオ生成方法。
  10.  楽曲再生と同期した映像再生に用いられるシナリオを生成するシナリオ生成プログラムであって、
     楽曲が表わす状況を推定する処理と、
     時系列の順序を有するシーンにより構成される映像のうち、前記推定された状況に合った映像を少なくとも一つ特定する処理と、
     前記特定された映像を構成するシーンを前記楽曲の区分ごとに関連付けたシナリオを生成する処理と、をコンピュータに実行させることを特徴とするシナリオ生成プログラム。
     
     
PCT/JP2014/066800 2013-06-26 2014-06-25 シナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラム Ceased WO2014208581A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US14/901,348 US10104356B2 (en) 2013-06-26 2014-06-25 Scenario generation system, scenario generation method and scenario generation program

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2013-133980 2013-06-26
JP2013133980A JP6159989B2 (ja) 2013-06-26 2013-06-26 シナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラム

Publications (1)

Publication Number Publication Date
WO2014208581A1 true WO2014208581A1 (ja) 2014-12-31

Family

ID=52141913

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2014/066800 Ceased WO2014208581A1 (ja) 2013-06-26 2014-06-25 シナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラム

Country Status (3)

Country Link
US (1) US10104356B2 (ja)
JP (1) JP6159989B2 (ja)
WO (1) WO2014208581A1 (ja)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019003270A (ja) * 2017-06-12 2019-01-10 日本電信電話株式会社 学習装置、映像検索装置、方法、及びプログラム

Families Citing this family (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP6481354B2 (ja) * 2014-12-10 2019-03-13 セイコーエプソン株式会社 情報処理装置、装置を制御する方法、コンピュータープログラム
JP6721570B2 (ja) * 2015-03-12 2020-07-15 株式会社Cotodama 楽曲再生システム、データ出力装置、及び楽曲再生方法
CN105159639B (zh) 2015-08-21 2018-07-27 小米科技有限责任公司 音频封面显示方法及装置
CN110807126B (zh) * 2018-08-01 2023-05-26 腾讯科技(深圳)有限公司 文章转换成视频的方法、装置、存储介质及设备
CN109802987B (zh) * 2018-09-11 2021-05-18 北京京东方技术开发有限公司 用于显示装置的内容推送方法、推送装置和显示设备
CN111935537A (zh) * 2020-06-30 2020-11-13 百度在线网络技术(北京)有限公司 音乐短片视频生成方法、装置、电子设备和存储介质
CN114117086A (zh) 2020-08-31 2022-03-01 脸萌有限公司 多媒体作品的制作方法、装置及计算机可读存储介质

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2003295870A (ja) * 2002-03-29 2003-10-15 Daiichikosho Co Ltd カラオケ装置における背景映像選出システム
JP2004301885A (ja) * 2003-03-28 2004-10-28 Daiichikosho Co Ltd カラオケ映像再生装置における映像更新方法
JP2012014595A (ja) * 2010-07-02 2012-01-19 Kddi Corp 楽曲選択装置、楽曲選択方法および楽曲選択プログラム

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3597160B2 (ja) 2001-09-26 2004-12-02 株式会社第一興商 カラオケ演奏とともに長編ストーリー映像を断続再生する映像音響娯楽システム
JP2004354423A (ja) 2003-05-27 2004-12-16 Xing Inc 音楽再生装置及びその映像表示方法
US20090307207A1 (en) * 2008-06-09 2009-12-10 Murray Thomas J Creation of a multi-media presentation
JP5055223B2 (ja) 2008-08-11 2012-10-24 Kddi株式会社 映像コンテンツ生成装置及びコンピュータプログラム
JP5306114B2 (ja) * 2009-08-28 2013-10-02 Kddi株式会社 クエリ抽出装置、クエリ抽出方法およびクエリ抽出プログラム
JP5431094B2 (ja) 2009-09-29 2014-03-05 株式会社エクシング カラオケ送信システム、カラオケ送信方法、およびコンピュータプログラム
JP2012220582A (ja) * 2011-04-05 2012-11-12 Sony Corp 音楽再生装置、音楽再生方法、プログラム、およびデータ作成装置

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2003295870A (ja) * 2002-03-29 2003-10-15 Daiichikosho Co Ltd カラオケ装置における背景映像選出システム
JP2004301885A (ja) * 2003-03-28 2004-10-28 Daiichikosho Co Ltd カラオケ映像再生装置における映像更新方法
JP2012014595A (ja) * 2010-07-02 2012-01-19 Kddi Corp 楽曲選択装置、楽曲選択方法および楽曲選択プログラム

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
TSUTOMU TERADA: "A System for Generating Background Scenes of Karaoke Using an Active Database System", IPSJ JOURNAL, vol. 44, no. 2, 15 February 2003 (2003-02-15), pages 235 - 244 *

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2019003270A (ja) * 2017-06-12 2019-01-10 日本電信電話株式会社 学習装置、映像検索装置、方法、及びプログラム

Also Published As

Publication number Publication date
JP2015012322A (ja) 2015-01-19
US20160134855A1 (en) 2016-05-12
JP6159989B2 (ja) 2017-07-12
US10104356B2 (en) 2018-10-16

Similar Documents

Publication Publication Date Title
JP6159989B2 (ja) シナリオ生成システム、シナリオ生成方法およびシナリオ生成プログラム
US11669296B2 (en) Computerized systems and methods for hosting and dynamically generating and providing customized media and media experiences
US11328013B2 (en) Generating theme-based videos
US11775580B2 (en) Playlist preview
Berry ‘Just because you play a guitar and are from Nashville doesn’t mean you are a country singer’: the emergence of medium identities in podcasting
KR101326897B1 (ko) 텔레비전 시퀀스를 제공하는 장치 및 방법
US9014832B2 (en) Augmenting media content in a media sharing group
JP2023036777A (ja) メディアコンテンツアイテムと組み合わされたインタースティシャルを含むメディアコンテンツプレイリストの生成
US12368909B2 (en) Image analysis system
US10560657B2 (en) Systems and methods for intelligently synchronizing events in visual content with musical features in audio content
CN114173067B (zh) 一种视频生成方法、装置、设备及存储介质
EP2868112A1 (en) Video remixing system
US11989231B2 (en) Audio recommendation based on text information and video content
CN105224581A (zh) 在播放音乐时呈现图片的方法和装置
Hills et al. Cult cinema and technological change
CN118828139B (zh) 一种ai音乐创作信息处理方法及系统
US20210225408A1 (en) Content Pushing Method for Display Device, Pushing Device and Display Device
JP2011164865A (ja) 画像選定装置、画像選定方法および画像選定プログラム
CN118044206A (zh) 事件源内容和远程内容同步
US20250111675A1 (en) Media trend detection and maintenance at a content sharing platform
TWI780333B (zh) 動態處理並播放多媒體內容的方法及多媒體播放裝置
Singh Bridging data and electronic dance music
CN121743529A (zh) 一种融合深度音频语义与协同过滤的音乐推荐方法及系统
JP2014075662A (ja) スライドショー作成サーバ、ユーザ端末およびスライドショー作成方法
Borsche et al. Location

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 14817193

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

WWE Wipo information: entry into national phase

Ref document number: 14901348

Country of ref document: US

122 Ep: pct application non-entry in european phase

Ref document number: 14817193

Country of ref document: EP

Kind code of ref document: A1