WO2020085675A1 - 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법 - Google Patents
컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법 Download PDFInfo
- Publication number
- WO2020085675A1 WO2020085675A1 PCT/KR2019/013138 KR2019013138W WO2020085675A1 WO 2020085675 A1 WO2020085675 A1 WO 2020085675A1 KR 2019013138 W KR2019013138 W KR 2019013138W WO 2020085675 A1 WO2020085675 A1 WO 2020085675A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- video
- information
- user terminal
- ugc
- server
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/2343—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements
- H04N21/234336—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements by media transcoding, e.g. video is transformed into a slideshow of still pictures or audio is converted into text
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q50/00—Information and communication technology [ICT] specially adapted for implementation of business processes of specific business sectors, e.g. utilities or tourism
- G06Q50/10—Services
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/233—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/25—Management operations performed by the server for facilitating the content distribution or administrating data related to end-users or client devices, e.g. end-user or client device authentication, learning user preferences for recommending movies
- H04N21/266—Channel or content management, e.g. generation and management of keys and entitlement messages in a conditional access system, merging a VOD unicast channel into a multicast channel
- H04N21/26603—Channel or content management, e.g. generation and management of keys and entitlement messages in a conditional access system, merging a VOD unicast channel into a multicast channel for automatically generating descriptors from content, e.g. when it is not made available by its provider, using content analysis techniques
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/27—Server based end-user applications
- H04N21/274—Storing end-user multimedia data in response to end-user request, e.g. network recorder
- H04N21/2743—Video hosting of uploaded data from client
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/41—Structure of client; Structure of client peripherals
- H04N21/422—Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS]
- H04N21/42203—Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS] sound input device, e.g. microphone
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/431—Generation of visual interfaces for content selection or interaction; Content or additional data rendering
- H04N21/4312—Generation of visual interfaces for content selection or interaction; Content or additional data rendering involving specific graphical features, e.g. screen layout, special fonts or colors, blinking icons, highlights or animations
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/435—Processing of additional data, e.g. decrypting of additional data, reconstructing software from modules extracted from the transport stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
- H04N21/4394—Processing of audio elementary streams involving operations for analysing the audio stream, e.g. detecting features or characteristics in audio streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/472—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content
- H04N21/47217—End-user interface for requesting content, additional data or services; End-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification, for manipulating displayed content for controlling playback functions for recorded or on-demand content, e.g. using progress bars, mode or play-point indicators or bookmarks
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/478—Supplemental services, e.g. displaying phone caller identification, shopping application
- H04N21/4788—Supplemental services, e.g. displaying phone caller identification, shopping application communicating with other users, e.g. chatting
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/488—Data services, e.g. news ticker
- H04N21/4884—Data services, e.g. news ticker for displaying subtitles
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/81—Monomedia components thereof
- H04N21/8146—Monomedia components thereof involving graphical data, e.g. 3D object, 2D graphics
- H04N21/8153—Monomedia components thereof involving graphical data, e.g. 3D object, 2D graphics comprising still images, e.g. texture, background image
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/85—Assembly of content; Generation of multimedia applications
- H04N21/854—Content authoring
- H04N21/8549—Creating video summaries, e.g. movie trailer
Definitions
- the present invention relates to a content providing server, a content providing terminal, and a content providing method, and more specifically, to provide a video slide service (Video Slide Service) using information extracted from user generated content (User Generated Content, UGC) It relates to a content providing server, a content providing terminal and a content providing method.
- Video Slide Service Video Slide Service
- UGC User Generated Content
- the most representative method for controlling the playback time of a video is control using a progress bar. This is a method in which when a viewer selects an arbitrary point of the progress bar, the playback point of the video moves to the selected point.
- the progress bar has a constant length regardless of the playback time of the video, when the video playback time is long, the playback time of the video is greatly changed even with a small movement in the progress bar, making it difficult to finely control the playback time. .
- the size of the display is small, and the progress bar needs to be controlled with a finger, so it is more difficult to control the playback time of the video.
- the present invention aims to solve the above and other problems. Another object is to provide a content providing server, a content providing terminal, and a content providing method capable of generating scene meta information for each playback section based on information extracted from user-generated content (UGC).
- ULC user-generated content
- Another object is to provide a content providing server, a content providing terminal, and a content providing method capable of providing a video slide service based on scene meta information for each play section on user-generated content (UGC).
- ULC user-generated content
- the step of uploading a UGC video to the server receiving scene meta information for each play section on the UGC video from the server; Generating a video slide file based on scene meta information for each play section; Displaying an item corresponding to the video slide file; And when selecting the item, displaying a page screen composed of representative image information and subtitle information for each playback section of the UGC video.
- the communication unit for providing a communication interface with the server; A display unit displaying a predefined user interface; And uploading a UGC video to the server using the upload menu included in the user interface, and receiving scene meta information for each play section of the UGC video from the server, based on the scene meta information for each play section.
- Control unit for generating a video slide file, displaying an item corresponding to the video slide file, and displaying a page screen consisting of representative image information and subtitle information for each playback section of the UGC video when the item is selected It provides a user terminal comprising a.
- the process of uploading the UGC video to the server Receiving scene meta information for each play section of the UGC video from the server; Generating a video slide file based on scene meta information for each play section; Displaying an item corresponding to the video slide file; And a computer program recorded on a computer-readable storage medium so that a process of displaying a page screen consisting of representative image information and subtitle information for each playback section of the UGC video is executed on a computer when the item is selected.
- the present invention by generating scene meta information for each playback section using information extracted from a UGC video, and providing a video slide service based on the scene meta information, it enables a viewer (user) to It has the advantage of allowing you to watch UGC videos page by page like a book.
- the user by controlling various operations related to the video slide service through a voice command of the viewer (user), the user can easily use the video slide service even in a situation where the user is difficult to use the hand.
- FIG. 1 is a diagram showing the configuration of a content providing system according to an embodiment of the present invention.
- FIG. 2 is a block diagram showing the configuration of a server according to an embodiment of the present invention.
- FIG. 3 is a block diagram showing the configuration of a user terminal according to an embodiment of the present invention.
- FIG. 4 is a block diagram showing the configuration of a scene meta information generating apparatus according to an embodiment of the present invention.
- FIG. 5 is a view showing the configuration of a scene meta information frame according to an embodiment of the present invention.
- FIG. 6 is a signaling flow diagram of a content providing system according to an embodiment of the present invention.
- FIG. 7 is a flowchart illustrating an operation of a user terminal according to an embodiment of the present invention.
- FIG. 8 is a view referenced to describe an operation of a user terminal displaying a main screen of a video slide application
- 9 and 10 are views referenced to describe the operation of a user terminal that receives scene meta information for each playback section by uploading a UGC video;
- FIG. 11 is a view referred to for describing an operation of a user terminal sharing a video slide file
- 12 to 16 are views for explaining an operation of a user terminal displaying a UGC video on a page-by-page basis.
- 'part' refers to components such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, Includes subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, database, data structures, tables, arrays and variables.
- the functionality provided within components and 'parts' may be combined into a smaller number of components and 'parts' or further separated into additional components and 'parts'.
- the present invention proposes a content providing server, a content providing terminal, and a content providing method capable of generating scene meta information for each playback section based on information extracted from user-generated content (UGC).
- the present invention proposes a content providing server, a content providing terminal, and a content providing method capable of providing a video slide service by utilizing scene meta information for each play section related to user-generated content (UGC).
- the user-generated content (UGC) described herein is a video content produced by a terminal user, and means a moving image composed of one or more video frames and audio frames.
- the user-generated content is not common, but may include subtitle files (or subtitle information).
- the video slide service refers to a video service that allows viewers (users) to easily and quickly grasp the contents of a video by turning the video in pages like a book.
- the scene meta information is information for identifying scenes constituting the video content (that is, a video), and includes at least one of timecode, representative image information, subtitle information, and voice information.
- the time code is information on the subtitle section or the audio section of the video content
- the representative image information is information on the representative image of the subtitle or the audio section
- the voice information is the unit audio information corresponding to the subtitle or the audio section
- the subtitle The information is unit subtitle information corresponding to a subtitle or an audio section.
- the audio section is information on a time section during which a unit voice is output among the playback sections of the video content.
- 'Voice start time information' regarding the playback time of the video content where the output of each unit voice starts, and the output of each unit audio are It may be composed of 'voice end time information' regarding the reproduction time of the video content that is finished, and 'voice output time information' regarding the time that the output of each unit voice is maintained.
- the voice section may consist of only 'voice start time information' and 'voice end time information'.
- the subtitle section is information on a section in which the unit subtitles are displayed among the playback section of the video content, 'subtitle start time information' regarding the playback time of the video content in which the display of each subtitle starts, and the display of each subtitle ends It may be composed of 'subtitle end time information' regarding the playback time of the video content, and 'subtitle display time information' regarding the time that the display of each subtitle is maintained.
- the subtitle section may consist of only 'subtitle start time information' and 'subtitle end time information'.
- FIG. 1 is a diagram showing the configuration of a content providing system according to an embodiment of the present invention.
- the content providing system 10 may include a communication network 100, a server 200 and a user terminal 300.
- the server 200 and the user terminal 300 may be connected to each other through the communication network 100.
- the communication network 100 may include a wired network and a wireless network, specifically, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), and the like. It can include a variety of networks.
- the communication network 100 may include a known World Wide Web (WWW).
- WWW World Wide Web
- the communication network 100 according to the present invention is not limited to the networks listed above, and may include at least one of a known wireless data network, a known telephone network, and a known wired / wireless television network.
- the server 200 is a service providing server or a content providing server, and may perform a function of providing a communication service requested by the user terminal 300.
- the server 200 may configure the content requested by the user terminal 300 in the form of a web page and provide it to the corresponding terminal 300.
- the server 200 may configure the multimedia content requested by the user terminal 300 in the form of a transmission file and provide it to the corresponding terminal 300.
- the server 200 generates scene meta information for each playback section including at least one of time code, representative image information, subtitle information, and voice information based on the video content stored in the database or the video content uploaded from the user terminal 300. And, it is possible to provide scene meta information about the video content to the user terminal (300).
- the reproduction section for generating scene meta information may be a subtitle section or a voice section. Accordingly, the 'scene meta information per reproduction section' may be referred to as 'scene meta information per subtitle section' or 'scene meta information per voice section'.
- the server 200 may also provide a video slide service to the user terminal 300 by utilizing scene meta information regarding video content.
- the server 200 may generate a plurality of page information based on scene meta information (ie, time code, representative image information, subtitle information, and voice information) for each playback section related to video content.
- scene meta information ie, time code, representative image information, subtitle information, and voice information
- the page information is information for providing a video slide service, and is information comprising at least one of representative image information, unit subtitle information, and unit audio information in a page form.
- the server 200 may generate a video slide file including a plurality of page information and provide it to the user terminal 300.
- the user terminal 300 may provide a communication service based on information provided from the server 200.
- the server 200 is a web server
- the user terminal 300 may provide a web service based on the content provided from the server 200.
- the server 200 is a multimedia providing server
- the user terminal 300 may provide a multimedia service based on the content provided from the server 200.
- the user terminal 300 may download and install an application for playing video content and / or providing an additional service related to video content (eg, video slide service).
- the user terminal 300 may access an app store, a play store, etc. to download the corresponding application, or download the corresponding application through a separate storage medium.
- the user terminal 300 may download the corresponding application through wired / wireless communication with the server 200 or another device.
- the user terminal 300 may upload video content (eg, UGC video) to the server 200 according to a user command.
- the user terminal 300 may receive at least one of a video slide file including video content, scene meta information for each playback section on the video content, and page information corresponding to the scene meta information from the server 200.
- the user terminal 300 may generate a plurality of page information based on scene meta information about the image content received from the server 200 or scene meta information about the image content stored in the memory. In addition, the user terminal 300 generates scene meta information for each playback section based on the image content received from the server 200 or the image content stored in the memory, and generates a plurality of page information using the scene meta information. It might be.
- the user terminal 300 may provide a video playback service based on the video content received from the server 200 or stored in the memory. In addition, the user terminal 300 may provide a video slide service based on scene meta information for each playback section related to video content.
- the user terminal 300 described herein includes a mobile phone, a smart phone, a laptop computer, a desktop computer, a terminal for digital broadcasting, personal digital assistants (PDAs), and a portable multimedia player (PMP). ), Slate PC, tablet PC, ultrabook, wearable device, e.g., watch-type terminal (smartwatch), glass-type terminal (smart glass), HMD (head) mounted display)).
- PDAs personal digital assistants
- PMP portable multimedia player
- the user terminal 300 provides a video slide service in conjunction with the server 200, but is not limited thereto, and independently provides video slide services without interworking with the server 200. It will be apparent to those skilled in the art that they can.
- FIG. 2 is a block diagram showing the configuration of a server 200 according to an embodiment of the present invention.
- the server 200 may include a communication unit 210, a database 220, a scene metadata information generation unit 230, a page generation unit 240, and a control unit 250. .
- the components shown in FIG. 2 are not essential for implementing the server 200, so the server described herein may have more or fewer components than those listed above.
- the communication unit 210 may include a wired communication module for supporting wired communication and a wireless communication module for supporting wireless communication.
- the wired communication module is on a wired communication network constructed according to technical standards or communication methods for wired communication (for example, Ethernet, Power Line Communication (PLC), Home PNA (Home PNA), IEEE 1394, etc.). Send and receive wired signals with at least one of other servers, base stations, and access points (APs).
- PLC Power Line Communication
- Home PNA Home PNA
- IEEE 1394 etc.
- the wireless communication module includes technical standards or communication methods for wireless communication (eg, Wireless LAN (WLAN), Wireless-Fidelity (Wi-Fi), Digital Living Network Alliance (DLNA), Global System for Mobile communication (GSM) ), Code Division Multi Access (CDMA), Wideband CDMA (WCDMA), Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), etc. Send and receive radio signals with one.
- wireless communication eg, Wireless LAN (WLAN), Wireless-Fidelity (Wi-Fi), Digital Living Network Alliance (DLNA), Global System for Mobile communication (GSM) ), Code Division Multi Access (CDMA), Wideband CDMA (WCDMA), Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), etc.
- WLAN Wireless LAN
- Wi-Fi Wireless-Fidelity
- DLNA Digital Living Network Alliance
- GSM Global System for Mobile communication
- CDMA Code Division Multi Access
- WCDMA Wideband CDMA
- LTE Long Term Evolution-Advance
- the communication unit 210 is a user terminal (such as a video slide file including video content stored in the database 220, scene meta information for each playback section of the video content, page information corresponding to the scene meta information, etc. 300).
- the communication unit 210 may receive video content uploaded from the user terminal 300, information about a video slide service requested by the user terminal 300, and the like.
- the database 220 includes information (or data) received from the user terminal 300 or another server (not shown), information (or data) generated by the server 200 itself, the user terminal 300 or another server It can perform the function of storing information (or data) to be provided.
- the database 200 may store a plurality of video contents, scene meta information for each playback section for a plurality of video contents, a video slide file including page information corresponding to the scene meta information, and the like.
- the scene meta information generating unit 230 plays at least one of time code, representative image information, subtitle information, and voice information based on the video content stored in the database 220 or the video content uploaded from the user terminal 300. It is possible to generate scene meta information for each section. To this end, the scene meta information generating unit 230 extracts a plurality of audio sections based on audio information extracted from the video content, and recognizes audio information of each audio section by voice recognition, thereby subtracting audio information and subtitles corresponding to each audio section Information can be generated. In addition, the scene meta information generation unit 230 extracts a plurality of voice sections based on audio information extracted from the video content, and each voice section through voice recognition and image recognition for audio information and image information of each voice section It can generate representative image information.
- the representative image information is image information representing subtitles or audio sections of video content, and may include at least one of consecutive video frames of video content played within the subtitles or audio section. More specifically, the representative image information is a video frame randomly selected among video frames in the subtitle or audio section or a video frame selected according to a predetermined rule among the video frames (for example, the most advanced order of the subtitle or audio section). It may be a video frame, a video frame in the middle order, a video frame in the last order, a video frame most similar to subtitle information, and the like.
- the page generator 240 may generate a plurality of page information based on scene meta information for each reproduction section related to the video content. That is, the page generator 240 may generate page information using at least one of representative image information, subtitle information, and voice information. In addition, the page generator 240 may generate a video slide file including a plurality of page information corresponding to scene meta information for each playback section. Meanwhile, when page information corresponding to scene meta information for each playback section is generated by the user terminal 300 instead of the server 200, the corresponding page generation unit 240 may be configured to be omitted.
- the control unit 250 controls the overall operation of the server 200. Furthermore, in order to implement various embodiments described below on the server 200 according to the present invention, the controller 250 may control by combining at least one of the components described above.
- the control unit 250 may provide a communication service requested by the user terminal 300.
- the controller 250 may provide a video playback service or a video slide service to the user terminal 300.
- the controller 250 may provide the video content stored in the database 220 to the user terminal 300.
- the controller 250 may generate scene meta information for each playback section based on information extracted from the video content and provide it to the user terminal 300.
- the controller 250 may generate a video slide file including a plurality of page information based on scene meta information for each playback section related to the video content and provide the video slide file to the user terminal 300.
- FIG. 3 is a block diagram illustrating the configuration of a user terminal 300 according to an embodiment of the present invention.
- the user terminal 300 includes a communication unit 310, an input unit 320, an output unit 330, a memory 340, a voice recognition unit 350, and a control unit 360. It can contain.
- the components shown in FIG. 3 are not essential for implementing a user terminal, so the user terminal described herein may have more or fewer components than those listed above.
- the communication unit 310 may include a wired communication module for supporting a wired network and a wireless communication module for supporting a wireless network.
- the wired communication module is external on the wired communication network built according to technical standards or communication methods for wired communication (for example, Ethernet, Power Line Communication (PLC), Home PNA (Home PNA), IEEE 1394, etc.) Send and receive wired signals with at least one of the server and other terminals.
- wired communication for example, Ethernet, Power Line Communication (PLC), Home PNA (Home PNA), IEEE 1394, etc.
- the wireless communication module includes technical standards or communication methods for wireless communication (eg, Wireless LAN (WLAN), Wireless-Fidelity (Wi-Fi), Digital Living Network Alliance (DLNA), Global System for Mobile communication (GSM) ), Code Division Multi Access (CDMA), Wideband CDMA (WCDMA), Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), etc. Send and receive radio signals with one.
- wireless communication eg, Wireless LAN (WLAN), Wireless-Fidelity (Wi-Fi), Digital Living Network Alliance (DLNA), Global System for Mobile communication (GSM) ), Code Division Multi Access (CDMA), Wideband CDMA (WCDMA), Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), etc.
- WLAN Wireless LAN
- Wi-Fi Wireless-Fidelity
- DLNA Digital Living Network Alliance
- GSM Global System for Mobile communication
- CDMA Code Division Multi Access
- WCDMA Wideband CDMA
- LTE Long Term Evolution-Advance
- the communication unit 310 includes a video slide file including video content from the server 200, scene meta information for each play section of the video content, and a plurality of page information corresponding to the scene meta information for each play section. Can receive.
- the communication unit 310 may transmit the video content uploaded from the user terminal 300, information about a video slide service requested by the user terminal 300, and the like to the server 200.
- the output unit 330 is for generating output related to visual, auditory, or tactile senses, and may include at least one of a display unit, an audio output unit, a hap tip module, and an optical output unit.
- the display unit displays (outputs) information processed by the user terminal 300.
- the display unit displays the execution screen information of the video playback program driven by the user terminal 300, the execution screen information of the video slide program, or the user interface (UI) information according to the execution screen information, a graphical user interface (GUI) Information can be displayed.
- UI user interface
- the display unit may form a mutual layer structure with the touch sensor or may be integrally formed, thereby realizing a touch screen.
- the touch screen may function as a user input unit that provides an input interface between the user terminal 300 and the viewer, and at the same time, provide an output interface between the user terminal 300 and the viewer.
- the audio output unit may output audio data received from the communication unit 310 or stored in the memory 340.
- the sound output unit may output a sound signal related to a video playback service or video slide service provided by the user terminal 300.
- the memory 340 stores data supporting various functions of the user terminal 300.
- the memory 340 may store data and instructions for the operation of the user terminal 300, a video playback program (or application), a video slide program (or application), and a user terminal 300. have.
- the memory 340 may store a plurality of video contents, scene meta information for each playback section related to the plurality of image contents, and a video slide file including a plurality of page information corresponding to the scene meta information.
- the memory 340 is a flash memory type, a hard disk type, a solid state disk type, an SDD type (Silicon Disk Drive type), a multimedia card micro type ), Card-type memory (e.g. SD or XD memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read (EPMROM) It may include a storage medium of at least one type of -only memory), PROM (programmable read-only memory), magnetic memory, magnetic disk, and optical disk.
- the voice recognition unit 350 may classify a voice signal by analyzing characteristics of an acoustic signal input through a microphone, and perform voice recognition on the voice signal to detect textualized voice information. At this time, the speech recognition unit 350 may use a predetermined speech recognition algorithm.
- the voice recognition unit 350 may provide the detected textual voice information to the control unit 360.
- voice information detected through the voice recognition unit 360 may be used as a user's control command.
- the controller 360 controls operations related to a video playback program or video slide program stored in the memory 340, and generally controls the overall operation of the user terminal 300. Furthermore, in order to implement various embodiments described below on the user terminal 300 according to the present invention, the control unit 360 may control by combining at least one of the above-described components.
- control unit 360 may provide a video playback service based on the image content received from the server 200 or stored in the memory 340.
- control unit 360 may provide a video slide service based on scene meta information for each playback section related to image content received from the server 200.
- control unit 360 may provide a video slide service based on a video slide file related to image content received from the server 200.
- control unit 360 directly generates scene meta information for each play section using the video content received from the server 200 or stored in the memory 340, and corresponds to the scene meta information for each play section.
- control unit 360 may directly generate a plurality of page information, and may provide a video slide service based on the plurality of page information.
- FIG. 4 is a block diagram showing the configuration of a scene meta information generating apparatus according to an embodiment of the present invention.
- the scene meta information generating apparatus 400 includes a voice information generating unit 410, a caption information generating unit 420, an image information generating unit 430, and a scene meta information configuring unit 440. ).
- the components shown in FIG. 4 are not essential for implementing the scene meta information generating apparatus 400, so the scene meta information generating apparatus described herein has more or fewer components than those listed above. Can have
- the scene meta information generating apparatus 400 may be implemented through the scene meta information generating unit 230 of the server 200 or may be implemented through the control unit 360 of the user terminal 300, but is not limited thereto. .
- the audio information generating unit 410 may detect a plurality of audio sections based on audio information extracted from the video content, and generate a plurality of audio information corresponding to the detected audio sections.
- the voice information generation unit 410 may generate textualized voice information by performing voice recognition on audio information of each voice section.
- the audio information generating unit 410 includes an audio stream extracting unit 411 for detecting audio information of video content, an audio section analysis unit 413 for detecting audio sections of video content, and audio information of each audio section.
- a voice recognition unit 415 for voice recognition may be included.
- the audio stream extractor 411 may extract an audio stream based on the audio file included in the video content.
- the audio stream extracting unit 411 may divide the audio stream into a plurality of audio frames suitable for signal processing.
- the audio stream may include a voice stream and a non-voice stream.
- the voice section analysis unit 413 analyzes frequency components, pitch components, mel-frequency cepstral coefficients (MFCC) coefficients, and linear predictive coding (LPC) coefficients of each audio frame to extract characteristics of the corresponding audio frame You can.
- the voice section analysis unit 413 may determine whether each audio frame is a voice section using characteristics of each audio frame and a predetermined voice model.
- the voice model includes at least one of a support vector machine (SVM) model, a hidden Markov Model (HMM) model, a Gaussian mixture model (GMM) model, a Recurrent Neural Networks (RNN) model, and a Long Short-Term Memory (LSTM) model.
- SVM support vector machine
- HMM hidden Markov Model
- GMM Gaussian mixture model
- RNN Recurrent Neural Networks
- LSTM Long Short-Term Memory
- the voice section analyzer 413 may combine audio frames corresponding to the voice section to detect the start and end times of each voice section.
- the start time of each audio section corresponds to the playback time of the video content in which the audio output starts in the corresponding section
- the end time of each audio section corresponds to the playback time of the video content in which the audio output ends in the corresponding section.
- the audio section analysis unit 413 may provide information on the audio section of the video content to the caption information generation unit 420 and / or the image information generation unit 430.
- the voice recognition unit 415 analyzes frequency components, pitch components, energy components, zero crossing components, MFCC coefficients, LPC coefficients, and Perceptual Linear Predictive (PLP) coefficients of voice information corresponding to each voice section. Feature vectors of speech information can be detected.
- the speech recognition unit 415 may classify the patterns of the detected feature vectors using a predetermined acoustic model, and recognize one or more candidate words by recognizing speech through the pattern classification.
- the speech recognition unit 415 may generate textual speech information by composing candidate words into sentences based on a predetermined language model.
- the voice recognition unit 415 may provide textual voice information to the caption information generation unit 420 and / or the image information generation unit 430.
- the caption information generation unit 420 may generate a plurality of caption information corresponding to voice sections of the video content based on the textized voice information received from the voice information generation unit 410. That is, if there is no subtitle file in the video content, the subtitle information generation unit 420 may generate new subtitle information by voice recognition of the audio information included in the video content.
- the subtitle information generation unit 420 may detect a plurality of subtitle sections based on the subtitle file, and detect subtitle information corresponding to the subtitle sections. In this case, the caption information generation unit 420 may correct a plurality of caption sections and / or caption information using audio information extracted from video content.
- the image information generation unit 430 detects a video section corresponding to each audio section, and among the plurality of scene images existing in the video section, a scene image most similar to subtitle information or textual audio information (that is, a representative image) You can choose
- the image information generation unit 430 includes a video stream extraction unit 431 for detecting image information constituting image content, a video section detection unit 433 for detecting video sections corresponding to each audio section, and each video section It may include an image tagging unit 435 for generating tag information from the images of and a scene selection unit 437 for selecting a representative image from the images of each video section.
- the video stream extractor 431 may extract a video stream based on a video file included in video content.
- the video stream may be composed of consecutive video frames.
- the video section extraction unit 433 may detect a video section corresponding to each audio section in the video stream. This is to reduce the time and cost of image processing by excluding video sections that are relatively less important (ie, video sections corresponding to non-speech sections).
- the image tagging unit 435 may generate image tag information by performing image recognition on each of a plurality of images (ie, image frames) existing in each video section. That is, the image tagging unit 435 may generate image tag information by recognizing objects (eg, people, objects, text, etc.) present in each image frame.
- the image tag information may include information about all objects present in each image frame.
- the scene selector 437 may measure the similarity between the first vector information corresponding to the image tag information and the second vector information corresponding to the textual speech information using a predetermined similarity measurement technique.
- the similarity measurement technique includes cosine similarity measurement technique, Euclidean similarity measurement technique, jacquard coefficient similarity measurement technique, Pearson correlation coefficient similarity measurement technique, and Manhattan distance similarity measurement technique. At least one of them can be used.
- the scene selection unit 437 detects an image corresponding to textual voice information and image tag information having the highest similarity among a plurality of images existing in each video section, and represents the detected image as a representative of the section You can choose as an image.
- the scene selector 437 among a plurality of images present in each video section, detects an image corresponding to the image tag information having the highest similarity with subtitle information, and corresponds to the detected image It can also be selected as a representative image of the section.
- the scene meta information construction unit 440 includes audio section information, unit subtitle information, unit audio information, and representative image information obtained from the audio information generation unit 410, the subtitle information generation unit 420, and the image information generation unit 430. Based on the, it is possible to configure scene meta information for each playback section.
- the scene meta information configuration unit 440 includes an ID field 510, a time code field 520, a representative image field 530, a voice field 540, and a subtitle field 550. ) And an image tag field 560 may be generated. At this time, the scene metadata information construction unit 440 may generate scene metadata frames as many as the number of subtitles or voice sections.
- the ID field 510 is a field for identifying scene meta information for each reproduction section
- the time code field 520 is a field representing a subtitle or audio section corresponding to the scene meta information. More preferably, the time code field 520 is a field indicating a voice section corresponding to scene meta information.
- the representative image field 530 is a field representing a representative image for each voice section
- the voice field 540 is a field representing voice information for each voice section.
- the subtitle field 550 is a field representing subtitle information for each voice section
- the image tag field 860 is a field representing image tag information for each voice section.
- the scene meta information construction unit 440 may merge the scene meta information into one scene meta information when the representative images of the scene meta information corresponding to the reproduction sections adjacent to each other are similar. At this time, the scene meta information constructing unit 440 may determine whether similarity between representative images is performed using a predetermined similarity measurement algorithm (eg, a cosine similarity measurement algorithm, an Euclidean similarity measurement algorithm, etc.).
- a predetermined similarity measurement algorithm eg, a cosine similarity measurement algorithm, an Euclidean similarity measurement algorithm, etc.
- the apparatus for generating scene meta information may generate scene meta information for each reproduction section based on information extracted from video content.
- the scene meta information for each playback section may be used to provide a video slide service.
- FIG. 6 is a signaling flowchart of a content providing system according to an embodiment of the present invention.
- the user terminal 300 may execute a video slide application according to a user command or the like (S605).
- the video slide application is an application that provides a user interface for viewing videos by turning them into pages like a book.
- the user terminal 300 may display a predefined user interface (UI) on the display unit.
- UI user interface
- the user terminal 300 may upload the selected UGC video to the server 200 (S610).
- the server 200 may detect audio information from the UGC video uploaded from the user terminal 300, and extract information on voice sections of the corresponding video based on the detected audio information (S615).
- the server 200 may generate first scene meta information including time code for each reproduction section, representative image information, and voice information using the extracted voice section information in operation S620.
- the server 200 may store UGC video uploaded from the user terminal 300 and first scene meta information regarding the UGC video in the database.
- the server 200 may transmit the first scene meta information to the user terminal 300 (S625). At this time, the server 200 may transmit the corresponding data in a streaming method.
- the user terminal 300 may generate a plurality of pages based on the first scene meta information received from the server 200 (S630). At this time, the plurality of pages do not include subtitle information.
- the server 200 may generate textualized voice information by voice recognition of voice information corresponding to each voice section (S635).
- the server 200 may generate subtitle information for each playback section based on textual voice information, and may generate second scene meta information including the subtitle information for each playback section (S640).
- the server 200 may store second scene meta information about the UGC video in the database.
- the server 200 may transmit the second scene meta information to the user terminal 300 (S645). Similarly, the server 200 may transmit the corresponding data in a streaming method. Meanwhile, in the present exemplary embodiment, after the server 200 completes voice recognition for all voice sections, subtitle information for all voice sections is illustrated, but is not limited thereto, and voice recognition of each voice section is performed. It will be apparent to those skilled in the art that subtitle information corresponding to a corresponding voice section can be transmitted each time.
- the user terminal 300 may add subtitle information to each page using the second scene meta information received from the server 200 (S650). That is, the user terminal 300 may generate a plurality of page information based on the first and second scene meta information.
- the user terminal 300 may generate a video slide file including a plurality of page information (S655).
- each page information is information for providing a video slide service, and is information comprising at least one of representative image information, subtitle information, and audio information in a page form.
- each page information may be composed of representative image information and caption information, or may be composed of representative image information, caption information, and voice information.
- the user terminal 300 may store the video slide file in memory.
- the user terminal 300 may provide a video slide service based on the video slide file stored in the memory (S660). Accordingly, the terminal user can watch the UGC video produced by the user in units of pages like a book.
- the server 200 side due to the difference between the time required for voice section analysis and the time required for voice recognition, the server 200 side generates and transmits first and second scene meta information sequentially.
- first and second scene meta information sequentially.
- it is not limited. Accordingly, it will be apparent to those skilled in the art that after both the voice section analysis and voice recognition are completed, one scene meta information can be generated and transmitted to the user terminal.
- the server 200 generates a video slide file including a plurality of page information corresponding to the scene meta information for each playback section, as well as scene meta information for each playback section for a UGC video, to the user terminal ( 200).
- FIG. 7 is a flowchart illustrating an operation of a user terminal according to an embodiment of the present invention.
- the user terminal 300 may execute a video slide application according to a user command (S705).
- the user terminal 300 may display a predefined user interface on the display unit.
- the user interface may be composed of an image list area including thumbnail images corresponding to video slide files and a menu area including operation menus of a video slide application, but is not limited thereto.
- the user terminal 300 may display a selection list screen including thumbnail images corresponding to UGC videos stored in the memory on the display unit.
- the user terminal 300 may upload the selected UGC video to the server 200 (S715).
- the user terminal 300 when selecting the upload menu, enters the video shooting mode may generate a new UGC video in real time. When the video shooting is completed, the user terminal 300 may upload the newly generated UGC video to the server 200.
- the server 200 may generate scene meta information for each playback section based on information included in the UGC video uploaded from the user terminal 300.
- the server 200 may transmit scene meta information for each playback section to the user terminal 300.
- the user terminal 300 may receive scene meta information for each playback section of the uploaded UGC video from the server 200 in operation S720.
- the user terminal 300 may generate a video slide file composed of page information corresponding to scene meta information for each playback section and store it in a memory.
- the user terminal 300 may display a thumbnail image corresponding to the video slide file stored in the memory in the image list area of the user interface.
- the user terminal 300 may display a pop-up window for inquiring whether or not to share the video slide file stored in the memory with others (S725).
- the sharing menu is selected through the pop-up window, the user terminal 300 may share the video slide file to other people.
- the people who share the file may be people who have subscribed to a specific website or people who have installed a video slide application, but are not limited thereto.
- the user terminal 300 may enter a video slide mode and execute a video slide file corresponding to the selected thumbnail image (S740).
- the user terminal 300 may display the first page screen among a plurality of pages constituting the video slide file on the display unit.
- the page screen may be composed of representative image information and subtitle information for each playback section of the UGC video.
- the user terminal 300 may output voice information corresponding to the corresponding page.
- the user terminal 300 may display a next page screen or a previous page screen on the display unit in response to a predetermined gesture input (eg, a directional flicking input).
- a predetermined gesture input eg, a directional flicking input
- the user terminal 300 may switch the video slide mode to the video playback mode and play the UGC video corresponding to the voice section of the current page (S750). .
- the user terminal 300 may activate the microphone to enter the voice recognition mode.
- the user terminal 300 may execute a video slide operation corresponding to the voice command (S760).
- the user terminal 300 may execute a page turning function, a mode switching function (video slide mode ⁇ video mode), a subtitle search function, an automatic playback function, a delete / edit / share function, etc. through voice commands.
- the user terminal 300 may execute a high-speed search function (S770). That is, the user terminal 300 may switch (move) the page screen related to the UGC video at a high speed. Meanwhile, in addition, the user terminal 300 may execute a subtitle search function, a tag search function, and the like.
- the user terminal 300 may repeatedly perform the operations of steps 710 to 770 described above until the video slide application ends.
- it is exemplified to provide a video slide service as a separate independent application, but is not limited thereto, and it will be apparent to those skilled in the art that a video slide service can be provided as an additional function of a general video playback application. .
- FIG. 8 is a view referred to for describing an operation of a user terminal displaying a main screen of a video slide application.
- the user terminal 300 may display the home screen 710 on the display unit according to a user command or the like.
- the home screen 810 includes an app icon 815 corresponding to a video slide application.
- the user terminal 300 may execute a video slide application corresponding to the selected app icon 815.
- the user terminal 300 may display a predefined user interface 820 on the display unit.
- the user interface 820 may include an image list area 820 including thumbnail images corresponding to video slide files, and a menu area 830 displayed on the top of the image list area 820.
- the menu area 830 may include a shared list menu 831 and a my list menu 832. Title information of a corresponding video slide file may be displayed at the bottom of each thumbnail image.
- the user terminal 300 may display thumbnail images corresponding to the video slide files in the sharing state in the image list area 820. Meanwhile, when the user list 300 selects the My List menu 832, the thumbnail images corresponding to the video slide files stored in the memory may be displayed in the image list area 820.
- 9 and 10 are diagrams referenced to describe the operation of a user terminal that uploads UGC video and receives scene meta information for each playback section.
- the user terminal 300 may display a predefined user interface 910 on the display unit.
- the user terminal 300 may display a pop-up window 920 for selecting a method of uploading a UGC video on the display unit.
- the pop-up window 920 may include an album menu 921 and a camera menu 922.
- the user terminal 300 may display a selection list screen (not shown) including thumbnail images corresponding to UGC videos stored in the memory on the display unit.
- a selection list screen including thumbnail images corresponding to UGC videos stored in the memory on the display unit.
- the user terminal 300 may upload a UGC video corresponding to the selected thumbnail image to the server 200.
- the user terminal 300 may enter a video shooting mode and generate a new UGC video.
- the user terminal 300 may upload the newly generated UGC video to the server 200.
- the user terminal 300 When the UGC video 930 selected through the pop-up window 920 is uploaded to the server 200, the user terminal 300 has notification information (940, 950, which explains the process of converting the UGC video to a video slide file). 960) may be displayed on the display unit. For example, as illustrated in FIG. 10, the user terminal 300 may sequentially display notification messages such as “uploading”, “extracting a voice section”, and “recognizing voice”. When the conversion process is completed, the user terminal 300 may receive scene meta information for each playback section of the uploaded UGC video from the server 200.
- notification information 940, 950, which explains the process of converting the UGC video to a video slide file.
- 960 may be displayed on the display unit. For example, as illustrated in FIG. 10, the user terminal 300 may sequentially display notification messages such as “uploading”, “extracting a voice section”, and “recognizing voice”.
- the user terminal 300 may receive scene meta information for each playback section of the uploaded UGC video from the server 200.
- FIG. 11 is a view referred to for describing an operation of a user terminal sharing a video slide file.
- the user terminal 300 may receive scene meta information for each playback section of the UGC video from the server 200 in real time.
- the user terminal 300 may generate a video slide file including a plurality of pages corresponding to scene meta information for each playback section.
- the user terminal 300 may display a pop-up window 1110 for inquiring whether to share the video slide file to other people on the display unit.
- the user terminal 300 may share the video slide file to other people.
- the user terminal 300 may display the thumbnail image 1125 corresponding to the video slide file on the user interface 1120.
- 12 to 16 are views for explaining an operation of a user terminal displaying a UGC video on a page-by-page basis.
- the user terminal 300 may display a user interface including a plurality of thumbnail images on the display unit.
- the plurality of thumbnail images correspond to video slide files related to UGC videos.
- the user terminal 300 may execute (play) a video slide file corresponding to the selected thumbnail image. That is, the user terminal 300 may enter a video slide mode that shows UGC videos in units of pages like a book.
- the user terminal 300 may display the predetermined page screen 1200 on the display unit.
- the page screen 1200 may include an image display area 1210, a subtitle display area 1220, a first menu area 1230, a second menu area 1240, and the like, but is not limited thereto. Does not.
- the image display area 1210 may include a representative image corresponding to the current page.
- the subtitle display area 1220 may include subtitle information corresponding to the current page.
- the first and second menu areas 1230 and 1240 may include a plurality of menus for executing functions related to the video slide mode.
- the first menu area 1230 includes a first operation menu (main menu, 1231) for moving to the main screen, a second operation menu (mode switching menu, 1232) for viewing in a video playback mode, subtitles and / or tags It may include a third action menu (search menu, 1233) to search for, and a fourth action menu (more menu, 1234) to view other menus.
- the second menu area 1240 includes a fifth operation menu (automatic switching menu, 1241) for automatically switching the page screen, a sixth operation menu (preview menu, 1242) for previewing before and after pages of the current page, and voice. And a seventh operation menu (microphone menu, 1243) for activating the recognition mode.
- the user terminal 300 may execute a function of previewing before and after pages of the current page. For example, as illustrated in FIG. 13, the user terminal 300 may display a scroll area 1250 including a plurality of thumbnail images corresponding to pages present before and after the current page on the bottom of the display unit. have.
- the plurality of thumbnail images are images in which representative images corresponding to a plurality of pages are reduced to a predetermined size.
- the plurality of thumbnail images may be sequentially arranged according to the time code of the pages.
- the plurality of thumbnail images may be configured to be scrollable according to a predetermined gesture input.
- the thumbnail image 1251 of the current page may be located at the center of the scroll area 1250. That is, a page currently viewed by a viewer may be located in the central portion of the scroll area 1250. The viewer can go directly to a page corresponding to the thumbnail image by selecting any one of the thumbnail images located in the scroll area 1250.
- the user terminal 300 may move to the initial screen of the corresponding application after exiting the video slide mode.
- the user terminal 300 switches a video slide mode to a video playback mode, and then plays the UGC video corresponding to the voice section of the current page. Can play.
- the user terminal 300 may display a video playback screen 1260 corresponding to the current page on the display unit.
- the user terminal 300 may display a search window (not shown) for searching for subtitles or tags.
- a search window for searching for subtitles or tags.
- the user terminal 300 may search for a page including subtitle or tag information corresponding to the text information and display it on the display unit.
- the user terminal 300 may activate the microphone to enter the voice recognition mode.
- the user terminal 300 may display the notification information 1270 as shown in FIG. 15 on the display unit.
- the voice recognition mode when the user's voice command is input through the microphone, the user terminal 300 may execute a video slide operation corresponding to the voice command. Accordingly, the terminal user can easily use the video slide service even in a situation where it is difficult to use the hand.
- the user terminal 300 may display the next page screen of the current page on the display unit. Conversely, when a flicking input having a second directionality is received through the display unit, the user terminal 300 may display the previous page screen of the current page on the display unit. In this way, the user terminal 300 can easily switch the page screen through a predetermined gesture input.
- the automatic switching menu 1241 of the second menu area 1240 is selected, the user terminal 300 may automatically switch to the next page screen after displaying the current page screen for a certain period of time.
- the user terminal 300 may perform a high-speed search function in response to a predetermined gesture input. For example, as illustrated in FIG. 16, when a long touch input 1280 is received for a certain time on the right area of the page screen 1200, the user terminal 300 displays a page screen related to a UGC video. Can be switched (moved) at high speed. In this case, the user terminal 300 may display information 1290 about the total number of pages and information 1295 about the page numbers to be switched on a display area.
- the user terminal 300 may divide a display area into a predetermined number and display a plurality of pages in the divided area in response to a viewer (user) screen division request.
- the user terminal 300 may play or stop the voice information corresponding to the current page in response to a viewer / user request to play / stop.
- the user terminal 300 may provide a video slide service capable of viewing UGC videos in a page unit like a book in conjunction with the server 200.
- the above-described present invention can be embodied as computer readable codes on a medium on which a program is recorded.
- the computer-readable medium continue to store executable computer program, or may be temporarily stored for a run or download.
- the medium may be one which can be a variety of recording device or storage means of a single or several hardware combined form, is not limited to the medium to be directly connected to any computer system, there distributed to the network. Examples of the medium include magnetic media such as hard disks, floppy disks and magnetic tapes, optical recording media such as CD-ROMs and DVDs, and magneto-optical media such as floptical disks, And program instructions including ROM, RAM, flash memory, and the like.
- the media can be recorded to storage media managed as an example for other media, website, etc. supplied to the app store or other retail distributor of various software applications, and servers. Accordingly, the above detailed description should not be construed as limiting in all respects, but should be considered illustrative. The scope of the present invention should be determined by rational interpretation of the appended claims, and all changes within the equivalent scope of the present invention are included in the scope of the present invention.
Landscapes
- Engineering & Computer Science (AREA)
- Signal Processing (AREA)
- Multimedia (AREA)
- Databases & Information Systems (AREA)
- Business, Economics & Management (AREA)
- Tourism & Hospitality (AREA)
- Human Computer Interaction (AREA)
- Marketing (AREA)
- Economics (AREA)
- Physics & Mathematics (AREA)
- General Business, Economics & Management (AREA)
- General Physics & Mathematics (AREA)
- Primary Health Care (AREA)
- Theoretical Computer Science (AREA)
- Human Resources & Organizations (AREA)
- General Health & Medical Sciences (AREA)
- Strategic Management (AREA)
- Health & Medical Sciences (AREA)
- Computer Security & Cryptography (AREA)
- General Engineering & Computer Science (AREA)
- Computer Graphics (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Television Signal Processing For Recording (AREA)
Abstract
본 발명은 비디오 슬라이드 서비스를 제공하는 사용자 단말의 동작 방법에 관한 것으로, UGC 동영상을 서버로 업로드하는 단계; 상기 서버로부터 상기 UGC 동영상에 관한 재생 구간별 장면메타정보를 수신하는 단계; 상기 재생 구간별 장면메타정보를 기반으로 비디오 슬라이드 파일을 생성하는 단계; 상기 비디오 슬라이드 파일에 대응하는 항목을 표시하는 단계; 및 상기 항목 선택 시, 상기 UGC 동영상의 재생 구간별 대표 이미지 정보와 자막 정보로 구성되는 페이지 화면을 표시하는 단계를 포함한다.
Description
본 발명은 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법에 관한 것으로서, 보다 구체적으로는 사용자 제작 컨텐츠(User Generated Content, UGC)로부터 추출된 정보를 이용하여 비디오 슬라이드 서비스(Video Slide Service)를 제공할 수 있는 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법에 관한 것이다.
정보 통신 기술과 대중 문화의 발달로 인하여 다양한 영상 컨텐츠가 제작되어 세계 전역으로 전파되고 있다. 그러나 영상 컨텐츠는 책과 달리 시청자가 동영상의 진행 정도를 임의로 제어할 수 없어서 동영상에 대한 시청자의 이해 여부와 무관하게 해당 동영상을 감상해야 하는 문제점이 있다. 따라서 이와 같은 문제점을 해결하기 위해 동영상의 재생 시점을 제어하거나 동영상을 탐색하기 위한 다양한 방법이 제시되고 있다.
동영상의 재생 시점을 제어하기 위하여 가장 대표적으로 사용되는 방법으로는, 진행 바(progress bar)를 이용한 제어가 있다. 이는 시청자가 진행 바의 임의 지점을 선택하는 경우, 상기 선택된 지점으로 동영상의 재생 시점이 이동하게 되는 방식이다.
그런데, 이러한 진행 바는 동영상의 재생 시간에 상관없이 일정한 길이를 갖기 때문에, 동영상의 재생 시간이 긴 경우 상기 진행 바에서의 작은 이동만으로도 동영상의 재생 시점이 크게 변경되어 재생 시점의 미세한 제어가 어려워진다. 특히 모바일 환경에서 동영상을 감상하는 경우, 디스플레이의 크기가 작고, 손가락으로 진행 바를 제어해야 하는 경우가 많아 동영상의 재생 시점을 제어하는 것이 더욱 어려워지는 문제가 있다.
또한, 사용자의 통신 속도가 제한되는 환경에서 동영상의 내용을 파악하고자 할 때, 동영상이 대용량이거나 고화질인 경우 서버로부터 컨텐츠 제공 단말에 동영상이 원활히 제공될 수 없어 동영상의 모든 장면을 실시간으로 감상하는 것이 어려운 문제가 있다. 따라서, 시청자가 영상 컨텐츠를 책처럼 페이지 단위로 넘겨볼 수 있는 새로운 동영상 서비스가 필요하다.
본 발명은 전술한 문제 및 다른 문제를 해결하는 것을 목적으로 한다. 또 다른 목적은 사용자 제작 컨텐츠(UGC)로부터 추출된 정보를 기반으로 재생 구간별 장면메타정보를 생성할 수 있는 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법을 제공함에 있다.
또 다른 목적은 사용자 제작 컨텐츠(UGC)에 관한 재생 구간별 장면메타정보를 기반으로 비디오 슬라이드 서비스를 제공할 수 있는 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법을 제공함에 있다.
상기 또는 다른 목적을 달성하기 위해 본 발명의 일 측면에 따르면, UGC 동영상을 서버로 업로드하는 단계; 상기 서버로부터 상기 UGC 동영상에 관한 재생 구간별 장면메타정보를 수신하는 단계; 상기 재생 구간별 장면메타정보를 기반으로 비디오 슬라이드 파일을 생성하는 단계; 상기 비디오 슬라이드 파일에 대응하는 항목을 표시하는 단계; 및 상기 항목 선택 시, 상기 UGC 동영상의 재생 구간별 대표 이미지 정보와 자막 정보로 구성되는 페이지 화면을 표시하는 단계를 포함하는 사용자 단말의 동작 방법을 제공한다.
본 발명의 다른 측면에 따르면, 서버와의 통신 인터페이스를 제공하는 통신부; 미리 정의된 사용자 인터페이스를 표시하는 디스플레이부; 및 상기 사용자 인터페이스에 포함된 업로드 메뉴를 이용하여 UGC 동영상을 상기 서버로 업로드하고, 상기 서버로부터 상기 UGC 동영상에 관한 재생 구간별 장면메타정보를 수신한 경우, 상기 재생 구간별 장면메타정보를 기반으로 비디오 슬라이드 파일을 생성하며, 상기 비디오 슬라이드 파일에 대응하는 항목을 표시하고, 상기 항목 선택 시, 상기 UGC 동영상의 재생 구간별 대표 이미지 정보와 자막 정보로 구성되는 페이지 화면을 상기 디스플레이부에 표시하는 제어부를 포함하는 사용자 단말을 제공한다.
본 발명의 또 다른 측면에 따르면, UGC 동영상을 서버로 업로드하는 과정; 상기 서버로부터 상기 UGC 동영상에 관한 재생 구간별 장면메타정보를 수신하는 과정; 상기 재생 구간별 장면메타정보를 기반으로 비디오 슬라이드 파일을 생성하는 과정; 상기 비디오 슬라이드 파일에 대응하는 항목을 표시하는 과정; 및 상기 항목 선택 시, 상기 UGC 동영상의 재생 구간별 대표 이미지 정보와 자막 정보로 구성되는 페이지 화면을 표시하는 과정이 컴퓨터 상에서 실행되도록 컴퓨터로 판독 가능한 저장매체에 기록된 컴퓨터 프로그램을 제공한다.
본 발명의 실시 예들에 따른 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법의 효과에 대해 설명하면 다음과 같다.
본 발명의 실시 예들 중 적어도 하나에 의하면, UGC 동영상으로부터 추출된 정보를 이용하여 재생 구간별 장면메타정보를 생성하고, 상기 장면메타정보를 기반으로 비디오 슬라이드 서비스를 제공함으로써, 시청자(사용자)로 하여금 UGC 동영상을 책처럼 페이지 단위로 시청할 수 있도록 하는 장점이 있다.
또한, 본 발명의 실시 예들 중 적어도 하나에 의하면, 시청자(사용자)의 음성 명령을 통해 비디오 슬라이드 서비스와 관련된 다양한 동작을 제어함으로써, 사용자가 손을 사용하기 힘든 상황에서도 비디오 슬라이드 서비스를 간편하게 이용할 수 있다는 장점이 있다.
또한, 본 발명의 실시 예들 중 적어도 하나에 의하면, 비디오 슬라이드 페이지의 음성 구간에 대응하는 UGC 동영상의 재생 구간만을 재생함으로써, 일반적인 동영상 시청 모드에 비해 동영상 시청 시간을 크게 단축할 수 있다는 장점이 있다.
다만, 본 발명의 실시 예들에 따른 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법이 달성할 수 있는 효과는 이상에서 언급한 것들로 제한되지 않으며, 언급하지 않은 또 다른 효과들은 아래의 기재로부터 본 발명이 속하는 기술분야에서 통상의 지식을 가진 자에게 명확하게 이해될 수 있을 것이다.
도 1은 본 발명의 일 실시 예에 따른 컨텐츠 제공 시스템의 구성을 도시하는 도면;
도 2는 본 발명의 일 실시 예에 따른 서버의 구성을 도시하는 블록도;
도 3은 본 발명의 일 실시 예에 따른 사용자 단말의 구성을 도시하는 블록도;
도 4는 본 발명의 일 실시 예에 따른 장면메타정보 생성장치의 구성을 도시하는 블록도;
도 5는 본 발명의 일 실시 예에 따른 장면메타정보 프레임의 구성을 나타내는 도면;
도 6은 본 발명의 일 실시 예에 따른 컨텐츠 제공 시스템의 시그널링 흐름도;
도 7은 본 발명의 일 실시 예에 따른 사용자 단말의 동작을 설명하는 흐름도;
도 8은 비디오 슬라이드 애플리케이션의 메인 화면을 표시하는 사용자 단말의 동작을 설명하기 위해 참조되는 도면;
도 9 및 도 10은 UGC 동영상을 업로드하여 재생 구간별 장면메타정보를 수신하는 사용자 단말의 동작을 설명하기 위해 참조되는 도면;
도 11은 비디오 슬라이드 파일을 공유하는 사용자 단말의 동작을 설명하기 위해 참조되는 도면;
도 12 내지 도 16은 UGC 동영상을 페이지 단위로 표시하는 사용자 단말의 동작을 설명하기 위해 참조되는 도면.
이하, 첨부된 도면을 참조하여 본 명세서에 개시된 실시 예를 상세히 설명하되, 도면 부호에 관계없이 동일하거나 유사한 구성요소는 동일한 참조 번호를 부여하고 이에 대한 중복되는 설명은 생략하기로 한다. 이하의 설명에서 사용되는 구성요소에 대한 접미사 "모듈" 및 "부"는 명세서 작성의 용이함만이 고려되어 부여되거나 혼용되는 것으로서, 그 자체로 서로 구별되는 의미 또는 역할을 갖는 것은 아니다. 즉, 본 발명에서 사용되는 '부'라는 용어는 소프트웨어, FPGA 또는 ASIC과 같은 하드웨어 구성요소를 의미하며, '부'는 어떤 역할들을 수행한다. 그렇지만 '부'는 소프트웨어 또는 하드웨어에 한정되는 의미는 아니다. '부'는 어드레싱할 수 있는 저장 매체에 있도록 구성될 수도 있고 하나 또는 그 이상의 프로세서들을 재생시키도록 구성될 수도 있다. 따라서, 일 예로서 '부'는 소프트웨어 구성요소들, 객체지향 소프트웨어 구성요소들, 클래스 구성요소들 및 태스크 구성요소들과 같은 구성요소들과, 프로세스들, 함수들, 속성들, 프로시저들, 서브루틴들, 프로그램 코드의 세그먼트들, 드라이버들, 펌웨어, 마이크로 코드, 회로, 데이터, 데이터베이스, 데이터 구조들, 테이블들, 어레이들 및 변수들을 포함한다. 구성요소들과 '부'들 안에서 제공되는 기능은 더 작은 수의 구성요소들 및 '부'들로 결합되거나 추가적인 구성요소들과 '부'들로 더 분리될 수 있다.
또한, 본 명세서에 개시된 실시 예를 설명함에 있어서 관련된 공지 기술에 대한 구체적인 설명이 본 명세서에 개시된 실시 예의 요지를 흐릴 수 있다고 판단되는 경우 그 상세한 설명을 생략한다. 또한, 첨부된 도면은 본 명세서에 개시된 실시 예를 쉽게 이해할 수 있도록 하기 위한 것일 뿐, 첨부된 도면에 의해 본 명세서에 개시된 기술적 사상이 제한되지 않으며, 본 발명의 사상 및 기술 범위에 포함되는 모든 변경, 균등물 내지 대체물을 포함하는 것으로 이해되어야 한다.
본 발명은 사용자 제작 컨텐츠(UGC)로부터 추출된 정보를 기반으로 재생 구간별 장면메타정보를 생성할 수 있는 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법을 제안한다. 또한, 본 발명은 사용자 제작 컨텐츠(UGC)에 관한 재생 구간별 장면메타정보를 활용하여 비디오 슬라이드 서비스를 제공할 수 있는 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법을 제안한다.
한편, 본 명세서에서 설명하는 사용자 제작 콘텐츠(UGC)는 단말 사용자에 의해 제작된 영상 컨텐츠로서, 하나 이상의 영상 프레임 및 오디오 프레임으로 구성된 동영상(moving image)을 의미한다. 상기 사용자 제작 컨텐츠에는 일반적이지 않지만 자막 파일(또는 자막 정보)이 포함될 수도 있다.
비디오 슬라이드 서비스(video slide service)는, 시청자(사용자)로 하여금 동영상을 책처럼 페이지 단위로 넘겨서 동영상의 내용을 쉽고 빠르게 파악할 수 있는 비디오 서비스를 의미한다.
장면메타정보는 영상 컨텐츠(즉, 동영상)를 구성하는 장면들(scenes)을 식별하기 위한 정보로서, 타임코드(timecode), 대표 이미지 정보, 자막 정보, 음성 정보 중 적어도 하나를 포함한다. 여기서, 타임코드는 영상 컨텐츠의 자막 구간 또는 음성 구간에 관한 정보이고, 대표 이미지 정보는 자막 또는 음성 구간의 대표 이미지에 관한 정보이며, 음성 정보는 자막 또는 음성 구간에 대응하는 단위 음성 정보이고, 자막 정보는 자막 또는 음성 구간에 대응하는 단위 자막 정보이다.
음성 구간은 영상 컨텐츠의 재생 구간 중 단위 음성이 출력되는 시 구간에 관한 정보로서, 각 단위 음성의 출력이 시작되는 영상 컨텐츠의 재생 시점에 관한 '음성 시작 시간 정보'와, 각 단위 음성의 출력이 종료되는 영상 컨텐츠의 재생 시점에 관한 '음성 종료 시간 정보'와, 각 단위 음성의 출력이 유지되는 시간에 관한 '음성 출력 시간 정보'로 구성될 수 있다. 한편, 다른 실시 예로, 상기 음성 구간은 '음성 시작 시간 정보'와 '음성 종료 시간 정보'만으로 구성될 수도 있다.
자막 구간은 영상 컨텐츠의 재생 구간 중 단위 자막이 표시되는 구간에 관한 정보로서, 각 단위 자막의 표시가 시작되는 영상 컨텐츠의 재생 시점에 관한 '자막 시작 시간 정보'와, 각 단위 자막의 표시가 종료되는 영상 컨텐츠의 재생 시점에 관한 '자막 종료 시간 정보'와, 각 단위 자막의 표시가 유지되는 시간에 관한 '자막 표시 시간 정보'로 구성될 수 있다. 한편, 다른 실시 예로, 상기 자막 구간은 '자막 시작 시간 정보'와 '자막 종료 시간 정보'만으로 구성될 수도 있다.
이하에서는, 본 발명의 다양한 실시 예들에 대하여, 도면을 참조하여 상세히 설명한다.
도 1은 본 발명의 일 실시 예에 따른 컨텐츠 제공 시스템의 구성을 도시하는 도면이다.
도 1을 참조하면, 본 발명에 따른 컨텐츠 제공 시스템(10)은, 통신 네트워크(100), 서버(200) 및 사용자 단말(300) 등을 포함할 수 있다.
서버(200)와 사용자 단말(300)은 통신 네트워크(100)를 통해 서로 연결될 수 있다. 통신 네트워크(100)는 유선 네트워크와 무선 네트워크를 포함할 수 있으며, 구체적으로, 근거리 네트워크(LAN: Local Area Network), 도시권 네트워크(MAN: Metropolitan Area Network), 광역 네트워크(WAN: Wide Area Network) 등 다양한 네트워크를 포함할 수 있다. 또한, 통신 네트워크(100)는 공지의 월드 와이드 웹(WWW: World Wide Web)을 포함할 수도 있다. 그러나, 본 발명에 따른 통신 네트워크(100)는 상기 열거된 네트워크에 국한되지 않고, 공지의 무선 데이터 네트워크, 공지의 전화 네트워크, 공지의 유/무선 텔레비전 네트워크 중 적어도 하나를 포함할 수도 있다.
서버(200)는, 서비스 제공 서버 또는 컨텐츠 제공 서버로서, 사용자 단말(300)에서 요청하는 통신 서비스(communication service)를 제공하는 기능을 수행할 수 있다. 일 예로, 서버(200)가 웹 서버인 경우, 서버(200)는 사용자 단말(300)에서 요청하는 컨텐츠(content)를 웹 페이지 형태로 구성하여 해당 단말(300)로 제공할 수 있다. 한편, 다른 예로, 서버(200)가 멀티미디어 제공 서버인 경우, 서버(200)는 사용자 단말(300)에서 요청하는 멀티미디어 컨텐츠를 전송 파일 형태로 구성하여 해당 단말(300)로 제공할 수 있다.
서버(200)는 데이터베이스에 저장된 영상 컨텐츠 또는 사용자 단말(300)로부터 업로드된 영상 컨텐츠를 기반으로 타임코드, 대표 이미지 정보, 자막 정보 및 음성 정보 중 적어도 하나를 포함하는 재생 구간별 장면메타정보를 생성하고, 영상 컨텐츠에 관한 장면메타정보를 사용자 단말(300)로 제공할 수 있다. 여기서, 장면메타정보를 생성하기 위한 재생 구간은 자막 구간이거나 혹은 음성 구간일 수 있다. 따라서, '재생 구간별 장면메타정보'는 '자막 구간별 장면메타정보' 또는 '음성 구간별 장면메타정보'라 지칭될 수 있다.
서버(200)는 영상 컨텐츠에 관한 장면메타정보를 활용하여 비디오 슬라이드 서비스를 사용자 단말(300)로 제공할 수도 있다. 이를 위해, 서버(200)는 영상 컨텐츠에 관한 재생 구간 별 장면메타정보(즉, 타임코드, 대표 이미지 정보, 자막 정보 및 음성 정보)를 기반으로 복수의 페이지 정보를 생성할 수 있다. 여기서, 페이지 정보(또는 비디오 슬라이드 정보)는 비디오 슬라이드 서비스를 제공하기 위한 정보로서, 대표 이미지 정보, 단위 자막 정보 및 단위 음성 정보 중 적어도 하나를 페이지 형태로 구성한 정보이다. 또한, 서버(200)는 복수의 페이지 정보를 포함하는 비디오 슬라이드 파일을 생성하여 사용자 단말(300)로 제공할 수 있다.
사용자 단말(300)은 서버(200)로부터 제공받은 정보를 기반으로 통신 서비스를 제공할 수 있다. 일 예로, 서버(200)가 웹 서버인 경우, 사용자 단말(300)은 서버(200)로부터 제공받은 컨텐츠를 기반으로 웹 서비스를 제공할 수 있다. 한편, 다른 예로, 서버(200)가 멀티미디어 제공 서버인 경우, 사용자 단말(300)은 서버(200)로부터 제공받은 컨텐츠를 기반으로 멀티미디어 서비스를 제공할 수 있다.
사용자 단말(300)은 영상 컨텐츠의 재생 및/또는 영상 컨텐츠와 관련된 부가 서비스(가령, 비디오 슬라이드 서비스)를 제공하기 위한 애플리케이션을 다운로드하여 설치할 수 있다. 이때, 사용자 단말(300)은 앱 스토어(app store), 플레이 스토어(play store) 등에 접속하여 해당 애플리케이션을 다운로드하거나, 혹은 별도의 저장매체를 통해 해당 애플리케이션을 다운로드할 수 있다. 또한, 사용자 단말(300)은 서버(200) 또는 타 기기와의 유/무선 통신을 통해 해당 애플리케이션을 다운로드할 수도 있다.
사용자 단말(300)은 사용자 명령 등에 따라 영상 컨텐츠(예를 들어, UGC 동영상)를 서버(200)로 업로드할 수 있다. 사용자 단말(300)은 서버(200)로부터 영상 컨텐츠, 영상 컨텐츠에 관한 재생 구간별 장면메타정보 및 장면메타정보에 대응하는 페이지 정보들을 포함하는 비디오 슬라이드 파일 중 적어도 하나를 수신할 수 있다.
사용자 단말(300)은 서버(200)로부터 수신된 영상 컨텐츠에 관한 장면메타정보 또는 메모리에 저장된 영상 컨텐츠에 관한 장면메타정보를 기반으로 복수의 페이지 정보를 생성할 수 있다. 또한, 사용자 단말(300)은 서버(200)로부터 수신된 영상 컨텐츠 또는 메모리에 저장된 영상 컨텐츠를 기반으로 재생 구간별 장면메타정보를 생성하고, 상기 장면메타정보를 이용하여 복수의 페이지 정보를 생성할 수도 있다.
사용자 단말(300)은 서버(200)로부터 수신하거나 혹은 메모리에 저장된 영상 컨텐츠을 기반으로 동영상 재생 서비스를 제공할 수 있다. 또한, 사용자 단말(300)은 영상 컨텐츠에 관한 재생 구간별 장면메타정보를 기반으로 비디오 슬라이드 서비스를 제공할 수 있다.
본 명세서에서 설명되는 사용자 단말(300)에는 휴대폰, 스마트 폰(smart phone), 노트북 컴퓨터(laptop computer), 데스크톱 컴퓨터(desktop computer), 디지털방송용 단말기, PDA(personal digital assistants), PMP(portable multimedia player), 슬레이트 PC(slate PC), 태블릿 PC(tablet PC), 울트라북(ultrabook), 웨어러블 디바이스(wearable device, 예를 들어, 워치형 단말기 (smartwatch), 글래스형 단말기 (smart glass), HMD(head mounted display)) 등이 포함될 수 있다.
한편, 본 실시 예에서는 사용자 단말(300)이 서버(200)와 연동하여 비디오 슬라이드 서비스를 제공하는 것을 예시하고 있으나 반드시 이에 제한되지는 않으며, 서버(200)와 연동 없이 독립적으로 비디오 슬라이드 서비스들을 제공할 수 있음은 당업자에게 자명할 것이다.
도 2는 본 발명의 일 실시 예에 따른 서버(200)의 구성을 도시하는 블록도이다.
도 2를 참조하면, 본 발명에 따른 서버(200)는 통신부(210), 데이터베이스(220), 장면메타정보 생성부(230), 페이지 생성부(240) 및 제어부(250)를 포함할 수 있다. 도 2에 도시된 구성요소들은 서버(200)를 구현하는데 있어서 필수적인 것은 아니어서, 본 명세서상에서 설명되는 서버는 위에서 열거된 구성요소들보다 많거나, 또는 적은 구성요소들을 가질 수 있다.
통신부(210)는 유선 통신을 지원하기 위한 유선 통신 모듈과 무선 통신을 지원하기 위한 무선 통신 모듈을 포함할 수 있다. 유선 통신 모듈은, 유선 통신을 위한 기술표준들 또는 통신방식(예를 들어, 이더넷(Ethernet), PLC(Power Line Communication), 홈 PNA(Home PNA), IEEE 1394 등)에 따라 구축된 유선 통신망 상에서 타 서버, 기지국, AP(access point) 중 적어도 하나와 유선 신호를 송수신한다. 무선 통신 모듈은, 무선 통신을 위한 기술표준들 또는 통신방식(예를 들어, WLAN(Wireless LAN), Wi-Fi(Wireless-Fidelity), DLNA(Digital Living Network Alliance), GSM(Global System for Mobile communication), CDMA(Code Division Multi Access), WCDMA(Wideband CDMA), LTE(Long Term Evolution), LTE-A(Long Term Evolution-Advanced) 등)에 따라 구축된 무선 통신망 상에서 기지국, Access Point 및 중계기 중 적어도 하나와 무선 신호를 송수신한다.
본 실시 예에서, 통신부(210)는 데이터베이스(220)에 저장된 영상 컨텐츠, 영상 컨텐츠에 관한 재생 구간별 장면메타정보, 상기 장면메타정보에 대응하는 페이지 정보들을 포함하는 비디오 슬라이드 파일 등을 사용자 단말(300)로 전송할 수 있다. 또한, 통신부(210)는 사용자 단말(300)에서 업로드하는 영상 컨텐츠, 사용자 단말(300)에서 요청하는 비디오 슬라이드 서비스에 관한 정보 등을 수신할 수 있다.
데이터베이스(220)는 사용자 단말(300) 또는 타 서버(미도시)로부터 수신하는 정보(또는 데이터), 서버(200)에 의해 자체적으로 생성되는 정보(또는 데이터), 사용자 단말(300) 또는 타 서버로 제공할 정보(또는 데이터) 등을 저장하는 기능을 수행할 수 있다. 본 실시 예에서, 데이터베이스(200)는 복수의 영상 컨텐츠, 복수의 영상 컨텐츠에 관한 재생 구간별 장면메타정보, 상기 장면메타정보에 대응하는 페이지 정보들을 포함하는 비디오 슬라이드 파일 등을 저장할 수 있다.
장면메타정보 생성부(230)는 데이터베이스(220)에 저장된 영상 컨텐츠 또는 사용자 단말(300)로부터 업로드된 영상 컨텐츠를 기반으로 타임코드, 대표 이미지 정보, 자막 정보 및 음성 정보 중 적어도 하나를 포함하는 재생 구간별 장면메타정보를 생성할 수 있다. 이를 위해, 장면메타정보 생성부(230)는 영상 컨텐츠로부터 추출된 오디오 정보를 기반으로 복수의 음성 구간을 추출하고, 각 음성 구간의 오디오 정보를 음성 인식하여 각 음성 구간에 대응하는 음성 정보 및 자막 정보를 생성할 수 있다. 또한, 장면메타정보 생성부(230)는 영상 컨텐츠로부터 추출된 오디오 정보를 기반으로 복수의 음성 구간을 추출하고, 각 음성 구간의 오디오 정보와 이미지 정보에 대한 음성 인식 및 영상 인식을 통해 각 음성 구간의 대표 이미지 정보를 생성할 수 있다.
대표 이미지 정보는 영상 컨텐츠의 자막 또는 음성 구간들을 대표하는 이미지 정보로서, 자막 또는 음성 구간 내에서 재생되는 영상 컨텐츠의 연속된 영상 프레임들 중 적어도 하나를 포함할 수 있다. 좀 더 구체적으로, 대표 이미지 정보는 자막 또는 음성 구간 내의 영상 프레임들 중에서 임의로 선택된 영상 프레임이거나 혹은 상기 영상 프레임들 중에서 미리 결정된 규칙에 따라 선택된 영상 프레임(예를 들면, 자막 또는 음성 구간 중 가장 앞선 순서의 영상 프레임, 중간 순서의 영상 프레임, 마지막 순서의 영상 프레임, 자막 정보와 가장 유사한 영상 프레임 등)일 수 있다.
페이지 생성부(240)는 영상 컨텐츠에 관한 재생 구간별 장면메타정보를 기반으로 복수의 페이지 정보를 생성할 수 있다. 즉, 페이지 생성부(240)는 대표 이미지 정보, 자막 정보 및 음성 정보 중 적어도 하나를 이용하여 페이지 정보를 생성할 수 있다. 또한, 페이지 생성부(240)는 재생 구간별 장면메타정보에 대응하는 복수의 페이지 정보들을 포함하는 비디오 슬라이드 파일을 생성할 수 있다. 한편, 서버(200) 대신 사용자 단말(300)에서 재생 구간별 장면메타정보에 대응하는 페이지 정보를 생성하는 경우, 해당 페이지 생성부(240)는 생략 가능하도록 구성될 수 있다.
제어부(250)는 서버(200)의 전반적인 동작을 제어한다. 나아가 제어부(250)는 이하에서 설명되는 다양한 실시 예들을 본 발명에 따른 서버(200) 상에서 구현하기 위하여, 위에서 살펴본 구성요소들을 중 적어도 하나를 조합하여 제어할 수 있다.
본 실시 예에서, 제어부(250)는 사용자 단말(300)에서 요청하는 통신 서비스를 제공할 수 있다. 일 예로, 제어부(250)는 동영상 재생 서비스 또는 비디오 슬라이드 서비스 등을 사용자 단말(300)로 제공할 수 있다. 이를 위해, 제어부(250)는 데이터베이스(220)에 저장된 영상 컨텐츠를 사용자 단말(300)로 제공할 수 있다. 또한, 제어부(250)는 영상 컨텐츠로부터 추출된 정보를 기반으로 재생 구간별 장면메타정보를 생성하여 사용자 단말(300)로 제공할 수 있다. 또한, 제어부(250)는 영상 컨텐츠에 관한 재생 구간별 장면메타정보를 기반으로 복수의 페이지 정보를 포함하는 비디오 슬라이드 파일을 생성하여 사용자 단말(300)로 제공할 수도 있다.
도 3은 본 발명의 일 실시 예에 따른 사용자 단말(300)의 구성을 설명하기 위한 블록도이다.
도 3을 참조하면, 본 발명에 따른 사용자 단말(300)은 통신부(310), 입력부(320), 출력부(330), 메모리(340), 음성 인식부(350) 및 제어부(360) 등을 포함할 수 있다. 도 3에 도시된 구성요소들은 사용자 단말을 구현하는데 있어서 필수적인 것은 아니어서, 본 명세서상에서 설명되는 사용자 단말은 위에서 열거된 구성요소들보다 많거나 또는 적은 구성요소들을 가질 수 있다.
통신부(310)는 유선 네트워크를 지원하기 위한 유선 통신 모듈과, 무선 네트워크를 지원하기 위한 무선 통신 모듈을 포함할 수 있다. 유선 통신 모듈은 유선 통신을 위한 기술표준들 또는 통신방식(예를 들어, 이더넷(Ethernet), PLC(Power Line Communication), 홈 PNA(Home PNA), IEEE 1394 등)에 따라 구축된 유선 통신망 상에서 외부 서버 및 타 단말 중 적어도 하나와 유선 신호를 송수신한다. 무선 통신 모듈은, 무선 통신을 위한 기술표준들 또는 통신방식(예를 들어, WLAN(Wireless LAN), Wi-Fi(Wireless-Fidelity), DLNA(Digital Living Network Alliance), GSM(Global System for Mobile communication), CDMA(Code Division Multi Access), WCDMA(Wideband CDMA), LTE(Long Term Evolution), LTE-A(Long Term Evolution-Advanced) 등)에 따라 구축된 무선 통신망 상에서 기지국, Access Point 및 중계기 중 적어도 하나와 무선 신호를 송수신한다.
본 실시 예에서, 통신부(310)는 서버(200)로부터 영상 컨텐츠, 영상 컨텐츠에 관한 재생 구간별 장면메타정보, 상기 재생 구간별 장면 메타정보에 대응하는 복수의 페이지 정보를 포함하는 비디오 슬라이드 파일 등을 수신할 수 있다. 또한, 통신부(310)는 사용자 단말(300)에서 업로드하는 영상 컨텐츠, 사용자 단말(300)에서 요청하는 비디오 슬라이드 서비스에 관한 정보 등을 서버(200)로 전송할 수 있다.
입력부(320)는 영상 신호 입력을 위한 카메라, 오디오 신호 입력을 위한 마이크로폰(microphone), 사용자로부터 정보를 입력받기 위한 사용자 입력부(예를 들어, 키보드, 마우스, 터치키(touch key), 푸시키(mechanical key) 등) 등을 포함할 수 있다. 상기 입력부(320)에서 획득한 데이터는 분석되어 단말 사용자의 제어 명령으로 처리될 수 있다. 본 실시 예에서, 입력부(320)는 동영상 재생 서비스 및 비디오 슬라이드 서비스 등과 관련된 명령 신호들을 수신할 수 있다.
출력부(330)는 시각, 청각 또는 촉각 등과 관련된 출력을 발생시키기 위한 것으로, 디스플레이부, 음향 출력부, 햅팁 모듈 및 광 출력부 중 적어도 하나를 포함할 수 있다.
디스플레이부는 사용자 단말(300)에서 처리되는 정보를 표시(출력)한다. 본 실시 예에서, 디스플레이부는 사용자 단말(300)에서 구동되는 동영상 재생 프로그램의 실행화면 정보, 비디오 슬라이드 프로그램의 실행화면 정보 또는 이러한 실행화면 정보에 따른 UI(User Interface) 정보, GUI(Graphic User Interface) 정보를 표시할 수 있다.
디스플레이부는 터치 센서와 상호 레이어 구조를 이루거나 일체형으로 형성됨으로써, 터치 스크린을 구현할 수 있다. 이러한 터치 스크린은, 사용자 단말(300)과 시청자 사이의 입력 인터페이스를 제공하는 사용자 입력부로써 기능함과 동시에, 사용자 단말(300)과 시청자 사이의 출력 인터페이스를 제공할 수 있다.
음향 출력부는 통신부(310)로부터 수신되거나 메모리(340)에 저장된 오디오 데이터를 출력할 수 있다. 본 실시 예에서, 음향 출력부는 사용자 단말(300)에서 제공되는 동영상 재생 서비스 또는 비디오 슬라이드 서비스와 관련된 음향 신호를 출력할 수 있다.
메모리(340)는 사용자 단말(300)의 다양한 기능을 지원하는 데이터를 저장한다. 본 실시 예에서, 메모리(340)는 사용자 단말(300)에서 구동되는 동영상 재생 프로그램(또는 애플리케이션), 비디오 슬라이드 프로그램(또는 애플리케이션), 사용자 단말(300)의 동작을 위한 데이터들 및 명령어들을 저장할 수 있다. 또한, 메모리(340)는 복수의 영상 컨텐츠, 복수의 영상 컨텐츠에 관한 재생 구간별 장면메타정보, 상기 장면메타정보에 대응하는 복수의 페이지 정보들을 포함하는 비디오 슬라이드 파일 등을 저장할 수 있다.
메모리(340)는 플래시 메모리 타입(flash memory type), 하드디스크 타입(hard disk type), SSD 타입(Solid State Disk type), SDD 타입(Silicon Disk Drive type), 멀티미디어 카드 마이크로 타입(multimedia card micro type), 카드 타입의 메모리(예를 들어 SD 또는 XD 메모리 등), 램(random access memory; RAM), SRAM(static random access memory), 롬(read-only memory; ROM), EEPROM(electrically erasable programmable read-only memory), PROM(programmable read-only memory), 자기 메모리, 자기 디스크 및 광디스크 중 적어도 하나의 타입의 저장매체를 포함할 수 있다.
음성 인식부(350)는 마이크로폰을 통해 입력되는 음향 신호의 특성을 분석하여 음성 신호를 분류하고, 상기 음성 신호에 대해 음성 인식을 수행하여 텍스트화된 음성 정보를 검출할 수 있다. 이때, 상기 음성 인식부(350)는 미리 결정된 음성 인식 알고리즘을 사용할 수 있다.
음성 인식부(350)는 상기 검출된 텍스트화된 음성 정보를 제어부(360)로 제공할 수 있다. 본 실시 예에서, 음성 인식부(360)를 통해 검출된 음성 정보는 사용자의 제어 명령으로 사용될 수 있다.
제어부(360)는 메모리(340)에 저장된 동영상 재생 프로그램 또는 비디오 슬라이드 프로그램 등과 관련된 동작과, 통상적으로 사용자 단말(300)의 전반적인 동작을 제어한다. 나아가 제어부(360)는 이하에서 설명되는 다양한 실시 예들을 본 발명에 따른 사용자 단말(300) 상에서 구현하기 위하여, 위에서 살펴본 구성요소들을 중 적어도 하나를 조합하여 제어할 수 있다.
본 실시 예에서, 제어부(360)는 서버(200)로부터 수신하거나 혹은 메모리(340)에 저장된 영상 컨텐츠를 기반으로 동영상 재생 서비스를 제공할 수 있다. 또한, 제어부(360)는 서버(200)로부터 수신된 영상 컨텐츠에 관한 재생 구간별 장면메타정보를 기반으로 비디오 슬라이드 서비스를 제공할 수 있다. 또한, 제어부(360)는 서버(200)로부터 수신된 영상 컨텐츠에 관한 비디오 슬라이드 파일을 기반으로 비디오 슬라이드 서비스를 제공할 수도 있다.
한편, 다른 실시 예로, 제어부(360)는 서버(200)로부터 수신하거나 혹은 메모리(340)에 저장된 영상 컨텐츠를 이용하여 재생 구간별 장면메타정보를 직접 생성하고, 상기 재생 구간별 장면메타정보에 대응하는 복수의 페이지 정보들을 생성하며, 상기 복수의 페이지 정보들을 기반으로 비디오 슬라이드 서비스를 제공할 수도 있다.
도 4는 본 발명의 일 실시 예에 따른 장면메타정보 생성장치의 구성을 도시하는 블록도이다.
도 4를 참조하면, 본 발명에 따른 장면메타정보 생성장치(400)는 음성정보 생성부(410), 자막정보 생성부(420), 이미지정보 생성부(430) 및 장면메타정보 구성부(440)를 포함할 수 있다. 도 4에 도시된 구성요소들은 장면메타정보 생성장치(400)를 구현하는데 있어서 필수적인 것은 아니어서, 본 명세서상에서 설명되는 장면메타정보 생성장치는 위에서 열거된 구성요소들보다 많거나 또는 적은 구성요소들을 가질 수 있다.
이러한 장면메타정보 생성장치(400)는 서버(200)의 장면메타정보 생성부(230)를 통해 구현되거나 혹은 사용자 단말(300)의 제어부(360)를 통해 구현될 수 있으며 반드시 이에 제한되지는 않는다.
음성정보 생성부(410)는 영상 컨텐츠에서 추출된 오디오 정보를 기반으로 복수의 음성 구간들을 검출하고, 상기 검출된 음성 구간들에 대응하는 복수의 음성 정보들을 생성할 수 있다. 또한, 음성정보 생성부(410)는 각 음성 구간의 오디오 정보에 대해 음성 인식을 수행하여 텍스트화된 음성 정보를 생성할 수 있다.
이러한 음성정보 생성부(410)는 영상 컨텐츠의 오디오 정보를 검출하기 위한 오디오 스트림 추출부(411), 영상 컨텐츠의 음성 구간들을 검출하기 위한 음성 구간 분석부(413) 및 각 음성 구간의 오디오 정보를 음성 인식하기 위한 음성 인식부(415)를 포함할 수 있다.
오디오 스트림 추출부(411)는 영상 컨텐츠에 포함된 오디오 파일을 기반으로 오디오 스트림을 추출할 수 있다. 오디오 스트림 추출부(411)는 오디오 스트림을 신호 처리에 적합한 복수의 오디오 프레임으로 분할할 수 있다. 여기서, 상기 오디오 스트림은 음성 스트림과 비 음성 스트림을 포함할 수 있다.
음성 구간 분석부(413)는 각 오디오 프레임의 주파수 성분, 피치(pitch) 성분, MFCC(mel-frequency cepstral coefficients) 계수, LPC(linear predictive coding) 계수 등을 분석하여 해당 오디오 프레임의 특징들을 추출할 수 있다. 음성 구간 분석부(413)는 각 오디오 프레임의 특징들과 미리 결정된 음성 모델을 이용하여 각각의 오디오 프레임이 음성 구간인지 여부를 결정할 수 있다. 이때, 상기 음성 모델로는 SVM(support vector machine) 모델, HMM(hidden Markov model) 모델, GMM(Gaussian mixture model) 모델, RNN(Recurrent Neural Networks) 모델, LSTM(Long Short-Term Memory) 모델 중 적어도 하나가 사용될 수 있다.
음성 구간 분석부(413)는 음성 구간에 해당하는 오디오 프레임들을 결합하여 각 음성 구간의 시작 시점과 종료 시점을 검출할 수 있다. 여기서, 각 음성 구간의 시작 시점은 해당 구간에서 음성 출력이 시작되는 영상 컨텐츠의 재생 시점에 대응하고, 각 음성 구간의 종료 시점은 해당 구간에서 음성 출력이 종료되는 영상 컨텐츠의 재생 시점에 대응한다. 음성 구간 분석부(413)는 영상 컨텐츠의 음성 구간에 관한 정보를 자막정보 생성부(420) 및/또는 이미지정보 생성부(430)로 제공할 수 있다.
음성 인식부(415)는 각 음성 구간에 대응하는 음성 정보의 주파수 성분, 피치 성분, 에너지 성분, 제로 크로싱(zero crossing) 성분, MFCC 계수, LPC 계수, PLP(Perceptual Linear Predictive) 계수 등을 분석하여 음성 정보의 특징 벡터들을 검출할 수 있다. 음성 인식부(415)는 미리 결정된 음향 모델을 이용하여 상기 검출된 특징 벡터들의 패턴을 분류하고, 상기 패턴 분류를 통해 음성을 인식하여 하나 이상의 후보 단어들을 검출할 수 있다. 그리고, 음성 인식부(415)는 미리 결정된 언어 모델을 기반으로 후보 단어들을 문장으로 구성하여 텍스트화된 음성 정보를 생성할 수 있다. 음성 인식부(415)는 텍스트화된 음성 정보를 자막정보 생성부(420) 및/또는 이미지정보 생성부(430)로 제공할 수 있다.
자막정보 생성부(420)는 음성정보 생성부(410)로부터 수신된 텍스트화된 음성 정보를 기반으로 영상 컨텐츠의 음성 구간들에 대응하는 복수의 자막 정보들을 생성할 수 있다. 즉, 영상 컨텐츠에 자막 파일이 존재하지 않는 경우, 자막정보 생성부(420)는 영상 컨텐츠에 포함된 오디오 정보를 음성 인식하여 새로운 자막 정보를 생성할 수 있다.
한편, 영상 컨텐츠에 자막 파일이 존재하는 경우, 자막정보 생성부(420)는 자막 파일을 기반으로 복수의 자막 구간들을 검출하고, 상기 자막 구간들에 대응하는 자막 정보들을 검출할 수 있다. 이 경우, 자막정보 생성부(420)는 영상 컨텐츠에서 추출된 오디오 정보를 이용하여 복수의 자막 구간 및/또는 자막 정보들을 보정할 수 있다.
이미지정보 생성부(430)는 각 음성 구간에 대응하는 비디오 구간을 검출하고, 상기 비디오 구간에 존재하는 복수의 장면 이미지들 중에서 자막 정보 또는 텍스트된 음성 정보와 가장 유사한 장면 이미지(즉, 대표 이미지)를 선택할 수 있다.
이러한 이미지정보 생성부(430)는 영상 컨텐츠를 구성하는 이미지 정보를 검출하기 위한 비디오 스트림 추출부(431), 각 음성 구간에 대응하는 비디오 구간을 검출하기 위한 비디오 구간 검출부(433), 각 비디오 구간의 이미지들로부터 태그 정보를 생성하는 이미지 태깅부(435) 및 각 비디오 구간의 이미지들 중에서 대표 이미지를 선택하는 장면 선택부(437)를 포함할 수 있다.
비디오 스트림 추출부(431)는 영상 컨텐츠에 포함된 동영상 파일을 기반으로 비디오 스트림을 추출할 수 있다. 여기서, 비디오 스트림은 연속된 영상 프레임들로 구성될 수 있다.
비디오 구간 추출부(433)는 비디오 스트림에서 각 음성 구간에 대응하는 비디오 구간을 검출할 수 있다. 이는 상대적으로 중요도가 낮은 비디오 구간(즉, 비 음성 구간에 대응하는 비디오 구간)을 제외시킴으로써 영상 처리하는데 소요되는 시간과 비용을 줄이기 위함이다.
이미지 태깅부(435)는 각 비디오 구간 내에 존재하는 복수의 이미지들(즉, 영상 프레임들) 각각에 대해 영상 인식을 수행하여 이미지 태그 정보를 생성할 수 있다. 즉, 이미지 태깅부(435)는 각 영상 프레임 내에 존재하는 객체들(가령, 사람, 사물, 텍스트 등)을 인식하여 이미지 태그 정보를 생성할 수 있다. 여기서, 이미지 태그 정보는 각 영상 프레임에 존재하는 모든 객체에 관한 정보를 포함할 수 있다.
장면 선택부(437)는 미리 결정된 유사도 측정 기법을 이용하여 이미지 태그 정보에 대응하는 제1 벡터 정보와 텍스트화된 음성 정보에 대응하는 제2 벡터 정보 간의 유사도를 측정할 수 있다. 상기 유사도 측정 기법으로는 코사인 유사도(cosine similarity) 측정 기법, 유클리디안 유사도(Euclidean similarity) 측정 기법, 자카드 계수를 이용한 유사도 측정 기법, 피어슨 상관계수를 이용한 유사도 측정 기법, 맨하튼 거리를 이용한 유사도 측정 기법 중 적어도 하나가 사용될 수 있다.
장면 선택부(437)는, 각 비디오 구간 내에 존재하는 복수의 이미지들 중에서, 텍스트화된 음성 정보와 유사도가 가장 높은 이미지 태그 정보에 대응하는 이미지를 검출하고, 상기 검출된 이미지를 해당 구간의 대표 이미지로 선택할 수 있다.
한편, 다른 실시 예로, 장면 선택부(437)는, 각 비디오 구간 내에 존재하는 복수의 이미지들 중에서, 자막 정보와 유사도가 가장 높은 이미지 태그 정보에 대응하는 이미지를 검출하고, 상기 검출된 이미지를 해당 구간의 대표 이미지로 선택할 수도 있다.
장면메타정보 구성부(440)는 음성정보 생성부(410), 자막정보 생성부(420) 및 이미지정보 생성부(430)로부터 획득한 음성 구간 정보, 단위 자막 정보, 단위 음성 정보 및 대표 이미지 정보를 기반으로 재생 구간별 장면메타정보를 구성할 수 있다.
일 예로, 도 5에 도시된 바와 같이, 장면메타정보 구성부(440)는 ID 필드(510), 타임코드 필드(520), 대표 이미지 필드(530), 음성 필드(540), 자막 필드(550) 및 이미지 태그 필드(560)를 포함하는 장면메타정보 프레임(500)을 생성할 수 있다. 이때, 장면메타정보 구성부(440)는 자막 또는 음성 구간들의 개수만큼 장면메타정보 프레임들을 생성할 수 있다.
ID 필드(510)는 재생 구간별 장면메타정보를 식별하기 위한 필드이고, 타임코드 필드(520)는 장면메타정보에 해당하는 자막 또는 음성 구간을 나타내는 필드이다. 좀 더 바람직하게, 타임코드 필드(520)는 장면메타정보에 대응하는 음성 구간을 나타내는 필드이다.
대표 이미지 필드(530)는 음성 구간별 대표 이미지를 나타내는 필드이고, 음성 필드(540)는 음성 구간별 음성 정보를 나타내는 필드이다. 그리고, 자막 필드(550)는 음성 구간별 자막 정보를 나타내는 필드이고, 이미지 태그 필드(860)는 음성 구간별 이미지 태그 정보를 나타내는 필드이다.
장면메타정보 구성부(440)는 서로 인접한 재생 구간에 해당하는 장면메타정보들의 대표 이미지가 유사한 경우, 해당 장면메타정보들을 하나의 장면메타정보로 병합할 수 있다. 이때, 상기 장면메타정보 구성부(440)는 미리 결정된 유사도 측정 알고리즘(가령, 코사인 유사도 측정 알고리즘, 유클리안 유사도 측정 알고리즘 등)을 이용하여 대표 이미지 간 유사 여부를 결정할 수 있다.
이상 상술한 바와 같이, 본 발명에 따른 장면메타정보 생성장치는 영상 컨텐츠로부터 추출된 정보를 기반으로 재생 구간별 장면메타정보를 생성할 수 있다. 이러한 재생 구간별 장면메타정보는 비디오 슬라이드 서비스를 제공하기 위해 사용될 수 있다.
도 6은 본 발명의 일 실시 예에 따른 컨텐츠 제공 시스템의 시그널링 흐름도이다.
도 6을 참조하면, 사용자 단말(300)은 사용자 명령 등에 따라 비디오 슬라이드 애플리케이션을 실행할 수 있다(S605). 여기서, 비디오 슬라이드 애플리케이션은, 동영상을 책처럼 페이지 단위로 넘겨서 시청할 수 있는 사용자 인터페이스를 제공하는 애플리케이션이다.
사용자 단말(300)은, 해당 애플리케이션 실행 시, 미리 정의된 사용자 인터페이스(User Interface, UI)를 디스플레이부에 표시할 수 있다. 사용자 인터페이스의 업로드 메뉴를 통해 UGC 동영상이 선택되면, 사용자 단말(300)은 상기 선택된 UGC 동영상을 서버(200)로 업로드할 수 있다(S610).
서버(200)는 사용자 단말(300)로부터 업로드된 UGC 동영상에서 오디오 정보를 검출하고, 상기 검출된 오디오 정보를 기반으로 해당 동영상의 음성 구간들에 관한 정보를 추출할 수 있다(S615).
서버(200)는 상기 추출된 음성 구간 정보를 이용하여 재생 구간별 타임코드, 대표 이미지 정보 및 음성 정보를 포함하는 제1 장면메타정보를 생성할 수 있다(S620). 서버(200)는 사용자 단말(300)로부터 업로드된 UGC 동영상과 상기 UGC 동영상에 관한 제1 장면메타정보를 데이터베이스에 저장할 수 있다.
서버(200)는 제1 장면메타정보를 사용자 단말(300)로 전송할 수 있다(S625). 이때, 서버(200)는 해당 데이터를 스트리밍 방식으로 전송할 수 있다. 사용자 단말(300)은 서버(200)로부터 수신된 제1 장면메타정보를 기반으로 복수의 페이지를 생성할 수 있다(S630). 이때, 상기 복수의 페이지는 자막 정보를 포함하고 있지 않은 상태이다.
서버(200)는 각 음성 구간에 대응하는 음성 정보를 음성 인식하여 텍스트화된 음성 정보를 생성할 수 있다(S635). 서버(200)는 텍스트화된 음성 정보를 기반으로 재생 구간별 자막 정보를 생성하고, 상기 재생 구간별 자막 정보를 포함하는 제2 장면메타정보를 생성할 수 있다(S640). 서버(200)는 UGC 동영상에 관한 제2 장면메타정보를 데이터베이스에 저장할 수 있다.
서버(200)는 제2 장면메타정보를 사용자 단말(300)로 전송할 수 있다(S645). 마찬가지로, 서버(200)는 해당 데이터를 스트리밍 방식으로 전송할 수 있다. 한편, 본 실시 예에서는, 서버(200)가 모든 음성 구간에 대해 음성 인식을 완료한 이후에 모든 음성 구간에 대한 자막 정보를 전송하는 것을 예시하고 있으나 이를 제한하지는 않으며, 각 음성 구간을 음성 인식할 때마다 해당 음성 구간에 대응하는 자막 정보를 전송할 수 있음은 당업자에게 자명할 것이다.
사용자 단말(300)은 서버(200)로부터 수신된 제2 장면메타정보를 이용하여 각각의 페이지에 자막 정보를 추가할 수 있다(S650). 즉, 사용자 단말(300)은 제1 및 제2 장면메타정보를 기반으로 복수의 페이지 정보들을 생성할 수 있다.
사용자 단말(300)은 복수의 페이지 정보들을 포함하는 비디오 슬라이드 파일을 생성할 수 있다(S655). 여기서, 각 페이지 정보는 비디오 슬라이드 서비스를 제공하기 위한 정보로서, 대표 이미지 정보, 자막 정보 및 음성 정보 중 적어도 하나를 페이지 형태로 구성한 정보이다. 예를 들어, 각 페이지 정보는 대표 이미지 정보와 자막 정보로 구성되거나 혹은 대표 이미지 정보, 자막 정보 및 음성 정보로 구성될 수 있다.
사용자 단말(300)은 비디오 슬라이드 파일을 메모리에 저장할 수 있다. 사용자 단말(300)은 메모리에 저장된 비디오 슬라이드 파일을 기반으로 비디오 슬라이드 서비스를 제공할 수 있다(S660). 이에 따라, 단말 사용자는 자신이 제작한 UGC 동영상을 책처럼 페이지 단위로 시청할 수 있게 된다.
한편, 본 실시 예에서는, 음성 구간 분석에 소요되는 시간과 음성 인식에 소요되는 시간의 차이로 인하여, 서버(200) 측에서 제1 및 제2 장면메타정보를 순차적으로 생성하여 전송하는 것을 예시하고 있으나 이를 제한하지는 않는다. 따라서, 음성 구간 분석 및 음성 인식이 모두 완료 이후에 하나의 장면메타정보를 생성하여 사용자 단말로 전송할 수 있음은 당업자에게 자명할 것이다.
또한, 다른 실시 예로, 서버(200)는 UGC 동영상에 관한 재생 구간별 장면메타정보뿐만 아니라, 상기 재생 구간별 장면메타정보에 대응하는 복수의 페이지 정보를 포함하는 비디오 슬라이드 파일을 생성하여 사용자 단말(200)로 전송할 수도 있다.
도 7은 본 발명의 일 실시 예에 따른 사용자 단말의 동작을 설명하는 흐름도이다.
도 7을 참조하면, 사용자 단말(300)은 사용자 명령 등에 따라 비디오 슬라이드 애플리케이션을 실행할 수 있다(S705).
사용자 단말(300)은, 해당 애플리케이션 실행 시, 미리 정의된 사용자 인터페이스를 디스플레이부에 표시할 수 있다. 이때, 상기 사용자 인터페이스는 비디오 슬라이드 파일들에 대응하는 썸네일 이미지들을 포함하는 이미지 목록 영역과, 비디오 슬라이드 애플리케이션의 동작 메뉴들을 포함하는 메뉴 영역으로 구성될 수 있으며 반드시 이에 제한되지는 않는다.
이러한 사용자 인터페이스를 통해 업로드 메뉴가 선택되면(S710), 사용자 단말(300)은 메모리에 저장된 UGC 동영상들에 대응하는 썸네일 이미지들을 포함하는 선택 목록화면을 디스플레이부에 표시할 수 있다. 상기 선택 목록화면을 통해 하나 이상의 UGC 동영상들이 선택되면, 사용자 단말(300)은 상기 선택된 UGC 동영상을 서버(200)로 업로드할 수 있다(S715).
한편, 다른 실시 예로, 사용자 단말(300)은, 업로드 메뉴 선택 시, 동영상 촬영 모드로 진입하여 새로운 UGC 동영상을 실시간으로 생성할 수 있다. 동영상 촬영이 완료되면, 사용자 단말(300)은 새로 생성된 UGC 동영상을 서버(200)로 업로드할 수 있다.
서버(200)는 사용자 단말(300)로부터 업로드된 UGC 동영상에 포함된 정보를 기반으로 재생 구간별 장면메타정보를 생성할 수 있다. 서버(200)는 재생 구간별 장면메타정보를 사용자 단말(300)로 전송할 수 있다.
사용자 단말(300)은 업로드한 UGC 동영상에 관한 재생 구간별 장면메타정보를 서버(200)로부터 수신할 수 있다(S720). 사용자 단말(300)은 재생 구간별 장면메타정보에 대응하는 페이지 정보들로 구성된 비디오 슬라이드 파일을 생성하여 메모리에 저장할 수 있다. 사용자 단말(300)은 메모리에 저장된 비디오 슬라이드 파일에 대응하는 썸네일 이미지를 사용자 인터페이스의 이미지 목록 영역에 표시할 수 있다.
사용자 단말(300)은, 메모리에 저장된 비디오 슬라이드 파일을 다른 사람들에게 공유할지 여부를 문의하기 위한 팝업창을 디스플레이부에 표시할 수 있다(S725). 상기 팝업창을 통해 공유 메뉴가 선택되는 경우, 사용자 단말(300)은 비디오 슬라이드 파일을 다른 사람들에게 공유할 수 있다. 여기서, 해당 파일을 공유 받는 사람들은 특정 웹 사이트에 가입한 사람들이거나 혹은 비디오 슬라이드 애플리케이션을 설치한 사람들일 수 있으며 반드시 이에 제한되지는 않는다.
한편, 사용자 인터페이스에 표시된 썸네일 이미지가 선택되면(S735), 사용자 단말(300)은 비디오 슬라이드 모드로 진입하여 상기 선택된 썸네일 이미지에 대응하는 비디오 슬라이드 파일을 실행할 수 있다(S740).
사용자 단말(300)은, 비디오 슬라이드 모드 시, 비디오 슬라이드 파일을 구성하는 복수의 페이지들 중 첫 번째 페이지 화면을 디스플레이부에 표시할 수 있다. 이때, 상기 페이지 화면은 UGC 동영상의 재생 구간별 대표 이미지 정보와 자막 정보로 구성될 수 있다. 또한, 사용자 단말(300)은, 페이지 화면 표시 시, 해당 페이지에 대응하는 음성 정보를 출력할 수 있다.
사용자 단말(300)은, 미리 결정된 제스쳐 입력(가령, 방향성을 갖는 플리킹 입력)에 대응하여, 다음 페이지 화면 혹은 이전 페이지 화면을 디스플레이부에 표시할 수 있다.
이러한 비디오 슬라이드 모드에서, 모드 전환 메뉴가 선택되면(S745), 사용자 단말(300)은 비디오 슬라이드 모드를 동영상 재생 모드로 전환하고, 현재 페이지의 음성 구간에 대응하는 UGC 동영상을 재생할 수 있다(S750).
또한, 비디오 슬라이드 모드에서, 음성 인식 메뉴가 선택되면(S755), 사용자 단말(300)은 마이크로폰을 활성화하여 음성 인식 모드로 진입할 수 있다. 상기 음성 인식 모드에서 사용자의 음성 명령이 마이크로폰을 통해 입력되면, 사용자 단말(300)은 상기 음성 명령에 대응하는 비디오 슬라이드 동작을 실행할 수 있다(S760). 예컨대, 사용자 단말(300)은 음성 명령을 통해 페이지 넘김 기능, 모드 전환 기능(비디오 슬라이드 모드⇔동영상 모드), 자막 검색 기능, 자동 재생 기능, 삭제/편집/공유 기능 등을 실행할 수 있다.
또한, 비디오 슬라이드 모드에서, 미리 결정된 제스쳐 입력이 수신되면(S765), 사용자 단말(300)은 고속 탐색 기능을 실행할 수 있다(S770). 즉, 사용자 단말(300)은 UGC 동영상에 관한 페이지 화면을 빠른 속도로 전환(이동)할 수 있다. 한편, 이외에도, 사용자 단말(300)은 자막 검색 기능 및 태그 검색 기능 등을 실행할 수 있다.
사용자 단말(300)은, 비디오 슬라이드 애플리케이션이 종료될 때까지, 상술한 710 단계 내지 770 단계의 동작을 반복적으로 수행할 수 있다. 한편, 본 실시 예에서는, 별도의 독립적인 애플리케이션으로 비디오 슬라이드 서비스를 제공하는 것을 예시하고 있으나 이를 제한하지 않으며, 일반적인 동영상 재생 애플리케이션의 부가 기능으로서 비디오 슬라이드 서비스를 제공할 수 있음은 당업자에게 자명할 것이다.
도 8은 비디오 슬라이드 애플리케이션의 메인 화면을 표시하는 사용자 단말의 동작을 설명하기 위해 참조되는 도면이다.
도 8을 참조하면, 사용자 단말(300)은 사용자 명령 등에 따라 홈 화면(710)을 디스플레이부에 표시할 수 있다. 이때, 상기 홈 화면(810)은 비디오 슬라이드 애플리케이션에 대응하는 앱 아이콘(815)을 포함하고 있음을 가정한다.
단말 사용자에 의해 해당 앱 아이콘(815)이 선택되는 경우, 사용자 단말(300)은 상기 선택된 앱 아이콘(815)에 대응하는 비디오 슬라이드 애플리케이션을 실행할 수 있다.
사용자 단말(300)은, 해당 애플리케이션 실행 시, 미리 정의된 사용자 인터페이스(820)를 디스플레이부에 표시할 수 있다. 상기 사용자 인터페이스(820)는 비디오 슬라이드 파일들에 대응하는 썸네일 이미지들을 포함하는 이미지 목록 영역(820)과, 상기 이미지 목록 영역(820)의 상단에 표시되는 메뉴 영역(830)으로 구성될 수 있다. 상기 메뉴 영역(830)은, 공유 리스트 메뉴(831) 및 마이 리스트 메뉴(832) 등을 포함할 수 있다. 각각의 썸네일 이미지의 하단에는 해당 비디오 슬라이드 파일의 타이틀 정보가 표시될 수 있다.
사용자 단말(300)은, 공유 리스트 메뉴(831) 선택 시, 공유 상태에 있는 비디오 슬라이드 파일들에 대응하는 썸네일 이미지들을 이미지 목록 영역(820)에 표시할 수 있다. 한편, 사용자 단말(300)은, 마이 리스트 메뉴(832) 선택 시, 메모리에 저장된 비디오 슬라이드 파일들에 대응하는 썸네일 이미지들을 이미지 목록 영역(820)에 표시할 수 있다.
도 9 및 도 10은 UGC 동영상을 업로드하여 재생 구간별 장면메타정보를 수신하는 사용자 단말의 동작을 설명하기 위해 참조되는 도면이다.
도 9 및 도 10을 참조하면, 사용자 단말(300)은, 비디오 슬라이드 애플리케이션 실행 시, 미리 정의된 사용자 인터페이스(910)를 디스플레이부에 표시할 수 있다.
사용자 인터페이스(910)의 일 영역에 표시된 업로드 메뉴(915)가 선택되면, 사용자 단말(300)은 UGC 동영상의 업로드 방법을 선택하기 위한 팝업창(920)을 디스플레이부에 표시할 수 있다. 상기 팝업창(920)은 앨범 메뉴(921)와 카메라 메뉴(922)를 포함할 수 있다.
팝업창(920)의 앨범 메뉴(921)가 선택되면, 사용자 단말(300)은 메모리에 저장된 UGC 동영상들에 대응하는 썸네일 이미지들을 포함하는 선택 목록화면(미도시)을 디스플레이부에 표시할 수 있다. 상기 선택 목록화면을 통해 하나 이상의 썸네일 이미지가 선택되면, 사용자 단말(300)은 상기 선택된 썸네일 이미지에 대응하는 UGC 동영상을 서버(200)로 업로드할 수 있다.
한편, 팝업창(920)의 카메라 메뉴(922)가 선택되면, 사용자 단말(300)은 동영상 촬영 모드로 진입하여 새로운 UGC 동영상을 생성할 수 있다. 동영상 촬영이 완료되면, 사용자 단말(300)은 새로 생성된 UGC 동영상을 서버(200)로 업로드할 수 있다.
상기 팝업창(920)을 통해 선택된 UGC 동영상(930)이 서버(200)로 업로드된 경우, 사용자 단말(300)은 해당 UGC 동영상이 비디오 슬라이드 파일로 변환되는 과정을 설명하는 알림 정보(940, 950, 960)를 디스플레이부에 표시할 수 있다. 가령, 도 10에 도시된 바와 같이, 사용자 단말(300)은 "업로드 중", "음성 구간 추출 중", "음성 인식 중" 등과 같은 알림 메시지를 순차적으로 표시할 수 있다. 상기 변환 과정이 완료되면, 사용자 단말(300)은 업로드한 UGC 동영상에 관한 재생 구간별 장면메타정보를 서버(200)로부터 수신할 수 있다.
도 11은 비디오 슬라이드 파일을 공유하는 사용자 단말의 동작을 설명하기 위해 참조되는 도면이다.
도 11을 참조하면, 사용자 단말(300)은, UGC 동영상 업로드 시, 상기 UGC 동영상에 관한 재생 구간별 장면메타정보를 서버(200)로부터 실시간으로 수신할 수 있다. 사용자 단말(300)은 재생 구간별 장면메타정보에 대응하는 복수의 페이지들을 포함하는 비디오 슬라이드 파일을 생성할 수 있다. 이때, 사용자 단말(300)은, 비디오 슬라이드 파일을 다른 사람들에게 공유할지 여부를 문의하기 위한 팝업창(1110)을 디스플레이부에 표시할 수 있다.
상기 팝업창(1110)을 통해 확인 메뉴(1115)가 선택되면, 사용자 단말(300)은 비디오 슬라이드 파일을 다른 사람들에게 공유할 수 있다. 사용자 단말(300)은 비디오 슬라이드 파일에 대응하는 썸네일 이미지(1125)를 사용자 인터페이스(1120)에 표시할 수 있다.
도 12 내지 도 16은 UGC 동영상을 페이지 단위로 표시하는 사용자 단말의 동작을 설명하기 위해 참조되는 도면이다.
도 12 내지 도 16을 참조하면, 사용자 단말(300)은, 비디오 슬라이드 애플리케이션 실행 시, 복수의 썸네일 이미지들을 포함하는 사용자 인터페이스를 디스플레이부에 표시할 수 있다. 이때, 상기 복수의 썸네일 이미지들은 UGC 동영상들에 관한 비디오 슬라이드 파일들에 대응한다.
이러한 사용자 인터페이스를 통해 썸네일 이미지가 선택되면, 사용자 단말(300)은 상기 선택된 썸네일 이미지에 대응하는 비디오 슬라이드 파일을 실행(재생)할 수 있다. 즉, 사용자 단말(300)은 UGC 동영상을 책처럼 페이지 단위로 보여주는 비디오 슬라이드 모드로 진입할 수 있다.
사용자 단말(300)은, 비디오 슬라이드 모드 시, 미리 결정된 페이지 화면(1200)을 디스플레이부에 표시할 수 있다. 이때, 상기 페이지 화면(1200)은, 이미지 표시 영역(1210), 자막 표시 영역(1220), 제1 메뉴 영역(1230) 및 제2 메뉴 영역(1240) 등을 포함할 수 있으며 반드시 이에 제한되지는 않는다.
도 12에 도시된 바와 같이, 이미지 표시 영역(1210)은 현재 페이지에 대응하는 대표 이미지를 포함할 수 있다. 자막 표시 영역(1220)은 현재 페이지에 대응하는 자막 정보를 포함할 수 있다. 제1 및 제2 메뉴 영역(1230, 1240)은 비디오 슬라이드 모드와 관련된 기능들을 실행하기 위한 복수의 메뉴들을 포함할 수 있다.
제1 메뉴 영역(1230)은, 메인 화면으로 이동하기 위한 제1 동작 메뉴(메인 메뉴, 1231), 동영상 재생 모드로 시청하기 위한 제2 동작 메뉴(모드 전환 메뉴, 1232), 자막 및/또는 태그를 검색하기 위한 제3 동작 메뉴(검색 메뉴, 1233) 및 다른 메뉴들을 더 보기 위한 제4 동작 메뉴(더보기 메뉴, 1234) 등을 포함할 수 있다. 제2 메뉴 영역(1240)은 페이지 화면을 자동으로 전환하기 위한 제5 동작 메뉴(자동 전환 메뉴, 1241), 현재 페이지의 전후 페이지들을 미리보기 위한 제6 동작 메뉴(미리보기 메뉴, 1242) 및 음성 인식 모드를 활성화하기 위한 제7 동작 메뉴(마이크 메뉴, 1243) 등을 포함할 수 있다.
이러한 페이지 화면(1200)이 표시된 상태에서, 제2 메뉴 영역(1240)의 미리보기 메뉴(1242)가 선택되는 경우, 사용자 단말(300)은 현재 페이지의 전후 페이지들을 미리 보여주는 기능을 실행할 수 있다. 가령, 도 13에 도시된 바와 같이, 사용자 단말(300)은 현재 페이지를 기준으로 전후에 존재하는 페이지들에 대응하는 복수의 썸네일 이미지들을 포함하는 스크롤 영역(1250)을 디스플레이부의 하단에 표시할 수 있다. 상기 복수의 썸네일 이미지들은 복수의 페이지들에 대응하는 대표 이미지들을 미리 결정된 크기로 축소한 이미지들이다. 상기 복수의 썸네일 이미지들은 페이지들의 타임코드에 따라 순차적으로 배열될 수 있다. 또한, 상기 복수의 썸네일 이미지들은 미리 결정된 제스쳐 입력에 따라 스크롤 가능하도록 구성될 수 있다.
현재 페이지의 썸네일 이미지(1251)는 스크롤 영역(1250)의 중앙부에 위치할 수 있다. 즉, 스크롤 영역(1250)의 중앙부에는 현재 시청자가 보고 있는 페이지가 위치할 수 있다. 시청자는 스크롤 영역(1250)에 위치한 썸네일 이미지들 중 어느 하나를 선택함으로써 해당 썸네일 이미지에 대응하는 페이지로 바로 이동할 수 있다.
한편, 제1 메뉴 영역(1230)의 메인 메뉴(1231)가 선택되는 경우, 사용자 단말(300)은 비디오 슬라이드 모드를 종료한 후 해당 애플리케이션의 초기 화면으로 이동할 수 있다.
제1 메뉴 영역(1230)의 모드 전환 메뉴(1232)가 선택되는 경우, 사용자 단말(300)은, 비디오 슬라이드 모드를 동영상 재생 모드로 전환한 후 현재 페이지의 음성 구간에 대응하는 UGC 동영상의 재생 구간을 재생할 수 있다. 가령, 도 14에 도시된 바와 같이, 사용자 단말(300)은 현재 페이지에 대응하는 동영상 재생 화면(1260)을 디스플레이부에 표시할 수 있다.
제1 메뉴 영역(1230)의 검색 메뉴(1233)가 선택되는 경우, 사용자 단말(300)은 자막 또는 태그를 검색하기 위한 검색창(미도시)을 디스플레이부에 표시할 수 있다. 상기 검색창을 통해 소정의 텍스트 정보가 입력되면, 사용자 단말(300)은 상기 텍스트 정보에 대응하는 자막 또는 태그 정보를 포함하는 페이지를 검색하여 디스플레이부에 표시할 수 있다.
제2 메뉴 영역(1240)의 마이크 메뉴(1243)가 선택되는 경우, 사용자 단말(300)은 마이크로폰을 활성화하여 음성 인식 모드로 진입할 수 있다. 사용자 단말(300)은, 음성 인식 모드 진입 시, 도 15에 도시된 바와 같은 알림 정보(1270)를 디스플레이부에 표시할 수 있다. 이러한 음성 인식 모드에서 사용자의 음성 명령이 마이크로폰을 통해 입력되면, 사용자 단말(300)은 상기 음성 명령에 대응하는 비디오 슬라이드 동작을 실행할 수 있다. 이에 따라, 단말 사용자는 손을 사용하기 힘든 상황에서도 비디오 슬라이드 서비스를 간편하게 이용할 수 있다.
한편, 페이지 화면(1200)이 표시된 상태에서, 디스플레이부를 통해 제1 방향성을 갖는 플리킹 입력이 수신되는 경우, 사용자 단말(300)은 현재 페이지의 다음 페이지 화면을 디스플레이부에 표시할 수 있다. 반대로, 디스플레이부를 통해 제2 방향성을 갖는 플리킹 입력이 수신되는 경우, 사용자 단말(300)은 현재 페이지의 이전 페이지 화면을 디스플레이부에 표시할 수 있다. 이처럼, 사용자 단말(300)은 미리 결정된 제스쳐 입력을 통해 페이지 화면을 간편하게 전환할 수 있다. 제2 메뉴 영역(1240)의 자동 전환 메뉴(1241)가 선택되는 경우, 사용자 단말(300)은 현재 페이지 화면을 일정 시간 동안 표시한 후 다음 페이지 화면으로 자동 전환할 수 있다.
또한, 사용자 단말(300)은 미리 결정된 제스쳐 입력에 대응하여 고속 탐색 기능을 수행할 수 있다. 가령, 도 16에 도시된 바와 같이, 페이지 화면(1200)의 우측 영역을 일정 시간 동안 롱 터치하는 입력(long touch input, 1280)이 수신되는 경우, 사용자 단말(300)은 UGC 동영상에 관한 페이지 화면을 빠른 속도로 전환(이동)할 수 있다. 이때, 사용자 단말(300)은 전체 페이지 수에 관한 정보(1290)와 화면 전환될 페이지 번호에 관한 정보(1295)를 디스플레이부의 일 영역에 표시할 수 있다.
한편, 이외에도, 사용자 단말(300)은, 시청자(사용자)의 화면 분할 요청에 대응하여, 디스플레이 영역을 미리 결정된 개수로 분할하고, 상기 분할된 영역에 복수의 페이지를 표시할 수 있다. 또한, 사용자 단말(300)은, 시청자(사용자)의 재생/정지 요청에 대응하여, 현재 페이지에 대응하는 음성 정보를 재생하거나 정지할 수 있다.
이상 상술한 바와 같이, 사용자 단말(300)은 서버(200)와 연동하여 UGC 동영상을 책처럼 페이지 단위로 시청할 수 있는 비디오 슬라이드 서비스를 제공할 수 있다.
전술한 본 발명은, 프로그램이 기록된 매체에 컴퓨터가 읽을 수 있는 코드로서 구현하는 것이 가능하다. 컴퓨터가 읽을 수 있는 매체는, 컴퓨터로 실행 가능한 프로그램을 계속 저장하거나, 실행 또는 다운로드를 위해 임시 저장하는 것일 수도 있다. 또한, 매체는 단일 또는 수개 하드웨어가 결합된 형태의 다양한 기록수단 또는 저장수단일 수 있는데, 어떤 컴퓨터 시스템에 직접 접속되는 매체에 한정되지 않고, 네트워크 상에 분산 존재하는 것일 수도 있다. 매체의 예시로는, 하드 디스크, 플로피 디스크 및 자기 테이프와 같은 자기 매체, CD-ROM 및 DVD와 같은 광기록 매체, 플롭티컬 디스크(floptical disk)와 같은 자기-광 매체(magneto-optical medium), 및 ROM, RAM, 플래시 메모리 등을 포함하여 프로그램 명령어가 저장되도록 구성된 것이 있을 수 있다. 또한, 다른 매체의 예시로, 애플리케이션을 유통하는 앱 스토어나 기타 다양한 소프트웨어를 공급 내지 유통하는 사이트, 서버 등에서 관리하는 기록매체 내지 저장매체도 들 수 있다. 따라서, 상기의 상세한 설명은 모든 면에서 제한적으로 해석되어서는 아니되고 예시적인 것으로 고려되어야 한다. 본 발명의 범위는 첨부된 청구항의 합리적 해석에 의해 결정되어야 하고, 본 발명의 등가적 범위 내에서의 모든 변경은 본 발명의 범위에 포함된다.
Claims (20)
- UGC(User Generated Content) 동영상을 서버로 업로드하는 단계;상기 서버로부터 상기 UGC 동영상에 관한 재생 구간별 장면메타정보를 수신하는 단계;상기 수신된 재생 구간별 장면메타정보를 기반으로 비디오 슬라이드 파일을 생성하는 단계;상기 비디오 슬라이드 파일에 대응하는 항목을 표시하는 단계; 및상기 항목 선택 시, 상기 UGC 동영상의 재생 구간별 대표 이미지 정보와 자막 정보로 구성되는 페이지 화면을 표시하는 단계를 포함하는 사용자 단말의 동작 방법.
- 제1항에 있어서,상기 비디오 슬라이드 파일은, 상기 UGC 동영상의 재생 구간별 장면메타정보에 대응하는 페이지 정보들로 구성되는 것을 특징으로 하는 사용자 단말의 동작 방법.
- 제1항에 있어서,상기 자막 정보는, 상기 UGC 동영상에서 추출된 오디오 정보를 음성 인식하여 생성되는 것을 특징으로 하는 사용자 단말의 동작 방법.
- 제1항에 있어서,상기 항목은, 썸네일 이미지 형태로 표시되는 것을 특징으로 하는 사용자 단말의 동작 방법.
- 제1항에 있어서,상기 재생 구간별 장면메타정보는, 타임코드 정보, 대표 이미지 정보 및 자막 정보를 포함하는 것을 특징으로 하는 사용자 단말의 동작 방법.
- 제1항에 있어서, 상기 수신 단계는,상기 재생 구간별 장면메타정보를 스트리밍 방식으로 수신하는 것을 특징으로 하는 사용자 단말의 동작 방법.
- 제1항에 있어서, 상기 수신 단계는,상기 UGC 동영상에 관한 재생구간 별 타임코드 정보 및 대표 이미지 정보를 포함하는 제1 장면메타정보를 수신하는 단계; 및상기 UGC 동영상에 관한 재생구간 별 자막 정보를 포함하는 제2 장면메타정보를 수신하는 단계를 포함하는 것을 특징으로 하는 사용자 단말의 동작 방법.
- 제1항에 있어서,공유 메뉴 선택 시, 상기 비디오 슬라이드 파일을 공유하기 위한 제1 팝업창을 표시하는 단계를 더 포함하는 사용자 단말의 동작 방법.
- 제1항에 있어서,음성 인식 모드 시, 사용자의 음성 명령에 대응하는 비디오 슬라이드 동작을 수행하는 단계를 더 포함하는 사용자 단말의 동작 방법.
- 제1항에 있어서,미리 결정된 제스쳐 입력에 대응하여, 상기 UGC 동영상에 관한 페이지 화면을 빠른 속도로 전환하는 단계를 더 포함하는 사용자 단말의 동작 방법.
- 제1항에 있어서,모드 전환 메뉴 선택 시, 상기 페이지 화면의 음성 구간에 대응하는 UGC 동영상을 재생하는 단계를 더 포함하는 사용자 단말의 동작 방법.
- 제1항에 있어서,미리보기 메뉴 선택 시, 현재 페이지를 기준으로 전후에 존재하는 페이지들에 대응하는 복수의 썸네일 이미지들을 포함하는 스크롤 영역을 표시하는 단계를 더 포함하는 사용자 단말의 동작 방법.
- 제1항에 있어서, 상기 업로드 단계는,업로드 메뉴 선택 시, 상기 UGC 동영상의 업로드 방법을 선택하기 위한 제2 팝업창을 표시하는 단계를 더 포함하는 사용자 단말의 동작 방법.
- 제13항에 있어서,상기 제2 팝업창은, 메모리에 저장된 UGC 동영상을 업로드하기 위한 제1 메뉴와, 카메라를 통해 실 시간으로 촬영한 UGC 동영상을 업로드하기 위한 제2 메뉴를 포함하는 것을 특징으로 하는 사용자 단말의 동작 방법.
- 제1항에 있어서,상기 UGC 동영상이 상기 비디오 슬라이드 파일로 변환되는 과정을 설명하는 알림 정보를 표시하는 단계를 더 포함하는 사용자 단말의 동작 방법.
- 제1항에 있어서,상기 페이지 화면 표시 시, 해당 페이지에 대응하는 음성 정보를 출력하는 단계를 더 포함하는 사용자 단말의 동작 방법.
- 제1항에 있어서,비디오 슬라이드 애플리케이션에 대응하는 앱 아이콘을 표시하는 단계; 및상기 앱 아이콘 선택 시, 미리 정의된 사용자 인터페이스 화면을 표시하는 단계를 더 포함하는 사용자 단말의 동작 방법.
- 제17항에 있어서,상기 사용자 인터페이스 화면은, 비디오 슬라이드 파일들에 대응하는 썸네일 이미지들을 포함하는 파일 목록 영역과 상기 비디오 슬라이드 애플리케이션과 관련된 메뉴들을 포함하는 메뉴 영역을 포함하는 것을 특징으로 하는 사용자 단말의 동작 방법.
- 제1항 내지 제18항 중 어느 하나의 항에 따른 방법이 컴퓨터 상에서 실행되도록 컴퓨터로 판독 가능한 저장매체에 기록된 컴퓨터 프로그램.
- 서버와의 통신 인터페이스를 제공하는 통신부;미리 정의된 사용자 인터페이스를 표시하는 디스플레이부; 및상기 사용자 인터페이스에 포함된 업로드 메뉴를 이용하여 UGC(User Generated Content) 동영상을 상기 서버로 업로드하고, 상기 서버로부터 상기 UGC 동영상에 관한 재생 구간별 장면메타정보를 수신한 경우, 상기 재생 구간별 장면메타정보를 기반으로 비디오 슬라이드 파일을 생성하며, 상기 비디오 슬라이드 파일에 대응하는 항목을 표시하고, 상기 항목 시, 상기 UGC 동영상의 재생 구간별 대표 이미지 정보와 자막 정보로 구성되는 페이지 화면을 상기 디스플레이부에 표시하는 제어부를 포함하는 사용자 단말.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/236,494 US11496806B2 (en) | 2018-10-24 | 2021-04-21 | Content providing server, content providing terminal, and content providing method |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR1020180127336A KR102142623B1 (ko) | 2018-10-24 | 2018-10-24 | 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법 |
| KR10-2018-0127336 | 2018-10-24 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/236,494 Continuation US11496806B2 (en) | 2018-10-24 | 2021-04-21 | Content providing server, content providing terminal, and content providing method |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020085675A1 true WO2020085675A1 (ko) | 2020-04-30 |
Family
ID=70331529
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2019/013138 Ceased WO2020085675A1 (ko) | 2018-10-24 | 2019-10-07 | 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US11496806B2 (ko) |
| KR (1) | KR102142623B1 (ko) |
| WO (1) | WO2020085675A1 (ko) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117198291A (zh) * | 2023-11-08 | 2023-12-08 | 四川蜀天信息技术有限公司 | 一种语音控制终端界面的方法、装置及系统 |
Families Citing this family (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109889880B (zh) * | 2019-04-11 | 2021-02-12 | 北京字节跳动网络技术有限公司 | 关注用户的信息展示方法、装置、设备及存储介质 |
| CN111930973B (zh) * | 2020-08-14 | 2022-06-21 | 北京字节跳动网络技术有限公司 | 多媒体数据的播放方法、装置、电子设备及存储介质 |
| WO2022039301A1 (ko) * | 2020-08-20 | 2022-02-24 | 주식회사 누날 | 비디오 큐레이션 서비스 방법 |
| CN112866796A (zh) * | 2020-12-31 | 2021-05-28 | 北京字跳网络技术有限公司 | 视频生成方法、装置、电子设备和存储介质 |
| KR102848564B1 (ko) * | 2021-11-25 | 2025-08-22 | 동서대학교 산학협력단 | 인공지능 기술기반 스트리밍 영상 검색 시스템 및 방법 |
| CN115048597A (zh) * | 2022-05-31 | 2022-09-13 | 北京字跳网络技术有限公司 | 页面展示方法、装置及电子设备 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20120005000A (ko) * | 2009-05-01 | 2012-01-13 | 소니 주식회사 | 서버 장치, 전자 기기, 전자 서적 제공 시스템, 전자 서적 제공 방법, 전자 서적 표시 방법 및 프로그램 |
| KR20150009053A (ko) * | 2013-07-11 | 2015-01-26 | 엘지전자 주식회사 | 이동 단말기 및 그것의 제어 방법 |
| KR20150122673A (ko) * | 2013-03-06 | 2015-11-02 | 톰슨 라이센싱 | 비디오의 화상 요약 |
| KR20160044981A (ko) * | 2014-10-16 | 2016-04-26 | 삼성전자주식회사 | 동영상 처리 장치 및 방법 |
| KR101769071B1 (ko) * | 2016-05-10 | 2017-08-18 | 네이버 주식회사 | 비디오 태그 제작 및 활용을 위한 방법 및 시스템 |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20080288890A1 (en) * | 2007-05-15 | 2008-11-20 | Netbriefings, Inc | Multimedia presentation authoring and presentation |
| KR101008494B1 (ko) | 2008-06-17 | 2011-01-14 | 테크빌닷컴 주식회사 | 사용자 저작 콘텐츠 제공 시스템 및 방법 |
| US10095348B2 (en) * | 2014-06-25 | 2018-10-09 | Lg Electronics Inc. | Mobile terminal and method for controlling the same |
| KR102319462B1 (ko) * | 2014-12-15 | 2021-10-28 | 조은형 | 미디어 콘텐츠 플레이백 제어 방법 및 이를 수행하는 전자 기기 |
| KR101832966B1 (ko) * | 2015-11-10 | 2018-02-28 | 엘지전자 주식회사 | 이동 단말기 및 이의 제어방법 |
| KR20170087307A (ko) | 2016-01-20 | 2017-07-28 | 엘지전자 주식회사 | 디스플레이 디바이스 및 그 제어 방법 |
| KR101686425B1 (ko) | 2016-11-17 | 2016-12-14 | 주식회사 엘지유플러스 | 동영상 관리 서버 및 동영상 재생 장치, 이들을 이용한 등장 인물 정보 제공 방법 |
| US20180211556A1 (en) * | 2017-01-23 | 2018-07-26 | Rovi Guides, Inc. | Systems and methods for adjusting display lengths of subtitles based on a user's reading speed |
-
2018
- 2018-10-24 KR KR1020180127336A patent/KR102142623B1/ko active Active
-
2019
- 2019-10-07 WO PCT/KR2019/013138 patent/WO2020085675A1/ko not_active Ceased
-
2021
- 2021-04-21 US US17/236,494 patent/US11496806B2/en active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20120005000A (ko) * | 2009-05-01 | 2012-01-13 | 소니 주식회사 | 서버 장치, 전자 기기, 전자 서적 제공 시스템, 전자 서적 제공 방법, 전자 서적 표시 방법 및 프로그램 |
| KR20150122673A (ko) * | 2013-03-06 | 2015-11-02 | 톰슨 라이센싱 | 비디오의 화상 요약 |
| KR20150009053A (ko) * | 2013-07-11 | 2015-01-26 | 엘지전자 주식회사 | 이동 단말기 및 그것의 제어 방법 |
| KR20160044981A (ko) * | 2014-10-16 | 2016-04-26 | 삼성전자주식회사 | 동영상 처리 장치 및 방법 |
| KR101769071B1 (ko) * | 2016-05-10 | 2017-08-18 | 네이버 주식회사 | 비디오 태그 제작 및 활용을 위한 방법 및 시스템 |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117198291A (zh) * | 2023-11-08 | 2023-12-08 | 四川蜀天信息技术有限公司 | 一种语音控制终端界面的方法、装置及系统 |
| CN117198291B (zh) * | 2023-11-08 | 2024-01-23 | 四川蜀天信息技术有限公司 | 一种语音控制终端界面的方法、装置及系统 |
Also Published As
| Publication number | Publication date |
|---|---|
| KR102142623B1 (ko) | 2020-08-10 |
| US11496806B2 (en) | 2022-11-08 |
| US20210243502A1 (en) | 2021-08-05 |
| KR20200046327A (ko) | 2020-05-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020085675A1 (ko) | 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법 | |
| KR102085908B1 (ko) | 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법 | |
| WO2016028042A1 (en) | Method of providing visual sound image and electronic device implementing the same | |
| WO2018080162A1 (ko) | 음성 명령에 기초하여 애플리케이션을 실행하는 방법 및 장치 | |
| WO2014007502A1 (en) | Display apparatus, interactive system, and response information providing method | |
| WO2019112342A1 (en) | Voice recognition apparatus and operation method thereof cross-reference to related application | |
| WO2016093552A2 (en) | Terminal device and data processing method thereof | |
| WO2014193161A1 (ko) | 멀티미디어 콘텐츠 검색을 위한 사용자 인터페이스 방법 및 장치 | |
| WO2013187715A1 (en) | Server and method of controlling the same | |
| WO2016022002A1 (en) | Apparatus and method for controlling content by using line interaction | |
| WO2018038428A1 (en) | Electronic device and method for rendering 360-degree multimedia content | |
| WO2015088155A1 (en) | Interactive system, server and control method thereof | |
| WO2015147437A1 (ko) | 모바일 서비스 시스템, 그 시스템에서의 위치 기반 앨범 생성 방법 및 장치 | |
| WO2016080660A1 (en) | Content processing device and method for transmitting segment of variable size | |
| EP3230902A2 (en) | Terminal device and data processing method thereof | |
| WO2020075926A1 (ko) | 모바일 장치 및 모바일 장치의 제어 방법 | |
| WO2022196973A1 (ko) | 음악 컨텐츠를 추천하는 방법 및 장치 | |
| WO2020091249A1 (ko) | 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법 | |
| WO2020159047A1 (ko) | 보이스 어시스턴트 서비스를 이용한 컨텐츠 재생 장치 및 그 동작 방법 | |
| WO2024219903A1 (en) | Method and apparatus for providing service based on emotion information of user about content | |
| KR20180114769A (ko) | 내용을 기반으로 하는 동영상 검색시스템 | |
| WO2022139428A1 (ko) | 노래방 지원 서비스 제공 방법 및 그 시스템 | |
| WO2016093551A1 (en) | Server and method for generating slide show thereof | |
| WO2023074918A1 (ko) | 디스플레이 장치 | |
| KR100944958B1 (ko) | 특정 구간의 멀티미디어 데이터 및 캡션 데이터를 제공하는장치 및 서버 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19875659 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19875659 Country of ref document: EP Kind code of ref document: A1 |