EP2462536A1 - Systems and methods for three-dimensional video generation - Google Patents

Systems and methods for three-dimensional video generation

Info

Publication number
EP2462536A1
EP2462536A1 EP10807016A EP10807016A EP2462536A1 EP 2462536 A1 EP2462536 A1 EP 2462536A1 EP 10807016 A EP10807016 A EP 10807016A EP 10807016 A EP10807016 A EP 10807016A EP 2462536 A1 EP2462536 A1 EP 2462536A1
Authority
EP
European Patent Office
Prior art keywords
training image
video
rules
internal data
generating
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP10807016A
Other languages
German (de)
French (fr)
Other versions
EP2462536A4 (en
Inventor
Haohong Wang
Glenn Adler
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Shenzhen TCL New Technology Co Ltd
Original Assignee
Shenzhen TCL New Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Shenzhen TCL New Technology Co Ltd filed Critical Shenzhen TCL New Technology Co Ltd
Publication of EP2462536A1 publication Critical patent/EP2462536A1/en
Publication of EP2462536A4 publication Critical patent/EP2462536A4/en
Withdrawn legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/50Depth or shape recovery
    • G06T7/55Depth or shape recovery from multiple images
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/20Image signal generators
    • H04N13/261Image signal generators with monoscopic-to-stereoscopic image conversion
    • H04N13/264Image signal generators with monoscopic-to-stereoscopic image conversion using the relative movement of objects in two video frames or fields

Definitions

  • This disclosure relates to systems and methods for three- dimensional video generation.
  • 3D TV has recently been foreseen as part of a next wave of promising technologies for consumer electronics.
  • 3D technologies incorporate a third dimension of depth into an image, which may provide a stereographic perception to a viewer of the image.
  • a system for generating three-dimensional (3D) video based on a two-dimensional (2D) input image sequence including at least one 2D input image comprising: a rule generator configured to generate rules for 2D-to-3D conversion; and a 3D video converter coupled to the rule generator, and configured to obtain the rules from the rule generator and automatically convert the 2D input image sequence to the 3D video based on the obtained rules.
  • a computer-implemented method for generating three-dimensional (3D) video based on a two-dimensional (2D) input image sequence including at least one 2D input image comprising: generating rules for 2D-to-3D conversion; and automatically converting the 2D input image sequence to the 3D video based on the rules.
  • a computer-readable medium including instructions, executable by a processor of a three-dimensional (3D) video generating system, for performing a method for generating 3D video based on a two-dimensional (2D) input image sequence including at least one 2D input image, the method comprising:
  • FIG. 1 illustrates a block diagram of a system for generating 3D video, according to an exemplary embodiment.
  • Fig. 2 illustrates a block diagram of a rule generator, according to an exemplary embodiment.
  • Fig. 3 illustrates a block diagram of a 3D video convertor, according to an exemplary embodiment.
  • Fig. 4 illustrates a block diagram of a 3D video generator
  • Fig. 5 illustrates a block diagram of a 3D video generator
  • FIG. 6 illustrates a flowchart of a method for generating 3D video based on a 2D input image sequence, according to an exemplary embodiment.
  • Fig. 1 illustrates a block diagram of a system 100 for generating a three-dimensional (3D) video, according to an exemplary embodiment.
  • the system 100 may include a two-dimensional (2D) video content source, such as a video storage medium 102 or a media server 104 connected with a network 106.
  • the system 100 also includes a video device 108, a 3D video generator 110, and a display device 112.
  • the video storage medium 102 may be any medium for storing video content.
  • the video storage medium 102 may be provided as a compact disc (CD), a digital video disc (DVD), a hard disk, a magnetic tape, a flash memory card/drive, a volatile or non-volatile memory, a holographic data storage, or any other storage medium.
  • the video storage medium 102 may be located within the video device 108, local to the video device 108, or remote from the video device 108.
  • the media server 104 may be a computer server that receives a request for 2D video content from the video device 108, processes the request, and provides 2D video content to the video device 108 through the network 106.
  • the media server 104 may be a web server, an enterprise server, or any other type of computer server.
  • the media server 104 is configured to accept requests from the video device 108 based on, e.g., a hypertext transfer protocol (HTTP) or other protocols that may initiate a video session, and to serve the video device 108 with 2D video content.
  • HTTP hypertext transfer protocol
  • the network 106 may include a wide area network (WAN), a local area network (LAN), a wireless network suitable for packet-type communications, such as Internet communications, a broadcast network, or any combination thereof.
  • the network 106 is configured to distribute digital or non-digital video content.
  • the video device 108 is a hardware device such as a computer, a personal digital assistant (PDA), a mobile phone, a laptop, a desktop, a videocassette recorder (VCR), a laserdisc player, a DVD player, a blue ray disc player, or any electronic device configured to output 2D video, i.e., a 2D image sequence.
  • the video device 108 may include software applications that allow the video device 108 to communicate with and receive 2D video content from, e.g., the video storage medium 102 or the media server 104.
  • the video device 108 may, by means of included software
  • the 3D video generator 110 is configured to generate 3D video based on the 2D image sequence outputted by the video device 108.
  • the generator 110 may be implemented as a hardware device that is either stand-alone or incorporated into the video device 108, or software applications installed on the video device 108, or a combination thereof.
  • the 3D video generator 110 may include a processor to generate the 3D video.
  • the 3D video generator 110 may include a rule generator 110-1 configured to generate rules, e.g., semantic rules, for automatic 2D-to-3D conversion, based on a 2D training image sequence, and a 3D video convertor 110-2 configured to automatically convert a 2D input image sequence to 3D video based on the rules.
  • the rule generator 110-1 is also referred to as an interactive 2D-to-3D conversion environment (i3DV)
  • the 3D video convertor 110-2 is also referred to as an automatic 2D-to-3D conversion solution (a3DC).
  • i3DV interactive 2D-to-3D conversion environment
  • a3DC automatic 2D-to-3D conversion solution
  • the display device 112 is configured to display images in the training image sequence or the 2D input image sequence, and to present the generated 3D video.
  • the display device 112 may be provided as a monitor, a projector, or any other video display device.
  • the display device 112 may also be a part of the video device 108.
  • the user may view the images in the training image sequence to provide inputs or feedbacks to the 3D video generator 110.
  • the user may also watch the generated 3D video displayed by the display device 112. It is to be understood that devices shown in Fig. 1 are for illustrative purposes. Certain devices may be removed or combined, and additional devices may be added.
  • Fig. 2 illustrates a block diagram of a rule generator 200, according to an exemplary embodiment.
  • the rule generator 200 may be the rule generator 110-1 (Fig. 1).
  • the rule generator 200 is configured to receive a 2D training image sequence including at least one 2D training image, and to generate 3D video from the training image sequence for preview.
  • the rule generator 200 may further generate rules, e.g., semantic rules, for automatic 2D- to-3D conversion based on the training image sequence.
  • the rule generator 200 may include one or more of the following components: an image sequence analyzing module 202, a scene structure detection module 204, an object segmentation module 206, an object tracking module 208, an object classification module 210, and an object depth estimation module 212.
  • the rule generator 200 may further include a quality evaluation module 214, a user interface 216, an object evolving detection module 218, a depth-image based rendering (DIBR) module 220, a memory device 222 for storing internal data, and a rule database 224.
  • DIBR depth-image based rendering
  • the image sequence analyzing module 202 is a hardware device or software configured to perform analysis of the training image sequence, to determine one or more key frames in the training image sequence. For example, the image sequence analyzing module 202 may analyze the training image sequence to detect scene changes therein, and break the training image sequence into one or more chunks each starting with a key frame where a scene change occurs. The image sequence analyzing module 202 may then output key frame indexes 230 for the determined key frames, and store the key frame indexes 230 as the internal data in the memory device 222.
  • the scene structure detection module 204 is a hardware device or software configured to analyze the training image in the training image sequence to obtain a scene structure 232 for the training image, which is also stored as the internal data in the memory device 222.
  • the scene structure detection module 204 may perform the analysis based on linear perspective properties of the training image, by detecting a vanishing point and vanishing lines in the training image.
  • a vanishing point represents a farthest point in a scene shown in a 2D image, and vanishing lines each represent a direction of increasing depth. The vanishing lines converge at the vanishing point in the 2D image.
  • the scene structure detection module 204 obtains a projected far-to-near direction in the training image and, thus, obtains the scene structure 232.
  • the object segmentation module 206, the object tracking module 208, the object classification module 210, and the object depth estimation module 212 are configured to generate object data 234 regarding semantic objects in the training image, and store the object data 234 as the internal data in the memory device 222.
  • the object segmentation module 206 is a hardware device or software configured to separate scene content in the training image into one or more constituent parts, i.e., semantic objects, referred to hereafter as objects.
  • objects corresponds to a region or shape in the training image that represents an entity with certain semantic meaning, such as a tree, a lake, a house, etc.
  • the object segmentation module 206 may detect the objects in the training image and segment out the detected objects, thereby generating the object data 234.
  • the object segmentation module 206 may group pixels of the training image into different regions based on a homogeneous low-level feature, such as color, motion, or texture, each of the regions representing one of the objects. As a result, the object segmentation module 206 detects and segments out the objects in the training image. In addition, the object segmentation module 206 may cut one or more foreground objects out of the training image, and refine boundaries of the foreground objects.
  • a homogeneous low-level feature such as color, motion, or texture
  • the object tracking module 208 is a hardware device or software configured to use an object tracking algorithm to track the objects in the training image sequence. For example, the object tracking module 208 may detect how positions/sizes of the objects change with time in the training image, thereby generating the object data 234. In such manner, temporal coherence may be provided for generating the 3D video from the training image sequence.
  • the object classification module 210 is a hardware device or software configured to perform object classification for the objects in the training image based on a plurality of training data sets in an object database (not shown).
  • Each of the training data sets includes a group of training images previously processed by the rule generator 200 and representing an object category, such as a building category, a grass category, a tree category, a cow category, a sky category, a face category, a car category, a bicycle category, etc.
  • the object classification module 210 may classify or identify the objects in the present training image, and assign category labels for the classified objects.
  • a training data set may initially include no training images or a relatively small number of training images.
  • the object classification module 210 operates together with the user interface 216 to perform user interactive object classification, as described below. As a result, the object classification module 210 learns additional object categories.
  • the object depth estimation module 212 is a hardware device or software configured to estimate depths of the objects in the training image. Typically, each pixel of the objects in a 2D image
  • a depth of a pixel of an object represents a distance between a viewer and a part of the object corresponding to that pixel.
  • the object depth estimation module 212 may detect the vanishing point and the vanishing lines in the training image, and generate a depth map of the training image accordingly. For example, the object depth estimation module 212 may generate different depth gradient planes relative to the vanishing point and the vanishing lines. The object depth estimation module 212 then assigns depth levels to pixels of the training image according to the depth gradient planes, thereby generating the object data 234. The object depth estimation module 212 may additionally estimate orientation or thicknesses of the objects in the training image.
  • the quality evaluation module 214 is a hardware device or software configured to provide evaluation of visual quality of the 3D video generated from the training image sequence.
  • the quality evaluation module 214 may provide a quality value 236 representing the visual quality of the generated 3D video, and store the quality value 236 in the memory device 222.
  • the quality value 236 is also part of the internal data in the memory device 222.
  • the user interface 216 is configured to receive user input and to output the 3D video generated from the training image sequence for preview, such that the rule generator 200 operates in a user interactive manner.
  • the user interface 216 may be implemented with a hardware device, such as a keyboard or a mouse, to receive user input, and/or a software application to process the user input. By involving user interaction in the rule generating process, efficiency and accuracy may be improved.
  • the user interface 216 receives user input regarding key frame adjustment (240). For example, a user may view the training image sequence on a display device, such as the display device 112 (Fig. 1), and make adjustment to the key frames determined by the image sequence analyzing module 202. The user may add additional key frames or remove ones of the key frames determined by the image sequence analyzing module 202. As a result, the user interface 216 receives the user input regarding the key frame adjustment, and the key frame indexes 230 are interactively specified or adjusted by the user.
  • the user interface 216 receives user input regarding specification or adjustment of the vanishing point/vanishing lines/depth gradient planes in the training image (242).
  • the user may view the training image on the display device and provide scene structure information by specifying or adjusting the vanishing point/vanishing lines/depth gradient planes in the training image.
  • the user interface 216 receives the user input regarding the specification or adjustment of the vanishing point/vanishing lines/depth gradient planes in the training image, and the scene structure 232 is interactively specified or adjusted by the user.
  • the user interface 216 receives user input regarding object selection, object editing, specification of object depth and orientation, object labeling, and/or specification of a topological relationship for the objects in the training image (244).
  • the user may view the training image on the display device and provide semantic object information for the training image.
  • the user may select an object in the training image and edit a shape of the selected object.
  • the user may also specify depth and orientation of the selected object, and classify the selected object by labeling the selected object with a category.
  • the user may additionally specify the topological relationship for the objects in the training image.
  • the user interface 216 receives user input regarding the object selection, the object editing, the specification of the object depth and orientation, the object labeling, and/or the specification of the topological relationship for the objects in the training image, and the object data 234 is interactively specified or adjusted by the user.
  • the user interface 216 receives user input regarding an evaluation of the visual quality of the 3D video generated from the training image sequence (246). For example, the user may view the generated 3D video on the display device and check its visual quality. The user may determine the generated 3D video has relatively good or relatively poor visual quality, and specify/adjust the quality value 236. As a result, the user interface 216 receives the user input regarding the evaluation of the visual quality of the generated 3D video, and the quality value 236 is interactively specified or adjusted by the user.
  • the object evolving detection module 218 is a hardware device or software configured to detect for each of the objects in the training image sequence an object evolving path in the time domain, based on the internal data stored in the memory device 222, such as the key frame indexes 230, the scene structure 232, and/or the object data 234.
  • the object evolving detection module 218 is also configured to provide the detected object evolving path for user preview through the user interface 216.
  • the object evolving detection module 218 manages lifecycles of all the objects that have appeared in the scene, from their first appearance, being partially/wholly occluded, being moved or deformed, until final disappearance from the scene.
  • the DIBR module 220 is a hardware device or software configured to apply DIBR algorithms to generate the 3D video for the training image sequence, based on the internal data stored in the memory device 222, such as the key frame indexes 230, the scene structure 232, and/or the object data 234.
  • the DIBR module 220 is also configured to provide the generated 3D video for user preview through the user interface 216.
  • the DIBR algorithms may include 3D image warping.
  • 3D image warping changes a view direction and a viewpoint of an object, and transforms pixels in a reference image of the object to a destination view in a 3D environment based on depth levels of the pixels.
  • a function may be used to map pixels from the reference image to the destination view.
  • the DIBR module 220 may adjust and reconstruct the destination view to achieve a better effect.
  • the DIBR algorithms may also include plenoptic image modeling.
  • Plenoptic image modeling provides 3D scene information of an image visible from arbitrary viewpoints.
  • the 3D scene information may be obtained by a function based on a set of reference images with depth information. These reference images are warped and combined to form 3D representations of the scene from a particular viewpoint.
  • the DIBR module 220 may adjust and reconstruct the 3D scene information. Based on the 3D scene information, the DIBR module 220 may generate multi-view video frames for 3D displaying.
  • the rule database 224 may then generate rules for automatic 2D-to-3D conversion based on the internal data stored in the memory device 222, such as the key frame indexes 230, the scene structure 232, the object data 234, and/or the quality value 236.
  • the rule database 224 also stores the generated rules for further use.
  • Fig. 3 illustrates a block diagram of a 3D video convertor 300, according to an exemplary embodiment.
  • the 3D video convertor 300 may be the 3D video convertor 110-2 (Fig. 1).
  • the 3D video convertor 300 is configured to receive a 2D input image sequence including at least one 2D input image to be converted to 3D video, and to perform automatic 2D-to-3D conversion on the input image sequence based on the rules generated by the rule generator 200 (Fig. 2).
  • the 3D video convertor 300 may include a complexity control module 302, an object segmentation module 304, an object depth estimation module 306, and a DIBR module 308.
  • the 3D video convertor 300 may further include an interface 310, e.g., a plug-in interface, for connecting the 3D video convertor 300 to the rule generator 200 (Fig. 2) and for obtaining the rules therefrom.
  • the complexity control module 302 is a hardware device or software configured to specify a computational complexity for the 2D-to-3D conversion.
  • the user may configure or customize algorithms adopted by the object segmentation module 304 or the object depth estimation module 306.
  • the user may customize an algorithm according to a set of complexity levels, or provide a specific number of logical operation limitations as well as selected features in the algorithm, which may automatically generate a configuration for the algorithm that provides a balance between computational complexity and system performance.
  • the object segmentation module 304 is a hardware device or software configured to separate scene content in the input image into one or more objects, based on the rules obtained from the rule generator 200 (Fig. 2).
  • the object segmentation module 304 may detect the objects in the input image and segment out the detected objects, similar to the above description in connection with the object segmentation module 206 (Fig. 2).
  • the object depth estimation module 306 is a hardware device or software configured to estimate depths of the objects in the input image.
  • the object depth estimation module 306 may use a gravity-based depth generation algorithm to perform automatic 2D-to-3D conversion when, initially, the rules from the rule generator 200 (Fig. 2) are not sufficient for performing the conversion.
  • the object depth estimation module 306 may generate different depth gradient planes relative to a vanishing point and vanishing lines in the input image, thereby deriving a depth map for the input image, similar to the above description in connection with the object depth estimation module 212 (Fig. 2).
  • the DIBR module 308 is a hardware device or software configured to apply DIBR algorithms to generate 3D video for the input image sequence.
  • the DIBR algorithms may produce a 3D
  • the DIBR module 308 may utilize depth information of one or more neighboring images in the input image sequence.
  • the rule generator 200 (Fig. 2) and the 3D video convertor 300 (Fig. 3) may be connected in a loose connection manner or in a tight connection manner.
  • the 3D video convertor 300 is a well established, e.g., independent, system that may obtain the rules from the rule generator 200 to perform automatic 2D-to-3D conversion.
  • the 3D video convertor 300 may be software applications automatically developed by the rule generator 200 or may share one or more modules of the rule generator 200, to perform automatic 2D-to-3D conversion.
  • Fig. 4 illustrates a block diagram of a 3D video generator 400 configured to operate in the loose connection manner, according to an exemplary embodiment.
  • the 3D video generator 400 includes a rule generator 402, similar to the rule generator 200 (Fig. 2) described above, and a 3D video convertor 404, similar to the 3D video convertor 300 (Fig. 3) described above.
  • the 3D video convertor 404 is connected to the rule generator 402 through, e.g., a plug-in interface (not shown).
  • the object segmentation module and the object depth estimation module in the 3D video convertor 404 may obtain rules from the rule database in the rule generator 402, and perform automatic 2D-to-3D conversion on a 2D input image sequence based on the obtained rules.
  • Fig. 5 illustrates a block diagram of a 3D video generator 500 configured to operate in the tight connection manner, according to an exemplary embodiment.
  • the 3D video generator 500 includes a rule generator 502, similar to the rule generator 200 (Fig. 2) described above, and 3D video convertors 504-1 , 504-2, ... and 504-N, each similar to the 3D video convertor 300 (Fig. 3) described above.
  • the rule generator 502 further includes a software application, referred to herein as a
  • 3D-video-convertor generator 506 to generate additional software applications or computer programs for automatic 2D-to-3D conversion, referred to herein as 3D video convertors 504-1 , 504-2, ... and 504-N.
  • the 3D-video-convertor generator 506 may generate the 3D video convertors 504-1 , 504-2, ... and 504-N based on rules in the rule database in the rule generator 502.
  • Each of the 3D video convertors 504-1 , 504-2, ... and 504-N may perform automatic 2D-to-3D conversion for a different scenario in a 2D input image sequence.
  • Fig. 6 illustrates a flowchart of a method 600 for generating 3D video based on a 2D input image sequence including at least one input 2D image, according to an exemplary embodiment.
  • rules e.g., semantic rules, for automatic 2D-to-3D conversion are generated (602).
  • the 2D input image sequence may then be automatically converted to the 3D video based on the rules (604).
  • internal data is generated based on a 2D training image sequence including at least one 2D training image.
  • 3D video from the 2D training image sequence is then generated based on the internal data and depth-image based rendering. Quality evaluation for the generated 3D video may be further provided and an evaluation result is included in the internal data.
  • the rules are generated based on the internal data including the evaluation result, and are stored in a rule database.
  • the generating of the internal data may include at least one of: determining key frame indexes for the 2D training image sequence and storing the key frame indexes as the internal data; detecting a scene structure in the 2D training image and storing the scene structure as the internal data;
  • segmenting out objects in the 2D training image and storing a segmentation result as the internal data tracking the objects in the 2D training image sequence and storing a tracking result as the internal data; classifying the objects in the 2D training image and storing a classification result as the internal data; or
  • the rules are obtained, and objects in the 2D input image are segmented out based on the obtained rules.
  • Depth information for the 2D input image is also estimated based on the obtained rules.
  • 3D video is automatically generated from the 2D input image sequence based on the object segmentation and the depth information estimation.
  • the method disclosed herein may be implemented as a computer program product, i.e., a computer program tangibly embodied in an information carrier, e.g., a machine readable storage device, for execution by a data processing apparatus, e.g., a programmable processor, a computer, or multiple computers.
  • the computer program may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including stand-alone program, module, subroutine, or other unit suitable for use in a computing environment.
  • the computer program may be deployed to be executed on one computer, or on multiple computers.
  • a computer- readable medium including instructions, executable by a processor in a 3D video generating system, for performing the above described method for generating 3D video based on a 2D input image sequence.
  • a portion or all of the method disclosed herein may also be implemented by an application specific integrated circuit (ASIC), a field- programmable gate array (FPGA), a complex programmable logic device (CPLD), a printed circuit board (PCB), a digital signal processor (DSP), a combination of programmable logic components and programmable interconnects, a single central processing unit (CPU) chip, or a CPU chip combined on a motherboard.
  • ASIC application specific integrated circuit
  • FPGA field- programmable gate array
  • CPLD complex programmable logic device
  • PCB printed circuit board
  • DSP digital signal processor
  • CPU central processing unit
  • CPU central processing unit
  • CPU central processing unit

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Processing Or Creating Images (AREA)

Abstract

A system for generating three-dimensional (3D) video based on a two dimensional (2D) input image sequence including at least one 2D input image. The system includes: a rule generator configured to generate rules for 2D-to-3D conversion; and a 3D video converter coupled to the rule generator, and configured to obtain the rules from the rule generator and automatically convert the 2D input image sequence to the 3D video based on the obtained rules.

Description

SYSTEMS AND METHODS FOR THREE-DIMENSIONAL VIDEO
GENERATION
DESCRIPTION
Related Applications
[0001] This application is based upon and claims the benefit of priority from U.S. Provisional Patent Application No. 61/231 ,286, filed August 4, 2009, the entire contents of which are incorporated herein by reference.
Technical Field
[0002] This disclosure relates to systems and methods for three- dimensional video generation.
Background
[0003] Three-dimensional (3D) TV has recently been foreseen as part of a next wave of promising technologies for consumer electronics. Theoretically, 3D technologies incorporate a third dimension of depth into an image, which may provide a stereographic perception to a viewer of the image.
[0004] Currently, there are limited 3D video content sources in the market. Therefore, different methods to generate 3D video content have been studied and developed. One of the methods is to convert two-dimensional (2D) video content to 3D video content, which may fully employ existing 2D video content sources. However, some disclosed conversion techniques may not be ready for use due to their high computational complexity or unsatisfactory quality.
SUMMARY
[0005] According to a first aspect of the present disclosure, there is provided a system for generating three-dimensional (3D) video based on a two-dimensional (2D) input image sequence including at least one 2D input image, comprising: a rule generator configured to generate rules for 2D-to-3D conversion; and a 3D video converter coupled to the rule generator, and configured to obtain the rules from the rule generator and automatically convert the 2D input image sequence to the 3D video based on the obtained rules.
[0006] According to a second aspect of the present disclosure, there is provided a computer-implemented method for generating three-dimensional (3D) video based on a two-dimensional (2D) input image sequence including at least one 2D input image, comprising: generating rules for 2D-to-3D conversion; and automatically converting the 2D input image sequence to the 3D video based on the rules.
[0007] According to a third aspect of the present disclosure, there is provided a computer-readable medium including instructions, executable by a processor of a three-dimensional (3D) video generating system, for performing a method for generating 3D video based on a two-dimensional (2D) input image sequence including at least one 2D input image, the method comprising:
generating rules for 2D-to-3D conversion; and automatically converting the 2D input image sequence to the 3D video based on the rules.
[0008] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed.
BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0010] Fig. 1 illustrates a block diagram of a system for generating 3D video, according to an exemplary embodiment.
[0011] Fig. 2 illustrates a block diagram of a rule generator, according to an exemplary embodiment. [0012] Fig. 3 illustrates a block diagram of a 3D video convertor, according to an exemplary embodiment.
[0013] Fig. 4 illustrates a block diagram of a 3D video generator
configured to operate in a loose connection manner, according to an exemplary embodiment.
[0014] Fig. 5 illustrates a block diagram of a 3D video generator
configured to operate in a tight connection manner, according to an exemplary embodiment.
[0015] Fig. 6 illustrates a flowchart of a method for generating 3D video based on a 2D input image sequence, according to an exemplary embodiment.
DESCRIPTION OF THE EMBODIMENTS
[0016] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which the same numbers in different drawings represent the same or similar elements unless otherwise represented. The implementations set forth in the following description of exemplary embodiments do not represent all implementations consistent with the invention. Instead, they are merely examples of systems and methods
consistent with aspects related to the invention as recited in the appended claims.
[0017] Fig. 1 illustrates a block diagram of a system 100 for generating a three-dimensional (3D) video, according to an exemplary embodiment. The system 100 may include a two-dimensional (2D) video content source, such as a video storage medium 102 or a media server 104 connected with a network 106. The system 100 also includes a video device 108, a 3D video generator 110, and a display device 112.
[0018] In exemplary embodiments, the video storage medium 102 may be any medium for storing video content. For example, the video storage medium 102 may be provided as a compact disc (CD), a digital video disc (DVD), a hard disk, a magnetic tape, a flash memory card/drive, a volatile or non-volatile memory, a holographic data storage, or any other storage medium. The video storage medium 102 may be located within the video device 108, local to the video device 108, or remote from the video device 108.
[0019] In exemplary embodiments, the media server 104 may be a computer server that receives a request for 2D video content from the video device 108, processes the request, and provides 2D video content to the video device 108 through the network 106. For example, the media server 104 may be a web server, an enterprise server, or any other type of computer server. The media server 104 is configured to accept requests from the video device 108 based on, e.g., a hypertext transfer protocol (HTTP) or other protocols that may initiate a video session, and to serve the video device 108 with 2D video content.
[0020] In exemplary embodiments, the network 106 may include a wide area network (WAN), a local area network (LAN), a wireless network suitable for packet-type communications, such as Internet communications, a broadcast network, or any combination thereof. The network 106 is configured to distribute digital or non-digital video content.
[0021] In exemplary embodiments, the video device 108 is a hardware device such as a computer, a personal digital assistant (PDA), a mobile phone, a laptop, a desktop, a videocassette recorder (VCR), a laserdisc player, a DVD player, a blue ray disc player, or any electronic device configured to output 2D video, i.e., a 2D image sequence. The video device 108 may include software applications that allow the video device 108 to communicate with and receive 2D video content from, e.g., the video storage medium 102 or the media server 104. In addition, the video device 108 may, by means of included software
applications, transform the received 2D video content into digital format, if not already in digital format.
[0022] In exemplary embodiments, the 3D video generator 110 is configured to generate 3D video based on the 2D image sequence outputted by the video device 108. The generator 110 may be implemented as a hardware device that is either stand-alone or incorporated into the video device 108, or software applications installed on the video device 108, or a combination thereof. In addition, the 3D video generator 110 may include a processor to generate the 3D video.
[0023] In exemplary embodiments, the 3D video generator 110 may include a rule generator 110-1 configured to generate rules, e.g., semantic rules, for automatic 2D-to-3D conversion, based on a 2D training image sequence, and a 3D video convertor 110-2 configured to automatically convert a 2D input image sequence to 3D video based on the rules. In some exemplary embodiments, the rule generator 110-1 is also referred to as an interactive 2D-to-3D conversion environment (i3DV), and the 3D video convertor 110-2 is also referred to as an automatic 2D-to-3D conversion solution (a3DC). The rule generator 110-1 and the 3D video convertor 110-2 are described in detail below.
[0024] In exemplary embodiments, the display device 112 is configured to display images in the training image sequence or the 2D input image sequence, and to present the generated 3D video. For example, the display device 112 may be provided as a monitor, a projector, or any other video display device. The display device 112 may also be a part of the video device 108. The user may view the images in the training image sequence to provide inputs or feedbacks to the 3D video generator 110. The user may also watch the generated 3D video displayed by the display device 112. It is to be understood that devices shown in Fig. 1 are for illustrative purposes. Certain devices may be removed or combined, and additional devices may be added.
[0025] Fig. 2 illustrates a block diagram of a rule generator 200, according to an exemplary embodiment. For example, the rule generator 200 may be the rule generator 110-1 (Fig. 1). The rule generator 200 is configured to receive a 2D training image sequence including at least one 2D training image, and to generate 3D video from the training image sequence for preview. The rule generator 200 may further generate rules, e.g., semantic rules, for automatic 2D- to-3D conversion based on the training image sequence.
[0026] In exemplary embodiments, the rule generator 200 may include one or more of the following components: an image sequence analyzing module 202, a scene structure detection module 204, an object segmentation module 206, an object tracking module 208, an object classification module 210, and an object depth estimation module 212. The rule generator 200 may further include a quality evaluation module 214, a user interface 216, an object evolving detection module 218, a depth-image based rendering (DIBR) module 220, a memory device 222 for storing internal data, and a rule database 224.
[0027] In exemplary embodiments, the image sequence analyzing module 202 is a hardware device or software configured to perform analysis of the training image sequence, to determine one or more key frames in the training image sequence. For example, the image sequence analyzing module 202 may analyze the training image sequence to detect scene changes therein, and break the training image sequence into one or more chunks each starting with a key frame where a scene change occurs. The image sequence analyzing module 202 may then output key frame indexes 230 for the determined key frames, and store the key frame indexes 230 as the internal data in the memory device 222.
[0028] In exemplary embodiments, the scene structure detection module 204 is a hardware device or software configured to analyze the training image in the training image sequence to obtain a scene structure 232 for the training image, which is also stored as the internal data in the memory device 222. For example, the scene structure detection module 204 may perform the analysis based on linear perspective properties of the training image, by detecting a vanishing point and vanishing lines in the training image. A vanishing point represents a farthest point in a scene shown in a 2D image, and vanishing lines each represent a direction of increasing depth. The vanishing lines converge at the vanishing point in the 2D image. By detecting the vanishing point and the vanishing lines, the scene structure detection module 204 obtains a projected far-to-near direction in the training image and, thus, obtains the scene structure 232.
[0029] In exemplary embodiments, the object segmentation module 206, the object tracking module 208, the object classification module 210, and the object depth estimation module 212 are configured to generate object data 234 regarding semantic objects in the training image, and store the object data 234 as the internal data in the memory device 222.
[0030] More particularly, in exemplary embodiments, the object segmentation module 206 is a hardware device or software configured to separate scene content in the training image into one or more constituent parts, i.e., semantic objects, referred to hereafter as objects. For example, an object corresponds to a region or shape in the training image that represents an entity with certain semantic meaning, such as a tree, a lake, a house, etc. The object segmentation module 206 may detect the objects in the training image and segment out the detected objects, thereby generating the object data 234.
[0031] In one exemplary embodiment, in performing segmentation, the object segmentation module 206 may group pixels of the training image into different regions based on a homogeneous low-level feature, such as color, motion, or texture, each of the regions representing one of the objects. As a result, the object segmentation module 206 detects and segments out the objects in the training image. In addition, the object segmentation module 206 may cut one or more foreground objects out of the training image, and refine boundaries of the foreground objects.
[0032] In exemplary embodiments, the object tracking module 208 is a hardware device or software configured to use an object tracking algorithm to track the objects in the training image sequence. For example, the object tracking module 208 may detect how positions/sizes of the objects change with time in the training image, thereby generating the object data 234. In such manner, temporal coherence may be provided for generating the 3D video from the training image sequence.
[0033] In exemplary embodiments, the object classification module 210 is a hardware device or software configured to perform object classification for the objects in the training image based on a plurality of training data sets in an object database (not shown). Each of the training data sets includes a group of training images previously processed by the rule generator 200 and representing an object category, such as a building category, a grass category, a tree category, a cow category, a sky category, a face category, a car category, a bicycle category, etc. Based on a group of training images for an object category, the object classification module 210 may classify or identify the objects in the present training image, and assign category labels for the classified objects.
[0034] In exemplary embodiments, a training data set may initially include no training images or a relatively small number of training images. In such situation, the object classification module 210 operates together with the user interface 216 to perform user interactive object classification, as described below. As a result, the object classification module 210 learns additional object categories.
[0035] In exemplary embodiments, the object depth estimation module 212 is a hardware device or software configured to estimate depths of the objects in the training image. Typically, each pixel of the objects in a 2D image
corresponds to a depth. A depth of a pixel of an object represents a distance between a viewer and a part of the object corresponding to that pixel.
[0036] In exemplary embodiments, the object depth estimation module 212 may detect the vanishing point and the vanishing lines in the training image, and generate a depth map of the training image accordingly. For example, the object depth estimation module 212 may generate different depth gradient planes relative to the vanishing point and the vanishing lines. The object depth estimation module 212 then assigns depth levels to pixels of the training image according to the depth gradient planes, thereby generating the object data 234. The object depth estimation module 212 may additionally estimate orientation or thicknesses of the objects in the training image.
[0037] In exemplary embodiments, the quality evaluation module 214 is a hardware device or software configured to provide evaluation of visual quality of the 3D video generated from the training image sequence. For example, the quality evaluation module 214 may provide a quality value 236 representing the visual quality of the generated 3D video, and store the quality value 236 in the memory device 222. The quality value 236 is also part of the internal data in the memory device 222.
[0038] In exemplary embodiments, the user interface 216 is configured to receive user input and to output the 3D video generated from the training image sequence for preview, such that the rule generator 200 operates in a user interactive manner. For example, the user interface 216 may be implemented with a hardware device, such as a keyboard or a mouse, to receive user input, and/or a software application to process the user input. By involving user interaction in the rule generating process, efficiency and accuracy may be improved.
[0039] In one exemplary embodiment, the user interface 216 receives user input regarding key frame adjustment (240). For example, a user may view the training image sequence on a display device, such as the display device 112 (Fig. 1), and make adjustment to the key frames determined by the image sequence analyzing module 202. The user may add additional key frames or remove ones of the key frames determined by the image sequence analyzing module 202. As a result, the user interface 216 receives the user input regarding the key frame adjustment, and the key frame indexes 230 are interactively specified or adjusted by the user.
[0040] In one exemplary embodiment, the user interface 216 receives user input regarding specification or adjustment of the vanishing point/vanishing lines/depth gradient planes in the training image (242). For example, the user may view the training image on the display device and provide scene structure information by specifying or adjusting the vanishing point/vanishing lines/depth gradient planes in the training image. As a result, the user interface 216 receives the user input regarding the specification or adjustment of the vanishing point/vanishing lines/depth gradient planes in the training image, and the scene structure 232 is interactively specified or adjusted by the user.
[0041] In one exemplary embodiment, the user interface 216 receives user input regarding object selection, object editing, specification of object depth and orientation, object labeling, and/or specification of a topological relationship for the objects in the training image (244). For example, the user may view the training image on the display device and provide semantic object information for the training image. The user may select an object in the training image and edit a shape of the selected object. The user may also specify depth and orientation of the selected object, and classify the selected object by labeling the selected object with a category. The user may additionally specify the topological relationship for the objects in the training image. As a result, the user interface 216 receives user input regarding the object selection, the object editing, the specification of the object depth and orientation, the object labeling, and/or the specification of the topological relationship for the objects in the training image, and the object data 234 is interactively specified or adjusted by the user.
[0042] In one exemplary embodiment, the user interface 216 receives user input regarding an evaluation of the visual quality of the 3D video generated from the training image sequence (246). For example, the user may view the generated 3D video on the display device and check its visual quality. The user may determine the generated 3D video has relatively good or relatively poor visual quality, and specify/adjust the quality value 236. As a result, the user interface 216 receives the user input regarding the evaluation of the visual quality of the generated 3D video, and the quality value 236 is interactively specified or adjusted by the user.
[0043] In exemplary embodiments, the object evolving detection module 218 is a hardware device or software configured to detect for each of the objects in the training image sequence an object evolving path in the time domain, based on the internal data stored in the memory device 222, such as the key frame indexes 230, the scene structure 232, and/or the object data 234. The object evolving detection module 218 is also configured to provide the detected object evolving path for user preview through the user interface 216. In one exemplary embodiment, the object evolving detection module 218 manages lifecycles of all the objects that have appeared in the scene, from their first appearance, being partially/wholly occluded, being moved or deformed, until final disappearance from the scene.
[0044] In exemplary embodiments, the DIBR module 220 is a hardware device or software configured to apply DIBR algorithms to generate the 3D video for the training image sequence, based on the internal data stored in the memory device 222, such as the key frame indexes 230, the scene structure 232, and/or the object data 234. The DIBR module 220 is also configured to provide the generated 3D video for user preview through the user interface 216.
[0045] In exemplary embodiments, the DIBR algorithms may include 3D image warping. 3D image warping changes a view direction and a viewpoint of an object, and transforms pixels in a reference image of the object to a destination view in a 3D environment based on depth levels of the pixels. A function may be used to map pixels from the reference image to the destination view. The DIBR module 220 may adjust and reconstruct the destination view to achieve a better effect.
[0046] In exemplary embodiments, the DIBR algorithms may also include plenoptic image modeling. Plenoptic image modeling provides 3D scene information of an image visible from arbitrary viewpoints. The 3D scene information may be obtained by a function based on a set of reference images with depth information. These reference images are warped and combined to form 3D representations of the scene from a particular viewpoint. For an improved effect, the DIBR module 220 may adjust and reconstruct the 3D scene information. Based on the 3D scene information, the DIBR module 220 may generate multi-view video frames for 3D displaying.
[0047] In exemplary embodiments, the rule database 224 may then generate rules for automatic 2D-to-3D conversion based on the internal data stored in the memory device 222, such as the key frame indexes 230, the scene structure 232, the object data 234, and/or the quality value 236. The rule database 224 also stores the generated rules for further use.
[0048] Fig. 3 illustrates a block diagram of a 3D video convertor 300, according to an exemplary embodiment. For example, the 3D video convertor 300 may be the 3D video convertor 110-2 (Fig. 1). The 3D video convertor 300 is configured to receive a 2D input image sequence including at least one 2D input image to be converted to 3D video, and to perform automatic 2D-to-3D conversion on the input image sequence based on the rules generated by the rule generator 200 (Fig. 2). The 3D video convertor 300 may include a complexity control module 302, an object segmentation module 304, an object depth estimation module 306, and a DIBR module 308. The 3D video convertor 300 may further include an interface 310, e.g., a plug-in interface, for connecting the 3D video convertor 300 to the rule generator 200 (Fig. 2) and for obtaining the rules therefrom.
[0049] In exemplary embodiments, the complexity control module 302 is a hardware device or software configured to specify a computational complexity for the 2D-to-3D conversion. For example, through the complexity control module 302, the user may configure or customize algorithms adopted by the object segmentation module 304 or the object depth estimation module 306.
Depending on detailed implementation, the user may customize an algorithm according to a set of complexity levels, or provide a specific number of logical operation limitations as well as selected features in the algorithm, which may automatically generate a configuration for the algorithm that provides a balance between computational complexity and system performance.
[0050] In exemplary embodiments, the object segmentation module 304 is a hardware device or software configured to separate scene content in the input image into one or more objects, based on the rules obtained from the rule generator 200 (Fig. 2). The object segmentation module 304 may detect the objects in the input image and segment out the detected objects, similar to the above description in connection with the object segmentation module 206 (Fig. 2).
[0051] In exemplary embodiments, the object depth estimation module 306 is a hardware device or software configured to estimate depths of the objects in the input image. For example, the object depth estimation module 306 may use a gravity-based depth generation algorithm to perform automatic 2D-to-3D conversion when, initially, the rules from the rule generator 200 (Fig. 2) are not sufficient for performing the conversion. Also for example, the object depth estimation module 306 may generate different depth gradient planes relative to a vanishing point and vanishing lines in the input image, thereby deriving a depth map for the input image, similar to the above description in connection with the object depth estimation module 212 (Fig. 2).
[0052] In exemplary embodiments, the DIBR module 308 is a hardware device or software configured to apply DIBR algorithms to generate 3D video for the input image sequence. The DIBR algorithms may produce a 3D
representation based on the objection segmentation performed by the object segmentation module 304 and the depth estimation performed by the object depth estimation module 306. To achieve a better 3D effect, the DIBR module 308 may utilize depth information of one or more neighboring images in the input image sequence. [0053] In exemplary embodiments, the rule generator 200 (Fig. 2) and the 3D video convertor 300 (Fig. 3) may be connected in a loose connection manner or in a tight connection manner. In the loose connection manner, the 3D video convertor 300 is a well established, e.g., independent, system that may obtain the rules from the rule generator 200 to perform automatic 2D-to-3D conversion. In the tight connection manner, the 3D video convertor 300 may be software applications automatically developed by the rule generator 200 or may share one or more modules of the rule generator 200, to perform automatic 2D-to-3D conversion.
[0054] Fig. 4 illustrates a block diagram of a 3D video generator 400 configured to operate in the loose connection manner, according to an exemplary embodiment. For example, the 3D video generator 400 includes a rule generator 402, similar to the rule generator 200 (Fig. 2) described above, and a 3D video convertor 404, similar to the 3D video convertor 300 (Fig. 3) described above. The 3D video convertor 404 is connected to the rule generator 402 through, e.g., a plug-in interface (not shown). As a result, the object segmentation module and the object depth estimation module in the 3D video convertor 404 may obtain rules from the rule database in the rule generator 402, and perform automatic 2D-to-3D conversion on a 2D input image sequence based on the obtained rules.
[0055] Fig. 5 illustrates a block diagram of a 3D video generator 500 configured to operate in the tight connection manner, according to an exemplary embodiment. For example, the 3D video generator 500 includes a rule generator 502, similar to the rule generator 200 (Fig. 2) described above, and 3D video convertors 504-1 , 504-2, ... and 504-N, each similar to the 3D video convertor 300 (Fig. 3) described above. In the illustrated embodiment, the rule generator 502 further includes a software application, referred to herein as a
3D-video-convertor generator 506, to generate additional software applications or computer programs for automatic 2D-to-3D conversion, referred to herein as 3D video convertors 504-1 , 504-2, ... and 504-N. For example, the 3D-video-convertor generator 506 may generate the 3D video convertors 504-1 , 504-2, ... and 504-N based on rules in the rule database in the rule generator 502. Each of the 3D video convertors 504-1 , 504-2, ... and 504-N may perform automatic 2D-to-3D conversion for a different scenario in a 2D input image sequence.
[0056] Fig. 6 illustrates a flowchart of a method 600 for generating 3D video based on a 2D input image sequence including at least one input 2D image, according to an exemplary embodiment. Referring to Fig. 6, rules, e.g., semantic rules, for automatic 2D-to-3D conversion are generated (602). The 2D input image sequence may then be automatically converted to the 3D video based on the rules (604).
[0057] In exemplary embodiments, in generating the rules (602), internal data is generated based on a 2D training image sequence including at least one 2D training image. 3D video from the 2D training image sequence is then generated based on the internal data and depth-image based rendering. Quality evaluation for the generated 3D video may be further provided and an evaluation result is included in the internal data. The rules are generated based on the internal data including the evaluation result, and are stored in a rule database.
[0058] Further, the generating of the internal data may include at least one of: determining key frame indexes for the 2D training image sequence and storing the key frame indexes as the internal data; detecting a scene structure in the 2D training image and storing the scene structure as the internal data;
segmenting out objects in the 2D training image and storing a segmentation result as the internal data; tracking the objects in the 2D training image sequence and storing a tracking result as the internal data; classifying the objects in the 2D training image and storing a classification result as the internal data; or
estimating depth information for the 2D training image and storing the depth information as the internal data. Each of these operations may be facilitated with user interaction, similar to the above description in connection with the user interface 216 (Fig. 2).
[0059] In exemplary embodiments, in converting the 2D input image sequence (604), the rules are obtained, and objects in the 2D input image are segmented out based on the obtained rules. Depth information for the 2D input image is also estimated based on the obtained rules. 3D video is automatically generated from the 2D input image sequence based on the object segmentation and the depth information estimation.
[0060] The method disclosed herein may be implemented as a computer program product, i.e., a computer program tangibly embodied in an information carrier, e.g., a machine readable storage device, for execution by a data processing apparatus, e.g., a programmable processor, a computer, or multiple computers. The computer program may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including stand-alone program, module, subroutine, or other unit suitable for use in a computing environment. The computer program may be deployed to be executed on one computer, or on multiple computers.
[0061] In exemplary embodiments, there is also provided a computer- readable medium including instructions, executable by a processor in a 3D video generating system, for performing the above described method for generating 3D video based on a 2D input image sequence.
[0062] A portion or all of the method disclosed herein may also be implemented by an application specific integrated circuit (ASIC), a field- programmable gate array (FPGA), a complex programmable logic device (CPLD), a printed circuit board (PCB), a digital signal processor (DSP), a combination of programmable logic components and programmable interconnects, a single central processing unit (CPU) chip, or a CPU chip combined on a motherboard. For example, one or more of these hardware devices may be included in the 3D video generator 110 (Fig. 1). [0063] Other embodiments of the invention will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. The scope of the invention is intended to cover any variations, uses, or adaptations of the invention following the general principles thereof and including such departures from the present disclosure as come within known or customary practice in the art. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the invention being indicated by the following claims.
[0064] It will be appreciated that the present invention is not limited to the exact construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes can be made without departing from the scope thereof. It is intended that the scope of the invention only be limited by the appended claims.

Claims

WHAT IS CLAIMED IS:
1. A system for generating three-dimensional (3D) video based on a two-dimensional (2D) input image sequence including at least one 2D input image, comprising:
a rule generator configured to generate rules for 2D-to-3D conversion; and a 3D video converter coupled to the rule generator, and configured to obtain the rules from the rule generator and automatically convert the 2D input image sequence to the 3D video based on the obtained rules.
2. The system of claim 1 , wherein the rule generator generates the rules based on a 2D training image sequence including at least one 2D training image, the rule generator further comprising:
a memory device for storing internal data generated based on the 2D training image sequence;
a depth-image based rendering module coupled to the memory device, and configured to generate 3D video from the 2D training image sequence based on the internal data stored in the memory device;
a quality evaluation module coupled to the memory device, and configured to provide quality evaluation for the 3D video generated from the 2D training image sequence and store an evaluation result in the memory device, the evaluation result being part of the internal data; and
a rule database coupled to the memory device, and configured to generate the rules based on the internal data and store the generated rules.
3. The system of claim 2, wherein the rule generator further comprises at least one of: an image sequence analyzing module coupled to the memory device, and configured to determine key frame indexes for the 2D training image sequence and store the key frame indexes as the internal data in the memory device;
a scene structure detection module coupled to the memory device, and configured to detect a scene structure in the 2D training image and store the scene structure as the internal data in the memory device;
an object segmentation module coupled to the memory device, and configured to segment out objects in the 2D training image and store a segmentation result as the internal data in the memory device;
an object tracking module coupled to the memory device, and configured to track the objects in the 2D training image sequence and store a tracking result as the internal data in the memory device;
an object classification module coupled to the memory device, and configured to classify the objects in the 2D training image and store a
classification result as the internal data in the memory device; or
an object depth estimation module coupled to the memory device, and configured to estimate depth information for the 2D training image and store the depth information as the internal data in the memory device.
4. The system of claim 3, wherein the rule generator includes the image sequence analyzing module, the rule generator further comprising:
a user interface coupled to the memory device, to receive user input regarding specification or adjustment of the key frame indexes.
5. The system of claim 3, wherein the rule generator includes the scene structure detection module, the rule generator further comprising:
a user interface coupled to the memory device, to receive user input regarding specification or adjustment of the scene structure.
6. The system of claim 5, wherein the scene structure detection module is further configured to:
detect a vanishing point and one or more vanishing lines in the 2D training image, the vanishing point representing a farthest point in a scene shown in the 2D training image, and the one or more vanishing lines each representing a direction of increasing depth.
7. The system of claim 3, wherein the rule generator includes the object segmentation module, the rule generator further comprising:
a user interface coupled to the memory device, to receive user input regarding segmentation out of the objects in the 2D training image.
8. The system of claim 7, wherein the object segmentation module is further configured to:
group pixels of the 2D training image into different regions based on a homogeneous feature of the 2D training image.
9. The system of claim 3, wherein the rule generator includes the object classification module, the rule generator further comprising:
a user interface coupled to the memory device, to receive user input regarding classification of the objects in the 2D training image.
10. The system of claim 9, wherein the object classification module is further configured to:
classify the objects based on a plurality of training data sets in an object database, each of the plurality of training data sets representing one object category.
11. The system of claim 3, wherein the rule generator includes the object depth estimation module, the rule generator further comprising: a user interface coupled to the memory device, to receive user input regarding depths of the objects in the 2D training image.
12. The system of claim 2, wherein the rule generator further comprises:
an object evolving detection module coupled to the memory device, and configured to detect for an object in the 2D training image sequence an object evolving path in time domain.
13. The system of claim 2, wherein the rule generator further comprises:
a user interface coupled to the memory device, to receive user input regarding quality evaluation of the 3D video generated from the 2D training image sequence.
14. The system of claim 1 , wherein the 3D video convertor comprises:
an interface configured to obtain the rules from the rule generator;
an object segmentation module coupled to the interface, and configured to segment out objects in the 2D input image based on the obtained rules;
an object depth estimation module coupled to the interface, and
configured to estimate depth information for the 2D input image based on the obtained rules; and
a depth-image based rendering module coupled to the object
segmentation module and the object depth estimation module, and configured to generate the 3D video for the 2D input image sequence based on the object segmentation and the depth information estimation.
15. The system of claim 14, wherein the 3D video convertor further comprises: a complexity control module coupled to the object segmentation module and the object depth estimation module, and configured to customize algorithms used by the object segmentation module and the object depth estimation module.
16. A computer-implemented method for generating three-dimensional (3D) video based on a two-dimensional (2D) input image sequence including at least one 2D input image, comprising:
generating rules for 2D-to-3D conversion; and
automatically converting the 2D input image sequence to the 3D video based on the rules.
17. The method of claim 16, wherein generating the rules comprises:
generating internal data based on a 2D training image sequence including at least one 2D training image;
generating 3D video from the 2D training image sequence based on the internal data and depth-image based rendering;
providing quality evaluation for the 3D video generated from the 2D training image sequence and including an evaluation result in the internal data; generating the rules based on the internal data including the evaluation result; and
storing the generated rules.
18. The method of claim 17, wherein the generating of the internal data comprises at least one of:
determining key frame indexes for the 2D training image sequence and storing the key frame indexes as the internal data;
detecting a scene structure in the 2D training image and storing the scene structure as the internal data;
segmenting out objects in the 2D training image and storing a
segmentation result as the internal data; tracking the objects in the 2D training image sequence and storing a tracking result as the internal data;
classifying the objects in the 2D training image and storing a classification result as the internal data; or
estimating depth information for the 2D training image and storing the depth information as the internal data.
19. The method of claim 18, wherein the generating of the internal data includes the determining of the key frame indexes, generating the rules further comprising:
receiving user input regarding specification or adjustment of the key frame indexes.
20. The method of claim 18, wherein the generating of the internal data includes the detecting of the scene structure, generating the rules further comprising:
receiving user input regarding specification or adjustment of the scene structure.
21. The method of claim 20, wherein the detecting of the scene structure comprises:
detecting a vanishing point and one or more vanishing lines in the 2D training image, the vanishing point representing a farthest point in a scene shown in the 2D training image, and the one or more vanishing lines each representing a direction of increasing depth.
22. The method of claim 18, wherein the generating of the internal data includes the segmenting out of the objects, generating the rules further comprising: receiving user input regarding segmentation out of the objects in the 2D training image.
23. The method of claim 22, wherein the segmenting out of the objects comprises:
grouping pixels of the 2D training image into different regions based on a homogeneous feature of the 2D training image.
24. The method of claim 18, wherein the generating of the internal data includes the classifying of the objects, generating the rules further comprising: receiving user input regarding classification of the objects in the 2D training image.
25. The method of claim 24, wherein the classifying of the objects comprises: classifying the objects based on a plurality of training data sets in an object database, each of the plurality of training data sets representing one object category.
26. The method of claim 18, wherein the generating of the internal data includes the estimating of the depth information, generating the rules further comprising:
receiving user input regarding depths of the objects in the 2D training image.
27. The method of claim 17, wherein generating the rules further comprises: detecting for an object in the 2D training image sequence an object evolving path in time domain.
28. The method of claim 17, wherein generating the rules further comprises: receiving user input regarding quality evaluation of the 3D video generated from the 2D training image sequence.
29. The method of claim 16, wherein the automatically converting comprises: obtaining the rules;
segmenting out objects in the 2D input image based on the obtained rules; estimating depth information for the 2D input image based on the obtained rules; and
generating the 3D video for the 2D input image sequence based on the object segmentation and the depth information estimation.
30. A computer-readable medium including instructions, executable by a processor of a three-dimensional (3D) video generating system, for performing a method for generating 3D video based on a two-dimensional (2D) input image sequence including at least one 2D input image, the method comprising:
generating rules for 2D-to-3D conversion; and
automatically converting the 2D input image sequence to the 3D video based on the rules.
EP10807016.0A 2009-08-04 2010-08-03 Systems and methods for three-dimensional video generation Withdrawn EP2462536A4 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US23128609P 2009-08-04 2009-08-04
PCT/US2010/044227 WO2011017308A1 (en) 2009-08-04 2010-08-03 Systems and methods for three-dimensional video generation

Publications (2)

Publication Number Publication Date
EP2462536A1 true EP2462536A1 (en) 2012-06-13
EP2462536A4 EP2462536A4 (en) 2014-10-22

Family

ID=43544623

Family Applications (1)

Application Number Title Priority Date Filing Date
EP10807016.0A Withdrawn EP2462536A4 (en) 2009-08-04 2010-08-03 Systems and methods for three-dimensional video generation

Country Status (2)

Country Link
EP (1) EP2462536A4 (en)
WO (1) WO2011017308A1 (en)

Families Citing this family (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110025830A1 (en) 2009-07-31 2011-02-03 3Dmedia Corporation Methods, systems, and computer-readable storage media for generating stereoscopic content via depth map creation
WO2011014419A1 (en) 2009-07-31 2011-02-03 3Dmedia Corporation Methods, systems, and computer-readable storage media for creating three-dimensional (3d) images of a scene
US9380292B2 (en) 2009-07-31 2016-06-28 3Dmedia Corporation Methods, systems, and computer-readable storage media for generating three-dimensional (3D) images of a scene
US9344701B2 (en) 2010-07-23 2016-05-17 3Dmedia Corporation Methods, systems, and computer-readable storage media for identifying a rough depth map in a scene and for determining a stereo-base distance for three-dimensional (3D) content creation
WO2012061549A2 (en) * 2010-11-03 2012-05-10 3Dmedia Corporation Methods, systems, and computer program products for creating three-dimensional video sequences
US10200671B2 (en) 2010-12-27 2019-02-05 3Dmedia Corporation Primary and auxiliary image capture devices for image processing and related methods
US8274552B2 (en) 2010-12-27 2012-09-25 3Dmedia Corporation Primary and auxiliary image capture devices for image processing and related methods
EP2525581A3 (en) * 2011-05-17 2013-10-23 Samsung Electronics Co., Ltd. Apparatus and Method for Converting 2D Content into 3D Content, and Computer-Readable Storage Medium Thereof
KR101862543B1 (en) * 2011-09-08 2018-07-06 삼성전자 주식회사 Apparatus, meethod for generating depth information and computer-readable storage medium thereof
US10614292B2 (en) * 2018-02-06 2020-04-07 Kneron Inc. Low-power face identification method capable of controlling power adaptively

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2006034256A2 (en) * 2004-09-17 2006-03-30 Cyberextruder.Com, Inc. System, method, and apparatus for generating a three-dimensional representation from one or more two-dimensional images
EP2033164B1 (en) * 2006-06-23 2015-10-07 Imax Corporation Methods and systems for converting 2d motion pictures for stereoscopic 3d exhibition
US20080205791A1 (en) * 2006-11-13 2008-08-28 Ramot At Tel-Aviv University Ltd. Methods and systems for use in 3d video generation, storage and compression
US8655052B2 (en) * 2007-01-26 2014-02-18 Intellectual Discovery Co., Ltd. Methodology for 3D scene reconstruction from 2D image sequences

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
HARMAN P ET AL: "RAPID 2D TO 3D CONVERSION", PROCEEDINGS OF SPIE, S P I E - INTERNATIONAL SOCIETY FOR OPTICAL ENGINEERING, US, vol. 4660, 21 January 2002 (2002-01-21), pages 78-86, XP008021652, ISSN: 0277-786X, DOI: 10.1117/12.468020 *
See also references of WO2011017308A1 *

Also Published As

Publication number Publication date
WO2011017308A1 (en) 2011-02-10
EP2462536A4 (en) 2014-10-22

Similar Documents

Publication Publication Date Title
EP2462536A1 (en) Systems and methods for three-dimensional video generation
EP2481023B1 (en) 2d to 3d video conversion
US10529086B2 (en) Three-dimensional (3D) reconstructions of dynamic scenes using a reconfigurable hybrid imaging system
US9846960B2 (en) Automated camera array calibration
JP4938093B2 (en) System and method for region classification of 2D images for 2D-TO-3D conversion
US9414048B2 (en) Automatic 2D-to-stereoscopic video conversion
Karsch et al. Depth extraction from video using non-parametric sampling
Karsch et al. Depth transfer: Depth extraction from video using non-parametric sampling
US8861836B2 (en) Methods and systems for 2D to 3D conversion from a portrait image
US9237330B2 (en) Forming a stereoscopic video
JP7664891B2 (en) Method for generating layered depth data of a scene - Patents.com
US8897542B2 (en) Depth map generation based on soft classification
US20100220920A1 (en) Method, apparatus and system for processing depth-related information
US20120068996A1 (en) Safe mode transition in 3d content rendering
US9661307B1 (en) Depth map generation using motion cues for conversion of monoscopic visual content to stereoscopic 3D
CN117475105A (en) Open world three-dimensional scene reconstruction and perception method based on monocular image
US20150030233A1 (en) System and Method for Determining a Depth Map Sequence for a Two-Dimensional Video Sequence
Jang et al. Efficient disparity map estimation using occlusion handling for various 3D multimedia applications
WO2011017310A1 (en) Systems and methods for three-dimensional video generation
WO2008152607A1 (en) Method, apparatus, system and computer program product for depth-related information propagation
Jain et al. Efficient stereo-to-multiview synthesis
KR101511315B1 (en) Method and system for creating dynamic floating window for stereoscopic contents
Lee et al. Estimating scene-oriented pseudo depth with pictorial depth cues
Xu et al. Depth estimation algorithm based on data-driven approach and depth cues for stereo conversion in three-dimensional displays
Xu et al. Comprehensive depth estimation algorithm for efficient stereoscopic content creation in three-dimensional video systems

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

17P Request for examination filed

Effective date: 20120301

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO SE SI SK SM TR

DAX Request for extension of the european patent (deleted)
A4 Supplementary search report drawn up and despatched

Effective date: 20140919

RIC1 Information provided on ipc code assigned before grant

Ipc: G06K 9/00 20060101AFI20140915BHEP

Ipc: G06T 7/00 20060101ALI20140915BHEP

17Q First examination report despatched

Effective date: 20141002

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN

18D Application deemed to be withdrawn

Effective date: 20160628