WO2010124926A1 - Method and a device for generating and processing depth data - Google Patents
Method and a device for generating and processing depth data Download PDFInfo
- Publication number
- WO2010124926A1 WO2010124926A1 PCT/EP2010/054773 EP2010054773W WO2010124926A1 WO 2010124926 A1 WO2010124926 A1 WO 2010124926A1 EP 2010054773 W EP2010054773 W EP 2010054773W WO 2010124926 A1 WO2010124926 A1 WO 2010124926A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- depth data
- control signal
- generating
- image
- processing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/50—Depth or shape recovery
- G06T7/55—Depth or shape recovery from multiple images
- G06T7/593—Depth or shape recovery from multiple images from stereo images
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/30—Determination of transform parameters for the alignment of images, i.e. image registration
- G06T7/32—Determination of transform parameters for the alignment of images, i.e. image registration using correlation-based methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10004—Still image; Photographic image
- G06T2207/10012—Stereo images
Definitions
- the invention is related to a method and a device for generating and processing depth data.
- Said depth data is associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image.
- An image is a two dimensional (2D) projection on an image plane located at a view point wherein the view point has a location and an orientation.
- a second image may be used for preserving the depth information, implicitly, and for allowing a viewer to vary the orientation within a small range. That is, a scene is stereoscopically depicted by help of two images taken from two slightly different view points, like the eyes of a human, the difference depending on the focal length used. One of the images is provided to the right eye and the other is provided to the left eye. This allows a viewer for experiencing the depth of the scene and to vary the orientation of the view point within said small range. But the location of the viewpoint remains fixed.
- Depth data is an explicit representation of depth information, for instance in form of a depth map, in form of a three-dimensional model of structures comprised in a depicted scene or in form of a set of layered images or depth planes.
- views of a scene from arbitrary located and/or orientated viewpoints may be generated automatically in response to a viewer' s input determining a desired viewpoint.
- Depth data may be collected using dedicated depth scanning devices based on laser ranging, ultrasonic ranging or the like. Another way of generating depth data is based on a mapping between images without using dedicated depth sensors. Such mapping -also known as disparity mapping- is based on the fact that relative dislocation between a place of depiction of a feature in one of the images and a place of depiction of the same feature in the other of the images is reciprocally proportional to the distance of said feature to the images.
- Depth data may be corrupted in different ways. For instance, due to the reciprocal relationship minor errors in determining disparities of closed by features may result in severely erroneous depth data. Other sources of corruption are quantization during encoding and packet loss during transmission.
- the inventors recognized that such corruptions more or less impede the viewer' s viewing experience in dependency on the degree of attention the viewer spends on image aspects generated using said corrupted depth data.
- the invention proposes a method for generating and processing depth data associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image, said method comprising the features of independent claim 1.
- a corresponding device is proposed which comprises the features of independent claim 6.
- Said method comprises the steps of generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data, generating depth data using said mapping and processing the depth data wherein generating and/or processing of depth data is dependent on said control signal.
- Fig. 1 depicts an exemplarily flow chart of generating and processing of a depth map which may be used in single frame plus depth (2D+z) or a multiview plus depth (MVD) encoding, for instance, and
- Fig. 2 depicts an exemplarily flow chart of generating and processing of a mesh-based three dimensional model, for instance.
- FIG. 1 An exemplarily flow chart of generating and processing of an integrated depth map DMn according to an exemplary embodiment of the inventive principle is depicted in Fig. 1.
- the integrated depth map DMn is based on camera parameters CP and generated from disparity maps Dn-l/n and Dn+l/n which are generated from image In and image In-I, respectively, from image In and image In+1.
- Disparity maps Dn-l/n and Dn+l/n are not only used for generating the integrated depth map DMn but also for warping or compensating saliency maps Sn+1 and Sn-I which are saliency maps generated from image In+1 and image In-I. Compensation or warping results in compensated saliency maps CSn+1 and CSn-I which are aligned with saliency map Sn generated from image In and therefore can be fused with or integrated into said saliency map Sn generated from image In which results in an integrated saliency map ISMn.
- FIG. 2 Another exemplarily flow chart of the inventive principle is depicted in Fig. 2. From a set SoV of two or more views another set SoSM of a corresponding number of saliency maps is generated. The generated set SoSM of saliency maps as well as the set SoV of views are used for extracting one or more three dimensional models TdVM of one or more objects or structures at least partly depicted in the views of the set SoV of views .
- the modelling is based on depth information determined using two or more views of said set SoV of views and controlled by saliency information determined using corresponding saliency maps of said set SoSM of saliency maps. That is, the more salient a surface region of a modelled object or structure is, the more detailed said surface region is modelled. For instance, in a mesh model the number of facets, nodes and/or edges is increased such that the size of facets in salient areas is reduced and approximation of the surface of the modelled object or structure is more precise in salient areas.
- modelling is based on disparities of pairs of image regions of which one is comprised in a first view of the modelled object or structure and one is comprised in a second view of the modelled object or structure wherein said second view depicts the object or scene from a different viewpoint than said first view and the images regions of each pair are determined as depicting a same aspect of the object or structure to-be-modelled.
- the extracted three dimensional models TdM may then be encoded into an encoded bitstream EBS wherein a region-of- interest based encoding is used which is based on one or more three dimensional saliency models TdSM of the modelled objects or structures.
- the claimed method may be implemented on a processing device comprised in a notebook, a desktop computer, a server or any other computing device.
- the computing device may further comprise a storage medium like a hard disk, a BluRay-disk or a flash memory carrying image data of images .
- the processing device may be operated as a saliency value determining device which uses image data for determining corresponding saliency values.
- the processing device may further be operated as a control signal generator which generates control signals in dependency on determined saliency values.
- processing device may further be operated as a mapping device which maps first image data comprised in a first image onto second image data comprised in at least a second image.
- the processing device may also be operated as a depth data generator which generates depth data in dependency on determined mappings and the generated control signal and/or the processing device may also be operated as a depth data processor which processes depth data in dependency on the generated control signal.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
- Testing, Inspecting, Measuring Of Stereoscopic Televisions And Televisions (AREA)
Abstract
The invention proposes a method for generating and processing depth data associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image. Said method comprising the steps of generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data, generating depth data using said mapping and processing the depth data wherein generating and/or processing of depth data is dependent on said control signal. This allows for more accurate generation and/or processing of the depth data in regions which are visually attracting and thus more likely in the focus of a viewer's attention.
Description
Method and a Device for generating and processing depth data
Technical Field
The invention is related to a method and a device for generating and processing depth data. Said depth data is associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image.
Background
An image is a two dimensional (2D) projection on an image plane located at a view point wherein the view point has a location and an orientation.
Due to the projection, some or all depth information comprised in a scene depicted by an image is lost.
It is known that a second image may be used for preserving the depth information, implicitly, and for allowing a viewer to vary the orientation within a small range. That is, a scene is stereoscopically depicted by help of two images taken from two slightly different view points, like the eyes of a human, the difference depending on the focal length used. One of the images is provided to the right eye and the other is provided to the left eye. This allows a viewer for experiencing the depth of the scene and to vary the orientation of the view point within said small range. But the location of the viewpoint remains fixed.
There is a trend in the art towards flexible viewpoint applications like free viewpoint TV (FTV) which allow a viewer to freely vary not only the orientation of the viewpoint but also the location.
For making the location of the viewpoint freely variable as well, depth data has to be available. Depth data is an explicit representation of depth information, for instance in form of a depth map, in form of a three-dimensional model of structures comprised in a depicted scene or in form of a set of layered images or depth planes.
Using one or more images and depth data, views of a scene from arbitrary located and/or orientated viewpoints may be generated automatically in response to a viewer' s input determining a desired viewpoint.
Depth data may be collected using dedicated depth scanning devices based on laser ranging, ultrasonic ranging or the like. Another way of generating depth data is based on a mapping between images without using dedicated depth sensors. Such mapping -also known as disparity mapping- is based on the fact that relative dislocation between a place of depiction of a feature in one of the images and a place of depiction of the same feature in the other of the images is reciprocally proportional to the distance of said feature to the images.
Depth data may be corrupted in different ways. For instance, due to the reciprocal relationship minor errors in determining disparities of closed by features may result in severely erroneous depth data. Other sources of corruption are quantization during encoding and packet loss during transmission.
Invention
The inventors recognized that such corruptions more or less impede the viewer' s viewing experience in dependency on the degree of attention the viewer spends on image aspects generated using said corrupted depth data.
Therefore, the invention proposes a method for generating and processing depth data associated with a mapping of
first image data comprised in a first image onto second image data comprised in at least a second image, said method comprising the features of independent claim 1. A corresponding device is proposed which comprises the features of independent claim 6.
Said method comprises the steps of generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data, generating depth data using said mapping and processing the depth data wherein generating and/or processing of depth data is dependent on said control signal.
Making generation and/or processing of depth data dependent on saliency values which are determined for that image data whose mapping is associated with said depth data allows for more accurate generation and/or processing of the depth data in regions which are visually attracting and thus more likely in the focus of a viewer's attention.
Further advantageous embodiments of the method are characterized by the features of one of the claims depending on claim 1 and further advantageous embodiments of the device are characterized by the features of one of the claims depending on claim 6.
Drawings
Exemplary embodiments of the invention are illustrated in the drawings and are explained in more detail in the following description.
In the figures:
Fig. 1 depicts an exemplarily flow chart of generating and processing of a depth map which may be used in single frame plus depth (2D+z) or a multiview plus depth (MVD) encoding, for instance, and
Fig. 2 depicts an exemplarily flow chart of generating and processing of a mesh-based three dimensional model, for instance.
Exemplary embodiments
An exemplarily flow chart of generating and processing of an integrated depth map DMn according to an exemplary embodiment of the inventive principle is depicted in Fig. 1.
The integrated depth map DMn is based on camera parameters CP and generated from disparity maps Dn-l/n and Dn+l/n which are generated from image In and image In-I, respectively, from image In and image In+1.
Disparity maps Dn-l/n and Dn+l/n are not only used for generating the integrated depth map DMn but also for warping or compensating saliency maps Sn+1 and Sn-I which are saliency maps generated from image In+1 and image In-I. Compensation or warping results in compensated saliency maps CSn+1 and CSn-I which are aligned with saliency map Sn generated from image In and therefore can be fused with or integrated into said saliency map Sn generated from image In which results in an integrated saliency map ISMn.
Thus, there are the integrated depth map DMn for image n as well as an integrated saliency map which allows for a region-of-interest based encoding of the integrated depth map DMn into an encoded bitstream EBS.
Another exemplarily flow chart of the inventive principle is depicted in Fig. 2.
From a set SoV of two or more views another set SoSM of a corresponding number of saliency maps is generated. The generated set SoSM of saliency maps as well as the set SoV of views are used for extracting one or more three dimensional models TdVM of one or more objects or structures at least partly depicted in the views of the set SoV of views .
The modelling is based on depth information determined using two or more views of said set SoV of views and controlled by saliency information determined using corresponding saliency maps of said set SoSM of saliency maps. That is, the more salient a surface region of a modelled object or structure is, the more detailed said surface region is modelled. For instance, in a mesh model the number of facets, nodes and/or edges is increased such that the size of facets in salient areas is reduced and approximation of the surface of the modelled object or structure is more precise in salient areas.
Commonly, modelling is based on disparities of pairs of image regions of which one is comprised in a first view of the modelled object or structure and one is comprised in a second view of the modelled object or structure wherein said second view depicts the object or scene from a different viewpoint than said first view and the images regions of each pair are determined as depicting a same aspect of the object or structure to-be-modelled.
The extracted three dimensional models TdM may then be encoded into an encoded bitstream EBS wherein a region-of- interest based encoding is used which is based on one or more three dimensional saliency models TdSM of the modelled objects or structures.
The claimed method may be implemented on a processing device comprised in a notebook, a desktop computer, a server or any other computing device. The computing device
may further comprise a storage medium like a hard disk, a BluRay-disk or a flash memory carrying image data of images .
The processing device may be operated as a saliency value determining device which uses image data for determining corresponding saliency values.
The processing device may further be operated as a control signal generator which generates control signals in dependency on determined saliency values.
And, the processing device may further be operated as a mapping device which maps first image data comprised in a first image onto second image data comprised in at least a second image.
The processing device may also be operated as a depth data generator which generates depth data in dependency on determined mappings and the generated control signal and/or the processing device may also be operated as a depth data processor which processes depth data in dependency on the generated control signal.
The exemplary embodiments of the invention are described for illustrative purposes only and shall not be construed as limiting the extent of projection as requested which is solely defined by the independent claims.
Claims
1. Method for generating and processing depth data associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image, wherein processing of depth data comprises encoding of the depth data and said method comprises the steps of
- generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data,
- generating depth data using said mapping and
- processing the depth data wherein
- encoding of depth data is dependent on said control signal .
2. Method according to claim 1, wherein generating said depth data is also dependent on said control signal and comprises generating one or more nodes of a mesh model and assigning one or more further saliency values to the generated nodes with the number of generated nodes and/or values of the further saliency values being dependent on said control signal.
3. Method according to one of the preceding claims, wherein processing of the depth data further comprises lossy compression of the depth data and said control signal is further used for controlling a degree of compression loss.
4. Method according to one of the preceding claims, wherein encoding of the depth data comprises ordered encoding of the depth data together with further depth data in a code sequence and said control signal is used for determining a rank in the encoding order.
5. Method according to one of the preceding claims, wherein the control signal is generated dependent on a sum of said first saliency value and said at least a second saliency value.
6. Device for generating and processing depth data associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image, wherein processing comprises encoding of the depth data and said device comprises
- means for generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data,
— means for generating depth data using said mapping,
- means for processing depth data and
— a controller adapted for using said control signal for controlling said encoding of the depth data.
7. Device according to claim 6, wherein said means for generating depth data is adapted for using said control signal for controlling said generating one or more nodes of a mesh model and assigning one or more further saliency values to the generated nodes with the controller being adapted for controlling the number of generated nodes and/or values of the further saliency values in dependence on said control signal.
8. Device according to one of the claims 6 or 7, wherein processing of the depth data further comprises lossy compression of the depth data and the controller is further adapted for using said control signal for controlling a degree of compression loss.
9. Device according to one of the claims 6-8, wherein encoding of the depth data comprises ordered encoding of the depth data together with further depth data in a code sequence and the controller is adapted for using said control signal for determining a rank in an encoding order,
10. Device according to one of the claims 6-9, wherein means for generating a control signal is adapted for generating the control signal dependent on a sum of said first saliency value and said at least a second saliency value .
11. Device according to claim 10 or method according to claim 5, wherein said sum is a weighted sum.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP09305372.6 | 2009-04-29 | ||
| EP09305372A EP2247114A1 (en) | 2009-04-29 | 2009-04-29 | Method and a device for generating and processing depth data |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2010124926A1 true WO2010124926A1 (en) | 2010-11-04 |
Family
ID=41057425
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2010/054773 Ceased WO2010124926A1 (en) | 2009-04-29 | 2010-04-12 | Method and a device for generating and processing depth data |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP2247114A1 (en) |
| WO (1) | WO2010124926A1 (en) |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP1544792A1 (en) * | 2003-12-18 | 2005-06-22 | Thomson Licensing S.A. | Device and method for creating a saliency map of an image |
-
2009
- 2009-04-29 EP EP09305372A patent/EP2247114A1/en not_active Withdrawn
-
2010
- 2010-04-12 WO PCT/EP2010/054773 patent/WO2010124926A1/en not_active Ceased
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP1544792A1 (en) * | 2003-12-18 | 2005-06-22 | Thomson Licensing S.A. | Device and method for creating a saliency map of an image |
Non-Patent Citations (3)
| Title |
|---|
| FELDMANN I ET AL: "REAL-TIME SEGMENTATION FOR ADVANCES DISPARITY ESTIMATION IN IMMERSIVE VIDEOCONFERENCE APPLICATIONS", JOURNAL OF WSCG, VACLAV SKALA-UNION AGENCY, PILZEN, CZ, vol. 10, no. 1, 1 January 2002 (2002-01-01), pages 171 - 178, XP008017665, ISSN: 1213-6972 * |
| LE MEUR O ET AL: "Efficient Saliency-Based Repurposing Method", IMAGE PROCESSING, 2006 IEEE INTERNATIONAL CONFERENCE ON, IEEE, PI, 1 October 2006 (2006-10-01), pages 421 - 424, XP031048663, ISBN: 978-1-4244-0480-3 * |
| YONGQING ZENG: "Perceptual segmentation algorithm and its application to image coding", IMAGE PROCESSING, 1999. ICIP 99. PROCEEDINGS. 1999 INTERNATIONAL CONFE RENCE ON KOBE, JAPAN 24-28 OCT. 1999, PISCATAWAY, NJ, USA,IEEE, US, vol. 2, 24 October 1999 (1999-10-24), pages 820 - 824, XP010369077, ISBN: 978-0-7803-5467-8 * |
Also Published As
| Publication number | Publication date |
|---|---|
| EP2247114A1 (en) | 2010-11-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CA2668941C (en) | System and method for model fitting and registration of objects for 2d-to-3d conversion | |
| JP5561781B2 (en) | Method and system for converting 2D image data into stereoscopic image data | |
| JP5887267B2 (en) | 3D image interpolation apparatus, 3D imaging apparatus, and 3D image interpolation method | |
| KR101907945B1 (en) | Displaying graphics in multi-view scenes | |
| CA2723627C (en) | System and method for measuring potential eyestrain of stereoscopic motion pictures | |
| JP7664891B2 (en) | Method for generating layered depth data of a scene - Patents.com | |
| CN102404592B (en) | Image processing device and method, and stereoscopic image display device | |
| CN114830676B (en) | Video processing devices and manifest files for video streaming | |
| CN1925627A (en) | Apparatus for controlling depth of 3D picture and method therefor | |
| TWI523488B (en) | A method of processing parallax information comprised in a signal | |
| JPWO2012035783A1 (en) | 3D image creation apparatus and 3D image creation method | |
| KR20160135660A (en) | Method and apparatus for providing 3-dimension image to head mount display | |
| WO2011030399A1 (en) | Image processing method and apparatus | |
| JP5178538B2 (en) | Method for determining depth map from image, apparatus for determining depth map | |
| Knorr et al. | An image-based rendering (ibr) approach for realistic stereo view synthesis of tv broadcast based on structure from motion | |
| CN115997379A (en) | Image FOV Restoration for Stereoscopic Rendering | |
| JP2009530701A5 (en) | ||
| WO2019092398A1 (en) | Stereoscopic image data compression | |
| WO2008152607A1 (en) | Method, apparatus, system and computer program product for depth-related information propagation | |
| EP2247114A1 (en) | Method and a device for generating and processing depth data | |
| KR101634225B1 (en) | Device and Method for Multi-view image Calibration | |
| Gurrieri et al. | Efficient panoramic sampling of real-world environments for image-based stereoscopic telepresence | |
| US12375633B2 (en) | Low complexity multilayer images with depth | |
| CN101566784B (en) | Method for establishing depth of field data for three-dimensional image and system thereof | |
| CN110740308B (en) | Time consistent reliability delivery system |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 10713923 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 10713923 Country of ref document: EP Kind code of ref document: A1 |