WO2010124926A1 - Method and a device for generating and processing depth data - Google Patents

Method and a device for generating and processing depth data Download PDF

Info

Publication number
WO2010124926A1
WO2010124926A1 PCT/EP2010/054773 EP2010054773W WO2010124926A1 WO 2010124926 A1 WO2010124926 A1 WO 2010124926A1 EP 2010054773 W EP2010054773 W EP 2010054773W WO 2010124926 A1 WO2010124926 A1 WO 2010124926A1
Authority
WO
WIPO (PCT)
Prior art keywords
depth data
control signal
generating
image
processing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/EP2010/054773
Other languages
French (fr)
Inventor
Guillaume Boisson
Paul Kerbiriou
Olivier Le Meur
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Thomson Licensing SAS
Original Assignee
Thomson Licensing SAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Thomson Licensing SAS filed Critical Thomson Licensing SAS
Publication of WO2010124926A1 publication Critical patent/WO2010124926A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/50Depth or shape recovery
    • G06T7/55Depth or shape recovery from multiple images
    • G06T7/593Depth or shape recovery from multiple images from stereo images
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/30Determination of transform parameters for the alignment of images, i.e. image registration
    • G06T7/32Determination of transform parameters for the alignment of images, i.e. image registration using correlation-based methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10004Still image; Photographic image
    • G06T2207/10012Stereo images

Definitions

  • the invention is related to a method and a device for generating and processing depth data.
  • Said depth data is associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image.
  • An image is a two dimensional (2D) projection on an image plane located at a view point wherein the view point has a location and an orientation.
  • a second image may be used for preserving the depth information, implicitly, and for allowing a viewer to vary the orientation within a small range. That is, a scene is stereoscopically depicted by help of two images taken from two slightly different view points, like the eyes of a human, the difference depending on the focal length used. One of the images is provided to the right eye and the other is provided to the left eye. This allows a viewer for experiencing the depth of the scene and to vary the orientation of the view point within said small range. But the location of the viewpoint remains fixed.
  • Depth data is an explicit representation of depth information, for instance in form of a depth map, in form of a three-dimensional model of structures comprised in a depicted scene or in form of a set of layered images or depth planes.
  • views of a scene from arbitrary located and/or orientated viewpoints may be generated automatically in response to a viewer' s input determining a desired viewpoint.
  • Depth data may be collected using dedicated depth scanning devices based on laser ranging, ultrasonic ranging or the like. Another way of generating depth data is based on a mapping between images without using dedicated depth sensors. Such mapping -also known as disparity mapping- is based on the fact that relative dislocation between a place of depiction of a feature in one of the images and a place of depiction of the same feature in the other of the images is reciprocally proportional to the distance of said feature to the images.
  • Depth data may be corrupted in different ways. For instance, due to the reciprocal relationship minor errors in determining disparities of closed by features may result in severely erroneous depth data. Other sources of corruption are quantization during encoding and packet loss during transmission.
  • the inventors recognized that such corruptions more or less impede the viewer' s viewing experience in dependency on the degree of attention the viewer spends on image aspects generated using said corrupted depth data.
  • the invention proposes a method for generating and processing depth data associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image, said method comprising the features of independent claim 1.
  • a corresponding device is proposed which comprises the features of independent claim 6.
  • Said method comprises the steps of generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data, generating depth data using said mapping and processing the depth data wherein generating and/or processing of depth data is dependent on said control signal.
  • Fig. 1 depicts an exemplarily flow chart of generating and processing of a depth map which may be used in single frame plus depth (2D+z) or a multiview plus depth (MVD) encoding, for instance, and
  • Fig. 2 depicts an exemplarily flow chart of generating and processing of a mesh-based three dimensional model, for instance.
  • FIG. 1 An exemplarily flow chart of generating and processing of an integrated depth map DMn according to an exemplary embodiment of the inventive principle is depicted in Fig. 1.
  • the integrated depth map DMn is based on camera parameters CP and generated from disparity maps Dn-l/n and Dn+l/n which are generated from image In and image In-I, respectively, from image In and image In+1.
  • Disparity maps Dn-l/n and Dn+l/n are not only used for generating the integrated depth map DMn but also for warping or compensating saliency maps Sn+1 and Sn-I which are saliency maps generated from image In+1 and image In-I. Compensation or warping results in compensated saliency maps CSn+1 and CSn-I which are aligned with saliency map Sn generated from image In and therefore can be fused with or integrated into said saliency map Sn generated from image In which results in an integrated saliency map ISMn.
  • FIG. 2 Another exemplarily flow chart of the inventive principle is depicted in Fig. 2. From a set SoV of two or more views another set SoSM of a corresponding number of saliency maps is generated. The generated set SoSM of saliency maps as well as the set SoV of views are used for extracting one or more three dimensional models TdVM of one or more objects or structures at least partly depicted in the views of the set SoV of views .
  • the modelling is based on depth information determined using two or more views of said set SoV of views and controlled by saliency information determined using corresponding saliency maps of said set SoSM of saliency maps. That is, the more salient a surface region of a modelled object or structure is, the more detailed said surface region is modelled. For instance, in a mesh model the number of facets, nodes and/or edges is increased such that the size of facets in salient areas is reduced and approximation of the surface of the modelled object or structure is more precise in salient areas.
  • modelling is based on disparities of pairs of image regions of which one is comprised in a first view of the modelled object or structure and one is comprised in a second view of the modelled object or structure wherein said second view depicts the object or scene from a different viewpoint than said first view and the images regions of each pair are determined as depicting a same aspect of the object or structure to-be-modelled.
  • the extracted three dimensional models TdM may then be encoded into an encoded bitstream EBS wherein a region-of- interest based encoding is used which is based on one or more three dimensional saliency models TdSM of the modelled objects or structures.
  • the claimed method may be implemented on a processing device comprised in a notebook, a desktop computer, a server or any other computing device.
  • the computing device may further comprise a storage medium like a hard disk, a BluRay-disk or a flash memory carrying image data of images .
  • the processing device may be operated as a saliency value determining device which uses image data for determining corresponding saliency values.
  • the processing device may further be operated as a control signal generator which generates control signals in dependency on determined saliency values.
  • processing device may further be operated as a mapping device which maps first image data comprised in a first image onto second image data comprised in at least a second image.
  • the processing device may also be operated as a depth data generator which generates depth data in dependency on determined mappings and the generated control signal and/or the processing device may also be operated as a depth data processor which processes depth data in dependency on the generated control signal.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)
  • Testing, Inspecting, Measuring Of Stereoscopic Televisions And Televisions (AREA)

Abstract

The invention proposes a method for generating and processing depth data associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image. Said method comprising the steps of generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data, generating depth data using said mapping and processing the depth data wherein generating and/or processing of depth data is dependent on said control signal. This allows for more accurate generation and/or processing of the depth data in regions which are visually attracting and thus more likely in the focus of a viewer's attention.

Description

Method and a Device for generating and processing depth data
Technical Field
The invention is related to a method and a device for generating and processing depth data. Said depth data is associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image.
Background
An image is a two dimensional (2D) projection on an image plane located at a view point wherein the view point has a location and an orientation.
Due to the projection, some or all depth information comprised in a scene depicted by an image is lost.
It is known that a second image may be used for preserving the depth information, implicitly, and for allowing a viewer to vary the orientation within a small range. That is, a scene is stereoscopically depicted by help of two images taken from two slightly different view points, like the eyes of a human, the difference depending on the focal length used. One of the images is provided to the right eye and the other is provided to the left eye. This allows a viewer for experiencing the depth of the scene and to vary the orientation of the view point within said small range. But the location of the viewpoint remains fixed.
There is a trend in the art towards flexible viewpoint applications like free viewpoint TV (FTV) which allow a viewer to freely vary not only the orientation of the viewpoint but also the location. For making the location of the viewpoint freely variable as well, depth data has to be available. Depth data is an explicit representation of depth information, for instance in form of a depth map, in form of a three-dimensional model of structures comprised in a depicted scene or in form of a set of layered images or depth planes.
Using one or more images and depth data, views of a scene from arbitrary located and/or orientated viewpoints may be generated automatically in response to a viewer' s input determining a desired viewpoint.
Depth data may be collected using dedicated depth scanning devices based on laser ranging, ultrasonic ranging or the like. Another way of generating depth data is based on a mapping between images without using dedicated depth sensors. Such mapping -also known as disparity mapping- is based on the fact that relative dislocation between a place of depiction of a feature in one of the images and a place of depiction of the same feature in the other of the images is reciprocally proportional to the distance of said feature to the images.
Depth data may be corrupted in different ways. For instance, due to the reciprocal relationship minor errors in determining disparities of closed by features may result in severely erroneous depth data. Other sources of corruption are quantization during encoding and packet loss during transmission.
Invention
The inventors recognized that such corruptions more or less impede the viewer' s viewing experience in dependency on the degree of attention the viewer spends on image aspects generated using said corrupted depth data.
Therefore, the invention proposes a method for generating and processing depth data associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image, said method comprising the features of independent claim 1. A corresponding device is proposed which comprises the features of independent claim 6.
Said method comprises the steps of generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data, generating depth data using said mapping and processing the depth data wherein generating and/or processing of depth data is dependent on said control signal.
Making generation and/or processing of depth data dependent on saliency values which are determined for that image data whose mapping is associated with said depth data allows for more accurate generation and/or processing of the depth data in regions which are visually attracting and thus more likely in the focus of a viewer's attention.
Further advantageous embodiments of the method are characterized by the features of one of the claims depending on claim 1 and further advantageous embodiments of the device are characterized by the features of one of the claims depending on claim 6.
Drawings
Exemplary embodiments of the invention are illustrated in the drawings and are explained in more detail in the following description.
In the figures: Fig. 1 depicts an exemplarily flow chart of generating and processing of a depth map which may be used in single frame plus depth (2D+z) or a multiview plus depth (MVD) encoding, for instance, and
Fig. 2 depicts an exemplarily flow chart of generating and processing of a mesh-based three dimensional model, for instance.
Exemplary embodiments
An exemplarily flow chart of generating and processing of an integrated depth map DMn according to an exemplary embodiment of the inventive principle is depicted in Fig. 1.
The integrated depth map DMn is based on camera parameters CP and generated from disparity maps Dn-l/n and Dn+l/n which are generated from image In and image In-I, respectively, from image In and image In+1.
Disparity maps Dn-l/n and Dn+l/n are not only used for generating the integrated depth map DMn but also for warping or compensating saliency maps Sn+1 and Sn-I which are saliency maps generated from image In+1 and image In-I. Compensation or warping results in compensated saliency maps CSn+1 and CSn-I which are aligned with saliency map Sn generated from image In and therefore can be fused with or integrated into said saliency map Sn generated from image In which results in an integrated saliency map ISMn.
Thus, there are the integrated depth map DMn for image n as well as an integrated saliency map which allows for a region-of-interest based encoding of the integrated depth map DMn into an encoded bitstream EBS.
Another exemplarily flow chart of the inventive principle is depicted in Fig. 2. From a set SoV of two or more views another set SoSM of a corresponding number of saliency maps is generated. The generated set SoSM of saliency maps as well as the set SoV of views are used for extracting one or more three dimensional models TdVM of one or more objects or structures at least partly depicted in the views of the set SoV of views .
The modelling is based on depth information determined using two or more views of said set SoV of views and controlled by saliency information determined using corresponding saliency maps of said set SoSM of saliency maps. That is, the more salient a surface region of a modelled object or structure is, the more detailed said surface region is modelled. For instance, in a mesh model the number of facets, nodes and/or edges is increased such that the size of facets in salient areas is reduced and approximation of the surface of the modelled object or structure is more precise in salient areas.
Commonly, modelling is based on disparities of pairs of image regions of which one is comprised in a first view of the modelled object or structure and one is comprised in a second view of the modelled object or structure wherein said second view depicts the object or scene from a different viewpoint than said first view and the images regions of each pair are determined as depicting a same aspect of the object or structure to-be-modelled.
The extracted three dimensional models TdM may then be encoded into an encoded bitstream EBS wherein a region-of- interest based encoding is used which is based on one or more three dimensional saliency models TdSM of the modelled objects or structures.
The claimed method may be implemented on a processing device comprised in a notebook, a desktop computer, a server or any other computing device. The computing device may further comprise a storage medium like a hard disk, a BluRay-disk or a flash memory carrying image data of images .
The processing device may be operated as a saliency value determining device which uses image data for determining corresponding saliency values.
The processing device may further be operated as a control signal generator which generates control signals in dependency on determined saliency values.
And, the processing device may further be operated as a mapping device which maps first image data comprised in a first image onto second image data comprised in at least a second image.
The processing device may also be operated as a depth data generator which generates depth data in dependency on determined mappings and the generated control signal and/or the processing device may also be operated as a depth data processor which processes depth data in dependency on the generated control signal.
The exemplary embodiments of the invention are described for illustrative purposes only and shall not be construed as limiting the extent of projection as requested which is solely defined by the independent claims.

Claims

Claims
1. Method for generating and processing depth data associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image, wherein processing of depth data comprises encoding of the depth data and said method comprises the steps of
- generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data,
- generating depth data using said mapping and
- processing the depth data wherein
- encoding of depth data is dependent on said control signal .
2. Method according to claim 1, wherein generating said depth data is also dependent on said control signal and comprises generating one or more nodes of a mesh model and assigning one or more further saliency values to the generated nodes with the number of generated nodes and/or values of the further saliency values being dependent on said control signal.
3. Method according to one of the preceding claims, wherein processing of the depth data further comprises lossy compression of the depth data and said control signal is further used for controlling a degree of compression loss.
4. Method according to one of the preceding claims, wherein encoding of the depth data comprises ordered encoding of the depth data together with further depth data in a code sequence and said control signal is used for determining a rank in the encoding order.
5. Method according to one of the preceding claims, wherein the control signal is generated dependent on a sum of said first saliency value and said at least a second saliency value.
6. Device for generating and processing depth data associated with a mapping of first image data comprised in a first image onto second image data comprised in at least a second image, wherein processing comprises encoding of the depth data and said device comprises
- means for generating a control signal which depends on a first saliency value determined using the first image data, wherein said control signal further depends on at least a second saliency value determined using the second image data,
— means for generating depth data using said mapping,
- means for processing depth data and
— a controller adapted for using said control signal for controlling said encoding of the depth data.
7. Device according to claim 6, wherein said means for generating depth data is adapted for using said control signal for controlling said generating one or more nodes of a mesh model and assigning one or more further saliency values to the generated nodes with the controller being adapted for controlling the number of generated nodes and/or values of the further saliency values in dependence on said control signal.
8. Device according to one of the claims 6 or 7, wherein processing of the depth data further comprises lossy compression of the depth data and the controller is further adapted for using said control signal for controlling a degree of compression loss.
9. Device according to one of the claims 6-8, wherein encoding of the depth data comprises ordered encoding of the depth data together with further depth data in a code sequence and the controller is adapted for using said control signal for determining a rank in an encoding order,
10. Device according to one of the claims 6-9, wherein means for generating a control signal is adapted for generating the control signal dependent on a sum of said first saliency value and said at least a second saliency value .
11. Device according to claim 10 or method according to claim 5, wherein said sum is a weighted sum.
PCT/EP2010/054773 2009-04-29 2010-04-12 Method and a device for generating and processing depth data Ceased WO2010124926A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP09305372.6 2009-04-29
EP09305372A EP2247114A1 (en) 2009-04-29 2009-04-29 Method and a device for generating and processing depth data

Publications (1)

Publication Number Publication Date
WO2010124926A1 true WO2010124926A1 (en) 2010-11-04

Family

ID=41057425

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2010/054773 Ceased WO2010124926A1 (en) 2009-04-29 2010-04-12 Method and a device for generating and processing depth data

Country Status (2)

Country Link
EP (1) EP2247114A1 (en)
WO (1) WO2010124926A1 (en)

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1544792A1 (en) * 2003-12-18 2005-06-22 Thomson Licensing S.A. Device and method for creating a saliency map of an image

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1544792A1 (en) * 2003-12-18 2005-06-22 Thomson Licensing S.A. Device and method for creating a saliency map of an image

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
FELDMANN I ET AL: "REAL-TIME SEGMENTATION FOR ADVANCES DISPARITY ESTIMATION IN IMMERSIVE VIDEOCONFERENCE APPLICATIONS", JOURNAL OF WSCG, VACLAV SKALA-UNION AGENCY, PILZEN, CZ, vol. 10, no. 1, 1 January 2002 (2002-01-01), pages 171 - 178, XP008017665, ISSN: 1213-6972 *
LE MEUR O ET AL: "Efficient Saliency-Based Repurposing Method", IMAGE PROCESSING, 2006 IEEE INTERNATIONAL CONFERENCE ON, IEEE, PI, 1 October 2006 (2006-10-01), pages 421 - 424, XP031048663, ISBN: 978-1-4244-0480-3 *
YONGQING ZENG: "Perceptual segmentation algorithm and its application to image coding", IMAGE PROCESSING, 1999. ICIP 99. PROCEEDINGS. 1999 INTERNATIONAL CONFE RENCE ON KOBE, JAPAN 24-28 OCT. 1999, PISCATAWAY, NJ, USA,IEEE, US, vol. 2, 24 October 1999 (1999-10-24), pages 820 - 824, XP010369077, ISBN: 978-0-7803-5467-8 *

Also Published As

Publication number Publication date
EP2247114A1 (en) 2010-11-03

Similar Documents

Publication Publication Date Title
CA2668941C (en) System and method for model fitting and registration of objects for 2d-to-3d conversion
JP5561781B2 (en) Method and system for converting 2D image data into stereoscopic image data
JP5887267B2 (en) 3D image interpolation apparatus, 3D imaging apparatus, and 3D image interpolation method
KR101907945B1 (en) Displaying graphics in multi-view scenes
CA2723627C (en) System and method for measuring potential eyestrain of stereoscopic motion pictures
JP7664891B2 (en) Method for generating layered depth data of a scene - Patents.com
CN102404592B (en) Image processing device and method, and stereoscopic image display device
CN114830676B (en) Video processing devices and manifest files for video streaming
CN1925627A (en) Apparatus for controlling depth of 3D picture and method therefor
TWI523488B (en) A method of processing parallax information comprised in a signal
JPWO2012035783A1 (en) 3D image creation apparatus and 3D image creation method
KR20160135660A (en) Method and apparatus for providing 3-dimension image to head mount display
WO2011030399A1 (en) Image processing method and apparatus
JP5178538B2 (en) Method for determining depth map from image, apparatus for determining depth map
Knorr et al. An image-based rendering (ibr) approach for realistic stereo view synthesis of tv broadcast based on structure from motion
CN115997379A (en) Image FOV Restoration for Stereoscopic Rendering
JP2009530701A5 (en)
WO2019092398A1 (en) Stereoscopic image data compression
WO2008152607A1 (en) Method, apparatus, system and computer program product for depth-related information propagation
EP2247114A1 (en) Method and a device for generating and processing depth data
KR101634225B1 (en) Device and Method for Multi-view image Calibration
Gurrieri et al. Efficient panoramic sampling of real-world environments for image-based stereoscopic telepresence
US12375633B2 (en) Low complexity multilayer images with depth
CN101566784B (en) Method for establishing depth of field data for three-dimensional image and system thereof
CN110740308B (en) Time consistent reliability delivery system

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 10713923

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 10713923

Country of ref document: EP

Kind code of ref document: A1