EP4631249A1 - Offset low discrepancy spherical sampling for image rendering - Google Patents
Offset low discrepancy spherical sampling for image renderingInfo
- Publication number
- EP4631249A1 EP4631249A1 EP23837518.2A EP23837518A EP4631249A1 EP 4631249 A1 EP4631249 A1 EP 4631249A1 EP 23837518 A EP23837518 A EP 23837518A EP 4631249 A1 EP4631249 A1 EP 4631249A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- offset
- sampling
- cube map
- image processing
- processing system
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/597—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding specially adapted for multi-view video sequence encoding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/70—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
Definitions
- Foveated rendering refers to a collection of computer graphics techniques in which the resolution of a rendered image is tailored to match the visual acuity of the eyes of the viewer. That is, more detail is provided at the center of the viewer’s gaze, and progressively less detail is provided further from that point.
- This technique is implemented in virtual reality (VR) and augmented reality (AR) applications, as foveated rendering provides reductions in rendering time and data transmission bandwidth.
- Cube map texturing is a technique for representing a three-dimensional (3D) environment map such as a 360° still image or video frame. Particularly, a textured 3D cube is represented as six two-dimensional (2D) textures, one for each cube face.
- Modern graphics processing units offer hardware support for cube map texturing, including handling sampling across more than a single cube face.
- BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS [0005] Disclosed herein are various embodiments of image processing systems for encoding and decoding spherically-sampled data.
- One example embodiment provides an image processing system comprising a processor to perform encoding of data, the processor configured to receive a three-dimensional environment.
- the processor is configured to apply, for a set of cube map texture coordinates, a cube map function to the cube map texture coordinates to generate a cube map direction, apply a low discrepancy sequence sampling function to the cube map direction to generate a low discrepancy spherical sampling direction, and apply an offset cube mapping function to the low discrepancy spherical sampling direction to generate an offset sampling direction.
- the processor is configured to generate a pixel value for the cube map texture coordinates using the offset sampling direction and the three-dimensional environment.
- the processor is configured to receive a three-dimensional environment, apply, to a set of cube map texture coordinates, an offset low discrepancy sequence (OLDS) sampling function to generate an OLDS direction, and generate a pixel value for the cube map texture coordinates using the OLDS direction and the three-dimensional environment.
- an image processing system for decoding spherically-sampled data.
- the image processing system comprises a processor to perform decoding of data, the processor configured to receive, from a coded bitstream, an offset sampled cube map.
- the processor is configured to apply, for an offset sampling direction, an inverse offset cube mapping function to the offset sampling direction to generate a low discrepancy spherical sampling direction, apply an inverse low discrepancy sequence sampling function to the low discrepancy spherical sampling direction to generate a cube map direction, and apply an inverse cube map function to the cube map direction to generate cube map texture coordinates.
- the processor is configured to generate a pixel value for the offset sampling direction using the cube map texture coordinates and the offset sampled cube map.
- FIG.1 depicts an example process for a video delivery pipeline according to one embodiment.
- FIG.2 depicts an example representation of a vision field of an average viewer’s eye.
- FIG.3 depicts an example cube map within a 3D environment according to one embodiment.
- FIGS.4A-4C depict an example offset cube map within a 3D environment according to one embodiment.
- FIGS.5A-5C depict various example sampling density uniformity of a cube map.
- FIGS.6A-6C depict various example sampling density uniformity of an offset cube map.
- FIG.7 depicts triangle areas from Delaunay triangulation of an offset standard cube map spherical sampling.
- FIG.8 depicts triangle area from Delaunay triangulation of an offset unicube cube map spherical sampling.
- FIG.9 depicts an example process performed by the video delivery pipeline of FIG.1 according to one embodiment.
- FIG.10 depicts another example process performed by the video delivery pipeline of FIG.1 according to one embodiment.
- a video application as described herein may refer to any of: video display applications, VR applications, AR applications, automobile entertainment applications, remote presence applications, display applications, gaming applications, mobile applications, internet- based video streaming applications, and the like.
- the techniques can be applied to virtual reality cinema-immersive movie watching (e.g., for headmounted displays, etc.) for streaming video data between video streaming server(s) and video streaming client(s).
- Example video content may include, but are not necessarily limited to, any of: audiovisual programs, movies, video programs, TV broadcasts, computer games, AR content, VR content, automobile entertainment content, and the like.
- Example video streaming clients may include, but are not necessarily limited to, any of: display devices, a computing device with a near-eye display, a head-mounted display (HMD), a mobile device, a wearable display device, a set-top box with a display such as television, a video monitor, and the like.
- a “video streaming server” may refer to one or more upstream devices that prepare and stream video content to one or more video streaming clients in order to render at least a portion (e.g., corresponding to a user’s FOV or viewport, etc.) of the video content on one or more (target) displays.
- Example video streaming servers may include, but are not necessarily limited to, any of: cloud-based video streaming servers located remotely from video streaming client(s), local video streaming servers connected with video streaming client(s) over local wired or wireless networks, VR devices, AR devices, automobile entertainment devices, digital media devices, digital media receivers, set-top boxes, gaming machines, general purpose personal computers, tablets, dedicated digital media receivers, and the like.
- Video Coding according to Example Embodiments FIG.1 depicts an example process of a video delivery pipeline 100, showing various stages from video capture to video-content display according to an embodiment.
- a sequence of video frames 102 may be captured or generated using an image-generation block 105.
- the video frames 102 may be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data 107.
- the video frames 102 may be captured on film by a film camera.
- the film may be translated into a digital format to provide the video data 107.
- the video data 107 may be edited to provide a video production stream 112.
- the data of the video production stream 112 may then be provided to a processor (or one or more processors, such as a central processing unit, CPU) at a post- production block 115 for post-production editing.
- a processor or one or more processors, such as a central processing unit, CPU
- the post-production editing of the block 115 may include, e.g., adjusting or modifying colors or brightness in particular areas of an image to enhance the image quality or achieve a particular appearance for the image in accordance with the video creator’s creative intent. This part of post-production editing is sometimes referred to as “color timing” or “color grading.” Other editing (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, etc.) may be performed at the block 115 to yield a “final” version 117 of the production for distribution. For example, rendering techniques described herein, such as cube mapping techniques, cube map offset techniques, and sampling techniques may be performed during post-production editing 115.
- video images may be viewed on a reference display 125.
- the rendering techniques described herein may be performed at image generation block 105, production phase 110, or another suitable processing step within the video delivery pipeline 100.
- video data of the final version 117 may be delivered to a coding block 120 for being delivered downstream to decoding and playback devices, such as VR headsets, near-eye displays, and the like.
- the coding block 120 may include audio and video encoders, such as those defined by the ATSC, DVB, DVD, Blu-Ray, and other delivery formats, to generate a coded bitstream 122.
- the coded bitstream 122 is decoded by a decoding unit 130 to generate a corresponding decoded signal 132 representing a copy or a close approximation of the signal 117.
- the receiver may be attached to a target display 140 that may have somewhat or completely different characteristics than the reference display 125.
- a display management (DM) block 135 may be used to map the decoded signal 132 to the characteristics of the target display 140 by generating a display-mapped signal 137.
- the decoding unit 130 and display management block 135 may include individual processors or may be based on a single integrated processing unit.
- the coded bitstream 122 may include metadata to assist in the reconstruction of the final version 117.
- FIG.2 illustrates an example representation of a vision field of an average viewer’s eye. Cone and rod distributions (in the eye) can be segmented into different distribution ranges of cones and rods and further projected into an angular vision field representation (of the eye) as illustrated in FIG.2. Highest levels of visual perception are achieved in the eye’s foveal (vision field) region 202.
- the widest angular range in the eye’s vision field is along the horizontal direction of FIG.2, which is parallel to the inter-pupil line between the viewer’s two eyes, without considering visual constraints from facial anatomy, and may be approximately 180 angular degrees.
- Each of concentric circles e.g., labelled as 30°, 60°, 90°, etc.
- angles such as 30°, 60°, 90°, etc., are for illustration purposes only. Different values of angles or different sets of angles can be used to define or describe a viewer’s vision field.
- the view direction (not shown in FIG.2) is pointed vertically out of the plane of FIG.2 at the intersection of a transverse direction 212 and a vertical direction 214 in a foveal region 202 (the darkest fill pattern).
- the transverse direction 212 and the vertical direction 214 form a plane normal to the view direction.
- the vision field of the eye may be partitioned (e.g., logically, projected by certain partitions in the distributions of densities of rods/cones, etc.) into the foveal region 202 immediately surrounded by a paracentral region 204.
- the foveal region 202 may correspond to the viewer’s fovea vision and extend from zero (0) angular degrees to a first angle (e.g., 2-4 angular degrees, 3-7 angular degrees, 5-9 angular degrees, etc.) relative to the view direction.
- the paracentral region 204 may extend from the first angle to a second angle (e.g., 6-12 angular degrees, etc.) relative to the view direction. [0031]
- the paracentral region 204 is immediately surrounded by a near-peripheral region 206.
- the near-peripheral region 206 is immediately adjacent to the mid-peripheral region 208, which in turn is immediately adjacent to the rest of the vision field, a far-peripheral region 210.
- the near-peripheral region 206 may extend from the second angle to a third angle (e.g., 25-35 angular degrees, etc.) relative to the view direction.
- the mid-peripheral region 208 may extend from the third angle to a fourth angle (e.g., 50-65 angular degrees, etc.) relative to the view direction.
- the far-peripheral region 210 may extend from the fourth angle to the edge of the vision field.
- the first, second, third and fourth angles used in this example logical partition of the vision field may be defined or specified along the transverse direction 212.
- the transverse direction 212 may be the same as, or parallel to, the viewer’s interpupil line.
- different schemes of logically partitioning a viewer’s vision field may be used in addition to, or in place of, the scheme of logically partitioning the viewer’s vision field into foveal, paracentral, near-peripheral, mid-peripheral, far-peripheral, etc., regions based on angles as illustrated in FIG.2.
- the viewer’s vision field may be partitioned into more or fewer regions such as a combination of foveal region, a near-peripheral region and a far- peripheral region, etc., without a paracentral region and/or a mid-peripheral region.
- a spatially faithful representation (or high-fidelity) image portion may be used to cover from the foveal region up to some or all of the near-peripheral region in such logical partition of the viewer’s vision field.
- the viewer’s vision field may be partitioned based on other quantities other than angles as illustrated in FIG.2.
- the foveal region may be defined as a vision field region that corresponds a viewer’s foveal-vision.
- the paracentral region may be defined as a vision field region that corresponds to a viewer’s retina area where cone/rod densities exceed relatively high cone/rod density thresholds.
- the near-peripheral region may be defined as a vision field region that corresponds a viewer’s retina area where cone/rod densities does not exceed relatively high cone/rod density thresholds respectively but does exceed intermediate cone/rod density thresholds.
- the mid-peripheral region may be defined as a vision field region that corresponds a viewer’s retina area where cone/rod densities does not exceed intermediate cone/rod density thresholds respectively but does exceed relatively low cone/rod density thresholds.
- a focal- vision region as described herein may cover from the viewer’s foveal-vision up to some or all of a region (e.g., some or all of the viewer’s near-peripheral vision, etc.) based on threshold(s) (e.g., cone/rod density threshold(s), etc.) that are not necessarily angle-based.
- a combination of two or more different schemes of logically partitioning the viewer’s vision field and/or other human vision factors may be used to determine a focal-vision region of the viewer’s vision field.
- the focal-vision region as described herein may cover a larger angular value range along the transverse direction 212 than an angular value range covered by the focal-vision region along the vertical direction 214, as the human vision system may be more sensitive to image details along the transverse direction 212 than those along the vertical direction 214.
- a focal-vision region as described herein covers some or all of: a foveal region (e.g., plus a safety margin, etc.), a paracentral region (e.g., excluding and extending from the foveal region, etc.), a near-peripheral region (e.g., further excluding and extending from the paracentral region, etc.), a mid-peripheral region (e.g., further excluding and extending from the near peripheral region, etc.), etc.
- a vision field of an eye as described herein takes into consideration vision-related factors such as eye swiveling, viewing constraints from nose, corneal, eyelid, etc.
- Examples of a focal-vision region as described herein may include, but are not necessarily limited to, any combination of one or more of: circular shapes, oblong shapes, oval shapes, heart shapes, star shapes, round shapes, square shapes, polygonal shapes, etc.
- Gaze Tracking [0040] In some embodiments, only a (e.g., relatively small) focal-vision region of the eye’s vision field needs to be provided with pixel values with the highest dynamic range, the widest color gamut, the highest (or sharpest) spatial resolution, and the like. In some embodiments, the focal-vision region of the eye’s vision field may approximately correspond to the entirety of the foveal-vision of the eye up to some or all of near-peripheral vision of the eye.
- the focal-vision region of the eye’s vision field may additionally include a safety vision field region.
- the size and/or shape of the safety vision field region in the focal-vision region can be preconfigured to a fixed size (e.g., 0%, 5%, 10%, -5%, -10%, etc.) that does not vary with network bandwidth, image content, types of computing devices (e.g., helmet mounted display devices, wall displays, etc.) involved in video applications, types of rendering environments (e.g., cloud-based video streaming servers, etc.) involved in video applications, and the like.
- the size and/or shape of the safety vision field region in the focal-vision region can be dynamically reconfigured at runtime, and can vary in range. For example, in response to determining that network connections do not support a relatively high bandwidth, the size and/or shape of the safety vision field region may be dynamically shrunk at runtime from 10% to 5% over the eye’s foveal-vision. On the other hand, in response to determining that network connections support a relatively high bandwidth, the size and/or shape of the safety vision field region may be dynamically expanded at runtime from 5% to 10% over the eye’s foveal-vision. [0042] The size and/or shape of the safety vision field region may also be set in dependence on latency in eye tracking.
- eye tracking data as described herein can be used to predict where the viewer would look next and reduce the bandwidth/safety region based on the prediction.
- the user’s view direction at runtime may be tracked by a view direction tracking device.
- the view direction tracking device may operate in real time with a display on which a sequence of display mapped images is rendered.
- the view direction tracking device tracks and computes the viewing angles and/or viewing distances in a coordinate system in which the sequence of display mapped images is being rendered, generates a time sequence of view directions, and signals each view direction in the time sequence of view directions to a video streaming server as described herein.
- Each such signaled view direction of the viewer as received by the video streaming server may be indexed by a time point value.
- the time point value may be associated or correlated by a video streaming server as described herein with an offset LDSS cube map.
- View direction data to track the viewer’s view directions is collected while the viewer is viewing a 3D environment or a derivative version.
- Example view direction data may include, without limitation, linear displacements, angular displacements, linear motions or translations, angular motions or rotations, pitch, roll, yaw, sway, heave, surge, etc., that may be collected by any combination of gaze tracking devices, face tracking devices, FOV tracking devices, and the like.
- View direction data may be collected, analyzed, and/or shared/transmitted among view direction tracking devices and streaming devices with relatively low latency (e.g., within a fraction of one image frame time, within 5 milliseconds, etc.).
- the view direction tracking data may be shared among these devices using the lowest latency data/network connections where multiple data/network connections are available.
- the view direction data may be referred to herein as a gaze vector.
- a video streaming server may dynamically shrink the size and/or shape of the safety vision field region at runtime from 10% to 5% over the eye’s foveal-vision.
- a relatively small area e.g., within 20 angular degrees from the view direction, etc.
- the highest dynamic range, the widest color gamut, the highest spatial resolution, etc. may be sent in the video signal to the downstream recipient device.
- the video streaming server may dynamically expand the size and/or shape of the safety vision field region at runtime from 1% to 3%, 2% to 6%, 5% to 10%, etc., over the eye’s foveal vision.
- a relatively large area e.g., up to 30 angular degrees from the view direction, etc.
- the highest dynamic range, the widest color gamut, the highest spatial resolution, etc. may be sent in the video signal to the downstream recipient device.
- Embodiments described herein relate to 3D virtual environments generated by a computer graphics engine or captured using one or more cameras. The 3D virtual environment is then accessed using an immersive device, such as a virtual reality headset. To assist with rendering images where resolution is greatest at the center of the viewer’s gaze, foveated rendering techniques may be implemented.
- One technique for representing a 3D environment in one or more 2D textures includes mapping the 3D virtual environment to a cube, thereby generating a cube map.
- FIG.3 provides an example cube mapping technique for a 3D environment (e.g., a scene).
- a 3D environment 300 includes a textured cube map 305.
- a center of the cube map 305 is located at an arbitrary point within of the 3D environment 300.
- the 3D environment 300 is mapped onto the cube map 305.
- the cube map 305 includes a first face 315.
- a camera view 310 e.g., a virtual camera
- the portion of the 3D environment 300 captured by the camera view 310 is mapped onto the first face 315. This process is repeated for each face of the cube map 305 such that the entire 3D environment 300 is mapped onto the cube map 305.
- a virtual cube (e.g., the cube map 305) encloses the six-camera arrangement, with the center of the cube map 305 being coincident with the focal point of the cameras.
- the focal point of the six cameras may be anywhere within the scene, and is not necessarily located at the center of the scene.
- the three orthogonal axes pass through the centers of the faces of the cube map 305 (such as the first face 315). Uniform sampling of each face of the virtual cube produces the six images that form the textured cube map 305.
- Each image is a view of the scene as captured by one of the six cameras. In some instances, the field of view of each camera is 90 degrees.
- cube mapping may result in spherical samples (or a 3D model map) being more heavily concentrated near the spherical projections of the cube map’s edges and corners, and less so at the projections of the cube face centers. This results in higher resolution (in pixels per solid angle) at the corners and edges of the cube and a lower resolution at the center of each face (relative to the edges and corners). Such nonuniformity leads to area and angle distortions in environment maps and artifacts when the cube map is rendered in a viewport. Additionally, when rendering videos, standard cube map sampling wastes rendering time and bandwidth for directions far from the viewer’s gaze. [0051] To address such issues, offset cube mapping techniques may be implemented.
- FIGS.4A-4C provides an example offset cube mapping technique for a 3D environment.
- FIG.4A illustrates a 3D environment 400 which includes a first virtual camera 405, a textured cube map 412 comprised of six faces, four of which (front face 410, back face 415, and side faces 420) are shown in FIG.4A.
- FIG.4B illustrates a second virtual camera 425 in the 3D environment 400 offset by a bias distance 427 (e.g., an offset bias), and a textured offset cube map 432 comprised of six faces, four of which (front face 430, back face 435, and side faces 440) are shown in FIG.4B.
- the bias distance 427 may be a vector having a magnitude and direction away from the common centers of the textured cube map 412 and textured offset cube map 432.
- a common center of the textured cube map 412 and textured offset cube map 432 is located at a general point within the 3D environment 400.
- FIG.4C illustrates the second virtual camera 425 capturing a narrower view 445 of the environment 400 onto the front face 430 compared to the first virtual camera 405, as depicted by the smaller range of angles subtended by 430 relative to 410, consequently providing a higher level of detail in the front direction.
- the offset back-facing camera captures a wider view 450 of the environment 400 onto the back face 435 than the back-facing standard cube map camera (not shown), consequently providing a lower level of detail in that direction.
- a virtual cube e.g., the first textured cube map 405
- the cube center is offset, along one of the orthogonal axes, from the common camera focal point. Accordingly, the view of the scene of each camera is different, with some cameras having more narrow views and some having wider views of the scene.
- the offset cube mapping is performed by applying a cube map texture coordinate transformation applied in a shader before performing a cube map sampling function.
- a cube map texture coordinate transformation applied in a shader before performing a cube map sampling function.
- rendering time and required bandwidth are both reduced.
- offset cube mapping may experience a spatial sample warping due to the offset.
- Low Discrepancy Spherical Sampling Embodiments described herein provide for the application of the offset cube map’s virtual camera bias to a low discrepancy spherical sampling (LDSS).
- LDSS low discrepancy spherical sampling
- Discrepancy as described herein refers to a measure of the equi-distribution of a pointset on a sphere, where a lower discrepancy value implies a more uniform sampling. Further details on equi-distribution of a pointset on a sphere can be found in “J. Cui, et al., Equidistribution on the Sphere, SIAM Journal on Scientific Computing, 1997, Vol.18, No.2: pp.595-609”, incorporated herein by reference. LDSS techniques described herein exhibit lower discrepancy, area, and angle distortion than standard (e.g., non-offset) cube mapping (such as that shown in FIG.5).
- Isocube sampling includes first partitioning a spherical surface into four equatorial zones and two polar zones, the area of each equatorial zone and each polar zone being equal. Each zone is then partitioned into a number of N ⁇ N elements by a plurality of longitudinal curves and a plurality of latitudinal curves. The resulting N ⁇ N elements, when mapped to a cube, are relatively curved (e.g., warped) compared to samples of the standard cube map. Samples on the isocube map are indexed on rings that run parallel to the equator of the sampled sphere.
- Unicube sampling includes partitioning a cube face based on two sets of parallel grid lines. The parallel lines are nonuniformly distributed such that the partition on the projected spherical surface maintains a uniform rectilinear structure. The nonuniform partition of the cube face is then mapped to a uniform partition.
- FIG.5A illustrates a unicube cube map 500 having unicube sampling density.
- FIG.5B illustrates a standard cube map 510 having a non-uniform (e.g., no LDSS technique applied) sampling density.
- FIG.5C illustrates an isocube cube map 520 having an isocube sampling density.
- FIG.6A illustrates a spherical projection of an offset unicube cube map 600 having both unicube sampling density and offset bias.
- FIG.6B illustrates a spherical projection of an offset standard cube map 610 having both standard cube map sampling density and offset bias.
- FIG.6C illustrates a spherical projection of an offset isocube cube map 620 having both isocube sampling density and offset bias.
- Sampling of the offset unicube cube map 600 and the offset isocube cube map 620 is more uniform compared to the offset cube map 610 (ignoring intended spatial warping of the sample coordinates due to offset). Accordingly, implementing both offset bias and LDSS techniques improves the resulting spatial sample distribution relative to using only offset bias. Additionally, implementing both offset bias and LDSS techniques results in better image quality, reduced transmission bandwidth requirements, and (in some instances) better video compression ratios.
- FIG.7 illustrates the area of triangles from Delaunay triangulation of an offset cube map having no LDSS technique applied.
- triangles having a smaller area are located along the spherical projection regions corresponding to the edges and corners of a face of the cube map.
- FIG.8, for comparison illustrates the area of triangles from Delaunay triangulations of an offset unicube cube map. The triangle areas in FIG.8 decrease monotonically closer to the center of the offset face compared to FIG.7.
- a direction vector as used herein refers to a 3-element vector within a 3D environment.
- a direction vector may be referred to simply as a direction (for example, a unit length direction, a sampling direction, etc.). Equivalent or corresponding spherical angles may be computed from x.
- u (u, v, f) be cube map texture coordinates, where (u, v) are horizontal and vertical texture coordinates and f is a face index in the range of [1, 6].
- Offset cube mapping ocm(x, d) is a function taking a direction vector x and an offset d and returning a modified direction vector x o .
- the offset cube map pixel at u i.e., the pixel at location (u, v) on face f
- OLDS offset low discrepancy sequence
- This OLDS texture may be encoded and transmitted to a receiver.
- the offset d is also transmitted to the receiver, or an agreed-upon value by both the sender and the receiver is used.
- the receiver When the receiver is rendering a scene for display in, for example, an AR or VR headset, the receiver obtains samples of the scene in a plurality of directions.
- FIG.9 illustrates a method 900 for generating pixel values for offset sampling directions.
- the method 900 may be performed by an electronic processor included in the post- production block 115, an electronic processor included in the image generation block 105, an electronic processor included in the production block 110, and the like.
- the electronic processor receives a 3D environment.
- the 3D environment may be, for example, a 3D environment map, a 3D game engine scene, a ray-traced environment (e.g., ray tracers), texture- mapped spheres, and other virtual 3D environments capable of being mapped onto a cube map.
- the electronic processor applies, for a set of cube map texture coordinates, a cube map function to the cube map texture coordinates to generate a cube map direction.
- the electronic processor applies a low discrepancy sequence sampling function to the cube map direction to generate a low discrepancy spherical sampling direction.
- the electronic processor applies an offset cube mapping function to the low discrepancy spherical sampling direction to generate an offset sampling direction.
- the electronic processor generates a pixel value for the cube map texture coordinates using the offset sampling direction and the 3D environment.
- FIG.10 illustrates a method 1000 for decoding an OLDS texture (e.g., an offset sampled cube map).
- the method 1000 may be performed by an electronic processor included in the decoding block 130.
- the electronic processor receives an offset sampled cube map.
- the decoding block 130 decodes a coded bitstream 122 to obtain the offset sampled cube map.
- the electronic processor applies, for an offset sampling direction, an inverse offset cube mapping function to the offset sampled cube map to generate a low discrepancy spherical sampling direction.
- the coded bitstream 122 may include metadata to assist with reconstructing the 3D environment.
- the metadata includes the distance and direction of the bias distance 427 that define ocm -1 ().
- the electronic processor reverses the distance and direction of the bias distance 427 to undo the bias distance 427 by applying ocm -1 () to the offset sampling direction xC to obtain xL, as previously described.
- ocm -1 () to the offset sampling direction xC to obtain xL, as previously described.
- the electronic processor applies an inverse low discrepancy sequence sampling function to the low discrepancy spherical sampling direction to generate a cube map direction.
- the electronic processor applies lds -1 () to the inverse low discrepancy sequence sampling direction x L to obtain a standard direction x, as previously described.
- the electronic processor applies an inverse cube mapping function to the cube map direction to generate cube map texture coordinates.
- the electronic processor applies cm -1 () to the cube map direction x to obtain cube map texture coordinates u.
- the electronic processor generates a pixel value for the offset sampling direction using the cube map texture coordinates and the offset sampled cube map.
- the bias vector (d) e.g., the magnitude and/or the direction
- the magnitude of the bias vector may be varied based on a rate of change of the gaze vector. For example, when the viewing direction is moving quickly, the magnitude of the bias vector may be reduced such that visual quality is more consistent across quickly-changing viewing angles.
- the magnitude of the bias vector (e.g., the amount of offset) is determined based on a linear function.
- the magnitude of the bias vector is increased as the rate of change of the gaze vector increases beyond thresholds. For example, when the rate of change is below a first threshold, a first bias vector magnitude is selected. When the rate of change is above the first threshold but less than a second threshold, a second bias vector magnitude is selected. When the rate of change is above the second threshold, a third bias vector magnitude is selected.
- the magnitude of the bias vector may be varied based on the likelihood that a viewer will be looking in a particular direction, or on the importance of action at a particular orientation. For example, for video content including a story line, it is highly likely that viewers will be looking in the specific direction of the story.
- the axis of offset e.g., the direction of the bias vector
- embodiments described herein have primarily referred to a full 360° environment, in some instances, less than the full 360° environment is rendered. For example, a subset of the six cube faces (for example, five faces) may be rendered.
- An image processing system for encoding spherically-sampled data comprising: a processor to perform encoding of data, the processor configured to: receive a three-dimensional environment; apply, for a set of cube map texture coordinates, a cube map function to the cube map texture coordinates to generate a cube map direction; apply a low discrepancy sequence sampling function to the cube map direction to generate a low discrepancy spherical sampling direction; apply an offset cube mapping function to the low discrepancy spherical sampling direction to generate an offset sampling direction; and generate a pixel value for the cube map texture coordinates using the offset sampling direction and the three-dimensional environment.
- An image processing system for encoding spherically-sampled data comprising: a processor to perform encoding of data, the processor configured to: receive a three-dimensional environment; apply, to a set of cube map texture coordinates, an offset low discrepancy sequence (OLDS) sampling function to generate an OLDS direction; and generate a pixel value for the cube map texture coordinates using the OLDS direction and the three-dimensional environment.
- a processor to perform encoding of data
- the processor configured to: receive a three-dimensional environment; apply, to a set of cube map texture coordinates, an offset low discrepancy sequence (OLDS) sampling function to generate an OLDS direction; and generate a pixel value for the cube map texture coordinates using the OLDS direction and the three-dimensional environment.
- the OLDS sampling function includes a unicube sampling function.
- the OLDS sampling function includes an isocube uniform sampling function.
- An image processing system for decoding spherically-sampled data comprising: a processor to perform decoding of data, the processor configured to: receive, from a coded bitstream, an offset sampled cube map; apply, for an offset sampling direction, an inverse offset cube mapping function to the offset sampling direction to generate a low discrepancy spherical sampling direction; apply an inverse low discrepancy sequence sampling function to the low discrepancy spherical sampling direction to generate a cube map direction; apply an inverse cube map function to the cube map direction to generate cube map texture coordinates; and generate a pixel value for the offset sampling direction using the cube map texture coordinates and the offset sampled cube map.
- Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s).
- Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and/or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s).
- the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”
- the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.
- the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard, and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard.
- the compatible element does not need to operate internally in a manner specified by the standard.
- the functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and/or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared.
- processor or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and/or custom, may also be included.
- DSP digital signal processor
- ASIC application specific integrated circuit
- FPGA field programmable gate array
- ROM read only memory
- RAM random access memory
- nonvolatile storage nonvolatile storage.
- Other hardware conventional and/or custom, may also be included.
- any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.
- circuit may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and/or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and/or digital hardware circuit(s) with software/firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.”
- This definition of circuitry applies to all uses of this term in this application, including in any claims.
- circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware.
- circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Image Generation (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Description
Claims
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263386559P | 2022-12-08 | 2022-12-08 | |
| EP23166550 | 2023-04-04 | ||
| PCT/US2023/082743 WO2024123915A1 (en) | 2022-12-08 | 2023-12-06 | Offset low discrepancy spherical sampling for image rendering |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4631249A1 true EP4631249A1 (en) | 2025-10-15 |
Family
ID=89508966
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23837518.2A Pending EP4631249A1 (en) | 2022-12-08 | 2023-12-06 | Offset low discrepancy spherical sampling for image rendering |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4631249A1 (en) |
| JP (1) | JP2026503369A (en) |
| CN (1) | CN120584491A (en) |
| WO (1) | WO2024123915A1 (en) |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109511284B (en) * | 2016-05-26 | 2023-09-01 | Vid拓展公司 | Method and device for window-adaptive 360-degree video transmission |
| WO2019083943A1 (en) * | 2017-10-24 | 2019-05-02 | Vid Scale, Inc. | Hybrid angular cubemap projection for 360-degree video coding |
| US10460509B2 (en) * | 2017-11-07 | 2019-10-29 | Dolby Laboratories Licensing Corporation | Parameterizing 3D scenes for volumetric viewing |
-
2023
- 2023-12-06 JP JP2025531804A patent/JP2026503369A/en active Pending
- 2023-12-06 EP EP23837518.2A patent/EP4631249A1/en active Pending
- 2023-12-06 CN CN202380093348.7A patent/CN120584491A/en active Pending
- 2023-12-06 WO PCT/US2023/082743 patent/WO2024123915A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| CN120584491A (en) | 2025-09-02 |
| JP2026503369A (en) | 2026-01-29 |
| WO2024123915A1 (en) | 2024-06-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20220414823A1 (en) | Virtual reality cinema-immersive movie watching for headmounted displays | |
| US11849104B2 (en) | Multi-resolution multi-view video rendering | |
| US11210838B2 (en) | Fusing, texturing, and rendering views of dynamic three-dimensional models | |
| US10839591B2 (en) | Stereoscopic rendering using raymarching and a virtual view broadcaster for such rendering | |
| US10346950B2 (en) | System and method of capturing and rendering a stereoscopic panorama using a depth buffer | |
| EP3130143B1 (en) | Stereo viewing | |
| US7656403B2 (en) | Image processing and display | |
| CN112470484A (en) | Partial shadow and HDR | |
| EP3564905A1 (en) | Conversion of a volumetric object in a 3d scene into a simpler representation model | |
| US20250037356A1 (en) | Augmenting a view of a real-world environment with a view of a volumetric video object | |
| US10891711B2 (en) | Image processing method and apparatus | |
| EP4631249A1 (en) | Offset low discrepancy spherical sampling for image rendering | |
| JP7556352B2 (en) | Image characteristic pixel structure generation and processing | |
| GB2638245A (en) | Telepresence system | |
| WO2025172728A1 (en) | Telepresence system | |
| CN118648284A (en) | Volumetric immersive experience with multiple perspectives | |
| TW202046716A (en) | Image signal representing a scene | |
| HK1233091B (en) | Stereo viewing |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250616 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| P01 | Opt-out of the competence of the unified patent court (upc) registered |
Free format text: CASE NUMBER: UPC_APP_0012842_4631249/2025 Effective date: 20251111 |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |