WO2025214613A1 - Warp support for 2-dimensional (2d) video coding - Google Patents
Warp support for 2-dimensional (2d) video codingInfo
- Publication number
- WO2025214613A1 WO2025214613A1 PCT/EP2024/060046 EP2024060046W WO2025214613A1 WO 2025214613 A1 WO2025214613 A1 WO 2025214613A1 EP 2024060046 W EP2024060046 W EP 2024060046W WO 2025214613 A1 WO2025214613 A1 WO 2025214613A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- graphics
- objects
- graphics region
- network node
- region
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/174—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a slice, e.g. a line of blocks or a group of blocks
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/597—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding specially adapted for multi-view video sequence encoding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/85—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using pre-processing or post-processing specially adapted for video compression
- H04N19/86—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using pre-processing or post-processing specially adapted for video compression involving reduction of coding artifacts, e.g. of blockiness
Definitions
- the present disclosure relates generally to extended Reality (XR) applications, and more particularly, to a process for avoiding disocclusion artifacts when encoding Composite Video Frames (CVFs).
- XR extended Reality
- CVFs Composite Video Frames
- Warp is a 2D image postprocessing technique for modifying a rendered 2D image of a 3D scene prior to displaying the image on a Head Mounted Device (HMD).
- HMD Head Mounted Device
- the warp process compensates for the effects of user motion (e.g., head movement) and/or object motion (e.g., the movement of an object through the scene) in situations where there is a time difference between the 3D rendering and displaying the image.
- some current warp techniques transform the entire 2D image.
- other more sophisticated warp techniques use a plurality of previous 2D images to predict the movement of an object in the image.
- the present disclosure provides a method and corresponding network node for supporting a warp function at a Head mounted Device (HMD) by configuring the network node to reduce or prevent disocclusion artifacts when encoding Composite Video Frames (CVFs).
- HMD Head mounted Device
- CVFs Composite Video Frames
- embodiments of the present disclosure provide a method, implemented by a network node, for avoiding disocclusion artifacts in CVFs.
- the method comprises the network node receiving first and second 2D graphics regions, with each of the first and second 2D graphics regions comprising an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object.
- the method further comprises the network node receiving, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters.
- the method further comprises the network node determining an un-occluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects.
- the method comprises the network node generating, based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region.
- the method then comprises the network node encoding a CVF to include the modified 2D graphics region and sending the encoded CVF to a decoder in a bitstream.
- the present embodiments provide a network node configured to avoid disocclusion artifacts in Composite Video Frames (CVFs).
- the network node is configured to receive first and second 2D graphics regions, wherein each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object, receive, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters, determine an unoccluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects, generate, based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region, encode a CVF to include the modified 2D graphics region, and send the encoded CVF to a decoder in a bitstream
- CVFs Composite Video
- the present embodiments provide a network node configured to avoid disocclusion artifacts in Composite Video Frames (CVFs).
- the network node comprises communications circuitry configured to communicate with a client device and processing circuitry operatively connected to the communications circuitry.
- the processing circuitry in this aspect is configured to receive first and second 2D graphics regions, wherein each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object, receive, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters, determine an un-occluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects, generate, based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region, encode a CVF to include the modified 2D graphics region, and send the encoded CVF to a decoder in a bitstream.
- the present embodiments provide a non-transitory computer-readable storage medium for a network node.
- the non-transitory computer-readable storage medium in this aspect comprises a computer program stored thereon.
- the computer program comprises executable instructions that, when executed by processing circuitry in the network node, causes the network node to receive first and second 2D graphics regions, wherein each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object, receive, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters, determine an un- occluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects, generate, based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics
- Figure 1 is a functional block diagram illustrating a communications system configured according to one embodiment of the present disclosure.
- Figure 2 illustrates the encoding of a Composite Video Frame (CVF) into a 2D image for display at a client device.
- CVF Composite Video Frame
- Figures 3A-3B illustrate some exemplary disocclusion artifacts that are addressed according to embodiments of the present disclosure.
- Figure 4 is a flow diagram illustrating a method for avoiding disocclusion artifacts at an encoder to support warp processing at a client device (e.g., a HMD) according to embodiments of the present disclosure.
- a client device e.g., a HMD
- Figures 5A-5B Illustrate the avoidance of disocclusion artifacts in situations where 2D graphics regions overlap according to embodiments of the present disclosure.
- Figure 6 illustrates a grouping (e.g., a slice group) of overlapping 2D graphics regions according to with embodiments of the present disclosure.
- Figures 7A-7B are flow diagrams illustrating a method for avoiding disocclusion artifacts in CVFs and for supporting a warp process at a client device (e.g., a HMD) according to embodiments of the present disclosure.
- a client device e.g., a HMD
- Figure 8 is a functional block diagram illustrating some of the components of a network node configured to support a warp process at a client device (e.g., a HMD) by reducing or preventing disocclusion artifacts when encoding CVFs according to embodiments of the present disclosure.
- a client device e.g., a HMD
- Figure 9 is a functional block diagram illustrating some of the components of a client device (e.g., a HMD) configured to reduce or prevent disocclusion artifacts in CVFs according to embodiments of the present disclosure.
- a client device e.g., a HMD
- Embodiments of the present disclosure provide a method for supporting the warp process at a Head mounted Device (HMD) by configuring a computing device, such as a network node disposed in the Cloud, for example, to reduce or prevent “disocclusion” artifacts when encoding Composite Video Frames (CVFs).
- HMD Head mounted Device
- CVFs Composite Video Frames
- the processing logic for whichever warp technique is being employed e.g., a positional time-warp technique or an asynchronous space warp technique
- needs something to fill the “empty” space the object leaves behind i.e., the previously occluded region of the object.
- the present embodiments configure a computing device to receive a plurality of 2D graphics regions - each of which comprises an image of a 3D object in a 3D scene.
- the first 3D object in the first 2D graphics region at least partially occludes the second 3D object in the second 2D graphics region.
- the present embodiments also configure the computing device to receive a corresponding set of 3D coordinates and motion parameters for each 3D object. Then, based on the received information, the computing device identifies a disoccluded area of the second 2D graphics region.
- the identified disoccluded area comprises an area or region of the second 3D object that was previously occluded (i.e., hidden) by the first 3D object, but is now disoccluded (i.e., visible) due to the motion of one or both of the first and second 3D objects relative to each other and/or the movement of a user’s head.
- the computing device Once determined, the computing device generates a modified 2D graphics region based on the motion parameters of the first and second 3D objects and an estimated display latency (i.e., an estimated time-period between when a simulation engine (SE), such as a game engine, for example, renders an image and when the resultant image is updated on a display of the client device).
- SE simulation engine
- the present embodiments configure the computing device to insert at least part of the previously occluded area of the second 3D object into the now disoccluded area of the corresponding 2D graphics region. So generated, the computing device encodes a CVF to include the modified 2D graphics region and sends the encoded CVF, along with the 3D coordinates, motion information, and latency information, in a bitstream to a decoder at the HMD.
- FIG. 1 illustrates a communications network 10 configured according to one embodiment of the present disclosure.
- network 10 comprises an access network 12 communicatively connecting a client device 300 (e.g., a HMD) with a network node 200 (e.g., a server node) disposed in a cloud network 14.
- a computing device 20 is disposed between client device 300 and network node 200, and is configured to perform at least some of the processing functions of client device 300.
- the access network 12 may be any type of communications network (e.g., WiFi, ETHERNET, Wireless LAN (WLAN), 3G, LTE, etc.), and functions to connect subscriber devices, such as client device 300, to one or more service provider nodes, such as network node 200.
- the cloud network 14 provides such subscriber devices with “on-demand” availability of computer resources (e.g., memory, data storage, processing power, etc.) without requiring the user to directly, actively manage those resources.
- such resources include, but are not limited to, one or more XR applications being executed on network node 200.
- the XR applications may comprise, for example, gaming applications and/or simulation applications used for training.
- one or more sensors (not shown) on client device 300 measure the translational and/or rotational movement of the user’s head as the user views images rendered by network node 200 on client device 300. Signals representing the detected and measured movement are then sent to network node 200. Upon receipt, network node 200 utilizes those signals to compensate the images for the user’s movement, and sends the compensated images to client device 300. In some embodiments, to help reduce and/or eliminate latency associated with the communications between client device 300 and network node 200, the present embodiments may place network node 200 in an Edge Data Network (EDN).
- EDN Edge Data Network
- warp 2D image postprocessing technique
- a client device 300 such as a HMD
- warp processes compensate for the effects of user motion (e.g., head movement) and object movement (e.g., through a scene) in cases where there is a time difference between the 3D rendering of an image containing the object at a SE operating on network node 200 and the displaying of the image at the client device 300.
- This time difference typically occurs in connection with so-called “remote rendering” schemes, in which the processing and rendering of 3D scenes occur in a remote cloud processing environment rather than at the client device.
- remote rendering the resources are more abundant at network node 200 than they are at the client device 300, this time difference can, unfortunately, be quite problematic as it gives rise to latency that the user will experience in the form of various undesirable “visual artifacts.”
- a CVF is encoded into a single bitstream (e.g., for video) and compressed with a state-of-the-art video encoder.
- a second technique illustrated in Figure 2 as technique 30, multiple arbitrarily shaped objects (e.g., 34a, 34b) can also be encoded into a CVF 30.
- the creation of the final 2D image is described by a 3D scene descriptor such as the Virtual Reality Modeling Language (VRML), Moving Picture Experts Group (MPEG-4) Binary Format for Scenes (BIFS), and the like.
- VRML Virtual Reality Modeling Language
- MPEG-4 Moving Picture Experts Group
- BIFS Binary Format for Scenes
- CVFs are encoded as 2D frames.
- a CVF 40 comprises three 2D graphics regions 42, 44, 46 having respective Z-orders of 0, 1 , and 2.
- Each graphics region 42, 44, 46 is an image of an object in a 3D scene.
- 2D graphics regions 44 and 46 are moving relative to each other and to 2D graphics region 42 at respective velocities vi and v 2 .
- 2D graphics region 42 occludes a portion of 2D graphics region 44 (i.e., a portion of the object in graphics region 44), and 2D graphics region 44 occludes a portion of 2D graphics region 46 (i.e., a portion of the object in graphics region 46).
- the movement of 2D graphics regions 44 and 46 causes those previously occluded portions of 2D graphics regions 44, 46 (i.e., the objects in 2D graphics regions 44, 46) to become disoccluded (i.e., visible).
- “empty” areas of space 48, 50 are introduced into the 2D CVF frame 40, with which the warp functionality performed at the client device must contend.
- full 2D graphics regions are encoded into a separate video stream, including the overlapped part of the regions (e.g., the occluded areas of 2D graphics regions 44, 46). Further, while each stream contains its own overhead (e.g. header information), additional multiplex overhead (e.g., 3D scene description) is transmitted describing how the streams are related. The streams are then multiplexed with the additional multiplex overhead for transmission to the client device.
- overhead e.g. header information
- additional multiplex overhead e.g., 3D scene description
- the present disclosure provides a method, and configures a corresponding computing device to, reduce the overhead used in warp processing and to increase the quality of object-based warp.
- embodiments of the present disclosure configure a computing device, such as a network node in the cloud, for example, to encode the overlapping areas of 2D graphics regions (i.e., the objects in a 3D scene) into a single video stream.
- a computing device such as a network node in the cloud
- embodiments of the present disclosure allocate a RegionlD for each 2D graphics region based on the Z-order and the ObjectID of the object in the 2D graphics region.
- the present embodiments then calculate the un-occluded area of the 2D graphics region (e.g., an area in the 2D graphics region that was previously occluded by the object in another 2D graphics region but is now visible due to movement). Then, based on motion parameters and an expected “latency-to-display” (i.e., the estimated time between the 3D rendering of the 2D graphics region at the SE and the displaying of 2D graphics region at the client device), the present embodiments add some of the previously occluded area of the 2D graphics region back in to the newly disoccluded area of the 2D graphics region.
- an expected “latency-to-display” i.e., the estimated time between the 3D rendering of the 2D graphics region at the SE and the displaying of 2D graphics region at the client device
- embodiments of the present disclosure allocate the newly disoccluded 2D graphics region to a “slice group.”
- a single pixel position in the newly disoccluded areas of an object in a 2D graphics region can contain different information for different, overlapping objects across multiple 2D graphics regions. Therefore, embodiments of the present disclosure extend the slice group syntax to incorporate the RegionlD.
- the RegionlD is a parameter generated from an object’s (i.e., 2D graphics region’s) own Z-order and ObjectID. It’s use, according to the present embodiments, allows a downstream warping process to uniquely identify and distinguish between the information belonging to multiple objects (i.e., multiple 2D graphics regions) in any given pixel position.
- the slice groups are further encodable using state-of-the art video encoding tools that are extended to support both the RegionlD and the overlapping regions in a slice.
- embodiments of the present disclosure add additional information to a frame.
- extra information may include, but is not limited to, the Z-order information for an object (i.e., 2D graphics region), the 3D coordinates of the object, motion parameters indicating the motion (or non-motion) of the object, and the like.
- the present embodiments provide advantages and benefits that conventional systems and methods cannot or do not provide. For example, as will be explained in more detail below, the present embodiments reduce the amount of overhead while also increasing the quality of object-based warp.
- the present embodiments encode the 2D graphics regions into a single 2D image with overlapped areas, and further, transmit only the newly disoccluded areas of the graphics regions.
- the size of the newly disoccluded area of the graphics region is calculated based on the motion of a 3D object as originally captured and the expected latency-to-display (i.e., the estimated time-period between when the 2D graphics region is 3D rendered and when the 2D graphics region is to be displayed at a client device). This helps avoid disocclusion artifacts in warp.
- the present embodiments reduce computational complexity by using only a single encoder and decoder context instead of using 3D scene rendering based on 3D scene descriptions.
- the present embodiments further extend the capabilities of conventional video encoding methodologies to handle overlapped areas in 2D graphics regions, thereby again lowering complexity. This lower complexity equates to lower computation times, which in turn, decreases latency and allows for the use of smaller warp processes that can handle overlapped areas of a graphics region.
- the present embodiments avoid disocclusion artifacts by adding a previously occluded part of an object in a 2D graphics region back into the newly disoccluded area of the object in the 2D graphics region.
- the particular disoccluded area of the 2D graphics region is calculated based on the 3D motion of the object and an expected “latency-to-display,” which as defined above, is the estimated period of time between when the 2D graphics region is 3D rendered by the simulation engine and when the 2D graphics region is to be displayed at the client device.
- an expected “latency-to-display,” which as defined above is the estimated period of time between when the 2D graphics region is 3D rendered by the simulation engine and when the 2D graphics region is to be displayed at the client device.
- embodiments of the present disclosure extend, for example, video encoding and decoding methodologies to enable the encoding and decoding of overlapped 2D regions in a given slice and extends the syntax to handle the RegionlD of the present disclosure.
- FIG 4 is a flow diagram illustrating a method 60 for avoiding disocclusion artifacts to support warp processing at the client device 300 according to embodiments of the present disclosure.
- method 60 is implemented by a simulation engine (SE) executing on network node 200 in cloud network 14, and more particularly, by Renderer functionality of the SE.
- SE simulation engine
- Renderer functionality of the SE Renderer functionality of the SE.
- this is for illustrative purposes only, and that method 60 may be implemented by the renderer functions of an SE that is executing on a node that is different than network node 200, such as computing device 20 seen in Figure 1 , for example.
- the system receives input from the simulation processing performed by the SE.
- the input includes a plurality of separate 2D graphics regions generated by the SE from 3D objects captured in an image by a camera, for example.
- each 2D graphics region has a 3D object and is augmented with information from the 3D simulation.
- information may comprise, for example, an object identifier (i.e., ObjectID) for the 3D object in the 2D graphics region, the 3D Coordinate information (e.g., the Z-layer or Z-order information) for the 3D object in the 2D graphics region, and the motion parameters (e.g., the velocity and direction of motion) for the 3D object in the 2D graphics region.
- embodiments of the present disclosure first sort the 2D graphics regions by Z layer (i.e., Z-order) and distance of the 3D objects in those 2D graphics regions from the camera (box 62).
- the sorting begins with the 2D graphics region having the lowest Z-order (i.e., that was closest to the camera when the image was captured).
- an object is selected (box 64).
- a RegionlD is then allocated to the 2D graphics region having the selected object based on the Z-order and the ObjectID of the selected object (box 66).
- the present embodiments use the same RegionlD for the same 2D graphics region and selected object so as to follow the selected object and the 2D graphics region over successive video frames.
- the present embodiments calculate the un-occluded area(s) of the 2D graphics region using the Z-order and 3D coordinate information received with the 2D graphics region (box 68). At least one un-occluded area of the 2D graphics region coincides with an area of the object in the 2D graphics region that was previously hidden by another object in a different 2D graphics region, but is now visible due to the movement of one or both of those objects relative to one another.
- the present embodiments “extend” the un-occluded areas of the 2D graphics region (box 70). For example, as was explained above in Figure 3 above, the movement of 2D graphics regions 44 and 46 relative to each other reveals corresponding “empty” spaces 48, 50, which must be filled-in prior to display at the client device 300.
- the SE knows the characteristics and parameters of the occluded areas of the object in the 2D graphics region. Therefore, in this embodiment, the Renderer functionality at the SE extends the un-occluded areas of the 2D graphics region by adding or inserting previously occluded parts of the object back into the now un-occluded area of that object based on the motion parameters and the expected latency-to-display related to the 2D graphics region. This allows the warp processing functions to avoid dis-occlusion artifacts.
- the present embodiments add the 2D graphics region to a “slice group” (box 72). So added, the present embodiments repeat the process until no more objects can be selected (box 74). Once the last object has been processed, the present embodiments encode the frame (e.g., a CVF) (box 76) and then add the encoded frame (box 78), along with other information as stated above, into a bitstream to send to a decoder at the client device 300.
- a CVF e.g., a CVF
- disocclusion is an artifact caused by an object-based warp algorithm when it moves objects/2D graphics regions out of the way of another object/2D graphics region.
- the present embodiments “extend” the 2D graphics regions to include their own previously occluded parts that, due to movement, have become, or are expected to become, visible.
- Figures 5A-5B illustrate how the Renderer function at the SE “extends” the un-occluded areas of the 2D graphics region according to embodiments of the present disclosure.
- Figures 5A-5B illustrate a frame 80 having 2D graphics regions 82, 84, 86. Each comprising a respective 2D object.
- 2D graphics region 82 is un- occluded (i.e., entirely visible).
- 2D graphics region 82 occludes a part 88 of 2D graphics region 84, which in turn, occludes a part 90 of 2D graphics region 86. Due to the motion of 2D graphics regions 84, 86 relative to the other 2D graphics regions (identified by vectors Vi, v 2 , respectively), some or all of these previously occluded parts 88, 90 become visible.
- these previously occluded parts appear as empty spaces 48, 50 (i.e., Figures 3A-3B).
- Conventional systems handle filling in these so-called empty spaces at the warp processing functions. However, doing so taxes the warp processing functionality at the client device 300 and undesirably increases the need for more available resources. Therefore, according to the present embodiments, the Renderer function at the SE fills-in (i.e., extends) these empty spaces 48, 50 by adding or inserting at least some of the previously occluded parts 92, 94 of 2D graphics regions 84, 86 into those empty spaces 48, 50. Therefore, when the frame 40 arrives for warp processing at the client device 300, the warp function need not “fill-in” the empty spaces 48, 50. Rather, they would already be filled-in by the entity that has the most knowledge of 2D graphics regions 82, 84, and 86.
- the size of the part that is being added into the empty space created by disocclusion depends on a number of parameters.
- One such parameter is the expected latency-to-display (i.e., the time between the simulation/renderer processing at the SE and display at the client device 300).
- Another parameter is the motion of the object/2D graphics region relative to one or more other objects/2D graphics regions.
- slice groups can have different information associated with the same pixel position. Accordingly, the present embodiments extend the slice syntax to include the previously-described RegionlD. With this parameter, each pixel position can be unambiguously identified.
- each slice group 110, 120, 130 is a corresponding 2D graphics region (e.g., 2D graphics regions 42, 44, 46) and comprises a set of Coding Units (CUs) 112, 122, 132, respectively, defined by a CU-to-slice-group-map 100.
- CUs 112, 122, 132 in this case can cover any part of the current frame.
- CUs 112, 122, 132 cover the non-overlapping areas of their respective slice group 110, 120, 130.
- CUs 140 and 150 cover the overlapping areas of slice groups 110, 120 and 120, 130, respectively.
- CUs 112, 122, and 132 may contain information for their respective slice group 110, 120, 130, respectively, However, because they at least partially overlap, CUs 140 may contain information for both groups 110, 120, while CUs 150 may contain information for both groups 120, 130.
- the RegionlD allows the present embodiments to identify and match the information in a given CU 140 or 150 to its appropriate slice group 110, 120, 130.
- this information is added to the bitstream being sent to the decoder for each frame.
- the information may be sent as a special Network Abstraction Layer (NAL) unit having a defined syntax, or it may be added as user data to a NAL unit.
- NAL Network Abstraction Layer
- Figures 7A-7B are flow diagrams illustrating a method 160 for avoiding disocclusion artifacts in CVFs and supporting the warp process at the client device 300 according to embodiments of the present disclosure.
- method 160 is implemented by the Renderer functionality at an SE, which for illustrative purposes only, executes on network node 200.
- the Renderer function at the SE receives first and second 2D graphics regions.
- Each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object (box 162).
- the Renderer function also receives, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters (box 164).
- the Renderer function then sorts the first and second 2D graphics regions based on a Z-order of the first and second 3D objects (box 166). In one embodiment, for example, the Z-order of the first and second 3D objects is based on the 3D coordinates of the first and second 3D objects.
- a 3D object i.e., a 2D graphics region
- a RegionlD is allocated to the 2D graphics region (box 170).
- this embodiment of the present disclosure first generates the RegionlD based on the Z-order and the Object ID of the selected 3D object (box 172).
- the present disclosure may concatenate the Z-order and Object ID.
- other information may be used in addition to, or in lieu of, the Z-order and/or the ObjectID when generating the RegionlD as needed or desired.
- the RegionlD is associated with the 2D graphics region and the 3D object it represents (box 174).
- the RegionlD will allow a warp function at client device 300 to distinguish between the pixel information of two different overlapping 3D objects represented in corresponding 2D graphics regions. Then, based on the 3D coordinates of each of the first and second 3D objects, the Renderer function at the SE determines an un-occluded area of the 2D graphics region (box 176). These so-called un-occluded areas of the first and second 3D objects are the “empty” spaces 48, 50 seen in Figures 3A-3B.
- the Renderer function of the SE generates a modified 2D graphics region (box 178). For example, in one embodiment, the Renderer function inserts at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region (i.e., the now un-occluded part of the second 3D object). The insertion is based on the motion parameters of each of the first and second 3D objects and the estimated display latency (i.e., the expected latency-to-display). So generated, the Renderer function groups the modified 2D graphics region into a corresponding slice group comprising a plurality of slices (box 180).
- Each slice in the slice group comprises a sequence of Coding Units (CUs) and is decodable independently of the other slices in the slice group. As previously described, each CU is also decodable independently of the other CUs in the sequence.
- the Renderer function then checks to see whether any more objects (i.e., 2D graphics regions) exist (box 182). If the Renderer function determines that additional objects/2D graphics regions exist, processing returns to box 168 in Figure 7A.
- the CVF is encoded (box 184) to include the modified 2D graphics region, the 3D coordinates, the motion parameters, and other information (e.g., Z-order information) as needed or desired, and sent in a bitstream to a decoder at client device 300 (box 186).
- Figures 7A-7B described determining the un-occluded area of the 3D object represented by a 2D graphics region as being based on the 3D coordinates of the first and second 3D objects.
- the present disclosure is not so limited.
- the un-occluded area of the 2D graphics region is further based on the Z-order of each of the first and second 3D objects.
- first and second 2D graphics regions representing the first and second 3D objects overlap each other.
- a first CU in an overlap area of the first 2D graphics region comprises image data for the first 3D object
- a second CU in the overlap area of the second 2D graphics region comprises image data for the second 3D object.
- the first and second CUs are associated with a same position of the overlap area.
- the RegionlD distinguishes between information associated with a corresponding slice group and information associated with one or more other slice groups.
- the RegionlD further associates the 3D coordinates of a 3D object represented by a 2D graphics region with the motion parameters of the 3D object represented by the 2D graphics region.
- the un-occluded area of the 2D graphics region comprises an un-occluded area of the second 3D object that was previously occluded by the first 3D object.
- a size of the un-occluded area of the second 3D object is based on the estimated display latency for the first and second 3D objects, and the movement of the first and second 3D objects relative to each other.
- the estimated display latency is an estimated time between when the first and second 3D objects are rendered at the simulation engine and when the first and second 3D objects are expected to be displayed on a display of a client device.
- the Z-order, the 3D coordinates, and the motion parameters of the modified 2D graphics region and each of the first and second 3D objects is encoded for each frame.
- the Z-order, the 3D coordinates, and the motion parameters of the modified 2D graphics region is encoded into a user data Network Abstraction Layer (NAL) unit.
- NAL Network Abstraction Layer
- an apparatus can perform any of the methods herein described by implementing any functional means, modules, units, or circuitry.
- the apparatuses comprise respective circuits or circuitry configured to perform the steps shown in the method figures.
- the circuits or circuitry in this regard may comprise circuits dedicated to performing certain functional processing and/or one or more microprocessors in conjunction with memory.
- the circuitry may include one or more microprocessors or microcontrollers, as well as other digital hardware, which may include Digital Signal Processors (DSPs), special-purpose digital logic, and the like.
- the processing circuitry may be configured to execute program code stored in memory, which may include one or several types of memory such as read-only memory (ROM), random-access memory, cache memory, flash memory devices, optical storage devices, etc.
- Program code stored in memory may include program instructions for executing one or more telecommunications and/or data communications protocols as well as instructions for carrying out one or more of the techniques described herein, in several embodiments.
- the memory stores program code that, when executed by the one or more processors, carries out the techniques described herein.
- FIG 8 is a functional block diagram illustrating some of the components of a network node 200 configured to avoid disocclusion artifacts and support the warp functionality at client device 300 according to embodiments of the present disclosure.
- the network node 200 in this embodiment is configured to execute the SE and comprises, inter alia, communication circuitry 202, processing circuitry 204, and memory 206.
- the communication circuitry 202 comprises both radio frequency (RF) circuitry 202a and network interface circuitry (NIC) 202b.
- the network node may comprise only NIC 202b.
- the RF circuitry 202a can be located at one or more TRPs and comprises the RF components necessary for communicating with various client devices 300, directly or indirectly, via a wireless communication link.
- the RF circuitry 202a may comprise, for example, a transmitter and receiver configured to operate according to the 5G standards or other wireless communication standard.
- the communication circuitry 202 also comprises network interface circuitry (e.g., NIC 202b) for communication with other RAN nodes, OA&M nodes, core network nodes, and/or other nodes in external systems.
- the network interface circuitry 202b in this regard may, for example, comprise an Ethernet interface, optical network interface, or a wireless interface.
- the processing circuitry 204 comprises one or more microprocessors, hardware, firmware, or a combination thereof that controls the overall operation of the network node 200.
- the processing circuitry 204 in this regard can be configured by software to perform the functionality described with respect to Figures 4 and 7A-7B, and one or more of the methods herein described, including methods 60 and 160 as shown in Figures 4 and 7A-7B, respectively.
- Memory 206 comprises both volatile and non-volatile memory for storing computer program code and data needed by the processing circuitry 204 for operation.
- Memory 206 may comprise any tangible, non-transitory computer-readable storage medium for storing data including electronic, magnetic, optical, electromagnetic, or semiconductor data storage.
- Memory 206 stores one or more computer programs 208 comprising executable instructions that configure the processing circuitry 204 in the network node 200 to perform one or more of the methods herein described, including methods 60 and 160 as shown in Figures 4 and 7A-7B, respectively.
- a computer program 208 in this regard may comprise one or more code modules corresponding to the means or units described above.
- computer program instructions and configuration information are stored in a non-volatile memory, such as a ROM, erasable programmable read only memory (EPROM) or flash memory. Temporary data generated during operation may be stored in a volatile memory, such as a random access memory (RAM).
- computer program 208 for configuring the processing circuitry 204 as herein described may be stored in a removable memory, such as a portable compact disc, portable digital video disc, or other removable media.
- the computer program 208 may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer readable storage medium.
- a computing device in these embodiments may, for example, comprise dedicated hardware circuitry, such as one or more graphics processing units (GPU), and be capable of executing games and other software programs associated with an XR environment.
- GPU graphics processing units
- Figure 9 is a functional block diagram illustrating some of the components of a client device 300 configured to avoid disocclusion artifacts and support warp processing at client device 300 according to embodiments of the present disclosure.
- the client device 300 in this embodiment may be, for example, a HMD worn by a user or a computing device capable of executing the SE. Regardless, though, the client device 300 in this embodiment is configured to perform these functions and comprises, inter alia, communication circuitry 302, processing circuitry 304, and memory 306.
- the communication circuitry 302 in this embodiment comprises both radio frequency (RF) circuitry 302a and network interface circuitry (NIC) 302b.
- client device 300 may comprise only the NIC 302b.
- the RF circuitry 302a can be located at one or more TRPs and comprises the RF components necessary for communicating with network node 200 and/or other devices over a wireless communication link.
- the RF circuitry 302a may comprise, for example, a transmitter and receiver configured to operate according to the 5G standards or other wireless communication standard.
- the communication circuitry 302 also comprises network interface circuitry (e.g., NIC 302b) for communication with other RAN nodes, OA&M nodes, core network nodes, and/or other nodes in external systems.
- the network interface circuitry 302b in this regard may, for example, comprise an ETHERNET interface, optical network interface, or a wireless interface.
- the processing circuitry 304 comprises one or more microprocessors, hardware, firmware, or a combination thereof that controls the overall operation of the client device 300.
- the processing circuitry 304 in this regard can be configured by software to perform the functionality described with respect to Figures 4 and 7A-7B, and one or more of the methods herein described, including methods 60 and 160 as shown in Figures 4 and 7A-7B, respectively.
- Memory 306 comprises both volatile and non-volatile memory for storing computer program code and data needed by the processing circuitry 304 for operation.
- Memory 306 may comprise any tangible, non-transitory computer-readable storage medium for storing data including electronic, magnetic, optical, electromagnetic, or semiconductor data storage.
- Memory 306 stores one or more computer programs 308 (e.g., an SE) comprising executable instructions that configure the processing circuitry 304 in the client device 300 (or other computing device) to perform one or more of the methods herein described, including methods 60 and 160 as shown in Figures 4 and 7A-7B, respectively.
- a computer program 308 in this regard may comprise one or more code modules corresponding to the means or units described above.
- computer program instructions and configuration information are stored in a non-volatile memory, such as a ROM, erasable programmable read only memory (EPROM) or flash memory. Temporary data generated during operation may be stored in a volatile memory, such as a random access memory (RAM).
- computer programs 308 for configuring the processing circuitry 304 as herein described may be stored in a removable memory, such as a portable compact disc, portable digital video disc, or other removable media.
- the computer program 308 may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer readable storage medium.
- the present embodiments may, of course, be carried out in other ways than those specifically set forth herein without departing from characteristics described herein. The present embodiments are therefore to be considered in all respects as illustrative and not restrictive, and all changes coming within the meaning and equivalency range of the appended claims are intended to be embraced therein.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Processing Or Creating Images (AREA)
Abstract
A computing device (200) and method (160) is provided for reducing or preventing "disocclusion" artifacts in Composite Video Frames, and for supporting warp functionality at a Head mounted Device (HMD). The computing device receives (162) first and second 2D graphics regions, each representing a respective 3D object with the 3D object represented in the first 2D graphics region overlapping the 3D object represented in the second 2D graphics region. The device also receives (164), for each 3D object, corresponding 3D coordinates and motion parameters. Based on the 3D coordinates of each 3D object, the device determines (176) an un-occluded area of the second 2D graphics region. Then, based on the motion parameters of each of the first and second 3D objects and an estimated display latency, the device generates (178) a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region. The device then encodes (184) a CVF to include the modified 2D graphics region and sends (186) the encoded CVF to a decoder in a bitstream.
Description
WARP SUPPORT FOR 2-DIMENSIONAL (2D) VIDEO CODING
TECHNICAL FIELD
The present disclosure relates generally to extended Reality (XR) applications, and more particularly, to a process for avoiding disocclusion artifacts when encoding Composite Video Frames (CVFs).
BACKGROUND
Warp is a 2D image postprocessing technique for modifying a rendered 2D image of a 3D scene prior to displaying the image on a Head Mounted Device (HMD). As is known in the art, the warp process compensates for the effects of user motion (e.g., head movement) and/or object motion (e.g., the movement of an object through the scene) in situations where there is a time difference between the 3D rendering and displaying the image. To address such situations, some current warp techniques transform the entire 2D image. However, other more sophisticated warp techniques use a plurality of previous 2D images to predict the movement of an object in the image.
SUMMARY
The present disclosure provides a method and corresponding network node for supporting a warp function at a Head mounted Device (HMD) by configuring the network node to reduce or prevent disocclusion artifacts when encoding Composite Video Frames (CVFs).
In a first aspect, embodiments of the present disclosure provide a method, implemented by a network node, for avoiding disocclusion artifacts in CVFs. In this aspect, the method comprises the network node receiving first and second 2D graphics regions, with each of the first and second 2D graphics regions comprising an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object. The method further comprises the network node receiving, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters. The method further comprises the network node determining an un-occluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects. So determined, the method comprises the network node generating, based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region. The method then comprises the network node encoding a CVF to include the modified 2D graphics region and sending the encoded CVF to a decoder in a bitstream.
In a second aspect, the present embodiments provide a network node configured to avoid disocclusion artifacts in Composite Video Frames (CVFs). In this aspect, the network node is configured to receive first and second 2D graphics regions, wherein each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D
object with the first 3D object occluding the second 3D object, receive, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters, determine an unoccluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects, generate, based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region, encode a CVF to include the modified 2D graphics region, and send the encoded CVF to a decoder in a bitstream.
In a third aspect, the present embodiments provide a network node configured to avoid disocclusion artifacts in Composite Video Frames (CVFs). In this aspect, the network node comprises communications circuitry configured to communicate with a client device and processing circuitry operatively connected to the communications circuitry. The processing circuitry in this aspect is configured to receive first and second 2D graphics regions, wherein each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object, receive, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters, determine an un-occluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects, generate, based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region, encode a CVF to include the modified 2D graphics region, and send the encoded CVF to a decoder in a bitstream.
In a fourth aspect, the present embodiments provide a non-transitory computer-readable storage medium for a network node. The non-transitory computer-readable storage medium in this aspect comprises a computer program stored thereon. The computer program comprises executable instructions that, when executed by processing circuitry in the network node, causes the network node to receive first and second 2D graphics regions, wherein each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object, receive, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters, determine an un- occluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects, generate, based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region, encode a CVF to include the modified 2D graphics region, and send the encoded CVF to a decoder in a bitstream.
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1 is a functional block diagram illustrating a communications system configured according to one embodiment of the present disclosure.
Figure 2 illustrates the encoding of a Composite Video Frame (CVF) into a 2D image for display at a client device.
Figures 3A-3B illustrate some exemplary disocclusion artifacts that are addressed according to embodiments of the present disclosure.
Figure 4 is a flow diagram illustrating a method for avoiding disocclusion artifacts at an encoder to support warp processing at a client device (e.g., a HMD) according to embodiments of the present disclosure.
Figures 5A-5B Illustrate the avoidance of disocclusion artifacts in situations where 2D graphics regions overlap according to embodiments of the present disclosure.
Figure 6 illustrates a grouping (e.g., a slice group) of overlapping 2D graphics regions according to with embodiments of the present disclosure.
Figures 7A-7B are flow diagrams illustrating a method for avoiding disocclusion artifacts in CVFs and for supporting a warp process at a client device (e.g., a HMD) according to embodiments of the present disclosure.
Figure 8 is a functional block diagram illustrating some of the components of a network node configured to support a warp process at a client device (e.g., a HMD) by reducing or preventing disocclusion artifacts when encoding CVFs according to embodiments of the present disclosure.
Figure 9 is a functional block diagram illustrating some of the components of a client device (e.g., a HMD) configured to reduce or prevent disocclusion artifacts in CVFs according to embodiments of the present disclosure.
DETAILED DESCRIPTION
Embodiments of the present disclosure provide a method for supporting the warp process at a Head mounted Device (HMD) by configuring a computing device, such as a network node disposed in the Cloud, for example, to reduce or prevent “disocclusion” artifacts when encoding Composite Video Frames (CVFs). Disocclusion is a term commonly understood by those in the graphics arts and describes a situation where a region of one object in an image, previously occluded by another object in the image, becomes visible as the result of its movement relative to the other object (and/or vice versa). Generally, as the previously occluded region of the object becomes visible, the processing logic for whichever warp technique is being employed (e.g., a positional time-warp technique or an asynchronous space warp technique) needs something to fill the “empty” space the object leaves behind (i.e., the previously occluded region of the object). Embodiments of the present disclosure address such situations.
In more detail, the present embodiments configure a computing device to receive a plurality of 2D graphics regions - each of which comprises an image of a 3D object in a 3D
scene. According to the present disclosure, the first 3D object in the first 2D graphics region at least partially occludes the second 3D object in the second 2D graphics region. The present embodiments also configure the computing device to receive a corresponding set of 3D coordinates and motion parameters for each 3D object. Then, based on the received information, the computing device identifies a disoccluded area of the second 2D graphics region. The identified disoccluded area comprises an area or region of the second 3D object that was previously occluded (i.e., hidden) by the first 3D object, but is now disoccluded (i.e., visible) due to the motion of one or both of the first and second 3D objects relative to each other and/or the movement of a user’s head. Once determined, the computing device generates a modified 2D graphics region based on the motion parameters of the first and second 3D objects and an estimated display latency (i.e., an estimated time-period between when a simulation engine (SE), such as a game engine, for example, renders an image and when the resultant image is updated on a display of the client device). More particularly, the present embodiments configure the computing device to insert at least part of the previously occluded area of the second 3D object into the now disoccluded area of the corresponding 2D graphics region. So generated, the computing device encodes a CVF to include the modified 2D graphics region and sends the encoded CVF, along with the 3D coordinates, motion information, and latency information, in a bitstream to a decoder at the HMD.
Turning now to the drawings, Figure 1 illustrates a communications network 10 configured according to one embodiment of the present disclosure. As seen in Figure 1 , network 10 comprises an access network 12 communicatively connecting a client device 300 (e.g., a HMD) with a network node 200 (e.g., a server node) disposed in a cloud network 14. In some embodiments, a computing device 20 is disposed between client device 300 and network node 200, and is configured to perform at least some of the processing functions of client device 300.
The access network 12 may be any type of communications network (e.g., WiFi, ETHERNET, Wireless LAN (WLAN), 3G, LTE, etc.), and functions to connect subscriber devices, such as client device 300, to one or more service provider nodes, such as network node 200. The cloud network 14 provides such subscriber devices with “on-demand” availability of computer resources (e.g., memory, data storage, processing power, etc.) without requiring the user to directly, actively manage those resources. According to embodiments of the present disclosure, such resources include, but are not limited to, one or more XR applications being executed on network node 200. The XR applications may comprise, for example, gaming applications and/or simulation applications used for training.
In general, one or more sensors (not shown) on client device 300 measure the translational and/or rotational movement of the user’s head as the user views images rendered by network node 200 on client device 300. Signals representing the detected and measured movement are then sent to network node 200. Upon receipt, network node 200 utilizes those signals to compensate the images for the user’s movement, and sends the compensated
images to client device 300. In some embodiments, to help reduce and/or eliminate latency associated with the communications between client device 300 and network node 200, the present embodiments may place network node 200 in an Edge Data Network (EDN).
As stated above, many conventional systems use a 2D image postprocessing technique known as “warp” for modifying a rendered 2D image of a 3D scene prior to displaying the image on a client device 300, such as a HMD, for example. As is known in the art, warp processes compensate for the effects of user motion (e.g., head movement) and object movement (e.g., through a scene) in cases where there is a time difference between the 3D rendering of an image containing the object at a SE operating on network node 200 and the displaying of the image at the client device 300. This time difference typically occurs in connection with so-called “remote rendering” schemes, in which the processing and rendering of 3D scenes occur in a remote cloud processing environment rather than at the client device. Although the resources are more abundant at network node 200 than they are at the client device 300, this time difference can, unfortunately, be quite problematic as it gives rise to latency that the user will experience in the form of various undesirable “visual artifacts.”
To address latency and reduce the related issues with warp processing, co-pending applications - i.e., PCT/EP2020/071947 entitled “Improved Split Rendering for Extended Reality (XR) Applications” and filed August 5, 2020, and PCT/EP2022/064366 entitled “Split Transport for Warping” and filed May 26, 2022, both of which are incorporated herein by reference in their entirety - introduce an improved method of split rendering. With this method, the SE generates separate 2D graphics regions from the 3D objects in an image. The graphics regions are then augmented with information, such as Z-layer information and the speed and direction of the motion of the 3D object, from the 3D simulation process performed by the SE. The computing device then encodes the graphics regions into a media stream as a Composite Video Frame (CVF). Warp functions at the client device 300 then move the 2D graphics regions based on object motion and latency.
Currently, there are two main techniques for creating a stream of CVFs. In a first technique, a CVF is encoded into a single bitstream (e.g., for video) and compressed with a state-of-the-art video encoder. In a second technique, illustrated in Figure 2 as technique 30, multiple arbitrarily shaped objects (e.g., 34a, 34b) can also be encoded into a CVF 30. The creation of the final 2D image is described by a 3D scene descriptor such as the Virtual Reality Modeling Language (VRML), Moving Picture Experts Group (MPEG-4) Binary Format for Scenes (BIFS), and the like.
While useful, both techniques can be problematic where warp processing is concerned. For example, in a single video stream, CVFs are encoded as 2D frames. However, as previously described, using a single 2D image as a source for warp processing can cause disocclusion artifacts. Figures 3A-3B illustrate such artifacts that result from such movement. Particularly, as seen in these figures, a CVF 40 comprises three 2D graphics regions 42, 44, 46
having respective Z-orders of 0, 1 , and 2. Each graphics region 42, 44, 46 is an image of an object in a 3D scene. Further, 2D graphics regions 44 and 46 are moving relative to each other and to 2D graphics region 42 at respective velocities vi and v2. As seen in Figure 3A, 2D graphics region 42 (i.e., the object in graphics region 42) occludes a portion of 2D graphics region 44 (i.e., a portion of the object in graphics region 44), and 2D graphics region 44 occludes a portion of 2D graphics region 46 (i.e., a portion of the object in graphics region 46). However, as seen in Figure 3B, the movement of 2D graphics regions 44 and 46 causes those previously occluded portions of 2D graphics regions 44, 46 (i.e., the objects in 2D graphics regions 44, 46) to become disoccluded (i.e., visible). When this occurs, “empty” areas of space 48, 50 are introduced into the 2D CVF frame 40, with which the warp functionality performed at the client device must contend.
By using multiple streams, full 2D graphics regions are encoded into a separate video stream, including the overlapped part of the regions (e.g., the occluded areas of 2D graphics regions 44, 46). Further, while each stream contains its own overhead (e.g. header information), additional multiplex overhead (e.g., 3D scene description) is transmitted describing how the streams are related. The streams are then multiplexed with the additional multiplex overhead for transmission to the client device.
However, with this latter method, multiple layers of overhead are necessarily added to the header data that is normally or typically communicated. Therefore, more information than is necessary is transmitted, thereby undesirably increasing the need for an increased amount of bitrate and/or bandwidth resources. Additionally, the computational complexity is also higher for methods using multiple streams. Particularly, the encoding and decoding of multiple video streams require multiple encoder and decoder sessions. This, at the least, severely taxes, and in some cases may deplete, the limited number of hardware encoder and decoder resources that are available for allocation. Additionally, the multiplexing and demultiplexing processes themselves have computational overhead, as does the rendering of the final 2D image based on the 3D description.
Accordingly, the present disclosure provides a method, and configures a corresponding computing device to, reduce the overhead used in warp processing and to increase the quality of object-based warp. In more detail, embodiments of the present disclosure configure a computing device, such as a network node in the cloud, for example, to encode the overlapping areas of 2D graphics regions (i.e., the objects in a 3D scene) into a single video stream. To distinguish between different objects in those overlapping areas, embodiments of the present disclosure allocate a RegionlD for each 2D graphics region based on the Z-order and the ObjectID of the object in the 2D graphics region. The present embodiments then calculate the un-occluded area of the 2D graphics region (e.g., an area in the 2D graphics region that was previously occluded by the object in another 2D graphics region but is now visible due to movement). Then, based on motion parameters and an expected “latency-to-display” (i.e., the
estimated time between the 3D rendering of the 2D graphics region at the SE and the displaying of 2D graphics region at the client device), the present embodiments add some of the previously occluded area of the 2D graphics region back in to the newly disoccluded area of the 2D graphics region.
To accomplish this function, embodiments of the present disclosure allocate the newly disoccluded 2D graphics region to a “slice group.” However, with the present embodiments, a single pixel position in the newly disoccluded areas of an object in a 2D graphics region can contain different information for different, overlapping objects across multiple 2D graphics regions. Therefore, embodiments of the present disclosure extend the slice group syntax to incorporate the RegionlD. The RegionlD, as stated previously, is a parameter generated from an object’s (i.e., 2D graphics region’s) own Z-order and ObjectID. It’s use, according to the present embodiments, allows a downstream warping process to uniquely identify and distinguish between the information belonging to multiple objects (i.e., multiple 2D graphics regions) in any given pixel position. The slice groups are further encodable using state-of-the art video encoding tools that are extended to support both the RegionlD and the overlapping regions in a slice. In some cases, embodiments of the present disclosure add additional information to a frame. Such extra information may include, but is not limited to, the Z-order information for an object (i.e., 2D graphics region), the 3D coordinates of the object, motion parameters indicating the motion (or non-motion) of the object, and the like.
It should be noted here that the present embodiments provide advantages and benefits that conventional systems and methods cannot or do not provide. For example, as will be explained in more detail below, the present embodiments reduce the amount of overhead while also increasing the quality of object-based warp. To reduce the overhead in both bitrate and bandwidth, the present embodiments encode the 2D graphics regions into a single 2D image with overlapped areas, and further, transmit only the newly disoccluded areas of the graphics regions. The size of the newly disoccluded area of the graphics region is calculated based on the motion of a 3D object as originally captured and the expected latency-to-display (i.e., the estimated time-period between when the 2D graphics region is 3D rendered and when the 2D graphics region is to be displayed at a client device). This helps avoid disocclusion artifacts in warp.
Additionally, the present embodiments reduce computational complexity by using only a single encoder and decoder context instead of using 3D scene rendering based on 3D scene descriptions. The present embodiments further extend the capabilities of conventional video encoding methodologies to handle overlapped areas in 2D graphics regions, thereby again lowering complexity. This lower complexity equates to lower computation times, which in turn, decreases latency and allows for the use of smaller warp processes that can handle overlapped areas of a graphics region.
Further, the present embodiments avoid disocclusion artifacts by adding a previously occluded part of an object in a 2D graphics region back into the newly disoccluded area of the object in the 2D graphics region. The particular disoccluded area of the 2D graphics region is calculated based on the 3D motion of the object and an expected “latency-to-display,” which as defined above, is the estimated period of time between when the 2D graphics region is 3D rendered by the simulation engine and when the 2D graphics region is to be displayed at the client device. Moreover, embodiments of the present disclosure extend, for example, video encoding and decoding methodologies to enable the encoding and decoding of overlapped 2D regions in a given slice and extends the syntax to handle the RegionlD of the present disclosure.
Figure 4 is a flow diagram illustrating a method 60 for avoiding disocclusion artifacts to support warp processing at the client device 300 according to embodiments of the present disclosure. In this embodiment, method 60 is implemented by a simulation engine (SE) executing on network node 200 in cloud network 14, and more particularly, by Renderer functionality of the SE. However, those of ordinary skill in the art will readily appreciate that this is for illustrative purposes only, and that method 60 may be implemented by the renderer functions of an SE that is executing on a node that is different than network node 200, such as computing device 20 seen in Figure 1 , for example.
As seen in Figure 4, the system receives input from the simulation processing performed by the SE. The input includes a plurality of separate 2D graphics regions generated by the SE from 3D objects captured in an image by a camera, for example. According to the present disclosure, each 2D graphics region has a 3D object and is augmented with information from the 3D simulation. Such information may comprise, for example, an object identifier (i.e., ObjectID) for the 3D object in the 2D graphics region, the 3D Coordinate information (e.g., the Z-layer or Z-order information) for the 3D object in the 2D graphics region, and the motion parameters (e.g., the velocity and direction of motion) for the 3D object in the 2D graphics region.
So received, embodiments of the present disclosure first sort the 2D graphics regions by Z layer (i.e., Z-order) and distance of the 3D objects in those 2D graphics regions from the camera (box 62). In this embodiment, the sorting begins with the 2D graphics region having the lowest Z-order (i.e., that was closest to the camera when the image was captured). Once sorted, an object is selected (box 64). A RegionlD is then allocated to the 2D graphics region having the selected object based on the Z-order and the ObjectID of the selected object (box 66). The present embodiments use the same RegionlD for the same 2D graphics region and selected object so as to follow the selected object and the 2D graphics region over successive video frames. Then, the present embodiments calculate the un-occluded area(s) of the 2D graphics region using the Z-order and 3D coordinate information received with the 2D graphics region (box 68). At least one un-occluded area of the 2D graphics region coincides with an area
of the object in the 2D graphics region that was previously hidden by another object in a different 2D graphics region, but is now visible due to the movement of one or both of those objects relative to one another.
Next, the present embodiments “extend” the un-occluded areas of the 2D graphics region (box 70). For example, as was explained above in Figure 3 above, the movement of 2D graphics regions 44 and 46 relative to each other reveals corresponding “empty” spaces 48, 50, which must be filled-in prior to display at the client device 300. However, the SE knows the characteristics and parameters of the occluded areas of the object in the 2D graphics region. Therefore, in this embodiment, the Renderer functionality at the SE extends the un-occluded areas of the 2D graphics region by adding or inserting previously occluded parts of the object back into the now un-occluded area of that object based on the motion parameters and the expected latency-to-display related to the 2D graphics region. This allows the warp processing functions to avoid dis-occlusion artifacts.
Next, the present embodiments add the 2D graphics region to a “slice group” (box 72). So added, the present embodiments repeat the process until no more objects can be selected (box 74). Once the last object has been processed, the present embodiments encode the frame (e.g., a CVF) (box 76) and then add the encoded frame (box 78), along with other information as stated above, into a bitstream to send to a decoder at the client device 300.
As previously described, disocclusion is an artifact caused by an object-based warp algorithm when it moves objects/2D graphics regions out of the way of another object/2D graphics region. To avoid having regions of a 2D image without encoded information becoming visible, the present embodiments “extend” the 2D graphics regions to include their own previously occluded parts that, due to movement, have become, or are expected to become, visible.
Figures 5A-5B, for example, illustrate how the Renderer function at the SE “extends” the un-occluded areas of the 2D graphics region according to embodiments of the present disclosure. Particularly, Figures 5A-5B illustrate a frame 80 having 2D graphics regions 82, 84, 86. Each comprising a respective 2D object. As seen in Figure 5A, 2D graphics region 82 is un- occluded (i.e., entirely visible). However, 2D graphics region 82 occludes a part 88 of 2D graphics region 84, which in turn, occludes a part 90 of 2D graphics region 86. Due to the motion of 2D graphics regions 84, 86 relative to the other 2D graphics regions (identified by vectors Vi, v2, respectively), some or all of these previously occluded parts 88, 90 become visible.
By way of example, these previously occluded parts appear as empty spaces 48, 50 (i.e., Figures 3A-3B). Conventional systems handle filling in these so-called empty spaces at the warp processing functions. However, doing so taxes the warp processing functionality at the client device 300 and undesirably increases the need for more available resources. Therefore, according to the present embodiments, the Renderer function at the SE fills-in (i.e., extends)
these empty spaces 48, 50 by adding or inserting at least some of the previously occluded parts 92, 94 of 2D graphics regions 84, 86 into those empty spaces 48, 50. Therefore, when the frame 40 arrives for warp processing at the client device 300, the warp function need not “fill-in” the empty spaces 48, 50. Rather, they would already be filled-in by the entity that has the most knowledge of 2D graphics regions 82, 84, and 86.
The size of the part that is being added into the empty space created by disocclusion depends on a number of parameters. One such parameter is the expected latency-to-display (i.e., the time between the simulation/renderer processing at the SE and display at the client device 300). Another parameter is the motion of the object/2D graphics region relative to one or more other objects/2D graphics regions.
Further, current video compression algorithms do not support the encoding of overlapped regions, and therefore, according to the present embodiment, are extended to handle the encoding of overlapped regions. To assist in this effort, the present embodiments use “slice group” terminology to describe a 2D graphics region. Slice group terminology is well- known to those of ordinary skill in the art and is taken from the H.264 specification (i.e., ITU H.264 (08/2021): “Advanced video coding for generic audiovisual services”), which is expressly incorporated herein by reference.
However, additional extensions are needed to describe different slices (i.e., 2D regions) that cover, at least partially, the same area of an encoded frame. Particularly, because such extended un-occluded parts overlap, slice groups can have different information associated with the same pixel position. Accordingly, the present embodiments extend the slice syntax to include the previously-described RegionlD. With this parameter, each pixel position can be unambiguously identified.
In more detail, Figure 6 illustrates how the present disclosure incorporates the use of slice groups in avoiding dis-occlusion artifacts according to one embodiment. As seen in Figure 6, each slice group 110, 120, 130 is a corresponding 2D graphics region (e.g., 2D graphics regions 42, 44, 46) and comprises a set of Coding Units (CUs) 112, 122, 132, respectively, defined by a CU-to-slice-group-map 100. CUs 112, 122, 132 in this case can cover any part of the current frame. Thus, CUs 112, 122, 132 cover the non-overlapping areas of their respective slice group 110, 120, 130. CUs 140 and 150, however, cover the overlapping areas of slice groups 110, 120 and 120, 130, respectively. As previously described, CUs 112, 122, and 132 may contain information for their respective slice group 110, 120, 130, respectively, However, because they at least partially overlap, CUs 140 may contain information for both groups 110, 120, while CUs 150 may contain information for both groups 120, 130. The RegionlD, as previously described, allows the present embodiments to identify and match the information in a given CU 140 or 150 to its appropriate slice group 110, 120, 130.
Additionally, extra information, such as the Z-order information, the 3D coordinates, and the motion parameters, for each 2D graphics region and for each related 3D object in a 2D
graphics region is required for an object-based warp algorithm. Therefore, in at least one embodiment, this information is added to the bitstream being sent to the decoder for each frame. For example, the information may be sent as a special Network Abstraction Layer (NAL) unit having a defined syntax, or it may be added as user data to a NAL unit.
It should be noted here that support for slice groups has been removed from later video compression standards. However, similar currently-available tools can be used to implement this function. Additionally, or alternatively, newer video compression algorithms can be created to include slice group support.
Figures 7A-7B are flow diagrams illustrating a method 160 for avoiding disocclusion artifacts in CVFs and supporting the warp process at the client device 300 according to embodiments of the present disclosure. In this embodiment, method 160 is implemented by the Renderer functionality at an SE, which for illustrative purposes only, executes on network node 200.
As seen in Figure 7A, the Renderer function at the SE receives first and second 2D graphics regions. Each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object (box 162). The Renderer function also receives, for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters (box 164). The Renderer function then sorts the first and second 2D graphics regions based on a Z-order of the first and second 3D objects (box 166). In one embodiment, for example, the Z-order of the first and second 3D objects is based on the 3D coordinates of the first and second 3D objects.
Next, a 3D object (i.e., a 2D graphics region) is selected for processing (box 168), and as previously described, a RegionlD is allocated to the 2D graphics region (box 170). To allocate the RegionlD, this embodiment of the present disclosure first generates the RegionlD based on the Z-order and the Object ID of the selected 3D object (box 172). By way of example only, the present disclosure may concatenate the Z-order and Object ID. Of course, other information may be used in addition to, or in lieu of, the Z-order and/or the ObjectID when generating the RegionlD as needed or desired. Regardless, however, the RegionlD is associated with the 2D graphics region and the 3D object it represents (box 174). As previously described, the RegionlD will allow a warp function at client device 300 to distinguish between the pixel information of two different overlapping 3D objects represented in corresponding 2D graphics regions. Then, based on the 3D coordinates of each of the first and second 3D objects, the Renderer function at the SE determines an un-occluded area of the 2D graphics region (box 176). These so-called un-occluded areas of the first and second 3D objects are the “empty” spaces 48, 50 seen in Figures 3A-3B.
Next, the Renderer function of the SE generates a modified 2D graphics region (box 178). For example, in one embodiment, the Renderer function inserts at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics
region (i.e., the now un-occluded part of the second 3D object). The insertion is based on the motion parameters of each of the first and second 3D objects and the estimated display latency (i.e., the expected latency-to-display). So generated, the Renderer function groups the modified 2D graphics region into a corresponding slice group comprising a plurality of slices (box 180). Each slice in the slice group comprises a sequence of Coding Units (CUs) and is decodable independently of the other slices in the slice group. As previously described, each CU is also decodable independently of the other CUs in the sequence. The Renderer function then checks to see whether any more objects (i.e., 2D graphics regions) exist (box 182). If the Renderer function determines that additional objects/2D graphics regions exist, processing returns to box 168 in Figure 7A. However, if the Renderer function determines that no further objects/2D graphics regions exist, the CVF is encoded (box 184) to include the modified 2D graphics region, the 3D coordinates, the motion parameters, and other information (e.g., Z-order information) as needed or desired, and sent in a bitstream to a decoder at client device 300 (box 186).
The embodiment of Figures 7A-7B described determining the un-occluded area of the 3D object represented by a 2D graphics region as being based on the 3D coordinates of the first and second 3D objects. However, the present disclosure is not so limited. In another embodiment, the un-occluded area of the 2D graphics region is further based on the Z-order of each of the first and second 3D objects.
Additionally, in one embodiment, the first and second 2D graphics regions representing the first and second 3D objects overlap each other. In these embodiments, a first CU in an overlap area of the first 2D graphics region comprises image data for the first 3D object, and a second CU in the overlap area of the second 2D graphics region comprises image data for the second 3D object.
In at least one embodiment, the first and second CUs are associated with a same position of the overlap area.
In one embodiment, the RegionlD distinguishes between information associated with a corresponding slice group and information associated with one or more other slice groups.
In one embodiment, the RegionlD further associates the 3D coordinates of a 3D object represented by a 2D graphics region with the motion parameters of the 3D object represented by the 2D graphics region.
In one or more embodiments, the un-occluded area of the 2D graphics region comprises an un-occluded area of the second 3D object that was previously occluded by the first 3D object.
Further, in one embodiment, a size of the un-occluded area of the second 3D object is based on the estimated display latency for the first and second 3D objects, and the movement of the first and second 3D objects relative to each other.
In one embodiment, the estimated display latency is an estimated time between when the first and second 3D objects are rendered at the simulation engine and when the first and second 3D objects are expected to be displayed on a display of a client device.
In one embodiment, the Z-order, the 3D coordinates, and the motion parameters of the modified 2D graphics region and each of the first and second 3D objects is encoded for each frame.
In one embodiment, the Z-order, the 3D coordinates, and the motion parameters of the modified 2D graphics region is encoded into a user data Network Abstraction Layer (NAL) unit.
An apparatus can perform any of the methods herein described by implementing any functional means, modules, units, or circuitry. In one embodiment, for example, the apparatuses comprise respective circuits or circuitry configured to perform the steps shown in the method figures. The circuits or circuitry in this regard may comprise circuits dedicated to performing certain functional processing and/or one or more microprocessors in conjunction with memory. For instance, the circuitry may include one or more microprocessors or microcontrollers, as well as other digital hardware, which may include Digital Signal Processors (DSPs), special-purpose digital logic, and the like. The processing circuitry may be configured to execute program code stored in memory, which may include one or several types of memory such as read-only memory (ROM), random-access memory, cache memory, flash memory devices, optical storage devices, etc. Program code stored in memory may include program instructions for executing one or more telecommunications and/or data communications protocols as well as instructions for carrying out one or more of the techniques described herein, in several embodiments. In embodiments that employ memory, the memory stores program code that, when executed by the one or more processors, carries out the techniques described herein.
Figure 8 is a functional block diagram illustrating some of the components of a network node 200 configured to avoid disocclusion artifacts and support the warp functionality at client device 300 according to embodiments of the present disclosure. As seen in Figure 8, the network node 200 in this embodiment is configured to execute the SE and comprises, inter alia, communication circuitry 202, processing circuitry 204, and memory 206.
In some embodiments, the communication circuitry 202 comprises both radio frequency (RF) circuitry 202a and network interface circuitry (NIC) 202b. In other embodiments, however, the network node may comprise only NIC 202b. More particularly, the RF circuitry 202a can be located at one or more TRPs and comprises the RF components necessary for communicating with various client devices 300, directly or indirectly, via a wireless communication link. According to the present embodiments, the RF circuitry 202a may comprise, for example, a transmitter and receiver configured to operate according to the 5G standards or other wireless communication standard.
The communication circuitry 202 also comprises network interface circuitry (e.g., NIC 202b) for communication with other RAN nodes, OA&M nodes, core network nodes, and/or
other nodes in external systems. The network interface circuitry 202b in this regard may, for example, comprise an Ethernet interface, optical network interface, or a wireless interface.
The processing circuitry 204 comprises one or more microprocessors, hardware, firmware, or a combination thereof that controls the overall operation of the network node 200. The processing circuitry 204 in this regard can be configured by software to perform the functionality described with respect to Figures 4 and 7A-7B, and one or more of the methods herein described, including methods 60 and 160 as shown in Figures 4 and 7A-7B, respectively.
Memory 206 comprises both volatile and non-volatile memory for storing computer program code and data needed by the processing circuitry 204 for operation. Memory 206 may comprise any tangible, non-transitory computer-readable storage medium for storing data including electronic, magnetic, optical, electromagnetic, or semiconductor data storage. Memory 206 stores one or more computer programs 208 comprising executable instructions that configure the processing circuitry 204 in the network node 200 to perform one or more of the methods herein described, including methods 60 and 160 as shown in Figures 4 and 7A-7B, respectively. A computer program 208 in this regard may comprise one or more code modules corresponding to the means or units described above.
In general, computer program instructions and configuration information are stored in a non-volatile memory, such as a ROM, erasable programmable read only memory (EPROM) or flash memory. Temporary data generated during operation may be stored in a volatile memory, such as a random access memory (RAM). In some embodiments, computer program 208 for configuring the processing circuitry 204 as herein described may be stored in a removable memory, such as a portable compact disc, portable digital video disc, or other removable media. The computer program 208 may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer readable storage medium.
It should be noted here that the previously described embodiment illustrates the functionality described herein on network node 200. However, this is merely illustrative, and the present embodiments are not so limited. In other embodiments, the functionality described herein may be implemented on client device 300 (e.g., on the user’s HMD) or on any computing device, such as computing device 20, the user’s laptop computer, notebook computer, desktop computer, local server computer, mobile communications device (e.g., a SMARTPHONE), tablet computer, and the like. Additionally, as with the network node 200, a computing device in these embodiments may, for example, comprise dedicated hardware circuitry, such as one or more graphics processing units (GPU), and be capable of executing games and other software programs associated with an XR environment.
For example, Figure 9 is a functional block diagram illustrating some of the components of a client device 300 configured to avoid disocclusion artifacts and support warp processing at client device 300 according to embodiments of the present disclosure. The client device 300 in this embodiment may be, for example, a HMD worn by a user or a computing device capable of
executing the SE. Regardless, though, the client device 300 in this embodiment is configured to perform these functions and comprises, inter alia, communication circuitry 302, processing circuitry 304, and memory 306.
The communication circuitry 302 in this embodiment comprises both radio frequency (RF) circuitry 302a and network interface circuitry (NIC) 302b. In other embodiments, however, client device 300 may comprise only the NIC 302b. More particularly, the RF circuitry 302a can be located at one or more TRPs and comprises the RF components necessary for communicating with network node 200 and/or other devices over a wireless communication link. According to the present embodiments, the RF circuitry 302a may comprise, for example, a transmitter and receiver configured to operate according to the 5G standards or other wireless communication standard.
The communication circuitry 302 also comprises network interface circuitry (e.g., NIC 302b) for communication with other RAN nodes, OA&M nodes, core network nodes, and/or other nodes in external systems. The network interface circuitry 302b in this regard may, for example, comprise an ETHERNET interface, optical network interface, or a wireless interface.
The processing circuitry 304 comprises one or more microprocessors, hardware, firmware, or a combination thereof that controls the overall operation of the client device 300. The processing circuitry 304 in this regard can be configured by software to perform the functionality described with respect to Figures 4 and 7A-7B, and one or more of the methods herein described, including methods 60 and 160 as shown in Figures 4 and 7A-7B, respectively.
Memory 306 comprises both volatile and non-volatile memory for storing computer program code and data needed by the processing circuitry 304 for operation. Memory 306 may comprise any tangible, non-transitory computer-readable storage medium for storing data including electronic, magnetic, optical, electromagnetic, or semiconductor data storage. Memory 306 stores one or more computer programs 308 (e.g., an SE) comprising executable instructions that configure the processing circuitry 304 in the client device 300 (or other computing device) to perform one or more of the methods herein described, including methods 60 and 160 as shown in Figures 4 and 7A-7B, respectively. A computer program 308 in this regard may comprise one or more code modules corresponding to the means or units described above.
In general, computer program instructions and configuration information are stored in a non-volatile memory, such as a ROM, erasable programmable read only memory (EPROM) or flash memory. Temporary data generated during operation may be stored in a volatile memory, such as a random access memory (RAM). In some embodiments, computer programs 308 for configuring the processing circuitry 304 as herein described may be stored in a removable memory, such as a portable compact disc, portable digital video disc, or other removable media. The computer program 308 may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer readable storage medium.
The present embodiments may, of course, be carried out in other ways than those specifically set forth herein without departing from characteristics described herein. The present embodiments are therefore to be considered in all respects as illustrative and not restrictive, and all changes coming within the meaning and equivalency range of the appended claims are intended to be embraced therein.
Claims
1 . A method (160), implemented by a network node (200), for avoiding disocclusion artifacts in Composite Video Frames (CVFs), the method comprising: receiving (162) first and second 2D graphics regions, wherein each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object; receiving (164), for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters; determining (176) an un-occluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects; generating (178), based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region; encoding (184) a CVF to include the modified 2D graphics region; and sending (186) the encoded CVF to a decoder in a bitstream.
2. The method of claim 1 , further comprising sorting (166) the first and second 2D graphics regions based on a Z-order of the first and second 3D objects.
3. The method of claim 2, wherein the Z-order of the first and second 3D objects is based on the 3D coordinates of the first and second 3D objects.
4. The method of any of claims 1-3, further comprising: for each of the first and second 3D objects: generating (172) a RegionlD based on the Z-order of the 3D object and an object ID of the 3D object; and associating (174) the RegionlD with the 2D graphics region representing the 3D object.
5. The method of any of claims 1 -4, wherein determining an un-occluded area of the second 2D graphics region is further based on the Z-order of each of the first and second 3D objects.
6. The method of any of claims 1 -5, further comprising grouping (180) the modified 2D graphics region into a corresponding slice group comprising a plurality of slices, wherein each slice in the slice group comprises a sequence of Coding Units (CUs) and is decodable independently of the other slices in the slice group.
7. The method of claims 1-6, wherein the first and second 2D graphics regions representing the first and second 3D objects overlap each other such that a first CU in an overlap area of the first 2D graphics region comprises image data for the first 3D object, and a second CU in the overlap area of the second 2D graphics region comprises image data for the second 3D object.
8. The method of claim 7, wherein the first and second CUs are associated with a same position of the overlap area.
9. The method of claims 3-8, wherein the RegionlD distinguishes between information associated with the corresponding slice group and information associated with one or more other slice groups.
10. The method of claim 3-9, wherein the RegionlD further associates the 3D coordinates of a 3D object represented by a 2D graphics region with the motion parameters of the 3D object represented by the 2D graphics region.
11 . The method of any of the preceding claims, wherein the un-occluded area of the 2D graphics region comprises an un-occluded area of the second 3D object that was previously occluded by the first 3D object.
12. The method of any of the preceding claims, wherein a size of the un-occluded area of the second 3D object is based on the estimated display latency for the first and second 3D objects, and the movement of the first and second 3D objects relative to each other.
13. The method of any of the preceding claims, wherein the estimated display latency is an estimated time between when the first and second 3D objects are rendered at the simulation engine and when the first and second 3D objects are expected to be displayed on a display of a client device.
14. The method of any of the preceding claims, wherein the Z-order, the 3D coordinates, and the motion parameters of the modified 2D graphics region and each of the first and second 3D objects is encoded for each frame.
15. The method of claim 14, wherein the Z-order, the 3D coordinates, and the motion parameters of the modified 2D graphics region is encoded into a user data Network Abstraction Layer (NAL) unit.
16. A network node (200) configured to avoid disocclusion artifacts in Composite Video Frames, the network node configured to: receive (162) first and second 2D graphics regions, wherein each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object; receive (164) , for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters; determine (176) an un-occluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects; generate (178), based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region; encode (184) a CVF to include the modified 2D graphics region; and send (186) the encoded CVF to a decoder in a bitstream.
17. The network node of claim 16, wherein the network node is further configured to perform the method according to any one of claims 2-15.
18. A network node (200) configured to avoid disocclusion artifacts in Composite Video Frames, the network node comprising: communications circuitry (202) configured to communicate with a client device (300); and processing circuitry (204) operatively connected to the communications circuitry and configured to: receive (162) first and second 2D graphics regions, wherein each of the first and second 2D graphics regions comprise an image representing a respective first and second 3D object with the first 3D object occluding the second 3D object; receive (164), for each of the first and second 3D objects, corresponding 3D coordinates and motion parameters; determine (176) an un-occluded area of the second 2D graphics region based on the 3D coordinates of each of the first and second 3D objects; generate (178), based on the motion parameters of each of the first and second 3D objects and an estimated display latency, a modified 2D graphics region by inserting at least part of an occluded area of the second 3D object into the un-occluded area of the second 2D graphics region; encode (184) a CVF to include the modified 2D graphics region; and send (186) the encoded CVF to a decoder in a bitstream.
19. The network node of claim 18, wherein the network node is further configured to perform the method according to any one of claims 2-15.
20. A computer program (208) comprising instructions that, when executed on processing circuitry (204) of a network node (200), cause the network node to perform the method according to any of claims 1-15.
21 . A carrier containing the computer program of claim 20, wherein the carrier is one of an electronic signal, optical signal, radio signal, or computer readable storage medium.
22. A non-transitory computer-readable storage medium (206) comprising a computer program (208) stored thereon, the computer program comprising executable instructions that, when executed by processing circuitry (204) in a network node (200), causes the network node to perform the method of any one of claims 1-15.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/EP2024/060046 WO2025214613A1 (en) | 2024-04-12 | 2024-04-12 | Warp support for 2-dimensional (2d) video coding |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/EP2024/060046 WO2025214613A1 (en) | 2024-04-12 | 2024-04-12 | Warp support for 2-dimensional (2d) video coding |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025214613A1 true WO2025214613A1 (en) | 2025-10-16 |
Family
ID=90731422
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2024/060046 Pending WO2025214613A1 (en) | 2024-04-12 | 2024-04-12 | Warp support for 2-dimensional (2d) video coding |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025214613A1 (en) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2022028684A1 (en) * | 2020-08-05 | 2022-02-10 | Telefonaktiebolaget Lm Ericsson (Publ) | Improved split rendering for extended reality (xr) applications |
| WO2023227223A1 (en) * | 2022-05-26 | 2023-11-30 | Telefonaktiebolaget Lm Ericsson (Publ) | Split transport for warping |
-
2024
- 2024-04-12 WO PCT/EP2024/060046 patent/WO2025214613A1/en active Pending
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2022028684A1 (en) * | 2020-08-05 | 2022-02-10 | Telefonaktiebolaget Lm Ericsson (Publ) | Improved split rendering for extended reality (xr) applications |
| WO2023227223A1 (en) * | 2022-05-26 | 2023-11-30 | Telefonaktiebolaget Lm Ericsson (Publ) | Split transport for warping |
Non-Patent Citations (1)
| Title |
|---|
| THOMAS WIEGAND ET AL: "Overview of the H.264/AVC Video Coding Standard", IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, IEEE, USA, vol. 13, no. 7, 1 July 2003 (2003-07-01), pages 560 - 576, XP008129745, ISSN: 1051-8215, DOI: 10.1109/TCSVT.2003.815165 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11122102B2 (en) | Point cloud data transmission apparatus, point cloud data transmission method, point cloud data reception apparatus and point cloud data reception method | |
| US11315270B2 (en) | Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method | |
| JP7434574B2 (en) | Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method | |
| CN115918093B (en) | Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device and point cloud data receiving method | |
| US11979607B2 (en) | Apparatus and method for processing point cloud data | |
| CN115443652B (en) | Point cloud data sending device, point cloud data sending method, point cloud data receiving device and point cloud data receiving method | |
| JP7798974B2 (en) | Point cloud data processing apparatus and method | |
| US20240137578A1 (en) | Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method | |
| CN114930813A (en) | Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method | |
| US8872895B2 (en) | Real-time video coding using graphics rendering contexts | |
| JP7440546B2 (en) | Point cloud data processing device and method | |
| US12149579B2 (en) | Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method | |
| JP7425207B2 (en) | Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method | |
| WO2020055655A1 (en) | Scalability of multi-directional video streaming | |
| US20240386615A1 (en) | Point cloud data transmission device and method, and point cloud data reception device and method | |
| JP7640542B2 (en) | Point cloud data transmission device, point cloud data transmission method, point cloud data receiving device, and point cloud data receiving method | |
| US12260600B2 (en) | Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method | |
| US12400369B2 (en) | Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method | |
| CN115380528A (en) | Point cloud data transmitting device, point cloud data transmitting method, point cloud data receiving device, and point cloud data receiving method | |
| US20240420377A1 (en) | Point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device | |
| US20220230360A1 (en) | Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method | |
| US20250133233A1 (en) | Point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device | |
| US20250088659A1 (en) | Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method | |
| US20240422355A1 (en) | Point cloud data transmission device and method, and point cloud data reception device and method | |
| US20240276013A1 (en) | Point cloud data transmission device, point cloud data transmission method, point cloud data reception device and point cloud data reception method |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24719157 Country of ref document: EP Kind code of ref document: A1 |