WO2025005646A1 - 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 - Google Patents
메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 Download PDFInfo
- Publication number
- WO2025005646A1 WO2025005646A1 PCT/KR2024/008873 KR2024008873W WO2025005646A1 WO 2025005646 A1 WO2025005646 A1 WO 2025005646A1 KR 2024008873 W KR2024008873 W KR 2024008873W WO 2025005646 A1 WO2025005646 A1 WO 2025005646A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- mesh
- displacement
- displacement vector
- information
- bitstream
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/70—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/18—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being a set of transform coefficients
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/46—Embedding additional information in the video signal during the compression process
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/537—Motion estimation other than block-based
- H04N19/54—Motion estimation other than block-based using feature points or meshes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/593—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving spatial prediction techniques
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/597—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding specially adapted for multi-view video sequence encoding
Definitions
- the embodiments provide a method for providing 3D content to provide users with various services such as Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), and autonomous driving services.
- VR Virtual Reality
- AR Augmented Reality
- MR Mixed Reality
- autonomous driving services such as autonomous driving services.
- point cloud data or mesh data is a collection of points in 3D space.
- the technical problem according to the embodiments is to provide a device and method for efficiently transmitting and receiving mesh data in order to solve the problems described above.
- a technical problem according to embodiments is to provide a device and method for resolving latency and encoding/decoding complexity of mesh data.
- the technical problem according to the embodiments is to provide a device and method for efficiently performing encoding and decoding of a displacement vector.
- a mesh data encoding method may include a step of encoding an original mesh, and a step of transmitting a bitstream including the encoded mesh and signaling information.
- the encoding step may include a base mesh processing step of encoding a base mesh generated by simplifying the original mesh to generate a base mesh bitstream, a displacement information processing step of encoding displacement information generated based on the base mesh to generate a displacement vector bitstream, a mesh restoration step of restoring a mesh based on the encoded base mesh and the encoded displacement information, and a texture map processing step of encoding a texture map generated based on the original mesh and the restored mesh to generate a texture map bitstream.
- the displacement information processing step may include a step of converting a coordinate system of the displacement information into a local coordinate system, a step of performing packing on at least one of a normal component, a tangential component, or a bi-tangential component included in the displacement information of the local coordinate system based on one of a first format, a second format, or a third format, and a step of encoding the packed displacement information.
- the packing step selectively performs packing on a bi-tangential component of the displacement information
- the signaling information may include information for identifying whether packing of the bi-tangential component is skipped.
- the signaling information may include information for identifying a format applied to the displacement information.
- a mesh data encoding device may include an encoder that encodes an original mesh, and a transmission unit that transmits a bitstream including the encoded mesh and signaling information.
- the encoder may include a base mesh processing unit that generates a base mesh bitstream by encoding a base mesh generated by simplifying the original mesh, a displacement information processing unit that generates a displacement vector bitstream by encoding displacement information generated based on the base mesh, a mesh restoration unit that restores a mesh based on the encoded base mesh and the encoded displacement information, and a texture map processing unit that generates a texture map bitstream by encoding a texture map generated based on the original mesh and the restored mesh.
- the displacement information processing unit may include a displacement information coordinate system conversion unit that converts a coordinate system of the displacement information into a local coordinate system, a displacement information packing unit that performs packing based on one of the first format, the second format, and the third format for at least one of the normal component, the tangential component, or the bi-tangential component included in the displacement information of the local coordinate system, and a displacement information encoding unit that encodes the packed displacement information.
- a displacement information coordinate system conversion unit that converts a coordinate system of the displacement information into a local coordinate system
- a displacement information packing unit that performs packing based on one of the first format, the second format, and the third format for at least one of the normal component, the tangential component, or the bi-tangential component included in the displacement information of the local coordinate system
- a displacement information encoding unit that encodes the packed displacement information.
- the displacement information packing unit selectively performs packing on a bi-tangential component of the displacement information
- the signaling information may include information for identifying whether packing of the bi-tangential component has been skipped.
- the signaling information may include information for identifying a format applied to the displacement information.
- a method for decoding mesh data may include a step of receiving a base mesh bitstream, a displacement vector bitstream, a texture map bitstream, and signaling information, a base mesh processing step of restoring a base mesh from the base mesh bitstream, a displacement information processing step of restoring displacement information from the displacement vector bitstream, a restoration step of restoring a mesh based on the base mesh and the displacement information, and a texture map processing step of restoring a texture map from the texture map bitstream.
- the displacement information processing step may include a step of decoding the displacement vector bitstream into displacement information, a step of identifying a format applied to the displacement information among the first format, the second format, and the third format, a step of depacking at least one of a normal component, a tangential component, or a bi-tangential component included in the displacement information based on the identified format, and a step of detransforming a local coordinate system of the depacked displacement information into an original coordinate system.
- the signaling information may include information for identifying a format applied to the displacement information.
- the signaling information further includes information for identifying whether packing of the bi-tangential component is skipped, and the depacking step can determine a depacking method of at least one of the normal component, the tangential component, or the bi-tangential component based on the signaling information.
- a mesh data decoding device may include a receiving unit that receives a base mesh bitstream, a displacement vector bitstream, a texture map bitstream, and signaling information, a base mesh processing unit that restores a base mesh from the base mesh bitstream, a displacement information processing unit that restores displacement information from the displacement vector bitstream, a restoration unit that restores a mesh based on the base mesh and the displacement information, and a texture map processing unit that restores a texture map from the texture map bitstream.
- the displacement information processing unit may include a displacement information decoding unit that decodes the displacement vector bitstream into displacement information, a displacement information depacking unit that identifies a format applied to the displacement information among the first format, the second format, and the third format, and depacks at least one of a normal component, a tangential component, or a bi-tangential component included in the displacement information based on the identified format, and a coordinate system inverse transformation unit that inversely transforms a local coordinate system of the depacked displacement information into an original coordinate system.
- a computer program stored on a computer-readable recording medium can be combined with a computer, which is hardware, to perform the above method.
- a computer program stored on a computer-readable recording medium can be combined with a computer, which is hardware, to perform the above method.
- the mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device can provide a quality 3D service.
- the mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device can achieve various video codec methods.
- the mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device can provide general-purpose 3D content such as autonomous driving services.
- the mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device encode and transmit only the normal component and the tangential component excluding the bi-tangential component among the normal component, the tangential component, and the bi-tangential component converted to a local coordinate system, and cause a decoder of the receiving device to decode up to the bi-tangential component, thereby obtaining mesh data with an efficient bit saving effect and thus better image quality.
- Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.
- Figure 2 illustrates a V-MESH compression method according to embodiments.
- FIG. 3 illustrates pre-processing of V-MESH compression according to embodiments.
- Figure 4 illustrates a mid-edge subdivision method according to embodiments.
- Figure 5 illustrates a displacement generation process according to embodiments.
- Figure 6 illustrates an intra frame encoding process of V-MESH data according to embodiments.
- Figure 7 illustrates an inter-frame encoding process of V-MESH data according to embodiments.
- Figure 8 shows a lifting conversion process for displacement according to embodiments.
- Figure 9 illustrates a process of packing transformation coefficients into a 2D image according to embodiments.
- Figure 10 illustrates the attribute transfer process of the V-MESH compression method according to embodiments.
- FIG. 11 illustrates an intra-frame decoding process of V-MESH data according to embodiments.
- Figure 12 shows an inter-frame decoding processor of V-MESH data.
- FIG. 13 is a drawing showing an example of a transmitter device according to embodiments.
- FIG. 14 is a drawing showing an example of a receiving device according to embodiments.
- FIG. 15 is a drawing showing another example of a transmitter device according to embodiments.
- Figure 16 is a detailed block diagram of a displacement vector encoder according to embodiments.
- FIG. 17 is a diagram showing an example of a displacement vector coefficient structure according to embodiments.
- FIG. 18 is a diagram showing an example of packing displacement vector coefficients according to embodiments.
- FIG. 19 is a diagram showing an example of packing displacement vector coefficient blocks by LoD according to embodiments.
- FIG. 20(a) and FIG. 20(b) are diagrams showing examples of a method for packing displacement vector coefficients in a displacement vector coefficient block according to embodiments.
- FIG. 21(a) and FIG. 21(b) are diagrams showing an example of packing displacement vector coefficients based on the YUV 4:4:4 format according to embodiments.
- FIG. 22(a) and FIG. 22(b) are diagrams showing another example of packing displacement vector coefficients based on the YUV 4:4:4 format according to embodiments.
- FIG. 23(a) and FIG. 23(b) are diagrams showing an example of packing displacement vector coefficients based on the YUV 4:2:0 format according to embodiments.
- FIG. 24(a) and FIG. 24(b) are diagrams showing another example of packing displacement vector coefficients based on the YUV 4:2:0 format according to embodiments.
- FIG. 25(a) and FIG. 25(b) are diagrams showing an example of packing displacement vector coefficients based on the YUV 4:0:0 format according to embodiments.
- FIG. 26(a) and FIG. 26(b) are diagrams showing another example of packing displacement vector coefficients based on the YUV 4:0:0 format according to embodiments.
- FIG. 27 is a drawing showing another example of a receiving device according to embodiments.
- Fig. 28 is an example of a detailed block diagram of a displacement vector decoder according to embodiments.
- FIG. 29(a) and FIG. 29(b) are diagrams showing an example of reverse packing displacement vector coefficients based on the YUV 4:4:4 format according to embodiments.
- FIG. 30(a) and FIG. 30(b) are diagrams showing another example of reverse packing displacement vector coefficients based on the YUV 4:4:4 format according to embodiments.
- FIG. 31(a) and FIG. 31(b) are diagrams showing an example of reverse packing displacement vector coefficients based on the YUV 4:2:0 format according to embodiments.
- FIG. 32(a) and FIG. 32(b) are diagrams showing another example of depacking displacement vector coefficients based on the YUV 4:2:0 format according to embodiments.
- FIG. 33(a) and FIG. 33(b) are diagrams showing an example of reverse packing displacement vector coefficients based on the YUV 4:0:0 format according to embodiments.
- FIG. 34(a) and FIG. 34(b) are diagrams showing another example of depacking displacement vector coefficients based on the YUV 4:0:0 format according to embodiments.
- Figures 35(a) and 35(b) are diagrams showing another example of depacking displacement vector coefficients based on the YUV 4:0:0 format according to embodiments.
- FIG. 36 is a diagram showing an example of a syntax structure of signaling information according to embodiments.
- FIG. 37 is a diagram showing another example of a syntax structure of signaling information according to embodiments.
- Figure 38 is a flowchart showing an example of a transmission method according to embodiments.
- Figure 39 is a flowchart showing an example of a receiving method according to embodiments.
- 3D data can be expressed as a point cloud, mesh, etc., depending on the expression format.
- a mesh is composed of geometric information expressing the coordinate values of each vertex (vertex or point), connection information indicating the connection relationship between vertices, a texture map expressing the color information of the mesh surface as 2D image data, and texture coordinates indicating mapping information between the surface of the mesh and the texture map.
- a mesh is defined as a dynamic mesh when at least one of the elements constituting the mesh changes over time, and a static mesh when it does not change.
- Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.
- the system of FIG. 1 includes a transmitting device (100) and a receiving device (110) according to embodiments.
- the transmitting device (100) may include a mesh video acquisition unit (101), a mesh video encoder (102), a file/segment encapsulator (103), and a transmitter (104).
- the receiving device (110) may include a receiving unit (111), a file/segment decapsulator (112), a mesh video decoder (113), and a renderer (114).
- Each component of FIG. 1 may correspond to hardware, software, a processor, and/or a combination thereof.
- the mesh data transmitting device may be interpreted as a term referring to a 3D data transmitting device or the transmitting device (100), or a mesh video encoder (hereinafter, referred to as an encoder) (102).
- the mesh data receiving device may be interpreted as a term referring to a 3D data receiving device or receiving device (110), or a mesh video decoder (hereinafter, decoder) (113).
- the system of Fig. 1 can perform video-based dynamic mesh compression and decompression.
- 3D content such as AR, XR, metaverse, and holograms
- 3D contents express objects more precisely and realistically so that users can enjoy immersive experiences, and for this purpose, a large amount of data is required to create and use 3D models.
- 3D mesh is widely used for efficient data utilization and realistic object expression. Embodiments include a series of processing processes in a system that uses such mesh content.
- Point cloud data is data having color information at the coordinates (X, Y, Z) of a vertex (or point).
- the coordinates (i.e., position information) of a vertex are referred to as geometry information
- the color information of a vertex is referred to as attribute information
- the geometry information and the attribute information are referred to as vertex information or point cloud data.
- the vertex information to which connectivity information between vertices is added is referred to as mesh data.
- it can be created in the form of mesh data from the beginning. Or, it can be used by converting it into mesh data by adding connectivity information to point cloud data.
- the MPEG standards body defines the data types of dynamic mesh data as the following two types.
- Category 1 Mesh data with texture maps as color information.
- Category 2 Mesh data with vertex colors as color information.
- the entire process for providing mesh content services may include acquisition processes, encoding processes, transmission processes, decoding processes, rendering processes, and/or feedback processes, as shown in Fig. 1.
- 3D data acquired through multiple cameras or special cameras can be processed into a mesh data type through a series of processes and then generated as a video.
- the generated mesh video is transmitted through a series of processes, and the receiving end can process the received data again into a mesh video and render it.
- a mesh video is provided to the user, and the user can use the mesh content according to his/her intention through interaction.
- a mesh compression system may include a transmitter (100) and a receiver (110) as shown in Fig. 1.
- the transmitter (100) may encode a mesh video to output a bitstream, and transmit it to the receiver (110) in the form of a file or streaming (streaming segment) through a digital storage medium or a network.
- the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
- the encoder may be called a Mesh video/video/picture/frame encoding device
- the decoder may be called a Mesh video/video/picture/frame decoding device.
- the transmitter may be included in a Mesh video encoder.
- the receiver may be included in a Mesh video decoder.
- the renderer (114) may include a display unit, and the renderer and/or the display unit may be configured as separate devices or external components.
- the transmitting device (100) and the receiving device (110) may further include separate internal or external modules/units/components for a feedback process.
- Mesh data represents the surface of an object as a number of polygons. Each polygon is defined by vertices in 3D space and connection information indicating how the vertices are connected. It can also include vertex attributes such as vertex color, normal, etc. Mapping information that allows the surface of the mesh to be mapped to a 2D planar area can also be included in the attributes of the mesh. The mapping can be described as a set of parameter coordinates, generally called UV coordinates or texture coordinates, associated with the mesh vertices.
- the mesh contains a 2D attribute map, which can be used to store high-resolution attribute information such as texture, normal, displacement, etc.
- displacement can be used interchangeably with displacement, displacement information, or displacement vector (i.e., displacement vector).
- the mesh video acquisition unit (101) may include processing 3D object data acquired through a camera, etc. into a mesh data type having the attributes described above through a series of processes and generating a video composed of such mesh data.
- the mesh video may have attributes of the mesh, such as vertices, polygons, connection information between vertices, colors, normals, etc., that may change over time.
- a mesh video having attributes and connection information that change over time in this way may be expressed as a dynamic mesh video.
- the mesh video encoder (102) can encode an input mesh video into one or more video streams.
- One video can include multiple frames, and one frame can correspond to a still image/picture.
- the mesh video can include mesh images/frames/pictures, and the mesh video can be used interchangeably with the mesh images/frames/pictures.
- the mesh video encoder (102) can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure.
- the mesh video encoder (102) can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency.
- the encoded data (encoded video/image information) can be output in the form of a bitstream.
- the file/segment encapsulator (103) can encapsulate encoded mesh video data and/or mesh video related metadata in the form of a file, etc.
- the mesh video related metadata may be received from a metadata processing unit, etc.
- the metadata processing unit may be included in the mesh video encoder (102) or may be configured as a separate component/module.
- the file/segment encapsulator (103) can encapsulate the corresponding data in a file format such as ISOBMFF, or process it in the form of other DASH segments, etc.
- the file/segment encapsulator (103) may include mesh video related metadata in the file format according to an embodiment.
- the mesh video metadata may be included in boxes at various levels in the ISOBMFF file format, for example, or may be included as data in a separate track within the file.
- the file/segment encapsulator (103) may encapsulate the mesh video related metadata itself into a file.
- the transmission processing unit can process encapsulated mesh video data for transmission according to the file format.
- the transmission processing unit can be included in the transmission unit (104) or can be configured as a separate component/module.
- the transmission processing unit can process mesh video data according to any transmission protocol.
- the processing for transmission can include processing for transmission through a broadcast network and processing for transmission through broadband.
- the transmission processing unit can receive not only mesh video data but also mesh video-related metadata from the metadata processing unit and process it for transmission.
- the transmission unit (104) can transmit encoded video/image information or data output in the form of a bitstream to the reception unit (111) of the reception device (110) through a digital storage medium or network in the form of a file or streaming.
- the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
- the transmission unit (104) can include an element for generating a media file through a predetermined file format and can include an element for transmission through a broadcasting/communication network.
- the reception unit (111) can extract the bitstream and transmit it to a decoding device.
- the receiving unit (111) can receive mesh video data transmitted by the mesh data transmitting device. Depending on the channel through which it is transmitted, the receiving unit (111) can receive mesh video data through a broadcast network, through a broadband, or through a digital storage medium.
- the receiving processing unit can perform processing according to a transmission protocol on the received mesh video data.
- the receiving processing unit can be included in the receiving unit (111), or can be configured as a separate component/module. In order to correspond to the processing performed for transmission on the transmitting side, the receiving processing unit can perform the reverse process of the aforementioned transmission processing unit.
- the receiving processing unit can transfer the acquired mesh video data to the file/segment decapsulator (112), and transfer the acquired mesh video-related metadata to the metadata parser.
- the mesh video-related metadata acquired by the receiving processing unit can be in the form of a signaling table.
- the file/segment decapsulator (112) can decapsulate mesh video data in the form of a file received from a receiving processing unit.
- the file/segment decapsulator (112) can decapsulate files according to ISOBMFF, etc., to obtain a mesh video bitstream or mesh video related metadata (metadata bitstream).
- the obtained mesh video bitstream can be transmitted to the mesh video decoder (113), and the obtained mesh video related metadata (metadata bitstream) can be transmitted to the metadata processing unit.
- the mesh video bitstream may include metadata (metadata bitstream).
- the metadata processing unit may be included in the mesh video decoder (113) or may be configured as a separate component/module.
- the mesh video related metadata obtained by the file/segment decapsulator (112) may be in the form of a box or track within a file format.
- the file/segment decapsulator (112) may receive metadata required for decapsulation from the metadata processing unit, if necessary.
- the mesh video related metadata may be passed to the mesh video decoder (113) and used in the mesh video decoding procedure, or may be passed to the renderer (114) and used in the mesh video rendering procedure.
- the mesh video decoder (113) can receive a bitstream and perform a reverse operation corresponding to the operation of the mesh video encoder (102) to decode the video/image.
- the decoded mesh video/image can be displayed through the display unit of the renderer (114).
- the user can view all or part of the rendered result through a VR/AR display or a general display.
- the feedback process may include a process of transmitting various feedback information that may be acquired during the rendering/display process to the transmitter or to the decoder of the receiver. Interactivity may be provided in mesh video consumption through the feedback process. According to an embodiment, head orientation information, viewport information indicating an area that the user is currently viewing, etc. may be transmitted during the feedback process. According to an embodiment, the user may interact with things implemented in the VR/AR/MR/autonomous driving environment, in which case information related to the interaction may be transmitted to the transmitter or the service provider during the feedback process. Depending on the embodiment, the feedback process may not be performed.
- Head orientation information can mean information about the user's head position, angle, movement, etc. Based on this information, information about the area the user is currently viewing within the mesh video, i.e. viewport information, can be calculated.
- Viewport information can be information about the area that the current user is viewing in the mesh video. This can be used to perform gaze analysis to determine how the user consumes the mesh video, which area of the mesh video they gaze at and for how long. Gaze analysis can be performed on the receiving side and transmitted to the transmitting side through a feedback channel.
- Devices such as VR/AR/MR displays can extract the viewport area based on the user's head position/orientation, the vertical or horizontal FOV supported by the device, etc.
- the aforementioned feedback information may be consumed by the receiver as well as transmitted to the transmitter. That is, the decoding and rendering processes of the receiver may be performed using the aforementioned feedback information. For example, only the mesh video for the area currently being viewed by the user may be preferentially decoded and rendered using head orientation information and/or viewport information.
- Dynamic mesh video compression is a method for processing mesh connection information and attributes that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-viewpoint video, and AR/VR.
- the dynamic mesh video compression method described below is based on MPEG's V-Mesh method.
- picture/frame can generally mean a unit representing one video image of a specific time period.
- a pixel or pel can mean the smallest unit that constitutes a picture (or image).
- a 'sample' can be used as a term corresponding to a pixel.
- a sample can generally represent a pixel or a pixel value, and can represent only a pixel/pixel value of a luma component, only a pixel/pixel value of a chroma component, or only a pixel/pixel value of a depth component.
- a unit may represent a basic unit of image processing.
- a unit may include at least one of a specific region of a picture and information related to the region.
- a unit may be used interchangeably with terms such as block or area, depending on the case.
- an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
- the video-based dynamic mesh compression (V-Mesh) compression method can provide a method of compressing dynamic mesh video data based on a 2D video codec such as HEVC (High Efficiency Video Coding) and VVC (Versatile Video Coding).
- HEVC High Efficiency Video Coding
- VVC Very Video Coding
- Input mesh Contains the 3D coordinates of the vertices that make up the mesh, normal information for each vertex, mapping information that maps the mesh surface to a 2D plane, connection information between the vertices that make up the surface, etc.
- the mesh surface can be expressed as triangles or more polygons, and connection information between the vertices that make up each surface is stored according to a set shape.
- the input mesh can be saved in the OBJ file format.
- Attribute map (Hereinafter, texture map is also used with the same meaning): Contains information on the attributes (color, normal, displacement, etc.) of the mesh, and stores data in the form of mapping the surface of the mesh onto a 2D image. Mapping which part (surface or vertex) of the mesh each data of this attribute map corresponds to is based on the mapping information included in the input mesh. Since the attribute map has data for each frame of the mesh video, it can also be expressed as an attribute map video.
- the attribute map in the V-Mesh compression method mainly contains color information of the mesh, and is stored in an image file format (PNG, BMP, etc.).
- Material Library File Contains information about the material attributes used in the mesh, and in particular, information that links the input mesh to its corresponding attribute map. It is saved in the Wavefront Material Template Library (MTL) file format.
- TTL Wavefront Material Template Library
- the following data and information can be generated through the compression process.
- Base mesh The input mesh is simplified (decimated) through a pre-processing process to express the objects of the input mesh using the minimum number of vertices determined by the user's criteria.
- Displacement information used to express the input mesh as similarly as possible using the base mesh, and is expressed in the form of three-dimensional coordinates.
- Atlas information This is metadata required to reconstruct the mesh using the base mesh, displacement, and attribute map information. This can be created and utilized as a sub-unit (sub-mesh, patch, etc.) that constitutes the mesh.
- FIGS. 2 to 7 a method for encoding mesh position information (or vertex position information) is described, and referring to FIGS. 6 to 10, etc., a method for encoding attribute information (attribute map) by restoring mesh position information is described.
- Figure 2 illustrates a V-MESH compression method according to embodiments.
- FIG. 2 illustrates the encoding process of FIG. 1, and the encoding process may include a pre-processing process and an encoding process.
- the mesh video encoder (102) of FIG. 1 may include a pre-processor (200) and an encoder (201) as in FIG. 2.
- the transmitting device of FIG. 1 may be broadly referred to as an encoder, and the mesh video encoder (102) of FIG. 1 may be referred to as an encoder.
- the V-Mesh compression method may include a pre-processing process (Pre-processing, 200) and an encoding process (Encoding, 201) as in FIG. 2.
- the pre-processor (200) of FIG. 2 may be located in front of the encoder (201) of FIG. 2.
- the pre-processor (200) and the encoder (201) of FIG. 2 may be referred to as one encoder.
- the pre-processor (200) can receive a static of a dynamic mesh (M(i)) and/or an attribute map (A(i)).
- the pre-processor (200) can generate a base mesh (m(i)) and/or a displacement (or displacement) (d(i)) through pre-processing.
- the pre-processor (200) can receive feedback information from an encoder (201) and generate the base mesh and/or the displacement based on the feedback information.
- the encoder (201) can receive a base mesh (m(i)), a displacement (d(i)), a static of a dynamic mesh (M(i)), and/or an attribute map (A(i)).
- a base mesh (m(i)), a displacement (d(i)), a static of a dynamic mesh (M(i)), and/or an attribute map (A(i)) may be referred to as mesh-related data.
- the encoder (201) can encode the mesh-related data to generate a compressed bitstream.
- FIG. 3 illustrates the pre-processing process of V-MESH compression according to embodiments.
- an input mesh may include a static of a dynamic mesh (M(i)) and/or an attribute map (A(i)).
- the input mesh may include three-dimensional coordinates of vertices constituting the mesh, normal information of each vertex, mapping information for mapping the mesh surface to a 2D plane, connection information between vertices constituting the surface, etc.
- FIG. 3 shows a process of performing pre-processing on an input mesh.
- the pre-processing process (200) may largely include four steps: 1) GoF (Group of Frame) generation, 2) mesh simplification (Mesh Decimation), 3) UV parameterization, and 4) fitting subdivision surface (300).
- GoF generation may be referred to as a GoF generation process or a GoF generation unit
- mesh simplification may be referred to as a mesh simplification process or a mesh simplification unit
- UV parameterization may be referred to as a UV parameterization process or a UV parameterization unit
- the fitting subdivision surface may be referred to as a fitting subdivision surface process or a fitting subdivision surface unit.
- the pre-processor (200) can generate displacement and/or base meshes from the received input mesh and transmit them to the encoder (201).
- the pre-processor (200) can transmit GoF information associated with GoF generation to the encoder (201).
- GoF Generation This is the process of generating a reference structure for mesh data. If the number of vertices, the number of texture coordinates, the vertex connection information, and the texture coordinate connection information of the mesh of the previous frame and the current mesh are all the same, the previous frame can be set as the reference frame. That is, if only the vertex coordinate values are different between the current input mesh and the reference input mesh, the encoder (201) can perform inter frame encoding. Otherwise, intra frame encoding is performed for the corresponding frame.
- Mesh Decimation This is the process of simplifying the input mesh to create a simplified mesh, or base mesh. After selecting vertices to be removed from the original mesh based on criteria defined by the user, the selected vertices and triangles connected to the selected vertices can be removed.
- the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) information are passed as input, and the simplified mesh (decimated mesh) can be obtained as output.
- the simplified mesh (decimated mesh) can be obtained as output.
- connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.
- UV parameterization This is the process of mapping a 3D surface to a texture domain for a decimated mesh. Parameterization can be performed using the UVAtlas tool. Through this process, mapping information is generated that shows where each vertex of the decimated mesh can be mapped to on a 2D image. The mapping information is expressed and saved as texture coordinates, and the final base mesh is generated through this process.
- Fitting subdivision surface (300) This is a process of performing subdivision on a decimated mesh (i.e., a simplified mesh having texture coordinates).
- the displacement and base mesh generated through this process are output to the encoder (201).
- a user-defined method such as a mid-edge method, may be applied as the subdivision method.
- a fitting process is performed so that the input mesh and the mesh on which the subdivision is performed are similar to each other.
- a mesh on which the fitting process is performed is referred to as a fitted subdivision mesh (or a fitted subdivision mesh).
- Figure 4 illustrates a mid-edge subdivision method according to embodiments.
- Fig. 4 shows the mid-edge method of the fitting subdivision surface described in Fig. 3.
- an original mesh including four vertices is subdivided to generate a sub-mesh.
- a sub-mesh can be generated by generating a new vertex in the middle of an edge between vertices. Then, a fitting process is performed so that the input mesh and the sub-mesh become similar to each other, and a fitted sub-division mesh is generated.
- the fitted subdivided mesh When a fitted subdivided mesh (hereinafter referred to as the fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decoded base mesh (hereinafter referred to as the reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface.
- the difference in position of each vertex between this result and the fitted subdivided mesh is the displacement for each vertex. Since the displacement represents the position difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of the Cartesian coordinate system.
- the (x, y, z) coordinate values can be converted to (normal, tangential, bi-tangential) coordinate values in the local coordinate system.
- Fig. 5 illustrates a displacement generation process according to embodiments.
- the displacement generation process of Fig. 5 may be performed in a pre-processor (200) or may be performed in an encoder (201).
- FIG. 5 illustrates in detail the displacement calculation method of the fitting subdivision surface (300) as described in FIG. 4.
- the encoder and/or pre-processor may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit.
- the subdivision unit may perform a subdivision on a restored base mesh to generate a subdivided restored base mesh.
- the restoration of the base mesh may be performed in the pre-processor (200) or may be performed in the encoder (201).
- the local coordinate system calculation unit may receive the fitted subdivision mesh and the subdivided restored base mesh, and may convert a coordinate system of the mesh into a local coordinate system based on the same.
- the local coordinate system calculation operation may be optional.
- the displacement calculation unit may calculate a position difference between the fitted subdivision mesh and the subdivided restored base mesh. For example, a position difference value between vertices of two input meshes may be generated. The vertex position difference value becomes a displacement.
- the mesh data transmission method and device can encode the mesh data as follows.
- Mesh data is a term including point cloud data.
- the point cloud data (which may be referred to as point cloud for short) according to the embodiments can refer to data including vertex coordinates (or referred to as geometry information) and color information (or referred to as attribute information).
- vertex coordinates or referred to as geometry information
- color information or referred to as attribute information
- patch information the geometry image, attribute image, accuracy map, and additional information generated through patch generation and packing based on the vertex coordinates and the color information
- the point cloud data including the connection information can be referred to as mesh data.
- the point cloud and mesh data can be used interchangeably.
- the V-Mesh compression (decompression) method may include intra frame encoding (FIG. 6) and inter frame encoding (FIG. 7).
- intra frame encoding or inter frame encoding is performed.
- the data to be compressed can be a base mesh, displacement, attribute map, etc.
- the data to be compressed can be a displacement, an attribute map, and a motion field between a reference base mesh and a current base mesh.
- Fig. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments.
- Each component for the intra-frame encoding process of Fig. 6 corresponds to hardware, software, a processor, and/or a combination thereof.
- the encoding process of Fig. 6 illustrates the encoding of the mesh video encoder (102) of Fig. 1 in detail. That is, it illustrates the configuration of the mesh video encoder (102) when the encoding of Fig. 1 is an intra-frame method.
- the encoder of Fig. 6 may include a pre-processor (200) and/or an encoder (201).
- the pre-processor (200) and the encoder (201) of Fig. 6 may correspond to the pre-processor (200) and the encoder (201) of Fig. 3.
- the pre-processor (200) can receive an input mesh and perform the pre-processing described above.
- the pre-processing can generate a base mesh and/or a fitted sub-divided mesh.
- the quantizer (411) of the encoder (201) can quantize the base mesh and/or the fitted subdivided mesh.
- the static mesh encoder (412) can encode the static mesh (i.e., the quantized base mesh) and generate a bitstream (i.e., a compressed base mesh bitstream) including the encoded base mesh.
- the static mesh decoder (413) can decode the encoded static mesh (i.e., the encoded base mesh).
- the inverse quantizer (414) can inversely quantize the quantized static mesh (i.e., the base mesh) to output a reconstructed (or restored) base mesh.
- the displacement calculation unit (415) can generate displacements (or displacements) based on the reconstructed static mesh (i.e., the base mesh) and the fitted subdivided mesh. According to embodiments, the displacement calculation unit (415) calculates displacement, which is the position difference between each vertex of the subdivided base mesh and the fitted subdivided mesh after subdividing (or refining) the restored base mesh. In other words, the displacement is a displacement vector, which is the position difference between the vertices of the two meshes so that the fitted subdivided (or refining) mesh becomes similar to the original mesh.
- the forward linear lifting unit (416) can perform lifting transformation on the input displacement to generate lifting coefficients (or transform coefficients).
- the quantizer (417) can quantize the lifting coefficients.
- the image packing unit (418) can pack an image based on the quantized lifting coefficients.
- the video encoder (419) can encode the packed image. That is, the quantized lifting coefficients are packed into one frame as a 2D image by the image packing unit (418), compressed through the video encoder (419), and output as a displacement bitstream (i.e., compressed displacement bitstream).
- a video decoder (420) decodes a compressed displacement bitstream.
- An image unpacking unit (421) can perform unpacking on a decoded displacement frame to output quantized lifting coefficients.
- An inverse quantizer (422) can inverse quantize the quantized lifting coefficients.
- An inverse linear lifting unit (423) applies inverse lifting to the inverse quantized lifting coefficients to generate restored displacement.
- a mesh restoration unit (424) restores a reconstructed and deformed mesh through the restored displacement output from the inverse linear lifting unit (423) and the restored base mesh (or referred to as a subdivided restored base mesh) output from the inverse quantization unit (414).
- the present disclosure refers to the reconstructed and deformed mesh as a restored deformed mesh.
- the attribute transfer (425) receives an input mesh and/or an input attribute map, and regenerates an attribute map based on the restored deformed mesh.
- the attribute map means a texture map corresponding to attribute information among mesh data components, and in the present disclosure, the attribute map and the texture map may be used interchangeably.
- the push-pull padding (426) may pad data in the attribute map based on the push-pull method.
- the color space conversion unit (427) may convert the space of the color component of the attribute map. For example, the attribute map may be converted from an RGB color space to a YUV color space.
- the video encoder (428) may encode the attribute map and output it as a compressed attribute bitstream.
- a multiplexer (430) can multiplex a compressed base mesh bitstream, a compressed displacement bitstream, and a compressed attribute bitstream to generate a compressed bitstream.
- the displacement calculation unit (415) may be included in the pre-processor (200).
- at least one of the quantizer (411), the static mesh encoder (412), the static mesh decoder (413), and the inverse quantizer (414) may be included in the pre-processor (200).
- the intra frame encoding method includes base mesh encoding (also called static mesh encoding). That is, when performing intra frame encoding on the current input mesh frame, the base mesh generated in the pre-processing process of the pre-processor (200) may be encoded using a static mesh compression technique in a static mesh encoder (412) after undergoing a quantization process in a quantizer (411).
- base mesh encoding also called static mesh encoding
- Draco technology is applied to base mesh encoding, and vertex position information, mapping information (texture coordinates), vertex connection information, etc. of the base mesh become compression targets.
- the encoder of Fig. 6 generates a bitstream by compressing the base mesh, displacement, and attributes within a frame
- the encoder of Fig. 7 generates a bitstream by compressing the motion, displacement, and attributes between the current frame and the reference frame.
- Fig. 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments.
- Each component for the inter-frame encoding process of Fig. 7 corresponds to hardware, software, a processor, and/or a combination thereof.
- the encoding process of Fig. 7 illustrates the encoding of Fig. 1 in detail. That is, it illustrates the configuration of an encoder when the encoding of Fig. 1 is an inter-frame method.
- the encoder of Fig. 7 may include a pre-processor (200) and/or an encoder (201).
- the pre-processor (200) and the encoder (201) of Fig. 7 may correspond to the pre-processor (200) and the encoder (201) of Fig. 3.
- the motion encoder (512) can obtain a motion vector between two base meshes based on the restored quantized reference base mesh and the quantized current base mesh, and then encode the motion vector to output a compressed motion bitstream.
- the motion encoder (512) can be referred to as a motion vector encoder.
- the base mesh restoration unit (513) can restore the base mesh based on the restored quantized reference base mesh and the encoded motion vector.
- the restored base mesh is dequantized in the dequantizer (514) and then output to the displacement calculation unit (515).
- the displacement calculation unit (515) may be included in the pre-processor (200). Additionally, at least one of the quantizer (511), the motion encoder (512), the base mesh restoration unit (513), and the inverse quantizer (514) may be included in the pre-processor (200).
- the inter-frame encoding method may include motion field encoding (or motion vector encoding).
- Inter-frame encoding may be performed when a reference mesh and a current input mesh have a one-to-one correspondence of vertices and only the position information of the vertices is different.
- the difference between the vertices of the reference base mesh and the current base mesh i.e., the motion field (or motion vector) may be calculated and this information may be encoded.
- the reference base mesh is the result of quantizing the already decoded base mesh data and is determined according to the reference frame index determined in the GoF generation.
- the motion field may also be encoded as a value.
- the predicted motion field may be calculated by averaging the motion fields of the restored vertices among the vertices connected to the current vertex, and the residual motion field, which is the difference between the predicted motion field value and the motion field value of the current vertex, may be encoded.
- the residual motion field value may be encoded using entropy coding.
- the process of encoding the displacement and attribute map, excluding the motion field encoding process of the inter frame encoding is the same as the remaining structure of the intra frame encoding method excluding the base mesh encoding.
- Figure 8 shows a lifting conversion process for displacement according to embodiments.
- Figure 9 illustrates a process of packing transformation coefficients (or lifting coefficients) according to embodiments into a 2D image.
- Figures 8 and 9 illustrate the process of transforming displacement and packing transform coefficients of the encoding process of Figures 6 and 7, respectively.
- the encoding method according to the embodiments includes displacement encoding.
- a reconstructed base mesh is generated through restoration and dequantization, and displacement between the result of performing subdivision on the reconstructed base mesh and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated (see 415 of FIG. 6 or 515 of FIG. 7).
- a data transform process such as wavelet transform can be applied to the displacement information (see 416 of FIG. 6 or 516 of FIG. 7).
- FIG. 8 shows a process of transforming displacement information using a lifting transform in the forward linear lifting unit (416) of FIG. 6 or the wavelet transformer (516) of FIG. 7.
- a linear wavelet-based lifting transform may be performed.
- the transform coefficients generated through the transform process are quantized in a quantizer (417 or 517) and then packed into a 2D image through an image packing unit (418 or 518) as in FIG. 9.
- the horizontal number of blocks is fixed to 16, but the vertical number of blocks can be determined according to the number of vertices of a subdivided base mesh.
- Transform coefficients can be packed by aligning them with Morton code within a block.
- the packed images generate a displacement video for each GoF unit, and the displacement video can be encoded using a conventional video compression codec in a video encoder (419 or 519).
- a base mesh (original) may include vertices and edges for LoD0.
- a first sub-division mesh generated by dividing (or subdividing) the base mesh includes vertices generated by further dividing (or subdividing) edges of the base mesh.
- the first sub-division mesh includes vertices for LoD0 and vertices for LoD1.
- LoD1 includes the sub-divided vertices and the vertices (LoD0) of the base mesh.
- the first sub-division mesh may be divided (or subdivided) again to generate a second sub-division mesh.
- the second sub-division mesh includes LoD2.
- LoD2 includes base mesh vertices (LoD0), LoD1 including vertices further divided (or subdivided) from LoD0, and vertices further divided (or subdivided) from LoD1.
- LoD is a level of detail of mesh data content, and as the index of the level increases, the distance between vertices becomes closer and the level of detail increases. In other words, the smaller the LoD value, the lower the detail of the mesh data content, and the larger the LoD value, the higher the detail of the mesh data content.
- LoD N includes the vertices included in the previous LoDN-1 as they are.
- the mesh When the mesh (or vertex) is further divided through subdivision, considering the previous vertices v1, v2 and the subdivided vertex v, the mesh can be encoded based on the prediction and/or update method. Instead of encoding the information about the current LoD N as it is, the residual value between the previous LoD N-1 can be generated and the mesh can be encoded through the residual value to reduce the size of the bitstream.
- the prediction process means the operation of predicting the current vertex v through the previous vertices v1 and v2. Since adjacent subdivision meshes have similar data, efficient encoding can be performed by utilizing this property.
- the current vertex position information is predicted as a residual for the previous vertex position information, and the previous vertex position information is updated through the residual.
- vertex, apex, and point may be used with the same meaning.
- LoDs may be defined in the subdivision process of the base mesh. According to embodiments, the subdivision process of the base mesh may be performed in the pre-processor (200) or may be performed in a separate component/module.
- a vertex has a transform coefficient (or lifting coefficient) generated through a lifting transformation.
- the transform coefficient of a vertex related to a lifting transformation can be packed into an image by an image packing unit (418 or 518) and then encoded by a video encoder (419 or 519).
- Figure 10 illustrates the attribute transfer process of the V-MESH compression method according to embodiments.
- FIG. 10 illustrates the detailed operation of attribute transfer (425 or 525) of encodings such as FIG. 6, FIG. 7, etc.
- Encoding according to embodiments includes attribute map encoding.
- attribute map encoding may be performed in the video encoder (428) of FIG. 6 or the video encoder (528) of FIG. 7.
- the encoder compresses information about an input mesh through base mesh encoding (i.e., intra encoding), motion field encoding (i.e., inter encoding), and displacement encoding.
- base mesh encoding i.e., intra encoding
- motion field encoding i.e., inter encoding
- displacement encoding the compressed input mesh is reconstructed through base mesh decoding (intra frame), motion field decoding (inter frame), and displacement video decoding, and the reconstructed result, the reconstructed deformed mesh (hereinafter referred to as Recon. deformed mesh), is used to compress an input attribute map, as shown in FIGS. 6 and 7.
- Recon. deformed mesh is used to compress an input attribute map, as shown in FIGS. 6 and 7.
- the reconstructed deformed mesh (Recon.
- deformed mesh has position information of vertices, texture coordinates, and connection information corresponding thereto, but does not have color information corresponding to the texture coordinates. Therefore, as shown in Fig. 10, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is regenerated through the attribute transfer process of attribute transfer (425 or 525).
- attribute transfer (425 or 525) first checks whether each point P(u, v) of a 2D texture domain belongs to a texture triangle of a reconstructed deformed mesh, and if it exists in a texture triangle T, the barycentric coordinate of P(u, v) according to the triangle T ( , , ) is calculated. And the 3D vertex positions of triangle T and ( , , ) is used to compute the 3D coordinates M(x, y, z) of P(u, v). Find the vertex coordinates M'(x', y', z') and the triangle T' containing this vertex, which corresponds to the most similar position to the computed M(x, y, z) in the input mesh domain.
- a new attribute map generated via attribute transfer (425 or 525) is grouped into GoF units to form an attribute map video, which is compressed using the video codec of the video encoder (428 or 528).
- the decoding processing of Fig. 1 can perform the reverse process of the corresponding process of the encoding process of Fig. 1.
- the specific decoding process is as follows.
- Figure 11 illustrates an intra frame decoding (or intra decoding) process of V-Mesh technology according to embodiments.
- Fig. 11 illustrates the configuration and operation of the mesh video decoder (113) of the receiving device of Fig. 1.
- Fig. 11 can restore mesh data by performing the reverse process of the intra frame encoding process of Fig. 6.
- Each component for the intra frame decoding process of Fig. 11 corresponds to hardware, software, and/or a combination thereof.
- the bitstream (i.e., compressed bitstream) received and input to the demultiplexer (611) of the intra frame decoding unit (610) can be separated into a mesh sub-stream, a displacement sub-stream, an attribute map sub-stream, and a sub-stream including patch information of the mesh such as V-PCC/V3C.
- V-PCC Video-based Point Cloud Compression
- V3C Visual Volumetric Video-based Coding
- the mesh sub-stream may be input to a static mesh decoder (612) and decoded
- the displacement sub-stream may be input to a video decoder (613) and decoded
- the attribute map sub-stream may be input to a video decoder (617) and decoded.
- the mesh sub-stream is decoded through a decoder (612) of a static mesh codec used in encoding, such as Google Draco, to thereby reconstruct a quantized base mesh, for example, connection information of the base mesh, vertex geometry information, vertex texture coordinates, etc.
- a decoder 612 of a static mesh codec used in encoding, such as Google Draco
- a displacement sub-stream is decoded into a displacement video through a decoder (613) of a video compression codec used in encoding, and is restored as displacement information for each vertex through an image unpacking process of an image unpacking unit (614), an inverse quantization process of an inverse quantizer (615), and an inverse transform process of an inverse linear lifting unit (616) (i.e., Recon. displacements).
- a base mesh restored by a static mesh decoder (612) is inverse quantized by an inverse quantizer (620) and then output to a mesh restoration unit (630).
- the mesh restoration unit (630) reconstructs and restores a deformed mesh (i.e., decoded mesh) through the restored displacement output from the inverse linear lifting unit (616) and the restored base mesh output from the inverse quantization unit (620). That is, the inverse quantized restored base mesh is combined with the restored displacement information to generate a final decoded mesh.
- the final decoded mesh is referred to as a reconstructed deformed mesh.
- an attribute map sub-stream is decoded through a decoder (617) corresponding to a video compression codec used in encoding, and then restored to a final attribute map (i.e., decoded attribute map) through processes such as color format conversion and color space conversion in a color conversion unit (640).
- the restored decoded mesh and decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.
- the received compressed bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream.
- the substream is interpreted as a term referring to some bitstreams included in the bitstream.
- the bitstream includes patch information (data), mesh information (data), displacement information (data), and attribute map information (data).
- the decoder of FIG. 11 performs the following intra-frame decoding operations.
- the static mesh decoder (612) decodes the mesh sub-stream to generate a reconstructed quantized base mesh, and the inverse quantizer (620) applies the quantization parameters of the quantizer inversely to generate the reconstructed base mesh.
- the video decoder (613) decodes the displacement sub-stream, the image unpacking unit (614) unpacks the images of the decoded displacement video, and the inverse quantizer (615) inversely quantizes the quantized images.
- the inverse linear lifting unit (616) applies a lifting transform in the reverse process of the encoder to generate the reconstructed displacement.
- the mesh restoration unit (630) generates a reconstructed deformed mesh based on the reconstructed base mesh and the reconstructed displacement.
- a video decoder (617) decodes an attribute map sub-stream, and a color converter (640) converts a color format and/or space of the decoded attribute map to generate a decoded attribute map.
- Figure 12 illustrates the inter-frame decoding (or inter-decoding) process of V-Mesh technology.
- Fig. 12 shows the configuration and operation of the mesh video decoder (113) of the receiving device of Fig. 1.
- Fig. 12 can restore mesh data by performing the reverse process of the inter-frame encoding process of Fig. 7.
- Each component for the inter-frame decoding process of Fig. 12 corresponds to hardware, software, and/or a combination thereof.
- the bitstream received and input to the demultiplexer (711) of the intra frame decoding unit (710) can be separated into a motion sub-stream (or motion vector sub-stream), a displacement sub-stream, an attribute map sub-stream, and a sub-stream including patch information of a mesh such as V3C/V-PCC.
- a motion sub-stream may be input to a motion decoder (712) and decoded
- a displacement sub-stream may be input to a video decoder (713) and decoded
- an attribute map sub-stream may be input to a video decoder (717) and decoded.
- a motion sub-stream is decoded through an entropy decoding and inverse prediction process in a motion decoder (712) and restored into motion information (or motion vector information).
- a base mesh restoration unit (718) generates a reconstructed quantized base mesh for a current frame by combining the reconstructed motion information and a reference base mesh that has already been restored and stored.
- An inverse quantizer (720) generates a reconstructed base mesh by applying inverse quantization to the reconstructed quantized base mesh.
- a video decoder (713) decodes a displacement sub-stream, an image unpacking unit (714) unpacks an image of a decoded displacement video, and an inverse quantizer (715) inversely quantizes a quantized image.
- the inverse linear lifting unit (716) applies a lifting transformation in the reverse process of the encoder to generate a restored displacement.
- the mesh restoration unit (730) generates a reconstructed deformed mesh, i.e., a final decoded mesh, based on the reconstructed base mesh and the restored displacement.
- the video decoder (717) decodes an attribute map sub-stream in the same manner as intra decoding, and the color conversion unit (740) converts a color format and/or space of the decoded attribute map to generate a decoded attribute map.
- the decoded mesh and the decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.
- the bitstream includes motion information (or motion vector), displacement, and attribute map. Since Fig. 12 performs inter-frame decoding, it further includes a process of decoding inter-frame motion information.
- the motion information is decoded, and a restored quantized base mesh for the motion information is generated based on a reference base mesh to generate a restored base mesh.
- Fig. 12 which is identical to Fig. 11, refer to the description of Fig. 11.
- Fig. 13 illustrates a mesh data transmission device according to embodiments.
- FIG. 13 corresponds to the transmitting device (100) or mesh video encoder (102) of FIG. 1, the encoder (pre-processor and encoder) of FIG. 2, FIG. 6, or FIG. 7, and/or the transmitting encoding device corresponding thereto.
- Each component of FIG. 13 corresponds to hardware, software, a processor, and/or a combination thereof.
- the operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in Fig. 13.
- the transmitter of Fig. 13 may perform an intra-frame encoding (or intra-encoding or intra-screen encoding) process and/or an inter-frame encoding (or inter-encoding or inter-screen encoding) process.
- the pre-processor (811) receives an original mesh as input and generates a simplified mesh (or base mesh) and a fitted decimated mesh (or subdivision). Simplification can be performed based on the target number of vertices or target number of polygons constituting the mesh. Parameterization, which generates texture coordinates and texture connection information per vertex, can be performed on the simplified mesh. For example, parameterization is a process of mapping a 3D surface to a texture domain for the decimated mesh. If parameterization is performed using the UVAtlas tool, mapping information that can identify where each vertex of the decimated mesh can be mapped on a 2D image is generated. The mapping information is expressed and stored as texture coordinates, and the final base mesh is generated through this process.
- a task of quantizing floating-point type mesh information into fixed-point type can be performed.
- This result can be output as a base mesh to a motion vector encoder (813) or a static mesh encoder (814) through a switching unit (812).
- mesh subdivision can be performed on the base mesh to generate additional vertices.
- vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated.
- the pre-processor (811) can generate a fitted subdivided mesh by adjusting vertex positions so that the subdivided mesh becomes similar to the original mesh.
- the base mesh is output to a motion vector encoder (813) through a switching unit (812) when performing inter encoding for the corresponding mesh frame, and is output to a static mesh encoder (814) through a switching unit (812) when performing intra encoding for the corresponding mesh frame.
- the motion vector encoder (813) may be referred to as a motion encoder.
- the base mesh when performing intra encoding or intra frame encoding for the corresponding mesh frame, can be compressed through a static mesh encoder (814).
- encoding can be performed on connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh.
- the base mesh bitstream generated through encoding is transmitted to a multiplexer (823).
- the motion vector encoder (813) when performing inter encoding (or inter frame encoding) for the corresponding mesh frame, can receive a base mesh and a reference reconstructed base mesh (or a reconstructed quantized reference base mesh) as inputs, calculate a motion vector between the two meshes, and encode the value. In addition, the motion vector encoder (813) can perform prediction based on connection information using a previously encoded/decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through the encoding is transmitted to the multiplexer (823).
- the base mesh restoration unit (815) can receive the base mesh encoded by the static mesh encoder (814) or the motion vector encoded by the motion vector encoder (813) and generate a reconstructed base mesh.
- the base mesh restoration unit (815) can perform static mesh decoding on the base mesh encoded by the static mesh encoder (814) to restore the base mesh.
- quantization can be applied before the static mesh decoding, and inverse quantization can be applied after the static mesh decoding.
- the base mesh restoration unit (815) can restore the base mesh based on the reconstructed quantized reference base mesh and the motion vector encoded by the motion vector encoder (813).
- the reconstructed base mesh is output to the displacement calculation unit (816) and the mesh restoration unit (820).
- the displacement calculation unit (816) can perform mesh refinement on the restored base mesh.
- the displacement calculation unit (816) can calculate a displacement vector, which is a difference value of vertex positions between the restored base mesh that has been refined and the fitted subdivision (or refined) mesh generated by the pre-processor (811). At this time, the displacement vector can be calculated as many times as the number of vertices of the refined mesh.
- the displacement calculation unit (816) can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.
- the displacement vector video generation unit (817) may include a linear lifting unit, a quantizer, and an image packing unit. That is, in the displacement vector video generation unit (817), the linear lifting unit may transform the displacement vector for effective encoding. The transformation may be performed by a lifting transformation, a wavelet transformation, etc. according to embodiments. In addition, quantization may be performed on the transformed displacement vector value, that is, the transform coefficient, in a quantizer. At this time, a different quantization parameter may be applied to each axis of the transform coefficient, and the quantization parameter may be derived by a promise of the encoder/decoder.
- the displacement vector information that has undergone transformation and quantization may be packed into a 2D image in the image packing unit.
- the displacement vector video generation unit (817) may generate a displacement vector video by bundling packed 2D images for each frame, and the displacement vector video may be generated for each GoF (Group of Frame) unit of the input mesh.
- GoF Group of Frame
- the displacement vector video encoder (818) can encode the generated displacement vector video using a video compression codec.
- the generated displacement vector video bitstream is transmitted to a multiplexer (823).
- the displacement vector restoration unit (819) may include a video decoder, an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. That is, the displacement vector restoration unit (819) performs decoding on an encoded displacement vector in the video decoder, performs image unpacking in the image unpacking unit, performs inverse quantization in the inverse quantizer, and then performs inverse transformation in the inverse linear lifting unit to restore the displacement vector.
- the restored displacement vector is output to the mesh restoration unit (820).
- the mesh restoration unit (820) restores a deformed mesh based on the base mesh restored by the base mesh restoration unit (815) and the displacement vector restored by the displacement vector restoration unit (819).
- the restored mesh (or restored deformed mesh) has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
- the texture map video generation unit (821) can regenerate a texture map based on a texture map (or attribute map) of an original mesh and a restored deformed mesh output from a mesh restoration unit (820). According to embodiments, the texture map video generation unit (821) can assign color information per vertex of a texture map of an original mesh to texture coordinates of a restored deformed mesh. According to embodiments, the texture map video generation unit (821) can generate a texture map video by grouping regenerated texture maps by GoF units for each frame.
- the generated texture map video can be encoded using a video compression codec of the texture map video encoder (822).
- the texture map video bitstream generated through encoding is transmitted to a multiplexer (823).
- a multiplexer (823) multiplexes a motion vector bitstream (e.g., in case of inter encoding), a base mesh bitstream (e.g., in case of intra encoding), a displacement vector bitstream, and a texture map bitstream into one bitstream.
- the one bitstream can be transmitted to a receiver via a transmitter (824).
- the motion vector bitstream, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream can be generated as a file with one or more track data or encapsulated into segments and transmitted to a receiver via the transmitter (824).
- a transmitting device can encode a mesh in an intra-frame or inter-frame manner.
- a transmitting device according to intra-encoding can generate a base mesh, a displacement vector (or referred to as displacement), and a texture map (or referred to as attribute map).
- a transmitting device according to inter-encoding can generate a motion vector (or referred to as motion), a displacement vector (or referred to as displacement), and a texture map (or referred to as attribute map).
- a texture map obtained from a data input unit is generated and encoded based on a restored mesh. Displacement is generated and encoded through a vertex position difference between a base mesh and a divided (or subdivided or subdivided) mesh.
- the displacement is a position difference between a fitted subdivided mesh and a subdivided restored base mesh, that is, a vertex position difference value between the two meshes.
- the base mesh is generated by simplifying and encoding an original mesh through pre-processing.
- Motion is generated as motion vectors for the mesh of the current frame based on the reference base mesh of the previous frame.
- Fig. 14 illustrates a mesh data receiving device according to embodiments.
- Fig. 14 corresponds to the receiving device (110) or mesh video decoder (113) of Fig. 1, the decoder of Fig. 11 or Fig. 12, and/or the receiving decoding device corresponding thereto.
- Each component of Fig. 14 corresponds to hardware, software, a processor, and/or a combination thereof.
- the receiving (decoding) operation of Fig. 14 can follow the reverse process of the corresponding process of the transmitting (encoding) operation of Fig. 13.
- the bitstream of the mesh data received by the receiver (910) is demultiplexed into a compressed motion vector bitstream (e.g., inter decoding) or a base mesh bitstream (e.g., intra decoding), a displacement vector bitstream, and a texture map bitstream after file/segment decapsulation by the demultiplexer (911).
- a compressed motion vector bitstream e.g., inter decoding
- a base mesh bitstream e.g., intra decoding
- a displacement vector bitstream e.g., a displacement vector bitstream
- a texture map bitstream after file/segment decapsulation
- the motion vector decoder (913) may be referred to as a motion decoder.
- the motion vector decoder (913) can perform decoding on the motion vector bitstream. According to embodiments, the motion vector decoder (913) can reconstruct the final motion vector by adding the previously decoded motion vector as a predictor and the residual motion vector decoded from the bitstream.
- the static mesh decoder (914) can decode the base mesh bitstream to restore connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh.
- the base mesh restoration unit (915) can restore the current base mesh based on the decoded motion vector or the decoded base mesh. For example, if the current mesh is subject to inter-screen encoding, the base mesh restoration unit (915) can generate a restored base mesh by adding the decoded motion vector to the reference base mesh and then performing inverse quantization. As another example, if the current mesh is subject to intra-screen encoding, the base mesh restoration unit (915) can generate a restored base mesh by performing inverse quantization on the decoded base mesh through the static mesh decoder (914).
- the displacement vector video decoder (917) can decode the displacement vector bitstream as a video bitstream using a video codec.
- the displacement vector restoration unit (918) extracts displacement vector transform coefficients from the decoded displacement vector video, and applies inverse quantization and inverse transformation processes to the extracted displacement vector transform coefficients to restore the displacement vector.
- the displacement vector restoration unit (918) may include an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. If the restored displacement vector is a value of a local coordinate system, a process of inversely transforming it to a Cartesian coordinate system may be performed.
- the mesh restoration unit (916) can perform subdivision on the restored base mesh to generate additional vertices. Through subdivision, vertex connection information including the added vertices, texture coordinates, and texture coordinate connection information, etc. can be generated. At this time, the mesh restoration unit (916) can combine the subdivided restored base mesh with the restored displacement vector to generate a final restored mesh (or a restored deformed mesh).
- the texture map video decoder (919) can decode the texture map bitstream as a video bitstream using a video codec to restore the texture map.
- the restored texture map has color information for each vertex contained in the restored mesh, and the color value of each vertex can be obtained from the texture map using the texture coordinate of each vertex.
- the mesh restored by the mesh restoration unit (916) and the texture map restored by the texture map video decoder (919) are shown to the user through a rendering process in the mesh data renderer (920).
- a receiving device can decode a mesh in an intra-frame or inter-frame manner.
- a receiving device according to intra-decoding can receive a base mesh, a displacement vector (or referred to as displacement), a texture map (or referred to as attribute map), and render mesh data based on the restored mesh and the restored texture map.
- a receiving device according to inter-decoding can receive a motion vector (or referred to as motion), a displacement vector (or referred to as displacement), a texture map (or referred to as attribute map), and render mesh data based on the restored mesh and the restored texture map.
- a mesh data transmission device and method can pre-process mesh data, encode the pre-processed mesh data, and transmit a bitstream including the encoded mesh data.
- a point mesh data reception device and method can receive a bitstream including mesh data and decode the mesh data.
- the mesh data transmission and reception method/device according to embodiments may be referred to as the method/device according to embodiments.
- the mesh data transmission and reception method/device according to embodiments may also be referred to as a 3D data transmission and reception method/device or a point cloud data transmission and reception method/device.
- the V-Mesh method converts displacement information generated during the encoding process into a video format and then compresses it using an existing 2D video codec. Then, the compressed displacement information is restored by performing the reverse process of the 2D video codec.
- the V-DMC encoder calculates a displacement vector which is the difference between the mesh restored from the base mesh and the mesh fitted in the pre-processing step, and converts the calculated displacement vector in the canonical coordinate system (i.e., in the form of x, y, and z) into a displacement vector in the local coordinate system (i.e., in the form of normal, tangential, and bi-tangential), and then performs lifting transform and quantization on the displacement vector in the local coordinate system to encode it into a displacement vector bitstream.
- the V-DMC decoder (or decoder or decoding device) performs the reverse process of the V-DMC encoder to restore the displacement vector.
- the encoder performs coordinate transformation calculations for the three components of normal, tangential, and bi-tangential, and performs a series of displacement vector encoding processes like the above for each component, which may be inefficient in terms of compression efficiency and transmission speed.
- the present disclosure proposes a method for more efficient compression and transmission, in which an encoder of a transmitting device packs only the normal and tangential components of a displacement vector and skips packing the bi-tangential component to perform encoding and transmit the same as a bitstream, and a decoder of a receiving device calculates the bi-tangential component using the two normal and tangential components received, and decodes the bi-tangential component into a final displacement vector based on the calculated bi-tangential component.
- the encoder of the transmitting device encodes and transmits only the normal component and the tangential component excluding the bi-tangential component among the normal, tangential, and bi-tangential components converted to the local coordinate system
- the decoder of the receiving device proposes a method of decoding up to the bi-tangential component using the received normal and tangential components.
- the present disclosure proposes a packing method and a signaling method according to various image packing formats. In this way, compared to transmitting all of the existing normal, tangential, and bi-tangential components, by compressing and transmitting only two components, an efficient bit saving effect and better image quality mesh data can be obtained.
- the present disclosure relates to V-DMC, a method for compressing 3D dynamic mesh data based on an existing 2D video codec, and describes a device and method for encoding/decoding in units of displacement vectors in the displacement vector transformation and quantization steps, and syntax and semantics information related thereto.
- the present disclosure describes a method for encoding and decoding by packing only two normal and tangential components among three normal, tangential, and bi-tangential components of a displacement vector, a 2D image packing method according to a displacement vector video format, whether to skip packing for the bi-tangential component of a displacement vector, and a signaling method according to an image packing format.
- the operations of a transmitter and a receiver to which the same are applied are described.
- geometry information is one of the elements constituting a mesh, and includes a vertex (or point), an edge, a polygon, etc.
- a vertex defines a position in 3D space
- an edge represents connection information between vertices
- a polygon forms a surface of the mesh with a combination of edges and vertices. That is, each vertex constituting the mesh represents a position in 3D space, and is expressed as, for example, x, y, z coordinates (i.e., canonical coordinate system).
- a polygon can be a triangle or a square. That is, geometry forms the skeleton of a 3D model, and thereby defines the shape of the model and is visually expressed when rendered.
- vertex, apex, and point may be used with the same meaning. That is, a vertex has a coordinate in 3D space, and a polygon of a triangle or a square can be generated through a connection between a plurality of vertices.
- V-DMC referred to in the present disclosure may also be referred to as V-mesh, and the two terms are expressions used with the same meaning.
- displacement information can be acquired based on a refined mesh (or referred to as a sub-mesh). That is, a fitting process is performed so that the input mesh and the sub-mesh become similar to each other, and the difference in each vertex position between the fitted sub-division mesh generated by performing refinement on the restored base mesh and the generated refined restored base mesh is calculated.
- the present disclosure calls this vertex position difference value a displacement vector.
- the displacement vector may be used interchangeably with the same meaning as displacement or displacement information.
- the displacement video may be used interchangeably with the same meaning as displacement vector video or displacement vector conversion coefficient video
- the displacement vector may be used interchangeably with the same meaning as displacement vector conversion coefficient or displacement vector coefficient.
- Fig. 15 illustrates a transmitting device according to embodiments.
- the transmitting device of Fig. 15 may be referred to as a mesh data transmitting device or an encoder or an encoder of a transmitting device or a V-Mesh encoder or a dynamic mesh encoder.
- FIG. 15 corresponds to the transmitting device (100) or mesh video encoder (102) of FIG. 1, the encoder (preprocessor and encoder) of FIG. 2, FIG. 6, or FIG. 7, the transmitting device of FIG. 13, and/or the transmitting encoding device corresponding thereto.
- Each component of FIG. 15 corresponds to hardware, software, a processor, and/or a combination thereof. The execution order of each block in FIG. 15 may be changed, some blocks may be omitted, and some blocks may be newly added.
- the operation process of a transmitter for compressing and transmitting dynamic mesh data using the V-Mesh compression technology may be as shown in FIG. 15.
- the transmitter of FIG. 15 may support both an intra-frame encoding (or intra-encoding or intra-screen encoding) process and/or an inter-frame encoding (or inter-encoding or inter-screen encoding) process.
- the mesh simplification unit (11011) simplifies the input original mesh through a mesh simplification algorithm to generate a base mesh (or a simplified base mesh or a simplified mesh).
- mesh simplification can be performed based on the number of target vertices or target polygons constituting the mesh.
- a method such as decimation can be used as a mesh simplification algorithm that simplifies the original mesh. That is, the decimation method can be a process of selecting a vertex to be removed from the original mesh with a certain reference point, and then removing the selected vertex and the triangle connected to the selected vertex.
- the mesh simplification unit (11011) can perform simplification of the input mesh by the target number of vertices or the target number of faces. At this time, the simplification process can be performed through various methods such as triangle collapse and edge collapse.
- a base mesh simplified in a mesh simplification unit (11011) is provided to a mesh parameterization unit (11012) and a mesh refinement unit (11018).
- the above mesh parameterization unit (11012) performs a process of mapping a 3D surface to a texture domain for a decimated mesh. That is, the mesh parameterization unit (11012) generates texture coordinates and texture connection information of an input mesh.
- the mesh parameterization unit (11012) may perform parameterization using a UV Atlas tool. Through this process, mapping information is generated regarding which location on a 2D image each vertex of the decimated mesh can be mapped to. The mapping information is expressed and stored as texture coordinates, and a final base mesh is generated through this process. That is, the mesh parameterization unit (11012) performs parameterization that generates texture coordinates (UV coordinates) and texture connection information per vertex of an input mesh (i.e., a simplified mesh or a simplified base mesh).
- the final base mesh (or base mesh having a texture map) generated in the above parameterization unit (11012) is input to the mesh quantization unit (11013) and quantized.
- the mesh quantization unit (11013) can perform a task of quantizing mesh information in a floating-point form (e.g., geometry information (x, y, z) or/and texture coordinates (u, v), normal information (nx, ny, nz), etc.) into a fixed-point form. That is, the mesh quantization unit (11013) can quantize vertex coordinates and texture coordinates of the base mesh.
- quantization for a specific component may be omitted.
- the above mesh subdivision unit (11018) subdivides the base mesh simplified by the mesh simplification unit (11011). That is, the mesh subdivision unit (11018) can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. At this time, depending on the subdivision method, geometry information connection information, texture coordinate connection information, and texture coordinates can be implicitly derived and generated. According to embodiments, the mesh subdivision unit (11018) can perform subdivision through a method such as mid-edge, Loop, or Catmul&Clark.
- the mesh fitting unit (11019) can perform fitting by adjusting vertex positions so that the mesh subdivided in the mesh subdivision unit (11018) becomes similar to the original mesh, thereby generating a fitted subdivided mesh.
- the present disclosure may be referred to as a pre-processor including a mesh simplification unit (11011), a mesh parameterization unit (11012), a mesh refinement unit (11018), and a mesh fitting unit (11019).
- the pre-processor may further include a displacement vector calculation unit (11020).
- the quantized base mesh in the mesh quantization unit (11013) can be output to a motion vector encoder (11015) or a static mesh encoder (11016) through a switching unit (11014).
- the base mesh is output to a motion vector encoder (11015) through a switching unit (11014) when performing inter encoding for the corresponding mesh frame, and is output to a static mesh encoder (11016) through a switching unit (11014) when performing intra encoding for the corresponding mesh frame.
- the motion vector encoder (11015) may be referred to as a motion encoder.
- the base mesh when performing intra encoding or intra frame encoding for the corresponding mesh frame, can be compressed through a static mesh encoder (11016).
- encoding can be performed on connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. That is, vertex coordinates, vertex connection information, texture coordinates, texture connection information, etc. of the mesh can be encoded in the static mesh encoder (11016).
- the base mesh bitstream generated through encoding is transmitted to a multiplexer (not shown).
- the motion vector encoder (11015) when performing inter encoding (or inter frame encoding) for the corresponding mesh frame, can receive a base mesh and a reference reconstructed base mesh (or a reconstructed quantized reference base mesh) as inputs, calculate a motion vector between the two meshes, and encode the value thereof.
- the motion vector encoder (11015) can perform prediction based on connection information using a previously encoded/decoded motion vector as a predictor, and entropy encode a differential motion vector (or residual motion vector) obtained by subtracting the predicted motion vector from the current motion vector.
- the motion vector bitstream generated through the encoding is transmitted to a multiplexer (not shown) as a base mesh bitstream.
- the static mesh bitstream is input to the multiplexer as a base mesh bitstream
- the motion vector bitstream is input to the multiplexer as a base mesh bitstream
- the base mesh decoder (11017) can receive a base mesh encoded by a static mesh encoder (11016) or a motion vector encoded by a motion vector encoder (11015) and generate a reconstructed base mesh.
- the base mesh decoder (11017) performs reconstruction of the base mesh according to the encoding type (inter-screen encoding or intra-screen encoding) of the current mesh.
- the base mesh decoder (11017) can perform static mesh decoding on the base mesh encoded by the static mesh encoder (11016) to reconstruct the base mesh.
- quantization can be applied before static mesh decoding, and inverse quantization can be applied after static mesh decoding.
- the base mesh decoder (11017) can restore the base mesh based on the restored quantized reference base mesh and the motion vector encoded by the motion vector encoder (11015). That is, when inter-screen encoding is performed, the current base mesh can be generated by decoding the motion vector by the motion vector decoding method and then applying (i.e., adding) the decoded motion vector to the reference restored base mesh. At this time, when the motion vector is not quantized, the motion vector restoration process is omitted and the current base mesh can be restored using the motion vector calculated by the motion vector encoder (11015). The restored base mesh is output to the displacement vector calculation unit (11020) and the mesh dequantization unit (11024).
- the displacement vector calculation unit (11020) can perform mesh refinement on the restored base mesh.
- the displacement vector calculation unit (11020) can calculate a displacement vector, which is a difference value of vertex positions between the restored base mesh that has been refined and the fitted subdivision (or refined) mesh generated by the mesh fitting unit (11019).
- the displacement vector can be calculated as many times as the number of vertices of the refined mesh. That is, the displacement vector of the number of vertices of the refined mesh can be calculated through the displacement vector calculation unit (11020).
- the displacement vector coordinate system transformation unit (11021) can output a vertex displacement vector calculated in a three-dimensional Cartesian coordinate system (i.e., (x, y, z) space) (or canonical coordinate system) as it is, or can transform it into a local coordinate system (i.e., normal, tangential, bi-tangential coordinate system) based on a normal vector of each vertex.
- a local coordinate system i.e., normal, tangential, bi-tangential coordinate system
- the normal vector can be calculated based on the geometry information and connection information of the surrounding vertices for each subdivided vertex.
- whether the displacement vector coordinate system is converted can be determined by an encoder (i.e., a transmitting device)/decoder (i.e., a receiving device) agreement, or a displacement vector coordinate system conversion status flag (asps_vmc_ext_displacement_coordinate_system), which is information that can identify whether the displacement vector coordinate system is converted, can be signaled to signaling information (e.g., atlas sequence parameter set, ASPS) and transmitted to the receiving device.
- an encoder i.e., a transmitting device
- decoder i.e., a receiving device
- a displacement vector coordinate system conversion status flag e.g., asps_vmc_ext_displacement_coordinate_system
- the canonical coordinate system is used as is, and if it is 1, it can indicate that the conversion to the local coordinate system is performed.
- the displacement vector in the canocial coordinate system or the displacement vector converted into the local coordinate system in the displacement vector coordinate system transformation unit (11021) is encoded into a displacement vector bitstream (or displacement vector video bitstream) in the displacement vector encoder (11022).
- the displacement vector encoder (11022) can encode the displacement vector in the canocial coordinate system or the displacement vector converted into the local coordinate system into the displacement vector bitstream using a 2D video encoder such as H.264, HEVC, or VVC.
- the displacement vector encoder (11022) can perform encoding on the displacement vector or the displacement vector coefficients (or displacement vector transform coefficients).
- the displacement vector encoder (11022) can perform encoding through a video codec-based encoder, a zero run-length encoder, an arithmetic encoder, etc.
- the displacement vector encoder (11022) can encode the displacement vector or the displacement vector coefficients by packing them into frames.
- the displacement vector coefficients can be packed into a 2D image and then encoded using a 2D video codec (i.e., a video compression codec), or zero run-length encoded, or arithmetic encoded to generate a displacement vector video bitstream.
- a 2D video codec i.e., a video compression codec
- a displacement vector video bitstream encoded and generated by a displacement vector encoder (11022) is transmitted to a multiplexer (not shown).
- a method for selecting encoding of the displacement vector encoder (11022) may use a displacement vector encoder promised in an encoder (i.e., a transmitting side)/decoder (i.e., a receiving side), or may analyze the characteristics of a displacement vector in an encoder on the transmitting side and transmit the type of a selected displacement vector encoder to a decoder on the receiving side.
- the displacement vector restoration unit (11023) may restore the displacement vector by performing the reverse process of displacement vector encoding on the displacement vector or displacement vector coefficient encoded by the displacement vector encoder (11022). That is, the displacement vector restoration unit (11023) may perform displacement vector depacking (or unpacking) depending on the method of encoding the displacement vector, for example, if the displacement vector is encoded based on a video codec. In addition, the displacement vector restoration unit (11023) may additionally perform dequantization, detransformation, etc. depending on whether quantization and transformation processes are performed in the displacement vector encoding process.
- the mesh dequantization unit (11024) can dequantize vertex coordinates or texture coordinates of the restored base mesh by the reverse process of quantization. If the quantization process is omitted in the mesh quantization unit (11013), the dequantization process is also omitted in the mesh dequantization unit (11024).
- the mesh restoration unit (11025) can restore a mesh based on a restored displacement vector output from a displacement vector restoration unit (11023) and a restored base mesh (or a dequantized restored base mesh) output from a mesh inverse quantization unit (11024). More specifically, the mesh restoration unit (11025) can perform subdivision on the restored base mesh output from the mesh inverse quantization unit (11024) and add the restored displacement vector from the displacement vector restoration unit (11023) to generate a reconstructed deformed mesh.
- the mesh restored by the mesh restoration unit (11025) (or referred to as a restored mesh or a restored deformed mesh) has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
- the restored mesh (or restored mesh or restored deformed mesh) generated in the above mesh restoration unit (11025) is provided to the texture map generation unit (11026).
- the texture map generation unit (11026) can regenerate the texture map of the current mesh based on the texture map (or attribute map) of the original mesh and the mesh restored by the mesh restoration unit (11025). That is, the texture map generation unit (11026) can generate the texture map of the restored mesh through the texture map of the original mesh and the relationship between the original mesh and the restored mesh.
- the texture map generation unit (11026) can assign color information per vertex of the texture map of the original mesh to the texture coordinates of the restored base mesh (or the restored deformed mesh). According to embodiments, the texture map generation unit (11026) can generate a texture map (or texture map video) by grouping the regenerated texture maps by GoF units for each frame.
- the texture map generated in the above texture map generation unit (11026) can be encoded in the texture map encoder (11027).
- the texture map encoder (11027) can encode the texture map using a 2D video codec-based encoder, a zero run length encoder, an entropy coding-based arithmetic encoder, etc.
- the texture map encoder (11027) can further perform color space conversion of the texture map. Then, the texture map substream (or texture map video bitstream) generated through the texture map encoding is transmitted to a multiplexer (not shown).
- the type of texture map encoder (11027) may include a video encoder (e.g., VVC, HEVC, etc.), an entropy coding-based encoder, etc.
- a method for selecting a texture map encoder (11027) may use a texture map encoder promised in an encoder (i.e., a transmitting side)/decoder (i.e., a receiving side), or may transmit the type of a texture map encoder selected by an encoder on the transmitting side to a decoder on the receiving side.
- a multiplexer may multiplex an input base mesh bitstream, a displacement vector bitstream, and a texture map bitstream into a single bitstream and then transmit the same to a receiving device.
- the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream may be encapsulated into a file/segment and transmitted to the receiving device.
- the bitstream multiplexed in the multiplexer may be transmitted over a network or stored in a digital storage medium.
- the network may include a broadcasting network and/or a communication network
- the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
- the V-DMC encoding device simplifies the original mesh and generates a base mesh through a mesh parameterization process.
- the generated base mesh is quantized, and in the case of an inter-frame, a motion vector is calculated from a previously referenced restored base mesh and the motion vector is encoded, and in the case of an intra-frame, it is transmitted as a base mesh bit stream through a static mesh encoder.
- the mesh data after the process of subdividing and fitting the simplified mesh from the original mesh and the mesh data restored from the previously encoded mesh are compared to calculate the displacement vector, which is the difference between each vertex.
- the displacement vector coordinate system is converted to a local coordinate system, and the displacement vector converted to the local coordinate system in the displacement vector encoder (12022) is converted and quantized into displacement vector coefficients, and then encoded and transmitted as a displacement vector bit stream.
- FIG. 16 is a block diagram showing an example of a displacement vector encoder according to embodiments.
- the displacement vector encoder (11022) may include a displacement vector conversion unit (12011), a displacement vector coefficient quantization unit (12012), a displacement vector coefficient packing unit (12013), and a displacement vector image/video encoding unit (12014).
- Each component of FIG. 16 corresponds to hardware, software, a processor, and/or a combination thereof.
- the execution order of each block in FIG. 16 may be changed, some blocks may be omitted, and some blocks may be newly added.
- the displacement vector encoder performs displacement vector coefficient image packing for each video image format and then encodes the image-packed displacement vector coefficients using a 2D video codec.
- the displacement vector transformation unit (12011) can perform linear lifting transformation, butterfly lifting transformation, wavelet transformation, etc. on the displacement vector of the (x, y, z) or (n, t, bt) coordinate system to transform the displacement vector of the (x, y, z) or (n, t, bt) coordinate system into displacement vector coefficients.
- n means normal
- t means tangential
- bt means bi-tangential.
- an average or distance-based weighted average prediction can be performed on n points near the vertex having a lower subdivision level than the current vertex based on connection information.
- a prediction can be performed based on a displacement vector of n vertices used to generate the current vertex in the mesh subdivision step.
- a process of updating a displacement vector of a vertex used for prediction can be performed using a residual signal generated by the prediction.
- the displacement vector conversion unit (12011) can perform displacement vector conversion only for the normal component and tangential component excluding the bi-tangential component.
- the displacement vector quantization unit (12012) can perform quantization on the displacement vector value, i.e., the displacement vector or the displacement vector coefficient, converted by the displacement vector transformation unit (12011). For example, if the displacement vector transformation unit (12011) performs transformation on all of the normal component, the tangential component, and the bi-tangential component, the displacement vector quantization unit (12012) performs quantization on the displacement vector coefficients of the normal component, the tangential component, and the bi-tangential component. As another example, if the displacement vector transformation unit (12011) performs transformation on only the normal component and the tangential component, the displacement vector quantization unit (12012) performs quantization on the displacement vector coefficients of the normal component and the tangential component. In the present disclosure, the displacement vector coefficient is used interchangeably with the displacement vector transformation coefficient to have the same meaning.
- the displacement vector quantization unit (12012) can derive a quantized value (quant) for each channel by multiplying a displacement vector coefficient (value) by a scale and adding an offset as in the following mathematical expression 1.
- each channel can be x, y, z or n, t, b according to the coordinate system of the displacement vector.
- displacement vector coefficient quantization can be performed only for the normal and tangential components.
- the mathematical expression 1 below is an equation for obtaining a quantized value for each channel
- the mathematical expression 2 is an equation for obtaining a scale value for each channel of the mathematical expression 1.
- the offset is in units of sequence or frame, and a fixed value can be used for each channel.
- the scale of mathematical expression 1 can be determined by the quantization parameter (QP) and the level-by-level scale (level_scale) as in mathematical expression 2.
- the level_scale in mathematical expression 2 can use a value determined by frame or sequence unit for each level.
- the scale can also be set as an individual value for each channel of the displacement vector.
- the displacement vector coefficient packing unit (12013) can pack the quantized displacement vector coefficients into a 2D image of the size of WxH. That is, if the displacement vector encoder (11022) encodes the displacement vector coefficients using a video codec-based encoding method, the process of packing the displacement vector coefficients into a frame as a 2D image is performed in the displacement vector coefficient packing unit (12013). In other words, the displacement vector coefficient packing unit (12013) packs the displacement vector coefficients into a 2D image, and the displacement vector image/video encoding unit (12014) can encode the packed 2D images using a video compression codec.
- FIG. 17 is a diagram showing an example of a displacement vector coefficient structure according to embodiments.
- the number of displacement vector coefficients included in LoD1 is N1 (LoD1)
- the number of displacement vector coefficients of level 0 (R 0 ) is N0
- the number of displacement vector coefficients of level 1 (R 1 ) is N1-N0.
- the number (N1) of displacement vector coefficients included in LoD1 is the sum of the number (N0) of displacement vector coefficients of level 0 (R 0 ) and the number (N1-N0) of displacement vector coefficients of level 1 (R 1 ).
- the displacement vector coefficient packing unit (12013) can pack N displacement vector coefficients (e.g., N quantized displacement vector coefficients) into an image of size W ⁇ H.
- W ⁇ H can be the size of a frame in which the displacement vector coefficients are packed as a 2D image.
- packing can be performed separately for the normal, tangential, and bi-tangential components. Or, some components can be skipped without packing. For example, the bi-tangential component can be skipped without packing.
- Fig. 18 is a drawing showing an example of packing displacement vector coefficients according to embodiments. That is, the displacement vector coefficients in the 1D form of Fig. 17 can be packed as a 2D image as in Fig. 18.
- the displacement vector coefficient packing unit (12013) can configure displacement vector coefficients (or transform coefficient levels) of a 1D vector or a scalar into blocks of a bx*by size, and pack each block according to a scanning order promised in an encoder/decoder, such as a z-scan order, a zig-zag scanning order, or a 2D Morton code order. That is, L*M displacement vector coefficient blocks can be block-packed into an image of a (bx*L) ⁇ (by*M) size according to an order defined in the encoder/decoder.
- L*M is the number of blocks having the bx*by size, and bx and by can each be 16.
- the displacement vector coefficients are configured as one block every 256 (16*16) units, and each block can be packed in a z-scan order or a zig-zag scan order.
- displacement vector coefficients within one block can be packed according to the z-scan order, the zig-zag scanning order, the 2D Morton code order, etc.
- one block can be configured with the size bx*by, and can be configured with L ⁇ M blocks determined according to the number N of displacement vector coefficients. At this time, packing can be performed into a 2D image of the size W ⁇ H according to the 2D Morton code, the zig-zag scanning order, etc.
- displacement vector coefficients of all levels are packed into one frame as a 2D image
- the packing can be performed sequentially in the scanning order from the displacement vector coefficient block of level 0 (R 0 ).
- the total size of the displacement vector coefficients is smaller than L * M
- padding is performed so that the size is (bx * L) ⁇ (by * M), as shown in Fig. 18. That is, if the total number of displacement vector coefficients is smaller than (bx * L) ⁇ (by * M), padding can be performed and filled so that the displacement vector image has the size of (bx * L) ⁇ (by * M).
- Fig. 18 shows an example in which padding is performed on five blocks.
- padding means filling the corresponding block with a meaningless value (e.g., 0).
- a meaningless value e.g., 0
- the packing can be performed sequentially in the scanning order from the displacement vector coefficient block of a smaller level.
- a 2D image can be packed by performing packing and padding in the reverse order of Fig. 18.
- L and M are determined according to the number (N) of displacement vector coefficients, or L (or M) is defined according to a convention of an encoder/decoder, and then M (or L) can be derived according to the number (N) of displacement vector coefficients as in the following mathematical expression 3.
- the following mathematical expression 3 is an example of deriving M when L is defined according to an encoder/decoder convention.
- the round function is a function that rounds the number in parentheses to the nearest integer.
- the packing method of the displacement vector coefficients of the displacement vector coefficient packing unit (12013) may be determined by a promise in the encoder/decoder, or the packing method performed in the displacement vector coefficient packing unit (12013) may be transmitted to the decoder of the receiving device.
- the displacement vector coefficients of all levels may be packed into one image (i.e., one frame) and the displacement vector image/video encoding unit (12014) may perform encoding, or the displacement vector coefficients may be packed into each image (i.e., each frame) for each subdivision level (R) and encoding may be performed respectively.
- the displacement vector coefficients of a specific level may be packed into one image, and the displacement vector coefficients of multiple or more different levels may be packed into another image.
- FIG. 19 is a diagram showing an example of packing displacement vector coefficient blocks by LoD according to embodiments.
- packing of displacement vector coefficient blocks (bx ⁇ by) may be performed as in FIG. 19 according to the CTU (Coding Tree Unit) size of a 2D video encoder encoded for each LoD level. At this time, it may be performed using an intermediate value or the last displacement vector coefficient value of an image to match the CTU size for each LoD.
- CTU is used in the process of dividing a video frame into coding tree units, and is a unit used particularly when encoding by dividing a video frame into a tree structure.
- Fig. 19 shows an example in which displacement vector coefficients of three levels (R 0 -R 2 ) are packed into one image. That is, it is an example in which displacement vector coefficients of R 0 , R 1 , and R 2 are packed into one frame. In this case, if the number of displacement vector coefficients of R 0 , R 1 , and R 2 is smaller than the size of the frame, a padding area may exist.
- the displacement vector coefficients of each level When packing the displacement vector coefficients of each level into one or more frames as in Fig. 19, the displacement vector coefficients can be packed into blocks of size bx ⁇ by.
- bx*by displacement vector coefficients within one displacement vector coefficient block may be packed according to a zig-zag scan order, or may be packed according to a 2D Moulton code order as shown in Fig. 20(a) and Fig. 20(b).
- FIG. 20(a) and FIG. 20(b) are diagrams showing examples of a method for packing displacement vector coefficients in a displacement vector coefficient block according to embodiments.
- the size of a frame into which displacement vector coefficients are packed can be obtained as (bx ⁇ L) ⁇ (by ⁇ M), where the bx and by parameters determine the width and height of a block, respectively, and the L and M parameters determine the number of blocks included in a row and column of the frame, respectively.
- the L and M parameters may be determined by a convention in the encoder/decoder, or L (or M) may be defined by a convention of the encoder/decoder, and then M (or L) may be derived according to the number of displacement vector coefficients (see Equation 3), or may be derived according to a level.
- displacement vector coefficients transformed into a local coordinate system may have normal components, tangential components, and bi-tangential components.
- the displacement vector coefficient packing unit (12013) may select a format such as YUV 4:4:4, YUV 4:2:0, YUV 4:0:0, etc. to perform image packing, and may configure displacement vector coefficients according to each format. In this process, packing of some components may be skipped.
- YUV 4:4:4, YUV 4:2:0, and YUV 4:0:0 formats represent formats for packing normal components, tangential components, and bi-tangential components of displacement vector coefficients into an image.
- the present disclosure can perform sampling of displacement vector coefficients V1 (N1, T1, B1) to V4 (N4, T4, B4) differently for each format in units of four vertices.
- the YUV 4:4:4 format means that the sizes of the Y channel (w y *h y ) , the U channel (w u *h u ), and the V channel (w v *h v ) are equal.
- the YUV 4:2:0 format means that the sizes of the U channel (w u *h u ) and the V channel (w v *h v ) are a predetermined multiple (e.g., 4 times) smaller than the size of the Y channel (w y * h y ) .
- the YUV 4:0:0 format means that only the size of the Y channel (w y *h y ) exists, that is, when packing is performed only in the Y channel. That is, among the Y, U, and V channels, only the Y channel exists.
- the YUV 4:4:4 format will be used interchangeably with the first format
- the YUV 4:2:0 format will be used interchangeably with the second format
- the YUV 4:0:0 format will be used interchangeably with the third format.
- the Y channel may be referred to as the first channel
- the U channel may be referred to as the second channel
- the V channel may be referred to as the third channel.
- first, second, third, etc. may be used to describe various components, but the components should not be limited by the terms. The terms are used to distinguish one component from another.
- the first component may be referred to as the second component, and similarly, the second component may also be referred to as the first component.
- the first component may be referred to as the third component.
- the encoder may signal image packing format information (ColourSpace_displacement_video) of a given displacement vector coefficient and information for identifying whether to skip bi-tangent component packing (Bi_tangent_skip_flag) in signaling information (e.g., ASPS) and transmit them to a receiving device (or a decoder of the receiving device).
- Image packing format information ColdSpace_displacement_video
- Bi_tangent_skip_flag bi-tangent component packing
- the image packing format information (ColourSpace_displacement_video) of the displacement vector coefficients can indicate whether the format used for packing the displacement vector coefficients is a YUV 4:4:4 format (i.e., a first format), a YUV 4:2:0 format (i.e., a second format), or a YUV 4:0:0 format (i.e., a third format).
- a YUV 4:4:4 format i.e., a first format
- a YUV 4:2:0 format i.e., a second format
- a YUV 4:0:0 format i.e., a third format.
- the value of the information (Bi_tangent_skip_flag) for identifying whether to skip bi-tangential component packing is 0, it can indicate that the bi-tangential component is not skipped, and if it is 1, it can indicate that the bi-tangential component is skipped.
- FIG. 21(a) and FIG. 21(b) are diagrams showing an example of packing displacement vector coefficients based on the YUV 4:4:4 format according to embodiments. That is, FIG. 21(a) and FIG. 21(b) are examples of not skipping packing of bi-tangential components when packing displacement vector coefficients based on the YUV 4:4:4 format.
- the value of the information (Bi_tangent_skip_flag) for identifying whether to skip bi-tangential component packing is 0, as an embodiment.
- the normal component of the displacement vector coefficient of each vertex is stored (or packed) in the Y channel
- the tangential component of the displacement vector coefficient of each vertex is stored (or packed) in the U channel
- the bi-tangential component of the displacement vector coefficient of each vertex is stored in the V channel. That is, the normal, tangential, and bi-tangential components can be stored (or packed) in order in the Y, U, and V channels, respectively.
- 21(b) is an example in which, when the displacement vector coefficients of four vertices are taken as an example, the normal components (N 1 -N 4 ), tangential components (T 1 -T 4 ), and bi-tangential components (B 1 -B 4 ) of the displacement vector coefficients of four vertices are packed in the Y channel, the U channel, and the V channel, respectively.
- FIG. 22(a) and FIG. 22(b) are diagrams showing other examples of packing displacement vector coefficients based on the YUV 4:4:4 format according to embodiments. That is, FIG. 22(a) and FIG. 22(b) are examples of skipping packing of bi-tangential components when packing displacement vector coefficients based on the YUV 4:4:4 format.
- the normal component of the displacement vector coefficients is stored (or packed) in the Y channel, and the tangential component of the displacement vector coefficients is stored (or packed) in the U channel. That is, the normal component is sequentially packed in the Y channel, and the tangential component is sequentially packed in the U channel, and then transmitted to the displacement vector image/video encoding unit (12014).
- the bi-tangential component of the displacement vector coefficient of each vertex is skipped from packing. That is, the bi-tangential component is skipped and not packed in the V channel.
- the present disclosure can fill the V channel with an intermediate value (or a fixed value) for efficient encoding/decoding. That is, in order to maintain the YUV 4:4:4 format, the V channel is filled with the intermediate value instead of the skipped bi-tangential component.
- FIG. 22(b) is an example in which, when displacement vector coefficients of four vertices are taken as an example, normal components (N 1 -N 4 ) and tangential components (T 1 -T 4 ) of the displacement vector coefficients of four vertices are packed into the Y channel and the U channel, respectively, and the intermediate value (M) instead of the bi-tangential component is packed into the V channel.
- FIG. 23(a) and FIG. 23(b) are diagrams showing an example of packing displacement vector coefficients based on the YUV 4:2:0 format according to embodiments. That is, FIG. 23(a) and FIG. 23(b) are examples of not skipping packing of bi-tangential components when packing displacement vector coefficients based on the YUV 4:2:0 format.
- the normal components of the displacement vector coefficients are stored (or packed) as they are in the Y channel.
- the tangential components of the displacement vector coefficients of one vertex per four vertices are packed in the U channel
- the bi-tangential components of the displacement vector coefficients of one vertex per four vertices are packed in the V channel. That is, the normal component can be stored in the Y channel, the tangential component in the U channel, and the bi-tangential component in the V channel, respectively, in that order.
- the tangential component of the U channel and the bi-tangential component of the V channel can be sampled according to the encoder/decoder agreement.
- 23(b) is an example in which, when displacement vector coefficients of four vertices are taken as an example, the normal components (N 1 -N 4 ) of the displacement vector coefficients of four vertices are packed into the Y channel, the tangential component (T) of the displacement vector coefficient of one vertex per four vertices is packed into the U channel, and the bi-tangential component (B) of the displacement vector coefficient of one vertex per four vertices is packed into the V channel.
- the tangential component of the displacement vector coefficient of one vertex packed into the U channel may be the average value of the tangential components of the displacement vector coefficients of four vertices, or may be the tangential component (i.e., representative value) of a specific displacement vector coefficient among the displacement vector coefficients of the four vertices.
- the bi-tangential component of the displacement vector coefficient of one vertex packed into the V channel may be the average of the bi-tangential components of the displacement vector coefficients of the four vertices, or may be the bi-tangential component (i.e., representative value) of a specific displacement vector coefficient among the displacement vector coefficients of the four vertices.
- the present disclosure increases the height of the Y channel by twice the packing frame image height (H) calculated by the displacement vector coefficient packing unit (12013), and stores (i.e. packs) the normal component, the tangential component, and the bi-tangential component in the Y channel component by component.
- the U channel and the V channel can each be filled with an intermediate value.
- FIG. 24(a) and FIG. 24(b) are diagrams showing another example of packing displacement vector coefficients based on the YUV 4:2:0 format according to embodiments. That is, FIG. 24(a) and FIG. 24(b) are examples of skipping packing of bi-tangential components when packing displacement vector coefficients based on the YUV 4:2:0 format.
- the height of the Y channel is increased by twice the packing frame image height (H) calculated by the displacement vector coefficient packing unit (12013), and then the displacement vector coefficients of the normal component and the displacement vector coefficients of the tangential component can be stored (i.e., packed) component-by-component in the Y channel.
- the U and V channels can be filled with intermediate values for efficient encoding.
- 24(b) is an example in which, when the displacement vector coefficients of four vertices are taken as an example, the normal components (N 1 -N 4 ) and the tangential components (T 1 -T 4 ) of the displacement vector coefficients of four vertices are packed in the Y channel, and the intermediate values (M) corresponding to two vertices are packed in the U and V channels respectively in order to maintain the YUV 4:2:0 format.
- the normal components (N 1 -N 4 ) of the displacement vector coefficients of the four vertices can be packed into a Y channel
- the tangential components (T) of the displacement vector coefficients of one vertex per four vertices can be packed into a U channel
- the median value can be filled into the V channel.
- FIG. 25(a) and FIG. 25(b) are diagrams showing an example of packing displacement vector coefficients based on the YUV 4:0:0 format according to embodiments. That is, FIG. 25(a) and FIG. 25(b) are examples of not skipping packing of bi-tangential components when packing displacement vector coefficients based on the YUV 4:0:0 format.
- the height of the Y channel is increased by three times the packing frame image height (H) calculated by the displacement vector coefficient packing unit (12013), and thereafter, the normal component, tangential component, and bi-tangential component of the displacement vector coefficients can be stored (i.e., packed) component by component in the Y channel.
- Fig. 25(b) is an example in which, when the displacement vector coefficients of four vertices are taken as an example, the normal components (N 1 -N 4 ), tangential components (T 1 -T 4 ), and bi-tangential components (B 1 -B 4 ) of the displacement vector coefficients of four vertices are packed component by component in the Y channel.
- FIG. 26(a) and FIG. 26(b) are diagrams showing other examples of packing displacement vector coefficients based on the YUV 4:0:0 format according to embodiments. That is, FIG. 26(a) and FIG. 26(b) are examples of skipping packing of bi-tangential components when packing displacement vector coefficients based on the YUV 4:0:0 format.
- the height of the Y channel is increased by twice the packing frame image height (H) calculated by the displacement vector coefficient packing unit (12013), and thereafter, the normal components and tangential components of the displacement vector coefficients can be stored (i.e., packed) component by component in the Y channel.
- Fig. 26(b) is an example in which, when the displacement vector coefficients of four vertices are taken as an example, the normal components (N 1 -N 4 ) and tangential components (T 1 -T 4 ) of the displacement vector coefficients of four vertices are packed component by component in the Y channel. That is, the bi-tangential component is skipped and therefore not packed or transmitted.
- the displacement vector image/video encoding unit (12014) can perform encoding on the packed 2D image through a 2D video encoder such as H.264, HEVC, or VVC.
- Fig. 27 illustrates a receiving device according to embodiments.
- the receiving device of Fig. 27 may be referred to as a mesh data receiving device or a decoder or a decoder of a receiving device or a V-Mesh decoder or a dynamic mesh decoder.
- FIG. 27 corresponds to the receiving device (110) or mesh video decoder (113) of FIG. 1, the decoder of FIG. 11 or FIG. 12, the receiving device of FIG. 14, and/or the receiving decoding device corresponding thereto.
- Each component of FIG. 27 corresponds to hardware, software, a processor, and/or a combination thereof.
- the receiving (decoding) operation of FIG. 27 can follow the reverse process of the corresponding process of the transmitting (encoding) operation of FIG. 15.
- the execution order of each block in FIG. 27 can be changed, some blocks can be omitted, and some blocks can be newly added.
- FIG. 27 may largely include a base mesh decoding unit, a displacement information decoding unit, and a texture map decoding unit.
- the base mesh decoding unit may include a switching unit (15011), a motion vector decoder (15012), a static mesh decoder (15013), a base mesh restoration unit (15014), a mesh subdivision unit (15015), and a mesh restoration unit (15016).
- the displacement information decoding unit may include a displacement vector decoder (15017) and a displacement vector coordinate system inverse transformation unit (15020).
- a bitstream of mesh data received by a receiver may be demultiplexed into a base mesh bitstream, a displacement vector bitstream, and a texture map bitstream in a demultiplexer (not shown) after file/segment decapsulation.
- the base mesh bitstream may be a motion vector bitstream.
- the base mesh bitstream is provided to a motion vector decoder (15012) via a switching unit (15011) or to a static mesh decoder (15013).
- the motion vector decoder (15012) may be referred to as a motion decoder.
- the motion vector decoder (15012) can perform decoding on a motion vector bitstream on a vertex-by-vertex basis or a subgroup basis.
- the motion vector decoder (15012) can reconstruct a final motion vector by adding a differential motion vector (i.e., a residual motion vector) decoded from a bitstream using a previously decoded motion vector as a predictor. That is, the motion vector decoder (15012) decodes a differential motion vector (or a residual motion vector) in units of a vertex or a subgroup (or a subblock) through a motion vector bitstream, and performs prediction based on connection information using a previously decoded motion vector as a predictor to decode the motion vector by adding it to the residual motion vector.
- a differential motion vector i.e., a residual motion vector
- the static mesh decoder (15013) can decode the base mesh bitstream to restore connection information, vertex geometry information, texture coordinates (i.e., attribute geometry information), normal information, etc. of the base mesh.
- the base mesh restoration unit (15014) can restore the current base mesh based on the decoded motion vector or the decoded base mesh. For example, if the current mesh is subject to inter-screen encoding, the base mesh restoration unit (15014) can add the decoded (or restored) motion vector to the reference base mesh and then perform inverse quantization to generate a restored base mesh (i.e., the current base mesh). As another example, if the current mesh is subject to intra-screen encoding, the base mesh restoration unit (15014) can perform inverse quantization on the decoded (or restored) base mesh through the static mesh decoder (15012) to generate a restored base mesh (i.e., the current base mesh).
- the mesh subdivision unit (15015) can perform subdivision on the base mesh to generate additional vertices.
- the present disclosure can implicitly derive and generate geometry information connection information, texture coordinate connection information, and texture coordinates according to the subdivision method.
- the mesh subdivision unit (15015) can perform subdivision through methods such as mid-edge, Loop, Catmul&Clark, etc.
- the displacement vector decoder may perform video codec-based decoding on the demultiplexed displacement vector bitstream as a video bitstream, or perform zero run-length decoding, or perform arithmetic decoding.
- the displacement vector decoder may be used interchangeably with the displacement vector transform decoder.
- the displacement vector decoder (15017) can restore the displacement vector by decoding the displacement vector in a reverse process of the displacement vector encoding method of the transmitter side.
- the displacement vector coordinate system inversion unit (15020) can perform a process of inversion into a Cartesian (or canonical) coordinate system (x, y, z) if the displacement vector decoded by the displacement vector decoder (15017) is a value of a local coordinate system (n, t, bt).
- the output of the displacement vector coordinate system inversion unit (15020) is provided to the mesh restoration unit (15016).
- the vertex displacement vector calculated in the (x, y, z) space can be converted to the (normal, tangential, bi-tangential) coordinate system (or local coordinate system) based on the normal vector of each vertex.
- the normal vector can be calculated for each subdivided vertex based on the geometry information and connection information of the surrounding vertices.
- the mesh restoration unit (15016) restores the mesh based on the mesh refined in the mesh refinement unit (15015) and the restored displacement vector output from the displacement vector coordinate system inverse transformation unit (15020).
- the received and demultiplexed texture map bitstream is input to a texture map decoder (15021).
- the texture map decoder (15021) can decode the texture map through a 2D scalable decoder. That is, the texture map decoder (15021) can restore the texture map by applying 2D scalable decoding to the texture map.
- the decoder of the receiving device goes through the process of decoding each bitstream to restore the mesh.
- the base mesh is decoded in the motion vector or static mesh decoder depending on whether it is inter or intra frame, and the geometry information is restored together with the decoded displacement vector information through subdivision.
- a decoding method for restoring displacement vectors of normal and tangential components and calculating bi-tangential components from a 2D image frame packed and transmitted in various formats e.g., YUV 4:4:4, YUV 4:2:0, YUV 4:0:0
- a displacement vector coefficient decoding unit e.g., YUV 4:4:4, YUV 4:2:0, YUV 4:0:0
- FIG. 28 is a block diagram showing an example of a displacement vector decoder according to embodiments.
- the displacement vector decoder (15017) may include a video decoding unit (16011), a displacement vector coefficient depacking unit (16012), a displacement vector coefficient dequantization unit (16013), and a displacement vector inverse transform unit (16014).
- Each component in FIG. 28 corresponds to hardware, software, a processor, and/or a combination thereof.
- the execution order of each block in FIG. 28 may be changed, some blocks may be omitted, and some blocks may be newly added.
- the video decoding unit (16011) receives a displacement vector bitstream as input and performs decoding on a displacement vector coefficient image/video through a 2D video codec. Then, the restored displacement vector coefficient video restored through the video decoding unit (16011) can perform displacement vector coefficient assignment corresponding to each vertex of the restored mesh by performing displacement vector coefficient unpacking (or unpacking) for each frame in the displacement vector coefficient unpacking unit (16012).
- the displacement vector coefficient depacking unit (16012) can perform depacking according to a scanning order defined by an encoder/decoder agreement from a restored displacement vector coefficient image corresponding to a current mesh frame or a scanning order parsed into a higher-level unit (sequence, frame, etc.). That is, the displacement vector coefficient depacking unit (16012) can perform depacking according to a scanning order defined by an encoder/decoder agreement, or can perform depacking by deriving a scanning order according to the characteristics of displacement vector coefficients, or can receive a scanning order from an encoder of a transmitting device.
- a displacement vector coefficient block packed in a specific scanning order in units of bx*by, and a displacement vector coefficient packed in a specific order (such as 2D Morton Code or Zig-zag scan) within a block can derive the displacement vector coefficient of the kth vertex according to the encoder/decoder promise or the sizes of the parsed block bx, by and L, M and the scanning order.
- the following describes a method of restoring displacement vector coefficients of normal and tangential components from a 2D image frame by performing inverse packing in a displacement vector coefficient inverse packing unit (16012) when displacement vector coefficients are packed and transmitted in a 2D image frame in various formats (e.g., YUV 4:4:4, YUV 4:2:0, YUV 4:0:0) as shown in FIGS. 21 to 26, and restoring bi-tangential components based on the same.
- the value of the information (Bi_tangent_ski
- the displacement vector coefficient depacking unit (16012) can perform restoration on the normal, tangential, and bi-tangential components of the Y, U, and V channels, respectively, in the YUV 4:4:4 format, as shown in Fig. 29(b). That is, the restored displacement vector coefficients of the Y, U, and V channels can be restored as normal, tangential, and bi-tangential components, respectively, according to a specific order (displacement_scan_method) within the block.
- the transmitting device since the transmitting device transmits all values of the normal, tangential, and bi-tangential components, the data before encoding and the decoded result can be the same.
- this is a case where packing of bi-tangential components is skipped in a transmitting device (e.g., displacement vector coefficient packing unit (12013)).
- a transmitting device e.g., displacement vector coefficient packing unit (12013)
- the displacement vector coefficients of four vertices are taken as an example, the normal components (N 1 -N 4 ) and tangential components (T 1 -T 4 ) of the displacement vector coefficients of the four vertices are packed into the Y channel and the U channel, respectively, and the intermediate value (M) is packed into the V channel instead of the bi-tangential component (B 1 -B 4 ).
- the Y and U channels may be composed of the displacement vector coefficients of the normal and tangential components, respectively, as in Fig. 30(a), and the V channel may be composed of the intermediate value (M). If it is assumed that the bitDepth of the displacement vector video image is 10 bits, the intermediate value of 1024, 512, can be the M value of the V channel.
- the displacement vector coefficient depacking unit (16012) can perform restoration on the normal, tangential, and bi-tangential components of the Y, U, and V channels, respectively, in the YUV 4:4:4 format, as shown in FIG. 30(b). That is, the displacement vector coefficients of the Y and U channels are restored as normal and tangential components, respectively, according to a specific order (displacement_scan_method) within the block, and the displacement vector coefficients having the intermediate value of the V channel can be restored as bi-tangential components with a value of 0. Alternatively, the bi-tangential component values may not be restored and may have 0 as the initial value. In addition, since the transmitting device transmits all values of the normal and tangential components, the data before encoding and the decoded results of the normal and tangential components may be identical.
- image packing format information ColdSpace_displacement_video
- this is a case where packing of bi-tangential components is not skipped in a transmitting device (e.g., displacement vector coefficient packing unit (12013)).
- a transmitting device e.g., displacement vector coefficient packing unit (12013)
- the normal components (N 1 -N 4 ) of the displacement vector coefficients of four vertices are packed into the Y channel
- the tangential component (T) of the displacement vector coefficient of one vertex per four vertices is packed into the U channel
- the bi-tangential component (B) of the displacement vector coefficient of one vertex per four vertices is packed into the V channel.
- the displacement vector coefficient depacking unit (16012) can perform restoration of the normal, tangential, and bi-tangential components of the Y, U, and V channels, respectively, in the YUV 4:2:0 format, as shown in FIG. 31(b).
- the tangential and bi-tangential components of the U and V channels are sampled and transmitted by the transmitter (e.g., 1 vertex per 4 vertices is sampled)
- the tangential component (T' 1 -T' 4 ) of the U channel and the bi-tangential component (B' 1 -B' 4 ) of the V channel can be restored through a sampling method agreed upon in advance with the encoder.
- all of the normal, tangential, and bi-tangential components can be restored from the Y channel by a method agreed upon in advance with the encoder.
- the data before encoding of the normal component and the decoded result can be identical.
- the displacement vector coefficients of the restored Y, U, and V channels can be restored as normal, tangential, and bi-tangential components, respectively, according to a specific order (displacement_scan_method) within the block.
- the transmitting device e.g., the displacement vector coefficient packing unit (12013)
- the transmitting device increases the height of the Y channel by twice the packing frame image height (H) and then performs packing.
- the displacement vector coefficient inverse packing unit (16012) can restore the normal component from 0 to H of the Y channel and the tangential component from H+1 to 2H using the height (H) information of the previously calculated packing frame, as shown in Fig. 32(b). Then, the displacement vector coefficients of the normal and tangential components of the Y channel can be restored according to a specific order (displacement_scan_method) within the block, respectively. At this time, the intermediate value of the U channel may not be used in the restoration process. This is because the displacement vector coefficients of the tangential component are restored from the Y channel. In addition, the displacement vector coefficients having the intermediate value of the V channel can restore the value of 0 as the bi-tangential component.
- the bi-tangential component value may not be restored from the V channel and may have 0 as the initial value.
- the values of the bi-tangential components may all become 0 after inverse packing.
- the bi-tangential component whose packing is skipped in the displacement vector coefficient packing unit (12013) of the transmitter can be restored to 0 in the displacement vector coefficient depacking unit (16012) of the receiver.
- the data before encoding and the decoded result of the normal component and the tangential component can be the same.
- image packing format information ColdSpace_displacement_video
- a transmitting device e.g., displacement vector coefficient packing unit (12013)
- the displacement vector coefficients of four vertices are taken as an example, only the normal components (N 1 -N 4 ) of the displacement vector coefficients of four vertices are packed component-by-component into the Y channel. At this time, the tangential components and bi-tangential components are not packed and transmitted. Therefore, the inverse packing of the tangential and bi-tangential components is not performed.
- the displacement vector coefficient inverse packing unit (16012) can perform restoration only on the normal component of the Y channel in the YUV 4:0:0 format as shown in Fig. 33(b). At this time, the restored displacement vector coefficients of the Y channel can be restored as normal components according to a specific order (displacement_scan_method) within the block. In addition, since all values of the normal components are transmitted from the transmitting device, the data before encoding of the normal components and the decoded results can be identical.
- the value of the information (Bi_tangent_
- the transmitting device e.g., the displacement vector coefficient packing unit (12013)
- the normal components (N 1 -N 4 ), tangential components (T 1 -T 4 ), and bi-tangential components (B 1 -B 4 ) of the displacement vector coefficients of four vertices are all packed component-by-component into the Y channel.
- the transmitting device increases the height of the Y channel by three times the packing frame image height (H) and then performs packing.
- the displacement vector coefficient inverse packing unit (16012) can restore normal components from 0 to H of the Y channel, tangential components from H+1 to 2H, and bi-tangential components from 2H+1 to 3H using the height (H) information of the previously calculated packing frame, as shown in FIG. 34(b).
- the displacement vector coefficients of the normal, tangential, and bi-tangential components of the Y channel can be restored according to a specific order (displacement_scan_method) within the block. That is, the normal, tangential, and bi-tangential components can be transmitted to the Y channel by agreement with the encoder, and the displacement vector coefficient inverse packing unit (16012) can restore normal, tangential, and bi-tangential components from the Y channel, respectively.
- a transmitting device e.g., displacement vector coefficient packing unit (12013)
- the transmitting device increases the height of the Y channel by twice the packing frame image height (H) and then performs packing.
- the displacement vector coefficient inverse packing unit (16012) can restore the normal component from 0 to H of the Y channel and the tangential component from H+1 to 2H using the height (H) information of the previously calculated packing frame as shown in Fig. 35(b).
- the displacement vector coefficients of the Y channel can be restored according to a specific order (displacement_scan_method) within the block. That is, the normal and tangential components can be transmitted to the Y channel by agreement with the encoder, and the displacement vector coefficient inverse packing unit (16012) can restore the normal component and the tangential component from the Y channel, respectively.
- the data before encoding of the normal component and the tangential component and the decoded result can be the same.
- Displacement vector coefficients depacked by at least one of the depacking methods of FIGS. 29 to 35 are provided to a displacement vector coefficient dequantization unit (16013).
- the displacement vector coefficient inverse quantization unit (16013) can inverse quantize the displacement vector coefficients restored by the displacement vector coefficient inverse packing unit (16012).
- the displacement vector inverse transform unit (16014) performs an inverse transform of the transform performed in the encoder of the transmitter on the inverse quantized displacement vector coefficients to output displacement vectors.
- a lifting inverse transform, a wavelet inverse transform, etc. may be performed. If a lifting inverse transform is performed in the displacement vector inverse transform unit (16014), a process of updating a displacement vector of a vertex used for prediction in the encoder may be performed through a parsed residual signal.
- the displacement vector inverse transform unit (16014) performs an inverse transform of the displacement vector coefficients or inverse quantized displacement vector coefficients allocated per vertex through the displacement vector coefficient inverse packing unit (16012).
- the transform may be applied as a linear lifting transform, a butterfly lifting transform, a wavelet transform, etc. For example, if the value of asps_vmc_ext_transform_method signaled in the signaling information (ASPS) is 1, it may indicate that Linear_Lifting was used as the transform method.
- an average or distance-based weighted average prediction can be performed on n nearby points based on connection information among vertices with a lower level of detail than the current vertex.
- prediction can be performed based on the displacement vector of n vertices used to generate the current vertex in the mesh refinement step.
- a process of updating the displacement vector of the vertex used for prediction in the encoder can be performed through the parsed residual signal.
- the displacement vector coefficient inverse quantization unit (16013) performs inverse quantization on the displacement vector coefficients allocated to each vertex or the displacement vector coefficients on which inverse transformation has been performed through the displacement vector coefficient inverse packing unit (16012).
- the quantization parameter (QP) for each component may be transmitted in units of sequences or frames to determine a quantization rate.
- the displacement vector coefficient inverse quantization unit (16013) can perform inverse quantization only on the normal and tangential components and skip inverse quantization on the bi-tangential component.
- displacement vector coefficients can be quantized through different quantization parameters for each axis, and the quantization rate can be determined for each LoD level by deriving quantization parameters or scaling parameters by encoder/decoder agreement.
- the inversely transformed displacement vector or the inversely quantized displacement vector coefficients are provided to the displacement vector coordinate system inverse transformation unit (15020).
- the coordinate system transformation status flag (asps_vmc_ext_displacement_coordinate_system) included in the signaling information in units of frame sequence or GOF (Group of Frames) or frame or submesh is parsed, and if its value is 1, the inverse quantized (or inversely transformed) restored displacement vector can be inversely transformed from the local coordinate system (n, t, b) to the canonical coordinate system (x, y, z).
- a normal vector per vertex is calculated based on the restored vertex position information of the restored base mesh, and normal values of newly created vertices can be assigned by interpolating the normal vector of the restored base mesh for the vertices additionally created through the subdivision process.
- interpolation can be performed by averaging or weighting the normal information of the base mesh used for subdivision.
- the normal information of the base mesh can be used as is for subdivided vertices on the same plane.
- the tangential and bi-tangential vectors orthogonal to the normal vector can be calculated and the displacement vector coordinate system inverse transformation can be performed.
- disp n [0] and disp n [1] represent the results of the normal and tangential components obtained by performing the inverse transformation and inverse quantization.
- the displacement vector of the bi-tangential component can be calculated through the outer product of the final displacement vectors of the normal component and the tangential component.
- the displacement vector of the bi-tangential component can be derived and calculated through a method such as Linear Regression or Multiple Regression.
- the coordinate system inverse transformation can be performed by multiplying the result of performing inverse quantization and inverse transformation of the n component and the calculated normal vector per vertex.
- coordinate system inverse transformation can always be performed without sending a flag.
- the mesh restoration unit (15016) can calculate and restore vertex geometry information of the restoration mesh by adding a restoration displacement vector to vertices generated through the subdivision process in the mesh subdivision unit (15015).
- signaling information may be generated in a metadata processing unit (not shown, may be referred to as a metadata generator, etc.) and provided to corresponding blocks in the transmitting device and/or a receiving device (or a decoder of the receiving device), and a metadata parser (not shown) of the receiving device may parse the received signaling information and provide it to the corresponding blocks.
- a metadata processing unit not shown, may be referred to as a metadata generator, etc.
- a metadata parser not shown
- each block of the receiving device may perform each operation based on the signaling information.
- FIG. 36 is a diagram showing an example of a structure of an atlas sequence parameter set (ASPS) among signaling information in a bitstream according to embodiments.
- FIG. 36 is a diagram showing an example of an atlas sequence parameter set extension RBSP syntax and semantics structure. That is, ASPS can be extended from atlas sequence parameters to further include parameters related to displacement vector encoding.
- the displacement vector coordinate system transformation status flag (asps_vmc_ext_displacement_coordinate_system) is information that can identify whether the displacement vector coordinate system is transformed, and indicates the type of coordinate system for displacement vector encoding. For example, if the value of the displacement vector coordinate system transformation status flag (asps_vmc_ext_displacement_coordinate_system) syntax (or field) is 0, it can indicate that the transmitter uses the canonical coordinate system as it is, and if it is 1, it can indicate that the transformation to the local coordinate system has been performed.
- the displacement vector coordinate system inverse transformation unit 15020
- the value of the coordinate system transformation status flag (asps_vmc_ext_displacement_coordinate_system) is 1
- the inverse quantized (or inversely transformed) restored displacement vector can be transformed from the local coordinate system (n, t, b) to the canonical coordinate system (x, y, z).
- the displacement vector coefficient packing method indicates the packing method of the displacement vector coefficients. For example, if this value is 0, it indicates that the displacement vector coefficients are packed in ascending order, and if it is 1, it indicates that the displacement vector coefficients are packed in descending order.
- Displacement vector encoding method indicates the encoding method of the displacement vector. For example, if this value is 0, it indicates None (no displacement vector encoding), 1 indicates that the displacement vector is encoded using the Arithmetic Coding method, and 2 indicates that the displacement vector is encoded using the Video Coding method.
- the displacement vector packing method for each LoD (asps_vmc_ext_displacement_LoD_packing_method) and displacement vector-related packing information (asps_vmc_ext_displacement_packing_info()) may be included.
- LoD-specific displacement vector packing method indicates the LoD-specific packing method of displacement vectors. For example, if this value is 0, it indicates that the displacement vectors are packed based on the global level packing frame configuration, and if it is 1, it indicates that the displacement vectors are packed based on the level-specific packing frame configuration.
- Displacement vector related packing information may include, as shown in FIG. 37, a displacement vector scan method (displacement_scan_method), the number of unit blocks of displacement vector video images (displacementVideoBlockSize), unit block size information (geometryVideoBlockSize), bit depth information (geometryVideoBitDepth), image packing format information (ColourSpace_displacement_video), and information for identifying whether to skip displacement vector bi-tangential component image packing (bi_tangent_skip_flag).
- FIG. 37 is a diagram showing another example of the structure of an atlas sequence parameter set (ASPS) among signaling information in a bitstream according to embodiments. That is, FIG. 37 is an example of the syntax structure of displacement vector related packing information (asps_vmc_ext_displacement_packing_info()) included in ASPS, which may be included in asps_vmc_extension( ) of FIG. 36.
- ASPS atlas sequence parameter set
- the displacement vector scan method indicates the scan order method when packing displacement vectors. For example, if the value of displacement_scan_method is 0, it indicates that the displacement vector is scanned with 2D Morton Code, and if it is 1, it indicates that the displacement vector is scanned in Zig-zag scan order.
- geometryVideoBlockSize represents the displacement vector video image block size.
- this value can represent the number of bx*by blocks, and the default value can be 16.
- geometryVideoBitDepth represents the displacement vector video image bit depth unit.
- the default value can be 10 bits. This value can determine the intermediate value when packing displacement vector coefficients.
- Image packing format information indicates the displacement vector video packing image format. For example, if this value is 0, it indicates that displacement vector coefficients are packed based on None (there is no image format of displacement vector video packing), 1 indicates that the displacement vector coefficients are packed based on the yuv400 format, 2 indicates that the displacement vector coefficients are packed based on the yuv420 format, and 3 indicates that the displacement vector coefficients are packed based on the yuv444 format.
- the bi-tangent component packing skip flag (Bi_tangent_skip_flag) is information for identifying whether the displacement vector bi-tangential component image packing is skipped. If this value is 0, it indicates no skipping, and if it is 1, it indicates skipping.
- Fig. 38 is a flowchart showing an example of a transmission method according to embodiments.
- the transmission method according to embodiments may include a step (21011) of encoding mesh data and a step (21012) of transmitting a bitstream including the encoded mesh data.
- the bitstream transmitted in step (21012) includes a base mesh bitstream, a displacement vector bitstream, and a texture map bitstream.
- the step of encoding mesh data (21011) may include a process of encoding a base mesh, a process of encoding displacement vectors or displacement vector transform coefficients, and a process of encoding a texture map.
- the original mesh to be transmitted is first simplified and mesh parameterized to generate a base mesh.
- the generated base mesh is quantized, and in the case of an inter-frame, a motion vector is calculated from the previously referenced restored base mesh to encode the motion vector, and in the case of an intra-frame, it is encoded through static mesh encoding and transmitted as a base mesh bitstream.
- the displacement vector between the mesh data that has been simplified through mesh simplification, refined and fitted, and the mesh data restored from the previously encoded base mesh is calculated.
- the displacement vector coordinate system is converted to a local coordinate system, and the displacement vector in the local coordinate system is converted and quantized into displacement vector coefficients, and then encoded into a displacement vector bitstream and transmitted.
- the displacement vectors converted to the local coordinate system can be converted into displacement vector coefficients by performing linear lifting transformation, butterfly lifting transformation, etc. in the displacement vector conversion unit (12011).
- the displacement vectors of (n, t, b) expressed in the local coordinate system can be converted for each normal, tangential, and bi-tangential component.
- the displacement vector coefficients converted in the displacement vector conversion unit (12011) can be quantized in the displacement vector quantization unit (12012), and at this time, each channel can be quantized into an individual value.
- bi_tangent_skip_flag is signaled and transmitted in signaling information (e.g., atlas sequence parameter set) as in FIG. 37.
- the quantized displacement vector coefficients go through a process of packing into a 2D image of the size of W ⁇ H in the displacement vector coefficient packing unit (12013). For example, if the displacement vector coefficients are configured as a 1D LoD ascending order as shown in FIG. 17, image packing is performed in a 2D form as shown in FIG. 18. At this time, the size of Bx*by can be configured as one block, and the displacement vector video image can be configured with L*M blocks determined according to the number N of displacement vector coefficients.
- the number of unit blocks (displacementVideoBlockSize) of the displacement vector video image, the size information of the unit block (geometryVideoBlockSize), and the bit depth information (geometryVideoBitDepth) can be signaled to signaling information (e.g., atlas sequence parameter set) as shown in FIG. 37.
- signaling information e.g., atlas sequence parameter set
- Displacement vector coefficients can be packed inside a block in a zig-zag scan order or a 2D Morton code order
- displacement vector packing order information (displacement_scan_method) can be signaled in the signaling information (e.g., atlas sequence parameter set) as in Fig. 37.
- padding can be performed with the median value or the last displacement vector coefficient value of the image to match the size of the basic block or the entire 2D video image for each LoD.
- Each displacement vector coefficient can be composed of a normal component, a tangential component, and a bi-tangential component, and some components can be skipped during the packing process.
- a format such as YUV 4:4:4, YUV 4:2:0, or YUV 4:0:0 can be selected and image packing format information (ColourSpace_displacement_video) can be signaled.
- the present disclosure refers to YUV 4:4:4 as a first format, YUV 4:2:0 as a second format, and YUV 4:0:0 as a third format.
- packing is performed as is for the normal component value as in Fig. 23(a) and Fig. 23(b) as Y channel, and the values of the tangential and bi-tangential components can be sampled and packed according to the promise of the encoder/decoder.
- the tangential and bi-tangential components can also be packed together in the Y channel, and packing can be performed with intermediate values in the U channel and the V channel.
- displacement vector coefficients can be imaged for each channel to form an encoding unit in the form of a packing frame.
- the entire level within the mesh frame can be packed into one packing frame, or each LoD within the mesh frame can be packed into a separate packing frame.
- the asps_vmc_ext_displacement_LoD_packing_method syntax indicating the packing method for each mesh sequence unit can be signaled to signaling information (e.g., ASPS) as in Fig. 36.
- the 2D image packed through the displacement vector coefficient packing unit (12013) is encoded by a 2D video encoder such as H.264, HEVC, or VVC in the displacement vector image/video encoding unit (12014) to generate a displacement vector bitstream.
- a 2D video encoder such as H.264, HEVC, or VVC in the displacement vector image/video encoding unit (12014) to generate a displacement vector bitstream.
- a new texture map having color information corresponding to the texture coordinates of the restored mesh is generated through a texture map generation unit (11026), and the generated texture map is encoded through a texture map encoder (i.e., a 2D video encoder) (11027) and transmitted as a texture bitstream.
- a texture map encoder i.e., a 2D video encoder
- the base mesh bitstream, displacement vector bitstream, and texture bitstream generated as described above in the step (21011) of encoding the mesh data are generated into a single bitstream through a multiplexing unit, and transmitted to a receiving device through a transmitting unit.
- FIG. 39 is a flowchart showing an example of a receiving method according to embodiments.
- the receiving method according to embodiments may include a step (22011) of receiving a bitstream including mesh data and a step (22012) of decoding mesh data included in the bitstream.
- the step (22011) of receiving a bitstream including mesh data receives a bitstream including a base mesh bitstream, a displacement vector bitstream, and a texture map bitstream, as an example.
- the step (22011) of receiving a bitstream including mesh data also receives signaling information including an atlas sequence parameter set (ASPS). At this time, the signaling information may be received while being included in the bitstream and may also be referred to as metadata.
- ASS atlas sequence parameter set
- the step of decoding mesh data (22012) may include a process of decoding a base mesh bitstream, a process of decoding a displacement vector bitstream, and a process of decoding a texture map bitstream.
- the received bitstream is demultiplexed into a base mesh bitstream, a displacement vector bitstream, and a texture map bitstream through a demultiplexing unit, and then a process of decoding each of the bitstreams is performed.
- the base mesh bitstream is decoded through a motion vector encoder (15012) for inter-frames and through a static mesh decoder (15013) for intra-frames.
- the decoded base mesh is then subjected to mesh refinement through a base mesh restoration unit (15014).
- the displacement vector bitstream decodes displacement vector coefficients in the reverse order of encoding, performs inverse quantization and inverse transformation, and then is inversely transformed to a coordinate system, and then mesh geometry information is restored together with the base mesh data.
- the video decoding unit (16011) receives a displacement vector bitstream as input and performs decoding on a displacement vector coefficient image/video through a 2D video codec. Then, the displacement vector coefficient inverse packing unit (16012) can perform inverse packing from a restored displacement vector coefficient image corresponding to the current mesh frame.
- the scan order information (displacement_scan_method) at the time of displacement vector packing, the number of unit blocks of the displacement vector video image (displacementVideoBlockSize), the size information of the unit block (geometryVideoBlockSize), and the bit depth information (geometryVideoBitDepth) are parsed from the signaling information (e.g., ASPS) of FIGS.
- the packing image format information ColdSpace_displacement_video
- whether to skip the bi-tangential component packing Bi_tangent_skip_flag
- ASPS signaling information
- the normal, tangential, and bi-tangential components can be restored from the Y, U, and V channels in sequence, as shown in Fig. 31(a) and Fig. 31(b), using the YUV 4:2:0 format.
- the tangential and bi-tangential components of the U and V channels may be the results sampled by the encoder/decoder agreement, and the restored values or positions may also be restored in a fixed state.
- the normal, tangential, and bi-tangential components may all be packed in the Y channel, and may be restored as each component from the Y channel, and the U and V channels may be composed of intermediate values.
- the normal component can be restored from the Y channel using the YUV 4:0:0 format as shown in Fig. 33(a) and Fig. 33(b).
- the normal, tangential, and bi-tangential components may all be packed in the Y channel, in which case the normal, tangential, and bi-tangential components can be restored in sequence from the Y channel as shown in Fig. 34(a) and Fig. 34(b).
- the displacement vector coefficients on which the above inverse quantization is performed are inversely transformed in the displacement vector inverse transform unit (16014) based on the transformation method parsed in the asps_vmc_ext_transform_method signaled in the signaling information, and the restored displacement vector is calculated.
- the inverse transform process can be performed on the inverse quantization results of the normal and tangential components excluding the bi-tangential component. Since the normal, tangential, and bi-tangential components are all orthogonal to each other, the displacement vector of the bi-tangential component can be calculated by taking the outer product of the final displacement vectors of the normal component and the tangential component.
- the displacement vector of the bi-tangential component can be derived using a method such as Linear Regression or Multiple Regression.
- the displacement vector coordinate system inverse transformation unit (15020) parses the asps_vmc_ext_displacement_coordinate_system signaled in the signaling information as shown in FIG. 36, and if its value is 1, the restored displacement vector is inversely transformed from the local coordinate system (n, t, b) to the canonical coordinate system (x, y, z).
- the displacement vector coordinate system inverse transformation can be performed by calculating the normal vector per vertex based on the restored vertex position information of the restored base mesh, and calculating the tangential and bi-tangential vectors orthogonal to the normal vector through the calculated normal vector per vertex.
- the mesh restoration unit (15016) can calculate the vertex geometry information of the restoration mesh by adding the restoration displacement vector to the vertices generated through the mesh subdivision process, thereby restoring the final geometric information.
- the received texture map bitstream is decoded through a texture map decoder (15021).
- the decoded texture map is used to generate a final restored mesh together with the restored geometry information in the mesh wall element (15016).
- displacement vectors are transformed into a local coordinate system for compression efficiency, and then transformed into displacement vector coefficients in a simple form through lifting transformation and quantization before compression is performed.
- this process is calculated in the encoder for all components of normal, tangential, and bi-tangential generated by the coordinate system transformation, and the resulting displacement vector coefficients of each component are packed into a 2D image and transmitted.
- the present disclosure proposes a method of packing and transmitting displacement vector coefficients of two components, normal and tangential, among three components in order to transmit data quickly with more efficient capacity, and calculating bi-tangential components after decoding displacement vectors of normal and tangential components transmitted from a decoder, as described in FIGS. 15 to 39.
- the present disclosure encodes and transmits only two components of the three displacement vector components, so that a capacity reduction effect of about two-thirds compared to the conventional displacement vector sub-bitstream can be observed.
- Each of the parts, modules or units described above may be software, processor or hardware parts that execute sequential execution processes stored in a memory (or storage unit). Each of the steps described in the above-described embodiments may be performed by processor, software or hardware parts. Each of the modules/blocks/units described in the above-described embodiments may operate as a processor, software or hardware. In addition, the methods presented in the embodiments may be executed as code. The code may be written in a processor-readable storage medium and thus may be read by a processor provided by an apparatus.
- the devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.
- the various components of the device of the embodiments may be performed by hardware, software, firmware, or a combination thereof.
- the various components of the embodiments may be implemented by one chip, for example, one hardware circuit.
- the components according to the embodiments may be implemented by separate chips, respectively.
- At least one of the components of the device of the embodiments may be configured by one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations/methods according to the embodiments.
- the executable instructions for performing the methods/operations of the device of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors.
- the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may include implementations in the form of carrier waves, such as transmission via the Internet. Additionally, the processor-readable recording medium may be distributed across network-connected computer systems, so that the processor-readable code may be stored and executed in a distributed manner.
- Various elements of the embodiments may be performed by hardware, software, firmware, or a combination thereof. Various elements of the embodiments may be performed on a single chip, such as a hardware circuit. In some embodiments, the embodiments may optionally be performed on separate chips. In some embodiments, at least one of the elements of the embodiments may be performed within one or more processors that include instructions for performing operations according to the embodiments.
- the operations according to the embodiments described in this document may be performed by a transceiver device including one or more memories and/or one or more processors according to the embodiments.
- the one or more memories may store programs for processing/controlling the operations according to the embodiments, and the one or more processors may control various operations described in this document.
- the one or more processors may be referred to as a controller, etc.
- the operations according to the embodiments may be performed by firmware, software, and/or a combination thereof, and the firmware, software, and/or a combination thereof may be stored in a processor or a memory.
- first, second, etc. may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be limited in their interpretation by the above terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a second user input signal. Similarly, a second user input signal may be referred to as a first user input signal. The use of these terms should be construed as not departing from the scope of the various embodiments. Although the first user input signal and the second user input signal are both user input signals, they do not mean the same user input signals unless the context clearly indicates otherwise.
- the embodiments can be applied in whole or in part to 3D data transmission and reception devices and systems.
- Those skilled in the art can variously change or modify the embodiments within the scope of the embodiments.
- the embodiments can include changes/modifications, and the changes/modifications do not depart from the scope of the claims and their equivalents.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Description
Claims (16)
- 베이스 메쉬 비트스트림, 변위 벡터 비트스트림, 텍스처 맵 비트스트림, 및 시그널링 정보를 수신하는 단계;상기 베이스 메쉬 비트스트림으로부터 베이스 메쉬를 복원하는 베이스 메쉬 처리 단계;상기 변위 벡터 비트스트림으로부터 변위 정보를 복원하는 변위 정보 처리 단계;상기 베이스 메쉬와 상기 변위 정보를 기반으로 메쉬를 복원하는 복원 단계; 및상기 텍스처 맵 비트스트림으로부터 텍스처 맵을 복원하는 텍스처 맵 처리 단계를 포함하는 메쉬 데이터 디코딩 방법.
- 제 1 항에 있어서, 상기 변위 정보 처리 단계는상기 변위 벡터 비트스트림을 변위 정보로 디코딩하는 단계;제1 포맷, 제2 포맷, 제3 포맷 중 상기 변위 정보에 적용된 포맷을 식별하고, 상기 식별된 포맷을 기반으로 상기 변위 정보에 포함된 노말 성분, 탄젠셜 성분, 또는 바이-탄젠셜 성분 중 적어도 하나의 성분을 역패킹하는 단계; 및상기 역패킹된 변위 정보의 로컬 좌표계를 오리지날 좌표계로 역변환하는 단계를 포함하는 메쉬 데이터 디코딩 방법.
- 제 2 항에 있어서,상기 시그널링 정보는 상기 변위 정보에 적용된 포맷을 식별하기 위한 정보를 포함하는 메쉬 데이터 디코딩 방법.
- 제 3 항에 있어서,상기 시그널링 정보는 상기 바이-탄젠셜 성분의 패킹이 스킵되었는지 여부를 식별하기 위한 정보를 더 포함하고,상기 역패킹 단계는 상기 시그널링 정보를 기반으로 상기 노말 성분, 탄젠셜 성분, 또는 바이-탄젠셜 성분 중 적어도 하나의 성분의 역패킹 방법을 결정하는 메쉬 데이터 디코딩 방법.
- 베이스 메쉬 비트스트림, 변위 벡터 비트스트림, 텍스처 맵 비트스트림, 및 시그널링 정보를 수신하는 수신부;상기 베이스 메쉬 비트스트림으로부터 베이스 메쉬를 복원하는 베이스 메쉬 처리부;상기 변위 벡터 비트스트림으로부터 변위 정보를 복원하는 변위 정보 처리부;상기 베이스 메쉬와 상기 변위 정보를 기반으로 메쉬를 복원하는 복원부; 및상기 텍스처 맵 비트스트림으로부터 텍스처 맵을 복원하는 텍스처 맵 처리부를 포함하는 메쉬 데이터 디코딩 장치.
- 제 5 항에 있어서, 상기 변위 정보 처리부는상기 변위 벡터 비트스트림을 변위 정보로 디코딩하는 변위 정보 디코딩부;제1 포맷, 제2 포맷, 제3 포맷 중 상기 변위 정보에 적용된 포맷을 식별하고, 상기 식별된 포맷을 기반으로 상기 변위 정보에 포함된 노말 성분, 탄젠셜 성분, 또는 바이-탄젠셜 성분 중 적어도 하나의 성분을 역패킹하는 변위 정보 역패킹부; 및상기 역패킹된 변위 정보의 로컬 좌표계를 오리지날 좌표계로 역변환하는 좌표계 역변환부를 포함하는 메쉬 데이터 디코딩 장치.
- 원본 메쉬를 인코딩하는 단계; 및상기 인코딩된 메쉬와 시그널링 정보를 포함하는 비트스트림을 전송하는 단계를 포함하는 메쉬 데이터 인코딩 방법.
- 제 7 항에 있어서, 상기 인코딩 단계는상기 원본 메쉬를 단순화하여 생성된 베이스 메쉬를 인코딩하여 베이스 메쉬 비트스트림을 생성하는 베이스 메쉬 처리 단계;상기 베이스 메쉬를 기반으로 생성된 변위 정보를 인코딩하여 변위 벡터 비트스트림을 생성하는 변위 정보 처리 단계;상기 인코딩된 베이스 메쉬와 상기 인코딩된 변위 정보를 기반으로 메쉬를 복원하는 메쉬 복원 단계; 및상기 원본 메쉬와 상기 복원된 메쉬를 기반으로 생성된 텍스처 맵을 인코딩하여 텍스처 맵 비트스트림을 생성하는 텍스처 맵 처리 단계를 포함하는 메쉬 데이터 인코딩 방법.
- 제 8 항에 있어서, 상기 변위 정보 처리 단계는상기 변위 정보의 좌표계를 로컬 좌표계로 변환하는 단계;상기 로컬 좌표계의 변위 정보에 포함된 노말 성분, 탄젠셜 성분, 또는 바이-탄젠셜 성분 중 적어도 하나의 성분에 대해 제1 포맷, 제2 포맷, 또는 제3 포맷 중 하나의 포맷을 기반으로 패킹을 수행하는 단계; 및상기 패킹된 변위 정보를 인코딩하는 단계를 포함하는 메쉬 데이터 인코딩 방법.
- 제 8 항에 있어서, 상기 패킹 단계는상기 변위 정보의 바이-탄젠셜 성분에 대해 선택적으로 패킹을 수행하며,상기 시그널링 정보는 상기 바이-탄젠셜 성분의 패킹이 스킵되었는지 여부를 식별하기 위한 정보를 포함하는 메쉬 데이터 인코딩 방법.
- 제 8 항에 있어서,상기 시그널링 정보는 상기 변위 정보에 적용된 포맷을 식별하기 위한 정보를 포함하는 메쉬 데이터 인코딩 방법.
- 원본 메쉬를 인코딩하는 인코더; 및상기 인코딩된 메쉬와 시그널링 정보를 포함하는 비트스트림을 전송하는 전송부를 포함하는 메쉬 데이터 인코딩 장치.
- 제 12 항에 있어서, 상기 인코더는상기 원본 메쉬를 단순화하여 생성된 베이스 메쉬를 인코딩하여 베이스 메쉬 비트스트림을 생성하는 베이스 메쉬 처리부;상기 베이스 메쉬를 기반으로 생성된 변위 정보를 인코딩하여 변위 벡터 비트스트림을 생성하는 변위 정보 처리부;상기 인코딩된 베이스 메쉬와 상기 인코딩된 변위 정보를 기반으로 메쉬를 복원하는 메쉬 복원부; 및상기 원본 메쉬와 상기 복원된 메쉬를 기반으로 생성된 텍스처 맵을 인코딩하여 텍스처 맵 비트스트림을 생성하는 텍스처 맵 처리부를 포함하는 메쉬 데이터 인코딩 장치.
- 제 13 항에 있어서, 상기 변위 정보 처리부는상기 변위 정보의 좌표계를 로컬 좌표계로 변환하는 변위 정보 좌표계 변환부;상기 로컬 좌표계의 변위 정보에 포함된 노말 성분, 탄젠셜 성분, 또는 바이-탄젠셜 성분 중 적어도 하나의 성분에 대해 제1 포맷, 제2 포맷, 또는 제3 포맷 중 하나의 포맷을 기반으로 패킹을 수행하는 변위 정보 패킹부; 및상기 패킹된 변위 정보를 인코딩하는 변위 정보 인코딩부를 포함하는 메쉬 데이터 인코딩 장치.
- 제 13 항에 있어서, 상기 변위 정보 패킹부는상기 변위 정보의 바이-탄젠셜 성분에 대해 선택적으로 패킹을 수행하며,상기 시그널링 정보는 상기 바이-탄젠셜 성분의 패킹이 스킵되었는지 여부를 식별하기 위한 정보를 포함하는 메쉬 데이터 인코딩 장치.
- 제 13 항에 있어서,상기 시그널링 정보는 상기 변위 정보에 적용된 포맷을 식별하기 위한 정보를 포함하는 메쉬 데이터 인코딩 장치.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24832424.6A EP4734518A1 (en) | 2023-06-26 | 2024-06-26 | Mesh data transmission device, mesh data transmission method, mesh data reception device and mesh data reception method |
| CN202480042395.3A CN121420558A (zh) | 2023-06-26 | 2024-06-26 | 网格数据发送设备、网格数据发送方法、网格数据接收设备和网格数据接收方法 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR10-2023-0082062 | 2023-06-26 | ||
| KR20230082062 | 2023-06-26 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025005646A1 true WO2025005646A1 (ko) | 2025-01-02 |
Family
ID=93939215
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2024/008873 Ceased WO2025005646A1 (ko) | 2023-06-26 | 2024-06-26 | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4734518A1 (ko) |
| CN (1) | CN121420558A (ko) |
| WO (1) | WO2025005646A1 (ko) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210112240A1 (en) * | 2018-02-23 | 2021-04-15 | Nokia Technologies Oy | Encoding and decoding of volumetric video |
| US20210287431A1 (en) * | 2020-03-15 | 2021-09-16 | Intel Corporation | Apparatus and method for displaced mesh compression |
| KR102373833B1 (ko) * | 2020-01-09 | 2022-03-14 | 엘지전자 주식회사 | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 |
| US20220108483A1 (en) * | 2020-10-06 | 2022-04-07 | Sony Group Corporation | Video based mesh compression |
| KR20220126225A (ko) * | 2021-03-08 | 2022-09-15 | 현대자동차주식회사 | 포인트 클라우드 압축을 이용하는 메시 압축 방법 및 장치 |
-
2024
- 2024-06-26 EP EP24832424.6A patent/EP4734518A1/en active Pending
- 2024-06-26 WO PCT/KR2024/008873 patent/WO2025005646A1/ko not_active Ceased
- 2024-06-26 CN CN202480042395.3A patent/CN121420558A/zh active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210112240A1 (en) * | 2018-02-23 | 2021-04-15 | Nokia Technologies Oy | Encoding and decoding of volumetric video |
| KR102373833B1 (ko) * | 2020-01-09 | 2022-03-14 | 엘지전자 주식회사 | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 |
| US20210287431A1 (en) * | 2020-03-15 | 2021-09-16 | Intel Corporation | Apparatus and method for displaced mesh compression |
| US20220108483A1 (en) * | 2020-10-06 | 2022-04-07 | Sony Group Corporation | Video based mesh compression |
| KR20220126225A (ko) * | 2021-03-08 | 2022-09-15 | 현대자동차주식회사 | 포인트 클라우드 압축을 이용하는 메시 압축 방법 및 장치 |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4734518A1 (en) | 2026-04-29 |
| CN121420558A (zh) | 2026-01-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2024063544A1 (ko) | 3d 데이터 송신 장치, 3d 데이터 송신 방법, 3d 데이터 수신 장치 및 3d 데이터 수신 방법 | |
| WO2020190075A1 (ko) | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 | |
| WO2020190114A1 (ko) | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 | |
| WO2020189895A1 (ko) | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 | |
| WO2024049197A1 (ko) | 3d 데이터 송신 장치, 3d 데이터 송신 방법, 3d 데이터 수신 장치 및 3d 데이터 수신 방법 | |
| WO2021029511A1 (ko) | 포인트 클라우드 데이터 전송 장치, 포인트 클라우드 데이터 전송 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 | |
| WO2023172098A1 (ko) | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 | |
| WO2020190097A1 (ko) | 포인트 클라우드 데이터 수신 장치, 포인트 클라우드 데이터 수신 방법, 포인트 클라우드 데이터 처리 장치 및 포인트 클라우드 데이터 처리 방법 | |
| WO2023136653A1 (ko) | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 | |
| WO2024123039A1 (ko) | 3d 데이터 송신 장치, 3d 데이터 송신 방법, 3d 데이터 수신 장치 및 3d 데이터 수신 방법 | |
| WO2022050688A1 (ko) | 3d 데이터 송신 장치, 3d 데이터 송신 방법, 3d 데이터 수신 장치 및 3d 데이터 수신 방법 | |
| WO2025048542A1 (ko) | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 | |
| WO2024215096A1 (ko) | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 | |
| WO2024191257A1 (ko) | 3d 데이터 송신 장치, 3d 데이터 송신 방법, 3d 데이터 수신 장치 및 3d 데이터 수신 방법 | |
| WO2022098140A1 (ko) | 포인트 클라우드 데이터 전송 방법, 포인트 클라우드 데이터 전송 장치, 포인트 클라우드 데이터 수신 방법 및 포인트 클라우드 데이터 수신 장치 | |
| WO2024186127A1 (ko) | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 | |
| WO2024191192A1 (ko) | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 | |
| WO2024185940A1 (ko) | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 | |
| WO2025005646A1 (ko) | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 | |
| WO2025230333A1 (ko) | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 | |
| WO2025263884A1 (ko) | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 | |
| WO2025193057A1 (ko) | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 | |
| WO2025071369A1 (ko) | 메쉬 데이터 송신 장치, 메쉬 데이터 송신 방법, 메쉬 데이터 수신 장치 및 메쉬 데이터 수신 방법 | |
| WO2021201386A1 (ko) | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 | |
| WO2024205193A2 (ko) | 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24832424 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2024832424 Country of ref document: EP Effective date: 20260126 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024832424 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2024832424 Country of ref document: EP Effective date: 20260126 |
|
| ENP | Entry into the national phase |
Ref document number: 2024832424 Country of ref document: EP Effective date: 20260126 |
